AI Nations
Typed skill catalogues that AI agents can actually compose
Agent skills are usually prose in a prompt. Here they are typed edges: every entry declares what it consumes and what it produces, so a solution is a path through a dependency graph rather than a hopeful paragraph. The catalogue is derived from recorded work, never hand-authored, and it is graded by verifiers and certification graders that do not belong to me.
- 100%scored by an external certification grader on real, time-boxed infrastructure
- 19/19mutation targets killed ยท 0 survived
- ~70recorded episodes, each with its gates, verdicts and cost
- $1.7428cost of one episode โ the real figure, not its cheapest run
The mechanism
Swipe the diagram sideways to read it.
list reveals a resource; it does not create one. Read-only entries
are provably unable to satisfy a prerequisite, which is why the dependency graph stays
acyclic and why a solver cannot “satisfy” a step by looking at it.Swipe the diagram sideways to read it.
Two episodes
Unedited screen recordings of the system doing the thing described above. Each one is a full run: the gates, the verdicts and the cost are all on screen.
Where this is going
Everything above this line is built and measured. This part is not, and the distinction matters more to me than the idea does.
The hypothesis. A typed catalogue is only worth building once if it can be reused.
The entries above carry input_schema_id and output_schema_id precisely so that two
catalogues built by two people in two domains can be checked for compatibility
mechanically rather than by reading them. If that holds, a catalogue stops being a private
convenience and becomes something that can be exchanged, composed with someone else's, and
recombined into a capability neither author built โ the same way a command sequence already
composes into a solution nobody wrote down in advance.
Why the biological vocabulary. I have been calling these AI chromosomes โ domain catalogues of reusable genes โ since long before the current wave, and the reason is structural rather than decorative. Nature solved "how do many small agents cooperate at scale" twice: once with a single-chromosome loop, and once with a nucleus, a power plant and combinatorial gene engines on top of it. Both patterns have direct analogues in agent architecture, and the interesting question is not which is right but which wins which job.
The nearest thing I have to evidence is unglamorous and already running: the same command
genome expresses differently on different hosts. One manifest installs kubectl
shortcuts on a machine that has kubectl and omits the Firebase ones because that CLI is
absent; another, for an ephemeral cloud VM, is close to its inverse. Every exclusion
carries its reason, and a survey error that wrongly dropped kubectl is corrected in the
file and dated. Environment-conditional expression of a shared genome is not a metaphor
there โ it is what the manifest is.
What AI Nations would be. A place where many practitioners' catalogues can be published, type-checked against each other, recombined, and ranked by what they actually solved โ with the grading kept external, as it is above. The ambition is that the recombining gets fast enough, and honest enough, to be pointed at problems worth the compute.
The status, plainly. The two engine phases that would carry this โ the single-chromosome
loop and the multi-chromosome runtime above it โ are written up in full and are still marked
Stub v0: scope sketched, tasks not decomposed, nothing running. The catalogue, the solver,
the verifiers, the episodes and the external score are real and you can watch them. The
nation is not. I would rather say that here than have you find it out later.
This is a design hypothesis under test, not a manifesto.