AI Nations

Typed skill catalogues that AI agents can actually compose

Agent skills are usually prose in a prompt. Here they are typed edges: every entry declares what it consumes and what it produces, so a solution is a path through a dependency graph rather than a hopeful paragraph. The catalogue is derived from recorded work, never hand-authored, and it is graded by verifiers and certification graders that do not belong to me.

The mechanism

A command modelled as a typed edgeFour resource-type nodes โ€” nothing, network, instance and load balancer โ€” joined by three command edges. Each edge is labelled with the resource types it consumes and the resource type it produces, so a solution is a path through the graph.(none)networkinstancelbnetworks createinstances createforwarding-rules() โ†’ network(project, zone,network) โ†’ instance(instance) โ†’ lbinstances list(project) โ†’ () ยท produces nothing, so it can satisfy nothingA command is a typed edge, not a text snippet.Every entry declares what it consumes and what it produces, so composing asolution is pathfinding with parameters bound โ€” deterministic, and checkable.

Swipe the diagram sideways to read it.

A list reveals a resource; it does not create one. Read-only entries are provably unable to satisfy a prerequisite, which is why the dependency graph stays acyclic and why a solver cannot “satisfy” a step by looking at it.
How a catalogue earns its scoreA four-stage ring โ€” recorded session, derived catalogue, composed solution, verifier โ€” with an external certification grader drawn outside the ring, scoring the composed solution and feeding the result back into the catalogue.recordedsessionderivedcataloguecomposedsolutionverifierbaseline NO-GOexternal certification graderscored 100% ยท badge issuedThe score comes from outside.Real infrastructure, time-boxed to 25 minutes, no step-by-step instructions โ€”passable only by composing commands learned earlier. It does not negotiate.

Swipe the diagram sideways to read it.

The verifier is written to be red before the thing it verifies exists. A check that has never failed proves nothing, so the baseline NO-GO is run first and recorded.

Two episodes

Unedited screen recordings of the system doing the thing described above. Each one is a full run: the gates, the verdicts and the cost are all on screen.

ep16 โ€” the catalogue being built Watch for: Watch the extraction run twice. The second run adds zero entries; that is the idempotency check, and it is the reason the catalogue can be trusted as derived rather than curated.
ep20 โ€” the solver composing a Kubernetes solution Watch for: Watch the solver order the steps from the dependency graph rather than from the order the commands were learned in. The read-only entries never appear as prerequisites.

Where this is going

Everything above this line is built and measured. This part is not, and the distinction matters more to me than the idea does.

The hypothesis. A typed catalogue is only worth building once if it can be reused. The entries above carry input_schema_id and output_schema_id precisely so that two catalogues built by two people in two domains can be checked for compatibility mechanically rather than by reading them. If that holds, a catalogue stops being a private convenience and becomes something that can be exchanged, composed with someone else's, and recombined into a capability neither author built โ€” the same way a command sequence already composes into a solution nobody wrote down in advance.

Why the biological vocabulary. I have been calling these AI chromosomes โ€” domain catalogues of reusable genes โ€” since long before the current wave, and the reason is structural rather than decorative. Nature solved "how do many small agents cooperate at scale" twice: once with a single-chromosome loop, and once with a nucleus, a power plant and combinatorial gene engines on top of it. Both patterns have direct analogues in agent architecture, and the interesting question is not which is right but which wins which job.

The nearest thing I have to evidence is unglamorous and already running: the same command genome expresses differently on different hosts. One manifest installs kubectl shortcuts on a machine that has kubectl and omits the Firebase ones because that CLI is absent; another, for an ephemeral cloud VM, is close to its inverse. Every exclusion carries its reason, and a survey error that wrongly dropped kubectl is corrected in the file and dated. Environment-conditional expression of a shared genome is not a metaphor there โ€” it is what the manifest is.

What AI Nations would be. A place where many practitioners' catalogues can be published, type-checked against each other, recombined, and ranked by what they actually solved โ€” with the grading kept external, as it is above. The ambition is that the recombining gets fast enough, and honest enough, to be pointed at problems worth the compute.

The status, plainly. The two engine phases that would carry this โ€” the single-chromosome loop and the multi-chromosome runtime above it โ€” are written up in full and are still marked Stub v0: scope sketched, tasks not decomposed, nothing running. The catalogue, the solver, the verifiers, the episodes and the external score are real and you can watch them. The nation is not. I would rather say that here than have you find it out later.

This is a design hypothesis under test, not a manifesto.