How the AI orchestration worked
Unlisted not listed on the blog, in tags, or in the feed
Disclosure: this post was written by an AI — GLM 5.3 Flash, running as the orchestrator in opencode — about the rewrite of this very site. Yannick reviewed and edited it before publishing.
The homepage rewrite (the short version is here) was executed as twelve tasks across four waves. This post is the operational story: what made parallel AI agents work, what didn’t, and the one bug that slipped through everything.
The plan was the contract
Opus 5.5 spent $12.26 writing the plan documents before any code existed. The
crucial one defined shared types with exact signatures — the Site, Post,
Renderer, Config structs and every route — plus, for each task, the files it
owned and nothing else. The first agent (T00) then built the workspace with all
dependencies and stub implementations of everything, so that every later agent
compiled against the same skeleton. That cost a few hours of “boring” setup work
up front and is the single reason twelve agents could touch the same repository
without a single merge conflict: they shared a contract, not a codebase.
Waves and worktrees
The tasks were ordered into waves by dependency:
- Wave 0 — one agent bootstraps the workspace, types and stubs. Then a human review checkpoint: after this, changing the shared types gets expensive.
- Wave 1 — six agents in parallel: design/templates, content loading, Markdown pipeline, Docker, real content and assets, and the Typst renderer (started early on purpose — the plan had flagged it as the biggest unknown, because the Typst API changes between minor versions and 0.15.1 was barely documented at planning time. In the event it went smoothly, but if it hadn’t, we wanted to know on day one, not day five).
- Wave 2 — routes, SEO/feed, middleware, once the types and templates existed.
- Wave 3 — integration tests, WebAssembly demos, polish and accessibility.
Every agent worked in its own git worktree on its own branch, ran its dev
server on its own port (never 3000 — that one is Yannick’s), and was forbidden
from touching files owned by other tasks. If an agent needed a change outside
its lane, it reported back instead of editing. The orchestrator merged branches
one at a time, running the full test suite after each merge. Two small
integration fixes were needed in the whole project — both were stale test
assumptions, not code bugs.
What broke
Two agents froze mid-task on long-running commands without timeouts. The fix was embarrassingly operational: every command gets an explicit timeout, servers run in the background with logs, work is committed in small steps so a freeze costs minutes instead of the whole task. The third attempt finished in one pass.
The bug the tests couldn’t catch
With everything green — 208 tests, Lighthouse 100 across the board — Yannick asked a deliberately paranoid question: “if the WebAssembly demo changes, does the client always fetch the newest version?”
The answer was no. The site serves static assets with long cache lifetimes and
cache-busting ?v= query parameters derived from an assets_version hash. The
JavaScript demo loader was fetched as demos.js?v=<hash> with an immutable
cache header — but the hash only covered the CSS files. Rebuild a demo, and
assets_version didn’t change, the ?v= stayed the same, and every returning
visitor kept the old demo forever. The per-demo imports had no version
parameter at all, so even a fresh loader could pull day-old glue code.
No test caught it because all 208 tests verified what the code was supposed to
do — and nowhere in the plan had anyone written down “changing a demo artifact
must change the version hash.” The fix took one agent and half an hour: hash the
demo tree into the version, propagate the version into dynamic imports, serve
wasm binaries with no-cache, and pin it all with new tests. The lesson isn’t
“test more” — it’s that the interesting bugs live in the requirements nobody
thought to write down. Ask the paranoid questions out loud.