About Harness Firmware: the 77-second film
This film needs WebGL2, which this browser does not offer. It shows a coding agent that forgets a correction, then a repository that remembers it, trims what the agent reads and writes, and has a second model check the work.
Field log: seven sessions that built a production site
Seven sessions, three days, one production site
Each session worked its own lane. The tunnel and the hero ran
/long-horizon,/init-projectand/wow-loop; the masthead ran/perf-loopand/refine; the About page took 80 rounds of/long-horizon-workflows. An orchestrator ran/codex-reviewbefore every merge, then a performance pass (/perf-loop,/codex-fullreview) and a study engine (/why,/long-horizon-swarm) came last.Each session hands the plan on; one lands it all
Each session wrote the brief for the next, and every handoff was checked against GitHub;
/codex-reviewcaught one bug on the way. One pull request of 575 files landed the work, and six PRs landed in all.A lesson saved in one session protects the next
At 16:52 the performance pass ran recall, which read the pitfalls, architecture and commands notes. At 22:40 a forced removal wiped the shared
node_modules. The session restored it and saved two pitfalls: never force-remove a worktree whosenode_modulesis a junction, and Git Bash rewrites/aboutinto a path, so setMSYS_NO_PATHCONV=1. At 22:51 the study engine's recall readpitfalls.md, and at 23:49 it met the same trap and avoided it, 69 minutes after the wipe.Another model catches what the builder missed
Six same-model audits in
/long-horizonpassed the page./codex-review, a model from another vendor, found a door-hole layering bug; it was fixed, then reviewed again. On one page,/codex-reviewand/astra-reviewfound 10 real bugs, all fixed.Judged blind, measured before and after
In
/wow-loopa blind judge compared new against old, and the new version won 23 of 24 rounds./perf-loopmade one measured change per round: the masthead went from 6.01 to 0.06 ms per frame, the home page from 4.10 to 2.43 MB per full scroll, and the About page from 50.8 to 64.9 fps on an integrated GPU. In/long-horizon-workflowsa judge flagged a round that had tuned its code to pass its own check.The work showed what had gone stale
Project rules and notes had gone stale, so
CLAUDE.mdandAGENTS.mdwere reviewed in the firmware. A review skill named a retired model, so/refineupdated the review skills. A reviewer with write access deleted the build, so reviewers are now read-only. Project setup had never been run, so/init-projectnow runs before release work.Rules load every turn; detail loads when it is needed
Loading every note and playbook on every turn would cost about 87,000 tokens. The firmware loads about 3,400:
CLAUDE.mdand one index line for each of its 36 skills. For a build bug, recall adds the pitfalls and architecture notes (4,100 tokens), then one playbook when the task calls for it (2,160), about 9,700 in all.What one project learns, the firmware keeps
Every new repository starts from the fixes, and
/sync-starterbrings them into existing ones.
MIT licensed, for Claude Code and Codex. Already have a repository? Copy the skills you want from the firmware repository ↗.