Better outcomes,
every round

Memory, workflows and review for Claude Code and Codex, built into the repository.

Your goal Long HorizonContinues across rounds
01 / Recall

Start with project memory

Illustrative workflow
·

Candidate
awaiting checks

Human
approval

Release stays
your decision
Illustrative system · choose skills for the task

The same agent, with a repository that remembers and checks

  1. 01 / The problem

    Without Harness Firmware, every session starts from scratch

    • Lessons vanish when the chat ends
    • Long tasks lose progress to compaction
    • "Done" is the agent grading its own work
    • A model reviewing itself shares its own blind spots
  2. 02 / Project memory

    Project knowledge lives in git, reviewable like code

    • Memory stuck on one machine
    • Travels with git clone to every machine and sandbox
    • Anything gets saved
    • Each save passes three tests; task status never gets in
    • Old facts contradict new ones
    • A changed fact gets one dated "retired" line
  3. 03 / Recall

    Before unfamiliar work, the agent reads the project's own pitfalls

    per pitfalls.md (2026-03-14): reset the test database first

    A fact that changes an action is cited with its file and date, so you can correct a stale one in one reply.

    Skills cost one index line until a task calls the full playbook.

  4. 04 / Plan

    Big tasks run against a written contract

    • Checks: Numbered acceptance checks, never weakened to make a round pass
    • Steps: Each sized for one fresh context
    • Memory: Progress and dead ends saved to a file that survives compaction
    • Stalls: A step that fails twice must change approach
  5. 05 / Checked apart

    In audited rounds, the builder never grades its own work

    1. The auditor's brief is written before the builder exists
    2. A content-hash baseline shows exactly what changed
    3. The auditor runs the checks itself
    4. Only complete, clean and aligned counts as progress
  6. 06 / Review and your call

    Reviews check every finding against the code before you see it

    • On request, the other vendor's model reviews the diff
    • Dismissed findings stay listed, so you can overrule them
    • Running a skill never authorizes a merge, release or deploy
  7. 07 / Refine

    Rules change on evidence, never on a one-off slip

    • Three checks must tie a failure to a rule; zero changes is valid
    • Only your own words count as a preference
    • An audit flags never-read notes and never-used skills to prune
  8. 08 / Next project

    With your approval, one project's fix improves the template

    /sync-starter sends back only fixes that can be made generic, and you pick what moves; keeping a lesson local stays the default.

  9. 09 / Next session

    The next session starts from what the project has learned

    • Decisions and pitfalls, read before unfamiliar work
    • In audited rounds, "done" means an independent check passed
    • Plain files, MIT licensed, read by Claude Code and Codex
Illustrative sequence · plain files in your repository

Coding agents start every session from zero

This one starts where you left off

Only what the task needs

Token-efficient

Skill names sit in a small index. The full instructions, project notes and checks load when a task calls for them, so the context window stays free for your code.

See how memory loads

Some of it runs by itself

You never type these. They run in every session.

always on

caveman

Replies drop filler and pleasantries and keep every technical detail. On from the first message.

automatic

recall

Before unfamiliar work, the agent reads the project notes that apply and skips the rest.

auto-saved

pitfalls.md

A mistake that cost a retry gets written down, so the next session avoids it.

Planned

Question the problem, try more than one answer, and tune it before it ships.

/dare

Takes the problem apart from first principles in four fresh passes: break it down, test each assumption, rebuild, check the result against reality.

/arena

Runs several attempts in parallel, has a blind judge pick the strongest as the base, and folds in the best parts of the rest.

/lab

Builds a live prototype with controls, so you tune motion and layout by feel before anything is final.

Audited

A second opinion on every change, from a reviewer with no stake in the work.

Fresh evidence / illustrativeInspect → recheck
Work under reviewChecked result
01

Inspect

A fresh reviewer gets the change, its scope and the checks it must pass.

02

Repair

Each finding is checked against the actual code before anything changes. Confirmed gaps get fixed, then reviewed again.

03

Decide

You get the evidence. Nothing merges, publishes or ships without your approval.

  • /codex-review

    A different model family, OpenAI's Codex, reviews the diff.

  • /impartial-review

    Fresh agents that never saw the work review it.

  • /handoff-audit

    Writes an audit brief with exact scope and pass/fail checks that another session can run.

Remembered

Decisions, pitfalls and commands live in plain files next to your code. They go wherever the repository goes.

Explore memory

/long-horizon

Your session holds the planA fresh agent builds each roundA separate auditor checks the real filesOnly work that passes moves forward

Building and checking happen in fresh contexts, so your main conversation stays small and quality holds on tasks too big for one context window.

/long-horizon · Claude Code and Codex

/long-horizon-workflows · Claude Code · the ultra version

Runs the same rounds through Claude Code's built-in workflows, with a fresh builder, an inspector and a panel of judges each round, written verdicts, and a run journal.

Production-ready

Long tasks run in audited rounds. Finished work gets polished and measured before it ships.

/showpiece

Pushes a page, deck or document past the generic look toward work you'd put in a portfolio.

/wow-loop

Reviews and repairs one piece against the goal, round after round, and keeps only changes that prove better.

/perf-loop

Measures speed against a baseline, changes one thing, and measures again with an independent check.

Handing work to another agent: /enhance-prompt writes the prompt it needs, and /handoff-audit writes the checks it must pass.

Self-improving

When a workflow stumbles, /refine finds the cause and puts the fix into the skill itself. Preferences you state become rules the same way.

Browse the skills

Start
Optimized

Use the template for a new project, or add the skills to one you already have. Built for Claude Code and Codex.

Questions

What am I installing? +

A repository template with agent instructions, reusable skills, and project references. Read and change the files like the rest of your code.

Does it remember automatically? +

Memory lives in project files. Agents can read those files and record approved, useful lessons. What gets retained depends on your workflow and instructions.

Does every improvement apply to every project?+

Keep project-specific lessons with the project. Review broadly useful improvements before sharing them through the template.

Where does the human stay involved? +

You set the goal and decide what can run. Reviews and checks provide evidence; your authorization controls publishing, merging, and other external actions.

Can I use it with an existing project? +

The skills-only plugin works with Claude Code in an existing repository. For a new project or Codex support, use the full template. See the setup instructions .

How is this different from a CLAUDE.md file?+

A CLAUDE.md holds rules. Harness Firmware adds the rest around it: project memory the agent reads when it applies, skills for planning, review and long tasks, and checks that run on every change.

Does it add token overhead?+

Very little up front. Each session loads the core rules and a short index of skill names. Full skills and project notes load only when a task calls for them.