Simple ai harness blueprint
Reusable AI agent collaboration scaffold system (AGENTS.md / CLAUDE.md / .agents/). Three additive sizes — S/M/L. Works with Claude Code, Codex Desktop, Cursor, Windsurf.
npx -y skills add GuiomeB/simple-ai-harness-blueprintAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Install, audit, or extend an AI-collaboration blueprint on a code repository — the AGENTS.md / CLAUDE.md / .agents/ structure that lets AI agents (Cursor, Claude Code, Codex Desktop/CLI, Windsurf) collaborate without drift. Three additive sizes (S/M/L) plus an opt-in L+ autonomy profile. Trigger when bootstrapping a new repo's AI scaffold, promoting an existing harness to the next size, or auditing one for sprawl.
SKILL.md
19.3 KB, ~4.6k tokens by cl100k_base, as published. Nobody here has run it
Simple AI Harness Blueprint
A scaffold for repos where AI agents and humans collaborate. Three sizes (S/M/L), strictly additive — filenames never change between levels — plus an opt-in L+ autonomy profile (ADR-gated; not a fourth size).
When NOT to use
Skip for product code, PR reviews, IDE configuration, or generic project scaffolding.
Doctrine — the 5 Karpathy rules + M0
Every AGENTS.md generated opens with these. They override convenience and other rules in conflict. The 5 rules are posture; M0 is the verification mechanism.
Version stamp. Every generated AGENTS.md carries, right under its title, the line:
> Doctrine: v5 (5 règles Karpathy + M0) — appliquée <YYYY-MM-DD>
Fill the date with the day the doctrine was applied to that repo. The stamp is what makes doctrine drift greppable across a fleet (scripts/audit_fleet.py reads it). A repo deliberately left on an older doctrine states it the same way (e.g. > Doctrine: v4 figée — …) — a known lag is fine, an invisible one is not.
- Ask, don't assume. Ambiguity → ask before coding. Only in an explicitly activated autonomous mode (L+): pick the most reasonable interpretation, proceed, record the assumption.
- Simplest solution for simple problems, stronger for hard ones. No over-engineering, no preventive abstractions.
- Don't touch unrelated code — but surface what you find. Diff = scope of the ticket; raise smells as a separate issue.
- Flag uncertainty explicitly. Unsure → see rule 1; a small low-risk experiment + results beats false confidence.
- Suggest better ways. Propose lasting improvements over tactical fixes.
M0 — Verification. Before acting, state a verifiable success criterion and loop until it holds: trigger · stop criterion · validation · budget · stop/no-progress. The validation matrix, DoD, /tdd-loop, and (at L+) /loop all inherit from M0.
Pick the size
Depths of harness, not codebase sizes. Match signals:
| Signal | S | M | L |
|---|---|---|---|
| AI agents in regular use | 1+ | 1–3 | 2+ rotating |
| Human contributors | Solo | Solo or 2–3 | 2+ |
| Deployment criticality | Local / hobby / early stage | Public / users | Production / regulated |
| Project age with active agent edits | < 3 months | 3–12 months | 12+ months |
| Friction "the agent forgot rule X" | Rare or absent | Recurring | Painful or routine |
| Critical zones to protect | 0–1 | 2–5 | 5+ |
Existing AGENTS.md state | None, or close to the ~100–150 line S baseline | Growing past 150 lines | Past 250 lines, hard to navigate |
Two or more signals at a level → adopt that level. Never jump straight to L on a fresh repo.
Note on the AGENTS.md signal: a freshly-generated S AGENTS.md is itself ~100–150 lines (5 Karpathy rules + M0 + critical zones + commands + rail rule + load order). When reading this signal, look at project-specific accretion on top of the baseline, not the raw line count.
L+ — the autonomy profile (opt-in, not a fourth size)
L+ is not the next rung on this matrix. It is a specialised profile layered on an L harness: the same L, plus the wiring to run bounded loops unattended. The size matrix never recommends it — promote only when all five are true, and record the decision in an ADR:
- repeated loops with a verifiable goal,
- real unattended execution (headless / scheduled),
- parallel subagents or worktrees in use,
- a CI / headless runner exists,
- budgets and stop-conditions are genuinely needed.
None of these → stay at L. Adding L+ to a repo that doesn't need it re-creates the sprawl this blueprint fights. A loop without its three hard brakes (budget · no-progress detection · kill-switch) is forbidden.
Tree per size
Every level is strictly additive: filenames never change between sizes, you only add.
S — minimum with all 3 mechanisms present
your-repo/
├── README.md
├── AGENTS.md 5 Karpathy + M0 verification + M1 (load order) + critical zones + commands + M3 rail rule
├── CLAUDE.md optional adapter — include only if Claude is a used agent
├── WORKFLOW.md process + when to invoke /learn (refers to AGENTS.md for rail discipline)
└── .agents/
└── workflows/
└── learning-loop.md M2 — /learn per-event capture
M (additive over S) — routing + status + formalized PR flow
+ STATUS_APP.md living log: current stage, recent decisions, known debt
+ .agents/
│ ├── ROUTER.md task family → context capsule matrix
│ ├── context/
│ │ └── <critical-domain>.md 1–3 capsules per project; 80–150 lines each
│ └── workflows/
│ ├── retro.md post-release retrospective workflow
│ └── tdd-loop.md optional but typical at M
+ docs/retro/ output of /retro
+ .github/
└── pull_request_template.md M3 formalized: risk rail declaration in PR body
L (additive over M) — full system, machine-enforced
+ .agents/
│ ├── patterns/
│ │ ├── INDEX.md registry: pattern ↔ pivot files in the codebase
│ │ └── <action-pattern>.md short copyable procedures (≤ 120 lines)
│ ├── rules/
│ │ └── <tech-convention>.md narrow technical rules; may carry `globs:` to auto-load
│ └── skills/
│ └── <project-skill>/SKILL.md project-specific skills / Claude skills where applicable
+ .claude/ Claude-Code runtime adapter (concept lives in AGENTS.md)
│ └── agents/
│ └── <reviewer>.md read-only subagent, tools allowlist, model: sonnet
+ docs/
│ ├── adr/ Architecture Decision Records
│ └── learn/ per-event /learn output (auto-created on first run)
+ scripts/
│ ├── validate_agent_context.* meta-CI: the harness validates its own coherence
│ └── check_pr_rail_consistency.* rail-guard: fails `green` PRs that touch CODEOWNERS paths
+ .github/
│ ├── CODEOWNERS source of truth for "critical paths"
│ └── workflows/
│ ├── agent-context.yml non-blocking signal on .agents/** changes
│ └── pr-rail-guard.yml blocking gate on PR risk rail consistency
+ _local/CI_RULEBOOK.md optional, gitignored: rationale of CI/agent choices
L+ (additive over L) — the opt-in autonomy profile (ADR-gated, not a size)
+ .agents/
│ └── workflows/
│ └── loop.md /loop unit-of-work: 3 hard brakes + reviewer verification
+ .claude/
│ ├── settings.json hooks: post-tool formatter, push-to-main defer, denied log
│ └── hooks/
│ └── gate_git_push.sh defers `git push … main` for human approval (fail-open)
+ .github/
│ └── workflows/
│ └── headless-loop.yml headless runner — manual-dispatch until brakes are exercised
+ docs/adr/
└── ADR-0002-*.md records the decision to run autonomously (mandatory)
Promotion to L+ is gated (see §Pick the size → L+). .claude/logs/ is gitignored. A loop without its three hard brakes is forbidden.
File size budgets
| File | Soft target | Hard ceiling |
|---|---|---|
AGENTS.md | 100–150 | 300 |
CLAUDE.md | 40–60 | 100 |
WORKFLOW.md | 150–200 | 300 |
STATUS_APP.md | unbounded (log) | n/a |
.agents/ROUTER.md | 40–60 | 100 |
.agents/context/*.md | 80–150 | 200 |
.agents/patterns/*.md | 40–120 | 200 |
.agents/workflows/*.md | 80–200 | 300 |
.agents/skills/*/SKILL.md | ≤ 200 | 500 |
.claude/agents/*.md | 30–60 | 120 |
.claude/hooks/*.sh (L+) | ≤ 40 | 80 |
Session-init budget: keep total auto-loaded context ≤ 6 500 lines across all files an agent reads before its first action. Trade per-doc ceilings against this total — a leaner ROUTER.md leaves room for a fatter WORKFLOW.md, and vice versa.
MCP / tool-schema budget: every connected MCP server spends context tokens on every turn (≈ 10–20k tokens for 50 tools without lazy loading). Cap a serious setup at ≤ 5 servers; prefer a code-graph/memory server, a Git server, a filesystem server, a web-search server, and a docs server. Fewer servers beats lazy-loading. Subagents with a narrow tools: allowlist keep the main loop's schema cost down.
The mechanisms (M0–M3) — templates only
M0. Verification (lives in AGENTS.md §Doctrine, right after the 5 rules)
The mechanism every task inherits — the executable backbone of rule 4. State a verifiable success criterion, then loop until it holds.
## M0 — Verification (the mechanism behind every task)
Before acting, state a verifiable success criterion, then loop until it holds:
- Trigger — what starts the work.
- Stop criterion — the verifiable signal it's done (red test → green, lint clean, smoke passes).
- Validation — the minimum commands from §Minimal validation matrix.
- Budget — a ceiling on time / iterations / tokens.
- Stop / no-progress — if not converging, stop and surface the blocker (rule 1), don't loop blindly.
The validation matrix, the DoD, /tdd-loop, and (at L+ only) /loop all inherit from M0.
M1. Init: ordered context load (lives in CLAUDE.md or the agent's adapter file)
Two progressive forms — the order grows one hop once the ROUTER exists.
At S (no ROUTER yet):
Before generating code for any new request, load context in this order:
1. AGENTS.md (project contract — implicit, never skip)
2. files directly touched by the request
3. additional documentation only if the task obviously requires it
Never load large documents "just in case".
At M and L (ROUTER present):
Before generating code for any new request, load context in this order:
1. AGENTS.md (project contract — implicit, never skip)
2. .agents/ROUTER.md (identify task family, pick minimum context)
3. the capsule / pattern the router points to (if any)
4. files directly touched by the request
5. additional documentation only if the router points to it
Never load large documents "just in case".
Always restate the active form in AGENTS.md §Role so the rule survives careless edits to the adapter.
M2. Learning: per-event capture (lives in .agents/workflows/learning-loop.md)
# /learn workflow
Invocation: /learn <family> <slug>
Families: release | candidate | incident | friction | refactor
Output: docs/learn/LEARN_<family>_<slug>_<YYYY-MM-DD>.md (≤ 40 lines):
- What helped (2–4 bullets)
- What slowed us down (2–4 bullets)
- ONE action retained (imperative, < 20 words)
- "Lands in: <target artefact>"
- Diffusion: files actually changed
Then update the target artefact and run `validate:agent-context`.
Hard constraint: ONE action per /learn.
M3. Risk rail enforcement (3 progressive forms)
The rail concept lives at every level — only the enforcement strength grows.
At S — agent self-declaration in AGENTS.md (no PR flow required):
## Risk rail (declare after every task)
After completing any task or before pushing/committing changes, declare a rail
in your message:
- Rail: green | amber | red
green = small/local; no critical zone touched; safe to merge fast
amber = behavioural or transverse; the user should scan the diff
red = critical path or production risk; the user must review
The rail is informational at this stage (no CI enforcement). Its value is
forcing the agent to self-assess sensitivity, and the user to see it.
At M — formalized PR template at .github/pull_request_template.md:
## Risk rail
- Rail: green | amber | red
green = small/local; no CODEOWNERS path; no runtime/auth/release touch
amber = behavioural or transverse; short human review or owner approval
red = critical path or production risk; full review mandatory
Touching a .github/CODEOWNERS path → minimum amber.
P1/P2 finding during review → automatic red until resolved.
At L — machine enforcement:
- CI job
pr-rail-guard: fail PRs declaredgreenthat touch any path listed in.github/CODEOWNERS - Meta-validator
scripts/validate_agent_context.*: check everynpm run Xcited inAGENTS.mdexists inpackage.json, every internal link in.agents/**resolves, andpatterns/INDEX.mdcovers every pattern file
Codex Desktop compatibility
This repository is both a Claude skill and a Codex Desktop skill. Codex reads the same SKILL.md frontmatter/body; agents/openai.yaml only adds optional UI metadata. Do not fork the instructions by agent. Keep AGENTS.md canonical and use thin adapters (CLAUDE.md, GEMINI.md, etc.) only when a tool requires a specific entry file.
Workflow for invoking this skill
- Inventory. List existing blueprint files in the target repo.
- Infer and propose — don't interrogate.
- Size: match the matrix above to repo signals —
AGENTS.mdline count, presence of.agents/, multi-agent traces,.github/, codebase age. - Agent adapters: scan for markers —
CLAUDE.md,.cursorrules,.windsurfrules,GEMINI.md, etc. Include each adapter only if its marker is found, the user names that agent explicitly, or the user explicitly says "Claude" / "Cursor" / etc. is in use.AGENTS.mdis always included. - PR workflow: detect
.github/,.gitlab/, branch-protection rulesets. Add.github/pull_request_template.mdonly at M+ and if a GitHub-compatible PR flow is detected. - Propose in one sentence with reasoning, e.g. "Based on no existing
AGENTS.md, a.cursorrulesfile, and a.github/folder, I'd target M, withCLAUDE.md(you mentioned Claude) and.github/pull_request_template.md. Confirm or adjust?" - Ask explicit questions only when inference is genuinely ambiguous (multiple plausible sizes, conflicting agent markers, or the user contradicts the inference).
- If existing artefacts exceed the requested level, warn about downgrade risk before removing anything.
- Size: match the matrix above to repo signals —
- Generate the delta. For each missing file, propose the matching template from
references/templates/<size>/. Confirm before writing. Never overwrite silently. Follow placeholder discipline (see below): do not invent commands, files, or critical zones the target repo doesn't have. In the generatedAGENTS.md, fill the doctrine version stamp's<YYYY-MM-DD>with the current date — it is the one placeholder you always resolve at generation time.- The M templates include one worked example (a
data-mutations.mdcapsule and a matching row inROUTER.md) to model after. Replace or delete the example if mutations aren't a critical domain in the target project, and route the user's own capsules instead.
- The M templates include one worked example (a
- Validate. Print the created tree, run any present
validate:agent-context, ask the user to skimAGENTS.mdfirst. - Plant the next promotion criterion AND report remaining placeholders. Add a one-line "promote to next size when X" at the bottom of
AGENTS.md. Then list every<placeholder>left in the generated files (commands, critical zones, project description, etc.) so the user knows exactly what to fill in before the harness becomes operational. An unfilled placeholder is the correct state — a filled one with invented content is a lie.
Adopting on an existing repo
If the target repo already has its own agent doctrine, never overwrite:
- Don't replace an existing
AGENTS.md. Propose an additive merge: keep the project's accreted rules, insert universal sections (5 Karpathy rules + M0, load order, rail discipline) only where they're missing. - Don't create a parallel
.agents/if the repo uses a different layout (.cursor/rules/,.windsurf/, custom paths). Adapt: place equivalent content there and note the divergence at the top ofAGENTS.md. - Look for what's missing, not for what doesn't match this template. The blueprint's value on a mature repo is usually three things: navigation (ROUTER at M), drift check (
validate_agent_context.*at L), per-event learning loop (/learn). Add only those.
A mature repo with a working memory should be enhanced, not rescaffolded.
Placeholder discipline
Templates ship with <...> placeholders (e.g. <run dev>, <typecheck>, <critical-zone>, <path/to/file.ts>). They are intentional — they mark what the project hasn't decided or doesn't have yet.
Hard rule: never replace a placeholder with an inferred value. Replace only when one of these is true:
- The actual file or tool already exists in the target repo, and you have read or listed it.
- The user has explicitly named the value (e.g. "I use
ruffandpytest" → fill<lint>and<test>). - The user explicitly asked you to scaffold product code too (rare; outside this skill's default boundary).
If you don't know, leave the placeholder. Empty placeholders tell whoever opens the file what's not yet decided. A filled placeholder with invented content is a lie that compounds over time and rots the harness quickly.
A user prompt mentioning a language or framework ("Python CLI", "Next.js side-project", "React + Node API") is a hint about direction, not a guarantee that tooling exists. The repo is the source of truth; the user's words are clues for what to ask about, not licenses to assume.
When the bootstrap is done, the report from step 5 of the workflow must enumerate every remaining <placeholder> so the user can fill them deliberately.
Health metrics
| Symptom | Cause | First fix |
|---|---|---|
| Agents ignore the router | Init rule buried | Restate at top of AGENTS.md §Role |
| Same friction reported 3+ times | /learn never ran | Run one /learn friction retroactively |
AGENTS.md > 300 lines | Domain doctrine in constitution | Extract to .agents/context/<domain>.md |
.agents/ files contradict | No validator | Add validate_agent_context.*; schedule a /retro |
| Token-heavy loads for trivial tasks | ROUTER too permissive | Tighten rules; mark files "do not load by default" |
| Capsule untouched 3+ months | Doctrine stable or zone dead | Inline into AGENTS.md or delete with an ADR |
| New agent ignores rules | Reads a different config file | Symlink or copy AGENTS.md to its expected name |
Two or more symptoms simultaneously → schedule rework.
References
references/templates/S/,M/,L/— copyable file templates per levelreferences/health-metrics.md— extended diagnostics
Boundaries
Operates only on the documentation surface agents read. Does not touch product code, run tests, configure IDEs, or decide architecture.
Gives 0 of the 12 instructions most memory context skills give in ~4.6k tokens
Counted across 674 of the 847 authors here whose files we hold, read 2026-08-06
- inform the user when setup is completein 21 of 674, across 6 files
- confirm the draft with the user before writingin 21 of 674, across 6 files
- update the agent skills block in place if it existsin 21 of 674, across 6 files
- present findings to the userin 20 of 674, across 5 files
- write the three docs files from seed templatesin 20 of 674, across 5 files
- ask the user about each decision one at a timein 19 of 674, across 4 files
- edit CLAUDE.md if it existsin 18 of 674, across 3 files
- explore current repo statein 18 of 674, across 3 files
- do not overwrite user edits to surrounding sectionsin 18 of 674, across 3 files
- back up the original file before overwritingin 16 of 674, across 8 files
- keep the memory index under 200 linesin 15 of 674
- Provide actionable steps and verificationin 13 of 674, across 2 files
Said here and by no other author read
- Add files additively across harness sizes
- State a verifiable success criterion before acting
- Adopt a harness size matching two or more signals
- Record the decision to run autonomously in an ADR
- Keep total auto-loaded context under 6500 lines
- Cap connected MCP servers at five
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.