agentsclimarketplace

Simple ai harness blueprint

Skill GuiomeB/simple-ai-harness-blueprint

Reusable AI agent collaboration scaffold system (AGENTS.md / CLAUDE.md / .agents/). Three additive sizes — S/M/L. Works with Claude Code, Codex Desktop, Cursor, Windsurf.

Install
npx -y skills add GuiomeB/simple-ai-harness-blueprint

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Install, audit, or extend an AI-collaboration blueprint on a code repository — the AGENTS.md / CLAUDE.md / .agents/ structure that lets AI agents (Cursor, Claude Code, Codex Desktop/CLI, Windsurf) collaborate without drift. Three additive sizes (S/M/L) plus an opt-in L+ autonomy profile. Trigger when bootstrapping a new repo's AI scaffold, promoting an existing harness to the next size, or auditing one for sprawl.

SKILL.md

19.3 KB, ~4.6k tokens by cl100k_base, as published. Nobody here has run it

Simple AI Harness Blueprint

A scaffold for repos where AI agents and humans collaborate. Three sizes (S/M/L), strictly additive — filenames never change between levels — plus an opt-in L+ autonomy profile (ADR-gated; not a fourth size).

When NOT to use

Skip for product code, PR reviews, IDE configuration, or generic project scaffolding.

Doctrine — the 5 Karpathy rules + M0

Every AGENTS.md generated opens with these. They override convenience and other rules in conflict. The 5 rules are posture; M0 is the verification mechanism.

Version stamp. Every generated AGENTS.md carries, right under its title, the line:

> Doctrine: v5 (5 règles Karpathy + M0) — appliquée <YYYY-MM-DD>

Fill the date with the day the doctrine was applied to that repo. The stamp is what makes doctrine drift greppable across a fleet (scripts/audit_fleet.py reads it). A repo deliberately left on an older doctrine states it the same way (e.g. > Doctrine: v4 figée — …) — a known lag is fine, an invisible one is not.

  1. Ask, don't assume. Ambiguity → ask before coding. Only in an explicitly activated autonomous mode (L+): pick the most reasonable interpretation, proceed, record the assumption.
  2. Simplest solution for simple problems, stronger for hard ones. No over-engineering, no preventive abstractions.
  3. Don't touch unrelated code — but surface what you find. Diff = scope of the ticket; raise smells as a separate issue.
  4. Flag uncertainty explicitly. Unsure → see rule 1; a small low-risk experiment + results beats false confidence.
  5. Suggest better ways. Propose lasting improvements over tactical fixes.

M0 — Verification. Before acting, state a verifiable success criterion and loop until it holds: trigger · stop criterion · validation · budget · stop/no-progress. The validation matrix, DoD, /tdd-loop, and (at L+) /loop all inherit from M0.

Pick the size

Depths of harness, not codebase sizes. Match signals:

SignalSML
AI agents in regular use1+1–32+ rotating
Human contributorsSoloSolo or 2–32+
Deployment criticalityLocal / hobby / early stagePublic / usersProduction / regulated
Project age with active agent edits< 3 months3–12 months12+ months
Friction "the agent forgot rule X"Rare or absentRecurringPainful or routine
Critical zones to protect0–12–55+
Existing AGENTS.md stateNone, or close to the ~100–150 line S baselineGrowing past 150 linesPast 250 lines, hard to navigate

Two or more signals at a level → adopt that level. Never jump straight to L on a fresh repo.

Note on the AGENTS.md signal: a freshly-generated S AGENTS.md is itself ~100–150 lines (5 Karpathy rules + M0 + critical zones + commands + rail rule + load order). When reading this signal, look at project-specific accretion on top of the baseline, not the raw line count.

L+ — the autonomy profile (opt-in, not a fourth size)

L+ is not the next rung on this matrix. It is a specialised profile layered on an L harness: the same L, plus the wiring to run bounded loops unattended. The size matrix never recommends it — promote only when all five are true, and record the decision in an ADR:

  • repeated loops with a verifiable goal,
  • real unattended execution (headless / scheduled),
  • parallel subagents or worktrees in use,
  • a CI / headless runner exists,
  • budgets and stop-conditions are genuinely needed.

None of these → stay at L. Adding L+ to a repo that doesn't need it re-creates the sprawl this blueprint fights. A loop without its three hard brakes (budget · no-progress detection · kill-switch) is forbidden.

Tree per size

Every level is strictly additive: filenames never change between sizes, you only add.

S — minimum with all 3 mechanisms present

your-repo/
├── README.md
├── AGENTS.md                            5 Karpathy + M0 verification + M1 (load order) + critical zones + commands + M3 rail rule
├── CLAUDE.md                            optional adapter — include only if Claude is a used agent
├── WORKFLOW.md                          process + when to invoke /learn (refers to AGENTS.md for rail discipline)
└── .agents/
    └── workflows/
        └── learning-loop.md             M2 — /learn per-event capture

M (additive over S) — routing + status + formalized PR flow

+ STATUS_APP.md                           living log: current stage, recent decisions, known debt
+ .agents/
│   ├── ROUTER.md                         task family → context capsule matrix
│   ├── context/
│   │   └── <critical-domain>.md          1–3 capsules per project; 80–150 lines each
│   └── workflows/
│       ├── retro.md                      post-release retrospective workflow
│       └── tdd-loop.md                   optional but typical at M
+ docs/retro/                             output of /retro
+ .github/
    └── pull_request_template.md          M3 formalized: risk rail declaration in PR body

L (additive over M) — full system, machine-enforced

+ .agents/
│   ├── patterns/
│   │   ├── INDEX.md                      registry: pattern ↔ pivot files in the codebase
│   │   └── <action-pattern>.md           short copyable procedures (≤ 120 lines)
│   ├── rules/
│   │   └── <tech-convention>.md          narrow technical rules; may carry `globs:` to auto-load
│   └── skills/
│       └── <project-skill>/SKILL.md      project-specific skills / Claude skills where applicable
+ .claude/                                Claude-Code runtime adapter (concept lives in AGENTS.md)
│   └── agents/
│       └── <reviewer>.md                 read-only subagent, tools allowlist, model: sonnet
+ docs/
│   ├── adr/                              Architecture Decision Records
│   └── learn/                            per-event /learn output (auto-created on first run)
+ scripts/
│   ├── validate_agent_context.*          meta-CI: the harness validates its own coherence
│   └── check_pr_rail_consistency.*       rail-guard: fails `green` PRs that touch CODEOWNERS paths
+ .github/
│   ├── CODEOWNERS                        source of truth for "critical paths"
│   └── workflows/
│       ├── agent-context.yml             non-blocking signal on .agents/** changes
│       └── pr-rail-guard.yml             blocking gate on PR risk rail consistency
+ _local/CI_RULEBOOK.md                   optional, gitignored: rationale of CI/agent choices

L+ (additive over L) — the opt-in autonomy profile (ADR-gated, not a size)

+ .agents/
│   └── workflows/
│       └── loop.md                        /loop unit-of-work: 3 hard brakes + reviewer verification
+ .claude/
│   ├── settings.json                      hooks: post-tool formatter, push-to-main defer, denied log
│   └── hooks/
│       └── gate_git_push.sh               defers `git push … main` for human approval (fail-open)
+ .github/
│   └── workflows/
│       └── headless-loop.yml              headless runner — manual-dispatch until brakes are exercised
+ docs/adr/
    └── ADR-0002-*.md                      records the decision to run autonomously (mandatory)

Promotion to L+ is gated (see §Pick the size → L+). .claude/logs/ is gitignored. A loop without its three hard brakes is forbidden.

File size budgets

FileSoft targetHard ceiling
AGENTS.md100–150300
CLAUDE.md40–60100
WORKFLOW.md150–200300
STATUS_APP.mdunbounded (log)n/a
.agents/ROUTER.md40–60100
.agents/context/*.md80–150200
.agents/patterns/*.md40–120200
.agents/workflows/*.md80–200300
.agents/skills/*/SKILL.md≤ 200500
.claude/agents/*.md30–60120
.claude/hooks/*.sh (L+)≤ 4080

Session-init budget: keep total auto-loaded context ≤ 6 500 lines across all files an agent reads before its first action. Trade per-doc ceilings against this total — a leaner ROUTER.md leaves room for a fatter WORKFLOW.md, and vice versa.

MCP / tool-schema budget: every connected MCP server spends context tokens on every turn (≈ 10–20k tokens for 50 tools without lazy loading). Cap a serious setup at ≤ 5 servers; prefer a code-graph/memory server, a Git server, a filesystem server, a web-search server, and a docs server. Fewer servers beats lazy-loading. Subagents with a narrow tools: allowlist keep the main loop's schema cost down.

The mechanisms (M0–M3) — templates only

M0. Verification (lives in AGENTS.md §Doctrine, right after the 5 rules)

The mechanism every task inherits — the executable backbone of rule 4. State a verifiable success criterion, then loop until it holds.

## M0 — Verification (the mechanism behind every task)

Before acting, state a verifiable success criterion, then loop until it holds:
- Trigger — what starts the work.
- Stop criterion — the verifiable signal it's done (red test → green, lint clean, smoke passes).
- Validation — the minimum commands from §Minimal validation matrix.
- Budget — a ceiling on time / iterations / tokens.
- Stop / no-progress — if not converging, stop and surface the blocker (rule 1), don't loop blindly.

The validation matrix, the DoD, /tdd-loop, and (at L+ only) /loop all inherit from M0.

M1. Init: ordered context load (lives in CLAUDE.md or the agent's adapter file)

Two progressive forms — the order grows one hop once the ROUTER exists.

At S (no ROUTER yet):

Before generating code for any new request, load context in this order:
1. AGENTS.md (project contract — implicit, never skip)
2. files directly touched by the request
3. additional documentation only if the task obviously requires it
Never load large documents "just in case".

At M and L (ROUTER present):

Before generating code for any new request, load context in this order:
1. AGENTS.md (project contract — implicit, never skip)
2. .agents/ROUTER.md (identify task family, pick minimum context)
3. the capsule / pattern the router points to (if any)
4. files directly touched by the request
5. additional documentation only if the router points to it
Never load large documents "just in case".

Always restate the active form in AGENTS.md §Role so the rule survives careless edits to the adapter.

M2. Learning: per-event capture (lives in .agents/workflows/learning-loop.md)

# /learn workflow
Invocation: /learn <family> <slug>
Families: release | candidate | incident | friction | refactor

Output: docs/learn/LEARN_<family>_<slug>_<YYYY-MM-DD>.md (≤ 40 lines):
- What helped (2–4 bullets)
- What slowed us down (2–4 bullets)
- ONE action retained (imperative, < 20 words)
- "Lands in: <target artefact>"
- Diffusion: files actually changed

Then update the target artefact and run `validate:agent-context`.
Hard constraint: ONE action per /learn.

M3. Risk rail enforcement (3 progressive forms)

The rail concept lives at every level — only the enforcement strength grows.

At S — agent self-declaration in AGENTS.md (no PR flow required):

## Risk rail (declare after every task)

After completing any task or before pushing/committing changes, declare a rail
in your message:

- Rail: green | amber | red

green = small/local; no critical zone touched; safe to merge fast
amber = behavioural or transverse; the user should scan the diff
red   = critical path or production risk; the user must review

The rail is informational at this stage (no CI enforcement). Its value is
forcing the agent to self-assess sensitivity, and the user to see it.

At M — formalized PR template at .github/pull_request_template.md:

## Risk rail
- Rail: green | amber | red

green = small/local; no CODEOWNERS path; no runtime/auth/release touch
amber = behavioural or transverse; short human review or owner approval
red   = critical path or production risk; full review mandatory

Touching a .github/CODEOWNERS path → minimum amber.
P1/P2 finding during review → automatic red until resolved.

At L — machine enforcement:

  • CI job pr-rail-guard: fail PRs declared green that touch any path listed in .github/CODEOWNERS
  • Meta-validator scripts/validate_agent_context.*: check every npm run X cited in AGENTS.md exists in package.json, every internal link in .agents/** resolves, and patterns/INDEX.md covers every pattern file

Codex Desktop compatibility

This repository is both a Claude skill and a Codex Desktop skill. Codex reads the same SKILL.md frontmatter/body; agents/openai.yaml only adds optional UI metadata. Do not fork the instructions by agent. Keep AGENTS.md canonical and use thin adapters (CLAUDE.md, GEMINI.md, etc.) only when a tool requires a specific entry file.

Workflow for invoking this skill

  1. Inventory. List existing blueprint files in the target repo.
  2. Infer and propose — don't interrogate.
    • Size: match the matrix above to repo signals — AGENTS.md line count, presence of .agents/, multi-agent traces, .github/, codebase age.
    • Agent adapters: scan for markers — CLAUDE.md, .cursorrules, .windsurfrules, GEMINI.md, etc. Include each adapter only if its marker is found, the user names that agent explicitly, or the user explicitly says "Claude" / "Cursor" / etc. is in use. AGENTS.md is always included.
    • PR workflow: detect .github/, .gitlab/, branch-protection rulesets. Add .github/pull_request_template.md only at M+ and if a GitHub-compatible PR flow is detected.
    • Propose in one sentence with reasoning, e.g. "Based on no existing AGENTS.md, a .cursorrules file, and a .github/ folder, I'd target M, with CLAUDE.md (you mentioned Claude) and .github/pull_request_template.md. Confirm or adjust?"
    • Ask explicit questions only when inference is genuinely ambiguous (multiple plausible sizes, conflicting agent markers, or the user contradicts the inference).
    • If existing artefacts exceed the requested level, warn about downgrade risk before removing anything.
  3. Generate the delta. For each missing file, propose the matching template from references/templates/<size>/. Confirm before writing. Never overwrite silently. Follow placeholder discipline (see below): do not invent commands, files, or critical zones the target repo doesn't have. In the generated AGENTS.md, fill the doctrine version stamp's <YYYY-MM-DD> with the current date — it is the one placeholder you always resolve at generation time.
    • The M templates include one worked example (a data-mutations.md capsule and a matching row in ROUTER.md) to model after. Replace or delete the example if mutations aren't a critical domain in the target project, and route the user's own capsules instead.
  4. Validate. Print the created tree, run any present validate:agent-context, ask the user to skim AGENTS.md first.
  5. Plant the next promotion criterion AND report remaining placeholders. Add a one-line "promote to next size when X" at the bottom of AGENTS.md. Then list every <placeholder> left in the generated files (commands, critical zones, project description, etc.) so the user knows exactly what to fill in before the harness becomes operational. An unfilled placeholder is the correct state — a filled one with invented content is a lie.

Adopting on an existing repo

If the target repo already has its own agent doctrine, never overwrite:

  • Don't replace an existing AGENTS.md. Propose an additive merge: keep the project's accreted rules, insert universal sections (5 Karpathy rules + M0, load order, rail discipline) only where they're missing.
  • Don't create a parallel .agents/ if the repo uses a different layout (.cursor/rules/, .windsurf/, custom paths). Adapt: place equivalent content there and note the divergence at the top of AGENTS.md.
  • Look for what's missing, not for what doesn't match this template. The blueprint's value on a mature repo is usually three things: navigation (ROUTER at M), drift check (validate_agent_context.* at L), per-event learning loop (/learn). Add only those.

A mature repo with a working memory should be enhanced, not rescaffolded.

Placeholder discipline

Templates ship with <...> placeholders (e.g. <run dev>, <typecheck>, <critical-zone>, <path/to/file.ts>). They are intentional — they mark what the project hasn't decided or doesn't have yet.

Hard rule: never replace a placeholder with an inferred value. Replace only when one of these is true:

  • The actual file or tool already exists in the target repo, and you have read or listed it.
  • The user has explicitly named the value (e.g. "I use ruff and pytest" → fill <lint> and <test>).
  • The user explicitly asked you to scaffold product code too (rare; outside this skill's default boundary).

If you don't know, leave the placeholder. Empty placeholders tell whoever opens the file what's not yet decided. A filled placeholder with invented content is a lie that compounds over time and rots the harness quickly.

A user prompt mentioning a language or framework ("Python CLI", "Next.js side-project", "React + Node API") is a hint about direction, not a guarantee that tooling exists. The repo is the source of truth; the user's words are clues for what to ask about, not licenses to assume.

When the bootstrap is done, the report from step 5 of the workflow must enumerate every remaining <placeholder> so the user can fill them deliberately.

Health metrics

SymptomCauseFirst fix
Agents ignore the routerInit rule buriedRestate at top of AGENTS.md §Role
Same friction reported 3+ times/learn never ranRun one /learn friction retroactively
AGENTS.md > 300 linesDomain doctrine in constitutionExtract to .agents/context/<domain>.md
.agents/ files contradictNo validatorAdd validate_agent_context.*; schedule a /retro
Token-heavy loads for trivial tasksROUTER too permissiveTighten rules; mark files "do not load by default"
Capsule untouched 3+ monthsDoctrine stable or zone deadInline into AGENTS.md or delete with an ADR
New agent ignores rulesReads a different config fileSymlink or copy AGENTS.md to its expected name

Two or more symptoms simultaneously → schedule rework.

References

  • references/templates/S/, M/, L/ — copyable file templates per level
  • references/health-metrics.md — extended diagnostics

Boundaries

Operates only on the documentation surface agents read. Does not touch product code, run tests, configure IDEs, or decide architecture.

Gives 0 of the 12 instructions most memory context skills give in ~4.6k tokens

Counted across 674 of the 847 authors here whose files we hold, read 2026-08-06

  • inform the user when setup is completein 21 of 674, across 6 files
  • confirm the draft with the user before writingin 21 of 674, across 6 files
  • update the agent skills block in place if it existsin 21 of 674, across 6 files
  • present findings to the userin 20 of 674, across 5 files
  • write the three docs files from seed templatesin 20 of 674, across 5 files
  • ask the user about each decision one at a timein 19 of 674, across 4 files
  • edit CLAUDE.md if it existsin 18 of 674, across 3 files
  • explore current repo statein 18 of 674, across 3 files
  • do not overwrite user edits to surrounding sectionsin 18 of 674, across 3 files
  • back up the original file before overwritingin 16 of 674, across 8 files
  • keep the memory index under 200 linesin 15 of 674
  • Provide actionable steps and verificationin 13 of 674, across 2 files

Said here and by no other author read

  • Add files additively across harness sizes
  • State a verifiable success criterion before acting
  • Adopt a harness size matching two or more signals
  • Record the decision to run autonomously in an ADR
  • Keep total auto-loaded context under 6500 lines
  • Cap connected MCP servers at five

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.