Karpathy coding loop
Skill yshms/karpathy-claude-skills/skills/karpathy-coding-loop
Claude Code skills that make Claude follow Andrej Karpathy's published AI-usage workflow — distilled from his X threads, blog, talks, and interviews
npx -y skills add yshms/karpathy-claude-skills --skill karpathy-coding-loopAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
2 things to look at
- no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Use before writing, modifying, or refactoring any code, and at each temptation point inside it: starting an implementation, proceeding on an unstated assumption, choosing between two readings of a request, adding an abstraction, config option, or error handler nobody asked for, editing a line or comment the task did not require, or reporting work done.
SKILL.md
7.3 KB, ~1.6k tokens by cl100k_base, as published. Nobody here has run it
Karpathy Coding Loop
Agent code no longer fails on syntax. It fails with "subtle conceptual errors that a slightly sloppy, hasty junior dev might do" (X status/2015883857489522876, 26 Jan 2026, agent-coding notes): wrong assumptions run with silently, bloated abstractions, dead code left behind, comments and code changed as side effects, no pushback when pushback is warranted. Countermeasure: think -> minimal -> surgical -> verified.
Scope. These rules bias caution over speed. For a typo fix or an obvious one-liner, use judgment — not every change needs the full rigor (tradeoff caveat adapted from multica-ai/andrej-karpathy-skills, MIT).
Scaling down changes the FORM of a rule, never whether it applies. On a trivial edit you skip the written plan and the Assumptions: block. You never skip stating a material assumption in one inline line, and you never skip tracing the diff to the request.
"We're behind schedule" or "no questions" converts assumptions from a question into a one-line preamble. It does not delete them.
The Loop
1. Think before coding
- State assumptions BEFORE code, not after. A material choice — a magic number, a scope boundary, a cache-invalidation policy — is an assumption.
- Multiple reasonable interpretations exist -> present them; do not pick silently.
- A simpler approach exists -> say so. The request looks wrong -> push back with the tradeoff.
- Something is unclear -> stop, name what is confusing, ask.
- Cannot ask -> lead the response with an
Assumptions:block and make each one trivially reversible.
2. Simplicity first
- Minimum code for the stated problem. No features beyond the ask, no abstraction for single-use code, no configurability nobody requested, no error handling for impossible scenarios.
- Two heuristics: if you write 200 lines and it could be 50, rewrite it; and Karpathy's observation that agents ship a bloated 1000-line construction and cut it to 100 the moment you ask whether there is a simpler way (Jan 2026 agent-coding notes).
- Senior-engineer test: would a senior engineer call this overcomplicated? If yes, simplify.
- After it works, run one explicit simplify pass: cut abstractions, delete only the dead code YOUR change orphaned.
- The bloated versions are not wrong patterns. They are premature ones.
3. Surgical changes
- Every changed line traces to the request. No drive-by refactors, no improving adjacent code, no altering comments or code you do not fully understand.
- Match existing style even where you would do it differently — including quote style and whether the file uses type hints. Adding type hints to an untyped file while fixing a bug is a side-effect edit.
- Notice unrelated dead code -> mention it, do not delete it.
- Remove imports, variables and functions YOUR change orphaned. Do not remove pre-existing dead code unless asked.
- Keep the diff small enough for a human to audit in one sitting. Karpathy's framing: keep the AI on a leash, work in small incremental chunks (YC AI Startup School / Software 3.0, announced in X status/1935518272667217925, 19 Jun 2025; transcript write-up at latent.space/p/s3).
4. Goal-driven verification
- Define the machine-checkable criterion BEFORE coding. Transform the task:
- "add validation" -> write tests for invalid inputs, then make them pass
- "fix the bug" -> write a test that reproduces it, then make it pass
- "refactor X" -> tests pass before and after
- Multi-step work states a plan first, one verify per step:
1. [Step] -> verify: [check]
2. [Step] -> verify: [check]
3. [Step] -> verify: [check]
- Loop generate -> verify until the check passes. Never claim done without running it.
- Strong criteria let you loop independently; "make it work" forces constant clarification.
- Aim tenacity at the criterion, not at expanding scope.
Before an external API
Do not write against a library or API from recollection. Get its current official docs into context first -> karpathy-context-engineering.
Red flags — STOP
- About to turn an unstated material assumption into code.
- Thinking "I'll note my questions at the end."
- Adding a class, config value, or option "for flexibility."
- Editing a comment or line the task did not require.
- Changing quote style or adding type hints while doing something else.
- Reporting done without having run anything.
Excuse -> Reality
| Excuse | Reality |
|---|---|
| "Team said no questions" | Then state assumptions inline before the code. |
| "It's a reasonable default" | Unstated != reasonable. One line to state it. |
| "More robust this way" | Robustness nobody asked for is bloat — the first-listed complaint in the Jan 2026 notes. |
| "I cleaned it up while there" | Side-effect edits are the second. Revert them. |
| "I added type hints while I was there" | Style drift is a side-effect edit. Revert. |
| "The strategy pattern is best practice" | Best practice at the wrong time is bloat. One function until a second case exists. |
| "My plan step is review and improve" | That is not a criterion. Name the check that will run. |
| "I'll write the test after" | Then the test encodes the code, not the bug. |
Worked examples
Before adding any abstraction, config flag, or class, and before touching a line outside the requested change, open references/examples.md and find the matching BAD/GOOD pair.
Sections map to the four rules: 1 hidden assumptions and multiple interpretations, 2 over-abstraction and speculative features, 3 drive-by refactoring and style drift, 4 vague-vs-verifiable, multi-step verification, and test-first.
Working if
Diffs contain only requested changes; no rewrites caused by overcomplication; clarifying questions arrive before implementation rather than after mistakes.
Sources
X status/2015883857489522876 (26 Jan 2026, agent-coding notes) — the failure-mode list and the 1000-lines-to-100 observation. YC AI Startup School "Software 3.0" talk, announced in X status/1935518272667217925 (19 Jun 2025, talk announcement), annotated transcript at latent.space/p/s3 — the leash and small-incremental-chunks framing. nanochat release thread (HN, Oct 2025) — Karpathy reported coding agents were "net unhelpful" on that repo, possibly because it sits too far off the training distribution.
Rules 1-4, the plan template, the transform table, the tradeoff caveat, the working-if signal, and all of references/examples.md adapt multica-ai/andrej-karpathy-skills (formerly forrestchang/andrej-karpathy-skills), MIT per its README.md and .claude-plugin/plugin.json. Full notice at the head of references/examples.md. That README states the rationale for rule 4 in two fragments it attributes to Karpathy — "LLMs are exceptionally good at looping until they meet specific goals" and "Don't tell it what to do, give it success criteria and watch it go" (README.md, "Key Insight" section) — but pins them to no primary artifact, so treat the attribution as the repo's, not as a verified Karpathy citation.
What ships with it: 1 file
20.7 KB alongside SKILL.md
references/
- examples.md20.7 KB