agentsclimarketplace

Pay down agentic debt

Skill impactbrussels/AINativeOS/skills/pay-down-agentic-debt

The open OS for building an AI-native company in hard-mode sectors. A 15-chapter Handbook, a 90-term Dictionary, and 24 runnable skills for Claude Code, Codex, Cursor, and Gemini. CC-BY-4.0 / Apache-2.0.

Install
npx -y skills add impactbrussels/AINativeOS --skill pay-down-agentic-debt

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Use when an agent-built codebase has grown faster than anyone understands it and every change feels risky - when they say "the codebase is a mess", "I'm scared to touch this", "one fix broke three other things", "it worked when we built it", "we have a God Agent", "there's a ghost in the system", "tech debt is killing us", or a small prompt edit caused a regression three hops away. Produces a ranked debt ledger (interest x principal), the CACE-safe paydown order, and the eval gate each change must pass. Invoke after architect-before-code, with eval-and-safety-harness.

The file declares its own license as CC-BY-4.0. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

6.3 KB, as published. Nobody here has run it

Pay Down Agentic Debt

In a machine-learning system, changing anything changes everything. The debt an agent leaves behind does not sit in one place and wait like the old kind; the intelligence at the centre is a learned thing, and a learned thing entangles whatever it touches. So a surgical, obviously correct fix in one corner degrades a model three hops away you would swear is unrelated, and you are debugging at midnight certain the system has a ghost in it. It does not have a ghost. It has agentic technical debt, and the founders who survive growth are the ones who rank it by interest rate and pay it down without breaking ten things to fix one.

The method

Find it, rank it by interest, pay it down CACE-safe through the eval gate. Full framework: references/debt-method.md. Source: Handbook Chapter 07.

Step 1: Audit the repo for the six debts

Point your agentic-coding tool at the codebase and force it to find the categories below, not a vague "what's messy". An unranked list paralyses; a named taxonomy you can score.

Debt categoryWhat to grep forWhy it bites
Prompt sprawlOne mega-prompt; rules added per caseMaximum entanglement; nothing testable in isolation
Entangled contextA signal feeding more than one learned componentA fix here moves outputs you cannot predict
Untested pathsCode with no eval asserting on its contentYou ship blind; silent failure has no alarm
God AgentOne agent doing many jobs in one contextUnauditable, unchangeable, brittle by construction
Missing evalsNo held-out set, no eval suite at allYou cannot prove any change is safe, only hope
Dead retrieval / broken loopA signal generated but never storedThe leak in the moat: thrown-away learning

Step 2: Score each item by interest x principal

Interest is how fast the debt compounds; principal is the cost to fix. Rank by interest first, always. See the scoring grid in references.

Interest (rank highest first)Score
Blocks data collection (a silently open loop, a dropped signal)Highest
Blocks scaling (an architectural flaw growth will hit)High
Risks a CACE regression (entangled, touches many learned parts)Medium
Merely cosmetic (ugly code that works and never changes)Lowest, leave it

The debt that screams is rarely the debt that kills you. The quiet leak nobody is complaining about, the broken feedback loop, that is usually item one.

Step 3: Pay it down CACE-safe, one thing at a time

Change one thing. Run the full eval. Measure. Keep or revert. Never bundle. The diff is not the blast radius; the eval is the only thing that knows what moved.

Step 4: Gate every change on the eval suite

No paydown ships without passing the eval that protects the parts you did not touch. If you have no eval suite, that is your highest-interest debt; build it before you refactor anything.

Output

  • A ranked debt ledger: each item scored interest x principal, highest-interest at the top.
  • The safe paydown order: which item first, and the one isolated change that pays it.
  • The eval gate for each change: the assertion that proves you broke nothing downstream.
  • Next: log the trade-offs you deferred via capture-learning; re-run this audit monthly.

Constraints

  • No big-bang refactor. Rewriting the God Agent in one swing is the faster mess, not the faster company. Extract one specialist, gate it, repeat.
  • No fix without an eval. CACE means you cannot reason a change safe, you can only measure it. Trusting the green diff is how the regression ships.
  • Pay the highest-interest debt first even when the ugly code is screaming. Especially then.
  • Stay theme-agnostic; the founder supplies the domain, you supply the rigour.

Dictionary

CACE · agentic technical debt · entanglement · eval · the God Agent · infinite loop

Copy-paste version

For non-coders: paste into a coding tool with repo access (Claude Code, Cursor, Codex). For the ranking and plan alone, paste into any chatbot (Claude.ai, ChatGPT).

Act as my staff engineer doing an agentic-debt audit. My product: [ONE_LINER]. Domain: [DOMAIN].
The codebase was largely AI-generated and now every change feels risky.
1. Audit for six debts: prompt sprawl (one mega-prompt), entangled context (a signal feeding more
   than one learned component), untested paths (code no eval checks), a God Agent (one agent doing
   many jobs), missing evals (no held-out test set), and dead retrieval / broken feedback loops (a
   signal generated but never stored).
2. Score each by INTEREST x PRINCIPAL. Rank highest interest first: blocks data collection > blocks
   scaling > risks a CACE regression > merely cosmetic. Tell me ugly code that works and never
   changes is almost free; leave it.
3. Give me the safe paydown order: ONE isolated change per item, the eval that must pass before it
   ships, and what breaks if I bundle changes. Remember CACE: changing anything changes everything.
Do not propose a big-bang rewrite. If I have no eval suite, tell me that is debt item one.
End with the single first change to make this week and the eval that gates it.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.