Ground
Armor for coding agents: a playbook of Claude Code skills that make an agent prove its work, not describe it.
npx -y skills add yiyaw-lab/agent-armor --skill groundAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Demand-evidence gate on the build decision — bind every product bet to real demand evidence so craft does not drift into impressive-but-unwanted. Classifies evidence against a hard hierarchy (revealed-&-costly > specific-stated > vanity proxy) and REFUSES to license a build on anything weaker than specific stated demand: defaults to UNKNOWN/PUSH and ships the cheapest test that would settle it. The demand-axis sibling of /frontier (which is supply/craft); makes "someone wants this" a falsifiable graded belief via your prediction ledger, and the verdict gates ship (redline-enforceable). Modes: <bet|product> (extract + falsify + verdict) | audit | apply. Alias: /pmf.
SKILL.md
11.1 KB, as published. Nobody here has run it
The law this skill runs on: a product dies not from bad craft but from being aimed at no one. Your whole suite optimizes SUPPLY — /frontier pushes craft to the 2026 standard, /harmonize enforces consistency, /raze and /moonwalk perfect the code, /taste encodes your aesthetic — and every one of them makes the product more impressive. None checks whether anyone wants it. Run as designed, the suite manufactures the exact failure you fear: a frontier-grade thing the market is indifferent to ("generation is free; judgment isn't" — and the deepest judgment is what to build). ground is the missing demand anchor. It forces each bet's unstated demand assumption explicit, classifies the evidence against a hard hierarchy that refuses to let attention masquerade as demand, and — in the common case — hands you the cheapest test that would settle it rather than a greenlight. It makes "someone wants this" a falsifiable, graded belief instead of a founder's conviction.
Three hard invariants:
- PULL is earned, never assumed — cite the demand or declare UNKNOWN. Evidence is ranked: Tier 0 (revealed & costly — someone paid, switched, retained past week 1; a competitor's real paying/churn numbers; usage on your own thing), Tier 1 (specific, recent, unprompted stated pain a named person would pay to fix), Tier 2 (vanity — search volume, funding, hype, "market size", upvotes, competitor marketing). A PULL verdict requires ≥ Tier-1 evidence, cited. Vanity proxies establish attention, never demand — they can only DOWNGRADE; alone they cap the verdict at UNKNOWN. The default is UNKNOWN/PUSH, and the common deliverable is the cheapest test that would upgrade it — not a verdict.
- Falsification framing + founder-bias guard. The research runs as the BEAR case — argue it flops, find the graveyard of similar dead products, find who does NOT want it; the proposer is never the judge. ground names the feature YOU find most impressive precisely to flag it: supply-pull is not demand-pull, and the more impressive you find it, the higher the PUSH suspicion. PULL is earned by surviving disconfirmation, never by a fluent market narrative.
- Gates and grades — never just a document. The verdict is a required input to the ship/build decision: a build of a PUSH or UNKNOWN bet without a logged demand test is flagged (enforce the gate via /redline). Each bet carries a falsifiable demand-prediction: a grader-measurable one is emitted to your prediction ledger (auto-flags on a FALSE grade); a MANUAL-measured one (retention, install→2nd-run, acted-on-push rate — which an automated grader cannot read) is recorded in the playbook with a manual-grade date. Never emit a metric the grader cannot read (the placement-split trap). The playbook is outcome-bound, not a passive doc nobody reads at decision time (the false-CLEAN trap).
/ground <bet|product> (default — extract + falsify + verdict)
Fan the research to a scout subagent; only the cited evidence + verdict return to main context.
- Extract the implicit demand bet — and decompose a portfolio. Restate what is being built/planned AND the unstated assumption underneath it: "we believe the market wants Y / will pay / switch / adopt because Z." If the target is a PRODUCT or a bundle (not a single feature), DECOMPOSE it into its constituent bets and ground each SEPARATELY — never average a bundle into one verdict; the bundle-as-one-bet framing is the weakest and hides which parts are PULL inside a PUSH (and vice versa). Ground the most supply-loved / riskiest bet first. Name the feature you find most impressive — the founder-bias flag (invariant 2).
- Falsify it (the bear case). Run /deep-research as the executor, framed to DISPROVE demand: the dead similar products and why they died, who does not want this, the reasons it stays a demo. Build the strongest disconfirmation, not a balanced survey.
- Classify the evidence by tier, citing each piece. Tier 0 / Tier 1 / Tier 2. Vanity (Tier 2) can only downgrade. No citation = it does not count.
- Render the verdict — exactly one, with its evidence:
- PULL — ≥ Tier-1 demand evidence, cited. The only verdict that licenses a build now.
- PUSH — supply-attractive but demand-unsupported (you looked with a falsification frame and found no Tier-0/1 support). The danger profile — flag loudly; not a proof of doom, a reason to test.
- UNKNOWN — only vanity / insufficient signal.
- DRIFTED — prior support no longer present in recent Tier-0 behavior (demand moved since the playbook).
- When sub-PULL, name the settling signal, then design the cheapest test (the common deliverable). First name the SINGLE revealed-behavior micro-signal that would flip the verdict — the one number that settles it (e.g. install→ran-2nd-command, free→paid, week-4 acted-on-push retention, your own open-and-act rate). Then from the fixed cheap-test menu — landing/smoke test, pre-sale or LOI, fake-door, 5 specific-pain interviews, concierge MVP — pick the cheapest experiment that produces THAT signal. This is the output in the common case, NOT a greenlight.
- Propose the playbook delta + the prediction. Write nothing yet: surface the verdict, the tier-cited evidence, the test (or, for PULL, the bet to commit) and the falsifiable demand-prediction it would emit. If a SHIP is in play, route the gate to /redline.
/ground audit
Read-only. Re-grade the existing playbook against the LATEST signals: surface PUSH bets currently being built or planned (impressive-but-unwanted), DRIFTED bets (demand moved), and demand-predictions now due for grading. Each with its tier-cited evidence or the explicit "only vanity / no signal." Propose and write nothing. The "is our roadmap still aimed at demand, and has the market moved under us" check — mirrors /raze, /hone, /harmonize audit.
/ground apply
Execute ONLY approved updates from a prior scan. Commit the playbook delta — the demand thesis, dated, each bet tagged PULL/PUSH/UNKNOWN/DRIFTED with its tier-cited evidence — and emit GRADER-MEASURABLE demand-predictions into your prediction ledger (ride an existing scheduled grader; a new metric kind should need zero new scheduler wiring); a MANUAL-measured metric (retention / install→2nd-run / acted-on-push) is recorded in the playbook with a manual-grade date instead, never handed to a grader that would never score it. For a PUSH/UNKNOWN bet being built anyway, install the demand-test gate via /redline. One bet = one commit. Record the running metric (PUSH-caught, prediction hit-rate) in the playbook's ledger.
Rules
- Evidence hierarchy is law: PULL ≥ Tier-1, vanity caps at UNKNOWN, cite or declare UNKNOWN — never infer demand from attention. Search / funding / hype establish that people are LOOKING, never that they will PAY, SWITCH, or STAY.
- The default is UNKNOWN, and UNKNOWN ships a test, not a greenlight: generate the signal with the cheapest experiment; do not read the vanity and call it demand. For a pre-PMF project this is the point, not a limitation — conviction becomes evidence.
- A sub-PULL verdict ends the craft lane in-session: once PUSH/UNKNOWN/DRIFTED is rendered, the only next artifact for that bet is the named cheapest test (or its result) — no further drafts, polish, essays, repoints, or analyses "to be sure." Producing more craft after the verdict is the gate violation happening live: name it and stop. Needing ever-more analysis to keep the bet standing is the kill signal, not an analysis gap.
- Decompose portfolios; never average: a product is a set of bets — ground each separately and report per-bet verdicts; a single averaged verdict on a bundle hides the PULL inside a PUSH (and the PUSH inside a PULL). Always name the settling micro-signal before designing a test.
- Falsification + founder-bias guard: the research builds the bear case; the proposer is not the judge; the more impressive you find it, the higher the PUSH suspicion; PULL is earned by surviving disconfirmation.
- Gates and grades, not a passive doc: the verdict is a required input to ship (a build of a PUSH/UNKNOWN bet without a logged demand test is flagged — enforce via /redline); each bet carries a falsifiable demand-prediction — emitted to the prediction ledger if grader-measurable (auto-flag on FALSE), else recorded in the playbook for manual grading.
- Recency through the same hierarchy: DRIFTED needs Tier-0 revealed movement; a recent Tier-2 spike is noise → UNKNOWN, not DRIFTED. No special pass for "but it's recent."
- Lane discipline: ground is the DEMAND axis; /frontier is the SUPPLY/craft axis (is it good? — they pipeline: frontier makes it excellent, ground aims the excellence at demand); /deep-research is ground's research executor (do not re-roll it); /redline enforces the ship gate; your prediction ledger grades the demand bets; /taste is YOUR internal aesthetic (a supply preference), never the market's demand. Your ship/no-ship decision consumes ground's verdict. A project fact →
/burn rent. - Propose-on-approval, reversible: default writes nothing; playbook, prediction, and gate land only in
apply, one bet per commit, dated and tier-cited. - Essentialness metric (honesty hook): PUSH bets caught before build (impressive-unwanted features killed/redirected pre-investment) + demand-prediction hit-rate. Never flags a PUSH and never beats a coin flip on predictions = astrology → /raze it. It earns its keep the first time it stops you sinking a month into a frontier-grade feature the market was never reaching for.
Pipeline placement: /frontier (make the craft excellent) → ground (prove the excellence is aimed at real demand, or hand back the cheapest test) → build → prediction grading (grade the demand bet against reality) → re-ground on drift. ground is the demand anchor on the whole flywheel; without it the suite's supply-side power compounds in a direction nobody asked for. Alias: /pmf runs this skill — ground is the self-documenting verb (stay grounded in evidence, do not float on proxy and conviction); pmf is retained as the recognizable term.