Deepen
Use when growing an already-playable Godot game along a depth axis (systemic / content / run-meta) without regressing proven behavior. The first ITERATION skill in the GameForge loop — it EXTENDS the rules engine (the deliberate inverse of the asset re-skin's frozen-logic rule), TDD-ing each new sub-system against selftest.gd as a regression guard, records manifest.depth_pass, loops back through validator, and does NOT advance status.From its SKILL.md
npx -y skills add qmertesdorf/GameForge --skill deepenAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
26.4 KB, ~6.6k tokens by cl100k_base, as published. Nobody here has run it
deepen
Take an already-playable game and grow it along ONE depth axis — systemic (new
interacting mechanics), content (more of the same), or run-meta (map / events /
economy / progression) — without regressing what already works, and prove the new
depth landed. Every other GameForge skill is build-once; deepen is the loop's first
in-place ITERATION skill.
Loop position
prompt → concept → builder → validator → playtest → ( deepen → validator → playtest )* → asset → visual-audit → audio → packager
deepen operates in place on a validated/playable game and loops back through the
validator. It does not advance status — a deeper game is the same status, just
bigger.
When deepen is REQUIRED (the * is not zero)
The loop writes ( deepen → validator → playtest )*, and the POC read the * as
"optional" — so it ran once across eleven titles (only deckbuilder-0001 has a
depth_pass). That is the single biggest reason the catalogue plays "weak": every other
game shipped at first-playable depth — a thin prototype with a solved loop and a content
ceiling a minute in. A playable game is not a finished game; it is a prototype that
proved its loop runs. At least one depth pass is REQUIRED before a title is a
candidate for asset/polish — the asset skill now gates on the presence of a
depth_pass and bounces a never-deepened game back here. Spend iterations deepening one
good loop, not restarting near-duplicates (the match3-survival 0001→0002→0003 churn re-
attempted a flawed concept three times instead of deepening one). Run deepen until the
content-ceiling and dominant-strategy tests below both pass at least once.
Inputs
- A game at status ≥
validatedwith a workinggames/<id>/selftest.gd. - A chosen depth axis + scope (spec-given, or assessed per the method below).
Outputs
- Extended game code; a grown
selftest.gd; amanifest.depth_passrecord. - Durable lessons folded back into this skill.
The method
- Assess depth — diagnose where it's shallow before choosing what to add. Run
these two tests on the current game; they pinpoint the leak nearly every weak GameForge
title shares:
- Content-ceiling test: play (or read the loop) and ask "on what beat does the game stop introducing anything new?" — the last new mechanic, enemy, recipe, or rule the player meets. Both shipped POC games hit their ceiling early (shopkeep introduces nothing after day 4; match3-survival's only escalation is a shrinking timer on one static threat). Everything past the ceiling is the same system replayed — that's the "weak" the player feels. Your job is to push the ceiling out: stage in newness over the session.
- Progression-payoff test (reject DEPTH-AS-MULTIPLIER). "Newness" must be new gameplay, not a bigger number on the same decision. For every progression vector the game offers — going deeper, levelling, scoring, a higher tier/day/wave — name what the player DOES at tier N that they could not do at tier 1. If the only answer is "the same action, for more points/value," the ceiling has not moved — that progression is a score multiplier wearing a costume, and it is the single most common reason a game with "progression" still feels shallow (diver-0001 shipped exactly this: depth only scaled treasure value, so there was no gameplay reason to descend — owner-rejected on playtest). A progression vector earns its keep only when reaching it unlocks a new decision, a new interaction, or gated content that plays differently — a destination, not a dial.
- Dominant-strategy test: name the single move a skilled player repeats every beat. If
it's strictly best (most reward and safest), the loop is solved and no amount of
content fixes it — you must add a cost / tradeoff that makes the dominant move
situational (this is the systemic axis, and the inverse of
concept's tradeoff gate: whenconceptlet a solved loop through,deepenis where it gets repaired). match3 -survival shipped solved (purge = both the safe move and the high-score move); fixing that is higher-leverage than any new content on top. - Core-before-meta gate (don't bolt a wrapper onto a shallow core). A run-meta or
content layer multiplies whatever the core loop already is. If the core loop is the thin
part — fails the dominant-strategy test, or offers one real decision — adding upgrades / a
map / an economy on top just yields a longer shallow game, and the meta layer reads as
busywork because the thing it wraps isn't worth repeating. Before choosing
run-metaorcontent, confirm the core loop itself clears the dominant-strategy + progression-payoff bars. If it doesn't, the axis is systemic (fix the loop first) no matter how tempting the shiny meta layer is. (diver-0001's first deepen got this wrong: it added a run-meta upgrade economy over a one-decision core, so the upgrades had nothing meaningful to deepen.) - Upgrade / reward-coherence test (for every unlock, upgrade, or reward you add). Two questions per item: (1) "what does the player DO differently after getting this?" — if the answer is only "the same thing, slightly better," it's a dead stat, not a decision; prefer upgrades that open new play or change a choice (gate access to new content, enable a new tactic, flip a risk calculus) over flat ±X nudges. (2) "can the player FEEL it on the very next run/dive?" — a buff that needs 3–4 stacked levels before it's perceptible is invisible and reads as pointless (diver-0001: +16 air/level barely changed reachable depth, so the shop felt meaningless). Tune so one purchase visibly changes what the player can do. Then pick ONE axis — the one that addresses the leak the tests found — and the single highest-leverage expansion on it. Don't widen three axes at once. Per-axis playbook:
- systemic (fixes a solved/shallow loop): add an interacting mechanic that creates a new decision — a cost on the dominant move, a second resource that contends with the first, a threat/opportunity that rewards a different response than the default. The test: after it lands, the dominant-strategy test must no longer have a single answer.
- content (fixes a low ceiling on a loop that's already a real decision): more of the same kind — new recipes/enemies/cards/tiles — staged in over the session, not all at start, so the ceiling moves. Only reach for this once the dominant-strategy test passes; piling content onto a solved loop just makes a longer solved loop.
- run-meta (fixes "no reason to start run 2"): a wrapper that makes sessions differ and accrue — a map/event/economy/unlock track, escalating modifiers, a persisted best or prestige. Gives the loop somewhere to go across plays.
- Decompose into sub-systems. Each one purpose, with a clean interface (data layer + logic + screen). Find the extension seams: where the existing code already supports growth (e.g. a static-func data table) vs. where you must refactor to create a seam first (e.g. a hardcoded linear sequence → a data-driven state machine). Refactor-for-seam before adding content.
- TDD each new system on the self-test — the validation spine:
deepenEXTENDS the logic; it does NOT freeze it. This is the deliberate inverse of theasset/re-skin "logic FROZEN" rule. Confusing the two is the classic mistake — re-skinning must not touch rules; deepening is all about touching them, safely.- Existing assertions are the regression guard.
SELFTEST OKmust hold after every change — andUITEST OKtoo if auitest.gdexists: deepening adds screens and controls, and a new view that renders fine can still swallow taps or skip its rebuild event (invisible to selftest, which bypasses the view). New tappable screens get newuitest.gdchecks, same RED→GREEN discipline. PLAYTEST OKtoo if agames/<id>/playtest.gdexists — deepening changes TUNING, which is exactly what breaks winnability. A depth pass that retunes costs, gate depths, spawn geometry, or the ramp can make the game unwinnable while every logic assertion stays green (theplaytest-auditskill exists because of exactly this). Re-run the balance bot after the pass; aPLAYTEST FAILon re-validation is attributed to thisdeepenpass, and the fix is tuning, never weakeningselftest.- Keep the generate-and-verify gate green if the game has one (
make_verified+Solver, perbuilder). Adding content on the content axis is the classic way to silently introduce unsolvable instances: every new tile / recipe / wave type / map piece widens the instance space the generator can deal, and the solver guarantee only held over the old space. Re-run the generate-and-verify selftest assertions (everymake_verifiedsolvable over K seeds; fallback rate still rare; fallback still solvable) after the pass — a regression here is attributed to thisdeepenpass, and the fix is the generator / new content, never weakeningSolver.is_solvable. If the depth pass adds discrete generated content to a game that didn't have the gate (e.g. the content axis turns a fixed layout into a procedural one), that is exactly when to introducemake_verified— treat it as a new sub-system with its own RED→GREEN assertions. - For each new system, write its assertion first (RED) → implement → GREEN. Prove new mechanics the same deterministic, headless way the original logic was.
- Never weaken or delete an existing assertion to make room. If a new system genuinely changes old behavior, surface it — call it out and confirm it's intended — never silently overwrite the guard.
- A pure refactor adds no new behavior. If the behavior it restructures isn't already covered, pin it with a characterization assertion first (one that passes both before and after the refactor), then refactor. Pure refactors add no new-behavior assertions.
Balance tuning (parameter search) — propose a config, don't hand-guess
A pass that changes tuning (systemic/run-meta retunes costs, gate depths, drain,
spawn geometry, the ramp) is exactly what makes a game unwinnable or unfair while every
logic assertion stays green. Instead of hand-guessing constants and re-running the bot,
search the tuning space against playtest-audit's metrics:
- Add the
GF_TUNE/GF_SEEDseam (apreload-ableTunestatic, pergames/diver-0001/Tune.gd): the data layer reads each tunable fromTune.num(...)defaulting to itsconst. UNSET env → identical production behavior. Only seam the constants you intend to search. - Emit the metrics contract:
playtest.gdmust print onePLAYTEST METRICS {json}line (seeplaytest-audit) carrying the numbers it already computes + the invariant booleans. - Declare a
balance.spec.json(search space + objective). The objective HARD-REJECTS any config failing the playtest invariants, then scores survivors by distance OUTSIDE target BANDS — never by maximization (maximizing earnings/clear-rate yields a trivially easy game). Use single-player metrics + a retention/engagement proxy (low time-to-first-goal, accruing-but-not-instant economy, a smooth/tight difficulty curve via the air/HP margin, did pushing pay off). Pick floor vs. two-sided band per metric's meaning (see Lesson 1): a cautious-bot solvency rate (does a careful player reliably succeed?) is a hard floor inrequire, NOT a two-sided band — banding it penalizes the very robustness you want. Two-sided bands fit metrics whose value reflects difficulty/pacing (time-to-first-goal, air margin, commissions filled), or a win-rate that genuinely reflects challenge (a roguelike clear-rate). NOT win-rate disparity. The realistic retention bar is top-quartile ~7-8% D7 (GameAnalytics def) — do NOT anchor on the old unverified "20%"; and we do not literally measure D7, so the proxy is a heuristic. - Run
node tools/balance.mjs <game-dir> <spec.json>(each candidate is run across K seeds so "clear-rate" is meaningful and the config isn't seed-lucky). - READ the per-focus-point breakdown + the non-dominated shortlist and CHOOSE — weigh the tradeoffs yourself (great pacing vs. borderline economy); do not blindly take the lowest composite. The composite is a heuristic sort key, not a verdict.
- Apply the chosen config to the defaults, then re-run the full gate set
(
SELFTEST/UITEST/PLAYTEST) with env unset.
Honesty rule (load-bearing): the tool proposes; the human playtest decides fun.
No validated automated fun proxy exists — the search guarantees winnable/fair/well-paced,
never fun. An owner "this isn't fun" verdict overrides any proxy win. Record the chosen
config + why in depth_pass.notes.
Lessons from first use (diver-0001 dogfood):
- Lesson 1 — solvency is a FLOOR, not a band. The first objective two-sided-banded
clear_rateat[0.6, 0.9]and the search penalised the diver for the cautious bot always banking (100%). But for push-your-luck (and most solo games) a careful player reliably succeeding is exactly the property you want — the risk lives in the player choosing to push deep, captured by the commission / margin metrics. Model cautious-bot solvency asrequire: { clear_rate: ≥X }and leave it out ofbands. - Lesson 2 — a result that's FLAT across the whole space is a finding, not a failure.
When every config ties on a residual penalty (diver:
commissions_filled = 1for all 36 configs), the gap is not tuning-fixable — it is bot-skill- or structure-limited (here the competent bot can't grab sparse deep qualifying treasures, exactly the limitationplaytest-auditsays to report, not gate). Do NOT change constants to chase a bot-unreachable metric — that is the maximize-the-proxy mistake. Report it as a human-playtest item and move on; "no change warranted" is a legitimate, honest outcome of a balance pass.
- One sub-system at a time, each independently self-tested and committed. Don't batch five then debug the soup. Keep a playable game at every step.
- Grow the UI per system, reusing established chrome. Hand composited-screen
judgment to
visual-auditand correctness tovalidator.deepenowns systems & content — not pixels, not the gate mechanics. New screens aren't self-test-gated (the headless self-test never instantiates the scene tree), so gate them two other ways: a headless boot check (godot --headless --path … --quit-after N) that proves the router/view code parses and runs with noSCRIPT ERROR, plus a throwaway real-renderer harness (aSceneTreescript that builds the relevant state, instantiates the view, waits ~200 frames, saves a PNG) for a visual sanity glance. Delete the throwaway; keep the PNG as a probe-data artifact. - Verify the depth landed — INDEPENDENT design-depth audit (REQUIRED). The agent that
did the deepening cannot grade its own depth: it knows what it intended to add and reads
the diff as proof, so it ships "bigger" believing it shipped "deeper" (diver-0001's first
deepen passed its own assessment and was still owner-rejected as shallow). Fix it the way
visual-auditfixes the screen — with fresh, adversarial eyes, but pointed at the systems instead of the pixels. Dispatch a fresh subagent (no knowledge of what you set out to add — give it only the running game + the concept) to play/read it and answer, bluntly:- Is each progression vector a destination or a dial? For going deeper / levelling / scoring: what does the player DO at the top that they couldn't at the bottom? "Same action, more points" = FAIL (depth-as-multiplier).
- Do the new systems change decisions? Name a concrete moment the new mechanic/upgrade made the player choose differently. If none, it's inert.
- Is it more fun, or just more? One sentence: did this make the game deeper, or longer?
Treat a "just bigger / just longer" verdict as a failed pass — iterate (often the real
fix is a different axis: the auditor saying "the upgrades are meaningless because the core
loop is one decision" means you picked run-meta when the answer was systemic). Record the
auditor's verdict in
depth_pass.notes. Scale the audit to the change: one skeptic for a small content add, a fuller play-and-critique for a systemic/run-meta pass.
- Record + codify. Write
manifest.depth_pass(axis, systems added, new-assertion count, and the independent audit's verdict). Fold durable lessons back into this skill.
manifest.depth_pass
"depth_pass": {
"axis": "run-meta | systemic | content",
"systems_added": ["..."],
"selftest_assertions_added": 0,
"notes": "what changed, and any surfaced behavior-changes to previously-frozen logic"
}
Boundaries / non-goals
- Not a re-skin (
asset+visual-audit) and not the audio pass. - Does not invent a new status or touch the packaging gate.
- Does not redesign from scratch — it grows what exists along one axis.
Project gotchas (carry these)
- Headless
godot --scriptdoes NOT instantiate autoloads → data layers viapreload+static func. - Seed every RNG; Fisher–Yates, never
Array.shuffle(). - Reset
user://save.jsonbefore asserting on meta writes (stale-file false positives). - A growing
selftest.gdruns all stages in one function scope → give each stage's locals unique names (e.g. suffix with the stage number) or you get redeclaration parse errors as you append. - Avoid GDScript method names that collide with
Objectbuilt-ins (connect,draw,set, …) on your data/model classes — they parse-error or shadow silently. - A new
manifest.depth_passfield is not free: the manifest schema isadditionalProperties: false, so add the field toschema/manifest.schema.jsonand re-runnode tools/manifest.mjs validate <id>+ the vitest suite before committing.
Lessons from first use (run-layer dogfood)
- The "surface, don't swallow" rule earns its keep. Two changes touched
previously-frozen combat logic — threading run-persistent HP through
setup(), and giving a relic that had silently been a no-op a real effect. Both were named indepth_pass.notesrather than slipped in. When deepening forces a change to old behavior, that is normal — make it loud. - Prove the system headless first, wire the screen second. Every sub-system landed
its self-test assertion before any view existed, so the regression gate never
depended on rendering. This ordering is what let view work stay a separate, lower-risk
concern handed to
visual-audit. - Check the acquisition path, not just the hook. A hook that only fires at one moment (e.g. run-start) is dormant for anything acquired after that moment. When you add hook points, confirm the real in-game path that grants the thing actually triggers the hook — or record the limitation explicitly instead of shipping a dead feature.
Lessons from second use (shopkeep-0001 systemic dogfood — a "triage" loop)
The stated hook (a cashier triage: "who do I serve next?") was INERT, and it took THREE iterations under the independent audit to fix. Each round failed for a different reason the implementer couldn't see — the audit is what caught them. The durable lessons:
- A "choice" loop needs CONTENTION — a scarce resource the action itself consumes —
before any decision exists.
serve()was an instant, free, exact-match action on independent shelves, so "who next?" had a trivially optimal answer (serve whoever's about to leave) and was a reflex. Adding more options or values does NOT help while you can satisfy everyone. Decoupling value from urgency (tourist = cheap+impatient vs. regular = rich+patient) was completely inert until a serve-time cost (a register cooldown = throughput) made serving A literally spend a resource B needed. Diagnose missing contention FIRST when a "decision" loop feels flat; a serve/triage/allocation loop with no action budget, cooldown, or contested stock is a reflex no matter how many patron types you add. - A "destination" must add a VERB or gated content — a price multiplier is a dial in a costume, even pre-announced or flavored. Reputation "unlocking Regulars" who wanted the same demand items you'd already craft, paying 2×, was just a multiplier. It only became a real destination when Regulars placed pre-announced standing orders for top-tier goods you must deliberately pre-stock — i.e. it changed the CRAFT decision, a new verb.
- A reward/commitment with no penalty for ignoring it is optional flavor, not a decision. Standing orders changed nothing until an unfilled order cost reputation — only then did spending scarce materials to pre-stock them become a genuine bet.
- Structure can LAND while the verdict stays "conditional on tuning" — and that is the
STOP signal, not a cue to keep twiddling constants. After three iterations the audit went
from "inert" to "genuinely well-constructed decision … JUST BIGGER/LONGER, conditional on
tuning": the mechanism was right, but whether the contention BITES often enough is a
tuning question (serve time vs. patience fuses vs. queue size vs. spawn rate). That belongs
to a
playtest.gdbot +tools/balance.mjssearch + the human fun check, NOT to blind hand-tuning toward the proxy auditor (that is the maximize-the-proxy mistake). Recognize the boundary: once the structure is sound, stop iterating it and hand tuning to the balance pass. - The independent audit pays for itself every single round. Three rounds, three "just bigger" verdicts, three different real flaws pinpointed (dial-reputation → no-contention → toothless-orders). Self-assessment would have shipped after round 1 believing it was deep. Re-dispatch a FRESH subagent each iteration — a re-used one anchors on its prior read.
- A packaged game with no
depth_passand noplaytest.gdis the loud symptom of the original POC gap (polish shipped on an unverified-deep, unverified-winnable core). When you re-open one, expect the depth pass to also surface the missing winnability bot — record it as a required follow-up even if you don't build it in the same pass.
Lessons from the shopkeep-0001 BALANCE pass (building the playtest bot + running the search)
The required-next from the systemic pass — build playtest.gd, add the GF_TUNE seam, run
tools/balance.mjs — and what the bot actually found turned a "conditional on tuning"
verdict into a live decision. The durable lessons:
- A new contention mechanic can be silently MASKED by a co-located scarcity — measure WHICH
constraint binds before assuming your mechanic drives the decision. The serve-time cooldown
was completely inert at the shipped tuning, not because the cooldown was too short, but
because a different scarce resource bound first: the shop is shelf-stock-limited and the
"empty shelves end the day" rule is a soft landing that sells out before the register ever
saturates (
blocked_ticks = 0,soldout_days = every day). The fix needed dense arrivals (spawn_interval) to make the register the bottleneck. Instrument the candidate bottlenecks (a "register-busy-while-a-servable-patron-waits" counter, a "sold-out" flag) and confirm your new mechanic is the one that binds — a green winnability gate hides which constraint is live. forced_walkoutswas the WRONG oracle; prove "it's a real decision" by POLICY DIVERGENCE. The intuitive metric (a patron times out with their item still on a shelf) ~never fired, because the stockout ended the day first. The honest test: run TWO reasonable policies on the same seed (here value-first vs urgent-first cashier) and compare outcomes. Convergence (Δ≈0) = the choice is inert; bidirectional divergence (policy A wins some seeds, B wins others — no dominant policy) is the signature of a genuine tradeoff. This is stronger and more honest than any single bot's score, and it directly answers the audit's "do the new systems change decisions?" — empirically, by playing, not on paper.- Decision-pressure is often SEED-DEPENDENT — judge it in AGGREGATE (a banded mean), never as
a per-seed hard gate. Whether a given run pits value against urgency depends on the random
patron mix; ~half the seeds were inert even on a good config. Gating
no_trivial_dominanton per-seed divergence flaked the whole search to "0 configs survive." Fix: hard-floor the ROBUST winnability invariants per-seed (solvent / first-goal / no-death-spiral / excess- demand-exists), and put the intermittent decision-quality metric in a banded mean the search optimises across seeds. (Pairs with Lesson 1 from the diver dogfood: floors vs bands.) - A balance search can park a real decision on a fragile tuning KNIFE-EDGE — record that as a
limitation, not a win, and do NOT chase it with more tuning. The independent audit confirmed
the triage is real but only in
serve_time ≈ 2.5–3.5(it vanishes at 2.0), and that the sharpest tension turned out to be a strategic gold-now-vs-reputation-later axis, not the split-second patience triage the concept advertised. That fragility is honest signal for the HUMAN playtest and a possible future structural iteration (e.g. remove the sellout soft- landing so stock-scarcity stops masking the register) — not a cue to keep twiddling constants toward the proxy. Once the structure is proven real, the remaining "is it FUN / does the knife-edge feel good" is the human's call, full stop.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.