agentsclimarketplace

Plan execution feedback

Skill lubochka/xiigen-mvp-engine/.agents/skills/plan-execution-feedback

Self-building AI code generation engine that generates application flows instead of implementing them. AGPL-3.0.From the repository description

Install
npx -y skills add lubochka/xiigen-mvp-engine --skill plan-execution-feedback

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • 19 days oldThe repository was created 19 days ago. New is not bad, but a brand new repository carrying a familiar-sounding name is the shape a typosquat arrives in, and there has been no time for anyone else to find a problem with it.
  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

4.4 KB, ~1.0k tokens by cl100k_base, as published. Nobody here has run it

Plan Execution Feedback Skill

After each session, compares planned vs actual (write-time fixes, test counts, API assumptions). Tracks discovery rate over time. When discovery rate consistently exceeds 120%, flags the planning process as under-investing in G0 infrastructure discovery.

When to Invoke

  • At session end, after Section A is written
  • Before writing the next session's plan (to calibrate G0 depth)

Metrics

WF discovery rate = (WF actual) / (WF planned) × 100%
Test delivery rate = (tests delivered) / (tests planned) × 100%

Interpretation Table

WF Discovery RateSignalRequired Action for Next Plan
≤ 120%NormalProceed with standard G0
121–200%G0 under-investedAdd explicit api-shape-verification table to next plan
> 200%API assumptions are guessesDo NOT write plan without live source reads first

XIIGen Execution History

SessionWF PlannedWF ActualDiscovery RateSignal
S100Normal
S205∞%G0 missed API shapes entirely
S347175%G0 under-invested
S47?TBDMust complete api-shape-verification table at G0

Pattern: Sessions 2 and 3 both had high discovery rates. This means the planning process consistently underestimates API mismatches. Root cause: plan authors write tests against assumed APIs without reading source files first.

Rules

  1. Record WF planned vs actual in STATE.json after every session
  2. If discovery rate > 120% for two consecutive sessions: the next session plan MUST include a completed api-shape-verification table before any code blocks are written
  3. If test delivery rate < 80%: investigate whether the plan's test count was based on counting describe blocks vs it() blocks (common source of over-counting)
  4. Report the discovery rate metric as part of Section A: WF planned: N, WF actual: M, discovery rate: M/N×100%

Integration

  • Invoked by agent-constitution at Session End (after Section A)
  • Feeds into planning-skill G0 depth calibration for the next session
  • Results stored in STATE.json under phases.N.wf_discovery_rate

UNIVERSAL STANDARD ADDENDUM — Human-Readable Log Gate + User Prompt Source update (ported from llm_mvp_core)

Added by the Universal-Skills refresh (UpdateUniversalSkills). The 120%/200% thresholds above are kept. The universal layer adds a readability gate on the feedback record and a per-phase update to the User Prompt Source Ledger, so the metric is never machine-only.

Human-Readable Plan/Execution Log Gate (Section A is not just machine rows)

A per-phase feedback record FAILS this gate if a human cannot tell from it: what the phase did, what is done, what remains, which claims stayed closed and why, and the next bounded action. Numbers (discovery rate, test delivery rate) are secondary evidence — they appear AFTER a plain-language line.

Phase N of M: <human-readable name>. Done: <…>. Remaining: <…>.
Discovery rate: WF planned N, WF actual M, M/N×100%.   # machine row AFTER the human line
Test delivery: delivered/planned ×100%.
Next action: <next bounded step, or NONE>.

A discovery rate / test delivery row with no preceding human sentence is a machine-only record and does not satisfy Section A.

User Prompt Source Ledger — update each phase

After each phase, append to the User Prompt Source Ledger any requirement that was discovered or strengthened at write-time (a wrong assumption about a NestJS service shape, a missing React/FastAPI contract, a DTO mismatch). A write-time fix that revealed a new requirement is recorded as its source so the next plan's Gate-0 discovery covers it. This closes the loop between plan-execution-feedback and planning-session-startup (Authority/Source/Delta ledgers).

Discovery-rate over-investment signal (>120% restated as a learning signal)

WF discovery rate > 120% for two consecutive phases is not just "add an api-shape-verification table". It is evidence that Gate-0 discovery under-read live TS source (server/src/**.ts, server/src/engine-contracts/*.ts). The next plan MUST read those sources live before writing test code (Jest/Playwright), and the feedback record must say so in plain language.

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Gives 0 of the 12 instructions most plan spec skills give in ~1.0k tokens

Counted across 1,099 of the 1,860 authors here whose files we hold, read 2026-08-07

  • Ask one question at a timein 51 of 1099
  • Break plans into vertical slicesin 29 of 1099, across 11 files
  • Publish issues in dependency orderin 27 of 1099, across 9 files
  • Iterate until user approves the breakdownin 25 of 1099, across 7 files
  • Explore the repository to understand the codebase statein 24 of 1099, across 7 files
  • Use domain glossary vocabularyin 23 of 1099, across 5 files
  • Apply correct triage labels to published issuesin 23 of 1099, across 5 files
  • Prefer AFK slices over HITLin 22 of 1099, across 7 files
  • Write a specification before writing any codein 22 of 1099, across 14 files
  • Write failing tests before implementation codein 22 of 1099, across 20 files
  • Ask clarifying questions until requirements are concretein 21 of 1099, across 13 files
  • Respect existing architecture decision recordsin 20 of 1099, across 5 files

Said here and by no other author read

  • calculate write-time discovery rate after every session
  • calculate test delivery rate after every session
  • report metrics with a preceding plain-language summary
  • append uncovered write-time requirements to the source ledger
  • include a completed api-shape-verification table if discovery rate exceeds 120 percent
  • read live source files before writing test code if discovery rate exceeds 120 percent

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 326,696. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.