Agentic test update
Skill jpbaking/agentic-tests/skills/shared/agentic-test-update
Reconcile agent tests after the user INTENTIONALLY changed main-code behavior — show each failing agent test as old-locked vs new-actual behavior, and update tests only with per-diff user confirmation. Use when agent tests fail after a deliberate behavior change.From its SKILL.md
npx -y skills add jpbaking/agentic-tests --skill agentic-test-updateAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
2.4 KB, 522 tokens by cl100k_base, as published. Nobody here has run it
Agentic Test Update
Agent tests lock old behavior; the user changed behavior on purpose. This skill re-locks — but only where the user confirms the change was intended. Unconfirmed failures are potential regressions and must stay failing.
RULES
- NEVER edit main code or user tests. Only agent tests and plan files.
- Never update a test without the user's explicit per-diff confirmation. No batch "update all".
- An updated test must assert the NEW actual behavior (test-quality rules from
agentic-unit-testapply). Never weaken a test to pass both old and new behavior.
Step 1 — Collect failures
- Run the full agent-test suite.
- All green? Report "nothing to update" and STOP.
- For each failing test, capture: test name, file, expected (old locked behavior), actual (new behavior). Group by source file.
Step 2 — Confirm with the user, per behavior change
For each group, show a compact diff and ask ONE question:
src/pricing.ts — 3 agent tests failing
old locked: priceWithTax(0, 0.2) → 0
new actual: priceWithTax(0, 0.2) → throws RangeError
Intended change? (yes = update tests to lock new behavior / no = keep failing as regression)
Record every answer in agentic-test-plan.md under a ## Behavior updates <YYYY-MM-DD> section: confirmed or regression.
Step 3 — Update confirmed tests
For each confirmed group:
- Rewrite the failing assertions to lock the NEW behavior. Keep test names honest (rename if the name describes old behavior).
- Run the test file 3× (flakiness check). 3 attempts; after the 3rd failure revert the test and mark it
FAILED to update: <reason>in the plan. - Lint if configured.
Leave every regression test untouched and failing.
Step 4 — Report
- Updated: tests re-locked to new behavior (per file).
- Regressions: failing tests the user did NOT confirm — listed loudly; the suite is intentionally left red until the user fixes main code or re-runs this skill.
- Failed to update: 3-attempt casualties with reasons (or "none").
- Final suite status: green, or red-with-known-regressions.