agentsclimarketplace

Ai regression testing

Skill yeaight7/agent-powerups/skills/ai-regression-testing

Curated power-ups for coding agents: skills, slash commands, MCP configs, hooks, AGENTS.md templates, and workflows for serious software engineering. Claude Code, Codex, Antigravity CLI, Cursor and more

Install
npx -y skills add yeaight7/agent-powerups --skill ai-regression-testing

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 6 stars6 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Use when an agent has modified logic, API routes, or data-transformation code, a bug was just fixed and must not be reintroduced, or behavior must stay identical across execution paths such as sandbox versus production.

SKILL.md

4.1 KB, as published. Nobody here has run it

AI Regression Testing

When an agent writes code and then reviews it, it carries the same assumptions into both steps. Automated tests break this cycle.

When to Use

  • An agent has modified logic, API routes, or data transformation code
  • A bug was found — need to prevent re-introduction
  • Running /bug-check after a change session
  • Multiple execution paths exist (feature flags, sandbox vs production, env variants)

The Core Problem

Agent writes fix → Agent reviews fix → Agent says "looks correct" → Bug still present

The most common blind spot: an agent fixes the production path but leaves the sandbox/mock path unchanged, or vice versa.

Workflow

Run in order. Do not skip to agent review if automated steps fail.

Step 1 — Run Tests (mandatory)

npm test      # or: pytest, cargo test, go test ./...
npm run build # TypeScript build / type check
  • Test fail → highest priority; fix before anything else
  • Build fail → report type errors as highest priority
  • Both pass → continue to Step 2

Step 2 — Agent Code Review

With tests passing, do a focused review for patterns agents commonly miss:

  1. Execution path parity: Do all code paths (sandbox, production, feature-flag on/off) return the same response shape?
  2. Query completeness: Are all fields used in the response present in the query or selection?
  3. Error state cleanup: On error, is stale state cleared before the error is surfaced?
  4. Optimistic update rollback: If an API call fails, is the optimistic UI change reverted?

Step 3 — Write a Regression Test for Each Bug Fixed

For every bug found and fixed, add a test immediately:

Bug: <description>
File: <path>
Regression test: <test name and what it asserts>

If you cannot write a test, document why:

Bug: <description>
Regression test: DEFERRED — <reason> (e.g., requires E2E harness not yet in place)

Do not silently skip. Every real bug should either have a test or an explicit deferral note.

Writing Effective Regression Tests

Test the contract, not the implementation:

// Test what the consumer receives, not how it's computed
const REQUIRED_RESPONSE_FIELDS = ["id", "email", "settings", "created_at"];

it("profile endpoint returns all required fields", async () => {
  const res = await GET(createRequest("/api/user/profile"));
  const json = await res.json();
  for (const field of REQUIRED_RESPONSE_FIELDS) {
    expect(json.data).toHaveProperty(field);
  }
});

Name tests after the bug category, not the fix:

it("sandbox path returns same field set as production path (BUG-CLASS: path-parity)")
it("notification_settings is not undefined after SELECT * removal (regression)")

Common AI Regression Patterns

PatternCheckPriority
Execution path paritySame response shape across all pathsHigh
Query field omissionAll response fields present in DB queryHigh
Error state leakageState cleared before error is returnedMedium
Missing rollbackPrevious state restored on API failureMedium

Strategy

Do not aim for coverage percentage. Write tests only for bugs that were found. Bug clusters naturally: if three bugs appeared in /api/user/profile, that endpoint needs tests. An endpoint that has never had a bug does not need tests yet.

Tests added this way grow organically with the bug history and cannot be gamed by coverage metrics.

Verification

  • Test suite and build/type check were run and pass — before any agent review started
  • The four review checks were applied: path parity, query completeness, error-state cleanup, optimistic-update rollback
  • Every bug fixed has a regression test, or an explicit DEFERRED note with the reason
  • New tests assert the contract (response shape, required fields), not the implementation

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.