agentsclimarketplace

Code health

Skill asong56/skills/04-assure/code-health

Code quality dashboard: runs all available tools (type check, lint, tests, coverage, dead-code, bundle size), scores each category, presents a composite dashboard, and tracks trends in memory/metrics.db. HARD GATE: never fixes issues.From its SKILL.md

Install
npx -y skills add asong56/skills --skill code-health

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • 21 days oldThe repository was created 21 days ago. New is not bad, but a brand new repository carrying a familiar-sounding name is the shape a typosquat arrives in, and there has been no time for anyone else to find a problem with it.
  • 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

19.4 KB, ~5.3k tokens by cl100k_base, as published. Nobody here has run it

Step 0: Gather project context

_ROOT=$(git rev-parse --show-toplevel 2>/dev/null || pwd)
_SLUG=$(basename "$_ROOT")
_BRANCH=$(git branch --show-current 2>/dev/null || echo "main")
_MEM="$_ROOT/memory"
mkdir -p "$_MEM" "$_MEM/sessions" "$_MEM/checkpoints" "$_MEM/retros" "$_MEM/reviews" "$_MEM/specs"
echo "=== Context: $_SLUG / $_BRANCH ==="
[ -f "$_MEM/context.md" ]      && echo "--- last context ---"     && tail -30 "$_MEM/context.md"
[ -f "$_MEM/learnings.jsonl" ] && echo "--- recent learnings ---" && tail -5  "$_MEM/learnings.jsonl"
[ -f "$_MEM/timeline.jsonl" ]  && echo "--- recent timeline ---"  && tail -5  "$_MEM/timeline.jsonl"

Memory dir (memory/): replaces gbrain. grep -r "X" memory/gbrain search X · echo '...' >> memory/timeline.jsonlgbrain store

Metrics Storage (SQLite)

_DB="$_MEM/metrics.db"
sqlite3 "$_DB" "
CREATE TABLE IF NOT EXISTS health_scores (
  ts TEXT, repo TEXT, category TEXT, score INTEGER, details TEXT
);
CREATE TABLE IF NOT EXISTS perf_benchmarks (
  ts TEXT, repo TEXT, url TEXT, load_ms INTEGER, dom_ms INTEGER, notes TEXT
);" 2>/dev/null

/health -- Code Quality Dashboard

You are a Staff Engineer who owns the CI dashboard. You know that code quality isn't one metric -- it's a composite of type safety, lint cleanliness, test coverage, dead code, and script hygiene. Your job is to run every available tool, score the results, present a clear dashboard, and track trends so the team knows if quality is improving or slipping.

HARD GATE: Do NOT fix any issues. Produce the dashboard and recommendations only. The user decides what to act on.

User-invocable

When the user types /health, run this skill.


Step 1: Detect Health Stack

Read CLAUDE.md and look for a ## Health Stack section. If found, parse the tools listed there and skip auto-detection.

If no ## Health Stack section exists, auto-detect available tools:

# Type checker
[ -f tsconfig.json ] && echo "TYPECHECK: tsc --noEmit"

# Linter
[ -f biome.json ] || [ -f biome.jsonc ] && echo "LINT: biome check ."
setopt +o nomatch 2>/dev/null || true
ls eslint.config.* .eslintrc.* .eslintrc 2>/dev/null | head -1 | xargs -I{} echo "LINT: eslint ."
[ -f .pylintrc ] || [ -f pyproject.toml ] && grep -q "pylint\|ruff" pyproject.toml 2>/dev/null && echo "LINT: ruff check ."

# Test runner
[ -f package.json ] && grep -q '"test"' package.json 2>/dev/null && echo "TEST: $(node -e "console.log(JSON.parse(require('fs').readFileSync('package.json','utf8')).scripts.test)" 2>/dev/null)"
[ -f pyproject.toml ] && grep -q "pytest" pyproject.toml 2>/dev/null && echo "TEST: pytest"
[ -f Cargo.toml ] && echo "TEST: cargo test"
[ -f go.mod ] && echo "TEST: go test ./..."

# Dead code
command -v knip >/dev/null 2>&1 && echo "DEADCODE: knip"
[ -f package.json ] && grep -q '"knip"' package.json 2>/dev/null && echo "DEADCODE: npx knip"

# Shell linting
command -v shellcheck >/dev/null 2>&1 && ls *.sh scripts/*.sh bin/*.sh 2>/dev/null | head -1 | xargs -I{} echo "SHELL: shellcheck"

# GBrain presence (D6) — only report as a dimension if gbrain is actually
# set up; otherwise skip so machines without gbrain aren't penalized.
if command -v gbrain >/dev/null 2>&1 && [ -f "$HOME/.gbrain/config.json" ]; then
  echo "GBRAIN: gbrain doctor --json (wrapped in timeout 5s)"
fi

Use Glob to search for shell scripts:

  • **/*.sh (shell scripts in the repo)

After auto-detection, present the detected tools via AskUserQuestion:

"I detected these health check tools for this project:

  • Type check: tsc --noEmit
  • Lint: biome check .
  • Tests: bun test
  • Dead code: knip
  • Shell lint: shellcheck *.sh

A) Looks right -- persist to CLAUDE.md and continue B) I need to adjust some tools (tell me which) C) Skip persistence -- just run these"

If the user chooses A or B (after adjustments), append or update a ## Health Stack section in CLAUDE.md:

## Health Stack

- typecheck: tsc --noEmit
- lint: biome check .
- test: bun test
- deadcode: knip
- shell: shellcheck *.sh scripts/*.sh

Step 2: Run Tools

Run each detected tool. For each tool:

  1. Record the start time
  2. Run the command, capturing both stdout and stderr
  3. Record the exit code
  4. Record the end time
  5. Capture the last 50 lines of output for the report
# Example for each tool — run each independently
START=$(date +%s)
tsc --noEmit 2>&1 | tail -50
EXIT_CODE=$?
END=$(date +%s)
echo "TOOL:typecheck EXIT:$EXIT_CODE DURATION:$((END-START))s"

Run tools sequentially (some may share resources or lock files). If a tool is not installed or not found, record it as SKIPPED with reason, not as a failure.


Step 3: Score Each Category

Score each category on a 0-10 scale using this rubric:

CategoryWeight10740
Type check22%Clean (exit 0)<10 errors<50 errors>=50 errors
Lint18%Clean (exit 0)<5 warnings<20 warnings>=20 warnings
Tests28%All pass (exit 0)>95% pass>80% pass<=80% pass
Dead code13%Clean (exit 0)<5 unused exports<20 unused>=20 unused
Shell lint9%Clean (exit 0)<5 issues>=5 issuesN/A (skip)
GBrain (D6)10%doctor=ok, queue<10, pushed <24hdoctor=warnings OR queue<100 OR pushed <72hdoctor broken OR queue>=100 OR pushed >=72hN/A (gbrain not installed)

Parsing tool output for counts:

  • tsc: Count lines matching error TS in output.
  • biome/eslint/ruff: Count lines matching error/warning patterns. Parse the summary line if available.
  • Tests: Parse pass/fail counts from the test runner output. If the runner only reports exit code, use: exit 0 = 10, exit non-zero = 4 (assume some failures).
  • knip: Count lines reporting unused exports, files, or dependencies.
  • shellcheck: Count distinct findings (lines starting with "In ... line").

Composite score:

composite = (typecheck_score * 0.22) + (lint_score * 0.18) + (test_score * 0.28) + (deadcode_score * 0.13) + (shell_score * 0.09) + (gbrain_score * 0.10)

If a category is skipped (tool not available — includes GBrain when gbrain is not installed), redistribute its weight proportionally among the remaining categories.

GBrain sub-score computation (D6):

doctor_component: 10 if `gbrain doctor --json | jq -r .status` == "ok";
                   7 if "warnings"; 0 otherwise (or command times out after 5s).
queue_component:   10 if $_MEM/.brain-queue.jsonl has <10 lines;
                    7 if 10-100; 0 if >=100 (suggests secret-scan rejections
                    piling up). N/A if artifacts_sync_mode == off.
push_component:    10 if (now - mtime of $_MEM/.brain-last-push) < 24h;
                    7 if <72h; 0 if >=72h. N/A if artifacts_sync_mode == off.
gbrain_score     = 0.5 * doctor_component + 0.3 * queue_component + 0.2 * push_component
                   (redistribute 0.3 + 0.2 into doctor when sync_mode is off:
                   gbrain_score = doctor_component in that case)

The gbrain doctor --json call MUST be wrapped in timeout 5s so a hung or misconfigured gbrain doesn't stall the entire /health dashboard.


Step 4: Present Dashboard

Present results as a clear table:

CODE HEALTH DASHBOARD
=====================

Project: <project name>
Branch:  <current branch>
Date:    <today>

Category      Tool              Score   Status     Duration   Details
----------    ----------------  -----   --------   --------   -------
Type check    tsc --noEmit      10/10   CLEAN      3s         0 errors
Lint          biome check .      8/10   WARNING    2s         3 warnings
Tests         bun test          10/10   CLEAN      12s        47/47 passed
Dead code     knip               7/10   WARNING    5s         4 unused exports
Shell lint    shellcheck        10/10   CLEAN      1s         0 issues
GBrain        gbrain doctor     10/10   CLEAN      <1s        doctor=ok, queue=3, pushed 2h ago

COMPOSITE SCORE: 9.1 / 10

Duration: 23s total

Use these status labels:

  • 10: CLEAN
  • 7-9: WARNING
  • 4-6: NEEDS WORK
  • 0-3: CRITICAL

If any category scored below 7, list the top issues from that tool's output:

DETAILS: Lint (3 warnings)
  biome check . output:
    src/utils.ts:42 — lint/complexity/noForEach: Prefer for...of
    src/api.ts:18 — lint/style/useConst: Use const instead of let
    src/api.ts:55 — lint/suspicious/noExplicitAny: Unexpected any

Step 5: Persist to Health History

eval "$(echo "$_SLUG" 2>/dev/null)" && mkdir -p $_MEM/projects/$SLUG

Append one JSONL line to $_MEM/health-history.jsonl:

{"ts":"2026-03-31T14:30:00Z","branch":"main","score":9.1,"typecheck":10,"lint":8,"test":10,"deadcode":7,"shell":10,"gbrain":10,"duration_s":23}

Fields:

  • ts -- ISO 8601 timestamp
  • branch -- current git branch
  • score -- composite score (one decimal)
  • typecheck, lint, test, deadcode, shell, gbrain -- individual category scores (integer 0-10)
  • duration_s -- total time for all tools in seconds

If a category was skipped, set its value to null. Pre-D6 history entries won't have a gbrain field — treat them as null for trend comparison and start new tracking from the first post-D6 run.


Step 6: Trend Analysis + Recommendations

Read the last 10 entries from $_MEM/health-history.jsonl (if the file exists and has prior entries).

eval "$(echo "$_SLUG" 2>/dev/null)" && mkdir -p $_MEM/projects/$SLUG
tail -10 $_MEM/health-history.jsonl 2>/dev/null || echo "NO_HISTORY"

If prior entries exist, show the trend:

HEALTH TREND (last 5 runs)
==========================
Date          Branch         Score   TC   Lint  Test  Dead  Shell  GBrain
----------    -----------    -----   --   ----  ----  ----  -----  ------
2026-03-28    main           9.4     10   9     10    8     10     10
2026-03-29    feat/auth      8.8     10   7     10    7     10     10
2026-03-30    feat/auth      8.2     10   6     9     7     10      7
2026-03-31    feat/auth      9.1     10   8     10    7     10     10

Trend: IMPROVING (+0.9 since last run)

If score dropped vs the previous run:

  1. Identify WHICH categories declined
  2. Show the delta for each declining category
  3. Correlate with tool output -- what specific errors/warnings appeared?
REGRESSIONS DETECTED
  Lint: 9 -> 6 (-3) — 12 new biome warnings introduced
    Most common: lint/complexity/noForEach (7 instances)
  Tests: 10 -> 9 (-1) — 2 test failures
    FAIL src/auth.test.ts > should validate token expiry
    FAIL src/auth.test.ts > should reject malformed JWT

Health improvement suggestions (always show these):

Prioritize suggestions by impact (weight * score deficit):

RECOMMENDATIONS (by impact)
============================
1. [HIGH]  Fix 2 failing tests (Tests: 9/10, weight 30%)
   Run: bun test --verbose to see failures
2. [MED]   Address 12 lint warnings (Lint: 6/10, weight 20%)
   Run: biome check . --write to auto-fix
3. [LOW]   Remove 4 unused exports (Dead code: 7/10, weight 15%)
   Run: knip --fix to auto-remove

Rank by weight * (10 - score) descending. Only show categories below 10.


Important Rules

  1. Wrap, don't replace. Run the project's own tools. Never substitute your own analysis for what the tool reports.
  2. Read-only. Never fix issues. Present the dashboard and let the user decide.
  3. Respect CLAUDE.md. If ## Health Stack is configured, use those exact commands. Do not second-guess.
  4. Skipped is not failed. If a tool isn't available, skip it gracefully and redistribute weight. Do not penalize the score.
  5. Show raw output for failures. When a tool reports errors, include the actual output (tail -50) so the user can act on it without re-running.
  6. Trends require history. On first run, say "First health check -- no trend data yet. Run /health again after making changes to track progress."
  7. Be honest about scores. A codebase with 100 type errors and all tests passing is not healthy. The composite score should reflect reality.

CodeScene MCP Mode (Real-Time Structural Health)

Code Health MCP (CodeScene)

Structural maintainability feedback for AI-assisted coding. Complements style/lint skills (coding-standards, WTQ-code-quality) with design-level health scores and regression gates.

Upstream: codescene-oss/codescene-mcp-server Package: @codescene/codehealth-mcp (stdio via npx)

Security and boundaries

Opt-in (SKC): The codescene block in mcp-configs/mcp-servers.json is a template only. SKC plugin installs do not auto-enable bundled MCP servers. Copy the entry into your config only if you want it. You can exclude it during SKC install/sync with ECC_DISABLED_MCPS=codescene,....

Credentials: No bundled token. Set CS_ACCESS_TOKEN yourself (see getting-a-personal-access-token.md in the upstream repo). Never commit tokens to the repo.

What the tools read: When invoked, tools analyze files and git state in the local repository you point them at (paths you pass, plus branch context for analyze_change_set). They do not run by themselves. For standalone mode, follow upstream privacy docs: codescene-mcp-server README and CodeScene policies. Do not use this skill for secrets, credentials, or paths you do not want analyzed.

If the MCP is unavailable (offline, bad token, server crash): Do not invent Code Health scores. Tell the user the check was skipped. Continue only with explicit user approval. Prefer lint/tests/verification-loop for gating when MCP is down. Re-enable checks once the server connects.

When to Use

  • User asks to review code quality, refactor a file, or check if AI changes degraded maintainability
  • Before editing a hotspot, legacy module, or unfamiliar file
  • Before commit or pull request when you need a maintainability safeguard
  • After a large agent-written diff — verify Code Health did not regress
  • Pair with verification-loop, tdd-workflow, or /quality-gate as a structural check (not a replacement for tests/lint)

When to use

Same triggers as When to Use above — this heading is what SKC uses for skill auto-activation.

How It Works

1. Connect the MCP server

Copy the codescene entry from mcp-configs/mcp-servers.json into your framework MCP config.

Claude Code (~/.claude.jsonmcpServers):

"codescene": {
  "command": "npx",
  "args": ["-y", "@codescene/codehealth-mcp"],
  "env": {
    "CS_ACCESS_TOKEN": "YOUR_CS_ACCESS_TOKEN_HERE"
  }
}

Project-scoped: merge the same block into .mcp.json at the repo root.

Token setup is documented in the upstream repo (link above). Standalone mode does not require a paid CodeScene platform account for the four tools listed below. Restart the session and confirm the codescene server is connected before relying on scores.

2. Call standalone tools only

ToolWhen to use
code_health_reviewFull structural analysis before modifying a file
code_health_scoreQuick numeric score after each change (delta check)
pre_commit_code_health_safeguardBlock commits that introduce Code Health regressions
analyze_change_setBranch-level check before opening a PR

Do not call platform-only tools (e.g. repository-wide technical debt hotspot lists). Do not reference delta_analysis — not available on standalone.

3. Interpret scores (1–10)

RangeMeaningAgent behavior
9.0–10.0Green — healthySafer to extend; still prefer vertical slices
4.0–8.9Yellow — debtTread carefully; no drive-by refactors
1.0–3.9Red — severe debtNarrow scope only

4. Run the feedback loop

Before touching a file

  1. Run code_health_review on the target path.
  2. Record baseline score and listed code smells.
  3. Plan the smallest change that addresses the task.

Scope by score: below 5 — minimal diff only; 5–7 — no broad refactors; above 7 — safer to refactor, still verify after each edit.

After each change

  1. Run code_health_score on the same file.
  2. Compare to the baseline from code_health_review.
  3. If the score regressed, fix before continuing. Never mark the task done while the score is lower than when you started.

Before every commit — run pre_commit_code_health_safeguard on the repository path.

Before a PR — run analyze_change_set against the base branch (e.g. main).

Examples

Example: Flask maintainability improvement

On pallets/flask, an agent loop using only standalone tools:

  1. code_health_review on a target module (baseline 4.82)
  2. Targeted refactor addressing listed smells
  3. code_health_score after each edit
  4. pre_commit_code_health_safeguard before commit
  5. analyze_change_set before PR

Result: Code Health 4.82 → 9.1 (free standalone token only).

Example: AGENTS.md enforcement block

Paste into the project AGENTS.md or CLAUDE.md:

## Code Health (CodeScene MCP)

Before modifying any file: run `code_health_review`, note score and issues.

- Score below 5: problematic range — scope changes narrowly.
- Score 5–7: warning range — no broad refactors.

After each change: run `code_health_score` to verify delta.

- If score regressed: fix before continuing; never declare done if score dropped.

Before every commit: run `pre_commit_code_health_safeguard`.

Before PR: run `analyze_change_set`.

Example: anti-patterns vs correct loop

# BAD: Edit first, check later
[large refactor without code_health_review]

# BAD: Ignore score drop
"Tests pass" → mark task done while Code Health decreased

# BAD: Broad refactor on red-score file (below 5)
Drive-by cleanup across the module

# GOOD: review → small change → score → commit safeguard → analyze_change_set

Pairing with SKC

SKC skill / flowCode Health MCP role
coding-standardsStyle/naming; Code Health = structure/complexity
WTQ-code-qualityWrite-time lint/format; Code Health = pre/post edit structural gate
verification-loop / /quality-gateAdd structural regression check before "done"
security-reviewSecurity vs maintainability — use both when relevant
tdd-workflowTests pass ≠ healthy design — check score after refactors

Context tip: SKC recommends keeping MCP count low. Enable codescene when doing substantive edits; disable when not needed.

Related Skills

  • coding-standards — baseline conventions
  • WTQ-code-quality — write-time lint/format hooks
  • verification-loop — build/test/lint gate
  • tdd-workflow — test-first development
  • security-review — security checklist
  • documentation-lookup — library docs via Context7 (orthogonal)

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 326,851. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.