Code health
Code quality dashboard: runs all available tools (type check, lint, tests, coverage, dead-code, bundle size), scores each category, presents a composite dashboard, and tracks trends in memory/metrics.db. HARD GATE: never fixes issues.From its SKILL.md
npx -y skills add asong56/skills --skill code-healthAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
2 things to look at
- 21 days oldThe repository was created 21 days ago. New is not bad, but a brand new repository carrying a familiar-sounding name is the shape a typosquat arrives in, and there has been no time for anyone else to find a problem with it.
- 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
19.4 KB, ~5.3k tokens by cl100k_base, as published. Nobody here has run it
Step 0: Gather project context
_ROOT=$(git rev-parse --show-toplevel 2>/dev/null || pwd)
_SLUG=$(basename "$_ROOT")
_BRANCH=$(git branch --show-current 2>/dev/null || echo "main")
_MEM="$_ROOT/memory"
mkdir -p "$_MEM" "$_MEM/sessions" "$_MEM/checkpoints" "$_MEM/retros" "$_MEM/reviews" "$_MEM/specs"
echo "=== Context: $_SLUG / $_BRANCH ==="
[ -f "$_MEM/context.md" ] && echo "--- last context ---" && tail -30 "$_MEM/context.md"
[ -f "$_MEM/learnings.jsonl" ] && echo "--- recent learnings ---" && tail -5 "$_MEM/learnings.jsonl"
[ -f "$_MEM/timeline.jsonl" ] && echo "--- recent timeline ---" && tail -5 "$_MEM/timeline.jsonl"
Memory dir (
memory/): replaces gbrain.grep -r "X" memory/≡gbrain search X·echo '...' >> memory/timeline.jsonl≡gbrain store
Metrics Storage (SQLite)
_DB="$_MEM/metrics.db"
sqlite3 "$_DB" "
CREATE TABLE IF NOT EXISTS health_scores (
ts TEXT, repo TEXT, category TEXT, score INTEGER, details TEXT
);
CREATE TABLE IF NOT EXISTS perf_benchmarks (
ts TEXT, repo TEXT, url TEXT, load_ms INTEGER, dom_ms INTEGER, notes TEXT
);" 2>/dev/null
/health -- Code Quality Dashboard
You are a Staff Engineer who owns the CI dashboard. You know that code quality isn't one metric -- it's a composite of type safety, lint cleanliness, test coverage, dead code, and script hygiene. Your job is to run every available tool, score the results, present a clear dashboard, and track trends so the team knows if quality is improving or slipping.
HARD GATE: Do NOT fix any issues. Produce the dashboard and recommendations only. The user decides what to act on.
User-invocable
When the user types /health, run this skill.
Step 1: Detect Health Stack
Read CLAUDE.md and look for a ## Health Stack section. If found, parse the tools
listed there and skip auto-detection.
If no ## Health Stack section exists, auto-detect available tools:
# Type checker
[ -f tsconfig.json ] && echo "TYPECHECK: tsc --noEmit"
# Linter
[ -f biome.json ] || [ -f biome.jsonc ] && echo "LINT: biome check ."
setopt +o nomatch 2>/dev/null || true
ls eslint.config.* .eslintrc.* .eslintrc 2>/dev/null | head -1 | xargs -I{} echo "LINT: eslint ."
[ -f .pylintrc ] || [ -f pyproject.toml ] && grep -q "pylint\|ruff" pyproject.toml 2>/dev/null && echo "LINT: ruff check ."
# Test runner
[ -f package.json ] && grep -q '"test"' package.json 2>/dev/null && echo "TEST: $(node -e "console.log(JSON.parse(require('fs').readFileSync('package.json','utf8')).scripts.test)" 2>/dev/null)"
[ -f pyproject.toml ] && grep -q "pytest" pyproject.toml 2>/dev/null && echo "TEST: pytest"
[ -f Cargo.toml ] && echo "TEST: cargo test"
[ -f go.mod ] && echo "TEST: go test ./..."
# Dead code
command -v knip >/dev/null 2>&1 && echo "DEADCODE: knip"
[ -f package.json ] && grep -q '"knip"' package.json 2>/dev/null && echo "DEADCODE: npx knip"
# Shell linting
command -v shellcheck >/dev/null 2>&1 && ls *.sh scripts/*.sh bin/*.sh 2>/dev/null | head -1 | xargs -I{} echo "SHELL: shellcheck"
# GBrain presence (D6) — only report as a dimension if gbrain is actually
# set up; otherwise skip so machines without gbrain aren't penalized.
if command -v gbrain >/dev/null 2>&1 && [ -f "$HOME/.gbrain/config.json" ]; then
echo "GBRAIN: gbrain doctor --json (wrapped in timeout 5s)"
fi
Use Glob to search for shell scripts:
**/*.sh(shell scripts in the repo)
After auto-detection, present the detected tools via AskUserQuestion:
"I detected these health check tools for this project:
- Type check:
tsc --noEmit - Lint:
biome check . - Tests:
bun test - Dead code:
knip - Shell lint:
shellcheck *.sh
A) Looks right -- persist to CLAUDE.md and continue B) I need to adjust some tools (tell me which) C) Skip persistence -- just run these"
If the user chooses A or B (after adjustments), append or update a ## Health Stack
section in CLAUDE.md:
## Health Stack
- typecheck: tsc --noEmit
- lint: biome check .
- test: bun test
- deadcode: knip
- shell: shellcheck *.sh scripts/*.sh
Step 2: Run Tools
Run each detected tool. For each tool:
- Record the start time
- Run the command, capturing both stdout and stderr
- Record the exit code
- Record the end time
- Capture the last 50 lines of output for the report
# Example for each tool — run each independently
START=$(date +%s)
tsc --noEmit 2>&1 | tail -50
EXIT_CODE=$?
END=$(date +%s)
echo "TOOL:typecheck EXIT:$EXIT_CODE DURATION:$((END-START))s"
Run tools sequentially (some may share resources or lock files). If a tool is not
installed or not found, record it as SKIPPED with reason, not as a failure.
Step 3: Score Each Category
Score each category on a 0-10 scale using this rubric:
| Category | Weight | 10 | 7 | 4 | 0 |
|---|---|---|---|---|---|
| Type check | 22% | Clean (exit 0) | <10 errors | <50 errors | >=50 errors |
| Lint | 18% | Clean (exit 0) | <5 warnings | <20 warnings | >=20 warnings |
| Tests | 28% | All pass (exit 0) | >95% pass | >80% pass | <=80% pass |
| Dead code | 13% | Clean (exit 0) | <5 unused exports | <20 unused | >=20 unused |
| Shell lint | 9% | Clean (exit 0) | <5 issues | >=5 issues | N/A (skip) |
| GBrain (D6) | 10% | doctor=ok, queue<10, pushed <24h | doctor=warnings OR queue<100 OR pushed <72h | doctor broken OR queue>=100 OR pushed >=72h | N/A (gbrain not installed) |
Parsing tool output for counts:
- tsc: Count lines matching
error TSin output. - biome/eslint/ruff: Count lines matching error/warning patterns. Parse the summary line if available.
- Tests: Parse pass/fail counts from the test runner output. If the runner only reports exit code, use: exit 0 = 10, exit non-zero = 4 (assume some failures).
- knip: Count lines reporting unused exports, files, or dependencies.
- shellcheck: Count distinct findings (lines starting with "In ... line").
Composite score:
composite = (typecheck_score * 0.22) + (lint_score * 0.18) + (test_score * 0.28) + (deadcode_score * 0.13) + (shell_score * 0.09) + (gbrain_score * 0.10)
If a category is skipped (tool not available — includes GBrain when gbrain is not installed), redistribute its weight proportionally among the remaining categories.
GBrain sub-score computation (D6):
doctor_component: 10 if `gbrain doctor --json | jq -r .status` == "ok";
7 if "warnings"; 0 otherwise (or command times out after 5s).
queue_component: 10 if $_MEM/.brain-queue.jsonl has <10 lines;
7 if 10-100; 0 if >=100 (suggests secret-scan rejections
piling up). N/A if artifacts_sync_mode == off.
push_component: 10 if (now - mtime of $_MEM/.brain-last-push) < 24h;
7 if <72h; 0 if >=72h. N/A if artifacts_sync_mode == off.
gbrain_score = 0.5 * doctor_component + 0.3 * queue_component + 0.2 * push_component
(redistribute 0.3 + 0.2 into doctor when sync_mode is off:
gbrain_score = doctor_component in that case)
The gbrain doctor --json call MUST be wrapped in timeout 5s so a hung
or misconfigured gbrain doesn't stall the entire /health dashboard.
Step 4: Present Dashboard
Present results as a clear table:
CODE HEALTH DASHBOARD
=====================
Project: <project name>
Branch: <current branch>
Date: <today>
Category Tool Score Status Duration Details
---------- ---------------- ----- -------- -------- -------
Type check tsc --noEmit 10/10 CLEAN 3s 0 errors
Lint biome check . 8/10 WARNING 2s 3 warnings
Tests bun test 10/10 CLEAN 12s 47/47 passed
Dead code knip 7/10 WARNING 5s 4 unused exports
Shell lint shellcheck 10/10 CLEAN 1s 0 issues
GBrain gbrain doctor 10/10 CLEAN <1s doctor=ok, queue=3, pushed 2h ago
COMPOSITE SCORE: 9.1 / 10
Duration: 23s total
Use these status labels:
- 10:
CLEAN - 7-9:
WARNING - 4-6:
NEEDS WORK - 0-3:
CRITICAL
If any category scored below 7, list the top issues from that tool's output:
DETAILS: Lint (3 warnings)
biome check . output:
src/utils.ts:42 — lint/complexity/noForEach: Prefer for...of
src/api.ts:18 — lint/style/useConst: Use const instead of let
src/api.ts:55 — lint/suspicious/noExplicitAny: Unexpected any
Step 5: Persist to Health History
eval "$(echo "$_SLUG" 2>/dev/null)" && mkdir -p $_MEM/projects/$SLUG
Append one JSONL line to $_MEM/health-history.jsonl:
{"ts":"2026-03-31T14:30:00Z","branch":"main","score":9.1,"typecheck":10,"lint":8,"test":10,"deadcode":7,"shell":10,"gbrain":10,"duration_s":23}
Fields:
ts-- ISO 8601 timestampbranch-- current git branchscore-- composite score (one decimal)typecheck,lint,test,deadcode,shell,gbrain-- individual category scores (integer 0-10)duration_s-- total time for all tools in seconds
If a category was skipped, set its value to null. Pre-D6 history entries
won't have a gbrain field — treat them as null for trend comparison
and start new tracking from the first post-D6 run.
Step 6: Trend Analysis + Recommendations
Read the last 10 entries from $_MEM/health-history.jsonl (if the
file exists and has prior entries).
eval "$(echo "$_SLUG" 2>/dev/null)" && mkdir -p $_MEM/projects/$SLUG
tail -10 $_MEM/health-history.jsonl 2>/dev/null || echo "NO_HISTORY"
If prior entries exist, show the trend:
HEALTH TREND (last 5 runs)
==========================
Date Branch Score TC Lint Test Dead Shell GBrain
---------- ----------- ----- -- ---- ---- ---- ----- ------
2026-03-28 main 9.4 10 9 10 8 10 10
2026-03-29 feat/auth 8.8 10 7 10 7 10 10
2026-03-30 feat/auth 8.2 10 6 9 7 10 7
2026-03-31 feat/auth 9.1 10 8 10 7 10 10
Trend: IMPROVING (+0.9 since last run)
If score dropped vs the previous run:
- Identify WHICH categories declined
- Show the delta for each declining category
- Correlate with tool output -- what specific errors/warnings appeared?
REGRESSIONS DETECTED
Lint: 9 -> 6 (-3) — 12 new biome warnings introduced
Most common: lint/complexity/noForEach (7 instances)
Tests: 10 -> 9 (-1) — 2 test failures
FAIL src/auth.test.ts > should validate token expiry
FAIL src/auth.test.ts > should reject malformed JWT
Health improvement suggestions (always show these):
Prioritize suggestions by impact (weight * score deficit):
RECOMMENDATIONS (by impact)
============================
1. [HIGH] Fix 2 failing tests (Tests: 9/10, weight 30%)
Run: bun test --verbose to see failures
2. [MED] Address 12 lint warnings (Lint: 6/10, weight 20%)
Run: biome check . --write to auto-fix
3. [LOW] Remove 4 unused exports (Dead code: 7/10, weight 15%)
Run: knip --fix to auto-remove
Rank by weight * (10 - score) descending. Only show categories below 10.
Important Rules
- Wrap, don't replace. Run the project's own tools. Never substitute your own analysis for what the tool reports.
- Read-only. Never fix issues. Present the dashboard and let the user decide.
- Respect CLAUDE.md. If
## Health Stackis configured, use those exact commands. Do not second-guess. - Skipped is not failed. If a tool isn't available, skip it gracefully and redistribute weight. Do not penalize the score.
- Show raw output for failures. When a tool reports errors, include the actual output (tail -50) so the user can act on it without re-running.
- Trends require history. On first run, say "First health check -- no trend data yet. Run /health again after making changes to track progress."
- Be honest about scores. A codebase with 100 type errors and all tests passing is not healthy. The composite score should reflect reality.
CodeScene MCP Mode (Real-Time Structural Health)
Code Health MCP (CodeScene)
Structural maintainability feedback for AI-assisted coding. Complements style/lint skills (coding-standards, WTQ-code-quality) with design-level health scores and regression gates.
Upstream: codescene-oss/codescene-mcp-server
Package: @codescene/codehealth-mcp (stdio via npx)
Security and boundaries
Opt-in (SKC): The codescene block in mcp-configs/mcp-servers.json is a template only. SKC plugin installs do not auto-enable bundled MCP servers. Copy the entry into your config only if you want it. You can exclude it during SKC install/sync with ECC_DISABLED_MCPS=codescene,....
Credentials: No bundled token. Set CS_ACCESS_TOKEN yourself (see getting-a-personal-access-token.md in the upstream repo). Never commit tokens to the repo.
What the tools read: When invoked, tools analyze files and git state in the local repository you point them at (paths you pass, plus branch context for analyze_change_set). They do not run by themselves. For standalone mode, follow upstream privacy docs: codescene-mcp-server README and CodeScene policies. Do not use this skill for secrets, credentials, or paths you do not want analyzed.
If the MCP is unavailable (offline, bad token, server crash): Do not invent Code Health scores. Tell the user the check was skipped. Continue only with explicit user approval. Prefer lint/tests/verification-loop for gating when MCP is down. Re-enable checks once the server connects.
When to Use
- User asks to review code quality, refactor a file, or check if AI changes degraded maintainability
- Before editing a hotspot, legacy module, or unfamiliar file
- Before commit or pull request when you need a maintainability safeguard
- After a large agent-written diff — verify Code Health did not regress
- Pair with
verification-loop,tdd-workflow, or/quality-gateas a structural check (not a replacement for tests/lint)
When to use
Same triggers as When to Use above — this heading is what SKC uses for skill auto-activation.
How It Works
1. Connect the MCP server
Copy the codescene entry from mcp-configs/mcp-servers.json into your framework MCP config.
Claude Code (~/.claude.json → mcpServers):
"codescene": {
"command": "npx",
"args": ["-y", "@codescene/codehealth-mcp"],
"env": {
"CS_ACCESS_TOKEN": "YOUR_CS_ACCESS_TOKEN_HERE"
}
}
Project-scoped: merge the same block into .mcp.json at the repo root.
Token setup is documented in the upstream repo (link above). Standalone mode does not require a paid CodeScene platform account for the four tools listed below. Restart the session and confirm the codescene server is connected before relying on scores.
2. Call standalone tools only
| Tool | When to use |
|---|---|
code_health_review | Full structural analysis before modifying a file |
code_health_score | Quick numeric score after each change (delta check) |
pre_commit_code_health_safeguard | Block commits that introduce Code Health regressions |
analyze_change_set | Branch-level check before opening a PR |
Do not call platform-only tools (e.g. repository-wide technical debt hotspot lists). Do not reference delta_analysis — not available on standalone.
3. Interpret scores (1–10)
| Range | Meaning | Agent behavior |
|---|---|---|
| 9.0–10.0 | Green — healthy | Safer to extend; still prefer vertical slices |
| 4.0–8.9 | Yellow — debt | Tread carefully; no drive-by refactors |
| 1.0–3.9 | Red — severe debt | Narrow scope only |
4. Run the feedback loop
Before touching a file
- Run
code_health_reviewon the target path. - Record baseline score and listed code smells.
- Plan the smallest change that addresses the task.
Scope by score: below 5 — minimal diff only; 5–7 — no broad refactors; above 7 — safer to refactor, still verify after each edit.
After each change
- Run
code_health_scoreon the same file. - Compare to the baseline from
code_health_review. - If the score regressed, fix before continuing. Never mark the task done while the score is lower than when you started.
Before every commit — run pre_commit_code_health_safeguard on the repository path.
Before a PR — run analyze_change_set against the base branch (e.g. main).
Examples
Example: Flask maintainability improvement
On pallets/flask, an agent loop using only standalone tools:
code_health_reviewon a target module (baseline 4.82)- Targeted refactor addressing listed smells
code_health_scoreafter each editpre_commit_code_health_safeguardbefore commitanalyze_change_setbefore PR
Result: Code Health 4.82 → 9.1 (free standalone token only).
Example: AGENTS.md enforcement block
Paste into the project AGENTS.md or CLAUDE.md:
## Code Health (CodeScene MCP)
Before modifying any file: run `code_health_review`, note score and issues.
- Score below 5: problematic range — scope changes narrowly.
- Score 5–7: warning range — no broad refactors.
After each change: run `code_health_score` to verify delta.
- If score regressed: fix before continuing; never declare done if score dropped.
Before every commit: run `pre_commit_code_health_safeguard`.
Before PR: run `analyze_change_set`.
Example: anti-patterns vs correct loop
# BAD: Edit first, check later
[large refactor without code_health_review]
# BAD: Ignore score drop
"Tests pass" → mark task done while Code Health decreased
# BAD: Broad refactor on red-score file (below 5)
Drive-by cleanup across the module
# GOOD: review → small change → score → commit safeguard → analyze_change_set
Pairing with SKC
| SKC skill / flow | Code Health MCP role |
|---|---|
coding-standards | Style/naming; Code Health = structure/complexity |
WTQ-code-quality | Write-time lint/format; Code Health = pre/post edit structural gate |
verification-loop / /quality-gate | Add structural regression check before "done" |
security-review | Security vs maintainability — use both when relevant |
tdd-workflow | Tests pass ≠ healthy design — check score after refactors |
Context tip: SKC recommends keeping MCP count low. Enable codescene when doing substantive edits; disable when not needed.
Related Skills
coding-standards— baseline conventionsWTQ-code-quality— write-time lint/format hooksverification-loop— build/test/lint gatetdd-workflow— test-first developmentsecurity-review— security checklistdocumentation-lookup— library docs via Context7 (orthogonal)
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.