Shakedown
Skill belousov-petr/shakedown
Full-stack review of any project - multi-agent systems, pipelines, codebases, or applications. Dynamically discovers project structure, reads all config/code/data, queries databases, tests backup integrity, and analyzes architecture, reliability, efficiency, and security. Produces ranked actionable recommendations. Use when asked to review a project, do a project health check, stress-test a codebase, do a gap analysis, or assess technical debt. Also when user asks 'how solid is this project', 'what is missing', 'find the weak spots', 'what would break first', 'where does this need tightening', 'what's wrong with this project', or 'how mature is this project'. Do NOT activate for simple code reviews, PR reviews, or single-file analysis.From its SKILL.md
npx -y skills add belousov-petr/shakedownAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its file declares
Copied from the file, not written here
The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.
SKILL.md
24.6 KB, ~5.7k tokens by cl100k_base, as published. Nobody here has run it
Shakedown
Comprehensive review of any project. Discovers the structure dynamically, reads everything, challenges every layer, finds bottlenecks, blind spots, and waste, then produces ranked recommendations you can implement in the same session.
Works on: multi-agent platforms (Paperclip, CrewAI, AutoGen), data pipelines, web applications, CLI tools, monorepos, microservices - any project with files to read and architecture to challenge.
Phase 1: Discover Project Structure
Before reading anything, map the project. Do NOT assume folder names, frameworks, or conventions - discover them.
1a. Identify What Kind of Project This Is
List files in the working directory and its immediate subdirectories to understand the project layout. Use your platform's tools (Glob, LS, or shell commands) to discover:
# Example commands (adapt to your platform):
ls -la
find . -maxdepth 2 -type f | head -50
Check for framework indicators by looking for common config files:
package.json, pyproject.toml, Cargo.toml, go.mod, Makefile,
docker-compose.yml, .env, *.config.*
Check for agent/pipeline platforms by looking for platform-specific
directories (.paperclip, .crew, .autogen, .langchain, etc.)
Check for agent skill indicators:
SKILL.mdin the project root -> this project IS an agent skill.agents/skills/directory -> project contains or uses skills- Skill-like YAML frontmatter (
name:,description:fields) - References to skill platforms (Claude Code, Copilot, Gemini CLI, Codex)
If agent skill detected, set a flag. This triggers section 4.8 (Agent Skill Standards Compliance) during Phase 4.
Check for running databases by listing active network listeners on common database ports (5432, 3306, 27017, 6379).
From this, determine:
- Project type: multi-agent system, web app, data pipeline, CLI tool, library, etc.
- Tech stack: languages, frameworks, platforms
- Data stores: databases, file-based storage, caches
- External dependencies: APIs, services, scheduled jobs
1b. Map the Full Directory Tree
List all directories up to 4 levels deep to understand the project
structure. Then find all config, code, and documentation files by
searching for common extensions: .md, .json, .yaml, .yml,
.py, .js, .ts, .toml, .env, .sh
# Example:
find {projectRoot} -maxdepth 4 -type d | head -100
Identify all:
- Config files: how is the system configured?
- Instruction/prompt files: agent instructions, system prompts, templates
- Code files: source, scripts, utilities
- Output directories: where does the system write its results?
- Data directories: storage, pipelines, caches, logs
- Documentation: READMEs, design docs, ADRs, runbooks
1c. Check Git History
Review recent git history to assess project health:
git log --oneline -20 # recent commits
git log --oneline --since="30 days ago" | wc -l # activity level
git shortlog -sn --since="6 months ago" # contributors
git branch -a # branch hygiene
git stash list # abandoned work?
Determine:
- Is the project actively developed or stalled?
- Single developer or team?
- Are commits atomic and well-described, or sprawling "fix stuff" patches?
- Any long-lived branches that suggest unfinished features?
1d. Read the Project's Stated Goal
Look for and read (if they exist):
- README.md - what does the project claim to do?
- PRD, spec, or design doc - what was the intended scope?
- CLAUDE.md / AGENTS.md - project-level instructions
- package.json description, pyproject.toml metadata
Capture the stated objective in one sentence. This becomes the benchmark for section 5.4 (does the system actually achieve its goal?).
1e. Confirm Scope with User
Before proceeding to Phase 2, present a brief summary:
Discovered: {project type} with {N} components/agents
- {tech stack}
- {data stores}
- {file count} files across {dir count} directories
- Stated goal: "{one sentence}"
Scope: review everything above. Proceed?
Wait for user confirmation. This prevents wasting tokens reviewing the wrong project, wrong subdirectory, or wrong scope.
If user declines scope: ask what to narrow to or exclude. Do not abort - adjust the scope and re-confirm.
1f. Choose Delivery Mode
After scope confirmation, ask the user how they want the report delivered:
Deliver the report:
[1] Inline (default) - full report in this session. Nothing saved to disk.
[2] File - full report in this session AND saved to
shakedown-YYYY-MM-DD.md at the project root. Pick this when you want
a standalone artifact to archive, share, or feed into a fresh session
later.
Which? (default: inline)
Capture the choice. This only changes whether the report is persisted to disk - both modes render the complete report in the current session, so follow-up handlers (Phase 7) work identically in either mode.
-
Inline mode: render the complete report in the session per the Output Format section below. Do not write any file.
-
File mode: render the complete report in the session AS WELL, then also write the identical report to
shakedown-YYYY-MM-DD.mdat the project root (use today's date).When to write the file: defer the write until ALL report phases complete (after Phase 6 has been rendered in-session, before any Phase 7 follow-up). The file write is the LAST action before surfacing the Phase 7 follow-up menu.
What to write: the identical report rendered in the session, byte-for-byte. Do NOT write a different shape, summary, or partial report. Use a single
Writecall withfile_path=<projectRoot>/shakedown-YYYY-MM-DD.md(today's date in Europe/Amsterdam if no other timezone is established).After the write, append exactly one line to the session:
Report saved: ./shakedown-YYYY-MM-DD.mdThen proceed directly to the Phase 7 follow-up menu.
If
Writeor file-creation is not available on the platform, fall back to inline mode and say so once. Do not silently skip.
Phase 2: Full Content Read (Parallel)
Dispatch 4 parallel agents, each reading one slice of the project. Adapt slices to whatever Phase 1 discovered.
Fallbacks: If no subagent support, run slices sequentially. If
1000 files, sample 2-3 per directory and count the rest. If an agent fails, continue with the rest and note the gap.
Agent 1: Architecture & Configuration
Read all configuration and architectural files:
- Config files (JSON, YAML, TOML, .env - note secrets, don't expose values)
- Agent/worker definitions (instruction files, prompt templates, role definitions)
- Pipeline/workflow definitions
- Infrastructure config (Docker, CI/CD, deployment)
Agent 2: Execution Logic & Coordination
Read all files that define HOW the system runs:
- Heartbeats, schedulers, cron definitions, event handlers
- Coordination mechanisms (handoffs, queues, barriers, locks)
- Error handling, retry logic, fallback paths
- Entry points, main loops, orchestrators
Agent 3: Outputs, Docs & Reviews
Read all output and documentation:
- Recent outputs (latest 3-5 of each type)
- Documentation (design docs, runbooks, ADRs, project memory)
- Review files if they exist
- Scripts (build, deploy, report generation)
Agent 4: Data, Infrastructure & Supporting Files
Explore and quantify (count, don't read every file):
- All data directories - count files, read 2-3 samples to understand schema
- Logs - check latest entries for errors, check if structured (JSON) or unstructured
- Backups - list, check sizes and dates
- Skills/plugins/extensions - read definitions
- Memory/state files - understand persistence model
Phase 3: Data Store Diagnostics
If the project has a database (SQL, NoSQL, or embedded), run diagnostics.
If no database found: skip this phase. Note "No data store detected" in the report and proceed to Phase 4.
If database found but credentials unavailable: skip queries. Note "Database detected but could not access - manual credential needed" and proceed to Phase 4.
If shell execution is not available: skip this phase entirely. Note "Data store diagnostics skipped - no shell access" and proceed to Phase 4.
See DB Diagnostics for database-specific queries and inspection guidance.
Phase 4: Quality Analysis
Synthesize everything from Phases 1-3.
4.1 What the Project Is
One paragraph: purpose, who it serves, architecture summary, tech stack. Derived from what you read, not assumed.
4.2 What Works Well
Genuine strengths with evidence. Be specific - cite files, patterns, design decisions that are genuinely good.
4.3 Critical Issues
Things that will cause failures soon. Must include evidence:
- Reliability data (failure rates, error messages)
- Broken coordination (instructions reference things that don't exist)
- Dead code, placeholder files, unfinished features presented as complete
- Pipeline stages that are out of sync
4.4 Architecture & Code Quality
Evaluate structural analysis, design pattern coherence (MECE check, contradictions), algorithm efficiency, code quality (dead code, complexity hotspots, naming), dependency graph, and test coverage.
See Architecture & Code Quality for the full checklist. Key outputs: biggest structural risk, MECE gaps, test coverage assessment, and complexity hotspots.
4.5 Error Handling, Resilience & Failure Modes
Trace actual behavior - don't just note that mechanisms exist. Cover crash scenarios, timeout coverage, silent failures, data integrity, edge cases, system-level failure modes, retry patterns, and graceful degradation.
See Error Handling & Resilience for the full checklist. Key outputs: failure paths traced, silent failure inventory, data loss window, and recovery readiness.
4.6 Performance & Bottleneck Analysis
See Performance Analysis for the full checklist covering timing, parallelism, scaling, resource waste, cost analysis, and optimization opportunities.
Summarize key findings here: biggest bottleneck, biggest waste, estimated cost, and top 3 optimization opportunities with expected impact.
4.7 Code & Storage Efficiency
Assess how lean the project's file footprint is. Check for: empty (0-byte) files, duplicate files across directories, build artifacts or temp files committed to git, copy-pasted code blocks, dead dependencies (declared but unused), and storage bloat (large binaries, oversized logs).
See Storage Efficiency for the full checklist with commands. Quantify: total waste in file count and bytes.
4.8 Agent Skill Standards Compliance
Conditional - only run if Phase 1 detected this project is an agent skill. If not a skill, note "Not an agent skill - section skipped."
Evaluate against the Agent Skills specification (agentskills.io) and platform best practices. See Skill Standards for the full compliance checklist.
Summarize as: spec conformance, description quality, instruction quality, script quality, eval framework, and progressive disclosure - each rated PASS / PARTIAL / FAIL with specific issues noted.
Phase 5: Security, Readiness & Recommendations
5.1 Security & Data Exposure
The security reference is split into 4 files for progressive disclosure. Load only the files that match the project's surface area:
| Reference | When to load | Contents |
|---|---|---|
| Traditional | Always | Section A: secrets, injection, PII, supply chain, workflow, network, licensing |
| LLM & GenAI | If project uses LLMs, GenAI APIs, vector stores, or MCP tools | Sections B (OWASP LLM Top 10 2025), D (MCP Security), E (GenAI Data Security) |
| Agentic | If project uses autonomous agents, multi-agent coordination, or tool-using agents | Sections C (OWASP Agentic Top 10 2026), H (23 concrete test procedures) |
| Governance & Red Teaming | If project must comply with EU AI Act / NIST AI RMF / ISO 42001 or undergo formal red-team review | Sections F (Governance & Compliance), G (Red Teaming Readiness) |
Detection rules:
- LLM file triggers:
openai/anthropic/google-genai/mistralai/ollamaimports;langchain/llama-index/crewai/autogenpackages;*.prompts.*/*.system_prompt*files;mcp_servers.*.json/claude_desktop_config.json. - Agentic triggers: anything from the LLM list PLUS multi-agent
patterns (worker pools, orchestrator, agent-to-agent message bus,
TaskCreate/spawn loops,
.crew/.autogen/.langgraphdirectories). - Governance triggers: explicit references to EU AI Act, NIST AI RMF,
ISO 42001/27001, GDPR DPIA, internal AI policy documents, or any
compliance/,governance/,audit/directory.
If none of the LLM/agentic/governance triggers match, load only
security-traditional.md and skip the other three files.
Summarize key findings here: any secrets exposed, injection risks found, PII handling issues, dependency vulnerabilities, LLM/agent-specific risks.
5.2 Logging & Observability
Assess log existence, quality (structured vs free text), traceability (can you follow a request end-to-end?), lifecycle (rotation, retention), and monitoring/alerting.
See Operational Health for the full checklist.
5.3 Documentation Quality
Compare documentation against the actual system: accuracy (docs vs reality), completeness (could someone else operate this?), maintenance (when last updated?), and onboarding quality.
See Operational Health for the full checklist. Flag any doc that describes nonexistent features or omits existing ones.
5.4 Goal Fulfillment
Scope: Technical comparison - does the code match the docs? For strategic assessment (should this exist?), see 5.10 Value Assessment.
Compare the stated objective (captured in Phase 1d) against actual behavior:
- Does the system do what it claims to do?
- Are there features described in docs/README that don't work?
- Are there capabilities the system has that aren't documented?
- Is the objective achievable with the current architecture?
5.5 Blind Spots
What nobody is monitoring: feedback loops, cost tracking, SLA adherence, input health, error visibility, graceful degradation, and drift/rot.
See Operational Health for the full checklist.
5.6 Objective Clarity Assessment
Scope: Operational dimensions (does it work reliably?). For strategic dimensions (is it worth building?), see 5.10 Value Assessment.
| Dimension | Rating (1-10) | Evidence |
|---|---|---|
| Objective clarity | Is the goal well-defined? | |
| Goal fulfillment | Does the system actually achieve it? | |
| Delivery reliability | Does it work consistently? | |
| Output quality | Is the output actually good? | |
| Automation maturity | How much runs unattended? | |
| Self-improvement | Does it learn from failures? | |
| Operational visibility | Can you see what's happening? | |
| Resource efficiency | Is it wasteful or lean? |
Adapt dimensions to the project type. Drop irrelevant ones, add project-specific ones.
5.7 Overall Rating
X/10 with one-sentence justification.
5.8 Production Readiness Assessment
Rate each gate as PASS / PARTIAL / FAIL:
| Gate | Status | Evidence |
|---|---|---|
| Functionality - does it do what it promises? | ||
| Reliability - does it work consistently without manual intervention? | ||
| Error handling - does it recover from failures gracefully? | ||
| Security - no exposed secrets, injection risks, or PII leaks? | ||
| Testing - are critical paths covered by tests? | ||
| Monitoring - can you tell when something breaks? | ||
| Documentation - can someone else operate this? | ||
| Scalability - will it handle growth without redesign? | ||
| Data integrity - is data consistent, backed up, recoverable? | ||
| Dependency health - are deps maintained, pinned, vulnerability-free? |
Readiness summary: State the count of PASS / PARTIAL / FAIL gates, then list the specific blockers in priority order. Do not issue a "ready to ship" or "not production-ready" verdict - that call is the user's, not yours.
5.9 Top 10 Ranked Recommendations
| # | Action | Impact | Effort | Who Implements |
|---|---|---|---|---|
| 1 | ... | Critical | Low | ... |
For "Who Implements": can the system's own agents/workers fix this, or does the human need to intervene directly?
5.10 Value Assessment
Scope: Strategic assessment - should this project exist in its current form? For technical goal comparison (code vs docs), see 5.4. For operational ratings (reliability, automation), see 5.6.
Assess: problem clarity, target audience definition, maturity vs. claims, measurable value, differentiation from alternatives, and adoption readiness.
See Value Assessment for the full framework with questions per dimension.
Summarize as a table:
| Dimension | Rating (1-5) | Evidence |
|---|---|---|
| Problem clarity | ||
| Audience definition | ||
| Maturity vs. claims | ||
| Measurable value | ||
| Differentiation | ||
| Adoption readiness |
5.11 The Uncomfortable Question
The one thing the project owner needs to hear but probably doesn't want to. Infrastructure gaps, fundamental design flaws, or unstated assumptions that undermine everything else.
Phase 6: Resilience Testing
If no backups exist: note as critical gap in the report. Skip restore testing but still assess operational resilience.
If shell execution is not available: skip restore testing. Instead, assess resilience from code analysis only - check whether backup logic exists in the code, whether recovery procedures are documented, and whether there are any disaster recovery references. Note "Resilience testing limited - no shell access for restore verification."
See Resilience Testing for backup validation steps and operational resilience checks.
Phase 7: Follow-Up Handlers
After the report is rendered (and the file is written if File mode was chosen in Phase 1f), surface the follow-up menu exactly as below, then wait for the user to pick:
Follow-up actions:
[A] Fix them all - implement all Top 10 recommendations now
[B] Fix selected - "fix #2, #5, #7" or "fix all Critical/High"
[C] Re-run Phase 5 - refresh security / readiness / recs only
[D] Compare with last - diff this report against the previous one
[E] Create GitHub issues - one issue per recommendation
[F] Done - exit, no further action
Pick one (default: F):
If the user picks nothing within their next message, treat it as F.
Handler A - Fix them all
Walk the Top 10 Ranked Recommendations table in order. For each row:
- Mark the row's recommendation as
in_progress(use TaskCreate / TodoWrite if available; otherwise an inline checklist). - Apply the change. Read every file before editing it. Make the smallest change that satisfies the recommendation - no extra cleanup or refactors.
- Verify: re-run any test, validator, or grep that proves the fix.
If verification fails, leave the recommendation
in_progressand surface the failure; do not silently advance. - Mark the row
completedand move on.
Stop conditions: stop early and surface the situation if (a) any recommendation requires destructive action (delete data, force-push, modify shared infrastructure) - confirm before proceeding, (b) two consecutive recommendations fail verification, (c) the user interrupts.
After the walk, print a one-line summary: N/M recommendations applied. Skipped: [list]. Failed: [list].
Handler B - Fix selected
Same as Handler A, but operate only on the rows the user names. Accept
either explicit IDs (fix #2, #5, #7) or a band (fix all Critical/High). If the request is ambiguous, ask once.
Handler C - Re-run Phase 5
Re-execute Phase 5 only (Security, Logging, Documentation, Goal Fulfillment, Blind Spots, Objective Clarity, Overall Rating, Production Readiness, Top 10 Recommendations, Value Assessment, Uncomfortable Question). Skip Phases 1-4 and 6 unless the user asks for a full re-run. Render in the same delivery mode as the original report.
Handler D - Compare with last
Find the most recent prior shakedown-YYYY-MM-DD.md at the project
root (use Glob for shakedown-*.md, sort by date in filename, pick
second-most-recent if today's exists). If none exists, say so.
Otherwise produce a diff table showing: ratings deltas (5.6 / 5.7 /
5.8), recommendations resolved since last run, recommendations still
open, and any new issues that appeared.
Handler E - Create GitHub issues
Verify gh is on PATH. If not, abort and say so. Otherwise, for each
row in the Top 10 Ranked Recommendations:
- Build an issue title:
[shakedown] {short Action}(truncate at 70 chars). - Build a body that includes: the Action, Impact, Effort, Who Implements, plus the evidence cite from the report.
- Run
gh issue create --title ... --body ...with proper escaping (HEREDOC for the body). - Print the URL of each created issue.
Do not push code, do not open PRs - issues only. The user runs the fix loop themselves or invokes Handler A separately.
Handler F - Done
Exit cleanly. No action taken.
Output Format
Structured markdown report with:
- Tables for ratings, recommendations, and reliability metrics
- Direct, challenging tone - find what's wrong, not just what's right
- Evidence for every claim (file paths, query results, counts)
- Actionable: every recommendation must be specific enough that Phase 7 Handler A can fix it without re-asking the user
For shape: see examples/benchmark-review-a.md,
examples/benchmark-review-b.md, examples/benchmark-review-c.md for
full-form exemplars. The validator at scripts/validate-output.py
encodes the structural contract (sections, tables, columns, rating
ranges); run it on any output to confirm conformance. The eval set at
evals/evals.json defines the test matrix and assertion rules.
Apply the delivery mode captured in Phase 1f:
- Inline: stream the full report into the session.
- File: stream the full report into the session, then write the
identical report to
shakedown-YYYY-MM-DD.mdAFTER Phase 6 completes and BEFORE the Phase 7 follow-up menu (see Phase 1f for the exact write contract).
Gotchas
See Gotchas for the full list of 10 common agent mistakes with examples and fixes. Read before starting Phase 4.
Key ones: don't praise by default, evaluate don't just list features, read files before referencing them, treat code as truth over README, go deep on security, always estimate cost, make recommendations specific, and check what happens at scale.
Key Principles
- Discover, don't assume - map structure before reading content
- Confirm scope - present discovery to user before deep read
- Read before judging - never critique what you haven't read
- Quantify - "23% duplicate rate" not "some duplicates"
- Compare stated vs actual - docs say X, system does Y
- Every recommendation needs: what, why, effort, who implements
- Actionable in same session - user says "fix them all", you proceed
- Test resilience - verify backups restore, trace failure paths
- Challenge the project - the goal is to make it better, not praise it
- Surface constraints - subscription limits, peak hours, budgets, SLAs
What ships with it: 28 files
2066.4 KB alongside SKILL.md, 2 of them executable
evals/
- evals.json16.9 KB
- README.md3.3 KB
examples/
- baseline-review-a.md10.6 KB
- baseline-review-b.md13.7 KB
- baseline-review-c.md12.4 KB
- benchmark-review-a.md68.0 KB
- benchmark-review-b.md62.3 KB
- benchmark-review-c.md40.7 KB
references/
- architecture-quality.md7.9 KB
- db-diagnostics.md1.3 KB
- error-resilience.md5.8 KB
- gotchas.md6.5 KB
- operational-health.md6.6 KB
- performance-analysis.md7.2 KB
- resilience-testing.md8.7 KB
- security-agentic.md44.8 KB
- security-governance.md9.1 KB
- security-llm.md32.2 KB
- security-traditional.md4.7 KB
- skill-standards.md8.2 KB
- storage-efficiency.md3.2 KB
- value-assessment.md2.2 KB
scripts/
- parity-check.pyruns4.6 KB
- validate-output.pyruns7.6 KB
- .gitignore245 B
- LICENSE1.0 KB
- README.md25.0 KB
- shakedown.png1651.5 KB