agentsclimarketplace

Security review

Skill kklimuk/docx-cli/.claude/skills/security-review

Review code for security vulnerabilities. Use when the user says 'security review', 'security audit', 'check for vulnerabilities', 'pentest the code', 'OWASP check', or any variation of wanting a security assessment.From its SKILL.md

Install
npx -y skills add kklimuk/docx-cli --skill security-review

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • runs commandsInstructs the agent to run 3 commands, including `git diff main...HEAD --name-only` and 2 more.

SKILL.md

6.8 KB, ~1.4k tokens by cl100k_base, as published. Nobody here has run it

Security Review

Audit changed files for security vulnerabilities, focusing on the OWASP Top 10 and issues specific to the project's stack.

When running locally as a forked subagent, the main session does not see any files you read or any reasoning you do — only the final report you return. When running in CI (e.g. via claude-code-action), the workflow takes the report and turns it into GitHub PR review comments. Either way, take your time, read every changed file completely, and produce a thorough, actionable report. The consumer of this report uses it as a worklist, so it must be complete and self-contained.

Scope

Determine the diff to review:

  1. Run git diff main...HEAD --name-only to get files changed on this branch vs main.
  2. If that fails (no main, detached worktree, etc.), fall back to git diff HEAD --name-only for uncommitted changes, then git diff --cached --name-only for staged files.
  3. If no diff is available, ask the user which files to review.

Read every changed file completely before starting the review. Read CLAUDE.md first to understand the project's stack and any subsystems with security-sensitive surface area (auth, real-time, payments, file uploads).

What to Look For

Injection & Input Handling

  • SQL injection: Raw SQL with string interpolation instead of parameterized queries. Check for any template literals that build SQL, and verify that database libraries are being used in their parameterized form.
  • Command injection: User input passed to shell commands (Bun.$, child_process, subprocess, os.system) without sanitization.
  • XSS: User-controlled data rendered as dangerouslySetInnerHTML, or reflected into HTML/JS without escaping. Check contentEditable fields that accept pasted HTML.
  • Path traversal: User input used in file paths without validation. Check for .. traversal.
  • Prototype pollution: Object.assign or spread on user-controlled objects without allowlisting keys.

Authentication & Authorization

  • Missing auth checks: Endpoints that read/write data without verifying the caller's identity or org membership.
  • IDOR (Insecure Direct Object Reference): Endpoints that accept an ID parameter and return/modify the resource without verifying the caller has access. Particularly dangerous when URL params (like a slug) aren't validated against the actual resource ownership.
  • Privilege escalation: Actions that should be restricted (delete, move, admin operations) but aren't gated on role/permission.

Data Exposure

  • Over-fetching: API responses that include more data than the client needs (e.g., internal IDs, secrets, full document state when only a title is needed).
  • Error leakage: Stack traces, SQL errors, or internal paths exposed in error responses.
  • Sensitive data in logs: Passwords, tokens, or PII logged to console.

Real-time / WebSocket Security

(Only relevant if the project has a WebSocket layer — see CLAUDE.md.)

  • Channel authorization: Can a client subscribe to any channel by guessing the name? Are channel subscriptions validated against user permissions?
  • Message spoofing: Can a client broadcast messages to channels they shouldn't have write access to?
  • Payload validation: Are incoming WebSocket messages validated before processing?

Denial of Service

  • Unbounded queries: Endpoints that return all records without pagination or limits.
  • Regex DoS: User input used in regex patterns without sanitization.
  • Resource exhaustion: File uploads, large request bodies, or expensive operations without rate limiting or size limits.

Cryptography & Secrets

  • Hardcoded secrets: API keys, passwords, or tokens in source code.
  • Weak randomness: Math.random() / random.random() used for security-sensitive operations instead of crypto.randomUUID() / secrets.token_*().
  • Missing TLS: WebSocket connections using ws:// in production contexts.

Dependencies

  • Known vulnerabilities: If bun audit / npm audit / pip-audit is available, check for known CVEs.
  • Prototype pollution via deps: Libraries that merge user input deeply.

Report Format

Return the complete formatted report as your final message — not a summary or TL;DR. Whatever consumes the report (a main Claude session locally, or a CI workflow that posts inline GitHub PR comments) uses it as a worklist, so it must be self-contained.

Organize findings by severity:

Critical

Exploitable now with no authentication required. Data loss, unauthorized access, or remote code execution.

High

Exploitable with some preconditions (e.g., needs authenticated user, specific timing). Privilege escalation, significant data leakage.

Medium

Defense-in-depth issues. Missing validation that's currently protected by another layer but shouldn't rely on it.

Low

Hardening recommendations. Not exploitable today but reduce attack surface.

For each finding, include enough detail that the consumer can apply the fix without re-reading the entire file:

  1. File and line — exact path:line (or path:start-end for ranges); list every site for cross-file findings
  2. Severity — Critical / High / Medium / Low
  3. Vulnerability type — OWASP category or CWE
  4. Current code — short snippet of the vulnerable code (not just a description)
  5. Exploit scenario — concrete steps showing how an attacker would use this
  6. Fix — specific code change, ideally as a before/after snippet
  7. Surrounding context — callers, related files that must change in lockstep, validation layers the fix depends on, tests that should be added

What NOT to Do

  • Don't flag style issues — that's the code review's job.
  • Don't suggest adding WAFs, rate limiters, or infrastructure changes unless the code-level fix is insufficient.
  • Don't report theoretical issues that require physical access or compromised infrastructure.
  • Don't pile on — prioritize the top findings that matter most.

After the Report

The fix phase (or PR-comment-posting phase) happens in whatever consumes this report — not here. Your job ends when you return the report. Make sure it has enough information for that consumer to act on findings without re-reading the codebase. Locally, the main session will work through findings in severity order with minimal, targeted fixes and run bun run check + bun test after each. In CI, the workflow will turn each finding into a GitHub PR review comment.

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Gives 0 of the 12 instructions most review quality skills give in ~1.4k tokens

Counted across 1,273 of the 2,403 authors here whose files we hold, read 2026-09-06

  • Ask one question at a timein 63 of 1273, across 62 files
  • Provide a recommended answer for each questionin 47 of 1273, across 45 files
  • Rank findings by severityin 44 of 1273
  • Use parameterized queries for database accessin 38 of 1273, across 20 files
  • Validate all user input with schemasin 33 of 1273, across 15 files
  • Store secrets in environment variablesin 32 of 1273, across 14 files
  • Explore the codebase to answer questionsin 31 of 1273, across 29 files
  • Store tokens in httpOnly cookiesin 30 of 1273, across 12 files
  • Implement rate limiting on API endpointsin 30 of 1273, across 12 files
  • Sanitize user-provided HTMLin 29 of 1273, across 11 files
  • Return generic error messages to usersin 28 of 1273, across 10 files
  • Cite file and line for every findingin 28 of 1273, across 25 files

Said here and by no other author read

  • Read every changed file completely before starting the review
  • Return a complete and self-contained report

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.