agentsclimarketplace

Ai pair hunting with claude

Skill ShulkwiSEC/bb-huge/skills/curated/ai-pair-hunting-with-claude

bb-huge πŸ€— , Personal bug bounty findings hub and bug bounty orchestration for multiple agents

Install
npx -y skills add ShulkwiSEC/bb-huge --skill ai-pair-hunting-with-claude

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 18 stars18 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Configure Claude as a "Pair Hunter" β€” autonomous overnight hacking, context management via per-target .claudemd files, sub-agent compaction avoidance, and scope enforcement. Based on Critical Thinking Bug Bounty Podcast Episode 166.

The file declares its own license as Apache-2.0. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

12.8 KB, as published. Nobody here has run it

AI Pair Hunting with Claude

When to Use

  • When setting up Claude Code CLI to autonomously test a bug bounty target.
  • When running overnight / multi-hour autonomous hacking sessions.
  • When managing scope and context across multiple bug bounty programs.
  • When Claude is getting stuck in compaction loops or losing context mid-session.
  • When you need Claude to stay strictly within program scope.

Prerequisites

  • Claude Code CLI installed and authenticated
  • A defined bug bounty target with program policy
  • Separate workspace directory per target program
  • Understanding of Claude's context window limitations

Core Concept: Claude as a Pair Hunter

"AI is not your replacement. It is your Pair Hunter β€” it brings determinism (accuracy) and speed." β€” Critical Thinking Podcast, Ep. 166

Claude excels at:

  • Deterministic tasks: Testing every parameter in a 200-param API
  • Speed: Fuzzing/enumerating faster than manual testing
  • Pattern recognition: Spotting anomalies in large response sets
  • Documentation: Auto-generating reports from findings

Claude struggles with:

  • Creative intuition: "This feels wrong" β€” that is YOUR job
  • Out-of-scope judgment: Without explicit policy, Claude will test everything
  • Long context retention: After ~100k tokens, context degrades

Workflow

Phase 1: Per-Target Context Management

"For every target, create a separate folder. Put a .claudemd file in it with the program's policy and scope." β€” Episode 166 [51:20]

Directory structure:

programs/
β”œβ”€β”€ example-corp/
β”‚   β”œβ”€β”€ .claudemd           # ← THIS IS THE KEY FILE
β”‚   β”œβ”€β”€ notes/
β”‚   β”œβ”€β”€ leads/
β”‚   β”œβ”€β”€ findings/
β”‚   └── scripts/
β”œβ”€β”€ another-target/
β”‚   β”œβ”€β”€ .claudemd
β”‚   └── ...

.claudemd template:

# Target: Example Corp Bug Bounty Program

## Program URL
https://hackerone.com/example-corp

## Scope β€” IN
- *.example.com
- api.example.com
- app.example.com (authenticated testing allowed)
- mobile-api.example.com

## Scope β€” OUT (DO NOT TEST)
- blog.example.com (third-party WordPress)
- status.example.com (StatusPage hosted)
- *.example.dev (staging β€” explicitly excluded)
- Any domain not listed above

## Rules of Engagement
- NO denial of service testing
- NO social engineering of employees
- NO accessing other users' data beyond proof of concept (read 1 record, stop)
- Rate limit: Max 10 requests/second
- Report vulnerabilities within 24 hours of confirmation

## Authentication
- Test account 1: [email protected] / [use env var AUTH_TOKEN_1]
- Test account 2: [email protected] / [use env var AUTH_TOKEN_2]
- API Key: [use env var EXAMPLE_CORP_API_KEY]

## Tech Stack (Known)
- Frontend: React 18 + Next.js
- Backend: Node.js + Express
- Database: PostgreSQL (inferred from error messages)
- CDN: Cloudflare
- Auth: OAuth 2.0 + JWT

## Priority Targets
1. `/api/v2/users/*` β€” IDOR testing
2. `/api/v2/billing/*` β€” Payment logic flaws
3. `/upload/*` β€” File upload vulnerabilities
4. `/auth/*` β€” Authentication bypass

## Previous Findings (to avoid duplicates)
- XSS in /search β€” reported 2024-12, resolved
- IDOR in /api/v1/users/{id} β€” reported 2025-01, resolved (v2 untested)

Why this works: When you run claude from inside programs/example-corp/, Claude automatically reads .claudemd and constrains itself to the defined scope.

Phase 2: Autonomous Overnight Hacking

"I am going to bed. Don't ask for input. Keep hacking." β€” Episode 166 [40:01]

The overnight prompt:

I am going to sleep. Do not ask me for input or confirmation.

Your mission for the next 4 hours:
1. Read the .claudemd file for scope and rules.
2. Map all API endpoints on api.example.com using the scripts in scripts/.
3. For each endpoint, test the following:
   - IDOR: Replace user IDs / resource IDs with other values.
   - Broken auth: Access endpoints without auth token, with expired token, with wrong role.
   - Input validation: Fuzz all parameters with payloads from scripts/fuzz-payloads.ts.
4. Log ALL findings to findings/ using the finding template.
5. Log all interesting observations to notes/.
6. Do NOT exceed 10 requests per second.
7. Do NOT test anything outside the scope defined in .claudemd.

When you finish, write a summary to overnight-report-[date].md with:
- Endpoints tested (count)
- Findings discovered (count + severity)
- Areas that need manual follow-up
- Errors encountered

Key considerations for overnight sessions:

FactorGuidance
Token usageThe session may cost $5-20+ in API tokens for 4 hours
Rate limitsEnforce in your scripts, not just in the prompt
Scope violationsThe .claudemd scope section is critical safety net
False positivesExpect ~30-50% of "interesting" results to be false positives β€” triage in morning
Tool permissionsPre-approve network access and file write permissions before sleeping

Phase 3: Compaction Avoidance

"If you use too many sub-agents (4+), Claude gets stuck in a compacting loop. Keep it to 2-3 sub-agents max." β€” Episode 166 [41:11]

What is compaction? When Claude's context window fills up, it "compacts" by summarizing older conversation turns. If too many sub-agents are running, the compaction process itself fills the context, creating a death spiral.

Rules:

❌ BAD: Spawning 5+ parallel sub-agents
   "Run these 5 separate tasks simultaneously..."
   Result: Context fills β†’ compaction loops β†’ Claude freezes

βœ… GOOD: Sequential tasks with 2-3 sub-agents max
   "First enumerate endpoints, then test the top 10 for IDOR"
   Result: Clean context, focused execution

βœ… GOOD: Use scripts instead of sub-agents for parallel work
   "Run scripts/parallel-fuzz.ts which handles 50 endpoints internally"
   Result: Claude manages 1 task; the script handles parallelism

Compaction warning signs:

  • Claude starts repeating itself
  • Responses become shorter and lose detail
  • Claude "forgets" earlier findings in the same session
  • Session time between responses increases dramatically

Mitigation strategies:

  1. Break long sessions into phases: Run 1-hour focused sessions instead of 4-hour marathons
  2. Offload parallelism to scripts: TypeScript scripts handle concurrency, Claude orchestrates
  3. Use the funnel: Write findings to disk immediately (see bug-bounty-workflow-funnel skill) so data survives compaction
  4. Start fresh sessions: If compaction is occurring, start a new Claude session with a summary of progress

Phase 4: Effective Prompting Patterns

Directive prompts (high autonomy):

Test all endpoints in api.example.com/v2/ for IDOR vulnerabilities. 
Use the authenticated tokens from .claudemd. Log findings to findings/.
Do not stop until all endpoints are tested.

Constraint prompts (safety rails):

You are ONLY allowed to test the following 3 endpoints:
- GET /api/v2/users/{id}
- POST /api/v2/users/{id}/update
- DELETE /api/v2/users/{id}

Do NOT test any other endpoint. Do NOT make more than 100 total requests.
After each test, write the result to notes/idor-test-results.md.

Chain prompts (multi-step):

Phase 1: Enumerate all JavaScript files on app.example.com. Extract API endpoints, 
         secrets, and interesting variables. Save to notes/js-analysis.md.

Phase 2: For each API endpoint found in Phase 1, test for:
         - Missing authentication
         - IDOR via ID manipulation  
         - SQL injection via single-quote test
         Save results to leads/.

Phase 3: For any confirmed vulnerability from Phase 2, create a full finding 
         document in findings/ with PoC.

Phase 5: Security & Permissions

"We use dangerouslySkipPermissions β€” but Claude has NO access to 1Password or personal email." β€” Episode 166

When running autonomous sessions, the --dangerously-skip-permissions flag prevents Claude from asking for confirmation on every file write and network request. However, you MUST harden the environment:

βœ… Allow❌ Block
Network access to in-scope targets onlyPassword managers (1Password, Bitwarden)
Write to findings/notes/leads directoriesPersonal email clients
Execute scripts in the target workspaceSSH keys to production systems
Read source code directoriesCloud provider admin CLIs

Best practice: Run Claude in a sandboxed VPS user account with restricted network access (firewall to scope IPs only). See the remote-hunting-workflow skill for full VPS setup, Discord bot, and tmux multi-target configurations.

Decision Point πŸ”€

flowchart TD
    A[Starting a new target] --> B[Create target directory]
    B --> C[Write .claudemd with scope + policy]
    C --> D{How long will you hunt?}
    D -->|"< 1 hour"| E[Interactive session with constraint prompts]
    D -->|"1-4 hours"| F[Focused directive prompt, 2-3 sub-agents max]
    D -->|"Overnight"| G[Overnight prompt with full autonomy]
    E --> H{Claude losing context?}
    F --> H
    G --> H
    H -->|Yes| I[Start new session with progress summary]
    H -->|No| J[Continue hunting]
    J --> K{Findings discovered?}
    K -->|Yes| L[Promote through Funnel stages]
    K -->|No| M[Expand test surface or switch targets]

Anti-Patterns

Anti-PatternConsequenceFix
No .claudemd fileClaude tests out-of-scope targets β†’ policy violation β†’ banAlways create scope file first
5+ sub-agents at onceCompaction death spiral, lost contextMax 2-3 sub-agents; offload to scripts
No rate limitingTarget's WAF blocks you; program bans your accountEnforce in scripts + prompt
Hardcoded auth tokensTokens leaked if skill/prompt is sharedUse environment variables
No disk persistenceFindings lost when context compactsWrite to files immediately
Vague overnight promptClaude wanders, tests irrelevant thingsSpecific objectives + strict scope

Output Format

Pair Hunting Session Report
===========================
Target: example-corp
Duration: 3h 45m
Mode: Overnight autonomous

Endpoints Tested: 47
Requests Made: 1,842

Findings:
  - [HIGH] IDOR in GET /api/v2/users/{id} β€” any user profile accessible
  - [MEDIUM] Missing rate limit on POST /api/v2/auth/login β€” brute force possible
  - [LOW] Verbose error messages on /api/v2/billing/charge with invalid card

Leads Requiring Manual Follow-up: 5
Notes Generated: 12 files

Context Health: βœ… No compaction issues
Scope Violations: ❌ None

Creativity Directive

IMPORTANT: The instructions above are a STARTING POINT, not a boundary. You are expected to go beyond these steps:

  • Develop your own prompting patterns for specific vulnerability classes.
  • Create meta-skills that orchestrate other skills in sequence.
  • Build monitoring scripts that alert you when Claude finds something critical overnight.
  • Experiment with different context management strategies.

Think like an attacker. Adapt. Improvise.

πŸ”΅ Blue Team

  • Deploy robust WAF rules to detect anomalies.
  • Monitor logs for unusual access patterns.

πŸ“š Shared Resources

For cross-cutting methodology applicable to all vulnerability classes, see:

References

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.