agentsclimarketplace

Pwnote offsec osai

Skill Pwnote/skills/skills/pwnote-offsec-osai

AI agent skills for Pwnote Pentest Notebook

Install
npx -y skills add Pwnote/skills --skill pwnote-offsec-osai

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • 14 days oldThe repository was created 14 days ago. New is not bad, but a brand new repository carrying a familiar-sounding name is the shape a typosquat arrives in, and there has been no time for anyone else to find a problem with it.
  • 3 stars3 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Use whenever the user is working on Offsec's OSAI / AI Red Teaming certification track, or on AI/LLM/agentic security engagements generally — prompt injection findings, tool-use abuse, agent trajectory documentation, or writing up AI-specific security findings that don't map cleanly to traditional CVSS. Trigger on "OSAI", "AI Red Teaming", "LLM pentest", "prompt injection engagement", or "agentic AI exploitation", even without the word "skill".

SKILL.md

4.1 KB, as published. Nobody here has run it

Offsec OSAI / AI Red Teaming Workflow

Reference for running and documenting an AI/LLM security assessment — prompt injection, tool-use abuse, and agent-specific findings that need a different taxonomy and severity model than traditional web/infra findings.

1. AI-Specific Finding Taxonomy

Categorize findings by the underlying attack predicate rather than a generic "prompt injection" label — this keeps findings comparable across engagements and maps directly onto a structured attack-algorithm taxonomy if one is in use:

CategoryDescription
EXFILTRATIONThe system is induced to leak data it shouldn't (system prompt, other users' context, tool outputs, secrets in context) to the attacker or an external destination
UNTRUSTED_TO_ACTIONUntrusted input (a document, webpage, tool result) is treated as an instruction and drives a consequential action the legitimate user didn't request
PRIVILEGE_ESCALATIONThe agent is induced to use a tool/permission beyond what the current user/context should allow
PERSISTENCEThe injected behavior survives beyond the single turn/session (e.g., written to memory, a file, or a scheduled task the agent later reads back)
DENIAL_OF_SERVICEThe agent is induced into a resource-exhausting loop, excessive tool calls, or a stuck state
JAILBREAK / POLICY_BYPASSModel safety/policy constraints are circumvented, independent of any tool-use consequence

Tag every finding with a primary category (and secondary if it chains, e.g. UNTRUSTED_TO_ACTION → EXFILTRATION).

2. Test Harness / Trajectory Documentation

Agent findings need the full trajectory, not just an input/output pair, since the vulnerability is often in how the agent got there:

- Turn-by-turn transcript (user input, model reasoning if visible, tool calls + arguments, tool results)
- Point of deviation — the exact turn where the agent's behavior diverged from expected/authorized behavior
- Root cause — which trust boundary was crossed (e.g., tool output treated as trusted instruction)
- Reproducibility — does it require a specific model version/temperature, or is it deterministic

If session replay/recording tooling is available in the environment being tested, capture the full replay artifact alongside the transcript — a static screenshot loses the tool-call sequence that's usually the actual finding.

3. Report Format

AI findings often don't map cleanly to CVSS since impact depends heavily on what tools/permissions the agent has, which varies by deployment. Use a two-part severity model:

## Finding: [Name]
Category: [EXFILTRATION / UNTRUSTED_TO_ACTION / etc.]

### Technical Severity
[How reliably the injection/bypass works, in isolation — independent of deployment]

### Deployment Impact Severity
[What this actually enables GIVEN the tools/permissions/data this specific agent has access to]

### Trajectory
[turn-by-turn evidence]

### Root Cause
[trust boundary crossed]

### Remediation
[e.g., input/output trust segmentation, tool permission scoping, human-in-the-loop gating for consequential actions]

Separating technical severity from deployment impact avoids both over- and under-stating risk — the same injection technique can be a non-issue in a read-only agent and critical in one with write/send/purchase tools.

See references/ai-severity-model.md for the full severity rubric.

4. Reusing Existing Taxonomy Work

If prior work already defines an attack-algorithm class or predicate taxonomy for this kind of engagement, reuse that taxonomy directly for consistency rather than inventing a parallel one per engagement.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.