Easyjailbreak pair attack pipeline
Skill kjuhwa/skills-hub/skills/security/easyjailbreak-pair-attack-pipeline
Use EasyJailbreak to run the PAIR attack against a target LLM using an attack model, target model, and evaluator.From its SKILL.md
npx -y skills add kjuhwa/skills-hub --skill easyjailbreak-pair-attack-pipelineAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
1.9 KB, 377 tokens by cl100k_base, as published. Nobody here has run it
EasyJailbreak PAIR Attack Pipeline
When to use
- Researching LLM safety via automated jailbreak attacks.
- PAIR: attack LLM iteratively refines jailbreak prompts against a target LLM.
- Reproducible framework with Selector, Mutator, Constraint, Evaluator modules.
Steps
- Install EasyJailbreak:
pip install easyjailbreak
- Load models:
from easyjailbreak.models.huggingface_model import from_pretrained, HuggingfaceModel
attack_model = from_pretrained('lmsys/vicuna-13b-v1.5', model_name='vicuna_v1.1')
target_model = HuggingfaceModel('meta-llama/Llama-2-7b-chat-hf', model_name='llama-2')
- Load dataset and initialize seeds:
from easyjailbreak.datasets import JailbreakDataset
from easyjailbreak.seed.seed_random import SeedRandom
dataset = JailbreakDataset(dataset='AdvBench')
seeder = SeedRandom()
seeder.new_seeds()
- Run PAIR attack:
from easyjailbreak.attacker.PAIR_chao_2023 import PAIR
attacker = PAIR(
attack_model=attack_model,
target_model=target_model,
eval_model=eval_model,
jailbreak_datasets=dataset
)
attacker.attack(save_path='results.jsonl')
Pitfalls
- PAIR requires three models (attack, target, eval); GPU memory scales accordingly.
- GPT-4 as eval model requires OpenAI API access (VPN in China).
Source
- Chapter 6 of dive-into-llms - documents/chapter6/README.md
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.
Gives 0 of the 12 instructions most security skills give in 377 tokens
Counted across 648 of the 828 authors here whose files we hold, read 2026-08-07
- Parameterize all database queriesin 68 of 648, across 51 files
- Hash passwords using bcrypt, scrypt, or argon2in 49 of 648, across 36 files
- Apply rate limiting to authentication endpointsin 48 of 648, across 24 files
- Configure security headersin 35 of 648, across 19 files
- Validate all inputsin 32 of 648, across 24 files
- Validate all external input at the system boundaryin 29 of 648, across 19 files
- Run containers as a non-root userin 28 of 648, across 15 files
- Use httponly secure samesite cookies for sessionsin 26 of 648, across 15 files
- Run dependency audits before every releasein 21 of 648, across 10 files
- Encode output to prevent cross-site scriptingin 21 of 648, across 11 files
- Copy dependencies before source codein 20 of 648, across 9 files
- Store secrets in environment variablesin 20 of 648, across 18 files
Said here and by no other author read
- install easyjailbreak
- load attack target and eval models
- load jailbreak dataset and generate seeds
- initialize the PAIR attacker
- run the attack
- save results to a jsonl file
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.