agentsclimarketplace

Ablate refusal representations

Skill gaelic-ghost/socket/plugins/model-lab-skills/skills/ablate-refusal-representations

The Source for macOS Agent Workflows

Install
npx -y skills add gaelic-ghost/socket --skill ablate-refusal-representations

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 6 stars6 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Reproduce and evaluate refusal-direction ablation on controlled checkpoints. Use when estimating refusal representations, applying activation or weight orthogonalization, measuring behavior change, or comparing ablated artifacts.

SKILL.md

2.2 KB, as published. Nobody here has run it

Ablate Refusal Representations

Research Boundary

Treat ablation as controlled model-internals research. Use models and evaluation targets the operator is authorized to modify and test. Do not turn a checkpoint experiment into unapproved probing of a third-party production service.

Workflow

  1. Invoke design-model-experiment; state the mechanism claim and the behavior claim separately.
  2. Pin the base checkpoint and reproduce its baseline refusal, compliance, capability, and safety behavior.
  3. Build matched harmful/harmless or refusal/compliance contrast sets with held-out topics and surface-form controls.
  4. Estimate candidate directions across declared layers and token positions; retain searched candidates, not only the winner.
  5. Test reversible activation ablation before persistent weight orthogonalization when the hypothesis permits it.
  6. Compare zero, sign-reversed, random norm-matched, unrelated-concept, and shuffled-label interventions.
  7. Measure refusal reduction on the target set plus benign compliance, general capability, calibration, fluency, and hazardous-behavior guardrails. Invoke evaluate-jailbreak-resilience whenever the claim includes harmful compliance or safety resilience.
  8. Preserve the base checkpoint as immutable, write every weight edit to a new artifact, and record both checksums plus the exact transformation; never overwrite the base weights in place.
  9. Evaluate the exact saved artifact after any weight edit, merge, quantization, or conversion.
  10. Use compare-model-checkpoints and report collateral behavior changes with the headline result.

Evidence Discipline

Reducing refusal strings is not sufficient: distinguish shallow response-format changes from increased task completion. A direction that generalizes only to the extraction template or searched topics is not a general refusal mechanism.

References

Read references/refusal-ablation.md for the source research and required controls.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.