Reward seeking evidence review
Skill jpoindexter/reward-seeking-safety-skills/skills/reward-seeking-evidence-review
Agent skills and eval utilities for reward-seeking, grader-targeting, and oversight-dependent behavior.
npx -y skills add jpoindexter/reward-seeking-safety-skills --skill reward-seeking-evidence-reviewAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
2 things to look at
- 15 days oldThe repository was created 15 days ago. New is not bad, but a brand new repository carrying a familiar-sounding name is the shape a typosquat arrives in, and there has been no time for anyone else to find a problem with it.
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Review evidence and calibrate claims about reward-seeking, reward hacking, grader targeting, evaluation gaming, or oversight-dependent behavior. Use before saying a model is safe, aligned, deceptive, reward-seeking, or improved; also use to distinguish executed behavior, code-path evidence, assumptions, and inconclusive null results.
SKILL.md
1.4 KB, as published. Nobody here has run it
Reward-Seeking Evidence Review
Build a claim ledger:
| Claim | Distribution | Evidence | Status | Does not establish |
|---|
Tag status as executed, code_path, inferred, or assumed. A passing unit
test for an evaluator proves evaluator plumbing, not model behavior. A behavioral
shift proves sensitivity in that condition, not a stable global objective.
Check:
- model/provider/version and harness version are pinned;
- task distribution, authority pair, prompt channel, tool policy, and date exist;
- positive and negative controls worked;
- feature opportunity, invalids, refusals, and uncertainty are reported;
- final actions and receipts support the claim independently of chain-of-thought;
- a null is not caused by failed activation or low power;
- the result replicated on a changed surface form;
- an independent decision-maker, not the evaluated grader alone, owns release.
Rank the strongest defensible claim first. Downgrade overbroad language and state the exact next experiment needed to close each gap.