Rlhf systems
Skill a5c-ai/babysitter/library/specializations/data-science-ml/skills/rlhf-systems
Human-feedback-driven model optimization — preference data collection, reward modeling, policy updates, and alignment evaluation.From its SKILL.md
npx -y skills add a5c-ai/babysitter --skill rlhf-systemsAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
SKILL.md
0.4 KB, 12 tokens by cl100k_base, as published. Nobody here has run it
RLHF Skill
Stub — implementation pending.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.