Run2 simpo implementation
Detailed SimPO implementation logic covering reward formulation and margin-based loss computation.From its SKILL.md
npx -y skills add cxcscmu/SkillLearnBench --skill run2_simpo-implementationAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
SKILL.md
0.5 KB, 105 tokens by cl100k_base, as published. Nobody here has run it
SimPO Loss Definition
SimPO reward formulation:
chosen_rewards = beta * log_prob_chosenrejected_rewards = beta * log_prob_rejected
Loss for sigmoid-based SimPO (Equation 6):
loss = -log_sigmoid((chosen_rewards - rejected_rewards - gamma) / beta)
- Ensure
betais set to control the reward scale. - The
gamma_beta_ratiocorresponds togamma/betain the paper implementation.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.