Simpo loss
Skill cxcscmu/SkillLearnBench/skills/b1-one-shot-gemini-3.1-pro-preview/nlp-paper-reproduction/simpo-loss
Guide on understanding and implementing the SimPO (Simple Preference Optimization) loss function for language models.From its SKILL.md
npx -y skills add cxcscmu/SkillLearnBench --skill simpo-lossAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
SKILL.md
1.2 KB, 290 tokens by cl100k_base, as published. Nobody here has run it
SimPO Loss
SimPO (Simple Preference Optimization) is an alternative to DPO for preference learning. It optimizes the reward directly on the sequence level without requiring a reference model, using the average log probability of the generated sequence.
Core Formula
The SimPO loss involves:
- Calculating the average log probability per token for the chosen and rejected sequences.
- Adding a target reward margin $\gamma$ (gamma).
- Passing the difference through a sigmoid function (specifically,
-log(sigmoid(diff))).
Formula: $$L_{\text{SimPO}} = - \log \sigma \left( \beta \left( \frac{1}{|y_w|} \sum_{i} \log p(y_{w,i}|x) - \frac{1}{|y_l|} \sum_{i} \log p(y_{l,i}|x) \right) - \gamma \right)$$
where:
- $\beta$ is a scaling factor.
- $\gamma$ is the reward margin.
- $y_w$ is the chosen/winning sequence.
- $y_l$ is the rejected/losing sequence.
Usage
In a trainer like DPOTrainer, this can be implemented by computing the average log probability for both policy and reference (though SimPO only needs the policy logprobs, they might be formatted similar to DPO).
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.