Simpo loss implementation
Guidelines for implementing the SimPO (Simple Preference Optimization) loss function for NLP model training. Use this skill when modifying trainers or implementing preference-based loss functions in transformer projects.From its SKILL.md
npx -y skills add cxcscmu/SkillLearnBench --skill simpo-loss-implementationAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
SKILL.md
1.3 KB, 256 tokens by cl100k_base, as published. Nobody here has run it
SimPO Loss Implementation
The SimPO loss is designed to optimize language models based on preference data without a separate reward model.
Implementation Principles
- Mathematical correctness: Implement the log-likelihood difference with the margin, normalized by the length of the sequences.
- Numerical stability: Use log-sum-exp or similar techniques if necessary to avoid overflow.
- Loss formula:
- For a sequence pair (chosen, rejected) $y_w, y_l$:
- Reward $r(x, y) = \beta \log P_\theta(y|x)$ (simplified, check paper for exact implementation details).
- SimPO Loss = $-\log \sigma (\beta(\log P_\theta(y_w|x) - \log P_\theta(y_l|x)) - \gamma)$
- $\gamma$ is the target margin.
Checklist
- Ensure
logitsare masked properly for padding tokens. - Calculate
log_probsusing the model's output and target labels. - Normalize the
log_probsby length if the paper requires it. - Ensure the margin $\gamma$ is configurable.
- Verify the loss is averaged over the batch.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.