Simpo loss
Skill cxcscmu/SkillLearnBench/skills/b4-skill-creator-claude-opus-4-6/nlp-paper-reproduction/simpo-loss
Implement the SimPO (Simple Preference Optimization) loss function from the NeurIPS 2024 paper. Use this skill whenever implementing SimPO, preference optimization losses, or DPO variants that use length-normalized rewards and reward margins.From its SKILL.md
npx -y skills add cxcscmu/SkillLearnBench --skill simpo-lossAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
SKILL.md
2.0 KB, 508 tokens by cl100k_base, as published. Nobody here has run it
SimPO Loss Implementation
Core Formula (Eq. 6 from the paper)
The SimPO objective:
L_SimPO(πθ) = -E log σ( (β/|yw|) log πθ(yw|x) - (β/|yl|) log πθ(yl|x) - γ )
Where:
β(beta): scaling factor for reward difference (default: 2.0)γ(gamma): target reward margin, computed asgamma_beta_ratio * beta(default ratio: 0.25)|yw|,|yl|: lengths of winning/losing responsesσ: sigmoid function
Key Design Decisions
- Length-normalized reward:
r(x,y) = (β/|y|) * log πθ(y|x)— the average log probability scaled by β - Target reward margin γ: ensures winning response reward exceeds losing by at least γ
- No reference model needed — unlike DPO
Implementation Pattern
Given policy_chosen_logps and policy_rejected_logps (already length-normalized average log probs):
pi_logratios = policy_chosen_logps - policy_rejected_logps
gamma = self.gamma_beta_ratio * self.beta
logits = pi_logratios * self.beta - gamma
# Sigmoid loss (with optional label smoothing)
losses = (
-F.logsigmoid(logits) * (1 - self.label_smoothing)
- F.logsigmoid(-logits) * self.label_smoothing
)
# Hinge loss variant
# losses = torch.relu(1 - logits)
chosen_rewards = self.beta * policy_chosen_logps.detach()
rejected_rewards = self.beta * policy_rejected_logps.detach()
Config Parameters (SimPOConfig)
| Parameter | Default | Description |
|---|---|---|
| beta | 2.0 | Reward scaling factor |
| gamma_beta_ratio | 0.25 | γ/β ratio |
| label_smoothing | 0.0 | Label smoothing factor |
| loss_type | "sigmoid" | "sigmoid" or "hinge" |
| sft_weight | 0.0 | Optional SFT loss weight |
Returns
Tuple of (losses, chosen_rewards, rejected_rewards) — all tensors of shape (batch_size,).
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.