agentsclimarketplace

Simpo loss

Skill cxcscmu/SkillLearnBench/skills/b1-one-shot-gemini-3.1-pro-preview/nlp-paper-reproduction/simpo-loss

Guide on understanding and implementing the SimPO (Simple Preference Optimization) loss function for language models.From its SKILL.md

Install
npx -y skills add cxcscmu/SkillLearnBench --skill simpo-loss

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

SKILL.md

1.2 KB, 290 tokens by cl100k_base, as published. Nobody here has run it

SimPO Loss

SimPO (Simple Preference Optimization) is an alternative to DPO for preference learning. It optimizes the reward directly on the sequence level without requiring a reference model, using the average log probability of the generated sequence.

Core Formula

The SimPO loss involves:

  1. Calculating the average log probability per token for the chosen and rejected sequences.
  2. Adding a target reward margin $\gamma$ (gamma).
  3. Passing the difference through a sigmoid function (specifically, -log(sigmoid(diff))).

Formula: $$L_{\text{SimPO}} = - \log \sigma \left( \beta \left( \frac{1}{|y_w|} \sum_{i} \log p(y_{w,i}|x) - \frac{1}{|y_l|} \sum_{i} \log p(y_{l,i}|x) \right) - \gamma \right)$$

where:

  • $\beta$ is a scaling factor.
  • $\gamma$ is the reward margin.
  • $y_w$ is the chosen/winning sequence.
  • $y_l$ is the rejected/losing sequence.

Usage

In a trainer like DPOTrainer, this can be implemented by computing the average log probability for both policy and reference (though SimPO only needs the policy logprobs, they might be formatted similar to DPO).

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.