agentsclimarketplace

Simpo implementation

Skill cxcscmu/SkillLearnBench/skills/b1-one-shot-gemini-3.1-flash-lite-preview/nlp-paper-reproduction/simpo-implementation

Guidelines for implementing the SimPO loss function as per the SimPO paper.From its SKILL.md

Install
npx -y skills add cxcscmu/SkillLearnBench --skill simpo-implementation

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

SKILL.md

1.4 KB, 362 tokens by cl100k_base, as published. Nobody here has run it

SimPO Loss Implementation

The SimPO loss is designed to optimize language models by maximizing the gap between positive and negative response log-probabilities, normalized by the length of the response.

Mathematical Formulation

The SimPO loss $L_{SimPO}$ is defined as: $L_{SimPO} = -\mathbb{E}{(x, y_w, y_l) \sim D} \left[ \log \sigma \left( \beta \left( \frac{1}{|y_w|} \log \pi\theta(y_w|x) - \frac{1}{|y_l|} \log \pi_\theta(y_l|x) - \gamma \right) \right) \right]$

Where:

  • $y_w$ is the winning response.
  • $y_l$ is the losing response.
  • $\beta$ is the reward margin scale.
  • $\gamma$ is the reward margin threshold.
  • $\pi_\theta(y|x)$ is the model's likelihood for sequence $y$ given $x$.

Implementation Steps

  1. Calculate Log Probabilities: Use the model to compute the log-likelihood of tokens in $y_w$ and $y_l$.
  2. Sum Log Probabilities: Sum the log-likelihoods over the generated sequence.
  3. Normalize: Divide the sum by the sequence length (or use a length-normalized log-prob implementation).
  4. Compute Margin: Calculate the difference between normalized log-probs of $y_w$ and $y_l$, then subtract $\gamma$.
  5. Apply Sigmoid and Log: Apply the sigmoid function to $\beta \times \text{margin}$, take the negative log, and compute the mean over the batch.

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.