agentsclimarketplace

Simpo loss

Skill cxcscmu/SkillLearnBench/skills/b1-one-shot-claude-opus-4-6/nlp-paper-reproduction/simpo-loss

SimPO (Simple Preference Optimization) loss computation for LLM alignment without a reference model.From its SKILL.md

Install
npx -y skills add cxcscmu/SkillLearnBench --skill simpo-loss

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

SKILL.md

1.1 KB, 300 tokens by cl100k_base, as published. Nobody here has run it

SimPO Loss

Overview

SimPO is a reference-free preference optimization algorithm. Its key innovation is using the average log probability of a sequence as the implicit reward, plus a target reward margin γ.

Loss Formula (Eq. 6 from the paper)

L_SimPO = -E log σ(β/|yw| · log πθ(yw|x) - β/|yl| · log πθ(yl|x) - γ)

Since the log probabilities passed to simpo_loss are already length-normalized (average log prob), the loss simplifies to:

logits = β * policy_chosen_logps - β * policy_rejected_logps - γ

where γ = gamma_beta_ratio * beta.

Loss Types

  • sigmoid (default): losses = -log σ(logits) * (1 - label_smoothing) - log σ(-logits) * label_smoothing
  • hinge: losses = relu(1 - logits)

Rewards

  • chosen_rewards = β * policy_chosen_logps
  • rejected_rewards = β * policy_rejected_logps

Default Hyperparameters

  • β = 2.0
  • gamma_beta_ratio = 0.25 (so γ = 0.5)
  • label_smoothing = 0.0
  • loss_type = "sigmoid"

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.