agentsclimarketplace

Simpo loss implementation

Skill cxcscmu/SkillLearnBench/skills/b4-skill-creator-gemini-3.1-flash-lite-preview/nlp-paper-reproduction/simpo-loss-implementation

Guidelines for implementing the SimPO (Simple Preference Optimization) loss function for NLP model training. Use this skill when modifying trainers or implementing preference-based loss functions in transformer projects.From its SKILL.md

Install
npx -y skills add cxcscmu/SkillLearnBench --skill simpo-loss-implementation

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

SKILL.md

1.3 KB, 256 tokens by cl100k_base, as published. Nobody here has run it

SimPO Loss Implementation

The SimPO loss is designed to optimize language models based on preference data without a separate reward model.

Implementation Principles

  1. Mathematical correctness: Implement the log-likelihood difference with the margin, normalized by the length of the sequences.
  2. Numerical stability: Use log-sum-exp or similar techniques if necessary to avoid overflow.
  3. Loss formula:
    • For a sequence pair (chosen, rejected) $y_w, y_l$:
    • Reward $r(x, y) = \beta \log P_\theta(y|x)$ (simplified, check paper for exact implementation details).
    • SimPO Loss = $-\log \sigma (\beta(\log P_\theta(y_w|x) - \log P_\theta(y_l|x)) - \gamma)$
    • $\gamma$ is the target margin.

Checklist

  • Ensure logits are masked properly for padding tokens.
  • Calculate log_probs using the model's output and target labels.
  • Normalize the log_probs by length if the paper requires it.
  • Ensure the margin $\gamma$ is configurable.
  • Verify the loss is averaged over the batch.

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.