Run2 implement simpo loss
Implements the SimPO (Simple Preference Optimization) loss function with length normalization and reward margin as specified in the research paper.From its SKILL.md
npx -y skills add cxcscmu/SkillLearnBench --skill run2_implement_simpo_lossAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
SKILL.md
1.0 KB, 241 tokens by cl100k_base, as published. Nobody here has run it
- Implement the Logic: In
SimPOTrainer.simpo_losswithin/root/SimPO/scripts/simpo_trainer.py, implement the following mathematical steps:- Calculate the log probabilities for the winning (chosen) and losing (rejected) responses.
- Length Normalization: Divide the log probabilities of each sequence by its length ($L$): $p_{norm} = \frac{1}{L} \log \pi(x, y)$.
- Reward Calculation: Calculate rewards as $R = \beta \cdot p_{norm}$.
- SimPO Loss: Use the formula: $\mathcal{L}{SimPO} = -\mathbb{E}{(x, y_w, y_l)} [\log \sigma(\beta p_{norm}(y_w|x) - \beta p_{norm}(y_l|x) - \gamma)]$ where $\gamma$ is the target reward margin and $\beta$ is the scale.
- Verification:
- Import the required libraries (torch, numpy).
- Ensure the function returns the loss tensor in the format expected by the trainer.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.