agentsclimarketplace

Run1 skill 1

Skill cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-gemini-3.1-pro-preview/nlp-paper-reproduction/run1_skill-1

[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.From the repository description

Install
npx -y skills add cxcscmu/SkillLearnBench --skill run1_skill-1

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

SKILL.md

0.8 KB, 154 tokens by cl100k_base, as published. Nobody here has run it

[SKILL]

name: simpo-preference-optimization description: Detailed explanation and implementation guide for the SimPO (Simple Preference Optimization) loss function, which uses length-normalized log probabilities and a reference-free reward margin.

SimPO (Simple Preference Optimization)

SimPO is a reference-free preference optimization algorithm that simplifies alignment compared to methods like DPO (Direct Preference Optimization). It eliminates the need for a reference model by defining the reward directly using the policy model's length-normalized log probabilities.

Mathematical Formulation

In SimPO, the reward $r_\theta(y)$ for a generated response $y$ given a prompt $x$ is defined as the average log probability of the response tokens, scaled by a constant $\beta$:

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.