agentsclimarketplace

Askit build samples

Skill product-on-purpose/agent-skills-toolkit/skills/askit-build-samples

Toolkit and Standard for building, grading, and scaling cross-agent skill libraries (Claude Code + Codex) to a tiered Bronze/Silver/Gold quality bar.

Install
npx -y skills add product-on-purpose/agent-skills-toolkit --skill askit-build-samples

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Creates and validates a skill's sample sets and eval sets (golden examples, anti-examples, and triggering cases) and detects drift against current behavior. Use when generating samples for a skill, building an eval set, or checking samples for drift.

SKILL.md

2.6 KB, as published. Nobody here has run it

askit-build-samples

Purpose

Author and validate the evidence a skill carries, following the builder pattern (../../docs/reference/builder-pattern.md). create generates a skill's samples (at least 3 golden examples plus at least 1 anti-example, Standard sec 7.2) under examples/, and its triggering eval set (at least 20 {query, should_trigger} cases, sec 8.3) plus any chain/hook behavior cases under evals/ in the eval-set format the G3 library-regression check and the askit-evaluate behavioral mode consume. validate detects drift: a sample or eval that no longer matches the skill's current behavior is an error, not a silent staleness. Format and the example-threads convention are in references/samples-format.md.

When to use

When generating samples or an eval set for a skill, or checking existing samples and evals for drift after a behavior change.

create mode

  1. Read the target skill (its description, triggers, and behavior).
  2. Generate examples/ golden samples (>= 3 realistic input/output pairs) and an anti-example (>= 1 case the skill should NOT handle), per sec 7.2.
  3. Generate evals/<name>.eval.json: a triggering set of >= 20 {query, should_trigger} cases (fires when it should, silent when it should not, sec 8.3), plus {given, expect} behavior cases for any chain the skill participates in (so the G3 check finds coverage).

validate mode

  1. Re-run the samples and evals against the skill's current behavior.
  2. Flag drift: a golden sample whose output changed, an anti-example that now triggers, or an eval whose expectation no longer holds. Drift is an error so samples stay honest (sec 7.2, 8.3).

Scope

Samples and eval sets are the evidence layer. The deterministic G3 baseline (presence + the regression signal) is enforced by library-regression; behavioral judging of the cases is the opt-in askit-evaluate behavioral mode (delegated to askit-quality-grader), never the CI gate (Design Principle 3). Example-threads (the bounded validation triad of ADR 0021: a greenfield Bronze plugin, the pm-skills adopt-and-grade thread, and the toolkit itself as Gold) anchor samples to real end-to-end arcs rather than isolated snippets.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.