agentsclimarketplace

Evaluate with braintrust

Skill hiteshbandhu/skills-i-use/skills/ai-engineer-talks/evaluate-with-braintrust

Drop-in skills and plugins for your AI development workflows

Install
npx -y skills add hiteshbandhu/skills-i-use --skill evaluate-with-braintrust

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Evaluates LLM and agent products with Braintrust—offline/online evals, SDK/CI workflows, eval platform design, Zapier/Notion product ops, and Loop-style optimization. Use when the user mentions Braintrust, eval playground, autoevals, production log evals, or agent quality platforms.

SKILL.md

1.9 KB, as published. Nobody here has run it

Evaluate with Braintrust

Action playbook from seven Braintrust @ AI Engineer talks. Do not summarize — pick a workflow.

Supporting files: workflows.md · source-index.md

Related skill: run-llm-evals (cross-vendor eval theory).

Optional: ./skill-outputs/evaluate-with-braintrust/


Step 0 — Pick workflow

What is the user trying to do?
├─ Eval fundamentals (datasets, scorers, modes)        → evals-101
├─ Build/buy eval platform architecture                → platform-design
├─ Five lessons / future of evals (velocity, Loop)     → eval-ops
├─ Zapier-style product + eval integration             → product-ops
├─ Notion-class world-class AI quality bar             → world-class-products
└─ Complex app workshop (Trainline patterns)           → complex-apps

Install

cp -r skills/evaluate-with-braintrust ~/.cursor/skills/
cp -r skills/evaluate-with-braintrust ~/.codex/skills/

Source: playlists/braintrust-ai-engineer/.


Cross-cutting rules

RuleSource
Eval = data + task + scorers; offline and online[src-001 @ 0:07:12]
Prod logs must feed datasets and online evals[src-001 @ 0:13:13]
Unify offline experiments with trace replay[src-002 @ 0:14:15]
24h model swaps when evals healthy (Notion bar)[src-003 @ 0:40]
Human-in-loop on Loop optimizations[src-004 @ 0:03:33]

Output

Name workflow; artifacts under ./skill-outputs/evaluate-with-braintrust/ when requested.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.