Measure experiment design
Skill product-on-purpose/pm-skills/skills/measure-experiment-design
Designs an A/B test or experiment with variants, success metrics, sample size, and duration for an existing hypothesis. Use when planning an experiment to validate a product change or test an assumption you have already framed. To articulate the hypothesis itself first, use define-hypothesis.From its SKILL.md
npx -y skills add product-on-purpose/pm-skills --skill measure-experiment-designAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
What its file declares
Copied from the file, not written here
The file declares its own license as Apache-2.0. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.
SKILL.md
4.2 KB, 765 tokens by cl100k_base, as published. Nobody here has run it
Experiment Design
An experiment design document defines all parameters needed to run a rigorous A/B test or controlled experiment. It ensures the team aligns on what you're testing, how you'll measure success, and how long to run the test before drawing conclusions. Good experiment design prevents common pitfalls: underpowered tests, unclear success criteria, and decisions based on noise rather than signal.
When to Use
- Before launching an A/B test to validate a product change
- When testing a hypothesis that requires quantitative validation
- After solution design to validate assumptions before full rollout
- When stakeholders want data-driven evidence for a decision
- To establish a culture of experimentation and learning
When NOT to Use
- The hypothesis itself is not yet articulated -> use
define-hypothesisfirst; this skill designs the test for a claim you already have - You are analyzing a completed experiment -> use
measure-experiment-results - You need the event tracking that will measure the experiment -> use
measure-instrumentation-spec - You are gathering opinions rather than running a controlled test -> use
measure-survey-analysis
Instructions
When asked to design an experiment, follow these steps:
-
Articulate the Hypothesis Write a clear, testable hypothesis in the format: "We believe [change] for [users] will [outcome] as measured by [metric]." One hypothesis per experiment - if you're testing multiple things, run multiple experiments.
-
Define the Variants Describe the control (current experience) and treatment (new experience) in sufficient detail. Include screenshots, mockups, or precise descriptions so anyone can understand what users will see.
-
Choose Primary and Secondary Metrics Select one primary metric that will determine success or failure. Add 2-3 secondary metrics to understand the broader impact. Include guardrail metrics to catch unintended negative effects.
-
Calculate Sample Size Determine how many users you need per variant to detect your minimum detectable effect (MDE) with statistical significance. Specify your significance level (typically 0.05) and power (typically 0.80).
-
Estimate Duration Based on sample size and available traffic, calculate how long the experiment needs to run. Account for weekly patterns - avoid ending mid-week if behavior varies by day.
-
Define Targeting and Allocation Specify which users are eligible for the experiment and how traffic is split between variants. Document any exclusions (e.g., employees, specific segments).
-
Set Success Criteria Define upfront what constitutes a win, a loss, or an inconclusive result. This prevents post-hoc rationalization and moving goalposts.
-
Document Risks and Mitigations Identify what could go wrong and how you'll detect/address it. Include monitoring plans and rollback criteria.
Output Format
Use the template in references/TEMPLATE.md to structure the output. A complete design fills every template section: Overview; Hypothesis; Background; Variants; Metrics; Sample Size & Duration; Audience Targeting; Success Criteria; Risks & Mitigations; Implementation Notes; and References.
Quality Checklist
Before finalizing, verify:
- Hypothesis is falsifiable and specific
- Only one primary metric is defined
- Sample size calculation is documented with assumptions
- Duration accounts for traffic patterns and statistical requirements
- Success criteria are defined before the experiment starts
- Guardrail metrics protect against unintended harm
Examples
See references/EXAMPLE.md for a completed example.
What ships with it: 6 files
21.4 KB alongside SKILL.md
evals/
references/
- EXAMPLE.md7.1 KB
- TEMPLATE.md4.3 KB
- HISTORY.md1.1 KB
Gives 0 of the 12 instructions most analytics metrics skills give in 765 tokens
Counted across 333 of the 342 authors here whose files we hold, read 2026-09-06
- Read product marketing context before asking questionsin 37 of 333, across 16 files
- Test one variable at a timein 26 of 333, across 11 files
- Pre-determine sample size before launchin 24 of 333, across 16 files
- Verify tracking and QA variants before launchin 17 of 333, across 8 files
- Monitor for technical issues during the testin 14 of 333, across 6 files
- Match each save offer to the cancel reasonin 14 of 333, across 5 files
- Start every test with a specific hypothesisin 14 of 333, across 7 files
- Keep the continue-cancelling option visiblein 13 of 333, across 4 files
- Document every test with hypothesis, variants, results, and learningsin 13 of 333, across 6 files
- Gather churn, billing, product, usage, and constraint context firstin 12 of 333, across 3 files
- Build a health score from weighted signalsin 12 of 333, across 3 files
- Retry soft declines 3-5 times over 7-10 daysin 12 of 333, across 3 files
Said here and by no other author read
- Describe control and treatment variants in detail
- Select a single primary success metric
- Specify significance level and statistical power
- Define user targeting and traffic allocation
- Set success criteria before the experiment starts
- Document risks, monitoring plans, and rollback criteria
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.