agentsclimarketplace

Direct response creative testing

Skill scumunna/programmatic-skills/skills/direct-response-creative-testing

Design, run, and read direct-response creative tests across programmatic and paid-social campaigns. Use when the user asks how to test creative, compare hooks or offers, read asset-level performance, avoid false winners, manage fatigue, set a creative learning agenda, run holdouts, decide when a test is significant enough to act, or brief the next batch of conversion creative.From its SKILL.md

Install
npx -y skills add scumunna/programmatic-skills --skill direct-response-creative-testing

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

10.8 KB, ~2.3k tokens by cl100k_base, as published. Nobody here has run it

Direct-response creative testing

Run creative tests that help a trader improve a campaign without pretending the ad platform is a lab. The job is to decide what to test, hold enough variables still, read results against the business KPI, and turn the learning into the next creative brief. This skill helps programmatic operators make better campaign decisions. It does not replace the operator, approve creative, or change live budgets on its own.

This is the cross-platform creative testing playbook. For platform-specific campaign build and reporting, hand off to the relevant campaign and reporting skills. For lift testing and holdout math, use incrementality-and-experimentation.

When to use this skill

  • "Which creative won?" / "Can I kill this ad?"
  • "How do I test hooks, offers, CTAs, formats, or landing pages?"
  • "My CPA improved but CTR fell. What does that mean?"
  • "How many creatives do I need for a real test?"
  • "How do I read asset-level reporting in PMax, Smart+, Advantage+, or a DSP?"
  • "The platform says this asset is best. Should I trust it?"
  • "How do I avoid creative fatigue?"
  • "Build a creative learning agenda for a direct-response campaign."

Boundaries with sibling skills:

  • Incrementality, holdout design, sample size, and geo tests: incrementality-and-experimentation.
  • Creative trafficking and specs: dv360-creative-trafficking, dv360-video-creative-specs, ttd-creative-and-formats, stackadapt-creative-and-formats, amazon-dsp-creative-and-formats.
  • Platform campaign builds: google-ads-performance-max, meta-campaign-setup-and-optimization, tiktok-creative-and-catalog-optimization, microsoft-campaign-setup-and-optimization, or the DSP-specific campaign skills.
  • Client-facing writeups: client-deliverable-templates.

Quick reference

QuestionBetter testWhy
Which hook converts better?Same audience, same offer, same landing page, two to four hook variantsIsolates message angle
Is the offer better?Same creative shell, two offers, equal budget or a platform experimentTests economics, not design taste
Is video better than static?Format test with matched message and landing pageSeparates format lift from message lift
Is this new concept incremental?Holdout, conversion lift, or geo liftPlatform-reported ROAS can over-credit easy conversions
Is an asset fatigued?Time-series read on frequency, CPA, CVR, CTR, and win rateA tired creative usually shows declining response at rising frequency
Can I scale the winner?Step budget after the result clears KPI and volume checksScaling changes auction mix, so a small winner can break at scale

Core process

  1. Define the decision before launching. Write the sentence: "If creative A beats creative B on primary KPI X by Y% at volume Z, we will do W." A test without a decision rule becomes dashboard browsing.
  2. Pick one primary KPI and two guardrails. For direct response the primary KPI is usually CPA, ROAS, conversion rate, or profit per impression. Guardrails can be CTR, CPC, frequency, landing-page conversion rate, or quality score. Do not crown a winner on CTR when the campaign buys conversions.
  3. Hold the non-tested variables still. Keep audience, geo, budget, bid strategy, placement eligibility, landing page, and conversion goal as similar as the platform allows. If you change hook, offer, and landing page at once, you tested a bundle, not a creative idea.
  4. Choose the test structure that the platform can actually run. Use a native experiment when available. If not, use matched ad groups or asset groups with similar budget, traffic, and targeting. In highly automated campaigns, treat the platform's asset labels as directional and look for repeated patterns, not single-asset truth.
  5. Set minimum evidence before reading. Require enough impressions and clicks to clear delivery noise, enough conversions for the KPI to stabilize, and at least one full purchase or lead cycle if the conversion is delayed. For low-volume campaigns, read directional learning and do not claim statistical proof.
  6. Monitor delivery bias. Platforms often give more impressions to one creative before the test has enough evidence. If one variant receives most delivery, the test is no longer equal exposure; read it as an algorithm preference, not a clean head-to-head.
  7. Read the whole path. Split the funnel into impression, click, landing-page action, conversion, and value. A high-CTR creative with poor post-click conversion is often curiosity, not demand. A lower-CTR creative with higher conversion rate may be pre-qualifying the right buyer.
  8. Separate fatigue from bad creative. Fatigue is performance decay after repeated exposure. Bad creative underperforms from the start. Use frequency, recency, trend, and audience saturation before blaming fatigue.
  9. Turn the result into the next brief. Every test should produce a creative instruction: "Make three more variants using hook X, keep offer Y, shorten the first three seconds, and avoid claim Z." The point is compounding learning, not one-off winners.
  10. Human-gate spend and creative changes. Recommend what to pause, scale, or brief next. A trader or account owner approves any live budget or status change.

Decision rules and thresholds

  • Pre-declare the primary KPI. A creative that wins on CTR but loses on CPA is not the winner for a conversion campaign.
  • Do not call a winner on tiny conversion counts. If each variant has only a handful of conversions, label the read directional and keep collecting data or broaden the test.
  • Watch delivery share. If the platform gives one variant most impressions, you learned what the platform preferred under its model, not necessarily what the audience preferred under equal exposure.
  • Use lift testing for incrementality. Platform-reported conversion lift from a new creative can be retargeting capture or attribution bias. For budget-moving decisions, validate with holdout or geo lift.
  • Keep budget changes modest after a win. Scaling a creative changes auction composition and frequency. Step up, then re-read CPA or ROAS before declaring the winner scalable.
  • Rotate before saturation. Rising frequency plus falling CTR or CVR plus rising CPA is a creative-refresh signal. Refresh the hook or offer before the line item collapses.
  • Protect the brand and legal review. No automated system should generate or launch unapproved claims in regulated categories.

Reading asset-level reports in automated campaigns

Automated campaigns such as Performance Max, Smart+, Advantage+, and Microsoft PMax often mix assets, audiences, and placements. Their asset reports are useful, but they are not clean randomized experiments.

  • Treat "best" and "low" labels as model feedback, not final truth.
  • Look for repeated patterns across asset groups, audiences, platforms, and weeks.
  • Compare creative concepts more than individual files when the platform recombines assets.
  • Use channel or placement reporting when available to avoid comparing a YouTube-heavy asset against a Search-heavy one.
  • When the next budget decision is large, run a holdout or matched experiment instead of trusting asset labels alone.

Test design menu

TestUse it whenHold still
Hook testYou need to find the strongest opening messageOffer, landing page, audience, format
Offer testYou need to know whether incentive or positioning changes demandAudience, format, landing page
Format testYou need static vs video vs native vs CTV evidenceMessage, offer, landing page
Landing-page paired testCreative changes promise different intentAudience, offer economics
Sequential refreshVolume is too low for parallel cellsFlight timing, budget, audience, reporting window
Platform experimentThe platform supports clean split testingNothing outside the experiment setup
Holdout or geo liftYou need incremental proofMarket selection, baseline, conversion source

Templates and examples

Creative learning agenda:

Business goal: Acquire new paid subscribers at target CPA below $85.
Primary KPI: CPA.
Guardrails: landing-page conversion rate above 3.5%, frequency below 4 per 7 days.
Hypothesis 1: Pain-point hook beats product-feature hook for cold prospecting.
Test cells: 4 videos, same offer, same landing page, same audience, equal launch budget.
Minimum read: 20,000 impressions per cell and 30 conversions total, then directional read.
Action rule: If one concept beats the next best by 20%+ CPA with guardrails intact, brief 3 more variants of that concept and step budget by 20%.
Human gate: Trader approves any pause, budget increase, or claim change.

Funnel read:

Variant A: CTR high, CVR low, CPA high. Likely curiosity click or mismatched promise.
Variant B: CTR lower, CVR high, CPA low. Better qualified traffic.
Decision: Keep B as the conversion lead, rewrite A's landing-page promise before retesting.

Fatigue read:

Signal: 7-day frequency rose from 2.1 to 5.8, CTR fell 31%, CVR fell 18%, CPA rose 42%.
Decision: Refresh hook and first frame, keep offer and landing page, cap or rotate the tired asset after approval.

Common pitfalls

  • Optimizing to CTR when the campaign buys CPA or ROAS.
  • Changing audience, bid strategy, landing page, and creative in the same "test."
  • Calling platform delivery preference a clean winner.
  • Pausing a slow-starting creative before delayed conversions arrive.
  • Declaring creative fatigue without checking frequency and audience saturation.
  • Reading retargeting and prospecting creative together.
  • Letting generated creative or claims go live without human and legal review.
  • Reporting a "winner" without saying what will be briefed, paused, scaled, or retested next.

Sources

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 326,736. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.