Claude prompt optimizer
A Claude skill that audits and rewrites prompts. Validated by blind A/B test: 80% pair-level win rate, +5.37 verdict-aligned margin on a 50-point rubric.
npx -y skills add viktor-milev/claude-prompt-optimizerAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Rigorous mode prompt optimizer — audits a polished prompt and rewrites it to elicit state-of-the-art output from Claude. Auto-triages every input into OUT OF PURVIEW / BORDERLINE / IN PURVIEW and routes to a band-appropriate output path; rigorous optimization (9-dimension diagnostic, archetype detection, Claude-specific capability activation) is applied ONLY to IN PURVIEW prompts. Simple prompts are returned substantially unchanged with a triage note. Use whenever the user pastes a prompt and asks to optimize, improve, rewrite, restructure, sharpen, audit, score, or diagnose it. Trigger phrases include "optimize this prompt", "audit this prompt", "score this prompt", "diagnose this prompt", "PO this", "run this through the optimizer", "sharpen this prompt", "level this prompt up". Best for prompts that deserve rigorous evaluation and a full diagnostic. For messy dictated input or flow-state work, use prompt-optimizer-flow instead. Do NOT trigger when the user wants the prompt EXECUTED rather than optimized.
SKILL.md
38.0 KB, ~8.1k tokens by cl100k_base, as published. Nobody here has run it
Prompt Optimizer (State-of-the-Art Edition)
CHANGELOG
v3 — Apr 2026. Fixes three real bugs surfaced by the 20-prompt A/B test (see full_results.json, Apr 24 2026):
- Signal preservation (new Step 4.5). The single largest skill-level failure mode in v2 was structural overwrite of raw-prompt signals. The optimizer wrapped unfilled placeholders (
[CV],[JD]) as if material were present; converted ambiguous hints ("I have a draft to share") into locked workflows; and dropped explicit format directives ("both analyses side by side") during restructuring. Step 4.5 enumerates every signal in the raw prompt before any restructuring and requires each to be preserved or explicitly overridden with rationale. - Feasibility pass (revised Step 6, replacing the old Ceiling Check). The old step asked only "what could we add?" It never asked "what would crash the turn if we added it?" The revised step runs four explicit checks — single-turn output budget, context-dependency honesty, user-format preservation, input-presence — and is allowed to cut the optimization rather than only extend it.
- Calibrated capability activation (Section 4).
[Confirmed]/[Inferred]confidence tags now require explicit source material in the prompt.<research_activation>now explicitly degrades to "state what you don't know" when no search tool is present. These two patterns drove the only two "won on rubric, lost on verdict" cases in the A/B test — both caused Claude to fabricate specifics with false confidence markers.
Issues surfaced in the same A/B test that are not skill bugs and are flagged for separate work:
- avg_margin aggregation sign-mismatch in the test harness (margin uses rubric totals; pair winner uses the judge's verdict field — they disagreed on career_cv and research_synthesis_banks, inflating reported margin from ~+0.3 to +3.15).
- Parser-level truncation of optimized prompts for
framework_application_private_debtandagentic_research_watchlist(captured only the output-format tail; the -35 and -11 "losses" are extraction failures, not skill output). - Per-category n=1 statistical weakness; judge length-bias (r ≈ 0.49 between response length and depth score).
Optimize prompts to reach the ceiling of what Claude can produce on a given task — but ONLY when the task warrants it. Rigorous optimization on a prompt that doesn't need it produces worse output, not better. The skill's first job is to decide whether to optimize at all.
Core philosophy
Stop optimizing prompts. Start optimizing the Claude session that the prompt initiates.
A prompt is not a request. It is the keystone of an interaction. The job of optimization is not to make the prompt well-formed — it is to make Claude perform at the top of its distribution on the task the prompt describes.
But this logic only applies when the task has a ceiling worth reaching. A PTO email does not. Forcing XML scaffolding, capability activation, and 6-section diagnostic outputs onto casual communication tasks makes the downstream output worse, not better — it bloats the prompt, buries the actual ask, and ships Claude a bureaucratic commission when the user wanted a 3-sentence note. This is a SEVERE failure mode (see Failure Modes section).
The skill therefore operates in three bands, triaged before any optimization work begins.
A second, equally important philosophy is added in v3: preserve the user's signals; restructure only what the user didn't specify. The raw prompt is the user's compressed intent. Every format directive, placeholder, ambiguity marker, and workflow hint is a signal. Optimization that overwrites those signals produces a prompt Claude follows perfectly — to the wrong target.
Step 0 — TRIAGE (mandatory; runs before every other step)
Classify the input into one of three bands using the four tests below. State the verdict explicitly in one line before proceeding. The user may override (e.g., "treat as IN PURVIEW") if they disagree.
The four triage tests
Run all four. Score each with OUT / BORDERLINE / IN leaning.
Test 1 — Artifact type. What is the downstream output?
- OUT leaning: email, text message, Slack message, tweet/post, caption, subject line, short note, simple list, simple lookup, simple formatting task, casual personal communication
- BORDERLINE leaning: article, blog post, cold outreach, summary of provided material, standard business document (meeting agenda, status update), recipe / single-workout plan, single-topic explainer
- IN leaning: report, memo, analysis, evaluation, research synthesis, strategic document, session keystone, multi-constraint creative work, agentic task spec, decision framework, published artifact intended for external stakeholders
Test 2 — Raw prompt word count.
- <30 words: OUT leaning
- 30–100 words: BORDERLINE leaning
-
100 words: IN leaning (content usually dominates, but size correlates)
Test 3 — Analytical load. Does the task require reasoning, decomposition, research, evaluation, or synthesis?
- None (pure generation, formatting, lookup): OUT leaning
- Light (single narrow judgment, e.g., "make this tone warmer"): BORDERLINE leaning
- Present (multi-step reasoning, tradeoff analysis, research integration, evaluation against criteria): IN leaning
Test 4 — Consequentiality. What is the downstream cost of a mediocre output?
- OUT leaning: Low-consequence communication — casual content, familiar or non-specific audience, and a mediocre version still accomplishes the task (PTO email to known boss, tweet about your day, text to a friend, routine Slack update)
- BORDERLINE leaning: Externally-facing or reputationally-significant artifact — first-impression content, cold outreach, professional networking, short business documents where tone and specificity materially affect whether the artifact achieves its purpose. The user sends it once, but getting it wrong has real professional cost.
- IN leaning: Session keystone, published document, work product shared with stakeholders, or anything that will be iterated on
Verdict rules
- 3 or 4 OUT leanings → OUT OF PURVIEW. Short-circuit. Do not run the 9-dim diagnostic.
- 3 or 4 IN leanings → IN PURVIEW. Run the full workflow below.
- Mixed or 3+ BORDERLINE leanings → BORDERLINE. Run abbreviated workflow.
- Override rule: if analytical load is PRESENT (a real reasoning task), promote at least one band up regardless of other scores. "Should I take this job?" at 8 words is IN PURVIEW, not OUT.
- Override rule: if the user explicitly flags the prompt as a keystone or high-stakes ("this kicks off my session", "this goes to my CIO", "this is the prompt for my automated pipeline"), promote to IN PURVIEW.
Output the verdict in one line
Before any further work, state:
TRIAGE: [OUT OF PURVIEW | BORDERLINE | IN PURVIEW] — [one-sentence rationale citing the two or three tests that drove the call]
If the user disagrees, they can override in their next message. Proceed to the band-appropriate output path below.
Band-specific output paths
Path A — OUT OF PURVIEW output format
Return exactly this, and nothing else:
TRIAGE: OUT OF PURVIEW — [rationale]
Verdict: this prompt is already well-calibrated for its task. Rigorous optimization would over-engineer it and degrade the downstream output.
Surgical note (optional, only if genuinely missing): [one sentence naming the single missing element, if any — typically a date, audience, or length spec. If nothing is missing, write "None — send as is."]
That is the entire response. Do NOT add a diagnostic table, archetype detection, XML scaffolding, change log, ceiling check, or use case guidance. Doing so defeats the purpose of the triage.
Self-check before finalizing Path A. If your surgical note names 2 or more distinct gaps (e.g., "you need to specify X, Y, and Z"), the triage miscalled. A prompt with 2+ real gaps is BORDERLINE by the skill's own definition. Re-classify and route to Path B. Path A is only valid when the prompt is genuinely one-gap-or-less from ready.
Path B — BORDERLINE output format
Return exactly three sections:
SECTION 1 — TRIAGE & DIAGNOSIS One line of triage verdict. Then 2–4 bullets naming the specific gaps worth closing (typically: audience, output format, length, tone, one missing constraint). Do NOT run the full 9-dimension scoring — it's theater at this band.
SECTION 2 — OPTIMIZED PROMPT The original prompt with surgical additions — typically 1 to 3 added sentences or constraints, inline. NO XML scaffolding. NO capability activation layer (no reasoning scaffolds, no anti-sycophancy permissions, no uncertainty flagging, no working principles). If the original works as a paragraph, it stays a paragraph. Copy-paste ready.
SECTION 3 — CHANGE LOG Two to four bullets. Each names what changed and why, in plain language. No diagnostic dimension tags — they don't earn their place here.
No ceiling check, no use case guidance. BORDERLINE outputs aren't the kind of thing that has a ceiling.
Path C — IN PURVIEW output format (the full rigorous treatment)
Proceed through Steps 1–8 below, producing the seven-section output described under "IN PURVIEW output format" at the end of this document.
IN PURVIEW workflow (Steps 1–8)
This is the full rigorous treatment. Apply ONLY when Step 0 returned IN PURVIEW.
1. Intent extraction
Before changing anything, state in 2–3 sentences: what this prompt is trying to accomplish, who would use it, and what a successful output looks like. This is the north star — every optimization must serve this intent.
2. Archetype detection
Identify which archetype the prompt belongs to. Different archetypes require different optimization patterns:
- Keystone prompt — the first prompt in a new chat or project. Sets persistent context, working norms, and defines the relationship for the entire session. Needs the heaviest optimization treatment: role depth, working principles, output norms, escalation paths.
- One-shot task — a single self-contained request. Needs tight scope, clear output format, and constraints. Lighter optimization. (Note: most one-shot tasks are triaged out by Step 0. If you're here, the one-shot is genuinely complex.)
- Iterative refinement — a prompt designed to be run repeatedly with variable input (templates, generators, evaluators). Needs strong input placeholders, reusability, and consistency mechanisms.
- Agentic task — a prompt that initiates multi-step work involving tools, research, or extended reasoning. Needs explicit reasoning scaffolds, tool permission, failure recovery, and progress checkpoints.
- Creative generation — prompts for writing, ideation, or aesthetic output. Needs voice calibration, anti-generic constraints, and freedom-of-form preservation.
State the archetype explicitly. If the prompt straddles multiple, name the primary one and note the secondary.
3. Diagnostic (9 dimensions)
Score the original prompt on these nine dimensions (1–10 each). For any dimension scoring below 7, identify the specific deficiency:
Structural dimensions (the foundation):
- Clarity — Are instructions unambiguous? Could two different LLMs interpret this the same way?
- Structure — Is information organized logically with clear sections? Are steps sequenced by dependency?
- Constraints — Are boundaries defined? Does the prompt prevent common failure modes (filler, hedging, scope creep, generic output)?
- Output format — Is the expected response shape specified? Would the user know what "done well" looks like?
- Specificity — Are instructions concrete enough to act on? Could "improve X" be replaced with "do Y to achieve X"?
Capability dimensions (the ceiling):
- Reasoning depth — Does the prompt elicit analytical thinking or just surface-level completion? Are there explicit mechanisms forcing genuine reasoning (decomposition, comparison, evaluation, self-critique)?
- Capability activation — Does the prompt use Claude-specific techniques to reach deeper modes? Reasoning scaffolds, anti-sycophancy permissions, uncertainty flagging, research/search activation where appropriate?
- Context economy — Does the prompt use the context window strategically? No bloat, no redundancy, but also no false economy that strips needed context?
- Calibration — Does the prompt's complexity match the actual difficulty of the task? Heavy scaffolding on a trivial task is waste. Light scaffolding on a hard task is failure. (If you're here, Step 0 already confirmed the task is non-trivial. This dimension is now about matching scaffolding to the specific complexity, not deciding whether scaffolding belongs at all.)
4. Capability activation layer
4.7 calibration note. Claude 4.7 has raised the floor on three techniques below — reasoning activation, anti-sycophancy, and research activation. The model now decomposes on genuinely hard analytical tasks by default, pushes back more readily on weak premises, and searches more aggressively on present-tense factual questions. This does not make these techniques obsolete. It makes their application more selective:
- Reasoning activation now earns its place primarily when the task's surface difficulty understates its actual difficulty — analytical questions disguised as simple ones, strategic decisions phrased as lookups, tradeoff problems framed as "just tell me which." The default decomposition behavior covers obviously-hard tasks; the scaffold covers tasks where Claude might otherwise answer shallowly.
- Anti-sycophancy permissions are now a reinforcement of default behavior rather than a reversal of it. Still valuable on evaluation, critique, and decision-support tasks — especially where the user's framing is subtly flawed in ways the model might work around rather than flag. Effect size is smaller than it was on 4.5 and earlier.
- Research activation now earns its place when the prompt's phrasing wouldn't naturally trigger search — "what should I think about X" vs. "what is the latest on X," or tasks where the user would benefit from current grounding but hasn't signaled it. Where the question is already present-tense factual, the model will search anyway.
The other capability techniques — epistemic calibration, pushback permission, working principles, self-critique — are unchanged in their value. Apply all techniques where they earn their place, not where they're traditionally expected.
This is where most prompt optimizers fall short. Apply these techniques where the archetype and task warrant them. Do not apply them mechanically — apply them where they earn their place.
Reasoning activation (for analytical, strategic, or complex tasks):
- Add a "think before responding" instruction with explicit decomposition steps
- Require Claude to identify and consider alternatives before committing to a recommendation
- For high-stakes outputs, add a self-critique pass: "Before finalizing, identify the strongest objection to your own answer and address it"
- Add chain-of-verification for factual or numerical claims
Anti-sycophancy permissions (for any prompt seeking analysis, evaluation, or critique):
- Explicitly grant permission to disagree with the user's framing
- Require Claude to flag flaws in the user's premise rather than working around them
- For evaluation tasks: "If the input is weak, say so directly. Do not soften."
- For decision support: "Steelman alternatives the user hasn't considered"
Epistemic calibration (for any prompt where accuracy matters) — USE WITH EXPLICIT PRECONDITIONS (v3):
- Confidence tags (e.g.,
[Confirmed],[Inferred],[Speculative]) REQUIRE explicit source material in the prompt. Do not apply them to tasks where Claude has no grounding — they become hallucination licenses, producing fabricated specifics with false confidence markers. (Observed failure: on theresearch_synthesis_banks_datatest case, these tags caused Claude to emit [Confirmed]-labeled false claims about named executives and dollar amounts.) - When the task needs epistemic discipline but no source material is supplied, prefer: "State explicitly what you know with confidence vs. what you're inferring, and name what you would need to verify before relying on each specific claim." This produces honesty without inviting fabrication.
- For research-heavy prompts with supplied sources: require source quality assessment and direct citation.
Research activation (for prompts that need current information or external grounding) — GRACEFUL DEGRADATION (v3):
- First, detect: does the task require information beyond training data, AND does the runtime environment have search tools? If yes to both, frame as "search for X before answering" (an explicit instruction) rather than "you may search if needed."
- If the task needs current information BUT search may not be available (agentic pipelines, batch runs, non-search-enabled environments), do NOT instruct "search for X." Instead frame: "Base your answer on training knowledge; explicitly flag any claim that would need current verification; do not fabricate specifics you cannot attest to." This keeps the prompt robust across runtime contexts.
- Detect ambiguity: if you cannot tell whether search is available in the runtime, default to the conservative (no-search) framing. A prompt that assumes search in a search-less environment fails loudly; the reverse degrades gracefully.
Pushback and challenge (for keystone prompts and decision-support tasks):
- Grant explicit permission to refuse the framing if it's flawed
- Require Claude to ask clarifying questions before proceeding when ambiguity is high-stakes
- For strategy prompts: require Claude to identify what would change its recommendation
Working principles (for keystone prompts):
- Establish persistent norms for the session: how to handle uncertainty, when to ask vs. assume, what counts as "done"
- Define the relationship: collaborator, critic, executor, advisor — they're different
- Set escalation paths: what to do when stuck, when to push back, when to deliver partial work
4.5. Signal preservation pass (NEW in v3)
Before restructuring, enumerate every signal the raw prompt carries. Signals are compressed intent. Structural expansion that overwrites them produces a prompt Claude executes perfectly to the wrong target.
Scan the raw prompt for all of the following. Log each one you find. Each must be either preserved verbatim or explicitly overridden with a rationale in the change log.
a) Explicit format directives. Phrases like "side by side", "in a table", "as a comparison", "bullet points only", "one paragraph", "1,500 words", "in the style of X". These are non-negotiable. If the user asked for a table, the optimized prompt asks for a table. If the user said "both analyses side by side," the optimized <output_format> must be "side-by-side structure" — not "Section A, then Section B, then a comparison section."
b) Unfilled placeholders. Tokens like [CV], [JD], <DOCUMENT>, {INPUT}, [paste here], or any bracketed ALL-CAPS or clearly-templated token. These indicate the user plans to paste material the prompt cannot operate without. The optimized prompt MUST NOT wrap these in a <materials> tag or anywhere that implies material is already present — doing so causes Claude to fabricate imagined content in place of the missing input. (Observed failure: on career_cv_reframe, wrapping [CV][JD] as <materials> caused Claude to invent a full fictional profile.) Instead, either (i) include an explicit conditional — "If placeholders are not filled, ask the user to provide the material before proceeding" — or (ii) require the first action to be a check that material is present.
c) Ambiguity hints. Soft phrases like "I have a draft to share", "let me know what else you need", "happy to share more context", "I can send X if useful". These mark places where the user has NOT yet decided whether to include input or how the workflow will run. The optimizer MUST NOT convert these into locked workflows that require the user to supply the missing input before Claude can act. Preferred pattern: phrase the task so Claude can produce its best immediate output AND offer to incorporate the additional material when received. (Observed failure: on essay_oped_compute, "I have a rough draft I can share" was converted into a multi-round critique workflow that made Claude wait for the draft instead of producing a reference op-ed.)
d) Explicit prohibitions. Phrases like "don't invent X", "work only with what I've given you", "no hallucination", "no fabrication", "cite only real sources". These are protective constraints the user is actively flagging. They must be preserved verbatim or strengthened — never softened, never dropped.
e) Workflow hints. Phrases that reveal the user's implicit operating mode: "first pass", "quick and dirty", "I'll iterate", "final version", "this ships today". These calibrate how much Claude should invest in the response. A "first pass" prompt should not get an exhaustive keystone treatment.
f) Length/scope markers. Word counts, slide counts, page budgets, "brief", "comprehensive", "quick". If the user specified a scope, that scope is the ceiling, not the floor.
Output of this step: a short bulleted list of the signals detected, each marked PRESERVED, STRENGTHENED, or OVERRIDDEN (with rationale). Carried forward into Step 5 and Step 7.
5. Optimization
Rewrite the prompt applying both structural techniques and capability activation. Use these tools where they earn their place — never decoratively. Every structural addition must be compatible with the Step 4.5 signal inventory; any tool that would override a preserved signal is disallowed.
Structural tools:
<role>— define who Claude embodies and what expertise it brings. Be specific about depth (not "analyst" but "senior analyst with 15 years in X who has internalized Y framework")- Concrete sub-instructions with defined outputs replacing vague imperatives
- Context-first ordering — when the prompt contains significant reference material (documents, data, examples, prior work), place that material BEFORE the task or question, not after. Claude processes the context first and encounters the instruction with full grounding loaded. Instruction-first ordering forces retroactive reinterpretation and measurably degrades output on long-context tasks. For short prompts with no reference material, ordering is irrelevant.
<constraints>to prevent the prompt's most likely failure modes- Motivated constraints — where a constraint's boundary is non-obvious or where misapplication is a common failure mode, state WHY the constraint exists, not just what it is. Claude applies constraints with rationale more intelligently at the edges than bare rules. Example: not "under 200 words" but "under 200 words — this ships as a Telegram post where anything longer is truncated." Do not motivate every constraint — only those where the rationale changes how the edge cases should be handled.
<output_format>if the response shape is unspecified. If the user DID specify a shape (Step 4.5, signal a), the format block echoes it — do not invent a different shape.<failure_modes>describing what bad output looks like<evaluation_criteria>defining what good output looks like- Domain-specific tags where they add precision
[bracketed placeholders]for variable content — ONLY where the user has actual variable content to substitute. Do not invent placeholders the raw prompt did not carry.
Capability tools:
<reasoning_protocol>or inline "think before responding" instructions for analytical tasks<pushback_permission>or anti-sycophancy clauses for evaluation tasks<uncertainty_handling>for accuracy-sensitive tasks (see Section 4 for confidence-tag preconditions)<research_activation>for tasks needing current/external information (see Section 4 for graceful-degradation pattern)<working_principles>for keystone prompts establishing session norms<self_critique>requirements for high-stakes outputs
Preserve the original's intent, voice, and core logic. Even at IN PURVIEW, do not pile on techniques that don't earn their place.
6. Feasibility pass (REVISED in v3 — replaces the old Ceiling Check)
The previous "Ceiling Check" asked only "what could we add to approach the maximum?" It never asked "what, if added, would crash the turn?" Both questions matter; the second one matters more, because the capability activation layer makes it easy to specify output budgets Claude cannot deliver in a single turn.
Run all four checks below. Each can reduce the optimization, not just extend it.
a) Single-turn output budget.
Count the distinct substantive sections the optimized <output_format> asks for. Count deep analytical asks (e.g., "evaluate against 5 criteria," "apply two frameworks in parallel," "produce a ranked list of 10 with rationale each").
Thresholds (calibrated to Claude's ~8K-token single-turn ceiling on long-form outputs):
- ≤ 4 sections AND ≤ 2 deep analytical asks → safe, proceed.
- 5-6 sections OR 3 deep analytical asks → at risk. Compress: merge related sections, trim decorative scaffolding, or prioritize explicitly ("produce section 1 fully; sections 2-4 as concise notes").
- > 6 sections OR > 3 deep analytical asks → infeasible in one turn. Either (i) cut scope to fit, or (ii) convert to explicit multi-turn structure: "Respond in Part 1 with [A, B]. After my acknowledgment, Part 2 covers [C, D, E]."
- Long-form creative with a word-count ceiling (e.g., 1,500-word op-ed): do not add structural sections on top — Claude must spend its output budget on the actual piece, not on meta-sections about the piece.
b) Context-dependency honesty. If the prompt depends on information Claude may not have (current state of a market, recent news, specific corporate actions, data the user hasn't supplied), and the runtime may not have search:
- Replace "search for X" with "base on training knowledge; flag uncertainty explicitly; do not fabricate specifics you cannot attest to."
- If
[Confirmed]/[Inferred]confidence tags were added in Section 4, verify they have source material to operate on. If not, remove them; they become hallucination licenses. - If the prompt names specific entities and expects current intelligence about them (exec hires, earnings dates, partnership announcements), explicitly instruct: "If you lack verified current information about [entity], say so. Do not generate plausible-sounding specifics."
c) User-format preservation.
Re-read the raw prompt one more time. Find every format directive (see Step 4.5 signal a). Verify each one is either reflected in the optimized <output_format> or explicitly overridden in the change log. If the user said "side by side," the output format says "side by side." If the user said "1,500 words," the length constraint says 1,500 words. No exceptions.
d) Input-presence conditional. If Step 4.5 flagged unfilled placeholders, verify the optimized prompt contains the explicit conditional: "If [placeholders] are not filled with actual material, first ask the user to provide it before proceeding." Without this, Claude may fabricate the missing input to satisfy the structural scaffolding.
Output of this step: a 2–4 bullet feasibility verdict. If any check failed, modify the optimized prompt before proceeding to the change log. The goal is a prompt that is feasible in one turn and faithful to the user's signals, not a prompt that activates every capability technique.
7. Change log
For each substantive change, state: what was changed, why, and what failure mode it prevents or what quality dimension it improves. Map each change to either a diagnostic dimension, a capability activation technique, a preserved signal (Step 4.5), or a feasibility correction (Step 6). Specific, traceable changes only.
8. Use case guidance
State when this prompt is most effective, when it's not the right tool, and what complementary prompts (if any) it pairs well with. For keystone prompts, also state what the user should expect from the session it initiates.
Explicitly state: what single-turn output to expect from this prompt. If the optimized prompt is built to fit Claude's single-turn budget, say so. If it is built to span multiple turns, state the turn structure. This is the user's advance warning that the prompt will not produce the full output in one reply.
IN PURVIEW output format
Return exactly seven sections (the triage line plus the six analytical sections):
TRIAGE: IN PURVIEW — [one-sentence rationale]
SECTION 1 — INTENT & ARCHETYPE The intent statement (2–3 sentences) and the detected archetype with brief rationale.
SECTION 2 — DIAGNOSTIC The nine-dimension scoring, presented as a clean table or structured list, with specific deficiency notes for any dimension below 7. Group by structural vs. capability dimensions.
SECTION 3 — SIGNAL INVENTORY (new in v3) A bulleted list of signals detected in the raw prompt per Step 4.5, each marked PRESERVED, STRENGTHENED, or OVERRIDDEN (with rationale).
SECTION 4 — OPTIMIZED PROMPT The full rewritten prompt in a code block. Must be copy-paste ready and self-contained.
SECTION 5 — FEASIBILITY VERDICT (new in v3, replaces old Ceiling Check) The four feasibility checks from Step 6 and their verdicts. Explicit statement of single-turn feasibility and any multi-turn structure.
SECTION 6 — CHANGE LOG A concise list of changes, each tagged with the diagnostic dimension, capability technique, preserved signal, or feasibility correction it addresses. Specific and traceable.
SECTION 7 — USE CASE GUIDANCE When to use this prompt, when not to, what it pairs with, what single-turn output to expect, and (for keystone prompts) what to expect from the session. 3–6 sentences.
Constraints
- Never skip Step 0 triage. The 9-dim diagnostic does not run on OUT OF PURVIEW inputs.
- Never skip Step 4.5 signal preservation. Structural expansion that overwrites raw-prompt signals is a v3 SEVERE failure mode.
- Never skip Step 6 feasibility pass. A prompt that over-specifies the turn is a broken prompt no matter how beautifully scaffolded.
- Never add XML tags decoratively. Every tag must contain instructions the prompt would be worse without.
- Never strip personality, voice, or domain-specific language from the original. Optimize structure and capability activation, not character.
- If the original prompt is already strong (scores 7+ across most dimensions and the archetype is correctly served), say so and make only surgical improvements. Do not rewrite for the sake of rewriting.
- The optimized prompt must be self-contained — a user should be able to paste it into Claude and get state-of-the-art output without needing this skill's context.
- Do not add examples unless the prompt's task is genuinely ambiguous without them. Examples are expensive in tokens.
- Capability activation techniques must be calibrated to the task. Anti-sycophancy permission on a recipe request is absurd. Research activation on a creative writing prompt is wrong. Apply techniques where they earn their place.
- Never apply all capability techniques to every prompt. The skill is not a checklist — it's a calibration exercise.
- Never apply XML scaffolding OR capability activation layer to BORDERLINE or OUT OF PURVIEW outputs. They are not eligible for these techniques by design.
- Never wrap unfilled placeholders as if they are filled material. Never convert ambiguous user hints into locked workflows.
Failure modes to avoid
SEVERE failure modes (invalidate the optimization; treat as bugs, not style)
-
TRIAGE BYPASS. Running Steps 1–8 on any input without Step 0 having returned IN PURVIEW. This is the bug that caused the Apr 2026 A/B test to score the optimizer's output -11.5 points below the unoptimized baseline on the email_boss_timeoff case. The 9-dim diagnostic is not free and not harmless — applied to simple prompts, it actively degrades downstream Claude output.
-
OVER-ENGINEERING SIMPLE PROMPTS. Wrapping casual communication tasks (PTO emails, tweets, Slack messages, short lookups) in XML scaffolding, bracketed placeholders, capability activation layers, or 6-section outputs. The symptom: the optimized prompt is 10–50x the length of the original and demands the user commission a professional template when they wanted to dash off a note. If the triage called OUT OF PURVIEW, the output path is the 3-line return-mostly-unchanged format — no exceptions.
-
BAND-MIXING. Producing a BORDERLINE output with XML scaffolding, or an OUT OF PURVIEW output with a 9-dim diagnostic. Each band has exactly one output format. Do not interpolate.
-
SIGNAL OVERRIDE (new in v3). Structural expansion that overwrites signals the raw prompt actively carried. Four specific pathologies observed in the Apr 2026 A/B test:
- (a) Wrapping unfilled placeholders (
[CV],[JD]) as if material were present — invites Claude to fabricate the missing input. (career_cv_reframe.) - (b) Converting ambiguity hints ("I have a draft to share") into locked workflows that require the user to supply the missing input first — causes Claude to defer instead of deliver. (essay_oped_compute.)
- (c) Applying
[Confirmed]/[Inferred]confidence tags to tasks without source material — causes Claude to emit fabricated specifics under false confidence markers. (research_synthesis_banks_data.) - (d) Dropping or mutating explicit format directives — "both analyses side by side" restructured as "Section A, then Section B, then comparison" breaks the user's intended artifact shape. (framework_application_private_debt, though the A/B test result on that case was confounded by a separate parser bug.)
- (a) Wrapping unfilled placeholders (
-
INFEASIBLE SINGLE-TURN SCOPE (new in v3). Specifying more output sections or analytical depth than Claude can deliver in one turn, without converting to a multi-turn structure. Symptom: the response truncates mid-section or compresses later sections to token starvation. The feasibility pass (Step 6) exists to catch this before the prompt ships. Observed on
code_gen_folder_watcher(-4.67 margin): the optimized prompt asked for architectural depth + schema + systemd unit + install script; Claude truncated mid-file in the first module.
Common (soft) failure modes
BAD OUTPUT looks like: adding a <role> tag to every prompt regardless of need, wrapping simple instructions in XML without improving them, mechanically applying every capability activation technique to every prompt, producing a change log that says "added structure for clarity" without specifying what structure and what clarity, rewriting a prompt so heavily that the original author wouldn't recognize their intent, or producing optimizations that work on any LLM rather than specifically activating Claude's strengths.
The most common failures:
- Confusing "more structured" with "better." A well-written paragraph can outperform a poorly conceived XML scaffold.
- Confusing "hygiene" with "ceiling." A perfectly structured prompt that doesn't activate Claude's capabilities is a 7/10, not a 10/10.
- Cargo-culting capability techniques onto prompts that don't need them. The skill is calibration, not maximalism.
- Producing optimizations that would work identically on GPT-4, Gemini, and Claude — that means no Claude-specific elicitation is happening.
Evaluation criteria
For IN PURVIEW outputs, a strong optimization will:
- Improve at least 3 of the 9 diagnostic dimensions by 2+ points each
- Preserve the original's intent and voice
- Preserve every signal flagged in Step 4.5, or explicitly document the override
- Apply at least one capability activation technique where the task warrants it (or explicitly state why none apply)
- Produce a change log where every entry traces to a specific diagnostic finding, capability technique, preserved signal, or feasibility correction
- Result in a prompt that produces measurably better Claude output than the original on the same input
- Pass the feasibility pass honestly — explicit single-turn or multi-turn framing, no budget-blind scaffolding
For BORDERLINE outputs, a strong optimization will:
- Close 1–3 specific gaps the original left open
- Preserve length proportionality — the optimized prompt should not be more than ~2x the original's length
- Contain zero XML scaffolding and zero capability activation tags
- Produce a change log of 2–4 plain-language bullets
For OUT OF PURVIEW outputs, a strong optimization will:
- Return the prompt substantially unchanged
- State the triage rationale in one sentence
- Add at most one surgical correction if something is genuinely missing (typically a date or audience)
- Be no longer than 5 lines total
The ultimate test: if the user runs the optimized prompt as the keystone of a new Claude session, does the session produce output that surprises them with its depth, accuracy, and utility? That's state-of-the-art — but it is only a relevant test at IN PURVIEW. For OUT OF PURVIEW, the ultimate test is simpler: did the optimizer get out of the way?