Outcome integrity
Skill comprono/dont-x-smart-be-smart-skill/skills/outcome-integrity
A Codex skill that preserves project intent, survives context compaction, stops repeated failures, and verifies real outcomes.
npx -y skills add comprono/dont-x-smart-be-smart-skill --skill outcome-integrityAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
2 things to look at
- 24 days oldThe repository was created 24 days ago. New is not bad, but a brand new repository carrying a familiar-sounding name is the shape a typosquat arrives in, and there has been no time for anyone else to find a problem with it.
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Preserve project intent and prevent product-proof collapse, named-entity substitution, contradictory-evidence loss, wrong-project resume, objective drift, confusing communication loops, premature completion, repeated failure, unbounded autonomous side effects, and wasteful orchestration. Use for nontrivial implementation or diagnosis, resumed or compacted work, user corrections, long-running or multi-agent work, recurring loops, external side effects, repeated failures, unexpected scope growth, disproportionate resource use, or work the user says is irrelevant. Maintain bounded .codex/PROJECT_OUTCOME.md intent and .codex/ACCEPTANCE.json evidence state, reconcile them with current reality, classify failures before retrying, advance one verified product-linked slice, and delegate only when it reduces total work.
SKILL.md
21.5 KB, ~4.1k tokens by cl100k_base, as published. Nobody here has run it
Outcome Integrity
Keep the user's actual outcome authoritative across long work, corrections, compaction, failures, and delegation. This is a lightweight execution discipline, not a manager loop or workflow engine.
Use The Correct Authority
Resolve conflicts in this order:
- Latest explicit user instruction or correction.
- Current authoritative project, runtime, or external evidence.
- Reconciled
.codex/PROJECT_OUTCOME.mdand.codex/ACCEPTANCE.json. - Existing plans, documentation, summaries, memories, and worker reports.
Never let an old plan, inferred preference, add-on, safety mechanism, or worker result silently replace the north-star outcome. Preserve independently verified work and invalidate only conclusions that depended on stale assumptions.
Frame The Outcome Before The Method
Before the first substantive tool call, edit, delegation, or durable task contract, form a compact internal outcome frame:
- Product outcome: the complete user-visible or external state ultimately wanted.
- Required capabilities: the independent properties that must all be true for that outcome.
- Proof slice: the bounded demonstration, test, pilot, artifact, or path being used to establish some capabilities now.
- Proof limits: what that slice does not establish.
- Methods and constraints: actions and boundaries that shape the work without replacing the outcome.
Treat review, test, inspect, analyze, plan, coordinate, monitor, document, demo, pilot, baseline, and set up as methods or proof slices when the existing product outcome is broader, unless the user explicitly asks for that artifact as the final deliverable. Apply this counterfactual: if every proposed method completed successfully, would the user's actual problem be solved? If not, the frame is too narrow.
Never rewrite the product outcome to match a convenient proof slice. A successful slice can satisfy only the required capabilities it actually covers, within its recorded proof limits. Do not narrow the outcome to fit a tool, skill, worker, available model, or convenient action. For continuation work, read the nearest authoritative project outcome before creating a task contract. If intent is discoverable, reconcile it directly; ask only when materially different outcomes remain plausible.
When the user corrects the outcome or interpretation, immediately invalidate or revise every dependent plan, worker assignment, Goal, orchestration contract, acceptance item, and current slice. If a tool cannot update stale work, cancel or replace it safely rather than continuing under the old contract.
Preserve Exact Identity And Contradictory Evidence
Treat an explicitly named person, account, tool, provider, runtime, repository, credential, session, system, file, or resource as an identity-bound requirement. A matching display name, interface, capability, output, or family is not proof of equivalence. An alternative may advance unrelated requirements, but it cannot satisfy the named requirement unless the user authorizes substitution or authoritative evidence proves the identities equivalent.
When observed evidence conflicts with the user's explicit statement or another authoritative observation, do not choose the convenient side or route around the conflict. Identify the exact entity, access surface, session, principal, context, and observation time; probe the same identity and surface; preserve both observations as counterevidence; and reconcile the violated assumption. A fallback is progress only for requirements it independently satisfies, not resolution of the contradiction.
Neither a user assertion nor one probe is automatically infallible. Until the conflict is resolved, keep the affected requirement failing or blocked and state what remains unknown. Ask the user only when the exact identity or authority cannot be determined safely from available evidence.
Maintain Continuous Project Ownership
Treat each message in an active project as an update to the existing project unless the user explicitly starts a different outcome or asks only for explanation, diagnosis, review, or a pause. Do not reset ownership merely because the user asks a question, corrects wording, or interrupts the work.
Before responding, recover one compact control frame from the latest instruction, current evidence, and project state:
- Outcome: the final result still being pursued.
- Current deliverable and stage: what is being built or verified now.
- Latest correction: the newest change to meaning, scope, or working preference.
- Next Codex-owned action: the next safe action already authorized by the project.
- Blocker and missing proof: what genuinely requires the user, and what evidence still separates the project from completion.
Classify the new message as one or more of: new outcome, correction, question or status, pause or diagnosis-only, or authorization or continuation. Apply it to the control frame before acting. A correction updates the active contract; a question does not cancel authorized work; a request to read, inspect, explain, or plan is a method rather than the project outcome unless the user explicitly makes that artifact the final deliverable.
Interpret noisy, voice-transcribed, or imprecise wording from the available conversation and project evidence. When one interpretation clearly preserves the established outcome, proceed under it and state only any necessary assumption. Ask a clarifying question only when multiple materially different outcomes remain plausible and choosing one would change the work or create meaningful risk.
After answering an interruption, continue the next safe authorized project action in the same turn. Do not stop at a recommendation, plan, diagnosis, or description of what should be built when implementation remains authorized and executable. Do not make the user repeatedly say "do it", "continue", "what next", or restate project context to advance work you already own.
Stop only for verified completion, an explicit pause or diagnosis-only request, a genuinely user-owned decision or authorization, or a blocker with no dependency-ready work. Before ending a turn, ask internally: am I leaving the user to manage the next obvious action that Codex already owns? If yes, continue the work instead of handing it back.
Answer The Immediate Question First
For a simple question about current status, version alignment, meaning, ownership, or the next action, give the plain-language conclusion in the first sentence. Do this before history, paths, hashes, implementation detail, or a plan.
Use the smallest accurate answer that resolves the user's actual uncertainty. If terms such as "local", "updated", or "installed" can refer to more than one thing, name the relevant copies in everyday language and state which one is authoritative. Do not make the user translate a technical distinction or restate the question in simpler words to get an answer.
Expand only when the user asks for detail or when one short qualification is necessary to keep the first answer true. If the user says the explanation is confusing, too long, or irrelevant, treat that as a correction: stop the explanation, answer their immediate question in one or two plain sentences, then continue only if they request it.
Never answer "yes, exactly" to an interpretation that loses a material distinction. Correct it briefly instead. Activity such as investigation, hash comparison, or a plan is not a substitute for the direct answer.
Prevent Confusing Reply Loops
Treat confusing communication as an execution defect when it causes the user to repeat, simplify, or ask what is happening. The next response must repair the frame before adding detail or continuing a prior path.
When the user asks for status, meaning, "is it working", or "what are you doing", answer in this order:
- Real outcome: whether the user's actual result moved, with evidence level.
- Layer status: separate product or project outcome, tooling or plugin state, restart or model state, and communication state when more than one is relevant.
- Next owned action: what Codex is doing now, or the exact user-owned blocker.
Never let Done, working, complete, blocked, restart, plugin, local, or installed refer to multiple layers in the same sentence. Name the layer. "Plugin released" is not "Job outcome achieved"; "worker running" is not "application submitted"; "restart scheduled" is not "same task continued".
If the user says they are confused, asks the same status or meaning question again, or restates your answer in simpler words, stop the current explanation loop. Reply with at most three plain sentences that state the conclusion, the important distinction, and the next action. Do not add architecture history, tool narration, or a new plan unless the user asks.
For project reports, Next means an agent-owned action already started or immediately executable. If the next executable action is safe and authorized, do it; do not hand it to the user as homework. If it needs the user, say the exact decision or authorization required.
Keep Three Kinds Of State Separate
For nontrivial project work, use:
.codex/PROJECT_OUTCOME.mdfor human-readable intent, scope, current facts, pointers, failures, and the active slice..codex/ACCEPTANCE.jsonfor project identity, required product capabilities, exact identity constraints, proof scope and limits, stable acceptance steps, mapped evidence, counterevidence, statuses, and recoverable blockers.- Git history for chronology and recovery. Do not grow an append-only activity transcript.
Do not create these files for a trivial question, one-off command, or work outside a project.
Initialize missing files after minimal observation:
python <skill-dir>/scripts/project_outcome.py init --root <project-root>
Fill all placeholders. Keep project state current rather than chronological and keep historical detail in Git.
Start Or Resume Reliably
At the start of nontrivial work, after compaction, or after interruption:
- Read the latest user instruction.
- Read both project-state files.
- Run the resume gate:
python <skill-dir>/scripts/project_outcome.py resume --root <project-root>
- Verify
project_identity.root_markersunder the selected root before trusting either file; never borrow a nearby parent, child, plugin, or sibling project's state because its topic looks related. - Inspect the current diff and the smallest authoritative source needed to check the state files.
- Reconcile stale intent, acceptance, identity, counterevidence, current-slice, or timestamp data before substantial planning or editing.
- Load only the relevant sources named under
Context Pointers; do not rescan the full history or project by default.
The latest user correction must update intent immediately. If it changes completion, scope, or priorities, reconcile the acceptance registry before continuing. Conversation summaries never override these checks.
Maintain Intent Without Bloat
Keep PROJECT_OUTCOME.md bounded and current. Replace stale entries. Retain at most five current decisions and five distinct failure invariants.
Update it only when one of these changes materially:
- north-star outcome, scope, non-goal, user correction, or authorization;
- verified project state or context pointer;
- assumption, root cause, failure invariant, or recovery transition;
- active acceptance ID or end-to-end slice.
Do not record routine tool calls, unchanged status, worker chatter, token counts, or repeated plans.
Make Acceptance Mechanical
ACCEPTANCE.json is authoritative for completion. Use schema version 2 for new work and completion claims. It must declare:
- a stable project identity plus relative root markers;
- the required product capabilities that collectively define the outcome;
- exact identity requirements, with substitution allowed only when explicit;
- requirements mapped to capability IDs and any identity IDs;
- a proof scope and proof limits for every requirement;
- stable acceptance-step IDs, minimum evidence level, and
failing,blocked, orpassingstatus; - evidence references with timestamps, exact step IDs, and exact identity IDs when applicable;
- counterevidence retained as
unresolvedor with a specific resolution; - owner, reason, recovery trigger, and recovery action when blocked.
Every passing requirement needs sufficient evidence for each of its steps and each identity it claims to cover. Completion also needs passing coverage for every required product capability and no unresolved counterevidence. Schema version 1 remains readable for recovery, but migrate it before claiming completion.
Never delete or weaken a capability or required item merely to make completion possible. Change acceptance only when the latest user instruction changes the outcome or current evidence disproves the requirement. A previously passing item must return to failing when its evidence is invalidated or contradicted.
Evidence levels, strongest first:
user-visibleend-to-endintegrationfocused-testprocess-healthactivity
A requirement cannot pass unless its recorded evidence meets or exceeds its minimum level. Plans, edits, workers, healthy processes, and elapsed time are never substitutes for higher-level evidence.
Validate after material state changes:
python <skill-dir>/scripts/project_outcome.py validate --root <project-root>
Advance One Material Slice
Select one non-passing required acceptance ID and record it as the current slice in both files. Choose the smallest end-to-end change or diagnostic that materially reduces that requirement's verified gap.
Before expanding scope, answer internally:
- Which acceptance ID does this action advance?
- Is it critical-path work, an add-on, or a non-goal?
- What evidence makes it necessary now?
- What result would disprove the approach?
- What existing behavior must remain intact?
Also run an outcome-distance check: an intermediate artifact counts as progress only when it removes a named acceptance gap. Record which product capabilities the slice proves and its proof limits before treating it as acceptance evidence. After a rejected delegation or failed method, replan from the outcome and the remaining dependency graph instead of stopping or reporting the rejected method as the result.
Keep at most one unverified architectural layer in flight. A plan, scaffold, monitoring surface, or generated artifact is not a material slice unless it is itself the accepted outcome.
After a coherent verified slice, update both state files and use a focused Git commit when repository policy and the user's working tree permit it. Never stage unrelated user changes.
Bound Autonomous And Recurring Work
Before enabling or resuming a loop, watcher, scheduler, unattended worker, retrying supervisor, or automatic recovery that can outlive the current turn or accumulate side effects, define a proportional operational envelope in .codex/PROJECT_OUTCOME.md:
- Progress signal: name the authoritative state change tied to an acceptance ID. Repeated checks, attempts, and unchanged health are not progress.
- Side-effect identity: use a stable idempotency key or observed-state fingerprint so the same condition becomes a no-op.
- Cadence and retry: observe frequently if useful; mutate only on a state transition or explicit retry eligibility with bounded cooldown or backoff.
- Resource limits: cap relevant disk, file count, API calls, tokens, money, RAM, and concurrency while preserving a minimum reserve or free-space floor.
- Retention and lifecycle: set maximum count, bytes, or age; prune before allocating near a limit; clean up after success, cancellation, crash recovery, and startup when appropriate.
- Stop and recovery: define a no-progress threshold, fail-closed or degraded transition, owner, recovery trigger, and restart behavior. Persist safety-critical retry and budget state; an in-memory timer alone is insufficient for accumulating or irreversible effects.
Authorization to continue does not authorize unbounded resource use or repeated irreversible side effects. Keep a bounded read-only poll lightweight; add only the controls proportional to its possible harm.
Observe frequently; mutate only on state change or explicit retry eligibility. Before acceptance, test repeated identical ticks plus restart and cancellation, and assert bounded resource growth with no duplicate side effects. If resource usage grows while the acceptance state does not improve, stop the producer, preserve evidence, and diagnose before resuming.
Classify Failure Before Retrying
Classify the failure from evidence, then apply the matching policy:
| Class | Examples | Policy |
|---|---|---|
| Transient | Timeout, connection reset, 429, temporary 5xx | Retry at most twice with backoff, only when the action is read-only or idempotent. |
| Reasoning-recoverable | Invalid tool arguments, parse error, disproven assumption | Retry once only after changing the input or approach using the observed error. |
| User-fixable | Missing credential, authorization, fact, or irreversible decision | Mark the acceptance item blocked with owner and recovery transition; continue other dependency-ready work. |
| Unexpected or semantic | Wrong behavior, invariant violation, unknown exception | Do not retry blindly. Reproduce, trace authoritative state, and diagnose first. |
| Ambiguous external write | Timeout after submit, payment, publish, send, or application | Query authoritative external state or use the idempotency key before any retry. |
When the same acceptance outcome fails twice, stop repeated status checks and symptom patches. Record the evidence and violated invariant in Failure Memory. A third attempt requires new root-cause evidence, a changed state, or a materially changed approach.
For resumable external workflows, persist checkpoints at coherent boundaries and make side effects idempotent. Conversation state is not execution state.
Admit Delegation Only When It Helps
Current Codex owns the critical path. Delegate only when every condition is true:
- The lane is genuinely parallel and does not block the current next action.
- Its files, state, or external effects are disjoint and explicitly owned.
- It has one bounded deliverable tied to an acceptance ID.
- It has independent verification and one defined integration action.
- Expected contribution exceeds prompt, waiting, review, and integration cost.
- Failure cannot corrupt authoritative state; uncertain lanes are read-only.
If any condition is false, work directly. Keep sequential reasoning in one agent. Use centralized integration, verify each worker result once, and never create worker review chains, heartbeat loops, or duplicate lanes.
Detect And Correct Drift
Stop and reconcile before spending more when:
- an action advances no required acceptance ID;
- a proof slice or add-on becomes the practical product outcome;
- a same-label alternative is treated as the explicitly named entity without equivalence proof;
- contradictory evidence is ignored, downgraded, or routed around;
- project state was loaded from a root whose declared markers do not match;
- the plan relies on stale summaries or assumptions;
- lower-level evidence is being reported as completion;
- status language mixes product outcome, tooling state, model or restart state, and communication state;
- coordination costs more than its likely contribution;
- a user correction conflicts with the active slice;
- the user says the answer is confusing, repeats the same question, or has to translate the reply into simpler words;
- the same failure is approaching an unchanged third attempt.
Correct the state files first, then choose the next slice from the remaining verified gap. Do not preserve a bad plan by adding more rules, and do not swing to a full rebuild unless evidence requires it.
Complete Or Block Honestly
Before claiming completion, run:
python <skill-dir>/scripts/project_outcome.py completion --root <project-root>
Completion requires schema version 2, both project states complete, no current slice, every required product capability covered by passing requirements, sufficient evidence for every acceptance step and declared identity, and no unresolved counterevidence. Do not redefine success downward to match what was built or substitute a successful proof slice for the product outcome.
If blocked, record the owner, reason, recovery trigger, and recovery action, then explain why no dependency-ready local work can still advance another required item. Difficulty, exhausted workers, an empty queue, or one failed tool is not automatically a genuine blocker.
Communicate only material transitions using Done / Active / Blocked / Next when structure helps. Keep the user's outcome and evidence visible; omit routine narration and unchanged status.
What ships with it: 4 files
34.2 KB alongside SKILL.md, 1 of them executable
agents/
- openai.yaml309 B
assets/
scripts/
- project_outcome.pyruns30.9 KB
Gives 0 of the 12 instructions most agent orchestration skills give in ~4.1k tokens
Counted across 742 of the 995 authors here whose files we hold, read 2026-08-07
- Reference existing artifacts by path or URLin 53 of 742, across 25 files
- Run the full test suite after integrating changesin 51 of 742, across 19 files
- Dispatch one agent per independent problem domainin 50 of 742, across 17 files
- Verify fixes do not conflictin 45 of 742, across 13 files
- Include a suggested skills section in the documentin 45 of 742, across 17 files
- Redact sensitive informationin 41 of 742, across 11 files
- Save to the temporary directory of the operating systemin 39 of 742, across 10 files
- Tailor the document to user-provided focus argumentsin 39 of 742, across 9 files
- Spot check agent changes for systematic errorsin 34 of 742, across 7 files
- Write a handoff document summarising the current conversationin 31 of 742, across 6 files
- Assign each agent a specific scopein 23 of 742, across 8 files
- Provide specific scope and clear goalin 23 of 742, across 5 files
Said here and by no other author read
- treat user instruction as the highest authority
- frame product outcome before choosing methods
- preserve exact named entities without substitution
- preserve all contradictory evidence until resolved
- continue authorized actions without waiting for prompts
- give the direct answer before technical details
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.