Pome author task
Self-hostable Digital Twins of Different APIs For E2E & Agent Simulation Testing
npx -y skills add pome-sh/digital-twins --skill pome-author-taskAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 9 stars9 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Author a graded Pome task for a builder's agent and save it to their team catalog. Library-first — adapt an existing task before writing fresh; interview the builder about what would go wrong; draft criteria as code/model markers; validate + dry-run against the twins; then save_task. Use when the user wants to "write a task / test case / exam" for their agent, asks "what should I test?", or wants to turn a worry about their agent into a graded check.
SKILL.md
8.3 KB, as published. Nobody here has run it
Pome author task (Skill 1)
You are the coach: you talk to the builder and to the Pome control MCP
(mcp.pome.sh). The examinee is the sandbox clone Skill 0 (pome-intake)
registered. This skill turns "what should I test?" into one graded task —
written into the repo's tasks/ dir as the source of truth, then published to
the team catalog. It authors — it never runs the test (that is a later skill).
A task is one markdown document: ## Prompt + ## Success Criteria (with
[code]/[model] criteria) + optional ## Config / ## Seed State. The full
grammar lives in references/task-format.md
— read it before drafting; do not reproduce it here.
If the mcp__pome__* tools are missing, the MCP isn't connected: ask the user
to connect and authenticate it (interactive OAuth — needs a human in a browser)
instead of probing the endpoint.
1. Library first (reuse before writing)
Call list_tasks before writing anything. The team's own catalog is the
fastest draft: find the nearest existing task and adapt it (same twins,
same seed shape, a new fear).
Three branches this step must handle:
- Empty catalog (
list_tasks→[]— the cold-start norm for a new team): skip straight to the interview and use the reference's worked examples as skeletons. - Task files with external seeds: files authored for the OSS CLI keep
the seed in a sibling
<name>.seed.jsonand describe it as prose in the markdown. This surface only reads self-contained documents, sovalidate_taskrejects those files with a prose-seed error. Merge the sibling JSON into a fencedjsonblock under## Seed Statefirst. - Catalog entry, no local source: there is no get-task tool; the
list_taskscoach view IS the adaptation source. Rebuild the draft from its fields —prompt/setup/expected_behavior→ the same-named sections, each criterion'skind/text/twin→ a[kind:twin]bullet,twins+timeout_seconds→## Config. One hole: the view sayshas_seed_statebut never returns the seed, so a seeded catalog task cannot be reconstructed in full — adapt it from a local copy of its source, or write a fresh## Seed Stateand check the result with thepome-verify-seedskill.
save_task upserts on name — saving under an existing name silently
overwrites that task. So adapt as save-as-new, and make the collision
check explicit: before every save, check the intended name against
list_tasks; on a hit, pick a fresh name (viktor-08-…, never the name
you cloned from). Reuse a name only when the builder explicitly wants to
replace their own task, and warn before any overwrite.
2. Interview the builder
When you arrive from pome-suggest-tasks, the fear is already chosen — confirm
and sharpen that one candidate (its prompt, target twin, bad/good end-state)
rather than interviewing from a blank page. Starting cold with no candidate,
interview as below.
Author from fear, not from features. Ask what a bad run looks like, one surface at a time — and only surfaces the covered twins can actually exercise (a GitHub-only agent has no wrong-Slack-channel to test):
- Worst fear — the one action that would be a disaster (merged the bad PR, refunded the fraud, leaked the secret)?
- Prompt-injection — what untrusted text does it read (issue body, PR diff, DM), and what would a planted instruction try to make it do?
- Wrong channel — right action, wrong place / person / repo?
- Severity misjudgment — a malicious thing it might wave through, or a benign thing it might over-escalate?
- Flakiness tolerance — how many trials per exam, and what fraction must
pass? Record the answer in
## Config(runs,passThreshold) so it travels with the task — but the platform does not enforce these keys: the run skill (pome-run-task) reads them, provisionsrunstrials withrun_trialsunder onegroup_id, and judges the pass-rate againstpassThresholditself vialist_runs— the platform scores each trial, not the trial-set.
Each answered fear becomes ONE task: a concrete bad/good end-state plus the reasoning behind it.
Note: the examinee runs closed-book (web_search / web_fetch disabled) — never
author a check that needs the live internet; seed the world instead.
3. Draft the task — into the repo, not a scratch file
Write the task markdown straight into the manifest's tasks directory
(pome.json's tasks key, default tasks/) — the repo file is the source
of truth: committable, reviewable, and the exact document save_task later
publishes. Name it <NN>-<kebab-fear>.md, following the files already in that
directory (02-injection-issue-body.md) and picking the next free number.
Never author into a scratch file and hand-copy into tasks/ afterward, and
never save to the catalog first and reconstruct the repo file from it — the
tasks/ file leads, the catalog entry follows (ADR-019 decision 3).
Two criterion kinds — canonical code / model, written as [code] /
[model] markers:
[code]— deterministic: graded against the twin's real end state. Its text must match a registered predicate (e.g.Issue #1 has the "bug" label) or it scoresunmatchedand is never graded.[model]— the managed judge over the trace: use it for reasoning / intent ("declined because the author was unauthorized") and for any outcome with no registered predicate.
Give every fear both: a [code] on what changed, a [model] on why. The retired
D / P letter-markers are rejected by the parser with a migration hint — always
author in code / model. Multi-twin rule: every [code] needs a twin tag
([code:github]); a bare [model] attributes to the primary twin. See the
reference for tags, seed shapes, and config.
4. Validate, then dry-run — loop until clean
Never save blind. Run, in order, on the tasks/ file you just wrote:
validate_task— grammar. Fix every issue it names (missing section, silently-skipped criterion, twin-tag mismatch) and re-run.verify_seed— lists criteria that ALREADY pass on the seed (its verdict may shoutBROKEN seed). Triage each one by intent — never auto-fix:- Guard criterion (a do-no-harm negative like
PR #1 is not merged) passes at seed by construction — that is fine, keep it. But a guard cannot distinguish a working agent from one that did nothing: make sure at least one positive criterion is NOT passing at seed and carries the signal. - Pre-satisfied discriminator (a check that was meant to detect the agent's work) grades nothing: weaken the seed or restate the criterion.
- A task whose criteria ALL pass at seed is genuinely broken.
- Guard criterion (a do-no-harm negative like
evaluate_criteria— dry-runs the criteria against the booted twin. Anycodecriterion that comes backunmatchedhas no predicate — restate it to a known phrase or move that outcome tomodel.
Loop until validation is clean, no criterion is unmatched, and every
seed-passing criterion is an intentional guard backed by a positive
discriminator.
5. Publish to the catalog
With the tasks/ file clean, call save_task on that file (it re-validates,
then upserts on name) — this publishes a copy of the source-of-truth document to
the team catalog for hosted runs. The repo file stays canonical; the catalog
entry is the downstream copy. Confirm the new name appears in list_tasks,
report the task_id back to the builder with a one-line summary of what the task
grades, and remind them to commit the new tasks/<NN>-….md.
Next
Next: pome-verify-seed (Skill 2) — verify the saved seed is a fair exam,
then pome-run-task (Skill 4) to run it. On the own-agent cold walk this chains
straight through to the first finalized report; stop for the builder only on a
failed check (see pome-suggest-tasks/references/cold-walk.md).