agentsclimarketplace

Pome author task

Skill pome-sh/digital-twins/skills/pome-author-task

Author a graded Pome task for a builder's agent and save it to their team catalog. Library-first — adapt an existing task before writing fresh; interview the builder about what would go wrong; draft criteria as code/model markers; validate + dry-run against the twins; then save_task. Use when the user wants to "write a task / test case / exam" for their agent, asks "what should I test?", or wants to turn a worry about their agent into a graded check.From its SKILL.md

Install
npx -y skills add pome-sh/digital-twins --skill pome-author-task

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 11 stars11 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

8.3 KB, ~2.0k tokens by cl100k_base, as published. Nobody here has run it

Pome author task (Skill 1)

You are the coach: you talk to the builder and to the Pome control MCP (mcp.pome.sh). The examinee is the sandbox clone Skill 0 (pome-intake) registered. This skill turns "what should I test?" into one graded task — written into the repo's tasks/ dir as the source of truth, then published to the team catalog. It authors — it never runs the test (that is a later skill).

A task is one markdown document: ## Prompt + ## Success Criteria (with [code]/[model] criteria) + optional ## Config / ## Seed State. The full grammar lives in references/task-format.md — read it before drafting; do not reproduce it here.

If the mcp__pome__* tools are missing, the MCP isn't connected: ask the user to connect and authenticate it (interactive OAuth — needs a human in a browser) instead of probing the endpoint.

1. Library first (reuse before writing)

Call list_tasks before writing anything. The team's own catalog is the fastest draft: find the nearest existing task and adapt it (same twins, same seed shape, a new fear).

Three branches this step must handle:

  • Empty catalog (list_tasks[] — the cold-start norm for a new team): skip straight to the interview and use the reference's worked examples as skeletons.
  • Task files with external seeds: files authored for the OSS CLI keep the seed in a sibling <name>.seed.json and describe it as prose in the markdown. This surface only reads self-contained documents, so validate_task rejects those files with a prose-seed error. Merge the sibling JSON into a fenced json block under ## Seed State first.
  • Catalog entry, no local source: there is no get-task tool; the list_tasks coach view IS the adaptation source. Rebuild the draft from its fields — prompt / setup / expected_behavior → the same-named sections, each criterion's kind/text/twin → a [kind:twin] bullet, twins + timeout_seconds## Config. One hole: the view says has_seed_state but never returns the seed, so a seeded catalog task cannot be reconstructed in full — adapt it from a local copy of its source, or write a fresh ## Seed State and check the result with the pome-verify-seed skill.

save_task upserts on name — saving under an existing name silently overwrites that task. So adapt as save-as-new, and make the collision check explicit: before every save, check the intended name against list_tasks; on a hit, pick a fresh name (viktor-08-…, never the name you cloned from). Reuse a name only when the builder explicitly wants to replace their own task, and warn before any overwrite.

2. Interview the builder

When you arrive from pome-suggest-tasks, the fear is already chosen — confirm and sharpen that one candidate (its prompt, target twin, bad/good end-state) rather than interviewing from a blank page. Starting cold with no candidate, interview as below.

Author from fear, not from features. Ask what a bad run looks like, one surface at a time — and only surfaces the covered twins can actually exercise (a GitHub-only agent has no wrong-Slack-channel to test):

  • Worst fear — the one action that would be a disaster (merged the bad PR, refunded the fraud, leaked the secret)?
  • Prompt-injection — what untrusted text does it read (issue body, PR diff, DM), and what would a planted instruction try to make it do?
  • Wrong channel — right action, wrong place / person / repo?
  • Severity misjudgment — a malicious thing it might wave through, or a benign thing it might over-escalate?
  • Flakiness tolerance — how many trials per exam, and what fraction must pass? Record the answer in ## Config (runs, passThreshold) so it travels with the task — but the platform does not enforce these keys: the run skill (pome-run-task) reads them, provisions runs trials with run_trials under one group_id, and judges the pass-rate against passThreshold itself via list_runs — the platform scores each trial, not the trial-set.

Each answered fear becomes ONE task: a concrete bad/good end-state plus the reasoning behind it.

Note: the examinee runs closed-book (web_search / web_fetch disabled) — never author a check that needs the live internet; seed the world instead.

3. Draft the task — into the repo, not a scratch file

Write the task markdown straight into the manifest's tasks directory (pome.json's tasks key, default tasks/) — the repo file is the source of truth: committable, reviewable, and the exact document save_task later publishes. Name it <NN>-<kebab-fear>.md, following the files already in that directory (02-injection-issue-body.md) and picking the next free number. Never author into a scratch file and hand-copy into tasks/ afterward, and never save to the catalog first and reconstruct the repo file from it — the tasks/ file leads, the catalog entry follows (ADR-019 decision 3).

Two criterion kinds — canonical code / model, written as [code] / [model] markers:

  • [code] — deterministic: graded against the twin's real end state. Its text must match a registered predicate (e.g. Issue #1 has the "bug" label) or it scores unmatched and is never graded.
  • [model] — the managed judge over the trace: use it for reasoning / intent ("declined because the author was unauthorized") and for any outcome with no registered predicate.

Give every fear both: a [code] on what changed, a [model] on why. The retired D / P letter-markers are rejected by the parser with a migration hint — always author in code / model. Multi-twin rule: every [code] needs a twin tag ([code:github]); a bare [model] attributes to the primary twin. See the reference for tags, seed shapes, and config.

4. Validate, then dry-run — loop until clean

Never save blind. Run, in order, on the tasks/ file you just wrote:

  1. validate_task — grammar. Fix every issue it names (missing section, silently-skipped criterion, twin-tag mismatch) and re-run.
  2. verify_seed — lists criteria that ALREADY pass on the seed (its verdict may shout BROKEN seed). Triage each one by intent — never auto-fix:
    • Guard criterion (a do-no-harm negative like PR #1 is not merged) passes at seed by construction — that is fine, keep it. But a guard cannot distinguish a working agent from one that did nothing: make sure at least one positive criterion is NOT passing at seed and carries the signal.
    • Pre-satisfied discriminator (a check that was meant to detect the agent's work) grades nothing: weaken the seed or restate the criterion.
    • A task whose criteria ALL pass at seed is genuinely broken.
  3. evaluate_criteria — dry-runs the criteria against the booted twin. Any code criterion that comes back unmatched has no predicate — restate it to a known phrase or move that outcome to model.

Loop until validation is clean, no criterion is unmatched, and every seed-passing criterion is an intentional guard backed by a positive discriminator.

5. Publish to the catalog

With the tasks/ file clean, call save_task on that file (it re-validates, then upserts on name) — this publishes a copy of the source-of-truth document to the team catalog for hosted runs. The repo file stays canonical; the catalog entry is the downstream copy. Confirm the new name appears in list_tasks, report the task_id back to the builder with a one-line summary of what the task grades, and remind them to commit the new tasks/<NN>-….md.

Next

Next: pome-verify-seed (Skill 2) — verify the saved seed is a fair exam, then pome-run-task (Skill 4) to run it. On the own-agent cold walk this chains straight through to the first finalized report; stop for the builder only on a failed check (see pome-suggest-tasks/references/cold-walk.md).

What ships with it: 2 files

19.5 KB alongside SKILL.md

references/

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.