FRM design
Design task-local frameworks, optimize agent action spaces, and extract reusable skills. Covers action space design, observation formatting, error recovery contracts, context budgeting, task-local framework templates, eval gates, shared skill extraction, and control pane checkpoints. Incorporates former agent-FRM-construction and dynamic-workflow-mode.From its SKILL.md
npx -y skills add asong56/skills --skill FRM-designAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
2 things to look at
- 25 days oldThe repository was created 25 days ago. New is not bad, but a brand new repository carrying a familiar-sounding name is the shape a typosquat arrives in, and there has been no time for anyone else to find a problem with it.
- 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
7.2 KB, ~1.4k tokens by cl100k_base, as published. Nobody here has run it
Agent Harness Construction
Use this skill when you are improving how an agent plans, calls tools, recovers from errors, and converges on completion.
Core Model
Agent output quality is constrained by:
- Action space quality
- Observation quality
- Recovery quality
- Context budget quality
Action Space Design
- Use stable, explicit tool names.
- Keep inputs schema-first and narrow.
- Return deterministic output shapes.
- Avoid catch-all tools unless isolation is impossible.
Granularity Rules
- Use micro-tools for high-risk operations (deploy, migration, permissions).
- Use medium tools for common edit/read/search loops.
- Use macro-tools only when round-trip overhead is the dominant cost.
Observation Design
Every tool response should include:
status: success|warning|errorsummary: one-line resultnext_actions: actionable follow-upsartifacts: file paths / IDs
Error Recovery Contract
For every error path, include:
- root cause hint
- safe retry instruction
- explicit stop condition
Context Budgeting
- Keep system prompt minimal and invariant.
- Move large guidance into skills loaded on demand.
- Prefer references to files over inlining long documents.
- Compact at phase boundaries, not arbitrary token thresholds.
Architecture Pattern Guidance
- ReAct: best for exploratory tasks with uncertain path.
- Function-calling: best for structured deterministic flows.
- Hybrid (recommended): ReAct planning + typed tool execution.
Benchmarking
Track:
- completion rate
- retries per task
- pass@1 and pass@3
- cost per successful task
Anti-Patterns
- Too many tools with overlapping semantics.
- Opaque tool output with no recovery hints.
- Error-only output without next steps.
- Context overloading with irrelevant references.
Dynamic Workflow Mode
Use this skill when a coding agent can generate or adapt a task-local framework instead of only following a static command flow. The goal is to turn dynamic workflow mode into a disciplined system: temporary frameworks for one-off work, shared skill extraction for repeated work, and observable control pane checkpoints for teams.
When to use
- The user mentions dynamic workflows, custom frameworks, framework-per-task, adaptive workflows, or Claude Code dynamic workflow mode.
- A task needs a custom loop, evaluator, crawler, fixture generator, watcher, or local dashboard.
- Multiple agents need the same repeatable process but the process is not yet captured as a shared skill.
- A workflow needs durable handoff artifacts, eval evidence, or operator approval before merge.
Core Contract
Dynamic workflow mode should produce a task-local framework only when the harness is cheaper and safer than manually driving the same steps. The harness must have:
- Objective: the outcome it owns and the outcome it explicitly does not own.
- Inputs: files, URLs, prompts, data sources, credentials policy, and user-provided constraints.
- Outputs: commits, reports, screenshots, status files, or control pane snapshots.
- Eval: at least one pass/fail check tied to the task, not only "it ran".
- Handoff: a short artifact that tells the next operator what happened, what is blocked, and how to resume.
Dynamic Harness Decision Tree
- One-shot task: keep it inline. Do not invent a harness.
- Repeated task with changing inputs: create a task-local framework and keep it under a temp or project-local working area.
- Repeated task across teammates or repos: extract the pattern into a shared skill.
- Task with external state, queueing, or approvals: add control pane visibility before adding more automation.
- Task with safety risk: add an eval gate and a human merge gate before autonomous execution.
Task-Local Harness Template
Use this structure before writing code:
# Dynamic Workflow Harness
Objective:
- Ship:
- Do not ship:
Inputs:
- Repo or workspace:
- External systems:
- Credentials policy:
Loop:
1. Discover current state.
2. Generate or update the smallest useful artifact.
3. Run eval checks.
4. Record status and handoff.
5. Stop on failed gate, unclear ownership, or unsafe external action.
Eval:
- Command:
- Expected pass signal:
- Failure owner:
Handoff:
- Status:
- Evidence:
- Next action:
Shared Skill Extraction
Promote a task-local framework into a shared skill only when at least two of these are true:
- The same workflow appears in multiple sessions, repos, teams, or launches.
- The workflow needs specific language, tool, or safety sequencing.
- Failures repeat because operators skip a gate or lose context.
- The workflow has a stable input/output contract.
- The workflow benefits from a control pane, status board, or team handoff.
When extracting, write the skill first in skills/<name>/SKILL.md. Add command shims only if a legacy slash-entry surface is still required.
Control Pane Checkpoints
Dynamic workflow mode becomes team-usable when it exposes state. Record these checkpoints whenever the task spans more than one session:
- Plan: objective, owner, acceptance criteria, and risky external systems.
- Queue: work items, assigned agent role, branch/worktree, and dependency edges.
- Run: active harness, current loop step, recent eval result, and token/cost signal if available.
- Gate: test results, browser screenshots, security review, and merge readiness.
- Handoff: what is done, what failed, what needs a human decision.
If the repo has ECC2 state enabled, prefer adding or reading checkpoints through the SKC control pane or state-store-backed scripts instead of scattering untracked notes.
Eval Gates
Every dynamic framework needs a task-specific eval. Pick the cheapest reliable gate:
| Work Type | Eval Gate |
|---|---|
| Code feature | Focused test, lint, coverage, and one integration path |
| UI/control pane | Browser smoke with screenshot and overflow/error checks |
| Agent workflow | Fixture transcript or seeded work item with expected routing |
| Research/content | Source-neutral brief, claim checklist, and publish-ready outline |
| Integration | Dry-run command, config validation, and no-secret scan |
Do not claim a dynamic workflow is reusable until the eval can be rerun by another teammate.
Anti-Patterns
- Generating scripts that hide the real decision logic from the operator.
- Treating dynamic workflow mode as permission to skip tests.
- Creating one-off docs when a shared skill or status artifact is the real product.
- Running multiple agents without ownership, merge gate, or conflict policy.
- Letting raw private research data leak into public docs.
Output Standard
Finish with:
- The harness or skill path.
- The eval commands and results.
- The control pane or handoff artifact path.
- The next reusable extraction candidate.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.