Agentic architecture
A Claude Skill for engineering agentic systems. Every pattern sourced to Anthropic & OpenAI docs, certainty-labelled, and kept honest by a bundled drift-checker.
npx -y skills add tdrapp/agentic-architectureAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Best practices for DESIGNING and STRUCTURING Claude agent systems — anchored on Anthropic's 'Building Effective Agents' taxonomy (workflows vs agents; augmented LLM; routing, parallelization, orchestrator-workers, evaluator-optimizer) and Claude Code / Agent SDK guidance, extended with validated cross-project patterns. Use this whenever architecting a multi-agent or subagent system, deciding workflow-vs-autonomous-agent, structuring skills/subagents/tools, wiring context isolation, adding a verification/red-team pass, encoding lessons as reusable memory, or packaging a Claude Code workspace as an Agent SDK app (dual-mode). Trigger even if the user only says 'how should I structure this agent', 'should this be a subagent or a tool', 'how do I split this into agents', 'design the orchestration', or 'turn my Claude Code repo into a real app'. NOT for generic application code (use toti-engineering) or a specific product stack (use genai-saas).
SKILL.md
10.3 KB, ~2.1k tokens by cl100k_base, as published. Nobody here has run it
Agentic Architecture
Reusable, Anthropic-grounded patterns for building Claude agent systems that stay simple, transparent, and cheap to run — and that compound across sessions. Anchored on Anthropic's Building Effective Agents framework, extended with patterns validated in production personal projects.
How to use this skill
- Start from the taxonomy below to name what you're building — a workflow (predefined code paths) or an agent (the model directs its own process). Most robust systems are workflows, or agents with workflow-shaped internals.
- Run the design checklist before writing any orchestration code. It forces the "start simple, add complexity only when it demonstrably helps" discipline.
- Pull specific patterns from
references/patterns.md(context isolation, generation/verification split, filesystem-first dual-mode, tool honesty, memory-as-files, orchestration shape). Each carries its Anthropic grounding + a why-it-transfers note.
The Anthropic anchor — Building Effective Agents
Source: Anthropic, "Building Effective Agents" (anthropic.com/engineering/building-effective-agents). Represent it faithfully; this taxonomy is the shared vocabulary.
Workflows vs agents. Workflows orchestrate LLMs and tools through predefined code paths. Agents let the LLM dynamically direct its own process and tool use. Prefer the most predictable form that solves the problem.
The building block — the augmented LLM. An LLM enhanced with retrieval, tools, and memory, generating its own queries, selecting tools, and deciding what to retain. Everything below composes this block.
The five workflow patterns:
- Prompt chaining — decompose into sequential steps, each processing the previous output; add programmatic checkpoints between steps.
- Routing — classify the input and dispatch to a specialised follow-up. Enables separation of concerns and per-route optimisation.
- Parallelization — run tasks simultaneously and aggregate. Two variants: sectioning (independent subtasks in parallel) and voting (same task run several times for diverse outputs).
- Orchestrator-workers — a central LLM breaks a task down, delegates to workers, and synthesises their results (task decomposition is dynamic, not fixed).
- Evaluator-optimizer — one LLM generates, a second evaluates and feeds back in a loop for iterative refinement.
Autonomous agent. Begins from a command or discussion; once the task is clear, it plans and operates independently with environmental feedback, optional human checkpoints, and explicit stopping conditions.
Three core principles (non-negotiable):
- Simplicity — the fewest moving parts that work.
- Transparency — explicitly show the agent's planning steps.
- Tool craftsmanship — the agent-computer interface (tool docs + testing) deserves as much care as the prompts.
When NOT to use an agent. Agents add cost and compounding-error risk. Start with a single optimised LLM call + retrieval + in-context examples; add agentic complexity only when it demonstrably improves outcomes. Many production systems are one good workflow, not an autonomous agent.
Also cross-checked against OpenAI
Where Anthropic and OpenAI independently agree, treat the practice as fundamental (not fashion):
gate agents against a simpler baseline; single-agent first, split on measured tool-overload (not
raw count — ">15 distinct tools ok, <10 overlapping fail"); a central orchestrator/manager;
layered guardrails + human-in-the-loop on irreversible actions; evals-first. The one real
divergence to choose consciously: OpenAI makes handoffs (peer control transfer) a first-class
primitive, whereas Anthropic keeps an orchestrator in control. Full table + sources in
references/openai-crosscheck.md.
Design checklist (run at project start)
Work top to bottom; stop as soon as a simpler tier solves the problem.
- Can one augmented LLM call do it? (prompt + retrieval + examples + a couple of tools). If yes, stop here — do not build an agent.
- If not, is the control flow predictable? If yes, build a workflow (chaining / routing / parallelization / orchestrator-workers) in code, not an autonomous agent. Reserve autonomy for genuinely open-ended tasks with good feedback signals.
- Draw the pipeline as named stages — e.g.
intake → routing → parallel fan-out → synthesis → independent verification → log. Each stage = one responsibility. - Assign context boundaries. Which stages should run in an isolated sub-context that returns only a synthesis (to protect the main thread from raw tool output)? See pattern C1.
- Separate generation from verification. Whatever produces the decision must NOT also bless it — add an independent evaluator/red-team that sees the raw evidence, not the builder's narrative. See pattern C2.
- Design the tools as an interface. Structured outputs, an explicit uncertainty/quality flag, and a no-invention rule (return
null+ a documented gap, never a guess). Test each tool branch. See pattern C4. - Decide what must persist and where. Live state vs immutable record vs cross-session memory. Encode recurring discipline as files (skills / guardrails / patterns), not as prompt-of-the-moment. See patterns C5–C6.
- Ground before acting. A pre-flight step reads real state before the system emits any number/decision. See pattern C7.
- Plan for two runtimes from day one. Keep config filesystem-first (
CLAUDE.md,.claude/agents/*.md, skills) so the same source serves interactive Claude Code AND the Agent SDK. See pattern C3 +references/patterns.md#sdk-packaging. - Justify every added stage against principle 1 (simplicity). If a stage doesn't demonstrably improve the output, cut it.
Pattern catalog
Read references/patterns.md for the full catalog. Each pattern is stated as what / why it transfers / Anthropic grounding, with a concrete implementation note.
| Need | Pattern |
|---|---|
| Keep the main context clean during wide exploration | C1 — Subagent context isolation (wide-then-narrow) |
| Stop motivated reasoning in a decision pipeline | C2 — Generation/verification separation |
| One codebase, both interactive CLI and SDK app | C3 — Filesystem-first dual-mode |
| Tools that never silently lie | C4 — Structured tools + uncertainty flag + no-invention |
| Lessons that compound across sessions | C5 — Encode discipline as files, not prompts |
| Recall without re-reading everything | C6 — Layered persistence / just-in-time memory |
| Don't optimise in a vacuum | C7 — Pre-flight state read (ground first) |
| Load only what the moment needs | C8 — Progressive disclosure |
| Measure that it works + catch regressions | C9 — Evals as a first-class loop |
| Fuse many workers without regex-parsing prose | C10 — Structured subagent output schemas |
| Remember across sessions, portably | C11 — Filesystem layers vs the Memory tool API |
| Defend in layers; gate irreversible actions | C12 — Layered guardrails + human-in-the-loop |
Tools
scripts/doc_freshness_check.py — a zero-dependency script that fetches the primary Anthropic + OpenAI
docs this skill is grounded in (scripts/sources.json) and reports whether the specific claims the
skill relies on are still present. Run it when the user asks whether best practices have evolved, or
before publishing/citing a mechanism:
python scripts/doc_freshness_check.py # report drift (claim-missing / page-changed / unreachable)
python scripts/doc_freshness_check.py --update # re-pin hashes after verifying claims
It dogfoods C4: structured JSON output with a _meta.data_quality flag; an unreachable source becomes
unavailable, never a fabricated "unchanged".
Runnable proofs (not just prose): examples/minimal_agent/run.py is a zero-dep reference of the
orchestration shape (orchestrator → workers → independent verifier, C1/C2/C4/C10); evals/run_evals.py
is a real eval loop (pattern C9) over evals/tasks.json.
Optional reviewer subagent
agents/architecture-reviewer.md is an independent, skeptical reviewer that runs the design checklist
as a cold pass and argues against added complexity (dogfooding C2). Use it to review a proposed
design — the agent that produced a design should not be the one that blesses it. Usable in Claude Code
or wired via the Agent SDK agents= parameter.
What this skill is NOT
- Not general reasoning directives. "Never invent, say I-don't-know, cite sources" belong in the consuming project's guardrails, not here. This skill assumes them; it doesn't restate them.
- Not code quality. Planning discipline, refactoring, quality gates, git conventions live in
toti-engineering. - Not a product stack. Framework/stack specifics live in a stack skill (e.g.
genai-saas). - Scope = design & structure of agent systems. How to shape, split, verify, ground, and package them — nothing more.
A caution on certainty
When you cite a mechanism as "Anthropic best practice", distinguish what is verified in current docs (taxonomy, principles, SDK option names — which drift across versions, so pin them) from what is idiomatic guidance and from a team's own validated discipline. Never present house conventions as official Anthropic doctrine; label the source.
What ships with it: 21 files
77.2 KB alongside SKILL.md, 6 of them executable
agents/
- architecture-reviewer.md4.0 KB
evals/
- README.md2.1 KB
- run_evals.pyruns2.8 KB
- tasks/design-judgment.md3.7 KB
- tasks.json1.6 KB
examples/
- minimal_agent/run.pyruns3.7 KB
- reference-agent-design.md4.1 KB
references/
- authoring-and-packaging.md4.6 KB
- openai-crosscheck.md6.5 KB
- patterns.md16.5 KB
scripts/
- doc_freshness_check.pyruns7.4 KB
- sources.json3.8 KB
tests/
- test_doc_freshness_check.pyruns4.0 KB
- test_minimal_agent.pyruns1.1 KB
- test_run_evals.pyruns1014 B
- CHANGELOG.md2.8 KB
- CONTRIBUTING.md1.7 KB
- DISCLAIMER.md1.0 KB
- .gitignore114 B
- LICENSE1.0 KB
- README.md3.6 KB
Gives 0 of the 12 instructions most architecture codebase skills give in ~2.1k tokens
Counted across 811 of the 1,134 authors here whose files we hold, read 2026-08-07
- Ask the user which candidate to explorein 45 of 811, across 15 files
- Apply the deletion test to suspected shallow modulesin 43 of 811, across 15 files
- Read any relevant architecture decision records firstin 31 of 811, across 8 files
- Use exact glossary terms in every suggestionin 30 of 811, across 10 files
- Accept dependencies instead of creating themin 24 of 811, across 5 files
- Include before and after visualisations for each candidatein 24 of 811, across 5 files
- Read the domain glossary before exploringin 24 of 811, across 6 files
- Return results instead of producing side effectsin 23 of 811, across 4 files
- Explore the codebase for shallow modules and frictionin 23 of 811, across 3 files
- Introduce seams only where things varyin 22 of 811, across 3 files
- Reduce the number of methodsin 21 of 811, across 2 files
- Design deep modules with small interfacesin 21 of 811, across 3 files
Said here and by no other author read
- start with a single augmented LLM call
- do not build an agent if simpler solutions exist
- prefer predictable workflows over autonomous agents
- use the simplest effective system architecture
- show agent planning steps explicitly
- design tools as a tested structured interface
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.