Mcp builder
Use when building an MCP server in Python (FastMCP) or Node/TypeScript (MCP SDK) — agent-centric tool design, input schemas, error handling, and the 10-question evaluation harness.From its SKILL.md
npx -y skills add event4u-app/agent-config --skill mcp-builderAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 7 stars7 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
6.7 KB, ~1.6k tokens by cl100k_base, as published. Nobody here has run it
mcp-builder
Author MCP servers that LLMs can drive end-to-end. The quality bar is can the agent finish the workflow, not does the endpoint return 200. This skill is the server-author counterpart to the existing mcp consumer skill.
When to use
- Wrapping an external API or service as MCP tools for an LLM client.
- Adding tools to an existing MCP server (Python FastMCP or TypeScript SDK).
- Reviewing an MCP server before shipping — Phase 4 evaluation gate below.
Do NOT use when:
- You only need to call an MCP server — route to
mcp. - The integration belongs in the host process — write a regular service, not an MCP server.
- The "server" wraps one endpoint with no workflow — a CLI wrapper is enough.
Procedure: Four phases, one tool at a time
Phase 1 — Research & plan
- Agent-centric design. Tools encode workflows, not raw endpoints. Consolidate (
schedule_eventchecks availability and creates the event). Default to human-readable names over IDs. Errors are educational, not just diagnostic ("retry withfilter='active_only'to reduce results"). - Load the protocol. Fetch
https://modelcontextprotocol.io/llms-full.txtonce into context — the canonical spec. - Load the SDK README for the chosen language:
- Python:
https://raw.githubusercontent.com/modelcontextprotocol/python-sdk/main/README.md - TypeScript:
https://raw.githubusercontent.com/modelcontextprotocol/typescript-sdk/main/README.md
- Python:
- Read the target service's API docs in full — auth, rate limits, pagination, error codes, schemas. Skipping this produces incomplete mocks (see
testing-anti-patterns§ Anti-Pattern 4). - Write the plan: tool list with priority, shared utilities (request helper, pagination, formatter), input/output schemas, error strategy, response-detail levels (concise vs detailed), character limits (default 25 000 tokens).
Phase 2 — Implement
- Project layout. Python: single
.pyor modular package; Pydantic v2 withmodel_config. TypeScript: standardpackage.json+tsconfig.jsonstrict mode; Zod schemas with.strict(). - Shared utilities first. API request helper with retry/timeout, error formatter, JSON-vs-Markdown response builder, pagination cursor handling, auth/token cache.
- Per tool:
- Input schema (Pydantic / Zod) with constraints, descriptions, and examples.
- One-line summary + detailed docstring covering purpose, parameters, return shape, when-to-use, when-NOT-to-use, error handling.
- Tool annotations:
readOnlyHint,destructiveHint,idempotentHint,openWorldHint. - Async/await for all I/O. Honor pagination. Truncate to the character limit and signal truncation in the response.
Phase 3 — Review & test
- Code-quality pass: DRY across tools, shared helpers extracted, consistent response shapes, all external calls have error handling, full type coverage.
- Build & syntax:
- Python:
python -m py_compile server.py. - TypeScript:
npm run build; verifydist/index.js.
- Python:
- Run the server safely. MCP servers block on stdio. Either run inside
tmuxand drive from the harness, or wrap withtimeout 5s python server.pyfor a smoke check. Do NOT block your own session by running it in-process.
Phase 4 — Evaluations (10-question harness)
Each evaluation is a question the agent must answer using only the new tools.
Requirements per question — independent, read-only, complex (multiple tool calls), realistic, verifiable (string-comparable answer), stable (answer does not drift over time).
<evaluation>
<qa_pair>
<question>...</question>
<answer>...</answer>
</qa_pair>
<!-- 9 more -->
</evaluation>
Process: enumerate the tools, explore READ-ONLY data, draft 10 questions, solve each yourself first to confirm the answer is reachable and stable.
Output format
- The server source plus the 10-question evaluation XML.
- A README with: install, env vars, transport mode (stdio / sse / http), example tool call.
- A line in
agents/settings/contexts/skills-provenance.ymlif the server was forked from an upstream, or a note that it was authored from scratch.
Gotcha
- "Wrap every endpoint" is the failure mode — agents cannot orchestrate 60 thin tools as well as 12 workflow tools.
- Returning the full upstream payload blows the agent's context. Default to a concise shape with an opt-in detailed mode.
- Pydantic / Zod descriptions are the only documentation the LLM sees at runtime — write them like usage docs, not comments.
- A server that hangs your session usually means stdio transport ran in the main process — move it under
tmuxor use atimeout. - Inflated token claims are not credible without an evaluation harness — Phase 4 is the validation gate, not optional.
Do NOT
- Do NOT mirror REST routes 1:1.
- Do NOT use
any(TypeScript) or untypeddict(Python) in tool I/O. - Do NOT skip the 10-question evaluation — Phase 4 IS the quality bar.
- Do NOT run the MCP server in your main process during testing — it will block.
- Do NOT log tokens, API keys, or full request bodies — sanitize before logging.
Auto-trigger keywords
- mcp server
- model context protocol
- fastmcp
- mcp builder
- agent-centric tools
Provenance
- Upstream protocol: https://modelcontextprotocol.io
- Upstream SDKs: https://github.com/modelcontextprotocol/python-sdk · https://github.com/modelcontextprotocol/typescript-sdk
- Adopted from: an external reference (internal provenance, redacted) — external
./reference/*.mdfile links replaced with inline guidance + upstream URLs. - Cross-linked:
mcp,testing-anti-patterns,api-design. - Provenance registry:
agents/settings/contexts/skills-provenance.yml(entry:mcp-builder). - Iron-Law floor:
verify-before-complete,tool-safety,skill-quality.
Encode usage policy in the description
Workflow sequencing, preconditions, ID/output provenance ("copy ids verbatim,
never from memory"), a mandatory "why" intent field, and turn-end contracts
belong INSIDE this artifact's description/frontmatter — where they fire at the
decision point — not in always-on prose. See
tool-description-as-policy.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.