agentsclimarketplace

Sherma

Skill MadaraUchiha-314/sherma/skills/sherma

Build LLM-powered agents with sherma — declarative YAML or programmatic Python, with multi-agent orchestration, skills, hooks, and A2A integration.From its SKILL.md

Install
npx -y skills add MadaraUchiha-314/sherma --skill sherma

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its file declares

Copied from the file, not written here

The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

32.9 KB, ~8.2k tokens by cl100k_base, as published. Nobody here has run it

Sherma Agent Builder

You are a sherma agent builder. Your job is to help the user build LLM-powered agents using the sherma framework. You produce working code — either declarative YAML, programmatic Python, or both — based on the user's requirements.

Input

Parse $ARGUMENTS for a description of the agent to build. If $ARGUMENTS is empty, ask the user what agent they'd like to build.

Decision Tree

When the user describes what they want, gather information in at most 2 rounds of questions.

MUST ASK (if not already clear from input)

  • What does the agent do? (core task / purpose)
  • What tools does the agent need? (existing Python functions, APIs, MCP servers)

ASK IF AMBIGUOUS

  • Declarative (YAML) vs Programmatic (Python)? — Default: declarative YAML. Use programmatic when the user needs custom graph logic, non-standard state management, or advanced LangGraph features.
  • Multi-agent? — Default: single agent. Use multi-agent when the user describes delegation, sub-tasks, or multiple specialized agents.
  • Skills? — Default: no skills. Use skills when the agent should discover and load capabilities on demand.
  • Hooks? — Default: no hooks. Use hooks when the user mentions logging, guardrails, model swapping, or lifecycle customization.
  • Human-in-the-loop? — Default: no interrupts. Use interrupts when the user needs confirmation or input mid-flow.

DEFAULTS (use unless user specifies otherwise)

  • LLM: OpenAI (gpt-4o-mini), provider openai
  • State: messages list (type list, default [])
  • Checkpointer: in-memory (MemorySaver, auto-configured)
  • Version: "1.0.0" for all entities

Quick Reference: Declarative YAML Schema

Top-level keys

data:          # Custom static data referenceable from CEL (free-form mapping)
prompts:       # Prompt definitions (id, version, instructions | instructions_path)
llms:          # LLM declarations (id, version, provider, model_name)
tools:         # Tool imports (id, version, import_path)
skills:        # Skill card references (id, version, skill_card_path or url)
sub_agents:    # Sub-agent declarations (id, version, yaml_path/import_path/url)
mcp_servers:   # MCP server declarations (auto-registers tools into the registry)
hooks:         # Hook executor imports (import_path or url)
checkpointer:  # Checkpointer config (type: memory)
default_llm:   # Default LLM for call_llm nodes (id reference)
agents:        # Agent graph definitions

Node types

TypePurposeKey args
call_llmCall an LLM with prompt + optional toolsllm, prompt, tools, use_tools_from_registry, use_tools_from_loaded_skills, use_sub_agents_as_tools, state_updates
tool_nodeExecute tool calls from last AIMessagetools (optional, restrict to specific tools)
call_agentInvoke another registered agentagent (id+version), input (CEL expression)
data_transformTransform state via CEL → dictexpression (CEL returning a dict)
set_stateSet individual state variablesvalues (map of field → CEL expression)
interruptPause for human inputvalue (required CEL expression)
load_skillsProgrammatically load skillsskill_ids (CEL → list of {id, version})
customLogic defined entirely by hooksmetadata (optional dict for hooks)

Error handling (on_error)

Any call_llm, tool_node, call_agent, or custom node can declare on_error:

on_error:
  retry:              # call_llm only
    max_attempts: 3   # total attempts
    strategy: exponential  # or "fixed"
    delay: 1.0        # base delay (seconds)
    max_delay: 30.0   # cap
  fallback: handler   # node to route to on failure
  • retry wraps only model.ainvoke() (safe, stateless)
  • fallback routes to a recovery node when retries exhaust
  • Error info stored in state["__sherma__"]["last_error"]
  • on_node_error hook only fires if no fallback handles it

Edge types

# Static edge
edges:
  - source: node_a
    target: node_b          # Use __end__ to terminate

# Conditional edge (CEL)
edges:
  - source: reflect
    branches:
      - condition: 'state.messages[size(state.messages) - 1]["content"].contains("DONE")'
        target: __end__
      - condition: 'state.retry_count < 3'
        target: retry
    default: fallback_node  # If no branch matches

# Branching at entry: use `__start__` as source (static or conditional).
# When used, omit the top-level `entry_point` (setting both is an error).
edges:
  - source: __start__
    branches:
      - condition: 'state.mode == "new"'
        target: handle_new
    default: passthrough

Prompt format in call_llm

prompt:
  - role: system              # CEL → string → SystemMessage
    content: 'prompts["my-prompt"]["instructions"]'
  - role: messages            # CEL → list → spliced in place (preserves original roles)
    content: 'state.messages'
  - role: human               # CEL → string → HumanMessage
    content: '"Summarize the above."'

Roles: system, human, ai, messages. The messages role splices conversation history — it is never auto-injected, you must include it explicitly.

Tool binding modes

ModeDescription
tools: [{id, version}]Bind specific tools
use_tools_from_registry: trueBind ALL registered tools
use_tools_from_loaded_skills: trueBind tools from loaded skills only
use_sub_agents_as_tools: trueBind all declared sub-agents as tools
use_sub_agents_as_tools: [{id, version}]Bind specific sub-agents

Dynamic flags are mutually exclusive with each other, but an explicit tools list can combine with any single flag.

Auto-injected tool_node: When a call_llm node has tools, sherma auto-injects a tool_node after it with conditional edges. You do NOT wire this manually.

state_updates (required): Map the LLM response to state field(s) using CEL expressions with llm_response.content and llm_response.tool_calls. Values are deltas passed to LangGraph reducers. The standard pattern is messages: '[llm_response]'. A warning is emitted if tools are bound but messages is not in the mapping.

MCP servers

Top-level mcp_servers: declarations connect to MCP servers and register their tools into the global tool registry, where they are picked up by tools: [{id, version}] or use_tools_from_registry: true.

mcp_servers:
  - id: factset
    transport: streamable_http     # or "sse" / "stdio"
    url: ${FACTSET_MCP_URL}
    headers: { Authorization: "Bearer ${FACTSET_TOKEN}" }
    tool_prefix: "factset__"        # optional, namespaces tool ids
  - id: filesystem
    transport: stdio
    command: uvx
    args: ["mcp-server-filesystem"]

Required fields: url for streamable_http / sse; command for stdio. Each registered tool gets id = <tool_prefix><tool_name> and version = <server.version>.

Environment-variable interpolation

Every YAML string supports ${VAR} and ${VAR:-default} for UPPERCASE_WITH_UNDERSCORES names. Use $$ for a literal $. Lowercase placeholders like ${available_skills} are not substituted at YAML-load time and remain available to the CEL template() function. Missing required env vars raise DeclarativeConfigError listing all unresolved names. Interpolation runs before Pydantic validation.

Agent input/output schema (JSON Schema)

agents.<name>.input_schema: and agents.<name>.output_schema: accept raw JSON Schema dicts. Incoming DataParts tagged agent_input: true are validated against input_schema; outgoing DataParts tagged agent_output: true are validated against output_schema. Both are also published as A2A capability extensions on the agent card. Bad input raises SchemaValidationError; bad output is reported as a failed task event with the validator's message attached.

agents:
  reader:
    output_schema:
      type: object
      required: [ticker]
      properties:
        ticker: { type: string, pattern: "^[A-Z.]+$" }

A programmatic Pydantic schema (passed via the Agent constructor) takes precedence over the YAML form.

Quick Reference: Programmatic Agent

Subclass LangGraphAgent and implement get_graph():

from langgraph.graph import START, MessagesState, StateGraph
from langgraph.graph.state import CompiledStateGraph
from langgraph.prebuilt import ToolNode, tools_condition
from langchain_openai import ChatOpenAI

from sherma.langgraph.agent import LangGraphAgent

class MyAgent(LangGraphAgent):
    api_key: str

    async def get_graph(self) -> CompiledStateGraph:
        llm = ChatOpenAI(model="gpt-4o-mini", api_key=self.api_key)
        llm_with_tools = llm.bind_tools([my_tool])

        async def call_model(state: MessagesState) -> dict:
            system = {"role": "system", "content": "You are helpful."}
            response = await llm_with_tools.ainvoke([system, *state["messages"]])
            return {"messages": [response]}

        graph = StateGraph(MessagesState)
        graph.add_node("agent", call_model)
        graph.add_node("tools", ToolNode([my_tool]))
        graph.add_edge(START, "agent")
        graph.add_conditional_edges("agent", tools_condition)
        graph.add_edge("tools", "agent")
        return graph.compile()

send_message and cancel_task are auto-implemented — you only write get_graph().

Templates

Minimal declarative agent

prompts:
  - id: system-prompt
    version: "1.0.0"
    instructions: >
      You are a helpful assistant.
    # Alternatively, load from a file (relative to the YAML's base_path):
    # instructions_path: "prompts/system-prompt.md"

llms:
  - id: openai-gpt-4o-mini
    version: "1.0.0"
    provider: openai
    model_name: gpt-4o-mini

agents:
  my-agent:
    state:
      fields:
        - name: messages
          type: list
          default: []

    graph:
      entry_point: agent
      nodes:
        - name: agent
          type: call_llm
          args:
            llm:
              id: openai-gpt-4o-mini
              version: "1.0.0"
            prompt:
              - role: system
                content: 'prompts["system-prompt"]["instructions"]'
              - role: messages
                content: 'state.messages'
            state_updates:
              messages: '[llm_response]'

      edges:
        - source: agent
          target: __end__

Declarative agent with tools

prompts:
  - id: system-prompt
    version: "1.0.0"
    instructions: >
      You are a helpful assistant. Use tools when needed.

llms:
  - id: openai-gpt-4o-mini
    version: "1.0.0"
    provider: openai
    model_name: gpt-4o-mini

tools:
  - id: my_tool
    version: "1.0.0"
    import_path: my_package.tools.my_tool

agents:
  my-agent:
    state:
      fields:
        - name: messages
          type: list
          default: []

    graph:
      entry_point: agent
      nodes:
        - name: agent
          type: call_llm
          args:
            llm:
              id: openai-gpt-4o-mini
              version: "1.0.0"
            prompt:
              - role: system
                content: 'prompts["system-prompt"]["instructions"]'
              - role: messages
                content: 'state.messages'
            tools:
              - id: my_tool
                version: "1.0.0"
            state_updates:
              messages: '[llm_response]'

      edges:
        - source: agent
          target: __end__

Multi-agent supervisor

prompts:
  - id: supervisor-prompt
    version: "1.0.0"
    instructions: >
      You are a supervisor. Delegate tasks to sub-agents as needed.

llms:
  - id: openai-gpt-4o-mini
    version: "1.0.0"
    provider: openai
    model_name: gpt-4o-mini

sub_agents:
  - id: worker-agent
    version: "1.0.0"
    yaml_path: worker-agent.yaml   # Relative to this YAML file

agents:
  supervisor:
    state:
      fields:
        - name: messages
          type: list
          default: []

    graph:
      entry_point: planner
      nodes:
        - name: planner
          type: call_llm
          args:
            llm:
              id: openai-gpt-4o-mini
              version: "1.0.0"
            prompt:
              - role: system
                content: 'prompts["supervisor-prompt"]["instructions"]'
              - role: messages
                content: 'state.messages'
            use_sub_agents_as_tools: true
            state_updates:
              messages: '[llm_response]'

      edges:
        - source: planner
          target: __end__

Skill-based agent

prompts:
  - id: discover-skills
    version: "1.0.0"
    instructions: >
      Here are the available skills: ${available_skills}
      1. Call load_skill_md for the most relevant skill.
      2. Respond with a brief summary.
      When a skill is no longer needed, call unload_skill to free
      context window space.

  - id: plan-and-execute
    version: "1.0.0"
    instructions: >
      Use the loaded skill tools to accomplish the user's request.

  - id: reflect
    version: "1.0.0"
    instructions: >
      If complete, respond with "TASK_COMPLETE" followed by the answer.
      Otherwise respond "NEEDS_MORE_WORK".

llms:
  - id: openai-gpt-4o-mini
    version: "1.0.0"
    provider: openai
    model_name: gpt-4o-mini

skills:
  - id: my-skill
    version: "1.0.0"
    skill_card_path: ../skills/my-skill/skill-card.json

agents:
  skill-agent:
    langgraph_config:
      recursion_limit: 50
    state:
      fields:
        - name: messages
          type: list
          default: []

    graph:
      entry_point: discover_skills
      nodes:
        - name: discover_skills
          type: call_llm
          args:
            llm: { id: openai-gpt-4o-mini, version: "1.0.0" }
            prompt:
              - role: system
                content: 'template(prompts["discover-skills"]["instructions"], {"available_skills": string(skills)})'
              - role: messages
                content: 'state.messages'
            tools:
              - id: load_skill_md
              - id: unload_skill
            state_updates:
              messages: '[llm_response]'

        - name: execute
          type: call_llm
          args:
            llm: { id: openai-gpt-4o-mini, version: "1.0.0" }
            prompt:
              - role: system
                content: 'prompts["plan-and-execute"]["instructions"]'
              - role: messages
                content: 'state.messages'
            use_tools_from_loaded_skills: true
            state_updates:
              messages: '[llm_response]'

        - name: reflect
          type: call_llm
          args:
            llm: { id: openai-gpt-4o-mini, version: "1.0.0" }
            prompt:
              - role: system
                content: 'prompts["reflect"]["instructions"]'
              - role: messages
                content: 'state.messages'
            state_updates:
              messages: '[llm_response]'

      edges:
        - source: discover_skills
          target: execute
        - source: execute
          target: reflect
        - source: reflect
          branches:
            - condition: 'state.messages[size(state.messages) - 1]["content"].contains("TASK_COMPLETE")'
              target: __end__
          default: execute

Agent with hooks

hooks:
  - import_path: my_package.hooks.LoggingHook
  - import_path: my_package.hooks.GuardrailHook
  # Remote hooks over JSON-RPC (HTTP):
  # - url: http://localhost:8000/hooks
  # Remote hooks over MCP (stdio or HTTP):
  # - mcp:
  #     command: python
  #     args: ["-m", "my_package.mcp_hook_server"]
  # - mcp:
  #     url: http://localhost:9000/mcp

prompts:
  - id: system-prompt
    version: "1.0.0"
    instructions: >
      You are a helpful assistant.

llms:
  - id: openai-gpt-4o-mini
    version: "1.0.0"
    provider: openai
    model_name: gpt-4o-mini

agents:
  my-agent:
    state:
      fields:
        - name: messages
          type: list
          default: []

    graph:
      entry_point: agent
      nodes:
        - name: agent
          type: call_llm
          args:
            llm: { id: openai-gpt-4o-mini, version: "1.0.0" }
            prompt:
              - role: system
                content: 'prompts["system-prompt"]["instructions"]'
              - role: messages
                content: 'state.messages'
            state_updates:
              messages: '[llm_response]'

      edges:
        - source: agent
          target: __end__

Hook executor pattern:

from sherma import BaseHookExecutor
from sherma.hooks.types import BeforeLLMCallContext

class GuardrailHook(BaseHookExecutor):
    async def before_llm_call(self, ctx: BeforeLLMCallContext) -> BeforeLLMCallContext | None:
        ctx.system_prompt += "\n\nIMPORTANT: Be accurate. Never fabricate data."
        return ctx  # Return modified context

Human-in-the-loop approval (message metadata routing)

Routes based on additional_kwargs and type fields on messages. A hook or custom node tags messages with metadata; downstream edges inspect it via CEL.

prompts:
  - id: draft-prompt
    version: "1.0.0"
    instructions: >
      Draft a response to the user's request. Be thorough.

  - id: revise-prompt
    version: "1.0.0"
    instructions: >
      The reviewer asked for changes. Revise your draft accordingly.

llms:
  - id: openai-gpt-4o-mini
    version: "1.0.0"
    provider: openai
    model_name: gpt-4o-mini

hooks:
  - import_path: my_package.hooks.ApprovalTaggingHook

agents:
  approval-agent:
    state:
      fields:
        - name: messages
          type: list
          default: []

    graph:
      entry_point: draft
      nodes:
        - name: draft
          type: call_llm
          args:
            llm: { id: openai-gpt-4o-mini, version: "1.0.0" }
            prompt:
              - role: system
                content: 'prompts["draft-prompt"]["instructions"]'
              - role: messages
                content: 'state.messages'
            state_updates:
              messages: '[llm_response]'

        # Pause for human review — pass the draft as interrupt value
        - name: get_approval
          type: interrupt
          args:
            value: >
              {"type": "approval", "draft": state.messages[size(state.messages) - 1]["content"]}

        # Route based on additional_kwargs set by ApprovalTaggingHook
        # The hook tags the human response with additional_kwargs["decision"]
        - name: revise
          type: call_llm
          args:
            llm: { id: openai-gpt-4o-mini, version: "1.0.0" }
            prompt:
              - role: system
                content: 'prompts["revise-prompt"]["instructions"]'
              - role: messages
                content: 'state.messages'
            state_updates:
              messages: '[llm_response]'

      edges:
        - source: draft
          target: get_approval

        - source: get_approval
          branches:
            # Route using additional_kwargs metadata set by a hook
            - condition: >
                state.messages[size(state.messages) - 1]["additional_kwargs"]["decision"] == "approve"
              target: __end__
          default: revise

        # After revision, go back for another review
        - source: revise
          target: get_approval

The ApprovalTaggingHook sets additional_kwargs["decision"] on the human message during node_exit of the interrupt node. You can also route on the message type field (e.g., state.messages[0]["type"] == "human").

A2A server

from a2a.server.apps import A2AStarletteApplication
from a2a.server.request_handlers import DefaultRequestHandler
from a2a.server.tasks import InMemoryTaskStore
from a2a.types import AgentCard, AgentCapabilities

from sherma import DeclarativeAgent
from sherma.a2a import ShermaAgentExecutor

agent = DeclarativeAgent(
    id="my-agent",
    version="1.0.0",
    yaml_path="agent.yaml",
)

executor = ShermaAgentExecutor(agent=agent)
handler = DefaultRequestHandler(
    agent_executor=executor,
    task_store=InMemoryTaskStore(),
)
card = AgentCard(
    name="My Agent",
    description="Does useful things",
    url="http://localhost:8000",
    version="1.0.0",
    capabilities=AgentCapabilities(streaming=False),
)
app = A2AStarletteApplication(agent_card=card, http_handler=handler).build()
# Serve with: uvicorn main:app --port 8000

Programmatic agent

from sherma.langgraph.agent import LangGraphAgent
from langgraph.graph import START, MessagesState, StateGraph
from langgraph.graph.state import CompiledStateGraph
from langgraph.prebuilt import ToolNode, tools_condition
from langchain_openai import ChatOpenAI
from langchain_core.tools import tool

@tool
def my_tool(query: str) -> str:
    """Description of what the tool does."""
    return "result"

class MyAgent(LangGraphAgent):
    api_key: str

    async def get_graph(self) -> CompiledStateGraph:
        llm = ChatOpenAI(model="gpt-4o-mini", api_key=self.api_key)
        llm_with_tools = llm.bind_tools([my_tool])

        async def call_model(state: MessagesState) -> dict:
            system = {"role": "system", "content": "You are helpful."}
            response = await llm_with_tools.ainvoke([system, *state["messages"]])
            return {"messages": [response]}

        graph = StateGraph(MessagesState)
        graph.add_node("agent", call_model)
        graph.add_node("tools", ToolNode([my_tool]))
        graph.add_edge(START, "agent")
        graph.add_conditional_edges("agent", tools_condition)
        graph.add_edge("tools", "agent")
        return graph.compile()

# Usage:
# agent = MyAgent(id="my-agent", version="1.0.0", api_key="sk-...")
# async for event in agent.send_message(message): ...

CEL Cheat Sheet

CEL (Common Expression Language) is used in YAML for dynamic behavior.

Common patterns

# Access last message content
'state.messages[size(state.messages) - 1]'

# Access last message's content field (for AIMessage objects)
'state.messages[size(state.messages) - 1]["content"]'

# String literal (MUST use inner quotes)
'"hello world"'

# Integer literal
'42'

# Build a dict for data_transform
'{"count": state.count + 1, "status": "done"}'

# Conditional check
'state.messages[size(state.messages) - 1]["content"].contains("COMPLETE")'

# Reference a registered prompt (top-level, no state prefix)
'prompts["my-prompt"]["instructions"]'

# String concatenation
'"Hello, " + state.user_name'

# List size
'size(state.messages)'

# Boolean check
'state.retry_count < 3 && state.status != "failed"'

# Message type (returns "ai", "human", "system", "tool")
'state.messages[size(state.messages) - 1]["type"] == "ai"'

# Message metadata via additional_kwargs
'state.messages[0]["additional_kwargs"]["type"] == "approval_decision"'

# Reference custom static data from the top-level `data:` section
'data.thresholds.max_retries'                                  # nested map access
'data["labels"][state.sentiment]'                             # state-driven lookup
'state.retry_count < data.thresholds.max_retries'             # threshold check

Data binding (what CEL can reference)

All bindings are top-level except state (which uses the state. prefix):

VariableSourceScope
stateagent state.fieldsevery expression
prompts["<id>"]["instructions"]top-level prompts:every expression
llms["<id>"]["model_name"]top-level llms:every expression
skills["<id>"]id/version/name/descriptiontop-level skills: (when declared)every expression
datatop-level data: (when declared) — free-formevery expression
llm_responsecontent/tool_callsthe LLM replyonly in call_llm state_updates

llm_response is the only runtime-scoped binding — it is unavailable in prompts, edges, or non-call_llm nodes.

List macros (built-in)

# Filter: keep elements matching a predicate
'state.messages.filter(m, m["type"] == "human")'               # filter by type
'state.items.filter(x, x > 0)'                                # filter primitives

# Exists: check if any element matches
'state.messages.exists(m, m["type"] == "ai")'                  # any AI message?
'state.messages.exists(m, m["additional_kwargs"]["type"] == "approval_decision")'

# All: check if all elements match
'state.items.all(x, x > 0)'                                   # all positive?

# Map: transform each element
'state.messages.map(m, m["type"])'                             # extract field from each
'state.items.map(x, x * 2)'                                   # double each

# Count matching elements
'size(state.messages.filter(m, m["type"] == "human")) > 0'

# findLast pattern: last() + filter()
'last(state.messages.filter(m, m["type"] == "human"))'         # last human message
'default(last(state.messages.filter(m, m["additional_kwargs"]["type"] == "approval_decision"))["content"], "")'

Custom functions

# JSON: parse structured data from strings
'json(state.response)["status"]'                              # parse JSON, access field
'json(state.messages[0]["content"])["action"]'                 # parse message content as JSON
'jsonValid(state.data)'                                        # check if string is valid JSON
'jsonValid(state.data) && json(state.data)["ok"] == true'      # guard + parse

# Safe access: fallback on errors
'default(json(state.response)["action"], "continue")'          # fallback if parse/key fails
'default(state.missing_field, 0)'                              # fallback for missing state

# List utilities
'last(state.items)'                                            # last element (error if empty)
'last(state.messages.filter(m, m["type"] == "ai"))'            # findLast pattern
'default(last(state.messages.filter(m, ...))["content"], "")'  # safe findLast with fallback

# String extensions (cel-go compatible)
'state.tags.split(",")'                                        # split → list
'"  hello  ".trim()'                                           # strip whitespace
'"HELLO".lowerAscii()'                                         # lowercase
'"hello".upperAscii()'                                         # uppercase
'"hello world".replace("world", "CEL")'                        # replace substrings
'"hello".indexOf("ll")'                                        # find index (or -1)
'["a", "b", "c"].join(", ")'                                   # join list → string
'"hello world".substring(0, 5)'                                # extract substring

# Templating: ${key} substitution from a map
'template("Hello ${name}!", {"name": state.user})'             # basic substitution
'template(prompts["plan"]["instructions"], {"skills": state.skill_list})'  # prompt templating

# Combining functions
'json(state.data.trim())["name"].lowerAscii()'                 # trim → parse → lowercase
'state.tags.split(",").join(" | ")'                            # split → rejoin

Gotchas

  • State access requires state. prefix: Use state.messages, state["counter"] — not bare messages or counter. Extra vars like prompts, llms, skills, and data are top-level.
  • String literals need inner quotes: '"hello"' not 'hello'. Without inner quotes, CEL treats it as a variable name.
  • size() not len(): CEL uses size() for list/string length.
  • Map syntax: {"key": value} — keys must be strings in double quotes.
  • No f-strings: Use template() for substitution or + for concatenation: template("Count: ${n}", {"n": state.count}) or '"Count: " + string(count)'
  • YAML quoting: Always single-quote the outer CEL expression to avoid YAML parsing issues.
  • default() is expression-level: default(expr, fallback) must be the outermost call — it catches errors in expr and returns fallback.

Common Gotchas

  1. Auto-injected tool_node: When call_llm has tools, sherma auto-injects a tool_node with conditional edges. Do NOT add your own tool_node for tool-equipped call_llm nodes.

  2. Prompt messages role is not auto-injected: You MUST explicitly include role: messages in the prompt to get conversation history. Without it, the LLM has no context of previous messages.

  3. import_path must be importable: Tool, hook, and agent import_path values must be importable Python paths (dot-separated). The module must be on sys.path. For tools, the path must point to a @tool-decorated function.

  4. String values in set_state: Each value is a CEL expression. String literals require inner quotes: '"ready"' not 'ready'.

  5. Relative paths resolve from YAML file directory: skill_card_path, sub-agent yaml_path, etc. resolve relative to the YAML file's parent directory, not the working directory.

  6. default_llm requires the LLM to be declared in llms: The default_llm field references an LLM by id — it must exist in the llms list.

  7. Interrupt value: The interrupt node requires a value CEL expression that is evaluated against state. Use a string literal (e.g., '"question"') or reference state (e.g., state.messages[size(state.messages) - 1].content).

  8. use_tools_from_loaded_skills: Only works after skills have been loaded via load_skill_md. Put the discovery node before the execution node.

  9. custom node hooks can access registries: NodeExecuteContext exposes ctx.registries (a RegistryBundle) so node_execute hooks can call chat models (ctx.registries.chat_models[llm_id]), resolve tools, render prompts, etc. without wiring dependencies at agent-init time. Remote (JSON-RPC and MCP) hooks receive registries=None because it contains live Python objects.

  10. MCP hook tool names use . not /: MCP restricts tool names to [A-Za-z0-9_.-]+, so MCP-based hook tools are named hooks.before_llm_call (dot separator), not hooks/before_llm_call. The executor discovers tools with the hooks. prefix at connect time and only dispatches the hooks the server has registered.

Process

Follow these steps when building an agent:

  1. Read the input — Parse $ARGUMENTS for the agent description.
  2. Ask clarifying questions — At most 2 rounds. Use defaults for anything not specified.
  3. Choose approach — Declarative YAML (default) or programmatic Python.
  4. Generate files:
    • For declarative: agent.yaml + main.py (entry point) + tool files if needed
    • For programmatic: agent.py (LangGraphAgent subclass) + main.py + tool files
    • For A2A server: add server.py with A2A boilerplate
  5. Verify — Check that:
    • All import_path values point to real modules
    • Prompt references match declared prompt IDs
    • LLM references match declared LLM IDs
    • Tool references match declared tool IDs
    • Edge targets match declared node names
    • __end__ is reachable from every path

Full API Surface

Entities

EntityBase, Prompt, LLM, Tool, Skill, SkillCard, SkillFrontMatter, LocalToolDef, MCPServerDef

Agents

Agent, LocalAgent, RemoteAgent, LangGraphAgent, DeclarativeAgent

Registries

Registry, RegistryEntry, RegistryBundle, TenantRegistryManager, PromptRegistry, LLMRegistry, ToolRegistry, SkillRegistry, AgentRegistry

Hooks

HookExecutor, BaseHookExecutor, HookManager, HookType, HookFastAPIApplication, HookStarletteApplication, RemoteHookExecutor, MCPHookExecutor, MCPHookTransport, MCPHookServer. Remote hook servers (JSON-RPC and MCP) are authored as BaseHookExecutor subclasses with typed contexts — same as in-process hooks.

Declarative

DeclarativeConfig, load_declarative_config

Skills

create_skill_tools

Schema Utilities

SCHEMA_INPUT_URI, SCHEMA_OUTPUT_URI, validate_data, schema_to_extension, make_schema_data_part, create_agent_input_as_message_part, create_agent_output_as_message_part, get_agent_input_from_message_part, get_agent_output_from_message_part

Types

EntityType, Markdown, Protocol, DEFAULT_TENANT_ID

Exceptions

ShermaError, EntityNotFoundError, VersionNotFoundError, RegistryError, RemoteEntityError, DeclarativeConfigError, GraphConstructionError, CelEvaluationError, SchemaValidationError

A2A Integration

ShermaAgentExecutor, a2a_to_langgraph, langgraph_to_a2a, combine_ai_messages

HTTP Client (sherma.http)

get_http_client, HttpClientFactory, mcp_http_client_factory, McpHttpClientFactory. A single shared httpx.AsyncClient is cached in a ContextVar and used by LLM, A2A, registries, remote hooks (JSON-RPC), and MCP hooks (streamable_http / sse). Configure headers/auth/event hooks/transport on it once via get_http_client(...).

For detailed signatures, see references/api-reference.md.

What ships with it: 13 files

155.4 KB alongside SKILL.md, 1 of them executable

Gives 0 of the 12 instructions most agent orchestration skills give in ~8.2k tokens

Counted across 742 of the 995 authors here whose files we hold, read 2026-08-07

  • Reference existing artifacts by path or URLin 53 of 742, across 25 files
  • Run the full test suite after integrating changesin 51 of 742, across 19 files
  • Dispatch one agent per independent problem domainin 50 of 742, across 17 files
  • Verify fixes do not conflictin 45 of 742, across 13 files
  • Include a suggested skills section in the documentin 45 of 742, across 17 files
  • Redact sensitive informationin 41 of 742, across 11 files
  • Save to the temporary directory of the operating systemin 39 of 742, across 10 files
  • Tailor the document to user-provided focus argumentsin 39 of 742, across 9 files
  • Spot check agent changes for systematic errorsin 34 of 742, across 7 files
  • Write a handoff document summarising the current conversationin 31 of 742, across 6 files
  • Assign each agent a specific scopein 23 of 742, across 8 files
  • Provide specific scope and clear goalin 23 of 742, across 5 files

Said here and by no other author read

  • default to declarative yaml configuration
  • use programmatic python for custom graph logic
  • explicitly splice conversation history in prompt definitions
  • include state updates on call_llm nodes
  • map llm responses to state fields using CEL expressions
  • set version to 1.0.0 for all entities

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 326,782. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.