agentsclimarketplace

Add derivation ai

Skill nyuta01/folio/skills/add-derivation-ai

Portable, AI-native data sheets.

Install
npx -y skills add nyuta01/folio --skill add-derivation-ai

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Wire up a Folio `kind: ai` derivation — the YAML file under `derivations/`, the prompt template (or `prompt_ref`), `output: text` vs `output: json` with `output_schema`, and a one-row materialize smoke. Invoke when the user asks to "have an LLM fill a column", "auto-classify", "summarize each row", or anything that maps a free-text field to a structured value via Claude/OpenAI/etc.

SKILL.md

6.4 KB, ~1.6k tokens by cl100k_base, as published. Nobody here has run it

Add an ai derivation to a Folio sheet

Author a derivations/<target>.yaml of kind: ai, declare the contract field as x-derived: true, and verify with one folio materialize call.

When this skill applies

  • The user says "make Folio fill this column with an LLM" or "classify / summarize / extract / translate every row".
  • The answer is fundamentally fuzzy — free-text classification, summarization, extraction from messy inputs.
  • The user wants Folio to manage the cache, retries, prompt versioning, and cost reporting — i.e. they don't just want a for loop calling the SDK.

This skill does not apply when:

  • The answer is deterministic (use kind: python or kind: sql).
  • The data already exists in a CSV / JSON file (use kind: import).
  • The data lives in another Folio sheet keyed by the same PK (use kind: cross_sheet — see the add-derivation-cross-sheet skill).

Prerequisites

  • A working Folio sheet (see folio-quickstart). You should already have contract.yaml and records.jsonl and folio validate <sheet> exits 0.
  • An ANTHROPIC_API_KEY (or whichever provider your AIClient targets) in the environment, or a StubAIClient for offline runs.
  • folio --help shows the materialize subcommand.

Procedure

  1. Pick the target field name — what column the LLM will fill. Convention: snake_case ASCII, the same as other fields. Add it to contract.yaml with x-derived: true so it's clearly not a human-edited column:

    - name: industry_tag
      logicalType: string
      x-derived: true
      x-inputs: [company_name]      # mirrors `inputs:` in the derivation
    
  2. List inputs — every field the prompt reads. The cache hashes these, so an honest list is what makes the cache correct. If the prompt reads no field (very rare for ai), use inputs: [].

  3. Choose output: text or output: json. One target → text. Multi-target → json plus output_schema. There is no "list" or "tuple" output — you express that as a JSON object.

  4. Write the prompt. Use prompt: for one-liners; prompt_ref: for anything multi-paragraph. prompt_ref paths are relative to the sheet root and the file's bytes go into input_hash, so editing the prompt invalidates the cache (correct behaviour).

  5. Create derivations/<target>.yaml. Skeleton (single target, text output):

    # derivations/industry_tag.yaml
    targets: [industry_tag]
    inputs: [company_name]
    kind: ai
    model: claude-sonnet-4-6
    prompt: |
      Industry of {{ company_name }} in one word.
    output: text
    

    Multi-target with structured JSON output:

    # derivations/enrich.yaml
    targets: [industry, employee_count]
    inputs: [company_name, country]
    kind: ai
    model: claude-sonnet-4-6
    prompt_ref: prompts/enrich.md
    output: json
    output_schema:
      industry: string
      employee_count: integer
    
  6. (Optional) Tune the loop. materialization: is a sub-block:

    materialization:
      respect_human_override: true   # default — skip cells a human edited
      retries: 0                     # default
      retry_delay_seconds: 1.0
    

    Retries cover AIClient errors only (timeouts, rate-limits). A missing output_schema key is a deterministic content error — it does not retry.

  7. Validate, then run materialize on one record. Always smoke-test on a single record before letting the LLM loose on every row:

    folio validate ./customers
    folio materialize ./customers industry_tag \
      --actor agent:demo \
      --ids cust_001
    

    folio materialize takes the target as a positional argument (one at a time; omit it to materialize every derivation), and --ids is comma-separated or repeated.

    The output is the §10.6 envelope:

    {"materialized": 1, "skipped": 0, "failures": [], "total_cost": 0.0021}
    
  8. Inspect the value & provenance. Read it back:

    folio list ./customers --filter "id = ?" --param cust_001
    folio provenance ./customers cust_001 industry_tag
    

    folio provenance takes the record ID and field as positional arguments. The provenance line includes model, input_hash, and cost_usd.

Verify

folio validate <sheet>
folio materialize <sheet> <field> --actor agent:demo --ids <one_id>

Both should exit 0 and the envelope's failures should be [].

Tips & idioms

  • Substitution shape. {{ field }} substitutes the JSON-encoded value — strings get quotes, integers stay bare, arrays become bracket lists. This keeps prompts safe against quotes / newlines in the data.
  • Deterministic first, AI fallback. Common pattern: a python derivation that handles the easy cases, then an ai derivation on the long tail. Keep the AI rows scarce — the cache is the savings.
  • Prompts in their own file. Once a prompt is more than ~3 lines, use prompt_ref: and put the file under prompts/. Reviewers hate diffs of multi-line YAML strings.
  • Multi-target = one cache key. All targets in a multi-target derivation share an input_hash. They update together or stay cached together — that's the invariant.

Common mistakes (don't make them)

  • Forgetting x-derived: true in contract.yaml. The materialize loop still runs, but folio status and human editors will treat the field as a normal user-editable column.
  • output: text with multiple targets. Folio will reject the contract — text output writes one cell. Use output: json + output_schema.
  • Listing inputs the prompt doesn't actually read. The cache re-runs every time those (irrelevant) fields change. Keep inputs: honest.
  • Editing the prompt without expecting a re-run. The prompt body is in the input_hash. That's the design — it means an old, outdated answer cannot stick around silently.
  • Hardcoding an API key in derivations/*.yaml. Folio reads keys from the environment. The YAML stays clean and shippable.

Gives 0 of the 12 instructions most prompt engineering skills give in ~1.6k tokens

Counted across 563 of the 626 authors here whose files we hold, read 2026-08-06

  • ask at most three clarifying questionsin 22 of 563, across 15 files
  • respond in the user input languagein 14 of 563, across 9 files
  • preserve the original intentin 13 of 563, across 11 files
  • Establish baseline metrics and collect representative examplesin 12 of 563, across 2 files
  • Identify failure modes and prioritize high-impact fixesin 12 of 563, across 2 files
  • Apply prompt and workflow improvements with measurable goalsin 12 of 563, across 2 files
  • Roll back quickly if quality or safety metrics regressin 12 of 563, across 2 files
  • validate changes with tests and roll out in controlled stagesin 12 of 563, across 2 files
  • generate quantitative baseline performance reportsin 12 of 563, across 2 files
  • create representative test scenariosin 12 of 563, across 2 files
  • treat prompts as codein 12 of 563, across 5 files
  • test prompts on diverse inputsin 12 of 563, across 8 files

Said here and by no other author read

  • pick target field and mark it x-derived
  • list every field the prompt reads in inputs
  • choose text or json output format
  • write prompt inline or use prompt_ref
  • create derivation yaml file
  • validate the sheet

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.