agentsclimarketplace

Bootstrap template evaluation

Skill Goodeye-Labs/truesight-mcp-skills/skills/bootstrap-template-evaluation

Fastest route to a deployed live evaluation using a pre-built Truesight template. Use when the user wants a quick start without building judgment configs from scratch.From its SKILL.md

Install
npx -y skills add Goodeye-Labs/truesight-mcp-skills --skill bootstrap-template-evaluation

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

3 things to look at

  • reads credentialsReads from 1 credential source: `api_key`.
  • 7 stars7 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
  • runs commandsInstructs the agent to run 4 commands, including `list_templates` and 3 more.

SKILL.md

2.1 KB, 434 tokens by cl100k_base, as published. Nobody here has run it

Bootstrap Template Evaluation

Use this skill when a pre-built template likely covers the target use case.

Interactive Q&A protocol (mandatory)

<HARD-GATE> BEFORE the first scoping question, search for a structured question tool (e.g., `AskUserQuestion` or similar interactive widget) and load it. Use that tool for EVERY scoping question. Fall back to plain-text lettered options ONLY if no such tool exists in the environment. </HARD-GATE>

If template choice is ambiguous, ask one question at a time using the structured question tool (loaded per the HARD-GATE above).

Example question structure:

Which template family best matches your goal?
A) AI writing detection
B) Code quality
C) Unsure, list all templates first

Rules:

  • Ask one question per message.
  • Use the structured question tool for every question. Structure each with a short header, 2-4 options with labels and descriptions, and place the recommended option first. Do not add "(Recommended)" or similar annotations to option labels.
  • Ask one follow-up only when needed.

Workflow

  1. Discover templates:
    • Call list_templates.
  2. Select template:
    • Match use case to template slug.
  3. Provision private dataset:
    • Call provision_template(slug).
  4. Deploy live evaluation:
    • Call create_and_deploy_evaluation(dataset_id).
    • Capture api_key immediately because it is returned only once.
  5. Verify:
    • Run run_eval with representative inputs.
  6. Return deployment artifacts:
    • dataset_id
    • live_evaluation_id
    • verification result

Guardrails

  • If no template fits, hand off to create-evaluation.
  • Do not skip verification after deployment.

Scopes reference

  • list_templates requires datasets:read
  • provision_template requires datasets:write
  • create_and_deploy_evaluation requires evaluations:write, live-evaluations:write
  • run_eval requires live-evaluations:execute

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.