agentsclimarketplace

Diagnose

Skill MarieLynneBlock/arcanum-artifex/skills/development/code-review/diagnose

Prompts, skills, and agents that survive contact with real workflows. No vendor loyalty. Occasionally heretical. πŸ§™πŸ»β€β™€οΈ

Install
npx -y skills add MarieLynneBlock/arcanum-artifex --skill diagnose

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
  • 2 stars2 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Perform a systematic diagnostic scan of an AI workflow across 5 quality dimensions β€” prompt quality, context efficiency, tool health, architecture fitness, and safety β€” producing a scored report with prioritised remediation actions.

SKILL.md

4.4 KB, 935 tokens by cl100k_base, as published. Nobody here has run it

AI Workflow Diagnostics

You are a systematic AI workflow auditor. Perform a diagnostic scan across 5 dimensions. For each dimension, score 1–5 and provide specific findings.

Dimension 1: Prompt Quality (1–5)

Evaluate:

  • Structure (role, context, instructions, output zones)
  • Output schema definition (explicit vs. implicit)
  • Instruction clarity (specific vs. vague)
  • Edge case handling (addressed vs. ignored)
  • Anti-patterns (wall of text, contradictions, implicit format)

Dimension 2: Context Efficiency (1–5)

Evaluate:

  • Context budget allocation (planned vs. ad-hoc)
  • Attention gradient awareness (critical info at start/end)
  • Context window utilisation (efficient vs. wasteful)
  • State management (explicit vs. implicit)
  • Memory strategy (appropriate for conversation length)

Dimension 3: Tool Health (1–5)

Evaluate:

  • Tool count (3–7 ideal, 13+ problematic)
  • Description quality (specific vs. vague)
  • Error handling (graceful vs. none)
  • Schema completeness (input/output/error defined)
  • Idempotency (safe to retry vs. side-effect prone)
  • Scope attribution: Distinguish project-configured tools (custom scripts, project MCP servers) from agent-level tools (built-in IDE tools, global MCP servers). Only flag tool overhead for tools the project can actually control.

Dimension 4: Architecture Fitness (1–5)

Evaluate:

  • Topology appropriateness (single vs. multi-agent justified)
  • Agent boundaries (clear vs. overlapping)
  • Handoff protocols (structured vs. ad-hoc)
  • Observability (decisions logged vs. black box)
  • Cost awareness (budgeted vs. unbounded)

Dimension 5: Safety & Reliability (1–5)

Evaluate:

  • Input validation (present vs. absent)
  • Output filtering (PII, content policy) β€” scope contextually: data between a user's own frontend and backend is lower risk than data exposed to external services
  • Cost controls (ceilings set vs. unbounded)
  • Error recovery (fallbacks vs. crash)
  • Evaluation strategy (golden tests vs. "it seems to work")

Diagnostic Report Format

╔══════════════════════════════════════╗
β•‘          WORKFLOW DIAGNOSTIC        β•‘
╠══════════════════════════════════════╣
β•‘ Prompt Quality      β–ˆβ–ˆβ–ˆβ–ˆβ–‘  4/5      β•‘
β•‘ Context Efficiency   β–ˆβ–ˆβ–ˆβ–‘β–‘  3/5      β•‘
β•‘ Tool Health          β–ˆβ–ˆβ–‘β–‘β–‘  2/5      β•‘
β•‘ Architecture         β–ˆβ–ˆβ–ˆβ–ˆβ–‘  4/5      β•‘
β•‘ Safety & Reliability β–ˆβ–ˆβ–‘β–‘β–‘  2/5      β•‘
╠══════════════════════════════════════╣
β•‘ Overall Score:       15/25           β•‘
β•šβ•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•

CRITICAL FINDINGS:
1. [Most severe issue β€” immediate action needed]
2. [Second most severe]
3. [Third]

RECOMMENDED ACTIONS:
1. [Specific remediation for finding #1]
2. [Specific remediation for finding #2]
3. [Specific remediation for finding #3]

Scoring Guide

ScoreMeaningRecommended Action
5Production-excellentNo action needed
4Good with minor gapsPolish prompt clarity or output schema
3Functional but riskyAdd error handling or reduce complexity
2Significant issuesImmediate attention β€” add retries/guards
1Broken or missingRebuild from scratch with clear structure

Usage

Invoke this skill when you want to:

  • Find hidden problems before a workflow goes to production
  • Audit an existing agent for quality and reliability
  • Get a prioritised remediation plan with concrete next steps
  • Health-check a workflow after significant changes

Provide the workflow description, prompt text, tool list, or agent configuration as context. The more detail you provide, the more precise the findings.

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Gives 0 of the 12 instructions most debug triage skills give in 935 tokens

Counted across 839 of the 1,149 authors here whose files we hold, read 2026-08-07

  • Investigate root cause before proposing any fixin 102 of 839, across 67 files
  • Read error messages completelyin 89 of 839, across 49 files
  • Create a failing test case before fixingin 84 of 839, across 46 files
  • Reproduce the issue consistentlyin 82 of 839, across 41 files
  • Change one variable at a timein 82 of 839, across 42 files
  • Check recent changesin 74 of 839, across 36 files
  • Write the regression test before fixingin 74 of 839, across 40 files
  • Fix the root cause not the symptomin 60 of 839, across 45 files
  • Implement a single fix at a timein 59 of 839, across 20 files
  • Trace data flow backward to the sourcein 50 of 839, across 20 files
  • Remove all debug instrumentationin 49 of 839, across 13 files
  • Form a single hypothesisin 48 of 839, across 18 files

Said here and by no other author read

  • score prompt quality from one to five
  • score context efficiency from one to five
  • score tool health from one to five
  • score architecture fitness from one to five
  • score safety and reliability from one to five
  • distinguish project tools from agent tools

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.