agentsclimarketplace

Enterprise readiness scorer

Skill AnthonyAlcaraz/agentic-graph-rag-skills/skills/crisis/enterprise-readiness-scorer

Companion repo for Agentic Graph RAG (O'Reilly, Anthony Alcaraz & Sam Julien) — 50 runnable skills + 8 pedagogical notebooks covering all eight chapters, on one moto-mocked AWS DevOps scenario

Install
npx -y skills add AnthonyAlcaraz/agentic-graph-rag-skills --skill enterprise-readiness-scorer

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 2 stars2 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Score a proposed or deployed enterprise agent against the architectural requirements Ch1 argues are non-negotiable: absence of the five fatal flaws of naive vector RAG (context amnesia / relationship blindness / temporal ignorance / reasoning paralysis / tool chaos), calibration of the three agency dimensions (autonomy / action / authority), presence of the four emergent capabilities, and the decision-trace test that separates a real context graph from a relabeled search index. Produces a 0-100 score, a band (PRODUCTION-READY / PILOT-READY / PROTOTYPE / NAIVE-VECTOR), and a gap-closing recommendation per open flaw. Use before greenlighting an enterprise agent for production. NOT for ranking models (Ch1 says the flaws are architectural, not model-quality), NOT for consumer FAQ bots where vector RAG is a fine fit.

SKILL.md

8.0 KB, as published. Nobody here has run it

Enterprise Agentic-Readiness Scorer

Overview

Ch1 opens with a promise and a trap. The promise: an agent that pursues goals instead of answering questions. The trap: "you fire up your favorite LLM, add it to your agentic framework of choice, and connect it to a vector-based RAG system. Should be easy, right? Wrong." A naive vector-only approach creates five fatal flaws that are not bugs but an architectural failure preventing the system from becoming truly agentic.

This skill turns that diagnosis into a score. It checks four things the chapter argues are required for enterprise agency:

  1. The five fatal flaws are cured. Each flaw is cured only by a specific graph capability (context amnesia by evolving memory, relationship blindness by entity relationships, temporal ignorance by temporal evolution, reasoning paralysis by multi-hop reasoning, tool chaos by tool orchestration).
  2. The three agency dimensions are calibrated. Autonomy, action, and authority are sliding scales, not binary — and Ch1's point is calibration, not maximization (a real-estate agent has high autonomy but deliberately low pricing authority).
  3. The four emergent capabilities are present. Autonomous decision-making, contextual understanding, strategic tool utilization, memory persistence.
  4. The decision-trace test passes. Per Marple's test in the Enterprise Context Graphs section: can the system tell you not just what happened, but what alternatives were considered and rejected?

When to Use

  • Before greenlighting an enterprise agent for production deployment
  • Reviewing a vendor's "context graph" claim against the rejected-alternatives test
  • Comparing a naive-vector prototype to a graph-augmented redesign
  • Architecture review where someone proposes "just add a bigger vector store"

Phrases: "is this agent production-ready", "enterprise agentic readiness", "score my RAG architecture", "are we naive vector RAG", "context graph vs search index".

When NOT to Use

  • Ranking LLMs by quality — Ch1 is explicit that the flaws are architectural, not model-quality ("cannot be addressed by expanding context windows or refining embedding techniques")
  • Consumer FAQ / support bots where queries map to a text snippet — Ch1 names these as a great fit for plain vector RAG
  • Single-turn Q&A with no actions, state, or temporal evolution

Process

StepInputActionOutputVerification
1profile JSON (graph_capabilities map)lib.score_flaws(caps)(points, cured, open_flaws)each flaw cured only by its mapped capability
2profile JSON (agency map)lib.score_agency(agency)(points, missing dims)scores coverage/calibration, not magnitude
3profile JSON (capabilities map)lib.score_capabilities(caps)(points, missing)proportional to capabilities present
4captures_rejected_alternatives boollib.decision_trace_test(b)(15 or 0, note)binary test per Marple
5full profilelib.assess(profile)score + band + recommendationsscore bounded 0-100; band matches thresholds

Rationalizations

Agent rationalizationDocumented rebuttal
"Our LLM is frontier-grade, so we are production-ready."Ch1: the five flaws "aren't bugs. They represent an architectural failure" and "cannot be addressed by optimizing search or tweaking embedding models." Model quality does not cure an architecture gap.
"We have a vector store, that covers retrieval."A vector store cures none of the five flaws by itself. Score it: all five stay open, band is NAIVE-VECTOR. The flaws are cured by graph capabilities, not by a vector index.
"We logged everything, so we have a context graph."Marple's test (Ch1, The Context Graph): can it tell you what alternatives were rejected? Logging final states is read-time data. The decision-trace test is worth 15 points precisely to catch relabeled search indexes.
"Max out autonomy and authority for a powerful agent."Ch1: agency dimensions are sliding scales and must be calibrated, not maximized. The real-estate agent has high autonomy, near-zero pricing authority. This scorer rewards calibration coverage, not magnitude.

Red Flags

  • Score is high but decision_trace is 0. The agent may pass the capability checklist while recording only outcomes; it will become "a high-fidelity log of failure" if it ever hallucinates (Ch1 counter-thesis).
  • All five flaws open but band is not NAIVE-VECTOR. Scoring bug — open flaws should dominate; recheck the FLAW_CURE mapping.
  • Agency magnitude drives the score. Misreads Ch1: a deliberately low-authority agent is correct design, not a deficiency.

Non-Negotiable Verification

  1. Run the benchmark battery. python cli.py benchmark must report:
    • all-open flaws score 0; all-cured score the full 40
    • each flaw cured only by its mapped graph capability
    • decision-trace test is binary 15/0
    • a perfect profile is exactly 100 and PRODUCTION-READY; empty is NAIVE-VECTOR
    • agency scores calibration coverage, not magnitude
  2. Verify CLI help. python cli.py --help exits 0 and prints the SKILL.md description.

Security Posture

  • Prompt injection. The profile JSON is untrusted input - a vendor can self-report flattering capability booleans to inflate the score. The scorer only reads fixed keys against a fixed rubric; adversarial field values can skew the score but never execute, so verify claimed capabilities (especially captures_rejected_alternatives) against evidence before trusting a band.
  • Data exfiltration. No network calls, no file writes. Architecture profiles may describe internal systems; they stay in-process and surface only in the stdout report the caller owns.
  • Privilege escalation. No shell invocation, no eval, no dynamic import. A PRODUCTION-READY band is advisory - it authorizes nothing; the deployment gate that consumes the score owns the actual go/no-go decision.

Source Attribution

Distilled from Agentic GraphRAG (O'Reilly) by Anthony Alcaraz and Sam Julien — Ch1: The Crisis of Agentic AI, specifically the five-fatal-flaws opening (the naive-vector failure), the "Defining Agentic AI" agency dimensions and capabilities, and the "Enterprise Context Graphs" decision-trace test. Supporting references: Singhal 2012 (strings to things), Lilian Weng 2023 (LLM-as-brain), Arvind Jain / Glean context-data-platform, Kirk Marple / Graphlit rejected-alternatives test.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.