agentsclimarketplace

Choose observability stack

Skill ContextJet-ai/awesome-llm-observability/skills/choose-observability-stack

50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.

Install
npx -y skills add ContextJet-ai/awesome-llm-observability --skill choose-observability-stack

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.

What its author says it does

Copied from the file, not written here

Use this to recommend an LLM observability / evaluation tool or stack for a specific situation. Trigger on "which observability tool should I use", "compare Langfuse vs Phoenix vs LangSmith", "what's the best LLM monitoring for us", or picking an eval/tracing/gateway tool given constraints (self-hosting, budget, compliance, existing stack). Ask about constraints, then recommend from the curated list - don't just name the most popular tool.

The file declares its own license as CC0-1.0. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

3.3 KB, as published. Nobody here has run it

Choose an LLM observability stack

There's no single best tool - the right choice depends on constraints. Gather them, then map to a recommendation. Base recommendations on this repo's curated list (verified tools + licenses), not on hype.

Ask these constraints first

  1. Deployment: SaaS OK, or must self-host / on-prem (data residency, regulated industry)?
  2. Primary need: tracing/cost, evaluation (quality testing), or both? Prompt management too?
  3. Existing stack: already on Datadog/Grafana/OTel? On LangChain? Using a gateway?
  4. Budget/licensing: need a permissive OSS license (MIT/Apache), or is a commercial tier fine? (Note AGPL/Elastic-license implications for embedding.)
  5. Code-change tolerance: want zero-code (proxy) or fine to add an SDK?
  6. Team: engineers, or also non-technical PMs who need a UI?

Map constraints → recommendation

  • Must self-host, permissive license, want everythingLangfuse (MIT core: tracing + evals + prompts) or Comet Opik (Apache-2.0). For eval-heavy local work, Arize Phoenix.
  • Zero code changes, just want cost + logs → a gateway/proxy: Helicone (change base URL), LiteLLM or Portkey (also routing).
  • Already on OTel / want vendor-neutral, future-proof → emit OpenTelemetry GenAI semantic conventions via OpenLLMetry or OpenInference; export to your existing backend.
  • Deep in the LangChain ecosystemLangSmith (tightest integration; SDK OSS, backend commercial).
  • Enterprise APM already (Datadog/New Relic) → use their LLM Observability product to keep one pane of glass.
  • Primary need is evaluation/testing, not dashboardspromptfoo (prompt/RAG + CI), DeepEval (pytest-style), Ragas (RAG metrics). Pair with a tracing tool for online scoring.
  • Regulated / finance / must audit + guardrail → self-hosted tracing (Langfuse/Phoenix) + guardrails (Guardrails AI, LLM Guard for PII/prompt-injection) + strict prompt/PII redaction.

Common production shape

A gateway (cost + routing) + an evaluation framework (quality) + an OTel-native tracing backbone. This keeps cost, quality, and traces decoupled and swappable.

Deliver the recommendation

  • Name a primary tool + a runner-up, each with a one-line why it fits these constraints.
  • Call out license/self-hosting implications explicitly (especially AGPL / Elastic-license for embedding, and SaaS data-egress for regulated data).
  • Link to the tool's row in this repo's README so they can compare stars/license.

Anti-pattern

Recommending the highest-star tool by default. LiteLLM has the most stars but is a gateway - it's the wrong answer for someone who asked for an evaluation framework. Match the tool category to the stated need.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.