agentsclimarketplace

Observe ai production

Skill hiteshbandhu/skills-i-use/skills/ai-engineer-talks/observe-ai-production

Drop-in skills and plugins for your AI development workflows

Install
npx -y skills add hiteshbandhu/skills-i-use --skill observe-ai-production

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Observes and improves production LLM apps with Arize Phoenix—tracing, layered agent evals, prompt learning loops, and PM/engineering eval pipelines. Use when the user mentions Arize, Phoenix, OpenInference, agent router evals, prompt optimization from traces, or shipping AI with production observability.

SKILL.md

2.0 KB, as published. Nobody here has run it

Observe AI production (Arize)

Action playbook from six Arize @ AI Engineer talks. Do not summarize talks — pick a workflow and execute it.

Supporting files: workflows.md · source-index.md

Optional: ./skill-outputs/observe-ai-production/


Step 0 — Pick workflow

What is the user trying to do?
├─ Layered agent evals (router, skills, convergence)     → agent-evals
├─ Hands-on Phoenix agent eval workshop                  → phoenix-workshop
├─ Prompt learning from production traces                → prompt-learning
├─ PM-facing eval frameworks and release gates           → pm-evals
└─ Org-scale eval pipelines (CI, versioning)             → eval-pipelines

See workflows.md. Related: run-llm-evals for cross-vendor eval theory.


Install

cp -r skills/observe-ai-production ~/.claude/skills/
cp -r skills/observe-ai-production ~/.cursor/skills/
cp -r skills/observe-ai-production ~/.codex/skills/

Source: playlists/arize-ai-engineer/.


Cross-cutting rules

RuleSource
Eval router decisions, not only final answers[src-001 @ 0:08:21]
Trace every tool/LLM call for agent eval substrate[src-002 @ 0:06:18]
Calibrate LLM judges on human labels before automation[src-002 @ 0:89:37]
Close loop: observability → dataset → prompt patch[src-003 @ 0:20:00]
Version datasets/scorers like production code[src-005 @ 0:10:00]

Output

Name workflow; save artifacts to ./skill-outputs/observe-ai-production/ when requested; do not auto-commit.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.