Home / Stacks / LLM Observability & Evals
📊

LLM Observability & Evals

See inside your LLM apps and prove they work — trace every call with Arize Phoenix and LangSmith, monitor cost and quality with Langfuse, score outputs with agent evaluation, and enforce mechanical guardrails

5 skills · Works with Claude Code, Codex, Cursor & more

⚙️ Engineering
RARE

Arize Phoenix Observability

Trace, evaluate, and debug LLM and RAG applications with Arize Phoenix, an open-source AI observability platform. Capture OpenTelemetry spans, run evals on retrieval and responses, and find where your pipeline breaks.

Arize AI 1.9K
Scanned
arize-phoenix observability llm-eval
mkdir -p ~/.claude/skills/arize-phoenix && curl -fsSL https://raw.githubusercontent.com/diegosouzapw/awesome-omni-skill/HEAD/skills/ai-agents/arize-phoenix/SKILL.md -o ~/.claude/skills/arize-phoenix/SKILL.md
⚙️ Engineering
RARE

LangSmith Tracing & Testing

Trace, test, and evaluate LLM and RAG apps with LangSmith. Validate that every query is traced, build evaluation datasets, catch regressions, and turn production traces into repeatable tests.

LangChain 2.5K
Scanned
langsmith llm-eval tracing
mkdir -p ~/.claude/skills/langsmith-testing && curl -fsSL https://raw.githubusercontent.com/diegosouzapw/awesome-omni-skill/HEAD/skills/ai-agents/langsmith-testing/SKILL.md -o ~/.claude/skills/langsmith-testing/SKILL.md
⚙️ Engineering
RARE

Langfuse LLM Observability

Instrument LLM apps with Langfuse for tracing, prompt management, evaluation, and datasets. Captures nested traces across LangChain, LlamaIndex, and OpenAI calls to debug latency, cost, and quality in production.

Community 2.4K
Scanned
langfuse observability tracing
mkdir -p ~/.claude/skills/langfuse && curl -fsSL https://raw.githubusercontent.com/diegosouzapw/awesome-omni-skill/HEAD/skills/ai-agents/langfuse/SKILL.md -o ~/.claude/skills/langfuse/SKILL.md
⚙️ Engineering
RARE

AI Agent Evaluation

Test and benchmark LLM agents with behavioral tests, capability assessments, reliability metrics, and production monitoring — catching the failures that pass benchmarks but break in the real world.

Community 1.7K
Scanned
agent-evaluation testing benchmarking
mkdir -p ~/.claude/skills/agent-evaluation && curl -fsSL https://raw.githubusercontent.com/diegosouzapw/awesome-omni-skill/HEAD/skills/ai-agents/agent-evaluation/SKILL.md -o ~/.claude/skills/agent-evaluation/SKILL.md
⚙️ Engineering
RARE

AI Agent Guardrails

Add mechanical guardrails that stop AI agents from bypassing your rules. Enforces policies with git hooks, secret detection, deployment verification, and import registries — born from real production incidents.

Community 1.6K
Scanned
guardrails ai-safety enforcement
mkdir -p ~/.claude/skills/agent-guardrails && curl -fsSL https://raw.githubusercontent.com/diegosouzapw/awesome-omni-skill/HEAD/skills/ai-agents/agent-guardrails/SKILL.md -o ~/.claude/skills/agent-guardrails/SKILL.md

More Stacks

View all stacks →

Added to wishlist