Best LLM evals & observability tools (2026)
The verdict
Langfuse — open-source LLM observability you can self-host, with tracing, evals, and cost tracking; the default when you want ownership. LangSmith is the native path if you're already in the LangChain ecosystem; Braintrust leads when eval-first development is the culture you want.
The contenders — v2026.Q3
| Tool | Price | Best for |
|---|---|---|
| Langfuse | Hobby free (50k units/mo) · Core $29/mo · Pro $199/mo · Enterprise $2,499/mo · self-host free (MIT core) | The one-tool default: tracing, prompt management, LLM-as-judge evals, and datasets in a single open-source platform you can self-host in minutes or run on a $29 cloud plan. |
| Braintrust | Starter free (1GB data, 10k scores) · Pro $249/mo · Enterprise custom (on-prem or hosted); 6–12 mo free for startups | Teams that treat evals as the product-development loop itself — experiments, datasets, human review, and a purpose-built trace store (Brainstore) with the best eval workflow in the category. |
| LangSmith | Developer free (5k traces/mo) · Plus $39/seat/mo (10k traces) · then pay-as-you-go · Enterprise custom (self-host/hybrid) | Anyone building on LangChain or LangGraph — zero-config tracing of the framework 35% of the Fortune 500 touches, plus deployment and agent infra in the same subscription. |
| Arize Phoenix | OSS free to self-host (ELv2) · free Phoenix Cloud instances · Arize AX (commercial platform) custom | OTel purists and self-hosters — vendor-neutral OpenTelemetry tracing plus evals that stay entirely in your infrastructure, with a funded enterprise platform (AX) behind it when you outgrow free. |
| W&B Weave | Free tier (1GB/mo Weave ingestion) · Pro from $60/mo (1.5GB/mo incl., then usage) · Enterprise custom incl. self-managed | Teams already in the W&B ecosystem, or running multi-agent systems that want sessions/turns/sub-agents as first-class trace concepts rather than bolted-on spans. |
How to choose
Forty percent of agent projects die from unproven ROI — evals are how the other sixty percent prove it. Non-negotiable for anything customer-facing: trace every production call, score a sample weekly, and keep a regression suite of your worst failures. Pair with Guardrails (Gd).
- If you want one tool, no lock-in, and a price that doesn't scale with panic
- Langfuse — MIT-core self-host or $29/mo cloud, now with ClickHouse's balance sheet behind it (Jan 2026).
- If evals drive your product decisions and you have real budget
- Braintrust — the eval UX Notion, Replit, and Cloudflare standardized on; take the 6–12 months free startup credit.
- If your stack is LangChain/LangGraph
- LangSmith — zero-config tracing plus deployment in one bill; its 12x YoY trace growth is the ecosystem talking.
- If you require traces to never leave your infra, at zero license cost
- Arize Phoenix (ELv2, OTel-native) or Opik (Apache-2.0, fully self-hostable) — the two credible free self-host paths.
- If you're shipping anything customer-facing without evals in CI
- Stop — wire pytest-style checks (DeepEval, promptfoo) into CI this week. Gartner's 40%-canceled-by-2027 cohort is made of teams that couldn't prove ROI.
Beyond these five, we track 21 more tools in this category — including 6 dead, renamed, or sunsetting. The full field, the comparison matrix, and every source live on the Evals & Observability element page.
Also consider — combining elements
Method
From edition v2026.Q3 of the elems table, verified 2026-08-06. Elements are jobs, not brands; picks are editorial and never paid for — see the independence charter. When a tool loses its seat, the changelog records the succession.
Answer five questions and get this personalized to your stage, budget, and focus — no email required to see your stack.
Build your stack →