Field guide · Trust & Compliance · v2026.Q3 · updated 2026-08-06

Best LLM evals & observability tools (2026)

The pick — v2026.Q3
Langfuse
the open-source default  ·  then Braintrust · LangSmith · Arize Phoenix · W&B Weave

The verdict

Langfuse — open-source LLM observability you can self-host, with tracing, evals, and cost tracking; the default when you want ownership. LangSmith is the native path if you're already in the LangChain ecosystem; Braintrust leads when eval-first development is the culture you want.

The contenders — v2026.Q3

The contenders — v2026.Q3, pricing verified 2026-08-06.
ToolPriceBest for
LangfuseHobby free (50k units/mo) · Core $29/mo · Pro $199/mo · Enterprise $2,499/mo · self-host free (MIT core)The one-tool default: tracing, prompt management, LLM-as-judge evals, and datasets in a single open-source platform you can self-host in minutes or run on a $29 cloud plan.
BraintrustStarter free (1GB data, 10k scores) · Pro $249/mo · Enterprise custom (on-prem or hosted); 6–12 mo free for startupsTeams that treat evals as the product-development loop itself — experiments, datasets, human review, and a purpose-built trace store (Brainstore) with the best eval workflow in the category.
LangSmithDeveloper free (5k traces/mo) · Plus $39/seat/mo (10k traces) · then pay-as-you-go · Enterprise custom (self-host/hybrid)Anyone building on LangChain or LangGraph — zero-config tracing of the framework 35% of the Fortune 500 touches, plus deployment and agent infra in the same subscription.
Arize PhoenixOSS free to self-host (ELv2) · free Phoenix Cloud instances · Arize AX (commercial platform) customOTel purists and self-hosters — vendor-neutral OpenTelemetry tracing plus evals that stay entirely in your infrastructure, with a funded enterprise platform (AX) behind it when you outgrow free.
W&B WeaveFree tier (1GB/mo Weave ingestion) · Pro from $60/mo (1.5GB/mo incl., then usage) · Enterprise custom incl. self-managedTeams already in the W&B ecosystem, or running multi-agent systems that want sessions/turns/sub-agents as first-class trace concepts rather than bolted-on spans.

How to choose

Forty percent of agent projects die from unproven ROI — evals are how the other sixty percent prove it. Non-negotiable for anything customer-facing: trace every production call, score a sample weekly, and keep a regression suite of your worst failures. Pair with Guardrails (Gd).

If you want one tool, no lock-in, and a price that doesn't scale with panic
Langfuse — MIT-core self-host or $29/mo cloud, now with ClickHouse's balance sheet behind it (Jan 2026).
If evals drive your product decisions and you have real budget
Braintrust — the eval UX Notion, Replit, and Cloudflare standardized on; take the 6–12 months free startup credit.
If your stack is LangChain/LangGraph
LangSmith — zero-config tracing plus deployment in one bill; its 12x YoY trace growth is the ecosystem talking.
If you require traces to never leave your infra, at zero license cost
Arize Phoenix (ELv2, OTel-native) or Opik (Apache-2.0, fully self-hostable) — the two credible free self-host paths.
If you're shipping anything customer-facing without evals in CI
Stop — wire pytest-style checks (DeepEval, promptfoo) into CI this week. Gartner's 40%-canceled-by-2027 cohort is made of teams that couldn't prove ROI.

Beyond these five, we track 21 more tools in this category — including 6 dead, renamed, or sunsetting. The full field, the comparison matrix, and every source live on the Evals & Observability element page.

Also consider — combining elements

Method

From edition v2026.Q3 of the elems table, verified 2026-08-06. Elements are jobs, not brands; picks are editorial and never paid for — see the independence charter. When a tool loses its seat, the changelog records the succession.

Answer five questions and get this personalized to your stage, budget, and focus — no email required to see your stack.

Build your stack →
Full element page: Ev · Evals & Observability →