Vital signs
Why it's on the table
On the table, Evals & Observability (Ev) is seat 47 of 58, in the Trust & Compliance family. It is an emerging element — the job is real and here to stay, but the leaderboard still changes quarterly. Choose for this quarter, hold loosely, and watch the changelog. It is marked essential: a company that leaves this seat empty is running with a gap that compounds. It sits in the lowest paid band — lunch money against the hours it returns.
Evals & Observability: the top 5 — v2026.Q3
Edition v2026.Q3 · ranking, pricing and status verified 2026-08-06.
1LangfuseLangfuse (by ClickHouse since Jan 2026)
Hobby free (50k units/mo) · Core $29/mo · Pro $199/mo · Enterprise $2,499/mo · self-host free (MIT core)Best for The one-tool default: tracing, prompt management, LLM-as-judge evals, and datasets in a single open-source platform you can self-host in minutes or run on a $29 cloud plan.
The open-source category leader — 32.6k GitHub stars, 300+ contributors, MIT-licensed core, batteries included (traces, prompt versioning, evals, playground, datasets). The ClickHouse acquisition (Jan 2026) removed the two classic OSS risks at once: funding runway and query-performance ceilings — Langfuse v4 shipped 'real-time, up to 165× faster' on the new backend. The $29 Core tier is the cheapest credible paid entry among the leaders.
Watch Eval UX and experiment workflows trail Braintrust's — Langfuse is observability-first, evals-second. Post-acquisition roadmap now serves ClickHouse's platform ambitions too; the /ee folders are not MIT, so 'fully open source' has an asterisk.
2BraintrustBraintrust Data
Starter free (1GB data, 10k scores) · Pro $249/mo · Enterprise custom (on-prem or hosted); 6–12 mo free for startupsBest for Teams that treat evals as the product-development loop itself — experiments, datasets, human review, and a purpose-built trace store (Brainstore) with the best eval workflow in the category.
Eval-first where everyone else is tracing-first, and the logo wall proves it: Notion, Replit, Cloudflare, Ramp, and Dropbox are named customers of record. $80M Series B led by ICONIQ (Feb 17, 2026) on top of a $36M a16z Series A. Shipping velocity is the tell — Loop (agent that turns production data into evals, Nov 2025), Topics auto-discovery GA (Jun 2026), CLI + MCP support (Apr 2026).
Watch Closed source with a proprietary data store — the deepest lock-in of the top five. $249/mo Pro plus usage ($3/GB, $1.50/1k scores) makes it the priciest non-enterprise entry; the free tier's 14-day retention is the shortest here.
3LangSmithLangChain
Developer free (5k traces/mo) · Plus $39/seat/mo (10k traces) · then pay-as-you-go · Enterprise custom (self-host/hybrid)Best for Anyone building on LangChain or LangGraph — zero-config tracing of the framework 35% of the Fortune 500 touches, plus deployment and agent infra in the same subscription.
The gravity play: LangChain + LangGraph pull 90M monthly downloads, and LangSmith commercial trace volume grew 12x year-over-year (Oct 2025) — the framework funnel works. The $125M Series B at $1.25B (IVP, Oct 20, 2025) bought a genuine platform: observability, evals, prompt engineering, and agent deployment under one roof, with self-hosted enterprise as a real option.
Watch Closed source, and outside the LangChain ecosystem it's just another good tracer — the instrumentation advantage evaporates. Per-seat + LCU/LSU metered pricing is the hardest here to forecast. Framework coupling cuts both ways if you migrate off LangGraph.
4Arize PhoenixArize AI
OSS free to self-host (ELv2) · free Phoenix Cloud instances · Arize AX (commercial platform) customBest for OTel purists and self-hosters — vendor-neutral OpenTelemetry tracing plus evals that stay entirely in your infrastructure, with a funded enterprise platform (AX) behind it when you outgrow free.
The standards bet: Phoenix's OpenInference instrumentation rides OpenTelemetry, so traces are portable by construction — 3M+ monthly Phoenix downloads and 22M+ monthly OTel instrumentation downloads (Aug 2026). Backing is the deepest in pure observability: Arize's $70M Series C (Feb 2025) was billed as the largest-ever investment in AI observability, with Datadog and M12 on the cap table.
Watch ELv2, not MIT/Apache — fine for self-hosting, but not fully open. Prompt management and eval workflow are thinner than Langfuse/Braintrust; the upsell path to Arize AX is where the polish (and the price) lives.
5W&B WeaveWeights & Biases (a CoreWeave company)
Free tier (1GB/mo Weave ingestion) · Pro from $60/mo (1.5GB/mo incl., then usage) · Enterprise custom incl. self-managedBest for Teams already in the W&B ecosystem, or running multi-agent systems that want sessions/turns/sub-agents as first-class trace concepts rather than bolted-on spans.
The most agent-native data model of the leaders — sessions, turns, steps, tools, and sub-agents are primitives, with pre-built safety scorers (toxicity, PII, hallucination) and an API that lets coding agents like Claude Code read production data and run eval loops autonomously. CoreWeave's acquisition of W&B (closed 2025) gives it GPU-cloud distribution and staying power the point solutions lack.
Watch It's a module inside a bigger ML platform — if you don't want experiment tracking and model registry, you're navigating around them. Post-acquisition, W&B's roadmap now serves CoreWeave's cloud strategy; ingestion overage pricing punishes verbose traces.
Evals & Observability: the top 8 compared
Edition v2026.Q3 · ranking, pricing and status verified 2026-08-06.
| Tool | Entry price | Open source | Self-host | OTel-native | Eval workflow | Prompt mgmt | Sweet spot |
|---|---|---|---|---|---|---|---|
| Langfuse | Free · $29/mo | Yes (MIT core) | Yes, free | Yes | Good (judge + code + human) | Yes | one-tool default |
| Braintrust | Free · $249/mo | No | Enterprise only | Partial | Best-in-class | Yes | eval-driven product loops |
| LangSmith | Free · $39/seat | No | Enterprise only | Partial (OTel export) | Good | Yes | LangChain/LangGraph shops |
| Arize Phoenix | Free | Yes (ELv2) | Yes, free | Yes (OpenInference) | Good | Basic | OTel-standard self-hosting |
| W&B Weave | Free · $60/mo | No (SDK only) | Enterprise only | Yes | Good + safety scorers | Basic | multi-agent tracing |
| Opik (Comet) | Free | Yes (Apache-2.0) | Yes, free | Yes | Good | Yes | unrestricted OSS self-host |
| DeepEval / Confident AI | Free · $200/mo | Framework (Apache-2.0) | Enterprise | Partial | Strong (pytest-style) | Yes | evals-as-unit-tests in CI |
| Helicone | Free · $79/mo | Yes | Yes | Via gateway | Basic | Basic | cost/usage logging via proxy |
How to choose your evals & observability
- If you want one tool, no lock-in, and a price that doesn't scale with panic
- Langfuse — MIT-core self-host or $29/mo cloud, now with ClickHouse's balance sheet behind it (Jan 2026).
- If evals drive your product decisions and you have real budget
- Braintrust — the eval UX Notion, Replit, and Cloudflare standardized on; take the 6–12 months free startup credit.
- If your stack is LangChain/LangGraph
- LangSmith — zero-config tracing plus deployment in one bill; its 12x YoY trace growth is the ecosystem talking.
- If you require traces to never leave your infra, at zero license cost
- Arize Phoenix (ELv2, OTel-native) or Opik (Apache-2.0, fully self-hostable) — the two credible free self-host paths.
- If you're shipping anything customer-facing without evals in CI
- Stop — wire pytest-style checks (DeepEval, promptfoo) into CI this week. Gartner's 40%-canceled-by-2027 cohort is made of teams that couldn't prove ROI.
Evals & Observability: the whole field
21 more tools tracked in this category, including 6 dead, renamed, or sunsetting — a reference that hides the graveyard isn't one. Verified 2026-08-06.
| Tool | Maker | What it is | Entry | Status |
|---|---|---|---|---|
| Opik | Comet | The other big OSS platform — 20.7k stars, Apache-2.0, fully self-hostable including backend; the near-miss for the top 5 | free | active |
| DeepEval / Confident AI | Confident AI | Pytest-for-LLMs framework (15.4k stars, Apache-2.0) + commercial platform from $200/mo; the CI-evals standard | free · $200/mo | active |
| promptfoo | Promptfoo Inc. | MIT eval + red-teaming CLI, 21.3k stars, local-first; its security half overlaps element Gd · Guardrails | free | active |
| Helicone | Helicone | One-line proxy for LLM logging and cost analytics (5.8k stars); observability via gateway, thinner on evals | free · $79/mo | active |
| OpenAI Evals | OpenAI | The 2023 pioneer repo (18.6k stars, MIT) is quiet; the live product is the Evals API/dashboard inside the OpenAI platform — OpenAI-centric by design | free + tokens | fading |
| Arize AX | Arize AI | Commercial sibling of Phoenix — enterprise agent engineering platform; $70M Series C (Feb 2025) | custom | active |
| Galileo | Galileo (Rungalileo) | Enterprise 'eval engineering' — Luna distilled evaluator models claim 96% cheaper production scoring; NVIDIA, HP, MongoDB testimonials | free tier · custom | active |
| Datadog LLM Observability | Datadog | The incumbent play: agent tracing + experiments + evaluators inside the APM you already pay for; free to 40k spans, Pro from $160/mo | free · $160/mo | active |
| New Relic AI Monitoring | New Relic | APM-bundled LLM observability; fine if you're already a customer, nobody's first choice for evals | usage-based | active |
| Patronus AI | Patronus AI | Evaluation API and research-grade judges (Lynx hallucination model, Percival agent debugger) | free tier | active |
| Ragas | Ragas (ex-Exploding Gradients) | The default OSS metric library for RAG evals (faithfulness, context precision); framework, not platform | free | active |
| Maxim AI | Maxim | Agent simulation + evals + observability platform; aggressive on agent-testing use cases | free tier | active |
| LangWatch | LangWatch | EU-based open-core LLM ops with optimization studio (DSPy-powered) | free tier | active |
| Lunary | Lunary | Lightweight open-source observability + prompt management; small but steady | free tier | active |
| PromptLayer | PromptLayer | Prompt-management-first with evals attached; popular with non-engineer prompt owners | free tier | active |
| Traceloop (OpenLLMetry) | Traceloop | Maintainer of OpenLLMetry, the OTel LLM instrumentation many platforms ingest; thin commercial layer | free tier | active |
| Humanloop | Humanloop → Anthropic | First-mover LLM eval/prompt platform; team acqui-hired by Anthropic, platform sunset with migration guides (2025) | — | dead |
| TruEra / TruLens | TruEra → Snowflake | ML-observability pioneer acquired by Snowflake (May 2024); TruLens OSS evals live on under Snowflake | free (OSS) | acquired |
| Aporia | Aporia → Coralogix | ML/LLM observability + guardrails, acquired by Coralogix (Dec 2024); now Coralogix's AI Center | — | acquired |
| Weights & Biases (company) | W&B → CoreWeave | The whole company was acquired by CoreWeave (2025, reported ~$1.7B) — Weave continues as its LLM-ops arm | — | acquired |
| Langfuse (company) | Langfuse → ClickHouse | Acquired by ClickHouse (Jan 2026); product continues under its own brand — see top 5 | — | acquired |
Evals & Observability: the category in numbers
Edition v2026.Q3 · ranking, pricing and status verified 2026-08-06.
- Gartner: 40%+ of agentic AI projects will be canceled by end of 2027 — escalating costs, unclear business value, inadequate risk controls (Jun 25, 2025). This is the demand driver for the whole category [src]
- Braintrust $80M Series B led by ICONIQ (Feb 17, 2026); ~$800M valuation per press reporting; named customers Notion, Replit, Cloudflare, Ramp, Dropbox [src]
- LangChain $125M at $1.25B (IVP, Oct 20, 2025); LangSmith commercial trace volume 12x YoY — evals/observability is the monetization layer of the framework business [src]
- Consolidation wave: Langfuse → ClickHouse (Jan 2026), W&B → CoreWeave (2025), Humanloop team → Anthropic with platform sunset (2025), Aporia → Coralogix (Dec 2024), TruEra → Snowflake (May 2024) [src]
- Arize $70M Series C (Feb 20, 2025) billed as the largest-ever investment in AI observability; Datadog and Microsoft's M12 on the cap table — incumbents buying visibility into the category [src]
- APM incumbents now bundle the job: Datadog LLM Observability ships datasets/experiments/evaluators from $160/mo, free to 40k spans (Aug 2026) — squeezing point solutions from above [src]
Evals & Observability: method & sources
Ranking criteria: evidence of production adoption (named customers, download/star trajectories), pricing accessibility for a small team, lock-in risk (license + data portability), and eval-workflow depth — not just tracing. Conflicts resolved: Braintrust's $800M valuation comes from press reporting around the round, not the company's own Feb 17, 2026 post (which discloses no valuation) — treat as reported, not confirmed. The W&B/CoreWeave ~$1.7B price is likewise reported, never officially confirmed. OpenAI's evals repo showed commit activity on fetch but community consensus is it's effectively frozen as a benchmark registry; the maintained product is the platform Evals API — we mark the repo 'fading', which a human should sanity-check. Langfuse GitHub shows 32.4k stars vs 32.6k on its pricing page; we cite the vendor page number, the delta is cache lag. Adjacent elements: LLM gateways with logging (Portkey, OpenRouter analytics, LiteLLM) belong to Rt · Model Router even though they sell 'observability'; promptfoo's red-teaming half and Galileo's runtime guardrails overlap Gd · Guardrails; token-cost dashboards alone are Sp · Spend Management. Uptime/APM monitoring of the app around the model stays with classic observability, out of scope here. Ranking criteria: verified commercial traction, independent satisfaction surveys, agent benchmarks, and founder-fit (price floor, lock-in, surfaces). Editorial, never paid — the charter. Machine-readable twin: ev.json.
All sources (22)
- https://langfuse.com/pricing
- https://github.com/langfuse/langfuse
- https://langfuse.com/blog
- https://www.braintrust.dev/pricing
- https://www.braintrust.dev/blog/announcing-series-b
- https://www.braintrust.dev/blog
- https://www.langchain.com/pricing
- https://www.langchain.com/blog/series-b
- https://arize.com/phoenix
- https://arize.com/blog/arize-ai-raises-70m-series-c/
- https://wandb.ai/site/weave
- https://wandb.ai/site/pricing/
- https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
- https://github.com/comet-ml/opik
- https://github.com/promptfoo/promptfoo
- https://github.com/confident-ai/deepeval
- https://github.com/openai/evals
- https://www.helicone.ai/pricing
- https://www.confident-ai.com/pricing
- https://www.datadoghq.com/product/llm-observability/
- https://galileo.ai
- https://humanloop.com
Our take
Forty percent of agent projects die from unproven ROI. Evals are how the other sixty percent prove it. Non-negotiable for anything customer-facing.
Combines with
Appears in compounds
The Agent Stack · The Trust Stack
Guides
This is element 47 of 58. The table is versioned quarterly — when a tool loses its seat, the changelog records the succession.
Explore the full table →