47● essential Ev Evals & Observability
Group 11 · Trust & Compliance · element 47 of 58

Evals & Observability

Prove the machine works.

Turns 'it seems fine' into measured quality.

Holders this quarterLangfuse · Braintrust · LangSmith · Arize Phoenix · W&B Weave

Vital signs

NecessityEssential
Price band$ · under $30/mo
MaturityEmerging
Editionv2026.Q3
Last verified2026-08-06

Why it's on the table

On the table, Evals & Observability (Ev) is seat 47 of 58, in the Trust & Compliance family. It is an emerging element — the job is real and here to stay, but the leaderboard still changes quarterly. Choose for this quarter, hold loosely, and watch the changelog. It is marked essential: a company that leaves this seat empty is running with a gap that compounds. It sits in the lowest paid band — lunch money against the hours it returns.

The verdict — v2026.Q3 · verified 2026-08-06
Langfuse
Langfuse, for the default seat — the most-starred open-source LLM engineering platform (32.6k stars, MIT core), a $29/mo cloud entry, and since January 2026 the resources of ClickHouse behind it without the pricing games of a VC burn clock. Braintrust is the pick when evals ARE your product loop and you'll pay for the best eval UX — the platform Notion, Replit, Cloudflare, Ramp, and Dropbox standardized on, freshly armed with an $80M Series B. LangSmith wins if you build on LangChain/LangGraph; Arize Phoenix if you want OpenTelemetry-standard tracing self-hosted for free; W&B Weave if you're already in the W&B/CoreWeave orbit and want agent-native sessions-and-turns tracing. Whatever you pick, pick one: Gartner projects 40%+ of agentic AI projects canceled by 2027 for unproven ROI — evals are how you avoid being in that cohort.

Evals & Observability: the top 5 — v2026.Q3

Edition v2026.Q3 · ranking, pricing and status verified 2026-08-06.

  1. 1LangfuseLangfuse (by ClickHouse since Jan 2026)

    Hobby free (50k units/mo) · Core $29/mo · Pro $199/mo · Enterprise $2,499/mo · self-host free (MIT core)

    Best for The one-tool default: tracing, prompt management, LLM-as-judge evals, and datasets in a single open-source platform you can self-host in minutes or run on a $29 cloud plan.

    The open-source category leader — 32.6k GitHub stars, 300+ contributors, MIT-licensed core, batteries included (traces, prompt versioning, evals, playground, datasets). The ClickHouse acquisition (Jan 2026) removed the two classic OSS risks at once: funding runway and query-performance ceilings — Langfuse v4 shipped 'real-time, up to 165× faster' on the new backend. The $29 Core tier is the cheapest credible paid entry among the leaders.

    Watch Eval UX and experiment workflows trail Braintrust's — Langfuse is observability-first, evals-second. Post-acquisition roadmap now serves ClickHouse's platform ambitions too; the /ee folders are not MIT, so 'fully open source' has an asterisk.

    32.6k GitHub stars, 300+ contributors; MIT core (Aug 2026) [src] · Joined ClickHouse Jan 2026; Langfuse v4 'up to 165× faster' [src] · Core $29/mo with 100k units incl.; Hobby free 50k units (verified Aug 2026) [src]
  2. 2BraintrustBraintrust Data

    Starter free (1GB data, 10k scores) · Pro $249/mo · Enterprise custom (on-prem or hosted); 6–12 mo free for startups

    Best for Teams that treat evals as the product-development loop itself — experiments, datasets, human review, and a purpose-built trace store (Brainstore) with the best eval workflow in the category.

    Eval-first where everyone else is tracing-first, and the logo wall proves it: Notion, Replit, Cloudflare, Ramp, and Dropbox are named customers of record. $80M Series B led by ICONIQ (Feb 17, 2026) on top of a $36M a16z Series A. Shipping velocity is the tell — Loop (agent that turns production data into evals, Nov 2025), Topics auto-discovery GA (Jun 2026), CLI + MCP support (Apr 2026).

    Watch Closed source with a proprietary data store — the deepest lock-in of the top five. $249/mo Pro plus usage ($3/GB, $1.50/1k scores) makes it the priciest non-enterprise entry; the free tier's 14-day retention is the shortest here.

    $80M Series B led by ICONIQ; customers Notion, Replit, Cloudflare, Ramp, Dropbox (Feb 17, 2026) [src] · Loop launched Nov 24, 2025; Topics GA Jun 1, 2026 [src] · Pro $249/mo incl. $249 model credits, 5GB data, 50k scores (verified Aug 2026) [src]
  3. 3LangSmithLangChain

    Developer free (5k traces/mo) · Plus $39/seat/mo (10k traces) · then pay-as-you-go · Enterprise custom (self-host/hybrid)

    Best for Anyone building on LangChain or LangGraph — zero-config tracing of the framework 35% of the Fortune 500 touches, plus deployment and agent infra in the same subscription.

    The gravity play: LangChain + LangGraph pull 90M monthly downloads, and LangSmith commercial trace volume grew 12x year-over-year (Oct 2025) — the framework funnel works. The $125M Series B at $1.25B (IVP, Oct 20, 2025) bought a genuine platform: observability, evals, prompt engineering, and agent deployment under one roof, with self-hosted enterprise as a real option.

    Watch Closed source, and outside the LangChain ecosystem it's just another good tracer — the instrumentation advantage evaporates. Per-seat + LCU/LSU metered pricing is the hardest here to forecast. Framework coupling cuts both ways if you migrate off LangGraph.

    $125M at $1.25B led by IVP; LangSmith trace volume 12x YoY; 35% of Fortune 500 use LangChain products (Oct 20, 2025) [src] · 90M combined monthly downloads for LangChain + LangGraph (Oct 2025) [src] · Plus $39/seat/mo, 10k base traces then usage (verified Aug 2026) [src]
  4. 4Arize PhoenixArize AI

    OSS free to self-host (ELv2) · free Phoenix Cloud instances · Arize AX (commercial platform) custom

    Best for OTel purists and self-hosters — vendor-neutral OpenTelemetry tracing plus evals that stay entirely in your infrastructure, with a funded enterprise platform (AX) behind it when you outgrow free.

    The standards bet: Phoenix's OpenInference instrumentation rides OpenTelemetry, so traces are portable by construction — 3M+ monthly Phoenix downloads and 22M+ monthly OTel instrumentation downloads (Aug 2026). Backing is the deepest in pure observability: Arize's $70M Series C (Feb 2025) was billed as the largest-ever investment in AI observability, with Datadog and M12 on the cap table.

    Watch ELv2, not MIT/Apache — fine for self-hosting, but not fully open. Prompt management and eval workflow are thinner than Langfuse/Braintrust; the upsell path to Arize AX is where the polish (and the price) lives.

    3M+ monthly downloads, 10k+ GitHub stars, 22M+ monthly OTel instrumentation downloads (Aug 2026) [src] · Arize $70M Series C, 'largest-ever investment in AI observability' (Feb 20, 2025) [src]
  5. 5W&B WeaveWeights & Biases (a CoreWeave company)

    Free tier (1GB/mo Weave ingestion) · Pro from $60/mo (1.5GB/mo incl., then usage) · Enterprise custom incl. self-managed

    Best for Teams already in the W&B ecosystem, or running multi-agent systems that want sessions/turns/sub-agents as first-class trace concepts rather than bolted-on spans.

    The most agent-native data model of the leaders — sessions, turns, steps, tools, and sub-agents are primitives, with pre-built safety scorers (toxicity, PII, hallucination) and an API that lets coding agents like Claude Code read production data and run eval loops autonomously. CoreWeave's acquisition of W&B (closed 2025) gives it GPU-cloud distribution and staying power the point solutions lack.

    Watch It's a module inside a bigger ML platform — if you don't want experiment tracking and model registry, you're navigating around them. Post-acquisition, W&B's roadmap now serves CoreWeave's cloud strategy; ingestion overage pricing punishes verbose traces.

    Free 1GB/mo Weave ingestion; Pro from $60/mo (verified Aug 2026) [src] · Agent-native tracing: sessions/turns/sub-agents first-class; autonomous eval loops via coding agents (Aug 2026) [src] · CoreWeave completed W&B acquisition 2025 (reported ~$1.7B) [src]

Evals & Observability: the top 8 compared

Edition v2026.Q3 · ranking, pricing and status verified 2026-08-06.

Evals & Observability — the top 8 compared. Edition v2026.Q3, verified 2026-08-06.
ToolEntry priceOpen sourceSelf-hostOTel-nativeEval workflowPrompt mgmtSweet spot
LangfuseFree · $29/moYes (MIT core)Yes, freeYesGood (judge + code + human)Yesone-tool default
BraintrustFree · $249/moNoEnterprise onlyPartialBest-in-classYeseval-driven product loops
LangSmithFree · $39/seatNoEnterprise onlyPartial (OTel export)GoodYesLangChain/LangGraph shops
Arize PhoenixFreeYes (ELv2)Yes, freeYes (OpenInference)GoodBasicOTel-standard self-hosting
W&B WeaveFree · $60/moNo (SDK only)Enterprise onlyYesGood + safety scorersBasicmulti-agent tracing
Opik (Comet)FreeYes (Apache-2.0)Yes, freeYesGoodYesunrestricted OSS self-host
DeepEval / Confident AIFree · $200/moFramework (Apache-2.0)EnterprisePartialStrong (pytest-style)Yesevals-as-unit-tests in CI
HeliconeFree · $79/moYesYesVia gatewayBasicBasiccost/usage logging via proxy

How to choose your evals & observability

If you want one tool, no lock-in, and a price that doesn't scale with panic
Langfuse — MIT-core self-host or $29/mo cloud, now with ClickHouse's balance sheet behind it (Jan 2026).
If evals drive your product decisions and you have real budget
Braintrust — the eval UX Notion, Replit, and Cloudflare standardized on; take the 6–12 months free startup credit.
If your stack is LangChain/LangGraph
LangSmith — zero-config tracing plus deployment in one bill; its 12x YoY trace growth is the ecosystem talking.
If you require traces to never leave your infra, at zero license cost
Arize Phoenix (ELv2, OTel-native) or Opik (Apache-2.0, fully self-hostable) — the two credible free self-host paths.
If you're shipping anything customer-facing without evals in CI
Stop — wire pytest-style checks (DeepEval, promptfoo) into CI this week. Gartner's 40%-canceled-by-2027 cohort is made of teams that couldn't prove ROI.

Evals & Observability: the whole field

21 more tools tracked in this category, including 6 dead, renamed, or sunsetting — a reference that hides the graveyard isn't one. Verified 2026-08-06.

Evals & Observability — every tool we track, including 6 dead, renamed, or sunsetting. Edition v2026.Q3, verified 2026-08-06.
ToolMakerWhat it isEntryStatus
OpikCometThe other big OSS platform — 20.7k stars, Apache-2.0, fully self-hostable including backend; the near-miss for the top 5freeactive
DeepEval / Confident AIConfident AIPytest-for-LLMs framework (15.4k stars, Apache-2.0) + commercial platform from $200/mo; the CI-evals standardfree · $200/moactive
promptfooPromptfoo Inc.MIT eval + red-teaming CLI, 21.3k stars, local-first; its security half overlaps element Gd · Guardrailsfreeactive
HeliconeHeliconeOne-line proxy for LLM logging and cost analytics (5.8k stars); observability via gateway, thinner on evalsfree · $79/moactive
OpenAI EvalsOpenAIThe 2023 pioneer repo (18.6k stars, MIT) is quiet; the live product is the Evals API/dashboard inside the OpenAI platform — OpenAI-centric by designfree + tokensfading
Arize AXArize AICommercial sibling of Phoenix — enterprise agent engineering platform; $70M Series C (Feb 2025)customactive
GalileoGalileo (Rungalileo)Enterprise 'eval engineering' — Luna distilled evaluator models claim 96% cheaper production scoring; NVIDIA, HP, MongoDB testimonialsfree tier · customactive
Datadog LLM ObservabilityDatadogThe incumbent play: agent tracing + experiments + evaluators inside the APM you already pay for; free to 40k spans, Pro from $160/mofree · $160/moactive
New Relic AI MonitoringNew RelicAPM-bundled LLM observability; fine if you're already a customer, nobody's first choice for evalsusage-basedactive
Patronus AIPatronus AIEvaluation API and research-grade judges (Lynx hallucination model, Percival agent debugger)free tieractive
RagasRagas (ex-Exploding Gradients)The default OSS metric library for RAG evals (faithfulness, context precision); framework, not platformfreeactive
Maxim AIMaximAgent simulation + evals + observability platform; aggressive on agent-testing use casesfree tieractive
LangWatchLangWatchEU-based open-core LLM ops with optimization studio (DSPy-powered)free tieractive
LunaryLunaryLightweight open-source observability + prompt management; small but steadyfree tieractive
PromptLayerPromptLayerPrompt-management-first with evals attached; popular with non-engineer prompt ownersfree tieractive
Traceloop (OpenLLMetry)TraceloopMaintainer of OpenLLMetry, the OTel LLM instrumentation many platforms ingest; thin commercial layerfree tieractive
HumanloopHumanloop → AnthropicFirst-mover LLM eval/prompt platform; team acqui-hired by Anthropic, platform sunset with migration guides (2025)dead
TruEra / TruLensTruEra → SnowflakeML-observability pioneer acquired by Snowflake (May 2024); TruLens OSS evals live on under Snowflakefree (OSS)acquired
AporiaAporia → CoralogixML/LLM observability + guardrails, acquired by Coralogix (Dec 2024); now Coralogix's AI Centeracquired
Weights & Biases (company)W&B → CoreWeaveThe whole company was acquired by CoreWeave (2025, reported ~$1.7B) — Weave continues as its LLM-ops armacquired
Langfuse (company)Langfuse → ClickHouseAcquired by ClickHouse (Jan 2026); product continues under its own brand — see top 5acquired

Evals & Observability: the category in numbers

Edition v2026.Q3 · ranking, pricing and status verified 2026-08-06.

  • Gartner: 40%+ of agentic AI projects will be canceled by end of 2027 — escalating costs, unclear business value, inadequate risk controls (Jun 25, 2025). This is the demand driver for the whole category [src]
  • Braintrust $80M Series B led by ICONIQ (Feb 17, 2026); ~$800M valuation per press reporting; named customers Notion, Replit, Cloudflare, Ramp, Dropbox [src]
  • LangChain $125M at $1.25B (IVP, Oct 20, 2025); LangSmith commercial trace volume 12x YoY — evals/observability is the monetization layer of the framework business [src]
  • Consolidation wave: Langfuse → ClickHouse (Jan 2026), W&B → CoreWeave (2025), Humanloop team → Anthropic with platform sunset (2025), Aporia → Coralogix (Dec 2024), TruEra → Snowflake (May 2024) [src]
  • Arize $70M Series C (Feb 20, 2025) billed as the largest-ever investment in AI observability; Datadog and Microsoft's M12 on the cap table — incumbents buying visibility into the category [src]
  • APM incumbents now bundle the job: Datadog LLM Observability ships datasets/experiments/evaluators from $160/mo, free to 40k spans (Aug 2026) — squeezing point solutions from above [src]

Evals & Observability: method & sources

Ranking criteria: evidence of production adoption (named customers, download/star trajectories), pricing accessibility for a small team, lock-in risk (license + data portability), and eval-workflow depth — not just tracing. Conflicts resolved: Braintrust's $800M valuation comes from press reporting around the round, not the company's own Feb 17, 2026 post (which discloses no valuation) — treat as reported, not confirmed. The W&B/CoreWeave ~$1.7B price is likewise reported, never officially confirmed. OpenAI's evals repo showed commit activity on fetch but community consensus is it's effectively frozen as a benchmark registry; the maintained product is the platform Evals API — we mark the repo 'fading', which a human should sanity-check. Langfuse GitHub shows 32.4k stars vs 32.6k on its pricing page; we cite the vendor page number, the delta is cache lag. Adjacent elements: LLM gateways with logging (Portkey, OpenRouter analytics, LiteLLM) belong to Rt · Model Router even though they sell 'observability'; promptfoo's red-teaming half and Galileo's runtime guardrails overlap Gd · Guardrails; token-cost dashboards alone are Sp · Spend Management. Uptime/APM monitoring of the app around the model stays with classic observability, out of scope here. Ranking criteria: verified commercial traction, independent satisfaction surveys, agent benchmarks, and founder-fit (price floor, lock-in, surfaces). Editorial, never paid — the charter. Machine-readable twin: ev.json.

All sources (22)
  1. https://langfuse.com/pricing
  2. https://github.com/langfuse/langfuse
  3. https://langfuse.com/blog
  4. https://www.braintrust.dev/pricing
  5. https://www.braintrust.dev/blog/announcing-series-b
  6. https://www.braintrust.dev/blog
  7. https://www.langchain.com/pricing
  8. https://www.langchain.com/blog/series-b
  9. https://arize.com/phoenix
  10. https://arize.com/blog/arize-ai-raises-70m-series-c/
  11. https://wandb.ai/site/weave
  12. https://wandb.ai/site/pricing/
  13. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
  14. https://github.com/comet-ml/opik
  15. https://github.com/promptfoo/promptfoo
  16. https://github.com/confident-ai/deepeval
  17. https://github.com/openai/evals
  18. https://www.helicone.ai/pricing
  19. https://www.confident-ai.com/pricing
  20. https://www.datadoghq.com/product/llm-observability/
  21. https://galileo.ai
  22. https://humanloop.com

Our take

Forty percent of agent projects die from unproven ROI. Evals are how the other sixty percent prove it. Non-negotiable for anything customer-facing.

Combines with

Appears in compounds

The Agent Stack · The Trust Stack

Guides

Best LLM evals & observability tools (2026)

This is element 47 of 58. The table is versioned quarterly — when a tool loses its seat, the changelog records the succession.

Explore the full table →