# Ev · Evals & Observability — element 47 of 58

> Prove the machine works. Turns 'it seems fine' into measured quality.

- **Group:** 11 · Trust & Compliance
- **Necessity:** Essential
- **Price band:** $ · under $30/mo
- **Maturity:** Emerging
- **Edition:** v2026.Q3 · verified 2026-09-13

## Leading tools (v2026.Q3)

- **Langfuse** — the open-source default
- **Braintrust** — eval-first, big-logo favorite
- **LangSmith** — the langchain-native path
- **Arize Phoenix** — otel-native, self-host free
- **W&B Weave** — agent-native, gpu-cloud backed

## Our take

Forty percent of agent projects die from unproven ROI. Evals are how the other sixty percent prove it. Non-negotiable for anything customer-facing.

## Combines with

Ca, Ag, Gd
- Guide: [Best LLM evals & observability tools (2026)](https://elems.ai/best/best-llm-evals-observability-tools.html)

## The top 5 — deep dossier (verified 2026-09-13)

Langfuse, for the default seat — the most-starred open-source LLM engineering platform (32.6k stars, MIT core), a $29/mo cloud entry, and since January 2026 the resources of ClickHouse behind it without the pricing games of a VC burn clock. Braintrust is the pick when evals ARE your product loop and you'll pay for the best eval UX — the platform Notion, Replit, Cloudflare, Ramp, and Dropbox standardized on, freshly armed with an $80M Series B. LangSmith wins if you build on LangChain/LangGraph; Arize Phoenix if you want OpenTelemetry-standard tracing self-hosted for free; W&B Weave if you're already in the W&B/CoreWeave orbit and want agent-native sessions-and-turns tracing. Whatever you pick, pick one: Gartner projects 40%+ of agentic AI projects canceled by 2027 for unproven ROI — evals are how you avoid being in that cohort.

1. **Langfuse** (Langfuse (by ClickHouse since Jan 2026)) — Hobby free (50k units/mo) · Core $29/mo · Pro $199/mo · Enterprise $2,499/mo · self-host free (MIT core). Best for: The one-tool default: tracing, prompt management, LLM-as-judge evals, and datasets in a single open-source platform you can self-host in minutes or run on a $29 cloud plan. Why: The open-source category leader — 32.6k GitHub stars, 300+ contributors, MIT-licensed core, batteries included (traces, prompt versioning, evals, playground, datasets). The ClickHouse acquisition (Jan 2026) removed the two classic OSS risks at once: funding runway and query-performance ceilings — Langfuse v4 shipped 'real-time, up to 165× faster' on the new backend. The $29 Core tier is the cheapest credible paid entry among the leaders. Watch: Eval UX and experiment workflows trail Braintrust's — Langfuse is observability-first, evals-second. Post-acquisition roadmap now serves ClickHouse's platform ambitions too; the /ee folders are not MIT, so 'fully open source' has an asterisk. [https://langfuse.com](https://langfuse.com)
2. **Braintrust** (Braintrust Data) — Starter free (1GB data, 10k scores) · Pro $249/mo · Enterprise custom (on-prem or hosted); 6–12 mo free for startups. Best for: Teams that treat evals as the product-development loop itself — experiments, datasets, human review, and a purpose-built trace store (Brainstore) with the best eval workflow in the category. Why: Eval-first where everyone else is tracing-first, and the logo wall proves it: Notion, Replit, Cloudflare, Ramp, and Dropbox are named customers of record. $80M Series B led by ICONIQ (Feb 17, 2026) on top of a $36M a16z Series A. Shipping velocity is the tell — Loop (agent that turns production data into evals, Nov 2025), Topics auto-discovery GA (Jun 2026), CLI + MCP support (Apr 2026). Watch: Closed source with a proprietary data store — the deepest lock-in of the top five. $249/mo Pro plus usage ($3/GB, $1.50/1k scores) makes it the priciest non-enterprise entry; the free tier's 14-day retention is the shortest here. [https://www.braintrust.dev](https://www.braintrust.dev)
3. **LangSmith** (LangChain) — Developer free (5k traces/mo) · Plus $39/seat/mo (10k traces) · then pay-as-you-go · Enterprise custom (self-host/hybrid). Best for: Anyone building on LangChain or LangGraph — zero-config tracing of the framework 35% of the Fortune 500 touches, plus deployment and agent infra in the same subscription. Why: The gravity play: LangChain + LangGraph pull 90M monthly downloads, and LangSmith commercial trace volume grew 12x year-over-year (Oct 2025) — the framework funnel works. The $125M Series B at $1.25B (IVP, Oct 20, 2025) bought a genuine platform: observability, evals, prompt engineering, and agent deployment under one roof, with self-hosted enterprise as a real option. Watch: Closed source, and outside the LangChain ecosystem it's just another good tracer — the instrumentation advantage evaporates. Per-seat + LCU/LSU metered pricing is the hardest here to forecast. Framework coupling cuts both ways if you migrate off LangGraph. [https://www.langchain.com/langsmith](https://www.langchain.com/langsmith)
4. **Arize Phoenix** (Arize AI) — OSS free to self-host (ELv2) · free Phoenix Cloud instances · Arize AX (commercial platform) custom. Best for: OTel purists and self-hosters — vendor-neutral OpenTelemetry tracing plus evals that stay entirely in your infrastructure, with a funded enterprise platform (AX) behind it when you outgrow free. Why: The standards bet: Phoenix's OpenInference instrumentation rides OpenTelemetry, so traces are portable by construction — 3M+ monthly Phoenix downloads and 22M+ monthly OTel instrumentation downloads (Aug 2026). Backing is the deepest in pure observability: Arize's $70M Series C (Feb 2025) was billed as the largest-ever investment in AI observability, with Datadog and M12 on the cap table. Watch: ELv2, not MIT/Apache — fine for self-hosting, but not fully open. Prompt management and eval workflow are thinner than Langfuse/Braintrust; the upsell path to Arize AX is where the polish (and the price) lives. [https://arize.com/phoenix](https://arize.com/phoenix)
5. **W&B Weave** (Weights & Biases (a CoreWeave company)) — Free tier (1GB/mo Weave ingestion) · Pro from $60/mo (1.5GB/mo incl., then usage) · Enterprise custom incl. self-managed. Best for: Teams already in the W&B ecosystem, or running multi-agent systems that want sessions/turns/sub-agents as first-class trace concepts rather than bolted-on spans. Why: The most agent-native data model of the leaders — sessions, turns, steps, tools, and sub-agents are primitives, with pre-built safety scorers (toxicity, PII, hallucination) and an API that lets coding agents like Claude Code read production data and run eval loops autonomously. CoreWeave's acquisition of W&B (closed 2025) gives it GPU-cloud distribution and staying power the point solutions lack. Watch: It's a module inside a bigger ML platform — if you don't want experiment tracking and model registry, you're navigating around them. Post-acquisition, W&B's roadmap now serves CoreWeave's cloud strategy; ingestion overage pricing punishes verbose traces. [https://wandb.ai/site/weave](https://wandb.ai/site/weave)

### How to choose
- If You want one tool, no lock-in, and a price that doesn't scale with panic → Langfuse — MIT-core self-host or $29/mo cloud, now with ClickHouse's balance sheet behind it (Jan 2026).
- If Evals drive your product decisions and you have real budget → Braintrust — the eval UX Notion, Replit, and Cloudflare standardized on; take the 6–12 months free startup credit.
- If Your stack is LangChain/LangGraph → LangSmith — zero-config tracing plus deployment in one bill; its 12x YoY trace growth is the ecosystem talking.
- If You require traces to never leave your infra, at zero license cost → Arize Phoenix (ELv2, OTel-native) or Opik (Apache-2.0, fully self-hostable) — the two credible free self-host paths.
- If You're shipping anything customer-facing without evals in CI → Stop — wire pytest-style checks (DeepEval, promptfoo) into CI this week. Gartner's 40%-canceled-by-2027 cohort is made of teams that couldn't prove ROI.

### The field (21 more)

Opik, DeepEval / Confident AI, promptfoo, Helicone, OpenAI Evals (fading), Arize AX, Galileo, Datadog LLM Observability, New Relic AI Monitoring, Patronus AI, Ragas, Maxim AI, LangWatch, Lunary, PromptLayer, Traceloop (OpenLLMetry), Humanloop (dead), TruEra / TruLens (acquired), Aporia (acquired), Weights & Biases (company) (acquired), Langfuse (company) (acquired)

Full dossier data: https://elems.ai/e/ev.json

---
Source: [elems.ai](https://elems.ai/e/ev.html) — the periodic table of the AI-led startup. Data: https://elems.ai/elements.json (CC BY 4.0, cite elems.ai).
