{
 "sym": "Ev",
 "updated": "2026-08-06",
 "verdict": "Langfuse, for the default seat \u2014 the most-starred open-source LLM engineering platform (32.6k stars, MIT core), a $29/mo cloud entry, and since January 2026 the resources of ClickHouse behind it without the pricing games of a VC burn clock. Braintrust is the pick when evals ARE your product loop and you'll pay for the best eval UX \u2014 the platform Notion, Replit, Cloudflare, Ramp, and Dropbox standardized on, freshly armed with an $80M Series B. LangSmith wins if you build on LangChain/LangGraph; Arize Phoenix if you want OpenTelemetry-standard tracing self-hosted for free; W&B Weave if you're already in the W&B/CoreWeave orbit and want agent-native sessions-and-turns tracing. Whatever you pick, pick one: Gartner projects 40%+ of agentic AI projects canceled by 2027 for unproven ROI \u2014 evals are how you avoid being in that cohort.",
 "top5": [
  {
   "rank": 1,
   "name": "Langfuse",
   "maker": "Langfuse (by ClickHouse since Jan 2026)",
   "url": "https://langfuse.com",
   "docs": "https://langfuse.com/docs",
   "pricing": "Hobby free (50k units/mo) \u00b7 Core $29/mo \u00b7 Pro $199/mo \u00b7 Enterprise $2,499/mo \u00b7 self-host free (MIT core)",
   "best_for": "The one-tool default: tracing, prompt management, LLM-as-judge evals, and datasets in a single open-source platform you can self-host in minutes or run on a $29 cloud plan.",
   "why": "The open-source category leader \u2014 32.6k GitHub stars, 300+ contributors, MIT-licensed core, batteries included (traces, prompt versioning, evals, playground, datasets). The ClickHouse acquisition (Jan 2026) removed the two classic OSS risks at once: funding runway and query-performance ceilings \u2014 Langfuse v4 shipped 'real-time, up to 165\u00d7 faster' on the new backend. The $29 Core tier is the cheapest credible paid entry among the leaders.",
   "watch": "Eval UX and experiment workflows trail Braintrust's \u2014 Langfuse is observability-first, evals-second. Post-acquisition roadmap now serves ClickHouse's platform ambitions too; the /ee folders are not MIT, so 'fully open source' has an asterisk.",
   "evidence": [
    {
     "stat": "32.6k GitHub stars, 300+ contributors; MIT core (Aug 2026)",
     "src": "https://github.com/langfuse/langfuse"
    },
    {
     "stat": "Joined ClickHouse Jan 2026; Langfuse v4 'up to 165\u00d7 faster'",
     "src": "https://langfuse.com/blog"
    },
    {
     "stat": "Core $29/mo with 100k units incl.; Hobby free 50k units (verified Aug 2026)",
     "src": "https://langfuse.com/pricing"
    }
   ],
   "tile_note": "the open-source default"
  },
  {
   "rank": 2,
   "name": "Braintrust",
   "maker": "Braintrust Data",
   "url": "https://www.braintrust.dev",
   "docs": "https://www.braintrust.dev/docs",
   "pricing": "Starter free (1GB data, 10k scores) \u00b7 Pro $249/mo \u00b7 Enterprise custom (on-prem or hosted); 6\u201312 mo free for startups",
   "best_for": "Teams that treat evals as the product-development loop itself \u2014 experiments, datasets, human review, and a purpose-built trace store (Brainstore) with the best eval workflow in the category.",
   "why": "Eval-first where everyone else is tracing-first, and the logo wall proves it: Notion, Replit, Cloudflare, Ramp, and Dropbox are named customers of record. $80M Series B led by ICONIQ (Feb 17, 2026) on top of a $36M a16z Series A. Shipping velocity is the tell \u2014 Loop (agent that turns production data into evals, Nov 2025), Topics auto-discovery GA (Jun 2026), CLI + MCP support (Apr 2026).",
   "watch": "Closed source with a proprietary data store \u2014 the deepest lock-in of the top five. $249/mo Pro plus usage ($3/GB, $1.50/1k scores) makes it the priciest non-enterprise entry; the free tier's 14-day retention is the shortest here.",
   "evidence": [
    {
     "stat": "$80M Series B led by ICONIQ; customers Notion, Replit, Cloudflare, Ramp, Dropbox (Feb 17, 2026)",
     "src": "https://www.braintrust.dev/blog/announcing-series-b"
    },
    {
     "stat": "Loop launched Nov 24, 2025; Topics GA Jun 1, 2026",
     "src": "https://www.braintrust.dev/blog"
    },
    {
     "stat": "Pro $249/mo incl. $249 model credits, 5GB data, 50k scores (verified Aug 2026)",
     "src": "https://www.braintrust.dev/pricing"
    }
   ],
   "tile_note": "eval-first, big-logo favorite"
  },
  {
   "rank": 3,
   "name": "LangSmith",
   "maker": "LangChain",
   "url": "https://www.langchain.com/langsmith",
   "docs": "https://docs.langchain.com/langsmith",
   "pricing": "Developer free (5k traces/mo) \u00b7 Plus $39/seat/mo (10k traces) \u00b7 then pay-as-you-go \u00b7 Enterprise custom (self-host/hybrid)",
   "best_for": "Anyone building on LangChain or LangGraph \u2014 zero-config tracing of the framework 35% of the Fortune 500 touches, plus deployment and agent infra in the same subscription.",
   "why": "The gravity play: LangChain + LangGraph pull 90M monthly downloads, and LangSmith commercial trace volume grew 12x year-over-year (Oct 2025) \u2014 the framework funnel works. The $125M Series B at $1.25B (IVP, Oct 20, 2025) bought a genuine platform: observability, evals, prompt engineering, and agent deployment under one roof, with self-hosted enterprise as a real option.",
   "watch": "Closed source, and outside the LangChain ecosystem it's just another good tracer \u2014 the instrumentation advantage evaporates. Per-seat + LCU/LSU metered pricing is the hardest here to forecast. Framework coupling cuts both ways if you migrate off LangGraph.",
   "evidence": [
    {
     "stat": "$125M at $1.25B led by IVP; LangSmith trace volume 12x YoY; 35% of Fortune 500 use LangChain products (Oct 20, 2025)",
     "src": "https://www.langchain.com/blog/series-b"
    },
    {
     "stat": "90M combined monthly downloads for LangChain + LangGraph (Oct 2025)",
     "src": "https://www.langchain.com/blog/series-b"
    },
    {
     "stat": "Plus $39/seat/mo, 10k base traces then usage (verified Aug 2026)",
     "src": "https://www.langchain.com/pricing"
    }
   ],
   "tile_note": "the langchain-native path"
  },
  {
   "rank": 4,
   "name": "Arize Phoenix",
   "maker": "Arize AI",
   "url": "https://arize.com/phoenix",
   "docs": "https://arize.com/docs/phoenix",
   "pricing": "OSS free to self-host (ELv2) \u00b7 free Phoenix Cloud instances \u00b7 Arize AX (commercial platform) custom",
   "best_for": "OTel purists and self-hosters \u2014 vendor-neutral OpenTelemetry tracing plus evals that stay entirely in your infrastructure, with a funded enterprise platform (AX) behind it when you outgrow free.",
   "why": "The standards bet: Phoenix's OpenInference instrumentation rides OpenTelemetry, so traces are portable by construction \u2014 3M+ monthly Phoenix downloads and 22M+ monthly OTel instrumentation downloads (Aug 2026). Backing is the deepest in pure observability: Arize's $70M Series C (Feb 2025) was billed as the largest-ever investment in AI observability, with Datadog and M12 on the cap table.",
   "watch": "ELv2, not MIT/Apache \u2014 fine for self-hosting, but not fully open. Prompt management and eval workflow are thinner than Langfuse/Braintrust; the upsell path to Arize AX is where the polish (and the price) lives.",
   "evidence": [
    {
     "stat": "3M+ monthly downloads, 10k+ GitHub stars, 22M+ monthly OTel instrumentation downloads (Aug 2026)",
     "src": "https://arize.com/phoenix"
    },
    {
     "stat": "Arize $70M Series C, 'largest-ever investment in AI observability' (Feb 20, 2025)",
     "src": "https://arize.com/blog/arize-ai-raises-70m-series-c/"
    }
   ],
   "tile_note": "otel-native, self-host free"
  },
  {
   "rank": 5,
   "name": "W&B Weave",
   "maker": "Weights & Biases (a CoreWeave company)",
   "url": "https://wandb.ai/site/weave",
   "docs": "https://weave-docs.wandb.ai",
   "pricing": "Free tier (1GB/mo Weave ingestion) \u00b7 Pro from $60/mo (1.5GB/mo incl., then usage) \u00b7 Enterprise custom incl. self-managed",
   "best_for": "Teams already in the W&B ecosystem, or running multi-agent systems that want sessions/turns/sub-agents as first-class trace concepts rather than bolted-on spans.",
   "why": "The most agent-native data model of the leaders \u2014 sessions, turns, steps, tools, and sub-agents are primitives, with pre-built safety scorers (toxicity, PII, hallucination) and an API that lets coding agents like Claude Code read production data and run eval loops autonomously. CoreWeave's acquisition of W&B (closed 2025) gives it GPU-cloud distribution and staying power the point solutions lack.",
   "watch": "It's a module inside a bigger ML platform \u2014 if you don't want experiment tracking and model registry, you're navigating around them. Post-acquisition, W&B's roadmap now serves CoreWeave's cloud strategy; ingestion overage pricing punishes verbose traces.",
   "evidence": [
    {
     "stat": "Free 1GB/mo Weave ingestion; Pro from $60/mo (verified Aug 2026)",
     "src": "https://wandb.ai/site/pricing/"
    },
    {
     "stat": "Agent-native tracing: sessions/turns/sub-agents first-class; autonomous eval loops via coding agents (Aug 2026)",
     "src": "https://wandb.ai/site/weave"
    },
    {
     "stat": "CoreWeave completed W&B acquisition 2025 (reported ~$1.7B)",
     "src": "https://www.coreweave.com/news/coreweave-completes-acquisition-of-weights-biases"
    }
   ],
   "tile_note": "agent-native, gpu-cloud backed"
  }
 ],
 "matrix": {
  "cols": [
   "Entry price",
   "Open source",
   "Self-host",
   "OTel-native",
   "Eval workflow",
   "Prompt mgmt",
   "Sweet spot"
  ],
  "rows": [
   [
    "Langfuse",
    "Free \u00b7 $29/mo",
    "Yes (MIT core)",
    "Yes, free",
    "Yes",
    "Good (judge + code + human)",
    "Yes",
    "one-tool default"
   ],
   [
    "Braintrust",
    "Free \u00b7 $249/mo",
    "No",
    "Enterprise only",
    "Partial",
    "Best-in-class",
    "Yes",
    "eval-driven product loops"
   ],
   [
    "LangSmith",
    "Free \u00b7 $39/seat",
    "No",
    "Enterprise only",
    "Partial (OTel export)",
    "Good",
    "Yes",
    "LangChain/LangGraph shops"
   ],
   [
    "Arize Phoenix",
    "Free",
    "Yes (ELv2)",
    "Yes, free",
    "Yes (OpenInference)",
    "Good",
    "Basic",
    "OTel-standard self-hosting"
   ],
   [
    "W&B Weave",
    "Free \u00b7 $60/mo",
    "No (SDK only)",
    "Enterprise only",
    "Yes",
    "Good + safety scorers",
    "Basic",
    "multi-agent tracing"
   ],
   [
    "Opik (Comet)",
    "Free",
    "Yes (Apache-2.0)",
    "Yes, free",
    "Yes",
    "Good",
    "Yes",
    "unrestricted OSS self-host"
   ],
   [
    "DeepEval / Confident AI",
    "Free \u00b7 $200/mo",
    "Framework (Apache-2.0)",
    "Enterprise",
    "Partial",
    "Strong (pytest-style)",
    "Yes",
    "evals-as-unit-tests in CI"
   ],
   [
    "Helicone",
    "Free \u00b7 $79/mo",
    "Yes",
    "Yes",
    "Via gateway",
    "Basic",
    "Basic",
    "cost/usage logging via proxy"
   ]
  ]
 },
 "rules": [
  {
   "if": "You want one tool, no lock-in, and a price that doesn't scale with panic",
   "then": "Langfuse \u2014 MIT-core self-host or $29/mo cloud, now with ClickHouse's balance sheet behind it (Jan 2026)."
  },
  {
   "if": "Evals drive your product decisions and you have real budget",
   "then": "Braintrust \u2014 the eval UX Notion, Replit, and Cloudflare standardized on; take the 6\u201312 months free startup credit."
  },
  {
   "if": "Your stack is LangChain/LangGraph",
   "then": "LangSmith \u2014 zero-config tracing plus deployment in one bill; its 12x YoY trace growth is the ecosystem talking."
  },
  {
   "if": "You require traces to never leave your infra, at zero license cost",
   "then": "Arize Phoenix (ELv2, OTel-native) or Opik (Apache-2.0, fully self-hostable) \u2014 the two credible free self-host paths."
  },
  {
   "if": "You're shipping anything customer-facing without evals in CI",
   "then": "Stop \u2014 wire pytest-style checks (DeepEval, promptfoo) into CI this week. Gartner's 40%-canceled-by-2027 cohort is made of teams that couldn't prove ROI."
  }
 ],
 "field": [
  {
   "name": "Opik",
   "maker": "Comet",
   "note": "The other big OSS platform \u2014 20.7k stars, Apache-2.0, fully self-hostable including backend; the near-miss for the top 5",
   "url": "https://github.com/comet-ml/opik",
   "oss": true,
   "entry": "free",
   "status": "active"
  },
  {
   "name": "DeepEval / Confident AI",
   "maker": "Confident AI",
   "note": "Pytest-for-LLMs framework (15.4k stars, Apache-2.0) + commercial platform from $200/mo; the CI-evals standard",
   "url": "https://github.com/confident-ai/deepeval",
   "oss": true,
   "entry": "free \u00b7 $200/mo",
   "status": "active"
  },
  {
   "name": "promptfoo",
   "maker": "Promptfoo Inc.",
   "note": "MIT eval + red-teaming CLI, 21.3k stars, local-first; its security half overlaps element Gd \u00b7 Guardrails",
   "url": "https://github.com/promptfoo/promptfoo",
   "oss": true,
   "entry": "free",
   "status": "active"
  },
  {
   "name": "Helicone",
   "maker": "Helicone",
   "note": "One-line proxy for LLM logging and cost analytics (5.8k stars); observability via gateway, thinner on evals",
   "url": "https://www.helicone.ai",
   "oss": true,
   "entry": "free \u00b7 $79/mo",
   "status": "active"
  },
  {
   "name": "OpenAI Evals",
   "maker": "OpenAI",
   "note": "The 2023 pioneer repo (18.6k stars, MIT) is quiet; the live product is the Evals API/dashboard inside the OpenAI platform \u2014 OpenAI-centric by design",
   "url": "https://github.com/openai/evals",
   "oss": true,
   "entry": "free + tokens",
   "status": "fading"
  },
  {
   "name": "Arize AX",
   "maker": "Arize AI",
   "note": "Commercial sibling of Phoenix \u2014 enterprise agent engineering platform; $70M Series C (Feb 2025)",
   "url": "https://arize.com",
   "oss": false,
   "entry": "custom",
   "status": "active"
  },
  {
   "name": "Galileo",
   "maker": "Galileo (Rungalileo)",
   "note": "Enterprise 'eval engineering' \u2014 Luna distilled evaluator models claim 96% cheaper production scoring; NVIDIA, HP, MongoDB testimonials",
   "url": "https://galileo.ai",
   "oss": false,
   "entry": "free tier \u00b7 custom",
   "status": "active"
  },
  {
   "name": "Datadog LLM Observability",
   "maker": "Datadog",
   "note": "The incumbent play: agent tracing + experiments + evaluators inside the APM you already pay for; free to 40k spans, Pro from $160/mo",
   "url": "https://www.datadoghq.com/product/llm-observability/",
   "oss": false,
   "entry": "free \u00b7 $160/mo",
   "status": "active"
  },
  {
   "name": "New Relic AI Monitoring",
   "maker": "New Relic",
   "note": "APM-bundled LLM observability; fine if you're already a customer, nobody's first choice for evals",
   "url": "https://newrelic.com/platform/ai-monitoring",
   "oss": false,
   "entry": "usage-based",
   "status": "active"
  },
  {
   "name": "Patronus AI",
   "maker": "Patronus AI",
   "note": "Evaluation API and research-grade judges (Lynx hallucination model, Percival agent debugger)",
   "url": "https://www.patronus.ai",
   "oss": false,
   "entry": "free tier",
   "status": "active"
  },
  {
   "name": "Ragas",
   "maker": "Ragas (ex-Exploding Gradients)",
   "note": "The default OSS metric library for RAG evals (faithfulness, context precision); framework, not platform",
   "url": "https://github.com/explodinggradients/ragas",
   "oss": true,
   "entry": "free",
   "status": "active"
  },
  {
   "name": "Maxim AI",
   "maker": "Maxim",
   "note": "Agent simulation + evals + observability platform; aggressive on agent-testing use cases",
   "url": "https://www.getmaxim.ai",
   "oss": false,
   "entry": "free tier",
   "status": "active"
  },
  {
   "name": "LangWatch",
   "maker": "LangWatch",
   "note": "EU-based open-core LLM ops with optimization studio (DSPy-powered)",
   "url": "https://langwatch.ai",
   "oss": true,
   "entry": "free tier",
   "status": "active"
  },
  {
   "name": "Lunary",
   "maker": "Lunary",
   "note": "Lightweight open-source observability + prompt management; small but steady",
   "url": "https://lunary.ai",
   "oss": true,
   "entry": "free tier",
   "status": "active"
  },
  {
   "name": "PromptLayer",
   "maker": "PromptLayer",
   "note": "Prompt-management-first with evals attached; popular with non-engineer prompt owners",
   "url": "https://www.promptlayer.com",
   "oss": false,
   "entry": "free tier",
   "status": "active"
  },
  {
   "name": "Traceloop (OpenLLMetry)",
   "maker": "Traceloop",
   "note": "Maintainer of OpenLLMetry, the OTel LLM instrumentation many platforms ingest; thin commercial layer",
   "url": "https://www.traceloop.com",
   "oss": true,
   "entry": "free tier",
   "status": "active"
  },
  {
   "name": "Humanloop",
   "maker": "Humanloop \u2192 Anthropic",
   "note": "First-mover LLM eval/prompt platform; team acqui-hired by Anthropic, platform sunset with migration guides (2025)",
   "url": "https://humanloop.com",
   "oss": false,
   "entry": "\u2014",
   "status": "dead"
  },
  {
   "name": "TruEra / TruLens",
   "maker": "TruEra \u2192 Snowflake",
   "note": "ML-observability pioneer acquired by Snowflake (May 2024); TruLens OSS evals live on under Snowflake",
   "url": "https://www.trulens.org",
   "oss": true,
   "entry": "free (OSS)",
   "status": "acquired"
  },
  {
   "name": "Aporia",
   "maker": "Aporia \u2192 Coralogix",
   "note": "ML/LLM observability + guardrails, acquired by Coralogix (Dec 2024); now Coralogix's AI Center",
   "url": "https://coralogix.com/ai-blog/coralogix-acquires-aporia/",
   "oss": false,
   "entry": "\u2014",
   "status": "acquired"
  },
  {
   "name": "Weights & Biases (company)",
   "maker": "W&B \u2192 CoreWeave",
   "note": "The whole company was acquired by CoreWeave (2025, reported ~$1.7B) \u2014 Weave continues as its LLM-ops arm",
   "url": "https://wandb.ai",
   "oss": false,
   "entry": "\u2014",
   "status": "acquired"
  },
  {
   "name": "Langfuse (company)",
   "maker": "Langfuse \u2192 ClickHouse",
   "note": "Acquired by ClickHouse (Jan 2026); product continues under its own brand \u2014 see top 5",
   "url": "https://langfuse.com",
   "oss": true,
   "entry": "\u2014",
   "status": "acquired"
  }
 ],
 "signals": [
  {
   "fact": "Gartner: 40%+ of agentic AI projects will be canceled by end of 2027 \u2014 escalating costs, unclear business value, inadequate risk controls (Jun 25, 2025). This is the demand driver for the whole category",
   "src": "https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027"
  },
  {
   "fact": "Braintrust $80M Series B led by ICONIQ (Feb 17, 2026); ~$800M valuation per press reporting; named customers Notion, Replit, Cloudflare, Ramp, Dropbox",
   "src": "https://www.braintrust.dev/blog/announcing-series-b"
  },
  {
   "fact": "LangChain $125M at $1.25B (IVP, Oct 20, 2025); LangSmith commercial trace volume 12x YoY \u2014 evals/observability is the monetization layer of the framework business",
   "src": "https://www.langchain.com/blog/series-b"
  },
  {
   "fact": "Consolidation wave: Langfuse \u2192 ClickHouse (Jan 2026), W&B \u2192 CoreWeave (2025), Humanloop team \u2192 Anthropic with platform sunset (2025), Aporia \u2192 Coralogix (Dec 2024), TruEra \u2192 Snowflake (May 2024)",
   "src": "https://langfuse.com/blog"
  },
  {
   "fact": "Arize $70M Series C (Feb 20, 2025) billed as the largest-ever investment in AI observability; Datadog and Microsoft's M12 on the cap table \u2014 incumbents buying visibility into the category",
   "src": "https://arize.com/blog/arize-ai-raises-70m-series-c/"
  },
  {
   "fact": "APM incumbents now bundle the job: Datadog LLM Observability ships datasets/experiments/evaluators from $160/mo, free to 40k spans (Aug 2026) \u2014 squeezing point solutions from above",
   "src": "https://www.datadoghq.com/product/llm-observability/"
  }
 ],
 "notes": "Ranking criteria: evidence of production adoption (named customers, download/star trajectories), pricing accessibility for a small team, lock-in risk (license + data portability), and eval-workflow depth \u2014 not just tracing. Conflicts resolved: Braintrust's $800M valuation comes from press reporting around the round, not the company's own Feb 17, 2026 post (which discloses no valuation) \u2014 treat as reported, not confirmed. The W&B/CoreWeave ~$1.7B price is likewise reported, never officially confirmed. OpenAI's evals repo showed commit activity on fetch but community consensus is it's effectively frozen as a benchmark registry; the maintained product is the platform Evals API \u2014 we mark the repo 'fading', which a human should sanity-check. Langfuse GitHub shows 32.4k stars vs 32.6k on its pricing page; we cite the vendor page number, the delta is cache lag. Adjacent elements: LLM gateways with logging (Portkey, OpenRouter analytics, LiteLLM) belong to Rt \u00b7 Model Router even though they sell 'observability'; promptfoo's red-teaming half and Galileo's runtime guardrails overlap Gd \u00b7 Guardrails; token-cost dashboards alone are Sp \u00b7 Spend Management. Uptime/APM monitoring of the app around the model stays with classic observability, out of scope here.",
 "sources": [
  "https://langfuse.com/pricing",
  "https://github.com/langfuse/langfuse",
  "https://langfuse.com/blog",
  "https://www.braintrust.dev/pricing",
  "https://www.braintrust.dev/blog/announcing-series-b",
  "https://www.braintrust.dev/blog",
  "https://www.langchain.com/pricing",
  "https://www.langchain.com/blog/series-b",
  "https://arize.com/phoenix",
  "https://arize.com/blog/arize-ai-raises-70m-series-c/",
  "https://wandb.ai/site/weave",
  "https://wandb.ai/site/pricing/",
  "https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027",
  "https://github.com/comet-ml/opik",
  "https://github.com/promptfoo/promptfoo",
  "https://github.com/confident-ai/deepeval",
  "https://github.com/openai/evals",
  "https://www.helicone.ai/pricing",
  "https://www.confident-ai.com/pricing",
  "https://www.datadoghq.com/product/llm-observability/",
  "https://galileo.ai",
  "https://humanloop.com"
 ],
 "element": {
  "number": 47,
  "name": "Evals & Observability",
  "group": "Trust & Compliance",
  "essential": true,
  "edition": "v2026.Q3",
  "revision": "r7",
  "license": "CC BY 4.0 \u2014 cite elems.ai",
  "url": "https://elems.ai/e/ev.html"
 }
}