Langfuse alternatives, 2026.Q3: every real option, ranked
The short answer
Edition v2026.Q3 · pricing and status verified 2026-08-06.
Braintrust is the strongest Langfuse alternative for most — teams that treat evals as the product-development loop itself — experiments, datasets, human review, and a purpose-built trace store (Brainstore) with the best eval workflow in the category. Then: LangSmith · Arize Phoenix · W&B Weave. Below, all 19 real options in the evals & observability category, with pricing and honest watch-outs — plus the 6 "alternatives" other lists still recommend that are dead, renamed, or sunsetting.
Why people look past Langfuse at all: Eval UX and experiment workflows trail Braintrust's — Langfuse is observability-first, evals-second. Post-acquisition roadmap now serves ClickHouse's platform ambitions too; the /ee folders are not MIT, so 'fully open source' has an asterisk.
The top Langfuse alternatives, ranked
2BraintrustBraintrust Data
Starter free (1GB data, 10k scores) · Pro $249/mo · Enterprise custom (on-prem or hosted); 6–12 mo free for startupsBest for Teams that treat evals as the product-development loop itself — experiments, datasets, human review, and a purpose-built trace store (Brainstore) with the best eval workflow in the category.
Watch Closed source with a proprietary data store — the deepest lock-in of the top five. $249/mo Pro plus usage ($3/GB, $1.50/1k scores) makes it the priciest non-enterprise entry; the free tier's 14-day retention is the shortest here.
3LangSmithLangChain
Developer free (5k traces/mo) · Plus $39/seat/mo (10k traces) · then pay-as-you-go · Enterprise custom (self-host/hybrid)Best for Anyone building on LangChain or LangGraph — zero-config tracing of the framework 35% of the Fortune 500 touches, plus deployment and agent infra in the same subscription.
Watch Closed source, and outside the LangChain ecosystem it's just another good tracer — the instrumentation advantage evaporates. Per-seat + LCU/LSU metered pricing is the hardest here to forecast. Framework coupling cuts both ways if you migrate off LangGraph.
4Arize PhoenixArize AI
OSS free to self-host (ELv2) · free Phoenix Cloud instances · Arize AX (commercial platform) customBest for OTel purists and self-hosters — vendor-neutral OpenTelemetry tracing plus evals that stay entirely in your infrastructure, with a funded enterprise platform (AX) behind it when you outgrow free.
Watch ELv2, not MIT/Apache — fine for self-hosting, but not fully open. Prompt management and eval workflow are thinner than Langfuse/Braintrust; the upsell path to Arize AX is where the polish (and the price) lives.
5W&B WeaveWeights & Biases (a CoreWeave company)
Free tier (1GB/mo Weave ingestion) · Pro from $60/mo (1.5GB/mo incl., then usage) · Enterprise custom incl. self-managedBest for Teams already in the W&B ecosystem, or running multi-agent systems that want sessions/turns/sub-agents as first-class trace concepts rather than bolted-on spans.
Watch It's a module inside a bigger ML platform — if you don't want experiment tracking and model registry, you're navigating around them. Post-acquisition, W&B's roadmap now serves CoreWeave's cloud strategy; ingestion overage pricing punishes verbose traces.
Every other live option in evals & observability
| Tool | Maker | What it is | Entry |
|---|---|---|---|
| Opik | Comet | The other big OSS platform — 20.7k stars, Apache-2.0, fully self-hostable including backend; the near-miss for the top 5 | free |
| DeepEval / Confident AI | Confident AI | Pytest-for-LLMs framework (15.4k stars, Apache-2.0) + commercial platform from $200/mo; the CI-evals standard | free · $200/mo |
| promptfoo | Promptfoo Inc. | MIT eval + red-teaming CLI, 21.3k stars, local-first; its security half overlaps element Gd · Guardrails | free |
| Helicone | Helicone | One-line proxy for LLM logging and cost analytics (5.8k stars); observability via gateway, thinner on evals | free · $79/mo |
| Arize AX | Arize AI | Commercial sibling of Phoenix — enterprise agent engineering platform; $70M Series C (Feb 2025) | custom |
| Galileo | Galileo (Rungalileo) | Enterprise 'eval engineering' — Luna distilled evaluator models claim 96% cheaper production scoring; NVIDIA, HP, MongoDB testimonials | free tier · custom |
| Datadog LLM Observability | Datadog | The incumbent play: agent tracing + experiments + evaluators inside the APM you already pay for; free to 40k spans, Pro from $160/mo | free · $160/mo |
| New Relic AI Monitoring | New Relic | APM-bundled LLM observability; fine if you're already a customer, nobody's first choice for evals | usage-based |
| Patronus AI | Patronus AI | Evaluation API and research-grade judges (Lynx hallucination model, Percival agent debugger) | free tier |
| Ragas | Ragas (ex-Exploding Gradients) | The default OSS metric library for RAG evals (faithfulness, context precision); framework, not platform | free |
| Maxim AI | Maxim | Agent simulation + evals + observability platform; aggressive on agent-testing use cases | free tier |
| LangWatch | LangWatch | EU-based open-core LLM ops with optimization studio (DSPy-powered) | free tier |
| Lunary | Lunary | Lightweight open-source observability + prompt management; small but steady | free tier |
| PromptLayer | PromptLayer | Prompt-management-first with evals attached; popular with non-engineer prompt owners | free tier |
| Traceloop (OpenLLMetry) | Traceloop | Maintainer of OpenLLMetry, the OTel LLM instrumentation many platforms ingest; thin commercial layer | free tier |
The "Langfuse alternatives" to avoid — no longer what they were
Listicles still recommend these. As of 2026-08-06, they are not what the listicles think.
| Tool | Status | What happened |
|---|---|---|
| OpenAI Evals | fading | The 2023 pioneer repo (18.6k stars, MIT) is quiet; the live product is the Evals API/dashboard inside the OpenAI platform — OpenAI-centric by design |
| Humanloop | dead | First-mover LLM eval/prompt platform; team acqui-hired by Anthropic, platform sunset with migration guides (2025) |
| TruEra / TruLens | acquired | ML-observability pioneer acquired by Snowflake (May 2024); TruLens OSS evals live on under Snowflake |
| Aporia | acquired | ML/LLM observability + guardrails, acquired by Coralogix (Dec 2024); now Coralogix's AI Center |
| Weights & Biases (company) | acquired | The whole company was acquired by CoreWeave (2025, reported ~$1.7B) — Weave continues as its LLM-ops arm |
| Langfuse (company) | acquired | Acquired by ClickHouse (Jan 2026); product continues under its own brand — see top 5 |
How to choose
- If you want one tool, no lock-in, and a price that doesn't scale with panic
- Langfuse — MIT-core self-host or $29/mo cloud, now with ClickHouse's balance sheet behind it (Jan 2026).
- If evals drive your product decisions and you have real budget
- Braintrust — the eval UX Notion, Replit, and Cloudflare standardized on; take the 6–12 months free startup credit.
- If your stack is LangChain/LangGraph
- LangSmith — zero-config tracing plus deployment in one bill; its 12x YoY trace growth is the ecosystem talking.
- If you require traces to never leave your infra, at zero license cost
- Arize Phoenix (ELv2, OTel-native) or Opik (Apache-2.0, fully self-hostable) — the two credible free self-host paths.
- If you're shipping anything customer-facing without evals in CI
- Stop — wire pytest-style checks (DeepEval, promptfoo) into CI this week. Gartner's 40%-canceled-by-2027 cohort is made of teams that couldn't prove ROI.
This analysis is drawn from the Evals & Observability element dossier — the ranked top 5, the comparison matrix, and the complete field of 21 more tools live there, with every source. Data: ev.json (CC BY 4.0).
Every claim above is dated and sourced from the elems dossiers — 1,421 tools tracked across 58 categories, verified 2026-08-06, including the 276 we found dead, renamed, acquired, or sunsetting. Rankings are editorial, never paid — the charter.
Build your stack in 5 questions →