ElevenLabs alternatives, 2026.Q3: every real option, ranked
The short answer
Edition v2026.Q3 · pricing and status verified 2026-08-06.
Cartesia (Sonic) is the strongest ElevenLabs alternative for most — real-time voice — agents, IVR, live products — where sub-100ms latency and turn-taking accuracy matter more than a giant voice library. Then: OpenAI TTS / Realtime · Hume (Octave / EVI) · Fish Audio (OpenAudio). Below, all 22 real options in the voice category, with pricing and honest watch-outs — plus the 6 "alternatives" other lists still recommend that are dead, renamed, or sunsetting.
Why people look past ElevenLabs at all: The quality crown is gone: Eleven v3 sits #11 on the Artificial Analysis blind arena (Elo 1171), behind Qwen, Speechify, Gemini, and Cartesia. Credit pricing gets expensive at scale against $0.65–25/1M-char rivals, and the platform sprawl — music, agents, 'Reception AI' — is a focus risk.
The top ElevenLabs alternatives, ranked
2Cartesia (Sonic)Cartesia
Free 20k credits · Pro $5/mo · Startup $49 · Scale $299 · Enterprise custom; Line agent calls $0.06/minBest for Real-time voice — agents, IVR, live products — where sub-100ms latency and turn-taking accuracy matter more than a giant voice library.
Watch A fraction of ElevenLabs' scale ($64M Series A, Mar 2025 — no larger round announced by Aug 2026). Marketing says '#1 for naturalness'; the overall arena board says #5. Voice library and creator tooling are thin — it's a developer product, not a studio.
3OpenAI TTS / RealtimeOpenAI
API only: gpt-4o-mini-tts ≈$0.015/min (est.) · gpt-realtime-2.1-mini $10/1M audio-in, $20/1M audio-out · full gpt-realtime-2.1 $32/$64Best for Teams already on the OpenAI stack that want instructable, good-enough voice — 13 voices, steerable tone — with zero new vendors and near-zero cost.
Watch No voice cloning at all (a deliberate safety stance) and no custom voice design; quality ranks #30 on the blind arena (TTS-1 HD, Elo 1097). Dedicated TTS pricing has been folded into the audio/realtime page — line-item costs are harder to pin down than a year ago.
4Hume (Octave / EVI)Hume AI
Free · Starter $3/mo · Creator $14 · Pro $70 · Scale $200 · Business $500 · Enterprise; TTS overage $0.05–0.15/1k chars, EVI $0.04–0.07/minBest for Voice as performance — Octave 'acts' a line from context and direction rather than reading it, and EVI adds emotionally-aware real-time speech-to-speech.
Watch The emotion thesis hasn't won blind listening: Octave 2 ranks #51 (Elo 1052) on the AA arena. Hume's own April 2026 essay concedes 'voice models are commoditizing' — its value case now rests on the empathic layer, a much narrower moat than a platform.
5Fish Audio (OpenAudio)Hanabi AI
Freemium studio · S2.1 Pro free developer tier · paid plans + usage-based API · commercial use on paid plansBest for Cost-sensitive builders who want near-frontier quality with an open-weight lineage — and a 2M-voice community library to draw from.
Watch Seed-stage company carrying real trust-and-safety surface: a 2M-voice community library makes consent provenance hard to police. Public pricing is opaque next to rivals, and enterprise compliance machinery (SOC 2, BAAs) trails the leaders.
Every other live option in voice
| Tool | Maker | What it is | Entry |
|---|---|---|---|
| Resemble AI | Resemble AI | Cloning pioneer now leading with deepfake Detect + watermarking (per-second pricing); open-sourced Chatterbox — a telling category pivot from generation to trust | Flex $0/mo pay-go · Team $350/mo |
| Chatterbox | Resemble AI (MIT) | Leading open TTS/cloning family — 25.9k stars, Multilingual V3 (23+ languages), [laugh]-style tags, Turbo/Nano variants for CPU | free |
| Kokoro-82M | hexgrad (Apache-2.0) | 82M-param open model; cheapest serious TTS on the AA index at $0.65/1M chars hosted — the commodity floor for clean narration | free |
| MiniMax Audio (Speech 2.8) | MiniMax | Top-10 arena quality (HD Elo 1172); listed on HKEX Jan 9, 2026 — but carries Disney/Universal/WBD copyright suit (Sep 2025) and an Anthropic fraud accusation (Feb 2026) | usage-based |
| Speechify (Simba 3.2) | Speechify AI | Consumer reading app turned API vendor — Simba 3.2 ranks #2 on the blind arena (Elo 1227), sub-100ms streaming | free tier · API usage |
| Qwen-Audio-3.0-TTS-Plus | Alibaba Cloud | #1 on the AA arena (Elo 1229, Aug 2026) — proof that quality leadership has left the pure-play vendors | usage-based |
| Gemini TTS (3.1 Flash) | #3 on the arena (Elo 1210), bundled into the Gemini API — the big-cloud default for GCP shops | usage-based | |
| Azure AI Speech | Microsoft | Enterprise incumbent: custom neural voice with strict consent gates, compliance depth over quality rank | usage-based + free tier |
| Deepgram Aura-2 | Deepgram | STT leader's agent-focused TTS at $0.030/1k chars — bought for the bundle, not the voice | $0.030/1k chars pay-go |
| Inworld TTS | Inworld AI | Gaming-heritage vendor with two models in the arena top 10; aggressive $5–25/1M-char pricing ladder | $15–35/1M chars on-demand |
| Smallest.ai (Lightning) | Smallest.ai | India-based real-time specialist; Lightning V3.1 Pro ranks #8 on the arena (Elo 1192) | usage-based |
| Murf AI | Murf | SMB studio standard for e-learning/corporate VO; Falcon 2 sits on AA's price-quality frontier, but flagship studio voices rank #74 (Elo 975) | free · Creator from $19/mo (annual) |
| WellSaid | WellSaid (now wellsaid.io) | Enterprise L&D voice-over with 280+ licensed voice actors — the consent-first counterpoint to community cloning | free trial · seat plans |
| LOVO / Genny | LOVO | 2M-user creator studio; still shadowed by the Lehrman v. Lovo voice-actor consent lawsuit — the category's precedent case | free trial · sub plans |
| Sesame (CSM) | Sesame AI | Open-sourced CSM-1B (Mar 2025) and the eerily-natural Maya demo; company has pivoted to a consumer personal agent + 2027 eyewear — voice model now a means, not the product | free preview |
| Respeecher | Respeecher | Film/TV-grade voice cloning with per-project consent contracts (Hollywood credits); services-heavy, not self-serve-first | project pricing · marketplace plans |
| Camb.ai (MARS) | Camb.ai | Dubbing/translation specialist (sports leagues, 140+ languages claimed) — overlaps localization more than product voice | freemium |
| VUI Labs (Luna) | VUI Labs | Newcomer at #4 on the AA arena (Elo 1209, Aug 2026) — little public company footprint yet; watch or wait | unverified |
The "ElevenLabs alternatives" to avoid — no longer what they were
Listicles still recommend these. As of 2026-08-06, they are not what the listicles think.
| Tool | Status | What happened |
|---|---|---|
| PlayHT / PlayAI | acquired | Acquired by Meta Jul 11, 2025 (Bloomberg; undisclosed) — team absorbed into Meta's voice efforts; play.ht product still online but effectively in maintenance |
| Amazon Polly | fading | The 2016-era workhorse — fastest on AA's index (1,149 chars/sec) and cheap, but generative-voice quality has passed it by |
| Dia | fading | Viral open 1.6B dialogue model (Apr 2025) with nonverbal tags; momentum cooled as Chatterbox and Fish absorbed the OSS mindshare |
| OpenVoice | fading | 2024's instant-cloning breakout; little movement since — superseded by Chatterbox-class models |
| Bark | dead | Early generative-audio OSS star; unmaintained since Suno went all-in on music |
| Tortoise-TTS | dead | The 2022 OSS ancestor of the modern wave; author joined OpenAI, repo archived in spirit |
How to choose
- If you're producing narration, audiobooks, dubbing, or a branded product voice and want one vendor for the whole workflow
- ElevenLabs — the Studio + 70-language + pro-cloning stack is unmatched, and 41% of the Fortune 500 already cleared it through procurement.
- If you're building a real-time voice agent and the latency budget is under 100ms
- Cartesia — Sonic 3.5 (<90ms claimed, #5 arena Elo) plus Ink-2's #1-ranked turn detection; hand the agent orchestration layer to element Vs.
- If you're already on OpenAI and voice is a feature, not the product
- gpt-4o-mini-tts at ~$0.015/min — accept no cloning and mid-pack quality for zero new vendors and one-line integration.
- If the line has to be performed — empathy, sarcasm, de-escalation — not just read
- Hume Octave (or ElevenLabs v3 audio tags); judge by ear on your script, because blind-arena Elo says the emotion premium is not free.
- If volume is huge, margins are thin, or audio must stay on your infra
- Open weights: Chatterbox V3 (MIT, 23+ languages) for cloning, Kokoro-82M at $0.65/1M chars for cheap clean narration, Fish's S1/S2 for quality self-hosting.
This analysis is drawn from the Voice element dossier — the ranked top 5, the comparison matrix, and the complete field of 24 more tools live there, with every source. Data: vo.json (CC BY 4.0).
Every claim above is dated and sourced from the elems dossiers — 1,421 tools tracked across 58 categories, verified 2026-08-06, including the 276 we found dead, renamed, acquired, or sunsetting. Rankings are editorial, never paid — the charter.
Build your stack in 5 questions →