Vital signs
Why it's on the table
On the table, Voice (Vo) is seat 22 of 58, in the Creative & Design family. It is a stable element — the category has settled, choosing is cheap, and switching is rare. Pick a holder and move on; this is not where your decision budget should go. It is optional: plenty of companies run without it — until a specific trigger (scale, regulation, cost, or customers) makes it essential for them. It sits in the lowest paid band — lunch money against the hours it returns.
Voice: the top 5 — v2026.Q3
Edition v2026.Q3 · ranking, pricing and status verified 2026-08-06.
1ElevenLabsElevenLabs
Free 10k credits · Starter $6/mo · Creator $22 · Pro $99 · Scale $299 · Business $990 · Enterprise custom; TTS ≈$0.17–0.36/minBest for Production narration, audiobooks, dubbing, and product voices where you want the full stack — voice library, pro cloning, Studio, 70+ languages — behind one API.
The category owner by revenue and reach: ~$500M ARR by April 2026 (from $350M at end-2025), 41% of the Fortune 500, a $500M Sequoia-led Series D at $11B (Feb 4, 2026), and distribution wins like Spotify's ElevenLabs-powered audiobook tool (May 2026). Eleven v3's audio tags ([whispers], [sighs]) and 70+ languages made expressive, directable speech mainstream.
Watch The quality crown is gone: Eleven v3 sits #11 on the Artificial Analysis blind arena (Elo 1171), behind Qwen, Speechify, Gemini, and Cartesia. Credit pricing gets expensive at scale against $0.65–25/1M-char rivals, and the platform sprawl — music, agents, 'Reception AI' — is a focus risk.
2Cartesia (Sonic)Cartesia
Free 20k credits · Pro $5/mo · Startup $49 · Scale $299 · Enterprise custom; Line agent calls $0.06/minBest for Real-time voice — agents, IVR, live products — where sub-100ms latency and turn-taking accuracy matter more than a giant voice library.
The performance challenger: Sonic 3.5 is the top-ranked major API-first voice vendor on the blind arena (#5, Elo 1203 — above ElevenLabs), claims sub-90ms latency across 40+ languages, and pairs with Ink-2, the #1-ranked streaming STT for voice agents (Jul 2026). Built by the state-space-model (Mamba) researchers, so the speed claims have architectural teeth.
Watch A fraction of ElevenLabs' scale ($64M Series A, Mar 2025 — no larger round announced by Aug 2026). Marketing says '#1 for naturalness'; the overall arena board says #5. Voice library and creator tooling are thin — it's a developer product, not a studio.
3OpenAI TTS / RealtimeOpenAI
API only: gpt-4o-mini-tts ≈$0.015/min (est.) · gpt-realtime-2.1-mini $10/1M audio-in, $20/1M audio-out · full gpt-realtime-2.1 $32/$64Best for Teams already on the OpenAI stack that want instructable, good-enough voice — 13 voices, steerable tone — with zero new vendors and near-zero cost.
The default, not the best: gpt-4o-mini-tts lets you steer delivery in plain English ('sound like a sympathetic agent') at commodity prices, and the Realtime API (gpt-realtime-2.1) handles full speech-to-speech for voice products. Distribution is the moat — every OpenAI developer has it one endpoint away.
Watch No voice cloning at all (a deliberate safety stance) and no custom voice design; quality ranks #30 on the blind arena (TTS-1 HD, Elo 1097). Dedicated TTS pricing has been folded into the audio/realtime page — line-item costs are harder to pin down than a year ago.
4Hume (Octave / EVI)Hume AI
Free · Starter $3/mo · Creator $14 · Pro $70 · Scale $200 · Business $500 · Enterprise; TTS overage $0.05–0.15/1k chars, EVI $0.04–0.07/minBest for Voice as performance — Octave 'acts' a line from context and direction rather than reading it, and EVI adds emotionally-aware real-time speech-to-speech.
The emotional-intelligence specialist: Octave generates delivery from meaning (frustration, sarcasm, comfort) with director-style prompts, EVI is one of the few production speech-to-speech APIs, and the team publishes real research (RW-Voice-EQ benchmark for the human quality of voice AI, Jul 2026). Cheapest serious entry point in the category at $3/mo.
Watch The emotion thesis hasn't won blind listening: Octave 2 ranks #51 (Elo 1052) on the AA arena. Hume's own April 2026 essay concedes 'voice models are commoditizing' — its value case now rests on the empathic layer, a much narrower moat than a platform.
5Fish Audio (OpenAudio)Hanabi AI
Freemium studio · S2.1 Pro free developer tier · paid plans + usage-based API · commercial use on paid plansBest for Cost-sensitive builders who want near-frontier quality with an open-weight lineage — and a 2M-voice community library to draw from.
The open-ecosystem value play: Fish Speech → S1 → S2 shipped as open models, S2.1 Pro ranks #15 on the blind arena (Elo 1138) — above OpenAI and Hume — and the company made it free for developers by rebuilding its inference stack. $52M seed, 8M+ builders, 30+ languages, 2M+ community voices.
Watch Seed-stage company carrying real trust-and-safety surface: a 2M-voice community library makes consent provenance hard to police. Public pricing is opaque next to rivals, and enterprise compliance machinery (SOC 2, BAAs) trails the leaders.
Voice: the top 8 compared
Edition v2026.Q3 · ranking, pricing and status verified 2026-08-06.
| Tool | Entry price | Typical API cost | Arena Elo (AA, Aug '26) | Latency claim | Voice cloning | Languages | Open weights |
|---|---|---|---|---|---|---|---|
| ElevenLabs | $6/mo | ≈$0.17–0.36/min (credits) | 1171 · #11 | low-lat. mode $0.05/min (Business) | Instant + professional | 70+ | No |
| Cartesia Sonic 3.5 | $5/mo | credits; agents $0.06/min | 1203 · #5 | <90 ms (claimed) | Instant (10s) + pro | 40+ | No |
| OpenAI gpt-4o-mini-tts | usage only | ≈$0.015/min (est.) | 1097 · #30 (TTS-1 HD) | Realtime API for live | None (policy) | Multilingual | No |
| Hume Octave 2 | $3/mo | $0.05–0.15/1k chars | 1052 · #51 | EVI real-time | Instant | Multilingual | No |
| Fish Audio S2.1 Pro | free | usage API; free dev tier | 1138 · #15 | Real-time streaming | Community + instant | 30+ | Yes (S1/S2) |
| MiniMax Speech 2.8 | usage | per-char API | 1172 · #10 (HD) | Streaming | Instant | 30+ | No |
| Chatterbox V3 | free (MIT) | self-host | n/a (not listed) | Self-host dependent | Zero-shot | 23+ | Yes |
| Kokoro-82M | free (Apache) | $0.65/1M chars (hosted) | n/a (not listed) | Fast, runs on CPU | No | 9 | Yes (82M) |
How to choose your voice
- If you're producing narration, audiobooks, dubbing, or a branded product voice and want one vendor for the whole workflow
- ElevenLabs — the Studio + 70-language + pro-cloning stack is unmatched, and 41% of the Fortune 500 already cleared it through procurement.
- If you're building a real-time voice agent and the latency budget is under 100ms
- Cartesia — Sonic 3.5 (<90ms claimed, #5 arena Elo) plus Ink-2's #1-ranked turn detection; hand the agent orchestration layer to element Vs.
- If you're already on OpenAI and voice is a feature, not the product
- gpt-4o-mini-tts at ~$0.015/min — accept no cloning and mid-pack quality for zero new vendors and one-line integration.
- If the line has to be performed — empathy, sarcasm, de-escalation — not just read
- Hume Octave (or ElevenLabs v3 audio tags); judge by ear on your script, because blind-arena Elo says the emotion premium is not free.
- If volume is huge, margins are thin, or audio must stay on your infra
- Open weights: Chatterbox V3 (MIT, 23+ languages) for cloning, Kokoro-82M at $0.65/1M chars for cheap clean narration, Fish's S1/S2 for quality self-hosting.
Voice: the whole field
24 more tools tracked in this category, including 6 dead, renamed, or sunsetting — a reference that hides the graveyard isn't one. Verified 2026-08-06.
| Tool | Maker | What it is | Entry | Status |
|---|---|---|---|---|
| PlayHT / PlayAI | PlayAI → Meta | Acquired by Meta Jul 11, 2025 (Bloomberg; undisclosed) — team absorbed into Meta's voice efforts; play.ht product still online but effectively in maintenance | freemium | acquired |
| Resemble AI | Resemble AI | Cloning pioneer now leading with deepfake Detect + watermarking (per-second pricing); open-sourced Chatterbox — a telling category pivot from generation to trust | Flex $0/mo pay-go · Team $350/mo | active |
| Chatterbox | Resemble AI (MIT) | Leading open TTS/cloning family — 25.9k stars, Multilingual V3 (23+ languages), [laugh]-style tags, Turbo/Nano variants for CPU | free | active |
| Kokoro-82M | hexgrad (Apache-2.0) | 82M-param open model; cheapest serious TTS on the AA index at $0.65/1M chars hosted — the commodity floor for clean narration | free | active |
| MiniMax Audio (Speech 2.8) | MiniMax | Top-10 arena quality (HD Elo 1172); listed on HKEX Jan 9, 2026 — but carries Disney/Universal/WBD copyright suit (Sep 2025) and an Anthropic fraud accusation (Feb 2026) | usage-based | active |
| Speechify (Simba 3.2) | Speechify AI | Consumer reading app turned API vendor — Simba 3.2 ranks #2 on the blind arena (Elo 1227), sub-100ms streaming | free tier · API usage | active |
| Qwen-Audio-3.0-TTS-Plus | Alibaba Cloud | #1 on the AA arena (Elo 1229, Aug 2026) — proof that quality leadership has left the pure-play vendors | usage-based | active |
| Gemini TTS (3.1 Flash) | #3 on the arena (Elo 1210), bundled into the Gemini API — the big-cloud default for GCP shops | usage-based | active | |
| Azure AI Speech | Microsoft | Enterprise incumbent: custom neural voice with strict consent gates, compliance depth over quality rank | usage-based + free tier | active |
| Amazon Polly | AWS | The 2016-era workhorse — fastest on AA's index (1,149 chars/sec) and cheap, but generative-voice quality has passed it by | usage-based + free tier | fading |
| Deepgram Aura-2 | Deepgram | STT leader's agent-focused TTS at $0.030/1k chars — bought for the bundle, not the voice | $0.030/1k chars pay-go | active |
| Inworld TTS | Inworld AI | Gaming-heritage vendor with two models in the arena top 10; aggressive $5–25/1M-char pricing ladder | $15–35/1M chars on-demand | active |
| Smallest.ai (Lightning) | Smallest.ai | India-based real-time specialist; Lightning V3.1 Pro ranks #8 on the arena (Elo 1192) | usage-based | active |
| Murf AI | Murf | SMB studio standard for e-learning/corporate VO; Falcon 2 sits on AA's price-quality frontier, but flagship studio voices rank #74 (Elo 975) | free · Creator from $19/mo (annual) | active |
| WellSaid | WellSaid (now wellsaid.io) | Enterprise L&D voice-over with 280+ licensed voice actors — the consent-first counterpoint to community cloning | free trial · seat plans | active |
| LOVO / Genny | LOVO | 2M-user creator studio; still shadowed by the Lehrman v. Lovo voice-actor consent lawsuit — the category's precedent case | free trial · sub plans | active |
| Sesame (CSM) | Sesame AI | Open-sourced CSM-1B (Mar 2025) and the eerily-natural Maya demo; company has pivoted to a consumer personal agent + 2027 eyewear — voice model now a means, not the product | free preview | active |
| Dia | Nari Labs | Viral open 1.6B dialogue model (Apr 2025) with nonverbal tags; momentum cooled as Chatterbox and Fish absorbed the OSS mindshare | free | fading |
| OpenVoice | MyShell (MIT) | 2024's instant-cloning breakout; little movement since — superseded by Chatterbox-class models | free | fading |
| Bark | Suno | Early generative-audio OSS star; unmaintained since Suno went all-in on music | free | dead |
| Tortoise-TTS | neonbjb (James Betker) | The 2022 OSS ancestor of the modern wave; author joined OpenAI, repo archived in spirit | free | dead |
| Respeecher | Respeecher | Film/TV-grade voice cloning with per-project consent contracts (Hollywood credits); services-heavy, not self-serve-first | project pricing · marketplace plans | active |
| Camb.ai (MARS) | Camb.ai | Dubbing/translation specialist (sports leagues, 140+ languages claimed) — overlaps localization more than product voice | freemium | active |
| VUI Labs (Luna) | VUI Labs | Newcomer at #4 on the AA arena (Elo 1209, Aug 2026) — little public company footprint yet; watch or wait | unverified | active |
Voice: the category in numbers
Edition v2026.Q3 · ranking, pricing and status verified 2026-08-06.
- ElevenLabs: ~$500M ARR by Apr 2026 (from $350M end-2025, +$100M in Q1 alone); 41% of Fortune 500; ~50/50 self-serve vs enterprise [src]
- ElevenLabs $500M Series D at $11B led by Sequoia (Feb 4, 2026); May 2026 third close added BlackRock, NVIDIA, Salesforce, Santander [src]
- Quality commoditized: 91 models on the AA Speech Arena; #1 is Alibaba's Qwen-Audio-3.0-TTS-Plus (Elo 1229) and ElevenLabs ranks #11 — the leader no longer has the best model (Aug 2026) [src]
- Big-tech consolidation of voice talent: Meta acquired PlayAI (Jul 11, 2025, undisclosed); MiniMax listed on HKEX Jan 9, 2026 [src]
- Hume's own thesis, Apr 15, 2026: 'Voice models are commoditizing. The value is moving up the stack' — vendors are repositioning around agents, emotion layers, and trust tooling [src]
- The trust flank is monetizing: cloning pioneer Resemble now leads its pricing page with deepfake Detect and watermarking at per-second rates (Aug 2026) [src]
Voice: method & sources
Ranking charter: we weight production-workflow completeness, verified traction, and API economics over single-utterance arena Elo — blind arenas measure naturalness of short clips, not steerability, cloning quality, or dubbing workflow; that is how ElevenLabs stays #1 while ranking #11 on Elo, and we state the Elo plainly. Conflicts resolved: (1) ElevenLabs ARR — CEO said 'crossed $330M last year' (Jan 13, 2026, TechCrunch) vs Sacra's $350M end-2025 / $500M Apr 2026; we cite Sacra as the fuller series, consistent with the CEO figure. (2) Cartesia claims '#1 for naturalness'; the overall AA board has it #5 — we use the board. (3) The reported July 2026 ElevenLabs tender/secondary talks at ~$22B could not be traced to a fetchable primary source by press time; treat as unconfirmed and excluded from evidence. (4) OpenAI's gpt-4o-mini-tts no longer appears as a line item on the current pricing page; the ~$0.015/min figure is the launch-era estimate — re-verify before budgeting. Aggregator caveat: Sacra figures are single-source estimates. Adjacent elements: voice support/call agents (Vapi, Retell, Bland, ElevenAgents-as-agent-platform) → Vs; talking-head video → Av; meeting transcription → Mt; Eleven Music and Suno's music work have no element and are noted only as ElevenLabs platform sprawl. PlayHT's site remains up post-acquisition — don't build on it. Ranking criteria: verified commercial traction, independent satisfaction surveys, agent benchmarks, and founder-fit (price floor, lock-in, surfaces). Editorial, never paid — the charter. Machine-readable twin: vo.json.
All sources (29)
- https://elevenlabs.io/pricing
- https://elevenlabs.io/docs
- https://sacra.com/c/elevenlabs/
- https://www.cnbc.com/2026/02/04/nvidia-backed-ai-startup-elevenlabs-11-billion-valuation.html
- https://techcrunch.com/tag/elevenlabs/
- https://en.wikipedia.org/wiki/ElevenLabs
- https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice
- https://artificialanalysis.ai/text-to-speech
- https://cartesia.ai/pricing
- https://cartesia.ai/blog
- https://cartesia.ai/sonic
- https://docs.cartesia.ai/
- https://www.hume.ai/pricing
- https://www.hume.ai/blog
- https://dev.hume.ai/docs
- https://developers.openai.com/api/docs/pricing
- https://developers.openai.com/api/docs/guides/text-to-speech
- https://fish.audio/
- https://docs.fish.audio/
- https://github.com/resemble-ai/chatterbox
- https://github.com/hexgrad/kokoro
- https://www.resemble.ai/pricing/
- https://en.wikipedia.org/wiki/List_of_mergers_and_acquisitions_by_Meta_Platforms
- https://en.wikipedia.org/wiki/MiniMax_(company)
- https://deepgram.com/pricing
- https://www.inworld.ai/pricing
- https://murf.ai/pricing
- https://speechify.ai/
- https://www.wellsaid.io/
Our take
Voice quality crossed the uncanny valley in 2025. Narration, support lines, and product voices are now table stakes.
Combines with
Guides
This is element 22 of 58. The table is versioned quarterly — when a tool loses its seat, the changelog records the succession.
Explore the full table →