22 Vo Voice
Group 5 · Creative & Design · element 22 of 58

Voice

Sound like anyone. Ethically, you.

Turns text into production-ready audio.

Holders this quarterElevenLabs · Cartesia (Sonic) · OpenAI TTS / Realtime · Hume (Octave / EVI) · Fish Audio (OpenAudio)

Vital signs

NecessityOptional
Price band$ · under $30/mo
MaturityStable
Editionv2026.Q3
Last verified2026-08-06

Why it's on the table

On the table, Voice (Vo) is seat 22 of 58, in the Creative & Design family. It is a stable element — the category has settled, choosing is cheap, and switching is rare. Pick a holder and move on; this is not where your decision budget should go. It is optional: plenty of companies run without it — until a specific trigger (scale, regulation, cost, or customers) makes it essential for them. It sits in the lowest paid band — lunch money against the hours it returns.

The verdict — v2026.Q3 · verified 2026-08-06
ElevenLabs
ElevenLabs, still — but for platform and distribution, not raw quality. It owns the workflow (voices, cloning, dubbing, Studio, API), hit ~$500M ARR by April 2026, serves 41% of the Fortune 500, and raised a $500M Series D at $11B in February — yet its flagship Eleven v3 ranks only #11 on the blind Artificial Analysis arena, behind Alibaba, Speechify, and Cartesia. Cartesia wins when you're building real-time voice agents and every millisecond counts; OpenAI's gpt-4o-mini-tts wins when 'good enough everywhere' at ~$0.015/min beats best-in-class; Hume wins when you need a voice that acts, not reads; Fish Audio wins when you want open-weight economics with hosted convenience. Voice quality crossed the uncanny valley in 2025 — in 2026 the models are commoditizing and the moat moved to workflow, ecosystem, and trust.

Voice: the top 5 — v2026.Q3

Edition v2026.Q3 · ranking, pricing and status verified 2026-08-06.

  1. 1ElevenLabsElevenLabs

    Free 10k credits · Starter $6/mo · Creator $22 · Pro $99 · Scale $299 · Business $990 · Enterprise custom; TTS ≈$0.17–0.36/min

    Best for Production narration, audiobooks, dubbing, and product voices where you want the full stack — voice library, pro cloning, Studio, 70+ languages — behind one API.

    The category owner by revenue and reach: ~$500M ARR by April 2026 (from $350M at end-2025), 41% of the Fortune 500, a $500M Sequoia-led Series D at $11B (Feb 4, 2026), and distribution wins like Spotify's ElevenLabs-powered audiobook tool (May 2026). Eleven v3's audio tags ([whispers], [sighs]) and 70+ languages made expressive, directable speech mainstream.

    Watch The quality crown is gone: Eleven v3 sits #11 on the Artificial Analysis blind arena (Elo 1171), behind Qwen, Speechify, Gemini, and Cartesia. Credit pricing gets expensive at scale against $0.65–25/1M-char rivals, and the platform sprawl — music, agents, 'Reception AI' — is a focus risk.

    ~$500M ARR by Apr 2026 (+$100M net-new in Q1); 41% of Fortune 500 [src] · $500M Series D at $11B led by Sequoia (Feb 4, 2026) [src] · Eleven v3 ranks #11 (Elo 1171) on the AA Speech Arena (Aug 2026) [src]
  2. 2Cartesia (Sonic)Cartesia

    Free 20k credits · Pro $5/mo · Startup $49 · Scale $299 · Enterprise custom; Line agent calls $0.06/min

    Best for Real-time voice — agents, IVR, live products — where sub-100ms latency and turn-taking accuracy matter more than a giant voice library.

    The performance challenger: Sonic 3.5 is the top-ranked major API-first voice vendor on the blind arena (#5, Elo 1203 — above ElevenLabs), claims sub-90ms latency across 40+ languages, and pairs with Ink-2, the #1-ranked streaming STT for voice agents (Jul 2026). Built by the state-space-model (Mamba) researchers, so the speed claims have architectural teeth.

    Watch A fraction of ElevenLabs' scale ($64M Series A, Mar 2025 — no larger round announced by Aug 2026). Marketing says '#1 for naturalness'; the overall arena board says #5. Voice library and creator tooling are thin — it's a developer product, not a studio.

    Sonic 3.5: #5 overall, Elo 1203 — highest-ranked major voice API (Aug 2026) [src] · Ink-2 STT ranked #1 on AA streaming leaderboard (Jul 9, 2026) [src] · $64M Series A led by Kleiner Perkins (Mar 11, 2025) [src]
  3. 3OpenAI TTS / RealtimeOpenAI

    API only: gpt-4o-mini-tts ≈$0.015/min (est.) · gpt-realtime-2.1-mini $10/1M audio-in, $20/1M audio-out · full gpt-realtime-2.1 $32/$64

    Best for Teams already on the OpenAI stack that want instructable, good-enough voice — 13 voices, steerable tone — with zero new vendors and near-zero cost.

    The default, not the best: gpt-4o-mini-tts lets you steer delivery in plain English ('sound like a sympathetic agent') at commodity prices, and the Realtime API (gpt-realtime-2.1) handles full speech-to-speech for voice products. Distribution is the moat — every OpenAI developer has it one endpoint away.

    Watch No voice cloning at all (a deliberate safety stance) and no custom voice design; quality ranks #30 on the blind arena (TTS-1 HD, Elo 1097). Dedicated TTS pricing has been folded into the audio/realtime page — line-item costs are harder to pin down than a year ago.

    gpt-4o-mini-tts current with 13 voices; tts-1/tts-1-hd now legacy (Aug 2026) [src] · gpt-realtime-2.1-mini: $10/1M audio input, $20/1M audio output (Aug 2026) [src] · OpenAI TTS-1 HD ranks #30 (Elo 1097) on AA arena (Aug 2026) [src]
  4. 4Hume (Octave / EVI)Hume AI

    Free · Starter $3/mo · Creator $14 · Pro $70 · Scale $200 · Business $500 · Enterprise; TTS overage $0.05–0.15/1k chars, EVI $0.04–0.07/min

    Best for Voice as performance — Octave 'acts' a line from context and direction rather than reading it, and EVI adds emotionally-aware real-time speech-to-speech.

    The emotional-intelligence specialist: Octave generates delivery from meaning (frustration, sarcasm, comfort) with director-style prompts, EVI is one of the few production speech-to-speech APIs, and the team publishes real research (RW-Voice-EQ benchmark for the human quality of voice AI, Jul 2026). Cheapest serious entry point in the category at $3/mo.

    Watch The emotion thesis hasn't won blind listening: Octave 2 ranks #51 (Elo 1052) on the AA arena. Hume's own April 2026 essay concedes 'voice models are commoditizing' — its value case now rests on the empathic layer, a much narrower moat than a platform.

    Octave 2 ranks #51 (Elo 1052) on AA arena (Aug 2026) [src] · Published RW-Voice-EQ bench for voice-AI human quality (Jul 2026) [src] · TTS from $3/mo; EVI overage down to $0.04/min at Business tier (Aug 2026) [src]
  5. 5Fish Audio (OpenAudio)Hanabi AI

    Freemium studio · S2.1 Pro free developer tier · paid plans + usage-based API · commercial use on paid plans

    Best for Cost-sensitive builders who want near-frontier quality with an open-weight lineage — and a 2M-voice community library to draw from.

    The open-ecosystem value play: Fish Speech → S1 → S2 shipped as open models, S2.1 Pro ranks #15 on the blind arena (Elo 1138) — above OpenAI and Hume — and the company made it free for developers by rebuilding its inference stack. $52M seed, 8M+ builders, 30+ languages, 2M+ community voices.

    Watch Seed-stage company carrying real trust-and-safety surface: a 2M-voice community library makes consent provenance hard to police. Public pricing is opaque next to rivals, and enterprise compliance machinery (SOC 2, BAAs) trails the leaders.

    $52M seed; 8M+ builders; S2.1 Pro made free for developers (2026) [src] · Fish Audio S2.1 Pro ranks #15 (Elo 1138), above OpenAI and Hume (Aug 2026) [src] · Open models S1/S2 published; 30+ languages, 2M+ voices (Aug 2026) [src]

Voice: the top 8 compared

Edition v2026.Q3 · ranking, pricing and status verified 2026-08-06.

Voice — the top 8 compared. Edition v2026.Q3, verified 2026-08-06.
ToolEntry priceTypical API costArena Elo (AA, Aug '26)Latency claimVoice cloningLanguagesOpen weights
ElevenLabs$6/mo≈$0.17–0.36/min (credits)1171 · #11low-lat. mode $0.05/min (Business)Instant + professional70+No
Cartesia Sonic 3.5$5/mocredits; agents $0.06/min1203 · #5<90 ms (claimed)Instant (10s) + pro40+No
OpenAI gpt-4o-mini-ttsusage only≈$0.015/min (est.)1097 · #30 (TTS-1 HD)Realtime API for liveNone (policy)MultilingualNo
Hume Octave 2$3/mo$0.05–0.15/1k chars1052 · #51EVI real-timeInstantMultilingualNo
Fish Audio S2.1 Profreeusage API; free dev tier1138 · #15Real-time streamingCommunity + instant30+Yes (S1/S2)
MiniMax Speech 2.8usageper-char API1172 · #10 (HD)StreamingInstant30+No
Chatterbox V3free (MIT)self-hostn/a (not listed)Self-host dependentZero-shot23+Yes
Kokoro-82Mfree (Apache)$0.65/1M chars (hosted)n/a (not listed)Fast, runs on CPUNo9Yes (82M)

How to choose your voice

If you're producing narration, audiobooks, dubbing, or a branded product voice and want one vendor for the whole workflow
ElevenLabs — the Studio + 70-language + pro-cloning stack is unmatched, and 41% of the Fortune 500 already cleared it through procurement.
If you're building a real-time voice agent and the latency budget is under 100ms
Cartesia — Sonic 3.5 (<90ms claimed, #5 arena Elo) plus Ink-2's #1-ranked turn detection; hand the agent orchestration layer to element Vs.
If you're already on OpenAI and voice is a feature, not the product
gpt-4o-mini-tts at ~$0.015/min — accept no cloning and mid-pack quality for zero new vendors and one-line integration.
If the line has to be performed — empathy, sarcasm, de-escalation — not just read
Hume Octave (or ElevenLabs v3 audio tags); judge by ear on your script, because blind-arena Elo says the emotion premium is not free.
If volume is huge, margins are thin, or audio must stay on your infra
Open weights: Chatterbox V3 (MIT, 23+ languages) for cloning, Kokoro-82M at $0.65/1M chars for cheap clean narration, Fish's S1/S2 for quality self-hosting.

Voice: the whole field

24 more tools tracked in this category, including 6 dead, renamed, or sunsetting — a reference that hides the graveyard isn't one. Verified 2026-08-06.

Voice — every tool we track, including 6 dead, renamed, or sunsetting. Edition v2026.Q3, verified 2026-08-06.
ToolMakerWhat it isEntryStatus
PlayHT / PlayAIPlayAI → MetaAcquired by Meta Jul 11, 2025 (Bloomberg; undisclosed) — team absorbed into Meta's voice efforts; play.ht product still online but effectively in maintenancefreemiumacquired
Resemble AIResemble AICloning pioneer now leading with deepfake Detect + watermarking (per-second pricing); open-sourced Chatterbox — a telling category pivot from generation to trustFlex $0/mo pay-go · Team $350/moactive
ChatterboxResemble AI (MIT)Leading open TTS/cloning family — 25.9k stars, Multilingual V3 (23+ languages), [laugh]-style tags, Turbo/Nano variants for CPUfreeactive
Kokoro-82Mhexgrad (Apache-2.0)82M-param open model; cheapest serious TTS on the AA index at $0.65/1M chars hosted — the commodity floor for clean narrationfreeactive
MiniMax Audio (Speech 2.8)MiniMaxTop-10 arena quality (HD Elo 1172); listed on HKEX Jan 9, 2026 — but carries Disney/Universal/WBD copyright suit (Sep 2025) and an Anthropic fraud accusation (Feb 2026)usage-basedactive
Speechify (Simba 3.2)Speechify AIConsumer reading app turned API vendor — Simba 3.2 ranks #2 on the blind arena (Elo 1227), sub-100ms streamingfree tier · API usageactive
Qwen-Audio-3.0-TTS-PlusAlibaba Cloud#1 on the AA arena (Elo 1229, Aug 2026) — proof that quality leadership has left the pure-play vendorsusage-basedactive
Gemini TTS (3.1 Flash)Google#3 on the arena (Elo 1210), bundled into the Gemini API — the big-cloud default for GCP shopsusage-basedactive
Azure AI SpeechMicrosoftEnterprise incumbent: custom neural voice with strict consent gates, compliance depth over quality rankusage-based + free tieractive
Amazon PollyAWSThe 2016-era workhorse — fastest on AA's index (1,149 chars/sec) and cheap, but generative-voice quality has passed it byusage-based + free tierfading
Deepgram Aura-2DeepgramSTT leader's agent-focused TTS at $0.030/1k chars — bought for the bundle, not the voice$0.030/1k chars pay-goactive
Inworld TTSInworld AIGaming-heritage vendor with two models in the arena top 10; aggressive $5–25/1M-char pricing ladder$15–35/1M chars on-demandactive
Smallest.ai (Lightning)Smallest.aiIndia-based real-time specialist; Lightning V3.1 Pro ranks #8 on the arena (Elo 1192)usage-basedactive
Murf AIMurfSMB studio standard for e-learning/corporate VO; Falcon 2 sits on AA's price-quality frontier, but flagship studio voices rank #74 (Elo 975)free · Creator from $19/mo (annual)active
WellSaidWellSaid (now wellsaid.io)Enterprise L&D voice-over with 280+ licensed voice actors — the consent-first counterpoint to community cloningfree trial · seat plansactive
LOVO / GennyLOVO2M-user creator studio; still shadowed by the Lehrman v. Lovo voice-actor consent lawsuit — the category's precedent casefree trial · sub plansactive
Sesame (CSM)Sesame AIOpen-sourced CSM-1B (Mar 2025) and the eerily-natural Maya demo; company has pivoted to a consumer personal agent + 2027 eyewear — voice model now a means, not the productfree previewactive
DiaNari LabsViral open 1.6B dialogue model (Apr 2025) with nonverbal tags; momentum cooled as Chatterbox and Fish absorbed the OSS mindsharefreefading
OpenVoiceMyShell (MIT)2024's instant-cloning breakout; little movement since — superseded by Chatterbox-class modelsfreefading
BarkSunoEarly generative-audio OSS star; unmaintained since Suno went all-in on musicfreedead
Tortoise-TTSneonbjb (James Betker)The 2022 OSS ancestor of the modern wave; author joined OpenAI, repo archived in spiritfreedead
RespeecherRespeecherFilm/TV-grade voice cloning with per-project consent contracts (Hollywood credits); services-heavy, not self-serve-firstproject pricing · marketplace plansactive
Camb.ai (MARS)Camb.aiDubbing/translation specialist (sports leagues, 140+ languages claimed) — overlaps localization more than product voicefreemiumactive
VUI Labs (Luna)VUI LabsNewcomer at #4 on the AA arena (Elo 1209, Aug 2026) — little public company footprint yet; watch or waitunverifiedactive

Voice: the category in numbers

Edition v2026.Q3 · ranking, pricing and status verified 2026-08-06.

  • ElevenLabs: ~$500M ARR by Apr 2026 (from $350M end-2025, +$100M in Q1 alone); 41% of Fortune 500; ~50/50 self-serve vs enterprise [src]
  • ElevenLabs $500M Series D at $11B led by Sequoia (Feb 4, 2026); May 2026 third close added BlackRock, NVIDIA, Salesforce, Santander [src]
  • Quality commoditized: 91 models on the AA Speech Arena; #1 is Alibaba's Qwen-Audio-3.0-TTS-Plus (Elo 1229) and ElevenLabs ranks #11 — the leader no longer has the best model (Aug 2026) [src]
  • Big-tech consolidation of voice talent: Meta acquired PlayAI (Jul 11, 2025, undisclosed); MiniMax listed on HKEX Jan 9, 2026 [src]
  • Hume's own thesis, Apr 15, 2026: 'Voice models are commoditizing. The value is moving up the stack' — vendors are repositioning around agents, emotion layers, and trust tooling [src]
  • The trust flank is monetizing: cloning pioneer Resemble now leads its pricing page with deepfake Detect and watermarking at per-second rates (Aug 2026) [src]

Voice: method & sources

Ranking charter: we weight production-workflow completeness, verified traction, and API economics over single-utterance arena Elo — blind arenas measure naturalness of short clips, not steerability, cloning quality, or dubbing workflow; that is how ElevenLabs stays #1 while ranking #11 on Elo, and we state the Elo plainly. Conflicts resolved: (1) ElevenLabs ARR — CEO said 'crossed $330M last year' (Jan 13, 2026, TechCrunch) vs Sacra's $350M end-2025 / $500M Apr 2026; we cite Sacra as the fuller series, consistent with the CEO figure. (2) Cartesia claims '#1 for naturalness'; the overall AA board has it #5 — we use the board. (3) The reported July 2026 ElevenLabs tender/secondary talks at ~$22B could not be traced to a fetchable primary source by press time; treat as unconfirmed and excluded from evidence. (4) OpenAI's gpt-4o-mini-tts no longer appears as a line item on the current pricing page; the ~$0.015/min figure is the launch-era estimate — re-verify before budgeting. Aggregator caveat: Sacra figures are single-source estimates. Adjacent elements: voice support/call agents (Vapi, Retell, Bland, ElevenAgents-as-agent-platform) → Vs; talking-head video → Av; meeting transcription → Mt; Eleven Music and Suno's music work have no element and are noted only as ElevenLabs platform sprawl. PlayHT's site remains up post-acquisition — don't build on it. Ranking criteria: verified commercial traction, independent satisfaction surveys, agent benchmarks, and founder-fit (price floor, lock-in, surfaces). Editorial, never paid — the charter. Machine-readable twin: vo.json.

All sources (29)
  1. https://elevenlabs.io/pricing
  2. https://elevenlabs.io/docs
  3. https://sacra.com/c/elevenlabs/
  4. https://www.cnbc.com/2026/02/04/nvidia-backed-ai-startup-elevenlabs-11-billion-valuation.html
  5. https://techcrunch.com/tag/elevenlabs/
  6. https://en.wikipedia.org/wiki/ElevenLabs
  7. https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice
  8. https://artificialanalysis.ai/text-to-speech
  9. https://cartesia.ai/pricing
  10. https://cartesia.ai/blog
  11. https://cartesia.ai/sonic
  12. https://docs.cartesia.ai/
  13. https://www.hume.ai/pricing
  14. https://www.hume.ai/blog
  15. https://dev.hume.ai/docs
  16. https://developers.openai.com/api/docs/pricing
  17. https://developers.openai.com/api/docs/guides/text-to-speech
  18. https://fish.audio/
  19. https://docs.fish.audio/
  20. https://github.com/resemble-ai/chatterbox
  21. https://github.com/hexgrad/kokoro
  22. https://www.resemble.ai/pricing/
  23. https://en.wikipedia.org/wiki/List_of_mergers_and_acquisitions_by_Meta_Platforms
  24. https://en.wikipedia.org/wiki/MiniMax_(company)
  25. https://deepgram.com/pricing
  26. https://www.inworld.ai/pricing
  27. https://murf.ai/pricing
  28. https://speechify.ai/
  29. https://www.wellsaid.io/

Our take

Voice quality crossed the uncanny valley in 2025. Narration, support lines, and product voices are now table stakes.

Combines with

Guides

Best AI voice generator (2026)

This is element 22 of 58. The table is versioned quarterly — when a tool loses its seat, the changelog records the succession.

Explore the full table →