# Vo · Voice — element 22 of 58

> Sound like anyone. Ethically, you. Turns text into production-ready audio.

- **Group:** 5 · Creative & Design
- **Necessity:** Optional
- **Price band:** $ · under $30/mo
- **Maturity:** Stable
- **Edition:** v2026.Q3 · verified 2026-09-13

## Leading tools (v2026.Q3)

- **ElevenLabs** — the category owner
- **Cartesia (Sonic)** — real-time speed king
- **OpenAI TTS / Realtime** — good enough, everywhere
- **Hume (Octave / EVI)** — emotion as an api
- **Fish Audio (OpenAudio)** — open-weights value play

## Our take

Voice quality crossed the uncanny valley in 2025. Narration, support lines, and product voices are now table stakes.

## Combines with

Av, Vs
- Guide: [Best AI voice generator (2026)](https://elems.ai/best/best-ai-voice-generator.html)

## The top 5 — deep dossier (verified 2026-09-13)

ElevenLabs, still — but for platform and distribution, not raw quality. It owns the workflow (voices, cloning, dubbing, Studio, API), hit ~$500M ARR by April 2026, serves 41% of the Fortune 500, and raised a $500M Series D at $11B in February — yet its flagship Eleven v3 ranks only #11 on the blind Artificial Analysis arena, behind Alibaba, Speechify, and Cartesia. Cartesia wins when you're building real-time voice agents and every millisecond counts; OpenAI's gpt-4o-mini-tts wins when 'good enough everywhere' at ~$0.015/min beats best-in-class; Hume wins when you need a voice that acts, not reads; Fish Audio wins when you want open-weight economics with hosted convenience. Voice quality crossed the uncanny valley in 2025 — in 2026 the models are commoditizing and the moat moved to workflow, ecosystem, and trust.

1. **ElevenLabs** (ElevenLabs) — Free 10k credits · Starter $6/mo · Creator $22 · Pro $99 · Scale $299 · Business $990 · Enterprise custom; TTS ≈$0.17–0.36/min. Best for: Production narration, audiobooks, dubbing, and product voices where you want the full stack — voice library, pro cloning, Studio, 70+ languages — behind one API. Why: The category owner by revenue and reach: ~$500M ARR by April 2026 (from $350M at end-2025), 41% of the Fortune 500, a $500M Sequoia-led Series D at $11B (Feb 4, 2026), and distribution wins like Spotify's ElevenLabs-powered audiobook tool (May 2026). Eleven v3's audio tags ([whispers], [sighs]) and 70+ languages made expressive, directable speech mainstream. Watch: The quality crown is gone: Eleven v3 sits #11 on the Artificial Analysis blind arena (Elo 1171), behind Qwen, Speechify, Gemini, and Cartesia. Credit pricing gets expensive at scale against $0.65–25/1M-char rivals, and the platform sprawl — music, agents, 'Reception AI' — is a focus risk. [https://elevenlabs.io](https://elevenlabs.io)
2. **Cartesia (Sonic)** (Cartesia) — Free 20k credits · Pro $5/mo · Startup $49 · Scale $299 · Enterprise custom; Line agent calls $0.06/min. Best for: Real-time voice — agents, IVR, live products — where sub-100ms latency and turn-taking accuracy matter more than a giant voice library. Why: The performance challenger: Sonic 3.5 is the top-ranked major API-first voice vendor on the blind arena (#5, Elo 1203 — above ElevenLabs), claims sub-90ms latency across 40+ languages, and pairs with Ink-2, the #1-ranked streaming STT for voice agents (Jul 2026). Built by the state-space-model (Mamba) researchers, so the speed claims have architectural teeth. Watch: A fraction of ElevenLabs' scale ($64M Series A, Mar 2025 — no larger round announced by Aug 2026). Marketing says '#1 for naturalness'; the overall arena board says #5. Voice library and creator tooling are thin — it's a developer product, not a studio. [https://cartesia.ai](https://cartesia.ai)
3. **OpenAI TTS / Realtime** (OpenAI) — API only: gpt-4o-mini-tts ≈$0.015/min (est.) · gpt-realtime-2.1-mini $10/1M audio-in, $20/1M audio-out · full gpt-realtime-2.1 $32/$64. Best for: Teams already on the OpenAI stack that want instructable, good-enough voice — 13 voices, steerable tone — with zero new vendors and near-zero cost. Why: The default, not the best: gpt-4o-mini-tts lets you steer delivery in plain English ('sound like a sympathetic agent') at commodity prices, and the Realtime API (gpt-realtime-2.1) handles full speech-to-speech for voice products. Distribution is the moat — every OpenAI developer has it one endpoint away. Watch: No voice cloning at all (a deliberate safety stance) and no custom voice design; quality ranks #30 on the blind arena (TTS-1 HD, Elo 1097). Dedicated TTS pricing has been folded into the audio/realtime page — line-item costs are harder to pin down than a year ago. [https://openai.com/api/](https://openai.com/api/)
4. **Hume (Octave / EVI)** (Hume AI) — Free · Starter $3/mo · Creator $14 · Pro $70 · Scale $200 · Business $500 · Enterprise; TTS overage $0.05–0.15/1k chars, EVI $0.04–0.07/min. Best for: Voice as performance — Octave 'acts' a line from context and direction rather than reading it, and EVI adds emotionally-aware real-time speech-to-speech. Why: The emotional-intelligence specialist: Octave generates delivery from meaning (frustration, sarcasm, comfort) with director-style prompts, EVI is one of the few production speech-to-speech APIs, and the team publishes real research (RW-Voice-EQ benchmark for the human quality of voice AI, Jul 2026). Cheapest serious entry point in the category at $3/mo. Watch: The emotion thesis hasn't won blind listening: Octave 2 ranks #51 (Elo 1052) on the AA arena. Hume's own April 2026 essay concedes 'voice models are commoditizing' — its value case now rests on the empathic layer, a much narrower moat than a platform. [https://www.hume.ai](https://www.hume.ai)
5. **Fish Audio (OpenAudio)** (Hanabi AI) — Freemium studio · S2.1 Pro free developer tier · paid plans + usage-based API · commercial use on paid plans. Best for: Cost-sensitive builders who want near-frontier quality with an open-weight lineage — and a 2M-voice community library to draw from. Why: The open-ecosystem value play: Fish Speech → S1 → S2 shipped as open models, S2.1 Pro ranks #15 on the blind arena (Elo 1138) — above OpenAI and Hume — and the company made it free for developers by rebuilding its inference stack. $52M seed, 8M+ builders, 30+ languages, 2M+ community voices. Watch: Seed-stage company carrying real trust-and-safety surface: a 2M-voice community library makes consent provenance hard to police. Public pricing is opaque next to rivals, and enterprise compliance machinery (SOC 2, BAAs) trails the leaders. [https://fish.audio](https://fish.audio)

### How to choose
- If You're producing narration, audiobooks, dubbing, or a branded product voice and want one vendor for the whole workflow → ElevenLabs — the Studio + 70-language + pro-cloning stack is unmatched, and 41% of the Fortune 500 already cleared it through procurement.
- If You're building a real-time voice agent and the latency budget is under 100ms → Cartesia — Sonic 3.5 (<90ms claimed, #5 arena Elo) plus Ink-2's #1-ranked turn detection; hand the agent orchestration layer to element Vs.
- If You're already on OpenAI and voice is a feature, not the product → gpt-4o-mini-tts at ~$0.015/min — accept no cloning and mid-pack quality for zero new vendors and one-line integration.
- If The line has to be performed — empathy, sarcasm, de-escalation — not just read → Hume Octave (or ElevenLabs v3 audio tags); judge by ear on your script, because blind-arena Elo says the emotion premium is not free.
- If Volume is huge, margins are thin, or audio must stay on your infra → Open weights: Chatterbox V3 (MIT, 23+ languages) for cloning, Kokoro-82M at $0.65/1M chars for cheap clean narration, Fish's S1/S2 for quality self-hosting.

### The field (24 more)

PlayHT / PlayAI (fading), Resemble AI, Chatterbox, Kokoro-82M, MiniMax Audio (Speech 2.8), Speechify (Simba 3.2), Qwen-Audio-3.0-TTS-Plus, Gemini TTS (3.1 Flash), Azure AI Speech, Amazon Polly (fading), Deepgram Aura-2, Inworld TTS, Smallest.ai (Lightning), Murf AI, WellSaid, LOVO / Genny, Sesame (CSM), Dia (fading), OpenVoice (fading), Bark (dead), Tortoise-TTS (dead), Respeecher, Camb.ai (MARS), VUI Labs (Luna)

Full dossier data: https://elems.ai/e/vo.json

---
Source: [elems.ai](https://elems.ai/e/vo.html) — the periodic table of the AI-led startup. Data: https://elems.ai/elements.json (CC BY 4.0, cite elems.ai).
