{
 "sym": "Vo",
 "updated": "2026-08-06",
 "verdict": "ElevenLabs, still \u2014 but for platform and distribution, not raw quality. It owns the workflow (voices, cloning, dubbing, Studio, API), hit ~$500M ARR by April 2026, serves 41% of the Fortune 500, and raised a $500M Series D at $11B in February \u2014 yet its flagship Eleven v3 ranks only #11 on the blind Artificial Analysis arena, behind Alibaba, Speechify, and Cartesia. Cartesia wins when you're building real-time voice agents and every millisecond counts; OpenAI's gpt-4o-mini-tts wins when 'good enough everywhere' at ~$0.015/min beats best-in-class; Hume wins when you need a voice that acts, not reads; Fish Audio wins when you want open-weight economics with hosted convenience. Voice quality crossed the uncanny valley in 2025 \u2014 in 2026 the models are commoditizing and the moat moved to workflow, ecosystem, and trust.",
 "top5": [
  {
   "rank": 1,
   "name": "ElevenLabs",
   "maker": "ElevenLabs",
   "url": "https://elevenlabs.io",
   "docs": "https://elevenlabs.io/docs",
   "pricing": "Free 10k credits \u00b7 Starter $6/mo \u00b7 Creator $22 \u00b7 Pro $99 \u00b7 Scale $299 \u00b7 Business $990 \u00b7 Enterprise custom; TTS \u2248$0.17\u20130.36/min",
   "best_for": "Production narration, audiobooks, dubbing, and product voices where you want the full stack \u2014 voice library, pro cloning, Studio, 70+ languages \u2014 behind one API.",
   "why": "The category owner by revenue and reach: ~$500M ARR by April 2026 (from $350M at end-2025), 41% of the Fortune 500, a $500M Sequoia-led Series D at $11B (Feb 4, 2026), and distribution wins like Spotify's ElevenLabs-powered audiobook tool (May 2026). Eleven v3's audio tags ([whispers], [sighs]) and 70+ languages made expressive, directable speech mainstream.",
   "watch": "The quality crown is gone: Eleven v3 sits #11 on the Artificial Analysis blind arena (Elo 1171), behind Qwen, Speechify, Gemini, and Cartesia. Credit pricing gets expensive at scale against $0.65\u201325/1M-char rivals, and the platform sprawl \u2014 music, agents, 'Reception AI' \u2014 is a focus risk.",
   "evidence": [
    {
     "stat": "~$500M ARR by Apr 2026 (+$100M net-new in Q1); 41% of Fortune 500",
     "src": "https://sacra.com/c/elevenlabs/"
    },
    {
     "stat": "$500M Series D at $11B led by Sequoia (Feb 4, 2026)",
     "src": "https://www.cnbc.com/2026/02/04/nvidia-backed-ai-startup-elevenlabs-11-billion-valuation.html"
    },
    {
     "stat": "Eleven v3 ranks #11 (Elo 1171) on the AA Speech Arena (Aug 2026)",
     "src": "https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice"
    }
   ],
   "tile_note": "the category owner"
  },
  {
   "rank": 2,
   "name": "Cartesia (Sonic)",
   "maker": "Cartesia",
   "url": "https://cartesia.ai",
   "docs": "https://docs.cartesia.ai",
   "pricing": "Free 20k credits \u00b7 Pro $5/mo \u00b7 Startup $49 \u00b7 Scale $299 \u00b7 Enterprise custom; Line agent calls $0.06/min",
   "best_for": "Real-time voice \u2014 agents, IVR, live products \u2014 where sub-100ms latency and turn-taking accuracy matter more than a giant voice library.",
   "why": "The performance challenger: Sonic 3.5 is the top-ranked major API-first voice vendor on the blind arena (#5, Elo 1203 \u2014 above ElevenLabs), claims sub-90ms latency across 40+ languages, and pairs with Ink-2, the #1-ranked streaming STT for voice agents (Jul 2026). Built by the state-space-model (Mamba) researchers, so the speed claims have architectural teeth.",
   "watch": "A fraction of ElevenLabs' scale ($64M Series A, Mar 2025 \u2014 no larger round announced by Aug 2026). Marketing says '#1 for naturalness'; the overall arena board says #5. Voice library and creator tooling are thin \u2014 it's a developer product, not a studio.",
   "evidence": [
    {
     "stat": "Sonic 3.5: #5 overall, Elo 1203 \u2014 highest-ranked major voice API (Aug 2026)",
     "src": "https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice"
    },
    {
     "stat": "Ink-2 STT ranked #1 on AA streaming leaderboard (Jul 9, 2026)",
     "src": "https://cartesia.ai/blog"
    },
    {
     "stat": "$64M Series A led by Kleiner Perkins (Mar 11, 2025)",
     "src": "https://cartesia.ai/blog"
    }
   ],
   "tile_note": "real-time speed king"
  },
  {
   "rank": 3,
   "name": "OpenAI TTS / Realtime",
   "maker": "OpenAI",
   "url": "https://openai.com/api/",
   "docs": "https://developers.openai.com/api/docs/guides/text-to-speech",
   "pricing": "API only: gpt-4o-mini-tts \u2248$0.015/min (est.) \u00b7 gpt-realtime-2.1-mini $10/1M audio-in, $20/1M audio-out \u00b7 full gpt-realtime-2.1 $32/$64",
   "best_for": "Teams already on the OpenAI stack that want instructable, good-enough voice \u2014 13 voices, steerable tone \u2014 with zero new vendors and near-zero cost.",
   "why": "The default, not the best: gpt-4o-mini-tts lets you steer delivery in plain English ('sound like a sympathetic agent') at commodity prices, and the Realtime API (gpt-realtime-2.1) handles full speech-to-speech for voice products. Distribution is the moat \u2014 every OpenAI developer has it one endpoint away.",
   "watch": "No voice cloning at all (a deliberate safety stance) and no custom voice design; quality ranks #30 on the blind arena (TTS-1 HD, Elo 1097). Dedicated TTS pricing has been folded into the audio/realtime page \u2014 line-item costs are harder to pin down than a year ago.",
   "evidence": [
    {
     "stat": "gpt-4o-mini-tts current with 13 voices; tts-1/tts-1-hd now legacy (Aug 2026)",
     "src": "https://developers.openai.com/api/docs/guides/text-to-speech"
    },
    {
     "stat": "gpt-realtime-2.1-mini: $10/1M audio input, $20/1M audio output (Aug 2026)",
     "src": "https://developers.openai.com/api/docs/pricing"
    },
    {
     "stat": "OpenAI TTS-1 HD ranks #30 (Elo 1097) on AA arena (Aug 2026)",
     "src": "https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice"
    }
   ],
   "tile_note": "good enough, everywhere"
  },
  {
   "rank": 4,
   "name": "Hume (Octave / EVI)",
   "maker": "Hume AI",
   "url": "https://www.hume.ai",
   "docs": "https://dev.hume.ai/docs",
   "pricing": "Free \u00b7 Starter $3/mo \u00b7 Creator $14 \u00b7 Pro $70 \u00b7 Scale $200 \u00b7 Business $500 \u00b7 Enterprise; TTS overage $0.05\u20130.15/1k chars, EVI $0.04\u20130.07/min",
   "best_for": "Voice as performance \u2014 Octave 'acts' a line from context and direction rather than reading it, and EVI adds emotionally-aware real-time speech-to-speech.",
   "why": "The emotional-intelligence specialist: Octave generates delivery from meaning (frustration, sarcasm, comfort) with director-style prompts, EVI is one of the few production speech-to-speech APIs, and the team publishes real research (RW-Voice-EQ benchmark for the human quality of voice AI, Jul 2026). Cheapest serious entry point in the category at $3/mo.",
   "watch": "The emotion thesis hasn't won blind listening: Octave 2 ranks #51 (Elo 1052) on the AA arena. Hume's own April 2026 essay concedes 'voice models are commoditizing' \u2014 its value case now rests on the empathic layer, a much narrower moat than a platform.",
   "evidence": [
    {
     "stat": "Octave 2 ranks #51 (Elo 1052) on AA arena (Aug 2026)",
     "src": "https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice"
    },
    {
     "stat": "Published RW-Voice-EQ bench for voice-AI human quality (Jul 2026)",
     "src": "https://www.hume.ai/blog"
    },
    {
     "stat": "TTS from $3/mo; EVI overage down to $0.04/min at Business tier (Aug 2026)",
     "src": "https://www.hume.ai/pricing"
    }
   ],
   "tile_note": "emotion as an api"
  },
  {
   "rank": 5,
   "name": "Fish Audio (OpenAudio)",
   "maker": "Hanabi AI",
   "url": "https://fish.audio",
   "docs": "https://docs.fish.audio",
   "pricing": "Freemium studio \u00b7 S2.1 Pro free developer tier \u00b7 paid plans + usage-based API \u00b7 commercial use on paid plans",
   "best_for": "Cost-sensitive builders who want near-frontier quality with an open-weight lineage \u2014 and a 2M-voice community library to draw from.",
   "why": "The open-ecosystem value play: Fish Speech \u2192 S1 \u2192 S2 shipped as open models, S2.1 Pro ranks #15 on the blind arena (Elo 1138) \u2014 above OpenAI and Hume \u2014 and the company made it free for developers by rebuilding its inference stack. $52M seed, 8M+ builders, 30+ languages, 2M+ community voices.",
   "watch": "Seed-stage company carrying real trust-and-safety surface: a 2M-voice community library makes consent provenance hard to police. Public pricing is opaque next to rivals, and enterprise compliance machinery (SOC 2, BAAs) trails the leaders.",
   "evidence": [
    {
     "stat": "$52M seed; 8M+ builders; S2.1 Pro made free for developers (2026)",
     "src": "https://fish.audio/"
    },
    {
     "stat": "Fish Audio S2.1 Pro ranks #15 (Elo 1138), above OpenAI and Hume (Aug 2026)",
     "src": "https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice"
    },
    {
     "stat": "Open models S1/S2 published; 30+ languages, 2M+ voices (Aug 2026)",
     "src": "https://docs.fish.audio/"
    }
   ],
   "tile_note": "open-weights value play"
  }
 ],
 "matrix": {
  "cols": [
   "Entry price",
   "Typical API cost",
   "Arena Elo (AA, Aug '26)",
   "Latency claim",
   "Voice cloning",
   "Languages",
   "Open weights"
  ],
  "rows": [
   [
    "ElevenLabs",
    "$6/mo",
    "\u2248$0.17\u20130.36/min (credits)",
    "1171 \u00b7 #11",
    "low-lat. mode $0.05/min (Business)",
    "Instant + professional",
    "70+",
    "No"
   ],
   [
    "Cartesia Sonic 3.5",
    "$5/mo",
    "credits; agents $0.06/min",
    "1203 \u00b7 #5",
    "<90 ms (claimed)",
    "Instant (10s) + pro",
    "40+",
    "No"
   ],
   [
    "OpenAI gpt-4o-mini-tts",
    "usage only",
    "\u2248$0.015/min (est.)",
    "1097 \u00b7 #30 (TTS-1 HD)",
    "Realtime API for live",
    "None (policy)",
    "Multilingual",
    "No"
   ],
   [
    "Hume Octave 2",
    "$3/mo",
    "$0.05\u20130.15/1k chars",
    "1052 \u00b7 #51",
    "EVI real-time",
    "Instant",
    "Multilingual",
    "No"
   ],
   [
    "Fish Audio S2.1 Pro",
    "free",
    "usage API; free dev tier",
    "1138 \u00b7 #15",
    "Real-time streaming",
    "Community + instant",
    "30+",
    "Yes (S1/S2)"
   ],
   [
    "MiniMax Speech 2.8",
    "usage",
    "per-char API",
    "1172 \u00b7 #10 (HD)",
    "Streaming",
    "Instant",
    "30+",
    "No"
   ],
   [
    "Chatterbox V3",
    "free (MIT)",
    "self-host",
    "n/a (not listed)",
    "Self-host dependent",
    "Zero-shot",
    "23+",
    "Yes"
   ],
   [
    "Kokoro-82M",
    "free (Apache)",
    "$0.65/1M chars (hosted)",
    "n/a (not listed)",
    "Fast, runs on CPU",
    "No",
    "9",
    "Yes (82M)"
   ]
  ]
 },
 "rules": [
  {
   "if": "You're producing narration, audiobooks, dubbing, or a branded product voice and want one vendor for the whole workflow",
   "then": "ElevenLabs \u2014 the Studio + 70-language + pro-cloning stack is unmatched, and 41% of the Fortune 500 already cleared it through procurement."
  },
  {
   "if": "You're building a real-time voice agent and the latency budget is under 100ms",
   "then": "Cartesia \u2014 Sonic 3.5 (<90ms claimed, #5 arena Elo) plus Ink-2's #1-ranked turn detection; hand the agent orchestration layer to element Vs."
  },
  {
   "if": "You're already on OpenAI and voice is a feature, not the product",
   "then": "gpt-4o-mini-tts at ~$0.015/min \u2014 accept no cloning and mid-pack quality for zero new vendors and one-line integration."
  },
  {
   "if": "The line has to be performed \u2014 empathy, sarcasm, de-escalation \u2014 not just read",
   "then": "Hume Octave (or ElevenLabs v3 audio tags); judge by ear on your script, because blind-arena Elo says the emotion premium is not free."
  },
  {
   "if": "Volume is huge, margins are thin, or audio must stay on your infra",
   "then": "Open weights: Chatterbox V3 (MIT, 23+ languages) for cloning, Kokoro-82M at $0.65/1M chars for cheap clean narration, Fish's S1/S2 for quality self-hosting."
  }
 ],
 "field": [
  {
   "name": "PlayHT / PlayAI",
   "maker": "PlayAI \u2192 Meta",
   "note": "Acquired by Meta Jul 11, 2025 (Bloomberg; undisclosed) \u2014 team absorbed into Meta's voice efforts; play.ht product still online but effectively in maintenance",
   "url": "https://play.ht",
   "oss": false,
   "entry": "freemium",
   "status": "acquired"
  },
  {
   "name": "Resemble AI",
   "maker": "Resemble AI",
   "note": "Cloning pioneer now leading with deepfake Detect + watermarking (per-second pricing); open-sourced Chatterbox \u2014 a telling category pivot from generation to trust",
   "url": "https://www.resemble.ai",
   "oss": false,
   "entry": "Flex $0/mo pay-go \u00b7 Team $350/mo",
   "status": "active"
  },
  {
   "name": "Chatterbox",
   "maker": "Resemble AI (MIT)",
   "note": "Leading open TTS/cloning family \u2014 25.9k stars, Multilingual V3 (23+ languages), [laugh]-style tags, Turbo/Nano variants for CPU",
   "url": "https://github.com/resemble-ai/chatterbox",
   "oss": true,
   "entry": "free",
   "status": "active"
  },
  {
   "name": "Kokoro-82M",
   "maker": "hexgrad (Apache-2.0)",
   "note": "82M-param open model; cheapest serious TTS on the AA index at $0.65/1M chars hosted \u2014 the commodity floor for clean narration",
   "url": "https://github.com/hexgrad/kokoro",
   "oss": true,
   "entry": "free",
   "status": "active"
  },
  {
   "name": "MiniMax Audio (Speech 2.8)",
   "maker": "MiniMax",
   "note": "Top-10 arena quality (HD Elo 1172); listed on HKEX Jan 9, 2026 \u2014 but carries Disney/Universal/WBD copyright suit (Sep 2025) and an Anthropic fraud accusation (Feb 2026)",
   "url": "https://www.minimax.io/audio",
   "oss": false,
   "entry": "usage-based",
   "status": "active"
  },
  {
   "name": "Speechify (Simba 3.2)",
   "maker": "Speechify AI",
   "note": "Consumer reading app turned API vendor \u2014 Simba 3.2 ranks #2 on the blind arena (Elo 1227), sub-100ms streaming",
   "url": "https://speechify.ai",
   "oss": false,
   "entry": "free tier \u00b7 API usage",
   "status": "active"
  },
  {
   "name": "Qwen-Audio-3.0-TTS-Plus",
   "maker": "Alibaba Cloud",
   "note": "#1 on the AA arena (Elo 1229, Aug 2026) \u2014 proof that quality leadership has left the pure-play vendors",
   "url": "https://www.alibabacloud.com/en/product/modelstudio",
   "oss": false,
   "entry": "usage-based",
   "status": "active"
  },
  {
   "name": "Gemini TTS (3.1 Flash)",
   "maker": "Google",
   "note": "#3 on the arena (Elo 1210), bundled into the Gemini API \u2014 the big-cloud default for GCP shops",
   "url": "https://ai.google.dev",
   "oss": false,
   "entry": "usage-based",
   "status": "active"
  },
  {
   "name": "Azure AI Speech",
   "maker": "Microsoft",
   "note": "Enterprise incumbent: custom neural voice with strict consent gates, compliance depth over quality rank",
   "url": "https://azure.microsoft.com/en-us/products/ai-services/ai-speech",
   "oss": false,
   "entry": "usage-based + free tier",
   "status": "active"
  },
  {
   "name": "Amazon Polly",
   "maker": "AWS",
   "note": "The 2016-era workhorse \u2014 fastest on AA's index (1,149 chars/sec) and cheap, but generative-voice quality has passed it by",
   "url": "https://aws.amazon.com/polly/",
   "oss": false,
   "entry": "usage-based + free tier",
   "status": "fading"
  },
  {
   "name": "Deepgram Aura-2",
   "maker": "Deepgram",
   "note": "STT leader's agent-focused TTS at $0.030/1k chars \u2014 bought for the bundle, not the voice",
   "url": "https://deepgram.com",
   "oss": false,
   "entry": "$0.030/1k chars pay-go",
   "status": "active"
  },
  {
   "name": "Inworld TTS",
   "maker": "Inworld AI",
   "note": "Gaming-heritage vendor with two models in the arena top 10; aggressive $5\u201325/1M-char pricing ladder",
   "url": "https://www.inworld.ai",
   "oss": false,
   "entry": "$15\u201335/1M chars on-demand",
   "status": "active"
  },
  {
   "name": "Smallest.ai (Lightning)",
   "maker": "Smallest.ai",
   "note": "India-based real-time specialist; Lightning V3.1 Pro ranks #8 on the arena (Elo 1192)",
   "url": "https://smallest.ai",
   "oss": false,
   "entry": "usage-based",
   "status": "active"
  },
  {
   "name": "Murf AI",
   "maker": "Murf",
   "note": "SMB studio standard for e-learning/corporate VO; Falcon 2 sits on AA's price-quality frontier, but flagship studio voices rank #74 (Elo 975)",
   "url": "https://murf.ai",
   "oss": false,
   "entry": "free \u00b7 Creator from $19/mo (annual)",
   "status": "active"
  },
  {
   "name": "WellSaid",
   "maker": "WellSaid (now wellsaid.io)",
   "note": "Enterprise L&D voice-over with 280+ licensed voice actors \u2014 the consent-first counterpoint to community cloning",
   "url": "https://www.wellsaid.io",
   "oss": false,
   "entry": "free trial \u00b7 seat plans",
   "status": "active"
  },
  {
   "name": "LOVO / Genny",
   "maker": "LOVO",
   "note": "2M-user creator studio; still shadowed by the Lehrman v. Lovo voice-actor consent lawsuit \u2014 the category's precedent case",
   "url": "https://lovo.ai",
   "oss": false,
   "entry": "free trial \u00b7 sub plans",
   "status": "active"
  },
  {
   "name": "Sesame (CSM)",
   "maker": "Sesame AI",
   "note": "Open-sourced CSM-1B (Mar 2025) and the eerily-natural Maya demo; company has pivoted to a consumer personal agent + 2027 eyewear \u2014 voice model now a means, not the product",
   "url": "https://www.sesame.com",
   "oss": true,
   "entry": "free preview",
   "status": "active"
  },
  {
   "name": "Dia",
   "maker": "Nari Labs",
   "note": "Viral open 1.6B dialogue model (Apr 2025) with nonverbal tags; momentum cooled as Chatterbox and Fish absorbed the OSS mindshare",
   "url": "https://github.com/nari-labs/dia",
   "oss": true,
   "entry": "free",
   "status": "fading"
  },
  {
   "name": "OpenVoice",
   "maker": "MyShell (MIT)",
   "note": "2024's instant-cloning breakout; little movement since \u2014 superseded by Chatterbox-class models",
   "url": "https://github.com/myshell-ai/OpenVoice",
   "oss": true,
   "entry": "free",
   "status": "fading"
  },
  {
   "name": "Bark",
   "maker": "Suno",
   "note": "Early generative-audio OSS star; unmaintained since Suno went all-in on music",
   "url": "https://github.com/suno-ai/bark",
   "oss": true,
   "entry": "free",
   "status": "dead"
  },
  {
   "name": "Tortoise-TTS",
   "maker": "neonbjb (James Betker)",
   "note": "The 2022 OSS ancestor of the modern wave; author joined OpenAI, repo archived in spirit",
   "url": "https://github.com/neonbjb/tortoise-tts",
   "oss": true,
   "entry": "free",
   "status": "dead"
  },
  {
   "name": "Respeecher",
   "maker": "Respeecher",
   "note": "Film/TV-grade voice cloning with per-project consent contracts (Hollywood credits); services-heavy, not self-serve-first",
   "url": "https://www.respeecher.com",
   "oss": false,
   "entry": "project pricing \u00b7 marketplace plans",
   "status": "active"
  },
  {
   "name": "Camb.ai (MARS)",
   "maker": "Camb.ai",
   "note": "Dubbing/translation specialist (sports leagues, 140+ languages claimed) \u2014 overlaps localization more than product voice",
   "url": "https://camb.ai",
   "oss": false,
   "entry": "freemium",
   "status": "active"
  },
  {
   "name": "VUI Labs (Luna)",
   "maker": "VUI Labs",
   "note": "Newcomer at #4 on the AA arena (Elo 1209, Aug 2026) \u2014 little public company footprint yet; watch or wait",
   "url": "https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice",
   "oss": false,
   "entry": "unverified",
   "status": "active"
  }
 ],
 "signals": [
  {
   "fact": "ElevenLabs: ~$500M ARR by Apr 2026 (from $350M end-2025, +$100M in Q1 alone); 41% of Fortune 500; ~50/50 self-serve vs enterprise",
   "src": "https://sacra.com/c/elevenlabs/"
  },
  {
   "fact": "ElevenLabs $500M Series D at $11B led by Sequoia (Feb 4, 2026); May 2026 third close added BlackRock, NVIDIA, Salesforce, Santander",
   "src": "https://www.cnbc.com/2026/02/04/nvidia-backed-ai-startup-elevenlabs-11-billion-valuation.html"
  },
  {
   "fact": "Quality commoditized: 91 models on the AA Speech Arena; #1 is Alibaba's Qwen-Audio-3.0-TTS-Plus (Elo 1229) and ElevenLabs ranks #11 \u2014 the leader no longer has the best model (Aug 2026)",
   "src": "https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice"
  },
  {
   "fact": "Big-tech consolidation of voice talent: Meta acquired PlayAI (Jul 11, 2025, undisclosed); MiniMax listed on HKEX Jan 9, 2026",
   "src": "https://en.wikipedia.org/wiki/List_of_mergers_and_acquisitions_by_Meta_Platforms"
  },
  {
   "fact": "Hume's own thesis, Apr 15, 2026: 'Voice models are commoditizing. The value is moving up the stack' \u2014 vendors are repositioning around agents, emotion layers, and trust tooling",
   "src": "https://www.hume.ai/blog"
  },
  {
   "fact": "The trust flank is monetizing: cloning pioneer Resemble now leads its pricing page with deepfake Detect and watermarking at per-second rates (Aug 2026)",
   "src": "https://www.resemble.ai/pricing/"
  }
 ],
 "notes": "Ranking charter: we weight production-workflow completeness, verified traction, and API economics over single-utterance arena Elo \u2014 blind arenas measure naturalness of short clips, not steerability, cloning quality, or dubbing workflow; that is how ElevenLabs stays #1 while ranking #11 on Elo, and we state the Elo plainly. Conflicts resolved: (1) ElevenLabs ARR \u2014 CEO said 'crossed $330M last year' (Jan 13, 2026, TechCrunch) vs Sacra's $350M end-2025 / $500M Apr 2026; we cite Sacra as the fuller series, consistent with the CEO figure. (2) Cartesia claims '#1 for naturalness'; the overall AA board has it #5 \u2014 we use the board. (3) The reported July 2026 ElevenLabs tender/secondary talks at ~$22B could not be traced to a fetchable primary source by press time; treat as unconfirmed and excluded from evidence. (4) OpenAI's gpt-4o-mini-tts no longer appears as a line item on the current pricing page; the ~$0.015/min figure is the launch-era estimate \u2014 re-verify before budgeting. Aggregator caveat: Sacra figures are single-source estimates. Adjacent elements: voice support/call agents (Vapi, Retell, Bland, ElevenAgents-as-agent-platform) \u2192 Vs; talking-head video \u2192 Av; meeting transcription \u2192 Mt; Eleven Music and Suno's music work have no element and are noted only as ElevenLabs platform sprawl. PlayHT's site remains up post-acquisition \u2014 don't build on it.",
 "sources": [
  "https://elevenlabs.io/pricing",
  "https://elevenlabs.io/docs",
  "https://sacra.com/c/elevenlabs/",
  "https://www.cnbc.com/2026/02/04/nvidia-backed-ai-startup-elevenlabs-11-billion-valuation.html",
  "https://techcrunch.com/tag/elevenlabs/",
  "https://en.wikipedia.org/wiki/ElevenLabs",
  "https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice",
  "https://artificialanalysis.ai/text-to-speech",
  "https://cartesia.ai/pricing",
  "https://cartesia.ai/blog",
  "https://cartesia.ai/sonic",
  "https://docs.cartesia.ai/",
  "https://www.hume.ai/pricing",
  "https://www.hume.ai/blog",
  "https://dev.hume.ai/docs",
  "https://developers.openai.com/api/docs/pricing",
  "https://developers.openai.com/api/docs/guides/text-to-speech",
  "https://fish.audio/",
  "https://docs.fish.audio/",
  "https://github.com/resemble-ai/chatterbox",
  "https://github.com/hexgrad/kokoro",
  "https://www.resemble.ai/pricing/",
  "https://en.wikipedia.org/wiki/List_of_mergers_and_acquisitions_by_Meta_Platforms",
  "https://en.wikipedia.org/wiki/MiniMax_(company)",
  "https://deepgram.com/pricing",
  "https://www.inworld.ai/pricing",
  "https://murf.ai/pricing",
  "https://speechify.ai/",
  "https://www.wellsaid.io/"
 ],
 "element": {
  "number": 22,
  "name": "Voice",
  "group": "Creative & Design",
  "essential": false,
  "edition": "v2026.Q3",
  "revision": "r7",
  "license": "CC BY 4.0 \u2014 cite elems.ai",
  "url": "https://elems.ai/e/vo.html"
 }
}