AI Voice

ElevenLabs Alternatives (2026): Pick the Right Layer, Not Just a Cheaper Voice

By Urvil Dhanani · Jul 24, 2026 · 11 min read

Urvil Dhanani
Urvil Dhanani
Jul 24, 202611 min read
Diagram separating the TTS engine layer from the conversational AI platform layer

The best ElevenLabs alternative depends on which ElevenLabs you’re replacing. For text-to-speech, Cartesia leads on latency (~40–90ms), Deepgram Aura-2 on enterprise compliance, and Fish Audio or Kokoro if you want open-source. For ElevenAgents — the voice-agent platform — the real alternatives are conversational AI platforms like SuperMIA, Vapi, Retell, or Bland, because you’re buying telephony, CRM, and compliance, not a voice.

See a voice agent answer a live call — book a 15-minute demo →

Key takeaways

  • ‘ElevenLabs’ is two products. A TTS engine priced per character, and ElevenAgents priced per minute. Pick your layer before your vendor.
  • For TTS, cheaper and faster both exist. Cartesia (~40–90ms), Deepgram Aura-2 ($30/1M, HIPAA + on-prem), OpenAI tts-1 ($15/1M), Fish Audio (~$15/1M).
  • Free and open-source is genuinely viable now. Chatterbox, Fish Speech, Kokoro, XTTS v2 — you trade money for GPU time.
  • ElevenAgents’ headline rate isn’t the invoice. Their own pricing page bills LLM and telephony separately, caps concurrency, and doubles the rate on burst.
  • Running a business phone line is a platform purchase. That’s a different thing from buying a voice.

First, which ElevenLabs are you replacing?

This is the question almost every ‘ElevenLabs alternatives’ article skips, and getting it wrong wastes weeks. ElevenLabs sells two fundamentally different products, at two different layers of the stack, on two different pricing models.

Diagram separating the TTS engine layer from the conversational AI platform layer
‘ElevenLabs’ is two products — know which one you’re replacing.

Layer 1 — the voice itself (ElevenCreative / the TTS API). You send text, you get audio. Priced per character. This is what you use for audiobooks, YouTube voiceovers, dubbing, podcasts, and in-app narration. Competitors here are Cartesia, Deepgram, OpenAI, Fish Audio, Google, AWS, Azure, Murf, PlayHT.

Layer 2 — the agent (ElevenAgents). A platform for building voice agents that hold live phone conversations. Priced per minute. Competitors here are conversational AI platforms: Vapi, Retell, Bland, PolyAI, SuperMIA. Notably, platforms at this layer often consume a TTS engine underneath — sometimes ElevenLabs’ own.

The 10-second test: Are you generating audio from a script you already have? → You’re shopping at Layer 1 (TTS). Are you trying to hold a two-way conversation with a caller? → You’re shopping at Layer 2 (platform). Different products, different vendors, different budgets.

Why teams look for an ElevenLabs alternative

ElevenLabs is very good. The reasons people leave are usually structural, not quality-related. Across the buyer conversations and public comparisons, five reasons dominate:

  • Cost at volume. ElevenLabs’ top-tier v3 model is around $100 per million characters on published rates — the premium end of the market. At scale, that gap compounds.
  • The free tier has no commercial rights. ElevenLabs’ free tier is 10,000 characters per month and cannot be used commercially, so any business use starts on a paid plan.
  • Latency for real-time agents. For live conversation, time-to-first-audio is the constraint. Purpose-built streaming engines target lower TTFA than general-purpose narration models.
  • The agent layer needs assembly. ElevenAgents gives you the agent runtime — but you still connect the LLM, telephony, CRM, and compliance yourself.
  • Compliance gating. On ElevenLabs’ published ElevenAgents pricing, BAAs for HIPAA customers appear on the Enterprise tier. If you’re in healthcare, that shapes the decision early.

Best ElevenLabs alternatives for text-to-speech

If you’re replacing the voice engine, the field is genuinely strong — and mostly cheaper. Here are the published list rates. Prices move; verify on each vendor’s pricing page before you commit.

Bar chart comparing published TTS prices per million characters across nine providers
TTS cost per 1M characters (published rates, verified June–July 2026).
ProviderPublished price / 1M charsLatency (TTFA)Voice cloningBest for
Cartesia (Sonic-3)~$38–$50~40–90 msYes (instant)Real-time voice agents — the latency leader
Deepgram (Aura-2)$30~90 ms optimizedNoRegulated industries — SOC 2, HIPAA, on-prem
OpenAI (tts-1)$15Not publishedNoTeams already building on OpenAI
OpenAI (tts-1-hd)$30Not publishedNoHigher-fidelity batch narration
Fish Audio (S2 Pro)~$15LowYesQuality per dollar; open-weights, self-hostable
AWS Polly (Neural)$16Cloud-typicalNoAWS-native stacks
AWS Polly / Google (Standard)~$4Cloud-typicalNoHighest-volume, cost-first workloads
Murf / PlayHTSubscription-basedCloud-typicalYesCreator and video workflows, studio UI
ElevenLabs (v3)~$100~75 ms (Flash)YesReference point — the expressiveness ceiling

Cartesia — best for real-time latency

Built on state-space models rather than transformers, and purpose-built for live conversation. Independent comparisons consistently put it at the front of the field on time-to-first-audio, which is the metric that decides whether a phone conversation feels natural or broken. If you’re building a real-time agent and latency is your constraint, start here. Check current rates on Cartesia’s pricing page.

Deepgram Aura-2 — best for regulated industries

$0.030 per 1,000 characters, with SOC 2, HIPAA and — crucially — on-premises deployment. It’s trained on call-centre audio, so it handles clinical, financial and legal terminology without markup wrangling. The trade-offs are honest ones: fewer languages than ElevenLabs, and no voice cloning. If you’re in healthcare or finance and your blocker is compliance rather than expressiveness, Deepgram Aura-2 is usually the default answer.

OpenAI TTS — best if you’re already on OpenAI

tts-1 at $15 per million characters and tts-1-hd at $30. No voice cloning, and OpenAI doesn’t publish latency specs, which matters if you’re latency-sensitive. But if your stack already authenticates against OpenAI, this is one fewer vendor, one fewer key, one fewer invoice.

Fish Audio — best quality-per-dollar

Around $15 per million characters via API, and the models are open-weights — so you can self-host if you’d rather trade cloud spend for GPU spend. It scores well in community blind tests. A strong pick when you need good voice at high volume without premium pricing.

Google, AWS and Azure — best for cloud alignment

Standard voices from Google Cloud and AWS Polly sit around $4 per million characters — an order of magnitude below the premium tier. Polly Neural is $16. They won’t win an expressiveness shootout, but for high-volume IVR prompts, notifications and accessibility, the economics are hard to argue with, and they’re already inside your cloud contract.

Murf and PlayHT — best for creator workflows

Subscription-priced studio tools rather than raw APIs. If your team is producing video voiceovers or podcasts and wants a timeline editor, pronunciation controls and a UI — not a REST endpoint — these are the ones to look at.

Best free and open-source ElevenLabs alternatives

Yes, there are genuinely good free alternatives — you just pay in GPU time and setup instead of dollars. The open-source TTS field moved fast in 2025–26.

ModelLicence / accessVoice cloningTrade-off
Chatterbox (Resemble AI)MIT — fully permissiveYes (~5–10s sample)Needs a GPU; Resemble reports it beat ElevenLabs in a blind preference test — their study, worth verifying yourself
Fish Speech / Fish AudioOpen-weightsYesSelf-host or use the paid API; strong community blind-test results
KokoroOpenLimitedVery lightweight; good for constrained hardware
XTTS v2 (Coqui)OpenYes (~6s sample)Mature, widely used; setup effort
Cloud free tiersFree tierVariesGoogle/AWS/Azure give generous starting credits before pay-as-you-go

⚠️ One thing to check before you self-host: ‘free’ means no licence fee, not no cost. You’re taking on GPU rental or hardware, deployment, scaling, uptime and updates. For a side project that’s a fine trade. For a production phone line answering real customers at 2am, it usually isn’t. Also note ElevenLabs’ own free tier grants no commercial rights — so it isn’t an option for business use at all.

Best ElevenAgents alternatives (the voice-agent layer)

If you’re evaluating ElevenAgents, you’re not shopping for a voice — you’re shopping for a platform. That means the comparison set changes completely. The question stops being ‘which voice sounds best’ and becomes ‘how much of the stack do I have to assemble myself?’

Every voice agent needs seven things: speech-to-text, an LLM, text-to-speech, an agent runtime, telephony, integrations (CRM/calendar/helpdesk), and compliance. Platforms differ mainly in how many of those they hand you versus how many you wire up.

Comparison of voice agent platform components across ElevenAgents, Vapi/Retell/Bland, PolyAI, and SuperMIA
PlatformModelYou assembleBest for
ElevenAgentsDeveloper platform, best-in-class voiceLLM (billed separately), telephony (at cost), CRM, compliance plumbingEngineering teams building a voice product
Vapi / Retell / BlandDeveloper-first agent infrastructureVaries — modular components, BYO keysTechnical teams wanting control over each layer
PolyAIManaged enterpriseLittle — but enterprise scoping and costLarge contact-centre deployments
SuperMIABundled platform (voice + chat)Little — telephony, analytics, compliance bundledOps/growth teams deploying without an AI engineering crew

Retell and Bland deserve their own maths — we’ve done both: here’s the full Retell pricing breakdown and how Bland’s subscription-plus-usage math works. And if you want the wide-angle view of the whole category rather than just the ElevenLabs comparison, see our full comparison of 12 voice agent platforms.

What ElevenAgents actually costs

ElevenLabs publishes ElevenAgents pricing openly, which is to their credit. Here it is, straight from ElevenLabs’ published ElevenAgents pricing (verified July 2026).

PlanPrice / moMinutes includedConcurrent calls
Free$0154
Starter$6756
Creator$2227510
Pro$991,23820
Scale$2993,73830
Business$99012,37540
EnterpriseCustomCustomElevated

Three things sit outside those numbers — and they’re all listed on ElevenLabs’ own pricing page:

  1. The LLM is billed separately. Under ‘External Providers’, their page lists the LLM as ‘based on usage and varies by model.’ Your reasoning layer is a separate line item.
  2. Telephony is billed at cost. Also under ‘External Providers.’ The phone connection is not in the plan price.
  3. Concurrency is capped, and burst pricing doubles the rate. Additional minutes are $0.080. Exceed your plan’s concurrent-call limit and burst pricing charges $0.160/min — exactly double.
Bar chart of ElevenAgents concurrent call limits and included minutes across six plan tiers
ElevenAgents’ published concurrency limits by plan. Source: elevenlabs.io/pricing/agents.

That concurrency ceiling is the line most teams miss. On Pro ($99/mo) you can run 20 simultaneous calls. If you’re a clinic on a Monday morning, a home-services company after a storm, or a retailer on Black Friday, 20 is not a lot — and call 21 either queues or bursts at double rate. It’s not a flaw, it’s a design choice for a developer platform. But it’s the kind of thing you want to know before launch, not after. If predictable, all-in billing matters more to you than component-level control, that’s the argument for a platform that bundles telephony and compliance.

⚠️ Vendor pricing changes frequently — ElevenLabs restructured Agents pricing as recently as late 2025. Every figure above is from their live pricing page as of July 2026. Confirm current rates before you make a decision.

How to choose — 3 questions

Three questions get you to the right answer in about a minute.

1. Script or conversation?

Generating audio from text you already have → you need a TTS engine (Layer 1). Holding a live two-way conversation with a caller → you need a platform (Layer 2). This single question eliminates half the market.

2. (If TTS) What’s your binding constraint?

Latency → Cartesia. Compliance or on-prem → Deepgram Aura-2. Budget at volume → Google/Polly Standard, or self-host Fish/Kokoro. Already on OpenAI → OpenAI TTS. Creator workflow → Murf or PlayHT.

3. (If platform) Do you have engineers to assemble the stack?

Yes, and you want control over every component → ElevenAgents, Vapi, Retell, or Bland. No — you want a working phone line configured by an ops team, with telephony, integrations and compliance included → a bundled platform. Compare plans on SuperMIA’s plans and pricing.

Where SuperMIA fits (and where it doesn’t)

Let’s be direct about this, because the rest of the page is only useful if this part is honest.

SuperMIA is NOT an ElevenLabs TTS alternative. We don’t sell text-to-speech by the character. If you want a voice for an audiobook, a YouTube video, or a dubbing pipeline, use one of the engines in the table above — Cartesia, Deepgram, Fish Audio, OpenAI. That’s a genuinely better answer than us, and we’d rather tell you that than waste your time.

SuperMIA IS an ElevenAgents alternative. If you’re evaluating ElevenAgents to run a real business phone line, that’s the same purchase we serve. SuperMIA’s AI voice agent is a bundled conversational AI platform rather than a developer toolkit. Telephony, analytics, and compliance come in the plan instead of arriving as separate invoices, and the same platform runs chat agents on the same platform — which ElevenAgents, being voice-first, doesn’t bundle the same way. The trade-off is real and worth naming: you get less component-level control than you would assembling Vapi or ElevenAgents yourself. If you have an AI engineering team and want to pick your own LLM, your own TTS engine and your own telephony, a developer platform is genuinely the better fit. If you want a phone line answering customers next week without hiring for it, that’s our lane.

See a voice agent answer a live call — book a 15-minute demo →

Frequently asked questions

What is the best ElevenLabs alternative?

It depends which ElevenLabs product you are replacing. For text-to-speech, Cartesia leads on latency, Deepgram Aura-2 on enterprise compliance and on-premises deployment, OpenAI tts-1 on price for teams already using OpenAI, and Fish Audio on quality per dollar. For ElevenAgents, the voice agent platform, the real alternatives are conversational AI platforms such as SuperMIA, Vapi, Retell, and Bland, because at that layer you are buying telephony, integrations, and compliance rather than a voice.

Is there a free alternative to ElevenLabs?

Yes. Open-source models including Chatterbox from Resemble AI, Fish Speech, Kokoro, and XTTS v2 can be run locally at no licence cost, and several support voice cloning from a short audio sample. The trade-off is that you supply the GPU and the engineering time. Cloud providers also offer generous free tiers, and ElevenLabs itself has a free tier of 10,000 characters per month, though it carries no commercial rights.

What is cheaper than ElevenLabs for text-to-speech?

Most alternatives are cheaper per character. Published list rates put Google Cloud Standard and AWS Polly Standard around $4 per million characters, OpenAI tts-1 and Fish Audio around $15 per million, AWS Polly Neural around $16 per million, Deepgram Aura-2 at $30 per million, and Cartesia in the $38 to $50 range, against roughly $100 per million for ElevenLabs’ top-tier v3 model. Confirm current rates on each vendor’s pricing page before committing.

How much does ElevenAgents actually cost?

ElevenLabs publishes ElevenAgents plans from Free (15 minutes, 4 concurrent calls) up to Business at $990 per month (12,375 minutes, 40 concurrent calls), with additional minutes at $0.080. Two things sit outside that figure on their own pricing page: the LLM is billed based on usage and varies by model, and telephony is billed at cost through an external provider. Exceeding your plan’s concurrency limit triggers burst pricing at $0.160 per minute, which is double the standard rate.

Is ElevenLabs good for running a business phone line?

ElevenAgents is a capable developer platform for building voice agents, and its voice quality is widely regarded as the best available. It is a weaker fit if you want a business phone line configured by an operations team rather than assembled by engineers, because you still connect the LLM, telephony, CRM, and compliance layers yourself, and HIPAA BAAs are only offered on the Enterprise tier. Teams in that position usually compare full conversational AI platforms instead.

Share this article:
Urvil Dhanani

Urvil Dhanani

Urvil Dhanani is the AI/ML Lead at SuperMIA, focused on the architecture behind reliable conversational AI — agent design patterns, voice and chat orchestration, and platform evaluation. He writes practical, vendor-neutral guides that help technical teams build and choose AI systems that hold up in production.