Skip to main content
Retell supports seven voice providers: platform voices (Retell’s own library), MiniMax, Cartesia, Inworld, Fish Audio, ElevenLabs, and OpenAI. The provider behind your agent’s voice affects how it sounds, which languages it can speak, and which features you can use. Start with a platform voice unless you have a specific reason not to. Platform voices are fine-tuned for phone conversations, handle TTS fallback automatically, and are the only voices that support Expressive Mode. Reach for a third-party provider voice when you need a specific voice, accent, or capability a platform voice doesn’t cover. Many voices in the library are available from more than one provider — the voice ID prefix tells you which provider serves it (for example, minimax-Cimo, cartesia-Cimo, and 11labs-Cimo are the same persona on different providers). This makes it easy to compare providers while keeping a similar voice.

Feature comparison

For the languages each provider supports, see Language support by provider, including the model-level restrictions where a provider gates certain languages to specific models.

Provider notes

Subjective observations below come from our internal testing. For an independent read on raw voice quality, the Artificial Analysis Speech Arena runs blind pairwise listening tests (English only); as of August 2026 it places Inworld’s and ElevenLabs’ newest models near the top, with MiniMax and Fish Audio close behind. Results vary by voice, model, and language — preview voices in the dashboard before committing.

Platform voices

Retell’s curated library, fine-tuned for conversational AI over the phone: natural fillers, pacing, and clarity at telecom bitrates. Retell routes each platform voice to the best underlying model for your agent’s language, and fallback during a provider outage is automatic, with no fallback plan to configure. Platform voices are also the only way to use Expressive Mode emotion and effect tags.

MiniMax

Strongest at reading spelled-out words (“W - O - R - D”) and the most consistent pacing and tone in our testing. Strong coverage of Asian languages: Cantonese, for example, is only available through MiniMax and platform voices. Supports the voice_emotion API field. Voices can sound somewhat more robotic than other providers.

Cartesia

Natural-sounding voices with stronger spelling accuracy than ElevenLabs in our testing, and among the lowest synthesis latency of any provider in published benchmarks. Supports the voice_emotion API field. Pacing and tone can be less consistent, and localization is weaker for some accents.

Inworld

The newest provider on Retell. Inworld’s models currently rank first in the Artificial Analysis Speech Arena’s blind listening tests, and inworld-tts-2-flash trades some quality for lower latency. Language coverage is narrower than most other providers: 15 languages, including English, Chinese, Japanese, Korean, Hindi, and Arabic. Supports voice cloning.

Fish Audio

Broad language coverage across its s2-pro and s2.1-pro models; the older s1 model supports a smaller language set. Its s2-pro model is the highest-ranked open-weight model in the Artificial Analysis Speech Arena. Supports voice cloning.

ElevenLabs

A large voice library with wide accent variety; community voices added through custom voice search come from ElevenLabs. In our testing, exact spelling is less reliable and you may notice occasional pacing or tone quirks. Note that eleven_flash_v2 is English-only, and the older Turbo models are deprecated in favor of their Flash equivalents.

OpenAI

Among the broadest language coverage of any provider, with the same language set on both of its models. Does not support voice cloning.

How to choose

  • Default choice, or calls over the phone → a platform voice
  • Most natural sound in blind listening tests → Inworld, then confirm with your own previews
  • Emotion tags, sighs, pauses, emphasis → a platform voice with Expressive Mode
  • Spelling accuracy (confirmation codes, emails, names) → MiniMax or Cartesia
  • Most consistent pacing and tone → MiniMax
  • Asian languages, including Cantonese → MiniMax or a platform voice
  • Emotional tone set per agent via the API (voice_emotion) → Cartesia or MiniMax
  • Voice cloning with zero fallback setup → clone to platform with the Clone Voice API
  • A specific language → check Language support by provider first, then pick among the providers that cover it

FAQ

No. Retell hosts all seven providers — you pick a voice and Retell handles the provider behind it. There is nothing to sign up for or configure on the provider’s side.
With a platform voice, Retell fails over automatically and the agent keeps talking. With any other provider’s voice, configure a TTS fallback plan so the agent can switch to a backup voice from a different provider mid-call.
Often, yes. Many library voices exist on several providers under the same persona name — swap cartesia-Cimo for minimax-Cimo, for example. Preview both in the dashboard, since the same persona can still sound slightly different across providers.