Now available

Flux TTS. Keep talking.

Text-to-speech built for real-time conversations. Flux TTS reads the whole conversation, not just the current line, carrying context from first response to last.

Meet the only voices that think before they speak.

Trusted by those shipping the best in voice AI.

Twilio logo
Pipecat
livekit
jambonz
vapi
Cloudflare
Nice-cognigy trustbarCoval
Twilio logo
Pipecat
livekit
jambonz
vapi
Cloudflare
Nice-cognigy trustbarCoval

One conversation. Two models

Flux STT handles the listening. Flux TTS handles the speaking. Both are purpose-built for interruptions, conversational cues, and the flow of real dialogue.

Flux architecture
Flux architectureFlux STTLLMFlux TTS

Today, Flux STT and Flux TTS run as one integrated stack on a single connection, with Deepgram orchestrating the timing between listening and speaking.

The right emotion for every moment

Pick a voice built to stay consistent across the entire conversation, then hear it handle the moments agents actually face: empathetic when it matters, precise when details count, warm from open to close.

One call · 0:51 · Haleyhover a line, click to hear it
Warm greetingLine 1 of 6 · 00:00

A month of free conversation

Through September 12, 2026, developers can build with Flux TTS free with up to 45 concurrent streaming connections globally (5 in EU/AU). Standard pricing applies beginning September 13, 2026.

It just sounds right

Flux TTS is naturally expressive. It delivers the right pacing, emphasis, and emotional register based on the context of the conversation. No prompt engineering, SSML, or style tags required.

Best-voice win rate, top models

  1. #1Deepgram Flux TTS73.4%
  2. #2Cartesia Sonic 3.572.0%
  3. #3Google Gemini 2.5 Flash62.4%
  4. #4Rime Arcana61.4%
  5. #5ElevenLabs Eleven v361.2%
  6. #6 Inworld TTS-1.5 MAX57.9%
  7. #7ElevenLabs Eleven Flash v2.552.6%
  8. #8 Inworld TTS-250.5%
  9. #9xAI Grok TTS49.5%
  10. #10OpenAI GPT Realtime49.4%
  11. #11OpenAI GPT-4o mini TTS33.5%
  12. #12AWS Polly Generative22.4%

Ranked by blind head-to-head win rate, scored on each model's best voice.

MethodologyBlind two-alternative forced choice: two voices read the same line, the rater picked the one that sounded more natural. ~14,400 paired comparisons across 12 models and ~73 voices, on 100 scripts drawn from real use cases. July 2026.

Best-voice expressiveness win rate

  1. #1Deepgram Flux TTS77.5%
  2. #2Cartesia Sonic 3.577.0%
  3. #3Google Gemini 2.5 Flash75.1%
  4. #4ElevenLabs Eleven Flash v2.570.1%
  5. #5ElevenLabs Eleven v368.9%
  6. #6 Inworld TTS-1.5 MAX63.9%
  7. #7Rime Arcana63.5%
  8. #8OpenAI GPT Realtime58.8%
  9. #9 Inworld TTS-250.9%
  10. #10OpenAI GPT-4o mini TTS48.8%
  11. #11xAI Grok TTS46.4%
  12. #12AWS Polly Generative24.0%

Ranked by blind head-to-head win rate, scored on each model's best voice.

MethodologyBlind two-alternative forced choice: two voices read the same line, the rater picked the one that sounded more expressive. ~14,400 paired comparisons across 12 models and ~73 voices, on 100 scripts drawn from real use cases. July 2026.

Flux TTS holds a conversation

The API is a conversational state machine, with a defined lifecycle for every turn. Carry context, handle interruptions, and respond without the need for additional developer orchestration.

Live conversation
Agent
Your balance is $2,847
Barge-in detected
Caller
Wait — is that before the payment?
Agent
As of this morning, yes. Want me to pull up the full statement?
WebSocket stream · wss://api.deepgram.com/v2/speak
SpeechStarted { speech_id: "dg_sp_a1b2c3d4e5f6" }
audio 
audio 
audio 
INTERRUPT
SpeechInterrupted {
  speech_id: "dg_sp_a1b2c3d4e5f6",
  text_spoken: "Your balance is $2,847",
  text_remaining: "as of this morning." // ✕ not played
}
Configure { speed: 1.05 }
SpeechStarted { speech_id: "dg_sp_b2c3d4e5f6a1" }
audio 
audio 
SpeechMetadata { speech_id: "dg_sp_b2c3d4e5f6a1" }

Works where your conversations happen

Flux TTS drops into the voice-agent stack you already run, and deploys wherever your data has to live — cloud, VPC, or on-prem.

  • PlugNative in your stackFirst-class integrations for Vapi, Pipecat, LiveKit, Jambonz, and Cloudflare.
  • ShieldFlexible deploymentSelf-host in your own cloud or on-prem for data residency, security, and regulated workloads.
  • VoiceStable voice identity, 24/7Consistent delivery across thousands of words for always-on production.

Your voice agent stack

VapiVapi
PipecatPipecat
LivekitLivekit
jambonzJambonz
CloudflareCloudflare
MicFlux TTSspeak.v2 · wss /v2/speak
cloudManaged cloud
VPCYour VPC
PremOn-prem

Gets the hard stuff right

9.8%

Word error rate on hard prompts

  • Drug names
  • Tracking numbers
  • IVR menus
  • $ amounts
  • Technical strings
  • Alphanumerics
01
Flux TTS
3.4%
02
Inworld 1.5 Max
5.0%
03
ElevenLabs v3
6.6%
04
Cartesia sonic-3.5
9.8%

Lower is better. Flux TTS (Haley).

The capabilities that set Flux TTS apart

Capabilities comparison
CapabilityFlux TTSElevenLabsCartesiaInworldRime
Cross-turn contextXX
No markup required (context-aware delivery)XXX
Interrupt feedback (text_spoken on barge-in)XX??
Server-side flushing??
Production accuracy on structured contentX???
Conversation-native WebSocket protocol
Self-host / on-prem
HIPAA + SOC 2 included
Concurrency at enterprise scale?

Built for the teams shipping voice AI

Hear how companies are building faster, more natural voice experiences with Flux TTS.

Frequently asked questions

What is Flux TTS?
What is the best text-to-speech for voice agents?
How does Flux TTS compare to ElevenLabs?
How does Flux TTS compare to Cartesia?
Is Flux TTS an alternative to Rime or Inworld?
How does Flux TTS handle interruptions?
How does Flux TTS stay consistent across turns?
What is the latency of Flux TTS?
Does Flux TTS support SSML or style tags?
How do I migrate to Flux TTS from another provider?
Can I self-host or deploy Flux TTS on-prem?
What languages does Flux TTS support?
How is Flux TTS different from Aura-2, and is Aura-2 going away?
What is the Flux family?
How do I get started with Flux TTS?
Is Flux TTS HIPAA compliant?

Keep talking

Build your next voice agent with Flux TTS and hear the difference across the whole conversation. Switch by changing the model parameter to flux-{voice}-en