Now available
Flux TTS. Keep talking.
Text-to-speech built for real-time conversations. Flux TTS reads the whole conversation, not just the current line, carrying context from first response to last.
Meet the only voices that think before they speak.
Trusted by those shipping the best in voice AI.
One conversation. Two models
Flux STT handles the listening. Flux TTS handles the speaking. Both are purpose-built for interruptions, conversational cues, and the flow of real dialogue.
Today, Flux STT and Flux TTS run as one integrated stack on a single connection, with Deepgram orchestrating the timing between listening and speaking.
The right emotion for every moment
Pick a voice built to stay consistent across the entire conversation, then hear it handle the moments agents actually face: empathetic when it matters, precise when details count, warm from open to close.
A month of free conversation
Through September 12, 2026, developers can build with Flux TTS free with up to 45 concurrent streaming connections globally (5 in EU/AU). Standard pricing applies beginning September 13, 2026.
It just sounds right
Flux TTS is naturally expressive. It delivers the right pacing, emphasis, and emotional register based on the context of the conversation. No prompt engineering, SSML, or style tags required.
Best-voice win rate, top models
- #1Deepgram Flux TTS73.4%
- #2Cartesia Sonic 3.572.0%
- #3Google Gemini 2.5 Flash62.4%
- #4Rime Arcana61.4%
- #5ElevenLabs Eleven v361.2%
- #6 Inworld TTS-1.5 MAX57.9%
- #7ElevenLabs Eleven Flash v2.552.6%
- #8 Inworld TTS-250.5%
- #9xAI Grok TTS49.5%
- #10OpenAI GPT Realtime49.4%
- #11OpenAI GPT-4o mini TTS33.5%
- #12AWS Polly Generative22.4%
Best-voice expressiveness win rate
- #1Deepgram Flux TTS77.5%
- #2Cartesia Sonic 3.577.0%
- #3Google Gemini 2.5 Flash75.1%
- #4ElevenLabs Eleven Flash v2.570.1%
- #5ElevenLabs Eleven v368.9%
- #6 Inworld TTS-1.5 MAX63.9%
- #7Rime Arcana63.5%
- #8OpenAI GPT Realtime58.8%
- #9 Inworld TTS-250.9%
- #10OpenAI GPT-4o mini TTS48.8%
- #11xAI Grok TTS46.4%
- #12AWS Polly Generative24.0%
Flux TTS holds a conversation
The API is a conversational state machine, with a defined lifecycle for every turn. Carry context, handle interruptions, and respond without the need for additional developer orchestration.
Works where your conversations happen
Flux TTS drops into the voice-agent stack you already run, and deploys wherever your data has to live — cloud, VPC, or on-prem.
Native in your stackFirst-class integrations for Vapi, Pipecat, LiveKit, Jambonz, and Cloudflare.
Flexible deploymentSelf-host in your own cloud or on-prem for data residency, security, and regulated workloads.
Stable voice identity, 24/7Consistent delivery across thousands of words for always-on production.
Your voice agent stack
Gets the hard stuff right
9.8%
Word error rate on hard prompts
- Drug names
- Tracking numbers
- IVR menus
- $ amounts
- Technical strings
- Alphanumerics
Lower is better. Flux TTS (Haley).
The capabilities that set Flux TTS apart
| Capability | Flux TTS | ElevenLabs | Cartesia | Inworld | Rime |
|---|---|---|---|---|---|
| Cross-turn context | ✓ | X | X | ✓ | ◒ |
| No markup required (context-aware delivery) | ✓ | X | X | ◒ | X |
| Interrupt feedback (text_spoken on barge-in) | ✓ | X | X | ? | ? |
| Server-side flushing | ✓ | ◒ | ◒ | ? | ? |
| Production accuracy on structured content | ✓ | X | ? | ? | ? |
| Conversation-native WebSocket protocol | ✓ | ✓ | ✓ | ◒ | ✓ |
| Self-host / on-prem | ✓ | ◒ | ✓ | ◒ | ◒ |
| HIPAA + SOC 2 included | ✓ | ◒ | ✓ | ✓ | ◒ |
| Concurrency at enterprise scale | ✓ | ◒ | ◒ | ? | ◒ |
Built for the teams shipping voice AI
Hear how companies are building faster, more natural voice experiences with Flux TTS.
Frequently asked questions
Keep talking
Build your next voice agent with Flux TTS and hear the difference across the whole conversation. Switch by changing the model parameter to flux-{voice}-en