Table of Contents
A patient in Madrid describes chest pain that started three days ago. A clinician in Munich moves a metformin dose from 500 to 850 milligrams. A family member translates for a parent during intake, switching between Hindi and English mid-sentence. Each of these conversations carries the information a healthcare application depends on: conditions, medications, doses, dates and numbers.
A general multilingual speech model can transcribe all three. Getting the medical details right is the harder problem. A model can support Spanish, German and Hindi and still miss the drug name, the dose or the diagnosis that everything downstream relies on. Supporting a language isn't the same as supporting medicine in that language.
Nova-3 Medical, which we launched in 2025, was built around a simple idea: medical speech needs a model trained on the language of medicine. Teams serving patients in other languages often rely on general multilingual models instead, which cover the language but miss more of the vocabulary that matters most.
Today, we're extending Nova-3 Medical to 10 languages. Nova-3 Medical Multilingual brings medical-specialized speech recognition to Dutch, English, French, German, Hindi, Italian, Japanese, Portuguese, Russian and Spanish, in both batch and streaming, with automatic code switching between supported languages. One model, one parameter: model=nova-3-medical&language=multi.
Medical Specialization That Holds Across Languages
To measure what specialization adds, we compared Nova-3 Medical Multilingual against Nova-3 Multilingual, our general-purpose multilingual model, on more than 115 hours of curated medical audio across all 10 languages, averaging results across languages. In batch transcription, average word error rate (WER) dropped from 6.62% to 4.96%, a 25% reduction. In streaming, it dropped from 8.32% to 5.86%, a 30% reduction.
The gains are broad. WER was significantly lower in 8 of 10 languages in batch and 9 of 10 in streaming. Russian, Spanish and Hindi saw the largest batch improvements, at 58%, 32% and 30%. English, where Nova-3 Medical already had a dedicated model, was essentially flat.
Not All Transcription Errors Are Equal
WER counts every word the same. Missing "the" and missing "metformin" each add one error. For an application drafting a clinical note, filling an intake form or deciding what a voice agent says next, those are very different mistakes.
That's why we also measure entity error rate (EER), which tracks how accurately a model recognizes the important entities in speech: conditions, drugs, doses, dates, numbers, names and more. Across the 10-language evaluation, Nova-3 Medical Multilingual cut EER from 14.46% to 9.11% in batch, a 37% reduction, and from 16.31% to 9.78% in streaming, a 40% reduction.
The largest gains land in the categories healthcare applications act on. Compared with Nova-3 Multilingual in batch transcription:
- Doses: 64% fewer errors (81% fewer in streaming)
- Conditions: 43% fewer errors
- Medical processes: 41% fewer errors
- Drugs: 31% fewer errors
Specialization shows up most where the words carry medical meaning. That's where a transcript turns into structured data, documentation or an action, and where an error costs the most.
Competitive Accuracy Across the Full Language Set
We also ran other speech recognition models through their APIs on the same medical dataset. Among models evaluated in all 10 languages, it posted the lowest average batch WER:
- Nova-3 Medical Multilingual: 4.96%
- ElevenLabs Scribe v2 Medical: 5.49%
- ElevenLabs Scribe v2: 5.58%
- Microsoft MAI 1.5: 7.80%
Nova-3 Medical Multilingual had the lowest WER in 7 of the 10 individual languages. No model wins every language, ours included. That's why we report results across the full language set: an application serving patients in five languages needs accuracy in all five.
Built for Real-Time Medical Applications
More healthcare speech applications now need to understand the conversation while it's happening. In streaming, Nova-3 Medical Multilingual averaged 5.86% WER and 9.78% EER across all 10 languages, and recorded the lowest WER in every language against the competing streaming models evaluated for each one.
That opens up a wider range of workloads on a single model:
Ambient documentation. Capture the visit as it happens, so note drafts reflect the medications, doses and conditions discussed, whichever supported language the patient and clinician speak.
Healthcare voice agents. Patient-facing agents for scheduling, refills and intake need the drug name and dose right before they take the next step. Streaming medical accuracy gives the agent a cleaner input to act on.
Dictation and clinical workflows. Clinicians can dictate in the language they work in, and switch languages when a conversation calls for it, without changing models.
How It Works
Set model=nova-3-medical and language=multi on any pre-recorded or streaming request. The model detects the language as it's spoken and follows the conversation when speakers switch, so you don't need to know the language before the audio arrives.
curl -X POST "https://api.deepgram.com/v1/listen?model=nova-3-medical&language=multi" \
-H "Authorization: Token $DEEPGRAM_API_KEY" \
-H "Content-Type: audio/wav" \
--data-binary @patient_visit.wav
For real-time audio, use the same parameters on the streaming endpoint (https://api.deepgram.com/v1/listen). Nova-3 Medical Multilingual is also available for self-hosted deployments. Request the models from your Deepgram account representative.
Start Building with Nova-3 Medical Multilingual
Nova-3 Medical Multilingual is available now through the Deepgram API.
Healthcare is one of the clearest cases for specialized speech models. The more consequential the vocabulary, the more a model built for that domain matters. Nova-3 Medical brought that specialization to English. Nova-3 Medical Multilingual brings it to 10 languages, so you can support more of the people your application serves without giving up medical transcription accuracy.

