Meeting assistants
Give every participant notes and action items in their own language, wherever your global team is based.
Learn more →Audio intelligence
Add speech-to-text translation to any transcription request and get audio in 100+ languages in one API call. Gladia returns the original transcript and every target language together, with timestamps and speaker labels kept, both for live streams and recorded files.
Translate audio from your own multilingual recordings. No expiry, no credit card.
Translate speech-to-text in any supported language, including calls where speakers switch mid-sentence.
Real-time translations that keep pace with live conversations.
Gladia has been multilingual from day one. Transcribe and translate from any to any of the 100+ supported languages, in a single API.
Skip stitching together transcription and translation APIs. Translate recorded and live audio in one request, with timestamps already aligned.
Upload a file, pass a URL, or open a real-time WebSocket session.
Set "translation": true, list your target_languages, and choose the base or enhanced model.
Each target language returns its full transcript, utterances, and word timestamps. In live sessions, every utterance is translated as it's finalized.
Translation is only as accurate as the transcript underneath it. Both Gladia models support speech-to-text translation, so pick the one that best transcribes the languages your users speak.
Highest accuracy on European real-world audio
Maximum language coverage across any domain
Set the model, context, and style per request, because support calls, dubbed video, and translated subtitles each need something different.
| Setting | What it does | Use it for |
|---|---|---|
model | base is fast and covers most use cases. enhanced is slower, with higher quality and context awareness | base for live calls, enhanced for domain-heavy content |
context | Describes the conversation to improve terminology, proper nouns, and disambiguation | Medical, legal, or product-specific recordings |
match_original_utterances | Keeps translated segments aligned with the original speech segments | Subtitles and dubbing |
lipsync | Aligns translated output with the speaker’s lip movements | Dubbed video content |
informal | Uses informal register where a language has one, like “tu” in French or “du” in German | Consumer apps and chat-style products |
Use cases
Serve users in their own language, from live support calls to dubbed video, without adding a separate translation service to your stack.
Give every participant notes and action items in their own language, wherever your global team is based.
Learn more →Bring calls from every market into one language, so QA, sentiment, and analytics work across your whole operation.
Learn more →Let callers speak their own language while your agent understands and responds in real time.
Learn more →Turn interviews, podcasts, and videos into subtitles for audiences in 100+ languages.
Learn more →Your rate stays the same whether you translate into one language or ten. There's no separate translation bill and no per-language fees.
Flexible pay-as-you-go for moderate audio volumes. Get started immediately.
Async at $0.61/hr
Real-time at $0.75/hr
* 50€ in free credits
Lower unit pricing for fast-growing teams. Commit upfront to unlock savings.
Async as low as $0.20/hr
Real-time as low as $0.25/hr
* 67% less than Starter
Annual plan with custom models, fine-tuning, debundled pricing, and more.
Custom
Start free with 50€ in credits, or book a demo to test speech-to-text translation on your own audio.