real-time

The real-time speech-to-text API built for live conversations. Sub-300ms latency. Accurate from the first word.

<103ms

Partial transcripts

Fastest on the market.

#1

On real-life speech

Best accuracy on conversational audio benchmarks.

100+

Languages covered

42 unique to Gladia.

Trusted by 300,000+ developers worldwide

Klarna Aircall HeyGen VEED Recall Attention Livestorm Method Sana Mojo Adversus Claap Attio Carv Selectra Spoke
Klarna Aircall HeyGen VEED Recall Attention Livestorm Method Sana Mojo Adversus Claap Attio Carv Selectra Spoke

Lowest error rate in the industry.

Gladia's Solaria-1 outperforms every provider on Switchboard, the toughest conversational benchmark. Our benchmark methodology is open-source, so you can reproduce the results.

Compare models Compare models
33.9%
Solaria-3
37.3%
Solaria-1
42.3%
AssemblyAI
46%
Speechmatics
48.1%
Mistral
49.8%
Deepgram
55.2%
ElevenLabs
LATENCY

Fast enough for real conversation

Low enough latency to feel instant, with controls to tune speed and accuracy to your use case.

  • 103ms partial transcripts: words appear as they're spoken
  • Sub-300ms end-to-end latency for uninterrupted dialogue
  • Adjustable endpointing to tune speed vs. accuracy per use case
  • Stable finals: partials resolve into accurate, clean utterances
Sub-300ms latency with adjustable speed and accuracy tradeoff
ACCURACY

Get the important stuff right

Trained on real, noisy audio, so accuracy holds up where it matters most.

  • #1 on Switchboard: the lowest error rate on the toughest conversational benchmark
  • Custom vocabulary for product names, acronyms, and domain-specific terms
  • Audio enhancer: pre-processing for noisy streams and variable network conditions
  • Word-level timestamps on every utterance
Accurate real-time transcription with entity recognition
MULTILINGUAL

Speak any language, switch mid-sentence

Handles mid-sentence language switches automatically, with broader language coverage than any other provider.

  • 100+ languages, 42 exclusive to Gladia
  • Code-switching detected mid-sentence, automatically
  • Auto language detection, no config needed
  • Multi-channel transcription with separate streams per participant over one WebSocket
Real-time transcription and insights across 100+ languages
COMPATIBILITY

One API, any stack

Works with the frameworks, telephony providers, and automation tools your team already uses.

  • WebSockets, SIP, VoIP for any protocol
  • Pipecat, LiveKit, Vapi for voice agent pipelines
  • Twilio, Recall, Meeting BaaS for telephony and meeting bots
  • 99.99% uptime SLA
API compatibility with SIP, VoIP, WebSockets, and framework integrations

Everything you need at no extra cost

Skip the tool-stitching. Get raw audio to accurate data in one place.

Check the audio-intelligence suiteCheck the audio-intelligence suite

Custom vocabulary

Get accurate transcripts on the terms that matter most with keyterm prompting: product names, acronyms, and domain-specific language.

Learn more

Code-switching

Transcribe conversations that shift between languages mid-sentence — no manual configuration required.

Learn more

Named entity recognition

Automatically identify and extract people, organizations, locations, and key terms directly from your transcripts.

Learn more
USE CASES

Purpose-built for enterprise teams

Teams use the real-time API to turn live audio into accurate text the moment it's spoken.

Voice agents

Transcribe caller speech as it happens to power conversational AI agents that respond without dead air.

Learn more

Live captioning

Generate accurate, low-latency captions for live events, broadcasts, and accessibility compliance.

Meeting copilots

Stream transcripts into note-taking and meeting assistant tools as the conversation happens, not after.

Learn more

Customer support

Give live agents real-time transcripts and prompts during calls for faster, more accurate resolutions.

Learn more
With world-class language auto-detection, translation across 100+ languages, and outstanding performance in French, we're proud to partner with Gladia.
Kwin Kramer
Kwin Kramer Co-founder at Daily

Compare pricing

Pay only for the audio you transcribe, with pricing that drops automatically as volume grows.

Starter

Flexible pay-as-you-go for moderate audio volumes. Get started immediately.

Async at $0.61/hr

Real-time at $0.75/hr

  • What's included
  • 100+ languages supported
  • Speaker diarization
  • Word-level timestamps
  • Community support

Enterprise

Annual plan with custom models, fine-tuning, debundled pricing, and more.

Custom

  • What's included
  • Everything in Growth
  • Dedicated infrastructure
  • GDPR & SOC 2 compliance support
  • 99.9% uptime SLA
  • Named account manager
  • On-prem deployment available
Explore the pricing Explore the pricing

FAQs

Start transcribing in minutes

Sign up for free and get an API key, or book a demo to see Gladia's real-time transcription handle your own audio.