Platform overview

Speech-to-text API

One speech-to-text API for meeting assistants, contact centers, and voice agents. Built for regulated industries, with hosting options in the EU and US. Turn live conversations and recorded audio into accurate, structured transcripts, in any language.

#1
On real-life speech
9.6% WER on real conversational audio, the lowest among 8 providers.
<300ms
Latency
Fast enough for natural, real-time voice experiences, with no waiting on the final word.
100+
Languages covered
Detect and transcribe mid-sentence switches across every supported language.
Trusted by over 350,000 users and 2,000+ enterprise teams
HeyGen Livestorm Adversus Selectra Circleback Recall.ai
Noota Robonote RapidSOS OVHcloud SFR Clariane

Transcription modes

Two ways to transcribe

Pick the mode that matches how your audio actually arrives, whether that's live conversations or pre-recorded files.

Real-time

Transcribe as people talk

For voice agents, live captioning, and call center tooling where a few hundred milliseconds is the difference between a natural conversation and an awkward pause. Streams audio in, returns text in near real time.

Async

Transcribe what’s already recorded

For meeting recordings, podcasts, call center QA, and anything where you'd rather optimize for accuracy and rich output (speaker labels, summaries, formatting) than for speed.

Audio intelligence

Go beyond transcription

A native suite of audio intelligence features helps you understand who spoke when, extract key entities and sentiment, and generate summaries or action items in a single pipeline.

Speaker diarization

Know who said what, with the lowest diarization error rate of any speech-to-text provider, proven by benchmarks. Powered by pyannoteAI, Precision-2.

Learn more

Audio-to-LLM

Turn audio into structured insights, including summaries and action items, in a single API call, with access to 400+ LLM models.

Learn more

Code-switching

Capture every word when speakers mix languages, with real-time language detection that follows the conversation wherever it goes.

Learn more

Speech-to-text translation

Translate each utterance into 100+ languages, keeping timestamps and speaker labels attached.

Learn more

Custom vocabulary

Boost recognition accuracy for product names, acronyms, and domain-specific terms with keyterm prompting.

Learn more

PII redaction

Automatically mask names, emails, phone numbers, and other sensitive data before it ever leaves your pipeline.

Learn more

Sentiment analysis

Detect positive, negative, and neutral tone across every call, with sentiment tied to each speaker and timestamp.

Learn more

Named entity recognition

Pull out the names, dates, and organizations that matter, turning raw speech into structured, queryable data.

Learn more

Audio summarization

Condense hours of conversation into clear recaps, in concise, general, or bullet-point formats.

Learn more
See the full suite of features

Supported languages

Multilingual since day one

Being multilingual has been our priority since day one. Gladia supports 100 languages, with code-switching and automatic language detection integrated into all models by default.

EnglishEN
SpanishES
FrenchFR
GermanDE
ItalianIT
PortuguesePT
ChineseZH
HindiHI
ArabicAR
JapaneseJA
DutchNL
PolishPL
AfrikaansAF
AlbanianSQ
AmharicAM
ArmenianHY
AssameseAS
AzerbaijaniAZ
BashkirBA
BasqueEU
BelarusianBE
BengaliBN
BosnianBS
BretonBR
BulgarianBG
CatalanCA
CroatianHR
CzechCS
DanishDA
EstonianET
FaroeseFO
FinnishFI
GalicianGL
GeorgianKA
GreekEL
GujaratiGU
Haitian CreoleHT
HausaHA
HawaiianHAW
HebrewHE
HungarianHU
IcelandicIS
IndonesianID
JavaneseJW
KannadaKN
KazakhKK
KhmerKM
KoreanKO
LaoLO
LatinLA
LatvianLV
LingalaLN
LithuanianLT
LuxembourgishLB
MacedonianMK
MalagasyMG
MalayMS
MalayalamML
MalteseMT
MaoriMI
MarathiMR
MongolianMN
MyanmarMY
NepaliNE
NorwegianNO
NynorskNN
OccitanOC
PashtoPS
PersianFA
PunjabiPA
RomanianRO
RussianRU
SanskritSA
SerbianSR
ShonaSN
SindhiSD
SinhalaSI
SlovakSK
SlovenianSL
SomaliSO
SundaneseSU
SwahiliSW
SwedishSV
TagalogTL
TajikTG
TamilTA
TatarTT
TeluguTE
ThaiTH
TibetanBO
TurkishTR
TurkmenTK
UkrainianUK
UrduUR
UzbekUZ
VietnameseVI
WelshCY
WolofWO
YiddishYI
YorubaYO

Benchmarks

How it compares

Anyone can claim they are the best. We’d rather put the numbers side by side and let you decide which trade-offs matter for your product.
WER and DER. Lower is better.

Conversational speech

Spontaneous telephone conversation. WER results.

Diarization

Weighted avg. Diarization Error Rate across 10 domains.

Real customer calls

Gladia’s internal production dataset, human-annotated. WER results.

Use cases

Built for real world use cases

Teams use Gladia to turn live and recorded conversations into accurate, speaker-labeled data their products can act on.

Meeting assistants

Accurate, speaker-labeled transcripts that keep AI summaries, action items, and CRM notes grounded in what was said.

Learn more →

Contact centers (CCaaS)

Every call transcribed and ready for QA, compliance, and coaching, even when callers switch languages mid-sentence.

Learn more →

Voice agents

Your agent hears callers as they speak and responds without dead air, in 100+ languages.

Learn more →

Media & content

Accurate, time-stamped subtitles and transcripts for video and podcasts, including multi-channel audio.

Learn more →

SDKs & integrations

Built for developers

Go from API key to production in an afternoon, with the tools and frameworks you already use.

SDKs

Official SDKs for Python (gladiaio-sdk) and JavaScript/TypeScript (@gladiaio/sdk). REST and WebSocket APIs for everything else.

SDK docs

Webhooks

Get results pushed the moment an async job completes, with no polling.

Webhooks docs

Agent-ready

Connect Claude, Cursor, or Codex to your Gladia account with the open-source MCP server, or install Gladia Skills so your coding agent writes the integration for you.

Integrations

Pipecat · LiveKit · Vapi · Twilio · SIP / VoIP · Recall · Meeting BaaS

All integrations
POST /v2/pre-recorded
200 OK
{
  "result": {
    "transcription": {
      "languages": ["en", "es"],
      "utterances": [
        {
          "speaker": 0,
          "start": 0.42,
          "text": "Hi, this is Ana from support."
        },
        {
          "speaker": 1,
          "start": 3.10,
          "text": "Hola, necesito ayuda con la factura."
        }
      ]
    }
  }
}

Transcribe in minutes

Start free with 50€ in credits, or book a demo to test Gladia on your own audio.

FAQs