Meeting assistants
Attribute every decision and action item to the person who said it, so AI summaries and follow-ups are trustworthy.
Learn more →Audio intelligence
The most accurate speaker diarization on the market, powered by pyannoteAI Precision-2 and built into one async API for meeting assistants, contact centers, and media platforms. Turn multi-speaker recordings into transcripts that show exactly who said what, in any language.
Test speaker diarization on 80+ hours of your own audio. No expiry, no credit card.
3x fewer speaker errors than leading APIs, and the lowest diarization error rate of any provider.
Speaker labels stay accurate in any language, even when people switch mid-sentence.
Gladia has the lowest diarization error rate (DER) of 9 providers tested on DIHARD III, the standard benchmark for real-world multi-speaker audio. Every provider was tested on identical files through its production API with default settings, and the full methodology is open-sourced so anyone can reproduce the results.
*Note: DER measures the share of audio time with a diarization error: missed speech, false alarms, or speech attributed to the wrong speaker. Lower is better. More about DER on our blog.
We pair the model from pyannote’s creators, the team behind the open-source standard for speaker diarization, with Gladia’s Solaria transcription, so you get every word with the right speaker and timestamp.
Speaker diarization runs in the same async request as transcription. There’s no second model to host, no separate pipeline to orchestrate, and no timestamp reconciliation between services.
Upload a file or pass a URL to the async endpoint. Mono, stereo, and multichannel audio are all supported.
Set "diarization": true. Optionally pass the exact number of speakers, or a min and max range, to tighten accuracy.
Each utterance comes back with a speaker index, text, language, and start and end timestamps.
Use cases
From sales calls to panel interviews, teams rely on Gladia to know exactly who said what, and to build summaries, QA, and analytics on top of it.
Attribute every decision and action item to the person who said it, so AI summaries and follow-ups are trustworthy.
Learn more →Separate agent from customer on mono recordings to power call QA, compliance checks, and coaching.
Learn more →Produce speaker-labeled transcripts, captions, and show notes for interviews, panels, and debates.
Learn more →Learn how speaker diarization works, what drives errors such as overlapping speech, and how accuracy is measured.
Read the guide →Every plan comes with speaker labels built in. Your bill stays the same whether it’s a two-person call or a ten-person meeting.
Flexible pay-as-you-go for moderate audio volumes. Get started immediately.
Async at $0.61/hr
Real-time at $0.75/hr
* 50€ in free credits
Lower unit pricing for fast-growing teams. Commit upfront to unlock savings.
Async as low as $0.20/hr
Real-time as low as $0.25/hr
* 67% less than Starter
Annual plan with custom models, fine-tuning, debundled pricing, and more.
Custom
Start free with 50€ in credits, or book a demo to test Gladia on your own audio.