Audio intelligence

Speaker diarization

The most accurate speaker diarization on the market, powered by pyannoteAI Precision-2 and built into one async API for meeting assistants, contact centers, and media platforms. Turn multi-speaker recordings into transcripts that show exactly who said what, in any language.

50€

Transcription credits

Test speaker diarization on 80+ hours of your own audio. No expiry, no credit card.

#1

Speaker diarization

3x fewer speaker errors than leading APIs, and the lowest diarization error rate of any provider.

100+

Languages

Speaker labels stay accurate in any language, even when people switch mid-sentence.

Trusted by over 350,000 users and 2,000+ enterprise teams
HeyGen Livestorm Adversus Selectra Circleback Recall.ai
Noota Robonote RapidSOS OVHcloud SFR Clariane

#1 on an open, reproducible benchmark

Gladia has the lowest diarization error rate (DER) of 9 providers tested on DIHARD III, the standard benchmark for real-world multi-speaker audio. Every provider was tested on identical files through its production API with default settings, and the full methodology is open-sourced so anyone can reproduce the results.

16.6%
Solaria-1
20.4%
NVIDIA NeMo
23%
pyannoteAI
33.8%
AWS
37.8%
Soniox
39.5%
ElevenLabs
43.9%
AssemblyAI
46.9%
Deepgram

*Note: DER measures the share of audio time with a diarization error: missed speech, false alarms, or speech attributed to the wrong speaker. Lower is better. More about DER on our blog.

Powered by pyannoteAI Precision-2

We pair the model from pyannote’s creators, the team behind the open-source standard for speaker diarization, with Gladia’s Solaria transcription, so you get every word with the right speaker and timestamp.

Add diarization in one parameter

Speaker diarization runs in the same async request as transcription. There’s no second model to host, no separate pipeline to orchestrate, and no timestamp reconciliation between services.

Send your recording

Upload a file or pass a URL to the async endpoint. Mono, stereo, and multichannel audio are all supported.

Turn on diarization

Set "diarization": true. Optionally pass the exact number of speakers, or a min and max range, to tighten accuracy.

Get speaker-labeled utterances

Each utterance comes back with a speaker index, text, language, and start and end timestamps.

Read the diarization docs

Use cases

Built for multi-speaker products

From sales calls to panel interviews, teams rely on Gladia to know exactly who said what, and to build summaries, QA, and analytics on top of it.

Meeting assistants

Attribute every decision and action item to the person who said it, so AI summaries and follow-ups are trustworthy.

Learn more →

Contact centers (CCaaS)

Separate agent from customer on mono recordings to power call QA, compliance checks, and coaching.

Learn more →

Media & content

Produce speaker-labeled transcripts, captions, and show notes for interviews, panels, and debates.

Learn more →

New to diarization?

Learn how speaker diarization works, what drives errors such as overlapping speech, and how accuracy is measured.

Read the guide →

We are very excited to collaborate with Gladia to integrate our advanced speaker diarization models. Together, we are pushing the boundaries of speech by making voice separation faster and more accurate than ever.

Speaker diarization included at no extra cost

Every plan comes with speaker labels built in. Your bill stays the same whether it’s a two-person call or a ten-person meeting.

Starter

Flexible pay-as-you-go for moderate audio volumes. Get started immediately.

Async at $0.61/hr

Real-time at $0.75/hr

* 50€ in free credits

Enterprise

Annual plan with custom models, fine-tuning, debundled pricing, and more.

Custom

Explore the pricing

Transcribe in minutes

Start free with 50€ in credits, or book a demo to test Gladia on your own audio.

FAQs