Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.
Migrating from Azure Speech to Gladia: a step-by-step switching guide
TL;DR: Migrating from Azure Speech to Gladia removes the overhead of custom training pipelines and fragmented per-feature billing. Azure routes diarization, translation, and sentiment through separate services with separate billing meters. We bundle all audio intelligence into one per-hour rate on Starter and Growth plans. Solaria-3 ranks #1 for real-world European business audio, Solaria-1 covers 100+ languages with native code-switching. Both deliver out-of-the-box accuracy that eliminates custom training for most production audio. Most engineering teams complete the API refactoring in under 24 hours.
How to evaluate a speech-to-text API: a technical buyer's framework
TL;DR: Choosing an STT API on vendor benchmarks alone is how teams end up with transcription that looks fine in staging and breaks on production audio. A rigorous evaluation requires building a test set from your own calls, measuring word error rate (WER) on your specific audio distribution, stress-testing latency under concurrent load, and auditing data retraining terms before signing. This guide gives you a reusable engineering blueprint to run that evaluation end-to-end, the same methodology behind our own open async benchmark, which covers 7 datasets and 74+ hours of audio across 8 providers.
The contact center QA scorecard: what to measure and how transcription feeds it
TL;DR: Manual QA teams review as little as 1% to 2% of contact center calls, leaving the vast majority of interactions unreviewed and exposing systemic compliance risks that sampling never surfaces. Scaling to automated coverage requires transcription accurate enough to power LLM-based scoring without silent failures. If your speech-to-text engine misattributes a speaker or drops a compliance disclosure, every downstream scorecard, CRM entry, and coaching flag is wrong. French CCaaS platform Gravite cut per-call review time from 15 minutes to 1 minute (93% reduction) while automating coverage across their full 50,000 hours of annual call volume on infrastructure built for real-world contact center audio.
Gladia x pyannoteAI: Speaker diarization and the future of voice AI
Published on Mar 11, 2025
Speaker recognition is advancing rapidly. Beyond merely capturing what is said, it reveals who is speaking and how they communicate, paving the way for more advanced communication platforms and assistant apps
Jean‑Louis Queguiner, CEO at Gladia, met our partners at pyannoteAI, a leading provider of cutting‑edge speaker diarization and identification models, to explore how speaker insights are continuing to evolve and contribute to better transcription accuracy and analytics.
Our speaker diarization pipeline is now powered by pyannoteAI’s Precision‑2, their most accurate model to date, bringing state-of-the-art accuracy and robustness to multi-speaker audio transcription. By integrating Precision-2, Gladia now delivers sharper speaker boundaries, better handling of overlaps, and higher consistency across languages.
Watch the webinar directly or read through a summary of key insights shared below.
Key takeaways
Gladia’s speaker diarization pipeline now uses pyannoteAI’s Precision‑2 for state‑of‑the‑art accuracy, sharper boundaries, and better overlap handling.
Speaker diarization answers “who spoke when,” enabling accurate STT, meeting notes, analytics, and voice AI.
pyannoteAI evolved from open‑source pyannote to a commercial platform with improved performance, reduced compute, and enterprise support.
Real‑time diarization, overlap detection, and speaker re‑identification are active innovation areas.
Diarization drives measurable value across meetings, sales/support, media, and regulated industries.
What is speaker diarization?
Speaker diarization is the process of identifying and segmenting different speakers in an audio recording. As Hervé Bredin, the creator of the Pyannote open‑source library, explained, it is a unique and complex machine learning problem. Unlike traditional supervised learning tasks, diarization must determine the number of speakers dynamically, cluster their voices, and handle overlapping speech.
With Pyannote, an open‑source tool widely used in the speech AI community, and pyannoteAI, a commercial product offering enhanced diarization accuracy and speed, speaker identification is becoming more accessible and reliable than ever.
The evolution of pyannoteAI
Pyannote started as an open‑source project designed to make speech research reproducible and accessible. Today, it has grown into an essential component of voice AI, with over 100,000 unique users and 30 million downloads per month on Hugging Face.
When OpenAI released Whisper in 2022, Pyannote's popularity surged as it became the go‑to tool for speaker diarization alongside Whisper’s transcription capabilities. This demand led to the creation of pyannoteAI, a commercial solution that offers improved performance, reduced computation time, and enterprise‑grade support.
With Precision‑2, pyannoteAI further improves diarization with higher accuracy, more precise speaker boundaries, stronger overlap handling, and consistency across languages—benefits that now power Gladia’s diarization pipeline. Precision‑2 also improves speaker counting and time alignment, and supports operational controls like bounding expected speaker counts for calls, clinical visits, and multi‑party meetings.
Why speaker diarization matters
Speaker diarization has a profound impact on multiple industries, including:
Speech‑to‑text & meeting notes: Companies like Circleback use pyannoteAI to accurately transcribe meetings, distinguishing different speakers for better insights.
Dubbing & localization: pyannoteAI helps streamline the dubbing process by ensuring the right voice is assigned to the correct character.
Voice AI training: AI models, such as Moshy by QAI, leverage Pyannote for clean, speaker‑separated datasets, ensuring higher accuracy in voice recognition systems.
Sales & customer support: Diarization plays a crucial role in call analytics and CRM integrations, ensuring that customer interactions are correctly attributed.
Healthcare & legal transcription: Misattributed speech can have critical consequences. In medical settings, diarization ensures accuracy in doctor‑patient interactions.
Challenges & innovations in speaker diarization
Despite its benefits, speaker diarization is one of the hardest problems in machine learning due to several challenges:
Handling overlapping speech: Real‑world conversations involve interruptions and overlaps. pyannoteAI has made significant progress in detecting and distinguishing overlapping speakers.
Real‑time diarization: While offline diarization is well‑optimized, real‑time processing is still evolving. pyannoteAI is actively developing a streaming solution to power live captioning and voice assistants.
Speaker re‑identification: Gladia is experimenting with speaker tracking across multiple recordings using embedding‑based recognition, allowing seamless continuity in multi‑session interactions.
Background noise & audio quality: Background noise, music, and different audio formats impact diarization accuracy. pyannoteAI continuously improves robustness against such factors.
The future of speaker insights
Audio intelligence is advancing, with speaker diarization playing a critical role. Jean‑Louis from Gladia highlighted that every day, people generate the equivalent of a Tolkien book in spoken words. Unlocking insights from this vast data pool requires more than just transcription—it demands accurate speaker identification, emotion detection, and contextual understanding.
Key trends shaping the future of speech AI
Voice agents: AI‑powered voice agents will revolutionize customer service, sales, and virtual assistants by providing real‑time, speaker‑aware responses.
Prosody & emotion recognition: Understanding not just words, but how they are spoken, will enhance AI interactions.
Non‑speech vocalization: Detecting laughter, sighs, and hesitations will add another layer of intelligence to voice AI.
AI‑powered personalization: AI systems will tailor interactions based on voice traits, improving accessibility and user experience.
Final thoughts
Speaker diarization is no longer just a niche problem; it is essential for the future of voice intelligence. Companies like Gladia and pyannoteAI are pushing the boundaries of what’s possible, making voice AI more accurate, efficient, and insightful.
As voice technology continues to evolve, speaker insights will become just as valuable as the words themselves. Whether it’s improving customer service, enhancing transcription accuracy, or creating lifelike AI assistants, diarization will be at the heart of the voice AI revolution.
If you're interested in integrating pyannoteAI or Gladia’s solutions into your workflow, now is the time to explore the possibilities!
Want to learn more? Watch the full webinar recording and reach out to Gladiaor pyannoteAI for partnership opportunities.
Contact us
Your request has been registered
A problem occurred while submitting the form.
Read more
Speech-To-Text
Migrating from Azure Speech to Gladia: a step-by-step switching guide
Speech-To-Text
How to evaluate a speech-to-text API: a technical buyer's framework
Speech-To-Text
The contact center QA scorecard: what to measure and how transcription feeds it
From audio to knowledge
Subscribe to receive latest news, product updates and curated AI content.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.