Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.
TL;DR: Your transcription model might achieve a 5% Word Error Rate, but your meeting summaries can still be completely unreliable if Diarization Error Rate (DER) spikes. DER is the metric that determines whether your system correctly identifies who spoke each word, measured as the sum of three error types: Missed Speech, False Alarm, and Speaker Confusion. For production multi-speaker pipelines, a DER below 15% is the threshold for reliable speaker-labeled analytics; below 10% is the target for clean audio with controlled conditions, such as high-quality meeting assistant output. Accurate speaker attribution directly determines the reliability of downstream LLM summaries and CRM data. Our async pipeline, powered by pyannoteAI's Precision-2 model, delivers up to 3x lower DER than alternatives on conversational speech.
Latency benchmarks for streaming speech-to-text (TTLB and P99)
TL;DR: A voice agent with a 150ms average STT latency sounds fast in a slide deck, but if its P99 spikes to 1.2 seconds, one in every hundred conversational turns breaks. This piece maps the full end-to-end streaming latency budget (network Round Trip Time (RTT), audio buffering, model inference, Voice Activity Detection (VAD) endpointing), explains why P99 and TTLB are the metrics that matter most for production user experience, and shows how to build a reproducible test harness. We cover how our Solaria-1 model delivers first partials under 103ms and final transcripts around 300ms, backed by an open, reproducible benchmark methodology.
Multi-tenant, white-label speech-to-text for platforms
TL;DR: Building a compliant multi-tenant STT layer requires strict data isolation at the key level, granular cost attribution per tenant, and contractual infrastructure guarantees that flow through to your own SLAs. Self-hosting open-source models introduces DevOps overhead, scaling unpredictability, and the absence of built-in tenant isolation features that a managed API provides by default. Managed infrastructure with per-client keys, certified data handling, and all-inclusive pricing removes most of that build cost ($0.20–$0.61/hr with diarization, translation, and entity recognition included, compared with $240K–$480K/yr in dedicated engineering to self-host) but the isolation and attribution architecture still has to be designed correctly regardless of which vendor provides it.
Gladia and Pipecat partner to push the boundaries of real-time voice AI
Published on May 14, 2025
We’re thrilled to announce a strategic partnership between Gladia and Daily, the team behind Pipecat, aimed at revolutionizing real-time conversational AI. This collaboration combines our cutting-edge audio intelligence capabilities with their flexible 100% open-source framework, empowering developers to create more dynamic, multilingual, and context-aware voice AI applications.
We’re thrilled to announce a strategic partnership between Gladia and Daily, the team behind Pipecat, aimed at revolutionizing real-time conversational AI. This collaboration combines our cutting-edge audio intelligence capabilities with their flexible 100% open-source framework, empowering developers to create more dynamic, multilingual, and context-aware voice AI applications.
Pipecat is a vendor-neutral framework designed to simplify the creation of voice and multimodal conversational agents. It allows developers to orchestrate LLM models and AI services effortlessly, enabling the development of video and voice applications such as personal coaches, meeting assistants, and customer support bots.
Pipecat is maintained by Daily with the support of the global developer community. Daily is a leader in developer tooling and global WebRTC infrastructure since 2016. Earlier this year Daily announced Pipecat Cloud, the first open source voice AI cloud.
About the partnership
At Gladia, we believe the future of human-AI interaction lies in systems that understand and respond in real-time, just like humans do. Pushing the boundaries of ultra low latency conversational AI is key to bridging the gap between humans and machines, enabling more natural, intuitive communication. In today’s world, where seamless interactions are crucial, having AI that can understand diverse languages and contexts is essential for real collaboration in customer support, meetings, and beyond.
This partnership with Pipecat goes beyond technology—it empowers developers to easily create intelligent, adaptable, multilingual voice AI applications that break down barriers and foster meaningful interactions. By combining Gladia's language processing with Pipecat's flexible framework, we can enable the creation of robust voice platforms that meet the needs of a wide range of use cases.
A shared vision for the future
This partnership is more than just a technical integration; it's a shared commitment to pushing the boundaries of what's possible in real-time conversational AI. In the words of Daily's co-founder:
Jean-Louis Queguiner, CEO of Gladia, also shared his excitement for the partnership: "At Gladia, we believe in pushing the boundaries of what's possible in real-time conversational AI. Partnering with Pipecat allows us to extend that vision even further—combining our advanced language processing capabilities with Pipecat’s open-source platform to help developers create truly innovative, scalable voice AI solutions. This collaboration is about more than just technology; it's about shaping the future of human-AI interaction."
What this means for developers
Developers can now leverage the combined strengths of Pipecat and Gladia to build more sophisticated voice AI applications. Whether you're creating a multilingual customer support bot, a real-time meeting assistant, or an interactive storytelling agent, this partnership provides the tools and flexibility needed to bring your vision to life.
To get started, visit pipecat.ai to explore the framework and sign up to the Gladia Playground to try first-hand our newest STT model, Solaria.
Stay tuned for more updates as we continue to innovate and expand the possibilities of real-time conversational AI.
Contact us
Your request has been registered
A problem occurred while submitting the form.
Read more
Speech-To-Text
Diarization error rate (DER) explained
Speech-To-Text
Latency benchmarks for streaming speech-to-text (TTLB and P99)
Speech-To-Text
Multi-tenant, white-label speech-to-text for platforms
From audio to knowledge
Subscribe to receive latest news, product updates and curated AI content.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.