API Comparison Table

Heading 1

Heading 2

Heading 3

Heading 4

Heading 5
Heading 6

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Block quote

Ordered list

  1. Item 1
  2. Item 2
  3. Item 3

Unordered list

Text link

Bold text

Emphasis

Superscript

Subscript

Pricing
Get started
Get started

Read more

Speech-To-Text

Microsoft Teams transcription via API

TL;DR: Native Microsoft Teams transcription via the Graph API delivers transcripts only after a meeting ends, provides utterance-level (not word-level) timestamps, and degrades sharply on accented or multilingual speech. For any product that routes Teams audio to downstream AI systems, the more reliable architectural pattern is capturing raw audio via a custom WebRTC bot and routing it to a managed STT engine, one that delivers word-level timestamps, accurate multilingual handling, and predictable per-hour costs, none of which the native Graph API provides. Solaria-1 covers real-time streaming and broad language support. Solaria-3 is optimised for European business audio.

Speech-To-Text

Podcast transcription at scale: an API workflow for media platforms

TL;DR: Podcast audio is invisible to search without accurate, word-level transcripts, and transcription quality sets the ceiling for everything downstream, from content discovery to AI-generated show notes. A production-grade async pipeline (decoupled webhook ingestion, pyannoteAI-powered diarization, word-level timestamps) is what separates a searchable audio library from a title-and-description catalog. At 10,000 hours monthly, a managed API costs $2,000–$6,100 depending on plan.

Speech-To-Text

HIPAA-ready meeting assistants for healthcare and therapy sessions

TL;DR: Building a HIPAA-ready meeting assistant requires more than a generic transcription wrapper. Any API that processes Protected Health Information on your behalf must sign a Business Associate Agreement (BAA) before PHI flows to it, and transcription accuracy matters more than most teams expect: word error rate can more than double in noisy, multi-speaker clinical environments compared to controlled recordings, meaning errors compound into every SOAP note and EHR entry downstream. This guide covers the BAA requirements, encryption controls, and unit economics product teams need to evaluate before committing to an audio infrastructure provider for clinical or therapy use cases.

AI-powered healthcare assistant enhances medical transcription by 120% with Gladia

Published on Feb 28, 2025
AI-powered healthcare assistant enhances medical transcription by 120% with Gladia

Medical transcription is among the most critical and challenging verticals for ASR models to date.

Filled with drug names and medical jargon, medical consultations, dictations, and online conferences require versatile solutions, with custom vocabulary and specialized models needed to make speech-to-text solutions attuned to jargon. There’s the issue of security too, as audio from medical consultations is among the most sensitive confidential data out there.

A fast-growing healthcare generative AI startup, who prefers to remain anonymous, turned to Gladia for top-quality medical transcription at scale. Here’s how we helped them increase their accuracy and speed of transcription, all while ensuring 100% security of confidential user data.

Challenge

Doctors spend about 60% of their time on computers, doing non-clinical work. This startup is aiming to get that number to 15%, enabling doctors to allocate most of their time for consultation, diagnostics, and other high-value tasks with the help of AI.

They knew that having accurate transcription for note-taking during consultations was the first step in designing a holistic solution to achieve this milestone.

Indeed, the platform’s ability to understand and actively transcribe jargon-filled medical conversations is an essential prerequisite for LLM-powered notes, prescriptions, and intricate EHR enrichment that distinguish their AI co-pilot.

Speed is likewise a key factor for them, as the ability to generate notes shortly after the consultation is critical for efficient clinical workflows.

Moreover, they needed to ensure 100% protection of all user data in accordance with HIPAA and GDPR, which most of the US-based providers are generally not able to provide.

This is why their team took the task of choosing a speech-to-text provider very seriously. With regular evaluations in place, they have tested over 7 different providers before, including the Big Tech cloud solutions — all of which ultimately failed to strike the right balance between accuracy, speed, price, and security standards.

Solution

With Gladia, the team was able to implement:

Impact

Following a swift onboarding with our tech team, they began to use Gladia as its primary speech-to-text provider. The results did not take long to show.

By working with the Gladia team to iterate and scale up, they saw a noticeable impact on their system’s performance:

__wf_reserved_inherit
__wf_reserved_inherit

The team was likewise impressed by the quality of Gladia’s technical assistance, allowing them to not only set up their dedicated environment in a matter of hours but also benefit from Gladia’s in-house engineering expertise to optimize their infrastructure as a whole.

Given the initial success with Gladia API and its on-premise deployment, this innovative company is already considering how they will leverage our product in the future as they extend their platform to new stakeholders.

For instance, they look forward to experimenting more with multilingual transcription and translation, which would enable patients to consult physicians in their native language. They also intend to leverage speaker diarization for collective medical meetings.

About Gladia

Gladia provides a speech-to-text and audio intelligence API for building virtual meeting and note-taking apps, call center platforms, and media products, providing transcription, translation, and insights powered by best-in-class ASR, LLMs and GenAI models.

Having read this case study, do you feel like Gladia could be the right fit for your business too?

Don't hesitate to contact our sales team to explore this in more detail, and follow us on X and LinkedIn.

Contact us

280
Your request has been registered
A problem occurred while submitting the form.

Read more