API Comparison Table

Heading 1

Heading 2

Heading 3

Heading 4

Heading 5
Heading 6

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Block quote

Ordered list

  1. Item 1
  2. Item 2
  3. Item 3

Unordered list

Text link

Bold text

Emphasis

Superscript

Subscript

Pricing
Get started
Get started

Read more

Speech-To-Text

Microsoft Teams transcription via API

TL;DR: Native Microsoft Teams transcription via the Graph API delivers transcripts only after a meeting ends, provides utterance-level (not word-level) timestamps, and degrades sharply on accented or multilingual speech. For any product that routes Teams audio to downstream AI systems, the more reliable architectural pattern is capturing raw audio via a custom WebRTC bot and routing it to a managed STT engine, one that delivers word-level timestamps, accurate multilingual handling, and predictable per-hour costs, none of which the native Graph API provides. Solaria-1 covers real-time streaming and broad language support. Solaria-3 is optimised for European business audio.

Speech-To-Text

Podcast transcription at scale: an API workflow for media platforms

TL;DR: Podcast audio is invisible to search without accurate, word-level transcripts, and transcription quality sets the ceiling for everything downstream, from content discovery to AI-generated show notes. A production-grade async pipeline (decoupled webhook ingestion, pyannoteAI-powered diarization, word-level timestamps) is what separates a searchable audio library from a title-and-description catalog. At 10,000 hours monthly, a managed API costs $2,000–$6,100 depending on plan.

Speech-To-Text

HIPAA-ready meeting assistants for healthcare and therapy sessions

TL;DR: Building a HIPAA-ready meeting assistant requires more than a generic transcription wrapper. Any API that processes Protected Health Information on your behalf must sign a Business Associate Agreement (BAA) before PHI flows to it, and transcription accuracy matters more than most teams expect: word error rate can more than double in noisy, multi-speaker clinical environments compared to controlled recordings, meaning errors compound into every SOAP note and EHR entry downstream. This guide covers the BAA requirements, encryption controls, and unit economics product teams need to evaluate before committing to an audio infrastructure provider for clinical or therapy use cases.

March 2023 Roadmap its Speech-to-Text API: Speaker Diarization, Word-Level Timestamps and more

Published on Jun 2, 2023
March 2023 Roadmap its Speech-to-Text API: Speaker Diarization, Word-Level Timestamps and more

A glimpse into Gladia's roadmap for its Speech-to-Text API, starting with speaker diarization. We’re incredibly excited to be building our Audio Intelligence product in a community-led way, delivering a holistic final product adapted to the many needs and use cases brought to our attention.

Following Gladia’s Speech-to-Text AI alpha release two weeks ago, we’ve received dozens of new feature requests from the alpha users, to make our core real-time audio transcription API even more exciting and versatile.

We heard you and are happy to announce that the API is growing more robust by the minute and is now available with more capabilities — on top of its blazing speed and top-tier output quality.

We’re incredibly excited to be building our Audio Intelligence product in a community-led way, delivering a holistic final product adapted to the many needs and use cases brought to our attention.

Here’s what we have in store already

Speech-to-Text (STT) Transcription

Setting a new standard for the industry, our STT API is build on OpenAI’s Whisper and can transcribe audio in 10s/h at 3.52%WER. Tested and approved by thousands of alpha users across a range of use cases (e.g. call center, virtual meetings, YouTube videos, podcasts).

Speech-to-Text Translation

Upload your file, select an output language of your choice, and enjoy the final translated transcript free of errors. Currently available in 99 languages, and counting. If your language is not supported yet, drop us a message in this Twitter thread.

Transcription from YouTube URL

Drop a video URL and enjoy a highly accurate output file (.srt or JSON) that can be used as an alternative to YouTube’s auto-captions to improve the viewer’s experience on your channel. Transcription as subtitles file (.srt) will become available shortly too.

And here is the list of new most anticipated features we’re planning to release in March.

Speaker Diarization

You will now be able to automatically identify and recognize all speakers mixed in a single audio or video stream, including when multiple languages are used.

Word-Level Timestamps

A feature enabling Gladia users to produce a highly accurate JSON transcript with time stamps at every word.

Live-Streaming Transcription

We’re adding the ability to transcribe speech in real-time, using your microphone.

We’re preparing a series of deep dives on some of these new features to showcase how our tech works behind the scenes. Stay tuned!

As always, feel free to test the API and give us your feedback  Discord. We truly love iterating with the community.

Contact us

280
Your request has been registered
A problem occurred while submitting the form.

Read more