Pricing
Get started
Get started

Blog

Technical guides, customer stories, and product updates
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

Speech-To-Text

Automatic Speech Recognition (ASR): how speech-to-text models work and which one to use

TL;DR: Automatic speech recognition (ASR), or speech-to-text, converts spoken audio into text using one of five architecture families: encoder-decoder, CTC, encoder-transducer, and speech LLMs with continuous or discrete input. There is no single best ASR model. The right one depends on your accuracy target, real-time vs batch needs, language coverage, and cost and deployment constraints. This guide, based on our Speech Engineering Lead Bruno Hays's analysis, explains how each architecture works and how to weigh those trade-offs. Our own Solaria-1 and Solaria-3 sit in this landscape too, built for full language breadth and for real-world European business audio respectively.

Speech-To-Text

OpenAI Whisper vs Google STT vs Amazon Transcribe: the ASR rundown (2026 edition)

TL;DR: OpenAI's gpt-4o-transcribe, Google's Chirp 3, and Amazon Transcribe have each closed real gaps since earlier versions, but they specialize differently. OpenAI leads on clean-audio accuracy, Google on language breadth and native diarization, Amazon on call-center and medical vertical tooling. Published WER numbers barely separate the three on clean, single-speaker audio anymore, so the deciding factor is usually what's already in your stack and what use case you're solving for. The real gap shows up on conversational, multi-speaker, accented, or code-switched audio, conditions the standard benchmark tables don't test. We built Solaria-1 specifically for multi-speaker, multilingual, real-world audio with native code-switching, and Solaria-3, our newest speech model for real-world business audio.

Speech-To-Text

How do speech recognition models work?

TL;DR: Speech recognition, or automatic speech recognition (ASR), turns spoken audio into text through a pipeline: it captures and digitises the signal, extracts acoustic features, then predicts words using acoustic and language models. Traditional ASR chained separate acoustic, lexicon, and language models; modern end-to-end models replace that stack with a single sequence-to-sequence neural network that is more accurate and multilingual. Which architecture fits depends on your accuracy target, languages, noise conditions, and latency budget.

Speech-To-Text

Best ASR engines and the models powering them: a review

TL;DR: Automatic speech recognition (ASR) is now a crowded field of transformer-based engines, and the right one depends on your languages, your audio conditions, and whether you need audio intelligence on top of the transcript. This review walks through the major engines and the models powering them, Whisper, Google USM, Amazon, AssemblyAI, Deepgram, Speechmatics, and our own Solaria models, with the strengths and limitations of each. The short version: benchmark on your own audio, weigh accuracy against speed and language coverage, and do not trust headline language counts without testing. On real-world European business audio, our Solaria-3 model ranks #1 against AssemblyAI, ElevenLabs, Deepgram, Mistral, and Speechmatics.

Tutorials

How to build a Google Meet Bot for recording and video transcription

TL;DR: Building a Google Meet bot means solving two problems: getting the audio out of the meeting with a headless-browser bot, and turning that audio into structured transcripts with a speech-to-text API. The integration takes under a day; teams lose weeks evaluating the wrong things, like headline rates instead of total cost at scale and English word error rate instead of multilingual accuracy. This guide covers the Playwright capture code, the WebSocket streaming code, and the commercial architecture you need to model costs and ship in under a week.

Speech-To-Text

MiFID II and FCA call recording: compliance for voice transcription in finance

TL;DR: Financial firms operating under MiFID II and FCA jurisdictions must maintain searchable, high-accuracy records of all client-related voice communications, including remote and mobile calls under FCA Market Watch 66. Engineering and product teams building for these obligations typically implement dedicated cloud infrastructure with defined data residency, controls that prevent customer audio from being used to retrain models on Growth and Enterprise plans, and transcription accurate enough that the resulting records hold up under regulatory review. Transcription and speaker attribution errors are not product quality issues in this context. They are audit trail failures, and regulators treat them as such.

Speech-To-Text

Why speech-to-text accuracy matters upstream of your LLM

TL;DR: Downstream LLM performance is ceiling-bounded by upstream transcription accuracy. A transcript with a meaningful error rate doesn't produce proportionally degraded summaries or CRM entries. It produces outputs where hallucinated names, inverted logic, and misattributed speaker turns compound silently into every downstream system that reads them. Prompt engineering cannot recover information the STT layer never captured. Gravite cut call quality review time by 93%, from 15 minutes to 1 minute per call, once transcript accuracy was high enough to trust the output without manual verification.

Speech-To-Text

Telephony-audio robustness: why 8kHz narrowband calls break generic STT

TL;DR: Telephony audio is constrained to 8kHz narrowband frequencies, stripping away the high-frequency spectral energy that generic 16kHz STT models require. Standard upsampling cannot recover phonemes that were never captured at the source, leading to transcription errors that compound silently into downstream NLU failures. Solaria-3 ranks #1 on Switchboard, the most challenging conversational telephony dataset, ahead of AssemblyAI, ElevenLabs, Deepgram, Mistral, and Speechmatics, ensuring your downstream LLM pipelines receive clean, structured data from real-world noisy call audio.

Product News

Gladia to join OVH Groupe, accelerating Europe's sovereign voice AI ambitions

PARIS, France — Sept 21, 2026 — Gladia, the AI audio infrastructure provider trusted by more than 300,000 developers and 2,000 enterprise customers worldwide, announced July 31st, that it has been acquired by OVH Groupe, the parent company of OVHcloud and Europe's leading sovereign cloud provider.