API Comparison Table

Heading 1

Heading 2

Heading 3

Heading 4

Heading 5
Heading 6

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Block quote

Ordered list

  1. Item 1
  2. Item 2
  3. Item 3

Unordered list

Text link

Bold text

Emphasis

Superscript

Subscript

Pricing
Get started
Get started

Read more

Speech-To-Text

MiFID II and FCA call recording: compliance for voice transcription in finance

TL;DR: Financial firms operating under MiFID II and FCA jurisdictions must maintain searchable, high-accuracy records of all client-related voice communications, including remote and mobile calls under FCA Market Watch 66. Engineering and product teams building for these obligations typically implement dedicated cloud infrastructure with defined data residency, controls that prevent customer audio from being used to retrain models on Growth and Enterprise plans, and transcription accurate enough that the resulting records hold up under regulatory review. Transcription and speaker attribution errors are not product quality issues in this context. They are audit trail failures, and regulators treat them as such.

Speech-To-Text

Why speech-to-text accuracy matters upstream of your LLM

TL;DR: Downstream LLM performance is ceiling-bounded by upstream transcription accuracy. A transcript with a meaningful error rate doesn't produce proportionally degraded summaries or CRM entries. It produces outputs where hallucinated names, inverted logic, and misattributed speaker turns compound silently into every downstream system that reads them. Prompt engineering cannot recover information the STT layer never captured. Gravite cut call quality review time by 93%, from 15 minutes to 1 minute per call, once transcript accuracy was high enough to trust the output without manual verification.

Speech-To-Text

Telephony-audio robustness: why 8kHz narrowband calls break generic STT

TL;DR: Telephony audio is constrained to 8kHz narrowband frequencies, stripping away the high-frequency spectral energy that generic 16kHz STT models require. Standard upsampling cannot recover phonemes that were never captured at the source, leading to transcription errors that compound silently into downstream NLU failures. Solaria-3 ranks #1 on Switchboard, the most challenging conversational telephony dataset, ahead of AssemblyAI, ElevenLabs, Deepgram, Mistral, and Speechmatics, ensuring your downstream LLM pipelines receive clean, structured data from real-world noisy call audio.

Automatic Speech Recognition (ASR): how speech-to-text models work and which one to use

Published on Sep 30, 2026
by Anna Jelezovskaia
Automatic Speech Recognition (ASR): how speech-to-text models work and which one to use

TL;DR: Automatic speech recognition (ASR), or speech-to-text, converts spoken audio into text using one of five architecture families: encoder-decoder, CTC, encoder-transducer, and speech LLMs with continuous or discrete input. There is no single best ASR model. The right one depends on your accuracy target, real-time vs batch needs, language coverage, and cost and deployment constraints. This guide, based on our Speech Engineering Lead Bruno Hays's analysis, explains how each architecture works and how to weigh those trade-offs. Our own Solaria-1 and Solaria-3 sit in this landscape too, built for full language breadth and for real-world European business audio respectively.

ASR, also known as speech-to-text (STT) technology, is a constantly evolving field. Knowing which ASR model is right for your product or service can be challenging. The main architecture families, CTC, encoder-decoder, transducer, and speech LLMs, each carry distinct trade-offs. What does it all mean? And what do you choose?!

What is ASR?

ASR is a technology that intelligently recognizes human speech and converts it into written text. It's the foundation behind voice assistants, transcription tools, and real-time communication solutions. In short, ASR and STT describe the same thing: turning spoken audio into text.

When incorporating speech recognition technologies into your business and customer workflows, the more you know, the better you'll be able to select an ASR model that's right for your specific requirements.

So let's start with some key terminology. These are fundamental terms we use when discussing ASR models.

  • Text tokens: Text tokens are slices of sentences created by cutting text at a character, word, or sub-word level using algorithms like BPE (Byte Pair Encoding). BPE segments words at meaningful boundaries, such as separating prefixes and suffixes, making it more efficient than character-level tokenization.
  • Embeddings: Embeddings are vectors representing concepts. While humans use words, these are the language of the model. Textual embedding tables serve as dictionaries to translate tokens into embeddings that models can understand. Embeddings start as random vectors at the beginning of training and get optimized to store useful information during the training process.
  • Attention mechanism: Attention is a method for handling sequential data like text or audio by processing each token separately while adding contextual information. Each token passes through multiple encoder blocks, where it gets refined using context from previous tokens. The refinement process goes like this: first, the input embedding derived from the embedding table has no context. Then each encoder block adds contextual information, creating progressively better embeddings.
    This process not only increases resolution, but the model augments the vector, effectively 'sculpting' a generic token into a highly specialized representation of that word within its specific environment. Stacking multiple encoder blocks creates higher-level concepts at each step. Early blocks might identify word relationships, later blocks might identify sentiment or task-specific features.
  • Encoders as audio processors: Encoders are the "ears" of ASR models that transform raw audio waveforms into meaningful sequences of embeddings for the ASR task. Most encoders use Transformer architecture based on the attention mechanism. Audio doesn't need an embedding table like text does because spectrograms already provide vectors. Slices of spectrograms can be used directly as embeddings.
    Conceptually, encoder output might represent phonemes (like "sh" or "t" sounds) rather than raw audio slices.

Modern ASR models

Legacy ASR systems first appeared on the scene in the 1970s. Early models had poor ASR performance due to inadequate adapters connecting the audio encoder to the LM.

Modern ASR models take a different approach to speech processing with end-to-end deep learning. Where legacy systems trained the language, acoustic, and lexicon models separately, modern models train them together as a single neural network.

These improvements deliver several key benefits: they reduce development time and costs, improve accuracy and support multiple languages thanks to their advanced neural architecture, and minimize latency while drastically improving overall performance and accuracy.

Current ASR model design

All modern ASR architecture models have two components that need to work in harmony to succeed:

  1. An encoder to understand audio (the "ears").
  2. A language model to produce sensible text (the "brain").

Conceptually, the encoder turns the raw audio into a sequence of phoneme-like embeddings. For example "maaaï neme iz bonde". The language model, in theory, converts it into a coherent sentence: "My name is Bond."

When comparing different model architectures, the key difference is how the encoder and language model interact with each other. For a deeper walkthrough of that mechanism, see our guide on how speech recognition models work.

The 5 primary modern architecture models

There are five main ASR architecture families: encoder-decoder (Whisper, Canary), CTC (Wav2Vec2), encoder-transducer (Parakeet TDT), speech LLMs with continuous input (Voxtral), and speech LLMs with discrete input (GPT-4o, Moshi).

1. Encoder-decoder

Encoder-decoder models, like Whisper, use a separate decoder to generate text token by token. The decoder relies on self-attention to see the tokens it has already generated and cross-attention to pull audio information from the encoder. That combination means each new token benefits from language modeling, since it can see the start of the sentence, while staying grounded in the actual audio through cross-attention.

2. CTC architecture

CTC forces the encoder to output letters or tokens directly from each audio slice, using an alignment trick that allows repeated letters. For every slice, the encoder produces a probability distribution over the vocabulary, for example an 80% chance of T, 15% P, and 5% S. Greedy decoding just takes the highest-probability letter for each slice, but it performs poorly without language modeling. Adding a language model re-ranks the proposed letters by likelihood, which significantly improves accuracy, and beam search takes this further by letting the language model effectively "see the future," keeping multiple possible paths in memory before committing to one. Wav2Vec2 is the audio encoder family most widely used with CTC decoding.

3. Encoder-transducer

An encoder-transducer is best described as a "disk and read head" system, where the encoder output is the disk and the joiner is the reading head. The joiner reads encoder embeddings one by one, asks the language model, called the projector, what word should come next, then outputs a token or nothing before moving to the next embedding. That design makes transducers streamable by nature, though harder to batch effectively than other architectures. Parakeet TDT is a transducer model built with an optimization that makes decoding significantly faster.

4. Speech LLMs with continuous input

Speech LLMs with continuous input add "ears," an audio encoder, to a pre-trained text LLM that already has strong language modeling capabilities, the "brain." The pipeline pairs a pre-trained encoder, like Whisper's, with a pre-trained LLM, like Gemma, connected by a trainable adapter. That adapter transforms audio embeddings into word-like embeddings the LLM can understand, so the LLM never has to learn how to process audio directly. Voxtral by Mistral and Qwen Audio are both built this way. What the model outputs depends entirely on how it was trained and prompted: transcription, topic analysis, emotion detection, or any other audio understanding task.

5. Speech LLMs with discrete input

Speech LLMs with discrete input rely on discrete audio tokens, compressed numerical representations of audio that let the LLM process and generate audio the same way it processes text. This is currently the most effective and popular method for building speech-to-speech LLMs that need to both understand and generate audio. The LLM can output interlaced text and speech tokens, with the speech tokens decoded into actual sound through a separate speech decoder. GPT-4o (most likely), Moshi, and Kimi Audio are all believed to use this approach.

Advantages and disadvantages

When weighing up the pros and cons, Bruno Hays argues that speech LLMs and encoder-decoders are fundamentally the same mathematically, and therefore provide similar results. Encoder-decoders use cross-attention for audio access while speech LLMs use self-attention, but both approaches are mathematically equivalent. In practice, both reach very similar ASR performance on leaderboards.

However, the real difference is the training approach:

  • Encoder-decoders train the encoder and decoder together on audio data.
  • Speech LLMs train them separately, then teach them to work together.

Popular ASR architecture models, compared

It's important to stress that there is no one-size-fits-all answer to which architecture model you should use. However, based on Bruno's findings, we introduce the most popular ASR model families at the top of the ASR architecture leaderboard. We spotlight the functionality and highlights of each, so you can make an informed decision on what is best for your needs.

Wav2Vec2

Wav2Vec2, developed by Facebook in 2020, was the go-to ASR model from 2020 to 2022, before Whisper took over the conversation. It applies BERT's masked language modeling approach to audio: removing random slices of the signal and training the model to predict what's missing, the same self-supervised trick that let BERT learn from raw text without labeled examples. That pretraining produces a general-purpose encoder that doesn't need much task-specific data to become useful, since the model has already learned the structure of speech before it ever sees a transcript. Teams can fine-tune it for ASR with comparatively little labeled audio, for example by adding a CTC decoding head on top, which is why Wav2Vec2 became the reference architecture for the CTC decoding covered earlier in this guide.

Follow-up models Hubert and WaveLM pushed the same self-supervised idea further, refining how the model learns structure from unlabeled audio and improving downstream accuracy. Wav2Vec2 itself hasn't stood still either: in mid-November 2025, Facebook released an omnilingual version supporting around 2,000 languages, extending the architecture into exactly the low-resource, long-tail language territory where labeled training data is hardest to come by.

Whisper

OpenAI developed Whisper in late 2022 as, on paper, a standard family of encoder-decoder models, the same architecture explained earlier in this guide: a decoder generating text token by token, grounded in audio through cross-attention. What set Whisper apart wasn't the architecture itself but the training approach. The real innovation was proving that encoder-decoder models can train on much noisier data than CTC allows. Because the encoder and decoder aren't tightly coupled, encoder-decoder handles non-standardized text, like "$" for dollars or inconsistent punctuation, far better than CTC does, and that tolerance for messy transcripts matters in practice: the same cleanup effort that produces 1,000 hours of usable audio for CTC can produce 1 million hours for encoder-decoder, a difference in scale, not just convenience.

OpenAI put that difference to work by scraping YouTube and training Whisper on 700,000 hours of audio paired with human-written subtitles, a dataset that would have been far too noisy and inconsistent for a CTC model to learn from. The result was more robust than any alternative available at the time, with multilingual and translation capabilities that outperformed the competition. That combination of scale and openness is a large part of why Whisper became the default reference point for encoder-decoder ASR, and why so many newer architectures still get benchmarked against it.

Kyutai-STT

Kyutai built its Kyutai-STT model family in 2025 around delayed streams modeling, a technique it originally developed for its Moshi voice assistant and adapted here into an audio LLM purpose-built for real-time interaction. The problem it solves is a familiar one: traditional voice assistants have to wait for a full sentence to finish before they can understand the context and start speaking, which creates the awkward pause anyone who has used a voice assistant has experienced. Delayed streams modeling gets around that by processing audio and text in parallel with a slight, deliberate delay, just enough for the model to "peek" at incoming information so it can start generating a high-quality response before the speaker has finished talking, without giving up the context a full sentence would have provided.

In production, that architecture handles 400 concurrent real-time streams on a single H100, and because it's streamable and batchable by design, it doesn't force teams to trade off latency against throughput the way some real-time architectures do. It also supports text-to-speech, making it a candidate for both directions of a voice interaction, not just transcription. The 1B model pushes latency further with semantic VAD, voice activity detection that recognizes when a speaker has actually finished a thought rather than just paused for breath, with no added delay. That brings end-of-turn latency down to as low as 0.125 seconds, using what Kyutai calls the "flush trick": forcing the model to output whatever it's holding in its buffer instead of waiting for more context that may never come.

Nemotron-Speech-Streaming-En-0.6B

NVIDIA released Nemotron-Speech-Streaming-En-0.6B in 2026, an English-only encoder-transducer model built around a Cache-Aware FastConformer encoder, an architectural twist on the streamable transducer design covered earlier in this guide. Most streaming encoders have to reprocess overlapping chunks of audio as new frames arrive, which wastes compute and adds latency. The Cache-Aware design avoids that by processing audio frame by frame and caching what it has already computed, so each new frame only adds the work strictly necessary to account for it. Paired with an inherently streamable transducer decoder, that efficiency translates directly into real-time transcription with minimal latency, without the redundant computation that would otherwise cap how many streams a single deployment can handle.

The model also gives teams dynamic runtime flexibility they don't typically get with other architectures: the latency-accuracy trade-off can be adjusted at inference time, without retraining a new model for every use case. Configurable chunk sizes as low as 80 or 160 milliseconds support near-instant interaction for use cases like live captioning or voice agents, while chunk sizes up to 1.12 seconds trade some responsiveness for higher accuracy, useful when the product can tolerate a slightly longer pause in exchange for a cleaner transcript. Because none of that flexibility comes at the cost of redundant computation, the architecture scales efficiently for production, supporting a high volume of concurrent streams without the infrastructure overhead that would otherwise come with it.

How to select the right ASR model

We've put together a few decisive factors you should consider to help you in the selection process.

  • Word error rate (WER) is the ideal goal of any voice recognition application is to achieve zero error rates. However, practical considerations dictate variations beyond our control, so factor in the precision and accuracy you need in your system when selecting an ASR model. If accuracy is the deciding factor, benchmark the shortlisted models on a sample of your own audio rather than trusting published numbers, since WER shifts with accent, noise, and domain. Our explainer on what word error rate measures covers how to read those numbers.
  • End goals consider the requirements of your end users when choosing a model. How will they use the product or service? A post-call meeting summary can tolerate a few seconds of processing for higher accuracy, while a live voice agent cannot.
  • Input audio type factors in how varied your input audio will be and the languages and dialects the model will need to support. Clean read-speech, noisy call-center audio, and multilingual conversations with mid-sentence language switches each favour different architectures.
  • Monitor performance. Every model performs differently, so you'll need to evaluate it based on your specific benchmarks. If you need real-time speech-to-text conversions (such as in smart devices and wearables), choose a streamable architecture (encoder-transducer or a streaming speech LLM) with the lowest latency. For batch workloads like transcription and post-call analysis, an encoder-decoder or async pipeline usually wins on accuracy.

Deployment and cost: decide early whether you will self-host an open model or call a managed API. Self-hosting an open model like Whisper removes per-hour vendor fees but adds infrastructure, scaling, and DevOps overhead, so the total cost of ownership can exceed a managed API. Managed APIs typically price per hour of audio processed rather than per token or per request, which makes budgeting more predictable, our own Starter, Growth, and Enterprise tiers follow that model. Our Growth plan starts as low as $0.20/hr async for teams with volume commitments. If you are comparing hosted providers, our review of the best ASR engines breaks down the current options.

Where ASR is heading next

None of the five architecture families is going away, and none is on track to become the single default. Bruno Hays's read: encoder-decoders and speech LLMs are already converging mathematically, so the real competition isn't architecture family anymore, it's how well each approach handles three trade-offs: streaming without giving up batch-level accuracy, language breadth without losing accuracy in any one language, and lower latency without adding infrastructure a team has to run.

Picking a model is closer to a systems decision than a leaderboard read. Test WER on your own audio, not published numbers. Match latency to what your product actually needs, not the lowest number available. Confirm language and accent coverage against your real user base. Add up total cost, including the infrastructure a self-hosted option requires, not just the sticker price.

Start with €50 in free credits and have your integration in production in less than a day. Test Gladia on your own multilingual audio to see how it handles language detection, accent-heavy speech, and code-switching.

FAQs

What is an ASR model?

An ASR model is a machine learning model that turns audio into text. Modern ASR models pair an audio encoder (the "ears") with a language model (the "brain") and fall into five architecture families: encoder-decoder, CTC, encoder-transducer, and speech LLMs with continuous or discrete input.

What is the difference between ASR and STT?

There is no difference. ASR and STT are two names for the same task: converting spoken language into text.

Is Whisper an ASR model?

Yes. Whisper, released by OpenAI in 2022, is an encoder-decoder ASR model. Its main contribution was showing that encoder-decoder models can be trained on much larger, noisier datasets than earlier CTC approaches.

What is the best ASR model?

There is no single best ASR model. The right choice depends on your accuracy target, whether you need real-time or batch processing, your language coverage, and your deployment and cost constraints. Benchmark two or three candidates on a sample of your own audio before committing. For a quick, brand-blind gut check first, the blind comparison tool is a fun and useful way to test your own audio across six providers and shows ELO-ranked results before revealing which vendor produced which transcript.

Key terms glossary

BPE (Byte Pair Encoding): The algorithm most text tokenizers use to segment words at meaningful boundaries, such as separating prefixes and suffixes. It's more efficient than splitting at the character level.

CTC (Connectionist Temporal Classification): An ASR architecture where the encoder outputs a letter or token directly for each audio slice, using an alignment trick that allows repeated letters. Accuracy improves substantially once a language model re-ranks the output.

Embeddings: Vectors that represent concepts the way words represent them for humans. Models optimize these vectors during training to store useful information.

Encoder: The component of an ASR model that turns raw audio into a sequence of embeddings representing phoneme-like sounds. Most encoders use a Transformer architecture based on attention.

Encoder-decoder: An ASR architecture, used by Whisper and Canary, where a decoder generates text token by token. It uses cross-attention to pull audio information from the encoder and self-attention to track previously generated tokens.

Encoder-transducer: An ASR architecture, used by Parakeet TDT, where a joiner reads encoder embeddings one at a time and asks a language model component what to output next. Streamable by design, but harder to batch.

Speech LLM: An architecture that adds an audio encoder to a pre-trained text LLM. Continuous-input speech LLMs, like Voxtral, connect a pre-trained encoder and LLM with a trainable adapter. Discrete-input speech LLMs, like GPT-4o and Moshi, convert audio into discrete tokens the LLM can process and generate directly.

Text tokens: Slices of a sentence created by cutting text at the character, word, or sub-word level, typically using an algorithm like BPE.

Word error rate (WER): The standard metric for ASR accuracy. It shifts with accent, noise, and domain, so benchmarking on your own audio is more reliable than trusting a published number alone.

‍

Contact us

280
Your request has been registered
A problem occurred while submitting the form.

Read more