Automatic speech recognition (ASR) is a cornerstone of many business applications in domains ranging from call centers to smart device engineering. At their core, ASR models, also referred to as Speech-to-Text (STT), intelligently recognize human speech and convert it into a written format.
Modern ASR engines combine a combination of groundbreaking technologies, including Natural Language Processing (NLP), AI, ML, and LLMs. While ASR relies on all these technologies, it's very different from each one of them in fundamental ways.
In this article, we'll explore the inner workings of ASR models and compare legacy approaches with today's state-of-the-art sequence-to-sequence models. We'll walk through how speech recognition models work step by step, from the raw audio signal to decoded text, so you can see where accuracy is won or lost in the pipeline. If you're a developer, AI engineer, CTO, or CPO, you'll discover a wealth of insights into ASR models and how they're vital for transcriptions, captioning, and content creation in business environments.
How do speech recognition models work? The ASR pipeline
When incorporating speech recognition technologies into your business and customer workflows, it helps to understand how they work. Armed with these insights, you'll be better able to select an ASR model that best meets your specific requirements. If you're new to speech recognition, feel free to check out our introductory guide to speech-to-text AI before diving into this.
At a high level, every speech recognition model runs the same pipeline. It captures and digitises the audio signal, filters noise and extracts acoustic features from short slices of sound, then predicts the most likely words using an acoustic model and a language model, and finally decodes those predictions into text. The difference between traditional and modern ASR is how many separate models handle those stages, and whether they are trained independently or as one network.
The traditional speech recognition approach
Let's take a look at how speech recognition has historically worked. Legacy ASR systems function by converting analog audio signals into digital bits and processing them with a decoder to create sentences with words based on the data sequence and context. They remained mainstream for the last four decades until the introduction of end-to-end ASR models.
The decoder in a traditional speech recognition system analyzes the input data in conjunction with multiple ML models.
- Acoustic Models (AM): The decoder needs acoustic models like Hidden Markov Models (HMM) and Gaussian Mixture Models (GMM) that understand the natural speech patterns to predict the exact spoken sound.
- Language Models: Merely estimating the sound (phoneme) is not enough. You need a language model that can predict the right sequence of words based on a statistical analysis of the language.
- Lexicon Models: A lexicon model determines the phonetic variations in language. This helps distinguish between accents and similar-sounding phrases. These models all work in tandem to create the desired output (written text). An example of a lexicon model would be a Finite-State Transducer (FST) with a pronunciation dictionary mapping the word "SPEECH" to "S P IY CH." All natural language words are represented in the FST as subword units.
This approach necessitates training multiple models independently of each other, which is both time-consuming and expensive. One of the most significant drawbacks, however, is the reliance on the lexicon model.
The success of the traditional speech recognition models depends to a great extent on how well the lexicon models are crafted. Experienced phoneticians collaborate to create a custom set for the language at hand. The process must be repeated for every single language the model aims to support. This makes implementation challenging, especially in dynamic business environments expanding into new markets.
Modern ASR models with end-to-end (E2E) deep learning
Modern ASR models take a disruptive approach to speech processing with end-to-end deep learning. Essentially, a complex neural network in a modern ASR model replaces multi-stage models in legacy systems which minimizes latency and drastically improves performance and accuracy.
This architecture also does away with independent language, acoustic, and lexicon models, and the resultant modern ASR system functions as a single neural network (as against multiple models in legacy systems), reducing development time and costs. In addition, they also achieve higher accuracy levels and support multiple languages owing to their advanced neural architecture.
Modern end-to-end systems are also where multilingual accuracy is won or lost, since a single network can learn many languages at once. This is the stage where robustness to accents and code-switching, where a speaker changes language mid-sentence, is built in rather than bolted on.
The modern ASR models are trained on large datasets and often require self-supervised deep learning as it would be extremely challenging for human operators to manually process voluminous data. Engineers use large amounts of unlabeled data to build a foundation model, which is then fine-tuned to achieve the desired word error rate (WER). Models can also be benchmarked based on this metric to compare performance.
Two current examples of modern end-to-end ASR models are our own Solaria-1 and Solaria-3, built for different jobs. Solaria-1 supports 100+ languages, with true mid-conversation code-switching built in, making it the broader-coverage choice for multilingual products, and real-time streaming. Solaria-3 is tuned for real-world European business and contact-center audio, noisy, fast-paced, and conversational, and on production recordings ranks #1 across English and core European languages (EN, FR, DE, ES, IT), ahead of AssemblyAI, ElevenLabs, Deepgram, Mistral, and Speechmatics. Against Solaria-1, it's 26% more accurate on real English customer calls. Both are more robust than models trained on smaller, more specialized datasets.
- Encoder: The encoder uses neural networks to process the input sequence and create a fixed-size vector representation. The encoder will also capture context from the input sequence and pass it to the decoder.
- Decoder: A 'decoder' module takes the encoder output vector and creates an output sequence. The decoder will predict the next tokens of the output sequence based on the received context and its own previous predictions.
- Transformers: Most current end-to-end models use an end-to-end transformer architecture to understand context and meaning from the input audio. The model first splits the input audio into small chunks before passing them to the encoder. The decoder predicts the text caption.
- Recurrent Neural Networks (RNN): Sometimes both the encoder and the decoder use Recurrent Neural Networks (RNN), a specific type of neural network well-suited for sequence prediction. RNN cells can remember information about previously seen elements of the sequence through their internal memory and use it to determine current output. Unlike transformers, RNNs are sequential models. LSTM (Long Short-Term Memory) models, for example, use an RNN architecture.
Legacy ASR models vs seq2seq architecture
The shift from legacy ASR to seq2seq architecture isn't just an accuracy upgrade, it changes what's practical to build and maintain. Legacy systems need a lexicon model built and validated for every language you support, which is why traditional ASR vendors historically supported a handful of languages well and treated everything else as an afterthought. Seq2seq models learn pronunciation, grammar, and acoustic patterns jointly from data, so adding a language becomes a training and data problem rather than a multi-year linguistics project, and there's no separate lexicon model to fail when a speaker's pronunciation doesn't match the dictionary. The tradeoff is data and compute: seq2seq models need large, diverse training sets and meaningful infrastructure to train, where legacy pipelines could be assembled from smaller, independently trained components.
We've summarized the key differences between the traditional and modern ASR approaches below for easy reference.
| Characteristics |
Legacy ASR models |
Seq2Seq models |
| Internal structure |
Modular pipeline with acoustic models, lexicon models, and language models |
Use end-to-end neural networks with an encoder-decoder architecture to directly convert speech to text |
| Accuracy |
Limited accuracy, cannot reach human accuracy levels |
High accuracy, can reach human-level accuracy and beyond |
| Versatility |
Limited adaptability to diverse inputs (some accents can throw off transcriptions) |
Highly adaptable to different accents and languages |
| Training data type |
Needs labeled phonetic data to function properly |
Works with unlabeled or less labeled data |
| Training technique |
Each model needs to be independently trained |
The entire model is trained in one go |
| Suitability |
Best suited for simple speech-to-text functions |
Best suited for complex and real-time speech-to-text applications |
| Speed |
Generally slower due to multi-stage component interactions |
Typically faster due to parallel processing |
| Error detection and correction |
Limited error correction capabilities |
Complex error correction and control mechanisms |
| Scalability |
Limited scalability |
Highly scalable |
| Language support |
Support a limited number of languages |
Can support a large number of languages |
| Context awareness |
Limited context awareness. Probabilistic language models analyze context to predict the next word with limited use of machine learning. |
Highly context-aware, use deep learning for comprehensive context analysis |
Do legacy ASR models still have a place?
While legacy models suffer from drawbacks, they aren't in any way rendered obsolete. In fact, modern ASR models build upon several traditional fundamental speech recognition systems such as acoustic and language models. Furthermore, legacy systems are also deployed for some specific tasks.
Generally speaking, though, Seq2seq models are great at tasks that require natural language understanding. This includes speech recognition, translation, and caption generation. This is, in part, owing to their superior architecture that lends them advantages when dealing with sequences of varying lengths. That said, modern models do need very large datasets and compute resources to train and work. They can also face problems with long-range dependencies and when rare words are encountered.
How to select the right ASR model?
It can be challenging to select an optimal ASR model for your specific needs. We've put together a few decisive factors you should consider to guide the selection process.
- Word error rate: The ideal goal of any voice recognition application is to achieve zero error rates. However, practical considerations dictate variations beyond our control, so make sure you factor in the precision and accuracy you need in your system when selecting an ASR model. If your application needs uncompromising performance, choose a modern end-to-end ASR system.
- End goal: Consider the requirements of your end users when choosing a model. How will they use the product or service?
- Operating environment: ASR performance depends on the background noise to a great extent. If your end users will operate the product or application in noisy environments, you'll need advanced noise suppression support in your model.
- Audio properties: Consider the audio sample rate, bitrate, file format, channels, and duration when selecting a model.
- Input audio type: Factor in how varied your input audio will be and the languages and dialects the model will need to support.
- Performance: Every model performs differently, so you'll need to evaluate it based on your specific benchmarks. If you need real-time speech-to-text conversions (such as in smart devices and wearables), choose a model with the lowest latency possible.
If your use case involves multi-speaker audio, accents, code-switching, or noisy real-world conditions, general-purpose architectures aren't always built around those problems specifically. Our own Solaria-1 and Solaria-3 models are built for that: Solaria-1 for full language breadth and code-switching, Solaria-3 for real-world European business and contact-center audio, where it currently ranks #1 on Switchboard. For your own audio, the blind comparison tool is a fun and useful way to test providers by removing the brand bias, letting you test files across six providers with ELO-ranked results, no integration work required. It's a useful gut check, but follow it with a reproducible benchmark on your full audio distribution before making a production commitment.
Choosing the right ASR architecture
Unless you have a specific reason to stay on a legacy pipeline, an existing investment in a lexicon model for a narrow, well-defined vocabulary, for example, seq2seq architecture is the practical default for anything shipping today. The real decision isn't legacy versus modern anymore, most production systems settled that years ago, it's which end-to-end model fits your audio.
Match the model to the conditions your users actually produce. Clean, single-speaker, single-language audio is close to solved across most providers. Multi-speaker calls, accents, background noise, and code-switching are where architectures and training data still diverge, and where testing on your own audio matters more than any published benchmark. If speaker labels matter for your use case, remember diarization runs as a separate step on top of the transcript, not something the core ASR model does for free.
Test Gladia on your own multilingual audio to see how it handles language detection, accent-heavy speech, and code-switching. Start with €50 in free credits and have your integration in production in less than a day.
FAQs
How do speech recognition models work?
Speech recognition models turn spoken audio into text through a pipeline: they capture and digitise the audio signal, extract acoustic features from short slices of sound, then use an acoustic model and a language model to predict the most likely sequence of words, and finally decode that into text. Traditional systems ran these stages as separate models; modern end-to-end models do it in a single neural network. Our Solaria-1 and Solaria-3 models follow this modern approach: a single network takes raw audio in and returns decoded text, with no separate acoustic, language, or lexicon models to maintain.
What is an ASR model?
An ASR (automatic speech recognition) model is a machine learning model that converts spoken language into written text. It is the core component behind transcription tools, voice assistants, and captioning. The term is used interchangeably with speech-to-text (STT). We built our own ASR models, Solaria-1 and Solaria-3, specifically for products like meeting assistants and contact centers, where transcription accuracy determines the quality of every downstream system, from CRM entries to AI summaries.
What is the difference between an end-to-end model?
A traditional ASR model chains together several components, an acoustic model, a lexicon model, and a language model, to convert sound into words. An end-to-end model replaces that stack with a single neural network trained in one pass, which is generally more accurate, handles more languages, and is faster to build. We build both of our production versions across 100+ languages, with true mid-conversation code-switching. Solaria-3 is tuned for real-world contact-center audio, and ranks #1 on Switchboard, the most challenging dataset, ahead of AssemblyAI, Deepgram, ElevenLabs, and others.
What is the difference between ASR and STT?
There is no practical difference. ASR (automatic speech recognition) and STT (speech-to-text) both describe the task of converting spoken audio into a written transcript, and the terms are used interchangeably across the industry. We use the terms interchangeably too. Whether you call it ASR or STT, our API returns the same structured output from one call: word-level timestamps, translated text, summaries, named entities, and sentiment.
How accurate is speech recognition?
Accuracy is measured with word error rate (WER), the percentage of words a model gets wrong. Modern end-to-end models reach low-single-digit WER on clean audio, though accuracy drops with background noise, heavy accents, and multiple speakers, which is why the model and its training data matter as much as the architecture. Our own accuracy work follows the same logic: we test Solaria-1 and Solaria-3 on real production audio rather than only clean benchmark recordings, and publish the methodology at our async benchmark and Solaria-3 page so the numbers can be checked against real conditions.
Key terms glossary
ASR (Automatic Speech Recognition): the technology that converts spoken audio into written text.
STT (Speech-to-Text): used interchangeably with ASR to describe the same task.
WER (Word Error Rate): the percentage of words a transcript gets wrong compared to a human-verified reference. Lower is better.
DER (Diarization Error Rate): the percentage of a transcript where speaker labels are assigned incorrectly.
Lexicon model: a pronunciation dictionary that maps written words to their phonetic building blocks, used in traditional ASR to bridge the acoustic and language models.
Hidden Markov Model (HMM): a statistical model used in traditional acoustic models to represent how sounds change over time.
Gaussian Mixture Model (GMM): a statistical model often paired with HMMs in traditionalASR to represent the probability distribution of acoustic features.
Finite-State Transducer (FST): a structure used in lexicon models to represent valid pronunciations as paths through a set of states.
Seq2seq (sequence-to-sequence): an end-to-end neutral network architecture that takes asequence as input (audio) and produces a related sequence as output (text), replacing the separate acoustic, language, and lexicon model of traditional ASR.
LSTM (Long Short-Term Memory): a type of RNN designed to retain relevant context over
longer sequences without it fading out.
Diarization: identifying who spoke when in a multi-speaker recording. Async-only in Gladia's pipeline, powered by pyannoteAI's Precision-2 model.
Code-switching: when a speaker changes language mid-conversation. A common failure point for ASR systems not built to handle it.