async

The async speech-to-text API built for real conversations, noisy audio, and overlapping speakers.

#1

On real-life speech

Best accuracy on conversational audio benchmarks.

3x

Lower diarization errors

Than other vendors.

100+

Languages covered

42 unique to Gladia.

Trusted by 300,000+ developers worldwide

Klarna Aircall HeyGen VEED Recall Attention Livestorm Method Sana Mojo Adversus Claap Attio Carv Selectra Spoke
Klarna Aircall HeyGen VEED Recall Attention Livestorm Method Sana Mojo Adversus Claap Attio Carv Selectra Spoke

Fewer errors than any other model

Gladia's Solaria models outperform every provider on Switchboard, the toughest conversational benchmark. Our benchmark methodology is open-source, so you can reproduce the results.

Compare models Compare models
33.9%
Solaria-3
37.3%
Solaria-1
42.3%
AssemblyAI
46%
Speechmatics
48.1%
Mistral
49.8%
Deepgram
55.2%
ElevenLabs
ACCURACY

The most accurate speech-to-text API

Our Solaria-3 leads on European business audio; Solaria-1 maximizes coverage across 100+ languages. Both benchmarked on real customer recordings.

  • #1 on Switchboard, the toughest conversational benchmark
  • Best-in-class on real-world customer audio
  • Most accurate speaker diarization in the market
  • Custom vocabulary for product names, acronyms, and domain-specific terms
Accurate async transcription with speaker diarization and custom vocabulary
MULTILINGUAL

Works in any language

Gladia covers languages no other provider transcribes, and handles mid-sentence language switches without any manual configuration.

  • 100+ languages, 42 exclusive to Gladia
  • Auto language detection, no config needed
  • Translation to target languages as a built-in output
  • Broad dialect and accent coverage
Multilingual transcription and translation across 100+ languages
COMPLIANCE

Enterprise-grade security and compliance

Sensitive audio needs more than a privacy policy. Gladia is fully compliant with requirements that matter in regulated industries.

  • SOC 2 Type 2, HIPAA, GDPR, ISO 27001 certified
  • EU and US data residency options
  • On-premise deployment for both Solaria-3 and Solaria-1
  • Your data is never used to train models on paid plans
Enterprise security certifications including GDPR, SOC2, and HIPAA

Everything you need, already built in

All you need to turn raw audio into accurate data, without stitching together separate tools.

Check the audio-intelligence suite Check the audio-intelligence suite

Speaker diarization

Automatically detect and label who said what in multi-speaker audio, with 3x fewer errors than other vendors.

Learn more

Custom vocabulary

Boost recognition accuracy for product names, acronyms, and domain-specific terms with keyterm prompting.

Learn more

Audio-to-LLM

Turn audio into structured insights, including summaries and action items, in a single API call, with access to 400+ LLM models.

Learn more

Match the model for your workflow

Solaria-3 and Solaria-1 are built for different jobs:

Solaria-3

Highest accuracy on European real-world audio

  • #1 on European business audio and conversational speech
  • Optimized for EN, FR, DE, ES, IT
  • #1 on English production audio
Learn more

Solaria-1

Maximum language coverage across any domain

  • The most multilingual model: 100+ languages, 42 unique to Gladia
  • 3x lower diarization errors than other vendors
  • Code-switching, handled natively
Learn more
USE CASES

Built for real enterprise use cases

Teams use Gladia's async API to turn raw audio into accurate, structured, and speaker-labeled data they can trust.

Meeting assistants

Accurate, speaker-labeled transcription that powers reliable AI summaries, action items, and CRM sync, without the model inventing details that were never said.

Learn more

Contact centers (CCaaS)

Transcribe and analyze recorded calls at scale for QA, compliance review, and agent coaching, with the multilingual coverage global support teams need.

Learn more

Sales enablement

Turn recorded sales calls into searchable, speaker-tagged transcripts that feed conversation intelligence and CRM workflows.

Learn more

Media & content

Generate accurate, time-stamped subtitles and transcripts for video and podcast content, with support for multi-channel audio.

Learn more
We are 100% benchmark and evaluation driven. Gladia was one of the best providers selected on merit to transcribe user videos, especially for non-English languages. Their reactive customer support and data compliance make their offer really compelling.
Kojo Hinson
Kojo Hinson Group Engineering Manager

Compare pricing

Pay only for the audio you transcribe, with pricing that drops automatically as volume grows.

Starter

Flexible pay-as-you-go for moderate audio volumes. Get started immediately.

Async at $0.61/hr

Real-time at $0.75/hr

  • What's included
  • 100+ languages supported
  • Speaker diarization
  • Word-level timestamps
  • Community support

Enterprise

Annual plan with custom models, fine-tuning, debundled pricing, and more.

Custom

  • What's included
  • Everything in Growth
  • Dedicated infrastructure
  • GDPR & SOC 2 compliance support
  • 99.9% uptime SLA
  • Named account manager
  • On-prem deployment available
Explore the pricing Explore the pricing

FAQs

Start transcribing in minutes

Sign up for free and get an API key, or book a demo to see Gladia's async transcription handle your own audio.