Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.
MiFID II and FCA call recording: compliance for voice transcription in finance
TL;DR: Financial firms operating under MiFID II and FCA jurisdictions must maintain searchable, high-accuracy records of all client-related voice communications, including remote and mobile calls under FCA Market Watch 66. Engineering and product teams building for these obligations typically implement dedicated cloud infrastructure with defined data residency, controls that prevent customer audio from being used to retrain models on Growth and Enterprise plans, and transcription accurate enough that the resulting records hold up under regulatory review. Transcription and speaker attribution errors are not product quality issues in this context. They are audit trail failures, and regulators treat them as such.
Why speech-to-text accuracy matters upstream of your LLM
TL;DR: Downstream LLM performance is ceiling-bounded by upstream transcription accuracy. A transcript with a meaningful error rate doesn't produce proportionally degraded summaries or CRM entries. It produces outputs where hallucinated names, inverted logic, and misattributed speaker turns compound silently into every downstream system that reads them. Prompt engineering cannot recover information the STT layer never captured. Gravite cut call quality review time by 93%, from 15 minutes to 1 minute per call, once transcript accuracy was high enough to trust the output without manual verification.
TL;DR: Telephony audio is constrained to 8kHz narrowband frequencies, stripping away the high-frequency spectral energy that generic 16kHz STT models require. Standard upsampling cannot recover phonemes that were never captured at the source, leading to transcription errors that compound silently into downstream NLU failures. Solaria-3 ranks #1 on Switchboard, the most challenging conversational telephony dataset, ahead of AssemblyAI, ElevenLabs, Deepgram, Mistral, and Speechmatics, ensuring your downstream LLM pipelines receive clean, structured data from real-world noisy call audio.
A new open-source developer app for AI translation, dubbing and lip synching to try
Feb 1, 2024
Text-to-speech, voice cloning, and visual dubbing are some of the hottest trends in AI at the moment. Used in tandem with AI transcription and translation, they make it possible to generate hyper-realistic voiceovers, indistinguishable from the sound of the speaker’s natural voice and speech patterns — including in entirely new languages.
Our partners at Sync Labs have just published an open-source repo for building an app that translates any video to any language with perfectly matched lip movements. Its backbone leverages Gladia API for speech-to-text and translation, ElevenLabs for text-to-speech and voice cloning, and Sync Labs for visual dubbing.
Following a quick intro to all of the tech elements of this fantastic project, we’ll explain how you can test them first-hand using this app, which will be available for public access in a week.
Speech-to-text and translation
Speech-to-text or automatic speech recognition (ASR) converts spoken words into text. The process involves preprocessing audio data to enhance quality, employing advanced speech recognition algorithms to correctly identify words, and integrating language modeling to predict word sequences. Post-processing may be applied to refine the transcribed text, resulting in an accurate written representation of the spoken content. For a more detailed breakdown of how it works, feel free to visit our introduction to speech-to-text.
AI translation, also known as machine translation, employs AL/ML to automatically translate text or speech from one language to another. The process includes tokenization of input, utilizing natural language processing for context and grammar understanding, and employing machine learning models—often neural networks like the multilingual Whisper ASR—to predict accurate translations.
At Gladia, we rely on a hybrid ASR architecture, powered by optimized Whisper and other state-of-the-art models, supporting 99 languages for transcription and translation. Integrated into Sync’s app, our API allows us to transcribe what’s being said and translate it in near real-time, with the resulting transcript fed into the rest of the structure.
Text-to-speech and voice cloning
Text-to-speech technology does the opposite of speech-to-text by converting written text into spoken language. The system analyzes input text using natural language processing, understanding its structure and semantics. Prosody modeling is then applied to incorporate elements like intonation and rhythm, contributing to a natural and expressive synthesized speech. The synthesis engine generates speech based on the analyzed text and prosody modeling, resulting in a final output of synthesized voice that faithfully represents the spoken version of the input text.
Voice cloning comes into play to make the output as close to the human voice as possible. To yield realistic results, the process starts with collecting a substantial dataset of the target voice. Extracting relevant features like pitch and tone, machine learning models, often utilizing deep neural networks, are trained to mimic the unique characteristics of the voice across a wide emotional spectrum, i.e. confident speech, happy exclamations, angry rants, and so on.
ElevenLabs is among the top software out there for text-to-speech and voice cloning. The company leverages proprietary deep-learning tech to choose from a library of high-fidelity male and female voices (or produce them from scratch!), enabling seamless creation of custom videos, ebooks, and more in 29 languages.
Visual dubbing
Visual dubbing, or lip reanimation, is an AI technology that synchronizes translated or transcribed audio with realistic lip movements in video content.
By analyzing and replicating the original speaker's lip gestures, the system generates animated lip movements that align with the new audio. While the technology is raising obvious concerns about the use of deep fakes, it’s also a highly powerful tool to break down language barriers in video content, providing a high-fidelity alternative to traditional dubbing.
On a mission to break the language barriers in video content and reinvent dubbing, Sync Labs enables developers to seamlessly lip-sync a video to audio in near real-time using a single API.
How the translation app works
We invited you to dive into the x thread below for a detailed video tutorial and instructions. Theres's also this Medium tutorial available, and of course the link to the original repo to clone and launch the app yourself.
Conclusion
Thanks to this amazing open-source project, we can see just how powerful speech-to-text, text-to-speech, voice cloning, and lip-synching technologies can be when used together. We hope you enjoy this incredible free tool by Sync Labs, powered by Gladia. If you’re building a voice app using our API and would like us to spread the word, do not hesitate to reach out here.
Contact us
Your request has been registered
A problem occurred while submitting the form.
Read more
Speech-To-Text
MiFID II and FCA call recording: compliance for voice transcription in finance
Speech-To-Text
Why speech-to-text accuracy matters upstream of your LLM