API Comparison Table

Heading 1

Heading 2

Heading 3

Heading 4

Heading 5
Heading 6

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Block quote

Ordered list

  1. Item 1
  2. Item 2
  3. Item 3

Unordered list

Text link

Bold text

Emphasis

Superscript

Subscript

Pricing
Get started
Get started

Read more

Speech-To-Text

MiFID II and FCA call recording: compliance for voice transcription in finance

TL;DR: Financial firms operating under MiFID II and FCA jurisdictions must maintain searchable, high-accuracy records of all client-related voice communications, including remote and mobile calls under FCA Market Watch 66. Engineering and product teams building for these obligations typically implement dedicated cloud infrastructure with defined data residency, controls that prevent customer audio from being used to retrain models on Growth and Enterprise plans, and transcription accurate enough that the resulting records hold up under regulatory review. Transcription and speaker attribution errors are not product quality issues in this context. They are audit trail failures, and regulators treat them as such.

Speech-To-Text

Why speech-to-text accuracy matters upstream of your LLM

TL;DR: Downstream LLM performance is ceiling-bounded by upstream transcription accuracy. A transcript with a meaningful error rate doesn't produce proportionally degraded summaries or CRM entries. It produces outputs where hallucinated names, inverted logic, and misattributed speaker turns compound silently into every downstream system that reads them. Prompt engineering cannot recover information the STT layer never captured. Gravite cut call quality review time by 93%, from 15 minutes to 1 minute per call, once transcript accuracy was high enough to trust the output without manual verification.

Speech-To-Text

Telephony-audio robustness: why 8kHz narrowband calls break generic STT

TL;DR: Telephony audio is constrained to 8kHz narrowband frequencies, stripping away the high-frequency spectral energy that generic 16kHz STT models require. Standard upsampling cannot recover phonemes that were never captured at the source, leading to transcription errors that compound silently into downstream NLU failures. Solaria-3 ranks #1 on Switchboard, the most challenging conversational telephony dataset, ahead of AssemblyAI, ElevenLabs, Deepgram, Mistral, and Speechmatics, ensuring your downstream LLM pipelines receive clean, structured data from real-world noisy call audio.

MiFID II and FCA call recording: compliance for voice transcription in finance

Published on Sep 25, 2026
by Ani Ghazaryan
MiFID II and FCA call recording: compliance for voice transcription in finance

TL;DR: Financial firms operating under MiFID II and FCA jurisdictions must maintain searchable, high-accuracy records of all client-related voice communications, including remote and mobile calls under FCA Market Watch 66. Engineering and product teams building for these obligations typically implement dedicated cloud infrastructure with defined data residency, controls that prevent customer audio from being used to retrain models on Growth and Enterprise plans, and transcription accurate enough that the resulting records hold up under regulatory review. Transcription and speaker attribution errors are not product quality issues in this context. They are audit trail failures, and regulators treat them as such.

Most engineering and product leads at financial firms treat call recording as a passive storage task. Under MiFID II Article 16(7) and SYSC 10A.1 of the FCA Handbook, the real challenge is converting thousands of hours of noisy, multi-speaker audio into searchable, structured, tamper-proof compliance assets that hold up when a regulator calls.

Regulatory enforcement actions against major financial firms have consistently targeted failures to maintain and preserve electronic communications, including voice channels. Recent SEC actions have resulted in multi-million dollar fines for firms including Blackstone entities, KKR, and Charles Schwab for recordkeeping violations. The message from regulators is unambiguous: recording is not the compliance bar. Accurate, retrievable, and tamper-proof records are.

This guide explains the regulatory obligations in technical terms and shows how to build a voice transcription pipeline that satisfies them.

What are MiFID II and FCA call recording requirements?

Under MiFID II Article 16(7), investment firms and credit institutions must record telephone conversations and electronic communications relating to transactions concluded when dealing on own account, and those relating to the provision of client order services. The scope extends to communications "intended to result in" a transaction, meaning the test is whether a conversation could plausibly lead to a transaction, not whether it actually does. The European Securities and Markets Authority's (ESMA's) MiFID II Q&A confirms this broad interpretation and is understood to include pre-trade discussions, order amendments, and cancellation calls within scope.

Applying Article 16(7) to voice data

The requirement covers the full substance of every relevant trade communication: caller identity, conversation content, and accurate timestamps sufficient to establish when the communication took place. Firms must capture this information regardless of whether the conversation results in a completed transaction. The practical implication is that compliance capture must be on by default for all channels used by relevant personnel, not triggered after a deal is confirmed.

SYSC 10A.1 compliance for voice data

In the UK, the substantive recording obligations sit in SYSC 10A.1 of the FCA Handbook. The FCA's requirements mirror MiFID II on the core mandate: record relevant telephone conversations, retain relevant electronic communications, notify clients that calls are being recorded, and take reasonable steps to prevent the use of channels the firm cannot capture.

The practical implication of SYSC 10A.1's availability requirement matters here. Records that must be produced on demand to the FCA are functionally useless if the underlying audio is inaudible, incomplete, or inaccurately transcribed. Availability is only meaningful if the record is legible.

Which firms must record client calls

The recording obligation applies to investment firms, inter-dealer brokers, wealth managers, credit institutions conducting MiFID-scope business, systematic internalisers, and execution venues. In practice: if your platform executes, transmits, or receives orders in financial instruments, call recording is mandatory.

Retention rules for MiFID II voice audit trails

Mandatory retention periods: 5 to 7 years

MiFID II Article 16(7) sets a mandatory minimum retention period of five years for all relevant voice recordings and electronic communications, measured from the date of the recording. For FCA-regulated firms, the UK position is stricter: SYSC 10A.1.12AR already mandates that records be available to the FCA for seven years and to the client for five years as a standard rule. This is not a discretionary extension a firm requests. For continental EU firms, other national competent authorities such as BaFin can extend the MiFID II minimum to seven years upon request. Firms operating across both UK and EU jurisdictions should design their retention architecture to the seven-year ceiling as a baseline to avoid jurisdiction-specific gaps.

Audio format requirements for regulators

Regulators expect recordings to remain legible and tamper-proof across the full retention window. Uncompressed formats (WAV) and lossless compressed formats (FLAC) satisfy this requirement and are both accepted by our API, which processes files up to 135 minutes and 1,000MB. WAV retains all original sound data without any compression and offers the highest fidelity, making it the preferred choice where evidentiary integrity is essential. FLAC achieves smaller file sizes without discarding any audio data, making it a practical alternative. Lossy formats such as M4A permanently discard audio data and, given their proprietary nature and inconsistent compatibility across archival systems, are generally less suitable for long-term regulatory compliance storage. Archiving must prevent unauthorized alteration, with an auditable change log and cryptographic hashing at ingestion to prove tamper-resistance if an information request arrives years later.

Key recordkeeping rules for financial voice logs

Compliant storage is not just about keeping files. Regulators need to retrieve a specific conversation, with context, within a tight window when they issue an information request. That requires four things working together:

  1. Tamper-proof storage: Audio files and transcripts must be stored in a way that prevents unauthorized alteration or deletion, with an auditable change log.
  2. Metadata tagging: Every record needs timestamp, caller ID, counterparty identifier, and ideally a transaction or order ID attached at ingestion so compliance teams can run precise searches rather than listening through hours of audio.
  3. Transcript accuracy: If a transcript misrepresents a numerical value, an instrument name, or attributes a quote to the wrong speaker, the audit trail is legally compromised. Poor word error rate is a regulatory risk, not just a product quality issue.
  4. Retrievability on demand: FCA Market Watch 66 reinforced that recording obligations apply in full to remote and hybrid working environments. Firms must be able to locate and reconstruct a specific trade conversation regardless of whether it was recorded in the office, from a home network, or over a mobile channel.

Our audio intelligence features, including named entity recognition and structured output, let compliance teams attach instrument names, account numbers, and entity tags directly to transcripts at processing time, turning passive audio logs into queryable compliance records.

How transcription powers MiFID II compliance workflows

Converting audio to text is what makes a call archive operationally useful. Without accurate transcripts, every information request from the FCA or a national competent authority requires manual audio review, which is slow, expensive, and error-prone at scale.

Building searchable, audit-ready archives

High-quality transcripts let compliance teams run automated keyword searches across millions of call hours, flagging terms like "guarantee," "off-book," specific instrument tickers, or prohibited phrases. Our key data extraction pipeline, powered by named entity recognition included in the base rate on Starter and Growth plans, surfaces account numbers, names, and intents from calls without a separate enrichment layer.

Trade reconstruction workflows require compliance officers to package audio, transcripts, metadata, and speaker attribution into a single audit bundle. With speaker diarization powered by pyannoteAI's Precision-2 model, each segment of the transcript carries a speaker label and word-level timestamp, so you can answer who said what and when without any post-processing guesswork. The conversation intelligence capabilities this unlocks also power internal compliance monitoring, so your team flags issues before the regulator does.

Regulatory enforcement note: Recent SEC enforcement actions have resulted in substantial penalties across multiple firms for failures to maintain and preserve electronic communications, including unrecorded voice channels. Relying on poorly transcribed or unmonitored voice channels is not a defensible position for regulated firms.

Engineering controls for MiFID II data governance

GDPR-compliant data residency strategies

We host on dedicated cloud clusters in EU and US regions. EU-based firms route their API calls to the EU cluster by default, keeping audio processing and transcript storage within GDPR-compliant geographic boundaries. Our compliance hub provides certifications, DPA templates, and data flow documentation for legal and procurement review.

Audio retention vs. transcription retention

Configure two independent retention policies: one for raw audio files and one for transcripts, coordinated so GDPR over-retention risk is avoided when the retention window closes. Our data retention documentation details how to configure deletion via the API, and we provide explicit delete endpoints for both audio and transcript records.

On Growth and Enterprise plans, customer audio is never used for model training by default and requires no opt-out action. On our Starter plan, data may be used for training, which is an important distinction for any firm handling sensitive financial audio.

Tamper-evidence as a compliance control

Generating a SHA-256 hash of both the audio file and the returned transcript at ingestion, and storing those hashes separately from the files themselves, is one well-established engineering approach for demonstrating to regulators that a record has not been altered since it was created. This is not a vendor-provided feature, and it belongs in your compliance pipeline design from the start.

How to evaluate transcription accuracy for financial calls

Measuring WER for background noise

Word error rate is the standard metric for evaluating transcription engines, calculated by dividing insertions, deletions, and substitutions by the total words spoken. Trading floors, conference calls, and financial advisory sessions all present noisy, multi-speaker, fast-paced audio that breaks models trained primarily on clean read-speech.

Solaria-3 ranks #1 on Switchboard, ahead of AssemblyAI, ElevenLabs, Deepgram, Mistral, and Speechmatics. Solaria-3 also leads on Earnings22 financial call audio at 6.4% WER. Both benchmarks use real-world business audio conditions.

Diarization for multi-party calls

Diarization error rate (DER) is the sum of three error types (speaker confusion, missed speech, and false alarms) expressed as a percentage of total reference speaking time. For compliance purposes, attributing a trade instruction to the wrong participant is as serious as misquoting the instruction itself.

Our diarization is powered by pyannoteAI's Precision-2 model, benchmarked across standard multi-speaker datasets including AMI, CALLHOME, and DIHARD 3. Diarization is available in async workflows only, which is the correct architecture for post-call compliance pipelines where a few seconds of processing delay is acceptable in exchange for higher attribution accuracy. For real-time workflows, speaker attribution should be handled in post-processing to maintain accuracy.

Code-switching in multilingual teams

Global trading desks switch languages mid-call constantly, and English, French, German, Spanish, and Italian often appear in the same conversation. Our Solaria-1 model handles mid-conversation code-switching across 100+ supported languages without breaking the transcription session or returning garbled output. For financial institutions with BPO operations in Southeast Asia or South Asia, Solaria-1 also covers Tagalog, Bengali, Tamil, and Punjabi, languages no other API-level STT provider supports.

Route European telephony and contact-center calls to Solaria-3 for maximum accuracy. Route multilingual trading desk calls with mid-conversation code-switching to Solaria-1. Custom vocabulary is included in the base rate and is the mechanism for teaching the model financial terminology, instrument names, and internal shorthand. Our custom vocabulary guide explains how to configure term lists for financial jargon at the API call level.

Validating transcription with live samples

Vendor benchmark tables are a starting point, not a final evaluation. Our blind STT comparison tool is a fun and useful way to test your own audio: it lets you upload up to two minutes of your financial audio and pick the better transcript before seeing which provider produced it. It tests Gladia alongside Deepgram, AssemblyAI, ElevenLabs, Speechmatics, and Mistral. Use it as an initial test, then follow it with a reproducible benchmark on your full audio distribution before committing to a production pipeline.

Operationalizing MiFID II data governance

Audio reaches our API through confirmed telephony integrations including Twilio, Vonage, and Telnyx, as well as real-time infrastructure partners such as LiveKit, Pipecat, and Vapi. Once audio is flowing, the compliance controls below apply regardless of the capture layer.

Modeling transcription costs at scale

All-inclusive per-hour pricing removes the most common source of infrastructure cost surprise for compliance workflows. Diarization, sentiment analysis, named entity recognition, and translation are all included in the base rate on Starter and Growth plans, with no add-on fees. The table below shows projected monthly costs for a compliance pipeline running async transcription at our published rates.

Monthly audio volume Starter plan ($0.61/hr async) Growth plan (as low as $0.20/hr async) Included compliance features
100 hours $61.00 $20.00 Diarization, NER, sentiment, translation
1,000 hours $610.00 $200.00 Diarization, NER, sentiment, translation
10,000 hours $6,100.00 $2,000.00 Diarization, NER, sentiment, translation

‍

Growth plan rates require an upfront volume commitment. On Growth and Enterprise plans, customer data is never used for model training by default.

Competitors like Deepgram and AssemblyAI charge diarization and enrichment features as separate add-ons. AssemblyAI's equivalent feature stack (speaker identification, sentiment, and entity detection) runs approximately $0.33/hr before translation, per their current Universal-3.5 Pro model pricing. At 10,000 hours per month, stacked add-on pricing compounds quickly. Gravite, a French CCaaS quality-monitoring platform processing 50,000 hours of audio per year, saw a 93% reduction in call quality review time after switching to our API, from approximately 15 minutes to 1 minute per call.

Managing PII in voice transcriptions

PII redaction is available and must be explicitly enabled in the API payload. It is not active by default. For financial calls containing account numbers, national insurance numbers, payment card details, or personal identifiers, configuring redaction at the API call level masks those values in the transcript while leaving surrounding context intact. This satisfies GDPR's data minimization principle without a separate post-processing step. Our call recording compliance guide details the GDPR, PCI DSS, and HIPAA implications for contact centers.

Technical integration checklist for compliant voice pipelines

Use the checklist below to audit any transcription vendor against MiFID II and FCA requirements before committing to a production integration.

  • Verify data residency: Confirm your API calls route to dedicated EU-west or US-west cloud clusters to meet regional data sovereignty mandates.
  • Configure zero-retraining tiers: Confirm you are on a Growth or Enterprise plan so sensitive financial audio is never used for model training by default.
  • Enable speaker diarization: Set the diarization flag in your async API request to use pyannoteAI's Precision-2 model for speaker attribution.
  • Configure PII redaction: Explicitly enable PII redaction in your API payload to mask names, account numbers, and payment card details.
  • Implement cryptographic hashing: Generate a SHA-256 hash of both the audio file and the returned transcript at ingestion (application-layer step) to prove tamper-resistance in your audit logs.
  • Set retention deletion schedules: Use the delete endpoint to purge records precisely when the five-year or seven-year window closes.
  • Attach transaction metadata: Tag each API call with caller ID, counterparty identifier, and order or transaction ID so records are retrievable by reference, not by listening.
  • Review certifications: Confirm the vendor holds SOC 2 Type II, ISO 27001, General Data Protection Regulation (GDPR), and Health Insurance Portability and Accountability Act (HIPAA) certifications. Request the Data Processing Agreement before processing live audio.

Clarifying MiFID II call retention rules

Transcription obligations under MiFID II

MiFID II does not explicitly mandate transcription. What it mandates is that records are searchable, accessible, and retrievable on demand. At any meaningful scale, that operational requirement makes high-accuracy transcription the only viable solution. Manually reviewing raw audio files to respond to an FCA information request is not feasible when you are processing thousands of hours of calls per month.

Audio deletion policies under MiFID II

Retention and deletion must be coordinated across both the audio file and the transcript. Under GDPR, holding personal data beyond the legitimate retention purpose creates its own liability. Configure your deletion schedules so that audio and transcripts are purged at exactly the same time when the five-year or seven-year window closes, and document each purge with an audit log entry. Our automated call disposition capabilities can flag records approaching their deletion threshold as part of your after-call workflow.

Verification protocols for transcript errors

Low-confidence segments, common on noisy trading floor audio or heavily accented calls, should route to a manual review queue rather than enter the compliance archive unchecked. Build a confidence threshold into your processing pipeline and treat segments below that threshold as provisional until a human reviewer confirms them. Our audio intelligence pipeline returns word-level confidence scores that make this routing logic straightforward to implement.

Regulatory requirements for mobile recording

FCA Market Watch 66 clarified that call recording and retention obligations apply equally to remote, hybrid, and mobile working environments, and that firms must take reasonable steps to prevent staff from using communication channels the firm cannot capture. Mobile voice channels must pipe into the same compliant transcription workflow as office-based calls. Our multilingual customer support pipeline architecture covers the capture-to-transcript workflow for distributed teams.

Start with €50 in free credits on the Starter plan for evaluation, then move to the Growth or Enterprise plan for the zero-retraining guarantee before processing live client audio.

FAQs

What is the standard retention period for financial call recordings under MiFID II?

MiFID II Article 16(7) mandates a minimum retention period of five years for all client-related voice recordings, measured from the date of the recording. For FCA-regulated firms, SYSC 10A.1.12AR already sets seven years availability to the FCA and five years to the client as the standard rule, not a discretionary extension. For continental EU firms, other national competent authorities such as BaFin can extend the baseline to seven years upon request.

Does Gladia support on-premises deployment for highly regulated financial firms?

We do not support on-premises or air-gapped hosting on any plan. We deploy exclusively on dedicated cloud clusters in secure EU and US regions, which satisfy the data residency and compliance requirements of regulated fintechs and financial platforms, documented at our compliance hub.

Is speaker diarization available for real-time financial transcription workflows?

Speaker diarization is an async, batch-only feature powered by pyannoteAI's Precision-2 model. For real-time workflows, speaker attribution should be handled in post-processing to maintain high accuracy.

Does Gladia use sensitive financial voice data to train its AI models?

On Growth and Enterprise plans, customer data is never used for model training by default and requires no opt-out action. On the Starter plan, customer data may be used for training.

Which Gladia model is most accurate for European financial call audio?

Solaria-3 is built for real-world business audio across English, French, German, Spanish, and Italian, and ranks first on the Earnings22 financial call benchmark at 6.4% WER, the only model under 7%, ahead of AssemblyAI, ElevenLabs, Deepgram, Mistral, and Speechmatics.

How does MiFID II apply to calls made on personal mobile devices?

FCA Market Watch 66 confirmed that recording and retention obligations apply to all remote, hybrid, and mobile environments. Firms must capture mobile voice channels into the same compliant pipeline as office-based calls, and must take reasonable steps to prevent staff from using channels the firm cannot record.

Key terms glossary

Word error rate (WER): The standard metric for transcription accuracy, calculated by dividing the total number of insertions, deletions, and substitutions by the total number of words spoken. Lower WER means fewer errors in the audit trail.

Diarization error rate (DER): The metric for speaker attribution accuracy, calculated as the sum of speaker confusion, missed speech, and false alarm errors expressed as a percentage of total reference speaking time. High DER means your compliance records cannot reliably answer who said what.

Code-switching: The practice of alternating between two or more languages within a single conversation, which Solaria-1 automatically detects across 100+ supported languages without breaking the transcription session.

FCA Market Watch 66: A regulatory publication from the UK Financial Conduct Authority clarifying that call recording and retention obligations apply equally to remote, hybrid, and mobile working environments, published in January 2021.

SOC 2 Type II: An independent audit standard that verifies a service organization's security, availability, and confidentiality controls over a defined operating period, required by most enterprise financial procurement processes.

PII redaction: An optional feature that must be explicitly enabled in the API payload. It masks personal identifiers such as names, account numbers, and payment card details in the transcript while preserving surrounding context.

‍

Contact us

280
Your request has been registered
A problem occurred while submitting the form.

Read more