Dataset attributions

Gladia's speech-to-text models are built with the help of openly licensed datasets. We gratefully acknowledge the providers whose work makes our research possible.

ACKNOWLEDGMENTS

Gladia thanks the following dataset providers

The corpora below are distributed under the Creative Commons Attribution 4.0 International (CC-BY 4.0) license and are used to train and evaluate Gladia's speech-to-text models.

Google Research

FLEURS

Few-shot Learning Evaluation of Universal Representations of Speech — a parallel speech corpus spanning 102 languages, designed for multilingual speech recognition and language identification.

Tallinn University of Technology

VoxLingua107

A spoken language identification dataset covering 107 languages, totalling roughly 6,600 hours of speech automatically collected from YouTube.

University of Edinburgh

AMI Meeting Corpus

Around 100 hours of multi-modal meeting recordings with rich manual annotations, widely used for speech recognition and speaker diarization research.

These datasets are used under the terms of the Creative Commons Attribution 4.0 International License. Attribution does not imply that the dataset providers endorse Gladia or its products.