Speech data your models actually need.
Consented, quality-reviewed voice data, delivered ready to train.
What's in the catalogue
Every dataset here has at least one hour of reviewed audio. Listen to a sample on any dataset.
Quality score: share of reviewed recordings approved by native-speaker reviewers.
What you get
Each recording is delivered as a WAV audio file with a matching JSON file.
The JSON file records the language, the task it was recorded for, its review status and reviewer tags, the contributor's consent ID and consent date, and whether it is test data. For scripted reads, it also includes the script text the contributor read.
How quality is measured
Every recording is reviewed by a native speaker of the language before it is approved. Reviewers listen in full and tag what they hear:
- Audio clarity
- Clean, intelligible speech without heavy noise, clipping or dropouts.
- Accent authenticity
- The speaker genuinely has the accent or dialect the dataset is for.
- Content accuracy
- The speaker said what was asked, with no missing or added words.
A dataset's quality score is the share of its reviewed recordings that reviewers approved.
Consent and provenance
Every contributor gives explicit consent during sign-up, with the date and time recorded.
That consent is stored against the contributor's account and travels with each recording as a consent ID and date in its JSON file.
Ways to work with us
Off-the-shelf datasets
License a dataset from the catalogue above.
Custom collection
We design and run a collection brief for the language, dialect or domain you need.
Human review of your data
Our native-speaker reviewers transcribe, label and quality-check audio you already have.
Talk to us
Tell us what you're building and our team will get back to you.