Speech data your models actually need.

Consented, quality-reviewed voice data, delivered ready to train.

Browse datasets

What's in the catalogue

Every dataset here has at least one hour of reviewed audio. Listen to a sample on any dataset.

Quality score: share of reviewed recordings approved by native-speaker reviewers.

What you get

Each recording is delivered as a WAV audio file with a matching JSON file.

The JSON file records the language, the task it was recorded for, its review status and reviewer tags, the contributor's consent ID and consent date, and whether it is test data. For scripted reads, it also includes the script text the contributor read.

How quality is measured

Every recording is reviewed by a native speaker of the language before it is approved. Reviewers listen in full and tag what they hear:

Audio clarity
Clean, intelligible speech without heavy noise, clipping or dropouts.
Accent authenticity
The speaker genuinely has the accent or dialect the dataset is for.
Content accuracy
The speaker said what was asked, with no missing or added words.

A dataset's quality score is the share of its reviewed recordings that reviewers approved.

Consent and provenance

Every contributor gives explicit consent during sign-up, with the date and time recorded.

That consent is stored against the contributor's account and travels with each recording as a consent ID and date in its JSON file.

Ways to work with us

Off-the-shelf datasets

License a dataset from the catalogue above.

Custom collection

We design and run a collection brief for the language, dialect or domain you need.

Human review of your data

Our native-speaker reviewers transcribe, label and quality-check audio you already have.

Talk to us

Tell us what you're building and our team will get back to you.