Low-resource and niche speech for ASR, TTS and speech-to-speech, recorded with real speakers in real settings.
Most speech datasets cover the same few high-resource languages, so models slip the moment a user changes dialect, accent or setting. We record where that data doesn't exist yet: with recruited and partnered speakers in their homes, streets and workplaces, paid for every session and told how their audio will be used.
Current coverage includes [Language 1, Language 2, Language 3] and grows with demand. Tell us the languages, dialects and recording conditions your model needs, and we'll tell you what exists today and what we can collect on a defined timeline.
Hear exactly what you'd receive.
This is a real clip from our collection, shown with the transcript, metadata, quality scores and consent reference it ships with. The full sample set is shared during scoping.
Every batch is measured before it reaches you.
Thresholds are agreed with you in the spec. Each check uses a measure that is standard in speech research, so results can be compared against any other dataset you hold.
Reported per clip
Scores sit in each clip's metadata, so you can filter or re-weight the set in your own pipeline.
Summarised per batch
Each delivery includes a data card: coverage, score distributions, acceptance thresholds and known limitations, following the datasheet practice used for published ML datasets.
From a real speaker to your pipeline, in four steps.
Source
Speakers are recruited directly or through community partners where the language is spoken. Every clip is an original recording, never repurposed from another dataset or synthesised.
Consent
Speakers are briefed in the language they record in, paid for the session, and consent in writing to defined uses before anything is captured.
Record and refine
Prompts, environment and equipment follow the agreed protocol. Transcription and labelling run against your spec, and a pilot batch is accepted before collection scales.
Deliver
Audio, transcripts and metadata are packaged for your pipeline, with consent records and a processing log attached to every clip.
Rejected recordings are re-recorded, not patched. If a speaker withdraws, the request is traced through the same record, and the licence states what happens to data already delivered.