Skulp Contact
Programmes / Voice

Real voices, in the languages and accents your model still gets wrong.

Available Samples on this page Commissioned collection open

Low-resource and niche speech for ASR, TTS and speech-to-speech, recorded with real speakers in real settings.

Most speech datasets cover the same few high-resource languages, so models slip the moment a user changes dialect, accent or setting. We record where that data doesn't exist yet: with recruited and partnered speakers in their homes, streets and workplaces, paid for every session and told how their audio will be used.

Current coverage includes [Language 1, Language 2, Language 3] and grows with demand. Tell us the languages, dialects and recording conditions your model needs, and we'll tell you what exists today and what we can collect on a defined timeline.

Built for ASR · TTS · Speech-to-speech · Speech translation · Evaluation sets
Define what your model needs
01
Language and dialect coverage
Current languages are listed above. New languages and dialects can be commissioned.
02
Speech type
Read, prompted or spontaneous; monologue or conversational.
03
Recording conditions
Quiet rooms or in-the-field environments, and the device profile you need.
04
Speaker mix
Distribution across age, gender and region, set as part of the spec.
05
Transcription, translation and labelling
Agreed per engagement against your annotation spec. Transcripts in the source script, with optional English translation.
06
Licensing
Permitted use, term and exclusivity, with consent records attached.
Sample record

Hear exactly what you'd receive.

This is a real clip from our collection, shown with the transcript, metadata, quality scores and consent reference it ships with. The full sample set is shared during scoping.

Audio · transcript
[clip id]
[mm:ss]
[Audio file to be added]
Transcript · original
[transcript in source language]
Translation · English
[English translation]
Metadata
Clip ID[clip id]
Language[Language 1 · dialect]
Speech type[read / prompted / spontaneous]
Speaker[age band · gender · region]
Environment[room / field · device]
Format[sample rate · bit depth · channels]
Duration[mm:ss]
Quality[SNR in dB · MOS score]
Consent ref[consent record id]
Quality and grading

Every batch is measured before it reaches you.

Thresholds are agreed with you in the spec. Each check uses a measure that is standard in speech research, so results can be compared against any other dataset you hold.

01

Signal

Every clip is checked for signal-to-noise ratio, clipping, long silences, and the agreed sample rate and channel layout. Clips that fail are re-recorded.

SNR · clipping · silence · format
02

Transcript accuracy

A second reviewer transcribes a share of each batch independently. Disagreement is measured as word or character error rate, and anything above the agreed threshold goes back for correction.

WER · CER · reviewer agreement
03

Listening quality

Trained listeners rate a sample of clips on a five-point scale, following the mean opinion score method used across speech research.

MOS (ITU-T P.800)
04

Coverage

The delivered mix of speakers, dialects, speech types and environments is compared against the spec, so gaps show before a batch is accepted.

Speaker · dialect · condition balance

Reported per clip

Scores sit in each clip's metadata, so you can filter or re-weight the set in your own pipeline.

Summarised per batch

Each delivery includes a data card: coverage, score distributions, acceptance thresholds and known limitations, following the datasheet practice used for published ML datasets.

In practice

From a real speaker to your pipeline, in four steps.

01

Source

Speakers are recruited directly or through community partners where the language is spoken. Every clip is an original recording, never repurposed from another dataset or synthesised.

02

Consent

Speakers are briefed in the language they record in, paid for the session, and consent in writing to defined uses before anything is captured.

03

Record and refine

Prompts, environment and equipment follow the agreed protocol. Transcription and labelling run against your spec, and a pilot batch is accepted before collection scales.

04

Deliver

Audio, transcripts and metadata are packaged for your pipeline, with consent records and a processing log attached to every clip.

Rejected recordings are re-recorded, not patched. If a speaker withdraws, the request is traced through the same record, and the licence states what happens to data already delivered.

Tell us what your model is missing.

Send the languages, dialects, recording conditions, volume and timeline. You'll hear back within two working days, including if it's outside what we do.

Scope a collection hello@skulp.ai