Midcentury · Audio & Speech Data

Thousands of hours of
real, multilingual speech.

Off-the-shelf and custom-manufactured speech corpora — large-scale conversational audio, customer-service interactions, multilingual and code-mixed speech, with configurable transcription and annotation.

69.5k
Hours, off-the-shelf
25+
Languages & varieties
24 kHz
Recording quality
~2%
Est. word error rate
1,000 h
Custom, per month
Aligned
Transcripts + labels
The catalog

Off-the-shelf and made-to-order

We offer both off-the-shelf audio datasets and custom-manufactured corpora. Core capabilities span large-scale conversational speech, customer-service interactions, multilingual and code-mixed speech, and configurable transcription and annotation.

The ready catalog totals roughly 69,500 hours across Hindi, Bengali, Urdu, Tamil, Gujarati, English, Japanese, Korean, Luxembourgish, and Hinglish, plus vertical customer-service speech — including a 50,000-hour hotel customer-service corpus and a 3,500+ hour Hindi corpus recorded at 24 kHz with time-aligned transcripts.

For anything not in the catalog, we manufacture to spec — Arabic varieties, major European languages, the Indic languages, and code-mixed speech — configurable by sample rate, channels, speaker profile, domain, and annotation depth, at roughly 1,000 hours per month.

Off-the-shelf

Ready to license

Conversational and customer-service corpora available now, sized by delivered hours.

Hours by dataset
Delivered hours available now · log scale
Hotel customer service
50,000
Luxembourgish
10,000
Hindi
3,500
English
1,000
Japanese
1,000
Korean
1,000
Hinglish (Hindi–English)
500
Bengali
500
Urdu
500
Tamil
500
Gujarati
500
English customer support
500
Share of catalog hours
By dataset · of ~69,500 h
Hotel customer service71.9%
Luxembourgish14.4%
Hindi5.0%
English1.4%
Japanese1.4%
Korean1.4%
Hinglish (Hindi–English)0.7%
Bengali0.7%
Urdu0.7%
Tamil0.7%
Gujarati0.7%
English customer support0.7%
Catalog detail
Language and delivery specs by dataset
DatasetLanguageDelivery specs
Hotel customer serviceEnglishVertical customer-service; specs & annotation configurable by subset
LuxembourgishLuxembourgishLarge-scale corpus; recording & transcription detail on request
HindiHindi24 kHz · 2-speaker · single/multichannel · time-aligned transcripts
EnglishEnglishLarge-scale conversational
JapaneseJapaneseLarge-scale conversational
KoreanKoreanLarge-scale conversational
Hinglish (Hindi–English)Hindi–EnglishCode-mixed · ≥22 kHz · 2-speaker · multichannel available
BengaliBengali24 kHz · 2-speaker · single/multichannel · time-aligned transcripts
UrduUrdu24 kHz · 2-speaker · single/multichannel · time-aligned transcripts
TamilTamil24 kHz · 2-speaker · single/multichannel · time-aligned transcripts
GujaratiGujarati24 kHz · 2-speaker · single/multichannel · time-aligned transcripts
English customer supportEnglishVertical customer-support · fully annotated

Available across the catalog: time-aligned transcription, single- or multi-channel, sample rates up to 24 kHz, and verbatim or normalized text — delivered in your preferred format.

Hindi transcription
3,500+ hours · internally fine-tuned ASR
24 kHz
Sample rate
~2%
Est. WER
3–5 days
Full-corpus ASR
Aligned
Single/multichannel

Recorded at 24 kHz, delivered with single-channel or multichannel aligned transcripts. Transcripts are produced with our internally fine-tuned ASR (estimated ~2% WER; full human QA pending) and can be completed for the full corpus in roughly three to five days. Audio and time-aligned text delivered in your preferred format; human review, enhanced QA, or additional annotation scoped separately.

Made to order

Custom collection & annotation

We manufacture bespoke conversational speech to your specification — configurable end to end, from acoustics to annotation.

Configurable around
Sample rate & formatSingle- or multichannelNumber of speakersConversational or scriptedDomain & scenarioSpeaker demographics & geographyOverlap & noise limitsVerbatim or normalized transcriptionSpeaker labels & time alignmentHuman QALinguistic & acoustic annotation
Languages we manufacture
On demand · configurable coverage
Modern Standard Arabic (ar-001)UAE Arabic (ar-AE)Saudi Arabic (ar-SA)SpanishFrenchGermanPortugueseItalianHindiBengaliTamilUrduMarathiGujaratiMalayalamKannadaPunjabiMandarin ChineseVietnameseIndonesianTurkishHinglish (code-mixed)

Each collection is tailored by use case, speaker profile, recording environment, channel configuration, and annotation requirements. Arabic varieties support configurable demographic and regional coverage.

Capacity

Built to scale

As a reference point, we have historically scaled custom collection programs at roughly 500 hours in 15 days — about 1,000 hours per month — subject to language, recording specs, speaker requirements, annotation depth, and quality control.

500 h
In 15 days
1,000 h
Per month
24 kHz
Up to, sample rate
Multichannel
Single or multi

Final availability, delivery timing, and pricing depend on the requested language, volume, licensing terms, technical specifications, and quality-assurance requirements.

Ask the team

Questions about the data?

Languages, hours, licensing, recording specifications, transcription and annotation formats, or a custom collection. Send a question and we'll reply by email.