The hard part is the code-switch.
Real patients don't speak clean monolingual sentences. They start in Bantu and switch to English mid-thought, often on exactly the clinical word that matters. Native speakers record body parts and health phrases in two modes: Bantu-only for clean pronunciation and Bantu & English for code-switch speech, with the switch point captured when the second mode is recorded.
One take. The switch, captured.
Click any consented take — it loads into the player and replays exactly as a patient speaks it: Bantu, then English, the words swapping at the captured switch point. The data self-segments at switch_ms, and the English half is Bantu-accented English you can validate your own model on.
Two takes per item
Bantu-only and Bantu & English code-switch audio for the same BTS-BH100 item — 48 kHz mono, with switch timing recorded as structured metadata.
Real people, fairly paid
8 named native contributors so far — found, compensated, and consented. Relationship work, not a crawl.
Growing nightly, honest
9,050 live consented takes to date — including 4,449 Bantu & English code-switch takes — across 11 languages, in active collection toward the full roster. We're plain about coverage because a rigorous reviewer should be.
Monolingual ASR/TTS corpora exist. Aligned, consented, clinically scoped code-switch audio for body parts and health phrases, keyed to the same concept IDs as the text and grammar layer, is the scarce layer. That alignment is what lets a voice model learn the switch, not just transcribe around it.
Hear the full catalogue under a pilot →