License one clinical recording —
it does the work of six datasets.
A patient names a body part or symptom in their language, then crosses to the English clinical term — in one consented take. We capture the switch, the meaning, and the structure. In the one domain where a wrong word is a safety incident.
The economics of the clinical corpus
One clinical take does the work of six datasets.
A single consented, de-identified Bantu–English code-switch recording is not one labelled example. Captured in one continuous take, its switch point is logged as switch_ms — so the record self-segments with no manual annotation, and one take becomes the ground truth for six things frontier labs otherwise buy from six separate vendors.
Hear one take do the work.
The highlight swaps at the stored boundary as it plays. Play the halves separately — that's the segmentation a voice model needs. Open { } machine record to see it's real structured data, not a labelling job.
{
"bts_id": null,
"iso": "bem",
"variety_id": "BTS:VAR:0031",
"kind": "bantu_english",
"switch_ms": 13712,
"duration_ms": 17665,
"n_words": 13,
"segments": [
{
"label": "bantu",
"start_ms": 0,
"end_ms": 13712
},
{
"label": "english",
"start_ms": 13712,
"end_ms": 17665
}
],
"speaker": "Speaker BEM-A",
"consent": "granted",
"deidentified": true
}
What one clinical take is worth.
Each of these is a dataset a lab would otherwise buy or build separately. This one take is evidence for all of them — pre-structured, consented, de-identified.
The right Bantu word for the right body part or condition — homograph-safe, native-verified. Your model won't name the wrong organ.
Natural Bantu→English at the exact clinical term, pre-labelled with the stored switch point — training + eval data for how patients and clinicians actually speak.
The English half is spoken by native Bantu speakers — accented-English clinical audio that improves ASR fairness and robustness for Bantu-accented speech.
Each Bantu clinical word decomposes into its language's own Full Syllable Inventory — the tone-bearing units a forced aligner needs, grounded where tone actually lives.
Consented native pronunciations keyed to clinical meaning — a clean basis for patient-facing health voice and speech synthesis.
The English segment is isolated and machine-transcribable against a gold reference — a self-grading evaluation set. No labelling vendor, no drift.
Every concept is anchored to the interactive body figure — a concept→part→phrase grounding surface for multimodal clinical models. A seventh function no flat dataset can offer.
The benchmark grades itself.
We isolated the English half of each take and gave it to a frontier ASR (Google Cloud Speech-to-Text). It drops the discourse marker, mishears the clinical term, or returns nothing. That gap is what your model must close — and the gold answer key is built in.
Pre-segmented, consented, de-identified. A turnkey accent-robustness eval set — no labelling vendor, no drift.
The same structure, across the Bantu world.
Different languages, same captured switch — every take auto-extracts into Bantu · English · full code-switch.
{
"bts_id": null,
"iso": "loz",
"variety_id": "BTS:VAR:0340",
"kind": "bantu_english",
"switch_ms": 13707,
"duration_ms": 23168,
"n_words": 19,
"segments": [
{
"label": "bantu",
"start_ms": 0,
"end_ms": 13707
},
{
"label": "english",
"start_ms": 13707,
"end_ms": 23168
}
],
"speaker": "Speaker LOZ-A",
"consent": "granted",
"deidentified": true
}
{
"bts_id": null,
"iso": "lue",
"variety_id": "BTS:VAR:0350",
"kind": "bantu_english",
"switch_ms": 8913,
"duration_ms": 40034,
"n_words": 11,
"segments": [
{
"label": "bantu",
"start_ms": 0,
"end_ms": 8913
},
{
"label": "english",
"start_ms": 8913,
"end_ms": 40034
}
],
"speaker": "Speaker LUE-A",
"consent": "granted",
"deidentified": true
}
{
"bts_id": null,
"iso": "lug",
"variety_id": "BTS:VAR:0351",
"kind": "bantu_english",
"switch_ms": 18988,
"duration_ms": 25819,
"n_words": 18,
"segments": [
{
"label": "bantu",
"start_ms": 0,
"end_ms": 18988
},
{
"label": "english",
"start_ms": 18988,
"end_ms": 25819
}
],
"speaker": "Speaker LUG-A",
"consent": "granted",
"deidentified": true
}
{
"bts_id": null,
"iso": "lun",
"variety_id": "BTS:VAR:0354",
"kind": "bantu_english",
"switch_ms": 19468,
"duration_ms": 25574,
"n_words": 24,
"segments": [
{
"label": "bantu",
"start_ms": 0,
"end_ms": 19468
},
{
"label": "english",
"start_ms": 19468,
"end_ms": 25574
}
],
"speaker": "Speaker LUN-A",
"consent": "granted",
"deidentified": true
}
Language infrastructure — not patient data.
Be the lab whose health AI is actually safe in Bantu.
License the program, not a dataset — one consented clinical take, six functions, across the Bantu family.
Start a conversation