Body parts and health symptoms,
standardized across 21 Bantu languages.
BantuNomics maps body parts, symptom families, patient-report phrases, noun-class grammar, plural forms, concord patterns, and native audio into one governed substrate. It gives health AI the clinical language layer it cannot scrape: what hurts, where it hurts, how a patient says it, and how the sentence must agree in the language.
| Language | Country | Speakers | Body terms | Phrases |
|---|---|---|---|---|
🇰🇪 Swahili (Kenya) Kiswahili · G42 |
Kenya · Tanzania | 100M | 80 | 174 |
🇺🇬 Swahili (Uganda) Kiswahili · G42 |
Uganda | 100M | 80 | 174 |
🇷🇼 Kinyarwanda Ikinyarwanda · JD61 |
Rwanda · DR Congo · Uganda | 15M | 80 | 173 |
🇲🇼 Chewa / Chichewa Chichewa · N31 |
Malawi · Zambia · Mozambique | 14M | 80 | 175 |
🇱🇸 Sesotho Sesotho · S33 |
Lesotho · South Africa | 13.7M | 85 | 175 |
🇧🇮 Kirundi Ikirundi · JD62 |
Burundi · Rwanda | 12M | 81 | 172 |
🇿🇦 isiZulu isiZulu · S42 |
South Africa | 12M | 91 | 175 |
🇺🇬 Luganda Oluganda · JE15 |
Uganda | 10M | 80 | 175 |
🇿🇼 Shona chiShona · S10 |
Zimbabwe · Mozambique | 9M | 80 | 175 |
🇧🇼 Setswana Setswana · S31 |
Botswana · South Africa | 8.2M | 80 | 175 |
🇿🇦 isiXhosa isiXhosa · S41 |
South Africa | 8.2M | 85 | 175 |
🇰🇪 Gikuyu Gĩkũyũ · E51 |
Kenya | 8.1M | 87 | 175 |
🇿🇦 Tsonga (Xitsonga) Xitsonga · S53 |
South Africa · Mozambique · Zimbabwe | 7M | 80 | 175 |
🇿🇦 Sepedi (Northern Sotho) Sesotho sa Leboa · S32 |
South Africa | 4.7M | 80 | 175 |
🇿🇲 Bemba IciBemba · M42 |
Zambia · DR Congo | 4.1M | 100 | 185 |
🇺🇬 Runyankore Runyankore · JE13 |
Uganda | 3.4M | 80 | 174 |
🇸🇿 Siswati (siSwati) siSwati · S43 |
Eswatini · South Africa | 2.4M | 108 | 185 |
🇿🇼 Northern Ndebele (isiNdebele) isiNdebele · S408 |
Zimbabwe | 1.6M | 80 | 175 |
🇿🇲 Luvale Luvale · K14 |
Zambia · Angola | 1M | 80 | 175 |
🇿🇲 Chilunda (Lunda) Chilunda · K14 |
Zambia · Angola · DR Congo | 500K | 80 | 175 |
🇿🇲 Kaonde (Kikaonde) Kikaonde · L41 |
Zambia | 240K | 80 | 175 |
Views cycle on their own — front, back, internal, brain; the language rotates, or pick one above. Click a part to inspect it.
Fluent isn't safe.
Bantu languages aren't built like English. Meaning is carried by syllables, noun classes, and verb grammar that flat training text quietly throws away. So a model can sound fluent and still be unable to do the basic clinical things — and, worse, it doesn't know it can't. It answers confidently. In a consultation, confidently wrong is dangerous.
It names the wrong body part.
Ask for "calf" and a model hands back the word for a baby cow. One spelling, many meanings — and no way to tell which one a clinician means.
It can't form the sentence.
"The joints ache" needs noun-class agreement the model never learned. It produces grammatical nonsense a patient won't understand — or trust.
It can't hear the patient.
Real patients switch between Bantu and English mid-sentence. There is almost no data teaching a model to follow it — so it mishears the symptom.
We ran the same clinical prompts across frontier models. One translated “fainting” as the word for “to die.” In a consultation, a clinician cannot tell fabricated from correct.
This isn't a long-tail edge case. It's hundreds of millions of speakers — ~235M across these 21 languages alone — and the exact place your health AI most needs to be right.
A living platform, growing on three axes.
All aligned to the same English anchor — every one is Anchored. New languages join comparable from day one.
143 concepts + 185 phrases. A living standard — it widens by version.
Languages with native audio underway — deepening Anchored → Filled → Voiced → Verified.
There is no dataset for this.
Clinical Bantu, natively attested and grammatically correct, isn't sitting on the web waiting to be crawled. To build it yourself you would need all of this — and years:
A linguistic standard
one shared structure so 21 languages line up — not 21 incompatible word lists.
Native speakers, 20 countries
found, paid, and verified — for terms, grammar, and voice. Relationship work, not a crawl.
Consent infrastructure
informed, revocable, de-identified — the gate your own safety review will demand.
A recording program
native voice, including the code-switch your models have never heard.
We already did. Here's what your model gets ↓
One consented recording does the work of six datasets.
A patient names a symptom in their language, then crosses to the English clinical term — in one take. We capture the switch, the meaning, and the structure.
Not a glossary. A structured clinical corpus.
Every concept is addressable by a stable ID and carries layers — terms, phrases, verbs, grammar, attested evidence, voice. You license queryable infrastructure, not a flat file.
The clinical map
286 concepts spanning 7 domains — anatomy to symptoms to care.The full clinical domain is mapped. Anatomy is deep today; the symptom, condition and care layers are scaffolded and fill with each partnership cycle — which is exactly what a subscription funds.
Pick a body part. Watch your model's blind spot fill in.
One clinical concept, resolved across every language at the same time — the family-scale alignment your model needs and can't get anywhere else.
Every word carries a stable ID — so you can license and query this like infrastructure, not a spreadsheet. See all 286 concepts →
It doesn't just match a word. It builds the sentence.
A word list can't reason. This carries the verbs and grammar a model needs to produce a correct clinical sentence — and the attested evidence to check itself against.
Every part, correctly named
so it stops guessing the body
The sentences patients say
real clinical phrasing, not paraphrase
The verbs to build with
generate, don't echo
The grammar that makes it correct
concord, not nonsense
Verbs used in context
how a clinician actually says it
Dictionary-grounded evidence
attested, not invented
Your model doesn't retrieve these. It generates them.
Each below is a real clinical sentence built from the health-verb grammar — same machinery, any symptom, in every language. The morphology equation is the recipe.
193 applied verb frames · 561 attested lexemes · 11 training-ready views — a generative engine, not a lookup table.
different real ways to say the same clinical thing — kept, not collapsed. Your model learns how people actually speak across the family, each source-attributed to a named speaker.
native speakers building and verifying these languages so far — fairly paid, consent-first. This is sourced, not scraped.
And we're teaching it to hear the patient.
Native speakers record every concept and phrase twice — once in Bantu, once as a real patient speaks it: Bantu, then switching to English mid-sentence — with the exact switch point captured. It's the code-switch data your voice models have never had. Consent-first; in active collection, scaling with partnership.
- ✓ Native voice at 48 kHz · two modes per item
- ✓ The switch point captured — clean for ASR, TTS, translation
- ✓ Consent-first · de-identified · grows with the partnership
The Bantu and English halves are separately usable — exactly what a voice model needs to learn the switch.
It clears your safety review before you ask.
Body-part words, symptom phrasing and clinical language — never medical records or identifiable patient information. It sits outside HIPAA / medical-records regimes. It's the substrate your health AI is built on, not health data itself — so it won't drag a regulatory review.
Informed consent
Every contributor consents — including voice — and can revoke. Deletion-on-request is honored.
De-identified by contract
Speaker identity, raw audio paths and internal IDs never leave the server.
Fairly sourced
Native contributors fairly compensated; voluntary, with grievance and withdrawal channels.
This is the public view. There's a lot more inside.
A Pilot opens the data room across every domain: per-record provenance, the full per-language matrix, native audio samples (Bantu-only + code-switch), live coverage, and exports — to evaluate on your own infrastructure. A Full Annual Subscription license is the entire program, uncapped.
Be the lab whose health AI is actually safe in Bantu.
One substrate, three reasons it matters to you. Evaluate it on your own infrastructure — data access only, no black box — then license it family-wide.
Anthropic
A verified, consented clinical lexicon your model won't hallucinate — the grounding that makes Bantu health assistance safe to ship.
OpenAI
Family-wide breadth plus the code-switch audio your voice models need for real African clinical speech.
Google · DeepMind
Concept→phrase→verb structure and a body-figure grounding surface for multilingual clinical models and Africa health programs.
See your own model fail it — then fix it.
Scoped, de-identified data access — no black box. A Pilot opens every domain, sampled, including this one; the Full Annual Subscription license is the entire program.
Start your free evaluation →Language infrastructure — not patient data / not PHI. Native-speaker audio program in active, consent-first collection.