The frontier, scored on a patient's language.
The same three clinical prompts as the Clinical Test — name the body, understand the patient, produce the phrase a patient would say — run across the frontier and scored server-side against the attested BTS-BH100 corpus. Only scores are published; the answer key never leaves the server, so the task stays unleaked.
Headline findings of the cross-frontier run. One model translated “fainting” as the word for “to die.” In a consultation, a clinician cannot tell fabricated from correct — which is why fabrication is counted, not just misses.
Model × language, per subtest
Every cell hangs off the same English spine.
Every item in the protocol is an English-anchored BTS-BH100 concept: one English body part or patient-report sentence, aligned to an attested native form in each of the 21 languages. The English side is what every model sees — identical role in every column — so scores compare across languages, not just across models. The attested targets stay server-side.
bem is shown — the same public-demo norm as /explore. Every other attested form is held server-side: models are scored against them, never shown them.
These 18 English concepts are exactly what the scoreboard scores this language on — deterministic, band-stratified, identical for every model. The attested forms stay server-side.
Every concept carries a stable BTS id, so the corpus is English ↔ 21-language translation pairs out of the box — ready for alignment training, fine-tuning, and cross-lingual eval without any mapping work.
Because each language's items anchor to the same English spine, a Production cell in Bemba and one in isiZulu measure the same thing. Read the grid row-wise for a model, column-wise for a language — both are valid.
The same anchored concept indexes the text layer and both consented audio takes (clean Bantu + real code-switch, switch point captured). Hear one take do the work →
How the grid is produced
- Items are deterministic per language — band-stratified, alphabetical round-robin over the BTS-BH100 snapshot — so every model sees the identical 8 anatomy items and 10 phrases, and the run is reproducible.
- Prompts are exactly the three Clinical Test prompts (copyable at /clinical-test); no scaffolding, no examples, no retries.
- 5 scored reps per prompt, median reported — the same convention as the L26 alphabet benchmark. (Distinct benchmarks: L26 measures operating-alphabet recognition; this measures clinical language.)
- Scored server-side against the attested corpus. Verdicts: exact, fuzzy, fabricated (confident-wrong), missed. Scores only are released — the answer key, reveal rows, and corpus stay on the server.
A citable artifact
Plain citation:
BantuNomics (2026). The BTS-BH100 Clinical Benchmark (CB-1): frontier models on clinical language across 21 Bantu languages. 3MegaLabs. https://health.bantunomics.com/benchmark
BibTeX:
@misc{bantunomics2026clinical,
title = {The BTS-BH100 Clinical Benchmark (CB-1):
Frontier Models on Clinical Language in 21 Bantu Languages},
author = {{BantuNomics (3MegaLabs)}},
year = {2026},
url = {https://health.bantunomics.com/benchmark},
note = {Scores only; attested answer key held server-side. Machine-readable: /benchmark/data}
}
Machine-readable scores (no key): GET /benchmark/data
Your model isn't on the board? Score it yourself.
The same protocol is self-serve: three prompts, two minutes, scored against the same attested corpus. Then fix it — the consented reference the frontier is missing is what BTS-BH100 is.