BantuNomics Sign in
B
BantuNomics
Body & Health · BTS-BH100
Partnership
The Clinical Benchmark · CB-1 · citable

The frontier, scored on a patient's language.

The same three clinical prompts as the Clinical Test — name the body, understand the patient, produce the phrase a patient would say — run across the frontier and scored server-side against the attested BTS-BH100 corpus. Only scores are published; the answer key never leaves the server, so the task stays unleaked.

17 named frontier models 21 Bantu languages 3 subtests · Anatomy / Comprehension / Production 5 scored reps per prompt scored on the attested corpus
1 / 12 attested clinical phrases the best frontier model produced
0 what most models produced — while fabricating confidently

Headline findings of the cross-frontier run. One model translated “fainting” as the word for “to die.” In a consultation, a clinician cannot tell fabricated from correct — which is why fabrication is counted, not just misses.

The scoreboard

Model × language, per subtest

≥ 70% 40–69% 1–39% 0% — nothing attested produced — not yet published Hover a cell for the count breakdown (incl. fabrications).
The English anchor · why columns compare

Every cell hangs off the same English spine.

Every item in the protocol is an English-anchored BTS-BH100 concept: one English body part or patient-report sentence, aligned to an attested native form in each of the 21 languages. The English side is what every model sees — identical role in every column — so scores compare across languages, not just across models. The attested targets stay server-side.

One anchor · 21 attested targets

bem is shown — the same public-demo norm as /explore. Every other attested form is held server-side: models are scored against them, never shown them.

The spine, per language

These 18 English concepts are exactly what the scoreboard scores this language on — deterministic, band-stratified, identical for every model. The attested forms stay server-side.

🧭
Aligned pairs, by construction

Every concept carries a stable BTS id, so the corpus is English ↔ 21-language translation pairs out of the box — ready for alignment training, fine-tuning, and cross-lingual eval without any mapping work.

📐
Comparable columns

Because each language's items anchor to the same English spine, a Production cell in Bemba and one in isiZulu measure the same thing. Read the grid row-wise for a model, column-wise for a language — both are valid.

🔗
One anchor, six datasets

The same anchored concept indexes the text layer and both consented audio takes (clean Bantu + real code-switch, switch point captured). Hear one take do the work →

Protocol

How the grid is produced

  • Items are deterministic per language — band-stratified, alphabetical round-robin over the BTS-BH100 snapshot — so every model sees the identical 8 anatomy items and 10 phrases, and the run is reproducible.
  • Prompts are exactly the three Clinical Test prompts (copyable at /clinical-test); no scaffolding, no examples, no retries.
  • 5 scored reps per prompt, median reported — the same convention as the L26 alphabet benchmark. (Distinct benchmarks: L26 measures operating-alphabet recognition; this measures clinical language.)
  • Scored server-side against the attested corpus. Verdicts: exact, fuzzy, fabricated (confident-wrong), missed. Scores only are released — the answer key, reveal rows, and corpus stay on the server.
Cite this

A citable artifact

Plain citation:

BantuNomics (2026). The BTS-BH100 Clinical Benchmark (CB-1): frontier models on clinical language across 21 Bantu languages. 3MegaLabs. https://health.bantunomics.com/benchmark

BibTeX:

@misc{bantunomics2026clinical,
  title  = {The BTS-BH100 Clinical Benchmark (CB-1):
            Frontier Models on Clinical Language in 21 Bantu Languages},
  author = {{BantuNomics (3MegaLabs)}},
  year   = {2026},
  url    = {https://health.bantunomics.com/benchmark},
  note   = {Scores only; attested answer key held server-side. Machine-readable: /benchmark/data}
}

Machine-readable scores (no key): GET /benchmark/data

Your model isn't on the board? Score it yourself.

The same protocol is self-serve: three prompts, two minutes, scored against the same attested corpus. Then fix it — the consented reference the frontier is missing is what BTS-BH100 is.

Access

Start where you are.

Prove it free, validate it on your own data, or license the whole program.

01 · Free
Evaluation

Score your model against the foundational layer — the Alphabet Test and the L26 Lite suite, with saved results. Self-serve, no cost.

Start evaluating →
Most labs start here
02 · 75 days
Validation Pilot

Three languages you choose — and everything we hold for them. Measured on your own held-out data.

Scope a pilot →
03 · Program
Full Annual Subscription

Every product, every language, the full consented corpus — and everything curated while you're subscribed.

Start a Full Annual Subscription →