Benchmark
Public rankings, methodology, reports, and the African AI Trust Index.
benchmark.mlatho.comPublic evaluation · Preview edition
Rankings, methodology, and receipts for whether frontier AI systems actually work for African languages, local context, and safety — published by Mlatho.
Five headline scores journalists and product teams can cite — produced from a fixed, human-judged evaluation set.
Comprehension, instruction-following, and translation quality on African-language prompts (Preview focus: Chichewa ↔ English).
Checkable local geography, institutions, and norms — where fluent English often hides confident error.
Practical usefulness for African service and commerce contexts (channels, local constraints, usable answers).
Africa-relevant refusal and caution: scams, medical overclaim, election-adjacent heat, and demeaning content.
Weighted composite across language, culture, business readiness, and safety for this edition’s fixed set.
According to the Mlatho African AI Trust Index — models are scored under a frozen prompt template, without tools, by bilingual human raters.
Fetching the latest published scores.
The African AI Trust Index Preview is a small human-evaluated set focused on Chichewa and African context. Scores are for this set only, under a fixed prompt template without tools. They are not a universal ranking of model quality.
Lab invents. Benchmark proves. App monetizes. Learn trains the people who make it possible.
Public rankings, methodology, reports, and the African AI Trust Index.
benchmark.mlatho.comCommercial product: submit models, run jobs, download reports, manage projects.
app.mlatho.comResearch engine for new benchmarks, datasets, and experimental metrics.
lab.mlatho.comTrains the African evaluation workforce that produces gold and adjudication.
learn.mlatho.comProduct teams shipping to African users can commission a scoped private evaluation against an expanded set — same method discipline, confidential outputs.