African Languages and AI: Why the Gap Is Bigger Than Most Reports Admit

Benchmarks that report African language coverage tend to measure the wrong thing. What breaks is not translation quality but everything downstream of it.

Placeholder — written to give the site structure before launch. This is not reporting and it is not a finished article. It must be replaced with commissioned work before AI News Round goes live.

Model cards increasingly list African languages among those supported. The claim is usually technically true and practically misleading. Support in these announcements typically means the tokenizer handles the script and the model produces plausible-looking output, measured against a translation benchmark built from a narrow domain.

What that measurement misses is everything downstream. Tokenization efficiency for most African languages is poor, which means the same sentence costs several times more to process than its English equivalent — a direct financial penalty on building in those languages.

The data problem is not just volume

The available corpora are dominated by religious texts and translated news, which produces models that handle formal register and fail on the conversational speech that actual products need. For several major languages the largest clean corpus available is still measured in tens of millions of tokens.

Who is closing it

The most substantial work is being done by research collectives and university groups on the continent, generally on grant funding and often assembling data by hand. That work is feeding into models that commercial vendors then benefit from, largely without compensation or attribution.

Get the next one by email