Model cards increasingly list African languages among those supported. The claim is usually technically true and practically misleading. Support in these announcements typically means the tokenizer handles the script and the model produces plausible-looking output, measured against a translation benchmark built from a narrow domain.
What that measurement misses is everything downstream. Tokenization efficiency for most African languages is poor, which means the same sentence costs several times more to process than its English equivalent — a direct financial penalty on building in those languages.
The data problem is not just volume
The available corpora are dominated by religious texts and translated news, which produces models that handle formal register and fail on the conversational speech that actual products need. For several major languages the largest clean corpus available is still measured in tens of millions of tokens.
Who is closing it
The most substantial work is being done by research collectives and university groups on the continent, generally on grant funding and often assembling data by hand. That work is feeding into models that commercial vendors then benefit from, largely without compensation or attribution.