Skip to content

Why Low-Resource Languages Hallucinate More Often

10 min read · updated August 11, 2026

“Less data, worse output” is true and almost useless. It does not explain why the failure takes the specific form of confident invention rather than obvious incompetence, and it does not explain why the same model that hedges carefully in English will assert something flatly false in Amharic.

The claim, stated precisely

The claim worth defending is narrower than the usual one. It is not that low-resource output is lower quality — that is obvious. It is that the ratio of fluency to accuracy is worse, and that this is a predictable consequence of how the two training stages differ in coverage.

A model produces a distribution over next tokens. Fluency is a property of that distribution being well-shaped locally: the model knows what words go next to what words. Accuracy is a property of the distribution being pinned to something outside the text. Local shape needs far less data than global pinning, so as you move down the resource curve, the two degrade at different rates. You get grammatical, idiomatic, entirely invented output long before you get output that looks broken.

What thin coverage actually removes

Here is the mechanism, concretely. Suppose a fact appears in English text ten thousand times and in Yoruba text twice. In English, the model’s representation of that fact is over-determined: the correct continuation is reinforced from many contexts, and the many near-miss continuations it also saw are outweighed. The distribution has a sharp peak.

In Yoruba, the same fact has essentially no mass of its own. What fills the gap is not silence — the model must still produce a distribution — it is whatever the language model prior suggests is a plausible continuation. That prior is shaped by all the other Yoruba text it saw, which is fluent and correct as language and unrelated as fact. The output is therefore locally well-formed and globally unanchored. That is exactly the definition of a hallucination, and it arrives by construction rather than by malfunction.

The same argument explains why the effect is worse for rare constructions than for common ones. A rare construction in a low-resource language is doubly rare, and the number of corrective examples — instances where the model would have been penalised for the wrong continuation during training — goes to zero. Joshi et al.’s ACL 2020 survey put over two thousand languages in a class with essentially no usable labelled or unlabelled data at all (The State and Fate of Linguistic Diversity in the NLP World). For those, there is no corrective mass anywhere.

The second skew nobody mentions

This is the part of the argument that changes what you do about it. Everyone reasons about pretraining coverage. Far fewer people reason about the coverage of the second stage — instruction tuning and preference optimisation — and that stage is where the behaviour we call “not making things up” is actually installed.

Pretraining data is scraped, so it is skewed by whatever the web is skewed by. Alignment data is written and rated by people who are paid to do it, which makes it skewed by labour markets, annotator recruitment, and the language the specification was written in. There is no reason to expect those two skews to be the same, and every reason to expect the second to be sharper: hiring a hundred thousand crawlable Yoruba web pages is a scraping problem, and hiring a hundred fluent Yoruba raters who agree on a preference rubric is an organisational one. The Aya initiative published by Cohere For AI in 2024 exists precisely because that second problem was not being solved at scale (The Aya Dataset), and the fact that building a multilingual instruction corpus was itself a publishable research contribution tells you how thin the prior state was.

So the calibrated behaviours — hedging, saying “I do not know”, refusing to invent a citation, declining to answer outside the provided context — are trained on a distribution even more English-concentrated than the pretraining mix. They are behaviours, and behaviours are conditioned on the language they were demonstrated in. A model that has been shown ten thousand English examples of admitting uncertainty and forty Yoruba ones has learned to admit uncertainty in English. That is the sharper cause, and it predicts something the naive account does not: that the gap between English and low-resource accuracy should be wider than the gap in raw fluency. Which is what practitioners report.

Falling back to the nearest neighbour

There is a second failure mode with the same root and a different shape. Multilingual models share representations across languages — that sharing is what makes cross-lingual transfer work at all and why a model can do anything in a language it barely saw. But sharing cuts both ways. When coverage of the target language runs out, the nearest thing in representation space is a related, better-resourced language, and the model drifts towards it.

The result is output that a non-speaker reads as the target language and a speaker reads as something else: Zulu with Xhosa forms, Quechua with Spanish syntax, a regional dialect flattened into the standard variety. It is not detectable by language identification, because the output genuinely is mostly the target language. It is one of the more insidious versions of the problem, and it is covered further in what a model does when it does not know a language.

What follows for a system you are building

  • Do not rely on the model to abstain. The abstention behaviour is the least transferable thing in the model. If your English path depends on the model saying “not in the provided documents”, your low-resource path needs that enforced outside the model — a retrieval gate, a schema, a confidence threshold you control.
  • Ground harder, not longer. Retrieval helps, but only if retrieval works, and multilingual embedding quality degrades on the same curve. Verify recall in the target language before attributing an error to generation.
  • Evaluate factuality separately from fluency. A human reviewer asked “is this good Yoruba?” will answer a different question from “is this true?”. Ask both, on the same items, and record them as two scores.
  • Prefer tasks that do not require the model to know things. Summarising a document you supplied, extracting fields, classifying, reformatting — these degrade gracefully. Open-ended factual generation degrades treacherously.