Skip to content

Multilingual generation quality

Where AI-generated text in a language goes subtly wrong — gender, register, regional variety, spelling convention — and what causes each failure.

A model that generates broken French produces something you can see. The failures in this cluster are the other kind: output that is fluent, idiomatic, confidently punctuated, and wrong in one morpheme. A feminine adjective on a masculine noun. A plural form that is correct for 22 and wrong for 12. A job title that defaults to the masculine three sentences after the text said “she”. None of it trips a validator, and none of it survives contact with a native reader.

These pages work each failure back to the grammar that causes it, because the grammar is what tells you whether the problem is fixable by prompting, fixable by post-processing, or not fixable at all without knowing something the model was never given. Where a figure is involved it is sourced to the specification or the dataset that publishes it; where a claim would need a measurement nobody made, the page says what to measure instead.

Gendered Nouns in French AI-Generated Text

Why a model assigns the wrong gender to a French noun, and why the error usually shows up on the adjective rather than on the noun itself.

9 min read

Writing Gender-Neutral Text in Spanish With AI

How to pin one gender-neutral convention in Spanish output and keep it consistent, including the words the convention cannot handle.

10 min read

Why Slavic Languages Need More Than One Plural Category

The dual number and the genitive-after-numeral rule that gave Russian and Polish their few/many split, and why the two languages disagree at 21.

9 min read

Polish Plural Forms for 1, 2–4 and 5+ in Generated Text

The exact count boundaries for Polish's four CLDR plural categories, worked through the edge cases, plus the two distinctions CLDR does not encode.

9 min read

German Grammatical Gender Agreement in AI-Generated Text

Why generated German gets the article right and the pronoun wrong, and the three places gender collides with case in the same ending.

10 min read

Gender Agreement Errors in Machine-Translated Arabic

Which slots in an Arabic sentence are forced to carry gender that the English source never specified, and what to do when the answer is not recoverable.

10 min read

Pluralization Rules in Welsh and Why They Break Standard Libraries

Welsh needs six plural categories, keeps the noun singular after numerals, and forms some singulars from plurals — three assumptions that break generic i18n code.

9 min read

Grammatical Gender in Generated Job Titles Across Languages

Why job titles default to the masculine even when the context says otherwise, and why structured data makes it worse than prose does.

9 min read

Why Japanese Has No Grammatical Plural and What That Means for AI Text

How Japanese expresses number without marking it on the noun, and the over-correction that puts a plural suffix where none belongs.

9 min read

Hindi Gender Agreement Between Verb and Subject in AI Text

Why generated Hindi puts the wrong gender on the verb, and why the ergative construction makes the verb agree with the object instead.

10 min read

How Many Plural Forms a Language Actually Needs

Languages grouped by their CLDR cardinal plural-category count, from one form to six, with what the categories actually mean and what they do not cover.

9 min read

What LLM Support for Swahili Actually Looks Like Today

What 'supports Swahili' decomposes into, which public datasets and corpus statistics you can check, and a probe set built around noun-class concord.

10 min read

What LLM Support for Yoruba Actually Looks Like Today

Yoruba has around 45 million speakers and a vanishing share of the web text models are trained on, and the gap shows up first in the tone diacritics.

9 min read

What LLM Support for Hausa Actually Looks Like Today

Hausa has more speakers than Italian and a fraction of its web presence, and the mismatch is compounded by an orthography that ordinary text pipelines quietly destroy.

9 min read

What LLM Support for Amharic Actually Looks Like Today

Amharic combines a large speaker population with a script that is expensive to tokenise and an orthography with built-in redundancy, and models lean on English loanwords where they run out of coverage.

9 min read

What LLM Support for Zulu Actually Looks Like Today

isiZulu is South Africa's most widely spoken home language and a language whose noun-class agreement gives you an unusually precise test of whether a model has really learned it.

9 min read

Why Low-Resource Languages Hallucinate More Often

Thin training coverage does not just make output worse; it removes the corrective pressure that teaches a model to stop, and the alignment data that would have taught it is thinner still.

10 min read

Building an AI Product for a Language With No Digital Corpus

A worked strategy for the case where the training data does not exist: build the evaluation set first, pivot through a related high-resource language, and put the glossary in the loop rather than in the model.

11 min read

Why Benchmark Coverage Skips Most of the World's Languages

The largest public multilingual benchmarks cover a couple of hundred languages out of roughly seven thousand, and the ones they cover are evaluated on narrower tasks than English is.

10 min read

Why Bengali Support Lags Despite Its Enormous Speaker Population

Bengali is among the most spoken languages on earth and behaves like a low-resource one, because speaker count measures demand and training corpora are built from supply.

10 min read

Why African Languages Are Underrepresented in AI Training Data

The shortfall is a property of what got published on the web, not of the languages, and the clearest evidence is that speakers are online in large numbers while writing in someone else's language.

10 min read

Crowdsourcing Training Data for a Low-Resource Language

A worked plan for collecting usable text or speech from a speaker community, built around the review-pair step that separates a corpus from a pile of noise.

11 min read

What LLM Support for Indigenous American Languages Looks Like Today

For most Indigenous American languages a model produces confident output that is not the language, and for several of them the communities have said the data should not be collected at all.

10 min read

Why Machine Translation Into Low-Resource Languages Still Fails Often

Most low-resource translation routes through English, and English cannot carry the distinctions the target language is obliged to mark, so the second leg guesses fluently.

10 min read

How Many of the World's Living Languages Have No AI Support at All

Subtracts the languages any major provider documents support for from Ethnologue's living-language count, and shows why the remainder is larger than the headline numbers suggest.

9 min read

Why Tokenizer Vocabulary Size Is a Bottleneck for Low-Resource Languages

How a fixed merge budget, allocated by corpus frequency, ends up giving hundreds of languages almost none of the vocabulary and why that shows up on your bill.

10 min read

What Happens When You Ask an AI Model a Question in a Language It Barely Knows

The specific failure modes a model produces in a language it has almost no training data for, why each one happens, and how to detect them without speaking the language.

10 min read

Getting AI Output in Egyptian Arabic Instead of Formal Arabic

The specific morphological markers that separate Egyptian colloquial from Modern Standard Arabic, and a prompt that holds a model to the dialect across a long response.

10 min read

Getting AI Output in Brazilian Portuguese Instead of European Portuguese

The pronoun, gerund and clitic-placement differences that make Portuguese read as Brazilian or European, and how to specify one reliably.

9 min read

Getting AI Output in Mexican Spanish Instead of Castilian Spanish

The pronoun system, vocabulary and register differences that separate Mexican Spanish from Peninsular Spanish, and how to decide whether you want either.

10 min read

Getting Simplified Chinese Output Instead of Traditional Chinese

The character-set split, the regional expectations behind it, and why one-to-many mappings make naive post-conversion produce wrong characters.

10 min read

Getting Swiss German Vocabulary Instead of Standard German

The difference between Swiss Standard German and Swiss German dialect, the orthographic rules that separate Swiss written German from German German, and which one you actually want.

9 min read

Getting British Spelling Instead of American Spelling From AI

Why a model reverts to American spelling partway through a long document, the full list of word classes that differ, and the deterministic fix that actually holds.

10 min read

Whether AI Models Can Generate Cantonese Instead of Mandarin

What genuine written Cantonese requires grammatically, why models tend to produce Standard Written Chinese instead, and a diagnostic you can run in two minutes.

10 min read

Getting Flemish Vocabulary Instead of Netherlands Dutch From AI

Belgian Dutch and Netherlands Dutch share one written standard, so what differs is vocabulary and register — and that is exactly what a generic Dutch prompt collapses.

9 min read

Getting Quebec French Instead of France French From AI

The terminology, typography and register differences that separate Quebec French from Metropolitan French, and a prompt that produces the first rather than the second.

10 min read

Norwegian Bokmål and Nynorsk in AI-Generated Text

Norwegian has two official written standards and no neutral one, so asking a model for “Norwegian” silently picks the majority form.

8 min read

Getting Austrian German Vocabulary Instead of Standard German From AI

Austrian German is a codified national standard with terms written into EU treaty law, and a model asked for German produces the German one.

8 min read

Why AI Models Rarely Recognise Scots as a Distinct Language

Scots has an ISO code and treaty recognition, and almost no clean digital corpus — which is why models treat it as English spelled oddly.

9 min read

Regional Arabic Dialects an AI Model Cannot Reliably Tell Apart

Arabic dialect identification has a documented accuracy ceiling, and the pairs that get confused are predictable from which markers a sentence happens to contain.

9 min read

Why AI-Generated Spanish Sounds Neutral Instead of Regional

Neutral Spanish is not an average the model computed — it is an editorial standard that already dominates the Spanish text the model was trained on.

9 min read

Getting Nigerian English Instead of Generic English From AI

Build a prompt that produces Nigerian Standard English — the right register, lexis and conventions — instead of American English or Pidgin.

9 min read

Why Singapore English (Singlish) Confuses AI Grammar Checkers

A grammar checker rewrites Singlish because no variety was declared and its default target is standard American English — here is the prompt that stops it.

9 min read

Writing Prompts Natively Instead of Translating Them

A translated prompt carries English structure into the target language, and the model continues in the register the prompt establishes.

9 min read

Writing Few-Shot Examples in the Target Language

Build a few-shot block whose examples are native, whose identifiers stay English, and whose typography is the target language's.

9 min read

Chain-of-Thought Language Mismatch: Reasoning in One Language, Answering in Another

Reasoning quality and answer fluency are separate variables, and the language that maximises one is often not the language that maximises the other.

9 min read

Does System Prompt Language Change Output Quality?

What is documented about system prompt language, what is not, and a test design that isolates it as a single variable on your own task.

9 min read

Prompting Effectively in a Low-Resource Language

Compensate for thin training coverage by supplying, explicitly, the structure a model would have inferred on its own in English.

10 min read

Other topics