Multilingual generation quality
Where AI-generated text in a language goes subtly wrong — gender, register, regional variety, spelling convention — and what causes each failure.
A model that generates broken French produces something you can see. The failures in this cluster are the other kind: output that is fluent, idiomatic, confidently punctuated, and wrong in one morpheme. A feminine adjective on a masculine noun. A plural form that is correct for 22 and wrong for 12. A job title that defaults to the masculine three sentences after the text said “she”. None of it trips a validator, and none of it survives contact with a native reader.
These pages work each failure back to the grammar that causes it, because the grammar is what tells you whether the problem is fixable by prompting, fixable by post-processing, or not fixable at all without knowing something the model was never given. Where a figure is involved it is sourced to the specification or the dataset that publishes it; where a claim would need a measurement nobody made, the page says what to measure instead.
Gendered Nouns in French AI-Generated Text
Why a model assigns the wrong gender to a French noun, and why the error usually shows up on the adjective rather than on the noun itself.
9 min read
Writing Gender-Neutral Text in Spanish With AI
How to pin one gender-neutral convention in Spanish output and keep it consistent, including the words the convention cannot handle.
10 min read
Why Slavic Languages Need More Than One Plural Category
The dual number and the genitive-after-numeral rule that gave Russian and Polish their few/many split, and why the two languages disagree at 21.
9 min read
Polish Plural Forms for 1, 2–4 and 5+ in Generated Text
The exact count boundaries for Polish's four CLDR plural categories, worked through the edge cases, plus the two distinctions CLDR does not encode.
9 min read
German Grammatical Gender Agreement in AI-Generated Text
Why generated German gets the article right and the pronoun wrong, and the three places gender collides with case in the same ending.
10 min read
Gender Agreement Errors in Machine-Translated Arabic
Which slots in an Arabic sentence are forced to carry gender that the English source never specified, and what to do when the answer is not recoverable.
10 min read
Pluralization Rules in Welsh and Why They Break Standard Libraries
Welsh needs six plural categories, keeps the noun singular after numerals, and forms some singulars from plurals — three assumptions that break generic i18n code.
9 min read
Grammatical Gender in Generated Job Titles Across Languages
Why job titles default to the masculine even when the context says otherwise, and why structured data makes it worse than prose does.
9 min read
Why Japanese Has No Grammatical Plural and What That Means for AI Text
How Japanese expresses number without marking it on the noun, and the over-correction that puts a plural suffix where none belongs.
9 min read
Hindi Gender Agreement Between Verb and Subject in AI Text
Why generated Hindi puts the wrong gender on the verb, and why the ergative construction makes the verb agree with the object instead.
10 min read
How Many Plural Forms a Language Actually Needs
Languages grouped by their CLDR cardinal plural-category count, from one form to six, with what the categories actually mean and what they do not cover.
9 min read
What LLM Support for Swahili Actually Looks Like Today
What 'supports Swahili' decomposes into, which public datasets and corpus statistics you can check, and a probe set built around noun-class concord.
10 min read
What LLM Support for Yoruba Actually Looks Like Today
Yoruba has around 45 million speakers and a vanishing share of the web text models are trained on, and the gap shows up first in the tone diacritics.
9 min read
What LLM Support for Hausa Actually Looks Like Today
Hausa has more speakers than Italian and a fraction of its web presence, and the mismatch is compounded by an orthography that ordinary text pipelines quietly destroy.
9 min read
What LLM Support for Amharic Actually Looks Like Today
Amharic combines a large speaker population with a script that is expensive to tokenise and an orthography with built-in redundancy, and models lean on English loanwords where they run out of coverage.
9 min read
What LLM Support for Zulu Actually Looks Like Today
isiZulu is South Africa's most widely spoken home language and a language whose noun-class agreement gives you an unusually precise test of whether a model has really learned it.
9 min read
Why Low-Resource Languages Hallucinate More Often
Thin training coverage does not just make output worse; it removes the corrective pressure that teaches a model to stop, and the alignment data that would have taught it is thinner still.
10 min read
Building an AI Product for a Language With No Digital Corpus
A worked strategy for the case where the training data does not exist: build the evaluation set first, pivot through a related high-resource language, and put the glossary in the loop rather than in the model.
11 min read
Why Benchmark Coverage Skips Most of the World's Languages
The largest public multilingual benchmarks cover a couple of hundred languages out of roughly seven thousand, and the ones they cover are evaluated on narrower tasks than English is.
10 min read
Why Bengali Support Lags Despite Its Enormous Speaker Population
Bengali is among the most spoken languages on earth and behaves like a low-resource one, because speaker count measures demand and training corpora are built from supply.
10 min read
Why African Languages Are Underrepresented in AI Training Data
The shortfall is a property of what got published on the web, not of the languages, and the clearest evidence is that speakers are online in large numbers while writing in someone else's language.
10 min read
Crowdsourcing Training Data for a Low-Resource Language
A worked plan for collecting usable text or speech from a speaker community, built around the review-pair step that separates a corpus from a pile of noise.
11 min read
What LLM Support for Indigenous American Languages Looks Like Today
For most Indigenous American languages a model produces confident output that is not the language, and for several of them the communities have said the data should not be collected at all.
10 min read
Why Machine Translation Into Low-Resource Languages Still Fails Often
Most low-resource translation routes through English, and English cannot carry the distinctions the target language is obliged to mark, so the second leg guesses fluently.
10 min read
How Many of the World's Living Languages Have No AI Support at All
Subtracts the languages any major provider documents support for from Ethnologue's living-language count, and shows why the remainder is larger than the headline numbers suggest.
9 min read
Why Tokenizer Vocabulary Size Is a Bottleneck for Low-Resource Languages
How a fixed merge budget, allocated by corpus frequency, ends up giving hundreds of languages almost none of the vocabulary and why that shows up on your bill.
10 min read
What Happens When You Ask an AI Model a Question in a Language It Barely Knows
The specific failure modes a model produces in a language it has almost no training data for, why each one happens, and how to detect them without speaking the language.
10 min read
Getting AI Output in Egyptian Arabic Instead of Formal Arabic
The specific morphological markers that separate Egyptian colloquial from Modern Standard Arabic, and a prompt that holds a model to the dialect across a long response.
10 min read
Getting AI Output in Brazilian Portuguese Instead of European Portuguese
The pronoun, gerund and clitic-placement differences that make Portuguese read as Brazilian or European, and how to specify one reliably.
9 min read
Getting AI Output in Mexican Spanish Instead of Castilian Spanish
The pronoun system, vocabulary and register differences that separate Mexican Spanish from Peninsular Spanish, and how to decide whether you want either.
10 min read
Getting Simplified Chinese Output Instead of Traditional Chinese
The character-set split, the regional expectations behind it, and why one-to-many mappings make naive post-conversion produce wrong characters.
10 min read
Getting Swiss German Vocabulary Instead of Standard German
The difference between Swiss Standard German and Swiss German dialect, the orthographic rules that separate Swiss written German from German German, and which one you actually want.
9 min read
Getting British Spelling Instead of American Spelling From AI
Why a model reverts to American spelling partway through a long document, the full list of word classes that differ, and the deterministic fix that actually holds.
10 min read
Whether AI Models Can Generate Cantonese Instead of Mandarin
What genuine written Cantonese requires grammatically, why models tend to produce Standard Written Chinese instead, and a diagnostic you can run in two minutes.
10 min read
Getting Flemish Vocabulary Instead of Netherlands Dutch From AI
Belgian Dutch and Netherlands Dutch share one written standard, so what differs is vocabulary and register — and that is exactly what a generic Dutch prompt collapses.
9 min read
Getting Quebec French Instead of France French From AI
The terminology, typography and register differences that separate Quebec French from Metropolitan French, and a prompt that produces the first rather than the second.
10 min read
Norwegian Bokmål and Nynorsk in AI-Generated Text
Norwegian has two official written standards and no neutral one, so asking a model for “Norwegian” silently picks the majority form.
8 min read
Getting Austrian German Vocabulary Instead of Standard German From AI
Austrian German is a codified national standard with terms written into EU treaty law, and a model asked for German produces the German one.
8 min read
Why AI Models Rarely Recognise Scots as a Distinct Language
Scots has an ISO code and treaty recognition, and almost no clean digital corpus — which is why models treat it as English spelled oddly.
9 min read
Regional Arabic Dialects an AI Model Cannot Reliably Tell Apart
Arabic dialect identification has a documented accuracy ceiling, and the pairs that get confused are predictable from which markers a sentence happens to contain.
9 min read
Why AI-Generated Spanish Sounds Neutral Instead of Regional
Neutral Spanish is not an average the model computed — it is an editorial standard that already dominates the Spanish text the model was trained on.
9 min read
Getting Nigerian English Instead of Generic English From AI
Build a prompt that produces Nigerian Standard English — the right register, lexis and conventions — instead of American English or Pidgin.
9 min read
Why Singapore English (Singlish) Confuses AI Grammar Checkers
A grammar checker rewrites Singlish because no variety was declared and its default target is standard American English — here is the prompt that stops it.
9 min read
Writing Prompts Natively Instead of Translating Them
A translated prompt carries English structure into the target language, and the model continues in the register the prompt establishes.
9 min read
Writing Few-Shot Examples in the Target Language
Build a few-shot block whose examples are native, whose identifiers stay English, and whose typography is the target language's.
9 min read
Chain-of-Thought Language Mismatch: Reasoning in One Language, Answering in Another
Reasoning quality and answer fluency are separate variables, and the language that maximises one is often not the language that maximises the other.
9 min read
Does System Prompt Language Change Output Quality?
What is documented about system prompt language, what is not, and a test design that isolates it as a single variable on your own task.
9 min read
Prompting Effectively in a Low-Resource Language
Compensate for thin training coverage by supplying, explicitly, the structure a model would have inferred on its own in English.
10 min read
Other topics
- LLM fundamentals & architecture
- Tokens, tokenization & context windows
- Prompt engineering
- Reasoning models & test-time compute
- Multimodal AI: vision, audio, video
- RAG & retrieval
- Embeddings & vector search
- AI agents & tool use
- Structured output & function calling
- Fine-tuning & post-training
- Local inference errors, string by string
- Running local models day to day
- Testing code that calls an LLM
- Snapshot and property testing for model output
- Regression suites for prompts
- Eval gates in CI
- Flaky tests against a model
- Determinism and the cost of testing
- Contract and streaming tests
- Testing tool calls and retrieval
- Inference, serving & latency
- Rolling out a prompt change
- Testing AI systems in practice
- Forecasting a time series
- Machine learning on tabular data
- Geospatial data and models
- Understanding audio that is not speech
- Understanding video
- Core computer vision tasks
- Machine learning on graphs
- Point clouds and 3D
- Evaluation, benchmarks & LLM-as-judge
- Sensor and IoT data
- Logs and event streams
- Models over biological sequences
- Machine learning on molecules
- Embedding and searching code
- Extracting invoices and purchase orders
- Receipts, statements and tax forms
- Insurance policies and contracts
- Deeds, court filings and patents
- Extracting from medical records
- Observability & LLMOps
- CVs, certificates and identity documents
- Shipping, customs and technical documents
- Meetings, email, chat and filled-in forms
- Building an extraction pipeline
- Business, property and inspection documents
- Contract clauses and insurance claims
- Regulated and compliance documents
- Consumer, travel and closing documents
- Mapping one chat API onto another
- SDK and framework migrations
- Hallucination & failure modes
- Re-embedding and model deprecation
- Cutting over between providers
- Parity gaps, shims and legacy endpoints
- Moving between model versions
- Migrating vector stores and caches
- Mapping capabilities and parameters
- Migrating pipelines and agents
- Contracts, runbooks and rollback
- Auditing a codebase before a cutover
- Compliance and fine-tune migration
- LLM cost engineering
- Routing, cost tracking and multi-tenancy
- What a migration does to your prompts
- AI security & prompt injection
- Privacy, compliance & data residency
- AI governance, policy & society
- Building reliable AI applications
- AI hardware, GPUs & compute
- Open-weight models & local inference
- AI for developers & coding agents
- AI in industry: vertical playbooks
- AGI, superintelligence, alignment & the long future
- Machine learning foundations
- NLP fundamentals & classical tasks
- Data engineering for AI
- Synthetic data & dataset curation
- AI product design & UX
- Search, ranking & recommendation
- Enterprise adoption & change management
- AI careers, skills & teams
- Reading AI research
- AI in science & discovery
- Robotics & embodied AI
- AI economics, markets & business models
- AI myths, hype & media literacy
- Context engineering
- Shipping AI features: patterns & anti-patterns
- Build it: end-to-end AI tutorials
- Python for AI: hands-on recipes
- TypeScript, React and the web
- Frameworks and SDKs
- Errors and troubleshooting
- AI facts, numbers and statistics
- The history of AI
- The maths behind AI
- Architectures beyond the transformer
- Reinforcement learning
- Diffusion and generative media
- Speech, audio and voice engineering
- Benchmarks, one at a time
- AI search visibility
- Infrastructure and operations
- Databases and storage for AI
- Knowledge graphs and structured knowledge
- Classical ML in production
- Regulation, jurisdiction by jurisdiction
- Prompt recipes and pattern library
- AI for people who do not write code
- Writing, media and creative work
- Edge and on-device AI
- Interpretability and model internals
- Field notes
- OpenAI model behaviour
- Claude model behaviour
- Gemini model behaviour
- Llama model behaviour
- Mistral model behaviour
- Qwen model behaviour
- DeepSeek model behaviour
- Cohere model behaviour
- Grok model behaviour
- Small model behaviour
- Hybrid model architectures
- Token cost by language and script
- Transliteration, romanization and script handling
- Locale-correct output
- Multilingual pipelines
- The EU AI Act, article by article
- AI under the GDPR and EU data law
- US AI regulation, state and sector
- International AI governance and standards
- AI litigation and enforcement
- Running AI workloads on AWS
- Running AI workloads on Google Cloud
- Running AI workloads on Azure
- AI at the edge: Workers, Vercel and Netlify
- Serving models on Kubernetes
- Operating AI infrastructure
- Quantization formats and what they cost
- llama.cpp, flag by flag
- Ollama and the desktop local-model runtimes
- Local models on Apple Silicon
- Hardware for local inference
- Running speech and embedding models locally
- Model files, adapters and conversion
- VRAM arithmetic for local models