Which Languages Dominate AI Training Data, and by How Much
9 min read · updated August 11, 2026
The honest answer has two parts: the ranking is well established and extremely skewed, and the precise per-language shares for any recent frontier model are not public. Everything below is sourced to something that was actually published, with its date attached.
What is actually published
The cleanest composition figure for a frontier-scale model comes from the GPT-3 paper, “Language Models are Few-Shot Learners” (Brown et al., 2020). Its appendix reports the language breakdown of the training corpus, and English accounts for approximately 93% of the training words. The remaining 7% is spread across every other language in the mix, with German, French, Portuguese and Italian at the top of the remainder — each around 1% or less. The paper is on arXiv as 2005.14165.
That figure is from 2020 and it is the last one of its kind from a leading lab. Model cards for GPT-4, Claude, and Gemini do not publish a per-language composition table for their pretraining corpora, and the Llama 3 model card describes its multilingual data qualitatively rather than as a distribution. So a page that quotes a current “percentage of training data that is English” for a 2025 or 2026 model is quoting something nobody released.
The Common Crawl distribution
What is still published, continuously, is the composition of the web corpus that underlies most pretraining mixes. The Common Crawl Foundation runs language detection over the extracted text of each crawl and publishes the resulting distribution in its crawl statistics. Those statistics are public per crawl, which means the figure can be checked rather than trusted.
The stable findings across crawls, rather than any single crawl’s number:
- English is a plurality of the corpus by a wide margin — consistently over 40% of documents, and the largest single language in every crawl.
- The second tier is European plus Russian — Russian, German, Spanish, French, Japanese and Chinese each occupy a few percent, an order of magnitude below English.
- The tail falls off very fast. Below the top twenty or so languages, individual shares are fractions of a percent, and the distribution has the long tail characteristic of web content generally.
- The detector itself is part of the measurement. Common Crawl’s statistics come from an automatic language identifier, which under-detects exactly the languages this question is about: short documents, romanised text, code-switched pages and low-resource languages the detector does not have a label for. The tail is therefore under-counted in a direction that flatters the head.
The last point is worth holding on to. The measurement instrument shares the blind spot of the systems being measured, so published web composition figures are an upper bound on the skew being visible, not a neutral observation.
The gap against speaker populations
The interesting number is not the share itself but the ratio between share of text and share of speakers. English has on the order of 1.5 billion speakers including second-language speakers — roughly a fifth of humanity — and over 40% of the web text. Hindi and Bengali together have several hundred million native speakers and a share of web text in the fractions of a percent.
Three mechanisms drive that divergence, and none of them is going to resolve on its own:
- Web text is produced by people with reliable internet access and a keyboard for their script, which is not distributed like speaker population.
- Multilingual speakers often write in the dominant language. A Hindi speaker writing technical content frequently writes it in English, so their contribution appears in the English column. This is the same effect that makes code-switched text hard to count at all.
- Deduplication and quality filtering remove disproportionately from the tail. Filters are tuned on English, and heuristics like perplexity thresholds against an English reference model discard low-resource text at higher rates.
The consequence for Bengali specifically — a language whose speaker population would place it near the top of any list and whose model support does not — is worked through in why Bengali support lags its speaker population.
Why the skew compounds downstream
A share of training text does not translate linearly into quality, because the same imbalance is applied several times over.
The tokenizer is trained on the same skewed corpus, so its merges are optimised for the head languages, and a tail language is encoded in more tokens per unit of meaning. Petrov et al., in “Language Model Tokenizers Introduce Unfairness Between Languages” (NeurIPS 2023), document differences of up to roughly 15× in token count for the same content across languages. That paper is arXiv 2305.15425. More tokens per sentence means less content fits in the context window and each request costs more, so a tail language is charged a premium on top of receiving lower quality.
Instruction tuning and preference data are more skewed than pretraining data, not less, because they are produced deliberately by paid annotators rather than scraped. Safety and evaluation data are more skewed again. So the final model reflects the pretraining imbalance multiplied by the fine-tuning imbalance multiplied by the fact that nobody measured the difference.
What to check for your own language
- Look up the language in the current Common Crawl statistics rather than in a blog post. The share moves and the published table is authoritative for the corpus it describes.
- Check whether the language appears in FLORES-200, the evaluation set released with Meta’s No Language Left Behind work, which covers 200 languages. Presence there means at least a professionally translated evaluation set exists.
- Check the model card for an explicit supported-language list, and read it as a claim about what was evaluated rather than about what the model saw. The distinction is the whole of how many languages a model actually supports well.
- Measure token inflation yourself: encode a paragraph of your language and its English translation with the model’s tokenizer and take the ratio. It takes a minute and it is the single most predictive cheap number available.