Why Hindi Text Costs More Tokens Than English
9 min read · updated August 11, 2026
A Hindi word that looks like four letters is often eight Unicode code points and twenty-four bytes. That gap between what you see and what the tokenizer counts is the whole of the Devanagari token problem.
Count code points, not letters
Devanagari is an abugida. A consonant carries an inherent vowel, and any other vowel is written as a matra — a sign attached above, below, before or after the consonant. Crucially, the matra is a separate code point stored after the consonant, regardless of which side of it the mark is drawn on.
So सु is two code points, स (U+0938) and the matra ु (U+0941). मे is म plus े. The word मेमोरी, which looks like four letters, is six code points. Every Devanagari code point lies in U+0900–U+097F, which UTF-8 encodes in three bytes, so those six code points are eighteen bytes. The English word memory is six ASCII bytes and, with a leading space, very plausibly one token.
This is why Python’s len() on a Hindi string does not agree with what a reader would count, and why a character-based text splitter tuned on English produces the wrong sizes on Hindi in a way that is invisible until you look at bytes.
What a virama costs
The second multiplier is the conjunct. When two consonants meet with no vowel between them, Devanagari fuses them into a single ligature, and the fusion is encoded with the virama U+094D between them. क् plus ष renders as the single glyph क्ष, and it is three code points: क, virama, ष. Nine bytes for one visible letter.
शिक्षा, “education”, is a four-glyph word to a reader and six code points to a computer — श, ि, क, ्, ष, ा — which is eighteen bytes. Conjuncts are not rare or ornamental; they are ordinary Hindi orthography and appear in a large fraction of content words, especially Sanskritic vocabulary, which is exactly the register used in formal, legal and technical writing.
For the merge table, the virama is the damaging part. It is a code point whose only job is to sit between two others, so it fragments any merge that might have spanned the consonant pair. A tokenizer that has learned क as one token and ष as one token still has to emit the virama as something, and unless the whole three-code-point sequence earned its own merge, a single conjunct glyph costs three tokens.
The arithmetic on a labelled sentence
यह सुविधा बहुत मेमोरी का उपयोग करती है। "This feature uses a lot of memory."
Counted as code points rather than as perceived letters: यह is 2, सुविधा is 6, बहुत is 4, मेमोरी is 6, का is 2, उपयोग is 5, करती is 4, है is 2, and the danda । is 1. That is 32 Devanagari code points at three bytes each, 96 bytes, plus seven single-byte spaces: exactly 103 bytes. The English is 35 ASCII bytes, near 9 tokens.
Now the derived range, assumption stated. Assume the tokenizer holds merges for frequent short Hindi words and for common consonant-plus-matra pairs, so those cost one token, while less common sequences and every virama fall back to one token per code point. That puts the sentence roughly between 30 and 60 tokens against the English 9 — a derived multiplier of about 3× to 6×, with an unconditional ceiling of 103. This is a much harsher range than Russian or Arabic, and the reason is compounding: three bytes per code point, several code points per visible letter, and thinner merge coverage than any of the languages above.
What this shares with the other Indic scripts
Rather than repeat this derivation for each Brahmic script, here is what transfers and what does not. The structure transfers completely: Bengali, Gujarati, Odia, Telugu, Kannada, Malayalam, Punjabi’s Gurmukhi, Nepali’s Devanagari and Sinhala are all abugidas with separate matra code points, all in three-byte UTF-8 ranges, and all use a virama or its equivalent to form conjuncts. The code-point-per-glyph arithmetic on this page applies to every one of them.
What does not transfer is corpus share, and that is where the numbers diverge. Hindi has by far the largest web presence of the group, so its merge coverage is the best of them; Nepali shares the Devanagari script but not the corpus, so it inherits the structural cost with less of the compensating coverage. Malayalam and Telugu compound the problem further with longer average conjunct chains. The rule of thumb worth carrying is that within Indic scripts, the structural penalty is constant and the vocabulary penalty is what ranks them. See Bengali and Nepali for how the same structure lands with different corpus shares.
Measuring it, and the splitting trap
import sys
import unicodedata
import tiktoken
enc = tiktoken.get_encoding(sys.argv[1] if len(sys.argv) > 1 else "o200k_base")
hi = "यह सुविधा बहुत मेमोरी का उपयोग करती है।"
print("codepoints", len(hi))
print("bytes", len(hi.encode("utf-8")))
print("tokens", len(enc.encode(hi)))
# how many code points are marks rather than letters
marks = sum(1 for ch in hi if unicodedata.category(ch).startswith("M"))
print("combining marks", marks)The combining marks figure is the diagnostic one. Unicode classifies matras and the virama in the mark categories, so it tells you what fraction of your code points are doing no work a Latin script would have paid for. In ordinary Hindi prose it is commonly a quarter to a third of all code points.
Finally, the splitting trap, which is the bug this page exists to prevent. If you cut a Hindi string at an arbitrary code point offset — which is what a naive fixed-size chunker does — you can cut between a consonant and its matra, or between a consonant and a virama. The result is a chunk that ends in a dangling combining mark, renders as a dotted circle, tokenises as garbage and embeds as noise. Split on grapheme cluster boundaries as defined in Unicode Standard Annex #29, or on word boundaries, never on raw code point counts. The retrieval consequences are set out in splitting Devanagari text for embedding.
- A 4,000-token budget holds very little Hindi. On the derived range above, something in the order of 700–1,300 Devanagari code points, which is a few short paragraphs.
- Transliterated Hindi is a different language to the model. Roman-script Hindi tokenises cheaply and is handled as its own phenomenon; it is not a free compression of the same input.
- Nukta-bearing letters have two spellings. क़ can be one precomposed code point or क plus the nukta U+093C. Normalise, or the same word arrives with two different token counts.