Skip to content

Splitting Devanagari Text for Embedding Without Breaking Conjuncts

9 min read · updated August 11, 2026

Devanagari writes one visual unit out of several code points: a consonant, a vowel sign that may sit above or before it, and a virama that fuses it to the next consonant. Any splitter measuring in code points will eventually cut inside one of those units, and the result is not a slightly worse chunk — it is a string containing characters that cannot legally start a word.

One letter, several code points

Three mechanisms in Devanagari turn several code points into one written unit.

  • Matras. A consonant carries an inherent a vowel; any other vowel is written as a dependent sign attached to it. कि is क (U+0915) plus ि (U+093F) — and note the vowel sign is stored after the consonant while it renders before it. Storage order and visual order differ, which is exactly why you must never reason about these strings by what they look like.
  • The virama. U+094D suppresses the inherent vowel and joins two consonants into a conjunct. क्ष is three code points — क, ्, ष — displayed as one fused glyph.
  • Nasalisation and other marks. Anusvara ं (U+0902), candrabindu ँ (U+0901) and nukta ़ (U+093C) are non-spacing marks attached to the preceding consonant.

So the word लक्ष्मी is eight code points and four written units. A length check in len() tells you almost nothing about how long it looks, how it tokenises, or where you may cut.

A worked split that goes wrong

Take परीक्षा (“examination”) as a labelled sample and lay out its code points:

प  U+092A  DEVANAGARI LETTER PA
र  U+0930  DEVANAGARI LETTER RA
ी  U+0940  DEVANAGARI VOWEL SIGN II      (spacing mark, renders after)
क  U+0915  DEVANAGARI LETTER KA
्  U+094D  DEVANAGARI SIGN VIRAMA        (joins KA to the next consonant)
ष  U+0937  DEVANAGARI LETTER SSA
ा  U+093E  DEVANAGARI VOWEL SIGN AA

7 code points, 4 written units: प · री · क्षा

Cut at index 5, which is where a fixed-size splitter lands when its budget expires there:

chunk A: "परीक्"    ends on a VIRAMA with nothing to join to
chunk B: "षा"       begins with a bare SSA carrying a vowel sign

Rendering: chunk A shows a "half ka" — a consonant with a visible
dead-stroke, which is not a word-final form in Hindi.
Tokenising: the trailing virama is a sequence the tokenizer has
essentially never seen at end of string; it splits into byte
fallback tokens that carry no learned meaning.

Cut one position earlier, at index 4, and it is worse: chunk B begins with U+094D, a combining mark with no base. Some renderers show a dotted circle placeholder, some attach it to whatever precedes it in the surrounding layout, and every one of those is a chunk that will embed and display incorrectly.

The harm is not mainly aesthetic. The purpose of a chunk is to be a retrievable unit of meaning, and परीक् is not a word — it is the machine equivalent of cutting “examination” into “examinati” except that the fragment also renders wrong.

Splitting on grapheme clusters

Unicode defines the unit you actually want. An extended grapheme cluster is the standard’s model of what a user perceives as one character, and the rules for finding cluster boundaries are specified in UAX #29, the Unicode text-segmentation annex. A base consonant and every mark attached to it form one cluster, so iterating clusters instead of code points removes the entire class of orphaned-mark bugs.

Python’s standard library does not implement this; the regex module does, via \X.

import regex, unicodedata

def clusters(s):
    return regex.findall(r"\X", unicodedata.normalize("NFC", s))

def split_devanagari(text, budget, ntok):
    """Pack whole grapheme clusters to a token budget, never cutting inside one."""
    out, cur, n = [], [], 0
    for word in text.split(" "):
        k = ntok(word + " ")
        if k > budget:                       # single word over budget: cluster-pack it
            buf = ""
            for g in clusters(word):
                if ntok(buf + g) > budget:
                    out.append(buf); buf = ""
                buf += g
            if buf:
                out.append(buf)
            continue
        if cur and n + k > budget:
            out.append(" ".join(cur)); cur, n = [], 0
        cur.append(word); n += k
    if cur:
        out.append(" ".join(cur))
    return out

Devanagari does use spaces between words, so word packing is the primary path and the cluster loop is a fallback for a single over-budget token — a long compound, a URL, an unbroken run from bad extraction. That is the opposite arrangement from chunking Thai, where there are no word spaces at all, and it is why Hindi chunking is usually fine by accident and fails only on the awkward inputs.

Conjuncts need a newer rule than you think

There is a subtlety here that catches people who did everything right. Under the extended grapheme cluster rules as they stood for many years, a virama-joined conjunct was two clusters, not one: क् formed one cluster and ष another, because the rules kept a mark with its base but did not treat the virama as joining across the boundary. Unicode 15.1 added a rule specifically for Indic conjunct breaks so that a linking-consonant sequence stays together.

The practical consequence is that whether \X keeps क्ष intact depends on which Unicode version your regex or ICU library was built against. Check it rather than assuming, because a library on the older rules will still never orphan a mark — the severe bug is fixed either way — but it will happily split क् from ष, which is a legal cut in the old model and a visibly wrong one to a reader.

Verify with your own installed library rather than from documentation: count the clusters in क्ष. Two means your library predates the conjunct rule; one means it implements it. Then decide whether the difference matters for your corpus — it matters most for Sanskrit and for technical Hindi, which are conjunct-dense, and least for conversational text.

What to do in a real pipeline

  1. Normalise to NFC at ingest, and apply the identical normalisation to queries. Devanagari has precomposed forms for some nukta combinations, and a document in one form will not match a query in the other. The general mechanism is covered in the difference between NFC and NFKC.
  2. Split on words first; treat grapheme clusters as the fallback unit for anything a word split cannot handle.
  3. Budget in tokens, not code points. Devanagari tokenises less efficiently than Latin in most current vocabularies, so a code-point budget overshoots in the same direction it does for CJK.
  4. Assert that no chunk starts with a code point whose Unicode general category is Mn or Mc, and that no chunk ends with U+094D. Those two assertions catch every instance of this bug and run in microseconds.
  5. Round-trip: joining the chunks must reproduce the input, so a normalisation applied halfway through is caught immediately.

The same two assertions generalise across scripts. Any Indic script, Thai, Arabic with diacritics and Hebrew with niqqud all fail the leading-combining-mark check when a splitter has cut inside a cluster, which makes it the single most useful invariant to add to a multilingual ingest pipeline.