Skip to content

Building an AI Product for a Language With No Digital Corpus

11 min read · updated August 11, 2026

The genuinely hard case is not a language with little data. It is a language with none: no Wikipedia edition worth the name, no parallel corpus, no benchmark, and a model that will happily produce confident-sounding output in it anyway. Here is a way to build something real from that position.

What you are actually starting with

Be honest about the starting inventory, because the plan depends on it. For a language in the bottom class of the taxonomy published by Joshi et al. at ACL 2020 — a class holding over two thousand languages — you typically have: a small number of speakers who are literate in the language, some printed material that is not digitised, a religious text or two, possibly a dictionary, and a related language that is far better resourced (Joshi et al., 2020).

You also have a model that has seen the language essentially never but has seen its neighbours. That combination is the whole opportunity and the whole danger: the model will produce output, the output will be fluent, and it will be the neighbour language wearing a costume unless you build something to stop it.

Build the evaluation set first

This inverts the usual order and it is the single most important decision here. With no corpus, everything you try is a guess, and a guess you cannot score is not an experiment. So the first artefact is not a model, a fine-tune or a prompt — it is two hundred sentences with reference outputs, produced by speakers, in your domain.

Two hundred is not arbitrary. The FLORES benchmarks published by Meta AI use a few thousand sentences per language across dev and test splits precisely because that is roughly where translation scores stop being dominated by sampling noise (No Language Left Behind, 2022). You will not get thousands. Two hundred domain sentences, split into a hundred you iterate against and a hundred you touch only at decision points, is enough to distinguish a real improvement from a mood.

Cover the constructions you know are hard, not a random sample: numbers, dates, names, negation, questions, politeness levels if the language marks them, and whatever your product says most often. Record which variety and which speaker produced each item.

Choosing a pivot language

You will route through a related high-resource language, and which one you pick matters more than which model you pick. Three criteria, in order:

  • Typological closeness beats geographic closeness. A pivot that shares morphology and word order preserves structure that survives the second leg. A pivot that shares only a border does not.
  • The pivot must mark what the target marks. If the target has an inclusive/exclusive “we”, evidentiality, noun classes or an honorific system, and the pivot does not, the pivot deletes that information and the second leg has to guess. This is the dominant error source — see why machine translation into low-resource languages still fails.
  • Speakers must be bilingual in it. Your reviewers need to read the intermediate to debug the pipeline. A theoretically better pivot nobody on your team speaks is worse.

The default pivot is English and it is usually the wrong one. English marks very little: no grammatical gender to speak of, no evidentiality, no clusivity, no honorific system, minimal agreement. It is a lossy intermediate for almost every language pair that is not English-adjacent.

The glossary goes in the loop

You cannot fine-tune your way out of no data, but you can constrain generation with the small amount of ground truth you do have. Build a glossary as a first-class asset: source term, target term, part of speech, the variety it belongs to, and a usage note. Two hundred entries covering your product vocabulary will change output quality more than any prompt rewrite.

Inject only the relevant slice, matched against the input, rather than the whole file — a thousand-entry glossary in every system prompt costs tokens on every request and dilutes attention:

import json

GLOSSARY = json.load(open("glossary.json"))  # [{"src": ..., "tgt": ..., "note": ...}]

def relevant(text, glossary, limit=25):
    """Only the entries whose source term appears in this input."""
    hits = [g for g in glossary if g["src"].lower() in text.lower()]
    return sorted(hits, key=lambda g: -len(g["src"]))[:limit]

def build_prompt(text, target_language, pivot_language):
    terms = relevant(text, GLOSSARY)
    lines = "\n".join(f'- {g["src"]} -> {g["tgt"]}  ({g["note"]})' for g in terms)
    return (
        f"Translate into {target_language}.\n"
        f"Reason through {pivot_language} if it helps, but output only "
        f"{target_language}.\n"
        f"Use these terms exactly as given; they are authoritative:\n{lines}\n\n"
        f"If a term is not in the list and you are not confident of the "
        f"{target_language} word, keep the source word unchanged and mark it "
        f"with [?]. Do not invent vocabulary.\n\n"
        f"Source:\n{text}"
    )

The last instruction is the one that earns its place. A model with no coverage will invent a word rather than leave a gap, and an invented word is indistinguishable from a real one to everyone on your team. Ask for an explicit marker instead, and the review queue builds itself: every [?] is a glossary entry waiting to be written.

The build, in order

  1. Fix the variety. Name the dialect, region and orthography you are targeting, in writing. Languages in this situation usually have no standardised spelling, and a corpus that silently mixes two conventions is worth much less than half of each.
  2. Recruit two speakers who can review each other. One speaker gives you an opinion; two give you a disagreement rate, which is the only signal you have about whether your reference set is reliable.
  3. Build the 200-sentence evaluation set, split 100 dev / 100 held-out, with the constructions listed above deliberately represented.
  4. Establish the baseline. Run direct prompting in the target language, with no pivot and no glossary. Score it. This is the number everything else has to beat, and sometimes it is already adequate for your task — find out before building a pipeline.
  5. Add the pivot, then the glossary, then the [?] abstention instruction, scoring after each. Add one thing at a time; with a hundred dev sentences you cannot attribute a change you made in a batch.
  6. Harvest the abstentions. Every marked gap goes to a speaker, comes back as a glossary entry, and the entry is versioned. This loop is the product; the prompt is scaffolding around it.
  7. Freeze and re-score on the held-out hundred before any release, and only then. A dev set you have looked at fifty times has stopped measuring anything.

Shipping without lying to your users

Two rules, and both are about what you tell the reader. First, label machine output as machine output in the target language, not only in your interface language. Second, give a speaker a path to correct it that takes one click, and route corrections into the glossary. A community that can fix your output will; a community that can only be annoyed by it will be.

And keep the ownership question in view from the start. If the language belongs to a community that has a position on how its material may be used, that position governs, and it is much easier to honour before you have built on top of the data than afterwards — the ground is covered in LLM support for Indigenous American languages.