Translating and Localising Responsibly
9 min read · updated August 4, 2026
The right question is not whether machine translation is good enough. It is what happens if this particular text is wrong, and the answer puts every job into one of three tiers with different rules.
Three tiers, decided by consequence
| Tier | Description |
|---|---|
| gist | You need to know roughly what a document says. Incoming email, a supplier's website, a review. Machine output used directly, no review, and you treat every detail as provisional. Do not forward it as though it were the document. |
| published | Text that goes out under your name to people who will read it as yours. Product descriptions, help pages, internal announcements. Machine draft plus a fluent-speaker review, plus a glossary. The reviewer is checking meaning and register, not grammar. |
| signed-off | Anything where a mistranslation causes harm or liability: contracts, safety instructions, medical information, regulatory filings, anything with a legal deadline in it. A qualified human translator produces or certifies the final text. The machine draft is a cost saving for them, not a substitute for them. |
Tier is a property of the text, not of the language pair. A safety notice into Dutch is tier three even though English-Dutch is among the best-served pairs in existence; a marketing tagline into any language is tier two at least, because the failure mode is not error but embarrassment.
What machine translation is good at now
Being accurate about this matters in both directions. Quality on well-resourced pairs — English with the major European languages, Chinese, Japanese — is genuinely high for ordinary prose, and pretending otherwise wastes money. Quality on low-resource languages is much weaker, and it degrades quietly rather than visibly.
- Strong: ordinary informative prose, in common pairs, where the subject is general and the register is neutral. Long documents, because a language model carries context across paragraphs better than sentence-by-sentence systems did.
- Weaker: specialised terminology without a glossary; anything where a term of art has a different meaning in the target jurisdiction; register that is not neutral, including humour, deliberate informality and politeness levels.
- Unreliable: low-resource languages and non-standard varieties, where the output remains fluent and stops being accurate, which is the worst combination available. There is a comparison of model translation against dedicated systems per pair in AI in translation and localisation.
The glossary is most of the quality
The largest single improvement available for tier-two work is not a better model. It is a list of your terms with their agreed translations, supplied with every job. Fifty entries covers most organisations.
GLOSSARY - do not deviate from these account holder -> rekeninghouder claim (insurance) -> schadeclaim [never "claim"] policy (document) -> polis policy (rule) -> beleid [these are different] submit a claim -> een schade melden [idiom, not literal] DO NOT TRANSLATE product names, UI button labels shown in screenshots, legal entity names, the names of statutory forms REGISTER Formal second person throughout (u, not je). British-style measured tone; do not add enthusiasm.
The two-sense entries are the valuable ones. A single English word with two meanings is translated inconsistently across a document by every system including human ones, and the glossary is where you settle it once.
The back-translation check
If you do not speak the target language, you are not helpless. Translate the output back and compare it with your original. This is a real technique used in professional practice, particularly in survey and clinical research, and it catches a specific and important class of error.
- Translate the source into the target language.
- In a separate, fresh conversation, translate the result back into the source language. The separation is essential: in one conversation the original is still visible and the back-translation drifts towards it, which defeats the test entirely.
- Compare against your original, sentence by sentence. You are looking for meaning changes, not wording changes. Different words with the same meaning are expected and fine.
- For every difference that changes meaning, ask what happened at that point. Often it is a term with two senses, which then goes into the glossary.
What this catches: reversed negations, dropped conditions, a “must” that became a “should”, an omitted clause. What it does not catch: text that is accurate and reads as stilted, wrong register, and anything culturally inappropriate. It is a meaning check, not a quality check, and it is not a substitute for a speaker.
Where a human must sign it off
Not a hedge — a list. In these categories the machine draft is an input to a qualified person and never the deliverable.
- Anything with legal effect. Contracts, terms, privacy notices, employment documents. Legal terms of art do not map across jurisdictions even between two countries sharing a language.
- Safety and medical information. Dosages, warnings, operating instructions, allergen information. The failure is irreversible and it falls on the reader.
- Anything sworn, certified or filed. Court documents, immigration papers, official registrations. These usually require a certified translator as a matter of law, which settles the question before quality enters into it.
- Public-facing text under your brand in a market you do not operate in. Not for accuracy but for register and cultural fit, which no check performed in the source language can assess.
- Anything a person will rely on to make a decision about their own money, health or rights. The general principle behind the whole list, and the one to apply to cases not on it.
Briefing a human reviewer
A reviewer given a translation and no context does the wrong job: they fix grammar, which is the part that is already fine. Tell them what they are for.
REVIEW BRIEF This is a machine translation, first draft. It is grammatically clean; please do not spend time there. Please check, in this order: 1. MEANING. Anything that says something the English does not. Mark these MEANING and they block publication. 2. TERMINOLOGY. Deviations from the attached glossary, and terms the glossary should contain but does not. 3. REGISTER. Does this sound like a company writing to a customer in your market, or like a translation? Rewrite freely. 4. CULTURAL. Anything that will read as odd, presumptuous or inappropriate locally, including examples, names and currency. Audience: [who reads this and why] Purpose: [what they should do after reading] Where it appears: [website, letter, product screen] You have authority to depart from the English wording to achieve the purpose. Tell me where you did and why.
The last line is the one that makes localisation work. A reviewer who believes they must stay close to the source produces a faithful translation of a text that was written for a different audience, which is the commonest way this goes wrong at tier two.
Length, formatting and the things that break
Most of the practical trouble in a translation project is not linguistic. It is that the translated text no longer fits, or that something structural did not survive the trip.
- Text expands and contracts. English into German or Finnish typically grows; English into Chinese or Korean typically shrinks. The consequence is buttons that no longer fit, headings that wrap to three lines, and printed material that runs over. If length matters, say so — “the result must fit in 20 characters; give three options with their lengths” — and then count them yourself, because a stated character count is not a measured one.
- Placeholders and markup get translated. A merge field, an HTML tag or a template variable will sometimes be helpfully rendered into the target language, which breaks the thing it was standing in for. List them explicitly as untranslatable, and check each one afterwards.
- Numbers, dates and units change meaning silently. Decimal commas, thousands separators, day-month order and paper sizes all differ by locale. A price of 1,500 rendered into a locale that uses the comma as a decimal separator becomes one and a half, and nothing about the sentence looks wrong.
- Long documents drift. The terminology chosen in chapter one is not necessarily the terminology used in chapter nine. This is the strongest practical argument for a glossary, and for translating in consistent chunks rather than paragraph by paragraph.
- Names, addresses and legal entities. Never translate them, and check them character by character. A transposed character in an address is not a translation error and it still stops the delivery.
The failures that do not look like failures
Fluency and accuracy are separate properties, and machine output is reliably fluent. A confident, well-formed, wrong sentence produces no signal for a reader who does not speak the language, which is why the checks on this page are procedural rather than perceptual.
Two specific traps. Names, numbers and units are altered surprisingly often — check every one by eye, including decimal separators, which differ between locales in a way that turns 1,500 into 1.5. And politeness systems in languages that mark them grammatically are chosen by default rather than by instruction: getting this wrong in Japanese or Korean is not a subtlety, it is the first thing the reader notices.