Why the Same Prompt Gets Different Refusal Behavior in Japanese and English
10 min read · updated August 11, 2026
Japanese and English versions of one prompt often get different treatment, and most reports of it are unreliable for a reason that has nothing to do with the model: a Japanese refusal usually does not contain any of the strings people detect refusals with.
What the question usually means
Three different observations get reported in the same words, and they have different causes.
- The Japanese version is refused and the English is not. Often a false positive from a filter reacting to literal phrasing — the mechanism in cultural context and false-positive refusals.
- The Japanese version is answered and the English is refused. The safety-gap direction, and the one that shows up in red-teaming.
- Neither is refused, but the Japanese answer hedges. The most common case in practice and the one that is almost never counted, because it is neither a refusal nor a clean compliance.
The third case is why aggregate refusal rates for Japanese are usually understated. Japanese pragmatics prefer indirect decline over flat refusal, and a model trained on Japanese text reproduces that: it answers partially, adds a caution, and redirects, without ever producing a sentence any refusal classifier recognises.
Where Japanese differs structurally
These are the properties that make Japanese a distinct case rather than another instance of the general low-resource story — Japanese is not low-resource, and the boundary still moves.
Politeness is grammatical
Japanese encodes register in the verb. The same request in plain form and in polite form is two different social acts, and honorific registers change it again. English alignment data has no examples of this distinction, so whatever the model does with it was not chosen. A prompt written casually can be treated as a different kind of request from its polite equivalent, with no corresponding difference in the English translation. See politeness levels across languages.
Subjects are routinely omitted
Japanese drops arguments that are recoverable from context, including the subject. A harm classifier’s core job is to determine who is doing what to whom, and in Japanese that information is frequently not in the sentence at all. The same request can therefore be unclassifiable in Japanese and unambiguous in English, and any translated evaluation set silently supplies the missing subject in translation, making the Japanese items easier than real Japanese input.
One word, four scripts
The same word can be written in kanji, hiragana, katakana or romaji, and katakana is conventionally used for emphasis. A term list or a classifier keyed on one form does not fire on the others, so lexical safety signals are far leakier than in a single-script language. See Japanese script mixing.
No word spaces
Segmentation is model-specific, so what the tokeniser and the classifier treat as a unit varies between systems in a way it does not for English. This is the same boundary problem discussed in Japanese sentence segmentation.
Genre conventions
A large share of Japanese creative-writing traffic uses genre conventions with no equivalent in the English alignment data. The model has no learned notion of “this is a fiction request in a recognised genre” in Japanese, so fiction and instruction are harder for it to separate there than in English.
What public evaluation covers
Less than you would expect for a language of Japanese’s size. MultiJail, one of the most cited multilingual jailbreak sets, covers nine non-English languages — Chinese, Italian, Vietnamese, Arabic, Korean, Thai, Bengali, Swahili and Javanese — and Japanese is not among them, as the paper’s language grouping sets out. Japanese does appear in some of the broader toxicity sets such as RTP-LX, and Japanese-language safety evaluation is largely maintained by Japanese research efforts — the LLM-jp project and the JGLUE line of work — rather than by the international benchmark suites. Check the current release of those directly rather than trusting a figure quoted second-hand, including from this page.
The consequence is that for Japanese specifically you are unlikely to find a published number that answers your question, and the general coverage picture is in why public safety benchmarks are mostly English-only.
Your detector is the confound
This is the part that invalidates most informal comparisons. Refusal detection is usually implemented as a substring or regular-expression match over a list of English refusal openings — “I’m sorry”, “I cannot”, “As an AI”. That detector has recall of essentially zero on Japanese output. A Japanese refusal opens with 申し訳ありませんが or 申し訳ございませんが, declines with お答えできません or ご要望にはお応えできません, and frequently does not decline explicitly at all — it answers around the request.
Run an English-keyed detector over both arms of a bilingual comparison and you will measure a large gap in the direction of “Japanese refuses less”, in every case, on every model, regardless of what the model did. The number is a property of the detector.
// Wrong: recall ~0 on Japanese.
const REFUSAL = /^(i'm sorry|i cannot|i can't|as an ai)/i;
// Better: per-language markers, plus a judged fallback for soft refusals.
const MARKERS: Record<string, RegExp> = {
en: /\b(i'm sorry|i cannot|i can't|i am unable|as an ai)\b/i,
ja: /(申し訳|お答えできません|お応えできません|ご遠慮|できかねます)/,
};
// Markers catch hard refusals only. Soft refusals — an on-topic answer that
// withholds the requested content — need a rubric-scored judge, and the judge
// must be validated separately per language against human labels.Validating the judge is not optional here. Label a few hundred Japanese responses by hand against a written rubric with three classes — complied, refused, hedged — then measure the judge against those labels and report its own agreement figure alongside your result. If the judge agrees with humans at 0.7 on Japanese and 0.9 on English, differences smaller than that gap are not interpretable.
A method that produces a real number
- Build paired items from your own traffic, not from translated red-team sets. Each item exists in both languages, adapted by a native speaker rather than machine-translated, with a record of any item whose meaning changed.
- Include benign-but-sensitive items as well as harmful ones. Measuring only harmful items gives you the under-refusal rate and tells you nothing about the over-refusal your users are experiencing.
- Vary register deliberately. Include each Japanese item in plain and polite form as separate conditions; if they differ, that difference is a finding in itself and has no English counterpart.
- Run at temperature 0 with at least five trials per cell, and record the provider’s finish or stop reason alongside the text so that filter blocks are separated from model refusals rather than pooled.
- Classify with per-language markers plus a validated judge, and report the judge’s agreement with human labels per language in the same table as the result.
- Analyse as paired data. McNemar’s test on the discordant pairs is the right comparison and is far more sensitive at small samples than comparing two independent rates.
- Report four numbers per model: over-refusal and compliance rates in each language, with the model version and the date. Never a single “Japanese is N% more restrictive”.