Skip to content

Why AI-Generated Forms Break on Two-Character Surnames

9 min read · updated August 11, 2026

“Last name must be at least 2 characters.” “Please enter a valid last name.” “Name must contain only letters.” These three strings appear in generated form code almost every time, and each of them rejects real people who then have no way to complete the form.

The error the user sees

The report usually arrives as “the signup form will not accept my name”, and it is not reproducible for whoever receives it, because their own name passes. It is a hard failure rather than a degraded experience: the user cannot proceed, cannot work around it — there is no alternative spelling of their own surname — and the message tells them their name is invalid, which is an unusually unpleasant thing for software to say.

The rule that produces it

Ask a model for a registration form with validation and you will get some variant of this, in whatever stack you asked for:

// zod
lastName: z.string().min(2, "Last name must be at least 2 characters"),

// HTML
<input name="lastName" minlength="2" pattern="[A-Za-z]+" required>

// a hand-rolled check
if (!/^[a-zA-Z]{2,50}$/.test(lastName)) return "Please enter a valid last name";

Three separate assumptions are baked in there, and all three are wrong. That a surname has at least two characters. That surnames are written in unaccented Latin letters. And, implicitly, that everyone has one. The model is not inventing these; it is reproducing the overwhelming majority of tutorial and Stack Overflow code it saw, which was written by people whose own names fit.

There is a fourth assumption hiding in .length. In JavaScript, Java and C#, string length counts UTF-16 code units, not characters, so the rule is not even measuring the thing it claims to measure.

The names it rejects

Two-character and one-character surnames are not edge cases; they are held by an enormous number of people.

  • Chinese surnames are one character. 李, 王, 张, 刘, 陈 — the most common surnames in the world’s largest population, each a single character. Romanised they become Li, Wu, He, Xu, Yu: two letters. The single-character form fails min(2) outright, and fails [A-Za-z] as well.
  • Korean surnames are one Hangul syllable. 이, 김, 박, 오, 안 — romanised as Lee, Kim, Park, Oh, Ahn. Several standard romanisations are two letters (Oh, An, Ko, Ha, Yu), and at least one widely used romanisation of 오 is the single letter O.
  • Vietnamese surnames. Lê, Vũ, Hồ, Đỗ, Lý. Two characters, all of them carrying diacritics that a Latin-only character class rejects, and Đ is a letter a naive class does not contain at all.
  • Cantonese romanisations. Ng — the standard romanisation of 吳 — is two letters with no vowel, which additionally trips “must contain a vowel” heuristics that some validators add to catch typing noise.
  • European short names. Ek, Ax, Ås in Scandinavia; Ó in Irish orthography; Bo, Le, Da as full surnames elsewhere. Apostrophes and hyphens in O’Brien, D’Angelo and Anne-Marie fail the letters-only class for a separate reason.

Notice that the CJK cases fail twice over. A single Han or Hangul character is one character and so fails the length rule, and it is not in [A-Za-z] so it fails the character rule as well. Loosening only one of the two leaves the user exactly where they were.

The same name, two verdicts

The length rule has a failure mode that is worse than being too strict: it is inconsistent. Unicode allows the same accented letter to be encoded as a single precomposed code point or as a base letter plus a combining mark, and both are correct.

"Ó".normalize("NFC").length   // 1  -> fails min(2)
"Ó".normalize("NFD").length   // 2  -> passes min(2)

So one user is rejected and another accepted for the same surname, depending on which keyboard, operating system or clipboard produced the text. The decomposed form is what macOS filesystems have historically produced, so a name pasted from a Mac may behave differently from the same name typed on Windows. The Korean case behaves the same way: composed Hangul syllables are one code point each, and the decomposed jamo sequence for the same syllable is two or three.

Any rule that counts characters therefore has to normalise first and count grapheme clusters rather than code units — which is a real amount of machinery to build in service of a rule that should not exist.

And some people have one name

A required surname field is a stronger assumption than a short one. It is common in Indonesia — including at the level of former heads of state — for a person to have a single name and no family name at all. The same is true for many Burmese and Javanese names, for some South Indian naming conventions where what looks like a surname is a patronymic or a village name, and for Icelandic names, where the second element is a patronymic that changes each generation and is not a family name in the sense a database means.

The design consequence is that splitting a name into given and family parts is itself a locale assumption. The W3C Internationalization article on personal names is the standard reference here, and the GOV.UK Design System names pattern reaches the practical conclusion: prefer a single full-name field, and do not restrict the characters people may enter.

What a name field should validate

  • That it is not empty after trimming whitespace. That is the whole of the correctness check.
  • A generous maximum — 100 characters or more — to bound storage, not to judge the name. Say the limit in the message.
  • Normalise to NFC before storing so that comparison and search behave, and so that two spellings of the same name are the same bytes. This is the same normalisation that makes sorting accented names behave predictably.
  • No character class, no minimum length, no requirement for a space, no rejection of numbers or punctuation. If you need to catch abuse, do it with rate limiting and review, not by asserting what a name looks like.

When prompting for form code, the instruction that reliably works is concrete rather than general: “name fields must accept a single character, any Unicode letter, apostrophes and hyphens; do not set a minimum length or a character pattern.” Asking for “internationalised validation” typically returns the same Latin-only regular expression with \p{L} substituted for A-Za-z, which fixes the character class and leaves the length rule intact — a genuine improvement that still rejects every Chinese surname.