Fix name encoding issues caused by surrogate pairs

  • Published: 2026.10.04
  • Updated: 2026.10.04
  • Technology

A person’s name is entered correctly, but a character becomes ? after saving or appears as a square on another device. Supplementary Unicode characters expose assumptions about storage, string length, and fonts that ordinary test data may not reveal.

The Japanese character 𠮷 is one useful example. Similar issues affect supplementary-plane emoji and symbols, so the underlying checks apply to international applications too.

What a surrogate pair represents

UTF-16 uses two 16-bit code units to encode a Unicode scalar value outside the Basic Multilingual Plane (BMP). Those two units form a surrogate pair. It is an encoding mechanism, not two separate visible characters.

Many familiar characters fit in one UTF-16 code unit. Supplementary characters, such as 𠮷 (U+20BB7), require two. Not every unusual-looking Japanese character does: 髙 (U+9AD9) and 﨑 (U+FA11) are in the BMP.

Example Code point UTF-16 code units
𠮷 U+20BB7 2
髙 U+9AD9 1
🥺 U+1F97A 2
𝄞 U+1D11E 2

Distinguish lost data from a missing glyph

A square can mean that the selected font lacks a glyph even though the underlying value is intact. A replaced or truncated value can instead indicate an encoding or storage problem. Inspect the actual value before deciding where to fix the system.

  • Input: check the field’s value before submission.
  • Transport: check that requests and responses use the intended encoding.
  • Storage: compare the stored value with the submitted value.
  • Display: try a font with the necessary glyph coverage while keeping the original data unchanged.

Check UTF-16 length assumptions

JavaScript’s length counts UTF-16 code units. The following examples therefore produce different lengths despite each representing one Unicode code point:

console.log("𠮷".length); // 2
console.log("髙".length); // 1
console.log(Array.from("𠮷").length); // 1

Taking a substring at an arbitrary code unit boundary can split a surrogate pair. Counting code points avoids that particular split, but visible character sequences such as a family emoji can still contain several code points. For a visible-character limit, define the unit as a grapheme cluster and use an appropriate segmentation implementation.

See Counting characters in JavaScript for the differences between code units, code points, and grapheme clusters, with shortening examples.

Check the database and connection character sets

For MySQL, utf8mb3 cannot store supplementary characters; utf8mb4 supports them. Check both the relevant columns and the application’s connection settings. A compatible table does not compensate for an incompatible value elsewhere in the path.

Changing a column’s character set is a database migration with possible index and compatibility implications. Plan and test that change against the actual schema rather than applying a generic alteration command from a character-display example.

Test registration, search, and display together

Use a small test set containing a supplementary name character such as 𠮷, a supplementary emoji such as 🥺, and a BMP character such as 髙. Check the round trip through saving, retrieval, search, and display, including the application’s length validation.

When only display is affected, retain the entered value and improve font fallback. If a legacy system cannot accept the correct value, explain the limitation and offer a supplementary field or agreed alternative while planning the underlying fix. Do not silently replace a person’s name.

References