Len() disagrees with what users see as one character

Back to list2025-05-02
You

len() disagrees with what users see as one character.

ChatGPT

Because a user-perceived character can be several code points: an emoji with a skin tone modifier is one grapheme and up to four code points. Count grapheme clusters when the number is shown to a person, code points when it is a storage limit.

You

Two visually identical strings fail an equality check.

ChatGPT

Almost certainly composed versus decomposed forms: "é" as one code point or as "e" plus a combining accent. Normalise both sides with NFC before comparing, and store normalised so the problem does not come back through a different door.

You

What is the most common way people get this wrong?

ChatGPT

Doing it once and never verifying. The setup is the visible part, so it gets attention, and the check that would catch a silent failure never gets written.