len() disagrees with what users see as one character.
Because a user-perceived character can be several code points: an emoji with a skin tone modifier is one grapheme and up to four code points. Count grapheme clusters when the number is shown to a person, code points when it is a storage limit.
Two visually identical strings fail an equality check.
Almost certainly composed versus decomposed forms: "é" as one code point or as "e" plus a combining accent. Normalise both sides with NFC before comparing, and store normalised so the problem does not come back through a different door.
Any downside worth knowing before I commit?
It commits you to a format that is tedious to migrate away from later. The first weeks also look worse than doing nothing, which is when most people abandon it.