Emoji Normalization: Why NFC and Equivalence Classes Matter
Two visually identical emoji strings can differ at the byte level. Here is how normalization prevents duplicate entries and broken equality checks. Two strings that look identical can be different at the byte level, and emojis are the most common place this bites. The same visual emoji can be represented as a single code point or as a sequence with a variation selector, and string equality checks that look correct fail on the difference. This is the normalization problem, and understanding it is essential for any code that compares, deduplicates, or hashes emoji strings. I encountered this bug in a production system where users were reporting duplicate entries in their emoji recents list. The same emoji appeared twice, visually identical, and the user could not understand why. The cause was that the user had pasted the emoji from two different sources: one included the variation selector, the other did not. The database stored both as distinct strings, and the recents list showed both. The fix was to normalize emoji strings before comparison and storage. What Normalization Means for Text Unicode defines normalization forms that map multiple byte level representations of the same character to a single canonical form. NFC, Normalization Form C, composes characters into their precomposed forms where possible. NFD decomposes them. For most text, NFC is the default and the right choice, and most modern text is already in NFC. Read the full article on Emoji Reference, and copy any emoji mentioned from the catalog of 1303+ entries on the home page.