Writing Regex That Matches Emojis Across All Unicode Versions
A single Unicode range never matches all emojis. Here is how to build a regex that handles modifiers, ZWJ sequences, and flags. Matching emojis with a regular expression is harder than it looks, and most regex patterns I see in production are wrong. They either match too much, catching plain symbols that are not emojis, or too little, missing skin tone modifiers and ZWJ sequences. I have rewritten my emoji regex three times as the standard evolved, and the version I use now is still a compromise. Here is how to think about it and what to use. The first version I wrote was a simple range check: match everything from U+1F300 to U+1FAFF. It worked for about 80 percent of emojis and failed on the other 20 percent. The failures were all the interesting cases: the black heart, the flag emojis, the skin tone modifiers, the ZWJ sequences. I kept adding ranges and special cases until the regex was unreadable, and then I threw it away and started over. Why a Single Range Does Not Work The naive approach is to match everything from U+1F300 to U+1FAFF. This catches most pictographic emojis but misses the legacy symbols in the BMP, like the black heart at U+2764, and it misses modifier pairs and ZWJ sequences entirely. Read the full article on Emoji Reference, and copy any emoji mentioned from the catalog of 1303+ entries on the home page.