Emoji Metadata Standards: CLDR Annotations, emoji-data, and the Data Files
Every emoji tool is built on a small set of standardized data files. Here is what each file provides and why skipping one leaves a gap. Every emoji picker, search engine, and reference tool is built on metadata, and the metadata comes from a small number of standardized data files maintained by the Unicode Consortium and the CLDR project. Understanding these files is the difference between building a reliable emoji tool and building one that breaks when a new Unicode version ships. I spent more time than I care to admit reading these files before I understood the layout, so here is a map. I started building an emoji reference tool thinking I could just use a npm package with all the emoji data. Three months later, I had read every Unicode data file related to emojis, written my own parser for four of them, and realized that the npm packages were all derived from the same source files. Understanding the source files is understanding the foundation that every emoji tool is built on. The emoji-data File The core file is emoji-data.txt, published as part of each Unicode release. It lists every emoji code point and sequence, along with a set of properties for each. Read the full article on Emoji Reference, and copy any emoji mentioned from the catalog of 1303+ entries on the home page.