Unicode 18.0 Released with 13,007 New Characters
Unicode 18.0 is out with 13,007 new characters, three new scripts, new currency symbols, nine emoji, and updated text-processing rules.
The emoji get the coverage. The other 12,998 characters matter more.
Scripts
Three new ones. Each addition is the result of years of work by linguists and by communities who want to write their own language on a computer, and the practical effect is the difference between a script existing digitally and not.
The bulk of a 13,000-character release is almost always CJK ideographs, historic scripts, and specialist notation. It is unglamorous, slow, and it is the work that makes a text encoding genuinely universal rather than adequate for Western European languages.
Text-processing rules
The part that affects software.
Unicode is not only a list of characters. It defines how text behaves: which sequences form a single user-perceived character, where lines may break, how strings compare and sort, and how bidirectional text is laid out.
Updates here mean libraries need updating. ICU, HarfBuzz, and every language’s string implementation have to absorb the new data before the characters render and compare correctly.
Grapheme cluster boundaries are the rule that determines what counts as one character for cursor movement and backspace. Get it wrong and pressing backspace on an emoji with a skin-tone modifier deletes half of it, which is the classic symptom of a stale Unicode implementation.
What this means on a Linux system
The characters exist in the standard immediately. Whether you can see them depends on three separate things being updated.
Fonts. A character with no glyph renders as a box. Noto exists specifically to chase this, and even Noto lags a new release. Our fontconfig guide covers installing coverage and diagnosing missing glyphs.
fc-list :charset=1F600 # which fonts cover a given codepoint
Libraries. glibc, ICU, and HarfBuzz need the new tables for correct sorting, casing, and shaping. On a stable distribution these arrive with the next release rather than as an update.
Terminals and applications. Anything computing string widths itself, rather than asking a library, will miscount new characters and misalign columns. This is the usual cause of a TUI drawing a broken table when someone’s filename contains something unexpected.
Our locale guide covers the related question of why sorting differs between machines, which is the same data driving different behaviour.
Filenames are still bytes
Worth restating whenever Unicode comes up. The kernel treats filenames as byte strings and does not know or care about encoding. A filename containing a character your font lacks is stored perfectly and displays as a box, and a filename created under a different encoding displays as nonsense.
ls | cat -v # show the actual bytes
That separation is why a Unicode update changes what you can see and never what the filesystem can store.
Full details are at unicode.org.