Tags: web-dev concept

Character Encoding

Date: 2026-08-17


The mapping between characters and bytes. Unicode assigns every character a number; UTF-8 decides how those numbers become bytes — and the gap between “character”, “code point” and “byte” is why string length lies about emoji.


A character encoding is a scheme for representing text as bytes. Unicode is the catalogue that gives every character a number (a code point); UTF-8 is the most common way of turning those numbers into bytes.

Those are two separate things, and conflating them is the root of most encoding confusion.

UTF-8’s variable width

CHARACTER   CODE POINT   BYTES IN UTF-8
A           U+0041       1   41
é           U+00E9       2   C3 A9
€           U+20AC       3   E2 82 AC
😀          U+1F600      4   F0 9F 98 80

ASCII characters cost one byte, which is why UTF-8 is backwards-compatible with ASCII and why it won. Everything else costs two to four.

Why string length lies

Three different counts, all reasonable, all different:

"héllo 👨‍👩‍👧"

bytes (UTF-8)        25
code points          11
UTF-16 units (JS)    14   ← what .length returns
grapheme clusters     7   ← what a human sees

JavaScript strings are UTF-16, so .length counts 16-bit units. Anything above U+FFFF — emoji, many CJK extensions, historic scripts — takes two units, called a surrogate pair.

"😀".length          // 2, not 1
[..."😀"].length     // 1  — iterates code points
"😀".slice(0, 1)     // half a surrogate pair
                     // renders as "�"

That slice is the bug. Truncating a product name or a meta description by character count splits emoji and combining marks in half.

"Rosé 😀 Serum" cut to 6 "characters"

.slice(0, 6)        "Rosé �"   ← half a
                                 surrogate pair
by grapheme cluster "Rosé 😀"  ← safe

The emoji occupies units 5 and 6, so cutting at 6 lands in the middle of it. Whether it breaks depends on where the cut falls, which means it works in testing and fails on a real product name.

Use Intl.Segmenter for anything user-visible.

The family emoji problem

👨‍👩‍👧 is not one code point. It’s three emoji joined by two zero-width joiners (U+200D) — five code points, eight UTF-16 units, eighteen bytes, rendering as one glyph. Skin-tone modifiers and flags work the same way.

There is no simple definition of “one character”, which is why the grapheme cluster concept exists and why naive truncation is unreliable.

Mojibake, and how it happens

Text written in one encoding and read as another:

WRITTEN UTF-8, READ AS LATIN-1
café  →  café

READ AS UTF-8 WHEN IT'S LATIN-1
café  →  caf�

The fix is never at the display end. It’s declaring the encoding at every boundary:

HTTP        Content-Type: text/html;
              charset=utf-8
HTML        <meta charset="utf-8">   ← first
                                        1KB
Database    utf8mb4, not utf8
Files       read and write with explicit
              encoding

MySQL’s utf8 is not UTF-8. It’s a three-byte subset that cannot store emoji or anything above U+FFFF. The real one is utf8mb4, and a column set to utf8 either errors or truncates on an emoji in a customer’s name or review — one of the most common encoding bugs in ecommerce.

Normalisation

The same visible character can have more than one encoding:

é   U+00E9                    (composed, NFC)
é   U+0065 U+0301             (e + accent, NFD)

look identical
compare as UNEQUAL
have different lengths

macOS filesystems have historically produced NFD; most other sources produce NFC. Normalise before comparing or storing — str.normalize('NFC') — or a search for “café” misses the record spelled with the other form.

Practical rules

  • UTF-8 everywhere, declared at every boundary
  • utf8mb4 in MySQL, and check existing columns rather than assuming
  • Normalise to NFC on input, before storing or comparing
  • Never truncate by .length for anything a person reads
  • Slugs and URLs need transliteration, not truncation — Redirects and Link Equity
  • Never assume one character is one byte when sizing a database column or a cookie