Open a file in the wrong program and café becomes café. That garbled output — mojibake — is what happens when bytes written in one text encoding get read as another. Understanding the two encodings that matter explains nearly every text corruption you'll ever meet.

ASCII: the 128-character foundation

ASCII (1963) assigned numbers 0–127 to English letters, digits, punctuation, and control codes. One byte per character, with the top bit unused. It was beautifully simple and hopelessly Anglocentric — no accents, no non-Latin scripts, no emoji. Every competing extension that filled the upper 128 slots differently created files that were unreadable elsewhere: the original encoding chaos.

Related reading: How to Remain Valuable When Intelligence Becomes Cheap — a 224-page practical book on staying valuable as intelligence gets cheap. $3.84. Read it on Gumroad →

UTF-8: the encoding that won

UTF-8 solved the mess with a variable-width design: ASCII characters stay single bytes (identical to ASCII, so old files just work), while other characters use two to four bytes with unmistakable prefix patterns. It encodes every Unicode character — over 149,000 of them — with no ambiguity about byte order and no wasted space on English text.

Today UTF-8 carries over 98% of the web. When someone says "just use UTF-8 everywhere," they're right: declare it in your HTML (<meta charset="utf-8">), save files in it, configure databases for it (utf8mb4, which is real UTF-8 — MySQL's utf8 is a broken 3-byte subset that chokes on emoji).

Where mojibake comes from

Garbled text means a UTF-8 byte sequence was decoded as a legacy encoding (usually Windows-1252): the two bytes of é render as the two characters é. The black diamond � (U+FFFD) is the opposite failure — bytes that are invalid in the assumed encoding, replaced with the "I give up" glyph. Both are display bugs, not data loss: reinterpret the original bytes correctly and the text returns.

The invisible characters

Beyond visible corruption, Unicode contains zero-width spaces, direction overrides, and lookalike homoglyphs — characters that are invisible or misleading in editors but very present to parsers, password fields, and plagiarism checkers. Pasted text from the web is full of them. The invisible character remover exposes and strips these ghosts locally in your browser — run any suspicious paste through it before the characters cause a bug you can't see.