Developer

Why Special Characters Break: UTF-8, Latin-1, Byte Order Marks, and Mojibake Explained

Stop corrupted symbols and question marks in text files. Understand how UTF-8 multibyte encoding works, why Mojibake happens, and how to fix encoding mismatches.

The Synctoolo Team··10 min read
Macro view of microchips and integrated circuit boards conveying raw binary encoding

You export a customer database to a CSV file or import an API payload, open it in a spreadsheet or database viewer, and discover utter chaos: "café" has morphed into "café", an em dash has turned into "—", and curly quotes have become "“".

This character corruption phenomenon is known internationally as Mojibake (from the Japanese word meaning "character transformation"). It has plagued developers, data engineers, and content creators since the dawn of digital computing.

Contrary to popular belief, character corruption is rarely random data loss. It is a precise mathematical misunderstanding between how bytes were written and how they were decoded. This guide explains how UTF-8 variable-length encoding works at the bit level, why legacy encodings like Windows-1252 and ISO-8859-1 cause corruption, and how to fix corrupted strings. You can inspect raw binary bytes and Base64 representations in browser memory using Synctoolo's free Base64 Encoder / Decoder and format payloads with our JSON Formatter.

The Bitwise Mechanics of UTF-8

Computers store and transmit raw bytes (sequences of 8 bits, values from 0 to 255). An encoding standard defines how numbers map to human-readable characters.

The original ASCII standard only defined 128 characters (using 7 bits), covering English letters, numbers, and basic punctuation. When computing went global, thousands of accented letters, Cyrillic glyphs, Asian ideograms, and emoji needed representation.

Invented by Ken Thompson and Rob Pike in 1992, UTF-8 solved this problem through a backward-compatible, variable-length scheme:

Unicode Code Point Range Byte Length UTF-8 Bit Pattern (Binary) Character Types
U+0000 - U+007F 1 Byte 0xxxxxxx Standard ASCII (English letters, numbers)
U+0080 - U+07FF 2 Bytes 110xxxxx 10xxxxxx Accented Latin, Greek, Cyrillic, Arabic
U+0800 - U+FFFF 3 Bytes 1110xxxx 10xxxxxx 10xxxxxx Chinese, Japanese, Korean, symbols
U+10000 - U+10FFFF 4 Bytes 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx Emoji, historical scripts, math symbols

Notice the clever design: if a byte starts with 0, it is a single-byte ASCII character. If it starts with 110, it signals a 2-byte sequence. Continuation bytes always begin with 10.

Movable metal type printer blocks with individual character glyphs
Mojibake occurs when text encoded in multibyte UTF-8 is misread by legacy systems as single-byte Windows-1252 or ISO-8859-1. Photo by Florian Klauer on Unsplash.

Why "café" Becomes "café": The Exact Failure Mechanism

Let us trace the exact lifecycle of the word "café":

  1. The accented letter é has Unicode code point U+00E9.
  2. In UTF-8, U+00E9 is encoded as two distinct bytes: 0xC3 0xA9 (in decimal: 195 169).
  3. The source application correctly writes these bytes to disk: c (0x63), a (0x61), f (0x66), [0xC3, 0xA9].
  4. A legacy program (like older versions of Microsoft Excel or a misconfigured MySQL database) opens the file using Windows-1252 (Latin-1) encoding. Windows-1252 assumes every single byte represents one character.
  5. In Windows-1252:
    • Byte 0xC3 maps to à (Capital A with tilde).
    • Byte 0xA9 maps to © (Copyright symbol).
  6. The program displays café. No data was deleted; the bytes were simply interpreted through the wrong translation table.

The Rosetta Stone of Common Mojibake Glitches

Intended Character UTF-8 Hex Bytes Corrupted Mojibake String (Windows-1252)
é (Accented e) C3 A9 é
ü (Umlaut u) C3 BC ü
ñ (Tilde n) C3 B1 ñ
— (Long dash) E2 80 94 —
“ (Left curly quote) E2 80 9C “
’ (Apostrophe) E2 80 99 ’
€ (Euro symbol) E2 82 AC €

The Byte Order Mark (BOM) Dilemma

In UTF-16, a Byte Order Mark (BOM) indicates whether bytes are arranged in big-endian or little-endian format. In UTF-8, endianness is irrelevant because characters are processed sequentially byte-by-byte.

However, Microsoft software (specifically Excel on Windows) often requires a UTF-8 BOM - three initial bytes: 0xEF 0xBB 0xBF - to recognize that a CSV file is UTF-8 rather than system ANSI. If you omit the BOM, Excel defaults to Windows-1252 and corrupts international characters. Conversely, Unix tools and web compilers often crash when encountering an unexpected BOM in code files.

How to Prevent Character Encoding Bugs

  1. Set Explicit HTTP Content-Type Headers: Always serve text with Content-Type: text/html; charset=utf-8 or application/json; charset=utf-8.
  2. Configure Database Collations to utf8mb4: In MySQL, standard utf8 only supports 3-byte characters, causing silent failures on 4-byte emoji. Always use utf8mb4 and utf8mb4_unicode_ci.
  3. Include the Meta Charset Tag: Ensure <meta charset="utf-8" /> appears in the <head> of all HTML templates before any scripts or stylesheets.

Tools mentioned in this article

FAQ

How do I fix a file that is already corrupted with Mojibake?+

If the file was saved without losing data, you can reverse the corruption by reading the file as Windows-1252 or Latin-1 to extract the raw bytes, and then re-encoding those exact bytes back to UTF-8. Many text editors (like VS Code) allow you to 'Reopen with Encoding -> Windows-1252' and then 'Save with Encoding -> UTF-8'.

What causes the black diamond with a question mark ()?+

The black diamond question mark (Unicode character U+FFFD) is the official 'Replacement Character'. Unlike Mojibake (where valid bytes are read through the wrong table), the replacement character appears when an application encounters a byte sequence that violates UTF-8 syntax rules and discards the invalid byte entirely.

Why does MySQL have both utf8 and utf8mb4?+

Early MySQL versions implemented an incomplete UTF-8 standard that supported a maximum of 3 bytes per character. When Unicode expanded to 4 bytes for emoji and complex Asian characters, MySQL introduced utf8mb4 ('UTF-8 Most Bytes 4'). Using standard 'utf8' in MySQL will truncate strings or crash whenever an emoji is inserted.

Should I include a BOM in my UTF-8 files?+

For web development, JavaScript, JSON, and source code, never include a Byte Order Mark (BOM). The only scenario where a UTF-8 BOM is helpful is when generating downloadable CSV files specifically intended to open correctly in Microsoft Excel on Windows.

S
The Synctoolo Team

We build and review free, privacy-first tools at Synctoolo.

Keep reading