Why Special Characters Break: UTF-8, Latin-1, Byte Order Marks, and Mojibake Explained
Stop corrupted symbols and question marks in text files. Understand how UTF-8 multibyte encoding works, why Mojibake happens, and how to fix encoding mismatches.

You export a customer database to a CSV file or import an API payload, open it in a spreadsheet or database viewer, and discover utter chaos: "café" has morphed into "café", an em dash has turned into "—", and curly quotes have become "“".
This character corruption phenomenon is known internationally as Mojibake (from the Japanese word meaning "character transformation"). It has plagued developers, data engineers, and content creators since the dawn of digital computing.
Contrary to popular belief, character corruption is rarely random data loss. It is a precise mathematical misunderstanding between how bytes were written and how they were decoded. This guide explains how UTF-8 variable-length encoding works at the bit level, why legacy encodings like Windows-1252 and ISO-8859-1 cause corruption, and how to fix corrupted strings. You can inspect raw binary bytes and Base64 representations in browser memory using Synctoolo's free Base64 Encoder / Decoder and format payloads with our JSON Formatter.
The Bitwise Mechanics of UTF-8
Computers store and transmit raw bytes (sequences of 8 bits, values from 0 to 255). An encoding standard defines how numbers map to human-readable characters.
The original ASCII standard only defined 128 characters (using 7 bits), covering English letters, numbers, and basic punctuation. When computing went global, thousands of accented letters, Cyrillic glyphs, Asian ideograms, and emoji needed representation.
Invented by Ken Thompson and Rob Pike in 1992, UTF-8 solved this problem through a backward-compatible, variable-length scheme:
| Unicode Code Point Range | Byte Length | UTF-8 Bit Pattern (Binary) | Character Types |
|---|---|---|---|
U+0000 - U+007F |
1 Byte | 0xxxxxxx |
Standard ASCII (English letters, numbers) |
U+0080 - U+07FF |
2 Bytes | 110xxxxx 10xxxxxx |
Accented Latin, Greek, Cyrillic, Arabic |
U+0800 - U+FFFF |
3 Bytes | 1110xxxx 10xxxxxx 10xxxxxx |
Chinese, Japanese, Korean, symbols |
U+10000 - U+10FFFF |
4 Bytes | 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx |
Emoji, historical scripts, math symbols |
Notice the clever design: if a byte starts with 0, it is a single-byte ASCII character. If it starts with 110, it signals a 2-byte sequence. Continuation bytes always begin with 10.
Why "café" Becomes "café": The Exact Failure Mechanism
Let us trace the exact lifecycle of the word "café":
- The accented letter
éhas Unicode code pointU+00E9. - In UTF-8,
U+00E9is encoded as two distinct bytes:0xC3 0xA9(in decimal:195 169). - The source application correctly writes these bytes to disk:
c (0x63), a (0x61), f (0x66), [0xC3, 0xA9]. - A legacy program (like older versions of Microsoft Excel or a misconfigured MySQL database) opens the file using Windows-1252 (Latin-1) encoding. Windows-1252 assumes every single byte represents one character.
- In Windows-1252:
- Byte
0xC3maps to à (Capital A with tilde). - Byte
0xA9maps to © (Copyright symbol).
- Byte
- The program displays café. No data was deleted; the bytes were simply interpreted through the wrong translation table.
The Rosetta Stone of Common Mojibake Glitches
| Intended Character | UTF-8 Hex Bytes | Corrupted Mojibake String (Windows-1252) |
|---|---|---|
é (Accented e) |
C3 A9 |
é |
ü (Umlaut u) |
C3 BC |
ü |
ñ (Tilde n) |
C3 B1 |
ñ |
— (Long dash) |
E2 80 94 |
— |
“ (Left curly quote) |
E2 80 9C |
“ |
’ (Apostrophe) |
E2 80 99 |
’ |
€ (Euro symbol) |
E2 82 AC |
€ |
The Byte Order Mark (BOM) Dilemma
In UTF-16, a Byte Order Mark (BOM) indicates whether bytes are arranged in big-endian or little-endian format. In UTF-8, endianness is irrelevant because characters are processed sequentially byte-by-byte.
However, Microsoft software (specifically Excel on Windows) often requires a UTF-8 BOM - three initial bytes: 0xEF 0xBB 0xBF - to recognize that a CSV file is UTF-8 rather than system ANSI. If you omit the BOM, Excel defaults to Windows-1252 and corrupts international characters. Conversely, Unix tools and web compilers often crash when encountering an unexpected BOM in code files.
How to Prevent Character Encoding Bugs
- Set Explicit HTTP Content-Type Headers: Always serve text with
Content-Type: text/html; charset=utf-8orapplication/json; charset=utf-8. - Configure Database Collations to utf8mb4: In MySQL, standard
utf8only supports 3-byte characters, causing silent failures on 4-byte emoji. Always useutf8mb4andutf8mb4_unicode_ci. - Include the Meta Charset Tag: Ensure
<meta charset="utf-8" />appears in the<head>of all HTML templates before any scripts or stylesheets.
Tools mentioned in this article
FAQ
How do I fix a file that is already corrupted with Mojibake?+
If the file was saved without losing data, you can reverse the corruption by reading the file as Windows-1252 or Latin-1 to extract the raw bytes, and then re-encoding those exact bytes back to UTF-8. Many text editors (like VS Code) allow you to 'Reopen with Encoding -> Windows-1252' and then 'Save with Encoding -> UTF-8'.
What causes the black diamond with a question mark ()?+
The black diamond question mark (Unicode character U+FFFD) is the official 'Replacement Character'. Unlike Mojibake (where valid bytes are read through the wrong table), the replacement character appears when an application encounters a byte sequence that violates UTF-8 syntax rules and discards the invalid byte entirely.
Why does MySQL have both utf8 and utf8mb4?+
Early MySQL versions implemented an incomplete UTF-8 standard that supported a maximum of 3 bytes per character. When Unicode expanded to 4 bytes for emoji and complex Asian characters, MySQL introduced utf8mb4 ('UTF-8 Most Bytes 4'). Using standard 'utf8' in MySQL will truncate strings or crash whenever an emoji is inserted.
Should I include a BOM in my UTF-8 files?+
For web development, JavaScript, JSON, and source code, never include a Byte Order Mark (BOM). The only scenario where a UTF-8 BOM is helpful is when generating downloadable CSV files specifically intended to open correctly in Microsoft Excel on Windows.
We build and review free, privacy-first tools at Synctoolo.
Keep reading

Stop Regular Expression Denial of Service (ReDoS). Discover why nested quantifiers cause exponential execution times and how to test patterns safely.

Solve CORS blocking in web applications. Learn how Access-Control-Allow-Origin works, how to handle OPTIONS preflight requests, and how to debug headers locally.