Detect and Repair Double-Encoded UTF-8 Text
Recognize common double-encoded UTF-8 patterns, reverse a known Windows-1252 mis-decoding in PHP, and avoid corrupting already-correct rows.
The easy way to know
Recognize common double-encoded UTF-8 patterns, reverse a known Windows-1252 mis-decoding in PHP, and avoid corrupting already-correct rows.
Understand why visually identical Unicode strings can compare differently and when to normalize text to NFC or NFD in JavaScript and databases.
Find and repair malformed UTF-8 before JSON encoding in PHP, preserve error visibility, and avoid silently discarding corrupted input.
Diagnose and repair garbled text after a MySQL migration by separating stored bytes, column metadata, connection character sets, and application output.
Convert Windows-1252 text and files to UTF-8 without losing smart quotes, dashes, euro signs, or other characters commonly mistaken for ISO-8859-1.
A practical reference for reserved HTML characters, spacing, typography, currency, and arrows.
Convert legacy CSV files to UTF-8 while preserving quoted fields and delimiters.
Detect and safely remove a UTF-8 BOM when it breaks headers, JSON, CSV, or scripts.
Understand how the HTTP Content-Type header and HTML meta charset declaration control decoding.
Find Unicode code points correctly in JavaScript, including supplementary characters.