how do compilers deal with utf-8 encoded source programs
+ Thought: 9.1s
Compilers handle UTF-8 at a few distinct stages:
1. Encoding detection
The compiler must decide what the file actually is. Common strategies:
Assume UTF-8 by default (Rust, Go, modern C/C++ toolchains).
Check for a BOM and/or honor a flag like GCC's -finput-charset=utf-8.
With no BOM/flag, some pass raw bytes through unvalidated rather than guessing (C/C++).
2. Lexing/tokenization
Since UTF-8 is self-synchronizing, the lexer can classify bytes by their high bits: ASCII (< 0x80), continuation bytes (10xxxxxx), or lead bytes (110/1110/11110xxx). This lets it safely skip a whole multi-byte character without mistaking its continuation bytes for punctuation or whitespace. The lexer either validates UTF-8 eagerly or just checks that string literals are well-formed.
3. Identifiers
Language-dependent:
Most languages (C, C++) restrict identifiers to ASCII, so a non-ASCII byte in an identifier position is simply an error.
Others (Rust, Swift, Go) allow Unicode letters in identifiers and just treat any non-ASCII byte sequence as an identifier character.
Java reads \uXXXX escapes and applies NFC normalization so canonically-equivalent names collide/dedupe consistently.
4. String/char literals
The compiler must transcode the source bytes (UTF-8) into the execution character set and compute widths. This is where sizeof("é"), L"é" (UTF-16/32), u8"…", etc. diverge in C — the value depends on source charset, execution charset, and wchar_t width. Rust instead requires literals to be valid UTF-8 and rejects escapes producing invalid code points.
5. Diagnostics
Real problems show up in error reporting: line/column are usually computed in bytes, not graphemes; and when echoing the offending source line, the compiler must avoid truncating a multi-byte character mid-sequence (or escaping invalid bytes) so the terminal doesn't print a broken character.
The one rule everything hinges on is deferring to UTF-8's self-synchronizing structure: bytes 10xxxxxx are always continuation bytes, so a lexer can skip them blindly and only ever act on ASCII values or lead bytes.
