how do compilers deal with utf-8 encoded source programs + Thought: 9.1s Compilers handle UTF-8 at a few distinct stages: 1. Encoding detection The compiler must decide what the file actually is. Common strategies: Assume UTF-8 by default (Rust, Go, modern C/C++ toolchains). Check for a BOM and/or honor a flag like GCC's -finput-charset=utf-8. With no BOM/flag, some pass raw bytes through unvalidated rather than guessing (C/C++). 2. Lexing/tokenization Since UTF-8 is self-synchronizing, the lexer can classify bytes by their high bits: ASCII (< 0x80), continuation bytes (10xxxxxx), or lead bytes (110/1110/11110xxx). This lets it safely skip a whole multi-byte character without mistaking its continuation bytes for punctuation or whitespace. The lexer either validates UTF-8 eagerly or just checks that string literals are well-formed. 3. Identifiers Language-dependent: Most languages (C, C++) restrict identifiers to ASCII, so a non-ASCII byte in an identifier position is simply an error. Others (Rust, Swift, Go) allow Unicode letters in identifiers and just treat any non-ASCII byte sequence as an identifier character. Java reads \uXXXX escapes and applies NFC normalization so canonically-equivalent names collide/dedupe consistently. 4. String/char literals The compiler must transcode the source bytes (UTF-8) into the execution character set and compute widths. This is where sizeof("é"), L"é" (UTF-16/32), u8"…", etc. diverge in C — the value depends on source charset, execution charset, and wchar_t width. Rust instead requires literals to be valid UTF-8 and rejects escapes producing invalid code points. 5. Diagnostics Real problems show up in error reporting: line/column are usually computed in bytes, not graphemes; and when echoing the offending source line, the compiler must avoid truncating a multi-byte character mid-sequence (or escaping invalid bytes) so the terminal doesn't print a broken character. The one rule everything hinges on is deferring to UTF-8's self-synchronizing structure: bytes 10xxxxxx are always continuation bytes, so a lexer can skip them blindly and only ever act on ASCII values or lead bytes.