# Step: Unicode — UCHAR / UString Tags `v3-unicode-1`, `v3-unicode-2`. Suite **120/120**; self-hosting fixpoint **OK** (`bootstrap/fixpoint.sh`, image 2,062,392 bytes). ## Increment 1 — types, literals, conversions - **`UCHAR`** — 32-bit codepoint (`0..1114111`), QBE `w`, distinct from `CHAR`/`INTEGER` (class `ClUChar`, form `FUChar`). Usable in `VAR`/array declarations and designators (`ARRAY OF UCHAR`). - **`U'..'` / `U".."` literals** — Coco/R `ustring` token; the parser action strict-RFC3629-decodes the bytes between the quotes (`QbeGen.DeclUStr`). One codepoint → `UCHAR` immediate; otherwise a **UString** descriptor `data $ustrN = { l count, w cp0, … }` (`QbeGen.FlushUStrings`, emitted at `EndModule`). A UString descriptor *is* a V3 length-prefixed array, so a `U"…"` literal is a valid actual for an open `ARRAY OF UCHAR` formal (`SymTab.Assignable`/`VarParamOk`). - **Strict decoding** — rejects overlongs, surrogates (`D800..DFFF`), `>10FFFF`, and truncated sequences → semantic error **234 `invalid UTF-8 in U-literal`** (message in `compiler.frm`). - **Conversions** — `UCHR(x)`: `CHAR`→`UCHAR` (identity) *or* integer→`UCHAR` (codepoint constructor, cf. `CHR`); `CHR8(u:UCHAR):CHAR` with an `l`-domain `0..255` range check (traps via `$abort`, rc 134); `UORD(u):INTEGER`. ## Increment 2 — codec + output - **`runtime/syslib/Utf8`** — `Decode` (one scalar from `s[lo]` → `cp`,`next`, `FALSE` on malformed input) and `Encode` (`UCHAR` → bytes, returns the count). `MaxUTF8Length = 4`, `MaxCodePoint = 1114111`. - **`TextIO.UWriteChar` / `UWriteString(VAR u: ARRAY OF UCHAR)` / `UWriteLn`** — encode codepoints to UTF-8 and emit the bytes via the shim; raw `WriteChar` stays byte-oriented. A `U"héllo"` literal can be passed directly to `UWriteString`. ## Deviations from the locked design (flagged) - **`UORD` returns `INTEGER`, not `LONGCARD`.** V3 has no `LONGCARD` type name and no 64→32-bit conversion, so a `LONGCARD` result would be unusable in codegen (`CHR`, `DIV`/`MOD`, comparisons). A codepoint fits in 32 bits, so `INTEGER` is lossless and consistent with `ORD` for `CHAR`. - **`UCHR` accepts an integer family operand** as well as `CHAR`, so a computed codepoint can be built without a literal. - `.ssa` is still written when error 234 fires (image gate pending). ## Tests (`compiler/run_tests.sh`) - `t_unicode.mod` (exit 31): `U'é'`, `U'😀'`, `U'€'`, `UCHR`, `CHR8`, `UORD`. - `t_trap_chr8.mod`: `CHR8(U'😀')` → abort (rc 134). - `t_bad_overlong.mod`, `t_bad_surrogate.mod`, `t_bad_utf8big.mod`, `t_bad_truncutf.mod`: expect `invalid UTF-8 in U-literal` (234); raw invalid bytes in the sources, `.LST` output git-ignored. - `utf8_prog.mod` (exit 127): `Utf8.Encode`/`Decode` round-trips. - `utext_prog.mod`: `TextIO.UWriteString(U"héllo")` → stdout `héllo` (UTF-8 bytes verified). ## Deferred - `UString` whole-string assignment to a *fixed* array (`a := U"…"`): the runtime count check traps, exactly as `a := "…"` does for `CHAR` arrays. Passing a literal to an open `ARRAY OF UCHAR` formal works. - Concatenation/`=`/`+` for UStrings, `UCHAR` `CASE` labels. ## Files `compiler/src/{SymTab.def,SymTab.mod}` (UCHAR type, ordinal/equality rules, UString↔open-array compat), `compiler/src/{QbeGen.def,QbeGen.mod}` (`DeclUStr`/`FlushUStrings`), `compiler/src/M2.atg` (`ustring` token + `Fact` branch + `UCHR`/`CHR8`/`UORD`), `compiler/src/compiler.frm` (message 234), `runtime/syslib/Utf8.def`/`.mod`, `stdlib/textio.def`/`.mod`, `compiler/run_tests.sh`, `compiler/tests/`.