summary_unicode.md 3.6 KB

Step: Unicode — UCHAR / UString

Tags v3-unicode-1, v3-unicode-2. Suite 120/120; self-hosting fixpoint OK (bootstrap/fixpoint.sh, image 2,062,392 bytes).

Increment 1 — types, literals, conversions

  • UCHAR — 32-bit codepoint (0..1114111), QBE w, distinct from CHAR/INTEGER (class ClUChar, form FUChar). Usable in VAR/array declarations and designators (ARRAY OF UCHAR).
  • U'..' / U".." literals — Coco/R ustring token; the parser action strict-RFC3629-decodes the bytes between the quotes (QbeGen.DeclUStr). One codepoint → UCHAR immediate; otherwise a UString descriptor data $ustrN = { l count, w cp0, … } (QbeGen.FlushUStrings, emitted at EndModule). A UString descriptor is a V3 length-prefixed array, so a U"…" literal is a valid actual for an open ARRAY OF UCHAR formal (SymTab.Assignable/VarParamOk).
  • Strict decoding — rejects overlongs, surrogates (D800..DFFF), >10FFFF, and truncated sequences → semantic error 234 invalid UTF-8 in U-literal (message in compiler.frm).
  • Conversions — UCHR(x): CHAR→UCHAR (identity) or integer→UCHAR (codepoint constructor, cf. CHR); CHR8(u:UCHAR):CHAR with an l-domain 0..255 range check (traps via $abort, rc 134); UORD(u):INTEGER.

Increment 2 — codec + output

  • runtime/syslib/Utf8 — Decode (one scalar from s[lo] → cp,next, FALSE on malformed input) and Encode (UCHAR → bytes, returns the count). MaxUTF8Length = 4, MaxCodePoint = 1114111.
  • TextIO.UWriteChar / UWriteString(VAR u: ARRAY OF UCHAR) / UWriteLn — encode codepoints to UTF-8 and emit the bytes via the shim; raw WriteChar stays byte-oriented. A U"héllo" literal can be passed directly to UWriteString.

Deviations from the locked design (flagged)

  • UORD returns INTEGER, not LONGCARD. V3 has no LONGCARD type name and no 64→32-bit conversion, so a LONGCARD result would be unusable in codegen (CHR, DIV/MOD, comparisons). A codepoint fits in 32 bits, so INTEGER is lossless and consistent with ORD for CHAR.
  • UCHR accepts an integer family operand as well as CHAR, so a computed codepoint can be built without a literal.
  • .ssa is still written when error 234 fires (image gate pending).

Tests (compiler/run_tests.sh)

  • t_unicode.mod (exit 31): U'é', U'😀', U'€', UCHR, CHR8, UORD.
  • t_trap_chr8.mod: CHR8(U'😀') → abort (rc 134).
  • t_bad_overlong.mod, t_bad_surrogate.mod, t_bad_utf8big.mod, t_bad_truncutf.mod: expect invalid UTF-8 in U-literal (234); raw invalid bytes in the sources, .LST output git-ignored.
  • utf8_prog.mod (exit 127): Utf8.Encode/Decode round-trips.
  • utext_prog.mod: TextIO.UWriteString(U"héllo") → stdout héllo (UTF-8 bytes verified).

Deferred

  • UString whole-string assignment to a fixed array (a := U"…"): the runtime count check traps, exactly as a := "…" does for CHAR arrays. Passing a literal to an open ARRAY OF UCHAR formal works.
  • Concatenation/=/+ for UStrings, UCHAR CASE labels.

Files

compiler/src/{SymTab.def,SymTab.mod} (UCHAR type, ordinal/equality rules, UString↔open-array compat), compiler/src/{QbeGen.def,QbeGen.mod} (DeclUStr/FlushUStrings), compiler/src/M2.atg (ustring token + Fact branch + UCHR/CHR8/UORD), compiler/src/compiler.frm (message 234), runtime/syslib/Utf8.def/.mod, stdlib/textio.def/.mod, compiler/run_tests.sh, compiler/tests/.