Просмотр исходного кода

Runtime: five more emitter bugs, and the checks that found them

Every byte in Runtime.mod is hand-assembled, and a wrong ModRM byte decodes
cleanly. The structural check, the disassembly golden and "it compiles" all
accept a perfectly well-formed instruction that does the wrong thing, so
none of them can see this class of bug. What can see it is saying what the
instruction is supposed to do and comparing that against what the bytes
actually mean.

Five more bugs, all found that way:

  MovSiBx   89 DC -> 89 DE   89 DC is MOV SP,BX
  CmpSiBx   39 DC -> 39 DE   39 DC is CMP SP,BX
  MovSiAx   8B C0 -> 8B F0   8B C0 is MOV AX,AX, a no-op, so initmem's SI
                                was never set and the whole header read was
                                operating on whatever SI happened to hold
  initmem   its "zeroing" loop used MovAxDx, so it wrote the header offset
            into memory instead of 0; replaced with XorAxAx, and the dead
            MovAxAxDx deleted
  initmem   MovCxSi6 read header word +8 (hdrMax, patched to 0) instead of
            +6 (hdrHeap), so `cmp cx,dx; jbe done` was 0 <= 0 and the loop
            zeroed nothing at all; dead MovCxSi8 deleted

The last one is the instructive one. It is not a wrong byte -- every byte was
correct -- it is a correct loop that never runs, and no amount of
disassembly would have found it. Runtime is 385 -> 391 bytes, code region
0..359, 14 entries, 37 branch and call targets checked.

The checks, all of which are non-vacuous by construction:

  audit_helpers.py   disassembles every one-line emitter in Runtime.mod and
                     compares the FCML decode against the procedure's NAME.
                     61 helpers, 61 agree. The name grammar is an ordered
                     pattern table with an opcode column, because MovBpSp
                     is BP:=SP for 8B but MovDlSi is DL:=[SI] for 8A.
                     UNPARSED names count as failures, so a name the grammar
                     cannot explain is a red test rather than a silent pass.
  check_runtime.py   sweeps the built code region with FCML: no desync, every
                     entry and branch target on an instruction boundary, and
                     the entire disassembly equal to tests/runtime.golden
                     (193 instructions). --bless re-baselines, and the
                     procedure for doing that deliberately is documented.
  nonvacuity.sh      13 assertions that the checks above can actually fail,
                     each breaking the source in a way that was a real bug
                     and asserting the check goes red FOR THE STATED REASON.
  disasm16.py        FCML linear sweep. MAX_WINDOW is 15, because the Debian
                     fcml-disasm wrapper aborts with rc=134 on inputs of 16
                     bytes or more.
  fcml_vs_objdump.py 1155 comparisons against objdump -b binary -m i8086.
                     4 disagreements, all F3/82, all real FCML quirks.
                     Corrects an earlier conclusion in this repo that there
                     was no working 16-bit disassembler available.

Stale constants re-baselined with the rationale recorded next to them:
RTSZ 385 -> 391 and the initmem HEAD bytes in run_com_tests.sh, comtest.py
and CompileTest.mod. The checkers deliberately restate the layout constants
instead of asking the code under test, and assert assertion-by-assertion
that initmem's SI displacements equal the header word offsets they verify --
that cross-check is what would have caught the MovCxSi6 bug.

The ModR/M table is now measured rather than remembered, which is what makes
the above possible:

  probe/modrm19.s     mod=00/01/10 by EXECUTION on a real 8086 under qemu.
                      24 cases; each stores a marker through the encoding and
                      reports the address that received it. 23 of 24 cells
                      confirmed; mod=10 rm=001 lands outside the probe's scan
                      window and is recorded as a gap, not a result.
  probe/run_modrm19.py re-runs the above and requires the measured offsets to
                      equal the register arithmetic computed from first
                      principles.
  probe/modrm11.s     mod=11 by ENCODING: GNU as (.code16) is asked to encode
  probe/modrm11.py    the register moves, FCML decodes the result, and all of
                      it must agree with hard-coded EXPECT bytes, with four
                      anchors nobody types by hand (83 C4 08 ADD SP,8 /
                      83 C6 02 ADD SI,2 / 8B EC MOV BP,SP / 8B E5 MOV SP,BP),
                      and with the table in Runtime.mod. mod=11 names a
                      register rather than an address, so modrm19's method
                      cannot be extended to it; the reasons a naive extension
                      fails are in the .s header, because the first attempt
                      produced two impossible answers and a plausible third.

  probe/README.md     what each artifact establishes, and the mistake in full.

Writing that probe turned up a documentation bug, not a code bug. The mod=11
column in Runtime.mod read

    rm:  CX DX BX SP BP SI DI BX

which is the correct list with AX dropped off the front and a duplicate BX
invented at the end, so every code is one too low -- except 100, which lands
on SP either way, which is why it survived. The correct column is

    r/m:  AX CX DX BX SP BP SI DI

and it is the same list as the reg field; there is no second BX at code 7.
The three emitters above were genuinely wrong bytes and their fixes stand
(89 DE is MOV SI,BX; 8B F0 is MOV SI,AX, because for 8B the reg field is the
destination). The comment is corrected and now states the four anchors, so
the column can be checked without trusting anybody's memory.

The mod=11 check is deliberately not self-referential. Comparing the .s
against `as` cannot fail on its own -- edit the .s and as faithfully encodes
the new claim -- which was found by corrupting the .s and watching the check
stay green. EXPECT and the anchors close that; nonvacuity.sh exercises both,
along with a corruption that restores the exact table this project shipped.

Also untracked shell/comtest: it is a build artifact of tests/run_com_tests.sh
and shell/.gitignore already listed it, but 7dada57 committed it by mistake.

Full suite green, tpshell 163000 bytes, runtime 391 bytes, 24 fixtures linked
and byte-checked. No milestone tag: no TP3-compiled .COM has ever been
executed yet.
Eric Streit 2 недель назад
Родитель
Сommit
962018cd34

+ 5 - 0
shell/Runtime.def

@@ -40,4 +40,9 @@ PROCEDURE RT_Byte (i : CARDINAL) : BYTE ;
 PROCEDURE RT_Entry (i : CARDINAL) : CARDINAL ;
 (* offset of entry i (an E_* selector) inside the runtime; 0 if unknown *)
 
+PROCEDURE RT_CodeEnd () : CARDINAL ;
+(* offset where the code stops and the runtime's data block starts.  Bytes
+   from here on are strings and scratch, so disassembling them is
+   meaningless; tests use this to sweep the code region only. *)
+
 END Runtime.

+ 119 - 25
shell/Runtime.mod

@@ -178,6 +178,71 @@ END Dd ;
 
 (* --- 8-bit / 16-bit register and memory forms, one instruction each --- *)
 
+(*  THE 16-BIT ModR/M EFFECTIVE-ADDRESS TABLE, MEASURED NOT REMEMBERED.
+    Every address form below was confirmed by EXECUTING it on a real 8086
+    (qemu-system-i386) with a probe that stores a marker through the candidate
+    encoding and then reports which physical address received it; see
+    tests/probe/modrm19.s.  The mod=11 column is not an address at all -- it
+    names a register -- so it is confirmed instead by asking GNU as to
+    ENCODE the eight register moves and checking FCML's decode of the result
+    (tests/probe/modrm11.py, which also carries the hand-checking anchors).
+
+    Do not "fix" any of these from memory.  The version that comes to mind is
+    wrong in exactly the cells called out below, and every one of those
+    mistakes shipped as code that decoded cleanly:
+
+       r/m   mod=00        mod=01          mod=10          mod=11
+       000   [BX+SI]       [BX+SI]+disp    [BX+SI]+disp    AX
+       001   [BX+DI]       [BX+DI]+disp    [BX+DI]+disp    CX
+       010   [BP+SI]       [BP+SI]+disp    [BP+SI]+disp    DX
+       011   [BP+DI]       [BP+DI]+disp    [BP+DI]+disp    BX
+       100   [SI]          [SI]+disp       [SI]+disp       SP
+       101   [DI]          [DI]+disp       [DI]+disp       BP
+       110   disp16        [BP]+disp       [BP]+disp       SI
+       111   [BX]          [BX]+disp       [BX]+disp       DI
+
+    "disp" is disp8 for mod=01 and disp16 for mod=10.
+
+    In mod=11 BOTH fields name a register and BOTH use the same list,
+    AX CX DX BX SP BP SI DI -- the reg field and the r/m field do not
+    differ, and there is no second BX at code 7.  The memorable version
+    drops AX off the front and invents a duplicate BX at the end, which
+    shifts every code down by one; that is precisely how MovSiBx and
+    CmpSiBx came to be written "89 DC" and "39 DC", which are MOV SP,BX
+    and CMP SP,BX.  Note what the shifted table gets right by luck: SP is
+    at 100 either way, so the mistake is invisible until you check a
+    register below it.  rm=100 is SP, and SI is rm=110, not rm=100.
+
+    And reg/rm swap direction with the opcode, which is the other trap:
+       88 /r  MOV r/m8,r8    89 /r  MOV r/m16,r16    reg is the SOURCE
+       8A /r  MOV r8,r/m8    8B /r  MOV r16,r/m16    reg is the TARGET
+    So ModRM C6 names "DH and AL" either way, but 88 C6 is DH:=AL while
+    8A C6 is AL:=DH.  The 8-bit register list is also its own, and it is not
+    the word list:
+       000 001 010 011 100 101 110 111  =  AL CL DL BL AH CH DH BH
+    The two lists agree at every code except 100, where the byte form is AH
+    and the word form is SP.  Reading the same ModRM byte as the other size
+    silently swaps AH for SP, and 88 C2 is DL:=AL while 88 17 is [BX]<-DL.
+
+    The word-form list, and the four anchors nobody writes by hand:
+       83 C4 08   ADD SP, 8    rm=100 -> SP
+       83 C6 02   ADD SI, 2    rm=110 -> SI
+       8B EC      MOV BP, SP   reg=101 rm=100
+       8B E5      MOV SP, BP   reg=100 rm=101
+    Asking "which register is rm=100?" and answering SI is the single most
+    common error in this file's history.
+
+    Concrete recipes for opcode 8r / 9r (r/m = rm, 16-bit form):
+       mod=00 rm=110 -> 06  <disp16>     the only direct form
+       mod=00 rm=111 -> 07  = [BX]       2 bytes
+       mod=00 rm=101 -> 05  = [DI]       2 bytes
+       mod=01 rm=110 -> 46  disp8        = [BP]+disp8
+       mod=01 rm=111 -> 47  disp8        = [BX]+disp8
+       mod=01 rm=101 -> 45  disp8        = [DI]+disp8
+    There is no [SP] form in 16-bit mode: SIB bytes are 386-only.  Anything
+    wanting the top of the stack has to go through BP. *)
+
+
 PROCEDURE PushBp ; BEGIN B (55H) END PushBp ;
 PROCEDURE PopBp  ; BEGIN B (5DH) END PopBp ;
 PROCEDURE MovBpSp ; BEGIN B (8BH) ; B (0ECH) END MovBpSp ;  (* 8B EC: MOV BP,SP *)
@@ -212,8 +277,20 @@ PROCEDURE AddDi2 ; BEGIN B (83H) ; B (0C7H) ; B (2) END AddDi2 ;  (* 11 000 111
 PROCEDURE CmpAl (v : CARDINAL ) ; BEGIN B (3CH) ; B (v) END CmpAl ;
 PROCEDURE CmpAx0 ; BEGIN B (83H) ; B (0F8H) ; B (0) END CmpAx0 ;
 PROCEDURE CmpCx0 ; BEGIN B (83H) ; B (0F9H) ; B (0) END CmpCx0 ;
-PROCEDURE CmpSpW0 ; BEGIN B (83H) ; B (7CH) ; B (24H) ; B (0) ; B (0) END CmpSpW0 ;
-PROCEDURE CmpSiBx ; BEGIN B (39H) ; B (0DCH) END CmpSiBx ;  (* 11 011 100 *)
+PROCEDURE CmpArgW0 ;
+(* CMP WORD PTR [BP+2],0 - the one 16-bit argument of a BOOLEAN entry.
+   [SP] cannot be encoded in 16-bit mode, so borrow BP for the three
+   instructions and hand SP back untouched before the caller pops the
+   argument.  (Was 83 7C 24 00 00, which decodes as CMP WORD [SI+24h],0.)
+   The imm8 of the 83 form is sign-extended to 16 bits, so 0 really is a
+   16-bit zero. *)
+BEGIN
+   B (8BH) ; B (0ECH) ;                 (* MOV BP,SP          *)
+   B (83H) ; B (7EH) ; B (2) ; B (0) ;  (* CMP WORD [BP+2],0  *)
+   B (89H) ; B (0ECH)                   (* MOV SP,BP          *)
+END CmpArgW0 ;
+
+PROCEDURE CmpSiBx ; BEGIN B (39H) ; B (0DEH) END CmpSiBx ;  (* 39 DE: CMP SI,BX *)
 PROCEDURE CmpCxDx ; BEGIN B (39H) ; B (0D1H) END CmpCxDx ;  (* 11 010 001 *)
 PROCEDURE CmpDiCx ; BEGIN B (39H) ; B (0CFH) END CmpDiCx ;  (* 11 001 111 *)
 
@@ -232,10 +309,8 @@ PROCEDURE MovDxV (v : CARDINAL ) ; BEGIN B (0BAH) ; W (v) END MovDxV ;
 PROCEDURE MovBxD (delta : CARDINAL ) ; BEGIN B (0BBH) ; Dd (delta) END MovBxD ;
 PROCEDURE MovDxD (delta : CARDINAL ) ; BEGIN B (0BAH) ; Dd (delta) END MovDxD ;
 
-PROCEDURE MovAxSp  ; BEGIN B (8BH) ; B (44H) ; B (24H) ; B (0) END MovAxSp ;
-PROCEDURE MovSiAx  ; BEGIN B (8BH) ; B (0C0H) END MovSiAx ;
+PROCEDURE MovSiAx  ; BEGIN B (8BH) ; B (0F0H) END MovSiAx ;  (* 8B F0: MOV SI,AX *)
 PROCEDURE MovAxDi  ; BEGIN B (8BH) ; B (0C7H) END MovAxDi ;   (* 11 000 111 *)
-PROCEDURE MovAxDx  ; BEGIN B (8BH) ; B (0D2H) END MovAxDx ;
 PROCEDURE MovCxSi6 ; BEGIN B (8BH) ; B (4CH) ; B (6) END MovCxSi6 ;
 PROCEDURE MovDxSi2 ; BEGIN B (8BH) ; B (54H) ; B (2) END MovDxSi2 ;
 (* The program header is a block of words laid out by the compiler:
@@ -244,26 +319,30 @@ PROCEDURE MovDxSi2 ; BEGIN B (8BH) ; B (54H) ; B (2) END MovDxSi2 ;
       +4  hdrDS     base of the data area     <- initmem wants these two
       +6  hdrHeap   end of the data area      <-
       +8  hdrMax    max open files
-   MovDxSi4/MovCxSi8 read the two that matter here.  The obvious +2/+6 would
-   be the code end and the data end, i.e. initmem would zero from the end of
-   the code to the end of the data - a 4 KiB gap of nothing, and the globals
-   themselves untouched.  Same instruction length, so no entry offset moves. *)
-PROCEDURE MovCxSi8 ; BEGIN B (8BH) ; B (4CH) ; B (8) END MovCxSi8 ;
+   initmem must read the words the COMPILER WRITES, and those are +4 (hdrDS,
+   the data base) and +6 (hdrHeap, the data end).  It used to read +8, which
+   is hdrMax -- and the compiler patches that to 0 -- so CX came out as 0,
+   "cmp cx,dx / jbe im_done" fired immediately, and initmem silently zeroed
+   nothing at all.  A 3-byte instruction, so no entry offset moved either
+   way; only the header offset in the byte changed.  Nothing caught it
+   because the loop was well formed - it just did nothing.  The cross-check
+   that would have: tests/run_com_tests.sh now asserts the runtime's SI-relative
+   reads against the same header offsets it verifies the words at. *)
 PROCEDURE MovDxSi4 ; BEGIN B (8BH) ; B (54H) ; B (4) END MovDxSi4 ;
-PROCEDURE MovAxBp4 ; BEGIN B (8BH) ; B (45H) ; B (4) END MovAxBp4 ;
-PROCEDURE MovBxBp4 ; BEGIN B (8BH) ; B (5EH) ; B (4) END MovBxBp4 ;
-PROCEDURE MovAlDh  ; BEGIN B (8AH) ; B (0C0H) END MovAlDh ;
+PROCEDURE MovAxBp4 ; BEGIN B (8BH) ; B (46H) ; B (4) END MovAxBp4 ;  (* AX:=[BP+4] *)
+PROCEDURE MovBxBp4 ; BEGIN B (8BH) ; B (5EH) ; B (4) END MovBxBp4 ;  (* BX:=[BP+4] *)
+PROCEDURE MovAlDh  ; BEGIN B (8AH) ; B (0C6H) END MovAlDh ;  (* 8A C6: AL:=DH *)
 PROCEDURE MovDlSi  ; BEGIN B (8AH) ; B (14H) END MovDlSi ;
-PROCEDURE MovDlSp  ; BEGIN B (8AH) ; B (54H) ; B (24H) ; B (0) END MovDlSp ;
+PROCEDURE MovDlArg ; BEGIN B (8AH) ; B (56H) ; B (2) END MovDlArg ;  (* DL:=[BP+2] *)
 
-PROCEDURE StDiAx   ; BEGIN B (89H) ; B (7H) END StDiAx ;    (* [DI] := AX *)
-PROCEDURE StDiBx   ; BEGIN B (89H) ; B (1FH) END StDiBx ;   (* [DI] := BX *)
-PROCEDURE StBxCx   ; BEGIN B (89H) ; B (8BH) END StBxCx ;   (* [BX] := CX *)
-PROCEDURE StSiDl   ; BEGIN B (88H) ; B (14H) END StSiDl ;   (* [SI] := DL *)
-PROCEDURE StBxDl   ; BEGIN B (88H) ; B (93H) END StBxDl ;   (* [BX] := DL *)
+PROCEDURE StDiAx   ; BEGIN B (89H) ; B (5H) END StDiAx ;    (* 89 05: [DI]:=AX *)
+PROCEDURE StDiBx   ; BEGIN B (89H) ; B (1DH) END StDiBx ;   (* 89 1D: [DI]:=BX *)
+PROCEDURE StBxCx   ; BEGIN B (89H) ; B (0FH) END StBxCx ;   (* 89 0F: [BX]:=CX *)
+PROCEDURE StSiDl   ; BEGIN B (88H) ; B (14H) END StSiDl ;   (* 88 14: [SI]:=DL *)
+PROCEDURE StBxDl   ; BEGIN B (88H) ; B (17H) END StBxDl ;   (* 88 17: [BX]:=DL *)
 PROCEDURE MovDhAl  ; BEGIN B (88H) ; B (0C6H) END MovDhAl ;
 PROCEDURE MovDlAl  ; BEGIN B (88H) ; B (0C2H) END MovDlAl ;
-PROCEDURE MovSiBx  ; BEGIN B (89H) ; B (0DCH) END MovSiBx ;
+PROCEDURE MovSiBx  ; BEGIN B (89H) ; B (0DEH) END MovSiBx ;  (* 89 DE: MOV SI,BX *)
 PROCEDURE MovDiDx  ; BEGIN B (89H) ; B (0D7H) END MovDiDx ;
 PROCEDURE MovDiAx  ; BEGIN B (89H) ; B (0C7H) END MovDiAx ;
 PROCEDURE AddDiAx  ; BEGIN B (1H) ; B (0C7H) END AddDiAx ;
@@ -294,12 +373,15 @@ BEGIN
    M ("initmem") ;
    MovSiAx ;               (* SI = AX = the header offset the caller passed *)
    MovDxSi4 ;              (* DX = [SI+4] = hdrDS  = data base *)
-   MovCxSi8 ;              (* CX = [SI+8] = hdrHeap = data end   *)
+   MovCxSi6 ;              (* CX = [SI+6] = hdrHeap = data end   *)
    CmpCxDx ;
    Jbe8 ("im_done") ;
    MovDiDx ;               (* DI = data base *)
    M ("im_zero") ;
-   MovAxDx ;
+   XorAxAx ;               (* AX = 0: the value written into every global.
+                             The caller passes the header offset in AX, so
+                             without this the "zeroing" loop would write the
+                             header offset into all of them. *)
    StDiAx ;
    AddDi2 ;
    CmpDiCx ;
@@ -380,10 +462,14 @@ BEGIN
 END EmitWrInt ;
 
 PROCEDURE EmitWrChar ;
-(* the low byte of the pushed word *)
+(* the low byte of the one 16-bit argument, which sits above the return
+   address.  There is no [SP] addressing in 16-bit mode, so BP stands in for
+   the stack pointer and is handed straight back before the RET. *)
 BEGIN
    M ("wrchar") ;
-   MovDlSp ;
+   MovBpSp ;
+   MovDlArg ;
+   MovSpBp ;
    MovAh (2) ;
    Int21 ;
    RetR
@@ -427,7 +513,7 @@ END EmitWrInl ;
 PROCEDURE EmitWrBool ;
 BEGIN
    M ("wrbool") ;
-   CmpSpW0 ;
+   CmpArgW0 ;
    Jne8 ("wb_t") ;
    MovDxD (D_FALSE) ;
    J8 ("wb_o") ;
@@ -674,4 +760,12 @@ BEGIN
    RETURN LblOff (entNm [i])
 END RT_Entry ;
 
+PROCEDURE RT_CodeEnd () : CARDINAL ;
+BEGIN
+   IF NOT built THEN
+      RT_Build ()
+   END ;
+   RETURN dataAt
+END RT_CodeEnd ;
+
 END Runtime.


+ 1 - 1
shell/tests/CompileTest.mod

@@ -266,7 +266,7 @@ END DumpCode ;
 
 PROCEDURE DumpImage ;
 (* The whole linked image - [runtime][program header][program code] - which is
-   what a .COM would actually contain.  The runtime's own 385 bytes are shown
+   what a .COM would actually contain.  The runtime's own 391 bytes are shown
    only at the head and the tail: enough to prove the blob is really in there,
    without burying the program in 24 lines of library. *)
 VAR i, n, rtSz, total, from, lim : CARDINAL ;

+ 2 - 1
shell/tests/RtProbe.mod

@@ -5,7 +5,7 @@ MODULE RtProbe ;
    is wired to it. *)
 
 FROM Posix IMPORT write ;
-FROM Runtime IMPORT RT_Build, RT_Size, RT_Byte, RT_Entry ;
+FROM Runtime IMPORT RT_Build, RT_Size, RT_Byte, RT_Entry, RT_CodeEnd ;
 FROM SYSTEM IMPORT ADR, BYTE ;
 
 VAR
@@ -69,6 +69,7 @@ VAR i, n : CARDINAL ;
 BEGIN
    RT_Build () ;
    PCARD (RT_Size ()) ; PS (" bytes") ; NL ;
+   PS ("code ends at ") ; PCARD (RT_CodeEnd ()) ; NL ;
    i := 0 ;
    WHILE i <= 13 DO
       names [i] := "" ;

+ 344 - 0
shell/tests/audit_helpers.py

@@ -0,0 +1,344 @@
+#!/usr/bin/env python3
+"""audit_helpers.py -- check that Runtime.mod's one-liner emitter procedures
+emit the instruction their name claims.
+
+Why this exists
+---------------
+Runtime.mod hand-assembles 8086 by writing raw bytes.  Most of it is wrapped
+up as one-line helpers:
+
+    PROCEDURE MovSiBx ; BEGIN B (89H) ; B (0DCH) END MovSiBx ;
+
+and those names are the only documentation of what the bytes mean.  That is
+exactly the setup for a silent, expensive mistake: write the right opcode and
+the wrong ModRM, and nothing complains.  It happened twice here, in the same
+direction both times:
+
+    MovSiBx  emitted 89 DC   -- which is MOV SP,BX, not MOV SI,BX
+    CmpSiBx  emitted 39 DC   -- ditto, CMP SP,BX
+
+The comment next to CmpSiBx even spelled the ModRM out as "11 011 100" and
+still got it wrong, because 11 011 100 in mod=11 means rm=100=SP; SI is
+rm=100 only in mod=00, where it means [SI].  Both bugs shipped together into
+EmitWrInt, which then stored decimal digits through an SI register it had
+never initialised.  The structural checks in check_runtime.py could not see
+any of this: the bytes decoded cleanly, the sweep stayed in sync, the branch
+targets were all on boundaries and none of the entry prologues moved.  Only a
+*name* versus a *decode* comparison finds it, because only that knows what
+the author was trying to say.
+
+So: for every helper whose body is a literal byte list, disassemble those
+bytes with FCML and require the decode to match the name.  The name grammar
+is deliberately narrow and mechanical:
+
+    <Op><Dest><Src>        e.g. MovSiBx, StDiAx, CmpSpBx, MovDlSi
+    <Op><Reg>              e.g. XorAxAx, NegAx, NotDx, IncSi, DecSi
+
+with Op in {Mov, Lea, St, Ld, Cmp, Add, Sub, Xor, And, Or, Not, Inc, Dec,
+Push, Pop, Int, Jmp, Jne, Jge, Jnl, Jle, Jl, Je, Jae, Jbe, Jb, Ja, JaE...}
+and the register/operand words below.  A helper whose name does not parse is
+reported as UNPARSED rather than silently skipped -- a name the grammar does
+not understand is a name we are not checking, and that must be visible.
+
+This is a source-level check, so it needs no re-baselining: it is a function
+of the current source, and it gets stricter as more names are added to the
+grammar.  It also runs without building anything.
+
+Usage: audit_helpers.py [MODULE.mod]      # default ../Runtime.mod
+       audit_helpers.py -v                 # print every helper, passing or not
+"""
+
+import os
+import re
+import sys
+
+HERE = os.path.dirname(os.path.abspath(__file__))
+SHELL = os.path.dirname(HERE)
+sys.path.insert(0, HERE)
+import disasm16  # noqa: E402  (path set above)
+
+RE_HELPER = re.compile(
+    r"^PROCEDURE\s+(\w+)\s*;\s*BEGIN\s+(.*?)\s+END\s+\1\s*;", re.MULTILINE)
+# The H suffix is optional: the source mixes B (8AH) and B (0) for the same
+# kind of literal, and a byte written without H used to be silently dropped
+# from the audit, which made three helpers look like truncated prefixes.
+RE_BYTE = re.compile(r"B\s*\(\s*([0-9A-Fa-f]+)H?\s*\)")
+
+# --- name grammar -----------------------------------------------------
+#
+# An ordered table of (name regex, mnemonic, operand specs, allowed opcodes).
+#
+# Why the opcode column exists
+# ----------------------------
+# "MovDlSi" and "MovBpSp" are ambiguous from the name alone: SI and BP are
+# both in the register list and in the memory-base list.  What settles it is
+# the opcode, and the rule is the asymmetry that makes 16-bit hand-assembly
+# so error-prone:
+#
+#     88 /r  MOV r/m8, r8    reg is the SOURCE     (store)
+#     89 /r  MOV r/m16, r16  reg is the SOURCE     (store)
+#     8A /r  MOV r8, r/m8    reg is the DESTINATION (load)
+#     8B /r  MOV r16, r/m16  reg is the DESTINATION (load)
+#
+# A byte move (8A/88) to or from a bare base register can only be a memory
+# access, because "mov dl, si" is not an instruction.  A word move (8B/89)
+# could be either, so for word moves the register reading is tried first.
+# Encoding that rule in the table, rather than in pattern order alone, is what
+# stops a helper being "verified" against the wrong reading of its own name.
+#
+# Operand spec tokens:
+#
+#     "r:Name"     a bare register spelled Name (Ax, Al, Dx, Si, Ds, ...)
+#     "i:N"        the immediate N; the name spells it in DECIMAL
+#     "ih:N"       like i: but the name spells it in HEX (only Int, whose
+#                  vector 21h FCML prints as "21h" and which reads as decimal
+#                  33 if you do not notice)
+#     "m:Base"     memory at [Base], with no displacement
+#     "m:Base+D"   memory at [Base + D]
+#
+# "Arg" is the emitter's alias for BP: on entry BP points at the return
+# address, so a word argument starts at [BP+2].  Naming it "Arg" rather than
+# "Bp" is what stops these being misread as [BP] accesses, and it is also
+# how the grammar knows to expect the +2.
+#
+# A name matching NO pattern is reported UNPARSED, which counts as a failure
+# on purpose: a helper the audit cannot read is a helper nobody is checking.
+
+REGS = ["Ax", "Bx", "Cx", "Dx", "Si", "Di", "Bp", "Sp",
+        "Al", "Bl", "Cl", "Dl", "Ah", "Bh", "Ch", "Dh"]
+# Segment registers: the 8086 can only PUSH/POP them, never MOV to or from
+# one, so they need names of their own.
+SEGS = ["Ds", "Es", "Cs", "Ss"]
+JCCS = ["Je", "Jne", "Jz", "Jnz", "Jge", "Jnl", "Jle", "Jl", "Ja", "Jae",
+        "Jb", "Jbe", "Jg", "Jns", "Js", "Jo", "Jno", "Jp", "Jnp", "Jcxz",
+        "Jecxz", "Jrcxz", "Loop", "Loope", "Loopne"]
+
+_ALT = "|".join(REGS)
+_BASE = "Si|Di|Bx|Bp|Sp|Arg|Data"
+_STORE = (0x88, 0x89)      # reg is the source
+_LOAD = (0x8A, 0x8B)       # reg is the destination
+
+PATTERNS = [
+    # --- no-operand and fixed-operand forms -------------------------
+    (re.compile(r"^RetR?$"),            "ret",    [], None),
+    (re.compile(r"^LeaveR?$"),          "leave",  [], None),
+    (re.compile(r"^Int([0-9A-Fa-f]+)$"), "int",   ["ih:%(1)s"], {0xCD}),
+    (re.compile(r"^(Push|Pop)(%s)$" % "|".join(SEGS)),
+     None, ["r:%(2)s"], None),
+    (re.compile(r"^(Push|Pop)(%s)$" % _ALT), None, ["r:%(2)s"], None),
+    # --- jumps: the target is a fixup, so only the mnemonic is checked
+    (re.compile(r"^(%s)(8|16)?$" % "|".join(JCCS)), None, [], None),
+    (re.compile(r"^Jmp(%s)$" % _ALT),   "jmp",    ["r:%(1)s"], {0xFF}),
+    # --- MUL names one operand; 16-bit DIV names two because it always
+    #     divides DX:AX, so the decode has an AX the name has no room for
+    (re.compile(r"^Mul(%s)$" % _ALT),   "mul",    ["r:%(1)s"], {0xF7}),
+    (re.compile(r"^Div(%s)$" % _ALT),   "div",    ["r:Ax", "r:%(1)s"], {0xF7}),
+    # --- store: name is St<base><reg>, decode puts the register last --
+    (re.compile(r"^St(%s)(%s)$" % (_BASE, _ALT)),
+     "mov", ["m:%(1)s", "r:%(2)s"], _STORE),
+    # --- load a register from memory, with a displacement ------------
+    # The displacement makes the name unambiguous, so any load opcode works.
+    (re.compile(r"^Mov(%s)(%s)(\d+)$" % (_ALT, _BASE)),
+     "mov", ["r:%(1)s", "m:%(2)s+%(3)s"], _LOAD),
+    # --- two registers: MOV BP,SP / CMP SI,BX / XOR AX,AX / ADD DI,AX -
+    # Tried before the bare-base reading below, because for a word move
+    # "MovBpSp" means BP := SP and only "MovDlSi" means DL := [SI].  The two
+    # readings cannot both match, because one demands a register operand
+    # where the other demands a memory operand.
+    (re.compile(r"^(Mov|Cmp|Add|Sub|Xor|And|Or|Xchg)(%s)(%s)$" % (_ALT, _ALT)),
+     None, ["r:%(2)s", "r:%(3)s"], None),
+    # --- byte move from a bare base register: necessarily memory ------
+    # Restricted to 8A/88 on purpose: if this ever matched an 8B/89 it would
+    # be a register move already claimed by the pattern above.
+    (re.compile(r"^Mov(%s)(%s)$" % (_ALT, _BASE)),
+     "mov", ["r:%(1)s", "m:%(2)s"], (0x8A, 0x88)),
+    # --- immediate against a register --------------------------------
+    (re.compile(r"^(Add|Sub|Cmp|Xor|And|Or)(%s)(\d+)$" % _ALT),
+     None, ["r:%(2)s", "i:%(3)s"], {0x83}),
+    # --- one register: INC CX / DEC SI / NOT DX / NEG AX ------------
+    (re.compile(r"^(Inc|Dec|Not|Neg|Shl|Shr|Sar)(%s)$" % _ALT),
+     None, ["r:%(2)s"], None),
+]
+
+
+def register_word(tok):
+    """Map an FCML operand token to a REGS/SEGS word, or None."""
+    t = tok.strip().lower()
+    if not t or t.startswith("word ptr ") or t.startswith("byte ptr "):
+        return None
+    t = t.split()[-1]
+    for r in REGS + SEGS:
+        if t == r.lower():
+            return r
+    return None
+
+
+def match_operand(spec, tok):
+    """Does one FCML operand token satisfy one token spec?  Returns None on
+    success, or a human-readable reason on failure."""
+    kind, arg = spec.split(":", 1)
+    t = tok.strip()
+    if kind in ("i", "ih"):
+        # FCML prints an immediate as hex with an h suffix ("0h", "2h", "21h")
+        raw = t.lower().rstrip("h")
+        try:
+            v = int(raw, 16)
+        except ValueError:
+            return "%r is not an immediate" % t
+        want = int(arg, 16) if kind == "ih" else int(arg, 10)
+        if v != want:
+            return "immediate is %d, name says %d" % (v, want)
+        return None
+    if kind == "r":
+        got = register_word(t)
+        if got != arg:
+            return "%r is not the register %s" % (t, arg)
+        return None
+    if kind == "m":
+        base, _, disp = arg.partition("+")
+        real = {"Arg": "bp", "Data": "si"}.get(base, base.lower())
+        if base == "Arg" and not disp:
+            disp = "2"       # the first word argument lives at [BP+2]
+        u = t.lower()
+        u = re.sub(r"^(word|byte) ptr ", "", u)
+        if not (u.startswith("[") and u.endswith("]")):
+            return "%r is not a memory reference" % t
+        body = u[1:-1]
+        if not body.startswith(real):
+            return "memory base is %r, name says [%s]" % (body, real)
+        if not disp:
+            if body != real:
+                return ("memory is %r, name says [%s] with no displacement"
+                        % (body, real))
+            return None
+        got = body[len(real):].strip()
+        m = re.fullmatch(r"\+\s*([0-9a-f]+)h?", got)
+        if not m:
+            return "displacement %r is not a number" % got
+        if int(m.group(1), 16) != int(disp):
+            return "displacement is %d, name says %s" \
+                   % (int(m.group(1), 16), disp)
+        return None
+    return "bad spec %r" % spec
+
+
+def operands_of(decode):
+    """Split an FCML decode like "mov si,bx" or "cmp word ptr [bp+4h],0h"
+    into operand tokens.  Returns (mnemonic, [tokens])."""
+    i = decode.find(" ")
+    if i < 0:
+        return decode.strip(), []
+    return decode[:i].strip(), decode[i + 1:].split(",")
+
+
+def main(argv):
+    verbose = "-v" in argv
+    files = [a for a in argv[1:] if not a.startswith("-")]
+    path = files[0] if files else os.path.join(SHELL, "Runtime.mod")
+    with open(path) as f:
+        src = f.read()
+
+    n_ok = n_unparsed = n_bad = 0
+    problems = []
+    if verbose:
+        print("name                 bytes            FCML decode")
+
+    for m in RE_HELPER.finditer(src):
+        name, body = m.group(1), m.group(2)
+        by = [int(b, 16) for b in RE_BYTE.findall(body)]
+        if not by:
+            continue
+        raw = bytes(by)
+        # Disassemble; helpers are 1-2 bytes so one instruction is the norm,
+        # but StCx / Shl / Div style helpers can be longer, so decode all.
+        decodes = []
+        off = 0
+        while off < len(raw):
+            text, length = disasm16.decode(raw[off:], off)
+            if length == 0:
+                decodes.append(("<DECODE ERROR>", ""))
+                break
+            decodes.append(operands_of(text or "?"))
+            off += length
+
+        hexs = " ".join("%02X" % b for b in raw)
+        dec_text = "; ".join(d[0] + " " + ", ".join(d[1]) for d in decodes)
+        if verbose:
+            print("%-20s %-16s %s" % (name, hexs, dec_text))
+
+        # A name can have more than one reading (MovBpSp: BP := SP, or
+        # BP := [SP]).  Take the first reading whose operand specs the decode
+        # actually satisfies; if two readings both fit, the name is genuinely
+        # ambiguous and that is itself reported, because it would mean the
+        # audit could be satisfied by the wrong one.
+        reads = []
+        nopattern = True
+        for rx, pmnem, specs, ops in PATTERNS:
+            mm = rx.match(name)
+            if not mm:
+                continue
+            if ops is not None and by[0] not in ops:
+                continue
+            nopattern = False
+            gmap = {str(i): g for i, g in enumerate(mm.groups(), 1)}
+            mn = pmnem
+            if mn is None:
+                mn = mm.group(1).lower()
+            elif "%(" in mn:
+                mn = mn % gmap
+            if len(decodes) != 1:
+                reads.append((mm, mn, specs,
+                              "emits %d instructions, name describes one"
+                              % len(decodes)))
+                continue
+            gmnem, got = decodes[0]
+            if gmnem != mn:
+                reads.append((mm, mn, specs,
+                              "FCML decodes %r, name says %s" % (gmnem, mn)))
+                continue
+            if len(got) != len(specs):
+                reads.append((mm, mn, specs,
+                              "name asserts %d operand(s), FCML decoded %d"
+                              % (len(specs), len(got))))
+                continue
+            why = [w for w in
+                   (match_operand(sp % gmap if "%(" in sp else sp, tok)
+                    for sp, tok in zip(specs, got)) if w]
+            if why:
+                reads.append((mm, mn, specs, "; ".join(why)))
+                continue
+            reads.append((mm, mn, specs, None))
+
+        if nopattern:
+            n_unparsed += 1
+            problems.append("%s: no name pattern accepts it (emits %s %r) - "
+                            "extend PATTERNS or fix the name"
+                            % (name, hexs, dec_text))
+            continue
+        fitting = [r for r in reads if r[3] is None]
+        if len(fitting) > 1:
+            n_bad += 1
+            problems.append("%s: %d name readings fit the same bytes - the "
+                            "name is ambiguous" % (name, len(fitting)))
+            continue
+        if not fitting:
+            n_bad += 1
+            problems.append("%s: %s" % (name, reads[0][3]))
+            continue
+        n_ok += 1
+
+    total = n_ok + n_unparsed + n_bad
+    print("\n%d one-line emitter helpers audited: %d agree with their name, "
+          "%d disagree, %d outside the grammar"
+          % (total, n_ok, n_bad, n_unparsed))
+    if problems:
+        print("FAIL: %d problem(s)" % len(problems))
+        for p in problems:
+            print("  - %s" % p)
+        return 1
+    print("PASS: every audited helper emits what its name says")
+    return 0
+
+
+if __name__ == "__main__":
+    sys.exit(main(sys.argv))

+ 290 - 0
shell/tests/check_runtime.py

@@ -0,0 +1,290 @@
+#!/usr/bin/env python3
+"""check_runtime.py -- structural check of the 8086 runtime blob.
+
+The runtime is hand-assembled byte by byte in Runtime.mod, so the only way
+to know it is right is to look at the bytes.  This sweeps the code region
+with FCML (see disasm16.py) and asserts the properties that a wrong emitter
+cannot produce:
+
+  1. every instruction decodes, and the decoded lengths tile the code region
+     exactly - a single byte-length mistake desynchronises the sweep, so a
+     clean sweep means the lengths are all right;
+  2. every RT_Entry offset is an instruction boundary;
+  3. every relative branch and CALL target is an instruction boundary inside
+     the code region;
+  4. the full disassembly matches tests/runtime.golden byte for byte;
+  5. each entry still begins with the instruction sequence it is supposed to
+     begin with (semantic goldens, not positional ones, so they survive
+     legitimate size changes but catch a wrong ModRM byte).
+
+Points 1-3 exist because they give readable diagnostics for the two most
+common classes of mistake (wrong instruction length, wrong fixup).  They are
+NOT sufficient on their own - they all pass on code that decodes cleanly but
+means the wrong thing.  Concretely: with MovAlDh emitting `8A C0`
+(`mov al,al`) instead of `8A C6` (`mov al,dh`), the sweep stayed in sync, every
+branch target stayed on a boundary, and no entry's first bytes changed, so all
+three checks passed on provably broken code.  That is why check 4 exists: it
+catches any wrong-but-well-formed instruction, not just the ones that happen
+to break the structure.
+
+Re-baselining the golden
+------------------------
+Check 4 fails on *every* legitimate edit to Runtime.mod, because inserting an
+instruction moves every later offset.  That is intended: the golden forces a
+human to look at the whole new disassembly and say "yes, that is what I meant".
+
+    python3 tests/check_runtime.py --bless     # rewrite tests/runtime.golden
+    git diff tests/runtime.golden              # READ THE DIFF, then commit
+
+Do not use --bless to silence a red test.  Re-baseline only after reading the
+diff and confirming the change is the one you intended; write the reason in the
+commit message.  An fcml upgrade may also change instruction *wording* (not
+meaning) and require a re-baseline for the same reason.
+
+Usage: check_runtime.py [--bless] [-v] [RTPROBE-DUMP]
+       (default: build RtProbe and dump it)
+"""
+
+import os
+import re
+import subprocess
+import sys
+
+HERE = os.path.dirname(os.path.abspath(__file__))
+SHELL = os.path.dirname(HERE)
+GM2 = "/home/eric/bin/Modula2/Gm2/bin/gm2"
+RTPROBE = "/tmp/tp_check_rtprobe"
+GOLDEN_FILE = os.path.join(HERE, "runtime.golden")
+
+RE_SIZE = re.compile(r"^(\d+) bytes$")
+RE_CODEEND = re.compile(r"^code ends at (\d+)$")
+RE_ENTRY = re.compile(r"^  entry (\d+) = (\d+)\s+\((\w+)\)$")
+RE_BRANCH = re.compile(r"^(?:j\w+|call|jmp)\s+([0-9a-f]+)h$")
+
+GOLDEN_TEXT_HEADER = """\
+# runtime.golden -- full FCML disassembly of the runtime's code region.
+#
+# GENERATED by tests/check_runtime.py --bless, then READ THE DIFF and commit.
+# Every line is "OFFSET  BYTES  MNEMONIC".  This file is a deliberate
+# tripwire: any edit to Runtime.mod invalidates it, because inserting an
+# instruction renumbers everything after it, so you are forced to look at the
+# whole new listing and confirm you meant every change.  See the "Re-baselining
+# the golden" section of check_runtime.py.
+#
+# Region: bytes 0..codeEnd of the runtime image (see the code/data split in
+# EmitData).  Relative targets are absolute offsets within the runtime, which
+# the compiler adds to the runtime's load address when it emits a CALL.
+"""
+
+# The "mod=11" column is reg-only, so reg is the second ModRM field: for
+# 88/8A (byte moves) ModRM = 11 010 rrr, for 89/8B (word moves) 11 001 rrr.
+
+
+def describe_golden_diff(have, want):
+    """One-line summary of how the current disassembly differs from the
+    golden, naming the first few differing instructions so the report points
+    at the actual mistake instead of just saying 'mismatch'."""
+    a = have.splitlines()
+    b = want.splitlines()
+    out = []
+    for i in range(max(len(a), len(b))):
+        x = a[i] if i < len(a) else "<end of golden>"
+        y = b[i] if i < len(b) else "<end of disassembly>"
+        if x != y:
+            out.append("line %d: golden has %r, now %r" % (i + 1, x, y))
+            if len(out) == 3:
+                break
+    out.append("%d differing line(s) total" % sum(
+        1 for i in range(max(len(a), len(b)))
+        if (a[i] if i < len(a) else None) != (b[i] if i < len(b) else None)))
+    return "; ".join(out)
+
+# Semantic goldens: the byte sequence each entry must START with.  Stated in
+# terms of intent ("read the argument through BP") rather than offsets, so
+# they survive a legitimate size change but still catch a wrong ModRM byte.
+# This list is the *explanation* for the bytes; runtime.golden is the
+# backstop.  Keep the two consistent -- they read the same measured ModRM
+# table that audit_helpers.py enforces on the source.
+GOLDEN = {
+    "initmem": "8B F0 8B 54 04 8B 4C 06",   # SI:=AX(header); DX:=[SI+4]; CX:=[SI+6]
+    "stackchk": "C3",                       # the no-op: a bare RET, by design
+    "progend": "31 C0 B4 4C CD 21 C3",      # XOR AX,AX; AH:=$4C; INT 21h; RET
+    "wrint": "55 8B EC 8B 46 04",           # PUSH BP; MOV BP,SP; AX:=[BP+4]
+    "wrchar": "8B EC 8A 56 02 89 EC",       # BP:=SP; DL:=[BP+2]; SP:=BP
+    "wrbool": "8B EC 83 7E 02 00 89 EC",    # BP:=SP; CMP [BP+2],0; SP:=BP
+    "wrtinl": "5B 31 C9 8A 0F 43 B4 02",    # POP BX; XOR CX,CX; CL:=[BX]; INC BX
+    "rdln": "50",                           # PUSH AX, to keep the caller's
+}
+
+
+def build_and_dump():
+    subprocess.run([GM2, "-fiso", "-Wall", "-c", "Runtime.mod"],
+                   cwd=SHELL, check=True)
+    subprocess.run([GM2, "-fiso", "-o", RTPROBE,
+                    os.path.join("tests", "RtProbe.mod"),
+                    "Runtime.o", "Posix.o"],
+                   cwd=SHELL, check=True)
+    out = subprocess.run([RTPROBE], capture_output=True, text=True, check=True)
+    return out.stdout
+
+
+def parse(dump):
+    size = codeend = None
+    entries = {}
+    code = bytearray()
+    inhex = False
+    for line in dump.splitlines():
+        if RE_SIZE.match(line):
+            size = int(RE_SIZE.match(line).group(1))
+            continue
+        m = RE_CODEEND.match(line)
+        if m:
+            codeend = int(m.group(1))
+            continue
+        m = RE_ENTRY.match(line)
+        if m:
+            entries[m.group(3)] = int(m.group(2))
+            continue
+        if line.strip() == "hex:":
+            inhex = True
+            continue
+        if inhex:
+            parts = line.split()
+            if not parts or not re.fullmatch(r"[0-9A-F]{8}", parts[0]):
+                continue
+            for p in parts[1:]:
+                code.append(int(p, 16))
+    return size, codeend, entries, bytes(code)
+
+
+def main(argv):
+    bless = "--bless" in argv
+    verbose = "-v" in argv
+    files = [a for a in argv[1:] if not a.startswith("-")]
+    if len(files) > 1:
+        print("usage: check_runtime.py [--bless] [-v] [RTPROBE-DUMP]")
+        return 2
+    if files:
+        with open(files[0]) as f:
+            dump = f.read()
+    else:
+        dump = build_and_dump()
+
+    size, codeend, entries, code = parse(dump)
+    problems = []
+
+    if size is None or codeend is None:
+        print("FAIL: could not parse the runtime dump (size/codeend missing)")
+        return 1
+    if len(code) != size:
+        problems.append("hex dump has %d bytes, header says %d"
+                        % (len(code), size))
+    if not entries:
+        problems.append("no entry offsets parsed")
+
+    sys.path.insert(0, HERE)
+    import disasm16
+
+    # --- 1. clean sweep of the code region ---------------------------
+    body = code[:codeend]
+    starts = set()
+    listing = []          # (pc, bytes, text) - kept structured so neither the
+                          # branch scan nor the golden has to re-parse the
+                          # pretty-printed line
+    pc = 0
+    import io
+    buf = io.StringIO()
+    while pc < len(body):
+        starts.add(pc)
+        text, length = disasm16.decode(body[pc:], pc)
+        if length == 0:
+            buf.write("%04X: %s  <DECODE ERROR>\n"
+                      % (pc, " ".join("%02X" % b for b in body[pc:pc + 8])))
+            problems.append("decode error at %04X" % pc)
+            break
+        raw = body[pc:pc + length]
+        text = text or "?"
+        listing.append((pc, raw, text))
+        buf.write("%04X: %-24s %s\n"
+                  % (pc, " ".join("%02X" % b for b in raw), text))
+        pc += length
+    if pc != len(body):
+        problems.append("sweep ended at %04X, code region ends at %04X"
+                        % (pc, len(body)))
+
+    if verbose:
+        sys.stderr.write(buf.getvalue())
+
+    # --- 2. entries are instruction boundaries -----------------------
+    for name, off in sorted(entries.items()):
+        if off not in starts:
+            problems.append("entry %s at %d is not an instruction boundary"
+                            % (name, off))
+
+    # --- 3. branch targets are instruction boundaries -----------------
+    nbranch = 0
+    for _at, _raw, text in listing:
+        m = RE_BRANCH.match(text)
+        if not m:
+            continue
+        nbranch += 1
+        tgt = int(m.group(1), 16)
+        if tgt not in starts:
+            problems.append("%r targets %04X, not an instruction boundary"
+                            % (text, tgt))
+        elif tgt >= codeend:
+            problems.append("%r targets %04X, outside the code region"
+                            % (text, tgt))
+
+    # --- 4. full-disassembly golden -----------------------------------
+    # This is the only check that catches an instruction which decodes
+    # cleanly but means the wrong thing; see the module docstring for the
+    # MovAlDh example that motivated it.
+    golden_text = "".join(
+        "%04X  %-11s %s\n" % (at, " ".join("%02X" % b for b in raw), text)
+        for at, raw, text in listing)
+    if bless:
+        with open(GOLDEN_FILE, "w") as f:
+            f.write(GOLDEN_TEXT_HEADER)
+            f.write(golden_text)
+        print("BLESSED: wrote %s (%d instructions).  Read the diff before"
+              " committing." % (GOLDEN_FILE, len(listing)))
+    else:
+        if not os.path.exists(GOLDEN_FILE):
+            problems.append("no golden at %s - run with --bless once and"
+                            " commit the result" % GOLDEN_FILE)
+        else:
+            with open(GOLDEN_FILE) as f:
+                have = f.read()
+            want = GOLDEN_TEXT_HEADER + golden_text
+            if have != want:
+                problems.append("disassembly does not match %s (%s)"
+                                % (os.path.basename(GOLDEN_FILE),
+                                   describe_golden_diff(have, want)))
+
+    # --- 5. semantic goldens -----------------------------------------
+    for name, want in sorted(GOLDEN.items()):
+        if name not in entries:
+            problems.append("no entry named %s" % name)
+            continue
+        off = entries[name]
+        got = " ".join("%02X" % b for b in code[off:off + len(want.split())])
+        if got.upper() != want.upper():
+            problems.append("entry %s starts %s, expected %s"
+                            % (name, got.upper(), want.upper()))
+
+    print("runtime: %d bytes, code 0..%d (%d), %d entries, %d branches checked"
+          % (size, codeend - 1, codeend, len(entries), nbranch))
+    for name in sorted(entries):
+        print("  %-9s at %4d" % (name, entries[name]))
+    if problems:
+        print("FAIL: %d problem(s)" % len(problems))
+        for p in problems:
+            print("  - %s" % p)
+        return 1
+    print("PASS: runtime code region is self-consistent")
+    return 0
+
+
+if __name__ == "__main__":
+    sys.exit(main(sys.argv))

+ 22 - 6
shell/tests/comtest.py

@@ -35,9 +35,20 @@ FIXTURE = os.path.abspath(sys.argv[1]) if len(sys.argv) > 1 else os.path.join(
 COM = os.path.splitext(FIXTURE)[0] + ".COM"
 
 # Restated here on purpose - the checker must not ask the code under test.
-RTSZ = 385                    # Runtime.RT_Size()
+# RTSZ was re-baselined 385 -> 391 together with the runtime bug fixes; see
+# the rationale in tests/run_com_tests.sh and tests/runtime.golden.  HEAD is
+# initmem's prologue, whose two SI displacements are the header words this
+# checker goes on to verify, so the two are tied together by assertion rather
+# than by coincidence (that is how initmem came to read hdrMax instead of
+# hdrHeap and silently zero nothing).
+RTSZ = 391                    # Runtime.RT_Size()
 DATAB = RTSZ + 0x1000         # Compiler: data base = rtSz + 1000H
-HEAD = "8B C0 8B 54 04 8B 4C 08"   # initmem: MOV AX,AX / MOV DX,[SI+4] / MOV CX,[SI+8]
+HEAD = "8B F0 8B 54 04 8B 4C 06"   # MOV SI,AX / MOV DX,[SI+4] / MOV CX,[SI+6]
+HDR_DS_WORD = 4               # header word holding the data base
+HDR_HEAP_WORD = 6             # header word holding the data end
+assert [int(HEAD.split()[4], 16), int(HEAD.split()[7], 16)] == \
+       [HDR_DS_WORD, HDR_HEAP_WORD], \
+       "initmem no longer reads the two header words this checker verifies"
 
 
 def check_com(path, expected_src_len):
@@ -46,8 +57,9 @@ def check_com(path, expected_src_len):
     if not os.path.exists(path):
         return ["no .COM file was written"]
     d = open(path, "rb").read()
-    if d[:8].hex(" ").upper() != HEAD:
-        errs.append("runtime not at offset 0: first bytes %s" % d[:8].hex(" ").upper())
+    if d[:len(HEAD.split())].hex(" ").upper() != HEAD:
+        errs.append("runtime not at offset 0: first bytes %s, want %s"
+                    % (d[:len(HEAD.split())].hex(" ").upper(), HEAD))
     if len(d) < DATAB:
         errs.append("file is %d bytes, shorter than the data base %d - the "
                     "globals would be outside the file" % (len(d), DATAB))
@@ -58,12 +70,16 @@ def check_com(path, expected_src_len):
         if flag != 1:
             errs.append("hdrFlag=%d at the program header, want 1" % flag)
     if len(d) >= RTSZ + 10:
-        ds = int.from_bytes(d[RTSZ + 4:RTSZ + 6], "little")
-        heap = int.from_bytes(d[RTSZ + 6:RTSZ + 8], "little")
+        ds = int.from_bytes(d[RTSZ + HDR_DS_WORD:RTSZ + HDR_DS_WORD + 2], "little")
+        heap = int.from_bytes(d[RTSZ + HDR_HEAP_WORD:RTSZ + HDR_HEAP_WORD + 2], "little")
         if ds != DATAB:
             errs.append("hdrDS=%d, want %d" % (ds, DATAB))
         if heap < ds:
             errs.append("hdrHeap=%d < hdrDS=%d" % (heap, ds))
+        if [d[4], d[7]] != [HDR_DS_WORD, HDR_HEAP_WORD]:
+            errs.append("initmem reads header words +%d/+%d, but the data base "
+                        "and data end are at +%d/+%d"
+                        % (d[4], d[7], HDR_DS_WORD, HDR_HEAP_WORD))
     # the gap between the end of the code and the data area must be all zero
     cs = int.from_bytes(d[RTSZ + 2:RTSZ + 4], "little") if len(d) >= RTSZ + 4 else 0
     if cs and cs < DATAB:

+ 116 - 0
shell/tests/disasm16.py

@@ -0,0 +1,116 @@
+#!/usr/bin/env python3
+"""disasm16.py -- linear-sweep 16-bit x86 disassembler built on FCML.
+
+Why this exists
+---------------
+We need an *independent* check on the bytes our Modula-2 compiler and runtime
+emit.  The obvious candidates all turned out to be unusable:
+
+  * objdump / binutils:  has no 16-bit x86 disassembler at all.  `-m i8086`
+    silently falls back to the 32-bit i386 rules.
+  * unicorn 2.1.4:      UC_MODE_16 mis-decodes 16-bit ModRM memory operands.
+
+FCML (libfcml, Debian package `fcml`) does have a real 16-bit x86
+disassembler, and we validated it against GNU as's `.code16` *encoder*
+(assemble there, disassemble here).  See tests/FCML_PROBE.md.
+
+FCML is driven through the `fcml-disasm` CLI because the package ships no
+development headers, so ctypes struct layout would be guesswork.  The CLI
+happens to be a linear sweeper that reports "Instruction code length", which
+is all this driver needs.
+
+The length agreement is the important part: if we emit a 3-byte instruction
+where a 4-byte one was required, the sweep desynchronises and the very next
+line points straight at the offending offset.
+
+Usage:
+    disasm16.py FILE [BASE]        # BASE defaults to 0, decimal or 0x-hex
+    disasm16.py --bytes '89 6e 04' # inline bytes, for one-off probes
+"""
+
+import re
+import subprocess
+import sys
+
+FCML = "fcml-disasm"
+
+# The Debian 1.3.0 `fcml-disasm` wrapper aborts (SIGABRT, rc=134) on any
+# buffer of 16 bytes or more: it linearly sweeps the *whole* input and blows
+# up on the resulting instruction list.  Reproduce with 16 NOPs:
+#     fcml-disasm -m16 0x90909090909090909090909090909090   # rc=134
+# 15 bytes is safe, and 15 >= the longest real-mode instruction (7 bytes,
+# e.g. `EA off16 seg16`), so a 15-byte window always contains the whole first
+# instruction.  We only ever read the first instruction's length.
+MAX_WINDOW = 15
+
+RE_LEN = re.compile(r"^\s*Instruction code length:\s*(\d+)\s*$")
+RE_TEXT = re.compile(r"^\s*Disassembled instruction:\s*(.*?)\s*$")
+
+
+def decode(code, base):
+    """Decode the first instruction of `code` (bytes). Returns (text, length).
+
+    `code` is the remaining stream; at most MAX_WINDOW bytes are handed to
+    FCML.  No padding is invented: an under-length buffer must surface as a
+    decode error rather than being silently completed with made-up bytes.
+    """
+    win = code[:MAX_WINDOW]
+    hexs = "".join("%02X" % b for b in win)
+    out = subprocess.run(
+        [FCML, "-m16", "-rh", "-rz", "-ip", hex(base), "0x" + hexs],
+        capture_output=True, text=True)
+    if out.returncode != 0:
+        return None, 0
+    text = None
+    length = None
+    for line in out.stdout.splitlines():
+        m = RE_TEXT.match(line)
+        if m:
+            text = m.group(1)
+        m = RE_LEN.match(line)
+        if m:
+            length = int(m.group(1))
+    if length is None or length == 0 or length > len(code):
+        return text, 0
+    return text, length
+
+
+def disasm(code, base=0, out=sys.stdout):
+    """Linear-sweep `code` from `base`, printing one line per instruction."""
+    pc = 0
+    ninstr = 0
+    while pc < len(code):
+        text, length = decode(code[pc:], base + pc)
+        if length == 0:
+            out.write("%04X: %-24s <DECODE ERROR>\n"
+                      % (base + pc, " ".join("%02X" % b
+                                             for b in code[pc:pc + 8])))
+            return ninstr, False
+        shown = code[pc:pc + length]
+        out.write("%04X: %-24s %s\n"
+                  % (base + pc,
+                     " ".join("%02X" % b for b in shown),
+                     text if text else "<no text>"))
+        pc += length
+        ninstr += 1
+    return ninstr, True
+
+
+def main(argv):
+    if len(argv) < 2:
+        sys.stderr.write(__doc__)
+        return 2
+    if argv[1] == "--bytes":
+        code = bytes.fromhex(argv[2].replace(" ", ""))
+        base = int(argv[3], 0) if len(argv) > 3 else 0
+    else:
+        with open(argv[1], "rb") as f:
+            code = f.read()
+        base = int(argv[2], 0) if len(argv) > 2 else 0
+    ninstr, ok = disasm(code, base)
+    sys.stderr.write("-- %d instructions, %d bytes\n" % (ninstr, len(code)))
+    return 0 if ok else 1
+
+
+if __name__ == "__main__":
+    sys.exit(main(sys.argv))

+ 86 - 0
shell/tests/fcml_vs_objdump.py

@@ -0,0 +1,86 @@
+#!/usr/bin/env python3
+"""fcml_vs_objdump.py -- differential test between two 16-bit disassemblors.
+
+Compares the *length of the first instruction* over many byte windows.  A
+length disagreement is the sharpest possible signal that the two tools read
+the ModRM/displacement/SIB fields differently, because if the lengths agree
+the tools walked the same opcode table.
+
+Usage:  fcml_vs_objdump.py [trials] [--seed N]
+"""
+
+import os
+import random
+import re
+import subprocess
+import sys
+import tempfile
+
+RE_LEN = re.compile(r"^\s*Instruction code length:\s*(\d+)\s*$", re.M)
+RE_ADDR = re.compile(r"^\s*([0-9a-f]+):\t", re.M)
+
+
+def fcml_first_len(win):
+    hexs = "".join("%02X" % b for b in win)
+    out = subprocess.run(["fcml-disasm", "-m16", "-rh", "-rz", "0x" + hexs],
+                         capture_output=True, text=True)
+    if out.returncode != 0:
+        return None
+    m = RE_LEN.search(out.stdout)
+    return int(m.group(1)) if m else None
+
+
+def objdump_first_len(win, tmp):
+    with open(tmp, "wb") as f:
+        f.write(win)
+    out = subprocess.run(["objdump", "-D", "-b", "binary", "-m", "i8086", tmp],
+                         capture_output=True, text=True)
+    if out.returncode != 0:
+        return None
+    addrs = [int(a, 16) for a in RE_ADDR.findall(out.stdout)]
+    if not addrs:
+        return None
+    if len(addrs) == 1:
+        # single instruction covering the whole window
+        return len(win)
+    return addrs[1] - addrs[0]
+
+
+def main(argv):
+    trials = 3000
+    seed = 1
+    args = [a for a in argv[1:]]
+    if "--seed" in args:
+        i = args.index("--seed")
+        seed = int(args[i + 1])
+        del args[i:i + 2]
+    if args:
+        trials = int(args[0])
+    rnd = random.Random(seed)
+    tmp = os.path.join(tempfile.gettempdir(), "diffprobe.bin")
+    agree = 0
+    one_err = 0
+    diffs = []
+    for t in range(trials):
+        n = rnd.choice([4, 6, 8, 10, 12, 15])
+        win = bytes(rnd.randrange(256) for _ in range(n))
+        a = fcml_first_len(win)
+        b = objdump_first_len(win, tmp)
+        if a is None or b is None:
+            one_err += 1
+            continue
+        if a == b:
+            agree += 1
+        else:
+            diffs.append((win.hex(" ").upper(), a, b))
+    total = agree + len(diffs)
+    print("seed=%d trials=%d compared=%d agree=%d (%d skipped: one tool "
+          "rejected the input) disagreements=%d"
+          % (seed, trials, total, agree, one_err, len(diffs)))
+    for h, a, b in diffs[:25]:
+        print("  %-32s fcml=%-3s objdump=%-3s" % (h, a, b))
+    return 1 if diffs else 0
+
+
+if __name__ == "__main__":
+    sys.exit(main(sys.argv))

+ 215 - 0
shell/tests/nonvacuity.sh

@@ -0,0 +1,215 @@
+#!/bin/sh
+# nonvacuity.sh -- prove the runtime checks can actually fail.
+#
+# A test that has never been seen red is not a test.  This script breaks the
+# runtime on purpose, once per check, and asserts that the check goes red and
+# says something useful about the breakage.  Then it restores the source and
+# asserts everything is green again.
+#
+# Each mutation below is a real bug that was in this file at some point, not an
+# invented one.  That is the point: these are the mistakes we actually make
+# with 16-bit ModRM, so these are the ones the checks have to catch.
+#
+#   audit_helpers.py   name-versus-decode: catches a wrong ModRM that still
+#                      decodes cleanly
+#   check_runtime.py   golden:            catches the same thing in the built
+#                                           image
+#   check_runtime.py   decode sweep:      catches a wrong instruction LENGTH
+#   check_runtime.py   branch targets:    catches a wrong fixup
+#   check_runtime.py   entry goldens:     catches a broken prologue
+#   probe/modrm11.py   the mod=11 table:   catches the ModRM column itself
+#                      going wrong, which no amount of decoding will show
+#
+# The mod=11 cases do not need a rebuild -- they read the probe sources
+# directly -- so they are cheap, and they are the ones that matter most: the
+# table they guard is the one thing in this project that was wrong in the
+# documentation while the code was right, and a table that is wrong in the
+# code produces bytes that decode perfectly.
+#
+# Usage: tests/nonvacuity.sh        (from shell/; leaves Runtime.mod restored)
+
+set -u
+cd "$(dirname "$0")/.." || exit 1
+
+GM2=/home/eric/bin/Modula2/Gm2/bin/gm2
+SAVED=/tmp/opencode/nonvacuity.Runtime.mod
+PROBE=/tmp/opencode/nonvacuity.rtprobe
+DUMP=/tmp/opencode/nonvacuity.dump
+
+cp Runtime.mod "$SAVED" || exit 1
+trap 'cp "$SAVED" Runtime.mod; "$GM2" -fiso -c Runtime.mod >/dev/null 2>&1' EXIT
+
+pass=0
+fail=0
+
+# rebuild <label> -- re-emit the runtime and dump it
+rebuild () {
+    "$GM2" -fiso -c Runtime.mod >/dev/null 2>&1 || return 1
+    "$GM2" -fiso -o "$PROBE" tests/RtProbe.mod Runtime.o Posix.o \
+        >/dev/null 2>&1 || return 1
+    "$PROBE" > "$DUMP" || return 1
+    return 0
+}
+
+# expect_red <label> <pattern> <checker-cmd...>
+#   <pattern> is a grep the failure output must match, so a check cannot
+#   "pass" by failing for some unrelated reason.
+expect_red () {
+    label=$1
+    want=$2
+    shift 2
+    if out=$("$@" 2>&1); then
+        echo "NOT NON-VACUOUS: $label -- the check still passed"
+        fail=$((fail + 1))
+    elif ! printf '%s\n' "$out" | grep -qi "$want"; then
+        echo "WRONG FAILURE: $label -- went red, but not for the stated reason"
+        printf '%s\n' "$out" | sed 's/^/       /'
+        fail=$((fail + 1))
+    else
+        echo "  ok: $label"
+        printf '%s\n' "$out" | grep -im1 "$want" | sed 's/^/       /'
+        pass=$((pass + 1))
+    fi
+}
+
+echo "== each mutation must turn the named check red"
+echo
+
+# --- 1. name-versus-decode -------------------------------------------
+# MovSiBx was `89 DC`, which is MOV SP,BX.  Two bytes either way, decodes
+# cleanly, and no structural check can see it.
+cp "$SAVED" Runtime.mod
+sed -i 's|B (0DEH) END MovSiBx|B (0DCH) END MovSiBx|' Runtime.mod
+expect_red "audit_helpers catches MovSiBx emitting MOV SP,BX" \
+    "MovSiBx" python3 tests/audit_helpers.py
+
+# CmpSiBx had the identical mistake, which is how you know a single fix is
+# not enough -- the same misreading was written twice.
+cp "$SAVED" Runtime.mod
+sed -i 's|B (39H) ; B (0DEH) END CmpSiBx|B (39H) ; B (0DCH) END CmpSiBx|' Runtime.mod
+expect_red "audit_helpers catches CmpSiBx emitting CMP SP,BX" \
+    "CmpSiBx" python3 tests/audit_helpers.py
+
+# --- 2. golden, and entry goldens ------------------------------------
+# MovAlDh was `8A C0` = MOV AL,AL instead of MOV AL,DH.  This is the case
+# that motivated runtime.golden: the sweep stayed in sync, every branch
+# target stayed on a boundary, and no entry's first bytes moved.
+cp "$SAVED" Runtime.mod
+sed -i 's|B (0C6H) END MovAlDh|B (0C0H) END MovAlDh|' Runtime.mod
+rebuild
+expect_red "runtime.golden catches MOV AL,AL" \
+    "mov al,al" python3 tests/check_runtime.py "$DUMP"
+
+# initmem opened with the mis-emitted MovSiAx, so its entry golden was the
+# thing that noticed the prologue was a no-op.
+cp "$SAVED" Runtime.mod
+sed -i 's|B (0F0H) END MovSiAx|B (0C0H) END MovSiAx|' Runtime.mod
+rebuild
+expect_red "check_runtime catches a broken initmem prologue" \
+    "mov ax,ax" python3 tests/check_runtime.py "$DUMP"
+
+# --- 3. decode sweep / length ----------------------------------------
+# StBxDl was `88 97` = [BX],DL with mod=10, so the instruction needs a
+# disp16 and the sweep loses sync two bytes later.
+cp "$SAVED" Runtime.mod
+sed -i 's|B (88H) ; B (17H) END StBxDl|B (88H) ; B (97H) END StBxDl|' Runtime.mod
+rebuild
+expect_red "decode sweep catches a mod=10 byte move with no displacement" \
+    "585Bh" python3 tests/check_runtime.py "$DUMP"
+
+# --- 4. branch targets ------------------------------------------------
+# FixUp measures a rel8 from the end of the instruction, one byte past the
+# displacement field.  Drop the +1 and every short branch lands one byte into
+# its target, which for a 3-byte instruction means the middle of it.  The
+# bytes themselves are all perfectly well formed -- only the fixups are
+# wrong -- so this is the one failure mode the golden cannot be expected to
+# catch on its own.
+cp "$SAVED" Runtime.mod
+sed -i 's|rel := (t + 100H - (fix \[i\].place + 1)) MOD 100H|rel := (t + 100H - fix [i].place) MOD 100H|' \
+    Runtime.mod
+rebuild
+expect_red "branch check catches rel8 fixups measured from the wrong byte" \
+    "not an instruction boundary" \
+    python3 tests/check_runtime.py "$DUMP"
+
+echo
+echo "== everything restored and green again"
+cp "$SAVED" Runtime.mod
+if rebuild; then
+    if python3 tests/audit_helpers.py >/dev/null 2>&1 &&
+       python3 tests/check_runtime.py "$DUMP" >/dev/null 2>&1; then
+        echo "  ok: both checks pass on the restored source"
+        pass=$((pass + 1))
+    else
+        echo "NOT RESTORED: a check is red after restoring Runtime.mod"
+        fail=$((fail + 1))
+    fi
+else
+    echo "NOT RESTORED: the runtime would not rebuild"
+    fail=$((fail + 1))
+fi
+
+echo
+echo "== the mod=11 table (probe/modrm11.py)"
+# These mutate the probe's own sources, not the runtime, so there is no
+# rebuild in the loop.  SAVED_PY / SAVED_S are restored after each case.
+SAVED_PY=/tmp/opencode/nonvacuity.modrm11.py
+SAVED_S=/tmp/opencode/nonvacuity.modrm11.s
+cp tests/probe/modrm11.py "$SAVED_PY" || exit 1
+cp tests/probe/modrm11.s "$SAVED_S" || exit 1
+
+M11="python3 tests/probe/modrm11.py"
+restore_probe () {
+    cp "$SAVED_PY" tests/probe/modrm11.py
+    cp "$SAVED_S" tests/probe/modrm11.s
+}
+
+# 1. one cell of the table moved
+sed -i 's|"Si", "Di"\]$|"Bp", "Di"]|' tests/probe/modrm11.py
+expect_red "anchor pins a moved table cell" \
+    "anchor ADD SI, 2" $M11
+restore_probe
+
+# 2. the table this project actually shipped: AX dropped off the front and a
+#    duplicate BX invented at the end, which shifts every code down by one
+sed -i 's|^REG = .*$|REG = ["Cx", "Dx", "Bx", "Sp", "Bp", "Si", "Di", "Bx"]|' \
+    tests/probe/modrm11.py
+expect_red "the table shifted by one (AX dropped, BX duplicated)" \
+    "anchor MOV SP, BP" $M11
+restore_probe
+
+# 3. the .s edited to contradict the table.  This is the case that shows why
+#    the hard-coded EXPECT bytes exist: the assembler encodes the new claim
+#    correctly, so comparing the .s against `as` alone can never fail here.
+sed -i 's|movw    %sp, %di         # reg 100|movw    %bp, %di         # reg 100|' \
+    tests/probe/modrm11.s
+expect_red "probe source edited away from the recorded bytes" \
+    "expected 89 E7" $M11
+restore_probe
+
+# 4. the 8-bit list edited, which is a different table from the word one
+sed -i 's|movb    %al, %dl         # 88 C2  ->  DL := AL|movb    %al, %bl         # was DL|' \
+    tests/probe/modrm11.s
+expect_red "the 8-bit register list edited" \
+    "expected 88 C2" $M11
+restore_probe
+
+# 5. an anchor's recorded byte corrupted, so the anchor can no longer
+#    corroborate itself
+sed -i 's|"8B EC", "8B E5"|"8B ED", "8B E5"|' tests/probe/modrm11.py
+expect_red "anchor byte no longer matches the emitted code" \
+    "expected 8B ED" $M11
+restore_probe
+
+if $M11 >/dev/null 2>&1; then
+    echo "  ok: modrm11.py passes on the restored probe sources"
+    pass=$((pass + 1))
+else
+    echo "NOT RESTORED: modrm11.py is red after restoring its sources"
+    $M11 2>&1 | sed 's/^/       /'
+    fail=$((fail + 1))
+fi
+
+echo
+echo "non-vacuity: $pass ok, $fail failed"
+[ "$fail" -eq 0 ]

+ 156 - 0
shell/tests/probe/README.md

@@ -0,0 +1,156 @@
+# ModR/M probes
+
+The 16-bit ModR/M table in `../../Runtime.mod` is the input every hand-written
+emitter in that file depends on, and it was wrong once. The wrong version was
+not a code bug — every byte it produced decoded cleanly — it was a
+*documentation* bug, and it survived a compile, a disassembly golden, and a
+structural check, because a well-formed instruction that does the wrong thing
+is indistinguishable from a correct one if you never say what it should do.
+
+So the table is not written from memory here. It is measured, by two
+independent oracles that share no code, and each half is measured the way that
+half can be.
+
+## What each file establishes
+
+| file | oracle | establishes |
+|---|---|---|
+| `modrm19.s` | **execution** on a real 8086 under `qemu-system-i386` | the mod=00/01/10 effective addresses |
+| `modrm11.s` + `modrm11.py` | **encoding** by GNU `as` (`.code16`), decoded by FCML | the mod=11 register identities |
+
+The split is not arbitrary. mod=11 is not an address form at all — the r/m
+field names a register — so the probe that measures addresses by writing a
+marker and reading back the address that received it cannot be extended to
+it. See the header of `modrm11.s` for the three reasons a naive extension
+fails; they are worth reading, because the first version of that probe
+produced two impossible answers and a plausible-looking third.
+
+## mod=00/01/10 — measured by execution
+
+`modrm19.s` runs 24 cases (8 r/m values x 3 mod values). For each it clears
+low memory, stores a marker through the encoding under test, then scans for
+the word and reports the offset it landed at. `BX=1000 DI=2000 SI=0030
+BP=0040`, so every candidate address is distinct and the answer is the
+offset's arithmetic, not a judgement call.
+
+To run it (needs `qemu-system-i386`, `as`, `objcopy`):
+
+```sh
+cd shell
+as --32 -o /tmp/modrm19.o tests/probe/modrm19.s
+objcopy -O binary -j .text /tmp/modrm19.o /tmp/modrm19.bin
+python3 tests/probe/run_modrm19.py
+```
+
+`run_modrm19.py` assembles, wraps the code in a boot sector (`EB 3C` at
+offset 0, code at `0x3E`, `55 AA` at `0x1FE` — the whole thing must fit in
+`0x3E + len(code) <= 512`), boots it under qemu with the serial port
+captured to a file, and decodes the 24 groups of `lo hi 0x20` into the table
+in `Runtime.mod`, asserting it.
+
+The result, as measured:
+
+```
+mod=00 : [BX+SI] [BX+DI] [BP+SI] [BP+DI] [SI]    [DI]    disp16  [BX]
+mod=01 : ...+disp8 for each of the above, with rm=110 -> [BP]+disp8
+mod=10 : ...+disp16 for each of the above
+```
+
+The `rm=011` cells are `[BP+SI]` and `[BP+DI]`, not `[BX+SI]` and `[BX+DI]`;
+that is the single most common misreading of the table and it is invisible
+in a disassembly.
+
+**One cell is not covered by execution.** mod=10, rm=001 lands at
+`BX+DI+disp16` = `4234`, which is outside the probe's scan window, so that one
+case reports "not found". It is recorded as a gap, not as a result, and it is
+covered statically instead by the recipes in `Runtime.mod` and by
+`fcml_vs_objdump.py`.
+
+## mod=11 — measured by encoding
+
+`modrm11.py` assembles `modrm11.s`, disassembles the result with FCML
+(`../disasm16.py`), and requires four things to agree:
+
+1. what the `.s` asked for,
+2. what `as` emitted,
+3. the bytes hard-coded in `EXPECT` in `modrm11.py`,
+4. the table in `Runtime.mod`, via the ModRM field of each instruction.
+
+The result, as measured:
+
+```
+mod=11 r/m field   000   001  010  011  100  101  110  111
+                   AX    CX   DX   BX   SP   BP   SI   DI
+
+mod=11 reg field   same list, same order
+
+8-bit mod=11       AL    CL   DL   BL   AH   CH   DH   BH
+```
+
+and the direction, which depends on the opcode and not on the ModRM byte:
+
+```
+88 /r  MOV r/m8,  r8      89 /r  MOV r/m16, r16     reg is the SOURCE
+8A /r  MOV r8,  r/m8      8B /r  MOV r16, r/m16     reg is the TARGET
+```
+
+So `88 C6` and `8A C6` are the same ModRM byte and opposite instructions.
+The word and byte lists differ at exactly one cell: 100 is SP in one, AH in
+the other.
+
+### Why this check is not self-referential
+
+Comparing the `.s` against `as` cannot fail on its own: edit the `.s` and `as`
+faithfully re-encodes the new claim, so the two always agree. That failure
+mode was found by deliberately corrupting the `.s` and watching the check
+stay green. Two things close it:
+
+- **`EXPECT`**, a hard-coded byte sequence per instruction, asserted
+  independently of the `.s` text. Editing the `.s` away from the table now
+  fails.
+- **four anchors** — `83 C4 08` (ADD SP,8), `83 C6 02` (ADD SI,2), `8B EC`
+  (MOV BP,SP), `8B E5` (MOV SP,BP). These are encodings nobody types by hand,
+  so they carry information the `.s`'s author did not supply, and a table
+  shifted by one cell cannot satisfy them. They are spelled as literal
+  `.byte` directives because `as` picks opcode 89 over 8B for a
+  register-to-register move and would never emit `8B EC` for a mnemonic.
+
+`../nonvacuity.sh` exercises all of this: it corrupts a table cell, restores
+the exact table this project once shipped (AX dropped off the front, a
+duplicate BX invented at the end), edits the `.s` away from `EXPECT`, edits
+the 8-bit list, and corrupts an anchor byte — asserting each one goes red
+for the stated reason.
+
+## The mistake, for the record
+
+The table that shipped was
+
+```
+rm:  CX DX BX SP BP SI DI BX
+```
+
+which is the correct list with AX dropped off the front and a duplicate BX
+invented at the end. Every code is one too low except 100, which lands on SP
+either way. So:
+
+- `89 DC` reads as `MOV SP,BX` under both tables, but the emitter that
+  intended `MOV SI,BX` wrote it. That is `MovSiBx`, and `39 DC` is `CmpSiBx`.
+- The error is invisible at 100 and only shows up from 101 down, which is why
+  it survived as long as it did.
+
+The correct encodings: `MOV SI,BX` is `89 DE` (reg=BX is the source, r/m=SI
+is the destination) and `CMP SI,BX` is `39 DE`.
+
+A third case is instructive because it is the same mistake pointed the other
+way. `MovSiAx` was written `8B C0`. Under the shifted table that looks like
+`MOV AX,AX` — a no-op, so the fix was to reach for the byte the shifted table
+said was SI. It is worth being explicit that the fix `8B F0` is correct and
+`8B C0` is not, because the direction is the part that is easy to get
+backwards: for 8B the *reg* field is the destination, so `8B F0`
+(reg=110=SI, r/m=000=AX) is `MOV SI,AX`, which is what the name asks for.
+
+## Why no Unicorn
+
+`rt_exec.py` is not part of this. `unicorn 2.1.4` decodes 16-bit ModRM
+incorrectly on this machine, and it is the only component in the chain that
+is wrong, which makes it useless as a reference — see `SUMMARY.md`.

+ 287 - 0
shell/tests/probe/modrm11.py

@@ -0,0 +1,287 @@
+#!/usr/bin/env python3
+"""modrm11.py -- check the mod=11 half of the 16-bit ModR/M table.
+
+Establishes the register identities for ModRM mod=11 by asking GNU `as`
+(.code16) to ENCODE the register moves, FCML to decode what came out, and
+requiring all of that to agree with both the hard-coded byte sequences below
+and the table recorded in Runtime.mod.  See probe/modrm11.s for why the
+assembler is the right oracle here, and why the obvious qemu probe for this
+half cannot work.
+
+The table
+---------
+ModRM is  mod:b7b6  reg:b5b4b3  r/m:b2b1b0.  In mod=11 both fields name a
+register, in the same order, with the same list:
+
+    code    000  001  010  011  100  101  110  111
+    reg     AX   CX   DX   BX   SP   BP   SI   DI
+    r/m     AX   CX   DX   BX   SP   BP   SI   DI
+
+Which of the two operands a field names depends on the opcode, and this is
+the part that has to be right or the whole table comes out shifted:
+
+    88 /r  MOV r/m8, r8     89 /r  MOV r/m16, r16    reg is the SOURCE
+    8A /r  MOV r8, r/m8     8B /r  MOV r16, r/m16    reg is the TARGET
+
+So for `movw %bx, %si` (AT&T: source BX, destination SI) the reg field is
+BX and the r/m field is SI -- and SI lands in the low three bits as 110.
+
+The memorable version of the r/m column drops AX off the front and invents a
+duplicate BX at the end, which shifts every code down by one cell.  That is
+how "MovSiBx" came to be written 89 DC, which is MOV SP,BX.  The shifted
+version gets 100 right by luck (SP is at 100 either way), so the error hides
+until you check a register below it, and every wrong byte still decodes
+cleanly -- a structural check and a disassembly golden both accept it.  That
+is the entire reason this file exists.
+
+Non-vacuity
+-----------
+Groups 1-3 of the .s compare that file against the assembler, which on its
+own cannot fail for a reason other than "someone edited the claims": edit
+the .s and the assembler faithfully re-encodes the new claim.  Two things
+stop that.  Every instruction's bytes are also hard-coded in EXPECT below, so
+editing the .s to contradict the table fails even though the encoder agrees.
+And group 4 of the .s is four anchors nobody types by hand -- ADD SP,8 /
+ADD SI,2 / MOV BP,SP / MOV SP,BP -- whose cells are asserted against the
+table directly, so a table shifted by one cell cannot satisfy them.  Both
+failure modes are exercised by probe/README.md's non-vacuity runs.
+
+Usage: modrm11.py [-v]        # -v prints the table
+"""
+
+import os
+import re
+import subprocess
+import sys
+
+HERE = os.path.dirname(os.path.abspath(__file__))
+sys.path.insert(0, HERE)
+sys.path.insert(0, os.path.dirname(HERE))   # disasm16.py lives one level up
+import disasm16  # noqa: E402
+
+SRC = os.path.join(HERE, "modrm11.s")
+
+# The word-form register list, used for BOTH the reg field and the r/m field
+# in mod=11.  Index is the code.
+REG = ["Ax", "Cx", "Dx", "Bx", "Sp", "Bp", "Si", "Di"]
+# The 8-bit list, which differs from REG at index 4 only: AH, not SP.
+REG8 = ["Al", "Cl", "Dl", "Bl", "Ah", "Ch", "Dh", "Bh"]
+
+# Group boundaries in the .s, as (first, count).
+G_RM, G_REG, G_BYTE, G_ANCHOR = 0, 8, 16, 22
+
+# The bytes every instruction in the .s must assemble to, in order.  These are
+# the point of the file: they are asserted independently of the .s text, so
+# the .s cannot be edited into agreement with a wrong table.
+EXPECT = [
+    "89 D9", "89 DA", "89 DB", "89 DC", "89 DD", "89 DE", "89 DF", "89 DB",
+    "89 C7", "89 CF", "89 D7", "89 DF", "89 E7", "89 EF", "89 F7", "89 FF",
+    "88 C6", "88 F0", "88 C2", "88 D0", "88 C3", "88 D8",
+    "83 C4 08", "83 C6 02", "8B EC", "8B E5",
+]
+
+# The anchors, each with the decode its .s comment claims and the table cells
+# it pins.  (index into REG, register).  The opcode's reg field carries the
+# operation (/0 = ADD) for the first two, so only their r/m fields say
+# anything about the register list; the two MOV anchors pin both fields.
+ANCHORS = [
+    ("ADD SP, 8",  "add sp,8h",  [(4, "Sp")]),
+    ("ADD SI, 2",  "add si,2h",  [(6, "Si")]),
+    ("MOV BP, SP", "mov bp,sp",  [(5, "Bp"), (4, "Sp")]),
+    ("MOV SP, BP", "mov sp,bp",  [(4, "Sp"), (5, "Bp")]),
+]
+
+RE_INSN = re.compile(r"^\s*([a-z]+)\s+([^,]+),\s*([^#;]+?)\s*(?:#.*|;.*)?$")
+
+
+def build():
+    """assemble and objcopy, returning the .text bytes"""
+    o = "/tmp/modrm11.o"
+    b = "/tmp/modrm11.bin"
+    subprocess.run(["as", "--32", "-o", o, SRC], check=True)
+    subprocess.run(["objcopy", "-O", "binary", "-j", ".text", o, b], check=True)
+    with open(b, "rb") as f:
+        return f.read()
+
+
+def source_instructions():
+    """the (mnemonic, src, dst) list the .s file asks for, in order.
+
+    The four anchor instructions are literal .byte directives, so they do
+    not appear here; they are checked from EXPECT and ANCHORS instead.
+    """
+    out = []
+    with open(SRC) as f:
+        for line in f:
+            m = RE_INSN.match(line)
+            if m and m.group(1) in ("movw", "movb"):
+                # strip the AT&T sigils: FCML prints Intel order, without them
+                out.append((m.group(1),
+                            m.group(2).strip().lstrip("%"),
+                            m.group(3).strip().lstrip("%")))
+    return out
+
+
+def decode_all(text, problems):
+    """linear sweep, returning [(hexbytes, inteltext)] per instruction"""
+    got = []
+    off = 0
+    while off < len(text):
+        t, n = disasm16.decode(text[off:], off)
+        if n == 0:
+            problems.append("decode error at +%d (%s)"
+                            % (off, text[off:off + 6].hex(" ").upper()))
+            break
+        got.append((text[off:off + n].hex(" ").upper(), (t or "?").lower()))
+        off += n
+    return got
+
+
+def main(argv):
+    verbose = "-v" in argv
+    problems = []
+    want = source_instructions()
+    if not want:
+        print("FAIL: no instructions parsed out of %s" % SRC)
+        return 1
+    if len(want) != len(EXPECT) - len(ANCHORS):
+        print("FAIL: %s asks for %d mnemonic lines; EXPECT has %d entries "
+              "less %d anchors" % (SRC, len(want), len(EXPECT), len(ANCHORS)))
+        return 1
+    text = build()
+    got = decode_all(text, problems)
+    if len(got) != len(EXPECT):
+        problems.append("assembled %d instructions, expected %d"
+                        % (len(got), len(EXPECT)))
+    if problems:
+        print("FAIL: %d problem(s)" % len(problems))
+        for p in problems:
+            print("  - %s" % p)
+        return 1
+
+    def field(i, shift):
+        """the 3-bit ModRM field of instruction i"""
+        return (int(got[i][0].split()[1], 16) >> shift) & 7
+
+    # --- the bytes must be the ones recorded here, not merely self-consistent
+    for i, exp in enumerate(EXPECT):
+        if got[i][0] != exp:
+            # instructions at or past G_ANCHOR are literal .byte directives and
+            # so have no source line to name them by
+            label = (" ".join(want[i]) if i < G_ANCHOR
+                     else "anchor %d" % (i - G_ANCHOR))
+            problems.append("instruction %d (%s): assembled %s, expected %s"
+                            % (i, label, got[i][0], exp))
+
+    # --- every mnemonic line must decode to what the .s asked for ---------
+    for i in range(G_ANCHOR):
+        mne, src, dst = want[i]
+        want_txt = "mov %s,%s" % (dst.lower(), src.lower())
+        if got[i][1] != want_txt:
+            problems.append("instruction %d: asked for AT&T `%s %s, %s`, "
+                            "as+fcml gave Intel %r"
+                            % (i, mne, src, dst, got[i][1]))
+
+    # --- group 1: the r/m column, low three bits, naming the DESTINATION --
+    for i in range(G_RM, G_RM + 8):
+        mne, src, dst = want[i]
+        c = field(i, 0)
+        if REG[c].lower() != dst.lower():
+            problems.append("r/m %d: %s has ModRM r/m=%d, table says %s but "
+                            "the .s asked for %s"
+                            % (i - G_RM, got[i][0], c, REG[c], dst))
+
+    # --- group 2: the reg column, bits 5..3, naming the SOURCE -----------
+    for i in range(G_REG, G_REG + 8):
+        mne, src, dst = want[i]
+        c = field(i, 3)
+        if REG[c].lower() != src.lower():
+            problems.append("reg %d: %s has ModRM reg=%d, table says %s but "
+                            "the .s asked for %s"
+                            % (i - G_REG, got[i][0], c, REG[c], src))
+
+    # --- group 3: the 8-bit list, and the direction the opcode implies ---
+    for i in range(G_BYTE, G_BYTE + 6):
+        mne, src, dst = want[i]
+        op = int(got[i][0].split()[0], 16)
+        if op not in (0x88, 0x8A):
+            problems.append("byte move %d: opcode %02X, expected 88 or 8A"
+                            % (i - G_BYTE, op))
+            continue
+        # 88 is MOV r/m8,r8 (r/m is the destination); 8A is MOV r8,r/m8
+        # (r/m is the source).  The ModRM byte is identical either way, so
+        # only the opcode says which way the data goes.
+        rmf = field(i, 0)
+        rmf_reg = dst if op == 0x88 else src
+        if REG8[rmf].lower() != rmf_reg.lower():
+            problems.append("byte move %d: %s has r/m=%d, table says %s but "
+                            "the .s asked for %s"
+                            % (i - G_BYTE, got[i][0], rmf, REG8[rmf],
+                               rmf_reg))
+
+    # --- group 4: the anchors, against the table directly ----------------
+    # The anchors are literal .byte directives, so there is no source line to
+    # cross-check: their bytes come from EXPECT and their decode comes from
+    # here, and both are asserted.  What they add is a table cell that this
+    # file's author did not choose.
+    for k, (label, want_txt, cells) in enumerate(ANCHORS):
+        i = G_ANCHOR + k
+        ghex, gtext = got[i]
+        if gtext != want_txt:
+            problems.append("anchor %s (%s): FCML decodes %r, the .s comment "
+                            "claims %r" % (label, ghex, gtext, want_txt))
+        for c, reg in cells:
+            if REG[c].lower() != reg.lower():
+                problems.append("anchor %s (%s): ModRM code %d must be %s, "
+                                "table says %s"
+                                % (label, ghex, c, reg, REG[c]))
+        if not any(REG[c].lower() == reg.lower() for c, reg in cells):
+            problems.append("anchor %s (%s) pins no live cell -- the anchor "
+                            "has stopped testing anything" % (label, ghex))
+
+    if verbose:
+        print("mod=11 r/m field.  89 is MOV r/m,r, so reg is pinned to BX and")
+        print("the low three bits ARE the r/m code, naming the destination.")
+        for i in range(G_RM, G_RM + 8):
+            print("   r/m %03d  asked %-4s  %s  code %d  table %-4s %s"
+                  % (i - G_RM, want[i][2], got[i][0], field(i, 0),
+                     REG[field(i, 0)],
+                     "ok" if REG[field(i, 0)].lower() == want[i][2].lower()
+                     else "MISMATCH"))
+        print("mod=11 reg field.  rm is pinned to DI, bits 5..3 are the reg")
+        print("code, and for 89 that is the source.")
+        for i in range(G_REG, G_REG + 8):
+            print("   reg %03d  asked %-4s  %s  code %d  table %-4s %s"
+                  % (i - G_REG, want[i][1], got[i][0], field(i, 3),
+                     REG[field(i, 3)],
+                     "ok" if REG[field(i, 3)].lower() == want[i][1].lower()
+                     else "MISMATCH"))
+        print("8-bit mod=11: same shape, list differs at 100 (AH not SP),")
+        print("and the opcode alone says which way the data goes.")
+        for i in range(G_BYTE, G_BYTE + 6):
+            print("   asked %-4s -> %-4s  %s  r/m %d = %s"
+                  % (want[i][1], want[i][2], got[i][0], field(i, 0),
+                     REG8[field(i, 0)]))
+        print("anchors: encodings nobody types by hand, asserted against the")
+        print("table.  A table shifted by one cell cannot satisfy these.")
+        for k, (label, _, cells) in enumerate(ANCHORS):
+            i = G_ANCHOR + k
+            print("   %-10s %-9s %-12s pins %s"
+                  % (label, got[i][0], got[i][1],
+                     ", ".join("%d=%s" % (c, reg) for c, reg in cells)))
+
+    print("mod=11 table: %d r/m codes, %d reg codes, %d byte moves and %d "
+          "anchors all agree -- as encoding, FCML decode, the hard-coded"
+          % (8, 8, 6, len(ANCHORS)))
+    print("               EXPECT bytes, and the table in Runtime.mod")
+    if problems:
+        print("FAIL: %d problem(s)" % len(problems))
+        for p in problems:
+            print("  - %s" % p)
+        return 1
+    print("PASS: mod=11 register identities match the table in Runtime.mod")
+    return 0
+
+
+if __name__ == "__main__":
+    sys.exit(main(sys.argv))

+ 127 - 0
shell/tests/probe/modrm11.s

@@ -0,0 +1,127 @@
+# modrm11.s -- establish the mod=11 half of the 16-bit ModR/M table using
+# GNU as as an independent ENCODER.
+#
+# modrm19.s measures the memory forms (mod=00/01/10) by executing them on a
+# real 8086 under qemu: it stores a marker through each encoding and reports
+# which physical address received it.  mod=11 is not a memory form at all --
+# the r/m field names a register -- so it needs a different oracle, and this
+# file is it.
+#
+# Why the assembler rather than another qemu probe
+# ------------------------------------------------
+# The obvious extension of modrm19.s is "store into the register, then report
+# which register took it".  That does not work, and the reason is worth
+# recording because it is a trap rather than a puzzle:
+#
+#   * The comparison needs a register to hold the expected value, and every
+#     register is a candidate, so the check clobbers what it is measuring.
+#   * r/m=100 is SP, and the marker value is not a legal stack pointer.  The
+#     very next `call` pushes its return address through SS:BEEF and leaves
+#     SP at BEED, so by the time any check runs, the evidence is gone.  A
+#     probe that reports "not found" for SP is reporting its own design, not
+#     the hardware.  (This is not hypothetical: the first version of that
+#     probe did exactly this, and also cleared AX "because AX is never an
+#     r/m target" -- which is true of the mod=11 list people remember, and
+#     false of the real one, where r/m 000 IS AX.  It reported two
+#     impossible answers and a plausible-looking third.)
+#   * Repairing that needs the check inlined between the store and the next
+#     call, so every case gets its own hand-written code block and its own
+#     chance of a typo -- and a typo there produces a plausible wrong answer,
+#     which is the worst kind.
+#
+# The assembler has none of those problems.  `as --32` with `.code16` is a
+# correct 16-bit x86 *encoder*, and it was already the encoder half of the
+# FCML cross-validation (38/38 agreement on a separate probe).  Asking it to
+# encode `movw %bx, %si` and getting `89 DE` back is a direct, unambiguous
+# statement that in mod=11 the r/m field 110 means SI.
+#
+# So the two files are complementary halves of one argument.
+#   modrm19.s  execution  ->  mod=00/01/10 effective addresses
+#   this file  encoding   ->  mod=11 register identities
+# Nothing in the table is taken from memory, and the two oracles share no code.
+#
+# This file is in four groups, and modrm11.py checks all four differently:
+#
+#   1. the r/m column   8 x `movw %bx, <reg>`  -> low three bits of ModRM
+#   2. the reg column   8 x `movw <reg>, %di`  -> bits 5..3 of ModRM
+#   3. the byte list    6 x 8-bit moves        -> AL CL DL BL AH CH DH BH
+#   4. the anchors      4 hand-checking bytes  -> the table itself
+#
+# Group 4 is what makes the check non-vacuous.  Groups 1-3 are self
+# referential in a dangerous way: they compare this file against the
+# assembler, so editing this file just changes the claim and the assembler
+# faithfully re-encodes it.  A check like that cannot fail on a wrong table
+# unless the table is what moved.  The anchors are different -- they are
+# encodings nobody types by hand, so they carry information this file's
+# author did not supply, and a table shifted by one position cannot satisfy
+# them.
+#
+# Encode and check (see modrm11.py):
+#   as --32 -o modrm11.o modrm11.s
+#   objcopy -O binary -j .text modrm11.o modrm11.bin
+
+        .code16
+        .text
+
+# --- 1. the r/m field, read straight off the low three bits -------------
+# Each of these is `movw %bx, <reg>`, i.e. opcode 89 with reg=BX (011), so
+# ModRM = 11 011 rrr and the low three bits ARE the r/m code for the register
+# named on the right.  One instruction per r/m code, in r/m order.  Opcode 89
+# is MOV r/m,r, so the r/m field is the DESTINATION -- the register on the
+# right of the AT&T line.  Both of those directions matter, and getting
+# either backwards produces a table that is wrong everywhere.
+        movw    %bx, %cx         # r/m 000
+        movw    %bx, %dx         # r/m 010
+        movw    %bx, %bx         # r/m 011
+        movw    %bx, %sp         # r/m 100
+        movw    %bx, %bp         # r/m 101
+        movw    %bx, %si         # r/m 110
+        movw    %bx, %di         # r/m 111
+        movw    %bx, %bx         # r/m 011 again -- there is no second BX
+
+# --- 2. the reg field, read off bits 5..3 ------------------------------
+# `movw <reg>, %di` is opcode 89 with rm=DI (111), so ModRM = 11 rrr 111 and
+# bits 5..3 are the r/g code, which is the SOURCE.  All eight codes are
+# encodable, AX included: 89 with reg=AX and mod=11 is an ordinary
+# MOV r/m,r, not an accumulator short form -- unlike ADD/ADC/AND/OR/SBB/SUB/
+# XOR/CMP, where /0 means the accumulator and the ModRM byte does collapse.
+        movw    %ax, %di         # reg 000
+        movw    %cx, %di         # reg 001
+        movw    %dx, %di         # reg 010
+        movw    %bx, %di         # reg 011
+        movw    %sp, %di         # reg 100
+        movw    %bp, %di         # reg 101
+        movw    %si, %di         # reg 110
+        movw    %di, %di         # reg 111
+
+# --- 3. the 8-bit forms, where the list differs and direction flips -----
+# 88 /r is MOV r/m8,r8 (reg is the SOURCE); 8A /r is MOV r8,r/m8 (reg is the
+# DESTINATION).  Both use the identical ModRM byte for the same pair of
+# registers, so the byte alone cannot tell you the direction -- the opcode
+# can.  The 8-bit list is also its own: AL CL DL BL AH CH DH BH, which agrees
+# with the word list at every code except 100, where it is AH rather than
+# SP.  This is the other trap in the table, and Runtime.mod documents it.
+        movb    %al, %dh         # 88 C6  ->  DH := AL
+        movb    %dh, %al         # 88 F0  ->  AL := DH
+        movb    %al, %dl         # 88 C2  ->  DL := AL
+        movb    %dl, %al         # 88 D0  ->  AL := DL
+        movb    %al, %bl         # 88 C3  ->  BL := AL
+        movb    %bl, %al         # 88 D8  ->  AL := BL
+
+# --- 4. the anchors -----------------------------------------------------
+# Four instructions no one writes by hand, each of which pins one cell of the
+# table.  If the ModRM column in modrm11.py is shifted by one, every one of
+# these disagrees with it.  They are the reason this file can fail for a
+# reason other than "someone edited the claims".
+#
+# They are spelled as literal bytes, not as mnemonics, because the point is
+# the exact sequence: `as` picks opcode 89 rather than 8B for a
+# register-to-register move, so asking it for `movw %sp, %bp` gets 89 E5 and
+# never 8B EC.  Both mean MOV BP,SP and the two differ in which field holds
+# which register -- which is exactly what the anchor is here to pin.
+# modrm11.py supplies the expected decode for each, so a byte that does not
+# decode as its comment claims still fails.
+        .byte   0x83, 0xC4, 0x08     # ADD SP, 8      rm=100 -> SP
+        .byte   0x83, 0xC6, 0x02     # ADD SI, 2      rm=110 -> SI
+        .byte   0x8B, 0xEC           # MOV BP, SP     reg=101 rm=100
+        .byte   0x8B, 0xE5           # MOV SP, BP     reg=100 rm=101

+ 193 - 0
shell/tests/probe/modrm19.s

@@ -0,0 +1,193 @@
+# modrm19.s -- derive the complete 16-bit ModR/M effective-address table by
+# EXECUTION, not from memory.  24 cases: 8 r/m values x 3 mod values.
+#
+# For each case: clear 0x0000..0x4FFF, store 0xBEEF through the encoding under
+# test, then scan for the word and emit its offset as 2 raw bytes.
+# BX=0x1000 DI=0x2000 SI=0x0030 BP=0x0040, so every candidate is distinct.
+#
+#   mod=00 : no displacement follows
+#   mod=01 : disp8 = 0x44
+#   mod=10 : disp16 = 0x1234
+#
+# The three "mod=11" cases are register-to-register and are not addressed here.
+# Output: 24 groups of "lo hi 0x20", then 0x0A 0x0A.
+
+.code16
+.text
+.globl _start
+_start:
+        cli
+        xorw    %ax, %ax
+        movw    %ax, %es
+        movw    $0x1000, %bx
+        movw    $0x2000, %di
+        movw    $0x0030, %si
+        movw    $0x0040, %bp
+
+        # ---- mod = 00, r/m = 000..111 -------------------------------
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x00          # 00 000
+        call    rep_
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x01          # 00 001
+        call    rep_
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x02          # 00 010
+        call    rep_
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x03          # 00 011
+        call    rep_
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x04          # 00 100
+        call    rep_
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x05          # 00 101
+        call    rep_
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x06, 0x34, 0x12   # 00 110 direct
+        call    rep_
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x07          # 00 111
+        call    rep_
+
+        # ---- mod = 01, disp8 = 0x44 --------------------------------
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x40, 0x44    # 01 000
+        call    rep_
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x41, 0x44    # 01 001
+        call    rep_
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x42, 0x44    # 01 010
+        call    rep_
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x43, 0x44    # 01 011
+        call    rep_
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x44, 0x44    # 01 100
+        call    rep_
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x45, 0x44    # 01 101
+        call    rep_
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x46, 0x44    # 01 110
+        call    rep_
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x47, 0x44    # 01 111
+        call    rep_
+
+        # ---- mod = 10, disp16 = 0x1234 ------------------------------
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x80, 0x34, 0x12   # 10 000
+        call    rep_
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x81, 0x34, 0x12   # 10 001
+        call    rep_
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x82, 0x34, 0x12   # 10 010
+        call    rep_
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x83, 0x34, 0x12   # 10 011
+        call    rep_
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x84, 0x34, 0x12   # 10 100
+        call    rep_
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x85, 0x34, 0x12   # 10 101
+        call    rep_
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x86, 0x34, 0x12   # 10 110
+        call    rep_
+        call    clr
+        movw    $0xBEEF, %ax
+        .byte   0x26, 0x89, 0x87, 0x34, 0x12   # 10 111
+        call    rep_
+
+        movw    $0x3F8, %dx
+        movb    $10, %al
+        outb    %al, %dx
+        hlt
+
+# clr -- clear 0000:0000..0000:4FFF
+clr:
+        pushw   %di
+        pushw   %cx
+        pushw   %ax
+        xorw    %ax, %ax
+        movw    $0x0000, %di
+        movw    $0x2800, %cx
+        rep     stosw
+        popw    %ax
+        popw    %cx
+        popw    %di
+        ret
+
+# rep_ -- emit the address of the word just stored
+rep_:
+        pushw   %ax
+        pushw   %bx
+        pushw   %cx
+        pushw   %dx
+        pushw   %si
+        pushw   %di
+        xorw    %si, %si
+scan:
+        cmpw    $0x3600, %si
+        jae     none
+        movw    %es:(%si), %dx
+        cmpw    $0xBEEF, %dx
+        je      found
+        incw    %si
+        jmp     scan
+found:
+        movw    %si, %ax
+        call    putb
+        movb    %ah, %al
+        call    putb
+        jmp     done
+none:
+        movw    $0xFF, %ax
+        call    putb
+        movw    $0xFF, %ax
+        call    putb
+done:
+        movw    $0x3F8, %dx
+        movb    $' ', %al
+        outb    %al, %dx
+        popw    %di
+        popw    %si
+        popw    %dx
+        popw    %cx
+        popw    %bx
+        popw    %ax
+        ret
+
+putb:
+        pushw   %dx
+        movw    $0x3F8, %dx
+        outb    %al, %dx
+        popw    %dx
+        ret

+ 159 - 0
shell/tests/probe/run_modrm19.py

@@ -0,0 +1,159 @@
+#!/usr/bin/env python3
+"""run_modrm19.py -- re-run the mod=00/01/10 ModR/M probe and check the result.
+
+modrm19.s derives the 16-bit effective-address table by EXECUTION: it stores a
+marker through each candidate encoding on a real 8086 under qemu, then scans
+memory for the word and reports the offset it landed at.  This script builds
+the bootable image, runs it, and requires the measured offsets to equal the
+table computed from first principles here -- which is the point, because the
+computation in EXPECTED is written from the register arithmetic (BX+SI and so
+on) and not from a remembered table.
+
+    usage: run_modrm19.py [-v] [--keep]
+
+Requires qemu-system-i386, as and objcopy.  This is not part of run_all.sh:
+it needs a hand-built floppy image and a 60-second qemu timeout, and its
+result is already recorded in the table in Runtime.mod, which
+tests/probe/modrm11.py checks on every run.  Re-run this when you doubt the
+memory forms.
+
+mod=11 is not measured here -- it names a register, not an address.  See
+modrm11.s for why, and modrm11.py for how that half is established.
+"""
+
+import os
+import shutil
+import subprocess
+import sys
+
+HERE = os.path.dirname(os.path.abspath(__file__))
+SRC = os.path.join(HERE, "modrm19.s")
+WORK = "/tmp/opencode/modrm19"
+
+# The probe's register setup, from modrm19.s.  Distinct values, so the offset
+# a marker lands at identifies the effective address by arithmetic alone.
+BX, DI, SI, BP = 0x1000, 0x2000, 0x0030, 0x0040
+DISP8, DISP16 = 0x44, 0x1234
+
+# The effective address for (mod, rm), computed from the arithmetic.  This is
+# the prediction; the probe's output is the measurement, and the two are
+# compared.  mod=00 rm=110 is the direct disp16 form, so the address is the
+# displacement itself.
+def expected(mod, rm):
+    base = {0: BX + SI, 1: BX + DI, 2: BP + SI, 3: BP + DI,
+            4: SI, 5: DI, 7: BX}
+    if mod == 0:
+        return DISP16 if rm == 6 else base[rm]
+    disp = DISP8 if mod == 1 else DISP16
+    if rm == 6:                       # [BP]+disp
+        return BP + disp
+    return base[rm] + disp
+
+# The one cell execution cannot answer.  mod=10 rm=001 is BX+DI+disp16 =
+# 0x4234, and the probe's scan window stops at 0x3600, so the marker is
+# written but never seen.  Recorded as a gap, not as a result; the cell is
+# covered statically by the recipes in Runtime.mod.
+GAP = {(2, 1)}
+
+
+def build():
+    subprocess.run(["as", "--32", "-o", WORK + ".o", SRC], check=True)
+    subprocess.run(["objcopy", "-O", "binary", "-j", ".text",
+                    WORK + ".o", WORK + ".bin"], check=True)
+    with open(WORK + ".bin", "rb") as f:
+        code = f.read()
+    # A boot sector is 512 bytes: EB 3C at 0, code at 0x3E, 55 AA at 0x1FE.
+    if 0x3E + len(code) > 512:
+        raise SystemExit("probe code is %d bytes, does not fit after the "
+                         "0x3E header" % len(code))
+    img = bytearray(512)
+    img[0:2] = b"\xeb\x3c"
+    img[0x3E:0x3E + len(code)] = code
+    img[0x1FE:0x200] = b"\x55\xaa"
+    with open(WORK + ".img", "wb") as f:
+        f.write(bytes(img))
+    return len(code)
+
+
+def run():
+    """boot the image under qemu, return the captured serial bytes"""
+    ser = WORK + ".ser"
+    if os.path.exists(ser):
+        os.remove(ser)
+    # The guest does all its work in the first few milliseconds and then halts;
+    # qemu keeps running, so the timeout is what ends it.  rc=124 is the normal
+    # outcome.  Any other non-zero rc is a real failure.
+    rc = subprocess.run(["timeout", "10", "qemu-system-i386",
+                         "-drive", "file=%s.img,format=raw,if=floppy" % WORK,
+                         "-serial", "file:" + ser,
+                         "-display", "none", "-no-reboot"],
+                        stdout=subprocess.DEVNULL,
+                        stderr=subprocess.DEVNULL).returncode
+    if rc not in (0, 124):
+        raise SystemExit("qemu exited %d; the probe did not run to the end"
+                         % rc)
+    with open(ser, "rb") as f:
+        return f.read()
+
+
+def main(argv):
+    verbose = "-v" in argv
+    for tool in ("as", "objcopy", "qemu-system-i386"):
+        if not shutil.which(tool):
+            print("SKIP: %s not installed" % tool)
+            return 0
+    os.makedirs(os.path.dirname(WORK), exist_ok=True)
+    n = build()
+    data = run()
+
+    # 24 groups of "lo hi 0x20", then a blank line.  The probe writes 0A 0A
+    # but qemu can lose the last byte when it tears down, so require the tail
+    # to be newline(s) and not insist on both.
+    if len(data) < 24 * 3 + 1 or set(data[24 * 3:]) - {0x0A}:
+        print("FAIL: serial capture is %d bytes and does not end in the "
+              "expected blank line" % len(data))
+        return 1
+    got = [int.from_bytes(data[i:i + 2], "little")
+           for i in range(0, 24 * 3, 3)]
+    if any(data[i + 2] != 0x20 for i in range(0, 24 * 3, 3)):
+        print("FAIL: a group separator is not 0x20")
+        return 1
+
+    bad = []
+    for mod in range(3):
+        for rm in range(8):
+            g = got[mod * 8 + rm]
+            if (mod, rm) in GAP:
+                if g != 0xFFFF:
+                    bad.append("mod=%d rm=%d: expected the known gap "
+                               "(0xFFFF, outside the scan window), got %04X"
+                               % (mod, rm, g))
+                continue
+            w = expected(mod, rm)
+            if g != w:
+                bad.append("mod=%d rm=%d: measured %04X, arithmetic says %04X"
+                           % (mod, rm, g, w))
+    if verbose:
+        for mod in range(3):
+            row = []
+            for rm in range(8):
+                g = got[mod * 8 + rm]
+                row.append("  gap " if (mod, rm) in GAP else "%04X" % g)
+            print("mod=%02d : %s" % (mod, "  ".join(row)))
+
+    print("mod=00/01/10: %d of 24 cells measured by execution on a real 8086"
+          % (24 - len(GAP)))
+    print("               %d cell is a documented gap, not a result"
+          % len(GAP))
+    if bad:
+        print("FAIL: %d problem(s)" % len(bad))
+        for b in bad:
+            print("  - %s" % b)
+        return 1
+    print("PASS: measured effective addresses match the arithmetic in "
+          "Runtime.mod")
+    return 0
+
+
+if __name__ == "__main__":
+    sys.exit(main(sys.argv))

+ 46 - 0
shell/tests/run_all.sh

@@ -18,6 +18,49 @@
 #   RtProbe             dumps the runtime size and its 14 entry offsets, so a
 #                       runtime change that moves an entry is visible here.
 #
+# Three runtime suites run before all of those, because everything else
+# trusts the runtime's bytes:
+#
+#   audit_helpers       disassembles every one-line emitter in Runtime.mod
+#                       and compares the DECODE against the procedure's NAME.
+#                       This is the only check that can see a wrong ModRM
+#                       that still decodes cleanly, which is the mistake this
+#                       file actually makes: MovSiBx and CmpSiBx were both
+#                       `DC` (= MOV SP,BX / CMP SP,BX) for a long time.
+#   check_runtime       sweeps the built runtime's code region with FCML: no
+#                       desync, every entry and every branch target on an
+#                       instruction boundary, and the whole disassembly equal
+#                       to tests/runtime.golden.
+#   nonvacuity          breaks the runtime six ways and asserts each named
+#                       check goes red for the stated reason.  Slower, so it
+#                       is not in the default run: run it when a check's
+#                       sensitivity is in question, or after editing a check.
+#
+# nonvacuity.sh is NOT run by default: it rebuilds the runtime six times and
+# it is a meta-check, so it belongs to "is the harness honest" reviews rather
+# than to every change.  Run tests/nonvacuity.sh directly after touching any
+# of the checks above.
+#
+# Two more checks guard the ModR/M table itself, which is the input every hand
+# written emitter in Runtime.mod depends on:
+#
+#   probe/modrm11.py    asks GNU `as` (.code16) to encode the register moves
+#                       and FCML to decode them, and requires the result to
+#                       agree with the hard-coded bytes, the four
+#                       hand-checking anchors, and the table in Runtime.mod.
+#                       This is the only check that can catch the TABLE being
+#                       wrong rather than one emitter: the table shipped once
+#                       with AX dropped off the front of the mod=11 column
+#                       and a duplicate BX invented at the end, which shifts
+#                       every code down by one and makes `89 DC` (MOV SP,BX)
+#                       look like the SI move the name asks for.
+#   probe/modrm19.s     the execution probe for mod=00/01/10, kept as the
+#                       record of how the memory forms were measured.  It is
+#                       not run by default: it needs qemu and a hand-built
+#                       floppy image, and its result is already written into
+#                       the table in Runtime.mod, which modrm11.py checks.
+#                       See probe/README.md to re-run it.
+#
 # rt_exec.py is NOT run: it needs Unicorn, whose 16-bit ModRM decoding is
 # wrong on this machine (see SUMMARY.md), so its failures would be the
 # emulator's, not the runtime's.  Running it by default would be noise.
@@ -59,6 +102,9 @@ run "runtime probe"  bash -c '
        >/tmp/tp_rt 2>&1 || { grep -m5 error: /tmp/tp_rt; exit 1; }
    /tmp/tp_rtprobe'
 
+run "helper audit"    python3 tests/audit_helpers.py
+run "mod=11 table"     python3 tests/probe/modrm11.py
+run "runtime image"   python3 tests/check_runtime.py
 run "compile matrix"  tests/run_compile_tests.sh
 run "COM linker"      tests/run_com_tests.sh
 run "UI error path"   python3 tests/uitest.py

+ 40 - 7
shell/tests/run_com_tests.sh

@@ -64,10 +64,34 @@ out = sys.argv[1]
 
 # Compiler layout constants, restated here on purpose: the checker must not
 # ask the code under test what the answer is.
-RTSZ  = 385                    # Runtime.RT_Size()
+#
+# Re-baselined 385 -> 391 with the runtime bug fixes (see the re-baseline
+# rationale in tests/runtime.golden and the ModRM table in Runtime.mod).  The
+# size is what a change to the runtime shows up in first; if this constant
+# ever needs changing again, find out why the runtime moved rather than just
+# editing the number, and check tests/check_runtime.py for the reason.
+RTSZ  = 391                    # Runtime.RT_Size()
 DATAB = RTSZ + 0x1000         # Compiler: data base = rtSz + 1000H
-# initmem: MOV AX,AX / MOV DX,[SI+4] / MOV CX,[SI+8]
-HEAD  = '8B C0 8B 54 04 8B 4C 08'
+
+# initmem's prologue, which is the runtime's only reader of the program
+# header.  The displacements +4 and +6 below are the whole point of this
+# constant: the header word checks further down read hdrDS at +4 and hdrHeap
+# at +6, and initmem has to read the SAME two words or it clears the wrong
+# range.  It used to read +8 (hdrMax, which the compiler patches to 0), so it
+# zeroed nothing at all, and nothing here noticed -- the emitted loop was
+# perfectly well formed, it just never ran.  Asserting the bytes and the
+# header offsets together is what closes that gap.
+HEAD  = '8B F0 8B 54 04 8B 4C 06'   # 11 bytes of initmem, see below
+HDR_DS_WORD   = 4             # header word holding the data base
+HDR_HEAP_WORD = 6             # header word holding the data end
+# Byte 4 of HEAD is the displacement of initmem's MOV DX,[SI+?], and byte 7
+# the displacement of its MOV CX,[SI+?].  The checks below read the header
+# words at HDR_DS_WORD and HDR_HEAP_WORD, so tying those two displacements to
+# the same two constants is what makes the runtime and the compiler agree by
+# construction rather than by coincidence.
+assert [int(HEAD.split()[4], 16), int(HEAD.split()[7], 16)] == \
+       [HDR_DS_WORD, HDR_HEAP_WORD], \
+       'initmem no longer reads the two header words this checker verifies'
 
 raw = open(os.path.join(out, 'raw.txt')).read()
 rows = re.findall(r'(\S+\.pas)\s+OK\s+com=(\d+)\s+image=(\d+)\s+data=(\d+)\s+nonzeroInGap=(\d+)', raw)
@@ -89,8 +113,9 @@ for name, com, image, data, nzg in rows:
     else:
         d = open(path, 'rb').read()
 
-    if d[:8].hex(' ').upper() != HEAD:
-        errs.append('runtime not at offset 0 (first bytes %s)' % d[:8].hex(' ').upper())
+    if d[:len(HEAD.split())].hex(' ').upper() != HEAD:
+        errs.append('runtime not at offset 0 (first bytes %s, want %s)'
+                    % (d[:len(HEAD.split())].hex(' ').upper(), HEAD))
     if len(d) != com:
         errs.append('file is %d bytes, harness said %d' % (len(d), com))
     if DATAB + data != len(d):
@@ -110,8 +135,8 @@ for name, com, image, data, nzg in rows:
         hdr = d[RTSZ:RTSZ+16]
         flag = int.from_bytes(hdr[0:2], 'little')
         cs   = int.from_bytes(hdr[2:4], 'little')
-        ds   = int.from_bytes(hdr[4:6], 'little')
-        heap = int.from_bytes(hdr[6:8], 'little')
+        ds   = int.from_bytes(hdr[HDR_DS_WORD:HDR_DS_WORD+2], 'little')
+        heap = int.from_bytes(hdr[HDR_HEAP_WORD:HDR_HEAP_WORD+2], 'little')
         if flag != 1:
             errs.append('hdrFlag=%d' % flag)
         # hdrCS is pc, the end of the code, and the harness's `image` is
@@ -123,6 +148,14 @@ for name, com, image, data, nzg in rows:
             errs.append('hdrDS=%d, want %d' % (ds, DATAB))
         if heap != DATAB + data:
             errs.append('hdrHeap=%d, want %d' % (heap, DATAB + data))
+        # initmem must read hdrDS and hdrHeap, not some other pair of header
+        # words.  (It read +8, hdrMax, which the compiler patches to 0, so
+        # the range it cleared was empty and every global kept whatever the
+        # loader left in it.)
+        if [d[4], d[7]] != [HDR_DS_WORD, HDR_HEAP_WORD]:
+            errs.append('initmem reads header words +%d/+%d, but the data '
+                        'base and data end are at +%d/+%d'
+                        % (d[4], d[7], HDR_DS_WORD, HDR_HEAP_WORD))
 
     if errs:
         bad += 1

+ 205 - 0
shell/tests/runtime.golden

@@ -0,0 +1,205 @@
+# runtime.golden -- full FCML disassembly of the runtime's code region.
+#
+# GENERATED by tests/check_runtime.py --bless, then READ THE DIFF and commit.
+# Every line is "OFFSET  BYTES  MNEMONIC".  This file is a deliberate
+# tripwire: any edit to Runtime.mod invalidates it, because inserting an
+# instruction renumbers everything after it, so you are forced to look at the
+# whole new listing and confirm you meant every change.  See the "Re-baselining
+# the golden" section of check_runtime.py.
+#
+# Region: bytes 0..codeEnd of the runtime image (see the code/data split in
+# EmitData).  Relative targets are absolute offsets within the runtime, which
+# the compiler adds to the runtime's load address when it emits a CALL.
+0000  8B F0       mov si,ax
+0002  8B 54 04    mov dx,word ptr [si+4h]
+0005  8B 4C 06    mov cx,word ptr [si+6h]
+0008  39 D1       cmp cx,dx
+000A  76 0D       jbe 19h
+000C  89 D7       mov di,dx
+000E  31 C0       xor ax,ax
+0010  89 05       mov word ptr [di],ax
+0012  83 C7 02    add di,2h
+0015  39 CF       cmp di,cx
+0017  72 F5       jb 0eh
+0019  1E          push ds
+001A  07          pop es
+001B  C3          ret
+001C  31 C0       xor ax,ax
+001E  B4 4C       mov ah,4ch
+0020  CD 21       int 21h
+0022  C3          ret
+0023  C3          ret
+0024  55          push bp
+0025  8B EC       mov bp,sp
+0027  8B 46 04    mov ax,word ptr [bp+4h]
+002A  83 F8 00    cmp ax,0h
+002D  7D 0A       jnl 39h
+002F  50          push ax
+0030  B2 2D       mov dl,2dh
+0032  B4 02       mov ah,2h
+0034  CD 21       int 21h
+0036  58          pop ax
+0037  F7 D8       neg ax
+0039  B9 0A 00    mov cx,0ah
+003C  BB 70 01    mov bx,170h
+003F  89 DE       mov si,bx
+0041  31 D2       xor dx,dx
+0043  F7 F1       div ax,cx
+0045  80 C2 30    add dl,30h
+0048  4E          dec si
+0049  88 14       mov byte ptr [si],dl
+004B  83 F8 00    cmp ax,0h
+004E  75 F1       jne 41h
+0050  39 DE       cmp si,bx
+0052  74 09       je 5dh
+0054  8A 14       mov dl,byte ptr [si]
+0056  B4 02       mov ah,2h
+0058  CD 21       int 21h
+005A  46          inc si
+005B  EB F3       jmp 50h
+005D  89 EC       mov sp,bp
+005F  5D          pop bp
+0060  C3          ret
+0061  8B EC       mov bp,sp
+0063  8A 56 02    mov dl,byte ptr [bp+2h]
+0066  89 EC       mov sp,bp
+0068  B4 02       mov ah,2h
+006A  CD 21       int 21h
+006C  C3          ret
+006D  8B EC       mov bp,sp
+006F  83 7E 02 00 cmp word ptr [bp+2h],0h
+0073  89 EC       mov sp,bp
+0075  75 05       jne 7ch
+0077  BA 76 01    mov dx,176h
+007A  EB 03       jmp 7fh
+007C  BA 70 01    mov dx,170h
+007F  B4 09       mov ah,9h
+0081  CD 21       int 21h
+0083  C3          ret
+0084  BA 80 01    mov dx,180h
+0087  B4 09       mov ah,9h
+0089  CD 21       int 21h
+008B  C3          ret
+008C  BA 7D 01    mov dx,17dh
+008F  B4 09       mov ah,9h
+0091  CD 21       int 21h
+0093  C3          ret
+0094  5B          pop bx
+0095  31 C9       xor cx,cx
+0097  8A 0F       mov cl,byte ptr [bx]
+0099  43          inc bx
+009A  B4 02       mov ah,2h
+009C  E3 07       jcxz 0a5h
+009E  8A 07       mov al,byte ptr [bx]
+00A0  CD 21       int 21h
+00A2  43          inc bx
+00A3  E2 F9       loop 9eh
+00A5  FF E3       jmp bx
+00A7  55          push bp
+00A8  8B EC       mov bp,sp
+00AA  50          push ax
+00AB  53          push bx
+00AC  51          push cx
+00AD  52          push dx
+00AE  57          push di
+00AF  E8 B1 00    call 163h
+00B2  3C 20       cmp al,20h
+00B4  74 F9       je 0afh
+00B6  3C 09       cmp al,9h
+00B8  74 F5       je 0afh
+00BA  3C 0D       cmp al,0dh
+00BC  74 F1       je 0afh
+00BE  3C 0A       cmp al,0ah
+00C0  74 ED       je 0afh
+00C2  31 C9       xor cx,cx
+00C4  3C 2D       cmp al,2dh
+00C6  75 06       jne 0ceh
+00C8  41          inc cx
+00C9  E8 97 00    call 163h
+00CC  EB 07       jmp 0d5h
+00CE  3C 2B       cmp al,2bh
+00D0  75 03       jne 0d5h
+00D2  E8 8E 00    call 163h
+00D5  31 FF       xor di,di
+00D7  3C 30       cmp al,30h
+00D9  72 1C       jb 0f7h
+00DB  3C 39       cmp al,39h
+00DD  77 18       jnbe 0f7h
+00DF  2C 30       sub al,30h
+00E1  88 C6       mov dh,al
+00E3  8B C7       mov ax,di
+00E5  BB 0A 00    mov bx,0ah
+00E8  F7 E3       mul bx
+00EA  89 C7       mov di,ax
+00EC  B4 00       mov ah,0h
+00EE  8A C6       mov al,dh
+00F0  01 C7       add di,ax
+00F2  E8 6E 00    call 163h
+00F5  EB E0       jmp 0d7h
+00F7  83 F9 00    cmp cx,0h
+00FA  74 02       je 0feh
+00FC  F7 DF       neg di
+00FE  8B 5E 04    mov bx,word ptr [bp+4h]
+0101  89 1D       mov word ptr [di],bx
+0103  5F          pop di
+0104  5A          pop dx
+0105  59          pop cx
+0106  5B          pop bx
+0107  58          pop ax
+0108  89 EC       mov sp,bp
+010A  5D          pop bp
+010B  C3          ret
+010C  55          push bp
+010D  8B EC       mov bp,sp
+010F  50          push ax
+0110  53          push bx
+0111  E8 4F 00    call 163h
+0114  88 C2       mov dl,al
+0116  8B 5E 04    mov bx,word ptr [bp+4h]
+0119  88 17       mov byte ptr [bx],dl
+011B  5B          pop bx
+011C  58          pop ax
+011D  89 EC       mov sp,bp
+011F  5D          pop bp
+0120  C3          ret
+0121  55          push bp
+0122  8B EC       mov bp,sp
+0124  50          push ax
+0125  53          push bx
+0126  51          push cx
+0127  E8 39 00    call 163h
+012A  31 C9       xor cx,cx
+012C  3C 54       cmp al,54h
+012E  74 12       je 142h
+0130  3C 74       cmp al,74h
+0132  74 0E       je 142h
+0134  3C 59       cmp al,59h
+0136  74 0A       je 142h
+0138  3C 79       cmp al,79h
+013A  74 06       je 142h
+013C  3C 31       cmp al,31h
+013E  74 02       je 142h
+0140  EB 01       jmp 143h
+0142  41          inc cx
+0143  8B 5E 04    mov bx,word ptr [bp+4h]
+0146  89 0F       mov word ptr [bx],cx
+0148  59          pop cx
+0149  5B          pop bx
+014A  58          pop ax
+014B  89 EC       mov sp,bp
+014D  5D          pop bp
+014E  C3          ret
+014F  50          push ax
+0150  E8 10 00    call 163h
+0153  3C 0D       cmp al,0dh
+0155  74 0A       je 161h
+0157  3C 0A       cmp al,0ah
+0159  74 06       je 161h
+015B  3C 1A       cmp al,1ah
+015D  74 02       je 161h
+015F  EB EF       jmp 150h
+0161  58          pop ax
+0162  C3          ret
+0163  B4 08       mov ah,8h
+0165  CD 21       int 21h
+0167  C3          ret