SUMMARY.md 108 KB

TP3-comp — state summary

Recreating Turbo Pascal 3.0 (shell / editor / compiler / interpreter) in GNU Modula-2 (gm2 -fiso -Wall). Faithful to the original: screen layout, keys and behaviour are reconstructed from the disassembled TP3.0 source in Resources/turbopascal3source/TP3/ (TPSRC1-10) and the reference manual, not guessed.

Milestones

Milestone Tag State
TP3.0 main-menu shell v0.1-shell done
WordStar-style editor v0.2-editor done
Shell/editor polish, Ctrl-K-D / Ctrl-K-X quit v-TP3-SHELL-EDITOR-QUIT done
Compiler skeleton + build recipe v-TP3-SHELL-COMPILES done
Parser Skip bug class (9 sites) v-TP3-PARSER-FIXES done
Standard procedures + rel16 fix v-TP3-STDPROCS done
Runtime library + 8086 execution harness v-TP3-RUNTIME-BLOB assembled, executed under qemu
Inline string literals (writeln('hi')) v-TP3-STRLITERAL done, executed
Linker: real DOS .COM writer + independent byte checker v-TP3-COM-IMAGE done, executed
Measured encodings: ModR/M table, runtime audit, golden disassembly, [BP+off] v-TP3-MEASURED-EMITTERS done, executed
Execution: a boot sector, qemu, and a claim about behaviour v-TP3-EXECUTION done, 21/21 fixtures run, exact output
Execute the nine fixtures that only compiled v-TP3-DEAD-FIXTURES done, 30/30 run; four operator/scoping bugs found
Runtime entries under qemu, and the register contract stated v-TP3-BP-CONTRACT done, 36/36 entry checks; wrchar/wrbool no longer destroy BP
8086-legal conditional branches and SETcc v-TP3-8086-LOWERING done, the emitted code no longer contains an opcode the 8086 lacks
CmdRun (the R key) + Exec86, the in-process 8086 interpreter v-TP3-CMDRUN done, a second execution oracle: 33/33 fixtures agree with qemu byte for byte, and R runs one inside the shell
Image layout in one file; a check's sweep region measured, not restated v-TP3-IMAGELAYOUT done, every reader of a linked .COM imports tests/comimage.py, and the region check_framedisp sweeps is two readings of the file required to agree
Arguments pushed as they are parsed; the parameter frame read from BP+4 backwards v-TP3-ARGORDER done, t36_argclobber runs: a computed argument survives the parse of the next one, and the callee's slots match the order they were pushed in

Every row that names a tag has one, and every tag points at a commit on master; verified by diffing the rows against git tag -l, which is how v-TP3-COM-IMAGE was found to be missing. One tag does not follow the convention used for the rest: it names the commit that did the work (7dada57, where Linker.mod, ComTest.mod and run_com_tests.sh appear) rather than a commit whose tree contains this table's account of it, because the documentation for that milestone was not written until three commits later and by then it already said "executed". The reasoning is in the tag message.

Separate piece of work

CALC-PORT.md — porting Resources/TP3/CALC.PAS (the MicroCalc spreadsheet that shipped with TP3) to GNU Modula-2 as a native Linux program, using shell/Term.mod for the screen. Analysis only; nothing ported. It is not a milestone of this compiler and is deliberately kept out of the table above. It records a measured Gm2 accept/reject matrix (calc-port/gm2-probes.sh, 13 probes, self-asserting), and the 80-byte .MCS record layout decoded from the original's own CALCDEMO.MCS — which also contradicts RESUME-TP3.md twice, on the Real exponent bias and on array iteration order. Two artefacts in this repo are the reason it is worth doing: that .MCS file makes phase 3 byte-exact against the original, and shell/tests/ptyharness.py already knows how to assert on a VT100 screen.

Build

cd shell && make          # → shell/tpshell

Whole-program link must be two-phase; a single gm2 -o pass 3 silently caps identifier/error counts on a compiler-sized program:

gm2 -fiso -c Compiler.mod                                     # phase A
gm2 -fiso -fgen-module-list=modules.lst -o /dev/null \
        Compiler.mod Term.o TextBuf.o Posix.o Editor.o        # phase B1 (rc=1 expected)
gm2 -fiso -fuse-module-list=modules.lst -o tpshell \
        Shell.mod Compiler.mod Term.o TextBuf.o Posix.o Editor.o   # phase B2

Current clean build: make clean && make → rc=0, tpshell 163616 bytes. The one diagnostic is ./Compiler.mod: ParseExpr: too many errors in pass 3, which is the expected phase-1 rollup that the recipe tolerates — not a real error. shell/Makefile is the single authoritative build recipe.

Scratch and logs live in TP3-comp/tmp/, beside the tree that produced them — never in /tmp. Every test that writes a build log, a saved copy of a mutated source, a probe binary or a corpus puts it there (../tmp/, since the shell scripts cd into shell/ first), and the folder is gitignored, so a failed run's evidence sits next to the code it describes and cannot be committed by accident. What still touches the system temp area is only the per-run working directory that tempfile.mkdtemp / mktemp -d creates — an anonymous tree, not a named file anyone goes back and reads.

shell/build_tpshell.sh is now a three-line wrapper around make, and that is a fix, not a refactor. It used to be a second hand-maintained copy of the recipe, and a copy of a build recipe drifts — this one was wrong twice over: it ran gm2 -c Posix.mod, but Posix is a foreign C module built by cc -c Posix.c and there is no Posix.mod, so it died on the third module every time; and it printed phase 2's return code and then carried on regardless, so a failed link was reported as a success whenever an older tpshell was still lying around. The copy that is wrong is the one nobody runs, which is how it survived. The wrapper keeps the old invocation working and holds no build knowledge of its own.

What is verified, and how

Nothing here is "it compiles clean" — each claim below comes from a run. tests/run_all.sh runs every check below in one pass and prints OVERALL: ALL PASS, or it prints which one did not.

Compiler front end — shell/tests/run_compile_tests.sh

tests/CompileTest.mod links Compiler + TextBuf + Posix, reads fixture paths from stdin, loads each exactly the way LoadWorkFile does (LF→CR normalisation, ^Z ends the text), calls Compile, and prints a verdict plus a source excerpt with a caret at errPos. No pty, instant.

cd shell && tests/run_compile_tests.sh            # all fixtures
cd shell && tests/run_compile_tests.sh /some/dir  # another fixture set
printf '@dump \ntests/fixtures/t19_int1.pas\n' | ./compiletest  # hex-dump (trailing space! see TP3-COMPILER.md)

The matrix asserts; it does not just count. tests/fixtures/expected.tsv pins the verdict and the numbers per fixture — verdict plus code size plus data size, or error number plus position — and the runner compares. A wrong error position or a program that lost six bytes now fails the suite instead of needing a squint. It was checked for vacuousness by reverting the string scanner fix: the suite went red with exit 1, and came back green only when the fix was restored.

34 of 37 fixtures compile clean, up from 1 (the empty program) when the direct harness was first built.

compile matrix: 37 passed, 0 failed (of 37)

Compiling: t01 minimal · t04 var+assign+writeln · t06 two args · t07 10 assignments · t08 const · t09 if/then/else · t10 while · t11 for/to · t12 repeat/until · t13 procedure + value param · t15 label + goto · t16 writeln('a') · t18 bare writeln · t19 writeln(1) · t20 3 string args · t21 mixed args · t22 case with two labels · t23 writeln('') · t24 writeln('don''t') · t26 mixed scalar/string args · t02/t03/t05/t17 multi-char literals · t27 five locals · t28 a 70-parameter declaration · t29 readln from a supplied input file · t30 a counted for whose bounds come from an expression · t31 a value parameter and a call across statements · t32 EXIT out of a for body · t33 all six comparisons · t34 the operators nothing else uses (div mod and or, unary -, a variable *) · t35 not, both halves of TPSRC9's neglevel split · t36 a call whose arguments are computed in various positions, each one keeping its own value.

Failing, all deliberately: t14 array [1..5] of integer at its point of use and t25 a string literal used as a value (s := 'hi') both → ENoLib (102), the original's "not implemented" path; uierror is a deliberate syntax error (41) used by the UI test.

run_com_tests.sh additionally links all 34 that compile to a real .COM and re-verifies the bytes with an independent Python checker that measures the layout instead of restating it: independent .COM check: 34 checked, 0 failed. That last part was itself a bug fix — see "the two restated constants" below.

Shell + editor UI — shell/tests/uitest.py

Drives a real pty (tests/ptyharness.py supplies read-until-quiet, key sending and a small VT100 emulator) and asserts the TP3 compile-error jump: W load → C compile → ESC → editor opens with the cursor on the error position → Ctrl-K D back to the menu → Q exit 0. 10/10 pass, and the reported error number is 41 (not 0 — this is what caught the errNo := errNo self-assignment where Compile's formal shadowed the module variable).

The jump mirrors original TP3 kcwait + editor2 (TPSRC5:333-336, :919): BX:=txerrpos; DEC BX; JMP editor2, and editor2 does ADD BX,txbeg; INC BX — the DEC/INC cancel, so the net is txbeg+txerrpos = our 0-based errPos. Armed as a sticky position (Editor.GotoOffset) rather than by changing Run's signature, so Editor.def stays additive.

uitest.py uses a fixture with a deliberate syntax error (x := 1 + ; → error 41, line 9 col 12) rather than a missing library feature: the latter move as the compiler grows, and a test whose expectations drift with it stops being a test.

The encodings — checks that can each go red

Nothing above looks at machine code. It all stops at "the compiler produced what it intended to produce", which is exactly where the bugs in this project live: a wrong ModRM byte is not a compile error and not a wrong code size, it is a perfectly well-formed instruction that does something else. So the encodings get their own stack, and it is built so that each check fails on a different class of mistake. tests/run_all.sh runs all of them; tests/nonvacuity.sh proves each one can go red.

check what it asserts what it cannot see
probe/run_modrm19.py the mod=00/01/10 effective addresses, by executing 23 cases on a real 8086 under qemu and scanning for where the marker landed mod=11 — see below
probe/modrm11.py the mod=11 register identities, by encoding with GNU as and decoding with FCML, against hard-coded bytes the table agreeing with itself
audit_helpers.py every one-line emitter in both Runtime.mod and Compiler.mod decodes to what its name says — 101/101, from an inventory scanned independently of the parser anything longer than one instruction
check_comimage.py a linked image's layout is described in exactly one file: comimage.py defines each part once, every reader imports it, no second find_header exists anywhere in tests/ a check that hand-rolls the layout from literals (hdr = 3 + rt_size), which defines none of those names — the readers catch that a different way: each states the layout twice, from two independent sources, and requires the two to agree
check_runtime.py + runtime.golden the built runtime's code region (436 bytes, code ends at 405) sweeps cleanly through FCML, every entry and all branch targets land on an instruction boundary, and the whole disassembly is byte-for-byte the committed golden whether the golden is right
check_framedisp.py [BP+off] uses disp8 iff off <= 127, for locals (negative) and far parameters (>127) which of the two encodings was chosen, if the other also works
run_com_exec.py the emitted image executes and prints exactly the expected bytes semantics the fixture never exercises
run_exec86.py the same image runs in a second, independent 8086 and agrees with qemu byte for byte, 33/33 the .out file it shares with run_com_exec.py — a wrong expectation fails both at once
rt_exec.py the runtime entries themselves are called directly, one qemu boot per call, 35 cases plus a pre-flight, 36/36 — plus a check_bp_contract pre-flight over the built blob: an entry that borrows BP must open with 55 8B EC (exact) and contain a 5D (a screen, because 5D is also a displacement byte) whether the caller's half of a contract is right, where the only witness is a case that happens to make two calls in a row

Two of these deserve the detail, because the reason they exist is the reason they are hard.

audit_helpers.py exists because a wrong ModRM that still decodes is invisible. The mistake this project actually makes is not a malformed instruction — it is MovSiBx emitting 89 DC, which decodes perfectly as MOV SP,BX, and CmpSiBx emitting 39 DC, which decodes as CMP SP,BX. Both shipped for a long time. A structural check passes them. A golden passes them. Only asking "does this byte sequence mean what this procedure is called?" fails, which is what the audit does: it disassembles each one-line emitter and compares the decode against the name. Five bugs came out of it in one pass.

And then the audit itself turned out to have three independent ways of silently dropping subjects, which is worse, because a check that quietly checks nothing looks exactly like a check that passes:

  1. It swept Runtime.mod only — while EmXchgAxCx lives in Compiler.mod and was wrong (93h = XCHG AX,BX) for its entire life, with a correct byte count and a green matrix. The fix sweeps both modules.
  2. The procedure-header regex required a non-empty parameter list, so it matched 0 of Compiler.mod's 22 emitters.
  3. The coverage list was derived from the parser's own output. That is vacuous: any input that stops the parser also deletes the helper from the inventory, so the audit reports "100% covered" for a file it read nothing from. The inventory is now a deliberately dumb ^PROCEDURE\s+(\w+) scan with no parameter, BEGIN or comment awareness, and the audit fails if a documented emitter is missing from it. EmitData is the one named exclusion, with the reason recorded in the source.

audit_vx_helpers had a smaller version of the same disease: it printed len(VX_KINDS) = "4 Vx subjects" for Compiler.mod, which has none. It now reports the number actually checked.

modrm11.py is deliberately not self-referential. The obvious way to check a ModR/M table is to write the table in assembly and assemble it — but that can never fail, because editing the assembly makes as faithfully re-encode the new claim and the two then agree again. That failure mode was found by corrupting the .s and watching the check stay green. Two things close it: an EXPECT byte sequence hard-coded independently of the .s text, and four anchor encodings that are spelled as literal .byte directives because as would never choose them for a mnemonic (83 C4 08 ADD SP,8 · 83 C6 02 ADD SI,2 · 8B EC MOV BP,SP · 8B E5 MOV SP,BP). No table shifted by one cell can satisfy all four.

check_framedisp.py exists because the bug it guards was accidentally correct. The old code truncated the displacement with off MOD 100H, always emitting disp8. Locals are allocated downward from 0FFFEh, so a local's offset is negative and −32768..+127 — the whole range where truncating to a byte happens to be right. The bug was only visible above +127, reachable from the 63rd parameter onward, and no fixture had one. The fix is disp8 iff off <= 127 else disp16, with no overflow branch at all: disp16 covers the entire 16-bit range as a signed value. The check also asserts the rule rather than one encoding — always-disp16 is accepted, and nonvacuity.sh proves that by building it and requiring the check to stay green.

It had a second, unrelated blind spot of its own: the region it swept came from RT_SZ = 391, a runtime size that had drifted, so the sweep began inside the runtime, part way through an instruction — and the check passed anyway, because the byte patterns it searches for are in the image wherever they happen to be and a decode starting mid-instruction happened to reach the same offsets. Nothing in the check could see where its own region began, which is a wrong input that produces the right answers until the layout moves. The region is now two independent readings of the file — where the entry jump says execution starts, where the program header says the runtime ends — required to agree before a single byte is swept, with tests/comimage.py supplying both readings and tests/check_comimage.py asserting there is only one copy of each.

Execution under qemu — tests/run_com_exec.py

The check that cannot be written as a byte comparison, so also the one that finds the most: 34 fixtures are compiled to .COM, put on a floppy, booted, and their serial output compared to a committed .out file exactly — CRLF included — plus the exit code passed to INT 21h AH=4Ch.

execution: 34 passed, 0 failed (of 34)

This is no longer the only execution oracle: run_exec86.py below runs the same 34 images a second time, in a different machine, and requires the two to agree byte for byte.

It also found bug 33 — the branch polarity inverted in every conditional in every program — while the compile matrix and the .COM layout check were both perfectly happy. That is the third time in a row that a fault passed every byte-level check in this file and was caught only here.

The two fixtures it found nothing in are the interesting ones: t31_procparam and t32_forexit were the two most expensive bugs in the project, and neither was visible as a wrong byte count. See overProc below.

The .COM layout constants are measured, not restated — and now live in one file. run_com_tests.sh and comtest.py both used to hard-code RT_SZ = 391 against a runtime that had since grown to 432 bytes at the time, so they read the program header 41 bytes early and reported 30 false failures — a red suite that meant nothing, which is the most expensive kind of red. Both now locate the header by its own signature (hdrFlag = 1, hdrDS == hdrOff + 1000h + bias, hdrHeap > hdrDS, hdrCS leaves room, initmem length at ENT_SZ) and derive rtSz, prologAt and dataBase per file:

measured runtime size: 436 bytes (header at image offset 439)

A restated constant that has drifted is worse than a derived one, and the two checkers had drifted from each other as well as from the runtime — which is why there were two of them and why both were wrong in the same way.

So the layout itself moved into tests/comimage.py, and the third reader joined the first two. Four places now read a linked image — comtest.py, the independent checker inside run_com_tests.sh, check_framedisp.py and check_8086.py — and none of them says what a .COM looks like; they import it. tests/check_comimage.py is the assertion that this stays true: one definition of each part, imported by every reader, no second find_header anywhere in tests/. Its scope is stated in its own docstring, because it cannot see a check that re-derives the layout from inline numbers — that is what the readers do instead, each stating the layout twice from two different sources and refusing to proceed when the two disagree.

8086 legality — tests/check_8086.py

The only check that asks a question about the target rather than about the compiler: does the 8086 have this instruction at all?

8086 check: 38 comparison sites, 34 lowered to a Boolean value, 13 lowered to a branch
            value conditions  : = x6  <> x2  < x5  >= x3  <= x2  > x16
            branch conditions : IF / REPEAT x9  CASE x2  FOR downto x1  FOR to x3
            runtime: 436 bytes, 219 swept, 0 0F-prefixed
            program code: 31 of 34 fixtures swept end to end, 3373 bytes
            t33_cmpops: 13 comparisons matched against their source operators, in order
            clause H: 8 of 8 fixtures matched the branch conditions read off their source

It exists because bugs 32 and 33 got past everything else, and because no execution oracle can exist for this: qemu 10.0.11's lowest CPU model is 486, so 0F 84 is an ordinary JZ there and always was. Nor could the disassembler help — FCML's -m16 mode is a 386, so the project's independent checker was guaranteed to agree with the bug.

The design is mostly a list of things that do not work:

  • A linear sweep of the code region is not a sound oracle. Inline string literals are emitted into the code stream, so the sweep desynchronises and then aborts on text — t09_if decodes 16 real instructions and dies on FE E9 0D 00, which is ASCII.
  • A bare 3D anchor is unsound. It matches displacement and immediate bytes as readily as opcodes, and using one produced two false alarms: a 3D inside a CALL displacement in t11_for, and one inside a string in t21_mixed. The anchor is the three-byte 3D 00 00 or nothing.

So it anchors on comparison sites instead of instruction boundaries: 3B C1 (EmCmpAxCx) and 3D 00 00 (EmCmpAxi (0)) are the only sequences that can precede a lowering, and all seven EmJcc sites and the EmSetcc site sit immediately after one — which was verified by reading each site, not inferred. Branches are additionally found by shape (7x 03 E9, searching for the E9), which is what covers the CASE arm whose EmCmpAxi carries a label rather than 0. Where a region sweeps clean the sweep also asserts no 0F, and the fraction it reached (28 of 31) is printed rather than implied.

"Is this an 8086 shape" is a much weaker question than it looks — a SETG where a SETGE belongs is still a fine shape — so two further clauses compare against source. G matches t33_cmpops's 13 comparisons against the operators in source order, which catches a >/>= swap. That fixture exists partly for this: =, <> and <= had never been used as a comparison anywhere in the suite, and >/>= had already been swapped once, so the coverage hole and the bug it let through were the same hole. H pins, per fixture, the conditions its branch sites declare, read off the .pas sources; its need was measured (mutation M5 came back green before H existed) rather than anticipated.

What it does not do. It cannot assert on the CASE arm's label immediate, and it never executes anything: it proves the opcodes are 8086 and that conditions are attached to the right constructs, but the control flow those bytes produce is a question for execution — and until this milestone it could only be asked of qemu, on a machine with no 8086 model. It is now asked twice.

The second execution oracle — tests/run_exec86.py

shell/Exec86.mod is a flat 8086 interpreter over the linked image, and this check runs every fixture through it as well as through qemu and requires the two to agree byte for byte:

exec86: 34 passed, 0 failed (of 34), cross-checked against qemu

Three assertions per fixture, in order of how much they are worth: the output matches the hand-derived .out exactly; the output matches what qemu printed for the same bytes; and the guest left through INT 21h AH=4Ch with code 0. The .out files are written from Pascal's semantics and are never blessed from a machine's output — run_exec86.py has no --rebless at all — so agreeing with qemu is a third opinion, not a second vote on the same one.

It exists because the 8086-legality check could not close its own gap: qemu's lowest CPU model is a 486, where 0F 84 is an ordinary JZ, so an oracle built on qemu is structurally blind to that class of fault in either direction. This one was written against the 8086's own reference instead, and nonvacuity.sh proves it can go red on its own: inverting JE inside its Cond swaps the two arms of every = in every program, while the emitted bytes, the sizes and every byte-level check stay green — and qemu, executing the unchanged image, still agrees with itself.

What it does not do: FLAGS are never written back, PF/AF are not maintained (JP/JNP fault saying so rather than answering a value nobody computed), no string instruction exists, a non-zero segment register faults, and execution outside the loaded image faults. Each is deliberate; Exec86.mod's header states the measured instruction set it implements and refuses to guess past it, because an interpreter that guesses does not fail, it answers.

The R key, end to end — tests/runtest.py

The only check that exercises CmdRun, and the only one whose evidence comes from a program nobody here has read: drive the shell through a pty, W to load t34_arith, R, and compare the guest's own output against that fixture's hand-derived .out.

UI TEST (R): t34_arith.pas
------------------------------------------------------------
R reported a successful compile                            PASS
R reported poking the image at 0100h                       PASS
guest ran to a clean AH=4Ch exit with code 0               PASS
R reported a non-zero step count                           PASS
guest output matches the hand-derived .out                 PASS
ESC after the run returned to the main menu                PASS
shell exited cleanly (status 0)                            PASS
R wrote no .COM (it runs the image where it already is)    PASS
------------------------------------------------------------
child: EXIT 0
RESULT: ALL PASS

The last assertion is the one no other check can make: R writes no file, so the .COM on disk must be exactly as it was before — nothing else in this project observes R at all. It is proved able to fail by neutering the poke loop's count: with n = 0 nothing is copied in, Run86 faults on its first step (execution left the loaded image), and the guest's AH=4Ch assertion goes red before a single instruction has run.

Non-vacuity — tests/nonvacuity.sh

Every assertion in this file is proved able to fail: 62 deliberate breakages, each asserted to turn exactly one named check red for the stated reason, then restored and re-asserted green — non-vacuity: 62 ok, 0 failed. They cover the runtime's emitter audit and its restored source, the mod=11 table (including restoring the exact wrong table this project once shipped), the [BP+off] rule, the behavioural bugs, the emitter-name audit of Compiler.mod, the BP contract, the .COM layout checker, the image-layout helper and the three checks that read it, the 8086 lowering, the interpreter itself, and the R key.

One assertion is proved outside that suite, and the difference is stated rather than papered over: t36_argclobber's .out was compiled and run against the compiler as it stood before the argument-order fix, on both oracles, and came back red with every line whose arguments included a computed value wrong while the two lines built from plain names stayed correct — so the two that would catch a one-sided flip were seen green before the fix, not after it. The evidence is quoted in the commit that introduces the fixture, and green followed only with the fix in. What is not there yet is the in-suite form of that proof: reverting each of the three call parsers and the parameter-offset remap in turn and requiring run_com_exec.py t36_argclobber red. That is the first of the next steps, and until it exists the count above does not cover it.

Eight arrived with the previous milestone (v-TP3-CMDRUN), and each one is the same shape: code that compiled clean and passed every byte-level check, until it was run. Two are behavioural (below), two come from the new interpreter, two pin the two new grammar rows the helper audit gained for EmXorAl01, and two are the R key's mutation and its restored green.

Ten arrived with v-TP3-IMAGELAYOUT, and they are a different shape: a check whose own input had never been verified. check_framedisp swept a region built from a literal that had drifted, and the region is now two independent readings of the file required to agree — so the first three cases break each reading in turn: the entry jump answered with the old literal, and the header's own equation one byte out, seen red by two different readers with two different sentences. The next four break the single copy itself — a find_header copied out of comimage.py, a layout constant copied out of it, the definition deleted (which must be a red report rather than a silent green one, because a rule that matches nothing looks exactly like a rule that passes), and a caller that reaches the helper through another consumer instead of through comimage. The last three are restored greens, one per reader.

That last group is the newest and the least optional. Two of its five (M4, M5) invert the branch polarity, which is a legal 8086 opcode sequence, so no shape-based check can see it and only a source-derived expectation can. M5 came back green on its first run, which is why clause H of check_8086.py exists at all; see bug 33 above.

Four properties of the harness itself are enforced, because all four had already gone wrong silently:

  • A baseline assertion runs first, so a case that is already red is reported as NOT NON-VACUOUS and distinguished from one that went red. Without it, a harness broken by an earlier case would make every later case look like a success.
  • Every source mutation goes through mutate, which asserts the file actually changed. Four cases were found to be dead this way: a sed that matched nothing, a helper that had been renamed, a helper that had been reformatted, and a checker that correctly stayed green because the breakage it looked for was no longer the breakage the checker hunts. Two of the four (MovAlDh, StBxDl) were repaired rather than deleted.
  • The python mutation helpers fail LOUDLY. A non-zero exit from a helper makes if mutate_foo; then false, so the case is silently skipped — and a skipped case and a passing case are indistinguishable in the total. This is not hypothetical: the SETcc case lost its first run to an apostrophe in an assert message that closed a Python string, python died, the compiler was never broken, and the suite still reported 0 failed. Each helper now echoes a BROKEN CASE line and increments fail.
  • Each helper counts its targets before replacing. assert s != before only proves the file changed; with two edits it passes if either landed, and with two identical HideLocals call sites it would delete the wrong one — still a changed file, still a working compiler, still green for the wrong reason. So assert s.count(X) == 1 everywhere.
  • A failed rebuild is a FAILURE, not a skip. The 8086 group's first draft rebuilt with make but ran check_8086.py, which links against the comtest binary that make does not build — so the stale compiler was measured and the first mutation was reported as VACUOUS. The helper now checks the build and says so.

The one that produced the most information was restoring the original shifted mod=11 table: it turns three cells red rather than one, because the error is invisible at code 100 and only visible from 101 down. See below.

The executor problem (SOLVED — images boot and run)

The whole point of a Pascal→8086 compiler is that the output runs, and for the first four milestones it did not: no compiled image and no runtime entry had ever been executed on a CPU, correct or otherwise. Everything in the checks above was a claim about bytes. Whether the bytes work was the untested part, and it was the only part that could not be closed by writing another checker.

It is now closed, and the milestone is v-TP3-EXECUTION. 21 fixtures compile to .COM images, are booted on a floppy by a 512-byte hand-assembled boot sector, and are compared byte for byte against the exact output the fixture demands — tests/run_com_exec.py, wired into run_all.sh, 21/21 at that tag and 33/33 now.

Everything below is about establishing what can be believed, because the first attempt at this used an emulator that was wrong, and a wrong oracle is worse than none: it cannot distinguish "my codegen is broken" from "the machine is broken".

Unicorn 2.1.4 cannot be used for this. UC_MODE_16 mis-decodes 16-bit ModRM memory operands. The measurement, by loading a byte-pattern image so a load reveals its own effective address, then executing a single LEA AX,[r+disp8] and reading AX back with every register set to a distinct value:

mod=01, disp8=4            got        8086 says
  8D 40 04  LEA AX,[BX+4]  AX=0d04    0x504   (= BX+SI+4)   WRONG
  8D 41 04  LEA AX,[BX+SI+4] AX=0e04  0xd04                WRONG
  8D 42 04  LEA AX,[BX+DI+4] AX=0f04  0xe04                WRONG
  8D 43 04  LEA AX,[BP+4]  AX=1004    0x704                WRONG
  8D 45 04  LEA AX,[DI+4]  AX=0904    0x904                right
  8D 46 04  LEA AX,[BP+4]  AX=0704    0x704                right
  8D 47 04  LEA AX,[DI+4]  AX=0504    0x904                WRONG
mod=00 / mod=10 direct disp16
  8B 1E 00 20  MOV BX,[2000]  BX=1234                     right
  8D 06 34 12  LEA AX,[1234]  AX=1234                     right

So rm=5 and rm=6 decode correctly but rm=0,1,2,3,7 do not, and only base-register-free addressing (direct disp16) is trustworthy. That rules Unicorn out as an oracle for exactly the instruction forms generated code is made of — LEA AX,[BP+d], MOV AX,[SI+d], LODSW-style loops, everything with a frame pointer. It cannot distinguish "my codegen is wrong" from "the emulator is wrong", which makes it worse than no emulator at all.

pip install --upgrade unicorn resolves to the same 2.1.4, so this is not avoidable by upgrading.

This is no longer the whole picture: qemu-system-i386 is now a trusted oracle, and it was trusted by measurement, not by reputation. modrm19.s assembles with GNU as, is wrapped in a 512-byte boot sector, and is booted under /usr/bin/qemu-system-i386 with the serial port captured to a file. For each of 24 ModR/M encodings it stores a marker through the encoding under test and then scans memory for where the word landed, with BX=1000 DI=2000 SI=0030 BP=0040 so every candidate address is distinct. The answers come out as raw offsets, so this is arithmetic, not a judgement call, and the 23 cells it covers are the project's ground truth for effective addressing. It is also how we know qemu's 8086 is right where Unicorn's is wrong, on the same instruction class, by the same method.

Two facts about the tooling came out of that work and are worth recording because both were believed wrong at first:

  • objdump -D -b binary -m i8086 disassembles 16-bit code correctly. An earlier note in this file said there was no usable 16-bit disassembler available; that was wrong, and the cost of believing it was a hand-derived ModR/M table. FCML (fcml-disasm -m16) is the other decoder and the two are cross-checked by tests/fcml_vs_objdump.py. Note fcml-disasm linear-sweeps and aborts (rc=134) on inputs of 16 bytes or more through the Debian wrapper, which is why tests/disasm16.py windows input at 15.
  • A qemu execution probe for mod=11 is structurally impossible, not merely awkward. The comparison register is itself a candidate target, and SP is destroyed by the next call before any check can run — the first version of that probe pushed its return address through SS:0xBEEF, so the evidence it was about to collect had already been overwritten. It also "cleared AX because AX is never an r/m target", which is true of the wrong table and false of the right one. So mod=11 is measured by encoding instead, with as as an oracle independent of both the runtime and qemu.

Also ruled out, for the record:

  • DOSBox-X 2025.02.01 (installed, /usr/bin/dosbox-x, plain root-owned ELF, not a snap) starts cleanly headless under SDL_VIDEODRIVER=dummy SDL_AUDIODRIVER=dummy, but its -c/autoexec commands never observably execute: a md never appeared on the host, and a .COM that creates OUT.TXT via INT 21h AH=3Dh/40h never produced the file. Shell > is intercepted by dosbox-x's own wrapper (SHELL:Redirect output to out.txt). A .BAT route timed out. The user has since installed FreeDOS (freedos.qcow2, FD14-LiveCD) which is very likely the answer to this, and untried.

The boot sector — tests/exec/bootcom.s

The missing piece is now written. A 512-byte boot sector assembled with GNU as and objcopy-ed onto a 1.44 MB floppy image: it loads sector 0, reads the .COM off the same floppy with INT 13h AH=02h (disk geometry C0 H2 S18, 512-byte sectors, so .COM byte n is at disk offset 512 + n — file offset 512 maps to memory 0x100), sets DS=ES=0, SS=2000h, SP=2004h, builds the three GDT descriptors at 2000h/2002h/2004h (code, data, stack) in the way a .COM expects, installs an INT 21h shim that routes AH=02h/09h/4Ch/08h to the serial port at COM1, and JMP 0000:0100. qemu is invoked as qemu-system-i386 -fda disk.img -serial out.txt, and the fixture's expected text is compared to out.txt exactly — including CRLF, and including the exit code the program passes to INT 21h AH=4Ch.

The INT 21h shim has two distinct epilogues for one hook, which is a fact about the 8086 and not a design choice: AH=08h (read with no echo, EOF) must be able to return CF=1 with AL=0, so it exits through a different path (.Ldone8) from the AH=02h/09h/4Ch case. And the runtime's INCUR is an index into the input buffer, not a pointer, so the shim's EOF comparison is an index comparison — getting that wrong produces an "EOF at the first character" bug that looks exactly like a broken readln.

run_com_exec.py --show prints the serial file, so a failing fixture can be told apart from a broken harness without re-deriving anything.

The runtime, entry by entry — tests/rt_exec.py

run_com_exec.py proves emitted programs behave. This harness proves the runtime entries themselves do, which is a different question: a program never calls rdint without a readln behind it, never calls rdbool twice in a row, and never exercises the INT16 extremes individually.

35 cases, one fresh qemu boot each, plus a pre-flight that runs before any of them: initmem (data poisoned to AA first, so a zero is a result and not a leftover), wrint ×10 including both INT16 extremes, wrchar ×2, wrbool ×3, wrln, stackchk, a composed writeln(42) writeln TRUE sequence, rdint ×6 against supplied input, rdchar ×2, rdbool ×6, rdint at EOF, and wrtinl.

Three properties are load-bearing, and each of them was chosen against the easier alternative:

  • One boot per case. A machine that has already run a case has already run a runtime entry, and if that entry corrupts something the next case inherits it. About a tenth of a second a boot buys a machine that has provably never executed anything, and 35 of them cost under four.
  • One case record, written twice. rt_exec.py and exec/rtdrv.s each state the 40-byte layout, and the cases only pass if the two agree on every offset. A driver that silently read the wrong field would still boot, still run, and still print — so the duplication is the check.
  • A BP-contract pre-flight, before any machine starts. See below.

The harness reuses tests/exec/bootcom.s, the same 512-byte boot sector run_com_exec.py uses, so the boot machinery is written once.

Two expectations were wrong, and both were wrong in the harness's favour

Recording this because the instinct on a red test is to suspect the code, and in both cases the code was right:

  • rdchar stores one byte (StDiDl = MOV [DI],DL), so the old check dumped a two-byte word and compared it against ord(c). It could never pass. The case now reports two bytes and expects the character followed by the 0xEE poison still sitting in the second.
  • rdint at end of input stores nothing — TP3's xrdint returns to rnerr without touching the variable, and EmitRdInt's own comment says so. The old expectation said the variable was set to zero.

wrtinl also needed a harness change rather than a machine change: that entry's argument is not a stack word, the caller must place a length byte and the characters at the return address, so testing it also tests the encoding contract between Compiler.IoCall and Runtime.EmitWrInl.

What execution found that no byte check could

A 10-byte corruption of the BIOS data area. The boot sector DMA'd the image straight to 0000:0100, which is what DOS does — and 0400h-04FFh is the BDA, where SeaBIOS keeps live state. The transfer destroys it, SeaBIOS writes part of it back after the DMA, and ten bytes of BIOS data end up on top of the image. The read sets CF=0 and returns success. The image is the right length in the right place and ten bytes of it are wrong.

It was found by a 19-byte probe that read 0440h as its first instruction after the jump and got the BIOS's own bytes back, and then localised by dumping the whole loaded image against a pattern: one 10-byte run wrong inside an otherwise byte-perfect sector, which no partial-read or sector-count bug can produce. The fix is in bootcom.s: read to 8000h — clear of the IVT, the BDA, SeaBIOS's stack at 700h and the ROM window at C000h — and then REP MOVSW down to 0100h. The copy is executed code, so it is the last thing that touches the image and no BIOS call follows it.

CmpAl (20) was decimal twenty. rdint's lead-in skip is CmpAl (20H), JBE — skip everything at or below a space. Written as 20, Modula-2 read it as decimal and emitted 3C 14, so a leading space was never skipped and the scan ended with nothing read. Every other magic number in Runtime.mod is now written as hex (MovAh (02H), CmpAl (0DH), CmpAl (1AH)), because a literal that reads like hex but is decimal is silent.

wrchar and wrbool destroyed BP. Both borrow BP to reach their argument — [SP] cannot be encoded in 16-bit mode, so BP stands in — and neither saved it. BP is the one register an entry may keep, because the driver keeps its cursor into the case record there. So wrchar sent the next call to a garbage address, the machine triple-faulted, SeaBIOS rebooted, and the run printed the record header twice and hung. Symptom: record was never closed (EOT).

The interesting part is that every existing check called that shape correct. The bytes were well formed, the size did not change, the golden matched, all branch targets were on instruction boundaries, and the name audit said every helper emitted what its name said. check_runtime.py's entry golden had blessed it in five bytes: "wrchar": "8B EC 8A 46 02 89 EC". A positional golden blesses whatever is there.

So the rule is now stated, in two places that can disagree:

  • check_runtime.py's entry goldens carry the PUSH BP and the POP BP. What the bytes must be.
  • rt_exec.py has a check_bp_contract pre-flight over the built blob. What the bytes must mean. Its two halves are not equally strong and the code says so: "must start with 55 8B EC" is exact, and "must contain a 5D" is a screen, because 5D is also a displacement byte and this code does not disassemble. Requiring the POP to be contiguous with anything else is not an improvement — wrbool legitimately closes its frame after the INT 21h, and a rule insisting on 89 EC 5D fails a correct entry, which is worse because it teaches a reader to distrust the check.

And a third consequence, in the emitters themselves: reaching the argument at [BP+2] versus [BP+4] is a one-byte difference that reads as a plausible character, and the first version of the fix passed that displacement as a Modula-2 parameter — which is invisible to audit_helpers.py, because the audit reads the name. The emitters are now MovAlArg2/MovAlArg4 and CmpArg2W0/CmpArg4W0, with the displacement in the name, so the audit pins it: a helper called CmpArg4W0 that emitted the +2 form fails the run with "step 2: displacement is 2, name says 4".

The in-process interpreter — shell/Exec86.mod, and the R key

TP3's R runs a .COM in place, without the DOS loader, from the same 64 KB DOS would give it. Exec86 is that machine: a flat 8086 over ARRAY [0..65535] OF CARDINAL of bytes (one element per byte), with three entry points and nothing else:

Clear86                        wipe all 64 KB; AX..DI := 0; SP := 0FFFEh;
                               IP := 0100h; segments := 0; flags cleared;
                               loadHi := 0100h
Poke86 (addr, value)           mem[addr] := value MOD 256, and raise loadHi
Run86 (VAR exitCode : CARDINAL;
       VAR steps : LONGCARD) : CARDINAL
                               0 = the guest halted through INT 21h AH=4Ch,
                               1 = fault, 2 = the step limit was reached

CmdRun does what the original's krungo does with no loader: Compile → Clear86 → poke LinkSize() bytes at 0100h → Run86 → report the status and the step count → WaitEsc. No file is written, so R is identical with Destination = Memory and Destination = .COM, and the guest's own AH=4Ch exit code is printed rather than discarded.

The guest writes to this process's fd 1 through the runtime's INT 21h AH=02/AH=09/AH=08, so its output lands between the two status lines instead of being collected and replayed: Term writes one byte per write(2) and buffers nothing, which is why the two streams stay in order.

Why Clear86/Poke86/Run86 rather than Clear/Poke/Run. ISO Modula-2 has no import renaming (FROM M IMPORT x AS y is rejected) and no procedure-local import, and the shell's flat namespace already contains TextBuf.Clear and Editor.Run. The 86 suffix is the fix, applied identically in Exec86.def, Exec86.mod, tests/Exec86Run.mod and the module-level import in Shell.mod.

What it deliberately does not do, each stated in Exec86.mod's own header: FLAGS are never written back, so INT 21h "preserves the flags" by construction (a real INT pushes them and IRET pops them, so the handler's CLC/STC are discarded — bootcom.s relies on that); PF/AF are not maintained, so JP/JNP fault instead of answering a value nobody computed; DF/IF/TF are not maintained and no string instruction exists, so nothing can read DF; a non-zero segment register faults before every step, because a non-zero segment would silently alias onto the same 64 KB instead of faulting; an unknown INT 21h function faults; and MaxSteps = 2000000000 catches a runaway. Execution outside [0100h, loadHi) faults, which is what makes the R non-vacuity case fail on its very first step.

The instruction set is not "the 8086" but the measured union of (a) every byte Runtime.mod emits — 219 instructions, swept by check_8086.py — and (b) every byte Compiler.mod's Em* procedures can write (the generated region cannot be swept: inline string literals desynchronise a sweep, so the authority there is the source). Everything outside that union faults rather than guessing, because an interpreter that guesses does not fail — it answers.

Components

Shell — shell/Shell.mod, Term.mod, Posix.c, TextBuf.mod

Exact TP3.0 screen (Logged drive, Active directory, Work file, Main file, Edit Compile Run Save / Dir Quit compiler Options, Text: n bytes, Free: n bytes, >), command letters drawn bold. Keys L A W M E C R S D O Q; any other key redraws the menu, like TP3. W auto-.PAS with Loading/New File; S writes ^Z EOF and rotates the old file to .BAK (unlink+rename, TP3's order); D is a DOS *.* glob listing with k bytes free; O is the options submenu; Q confirms and prompts to save when the text changed.

E, C and R are wired. R compiles, pokes the linked image into the in-process interpreter at 0100h, runs it and reports the guest's exit status and the step count — see The in-process interpreter above.

Editor — shell/Editor.mod

WordStar-style full-screen editing over TextBuf. Status line Line n Col n Insert/Overwrite Indent X:FILENAME; text on rows 2-24.

Movement (^S/^D/^E/^X/^A/^F/^R/^C/^W/^Z, ^Q S/D/E/X/R/C/B/K/P, arrows, PgUp/PgDn/Home/End), editing (^V insert/overtype, ^G delete char, backspace joins lines, ^T/^Y word/line, ^Q-Y to EOL, ^N/CR break, TAB auto-indent to the word start above), block (^K B/T/H mark, word, show, ^K C/V/Y copy/move/delete, ^K R/W file read/write), search (^Q-F, ^Q-A, ^L repeat; options B/G/n/U/W, N = no-confirm; ^A any char, CR LF matches a line break), ^P literal control char, ^U/ESC aborts a prompt. ^K-D returns to the shell with the text still in memory and changed reported. Disk format matches TP3: CRLF plus trailing ^Z, load normalises, save re-expands.

Detail: TP3-EDITOR.md, TP3-EDITOR-PSEUDOCODE.md.

Compiler — shell/Compiler.mod

Single-pass Pascal → 8086, following TPSRC6 turbo / TPSRC7-10: one pass over the shared TextBuf emitting machine code into cbuf, a patch list for forward references, TP3-style error reporting (number + relative position), and code/data size accounting. Emitted image is a byte array (mode word, CS/DS, size words, CALL initmem, MOV BP,SP, generated code).

Working subset (v0.5): integer/char/boolean/byte scalars, constants with folding, globals, locals, value parameters, procedures and scalar-result functions, ARRAY[const..const] with constant indexing, control flow, GOTO/EXIT, the standard procedures WRITE, WRITELN, READ, READLN, HALT, and inline string literals as WRITE/WRITELN arguments.

Standard procedures are KBuiltin, not KProc, because they are not called generically. TP3 (TPSRC8 pwriteln/pwrloop/prdtyped) does not pass a descriptor to the runtime: it inspects each argument's class and emits a different call per type, so formatting is fixed at compile time and the runtime only ever sees a value. IoCall mirrors that — one call per argument, then a final call for the line break:

writeln(1)      MOV AX,1     ; PUSH AX ; CALL 20H ; ADD SP,2 ; CALL 40H
writeln('a')    MOV AX,'a'   ; PUSH AX ; CALL 28H ; ADD SP,2 ; CALL 40H
writeln('hi')                 ; CALL 70H ; 02 'h' 'i' ; CALL 40H
readln(x)       LEA AX,[0104]; PUSH AX ; CALL 48H ; ADD SP,2 ; CALL 60H

(Those TU_* names are the compiler's own; the offsets behind them are now assigned from Runtime.RT_Entry rather than written down, so the placeholder-versus-real distinction is gone. The third line is the inline-literal form, which differs in kind: no value is pushed and the ADD SP,2 is absent, because the length and the characters are the argument.)

READ/READLN push the address so the runtime can store (EmPushVarAddr: LEA AX,[BP+off] via EmBpDisp for locals, 8D 06 off for globals); a non-variable argument is ETypeErr (56), as in TP3.

Detail: TP3-COMPILER.md.

Runtime library — shell/Runtime.mod, shell/Runtime.def

The 8086 runtime, assembled byte by byte from Modula-2 — no external assembler and no checked-in binary, so the tree stays self-contained and the entry offsets are derived rather than guessed. Each emitter is one instruction with its ModRM byte spelled out in a comment so the encoding can be checked by hand against an 8086 table. This is the same approach Compiler.mod already takes (Ebyte/Eword/EmCall).

It mirrors the original's own mechanism: TPSRC7 copyrt copies the runtime into the front of the code buffer (SI=DI=0, REPZ MOVSB) and pc is then initialised past it (MOV pc,#$2D7C). Same shape here, and now actually done: pc := RT_Size, dc := RT_Size + 1000H, so the image is [runtime][program header][program code] and every emitted address is image-absolute. No relocation pass is needed — worth having paid for, since a linker that has to walk fixups is a linker that can get them wrong.

Current blob: 436 bytes, 14 entries, 38 branch targets, offsets read back out of the assembled bytes by tests/check_runtime.py rather than asserted by hand:

entry offset entry offset entry offset
initmem 0 wrint 36 rdint 171
progend 28 wrchar 97 rdchar 279
stackchk 35 wrbool 111 rdbool 302
halt 28 wrreal 136 rdln 350
wrln 144 wrtinl 152

Every number in that table is measured, and that is the point: it was last correct when the blob was 391 bytes, and the entries quietly moved as it grew to 436 — six of the fourteen offsets in the previous version of this table were stale, and the size was stale too. A restated constant that has drifted is worse than a derived one, which is why the compiler asks Runtime.RT_Entry instead (below) and why the checkers locate the runtime by its own signature rather than by a literal.

The TU_* constants in Compiler.mod are assigned from Runtime.RT_Entry in Inittur, not written down, so a runtime edit that moves an entry cannot leave the compiler calling the old address.

progend and halt deliberately share one address (XOR AX,AX / MOV AH,4C / INT 21h / RET): the compiler already zeroes AX before progend and discards the HALT argument at compile time, so both leave with exit code 0.

Conventions, matching Compiler.IoCall exactly:

entry argument notes
WrInt/WrChar/WrBool/WrReal one 16-bit value on the stack caller pops
RdInt/RdChar/RdBool one address on the stack caller pops
WrLn/RdLn/StackChk nothing
InitMem AX = offset of the program header a register, not a stack word
ProgEnd/Halt nothing exits, code 0
WrInl nothing — reads its own text via POP BX see below

InitMem receives the program-header offset in AX (not on the stack), reads the data base and end out of the header, and zeroes that range, because Pascal leaves globals undefined. StackChk is a bare RET — range and stack checking aren't compiled in yet, and the call site sits mid-expression, so it must not touch a register.

Five emitter bugs were fixed here in one session, all found by the audit_helpers.py name-vs-decode pass described above, and all of them were invisible to everything else that was already in place:

emitter emitted actually was correct
MovSiBx 89 DC MOV SP,BX 89 DE
CmpSiBx 39 DC CMP SP,BX 39 DE
MovSiAx 8B C0 MOV AX,AX (a no-op) 8B F0
initmem zeroing loop MovAxDx loaded a value it then discarded XOR AX,AX
initmem header read +8 hdrMax +6 (hdrHeap)

The third is the instructive one, because it is the same error pointed the other way. 8B C0 reads as MOV AX,AX under the shifted ModR/M table this project shipped, and the fix looked like it should be the byte that table said was SI. It is not: for opcode 8B the reg field is the destination, so 8B F0 (reg=110=SI, r/m=000=AX) is MOV SI,AX, which is what the name asks for. Getting the direction backwards nearly caused a correct fix to be reverted.

Two label-name plus fixup list: rel8, rel16 and runtime-data addresses are all patched after the blob is placed, so nothing depends on a hand-computed displacement.

It executes. Every entry is called directly under qemu by tests/rt_exec.py — 36 cases, all passing — with the register contract each entry must honour stated rather than assumed, which is how wrchar and wrbool were caught destroying the caller's BP. That milestone (v-TP3-BP-CONTRACT) is the reason this paragraph no longer says the runtime has never run; it did, for two releases. What is not covered is what a direct call cannot reach: the runtime is still measured entry-by-entry rather than against TP3's own binary, so this is a compatibility claim and not a byte-for-byte reproduction of the original's.

The program image — shell/Linker.mod

The runtime is copied to the front of the code buffer, then the program header, then the program. The header is our own format at rtSz:

off field
+0 hdrFlag 1, "header present"
+2 hdrCS
+4 hdrDS
+6 hdrHeap = dc, the end of the data area
+8 hdrMax
+10… max-open-files, input buffer, output buffer words

InitMem gets the header offset in AX and reads +6 for the heap limit; the independent checker in run_com_tests.sh restates these offsets as its own constants and asserts initmem's SI displacements equal them, assertion by assertion, rather than asking the compiler where it thinks the header is.

CmdCompile honours the Destination option: 0 = memory, 1 = .COM, 2 = .CHN (refused). A .COM is named after its source with the extension swapped at the last dot, padded with a zero gap to max(pc, dc); its stack sits at the segment top (SS = SP = CS:FFFE), which is where DOS puts it.

Known limit: the data area starts at a fixed rtSz + 1000H (436 + 4096 = 4532), so a program whose code exceeds 4 KiB runs into its own data. Every fixture's image is between 4539 and 4547 bytes. Documented rather than fixed, because the original has the same fixed-offset behaviour.

String literals — writeln('hi')

writeln('toto') was the last thing standing between the front end and a hello-world: string literals had no encoding at all, so anything past one character was ENoLib. The encoding is not invented — it is the original's.

TPSRC8 pwrinlin peeks at the character after the literal: if it is , or ) the literal is a WRITE argument, not an expression, and it emits (TPSRC10 estring) the length byte and the characters into the code stream right behind the call:

CALL wrtinl   <length byte> <character>...

TPSRC4 xwrtinl is what makes that self-delimiting: POP BX takes the return address — which is the address of the length byte — and the entry ends with JMP BX, returning to just past the last character. So the literal needs no terminator, no length table, and nothing at all in the data segment. The arithmetic confirms it: t26 emits 02 68 69 inline and its data size is unchanged from a program with no strings at all.

wrtinl is 19 bytes at offset 148 (5B POP BX · 31 C9 XOR CX,CX · 8A 0F MOV CL,[BX] · 43 INC BX · B4 02 MOV AH,2 · E3 07 JCXZ to the end label · 8A 07 MOV AL,[BX] · CD 21 · 43 · E2 F9 LOOP · FF E3 JMP BX) — hand-checked once, and now also covered by the golden disassembly and the audit.

A string literal is a value in exactly one place: a WRITE/WRITELN argument. Everywhere else it is a hard error, and it is enforced in a single place — LoadAtom — because assignment, IF, WHILE, FOR, REPEAT, CASE, array subscripts and every operator all reach their operand through LoadAtom, and none of them can use a counted string where a 16-bit word is expected. ParseFactor therefore marks a literal (kind = 3) rather than rejecting it, and IoCall handles it before LoadAtom is ever reached. The alternative — letting it through and producing a machine word that happens to be a pointer — would be a silently wrong program; t25 pins the error instead.

Literal text has to survive from the scan to IoCall, since the parser does not yet know it is writing rather than computing, so it is collected into a pool as it is read: strPool[0..4095], strOff/strLen[0..255], strTop, strCnt, plus StrNew/StrPut. The pool is reset in Inittur, so it is per-compilation.

Details that are deliberate, not incidental:

  • '' is a zero-length literal, reaching the runtime's JCXZ path. It used to be the scalar 39, so writeln('') printed a quote mark.
  • 'don''t' — the doubled quote becomes one character; t24 pins len 5.
  • A literal of 256 characters or more is EConstRange (45), not truncated. The length is one byte, so 300 characters would go out behind a length of 44 and the runtime would print 44 of them and silently drop the rest. TP3 strings are at most 255 characters, so refusing is the faithful answer.
  • A string variable is ENoLib, not wrong code. IoCall knows the difference and refuses. EmPushVarAddr's local form is fixed now (see bug 4 below), so the blocker is no longer the encoding — it is that there is no string type, no length word, no assignment path and no WrStr entry.

Honest limitations

  • runtest.py runs one fixture through R, not the whole suite. It drives t34_arith through the pty because a pty test is expensive and this one's job is to prove the path exists (compile → poke → run → report → no file written), which it does with eight assertions. The other 33 are covered by run_exec86.py, which calls the same interpreter directly.
  • Every fixture that compiles is now also run. 34 of 37 execute; the 3 that do not are t14 and t25 (ENoLib, by design) and uierror (a deliberate syntax error). The gap this replaces was nine fixtures that compiled and were never executed — t08 const, t09 if/then/else, t10 while, t11 for/to, t12 repeat/until, t13 procedure + value param, t15 label + goto, t27 five locals, t28 the 70-parameter declaration — and all nine contained no write/writeln at all. They assigned to a variable and fell off the end, so the only thing an empty .out could have asserted was "did not crash". That is why t12's loop ran exactly once for the whole life of the project with nothing to see it. The nine now print, each .out is derived by hand from the Pascal rather than from the machine, and nonvacuity.sh mutates each of the four bugs back in and requires the execution check to go red. Finding them cost four fixes; see the milestone section.
  • A call takes at most 16 arguments, and says "compiler overflow" if you exceed it. Each of the three call parsers counts as it parses and raises ECompOvf = 99 at the seventeenth — there is no args array left to overflow, since an argument is pushed the moment it has been parsed. So far (1, 2, ... , 70) is rejected with error 99, whose text in TP3 means the compiler's own table overflowed — an error about the compiler, for a program that merely has a long argument list. This is why t28_farparam can declare 70 parameters (which is what puts p8 at [BP+128] and exercises the disp16 encoding) but can only pass 16, and why its body reads p8/p1 into a variable whose value is deliberately absent from the .out: those two slots hold stack garbage, and a fixture that printed them would be testing the harness, not the compiler.
  • Parameters are separated by , and only by ,. procedure two (a : integer ; b : integer) is error 1 at the semicolon; procedure two (a : integer, b : integer) compiles. The reverse holds for variable declarations — var a, b : integer is error 1 at the comma and needs two var lines. Inconsistent, and neither spelling is wrong Pascal, so a program that compiles under one compiler may not under another.
  • A parameter may not shadow a global. DefProc calls DupTest on every parameter name, and DupTest reports error 41 for any name Search finds — including one declared at an outer level. procedure bump (x : integer) with a global x is rejected. Pascal allows the shadow and the inner one wins. HideLocals fixed siblings colliding with each other; this is the parent-vs-child direction of the same question and is untouched.
  • qemu-system-i386 cannot execute 8086 code, so the 8086 checks are static and stay static. The lowest CPU model qemu 10.0.11 offers is 486 (-cpu help confirms it; there is no 8086 model to ask for). Every image in run_com_exec.py therefore runs on hardware that did not exist when TP3 shipped, and no amount of fixture work changes that. This is why tests/check_8086.py exists and why it had to be written as a byte-level question rather than a behavioural one — and why its clauses G and H compare against source rather than against another run of the machine. It remains a weaker instrument than execution: it proves the opcodes are 8086 and that the conditions are attached to the right constructs, but the control flow those bytes produce is still only checked by qemu on a 486.
  • The 8086 check cannot assert on the CASE arm's comparison. CASE compares against a label, so its EmCmpAxi is 3D lo hi with a non-zero immediate, and no sound anchor can find it — a bare 3D also matches displacement and immediate bytes, which produced two false alarms before the 3-byte 3D 00 00 form replaced it. The arm is covered by shape (7x 03 E9, searching for the E9) and by t22_case's execution, but the label immediate itself is unchecked.
  • The runtime's entries are now each called directly, and two of its invariants are stated rather than implied. 436 bytes, 14 entries, 101 emitter helpers decoded against their own names across both modules, the whole code region golden-pinned, all branch targets on instruction boundaries — and 35 cases that call the entries one at a time under qemu (rt_exec.py, 36/36). What is not covered is what a direct call cannot see: wrtinl is checked, but only from the harness's side of the calling contract. (The nine fixtures that compiled but were never executed used to sit here as an extra gap; they are now run, so every fixture that reaches the runtime at all also reaches it through generated code.)
  • The pushback slot's address is a moving target. It sits at rtSz + LoadBias + dataAt + D_PUSH, so every runtime growth moves it, and every address derived from it must be recomputed. It is computed, not hard-coded — but the two RT_SZ constants that were hard-coded and had drifted (see the execution section) are the precedent for why this one gets stated every time.
  • 33 executed fixtures is a small sample of Pascal. They cover var, const, if, while, for, repeat, case over scalars, procedures with value parameters, goto/label, string literals, readln, all six comparisons, and the arithmetic and logical operators * + - div mod and or not on variables. They do not cover nested procedures, recursion, var parameters, with, records, sets, files, reals, or any type wider than 2 bytes — all still ENoLib. A green execution matrix says nothing about those.
  • wrreal is a deliberate stub. It writes the literal text ?REAL? — the string lives in the runtime's own data block at D_REAL=24, which is what makes it a real 9-byte routine rather than a trap. Reals are not formatted yet, so writeln(1.5) "works" and prints nonsense. A trap would be louder; neither choice is a real answer, and this is now reachable code rather than an unreachable one, which raises the stakes on the choice.
  • Code above 4 KiB overruns the data area. The data base is fixed at rtSz + 1000H = 4432 and a .COM is padded to max(pc, dc), so the fixed 4 KiB code window is real and not advisory. The largest fixture (t32_forexit, 135 bytes of code) is nowhere near it, so nothing has hit this and nothing tests it. rtSz also moves every time the runtime grows, so the window shrinks silently — another derived value that must never be restated.
  • [SP] cannot be encoded on an 8086, and the fix is a shape change. EmMovAxSp/EmMovCxSp now emit POP reg / PUSH reg — two instructions where there used to be one, so anything that assumed a one-instruction emitter has to be revisited. The audit handles them as a documented sequence; if another one appears, the sequence machinery is where it goes, not a name hack.
  • String literals work; string variables do not. A literal in a WRITE/WRITELN argument list is emitted inline and needs no runtime support beyond wrtinl. Declaring s : string, assigning to it and printing it are all ENoLib — there is no string type, no length word, no assignment path, no WrStr entry. The chr flag on ERes is what keeps writeln('a') calling the character writer instead of the integer writer; without it the compiler emitted the integer path and printed 97 while the test still said OK.
  • Comma-separated names are not supported. var i, c : integer; is a parse error. Pre-existing, unrelated to any of the above, and still open.
  • readln of a BYTE is a latent 1-byte overflow. It calls rdint, which stores a 2-byte word, so the high byte lands on the next variable. TP3 has a separate xrdbyte for exactly this. No fixture declares a BYTE, which is the only reason this has never been observed. Write the fixture first.
  • The program-header parameter loop is unguarded against non-advancing input. program p(1;) would loop forever. The fix needs a BOOLEAN flag and must not use EXIT, which ICEs gm2 in pass 3 (see the pitfalls list). Not yet done.
  • Not implemented (all ENoLib): real, set, record, file, string variables, and any type wider than 2 bytes. with is ENoLib. case is implemented (cascade CMP/JNZ per label, per RESUME-TP3.md §3.6) but only over scalar labels — subrange labels and label lists are untested.
  • gm2 string-literal → ARRAY OF CHAR assignment copies the literal plus a NUL and leaves the tail untouched, so NUL-terminated tables are safe — this was checked, not assumed.

Three bugs only a byte-level dump could find

None of these produced a compile error, and two produced passing tests. All three were found by hex-dumping the emitted image and decoding it by hand — code sizes looked perfectly plausible throughout.

  1. rel16 off by 2 in every direct CALL/JMP. EmCall/EmJmpNear/ EmJcc computed the displacement from pc at a point where pc already pointed past the opcode and at the displacement field; x86 measures from the end of the instruction (pc + 2). ResolvePatches, on the forward-patched path, was already right — which is why forward gotos looked fine and backward ones did not.

  2. EmMovAxSp emitted 8B 04 with no SIB byte. ModRM 04 means "a SIB byte follows", so the CMP AX,imm16 of the next instruction was eaten as that SIB byte, and the intended MOV AX,[SP] became MOV AX,[BP+DI+disp]. Every case label comparison therefore vanished and case compiled to a chain of loads from garbage addresses — while the fixture reported OK. Correct encoding is 8B 44 24 00 ([SP] cannot use mod=00, that computes BP+SP).

  3. The dump tool's own offset column was wrong. It printed 4 nibbles through a helper that formats a byte, so every row was labelled 16× too large (0010 shown as 00000100). A misleading tool is worse than none — it corrupts any offset arithmetic done from its output.

Bugs found by actually running code

Executing the hand-assembled runtime — even under a broken emulator — and running the emitted images back through the harness paid for itself immediately, because a wrong encoding executes rather than failing to assemble. Sixteen so far, none of which a compiler diagnostic would ever have reported.

  1. B() silently truncated multi-byte opcodes. PROCEDURE B emits exactly one byte and masks with MOD 100H, so B (8BE4H) — a two-byte opcode passed as one literal — emitted just E4. MOV BP,SP was missing from every frame in the runtime, so BP stayed 0 and every BP-relative access read address 4. Found by decoding the hex dump; the byte is gone, not wrong.
  2. MOV BP,SP encoded as 8B E4, which is MOV SP,SP — a no-op. The ModRM byte is mod·64 + reg·8 + rm, and I mis-derived it. Compiler.mod's EmMovBpSp had it right all along (8B 0CH); only the new module was wrong. This is why the encoding is now written as three explicit B calls with the arithmetic in a comment rather than as one hex literal.
  3. InitMem read its argument from [SP] — the return address. The convention is a register (AX), unlike the per-argument I/O entries which do take a stack word. Caught because the data area was never cleared.
  4. EmPushVarAddr computed the wrong base register for locals, and truncated the displacement (pre-existing; now fixed). It emitted 8D 46 disp, which is LEA AX,[SI+disp8], but intended LEA AX,[BP+disp8] = 8D 45 disp — so read into a local had always addressed the wrong cell, silently. It also masked the offset with off MOD 100H, losing displacements above
    1. The global form 8D 06 off (LEA AX,[disp16]) was always correct. The fix is EmBpDisp, one procedure that owns the choice: disp8 iff off <= 127, else disp16, with no overflow branch, because disp16 covers the whole 16-bit range as a signed value and every real offset is either negative (locals, allocated down from 0FFFEh) or small-positive. It is now shared by EmLoadVar, EmStoreVar and EmPushVarAddr rather than written three times, and guarded by check_framedisp.py. The old truncation was accidentally correct across −32768..+127, which is the whole range where every real variable lives — so the bug was unreachable from any fixture that existed. t28 exists to make it reachable.
  5. A string literal was eating the rest of the source. Every multi-character literal reported its error at exactly Length() — one past the last character of the buffer — so the editor landed past the final . of the program. The scanner loop tested CurCh # quote, but its "closing quote detected" branch consumed two characters (the content character and the quote), so the cursor moved past the quote and the next condition test saw the character after the literal, was satisfied, and scanned on to end-of-buffer. The IF after the loop that was meant to consume the closing quote was unreachable for any string of two or more characters — which is exactly why writeln('a') always worked and writeln('hi') never did. Worse than a bad caret: it destroyed the parse, so anything after a literal was consumed as string contents and a genuine later error was misattributed to end-of-file.
  6. Every program lost 6 bytes to an uninitialised flag. The FOR over IoCall's argument list needed a "did this argument push a value" flag, and it was never set, so the first argument's CALL and its ADD SP,2 were skipped. Nothing looked wrong — a slightly smaller image looks more plausible, not less. Only expected.tsv pinning code sizes caught it.
  7. writeln('hi') emitted 02 69 00 — i then NUL. StrNew recorded the first character of a literal but did not advance strTop, so the first StrPut landed on top of the seeded character and overwrote it. t17_two_str is what pinned it down: its third emitted character was e, the second literal's character, which had been written into that slot.

The last three were each found by a different means — (5) by noticing that every error position was exactly the buffer length, (6) by expected.tsv, (7) by hex-dumping the image — and it is worth being precise about why all three were invisible to a check that only asks "does it compile": (5) still produced a plausible error number, (6) a plausible code size, and (7) a plausible character. Each is precisely the shape of bug a compile-only fixture ships.

  1. The [BP+off] displacement was truncated to a byte (see 4 above). Found by reading EmLoadVar and asking what off MOD 100H means for a negative local offset — at which point the answer is "correct by accident, and unreachable from any fixture that exists", which is the most expensive kind of wrong.

9–13. Five runtime emitters were one ModRM byte off — MovSiBx 89 DC, CmpSiBx 39 DC, MovSiAx 8B C0, initmem's zeroing loop, and initmem's header word. Tabulated with their correct encodings in the runtime section above. All five decoded cleanly, all five passed a structural check, and all five passed a golden disassembly. They were found by the one check that asks a question the bytes can answer on their own: does this decode to what this procedure is called?

  1. The ModR/M table in the runtime's own documentation was wrong, and it had been wrong since the runtime was written. It read CX DX BX SP BP SI DI BX — the correct list with AX dropped off the front and a duplicate BX invented at the end. Every code was therefore one too low except 100, which lands on SP either way, so the error was invisible at exactly the cell anyone would check first. This was a documentation bug only: the emitters that followed the wrong table emitted 89 DE/39 DE/ 8B F0, which are right. The table is now measured, not remembered — see tests/probe/README.md, which is the fuller account.

  2. DataBytes() returned dc, the absolute end of the data area, not a size. Every program over-reported by 256, and every expected.tsv row had been baselined to agree. The field is documented as "emitted data size in bytes", so 4 is right and 260 was wrong. Re-baselining the whole matrix is exactly the move that can turn a red suite green by hiding a bug, so it was done with the semantic argument above written into the file, and the 6-byte rows (the fixtures declaring one global) are the ones that carry the claim.

  3. build_tpshell.sh was a broken duplicate of the Makefile — it built Posix as if it were Modula-2, and ignored its own link's return code. It was never run, because the Makefile is what everyone runs, and a build script nobody runs is documentation. The specific lesson: when two things must agree, keep one.

Then the image actually ran, and eleven more appeared

The sixteen above were found with byte dumps, structural checks and one broken emulator. Booting the emitted .COM under qemu and comparing exact output is a different kind of instrument: it catches bugs whose every byte is locally correct and whose only defect is that the program goes somewhere else. Eleven more, and the two worst in the project are here.

  1. FOR was emitted as a post-test loop — the increment sat outside the body and the test came after it, so the body ran once before the first comparison. TP3's own emitfor (TPSRC1, DoEmit) tests before the body and increments inside it. Every for fixture printed one line too many, and only run_com_exec.py could say so: the byte count was unchanged, the control flow was well-formed, and the matrix stayed green.
  2. EXIT inside a for body jumped to the increment, not to the exit. t32_forexit therefore re-tested the condition and could re-enter the body. Worse, the handler had an EmAddSp (2) left over from when it exited a with-style scope, which unbalanced the argument-cleanup stack. The patch-list sweep also had to be rewritten from FOR i := a TO exitCnt - 1 to WHILE i < exitCnt — exitCnt is a CARDINAL, so an empty range means TO 65535, and a program with no EXIT in the loop patched 65532 slots. That is a MODULA-2 idiom trap, not a codegen bug.
  3. The procedure-skip jump was missing, so procedure bodies executed as part of the main body. The compiler emits a procedure's body between the caller's prologue and its own main body, so a jump over it is required. DeclaresProc () now answers "does this program declare any procedure?" by a save/restore-srcPos lookahead, and Compile emits EmJmpNear (0) after the prolog when the answer is TRUE, patching it with SetPatTgt (overProc, pc) after DefPart. Four fixture code sizes grew by exactly 3 bytes — the E9 plus its rel16 — and each re-baseline is documented in expected.tsv rather than waved through.
  4. And then the jump's patch slot could not be told from "no jump". EmJmpNear (target) returns a patch slot index, and slot 0 is a legitimate slot. The sentinel was overProc := 0, so IF overProc # 0 THEN SetPatTgt (...) skipped the patch for a program whose procedure-skip jump happened to be the first patch in the image. The jump kept its placeholder target 0, so it landed at image offset 0 — the entry jump — and the program looped forever, printing nothing. The fix is a separate hasProc : BOOLEAN, not a magic slot number. t31_procparam hung rather than failed, which is how it was found: a fixture that never terminates is a louder signal than one that prints the wrong thing — but only if the harness has a timeout, and only if somebody reads a timeout as information.

    The general lesson is the one this project keeps re-learning: a plausible placeholder is indistinguishable from a real value. Slot 0, 0 as "no target", 0 as "no flag set" — each was correct until the first case where the real value was 0. A separate BOOLEAN has no collision to have.

  5. IF MatchKey (tok) AND (tok = TkElse) consumed the token it was rejecting. AND is not short-circuit in Modula-2, and a lookahead built out of a matching primitive is a parser bug, not a lookahead. MatchKey advanced past the keyword, so an else that should have been left for the enclosing statement was eaten. The shape that works is PeekKw (tok) for the question and DropB (MatchKey (tok)) for the commit — and IF <stmt> may not be the last thing in a compound, which is a separate ISO rule that this hit too.

  6. A getbyte with no pushback slot lost one character per call. TP3's getbyte has a char pre-read flag: it reads the next character while looking for digits and then hands it back. The first port had no such slot, so rdint consumed the delimiter and the following statement lost its first character. The fix is a one-character pushback slot D_PUSH plus ungetch, and the slot's address moves whenever the runtime grows (rtSz + LoadBias + dataAt + D_PUSH — at 436 bytes, image 0x1B4, memory 0x2B4, one pad byte before the program header at 0x1B7). It is computed, never hard-coded: the same class of restated-constant bug as RT_SZ above.

  7. rdint did not mirror TP3's xrdint/readnum. TP3 checks for ^Z first, then skips characters <= 20h, then an optional sign, then digits, then pushes the terminator back. Ours skipped leading whitespace only and did not push the terminator back, so two readlns misbehaved. MUL r/m16 also writes DX, so the running digit is parked in DI — an encoding fact that has to be written down, or the next person "simplifies" it back.

  8. INT 21h AH=02h was given the character in DL. It takes it in AL. Four emitters (MovAlD, MovAlSi, MovAlArg, MovAlDl) were wrong together, so every character written was the high half of something else — and the count of bytes written was right, which is why a byte-counting check passed.

  9. EmMovAxSp / EmMovCxSp tried to encode [SP], which the 8086 has no ModRM form for. They now emit POP reg / PUSH reg — a two-instruction shape, so the audit treats them as a sequence rather than trying to match a mnemonic.

  10. There is no [BX] in mode 16. mod=00 rm=111 is [BX+SI], not [BX]. A store through a pointer was writing to BX+SI. Stores now go through DI or SI, and absolute data addresses use mod=00 rm=110 = ModRM 1E. Worth stating as a rule because the name [BX] appears in half the 8086 documentation ever written.

  11. EmMoveAxDx emitted 92h — which is XCHG AX,DX. Behaviour was identical either way, so nothing was broken; only the name lied, and a name that lies is how the next one becomes a silent wrong value. It is now EmXchgAxDx. The same pass audited the emitter vocabulary and renamed the ambiguous ones so the distinction survives: MovBxImm/MovBxVx (ADDRESS) vs LdBxVx (CONTENTS) vs StVxBx; MovAlBl (8A C3, register) vs LdAlBx (8A 07, memory); MovAh0 → {0xB4} vs MovAl0 → {0xB0}.

Then nine fixtures were made to print, and four more appeared

Twenty-seven bugs, and every one of them had been found by looking at bytes, decoding them, or reading them. The remaining fixtures — nine of them, the control-flow ones — had never been executed, because each one assigned to a variable and fell off the end: no write, no writeln, nothing to observe. An empty .out would have asserted only "did not crash", which is a property of the runtime, not of the operator under test. t12_repeat could have been compiling repeat…until as a single-pass loop for the entire life of the project and no check would have said.

Rewriting the nine to print, with each .out derived by hand from the Pascal and not blessed from the machine, found four more. All four shipped with a green compile matrix, a green .COM layout check, a green golden and a green emitter audit.

  1. a * b emitted ADD AX,CX. ParseAdd numbered + as 1 and ParseMul numbered * as 1 as well, and both pass the bare number to BinOpEmit, which cannot see which precedence level called it. So every multiplication dispatched to the addition: a * 2 became a + 2, and 7 * 6 printed 13. The constant-folding arm of BinOpEmit was correct, which is the only reason anything looked right — t08_const's n * n has two constants, and the one fixture that ever multiplied was multiplying two constants. Now OpAdd/OpSub/OpMul are named constants, so two levels cannot collide on a bare 1, and t08 was given a variable multiply.

  2. > and >= had their SETcc opcodes swapped. Op 4 (>) emitted 9Dh = SETGE and op 5 (>=) emitted 9Fh = SETG, so a > b meant a >= b and a >= b meant a > b. Only the equality boundary could see it: 6>5, 4>5, 5<5, 4<=5 and -1>-2 were all already correct, and 5 > 5 and 5 >= 5 were both wrong. One letter apart in the mnemonic, and the CASE arm gave no hint which comparison it answered. Every arm now carries its mnemonic beside the hex.

  3. REPEAT…UNTIL looped back while the condition was TRUE — JNZ where the body should be re-entered only when the condition is FALSE. That is while, so the body ran once, the condition was tested, and it stopped. i := 0; repeat i := i + 1 until i > 5; writeln (i) printed 1; the hand-derived answer is 6. TP3's own sequence (TPSRC8 ~244-300) calls excond, which materialises the negated boolean, then a fixed brnchop of JZ, so a single branch shape serves IF, WHILE and REPEAT alike.

  4. Sibling procedures shared one parameter namespace. A finished procedure's symbols were left at a level that Search's level <= lexnest test still accepted, and every procedure body compiles at the same depth. So a second a : integer was a duplicate (error 41) and an unqualified a inside procedure two silently read procedure one's argument — passing 3 into one and then computing x := a + 1 in two printed the wrong number with no diagnostic at all. HideLocals relabels a procedure's own symbols to 0FFFFH when it closes, which fails the visibility test at every depth a later procedure can be at. Relabelled rather than popped, because symtab[old].resvar holds an index and a function's result variable is one of the entries being hidden.

Then the emitted code stopped being 8086 code, and two more appeared

Thirty-two bugs. Bug 32 had been present since the code generator was written and every check in the project was blind to it, which is a more interesting fault than the previous thirty-one and is worth setting out at length.

  1. Every conditional branch and every comparison in every compiled program was an illegal instruction on the target CPU. EmJcc emitted 0F 8x rel16 (Jcc near) and EmSetcc emitted 0F 9x rel8 (SETcc). 0F is a 386-and-later opcode-escape prefix. The 8086 — the machine TP3 targets, and the machine this compiler exists to emit for — has no such prefix. So the code was not wrong, it was not code.

    What makes this worth a section rather than a bullet is that everything was green while it was true. The compile matrix passed. The .COM layout checker passed. The runtime golden passed. The emitter audit passed. Every fixture booted in qemu and printed exactly its hand-derived expected bytes. Two reasons, and the second is the one to remember:

    • qemu-system-i386 has no 8086 model. Its lowest is 486, where 0F 84 is an ordinary JZ. So the execution oracle cannot see this class of fault, ever — not with a better fixture, not with a longer run. The suite that had found bugs 28–31 could not find this one by construction.
    • FCML is this project's independent disassembler, and FCML's 16-bit mode is a 386. The one tool whose entire job is to say "this is not a real instruction" was architecturally guaranteed to agree with the bug. An independent checker that shares the subject's blind spot is worse than no checker, because it converts an unknown into a false assurance.

    The fix is TP3's own idiom, read off the disassembly of the original rather than invented. TPSRC8 ~246-295 lays IF/WHILE/REPEAT out as MOV AL,brnchop ; MOV AH,#$03 ; CALL eword ; PUSH pc ; CALL ejump — a short Jcc of displacement 3, stepping over a 3-byte EJMP. And TPSRC9 ~412-424 (flgbool) lays a comparison-to-Boolean out as MOV AX,#0001 ; <Jcc> +1 ; DEC AX. Both are 8086 code, both are shorter than what they replace, and the condition is carried in the opcode's low nibble, which is why 70H + cc reproduces the identical condition and all seven EmJcc sites and all six EmSetcc arms go on passing the byte they always passed. The flags survive, which the FOR test needs: it emits CMP then Jcc with nothing in between.

  2. The first attempt at that port inverted every conditional in every program. EmJcc jumps to its target, but the 8086 shape steps over the EJMP — so writing 7x 03 and falling into the destination runs the two the wrong way round. TP3's brnchop is the branch taken when the condition is TRUE; the nibbles these call sites pass are the branch taken when the condition is FALSE (IF's EmJcc (84H) is JZ patched to the ELSE, so it must fire when the test failed). The two spellings are opposite, and the naive port took the wrong one. t09_if printed pos / nonpos / lt for a program whose hand-derived output is nonpos / pos / ge, and t12_repeat looped forever.

    Only the execution suite noticed. The compile matrix and the .COM layout check were both perfectly happy with every conditional inverted — third instance in a row of a byte-level-clean, behaviour-wrong fault. JccShortInv now does the negation (n XOR 1, spelled n + 1 - 2*(n MOD 2) because gm2 under -fiso has no XOR on integers at all), and it is one named function rather than open-coded at two call sites.

Fixing them needed a check that asks a question nothing else was asking: does the target CPU have this opcode at all? tests/check_8086.py.

Its design is mostly about what cannot be done. A linear sweep of the code region is not a sound oracle here: inline string literals are emitted into the code stream, so the sweep desynchronises and then aborts on text — t09_if decodes 16 real instructions and dies on FE E9 0D 00, which is ASCII. So the check anchors on comparison sites instead of boundaries: 3B C1 (EmCmpAxCx) and the three-byte 3D 00 00 (EmCmpAxi (0)) are the only sequences that can precede a lowering, and every one of the seven EmJcc sites and the one EmSetcc site sits immediately after one. A bare 3D anchor is unsound — it matches displacement and immediate bytes, and using one produced two false alarms (a 3D inside a CALL displacement in t11_for, one inside a string in t21_mixed) before it was replaced by the 3-byte form. Where a fixture's region does sweep clean, the sweep additionally asserts no 0F, and the fraction it reached is printed (28 of 31) rather than implied. The CASE arm, whose EmCmpAxi carries a label rather than 0, is found by shape (7x 03 E9, searching for the E9) — but the check cannot assert on the label immediate itself, and t22_case's execution is what covers that.

The check also carries two clauses about which condition each site means, because "is this an 8086 shape" is a much weaker question than it looks. A SETG where a SETGE belongs is still a perfectly good shape. Clause G matches the 13 comparisons of t33_cmpops against the operators in source order, which is what catches a >/>= swap (mutation M3, red with the swap visible in the message: … Dh Dh Fh Fh Dh where the source asks … Fh Fh Dh Dh Fh). That fixture exists partly for this: =, <> and <= had never been used as a comparison anywhere in the suite, and >/>= had already been swapped once, so the coverage hole and the bug it let through were the same hole.

Clause H exists because clause H's need was measured, not anticipated. As first written the check was blind to M5 — the polarity inversion of bug 33 applied to the IF and CASE sites only. IF declares nibble 4, CASE declares 5, they negate into each other, and both are declared, so nothing complained while every conditional took the wrong path. The whole-suite version (M4) was caught only by luck: FOR declares C and F, whose negations D and E are declared for nothing at all. So clause H pins, per fixture, the multiset of conditions its branch sites declare, read off the .pas sources and written out with the reasoning beside each row — a table measured from the image would agree with any behaviour including a wrong one.

Five mutations are now permanent cases in tests/nonvacuity.sh (its final non-vacuity: line reports the total across every section; these five were 44 ok when added): M1 and M2 restore each original defect, M3 swaps >/>=, M4 inverts the branch polarity everywhere, M5 inverts it for IF and CASE only. M5's first run was the one that came back green, and that is the case's whole reason for existing.

Then two fixtures were written for the untested operators, and two more appeared

  1. A computed left operand was destroyed while the right one was parsed. t34_arith exists because div, mod, and, or, unary - and a variable * had never been executed by anything — t08 was the only fixture that multiplied, and it multiplied two constants, which BinOpEmit folds away without emitting an instruction at all. Its and/or lines are a second first: (p > q) and (q > p) puts a value that exists only in AX on both sides of an operator, and nothing kept the left one alive while the right one was parsed. The or line printed FALSE where Pascal says TRUE.

    SaveLeft pushes it the moment the operator is recognised (kind := 4) and LoadPair then materialises AX = left, CX = right in whichever of three shapes the situation needs; the third is the original path and is still correct for its case. Both are named in Compiler.mod because the shape table is the fix — an inline "just reload it" at one call site would not survive the next operator.

    This bug is also why the fix had a stated boundary: the call path had the same hole (f (a > b, x)) and was written down as unfixed rather than quietly left out of a claim about "computed operands". A fix that claims the general case while covering one call site is worse than one that writes its edge down. That boundary is closed by bug 36 below, and the fixture that closes it is t36_argclobber.

  2. not was lowered identically for booleans and integers, so every boolean negation was wrong. TPSRC9's neglevel picks the instruction from the operand's type before it emits anything — NOT AX (F7 D0) for an integer, XOR AL,#01 (34 01) for a boolean, error 47 for anything else — and this compiler emitted NOT AX for both. So not (a = 17) computed 0FFFEh, and the runtime's wrbool tests [BP+4] <> 0, which reads 0FFFEh as TRUE. Pascal says FALSE.

    Exec86 and qemu agreed with each other and both disagreed with Pascal, which is what pins the fault on the compiler rather than on either interpreter: two machines built independently cannot share a bug in an emitter neither of them ever reads. It also made the new oracle worth having — the same disagreement would have been invisible in a suite whose only second opinion was a different execution of the same bytes.

    The fix is neglevel's own split, not a special case: ParseNeg dispatches on r.cls, so t35_not carries both arms — not a has to stay NOT AX and must not be dragged to XOR AL,#01 by a fix aimed at booleans. A narrow row in the helper audit pins the new emitter by its name: a helper called XorAl01 that emits an immediate of 02 fails with "immediate is 2, name says 1", and the matching opcode row refuses a byte no name claims — moving the 34H to 35H fails with "no name pattern accepts it", so the emitter cannot silently become a different instruction.

Then a computed argument met the parse of the next one, and one more appeared

  1. Every call parser read the whole argument list before pushing any of it, so an argument that existed only in AX did not survive the parse of the argument after it — and the callee numbered its parameter slots the opposite way round from the pushes. t36_argclobber exists because no fixture had ever called anything with a computed argument anywhere but last.

    Three parsers deferred the pushes: ParseCallArgs (a procedure call as a statement), ParseCall (a function called inside an expression) and IoCall (write/writeln/read). A kind 2 result means the value is in AX and nowhere else, and LoadAtom deliberately does nothing to it — so by the time a deferred loop reached argument i, AX held whatever argument i+1 had computed. p2 (a + b, x) printed 0 0, b2 (a > b, x > c) printed FALSE FALSE, and writeln (add2 (c * 2, a)) printed 10 instead of 19. IoCall carried a second loss underneath that one: it pushed after the whole list had been parsed, so an argument was also pushed after the runtime call made for an inline string literal in between — writeln (b > a, ' ', x + 1) printed TRUE 544, the first argument handed over as the third one's value and the third read out of an AX that a call had already had. read/readln could not hit any of this, because their arguments are always plain names and a name emits nothing while it is parsed.

    The other half was the callee. Parameter offsets were handed out in declaration order from BP+4, so the first declared parameter sat at BP+4; but an argument pushed first ends up farthest from BP, so the two halves disagreed about which parameter was where, and a one-argument call only ever worked because one slot and one push cannot disagree. Both halves now follow the original: TPSRC8 cproc/cprlp1/cprlp2 emits CALL exprsave then CALL epushax for each argument as it is read and only then the CALL, and RESUME-TP3.md §3.11 states that the last declared parameter is the one at BP+4 — which is exactly where the first-pushed argument ends up. Runtime.mod needed no change: every entry it has takes one stack argument, and one push and one slot are order-independent.

    So each argument is loaded and pushed the moment it has been parsed, in parse order; IoCall's parse loop and its emit loop are one loop; and once a procedure's parameter list has been read, its symbols are remapped with off := parmOff + 2 - off, which sends 4 + 2*(k-1) to 4 + 2*(n-k) and invents no offset in between. The range that remap sweeps is exactly the parameters: the parameter's own NewSym inside the loop is the only symbol created there, ParseType creates none, the procedure's own name predates nestMark, and a function's result variable comes after it.

    No expectation row moved, and that is worth as much as the fix: the reorder emits the same bytes in a different order, and the remap hands the same set of displacements to different names — t28's two disp16 reads are still at +128 and +142, now read through p8 and p1 where they were p63 and p70, and its printed sum still comes from the two slots the sixteen pushed arguments land in. So nothing was re-baselined; the t36 row was written while the compiler was still broken, and its numbers simply held.

    The red was seen before the fixture was encoded, not after: the pre-fix compiler got the two plain-name lines right and the other six wrong, on both oracles at once. Those two greens matter — they are what would catch a fix that flipped only one of the two halves — and they were green before the change as well as after it.

The bug family, stated once

Nine of the bugs recorded here are the same bug in different clothes: loading the address where the value was wanted, or picking the register one byte or one letter away from the right one. EmPushVarAddr had the right bytes for the wrong register. LdAlBx and MovAlBl are one letter apart. MovAh0 and MovAl0 are one letter apart and their opcodes differ by 04h. The response to that is not more checking — it is that the names now carry the distinction, so the next one is a compile error instead of a silent wrong value. Check emitted bytes against intended semantics whenever an emitter is added, and let the audit do it for one-liners.

And a second family, on the harness side rather than the compiler's: a check that quietly checks nothing looks exactly like a check that passes. The audit that swept one module, the regex that matched no emitters, the coverage list derived from the parser's own output, the four nonvacuity mutations that matched nothing, and two hard-coded RT_SZ constants — six instances, all green, all worthless. That is why this file has a non-vacuity section and an independently-scanned inventory at all.

gm2 / ISO Modula-2 pitfalls hit along the way

  • Two-phase link (above) — a single whole-program pass 3 caps identifiers/errors silently.
  • EXIT inside the program-header parameter WHILE ICEs gm2 in pass 3 (ExitStatement → PopExit → M2StackWord_PopWord → invalidloc). Extracting the loop into its own procedure did not help. Use a BOOLEAN "advanced" flag, never EXIT. (LOOP+EXIT is fine, and is used widely.)
  • A bare HALT aborts under -fiso (SIGABRT, exit 134); HALT (0) is correct.
  • CHAR is not the ZType: ch = 09H must be ORD (ch) = 09H.
  • gm2's ISO SYSTEM exports ORD but not Ord — the import is case-sensitive here despite gm2's usual case-insensitivity, so FROM SYSTEM IMPORT Ord fails with "unknown symbol" while plain ORD (c) compiles. Write ORD, never Ord.
  • ISO forbids dropping a function result in a statement → the DropCh/ DropB/DropC discard helpers wrap ~39 call sites.
  • AND/OR/NOT on 16-bit CARDINAL → BitAnd/BitOr/BitNot.
  • PROCEDURE f : T is invalid; the file's convention is PROCEDURE f () : T.
  • ISO will not index a plain string constant as an array (needs an array constructor) — compute hex digits with CHR instead.
  • Foreign modules (FOR "C") must use ADDRESS, not Modula-2 pointer types; cast at the call site. Posix has no Posix.mod (do not try to rebuild it) and no argv binding.
  • termios raw mode: every newline is an explicit CR LF.
  • Foreign modules and Posix cannot pass argv; the test harness reads fixture paths from stdin instead.
  • subprocess.run(input=…) needs bytes, not str, or it raises inside Python rather than reporting the real error.
  • as always picks opcode 89 for a register-to-register mov, so it will never emit 8B EC for the mnemonic mov bp,sp. Test anchors that must assert a specific opcode have to be written as literal .byte.
  • FCML prints immediates with an h suffix (add sp,8h), and the Debian fcml-disasm wrapper aborts with rc=134 on inputs of 16 bytes or more.

Reference material

  • TP3-COMPILER.md — the compiler: what was changed, the Skip bug class, the rel16 off-by-2, standard procedures, current matrix.
  • TP3-EDITOR.md / TP3-EDITOR-PSEUDOCODE.md — the original TP3 editor, from TPSRC5/TPSRC6.
  • RESUME-TP3.md — book-derived reference for the 8086 code TP3 generates per Pascal construct (types, skeletons, arithmetic, IF/CASE/REPEAT/WHILE/ FOR, procedures, parameters, functions, I/O, typed constants, absolutes), plus the TU_* runtime entry list.
  • shell/tests/probe/README.md — read this before touching any emitter. What each ModR/M artifact establishes, which oracle measures which half of the table, why the mod=11 execution probe is structurally impossible, and the shifted table this project shipped, in full.
  • shell/tests/run_all.sh / nonvacuity.sh headers — what runs, in what order, and which breakage is supposed to turn which check red.
  • Resources/turbopascal3source/TP3/ — TPSRC1-10, the disassembled original. This is the ground truth; when our behaviour and the book disagree, the disassembly wins.
  • Resources/coeur-tp-ocr/ — OCR of "Au coeur de Turbo Pascal".

Next steps

  1. The in-suite non-vacuity cases for bug 36. t36_argclobber's .out was proved able to fail by running the pre-fix compiler — red on both oracles, quoted in the commit that introduced the fixture — but nonvacuity.sh does not yet mutate the fix back. Revert ParseCallArgs, ParseCall, IoCall and the parameter-offset remap in turn, and require run_com_exec.py t36_argclobber red for the stated reason, green on restore. Until that exists, the 62 ok under "Non-vacuity" does not cover this fixture, and the text there says so.
  2. String *variables* — s : string, s := 'hi', writeln(s). The encoding blocker is gone (EmBpDisp); what is left is a length word, an assignment path, and a WrStr entry (TPSRC4 xwrtstr). IoCall currently refuses with ENoLib.
  3. Nested procedures / recursion, var parameters (the SEG:OFF push from RESUME-TP3.md §3.11), range/index checks (TU_RANGE_CHECK, TU_INDEX_CHECK), typed constants (RESUME-TP3.md §3.14), array at its point of use (t14), case with subrange labels.
  4. readln of a BYTE calls rdint, which stores 2 bytes and overflows into the next variable. TP3 has a separate xrdbyte; a TU_RdByte entry is the fix. No fixture exists yet, which is why it has not been done — write the fixture first, so the bug is red before the fix.
  5. Make the 4 KiB code window an enforced limit rather than a documented one: report an error when pc reaches dc, instead of writing over the data.
  6. Both spellings of a multi-name declaration. var i, c : integer; is error 1 at the comma and needs two var lines; procedure f (a : integer; b : integer) is error 1 at the semicolon and needs a comma. Neither is wrong Pascal, so a program that compiles under one compiler may not under another. A parameter may also not shadow a global (DupTest rejects any name Search finds at any level), which Pascal allows.
  7. Harden the program-header parameter loop against non-advancing input (program p(1;)) with a BOOLEAN flag — not EXIT, which ICEs gm2.
  8. FreeDOS (freedos.qcow2, FD14-LiveCD) is still untried. Now that R runs the image in-process, a real DOS is no longer needed for any claim above — but it is still the only way to get a third opinion on the INT 21h shim, and Exec86 deliberately implements only the three functions the runtime calls.