SUMMARY.md 46 KB

TP3-comp — state summary

Recreating Turbo Pascal 3.0 (shell / editor / compiler / interpreter) in GNU Modula-2 (gm2 -fiso -Wall). Faithful to the original: screen layout, keys and behaviour are reconstructed from the disassembled TP3.0 source in Resources/turbopascal3source/TP3/ (TPSRC1-10) and the reference manual, not guessed.

Milestones

Milestone Tag State
TP3.0 main-menu shell v0.1-shell done
WordStar-style editor v0.2-editor done
Shell/editor polish, Ctrl-K-D / Ctrl-K-X quit v-TP3-SHELL-EDITOR-QUIT done
Compiler skeleton + build recipe v-TP3-SHELL-COMPILES done
Parser Skip bug class (9 sites) v-TP3-PARSER-FIXES done
Standard procedures + rel16 fix v-TP3-STDPROCS done
Runtime library + 8086 execution harness v-TP3-RUNTIME-BLOB assembled, never run
Inline string literals (writeln('hi')) v-TP3-STRLITERAL done, never run
Linker: real DOS .COM writer + independent byte checker v-TP3-COM-IMAGE done, never run
Measured encodings: ModR/M table, runtime audit, golden disassembly, [BP+off] v-TP3-MEASURED-EMITTERS done, never run
CmdRun, and a .COM that has actually executed — not started

Build

cd shell && make          # → shell/tpshell

Whole-program link must be two-phase; a single gm2 -o pass 3 silently caps identifier/error counts on a compiler-sized program:

gm2 -fiso -c Compiler.mod                                     # phase A
gm2 -fiso -fgen-module-list=modules.lst -o /dev/null \
        Compiler.mod Term.o TextBuf.o Posix.o Editor.o        # phase B1 (rc=1 expected)
gm2 -fiso -fuse-module-list=modules.lst -o tpshell \
        Shell.mod Compiler.mod Term.o TextBuf.o Posix.o Editor.o   # phase B2

Current clean build: make clean && make → rc=0, tpshell 163032 bytes. The one diagnostic is ./Compiler.mod: ParseExpr: too many errors in pass 3, which is the expected phase-1 rollup that the recipe tolerates — not a real error. shell/Makefile is the single authoritative build recipe.

shell/build_tpshell.sh is now a three-line wrapper around make, and that is a fix, not a refactor. It used to be a second hand-maintained copy of the recipe, and a copy of a build recipe drifts — this one was wrong twice over: it ran gm2 -c Posix.mod, but Posix is a foreign C module built by cc -c Posix.c and there is no Posix.mod, so it died on the third module every time; and it printed phase 2's return code and then carried on regardless, so a failed link was reported as a success whenever an older tpshell was still lying around. The copy that is wrong is the one nobody runs, which is how it survived. The wrapper keeps the old invocation working and holds no build knowledge of its own.

What is verified, and how

Nothing here is "it compiles clean" — each claim below comes from a run.

Compiler front end — shell/tests/run_compile_tests.sh

tests/CompileTest.mod links Compiler + TextBuf + Posix, reads fixture paths from stdin, loads each exactly the way LoadWorkFile does (LF→CR normalisation, ^Z ends the text), calls Compile, and prints a verdict plus a source excerpt with a caret at errPos. No pty, instant.

cd shell && tests/run_compile_tests.sh            # all fixtures
cd shell && tests/run_compile_tests.sh /some/dir  # another fixture set
printf '@dump\ntests/fixtures/t19_int1.pas\n' | ./compiletest   # hex-dump the image

The matrix asserts; it does not just count. tests/fixtures/expected.tsv pins the verdict and the numbers per fixture — verdict plus code size plus data size, or error number plus position — and the runner compares. A wrong error position or a program that lost six bytes now fails the suite instead of needing a squint. It was checked for vacuousness by reverting the string scanner fix: 19/23 and exit 1, restored: 23/23 and exit 0.

27 of 29 fixtures compile, up from 1 (the empty program) when the direct harness was first built.

compile matrix: 29 passed, 0 failed (of 29)

Compiling: t01 minimal · t04 var+assign+writeln · t06 two args · t07 10 assignments · t08 const · t09 if/then/else · t10 while · t11 for/to · t12 repeat/until · t13 procedure + value param · t15 label + goto · t16 writeln('a') · t18 bare writeln · t19 writeln(1) · t20 3 string args · t21 mixed args · t22 case with two labels · t23 writeln('') · t24 writeln('don''t') · t26 mixed scalar/string args · t02/t03/t05/t17 multi-char literals · t27 five locals · t28 a 70-parameter declaration.

comtest additionally links every one of the 29 to a real .COM and re-verifies the bytes with an independent checker that restates the layout constants instead of asking the compiler: 26 checked, 0 failed.

Failing, all deliberately: t14 array [1..5] of integer at its point of use and t25 a string literal used as a value (s := 'hi') both → ENoLib (102), the original's "not implemented" path; uierror is a deliberate syntax error used by the UI test.

Shell + editor UI — shell/tests/uitest.py

Drives a real pty (tests/ptyharness.py supplies read-until-quiet, key sending and a small VT100 emulator) and asserts the TP3 compile-error jump: W load → C compile → ESC → editor opens with the cursor on the error position → Ctrl-K D back to the menu → Q exit 0. 10/10 pass, and the reported error number is 41 (not 0 — this is what caught the errNo := errNo self-assignment where Compile's formal shadowed the module variable).

The jump mirrors original TP3 kcwait + editor2 (TPSRC5:333-336, :919): BX:=txerrpos; DEC BX; JMP editor2, and editor2 does ADD BX,txbeg; INC BX — the DEC/INC cancel, so the net is txbeg+txerrpos = our 0-based errPos. Armed as a sticky position (Editor.GotoOffset) rather than by changing Run's signature, so Editor.def stays additive.

uitest.py uses a fixture with a deliberate syntax error (x := 1 + ; → error 41, line 9 col 12) rather than a missing library feature: the latter move as the compiler grows, and a test whose expectations drift with it stops being a test.

The encodings — five checks that can each go red

Nothing above looks at machine code. It all stops at "the compiler produced what it intended to produce", which is exactly where the bugs in this project live: a wrong ModRM byte is not a compile error and not a wrong code size, it is a perfectly well-formed instruction that does something else. So the encodings get their own stack, and it is built so that each check fails on a different class of mistake. tests/run_all.sh runs all of them; tests/nonvacuity.sh proves each one can go red.

check what it asserts what it cannot see
probe/run_modrm19.py the mod=00/01/10 effective addresses, by executing 23 cases on a real 8086 under qemu and scanning for where the marker landed mod=11 — see below
probe/modrm11.py the mod=11 register identities, by encoding with GNU as and decoding with FCML, against hard-coded bytes the table agreeing with itself
audit_helpers.py every one-line emitter in Runtime.mod decodes to what its name says — 61/61 anything longer than one instruction
check_runtime.py + runtime.golden the built runtime's 360-byte code region sweeps cleanly through FCML, every entry and all 37 branch targets land on an instruction boundary, and the whole disassembly is byte-for-byte the committed golden whether the golden is right
check_framedisp.py [BP+off] uses disp8 iff off <= 127, for locals (negative) and far parameters (>127) which of the two encodings was chosen, if the other also works

Two of these deserve the detail, because the reason they exist is the reason they are hard.

audit_helpers.py exists because a wrong ModRM that still decodes is invisible. The mistake this project actually makes is not a malformed instruction — it is MovSiBx emitting 89 DC, which decodes perfectly as MOV SP,BX, and CmpSiBx emitting 39 DC, which decodes as CMP SP,BX. Both shipped for a long time. A structural check passes them. A golden passes them. Only asking "does this byte sequence mean what this procedure is called?" fails, which is what the audit does: it disassembles each one-line emitter and compares the decode against the name. Five bugs came out of it in one pass.

modrm11.py is deliberately not self-referential. The obvious way to check a ModR/M table is to write the table in assembly and assemble it — but that can never fail, because editing the assembly makes as faithfully re-encode the new claim and the two then agree again. That failure mode was found by corrupting the .s and watching the check stay green. Two things close it: an EXPECT byte sequence hard-coded independently of the .s text, and four anchor encodings that are spelled as literal .byte directives because as would never choose them for a mnemonic (83 C4 08 ADD SP,8 · 83 C6 02 ADD SI,2 · 8B EC MOV BP,SP · 8B E5 MOV SP,BP). No table shifted by one cell can satisfy all four.

check_framedisp.py exists because the bug it guards was accidentally correct. The old code truncated the displacement with off MOD 100H, always emitting disp8. Locals are allocated downward from 0FFFEh, so a local's offset is negative and −32768..+127 — the whole range where truncating to a byte happens to be right. The bug was only visible above +127, reachable from the 63rd parameter onward, and no fixture had one. The fix is disp8 iff off <= 127 else disp16, with no overflow branch at all: disp16 covers the entire 16-bit range as a signed value. The check also asserts the rule rather than one encoding — always-disp16 is accepted, and nonvacuity.sh proves that by building it and requiring the check to stay green.

Non-vacuity — tests/nonvacuity.sh

Every assertion above is proved able to fail: 16 deliberate breakages, each asserted to turn exactly one named check red for the stated reason, then restored and re-asserted green. Six break the runtime, five attack the mod=11 table (including restoring the exact wrong table this project once shipped), and three target EmBpDisp — the truncation, the always-disp16 over-encoding that must stay green, and the restored source.

The one that produced the most information was restoring the original shifted table: it turns three cells red rather than one, because the error is invisible at code 100 and only visible from 101 down. See below.

The executor problem (blocking everything downstream)

The whole point of a Pascal→8086 compiler is that the output runs, and it still has not: no compiled image and no runtime entry has ever been executed on a CPU, correct or otherwise. Everything in the five checks above is a claim about bytes. Whether the bytes work is the untested part, and it is the only part that cannot be closed by writing another checker.

The rest of this section is about establishing what can be believed, because the first attempt at this used an emulator that was wrong, and a wrong oracle is worse than none: it cannot distinguish "my codegen is broken" from "the machine is broken".

Unicorn 2.1.4 cannot be used for this. UC_MODE_16 mis-decodes 16-bit ModRM memory operands. The measurement, by loading a byte-pattern image so a load reveals its own effective address, then executing a single LEA AX,[r+disp8] and reading AX back with every register set to a distinct value:

mod=01, disp8=4            got        8086 says
  8D 40 04  LEA AX,[BX+4]  AX=0d04    0x504   (= BX+SI+4)   WRONG
  8D 41 04  LEA AX,[BX+SI+4] AX=0e04  0xd04                WRONG
  8D 42 04  LEA AX,[BX+DI+4] AX=0f04  0xe04                WRONG
  8D 43 04  LEA AX,[BP+4]  AX=1004    0x704                WRONG
  8D 45 04  LEA AX,[DI+4]  AX=0904    0x904                right
  8D 46 04  LEA AX,[BP+4]  AX=0704    0x704                right
  8D 47 04  LEA AX,[DI+4]  AX=0504    0x904                WRONG
mod=00 / mod=10 direct disp16
  8B 1E 00 20  MOV BX,[2000]  BX=1234                     right
  8D 06 34 12  LEA AX,[1234]  AX=1234                     right

So rm=5 and rm=6 decode correctly but rm=0,1,2,3,7 do not, and only base-register-free addressing (direct disp16) is trustworthy. That rules Unicorn out as an oracle for exactly the instruction forms generated code is made of — LEA AX,[BP+d], MOV AX,[SI+d], LODSW-style loops, everything with a frame pointer. It cannot distinguish "my codegen is wrong" from "the emulator is wrong", which makes it worse than no emulator at all.

pip install --upgrade unicorn resolves to the same 2.1.4, so this is not avoidable by upgrading.

This is no longer the whole picture: qemu-system-i386 is now a trusted oracle, and it was trusted by measurement, not by reputation. modrm19.s assembles with GNU as, is wrapped in a 512-byte boot sector, and is booted under /usr/bin/qemu-system-i386 with the serial port captured to a file. For each of 24 ModR/M encodings it stores a marker through the encoding under test and then scans memory for where the word landed, with BX=1000 DI=2000 SI=0030 BP=0040 so every candidate address is distinct. The answers come out as raw offsets, so this is arithmetic, not a judgement call, and the 23 cells it covers are the project's ground truth for effective addressing. It is also how we know qemu's 8086 is right where Unicorn's is wrong, on the same instruction class, by the same method.

Two facts about the tooling came out of that work and are worth recording because both were believed wrong at first:

  • objdump -D -b binary -m i8086 disassembles 16-bit code correctly. An earlier note in this file said there was no usable 16-bit disassembler available; that was wrong, and the cost of believing it was a hand-derived ModR/M table. FCML (fcml-disasm -m16) is the other decoder and the two are cross-checked by tests/fcml_vs_objdump.py. Note fcml-disasm linear-sweeps and aborts (rc=134) on inputs of 16 bytes or more through the Debian wrapper, which is why tests/disasm16.py windows input at 15.
  • A qemu execution probe for mod=11 is structurally impossible, not merely awkward. The comparison register is itself a candidate target, and SP is destroyed by the next call before any check can run — the first version of that probe pushed its return address through SS:0xBEEF, so the evidence it was about to collect had already been overwritten. It also "cleared AX because AX is never an r/m target", which is true of the wrong table and false of the right one. So mod=11 is measured by encoding instead, with as as an oracle independent of both the runtime and qemu.

Also ruled out, for the record:

  • DOSBox-X 2025.02.01 (installed, /usr/bin/dosbox-x, plain root-owned ELF, not a snap) starts cleanly headless under SDL_VIDEODRIVER=dummy SDL_AUDIODRIVER=dummy, but its -c/autoexec commands never observably execute: a md never appeared on the host, and a .COM that creates OUT.TXT via INT 21h AH=3Dh/40h never produced the file. Shell > is intercepted by dosbox-x's own wrapper (SHELL:Redirect output to out.txt). A .BAT route timed out. The user has since installed FreeDOS (freedos.qcow2, FD14-LiveCD) which is very likely the answer to this, and untried.

So what is still missing is narrow: qemu can decode, assemble and execute, but nothing has yet booted an image that was produced by this compiler. The remaining piece is a boot sector that reads a .COM off the floppy with INT 13h, sets SS:SP at the segment top, hooks INT 21h for AH=02h/09h/4Ch to the serial port, and JMP 0x100. tests/rt_exec.py is the harness that will consume it: it loads the runtime, calls each entry with a known argument, and compares the bytes sent to INT 21h against expectations — 33 checks covering initmem, wrint (10 values incl. both INT16 extremes), wrchar, wrbool, wrln, stackchk, a composed writeln(42) writeln TRUE sequence, and the read entries against supplied input including EOF. It was written against Unicorn and currently fails 33 of 33 — the failures are Unicorn's, not the library's, and run_all.sh does not run it. Re-point it at qemu and those numbers become a verdict on the runtime. It has no wrtinl case yet, which needs a harness change rather than just a machine change: that entry's argument is not a stack word, the caller must place a length byte and the characters at the return address, so testing it also tests the encoding contract between Compiler.IoCall and Runtime.EmitWrInl.

Components

Shell — shell/Shell.mod, Term.mod, Posix.c, TextBuf.mod

Exact TP3.0 screen (Logged drive, Active directory, Work file, Main file, Edit Compile Run Save / Dir Quit compiler Options, Text: n bytes, Free: n bytes, >), command letters drawn bold. Keys L A W M E C R S D O Q; any other key redraws the menu, like TP3. W auto-.PAS with Loading/New File; S writes ^Z EOF and rotates the old file to .BAK (unlink+rename, TP3's order); D is a DOS *.* glob listing with k bytes free; O is the options submenu; Q confirms and prompts to save when the text changed.

E and C are wired. R is a stub — it prints "Interpreter pending".

Editor — shell/Editor.mod

WordStar-style full-screen editing over TextBuf. Status line Line n Col n Insert/Overwrite Indent X:FILENAME; text on rows 2-24.

Movement (^S/^D/^E/^X/^A/^F/^R/^C/^W/^Z, ^Q S/D/E/X/R/C/B/K/P, arrows, PgUp/PgDn/Home/End), editing (^V insert/overtype, ^G delete char, backspace joins lines, ^T/^Y word/line, ^Q-Y to EOL, ^N/CR break, TAB auto-indent to the word start above), block (^K B/T/H mark, word, show, ^K C/V/Y copy/move/delete, ^K R/W file read/write), search (^Q-F, ^Q-A, ^L repeat; options B/G/n/U/W, N = no-confirm; ^A any char, CR LF matches a line break), ^P literal control char, ^U/ESC aborts a prompt. ^K-D returns to the shell with the text still in memory and changed reported. Disk format matches TP3: CRLF plus trailing ^Z, load normalises, save re-expands.

Detail: TP3-EDITOR.md, TP3-EDITOR-PSEUDOCODE.md.

Compiler — shell/Compiler.mod

Single-pass Pascal → 8086, following TPSRC6 turbo / TPSRC7-10: one pass over the shared TextBuf emitting machine code into cbuf, a patch list for forward references, TP3-style error reporting (number + relative position), and code/data size accounting. Emitted image is a byte array (mode word, CS/DS, size words, CALL initmem, MOV BP,SP, generated code).

Working subset (v0.5): integer/char/boolean/byte scalars, constants with folding, globals, locals, value parameters, procedures and scalar-result functions, ARRAY[const..const] with constant indexing, control flow, GOTO/EXIT, the standard procedures WRITE, WRITELN, READ, READLN, HALT, and inline string literals as WRITE/WRITELN arguments.

Standard procedures are KBuiltin, not KProc, because they are not called generically. TP3 (TPSRC8 pwriteln/pwrloop/prdtyped) does not pass a descriptor to the runtime: it inspects each argument's class and emits a different call per type, so formatting is fixed at compile time and the runtime only ever sees a value. IoCall mirrors that — one call per argument, then a final call for the line break:

writeln(1)      MOV AX,1     ; PUSH AX ; CALL 20H ; ADD SP,2 ; CALL 40H
writeln('a')    MOV AX,'a'   ; PUSH AX ; CALL 28H ; ADD SP,2 ; CALL 40H
writeln('hi')                 ; CALL 70H ; 02 'h' 'i' ; CALL 40H
readln(x)       LEA AX,[0104]; PUSH AX ; CALL 48H ; ADD SP,2 ; CALL 60H

(Those TU_* names are the compiler's own; the offsets behind them are now assigned from Runtime.RT_Entry rather than written down, so the placeholder-versus-real distinction is gone. The third line is the inline-literal form, which differs in kind: no value is pushed and the ADD SP,2 is absent, because the length and the characters are the argument.)

READ/READLN push the address so the runtime can store (EmPushVarAddr: LEA AX,[BP+off] via EmBpDisp for locals, 8D 06 off for globals); a non-variable argument is ETypeErr (56), as in TP3.

Detail: TP3-COMPILER.md.

Runtime library — shell/Runtime.mod, shell/Runtime.def

The 8086 runtime, assembled byte by byte from Modula-2 — no external assembler and no checked-in binary, so the tree stays self-contained and the entry offsets are derived rather than guessed. Each emitter is one instruction with its ModRM byte spelled out in a comment so the encoding can be checked by hand against an 8086 table. This is the same approach Compiler.mod already takes (Ebyte/Eword/EmCall).

It mirrors the original's own mechanism: TPSRC7 copyrt copies the runtime into the front of the code buffer (SI=DI=0, REPZ MOVSB) and pc is then initialised past it (MOV pc,#$2D7C). Same shape here, and now actually done: pc := RT_Size, dc := RT_Size + 1000H, so the image is [runtime][program header][program code] and every emitted address is image-absolute. No relocation pass is needed — worth having paid for, since a linker that has to walk fixups is a linker that can get them wrong.

Current blob: 391 bytes, 14 entries, 37 branch targets, offsets read back out of the assembled bytes by tests/check_runtime.py rather than asserted by hand:

entry offset entry offset entry offset
initmem 0 wrint 36 rdint 167
progend 28 wrchar 97 rdchar 268
stackchk 35 wrbool 109 rdbool 289
halt 28 wrreal 132 rdln 335
wrln 140 wrtinl 148

The TU_* constants in Compiler.mod are assigned from Runtime.RT_Entry in Inittur, not written down, so a runtime edit that moves an entry cannot leave the compiler calling the old address.

progend and halt deliberately share one address (XOR AX,AX / MOV AH,4C / INT 21h / RET): the compiler already zeroes AX before progend and discards the HALT argument at compile time, so both leave with exit code 0.

Conventions, matching Compiler.IoCall exactly:

entry argument notes
WrInt/WrChar/WrBool/WrReal one 16-bit value on the stack caller pops
RdInt/RdChar/RdBool one address on the stack caller pops
WrLn/RdLn/StackChk nothing
InitMem AX = offset of the program header a register, not a stack word
ProgEnd/Halt nothing exits, code 0
WrInl nothing — reads its own text via POP BX see below

InitMem receives the program-header offset in AX (not on the stack), reads the data base and end out of the header, and zeroes that range, because Pascal leaves globals undefined. StackChk is a bare RET — range and stack checking aren't compiled in yet, and the call site sits mid-expression, so it must not touch a register.

Five emitter bugs were fixed here in one session, all found by the audit_helpers.py name-vs-decode pass described above, and all of them were invisible to everything else that was already in place:

emitter emitted actually was correct
MovSiBx 89 DC MOV SP,BX 89 DE
CmpSiBx 39 DC CMP SP,BX 39 DE
MovSiAx 8B C0 MOV AX,AX (a no-op) 8B F0
initmem zeroing loop MovAxDx loaded a value it then discarded XOR AX,AX
initmem header read +8 hdrMax +6 (hdrHeap)

The third is the instructive one, because it is the same error pointed the other way. 8B C0 reads as MOV AX,AX under the shifted ModR/M table this project shipped, and the fix looked like it should be the byte that table said was SI. It is not: for opcode 8B the reg field is the destination, so 8B F0 (reg=110=SI, r/m=000=AX) is MOV SI,AX, which is what the name asks for. Getting the direction backwards nearly caused a correct fix to be reverted.

Two label-name plus fixup list: rel8, rel16 and runtime-data addresses are all patched after the blob is placed, so nothing depends on a hand-computed displacement.

It has still never executed. The encodings are now audited, golden-pinned and non-vacuity-proved, which is a much stronger static claim than "it assembles" — but a static check cannot tell you the code works, only that it is what was intended. See the executor section.

The program image — shell/Linker.mod

The runtime is copied to the front of the code buffer, then the program header, then the program. The header is our own format at rtSz:

off field
+0 hdrFlag 1, "header present"
+2 hdrCS
+4 hdrDS
+6 hdrHeap = dc, the end of the data area
+8 hdrMax
+10… max-open-files, input buffer, output buffer words

InitMem gets the header offset in AX and reads +6 for the heap limit; the independent checker in run_com_tests.sh restates these offsets as its own constants and asserts initmem's SI displacements equal them, assertion by assertion, rather than asking the compiler where it thinks the header is.

CmdCompile honours the Destination option: 0 = memory, 1 = .COM, 2 = .CHN (refused). A .COM is named after its source with the extension swapped at the last dot, padded with a zero gap to max(pc, dc); its stack sits at the segment top (SS = SP = CS:FFFE), which is where DOS puts it.

Known limit: the data area starts at a fixed rtSz + 1000H (391 + 4096 = 4481), so a program whose code exceeds 4 KiB runs into its own data. Every fixture is at 4491 or 4493 bytes. Documented rather than fixed, because the original has the same fixed-offset behaviour.

String literals — writeln('hi')

writeln('toto') was the last thing standing between the front end and a hello-world: string literals had no encoding at all, so anything past one character was ENoLib. The encoding is not invented — it is the original's.

TPSRC8 pwrinlin peeks at the character after the literal: if it is , or ) the literal is a WRITE argument, not an expression, and it emits (TPSRC10 estring) the length byte and the characters into the code stream right behind the call:

CALL wrtinl   <length byte> <character>...

TPSRC4 xwrtinl is what makes that self-delimiting: POP BX takes the return address — which is the address of the length byte — and the entry ends with JMP BX, returning to just past the last character. So the literal needs no terminator, no length table, and nothing at all in the data segment. The arithmetic confirms it: t26 emits 02 68 69 inline and its data size is unchanged from a program with no strings at all.

wrtinl is 19 bytes at offset 148 (5B POP BX · 31 C9 XOR CX,CX · 8A 0F MOV CL,[BX] · 43 INC BX · B4 02 MOV AH,2 · E3 07 JCXZ to the end label · 8A 07 MOV AL,[BX] · CD 21 · 43 · E2 F9 LOOP · FF E3 JMP BX) — hand-checked once, and now also covered by the golden disassembly and the audit.

A string literal is a value in exactly one place: a WRITE/WRITELN argument. Everywhere else it is a hard error, and it is enforced in a single place — LoadAtom — because assignment, IF, WHILE, FOR, REPEAT, CASE, array subscripts and every operator all reach their operand through LoadAtom, and none of them can use a counted string where a 16-bit word is expected. ParseFactor therefore marks a literal (kind = 3) rather than rejecting it, and IoCall handles it before LoadAtom is ever reached. The alternative — letting it through and producing a machine word that happens to be a pointer — would be a silently wrong program; t25 pins the error instead.

Literal text has to survive from the scan to IoCall, since the parser does not yet know it is writing rather than computing, so it is collected into a pool as it is read: strPool[0..4095], strOff/strLen[0..255], strTop, strCnt, plus StrNew/StrPut. The pool is reset in Inittur, so it is per-compilation.

Details that are deliberate, not incidental:

  • '' is a zero-length literal, reaching the runtime's JCXZ path. It used to be the scalar 39, so writeln('') printed a quote mark.
  • 'don''t' — the doubled quote becomes one character; t24 pins len 5.
  • A literal of 256 characters or more is EConstRange (45), not truncated. The length is one byte, so 300 characters would go out behind a length of 44 and the runtime would print 44 of them and silently drop the rest. TP3 strings are at most 255 characters, so refusing is the faithful answer.
  • A string variable is ENoLib, not wrong code. IoCall knows the difference and refuses. EmPushVarAddr's local form is fixed now (see bug 4 below), so the blocker is no longer the encoding — it is that there is no string type, no length word, no assignment path and no WrStr entry.

Honest limitations

  • A compiled image has still never been executed. This is the one limitation that everything else is downstream of. CmdRun is a stub; the linker does now write a real .COM; qemu is a proven oracle for encodings but nothing has yet booted an image this compiler produced. Everything in the checks section is a claim about the bytes, and the bytes have been checked hard. Whether the bytes work is exactly the untested part.
  • The runtime is audited but unproven. 391 bytes, 14 entries, every one-line emitter decoded against its own name, the whole code region golden-pinned, all 37 branch targets on instruction boundaries — and still never run on a CPU. Nine of its own bugs have been found this way so far, so the prior is not reassuring. wrtinl is the newest entry and the only one no check touches even in principle: its argument lives at its own return address, so testing it needs a harness that models the caller's contract.
  • wrreal is a deliberate stub. It writes the literal text ?REAL? — the string lives in the runtime's own data block at D_REAL=24, which is what makes it a real 9-byte routine rather than a trap. Reals are not formatted yet, so writeln(1.5) "works" and prints nonsense. A trap would be louder; neither choice is a real answer, and this is now reachable code rather than an unreachable one, which raises the stakes on the choice.
  • Code above 4 KiB overruns the data area. The data base is fixed at rtSz + 1000H = 4481 and a .COM is padded to max(pc, dc), so the fixed 4 KiB code window is real and not advisory. Every fixture is 4491 or 4493 bytes, so nothing has hit this yet and nothing tests it.
  • EmMovAxSp still emits a 386-only SIB byte (8B 44 24 00). It is correct on any 386+ but the SIB byte did not exist in 1984, and the whole premise of this project is an 8086. Same class of bug as finding 2 below, unfixed.
  • String literals work; string variables do not. A literal in a WRITE/WRITELN argument list is emitted inline and needs no runtime support beyond wrtinl. Declaring s : string, assigning to it and printing it are all ENoLib — there is no string type, no length word, no assignment path, no WrStr entry. The chr flag on ERes is what keeps writeln('a') calling the character writer instead of the integer writer; without it the compiler emitted the integer path and printed 97 while the test still said OK.
  • Comma-separated names are not supported. var i, c : integer; is a parse error. Pre-existing, unrelated to any of the above, and still open.
  • Not implemented (all ENoLib): real, set, record, file, string variables, and any type wider than 2 bytes. with is ENoLib. case is implemented (cascade CMP/JNZ per label, per RESUME-TP3.md §3.6) but only over scalar labels — subrange labels and label lists are untested.
  • gm2 string-literal → ARRAY OF CHAR assignment copies the literal plus a NUL and leaves the tail untouched, so NUL-terminated tables are safe — this was checked, not assumed.

Three bugs only a byte-level dump could find

None of these produced a compile error, and two produced passing tests. All three were found by hex-dumping the emitted image and decoding it by hand — code sizes looked perfectly plausible throughout.

  1. rel16 off by 2 in every direct CALL/JMP. EmCall/EmJmpNear/ EmJcc computed the displacement from pc at a point where pc already pointed past the opcode and at the displacement field; x86 measures from the end of the instruction (pc + 2). ResolvePatches, on the forward-patched path, was already right — which is why forward gotos looked fine and backward ones did not.

  2. EmMovAxSp emitted 8B 04 with no SIB byte. ModRM 04 means "a SIB byte follows", so the CMP AX,imm16 of the next instruction was eaten as that SIB byte, and the intended MOV AX,[SP] became MOV AX,[BP+DI+disp]. Every case label comparison therefore vanished and case compiled to a chain of loads from garbage addresses — while the fixture reported OK. Correct encoding is 8B 44 24 00 ([SP] cannot use mod=00, that computes BP+SP).

  3. The dump tool's own offset column was wrong. It printed 4 nibbles through a helper that formats a byte, so every row was labelled 16× too large (0010 shown as 00000100). A misleading tool is worse than none — it corrupts any offset arithmetic done from its output.

Bugs found by actually running code

Executing the hand-assembled runtime — even under a broken emulator — and running the emitted images back through the harness paid for itself immediately, because a wrong encoding executes rather than failing to assemble. Sixteen so far, none of which a compiler diagnostic would ever have reported.

  1. B() silently truncated multi-byte opcodes. PROCEDURE B emits exactly one byte and masks with MOD 100H, so B (8BE4H) — a two-byte opcode passed as one literal — emitted just E4. MOV BP,SP was missing from every frame in the runtime, so BP stayed 0 and every BP-relative access read address 4. Found by decoding the hex dump; the byte is gone, not wrong.
  2. MOV BP,SP encoded as 8B E4, which is MOV SP,SP — a no-op. The ModRM byte is mod·64 + reg·8 + rm, and I mis-derived it. Compiler.mod's EmMovBpSp had it right all along (8B 0CH); only the new module was wrong. This is why the encoding is now written as three explicit B calls with the arithmetic in a comment rather than as one hex literal.
  3. InitMem read its argument from [SP] — the return address. The convention is a register (AX), unlike the per-argument I/O entries which do take a stack word. Caught because the data area was never cleared.
  4. EmPushVarAddr computed the wrong base register for locals, and truncated the displacement (pre-existing; now fixed). It emitted 8D 46 disp, which is LEA AX,[SI+disp8], but intended LEA AX,[BP+disp8] = 8D 45 disp — so read into a local had always addressed the wrong cell, silently. It also masked the offset with off MOD 100H, losing displacements above
    1. The global form 8D 06 off (LEA AX,[disp16]) was always correct. The fix is EmBpDisp, one procedure that owns the choice: disp8 iff off <= 127, else disp16, with no overflow branch, because disp16 covers the whole 16-bit range as a signed value and every real offset is either negative (locals, allocated down from 0FFFEh) or small-positive. It is now shared by EmLoadVar, EmStoreVar and EmPushVarAddr rather than written three times, and guarded by check_framedisp.py. The old truncation was accidentally correct across −32768..+127, which is the whole range where every real variable lives — so the bug was unreachable from any fixture that existed. t28 exists to make it reachable.
  5. A string literal was eating the rest of the source. Every multi-character literal reported its error at exactly Length() — one past the last character of the buffer — so the editor landed past the final . of the program. The scanner loop tested CurCh # quote, but its "closing quote detected" branch consumed two characters (the content character and the quote), so the cursor moved past the quote and the next condition test saw the character after the literal, was satisfied, and scanned on to end-of-buffer. The IF after the loop that was meant to consume the closing quote was unreachable for any string of two or more characters — which is exactly why writeln('a') always worked and writeln('hi') never did. Worse than a bad caret: it destroyed the parse, so anything after a literal was consumed as string contents and a genuine later error was misattributed to end-of-file.
  6. Every program lost 6 bytes to an uninitialised flag. The FOR over IoCall's argument list needed a "did this argument push a value" flag, and it was never set, so the first argument's CALL and its ADD SP,2 were skipped. Nothing looked wrong — a slightly smaller image looks more plausible, not less. Only expected.tsv pinning code sizes caught it.
  7. writeln('hi') emitted 02 69 00 — i then NUL. StrNew recorded the first character of a literal but did not advance strTop, so the first StrPut landed on top of the seeded character and overwrote it. t17_two_str is what pinned it down: its third emitted character was e, the second literal's character, which had been written into that slot.

The last three were each found by a different means — (5) by noticing that every error position was exactly the buffer length, (6) by expected.tsv, (7) by hex-dumping the image — and it is worth being precise about why all three were invisible to a check that only asks "does it compile": (5) still produced a plausible error number, (6) a plausible code size, and (7) a plausible character. Each is precisely the shape of bug a compile-only fixture ships.

  1. The [BP+off] displacement was truncated to a byte (see 4 above). Found by reading EmLoadVar and asking what off MOD 100H means for a negative local offset — at which point the answer is "correct by accident, and unreachable from any fixture that exists", which is the most expensive kind of wrong.

9–13. Five runtime emitters were one ModRM byte off — MovSiBx 89 DC, CmpSiBx 39 DC, MovSiAx 8B C0, initmem's zeroing loop, and initmem's header word. Tabulated with their correct encodings in the runtime section above. All five decoded cleanly, all five passed a structural check, and all five passed a golden disassembly. They were found by the one check that asks a question the bytes can answer on their own: does this decode to what this procedure is called?

  1. The ModR/M table in the runtime's own documentation was wrong, and it had been wrong since the runtime was written. It read CX DX BX SP BP SI DI BX — the correct list with AX dropped off the front and a duplicate BX invented at the end. Every code was therefore one too low except 100, which lands on SP either way, so the error was invisible at exactly the cell anyone would check first. This was a documentation bug only: the emitters that followed the wrong table emitted 89 DE/39 DE/ 8B F0, which are right. The table is now measured, not remembered — see tests/probe/README.md, which is the fuller account.

  2. DataBytes() returned dc, the absolute end of the data area, not a size. Every program over-reported by 256, and every expected.tsv row had been baselined to agree. The field is documented as "emitted data size in bytes", so 4 is right and 260 was wrong. Re-baselining the whole matrix is exactly the move that can turn a red suite green by hiding a bug, so it was done with the semantic argument above written into the file, and the 6-byte rows (the fixtures declaring one global) are the ones that carry the claim.

  3. build_tpshell.sh was a broken duplicate of the Makefile — it built Posix as if it were Modula-2, and ignored its own link's return code. It was never run, because the Makefile is what everyone runs, and a build script nobody runs is documentation. The specific lesson: when two things must agree, keep one.

gm2 / ISO Modula-2 pitfalls hit along the way

  • Two-phase link (above) — a single whole-program pass 3 caps identifiers/errors silently.
  • EXIT inside the program-header parameter WHILE ICEs gm2 in pass 3 (ExitStatement → PopExit → M2StackWord_PopWord → invalidloc). Extracting the loop into its own procedure did not help. Use a BOOLEAN "advanced" flag, never EXIT. (LOOP+EXIT is fine, and is used widely.)
  • A bare HALT aborts under -fiso (SIGABRT, exit 134); HALT (0) is correct.
  • CHAR is not the ZType: ch = 09H must be ORD (ch) = 09H.
  • gm2's ISO SYSTEM exports ORD but not Ord — the import is case-sensitive here despite gm2's usual case-insensitivity, so FROM SYSTEM IMPORT Ord fails with "unknown symbol" while plain ORD (c) compiles. Write ORD, never Ord.
  • ISO forbids dropping a function result in a statement → the DropCh/ DropB/DropC discard helpers wrap ~39 call sites.
  • AND/OR/NOT on 16-bit CARDINAL → BitAnd/BitOr/BitNot.
  • PROCEDURE f : T is invalid; the file's convention is PROCEDURE f () : T.
  • ISO will not index a plain string constant as an array (needs an array constructor) — compute hex digits with CHR instead.
  • Foreign modules (FOR "C") must use ADDRESS, not Modula-2 pointer types; cast at the call site. Posix has no Posix.mod (do not try to rebuild it) and no argv binding.
  • termios raw mode: every newline is an explicit CR LF.
  • Foreign modules and Posix cannot pass argv; the test harness reads fixture paths from stdin instead.
  • subprocess.run(input=…) needs bytes, not str, or it raises inside Python rather than reporting the real error.
  • as always picks opcode 89 for a register-to-register mov, so it will never emit 8B EC for the mnemonic mov bp,sp. Test anchors that must assert a specific opcode have to be written as literal .byte.
  • FCML prints immediates with an h suffix (add sp,8h), and the Debian fcml-disasm wrapper aborts with rc=134 on inputs of 16 bytes or more.

Reference material

  • TP3-COMPILER.md — the compiler: what was changed, the Skip bug class, the rel16 off-by-2, standard procedures, current matrix.
  • TP3-EDITOR.md / TP3-EDITOR-PSEUDOCODE.md — the original TP3 editor, from TPSRC5/TPSRC6.
  • RESUME-TP3.md — book-derived reference for the 8086 code TP3 generates per Pascal construct (types, skeletons, arithmetic, IF/CASE/REPEAT/WHILE/ FOR, procedures, parameters, functions, I/O, typed constants, absolutes), plus the TU_* runtime entry list.
  • shell/tests/probe/README.md — read this before touching any emitter. What each ModR/M artifact establishes, which oracle measures which half of the table, why the mod=11 execution probe is structurally impossible, and the shifted table this project shipped, in full.
  • shell/tests/run_all.sh / nonvacuity.sh headers — what runs, in what order, and which breakage is supposed to turn which check red.
  • Resources/turbopascal3source/TP3/ — TPSRC1-10, the disassembled original. This is the ground truth; when our behaviour and the book disagree, the disassembly wins.
  • Resources/coeur-tp-ocr/ — OCR of "Au coeur de Turbo Pascal".

Next steps

  1. Boot a compiled .COM and read its output. This gates everything and nothing else can honestly be claimed until it works. The executor is no longer the open question — qemu is a proven oracle and the FreeDOS image gives a real DOS to run in. Concretely: a boot sector that reads a .COM off the floppy with INT 13h to 0x100, sets SS:SP at the segment top, hooks INT 21h (AH=02h/09h/4Ch → serial), JMP 0x100; debug with qemu -d in_asm,exec -D trace.log; assert the exact stdout bytes per fixture against expected-output files under shell/tests/fixtures/ (writeln('hi') → hi). That single assertion is what turns "assembles" into "works". Either route works and both are worth having: under FreeDOS for realism, under bare qemu with an INT 21h shim for reproducibility.
  2. Re-point rt_exec.py at qemu and require all 33 of its checks to pass. They currently fail 33/33 under Unicorn, and those failures are the emulator's, not the runtime's. The expectations stay; only the machine changes. Add the missing wrtinl case, which needs a harness that models the caller's contract (length byte and characters at the return address).
  3. CmdRun as an in-process 8086 interpreter — the R menu key, and a fallback executor for environments with no DOS. Validate it against qemu on the same images, so the two oracles check each other.
  4. String *variables* — s : string, s := 'hi', writeln(s). The encoding blocker is gone (EmBpDisp); what is left is a length word, an assignment path, and a WrStr entry (TPSRC4 xwrtstr).
  5. Nested procedures / recursion, var parameters (the SEG:OFF push from RESUME-TP3.md §3.11), range/index checks (TU_RANGE_CHECK, TU_INDEX_CHECK), typed constants (RESUME-TP3.md §3.14), array at its point of use (t14), case with subrange labels.
  6. Fix EmMovAxSp, which still emits the 386-only 8B 44 24 00. On an 8086 there is no SIB byte, so the right encoding is 8B 46 00 (MOV AX,[BP+0]-adjacent form) or a register copy — this needs thinking against the measured table rather than a habit, and a fixture that reads [SP].
  7. Make the 4 KiB code window an enforced limit rather than a documented one: report an error when pc reaches dc, instead of writing over the data.
  8. Comma-separated names: var i, c : integer;.
  9. Harden the program-header parameter loop against non-advancing input (program p(1;)) with a BOOLEAN flag — not EXIT, which ICEs gm2.