#!/bin/bash # Run every check. Non-zero exit if anything fails. # # Six suites, in increasing order of "how much could be lying to me": # # 1. compile matrix every fixture in tests/fixtures/expected.tsv, verdict + # code size + data size ASSERTED from that file. Fast, # no pty. The count is deliberately not repeated here: # it changes whenever a fixture is added, and a number # that is only right until the next edit is how a # comment starts claiming things it cannot support. # 2. .COM linker links every fixture, then re-verifies the bytes with an # independent checker that restates the layout constants # instead of asking the compiler. # 3. UI error path pty: W, C on a bad file, editor opens on the error, # Ctrl-K D, Q. Covers the shell behaviour that only # exists interactively. # 4. UI success path pty: Options -> Destination=COM, C, then the .COM is # checked ON DISK. This is the end-to-end path from # keystroke to artifact. # # 5. frame displ links the two fixtures that touch a stack frame and # asserts the [BP+off] displacement encoding, which no # other check can see: disp8 and disp16 are both # well-formed and both decode cleanly, and only one of # them reads the variable the symbol table named. # # 6. EXECUTE (Exec86) the same fixtures a second time, through # shell/Exec86.mod - this project's own in-process 8086 # interpreter - and require its output to agree with # both the hand-derived .out and qemu, byte for byte. # This is the second, INDEPENDENT execution oracle: it # was written against the 8086's own reference, where # qemu-system-i386's lowest CPU model is a 486 (`0F 84' # is an ordinary JZ there) and FCML's -m16 is a 386. Two # checks that merely agreed with each other would be one # check wearing two hats. # # 7. UI run (R) pty: W, then R - the path that produces NO artifact. # The image is poked into the interpreter at 0100h and # run in-process, so the only evidence is the guest's # own output, asserted against the hand-derived .out, # plus the reported AH=4Ch status and a non-zero step # count. Nothing else in this file can observe R. # # RtProbe dumps the runtime size and its 14 entry offsets, so a # runtime change that moves an entry is visible here. # # Three runtime suites run before all of those, because everything else # trusts the runtime's bytes: # # audit_helpers disassembles every one-line emitter in BOTH Runtime.mod # and Compiler.mod and compares the DECODE against the # procedure's NAME. This is the only check that can see # a wrong ModRM that still decodes cleanly, which is the # mistake this project actually makes: MovSiBx and # CmpSiBx were both `DC` (= MOV SP,BX / CMP SP,BX) for a # long time, and EmXchgAxCx was `93` (XCHG BX,AX) under a # name that says XCHG AX,CX. # # It sweeps Compiler.mod because that is where the # emitter was. Auditing one module said "every helper # agrees with its name" about a module it had not # examined, which is why the XCHG fault survived a green # run. It also cross-checks its own coverage: a # procedure that emits bytes but that the audit never # matched is itself a failure, because a helper that # drops out of the audit when you add a comment is worse # than a helper that is checked and found wanting. # check_comimage asserts that a linked .COM's layout is described in # exactly one file and that every reader imports it. # Each reader used to keep its own copy: two had their # own find_header, one a single constant of it, and one # nothing but a drifted literal for where the runtime # ends - which let its sweep begin inside the runtime # and report PASS over bytes the program never executes. # Silent, because each copy is self-consistent on its # own terms. # check_runtime sweeps the built runtime's code region with FCML: no # desync, every entry and every branch target on an # instruction boundary, and the whole disassembly equal # to tests/runtime.golden. # nonvacuity breaks the runtime and the compiler in the ways each # named check is meant to catch, and asserts it goes red # for the STATED reason, not merely red. Slower, so it # is not in the default run: run it when a check's # sensitivity is in question, or after editing a check. # # nonvacuity.sh is NOT run by default: it rebuilds the runtime and the compiler # a dozen times over and it is a meta-check, so it belongs to "is the harness # honest" reviews rather than to every change. Run tests/nonvacuity.sh # directly after touching any of the checks above. (The count of rebuilds is # not stated here on purpose: it drifts every time a case is added, and a # number in a comment that is only right until the next edit is the kind of # restated constant that hides a fault here.) # # Two more checks guard the ModR/M table itself, which is the input every hand # written emitter in Runtime.mod depends on: # # probe/modrm11.py asks GNU `as` (.code16) to encode the register moves # and FCML to decode them, and requires the result to # agree with the hard-coded bytes, the four # hand-checking anchors, and the table in Runtime.mod. # This is the only check that can catch the TABLE being # wrong rather than one emitter: the table shipped once # with AX dropped off the front of the mod=11 column # and a duplicate BX invented at the end, which shifts # every code down by one and makes `89 DC` (MOV SP,BX) # look like the SI move the name asks for. # probe/modrm19.s the execution probe for mod=00/01/10, kept as the # record of how the memory forms were measured. It is # not run by default: it needs qemu and a hand-built # floppy image, and its result is already written into # the table in Runtime.mod, which modrm11.py checks. # See probe/README.md to re-run it. # # rt_exec.py DOES run. It used to run on Unicorn and fail 33 of 33, because # Unicorn 2.1.4 mis-decodes 16-bit ModRM memory operands - so every failure was # the emulator's, and a wrong oracle is worse than none: it cannot tell "my # codegen is broken" from "the machine is broken". It now boots the same # bootcom.s the .COM fixtures use, one fresh machine per case, so every case # gets a machine that has provably never run anything. See its own docstring # and the EXPECTATION FIXES section in tests/rt_exec.py. set -u D=/home/eric/Projets/Projets-Modula2/MyWork/TP3-comp/shell GM2=/home/eric/bin/Modula2/Gm2/bin/gm2 cd "$D" || exit 9 fail=0 run () { name=$1 shift echo echo "==============================================================" echo "== $name" echo "==============================================================" if "$@" ; then echo "-- $name: PASS" else echo "-- $name: FAIL (rc=$?)" fail=1 fi } # Scratch and logs go in the local tmp/ folder, beside the tree that produced # them, so a failing run's evidence can be read next to the code it describes. mkdir -p ../tmp echo "== build ==" make clean >/dev/null 2>&1 if ! make >../tmp/all.make.log 2>&1 ; then echo "BUILD FAIL" grep -m10 "error:" ../tmp/all.make.log exit 1 fi echo "make rc=0, tpshell $(stat -c%s tpshell) bytes" # gm2 reports a pass-3 rollup on the compiler during phase 1 of the two-phase # link; that is expected and is not a build failure (make rc=0 is the check). # The probe step asks rt_exec.py for its dump rather than building RtProbe a # second time by hand. Two recipes for one artifact is how a stale binary gets # believed: ensure_rtprobe() is the only one that knows Runtime.mod's mtime, and # a probe older than the runtime it reports on is worse than no probe, because # its failure -- a blob with the wrong addresses baked in -- is invisible at the # byte level and catastrophic at the behaviour level. run "runtime probe" python3 tests/rt_exec.py --probe run "helper audit" python3 tests/audit_helpers.py run "layout helper" python3 tests/check_comimage.py run "mod=11 table" python3 tests/probe/modrm11.py run "runtime image" python3 tests/check_runtime.py run "compile matrix" tests/run_compile_tests.sh run "COM linker" tests/run_com_tests.sh # The byte checks above can only ask "are these bytes well formed". This one # asks "does the program do the right thing": every fixture with a .out is # compiled, booted in qemu, and its output compared byte for byte, together # with the DOS exit code. It is the step that found eleven runtime bugs, the # FOR off-by-one, the missing procedure-skip jump and a parser bug - none of # which produced a malformed byte. run "EXECUTE under qemu" python3 tests/run_com_exec.py # The same images a second time, inside this project's own interpreter, which # must agree with qemu byte for byte. This is the only check that can fail on # something qemu cannot do at all: an 8086-only fault, or the interpreter # itself misreading one. It has no --rebless by design - the .out files are # hand-derived from Pascal's semantics, and an oracle that could bless its own # output would pass by construction. run "EXECUTE (Exec86)" python3 tests/run_exec86.py # And this one asks whether the runtime ENTRIES do what they claim, one qemu # boot per case plus a pre-flight. It is the only check that can catch # an entry that is well-formed, decodes cleanly, preserves every register the # driver happens to need - and computes the wrong answer: wrchar and wrbool # borrowed BP to reach their argument and never gave it back, which a byte check # and a register-convention check both call correct. One case (four calls back # to back) noticed; the BP-contract pre-flight is what stops the next entry from # depending on there being such a case. run "runtime entries" python3 tests/rt_exec.py run "frame displ" python3 tests/check_framedisp.py # The machine this compiler targets is an 8086, and the byte checks above all # ask "is this well formed" -- which `0F 84 lo hi' is. qemu-system-i386's # lowest CPU model is 486, so the execution suite could not object either. # This one asks "does the target have this opcode at all", and asserts the # lowering that replaced the two illegal ones is the original compiler's shape. run "8086 opcodes" python3 tests/check_8086.py run "UI error path" python3 tests/uitest.py run "UI success path" python3 tests/comtest.py # And the path that produces NO artifact: R compiles, pokes the image into the # in-process interpreter at 0100h and runs it here. Its evidence is the # guest's own output against the hand-derived .out, plus the reported INT 21h # status and step count - there is no file on disk for any other check to # read, which is why this one drives the shell itself. run "UI run (R)" python3 tests/runtest.py echo echo "==============================================================" if [ "$fail" -eq 0 ]; then echo "OVERALL: ALL PASS" else echo "OVERALL: FAILURES PRESENT" fi echo "==============================================================" exit $fail