#!/usr/bin/env python3 """comimage.py -- the layout of a linked .COM, measured from the file itself. Every check in this project that reads an emitted image has to answer the same three questions before it can say anything at all: where is the program header? find_header(d) where does the runtime end? find_header(d) - HDR_SZ where does execution start? entry_target(d) Until now each check answered for itself. comtest.py carried one find_header, the independent checker inside run_com_tests.sh carried a second copy of it, and check_framedisp.py carried neither: it wrote RT_SZ = 391 beside a comment saying it tracked Runtime.RT_Size(), and swept the region beginning ENT_SZ + RT_SZ - which is inside the RUNTIME, part way through an instruction. The check passed anyway, for the two reasons a restated offset always passes: the byte patterns it searched for are in the image wherever they happen to be, and the decode sweep reached the same answers from a start point that nothing in the check could tell was wrong. A region that is wrong in a way its assertions cannot see is the same fault as a constant that has drifted, only harder to notice, because the report it prints is a clean PASS. So the layout lives in one file, is measured rather than remembered, and every reader imports it. tests/check_comimage.py asserts that there is exactly one definition of each part and that every reader gets it from here, because a shared helper that somebody copies again is four sources of truth instead of one - and four is how this started. What is written down here, and what is not ------------------------------------------ The FORMAT is written down: the entry jump is three bytes, the header is sixteen, initmem's first bytes and the two header words it reads, and the load bias. Those are this project's knowledge of what it emits, and asking the code under test for them would let the code under test satisfy every check by agreeing with itself. The SIZES are not written down. The runtime's size was written down twice (385, then 391), beside a comment saying it tracked the runtime; it did not, and each stale value sent a checker into the middle of the code, where it reported a well-formed image as a compiler fault. A duplicated constant that has silently drifted is not an independent check, it is a second source of truth that lies, and it lies in the direction of looking like the thing under test is broken. Measuring is not the same as asking the compiler: everything here reads the emitted file. tests/check_runtime.py is where the runtime's size is pinned on purpose, and tests/run_com_tests.sh prints the size it measures on every run. Usage: this is a library; import it from a check. """ # ---- the format, written down ------------------------------------------- # ENT_SZ and HDR_SZ are the layout, and the load bias below is a decision this # project made about how a DOS .COM is loaded. Those stay written down: they # are knowledge about the format, not sizes that move when code is edited. ENT_SZ = 3 # E9 lo hi, the entry jump (Compiler.Inittur) HDR_SZ = 16 # 5 header words + 3 buffer words (Compiler) # The load bias. A DOS .COM's first byte is at CS:0100 and CS = DS, so an image # offset K is at DS:(K + 0100h); every ABSOLUTE address the image contains has # to carry it, or it points 0100h low and - since the code region and the # runtime are both below the bias - almost always lands inside the runtime # instead of inside the data. Relative encodings must not carry it, because # both of their operands shift together. # Restated, not asked of the code under test. See Runtime.LoadBias. LOAD_BIAS = 0x100 # initmem's prologue, which is the runtime's only reader of the program header. # The displacements +4 and +6 below are the whole point of this constant: the # header checks in every checker read hdrDS at +4 and hdrHeap at +6, and # initmem has to read the SAME two words or it clears the wrong range. It read # +8 (hdrMax, which the compiler patches to 0), so it zeroed nothing at all and # nothing here noticed - the emitted loop was perfectly well formed, it just # never ran. Asserting these bytes together with the header offsets is what # closes that gap. # # It is the first eleven bytes of the runtime, so it sits at ENT_SZ, not 0. HEAD = "8B F0 8B 54 04 8B 4C 06" # MOV SI,AX / MOV DX,[SI+4] / MOV CX,[SI+6] HDR_DS_WORD = 4 # header word holding the data base HDR_HEAP_WORD = 6 # header word holding the data end assert [int(HEAD.split()[4], 16), int(HEAD.split()[7], 16)] == \ [HDR_DS_WORD, HDR_HEAP_WORD], \ "initmem no longer reads the two header words the checks verify" # ---- what the file says, measured --------------------------------------- def find_header(d): """The image offset of the program header in `d`, or None. The header is eight words, and the layout says what they are (offsets here are BYTES into the header, which is why HDR_DS_WORD is 4 and not 2 - the words are two bytes each): +0 1 hdrFlag, always 1 +2 code end + bias hdrCS +4 data base + bias hdrDS, where data base = hdrOff + 1000h +6 data end + bias hdrHeap, which is hdrDS + dataBytes hdrDS ties the header to its OWN offset: the data base is header offset + 1000h, and hdrDS is that with the load bias added, so the offset is recoverable from the file with no remembered runtime size - which is the whole point. A candidate is accepted only if hdrFlag is 1, hdrDS satisfies that equation, hdrHeap is above hdrDS (a heap below its own base is not a layout, it is a coincidence), and hdrCS leaves room for the header itself. initmem is the only code in the image that reads the header, so its bytes are pinned at ENT_SZ: a candidate that also has them there is not a coincidence in the code stream. A header that cannot be found is returned as None rather than as the nearest guess, because a caller that guesses reports confidently about the wrong bytes - which is what the restated runtime size used to do. """ head = bytes(int(x, 16) for x in HEAD.split()) if d[ENT_SZ:ENT_SZ + len(head)] != head: return None # no runtime: nothing to measure for off in range(ENT_SZ, len(d) - HDR_SZ + 1): w = (lambda b: int.from_bytes(d[off + b:off + b + 2], "little")) if w(0) != 1: # hdrFlag continue if w(HDR_DS_WORD) != off + 0x1000 + LOAD_BIAS: # hdrDS continue if w(HDR_HEAP_WORD) <= w(HDR_DS_WORD): # hdrHeap continue if w(2) - LOAD_BIAS < off + HDR_SZ: # hdrCS continue return off return None def entry_target(d): """The image offset the entry JMP at file offset 0 lands on, or None. A .COM is entered at its first byte, so byte 0 is `E9 rel16' and the first instruction of the program is at ENT_SZ + rel16 (Compiler.Inittur). This is where EXECUTION starts, which is a different statement from where the runtime ends: the program header sits between the two, and a check that measures only one of them cannot tell that it has started its sweep in the wrong place. A check that measures both is asked to make them agree. """ if len(d) < ENT_SZ or d[0] != 0xE9: return None return ENT_SZ + int.from_bytes(d[1:ENT_SZ], "little", signed=True)