| 123456789101112131415161718192021222324252627282930313233343536373839404142434445464748495051525354555657585960616263646566676869707172737475767778798081828384858687888990919293949596979899100101102103104105106107108109110111112113114115116117118119120121122123124125126127128129130131132133134135136137138139140141142143144145 |
- #!/usr/bin/env python3
- """comimage.py -- the layout of a linked .COM, measured from the file itself.
- Every check in this project that reads an emitted image has to answer the same
- three questions before it can say anything at all:
- where is the program header? find_header(d)
- where does the runtime end? find_header(d) - HDR_SZ
- where does execution start? entry_target(d)
- Until now each check answered for itself. comtest.py carried one find_header,
- the independent checker inside run_com_tests.sh carried a second copy of it,
- and check_framedisp.py carried neither: it wrote RT_SZ = 391 beside a comment
- saying it tracked Runtime.RT_Size(), and swept the region beginning
- ENT_SZ + RT_SZ - which is inside the RUNTIME, part way through an instruction.
- The check passed anyway, for the two reasons a restated offset always passes:
- the byte patterns it searched for are in the image wherever they happen to be,
- and the decode sweep reached the same answers from a start point that nothing
- in the check could tell was wrong. A region that is wrong in a way its
- assertions cannot see is the same fault as a constant that has drifted, only
- harder to notice, because the report it prints is a clean PASS.
- So the layout lives in one file, is measured rather than remembered, and every
- reader imports it. tests/check_comimage.py asserts that there is exactly one
- definition of each part and that every reader gets it from here, because a
- shared helper that somebody copies again is four sources of truth instead of
- one - and four is how this started.
- What is written down here, and what is not
- ------------------------------------------
- The FORMAT is written down: the entry jump is three bytes, the header is
- sixteen, initmem's first bytes and the two header words it reads, and the load
- bias. Those are this project's knowledge of what it emits, and asking the code
- under test for them would let the code under test satisfy every check by
- agreeing with itself.
- The SIZES are not written down. The runtime's size was written down twice
- (385, then 391), beside a comment saying it tracked the runtime; it did not,
- and each stale value sent a checker into the middle of the code, where it
- reported a well-formed image as a compiler fault. A duplicated constant that
- has silently drifted is not an independent check, it is a second source of
- truth that lies, and it lies in the direction of looking like the thing under
- test is broken.
- Measuring is not the same as asking the compiler: everything here reads the
- emitted file. tests/check_runtime.py is where the runtime's size is pinned on
- purpose, and tests/run_com_tests.sh prints the size it measures on every run.
- Usage: this is a library; import it from a check.
- """
- # ---- the format, written down -------------------------------------------
- # ENT_SZ and HDR_SZ are the layout, and the load bias below is a decision this
- # project made about how a DOS .COM is loaded. Those stay written down: they
- # are knowledge about the format, not sizes that move when code is edited.
- ENT_SZ = 3 # E9 lo hi, the entry jump (Compiler.Inittur)
- HDR_SZ = 16 # 5 header words + 3 buffer words (Compiler)
- # The load bias. A DOS .COM's first byte is at CS:0100 and CS = DS, so an image
- # offset K is at DS:(K + 0100h); every ABSOLUTE address the image contains has
- # to carry it, or it points 0100h low and - since the code region and the
- # runtime are both below the bias - almost always lands inside the runtime
- # instead of inside the data. Relative encodings must not carry it, because
- # both of their operands shift together.
- # Restated, not asked of the code under test. See Runtime.LoadBias.
- LOAD_BIAS = 0x100
- # initmem's prologue, which is the runtime's only reader of the program header.
- # The displacements +4 and +6 below are the whole point of this constant: the
- # header checks in every checker read hdrDS at +4 and hdrHeap at +6, and
- # initmem has to read the SAME two words or it clears the wrong range. It read
- # +8 (hdrMax, which the compiler patches to 0), so it zeroed nothing at all and
- # nothing here noticed - the emitted loop was perfectly well formed, it just
- # never ran. Asserting these bytes together with the header offsets is what
- # closes that gap.
- #
- # It is the first eleven bytes of the runtime, so it sits at ENT_SZ, not 0.
- HEAD = "8B F0 8B 54 04 8B 4C 06" # MOV SI,AX / MOV DX,[SI+4] / MOV CX,[SI+6]
- HDR_DS_WORD = 4 # header word holding the data base
- HDR_HEAP_WORD = 6 # header word holding the data end
- assert [int(HEAD.split()[4], 16), int(HEAD.split()[7], 16)] == \
- [HDR_DS_WORD, HDR_HEAP_WORD], \
- "initmem no longer reads the two header words the checks verify"
- # ---- what the file says, measured ---------------------------------------
- def find_header(d):
- """The image offset of the program header in `d`, or None.
- The header is eight words, and the layout says what they are (offsets here
- are BYTES into the header, which is why HDR_DS_WORD is 4 and not 2 - the
- words are two bytes each):
- +0 1 hdrFlag, always 1
- +2 code end + bias hdrCS
- +4 data base + bias hdrDS, where data base = hdrOff + 1000h
- +6 data end + bias hdrHeap, which is hdrDS + dataBytes
- hdrDS ties the header to its OWN offset: the data base is header offset +
- 1000h, and hdrDS is that with the load bias added, so the offset is
- recoverable from the file with no remembered runtime size - which is the
- whole point. A candidate is accepted only if hdrFlag is 1, hdrDS satisfies
- that equation, hdrHeap is above hdrDS (a heap below its own base is not a
- layout, it is a coincidence), and hdrCS leaves room for the header itself.
- initmem is the only code in the image that reads the header, so its bytes
- are pinned at ENT_SZ: a candidate that also has them there is not a
- coincidence in the code stream.
- A header that cannot be found is returned as None rather than as the
- nearest guess, because a caller that guesses reports confidently about the
- wrong bytes - which is what the restated runtime size used to do.
- """
- head = bytes(int(x, 16) for x in HEAD.split())
- if d[ENT_SZ:ENT_SZ + len(head)] != head:
- return None # no runtime: nothing to measure
- for off in range(ENT_SZ, len(d) - HDR_SZ + 1):
- w = (lambda b: int.from_bytes(d[off + b:off + b + 2], "little"))
- if w(0) != 1: # hdrFlag
- continue
- if w(HDR_DS_WORD) != off + 0x1000 + LOAD_BIAS: # hdrDS
- continue
- if w(HDR_HEAP_WORD) <= w(HDR_DS_WORD): # hdrHeap
- continue
- if w(2) - LOAD_BIAS < off + HDR_SZ: # hdrCS
- continue
- return off
- return None
- def entry_target(d):
- """The image offset the entry JMP at file offset 0 lands on, or None.
- A .COM is entered at its first byte, so byte 0 is `E9 rel16' and the first
- instruction of the program is at ENT_SZ + rel16 (Compiler.Inittur). This
- is where EXECUTION starts, which is a different statement from where the
- runtime ends: the program header sits between the two, and a check that
- measures only one of them cannot tell that it has started its sweep in the
- wrong place. A check that measures both is asked to make them agree.
- """
- if len(d) < ENT_SZ or d[0] != 0xE9:
- return None
- return ENT_SZ + int.from_bytes(d[1:ENT_SZ], "little", signed=True)
|