| 123456789101112131415161718192021222324252627282930313233343536373839404142434445464748495051525354555657585960616263646566676869707172737475767778798081828384858687888990919293949596979899100101102103104105106107108109110111112113114115116117118119120121122123124125126127 |
- # modrm11.s -- establish the mod=11 half of the 16-bit ModR/M table using
- # GNU as as an independent ENCODER.
- #
- # modrm19.s measures the memory forms (mod=00/01/10) by executing them on a
- # real 8086 under qemu: it stores a marker through each encoding and reports
- # which physical address received it. mod=11 is not a memory form at all --
- # the r/m field names a register -- so it needs a different oracle, and this
- # file is it.
- #
- # Why the assembler rather than another qemu probe
- # ------------------------------------------------
- # The obvious extension of modrm19.s is "store into the register, then report
- # which register took it". That does not work, and the reason is worth
- # recording because it is a trap rather than a puzzle:
- #
- # * The comparison needs a register to hold the expected value, and every
- # register is a candidate, so the check clobbers what it is measuring.
- # * r/m=100 is SP, and the marker value is not a legal stack pointer. The
- # very next `call` pushes its return address through SS:BEEF and leaves
- # SP at BEED, so by the time any check runs, the evidence is gone. A
- # probe that reports "not found" for SP is reporting its own design, not
- # the hardware. (This is not hypothetical: the first version of that
- # probe did exactly this, and also cleared AX "because AX is never an
- # r/m target" -- which is true of the mod=11 list people remember, and
- # false of the real one, where r/m 000 IS AX. It reported two
- # impossible answers and a plausible-looking third.)
- # * Repairing that needs the check inlined between the store and the next
- # call, so every case gets its own hand-written code block and its own
- # chance of a typo -- and a typo there produces a plausible wrong answer,
- # which is the worst kind.
- #
- # The assembler has none of those problems. `as --32` with `.code16` is a
- # correct 16-bit x86 *encoder*, and it was already the encoder half of the
- # FCML cross-validation (38/38 agreement on a separate probe). Asking it to
- # encode `movw %bx, %si` and getting `89 DE` back is a direct, unambiguous
- # statement that in mod=11 the r/m field 110 means SI.
- #
- # So the two files are complementary halves of one argument.
- # modrm19.s execution -> mod=00/01/10 effective addresses
- # this file encoding -> mod=11 register identities
- # Nothing in the table is taken from memory, and the two oracles share no code.
- #
- # This file is in four groups, and modrm11.py checks all four differently:
- #
- # 1. the r/m column 8 x `movw %bx, <reg>` -> low three bits of ModRM
- # 2. the reg column 8 x `movw <reg>, %di` -> bits 5..3 of ModRM
- # 3. the byte list 6 x 8-bit moves -> AL CL DL BL AH CH DH BH
- # 4. the anchors 4 hand-checking bytes -> the table itself
- #
- # Group 4 is what makes the check non-vacuous. Groups 1-3 are self
- # referential in a dangerous way: they compare this file against the
- # assembler, so editing this file just changes the claim and the assembler
- # faithfully re-encodes it. A check like that cannot fail on a wrong table
- # unless the table is what moved. The anchors are different -- they are
- # encodings nobody types by hand, so they carry information this file's
- # author did not supply, and a table shifted by one position cannot satisfy
- # them.
- #
- # Encode and check (see modrm11.py):
- # as --32 -o modrm11.o modrm11.s
- # objcopy -O binary -j .text modrm11.o modrm11.bin
- .code16
- .text
- # --- 1. the r/m field, read straight off the low three bits -------------
- # Each of these is `movw %bx, <reg>`, i.e. opcode 89 with reg=BX (011), so
- # ModRM = 11 011 rrr and the low three bits ARE the r/m code for the register
- # named on the right. One instruction per r/m code, in r/m order. Opcode 89
- # is MOV r/m,r, so the r/m field is the DESTINATION -- the register on the
- # right of the AT&T line. Both of those directions matter, and getting
- # either backwards produces a table that is wrong everywhere.
- movw %bx, %cx # r/m 000
- movw %bx, %dx # r/m 010
- movw %bx, %bx # r/m 011
- movw %bx, %sp # r/m 100
- movw %bx, %bp # r/m 101
- movw %bx, %si # r/m 110
- movw %bx, %di # r/m 111
- movw %bx, %bx # r/m 011 again -- there is no second BX
- # --- 2. the reg field, read off bits 5..3 ------------------------------
- # `movw <reg>, %di` is opcode 89 with rm=DI (111), so ModRM = 11 rrr 111 and
- # bits 5..3 are the r/g code, which is the SOURCE. All eight codes are
- # encodable, AX included: 89 with reg=AX and mod=11 is an ordinary
- # MOV r/m,r, not an accumulator short form -- unlike ADD/ADC/AND/OR/SBB/SUB/
- # XOR/CMP, where /0 means the accumulator and the ModRM byte does collapse.
- movw %ax, %di # reg 000
- movw %cx, %di # reg 001
- movw %dx, %di # reg 010
- movw %bx, %di # reg 011
- movw %sp, %di # reg 100
- movw %bp, %di # reg 101
- movw %si, %di # reg 110
- movw %di, %di # reg 111
- # --- 3. the 8-bit forms, where the list differs and direction flips -----
- # 88 /r is MOV r/m8,r8 (reg is the SOURCE); 8A /r is MOV r8,r/m8 (reg is the
- # DESTINATION). Both use the identical ModRM byte for the same pair of
- # registers, so the byte alone cannot tell you the direction -- the opcode
- # can. The 8-bit list is also its own: AL CL DL BL AH CH DH BH, which agrees
- # with the word list at every code except 100, where it is AH rather than
- # SP. This is the other trap in the table, and Runtime.mod documents it.
- movb %al, %dh # 88 C6 -> DH := AL
- movb %dh, %al # 88 F0 -> AL := DH
- movb %al, %dl # 88 C2 -> DL := AL
- movb %dl, %al # 88 D0 -> AL := DL
- movb %al, %bl # 88 C3 -> BL := AL
- movb %bl, %al # 88 D8 -> AL := BL
- # --- 4. the anchors -----------------------------------------------------
- # Four instructions no one writes by hand, each of which pins one cell of the
- # table. If the ModRM column in modrm11.py is shifted by one, every one of
- # these disagrees with it. They are the reason this file can fail for a
- # reason other than "someone edited the claims".
- #
- # They are spelled as literal bytes, not as mnemonics, because the point is
- # the exact sequence: `as` picks opcode 89 rather than 8B for a
- # register-to-register move, so asking it for `movw %sp, %bp` gets 89 E5 and
- # never 8B EC. Both mean MOV BP,SP and the two differ in which field holds
- # which register -- which is exactly what the anchor is here to pin.
- # modrm11.py supplies the expected decode for each, so a byte that does not
- # decode as its comment claims still fails.
- .byte 0x83, 0xC4, 0x08 # ADD SP, 8 rm=100 -> SP
- .byte 0x83, 0xC6, 0x02 # ADD SI, 2 rm=110 -> SI
- .byte 0x8B, 0xEC # MOV BP, SP reg=101 rm=100
- .byte 0x8B, 0xE5 # MOV SP, BP reg=100 rm=101
|