|
|
4 일 전 | |
|---|---|---|
| .. | ||
| README.md | 4 일 전 | |
| modrm11.py | 4 일 전 | |
| modrm11.s | 2 주 전 | |
| modrm19.s | 2 주 전 | |
| run_modrm19.py | 4 일 전 | |
The 16-bit ModR/M table in ../../Runtime.mod is the input every hand-written
emitter in that file depends on, and it was wrong once. The wrong version was
not a code bug — every byte it produced decoded cleanly — it was a
documentation bug, and it survived a compile, a disassembly golden, and a
structural check, because a well-formed instruction that does the wrong thing
is indistinguishable from a correct one if you never say what it should do.
So the table is not written from memory here. It is measured, by two independent oracles that share no code, and each half is measured the way that half can be.
| file | oracle | establishes |
|---|---|---|
modrm19.s |
execution on a real 8086 under qemu-system-i386 |
the mod=00/01/10 effective addresses |
modrm11.s + modrm11.py |
encoding by GNU as (.code16), decoded by FCML |
the mod=11 register identities |
The split is not arbitrary. mod=11 is not an address form at all — the r/m
field names a register — so the probe that measures addresses by writing a
marker and reading back the address that received it cannot be extended to
it. See the header of modrm11.s for the three reasons a naive extension
fails; they are worth reading, because the first version of that probe
produced two impossible answers and a plausible-looking third.
modrm19.s runs 24 cases (8 r/m values x 3 mod values). For each it clears
low memory, stores a marker through the encoding under test, then scans for
the word and reports the offset it landed at. BX=1000 DI=2000 SI=0030
BP=0040, so every candidate address is distinct and the answer is the
offset's arithmetic, not a judgement call.
To run it (needs qemu-system-i386, as, objcopy):
cd shell
as --32 -o ../tmp/modrm19.o tests/probe/modrm19.s
objcopy -O binary -j .text ../tmp/modrm19.o ../tmp/modrm19.bin
python3 tests/probe/run_modrm19.py
run_modrm19.py assembles, wraps the code in a boot sector (EB 3C at
offset 0, code at 0x3E, 55 AA at 0x1FE — the whole thing must fit in
0x3E + len(code) <= 512), boots it under qemu with the serial port
captured to a file, and decodes the 24 groups of lo hi 0x20 into the table
in Runtime.mod, asserting it.
The result, as measured:
mod=00 : [BX+SI] [BX+DI] [BP+SI] [BP+DI] [SI] [DI] disp16 [BX]
mod=01 : ...+disp8 for each of the above, with rm=110 -> [BP]+disp8
mod=10 : ...+disp16 for each of the above
The rm=011 cells are [BP+SI] and [BP+DI], not [BX+SI] and [BX+DI];
that is the single most common misreading of the table and it is invisible
in a disassembly.
One cell is not covered by execution. mod=10, rm=001 lands at
BX+DI+disp16 = 4234, which is outside the probe's scan window, so that one
case reports "not found". It is recorded as a gap, not as a result, and it is
covered statically instead by the recipes in Runtime.mod and by
fcml_vs_objdump.py.
modrm11.py assembles modrm11.s, disassembles the result with FCML
(../disasm16.py), and requires four things to agree:
.s asked for,as emitted,EXPECT in modrm11.py,Runtime.mod, via the ModRM field of each instruction.The result, as measured:
mod=11 r/m field 000 001 010 011 100 101 110 111
AX CX DX BX SP BP SI DI
mod=11 reg field same list, same order
8-bit mod=11 AL CL DL BL AH CH DH BH
and the direction, which depends on the opcode and not on the ModRM byte:
88 /r MOV r/m8, r8 89 /r MOV r/m16, r16 reg is the SOURCE
8A /r MOV r8, r/m8 8B /r MOV r16, r/m16 reg is the TARGET
So 88 C6 and 8A C6 are the same ModRM byte and opposite instructions.
The word and byte lists differ at exactly one cell: 100 is SP in one, AH in
the other.
Comparing the .s against as cannot fail on its own: edit the .s and as
faithfully re-encodes the new claim, so the two always agree. That failure
mode was found by deliberately corrupting the .s and watching the check
stay green. Two things close it:
EXPECT, a hard-coded byte sequence per instruction, asserted
independently of the .s text. Editing the .s away from the table now
fails.83 C4 08 (ADD SP,8), 83 C6 02 (ADD SI,2), 8B EC
(MOV BP,SP), 8B E5 (MOV SP,BP). These are encodings nobody types by hand,
so they carry information the .s's author did not supply, and a table
shifted by one cell cannot satisfy them. They are spelled as literal
.byte directives because as picks opcode 89 over 8B for a
register-to-register move and would never emit 8B EC for a mnemonic.../nonvacuity.sh exercises all of this: it corrupts a table cell, restores
the exact table this project once shipped (AX dropped off the front, a
duplicate BX invented at the end), edits the .s away from EXPECT, edits
the 8-bit list, and corrupts an anchor byte — asserting each one goes red
for the stated reason.
The table that shipped was
rm: CX DX BX SP BP SI DI BX
which is the correct list with AX dropped off the front and a duplicate BX invented at the end. Every code is one too low except 100, which lands on SP either way. So:
89 DC reads as MOV SP,BX under both tables, but the emitter that
intended MOV SI,BX wrote it. That is MovSiBx, and 39 DC is CmpSiBx.The correct encodings: MOV SI,BX is 89 DE (reg=BX is the source, r/m=SI
is the destination) and CMP SI,BX is 39 DE.
A third case is instructive because it is the same mistake pointed the other
way. MovSiAx was written 8B C0. Under the shifted table that looks like
MOV AX,AX — a no-op, so the fix was to reach for the byte the shifted table
said was SI. It is worth being explicit that the fix 8B F0 is correct and
8B C0 is not, because the direction is the part that is easy to get
backwards: for 8B the reg field is the destination, so 8B F0
(reg=110=SI, r/m=000=AX) is MOV SI,AX, which is what the name asks for.
rt_exec.py is not part of this. unicorn 2.1.4 decodes 16-bit ModRM
incorrectly on this machine, and it is the only component in the chain that
is wrong, which makes it useless as a reference — see SUMMARY.md.