summary_topspeed-corpus.md 3.8 KB

Step: TopSpeed V3 corpus validation

Tag v3-topspeed-corpus. V3 suite 151/151, fixpoint OK (unchanged — this touches only the sidecar).

Why

After self-hosting, the next "normal" step was to exercise the front-end on a real corpus: the historical TopSpeed/JPI V3 tree at $HOME/bin/Dos/M2/TS-V3 — 662 files (352 .DEF + 310 .MOD).

V3's own M2.atg is the classic/Redux subset and would not take this dialect (pragmas (*#...), conditional compilation (*%...), H/C/ B literals, ::=, address constructors [seg:ofs]^, CLASS/IS, manifest constants, ...). The project already carries a sidecar grammar for it, compiler/src/TopSpeed-V3-M2.atg, whose deliverable is parsing the dialect (unlowered constructs end in the standard 230, parse-now / lower-later). It had never been compiled and run over the full corpus.

What was done

  1. Built the sidecar — tools/topspeed-grammar/build.sh copies the grammar + the frames + the V3 SymTab/QbeGen/FileIO, runs CR -m -C, gm2 -fiso, links build/TSM2.
  2. Fixed API drift vs the current SymTab/QbeGen: SymTab.NewArray → NewOpenArray; FixPending/SetProcRes now return BOOLEAN; QbeGen.EndModule now takes the module name.
  3. Grammar fixes the corpus exposed:
    • tagless variant CASE : T OF (anonymous selector);
    • manifest constants without CONST (NAME = value;);
    • trailing | in a variant list.
  4. Conditional compilation (*%T cond*) / (*%F cond*) / (*%E*) — the sidecar now uses its own scanner frame (tools/topspeed-grammar/scanner.frm; V3's compiler/src/scanner.frm is untouched). A CondFilter runs on the source buffer after it is read: it evaluates each directive against a define set and blanks the inactive regions (keeping CR/LF so line numbers survive). This matters because the corpus's %F/%T branches wrap partial statements (an IF … THEN in one branch, the body/END outside), so exactly one branch must be selected. The define set is {_fdata, _mthread, _fptr, _fcall} (in InitDefs), tuned against the corpus; it makes the prose/legacy %F blocks inactive while keeping the plain parseable branches. (Pragmas (*#...) were never a problem — they lex as ordinary comments.)
  5. Corpus harness — tools/topspeed-grammar/corpus.sh runs TSM2 over every file in a scratch tree (so listings never touch the corpus) and classifies each as parse-OK (only the deliberate 230/semantic markers) or syntax-bad (a real Coco/R error).

Result

files
corpus 662
parse-OK (230/semantic only) 644
syntax-bad 18

Coverage rose 590 → 627 → 631 → 640 → 644 as the fixes landed (grammar fixes then conditional compilation).

Remaining 18 syntax-bad

A mix of:

  • a few files whose conditional branch still selects the "wrong" side for a universal define set (FIO, FIOR, LOADER, SHTHEAP, …);
  • non-conditional gaps: a typed set constructor in an argument list (GetMenu("…", CharSet{'N','P','T'}, ch)), address constructors, and a few files with malformed/commented-out regions (END (*QMLB …).

Reproduce

cd tools/topspeed-grammar
./build.sh          # -> build/TSM2
./corpus.sh         # -> 644 parse-OK / 18 syntax-bad

Notes

  • V3's own compiler (M2.atg, compiler/src/scanner.frm) and the self-hosting path are untouched; the sidecar has its own scanner frame and is built on demand (build/ is git-ignored).
  • The sidecar now tracks the V3 SymTab/QbeGen API; if those drift again, build.sh is where the small patch belongs.

Files

compiler/src/TopSpeed-V3-M2.atg (three grammar fixes), tools/topspeed-grammar/{build.sh,corpus.sh,scanner.frm,README.md}, .gitignore (build dir), this doc.