After each session: did the served truth actually shape what shipped? Four honest labels, a scorekeeper that audits wins without ever minting them — and one call on how the hard cases get judged. Three independent review rounds (9/10).
Recommendation: D now, B pre-registered as the upgrade at scale.
Self-report · re-run the consistent subset · re-run everything · human judgment on the small pile (recommended). One line is enough.
A4 becomes the scorekeeper, not a fifth arm; the passively-served arms’ sealed hit rests on the mechanical endpoint the prereg already defines as their fallback, scored like-for-like against the control; consumption is measured and reported, never score-gating. Sealed pass/kill numbers untouched.
Reference —
the reframe & amendment ·
the four labels ·
the audit boundary ·
the four options ·
risks & counters.
Canonical source: docs/ai/SPEC-a4-mechanism.md (v2.1, three independent review rounds, 9/10,
ready-for-red-pen). Nothing builds until your word.