Your red-pen reshaped this: the defect-fixes build now, and the design calls compete as pre-registered experiment arms in cold sessions — four ways of serving settled truth, measured against a control, with the end-state value named so every arm aims at it.
Draft v4 · your 3 notes folded · floors await your ratify, arms await your roster
What your pen changed: v3 defended each design call by argument plus one review. v4 splits the spec — the defect-fixes build now (they make any experiment honest and commit us to no shape), and the shape choices compete empirically: pre-registered arms, cold-session venue, live instruments, pass bars and kills set before any arm can negotiate. The strongest-alternative bar includes one frontier scan for an approach we didn’t invent ourselves.
01
The end state — what this is ultimately for
At maturity, a verification act, once paid, stays paid. Any agent, in any session, on any machine, touching any file, can see what’s already proven about it — how fresh, what kind of proof, what it survived — without re-deriving any of it. Re-verification spends only where the ground actually shifted. Cold workers start productive instead of blind. The same mechanics carry to the practice genome, where “one fleet-wide change, receipts instead of per-site QA” is the business unlock.
Usefulness is three numbers: verification work honestly avoided · cold-start time-to-productive · false greens introduced by trusting served truth — and the third must stay zero; it’s a floor, not a goal. Every experiment below aims at the first two without ever touching the third.
02
Floors — build now, no shape committed
Each is a defect-fix or an instrument every arm needs equally.
Floor 1
The verifier can never overstate
The vacuous-HOLDS trap gets its born-red guard, and every verdict names its court — “ground unchanged” vs “proof re-run.” A verifier that can overstate poisons every arm equally.
Floor 2
Evidence that cannot rot
Hash-addressed, capped, secret-scanned, heal-immune blob machinery — built once as an instrument. Whether and how evidence is served is an arm variable; that it survives is not.
Floor 3
Lead verdicts stop being invisible
The deposit seam for merge gates and walks — without it every experiment undercounts the ledger. Machinery-vs-habit gets an observed coverage number before anything is hard-wired.
Floor 4
Consumption is measured, never vibes
Every candidate surface writes a consult edge tagged by arm. Arms ride kill-switch flags with a byte-identity-off proof — an experiment that can’t be turned off is a rollout.
03
The competition — EXP-READER
Call A, the highest-value question: what actually makes a cold agent consume settled truth? Four arms, sealed before running, in the honest venue — cold sessions whose task crosses verdict-bearing ground.
A0 — control: serve nothing new. Today’s count-glance only. The null arm keeps every other number honest — an arm that can’t beat A0 dies without ceremony.
A1 — session-start serving. Court-labeled verdict lines on the map every session already opens: what’s proven, how fresh, where the evidence lives.
A2 — decision-moment serving. The verdict arrives at the un-skippable moment — on the file being edited, or right before a verification leg is spent. (Affordable as an arm at all because of this week’s 106x hook fix.)
A3 — binding-first. No summary at all: the claim travels with its runnable receipt — the reader is handed something it can re-run, not something it must believe. The lineage favorite: agents refused asserted trust six out of six; this is the receipted counterpart, finally tested in the field.
↘ go deeper — instruments, venue, kills
Instruments: reads-plane consult edges (mechanical, per arm) · re-derivation avoided, judged by an independent grader — never the session, never the author · token/wall-clock delta vs baseline · safety leg: any wrong build traceable to trusting a served claim is a program STOP, not a data point. Arms rotate across qualifying sessions (no self-selection). Pass bars sealed before running; one frontier scan for the field’s strongest external alternative folds into the roster before sealing. Call B (evidence register: stored blob vs runnable-receipt-only vs inline quote) and Call C (lead-banking shape, decided by Floor 3’s observed coverage number) run alongside. Program kill: no arm moves cold behavior across the session budget ⇒ the bet resizes and the roadmap says so out loud.
04
Your calls
Ratify the floors building now
Four small PRs, no shape committed. Each is owed regardless of which arm wins — a defect-fix or an instrument.
★ The arm roster — add or kill
Especially Call A. This is the place to name an approach we haven’t thought of; the frontier scan adds one from outside, but your read on what would move an agent is exactly the kind of arm the roster wants.
★ Timebox and session budget
How many qualifying cold sessions per arm before the program-level kill engages. The fleet generates these naturally as work happens; the budget decides how long we let the question run.