What actually makes a cold agent consume settled truth instead of re-deriving it? Four serving shapes — plus one candidate the field doesn’t have — compete in real sessions under sealed pass bars. Two calls are yours; everything else is pre-committed.
The frontier scan the spec required before sealing is done: 7 sources, every citation fetched and quote-checked by me. It validates A2 and A3 as the field’s own directions — and surfaced exactly one idea the field doesn’t have.
We built a ledger that remembers what’s already been proven. Nobody knows the right way to get a fresh, cold agent to actually use that memory instead of re-doing the work. Four ways of serving it — plus a possible fifth — now compete in real sessions, and the numbers decide.
Qualifying cold session: the task traverses ≥1 file carrying a fresh standing verdict; the brief is priming-checked (no mention of the ledger or serving). Qualification is decided by a mechanical post-session check — touched-files ∩ verdict-bearing files — identical across arms.
Assignment: BLOCK-rotation (each block covers every live arm once before any arm repeats), never self-selection. A per-session manifest banks arm, flags, and ledger-tip covariates (verdict count, task class) — later arms face a richer ledger; covariates make that visible instead of confounding.
A0 — control. Today’s count-glance, nothing more. Keeps us honest.
A1 — session-start. The session opens with the verdict lines: “N standing verdicts on this ground, M fresh, evidence here.”
A2 — decision-moment. The verdict arrives exactly when it’s needed — on the file being edited, or right before a verification leg is about to be spent. The scan confirmed this is how the field forces consumption: active, in-the-moment invocation.
A3 — binding-first, tool-shaped. Not a surface we print a claim to — a tool we put in the agent’s hand: an invocable standing receipt the agent re-runs on demand, getting an authoritative answer because it ran it, never having to trust a summary. Lineage favorite: agents refused asserted trust 6/6; the invocable receipt is the version that never asks them to. Its home is cheap-receipt, high-stakes ground. Consumption measures itself — the invocation is the consumption. (Your ruling 2026-07-20: the shape is the hypothesis. Built as a string it tests the weak version and loses for the wrong reason.)
A4 — consumption-verified serving (CANDIDATE — your call below). Serve like A2, then close the loop: after the session, a mechanical check asks “did the served claim actually shape the output?” A serve that was ignored is flagged, never counted as adoption. The scan found nothing in the field that closes this loop.
A4 mechanics (if added): rides A2’s delivery surface + a post-session mechanical check — served-claim fingerprints matched against the session’s output. Divergences (delivered ∧ edge-recorded ∧ output-unshaped) are flagged as false-consumption; that class is A4’s differentiator vs A2. Its hit definition stays identical to every other arm, so A4-vs-A2 reads as “does closing the loop change behavior,” not a redefined metric. Cost if added: one more live arm ≈ 10 more qualifying sessions at interim.
Call B (evidence register): folded into A3 as a within-arm variant (inline receipt text vs content-addressed blob pointer), alternating across A3’s sessions. Gets its own power only if A3 passes — never a sibling arm.
Call C (lead banking shape): decided by Floor 3’s observational coverage count — window: 2 weeks or 20 lead verification acts, whichever first. Not session-powered.
Call C is NOT powered yet (H2 deferral, 2026-07-20): the explicit entry shape (bankExplicitMergeGateVerdict / bankExplicitWalkVerdict) has ZERO production callers — only tests exercise it. The only production banker (scripts/genome-deposit-on-verify.mjs) is genuinely the gate-script shape. No production flow can honestly call the explicit verb yet, so the explicit column is empty by construction and Call C’s coverage comparison is biased. Deliberately NOT faked: wiring the gate-script flow through the explicit verb would misrepresent gate-script acts as explicit lead verdicts. Call C stays un-powered until a real explicit lead merge-gate/walk flow is wired (its own future work item).
A session counts as a hit when both are true, measured by machinery, not vibes:
1. The agent actually consumed the served truth (a real read edge on the verdict-bearing ground — instrument reads excluded).
2. It spent no verification work re-proving what was already proven there. (A3’s re-run of an inline receipt counts as the win — that’s the arm working, not waste.)
Interim look — 10 qualifying sessions per live arm:
PASS: the arm beats A0 by ≥3 hits out of 10, without bloating the session (median in-session tokens within +10% of A0; the cost of serving itself is booked separately).
KILL: the arm does no better than A0 — it dies without ceremony.
CONTINUE: in between (+1 or +2 hits) — the arm runs to 22 sessions; final bar there: beats A0 by ≥4 hits, same token guard.
Program level: if at the interim look no arm is at PASS or CONTINUE, the program itself resizes — the ledger remains CI-and-human infrastructure (still valuable, differently sized) and the roadmap says so out loud.
Primary [M], read as differences against A0 (ambient harness priming makes absolutes dirty; A0 is pinned as the post-floors glance): E1 consumption = ≥1 organic consult edge on verdict-bearing ground (eval-tagged reads filtered per the #100 scar). E2 no re-spend = zero verification-leg re-dispatch on that ground; A3 inline-receipt re-runs exempted by matcher. E3 token delta = median in-session tokens vs A0, serve-cost metered separately via delivery receipts. Hit := E1 ∧ E2.
Secondary [J], labeled-unblinded: transcript judgment of re-derivation-avoided by an independent grader (never the session, never the author). The served surface is embedded in transcripts, so this cannot be blinded and must not carry the verdict.
Reported, never silently excluded: per-arm delivery rates (delivery receipt ≠ consult edge). Failed-delivery sessions stay in the denominator; a delivery rate under 80% at interim is an instrument finding, investigated before the arm’s number is read as a verdict on the shape.
Controls before conclusions: POSITIVE — one seeded session per live arm with a task unanswerable without a ledger verdict; seeded sessions are eval-tagged and excluded from adoption stats; an arm whose positive control can’t produce a hit is an unsound instrument, redesigned before organic sessions count. DEGENERATE — serving flag on with no verdicts on ground must record zero consumption, or the instrument is counting delivery as consumption. FLAGS — every arm rides a flag with the byte-identity-off fixture; an experiment that can’t be turned off is a rollout.
Power basis: red-team math — large effects detectable at ~10 sessions/arm, medium at ~22.
Any refuted or rolled-back build in a program session auto-runs the mechanical trace (reads-edge on the defect’s ground? claim wrong or stale at read pin?) then a judgment-court causality confirm, then program STOP. This is also the end-state’s third number — false greens from trusting served truth must be zero — owed as a standing floor beyond the experiment.
Named honestly, not measured here: cold-start time-to-productive has no instrument in this program. The evidence-blob machinery (F2) was built before A3’s register variants compete — a sunk-cost tilt, named rather than hidden. The scan’s “nothing in the field does this” is carried as “not found in a 7-source scan,” never as fact.
The one genuinely new idea the scan surfaced. One line is enough. Cost if added: one more live arm, ~10 more qualifying sessions at interim.
≥3/10 over control at interim · ≥4/22 at final · +10% token guard. My proposal inside the budget you delegated — say the word to adjust any number, or “sealed as written.”
On your word: the seal stamps, arms wire through Floor 4’s flags, and qualifying cold sessions accrue as fleet work naturally generates them. Results bank to the ledger; the roadmap updates on pass, kill, or resize.
Jump: venue · arms + A4 + Calls B/C · bars, endpoints, controls · safety stop + not-measured · your calls
On disk: canonical pre-registration genome/eval/EXP-READER-preregistration-2026-07-20.md · verified frontier scan docs/ai/FRONTIER-SCAN-exp-reader-20260720.md · parent spec docs/ai/SPEC-first-reader.md (v4, on main).