Spec · the consumption meter · for red-pen

The A4 Meter

After each session: did the served truth actually shape what shipped? Four honest labels, a scorekeeper that audits wins without ever minting them — and one call on how the hard cases get judged. Three independent review rounds (9/10).

Ready for your red-pen · one decision + one amendment ack · nothing builds until your word
  1. The gap. We’re about to spend real sessions testing ways of serving settled truth to cold agents. The scoreboard counts a “hit” when an agent read the served truth and didn’t re-prove it. But a read is not a use. An agent can open the page, ignore it, and still look like adoption. Worse: it can read the truth and ship the opposite — and today nothing would notice. A4 closes that gap: after each session, a check asks — did the served claim actually shape what shipped? Your frame, which this spec serves: A4 is the meter on whether the genome functions as capital. Without it, “workers born knowing” is unfalsifiable.
    ↘ go deeper — inputs & mechanism
    Inputs per session: the delivery receipt (served claim IDs + canonical text + serve cost), the session’s output artifacts (final diff, banked docs, final message), and the reads-plane edges. Stage 1 extracts (served claim, output claim) pairs on shared ground via the graph layer — pairs, never whole transcripts (the long-input weakness of these checkers is a verified field finding, dodged by construction). Stage 2 runs a pinned sentence-level consistency check: support / neutral / contradict. Stage 3 adjudicates the hard question — “consistent, but was it load-bearing?” — per your call below.
  2. The honest reframe (this changes what you’re deciding). The prereg penciled A4 in as a fifth racing arm costing ~10 more sessions. The independent review caught that this doesn’t hold together: in every version worth building, A4’s serving is identical to A2’s and the agent never knows the check exists — so A4 can’t race A2; nothing about it changes what the agent does. A4 is not another horse. It’s the scorekeeper — it runs over the other arms’ sessions and tells us which of their “hits” were real. Better news than the prereg assumed: near-zero extra sessions, and it’s exactly the piece that makes the passively-served arms measurable at all — the gap the seal was blocked on. One line of the prereg gets amended to say so, recorded, on your word; the sealed numbers themselves don’t move.
    ↘ go deeper — the amendment’s exact content
    (1) A4 is an instrument over A1/A2 sessions, not a roster arm — no extra qualifying sessions, no block-rotation cell. (2) A1/A2’s sealed hit = E2 (no re-spend) — the prereg’s own stated fallback for those arms, kept mechanical and blindable; the A1/A2-vs-A0 “beats by ≥N hits” comparison is scored E2-to-E2 (A0’s E2 half). Consumption is measured and REPORTED by the instrument — never hit-gating. (3) The one exception: the self-report option below would make A4 a behavioral arm again (asking agents to report changes what they do) at ~10 sessions with the comparison-contamination cost named. A named alternative — seating the mechanical screen as those arms’ E1 — is yours to choose instead, with two costs stated: different hit-courts across arms, and the screen alone counts “agreement” as consumption (the exact false positive the instrument exists to remove).
  3. What the meter says about each served claim — four honest labels:
    Used. The output leans on it, and it was genuinely load-bearing. The real hit.
    Ignored. Delivered, no trace in what shipped. Today this can masquerade as adoption; under A4 it can’t.
    Contradicted. The output fights the served claim — the agent had the truth and shipped the opposite. Flagged loudly, every time — but a flag alone never halts the program (that takes a real broken build plus a confirmed causal link).
    Independently derived. The output agrees but didn’t need the serve. Agreement is not consumption; counting it would flatter the serving. Screened out.
  4. The boundary the review held me to: the scorekeeper audits wins — it never mints them. The human judgment inside it can’t be blinded (the judge sees what was served), so it never becomes part of an official score; the arms’ sealed scores stay on the purely mechanical endpoint, and the meter’s labels ride alongside as the honesty column. If a winning arm’s wins turn out mostly hollow (“read it, didn’t need it”), the column says so out loud before anyone calls it a win — that check is pre-committed, not discretionary.
    ↘ go deeper — the tripwire’s fail-safe direction
    The hollowness tripwire (USED-fraction bar, set at calibration freeze) mirrors the prereg’s sealed delivery-rate rule: it pauses interpretation pending investigation; it never changes the mechanical count. Round-3 review stress-tested it as a possible back door for the label to gate the verdict: it isn’t — and it fails safe: the unblinding bias inflates USED (optimism), which fires the tripwire less, so it can never falsely block a good arm. At worst a hollow win passes unflagged — an audit-layer miss, never a false pass/kill of the sealed comparison.
  5. Your call — how “used” vs “would have figured it out anyway” gets decided. The cheap mechanical layer sorts every claim into consistent / no-trace / contradicted. The hard remaining question has four candidate answers:
    A — agent self-report. Ask agents to cite what they used; check citations mechanically. Why not: it changes the served text itself (the comparison stops being clean) and measures claimed use, not actual use. Also the one option that turns A4 back into a ~10-session behavioral arm.
    B — remove-and-re-run, on the consistent subset. Re-run the moment without the claim; if the output changes, it was load-bearing. Mechanical and blindable — but proven only for single-shot generation; our sessions are long trajectories where influence can arrive early and shape the whole path, which a single-step re-run can miss. Unproven ground, named.
    C — remove-and-re-run on everything. Same caveat, ~10x the compute.
    D — human judgment on the small remaining pile (my recommendation). At this experiment’s size that pile is ~50–150 judgments total, and the experiment already budgets an independent human grader. Avoids B’s unproven ground entirely; every judgment banks as labeled ground truth for the mechanical upgrade later. Its own costs, named just as plainly: the judge can’t be blinded (exactly why their call stays the audit column and never touches a score), it doesn’t scale past human hours, and two judges can disagree (an agreement check is part of calibration). B-vs-D is a real trade — a blindable machine on unproven ground, or a valid human judgment that can never gate a verdict.

    Recommendation: D now, B pre-registered as the upgrade at scale.

    ↘ go deeper — how the meter earns its numbers first
    Before any real session counts, the instrument passes a born-red battery: a seeded session that must read USED · an unrelated-task serve that must read IGNORED · a scripted opposite-shipper that must fire CONTRADICTED · a superficially-contradictory-but-fine pair that must NOT fire (the false-alarm control) · a re-derivable fact the screen must hand to adjudication · a diffuse session-start bundle and a jointly-load-bearing pair (the shapes the meter is most likely to miss) · a no-serve session that must emit zero labels · a gutted-checker control the battery must catch · a non-ASCII pair (byte-exact canonicalization). Separately — because passing your own test cases proves fit, not generalization — 10–15 real sessions get double-labeled by humans as a held-out check, with the agreement bar pre-set. The cited accuracy numbers for an off-the-shelf checker failed my source-verification and were struck; the battery replaces the card numbers.
  6. What could go wrong (named, with the counter built in):
    The meter itself could be wrong. Planted test cases with known-true answers plus a double-labeled spot-check on real sessions — both before anything counts, both with pre-set bars; a miss means redesign, never bar adjustment.
    The meter could be harsher on one arm. The session-start arm serves broad bundles whose influence is diffuse — easy to undercount vs the precise decision-moment arm. Named bias, dedicated test cases — and because labels never gate the sealed score, an undercount can’t kill an arm; it can only misdescribe the audit column.
    A false alarm on “contradicted.” Deciding whether two technical sentences really disagree is a judgment call. The checker is pinned, proven against a planted false-alarm case before it scores anything, and its flag is a signal for investigation — never an automatic stop.
  7. On your word. Hinge call (A / B / C / D) → the prereg’s one-line amendment records the scorekeeper reframe → the meter builds and passes its planted cases + organic spot-check → it wires in as the passively-served arms’ consumption measure and the honesty column → the seal’s last instrument gap closes. Report-only by construction; killing it later costs the column, nothing else.

The questions for your pen

The hinge: A, B, C, or D?

Self-report · re-run the consistent subset · re-run everything · human judgment on the small pile (recommended). One line is enough.

The prereg amendment — “recorded,” or red-pen it.

A4 becomes the scorekeeper, not a fifth arm; the passively-served arms’ sealed hit rests on the mechanical endpoint the prereg already defines as their fallback, scored like-for-like against the control; consumption is measured and reported, never score-gating. Sealed pass/kill numbers untouched.

Reference — the reframe & amendment · the four labels · the audit boundary · the four options · risks & counters. Canonical source: docs/ai/SPEC-a4-mechanism.md (v2.1, three independent review rounds, 9/10, ready-for-red-pen). Nothing builds until your word.