Experiment spec · v2 draft · Warp’s adversarial pass folded · for red-pen

The Habitat Experiments

Warp dreamed the factory as a habitat; Quanta judged the doctrine; you asked the sharper question — run it and see. Four bounded probes on the real factory: hunches that outlive sessions, an honesty ledger, a right to fix the wobbly table, a fleet board. Every probe carries its own death condition, written before it runs — and Warp’s adversarial pass already repaired four bars that couldn’t fail honestly.

Draft for Robert’s annotation · probes not approved · nothing runs yet

Locked, unless you unlock them

Probes, not beliefs. Approving this approves four bounded experiments — nobody is asked to believe in the habitat.
Muzzle before medium. A hunch may aim attention anywhere and count as evidence nowhere. Enforced by a machine check, not a norm.
Kill bars pre-named. Every probe's death condition is written here, before it runs, where it can't be negotiated later.
Default-dead. A probe with no verdict session by week 6 dies on its own. Nothing becomes furniture by lingering.
Verification changes 0% during E1. The honesty ledger only watches; pricing is a future, separate ratification.
Push-pheromone out of scope. Nothing whispers into a working agent's context in any of these four probes.

The journey of one hunch

Follow a single noticing through the habitat — each stop is one probe. Today, at 4pm, a worker deep in a job thinks: “the auth module is lying about timeouts.”

  1. Today: the hunch dies at 5pm. The session ends, the handoff carries only what was proven, and tomorrow's worker meets the auth module flat — no nose, just notes. Every primitive below exists because of this moment.
    ↘ go deeper — why this is the right first-principles gap
    Human engineers compound unproven noticings into taste; agents are reborn flat because the genome's truth-classes are facts-only (claim · receipt · currency stamp). The proposal's strongest idea is a truth-class for the unproven-but-felt — provided it can never masquerade as fact. Provenance: Warp's proposal §1; Quanta's position paper §1.
  2. E1 — the ledger starts keeping score (shadow mode, starts on approval). Every claim a worker makes carries its stated confidence — mandatory, bound to the receipt lines the report contract already defines — and verification continues exactly as today, full depth, nothing skipped. Skipping the confession is itself recorded. After four weeks we don't argue about whether trust-pricing is safe; we compute what it would have missed. Kill bars: pricing dies for any claim-class where the replay shows even one load-bearing miss (which classes count as load-bearing is pre-registered, not decided after the fact) — the whole ledger dies as UNINFORMATIVE if under 70% of receipt lines carry confidence by week 2 (an optional confessional collects only the easy claims — honest bars need denominators) — or if everyone's confidence clusters at one safe number (gaming) — and any class under 20 checked claims reads UNMEASURED, never a number.
    ↘ go deeper — ledger schema, classes, counterfactual replay
    Report contracts (bay brief templates, workflow verdict schemas) add MANDATORY per-claim confidence: 0.0–1.0, bound to the receipt lines the contract already defines — workers don't control receipt granularity, so they can't game which claims get scored (Warp B1). A receipt line without confidence is ledgered outcome: unconfessed. Join against verification outcomes into append-only JSONL: ~/.claude/engine-bay/calibration-ledger.jsonl. Classes: gate-count, fixture-verdict, root-cause (load-bearing, pre-registered), scope-claim, design-safety-claim, env-note (not load-bearing); unregistered classes get their stamp from the lead AT refutation time, before the replay. E2 suspicion confidences are EXCLUDED (censored outcomes would poison curves). Curves: per-class reliability + Brier, no-dep read-only script; a class under n=20 reads UNMEASURED, never a number. Counterfactual: replay window under pricing rule R, count refuted claims R would have skipped. Kill mechanics: missed load-bearing > 0 ⇒ dead per class; coverage <70% by week 2 ⇒ UNINFORMATIVE; >80% of confidences within ±0.05 of the mode ⇒ uninformative; Brier drift >0.1 week-over-week ⇒ history can't price the future. Live receipt for why novelty matters: the same builder session was calibrated on “gate 275/0” and wrong on “false refusals are rare” within one hour (PR #266, round 3).
  3. E2 — the hunch outlives its thinker (the load-bearing probe). The 4pm noticing becomes a real deposit with a decay clock: tomorrow's worker inherits it at birth, it stays alive only if someone freshly observes something (re-reading it doesn't count), and a machine check makes it structurally unable to pass or block anything — it aims probes, never verdicts. Kill bars: dead if under 10% of visited hunches convert into a receipted finding — visited means a worker actually entered the territory while the hunch lived; counting unvisited ones would make E2 die of fleet scheduling, not hunch quality (Warp's catch). Conversion credit counts only when the hunch was named at probe-dispatch time; stamping it afterward is laundering and inadmissible. Dead too if ghost-chases outnumber conversions 2:1. A muzzle breach that LANDS in a recorded verdict halts every probe; a breach the fence catches is the system working — data, never a halt.
    ↘ go deeper — schema, the muzzle as a gate, both-pole fixtures
    New deposit kind: {id: susp-<hash>, claim, confidence, provenance{session, agent, receipt_ptr?}, born, ttl_days:14, renewals[{ts, new_receipt_ptr}], territory[paths], status: alive|decayed|converted|refuted}. Birth-read quanta suspicions <path> expands territory via the reach graph (computable, never declared prose) and drops decayed items entirely. Renewal refuses duplicate or absent receipt pointers — citation ≠ observation, mechanically. THE MUZZLE [M]: a gate leg asserts no verdict-bearing artifact carries a susp- id in an evidence position; aimed_by: susp-X in a finding's provenance block is legal (that's how conversion is measured) — provenance ≠ evidence. Fixtures red-first, both poles: verdict-citing-suspicion ⇒ gate red; decayed item at birth ⇒ absent; echo-renewal ⇒ refused; zero-suspicion repo ⇒ zero new friction. The muzzle gate ships with an INVENTORY of evidence positions per verdict-artifact schema, pinned by fixtures — without the inventory the gate is vacuous or cry-wolf (Warp). Conversion = VISITED suspicions ending converted (territory intersected ≥1 dispatch while alive — computable from bay ledger paths × territory reach; unvisited excluded from the denominator). aimed_by admissible ONLY when recorded at probe-dispatch time; retro stamps are inadmissible. On refutation, a small receipted clean-check fact is minted so the same ghost isn't re-chased after decay. Ghost-cost = probe legs on refuted suspicions vs conversions.
  4. E3 — the wobbly table gets fixed without a ticket. The same worker notices a dead path in the docs. Instead of a Post-it, it takes a small declared lease — capped at 3 a week fleet-wide, mechanical rot only, visible to everyone, revocable — and the fix arrives as a normal PR through normal review. This is the garden lane you already ratified, with an actor attached. Kill bars: one warning for a lease that strays into norm territory, dead on the second; dead instantly if a lease ever breaks trunk; dead if half the leases get rejected as pointless.
    ↘ go deeper — the mechanical fence, scope classes, ledger
    Lease = declared deposit {lease_id, agent, scope_class, files[], opened, outcome}; scope_class ∈ {dead-path, stale-doc-pointer, reach-confirmed-dead-code, fixture-hygiene, map-freshness}. Fence [M], INVERTED to an allowlist (Warp M3): a lease is valid only if its diff falls entirely within its declared scope_class's file surface; the named denylist (rules/, gate scripts, .github/workflows/, CLAUDE.md) always refuses; anything else is simply not a lease — no judgment call about whether a file “encodes routing” is ever needed, undecidables fail closed by never qualifying. No self-merge, unchanged law. Outcomes ledgered: merged | rejected | breached, plus elapsed time per lease — E3's hidden tax on primary work is verdict input.
  5. E4 — the fleet sees each other (read-only; starts alongside E1 and E3). One live board: who holds what, working on what, since when. The six-hour standoff over one uncommitted file dies here — not because agents whisper to each other, but because they can look. It needs nothing from E2 (rows come from ledgers that exist today — Warp's catch), so it doesn't wait. Nothing pushes into a working agent's context; that half stays out of scope until E2's muzzles exist as gates. Kill bar: dead if nobody chooses to read it — under 1 human-initiated look/day after four weeks; automated pings are tagged and excluded, else the bar can't fail. Standoff-prevention at this scale is honestly anecdote: the verdict input is traffic plus at least one documented avoided collision, labeled as such.
    ↘ go deeper — sources, instrumentation, what stays fenced
    Pure aggregation over state that already exists: bay ledger (runs, branches), workflow journals (live legs), task list, clone/worktree registry. Surface: quanta fleet CLI + one served page — rows {agent, territory, claim, started, heartbeat}. Consultations counted (page hits + CLI calls → local JSONL). Standoff incidents lead-logged at occurrence, compared at verdict. The push half — live world-model updates into working contexts — is a FUTURE spec gated on the E2 muzzle plus a verifier independence fence (verification legs sealed from any live field for their whole leg; reviewer-is-never-the-author at environment altitude).
  6. The verdict bench — where probes come to die or earn. Each probe gets a verdict session at week 4 — scheduled at probe start with a named owner (the lead), so healthy probes can't die of calendar neglect: keep, kill, or extend once, judged ONLY against its pre-named bar, from raw numbers the lead recomputes, each verdict carrying the probe's cost line. No session by week 6 = dead by default. A floor breach that LANDS (a merged bad lease, a verdict recorded on a suspicion) halts everything pending your review — but a breach the fences CATCH is the system working: data, never a halt. Punishing fences for firing is the #266 cry-wolf mistake, and we're not making it here. The habitat runs as a guest of the truth layer, never its landlord.
    ↘ go deeper — global rules, named judges, build discipline
    Frame-floor judges, named in advance: (a) the kill bars — reality numbers set here before any probe can negotiate them; (b) Robert's red-pen on this page, plus Warp as the differently-framed reviewing mind. No probe is graded by instruments it defined: E-numbers are computed [M] from raw ledgers; keep/kill is Robert's [J]. Each build increment runs the normal sequence (graded review, red-first fixtures, gate). Step-back answered in writing: the deeper “experiment platform” plan is refused — four concrete probes on existing machinery; a platform only if a fifth probe ever arrives.

What could go wrong (honest)

Activity reads as value. A humming suspicion field feels alive regardless of worth. Only conversion and cost numbers are admissible at verdicts; hum is not.
The ratchet. Features that feel good survive dead bars. Verdict dates plus default-dead exist precisely to break it.
Sandbagging. Workers state 0.5 on everything and are never wrong. E1's clustering check makes uninformative confidence itself a kill condition.
Inherited paranoia. Stale hunches haunting every newborn. Decay, drop-at-birth, renewal-only-on-evidence, and the ghost-cost bar.
The habitat grading itself. The deep failure mode. The muzzle is mechanical, the halt rule is global, and no probe judges its own success.

Your calls

1 — Approve the four probes?

This is “approve these experiments,” not “believe in the habitat.” Each is bounded, muzzled, reversible, and dies on its bar.

2 — Set the kill-bar numbers.

Every number is a proposal: 10% conversion · 2:1 ghost ratio · 3 leases/week · one warning then dead · <1 read/day. Your red-pen aims at the one thing machinery can't see — whether these bars measure what you'd actually kill for.

3 — Sequence.

Proposed (revised after Warp's pass): E1, E3, and E4 start together on approval — all three near-zero risk, and the read-only board needs nothing from E2. E2 starts when its build lands. The push half of the pheromone idea stays out of scope entirely. Confirm or reverse.

4 — Verdict cadence.

One omnibus session at week 4, or per-probe verdicts as each matures? Default-dead fires at week 6 either way.

Reference walk: Probes — E1 ledger · E2 suspicions · E3 leases · E4 board  |  Rules — verdicts & halt  |  Risks  |  Your calls. Canonical source: docs/ai/SPEC-habitat-experiments.md · provenance: Warp proposal + Quanta position, both in ~/quanta-build-clones/.