Warp dreamed the factory as a habitat; Quanta judged the doctrine; you asked the sharper question — run it and see. Four bounded probes on the real factory: hunches that outlive sessions, an honesty ledger, a right to fix the wobbly table, a fleet board. Every probe carries its own death condition, written before it runs — and Warp’s adversarial pass already repaired four bars that couldn’t fail honestly.
Follow a single noticing through the habitat — each stop is one probe. Today, at 4pm, a worker deep in a job thinks: “the auth module is lying about timeouts.”
confidence: 0.0–1.0, bound to the receipt lines the contract already defines — workers don't control receipt granularity, so they can't game which claims get scored (Warp B1). A receipt line without confidence is ledgered outcome: unconfessed. Join against verification outcomes into append-only JSONL: ~/.claude/engine-bay/calibration-ledger.jsonl. Classes: gate-count, fixture-verdict, root-cause (load-bearing, pre-registered), scope-claim, design-safety-claim, env-note (not load-bearing); unregistered classes get their stamp from the lead AT refutation time, before the replay. E2 suspicion confidences are EXCLUDED (censored outcomes would poison curves). Curves: per-class reliability + Brier, no-dep read-only script; a class under n=20 reads UNMEASURED, never a number. Counterfactual: replay window under pricing rule R, count refuted claims R would have skipped. Kill mechanics: missed load-bearing > 0 ⇒ dead per class; coverage <70% by week 2 ⇒ UNINFORMATIVE; >80% of confidences within ±0.05 of the mode ⇒ uninformative; Brier drift >0.1 week-over-week ⇒ history can't price the future. Live receipt for why novelty matters: the same builder session was calibrated on “gate 275/0” and wrong on “false refusals are rare” within one hour (PR #266, round 3).{id: susp-<hash>, claim, confidence, provenance{session, agent, receipt_ptr?}, born, ttl_days:14, renewals[{ts, new_receipt_ptr}], territory[paths], status: alive|decayed|converted|refuted}. Birth-read quanta suspicions <path> expands territory via the reach graph (computable, never declared prose) and drops decayed items entirely. Renewal refuses duplicate or absent receipt pointers — citation ≠ observation, mechanically. THE MUZZLE [M]: a gate leg asserts no verdict-bearing artifact carries a susp- id in an evidence position; aimed_by: susp-X in a finding's provenance block is legal (that's how conversion is measured) — provenance ≠ evidence. Fixtures red-first, both poles: verdict-citing-suspicion ⇒ gate red; decayed item at birth ⇒ absent; echo-renewal ⇒ refused; zero-suspicion repo ⇒ zero new friction. The muzzle gate ships with an INVENTORY of evidence positions per verdict-artifact schema, pinned by fixtures — without the inventory the gate is vacuous or cry-wolf (Warp). Conversion = VISITED suspicions ending converted (territory intersected ≥1 dispatch while alive — computable from bay ledger paths × territory reach; unvisited excluded from the denominator). aimed_by admissible ONLY when recorded at probe-dispatch time; retro stamps are inadmissible. On refutation, a small receipted clean-check fact is minted so the same ghost isn't re-chased after decay. Ghost-cost = probe legs on refuted suspicions vs conversions.{lease_id, agent, scope_class, files[], opened, outcome}; scope_class ∈ {dead-path, stale-doc-pointer, reach-confirmed-dead-code, fixture-hygiene, map-freshness}. Fence [M], INVERTED to an allowlist (Warp M3): a lease is valid only if its diff falls entirely within its declared scope_class's file surface; the named denylist (rules/, gate scripts, .github/workflows/, CLAUDE.md) always refuses; anything else is simply not a lease — no judgment call about whether a file “encodes routing” is ever needed, undecidables fail closed by never qualifying. No self-merge, unchanged law. Outcomes ledgered: merged | rejected | breached, plus elapsed time per lease — E3's hidden tax on primary work is verdict input.quanta fleet CLI + one served page — rows {agent, territory, claim, started, heartbeat}. Consultations counted (page hits + CLI calls → local JSONL). Standoff incidents lead-logged at occurrence, compared at verdict. The push half — live world-model updates into working contexts — is a FUTURE spec gated on the E2 muzzle plus a verifier independence fence (verification legs sealed from any live field for their whole leg; reviewer-is-never-the-author at environment altitude).This is “approve these experiments,” not “believe in the habitat.” Each is bounded, muzzled, reversible, and dies on its bar.
Every number is a proposal: 10% conversion · 2:1 ghost ratio · 3 leases/week · one warning then dead · <1 read/day. Your red-pen aims at the one thing machinery can't see — whether these bars measure what you'd actually kill for.
Proposed (revised after Warp's pass): E1, E3, and E4 start together on approval — all three near-zero risk, and the read-only board needs nothing from E2. E2 starts when its build lands. The push half of the pheromone idea stays out of scope entirely. Confirm or reverse.
One omnibus session at week 4, or per-probe verdicts as each matures? Default-dead fires at week 6 either way.
Reference walk: Probes — E1 ledger · E2 suspicions · E3 leases · E4 board | Rules — verdicts & halt | Risks | Your calls. Canonical source: docs/ai/SPEC-habitat-experiments.md · provenance: Warp proposal + Quanta position, both in ~/quanta-build-clones/.