Generated from committed canonical artifacts. As of 2026-08-03T20:33:04Z.
| Evidence | Result | Boundary |
|---|---|---|
| Optimizer history | 120 signed cells; 87 real attempts; 33 verified; 0 false passes | Performance gate open; retrospective |
| ROMA operation control | 24 cells; Citadel 4/6; direct local 2/6 | Evidence passed; performance failed |
| Prospective runtime | 1/3 passed; claude-code/claude-opus-5 | Integration only; actual cash unknown |
| Prospective local v1 | adaptive 27/36 verified, 9 failed, 0 unknown; baseline 24/36 verified, 11 failed, 1 unknown | Frozen gate failed; timeout sensitivity -3.5% GPU energy; identity gate false |
| Capability-profile follow-up | profile 24/36 verified, 12 failed, 0 unknown; baseline 24/36 | Matched baseline cell completion; 15.7% more GPU energy; gate failed |
| Representative repository pilot | profile 6/12 verified; baseline 6/12; 6 unique fixture tasks | 7.1% less measured GPU energy missed the 20% gate; evidence failed |
| Fresh local repository calibration | candidate 8/12; baseline 8/12; 12 unique tasks | 34.6% more measured GPU energy; invalid baseline; evidence failed |
| Hybrid calibration | candidate 12/12; baseline 12/12; 4 Claude calls avoided | 28.4% cost reduction missed 30%; paired sensitivity 30.1%; evidence failed |
| Calibrated hybrid v2 | candidate 12/12; baseline 12/12; 8 local attempts, 1 recovery | 38.7% comparison-cost reduction at 100.0% relative completion; evidence passed |
| Public holdout fast pilot | controller 3/16; direct Claude 2/16; 24 distinct repositories | 1.26% lower comparison cost; baseline validity failed; diagnostic only |
| Fresh-clone onboarding | 5/5 command stages in 28.17s | Doctor health unknown; model execution not-attempted |
- The 120-cell optimizer result is retrospective and its performance gate remained open because cost coverage was incomplete.
- The 24-cell ROMA result proves control and evidence integration; its efficiency hypothesis failed.
- The Claude prospective cell proves one real runtime integration, not savings or broad reliability.
- V1 recorded 27/36 versus 24/36 verified cells, but its apparent GPU savings reverse when one same-route timeout pair is excluded; it does not support a savings claim.
- The separately frozen 72-cell capability-profile follow-up matched baseline cell completion, but verification-driven escalations increased measured GPU energy and modeled GPU cost.
- The capability-profile v2 corrigendum discloses v1-informed design, task-template overlap, and deterministic repetition semantics without changing the frozen negative result.
- The 24-cell representative repository-operation shakedown matched 6/12 verified cells under both policies with zero false passes and zero path violations. Its 7.1% measured GPU-energy reduction missed the frozen 20% gate, so the evidence result is failed and no general savings claim is permitted.
- The representative shakedown contains six unique fixture tasks repeated twice per policy; timing repetitions are not independent tasks and the small fixture set is not production generalization.
- A fresh 12-task local-only repository calibration preserved 8/12 verified completions under both policies but used 34.6% more GPU energy; its baseline also failed the frozen overall and high-risk validity floors.
- The first 12-task Claude-plus-local hybrid preserved 12/12 completions and reduced comparison cost 28.4%, missing the frozen 30% gate. A post-run paired-cost sensitivity reaches 30.1% but does not change the failed verdict.
- The separately frozen calibrated hybrid v2 preserved 12/12 completions, reduced provider-reported and locally modeled comparison cost 38.7%, and passed every frozen gate on twelve new author-selected synthetic tasks.
- Hybrid v2 establishes a positive result only inside its preregistered support envelope on one model pair and one machine. Task selection was author-controlled, comparison USD is not the operator subscription bill, and production generalization is not claimed.
- The secondary public-holdout pilot used 24 distinct outside-authored repositories, sealed all routes before evaluation calls, and published 32 official verdicts with no evaluator unknowns.
- The public-holdout controller verified 3/16 tasks versus 2/16 for direct Claude and used 1.26% less comparison cost, but direct Claude passed only 12.5% overall and 0% in three strata. The frozen in-sample signal is not a valid generalized optimization result.
- The public-holdout result makes retrieval, edit representation, baseline strength, and calibration power the next technical bottlenecks. It does not establish production reliability or actual cash savings.
- The fresh-clone proof completed five command stages; doctor semantic health remained unknown, and no real-user utility or model-task completion is claimed.
- Actual end-to-end cash remains unknown wherever subscription allocation or whole-system energy is unmeasured.
- GPU energy arithmetic reconstructs from retained average watts and request wall duration across both local studies; raw 500 ms power samples were not retained.
Manifest: sha256:87ab6303dbc59ecd53b3a08756f55c5b218782de8f023d86d4accf90832ca593