From 51bfa57067f343e63abc02e3fef41c8080ce0d62 Mon Sep 17 00:00:00 2001 From: Davide Cappellini Date: Sun, 27 Sep 2026 13:55:19 +0200 Subject: [PATCH] j163: RESULT - the firing floor is a clean negative, and the offline energy corpus missed the live game by 200x 450 runs/arm x 2 arms (15 opponents x 30 runs x 3 rounds, conc=6, 0 failed, 0 never started, 0 env mis-set), frozen binary d9a39c3b8472, TR_MOVEMENT=tfil. Round-win 40.30% -> 39.70% (-0.59 pp, CI -3.90..+2.72, sign-flip exact-2^15 p=0.7676, MDE 4.73 pp); wins/run 0.200 -> 0.182 (MDE 0.1420); damage/run 113.65 -> 112.70 (p=0.5298, MDE 4.19). Deviation disclosed: 30 runs/opponent instead of 42 (throughput 14-22 runs/min vs 22.7-23.2 assumed); full 15-opponent panel and both arms kept, runs reduced. Headline is the mechanism failure: firing was suppressed on 0.04% of ticks, not the 9.6% the offline ruler predicted (~200x smaller); only 4.8% of shots happen in the low-energy zone and the floor removed 14% of those; rounds ending at self energy <=0 were 40.6% live vs 61.2% implied by the offline corpus; median self energy at death 14.1 -> 15.5. The pre-registered 'we cannot measure the damage cost directly' call was correct. DO NOT ADOPT. TR_RAM_FLOOR_ENERGY stays 0.0, do not re-test. The null does not prove the knob inert - the mechanism barely fired. --- docs/ram_floor_exhaustion_ab.md | 104 ++++++++++++++++++++++++++++++++ 1 file changed, 104 insertions(+) diff --git a/docs/ram_floor_exhaustion_ab.md b/docs/ram_floor_exhaustion_ab.md index 31866b1..f7a0c9b 100644 --- a/docs/ram_floor_exhaustion_ab.md +++ b/docs/ram_floor_exhaustion_ab.md @@ -260,3 +260,107 @@ subsetting, no dropping opponents, no re-running to chase a p-value, no reinterpreting the bar after seeing the data. **A clean null is a fully acceptable result** — and a null here licenses only "no effect >= MDE is detectable at this design", never "the knob is harmless". + +--- + +## MEASURED + +*(appended after the battles — everything above was committed first, at +`6cfb169`)* + +### MEASURED — the live A/B, 450 runs/arm, 2700 rounds (j163) + +* **Provenance.** Frozen 15-opponent movement panel + (`tools/ab/panel_movement.txt`), 15 x 2 x **30 runs** x 3 rounds = + **450 runs/arm, 2700 rounds**, `conc=6`, **0 failed, 0 never started**. + Frozen binary `d9a39c3b8472…`, `TR_MOVEMENT=tfil` in both arms; the only + difference is `TR_RAM_FLOOR_ENERGY` `0` (`A_floor0`, reference) vs `5` + (`B_floor5`). +* **Env verification — 0 mis-set.** Every run's `[env]` boot block was checked + against its arm as the session progressed, live at **24 / 193 / 410 / 826 / + 900** runs completed: no disagreement at any checkpoint. Per-arm + `TR_ENV_FILE` in the session's own directory, botdir without a `.env`, loader + does not walk up parents. This is contamination control #2 from the + pre-registration, satisfied. +* **DEVIATION FROM THE PRE-REGISTRATION, DISCLOSED.** The design asked for + **42 runs/opponent** (1890 battles, MDE ~0.10 wins/run). Measured throughput + was **14-22 runs/min**, not the assumed 22.7-23.2, so the wall-clock cost of + the pre-registered n was not affordable. As the pre-registration required, + the **panel and the arms were NOT shrunk**: the full 15 opponents and both + arms were kept and the **runs per opponent were reduced to 30**. The MDE + actually reached is reported below and is the honest resolution limit of this + run. No opponent was dropped, no arm was re-run to chase a p-value. + +**Pooled dashboard (descriptive, NOT the verdict):** + +| arm | runs | dmg/run | wins/run | round wins | round win rate | +|---|---:|---:|---:|---:|---:| +| `A_floor0` | 450 | 113.65 | 0.200 | — | **40.30%** | +| `B_floor5` | 450 | 112.70 | 0.182 | — | **39.70%** | + +**Verdict layer** (per-opponent paired deltas, arm − reference; the sign-flip +is the exact 2^15 permutation the pre-registration names as the decision test): + +| metric | mean Δ | 95% CI | p(sign-flip) | sign test | Wilcoxon p | **MDE reached** | +|---|---:|---|---:|---:|---:|---:| +| **round-win rate** | **-0.59 pp** | [-3.90, +2.72] | **0.7676** | 1.00 | 0.84 | **4.73 pp** | +| wins/run | -0.0178 | [-0.117, +0.082] | 0.7676 | 1.00 | 0.84 | 0.1420 | +| damage/run | -0.95 | [-3.88, +1.98] | 0.5298 | 1.00 | 0.84 | 4.19 | + +### HEADLINE FINDING — the mechanism barely fired; the offline energy corpus did not survive contact with the live game + +1. **Suppression was 0.04% of ticks, not the 9.6% the offline ruler predicted** — + a **~200x** smaller effect. The pre-registration's own mechanism metric + (`share of ticks with firing suppressed`, predicted ~9.6%, median suppressed + run ~53 ticks) is the number that failed, and it failed by two orders of + magnitude. +2. **Only 4.8% of shots are ever taken in the low-energy zone, and the floor + removed 14% of those.** The gate therefore touches a small slice of a small + slice: a bot that almost never wants to fire at low energy. The pre- + registration's expected damage cost (~2 damage/run) was the arithmetic + consequence of the 9.6% figure; with 0.04% it is ~200x smaller still, which + is why the damage MDE (4.19) is unreachable by construction and not by bad + luck. +3. **Rounds ending at self energy <= 0: 40.6% live vs 61.2% implied by the + offline corpus** (the j162 baseline the pre-registration quoted). The + recorded-fixture corpus over-states how often we die broke by ~1.5x. This is + the second, independent way the same corpus mis-called the live game. +4. **Median self energy at death 14.1 -> 15.5** — the floor moved the death + energy by +1.4, real but tiny, and nowhere near the "we sit disabled" state + the floor was built for. +5. **The pre-registered blind spot was called correctly.** The pre-registration + states, in advance, that we expect to be unable to measure the damage cost + directly and that a damage null must not be re-read afterwards as evidence. + That call was right, and it is the reason this run cannot be misread. + +### VERDICT — DO NOT ADOPT + +1. **The pre-registered verdict rule is not met.** Adopt required round-win + rate favouring `B_floor5` with per-opponent sign-flip p < 0.05 AND the + effect at or above the reported MDE. Observed: **-0.59 pp against**, p = + **0.7676**, under the 4.73 pp MDE. Round wins and wins/run are the same + null (p = 0.7676, MDE 0.1420 wins/run). +2. **`TR_RAM_FLOOR_ENERGY` stays `0.0`.** Do not adopt, do not ship, and **do + not re-test this knob.** The design that could resolve a real effect does + not exist at an affordable run count, and the mechanism it was built to + suppress is nearly absent in the live game. Re-running buys resolution on an + effect that is not there. +3. **A null here does NOT prove the knob inert — the opposite.** The mechanism + fired on 0.04% of ticks. This run licenses only: "no effect >= 4.73 pp of + round-win rate at 450 runs/arm". It says nothing about the ~0.04% of ticks + it did suppress, because too few of them existed to measure. +4. **The generalisable finding is the corpus, not the knob.** The offline + energy corpus over-predicted both the size of the low-energy firing window + (9.6% -> 0.04%) and the rate of dying broke (61.2% -> 40.6%). An offline + ruler built on recorded fixtures is only as representative as the fixtures; + the death-energy corpus does not represent the live energy ledger. Future + offline rulers for exhaustion must be calibrated against a live + death-energy distribution before their predictions are pre-registered as + expectations, not just as a rationale. + +**Ship state: unchanged. `TR_RAM_FLOOR_ENERGY=0.0` and `TR_RAM_ENEMY_ENERGY=0.0` +remain the shipped defaults, and the ram path keeps its pre-j160 behaviour.** +This is the sixth mechanism-positive-or-presumed / outcome-not-positive result +in the campaign (j144, j145, j146, j147, j159, j163) — and the first where the +mechanism was not merely ineffective but **~200x smaller than the offline ruler +said it would be**.