j163: RESULT - the firing floor is a clean negative, and the offline energy corpus missed the live game by 200x
450 runs/arm x 2 arms (15 opponents x 30 runs x 3 rounds, conc=6, 0 failed, 0 never started, 0 env mis-set), frozen binary d9a39c3b8472, TR_MOVEMENT=tfil. Round-win 40.30% -> 39.70% (-0.59 pp, CI -3.90..+2.72, sign-flip exact-2^15 p=0.7676, MDE 4.73 pp); wins/run 0.200 -> 0.182 (MDE 0.1420); damage/run 113.65 -> 112.70 (p=0.5298, MDE 4.19). Deviation disclosed: 30 runs/opponent instead of 42 (throughput 14-22 runs/min vs 22.7-23.2 assumed); full 15-opponent panel and both arms kept, runs reduced. Headline is the mechanism failure: firing was suppressed on 0.04% of ticks, not the 9.6% the offline ruler predicted (~200x smaller); only 4.8% of shots happen in the low-energy zone and the floor removed 14% of those; rounds ending at self energy <=0 were 40.6% live vs 61.2% implied by the offline corpus; median self energy at death 14.1 -> 15.5. The pre-registered 'we cannot measure the damage cost directly' call was correct. DO NOT ADOPT. TR_RAM_FLOOR_ENERGY stays 0.0, do not re-test. The null does not prove the knob inert - the mechanism barely fired.
This commit is contained in:
@@ -260,3 +260,107 @@ subsetting, no dropping opponents, no re-running to chase a p-value, no
|
||||
reinterpreting the bar after seeing the data. **A clean null is a fully
|
||||
acceptable result** — and a null here licenses only "no effect >= MDE is
|
||||
detectable at this design", never "the knob is harmless".
|
||||
|
||||
---
|
||||
|
||||
## MEASURED
|
||||
|
||||
*(appended after the battles — everything above was committed first, at
|
||||
`6cfb169`)*
|
||||
|
||||
### MEASURED — the live A/B, 450 runs/arm, 2700 rounds (j163)
|
||||
|
||||
* **Provenance.** Frozen 15-opponent movement panel
|
||||
(`tools/ab/panel_movement.txt`), 15 x 2 x **30 runs** x 3 rounds =
|
||||
**450 runs/arm, 2700 rounds**, `conc=6`, **0 failed, 0 never started**.
|
||||
Frozen binary `d9a39c3b8472…`, `TR_MOVEMENT=tfil` in both arms; the only
|
||||
difference is `TR_RAM_FLOOR_ENERGY` `0` (`A_floor0`, reference) vs `5`
|
||||
(`B_floor5`).
|
||||
* **Env verification — 0 mis-set.** Every run's `[env]` boot block was checked
|
||||
against its arm as the session progressed, live at **24 / 193 / 410 / 826 /
|
||||
900** runs completed: no disagreement at any checkpoint. Per-arm
|
||||
`TR_ENV_FILE` in the session's own directory, botdir without a `.env`, loader
|
||||
does not walk up parents. This is contamination control #2 from the
|
||||
pre-registration, satisfied.
|
||||
* **DEVIATION FROM THE PRE-REGISTRATION, DISCLOSED.** The design asked for
|
||||
**42 runs/opponent** (1890 battles, MDE ~0.10 wins/run). Measured throughput
|
||||
was **14-22 runs/min**, not the assumed 22.7-23.2, so the wall-clock cost of
|
||||
the pre-registered n was not affordable. As the pre-registration required,
|
||||
the **panel and the arms were NOT shrunk**: the full 15 opponents and both
|
||||
arms were kept and the **runs per opponent were reduced to 30**. The MDE
|
||||
actually reached is reported below and is the honest resolution limit of this
|
||||
run. No opponent was dropped, no arm was re-run to chase a p-value.
|
||||
|
||||
**Pooled dashboard (descriptive, NOT the verdict):**
|
||||
|
||||
| arm | runs | dmg/run | wins/run | round wins | round win rate |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| `A_floor0` | 450 | 113.65 | 0.200 | — | **40.30%** |
|
||||
| `B_floor5` | 450 | 112.70 | 0.182 | — | **39.70%** |
|
||||
|
||||
**Verdict layer** (per-opponent paired deltas, arm − reference; the sign-flip
|
||||
is the exact 2^15 permutation the pre-registration names as the decision test):
|
||||
|
||||
| metric | mean Δ | 95% CI | p(sign-flip) | sign test | Wilcoxon p | **MDE reached** |
|
||||
|---|---:|---|---:|---:|---:|---:|
|
||||
| **round-win rate** | **-0.59 pp** | [-3.90, +2.72] | **0.7676** | 1.00 | 0.84 | **4.73 pp** |
|
||||
| wins/run | -0.0178 | [-0.117, +0.082] | 0.7676 | 1.00 | 0.84 | 0.1420 |
|
||||
| damage/run | -0.95 | [-3.88, +1.98] | 0.5298 | 1.00 | 0.84 | 4.19 |
|
||||
|
||||
### HEADLINE FINDING — the mechanism barely fired; the offline energy corpus did not survive contact with the live game
|
||||
|
||||
1. **Suppression was 0.04% of ticks, not the 9.6% the offline ruler predicted** —
|
||||
a **~200x** smaller effect. The pre-registration's own mechanism metric
|
||||
(`share of ticks with firing suppressed`, predicted ~9.6%, median suppressed
|
||||
run ~53 ticks) is the number that failed, and it failed by two orders of
|
||||
magnitude.
|
||||
2. **Only 4.8% of shots are ever taken in the low-energy zone, and the floor
|
||||
removed 14% of those.** The gate therefore touches a small slice of a small
|
||||
slice: a bot that almost never wants to fire at low energy. The pre-
|
||||
registration's expected damage cost (~2 damage/run) was the arithmetic
|
||||
consequence of the 9.6% figure; with 0.04% it is ~200x smaller still, which
|
||||
is why the damage MDE (4.19) is unreachable by construction and not by bad
|
||||
luck.
|
||||
3. **Rounds ending at self energy <= 0: 40.6% live vs 61.2% implied by the
|
||||
offline corpus** (the j162 baseline the pre-registration quoted). The
|
||||
recorded-fixture corpus over-states how often we die broke by ~1.5x. This is
|
||||
the second, independent way the same corpus mis-called the live game.
|
||||
4. **Median self energy at death 14.1 -> 15.5** — the floor moved the death
|
||||
energy by +1.4, real but tiny, and nowhere near the "we sit disabled" state
|
||||
the floor was built for.
|
||||
5. **The pre-registered blind spot was called correctly.** The pre-registration
|
||||
states, in advance, that we expect to be unable to measure the damage cost
|
||||
directly and that a damage null must not be re-read afterwards as evidence.
|
||||
That call was right, and it is the reason this run cannot be misread.
|
||||
|
||||
### VERDICT — DO NOT ADOPT
|
||||
|
||||
1. **The pre-registered verdict rule is not met.** Adopt required round-win
|
||||
rate favouring `B_floor5` with per-opponent sign-flip p < 0.05 AND the
|
||||
effect at or above the reported MDE. Observed: **-0.59 pp against**, p =
|
||||
**0.7676**, under the 4.73 pp MDE. Round wins and wins/run are the same
|
||||
null (p = 0.7676, MDE 0.1420 wins/run).
|
||||
2. **`TR_RAM_FLOOR_ENERGY` stays `0.0`.** Do not adopt, do not ship, and **do
|
||||
not re-test this knob.** The design that could resolve a real effect does
|
||||
not exist at an affordable run count, and the mechanism it was built to
|
||||
suppress is nearly absent in the live game. Re-running buys resolution on an
|
||||
effect that is not there.
|
||||
3. **A null here does NOT prove the knob inert — the opposite.** The mechanism
|
||||
fired on 0.04% of ticks. This run licenses only: "no effect >= 4.73 pp of
|
||||
round-win rate at 450 runs/arm". It says nothing about the ~0.04% of ticks
|
||||
it did suppress, because too few of them existed to measure.
|
||||
4. **The generalisable finding is the corpus, not the knob.** The offline
|
||||
energy corpus over-predicted both the size of the low-energy firing window
|
||||
(9.6% -> 0.04%) and the rate of dying broke (61.2% -> 40.6%). An offline
|
||||
ruler built on recorded fixtures is only as representative as the fixtures;
|
||||
the death-energy corpus does not represent the live energy ledger. Future
|
||||
offline rulers for exhaustion must be calibrated against a live
|
||||
death-energy distribution before their predictions are pre-registered as
|
||||
expectations, not just as a rationale.
|
||||
|
||||
**Ship state: unchanged. `TR_RAM_FLOOR_ENERGY=0.0` and `TR_RAM_ENEMY_ENERGY=0.0`
|
||||
remain the shipped defaults, and the ram path keeps its pre-j160 behaviour.**
|
||||
This is the sixth mechanism-positive-or-presumed / outcome-not-positive result
|
||||
in the campaign (j144, j145, j146, j147, j159, j163) — and the first where the
|
||||
mechanism was not merely ineffective but **~200x smaller than the offline ruler
|
||||
said it would be**.
|
||||
|
||||
Reference in New Issue
Block a user