Files
SirRoboGarage/docs/ram_floor_exhaustion_ab.md
SirStone 51bfa57067 j163: RESULT - the firing floor is a clean negative, and the offline energy corpus missed the live game by 200x
450 runs/arm x 2 arms (15 opponents x 30 runs x 3 rounds, conc=6, 0 failed,
0 never started, 0 env mis-set), frozen binary d9a39c3b8472, TR_MOVEMENT=tfil.
Round-win 40.30% -> 39.70% (-0.59 pp, CI -3.90..+2.72, sign-flip exact-2^15
p=0.7676, MDE 4.73 pp); wins/run 0.200 -> 0.182 (MDE 0.1420); damage/run
113.65 -> 112.70 (p=0.5298, MDE 4.19). Deviation disclosed: 30 runs/opponent
instead of 42 (throughput 14-22 runs/min vs 22.7-23.2 assumed); full 15-opponent
panel and both arms kept, runs reduced.

Headline is the mechanism failure: firing was suppressed on 0.04% of ticks,
not the 9.6% the offline ruler predicted (~200x smaller); only 4.8% of shots
happen in the low-energy zone and the floor removed 14% of those; rounds ending
at self energy <=0 were 40.6% live vs 61.2% implied by the offline corpus;
median self energy at death 14.1 -> 15.5. The pre-registered 'we cannot measure
the damage cost directly' call was correct.

DO NOT ADOPT. TR_RAM_FLOOR_ENERGY stays 0.0, do not re-test. The null does not
prove the knob inert - the mechanism barely fired.
2026-09-27 13:55:19 +02:00

367 lines
19 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# j162 — the FIRING FLOOR: exhaustion measurement + a re-sized A/B proposal
**No battle, server, GUI or A/B was started for this job.** Both knobs remain
default-`0.0`; `out/ModularBot` was not rebuilt. Everything below the divider is
a proposal awaiting the owner's explicit permission.
---
## MEASURED (`common_libs/tests/measure_ram_exhaustion`, offline, state only)
Same corpus as j160: **8149 closed-loop recordings / 35163 rounds / 34.46M
ticks**. No counterfactual replay (the offline harness scored 0/6 on
closed-loop questions, `docs/offline_harness_trust.md`).
### 1. We die BROKE, and it is a death event — not a state we sit disabled in
* **21865 rounds (61.2%) end with self energy crossing 0**; 99.0% of rounds end
in *some* death. Energy on the last tick we were alive:
**median 0.83, mean 2.50, p90 8.90, max 24.83**.
55.9% of self-deaths at <=1, 84.2% at <=5, 92.8% at <=10, **100% at <=20**.
* Time at energy <= 0 before the round ends: **median 1 tick** (p90 18). The
round ends on the crossing tick. The 1.72%-of-ticks figure is dominated by a
handful of recordings that hold a dead bot for hundreds of ticks — a recorder
artefact, not a lived state. **There is no recoverable disabled window to
defend.**
* The reserve that *would* have absorbed the killing blow (the overshoot of the
final hit): **median 0.40, p75 2.00, p90 6.90, p99 15.0**. A free 5-energy
reserve would have saved 85.1% of self-deaths.
So the prompt's hypothesis ("it dies at 40 energy, so the floor protects
against nothing") is **false**: the bot dies with nothing, every time. The
floor's premise is real.
### 2. We cannot climb back out of a low-energy dip
Energy rises on 0.598% of tick-pairs (~5.87 landed hits/round; **mean landed
power 1.42**, mode 1.0 — not 0.1). Per tick spent at a given level:
| energy | climb next tick | killed this tick | ratio |
|---|---|---|---|
| <= 3 | 0.130% | 0.830% | dying **6.4x** more likely |
| <= 5 | 0.232% | 0.663% | dying **2.9x** more likely |
| <= 10 | 0.365% | 0.437% | dying 1.2x more likely |
| <= 20 | 0.500% | 0.256% | recovering **2x** more likely |
Below ~10 energy a landed hit is not coming; below 20 it usually is. **A floor
at 20 would therefore block the only zone where recovery is actually
plausible.**
### 3. What the floor costs
| floor | % ticks blocked | rounds | med run | p90 run | mean run | % runs ending in death | bank @1.0p |
|---|---|---|---|---|---|---|---|
| 3 | 7.64% | 23086 | 34 | 299 | 98 | 80.8% | 8.2 |
| 5 | **9.58%** | 23739 | 53 | 330 | 118 | 78.0% | **9.9** |
| 10 | 14.52% | 25211 | 89 | 433 | 162 | 70.3% | 13.5 |
| 20 | 24.74% | 27799 | 139 | 621 | 236 | 60.0% | 19.6 |
"Bank" = mean suppressed run x `p/(10+2p)` energy/tick, the gun-heat ceiling
(`heat = 1 + p/5`, cool 0.1/tick). At 0.1 power it is 0.52 energy for floor 5;
at 2.0 power, 16.9.
**Measured caveat, and it matters:** the recorded energy ledger closes *exactly*
— `start + landed-gains - damage - end = -0.00` over 35065 rounds. **These
captures do not charge the firepower cost**, so the cost column is derived from
the game rules, not read off the data. The landed-hit *power* distribution is
read off the data (mode 1.0, mean 1.42) and is what sets the bracket.
### 4. Honest read — materially DIFFERENT from the geometry arm
Firing is **net energy-negative** for this bot on this panel: a landed hit
returns `3p` for `p` spent (break-even hit rate 1/3), and the hit rate cannot
exceed 5.87 hits / 76 shots-per-round heat ceiling = **7.7%**. So not firing
really does bank energy — about 9.9 at floor 5.
That is the same *kind* of trade the geometry arm made — spend offence, buy
protection — but a different *magnitude*:
* **Safety claim is stronger.** The geometry arm's safety gain did not convert
into wins. Here the hazard is measured directly: 0.66%/tick death at energy
<= 5, against a p75 overshoot of 2.0 and a bank of 9.9. The bank is above
p75 and near p90 — a genuinely material reserve, not a rounding error.
* **Damage cost is ~an order of magnitude smaller.** Floor 5 suppresses ~4.4
shots per median run; at 4 damage/hit and a 7.7% hit rate that is ~1.4
damage per suppressed run, ~2 damage/run. The geometry arm lost **8.83
damage/run** for its unconverted gain.
* **It is a light touch in time, not in behaviour**: 9.6% of ticks, median run
53 ticks. Not a blackout.
**Verdict: the floor is worth an A/B. It is not the clean negative.** But the
honest counterweight is on the record: 78% of suppressed runs still end in
death, and the bank is only reached *because* we stopped shooting.
**Chosen value: `TR_RAM_FLOOR_ENERGY=5`** — bank 9.9 (above p75 overshoot 2.0,
near p90 6.9) at 7.6%-vs-9.6% less tick cost than 10. Floor 20 is dropped on
the measurement: it costs 24.7% of ticks and sits on top of the <= 20 recovery
window.
---
## PROPOSAL — PRE-REGISTERED, **NOT RUN**
* Harness: `tools/ab/tournament_run.sh` + `tools/ab/tournament_analyze.py`,
unmodified. Panel: `tools/ab/panel_movement.txt` (frozen 15-opponent movement
panel). Unit of evidence is the opponent, not the battle. One frozen binary
from `git archive` of `j160-ramfloor`.
* **Contamination control — j159's three guarantees, unchanged**: (1) per-arm
`TR_ENV_FILE` in this job's own outdir (`/tmp/j162_floor/env/<arm>.env`);
(2) the per-run botdir holds only `.json`, `.sh` and a symlink to the frozen
binary — no `.env`, and the loader does not walk up; (3) **every** run's
`[env]` boot report is checked against its arm, and any disagreement voids
the session. **A session-record check runs BEFORE analysis**, not as a rewrite
afterwards (j159 had a mid-analysis `session.json` rewrite; not repeated here).
| arm | env | role |
|---|---|---|
| `A_off` | both unset | REFERENCE |
| `B_floor` | `TR_RAM_FLOOR_ENERGY=5` | the floor alone (value from the measurement) |
| `C_exhaust` | `TR_RAM_ENEMY_ENERGY=20` | the exhaustion trigger alone |
| `D_both` | `TR_RAM_FLOOR_ENERGY=5 TR_RAM_ENEMY_ENERGY=20` | the combined policy |
No arm is inert, so none is dropped: `C_exhaust=20` acts on 7.1% of ticks
(enemy <= 20 while we are > 20) and `D_both` is the only arm that answers the
composition rule. A floor *sweep* arm is deliberately omitted — the measurement
chose the value, and the budget is better spent on n.
### Size — corrected for the real throughput
j159 measured **420 battles in 1110 s = 23 battles/min** (not the ~7/min the
previous estimate assumed). MDE scales as `1/sqrt(n)`; the shipped default
movement gate resolved **0.17 wins/run at 210 runs/arm**. For a target MDE of
**0.10 wins/run**: `n = 210 * (0.17/0.10)^2 = 607` runs/arm, rounded up to
**42 runs/opponent = 630 runs/arm**.
* 15 opponents x 4 arms x 42 runs x 3 rounds = **3780 battles ≈ 2.7 h**.
* Decision-only 2-arm version (`A_off` vs `B_floor`): 15 x 2 x 42 x 3 =
**1890 battles ≈ 1.4 h**, still at MDE 0.10.
### Metrics (fixed now)
**Primaries: damage/run, round-win rate.** Mechanism, never a verdict: self
energy at death, ticks spent disabled, shots fired/run, ram-kill count.
Paired per-opponent deltas, mean/SD/SE/95% CI, sign test, sign-flip
permutation, Wilcoxon cross-check, reported MDE. Two-sided.
### Verdict rule (fixed now)
Adopt only if BOTH primaries favour the arm with `p(sign-flip) < 0.05` **and**
the effect is at or above the reported MDE. Otherwise do not ship; both knobs
stay `0.0`. A clean null is a fully acceptable result. No subsetting, no
dropping opponents, no re-running to chase a p-value.
---
# j163 — PRE-REGISTRATION: the FIRING FLOOR A/B (2 arms), BEFORE ANY BATTLE
**Written and committed before a single battle of this design was run.** Nothing
below was chosen after seeing data. Worktree `j160-ramfloor` @ `64e23e2`, one
frozen binary built by `tournament_run.sh` from `git archive HEAD`, both knobs
default-`0.0`, no code changed by this job.
## Hypothesis
The bot dies broke: **61.2% of rounds (21865/35753) end with self energy
crossing 0**, and energy on the last alive tick is median 0.83. Below **5**
energy the next tick brings death **2.9x** more often than a landed hit
(0.663%/tick vs 0.232%/tick); the bank of a suppressed run at floor 5 is ~9.9
energy against a p75 overshoot of 2.0 / p90 6.9 — a free 5-energy reserve would
have saved 85.1% of self-deaths. Firing is net energy-negative here (a landed
hit returns `3p` for `p` spent; the hit rate is capped at 5.87/76 = 7.7%), so
holding a reserve in the sub-5 zone should convert safety into **round wins**.
## Arms — two, differing in exactly one variable (`TR_MOVEMENT=tfil` pinned)
| arm | per-arm env file | role |
|---|---|---|
| `A_baseline` | `TR_RAM_FLOOR_ENERGY=0` | REFERENCE (today's shipped behaviour) |
| `B_floor5` | `TR_RAM_FLOOR_ENERGY=5` | treatment, the value chosen by the j162 measurement |
**The `TR_RAM_ENEMY_ENERGY` (ram-exhaustion) arm is DELIBERATELY EXCLUDED.**
Its own measurement found the trigger is rare at its literal threshold and that
whether it fires is close to a coin flip in direction — it is not a
well-founded mechanism. Excluding it here is a decision, **not an oversight**;
this job tests only the one well-founded mechanism. If the floor is adopted, the
exhaustion trigger needs its own design and its own A/B.
## Primaries (fixed now)
1. **round-win rate** (rounds won / rounds fought) — **the deciding primary**
2. **damage/run**
## THE DAMAGE MDE IS STATED UP FRONT, BECAUSE IT IS BIGGER THAN THE EFFECT
Expected damage cost of floor 5: **~2 damage/run** (4.4 suppressed shots per
median run x 4 damage/hit x 7.7% hit rate) — versus **8.83 damage/run** for the
already-rejected geometry arm. The design's **damage MDE is 7.65**. The damage
effect is therefore **~3.8x below what this design can resolve**.
> **Recorded before any data: we EXPECT TO BE UNABLE TO MEASURE THE DAMAGE
> COST DIRECTLY. A null on damage/run is the predicted outcome, not a surprise,
> and must NOT be re-read after the fact as evidence either for or against the
> floor.** The verdict is judged on **ROUND WINS**. The damage MDE is a
> one-sided blind spot of this design, fixed in advance.
## Counterweight (also on the record before any data)
**78.0% of suppressed runs still end in death.** The floor protects the tail of
the energy ledger; it is not a shield. A mechanism-positive / outcome-null
result is the fifth such in this campaign (j144, j145, j146, j147, j159).
## Mechanism metrics (reported, never a verdict)
* self energy at death (per round);
* share of rounds ending at self energy **<= 0** (baseline **61.2%**);
* shots/run;
* share of ticks with firing suppressed (**floor 5 predicts ~9.6% of ticks,
median suppressed run ~53 ticks**).
* Reported **per opponent as well as pooled**: j159's re-analysis showed a
pooled test hid a real per-opponent effect (safety signal p=0.0008
per-opponent, null pooled). The unit of evidence is the opponent.
## Size, MDE and the time floor
15 frozen opponents x 2 arms x **42 runs** x 3 rounds = **1890 battles**.
MDE ~**0.10 wins/run** (`210 x (0.17/0.10)^2`; 210 was the design that resolved
0.17 wins/run). At the measured 22.7-23.2 runs/min that is **~1.4 h — a floor
on elapsed time, not an estimate**: opponent heterogeneity does not average
down with added runs. If time runs short, the achieved n and the MDE actually
reached are reported exactly; the panel and the arms are **not** silently
shrunk.
## Contamination controls (all three, in order)
1. Each arm is launched with `TR_ENV_FILE` pointing at a **per-arm file this
job generated** in its own directory (`/tmp/j163_env/<arm>.env`) — never a
shell export, because the dotenv loader gives the FILE priority. The file
dir is deliberately **outside** `--outdir` (`tournament_run.sh` `rm -rf`s
the outdir). This matters: the owner has an 18 KB `.env` at
`ModularBot_garage/out/.env` in the main tree. The per-run botdir holds only
`.json`, `.sh` and a symlink to the frozen binary; the frozen binary's own
directory holds no `.env`; the loader's fallbacks are `./.env` then
exe-adjacent with **no parent walk**, so the owner's file is unreachable.
2. **Every run's `[env]` boot block is verified against its arm as runs
complete** — `TR_RAM_FLOOR_ENERGY` and `TR_MOVEMENT=tfil` — and `mis-set`
is counted and reported. A previous session was invalidated-risk because
this was checked too late.
3. The session record is **read before analysis, never rewritten**. j159 had a
mid-analysis `session.json` rewrite; declaring `TR_MOVEMENT` explicitly
disables the analyzer's leaked-`TR_MOVEMENT` fatal check, so the built-in
leak guard is **not trustworthy here** — the explicit per-run `[env]`
verification above is the primary control. If the guard misbehaves it is
**reported as a finding, not worked around**.
## Verdict rule (fixed now, two-sided)
**Adopt** only if round-win rate favours `B_floor5` with a per-opponent
**sign-flip permutation p < 0.05** AND the effect is at or above the reported
MDE. Otherwise **do not ship**; `TR_RAM_FLOOR_ENERGY` stays `0.0`. No
subsetting, no dropping opponents, no re-running to chase a p-value, no
reinterpreting the bar after seeing the data. **A clean null is a fully
acceptable result** — and a null here licenses only "no effect >= MDE is
detectable at this design", never "the knob is harmless".
---
## MEASURED
*(appended after the battles — everything above was committed first, at
`6cfb169`)*
### MEASURED — the live A/B, 450 runs/arm, 2700 rounds (j163)
* **Provenance.** Frozen 15-opponent movement panel
(`tools/ab/panel_movement.txt`), 15 x 2 x **30 runs** x 3 rounds =
**450 runs/arm, 2700 rounds**, `conc=6`, **0 failed, 0 never started**.
Frozen binary `d9a39c3b8472…`, `TR_MOVEMENT=tfil` in both arms; the only
difference is `TR_RAM_FLOOR_ENERGY` `0` (`A_floor0`, reference) vs `5`
(`B_floor5`).
* **Env verification — 0 mis-set.** Every run's `[env]` boot block was checked
against its arm as the session progressed, live at **24 / 193 / 410 / 826 /
900** runs completed: no disagreement at any checkpoint. Per-arm
`TR_ENV_FILE` in the session's own directory, botdir without a `.env`, loader
does not walk up parents. This is contamination control #2 from the
pre-registration, satisfied.
* **DEVIATION FROM THE PRE-REGISTRATION, DISCLOSED.** The design asked for
**42 runs/opponent** (1890 battles, MDE ~0.10 wins/run). Measured throughput
was **14-22 runs/min**, not the assumed 22.7-23.2, so the wall-clock cost of
the pre-registered n was not affordable. As the pre-registration required,
the **panel and the arms were NOT shrunk**: the full 15 opponents and both
arms were kept and the **runs per opponent were reduced to 30**. The MDE
actually reached is reported below and is the honest resolution limit of this
run. No opponent was dropped, no arm was re-run to chase a p-value.
**Pooled dashboard (descriptive, NOT the verdict):**
| arm | runs | dmg/run | wins/run | round wins | round win rate |
|---|---:|---:|---:|---:|---:|
| `A_floor0` | 450 | 113.65 | 0.200 | — | **40.30%** |
| `B_floor5` | 450 | 112.70 | 0.182 | — | **39.70%** |
**Verdict layer** (per-opponent paired deltas, arm − reference; the sign-flip
is the exact 2^15 permutation the pre-registration names as the decision test):
| metric | mean Δ | 95% CI | p(sign-flip) | sign test | Wilcoxon p | **MDE reached** |
|---|---:|---|---:|---:|---:|---:|
| **round-win rate** | **-0.59 pp** | [-3.90, +2.72] | **0.7676** | 1.00 | 0.84 | **4.73 pp** |
| wins/run | -0.0178 | [-0.117, +0.082] | 0.7676 | 1.00 | 0.84 | 0.1420 |
| damage/run | -0.95 | [-3.88, +1.98] | 0.5298 | 1.00 | 0.84 | 4.19 |
### HEADLINE FINDING — the mechanism barely fired; the offline energy corpus did not survive contact with the live game
1. **Suppression was 0.04% of ticks, not the 9.6% the offline ruler predicted** —
a **~200x** smaller effect. The pre-registration's own mechanism metric
(`share of ticks with firing suppressed`, predicted ~9.6%, median suppressed
run ~53 ticks) is the number that failed, and it failed by two orders of
magnitude.
2. **Only 4.8% of shots are ever taken in the low-energy zone, and the floor
removed 14% of those.** The gate therefore touches a small slice of a small
slice: a bot that almost never wants to fire at low energy. The pre-
registration's expected damage cost (~2 damage/run) was the arithmetic
consequence of the 9.6% figure; with 0.04% it is ~200x smaller still, which
is why the damage MDE (4.19) is unreachable by construction and not by bad
luck.
3. **Rounds ending at self energy <= 0: 40.6% live vs 61.2% implied by the
offline corpus** (the j162 baseline the pre-registration quoted). The
recorded-fixture corpus over-states how often we die broke by ~1.5x. This is
the second, independent way the same corpus mis-called the live game.
4. **Median self energy at death 14.1 -> 15.5** — the floor moved the death
energy by +1.4, real but tiny, and nowhere near the "we sit disabled" state
the floor was built for.
5. **The pre-registered blind spot was called correctly.** The pre-registration
states, in advance, that we expect to be unable to measure the damage cost
directly and that a damage null must not be re-read afterwards as evidence.
That call was right, and it is the reason this run cannot be misread.
### VERDICT — DO NOT ADOPT
1. **The pre-registered verdict rule is not met.** Adopt required round-win
rate favouring `B_floor5` with per-opponent sign-flip p < 0.05 AND the
effect at or above the reported MDE. Observed: **-0.59 pp against**, p =
**0.7676**, under the 4.73 pp MDE. Round wins and wins/run are the same
null (p = 0.7676, MDE 0.1420 wins/run).
2. **`TR_RAM_FLOOR_ENERGY` stays `0.0`.** Do not adopt, do not ship, and **do
not re-test this knob.** The design that could resolve a real effect does
not exist at an affordable run count, and the mechanism it was built to
suppress is nearly absent in the live game. Re-running buys resolution on an
effect that is not there.
3. **A null here does NOT prove the knob inert — the opposite.** The mechanism
fired on 0.04% of ticks. This run licenses only: "no effect >= 4.73 pp of
round-win rate at 450 runs/arm". It says nothing about the ~0.04% of ticks
it did suppress, because too few of them existed to measure.
4. **The generalisable finding is the corpus, not the knob.** The offline
energy corpus over-predicted both the size of the low-energy firing window
(9.6% -> 0.04%) and the rate of dying broke (61.2% -> 40.6%). An offline
ruler built on recorded fixtures is only as representative as the fixtures;
the death-energy corpus does not represent the live energy ledger. Future
offline rulers for exhaustion must be calibrated against a live
death-energy distribution before their predictions are pre-registered as
expectations, not just as a rationale.
**Ship state: unchanged. `TR_RAM_FLOOR_ENERGY=0.0` and `TR_RAM_ENEMY_ENERGY=0.0`
remain the shipped defaults, and the ram path keeps its pre-j160 behaviour.**
This is the sixth mechanism-positive-or-presumed / outcome-not-positive result
in the campaign (j144, j145, j146, j147, j159, j163) — and the first where the
mechanism was not merely ineffective but **~200x smaller than the offline ruler
said it would be**.