450 runs/arm x 2 arms (15 opponents x 30 runs x 3 rounds, conc=6, 0 failed, 0 never started, 0 env mis-set), frozen binary d9a39c3b8472, TR_MOVEMENT=tfil. Round-win 40.30% -> 39.70% (-0.59 pp, CI -3.90..+2.72, sign-flip exact-2^15 p=0.7676, MDE 4.73 pp); wins/run 0.200 -> 0.182 (MDE 0.1420); damage/run 113.65 -> 112.70 (p=0.5298, MDE 4.19). Deviation disclosed: 30 runs/opponent instead of 42 (throughput 14-22 runs/min vs 22.7-23.2 assumed); full 15-opponent panel and both arms kept, runs reduced. Headline is the mechanism failure: firing was suppressed on 0.04% of ticks, not the 9.6% the offline ruler predicted (~200x smaller); only 4.8% of shots happen in the low-energy zone and the floor removed 14% of those; rounds ending at self energy <=0 were 40.6% live vs 61.2% implied by the offline corpus; median self energy at death 14.1 -> 15.5. The pre-registered 'we cannot measure the damage cost directly' call was correct. DO NOT ADOPT. TR_RAM_FLOOR_ENERGY stays 0.0, do not re-test. The null does not prove the knob inert - the mechanism barely fired.
19 KiB
j162 — the FIRING FLOOR: exhaustion measurement + a re-sized A/B proposal
No battle, server, GUI or A/B was started for this job. Both knobs remain
default-0.0; out/ModularBot was not rebuilt. Everything below the divider is
a proposal awaiting the owner's explicit permission.
MEASURED (common_libs/tests/measure_ram_exhaustion, offline, state only)
Same corpus as j160: 8149 closed-loop recordings / 35163 rounds / 34.46M
ticks. No counterfactual replay (the offline harness scored 0/6 on
closed-loop questions, docs/offline_harness_trust.md).
1. We die BROKE, and it is a death event — not a state we sit disabled in
- 21865 rounds (61.2%) end with self energy crossing 0; 99.0% of rounds end in some death. Energy on the last tick we were alive: median 0.83, mean 2.50, p90 8.90, max 24.83. 55.9% of self-deaths at <=1, 84.2% at <=5, 92.8% at <=10, 100% at <=20.
- Time at energy <= 0 before the round ends: median 1 tick (p90 18). The round ends on the crossing tick. The 1.72%-of-ticks figure is dominated by a handful of recordings that hold a dead bot for hundreds of ticks — a recorder artefact, not a lived state. There is no recoverable disabled window to defend.
- The reserve that would have absorbed the killing blow (the overshoot of the final hit): median 0.40, p75 2.00, p90 6.90, p99 15.0. A free 5-energy reserve would have saved 85.1% of self-deaths.
So the prompt's hypothesis ("it dies at 40 energy, so the floor protects against nothing") is false: the bot dies with nothing, every time. The floor's premise is real.
2. We cannot climb back out of a low-energy dip
Energy rises on 0.598% of tick-pairs (~5.87 landed hits/round; mean landed power 1.42, mode 1.0 — not 0.1). Per tick spent at a given level:
| energy | climb next tick | killed this tick | ratio |
|---|---|---|---|
| <= 3 | 0.130% | 0.830% | dying 6.4x more likely |
| <= 5 | 0.232% | 0.663% | dying 2.9x more likely |
| <= 10 | 0.365% | 0.437% | dying 1.2x more likely |
| <= 20 | 0.500% | 0.256% | recovering 2x more likely |
Below ~10 energy a landed hit is not coming; below 20 it usually is. A floor at 20 would therefore block the only zone where recovery is actually plausible.
3. What the floor costs
| floor | % ticks blocked | rounds | med run | p90 run | mean run | % runs ending in death | bank @1.0p |
|---|---|---|---|---|---|---|---|
| 3 | 7.64% | 23086 | 34 | 299 | 98 | 80.8% | 8.2 |
| 5 | 9.58% | 23739 | 53 | 330 | 118 | 78.0% | 9.9 |
| 10 | 14.52% | 25211 | 89 | 433 | 162 | 70.3% | 13.5 |
| 20 | 24.74% | 27799 | 139 | 621 | 236 | 60.0% | 19.6 |
"Bank" = mean suppressed run x p/(10+2p) energy/tick, the gun-heat ceiling
(heat = 1 + p/5, cool 0.1/tick). At 0.1 power it is 0.52 energy for floor 5;
at 2.0 power, 16.9.
Measured caveat, and it matters: the recorded energy ledger closes exactly
— start + landed-gains - damage - end = -0.00 over 35065 rounds. These
captures do not charge the firepower cost, so the cost column is derived from
the game rules, not read off the data. The landed-hit power distribution is
read off the data (mode 1.0, mean 1.42) and is what sets the bracket.
4. Honest read — materially DIFFERENT from the geometry arm
Firing is net energy-negative for this bot on this panel: a landed hit
returns 3p for p spent (break-even hit rate 1/3), and the hit rate cannot
exceed 5.87 hits / 76 shots-per-round heat ceiling = 7.7%. So not firing
really does bank energy — about 9.9 at floor 5.
That is the same kind of trade the geometry arm made — spend offence, buy protection — but a different magnitude:
- Safety claim is stronger. The geometry arm's safety gain did not convert into wins. Here the hazard is measured directly: 0.66%/tick death at energy <= 5, against a p75 overshoot of 2.0 and a bank of 9.9. The bank is above p75 and near p90 — a genuinely material reserve, not a rounding error.
- Damage cost is ~an order of magnitude smaller. Floor 5 suppresses ~4.4 shots per median run; at 4 damage/hit and a 7.7% hit rate that is ~1.4 damage per suppressed run, ~2 damage/run. The geometry arm lost 8.83 damage/run for its unconverted gain.
- It is a light touch in time, not in behaviour: 9.6% of ticks, median run 53 ticks. Not a blackout.
Verdict: the floor is worth an A/B. It is not the clean negative. But the honest counterweight is on the record: 78% of suppressed runs still end in death, and the bank is only reached because we stopped shooting.
Chosen value: TR_RAM_FLOOR_ENERGY=5 — bank 9.9 (above p75 overshoot 2.0,
near p90 6.9) at 7.6%-vs-9.6% less tick cost than 10. Floor 20 is dropped on
the measurement: it costs 24.7% of ticks and sits on top of the <= 20 recovery
window.
PROPOSAL — PRE-REGISTERED, NOT RUN
- Harness:
tools/ab/tournament_run.sh+tools/ab/tournament_analyze.py, unmodified. Panel:tools/ab/panel_movement.txt(frozen 15-opponent movement panel). Unit of evidence is the opponent, not the battle. One frozen binary fromgit archiveofj160-ramfloor. - Contamination control — j159's three guarantees, unchanged: (1) per-arm
TR_ENV_FILEin this job's own outdir (/tmp/j162_floor/env/<arm>.env); (2) the per-run botdir holds only.json,.shand a symlink to the frozen binary — no.env, and the loader does not walk up; (3) every run's[env]boot report is checked against its arm, and any disagreement voids the session. A session-record check runs BEFORE analysis, not as a rewrite afterwards (j159 had a mid-analysissession.jsonrewrite; not repeated here).
| arm | env | role |
|---|---|---|
A_off |
both unset | REFERENCE |
B_floor |
TR_RAM_FLOOR_ENERGY=5 |
the floor alone (value from the measurement) |
C_exhaust |
TR_RAM_ENEMY_ENERGY=20 |
the exhaustion trigger alone |
D_both |
TR_RAM_FLOOR_ENERGY=5 TR_RAM_ENEMY_ENERGY=20 |
the combined policy |
No arm is inert, so none is dropped: C_exhaust=20 acts on 7.1% of ticks
(enemy <= 20 while we are > 20) and D_both is the only arm that answers the
composition rule. A floor sweep arm is deliberately omitted — the measurement
chose the value, and the budget is better spent on n.
Size — corrected for the real throughput
j159 measured 420 battles in 1110 s = 23 battles/min (not the ~7/min the
previous estimate assumed). MDE scales as 1/sqrt(n); the shipped default
movement gate resolved 0.17 wins/run at 210 runs/arm. For a target MDE of
0.10 wins/run: n = 210 * (0.17/0.10)^2 = 607 runs/arm, rounded up to
42 runs/opponent = 630 runs/arm.
- 15 opponents x 4 arms x 42 runs x 3 rounds = 3780 battles ≈ 2.7 h.
- Decision-only 2-arm version (
A_offvsB_floor): 15 x 2 x 42 x 3 = 1890 battles ≈ 1.4 h, still at MDE 0.10.
Metrics (fixed now)
Primaries: damage/run, round-win rate. Mechanism, never a verdict: self energy at death, ticks spent disabled, shots fired/run, ram-kill count. Paired per-opponent deltas, mean/SD/SE/95% CI, sign test, sign-flip permutation, Wilcoxon cross-check, reported MDE. Two-sided.
Verdict rule (fixed now)
Adopt only if BOTH primaries favour the arm with p(sign-flip) < 0.05 and
the effect is at or above the reported MDE. Otherwise do not ship; both knobs
stay 0.0. A clean null is a fully acceptable result. No subsetting, no
dropping opponents, no re-running to chase a p-value.
j163 — PRE-REGISTRATION: the FIRING FLOOR A/B (2 arms), BEFORE ANY BATTLE
Written and committed before a single battle of this design was run. Nothing
below was chosen after seeing data. Worktree j160-ramfloor @ 64e23e2, one
frozen binary built by tournament_run.sh from git archive HEAD, both knobs
default-0.0, no code changed by this job.
Hypothesis
The bot dies broke: 61.2% of rounds (21865/35753) end with self energy
crossing 0, and energy on the last alive tick is median 0.83. Below 5
energy the next tick brings death 2.9x more often than a landed hit
(0.663%/tick vs 0.232%/tick); the bank of a suppressed run at floor 5 is ~9.9
energy against a p75 overshoot of 2.0 / p90 6.9 — a free 5-energy reserve would
have saved 85.1% of self-deaths. Firing is net energy-negative here (a landed
hit returns 3p for p spent; the hit rate is capped at 5.87/76 = 7.7%), so
holding a reserve in the sub-5 zone should convert safety into round wins.
Arms — two, differing in exactly one variable (TR_MOVEMENT=tfil pinned)
| arm | per-arm env file | role |
|---|---|---|
A_baseline |
TR_RAM_FLOOR_ENERGY=0 |
REFERENCE (today's shipped behaviour) |
B_floor5 |
TR_RAM_FLOOR_ENERGY=5 |
treatment, the value chosen by the j162 measurement |
The TR_RAM_ENEMY_ENERGY (ram-exhaustion) arm is DELIBERATELY EXCLUDED.
Its own measurement found the trigger is rare at its literal threshold and that
whether it fires is close to a coin flip in direction — it is not a
well-founded mechanism. Excluding it here is a decision, not an oversight;
this job tests only the one well-founded mechanism. If the floor is adopted, the
exhaustion trigger needs its own design and its own A/B.
Primaries (fixed now)
- round-win rate (rounds won / rounds fought) — the deciding primary
- damage/run
THE DAMAGE MDE IS STATED UP FRONT, BECAUSE IT IS BIGGER THAN THE EFFECT
Expected damage cost of floor 5: ~2 damage/run (4.4 suppressed shots per median run x 4 damage/hit x 7.7% hit rate) — versus 8.83 damage/run for the already-rejected geometry arm. The design's damage MDE is 7.65. The damage effect is therefore ~3.8x below what this design can resolve.
Recorded before any data: we EXPECT TO BE UNABLE TO MEASURE THE DAMAGE COST DIRECTLY. A null on damage/run is the predicted outcome, not a surprise, and must NOT be re-read after the fact as evidence either for or against the floor. The verdict is judged on ROUND WINS. The damage MDE is a one-sided blind spot of this design, fixed in advance.
Counterweight (also on the record before any data)
78.0% of suppressed runs still end in death. The floor protects the tail of the energy ledger; it is not a shield. A mechanism-positive / outcome-null result is the fifth such in this campaign (j144, j145, j146, j147, j159).
Mechanism metrics (reported, never a verdict)
- self energy at death (per round);
- share of rounds ending at self energy <= 0 (baseline 61.2%);
- shots/run;
- share of ticks with firing suppressed (floor 5 predicts ~9.6% of ticks, median suppressed run ~53 ticks).
- Reported per opponent as well as pooled: j159's re-analysis showed a pooled test hid a real per-opponent effect (safety signal p=0.0008 per-opponent, null pooled). The unit of evidence is the opponent.
Size, MDE and the time floor
15 frozen opponents x 2 arms x 42 runs x 3 rounds = 1890 battles.
MDE ~0.10 wins/run (210 x (0.17/0.10)^2; 210 was the design that resolved
0.17 wins/run). At the measured 22.7-23.2 runs/min that is ~1.4 h — a floor
on elapsed time, not an estimate: opponent heterogeneity does not average
down with added runs. If time runs short, the achieved n and the MDE actually
reached are reported exactly; the panel and the arms are not silently
shrunk.
Contamination controls (all three, in order)
- Each arm is launched with
TR_ENV_FILEpointing at a per-arm file this job generated in its own directory (/tmp/j163_env/<arm>.env) — never a shell export, because the dotenv loader gives the FILE priority. The file dir is deliberately outside--outdir(tournament_run.shrm -rfs the outdir). This matters: the owner has an 18 KB.envatModularBot_garage/out/.envin the main tree. The per-run botdir holds only.json,.shand a symlink to the frozen binary; the frozen binary's own directory holds no.env; the loader's fallbacks are./.envthen exe-adjacent with no parent walk, so the owner's file is unreachable. - Every run's
[env]boot block is verified against its arm as runs complete —TR_RAM_FLOOR_ENERGYandTR_MOVEMENT=tfil— andmis-setis counted and reported. A previous session was invalidated-risk because this was checked too late. - The session record is read before analysis, never rewritten. j159 had a
mid-analysis
session.jsonrewrite; declaringTR_MOVEMENTexplicitly disables the analyzer's leaked-TR_MOVEMENTfatal check, so the built-in leak guard is not trustworthy here — the explicit per-run[env]verification above is the primary control. If the guard misbehaves it is reported as a finding, not worked around.
Verdict rule (fixed now, two-sided)
Adopt only if round-win rate favours B_floor5 with a per-opponent
sign-flip permutation p < 0.05 AND the effect is at or above the reported
MDE. Otherwise do not ship; TR_RAM_FLOOR_ENERGY stays 0.0. No
subsetting, no dropping opponents, no re-running to chase a p-value, no
reinterpreting the bar after seeing the data. A clean null is a fully
acceptable result — and a null here licenses only "no effect >= MDE is
detectable at this design", never "the knob is harmless".
MEASURED
(appended after the battles — everything above was committed first, at
6cfb169)
MEASURED — the live A/B, 450 runs/arm, 2700 rounds (j163)
- Provenance. Frozen 15-opponent movement panel
(
tools/ab/panel_movement.txt), 15 x 2 x 30 runs x 3 rounds = 450 runs/arm, 2700 rounds,conc=6, 0 failed, 0 never started. Frozen binaryd9a39c3b8472…,TR_MOVEMENT=tfilin both arms; the only difference isTR_RAM_FLOOR_ENERGY0(A_floor0, reference) vs5(B_floor5). - Env verification — 0 mis-set. Every run's
[env]boot block was checked against its arm as the session progressed, live at 24 / 193 / 410 / 826 / 900 runs completed: no disagreement at any checkpoint. Per-armTR_ENV_FILEin the session's own directory, botdir without a.env, loader does not walk up parents. This is contamination control #2 from the pre-registration, satisfied. - DEVIATION FROM THE PRE-REGISTRATION, DISCLOSED. The design asked for 42 runs/opponent (1890 battles, MDE ~0.10 wins/run). Measured throughput was 14-22 runs/min, not the assumed 22.7-23.2, so the wall-clock cost of the pre-registered n was not affordable. As the pre-registration required, the panel and the arms were NOT shrunk: the full 15 opponents and both arms were kept and the runs per opponent were reduced to 30. The MDE actually reached is reported below and is the honest resolution limit of this run. No opponent was dropped, no arm was re-run to chase a p-value.
Pooled dashboard (descriptive, NOT the verdict):
| arm | runs | dmg/run | wins/run | round wins | round win rate |
|---|---|---|---|---|---|
A_floor0 |
450 | 113.65 | 0.200 | — | 40.30% |
B_floor5 |
450 | 112.70 | 0.182 | — | 39.70% |
Verdict layer (per-opponent paired deltas, arm − reference; the sign-flip is the exact 2^15 permutation the pre-registration names as the decision test):
| metric | mean Δ | 95% CI | p(sign-flip) | sign test | Wilcoxon p | MDE reached |
|---|---|---|---|---|---|---|
| round-win rate | -0.59 pp | [-3.90, +2.72] | 0.7676 | 1.00 | 0.84 | 4.73 pp |
| wins/run | -0.0178 | [-0.117, +0.082] | 0.7676 | 1.00 | 0.84 | 0.1420 |
| damage/run | -0.95 | [-3.88, +1.98] | 0.5298 | 1.00 | 0.84 | 4.19 |
HEADLINE FINDING — the mechanism barely fired; the offline energy corpus did not survive contact with the live game
- Suppression was 0.04% of ticks, not the 9.6% the offline ruler predicted —
a ~200x smaller effect. The pre-registration's own mechanism metric
(
share of ticks with firing suppressed, predicted ~9.6%, median suppressed run ~53 ticks) is the number that failed, and it failed by two orders of magnitude. - Only 4.8% of shots are ever taken in the low-energy zone, and the floor removed 14% of those. The gate therefore touches a small slice of a small slice: a bot that almost never wants to fire at low energy. The pre- registration's expected damage cost (~2 damage/run) was the arithmetic consequence of the 9.6% figure; with 0.04% it is ~200x smaller still, which is why the damage MDE (4.19) is unreachable by construction and not by bad luck.
- Rounds ending at self energy <= 0: 40.6% live vs 61.2% implied by the offline corpus (the j162 baseline the pre-registration quoted). The recorded-fixture corpus over-states how often we die broke by ~1.5x. This is the second, independent way the same corpus mis-called the live game.
- Median self energy at death 14.1 -> 15.5 — the floor moved the death energy by +1.4, real but tiny, and nowhere near the "we sit disabled" state the floor was built for.
- The pre-registered blind spot was called correctly. The pre-registration states, in advance, that we expect to be unable to measure the damage cost directly and that a damage null must not be re-read afterwards as evidence. That call was right, and it is the reason this run cannot be misread.
VERDICT — DO NOT ADOPT
- The pre-registered verdict rule is not met. Adopt required round-win
rate favouring
B_floor5with per-opponent sign-flip p < 0.05 AND the effect at or above the reported MDE. Observed: -0.59 pp against, p = 0.7676, under the 4.73 pp MDE. Round wins and wins/run are the same null (p = 0.7676, MDE 0.1420 wins/run). TR_RAM_FLOOR_ENERGYstays0.0. Do not adopt, do not ship, and do not re-test this knob. The design that could resolve a real effect does not exist at an affordable run count, and the mechanism it was built to suppress is nearly absent in the live game. Re-running buys resolution on an effect that is not there.- A null here does NOT prove the knob inert — the opposite. The mechanism fired on 0.04% of ticks. This run licenses only: "no effect >= 4.73 pp of round-win rate at 450 runs/arm". It says nothing about the ~0.04% of ticks it did suppress, because too few of them existed to measure.
- The generalisable finding is the corpus, not the knob. The offline energy corpus over-predicted both the size of the low-energy firing window (9.6% -> 0.04%) and the rate of dying broke (61.2% -> 40.6%). An offline ruler built on recorded fixtures is only as representative as the fixtures; the death-energy corpus does not represent the live energy ledger. Future offline rulers for exhaustion must be calibrated against a live death-energy distribution before their predictions are pre-registered as expectations, not just as a rationale.
Ship state: unchanged. TR_RAM_FLOOR_ENERGY=0.0 and TR_RAM_ENEMY_ENERGY=0.0
remain the shipped defaults, and the ram path keeps its pre-j160 behaviour.
This is the sixth mechanism-positive-or-presumed / outcome-not-positive result
in the campaign (j144, j145, j146, j147, j159, j163) — and the first where the
mechanism was not merely ineffective but ~200x smaller than the offline ruler
said it would be.