13 KiB
j162 — the FIRING FLOOR: exhaustion measurement + a re-sized A/B proposal
No battle, server, GUI or A/B was started for this job. Both knobs remain
default-0.0; out/ModularBot was not rebuilt. Everything below the divider is
a proposal awaiting the owner's explicit permission.
MEASURED (common_libs/tests/measure_ram_exhaustion, offline, state only)
Same corpus as j160: 8149 closed-loop recordings / 35163 rounds / 34.46M
ticks. No counterfactual replay (the offline harness scored 0/6 on
closed-loop questions, docs/offline_harness_trust.md).
1. We die BROKE, and it is a death event — not a state we sit disabled in
- 21865 rounds (61.2%) end with self energy crossing 0; 99.0% of rounds end in some death. Energy on the last tick we were alive: median 0.83, mean 2.50, p90 8.90, max 24.83. 55.9% of self-deaths at <=1, 84.2% at <=5, 92.8% at <=10, 100% at <=20.
- Time at energy <= 0 before the round ends: median 1 tick (p90 18). The round ends on the crossing tick. The 1.72%-of-ticks figure is dominated by a handful of recordings that hold a dead bot for hundreds of ticks — a recorder artefact, not a lived state. There is no recoverable disabled window to defend.
- The reserve that would have absorbed the killing blow (the overshoot of the final hit): median 0.40, p75 2.00, p90 6.90, p99 15.0. A free 5-energy reserve would have saved 85.1% of self-deaths.
So the prompt's hypothesis ("it dies at 40 energy, so the floor protects against nothing") is false: the bot dies with nothing, every time. The floor's premise is real.
2. We cannot climb back out of a low-energy dip
Energy rises on 0.598% of tick-pairs (~5.87 landed hits/round; mean landed power 1.42, mode 1.0 — not 0.1). Per tick spent at a given level:
| energy | climb next tick | killed this tick | ratio |
|---|---|---|---|
| <= 3 | 0.130% | 0.830% | dying 6.4x more likely |
| <= 5 | 0.232% | 0.663% | dying 2.9x more likely |
| <= 10 | 0.365% | 0.437% | dying 1.2x more likely |
| <= 20 | 0.500% | 0.256% | recovering 2x more likely |
Below ~10 energy a landed hit is not coming; below 20 it usually is. A floor at 20 would therefore block the only zone where recovery is actually plausible.
3. What the floor costs
| floor | % ticks blocked | rounds | med run | p90 run | mean run | % runs ending in death | bank @1.0p |
|---|---|---|---|---|---|---|---|
| 3 | 7.64% | 23086 | 34 | 299 | 98 | 80.8% | 8.2 |
| 5 | 9.58% | 23739 | 53 | 330 | 118 | 78.0% | 9.9 |
| 10 | 14.52% | 25211 | 89 | 433 | 162 | 70.3% | 13.5 |
| 20 | 24.74% | 27799 | 139 | 621 | 236 | 60.0% | 19.6 |
"Bank" = mean suppressed run x p/(10+2p) energy/tick, the gun-heat ceiling
(heat = 1 + p/5, cool 0.1/tick). At 0.1 power it is 0.52 energy for floor 5;
at 2.0 power, 16.9.
Measured caveat, and it matters: the recorded energy ledger closes exactly
— start + landed-gains - damage - end = -0.00 over 35065 rounds. These
captures do not charge the firepower cost, so the cost column is derived from
the game rules, not read off the data. The landed-hit power distribution is
read off the data (mode 1.0, mean 1.42) and is what sets the bracket.
4. Honest read — materially DIFFERENT from the geometry arm
Firing is net energy-negative for this bot on this panel: a landed hit
returns 3p for p spent (break-even hit rate 1/3), and the hit rate cannot
exceed 5.87 hits / 76 shots-per-round heat ceiling = 7.7%. So not firing
really does bank energy — about 9.9 at floor 5.
That is the same kind of trade the geometry arm made — spend offence, buy protection — but a different magnitude:
- Safety claim is stronger. The geometry arm's safety gain did not convert into wins. Here the hazard is measured directly: 0.66%/tick death at energy <= 5, against a p75 overshoot of 2.0 and a bank of 9.9. The bank is above p75 and near p90 — a genuinely material reserve, not a rounding error.
- Damage cost is ~an order of magnitude smaller. Floor 5 suppresses ~4.4 shots per median run; at 4 damage/hit and a 7.7% hit rate that is ~1.4 damage per suppressed run, ~2 damage/run. The geometry arm lost 8.83 damage/run for its unconverted gain.
- It is a light touch in time, not in behaviour: 9.6% of ticks, median run 53 ticks. Not a blackout.
Verdict: the floor is worth an A/B. It is not the clean negative. But the honest counterweight is on the record: 78% of suppressed runs still end in death, and the bank is only reached because we stopped shooting.
Chosen value: TR_RAM_FLOOR_ENERGY=5 — bank 9.9 (above p75 overshoot 2.0,
near p90 6.9) at 7.6%-vs-9.6% less tick cost than 10. Floor 20 is dropped on
the measurement: it costs 24.7% of ticks and sits on top of the <= 20 recovery
window.
PROPOSAL — PRE-REGISTERED, NOT RUN
- Harness:
tools/ab/tournament_run.sh+tools/ab/tournament_analyze.py, unmodified. Panel:tools/ab/panel_movement.txt(frozen 15-opponent movement panel). Unit of evidence is the opponent, not the battle. One frozen binary fromgit archiveofj160-ramfloor. - Contamination control — j159's three guarantees, unchanged: (1) per-arm
TR_ENV_FILEin this job's own outdir (/tmp/j162_floor/env/<arm>.env); (2) the per-run botdir holds only.json,.shand a symlink to the frozen binary — no.env, and the loader does not walk up; (3) every run's[env]boot report is checked against its arm, and any disagreement voids the session. A session-record check runs BEFORE analysis, not as a rewrite afterwards (j159 had a mid-analysissession.jsonrewrite; not repeated here).
| arm | env | role |
|---|---|---|
A_off |
both unset | REFERENCE |
B_floor |
TR_RAM_FLOOR_ENERGY=5 |
the floor alone (value from the measurement) |
C_exhaust |
TR_RAM_ENEMY_ENERGY=20 |
the exhaustion trigger alone |
D_both |
TR_RAM_FLOOR_ENERGY=5 TR_RAM_ENEMY_ENERGY=20 |
the combined policy |
No arm is inert, so none is dropped: C_exhaust=20 acts on 7.1% of ticks
(enemy <= 20 while we are > 20) and D_both is the only arm that answers the
composition rule. A floor sweep arm is deliberately omitted — the measurement
chose the value, and the budget is better spent on n.
Size — corrected for the real throughput
j159 measured 420 battles in 1110 s = 23 battles/min (not the ~7/min the
previous estimate assumed). MDE scales as 1/sqrt(n); the shipped default
movement gate resolved 0.17 wins/run at 210 runs/arm. For a target MDE of
0.10 wins/run: n = 210 * (0.17/0.10)^2 = 607 runs/arm, rounded up to
42 runs/opponent = 630 runs/arm.
- 15 opponents x 4 arms x 42 runs x 3 rounds = 3780 battles ≈ 2.7 h.
- Decision-only 2-arm version (
A_offvsB_floor): 15 x 2 x 42 x 3 = 1890 battles ≈ 1.4 h, still at MDE 0.10.
Metrics (fixed now)
Primaries: damage/run, round-win rate. Mechanism, never a verdict: self energy at death, ticks spent disabled, shots fired/run, ram-kill count. Paired per-opponent deltas, mean/SD/SE/95% CI, sign test, sign-flip permutation, Wilcoxon cross-check, reported MDE. Two-sided.
Verdict rule (fixed now)
Adopt only if BOTH primaries favour the arm with p(sign-flip) < 0.05 and
the effect is at or above the reported MDE. Otherwise do not ship; both knobs
stay 0.0. A clean null is a fully acceptable result. No subsetting, no
dropping opponents, no re-running to chase a p-value.
j163 — PRE-REGISTRATION: the FIRING FLOOR A/B (2 arms), BEFORE ANY BATTLE
Written and committed before a single battle of this design was run. Nothing
below was chosen after seeing data. Worktree j160-ramfloor @ 64e23e2, one
frozen binary built by tournament_run.sh from git archive HEAD, both knobs
default-0.0, no code changed by this job.
Hypothesis
The bot dies broke: 61.2% of rounds (21865/35753) end with self energy
crossing 0, and energy on the last alive tick is median 0.83. Below 5
energy the next tick brings death 2.9x more often than a landed hit
(0.663%/tick vs 0.232%/tick); the bank of a suppressed run at floor 5 is ~9.9
energy against a p75 overshoot of 2.0 / p90 6.9 — a free 5-energy reserve would
have saved 85.1% of self-deaths. Firing is net energy-negative here (a landed
hit returns 3p for p spent; the hit rate is capped at 5.87/76 = 7.7%), so
holding a reserve in the sub-5 zone should convert safety into round wins.
Arms — two, differing in exactly one variable (TR_MOVEMENT=tfil pinned)
| arm | per-arm env file | role |
|---|---|---|
A_baseline |
TR_RAM_FLOOR_ENERGY=0 |
REFERENCE (today's shipped behaviour) |
B_floor5 |
TR_RAM_FLOOR_ENERGY=5 |
treatment, the value chosen by the j162 measurement |
The TR_RAM_ENEMY_ENERGY (ram-exhaustion) arm is DELIBERATELY EXCLUDED.
Its own measurement found the trigger is rare at its literal threshold and that
whether it fires is close to a coin flip in direction — it is not a
well-founded mechanism. Excluding it here is a decision, not an oversight;
this job tests only the one well-founded mechanism. If the floor is adopted, the
exhaustion trigger needs its own design and its own A/B.
Primaries (fixed now)
- round-win rate (rounds won / rounds fought) — the deciding primary
- damage/run
THE DAMAGE MDE IS STATED UP FRONT, BECAUSE IT IS BIGGER THAN THE EFFECT
Expected damage cost of floor 5: ~2 damage/run (4.4 suppressed shots per median run x 4 damage/hit x 7.7% hit rate) — versus 8.83 damage/run for the already-rejected geometry arm. The design's damage MDE is 7.65. The damage effect is therefore ~3.8x below what this design can resolve.
Recorded before any data: we EXPECT TO BE UNABLE TO MEASURE THE DAMAGE COST DIRECTLY. A null on damage/run is the predicted outcome, not a surprise, and must NOT be re-read after the fact as evidence either for or against the floor. The verdict is judged on ROUND WINS. The damage MDE is a one-sided blind spot of this design, fixed in advance.
Counterweight (also on the record before any data)
78.0% of suppressed runs still end in death. The floor protects the tail of the energy ledger; it is not a shield. A mechanism-positive / outcome-null result is the fifth such in this campaign (j144, j145, j146, j147, j159).
Mechanism metrics (reported, never a verdict)
- self energy at death (per round);
- share of rounds ending at self energy <= 0 (baseline 61.2%);
- shots/run;
- share of ticks with firing suppressed (floor 5 predicts ~9.6% of ticks, median suppressed run ~53 ticks).
- Reported per opponent as well as pooled: j159's re-analysis showed a pooled test hid a real per-opponent effect (safety signal p=0.0008 per-opponent, null pooled). The unit of evidence is the opponent.
Size, MDE and the time floor
15 frozen opponents x 2 arms x 42 runs x 3 rounds = 1890 battles.
MDE ~0.10 wins/run (210 x (0.17/0.10)^2; 210 was the design that resolved
0.17 wins/run). At the measured 22.7-23.2 runs/min that is ~1.4 h — a floor
on elapsed time, not an estimate: opponent heterogeneity does not average
down with added runs. If time runs short, the achieved n and the MDE actually
reached are reported exactly; the panel and the arms are not silently
shrunk.
Contamination controls (all three, in order)
- Each arm is launched with
TR_ENV_FILEpointing at a per-arm file this job generated in its own directory (/tmp/j163_env/<arm>.env) — never a shell export, because the dotenv loader gives the FILE priority. The file dir is deliberately outside--outdir(tournament_run.shrm -rfs the outdir). This matters: the owner has an 18 KB.envatModularBot_garage/out/.envin the main tree. The per-run botdir holds only.json,.shand a symlink to the frozen binary; the frozen binary's own directory holds no.env; the loader's fallbacks are./.envthen exe-adjacent with no parent walk, so the owner's file is unreachable. - Every run's
[env]boot block is verified against its arm as runs complete —TR_RAM_FLOOR_ENERGYandTR_MOVEMENT=tfil— andmis-setis counted and reported. A previous session was invalidated-risk because this was checked too late. - The session record is read before analysis, never rewritten. j159 had a
mid-analysis
session.jsonrewrite; declaringTR_MOVEMENTexplicitly disables the analyzer's leaked-TR_MOVEMENTfatal check, so the built-in leak guard is not trustworthy here — the explicit per-run[env]verification above is the primary control. If the guard misbehaves it is reported as a finding, not worked around.
Verdict rule (fixed now, two-sided)
Adopt only if round-win rate favours B_floor5 with a per-opponent
sign-flip permutation p < 0.05 AND the effect is at or above the reported
MDE. Otherwise do not ship; TR_RAM_FLOOR_ENERGY stays 0.0. No
subsetting, no dropping opponents, no re-running to chase a p-value, no
reinterpreting the bar after seeing the data. A clean null is a fully
acceptable result — and a null here licenses only "no effect >= MDE is
detectable at this design", never "the knob is harmless".