j162: measure whether the bot ever runs out of energy (offline, no battles)

Decisive measurement for the j160 firing floor, on the recorded closed-loop
corpus (8149 recordings / 35163 rounds / 34.46M ticks, state only):

* 61.2% of rounds end with self energy crossing 0. Energy at death: median
  0.83, p90 8.90, max 24.83 -- the bot dies BROKE, so the floor's premise is
  real. Time at energy<=0 is a median of 1 tick: the round ends on the
  crossing tick, there is no recoverable disabled window.
* Reserve that would have absorbed the killing blow: median 0.40, p75 2.00,
  p90 6.90.
* Cannot climb back out: at energy<=5 the next tick brings a landed hit 0.232%
  of the time and death 0.663% (2.9x). At <=20 recovery is 2x more likely,
  which is why a floor at 20 is the wrong value.
* Cost: floor 5 blocks 9.58% of ticks, median run 53, mean run 118, banking
  ~9.9 energy against a p75 overshoot of 2.0.
* Measured caveat: the recorded ledger closes exactly (residual -0.00 over
  35065 rounds), so these captures do NOT charge firepower cost; the cost
  column is derived from the game rules, and the landed-hit power (mode 1.0,
  mean 1.42) is what sets the bracket.

Honest read: worth an A/B, materially different in magnitude from the geometry
arm (smaller damage cost, stronger and directly measured safety claim), so NOT
the clean 'protects against nothing' negative.

docs/ram_floor_exhaustion_ab.md: replaces the draft with the measurement plus a
re-sized A/B proposal (4 arms, 42 runs/opponent/arm, 630 runs/arm for MDE 0.10
wins/run, ~2.7 h at the measured 23 battles/min). NOT RUN -- no battle, server
or GUI was started. Both knobs remain default 0.0.
This commit is contained in:
2026-09-27 12:58:14 +02:00
parent 23bce2dad5
commit 64e23e29d1
3 changed files with 481 additions and 139 deletions
+146 -139
View File
@@ -1,144 +1,151 @@
# j160 — PRE-REGISTRATION: the energy-reserve ram policy (floor + exhaustion)
# j162 — the FIRING FLOOR: exhaustion measurement + a re-sized A/B proposal
**Written and committed BEFORE any battle of this experiment ran. Nothing below
the divider has data in it, and none of it will unless the owner approves.**
## What was built (both default-OFF, both default = today's behaviour)
| knob | default | effect when set |
|---|---|---|
| `TR_RAM_FLOOR_ENERGY` | **0.0 = off** | at/below this self energy we do not start a NEW shot |
| `TR_RAM_ENEMY_ENERGY` | **0.0 = off** | last-scanned enemy energy <= this -> ram mode |
Mechanics that justify the design (all read from source, not assumed):
energy has **no cap and no regeneration**; the only gain in the game is
`+3 * power` per bullet hit **landed** (`server/rules/rules.kt`). Not firing
therefore denies the enemy its only refill *and* preserves our ram reserve —
the floor and the exhaustion trigger are the same bet, not two ideas.
`RAM_DAMAGE 0.6` is applied to **both** bots on every contact tick, so a
head-on contact is a symmetric bleed decided by who entered with the surplus.
The exhaustion trigger deliberately **keeps** the shipped finisher's
`selfEnergy > enemyEnergy` guard for that reason.
**Composition when they conflict: RAMMING WINS.** Once ram mode is engaged the
duel is over and the reserve is being *spent*, not held, so the floor is
bypassed. The floor therefore can never deadlock the ram it exists to enable.
The floor blocks only NEW shots: a bullet already in the air has `gunHeat > 0`
and `shouldFire` already gates on `gunHeat <= 0.0`, so no committed shot is
suppressed or cancelled.
Guards: `common_libs/tests/test_tfil_commit_env.nim` **136 -> 147 checks**, all
green. Eleven new checks cover default parity on both knobs (the exhaustion arm
is unreachable at 0.0 and the old finisher/desperation verdicts are unchanged
over 210 input combinations), the floor firing at exactly the threshold and not
one tick above it, the floor never blocking healthy energy in any configuration,
the trigger switching to ram exactly at the tolerance, and the composition.
## The open-loop measurement (`common_libs/tests/measure_ramfloor_energy`)
Recorded closed-loop corpus, **8149 recordings / 29871 rounds / 33.8M ticks**,
state only, no counterfactual replay (the offline harness scored 0/6 on
closed-loop questions, `docs/offline_harness_trust.md`).
**The owner's premise ("both low, nobody firing") is TRUE and COMMON — but it is
not the situation he thinks it is.** Both bots below 25 energy happens on
**15.1% of ticks**; below 20, **10.4%**; below 10, **2.9%**. Self energy is
below 20 on **23.5%** of ticks with a median continuous run of **131 ticks**.
**Neither side goes low first — it is a coin flip.** The enemy crosses 20 first
in 52.7% of rounds, we do in 47.3%. The premise that the *enemy* is the one
that runs dry is not supported; half the time we are the exhausted one, and
then there is nothing to ram *into* safely.
**The specific trigger the owner described — "so low that it doesn't fire
anywhere more" — is rare.** The server rejects a shot when `energy <= power`, so
a p=1.95 opponent (DrussGT) stops firing below **2.0 energy**: **2.5% of ticks**,
and only **1.1%** while we are healthy. Below 3.0: 3.2%. An exhaustion tolerance
sized to the literal premise (`TR_RAM_ENEMY_ENERGY` ~2) therefore acts on about
1% of ticks.
A looser tolerance is available and is what an A/B would actually sweep:
enemy <= 20 while we are above 20 is **7.1% of ticks**, enemy <= 30 is 13.2%.
Gun-heat note: a 0.1-power shot adds 1.02 heat and the gun cools 0.1/tick, so a
floor suppresses roughly one shot per 11 ticks. This is a *lost opportunity
rate*, not a lost-damage estimate, and no counterfactual damage number is
produced here.
## Arms
Both: `TR_MOVEMENT` pinned, one frozen binary built once from `git archive` of
the j160 branch, per-arm `TR_ENV_FILE` in this job's own outdir, and every
run's `[env]` boot report checked against its arm before any number is read
(same contamination control as j159 — the owner's personal `.env` gives the
FILE priority and would silently void arm A).
| arm | env | role |
|---|---|---|
| `A_off` | both knobs unset | **REFERENCE** — the shipped behaviour |
| `B_floor` | `TR_RAM_FLOOR_ENERGY=5` | the reserve half alone |
| `C_exhaust` | `TR_RAM_ENEMY_ENERGY=20` | the exhaustion half alone |
| `D_both` | `TR_RAM_FLOOR_ENERGY=5 TR_RAM_ENEMY_ENERGY=20` | the owner's full policy |
Four arms, not two, because the two mechanisms are separable: the floor can
plausibly help on its own (it is the only part whose premise the data supports)
while the exhaustion trigger acts on 1-7% of ticks. A two-arm A+B run could not
tell a win from one half.
## Design
* Harness: `tools/ab/tournament_run.sh` + `tools/ab/tournament_analyze.py`,
unmodified. Panel: `tools/ab/panel_movement.txt`, the frozen 15-opponent
movement panel. Unit of evidence is the opponent, not the battle.
* 15 opponents x 4 arms x **10 runs** x 3 rounds = **600 battles**, serialised
(`--wait-arena 45`). ~1.5 h.
### Primary metrics (pre-registered, fixed)
1. **damage/run**
2. **round-win rate**
### Mechanism metrics (reported, never a verdict) — specific to this one
* **self energy at death** (the floor's entire claim is about the reserve we
hold when it matters);
* **ram-kill count** and ram-contact count (the exhaustion trigger's channel);
* shots fired/run (the floor's direct cost).
### Statistical treatment
Per-opponent paired deltas (arm − reference), mean, SD, SE, 95% CI, sign test,
sign-flip permutation test, Wilcoxon as a cross-check, plus the MDE the
analyzer prints. **Direction is pre-registered as two-sided**: a regression is
as interesting as a win, and the offline literature (`docs/ramming_negative_
result.md`: proactive straight-line ramming converted 0/59) makes a regression
entirely plausible.
### MDE — large, stated now
At **10 runs/arm** the design resolves roughly **0.3 wins/run** and the
corresponding damage/run figure. Resolving 0.1 wins/run would need ~3x the
battles. **A null is the likely outcome — this would be the fifth consecutive
mechanism-positive / outcome-null result in this campaign.**
Consequences recorded before any data:
* a null **excludes only a large effect**; it does not show the knobs do nothing;
* if `B_floor` is null and `C_exhaust` is null, the *pair* still leaves the
combined policy untested, and `D_both` is the arm that would be shipped.
### Verdict rule (fixed now, not re-read later)
* **Adopt** only if BOTH primaries move in the arm's favour with
`p(sign-flip) < 0.05` and the effect is at or above the reported MDE.
* Otherwise **do not ship**; the knobs stay default `off`.
* Mechanism numbers are reported as measured, with no vote in the verdict.
* No subsetting, no dropping opponents, no re-running to chase a p-value. A
clean null is a fully acceptable result.
**No battle, server, GUI or A/B was started for this job.** Both knobs remain
default-`0.0`; `out/ModularBot` was not rebuilt. Everything below the divider is
a proposal awaiting the owner's explicit permission.
---
## MEASURED
## MEASURED (`common_libs/tests/measure_ram_exhaustion`, offline, state only)
*(appended after the battles — everything above was committed first. Nothing
below exists yet: **no battle, server, GUI or A/B was started for this job**.)*
Same corpus as j160: **8149 closed-loop recordings / 35163 rounds / 34.46M
ticks**. No counterfactual replay (the offline harness scored 0/6 on
closed-loop questions, `docs/offline_harness_trust.md`).
### 1. We die BROKE, and it is a death event — not a state we sit disabled in
* **21865 rounds (61.2%) end with self energy crossing 0**; 99.0% of rounds end
in *some* death. Energy on the last tick we were alive:
**median 0.83, mean 2.50, p90 8.90, max 24.83**.
55.9% of self-deaths at <=1, 84.2% at <=5, 92.8% at <=10, **100% at <=20**.
* Time at energy <= 0 before the round ends: **median 1 tick** (p90 18). The
round ends on the crossing tick. The 1.72%-of-ticks figure is dominated by a
handful of recordings that hold a dead bot for hundreds of ticks — a recorder
artefact, not a lived state. **There is no recoverable disabled window to
defend.**
* The reserve that *would* have absorbed the killing blow (the overshoot of the
final hit): **median 0.40, p75 2.00, p90 6.90, p99 15.0**. A free 5-energy
reserve would have saved 85.1% of self-deaths.
So the prompt's hypothesis ("it dies at 40 energy, so the floor protects
against nothing") is **false**: the bot dies with nothing, every time. The
floor's premise is real.
### 2. We cannot climb back out of a low-energy dip
Energy rises on 0.598% of tick-pairs (~5.87 landed hits/round; **mean landed
power 1.42**, mode 1.0 — not 0.1). Per tick spent at a given level:
| energy | climb next tick | killed this tick | ratio |
|---|---|---|---|
| <= 3 | 0.130% | 0.830% | dying **6.4x** more likely |
| <= 5 | 0.232% | 0.663% | dying **2.9x** more likely |
| <= 10 | 0.365% | 0.437% | dying 1.2x more likely |
| <= 20 | 0.500% | 0.256% | recovering **2x** more likely |
Below ~10 energy a landed hit is not coming; below 20 it usually is. **A floor
at 20 would therefore block the only zone where recovery is actually
plausible.**
### 3. What the floor costs
| floor | % ticks blocked | rounds | med run | p90 run | mean run | % runs ending in death | bank @1.0p |
|---|---|---|---|---|---|---|---|
| 3 | 7.64% | 23086 | 34 | 299 | 98 | 80.8% | 8.2 |
| 5 | **9.58%** | 23739 | 53 | 330 | 118 | 78.0% | **9.9** |
| 10 | 14.52% | 25211 | 89 | 433 | 162 | 70.3% | 13.5 |
| 20 | 24.74% | 27799 | 139 | 621 | 236 | 60.0% | 19.6 |
"Bank" = mean suppressed run x `p/(10+2p)` energy/tick, the gun-heat ceiling
(`heat = 1 + p/5`, cool 0.1/tick). At 0.1 power it is 0.52 energy for floor 5;
at 2.0 power, 16.9.
**Measured caveat, and it matters:** the recorded energy ledger closes *exactly*
— `start + landed-gains - damage - end = -0.00` over 35065 rounds. **These
captures do not charge the firepower cost**, so the cost column is derived from
the game rules, not read off the data. The landed-hit *power* distribution is
read off the data (mode 1.0, mean 1.42) and is what sets the bracket.
### 4. Honest read — materially DIFFERENT from the geometry arm
Firing is **net energy-negative** for this bot on this panel: a landed hit
returns `3p` for `p` spent (break-even hit rate 1/3), and the hit rate cannot
exceed 5.87 hits / 76 shots-per-round heat ceiling = **7.7%**. So not firing
really does bank energy — about 9.9 at floor 5.
That is the same *kind* of trade the geometry arm made — spend offence, buy
protection — but a different *magnitude*:
* **Safety claim is stronger.** The geometry arm's safety gain did not convert
into wins. Here the hazard is measured directly: 0.66%/tick death at energy
<= 5, against a p75 overshoot of 2.0 and a bank of 9.9. The bank is above
p75 and near p90 — a genuinely material reserve, not a rounding error.
* **Damage cost is ~an order of magnitude smaller.** Floor 5 suppresses ~4.4
shots per median run; at 4 damage/hit and a 7.7% hit rate that is ~1.4
damage per suppressed run, ~2 damage/run. The geometry arm lost **8.83
damage/run** for its unconverted gain.
* **It is a light touch in time, not in behaviour**: 9.6% of ticks, median run
53 ticks. Not a blackout.
**Verdict: the floor is worth an A/B. It is not the clean negative.** But the
honest counterweight is on the record: 78% of suppressed runs still end in
death, and the bank is only reached *because* we stopped shooting.
**Chosen value: `TR_RAM_FLOOR_ENERGY=5`** — bank 9.9 (above p75 overshoot 2.0,
near p90 6.9) at 7.6%-vs-9.6% less tick cost than 10. Floor 20 is dropped on
the measurement: it costs 24.7% of ticks and sits on top of the <= 20 recovery
window.
---
## PROPOSAL — PRE-REGISTERED, **NOT RUN**
* Harness: `tools/ab/tournament_run.sh` + `tools/ab/tournament_analyze.py`,
unmodified. Panel: `tools/ab/panel_movement.txt` (frozen 15-opponent movement
panel). Unit of evidence is the opponent, not the battle. One frozen binary
from `git archive` of `j160-ramfloor`.
* **Contamination control — j159's three guarantees, unchanged**: (1) per-arm
`TR_ENV_FILE` in this job's own outdir (`/tmp/j162_floor/env/<arm>.env`);
(2) the per-run botdir holds only `.json`, `.sh` and a symlink to the frozen
binary — no `.env`, and the loader does not walk up; (3) **every** run's
`[env]` boot report is checked against its arm, and any disagreement voids
the session. **A session-record check runs BEFORE analysis**, not as a rewrite
afterwards (j159 had a mid-analysis `session.json` rewrite; not repeated here).
| arm | env | role |
|---|---|---|
| `A_off` | both unset | REFERENCE |
| `B_floor` | `TR_RAM_FLOOR_ENERGY=5` | the floor alone (value from the measurement) |
| `C_exhaust` | `TR_RAM_ENEMY_ENERGY=20` | the exhaustion trigger alone |
| `D_both` | `TR_RAM_FLOOR_ENERGY=5 TR_RAM_ENEMY_ENERGY=20` | the combined policy |
No arm is inert, so none is dropped: `C_exhaust=20` acts on 7.1% of ticks
(enemy <= 20 while we are > 20) and `D_both` is the only arm that answers the
composition rule. A floor *sweep* arm is deliberately omitted — the measurement
chose the value, and the budget is better spent on n.
### Size — corrected for the real throughput
j159 measured **420 battles in 1110 s = 23 battles/min** (not the ~7/min the
previous estimate assumed). MDE scales as `1/sqrt(n)`; the shipped default
movement gate resolved **0.17 wins/run at 210 runs/arm**. For a target MDE of
**0.10 wins/run**: `n = 210 * (0.17/0.10)^2 = 607` runs/arm, rounded up to
**42 runs/opponent = 630 runs/arm**.
* 15 opponents x 4 arms x 42 runs x 3 rounds = **3780 battles ≈ 2.7 h**.
* Decision-only 2-arm version (`A_off` vs `B_floor`): 15 x 2 x 42 x 3 =
**1890 battles ≈ 1.4 h**, still at MDE 0.10.
### Metrics (fixed now)
**Primaries: damage/run, round-win rate.** Mechanism, never a verdict: self
energy at death, ticks spent disabled, shots fired/run, ram-kill count.
Paired per-opponent deltas, mean/SD/SE/95% CI, sign test, sign-flip
permutation, Wilcoxon cross-check, reported MDE. Two-sided.
### Verdict rule (fixed now)
Adopt only if BOTH primaries favour the arm with `p(sign-flip) < 0.05` **and**
the effect is at or above the reported MDE. Otherwise do not ship; both knobs
stay `0.0`. A clean null is a fully acceptable result. No subsetting, no
dropping opponents, no re-running to chase a p-value.