64e23e29d1
Decisive measurement for the j160 firing floor, on the recorded closed-loop corpus (8149 recordings / 35163 rounds / 34.46M ticks, state only): * 61.2% of rounds end with self energy crossing 0. Energy at death: median 0.83, p90 8.90, max 24.83 -- the bot dies BROKE, so the floor's premise is real. Time at energy<=0 is a median of 1 tick: the round ends on the crossing tick, there is no recoverable disabled window. * Reserve that would have absorbed the killing blow: median 0.40, p75 2.00, p90 6.90. * Cannot climb back out: at energy<=5 the next tick brings a landed hit 0.232% of the time and death 0.663% (2.9x). At <=20 recovery is 2x more likely, which is why a floor at 20 is the wrong value. * Cost: floor 5 blocks 9.58% of ticks, median run 53, mean run 118, banking ~9.9 energy against a p75 overshoot of 2.0. * Measured caveat: the recorded ledger closes exactly (residual -0.00 over 35065 rounds), so these captures do NOT charge firepower cost; the cost column is derived from the game rules, and the landed-hit power (mode 1.0, mean 1.42) is what sets the bracket. Honest read: worth an A/B, materially different in magnitude from the geometry arm (smaller damage cost, stronger and directly measured safety claim), so NOT the clean 'protects against nothing' negative. docs/ram_floor_exhaustion_ab.md: replaces the draft with the measurement plus a re-sized A/B proposal (4 arms, 42 runs/opponent/arm, 630 runs/arm for MDE 0.10 wins/run, ~2.7 h at the measured 23 battles/min). NOT RUN -- no battle, server or GUI was started. Both knobs remain default 0.0.
152 lines
7.4 KiB
Markdown
152 lines
7.4 KiB
Markdown
# j162 — the FIRING FLOOR: exhaustion measurement + a re-sized A/B proposal
|
|
|
|
**No battle, server, GUI or A/B was started for this job.** Both knobs remain
|
|
default-`0.0`; `out/ModularBot` was not rebuilt. Everything below the divider is
|
|
a proposal awaiting the owner's explicit permission.
|
|
|
|
---
|
|
|
|
## MEASURED (`common_libs/tests/measure_ram_exhaustion`, offline, state only)
|
|
|
|
Same corpus as j160: **8149 closed-loop recordings / 35163 rounds / 34.46M
|
|
ticks**. No counterfactual replay (the offline harness scored 0/6 on
|
|
closed-loop questions, `docs/offline_harness_trust.md`).
|
|
|
|
### 1. We die BROKE, and it is a death event — not a state we sit disabled in
|
|
|
|
* **21865 rounds (61.2%) end with self energy crossing 0**; 99.0% of rounds end
|
|
in *some* death. Energy on the last tick we were alive:
|
|
**median 0.83, mean 2.50, p90 8.90, max 24.83**.
|
|
55.9% of self-deaths at <=1, 84.2% at <=5, 92.8% at <=10, **100% at <=20**.
|
|
* Time at energy <= 0 before the round ends: **median 1 tick** (p90 18). The
|
|
round ends on the crossing tick. The 1.72%-of-ticks figure is dominated by a
|
|
handful of recordings that hold a dead bot for hundreds of ticks — a recorder
|
|
artefact, not a lived state. **There is no recoverable disabled window to
|
|
defend.**
|
|
* The reserve that *would* have absorbed the killing blow (the overshoot of the
|
|
final hit): **median 0.40, p75 2.00, p90 6.90, p99 15.0**. A free 5-energy
|
|
reserve would have saved 85.1% of self-deaths.
|
|
|
|
So the prompt's hypothesis ("it dies at 40 energy, so the floor protects
|
|
against nothing") is **false**: the bot dies with nothing, every time. The
|
|
floor's premise is real.
|
|
|
|
### 2. We cannot climb back out of a low-energy dip
|
|
|
|
Energy rises on 0.598% of tick-pairs (~5.87 landed hits/round; **mean landed
|
|
power 1.42**, mode 1.0 — not 0.1). Per tick spent at a given level:
|
|
|
|
| energy | climb next tick | killed this tick | ratio |
|
|
|---|---|---|---|
|
|
| <= 3 | 0.130% | 0.830% | dying **6.4x** more likely |
|
|
| <= 5 | 0.232% | 0.663% | dying **2.9x** more likely |
|
|
| <= 10 | 0.365% | 0.437% | dying 1.2x more likely |
|
|
| <= 20 | 0.500% | 0.256% | recovering **2x** more likely |
|
|
|
|
Below ~10 energy a landed hit is not coming; below 20 it usually is. **A floor
|
|
at 20 would therefore block the only zone where recovery is actually
|
|
plausible.**
|
|
|
|
### 3. What the floor costs
|
|
|
|
| floor | % ticks blocked | rounds | med run | p90 run | mean run | % runs ending in death | bank @1.0p |
|
|
|---|---|---|---|---|---|---|---|
|
|
| 3 | 7.64% | 23086 | 34 | 299 | 98 | 80.8% | 8.2 |
|
|
| 5 | **9.58%** | 23739 | 53 | 330 | 118 | 78.0% | **9.9** |
|
|
| 10 | 14.52% | 25211 | 89 | 433 | 162 | 70.3% | 13.5 |
|
|
| 20 | 24.74% | 27799 | 139 | 621 | 236 | 60.0% | 19.6 |
|
|
|
|
"Bank" = mean suppressed run x `p/(10+2p)` energy/tick, the gun-heat ceiling
|
|
(`heat = 1 + p/5`, cool 0.1/tick). At 0.1 power it is 0.52 energy for floor 5;
|
|
at 2.0 power, 16.9.
|
|
|
|
**Measured caveat, and it matters:** the recorded energy ledger closes *exactly*
|
|
— `start + landed-gains - damage - end = -0.00` over 35065 rounds. **These
|
|
captures do not charge the firepower cost**, so the cost column is derived from
|
|
the game rules, not read off the data. The landed-hit *power* distribution is
|
|
read off the data (mode 1.0, mean 1.42) and is what sets the bracket.
|
|
|
|
### 4. Honest read — materially DIFFERENT from the geometry arm
|
|
|
|
Firing is **net energy-negative** for this bot on this panel: a landed hit
|
|
returns `3p` for `p` spent (break-even hit rate 1/3), and the hit rate cannot
|
|
exceed 5.87 hits / 76 shots-per-round heat ceiling = **7.7%**. So not firing
|
|
really does bank energy — about 9.9 at floor 5.
|
|
|
|
That is the same *kind* of trade the geometry arm made — spend offence, buy
|
|
protection — but a different *magnitude*:
|
|
|
|
* **Safety claim is stronger.** The geometry arm's safety gain did not convert
|
|
into wins. Here the hazard is measured directly: 0.66%/tick death at energy
|
|
<= 5, against a p75 overshoot of 2.0 and a bank of 9.9. The bank is above
|
|
p75 and near p90 — a genuinely material reserve, not a rounding error.
|
|
* **Damage cost is ~an order of magnitude smaller.** Floor 5 suppresses ~4.4
|
|
shots per median run; at 4 damage/hit and a 7.7% hit rate that is ~1.4
|
|
damage per suppressed run, ~2 damage/run. The geometry arm lost **8.83
|
|
damage/run** for its unconverted gain.
|
|
* **It is a light touch in time, not in behaviour**: 9.6% of ticks, median run
|
|
53 ticks. Not a blackout.
|
|
|
|
**Verdict: the floor is worth an A/B. It is not the clean negative.** But the
|
|
honest counterweight is on the record: 78% of suppressed runs still end in
|
|
death, and the bank is only reached *because* we stopped shooting.
|
|
|
|
**Chosen value: `TR_RAM_FLOOR_ENERGY=5`** — bank 9.9 (above p75 overshoot 2.0,
|
|
near p90 6.9) at 7.6%-vs-9.6% less tick cost than 10. Floor 20 is dropped on
|
|
the measurement: it costs 24.7% of ticks and sits on top of the <= 20 recovery
|
|
window.
|
|
|
|
---
|
|
|
|
## PROPOSAL — PRE-REGISTERED, **NOT RUN**
|
|
|
|
* Harness: `tools/ab/tournament_run.sh` + `tools/ab/tournament_analyze.py`,
|
|
unmodified. Panel: `tools/ab/panel_movement.txt` (frozen 15-opponent movement
|
|
panel). Unit of evidence is the opponent, not the battle. One frozen binary
|
|
from `git archive` of `j160-ramfloor`.
|
|
* **Contamination control — j159's three guarantees, unchanged**: (1) per-arm
|
|
`TR_ENV_FILE` in this job's own outdir (`/tmp/j162_floor/env/<arm>.env`);
|
|
(2) the per-run botdir holds only `.json`, `.sh` and a symlink to the frozen
|
|
binary — no `.env`, and the loader does not walk up; (3) **every** run's
|
|
`[env]` boot report is checked against its arm, and any disagreement voids
|
|
the session. **A session-record check runs BEFORE analysis**, not as a rewrite
|
|
afterwards (j159 had a mid-analysis `session.json` rewrite; not repeated here).
|
|
|
|
| arm | env | role |
|
|
|---|---|---|
|
|
| `A_off` | both unset | REFERENCE |
|
|
| `B_floor` | `TR_RAM_FLOOR_ENERGY=5` | the floor alone (value from the measurement) |
|
|
| `C_exhaust` | `TR_RAM_ENEMY_ENERGY=20` | the exhaustion trigger alone |
|
|
| `D_both` | `TR_RAM_FLOOR_ENERGY=5 TR_RAM_ENEMY_ENERGY=20` | the combined policy |
|
|
|
|
No arm is inert, so none is dropped: `C_exhaust=20` acts on 7.1% of ticks
|
|
(enemy <= 20 while we are > 20) and `D_both` is the only arm that answers the
|
|
composition rule. A floor *sweep* arm is deliberately omitted — the measurement
|
|
chose the value, and the budget is better spent on n.
|
|
|
|
### Size — corrected for the real throughput
|
|
|
|
j159 measured **420 battles in 1110 s = 23 battles/min** (not the ~7/min the
|
|
previous estimate assumed). MDE scales as `1/sqrt(n)`; the shipped default
|
|
movement gate resolved **0.17 wins/run at 210 runs/arm**. For a target MDE of
|
|
**0.10 wins/run**: `n = 210 * (0.17/0.10)^2 = 607` runs/arm, rounded up to
|
|
**42 runs/opponent = 630 runs/arm**.
|
|
|
|
* 15 opponents x 4 arms x 42 runs x 3 rounds = **3780 battles ≈ 2.7 h**.
|
|
* Decision-only 2-arm version (`A_off` vs `B_floor`): 15 x 2 x 42 x 3 =
|
|
**1890 battles ≈ 1.4 h**, still at MDE 0.10.
|
|
|
|
### Metrics (fixed now)
|
|
|
|
**Primaries: damage/run, round-win rate.** Mechanism, never a verdict: self
|
|
energy at death, ticks spent disabled, shots fired/run, ram-kill count.
|
|
Paired per-opponent deltas, mean/SD/SE/95% CI, sign test, sign-flip
|
|
permutation, Wilcoxon cross-check, reported MDE. Two-sided.
|
|
|
|
### Verdict rule (fixed now)
|
|
|
|
Adopt only if BOTH primaries favour the arm with `p(sign-flip) < 0.05` **and**
|
|
the effect is at or above the reported MDE. Otherwise do not ship; both knobs
|
|
stay `0.0`. A clean null is a fully acceptable result. No subsetting, no
|
|
dropping opponents, no re-running to chase a p-value.
|