Files
SirRoboGarage/docs/ram_floor_exhaustion_ab.md
T
SirStone 23bce2dad5 j160 (default-off): the energy-reserve FIRING FLOOR + ENEMY-EXHAUSTION ram trigger
TR_RAM_FLOOR_ENERGY (0.0 = off): at/below this self energy we start no NEW
shot, holding back the reserve for a final ram exchange. Justified by the only
energy gain in the game being +3*power per bullet hit LANDED, so not firing
denies the enemy its only refill. Blocks only NEW shots (gunHeat already gates
committed ones) and is bypassed while ramming.

TR_RAM_ENEMY_ENERGY (0.0 = off): last-scanned enemy energy <= this -> ram
mode. Enemy energy IS observable (ScannedBotEvent.energy, schemas.nim:306),
1-8 ticks stale. This is the shipped finisher with its energy tolerance
promoted to a knob, keeping the self>enemy surplus guard because RAM_DAMAGE
0.6 applies to BOTH bots on every contact tick.

Open-loop measurement (measure_ramfloor_energy, 8149 recordings / 29871
rounds / 33.8M ticks): 'both low' is COMMON (10.4% of ticks below 20, 15.1%
below 25) but neither side goes low first (enemy 52.7% / us 47.3%), and the
owner's literal trigger - enemy so low it cannot fire (energy <= 1.95) - is
only 2.5% of ticks, 1.1% while we are healthy.

Guards 136 -> 147 in test_tfil_commit_env.nim, all green. A/B PRE-REGISTERED
in docs/ram_floor_exhaustion_ab.md and NOT RUN.
2026-09-27 12:46:46 +02:00

6.9 KiB
Raw Blame History

j160 — PRE-REGISTRATION: the energy-reserve ram policy (floor + exhaustion)

Written and committed BEFORE any battle of this experiment ran. Nothing below the divider has data in it, and none of it will unless the owner approves.

What was built (both default-OFF, both default = today's behaviour)

knob default effect when set
TR_RAM_FLOOR_ENERGY 0.0 = off at/below this self energy we do not start a NEW shot
TR_RAM_ENEMY_ENERGY 0.0 = off last-scanned enemy energy <= this -> ram mode

Mechanics that justify the design (all read from source, not assumed): energy has no cap and no regeneration; the only gain in the game is +3 * power per bullet hit landed (server/rules/rules.kt). Not firing therefore denies the enemy its only refill and preserves our ram reserve — the floor and the exhaustion trigger are the same bet, not two ideas.

RAM_DAMAGE 0.6 is applied to both bots on every contact tick, so a head-on contact is a symmetric bleed decided by who entered with the surplus. The exhaustion trigger deliberately keeps the shipped finisher's selfEnergy > enemyEnergy guard for that reason.

Composition when they conflict: RAMMING WINS. Once ram mode is engaged the duel is over and the reserve is being spent, not held, so the floor is bypassed. The floor therefore can never deadlock the ram it exists to enable. The floor blocks only NEW shots: a bullet already in the air has gunHeat > 0 and shouldFire already gates on gunHeat <= 0.0, so no committed shot is suppressed or cancelled.

Guards: common_libs/tests/test_tfil_commit_env.nim 136 -> 147 checks, all green. Eleven new checks cover default parity on both knobs (the exhaustion arm is unreachable at 0.0 and the old finisher/desperation verdicts are unchanged over 210 input combinations), the floor firing at exactly the threshold and not one tick above it, the floor never blocking healthy energy in any configuration, the trigger switching to ram exactly at the tolerance, and the composition.

The open-loop measurement (common_libs/tests/measure_ramfloor_energy)

Recorded closed-loop corpus, 8149 recordings / 29871 rounds / 33.8M ticks, state only, no counterfactual replay (the offline harness scored 0/6 on closed-loop questions, docs/offline_harness_trust.md).

The owner's premise ("both low, nobody firing") is TRUE and COMMON — but it is not the situation he thinks it is. Both bots below 25 energy happens on 15.1% of ticks; below 20, 10.4%; below 10, 2.9%. Self energy is below 20 on 23.5% of ticks with a median continuous run of 131 ticks.

Neither side goes low first — it is a coin flip. The enemy crosses 20 first in 52.7% of rounds, we do in 47.3%. The premise that the enemy is the one that runs dry is not supported; half the time we are the exhausted one, and then there is nothing to ram into safely.

The specific trigger the owner described — "so low that it doesn't fire anywhere more" — is rare. The server rejects a shot when energy <= power, so a p=1.95 opponent (DrussGT) stops firing below 2.0 energy: 2.5% of ticks, and only 1.1% while we are healthy. Below 3.0: 3.2%. An exhaustion tolerance sized to the literal premise (TR_RAM_ENEMY_ENERGY ~2) therefore acts on about 1% of ticks.

A looser tolerance is available and is what an A/B would actually sweep: enemy <= 20 while we are above 20 is 7.1% of ticks, enemy <= 30 is 13.2%.

Gun-heat note: a 0.1-power shot adds 1.02 heat and the gun cools 0.1/tick, so a floor suppresses roughly one shot per 11 ticks. This is a lost opportunity rate, not a lost-damage estimate, and no counterfactual damage number is produced here.

Arms

Both: TR_MOVEMENT pinned, one frozen binary built once from git archive of the j160 branch, per-arm TR_ENV_FILE in this job's own outdir, and every run's [env] boot report checked against its arm before any number is read (same contamination control as j159 — the owner's personal .env gives the FILE priority and would silently void arm A).

arm env role
A_off both knobs unset REFERENCE — the shipped behaviour
B_floor TR_RAM_FLOOR_ENERGY=5 the reserve half alone
C_exhaust TR_RAM_ENEMY_ENERGY=20 the exhaustion half alone
D_both TR_RAM_FLOOR_ENERGY=5 TR_RAM_ENEMY_ENERGY=20 the owner's full policy

Four arms, not two, because the two mechanisms are separable: the floor can plausibly help on its own (it is the only part whose premise the data supports) while the exhaustion trigger acts on 1-7% of ticks. A two-arm A+B run could not tell a win from one half.

Design

  • Harness: tools/ab/tournament_run.sh + tools/ab/tournament_analyze.py, unmodified. Panel: tools/ab/panel_movement.txt, the frozen 15-opponent movement panel. Unit of evidence is the opponent, not the battle.
  • 15 opponents x 4 arms x 10 runs x 3 rounds = 600 battles, serialised (--wait-arena 45). ~1.5 h.

Primary metrics (pre-registered, fixed)

  1. damage/run
  2. round-win rate

Mechanism metrics (reported, never a verdict) — specific to this one

  • self energy at death (the floor's entire claim is about the reserve we hold when it matters);
  • ram-kill count and ram-contact count (the exhaustion trigger's channel);
  • shots fired/run (the floor's direct cost).

Statistical treatment

Per-opponent paired deltas (arm − reference), mean, SD, SE, 95% CI, sign test, sign-flip permutation test, Wilcoxon as a cross-check, plus the MDE the analyzer prints. Direction is pre-registered as two-sided: a regression is as interesting as a win, and the offline literature (docs/ramming_negative_ result.md: proactive straight-line ramming converted 0/59) makes a regression entirely plausible.

MDE — large, stated now

At 10 runs/arm the design resolves roughly 0.3 wins/run and the corresponding damage/run figure. Resolving 0.1 wins/run would need ~3x the battles. A null is the likely outcome — this would be the fifth consecutive mechanism-positive / outcome-null result in this campaign.

Consequences recorded before any data:

  • a null excludes only a large effect; it does not show the knobs do nothing;
  • if B_floor is null and C_exhaust is null, the pair still leaves the combined policy untested, and D_both is the arm that would be shipped.

Verdict rule (fixed now, not re-read later)

  • Adopt only if BOTH primaries move in the arm's favour with p(sign-flip) < 0.05 and the effect is at or above the reported MDE.
  • Otherwise do not ship; the knobs stay default off.
  • Mechanism numbers are reported as measured, with no vote in the verdict.
  • No subsetting, no dropping opponents, no re-running to chase a p-value. A clean null is a fully acceptable result.

MEASURED

(appended after the battles — everything above was committed first. Nothing below exists yet: no battle, server, GUI or A/B was started for this job.)