From 4f332dcb95914b9d066e081f1c63d972ea6ec7a8 Mon Sep 17 00:00:00 2001 From: Davide Cappellini Date: Sat, 26 Sep 2026 10:55:14 +0200 Subject: [PATCH] learned movement (SBC): live panel results and the direct answer (module is a wash vs strafe on wins; state conditioning is a CI-separated hit-rate gain vs the same mover without it; the label is the wrong quantity) --- docs/movement_campaign.md | 137 ++++++++++++++++++++++++++++++++++++++ 1 file changed, 137 insertions(+) diff --git a/docs/movement_campaign.md b/docs/movement_campaign.md index 86c5cf6..bc1c639 100644 --- a/docs/movement_campaign.md +++ b/docs/movement_campaign.md @@ -1864,3 +1864,140 @@ against `strafe`'s 9.40%. What the mover **should** learn is the outcome: a counted/decayed SBC over states and bins labelled by HIT/MISS would estimate P(hit | state, bin) directly. That is the natural next experiment and it is NOT what was measured here.)* + +--- + +## Learned movement (SBC) — RESULTS (appended AFTER the battles) + +**Session `/tmp/ab/j128_learned`, frozen from the pre-registration commit +`a436e9f` (binary sha256 `60f2093b58b8…`): 15 opponents × 4 arms × 3 runs × +3 rounds = 180 battles, conc 6, 0 excluded, 0 failed starts.** Reference: +`strafe` (the shipped champion). Every delta is (arm − `strafe`). + +### Pooled dashboard (descriptive, NOT the verdict) + +| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | **incoming hit rate** | mean distance | +|---|---:|---:|---:|---:|---:|---:|---:|---:| +| `strafe` (champion) | 45 | 106.0 | 152.4 | 1.56 | 70/135 | 51.9% | 13.07% | 431 | +| `learned` (decay on) | 45 | 97.3 | 138.7 | 1.76 | 79/135 | 58.5% | **12.26%** | 408 | +| `learned_nodecay` | 45 | 104.7 | 146.3 | 1.71 | 77/135 | 57.0% | 12.73% | 410 | +| `learned_global` (state conditioning OFF) | 45 | 99.4 | 151.7 | 1.78 | 80/135 | 59.3% | 13.76% | 407 | + +### Cross-opponent aggregation (the verdict layer), reference `strafe` + +| arm | metric | mean Δ | spread (SD) | 95% CI | sign test | p(sign) | p(sign-flip) | MDE | +|---|---|---:|---:|---|---:|---:|---:|---:| +| `learned` | wins | +0.20 | 0.73 | **[−0.21, +0.61]** | 6/11 | 1 | 0.3662 | **0.53** | +| `learned` | damage | −8.67 | 30.54 | [−25.58, +8.25] | 7/15 | 1 | 0.2953 | 22.09 | +| `learned` | damage_taken | −13.64 | 41.46 | [−36.60, +9.33] | 7/15 | 1 | 0.2311 | 29.99 | +| `learned` | **hit_rate** | **−1.01 pp** | 4.05 | **[−3.26, +1.23]** | 5/15 | 0.3018 | 0.3437 | **2.93** | +| `learned_nodecay` | wins | +0.16 | 0.69 | [−0.23, +0.54] | 6/9 | 0.5078 | 0.4688 | 0.50 | +| `learned_nodecay` | hit_rate | −1.13 pp | 3.97 | [−3.33, +1.07] | 7/15 | 1 | 0.291 | 2.87 | +| `learned_global` | wins | +0.22 | 0.88 | [−0.26, +0.71] | 7/12 | 0.7744 | 0.394 | 0.64 | +| `learned_global` | hit_rate | **+0.30 pp** | 4.17 | [−2.01, +2.61] | 8/15 | 1 | 0.783 | 3.02 | + +**Pre-registered verdict vs `strafe`: NO ARM BEATS THE CHAMPION.** All three +learned arms are **not distinguishable** from `strafe` on round wins *and* on +damage, by both readings of the pre-registered rule. The win leg (rule 1) fails +for every arm: the sign-flip p-values are 0.37 / 0.47 / 0.39 and every 95% CI +contains 0. The mechanism leg (rule 2) also fails: the incoming hit rate is +−1.01 pp for `learned` (CI [−3.26, +1.23], MDE 2.93 pp) — pointing the right +way, but smaller than this batch can resolve. + +### The arm that DOES separate: state conditioning vs the same mover without it + +`learned` vs `learned_global` is a free pairwise comparison on the same 180 +battles (re-analyze with `--reference learned_global`): identical binary, +identical wave geometry, identical counted SBC and priors — the only difference +is that `learned_global` forces the state to a single cell (the old global +histogram). + +| `learned` − `learned_global` | mean Δ | 95% CI | sign-flip p | MDE | +|---|---:|---|---:|---:| +| **incoming hit rate** | **−1.31 pp** | **[−2.58, −0.05]** | **0.0444** | 1.65 | +| **damage taken/run** | **−12.93** | **[−25.81, −0.05]** | **0.0485** | 16.82 | +| wins/run | −0.02 | [−0.27, +0.22] | 1.0 | 0.32 | +| damage/run | −2.03 | [−12.13, +8.08] | 0.667 | 13.20 | + +**The state conditioning is a REAL, measurable dodging improvement** — the +incoming hit rate drops 1.31 pp with a CI that excludes 0 and sign-flip +p = 0.044, and damage taken drops 12.9/run with a CI that excludes 0 — **but it +does not move round wins at all** (Δwins −0.02). So the learned state +conditioning works as advertised and is simply too small to matter for the +score against this panel. + +### Cost (MEASURED, `-d:release`, 200k ticks, `git archive HEAD` clean build) + +| scenario | mean ms/tick | worst single tick observed | +|---|---:|---:| +| 1v1 (decision every tick a wave is live) | **0.0024** | 1.14 ms | +| 4 enemies | 0.0085 | 0.56 ms | + +Budget is 13.16 ms/tick; the module uses **0.02%** of it. Memory: one +`uint8` per (state × bin) = 16×16×31 = 7 936 B. It is not a cost problem. + +### Direct answer + +**Does state-conditional learned danger beat the hand-tuned `strafe` on dodging +and/or on wins? NO — on neither, by the pre-registered rules.** The point +estimates lean the module's way (wins +0.20/run, hit rate −1.01 pp, damage taken +−13.6/run) but every CI contains 0 and the win-leg MDE (0.53 wins/run) is 2.6× +the observed effect: this batch cannot resolve an effect of the measured size, +and a confirmation would need ~100 opponents or 4× the runs. The honest +statement is **"a wash on wins, a small unresolvable dodging gain"**, not a win. + +**Is the failure in the information or in the learner? MAINLY THE INFORMATION — +and specifically the LABEL.** Three independent measurements say so: + +1. **The observable state carries almost nothing (offline, MEASURED).** On + 63 held-out battles the state-conditional model beats the global histogram + and chance on log-loss, but the absolute skill is 3.93% top-1 (global 3.96%, + chance 3.23%) — ~no information about the wave-crossing bin. The strong + information in `docs/state_window_gate.md` (0.41 accuracy) came from a state + measured relative to the ENEMY'S BULLET LINE, which leaks the enemy's lead; + measured in the frame the mover can actually observe, that signal is gone. +2. **The map the mover minimises is the WRONG quantity (offline, MEASURED).** + `corr( P(arrival bin), P(hit | arrival bin) ) = −0.342` over the 31 bins: the + bins with the least mass (the clamped extremes) are where this corpus's gun + lands the MOST hits (bins 1–2 and 28–29: 14–17.5%; bins 23–25: 6.7–7.2%). + Minimising the resolved-position histogram steers INTO the bullets. No + learner can fix a mislabelled target, and this also explains j115. +3. **The learner itself is fine (live, MEASURED).** Against the identical mover + with the state removed, the state conditioning produces a CI-separated + −1.31 pp hit rate and −12.9 damage taken/run. The counted SBC learns and + extracts a real signal; the signal is just too small to beat `strafe`. + +**Two secondary findings.** (a) `learned` vs `learned_nodecay` is a wash live +(12.26% vs 12.73% hit rate, Δwins +0.05) — the forgetting mechanism is NOT the +binding constraint here, and offline the no-decay arm was even the better +predictor, i.e. this opponent did not adapt to us on the decay's timescale. +(b) `learned_global` (state conditioning OFF) has the BEST pooled wins/run of +the four arms (1.78) while dodging WORSE (13.76%) — a reminder that this panel's +win signal is noisy at 3 runs/arm and that the wave-surfing geometry, not the +learning, is where the movement value lives. + +### MEASURED vs INFERRED + +**MEASURED:** the session identity (commit, sha, 180 battles, 0 excluded); the +pooled dashboard; every cross-opponent mean/CI/sign/p/MDE above; the +`learned` vs `learned_global` and `learned` vs `learned_nodecay` pairwise +numbers (same 180 battles, no new fighting); the offline table, the recurrence +counts, the label-shuffle control and the danger-map alignment in "Gate A"; +the ms/tick cost; the clean-build verification. + +**INFERRED:** (i) that the danger-map misalignment is *the* cause of the +resolved-position surfer's weakness — it is consistent with j115 (13.51% vs +9.40%) and with the near-zero offline skill, but it is not a controlled +intervention; (ii) that the small live hit-rate gain is the same mechanism the +offline gate measured; (iii) that the win leg is unresolvable rather than +absent — the CI is wide on both sides. + +**Recommended follow-up (not done, not scheduled):** label by OUTCOME. A +counted+decayed SBC over (state, candidate bin) labelled HIT/MISS estimates +P(hit | state, bin) directly — the quantity the mover should minimise and the +one the alignment table shows is not the histogram. That is the single change +that the evidence here points at, and it is a different experiment from this +one. + +**Status: the default is UNCHANGED (`TR_MOVEMENT=strafe`); the module is +default-off behind `TR_MOVEMENT=learned`.** Revert = do not set the env var.