learned movement (SBC): live panel results and the direct answer (module is a wash vs strafe on wins; state conditioning is a CI-separated hit-rate gain vs the same mover without it; the label is the wrong quantity)

This commit is contained in:
2026-09-26 10:55:14 +02:00
parent ee98827a62
commit 4f332dcb95
+137
View File
@@ -1864,3 +1864,140 @@ against `strafe`'s 9.40%. What the mover **should** learn is the outcome: a
counted/decayed SBC over states and bins labelled by HIT/MISS would estimate
P(hit | state, bin) directly. That is the natural next experiment and it is NOT
what was measured here.)*
---
## Learned movement (SBC) — RESULTS (appended AFTER the battles)
**Session `/tmp/ab/j128_learned`, frozen from the pre-registration commit
`a436e9f` (binary sha256 `60f2093b58b8…`): 15 opponents × 4 arms × 3 runs ×
3 rounds = 180 battles, conc 6, 0 excluded, 0 failed starts.** Reference:
`strafe` (the shipped champion). Every delta is (arm − `strafe`).
### Pooled dashboard (descriptive, NOT the verdict)
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | **incoming hit rate** | mean distance |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| `strafe` (champion) | 45 | 106.0 | 152.4 | 1.56 | 70/135 | 51.9% | 13.07% | 431 |
| `learned` (decay on) | 45 | 97.3 | 138.7 | 1.76 | 79/135 | 58.5% | **12.26%** | 408 |
| `learned_nodecay` | 45 | 104.7 | 146.3 | 1.71 | 77/135 | 57.0% | 12.73% | 410 |
| `learned_global` (state conditioning OFF) | 45 | 99.4 | 151.7 | 1.78 | 80/135 | 59.3% | 13.76% | 407 |
### Cross-opponent aggregation (the verdict layer), reference `strafe`
| arm | metric | mean Δ | spread (SD) | 95% CI | sign test | p(sign) | p(sign-flip) | MDE |
|---|---|---:|---:|---|---:|---:|---:|---:|
| `learned` | wins | +0.20 | 0.73 | **[−0.21, +0.61]** | 6/11 | 1 | 0.3662 | **0.53** |
| `learned` | damage | −8.67 | 30.54 | [−25.58, +8.25] | 7/15 | 1 | 0.2953 | 22.09 |
| `learned` | damage_taken | −13.64 | 41.46 | [−36.60, +9.33] | 7/15 | 1 | 0.2311 | 29.99 |
| `learned` | **hit_rate** | **−1.01 pp** | 4.05 | **[−3.26, +1.23]** | 5/15 | 0.3018 | 0.3437 | **2.93** |
| `learned_nodecay` | wins | +0.16 | 0.69 | [−0.23, +0.54] | 6/9 | 0.5078 | 0.4688 | 0.50 |
| `learned_nodecay` | hit_rate | −1.13 pp | 3.97 | [−3.33, +1.07] | 7/15 | 1 | 0.291 | 2.87 |
| `learned_global` | wins | +0.22 | 0.88 | [−0.26, +0.71] | 7/12 | 0.7744 | 0.394 | 0.64 |
| `learned_global` | hit_rate | **+0.30 pp** | 4.17 | [−2.01, +2.61] | 8/15 | 1 | 0.783 | 3.02 |
**Pre-registered verdict vs `strafe`: NO ARM BEATS THE CHAMPION.** All three
learned arms are **not distinguishable** from `strafe` on round wins *and* on
damage, by both readings of the pre-registered rule. The win leg (rule 1) fails
for every arm: the sign-flip p-values are 0.37 / 0.47 / 0.39 and every 95% CI
contains 0. The mechanism leg (rule 2) also fails: the incoming hit rate is
−1.01 pp for `learned` (CI [−3.26, +1.23], MDE 2.93 pp) — pointing the right
way, but smaller than this batch can resolve.
### The arm that DOES separate: state conditioning vs the same mover without it
`learned` vs `learned_global` is a free pairwise comparison on the same 180
battles (re-analyze with `--reference learned_global`): identical binary,
identical wave geometry, identical counted SBC and priors — the only difference
is that `learned_global` forces the state to a single cell (the old global
histogram).
| `learned` − `learned_global` | mean Δ | 95% CI | sign-flip p | MDE |
|---|---:|---|---:|---:|
| **incoming hit rate** | **−1.31 pp** | **[−2.58, −0.05]** | **0.0444** | 1.65 |
| **damage taken/run** | **−12.93** | **[−25.81, −0.05]** | **0.0485** | 16.82 |
| wins/run | −0.02 | [−0.27, +0.22] | 1.0 | 0.32 |
| damage/run | −2.03 | [−12.13, +8.08] | 0.667 | 13.20 |
**The state conditioning is a REAL, measurable dodging improvement** — the
incoming hit rate drops 1.31 pp with a CI that excludes 0 and sign-flip
p = 0.044, and damage taken drops 12.9/run with a CI that excludes 0 — **but it
does not move round wins at all** (Δwins −0.02). So the learned state
conditioning works as advertised and is simply too small to matter for the
score against this panel.
### Cost (MEASURED, `-d:release`, 200k ticks, `git archive HEAD` clean build)
| scenario | mean ms/tick | worst single tick observed |
|---|---:|---:|
| 1v1 (decision every tick a wave is live) | **0.0024** | 1.14 ms |
| 4 enemies | 0.0085 | 0.56 ms |
Budget is 13.16 ms/tick; the module uses **0.02%** of it. Memory: one
`uint8` per (state × bin) = 16×16×31 = 7 936 B. It is not a cost problem.
### Direct answer
**Does state-conditional learned danger beat the hand-tuned `strafe` on dodging
and/or on wins? NO — on neither, by the pre-registered rules.** The point
estimates lean the module's way (wins +0.20/run, hit rate −1.01 pp, damage taken
−13.6/run) but every CI contains 0 and the win-leg MDE (0.53 wins/run) is 2.6×
the observed effect: this batch cannot resolve an effect of the measured size,
and a confirmation would need ~100 opponents or 4× the runs. The honest
statement is **"a wash on wins, a small unresolvable dodging gain"**, not a win.
**Is the failure in the information or in the learner? MAINLY THE INFORMATION —
and specifically the LABEL.** Three independent measurements say so:
1. **The observable state carries almost nothing (offline, MEASURED).** On
63 held-out battles the state-conditional model beats the global histogram
and chance on log-loss, but the absolute skill is 3.93% top-1 (global 3.96%,
chance 3.23%) — ~no information about the wave-crossing bin. The strong
information in `docs/state_window_gate.md` (0.41 accuracy) came from a state
measured relative to the ENEMY'S BULLET LINE, which leaks the enemy's lead;
measured in the frame the mover can actually observe, that signal is gone.
2. **The map the mover minimises is the WRONG quantity (offline, MEASURED).**
`corr( P(arrival bin), P(hit | arrival bin) ) = −0.342` over the 31 bins: the
bins with the least mass (the clamped extremes) are where this corpus's gun
lands the MOST hits (bins 1–2 and 28–29: 14–17.5%; bins 23–25: 6.7–7.2%).
Minimising the resolved-position histogram steers INTO the bullets. No
learner can fix a mislabelled target, and this also explains j115.
3. **The learner itself is fine (live, MEASURED).** Against the identical mover
with the state removed, the state conditioning produces a CI-separated
−1.31 pp hit rate and −12.9 damage taken/run. The counted SBC learns and
extracts a real signal; the signal is just too small to beat `strafe`.
**Two secondary findings.** (a) `learned` vs `learned_nodecay` is a wash live
(12.26% vs 12.73% hit rate, Δwins +0.05) — the forgetting mechanism is NOT the
binding constraint here, and offline the no-decay arm was even the better
predictor, i.e. this opponent did not adapt to us on the decay's timescale.
(b) `learned_global` (state conditioning OFF) has the BEST pooled wins/run of
the four arms (1.78) while dodging WORSE (13.76%) — a reminder that this panel's
win signal is noisy at 3 runs/arm and that the wave-surfing geometry, not the
learning, is where the movement value lives.
### MEASURED vs INFERRED
**MEASURED:** the session identity (commit, sha, 180 battles, 0 excluded); the
pooled dashboard; every cross-opponent mean/CI/sign/p/MDE above; the
`learned` vs `learned_global` and `learned` vs `learned_nodecay` pairwise
numbers (same 180 battles, no new fighting); the offline table, the recurrence
counts, the label-shuffle control and the danger-map alignment in "Gate A";
the ms/tick cost; the clean-build verification.
**INFERRED:** (i) that the danger-map misalignment is *the* cause of the
resolved-position surfer's weakness — it is consistent with j115 (13.51% vs
9.40%) and with the near-zero offline skill, but it is not a controlled
intervention; (ii) that the small live hit-rate gain is the same mechanism the
offline gate measured; (iii) that the win leg is unresolvable rather than
absent — the CI is wide on both sides.
**Recommended follow-up (not done, not scheduled):** label by OUTCOME. A
counted+decayed SBC over (state, candidate bin) labelled HIT/MISS estimates
P(hit | state, bin) directly — the quantity the mover should minimise and the
one the alignment table shows is not the histogram. That is the single change
that the evidence here points at, and it is a different experiment from this
one.
**Status: the default is UNCHANGED (`TR_MOVEMENT=strafe`); the module is
default-off behind `TR_MOVEMENT=learned`.** Revert = do not set the env var.