learned movement (SBC): live panel results and the direct answer (module is a wash vs strafe on wins; state conditioning is a CI-separated hit-rate gain vs the same mover without it; the label is the wrong quantity)
This commit is contained in:
@@ -1864,3 +1864,140 @@ against `strafe`'s 9.40%. What the mover **should** learn is the outcome: a
|
||||
counted/decayed SBC over states and bins labelled by HIT/MISS would estimate
|
||||
P(hit | state, bin) directly. That is the natural next experiment and it is NOT
|
||||
what was measured here.)*
|
||||
|
||||
---
|
||||
|
||||
## Learned movement (SBC) — RESULTS (appended AFTER the battles)
|
||||
|
||||
**Session `/tmp/ab/j128_learned`, frozen from the pre-registration commit
|
||||
`a436e9f` (binary sha256 `60f2093b58b8…`): 15 opponents × 4 arms × 3 runs ×
|
||||
3 rounds = 180 battles, conc 6, 0 excluded, 0 failed starts.** Reference:
|
||||
`strafe` (the shipped champion). Every delta is (arm − `strafe`).
|
||||
|
||||
### Pooled dashboard (descriptive, NOT the verdict)
|
||||
|
||||
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | **incoming hit rate** | mean distance |
|
||||
|---|---:|---:|---:|---:|---:|---:|---:|---:|
|
||||
| `strafe` (champion) | 45 | 106.0 | 152.4 | 1.56 | 70/135 | 51.9% | 13.07% | 431 |
|
||||
| `learned` (decay on) | 45 | 97.3 | 138.7 | 1.76 | 79/135 | 58.5% | **12.26%** | 408 |
|
||||
| `learned_nodecay` | 45 | 104.7 | 146.3 | 1.71 | 77/135 | 57.0% | 12.73% | 410 |
|
||||
| `learned_global` (state conditioning OFF) | 45 | 99.4 | 151.7 | 1.78 | 80/135 | 59.3% | 13.76% | 407 |
|
||||
|
||||
### Cross-opponent aggregation (the verdict layer), reference `strafe`
|
||||
|
||||
| arm | metric | mean Δ | spread (SD) | 95% CI | sign test | p(sign) | p(sign-flip) | MDE |
|
||||
|---|---|---:|---:|---|---:|---:|---:|---:|
|
||||
| `learned` | wins | +0.20 | 0.73 | **[−0.21, +0.61]** | 6/11 | 1 | 0.3662 | **0.53** |
|
||||
| `learned` | damage | −8.67 | 30.54 | [−25.58, +8.25] | 7/15 | 1 | 0.2953 | 22.09 |
|
||||
| `learned` | damage_taken | −13.64 | 41.46 | [−36.60, +9.33] | 7/15 | 1 | 0.2311 | 29.99 |
|
||||
| `learned` | **hit_rate** | **−1.01 pp** | 4.05 | **[−3.26, +1.23]** | 5/15 | 0.3018 | 0.3437 | **2.93** |
|
||||
| `learned_nodecay` | wins | +0.16 | 0.69 | [−0.23, +0.54] | 6/9 | 0.5078 | 0.4688 | 0.50 |
|
||||
| `learned_nodecay` | hit_rate | −1.13 pp | 3.97 | [−3.33, +1.07] | 7/15 | 1 | 0.291 | 2.87 |
|
||||
| `learned_global` | wins | +0.22 | 0.88 | [−0.26, +0.71] | 7/12 | 0.7744 | 0.394 | 0.64 |
|
||||
| `learned_global` | hit_rate | **+0.30 pp** | 4.17 | [−2.01, +2.61] | 8/15 | 1 | 0.783 | 3.02 |
|
||||
|
||||
**Pre-registered verdict vs `strafe`: NO ARM BEATS THE CHAMPION.** All three
|
||||
learned arms are **not distinguishable** from `strafe` on round wins *and* on
|
||||
damage, by both readings of the pre-registered rule. The win leg (rule 1) fails
|
||||
for every arm: the sign-flip p-values are 0.37 / 0.47 / 0.39 and every 95% CI
|
||||
contains 0. The mechanism leg (rule 2) also fails: the incoming hit rate is
|
||||
−1.01 pp for `learned` (CI [−3.26, +1.23], MDE 2.93 pp) — pointing the right
|
||||
way, but smaller than this batch can resolve.
|
||||
|
||||
### The arm that DOES separate: state conditioning vs the same mover without it
|
||||
|
||||
`learned` vs `learned_global` is a free pairwise comparison on the same 180
|
||||
battles (re-analyze with `--reference learned_global`): identical binary,
|
||||
identical wave geometry, identical counted SBC and priors — the only difference
|
||||
is that `learned_global` forces the state to a single cell (the old global
|
||||
histogram).
|
||||
|
||||
| `learned` − `learned_global` | mean Δ | 95% CI | sign-flip p | MDE |
|
||||
|---|---:|---|---:|---:|
|
||||
| **incoming hit rate** | **−1.31 pp** | **[−2.58, −0.05]** | **0.0444** | 1.65 |
|
||||
| **damage taken/run** | **−12.93** | **[−25.81, −0.05]** | **0.0485** | 16.82 |
|
||||
| wins/run | −0.02 | [−0.27, +0.22] | 1.0 | 0.32 |
|
||||
| damage/run | −2.03 | [−12.13, +8.08] | 0.667 | 13.20 |
|
||||
|
||||
**The state conditioning is a REAL, measurable dodging improvement** — the
|
||||
incoming hit rate drops 1.31 pp with a CI that excludes 0 and sign-flip
|
||||
p = 0.044, and damage taken drops 12.9/run with a CI that excludes 0 — **but it
|
||||
does not move round wins at all** (Δwins −0.02). So the learned state
|
||||
conditioning works as advertised and is simply too small to matter for the
|
||||
score against this panel.
|
||||
|
||||
### Cost (MEASURED, `-d:release`, 200k ticks, `git archive HEAD` clean build)
|
||||
|
||||
| scenario | mean ms/tick | worst single tick observed |
|
||||
|---|---:|---:|
|
||||
| 1v1 (decision every tick a wave is live) | **0.0024** | 1.14 ms |
|
||||
| 4 enemies | 0.0085 | 0.56 ms |
|
||||
|
||||
Budget is 13.16 ms/tick; the module uses **0.02%** of it. Memory: one
|
||||
`uint8` per (state × bin) = 16×16×31 = 7 936 B. It is not a cost problem.
|
||||
|
||||
### Direct answer
|
||||
|
||||
**Does state-conditional learned danger beat the hand-tuned `strafe` on dodging
|
||||
and/or on wins? NO — on neither, by the pre-registered rules.** The point
|
||||
estimates lean the module's way (wins +0.20/run, hit rate −1.01 pp, damage taken
|
||||
−13.6/run) but every CI contains 0 and the win-leg MDE (0.53 wins/run) is 2.6×
|
||||
the observed effect: this batch cannot resolve an effect of the measured size,
|
||||
and a confirmation would need ~100 opponents or 4× the runs. The honest
|
||||
statement is **"a wash on wins, a small unresolvable dodging gain"**, not a win.
|
||||
|
||||
**Is the failure in the information or in the learner? MAINLY THE INFORMATION —
|
||||
and specifically the LABEL.** Three independent measurements say so:
|
||||
|
||||
1. **The observable state carries almost nothing (offline, MEASURED).** On
|
||||
63 held-out battles the state-conditional model beats the global histogram
|
||||
and chance on log-loss, but the absolute skill is 3.93% top-1 (global 3.96%,
|
||||
chance 3.23%) — ~no information about the wave-crossing bin. The strong
|
||||
information in `docs/state_window_gate.md` (0.41 accuracy) came from a state
|
||||
measured relative to the ENEMY'S BULLET LINE, which leaks the enemy's lead;
|
||||
measured in the frame the mover can actually observe, that signal is gone.
|
||||
2. **The map the mover minimises is the WRONG quantity (offline, MEASURED).**
|
||||
`corr( P(arrival bin), P(hit | arrival bin) ) = −0.342` over the 31 bins: the
|
||||
bins with the least mass (the clamped extremes) are where this corpus's gun
|
||||
lands the MOST hits (bins 1–2 and 28–29: 14–17.5%; bins 23–25: 6.7–7.2%).
|
||||
Minimising the resolved-position histogram steers INTO the bullets. No
|
||||
learner can fix a mislabelled target, and this also explains j115.
|
||||
3. **The learner itself is fine (live, MEASURED).** Against the identical mover
|
||||
with the state removed, the state conditioning produces a CI-separated
|
||||
−1.31 pp hit rate and −12.9 damage taken/run. The counted SBC learns and
|
||||
extracts a real signal; the signal is just too small to beat `strafe`.
|
||||
|
||||
**Two secondary findings.** (a) `learned` vs `learned_nodecay` is a wash live
|
||||
(12.26% vs 12.73% hit rate, Δwins +0.05) — the forgetting mechanism is NOT the
|
||||
binding constraint here, and offline the no-decay arm was even the better
|
||||
predictor, i.e. this opponent did not adapt to us on the decay's timescale.
|
||||
(b) `learned_global` (state conditioning OFF) has the BEST pooled wins/run of
|
||||
the four arms (1.78) while dodging WORSE (13.76%) — a reminder that this panel's
|
||||
win signal is noisy at 3 runs/arm and that the wave-surfing geometry, not the
|
||||
learning, is where the movement value lives.
|
||||
|
||||
### MEASURED vs INFERRED
|
||||
|
||||
**MEASURED:** the session identity (commit, sha, 180 battles, 0 excluded); the
|
||||
pooled dashboard; every cross-opponent mean/CI/sign/p/MDE above; the
|
||||
`learned` vs `learned_global` and `learned` vs `learned_nodecay` pairwise
|
||||
numbers (same 180 battles, no new fighting); the offline table, the recurrence
|
||||
counts, the label-shuffle control and the danger-map alignment in "Gate A";
|
||||
the ms/tick cost; the clean-build verification.
|
||||
|
||||
**INFERRED:** (i) that the danger-map misalignment is *the* cause of the
|
||||
resolved-position surfer's weakness — it is consistent with j115 (13.51% vs
|
||||
9.40%) and with the near-zero offline skill, but it is not a controlled
|
||||
intervention; (ii) that the small live hit-rate gain is the same mechanism the
|
||||
offline gate measured; (iii) that the win leg is unresolvable rather than
|
||||
absent — the CI is wide on both sides.
|
||||
|
||||
**Recommended follow-up (not done, not scheduled):** label by OUTCOME. A
|
||||
counted+decayed SBC over (state, candidate bin) labelled HIT/MISS estimates
|
||||
P(hit | state, bin) directly — the quantity the mover should minimise and the
|
||||
one the alignment table shows is not the histogram. That is the single change
|
||||
that the evidence here points at, and it is a different experiment from this
|
||||
one.
|
||||
|
||||
**Status: the default is UNCHANGED (`TR_MOVEMENT=strafe`); the module is
|
||||
default-off behind `TR_MOVEMENT=learned`.** Revert = do not set the env var.
|
||||
|
||||
Reference in New Issue
Block a user