j128 learned movement: state-conditional counted-SBC wave danger (TR_MOVEMENT=learned, default-off), offline gate + pre-registered panel arms
This commit is contained in:
@@ -1694,3 +1694,144 @@ override.
|
||||
|---|---|---:|---|---|
|
||||
| `/tmp/ab/j122_v2` | `5146748` | 300 (0 failed, 0 excluded) | `strafe`, `tfil` (10 runs/arm) | **gate v2 primary PASSED** (sign-flip p=0.045, CI [+0.02,+0.58]); **default FLIPPED to `strafe`** |
|
||||
|
||||
|
||||
---
|
||||
|
||||
## Learned movement (SBC) — PRE-REGISTRATION (written BEFORE any battle)
|
||||
|
||||
**The design.** A new swappable movement module
|
||||
`common_libs/movements/learned_surfer.nim`, selected by `TR_MOVEMENT=learned`
|
||||
(the shipped default `strafe` is untouched). It replaces the *constant* danger
|
||||
map of `wave_surfer` (j115: one global 31-bin histogram, no conditioning, no
|
||||
decay — it lost to both `tfil` and `strafe`) with a **state-conditional** one:
|
||||
the danger of a guess-factor bin is learned separately for each **coarse
|
||||
wave-relative movement state**, using the **counted SBC with global fractional
|
||||
decay** from `common_libs/bitbrain` (jobs j102/j103, measured to forget a
|
||||
changed mapping and to give true probabilities).
|
||||
|
||||
* **Wave**: detected from the one-tick enemy energy drop (exactly as
|
||||
`wave_surfer`/`strafe` do — `WorldState` has no bullet bodies), origin = the
|
||||
enemy position at the fire tick, centre line = the bearing from that origin to
|
||||
us at the fire tick.
|
||||
* **Label**: a wave resolves at the **nominal arrival tick**
|
||||
`ceil(startDist/speed)` and the label is the 31-bin guess factor of our
|
||||
angular offset from the centre line at that tick (`gfToBin`, the same 31-bin
|
||||
quantisation `wave_surfer` uses). The nominal rule is used instead of
|
||||
"radius >= current distance" because the latter runs away to the clamped
|
||||
`±1` bins and was measured to carry even less information.
|
||||
* **State (ONE state, never a window — `docs/state_window_gate.md` measured
|
||||
windows dead)**: 4 fields x 4 symbols = **256 states**; `vlat` (lateral
|
||||
velocity in the wave frame, px/tick), `dist` (range at the fire tick), `room`
|
||||
(directional wall room along the direction we are running), `turn` (our own
|
||||
signed heading change). Bin edges are the corpus quantiles, frozen in the
|
||||
module. `lat` is deliberately NOT a field: at the fire tick the centre line
|
||||
passes through us, so it is identically zero.
|
||||
* **Learner**: `initCountedSbc` (saturating `uint8` per (state, bin), `c -= c
|
||||
shr shift` every `decayEvery` learns), read with `inferProb` (per-cell
|
||||
posterior), interpolated with the global histogram with weight `alpha`.
|
||||
* **Decision**: danger = the predicted probability of the GF bin we would
|
||||
arrive in, SUMMED over every live wave, plus a wall penalty, a travel penalty
|
||||
and a reversal penalty; the safest reachable bin wins. Reversals stay cheap
|
||||
(the mover must not become turn-heavy).
|
||||
|
||||
**The offline veto (Gate A) — see the table in this section when it is
|
||||
appended.** Harness `common_libs/tests/learned_surfer_gate.py`, corpus
|
||||
`/tmp/tfil_ab2/out` (70 recorded battles, 54 923 shots), split BY BATTLE 70/30,
|
||||
3 seeds, veto-only per `docs/offline_harness_trust.md`.
|
||||
|
||||
**Pre-registered arms** (`tools/ab/arms_movement_learned.txt`), all on the frozen
|
||||
panel `tools/ab/panel_movement.txt`, 3 runs x 3 rounds, `--reference strafe`:
|
||||
|
||||
| arm | env | isolates |
|
||||
|---|---|---|
|
||||
| `strafe` | `TR_MOVEMENT=strafe` | the champion to beat |
|
||||
| `learned` | `TR_MOVEMENT=learned` | the module (decay 128 learns, shift 1) |
|
||||
| `learned_nodecay` | `+ TR_LEARNED_DECAY_SHIFT=0` | the counted+decay forgetting mechanism |
|
||||
| `learned_global` | `+ TR_LEARNED_GLOBAL=1` | **the state conditioning itself** (same mover, same SBC, state forced to one cell = the old global histogram) |
|
||||
|
||||
**Pre-registered decision rules (fixed before any battle):**
|
||||
|
||||
1. **Win leg (primary, the standing rule).** Cross-opponent sign-flip
|
||||
permutation test on the paired per-opponent Δwins/run, two-sided p < 0.05,
|
||||
AND the pooled 95% CI excludes 0, AND the point estimate is positive in the
|
||||
challenger's favour. Only then does the challenger "beat" the reference.
|
||||
2. **Mechanism leg.** The same test on the **incoming hit rate** (the dodging
|
||||
metric, and here the mechanism being claimed). A hit-rate win with a flat
|
||||
win leg is reported as *"dodges better, wins the same"*, not as a win.
|
||||
3. **Information-vs-learner split (declared now, not after seeing the data).**
|
||||
* `learned` ≈ `learned_global` ⇒ the failure is in the **information**: the
|
||||
coarse observable state carries nothing the global histogram does not.
|
||||
* `learned` > `learned_global` but `learned` ≤ `strafe` ⇒ the state
|
||||
conditioning helps *relative to the old surfer* but the whole learned
|
||||
family is still behind the hand-tuned champion.
|
||||
* `learned` < `learned_nodecay` ⇒ the decay is hurting (the opponent does
|
||||
not in fact adapt on the timescale of the decay).
|
||||
4. **The default is NOT touched.** `strafe` stays shipped whatever the result.
|
||||
|
||||
**Pre-registered prediction (recorded before the battles; my honest prior).**
|
||||
The offline gate shows the state-conditional model beats the global histogram
|
||||
and chance on held-out log-loss (4.927 vs 4.974 vs 4.954 bits) in **63/63**
|
||||
held-out battles (sign-flip p = 5e-5) — but the absolute skill is tiny
|
||||
(top-1 3.93%, global 3.96%, chance 3.23%). **I therefore predict `learned` will
|
||||
NOT beat `strafe` on round wins, that its incoming hit rate will be within
|
||||
noise of `strafe`'s, and that `learned` ≈ `learned_global` — i.e. the failure
|
||||
is expected to be in the information, not in the learner.** A negative here is
|
||||
the expected outcome and is a fully successful result.
|
||||
|
||||
**Session:** `/tmp/ab/j128_learned`, frozen from the commit that contains this
|
||||
pre-registration.
|
||||
|
||||
### Gate A — offline prediction quality (MEASURED, before any battle)
|
||||
|
||||
Command: `python3 common_libs/tests/learned_surfer_gate.py --corpus
|
||||
/tmp/tfil_ab2/out --label nominal --report
|
||||
common_libs/tests/fixtures/learned_surfer_gate_report.txt --json
|
||||
common_libs/tests/fixtures/learned_surfer_gate.json` (70 battles, 54 923
|
||||
shots, split BY BATTLE 70/30, 3 seeds, ~1 min).
|
||||
|
||||
**Held-out prediction quality** (mean over the 3 battle splits; lower log-loss /
|
||||
higher accuracy is better):
|
||||
|
||||
| predictor | log-loss (bits) | top-1 | top-3 |
|
||||
|---|---:|---:|---:|
|
||||
| chance (uniform over 31 bins) | 4.9542 | 3.23% | 9.68% |
|
||||
| unconditional average / old global 31-bin histogram | 4.9739 | 3.96% | 12.15% |
|
||||
| majority bin (degenerate top-1) | 4.9739 | 4.63% | n/a |
|
||||
| **state-conditional counted SBC (Q4, decay 128/1)** | **4.9272** | 3.93% | **12.24%** |
|
||||
| state-conditional, no decay | 4.8408 | **6.33%** | 15.72% |
|
||||
| state-conditional, Q3 (81 states) | 4.9401 | 3.96% | 12.44% |
|
||||
| state-conditional, Q5 (625 states) | 4.9200 | 4.02% | 12.40% |
|
||||
| **label-shuffle control** (same states, train labels permuted) | 4.9480 | 3.84% | — |
|
||||
|
||||
* The unconditional average and "the 31-bin global histogram of the old surfer"
|
||||
are **the same estimator by construction** (both are the train marginal over
|
||||
bins), so they are one row. The old surfer's histogram is *worse than a
|
||||
uniform guess* on held-out log-loss because an unsmoothed 31-bin marginal is
|
||||
over-confident; that is a calibration fact, not a win for the learner.
|
||||
* **RECURRENCE IS NOT THE PROBLEM**: 256 declared states, ~255 distinct seen,
|
||||
**150 observations per state**, and **100.0%** of held-out shots fall in a
|
||||
state that occurred in training. The j115 failure was not a recurrence
|
||||
failure; neither is this.
|
||||
* The state-conditional model beats the global histogram and chance on
|
||||
held-out log-loss in **63/63** held-out battles: pooled Δlog-loss
|
||||
**−0.0467 bits**, 95% CI [−0.0481, −0.0453], sign 0/63, sign-flip
|
||||
p = 5e-5, MDE 0.0021.
|
||||
* The label-shuffle control collapses the gain to −0.0056 bits, so the gain is
|
||||
real and comes from the state.
|
||||
* **But the effect is TINY in absolute terms**: 0.047 bits out of 4.95, and
|
||||
top-1 3.93% vs chance 3.23% vs global 3.96% — the state buys ~27% relative
|
||||
top-1 over *chance* and **nothing over the global histogram on top-1**.
|
||||
* **The diagnosis of why.** At the fire tick the only strongly predictive
|
||||
quantity in the wave frame is the enemy's own lead (its bullet direction),
|
||||
which the mover cannot observe. Measured on the same corpus: an *oracle*
|
||||
state map (edges fitted on all data) reaches top-1 **20.8%** on the enemy's
|
||||
true AIM bin (marginal 19.0%) from the observable state, and the sign of our
|
||||
lateral velocity agrees with the enemy's aim bin only **64.1%** of the time
|
||||
(against **58.8%** for the resolved crossing bin the module can label). The
|
||||
observable state is nearly uninformative about where the wave crosses us.
|
||||
|
||||
**Gate A verdict: the veto does NOT fire** — the state-conditional model is
|
||||
better than the global histogram, the unconditional average and chance, with a
|
||||
consistent cross-battle sign. But it clears the bar by ~1% of a bit, so the
|
||||
live panel is the decider, and the pre-registered prediction above is that the
|
||||
module will NOT beat `strafe`.
|
||||
|
||||
Reference in New Issue
Block a user