j130 learned movement: outcome label (P(hit|state,g)) mode + Gate A; pre-registered outcome arms

This commit is contained in:
2026-09-26 11:27:03 +02:00
parent 4c934a67c1
commit 61def1c3e9
6 changed files with 695 additions and 24 deletions
+83
View File
@@ -2009,3 +2009,86 @@ one.
**Status: the default is UNCHANGED (`TR_MOVEMENT=strafe`); the module is
default-off behind `TR_MOVEMENT=learned`.** Revert = do not set the env var.
---
## Learned movement — outcome label (P(hit)) — PRE-REGISTRATION (written BEFORE any battle)
**The change.** j128 labelled a resolved wave by the 31-bin **GF bin we crossed
at**, and measured `corr( P(arrival bin), P(hit | arrival bin) ) = −0.342` over
the 31 bins (`learned_surfer_gate.py` section F): the least-visited bins are the
ones the gun lands the most hits in, so minimising the resolved-position
histogram steers **into** the bullets. j130 stops predicting *where* the wave
goes and learns the **outcome** directly:
> `hit(state, g) = hit and |g − b| <= window(wave)` — would this wave have hit
> me at candidate direction `g`?
where `b` is the bin the wave resolved at and `window` is the bot's body width
as an angle at that wave's distance, in GF bins
(`asin(18 / d) / asin(8 / speed) · (31−1)/2`). One resolved wave yields a label
for **every** candidate direction (dense), which attacks the volume/starvation
constraint. The learner stays the counted+decayed SBC (`common_libs/bitbrain`,
in a 2-class readout `P(hit | state, g)`), the geometry, penalties and mover are
j128's, so the two labels are isolated against each other.
**New knob:** `TR_LEARNED_LABEL=histogram` (default — today's behaviour) or
`outcome`; registered in `env_report.knownEnvNames()`. Both are default-off
behind `TR_MOVEMENT=learned`; the shipped `strafe` default is untouched.
**Gate A (offline veto) — `common_libs/tests/outcome_label_gate.py`, corpus
`/tmp/tfil_ab2/out`, 70 battles, 54 923 shots, split BY BATTLE 70/30, 3 seeds,
the module's canonical state edges.**
* **Alignment.** `corr( learned danger(g), P(hit | b_our=g) )` over the 31 bins:
histogram **−0.341**; the module's live outcome label (hit-window around the
resolved bin) **+0.566**; the pure geometric bullet-line label (needs bullet
bodies, not available live) −0.230. **The correlation flips positive, so the
veto does NOT fire.**
* **State-conditional information.** Held-out per-candidate log-loss of the
outcome label: state-free `P(hit | g)` **0.1873 bits**, state-conditional
`P(hit | state, g)` **0.3906 bits** (Δ **+0.203**, better in **0/3** splits):
under the outcome label the coarse state does **not** help — it overfits.
* **Open-loop decision counterfactual** (argmin danger, ground truth = the
recorded bullet line; veto-only): histogram 3.53%, outcome 3.33%, recorded
trajectory 10.17% — the counterfactual **barely moves**.
**Pre-registered arms** (`tools/ab/arms_movement_outcome.txt`), frozen panel
`tools/ab/panel_movement.txt`, 3 runs × 3 rounds, `--reference strafe`:
| arm | env | isolates |
|---|---|---|
| `strafe` | `TR_MOVEMENT=strafe` | the shipped champion — has to be beaten |
| `learned` | `TR_MOVEMENT=learned` | the **old label** (j128 arrival bin) |
| `learned_outcome` | `+ TR_LEARNED_LABEL=outcome` | the **new label** (dense P(hit)) |
| `learned_outcome_global` | `+ TR_LEARNED_LABEL=outcome TR_LEARNED_GLOBAL=1` | the information control: outcome label, state OFF |
**Pre-registered decision rules (fixed before any battle):**
1. **Win leg (primary, the standing rule).** Cross-opponent sign-flip
permutation test on the paired per-opponent Δwins/run, two-sided p < 0.05,
AND the pooled 95% CI excludes 0, AND the point estimate is positive in the
challenger's favour. Only then does an arm "beat" `strafe`.
2. **Mechanism leg.** The same test on the **incoming hit rate** (the dodging
metric, and the mechanism the outcome label claims). A hit-rate win with a
flat win leg is "dodges better, wins the same", not a win.
3. **Information-vs-learner split (declared now).**
* `learned_outcome` ≈ `learned_outcome_global` ⇒ the failure is the
**information** (the state is uninformative under the outcome label too).
* `learned_outcome` > `learned_outcome_global` but `learned_outcome` ≤
`strafe` ⇒ the state helps relative to its own ablation but the learned
family is still behind the hand-tuned champion.
* `learned_outcome` > `learned` (on hit rate) ⇒ the new label is a genuine
improvement over the old one, even if the family loses to `strafe`.
4. **The default is NOT touched.** `strafe` stays shipped whatever the result.
**Pre-registered prediction (honest prior).** Gate A's alignment flips positive
but the state buys no held-out information under the outcome label and the
decision counterfactual is flat, so I predict **`learned_outcome` will NOT beat
`strafe` on round wins**, that its hit rate will be within noise of `strafe`'s,
and that `learned_outcome` ≈ `learned_outcome_global` — i.e. the failure is in
the information, not in the learner or the label. A negative is the expected,
fully successful outcome.
**Session:** `/tmp/ab/j130_outcome`, frozen from the commit that contains this
pre-registration.