j130 learned movement outcome label: live panel results (no arm beats strafe; label swap is a dead heat; gap is information) + Gate A report

This commit is contained in:
2026-09-26 11:35:47 +02:00
parent 61def1c3e9
commit 8dd9b3b3b5
2 changed files with 234 additions and 0 deletions
@@ -0,0 +1,83 @@
# Outcome-label Gate A — offline sanity check (VETO ONLY)
corpus : /tmp/tfil_ab2/out
battles : 70
shots : 54923
base hit : 9.97%
state : vlat, dist, room, turn (the module's 4 fields, canonical edges)
label : outcome hit(state,g) = hit and |g - b_our| <= w (w = body width as an angle)
## A. is the danger map the mover MINIMISES aligned with the realised per-bin hit rate?
corr( danger(g) , P(hit | b_our = g) ) over the 31 bins:
| danger map | corr |
|---|---:|
| histogram label (j128) — P(arrival bin = g) | -0.341 |
| **outcome label (j130, the module's live label)** | **+0.566** |
| geometric bullet-line label (needs bullet bodies) | -0.230 |
Negative = minimising the danger steers INTO the bullets (the j128 defect). The histogram reproduces the ledger's -0.342.
| bin | P(hit) | P(arrival=bin) | outcome danger |
|---:|---:|---:|---:|
| 0 | 9.7% | 2.4% | 0.005 |
| 1 | 14.1% | 1.9% | 0.008 |
| 2 | 13.2% | 2.6% | 0.008 |
| 3 | 10.8% | 2.3% | 0.008 |
| 4 | 8.8% | 2.5% | 0.007 |
| 5 | 7.8% | 2.9% | 0.007 |
| 6 | 7.9% | 3.3% | 0.008 |
| 7 | 8.8% | 3.5% | 0.009 |
| 8 | 9.2% | 3.6% | 0.009 |
| 9 | 9.1% | 3.8% | 0.010 |
| 10 | 9.8% | 3.9% | 0.011 |
| 11 | 11.0% | 4.0% | 0.011 |
| 12 | 10.6% | 3.9% | 0.012 |
| 13 | 11.1% | 4.1% | 0.011 |
| 14 | 9.0% | 4.0% | 0.011 |
| 15 | 9.8% | 4.6% | 0.011 |
| 16 | 9.8% | 3.8% | 0.010 |
| 17 | 9.2% | 3.9% | 0.010 |
| 18 | 10.7% | 3.6% | 0.009 |
| 19 | 8.8% | 3.5% | 0.008 |
| 20 | 8.5% | 3.4% | 0.008 |
| 21 | 7.9% | 3.5% | 0.007 |
| 22 | 7.5% | 3.3% | 0.007 |
| 23 | 6.7% | 3.1% | 0.007 |
| 24 | 8.3% | 2.8% | 0.006 |
| 25 | 7.2% | 2.6% | 0.007 |
| 26 | 11.5% | 2.3% | 0.007 |
| 27 | 13.1% | 2.2% | 0.010 |
| 28 | 17.5% | 2.6% | 0.012 |
| 29 | 16.2% | 2.8% | 0.012 |
| 30 | 10.9% | 3.3% | 0.008 |
## B. state-conditional information under the OUTCOME label
held-out per-candidate log-loss (bits) of the outcome label, state-conditional vs state-free (same rows, same split):
| model | log-loss (bits) |
|---|---:|
| state-free P(hit | g) | 0.1873 |
| state-conditional P(hit | state, g) | 0.3906 |
| Δ (state − state-free) | +0.2033 |
state conditioning is better in 0/3 splits (negative Δ = better).
## C. open-loop decision counterfactual (VETO ONLY)
If the mover picks argmin_g danger, the fraction of held-out waves whose bullet line would still pass within a body width of g (ground truth = the recorded bullet line b_bullet).
| policy | held-out waves still hit |
|---|---:|
| histogram argmin (j128) | 3.53% |
| outcome argmin (j130) | 3.33% |
| recorded trajectory (floor/ceiling) | 10.17% |
The counterfactual is OPEN LOOP: the recorded bullet lines were fired at a different mover, so it cannot predict the live closed loop. It is a veto, not a selection.
## MEASURED vs INFERRED
* MEASURED: every number above, on the recorded corpus.
* INFERRED: that the offline alignment transfers live. It cannot — see docs/offline_harness_trust.md.
+151
View File
@@ -2092,3 +2092,154 @@ fully successful outcome.
**Session:** `/tmp/ab/j130_outcome`, frozen from the commit that contains this **Session:** `/tmp/ab/j130_outcome`, frozen from the commit that contains this
pre-registration. pre-registration.
### Gate A — MEASURED (offline, before the battle)
`python3 common_libs/tests/outcome_label_gate.py --corpus /tmp/tfil_ab2/out
--report common_libs/tests/fixtures/outcome_label_gate_report.txt`
(70 battles, 54 923 shots, canonical module edges, split BY BATTLE 70/30,
3 seeds).
| danger map | corr( danger(g) , P(hit \| b_our=g) ) |
|---|---:|
| histogram label (j128) — P(arrival bin = g) | **−0.341** |
| **outcome label (j130, the module's live label)** | **+0.566** |
| geometric bullet-line label (needs bullet bodies) | −0.230 |
The alignment **flips positive** — the veto does not fire. But:
* **State-conditional information is NEGATIVE.** Held-out per-candidate
log-loss of the outcome label: state-free `P(hit | g)` **0.1873 bits** vs
state-conditional `P(hit | state, g)` **0.3906 bits** (Δ **+0.203**, better in
**0/3** splits). Under the outcome label the coarse state does **not** help;
the state-free model is better.
* **Decision counterfactual barely moves** (open-loop, VETO ONLY): argmin danger
with the recorded bullet line as ground truth — histogram **3.53%**, outcome
**3.33%**, recorded trajectory **10.17%**.
### RESULTS (appended AFTER the battles)
**Session `/tmp/ab/j130_outcome`, frozen from the pre-registration commit
`61def1c` (binary sha256 `e74c6c788ddf…`): 15 opponents × 4 arms × 3 runs ×
3 rounds = 180 battles, conc 6, 0 excluded, 0 failed starts.** Reference:
`strafe`. Every delta is (arm − `strafe`).
### Pooled dashboard (descriptive, NOT the verdict)
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | **incoming hit rate** | mean distance |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| `strafe` (champion) | 45 | 113.9 | 153.9 | 1.62 | 73/135 | 54.1% | 12.84% | 427 |
| `learned` (old label) | 45 | 101.6 | 146.6 | 1.76 | 79/135 | 58.5% | 13.57% | 408 |
| `learned_outcome` (new label) | 45 | 96.3 | 141.7 | 1.71 | 77/135 | 57.0% | 13.09% | 406 |
| `learned_outcome_global` (state OFF) | 45 | 94.1 | 152.2 | 1.58 | 71/135 | 52.6% | 14.19% | 409 |
### Cross-opponent aggregation, reference `strafe` (the verdict layer)
| arm | metric | mean Δ | spread (SD) | 95% CI | sign test | p(sign) | p(sign-flip) | MDE |
|---|---|---:|---:|---|---:|---:|---:|---:|
| `learned` | wins | +0.13 | 0.65 | [−0.23, +0.49] | 6/9 | 0.508 | 0.523 | 0.47 |
| `learned` | damage | −12.4 | 21.8 | [−24.4, −0.3] | 6/15 | 0.607 | 0.047 | 15.8 |
| `learned` | hit_rate | −0.66 pp | 5.84 | [−3.89, +2.57] | 8/15 | 1 | 0.668 | 4.22 |
| **`learned_outcome`** | **wins** | **+0.09** | 0.53 | **[−0.20, +0.38]** | 6/12 | 1 | 0.645 | 0.38 |
| `learned_outcome` | damage | −17.6 | 25.2 | [−31.6, −3.6] | **3/15** | **0.035** | 0.017 | 18.3 |
| `learned_outcome` | hit_rate | −1.12 pp | 6.33 | [−4.63, +2.39] | 9/15 | 0.607 | 0.507 | 4.58 |
| `learned_outcome_global` | wins | −0.04 | 0.71 | [−0.44, +0.35] | 6/12 | 1 | 0.907 | 0.51 |
| `learned_outcome_global` | damage | −19.9 | 24.1 | [−33.2, −6.5] | 3/15 | 0.035 | 0.007 | 17.4 |
| `learned_outcome_global` | hit_rate | +0.24 pp | 6.63 | [−3.43, +3.91] | 9/15 | 0.607 | 0.891 | 4.79 |
**Pre-registered verdict vs `strafe`: NO ARM BEATS THE CHAMPION.** The win leg
(rule 1) fails for every arm — every Δwins/run 95% CI contains 0 and no
sign-flip p clears 0.05. `learned_outcome` is **not distinguishable** from
`strafe` on wins (+0.09) and on hit rate (−1.12 pp, CI [−4.63, +2.39], MDE
4.58) but is **detectably WORSE on damage** (−17.6/run, CI [−31.6, −3.6], sign
3/15 p = 0.035). Under the pre-registered substantive reading it is **WORSE**,
not a win.
### The two pairwise isolations (same 180 battles, no new fighting)
**(a) The LABEL, isolated: `learned_outcome` vs `learned`** (re-analyze with
`--reference learned`) — the only difference is
`TR_LEARNED_LABEL=histogram|outcome`:
| `learned_outcome` − `learned` | mean Δ | 95% CI | sign-flip p | MDE |
|---|---:|---|---:|---:|
| wins/run | −0.04 | [−0.43, +0.34] | 0.902 | 0.50 |
| incoming hit rate | −0.46 pp | [−2.67, +1.75] | 0.664 | 2.88 |
| damage/run | −5.26 | [−16.94, +6.42] | 0.347 | 15.3 |
| damage taken/run | −4.90 | [−24.72, +14.92] | 0.597 | 25.9 |
**The new label changes nothing measurable live.** Wins, hit rate and damage are
all statistically indistinguishable from the old arrival-bin label.
**(b) The STATE under the new label: `learned_outcome` vs
`learned_outcome_global`** (re-analyze with `--reference learned_outcome`):
| `learned_outcome_global` − `learned_outcome` | mean Δ | 95% CI | sign-flip p | MDE |
|---|---:|---|---:|---:|
| incoming hit rate | +1.36 pp | [−0.83, +3.55] | 0.204 | 2.86 |
| wins/run | −0.13 | [−0.36, +0.10] | 0.363 | 0.30 |
| damage/run | −2.25 | [−12.50, +7.99] | 0.667 | 13.4 |
| damage taken/run | +10.56 | [−7.85, +28.98] | 0.244 | 24.1 |
Turning the state conditioning OFF **costs 1.36 pp of incoming hit rate**
(state-conditional dodges better) — the same sign and roughly the same size as
j128's −1.31 pp, but again **not CI-separated at this n** and it does not move
round wins.
### Direct answer
**Does learning `P(hit | state, direction)` fix the inversion? OFFLINE, YES;
LIVE, IT DOES NOT CHANGE ANYTHING. Does it beat `strafe`? NO.**
* The **inversion is fixed in the correlation sense**: the danger the mover
minimises goes from `corr = −0.341` (histogram) to `+0.566` (outcome). The
outcome-labelled danger is no longer anti-aligned with where hits happen.
* But the **offline decision counterfactual barely moves** (3.53% → 3.33%) and,
**live, the label swap is a dead heat** with the old one (wins −0.04,
hit-rate −0.46 pp, all CIs far inside the MDE). The mechanism the label was
supposed to fix never reaches the score.
* **The remaining gap is INFORMATION, not the learner.** Three measurements say
so: (i) offline, the state-conditional outcome model is *worse* than the
state-free one on held-out log-loss (+0.203 bits, 0/3 splits) — the state buys
no information under the outcome label either; (ii) live, the state
conditioning is worth only ~1.4 pp of hit rate (`learned_outcome` vs its
state-free ablation), below this design's MDE (2.86 pp) and worth 0 wins;
(iii) the label swap itself (a pure supervision change) moves nothing. The
counterfactual hits are concentrated where the enemy's fixed bullet line is,
and within a single wave that line is **unobservable** to a bot with no bullet
bodies — neither the histogram label nor the outcome label creates the missing
information, it only re-weights it.
**Honest reading of the negative.** The campaign's champion `strafe` is a
hand-tuned wave-geometry mover; the learned family (both labels) matches it on
wins but pays a small damage cost and cannot separate. This is now the **third**
independent negative for the learned-surfer family (j115 hand-written, j128
histogram label, j130 outcome label), which is itself the answer to the honest
question: **hand-tuned movement is simply hard to beat on this panel**, and the
binding constraint is the observable state, not the label or the learner.
### MEASURED vs INFERRED
**MEASURED:** the session identity (commit, sha, 180 battles, 0 excluded); the
pooled dashboard; every cross-opponent mean/CI/sign/p/MDE above; the two
pairwise isolations (same 180 battles, no new fighting); the Gate A correlation
table, the state-conditional log-loss table and the decision counterfactual; the
module unit tests (14/14) and the false-premise scan (the histogram label's
−0.342 is reproduced exactly).
**INFERRED:** that the offline correlation/decision numbers transfer live (they
do not — the corpus is open-loop); that the ~1.4 pp state-conditioning hit-rate
gain is the true effect (it is below MDE and not separated).
**PREDICTION RECORDED AS PARTLY WRONG.** The pre-registration predicted that
`learned_outcome` would NOT beat `strafe` on wins (CORRECT: +0.09, CI includes
0), that its hit rate would be within noise of `strafe`'s (CORRECT: −1.12 pp,
CI [−4.63, +2.39]), and that `learned_outcome` ≈ `learned_outcome_global`
(CORRECT on wins, −0.13; **WRONG on the hit rate**: the state conditioning is
worth −1.36 pp, same sign as j128, though not CI-separated). I also did not
predict the detectably **worse** damage (−17.6, p = 0.035), which the
pre-registered rule records as WORSE.
**Status: the default is UNCHANGED (`TR_MOVEMENT=strafe`); the outcome mode is
default-off behind `TR_MOVEMENT=learned TR_LEARNED_LABEL=outcome`.** Revert = do
not set the env vars.