j130 learned movement outcome label: live panel results (no arm beats strafe; label swap is a dead heat; gap is information) + Gate A report

This commit is contained in:
2026-09-26 11:35:47 +02:00
parent 61def1c3e9
commit 8dd9b3b3b5
2 changed files with 234 additions and 0 deletions
@@ -0,0 +1,83 @@
# Outcome-label Gate A — offline sanity check (VETO ONLY)
corpus : /tmp/tfil_ab2/out
battles : 70
shots : 54923
base hit : 9.97%
state : vlat, dist, room, turn (the module's 4 fields, canonical edges)
label : outcome hit(state,g) = hit and |g - b_our| <= w (w = body width as an angle)
## A. is the danger map the mover MINIMISES aligned with the realised per-bin hit rate?
corr( danger(g) , P(hit | b_our = g) ) over the 31 bins:
| danger map | corr |
|---|---:|
| histogram label (j128) — P(arrival bin = g) | -0.341 |
| **outcome label (j130, the module's live label)** | **+0.566** |
| geometric bullet-line label (needs bullet bodies) | -0.230 |
Negative = minimising the danger steers INTO the bullets (the j128 defect). The histogram reproduces the ledger's -0.342.
| bin | P(hit) | P(arrival=bin) | outcome danger |
|---:|---:|---:|---:|
| 0 | 9.7% | 2.4% | 0.005 |
| 1 | 14.1% | 1.9% | 0.008 |
| 2 | 13.2% | 2.6% | 0.008 |
| 3 | 10.8% | 2.3% | 0.008 |
| 4 | 8.8% | 2.5% | 0.007 |
| 5 | 7.8% | 2.9% | 0.007 |
| 6 | 7.9% | 3.3% | 0.008 |
| 7 | 8.8% | 3.5% | 0.009 |
| 8 | 9.2% | 3.6% | 0.009 |
| 9 | 9.1% | 3.8% | 0.010 |
| 10 | 9.8% | 3.9% | 0.011 |
| 11 | 11.0% | 4.0% | 0.011 |
| 12 | 10.6% | 3.9% | 0.012 |
| 13 | 11.1% | 4.1% | 0.011 |
| 14 | 9.0% | 4.0% | 0.011 |
| 15 | 9.8% | 4.6% | 0.011 |
| 16 | 9.8% | 3.8% | 0.010 |
| 17 | 9.2% | 3.9% | 0.010 |
| 18 | 10.7% | 3.6% | 0.009 |
| 19 | 8.8% | 3.5% | 0.008 |
| 20 | 8.5% | 3.4% | 0.008 |
| 21 | 7.9% | 3.5% | 0.007 |
| 22 | 7.5% | 3.3% | 0.007 |
| 23 | 6.7% | 3.1% | 0.007 |
| 24 | 8.3% | 2.8% | 0.006 |
| 25 | 7.2% | 2.6% | 0.007 |
| 26 | 11.5% | 2.3% | 0.007 |
| 27 | 13.1% | 2.2% | 0.010 |
| 28 | 17.5% | 2.6% | 0.012 |
| 29 | 16.2% | 2.8% | 0.012 |
| 30 | 10.9% | 3.3% | 0.008 |
## B. state-conditional information under the OUTCOME label
held-out per-candidate log-loss (bits) of the outcome label, state-conditional vs state-free (same rows, same split):
| model | log-loss (bits) |
|---|---:|
| state-free P(hit | g) | 0.1873 |
| state-conditional P(hit | state, g) | 0.3906 |
| Δ (state − state-free) | +0.2033 |
state conditioning is better in 0/3 splits (negative Δ = better).
## C. open-loop decision counterfactual (VETO ONLY)
If the mover picks argmin_g danger, the fraction of held-out waves whose bullet line would still pass within a body width of g (ground truth = the recorded bullet line b_bullet).
| policy | held-out waves still hit |
|---|---:|
| histogram argmin (j128) | 3.53% |
| outcome argmin (j130) | 3.33% |
| recorded trajectory (floor/ceiling) | 10.17% |
The counterfactual is OPEN LOOP: the recorded bullet lines were fired at a different mover, so it cannot predict the live closed loop. It is a veto, not a selection.
## MEASURED vs INFERRED
* MEASURED: every number above, on the recorded corpus.
* INFERRED: that the offline alignment transfers live. It cannot — see docs/offline_harness_trust.md.