127 lines
5.2 KiB
Plaintext
127 lines
5.2 KiB
Plaintext
# Learned-surfer Gate A — offline prediction quality (VETO ONLY)
|
|
|
|
corpus : /tmp/tfil_ab2/out
|
|
battles : 70
|
|
shots used : 54923
|
|
state fields: vlat, dist, room, turn (5 fields, ONE state, no window)
|
|
label : 31-bin guess factor at wave resolution (wave_surfer.gfToBin), mode=nominal
|
|
learner : counted SBC (saturating uint8 + `c -= c shr shift` every decayEvery learns), inferProb readout, alpha=5 prior mix
|
|
|
|
## A. held-out prediction quality (mean over 3 battle splits)
|
|
|
|
| config | states | log-loss(state) | log-loss(global) | Δ | top-1 st | top-1 glob | top-3 st | top-3 glob |
|
|
|---|---:|---:|---:|---:|---:|---:|---:|---:|
|
|
| Q4 decay128/1 (primary) | 256 | 4.9272 | 4.9739 | -0.0467 | 0.0393 | 0.0396 | 0.1224 | 0.1215 |
|
|
| Q4 decay32/1 | 256 | 4.9594 | 5.0133 | -0.0539 | 0.0370 | 0.0312 | 0.1102 | 0.0970 |
|
|
| Q4 NO DECAY | 256 | 4.8408 | 4.9221 | -0.0813 | 0.0633 | 0.0463 | 0.1572 | 0.1255 |
|
|
| Q4 decay128/2 | 256 | 4.8991 | 4.9435 | -0.0444 | 0.0464 | 0.0463 | 0.1234 | 0.1240 |
|
|
| Q3 decay128/1 | 81 | 4.9401 | 4.9739 | -0.0338 | 0.0396 | 0.0396 | 0.1244 | 0.1215 |
|
|
| Q5 decay128/1 | 625 | 4.9200 | 4.9739 | -0.0539 | 0.0402 | 0.0396 | 0.1240 | 0.1215 |
|
|
|
|
## B. floors (same held-out test sets, primary config splits)
|
|
|
|
| predictor | log-loss (bits) | top-1 | top-3 |
|
|
|---|---:|---:|---:|
|
|
| chance (uniform 31) | 4.9542 | 0.0323 | 0.0968 |
|
|
| majority bin | 4.9739 | 0.0463 | n/a |
|
|
| global 31-bin histogram (the OLD surfer) | 4.9739 | 0.0396 | 0.1215 |
|
|
| state-conditional counted SBC | 4.9272 | 0.0393 | 0.1224 |
|
|
|
|
NOTE: the 'unconditional average' and the '31-bin global histogram of the
|
|
old surfer' are the SAME estimator by construction (both are the train
|
|
marginal over bins); they are therefore reported as one row. The majority
|
|
predictor is the degenerate top-1 version of the same marginal.
|
|
|
|
## C. primary config per split (recurrence + paired per-battle stats)
|
|
|
|
| seed | train shots | test shots | distinct states | mean count/state | test shots with a SEEN state | Δlog-loss (state-global) |
|
|
|---|---:|---:|---:|---:|---:|---:|
|
|
| 0 | 38172 | 16751 | 256 | 149.1 | 100.0% | -0.0454 |
|
|
| 1 | 38312 | 16611 | 256 | 149.7 | 100.0% | -0.0465 |
|
|
| 2 | 38373 | 16550 | 256 | 149.9 | 100.0% | -0.0482 |
|
|
|
|
### paired per-battle statistics (primary config, all splits pooled)
|
|
|
|
| metric | n battles | mean Δ | SD | 95% CI | sign | p(sign) | p(sign-flip) | MDE |
|
|
|---|---:|---:|---:|---|---:|---:|---:|---:|
|
|
| logloss(state-global), pooled | 63 | -0.0467 | 0.0058 | [-0.0481, -0.0453] | 0/63 | 2.168e-19 | 5e-05 | 0.0021 |
|
|
| logloss(state-global) seed0 | 21 | -0.0454 | 0.0048 | [-0.0475, -0.0433] | 0/21 | 9.537e-07 | 5e-05 | 0.0029 |
|
|
| logloss(state-global) seed1 | 21 | -0.0465 | 0.0060 | [-0.0490, -0.0439] | 0/21 | 9.537e-07 | 5e-05 | 0.0037 |
|
|
| logloss(state-global) seed2 | 21 | -0.0482 | 0.0065 | [-0.0510, -0.0454] | 0/21 | 9.537e-07 | 5e-05 | 0.0040 |
|
|
(negative Δ = the state-conditional model predicts better)
|
|
|
|
## D. label-shuffle control (same states, train labels permuted)
|
|
|
|
| arm | log-loss(state) | log-loss(global) | Δ | top-1(state) |
|
|
|---|---:|---:|---:|---:|
|
|
| shuffled labels | 4.9480 | 4.9535 | -0.0056 | 0.0384 |
|
|
| real labels (Q4 decay128) | 4.9272 | 4.9739 | -0.0467 | 0.0393 |
|
|
|
|
(chance log-loss floor = 4.9542 bits, target entropy = the
|
|
global-histogram log-loss above; a shuffled-label state model must fall
|
|
back to it.)
|
|
|
|
## E. canonical state-bin edges (4 symbols/field), hard-coded into the
|
|
module `common_libs/movements/learned_surfer.nim` (derived from the whole
|
|
corpus; the gate numbers above use TRAIN-only edges per split, so the
|
|
report is not conditioned on these)
|
|
|
|
```nim
|
|
vlat : -6.736, 0.000, 6.753
|
|
dist : 431.321, 487.612, 552.670
|
|
room : 137.965, 206.589, 296.753
|
|
turn : -0.142, 0.000, 0.105
|
|
```
|
|
|
|
cross-check with the canonical edges: log-loss(state) 4.9272, log-loss(global) 4.9739, Δ -0.0467, top-1 0.0393
|
|
|
|
## F. is the DANGER MAP the module minimises the RIGHT one? (MEASURED)
|
|
|
|
shots 54936, base hit rate 9.98%
|
|
corr( P(arrival bin) , P(hit | arrival bin) ) over the 31 bins = -0.342
|
|
safest bin by the histogram MASS the mover minimises: bin 1 (mass 1.9%, hit rate 14.1%)
|
|
safest bin by the ACTUAL hit rate: bin 23 (mass 3.1%, hit rate 6.8%)
|
|
|
|
| bin | P(hit) | P(arrival bin) |
|
|
|---:|---:|---:|
|
|
| 0 | 9.8% | 2.4% |
|
|
| 1 | 14.1% | 1.9% |
|
|
| 2 | 13.3% | 2.6% |
|
|
| 3 | 10.8% | 2.3% |
|
|
| 4 | 8.8% | 2.5% |
|
|
| 5 | 7.8% | 2.9% |
|
|
| 6 | 7.9% | 3.3% |
|
|
| 7 | 8.8% | 3.5% |
|
|
| 8 | 9.2% | 3.6% |
|
|
| 9 | 9.1% | 3.8% |
|
|
| 10 | 9.8% | 3.9% |
|
|
| 11 | 11.0% | 4.0% |
|
|
| 12 | 10.6% | 3.9% |
|
|
| 13 | 11.1% | 4.1% |
|
|
| 14 | 9.0% | 4.0% |
|
|
| 15 | 9.8% | 4.6% |
|
|
| 16 | 9.8% | 3.8% |
|
|
| 17 | 9.2% | 3.9% |
|
|
| 18 | 10.7% | 3.6% |
|
|
| 19 | 8.8% | 3.5% |
|
|
| 20 | 8.5% | 3.4% |
|
|
| 21 | 7.9% | 3.5% |
|
|
| 22 | 7.5% | 3.3% |
|
|
| 23 | 6.8% | 3.1% |
|
|
| 24 | 8.3% | 2.8% |
|
|
| 25 | 7.2% | 2.6% |
|
|
| 26 | 11.5% | 2.3% |
|
|
| 27 | 13.1% | 2.2% |
|
|
| 28 | 17.5% | 2.6% |
|
|
| 29 | 16.2% | 2.8% |
|
|
| 30 | 10.9% | 3.3% |
|
|
|
|
## MEASURED vs INFERRED
|
|
|
|
* MEASURED: every number in this file, produced by the command in the
|
|
module docstring on the recorded corpus.
|
|
* INFERRED: that this offline prediction-quality result transfers to the
|
|
LIVE closed loop. It cannot: the recorded trajectory was produced while
|
|
the enemy gun reacted to a DIFFERENT mover (see
|
|
docs/offline_harness_trust.md §0/§4). This gate is a veto only.
|