# Learned-surfer Gate A — offline prediction quality (VETO ONLY) corpus : /tmp/tfil_ab2/out battles : 70 shots used : 54923 state fields: vlat, dist, room, turn (5 fields, ONE state, no window) label : 31-bin guess factor at wave resolution (wave_surfer.gfToBin), mode=nominal learner : counted SBC (saturating uint8 + `c -= c shr shift` every decayEvery learns), inferProb readout, alpha=5 prior mix ## A. held-out prediction quality (mean over 3 battle splits) | config | states | log-loss(state) | log-loss(global) | Δ | top-1 st | top-1 glob | top-3 st | top-3 glob | |---|---:|---:|---:|---:|---:|---:|---:|---:| | Q4 decay128/1 (primary) | 256 | 4.9272 | 4.9739 | -0.0467 | 0.0393 | 0.0396 | 0.1224 | 0.1215 | | Q4 decay32/1 | 256 | 4.9594 | 5.0133 | -0.0539 | 0.0370 | 0.0312 | 0.1102 | 0.0970 | | Q4 NO DECAY | 256 | 4.8408 | 4.9221 | -0.0813 | 0.0633 | 0.0463 | 0.1572 | 0.1255 | | Q4 decay128/2 | 256 | 4.8991 | 4.9435 | -0.0444 | 0.0464 | 0.0463 | 0.1234 | 0.1240 | | Q3 decay128/1 | 81 | 4.9401 | 4.9739 | -0.0338 | 0.0396 | 0.0396 | 0.1244 | 0.1215 | | Q5 decay128/1 | 625 | 4.9200 | 4.9739 | -0.0539 | 0.0402 | 0.0396 | 0.1240 | 0.1215 | ## B. floors (same held-out test sets, primary config splits) | predictor | log-loss (bits) | top-1 | top-3 | |---|---:|---:|---:| | chance (uniform 31) | 4.9542 | 0.0323 | 0.0968 | | majority bin | 4.9739 | 0.0463 | n/a | | global 31-bin histogram (the OLD surfer) | 4.9739 | 0.0396 | 0.1215 | | state-conditional counted SBC | 4.9272 | 0.0393 | 0.1224 | NOTE: the 'unconditional average' and the '31-bin global histogram of the old surfer' are the SAME estimator by construction (both are the train marginal over bins); they are therefore reported as one row. The majority predictor is the degenerate top-1 version of the same marginal. ## C. primary config per split (recurrence + paired per-battle stats) | seed | train shots | test shots | distinct states | mean count/state | test shots with a SEEN state | Δlog-loss (state-global) | |---|---:|---:|---:|---:|---:|---:| | 0 | 38172 | 16751 | 256 | 149.1 | 100.0% | -0.0454 | | 1 | 38312 | 16611 | 256 | 149.7 | 100.0% | -0.0465 | | 2 | 38373 | 16550 | 256 | 149.9 | 100.0% | -0.0482 | ### paired per-battle statistics (primary config, all splits pooled) | metric | n battles | mean Δ | SD | 95% CI | sign | p(sign) | p(sign-flip) | MDE | |---|---:|---:|---:|---|---:|---:|---:|---:| | logloss(state-global), pooled | 63 | -0.0467 | 0.0058 | [-0.0481, -0.0453] | 0/63 | 2.168e-19 | 5e-05 | 0.0021 | | logloss(state-global) seed0 | 21 | -0.0454 | 0.0048 | [-0.0475, -0.0433] | 0/21 | 9.537e-07 | 5e-05 | 0.0029 | | logloss(state-global) seed1 | 21 | -0.0465 | 0.0060 | [-0.0490, -0.0439] | 0/21 | 9.537e-07 | 5e-05 | 0.0037 | | logloss(state-global) seed2 | 21 | -0.0482 | 0.0065 | [-0.0510, -0.0454] | 0/21 | 9.537e-07 | 5e-05 | 0.0040 | (negative Δ = the state-conditional model predicts better) ## D. label-shuffle control (same states, train labels permuted) | arm | log-loss(state) | log-loss(global) | Δ | top-1(state) | |---|---:|---:|---:|---:| | shuffled labels | 4.9480 | 4.9535 | -0.0056 | 0.0384 | | real labels (Q4 decay128) | 4.9272 | 4.9739 | -0.0467 | 0.0393 | (chance log-loss floor = 4.9542 bits, target entropy = the global-histogram log-loss above; a shuffled-label state model must fall back to it.) ## E. canonical state-bin edges (4 symbols/field), hard-coded into the module `common_libs/movements/learned_surfer.nim` (derived from the whole corpus; the gate numbers above use TRAIN-only edges per split, so the report is not conditioned on these) ```nim vlat : -6.736, 0.000, 6.753 dist : 431.321, 487.612, 552.670 room : 137.965, 206.589, 296.753 turn : -0.142, 0.000, 0.105 ``` cross-check with the canonical edges: log-loss(state) 4.9272, log-loss(global) 4.9739, Δ -0.0467, top-1 0.0393 ## F. is the DANGER MAP the module minimises the RIGHT one? (MEASURED) shots 54936, base hit rate 9.98% corr( P(arrival bin) , P(hit | arrival bin) ) over the 31 bins = -0.342 safest bin by the histogram MASS the mover minimises: bin 1 (mass 1.9%, hit rate 14.1%) safest bin by the ACTUAL hit rate: bin 23 (mass 3.1%, hit rate 6.8%) | bin | P(hit) | P(arrival bin) | |---:|---:|---:| | 0 | 9.8% | 2.4% | | 1 | 14.1% | 1.9% | | 2 | 13.3% | 2.6% | | 3 | 10.8% | 2.3% | | 4 | 8.8% | 2.5% | | 5 | 7.8% | 2.9% | | 6 | 7.9% | 3.3% | | 7 | 8.8% | 3.5% | | 8 | 9.2% | 3.6% | | 9 | 9.1% | 3.8% | | 10 | 9.8% | 3.9% | | 11 | 11.0% | 4.0% | | 12 | 10.6% | 3.9% | | 13 | 11.1% | 4.1% | | 14 | 9.0% | 4.0% | | 15 | 9.8% | 4.6% | | 16 | 9.8% | 3.8% | | 17 | 9.2% | 3.9% | | 18 | 10.7% | 3.6% | | 19 | 8.8% | 3.5% | | 20 | 8.5% | 3.4% | | 21 | 7.9% | 3.5% | | 22 | 7.5% | 3.3% | | 23 | 6.8% | 3.1% | | 24 | 8.3% | 2.8% | | 25 | 7.2% | 2.6% | | 26 | 11.5% | 2.3% | | 27 | 13.1% | 2.2% | | 28 | 17.5% | 2.6% | | 29 | 16.2% | 2.8% | | 30 | 10.9% | 3.3% | ## MEASURED vs INFERRED * MEASURED: every number in this file, produced by the command in the module docstring on the recorded corpus. * INFERRED: that this offline prediction-quality result transfers to the LIVE closed loop. It cannot: the recorded trajectory was produced while the enemy gun reacted to a DIFFERENT mover (see docs/offline_harness_trust.md §0/§4). This gate is a veto only.