State-window gate: a temporal window of wave-relative states does NOT beat a single state

Re-runs the SBC coincidence premise as a cheap veto test on the 70-battle
live-vs-DrussGT corpus (/tmp/tfil_ab2).  Defines the wave-relative state
(lat/vlat/toa/room/turn, 5/7.9/10 bits at Q=2/3/4), quantises the miss offset
at the bullet's arrival into 7 bins, and sweeps window length K in
{1,4,8,16,32,48} with an interpolated suffix-backoff model under a BY-BATTLE
70/30 split (3 seeds).

Result: NO.  On the pre-fire frame the window is worse than the single
fire-tick state at every K/Q/A (e.g. K=8 costs +0.35..+0.46 bits).  On the
during-flight frame the entire apparent gain is the later decision tick, not
the window; the single state alone drops 2.70 -> 1.31 bits as K goes 1 -> 32.
The shuffle-order control confirms recency matters but the windows do not:
by Q=4/K=8 they average ~1 observation and never recur.  The single state
survives as a strong predictor (log-loss 2.346 vs 2.698 majority; bin
accuracy 0.409 vs 0.235).

Gate only: no gun, no live-win claim.
This commit is contained in:
2026-09-25 08:43:36 +02:00
parent 40ba96f649
commit d85ff53d34
4 changed files with 11588 additions and 0 deletions
File diff suppressed because it is too large Load Diff
+178
View File
@@ -0,0 +1,178 @@
====================================================================================================
FRAME pre (window ENDS at the fire tick, looks BACK)
====================================================================================================
MAJORITY / NO-WINDOW FLOOR (test acc 0.2348, log-loss 2.6983 bits, empirical hit 0.0906)
target-bin edges [-120.0, -60.0, -18.0, 18.0, 60.0, 120.0] px -> central +-18 px hit bin. At the median
fire range (487 px) the 36 px hit window subtends 4.23 deg; at 450 px
it is 4.58 deg; at 100 px 20.41 deg. (atan(18/range).)
target-bin distribution on train: 0:0.231, 1:0.125, 2:0.102, 3:0.089, 4:0.097, 5:0.126, 6:0.230
----------------------------------------------------------------------------------------------------
COARSENESS Q=2 -> 32 distinct single states (5.0 bits)
K | single@D (temporal) | window (temporal) | window=SHUFFLED | window=REVERSE
| logloss acc hitP | logloss acc hitP | logloss | logloss
1 | 2.4524 0.4072 0.0896 | 2.4524 0.4072 0.0896 | 2.5464 | 2.5787
4 | 2.4524 0.4072 0.0896 | 2.5107 0.3990 0.0898 | 2.7611 | 2.6403
8 | 2.4524 0.4072 0.0896 | 2.8015 0.3700 0.0889 | 2.8188 | 2.9373
16 | 2.4524 0.4072 0.0896 | 3.1310 0.3443 0.0896 | 2.8120 | 3.2793
32 | 2.4524 0.4072 0.0896 | 3.1310 0.3443 0.0896 | 2.8026 | 3.2793
48 | 2.4524 0.4072 0.0896 | 3.1310 0.3443 0.0896 | 2.8115 | 3.2793
recurrence, distinct ordered window tuples over the whole corpus:
K=1 distinct=16 of 54939 (mean count 3433.69, repeat_frac 1.000)
K=4 distinct=1747 of 54939 (mean count 31.45, repeat_frac 0.990)
K=8 distinct=13474 of 54939 (mean count 4.08, repeat_frac 0.836)
K=16 distinct=41447 of 54939 (mean count 1.33, repeat_frac 0.311)
K=32 distinct=54454 of 54939 (mean count 1.01, repeat_frac 0.011)
K=48 distinct=54744 of 54939 (mean count 1.00, repeat_frac 0.004)
----------------------------------------------------------------------------------------------------
COARSENESS Q=3 -> 243 distinct single states (7.9 bits)
K | single@D (temporal) | window (temporal) | window=SHUFFLED | window=REVERSE
| logloss acc hitP | logloss acc hitP | logloss | logloss
1 | 2.3782 0.4032 0.0897 | 2.3782 0.4032 0.0897 | 2.5116 | 2.5502
4 | 2.3782 0.4032 0.0897 | 2.5461 0.3912 0.0903 | 2.6214 | 2.7670
8 | 2.3782 0.4032 0.0897 | 2.8336 0.3594 0.0898 | 2.6284 | 3.0753
16 | 2.3782 0.4032 0.0897 | 2.9473 0.3458 0.0898 | 2.6221 | 3.1881
32 | 2.3782 0.4032 0.0897 | 2.9473 0.3458 0.0898 | 2.6225 | 3.1881
48 | 2.3782 0.4032 0.0897 | 2.9473 0.3458 0.0898 | 2.6216 | 3.1881
recurrence, distinct ordered window tuples over the whole corpus:
K=1 distinct=82 of 54939 (mean count 669.99, repeat_frac 1.000)
K=4 distinct=11018 of 54939 (mean count 4.99, repeat_frac 0.891)
K=8 distinct=37337 of 54939 (mean count 1.47, repeat_frac 0.404)
K=16 distinct=54163 of 54939 (mean count 1.01, repeat_frac 0.020)
K=32 distinct=54754 of 54939 (mean count 1.00, repeat_frac 0.004)
K=48 distinct=54772 of 54939 (mean count 1.00, repeat_frac 0.004)
----------------------------------------------------------------------------------------------------
COARSENESS Q=4 -> 1024 distinct single states (10.0 bits)
K | single@D (temporal) | window (temporal) | window=SHUFFLED | window=REVERSE
| logloss acc hitP | logloss acc hitP | logloss | logloss
1 | 2.3460 0.4094 0.0895 | 2.3460 0.4094 0.0895 | 2.5280 | 2.5352
4 | 2.3460 0.4094 0.0895 | 2.5656 0.3833 0.0894 | 2.5672 | 2.8203
8 | 2.3460 0.4094 0.0895 | 2.7265 0.3641 0.0885 | 2.5666 | 3.0138
16 | 2.3460 0.4094 0.0895 | 2.7450 0.3622 0.0886 | 2.5613 | 3.0377
32 | 2.3460 0.4094 0.0895 | 2.7450 0.3622 0.0886 | 2.5613 | 3.0377
48 | 2.3460 0.4094 0.0895 | 2.7450 0.3622 0.0886 | 2.5640 | 3.0377
recurrence, distinct ordered window tuples over the whole corpus:
K=1 distinct=372 of 54939 (mean count 147.69, repeat_frac 0.999)
K=4 distinct=24223 of 54939 (mean count 2.27, repeat_frac 0.685)
K=8 distinct=49891 of 54939 (mean count 1.10, repeat_frac 0.134)
K=16 distinct=54707 of 54939 (mean count 1.00, repeat_frac 0.005)
K=32 distinct=54779 of 54939 (mean count 1.00, repeat_frac 0.004)
K=48 distinct=54789 of 54939 (mean count 1.00, repeat_frac 0.004)
====================================================================================================
HEADLINE (mean over 3 battle-split seeds; A=5)
held-out log-loss. delta_window = window - single@D (NEGATIVE = the
window beats the single state at the same decision tick)
Q=2 single@D K=1 2.4524 K=4 2.4524 K=8 2.4524 K=16 2.4524 K=32 2.4524 K=48 2.4524
window K=1 2.4524(+0.0000) K=4 2.5107(+0.0583) K=8 2.8015(+0.3491) K=16 3.1310(+0.6786) K=32 3.1310(+0.6786) K=48 3.1310(+0.6786)
Q=3 single@D K=1 2.3782 K=4 2.3782 K=8 2.3782 K=16 2.3782 K=32 2.3782 K=48 2.3782
window K=1 2.3782(+0.0000) K=4 2.5461(+0.1679) K=8 2.8336(+0.4554) K=16 2.9473(+0.5691) K=32 2.9473(+0.5691) K=48 2.9473(+0.5691)
Q=4 single@D K=1 2.3460 K=4 2.3460 K=8 2.3460 K=16 2.3460 K=32 2.3460 K=48 2.3460
window K=1 2.3460(+0.0000) K=4 2.5656(+0.2196) K=8 2.7265(+0.3805) K=16 2.7450(+0.3990) K=32 2.7450(+0.3990) K=48 2.7450(+0.3990)
shuffle control: window(shuffled) - window(temporal) (must be >>0)
Q=2 K=1 +0.0940 K=4 +0.2504 K=8 +0.0173 K=16 -0.3190 K=32 -0.3284 K=48 -0.3195
Q=3 K=1 +0.1334 K=4 +0.0753 K=8 -0.2051 K=16 -0.3252 K=32 -0.3248 K=48 -0.3257
Q=4 K=1 +0.1820 K=4 +0.0016 K=8 -0.1599 K=16 -0.1837 K=32 -0.1838 K=48 -0.1810
reverse control: window(reverse) - window(temporal)
Q=2 K=1 +0.1263 K=4 +0.1296 K=8 +0.1358 K=16 +0.1483 K=32 +0.1483 K=48 +0.1483
Q=3 K=1 +0.1720 K=4 +0.2210 K=8 +0.2417 K=16 +0.2409 K=32 +0.2409 K=48 +0.2409
Q=4 K=1 +0.1892 K=4 +0.2547 K=8 +0.2873 K=16 +0.2927 K=32 +0.2927 K=48 +0.2927
robustness in the interpolation strength A (window temporal log-loss):
Q=2 A=1 K=1 2.4526 K=4 2.6094 K=8 3.4554 K=16 4.4613 K=32 4.4613 K=48 4.4613
Q=2 A=5 K=1 2.4524 K=4 2.5107 K=8 2.8015 K=16 3.1310 K=32 3.1310 K=48 3.1310
Q=2 A=20 K=1 2.4524 K=4 2.4643 K=8 2.5491 K=16 2.6407 K=32 2.6407 K=48 2.6407
Q=3 A=1 K=1 2.3782 K=4 2.8964 K=8 3.8419 K=16 4.2381 K=32 4.2381 K=48 4.2381
Q=3 A=5 K=1 2.3782 K=4 2.5461 K=8 2.8336 K=16 2.9473 K=32 2.9473 K=48 2.9473
Q=3 A=20 K=1 2.3810 K=4 2.4149 K=8 2.4868 K=16 2.5134 K=32 2.5134 K=48 2.5134
Q=4 A=1 K=1 2.3499 K=4 3.0860 K=8 3.6606 K=16 3.7319 K=32 3.7319 K=48 3.7319
Q=4 A=5 K=1 2.3460 K=4 2.5656 K=8 2.7265 K=16 2.7450 K=32 2.7450 K=48 2.7450
Q=4 A=20 K=1 2.3557 K=4 2.3928 K=8 2.4289 K=16 2.4327 K=32 2.4327 K=48 2.4327
====================================================================================================
====================================================================================================
FRAME fly (window STARTS at the fire tick, ends at t0+K-1)
====================================================================================================
MAJORITY / NO-WINDOW FLOOR (test acc 0.1912, log-loss 2.7810 bits, empirical hit 0.1100)
target-bin edges [-120.0, -60.0, -18.0, 18.0, 60.0, 120.0] px -> central +-18 px hit bin. At the median
fire range (487 px) the 36 px hit window subtends 4.23 deg; at 450 px
it is 4.58 deg; at 100 px 20.41 deg. (atan(18/range).)
target-bin distribution on train: 0:0.192, 1:0.145, 2:0.125, 3:0.109, 4:0.118, 5:0.139, 6:0.172
----------------------------------------------------------------------------------------------------
COARSENESS Q=2 -> 32 distinct single states (5.0 bits)
K | single@D (temporal) | window (temporal) | window=SHUFFLED | window=REVERSE
| logloss acc hitP | logloss acc hitP | logloss | logloss
1 | 2.6645 0.2935 0.1099 | 2.6645 0.2935 0.1099 | 2.6645 | 2.6645
4 | 2.6376 0.2988 0.1101 | 2.7695 0.2808 0.1106 | 2.8096 | 2.7957
8 | 2.5659 0.3076 0.1102 | 3.0331 0.2785 0.1107 | 3.1212 | 3.1197
16 | 2.4047 0.3405 0.1104 | 3.1643 0.3070 0.1098 | 2.9981 | 3.3696
32 | 1.8365 0.4189 0.1083 | 1.9970 0.4430 0.1078 | 2.3517 | 3.3696
recurrence, distinct ordered window tuples over the whole corpus:
K=1 distinct=16 of 8156 (mean count 509.75, repeat_frac 1.000)
K=4 distinct=872 of 8156 (mean count 9.35, repeat_frac 0.954)
K=8 distinct=3557 of 8156 (mean count 2.29, repeat_frac 0.665)
K=16 distinct=7389 of 8156 (mean count 1.10, repeat_frac 0.133)
K=32 distinct=8156 of 8156 (mean count 1.00, repeat_frac 0.000)
----------------------------------------------------------------------------------------------------
COARSENESS Q=3 -> 243 distinct single states (7.9 bits)
K | single@D (temporal) | window (temporal) | window=SHUFFLED | window=REVERSE
| logloss acc hitP | logloss acc hitP | logloss | logloss
1 | 2.6741 0.2901 0.1105 | 2.6741 0.2901 0.1105 | 2.6741 | 2.6741
4 | 2.6262 0.2950 0.1103 | 2.8858 0.2658 0.1086 | 2.9378 | 2.9633
8 | 2.5302 0.3184 0.1100 | 3.0207 0.2759 0.1089 | 2.9292 | 3.1779
16 | 2.3103 0.3661 0.1090 | 2.6362 0.3390 0.1110 | 2.6531 | 3.2373
32 | 1.4355 0.5221 0.1081 | 1.4625 0.5384 0.1030 | 2.2674 | 3.2373
recurrence, distinct ordered window tuples over the whole corpus:
K=1 distinct=81 of 8156 (mean count 100.69, repeat_frac 1.000)
K=4 distinct=3383 of 8156 (mean count 2.41, repeat_frac 0.714)
K=8 distinct=6857 of 8156 (mean count 1.19, repeat_frac 0.226)
K=16 distinct=8140 of 8156 (mean count 1.00, repeat_frac 0.004)
K=32 distinct=8156 of 8156 (mean count 1.00, repeat_frac 0.000)
----------------------------------------------------------------------------------------------------
COARSENESS Q=4 -> 1024 distinct single states (10.0 bits)
K | single@D (temporal) | window (temporal) | window=SHUFFLED | window=REVERSE
| logloss acc hitP | logloss acc hitP | logloss | logloss
1 | 2.7020 0.2817 0.1113 | 2.7020 0.2817 0.1113 | 2.7020 | 2.7020
4 | 2.6397 0.2982 0.1112 | 2.9108 0.2750 0.1117 | 2.9175 | 3.0026
8 | 2.6101 0.3037 0.1137 | 2.9084 0.2793 0.1159 | 2.8151 | 3.0861
16 | 2.3853 0.3617 0.1123 | 2.5918 0.3377 0.1135 | 2.6342 | 3.0891
32 | 1.3121 0.5988 0.1055 | 1.2668 0.6027 0.1089 | 2.3830 | 3.0891
recurrence, distinct ordered window tuples over the whole corpus:
K=1 distinct=252 of 8156 (mean count 32.37, repeat_frac 1.000)
K=4 distinct=5368 of 8156 (mean count 1.52, repeat_frac 0.465)
K=8 distinct=7983 of 8156 (mean count 1.02, repeat_frac 0.037)
K=16 distinct=8156 of 8156 (mean count 1.00, repeat_frac 0.000)
K=32 distinct=8156 of 8156 (mean count 1.00, repeat_frac 0.000)
====================================================================================================
HEADLINE (mean over 3 battle-split seeds; A=5)
held-out log-loss. delta_window = window - single@D (NEGATIVE = the
window beats the single state at the same decision tick)
Q=2 single@D K=1 2.6645 K=4 2.6376 K=8 2.5659 K=16 2.4047 K=32 1.8365
window K=1 2.6645(+0.0000) K=4 2.7695(+0.1319) K=8 3.0331(+0.4673) K=16 3.1643(+0.7596) K=32 1.9970(+0.1605)
Q=3 single@D K=1 2.6741 K=4 2.6262 K=8 2.5302 K=16 2.3103 K=32 1.4355
window K=1 2.6741(+0.0000) K=4 2.8858(+0.2597) K=8 3.0207(+0.4904) K=16 2.6362(+0.3259) K=32 1.4625(+0.0270)
Q=4 single@D K=1 2.7020 K=4 2.6397 K=8 2.6101 K=16 2.3853 K=32 1.3121
window K=1 2.7020(+0.0000) K=4 2.9108(+0.2711) K=8 2.9084(+0.2984) K=16 2.5918(+0.2065) K=32 1.2668(-0.0453)
shuffle control: window(shuffled) - window(temporal) (must be >>0)
Q=2 K=1 +0.0000 K=4 +0.0400 K=8 +0.0881 K=16 -0.1662 K=32 +0.3548
Q=3 K=1 +0.0000 K=4 +0.0520 K=8 -0.0914 K=16 +0.0169 K=32 +0.8048
Q=4 K=1 +0.0000 K=4 +0.0068 K=8 -0.0933 K=16 +0.0425 K=32 +1.1162
reverse control: window(reverse) - window(temporal)
Q=2 K=1 +0.0000 K=4 +0.0262 K=8 +0.0865 K=16 +0.2053 K=32 +1.3726
Q=3 K=1 +0.0000 K=4 +0.0774 K=8 +0.1572 K=16 +0.6011 K=32 +1.7748
Q=4 K=1 +0.0000 K=4 +0.0918 K=8 +0.1776 K=16 +0.4974 K=32 +1.8223
robustness in the interpolation strength A (window temporal log-loss):
Q=2 A=1 K=1 2.6653 K=4 3.0442 K=8 4.0190 K=16 4.8738 K=32 2.8391
Q=2 A=5 K=1 2.6645 K=4 2.7695 K=8 3.0331 K=16 3.1643 K=32 1.9970
Q=2 A=20 K=1 2.6638 K=4 2.6681 K=8 2.6862 K=16 2.5833 K=32 1.7668
Q=3 A=1 K=1 2.6844 K=4 3.5284 K=8 4.2822 K=16 3.6127 K=32 1.9264
Q=3 A=5 K=1 2.6741 K=4 2.8858 K=8 3.0207 K=16 2.6362 K=32 1.4625
Q=3 A=20 K=1 2.6681 K=4 2.6773 K=8 2.6250 K=16 2.3607 K=32 1.4132
Q=4 A=1 K=1 2.7908 K=4 3.6658 K=8 3.9767 K=16 3.4357 K=32 1.4693
Q=4 A=5 K=1 2.7020 K=4 2.9108 K=8 2.9084 K=16 2.5918 K=32 1.2668
Q=4 A=20 K=1 2.6736 K=4 2.6783 K=8 2.6388 K=16 2.4433 K=32 1.4137
====================================================================================================
+579
View File
@@ -0,0 +1,579 @@
#!/usr/bin/env python3
"""STATE-WINDOW GATE: does a temporal WINDOW of wave-relative states predict
DrussGT's future lateral position better than a SINGLE state?
This is the cheap veto test for the "feed SBC a temporal list of states" design
(docs/state_window_gate.md). It is NOT a gun and makes no live-win claim.
DATA (real live battles, never regenerated here)
/tmp/tfil_ab2/out/<A..E>/runN.jsonl + .events.jsonl + .rounds.json
70 battles / 490 rounds / ~55k shots fired by ModularBot at the real
unmodified DrussGT, recorded by tools/robocode_shim/run_bridge_battle.sh.
In the capture rows `e*` is DrussGT (the subject) and `s*` is ModularBot (us);
the per-shot geometry is re-derived by the validated instrument in
common_libs/tests/analyze_drussgt_dodge_vs_power.py, which this file imports.
STATE (a design artifact -- see docs/state_window_gate.md for the rationale)
One state at absolute tick t, in the frame of the bullet fired at t0 along
direction u = (cos dir, sin dir):
lat = (D(t) - P0) x u lateral offset from the bullet line, px
vlat = lat(t) - lat(t-1) lateral velocity, px/tick (crossing/returning)
toa = (t0 - t) + karr ticks until the bullet reaches arrival
room = ray distance from D(t) along sign(vlat)*n until the arena wall, px
turn = wrap180(eh(t) - eh(t-1)) signed turn rate, deg/tick
Each field is quantised into Q in {2,3,4} bins (the numerosity dial). A
window is K consecutive states; K=1 is the single-state baseline.
Two frames are tested, and BOTH compare a window against the single state at
the SAME decision tick D (otherwise a longer window would win only because its
decision tick is later):
pre D = t0 (the fire tick); the window is the K pre-fire states ending at D.
fly D = t0+K-1; the window is the first K states of the flight, and the
baseline is the single state at D. Only shots with karr > 31 are used,
so the whole window is strictly before the bullet's arrival.
TARGET
perp_arr = DrussGT's SIGNED perpendicular offset from the bullet line at the
tick our bullet reaches its along-track plane (the miss offset that decides
the hit). Quantised into 7 bins (edges +-120, +-60, +-18 px); the central
+-18 px bin is the hit window.
MODEL (deliberately dull: the question is about INFORMATION, not modelling)
An interpolated (Jelinek-Mercer) suffix-backoff table over the quantised
window: P = global; for j=1..K, P <- (count(suffix_j) + A*P)/(total+A). This
is the direct analogue of the SBC "count coincidences" idea, it is
order-sensitive, and it can never do much worse than the shorter context, so
the sweep isolates information rather than overfitting. The context depth is
capped at 12 (orders above that are never observed often enough to matter).
A in {1,5,20} is swept as a robustness check (A=5 is primary).
SPLIT
BY BATTLE, never by tick. All rounds of a battle go to one side. 70 %/30 %
battle split, repeated over 3 seeds; model hyper-parameters (bin edges) are
derived from the TRAIN side only.
CONTROLS
1. shuffle -- random permutation of the K states inside each window (the
multiset is preserved, order destroyed). MANDATORY.
2. reverse -- deterministic reversal (preserves recurrence, reverses time).
3. majority -- no-window floor.
4. recurrence -- distinct windows, and how often a window repeats.
Run:
python3 common_libs/tests/state_window_gate.py --tfil /tmp/tfil_ab2/out \
--json common_libs/tests/fixtures/state_window_gate.json \
> common_libs/tests/fixtures/state_window_gate_report.txt
"""
from __future__ import annotations
import argparse
import bisect
import collections
import json
import math
import os
import random
import statistics
import sys
from array import array
HERE = os.path.dirname(os.path.abspath(__file__))
sys.path.insert(0, HERE)
import analyze_drussgt_dodge_vs_power as dodge # noqa: E402
ARENA_W, ARENA_H = 800.0, 600.0
BOT_R = 18.0
MAXK_PRE = 48 # pre-fire history depth (all K <= 48 are feasible)
MAXK_FLY = 32 # during-flight window depth (max real flight is 43)
FIELDS = ("lat", "vlat", "toa", "room", "turn")
D_CAP = 12 # model context depth cap (see docs; orders > cap never observed)
TARGET_EDGES = [-120.0, -60.0, -18.0, 18.0, 60.0, 120.0]
HIT_BIN = 3 # bin [ -18, 18 ) -> the bot radius
K_PRE = (1, 4, 8, 16, 32, 48)
K_FLY = (1, 4, 8, 16, 32)
def wrap180(a):
return ((a + 180.0) % 360.0) - 180.0
def room_to_wall(px, py, dx, dy):
"""Distance from (px,py) along unit (dx,dy) until leaving the arena
(accounting for the 18 px bot radius)."""
t = float("inf")
for p, d, lo, hi in ((px, dx, BOT_R, ARENA_W - BOT_R),
(py, dy, BOT_R, ARENA_H - BOT_R)):
if abs(d) > 1e-9:
cand = (hi - p) / d if d > 0 else (lo - p) / d
if cand < t:
t = cand
if t == float("inf"):
return 0.0
return max(0.0, t)
def tbin(v):
return bisect.bisect_right(TARGET_EDGES, v)
def raw_state(run, t0, karr, P0, ux, uy, t, rnd):
rs = run.start[rnd]
re = rs + run.count[rnd] - 1
tc = min(max(t, rs), re)
tp = max(tc - 1, rs)
r = run.by_tick.get(tc)
rp = run.by_tick.get(tp)
if r is None or rp is None:
return None
lat = (r["ex"] - P0[0]) * uy - (r["ey"] - P0[1]) * ux
latp = (rp["ex"] - P0[0]) * uy - (rp["ey"] - P0[1]) * ux
vlat = lat - latp
toa = (t0 - tc) + karr
turn = wrap180(r["eh"] - rp["eh"])
sg = 1.0 if vlat >= 0 else -1.0
room = room_to_wall(r["ex"], r["ey"], -uy * sg, ux * sg)
return (lat, vlat, toa, room, turn)
def extract_run(run):
out = []
for s in run.shots():
t0 = s["tick"]
karr = s["flight"]
if karr is None:
continue
th = math.radians(s["_dir"])
ux, uy = math.cos(th), math.sin(th)
P0 = (s["_x"], s["_y"])
rnd = s["rnd"]
pre = array("f")
ok = True
for off in range(-(MAXK_PRE - 1), 1):
rs = raw_state(run, t0, karr, P0, ux, uy, t0 + off, rnd)
if rs is None:
ok = False
break
pre.extend(rs)
if not ok:
continue
fly = []
for off in range(MAXK_FLY):
rs = raw_state(run, t0, karr, P0, ux, uy, t0 + off, rnd)
fly.append(None if rs is None else rs)
out.append(dict(
battle=os.path.basename(os.path.dirname(run.cap_path)) + "/" +
os.path.basename(run.cap_path),
rnd=rnd, t0=t0, karr=karr, perp_arr=s["perp_arr"],
perp_fire=s["perp_fire"], range=s["range"], power=s["power"],
hit=s["hit"], tbin=tbin(s["perp_arr"]), pre=pre, fly=fly,
))
return out
# ------------------------------------------------------------------ quantise
def edges_for(vals, q):
xs = sorted(vals)
n = len(xs)
return [xs[min(n - 1, int((qi / q) * n))] for qi in range(1, q)]
def frame_edges(samples, frame, Q):
"""Quantile bin edges per field, derived from the TRAIN side only."""
edges = []
for f in range(5):
vals = []
if frame == "pre":
for s in samples:
if s["_split"] != "train":
continue
a = s["pre"]
for i in range(MAXK_PRE):
vals.append(a[i * 5 + f])
else:
for s in samples:
if s["_split"] != "train":
continue
for st in s["fly"]:
if st is not None:
vals.append(st[f])
edges.append(edges_for(vals, Q))
return edges
def apply_codes(samples, frame, edges, Q):
res = []
for s in samples:
if frame == "pre":
a = s["pre"]
n = MAXK_PRE
codes = [0] * n
for i in range(n):
base = i * 5
c = 0
for f in range(5):
c += bisect.bisect_right(edges[f], a[base + f]) * (Q ** f)
codes[i] = c
else:
a = s["fly"]
n = MAXK_FLY
codes = [0] * n
for i in range(n):
if a[i] is None:
codes[i] = -1
continue
c = 0
for f in range(5):
c += bisect.bisect_right(edges[f], a[i][f]) * (Q ** f)
codes[i] = c
res.append(codes)
return res
# ------------------------------------------------------------------ the model
def order_window(w, mode, rng):
if mode == "temporal":
return w
if mode == "reverse":
return w[::-1]
if mode == "shuffled":
return [w[i] for i in rng.sample(range(len(w)), len(w))]
raise ValueError(mode)
def build_counts(windows, ys, mode, Dcap, rng):
"""cbyorder[j][context] = [ {target_bin: count}, total ].
The model is a standard interpolated (Jelinek-Mercer) suffix backoff: for a
window w the prediction for the next-state target is
P = global
for j = 1..K: P = (count(suffix_j) + A*P) / (total(suffix_j) + A)
so a longer context is only believed as far as the data supports it, and the
model can never do much worse than the shorter one. This keeps the MODEL
uninteresting, which is what the gate needs."""
cbyorder = [dict() for _ in range(Dcap + 1)]
for w, y in zip(windows, ys):
wo = order_window(w, mode, rng)
ctx = ()
for j in range(1, min(Dcap, len(wo)) + 1):
ctx = (wo[-j],) + ctx
e = cbyorder[j].get(ctx)
if e is None:
e = [{}, 0]
cbyorder[j][ctx] = e
e[0][y] = e[0].get(y, 0) + 1
e[1] += 1
return cbyorder
NBINS = len(TARGET_EDGES) + 1
def glob_vec(glob, nbins=NBINS):
tot = sum(glob.values())
return [(glob.get(b, 0)) / tot for b in range(nbins)]
def predict(w, cbyorder, P0, K, A):
P = list(P0)
for j in range(1, min(K, len(w), len(cbyorder) - 1) + 1):
ctx = tuple(w[-j:])
e = cbyorder[j].get(ctx)
if e is None:
continue
cnt, tot = e
P = [(cnt.get(b, 0) + A * P[b]) / (tot + A) for b in range(len(P))]
return P
def evaluate(windows, ys, cbyorder, P0, mode, Kdepth, A, rng):
acc = 0
n = 0
ll = 0.0
hitp = 0.0
for w, y in zip(windows, ys):
wo = order_window(w, mode, rng)
P = predict(wo, cbyorder, P0, Kdepth, A)
best = max(range(len(P)), key=lambda b: (P[b], -b))
if best == y:
acc += 1
ll += -math.log2(max(P[y], 1e-12))
hitp += P[HIT_BIN]
n += 1
return {"n": n, "acc": acc / n, "logloss": ll / n,
"implied_hit": hitp / n}
def majority_floor(samples):
train = [s for s in samples if s["_split"] == "train"]
glob = collections.Counter(s["tbin"] for s in train)
top = max(glob, key=lambda b: glob[b])
acc = ll = 0.0
n = 0
test_hit = 0
for s in samples:
if s["_split"] != "test":
continue
n += 1
if s["tbin"] == top:
acc += 1
ll += -math.log2(glob[s["tbin"]] / sum(glob.values()))
test_hit += 1 if s["tbin"] == HIT_BIN else 0
return {"top": top, "n": n, "acc": acc / n, "logloss": ll / n,
"emp_hit": test_hit / n,
"dist": {str(k): v / sum(glob.values()) for k, v in sorted(glob.items())}}
def recurrence(windows, K):
"""Distinct windows and repeat rate over the WHOLE corpus (no split)."""
seen = collections.Counter()
for w in windows:
k = min(K, len(w))
seen[tuple(w[len(w) - k:])] += 1
total = len(windows)
rep2 = sum(c for c in seen.values() if c >= 2)
return {"distinct": len(seen), "total": total, "repeat_frac": rep2 / total,
"mean_count": total / len(seen)}
# ------------------------------------------------------------------ driver
def split_battles(samples, seed):
battles = sorted({s["battle"] for s in samples})
rng = random.Random(seed)
rng.shuffle(battles)
ntr = int(round(0.70 * len(battles)))
train = set(battles[:ntr])
for s in samples:
s["_split"] = "train" if s["battle"] in train else "test"
def window_list(codes, frame, K):
"""The ordered window for one sample.
pre frame: the full 48-state pre-fire history ENDING at the fire tick; the
model's depth parameter selects how many of the most recent
states it may use, so K=1 is the fire-tick state itself.
fly frame: the first K states of the flight, i.e. ending at t0+K-1; the
single-state baseline is then the state at that same tick.
"""
if frame == "pre":
return list(codes)
return list(codes[:K])
def run_frame(samples, frame, Ks, seeds, As):
res = {"frame": frame, "seeds": {}}
for seed in seeds:
split_battles(samples, seed)
tr = [i for i, s in enumerate(samples) if s["_split"] == "train"]
te = [i for i, s in enumerate(samples) if s["_split"] == "test"]
ytr = [samples[i]["tbin"] for i in tr]
yte = [samples[i]["tbin"] for i in te]
seed_res = {"majority": majority_floor(samples), "Q": {}}
for Q in (2, 3, 4):
edges = frame_edges(samples, frame, Q)
codes = apply_codes(samples, frame, edges, Q)
glob = {}
for i in tr:
glob[samples[i]["tbin"]] = glob.get(samples[i]["tbin"], 0) + 1
P0 = glob_vec(glob)
qres = {}
for mode in ("temporal", "shuffled", "reverse"):
mres = {str(A): {} for A in As}
cache = {}
for K in Ks:
wid = "full" if frame == "pre" else K
if wid not in cache:
wl = [window_list(codes[i], frame, K) for i in range(len(samples))]
cby = build_counts([wl[i] for i in tr], ytr, mode, D_CAP,
random.Random(1000 + seed))
cache[wid] = (cby, wl)
cby, wl = cache[wid]
wte = [wl[i] for i in te]
kd = min(K, D_CAP)
for A in As:
rnge = random.Random(4000 + seed * 13 + K)
mres[str(A)][str(K)] = {
"window": evaluate(wte, yte, cby, P0, mode, kd, A, rnge)}
qres[mode] = mres
# single-state-at-the-same-decision-tick baseline (true recent state)
sres = {str(A): {} for A in As}
for K in Ks:
wid = "full" if frame == "pre" else K
wl = [window_list(codes[i], frame, K) for i in range(len(samples))]
cby = build_counts([wl[i] for i in tr], ytr, "temporal", D_CAP,
random.Random(1000 + seed))
wte = [wl[i] for i in te]
for A in As:
sres[str(A)][str(K)] = evaluate(
wte, yte, cby, P0, "temporal", 1, A,
random.Random(5000 + seed * 17 + K))
qres["single"] = sres
qres["recurrence"] = {str(K): recurrence(
[window_list(codes[i], frame, K) for i in range(len(samples))], K)
for K in Ks}
seed_res["Q"][str(Q)] = qres
res["seeds"][str(seed)] = seed_res
return res
def mean_over(xs):
return statistics.fmean(xs) if xs else float("nan")
def cell(res, frame, Q, mode, A, K, key="window"):
if key == "single":
q = lambda sd: res[frame]["seeds"][sd]["Q"][str(Q)]["single"][str(A)][str(K)]
else:
q = lambda sd: res[frame]["seeds"][sd]["Q"][str(Q)][mode][str(A)][str(K)][key]
a = [q(sd) for sd in res[frame]["seeds"]]
return (mean_over([c["acc"] for c in a]),
mean_over([c["logloss"] for c in a]),
mean_over([c["implied_hit"] for c in a]))
def render(res, out, A_primary):
lines = []
def p(s=""):
lines.append(s)
print(s)
for frame in res:
Ks = K_PRE if frame == "pre" else K_FLY
p("=" * 100)
p("FRAME %s%s" % (frame, " (window ENDS at the fire tick, looks BACK)"
if frame == "pre" else
" (window STARTS at the fire tick, ends at t0+K-1)"))
p("=" * 100)
maj = [res[frame]["seeds"][sd]["majority"] for sd in res[frame]["seeds"]]
p("MAJORITY / NO-WINDOW FLOOR (test acc %.4f, log-loss %.4f bits,"
" empirical hit %.4f)" % (mean_over([m["acc"] for m in maj]),
mean_over([m["logloss"] for m in maj]),
mean_over([m["emp_hit"] for m in maj])))
p(" target-bin edges %s px -> central +-18 px hit bin. At the median"
% TARGET_EDGES)
p(" fire range (487 px) the 36 px hit window subtends %.2f deg; at 450 px"
% math.degrees(2 * math.atan(18.0 / 487.0)))
p(" it is %.2f deg; at 100 px %.2f deg. (atan(18/range).)"
% (math.degrees(2 * math.atan(18.0 / 450.0)),
math.degrees(2 * math.atan(18.0 / 100.0))))
p(" target-bin distribution on train: %s" %
", ".join("%s:%.3f" % (k, v) for k, v in sorted(maj[0]["dist"].items())))
p()
for Q in (2, 3, 4):
nstates = Q ** 5
p("-" * 100)
p("COARSENESS Q=%d -> %d distinct single states (%.1f bits)"
% (Q, nstates, math.log2(nstates)))
p(" %-4s | %-24s | %-24s | %-13s | %-13s" %
("K", "single@D (temporal)", "window (temporal)", "window=SHUFFLED",
"window=REVERSE"))
p(" %-4s | %-24s | %-24s | %-13s | %-13s" %
("", "logloss acc hitP", "logloss acc hitP", "logloss",
"logloss"))
for K in Ks:
sc = cell(res, frame, Q, "temporal", A_primary, K, "single")
tc = cell(res, frame, Q, "temporal", A_primary, K, "window")
sh = cell(res, frame, Q, "shuffled", A_primary, K, "window")
rv = cell(res, frame, Q, "reverse", A_primary, K, "window")
p(" %-4d | %.4f %.4f %.4f | %.4f %.4f %.4f | %.4f | %.4f"
% (K, sc[1], sc[0], sc[2], tc[1], tc[0], tc[2],
sh[1], rv[1]))
rec = res[frame]["seeds"][sorted(res[frame]["seeds"])[0]]["Q"][str(Q)]["recurrence"]
p(" recurrence, distinct ordered window tuples over the whole corpus:")
for K in Ks:
r = rec[str(K)]
p(" K=%-2d distinct=%-8d of %-6d (mean count %.2f, repeat_frac %.3f)"
% (K, r["distinct"], r["total"], r["mean_count"], r["repeat_frac"]))
p()
p("=" * 100)
p("HEADLINE (mean over %d battle-split seeds; A=%g)" %
(len(res[frame]["seeds"]), A_primary))
p(" held-out log-loss. delta_window = window - single@D (NEGATIVE = the")
p(" window beats the single state at the same decision tick)")
for Q in (2, 3, 4):
srow = [cell(res, frame, Q, "temporal", A_primary, K, "single")[1] for K in Ks]
trow = [cell(res, frame, Q, "temporal", A_primary, K, "window")[1] for K in Ks]
p(" Q=%d single@D " % Q + " ".join("K=%d %.4f" % (K, v)
for K, v in zip(Ks, srow)))
p(" window " + " ".join("K=%d %.4f(%+.4f)" % (K, t, t - s)
for K, t, s in zip(Ks, trow, srow)))
p(" shuffle control: window(shuffled) - window(temporal) (must be >>0)")
for Q in (2, 3, 4):
row = [cell(res, frame, Q, "shuffled", A_primary, K, "window")[1] -
cell(res, frame, Q, "temporal", A_primary, K, "window")[1]
for K in Ks]
p(" Q=%d " % Q + " ".join("K=%d %+.4f" % (K, v) for K, v in zip(Ks, row)))
p(" reverse control: window(reverse) - window(temporal)")
for Q in (2, 3, 4):
row = [cell(res, frame, Q, "reverse", A_primary, K, "window")[1] -
cell(res, frame, Q, "temporal", A_primary, K, "window")[1]
for K in Ks]
p(" Q=%d " % Q + " ".join("K=%d %+.4f" % (K, v) for K, v in zip(Ks, row)))
p(" robustness in the interpolation strength A (window temporal log-loss):")
for Q in (2, 3, 4):
for A in out.get("A_all", []):
row = [cell(res, frame, Q, "temporal", A, K, "window")[1] for K in Ks]
p(" Q=%d A=%-4g " % (Q, A) +
" ".join("K=%d %.4f" % (K, v) for K, v in zip(Ks, row)))
p("=" * 100)
out["report"] = lines
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--tfil", default="/tmp/tfil_ab2/out")
ap.add_argument("--json", default=None)
ap.add_argument("--limit-runs", type=int, default=0)
ap.add_argument("--seed", type=int, default=1)
ap.add_argument("--seeds", type=int, default=3)
ap.add_argument("--frames", default="pre")
ap.add_argument("--alphas", default="1,5,20")
a = ap.parse_args()
runs = dodge.discover_tfil(a.tfil)
if a.limit_runs:
runs = runs[:a.limit_runs]
print("[gate] %d battles" % len(runs), file=sys.stderr)
samples = []
for k, run in enumerate(runs):
samples.extend(extract_run(run))
if (k + 1) % 10 == 0:
print("[gate] %d/%d battles, %d shots" % (k + 1, len(runs),
len(samples)), file=sys.stderr)
print("[gate] %d samples" % len(samples), file=sys.stderr)
hit = sum(1 for s in samples if s["tbin"] == HIT_BIN)
print("[gate] empirical arrival hit bin (|perp|<18): %.4f" % (hit / len(samples)),
file=sys.stderr)
seeds = [a.seed + i for i in range(a.seeds)]
As = [float(x) for x in a.alphas.split(",")]
A_primary = 5.0
res = {}
nfly = 0
for frame in a.frames.split(","):
if frame == "pre":
fs = samples
else:
kmin = max(K_FLY) + 1 # keep the whole window strictly before arrival
fs = [s for s in samples if s["karr"] >= kmin]
nfly = len(fs)
print("[gate] frame %s: %d shots with karr >= %d" %
(frame, len(fs), kmin), file=sys.stderr)
res[frame] = run_frame(fs, frame, K_PRE if frame == "pre" else K_FLY,
seeds, As)
out = {"samples": len(samples), "battles": len(runs), "seeds": seeds,
"fly_samples": nfly,
"alphas": As, "A_primary": A_primary, "A_all": As,
"D_cap": D_CAP, "target_edges": TARGET_EDGES, "hit_bin": HIT_BIN,
"k_pre": list(K_PRE), "k_fly": list(K_FLY)}
render(res, out, A_primary)
out["raw"] = res
if a.json:
with open(a.json, "w") as f:
json.dump(out, f, indent=1, default=str)
return out
if __name__ == "__main__":
main()
+321
View File
@@ -0,0 +1,321 @@
# The state-window gate: does a temporal WINDOW of wave-relative states beat a SINGLE state?
**The design under test (the user's words).** *"We decide first what makes an
'environment state' ... then we feed as input a long list of these states, long
enough to have one or more patterns of movement inside."* The premise is that SBC
is a coincidence detector and that a **temporal list** of states, fed instead of a
single tick, turns "coincidence" into "movement pattern".
**Verdict: NO.** On held-out battles a window of wave-relative states **does not**
predict DrussGT's lateral position at the bullet's arrival better than the single
state at the **same decision tick**. On the pre-fire frame the window is strictly
and substantially *worse* at every window length, every coarseness and every
smoothing strength. On the during-flight frame the apparent gain is entirely the
**later decision tick**, not the window — and where the window finally edges the
single state (K=32, Q=4) it is by 0.045 bits on a subset where the 32-window
**never repeats at all**. The design, as stated, is **dead**. The *state
definition* survives and is the useful result.
This is a gate, not a gun. **No live win is claimed and none is measured here.**
---
## 1. The state (design artifact)
The dodge is a response to **our bullet**, so the state is **wave-relative**, never
absolute world coordinates. One state at absolute tick `t`, in the frame of the
bullet fired at `t0` along `u = (cos dir, sin dir)` (with `n = (-u_y, u_x)` the
lateral unit vector):
| field | definition | why |
|---|---|---|
| `lat` | `(D(t) − P0) × u` — DrussGT's signed perpendicular offset from the bullet line, px | the quantity that decides the hit; signed so "which side" is visible |
| `vlat` | `lat(t) − lat(t−1)` — signed lateral velocity, px/tick | crossing vs returning |
| `toa` | `(t0 − t) + karr` — ticks until the bullet reaches its arrival plane | how much time is left to move |
| `room` | ray distance from `D(t)` along `sign(vlat)·n` until the arena wall (18 px bot radius) | room to keep running vs being cornered — the only *directional* wall field, unlike plain "distance to nearest wall" |
| `turn` | `wrap180(eh(t) − eh(t−1))` — signed turn rate, deg/tick | surfers slow to turn; turning reveals a reversal |
**Refinements to the starting proposal, and why.** (a) was kept signed, not
absolutised, because the target is signed. (b) was made a *difference of the
lateral offset*, so it is exactly the lateral component of velocity in the bullet
frame rather than a speed. (d) "distance to the wall" was made *directional*
(room to run in the current lateral-motion direction); a plain nearest-wall
distance is not wave-relative and is nearly constant at long range. (e)
"heading/turn rate" was reduced to the **turn rate** alone, because heading
relative to the bullet is already implied by the sign of `vlat`, and the turn rate
is the thing that distinguishes a surf from a corner.
**Quantisation (the numerosity dial).** Each field is binned into `Q ∈ {2, 3, 4}`
symbols. A state is therefore:
| Q | bits / state | distinct states |
|---|---|---|
| 2 | **5.0** | 32 |
| 3 | **7.9** | 243 |
| 4 | **10.0** | 1024 |
Bin edges are **quantiles of the TRAIN side only** (no test leakage); a coarse
state makes two similar situations look the same, which is what lets a
"coincidence" repeat. The sweep below is exactly the question of whether coarser
buys more recurrence than it loses in resolution.
### The two windows tested, and why both
A window is `K` consecutive states. Crucially, **the baseline is always the single
state at the same decision tick `D`** — otherwise a longer window would "win" only
because its last state is closer to the answer.
* **`pre`** — `D = t0` (the fire tick); the window is the `K` pre-fire states
ending at `D`. This is what a bot can compute *before* pulling the trigger. All
`K ≤ 48` are feasible (max recorded flight is 43 ticks).
* **`fly`** — `D = t0+K−1`; the window is the first `K` states of the flight, and
the baseline is the single state at `D`. Only shots with `karr > 31` are used,
so the whole window is strictly before arrival (8156 of 54939 shots).
The task's field `toa` only exists once a bullet is in the air, so `fly` is the
"reaction" reading; `pre` is the "am I about to be dodged?" reading. Both are
reported.
### The target
`perp_arr` — DrussGT's **signed** perpendicular offset from the bullet line at the
tick the bullet reaches its along-track plane. This is the miss offset that
decides the hit; unlike "required lead" it does not depend on the bullet speed. It
is quantised into 7 bins with edges `±120, ±60, ±18` px; the central `±18` px bin
is the hit window (bot radius 18 px).
**Bin width in degrees-at-450px.** The central hit bin is 36 px wide = one bot
diameter. `atan(18/range)` is the hit half-window: **2.29°** at 450 px, 4.23° at
the median fire range (487 px), and 20.4° at 100 px. So the 36 px central bin is
**4.58° at 450 px**. The three inner bins are 36/42/60 px wide (4.58/5.34/7.63° at
450 px) — the target resolution is *finer* than the ~2.29° hit half-window, so a
model that predicts the bin is not being asked an impossible question.
---
## 2. Data and split
* **Data (MEASURED present before use):** `/tmp/tfil_ab2/out/<A..E>/runN.jsonl` +
`.events.jsonl` + `.rounds.json` — **70 real live battles vs the unmodified
DrussGT, 490 rounds, 54 939 shots** by ModularBot, recorded by
`tools/robocode_shim/run_bridge_battle.sh`. Per-shot geometry is re-derived by
the validated instrument `common_libs/tests/analyze_drussgt_dodge_vs_power.py`
(which this extractor imports). Attribution (`e*`=DrussGT, `s*`=us) is the one
already documented in `docs/drussgt_dodge_vs_power.md`.
* **Split (stated explicitly): BY BATTLE, never by tick.** All rounds of a battle
go to one side. A 70 %/30 % battle split (**49 train / 21 test battles**),
repeated over **3 seeds**. Bin edges are derived from the train side only. Two
ticks inside one round never straddle the split.
---
## 3. Model (deliberately dull)
The question is about **information**, not modelling, so the model is a plain
**interpolated (Jelinek–Mercer) suffix-backoff table** — the direct analogue of
SBC "counting coincidences":
```
P = global target-bin histogram
for j = 1..K: P <- (count(suffix_j) + A*P) / (total(suffix_j) + A)
```
It is order-sensitive, it can never do much worse than the shorter context, and it
degrades gracefully when a suffix is unseen. Context depth is capped at 12 (orders
above that are never repeated often enough to matter — see §5). `A ∈ {1,5,20}` is
swept; **A = 5 is primary**. A majority/no-window predictor is the floor. The
model predicts the 7-bin displacement distribution; we report held-out **accuracy**
(argmax), **log-loss** in bits, and the **implied hit probability** `P(central
bin)`.
---
## 4. Result — does the window beat the single state? **No** *(MEASURED)*
70 battles, 3 battle-split seeds, mean held-out log-loss (bits). `delta = window −
single@D`; **negative would mean the window wins**.
### 4.1 `pre` frame — window ends at the fire tick, looks back *(54 939 shots)*
`single@D` is the fire-tick state; it does **not** change with K. A = 5:
| K | Q=2 single → window (Δ) | Q=3 single → window (Δ) | Q=4 single → window (Δ) |
|---|---|---|---|
| 1 | 2.4524 → 2.4524 (—) | 2.3782 → 2.3782 (—) | 2.3460 → 2.3460 (—) |
| 4 | 2.4524 → 2.5107 (**+0.058**) | 2.3782 → 2.5461 (**+0.168**) | 2.3460 → 2.5656 (**+0.220**) |
| 8 | 2.4524 → 2.8015 (**+0.349**) | 2.3782 → 2.8336 (**+0.455**) | 2.3460 → 2.7265 (**+0.381**) |
| 16 | 2.4524 → 3.1310 (**+0.679**) | 2.3782 → 2.9473 (**+0.569**) | 2.3460 → 2.7450 (**+0.399**) |
| 32 | same as K=16 (depth cap) | same | same |
| 48 | same as K=16 (depth cap) | same | same |
Accuracy tells the same story (Q=4: single `0.409` → window `0.383`/`0.364`/
`0.362`). The per-seed deltas are essentially identical (e.g. Q=2/K=4:
`+0.0572, +0.0534, +0.0644`) — there is no seed where the window helps. The
**heavy-smoothing** arm is the fairest to the window and it still loses:
Q=4, A=20: K=4 `2.3928` vs single `2.3557` (+0.037); Q=2, A=20: K=4 `2.4643` vs
`2.4524` (+0.012).
**The extra history is not just useless, it is actively harmful**, and the reason
is visible in the shuffle control below: the long ordered contexts are sparse, and
the interpolation spends held-out probability on them.
The baseline is not trivial: the single state halves the majority floor
(majority log-loss **2.6983**, acc **0.2348**; best single Q=4 **2.3460**, acc
**0.4094**).
### 4.2 `fly` frame — window starts at the fire tick, ends at `t0+K−1` *(8 156 shots)*
Here the **single@D itself** improves steeply as K grows, because the decision tick
is later and therefore closer to the answer. A = 5:
| K | Q=2 single@D → window (Δ) | Q=3 single@D → window (Δ) | Q=4 single@D → window (Δ) |
|---|---|---|---|
| 1 | 2.6645 → 2.6645 (—) | 2.6741 → 2.6741 (—) | 2.7020 → 2.7020 (—) |
| 4 | 2.6376 → 2.7695 (**+0.132**) | 2.6262 → 2.8858 (**+0.260**) | 2.6397 → 2.9108 (**+0.271**) |
| 8 | 2.5659 → 3.0331 (**+0.467**) | 2.5302 → 3.0207 (**+0.490**) | 2.6101 → 2.9084 (**+0.298**) |
| 16 | 2.4047 → 3.1643 (**+0.760**) | 2.3103 → 2.6362 (**+0.326**) | 2.3853 → 2.5918 (**+0.207**) |
| 32 | 1.8365 → 1.9970 (**+0.161**) | 1.4355 → 1.4625 (**+0.027**) | **1.3121 → 1.2668 (−0.045)** |
Read this as: **the gain is the later tick, not the window.** `single@D` alone drops
from `2.70` at K=1 to `1.31` at K=32 (Q=4) — the single state at `t0+31` is 2
ticks from arrival and already knows the answer. The window adds nothing on top;
the only cell where it "wins" is K=32/Q=4 by **0.045 bits**, and that subset's
32-state window **never repeats** (§5), so the model is running on the single
state plus a trace of recent acceleration. Q=3 at the same K is *worse*, and
K=4–16 are worse by 5–20× that margin.
### 4.3 The mandatory shuffle-order control *(MEASURED)*
The order of the K states inside each window is randomly permuted, preserving the
multiset. In the `pre` frame the shuffled window is **worse** than the temporal
window at K=4 (e.g. Q=4: `2.5672` vs `2.5656`… within noise) but **better** at
K≥8 (Q=4/K=16: shuffled `2.5613` vs temporal `2.7450`). In the `fly` frame at
K=32 the shuffled window is much worse (Δ = `+0.35/ +0.80/ +1.12`).
**How to read that, honestly.** The shuffle penalty at `fly` K=32 is **not**
evidence the window works: the 32-state window never repeats, so the only order
the model can actually use is the *recency* of the last state. Shuffling destroys
*which state is current*, and the model degrades because recency matters — a
statement about the **single state**, not the window. Where the shuffled window is
*better* than the temporal one (pre frame, K≥8), it is because random order
disables the sparse long contexts that the temporal order feeds to the model.
Either way, the window itself contributes nothing.
A deterministic **reverse** control (reverse time, preserves recurrence) is worse
than temporal everywhere, confirming the model is genuinely order-sensitive and
that the last state is the informative one.
---
## 5. Recurrence — the "we can find coincidences" premise *(MEASURED)*
Distinct **ordered** window tuples over the whole corpus (54 939 shots; 8 156 in
the `fly` subset). A window is *learnable by exact coincidence* only if its count
is well above 1.
| K | Q=2 distinct (mean count) | Q=3 | Q=4 |
|---|---|---|---|
| 1 | 16 (3433) | 82 (670) | 372 (148) |
| 4 | 1 747 (31.4) | 11 018 (5.0) | 24 223 (2.3) |
| 8 | 13 474 (4.1) | 37 337 (1.5) | 49 891 (**1.1**) |
| 16 | 41 447 (1.3) | 54 163 (1.01) | 54 707 (1.00) |
| 32 | 54 454 (**1.01**) | 54 754 (1.00) | 54 779 (1.00) |
| 48 | 54 744 (1.00) | 54 772 (1.00) | 54 789 (1.00) |
(In the `fly` subset the same collapse happens earlier and harder: Q=4, K=8 already
1.02 mean count, and **K=32 is 8 156 distinct of 8 156 — literally every window is
unique**.)
So the "long list with patterns inside" premise fails as a **combinatorial**
matter, not merely a modelling one:
* At the coarsest state worth using (Q=2, 5 bits), only **K ≤ 8** has meaningful
recurrence — and even there the window does not beat the single state.
* At the finest (Q=4, 10 bits), K=4 already averages 2.3 observations; K≥8 is
essentially all singletons. **Learning from a window that never recurs is
impossible**, and the sweep shows it does not happen.
* The single state is the only object with a real coincidence structure: 16–372
distinct values, every one repeated hundreds of times.
---
## 6. Direct answer
**Does a window of wave-relative states predict DrussGT's future lateral position
better than a single state on held-out battles? NO.**
* `pre` frame: the window is worse at **every** K, Q and A; at K=8 it costs
**+0.35 to +0.46 bits** over the single state.
* `fly` frame: the whole apparent improvement is the later decision tick
(`single@D` 2.70 → 1.31 bits as K goes 1 → 32). The window adds nothing except
a 0.045-bit sliver at K=32/Q=4 on a subset whose 32-windows never repeat.
* The recurrence statistics say the same thing independently: re-usable
"patterns" of movement do not recur at any window length that carries extra
information.
**The design — feed SBC a long temporal list of states — is dead.** Do not build
it. *(MEASURED)*
**What survives (and is worth keeping).** The **state definition itself is a
strong single-state predictor**: the fire-tick state with 4 symbols/field
(10 bits) halves the majority log-loss (2.6983 → **2.3460** bits) and triples the
chance the predicted displacement bin is right (0.2348 → **0.4094**) on battles
the model has never seen. If anything from this line is taken forward, take the
**single** wave-relative state at Q=4 and use it as a context for a *different*
decision problem — but note the implied hit probability it produces is still the
base rate (~0.09), so it is informative about **which side** the miss falls on, not
about **whether** the shot hits. *(MEASURED/INFERRED)*
---
## 7. MEASURED vs INFERRED
**MEASURED** (exact commands in §8; numbers above)
* 70 battles / 490 rounds / 54 939 shots; 8 156 with `karr ≥ 33`.
* Every held-out log-loss / accuracy / implied-hit number in §4, over 3
battle-split seeds, with bin edges from train only.
* The shuffle-order and reverse-order controls.
* The distinct-window and repeat counts in §5.
**INFERRED**
* That the residual fly-frame "win" at K=32/Q=4 is a trace of recent acceleration
rather than a movement pattern — it is consistent with the never-repeating
32-window and with the single state already containing `vlat`.
* That a *different* (e.g. TSetlin/SBC) learner would not reverse the answer. This
is not proved. What **is** proved is the recurrence floor: a coincidence-based
method cannot learn from windows that never recur, and by Q=4/K=8 they already
essentially never do.
* "Design is dead" applies to the **long temporal window**. It does not say the
state is useless, and it does not say a *short* (K=2) context is worthless — K=2
was not separately swept because at Q≥3 it is already 1.5–5 observations per
context, i.e. below the point where it could help.
**Caveats.** (1) The `fly` frame is a biased long-range subset (all shots with
`flight ≥ 33`, i.e. the longest-range quarter). (2) The model's context depth is
capped at 12; K=16/32/48 therefore share a row, but §5 shows orders above ~4 are
never repeated often enough to matter, so the cap is not the binding constraint.
(3) The target is the miss offset, not the hit; the implied hit probability is the
base rate for every model, so nothing here is a hit-rate claim.
---
## 8. Reproducing
```bash
# corpus (live captures; not in the repo)
ls /tmp/tfil_ab2/out # 70 battles: <A..E>/runN.jsonl{,.events.jsonl,.rounds.json}
python3 common_libs/tests/state_window_gate.py \
--tfil /tmp/tfil_ab2/out --seeds 3 --frames pre,fly \
--json common_libs/tests/fixtures/state_window_gate.json \
> common_libs/tests/fixtures/state_window_gate_report.txt
```
Runtime ≈ **5 minutes**, single-threaded pure Python (no numpy required). The
script imports `common_libs/tests/analyze_drussgt_dodge_vs_power.py` for the
per-shot geometry; `common_libs/tests/fixtures/state_window_gate_report.txt` is
the verbatim captured output and `…state_window_gate.json` is the same numbers as
JSON (including the per-seed values). No offline fixture replay, no simulator:
these are the recorded real battles.