State-window gate: a temporal window of wave-relative states does NOT beat a single state
Re-runs the SBC coincidence premise as a cheap veto test on the 70-battle
live-vs-DrussGT corpus (/tmp/tfil_ab2). Defines the wave-relative state
(lat/vlat/toa/room/turn, 5/7.9/10 bits at Q=2/3/4), quantises the miss offset
at the bullet's arrival into 7 bins, and sweeps window length K in
{1,4,8,16,32,48} with an interpolated suffix-backoff model under a BY-BATTLE
70/30 split (3 seeds).
Result: NO. On the pre-fire frame the window is worse than the single
fire-tick state at every K/Q/A (e.g. K=8 costs +0.35..+0.46 bits). On the
during-flight frame the entire apparent gain is the later decision tick, not
the window; the single state alone drops 2.70 -> 1.31 bits as K goes 1 -> 32.
The shuffle-order control confirms recency matters but the windows do not:
by Q=4/K=8 they average ~1 observation and never recur. The single state
survives as a strong predictor (log-loss 2.346 vs 2.698 majority; bin
accuracy 0.409 vs 0.235).
Gate only: no gun, no live-win claim.
This commit is contained in:
+10510
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,178 @@
|
||||
====================================================================================================
|
||||
FRAME pre (window ENDS at the fire tick, looks BACK)
|
||||
====================================================================================================
|
||||
MAJORITY / NO-WINDOW FLOOR (test acc 0.2348, log-loss 2.6983 bits, empirical hit 0.0906)
|
||||
target-bin edges [-120.0, -60.0, -18.0, 18.0, 60.0, 120.0] px -> central +-18 px hit bin. At the median
|
||||
fire range (487 px) the 36 px hit window subtends 4.23 deg; at 450 px
|
||||
it is 4.58 deg; at 100 px 20.41 deg. (atan(18/range).)
|
||||
target-bin distribution on train: 0:0.231, 1:0.125, 2:0.102, 3:0.089, 4:0.097, 5:0.126, 6:0.230
|
||||
|
||||
----------------------------------------------------------------------------------------------------
|
||||
COARSENESS Q=2 -> 32 distinct single states (5.0 bits)
|
||||
K | single@D (temporal) | window (temporal) | window=SHUFFLED | window=REVERSE
|
||||
| logloss acc hitP | logloss acc hitP | logloss | logloss
|
||||
1 | 2.4524 0.4072 0.0896 | 2.4524 0.4072 0.0896 | 2.5464 | 2.5787
|
||||
4 | 2.4524 0.4072 0.0896 | 2.5107 0.3990 0.0898 | 2.7611 | 2.6403
|
||||
8 | 2.4524 0.4072 0.0896 | 2.8015 0.3700 0.0889 | 2.8188 | 2.9373
|
||||
16 | 2.4524 0.4072 0.0896 | 3.1310 0.3443 0.0896 | 2.8120 | 3.2793
|
||||
32 | 2.4524 0.4072 0.0896 | 3.1310 0.3443 0.0896 | 2.8026 | 3.2793
|
||||
48 | 2.4524 0.4072 0.0896 | 3.1310 0.3443 0.0896 | 2.8115 | 3.2793
|
||||
recurrence, distinct ordered window tuples over the whole corpus:
|
||||
K=1 distinct=16 of 54939 (mean count 3433.69, repeat_frac 1.000)
|
||||
K=4 distinct=1747 of 54939 (mean count 31.45, repeat_frac 0.990)
|
||||
K=8 distinct=13474 of 54939 (mean count 4.08, repeat_frac 0.836)
|
||||
K=16 distinct=41447 of 54939 (mean count 1.33, repeat_frac 0.311)
|
||||
K=32 distinct=54454 of 54939 (mean count 1.01, repeat_frac 0.011)
|
||||
K=48 distinct=54744 of 54939 (mean count 1.00, repeat_frac 0.004)
|
||||
|
||||
----------------------------------------------------------------------------------------------------
|
||||
COARSENESS Q=3 -> 243 distinct single states (7.9 bits)
|
||||
K | single@D (temporal) | window (temporal) | window=SHUFFLED | window=REVERSE
|
||||
| logloss acc hitP | logloss acc hitP | logloss | logloss
|
||||
1 | 2.3782 0.4032 0.0897 | 2.3782 0.4032 0.0897 | 2.5116 | 2.5502
|
||||
4 | 2.3782 0.4032 0.0897 | 2.5461 0.3912 0.0903 | 2.6214 | 2.7670
|
||||
8 | 2.3782 0.4032 0.0897 | 2.8336 0.3594 0.0898 | 2.6284 | 3.0753
|
||||
16 | 2.3782 0.4032 0.0897 | 2.9473 0.3458 0.0898 | 2.6221 | 3.1881
|
||||
32 | 2.3782 0.4032 0.0897 | 2.9473 0.3458 0.0898 | 2.6225 | 3.1881
|
||||
48 | 2.3782 0.4032 0.0897 | 2.9473 0.3458 0.0898 | 2.6216 | 3.1881
|
||||
recurrence, distinct ordered window tuples over the whole corpus:
|
||||
K=1 distinct=82 of 54939 (mean count 669.99, repeat_frac 1.000)
|
||||
K=4 distinct=11018 of 54939 (mean count 4.99, repeat_frac 0.891)
|
||||
K=8 distinct=37337 of 54939 (mean count 1.47, repeat_frac 0.404)
|
||||
K=16 distinct=54163 of 54939 (mean count 1.01, repeat_frac 0.020)
|
||||
K=32 distinct=54754 of 54939 (mean count 1.00, repeat_frac 0.004)
|
||||
K=48 distinct=54772 of 54939 (mean count 1.00, repeat_frac 0.004)
|
||||
|
||||
----------------------------------------------------------------------------------------------------
|
||||
COARSENESS Q=4 -> 1024 distinct single states (10.0 bits)
|
||||
K | single@D (temporal) | window (temporal) | window=SHUFFLED | window=REVERSE
|
||||
| logloss acc hitP | logloss acc hitP | logloss | logloss
|
||||
1 | 2.3460 0.4094 0.0895 | 2.3460 0.4094 0.0895 | 2.5280 | 2.5352
|
||||
4 | 2.3460 0.4094 0.0895 | 2.5656 0.3833 0.0894 | 2.5672 | 2.8203
|
||||
8 | 2.3460 0.4094 0.0895 | 2.7265 0.3641 0.0885 | 2.5666 | 3.0138
|
||||
16 | 2.3460 0.4094 0.0895 | 2.7450 0.3622 0.0886 | 2.5613 | 3.0377
|
||||
32 | 2.3460 0.4094 0.0895 | 2.7450 0.3622 0.0886 | 2.5613 | 3.0377
|
||||
48 | 2.3460 0.4094 0.0895 | 2.7450 0.3622 0.0886 | 2.5640 | 3.0377
|
||||
recurrence, distinct ordered window tuples over the whole corpus:
|
||||
K=1 distinct=372 of 54939 (mean count 147.69, repeat_frac 0.999)
|
||||
K=4 distinct=24223 of 54939 (mean count 2.27, repeat_frac 0.685)
|
||||
K=8 distinct=49891 of 54939 (mean count 1.10, repeat_frac 0.134)
|
||||
K=16 distinct=54707 of 54939 (mean count 1.00, repeat_frac 0.005)
|
||||
K=32 distinct=54779 of 54939 (mean count 1.00, repeat_frac 0.004)
|
||||
K=48 distinct=54789 of 54939 (mean count 1.00, repeat_frac 0.004)
|
||||
|
||||
====================================================================================================
|
||||
HEADLINE (mean over 3 battle-split seeds; A=5)
|
||||
held-out log-loss. delta_window = window - single@D (NEGATIVE = the
|
||||
window beats the single state at the same decision tick)
|
||||
Q=2 single@D K=1 2.4524 K=4 2.4524 K=8 2.4524 K=16 2.4524 K=32 2.4524 K=48 2.4524
|
||||
window K=1 2.4524(+0.0000) K=4 2.5107(+0.0583) K=8 2.8015(+0.3491) K=16 3.1310(+0.6786) K=32 3.1310(+0.6786) K=48 3.1310(+0.6786)
|
||||
Q=3 single@D K=1 2.3782 K=4 2.3782 K=8 2.3782 K=16 2.3782 K=32 2.3782 K=48 2.3782
|
||||
window K=1 2.3782(+0.0000) K=4 2.5461(+0.1679) K=8 2.8336(+0.4554) K=16 2.9473(+0.5691) K=32 2.9473(+0.5691) K=48 2.9473(+0.5691)
|
||||
Q=4 single@D K=1 2.3460 K=4 2.3460 K=8 2.3460 K=16 2.3460 K=32 2.3460 K=48 2.3460
|
||||
window K=1 2.3460(+0.0000) K=4 2.5656(+0.2196) K=8 2.7265(+0.3805) K=16 2.7450(+0.3990) K=32 2.7450(+0.3990) K=48 2.7450(+0.3990)
|
||||
shuffle control: window(shuffled) - window(temporal) (must be >>0)
|
||||
Q=2 K=1 +0.0940 K=4 +0.2504 K=8 +0.0173 K=16 -0.3190 K=32 -0.3284 K=48 -0.3195
|
||||
Q=3 K=1 +0.1334 K=4 +0.0753 K=8 -0.2051 K=16 -0.3252 K=32 -0.3248 K=48 -0.3257
|
||||
Q=4 K=1 +0.1820 K=4 +0.0016 K=8 -0.1599 K=16 -0.1837 K=32 -0.1838 K=48 -0.1810
|
||||
reverse control: window(reverse) - window(temporal)
|
||||
Q=2 K=1 +0.1263 K=4 +0.1296 K=8 +0.1358 K=16 +0.1483 K=32 +0.1483 K=48 +0.1483
|
||||
Q=3 K=1 +0.1720 K=4 +0.2210 K=8 +0.2417 K=16 +0.2409 K=32 +0.2409 K=48 +0.2409
|
||||
Q=4 K=1 +0.1892 K=4 +0.2547 K=8 +0.2873 K=16 +0.2927 K=32 +0.2927 K=48 +0.2927
|
||||
robustness in the interpolation strength A (window temporal log-loss):
|
||||
Q=2 A=1 K=1 2.4526 K=4 2.6094 K=8 3.4554 K=16 4.4613 K=32 4.4613 K=48 4.4613
|
||||
Q=2 A=5 K=1 2.4524 K=4 2.5107 K=8 2.8015 K=16 3.1310 K=32 3.1310 K=48 3.1310
|
||||
Q=2 A=20 K=1 2.4524 K=4 2.4643 K=8 2.5491 K=16 2.6407 K=32 2.6407 K=48 2.6407
|
||||
Q=3 A=1 K=1 2.3782 K=4 2.8964 K=8 3.8419 K=16 4.2381 K=32 4.2381 K=48 4.2381
|
||||
Q=3 A=5 K=1 2.3782 K=4 2.5461 K=8 2.8336 K=16 2.9473 K=32 2.9473 K=48 2.9473
|
||||
Q=3 A=20 K=1 2.3810 K=4 2.4149 K=8 2.4868 K=16 2.5134 K=32 2.5134 K=48 2.5134
|
||||
Q=4 A=1 K=1 2.3499 K=4 3.0860 K=8 3.6606 K=16 3.7319 K=32 3.7319 K=48 3.7319
|
||||
Q=4 A=5 K=1 2.3460 K=4 2.5656 K=8 2.7265 K=16 2.7450 K=32 2.7450 K=48 2.7450
|
||||
Q=4 A=20 K=1 2.3557 K=4 2.3928 K=8 2.4289 K=16 2.4327 K=32 2.4327 K=48 2.4327
|
||||
====================================================================================================
|
||||
====================================================================================================
|
||||
FRAME fly (window STARTS at the fire tick, ends at t0+K-1)
|
||||
====================================================================================================
|
||||
MAJORITY / NO-WINDOW FLOOR (test acc 0.1912, log-loss 2.7810 bits, empirical hit 0.1100)
|
||||
target-bin edges [-120.0, -60.0, -18.0, 18.0, 60.0, 120.0] px -> central +-18 px hit bin. At the median
|
||||
fire range (487 px) the 36 px hit window subtends 4.23 deg; at 450 px
|
||||
it is 4.58 deg; at 100 px 20.41 deg. (atan(18/range).)
|
||||
target-bin distribution on train: 0:0.192, 1:0.145, 2:0.125, 3:0.109, 4:0.118, 5:0.139, 6:0.172
|
||||
|
||||
----------------------------------------------------------------------------------------------------
|
||||
COARSENESS Q=2 -> 32 distinct single states (5.0 bits)
|
||||
K | single@D (temporal) | window (temporal) | window=SHUFFLED | window=REVERSE
|
||||
| logloss acc hitP | logloss acc hitP | logloss | logloss
|
||||
1 | 2.6645 0.2935 0.1099 | 2.6645 0.2935 0.1099 | 2.6645 | 2.6645
|
||||
4 | 2.6376 0.2988 0.1101 | 2.7695 0.2808 0.1106 | 2.8096 | 2.7957
|
||||
8 | 2.5659 0.3076 0.1102 | 3.0331 0.2785 0.1107 | 3.1212 | 3.1197
|
||||
16 | 2.4047 0.3405 0.1104 | 3.1643 0.3070 0.1098 | 2.9981 | 3.3696
|
||||
32 | 1.8365 0.4189 0.1083 | 1.9970 0.4430 0.1078 | 2.3517 | 3.3696
|
||||
recurrence, distinct ordered window tuples over the whole corpus:
|
||||
K=1 distinct=16 of 8156 (mean count 509.75, repeat_frac 1.000)
|
||||
K=4 distinct=872 of 8156 (mean count 9.35, repeat_frac 0.954)
|
||||
K=8 distinct=3557 of 8156 (mean count 2.29, repeat_frac 0.665)
|
||||
K=16 distinct=7389 of 8156 (mean count 1.10, repeat_frac 0.133)
|
||||
K=32 distinct=8156 of 8156 (mean count 1.00, repeat_frac 0.000)
|
||||
|
||||
----------------------------------------------------------------------------------------------------
|
||||
COARSENESS Q=3 -> 243 distinct single states (7.9 bits)
|
||||
K | single@D (temporal) | window (temporal) | window=SHUFFLED | window=REVERSE
|
||||
| logloss acc hitP | logloss acc hitP | logloss | logloss
|
||||
1 | 2.6741 0.2901 0.1105 | 2.6741 0.2901 0.1105 | 2.6741 | 2.6741
|
||||
4 | 2.6262 0.2950 0.1103 | 2.8858 0.2658 0.1086 | 2.9378 | 2.9633
|
||||
8 | 2.5302 0.3184 0.1100 | 3.0207 0.2759 0.1089 | 2.9292 | 3.1779
|
||||
16 | 2.3103 0.3661 0.1090 | 2.6362 0.3390 0.1110 | 2.6531 | 3.2373
|
||||
32 | 1.4355 0.5221 0.1081 | 1.4625 0.5384 0.1030 | 2.2674 | 3.2373
|
||||
recurrence, distinct ordered window tuples over the whole corpus:
|
||||
K=1 distinct=81 of 8156 (mean count 100.69, repeat_frac 1.000)
|
||||
K=4 distinct=3383 of 8156 (mean count 2.41, repeat_frac 0.714)
|
||||
K=8 distinct=6857 of 8156 (mean count 1.19, repeat_frac 0.226)
|
||||
K=16 distinct=8140 of 8156 (mean count 1.00, repeat_frac 0.004)
|
||||
K=32 distinct=8156 of 8156 (mean count 1.00, repeat_frac 0.000)
|
||||
|
||||
----------------------------------------------------------------------------------------------------
|
||||
COARSENESS Q=4 -> 1024 distinct single states (10.0 bits)
|
||||
K | single@D (temporal) | window (temporal) | window=SHUFFLED | window=REVERSE
|
||||
| logloss acc hitP | logloss acc hitP | logloss | logloss
|
||||
1 | 2.7020 0.2817 0.1113 | 2.7020 0.2817 0.1113 | 2.7020 | 2.7020
|
||||
4 | 2.6397 0.2982 0.1112 | 2.9108 0.2750 0.1117 | 2.9175 | 3.0026
|
||||
8 | 2.6101 0.3037 0.1137 | 2.9084 0.2793 0.1159 | 2.8151 | 3.0861
|
||||
16 | 2.3853 0.3617 0.1123 | 2.5918 0.3377 0.1135 | 2.6342 | 3.0891
|
||||
32 | 1.3121 0.5988 0.1055 | 1.2668 0.6027 0.1089 | 2.3830 | 3.0891
|
||||
recurrence, distinct ordered window tuples over the whole corpus:
|
||||
K=1 distinct=252 of 8156 (mean count 32.37, repeat_frac 1.000)
|
||||
K=4 distinct=5368 of 8156 (mean count 1.52, repeat_frac 0.465)
|
||||
K=8 distinct=7983 of 8156 (mean count 1.02, repeat_frac 0.037)
|
||||
K=16 distinct=8156 of 8156 (mean count 1.00, repeat_frac 0.000)
|
||||
K=32 distinct=8156 of 8156 (mean count 1.00, repeat_frac 0.000)
|
||||
|
||||
====================================================================================================
|
||||
HEADLINE (mean over 3 battle-split seeds; A=5)
|
||||
held-out log-loss. delta_window = window - single@D (NEGATIVE = the
|
||||
window beats the single state at the same decision tick)
|
||||
Q=2 single@D K=1 2.6645 K=4 2.6376 K=8 2.5659 K=16 2.4047 K=32 1.8365
|
||||
window K=1 2.6645(+0.0000) K=4 2.7695(+0.1319) K=8 3.0331(+0.4673) K=16 3.1643(+0.7596) K=32 1.9970(+0.1605)
|
||||
Q=3 single@D K=1 2.6741 K=4 2.6262 K=8 2.5302 K=16 2.3103 K=32 1.4355
|
||||
window K=1 2.6741(+0.0000) K=4 2.8858(+0.2597) K=8 3.0207(+0.4904) K=16 2.6362(+0.3259) K=32 1.4625(+0.0270)
|
||||
Q=4 single@D K=1 2.7020 K=4 2.6397 K=8 2.6101 K=16 2.3853 K=32 1.3121
|
||||
window K=1 2.7020(+0.0000) K=4 2.9108(+0.2711) K=8 2.9084(+0.2984) K=16 2.5918(+0.2065) K=32 1.2668(-0.0453)
|
||||
shuffle control: window(shuffled) - window(temporal) (must be >>0)
|
||||
Q=2 K=1 +0.0000 K=4 +0.0400 K=8 +0.0881 K=16 -0.1662 K=32 +0.3548
|
||||
Q=3 K=1 +0.0000 K=4 +0.0520 K=8 -0.0914 K=16 +0.0169 K=32 +0.8048
|
||||
Q=4 K=1 +0.0000 K=4 +0.0068 K=8 -0.0933 K=16 +0.0425 K=32 +1.1162
|
||||
reverse control: window(reverse) - window(temporal)
|
||||
Q=2 K=1 +0.0000 K=4 +0.0262 K=8 +0.0865 K=16 +0.2053 K=32 +1.3726
|
||||
Q=3 K=1 +0.0000 K=4 +0.0774 K=8 +0.1572 K=16 +0.6011 K=32 +1.7748
|
||||
Q=4 K=1 +0.0000 K=4 +0.0918 K=8 +0.1776 K=16 +0.4974 K=32 +1.8223
|
||||
robustness in the interpolation strength A (window temporal log-loss):
|
||||
Q=2 A=1 K=1 2.6653 K=4 3.0442 K=8 4.0190 K=16 4.8738 K=32 2.8391
|
||||
Q=2 A=5 K=1 2.6645 K=4 2.7695 K=8 3.0331 K=16 3.1643 K=32 1.9970
|
||||
Q=2 A=20 K=1 2.6638 K=4 2.6681 K=8 2.6862 K=16 2.5833 K=32 1.7668
|
||||
Q=3 A=1 K=1 2.6844 K=4 3.5284 K=8 4.2822 K=16 3.6127 K=32 1.9264
|
||||
Q=3 A=5 K=1 2.6741 K=4 2.8858 K=8 3.0207 K=16 2.6362 K=32 1.4625
|
||||
Q=3 A=20 K=1 2.6681 K=4 2.6773 K=8 2.6250 K=16 2.3607 K=32 1.4132
|
||||
Q=4 A=1 K=1 2.7908 K=4 3.6658 K=8 3.9767 K=16 3.4357 K=32 1.4693
|
||||
Q=4 A=5 K=1 2.7020 K=4 2.9108 K=8 2.9084 K=16 2.5918 K=32 1.2668
|
||||
Q=4 A=20 K=1 2.6736 K=4 2.6783 K=8 2.6388 K=16 2.4433 K=32 1.4137
|
||||
====================================================================================================
|
||||
@@ -0,0 +1,579 @@
|
||||
#!/usr/bin/env python3
|
||||
"""STATE-WINDOW GATE: does a temporal WINDOW of wave-relative states predict
|
||||
DrussGT's future lateral position better than a SINGLE state?
|
||||
|
||||
This is the cheap veto test for the "feed SBC a temporal list of states" design
|
||||
(docs/state_window_gate.md). It is NOT a gun and makes no live-win claim.
|
||||
|
||||
DATA (real live battles, never regenerated here)
|
||||
/tmp/tfil_ab2/out/<A..E>/runN.jsonl + .events.jsonl + .rounds.json
|
||||
70 battles / 490 rounds / ~55k shots fired by ModularBot at the real
|
||||
unmodified DrussGT, recorded by tools/robocode_shim/run_bridge_battle.sh.
|
||||
In the capture rows `e*` is DrussGT (the subject) and `s*` is ModularBot (us);
|
||||
the per-shot geometry is re-derived by the validated instrument in
|
||||
common_libs/tests/analyze_drussgt_dodge_vs_power.py, which this file imports.
|
||||
|
||||
STATE (a design artifact -- see docs/state_window_gate.md for the rationale)
|
||||
One state at absolute tick t, in the frame of the bullet fired at t0 along
|
||||
direction u = (cos dir, sin dir):
|
||||
lat = (D(t) - P0) x u lateral offset from the bullet line, px
|
||||
vlat = lat(t) - lat(t-1) lateral velocity, px/tick (crossing/returning)
|
||||
toa = (t0 - t) + karr ticks until the bullet reaches arrival
|
||||
room = ray distance from D(t) along sign(vlat)*n until the arena wall, px
|
||||
turn = wrap180(eh(t) - eh(t-1)) signed turn rate, deg/tick
|
||||
Each field is quantised into Q in {2,3,4} bins (the numerosity dial). A
|
||||
window is K consecutive states; K=1 is the single-state baseline.
|
||||
|
||||
Two frames are tested, and BOTH compare a window against the single state at
|
||||
the SAME decision tick D (otherwise a longer window would win only because its
|
||||
decision tick is later):
|
||||
pre D = t0 (the fire tick); the window is the K pre-fire states ending at D.
|
||||
fly D = t0+K-1; the window is the first K states of the flight, and the
|
||||
baseline is the single state at D. Only shots with karr > 31 are used,
|
||||
so the whole window is strictly before the bullet's arrival.
|
||||
|
||||
TARGET
|
||||
perp_arr = DrussGT's SIGNED perpendicular offset from the bullet line at the
|
||||
tick our bullet reaches its along-track plane (the miss offset that decides
|
||||
the hit). Quantised into 7 bins (edges +-120, +-60, +-18 px); the central
|
||||
+-18 px bin is the hit window.
|
||||
|
||||
MODEL (deliberately dull: the question is about INFORMATION, not modelling)
|
||||
An interpolated (Jelinek-Mercer) suffix-backoff table over the quantised
|
||||
window: P = global; for j=1..K, P <- (count(suffix_j) + A*P)/(total+A). This
|
||||
is the direct analogue of the SBC "count coincidences" idea, it is
|
||||
order-sensitive, and it can never do much worse than the shorter context, so
|
||||
the sweep isolates information rather than overfitting. The context depth is
|
||||
capped at 12 (orders above that are never observed often enough to matter).
|
||||
A in {1,5,20} is swept as a robustness check (A=5 is primary).
|
||||
|
||||
SPLIT
|
||||
BY BATTLE, never by tick. All rounds of a battle go to one side. 70 %/30 %
|
||||
battle split, repeated over 3 seeds; model hyper-parameters (bin edges) are
|
||||
derived from the TRAIN side only.
|
||||
|
||||
CONTROLS
|
||||
1. shuffle -- random permutation of the K states inside each window (the
|
||||
multiset is preserved, order destroyed). MANDATORY.
|
||||
2. reverse -- deterministic reversal (preserves recurrence, reverses time).
|
||||
3. majority -- no-window floor.
|
||||
4. recurrence -- distinct windows, and how often a window repeats.
|
||||
|
||||
Run:
|
||||
python3 common_libs/tests/state_window_gate.py --tfil /tmp/tfil_ab2/out \
|
||||
--json common_libs/tests/fixtures/state_window_gate.json \
|
||||
> common_libs/tests/fixtures/state_window_gate_report.txt
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import bisect
|
||||
import collections
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import random
|
||||
import statistics
|
||||
import sys
|
||||
from array import array
|
||||
|
||||
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
sys.path.insert(0, HERE)
|
||||
import analyze_drussgt_dodge_vs_power as dodge # noqa: E402
|
||||
|
||||
ARENA_W, ARENA_H = 800.0, 600.0
|
||||
BOT_R = 18.0
|
||||
MAXK_PRE = 48 # pre-fire history depth (all K <= 48 are feasible)
|
||||
MAXK_FLY = 32 # during-flight window depth (max real flight is 43)
|
||||
FIELDS = ("lat", "vlat", "toa", "room", "turn")
|
||||
D_CAP = 12 # model context depth cap (see docs; orders > cap never observed)
|
||||
TARGET_EDGES = [-120.0, -60.0, -18.0, 18.0, 60.0, 120.0]
|
||||
HIT_BIN = 3 # bin [ -18, 18 ) -> the bot radius
|
||||
K_PRE = (1, 4, 8, 16, 32, 48)
|
||||
K_FLY = (1, 4, 8, 16, 32)
|
||||
|
||||
|
||||
def wrap180(a):
|
||||
return ((a + 180.0) % 360.0) - 180.0
|
||||
|
||||
|
||||
def room_to_wall(px, py, dx, dy):
|
||||
"""Distance from (px,py) along unit (dx,dy) until leaving the arena
|
||||
(accounting for the 18 px bot radius)."""
|
||||
t = float("inf")
|
||||
for p, d, lo, hi in ((px, dx, BOT_R, ARENA_W - BOT_R),
|
||||
(py, dy, BOT_R, ARENA_H - BOT_R)):
|
||||
if abs(d) > 1e-9:
|
||||
cand = (hi - p) / d if d > 0 else (lo - p) / d
|
||||
if cand < t:
|
||||
t = cand
|
||||
if t == float("inf"):
|
||||
return 0.0
|
||||
return max(0.0, t)
|
||||
|
||||
|
||||
def tbin(v):
|
||||
return bisect.bisect_right(TARGET_EDGES, v)
|
||||
|
||||
|
||||
def raw_state(run, t0, karr, P0, ux, uy, t, rnd):
|
||||
rs = run.start[rnd]
|
||||
re = rs + run.count[rnd] - 1
|
||||
tc = min(max(t, rs), re)
|
||||
tp = max(tc - 1, rs)
|
||||
r = run.by_tick.get(tc)
|
||||
rp = run.by_tick.get(tp)
|
||||
if r is None or rp is None:
|
||||
return None
|
||||
lat = (r["ex"] - P0[0]) * uy - (r["ey"] - P0[1]) * ux
|
||||
latp = (rp["ex"] - P0[0]) * uy - (rp["ey"] - P0[1]) * ux
|
||||
vlat = lat - latp
|
||||
toa = (t0 - tc) + karr
|
||||
turn = wrap180(r["eh"] - rp["eh"])
|
||||
sg = 1.0 if vlat >= 0 else -1.0
|
||||
room = room_to_wall(r["ex"], r["ey"], -uy * sg, ux * sg)
|
||||
return (lat, vlat, toa, room, turn)
|
||||
|
||||
|
||||
def extract_run(run):
|
||||
out = []
|
||||
for s in run.shots():
|
||||
t0 = s["tick"]
|
||||
karr = s["flight"]
|
||||
if karr is None:
|
||||
continue
|
||||
th = math.radians(s["_dir"])
|
||||
ux, uy = math.cos(th), math.sin(th)
|
||||
P0 = (s["_x"], s["_y"])
|
||||
rnd = s["rnd"]
|
||||
pre = array("f")
|
||||
ok = True
|
||||
for off in range(-(MAXK_PRE - 1), 1):
|
||||
rs = raw_state(run, t0, karr, P0, ux, uy, t0 + off, rnd)
|
||||
if rs is None:
|
||||
ok = False
|
||||
break
|
||||
pre.extend(rs)
|
||||
if not ok:
|
||||
continue
|
||||
fly = []
|
||||
for off in range(MAXK_FLY):
|
||||
rs = raw_state(run, t0, karr, P0, ux, uy, t0 + off, rnd)
|
||||
fly.append(None if rs is None else rs)
|
||||
out.append(dict(
|
||||
battle=os.path.basename(os.path.dirname(run.cap_path)) + "/" +
|
||||
os.path.basename(run.cap_path),
|
||||
rnd=rnd, t0=t0, karr=karr, perp_arr=s["perp_arr"],
|
||||
perp_fire=s["perp_fire"], range=s["range"], power=s["power"],
|
||||
hit=s["hit"], tbin=tbin(s["perp_arr"]), pre=pre, fly=fly,
|
||||
))
|
||||
return out
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ quantise
|
||||
def edges_for(vals, q):
|
||||
xs = sorted(vals)
|
||||
n = len(xs)
|
||||
return [xs[min(n - 1, int((qi / q) * n))] for qi in range(1, q)]
|
||||
|
||||
|
||||
def frame_edges(samples, frame, Q):
|
||||
"""Quantile bin edges per field, derived from the TRAIN side only."""
|
||||
edges = []
|
||||
for f in range(5):
|
||||
vals = []
|
||||
if frame == "pre":
|
||||
for s in samples:
|
||||
if s["_split"] != "train":
|
||||
continue
|
||||
a = s["pre"]
|
||||
for i in range(MAXK_PRE):
|
||||
vals.append(a[i * 5 + f])
|
||||
else:
|
||||
for s in samples:
|
||||
if s["_split"] != "train":
|
||||
continue
|
||||
for st in s["fly"]:
|
||||
if st is not None:
|
||||
vals.append(st[f])
|
||||
edges.append(edges_for(vals, Q))
|
||||
return edges
|
||||
|
||||
|
||||
def apply_codes(samples, frame, edges, Q):
|
||||
res = []
|
||||
for s in samples:
|
||||
if frame == "pre":
|
||||
a = s["pre"]
|
||||
n = MAXK_PRE
|
||||
codes = [0] * n
|
||||
for i in range(n):
|
||||
base = i * 5
|
||||
c = 0
|
||||
for f in range(5):
|
||||
c += bisect.bisect_right(edges[f], a[base + f]) * (Q ** f)
|
||||
codes[i] = c
|
||||
else:
|
||||
a = s["fly"]
|
||||
n = MAXK_FLY
|
||||
codes = [0] * n
|
||||
for i in range(n):
|
||||
if a[i] is None:
|
||||
codes[i] = -1
|
||||
continue
|
||||
c = 0
|
||||
for f in range(5):
|
||||
c += bisect.bisect_right(edges[f], a[i][f]) * (Q ** f)
|
||||
codes[i] = c
|
||||
res.append(codes)
|
||||
return res
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ the model
|
||||
def order_window(w, mode, rng):
|
||||
if mode == "temporal":
|
||||
return w
|
||||
if mode == "reverse":
|
||||
return w[::-1]
|
||||
if mode == "shuffled":
|
||||
return [w[i] for i in rng.sample(range(len(w)), len(w))]
|
||||
raise ValueError(mode)
|
||||
|
||||
|
||||
def build_counts(windows, ys, mode, Dcap, rng):
|
||||
"""cbyorder[j][context] = [ {target_bin: count}, total ].
|
||||
|
||||
The model is a standard interpolated (Jelinek-Mercer) suffix backoff: for a
|
||||
window w the prediction for the next-state target is
|
||||
P = global
|
||||
for j = 1..K: P = (count(suffix_j) + A*P) / (total(suffix_j) + A)
|
||||
so a longer context is only believed as far as the data supports it, and the
|
||||
model can never do much worse than the shorter one. This keeps the MODEL
|
||||
uninteresting, which is what the gate needs."""
|
||||
cbyorder = [dict() for _ in range(Dcap + 1)]
|
||||
for w, y in zip(windows, ys):
|
||||
wo = order_window(w, mode, rng)
|
||||
ctx = ()
|
||||
for j in range(1, min(Dcap, len(wo)) + 1):
|
||||
ctx = (wo[-j],) + ctx
|
||||
e = cbyorder[j].get(ctx)
|
||||
if e is None:
|
||||
e = [{}, 0]
|
||||
cbyorder[j][ctx] = e
|
||||
e[0][y] = e[0].get(y, 0) + 1
|
||||
e[1] += 1
|
||||
return cbyorder
|
||||
|
||||
|
||||
NBINS = len(TARGET_EDGES) + 1
|
||||
|
||||
|
||||
def glob_vec(glob, nbins=NBINS):
|
||||
tot = sum(glob.values())
|
||||
return [(glob.get(b, 0)) / tot for b in range(nbins)]
|
||||
|
||||
|
||||
def predict(w, cbyorder, P0, K, A):
|
||||
P = list(P0)
|
||||
for j in range(1, min(K, len(w), len(cbyorder) - 1) + 1):
|
||||
ctx = tuple(w[-j:])
|
||||
e = cbyorder[j].get(ctx)
|
||||
if e is None:
|
||||
continue
|
||||
cnt, tot = e
|
||||
P = [(cnt.get(b, 0) + A * P[b]) / (tot + A) for b in range(len(P))]
|
||||
return P
|
||||
|
||||
|
||||
def evaluate(windows, ys, cbyorder, P0, mode, Kdepth, A, rng):
|
||||
acc = 0
|
||||
n = 0
|
||||
ll = 0.0
|
||||
hitp = 0.0
|
||||
for w, y in zip(windows, ys):
|
||||
wo = order_window(w, mode, rng)
|
||||
P = predict(wo, cbyorder, P0, Kdepth, A)
|
||||
best = max(range(len(P)), key=lambda b: (P[b], -b))
|
||||
if best == y:
|
||||
acc += 1
|
||||
ll += -math.log2(max(P[y], 1e-12))
|
||||
hitp += P[HIT_BIN]
|
||||
n += 1
|
||||
return {"n": n, "acc": acc / n, "logloss": ll / n,
|
||||
"implied_hit": hitp / n}
|
||||
|
||||
|
||||
def majority_floor(samples):
|
||||
train = [s for s in samples if s["_split"] == "train"]
|
||||
glob = collections.Counter(s["tbin"] for s in train)
|
||||
top = max(glob, key=lambda b: glob[b])
|
||||
acc = ll = 0.0
|
||||
n = 0
|
||||
test_hit = 0
|
||||
for s in samples:
|
||||
if s["_split"] != "test":
|
||||
continue
|
||||
n += 1
|
||||
if s["tbin"] == top:
|
||||
acc += 1
|
||||
ll += -math.log2(glob[s["tbin"]] / sum(glob.values()))
|
||||
test_hit += 1 if s["tbin"] == HIT_BIN else 0
|
||||
return {"top": top, "n": n, "acc": acc / n, "logloss": ll / n,
|
||||
"emp_hit": test_hit / n,
|
||||
"dist": {str(k): v / sum(glob.values()) for k, v in sorted(glob.items())}}
|
||||
|
||||
|
||||
def recurrence(windows, K):
|
||||
"""Distinct windows and repeat rate over the WHOLE corpus (no split)."""
|
||||
seen = collections.Counter()
|
||||
for w in windows:
|
||||
k = min(K, len(w))
|
||||
seen[tuple(w[len(w) - k:])] += 1
|
||||
total = len(windows)
|
||||
rep2 = sum(c for c in seen.values() if c >= 2)
|
||||
return {"distinct": len(seen), "total": total, "repeat_frac": rep2 / total,
|
||||
"mean_count": total / len(seen)}
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ driver
|
||||
def split_battles(samples, seed):
|
||||
battles = sorted({s["battle"] for s in samples})
|
||||
rng = random.Random(seed)
|
||||
rng.shuffle(battles)
|
||||
ntr = int(round(0.70 * len(battles)))
|
||||
train = set(battles[:ntr])
|
||||
for s in samples:
|
||||
s["_split"] = "train" if s["battle"] in train else "test"
|
||||
|
||||
|
||||
def window_list(codes, frame, K):
|
||||
"""The ordered window for one sample.
|
||||
|
||||
pre frame: the full 48-state pre-fire history ENDING at the fire tick; the
|
||||
model's depth parameter selects how many of the most recent
|
||||
states it may use, so K=1 is the fire-tick state itself.
|
||||
fly frame: the first K states of the flight, i.e. ending at t0+K-1; the
|
||||
single-state baseline is then the state at that same tick.
|
||||
"""
|
||||
if frame == "pre":
|
||||
return list(codes)
|
||||
return list(codes[:K])
|
||||
|
||||
|
||||
def run_frame(samples, frame, Ks, seeds, As):
|
||||
res = {"frame": frame, "seeds": {}}
|
||||
for seed in seeds:
|
||||
split_battles(samples, seed)
|
||||
tr = [i for i, s in enumerate(samples) if s["_split"] == "train"]
|
||||
te = [i for i, s in enumerate(samples) if s["_split"] == "test"]
|
||||
ytr = [samples[i]["tbin"] for i in tr]
|
||||
yte = [samples[i]["tbin"] for i in te]
|
||||
seed_res = {"majority": majority_floor(samples), "Q": {}}
|
||||
for Q in (2, 3, 4):
|
||||
edges = frame_edges(samples, frame, Q)
|
||||
codes = apply_codes(samples, frame, edges, Q)
|
||||
glob = {}
|
||||
for i in tr:
|
||||
glob[samples[i]["tbin"]] = glob.get(samples[i]["tbin"], 0) + 1
|
||||
P0 = glob_vec(glob)
|
||||
qres = {}
|
||||
for mode in ("temporal", "shuffled", "reverse"):
|
||||
mres = {str(A): {} for A in As}
|
||||
cache = {}
|
||||
for K in Ks:
|
||||
wid = "full" if frame == "pre" else K
|
||||
if wid not in cache:
|
||||
wl = [window_list(codes[i], frame, K) for i in range(len(samples))]
|
||||
cby = build_counts([wl[i] for i in tr], ytr, mode, D_CAP,
|
||||
random.Random(1000 + seed))
|
||||
cache[wid] = (cby, wl)
|
||||
cby, wl = cache[wid]
|
||||
wte = [wl[i] for i in te]
|
||||
kd = min(K, D_CAP)
|
||||
for A in As:
|
||||
rnge = random.Random(4000 + seed * 13 + K)
|
||||
mres[str(A)][str(K)] = {
|
||||
"window": evaluate(wte, yte, cby, P0, mode, kd, A, rnge)}
|
||||
qres[mode] = mres
|
||||
# single-state-at-the-same-decision-tick baseline (true recent state)
|
||||
sres = {str(A): {} for A in As}
|
||||
for K in Ks:
|
||||
wid = "full" if frame == "pre" else K
|
||||
wl = [window_list(codes[i], frame, K) for i in range(len(samples))]
|
||||
cby = build_counts([wl[i] for i in tr], ytr, "temporal", D_CAP,
|
||||
random.Random(1000 + seed))
|
||||
wte = [wl[i] for i in te]
|
||||
for A in As:
|
||||
sres[str(A)][str(K)] = evaluate(
|
||||
wte, yte, cby, P0, "temporal", 1, A,
|
||||
random.Random(5000 + seed * 17 + K))
|
||||
qres["single"] = sres
|
||||
qres["recurrence"] = {str(K): recurrence(
|
||||
[window_list(codes[i], frame, K) for i in range(len(samples))], K)
|
||||
for K in Ks}
|
||||
seed_res["Q"][str(Q)] = qres
|
||||
res["seeds"][str(seed)] = seed_res
|
||||
return res
|
||||
|
||||
|
||||
def mean_over(xs):
|
||||
return statistics.fmean(xs) if xs else float("nan")
|
||||
|
||||
|
||||
def cell(res, frame, Q, mode, A, K, key="window"):
|
||||
if key == "single":
|
||||
q = lambda sd: res[frame]["seeds"][sd]["Q"][str(Q)]["single"][str(A)][str(K)]
|
||||
else:
|
||||
q = lambda sd: res[frame]["seeds"][sd]["Q"][str(Q)][mode][str(A)][str(K)][key]
|
||||
a = [q(sd) for sd in res[frame]["seeds"]]
|
||||
return (mean_over([c["acc"] for c in a]),
|
||||
mean_over([c["logloss"] for c in a]),
|
||||
mean_over([c["implied_hit"] for c in a]))
|
||||
|
||||
|
||||
def render(res, out, A_primary):
|
||||
lines = []
|
||||
|
||||
def p(s=""):
|
||||
lines.append(s)
|
||||
print(s)
|
||||
|
||||
for frame in res:
|
||||
Ks = K_PRE if frame == "pre" else K_FLY
|
||||
p("=" * 100)
|
||||
p("FRAME %s%s" % (frame, " (window ENDS at the fire tick, looks BACK)"
|
||||
if frame == "pre" else
|
||||
" (window STARTS at the fire tick, ends at t0+K-1)"))
|
||||
p("=" * 100)
|
||||
maj = [res[frame]["seeds"][sd]["majority"] for sd in res[frame]["seeds"]]
|
||||
p("MAJORITY / NO-WINDOW FLOOR (test acc %.4f, log-loss %.4f bits,"
|
||||
" empirical hit %.4f)" % (mean_over([m["acc"] for m in maj]),
|
||||
mean_over([m["logloss"] for m in maj]),
|
||||
mean_over([m["emp_hit"] for m in maj])))
|
||||
p(" target-bin edges %s px -> central +-18 px hit bin. At the median"
|
||||
% TARGET_EDGES)
|
||||
p(" fire range (487 px) the 36 px hit window subtends %.2f deg; at 450 px"
|
||||
% math.degrees(2 * math.atan(18.0 / 487.0)))
|
||||
p(" it is %.2f deg; at 100 px %.2f deg. (atan(18/range).)"
|
||||
% (math.degrees(2 * math.atan(18.0 / 450.0)),
|
||||
math.degrees(2 * math.atan(18.0 / 100.0))))
|
||||
p(" target-bin distribution on train: %s" %
|
||||
", ".join("%s:%.3f" % (k, v) for k, v in sorted(maj[0]["dist"].items())))
|
||||
p()
|
||||
for Q in (2, 3, 4):
|
||||
nstates = Q ** 5
|
||||
p("-" * 100)
|
||||
p("COARSENESS Q=%d -> %d distinct single states (%.1f bits)"
|
||||
% (Q, nstates, math.log2(nstates)))
|
||||
p(" %-4s | %-24s | %-24s | %-13s | %-13s" %
|
||||
("K", "single@D (temporal)", "window (temporal)", "window=SHUFFLED",
|
||||
"window=REVERSE"))
|
||||
p(" %-4s | %-24s | %-24s | %-13s | %-13s" %
|
||||
("", "logloss acc hitP", "logloss acc hitP", "logloss",
|
||||
"logloss"))
|
||||
for K in Ks:
|
||||
sc = cell(res, frame, Q, "temporal", A_primary, K, "single")
|
||||
tc = cell(res, frame, Q, "temporal", A_primary, K, "window")
|
||||
sh = cell(res, frame, Q, "shuffled", A_primary, K, "window")
|
||||
rv = cell(res, frame, Q, "reverse", A_primary, K, "window")
|
||||
p(" %-4d | %.4f %.4f %.4f | %.4f %.4f %.4f | %.4f | %.4f"
|
||||
% (K, sc[1], sc[0], sc[2], tc[1], tc[0], tc[2],
|
||||
sh[1], rv[1]))
|
||||
rec = res[frame]["seeds"][sorted(res[frame]["seeds"])[0]]["Q"][str(Q)]["recurrence"]
|
||||
p(" recurrence, distinct ordered window tuples over the whole corpus:")
|
||||
for K in Ks:
|
||||
r = rec[str(K)]
|
||||
p(" K=%-2d distinct=%-8d of %-6d (mean count %.2f, repeat_frac %.3f)"
|
||||
% (K, r["distinct"], r["total"], r["mean_count"], r["repeat_frac"]))
|
||||
p()
|
||||
p("=" * 100)
|
||||
p("HEADLINE (mean over %d battle-split seeds; A=%g)" %
|
||||
(len(res[frame]["seeds"]), A_primary))
|
||||
p(" held-out log-loss. delta_window = window - single@D (NEGATIVE = the")
|
||||
p(" window beats the single state at the same decision tick)")
|
||||
for Q in (2, 3, 4):
|
||||
srow = [cell(res, frame, Q, "temporal", A_primary, K, "single")[1] for K in Ks]
|
||||
trow = [cell(res, frame, Q, "temporal", A_primary, K, "window")[1] for K in Ks]
|
||||
p(" Q=%d single@D " % Q + " ".join("K=%d %.4f" % (K, v)
|
||||
for K, v in zip(Ks, srow)))
|
||||
p(" window " + " ".join("K=%d %.4f(%+.4f)" % (K, t, t - s)
|
||||
for K, t, s in zip(Ks, trow, srow)))
|
||||
p(" shuffle control: window(shuffled) - window(temporal) (must be >>0)")
|
||||
for Q in (2, 3, 4):
|
||||
row = [cell(res, frame, Q, "shuffled", A_primary, K, "window")[1] -
|
||||
cell(res, frame, Q, "temporal", A_primary, K, "window")[1]
|
||||
for K in Ks]
|
||||
p(" Q=%d " % Q + " ".join("K=%d %+.4f" % (K, v) for K, v in zip(Ks, row)))
|
||||
p(" reverse control: window(reverse) - window(temporal)")
|
||||
for Q in (2, 3, 4):
|
||||
row = [cell(res, frame, Q, "reverse", A_primary, K, "window")[1] -
|
||||
cell(res, frame, Q, "temporal", A_primary, K, "window")[1]
|
||||
for K in Ks]
|
||||
p(" Q=%d " % Q + " ".join("K=%d %+.4f" % (K, v) for K, v in zip(Ks, row)))
|
||||
p(" robustness in the interpolation strength A (window temporal log-loss):")
|
||||
for Q in (2, 3, 4):
|
||||
for A in out.get("A_all", []):
|
||||
row = [cell(res, frame, Q, "temporal", A, K, "window")[1] for K in Ks]
|
||||
p(" Q=%d A=%-4g " % (Q, A) +
|
||||
" ".join("K=%d %.4f" % (K, v) for K, v in zip(Ks, row)))
|
||||
p("=" * 100)
|
||||
out["report"] = lines
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--tfil", default="/tmp/tfil_ab2/out")
|
||||
ap.add_argument("--json", default=None)
|
||||
ap.add_argument("--limit-runs", type=int, default=0)
|
||||
ap.add_argument("--seed", type=int, default=1)
|
||||
ap.add_argument("--seeds", type=int, default=3)
|
||||
ap.add_argument("--frames", default="pre")
|
||||
ap.add_argument("--alphas", default="1,5,20")
|
||||
a = ap.parse_args()
|
||||
|
||||
runs = dodge.discover_tfil(a.tfil)
|
||||
if a.limit_runs:
|
||||
runs = runs[:a.limit_runs]
|
||||
print("[gate] %d battles" % len(runs), file=sys.stderr)
|
||||
samples = []
|
||||
for k, run in enumerate(runs):
|
||||
samples.extend(extract_run(run))
|
||||
if (k + 1) % 10 == 0:
|
||||
print("[gate] %d/%d battles, %d shots" % (k + 1, len(runs),
|
||||
len(samples)), file=sys.stderr)
|
||||
print("[gate] %d samples" % len(samples), file=sys.stderr)
|
||||
hit = sum(1 for s in samples if s["tbin"] == HIT_BIN)
|
||||
print("[gate] empirical arrival hit bin (|perp|<18): %.4f" % (hit / len(samples)),
|
||||
file=sys.stderr)
|
||||
|
||||
seeds = [a.seed + i for i in range(a.seeds)]
|
||||
As = [float(x) for x in a.alphas.split(",")]
|
||||
A_primary = 5.0
|
||||
res = {}
|
||||
nfly = 0
|
||||
for frame in a.frames.split(","):
|
||||
if frame == "pre":
|
||||
fs = samples
|
||||
else:
|
||||
kmin = max(K_FLY) + 1 # keep the whole window strictly before arrival
|
||||
fs = [s for s in samples if s["karr"] >= kmin]
|
||||
nfly = len(fs)
|
||||
print("[gate] frame %s: %d shots with karr >= %d" %
|
||||
(frame, len(fs), kmin), file=sys.stderr)
|
||||
res[frame] = run_frame(fs, frame, K_PRE if frame == "pre" else K_FLY,
|
||||
seeds, As)
|
||||
out = {"samples": len(samples), "battles": len(runs), "seeds": seeds,
|
||||
"fly_samples": nfly,
|
||||
"alphas": As, "A_primary": A_primary, "A_all": As,
|
||||
"D_cap": D_CAP, "target_edges": TARGET_EDGES, "hit_bin": HIT_BIN,
|
||||
"k_pre": list(K_PRE), "k_fly": list(K_FLY)}
|
||||
render(res, out, A_primary)
|
||||
out["raw"] = res
|
||||
if a.json:
|
||||
with open(a.json, "w") as f:
|
||||
json.dump(out, f, indent=1, default=str)
|
||||
return out
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,321 @@
|
||||
# The state-window gate: does a temporal WINDOW of wave-relative states beat a SINGLE state?
|
||||
|
||||
**The design under test (the user's words).** *"We decide first what makes an
|
||||
'environment state' ... then we feed as input a long list of these states, long
|
||||
enough to have one or more patterns of movement inside."* The premise is that SBC
|
||||
is a coincidence detector and that a **temporal list** of states, fed instead of a
|
||||
single tick, turns "coincidence" into "movement pattern".
|
||||
|
||||
**Verdict: NO.** On held-out battles a window of wave-relative states **does not**
|
||||
predict DrussGT's lateral position at the bullet's arrival better than the single
|
||||
state at the **same decision tick**. On the pre-fire frame the window is strictly
|
||||
and substantially *worse* at every window length, every coarseness and every
|
||||
smoothing strength. On the during-flight frame the apparent gain is entirely the
|
||||
**later decision tick**, not the window — and where the window finally edges the
|
||||
single state (K=32, Q=4) it is by 0.045 bits on a subset where the 32-window
|
||||
**never repeats at all**. The design, as stated, is **dead**. The *state
|
||||
definition* survives and is the useful result.
|
||||
|
||||
This is a gate, not a gun. **No live win is claimed and none is measured here.**
|
||||
|
||||
---
|
||||
|
||||
## 1. The state (design artifact)
|
||||
|
||||
The dodge is a response to **our bullet**, so the state is **wave-relative**, never
|
||||
absolute world coordinates. One state at absolute tick `t`, in the frame of the
|
||||
bullet fired at `t0` along `u = (cos dir, sin dir)` (with `n = (-u_y, u_x)` the
|
||||
lateral unit vector):
|
||||
|
||||
| field | definition | why |
|
||||
|---|---|---|
|
||||
| `lat` | `(D(t) − P0) × u` — DrussGT's signed perpendicular offset from the bullet line, px | the quantity that decides the hit; signed so "which side" is visible |
|
||||
| `vlat` | `lat(t) − lat(t−1)` — signed lateral velocity, px/tick | crossing vs returning |
|
||||
| `toa` | `(t0 − t) + karr` — ticks until the bullet reaches its arrival plane | how much time is left to move |
|
||||
| `room` | ray distance from `D(t)` along `sign(vlat)·n` until the arena wall (18 px bot radius) | room to keep running vs being cornered — the only *directional* wall field, unlike plain "distance to nearest wall" |
|
||||
| `turn` | `wrap180(eh(t) − eh(t−1))` — signed turn rate, deg/tick | surfers slow to turn; turning reveals a reversal |
|
||||
|
||||
**Refinements to the starting proposal, and why.** (a) was kept signed, not
|
||||
absolutised, because the target is signed. (b) was made a *difference of the
|
||||
lateral offset*, so it is exactly the lateral component of velocity in the bullet
|
||||
frame rather than a speed. (d) "distance to the wall" was made *directional*
|
||||
(room to run in the current lateral-motion direction); a plain nearest-wall
|
||||
distance is not wave-relative and is nearly constant at long range. (e)
|
||||
"heading/turn rate" was reduced to the **turn rate** alone, because heading
|
||||
relative to the bullet is already implied by the sign of `vlat`, and the turn rate
|
||||
is the thing that distinguishes a surf from a corner.
|
||||
|
||||
**Quantisation (the numerosity dial).** Each field is binned into `Q ∈ {2, 3, 4}`
|
||||
symbols. A state is therefore:
|
||||
|
||||
| Q | bits / state | distinct states |
|
||||
|---|---|---|
|
||||
| 2 | **5.0** | 32 |
|
||||
| 3 | **7.9** | 243 |
|
||||
| 4 | **10.0** | 1024 |
|
||||
|
||||
Bin edges are **quantiles of the TRAIN side only** (no test leakage); a coarse
|
||||
state makes two similar situations look the same, which is what lets a
|
||||
"coincidence" repeat. The sweep below is exactly the question of whether coarser
|
||||
buys more recurrence than it loses in resolution.
|
||||
|
||||
### The two windows tested, and why both
|
||||
|
||||
A window is `K` consecutive states. Crucially, **the baseline is always the single
|
||||
state at the same decision tick `D`** — otherwise a longer window would "win" only
|
||||
because its last state is closer to the answer.
|
||||
|
||||
* **`pre`** — `D = t0` (the fire tick); the window is the `K` pre-fire states
|
||||
ending at `D`. This is what a bot can compute *before* pulling the trigger. All
|
||||
`K ≤ 48` are feasible (max recorded flight is 43 ticks).
|
||||
* **`fly`** — `D = t0+K−1`; the window is the first `K` states of the flight, and
|
||||
the baseline is the single state at `D`. Only shots with `karr > 31` are used,
|
||||
so the whole window is strictly before arrival (8156 of 54939 shots).
|
||||
|
||||
The task's field `toa` only exists once a bullet is in the air, so `fly` is the
|
||||
"reaction" reading; `pre` is the "am I about to be dodged?" reading. Both are
|
||||
reported.
|
||||
|
||||
### The target
|
||||
|
||||
`perp_arr` — DrussGT's **signed** perpendicular offset from the bullet line at the
|
||||
tick the bullet reaches its along-track plane. This is the miss offset that
|
||||
decides the hit; unlike "required lead" it does not depend on the bullet speed. It
|
||||
is quantised into 7 bins with edges `±120, ±60, ±18` px; the central `±18` px bin
|
||||
is the hit window (bot radius 18 px).
|
||||
|
||||
**Bin width in degrees-at-450px.** The central hit bin is 36 px wide = one bot
|
||||
diameter. `atan(18/range)` is the hit half-window: **2.29°** at 450 px, 4.23° at
|
||||
the median fire range (487 px), and 20.4° at 100 px. So the 36 px central bin is
|
||||
**4.58° at 450 px**. The three inner bins are 36/42/60 px wide (4.58/5.34/7.63° at
|
||||
450 px) — the target resolution is *finer* than the ~2.29° hit half-window, so a
|
||||
model that predicts the bin is not being asked an impossible question.
|
||||
|
||||
---
|
||||
|
||||
## 2. Data and split
|
||||
|
||||
* **Data (MEASURED present before use):** `/tmp/tfil_ab2/out/<A..E>/runN.jsonl` +
|
||||
`.events.jsonl` + `.rounds.json` — **70 real live battles vs the unmodified
|
||||
DrussGT, 490 rounds, 54 939 shots** by ModularBot, recorded by
|
||||
`tools/robocode_shim/run_bridge_battle.sh`. Per-shot geometry is re-derived by
|
||||
the validated instrument `common_libs/tests/analyze_drussgt_dodge_vs_power.py`
|
||||
(which this extractor imports). Attribution (`e*`=DrussGT, `s*`=us) is the one
|
||||
already documented in `docs/drussgt_dodge_vs_power.md`.
|
||||
* **Split (stated explicitly): BY BATTLE, never by tick.** All rounds of a battle
|
||||
go to one side. A 70 %/30 % battle split (**49 train / 21 test battles**),
|
||||
repeated over **3 seeds**. Bin edges are derived from the train side only. Two
|
||||
ticks inside one round never straddle the split.
|
||||
|
||||
---
|
||||
|
||||
## 3. Model (deliberately dull)
|
||||
|
||||
The question is about **information**, not modelling, so the model is a plain
|
||||
**interpolated (Jelinek–Mercer) suffix-backoff table** — the direct analogue of
|
||||
SBC "counting coincidences":
|
||||
|
||||
```
|
||||
P = global target-bin histogram
|
||||
for j = 1..K: P <- (count(suffix_j) + A*P) / (total(suffix_j) + A)
|
||||
```
|
||||
|
||||
It is order-sensitive, it can never do much worse than the shorter context, and it
|
||||
degrades gracefully when a suffix is unseen. Context depth is capped at 12 (orders
|
||||
above that are never repeated often enough to matter — see §5). `A ∈ {1,5,20}` is
|
||||
swept; **A = 5 is primary**. A majority/no-window predictor is the floor. The
|
||||
model predicts the 7-bin displacement distribution; we report held-out **accuracy**
|
||||
(argmax), **log-loss** in bits, and the **implied hit probability** `P(central
|
||||
bin)`.
|
||||
|
||||
---
|
||||
|
||||
## 4. Result — does the window beat the single state? **No** *(MEASURED)*
|
||||
|
||||
70 battles, 3 battle-split seeds, mean held-out log-loss (bits). `delta = window −
|
||||
single@D`; **negative would mean the window wins**.
|
||||
|
||||
### 4.1 `pre` frame — window ends at the fire tick, looks back *(54 939 shots)*
|
||||
|
||||
`single@D` is the fire-tick state; it does **not** change with K. A = 5:
|
||||
|
||||
| K | Q=2 single → window (Δ) | Q=3 single → window (Δ) | Q=4 single → window (Δ) |
|
||||
|---|---|---|---|
|
||||
| 1 | 2.4524 → 2.4524 (—) | 2.3782 → 2.3782 (—) | 2.3460 → 2.3460 (—) |
|
||||
| 4 | 2.4524 → 2.5107 (**+0.058**) | 2.3782 → 2.5461 (**+0.168**) | 2.3460 → 2.5656 (**+0.220**) |
|
||||
| 8 | 2.4524 → 2.8015 (**+0.349**) | 2.3782 → 2.8336 (**+0.455**) | 2.3460 → 2.7265 (**+0.381**) |
|
||||
| 16 | 2.4524 → 3.1310 (**+0.679**) | 2.3782 → 2.9473 (**+0.569**) | 2.3460 → 2.7450 (**+0.399**) |
|
||||
| 32 | same as K=16 (depth cap) | same | same |
|
||||
| 48 | same as K=16 (depth cap) | same | same |
|
||||
|
||||
Accuracy tells the same story (Q=4: single `0.409` → window `0.383`/`0.364`/
|
||||
`0.362`). The per-seed deltas are essentially identical (e.g. Q=2/K=4:
|
||||
`+0.0572, +0.0534, +0.0644`) — there is no seed where the window helps. The
|
||||
**heavy-smoothing** arm is the fairest to the window and it still loses:
|
||||
Q=4, A=20: K=4 `2.3928` vs single `2.3557` (+0.037); Q=2, A=20: K=4 `2.4643` vs
|
||||
`2.4524` (+0.012).
|
||||
|
||||
**The extra history is not just useless, it is actively harmful**, and the reason
|
||||
is visible in the shuffle control below: the long ordered contexts are sparse, and
|
||||
the interpolation spends held-out probability on them.
|
||||
|
||||
The baseline is not trivial: the single state halves the majority floor
|
||||
(majority log-loss **2.6983**, acc **0.2348**; best single Q=4 **2.3460**, acc
|
||||
**0.4094**).
|
||||
|
||||
### 4.2 `fly` frame — window starts at the fire tick, ends at `t0+K−1` *(8 156 shots)*
|
||||
|
||||
Here the **single@D itself** improves steeply as K grows, because the decision tick
|
||||
is later and therefore closer to the answer. A = 5:
|
||||
|
||||
| K | Q=2 single@D → window (Δ) | Q=3 single@D → window (Δ) | Q=4 single@D → window (Δ) |
|
||||
|---|---|---|---|
|
||||
| 1 | 2.6645 → 2.6645 (—) | 2.6741 → 2.6741 (—) | 2.7020 → 2.7020 (—) |
|
||||
| 4 | 2.6376 → 2.7695 (**+0.132**) | 2.6262 → 2.8858 (**+0.260**) | 2.6397 → 2.9108 (**+0.271**) |
|
||||
| 8 | 2.5659 → 3.0331 (**+0.467**) | 2.5302 → 3.0207 (**+0.490**) | 2.6101 → 2.9084 (**+0.298**) |
|
||||
| 16 | 2.4047 → 3.1643 (**+0.760**) | 2.3103 → 2.6362 (**+0.326**) | 2.3853 → 2.5918 (**+0.207**) |
|
||||
| 32 | 1.8365 → 1.9970 (**+0.161**) | 1.4355 → 1.4625 (**+0.027**) | **1.3121 → 1.2668 (−0.045)** |
|
||||
|
||||
Read this as: **the gain is the later tick, not the window.** `single@D` alone drops
|
||||
from `2.70` at K=1 to `1.31` at K=32 (Q=4) — the single state at `t0+31` is 2
|
||||
ticks from arrival and already knows the answer. The window adds nothing on top;
|
||||
the only cell where it "wins" is K=32/Q=4 by **0.045 bits**, and that subset's
|
||||
32-state window **never repeats** (§5), so the model is running on the single
|
||||
state plus a trace of recent acceleration. Q=3 at the same K is *worse*, and
|
||||
K=4–16 are worse by 5–20× that margin.
|
||||
|
||||
### 4.3 The mandatory shuffle-order control *(MEASURED)*
|
||||
|
||||
The order of the K states inside each window is randomly permuted, preserving the
|
||||
multiset. In the `pre` frame the shuffled window is **worse** than the temporal
|
||||
window at K=4 (e.g. Q=4: `2.5672` vs `2.5656`… within noise) but **better** at
|
||||
K≥8 (Q=4/K=16: shuffled `2.5613` vs temporal `2.7450`). In the `fly` frame at
|
||||
K=32 the shuffled window is much worse (Δ = `+0.35/ +0.80/ +1.12`).
|
||||
|
||||
**How to read that, honestly.** The shuffle penalty at `fly` K=32 is **not**
|
||||
evidence the window works: the 32-state window never repeats, so the only order
|
||||
the model can actually use is the *recency* of the last state. Shuffling destroys
|
||||
*which state is current*, and the model degrades because recency matters — a
|
||||
statement about the **single state**, not the window. Where the shuffled window is
|
||||
*better* than the temporal one (pre frame, K≥8), it is because random order
|
||||
disables the sparse long contexts that the temporal order feeds to the model.
|
||||
Either way, the window itself contributes nothing.
|
||||
|
||||
A deterministic **reverse** control (reverse time, preserves recurrence) is worse
|
||||
than temporal everywhere, confirming the model is genuinely order-sensitive and
|
||||
that the last state is the informative one.
|
||||
|
||||
---
|
||||
|
||||
## 5. Recurrence — the "we can find coincidences" premise *(MEASURED)*
|
||||
|
||||
Distinct **ordered** window tuples over the whole corpus (54 939 shots; 8 156 in
|
||||
the `fly` subset). A window is *learnable by exact coincidence* only if its count
|
||||
is well above 1.
|
||||
|
||||
| K | Q=2 distinct (mean count) | Q=3 | Q=4 |
|
||||
|---|---|---|---|
|
||||
| 1 | 16 (3433) | 82 (670) | 372 (148) |
|
||||
| 4 | 1 747 (31.4) | 11 018 (5.0) | 24 223 (2.3) |
|
||||
| 8 | 13 474 (4.1) | 37 337 (1.5) | 49 891 (**1.1**) |
|
||||
| 16 | 41 447 (1.3) | 54 163 (1.01) | 54 707 (1.00) |
|
||||
| 32 | 54 454 (**1.01**) | 54 754 (1.00) | 54 779 (1.00) |
|
||||
| 48 | 54 744 (1.00) | 54 772 (1.00) | 54 789 (1.00) |
|
||||
|
||||
(In the `fly` subset the same collapse happens earlier and harder: Q=4, K=8 already
|
||||
1.02 mean count, and **K=32 is 8 156 distinct of 8 156 — literally every window is
|
||||
unique**.)
|
||||
|
||||
So the "long list with patterns inside" premise fails as a **combinatorial**
|
||||
matter, not merely a modelling one:
|
||||
|
||||
* At the coarsest state worth using (Q=2, 5 bits), only **K ≤ 8** has meaningful
|
||||
recurrence — and even there the window does not beat the single state.
|
||||
* At the finest (Q=4, 10 bits), K=4 already averages 2.3 observations; K≥8 is
|
||||
essentially all singletons. **Learning from a window that never recurs is
|
||||
impossible**, and the sweep shows it does not happen.
|
||||
* The single state is the only object with a real coincidence structure: 16–372
|
||||
distinct values, every one repeated hundreds of times.
|
||||
|
||||
---
|
||||
|
||||
## 6. Direct answer
|
||||
|
||||
**Does a window of wave-relative states predict DrussGT's future lateral position
|
||||
better than a single state on held-out battles? NO.**
|
||||
|
||||
* `pre` frame: the window is worse at **every** K, Q and A; at K=8 it costs
|
||||
**+0.35 to +0.46 bits** over the single state.
|
||||
* `fly` frame: the whole apparent improvement is the later decision tick
|
||||
(`single@D` 2.70 → 1.31 bits as K goes 1 → 32). The window adds nothing except
|
||||
a 0.045-bit sliver at K=32/Q=4 on a subset whose 32-windows never repeat.
|
||||
* The recurrence statistics say the same thing independently: re-usable
|
||||
"patterns" of movement do not recur at any window length that carries extra
|
||||
information.
|
||||
|
||||
**The design — feed SBC a long temporal list of states — is dead.** Do not build
|
||||
it. *(MEASURED)*
|
||||
|
||||
**What survives (and is worth keeping).** The **state definition itself is a
|
||||
strong single-state predictor**: the fire-tick state with 4 symbols/field
|
||||
(10 bits) halves the majority log-loss (2.6983 → **2.3460** bits) and triples the
|
||||
chance the predicted displacement bin is right (0.2348 → **0.4094**) on battles
|
||||
the model has never seen. If anything from this line is taken forward, take the
|
||||
**single** wave-relative state at Q=4 and use it as a context for a *different*
|
||||
decision problem — but note the implied hit probability it produces is still the
|
||||
base rate (~0.09), so it is informative about **which side** the miss falls on, not
|
||||
about **whether** the shot hits. *(MEASURED/INFERRED)*
|
||||
|
||||
---
|
||||
|
||||
## 7. MEASURED vs INFERRED
|
||||
|
||||
**MEASURED** (exact commands in §8; numbers above)
|
||||
|
||||
* 70 battles / 490 rounds / 54 939 shots; 8 156 with `karr ≥ 33`.
|
||||
* Every held-out log-loss / accuracy / implied-hit number in §4, over 3
|
||||
battle-split seeds, with bin edges from train only.
|
||||
* The shuffle-order and reverse-order controls.
|
||||
* The distinct-window and repeat counts in §5.
|
||||
|
||||
**INFERRED**
|
||||
|
||||
* That the residual fly-frame "win" at K=32/Q=4 is a trace of recent acceleration
|
||||
rather than a movement pattern — it is consistent with the never-repeating
|
||||
32-window and with the single state already containing `vlat`.
|
||||
* That a *different* (e.g. TSetlin/SBC) learner would not reverse the answer. This
|
||||
is not proved. What **is** proved is the recurrence floor: a coincidence-based
|
||||
method cannot learn from windows that never recur, and by Q=4/K=8 they already
|
||||
essentially never do.
|
||||
* "Design is dead" applies to the **long temporal window**. It does not say the
|
||||
state is useless, and it does not say a *short* (K=2) context is worthless — K=2
|
||||
was not separately swept because at Q≥3 it is already 1.5–5 observations per
|
||||
context, i.e. below the point where it could help.
|
||||
|
||||
**Caveats.** (1) The `fly` frame is a biased long-range subset (all shots with
|
||||
`flight ≥ 33`, i.e. the longest-range quarter). (2) The model's context depth is
|
||||
capped at 12; K=16/32/48 therefore share a row, but §5 shows orders above ~4 are
|
||||
never repeated often enough to matter, so the cap is not the binding constraint.
|
||||
(3) The target is the miss offset, not the hit; the implied hit probability is the
|
||||
base rate for every model, so nothing here is a hit-rate claim.
|
||||
|
||||
---
|
||||
|
||||
## 8. Reproducing
|
||||
|
||||
```bash
|
||||
# corpus (live captures; not in the repo)
|
||||
ls /tmp/tfil_ab2/out # 70 battles: <A..E>/runN.jsonl{,.events.jsonl,.rounds.json}
|
||||
|
||||
python3 common_libs/tests/state_window_gate.py \
|
||||
--tfil /tmp/tfil_ab2/out --seeds 3 --frames pre,fly \
|
||||
--json common_libs/tests/fixtures/state_window_gate.json \
|
||||
> common_libs/tests/fixtures/state_window_gate_report.txt
|
||||
```
|
||||
|
||||
Runtime ≈ **5 minutes**, single-threaded pure Python (no numpy required). The
|
||||
script imports `common_libs/tests/analyze_drussgt_dodge_vs_power.py` for the
|
||||
per-shot geometry; `common_libs/tests/fixtures/state_window_gate_report.txt` is
|
||||
the verbatim captured output and `…state_window_gate.json` is the same numbers as
|
||||
JSON (including the per-seed values). No offline fixture replay, no simulator:
|
||||
these are the recorded real battles.
|
||||
Reference in New Issue
Block a user