Files
SirRoboGarage/docs/state_window_gate.md
T
SirStone d85ff53d34 State-window gate: a temporal window of wave-relative states does NOT beat a single state
Re-runs the SBC coincidence premise as a cheap veto test on the 70-battle
live-vs-DrussGT corpus (/tmp/tfil_ab2).  Defines the wave-relative state
(lat/vlat/toa/room/turn, 5/7.9/10 bits at Q=2/3/4), quantises the miss offset
at the bullet's arrival into 7 bins, and sweeps window length K in
{1,4,8,16,32,48} with an interpolated suffix-backoff model under a BY-BATTLE
70/30 split (3 seeds).

Result: NO.  On the pre-fire frame the window is worse than the single
fire-tick state at every K/Q/A (e.g. K=8 costs +0.35..+0.46 bits).  On the
during-flight frame the entire apparent gain is the later decision tick, not
the window; the single state alone drops 2.70 -> 1.31 bits as K goes 1 -> 32.
The shuffle-order control confirms recency matters but the windows do not:
by Q=4/K=8 they average ~1 observation and never recur.  The single state
survives as a strong predictor (log-loss 2.346 vs 2.698 majority; bin
accuracy 0.409 vs 0.235).

Gate only: no gun, no live-win claim.
2026-09-25 08:43:36 +02:00

16 KiB
Raw Blame History

The state-window gate: does a temporal WINDOW of wave-relative states beat a SINGLE state?

The design under test (the user's words). "We decide first what makes an 'environment state' ... then we feed as input a long list of these states, long enough to have one or more patterns of movement inside." The premise is that SBC is a coincidence detector and that a temporal list of states, fed instead of a single tick, turns "coincidence" into "movement pattern".

Verdict: NO. On held-out battles a window of wave-relative states does not predict DrussGT's lateral position at the bullet's arrival better than the single state at the same decision tick. On the pre-fire frame the window is strictly and substantially worse at every window length, every coarseness and every smoothing strength. On the during-flight frame the apparent gain is entirely the later decision tick, not the window — and where the window finally edges the single state (K=32, Q=4) it is by 0.045 bits on a subset where the 32-window never repeats at all. The design, as stated, is dead. The state definition survives and is the useful result.

This is a gate, not a gun. No live win is claimed and none is measured here.


1. The state (design artifact)

The dodge is a response to our bullet, so the state is wave-relative, never absolute world coordinates. One state at absolute tick t, in the frame of the bullet fired at t0 along u = (cos dir, sin dir) (with n = (-u_y, u_x) the lateral unit vector):

field definition why
lat (D(t) − P0) × u — DrussGT's signed perpendicular offset from the bullet line, px the quantity that decides the hit; signed so "which side" is visible
vlat lat(t) − lat(t−1) — signed lateral velocity, px/tick crossing vs returning
toa (t0 − t) + karr — ticks until the bullet reaches its arrival plane how much time is left to move
room ray distance from D(t) along sign(vlat)·n until the arena wall (18 px bot radius) room to keep running vs being cornered — the only directional wall field, unlike plain "distance to nearest wall"
turn wrap180(eh(t) − eh(t−1)) — signed turn rate, deg/tick surfers slow to turn; turning reveals a reversal

Refinements to the starting proposal, and why. (a) was kept signed, not absolutised, because the target is signed. (b) was made a difference of the lateral offset, so it is exactly the lateral component of velocity in the bullet frame rather than a speed. (d) "distance to the wall" was made directional (room to run in the current lateral-motion direction); a plain nearest-wall distance is not wave-relative and is nearly constant at long range. (e) "heading/turn rate" was reduced to the turn rate alone, because heading relative to the bullet is already implied by the sign of vlat, and the turn rate is the thing that distinguishes a surf from a corner.

Quantisation (the numerosity dial). Each field is binned into Q ∈ {2, 3, 4} symbols. A state is therefore:

Q bits / state distinct states
2 5.0 32
3 7.9 243
4 10.0 1024

Bin edges are quantiles of the TRAIN side only (no test leakage); a coarse state makes two similar situations look the same, which is what lets a "coincidence" repeat. The sweep below is exactly the question of whether coarser buys more recurrence than it loses in resolution.

The two windows tested, and why both

A window is K consecutive states. Crucially, the baseline is always the single state at the same decision tick D — otherwise a longer window would "win" only because its last state is closer to the answer.

  • pre — D = t0 (the fire tick); the window is the K pre-fire states ending at D. This is what a bot can compute before pulling the trigger. All K ≤ 48 are feasible (max recorded flight is 43 ticks).
  • fly — D = t0+K−1; the window is the first K states of the flight, and the baseline is the single state at D. Only shots with karr > 31 are used, so the whole window is strictly before arrival (8156 of 54939 shots).

The task's field toa only exists once a bullet is in the air, so fly is the "reaction" reading; pre is the "am I about to be dodged?" reading. Both are reported.

The target

perp_arr — DrussGT's signed perpendicular offset from the bullet line at the tick the bullet reaches its along-track plane. This is the miss offset that decides the hit; unlike "required lead" it does not depend on the bullet speed. It is quantised into 7 bins with edges ±120, ±60, ±18 px; the central ±18 px bin is the hit window (bot radius 18 px).

Bin width in degrees-at-450px. The central hit bin is 36 px wide = one bot diameter. atan(18/range) is the hit half-window: 2.29° at 450 px, 4.23° at the median fire range (487 px), and 20.4° at 100 px. So the 36 px central bin is 4.58° at 450 px. The three inner bins are 36/42/60 px wide (4.58/5.34/7.63° at 450 px) — the target resolution is finer than the ~2.29° hit half-window, so a model that predicts the bin is not being asked an impossible question.


2. Data and split

  • Data (MEASURED present before use): /tmp/tfil_ab2/out/<A..E>/runN.jsonl + .events.jsonl + .rounds.json — 70 real live battles vs the unmodified DrussGT, 490 rounds, 54 939 shots by ModularBot, recorded by tools/robocode_shim/run_bridge_battle.sh. Per-shot geometry is re-derived by the validated instrument common_libs/tests/analyze_drussgt_dodge_vs_power.py (which this extractor imports). Attribution (e*=DrussGT, s*=us) is the one already documented in docs/drussgt_dodge_vs_power.md.
  • Split (stated explicitly): BY BATTLE, never by tick. All rounds of a battle go to one side. A 70 %/30 % battle split (49 train / 21 test battles), repeated over 3 seeds. Bin edges are derived from the train side only. Two ticks inside one round never straddle the split.

3. Model (deliberately dull)

The question is about information, not modelling, so the model is a plain interpolated (Jelinek–Mercer) suffix-backoff table — the direct analogue of SBC "counting coincidences":

P = global target-bin histogram
for j = 1..K:   P <- (count(suffix_j) + A*P) / (total(suffix_j) + A)

It is order-sensitive, it can never do much worse than the shorter context, and it degrades gracefully when a suffix is unseen. Context depth is capped at 12 (orders above that are never repeated often enough to matter — see §5). A ∈ {1,5,20} is swept; A = 5 is primary. A majority/no-window predictor is the floor. The model predicts the 7-bin displacement distribution; we report held-out accuracy (argmax), log-loss in bits, and the implied hit probability P(central bin).


4. Result — does the window beat the single state? No (MEASURED)

70 battles, 3 battle-split seeds, mean held-out log-loss (bits). delta = window − single@D; negative would mean the window wins.

4.1 pre frame — window ends at the fire tick, looks back (54 939 shots)

single@D is the fire-tick state; it does not change with K. A = 5:

K Q=2 single → window (Δ) Q=3 single → window (Δ) Q=4 single → window (Δ)
1 2.4524 → 2.4524 (—) 2.3782 → 2.3782 (—) 2.3460 → 2.3460 (—)
4 2.4524 → 2.5107 (+0.058) 2.3782 → 2.5461 (+0.168) 2.3460 → 2.5656 (+0.220)
8 2.4524 → 2.8015 (+0.349) 2.3782 → 2.8336 (+0.455) 2.3460 → 2.7265 (+0.381)
16 2.4524 → 3.1310 (+0.679) 2.3782 → 2.9473 (+0.569) 2.3460 → 2.7450 (+0.399)
32 same as K=16 (depth cap) same same
48 same as K=16 (depth cap) same same

Accuracy tells the same story (Q=4: single 0.409 → window 0.383/0.364/ 0.362). The per-seed deltas are essentially identical (e.g. Q=2/K=4: +0.0572, +0.0534, +0.0644) — there is no seed where the window helps. The heavy-smoothing arm is the fairest to the window and it still loses: Q=4, A=20: K=4 2.3928 vs single 2.3557 (+0.037); Q=2, A=20: K=4 2.4643 vs 2.4524 (+0.012).

The extra history is not just useless, it is actively harmful, and the reason is visible in the shuffle control below: the long ordered contexts are sparse, and the interpolation spends held-out probability on them.

The baseline is not trivial: the single state halves the majority floor (majority log-loss 2.6983, acc 0.2348; best single Q=4 2.3460, acc 0.4094).

4.2 fly frame — window starts at the fire tick, ends at t0+K−1 (8 156 shots)

Here the single@D itself improves steeply as K grows, because the decision tick is later and therefore closer to the answer. A = 5:

K Q=2 single@D → window (Δ) Q=3 single@D → window (Δ) Q=4 single@D → window (Δ)
1 2.6645 → 2.6645 (—) 2.6741 → 2.6741 (—) 2.7020 → 2.7020 (—)
4 2.6376 → 2.7695 (+0.132) 2.6262 → 2.8858 (+0.260) 2.6397 → 2.9108 (+0.271)
8 2.5659 → 3.0331 (+0.467) 2.5302 → 3.0207 (+0.490) 2.6101 → 2.9084 (+0.298)
16 2.4047 → 3.1643 (+0.760) 2.3103 → 2.6362 (+0.326) 2.3853 → 2.5918 (+0.207)
32 1.8365 → 1.9970 (+0.161) 1.4355 → 1.4625 (+0.027) 1.3121 → 1.2668 (−0.045)

Read this as: the gain is the later tick, not the window. single@D alone drops from 2.70 at K=1 to 1.31 at K=32 (Q=4) — the single state at t0+31 is 2 ticks from arrival and already knows the answer. The window adds nothing on top; the only cell where it "wins" is K=32/Q=4 by 0.045 bits, and that subset's 32-state window never repeats (§5), so the model is running on the single state plus a trace of recent acceleration. Q=3 at the same K is worse, and K=4–16 are worse by 5–20× that margin.

4.3 The mandatory shuffle-order control (MEASURED)

The order of the K states inside each window is randomly permuted, preserving the multiset. In the pre frame the shuffled window is worse than the temporal window at K=4 (e.g. Q=4: 2.5672 vs 2.5656… within noise) but better at K≥8 (Q=4/K=16: shuffled 2.5613 vs temporal 2.7450). In the fly frame at K=32 the shuffled window is much worse (Δ = +0.35/ +0.80/ +1.12).

How to read that, honestly. The shuffle penalty at fly K=32 is not evidence the window works: the 32-state window never repeats, so the only order the model can actually use is the recency of the last state. Shuffling destroys which state is current, and the model degrades because recency matters — a statement about the single state, not the window. Where the shuffled window is better than the temporal one (pre frame, K≥8), it is because random order disables the sparse long contexts that the temporal order feeds to the model. Either way, the window itself contributes nothing.

A deterministic reverse control (reverse time, preserves recurrence) is worse than temporal everywhere, confirming the model is genuinely order-sensitive and that the last state is the informative one.


5. Recurrence — the "we can find coincidences" premise (MEASURED)

Distinct ordered window tuples over the whole corpus (54 939 shots; 8 156 in the fly subset). A window is learnable by exact coincidence only if its count is well above 1.

K Q=2 distinct (mean count) Q=3 Q=4
1 16 (3433) 82 (670) 372 (148)
4 1 747 (31.4) 11 018 (5.0) 24 223 (2.3)
8 13 474 (4.1) 37 337 (1.5) 49 891 (1.1)
16 41 447 (1.3) 54 163 (1.01) 54 707 (1.00)
32 54 454 (1.01) 54 754 (1.00) 54 779 (1.00)
48 54 744 (1.00) 54 772 (1.00) 54 789 (1.00)

(In the fly subset the same collapse happens earlier and harder: Q=4, K=8 already 1.02 mean count, and K=32 is 8 156 distinct of 8 156 — literally every window is unique.)

So the "long list with patterns inside" premise fails as a combinatorial matter, not merely a modelling one:

  • At the coarsest state worth using (Q=2, 5 bits), only K ≤ 8 has meaningful recurrence — and even there the window does not beat the single state.
  • At the finest (Q=4, 10 bits), K=4 already averages 2.3 observations; K≥8 is essentially all singletons. Learning from a window that never recurs is impossible, and the sweep shows it does not happen.
  • The single state is the only object with a real coincidence structure: 16–372 distinct values, every one repeated hundreds of times.

6. Direct answer

Does a window of wave-relative states predict DrussGT's future lateral position better than a single state on held-out battles? NO.

  • pre frame: the window is worse at every K, Q and A; at K=8 it costs +0.35 to +0.46 bits over the single state.
  • fly frame: the whole apparent improvement is the later decision tick (single@D 2.70 → 1.31 bits as K goes 1 → 32). The window adds nothing except a 0.045-bit sliver at K=32/Q=4 on a subset whose 32-windows never repeat.
  • The recurrence statistics say the same thing independently: re-usable "patterns" of movement do not recur at any window length that carries extra information.

The design — feed SBC a long temporal list of states — is dead. Do not build it. (MEASURED)

What survives (and is worth keeping). The state definition itself is a strong single-state predictor: the fire-tick state with 4 symbols/field (10 bits) halves the majority log-loss (2.6983 → 2.3460 bits) and triples the chance the predicted displacement bin is right (0.2348 → 0.4094) on battles the model has never seen. If anything from this line is taken forward, take the single wave-relative state at Q=4 and use it as a context for a different decision problem — but note the implied hit probability it produces is still the base rate (~0.09), so it is informative about which side the miss falls on, not about whether the shot hits. (MEASURED/INFERRED)


7. MEASURED vs INFERRED

MEASURED (exact commands in §8; numbers above)

  • 70 battles / 490 rounds / 54 939 shots; 8 156 with karr ≥ 33.
  • Every held-out log-loss / accuracy / implied-hit number in §4, over 3 battle-split seeds, with bin edges from train only.
  • The shuffle-order and reverse-order controls.
  • The distinct-window and repeat counts in §5.

INFERRED

  • That the residual fly-frame "win" at K=32/Q=4 is a trace of recent acceleration rather than a movement pattern — it is consistent with the never-repeating 32-window and with the single state already containing vlat.
  • That a different (e.g. TSetlin/SBC) learner would not reverse the answer. This is not proved. What is proved is the recurrence floor: a coincidence-based method cannot learn from windows that never recur, and by Q=4/K=8 they already essentially never do.
  • "Design is dead" applies to the long temporal window. It does not say the state is useless, and it does not say a short (K=2) context is worthless — K=2 was not separately swept because at Q≥3 it is already 1.5–5 observations per context, i.e. below the point where it could help.

Caveats. (1) The fly frame is a biased long-range subset (all shots with flight ≥ 33, i.e. the longest-range quarter). (2) The model's context depth is capped at 12; K=16/32/48 therefore share a row, but §5 shows orders above ~4 are never repeated often enough to matter, so the cap is not the binding constraint. (3) The target is the miss offset, not the hit; the implied hit probability is the base rate for every model, so nothing here is a hit-rate claim.


8. Reproducing

# corpus (live captures; not in the repo)
ls /tmp/tfil_ab2/out            # 70 battles: <A..E>/runN.jsonl{,.events.jsonl,.rounds.json}

python3 common_libs/tests/state_window_gate.py \
    --tfil /tmp/tfil_ab2/out --seeds 3 --frames pre,fly \
    --json common_libs/tests/fixtures/state_window_gate.json \
    > common_libs/tests/fixtures/state_window_gate_report.txt

Runtime ≈ 5 minutes, single-threaded pure Python (no numpy required). The script imports common_libs/tests/analyze_drussgt_dodge_vs_power.py for the per-shot geometry; common_libs/tests/fixtures/state_window_gate_report.txt is the verbatim captured output and …state_window_gate.json is the same numbers as JSON (including the per-seed values). No offline fixture replay, no simulator: these are the recorded real battles.