Files
SirRoboGarage/docs/state_window_gate.md
SirStone d85ff53d34 State-window gate: a temporal window of wave-relative states does NOT beat a single state
Re-runs the SBC coincidence premise as a cheap veto test on the 70-battle
live-vs-DrussGT corpus (/tmp/tfil_ab2).  Defines the wave-relative state
(lat/vlat/toa/room/turn, 5/7.9/10 bits at Q=2/3/4), quantises the miss offset
at the bullet's arrival into 7 bins, and sweeps window length K in
{1,4,8,16,32,48} with an interpolated suffix-backoff model under a BY-BATTLE
70/30 split (3 seeds).

Result: NO.  On the pre-fire frame the window is worse than the single
fire-tick state at every K/Q/A (e.g. K=8 costs +0.35..+0.46 bits).  On the
during-flight frame the entire apparent gain is the later decision tick, not
the window; the single state alone drops 2.70 -> 1.31 bits as K goes 1 -> 32.
The shuffle-order control confirms recency matters but the windows do not:
by Q=4/K=8 they average ~1 observation and never recur.  The single state
survives as a strong predictor (log-loss 2.346 vs 2.698 majority; bin
accuracy 0.409 vs 0.235).

Gate only: no gun, no live-win claim.
2026-09-25 08:43:36 +02:00

322 lines
16 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# The state-window gate: does a temporal WINDOW of wave-relative states beat a SINGLE state?
**The design under test (the user's words).** *"We decide first what makes an
'environment state' ... then we feed as input a long list of these states, long
enough to have one or more patterns of movement inside."* The premise is that SBC
is a coincidence detector and that a **temporal list** of states, fed instead of a
single tick, turns "coincidence" into "movement pattern".
**Verdict: NO.** On held-out battles a window of wave-relative states **does not**
predict DrussGT's lateral position at the bullet's arrival better than the single
state at the **same decision tick**. On the pre-fire frame the window is strictly
and substantially *worse* at every window length, every coarseness and every
smoothing strength. On the during-flight frame the apparent gain is entirely the
**later decision tick**, not the window — and where the window finally edges the
single state (K=32, Q=4) it is by 0.045 bits on a subset where the 32-window
**never repeats at all**. The design, as stated, is **dead**. The *state
definition* survives and is the useful result.
This is a gate, not a gun. **No live win is claimed and none is measured here.**
---
## 1. The state (design artifact)
The dodge is a response to **our bullet**, so the state is **wave-relative**, never
absolute world coordinates. One state at absolute tick `t`, in the frame of the
bullet fired at `t0` along `u = (cos dir, sin dir)` (with `n = (-u_y, u_x)` the
lateral unit vector):
| field | definition | why |
|---|---|---|
| `lat` | `(D(t) − P0) × u` — DrussGT's signed perpendicular offset from the bullet line, px | the quantity that decides the hit; signed so "which side" is visible |
| `vlat` | `lat(t) − lat(t−1)` — signed lateral velocity, px/tick | crossing vs returning |
| `toa` | `(t0 − t) + karr` — ticks until the bullet reaches its arrival plane | how much time is left to move |
| `room` | ray distance from `D(t)` along `sign(vlat)·n` until the arena wall (18 px bot radius) | room to keep running vs being cornered — the only *directional* wall field, unlike plain "distance to nearest wall" |
| `turn` | `wrap180(eh(t) − eh(t−1))` — signed turn rate, deg/tick | surfers slow to turn; turning reveals a reversal |
**Refinements to the starting proposal, and why.** (a) was kept signed, not
absolutised, because the target is signed. (b) was made a *difference of the
lateral offset*, so it is exactly the lateral component of velocity in the bullet
frame rather than a speed. (d) "distance to the wall" was made *directional*
(room to run in the current lateral-motion direction); a plain nearest-wall
distance is not wave-relative and is nearly constant at long range. (e)
"heading/turn rate" was reduced to the **turn rate** alone, because heading
relative to the bullet is already implied by the sign of `vlat`, and the turn rate
is the thing that distinguishes a surf from a corner.
**Quantisation (the numerosity dial).** Each field is binned into `Q ∈ {2, 3, 4}`
symbols. A state is therefore:
| Q | bits / state | distinct states |
|---|---|---|
| 2 | **5.0** | 32 |
| 3 | **7.9** | 243 |
| 4 | **10.0** | 1024 |
Bin edges are **quantiles of the TRAIN side only** (no test leakage); a coarse
state makes two similar situations look the same, which is what lets a
"coincidence" repeat. The sweep below is exactly the question of whether coarser
buys more recurrence than it loses in resolution.
### The two windows tested, and why both
A window is `K` consecutive states. Crucially, **the baseline is always the single
state at the same decision tick `D`** — otherwise a longer window would "win" only
because its last state is closer to the answer.
* **`pre`** — `D = t0` (the fire tick); the window is the `K` pre-fire states
ending at `D`. This is what a bot can compute *before* pulling the trigger. All
`K ≤ 48` are feasible (max recorded flight is 43 ticks).
* **`fly`** — `D = t0+K−1`; the window is the first `K` states of the flight, and
the baseline is the single state at `D`. Only shots with `karr > 31` are used,
so the whole window is strictly before arrival (8156 of 54939 shots).
The task's field `toa` only exists once a bullet is in the air, so `fly` is the
"reaction" reading; `pre` is the "am I about to be dodged?" reading. Both are
reported.
### The target
`perp_arr` — DrussGT's **signed** perpendicular offset from the bullet line at the
tick the bullet reaches its along-track plane. This is the miss offset that
decides the hit; unlike "required lead" it does not depend on the bullet speed. It
is quantised into 7 bins with edges `±120, ±60, ±18` px; the central `±18` px bin
is the hit window (bot radius 18 px).
**Bin width in degrees-at-450px.** The central hit bin is 36 px wide = one bot
diameter. `atan(18/range)` is the hit half-window: **2.29°** at 450 px, 4.23° at
the median fire range (487 px), and 20.4° at 100 px. So the 36 px central bin is
**4.58° at 450 px**. The three inner bins are 36/42/60 px wide (4.58/5.34/7.63° at
450 px) — the target resolution is *finer* than the ~2.29° hit half-window, so a
model that predicts the bin is not being asked an impossible question.
---
## 2. Data and split
* **Data (MEASURED present before use):** `/tmp/tfil_ab2/out/<A..E>/runN.jsonl` +
`.events.jsonl` + `.rounds.json` — **70 real live battles vs the unmodified
DrussGT, 490 rounds, 54 939 shots** by ModularBot, recorded by
`tools/robocode_shim/run_bridge_battle.sh`. Per-shot geometry is re-derived by
the validated instrument `common_libs/tests/analyze_drussgt_dodge_vs_power.py`
(which this extractor imports). Attribution (`e*`=DrussGT, `s*`=us) is the one
already documented in `docs/drussgt_dodge_vs_power.md`.
* **Split (stated explicitly): BY BATTLE, never by tick.** All rounds of a battle
go to one side. A 70 %/30 % battle split (**49 train / 21 test battles**),
repeated over **3 seeds**. Bin edges are derived from the train side only. Two
ticks inside one round never straddle the split.
---
## 3. Model (deliberately dull)
The question is about **information**, not modelling, so the model is a plain
**interpolated (Jelinek–Mercer) suffix-backoff table** — the direct analogue of
SBC "counting coincidences":
```
P = global target-bin histogram
for j = 1..K: P <- (count(suffix_j) + A*P) / (total(suffix_j) + A)
```
It is order-sensitive, it can never do much worse than the shorter context, and it
degrades gracefully when a suffix is unseen. Context depth is capped at 12 (orders
above that are never repeated often enough to matter — see §5). `A ∈ {1,5,20}` is
swept; **A = 5 is primary**. A majority/no-window predictor is the floor. The
model predicts the 7-bin displacement distribution; we report held-out **accuracy**
(argmax), **log-loss** in bits, and the **implied hit probability** `P(central
bin)`.
---
## 4. Result — does the window beat the single state? **No** *(MEASURED)*
70 battles, 3 battle-split seeds, mean held-out log-loss (bits). `delta = window −
single@D`; **negative would mean the window wins**.
### 4.1 `pre` frame — window ends at the fire tick, looks back *(54 939 shots)*
`single@D` is the fire-tick state; it does **not** change with K. A = 5:
| K | Q=2 single → window (Δ) | Q=3 single → window (Δ) | Q=4 single → window (Δ) |
|---|---|---|---|
| 1 | 2.4524 → 2.4524 (—) | 2.3782 → 2.3782 (—) | 2.3460 → 2.3460 (—) |
| 4 | 2.4524 → 2.5107 (**+0.058**) | 2.3782 → 2.5461 (**+0.168**) | 2.3460 → 2.5656 (**+0.220**) |
| 8 | 2.4524 → 2.8015 (**+0.349**) | 2.3782 → 2.8336 (**+0.455**) | 2.3460 → 2.7265 (**+0.381**) |
| 16 | 2.4524 → 3.1310 (**+0.679**) | 2.3782 → 2.9473 (**+0.569**) | 2.3460 → 2.7450 (**+0.399**) |
| 32 | same as K=16 (depth cap) | same | same |
| 48 | same as K=16 (depth cap) | same | same |
Accuracy tells the same story (Q=4: single `0.409` → window `0.383`/`0.364`/
`0.362`). The per-seed deltas are essentially identical (e.g. Q=2/K=4:
`+0.0572, +0.0534, +0.0644`) — there is no seed where the window helps. The
**heavy-smoothing** arm is the fairest to the window and it still loses:
Q=4, A=20: K=4 `2.3928` vs single `2.3557` (+0.037); Q=2, A=20: K=4 `2.4643` vs
`2.4524` (+0.012).
**The extra history is not just useless, it is actively harmful**, and the reason
is visible in the shuffle control below: the long ordered contexts are sparse, and
the interpolation spends held-out probability on them.
The baseline is not trivial: the single state halves the majority floor
(majority log-loss **2.6983**, acc **0.2348**; best single Q=4 **2.3460**, acc
**0.4094**).
### 4.2 `fly` frame — window starts at the fire tick, ends at `t0+K−1` *(8 156 shots)*
Here the **single@D itself** improves steeply as K grows, because the decision tick
is later and therefore closer to the answer. A = 5:
| K | Q=2 single@D → window (Δ) | Q=3 single@D → window (Δ) | Q=4 single@D → window (Δ) |
|---|---|---|---|
| 1 | 2.6645 → 2.6645 (—) | 2.6741 → 2.6741 (—) | 2.7020 → 2.7020 (—) |
| 4 | 2.6376 → 2.7695 (**+0.132**) | 2.6262 → 2.8858 (**+0.260**) | 2.6397 → 2.9108 (**+0.271**) |
| 8 | 2.5659 → 3.0331 (**+0.467**) | 2.5302 → 3.0207 (**+0.490**) | 2.6101 → 2.9084 (**+0.298**) |
| 16 | 2.4047 → 3.1643 (**+0.760**) | 2.3103 → 2.6362 (**+0.326**) | 2.3853 → 2.5918 (**+0.207**) |
| 32 | 1.8365 → 1.9970 (**+0.161**) | 1.4355 → 1.4625 (**+0.027**) | **1.3121 → 1.2668 (−0.045)** |
Read this as: **the gain is the later tick, not the window.** `single@D` alone drops
from `2.70` at K=1 to `1.31` at K=32 (Q=4) — the single state at `t0+31` is 2
ticks from arrival and already knows the answer. The window adds nothing on top;
the only cell where it "wins" is K=32/Q=4 by **0.045 bits**, and that subset's
32-state window **never repeats** (§5), so the model is running on the single
state plus a trace of recent acceleration. Q=3 at the same K is *worse*, and
K=4–16 are worse by 5–20× that margin.
### 4.3 The mandatory shuffle-order control *(MEASURED)*
The order of the K states inside each window is randomly permuted, preserving the
multiset. In the `pre` frame the shuffled window is **worse** than the temporal
window at K=4 (e.g. Q=4: `2.5672` vs `2.5656`… within noise) but **better** at
K≥8 (Q=4/K=16: shuffled `2.5613` vs temporal `2.7450`). In the `fly` frame at
K=32 the shuffled window is much worse (Δ = `+0.35/ +0.80/ +1.12`).
**How to read that, honestly.** The shuffle penalty at `fly` K=32 is **not**
evidence the window works: the 32-state window never repeats, so the only order
the model can actually use is the *recency* of the last state. Shuffling destroys
*which state is current*, and the model degrades because recency matters — a
statement about the **single state**, not the window. Where the shuffled window is
*better* than the temporal one (pre frame, K≥8), it is because random order
disables the sparse long contexts that the temporal order feeds to the model.
Either way, the window itself contributes nothing.
A deterministic **reverse** control (reverse time, preserves recurrence) is worse
than temporal everywhere, confirming the model is genuinely order-sensitive and
that the last state is the informative one.
---
## 5. Recurrence — the "we can find coincidences" premise *(MEASURED)*
Distinct **ordered** window tuples over the whole corpus (54 939 shots; 8 156 in
the `fly` subset). A window is *learnable by exact coincidence* only if its count
is well above 1.
| K | Q=2 distinct (mean count) | Q=3 | Q=4 |
|---|---|---|---|
| 1 | 16 (3433) | 82 (670) | 372 (148) |
| 4 | 1 747 (31.4) | 11 018 (5.0) | 24 223 (2.3) |
| 8 | 13 474 (4.1) | 37 337 (1.5) | 49 891 (**1.1**) |
| 16 | 41 447 (1.3) | 54 163 (1.01) | 54 707 (1.00) |
| 32 | 54 454 (**1.01**) | 54 754 (1.00) | 54 779 (1.00) |
| 48 | 54 744 (1.00) | 54 772 (1.00) | 54 789 (1.00) |
(In the `fly` subset the same collapse happens earlier and harder: Q=4, K=8 already
1.02 mean count, and **K=32 is 8 156 distinct of 8 156 — literally every window is
unique**.)
So the "long list with patterns inside" premise fails as a **combinatorial**
matter, not merely a modelling one:
* At the coarsest state worth using (Q=2, 5 bits), only **K ≤ 8** has meaningful
recurrence — and even there the window does not beat the single state.
* At the finest (Q=4, 10 bits), K=4 already averages 2.3 observations; K≥8 is
essentially all singletons. **Learning from a window that never recurs is
impossible**, and the sweep shows it does not happen.
* The single state is the only object with a real coincidence structure: 16–372
distinct values, every one repeated hundreds of times.
---
## 6. Direct answer
**Does a window of wave-relative states predict DrussGT's future lateral position
better than a single state on held-out battles? NO.**
* `pre` frame: the window is worse at **every** K, Q and A; at K=8 it costs
**+0.35 to +0.46 bits** over the single state.
* `fly` frame: the whole apparent improvement is the later decision tick
(`single@D` 2.70 → 1.31 bits as K goes 1 → 32). The window adds nothing except
a 0.045-bit sliver at K=32/Q=4 on a subset whose 32-windows never repeat.
* The recurrence statistics say the same thing independently: re-usable
"patterns" of movement do not recur at any window length that carries extra
information.
**The design — feed SBC a long temporal list of states — is dead.** Do not build
it. *(MEASURED)*
**What survives (and is worth keeping).** The **state definition itself is a
strong single-state predictor**: the fire-tick state with 4 symbols/field
(10 bits) halves the majority log-loss (2.6983 → **2.3460** bits) and triples the
chance the predicted displacement bin is right (0.2348 → **0.4094**) on battles
the model has never seen. If anything from this line is taken forward, take the
**single** wave-relative state at Q=4 and use it as a context for a *different*
decision problem — but note the implied hit probability it produces is still the
base rate (~0.09), so it is informative about **which side** the miss falls on, not
about **whether** the shot hits. *(MEASURED/INFERRED)*
---
## 7. MEASURED vs INFERRED
**MEASURED** (exact commands in §8; numbers above)
* 70 battles / 490 rounds / 54 939 shots; 8 156 with `karr ≥ 33`.
* Every held-out log-loss / accuracy / implied-hit number in §4, over 3
battle-split seeds, with bin edges from train only.
* The shuffle-order and reverse-order controls.
* The distinct-window and repeat counts in §5.
**INFERRED**
* That the residual fly-frame "win" at K=32/Q=4 is a trace of recent acceleration
rather than a movement pattern — it is consistent with the never-repeating
32-window and with the single state already containing `vlat`.
* That a *different* (e.g. TSetlin/SBC) learner would not reverse the answer. This
is not proved. What **is** proved is the recurrence floor: a coincidence-based
method cannot learn from windows that never recur, and by Q=4/K=8 they already
essentially never do.
* "Design is dead" applies to the **long temporal window**. It does not say the
state is useless, and it does not say a *short* (K=2) context is worthless — K=2
was not separately swept because at Q≥3 it is already 1.5–5 observations per
context, i.e. below the point where it could help.
**Caveats.** (1) The `fly` frame is a biased long-range subset (all shots with
`flight ≥ 33`, i.e. the longest-range quarter). (2) The model's context depth is
capped at 12; K=16/32/48 therefore share a row, but §5 shows orders above ~4 are
never repeated often enough to matter, so the cap is not the binding constraint.
(3) The target is the miss offset, not the hit; the implied hit probability is the
base rate for every model, so nothing here is a hit-rate claim.
---
## 8. Reproducing
```bash
# corpus (live captures; not in the repo)
ls /tmp/tfil_ab2/out # 70 battles: <A..E>/runN.jsonl{,.events.jsonl,.rounds.json}
python3 common_libs/tests/state_window_gate.py \
--tfil /tmp/tfil_ab2/out --seeds 3 --frames pre,fly \
--json common_libs/tests/fixtures/state_window_gate.json \
> common_libs/tests/fixtures/state_window_gate_report.txt
```
Runtime ≈ **5 minutes**, single-threaded pure Python (no numpy required). The
script imports `common_libs/tests/analyze_drussgt_dodge_vs_power.py` for the
per-shot geometry; `common_libs/tests/fixtures/state_window_gate_report.txt` is
the verbatim captured output and `…state_window_gate.json` is the same numbers as
JSON (including the per-seed values). No offline fixture replay, no simulator:
these are the recorded real battles.