State-window gate: a temporal window of wave-relative states does NOT beat a single state
Re-runs the SBC coincidence premise as a cheap veto test on the 70-battle
live-vs-DrussGT corpus (/tmp/tfil_ab2). Defines the wave-relative state
(lat/vlat/toa/room/turn, 5/7.9/10 bits at Q=2/3/4), quantises the miss offset
at the bullet's arrival into 7 bins, and sweeps window length K in
{1,4,8,16,32,48} with an interpolated suffix-backoff model under a BY-BATTLE
70/30 split (3 seeds).
Result: NO. On the pre-fire frame the window is worse than the single
fire-tick state at every K/Q/A (e.g. K=8 costs +0.35..+0.46 bits). On the
during-flight frame the entire apparent gain is the later decision tick, not
the window; the single state alone drops 2.70 -> 1.31 bits as K goes 1 -> 32.
The shuffle-order control confirms recency matters but the windows do not:
by Q=4/K=8 they average ~1 observation and never recur. The single state
survives as a strong predictor (log-loss 2.346 vs 2.698 majority; bin
accuracy 0.409 vs 0.235).
Gate only: no gun, no live-win claim.
This commit is contained in:
@@ -0,0 +1,321 @@
|
||||
# The state-window gate: does a temporal WINDOW of wave-relative states beat a SINGLE state?
|
||||
|
||||
**The design under test (the user's words).** *"We decide first what makes an
|
||||
'environment state' ... then we feed as input a long list of these states, long
|
||||
enough to have one or more patterns of movement inside."* The premise is that SBC
|
||||
is a coincidence detector and that a **temporal list** of states, fed instead of a
|
||||
single tick, turns "coincidence" into "movement pattern".
|
||||
|
||||
**Verdict: NO.** On held-out battles a window of wave-relative states **does not**
|
||||
predict DrussGT's lateral position at the bullet's arrival better than the single
|
||||
state at the **same decision tick**. On the pre-fire frame the window is strictly
|
||||
and substantially *worse* at every window length, every coarseness and every
|
||||
smoothing strength. On the during-flight frame the apparent gain is entirely the
|
||||
**later decision tick**, not the window — and where the window finally edges the
|
||||
single state (K=32, Q=4) it is by 0.045 bits on a subset where the 32-window
|
||||
**never repeats at all**. The design, as stated, is **dead**. The *state
|
||||
definition* survives and is the useful result.
|
||||
|
||||
This is a gate, not a gun. **No live win is claimed and none is measured here.**
|
||||
|
||||
---
|
||||
|
||||
## 1. The state (design artifact)
|
||||
|
||||
The dodge is a response to **our bullet**, so the state is **wave-relative**, never
|
||||
absolute world coordinates. One state at absolute tick `t`, in the frame of the
|
||||
bullet fired at `t0` along `u = (cos dir, sin dir)` (with `n = (-u_y, u_x)` the
|
||||
lateral unit vector):
|
||||
|
||||
| field | definition | why |
|
||||
|---|---|---|
|
||||
| `lat` | `(D(t) − P0) × u` — DrussGT's signed perpendicular offset from the bullet line, px | the quantity that decides the hit; signed so "which side" is visible |
|
||||
| `vlat` | `lat(t) − lat(t−1)` — signed lateral velocity, px/tick | crossing vs returning |
|
||||
| `toa` | `(t0 − t) + karr` — ticks until the bullet reaches its arrival plane | how much time is left to move |
|
||||
| `room` | ray distance from `D(t)` along `sign(vlat)·n` until the arena wall (18 px bot radius) | room to keep running vs being cornered — the only *directional* wall field, unlike plain "distance to nearest wall" |
|
||||
| `turn` | `wrap180(eh(t) − eh(t−1))` — signed turn rate, deg/tick | surfers slow to turn; turning reveals a reversal |
|
||||
|
||||
**Refinements to the starting proposal, and why.** (a) was kept signed, not
|
||||
absolutised, because the target is signed. (b) was made a *difference of the
|
||||
lateral offset*, so it is exactly the lateral component of velocity in the bullet
|
||||
frame rather than a speed. (d) "distance to the wall" was made *directional*
|
||||
(room to run in the current lateral-motion direction); a plain nearest-wall
|
||||
distance is not wave-relative and is nearly constant at long range. (e)
|
||||
"heading/turn rate" was reduced to the **turn rate** alone, because heading
|
||||
relative to the bullet is already implied by the sign of `vlat`, and the turn rate
|
||||
is the thing that distinguishes a surf from a corner.
|
||||
|
||||
**Quantisation (the numerosity dial).** Each field is binned into `Q ∈ {2, 3, 4}`
|
||||
symbols. A state is therefore:
|
||||
|
||||
| Q | bits / state | distinct states |
|
||||
|---|---|---|
|
||||
| 2 | **5.0** | 32 |
|
||||
| 3 | **7.9** | 243 |
|
||||
| 4 | **10.0** | 1024 |
|
||||
|
||||
Bin edges are **quantiles of the TRAIN side only** (no test leakage); a coarse
|
||||
state makes two similar situations look the same, which is what lets a
|
||||
"coincidence" repeat. The sweep below is exactly the question of whether coarser
|
||||
buys more recurrence than it loses in resolution.
|
||||
|
||||
### The two windows tested, and why both
|
||||
|
||||
A window is `K` consecutive states. Crucially, **the baseline is always the single
|
||||
state at the same decision tick `D`** — otherwise a longer window would "win" only
|
||||
because its last state is closer to the answer.
|
||||
|
||||
* **`pre`** — `D = t0` (the fire tick); the window is the `K` pre-fire states
|
||||
ending at `D`. This is what a bot can compute *before* pulling the trigger. All
|
||||
`K ≤ 48` are feasible (max recorded flight is 43 ticks).
|
||||
* **`fly`** — `D = t0+K−1`; the window is the first `K` states of the flight, and
|
||||
the baseline is the single state at `D`. Only shots with `karr > 31` are used,
|
||||
so the whole window is strictly before arrival (8156 of 54939 shots).
|
||||
|
||||
The task's field `toa` only exists once a bullet is in the air, so `fly` is the
|
||||
"reaction" reading; `pre` is the "am I about to be dodged?" reading. Both are
|
||||
reported.
|
||||
|
||||
### The target
|
||||
|
||||
`perp_arr` — DrussGT's **signed** perpendicular offset from the bullet line at the
|
||||
tick the bullet reaches its along-track plane. This is the miss offset that
|
||||
decides the hit; unlike "required lead" it does not depend on the bullet speed. It
|
||||
is quantised into 7 bins with edges `±120, ±60, ±18` px; the central `±18` px bin
|
||||
is the hit window (bot radius 18 px).
|
||||
|
||||
**Bin width in degrees-at-450px.** The central hit bin is 36 px wide = one bot
|
||||
diameter. `atan(18/range)` is the hit half-window: **2.29°** at 450 px, 4.23° at
|
||||
the median fire range (487 px), and 20.4° at 100 px. So the 36 px central bin is
|
||||
**4.58° at 450 px**. The three inner bins are 36/42/60 px wide (4.58/5.34/7.63° at
|
||||
450 px) — the target resolution is *finer* than the ~2.29° hit half-window, so a
|
||||
model that predicts the bin is not being asked an impossible question.
|
||||
|
||||
---
|
||||
|
||||
## 2. Data and split
|
||||
|
||||
* **Data (MEASURED present before use):** `/tmp/tfil_ab2/out/<A..E>/runN.jsonl` +
|
||||
`.events.jsonl` + `.rounds.json` — **70 real live battles vs the unmodified
|
||||
DrussGT, 490 rounds, 54 939 shots** by ModularBot, recorded by
|
||||
`tools/robocode_shim/run_bridge_battle.sh`. Per-shot geometry is re-derived by
|
||||
the validated instrument `common_libs/tests/analyze_drussgt_dodge_vs_power.py`
|
||||
(which this extractor imports). Attribution (`e*`=DrussGT, `s*`=us) is the one
|
||||
already documented in `docs/drussgt_dodge_vs_power.md`.
|
||||
* **Split (stated explicitly): BY BATTLE, never by tick.** All rounds of a battle
|
||||
go to one side. A 70 %/30 % battle split (**49 train / 21 test battles**),
|
||||
repeated over **3 seeds**. Bin edges are derived from the train side only. Two
|
||||
ticks inside one round never straddle the split.
|
||||
|
||||
---
|
||||
|
||||
## 3. Model (deliberately dull)
|
||||
|
||||
The question is about **information**, not modelling, so the model is a plain
|
||||
**interpolated (Jelinek–Mercer) suffix-backoff table** — the direct analogue of
|
||||
SBC "counting coincidences":
|
||||
|
||||
```
|
||||
P = global target-bin histogram
|
||||
for j = 1..K: P <- (count(suffix_j) + A*P) / (total(suffix_j) + A)
|
||||
```
|
||||
|
||||
It is order-sensitive, it can never do much worse than the shorter context, and it
|
||||
degrades gracefully when a suffix is unseen. Context depth is capped at 12 (orders
|
||||
above that are never repeated often enough to matter — see §5). `A ∈ {1,5,20}` is
|
||||
swept; **A = 5 is primary**. A majority/no-window predictor is the floor. The
|
||||
model predicts the 7-bin displacement distribution; we report held-out **accuracy**
|
||||
(argmax), **log-loss** in bits, and the **implied hit probability** `P(central
|
||||
bin)`.
|
||||
|
||||
---
|
||||
|
||||
## 4. Result — does the window beat the single state? **No** *(MEASURED)*
|
||||
|
||||
70 battles, 3 battle-split seeds, mean held-out log-loss (bits). `delta = window −
|
||||
single@D`; **negative would mean the window wins**.
|
||||
|
||||
### 4.1 `pre` frame — window ends at the fire tick, looks back *(54 939 shots)*
|
||||
|
||||
`single@D` is the fire-tick state; it does **not** change with K. A = 5:
|
||||
|
||||
| K | Q=2 single → window (Δ) | Q=3 single → window (Δ) | Q=4 single → window (Δ) |
|
||||
|---|---|---|---|
|
||||
| 1 | 2.4524 → 2.4524 (—) | 2.3782 → 2.3782 (—) | 2.3460 → 2.3460 (—) |
|
||||
| 4 | 2.4524 → 2.5107 (**+0.058**) | 2.3782 → 2.5461 (**+0.168**) | 2.3460 → 2.5656 (**+0.220**) |
|
||||
| 8 | 2.4524 → 2.8015 (**+0.349**) | 2.3782 → 2.8336 (**+0.455**) | 2.3460 → 2.7265 (**+0.381**) |
|
||||
| 16 | 2.4524 → 3.1310 (**+0.679**) | 2.3782 → 2.9473 (**+0.569**) | 2.3460 → 2.7450 (**+0.399**) |
|
||||
| 32 | same as K=16 (depth cap) | same | same |
|
||||
| 48 | same as K=16 (depth cap) | same | same |
|
||||
|
||||
Accuracy tells the same story (Q=4: single `0.409` → window `0.383`/`0.364`/
|
||||
`0.362`). The per-seed deltas are essentially identical (e.g. Q=2/K=4:
|
||||
`+0.0572, +0.0534, +0.0644`) — there is no seed where the window helps. The
|
||||
**heavy-smoothing** arm is the fairest to the window and it still loses:
|
||||
Q=4, A=20: K=4 `2.3928` vs single `2.3557` (+0.037); Q=2, A=20: K=4 `2.4643` vs
|
||||
`2.4524` (+0.012).
|
||||
|
||||
**The extra history is not just useless, it is actively harmful**, and the reason
|
||||
is visible in the shuffle control below: the long ordered contexts are sparse, and
|
||||
the interpolation spends held-out probability on them.
|
||||
|
||||
The baseline is not trivial: the single state halves the majority floor
|
||||
(majority log-loss **2.6983**, acc **0.2348**; best single Q=4 **2.3460**, acc
|
||||
**0.4094**).
|
||||
|
||||
### 4.2 `fly` frame — window starts at the fire tick, ends at `t0+K−1` *(8 156 shots)*
|
||||
|
||||
Here the **single@D itself** improves steeply as K grows, because the decision tick
|
||||
is later and therefore closer to the answer. A = 5:
|
||||
|
||||
| K | Q=2 single@D → window (Δ) | Q=3 single@D → window (Δ) | Q=4 single@D → window (Δ) |
|
||||
|---|---|---|---|
|
||||
| 1 | 2.6645 → 2.6645 (—) | 2.6741 → 2.6741 (—) | 2.7020 → 2.7020 (—) |
|
||||
| 4 | 2.6376 → 2.7695 (**+0.132**) | 2.6262 → 2.8858 (**+0.260**) | 2.6397 → 2.9108 (**+0.271**) |
|
||||
| 8 | 2.5659 → 3.0331 (**+0.467**) | 2.5302 → 3.0207 (**+0.490**) | 2.6101 → 2.9084 (**+0.298**) |
|
||||
| 16 | 2.4047 → 3.1643 (**+0.760**) | 2.3103 → 2.6362 (**+0.326**) | 2.3853 → 2.5918 (**+0.207**) |
|
||||
| 32 | 1.8365 → 1.9970 (**+0.161**) | 1.4355 → 1.4625 (**+0.027**) | **1.3121 → 1.2668 (−0.045)** |
|
||||
|
||||
Read this as: **the gain is the later tick, not the window.** `single@D` alone drops
|
||||
from `2.70` at K=1 to `1.31` at K=32 (Q=4) — the single state at `t0+31` is 2
|
||||
ticks from arrival and already knows the answer. The window adds nothing on top;
|
||||
the only cell where it "wins" is K=32/Q=4 by **0.045 bits**, and that subset's
|
||||
32-state window **never repeats** (§5), so the model is running on the single
|
||||
state plus a trace of recent acceleration. Q=3 at the same K is *worse*, and
|
||||
K=4–16 are worse by 5–20× that margin.
|
||||
|
||||
### 4.3 The mandatory shuffle-order control *(MEASURED)*
|
||||
|
||||
The order of the K states inside each window is randomly permuted, preserving the
|
||||
multiset. In the `pre` frame the shuffled window is **worse** than the temporal
|
||||
window at K=4 (e.g. Q=4: `2.5672` vs `2.5656`… within noise) but **better** at
|
||||
K≥8 (Q=4/K=16: shuffled `2.5613` vs temporal `2.7450`). In the `fly` frame at
|
||||
K=32 the shuffled window is much worse (Δ = `+0.35/ +0.80/ +1.12`).
|
||||
|
||||
**How to read that, honestly.** The shuffle penalty at `fly` K=32 is **not**
|
||||
evidence the window works: the 32-state window never repeats, so the only order
|
||||
the model can actually use is the *recency* of the last state. Shuffling destroys
|
||||
*which state is current*, and the model degrades because recency matters — a
|
||||
statement about the **single state**, not the window. Where the shuffled window is
|
||||
*better* than the temporal one (pre frame, K≥8), it is because random order
|
||||
disables the sparse long contexts that the temporal order feeds to the model.
|
||||
Either way, the window itself contributes nothing.
|
||||
|
||||
A deterministic **reverse** control (reverse time, preserves recurrence) is worse
|
||||
than temporal everywhere, confirming the model is genuinely order-sensitive and
|
||||
that the last state is the informative one.
|
||||
|
||||
---
|
||||
|
||||
## 5. Recurrence — the "we can find coincidences" premise *(MEASURED)*
|
||||
|
||||
Distinct **ordered** window tuples over the whole corpus (54 939 shots; 8 156 in
|
||||
the `fly` subset). A window is *learnable by exact coincidence* only if its count
|
||||
is well above 1.
|
||||
|
||||
| K | Q=2 distinct (mean count) | Q=3 | Q=4 |
|
||||
|---|---|---|---|
|
||||
| 1 | 16 (3433) | 82 (670) | 372 (148) |
|
||||
| 4 | 1 747 (31.4) | 11 018 (5.0) | 24 223 (2.3) |
|
||||
| 8 | 13 474 (4.1) | 37 337 (1.5) | 49 891 (**1.1**) |
|
||||
| 16 | 41 447 (1.3) | 54 163 (1.01) | 54 707 (1.00) |
|
||||
| 32 | 54 454 (**1.01**) | 54 754 (1.00) | 54 779 (1.00) |
|
||||
| 48 | 54 744 (1.00) | 54 772 (1.00) | 54 789 (1.00) |
|
||||
|
||||
(In the `fly` subset the same collapse happens earlier and harder: Q=4, K=8 already
|
||||
1.02 mean count, and **K=32 is 8 156 distinct of 8 156 — literally every window is
|
||||
unique**.)
|
||||
|
||||
So the "long list with patterns inside" premise fails as a **combinatorial**
|
||||
matter, not merely a modelling one:
|
||||
|
||||
* At the coarsest state worth using (Q=2, 5 bits), only **K ≤ 8** has meaningful
|
||||
recurrence — and even there the window does not beat the single state.
|
||||
* At the finest (Q=4, 10 bits), K=4 already averages 2.3 observations; K≥8 is
|
||||
essentially all singletons. **Learning from a window that never recurs is
|
||||
impossible**, and the sweep shows it does not happen.
|
||||
* The single state is the only object with a real coincidence structure: 16–372
|
||||
distinct values, every one repeated hundreds of times.
|
||||
|
||||
---
|
||||
|
||||
## 6. Direct answer
|
||||
|
||||
**Does a window of wave-relative states predict DrussGT's future lateral position
|
||||
better than a single state on held-out battles? NO.**
|
||||
|
||||
* `pre` frame: the window is worse at **every** K, Q and A; at K=8 it costs
|
||||
**+0.35 to +0.46 bits** over the single state.
|
||||
* `fly` frame: the whole apparent improvement is the later decision tick
|
||||
(`single@D` 2.70 → 1.31 bits as K goes 1 → 32). The window adds nothing except
|
||||
a 0.045-bit sliver at K=32/Q=4 on a subset whose 32-windows never repeat.
|
||||
* The recurrence statistics say the same thing independently: re-usable
|
||||
"patterns" of movement do not recur at any window length that carries extra
|
||||
information.
|
||||
|
||||
**The design — feed SBC a long temporal list of states — is dead.** Do not build
|
||||
it. *(MEASURED)*
|
||||
|
||||
**What survives (and is worth keeping).** The **state definition itself is a
|
||||
strong single-state predictor**: the fire-tick state with 4 symbols/field
|
||||
(10 bits) halves the majority log-loss (2.6983 → **2.3460** bits) and triples the
|
||||
chance the predicted displacement bin is right (0.2348 → **0.4094**) on battles
|
||||
the model has never seen. If anything from this line is taken forward, take the
|
||||
**single** wave-relative state at Q=4 and use it as a context for a *different*
|
||||
decision problem — but note the implied hit probability it produces is still the
|
||||
base rate (~0.09), so it is informative about **which side** the miss falls on, not
|
||||
about **whether** the shot hits. *(MEASURED/INFERRED)*
|
||||
|
||||
---
|
||||
|
||||
## 7. MEASURED vs INFERRED
|
||||
|
||||
**MEASURED** (exact commands in §8; numbers above)
|
||||
|
||||
* 70 battles / 490 rounds / 54 939 shots; 8 156 with `karr ≥ 33`.
|
||||
* Every held-out log-loss / accuracy / implied-hit number in §4, over 3
|
||||
battle-split seeds, with bin edges from train only.
|
||||
* The shuffle-order and reverse-order controls.
|
||||
* The distinct-window and repeat counts in §5.
|
||||
|
||||
**INFERRED**
|
||||
|
||||
* That the residual fly-frame "win" at K=32/Q=4 is a trace of recent acceleration
|
||||
rather than a movement pattern — it is consistent with the never-repeating
|
||||
32-window and with the single state already containing `vlat`.
|
||||
* That a *different* (e.g. TSetlin/SBC) learner would not reverse the answer. This
|
||||
is not proved. What **is** proved is the recurrence floor: a coincidence-based
|
||||
method cannot learn from windows that never recur, and by Q=4/K=8 they already
|
||||
essentially never do.
|
||||
* "Design is dead" applies to the **long temporal window**. It does not say the
|
||||
state is useless, and it does not say a *short* (K=2) context is worthless — K=2
|
||||
was not separately swept because at Q≥3 it is already 1.5–5 observations per
|
||||
context, i.e. below the point where it could help.
|
||||
|
||||
**Caveats.** (1) The `fly` frame is a biased long-range subset (all shots with
|
||||
`flight ≥ 33`, i.e. the longest-range quarter). (2) The model's context depth is
|
||||
capped at 12; K=16/32/48 therefore share a row, but §5 shows orders above ~4 are
|
||||
never repeated often enough to matter, so the cap is not the binding constraint.
|
||||
(3) The target is the miss offset, not the hit; the implied hit probability is the
|
||||
base rate for every model, so nothing here is a hit-rate claim.
|
||||
|
||||
---
|
||||
|
||||
## 8. Reproducing
|
||||
|
||||
```bash
|
||||
# corpus (live captures; not in the repo)
|
||||
ls /tmp/tfil_ab2/out # 70 battles: <A..E>/runN.jsonl{,.events.jsonl,.rounds.json}
|
||||
|
||||
python3 common_libs/tests/state_window_gate.py \
|
||||
--tfil /tmp/tfil_ab2/out --seeds 3 --frames pre,fly \
|
||||
--json common_libs/tests/fixtures/state_window_gate.json \
|
||||
> common_libs/tests/fixtures/state_window_gate_report.txt
|
||||
```
|
||||
|
||||
Runtime ≈ **5 minutes**, single-threaded pure Python (no numpy required). The
|
||||
script imports `common_libs/tests/analyze_drussgt_dodge_vs_power.py` for the
|
||||
per-shot geometry; `common_libs/tests/fixtures/state_window_gate_report.txt` is
|
||||
the verbatim captured output and `…state_window_gate.json` is the same numbers as
|
||||
JSON (including the per-seed values). No offline fixture replay, no simulator:
|
||||
these are the recorded real battles.
|
||||
Reference in New Issue
Block a user