Files
SirRoboGarage/common_libs/tests
SirStone e180626b50 Horizon headroom: QUALIFIED PASS at h=10..50, but the signal is one feature
Measures the learnable headroom of the "where will the enemy be in h ticks"
problem on the DrussGT fixtures, for h=1..50, using the ANGULAR error (the aim
only cares about the angle - a distance-only error cannot change the shot).
Observer = ModularBot at t; naive guess = straight-line extrapolation, never
bounced off walls. Round-bounded via the .rounds.json sidecars; the tail h ticks
of each round are dropped, never labelled with garbage. N = 32,630 (h=1) down to
31,405 (h=50).

WHAT THE NUMBERS SAY
- **The question is FAIR.** Sign balance is dead-centre at EVERY horizon
  (47.7-50.2% left) with zero systematic bias (median signed error = 0.00 deg at
  every h). No de-biasing needed - a healthy symmetric question, which is what
  the two-binary output shape needs.
- **The dumb guess is exact short, wrong long.** naiveMiss (aim lands outside the
  18px body): 0% at h<=3, 3.3% at 4, 11.8% at 5, 27.6% at 6, 48.6% at 10, 62.6%
  at 15, 73.5% at 20, 85.9% at 30, 93.7% at 50. Error magnitudes: median 2.0 deg
  (h10), 7.0 (h20), 12.8 (h30), 18.5 (h40), 24.5 (h50). Body half-angle at 300px
  is 3.43 deg - so at long horizons the error is 4-7x the body size.
- **No trivial rule solves it.** Turn-direction accuracy is 53-54% at h<=3, drops
  through 50% at h~7, and INVERTS to 38-45% at long h (flipped = 55-62%, the best
  single rule anywhere). Causal 1-step persistence peaks ~65% at h2-4 and decays
  to ~49% by h50. Majority is 50-51.5%.
- **REAL signal exists in exactly ONE feature: the enemy's current turn/reversal
  direction.** dTurn swings from +8.4pp (h2) through 0 (h7) to **-24.0pp at h50",
  all |z|>6. Every other planned feature is WEAK: walls <=2-4pp, speed <=5pp,
  bullets <=4pp, reversal <=6pp, closing <=3pp.

TWO FINDINGS I DID NOT EXPECT
1. **The dead zone is h=5..9.** The naive guess starts missing there (12-44%) but
   NO tested feature shifts the left/right split by >=5pp - the sign is a
   featureless coin flip in that band. So those horizons carry no learnable signal
   despite looking promising.
2. **The bullet block - which we designed with enthusiasm - shows <=4pp of shift.**
   The enemy's dodging reaction to our bullet is NOT a strong conditioning signal
   at these horizons in this measurement. CAVEAT: the "bullet in flight" split
   used an energy-drop proxy with ~370 false positives in 1504 positives, so this
   is a weak negative, not a settled one - the block should be measured properly
   before being cut.

RECOMMENDED RANGE: **h = 10..50** (41 horizons, contiguous). Criteria: (a) 42-58%
left, (b) max(turnAcc, pers1Acc) < 75%, (c) some feature |delta| >= 5pp with
|z| >= 3, (d) naiveMiss >= 15%. h=1-4 fail (d); h=5-9 fail (c); from h=10 all four
hold and strengthen with h.

VERDICT - QUALIFIED PASS, and the job's own calibration is worth quoting: there is
genuine, non-degenerate structure, so the horizon-input design is not obviously
wasted; BUT the per-sample signal is weak and concentrated almost entirely in one
feature which already captures most of the easy structure (~60% sign accuracy).
**"I would not treat this as a green light for a big build; I would first check
that a model can beat 60% sign accuracy on a held-out round at h~15-25."**

CAVEATS (from the tool): DrussGT-only, and the enemy's movement at capture time
was a RESPONSE to our current movement, so the headroom is conditional on how we
move now; offline observation is perfect every tick while live we see the enemy
only on radar scans, so these numbers are an UPPER BOUND; and the replay is
open-loop even though the capture was closed-loop.
2026-09-22 21:46:05 +02:00
..