Gate 2b: the shift form WORKS - but break-even needs ~80% side accuracy and the

problem gives 60%. Thread closed with a number, not a shrug.

Gate 2 (fd2f7f6) said STOP because every correction lost hits. But the shifts it
applied were the conditional MEDIANS of the error (4-16 deg) while the gate metric
is HITS, and the naive guess is UNBIASED - so shifting an already-centred
distribution can only destroy near-target mass. That suggested the experiment was
MISCALIBRATED rather than the idea being dead. One cheap test settled it.

ARMS (pooled, both primary fixtures, 25 rounds, N~9433/h; deltas in pp vs naive):
  h    naive   TM-med   TM-hit  TM-x0.25   PERF-SIGN   shuffled
  15    39.3    29.0(-10.3)  35.1(-4.2)  37.9(-1.5)  **47.0 (+7.7)**  35.1(-4.2)
  20    28.2    19.2(-9.0)   22.8(-5.4)  25.6(-2.6)  **33.4 (+5.1)**  22.0(-6.2)
  25    20.9    13.1(-7.8)   15.4(-5.5)  18.8(-2.2)  **25.6 (+4.7)**  16.0(-4.9)
  30    16.4    10.5(-5.9)   11.4(-5.0)  14.1(-2.3)  **18.3 (+1.9)**  11.2(-5.2)

1. **THE FORM CAN BUY HITS.** Perfect side knowledge + the hit-optimal shift gains
   **+4.9 pp mean** (+1.9..+7.7), positive in ~19/25 rounds and in BOTH fixtures -
   and it improves the median AND p90 residual. So the premise is NOT structurally
   impossible.
2. **BUT THE TM'S 60% SIDE ACCURACY LOSES** (-4.2 pp quadrant, -3.3..-5.6 side).
   Shrinking the shift toward zero monotonically reduces the loss but never turns
   it positive.
3. **BREAK-EVEN IS ~80% SIDE ACCURACY** (synthetic sweep: 0.70 -> -1.6 pp, **0.80
   -> 0.0**, 1.00 -> +4.9 pp). The observed side signal tops out at **~60% (TM) /
   53-56% (turn-only rule)** against 50% chance, and the headroom study found it is
   essentially ONE WEAK FEATURE. **No realistic predictor of this problem clears
   the 80% wall.**

WHY IT IS SO SENSITIVE - the hits-vs-shift curve (h15, calibration): shift 0 ->
36.2%, +-2deg -> 27.4/26.8, +-4deg -> 14.2, +-6deg -> 11.7, +-10deg -> 8.3.
Shifting unconditionally is catastrophic. The whole value comes from shifting ONLY
when the side is known, because the signed-error distribution becomes ONE-SIDED
once you condition on the true side. And the hit-optimal PERF-SIGN shift is only
**+-2 to 3.5 deg** - an order of magnitude smaller than the 4-16 deg medians Gate 2
used, which is exactly why Gate 2's calibration could never win.

VERDICT: **MISCALIBRATED, but practically dead at achievable accuracy.**
Recommendation: stop the per-bucket-shift direction, and record the RIGHT reason -
an **accuracy wall at ~80%**, not an impossibility of the form. If ever revived,
the only viable path is a side predictor that materially exceeds 60%; a better
calibration cannot fix it (we already used the hit-optimal one).

METHOD NOTE: the hit-optimal shift is fitted on the calibration slice (the
pipeline's existing out-of-sample offset split) and applied on the later eval
slice - NOT fitted on the TM training slice, to avoid in-sample label leakage.
INTEGRITY: TM-med reproduces Gate 2's committed deltas to the decimal
(-10.3/-9.0/-7.8/-5.9); PS-shift0 and TM-x0.0 are exactly naive; the shuffled
control never improves. Guards all pass (test_tm_diag 48, test_tm_automata_diag 55,
test_tm_clause_shape 66, diag_synthetic, diag_automata_validation, test_gun_harness,
test_vbullet_metric, test_power_selection, test_adaptive_radar,
test_tfil_ring_weights, test_power_policy, test_ram_decision, test_rack_membership,
test_selector_tiebreak, test_tm_pattern_registration, test_vbullet_admit_gate).

Caveats: DrussGT-only; enemy movement is a closed-loop response to our CURRENT
movement; offline observation is perfect while live we see the enemy only on scans
-> all absolute hit fractions are optimistic upper bounds, only deltas are
meaningful.
This commit is contained in:
2026-09-22 22:21:32 +02:00
parent fd2f7f608c
commit be74e369eb
2 changed files with 2258 additions and 0 deletions
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff