Gate 2b: the shift form WORKS - but break-even needs ~80% side accuracy and the
problem gives 60%. Thread closed with a number, not a shrug.
Gate 2 (fd2f7f6) said STOP because every correction lost hits. But the shifts it
applied were the conditional MEDIANS of the error (4-16 deg) while the gate metric
is HITS, and the naive guess is UNBIASED - so shifting an already-centred
distribution can only destroy near-target mass. That suggested the experiment was
MISCALIBRATED rather than the idea being dead. One cheap test settled it.
ARMS (pooled, both primary fixtures, 25 rounds, N~9433/h; deltas in pp vs naive):
h naive TM-med TM-hit TM-x0.25 PERF-SIGN shuffled
15 39.3 29.0(-10.3) 35.1(-4.2) 37.9(-1.5) **47.0 (+7.7)** 35.1(-4.2)
20 28.2 19.2(-9.0) 22.8(-5.4) 25.6(-2.6) **33.4 (+5.1)** 22.0(-6.2)
25 20.9 13.1(-7.8) 15.4(-5.5) 18.8(-2.2) **25.6 (+4.7)** 16.0(-4.9)
30 16.4 10.5(-5.9) 11.4(-5.0) 14.1(-2.3) **18.3 (+1.9)** 11.2(-5.2)
1. **THE FORM CAN BUY HITS.** Perfect side knowledge + the hit-optimal shift gains
**+4.9 pp mean** (+1.9..+7.7), positive in ~19/25 rounds and in BOTH fixtures -
and it improves the median AND p90 residual. So the premise is NOT structurally
impossible.
2. **BUT THE TM'S 60% SIDE ACCURACY LOSES** (-4.2 pp quadrant, -3.3..-5.6 side).
Shrinking the shift toward zero monotonically reduces the loss but never turns
it positive.
3. **BREAK-EVEN IS ~80% SIDE ACCURACY** (synthetic sweep: 0.70 -> -1.6 pp, **0.80
-> 0.0**, 1.00 -> +4.9 pp). The observed side signal tops out at **~60% (TM) /
53-56% (turn-only rule)** against 50% chance, and the headroom study found it is
essentially ONE WEAK FEATURE. **No realistic predictor of this problem clears
the 80% wall.**
WHY IT IS SO SENSITIVE - the hits-vs-shift curve (h15, calibration): shift 0 ->
36.2%, +-2deg -> 27.4/26.8, +-4deg -> 14.2, +-6deg -> 11.7, +-10deg -> 8.3.
Shifting unconditionally is catastrophic. The whole value comes from shifting ONLY
when the side is known, because the signed-error distribution becomes ONE-SIDED
once you condition on the true side. And the hit-optimal PERF-SIGN shift is only
**+-2 to 3.5 deg** - an order of magnitude smaller than the 4-16 deg medians Gate 2
used, which is exactly why Gate 2's calibration could never win.
VERDICT: **MISCALIBRATED, but practically dead at achievable accuracy.**
Recommendation: stop the per-bucket-shift direction, and record the RIGHT reason -
an **accuracy wall at ~80%**, not an impossibility of the form. If ever revived,
the only viable path is a side predictor that materially exceeds 60%; a better
calibration cannot fix it (we already used the hit-optimal one).
METHOD NOTE: the hit-optimal shift is fitted on the calibration slice (the
pipeline's existing out-of-sample offset split) and applied on the later eval
slice - NOT fitted on the TM training slice, to avoid in-sample label leakage.
INTEGRITY: TM-med reproduces Gate 2's committed deltas to the decimal
(-10.3/-9.0/-7.8/-5.9); PS-shift0 and TM-x0.0 are exactly naive; the shuffled
control never improves. Guards all pass (test_tm_diag 48, test_tm_automata_diag 55,
test_tm_clause_shape 66, diag_synthetic, diag_automata_validation, test_gun_harness,
test_vbullet_metric, test_power_selection, test_adaptive_radar,
test_tfil_ring_weights, test_power_policy, test_ram_decision, test_rack_membership,
test_selector_tiebreak, test_tm_pattern_registration, test_vbullet_admit_gate).
Caveats: DrussGT-only; enemy movement is a closed-loop response to our CURRENT
movement; offline observation is perfect while live we see the enemy only on scans
-> all absolute hit fractions are optimistic upper bounds, only deltas are
meaningful.
This commit is contained in: