From 55e92bc5e08c7db4be461e1f8c4a19af17539365 Mon Sep 17 00:00:00 2001 From: pi Date: Sun, 27 Sep 2026 15:05:14 +0200 Subject: [PATCH] docs: the range-vs-approach ceiling - melee unavailable vs DrussGT, the 16 deg is a lead-model number, pre-multiply premise dead --- docs/range_vs_approach_ceiling.md | 283 ++++++++++++++++++++++++++++++ 1 file changed, 283 insertions(+) create mode 100644 docs/range_vs_approach_ceiling.md diff --git a/docs/range_vs_approach_ceiling.md b/docs/range_vs_approach_ceiling.md new file mode 100644 index 0000000..05e1993 --- /dev/null +++ b/docs/range_vs_approach_ceiling.md @@ -0,0 +1,283 @@ +# Range vs approach: melee is unavailable against DrussGT, so the gun at range is the only lever + +**Date:** 2026-09-27 · **Job:** j167 · **Branch:** `research/lead-targeting` +**Evidence base:** j166 (`worktrees/j166-aim` @ `a5a49bd`), j165 (`fe77056`), j159 +(`4a1f3e1`), j165/j151/j152/j154, j160/j163 (`0df7763` / `51bfa57`), j161 +(`docs/ram_floor_exhaustion_ab.md:219`). +**Instrument for the new numbers below:** `worktrees/j167-ceiling/j167_probe.py` +(branch `j167-ceiling`) — pure replay of the recorded corpora +(`/tmp/tfil_ab2/out/`, 60 252 of our own scored shots over 140 battles; and +`/tmp/firelag_live2/`, 1 700 shots / 1 664 incoming bullets over 4 battles). +**No battle, A/B, server or GUI was run for this document.** + +--- + +## 1. The ceiling + +**Melee is structurally unavailable against DrussGT, and the exhaust/ram line is a +niche rather than a lever.** The evidence is two-sided and independent: (a) *our +mover's own ruler* — 94% of forced (no-safe-tile) picks happen at range > 300 u +and only 6-7% of picks reach the chosen tile at the estimated arrival time, with +the destination hot on arrival 35-42% of the time; and (b) *the j166 pursuit +probe* — 35 windows × 250 ticks of open-loop kinematics in which **every** +steering law is equal-or-worse than doing nothing clever: + +| steering law (j166) | closing (u/tick) | contact % | TTI (ticks) | +|---|---:|---:|---:| +| current-position closing | 4.03 | 65.7 | 72.3 | +| body/barrel ray | 1.38 | 40.0 | 131.4 | +| velocity intercept (degenerate at equal speed) | — | 0 over 2 118 ticks | — | +| best case: lag-5 lead | 4.20 | 65.7 | 68.9 | + +The root cause of the historical **0/59 proactive-ram** result +(`docs/ramming_negative_result.md`) is not a bad gate: **DrussGT never let the +distance drop.** Per-round minimum distance 152-338 u, median ~490 u, and +`frac(dist < 50) = 0.000` in all four recorded rounds. A pursuit that never gets +below 152 u cannot make contact, whatever the gate says. It is also not a +gun-side problem: **the server never transmits the enemy's gun direction** +(`ScannedBotEvent` = `energy, x, y, direction` where `direction` is the BODY +heading, plus `speed`; `TurnProcessor.kt:313-323`). There is no aim-based lead, +no aim-based dodge and no early warning available. Against DrussGT the +body-to-bullet angle has median **90.1 deg**, and the body ray passes within +10 deg of us on **0.0% of 1 794 ticks** — its gun is always on us, its body +never is. + +**Recorded so the idea is not re-proposed:** the j166 lag-5 residue does improve +TTI (72.3 → 68.9) and **converts to contact 0% of the time**. A 4% TTI gain with +zero contact conversion is noise, not a lead. + +### The honest remaining niches for exhaust/ram + +1. **An opponent that closes on us.** Ram works whenever the other side comes to + us. Nothing here generalises away from that. +2. **A late-round exhaustion when they are already near.** The one conversion + ever recorded came from a *finisher* (enemy 16 → 1 energy), which is already + the default gate. +3. **Any 2v1+ mode**, where closing dynamics are not symmetric. + +Against DrussGT specifically none of these will move the score, and +`TR_RAM_FLOOR_ENERGY` is under test in j163 — do not duplicate it. + +--- + +## 2. What the ceiling implies + +**If range is held, the only remaining lever is the gun at range, and the binding +numbers are the gun's, not the tile picker's.** The long-range hit rate is +**~9-10%** (Pattern live: 12.3% at 300-450 px, 9.2% at 450+; overall 10.5% — +`docs/headon_longrange_live.md`), the live hit half-window at 450 px is +**`atan(18/450) = 2.29°`** (`docs/gun_campaign.md:59`), and the measured arrival +aim error is **16.19° mean-abs at 450+** (`docs/bitbrain_campaign.md:107`, +`docs/headon_longrange_live.md:85`). 16.19° is **7× the window**. The tile picker +cannot close a 7× gap that sits downstream of the gun. + +### New measurement — arrival aim error decomposed (j167, 23 275 shots at 450+ px) + +Arrival aim error is defined non-circularly: the angle between the fired bearing +and the bearing to where the target *actually is* when the bullet arrives +(`tof = 20 - 3·power`, so the flight time comes from the power, not from the +shot's own geometry). It splits **exactly**, as signed angles, into + +* **B, the model part** = the error the gun's own lead model leaves behind, and +* **C, manoeuvre** = the target's path curvature relative to the + constant-velocity extrapolation from the true state at fire time. + +| band (px) | n | mean|A| | mean|B| (model) | mean|C| (manoeuvre) | sd(B) | sd(C) | corr(B,C) | +|---|---:|---:|---:|---:|---:|---:|---:| +| 0-100 | 24 767 | 83.36 | 101.56 | 49.38 | 124.5 | 75.9 | −0.65 | +| 300-450 | 11 372 | 13.07 | 21.77 | 11.00 | 26.0 | 13.0 | −0.87 | +| **450+** | **23 275** | **11.26** | **17.18** | **7.73** | **20.7** | **9.3** | **−0.85** | + +At 450+ the model part's variance is **2.2× the manoeuvre part's**, and +`corr(B,C) = −0.85` means the two largely *cancel* — the net 11.26° is much +smaller than either part. **The 16° is a lead-model number, not a dodge number.** +Two supporting numbers: a naive constant-velocity extrapolation of a **2-tick-old** +position scores 8.45° mean-abs at 450+, and the time-of-flight implied by the +shot's own geometry (holding the current velocity) sits a **median 10 ticks short** +of the power-derived arrival tick (p10 −18, p90 +31) — i.e. the gun systematically +**under-leads in time**, consistent with `docs/lead_capture_by_range.md` +(capture 0.135 at 450+). + +> **Do not read "a stale-CV model scores 8.45°" as "simplify the gun".** This is +> exactly the offline-ruler trap that killed HeadOn: the ruler said a no-lead gun +> was equal-or-better at 300+ and live it hit **20×/23× less** +> (`docs/headon_longrange_live.md`). The corpus is closed-loop — the target's +> manoeuvre is a *reaction to our own bullet* — so (B) and (C) are not separable +> here, and per `docs/offline_harness_trust.md` (j89: 0/6 on closed-loop) this +> instrument ranks per-gun single-tick prediction, it does not predict a live A/B. + +### Cross-reference: what is still open in the gun docs + +| doc | finding | status after this ceiling | +|---|---|---| +| `docs/gun_campaign.md:59` | hit half-window 2.29° at 450 px; measured signal 4.6-7.6° | **STILL OPEN and now the load-bearing number.** The decomposition says the gap is in the *model*, and the model is systematically 10 ticks short in time-of-flight. | +| `docs/gun_campaign.md:40-45` | lead amplitude is dead (1.0/1.5/2.0/3.0 all worse); radial knobs are bearing-invariant by construction | **CLOSED.** | +| `docs/gun_campaign.md:737-753` | `len6` +0.49 wins/run (p=0.039, n=15) did not replicate on n=33 | **CLOSED.** | +| `docs/bitbrain_campaign.md:189` | BitBrain / TMHorizon corrector adds no measurable aim (16.199 vs 16.193) | **CLOSED.** | +| `docs/bitbrain_campaign.md:107` | Pattern's own lead correlation with the required lead is 0.165 at 450+ | **STILL OPEN.** It is the same defect the decomposition names. | +| `docs/state_window_gate.md` | single wave-relative state at Q=4 predicts the miss bin at 0.4094 vs 0.2348 majority, but bins are 4.58-7.63° wide | **STILL OPEN, and now the best-placed surviving idea** — it is a *model* correction, which is where the error is. | +| `docs/gun_rack_analysis.md:423-455` | the 16-candidate rack ranking A/B found no winner; knobs added, all neutral | **CLOSED** (13 guns, `onlyPattern` shipped). | + +--- + +## 3. Negative-results ledger — mechanisms closed by measurement + +Do not re-litigate any row. The unit of evidence is the **opponent**. + +| mechanism | knob / job | headline number | verdict | +|---|---|---|---| +| Geometry-weighted tile draw | `TR_TFIL_GEO_MODE/TAU`, j152 `38fbc6e`, A/B'd j159 `4a1f3e1` | **−8.83 damage/run, p=0.0061**; wins −0.05, p=0.46; +26.3 px mean distance on 15/15 opponents | **REJECTED.** Default off, stays off. | +| The bounded hold | `TR_TFIL_HOLD_MAX_TICKS`, j154 `2223ca6` | mechanism-positive, outcome-null (j146/j153) | **Default off.** No live win. | +| The proactive ram | `oldram` vs `base` gate `dist<200` | **p=0.69**, damage 279 vs 284, survival 17/49 vs 16/49; **0/59 opportunity→contact** | **CLOSED** (`docs/ramming_negative_result.md`). | +| The aim-based ram | j166 `a5a49bd` | body ray within 10° of us on **0.0% of 1 794 ticks**; body/barrel ray contact 40.0% vs 65.7% for doing nothing clever | **IMPOSSIBLE** — the server never sends gun direction (`TurnProcessor.kt:313-323`). | +| Arrival commitment (`tfil`) | j144 `d2005ab` | mechanism-positive, outcome-null | Default off. | +| Turn-cost tiebreak among safe tiles | j145 `39c90fd` | real but small mechanism, under-powered outcome null (300 battles, 5 arms) | Default off. | +| Field shape (safety) | j146 `de5d02b` | safe-set broken 63.5% → 30.4% offline; live null on damage and wins (375 battles, 5 arms) | **Default off.** | +| Corridor bound | j148 `5e213df` `TR_{TFIL,STRAFE}_CORRIDOR_TICKS` | never landed in a live A/B | Untested, not a candidate. | +| Ring arrival commitment | j165 `fe77056` `TR_TFIL_RING_COMMIT_ARRIVAL` | reach 0.24% → **3.05%**, picks 5 521 → 525, byte-for-byte default parity over 20 026 ticks, 148 guards | **Mechanism-positive, default off.** The strongest surviving movement mechanism. | +| Firing floor / enemy-exhaustion ram | j160 `23bce2d`, A/B'd j163 `51bfa57` | **clean negative**; the offline energy corpus missed the live game by 200× | **Under test in j163 — do not duplicate.** | +| Fire-detection lag | j147 `d21f7ce` `TR_FIRE_LAG` | displacement 19.06 → 5.37 px, deadline error 0.99 → 0.06 ticks; **live outcome-neutral**; ceiling ~10% of incoming damage (measured below) | **Default off, permanently.** | +| Hard arrival bound | j151 `a01141c` `TR_TFIL_ARRIVE_TICKS` | mechanism-positive, outcome-null | Default off. | + +> **Methodological caution (j161), binding on everything above.** Pooled tests +> can hide real per-opponent effects: j159's safety signal was **p=0.0008 +> per-opponent while the pooled test was null** (`docs/ram_floor_exhaustion_ab.md:219`). +> **Any future mechanism claim must report per-opponent mechanism metrics, not a +> pooled mean.** A pooled null is not evidence of absence; it is evidence that +> the heterogeneity was not averaged down. + +--- + +## 4. Lead-time lever 1 — what a 2-tick-stale ghost really costs + +`TR_FIRE_LAG` back-dates the bullet ghost (default 0). Energy-drop shot +detection lags **1.9 ticks mean**; median bullet flight is **19 ticks** +(`onHitByBullet` gives 82 hits / 7 421 ticks, one update per ~90 ticks). + +**Measured on 55 750 incoming bullets** (`/tmp/tfil_ab2/out/`). For each bullet: +the time to closest approach of the target's recorded path to the bullet line +(**median 9 ticks**, p10 1, p90 39), and the minimum number of ticks of lead time +a max-speed hard-turn dodge needs to build 17 px of lateral displacement: + +| minimum dodge lead time (ticks) | 0 | 1 | 2 | 3 | 4 | 5+ | +|---|---:|---:|---:|---:|---:|---:| +| share of incoming bullets | **57%** | 33% | 4% | 2% | 1% | 2% | + +**57% of incoming bullets are already undodgeable at the instant they are fired**, +and only **~10%** (need ≥ 2 ticks) are in a regime where a 2-tick detection lag +can change anything. Applying the lag to the open-loop dodge model: + +| ghost lag (ticks) | modelled hits | Δ vs perfect | share of all bullets whose hit/miss verdict flips | +|---|---:|---:|---:| +| 0 | 22 419 | — | — | +| **1.9 / 2** | **25 100** | **+2 681 (+12.0%)** | **10.18%** | +| 3 | 26 230 | +14.6% | 14.63% | +| 5 | 28 949 | +22.0% | 22.04% | + +**Verdict: the 1.9-tick lag costs on the order of 10% more incoming hits** — at +the measured ~200 damage/run, roughly **20 damage/run**, an order of magnitude +below the movement A/B damage MDE. This is consistent with `TR_FIRE_LAG`'s already +measured live outcome-neutral result. The ghost is *wrong*, but wrongness at +10% of incoming damage cannot be turned into wins at this sample size. + +**Recommendation: `TR_FIRE_LAG` stays off permanently.** It is a correctness fix +with a measured, bounded, sub-MDE payoff. + +--- + +## 5. Lead-time lever 2 — the 16° decomposed, component by component + +At 450+ px (23 275 shots), against the 11.26° net arrival error: + +| component | measured | addressable? | +|---|---|---| +| **(a) enemy body-gun decoupling** | `\|gun dir − body heading\|` median **89.9°** (p10 25.9, p90 154.0, n=60 928). Extrapolating the target along its **gun** instead of its **body** would put the arrival bearing **79.5° median** wrong. | **Not present, and not addressable.** The intercept model uses the target's *recorded position and velocity*, both of which are the true body quantities and both exactly observed. Body-gun decoupling therefore contributes **exactly 0** to our arrival error. It is fatal for *aim-based* leading and threat warning (j166) and irrelevant to *position-based* leading. | +| **(b) our own leading model** | mean|·| **17.18°**, sd **20.7**; implied time-of-flight a **median 10 ticks short** of the power-derived arrival tick | **DOMINANT, and addressable.** This is ~2.2× the manoeuvre variance and it is the whole of the 16°. | +| **(c) target manoeuvre between scan and fire** | mean|·| **7.73°**, sd **9.3** | Small relative to (b), and **irreducible** — it is the dodger's own unpredictability, exactly the ~half of the under-lead `docs/lead_capture_by_range.md` attributes to a trivial predictor's own ceiling. | +| **(d) gun turn rate / time-to-fire** | the correct solution drifts a **median 0.416°/tick** (p90 5.45). The gun turns at 10°/tick, so a 17° correction takes **1.7 ticks ≈ 0.40°** of drift. | **Not binding.** Contributes ~**0.4°, i.e. ~3% of the 11.26° error.** The gun can always reach the answer; it aims at the wrong answer. | + +**So the 16° is not (a), not (c) and not (d). It is (b) — the lead model's +time-of-flight, short by ~10 ticks.** Caveat, stated once and load-bearing: on a +closed-loop corpus (B) and (C) are not cleanly separable, since the target's +manoeuvre is a reaction to our own shot; the `corr(B,C) = −0.85` is exactly that +confound showing up. The *rank order* (b) ≫ (c) ≫ (d) > (a)=0 is robust to it +because (b) and (c) differ by 2.2× in variance and (d) is 3%. + +--- + +## 6. The proposed lever: pre-multiply before learning — PREMISE DEAD + +The design: aim error is largely a *product* (bearing-rate × time-of-flight), so +pre-multiply the two features and feed one small Tsetlin machine. **Measured on +the same corpus, the premise does not hold and the experiment should not be +built.** `y` = the required lead angle (current bearing → arrival bearing), i.e. +exactly the quantity the gun must predict; `b` = the observable 4-tick finite +difference of the bearing; `t = 20 − 3·power`. + +| band (px) | n | corr(**b·t**, y) | corr(b+t, y) | R² additive [1,b,t] | R² product [1,b·t] | held-out side acc, additive | held-out side acc, product | held-out residual rms (deg) | +|---|---:|---:|---:|---:|---:|---:|---:|---:| +| 0-200 | 24 934 | −0.1020 | −0.1022 | 0.0105 | 0.0104 | 0.579 | 0.580 | 75.98 | +| 200-300 | 671 | 0.1190 | 0.0766 | 0.0147 | 0.0142 | 0.488 | 0.487 | 13.70 | +| 300-450 | 11 372 | **0.1501** | 0.0299 | 0.0256 | 0.0225 | 0.544 | 0.534 | 8.43 | +| **450+** | **23 275** | **0.2833** | 0.1052 | **0.0831** | 0.0803 | **0.603** | **0.601** | **6.68** | +| pooled | 60 252 | −0.1005 | −0.1009 | 0.0102 | 0.0101 | — | — | — | + +*(side accuracy is 2-fold held-out and balanced; the TM record is ~0.47-0.49)* + +Two things are true and the second kills the idea: + +1. **As a single scalar, the product is much the better feature at range**: + `corr(b·t, y) = 0.283` vs `corr(b+t, y) = 0.105` at 450+ — 2.7× better, and + 5× better at 300-450. So the *premise* ("the error is a product, not a sum") + is **confirmed as a statement about correlation**. +2. **But it buys nothing a weight-sum cannot already express.** The best linear + additive model on the same two features reaches **R² 0.0831 vs the product's + 0.0803**, and the held-out balanced side accuracy is **0.603 (additive) vs + 0.601 (product)** — a 0.002 difference, i.e. nothing. Pooled, the two are + identical (−0.1005 vs −0.1009; R² 0.0102 vs 0.0101). A TM with two input + features **already reconstructs the product term**; the multiplication is what + the network was doing anyway. + +**Even the ceiling is out of reach.** The best held-out residual on the required +lead at 450+ is **6.68° rms**, against a live hit half-window of **2.29°** — a +2.9× shortfall. Pre-multiplying does not get a classifier to 2.29°; nothing in +this family does. **Do not build it.** The spec is recorded here so the idea is +closed on measurement rather than on taste. + +*(Had it survived, the spec would have been: one TM, ONE input feature `b·t` +binarised on sign, plus the 4-bit horizon one-hot as today; offline gate = +held-out balanced side accuracy above 0.55 and residual rms below 3° at 450+; +live gate = wins/run with CI excluding 0 and sign-flip p<0.05 at 210 runs/arm, +damage not detectably down, MDE 0.17 wins/run. Predicted accuracy was 0.60 side +accuracy, which is a real signal against the 0.47-0.49 record — and still not +close enough to the window to convert.)* + +--- + +## 7. The A/B queue, in priority order, with the MDE honestly restated + +Throughput **22.7-23.2 runs/min**; movement gate resolved **0.17 wins/run at 210 +runs/arm**; the `1/√n` extrapolation to 0.10 wins/run is **607 runs/arm ≈ 1.4 h — +a FLOOR on elapsed time, not an estimate**, because opponent heterogeneity does +not average down. A null at this sample size **only excludes a LARGE effect.** +(j163 additionally measured 14-22 runs/min, not 22.7-23.2, so even the floor is +optimistic.) + +| # | experiment | what it tests | cost | a null would license | +|---|---|---|---|---| +| **1** | **The lead-model time-of-flight correction** (j167's (b)): re-derive the gun's arrival prediction so the implied flight is the power-derived tick, not 10 ticks short. | The one component that carries 2.2× the error variance at 450+, and the only open axis in `docs/gun_campaign.md` (lead *information*, not amplitude). | Offline gate first: arrival aim error at 450+ must fall below 11.26° mean-abs on held-out battles, ideally <8°; only then 2 arms × 15 opponents × 14 runs = 420 battles ≈ **0.3-0.4 h** wall. | Closing the single open gun axis. Nothing left in the gun. | +| 2 | `TR_TFIL_RING_COMMIT_ARRIVAL` (j165, default off) | Whether the largest surviving *movement* mechanism (reach 0.24% → 3.05%, picks 5 521 → 525, 148 guards) converts to wins. | 210 runs/arm ≈ **1.4 h floor**. | Retiring the whole ring/approach programme: if even a 12× reach gain is outcome-null, the ceiling argument is confirmed end to end. | +| 3 | `TR_FIRE_LAG` (tfil/strafe, default off) | Nothing worth testing — its ceiling is now measured at **~10% of incoming damage ≈ 20 dmg/run**, below the MDE. | Would be 1.4 h to learn nothing. | Nothing. **Skip it**; the measurement has already answered it. | +| 4 | `TR_TFIL_ARRIVE_TICKS` (j151, default off) | Whether a hard arrival bound converts now that the ring is rehabilitated. | 1.4 h. | Retiring it with j151's own null attached. | +| 5 | `TR_RAM_FLOOR_ENERGY` (j160, j163) | **Under test in j163. DO NOT DUPLICATE.** | — | — | + +**Recommendation.** Run **only experiment 1**, and only after the *offline* gate +passes; if the offline gate does not move the 450+ arrival error below ~8°, run +nothing at all. Given five consecutive nulls or near-nulls (j144, j145, j146, +j147, j159) plus a clean negative in j163, spending 1.4 h of live time on +experiments 2-4 is not justified — those are mechanism-positive +mechanisms whose outcome nulls are already the standing record, and a null there +teaches nothing that the ledger does not already say. + +**"The ceiling is real and we should stop spending on movement" is the answer.** +The remaining budget belongs to the gun's lead model, or it is not spent.