docs: the range-vs-approach ceiling - melee unavailable vs DrussGT, the 16 deg is a lead-model number, pre-multiply premise dead

This commit is contained in:
pi
2026-09-27 15:05:14 +02:00
parent 486e2a69c6
commit 55e92bc5e0
+283
View File
@@ -0,0 +1,283 @@
# Range vs approach: melee is unavailable against DrussGT, so the gun at range is the only lever
**Date:** 2026-09-27 · **Job:** j167 · **Branch:** `research/lead-targeting`
**Evidence base:** j166 (`worktrees/j166-aim` @ `a5a49bd`), j165 (`fe77056`), j159
(`4a1f3e1`), j165/j151/j152/j154, j160/j163 (`0df7763` / `51bfa57`), j161
(`docs/ram_floor_exhaustion_ab.md:219`).
**Instrument for the new numbers below:** `worktrees/j167-ceiling/j167_probe.py`
(branch `j167-ceiling`) — pure replay of the recorded corpora
(`/tmp/tfil_ab2/out/`, 60 252 of our own scored shots over 140 battles; and
`/tmp/firelag_live2/`, 1 700 shots / 1 664 incoming bullets over 4 battles).
**No battle, A/B, server or GUI was run for this document.**
---
## 1. The ceiling
**Melee is structurally unavailable against DrussGT, and the exhaust/ram line is a
niche rather than a lever.** The evidence is two-sided and independent: (a) *our
mover's own ruler* — 94% of forced (no-safe-tile) picks happen at range > 300 u
and only 6-7% of picks reach the chosen tile at the estimated arrival time, with
the destination hot on arrival 35-42% of the time; and (b) *the j166 pursuit
probe* — 35 windows × 250 ticks of open-loop kinematics in which **every**
steering law is equal-or-worse than doing nothing clever:
| steering law (j166) | closing (u/tick) | contact % | TTI (ticks) |
|---|---:|---:|---:|
| current-position closing | 4.03 | 65.7 | 72.3 |
| body/barrel ray | 1.38 | 40.0 | 131.4 |
| velocity intercept (degenerate at equal speed) | — | 0 over 2 118 ticks | — |
| best case: lag-5 lead | 4.20 | 65.7 | 68.9 |
The root cause of the historical **0/59 proactive-ram** result
(`docs/ramming_negative_result.md`) is not a bad gate: **DrussGT never let the
distance drop.** Per-round minimum distance 152-338 u, median ~490 u, and
`frac(dist < 50) = 0.000` in all four recorded rounds. A pursuit that never gets
below 152 u cannot make contact, whatever the gate says. It is also not a
gun-side problem: **the server never transmits the enemy's gun direction**
(`ScannedBotEvent` = `energy, x, y, direction` where `direction` is the BODY
heading, plus `speed`; `TurnProcessor.kt:313-323`). There is no aim-based lead,
no aim-based dodge and no early warning available. Against DrussGT the
body-to-bullet angle has median **90.1 deg**, and the body ray passes within
10 deg of us on **0.0% of 1 794 ticks** — its gun is always on us, its body
never is.
**Recorded so the idea is not re-proposed:** the j166 lag-5 residue does improve
TTI (72.3 → 68.9) and **converts to contact 0% of the time**. A 4% TTI gain with
zero contact conversion is noise, not a lead.
### The honest remaining niches for exhaust/ram
1. **An opponent that closes on us.** Ram works whenever the other side comes to
us. Nothing here generalises away from that.
2. **A late-round exhaustion when they are already near.** The one conversion
ever recorded came from a *finisher* (enemy 16 → 1 energy), which is already
the default gate.
3. **Any 2v1+ mode**, where closing dynamics are not symmetric.
Against DrussGT specifically none of these will move the score, and
`TR_RAM_FLOOR_ENERGY` is under test in j163 — do not duplicate it.
---
## 2. What the ceiling implies
**If range is held, the only remaining lever is the gun at range, and the binding
numbers are the gun's, not the tile picker's.** The long-range hit rate is
**~9-10%** (Pattern live: 12.3% at 300-450 px, 9.2% at 450+; overall 10.5% —
`docs/headon_longrange_live.md`), the live hit half-window at 450 px is
**`atan(18/450) = 2.29°`** (`docs/gun_campaign.md:59`), and the measured arrival
aim error is **16.19° mean-abs at 450+** (`docs/bitbrain_campaign.md:107`,
`docs/headon_longrange_live.md:85`). 16.19° is **7× the window**. The tile picker
cannot close a 7× gap that sits downstream of the gun.
### New measurement — arrival aim error decomposed (j167, 23 275 shots at 450+ px)
Arrival aim error is defined non-circularly: the angle between the fired bearing
and the bearing to where the target *actually is* when the bullet arrives
(`tof = 20 - 3·power`, so the flight time comes from the power, not from the
shot's own geometry). It splits **exactly**, as signed angles, into
* **B, the model part** = the error the gun's own lead model leaves behind, and
* **C, manoeuvre** = the target's path curvature relative to the
constant-velocity extrapolation from the true state at fire time.
| band (px) | n | mean&#124;A&#124; | mean&#124;B&#124; (model) | mean&#124;C&#124; (manoeuvre) | sd(B) | sd(C) | corr(B,C) |
|---|---:|---:|---:|---:|---:|---:|---:|
| 0-100 | 24 767 | 83.36 | 101.56 | 49.38 | 124.5 | 75.9 | −0.65 |
| 300-450 | 11 372 | 13.07 | 21.77 | 11.00 | 26.0 | 13.0 | −0.87 |
| **450+** | **23 275** | **11.26** | **17.18** | **7.73** | **20.7** | **9.3** | **−0.85** |
At 450+ the model part's variance is **2.2× the manoeuvre part's**, and
`corr(B,C) = −0.85` means the two largely *cancel* — the net 11.26° is much
smaller than either part. **The 16° is a lead-model number, not a dodge number.**
Two supporting numbers: a naive constant-velocity extrapolation of a **2-tick-old**
position scores 8.45° mean-abs at 450+, and the time-of-flight implied by the
shot's own geometry (holding the current velocity) sits a **median 10 ticks short**
of the power-derived arrival tick (p10 −18, p90 +31) — i.e. the gun systematically
**under-leads in time**, consistent with `docs/lead_capture_by_range.md`
(capture 0.135 at 450+).
> **Do not read "a stale-CV model scores 8.45°" as "simplify the gun".** This is
> exactly the offline-ruler trap that killed HeadOn: the ruler said a no-lead gun
> was equal-or-better at 300+ and live it hit **20×/23× less**
> (`docs/headon_longrange_live.md`). The corpus is closed-loop — the target's
> manoeuvre is a *reaction to our own bullet* — so (B) and (C) are not separable
> here, and per `docs/offline_harness_trust.md` (j89: 0/6 on closed-loop) this
> instrument ranks per-gun single-tick prediction, it does not predict a live A/B.
### Cross-reference: what is still open in the gun docs
| doc | finding | status after this ceiling |
|---|---|---|
| `docs/gun_campaign.md:59` | hit half-window 2.29° at 450 px; measured signal 4.6-7.6° | **STILL OPEN and now the load-bearing number.** The decomposition says the gap is in the *model*, and the model is systematically 10 ticks short in time-of-flight. |
| `docs/gun_campaign.md:40-45` | lead amplitude is dead (1.0/1.5/2.0/3.0 all worse); radial knobs are bearing-invariant by construction | **CLOSED.** |
| `docs/gun_campaign.md:737-753` | `len6` +0.49 wins/run (p=0.039, n=15) did not replicate on n=33 | **CLOSED.** |
| `docs/bitbrain_campaign.md:189` | BitBrain / TMHorizon corrector adds no measurable aim (16.199 vs 16.193) | **CLOSED.** |
| `docs/bitbrain_campaign.md:107` | Pattern's own lead correlation with the required lead is 0.165 at 450+ | **STILL OPEN.** It is the same defect the decomposition names. |
| `docs/state_window_gate.md` | single wave-relative state at Q=4 predicts the miss bin at 0.4094 vs 0.2348 majority, but bins are 4.58-7.63° wide | **STILL OPEN, and now the best-placed surviving idea** — it is a *model* correction, which is where the error is. |
| `docs/gun_rack_analysis.md:423-455` | the 16-candidate rack ranking A/B found no winner; knobs added, all neutral | **CLOSED** (13 guns, `onlyPattern` shipped). |
---
## 3. Negative-results ledger — mechanisms closed by measurement
Do not re-litigate any row. The unit of evidence is the **opponent**.
| mechanism | knob / job | headline number | verdict |
|---|---|---|---|
| Geometry-weighted tile draw | `TR_TFIL_GEO_MODE/TAU`, j152 `38fbc6e`, A/B'd j159 `4a1f3e1` | **−8.83 damage/run, p=0.0061**; wins −0.05, p=0.46; +26.3 px mean distance on 15/15 opponents | **REJECTED.** Default off, stays off. |
| The bounded hold | `TR_TFIL_HOLD_MAX_TICKS`, j154 `2223ca6` | mechanism-positive, outcome-null (j146/j153) | **Default off.** No live win. |
| The proactive ram | `oldram` vs `base` gate `dist<200` | **p=0.69**, damage 279 vs 284, survival 17/49 vs 16/49; **0/59 opportunity→contact** | **CLOSED** (`docs/ramming_negative_result.md`). |
| The aim-based ram | j166 `a5a49bd` | body ray within 10° of us on **0.0% of 1 794 ticks**; body/barrel ray contact 40.0% vs 65.7% for doing nothing clever | **IMPOSSIBLE** — the server never sends gun direction (`TurnProcessor.kt:313-323`). |
| Arrival commitment (`tfil`) | j144 `d2005ab` | mechanism-positive, outcome-null | Default off. |
| Turn-cost tiebreak among safe tiles | j145 `39c90fd` | real but small mechanism, under-powered outcome null (300 battles, 5 arms) | Default off. |
| Field shape (safety) | j146 `de5d02b` | safe-set broken 63.5% → 30.4% offline; live null on damage and wins (375 battles, 5 arms) | **Default off.** |
| Corridor bound | j148 `5e213df` `TR_{TFIL,STRAFE}_CORRIDOR_TICKS` | never landed in a live A/B | Untested, not a candidate. |
| Ring arrival commitment | j165 `fe77056` `TR_TFIL_RING_COMMIT_ARRIVAL` | reach 0.24% → **3.05%**, picks 5 521 → 525, byte-for-byte default parity over 20 026 ticks, 148 guards | **Mechanism-positive, default off.** The strongest surviving movement mechanism. |
| Firing floor / enemy-exhaustion ram | j160 `23bce2d`, A/B'd j163 `51bfa57` | **clean negative**; the offline energy corpus missed the live game by 200× | **Under test in j163 — do not duplicate.** |
| Fire-detection lag | j147 `d21f7ce` `TR_FIRE_LAG` | displacement 19.06 → 5.37 px, deadline error 0.99 → 0.06 ticks; **live outcome-neutral**; ceiling ~10% of incoming damage (measured below) | **Default off, permanently.** |
| Hard arrival bound | j151 `a01141c` `TR_TFIL_ARRIVE_TICKS` | mechanism-positive, outcome-null | Default off. |
> **Methodological caution (j161), binding on everything above.** Pooled tests
> can hide real per-opponent effects: j159's safety signal was **p=0.0008
> per-opponent while the pooled test was null** (`docs/ram_floor_exhaustion_ab.md:219`).
> **Any future mechanism claim must report per-opponent mechanism metrics, not a
> pooled mean.** A pooled null is not evidence of absence; it is evidence that
> the heterogeneity was not averaged down.
---
## 4. Lead-time lever 1 — what a 2-tick-stale ghost really costs
`TR_FIRE_LAG` back-dates the bullet ghost (default 0). Energy-drop shot
detection lags **1.9 ticks mean**; median bullet flight is **19 ticks**
(`onHitByBullet` gives 82 hits / 7 421 ticks, one update per ~90 ticks).
**Measured on 55 750 incoming bullets** (`/tmp/tfil_ab2/out/`). For each bullet:
the time to closest approach of the target's recorded path to the bullet line
(**median 9 ticks**, p10 1, p90 39), and the minimum number of ticks of lead time
a max-speed hard-turn dodge needs to build 17 px of lateral displacement:
| minimum dodge lead time (ticks) | 0 | 1 | 2 | 3 | 4 | 5+ |
|---|---:|---:|---:|---:|---:|---:|
| share of incoming bullets | **57%** | 33% | 4% | 2% | 1% | 2% |
**57% of incoming bullets are already undodgeable at the instant they are fired**,
and only **~10%** (need ≥ 2 ticks) are in a regime where a 2-tick detection lag
can change anything. Applying the lag to the open-loop dodge model:
| ghost lag (ticks) | modelled hits | Δ vs perfect | share of all bullets whose hit/miss verdict flips |
|---|---:|---:|---:|
| 0 | 22 419 | — | — |
| **1.9 / 2** | **25 100** | **+2 681 (+12.0%)** | **10.18%** |
| 3 | 26 230 | +14.6% | 14.63% |
| 5 | 28 949 | +22.0% | 22.04% |
**Verdict: the 1.9-tick lag costs on the order of 10% more incoming hits** — at
the measured ~200 damage/run, roughly **20 damage/run**, an order of magnitude
below the movement A/B damage MDE. This is consistent with `TR_FIRE_LAG`'s already
measured live outcome-neutral result. The ghost is *wrong*, but wrongness at
10% of incoming damage cannot be turned into wins at this sample size.
**Recommendation: `TR_FIRE_LAG` stays off permanently.** It is a correctness fix
with a measured, bounded, sub-MDE payoff.
---
## 5. Lead-time lever 2 — the 16° decomposed, component by component
At 450+ px (23 275 shots), against the 11.26° net arrival error:
| component | measured | addressable? |
|---|---|---|
| **(a) enemy body-gun decoupling** | `\|gun dir − body heading\|` median **89.9°** (p10 25.9, p90 154.0, n=60 928). Extrapolating the target along its **gun** instead of its **body** would put the arrival bearing **79.5° median** wrong. | **Not present, and not addressable.** The intercept model uses the target's *recorded position and velocity*, both of which are the true body quantities and both exactly observed. Body-gun decoupling therefore contributes **exactly 0** to our arrival error. It is fatal for *aim-based* leading and threat warning (j166) and irrelevant to *position-based* leading. |
| **(b) our own leading model** | mean&#124;·&#124; **17.18°**, sd **20.7**; implied time-of-flight a **median 10 ticks short** of the power-derived arrival tick | **DOMINANT, and addressable.** This is ~2.2× the manoeuvre variance and it is the whole of the 16°. |
| **(c) target manoeuvre between scan and fire** | mean&#124;·&#124; **7.73°**, sd **9.3** | Small relative to (b), and **irreducible** — it is the dodger's own unpredictability, exactly the ~half of the under-lead `docs/lead_capture_by_range.md` attributes to a trivial predictor's own ceiling. |
| **(d) gun turn rate / time-to-fire** | the correct solution drifts a **median 0.416°/tick** (p90 5.45). The gun turns at 10°/tick, so a 17° correction takes **1.7 ticks ≈ 0.40°** of drift. | **Not binding.** Contributes ~**0.4°, i.e. ~3% of the 11.26° error.** The gun can always reach the answer; it aims at the wrong answer. |
**So the 16° is not (a), not (c) and not (d). It is (b) — the lead model's
time-of-flight, short by ~10 ticks.** Caveat, stated once and load-bearing: on a
closed-loop corpus (B) and (C) are not cleanly separable, since the target's
manoeuvre is a reaction to our own shot; the `corr(B,C) = −0.85` is exactly that
confound showing up. The *rank order* (b) ≫ (c) ≫ (d) > (a)=0 is robust to it
because (b) and (c) differ by 2.2× in variance and (d) is 3%.
---
## 6. The proposed lever: pre-multiply before learning — PREMISE DEAD
The design: aim error is largely a *product* (bearing-rate × time-of-flight), so
pre-multiply the two features and feed one small Tsetlin machine. **Measured on
the same corpus, the premise does not hold and the experiment should not be
built.** `y` = the required lead angle (current bearing → arrival bearing), i.e.
exactly the quantity the gun must predict; `b` = the observable 4-tick finite
difference of the bearing; `t = 20 − 3·power`.
| band (px) | n | corr(**b·t**, y) | corr(b+t, y) | R² additive [1,b,t] | R² product [1,b·t] | held-out side acc, additive | held-out side acc, product | held-out residual rms (deg) |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| 0-200 | 24 934 | −0.1020 | −0.1022 | 0.0105 | 0.0104 | 0.579 | 0.580 | 75.98 |
| 200-300 | 671 | 0.1190 | 0.0766 | 0.0147 | 0.0142 | 0.488 | 0.487 | 13.70 |
| 300-450 | 11 372 | **0.1501** | 0.0299 | 0.0256 | 0.0225 | 0.544 | 0.534 | 8.43 |
| **450+** | **23 275** | **0.2833** | 0.1052 | **0.0831** | 0.0803 | **0.603** | **0.601** | **6.68** |
| pooled | 60 252 | −0.1005 | −0.1009 | 0.0102 | 0.0101 | — | — | — |
*(side accuracy is 2-fold held-out and balanced; the TM record is ~0.47-0.49)*
Two things are true and the second kills the idea:
1. **As a single scalar, the product is much the better feature at range**:
`corr(b·t, y) = 0.283` vs `corr(b+t, y) = 0.105` at 450+ — 2.7× better, and
5× better at 300-450. So the *premise* ("the error is a product, not a sum")
is **confirmed as a statement about correlation**.
2. **But it buys nothing a weight-sum cannot already express.** The best linear
additive model on the same two features reaches **R² 0.0831 vs the product's
0.0803**, and the held-out balanced side accuracy is **0.603 (additive) vs
0.601 (product)** — a 0.002 difference, i.e. nothing. Pooled, the two are
identical (−0.1005 vs −0.1009; R² 0.0102 vs 0.0101). A TM with two input
features **already reconstructs the product term**; the multiplication is what
the network was doing anyway.
**Even the ceiling is out of reach.** The best held-out residual on the required
lead at 450+ is **6.68° rms**, against a live hit half-window of **2.29°** — a
2.9× shortfall. Pre-multiplying does not get a classifier to 2.29°; nothing in
this family does. **Do not build it.** The spec is recorded here so the idea is
closed on measurement rather than on taste.
*(Had it survived, the spec would have been: one TM, ONE input feature `b·t`
binarised on sign, plus the 4-bit horizon one-hot as today; offline gate =
held-out balanced side accuracy above 0.55 and residual rms below 3° at 450+;
live gate = wins/run with CI excluding 0 and sign-flip p<0.05 at 210 runs/arm,
damage not detectably down, MDE 0.17 wins/run. Predicted accuracy was 0.60 side
accuracy, which is a real signal against the 0.47-0.49 record — and still not
close enough to the window to convert.)*
---
## 7. The A/B queue, in priority order, with the MDE honestly restated
Throughput **22.7-23.2 runs/min**; movement gate resolved **0.17 wins/run at 210
runs/arm**; the `1/√n` extrapolation to 0.10 wins/run is **607 runs/arm ≈ 1.4 h —
a FLOOR on elapsed time, not an estimate**, because opponent heterogeneity does
not average down. A null at this sample size **only excludes a LARGE effect.**
(j163 additionally measured 14-22 runs/min, not 22.7-23.2, so even the floor is
optimistic.)
| # | experiment | what it tests | cost | a null would license |
|---|---|---|---|---|
| **1** | **The lead-model time-of-flight correction** (j167's (b)): re-derive the gun's arrival prediction so the implied flight is the power-derived tick, not 10 ticks short. | The one component that carries 2.2× the error variance at 450+, and the only open axis in `docs/gun_campaign.md` (lead *information*, not amplitude). | Offline gate first: arrival aim error at 450+ must fall below 11.26° mean-abs on held-out battles, ideally <8°; only then 2 arms × 15 opponents × 14 runs = 420 battles ≈ **0.3-0.4 h** wall. | Closing the single open gun axis. Nothing left in the gun. |
| 2 | `TR_TFIL_RING_COMMIT_ARRIVAL` (j165, default off) | Whether the largest surviving *movement* mechanism (reach 0.24% → 3.05%, picks 5 521 → 525, 148 guards) converts to wins. | 210 runs/arm ≈ **1.4 h floor**. | Retiring the whole ring/approach programme: if even a 12× reach gain is outcome-null, the ceiling argument is confirmed end to end. |
| 3 | `TR_FIRE_LAG` (tfil/strafe, default off) | Nothing worth testing — its ceiling is now measured at **~10% of incoming damage ≈ 20 dmg/run**, below the MDE. | Would be 1.4 h to learn nothing. | Nothing. **Skip it**; the measurement has already answered it. |
| 4 | `TR_TFIL_ARRIVE_TICKS` (j151, default off) | Whether a hard arrival bound converts now that the ring is rehabilitated. | 1.4 h. | Retiring it with j151's own null attached. |
| 5 | `TR_RAM_FLOOR_ENERGY` (j160, j163) | **Under test in j163. DO NOT DUPLICATE.** | — | — |
**Recommendation.** Run **only experiment 1**, and only after the *offline* gate
passes; if the offline gate does not move the 450+ arrival error below ~8°, run
nothing at all. Given five consecutive nulls or near-nulls (j144, j145, j146,
j147, j159) plus a clean negative in j163, spending 1.4 h of live time on
experiments 2-4 is not justified — those are mechanism-positive
mechanisms whose outcome nulls are already the standing record, and a null there
teaches nothing that the ledger does not already say.
**"The ceiling is real and we should stop spending on movement" is the answer.**
The remaining budget belongs to the gun's lead model, or it is not spent.