Files
SirRoboGarage/docs/range_vs_approach_ceiling.md
T

284 lines
20 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Range vs approach: melee is unavailable against DrussGT, so the gun at range is the only lever
**Date:** 2026-09-27 · **Job:** j167 · **Branch:** `research/lead-targeting`
**Evidence base:** j166 (`worktrees/j166-aim` @ `a5a49bd`), j165 (`fe77056`), j159
(`4a1f3e1`), j165/j151/j152/j154, j160/j163 (`0df7763` / `51bfa57`), j161
(`docs/ram_floor_exhaustion_ab.md:219`).
**Instrument for the new numbers below:** `worktrees/j167-ceiling/j167_probe.py`
(branch `j167-ceiling`) — pure replay of the recorded corpora
(`/tmp/tfil_ab2/out/`, 60 252 of our own scored shots over 140 battles; and
`/tmp/firelag_live2/`, 1 700 shots / 1 664 incoming bullets over 4 battles).
**No battle, A/B, server or GUI was run for this document.**
---
## 1. The ceiling
**Melee is structurally unavailable against DrussGT, and the exhaust/ram line is a
niche rather than a lever.** The evidence is two-sided and independent: (a) *our
mover's own ruler* — 94% of forced (no-safe-tile) picks happen at range > 300 u
and only 6-7% of picks reach the chosen tile at the estimated arrival time, with
the destination hot on arrival 35-42% of the time; and (b) *the j166 pursuit
probe* — 35 windows × 250 ticks of open-loop kinematics in which **every**
steering law is equal-or-worse than doing nothing clever:
| steering law (j166) | closing (u/tick) | contact % | TTI (ticks) |
|---|---:|---:|---:|
| current-position closing | 4.03 | 65.7 | 72.3 |
| body/barrel ray | 1.38 | 40.0 | 131.4 |
| velocity intercept (degenerate at equal speed) | — | 0 over 2 118 ticks | — |
| best case: lag-5 lead | 4.20 | 65.7 | 68.9 |
The root cause of the historical **0/59 proactive-ram** result
(`docs/ramming_negative_result.md`) is not a bad gate: **DrussGT never let the
distance drop.** Per-round minimum distance 152-338 u, median ~490 u, and
`frac(dist < 50) = 0.000` in all four recorded rounds. A pursuit that never gets
below 152 u cannot make contact, whatever the gate says. It is also not a
gun-side problem: **the server never transmits the enemy's gun direction**
(`ScannedBotEvent` = `energy, x, y, direction` where `direction` is the BODY
heading, plus `speed`; `TurnProcessor.kt:313-323`). There is no aim-based lead,
no aim-based dodge and no early warning available. Against DrussGT the
body-to-bullet angle has median **90.1 deg**, and the body ray passes within
10 deg of us on **0.0% of 1 794 ticks** — its gun is always on us, its body
never is.
**Recorded so the idea is not re-proposed:** the j166 lag-5 residue does improve
TTI (72.3 → 68.9) and **converts to contact 0% of the time**. A 4% TTI gain with
zero contact conversion is noise, not a lead.
### The honest remaining niches for exhaust/ram
1. **An opponent that closes on us.** Ram works whenever the other side comes to
us. Nothing here generalises away from that.
2. **A late-round exhaustion when they are already near.** The one conversion
ever recorded came from a *finisher* (enemy 16 → 1 energy), which is already
the default gate.
3. **Any 2v1+ mode**, where closing dynamics are not symmetric.
Against DrussGT specifically none of these will move the score, and
`TR_RAM_FLOOR_ENERGY` is under test in j163 — do not duplicate it.
---
## 2. What the ceiling implies
**If range is held, the only remaining lever is the gun at range, and the binding
numbers are the gun's, not the tile picker's.** The long-range hit rate is
**~9-10%** (Pattern live: 12.3% at 300-450 px, 9.2% at 450+; overall 10.5% —
`docs/headon_longrange_live.md`), the live hit half-window at 450 px is
**`atan(18/450) = 2.29°`** (`docs/gun_campaign.md:59`), and the measured arrival
aim error is **16.19° mean-abs at 450+** (`docs/bitbrain_campaign.md:107`,
`docs/headon_longrange_live.md:85`). 16.19° is **7× the window**. The tile picker
cannot close a 7× gap that sits downstream of the gun.
### New measurement — arrival aim error decomposed (j167, 23 275 shots at 450+ px)
Arrival aim error is defined non-circularly: the angle between the fired bearing
and the bearing to where the target *actually is* when the bullet arrives
(`tof = 20 - 3·power`, so the flight time comes from the power, not from the
shot's own geometry). It splits **exactly**, as signed angles, into
* **B, the model part** = the error the gun's own lead model leaves behind, and
* **C, manoeuvre** = the target's path curvature relative to the
constant-velocity extrapolation from the true state at fire time.
| band (px) | n | mean&#124;A&#124; | mean&#124;B&#124; (model) | mean&#124;C&#124; (manoeuvre) | sd(B) | sd(C) | corr(B,C) |
|---|---:|---:|---:|---:|---:|---:|---:|
| 0-100 | 24 767 | 83.36 | 101.56 | 49.38 | 124.5 | 75.9 | −0.65 |
| 300-450 | 11 372 | 13.07 | 21.77 | 11.00 | 26.0 | 13.0 | −0.87 |
| **450+** | **23 275** | **11.26** | **17.18** | **7.73** | **20.7** | **9.3** | **−0.85** |
At 450+ the model part's variance is **2.2× the manoeuvre part's**, and
`corr(B,C) = −0.85` means the two largely *cancel* — the net 11.26° is much
smaller than either part. **The 16° is a lead-model number, not a dodge number.**
Two supporting numbers: a naive constant-velocity extrapolation of a **2-tick-old**
position scores 8.45° mean-abs at 450+, and the time-of-flight implied by the
shot's own geometry (holding the current velocity) sits a **median 10 ticks short**
of the power-derived arrival tick (p10 −18, p90 +31) — i.e. the gun systematically
**under-leads in time**, consistent with `docs/lead_capture_by_range.md`
(capture 0.135 at 450+).
> **Do not read "a stale-CV model scores 8.45°" as "simplify the gun".** This is
> exactly the offline-ruler trap that killed HeadOn: the ruler said a no-lead gun
> was equal-or-better at 300+ and live it hit **20×/23× less**
> (`docs/headon_longrange_live.md`). The corpus is closed-loop — the target's
> manoeuvre is a *reaction to our own bullet* — so (B) and (C) are not separable
> here, and per `docs/offline_harness_trust.md` (j89: 0/6 on closed-loop) this
> instrument ranks per-gun single-tick prediction, it does not predict a live A/B.
### Cross-reference: what is still open in the gun docs
| doc | finding | status after this ceiling |
|---|---|---|
| `docs/gun_campaign.md:59` | hit half-window 2.29° at 450 px; measured signal 4.6-7.6° | **STILL OPEN and now the load-bearing number.** The decomposition says the gap is in the *model*, and the model is systematically 10 ticks short in time-of-flight. |
| `docs/gun_campaign.md:40-45` | lead amplitude is dead (1.0/1.5/2.0/3.0 all worse); radial knobs are bearing-invariant by construction | **CLOSED.** |
| `docs/gun_campaign.md:737-753` | `len6` +0.49 wins/run (p=0.039, n=15) did not replicate on n=33 | **CLOSED.** |
| `docs/bitbrain_campaign.md:189` | BitBrain / TMHorizon corrector adds no measurable aim (16.199 vs 16.193) | **CLOSED.** |
| `docs/bitbrain_campaign.md:107` | Pattern's own lead correlation with the required lead is 0.165 at 450+ | **STILL OPEN.** It is the same defect the decomposition names. |
| `docs/state_window_gate.md` | single wave-relative state at Q=4 predicts the miss bin at 0.4094 vs 0.2348 majority, but bins are 4.58-7.63° wide | **STILL OPEN, and now the best-placed surviving idea** — it is a *model* correction, which is where the error is. |
| `docs/gun_rack_analysis.md:423-455` | the 16-candidate rack ranking A/B found no winner; knobs added, all neutral | **CLOSED** (13 guns, `onlyPattern` shipped). |
---
## 3. Negative-results ledger — mechanisms closed by measurement
Do not re-litigate any row. The unit of evidence is the **opponent**.
| mechanism | knob / job | headline number | verdict |
|---|---|---|---|
| Geometry-weighted tile draw | `TR_TFIL_GEO_MODE/TAU`, j152 `38fbc6e`, A/B'd j159 `4a1f3e1` | **−8.83 damage/run, p=0.0061**; wins −0.05, p=0.46; +26.3 px mean distance on 15/15 opponents | **REJECTED.** Default off, stays off. |
| The bounded hold | `TR_TFIL_HOLD_MAX_TICKS`, j154 `2223ca6` | mechanism-positive, outcome-null (j146/j153) | **Default off.** No live win. |
| The proactive ram | `oldram` vs `base` gate `dist<200` | **p=0.69**, damage 279 vs 284, survival 17/49 vs 16/49; **0/59 opportunity→contact** | **CLOSED** (`docs/ramming_negative_result.md`). |
| The aim-based ram | j166 `a5a49bd` | body ray within 10° of us on **0.0% of 1 794 ticks**; body/barrel ray contact 40.0% vs 65.7% for doing nothing clever | **IMPOSSIBLE** — the server never sends gun direction (`TurnProcessor.kt:313-323`). |
| Arrival commitment (`tfil`) | j144 `d2005ab` | mechanism-positive, outcome-null | Default off. |
| Turn-cost tiebreak among safe tiles | j145 `39c90fd` | real but small mechanism, under-powered outcome null (300 battles, 5 arms) | Default off. |
| Field shape (safety) | j146 `de5d02b` | safe-set broken 63.5% → 30.4% offline; live null on damage and wins (375 battles, 5 arms) | **Default off.** |
| Corridor bound | j148 `5e213df` `TR_{TFIL,STRAFE}_CORRIDOR_TICKS` | never landed in a live A/B | Untested, not a candidate. |
| Ring arrival commitment | j165 `fe77056` `TR_TFIL_RING_COMMIT_ARRIVAL` | reach 0.24% → **3.05%**, picks 5 521 → 525, byte-for-byte default parity over 20 026 ticks, 148 guards | **Mechanism-positive, default off.** The strongest surviving movement mechanism. |
| Firing floor / enemy-exhaustion ram | j160 `23bce2d`, A/B'd j163 `51bfa57` | **clean negative**; the offline energy corpus missed the live game by 200× | **Under test in j163 — do not duplicate.** |
| Fire-detection lag | j147 `d21f7ce` `TR_FIRE_LAG` | displacement 19.06 → 5.37 px, deadline error 0.99 → 0.06 ticks; **live outcome-neutral**; ceiling ~10% of incoming damage (measured below) | **Default off, permanently.** |
| Hard arrival bound | j151 `a01141c` `TR_TFIL_ARRIVE_TICKS` | mechanism-positive, outcome-null | Default off. |
> **Methodological caution (j161), binding on everything above.** Pooled tests
> can hide real per-opponent effects: j159's safety signal was **p=0.0008
> per-opponent while the pooled test was null** (`docs/ram_floor_exhaustion_ab.md:219`).
> **Any future mechanism claim must report per-opponent mechanism metrics, not a
> pooled mean.** A pooled null is not evidence of absence; it is evidence that
> the heterogeneity was not averaged down.
---
## 4. Lead-time lever 1 — what a 2-tick-stale ghost really costs
`TR_FIRE_LAG` back-dates the bullet ghost (default 0). Energy-drop shot
detection lags **1.9 ticks mean**; median bullet flight is **19 ticks**
(`onHitByBullet` gives 82 hits / 7 421 ticks, one update per ~90 ticks).
**Measured on 55 750 incoming bullets** (`/tmp/tfil_ab2/out/`). For each bullet:
the time to closest approach of the target's recorded path to the bullet line
(**median 9 ticks**, p10 1, p90 39), and the minimum number of ticks of lead time
a max-speed hard-turn dodge needs to build 17 px of lateral displacement:
| minimum dodge lead time (ticks) | 0 | 1 | 2 | 3 | 4 | 5+ |
|---|---:|---:|---:|---:|---:|---:|
| share of incoming bullets | **57%** | 33% | 4% | 2% | 1% | 2% |
**57% of incoming bullets are already undodgeable at the instant they are fired**,
and only **~10%** (need ≥ 2 ticks) are in a regime where a 2-tick detection lag
can change anything. Applying the lag to the open-loop dodge model:
| ghost lag (ticks) | modelled hits | Δ vs perfect | share of all bullets whose hit/miss verdict flips |
|---|---:|---:|---:|
| 0 | 22 419 | — | — |
| **1.9 / 2** | **25 100** | **+2 681 (+12.0%)** | **10.18%** |
| 3 | 26 230 | +14.6% | 14.63% |
| 5 | 28 949 | +22.0% | 22.04% |
**Verdict: the 1.9-tick lag costs on the order of 10% more incoming hits** — at
the measured ~200 damage/run, roughly **20 damage/run**, an order of magnitude
below the movement A/B damage MDE. This is consistent with `TR_FIRE_LAG`'s already
measured live outcome-neutral result. The ghost is *wrong*, but wrongness at
10% of incoming damage cannot be turned into wins at this sample size.
**Recommendation: `TR_FIRE_LAG` stays off permanently.** It is a correctness fix
with a measured, bounded, sub-MDE payoff.
---
## 5. Lead-time lever 2 — the 16° decomposed, component by component
At 450+ px (23 275 shots), against the 11.26° net arrival error:
| component | measured | addressable? |
|---|---|---|
| **(a) enemy body-gun decoupling** | `\|gun dir − body heading\|` median **89.9°** (p10 25.9, p90 154.0, n=60 928). Extrapolating the target along its **gun** instead of its **body** would put the arrival bearing **79.5° median** wrong. | **Not present, and not addressable.** The intercept model uses the target's *recorded position and velocity*, both of which are the true body quantities and both exactly observed. Body-gun decoupling therefore contributes **exactly 0** to our arrival error. It is fatal for *aim-based* leading and threat warning (j166) and irrelevant to *position-based* leading. |
| **(b) our own leading model** | mean&#124;·&#124; **17.18°**, sd **20.7**; implied time-of-flight a **median 10 ticks short** of the power-derived arrival tick | **DOMINANT, and addressable.** This is ~2.2× the manoeuvre variance and it is the whole of the 16°. |
| **(c) target manoeuvre between scan and fire** | mean&#124;·&#124; **7.73°**, sd **9.3** | Small relative to (b), and **irreducible** — it is the dodger's own unpredictability, exactly the ~half of the under-lead `docs/lead_capture_by_range.md` attributes to a trivial predictor's own ceiling. |
| **(d) gun turn rate / time-to-fire** | the correct solution drifts a **median 0.416°/tick** (p90 5.45). The gun turns at 10°/tick, so a 17° correction takes **1.7 ticks ≈ 0.40°** of drift. | **Not binding.** Contributes ~**0.4°, i.e. ~3% of the 11.26° error.** The gun can always reach the answer; it aims at the wrong answer. |
**So the 16° is not (a), not (c) and not (d). It is (b) — the lead model's
time-of-flight, short by ~10 ticks.** Caveat, stated once and load-bearing: on a
closed-loop corpus (B) and (C) are not cleanly separable, since the target's
manoeuvre is a reaction to our own shot; the `corr(B,C) = −0.85` is exactly that
confound showing up. The *rank order* (b) ≫ (c) ≫ (d) > (a)=0 is robust to it
because (b) and (c) differ by 2.2× in variance and (d) is 3%.
---
## 6. The proposed lever: pre-multiply before learning — PREMISE DEAD
The design: aim error is largely a *product* (bearing-rate × time-of-flight), so
pre-multiply the two features and feed one small Tsetlin machine. **Measured on
the same corpus, the premise does not hold and the experiment should not be
built.** `y` = the required lead angle (current bearing → arrival bearing), i.e.
exactly the quantity the gun must predict; `b` = the observable 4-tick finite
difference of the bearing; `t = 20 − 3·power`.
| band (px) | n | corr(**b·t**, y) | corr(b+t, y) | R² additive [1,b,t] | R² product [1,b·t] | held-out side acc, additive | held-out side acc, product | held-out residual rms (deg) |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| 0-200 | 24 934 | −0.1020 | −0.1022 | 0.0105 | 0.0104 | 0.579 | 0.580 | 75.98 |
| 200-300 | 671 | 0.1190 | 0.0766 | 0.0147 | 0.0142 | 0.488 | 0.487 | 13.70 |
| 300-450 | 11 372 | **0.1501** | 0.0299 | 0.0256 | 0.0225 | 0.544 | 0.534 | 8.43 |
| **450+** | **23 275** | **0.2833** | 0.1052 | **0.0831** | 0.0803 | **0.603** | **0.601** | **6.68** |
| pooled | 60 252 | −0.1005 | −0.1009 | 0.0102 | 0.0101 | — | — | — |
*(side accuracy is 2-fold held-out and balanced; the TM record is ~0.47-0.49)*
Two things are true and the second kills the idea:
1. **As a single scalar, the product is much the better feature at range**:
`corr(b·t, y) = 0.283` vs `corr(b+t, y) = 0.105` at 450+ — 2.7× better, and
5× better at 300-450. So the *premise* ("the error is a product, not a sum")
is **confirmed as a statement about correlation**.
2. **But it buys nothing a weight-sum cannot already express.** The best linear
additive model on the same two features reaches **R² 0.0831 vs the product's
0.0803**, and the held-out balanced side accuracy is **0.603 (additive) vs
0.601 (product)** — a 0.002 difference, i.e. nothing. Pooled, the two are
identical (−0.1005 vs −0.1009; R² 0.0102 vs 0.0101). A TM with two input
features **already reconstructs the product term**; the multiplication is what
the network was doing anyway.
**Even the ceiling is out of reach.** The best held-out residual on the required
lead at 450+ is **6.68° rms**, against a live hit half-window of **2.29°** — a
2.9× shortfall. Pre-multiplying does not get a classifier to 2.29°; nothing in
this family does. **Do not build it.** The spec is recorded here so the idea is
closed on measurement rather than on taste.
*(Had it survived, the spec would have been: one TM, ONE input feature `b·t`
binarised on sign, plus the 4-bit horizon one-hot as today; offline gate =
held-out balanced side accuracy above 0.55 and residual rms below 3° at 450+;
live gate = wins/run with CI excluding 0 and sign-flip p<0.05 at 210 runs/arm,
damage not detectably down, MDE 0.17 wins/run. Predicted accuracy was 0.60 side
accuracy, which is a real signal against the 0.47-0.49 record — and still not
close enough to the window to convert.)*
---
## 7. The A/B queue, in priority order, with the MDE honestly restated
Throughput **22.7-23.2 runs/min**; movement gate resolved **0.17 wins/run at 210
runs/arm**; the `1/√n` extrapolation to 0.10 wins/run is **607 runs/arm ≈ 1.4 h —
a FLOOR on elapsed time, not an estimate**, because opponent heterogeneity does
not average down. A null at this sample size **only excludes a LARGE effect.**
(j163 additionally measured 14-22 runs/min, not 22.7-23.2, so even the floor is
optimistic.)
| # | experiment | what it tests | cost | a null would license |
|---|---|---|---|---|
| **1** | **The lead-model time-of-flight correction** (j167's (b)): re-derive the gun's arrival prediction so the implied flight is the power-derived tick, not 10 ticks short. | The one component that carries 2.2× the error variance at 450+, and the only open axis in `docs/gun_campaign.md` (lead *information*, not amplitude). | Offline gate first: arrival aim error at 450+ must fall below 11.26° mean-abs on held-out battles, ideally <8°; only then 2 arms × 15 opponents × 14 runs = 420 battles ≈ **0.3-0.4 h** wall. | Closing the single open gun axis. Nothing left in the gun. |
| 2 | `TR_TFIL_RING_COMMIT_ARRIVAL` (j165, default off) | Whether the largest surviving *movement* mechanism (reach 0.24% → 3.05%, picks 5 521 → 525, 148 guards) converts to wins. | 210 runs/arm ≈ **1.4 h floor**. | Retiring the whole ring/approach programme: if even a 12× reach gain is outcome-null, the ceiling argument is confirmed end to end. |
| 3 | `TR_FIRE_LAG` (tfil/strafe, default off) | Nothing worth testing — its ceiling is now measured at **~10% of incoming damage ≈ 20 dmg/run**, below the MDE. | Would be 1.4 h to learn nothing. | Nothing. **Skip it**; the measurement has already answered it. |
| 4 | `TR_TFIL_ARRIVE_TICKS` (j151, default off) | Whether a hard arrival bound converts now that the ring is rehabilitated. | 1.4 h. | Retiring it with j151's own null attached. |
| 5 | `TR_RAM_FLOOR_ENERGY` (j160, j163) | **Under test in j163. DO NOT DUPLICATE.** | — | — |
**Recommendation.** Run **only experiment 1**, and only after the *offline* gate
passes; if the offline gate does not move the 450+ arrival error below ~8°, run
nothing at all. Given five consecutive nulls or near-nulls (j144, j145, j146,
j147, j159) plus a clean negative in j163, spending 1.4 h of live time on
experiments 2-4 is not justified — those are mechanism-positive
mechanisms whose outcome nulls are already the standing record, and a null there
teaches nothing that the ledger does not already say.
**"The ceiling is real and we should stop spending on movement" is the answer.**
The remaining budget belongs to the gun's lead model, or it is not spent.