Files
SirRoboGarage/docs/lead_capture_by_range.md
SirStone f91e121965 lead capture by range: we under-lead (0.40->0.135) — a RANGE effect, not a power one
Measure, for every shot ModularBot fires at the real DrussGT, the lead we
actually applied vs the lead the enemy's motion required, from the recorded
live battles (/tmp/tfil_ab2, 70 battles / 490 rounds / 54926 shots, plus a
35-battle powtest replication of a different binary).

- requiredLead from an AIM-INDEPENDENT interception solve (bullet speed vs
  enemy truth), appliedLead from the server-recorded bullet bearing.
- capture = applied/required, guarded at 2px lateral lead (1.6% excluded);
  headline metric is the robust proportional slope.
- validation: hits 11.6px mean miss / 80.8% inside 18px, misses 134px,
  11.6x separation; 496/496 death + 70/70 owner attributions correct.

Direct answer: capture falls with RANGE (capSlp 0.401 -> 0.135, and
|err|/tolerance 1.27 -> 7.54) but is FLAT across fired POWER within a band
(450+, enemy alive: 0.154 / 0.127 / 0.127). The sub-0.5 long-range shots
(1.18% hit) are finishKill endgame shots at a near-dead DrussGT, not a
lead-capture failure. A naive linear predictor captures 0.29-0.60; we reach
46-67% of that, so the under-lead is real but capture=1.0 is unattainable
against a dodger (oracle required lead).
2026-09-24 22:47:18 +02:00

308 lines
17 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Lead capture by range: how much of the required lead do we actually apply?
**Question (the user's hypothesis).** *"DrussGT is dodging more when low power only
because DrussGT is moving faster, more far away from the head-on and our displacement
capacity, so we try to hit him but we miss as we are not capable of reaching the
correct angle."*
**Answer: half right, and the half that is wrong matters.** We genuinely **do not
apply the required lead, and the shortfall grows with range** — capture falls from
**0.40 at 0–200 px to 0.135 at 450+ px** on 54 926 shots, so at long range we put
**~13 %** of the lead a perfect gun would need into the bullet, and miss by ~137 px
against an 18 px bot. But this is a **RANGE** effect, **not a POWER** effect: inside a
range band, capture is the same at 0.5 power as at 1.0 power. Low-power shots miss
more because *they are long-range shots*, plus a separate endgame confound, **not**
because low power costs us displacement capacity.
The other half: capture below 1.0 is **partly unavoidable**. `requiredLead` uses the
enemy's *actual* future dodge (perfect information), which no real gun can know. A
trivial straight-line predictor — the same bullet, aimed at the intercept of the
enemy's last 4 ticks of velocity — captures only **0.29–0.60**. We capture **46–67 %**
of that trivial ceiling. So the under-lead is real and fixable, but roughly half of
the distance from us to "perfect lead" is the dodger's intrinsic unpredictability.
Measured on **70 real live battles / 490 rounds / 54 926 shots** against the real,
unmodified DrussGT (`/tmp/tfil_ab2/out/`), replicated on **35 more battles / 24 277
shots** of a *different* ModularBot build (`/tmp/powtest/`). All live Tank Royale
battles recorded through `tools/robocode_shim/run_bridge_battle.sh`; no offline
fixture replay.
---
## 1. The decomposition
For every shot **we (ModularBot)** fire at DrussGT, with power `p`, all angles in
degrees relative to the **line of sight (LOS) at the fire tick**:
* `O` = the firing tank's centre at the fire tick (= the bullet-line origin; the
server's fire `(x,y)` is the tank centre to 0.02 px, verified);
* `v = 20 − 3p` px/tick, so **low power is a faster bullet**;
* `t*` = the **aim-independent** interception tick: the first tick `k` with
`|E(t₀+k) − O| ≤ v·k`, where `E` is DrussGT's recorded true position. It depends
only on the enemy's truth and the bullet speed, **never on our aim**;
* **`requiredLead`** = `bearing(O → E(t*)) − LOS`;
* **`appliedLead`** = our bullet's server-recorded bearing `− LOS`;
* **`leadError`** = `appliedLead − requiredLead` (wrapped to ±180°);
* **`capture`** = `appliedLead / requiredLead`.
`capture` is guarded: it is defined only when the required lateral lead is at least
**2 px** at the fire range (`|requiredLead|·range ≥ 2 px`). This excludes **887 of
54 926 shots (1.6 %)**. The mean of the *ratio* is a noisy statistic (small
denominators); the headline capture is therefore the **proportional slope**
`Σ(applied·required)/Σ(required²)` computed cell-by-cell, labelled `capSlp`. The mean
ratio (`capt`) and its median (`medcap`) are printed too and tell the same story.
Arrival-adjacent quantities use two ticks, both non-circular:
* `miss_px` = perpendicular distance between DrussGT's true position and our bullet's
real line at the tick the bullet reaches DrussGT's along-track plane — the **same
definition used in `docs/drussgt_dodge_vs_power.md`**, so the validation numbers
match that job exactly;
* `flight` = `t*` (ticks in the air).
### Attribution (stated explicitly, this has bitten the project before)
* In the capture rows **`e*` is the SUBJECT = DrussGT** and **`s*` is the adversary =
ModularBot = us**, by construction of
`tools/robocode_shim/src/robocode_shim/TrBattleCapture.java` (`en` = the bot whose
name contains "DrussGT", written `e*`; `sh` = the other, written `s*`).
* The event sidecar's `owner` id is a Tank Royale id and **is not stable across runs**.
It is recovered **per battle** from the fire geometry: the owner's position equals
the fire event's `(x,y)` and its energy drops by exactly `power` on the next capture
row. **70/70 battles** resolve to the two sides `('e','s')`.
* Cross-check: for **496/496 death events** the mapped victim is the bot whose energy
is ~0 at the end of that round.
* Fingerprint check on the way round: our recovered side fires
`{1.00 ×32044, 0.50 ×14556, 0.10 ×807, …}` — the documented ModularBot policy
(spike at `TR_POWER_ENERGY_MIN = 0.5` under 20 energy, spike at the 1.0 range cap
beyond `TR_POWER_FAR_DIST = 200`, linear slope between). DrussGT's recovered side
fires `{0.15 ×27209, 0.95 ×22789, 0.45 ×5274, …}` — a completely different,
distance-quantised curve. **No power-value heuristic is used anywhere** (the old
`{1.0, 1.5, 2.0, 3.0}` assumption is exactly what mis-attributed ModularBot in a
previous job).
Because the four-decimal positions repeat when a tank is stationary, 12 112 fire
events have more than one *position* match inside the ±8-tick search window; the
energy-drop term of the match disambiguates all of them, and the 496/496 death check
plus the hit geometry below validate the result end to end.
---
## 2. Geometry validation: hits separate from misses
| | n | mean &#124;leadError&#124; (deg) | mean &#124;leadError&#124; (px) | miss_px mean | miss_px median | fraction < 18 px |
|---|---|---|---|---|---|---|
| **HITS** (server truth) | 5 480 | **1.478** | 11.4 | **11.6** | 10.8 | **80.8 %** |
| **MISSES** (hitwall/hitbullet) | 48 304 | **16.724** | 141.3 | **134.1** | 123.7 | **0.8 %** |
Miss separation ratio (misses/hits) = **11.59×**. The hits show a mean miss of
**11.6 px** and **80.8 % inside 18 px**, matching the previously recorded live
measurement (11.6 px, 80.6 %) exactly — the bearing recovery is correct. Overall
measured hit rate **9.98 %**.
---
## 3. Main result: by range band
`capt` = mean of the ratio, `capSlp` = proportional capture (robust), `capLin` = the
same slope for the **naive linear predictor control**, `|req|`/`|app|` = mean *absolute*
lead magnitude in degrees, `tol` = angular tolerance `atan(18/range)`.
| range | n | range px | flight | &#124;req&#124;° | &#124;app&#124;° | capt | **capSlp** | capLin | &#124;err&#124;° | &#124;err&#124;px | tol° | &#124;err&#124;/tol | hit % |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0–100 | 27 | 81 | 10.3 | 24.8 | 13.4 | 0.72 | **0.401** | 0.597 | 16.8 | 23.2 | 13.01 | 1.27 | 44.00 |
| 100–200 | 274 | 162 | 14.7 | 20.2 | 17.0 | 0.07 | **0.391** | 0.600 | 17.7 | 50.1 | 6.57 | 2.70 | 25.28 |
| 200–300 | 990 | 261 | 17.9 | 15.7 | 13.8 | 0.04 | **0.246** | 0.428 | 16.7 | 75.8 | 3.99 | 4.17 | 16.04 |
| 300–450 | 17 178 | 405 | 24.5 | 13.5 | 13.1 | 0.01 | **0.198** | 0.404 | 16.0 | 112.9 | 2.56 | 6.23 | 11.79 |
| 450+ | 36 457 | 535 | 31.0 | 11.6 | 11.4 | 0.15 | **0.135** | 0.291 | 14.7 | 137.0 | 1.95 | 7.54 | 9.14 |
Read this table twice. Two things happen at once and they are different:
1. **The angular error `|leadError|` is roughly constant (~15–18°) at every range**
(it does not shrink with distance).
2. **The window shrinks**: the 18 px bot spans 13° at 100 px but only 1.95° at 500 px,
because `tol = atan(18/range)`.
So `|err|/tol` grows from **1.27 to 7.54**: the same angular miss that was survivable
up close becomes fatal at distance. This is the "tighter angular window" half of the
story — but it is *not* the whole story, because `capSlp` genuinely **falls with range**
(0.40 → 0.135). We are not applying a constant fraction of a shrinking lead; we are
applying a *shrinking fraction* of a roughly constant-magnitude lead.
---
## 4. Power does NOT change capture (the 'range vs power' split)
Within the 450+ band, restricting to shots fired while **DrussGT still has ≥ 5 energy**
(which removes the endgame confound of §5):
| 450+, enemy alive | n | mean p | range px | **capSlp** | capLin | &#124;err&#124;px | hit % |
|---|---|---|---|---|---|---|---|
| 0.50–0.75 | 11 882 | 0.53 | 535 | **0.154** | 0.312 | 130.3 | 10.39 |
| 0.75–1.00 | 2 377 | 0.87 | 540 | **0.127** | 0.275 | 140.7 | 9.60 |
| 1.00–1.50 | 20 101 | 1.00 | 534 | **0.127** | 0.279 | 140.4 | 8.80 |
**Flat.** Same result in the 300–450 band (`capSlp` 0.220 / 0.205 / 0.191 for the same
three power bins). At 450+ our policy fires only 0.50 or 1.00 in this regime, so the
comparison is exactly "our fastest long-range bullet" (p = 0.5, v = 18.5, flight 29.5
ticks) versus "our standard bullet" (p = 1.0, v = 17, flight 32 ticks). The faster
low-power bullet has a **shorter** flight and a **smaller** required lead by
construction, and we capture **the same fraction** of it. Power is not the driver.
---
## 5. The endgame confound (why raw per-power hit rates lie)
Our power policy fires **sub-0.5 power only to finish a nearly-dead enemy**. At 450+,
**944 of the 950 sub-0.5 shots have the enemy's energy below 2** (mean **1.1**), because
`prFinishKill` cuts the power when the enemy's remaining energy is already tiny. Those
shots hit at **1.18 %**, which looks like a catastrophic low-power failure — it is not:
| 450+, by enemy energy at the fire tick | n | mean p | capSlp | hit % |
|---|---|---|---|---|
| [0, 2) | 1 017 | 0.22 | 0.111 | **1.10** |
| [2, 5) | 1 076 | 0.54 | 0.128 | 8.55 |
| [5, 10) | 2 614 | 0.59 | 0.135 | 9.69 |
| [10, 20) | 6 758 | 0.63 | 0.146 | 9.85 |
| [20, 40) | 10 733 | 0.83 | 0.147 | 9.43 |
| [40, 150) | 14 259 | 0.97 | 0.122 | 9.12 |
The hit rate collapses **only when the target is already effectively dead**, and it
collapses for **high-power shots too** (73 shots at p ≥ 0.5 against a < 2-energy
DrussGT hit 0.00 %). Every other enemy-energy stratum sits at 8.5–9.9 %. So the
"low power hits 1 %" reading is a *dead-target* artifact, not a lead-capture failure.
Capture itself is flat across enemy energy (0.111 → 0.147 → 0.122).
---
## 6. Do we move the aim in proportion to the required lead? (response curve)
At 450+, split the shots into deciles of `requiredLead` and look at the **mean applied
lead** in each:
| requiredLead decile (mean) | −21.6 | −13.7 | −7.8 | −2.5 | +2.5 | +8.1 | +14.3 | +22.3 |
|---|---|---|---|---|---|---|---|---|
| **mean appliedLead** | −4.09 | −1.51 | −1.00 | −0.55 | +0.36 | −0.16 | +0.14 | +3.26 |
| hit % | 9.27 | 7.92 | 10.29 | 10.08 | 9.95 | 8.75 | 7.22 | 9.69 |
The required lead swings across **±22°**, and our mean applied lead responds by about
**±4°** — and the hit rate does **not** depend on the size of the required lead at all.
This is the cleanest single number in the report: the bullet is aimed, to first order,
**at the line of sight, not at the interception point**. Consistently,
`mean|appliedLead| ≈ mean|requiredLead|` (11.4° vs 11.6° at 450+) — we have as much
*dispersion* of aim as the target has motion, but almost **no correlation** with it
(`capSlp = 0.135`).
---
## 7. The ceiling: what a trivial predictor would do
`requiredLead` is an **oracle** (it uses DrussGT's actual future dodge). To separate
"our gun is bad" from "the dodge is unpredictable", the analyzer also computes a
**naive linear-predictor control**: same bullet speed, aimed at the intercept of a
straight-line continuation of DrussGT's last 4 ticks of velocity. Its capture
(`capLin`) and the ratio are in §3:
| range | our capSlp | naive-linear capLin | **us / linear** |
|---|---|---|---|
| 0–100 | 0.401 | 0.597 | 0.67 |
| 100–200 | 0.391 | 0.600 | 0.65 |
| 200–300 | 0.246 | 0.428 | 0.58 |
| 300–450 | 0.198 | 0.404 | 0.49 |
| 450+ | 0.135 | 0.291 | **0.46** |
So a *trivial* predictor still only reaches 0.29 at long range — most of the oracle
lead is genuinely unattainable against a strong dodger. But we reach only **46 %** of
that trivial benchmark. The under-lead is therefore real and worth fixing, while the
gap from the trivial benchmark to 1.0 is the dodger's unpredictability and is not.
Corroboration: the powtest corpus — a **different** ModularBot binary — reproduces the
shape almost exactly (`capSlp` 0.529 / 0.354 / 0.283 / 0.219 / 0.136 by the same range
bands). It is a property of the architecture, not of one build.
---
## 8. Caveats and sample sizes
* **The oracle caveat.** `requiredLead` uses perfect future information. `capture = 1.0`
is *unattainable* against a bot that dodges; it is a diagnostic, not a target. Compare
our capture to the `capLin` control, not to 1.0.
* **Weak correlation, large dispersion.** At long range our applied lead is essentially
an aim with mean 0 and ~11.5° dispersion that is only weakly proportional to the
required lead (proportional slope 0.135). "We apply 13 % of the lead"
is a proportional-fit statement; per shot, what we actually do is *aim near the LOS
with a wide error*.
* **Stale live world state.** The recorded `dir` is what the **live** ModularBot fired
using its own stale between-scan world state; the capture's positions are perfect
truth. The measured `leadError` therefore folds in the bot's scan staleness. That is
the correct thing to measure (it is the error the bullet actually carries), but it is
not the same as the gun's internal prediction error.
* **Stratification.** Power is not randomised: it is assigned by the range cap, the
energy slope and the finishing rule. Every power comparison above is made **within a
range band** and (in §4) with the endgame removed. The raw cross-range power
distribution is reported in the captured output.
* **Sample sizes.** The short bands are thin (0–100: n = 27; 100–200: n = 274) — treat
their capture as indicative. Everything from 200 px out is large (990 → 36 457).
The 450+ sub-0.5 power cell *with the enemy alive* is n = 4 and should be ignored
(it is printed for completeness; the real sub-0.5 shots are the endgame of §5).
* **Rotation direction.** `requiredLead`/`appliedLead` are signed about the LOS;
`capSlp` additionally assumes a proportional (through-origin) relation, which is the
right first-order model for a lead gun but not exact for large angles.
---
## 9. Verdict
**MEASURED**
* Capture falls monotonically with range: **capSlp 0.401 → 0.391 → 0.246 → 0.198 →
0.135**. At 450+ we apply **~13 %** of the required lead and miss by **~137 px**
against an 18 px bot; `|leadError|/tolerance` grows **1.27 → 7.54**.
* Capture is **flat across fired power within a range band**
(450+, enemy alive: 0.154 / 0.127 / 0.127 for p = 0.53 / 0.87 / 1.00).
* The applied lead barely responds to the required lead: across required-lead deciles
spanning ±22°, the mean applied lead moves by ~±4° and the hit rate is flat.
* A naive linear predictor captures 0.29–0.60; we capture **46–67 %** of it.
* Hits validate the geometry (11.6 px mean, 80.8 % inside 18 px) and separate from
misses by **11.6×**; **496/496** deaths and **70/70** owner maps resolve correctly.
* The sub-0.5-power long-range shots (1.18 % hit) are **endgame shots at a near-dead
DrussGT** (mean enemy energy 1.1), and high-power shots there miss just as much.
**INFERRED**
* The gun is, to first order, **a line-of-sight / weak-lead gun against DrussGT**, not
an intercept-point gun: the bullet direction tracks the enemy's current line rather
than where the enemy will be, which is why capture is low and why it degrades with
range while the angular error stays ~constant.
* The user's mechanism — "we are not capable of reaching the correct angle at range" —
is **confirmed for range**, but **not through power**: low power does not reduce our
displacement capacity here (it is a faster bullet and we capture the same fraction).
The low-power/long-range association is a policy artifact (the range cap), so
*fixing the lead at range fixes the low-power miss rate too*.
* Where to look next: the gap between our capture and the trivial-linear control
(0.46×) is the actionable, fixable part. The gap from the trivial control to 1.0 is
the dodger's unpredictability and should not be chased.
---
## 10. Reproducing
```bash
# corpora (live captures; not in the repo, ~500 MB total)
# /tmp/tfil_ab2/out/<A..E>/run*.jsonl{,.events.jsonl,.rounds.json}
# /tmp/powtest/{cap_<arm>_r<n>.jsonl,events_<arm>_r<n>.json,cap_*_r<n>.jsonl.rounds.json}
python3 common_libs/tests/analyze_lead_capture_by_range.py \
--tfil /tmp/tfil_ab2/out --powtest /tmp/powtest \
--json common_libs/tests/fixtures/lead_capture_by_range_results.json \
| tee common_libs/tests/fixtures/lead_capture_by_range_output.txt
```
`common_libs/tests/fixtures/lead_capture_by_range_output.txt` is the verbatim captured
output every table above is taken from; `..._results.json` is the same numbers as JSON.
Runtime ≈ 15 s for both corpora, single-threaded pure Python (no numpy).
To rebuild a corpus, `tools/robocode_shim/make_botdir.sh /tmp/tr_bots/DrussGT` and then
`TR_EVENTS_OUT=<run>.events.jsonl tools/robocode_shim/run_bridge_battle.sh
<adversaryBotDir> 7 <run>.jsonl`.