lead capture by range: we under-lead (0.40->0.135) — a RANGE effect, not a power one
Measure, for every shot ModularBot fires at the real DrussGT, the lead we actually applied vs the lead the enemy's motion required, from the recorded live battles (/tmp/tfil_ab2, 70 battles / 490 rounds / 54926 shots, plus a 35-battle powtest replication of a different binary). - requiredLead from an AIM-INDEPENDENT interception solve (bullet speed vs enemy truth), appliedLead from the server-recorded bullet bearing. - capture = applied/required, guarded at 2px lateral lead (1.6% excluded); headline metric is the robust proportional slope. - validation: hits 11.6px mean miss / 80.8% inside 18px, misses 134px, 11.6x separation; 496/496 death + 70/70 owner attributions correct. Direct answer: capture falls with RANGE (capSlp 0.401 -> 0.135, and |err|/tolerance 1.27 -> 7.54) but is FLAT across fired POWER within a band (450+, enemy alive: 0.154 / 0.127 / 0.127). The sub-0.5 long-range shots (1.18% hit) are finishKill endgame shots at a near-dead DrussGT, not a lead-capture failure. A naive linear predictor captures 0.29-0.60; we reach 46-67% of that, so the under-lead is real but capture=1.0 is unattainable against a dodger (oracle required lead).
This commit is contained in:
@@ -0,0 +1,307 @@
|
||||
# Lead capture by range: how much of the required lead do we actually apply?
|
||||
|
||||
**Question (the user's hypothesis).** *"DrussGT is dodging more when low power only
|
||||
because DrussGT is moving faster, more far away from the head-on and our displacement
|
||||
capacity, so we try to hit him but we miss as we are not capable of reaching the
|
||||
correct angle."*
|
||||
|
||||
**Answer: half right, and the half that is wrong matters.** We genuinely **do not
|
||||
apply the required lead, and the shortfall grows with range** — capture falls from
|
||||
**0.40 at 0–200 px to 0.135 at 450+ px** on 54 926 shots, so at long range we put
|
||||
**~13 %** of the lead a perfect gun would need into the bullet, and miss by ~137 px
|
||||
against an 18 px bot. But this is a **RANGE** effect, **not a POWER** effect: inside a
|
||||
range band, capture is the same at 0.5 power as at 1.0 power. Low-power shots miss
|
||||
more because *they are long-range shots*, plus a separate endgame confound, **not**
|
||||
because low power costs us displacement capacity.
|
||||
|
||||
The other half: capture below 1.0 is **partly unavoidable**. `requiredLead` uses the
|
||||
enemy's *actual* future dodge (perfect information), which no real gun can know. A
|
||||
trivial straight-line predictor — the same bullet, aimed at the intercept of the
|
||||
enemy's last 4 ticks of velocity — captures only **0.29–0.60**. We capture **46–67 %**
|
||||
of that trivial ceiling. So the under-lead is real and fixable, but roughly half of
|
||||
the distance from us to "perfect lead" is the dodger's intrinsic unpredictability.
|
||||
|
||||
Measured on **70 real live battles / 490 rounds / 54 926 shots** against the real,
|
||||
unmodified DrussGT (`/tmp/tfil_ab2/out/`), replicated on **35 more battles / 24 277
|
||||
shots** of a *different* ModularBot build (`/tmp/powtest/`). All live Tank Royale
|
||||
battles recorded through `tools/robocode_shim/run_bridge_battle.sh`; no offline
|
||||
fixture replay.
|
||||
|
||||
---
|
||||
|
||||
## 1. The decomposition
|
||||
|
||||
For every shot **we (ModularBot)** fire at DrussGT, with power `p`, all angles in
|
||||
degrees relative to the **line of sight (LOS) at the fire tick**:
|
||||
|
||||
* `O` = the firing tank's centre at the fire tick (= the bullet-line origin; the
|
||||
server's fire `(x,y)` is the tank centre to 0.02 px, verified);
|
||||
* `v = 20 − 3p` px/tick, so **low power is a faster bullet**;
|
||||
* `t*` = the **aim-independent** interception tick: the first tick `k` with
|
||||
`|E(t₀+k) − O| ≤ v·k`, where `E` is DrussGT's recorded true position. It depends
|
||||
only on the enemy's truth and the bullet speed, **never on our aim**;
|
||||
* **`requiredLead`** = `bearing(O → E(t*)) − LOS`;
|
||||
* **`appliedLead`** = our bullet's server-recorded bearing `− LOS`;
|
||||
* **`leadError`** = `appliedLead − requiredLead` (wrapped to ±180°);
|
||||
* **`capture`** = `appliedLead / requiredLead`.
|
||||
|
||||
`capture` is guarded: it is defined only when the required lateral lead is at least
|
||||
**2 px** at the fire range (`|requiredLead|·range ≥ 2 px`). This excludes **887 of
|
||||
54 926 shots (1.6 %)**. The mean of the *ratio* is a noisy statistic (small
|
||||
denominators); the headline capture is therefore the **proportional slope**
|
||||
`Σ(applied·required)/Σ(required²)` computed cell-by-cell, labelled `capSlp`. The mean
|
||||
ratio (`capt`) and its median (`medcap`) are printed too and tell the same story.
|
||||
|
||||
Arrival-adjacent quantities use two ticks, both non-circular:
|
||||
|
||||
* `miss_px` = perpendicular distance between DrussGT's true position and our bullet's
|
||||
real line at the tick the bullet reaches DrussGT's along-track plane — the **same
|
||||
definition used in `docs/drussgt_dodge_vs_power.md`**, so the validation numbers
|
||||
match that job exactly;
|
||||
* `flight` = `t*` (ticks in the air).
|
||||
|
||||
### Attribution (stated explicitly, this has bitten the project before)
|
||||
|
||||
* In the capture rows **`e*` is the SUBJECT = DrussGT** and **`s*` is the adversary =
|
||||
ModularBot = us**, by construction of
|
||||
`tools/robocode_shim/src/robocode_shim/TrBattleCapture.java` (`en` = the bot whose
|
||||
name contains "DrussGT", written `e*`; `sh` = the other, written `s*`).
|
||||
* The event sidecar's `owner` id is a Tank Royale id and **is not stable across runs**.
|
||||
It is recovered **per battle** from the fire geometry: the owner's position equals
|
||||
the fire event's `(x,y)` and its energy drops by exactly `power` on the next capture
|
||||
row. **70/70 battles** resolve to the two sides `('e','s')`.
|
||||
* Cross-check: for **496/496 death events** the mapped victim is the bot whose energy
|
||||
is ~0 at the end of that round.
|
||||
* Fingerprint check on the way round: our recovered side fires
|
||||
`{1.00 ×32044, 0.50 ×14556, 0.10 ×807, …}` — the documented ModularBot policy
|
||||
(spike at `TR_POWER_ENERGY_MIN = 0.5` under 20 energy, spike at the 1.0 range cap
|
||||
beyond `TR_POWER_FAR_DIST = 200`, linear slope between). DrussGT's recovered side
|
||||
fires `{0.15 ×27209, 0.95 ×22789, 0.45 ×5274, …}` — a completely different,
|
||||
distance-quantised curve. **No power-value heuristic is used anywhere** (the old
|
||||
`{1.0, 1.5, 2.0, 3.0}` assumption is exactly what mis-attributed ModularBot in a
|
||||
previous job).
|
||||
|
||||
Because the four-decimal positions repeat when a tank is stationary, 12 112 fire
|
||||
events have more than one *position* match inside the ±8-tick search window; the
|
||||
energy-drop term of the match disambiguates all of them, and the 496/496 death check
|
||||
plus the hit geometry below validate the result end to end.
|
||||
|
||||
---
|
||||
|
||||
## 2. Geometry validation: hits separate from misses
|
||||
|
||||
| | n | mean |leadError| (deg) | mean |leadError| (px) | miss_px mean | miss_px median | fraction < 18 px |
|
||||
|---|---|---|---|---|---|---|
|
||||
| **HITS** (server truth) | 5 480 | **1.478** | 11.4 | **11.6** | 10.8 | **80.8 %** |
|
||||
| **MISSES** (hitwall/hitbullet) | 48 304 | **16.724** | 141.3 | **134.1** | 123.7 | **0.8 %** |
|
||||
|
||||
Miss separation ratio (misses/hits) = **11.59×**. The hits show a mean miss of
|
||||
**11.6 px** and **80.8 % inside 18 px**, matching the previously recorded live
|
||||
measurement (11.6 px, 80.6 %) exactly — the bearing recovery is correct. Overall
|
||||
measured hit rate **9.98 %**.
|
||||
|
||||
---
|
||||
|
||||
## 3. Main result: by range band
|
||||
|
||||
`capt` = mean of the ratio, `capSlp` = proportional capture (robust), `capLin` = the
|
||||
same slope for the **naive linear predictor control**, `|req|`/`|app|` = mean *absolute*
|
||||
lead magnitude in degrees, `tol` = angular tolerance `atan(18/range)`.
|
||||
|
||||
| range | n | range px | flight | |req|° | |app|° | capt | **capSlp** | capLin | |err|° | |err|px | tol° | |err|/tol | hit % |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| 0–100 | 27 | 81 | 10.3 | 24.8 | 13.4 | 0.72 | **0.401** | 0.597 | 16.8 | 23.2 | 13.01 | 1.27 | 44.00 |
|
||||
| 100–200 | 274 | 162 | 14.7 | 20.2 | 17.0 | 0.07 | **0.391** | 0.600 | 17.7 | 50.1 | 6.57 | 2.70 | 25.28 |
|
||||
| 200–300 | 990 | 261 | 17.9 | 15.7 | 13.8 | 0.04 | **0.246** | 0.428 | 16.7 | 75.8 | 3.99 | 4.17 | 16.04 |
|
||||
| 300–450 | 17 178 | 405 | 24.5 | 13.5 | 13.1 | 0.01 | **0.198** | 0.404 | 16.0 | 112.9 | 2.56 | 6.23 | 11.79 |
|
||||
| 450+ | 36 457 | 535 | 31.0 | 11.6 | 11.4 | 0.15 | **0.135** | 0.291 | 14.7 | 137.0 | 1.95 | 7.54 | 9.14 |
|
||||
|
||||
Read this table twice. Two things happen at once and they are different:
|
||||
|
||||
1. **The angular error `|leadError|` is roughly constant (~15–18°) at every range**
|
||||
(it does not shrink with distance).
|
||||
2. **The window shrinks**: the 18 px bot spans 13° at 100 px but only 1.95° at 500 px,
|
||||
because `tol = atan(18/range)`.
|
||||
|
||||
So `|err|/tol` grows from **1.27 to 7.54**: the same angular miss that was survivable
|
||||
up close becomes fatal at distance. This is the "tighter angular window" half of the
|
||||
story — but it is *not* the whole story, because `capSlp` genuinely **falls with range**
|
||||
(0.40 → 0.135). We are not applying a constant fraction of a shrinking lead; we are
|
||||
applying a *shrinking fraction* of a roughly constant-magnitude lead.
|
||||
|
||||
---
|
||||
|
||||
## 4. Power does NOT change capture (the 'range vs power' split)
|
||||
|
||||
Within the 450+ band, restricting to shots fired while **DrussGT still has ≥ 5 energy**
|
||||
(which removes the endgame confound of §5):
|
||||
|
||||
| 450+, enemy alive | n | mean p | range px | **capSlp** | capLin | |err|px | hit % |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| 0.50–0.75 | 11 882 | 0.53 | 535 | **0.154** | 0.312 | 130.3 | 10.39 |
|
||||
| 0.75–1.00 | 2 377 | 0.87 | 540 | **0.127** | 0.275 | 140.7 | 9.60 |
|
||||
| 1.00–1.50 | 20 101 | 1.00 | 534 | **0.127** | 0.279 | 140.4 | 8.80 |
|
||||
|
||||
**Flat.** Same result in the 300–450 band (`capSlp` 0.220 / 0.205 / 0.191 for the same
|
||||
three power bins). At 450+ our policy fires only 0.50 or 1.00 in this regime, so the
|
||||
comparison is exactly "our fastest long-range bullet" (p = 0.5, v = 18.5, flight 29.5
|
||||
ticks) versus "our standard bullet" (p = 1.0, v = 17, flight 32 ticks). The faster
|
||||
low-power bullet has a **shorter** flight and a **smaller** required lead by
|
||||
construction, and we capture **the same fraction** of it. Power is not the driver.
|
||||
|
||||
---
|
||||
|
||||
## 5. The endgame confound (why raw per-power hit rates lie)
|
||||
|
||||
Our power policy fires **sub-0.5 power only to finish a nearly-dead enemy**. At 450+,
|
||||
**944 of the 950 sub-0.5 shots have the enemy's energy below 2** (mean **1.1**), because
|
||||
`prFinishKill` cuts the power when the enemy's remaining energy is already tiny. Those
|
||||
shots hit at **1.18 %**, which looks like a catastrophic low-power failure — it is not:
|
||||
|
||||
| 450+, by enemy energy at the fire tick | n | mean p | capSlp | hit % |
|
||||
|---|---|---|---|---|
|
||||
| [0, 2) | 1 017 | 0.22 | 0.111 | **1.10** |
|
||||
| [2, 5) | 1 076 | 0.54 | 0.128 | 8.55 |
|
||||
| [5, 10) | 2 614 | 0.59 | 0.135 | 9.69 |
|
||||
| [10, 20) | 6 758 | 0.63 | 0.146 | 9.85 |
|
||||
| [20, 40) | 10 733 | 0.83 | 0.147 | 9.43 |
|
||||
| [40, 150) | 14 259 | 0.97 | 0.122 | 9.12 |
|
||||
|
||||
The hit rate collapses **only when the target is already effectively dead**, and it
|
||||
collapses for **high-power shots too** (73 shots at p ≥ 0.5 against a < 2-energy
|
||||
DrussGT hit 0.00 %). Every other enemy-energy stratum sits at 8.5–9.9 %. So the
|
||||
"low power hits 1 %" reading is a *dead-target* artifact, not a lead-capture failure.
|
||||
Capture itself is flat across enemy energy (0.111 → 0.147 → 0.122).
|
||||
|
||||
---
|
||||
|
||||
## 6. Do we move the aim in proportion to the required lead? (response curve)
|
||||
|
||||
At 450+, split the shots into deciles of `requiredLead` and look at the **mean applied
|
||||
lead** in each:
|
||||
|
||||
| requiredLead decile (mean) | −21.6 | −13.7 | −7.8 | −2.5 | +2.5 | +8.1 | +14.3 | +22.3 |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| **mean appliedLead** | −4.09 | −1.51 | −1.00 | −0.55 | +0.36 | −0.16 | +0.14 | +3.26 |
|
||||
| hit % | 9.27 | 7.92 | 10.29 | 10.08 | 9.95 | 8.75 | 7.22 | 9.69 |
|
||||
|
||||
The required lead swings across **±22°**, and our mean applied lead responds by about
|
||||
**±4°** — and the hit rate does **not** depend on the size of the required lead at all.
|
||||
This is the cleanest single number in the report: the bullet is aimed, to first order,
|
||||
**at the line of sight, not at the interception point**. Consistently,
|
||||
`mean|appliedLead| ≈ mean|requiredLead|` (11.4° vs 11.6° at 450+) — we have as much
|
||||
*dispersion* of aim as the target has motion, but almost **no correlation** with it
|
||||
(`capSlp = 0.135`).
|
||||
|
||||
---
|
||||
|
||||
## 7. The ceiling: what a trivial predictor would do
|
||||
|
||||
`requiredLead` is an **oracle** (it uses DrussGT's actual future dodge). To separate
|
||||
"our gun is bad" from "the dodge is unpredictable", the analyzer also computes a
|
||||
**naive linear-predictor control**: same bullet speed, aimed at the intercept of a
|
||||
straight-line continuation of DrussGT's last 4 ticks of velocity. Its capture
|
||||
(`capLin`) and the ratio are in §3:
|
||||
|
||||
| range | our capSlp | naive-linear capLin | **us / linear** |
|
||||
|---|---|---|---|
|
||||
| 0–100 | 0.401 | 0.597 | 0.67 |
|
||||
| 100–200 | 0.391 | 0.600 | 0.65 |
|
||||
| 200–300 | 0.246 | 0.428 | 0.58 |
|
||||
| 300–450 | 0.198 | 0.404 | 0.49 |
|
||||
| 450+ | 0.135 | 0.291 | **0.46** |
|
||||
|
||||
So a *trivial* predictor still only reaches 0.29 at long range — most of the oracle
|
||||
lead is genuinely unattainable against a strong dodger. But we reach only **46 %** of
|
||||
that trivial benchmark. The under-lead is therefore real and worth fixing, while the
|
||||
gap from the trivial benchmark to 1.0 is the dodger's unpredictability and is not.
|
||||
|
||||
Corroboration: the powtest corpus — a **different** ModularBot binary — reproduces the
|
||||
shape almost exactly (`capSlp` 0.529 / 0.354 / 0.283 / 0.219 / 0.136 by the same range
|
||||
bands). It is a property of the architecture, not of one build.
|
||||
|
||||
---
|
||||
|
||||
## 8. Caveats and sample sizes
|
||||
|
||||
* **The oracle caveat.** `requiredLead` uses perfect future information. `capture = 1.0`
|
||||
is *unattainable* against a bot that dodges; it is a diagnostic, not a target. Compare
|
||||
our capture to the `capLin` control, not to 1.0.
|
||||
* **Weak correlation, large dispersion.** At long range our applied lead is essentially
|
||||
an aim with mean 0 and ~11.5° dispersion that is only weakly proportional to the
|
||||
required lead (proportional slope 0.135). "We apply 13 % of the lead"
|
||||
is a proportional-fit statement; per shot, what we actually do is *aim near the LOS
|
||||
with a wide error*.
|
||||
* **Stale live world state.** The recorded `dir` is what the **live** ModularBot fired
|
||||
using its own stale between-scan world state; the capture's positions are perfect
|
||||
truth. The measured `leadError` therefore folds in the bot's scan staleness. That is
|
||||
the correct thing to measure (it is the error the bullet actually carries), but it is
|
||||
not the same as the gun's internal prediction error.
|
||||
* **Stratification.** Power is not randomised: it is assigned by the range cap, the
|
||||
energy slope and the finishing rule. Every power comparison above is made **within a
|
||||
range band** and (in §4) with the endgame removed. The raw cross-range power
|
||||
distribution is reported in the captured output.
|
||||
* **Sample sizes.** The short bands are thin (0–100: n = 27; 100–200: n = 274) — treat
|
||||
their capture as indicative. Everything from 200 px out is large (990 → 36 457).
|
||||
The 450+ sub-0.5 power cell *with the enemy alive* is n = 4 and should be ignored
|
||||
(it is printed for completeness; the real sub-0.5 shots are the endgame of §5).
|
||||
* **Rotation direction.** `requiredLead`/`appliedLead` are signed about the LOS;
|
||||
`capSlp` additionally assumes a proportional (through-origin) relation, which is the
|
||||
right first-order model for a lead gun but not exact for large angles.
|
||||
|
||||
---
|
||||
|
||||
## 9. Verdict
|
||||
|
||||
**MEASURED**
|
||||
|
||||
* Capture falls monotonically with range: **capSlp 0.401 → 0.391 → 0.246 → 0.198 →
|
||||
0.135**. At 450+ we apply **~13 %** of the required lead and miss by **~137 px**
|
||||
against an 18 px bot; `|leadError|/tolerance` grows **1.27 → 7.54**.
|
||||
* Capture is **flat across fired power within a range band**
|
||||
(450+, enemy alive: 0.154 / 0.127 / 0.127 for p = 0.53 / 0.87 / 1.00).
|
||||
* The applied lead barely responds to the required lead: across required-lead deciles
|
||||
spanning ±22°, the mean applied lead moves by ~±4° and the hit rate is flat.
|
||||
* A naive linear predictor captures 0.29–0.60; we capture **46–67 %** of it.
|
||||
* Hits validate the geometry (11.6 px mean, 80.8 % inside 18 px) and separate from
|
||||
misses by **11.6×**; **496/496** deaths and **70/70** owner maps resolve correctly.
|
||||
* The sub-0.5-power long-range shots (1.18 % hit) are **endgame shots at a near-dead
|
||||
DrussGT** (mean enemy energy 1.1), and high-power shots there miss just as much.
|
||||
|
||||
**INFERRED**
|
||||
|
||||
* The gun is, to first order, **a line-of-sight / weak-lead gun against DrussGT**, not
|
||||
an intercept-point gun: the bullet direction tracks the enemy's current line rather
|
||||
than where the enemy will be, which is why capture is low and why it degrades with
|
||||
range while the angular error stays ~constant.
|
||||
* The user's mechanism — "we are not capable of reaching the correct angle at range" —
|
||||
is **confirmed for range**, but **not through power**: low power does not reduce our
|
||||
displacement capacity here (it is a faster bullet and we capture the same fraction).
|
||||
The low-power/long-range association is a policy artifact (the range cap), so
|
||||
*fixing the lead at range fixes the low-power miss rate too*.
|
||||
* Where to look next: the gap between our capture and the trivial-linear control
|
||||
(0.46×) is the actionable, fixable part. The gap from the trivial control to 1.0 is
|
||||
the dodger's unpredictability and should not be chased.
|
||||
|
||||
---
|
||||
|
||||
## 10. Reproducing
|
||||
|
||||
```bash
|
||||
# corpora (live captures; not in the repo, ~500 MB total)
|
||||
# /tmp/tfil_ab2/out/<A..E>/run*.jsonl{,.events.jsonl,.rounds.json}
|
||||
# /tmp/powtest/{cap_<arm>_r<n>.jsonl,events_<arm>_r<n>.json,cap_*_r<n>.jsonl.rounds.json}
|
||||
|
||||
python3 common_libs/tests/analyze_lead_capture_by_range.py \
|
||||
--tfil /tmp/tfil_ab2/out --powtest /tmp/powtest \
|
||||
--json common_libs/tests/fixtures/lead_capture_by_range_results.json \
|
||||
| tee common_libs/tests/fixtures/lead_capture_by_range_output.txt
|
||||
```
|
||||
|
||||
`common_libs/tests/fixtures/lead_capture_by_range_output.txt` is the verbatim captured
|
||||
output every table above is taken from; `..._results.json` is the same numbers as JSON.
|
||||
Runtime ≈ 15 s for both corpora, single-threaded pure Python (no numpy).
|
||||
|
||||
To rebuild a corpus, `tools/robocode_shim/make_botdir.sh /tmp/tr_bots/DrussGT` and then
|
||||
`TR_EVENTS_OUT=<run>.events.jsonl tools/robocode_shim/run_bridge_battle.sh
|
||||
<adversaryBotDir> 7 <run>.jsonl`.
|
||||
Reference in New Issue
Block a user