Files
SirRoboGarage/docs/lead_capture_by_range.md
T
SirStone f91e121965 lead capture by range: we under-lead (0.40->0.135) — a RANGE effect, not a power one
Measure, for every shot ModularBot fires at the real DrussGT, the lead we
actually applied vs the lead the enemy's motion required, from the recorded
live battles (/tmp/tfil_ab2, 70 battles / 490 rounds / 54926 shots, plus a
35-battle powtest replication of a different binary).

- requiredLead from an AIM-INDEPENDENT interception solve (bullet speed vs
  enemy truth), appliedLead from the server-recorded bullet bearing.
- capture = applied/required, guarded at 2px lateral lead (1.6% excluded);
  headline metric is the robust proportional slope.
- validation: hits 11.6px mean miss / 80.8% inside 18px, misses 134px,
  11.6x separation; 496/496 death + 70/70 owner attributions correct.

Direct answer: capture falls with RANGE (capSlp 0.401 -> 0.135, and
|err|/tolerance 1.27 -> 7.54) but is FLAT across fired POWER within a band
(450+, enemy alive: 0.154 / 0.127 / 0.127). The sub-0.5 long-range shots
(1.18% hit) are finishKill endgame shots at a near-dead DrussGT, not a
lead-capture failure. A naive linear predictor captures 0.29-0.60; we reach
46-67% of that, so the under-lead is real but capture=1.0 is unattainable
against a dodger (oracle required lead).
2026-09-24 22:47:18 +02:00

17 KiB
Raw Blame History

Lead capture by range: how much of the required lead do we actually apply?

Question (the user's hypothesis). "DrussGT is dodging more when low power only because DrussGT is moving faster, more far away from the head-on and our displacement capacity, so we try to hit him but we miss as we are not capable of reaching the correct angle."

Answer: half right, and the half that is wrong matters. We genuinely do not apply the required lead, and the shortfall grows with range — capture falls from 0.40 at 0–200 px to 0.135 at 450+ px on 54 926 shots, so at long range we put ~13 % of the lead a perfect gun would need into the bullet, and miss by ~137 px against an 18 px bot. But this is a RANGE effect, not a POWER effect: inside a range band, capture is the same at 0.5 power as at 1.0 power. Low-power shots miss more because they are long-range shots, plus a separate endgame confound, not because low power costs us displacement capacity.

The other half: capture below 1.0 is partly unavoidable. requiredLead uses the enemy's actual future dodge (perfect information), which no real gun can know. A trivial straight-line predictor — the same bullet, aimed at the intercept of the enemy's last 4 ticks of velocity — captures only 0.29–0.60. We capture 46–67 % of that trivial ceiling. So the under-lead is real and fixable, but roughly half of the distance from us to "perfect lead" is the dodger's intrinsic unpredictability.

Measured on 70 real live battles / 490 rounds / 54 926 shots against the real, unmodified DrussGT (/tmp/tfil_ab2/out/), replicated on 35 more battles / 24 277 shots of a different ModularBot build (/tmp/powtest/). All live Tank Royale battles recorded through tools/robocode_shim/run_bridge_battle.sh; no offline fixture replay.


1. The decomposition

For every shot we (ModularBot) fire at DrussGT, with power p, all angles in degrees relative to the line of sight (LOS) at the fire tick:

  • O = the firing tank's centre at the fire tick (= the bullet-line origin; the server's fire (x,y) is the tank centre to 0.02 px, verified);
  • v = 20 − 3p px/tick, so low power is a faster bullet;
  • t* = the aim-independent interception tick: the first tick k with |E(t₀+k) − O| ≤ v·k, where E is DrussGT's recorded true position. It depends only on the enemy's truth and the bullet speed, never on our aim;
  • requiredLead = bearing(O → E(t*)) − LOS;
  • appliedLead = our bullet's server-recorded bearing − LOS;
  • leadError = appliedLead − requiredLead (wrapped to ±180°);
  • capture = appliedLead / requiredLead.

capture is guarded: it is defined only when the required lateral lead is at least 2 px at the fire range (|requiredLead|·range ≥ 2 px). This excludes 887 of 54 926 shots (1.6 %). The mean of the ratio is a noisy statistic (small denominators); the headline capture is therefore the proportional slope Σ(applied·required)/Σ(required²) computed cell-by-cell, labelled capSlp. The mean ratio (capt) and its median (medcap) are printed too and tell the same story.

Arrival-adjacent quantities use two ticks, both non-circular:

  • miss_px = perpendicular distance between DrussGT's true position and our bullet's real line at the tick the bullet reaches DrussGT's along-track plane — the same definition used in docs/drussgt_dodge_vs_power.md, so the validation numbers match that job exactly;
  • flight = t* (ticks in the air).

Attribution (stated explicitly, this has bitten the project before)

  • In the capture rows e* is the SUBJECT = DrussGT and s* is the adversary = ModularBot = us, by construction of tools/robocode_shim/src/robocode_shim/TrBattleCapture.java (en = the bot whose name contains "DrussGT", written e*; sh = the other, written s*).
  • The event sidecar's owner id is a Tank Royale id and is not stable across runs. It is recovered per battle from the fire geometry: the owner's position equals the fire event's (x,y) and its energy drops by exactly power on the next capture row. 70/70 battles resolve to the two sides ('e','s').
  • Cross-check: for 496/496 death events the mapped victim is the bot whose energy is ~0 at the end of that round.
  • Fingerprint check on the way round: our recovered side fires {1.00 ×32044, 0.50 ×14556, 0.10 ×807, …} — the documented ModularBot policy (spike at TR_POWER_ENERGY_MIN = 0.5 under 20 energy, spike at the 1.0 range cap beyond TR_POWER_FAR_DIST = 200, linear slope between). DrussGT's recovered side fires {0.15 ×27209, 0.95 ×22789, 0.45 ×5274, …} — a completely different, distance-quantised curve. No power-value heuristic is used anywhere (the old {1.0, 1.5, 2.0, 3.0} assumption is exactly what mis-attributed ModularBot in a previous job).

Because the four-decimal positions repeat when a tank is stationary, 12 112 fire events have more than one position match inside the ±8-tick search window; the energy-drop term of the match disambiguates all of them, and the 496/496 death check plus the hit geometry below validate the result end to end.


2. Geometry validation: hits separate from misses

n mean |leadError| (deg) mean |leadError| (px) miss_px mean miss_px median fraction < 18 px
HITS (server truth) 5 480 1.478 11.4 11.6 10.8 80.8 %
MISSES (hitwall/hitbullet) 48 304 16.724 141.3 134.1 123.7 0.8 %

Miss separation ratio (misses/hits) = 11.59×. The hits show a mean miss of 11.6 px and 80.8 % inside 18 px, matching the previously recorded live measurement (11.6 px, 80.6 %) exactly — the bearing recovery is correct. Overall measured hit rate 9.98 %.


3. Main result: by range band

capt = mean of the ratio, capSlp = proportional capture (robust), capLin = the same slope for the naive linear predictor control, |req|/|app| = mean absolute lead magnitude in degrees, tol = angular tolerance atan(18/range).

range n range px flight |req|° |app|° capt capSlp capLin |err|° |err|px tol° |err|/tol hit %
0–100 27 81 10.3 24.8 13.4 0.72 0.401 0.597 16.8 23.2 13.01 1.27 44.00
100–200 274 162 14.7 20.2 17.0 0.07 0.391 0.600 17.7 50.1 6.57 2.70 25.28
200–300 990 261 17.9 15.7 13.8 0.04 0.246 0.428 16.7 75.8 3.99 4.17 16.04
300–450 17 178 405 24.5 13.5 13.1 0.01 0.198 0.404 16.0 112.9 2.56 6.23 11.79
450+ 36 457 535 31.0 11.6 11.4 0.15 0.135 0.291 14.7 137.0 1.95 7.54 9.14

Read this table twice. Two things happen at once and they are different:

  1. The angular error |leadError| is roughly constant (~15–18°) at every range (it does not shrink with distance).
  2. The window shrinks: the 18 px bot spans 13° at 100 px but only 1.95° at 500 px, because tol = atan(18/range).

So |err|/tol grows from 1.27 to 7.54: the same angular miss that was survivable up close becomes fatal at distance. This is the "tighter angular window" half of the story — but it is not the whole story, because capSlp genuinely falls with range (0.40 → 0.135). We are not applying a constant fraction of a shrinking lead; we are applying a shrinking fraction of a roughly constant-magnitude lead.


4. Power does NOT change capture (the 'range vs power' split)

Within the 450+ band, restricting to shots fired while DrussGT still has ≥ 5 energy (which removes the endgame confound of §5):

450+, enemy alive n mean p range px capSlp capLin |err|px hit %
0.50–0.75 11 882 0.53 535 0.154 0.312 130.3 10.39
0.75–1.00 2 377 0.87 540 0.127 0.275 140.7 9.60
1.00–1.50 20 101 1.00 534 0.127 0.279 140.4 8.80

Flat. Same result in the 300–450 band (capSlp 0.220 / 0.205 / 0.191 for the same three power bins). At 450+ our policy fires only 0.50 or 1.00 in this regime, so the comparison is exactly "our fastest long-range bullet" (p = 0.5, v = 18.5, flight 29.5 ticks) versus "our standard bullet" (p = 1.0, v = 17, flight 32 ticks). The faster low-power bullet has a shorter flight and a smaller required lead by construction, and we capture the same fraction of it. Power is not the driver.


5. The endgame confound (why raw per-power hit rates lie)

Our power policy fires sub-0.5 power only to finish a nearly-dead enemy. At 450+, 944 of the 950 sub-0.5 shots have the enemy's energy below 2 (mean 1.1), because prFinishKill cuts the power when the enemy's remaining energy is already tiny. Those shots hit at 1.18 %, which looks like a catastrophic low-power failure — it is not:

450+, by enemy energy at the fire tick n mean p capSlp hit %
[0, 2) 1 017 0.22 0.111 1.10
[2, 5) 1 076 0.54 0.128 8.55
[5, 10) 2 614 0.59 0.135 9.69
[10, 20) 6 758 0.63 0.146 9.85
[20, 40) 10 733 0.83 0.147 9.43
[40, 150) 14 259 0.97 0.122 9.12

The hit rate collapses only when the target is already effectively dead, and it collapses for high-power shots too (73 shots at p ≥ 0.5 against a < 2-energy DrussGT hit 0.00 %). Every other enemy-energy stratum sits at 8.5–9.9 %. So the "low power hits 1 %" reading is a dead-target artifact, not a lead-capture failure. Capture itself is flat across enemy energy (0.111 → 0.147 → 0.122).


6. Do we move the aim in proportion to the required lead? (response curve)

At 450+, split the shots into deciles of requiredLead and look at the mean applied lead in each:

requiredLead decile (mean) −21.6 −13.7 −7.8 −2.5 +2.5 +8.1 +14.3 +22.3
mean appliedLead −4.09 −1.51 −1.00 −0.55 +0.36 −0.16 +0.14 +3.26
hit % 9.27 7.92 10.29 10.08 9.95 8.75 7.22 9.69

The required lead swings across ±22°, and our mean applied lead responds by about ±4° — and the hit rate does not depend on the size of the required lead at all. This is the cleanest single number in the report: the bullet is aimed, to first order, at the line of sight, not at the interception point. Consistently, mean|appliedLead| ≈ mean|requiredLead| (11.4° vs 11.6° at 450+) — we have as much dispersion of aim as the target has motion, but almost no correlation with it (capSlp = 0.135).


7. The ceiling: what a trivial predictor would do

requiredLead is an oracle (it uses DrussGT's actual future dodge). To separate "our gun is bad" from "the dodge is unpredictable", the analyzer also computes a naive linear-predictor control: same bullet speed, aimed at the intercept of a straight-line continuation of DrussGT's last 4 ticks of velocity. Its capture (capLin) and the ratio are in §3:

range our capSlp naive-linear capLin us / linear
0–100 0.401 0.597 0.67
100–200 0.391 0.600 0.65
200–300 0.246 0.428 0.58
300–450 0.198 0.404 0.49
450+ 0.135 0.291 0.46

So a trivial predictor still only reaches 0.29 at long range — most of the oracle lead is genuinely unattainable against a strong dodger. But we reach only 46 % of that trivial benchmark. The under-lead is therefore real and worth fixing, while the gap from the trivial benchmark to 1.0 is the dodger's unpredictability and is not.

Corroboration: the powtest corpus — a different ModularBot binary — reproduces the shape almost exactly (capSlp 0.529 / 0.354 / 0.283 / 0.219 / 0.136 by the same range bands). It is a property of the architecture, not of one build.


8. Caveats and sample sizes

  • The oracle caveat. requiredLead uses perfect future information. capture = 1.0 is unattainable against a bot that dodges; it is a diagnostic, not a target. Compare our capture to the capLin control, not to 1.0.
  • Weak correlation, large dispersion. At long range our applied lead is essentially an aim with mean 0 and ~11.5° dispersion that is only weakly proportional to the required lead (proportional slope 0.135). "We apply 13 % of the lead" is a proportional-fit statement; per shot, what we actually do is aim near the LOS with a wide error.
  • Stale live world state. The recorded dir is what the live ModularBot fired using its own stale between-scan world state; the capture's positions are perfect truth. The measured leadError therefore folds in the bot's scan staleness. That is the correct thing to measure (it is the error the bullet actually carries), but it is not the same as the gun's internal prediction error.
  • Stratification. Power is not randomised: it is assigned by the range cap, the energy slope and the finishing rule. Every power comparison above is made within a range band and (in §4) with the endgame removed. The raw cross-range power distribution is reported in the captured output.
  • Sample sizes. The short bands are thin (0–100: n = 27; 100–200: n = 274) — treat their capture as indicative. Everything from 200 px out is large (990 → 36 457). The 450+ sub-0.5 power cell with the enemy alive is n = 4 and should be ignored (it is printed for completeness; the real sub-0.5 shots are the endgame of §5).
  • Rotation direction. requiredLead/appliedLead are signed about the LOS; capSlp additionally assumes a proportional (through-origin) relation, which is the right first-order model for a lead gun but not exact for large angles.

9. Verdict

MEASURED

  • Capture falls monotonically with range: capSlp 0.401 → 0.391 → 0.246 → 0.198 → 0.135. At 450+ we apply ~13 % of the required lead and miss by ~137 px against an 18 px bot; |leadError|/tolerance grows 1.27 → 7.54.
  • Capture is flat across fired power within a range band (450+, enemy alive: 0.154 / 0.127 / 0.127 for p = 0.53 / 0.87 / 1.00).
  • The applied lead barely responds to the required lead: across required-lead deciles spanning ±22°, the mean applied lead moves by ~±4° and the hit rate is flat.
  • A naive linear predictor captures 0.29–0.60; we capture 46–67 % of it.
  • Hits validate the geometry (11.6 px mean, 80.8 % inside 18 px) and separate from misses by 11.6×; 496/496 deaths and 70/70 owner maps resolve correctly.
  • The sub-0.5-power long-range shots (1.18 % hit) are endgame shots at a near-dead DrussGT (mean enemy energy 1.1), and high-power shots there miss just as much.

INFERRED

  • The gun is, to first order, a line-of-sight / weak-lead gun against DrussGT, not an intercept-point gun: the bullet direction tracks the enemy's current line rather than where the enemy will be, which is why capture is low and why it degrades with range while the angular error stays ~constant.
  • The user's mechanism — "we are not capable of reaching the correct angle at range" — is confirmed for range, but not through power: low power does not reduce our displacement capacity here (it is a faster bullet and we capture the same fraction). The low-power/long-range association is a policy artifact (the range cap), so fixing the lead at range fixes the low-power miss rate too.
  • Where to look next: the gap between our capture and the trivial-linear control (0.46×) is the actionable, fixable part. The gap from the trivial control to 1.0 is the dodger's unpredictability and should not be chased.

10. Reproducing

# corpora (live captures; not in the repo, ~500 MB total)
#   /tmp/tfil_ab2/out/<A..E>/run*.jsonl{,.events.jsonl,.rounds.json}
#   /tmp/powtest/{cap_<arm>_r<n>.jsonl,events_<arm>_r<n>.json,cap_*_r<n>.jsonl.rounds.json}

python3 common_libs/tests/analyze_lead_capture_by_range.py \
    --tfil /tmp/tfil_ab2/out --powtest /tmp/powtest \
    --json common_libs/tests/fixtures/lead_capture_by_range_results.json \
    | tee common_libs/tests/fixtures/lead_capture_by_range_output.txt

common_libs/tests/fixtures/lead_capture_by_range_output.txt is the verbatim captured output every table above is taken from; ..._results.json is the same numbers as JSON. Runtime ≈ 15 s for both corpora, single-threaded pure Python (no numpy).

To rebuild a corpus, tools/robocode_shim/make_botdir.sh /tmp/tr_bots/DrussGT and then TR_EVENTS_OUT=<run>.events.jsonl tools/robocode_shim/run_bridge_battle.sh <adversaryBotDir> 7 <run>.jsonl.