Files
SirRoboGarage/docs/drussgt_dodge_vs_power.md
SirStone 1adefaba26 DrussGT dodge vs fired power: no movement response once range is controlled
Answers the user's hypothesis that DrussGT dodges low-power shots better.
Measured on 70 live battles / 490 rounds / 54939 real shots vs real DrussGT
(/tmp/tfil_ab2) and replicated on 35 more battles / 24280 shots (/tmp/powtest).

Power is not randomly assigned - our policy caps it by RANGE
(TR_POWER_FAR_DIST=200 -> 1.0) and by OUR OWN ENERGY (the slope), so inside a
range band power is almost a deterministic function of our energy and a naive
low-vs-high comparison is secretly a losing-vs-healthy comparison. Everything
is stratified by range band and backed by a within-band shuffled-label null
(arrival re-derived, so the null keeps the kinematic channel), a round-cluster
bootstrap, and a within-shot CONTROL window 40 ticks later when the bullet is
long gone.

RESULT: no behavioural response. In band 450+ the raw miss distance at arrival
is +8.25 px [+5.39,+11.25] for HIGH power - but per flight tick it is 4.52 vs
4.51 px/tick (delta -0.01 [-0.12,+0.10]), i.e. entirely the 2.13-tick longer
flight window of the slower bullet. Fixed-12-tick lateral displacement is flat
(55.63 vs 55.47, -0.15 [-1.15,+0.84]) and turn rate / speed are flat. The whole
difference is already present 5 ticks after the trigger pull (+4.2 px) and is
just as large in the bullet-free control window (+5.6 px), so it is a property
of the low-energy situation, not of the shot. Hit rate is flat (0.10 vs 0.09).

Corpus/attribution notes: e*=DrussGT (subject), s*=ModularBot, per
TrBattleCapture.java; the Tank-Royale owner id is NOT stable across runs and is
recovered per battle from fire geometry + the energy decrement, cross-checked on
496/496 death events. Geometry validated on the server's own hits (mean miss
11.6 px, 80.6% inside the 18 px radius).
2026-09-24 21:22:38 +02:00

18 KiB
Raw Permalink Blame History

DrussGT's movement response as a function of the bullet power we fire

Question (the user's hypothesis). "When low power, more bullets fly and I see DrussGT dodging easier. Is this a DrussGT hidden feature at low energy to suddenly increase the capability to dodge? I doubt."

Answer: no. Once range is controlled, DrussGT's movement does not respond to the power we fire. The perceived "better dodging at low power" is a property of the situation the low-power shots are fired in (we are nearly dead, so the whole engagement geometry differs), not of the bullet — and most of the raw difference in miss distance is pure flight-time kinematics: a low-power bullet is a faster bullet, so it arrives sooner and DrussGT simply has fewer ticks to drift off the line. Every measure that removes the flight window is flat.

This was measured on 70 real live battles / 490 rounds / 54 939 shots against the real, unmodified DrussGT (/tmp/tfil_ab2/) and replicated on 35 more battles / 245 rounds / 24 280 shots (/tmp/powtest/). Nothing here uses the offline fixture harness; these are real robot-vs-robot Tank Royale battles recorded through tools/robocode_shim/run_bridge_battle.sh.


1. What was measured, and how a shot is attributed

Each shot is analysed from the recorded per-tick worldstate plus the real fire/hit/wall event sidecar. For a shot fired by ModularBot (us) at DrussGT with power p in direction dir:

  • bullet speed v = 20 - 3p px/tick (so low power = faster bullet = shorter flight);
  • the bullet flies from the fire position along u = (cos dir, sin dir);
  • along(t) = (D(t) - P0)·u and perp(t) = (D(t) - P0)×u are DrussGT's along-track distance and perpendicular offset from our aim line at tick t;
  • arrival tick k* = the first tick where the bullet has travelled at least along(k*);
  • (a) miss distance at arrival = |perp(k*)| (bot radius is 18 px).

Attribution (this has bitten the project before, so it is stated explicitly).

  • In the capture rows, e* is the SUBJECT = DrussGT and s* is the adversary = ModularBot. That is by construction of tools/robocode_shim/src/robocode_shim/TrBattleCapture.java (en = the bot whose name contains "DrussGT", written as e*; sh = the other, written as s*).
  • The event sidecar's owner is the Tank Royale bot id, and it is not stable across runs (start order varies). It is therefore recovered per battle from the fire geometry (owner's position equals the fire event's x,y, and its energy drops by exactly power on the next capture row).
  • Cross-check: for 496/496 death events the mapped victim is the bot whose energy is ~0 at the end of that round. The mapping agrees with the death evidence in every battle.
  • Sanity check that the mapping is the right way round: the recovered ModularBot power histogram is exactly its documented policy (common_libs/gun_harness/virtual_bullets.nim): a spike at 0.50 (TR_POWER_ENERGY_MIN, when our energy ≤ 20), a spike at 1.00 (TR_POWER_FAR_CAP, beyond TR_POWER_FAR_DIST = 200), and the linear TR_POWER_ENERGY_SLOPE continuum in between.
  • No power-value heuristic is used anywhere (the old {1.0,1.5,2.0,3.0} assumption is exactly what mis-attributed ModularBot in a previous job; our shots here are mostly 0.50 and 1.00, plus a continuum 0.10–0.99).

Validation of the geometry (MEASURED). For shots the server recorded as hits, the measured miss distance is mean 11.6 px, median 10.8 px, 80.6 % below the 18 px bot radius; for shots that hit a wall it is 133 px. The measurement is therefore correct and the ~18 px figure is the right reference scale.


2. The confound is real and severe (MEASURED)

Power is not randomly assigned. The policy only ever caps the gun's preference, by range (TR_POWER_FAR_DIST=200 → 1.0), by our own energy (a linear slope from 0.5 at ≤20 to the cap at ≥80), by the finishing rule, and by the sub-average-chances rule. The consequence is stark — this is the same table for both corpora (tfil_ab2):

band power shots fire px OUR energy enemy energy flight ticks
300-450 0.50–0.75 5 272 411.3 13.8 19.2 22.36
300-450 0.75–1.00 880 407.7 29.0 28.5 23.49
300-450 1.00–1.50 10 696 402.0 69.2 62.6 23.70
450+ 0.50–0.75 12 960 536.1 13.7 18.6 28.38
450+ 0.75–1.00 2 424 540.4 28.8 26.4 30.18
450+ 1.00–1.50 20 126 534.2 63.0 53.8 30.52

Two things follow, and they drive the whole design:

  1. Range must be held fixed — hence everything below is stratified by range band, and the headline is the stratified result.
  2. Within a range band, our power is almost a deterministic function of our own energy. Low-power shots are shots fired when we are nearly dead. So a naive "low power vs high power" comparison inside a band is secretly a "losing badly vs healthy" comparison. That is why a within-shot control is needed to separate the two, and it is provided in §4.

Because the LOW/HIGH split is also a low-energy/high-energy split, the whole comparison must be read as an energy-conditioned contrast, and the burden of proof falls on the time-resolved tests, not on the raw miss distance.


3. Headline: within-range-band, LOW vs HIGH power

LOW = p ∈ [0.50, 0.75), HIGH = p ∈ [1.00, 1.50); 95 % CI from a round-cluster bootstrap (resampling battles, not shots — shots inside a battle are correlated). Corpus tfil_ab2, 70 battles.

metric band nLOW nHIGH LOW HIGH delta 95 % CI
(a) miss at arrival px 300-450 5 272 10 696 104.67 110.55 +5.88 [+2.64, +9.07]
(a) miss at arrival px 450+ 12 960 20 126 124.01 132.26 +8.25 [+5.39, +11.25]
(b) lateral disp. / flight tick 450+ 12 960 20 126 3.47 3.42 −0.04 [−0.12, +0.03]
(b′) lateral disp., FIXED 12 ticks 450+ 12 960 20 126 55.63 55.47 −0.15 [−1.15, +0.84]
(b′′) CONTROL, same shot, +40 ticks 450+ 12 828 20 120 52.03 52.64 +0.61 [−0.38, +1.58]
(f) flight window ticks 450+ 12 960 20 126 28.38 30.52 +2.13 [+1.81, +2.44]
(c) response latency ticks 450+ 8 682 13 884 13.60 14.27 +0.68 [+0.36, +0.99]
(d) turn rate deg/tick (flight) 450+ 12 960 20 126 1.49 1.46 −0.03 [−0.09, +0.02]
(d′) turn rate deg/tick (fixed 12) 450+ 12 960 20 126 1.43 1.28 −0.15 [−0.20, −0.10]
(d′′) turn rate, CONTROL (+40) 450+ 12 902 20 126 1.52 1.58 +0.06 [+0.01, +0.12]
(d) speed px/tick (flight) 450+ 12 960 20 126 5.46 5.48 +0.01 [−0.06, +0.08]
OUTCOME hit rate 450+ 12 960 20 126 0.10 0.09 −0.01 [−0.02, −0.00]

powtest, 35 battles (replication, same bins):

metric band nLOW nHIGH LOW HIGH delta 95 % CI
(a) miss at arrival px 300-450 1 388 8 151 106.22 109.93 +3.72 [−1.50, +8.65]
(a) miss at arrival px 450+ 2 249 11 085 122.00 131.86 +9.86 [+4.56, +15.25]
(b) lateral disp. / flight tick 450+ 2 249 11 085 3.53 3.48 −0.05 [−0.20, +0.11]
(b′) lateral disp., FIXED 12 ticks 450+ 2 249 11 085 55.53 55.85 +0.31 [−1.24, +1.87]
(f) flight window ticks 450+ 2 249 11 085 27.75 29.93 +2.18 [+1.63, +2.73]
OUTCOME hit rate 450+ 2 249 11 085 0.10 0.09 −0.01 [−0.02, +0.01]

Reading. The only metric with a non-trivial difference is the raw miss distance, and it is larger for high power — i.e. the opposite of the hypothesis ("low power is dodged better"). Two things explain it entirely, and neither is a behavioural response.


4. Why the miss-distance difference is not a dodge: two decisive tests

4.1 The difference is the flight window, not the dodging

A low-power bullet is faster, so it arrives 2.13 ticks sooner in band 450+ (28.38 vs 30.52). DrussGT drifts away from the aim line at roughly 1.7–1.8 px per tick; 2.13 ticks × ~3.9 px/tick of accumulated miss ≈ +8 px — exactly the observed +8.25 px. Normalise by the window and it disappears:

metric band LOW HIGH delta 95 % CI
miss / flight ticks 300-450 4.85 4.88 +0.03 [−0.13, +0.17]
miss / flight ticks 450+ 4.52 4.51 −0.01 [−0.12, +0.10]

(per-tick miss rate, px/tick). The per-tick dodge rate is identical. Same in powtest: 4.56 vs 4.59, delta +0.03 [−0.19, +0.25].

4.2 The difference is already present before the bullet can matter, and survives the bullet

The arrival tick depends on the power, so it is the wrong place to compare. Instead, measure |perp| — DrussGT's perpendicular offset from our aim line — at a fixed number of ticks after the fire, and again in a CONTROL window 40 ticks later, when the bullet is long gone. A response to our shot must be ~0 at k=5 and grow with k; a phase/geometry difference is present already at k=5 and identical in the control.

tfil_ab2, band 450+ (nLOW = 12 960, nHIGH = 20 126):

k (ticks after fire) LOW HIGH delta CONTROL (+40 ticks) LOW HIGH delta
5 94.9 99.1 +4.21 144.2 149.7 +5.57
10 95.0 98.9 +3.86 147.8 153.1 +5.23
15 100.7 104.3 +3.53 — — —
20 109.0 112.5 +3.50 — — —
25 118.7 122.6 +3.83 — — —
30 127.7 132.5 +4.74 — — —

powtest, band 450+ (nLOW = 2 249, nHIGH = 11 085): k=5 delta +5.97, control +6.59; identical story.

The whole difference is present 5 ticks after the trigger pull (before the miss can be a reaction to a bullet still in flight) and is at least as large in the window where the bullet has already gone. It is a standing offset between the two situations, not a dodge.

Also note the sign flip: in band 300-450 the turn rate is +0.03 deg/tick in the flight window and +0.13 deg/tick in the bullet-free control window — the control window shows more difference than the flight window. In band 450+ the flight-window turn rate leans one way (−0.15) and the control window the other (+0.06). A real response would behave the opposite way in both.

4.3 What the miss-distance metric actually contains (MEASURED)

|perp| is already 80–95 px only 5 ticks after the fire. The miss distance is dominated by our own gun's lead/aim error, not by DrussGT's dodge. It is therefore a weak instrument for "dodge quality" on its own, which is why the flight-normalised and fixed-window measures carry the conclusion.


5. Shuffled-label null (MEASURED)

Labels are permuted inside each range band; for the arrival-dependent metrics the arrival tick is re-derived for the permuted power, so the null keeps the kinematic channel and destroys only the response. 400 reps. Corpus tfil_ab2, per band, two-sided permutation p on delta = mean(HIGH) − mean(LOW):

metric 100-200 200-300 300-450 450+
miss d=+5.35, p=0.89 d=+0.35, p=0.39 d=+5.91, p=0.27 d=+8.24, p=0.005
miss / flight d=−0.35, p=0.62 d=−0.34, p=0.52 d=+0.03, p=0.01 d=−0.01, p=0.005
lateral / flight — — d=−0.08, p≈0.1 d=−0.04, p≈0.1
lateral, fixed 12 d=+16.2, p=0.02 d=+2.89, p=0.27 d=−1.45, p=0.005 d=−0.15, p=0.66
lateral, CONTROL (+40) d=+3.67, p=0.68 d=+5.53, p=0.10 d=−0.85, p=0.10 d=+0.61, p=0.08
turn rate, fixed 12 d=+0.57, p=0.15 d=+0.32, p=0.04 d=+0.03, p=0.13 d=−0.15, p=0.005
turn rate, CONTROL d=+0.84, p=0.06 d=+0.41, p=0.005 d=+0.14, p=0.005 d=+0.06, p=0.005
speed, fixed 12 d=+0.09, p=0.91 d=+0.02, p=0.85 d=−0.05, p=0.14 d=−0.01, p=0.76
hit rate d=+0.09, p=0.57 d=−0.02, p=0.67 d=+0.004, p=0.43 d=−0.013, p=0.005

The null is centred on the kinematic counterfactual, so a small p means "the observed difference is not fully explained by flight-time kinematics". Only the raw miss in 450+ (and the fixed-12 turn rate) reach that, and in both cases the sign is the opposite of "low power is dodged better": at matched conditions DrussGT is marginally further from the line and turns less on the high-power (slower) bullets, and the difference is present in the bullet-free control window as well.

Effective sample size. The two bands that carry the conclusion are band 450+ (12 960 LOW / 20 126 HIGH, 70 battles) and band 300-450 (5 272 / 10 696, 70 battles), replicated in a second corpus (2 249 / 11 085, 35 battles). The bootstrap is clustered by battle, so the CIs already carry the between-battle variance. Bands 0-100 and 100-200 have n = 6/2 and 53/23 shots — far too few for any claim, and they are labelled "too few for a contrast" in the report. The design has ample power to detect a modest effect (the CIs on the fixed-window lateral displacement, a ~55 px quantity, are ±1 px); it cannot rule out an effect smaller than ~2 % on any of these measures.


6. The response latency, turn rate and speed (MEASURED)

  • (c) Response latency (ticks from our fire until DrussGT's heading has turned

    15°): band 450+ 13.60 → 14.27 ticks, delta +0.68 [+0.36, +0.99]. But the measurement window is capped by the flight, which is itself 2.13 ticks longer for HIGH power. As a fraction of the flight window it is 0.479 (LOW) vs 0.468 (HIGH) — the high-power shots are responded to sooner relative to the bullet. In band 300-450 it is −0.09 [−0.47, +0.30]. Directionally inconsistent ⇒ no effect. powtest: +0.47 (450+), −0.29 (300-450).

  • (d) Turn rate in the flight window: 1.49 vs 1.46 deg/tick in 450+ (−0.03 [−0.09, +0.02]); +0.13 [+0.05, +0.21] in 300-450 with the same sign in the control window (+0.13). Surfers do slow to turn, but DrussGT's turn rate does not track our power.
  • Speed in the flight window: 5.46 vs 5.48 px/tick — flat, and the control window agrees. DrussGT is not slowing down more against low-power bullets.

7. Flight-window lengths (MEASURED) — the thing that would have to be true

speed = 20 − 3p, so a lower power is a faster bullet and a shorter flight. Measured, band 450+:

power 0.50–0.75 0.75–1.00 1.00–1.50
flight ticks 28.38 30.18 30.52

A genuine "dodge low power better" would have to overcome a 2.1-tick handicap and still produce a larger miss per unit time at low power. It does not: the per-tick miss rate is 4.52 at LOW and 4.51 at HIGH. If anything the extra reaction time on the slow, high-power bullets produces slightly more lateral drift (1.8 px/tick vs 1.7 px/tick between k=10 and k=30 in band 450+).


8. Outcome (MEASURED) — reproduces the known flat result

Hit rate from the real server events, per band:

band 0.50–0.75 1.00–1.50 delta 95 % CI
300-450 0.11 0.12 +0.004 [−0.01, +0.02]
450+ 0.10 0.09 −0.013 [−0.02, −0.00]

Flat, in agreement with the previously measured live outcome (9.55 % / 10.86 % / 10.35 % by fired power, Fisher p = 0.37). This job adds the mechanism: the outcome is flat because the movement is flat.


9. Verdict

MEASURED

  • DrussGT's miss-normalised, fixed-window and control-window movement metrics are flat across the power we fire, inside a range band, on 54 939 shots and replicated on 24 280 more.
  • The only between-power difference that is large is the raw miss distance at arrival (+8.2 px in band 450+ for a 2.13-tick longer flight), and that is exactly the flight-window kinematics; normalising by the window removes it.
  • The difference is present 5 ticks after the fire and is as large in the window 40 ticks later when the bullet is gone — so it is a property of the low-energy / high-energy situation, not of the shot.
  • Hit rate is flat.

INFERRED

  • The null results rule out a dodge-quality response to power of the size the visual impression suggests. They cannot exclude an effect that is exactly cancelled by the same energy confound in every metric — but that would require a coincidence, and the control window argues against it.
  • The user's visual impression ("more bullets fly at low power") is real but is a bullet-count effect, not a dodge effect: firing at low power gives a shorter reload interval (10 + 2p ticks) and a faster bullet, so more bullets are in the air and more of them are seen. That increased visual rate of activity is not DrussGT dodging differently.
  • There is no hidden DrussGT behaviour to exploit at low power. If anything, the (tiny, situational) residual points the other way: DrussGT is marginally further off the aim line and turns less against the slower high-power bullets, which is just the extra reaction time.
  • Practical reading for the energy-slope policy: the low-power arm loses nothing in terms of DrussGT's evasion; its hit rate is the same and its energy cost is lower. The flat outcome measured earlier is not hiding a movement-side penalty.

10. Reproducing

# corpora (live captures; not in the repo - ~490 MB total)
#   /tmp/tfil_ab2/out/<A..E>/run*.jsonl{,.events.jsonl,.rounds.json}
#   /tmp/powtest/{cap_<arm>_r<n>.jsonl,events_<arm>_r<n>.json,cap_*_r<n>.jsonl.rounds.json}

python3 common_libs/tests/analyze_drussgt_dodge_vs_power.py \
    --tfil /tmp/tfil_ab2/out --powtest /tmp/powtest --reps 400 \
    --json common_libs/tests/fixtures/dodge_vs_power_results.json \
    | tee common_libs/tests/fixtures/dodge_vs_power_report.txt

common_libs/tests/fixtures/dodge_vs_power_report.txt is the verbatim captured output the tables above were taken from; common_libs/tests/fixtures/dodge_vs_power_results.json is the same numbers as JSON. Runtime ≈ 2.5 minutes for both corpora, single-threaded pure Python (no numpy).

To rebuild a corpus, tools/robocode_shim/make_botdir.sh /tmp/tr_bots/DrussGT and then TR_EVENTS_OUT=<run>.events.jsonl tools/robocode_shim/run_bridge_battle.sh <adversaryBotDir> 7 <run>.jsonl (the analyzer also needs the .rounds.json sidecar, which the shim writes automatically).