# DrussGT's movement response as a function of the bullet power we fire **Question (the user's hypothesis).** *"When low power, more bullets fly and I see DrussGT dodging easier. Is this a DrussGT hidden feature at low energy to suddenly increase the capability to dodge? I doubt."* **Answer: no.** Once range is controlled, DrussGT's movement does **not** respond to the power we fire. The perceived "better dodging at low power" is a property of the *situation* the low-power shots are fired in (we are nearly dead, so the whole engagement geometry differs), not of the bullet — and most of the raw difference in miss distance is pure flight-time kinematics: a low-power bullet is a *faster* bullet, so it arrives **sooner** and DrussGT simply has fewer ticks to drift off the line. Every measure that removes the flight window is flat. This was measured on **70 real live battles / 490 rounds / 54 939 shots** against the real, unmodified DrussGT (`/tmp/tfil_ab2/`) and **replicated on 35 more battles / 245 rounds / 24 280 shots** (`/tmp/powtest/`). Nothing here uses the offline fixture harness; these are real robot-vs-robot Tank Royale battles recorded through `tools/robocode_shim/run_bridge_battle.sh`. --- ## 1. What was measured, and how a shot is attributed Each shot is analysed from the recorded per-tick worldstate plus the real fire/hit/wall event sidecar. For a shot fired by **ModularBot (us)** at DrussGT with power `p` in direction `dir`: * bullet speed `v = 20 - 3p` px/tick (so **low power = faster bullet = shorter flight**); * the bullet flies from the fire position along `u = (cos dir, sin dir)`; * `along(t) = (D(t) - P0)·u` and `perp(t) = (D(t) - P0)×u` are DrussGT's along-track distance and **perpendicular offset from our aim line** at tick `t`; * **arrival tick** `k*` = the first tick where the bullet has travelled at least `along(k*)`; * **(a) miss distance at arrival** = `|perp(k*)|` (bot radius is 18 px). **Attribution (this has bitten the project before, so it is stated explicitly).** * In the capture rows, **`e*` is the SUBJECT = DrussGT** and **`s*` is the adversary = ModularBot**. That is by construction of `tools/robocode_shim/src/robocode_shim/TrBattleCapture.java` (`en` = the bot whose name contains "DrussGT", written as `e*`; `sh` = the other, written as `s*`). * The event sidecar's `owner` is the Tank Royale bot id, and **it is not stable across runs** (start order varies). It is therefore recovered **per battle** from the fire geometry (`owner`'s position equals the fire event's `x,y`, and its energy drops by exactly `power` on the next capture row). * Cross-check: for **496/496 death events** the mapped victim is the bot whose energy is ~0 at the end of that round. The mapping agrees with the death evidence in every battle. * Sanity check that the mapping is the *right way round*: the recovered ModularBot power histogram is exactly its documented policy (`common_libs/gun_harness/virtual_bullets.nim`): a spike at **0.50** (`TR_POWER_ENERGY_MIN`, when our energy ≤ 20), a spike at **1.00** (`TR_POWER_FAR_CAP`, beyond `TR_POWER_FAR_DIST = 200`), and the linear `TR_POWER_ENERGY_SLOPE` continuum in between. * No power-value heuristic is used anywhere (the old {1.0,1.5,2.0,3.0} assumption is exactly what mis-attributed ModularBot in a previous job; our shots here are mostly **0.50 and 1.00**, plus a continuum 0.10–0.99). **Validation of the geometry (MEASURED).** For shots the *server* recorded as hits, the measured miss distance is **mean 11.6 px, median 10.8 px, 80.6 % below the 18 px bot radius**; for shots that hit a wall it is 133 px. The measurement is therefore correct and the ~18 px figure is the right reference scale. --- ## 2. The confound is real and severe (MEASURED) Power is **not** randomly assigned. The policy only ever *caps* the gun's preference, by range (`TR_POWER_FAR_DIST=200 → 1.0`), by **our own energy** (a linear slope from 0.5 at ≤20 to the cap at ≥80), by the finishing rule, and by the sub-average-chances rule. The consequence is stark — this is the *same* table for both corpora (tfil_ab2): | band | power | shots | fire px | OUR energy | enemy energy | flight ticks | |---|---|---|---|---|---|---| | 300-450 | 0.50–0.75 | 5 272 | 411.3 | **13.8** | 19.2 | 22.36 | | 300-450 | 0.75–1.00 | 880 | 407.7 | 29.0 | 28.5 | 23.49 | | 300-450 | 1.00–1.50 | 10 696 | 402.0 | **69.2** | 62.6 | 23.70 | | 450+ | 0.50–0.75 | 12 960 | 536.1 | **13.7** | 18.6 | 28.38 | | 450+ | 0.75–1.00 | 2 424 | 540.4 | 28.8 | 26.4 | 30.18 | | 450+ | 1.00–1.50 | 20 126 | 534.2 | **63.0** | 53.8 | 30.52 | Two things follow, and they drive the whole design: 1. **Range must be held fixed** — hence everything below is stratified by range band, and the headline is the stratified result. 2. **Within a range band, our power is almost a deterministic function of our own energy.** Low-power shots are shots fired when *we* are nearly dead. So a naive "low power vs high power" comparison inside a band is secretly a "losing badly vs healthy" comparison. That is why a *within-shot* control is needed to separate the two, and it is provided in §4. Because the LOW/HIGH split is also a low-energy/high-energy split, the **whole comparison must be read as an energy-conditioned contrast**, and the burden of proof falls on the *time-resolved* tests, not on the raw miss distance. --- ## 3. Headline: within-range-band, LOW vs HIGH power `LOW = p ∈ [0.50, 0.75)`, `HIGH = p ∈ [1.00, 1.50)`; 95 % CI from a **round-cluster bootstrap** (resampling battles, not shots — shots inside a battle are correlated). Corpus `tfil_ab2`, 70 battles. | metric | band | nLOW | nHIGH | LOW | HIGH | delta | 95 % CI | |---|---|---|---|---|---|---|---| | **(a) miss at arrival px** | 300-450 | 5 272 | 10 696 | 104.67 | 110.55 | **+5.88** | [+2.64, +9.07] | | **(a) miss at arrival px** | 450+ | 12 960 | 20 126 | 124.01 | 132.26 | **+8.25** | [+5.39, +11.25] | | (b) lateral disp. / flight tick | 450+ | 12 960 | 20 126 | 3.47 | 3.42 | −0.04 | [−0.12, +0.03] | | (b′) lateral disp., FIXED 12 ticks | 450+ | 12 960 | 20 126 | 55.63 | 55.47 | −0.15 | [−1.15, +0.84] | | (b′′) CONTROL, same shot, +40 ticks | 450+ | 12 828 | 20 120 | 52.03 | 52.64 | +0.61 | [−0.38, +1.58] | | (f) flight window ticks | 450+ | 12 960 | 20 126 | 28.38 | 30.52 | **+2.13** | [+1.81, +2.44] | | (c) response latency ticks | 450+ | 8 682 | 13 884 | 13.60 | 14.27 | +0.68 | [+0.36, +0.99] | | (d) turn rate deg/tick (flight) | 450+ | 12 960 | 20 126 | 1.49 | 1.46 | −0.03 | [−0.09, +0.02] | | (d′) turn rate deg/tick (fixed 12) | 450+ | 12 960 | 20 126 | 1.43 | 1.28 | −0.15 | [−0.20, −0.10] | | (d′′) turn rate, CONTROL (+40) | 450+ | 12 902 | 20 126 | 1.52 | 1.58 | +0.06 | [+0.01, +0.12] | | (d) speed px/tick (flight) | 450+ | 12 960 | 20 126 | 5.46 | 5.48 | +0.01 | [−0.06, +0.08] | | OUTCOME hit rate | 450+ | 12 960 | 20 126 | 0.10 | 0.09 | −0.01 | [−0.02, −0.00] | `powtest`, 35 battles (replication, same bins): | metric | band | nLOW | nHIGH | LOW | HIGH | delta | 95 % CI | |---|---|---|---|---|---|---|---| | (a) miss at arrival px | 300-450 | 1 388 | 8 151 | 106.22 | 109.93 | +3.72 | [−1.50, +8.65] | | (a) miss at arrival px | 450+ | 2 249 | 11 085 | 122.00 | 131.86 | **+9.86** | [+4.56, +15.25] | | (b) lateral disp. / flight tick | 450+ | 2 249 | 11 085 | 3.53 | 3.48 | −0.05 | [−0.20, +0.11] | | (b′) lateral disp., FIXED 12 ticks | 450+ | 2 249 | 11 085 | 55.53 | 55.85 | +0.31 | [−1.24, +1.87] | | (f) flight window ticks | 450+ | 2 249 | 11 085 | 27.75 | 29.93 | +2.18 | [+1.63, +2.73] | | OUTCOME hit rate | 450+ | 2 249 | 11 085 | 0.10 | 0.09 | −0.01 | [−0.02, +0.01] | **Reading.** The only metric with a non-trivial difference is the raw miss distance, and it is *larger* for **high** power — i.e. the opposite of the hypothesis ("low power is dodged better"). Two things explain it entirely, and neither is a behavioural response. --- ## 4. Why the miss-distance difference is not a dodge: two decisive tests ### 4.1 The difference is the flight window, not the dodging A low-power bullet is *faster*, so it arrives **2.13 ticks sooner** in band 450+ (28.38 vs 30.52). DrussGT drifts away from the aim line at roughly **1.7–1.8 px per tick**; 2.13 ticks × ~3.9 px/tick of accumulated miss ≈ **+8 px** — exactly the observed +8.25 px. Normalise by the window and it disappears: | metric | band | LOW | HIGH | delta | 95 % CI | |---|---|---|---|---|---| | miss / flight ticks | 300-450 | 4.85 | 4.88 | +0.03 | [−0.13, +0.17] | | miss / flight ticks | 450+ | 4.52 | 4.51 | −0.01 | [−0.12, +0.10] | (per-tick miss rate, px/tick). The per-tick dodge rate is **identical**. Same in `powtest`: 4.56 vs 4.59, delta +0.03 [−0.19, +0.25]. ### 4.2 The difference is already present before the bullet can matter, and survives the bullet The arrival tick depends on the power, so it is the wrong place to compare. Instead, measure `|perp|` — DrussGT's perpendicular offset from our aim line — at a **fixed number of ticks after the fire**, and again in a **CONTROL window 40 ticks later, when the bullet is long gone**. A response to *our shot* must be ~0 at k=5 and grow with k; a phase/geometry difference is present already at k=5 and identical in the control. `tfil_ab2`, band 450+ (nLOW = 12 960, nHIGH = 20 126): | k (ticks after fire) | LOW | HIGH | delta | CONTROL (+40 ticks) LOW | HIGH | delta | |---|---|---|---|---|---|---| | 5 | 94.9 | 99.1 | **+4.21** | 144.2 | 149.7 | **+5.57** | | 10 | 95.0 | 98.9 | +3.86 | 147.8 | 153.1 | +5.23 | | 15 | 100.7 | 104.3 | +3.53 | — | — | — | | 20 | 109.0 | 112.5 | +3.50 | — | — | — | | 25 | 118.7 | 122.6 | +3.83 | — | — | — | | 30 | 127.7 | 132.5 | +4.74 | — | — | — | `powtest`, band 450+ (nLOW = 2 249, nHIGH = 11 085): k=5 delta **+5.97**, control **+6.59**; identical story. The whole difference is present **5 ticks after the trigger pull** (before the miss can be a reaction to a bullet still in flight) and is *at least as large in the window where the bullet has already gone*. It is a standing offset between the two situations, not a dodge. Also note the **sign flip**: in band 300-450 the turn rate is +0.03 deg/tick in the flight window and +0.13 deg/tick in the bullet-free control window — the control window shows *more* difference than the flight window. In band 450+ the flight-window turn rate leans one way (−0.15) and the control window the other (+0.06). A real response would behave the opposite way in both. ### 4.3 What the miss-distance metric actually contains (MEASURED) `|perp|` is already 80–95 px only 5 ticks after the fire. **The miss distance is dominated by our own gun's lead/aim error, not by DrussGT's dodge.** It is therefore a weak instrument for "dodge quality" on its own, which is why the flight-normalised and fixed-window measures carry the conclusion. --- ## 5. Shuffled-label null (MEASURED) Labels are permuted **inside each range band**; for the arrival-dependent metrics the arrival tick is **re-derived** for the permuted power, so the null keeps the kinematic channel and destroys only the response. 400 reps. Corpus `tfil_ab2`, per band, two-sided permutation p on `delta = mean(HIGH) − mean(LOW)`: | metric | 100-200 | 200-300 | 300-450 | 450+ | |---|---|---|---|---| | miss | d=+5.35, p=0.89 | d=+0.35, p=0.39 | d=+5.91, p=0.27 | d=+8.24, p=0.005 | | miss / flight | d=−0.35, p=0.62 | d=−0.34, p=0.52 | d=+0.03, p=0.01 | d=−0.01, p=0.005 | | lateral / flight | — | — | d=−0.08, p≈0.1 | d=−0.04, p≈0.1 | | lateral, fixed 12 | d=+16.2, p=0.02 | d=+2.89, p=0.27 | d=−1.45, p=0.005 | d=−0.15, p=0.66 | | lateral, CONTROL (+40) | d=+3.67, p=0.68 | d=+5.53, p=0.10 | d=−0.85, p=0.10 | d=+0.61, p=0.08 | | turn rate, fixed 12 | d=+0.57, p=0.15 | d=+0.32, p=0.04 | d=+0.03, p=0.13 | d=−0.15, p=0.005 | | turn rate, CONTROL | d=+0.84, p=0.06 | d=+0.41, p=0.005 | d=+0.14, p=0.005 | d=+0.06, p=0.005 | | speed, fixed 12 | d=+0.09, p=0.91 | d=+0.02, p=0.85 | d=−0.05, p=0.14 | d=−0.01, p=0.76 | | hit rate | d=+0.09, p=0.57 | d=−0.02, p=0.67 | d=+0.004, p=0.43 | d=−0.013, p=0.005 | The null is **centred on the kinematic counterfactual**, so a *small* p means "the observed difference is not fully explained by flight-time kinematics". Only the raw miss in 450+ (and the fixed-12 turn rate) reach that, and in **both** cases the sign is the *opposite* of "low power is dodged better": at matched conditions DrussGT is marginally **further** from the line and turns **less** on the high-power (slower) bullets, and the difference is present in the bullet-free control window as well. *Effective sample size.* The two bands that carry the conclusion are band 450+ (12 960 LOW / 20 126 HIGH, 70 battles) and band 300-450 (5 272 / 10 696, 70 battles), replicated in a second corpus (2 249 / 11 085, 35 battles). The bootstrap is clustered by battle, so the CIs already carry the between-battle variance. Bands **0-100** and **100-200** have n = 6/2 and 53/23 shots — far too few for any claim, and they are labelled "too few for a contrast" in the report. The design has ample power to detect a modest effect (the CIs on the fixed-window lateral displacement, a ~55 px quantity, are ±1 px); it cannot rule out an effect smaller than ~2 % on any of these measures. --- ## 6. The response latency, turn rate and speed (MEASURED) * **(c) Response latency** (ticks from our fire until DrussGT's heading has turned >15°): band 450+ **13.60 → 14.27 ticks**, delta +0.68 [+0.36, +0.99]. But the measurement window is capped by the flight, which is itself 2.13 ticks longer for HIGH power. As a **fraction of the flight window** it is 0.479 (LOW) vs 0.468 (HIGH) — the high-power shots are responded to *sooner relative to the bullet*. In band 300-450 it is **−0.09** [−0.47, +0.30]. Directionally inconsistent ⇒ no effect. `powtest`: +0.47 (450+), −0.29 (300-450). * **(d) Turn rate** in the flight window: 1.49 vs 1.46 deg/tick in 450+ (−0.03 [−0.09, +0.02]); +0.13 [+0.05, +0.21] in 300-450 with the same sign in the control window (+0.13). Surfers do slow to turn, but DrussGT's turn rate does not track our power. * **Speed** in the flight window: 5.46 vs 5.48 px/tick — flat, and the control window agrees. DrussGT is not slowing down more against low-power bullets. --- ## 7. Flight-window lengths (MEASURED) — the thing that would have to be true `speed = 20 − 3p`, so a **lower** power is a **faster** bullet and a **shorter** flight. Measured, band 450+: | power | 0.50–0.75 | 0.75–1.00 | 1.00–1.50 | |---|---|---|---| | flight ticks | 28.38 | 30.18 | 30.52 | A genuine "dodge low power better" would have to overcome a **2.1-tick handicap** and still produce a *larger* miss per unit time at low power. It does not: the per-tick miss rate is 4.52 at LOW and 4.51 at HIGH. If anything the extra reaction time on the slow, high-power bullets produces slightly *more* lateral drift (1.8 px/tick vs 1.7 px/tick between k=10 and k=30 in band 450+). --- ## 8. Outcome (MEASURED) — reproduces the known flat result Hit rate from the real server events, per band: | band | 0.50–0.75 | 1.00–1.50 | delta | 95 % CI | |---|---|---|---|---| | 300-450 | 0.11 | 0.12 | +0.004 | [−0.01, +0.02] | | 450+ | 0.10 | 0.09 | −0.013 | [−0.02, −0.00] | Flat, in agreement with the previously measured live outcome (9.55 % / 10.86 % / 10.35 % by fired power, Fisher p = 0.37). This job adds the *mechanism*: the outcome is flat because the movement is flat. --- ## 9. Verdict **MEASURED** * DrussGT's miss-normalised, fixed-window and control-window movement metrics are **flat across the power we fire, inside a range band**, on 54 939 shots and replicated on 24 280 more. * The only between-power difference that is large is the **raw miss distance at arrival** (+8.2 px in band 450+ for a 2.13-tick longer flight), and that is **exactly the flight-window kinematics**; normalising by the window removes it. * The difference is present **5 ticks after the fire** and is **as large in the window 40 ticks later when the bullet is gone** — so it is a property of the low-energy / high-energy *situation*, not of the shot. * Hit rate is flat. **INFERRED** * The null results rule out a *dodge-quality* response to power of the size the visual impression suggests. They cannot exclude an effect that is exactly cancelled by the same energy confound in every metric — but that would require a coincidence, and the control window argues against it. * The user's visual impression ("more bullets fly at low power") is real but is a *bullet-count* effect, not a dodge effect: firing at low power gives a shorter reload interval (`10 + 2p` ticks) and a faster bullet, so more bullets are in the air and more of them are seen. That increased visual *rate of activity* is not DrussGT dodging differently. * There is **no hidden DrussGT behaviour** to exploit at low power. If anything, the (tiny, situational) residual points the other way: DrussGT is marginally *further* off the aim line and turns *less* against the slower high-power bullets, which is just the extra reaction time. * Practical reading for the energy-slope policy: the low-power arm loses nothing in terms of DrussGT's evasion; its hit rate is the same and its energy cost is lower. The flat outcome measured earlier is not hiding a movement-side penalty. --- ## 10. Reproducing ```bash # corpora (live captures; not in the repo - ~490 MB total) # /tmp/tfil_ab2/out//run*.jsonl{,.events.jsonl,.rounds.json} # /tmp/powtest/{cap__r.jsonl,events__r.json,cap_*_r.jsonl.rounds.json} python3 common_libs/tests/analyze_drussgt_dodge_vs_power.py \ --tfil /tmp/tfil_ab2/out --powtest /tmp/powtest --reps 400 \ --json common_libs/tests/fixtures/dodge_vs_power_results.json \ | tee common_libs/tests/fixtures/dodge_vs_power_report.txt ``` `common_libs/tests/fixtures/dodge_vs_power_report.txt` is the verbatim captured output the tables above were taken from; `common_libs/tests/fixtures/dodge_vs_power_results.json` is the same numbers as JSON. Runtime ≈ 2.5 minutes for both corpora, single-threaded pure Python (no numpy). To rebuild a corpus, `tools/robocode_shim/make_botdir.sh /tmp/tr_bots/DrussGT` and then `TR_EVENTS_OUT=.events.jsonl tools/robocode_shim/run_bridge_battle.sh 7 .jsonl` (the analyzer also needs the `.rounds.json` sidecar, which the shim writes automatically).