12 KiB
Wiring the wave surfer as TR_MOVEMENT=surf, and a live A/B vs TFIL and STRAFE
Session: commit 0f5cfe37, frozen binary 788d6a34…, 3 arms × 15 runs × 7
rounds = 45 battles vs real DrussGT (Tank Royale physics via the robocode
shim), tools/ab/ab_run.sh. All numbers below are MEASURED from those
captures unless a line says INFERRED.
common_libs/movements/wave_surfer.nim existed but had never been wired to
the bot and never been tested. This doc reports what was broken, the A/B, and
the answer to the only question that matters: does the plain surfer already
dodge and win better than what we ship?
1. What was broken in the module, and what was fixed
A full read found four real defects. All are fixed in this commit; the module
now compiles, is dispatched exactly like tfil/strafe, and its TR_SURF_*
knobs are registered in env_report.nim + knownEnvNames().
- The dodge was INVERTED (the critical one). The perpendicular direction
was built from the bot→enemy bearing
(
arctan2(enemyY-botY, enemyX-botX)), while the GF is measured in the enemy→bot frame. Rotating bot→enemy by +90° points the opposite way from rotating enemy→bot by +90°, so when the safest bin was at a higher GF the bot moved toward a lower GF — i.e. toward the bin that gets hit most. Fixed by building the perpendicular from the wave's ownorigin→botbearing;+90°then provably increases GF. Verified offline (scratch probe, not committed): with the whole high-GF side marked deadly the mover now picksstrafeDir = -1(decrease GF), and with the low-GF side deadly it picks+1. - The histogram was never reset.
resetRoundcleared the waves but leftbinsaccumulating for the whole battle — a global static average forever. It is now reset to the uniform prior every round. - Fire detection tracked only the current target's energy via one scalar
prevEnergy(init 100.0). In melee a target switch silently compared two different bots' energies. Now per-enemy (seq[(id, energy)], like STRAFE); a new enemy id is seeded without emitting a wave. - The wall penalty projected the wrong point. It placed the future
position at
bot + (wave-bearing + gf·MEA)— a direction from the wave origin, not from the bot — so the wall test was meaningless. Now it projects the point on the wave circle at the candidate GF and takes the direction from the bot.
Wave geometry already used the actual fired power correctly: the one-tick
energy drop is the firepower, so speed = 20 - 3·drop is exact. The GF is
normalised by that power's arcsin(8/speed).
Known limitation (not fixed, not hidden): a low-power hit on the enemy (our bullet doing ≤ 3.0 energy) is indistinguishable from a small firepower in the energy drop, so those shots can spawn a false wave. High-power hits are rejected by the drop window. This is the standard energy-drop ambiguity.
Default parity: common_libs/tests/test_tfil_commit_env.nim → 30/30 PASS
on the working tree and on a clean git archive HEAD extraction;
test_env_report.nim → PASS. The shipped TR_MOVEMENT=tfil path is untouched.
2. The live A/B (3 arms × 15 runs, 7 rounds)
Frozen binary built from git archive HEAD (one binary, env-selected arms).
Liveness: all 45 runs OK; TR_MOVEMENT=strafe/surf confirmed present in the
bot's own [env] boot report.
2a. PRIMARY — damage/run and ROUND WINS (ab_analyze.py)
| arm | dmg/run | dmg taken/run | round wins | win% | our shots/run | hits taken/run |
|---|---|---|---|---|---|---|
tfil |
293 | 224 | 45/105 | 42.9% | 771 | 95.4 |
strafe |
250 | 198 | 37/105 | 35.2% | 776 | 88.2 |
surf |
255 | 259 | 37/105 | 35.2% | 721 | 115.3 |
Per-run damage / round-wins (never just the mean):
tfil dmg : 262 286 312 325 276 264 302 291 264 301 354 282 260 334 275
wins: 1 4 4 3 3 1 3 5 3 3 4 4 1 4 2 (/7)
strafe dmg : 182 216 262 358 244 280 238 295 281 248 224 171 274 236 248
wins: 1 2 3 5 2 5 2 3 3 1 1 1 3 3 2 (/7)
surf dmg : 303 304 280 272 260 257 226 211 225 209 269 195 268 282 270
wins: 3 4 4 2 3 5 1 0 2 1 3 2 3 2 2 (/7)
Liveness (ab_analyze.py): all arms OK — tfil 15/15 no arm env; strafe
15/15 TR_MOVEMENT=strafe applied; surf 15/15 TR_MOVEMENT=surf applied
(each confirmed in the bot's own [env] boot report, so a setting that never
reached the process would have been a loud FAIL).
Pairwise permutation (per-run, exact/MC):
| metric | A vs B | diff (A−B) | p (perm) | MW p |
|---|---|---|---|---|
| dmg/run | tfil − strafe | +42.1 | 0.0047 | 0.0032 |
| round wins | tfil − strafe | +0.53 | 0.324 | 0.192 |
| dmg/run | tfil − surf | +37.2 | 0.0030 | 0.0114 |
| round wins | tfil − surf | +0.53 | 0.325 | 0.233 |
| dmg/run | strafe − surf | −4.95 | 0.740 | 0.619 |
| round wins | strafe − surf | 0.00 | 1.000 | 0.898 |
MDE (α=0.05 two-sided, 80% power, n=15): damage/run 29.4 (10.0% of the 292.6 control mean); round wins 1.28 (42.7% of the 3.0 control mean).
The damage gap tfil→{strafe,surf} is detectable; the round-win gap is not at 15 runs (1.28-win MDE ≫ the 0.53-win observed difference). So we say the round-win difference is not detectable, not absent.
2b. SECONDARY — does the engine DODGE? (the user's explicit question)
Incoming hit rate = DrussGT hits on us ÷ DrussGT shots (capture subject counters, cross-checked against the events sidecar; the two agree to ±1 hit):
| arm | DrussGT shots | DrussGT hits on us | incoming hit rate | damage taken | dmg per enemy shot |
|---|---|---|---|---|---|
tfil |
13755 | 1431 | 10.40% | 3366 | 0.245 |
strafe |
14070 | 1323 | 9.40% | 2964 | 0.211 |
surf |
12809 | 1730 | 13.51% | 3885 | 0.303 |
Permutation on the per-run incoming hit rate: strafe is 1.03 pp LOWER than tfil (p=0.0066); surf is 3.15 pp HIGHER than tfil (p<0.0001) and 4.1 pp higher than strafe. MDE: 1.23 / 0.66 / 1.36 pp (12%/7%/10% relative).
Per-shot dodge instrument, roles swapped to measure our dodge (same
validated math as analyze_drussgt_dodge_vs_power.py / ab_dodge_analyze.py:
perp offset of our tank from the enemy bullet's aim line at arrival). Higher =
dodged further:
| arm | miss at arrival | miss/tick | lateral disp | lat (fixed 12) | per-shot hit% | our dist to enemy |
|---|---|---|---|---|---|---|
tfil |
95.5 px | 4.01 | 47.7 | 30.8 | 10.12% | 458 px |
strafe |
126.7 px | 5.17 | 78.3 | 52.4 | 9.18% | 487 px |
surf |
98.9 px | 4.48 | 87.6 | 53.6 | 13.28% | 430 px |
Permutation vs tfil: strafe dodges +31.2 px further (p<0.0001) and takes −1.0 pp hits (p=0.0086); surf dodges +3.6 px further (p=0.0009) yet takes +3.2 pp hits (p<0.0001). MDE for miss-at-arrival is 2.58 px (2.7% rel). INFERRED reading: surf's higher hit rate despite a similar miss is explained by its 28 px shorter firing range (430 vs 458), so enemy bullets arrive sooner with less time to dodge — the miss proxy does not capture that.
Enemy-bullet proximity (ab_mechanism.py, nearest live enemy bullet):
strafe spends the fewest ticks within 100 px (21.1% vs surf 28.3%, tfil 30.0%).
This is dodging capability in the sense of "time spent in the bullet's path"
and agrees with the incoming-hit-rate ranking.
Range distribution (ab_mechanism.py, per-tick distance to enemy):
| arm | mean | 0–100 | 100–200 | 200–300 | 300–400 | 400+ | central-box % |
|---|---|---|---|---|---|---|---|
tfil |
452.5 | 0.19 | 0.84 | 2.87 | 20.8 | 75.3 | 2.3 |
strafe |
484.9 | 0.07 | 0.27 | 1.19 | 11.4 | 87.0 | 1.6 |
surf |
428.8 | 0.05 | 0.28 | 0.58 | 18.5 | 80.6 | 8.1 |
Our hit rate on DrussGT by range (ab_range_bands.py) — ALL: tfil 10.8%,
surf 10.6%, strafe 9.8% — so the surfer does not convert its aggression
into more damage. DrussGT's dodge of our bullets (ab_dodge_analyze.py) misses
us most against strafe (124.3 px) and least against surf (111.9 px), i.e. our
shots are hardest for DrussGT to dodge when we are farther away.
ab_dodge_analyze.pycrashes at its final between-arm step withKeyError: 'mix'(a hardcoded arm name when no arm is literally namedmix). The per-arm table above is the part that printed before the crash.
3. Direct answer
Does the surfer dodge better than strafe and TFIL? NO.
surf has the worst incoming hit rate (13.51% vs 9.40% strafe, 10.40%
tfil), takes the most damage (259 vs 198/224 per run), and spends slightly
more time within 100 px of enemy bullets than strafe. It sits ~28 px closer
than tfil and ~56 px closer than strafe, which is where the extra hits come
from.
Does it win more? NO. It ties strafe (37/105 each, p=1.00) and is below tfil (45/105). The win gap vs tfil is not detectable at 15 runs (MDE 1.28 wins), but the damage gap is: surf deals 37/run less than tfil (p=0.003).
A surprising secondary finding, stated plainly: the user's premise —
"strafe is dodging less against DrussGT" — is not supported by this
session. strafe is the best dodger of the three (lowest incoming hit
rate, lowest damage taken, fewest hits taken, most time spent far from enemy
bullets). It nevertheless wins fewer rounds and deals less damage than
tfil (37 vs 45 wins, 250 vs 293 dmg/run, p=0.005). This is exactly the j107
trap: an arm can take fewer hits and still win fewer rounds. Winning here is
set by damage output / engagement range, not by dodge quality.
Infrastructure note: even a perfect surfer cannot be judged on dodging alone in this harness — see the strafe result. The primary axis stays damage/run and round wins.
4. GUI recipe
Set only this (the server passes its environment to the bot):
TR_MOVEMENT=surf
TR_SURF_PREF_DIST (400), TR_SURF_DIST_BAND (50), TR_SURF_WALL_MARGIN
(48), TR_SURF_RADIAL_FRAC (0.35) and TR_SURF_LOG (presence) are optional,
all defaults are the shipped ones. Unset/bogus TR_MOVEMENT → tfil.
What to look for in the GUI log (grep '^\[env\]' confirms the engine;
TR_SURF_LOG=1 prints one [surf] wave … line per danger-resolution):
- the bot should orbit the enemy at ~300–450 px and not cross the arena centre (the A/B showed 8.1% central-box occupancy vs 1.6–2.3% for the others — if you see it hanging around the middle, the radial blend is fighting the dodge);
- it should never sit on a wall: the hard wall-escape blends toward the centre once within 48 px, and the probe found 0 wall-exit ticks in a 2000-tick closed-loop sim;
- watch the incoming hit rate, not the win count, when judging "dodging" — but remember the strafe result before drawing a conclusion.
5. What the next step should be
Do not start the BitBrain upgrade for the surfer, and do not ship
TR_MOVEMENT=surf. The plain surfer is already a step backward on dodging
and no better than strafe on wins/damage, and BitBrain is an aim correction
— it cannot fix a mover that positions itself 28–56 px too close.
The evidence says the real lever is engagement range / damage output, not
dodging: tfil wins the most while being hit the most; strafe dodges the best
while winning the least. The next experiment should be a range/aggression A/B
(e.g. raise TR_SURF_PREF_DIST well above 400 and/or cut the radial blend) to
test whether the surfer can reach tfil's damage without losing its (already
poor) dodge. If it cannot, the surfer should be retired rather than upgraded.
Commands
tools/ab/ab_run.sh --arms /tmp/arms_surf.txt --runs 15 --outdir /tmp/ab/surf --conc 7 --rounds 7
python3 tools/ab/ab_analyze.py /tmp/ab/surf --reference tfil
python3 tools/ab/ab_mechanism.py /tmp/ab/surf --reference tfil
python3 tools/ab/ab_range_bands.py /tmp/ab/surf --reference tfil
# incoming hit rate + swapped per-shot dodge instrument: scratch scripts in /tmp
nim c -r --path:common_libs common_libs/tests/test_tfil_commit_env.nim
arms.txt:
tfil |
strafe | TR_MOVEMENT=strafe
surf | TR_MOVEMENT=surf