Files
SirRoboGarage/docs/surfer_wiring_ab.md
T
SirStone 6d6648ccf8 docs: wave surfer wiring A/B vs TFIL and STRAFE (surf does not beat the shipped default)
Live 3-arm x 15-run x 7-round A/B vs real DrussGT on commit 0f5cfe37.
Primary: surf ties strafe on round wins (37/105) and damage (255 vs 250/run),
both below tfil (45/105, 293/run; damage p=0.003). Incoming hit rate: surf
13.51% (worst) vs strafe 9.40% (best) and tfil 10.40%. So the plain surfer does
NOT dodge better and does NOT win more. Also records the j107 trap: strafe
dodges best yet wins fewer rounds than tfil. Next step: range/aggression A/B,
not a BitBrain upgrade.
2026-09-26 00:27:34 +02:00

12 KiB
Raw Blame History

Wiring the wave surfer as TR_MOVEMENT=surf, and a live A/B vs TFIL and STRAFE

Session: commit 0f5cfe37, frozen binary 788d6a34…, 3 arms × 15 runs × 7 rounds = 45 battles vs real DrussGT (Tank Royale physics via the robocode shim), tools/ab/ab_run.sh. All numbers below are MEASURED from those captures unless a line says INFERRED.

common_libs/movements/wave_surfer.nim existed but had never been wired to the bot and never been tested. This doc reports what was broken, the A/B, and the answer to the only question that matters: does the plain surfer already dodge and win better than what we ship?


1. What was broken in the module, and what was fixed

A full read found four real defects. All are fixed in this commit; the module now compiles, is dispatched exactly like tfil/strafe, and its TR_SURF_* knobs are registered in env_report.nim + knownEnvNames().

  1. The dodge was INVERTED (the critical one). The perpendicular direction was built from the bot→enemy bearing (arctan2(enemyY-botY, enemyX-botX)), while the GF is measured in the enemy→bot frame. Rotating bot→enemy by +90° points the opposite way from rotating enemy→bot by +90°, so when the safest bin was at a higher GF the bot moved toward a lower GF — i.e. toward the bin that gets hit most. Fixed by building the perpendicular from the wave's own origin→bot bearing; +90° then provably increases GF. Verified offline (scratch probe, not committed): with the whole high-GF side marked deadly the mover now picks strafeDir = -1 (decrease GF), and with the low-GF side deadly it picks +1.
  2. The histogram was never reset. resetRound cleared the waves but left bins accumulating for the whole battle — a global static average forever. It is now reset to the uniform prior every round.
  3. Fire detection tracked only the current target's energy via one scalar prevEnergy (init 100.0). In melee a target switch silently compared two different bots' energies. Now per-enemy (seq[(id, energy)], like STRAFE); a new enemy id is seeded without emitting a wave.
  4. The wall penalty projected the wrong point. It placed the future position at bot + (wave-bearing + gf·MEA) — a direction from the wave origin, not from the bot — so the wall test was meaningless. Now it projects the point on the wave circle at the candidate GF and takes the direction from the bot.

Wave geometry already used the actual fired power correctly: the one-tick energy drop is the firepower, so speed = 20 - 3·drop is exact. The GF is normalised by that power's arcsin(8/speed).

Known limitation (not fixed, not hidden): a low-power hit on the enemy (our bullet doing ≤ 3.0 energy) is indistinguishable from a small firepower in the energy drop, so those shots can spawn a false wave. High-power hits are rejected by the drop window. This is the standard energy-drop ambiguity.

Default parity: common_libs/tests/test_tfil_commit_env.nim → 30/30 PASS on the working tree and on a clean git archive HEAD extraction; test_env_report.nim → PASS. The shipped TR_MOVEMENT=tfil path is untouched.


2. The live A/B (3 arms × 15 runs, 7 rounds)

Frozen binary built from git archive HEAD (one binary, env-selected arms). Liveness: all 45 runs OK; TR_MOVEMENT=strafe/surf confirmed present in the bot's own [env] boot report.

2a. PRIMARY — damage/run and ROUND WINS (ab_analyze.py)

arm dmg/run dmg taken/run round wins win% our shots/run hits taken/run
tfil 293 224 45/105 42.9% 771 95.4
strafe 250 198 37/105 35.2% 776 88.2
surf 255 259 37/105 35.2% 721 115.3

Pairwise permutation (per-run, exact/MC):

metric A vs B diff (A−B) p (perm) MW p
dmg/run tfil − strafe +42.1 0.0047 0.0032
round wins tfil − strafe +0.53 0.324 0.192
dmg/run tfil − surf +37.2 0.0030 0.0114
round wins tfil − surf +0.53 0.325 0.233
dmg/run strafe − surf −4.95 0.740 0.619
round wins strafe − surf 0.00 1.000 0.898

MDE (α=0.05 two-sided, 80% power, n=15): damage/run 29.4 (10.0% of the 292.6 control mean); round wins 1.28 (42.7% of the 3.0 control mean).

The damage gap tfil→{strafe,surf} is detectable; the round-win gap is not at 15 runs (1.28-win MDE ≫ the 0.53-win observed difference). So we say the round-win difference is not detectable, not absent.

2b. SECONDARY — does the engine DODGE? (the user's explicit question)

Incoming hit rate = DrussGT hits on us ÷ DrussGT shots (capture subject counters, cross-checked against the events sidecar; the two agree to ±1 hit):

arm DrussGT shots DrussGT hits on us incoming hit rate damage taken dmg per enemy shot
tfil 13755 1431 10.40% 3366 0.245
strafe 14070 1323 9.40% 2964 0.211
surf 12809 1730 13.51% 3885 0.303

Permutation on the per-run incoming hit rate: strafe is 1.03 pp LOWER than tfil (p=0.0066); surf is 3.15 pp HIGHER than tfil (p<0.0001) and 4.1 pp higher than strafe. MDE: 1.23 / 0.66 / 1.36 pp (12%/7%/10% relative).

Per-shot dodge instrument, roles swapped to measure our dodge (same validated math as analyze_drussgt_dodge_vs_power.py / ab_dodge_analyze.py: perp offset of our tank from the enemy bullet's aim line at arrival). Higher = dodged further:

arm miss at arrival miss/tick lateral disp lat (fixed 12) per-shot hit% our dist to enemy
tfil 95.5 px 4.01 47.7 30.8 10.12% 458 px
strafe 126.7 px 5.17 78.3 52.4 9.18% 487 px
surf 98.9 px 4.48 87.6 53.6 13.28% 430 px

Permutation vs tfil: strafe dodges +31.2 px further (p<0.0001) and takes −1.0 pp hits (p=0.0086); surf dodges +3.6 px further (p=0.0009) yet takes +3.2 pp hits (p<0.0001). MDE for miss-at-arrival is 2.58 px (2.7% rel). INFERRED reading: surf's higher hit rate despite a similar miss is explained by its 28 px shorter firing range (430 vs 458), so enemy bullets arrive sooner with less time to dodge — the miss proxy does not capture that.

Enemy-bullet proximity (ab_mechanism.py, nearest live enemy bullet): strafe spends the fewest ticks within 100 px (21.1% vs surf 28.3%, tfil 30.0%). This is dodging capability in the sense of "time spent in the bullet's path" and agrees with the incoming-hit-rate ranking.

Range distribution (ab_mechanism.py, per-tick distance to enemy):

arm mean 0–100 100–200 200–300 300–400 400+ central-box %
tfil 452.5 0.19 0.84 2.87 20.8 75.3 2.3
strafe 484.9 0.07 0.27 1.19 11.4 87.0 1.6
surf 428.8 0.05 0.28 0.58 18.5 80.6 8.1

Our hit rate on DrussGT by range (ab_range_bands.py) — ALL: tfil 10.8%, surf 10.6%, strafe 9.8% — so the surfer does not convert its aggression into more damage. DrussGT's dodge of our bullets (ab_dodge_analyze.py) misses us most against strafe (124.3 px) and least against surf (111.9 px), i.e. our shots are hardest for DrussGT to dodge when we are farther away.

ab_dodge_analyze.py crashes at its final between-arm step with KeyError: 'mix' (a hardcoded arm name when no arm is literally named mix). The per-arm table above is the part that printed before the crash.


3. Direct answer

Does the surfer dodge better than strafe and TFIL? NO. surf has the worst incoming hit rate (13.51% vs 9.40% strafe, 10.40% tfil), takes the most damage (259 vs 198/224 per run), and spends slightly more time within 100 px of enemy bullets than strafe. It sits ~28 px closer than tfil and ~56 px closer than strafe, which is where the extra hits come from.

Does it win more? NO. It ties strafe (37/105 each, p=1.00) and is below tfil (45/105). The win gap vs tfil is not detectable at 15 runs (MDE 1.28 wins), but the damage gap is: surf deals 37/run less than tfil (p=0.003).

A surprising secondary finding, stated plainly: the user's premise — "strafe is dodging less against DrussGT" — is not supported by this session. strafe is the best dodger of the three (lowest incoming hit rate, lowest damage taken, fewest hits taken, most time spent far from enemy bullets). It nevertheless wins fewer rounds and deals less damage than tfil (37 vs 45 wins, 250 vs 293 dmg/run, p=0.005). This is exactly the j107 trap: an arm can take fewer hits and still win fewer rounds. Winning here is set by damage output / engagement range, not by dodge quality.

Infrastructure note: even a perfect surfer cannot be judged on dodging alone in this harness — see the strafe result. The primary axis stays damage/run and round wins.


4. GUI recipe

Set only this (the server passes its environment to the bot):

TR_MOVEMENT=surf

TR_SURF_PREF_DIST (400), TR_SURF_DIST_BAND (50), TR_SURF_WALL_MARGIN (48), TR_SURF_RADIAL_FRAC (0.35) and TR_SURF_LOG (presence) are optional, all defaults are the shipped ones. Unset/bogus TR_MOVEMENT → tfil.

What to look for in the GUI log (grep '^\[env\]' confirms the engine; TR_SURF_LOG=1 prints one [surf] wave … line per danger-resolution):

  • the bot should orbit the enemy at ~300–450 px and not cross the arena centre (the A/B showed 8.1% central-box occupancy vs 1.6–2.3% for the others — if you see it hanging around the middle, the radial blend is fighting the dodge);
  • it should never sit on a wall: the hard wall-escape blends toward the centre once within 48 px, and the probe found 0 wall-exit ticks in a 2000-tick closed-loop sim;
  • watch the incoming hit rate, not the win count, when judging "dodging" — but remember the strafe result before drawing a conclusion.

5. What the next step should be

Do not start the BitBrain upgrade for the surfer, and do not ship TR_MOVEMENT=surf. The plain surfer is already a step backward on dodging and no better than strafe on wins/damage, and BitBrain is an aim correction — it cannot fix a mover that positions itself 28–56 px too close.

The evidence says the real lever is engagement range / damage output, not dodging: tfil wins the most while being hit the most; strafe dodges the best while winning the least. The next experiment should be a range/aggression A/B (e.g. raise TR_SURF_PREF_DIST well above 400 and/or cut the radial blend) to test whether the surfer can reach tfil's damage without losing its (already poor) dodge. If it cannot, the surfer should be retired rather than upgraded.


Commands

tools/ab/ab_run.sh --arms /tmp/arms_surf.txt --runs 15 --outdir /tmp/ab/surf --conc 7 --rounds 7
python3 tools/ab/ab_analyze.py    /tmp/ab/surf --reference tfil
python3 tools/ab/ab_mechanism.py  /tmp/ab/surf --reference tfil
python3 tools/ab/ab_range_bands.py /tmp/ab/surf --reference tfil
# incoming hit rate + swapped per-shot dodge instrument: scratch scripts in /tmp
nim c -r --path:common_libs common_libs/tests/test_tfil_commit_env.nim

arms.txt:

tfil |
strafe | TR_MOVEMENT=strafe
surf | TR_MOVEMENT=surf