# Wiring the wave surfer as `TR_MOVEMENT=surf`, and a live A/B vs TFIL and STRAFE **Session:** commit `0f5cfe37`, frozen binary `788d6a34…`, 3 arms × 15 runs × 7 rounds = 45 battles vs **real DrussGT** (Tank Royale physics via the robocode shim), `tools/ab/ab_run.sh`. All numbers below are **MEASURED** from those captures unless a line says INFERRED. `common_libs/movements/wave_surfer.nim` existed but had **never been wired to the bot and never been tested**. This doc reports what was broken, the A/B, and the answer to the only question that matters: *does the plain surfer already dodge and win better than what we ship?* --- ## 1. What was broken in the module, and what was fixed A full read found four real defects. All are fixed in this commit; the module now compiles, is dispatched exactly like `tfil`/`strafe`, and its `TR_SURF_*` knobs are registered in `env_report.nim` + `knownEnvNames()`. 1. **The dodge was INVERTED (the critical one).** The perpendicular direction was built from the **bot→enemy** bearing (`arctan2(enemyY-botY, enemyX-botX)`), while the GF is measured in the **enemy→bot** frame. Rotating bot→enemy by +90° points the *opposite* way from rotating enemy→bot by +90°, so when the safest bin was at a higher GF the bot moved toward a *lower* GF — i.e. **toward the bin that gets hit most**. Fixed by building the perpendicular from the wave's own `origin→bot` bearing; `+90°` then provably increases GF. *Verified offline* (scratch probe, not committed): with the whole high-GF side marked deadly the mover now picks `strafeDir = -1` (decrease GF), and with the low-GF side deadly it picks `+1`. 2. **The histogram was never reset.** `resetRound` cleared the waves but left `bins` accumulating for the whole battle — a global static average forever. It is now reset to the uniform prior every round. 3. **Fire detection tracked only the current target's energy** via one scalar `prevEnergy` (init 100.0). In melee a target switch silently compared two different bots' energies. Now per-enemy (`seq[(id, energy)]`, like STRAFE); a new enemy id is seeded without emitting a wave. 4. **The wall penalty projected the wrong point.** It placed the future position at `bot + (wave-bearing + gf·MEA)` — a direction from the *wave origin*, not from the bot — so the wall test was meaningless. Now it projects the point on the wave circle at the candidate GF and takes the direction from the bot. Wave geometry already used the **actual fired power** correctly: the one-tick energy drop *is* the firepower, so `speed = 20 - 3·drop` is exact. The GF is normalised by that power's `arcsin(8/speed)`. **Known limitation (not fixed, not hidden):** a *low-power hit on the enemy* (our bullet doing ≤ 3.0 energy) is indistinguishable from a small firepower in the energy drop, so those shots can spawn a false wave. High-power hits are rejected by the drop window. This is the standard energy-drop ambiguity. **Default parity:** `common_libs/tests/test_tfil_commit_env.nim` → 30/30 PASS on the working tree **and** on a clean `git archive HEAD` extraction; `test_env_report.nim` → PASS. The shipped `TR_MOVEMENT=tfil` path is untouched. --- ## 2. The live A/B (3 arms × 15 runs, 7 rounds) Frozen binary built from `git archive HEAD` (one binary, env-selected arms). Liveness: all 45 runs OK; `TR_MOVEMENT=strafe/surf` confirmed present in the bot's own `[env]` boot report. ### 2a. PRIMARY — damage/run and ROUND WINS (`ab_analyze.py`) | arm | dmg/run | dmg taken/run | round wins | win% | our shots/run | hits taken/run | |-----|--------:|--------------:|-----------:|-----:|--------------:|---------------:| | `tfil` | **293** | 224 | **45/105** | **42.9%** | 771 | 95.4 | | `strafe` | 250 | **198** | 37/105 | 35.2% | 776 | 88.2 | | `surf` | 255 | 259 | 37/105 | 35.2% | 721 | 115.3 | Pairwise permutation (per-run, exact/MC): | metric | A vs B | diff (A−B) | p (perm) | MW p | |--------|--------|-----------:|---------:|-----:| | dmg/run | tfil − strafe | +42.1 | **0.0047** | 0.0032 | | round wins | tfil − strafe | +0.53 | 0.324 | 0.192 | | dmg/run | tfil − surf | +37.2 | **0.0030** | 0.0114 | | round wins | tfil − surf | +0.53 | 0.325 | 0.233 | | dmg/run | strafe − surf | −4.95 | 0.740 | 0.619 | | round wins | strafe − surf | 0.00 | 1.000 | 0.898 | **MDE (α=0.05 two-sided, 80% power, n=15):** damage/run **29.4** (10.0% of the 292.6 control mean); round wins **1.28** (42.7% of the 3.0 control mean). > The damage gap tfil→{strafe,surf} is **detectable**; the round-win gap is > **not** at 15 runs (1.28-win MDE ≫ the 0.53-win observed difference). So we > say the round-win difference is **not detectable**, not absent. ### 2b. SECONDARY — does the engine DODGE? (the user's explicit question) **Incoming hit rate** = DrussGT hits on us ÷ DrussGT shots (capture subject counters, cross-checked against the events sidecar; the two agree to ±1 hit): | arm | DrussGT shots | DrussGT hits on us | **incoming hit rate** | damage taken | dmg per enemy shot | |-----|--------------:|-------------------:|----------------------:|-------------:|-------------------:| | `tfil` | 13755 | 1431 | 10.40% | 3366 | 0.245 | | `strafe` | 14070 | 1323 | **9.40%** | **2964** | **0.211** | | `surf` | 12809 | 1730 | **13.51%** | 3885 | 0.303 | Permutation on the per-run incoming hit rate: strafe is **1.03 pp LOWER** than tfil (p=0.0066); surf is **3.15 pp HIGHER** than tfil (p<0.0001) and **4.1 pp higher than strafe**. MDE: 1.23 / 0.66 / 1.36 pp (12%/7%/10% relative). **Per-shot dodge instrument, roles swapped to measure *our* dodge** (same validated math as `analyze_drussgt_dodge_vs_power.py` / `ab_dodge_analyze.py`: perp offset of our tank from the enemy bullet's aim line at arrival). Higher = dodged further: | arm | miss at arrival | miss/tick | lateral disp | lat (fixed 12) | per-shot hit% | our dist to enemy | |-----|----------------:|----------:|-------------:|---------------:|--------------:|------------------:| | `tfil` | 95.5 px | 4.01 | 47.7 | 30.8 | 10.12% | 458 px | | `strafe` | **126.7 px** | **5.17** | 78.3 | 52.4 | **9.18%** | 487 px | | `surf` | 98.9 px | 4.48 | 87.6 | 53.6 | **13.28%** | 430 px | Permutation vs tfil: strafe dodges **+31.2 px further** (p<0.0001) and takes **−1.0 pp** hits (p=0.0086); surf dodges +3.6 px further (p=0.0009) yet takes **+3.2 pp** hits (p<0.0001). MDE for miss-at-arrival is 2.58 px (2.7% rel). INFERRED reading: surf's higher hit rate despite a similar miss is explained by its **28 px shorter firing range** (430 vs 458), so enemy bullets arrive sooner with less time to dodge — the miss proxy does not capture that. **Enemy-bullet proximity** (`ab_mechanism.py`, nearest live enemy bullet): strafe spends the fewest ticks within 100 px (21.1% vs surf 28.3%, tfil 30.0%). This is *dodging capability in the sense of "time spent in the bullet's path"* and agrees with the incoming-hit-rate ranking. **Range distribution** (`ab_mechanism.py`, per-tick distance to enemy): | arm | mean | 0–100 | 100–200 | 200–300 | 300–400 | 400+ | central-box % | |-----|-----:|------:|--------:|--------:|--------:|-----:|--------------:| | `tfil` | 452.5 | 0.19 | 0.84 | 2.87 | 20.8 | 75.3 | 2.3 | | `strafe` | 484.9 | 0.07 | 0.27 | 1.19 | 11.4 | **87.0** | 1.6 | | `surf` | 428.8 | 0.05 | 0.28 | 0.58 | 18.5 | 80.6 | **8.1** | Our hit rate on DrussGT by range (`ab_range_bands.py`) — ALL: tfil 10.8%, surf 10.6%, strafe 9.8% — so the surfer does **not** convert its aggression into more damage. DrussGT's dodge of our bullets (`ab_dodge_analyze.py`) misses us most against strafe (124.3 px) and least against surf (111.9 px), i.e. our shots are hardest for DrussGT to dodge when we are *farther* away. > `ab_dodge_analyze.py` crashes at its final between-arm step with > `KeyError: 'mix'` (a hardcoded arm name when no arm is literally named > `mix`). The per-arm table above is the part that printed before the crash. --- ## 3. Direct answer **Does the surfer dodge better than strafe and TFIL? NO.** `surf` has the **worst** incoming hit rate (13.51% vs 9.40% strafe, 10.40% tfil), takes the **most** damage (259 vs 198/224 per run), and spends slightly more time within 100 px of enemy bullets than strafe. It sits ~28 px closer than tfil and ~56 px closer than strafe, which is where the extra hits come from. **Does it win more? NO.** It ties strafe (37/105 each, p=1.00) and is below tfil (45/105). The win gap vs tfil is **not detectable** at 15 runs (MDE 1.28 wins), but the damage gap is: surf deals 37/run **less** than tfil (p=0.003). **A surprising secondary finding, stated plainly:** the user's premise — *"strafe is dodging less against DrussGT"* — is **not supported** by this session. `strafe` is the **best** dodger of the three (lowest incoming hit rate, lowest damage taken, fewest hits taken, most time spent far from enemy bullets). It nevertheless **wins fewer rounds and deals less damage than tfil** (37 vs 45 wins, 250 vs 293 dmg/run, p=0.005). This is exactly the j107 trap: *an arm can take fewer hits and still win fewer rounds.* Winning here is set by damage output / engagement range, not by dodge quality. **Infrastructure note:** even a perfect surfer cannot be judged on dodging alone in this harness — see the strafe result. The primary axis stays damage/run and round wins. --- ## 4. GUI recipe Set **only** this (the server passes its environment to the bot): ``` TR_MOVEMENT=surf ``` `TR_SURF_PREF_DIST` (400), `TR_SURF_DIST_BAND` (50), `TR_SURF_WALL_MARGIN` (48), `TR_SURF_RADIAL_FRAC` (0.35) and `TR_SURF_LOG` (presence) are optional, all defaults are the shipped ones. Unset/bogus `TR_MOVEMENT` → `tfil`. **What to look for in the GUI log** (`grep '^\[env\]'` confirms the engine; `TR_SURF_LOG=1` prints one `[surf] wave …` line per danger-resolution): * the bot should **orbit** the enemy at ~300–450 px and *not* cross the arena centre (the A/B showed 8.1% central-box occupancy vs 1.6–2.3% for the others — if you see it hanging around the middle, the radial blend is fighting the dodge); * it should **never** sit on a wall: the hard wall-escape blends toward the centre once within 48 px, and the probe found 0 wall-exit ticks in a 2000-tick closed-loop sim; * watch the **incoming hit rate**, not the win count, when judging "dodging" — but remember the strafe result before drawing a conclusion. --- ## 5. What the next step should be **Do not start the BitBrain upgrade for the surfer, and do not ship `TR_MOVEMENT=surf`.** The plain surfer is already a step *backward* on dodging and no better than strafe on wins/damage, and BitBrain is an *aim* correction — it cannot fix a mover that positions itself 28–56 px too close. The evidence says the real lever is **engagement range / damage output**, not dodging: tfil wins the most while being hit the most; strafe dodges the best while winning the least. The next experiment should be a range/aggression A/B (e.g. raise `TR_SURF_PREF_DIST` well above 400 and/or cut the radial blend) to test whether the surfer can reach tfil's damage without losing its (already poor) dodge. If it cannot, the surfer should be retired rather than upgraded. --- ### Commands ```sh tools/ab/ab_run.sh --arms /tmp/arms_surf.txt --runs 15 --outdir /tmp/ab/surf --conc 7 --rounds 7 python3 tools/ab/ab_analyze.py /tmp/ab/surf --reference tfil python3 tools/ab/ab_mechanism.py /tmp/ab/surf --reference tfil python3 tools/ab/ab_range_bands.py /tmp/ab/surf --reference tfil # incoming hit rate + swapped per-shot dodge instrument: scratch scripts in /tmp nim c -r --path:common_libs common_libs/tests/test_tfil_commit_env.nim ``` `arms.txt`: ``` tfil | strafe | TR_MOVEMENT=strafe surf | TR_MOVEMENT=surf ```