254 lines
12 KiB
Markdown
254 lines
12 KiB
Markdown
# Wiring the wave surfer as `TR_MOVEMENT=surf`, and a live A/B vs TFIL and STRAFE
|
||
|
||
**Session:** commit `0f5cfe37`, frozen binary `788d6a34…`, 3 arms × 15 runs × 7
|
||
rounds = 45 battles vs **real DrussGT** (Tank Royale physics via the robocode
|
||
shim), `tools/ab/ab_run.sh`. All numbers below are **MEASURED** from those
|
||
captures unless a line says INFERRED.
|
||
|
||
`common_libs/movements/wave_surfer.nim` existed but had **never been wired to
|
||
the bot and never been tested**. This doc reports what was broken, the A/B, and
|
||
the answer to the only question that matters: *does the plain surfer already
|
||
dodge and win better than what we ship?*
|
||
|
||
---
|
||
|
||
## 1. What was broken in the module, and what was fixed
|
||
|
||
A full read found four real defects. All are fixed in this commit; the module
|
||
now compiles, is dispatched exactly like `tfil`/`strafe`, and its `TR_SURF_*`
|
||
knobs are registered in `env_report.nim` + `knownEnvNames()`.
|
||
|
||
1. **The dodge was INVERTED (the critical one).** The perpendicular direction
|
||
was built from the **bot→enemy** bearing
|
||
(`arctan2(enemyY-botY, enemyX-botX)`), while the GF is measured in the
|
||
**enemy→bot** frame. Rotating bot→enemy by +90° points the *opposite* way
|
||
from rotating enemy→bot by +90°, so when the safest bin was at a higher GF
|
||
the bot moved toward a *lower* GF — i.e. **toward the bin that gets hit
|
||
most**. Fixed by building the perpendicular from the wave's own
|
||
`origin→bot` bearing; `+90°` then provably increases GF.
|
||
*Verified offline* (scratch probe, not committed): with the whole high-GF
|
||
side marked deadly the mover now picks `strafeDir = -1` (decrease GF), and
|
||
with the low-GF side deadly it picks `+1`.
|
||
2. **The histogram was never reset.** `resetRound` cleared the waves but left
|
||
`bins` accumulating for the whole battle — a global static average forever.
|
||
It is now reset to the uniform prior every round.
|
||
3. **Fire detection tracked only the current target's energy** via one scalar
|
||
`prevEnergy` (init 100.0). In melee a target switch silently compared two
|
||
different bots' energies. Now per-enemy (`seq[(id, energy)]`, like STRAFE);
|
||
a new enemy id is seeded without emitting a wave.
|
||
4. **The wall penalty projected the wrong point.** It placed the future
|
||
position at `bot + (wave-bearing + gf·MEA)` — a direction from the *wave
|
||
origin*, not from the bot — so the wall test was meaningless. Now it
|
||
projects the point on the wave circle at the candidate GF and takes the
|
||
direction from the bot.
|
||
|
||
Wave geometry already used the **actual fired power** correctly: the one-tick
|
||
energy drop *is* the firepower, so `speed = 20 - 3·drop` is exact. The GF is
|
||
normalised by that power's `arcsin(8/speed)`.
|
||
|
||
**Known limitation (not fixed, not hidden):** a *low-power hit on the enemy*
|
||
(our bullet doing ≤ 3.0 energy) is indistinguishable from a small firepower in
|
||
the energy drop, so those shots can spawn a false wave. High-power hits are
|
||
rejected by the drop window. This is the standard energy-drop ambiguity.
|
||
|
||
**Default parity:** `common_libs/tests/test_tfil_commit_env.nim` → 30/30 PASS
|
||
on the working tree **and** on a clean `git archive HEAD` extraction;
|
||
`test_env_report.nim` → PASS. The shipped `TR_MOVEMENT=tfil` path is untouched.
|
||
|
||
---
|
||
|
||
## 2. The live A/B (3 arms × 15 runs, 7 rounds)
|
||
|
||
Frozen binary built from `git archive HEAD` (one binary, env-selected arms).
|
||
Liveness: all 45 runs OK; `TR_MOVEMENT=strafe/surf` confirmed present in the
|
||
bot's own `[env]` boot report.
|
||
|
||
### 2a. PRIMARY — damage/run and ROUND WINS (`ab_analyze.py`)
|
||
|
||
| arm | dmg/run | dmg taken/run | round wins | win% | our shots/run | hits taken/run |
|
||
|-----|--------:|--------------:|-----------:|-----:|--------------:|---------------:|
|
||
| `tfil` | **293** | 224 | **45/105** | **42.9%** | 771 | 95.4 |
|
||
| `strafe` | 250 | **198** | 37/105 | 35.2% | 776 | 88.2 |
|
||
| `surf` | 255 | 259 | 37/105 | 35.2% | 721 | 115.3 |
|
||
|
||
Per-run damage / round-wins (never just the mean):
|
||
|
||
```
|
||
tfil dmg : 262 286 312 325 276 264 302 291 264 301 354 282 260 334 275
|
||
wins: 1 4 4 3 3 1 3 5 3 3 4 4 1 4 2 (/7)
|
||
strafe dmg : 182 216 262 358 244 280 238 295 281 248 224 171 274 236 248
|
||
wins: 1 2 3 5 2 5 2 3 3 1 1 1 3 3 2 (/7)
|
||
surf dmg : 303 304 280 272 260 257 226 211 225 209 269 195 268 282 270
|
||
wins: 3 4 4 2 3 5 1 0 2 1 3 2 3 2 2 (/7)
|
||
```
|
||
|
||
Liveness (`ab_analyze.py`): all arms OK — `tfil` 15/15 no arm env; `strafe`
|
||
15/15 `TR_MOVEMENT=strafe` applied; `surf` 15/15 `TR_MOVEMENT=surf` applied
|
||
(each confirmed in the bot's own `[env]` boot report, so a setting that never
|
||
reached the process would have been a loud FAIL).
|
||
|
||
Pairwise permutation (per-run, exact/MC):
|
||
|
||
| metric | A vs B | diff (A−B) | p (perm) | MW p |
|
||
|--------|--------|-----------:|---------:|-----:|
|
||
| dmg/run | tfil − strafe | +42.1 | **0.0047** | 0.0032 |
|
||
| round wins | tfil − strafe | +0.53 | 0.324 | 0.192 |
|
||
| dmg/run | tfil − surf | +37.2 | **0.0030** | 0.0114 |
|
||
| round wins | tfil − surf | +0.53 | 0.325 | 0.233 |
|
||
| dmg/run | strafe − surf | −4.95 | 0.740 | 0.619 |
|
||
| round wins | strafe − surf | 0.00 | 1.000 | 0.898 |
|
||
|
||
**MDE (α=0.05 two-sided, 80% power, n=15):** damage/run **29.4** (10.0% of the
|
||
292.6 control mean); round wins **1.28** (42.7% of the 3.0 control mean).
|
||
|
||
> The damage gap tfil→{strafe,surf} is **detectable**; the round-win gap is
|
||
> **not** at 15 runs (1.28-win MDE ≫ the 0.53-win observed difference). So we
|
||
> say the round-win difference is **not detectable**, not absent.
|
||
|
||
### 2b. SECONDARY — does the engine DODGE? (the user's explicit question)
|
||
|
||
**Incoming hit rate** = DrussGT hits on us ÷ DrussGT shots (capture subject
|
||
counters, cross-checked against the events sidecar; the two agree to ±1 hit):
|
||
|
||
| arm | DrussGT shots | DrussGT hits on us | **incoming hit rate** | damage taken | dmg per enemy shot |
|
||
|-----|--------------:|-------------------:|----------------------:|-------------:|-------------------:|
|
||
| `tfil` | 13755 | 1431 | 10.40% | 3366 | 0.245 |
|
||
| `strafe` | 14070 | 1323 | **9.40%** | **2964** | **0.211** |
|
||
| `surf` | 12809 | 1730 | **13.51%** | 3885 | 0.303 |
|
||
|
||
Permutation on the per-run incoming hit rate: strafe is **1.03 pp LOWER** than
|
||
tfil (p=0.0066); surf is **3.15 pp HIGHER** than tfil (p<0.0001) and **4.1 pp
|
||
higher than strafe**. MDE: 1.23 / 0.66 / 1.36 pp (12%/7%/10% relative).
|
||
|
||
**Per-shot dodge instrument, roles swapped to measure *our* dodge** (same
|
||
validated math as `analyze_drussgt_dodge_vs_power.py` / `ab_dodge_analyze.py`:
|
||
perp offset of our tank from the enemy bullet's aim line at arrival). Higher =
|
||
dodged further:
|
||
|
||
| arm | miss at arrival | miss/tick | lateral disp | lat (fixed 12) | per-shot hit% | our dist to enemy |
|
||
|-----|----------------:|----------:|-------------:|---------------:|--------------:|------------------:|
|
||
| `tfil` | 95.5 px | 4.01 | 47.7 | 30.8 | 10.12% | 458 px |
|
||
| `strafe` | **126.7 px** | **5.17** | 78.3 | 52.4 | **9.18%** | 487 px |
|
||
| `surf` | 98.9 px | 4.48 | 87.6 | 53.6 | **13.28%** | 430 px |
|
||
|
||
Permutation vs tfil: strafe dodges **+31.2 px further** (p<0.0001) and takes
|
||
**−1.0 pp** hits (p=0.0086); surf dodges +3.6 px further (p=0.0009) yet takes
|
||
**+3.2 pp** hits (p<0.0001). MDE for miss-at-arrival is 2.58 px (2.7% rel).
|
||
INFERRED reading: surf's higher hit rate despite a similar miss is explained
|
||
by its **28 px shorter firing range** (430 vs 458), so enemy bullets arrive
|
||
sooner with less time to dodge — the miss proxy does not capture that.
|
||
|
||
**Enemy-bullet proximity** (`ab_mechanism.py`, nearest live enemy bullet):
|
||
strafe spends the fewest ticks within 100 px (21.1% vs surf 28.3%, tfil 30.0%).
|
||
This is *dodging capability in the sense of "time spent in the bullet's path"*
|
||
and agrees with the incoming-hit-rate ranking.
|
||
|
||
**Range distribution** (`ab_mechanism.py`, per-tick distance to enemy):
|
||
|
||
| arm | mean | 0–100 | 100–200 | 200–300 | 300–400 | 400+ | central-box % |
|
||
|-----|-----:|------:|--------:|--------:|--------:|-----:|--------------:|
|
||
| `tfil` | 452.5 | 0.19 | 0.84 | 2.87 | 20.8 | 75.3 | 2.3 |
|
||
| `strafe` | 484.9 | 0.07 | 0.27 | 1.19 | 11.4 | **87.0** | 1.6 |
|
||
| `surf` | 428.8 | 0.05 | 0.28 | 0.58 | 18.5 | 80.6 | **8.1** |
|
||
|
||
Our hit rate on DrussGT by range (`ab_range_bands.py`) — ALL: tfil 10.8%,
|
||
surf 10.6%, strafe 9.8% — so the surfer does **not** convert its aggression
|
||
into more damage. DrussGT's dodge of our bullets (`ab_dodge_analyze.py`) misses
|
||
us most against strafe (124.3 px) and least against surf (111.9 px), i.e. our
|
||
shots are hardest for DrussGT to dodge when we are *farther* away.
|
||
|
||
> `ab_dodge_analyze.py` crashes at its final between-arm step with
|
||
> `KeyError: 'mix'` (a hardcoded arm name when no arm is literally named
|
||
> `mix`). The per-arm table above is the part that printed before the crash.
|
||
|
||
---
|
||
|
||
## 3. Direct answer
|
||
|
||
**Does the surfer dodge better than strafe and TFIL? NO.**
|
||
`surf` has the **worst** incoming hit rate (13.51% vs 9.40% strafe, 10.40%
|
||
tfil), takes the **most** damage (259 vs 198/224 per run), and spends slightly
|
||
more time within 100 px of enemy bullets than strafe. It sits ~28 px closer
|
||
than tfil and ~56 px closer than strafe, which is where the extra hits come
|
||
from.
|
||
|
||
**Does it win more? NO.** It ties strafe (37/105 each, p=1.00) and is below
|
||
tfil (45/105). The win gap vs tfil is **not detectable** at 15 runs (MDE 1.28
|
||
wins), but the damage gap is: surf deals 37/run **less** than tfil (p=0.003).
|
||
|
||
**A surprising secondary finding, stated plainly:** the user's premise —
|
||
*"strafe is dodging less against DrussGT"* — is **not supported** by this
|
||
session. `strafe` is the **best** dodger of the three (lowest incoming hit
|
||
rate, lowest damage taken, fewest hits taken, most time spent far from enemy
|
||
bullets). It nevertheless **wins fewer rounds and deals less damage than
|
||
tfil** (37 vs 45 wins, 250 vs 293 dmg/run, p=0.005). This is exactly the j107
|
||
trap: *an arm can take fewer hits and still win fewer rounds.* Winning here is
|
||
set by damage output / engagement range, not by dodge quality.
|
||
|
||
**Infrastructure note:** even a perfect surfer cannot be judged on dodging
|
||
alone in this harness — see the strafe result. The primary axis stays
|
||
damage/run and round wins.
|
||
|
||
---
|
||
|
||
## 4. GUI recipe
|
||
|
||
Set **only** this (the server passes its environment to the bot):
|
||
|
||
```
|
||
TR_MOVEMENT=surf
|
||
```
|
||
|
||
`TR_SURF_PREF_DIST` (400), `TR_SURF_DIST_BAND` (50), `TR_SURF_WALL_MARGIN`
|
||
(48), `TR_SURF_RADIAL_FRAC` (0.35) and `TR_SURF_LOG` (presence) are optional,
|
||
all defaults are the shipped ones. Unset/bogus `TR_MOVEMENT` → `tfil`.
|
||
|
||
**What to look for in the GUI log** (`grep '^\[env\]'` confirms the engine;
|
||
`TR_SURF_LOG=1` prints one `[surf] wave …` line per danger-resolution):
|
||
|
||
* the bot should **orbit** the enemy at ~300–450 px and *not* cross the arena
|
||
centre (the A/B showed 8.1% central-box occupancy vs 1.6–2.3% for the
|
||
others — if you see it hanging around the middle, the radial blend is
|
||
fighting the dodge);
|
||
* it should **never** sit on a wall: the hard wall-escape blends toward the
|
||
centre once within 48 px, and the probe found 0 wall-exit ticks in a 2000-tick
|
||
closed-loop sim;
|
||
* watch the **incoming hit rate**, not the win count, when judging "dodging" —
|
||
but remember the strafe result before drawing a conclusion.
|
||
|
||
---
|
||
|
||
## 5. What the next step should be
|
||
|
||
**Do not start the BitBrain upgrade for the surfer, and do not ship
|
||
`TR_MOVEMENT=surf`.** The plain surfer is already a step *backward* on dodging
|
||
and no better than strafe on wins/damage, and BitBrain is an *aim* correction
|
||
— it cannot fix a mover that positions itself 28–56 px too close.
|
||
|
||
The evidence says the real lever is **engagement range / damage output**, not
|
||
dodging: tfil wins the most while being hit the most; strafe dodges the best
|
||
while winning the least. The next experiment should be a range/aggression A/B
|
||
(e.g. raise `TR_SURF_PREF_DIST` well above 400 and/or cut the radial blend) to
|
||
test whether the surfer can reach tfil's damage without losing its (already
|
||
poor) dodge. If it cannot, the surfer should be retired rather than upgraded.
|
||
|
||
---
|
||
|
||
### Commands
|
||
|
||
```sh
|
||
tools/ab/ab_run.sh --arms /tmp/arms_surf.txt --runs 15 --outdir /tmp/ab/surf --conc 7 --rounds 7
|
||
python3 tools/ab/ab_analyze.py /tmp/ab/surf --reference tfil
|
||
python3 tools/ab/ab_mechanism.py /tmp/ab/surf --reference tfil
|
||
python3 tools/ab/ab_range_bands.py /tmp/ab/surf --reference tfil
|
||
# incoming hit rate + swapped per-shot dodge instrument: scratch scripts in /tmp
|
||
nim c -r --path:common_libs common_libs/tests/test_tfil_commit_env.nim
|
||
```
|
||
|
||
`arms.txt`:
|
||
```
|
||
tfil |
|
||
strafe | TR_MOVEMENT=strafe
|
||
surf | TR_MOVEMENT=surf
|
||
```
|