Commit Graph

385 Commits

Author SHA1 Message Date
SirStone b0654d18eb TM verdict, settled: it loses LIVE and sits at/below its majority class - (c)
The user pushed back on "the TM can't be your best 1v1 gun", correctly, because two
decisive tests had never been run. Both are now run and they agree.

TASK 1 - THE GF HEAD vs ITS MAJORITY-CLASS BASELINE (offline, n=1,751,067):
  label histogram [254286, 284578, 678879, 297055, 236269]
  majority class = 2 (the CENTRE bucket) = 38.77%
  RAW head accuracy = 36.69%  ->  margin **-2.08 pp, BELOW majority**
  GATED head accuracy = 40.37% vs 38.75% majority -> +1.62 pp, BUT it predicts the
  majority class on 62.4% of ticks and its minority recall is 13.6% / 12.9% - a
  base-rate predictor wearing a classifier's clothes.
  Shuffled control sits at its own majority (20.04% vs 20.12%), confirming chance.
**THE OLD "46% vs 20% CHANCE" FIGURE I QUOTED WAS WRONG ON TWO COUNTS:** the
baseline is 38.8%, not 20%, and the 46% predated the deferred-label fix. Against
the correct baseline the head is BELOW it.

TASK 2 - THE FIRST-EVER LIVE A/B OF THE TM GUN (7 runs x 7 rounds per arm, one
frozen binary from git archive HEAD = eb74f9b2, sha256 cb66d66b..., real DrussGT,
every arm forced alone with TR_RACK_<GUN>=both and all 14 others off, liveness
confirmed per run):
  arm                     shots   real %   dmg/run   round wins
  onlyPattern              4610   10.74%     285      25/49
  onlyTMPATTERN (radial)   3374    3.50%      71       0/49
  onlyLinear               3218    3.23%      61       0/49
  Pattern vs TM:  +7.22 pp / +213.7 dmg, exact p=0.0006
  TM vs Linear:   +0.30 pp, p=0.659  (dmg p=0.438)
**The TM is statistically INDISTINGUISHABLE from its own Linear base live.** So it
is not "the TM works and we are aiming it wrong".

DIRECT ANSWER: **(c) It loses live AND sits at/below majority - the target carries
no learnable signal beyond the base rate, and that is the reason.** The reason is
not the machine, not the knobs, and not the application alone: the thing it was
asked to predict is dominated by the modal answer.

This closes the TM-as-gun thread. If a TM is wanted in the bot, a firing gate or a
movement decision is a better fit for a boolean-rule classifier than an aim point -
that is untested and is a different project.

A LIVE GF-MODE ARM WAS NOT RUN (stated as unmeasured): the task pinned one frozen
HEAD binary and HEAD registers the TM gun as radial only; Task 1 already makes GF
the unpromising candidate.

HARNESS FIX WORTH KEEPING: `tools/ab/which_gun_arm_env.sh` left the TARGET gun
unset, so with the now-Pattern-only default it silently fell back to the FULL rack
- an arm could appear to test a single gun while actually running the whole rack.
It now emits `TR_RACK_<GUN>=both` for the target and `=off` for all 14 others.
(Earlier which-gun results are unaffected: they ran before the Pattern-only default,
or - as in the melee/1v1 campaign - set the explicit `=both` themselves.)

tm_pattern.nim gains a per-class confusion matrix (warm samples only) to support the
majority baseline; no behaviour change. Adds Round 4 to
tm_pattern_sweep_results.md with both tasks and the interpretation rule.
2026-09-22 08:21:49 +02:00
SirStone eb74f9b2e3 Ram: finisher-only by default, and the bullet-rain abort now measures real energy
Follows the diagnosis that proactive straight-line ramming CANNOT work: both bots
have MAX_SPEED=8, so a pursuit cannot catch an evading equal-speed opponent.
Measured over 49 rounds per arm, opportunity -> contact was **0/6** (base), 0/40
(ring), 0/12 (ringhot). The only proactive conversion in the whole corpus came
from a FINISHER, and only because a <20-energy DrussGT stops fleeing (that episode
closed at 6-8 px/tick). Opportunity episodes never got below ~80px; one ran the
full 60-tick duration cap and closed only 198->171px; a perfectly aligned
full-speed one closed 195->114px then plateaued.

CHANGES
- **Finisher-only default.** `finisher` (<20 energy, dist<300, we are healthier)
  and the rare `desperation` (both <5, dist<150) are kept; `opportunity` and the
  speculative `plan` are OFF. Both are env-reenableable with no rebuild:
  `TR_RAM_OPPORTUNITY=1` (tune via TR_RAM_OPP_DIST/MARGIN) and `TR_RAM_PLAN=1`.
  Justification: it removes 100+ non-converting episodes per fixture at zero
  measured loss (oldram vs base was p=0.69, damage 279 vs 284, survival 17/49 vs
  16/49) - and each of those episodes spent up to 60 ticks driving STRAIGHT at
  the enemy, abandoning the mover's dodging and disrupting aim.
- **`desperation` KEPT** deliberately: it is cheap and rare, fires only when both
  bots are nearly dead at short range (a coin-flip where 0.6 contact can decide
  it), and it is not the refuted straight-line pursuit.
- **THE BULLET-RAIN ABORT WAS DEAD CODE AND IS NOW FIXED.** `onHitByBullet`
  accumulated raw bullet FIREPOWER while `TR_RAM_ABORT_DMG = 0.5` was documented
  as a DAMAGE rate - so the bar was implicitly "sum of power > 7.5 over 15 turns"
  and the maximum rate ever observed was 0.27. It now accumulates REAL ENERGY via
  a `bulletDamage(power)` helper matching the server's `4p` / `6p-2` formula, and
  `TR_RAM_ABORT_DMG` defaults to **2.0 energy/turn** (~30 HP over 15 turns):
  "abort an in-progress ram if we take > 2.0 energy per turn". Same effective bar
  for normal firepower, and it can now actually fire - the live run reports
  `dmgRate=1.07/turn` where the old units said 0.27.
- **`ramStuckTicks` REMOVED.** It required `dist < 5px`; contact occurs at ~36px
  (two 18px radii) and position rewind prevents getting closer, so it could never
  increment. Only the 60-tick duration cap can now self-end a ram.

LIVENESS (measured, default config, vs a charging Java RamFire, 3 rounds):
  default              -> `[ram] ON reason=finisher` x3, `reason=opportunity` x0
  TR_RAM_OPPORTUNITY=1 -> `reason=opportunity` x4, `reason=finisher` x2
So the opportunity states DID occur and are suppressed by the new default - the
removal is real, not an arm that never fires. A line also read
`[ram] OFF reason=duration dmgRate=1.07/turn`, confirming the new energy units.

Adds docs/ramming_negative_result.md (70 lines) recording the question, the five
diagnostic answers, the geometric reason, the finisher exception, the two dead
code paths, and an explicit "do not re-attempt a proactive straight-line ram; if
point-blank forcing is ever wanted it is an INTERCEPTION/cornering movement
problem" note - the same pattern that stopped the corpse bug recurring.

Guards: test_ram_decision 40 (was 28), test_gun_harness 39, test_vbullet_metric 11,
test_power_selection 3, test_adaptive_radar 41, test_tfil_ring_weights 24,
test_power_policy 26, test_rack_membership 48, test_selector_tiebreak 19,
test_tm_pattern_registration 20, test_vbullet_admit_gate 12,
acceptance_offline_vs_online 12/12. ModularBot compiles.

Honest note: the abort-threshold fix is a real (tiny) behaviour change, NOT
measured-neutral - it only bites while a finisher ram is under sustained fire,
which is exactly the user's stated wish. The finisher-only removal itself is
measured-neutral per the given A/B.
2026-09-22 08:14:03 +02:00
SirStone 99f9532f55 Melee vs 1v1 racks: mechanism built, but NO gun-level payoff - do not split
The user's plan was separate melee and 1v1 racks. The mechanism is built and
committed (TR_RACK_<GUN>=both|1v1|melee|off, mode from server truth). This job
produced the missing evidence: each candidate gun forced ALONE in BOTH modes,
one frozen binary, env-only arms, exact two-sided permutation tests.

1v1 vs the real DrussGT (6 arms x 5 runs x 7 rounds):
  Pattern      10.80%  292 dmg/run   11/35 round wins
  KNN           5.58%  128            1/35
  Linear        2.99%   59            0/35
  Circular      2.89%   60            0/35
  WallBounce    2.72%   57            0/35
  GuessFactor   2.14%   38            0/35
  -> every arm differs from Pattern at p=0.0079 (the 5v5 floor). Per-gun damage
     span 7.7x. In 1v1 the gun matters ENORMOUSLY.

Melee (+DrussGT, RandomMover, WaveSurfer, OscillatorBot; enemyCount 4):
  WallBounce   25.79%  643 dmg/run    7/25
  Linear       24.88%  624            7/25
  GuessFactor  24.48%  613            4/25
  Pattern      24.39%  599            5/25
  Circular     25.88%  596            5/25
  KNN          23.87%  539            3/25
  -> NO arm beats Pattern with significance (p=0.42-0.96, fully overlapping).
     WallBounce's nominal +7.3% is p=0.42. Per-gun damage span only 1.19x.

THE FINDING: in 1v1 the gun matters enormously (7.7x spread); in melee it barely
matters (1.19x). Melee is won by movement, survival and placement, not by which
gun you carry - every candidate lands in the same ~24-26% band. So a separate
melee rack has NO gun-level payoff and DefaultRackMembership stays Pattern-only.
The mechanism remains available if it is ever wanted.

Honest caveats: the melee comparison is UNDER-POWERED at 5 runs (detecting the
~44-damage WallBounce-Pattern gap would need ~2.5-3x the runs), so "no significant
difference" is NOT "no difference"; the only candidate worth re-testing is
WallBounce in melee, and it must not be shipped on this evidence.

SHIM LIMITATION FOUND, worth recording: run_bridge_battle.sh captures with
`--subject DrussGT`, and the shim DROPS every tick once the subject dies - losing
5-20% of ModularBot's shots in melee. The campaign was re-run with
`--subject ModularBot` so ModularBot's events are complete (fires match gun_stats
realShots exactly). Any future melee evidence through this shim must do the same
or it will silently under-count.

Adds docs/melee_vs_1v1_racks.md. No repository source changed.
2026-09-22 08:08:43 +02:00
SirStone bfdcdf8919 The three unmeasured features, A/B'd - and a methodology correction I had wrong
Five arms x 7 runs x 7 rounds (35 real-DrusGT battles, 8 concurrent), one frozen
binary from git archive HEAD at 185a32e (includes the shipped Pattern-only rack),
env knobs only, server-side event sidecar, exact two-sided permutation tests.

  arm                   real %   dmg/run   survival(rounds won)   p vs base
  base (shipped)         10.61    284.1     16/49 (32.7%)          --
  nopower (policy off)    7.88    247.1      9/49 (18.4%)          0.0012
  oldram (old gates)     10.78    279.4     17/49 (34.7%)          0.6888
  ring (tfil_ring)       20.28    267.4      6/49 (12.2%)          0.0006
  ringhot (orig heat)    11.73    285.3     15/49 (30.6%)          0.0303
Every round ends with exactly one death (0 timeouts), so survival = round win.

1. POWER POLICY HELPS - KEEP. Turning it off drops real hit rate 10.61 -> 7.88
   (p=0.0012), damage/run 284 -> 247, and wins FEWER rounds (16 -> 9). The cap
   trades per-shot damage for many more shots and a higher per-shot rate; that
   trade is a clear win. This was shipped on unit tests alone until now.

2. PROACTIVE RAMMING - INDISTINGUISHABLE. oldram vs base p=0.69, damage 279 vs
   284, survival 17/49 vs 16/49 (p=1.0), ram contacts 1 vs 2. The lever IS live
   (6 `opportunity` ON events vs 0 under the old 50px/+30 gates; base reached
   <40px on 10 ticks vs 0) but converts to essentially no extra collisions and no
   measurable outcome change. Safe to keep, but it is not earning its keep and
   reverting it is equally defensible.

3. RING MOVER HURTS THE OBJECTIVE - DO NOT SHIP. And this is the important one.

*** METHODOLOGY CORRECTION - I HAD THIS WRONG ALL NIGHT ***
The ring arm has the BEST hit rate of anything measured tonight: 20.28% vs 10.61%
(+9.67pp, p=0.0006, non-overlapping). Read alone it says "ship it immediately".
It is a CONFOUND. The ring halves engagement range (median 460 -> 240px), which
halves round length (1542 -> 636 ticks) and shots (4466 -> 1834). So damage/run is
FLAT (284 -> 267) while survival/round-win MORE THAN HALVES (16/49 -> 6/49,
p=0.0122). It is a GLASS CANNON: same damage dealt, twice as many deaths. The
hit-rate gain is a geometric artefact of fighting closer, not an improvement.
** For MOVEMENT arms, hit rate alone INVERTS the verdict. ** A movement change
alters range, shots fired and round length simultaneously, so the objective
metrics are DAMAGE/RUN and ROUND-WIN RATE (survival) - report all three. For
gun/selection arms, where range and round length are held fixed, real hit rate
remains the right ground truth. I had been enforcing the hit-rate-only rule
without qualification; it is now qualified.

HEAT TAMING is the knob that moves the tradeoff: tamed (corridor 5 / wall 10)
lets the range weighting pull to ~240px (20.28% / 6 wins); original (20/30) keeps
it at ~400px (11.73% / 15 wins). So heat taming buys hit rate at the cost of
survival - a knob to keep conservative, and the reason the ring is not the default.

Nothing to revert: the shipped movement is already `tfil` and the ring is opt-in.
Caveats recorded: only adversary is the real DrussGT jar (the shim hosts only
jk.mega.DrussGT), so the ring-vs-winning result needs re-checking elsewhere; and
base (10.61%) is consistent with the committed onlyPattern result (9.99%),
validating the frozen binary and pipeline.

Adds docs/feature_ab_results.md. No source files changed by this job.
2026-09-22 02:34:05 +02:00
SirStone 3142b70aa5 Gate virtual-bullet spawn on rack admission: +68% tick rate, selected gun unchanged
The default rack is now Pattern-only (31c7c01), but membership filters SELECTION,
not SPAWNING - so all 13 unselected guns still ran `predict` + `spawnBullets`
every tick to feed fitness tables nobody reads. Measured waste: Tsetlin alone
0.98 ms/tick, KNN 0.30, plus 10 more. This generalises the gate TMPATTERN already
had to every gun, behind `TR_VBULLET_ADMIT_ONLY` (default 1 = gate, 0 = old).

OFFLINE COST (release build, 400 ticks, min of 2 reps, 13-gun rack):
  OFF  2.14 ms/tick  (implied 467 ticks/s)
  ON   0.04 ms/tick  (implied 27667 ticks/s)
  -> reclaimed 2.10 ms/tick, ~98% of the virtual-bullet cost. Tsetlin's 0.98
     disappears, KNN's 0.30 disappears, only Pattern (0.017) survives.
Note the absolute scale is lower than an earlier unoptimised measurement (~6.1
ms/t) because this is a -d:release build; the ON-vs-OFF DELTA is the robust result.

LIVE TICK RATE (ONE frozen binary, 4 runs/arm x 4 rounds, all 8 concurrent, gate
varied by env only):
  ON   146.2 ticks/s  (142.3, 145.2, 147.9, 149.4)
  OFF   86.9 ticks/s  ( 95.0,  41.4,  99.1, 112.1)
  NO OVERLAP: ON min 142.3 > OFF max 112.1. Excluding a game-outcome outlier in
  the OFF arm, OFF max is still 112.1. Outside the noise.
**+68% tick rate**, which also makes every future A/B faster. Both arms still pay
the fixed 8192-slot ring scan in tickBullets.

SAFETY, verified not assumed: Pattern is admitted under the shipped rack, so its
own fitness keeps accumulating and the SELECTED gun is unchanged - Pattern 100% in
all 4 runs both arms, and Pattern was the ONLY gun with vShots>0 in the ON arm
while all 13 had vShots>0 in the OFF arm. So admission is the correct predicate.
Mid-round transitions are safe by construction: only predict/spawn are gated, while
tickBullets still resolves every active bullet and the feedback case still calls
the owning gun's onResult.

A TOOLING BUG THIS CAUGHT, and a correction to the task's assumption:
`acceptance_offline_vs_online.nim` IS affected (I had assumed it was not). It
compares the live per-gun vShots against an offline replay that always spawns all
guns, so under the default gate the online non-Pattern vShots are 0 while offline
is ~400 - a guaranteed mismatch. Fixed by pinning `TR_VBULLET_ADMIT_ONLY=0` inside
that test (same putEnv/defer pattern as TR_RECORD_WORLDSTATE), keeping 12/12. The
test is about offline/online METRIC parity, so it needs every gun spawning.
Unaffected (verified from source): audit_virtual_guns.nim,
measure_cornering_guns.nim, sweep_tm_pattern.nim - all offline, none read
GUN_STATS_PATH.

Guards: test_vbullet_admit_gate 12 (new, pure), test_gun_harness 39,
test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41,
test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 28,
test_rack_membership 48, test_selector_tiebreak 19, test_tm_pattern_registration 20,
test_tm_pattern_learning 3, acceptance_offline_vs_online 12/12. ModularBot compiles.
2026-09-22 02:24:18 +02:00
SirStone 185a32e9eb Radial offset: STRUCTURALLY incapable of helping, and Pattern does not overshoot
Follow-up to 9cd6e9b, which found the LINEAR base systematically overshoots (mean
radial error -71..-100px, enemy nearer in 63-81% of shots). The question was
whether the gun that actually ships, `Pattern`, overshoots too - because
correcting a systematic bias would be a cheap win.

1. PATTERN DOES NOT OVERSHOOT. Measured over the DrussGT fixtures (n=250,989):
     Pattern  mean -12.0 px, median  -3.2 px, nearer 52.8% / farther 45.0%
     Linear   mean -87.3 px, median -61.0 px, nearer 83.4% / farther 14.4%  (same states)
   So the overshoot was a property of the CONSTANT-VELOCITY BASE, not of our
   predictions in general. Pattern's pattern-matching does not have it, so there
   was nothing to correct. (All-10-fixture pooled: mean -14.0, median -4.2.)

2. THE AVENUE IS STRUCTURALLY DEAD, not merely unprofitable. The live aim is
   `aimAngle(self, pred)` and a RADIAL-only offset keeps the BEARING unchanged
   (proven exactly by a guard test: bearing is invariant). So a radial offset
   cannot change the fired bullet's direction at all. `bmPath` never scores the
   aim distance either - and measured, every offset arm is BYTE-IDENTICAL to plain
   Pattern on bmPath (33.9%/25.6%). The only real-effect channel is the `shouldFire`
   gate via `distPx`, which is indistinguishable from noise.

3. LIVE A/B CONFIRMS: one frozen binary (built from HEAD + only this change),
   env-only arms, 7 runs x 7 rounds, 8 concurrent, real DrussGT, server-side hit
   rate, exact two-sided permutation test.
     control (plain Pattern)  10.61% / 284 dmg-per-run
     s0.98                    10.89% / 302   (+0.28pp, p=0.62)
     s0.95                    10.02%         (p=0.35)
     o-20                     10.19%         (p=0.46)
   No significant winner.

VERDICT: STOP. This line cannot help the shipped configuration, and the reason is
structural rather than statistical - a radial correction is bearing-invariant, so
it is invisible to the actual shot. The bmPoint "win" the radial TM showed was a
metric artefact of that same irrelevance.

Incidental: the control arm (10.61% / 284) independently replicates the shipped
Pattern-only default's A/B numbers (10.36% / 264, 10.78% / 287, 9.99%).

Kept anyway: `TR_PATTERN_RAD_SCALE` / `TR_PATTERN_RAD_OFFSET` default to
(1.0, 0.0) and the default path is byte-identical (proven over 2400 predictions,
plus bearing invariance and unparsable-value fallback - 6 checks). Adds
measure_pattern_radial.nim, sweep_pattern_radial.nim, test_pattern_radial_offset.nim
and pattern_radial_results.md.

Guards: test_gun_harness 39, test_vbullet_metric 11, test_power_selection 3,
test_adaptive_radar 41, test_tfil_ring_weights 24, test_power_policy 26,
test_ram_decision 28, test_rack_membership 48, test_tm_pattern_registration 20,
acceptance_offline_vs_online 12/12 PASS.
2026-09-22 02:23:41 +02:00
SirStone 31c7c01d28 SHIPPED: the default rack is now Pattern-only (+49% hit rate, +66% damage on the boss)
`DefaultRackMembership` now admits Pattern (id 5) and marks all 14 other guns
`rmOff`. **The selector mechanism is untouched** - `chooseFromFit`, the floor/band
logic, the hysteresis and the virtual-fitness plumbing are all intact and
functional. Only the rack membership changed, so this is reverted by env alone.

Evidence (measured, replicated three times, 10 adversaries): Pattern alone gives
10.36% real hit rate / 264 damage per run vs the full rack's 6.93% / 159. Pattern
significantly wins on DrussGT, Corners, Crazy and PatternMover, ties on three, and
the full rack never significantly beats it on ANY adversary. Mechanism: the
virtual signal keeps ranking the wrong guns first (HeadOn 46% of ticks at 2.0%
real; Linear 57.7% at 6.0% real while Pattern sits at 11.1%).

**THIS CONTRADICTS THE USER'S STANDING DIRECTIVE** to keep virtual-fitness
selection. Recorded plainly in docs/selector_negative_value.md with a SHIPPED
DECISION banner rather than done quietly: the mechanism is retained and one env
var away, because the measurement says it is negative value on every rack size
tested and on 10/10 adversaries.

Revert one-liner (no rebuild):
  TR_RACK_PATTERN=both TR_RACK_HEADON=both TR_RACK_LINEAR=both TR_RACK_TSETLIN=both \
  TR_RACK_CIRCULAR=both TR_RACK_GUESSFACTOR=both TR_RACK_WALLBOUNCE=both \
  TR_RACK_ACCEL=both TR_RACK_STOPSHOT=both TR_RACK_DISPLACE=both TR_RACK_AVGLEAD=both \
  TR_RACK_DECAYGF=both TR_RACK_KNN=both TR_RACK_TMSELECT=both ./out/ModularBot
The unit test `testRevertOverrideRestoresFullRack` exercises exactly this table.

FLOOR PATH, verified not assumed: `chooseFromFit` already returns `admitted[0]` on
the floor path, so it respects admission by construction. Cold field + shipped
default -> floor returns Pattern (id 5), NOT HeadOn. With an explicit all-`both`
membership the same cold field returns gun 0 (HeadOn) - the old behaviour. Four
assertions in `testFloorRespectsAdmission`.

LIVENESS: one 1-round battle with NO overrides -> Pattern selected 105/105 = 100%,
every other gun 0 including TMPattern.

Honesty caveat retained in the doc: 4 of the 10 opponents were Tank Royale
sample-bot PORTS rather than the original classic jars (only DrussGT is a real
classic jar through the shim).

Guards: test_rack_membership 48 (was 38; new floor/revert/default checks),
test_tm_pattern_registration 20 (5 checks hard-coded the old default and were
updated to assert the new one, with the TMPATTERN parity proof moved onto an
explicit old-rack table), test_gun_harness 39, test_vbullet_metric 11,
test_power_selection 3, test_adaptive_radar 41, test_tfil_ring_weights 24,
test_power_policy 26, test_ram_decision 28, test_selector_tiebreak 19,
test_tm_pattern_rack_live 4, test_tm_pattern_learning 3,
acceptance_offline_vs_online 12/12 VERDICT PASS. ModularBot compiles.

FOLLOW-ON THIS EXPOSED: membership filters SELECTION but not virtual-bullet
SPAWNING, so under `onlyPattern` the 13 unselected guns still predict and spawn
every tick. Tsetlin alone is ~5.3 ms/tick (~41% of the 13.16 ms per-tick budget),
so we are still paying for it while never using it. Gating spawn on admission
would reclaim that; it was deliberately NOT done here because it would alter the
measurement protocol mid-A/B.
2026-09-22 02:10:07 +02:00
SirStone 9cd6e9b8ce Ablation: the radial TM is replaceable by a CONSTANT, and its avenue is dead on the
shipped metric

The radial TM beats Linear on bmPoint, but its head never beat the majority
baseline after the label bias was fixed - suggesting the win is a constant lean
rather than learning. So: sweep a stateless constant short-range offset (new
`common_libs/guns/radial_offset.nim`, no learning at all) against the learned TM.

VERDICT (measured, offline range, seeds=3, 18 paired runs, 231 TM rounds):
1. **bmPoint - REPLACE the TM with a constant.** `RO_s0.95` (aim distance x0.95)
   TIES it early (9/9, p=1.0) and BEATS it overall (15/3, p=0.0075; 7.47% vs
   6.89% per-run mean). A fixed -20px does the same. The head never beats its
   majority baseline (56.2% vs 57.2%).
2. **bmPath (the SHIPPED metric) - the radial avenue is a DEAD END.** TMRadial is
   a systematic LOSS there (2/16, p=0.0013); every constant is within +-0.2pp; the
   only real bmPath effect is the BotRadius clamp. So the radial shift cannot help
   the shipped configuration.
3. The per-adversary optimum DOES vary (fixed -10 for crazy, -30 for tr_crazy,
   scale 0.95 for three others) - but ONE GLOBAL CONSTANT still beats the
   adaptively-trained head, so the "fragility justifies learning" argument FAILS.

THE REAL FINDING UNDERNEATH, and it generalises beyond this gun: the base linear
prediction systematically OVERSHOOTS. Measured raw per-tick base radial error has
mean -71 to -100 px; the enemy is NEARER than the prediction in 63-81% of shots
and farther in only 4-14%, CONSISTENT ACROSS ALL SIX CAPTURES. Radial label
histogram [415166,126461,120325,44694,19351] = 57.2% majority class, mean label
-82.3 px, mean applied shift -37.6 px. So the net-short bias is a GENUINE property
of these range-holders against a constant-velocity extrapolation (they decelerate
and turn, so the true position is closer than the straight-line guess) - NOT a
fixture artefact. That is worth chasing for the guns that actually ship.

Caveat: bmPoint is not the shipped metric (bmPath won the real-hit-rate A/B for
SELECTION), so a bmPoint win is not yet evidence of a real win. That needs a live
test - and the natural target is Pattern, which is now the default and best gun.

Adds radial_offset.nim + sweep_radial_offset.nim; tm_pattern.nim gains additive
instrumentation only (radial label mean and applied-shift mean; no behaviour
change, and test_tm_pattern_registration still passes all 20 checks).
2026-09-22 02:08:38 +02:00
SirStone 589a230106 TM radial gun: registered (default OFF) + label-bias fix that removes the bias but
retracts its own earlier learning claim

=== TASK 1: REGISTERED AS GUN 14, DEFAULT `off` ===
The radial TM gun is now a first-class rack member (`TMPATTERN`, id 14), forceable
alone with `TR_RACK_TMPATTERN=both` plus every other `TR_RACK_*=off`.
DEFAULT IS `off`, and the justification matters: `both` would let it compete for
selection AND (because the shared VirtualTracker ring is order-sensitive) shift
every other gun's learning order, so it CANNOT leave the default path unchanged.
With `off` its predict and spawnBullets are additionally GATED on rack admission
(the only gun wired that way), so the shipped default never spawns it at all:
zero cost, zero ring perturbation.
Live proof: 1-round battle with only TMPATTERN racked ->
  `gun 14 (TMPattern): vShots=400 selected=104 other-gun selections=0`.
Default-path-unchanged proof: parity checks that the 15-gun default bestGun/
selectGun equals the old 14-gun rack RNG-draw-for-RNG-draw, that gun 14 is never
selected by default, and acceptance 12/12.
Cost: 0.36 ms/tick (predict 0.30 + onResult 0.05) ~= 3% of the 13.16 ms budget.
Tsetlin in the same harness is 1.62 ms/tick, so the new gun is ~4.5x cheaper.

=== TASK 2: THE LABEL-BIAS FIX - AND A RETRACTION ===
Root cause confirmed: under bmPoint a SHORT radial correction resolves the virtual
bullet BEFORE the base arrival tick, so the label was dropped (labelMisses).
Fix: defer the label in a pending queue and flush it once the arrival tick is
recorded; labels still come from the BASE arrival tick.
  labelMisses        4,281,695  ->  0
  training samples   1,071,824  ->  5,345,847  (x5)
  radial head acc         48.8% ->  57.0%   (shuffled control 20.0%)
  bmPoint hit rate     9.4/5.8% ->  9.1/5.7%  (unchanged, within noise)
So the fix IMPROVES LEARNING but NOT the metric.

**RETRACTION OF THE PREVIOUS JOB'S CLAIM.** It reported the radial head's 48.8%
against a 36.7% majority baseline and concluded "conditional learning, not a
constant bias". With the bias removed, the correctly-measured majority baseline is
**58.2%** - so the head at 57.0% is AT/BELOW majority. The earlier apparent
conditional learning was PARTLY AN ARTEFACT OF THE BIASED SAMPLE. The bmPoint
metric win is real (TMRadial > Linear early 16/2 p=0.0013, overall 18/0 p<0.0001;
> shuffled 18/0 p<0.0001) but it comes from a NET-POSITIVE AVERAGE RADIAL SHIFT,
not from beating a majority classifier. Recorded plainly rather than left standing.

Guards: test_tm_pattern_registration 20 (new), test_tm_pattern_rack_live 4 (new),
test_gun_harness 39, test_vbullet_metric 11, test_power_selection 3 (the SIGSEGV is
gone - the knn_gun rewrite is now committed), test_adaptive_radar 41,
test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 28,
test_rack_membership 38, test_selector_tiebreak 19, test_tm_pattern_learning 3,
acceptance_offline_vs_online 12/12. ModularBot compiles (release).

Note: `common_libs/tests/range_guns.nim` still builds 14 offline drivers (the
offline sweep constructs TmPatternGun directly and acceptance only inspects ids
0..13), so nothing breaks - but a future job wanting it in the offline rack must
add a 15th driver and mirror the live admission gating. gun_stats.jsonl now emits
15 rows; downstream tooling should ignore id 14.
2026-09-22 01:58:33 +02:00
SirStone 394b3deeed SETTLED: Pattern alone is the best single gun IN GENERAL, not just vs DrussGT
Closes the one-adversary caveat that blocked shipping `onlyPattern`. Two arms
(`full` vs `onlyPattern`, rack knobs only), one frozen binary from clean HEAD
(a54ae6a, sha256 04d63cd8...), 10 adversaries, 120 battles, real server-side hit
rate, exact two-sided permutation test per adversary.

  adversary      full %   onlyPattern %   diff    exact p   winner
  drussgt         6.88        9.99        +3.11   0.0023    Pattern
  corners        85.10       88.48        +3.38   0.0012    Pattern
  crazy          39.16       48.42        +9.26   0.0006    Pattern
  patternmover   59.97       67.36        +7.39   0.0159    Pattern
  spinbot        83.37       81.20        -2.17   0.4720    full (n.s.)
  ramfire        97.06       95.58        -1.48   0.2214    full (n.s.)
  randommover    46.06       43.15        -2.91   0.3968    full (n.s.)
  sittingduck    96.37       96.42        +0.04   0.9048    tie
  oscillator     65.13       67.08        +1.94   0.5238    tie
  wavesurfer     41.09       42.68        +1.59   0.5159    tie

Pattern significantly WINS on 4 adversaries, ties on 3, and the full rack's three
nominal "wins" are all NON-SIGNIFICANT and only on saturated bots (83-97% hit
rate, where any gun works). **The full rack never significantly beats Pattern
alone on any adversary.** DrussGT replicates across a different frozen binary
(full 6.88 vs 6.93 prior; onlyPattern 9.99 vs 10.36/10.78 prior).

HONEST LIMITATION on the opponents: the real classic SpinBot/Corners/Crazy/RamFire
JARS are NOT runnable through the shim - `tools/robocode_shim/BotHost.java`
hardcodes `loadClass("jk.mega.DrussGT")`, and generalising it needs source edits.
So those four are the Tank Royale SAMPLE-BOT PORTS (the same opponents the shim
README validated against classic captures), not the original jars. Only DrussGT is
a real classic jar via the shim. Stated explicitly rather than implied.

Also noted: `full` in the harness passes TR_RACK_<GUN>=off for every gun, which
yields an empty admitted set and hits the documented FULL fallback ("an empty
membership admits every gun") - i.e. the shipped all-`both` rack. Verified against
the source, not assumed. Liveness proven per arm ([rack] active=..., gun_stats
100% Pattern for onlyPattern) and owner identification validated in 120/120 battles.

Follow-up if a stricter generalisation is wanted: add a generic classic-bot host to
the shim (a source change) and re-run those four as real jars.
2026-09-22 01:49:57 +02:00
SirStone a54ae6a162 SETTLED: no small rack beats Pattern alone; the selector is negative value on a GOOD rack
Follow-up to e0666a5, which showed Pattern alone (10.78%) beats the full rack
(6.93%). That left two open questions: is a SMALL rack of good guns better than
Pattern alone, and does the selector add value on a good rack (rather than only
on the bloated one)? Both are now answered: NO and NO.

6 arms x 7 runs x 7 rounds, one frozen binary from CLEAN HEAD e0666a5 (built via
`git archive`, source verified byte-identical to the clean tree), rack knobs
only, 8 concurrent battles, real server-side hit rate vs the real DrussGT, exact
two-sided permutation test on per-run rates.

  arm          guns (selector active?)                    real %  dmg/run  p vs onlyPattern
  onlyPattern  Pattern, NO selection                       10.36    264     --
  lean8        HeadOn,Linear,Circular,Accel,Pattern,GF,KNN,WallBounce  6.31  146  0.0169
  lean6        lean8 - HeadOn                               8.83    212     0.0262
  pairPC       Pattern + Circular                           8.23    185     0.0460
  pairPK       Pattern + KNN                                9.80    264     0.3998
  pairPL       Pattern + Linear                             8.23    200     0.0035

The control replicates the prior run (10.36% vs 10.78% before; same binary tree,
different build path).

THE MECHANISM, from the per-arm selected-gun mix - the virtual signal keeps
ranking the WRONG guns first, even on a two-gun rack:
  lean8: HeadOn 46.2% of ticks at 2.0% REAL; Pattern only 14.4% (12.6% real)
  lean6: Pattern 29.2% (10.4% real) vs KNN 25.3% (8.1%) and Linear 15.2% (8.0%)
  pairPC: Circular 66.8% (6.9% real) vs Pattern 33.2% (11.2% real) - over-picks Circular
  pairPL: Linear 57.7% (6.0% real) vs Pattern 42.3% (11.1% real) - over-picks Linear
  pairPK: Pattern 86.4% - ties ONLY because the selector happens to pick Pattern
          most of the time; it is numerically lower with identical dmg/run
So the failure is NOT rack size. Pruning does not fix it; the ranking is wrong.

VERDICT: ship `onlyPattern` - Pattern alone with selection bypassed - at 10.36%
real and 264 dmg/run, vs lean8 6.31%/146 and the prior full rack 6.93%/159.
This DIRECTLY CONTRADICTS the standing user directive to keep virtual-fitness
selection, so it is recorded here plainly rather than quietly acted on: disable
the selector (`TR_RACK_<every gun but PATTERN>=off`) pending a better fitness
signal. The mechanism itself is left intact and functional so it can be re-enabled
with one env var, and so it can be fixed rather than discarded.

REMAINING CAVEAT: ONE ADVERSARY. All of this is vs DrussGT. Pattern as the default
must be re-checked against other bots first - that is the next job.

Extends the reusable harness (tools/ab/which_gun_arm_env.sh now has lean8/lean6/
pairPC/pairPK/pairPL; which_gun_analyze.py is parameterised by WHICHGUN_OUT and
compares against both `full` and `onlyPattern`).
2026-09-22 01:35:51 +02:00
SirStone 4657fe715e wave pairing: 36-58% of GF/DecayGF/KNN learning samples were MISLABELLED
The audit inferred (from code) that GF/DecayGF/KNN pop the OLDEST wave on
resolution, while under bmPath bullets leave the arena in NON-FIFO order - so an
outcome could be attached to the wrong wave. It also noted that `starved=0` does
NOT rule this out. Both halves are now MEASURED.

MISPAIRING RATE (10 DrussGT fixtures, real VirtualTracker, 344k resolutions/gun):
  gun         bmPath mispair   label err      bmPoint mispair   label err
  GuessFactor     36.48%         19.39%           18.24%          7.62%
  DecayGF         36.85%         19.52%           20.57%          8.64%
  KNN             57.91%         27.63%           29.75%         11.58%
  (starved = 0 everywhere, exactly as the audit predicted)
So ~1 in 5 GF/DecayGF learning samples and ~1 in 4 KNN samples carried a WRONG
guess-factor bin. This is a material corruption of the learning signal.

FIX: the same fireTick-keyed ring scheme `tsetlin.nim`/`tm_selector.nim` already
use - `slot = (fireTick*4 + bin) mod 1024` (period 256 ticks, longer than the
~91-tick max flight), looked up by exact key. Public interfaces unchanged; added
`waveResolved`/`waveMispaired` integrity counters. AFTER: mispaired = 0 and
starved = 0, both metrics, all three guns.

EFFECT ON HIT RATE: SMALL AND NOT SIGNIFICANT. bmPath 4000 samples/gun:
  GuessFactor 23.20% -> 23.02% (-0.18pp, per-run sign-flip p=0.750)
  DecayGF     23.80% -> 24.25% (+0.45pp, p=0.625)
  KNN         18.27% -> 18.80% (+0.53pp, p=0.547)
bmPoint: +0.05 / +0.33 / -0.15pp, p = 1.00 / 0.50 / 0.50. Per-run ranges overlap
almost completely. A bullet-level z-test is anti-conservative (bullets within a
fixture share a trajectory) and its KNN p=1.9e-16 cannot be trusted given ~10
effective independent runs.
PLAIN READING: this is a CORRECTNESS fix, not a measurable hit-rate win. It
removes a 36-58% mislabelling of the learning signal; the point estimates move by
at most ~0.5pp, within run-to-run noise. Stated plainly rather than oversold.

A REGRESSION IT CAUGHT IN ITSELF (and this explains the SIGSEGV another job saw
and correctly attributed to a concurrent knn_gun.nim rewrite): the first
implementation put an inline `array[1024, KNNWave]` (~100KB) inside each gun,
which overflowed the default 8MB stack and made `test_power_selection` SIGSEGV.
Causation was proven by stashing only the three gun files (test passed), then
fixed by making the rings heap-backed `seq`. Verified: `test_power_selection`
3 PASS on the default stack, and zero inline `array[1024]` remain.

Guards: test_wave_pairing 17 (new, pure), test_gun_harness 39,
test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41,
test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 28.
ModularBot compiles. Adds audit_wave_pairing.nim and compare_pairing.nim.
2026-09-22 01:33:31 +02:00
SirStone 1ea72c7f14 TM gun round 2: base was never behind; RADIAL target beats Linear on bmPoint
=== TASK 1: MY PREMISE WAS REFUTED ===
I instructed the job to "fix the baseline" because an earlier measurement said the
TM gun's base did not iterate flight time like `LinearGun`. MEASURED: the new gun's
base is BYTE-FOR-BYTE `LinearGun` - 18/18 runs tie exactly, p=1.000, every per-run
row byte-identical. The "non-iterating baseline" belonged to the OLD `tsetlin.nim`,
not this gun. So no fix was needed, and the earlier inference should not have been
generalised to the new gun. (It did still align the zero-correction clamp to
LinearGun's exact [0, arena] range, and reports the old BotRadius-inset base was a
wash/marginally better at 34.2%/24.7%.)

=== TASK 2: THE RADIAL TARGET - A CONTROL-VALIDATED WIN, BUT ONLY ON bmPoint ===
Instead of the lateral (GF-bucket) component - which the linear lead already
captures - the TM now predicts the RADIAL component: will the enemy be nearer or
farther than the base prediction when our bullet arrives? A 5-class radial head
sharing the same 40-bit context and TM core; the readout advances/retards the aim
distance along the base bearing.

  under bmPath (the SHIPPED metric): STRUCTURAL NO-OP
    synthetic 8/8 exact ties, p=1.0; real 33.9%/24.1% vs Linear 34.0%/24.3%
  under bmPoint: A WIN, control-validated
    TMRadial 9.4% (6013/63785) / 5.8% (42079/726652)
    Linear   7.2% / 4.7%          overall 17/1, p=0.0001
    Tsetlin  7.0% / 4.8%          overall 15/3, p=0.0075
    shuffled 7.0% / 3.6%          early 17/1 p=0.0001; overall 18/0, p<0.0001
  radial head online accuracy 48.8% vs 19.9% shuffled chance and 36.7% majority
  -> it is CONDITIONAL learning, not a constant short-range bias.
Best config: TM_RADIAL_RANGE=60, TM_RAD_MARGIN=0.25, 5 classes.

CAVEAT THAT MATTERS: a win on `bmPoint` is NOT yet evidence of a real win. `bmPath`
is the shipped SELECTION metric precisely because it beat `bmPoint` on real hit
rate (7.43% vs 4.70%). But that A/B was about which gun to PICK, not about gun
QUALITY - a gun can be better in reality while scoring worse on the selection
metric. So this needs a LIVE test, and it is the decisive one.

=== TASK 3: REVERSAL TARGET - CLEAN NEGATIVE ===
The label positive rate is only 9.7% (rev=[24772,2673]) and the head's 86.8%
accuracy is BELOW the 90.3% majority baseline: it does not learn the positive
class at all. Hit-rate effect neutral (bmPath 19.5%/18.4% vs shuffled 19.1%/17.8%,
p=0.24/0.82). Dropped.

=== OVERALL ===
Not competitive on the shipped bmPath metric (gated GF 28.3%/22.2% vs Linear
34.0%/24.3%, p=0.0075). Better than Linear on bmPoint via TMRadial (+2.2pp early,
+1.1pp overall). Per-enemy reset exists; a fresh gun per round; NO cross-battle
persistence (the user's non-negotiable).

MEASURED LIMITATION: radial mode has a high labelMiss because aiming short
resolves BEFORE the base arrival tick, biasing training toward resolvable samples.
The metric win is label-independent. A deferred-label fix is the next refinement.
INFERRED: the mechanism is surfers being NEARER than the base prediction
(range-holding); a constant-short-offset ablation would separate a learned
short-range bias from genuine per-tick conditional prediction.
2026-09-22 01:27:59 +02:00
SirStone 78975a35c4 cost: parallel per-enemy virtual bullets are NOT affordable as proposed
Benchmark driving the real rack and the real VirtualTracker over 7 recorded
DrussGT fixtures synthesised into an N-enemy melee. Answers "the virtual bullets
are cheap, why not keep fitness for every enemy in parallel?" (the user's idea,
motivated by making kill-stealing target switches free).

BASELINE: exactly 4.0 predict calls per gun per tick (one per power bin) - 52/tick
for the shipped 13-gun rack (TMSelect is compiled out). The task's 56/tick was
the 14-gun figure.

VERDICT: NOT AFFORDABLE. Budget is 13.16 ms/tick (76 ticks/s measured live).
  N=1  46% of budget
  N=2  94%          <- already at the edge
  N=4  189%
  N=6  274%
Marginal cost ~= 5.9 ms per extra target, linear.

TWO FINDINGS THE PROPOSAL MISSED:

1. `onResult` TRAINING dominates, not predict. Tsetlin's onResult alone is
   3.40 ms/tick - ~99.5% of all 13-gun onResult cost - doing ~174k rand() calls
   per resolved bullet. Every spawned bullet that resolves triggers it, so it
   scales 1:1 with targets. The per-target cost is the Tsetlin training pass.

2. `MaxBullets=8192` is a HARD BLOCKER, not just CPU. Spawn rate is 56*N/tick and
   path-metric bullets live until they hit a wall (40-90 ticks). Measured dropped
   bullets/tick: N=1 -> 0, N=2 -> ~3, N=4 -> ~180, N=6 -> ~300. At N=6 the ring
   wraps every ~24 ticks, so most bullets are silently clobbered and never scored.
   A working N=6 pipeline needs MaxBullets ~30k-50k (~4-6 MB, cheap RAM).

ALSO MEASURED: Tsetlin and KNN do NOT cache per tick - they redo the full TM
forward pass / full KNN scan for EACH of the 4 power bins (Tsetlin 1.94 ms/tick
of predict, KNN 0.45). The earlier "tick-only cache" fix never touched the two
most expensive predicts. Pattern and TMSelect do cache fully.

MITIGATIONS (measured predict+spawn at N=6 vs 13.59 ms baseline):
  nearest-K=1 only        45% budget
  nearest-K=2             95%
  rotate every 3 ticks    95%
  drop Tsetlin for extras ~68% (INFERRED from Tsetlin's measured 90% share)
Tsetlin is ~90% of the per-target cost, so excluding it from non-primary targets
makes N=6 fit. "Resolve less often" is not a separate lever - resolution IS when
training happens.

ARCHITECTURAL CAVEAT (correctness, not cost - and not priced into the proposal):
the shared-rack topology is broken for this. The guns are global singletons, so
predicting for enemy B ADVANCES/OVERWRITES enemy A's velocity tracker, KNN
feature history and Tsetlin frame window in the SAME instance. Per-enemy fitness
with correct histories therefore requires PER-ENEMY GUN INSTANCES, which is what
this benchmark measured. That multiplies the (already dominant) Tsetlin cost.

CONSEQUENCE FOR THE PLAN: combined with the measured finding that the selector is
negative value and the rack should shrink to a few good guns, this work is much
less valuable than assumed - with a small rack (Pattern's predict is 40us and
fully cached) the cost falls proportionally. Priority lowered accordingly.

Caveat: the host was heavily loaded (load 15/16), so absolute ms carry ~30-50%
noise; min-of-2 and two independent runs agree on the trend, the Tsetlin
dominance, and the ring overflow. No melee fixture exists in the repo, so the
7 enemies are 7 distinct recorded trajectories (stated in the file header).
2026-09-22 01:26:12 +02:00
SirStone e0666a562d The gun selector is NEGATIVE value: Pattern alone beats the full rack (p=0.0012)
MEASURED against the real DrussGT, one frozen binary built from clean HEAD, rack
knobs only (no source edits), 5 arms x 7 runs x 7 rounds, 8 concurrent battles,
judged ONLY on server-side real hit rate from the events sidecar, exact
two-sided permutation test on per-run rates.

  arm             runs  shots  hits  real %  dmg/run  p vs full
  full (shipped)     7   3898   270    6.93     159      --
  onlyPattern        7   4582   494   10.78     287      0.0012  <- BETTER
  onlyKNN            7   4033   207    5.13     119      0.1340
  onlyLinear         7   3215   105    3.27      65      0.0082
  onlyGF             7   3193    72    2.25      45      0.0012

Firing Pattern ALONE gives +3.85pp pooled hit rate and +80% damage per run, and
it fires MORE shots (4582 vs 3898) - it dominates on rate and volume. This is not
"any single gun wins" (full beats Linear, GF and KNN); it is specifically
"Pattern alone beats the rack".

WHY - the virtual fitness signal mis-ranks guns against real outcomes:
- HeadOn is massively over-selected: 31.4% of ticks, the most real shots (1070),
  but only 4.5% REAL. It alone drags the rack down.
- Pattern has the best virtual rank and near-best real rate (11.9%, rank 2), yet
  is selected only 22.6% of the time.
- Linear's apparent strength was SELECTION BIAS: conditional on being selected it
  looked like 15.2% (n=33), but its UNCONDITIONAL rate (onlyLinear) is 3.27%.
  Every earlier per-gun "real rate" in this repo is conditional on selection and
  is therefore confounded. This experiment is the clean measurement.

NOT YET SETTLED (do not overclaim):
- ONE ADVERSARY. All of this is vs DrussGT. Pattern must be re-checked against
  other bots before it becomes the default on this evidence alone.
- Whether a SMALL rack of good guns beats Pattern alone. The selector is negative
  value on the CURRENT bloated rack; that does not prove it is negative value on
  a rack of only good guns. That is the next experiment and it decides whether
  the selection apparatus is fixed or disabled.
- The user's standing directive is to KEEP virtual-fitness selection. This
  measurement conflicts with it, so the next step tests the selector on a small
  good rack rather than assuming either answer.

Context - three prior selection-side attempts all failed: hysteresis (7.02% ->
5.10%, p=0.002), commitment (7.17% -> 4.44%, p=0.0012), arrival-accuracy
tie-break (7.08%, p=0.88 null). The per-tick random draw is load-bearing on
three independent measurements. This experiment locates the real problem one
level up: which guns are in the rack, and that the virtual signal ranks them
wrongly.

Preserves the reusable harness (tools/ab/which_gun_run_one.sh,
which_gun_arm_env.sh, which_gun_analyze.py) and the full writeup
(docs/selector_negative_value.md).
2026-09-22 01:21:14 +02:00
SirStone 0ede6d12ec selector: arrival-accuracy tie-break measured NEGATIVE; randomness is load-bearing
Hypothesis under test (from the gun audit, which named the tie-band as "the
lever that matters most"): `bmPath` is deliberately generous (2.3-3.6x
`bmPoint`), so a gun can sit in the tied band on a ray that sweeps the target's
path while its bullets ARRIVE badly. So: keep the `path`-ranked band (path beat
point on real hit rate 7.43% vs 4.70%, z=5.56), but narrow the random draw
inside it using a parallel `point` (arrival-accuracy) window.

RESULT: NO EFFECT. Real DrussGT, ONE frozen binary (/tmp/ModularBot_tieband,
md5 2c0c56e6...), env knobs only, 7 runs x 7 rounds per arm, server-side events
sidecar, exact two-sided permutation test on per-run rates.

  arm                              runs  shots  real %  dmg/run   d      p
  tbbase (shipped)                    7   4128   7.17     175      --     --
  tbpt  path-rank + point-narrow      7   3938   7.08     165    +0.14  0.88
  tbpc  =commit control               7   3759   4.44      98    +2.74  0.0012
  tbpt25 point margin 0.25            7   3683   5.59     119    +1.65  0.20
  tbtie05 / tbtie40 (band width)      7   3937/3917  5.84/6.28  133/144  1.49/1.00  0.11/0.25
  tbwin50 (SelectorWindow=50)         7   3983   6.05     139    +1.20  0.11
  tbfloor10 (FloorPeakFrac=0.10)      7   3829   5.33     118    +2.12  0.11

tbpt vs base: fully overlapping ranges, p=0.88. This is a REAL null, not a dead
arm - the mechanism was live, and it visibly changed the selected-gun mix
(Pattern 24%->16%, Accel 6%->16%, Tsetlin ~0%->13%).

CONTROL VALIDATED, AND THIS IS THE THIRD TIME: removing the random draw inside
the band is SIGNIFICANTLY WORSE (4.44%, p=0.0012). Combined with the earlier
hysteresis A/B (7.02% -> 5.10% for commitment) and the light-hysteresis result,
the selector's per-tick randomness is now load-bearing on three independent
measurements. Narrowing the band on ANY second virtual statistic has not helped.

Every knob swept (band width, floor, window) is nominally worse than shipped at
n=7; that is "no credible win" rather than "proven harm" (sd ~1.8pp, ~1pp
resolution, underpowered).

Shipped default stays `GUN_SELECTOR_TIEBREAK=off`; the feature is opt-in, fully
guarded, and costs zero extra work on the default path (point windows are scored
only when the mode is on).

Guards: test_selector_tiebreak 19 (new, pure), test_gun_harness 39,
test_vbullet_metric 11, test_adaptive_radar 41, test_tfil_ring_weights 24,
test_power_policy 26, test_ram_decision 28, test_rack_membership 38,
acceptance_offline_vs_online 12/12 PASS (offline path calls neither
chooseFromFit nor the tie-break).

STRATEGIC CONCLUSION: three selection-side attempts have now failed (hysteresis,
commitment, point tie-break). The selector is at a local optimum and the
remaining lever is the QUALITY OF THE GUNS, not the selection among them.
2026-09-22 01:07:12 +02:00
SirStone ca82053a11 TM gun: the discrete-target diagnosis was RIGHT - it learns now. Still loses to Linear.
The user's goal: a TM gun that is the best 1v1 gun, starting from scratch every
battle but quickly overfitting the current enemy. The previous attempt (knob
tuning) failed: NO configuration beat its own shuffled-feedback control, and the
TM-off ablation scored the same as TM-on, i.e. the TM's correction was
near-zero-mean noise. Diagnosis then: a Tsetlin Machine is a CLASSIFIER, and we
were asking it for an absolute aim point - a regression target. So this attempt
gave it a DISCRETE target (multi-class over guess-factor buckets) with 40
binary/bucketed motion features, and measured it against Linear, the default
Tsetlin gun, and a MANDATORY shuffled control.

THE DIAGNOSIS IS CONFIRMED - THE TM LEARNS, DECISIVELY:
  online class accuracy     46.0%  vs shuffled control 20.0%   (2.3x chance)
  raw ungated argmax        21.2%/18.6% vs shuffled 15.2%/8.3% (18/18, p<0.0001)
  TMPattern > its shuffled control, overall   17/1 runs, p=0.0001
Compare the previous attempt, which could not beat shuffled feedback at all.
TMPattern also beats the default Tsetlin gun early (17/1, p=0.0001), so it is a
strictly better TM gun than the one in the rack.

BUT IT IS NOT COMPETITIVE WITH LINEAR ON REAL SURFERS:
  real DrussGT, bmPath (the shipped metric), 3 seeds, pooled early/overall
    Linear            34.0% (6358/18715)    24.3% (58297/239943)
    TMPattern (gated) 27.9% (15514/55535)   22.0% (158658/719681)
    TMPatternShuf     28.7%                 19.4%
  Linear > TMPattern: 15/18 early p=0.0075, 15/18 overall p=0.0075
  bmPoint: neutral (7.2%/4.6% vs Linear 7.2%/4.7%)
  synthetic controlled motion: matches/edges Linear (66.8%/60.6% vs 66.4%/59.6%,
    shuffled 55.7%/50.1%) - the mechanism works when motion is predictable.

So: the representation fix moved this from "learns nothing" to "learns strongly
but applies its knowledge badly". INFERRED reason for the residual loss: the
linear lead is already the modal GF bucket (the label histogram is centred), so
corrective excursions away from it are net-negative. The measured deficit lives
in the BASELINE and in RANGE, not in the TM knobs - which is why further knob
tuning was never going to work.

Best config: gated hard K=5, TM_CONF_MARGIN=0.25, TM_SHRINK=0.5.
NOT TRIED (time-boxed): the binary-reversal target, and a RADIAL (range-holding)
target - the latter is the top next step.

Adds `common_libs/guns/tm_pattern.nim` (NOT registered in the rack),
`common_libs/tests/sweep_tm_pattern.nim`, and a durable writeup at
`common_libs/tests/tm_pattern_sweep_results.md`.
2026-09-22 00:56:01 +02:00
SirStone a73de13458 racks: separate melee and 1v1 gun racks, plus per-mode real hit-rate data
The user's plan: "separate racks for melee and 1v1, so the bot switches from
those based on the situation, and we can put the guns we want in one or both
racks."

MECHANISM
- `RackMode` (rm1v1/rmMelee) derived from SERVER TRUTH: `rackMode(enemyCount)`
  = 1v1 when the count is 1, melee otherwise. This is the SAME `getEnemyCount()`
  value the radar already uses, so there is now ONE definition of the mode.
  (Using the tracker's known-enemy count was a previous bug in the radar: it
  read 1 before the second enemy was scanned.)
- `RackMembership` per gun: both (default) | 1v1 | melee | off.
- The selector ranks only admitted guns - including the floor path and the
  incumbent-hysteresis path.
- Empty filtered set FALLS BACK to the full rack, so the bot can never end up
  with no gun.
- Env-overridable at process start, no rebuild: `TR_RACK_<GUN>` for all 14 guns
  (TR_RACK_HEADON, TR_RACK_LINEAR, ... TR_RACK_TMSELECT), values
  both|1v1|melee|off. Empty/unknown -> both + a stderr warning, never fatal.
- `[rack] mode=<1v1|melee> active=<guns> overrides=<...>` logged once per mode
  change, never per tick.

DEFAULT IS UNCHANGED: every gun ships `rmBoth`, so behaviour is byte-identical
until the user re-racks anything. Verified by the unit test's default-config
selection parity (RNG draw for RNG draw) and by `test_gun_harness` 39 and
acceptance 12/12. `chooseFromFit` iterates the admitted list in ascending id
order, so the random tie-break draws are unchanged.

NO TUNING DONE, deliberately: we had no per-gun melee hit-rate data, and an
earlier 15-paired-run experiment found pruning neutral-to-negative on hit rate
(p=0.57/0.21). So all guns stay `both` and the membership pass waits for data.

PER-MODE DATA PLUMBING (this is what unblocks that pass): per-gun real shot
accounting is now split by the rack in force at fire time, adding to
gun_stats.jsonl: realShots1v1, realHits1v1, realHitRate1v1, realShotsMelee,
realHitsMelee, realHitRateMelee.

Verification: test_rack_membership 38/38 (new, pure, no battle); test_gun_harness
39, test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41,
test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 28;
acceptance_offline_vs_online 12/12 VERDICT PASS; ModularBot compiles. The live
`[rack]` line was observed switching 1v1 -> melee when the enemy died.

The offline range never calls the selector (only spawnBullets/tickBullets/
reportFor), so mode filtering cannot change the offline result and no offline
mode parameter was needed - confirmed by reasoning over the source and by 12/12.
2026-09-22 00:45:10 +02:00
SirStone ab86c0481f gun audit: the virtual system is sound; the rack is redundant, not broken
Audited all 14 guns offline over the committed DrussGT fixtures (~150k resolved
bullets/gun) plus 123 rounds of live gun_stats. Prompted by a GUI observation
that selected guns "fire dozens of pixels away" and a suspicion of reverse
selection.

MY HYPOTHESIS WAS WRONG. I expected guns to be ignoring `bulletSpeed`, which
would make their 4 power bins identical and the per-bin fitness pure noise.
MEASURED: only `HeadOn` is speed-blind (100% identical bins) and that is its
correct definition. Every other gun emits 91-94% DISTINCT per-bin predictions
(mean intra-tick bin spread 62-103px). The earlier "tick-only cache collapsed
all bins onto bin 0" fix is complete across the whole rack.

LEAD/SIGN/UNITS ARE CORRECT: replaying synthetic ground truth, all 14 guns score
100% on a stationary target (which also proves predictions are ABSOLUTE - a
relative or angle return would score 0), ~100% on constant-velocity for every
leaded gun, Circular 100% / Accel 99.8% on a 3deg/tick circle, WallBounce 98.2%
on a bounce. No missing lead, no sign inversion. Resolution is right (BotRadius
18, hit credited to the owning gun).

FEEDBACK IS INTACT: offline pushes=151260/starved=0, Tsetlin trained=149205/
traceMisses=0; live `vStarved=0` and `vDropped=0` across all 123 rounds.

THE "43% FLAT / 4x OPTIMISTIC" EVIDENCE I CITED IS NOT REPRODUCIBLE on current
code/data. Live virtual/real ratios against DrussGT are 0.8-2.0 for most guns
(Linear 10.9 virt / 13.4 real; KNN 8.1/7.1; DecayGF 10.5/9.8). The 43%-flat
session matches an older config or a weak opponent (SittingDuck), not DrussGT.
`bmPath` IS 2.3-3.6x `bmPoint` - but by design and documented: it asks "does the
ray eventually sweep the target's path", a deliberately generous relative
signal. So the flat tie is a RANKING artefact: many guns share the same base
forecast and, with a near-zero learned correction, collapse onto the same ray;
RelTieMargin=0.20 then treats the top ~half of the rack as tied.

DUTY AND OVERLAP (>=50% of ticks within 20px = redundant):
  Tsetlin   ~ StopShot 87%            -> duplicate pair
  DecayGF   ~ GuessFactor 92%         -> duplicate pair
  Accel     ~ Circular 65%            -> partial duplicate
  WallBounce~ Linear 58%
  AvgLead   = the MEAN of Linear+Circular+WallBounce (constructed redundancy)
  Displace  worst point% (6.6) AND worst real% (2.4); wins no bucket
  TMSelect  DEAD - never spawned (EnableTmSelector=false), 0 shots in every log
  Pattern   the ONLY gun competitive in every distance/speed bucket
  HeadOn/Linear/Tsetlin/StopShot are identical copies of each other on a real
  surfer (v<1 ~50.8%, everything else ~2%)

RECOMMENDED LEAN RACK (8): HeadOn, Linear, Circular, Accel, Pattern,
GuessFactor, KNN, WallBounce.
DROP (6): TMSelect (dead), AvgLead (constructed mean), Displace (worst), DecayGF
(92% GF), StopShot (87% Tsetlin), Tsetlin (the repo's own sweep already showed
it learns nothing on DrussGT).

HONEST HEADLINE: pruning is NOT expected to raise hit rate - an earlier
15-paired-run experiment found it neutral-to-negative (p=0.57/0.21). The
mechanism by which it could help is a SELECTOR effect (shrinking the tied band),
not a gun effect, and that is UNVERIFIED until A/B'd. The virtual system and the
rack are basically sound; the lever that matters most is the selector's
metric/tie-band, not deleting guns.

DESIGN SMELL FOUND (INFERRED, not measured): GF/DecayGF/KNN `onResult` pops the
OLDEST wave, but under bmPath bullets leave the arena in non-FIFO order, so a
resolution can be paired with a neighbouring tick's wave. starved=0 does not
rule this out. Candidate fix: key waves by fireTick, as Tsetlin/TMSelect do.

Adds common_libs/tests/audit_virtual_guns.nim (offline, no shipped file touched).
2026-09-22 00:44:19 +02:00
SirStone 994f88d7a7 ramming: make the decision PROACTIVE, with a bullet-rain abort
The user watched 1v1 and melee runs and saw ram opportunities arise that the bot
declined: "there were moments where the bot could jump over the enemy and shred
it but shot it down instead."

DIAGNOSIS - a chicken-and-egg loop. `ramOpportunity` required dist < 50px, but
an offline measurement over 15 rounds vs DrussGT found the closest approach was
118.7px and the <50px trigger had NEVER fired: the mover has no reason to close,
so the trigger waited for a proximity nothing created. The MECHANISM to close
already existed (the ring mover expresses a ram as band=(0,50)); what was
missing was a decision that fires at a range the bot can actually close from.

Changes:
- opportunity gate relaxed: dist 50 -> TR_RAM_OPP_DIST (200), energy margin
  +30 -> TR_RAM_OPP_MARGIN (15). Both env-tunable, no rebuild needed.
- New pure module `common_libs/movements/ram_decision.nim` holding the trigger
  and abort logic (no battle/API deps), so it is unit-testable.
- BULLET-RAIN ABORT (the user asked for this earlier): `onHitByBullet` now
  accumulates `e.bullet.power` into a 15-turn ring; damageRatePerTurn = sum/15;
  an in-progress ram aborts when rate > TR_RAM_ABORT_DMG (0.5/turn). On abort:
  isRamming=false, cooldown 30, TARGET KEPT, and the mover returns to the normal
  range band. It never stops the bot.
- Opt-in, DEFAULT-OFF `plan` trigger for "change of plan when the gun duel is
  failing" (dist<250, margin+20, selected gun's pooled virtual rate < 0.05).
  Left off because a cold gun reads 0.0 and would qualify - speculative.
- `TR_RAM_LOG=1` change-gated line: `[ram] ON reason=opportunity dist=143
  selfE=78 enemyE=41 cap=3.0 band=[0,50]` / `[ram] OFF reason=bulletRain`.
- Existing cooldown/duration/stuck machinery untouched (stuck>10 or duration>60
  -> abort + cooldown 30). A refactor bug that briefly DROPPED the
  `ramCooldownTicks == 0` gate was caught and fixed.

Trigger set (first match wins): finisher (dist<300, enemy<20, we are healthier);
opportunity (dist<200, we lead by 15+); desperation (both <5, dist<150); plan
(off). All require enemy>0, a valid target, and no cooldown.

Proof the intent now fires at a closable distance (28/28 unit checks):
  PASS: opportunity fires at dist 143 with a 37-energy lead   <- the exact case
  PASS: old gate (dist<50, margin+30) does NOT fire at 143    <- the old bug
  PASS: fires at 199px / does NOT fire at 201px
  + margin, finisher priority, desperation, plan on/off, window mean, abort
    threshold checks.

HONEST FRAMING: ram damage is 0.6 per CONTACT EVENT, one-shot (collision
resolution rewinds positions so contacts do not stream) - small next to a p=3.0
bullet hit (16). The payoff is that point-blank forces hit probability toward 1,
so heavy bullets stop missing and E[dE]=p(3P-1) turns positive above P=1/3; ram
damage also scores 2.0/point (highest in the game) and a ram kill carries a 0.30
bonus vs 0.20. So this is "force the fight to point-blank", not "the ram shreds
them". Base rate is rare (2 collisions in the whole fixture corpus).

Guards: test_ram_decision 28 (new), test_power_policy 26, test_gun_harness 39,
test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41,
test_tfil_ring_weights 24. Compiles (release).
UNVERIFIED: the live effect. No A/B has run, and whether the mover actually
reaches contact is unproven.
2026-09-22 00:24:45 +02:00
SirStone c9825dfb0b power policy: cap power by range and energy, gate 3.0 on above-average chances
Implements the user's energy management request: "firing from more than 200px
should be a 'not good chances zone' so faster bullets and more chances to hit
matters more than single hit damage with low chances. When we are lower than 50
health, same thing. I would like to use 3.0 power only when the chances of
hitting are higher than average."

Design: a CAP on top of the existing `bestPower`, not a rewrite. `bestPower`
still answers "which bin does this gun's own data prefer"; the policy caps it:

  ramming                                       -> 3.0  (reason ram, exempt)
  dist > TR_POWER_FAR_DIST (200)                -> 1.0  (far)
  elif selfEnergy < TR_POWER_LOW_ENERGY (50)    -> 1.0  (lowEnergy)
  elif pEst <= pRef                             -> 2.0  (belowAvg)
  else                                          -> 3.0  (full)
  power = min(gunPreferredBinPower, cap)   # can only LOWER power

p=1.0 is the right "low" tier on the measured mechanics: bullet speed 20-3p so
p=1.0 gives speed 17 vs 11 at p=3.0 (55% faster = less lead error), fire
interval 10+2p so 12 ticks vs 16 (33% more shots), and drain 0.083/turn vs
0.1875 (2.25x slower). All three things the user asked for at long range.
pEst = the chosen bin's virtual rate (gun aggregate when the bin is empty);
pRef = the gun's aggregate mean unless TR_POWER_REF > 0. No-data guns are
vacuously below-average -> cap 2.0 (conservative, documented).

Control arm: TR_POWER_POLICY=0 = uncapped = today's behaviour exactly.
Knobs: TR_POWER_POLICY, TR_POWER_FAR_DIST, TR_POWER_LOW_ENERGY,
TR_POWER_FAR_CAP, TR_POWER_MID_CAP, TR_POWER_REF, TR_POWER_LOG.
TR_POWER_MID_CAP exists because the user did not specify the middle case
(close + healthy + not-above-average); 2.0 is the default, flippable to 1.0.

Seam: the cap lives in a pure `applyPowerPolicy` and is applied only in
`selectShot` (the single place real shots are chosen), so the logic is testable
without a battle. Ram is wired from `shouldRam` - the same value the movement
dispatch uses for the (0,50) band.

CORRECTION TO AN ASSUMPTION IN THE TASK: `offline_range.nim` does NOT call
`bestPower`/`selectShot` - it only replays virtual-bullet spawn/resolve across
all power bins, independent of the real shot's power. So there is no offline
power-selection path that could diverge from the live one, and the acceptance
test guards the metric, not the policy. Policy coverage therefore comes from the
new unit test.

Verification: test_power_policy 26/26 in BOTH modes (default and TR_POWER_POLICY=0
control arm); test_gun_harness 39, test_vbullet_metric 11, test_power_selection 3,
test_adaptive_radar 41, test_tfil_ring_weights 24; acceptance_offline_vs_online
12/12 VERDICT PASS (live battle). ModularBot compiles.

UNVERIFIED: the live effect on damage/survival/score. No A/B has run.
2026-09-21 23:59:56 +02:00
SirStone 9caf1d3728 movement: range-weighted TFIL variant + tamed heat field (opt-in, default unchanged)
New mover `the_floor_is_lava_ring.nim`, a COPY of `the_floor_is_lava.nim` (which
stays byte-identical - the user explicitly wants the current TFIL preserved).
Selected only via `TR_MOVEMENT=tfil_ring`; the default stays `tfil`.

WHY: our measured real hit rate vs DrussGT is strongly range-dependent - 21.6%
at 0-100px, 27.1% at 100-200px, 19.3% at 200-300, 10.9% at 300-400, 6.8% at
400-600, 5.4% at 600-800 - but we shoot from ~450px on average. Plain TFIL has
no range preference at all.

THE ONE CHANGE: the final tile draw is re-weighted toward a target band.
  rangeW(d) = 1.0 if lo<=d<=hi; exp(-((lo-d)/K)^2) if d<lo; exp(-((d-hi)/K)^2) if d>hi
  w_i = rangeW(d_i)^(1/T);  chosen ~ Categorical(w)
FLAT TOP on purpose: a Gaussian centred on the band midpoint would collapse the
band to a point and destroy the within-band hedge. `T` is the only knob;
`TR_TFIL_RANGE_TEMP=0` gives plain `rand(candidates.high)` - the exact control
arm. Safety stays a HARD constraint: the weighting only reorders the draw among
the pool the old code already accepted, so it can never pick a tile the old code
rejected (monotone refinement). Small pools (<4) stay uniform.
Randomness is deliberately KEPT: a measured A/B showed committing to the "best"
tile made real hit rate WORSE (7.02% -> 5.10%), so the distribution is tilted,
never removed.

HEAT TAMING (ring copy only; env-overridable):
  TR_TFIL_CORRIDOR_HEAT  20.0 -> 5.0
  TR_TFIL_WALL_HOTNESS   30.0 -> 10.0
Rationale, measured: `CorridorHeat=20` is TWICE `PathDangerThreshold=10`, so a
single corridor could poison a path by itself; `WallHotness=30` with
`WallRadiance=10` put the outer two tile rings over threshold on their own.
Per-source shares of total lava: wall 60.6%, corridor 25.4%, pillar 7.4%,
everything else <3%.

MEASURED EFFECT (primary fixture, 20,026 ticks / 15 rounds, field identity
verified max diff 0.000e+00):
  metric                        original(20/30)   ring(5/10)
  band-weightable ticks              10.79%         26.45%
  mean safeTiles/tick                 15.19          85.04
  ticks with 0 safe (pre-fallback)    58.5%           7.5%
  safePool >= 4                       39.05%         92.47%
  >=1 safe tile in 100-200px          11.84%         26.64%
  MEAN CLOSEST-SAFE-TILE DISTANCE    397.78px       284.84px
  tiles > 10 threshold                 0.61           0.15
The 397.78px figure is why the bot stayed far away: the safety filter left
nothing safe near the target, and 397px is our WORST range. Control: setting
corridor=20 wall=30 reproduces the original baseline exactly.
CEILING, honestly: even at corridor 0 / wall 0 only ~40% of ticks are
band-weightable, so no constant tweak fully unlocks the range weighting.

Ram unification: the ring mover takes a `band` field; ramming becomes just
`band=(0,50)`, so there is one movement engine. The `tfil` path is unchanged.

Observability: magenta annulus at the band edges, candidates tinted by weight,
chosen tile marked; one `[tfil_ring]` log line on change (now including
corridorHeat/wallHotness).

Guards: test_tfil_ring_weights 24/24 (new, pure, no battle), test_gun_harness 39,
test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41.
UNVERIFIED: the mover's live effect. It has not been run in a battle yet.
2026-09-21 23:46:47 +02:00
SirStone 2daa519e15 TFIL heat field: the safety model keeps us ~400px away, which is our worst range
Offline diagnostic driving the REAL TFILModule.computeMove over the committed
DrussGT fixtures (re-derived field matched the module's own m.lava bit-for-bit,
max diff 0.000e+00). Answers "is the heat map too hot, and are the corridors to
blame?" - the user's suspicion after watching a GUI run stay far away.

PRIMARY FIXTURE (tr_drussgt_vs_modularbot, 20,026 ticks / 15 rounds):

Field saturation
  tiles == 0                 22%
  tiles > 0                  78%
  tiles > PathDangerThreshold(10)   61%   (worst tick 91%)
  median / p90 / max lava    20.06 / 44.53 / 77.51
  > 10 with NO bullets at all      44%   <- wall radiance + pillar alone
  early/mid/late frac > 10   0.61 / 0.63 / 0.60  (saturated from tick 0, not degrading)

Safe pool - THIS IS THE KEY NUMBER
  inside-hull tiles/tick     173.7
  safeTiles/tick              15.19
  ticks with ZERO tile passing the filter   58.5%   (2-tile promote fallback used 59.0%)
  ticks where the ring weighting is enabled (pool >= MinRingPool=4)   39.1%
  ticks with >=1 safe tile in the 100-200px band   11.84%
  MEAN DISTANCE TO THE CLOSEST SAFE TILE   397.8 px
  ticks both pool>=4 AND band present ("band-weightable")   10.79%

So the safety filter leaves nothing safe near the target: the closest safe tile
averages 398px away. Our measured hit rate is 27.1% at 100-200px and ~5% at
450px, so TFIL's danger model structurally parks us at our worst range. This -
not only the env-var issue - is why the bot stays far away.

Per-source attribution (share of total lava / of the over-10 set)
  wall        60.56% / 56.39%   <- saturates the RAW field
  corridor    25.43% / 21.35%   <- blocks the BAND
  pillar       7.44% /  5.06%
  bullet_aura  2.22% /  1.47%
  enemy_core   1.97% /  0.62%
  bullet_core  1.17% /  0.33%
  enemy_aura   1.21% /  0.95%
Note CorridorHeat=20 is TWICE PathDangerThreshold=10, so a single corridor can
poison a path on its own; WallHotness=30 with WallRadiance=10 puts the outer two
tile rings at/over threshold by themselves (38.6% of all tiles).

Counterfactuals (shipped constants NOT changed) - band-weightable ticks
  corridor 20 (shipped)   10.79%   band-safe 11.84%   pool 15.19
  corridor 10             16.47%                     pool 30.54
  corridor  5             23.80%   band-safe 24.21%   pool 43.93
  corridor  0             38.74%                     pool 62.14
  wall 30->10 only        pool 15.19 -> 24.80, band unchanged (12.31%)
  corridor 5 + wall 10    26.45%   band-safe 26.64%   pool 85.04

Reachability - NOT the blocker
  band inside the 50-tick reachable hull   64.94% of ticks
  0-300px inside hull                      90.34%
So the band is reachable 65% of the time but SAFE only 12%: the 53-point gap is
heat, not hull geometry.

VERDICT: heat saturation is the real blocker; the WALL is the largest raw-heat
source but the CORRIDORS are the band blocker (removing them multiplies
band-weightable ticks 3.6x, while taming walls leaves the band unchanged).
Even at corridor=0/wall=0 the band is weightable only 40% of ticks, so no
constant tweak fully unlocks the range weighting - the safe set against DrussGT
rarely reaches 100-200px at all. Recommended (NOT applied): CorridorHeat 20->5
and WallHotness 30->10, to be validated by a live A/B.

Adds common_libs/tests/measure_tfil_heat_field.nim (offline, no shipped file
touched; both movers byte-identical).
2026-09-21 23:39:14 +02:00
SirStone 391318a7bd cornering/ramming premise REFUTED on three independent measurements
Hypothesis (user's): pushing an enemy toward a wall makes it predictable, which
both enables a ram and raises our gun hit rate. Measured offline over the
committed DrussGT fixtures using REAL server event attribution (an events
sidecar survived: tools/robocode_shim/evidence/tr_drussgt_vs_modularbot.events.json,
2534 fires / 229 hits; per-round event tick joins the fixture global tick at
global = round.startTick + tick - 2, verified exact over all 2534 fires).

M1 - cornering does NOT raise hit rate.
  REAL attribution, all 1134 ModularBot shots, bucketed by DrussGT's distance
  to the nearest wall at fire time:
    <=30px   69 shots   5 hits   7.25%
    30-60   336        18        5.36%
    60-120  555        26        4.68%
    120-250 172        11        6.40%
    >250      2         0        0.00%
    TOTAL  1134        60        5.29%
  Adjacent(<=60) 5.68% vs Open(>60) 5.08%, z=+0.435 -> NOT significant.
  Per-round ranges fully overlap (adjacent 0-18.2%, open 0-9.6%).
  Virtual per-gun within-gun check: most guns neutral-to-negative; only DecayGF
  favours it. Across all 10 fixtures every one of 12 guns scores LOWER adjacent
  (range-confounded, directional only).

M2 - a wall-adjacent enemy is LESS predictable, not more.
  30-degree tolerance, uniform chance 16.7%, adjacent vs open:
    keep-direction (1 tick)      93.35% vs 95.88%   z=-13.82
    turn-persistence             87.6%  vs 90.9%
    constant-velocity err H=10   45.9%  vs 25.7%    (1.8x MORE deviation)
    "move away from nearest wall" 1.1%  vs 8.4%
    "move toward centre"          0.4%  vs 3.0%
  Wall-adjacent DrussGT reverses more and deviates from constant-velocity ~1.8x
  more. It does NOT flee the wall - it surfs perpendicular. Base rate of
  wall-adjacency: 20.2% of moving ticks.

M3 - the ram is a near-zero-frequency opportunity against DrussGT.
  Strict contact (<=36px): ZERO ticks in all 10 fixtures. Closest global
  approach 39.1px. In the primary fixture (ModularBot vs DrussGT) the closest
  approach was 118.7px - 0 ticks <=80px, 0 near-contact episodes, and ZERO ram
  collisions in 15 rounds. Real ram collisions anywhere in the corpus: 2 total
  (drussgt_vs_ramfire 1/20 rounds, tr_drussgt_vs_crazy 1/10), each a ONE-SHOT
  0.6 energy to both bots, no sustained multi-tick stream.

CORRECTION TO AN EARLIER CLAIM: ram damage is 0.6 per CONTACT EVENT, not
0.6/turn sustained. The efficiency ratio (0.6 damage for 0.6 energy taken,
scored 2.0/pt) still beats firing, but the magnitude is 0.6 vs 16 for a p=3
bullet hit, and against DrussGT the frequency is zero.

CAVEAT (from the analysis): the fixtures capture DrussGT's NATURAL wall
behaviour, not an enemy being actively pushed into a corner by a rammer, so the
exact scenario is not directly represented. But M3 shows we never get close
enough to push in the first place - ModularBot's closest approach in 15 rounds
was 118.7px, so the <50px ram trigger has never fired against this adversary.

Adds two reusable offline instruments:
- measure_cornering_guns.nim (replays a fixture through the real VirtualTracker,
  attributing each resolved virtual bullet to its fire-tick wall bucket)
- measure_cornering_ram.py (real-event join, predictability, ram base rate)
Neither edits offline_range.nim; the 12/12 deterministic-gun contract is
untouched and was not re-run (it requires a live battle).
2026-09-21 22:49:56 +02:00
SirStone 07f6f3af3f Tsetlin gun: NO configuration adapts faster than random feedback
The user's goal was "a TM gun that can learn fast and generalize better".
Swept offline over the real DrussGT fixtures (no live battles) by coordinate
descent, one lever at a time, with a SHUFFLED-FEEDBACK CONTROL - a TM trained
on randomised targets. That control is what settles the question.

Final confirmation, 4 seeds each (~74,600 first-100-tick bullets per config):

  config                          EARLY(first 100)   OVERALL
  Shuf_w3  (RANDOM feedback)          23.9%           20.0%
  win3_s1.1 (best real TM found)      23.7%           20.2%
  Shuf_w10 (RANDOM feedback)          23.1%           20.0%
  win3_st100 (prior job's edit)       23.0%           20.1%
  win3_off (TM correction ~= 0)       22.7%           20.2%
  def_w10  (shipped default)          22.1%           20.3%
  Linear (deterministic reference)    34.0%           24.3%

The best real config beats the default early (23.7% vs 22.1%, non-overlapping
per-seed ranges, z=+7.34, p=2e-13) - but its own SHUFFLED control scores 23.9%,
i.e. HIGHER, z=-0.91, p=0.37. Random targets do at least as well. So the early
gain is not learning.

Per-lever screens were flat: TM_N_CLAUSES 25/50/100/200 all 23.0% early,
completely flat; TM_N_STATES 4/32/100 all ~22-23% (unstable across seeds);
TM_S mildly monotonic (lower better early); TM_T flat; TM_WINDOW_SIZE 2/3/10
all within noise of each other and of the shuffled control.

Two further findings:
- The TM-off ablation (correction ~= 0) scores 22.7%/20.2%, essentially the
  same as TM-on. The TM's correction is near-zero-mean noise; the gun's
  one-shot internal linear baseline accounts for its accuracy.
- The TM gun is 10.3 pp behind Linear early and 4.1 pp behind overall. That
  deficit is in the BASELINE MODEL (LinearGun iterates flight time; this gun
  does not), not in the TM hyper-parameters. Tuning knobs cannot close it.

Conclusion: do not tune TM hyper-parameters further. Either the input
representation or the prediction target is what needs to change - the shuffled
control shows the TM is not extracting target information beyond its baseline.

Defaults left UNCHANGED (window=10/states=32/S=1.5/T=25/clauses=50); an
uncommitted prior edit (window=3/states=100) was reverted as unsupported.
Hyper-parameters are now compile-time overridable (-d:TM_WINDOW_SIZE=3 etc.)
so future sweeps need no gun edit.

NOT MEASURED: real hit rate vs DrussGT (offline only by design). The repo's own
docs/gun_rack_analysis.md 2 reports the virtual metric is a sign-unstable ranker
of real hit rate, so the comparison against "Linear 10.7% real" is not direct -
whether the TM is competitive live is INFERRED-unknown, not measured.

Guards: test_gun_harness 39/39, test_vbullet_metric, test_power_selection,
test_tsetlin_gun, test_tm_pattern_learning all green.
2026-09-21 22:45:28 +02:00
SirStone fb36a0a685 tracker: corpses do not exist - revert the fix and retire the workaround
The belief "BotDeathEvent never reaches ModularBot, so enemyTracker keeps dead
enemies alive forever" was written into a code comment and then believed twice.
It is FALSE. Measured in a 7-bot melee with a per-tick probe comparing
enemyTracker's alive count against the server's getEnemyCount():

  metric                          1.3.1 (20 rd)   0.35.5 (15 rd)
  observed enemy deaths                83              68
  ...non-round-ending              83 (100%)       66 (97%)
  ekBotDeath events DROPPED             0               0
  max dispatch lag (turns behind)       1               1
  phantom ticks                  1 / 16,820      1 / 12,596
  MAX CORPSE LIFETIME               0 ticks         0 ticks
  victims still alive at round end      0               0

onBotDeath fires for every death, including non-round-ending ones. The
API-level event-drop mechanism IS real (test_event_drop_mechanism.nim proves
it: ekBotDeath is not in isCritical and MAX_EVENTS_AGE=2) - the bot simply
never falls far enough behind for it to trigger (max lag 1 turn).

Removed:
- reconcileWithServer + ReconcilePersistTicks/mismatchTicks/sawServerAlive
  (uncommitted, and ON BY DEFAULT despite the premise being false). Its own
  comment admitted a shorter window once KILLED A LIVE ENEMY ("it fired three
  more times after the tracker marked it dead") - a latent mis-prune path
  defending against a bug that does not exist.
- The radar's CorpseTicks=40 filter and the same-class age>60 filter in
  recordRadarStats, both carrying the false comment. Removal changes no real
  behaviour: buildState feeds the radar enemyTracker.allAlive(), so a dead
  enemy never reaches computeScan.

Kept:
- The TR_TRACKER_PROBE instrument (default OFF), which produced the table above.
- test_event_drop_mechanism.nim - the drop mechanism is a genuine library
  behaviour worth guarding.
- isAlive/aliveCount on the tracker.

Added: docs/tracker_death_events.md (the durable negative, so this is not
re-invented a third time) and test_enemy_tracker_death.nim (13 checks) in place
of the test for the deleted feature.

Guards: test_gun_harness 39/39, test_vbullet_metric 11, test_power_selection 3,
test_adaptive_radar 41/41, test_event_drop_mechanism 6, test_enemy_tracker_death
13, acceptance 12/12, ModularBot compiles.
2026-09-21 22:41:11 +02:00
SirStone c091bf3c34 harness: upgrade to server 1.3.1, keep 0.35.5 selectable, re-baseline
All prior measurements ran on server 0.35.5. The default is now the current
1.3.1 jar, with the legacy jar kept and switchable via TR_SERVER_JAR (no code
edit). test_gauntlet_5bots.nim no longer clobbers a caller's TR_SERVER_JAR -
it used to putEnv() unconditionally, so an override was silently ignored.

RE-BASELINE (controlled RulesProbe battle, stationary bot, powers 0.1/0.5/1/2/3):

  dimension                    1.3.1              0.35.5            verdict
  bullet damage per hit        0.4/2/4/10/16       identical         SAME
  bullet speed (20-3p)         within noise        within noise      SAME
  post-fire gun heat (1+p/5)   identical           identical         SAME
  cooling                      0.1/tick            0.1/tick          SAME
  bulletDamage SCORE           exactly 100/round   104..113/round    DIFFERENT
  bulletKillBonus (20%)        20/round            20..23/round      DIFFERENT

LOUD FINDING - a SCORING rule changed, physics did not: 0.35.5 credits
OVERKILL to bulletDamage (the killing bullet's full damage even past 0 energy);
1.3.1 caps it at the energy actually removed. Every 0.35.5 score is therefore
inflated ~5-6%, and bulletKillBonus inherits the inflation. Gauntlet totals
shift accordingly (SittingDuck 1936 -> 1800, WaveSurfer 1886 -> 1669).

Consequence: score-based numbers recorded on 0.35.5 are NOT comparable to 1.3.1.
Our gun A/Bs used real HIT RATE, not score, so those conclusions stand.

Runner 1.0.2 (unchanged, no newer one on the box) is measured compatible with
the 1.3.1 server. Note TrBattleCapture uses the runner's EMBEDDED server, which
is 1.0.2 - so the capture path still runs an older engine than the gauntlet.

Also re-ran acceptance_offline_vs_online on the new default: 12/12.
2026-09-21 22:41:06 +02:00
SirStone 18f778056b gun selector: hysteresis measured NEGATIVE, shipped at the lightest setting
Hypothesis under test: the selector chatters (~54 switches/100 ticks) and that
chatter suppresses firing, so committing to the virtual-best gun should raise
real hit rate. MEASURED AGAINST THE REAL DRUSSGT: it does not.

  setting              switches/100t   real hit %   dmg/run   shots/run
  no hysteresis 0/0         54.29        7.02%        217       244.8
  light 10/0.05              1.95        6.22%        191       235.3
  moderate 30/0.15           1.33        5.10%        156       233.9
  aggressive 60/0.30           -         5.72%        175       242.5
(16 runs x 7 rounds per config except aggressive = 8; server-side events
sidecar; permutation test baseline-vs-moderate p=0.002, baseline-vs-light
p=0.18.)

Hysteresis cuts chatter 28-54x but every variant fires slightly FEWER shots and
deals LESS damage than baseline. Mechanism [INFERRED, consistent with
docs/gun_rack_analysis.md 2/4]: the per-tick random tie-break among the tied
band is a hedge, and hysteresis destroys it by committing to the virtual-best
gun - which is not the real-best, because the virtual metric is a weak,
sign-unstable ranker. The chattering was load-bearing.

Shipped: GunDwellTicks=10, GunSwitchMargin=0.05 (GUN_SELECTOR_DWELL /
GUN_SELECTOR_MARGIN) - the only setting within the baseline's run-to-run spread.
GUN_SELECTOR_DWELL=0 GUN_SELECTOR_MARGIN=0 reproduces the pre-change selector
exactly.

Seam: VirtualTracker, which already owns the other selection state (fitness,
the relative floor's peakRateRef), so the bot needs no new fields. bestGun and
chooseFromFit stay pure/memoryless, which is why the existing random-tiebreak
test needed no change.

Guards: test_gun_harness 39/39 (33 original + 6 new hysteresis checks),
test_vbullet_metric, test_power_selection, acceptance_offline_vs_online 12/12.
2026-09-21 22:41:01 +02:00
SirStone 68e0375be2 feat(radar): adaptive melee radar sweeps only the arc containing all enemies
Replaces melee_scan in the rack. melee_scan spun the radar at the 45 deg/tick
cap unconditionally, so a full 360 deg revolution took 8 ticks and every enemy
was scanned roughly every 8 ticks. The new module starts with the same full
spin, and once it is SURE it has covered every enemy it sweeps back and forth
over only the minimal covering arc of all enemy bearings.

MEASURED, real melees via the bridge, per-enemy onScannedBot counts:
  3-bot melee (2 enemies): 29.6 -> 61.1 scans/100 melee-ticks  (2.06x)
  4-bot melee (3 enemies): 36.0 -> 75.7 scans/100 melee-ticks  (2.10x)
Covering-arc widths observed: mostly <90 deg in the 2-enemy case, up to 240 deg
in the 3-enemy case, so the gain shrinks as the arc widens - and at the
ExitTrackWidthDeg=300 fallback it degenerates to exactly the old full spin, so
there is no loss when narrowing would not help.

TRADEOFF, recorded rather than hidden: a wider arc legitimately takes longer to
traverse, so the freshness window costs 5-7 points (fresh<=16: 93-95% vs
98-100%) and more at fresh<=8 (75-76% vs 97-100%). More scans per enemy, at
slightly staler individual fixes.

DESIGN: acquisition spins 360 until every live known enemy was seen within
FreshnessTicks=16 (two revolutions of slack), no new id appeared, and the live
count matches getEnemyCount(); that must hold FreshStreakTicks=3 consecutive
ticks. Tracking then bang-bang sweeps the wraparound-aware covering arc
(350+10 -> 20 through 0) widened by MarginDeg=20 each end, at up to 45 deg/tick.
Fallbacks return to acquisition: any stale enemy, any new id, or an arc >= 300
deg. Enter 270 / Exit 300 gives 30 deg of hysteresis so it cannot flap.

Adds EnemyInfo.lastSeenTick (additive) so coverage is judged on staleness, not
mere knowledge - without it an enemy that slipped behind the sweep would keep
contributing its own stale bearing, which is self-confirming. The offline range
now round-trips that field from the fixture 'lst'.

COMPANION FIX, and it matters: the radar-mode switch used the TRACKER's known
enemy count, so in melee the bot saw one enemy before scanning the second, locked
to 1v1, and the melee radar never ran at all. Now uses getEnemyCount() (server
truth), so melee mode persists until one enemy is genuinely left.

41 new unit checks (wraparound arcs, straddle at 0/360, single/empty enemies,
the 45 deg/tick cap, every phase transition and fallback). melee_scan is kept
but marked DEPRECATED; nothing in the rack imports it.

Non-regression: 33 gun-harness checks, vbullet metric, power selection, and
12/12 offline==online acceptance all pass.
2026-09-21 21:35:22 +02:00
SirStone d3b3c28cdf fix(adversaries): migrate to bot-api 1.0.7 - kills an intermittent crash that corrupted measurements
Four of the five adversaries imported the OLD package (tankroyale_botapi
1.0.1); only SittingDuck used robocode_tankroyale_botapi 1.0.7, which is what
the rest of the repo requires. A previous report claimed OscillatorBot was
already on 1.0.7 - that was WRONG, and OscillatorBot turned out to crash the
MOST (8 SIGSEGVs in the first reproduction, 15 in its historical /tmp logs).

THE CRASH, reproduced with an identical stack in every case:
  botThreadEntry -> run -> adversary run -> go -> dispatchPendingEvents ->
  tankroyale_botapi-1.0.1/event_queue.nim(89) addEvent -> realloc/rawDealloc ->
  SIGSEGV
Counts, old API: 60 melee battles x 8 rounds gave RandomMover 1, PatternMover 3,
WaveSurfer 0, OscillatorBot 8; 6 battles x 6 rounds vs SittingDuck gave 4/2/0/3.

ROOT CAUSE: the main->bot event hand-off. 1.0.1 passes a lock-protected
seq[BotEvent] (signalTick writes gPendingEvents, dispatchPendingEvents copies it
under lock). 1.0.7 uses a Channel[seq[BotEvent]] (send(move(pending)) /
tryRecv). The old path copied string-bearing BotEvent payloads across threads
every tick, churning ORC refcounts on the shared heap until the freelist was
corrupted. 1.0.7's own source documents this as the gdb-confirmed fix.

WHY IT MATTERED MORE THAN IT LOOKED: the crash silently corrupted measurements.
Against a stationary duck, crash contamination inflated WaveSurfer's rest
fraction from 12.4% (clean) to 20.7%; in a focused run the server logged
'Bot left: OscillatorBot' while the game continued and its score stopped
growing. So every gauntlet run tonight was fighting adversaries that were
partially dead - which is a second, independent reason the user's instinct that
these bots were bugged was correct, and why they should not be used as a
measurement baseline. (The per-gun REAL hit rates are unaffected: those came
from DrussGT battles.)

FIX: all four migrated to robocode_tankroyale_botapi 1.0.7. NO API adaptations
were needed beyond the module rename - every symbol these bots use is identical
in 1.0.7, verified by diffing the two packages (constants/utils/json_parse/
schemas semantically identical; the movement and intent procs in bot.nim are
byte-identical). The .nimble files now require robocode_tankroyale_botapi.

VERIFIED: 120 melee battles x 8 rounds plus 6x6 vs SittingDuck -> 0 SIGSEGV in
all four stderr logs (0 bytes). Behaviour unchanged: sub-1% absolute drift in
mean speed, rest fraction, reversal rate, mean range and perpendicular fraction,
all within run-to-run spread; the one >=3-sigma flag (WaveSurfer perpendicular
relative to DrussGT) was isolated against a stationary opponent and shown to be
the chaotic closed loop, not the migration. test_wavesurfer_velocity passes 7/7.

NOT migrated, reported only: GotoTest_garage, OscillatorBot_garage (archived
copy), PPO_Bot_garage, QBot_garage, SAC_LSTM_Bot_garage - older experiment
garages, left alone deliberately.
2026-09-21 09:14:47 +02:00
SirStone ab35540035 fix(logging): [config] reported a stale gun - it printed before selection ran
The user spotted this from the game itself: the [config] line always said
gun=HeadOn while the in-game turret and bullet COLOURS varied. The colours are
set at the selection site, so they were truthful and the log was not.

MEASURED, one 2-round battle, same process:
  [config] output      : 6 lines, ALL gun=HeadOn
  tracker selection    : Displace 25.1%, HeadOn 20.8%, Pattern 15.5%,
                         KNN 14.1%, Tsetlin 6.3%, WallBounce 6.1%, ...
Root cause: printConfig did GunNames[bot.currentGun] but EVERY call site ran
before the tick's gun selection - onRoundStarted right after currentGun = 0
(so HeadOn by construction), and the target/radar-change prints. The selection
that sets currentGun is ~490 lines later in the same tick. radarMode and
currentTargetId ARE updated before those sites, which is exactly why the radar
and target columns looked plausible while the gun column did not.

WHERE IT CAME FROM: git history shows commit 2cc2a3b ('cleaner logging')
removed the original printConfig call at the selection site while leaving the
prevGun highlight logic in place. That removal is when the regression appeared -
before it, the bot logged on round start AND on gun switch.

FIX: one emission at the end of the tick, after selectShot has run, gated by a
cfgDirty flag set on gun switch / target change / radar change. The index is
guarded (currentGun may be -1), the round-start line omits the gun field rather
than inventing one, and the existing output contract is preserved (all white,
only changed fields green, enemies= and target= kept).

VERIFIED: before 6 lines all HeadOn; after 782 lines over 2118 ticks with 13
distinct guns and ZERO violations - no printed gun was one that had not been
selected. Counts differ from the selection totals because the log prints only
on change, which is the behaviour the user asked for.

Also removes the 'currentGun = 0' initialisation in onRoundStarted, keeping the
existing -1 sentinel so 'no gun chosen yet' is representable.
2026-09-21 08:50:44 +02:00
SirStone 1e8f0d342a fix(adversaries): the launchers ran STALE binaries - this is why the earlier fix never took effect
P0. Three of the five launchers ran ./<Bot> (a tracked binary at the bot root)
while config.nims sets outdir=out and both the test framework's compileBots and
a manual 'nim c src/<Bot>.nim' write to out/. SittingDuck and OscillatorBot
correctly ran ./out/<Bot>; RandomMover, PatternMover and WaveSurfer did not.
cmp confirms the root and out binaries differed for all three.

Consequence: the previous session's adversary fixes were compiled into out/ and
never executed. Every gauntlet and every capture ran the OLD code. This is
almost certainly why the user's instinct that these bots were still bugged was
correct while the code claimed otherwise.

Fixed by pointing all five launchers at ./out/<Bot>, and by deleting the three
stale root binaries so the trap cannot recur. Verified end to end through the
booter: WaveSurfer went from standing still 96.2% of ticks with a 1398-tick
longest standstill, to rest 12.3% / mean speed 6.69 / longest zero run 18 /
perpendicular 0.845.

Also honours GUN_STATS_PATH in test_gauntlet_5bots.nim (same knob ModularBot
reads) so pooled gauntlet runs append to one file instead of clobbering the
default.

NOTE for a follow-up: the out/ binaries are still TRACKED build artifacts, which
is the same class of hazard that caused this. Untracking them (as was done for
ModularBot_garage/ModularBot) would remove the failure mode entirely.
2026-09-21 08:22:04 +02:00
SirStone c214abcfa8 fix(adversaries): repair four of the five sparring bots
The user suspected these were bugged. They were, and the verdicts are not
uniform - three genuinely broken, one merely sloppy, one fine:

- WaveSurfer: GENUINELY BUGGED, worst of the five. (a) The enemy velocity
  decomposition was sin/cos SWAPPED - enemyVx used sin and enemyVy used cos,
  while Tank Royale is 0 deg = East, CCW+, so it must be cos for X and sin for
  Y. Its linear-prediction gun was aiming at a reflected position. (b) The wall
  escape flipped strafeDir on EVERY tick the bot was inside the wall margin,
  so instead of turning away it flip-flopped in place: measured standing still
  (speed < 0.5) for 96.2% of ticks with a longest continuous standstill of 1398
  ticks. Fixed with a hysteretic wall-escape selection plus a corner escape,
  dead enemyLastDir removed, and per-round state reset.
  AFTER, measured through the booter: rest 12.3%, mean speed 6.69, full speed
  79.7%, longest zero run 18, perpendicular 0.845 / radial 0.012 - it now
  actually strafes. Gun sanity: lead error 1.0 px vs 106 px for head-on on a
  constant-velocity target; lead gun 45.8% hits vs 29.3% for head-on.
- PatternMover: GENUINELY BUGGED. Real deadlock - it decremented its step
  counter by the REQUESTED amount while issuing setTargetSpeed(8), so against a
  wall the counter never reached 0, advanceStep never ran and it was stuck
  forever (309-tick standstill). Now counts down by ACTUAL distance/turn with a
  STALL_LIMIT watchdog and steers toward the arena centre. Standstill 309 -> 19
  ticks; full-speed ticks 10.0% -> 28.4%.
- OscillatorBot: GENUINELY BUGGED, milder. No wall handling at all, so it
  ground along walls 53.4% of ticks and could pin in a corner. Added wall
  steering that preserves the fixed 25-tick reversal cadence. Wall-band 53.4%
  -> 18.6%, mean wall distance 72 -> 119.
- RandomMover: merely sloppy, not broken. Its turn intent saturated against the
  speed-dependent limit (18.4% of moving ticks clamped) and the fire gate was a
  very loose 10 deg. Now clamps to calcMaxTurnRate and fires within 3 deg.
  Saturation 18.4% -> 3.9%.
- SittingDuck: FINE. Speed 0 for 100% of ticks, zero shots. Left untouched -
  it is a duck by design.

Adds test_wavesurfer_velocity.nim, a direct assertion that the decomposition is
cos/sin and explicitly NOT the swapped form (7 cases).

KNOWN ISSUE, not fixed: RandomMover/PatternMover/WaveSurfer import
tankroyale_botapi 1.0.1 and intermittently SIGSEGV in
tankroyale_botapi/event_queue.nim:89 addEvent, freezing the bot for the rest of
the battle. It reproduces on old and new code and never occurs for SittingDuck/
OscillatorBot, which import robocode_tankroyale_botapi 1.0.7. Migrating the
three to 1.0.7 would likely fix it and is worth doing - it is a real
reliability risk for these as sparring partners.
2026-09-21 08:21:47 +02:00
SirStone e6653199bb fix(shim): make the generated DrussGT bot dir runnable by hand
Running /tmp/tr_bots/DrussGT/DrussGT.sh manually failed with:
  BotException: Required bot property 'name' is missing.
The Tank Royale Java bot API reads bot identity from BOT_* environment
variables (EnvVars: BOT_NAME, BOT_VERSION, BOT_AUTHORS, BOT_DESCRIPTION,
BOT_HOMEPAGE, BOT_COUNTRY_CODES, BOT_GAME_TYPES, BOT_PLATFORM, BOT_PROG_LANG,
BOT_INITIAL_POS). When the TR booter launches the directory it supplies those,
derived from the .json - which is why every headless battle worked - but
launching the script directly supplies nothing, so the API rejects the
handshake.

make_botdir.sh now exports them in the generated <dir>.sh (kept in sync with
the .json it writes), so the script works standalone as well as under the
booter.

Verified against a server started for the test: before, the exact error above;
after, 'Connected to: ws://localhost:4599' and the server logs
'Bot joined: DrussGT 3.1.4159'.
2026-09-21 08:16:09 +02:00
SirStone 013b9fe01e docs: correct the overstated 'inverted metric' claim; record tonight's fixes
Three corrections, all prompted by later measurements:

1. The virtual-vs-real rank correlation is NOT robustly negative. Six
   independent Spearman measurements now exist (-0.374, +0.335, +0.522,
   -0.371, -0.073, -0.037) and the sign flips on large samples, so it is near
   zero on average. The honest headline is that virtual hit rate is a POOR
   RANKER, not an inverted one. The report said 'not weak - it is inverted' in
   six places; it now says so in none. The practical conclusion (do not trust
   it for ranking) is unchanged; the mechanism claimed was wrong.
2. The offline==online acceptance is FIXED, not flaky. Root cause was that the
   replay spawned gun 13 (TMSelect) while the live rack has it disabled, and
   the shared VirtualTracker ring is ORDER-SENSITIVE, so gun 13's extra 4
   bullets/tick permuted the per-tick resolution order for every other gun and
   shifted the learning guns' observations. After closing gun 13's ready gate
   offline the live and offline KNN traces are byte-identical (904/904 lines,
   empty diff). 5/5 consecutive runs now report 12/12 exact with the death
   boundary included. Recorded with the lesson: a flaky proof was hiding a real
   bug. Also records the general A/B confound - disabling a gun removes its 4
   spawns/tick from the shared ring, perturbing resolution order for the rest.
3. Pruning was tested and does NOT help, so the verdict for Tsetlin and
   Displace changes from an implied drop to BELOW OVERALL - KEEP. 15 paired
   runs: baseline 6.18%, Tsetlin-off 5.76%, Tsetlin+Displace-off 5.46%;
   paired permutation p=0.57 and p=0.21; distributions completely overlap; a
   non-surfer control showed no separation. Being below average does not
   justify removal.

Also records the tie-break randomness fix, and quotes run counts with every
rate (6.95% over 13 runs vs 6.18% over 15 runs, same binary) rather than
presenting a single figure as definitive.
2026-09-21 06:59:00 +02:00
SirStone 4cd5618435 fix(test): repair the flaky offline==online acceptance; make the tie-break truly random
TASK 2 - THE FLAKY ACCEPTANCE TEST, root-caused. It was NOT a live/offline
boundary race as suspected. The replay spawned gun 13 (TMSelect) while live has
EnableTmSelector = false and never does. The shared VirtualTracker ring is
ORDER-SENSITIVE, so gun 13's extra 4 bullets/tick shift the ring head and
permute the per-tick RESOLUTION ORDER of every other gun. The learning guns
append observations in resolution order, so their predictions shifted and
produced small hit deltas that moved between runs.
Evidence: the first KNN divergence was at rtick=174 with the SAME resolution
set merely reordered (live ft133,138,139,142,148,150,151 vs offline
ft150,151,133,138,139,142,148); after closing gun 13's ready gate offline the
live and offline KNN traces became BYTE-IDENTICAL (diff empty, 904/904 lines).
Fix: mirror the live rack in the replay. No tick exclusion, no tolerance
loosening. Stability: 5/5 consecutive runs now report 12/12 exact, each with
enemyDied=true - the death boundary is included, not excluded. The proof is
now real rather than a lucky run.

TASK 1 - the tie-break was not random. randomize() was only reached
incidentally through initTsetlinGun(), so a rack without Tsetlin had a fixed
rand() stream and ties always resolved the same way across process restarts.
Added seedSelectorRng() after gun construction, honouring GUN_SELECTOR_SEED.
Evidence: unseeded, 6 separate processes gave different pick sequences; with
GUN_SELECTOR_SEED=42, 3 processes gave identical sequences.

TASK 3 - PRUNING DOES NOT HELP; keep the full rack. 15 PAIRED runs per variant
vs DrussGT, 8 rounds, identical seeds:
  baseline             3238 shots  6.18%  (events 6.16%)  200 dmg/run
  Tsetlin disabled     3522 shots  5.76%  (events 5.71%)  197 dmg/run
  Tsetlin+Displace     3478 shots  5.46%  (events 5.37%)  183 dmg/run
Paired permutation tests: -0.34pp p=0.57 and -0.70pp p=0.21. Per-run
distributions completely overlap (baseline range [2.68, 10.00]; 15/15 and 14/15
runs inside it). A Crazy control showed no separation either. So removing the
measured-worst real performers is neutral-to-slightly-negative, and with
sd ~1.8pp a definitive claim either way would need far more runs.

CORRECTION TO A CLAIM I MADE: the 'virtual metric is INVERTED' finding does NOT
reproduce. Job-24 measured Spearman -0.374; this job measures +0.335 over the
same 13 guns with a different but equally defensible aggregation. Two opposite
signs means the correlation is NOT robustly negative - it is WEAK AND
SIGN-UNSTABLE. The honest statement is that virtual hit rate is a poor ranker,
not an inverted one. The docs assert the inversion and need correcting.

Also adds per-process GUN_STATS_PATH/GUN_SHOTLOG_PATH so concurrent A/B runs do
not clobber each other, and an env-gated GUN_RACK_DISABLE for rack A/Bs. All
default behaviour is unchanged when the env vars are unset.
2026-09-21 06:56:44 +02:00
SirStone 19410164f1 docs: definitive gun-rack report on real measured numbers
Replaces the stale 2026-09-20 docs, which predated per-gun real attribution,
the offline gun range and the DrussGT boss, and whose verdicts were built on
virtual hit rates that turned out to be ANTI-correlated with reality.

docs/gun_rack_analysis.md (751 lines) covers: the test infrastructure
described honestly (offline range with its flaky-acceptance caveat, the 20
fixtures and what each set is good for, the live boss, and the A/B methodology
of per-run server-side real hit rate with an explicit overlap test); the
virtual-vs-real metric lesson with Spearman -0.374 and the point-vs-path A/B;
per-gun real performance and the 16-rule ranking A/B; the offline per-fixture
gun matrix; KEEP/MARGINAL/BELOW verdicts; and the five root-cause bugs with
before/after numbers.
docs/gun_rack_summary.md (58 lines) is the verdict table plus top actions.

The '~230 point' score-noise band that has been steering methodology all night
was re-derived from the artifacts rather than asserted: the 13 shipped-config
run scores span 175-526, s.d. ~105, i.e. a ~210-point 2-s.d. band.

Caveats recorded verbatim rather than softened: the offline==online acceptance
is flaky (typically 11/12 on unmodified HEAD), fixtures are perfect-information
and therefore optimistic vs live play, per-gun real N is small so single-gun
ordering is indicative, the headline numbers come from ONE wave-surfer
adversary, and HeadOn must stay despite being lowest because it is the floor
fallback (disabling it: 5.08% / 175 dmg vs 6.95% / 251 dmg).
2026-09-21 06:37:20 +02:00
SirStone 2c94dc221a test(selector): 16 ranking rules A/B'd against the boss - none beat the shipped config
Added runtime-tunable ranking knobs to the selector, all defaulting to the
shipped values so behaviour is byte-identical when unset: GUN_SELECTOR_WINDOW,
MINOBS, TIE, FLOOR, POOL, RANK, SHRINK, SEED. rankScore supports mean, Wilson
lower bound, UCB, Thompson and shrinkage. Also fixed hitRate's most-recent-N
read for sub-WindowSize windows (windowHits).

RESULT: NO candidate credibly beat the shipped config. 13 runs x 8 rounds vs
DrussGT, 3612 shots, base 6.95% at 251 dmg/run; every candidate's per-run
interval overlaps base, and the nominal 'winners' are <=0.6 SE apart on far
fewer shots. Kept the shipped default. Valid outcome, recorded plainly.

THE FINDING THAT MATTERS MORE: the virtual-bullet ranking is ANTI-correlated
with real hit rate - Spearman ~ -0.37 for the shipped config. It is not merely
weak, it is INVERTED. The guns with the highest VIRTUAL rates have among the
lowest REAL rates: Tsetlin 12.9% virtual / 5.8% real, WallBounce 12.9 / 6.2,
StopShot 12.6 / 6.1, AvgLead 12.3 / 7.0 - while Linear sits at 10.2 virtual /
10.7 real and KNN at 7.5 / 9.0. So what carries the selector is the floor/tie
HEDGING, not the ranking: removing the floor drops us to 5.08% / 175 dmg.
That also kills the 'exploration' hypothesis - every gun spawns virtual bullets
every tick, so sampling is uniform and the bottleneck is SIGNAL QUALITY, not
under-sampling.

FINAL PER-GUN REAL HIT RATE vs DrussGT (13 runs, 3612 shots, overall 6.95%):
  Linear 10.7 | Circular 9.9 | KNN 9.0 | Pattern 8.6 | Accel 7.3 | AvgLead 7.0
  GuessFactor 6.9 | DecayGF 6.4 | WallBounce 6.2 | StopShot 6.1 | Tsetlin 5.8
  Displace 5.3 | HeadOn 5.2
Keep: Linear, Circular, KNN, Pattern, Accel, AvgLead. Marginal: GuessFactor,
DecayGF, WallBounce, StopShot. Below overall: Tsetlin, Displace, HeadOn - but
HeadOn must STAY as the floor fallback, since disabling the floor measurably
hurt.

CORRECTION TO A CLAIM I HAVE BEEN MAKING: the 12/12 offline==online acceptance
is FLAKY. It fails 11/12 on the UNMODIFIED HEAD source (control: KNN 81 online
vs 71 offline), and the mismatching gun moves between runs (KNN, then
WallBounce) - a live/offline boundary race. So '12/12' was a lucky run, and
that proof should be treated as strong-but-not-exact until the race is fixed.
This diff does not touch replayFixture/spawnBullets/tickBullets and the
selector is never called during replay, so it is pre-existing.

SIDE FINDING, not fixed: the shipped live bot never calls randomize(), so the
'random tie-break' is a FIXED sequence across process restarts.

Overfitting guard vs a non-surfer (SpinBot): inconclusive - ModularBot fires
only 17-31 real shots/run against fast bots because the range-aware firing gate
is strict at long range, so the guard has little power. Wilson looked better
(18.5% vs 8.6%) but on 70-92 shots with a 5-33% spread. Not evidence either way.
2026-09-21 06:31:00 +02:00
SirStone 57b2ac3849 feat(guns): scale-aware power selection (+52% damage); TM classifier gun built, measured, DISABLED
TASK 2 - power selection, a clear win. bestPower used an ABSOLUTE
MinHitRate = 0.40 bar. Measured per-bin virtual rates (rolling-100 fraction)
show no bin ever clears 40%, so 11 of 14 guns were stuck at bin 0 (power 1.0)
even where higher bins were comparable:
  Linear  p1.0 44% p1.5 39% p2.0 30% p3.0 29%   old bin 0 -> new bin 3
  Accel   p1.0 44% p1.5 40% p2.0 26% p3.0 29%   old bin 1 -> new bin 3
  Pattern p1.0 50% p1.5 40% p2.0 27% p3.0 12%   old bin 1 -> new bin 2
Replaced with a scale-aware PowerBarFrac = 0.50 (a dimensionless FRACTION of
the gun's own best bin rate). 13 of 14 selections now pick heavier bullets.
Real effect vs DrussGT (8 rounds x 3 runs): hit rate unchanged (7.56% ->
7.47%) but damage dealt +52% (157 -> 239 per run) and rounds end faster.
Same accuracy, half the shots, half again more damage.

TASK 1 - the TM pattern-classifier gun does NOT earn its slot. It was built as
a mixture of experts with a corrected-Granmo TM as a multi-class gate over
HeadOn/Linear/Circular/WallBounce/Accel, labelled by which expert's prediction
was closest to the actual enemy position (an exact, supervised, per-shot
label - no delayed credit). Offline it loses to the best of its OWN experts on
essentially every fixture, and against DrussGT it cost real performance:
  baseline (path+relative)  7.56% real hit rate, damage 157
  + power fix               7.47%,                 damage 239
  + power fix + TM gun      5.59%,                 damage 133
The gun was selected on 806 ticks and fired 24 real shots at 4.2%.
So the tree ships with EnableTmSelector = false: code and wiring kept intact
for re-enabling, but it is not in the active rack.

Worth recording from the clause dump: the gate DOES latch onto meaningful
structure. On energy-threshold-turner, HeadOn's clauses key on the energy bits
(the rule's own driving variable) while Circular keys on distance/velocity. So
the TM is learning something real and interpretable - it simply cannot beat
'always pick the best expert'. Root cause (INFERRED): the closest-expert label
is noisy because several experts are near-tied, and under the path metric the
winner varies by power bin while the gate sees one shared per-tick input, so a
one-vs-rest gate over a saturated 870-bit clause space has no margin to exploit.
(Zero-padding the 2-frame window was tried first and saturated every clause at
256-755 included literals; alternating the two real frames fixed that.)

Also factors the corrected feedback into an exported tmLearnDir and exports the
encoding/TM primitives; the Tsetlin tests still reproduce the documented
mean=13.8 included literals, so the refactor is behaviour-preserving.

Verified: 33/33 guard checks, tsetlin tests green, metric checks green, new
power-selection guard green (13/14 selections change; relative bar still picks
bin 1 and not bin 3 for a [30,25,12,5]% profile), 12/12 offline==online
acceptance under the shipped default.
2026-09-21 05:19:07 +02:00
SirStone dea4dcb574 feat(gun_harness): scale-aware selector thresholds; default = path + relative
The selection thresholds were calibrated for a rate scale that does not exist.
MEASURED on an exact offline replay of a fogged live WorldState vs DrussGT
(1397 selection ticks), the 0.10 absolute floor fires on 53.0% of point-metric
ticks and forces HeadOn, which has a REAL hit rate of 2.0-4.4% - worst or
near-worst of 13 guns. HeadOn's selection share: 69.1% (abs+point) -> 43.5%
(rel+point). My earlier claim that the floor fires ALWAYS is REFUTED - it is
53%, because bestRate is a max over gun x bin and a >=50-sample bin
occasionally clears 10%. The mechanism is confirmed; the literal statement was
not.

Scale-aware mode (GUN_SELECTOR_MODE, absolute|relative, default relative):
  RelTieMargin   = 0.20  dimensionless FRACTION of bestRate, replacing the
                         fixed 2pp band so the band scales with the metric
  FloorPeakFrac  = 0.25  the floor fires iff bestRate < 0.25 * peakRateRef,
  SelectorWindow = 256   where peakRateRef is the field-best rate over the last
                         256 selection ticks - keeping the original 'don't trust
                         a collapsed field' purpose but only when the field is
                         bad RELATIVE TO ITS OWN RECENT BEST, and counting only
                         guns with >= MinObsBeforeCompete samples so cold-start
                         100% spikes cannot pin HeadOn
  also pools the rate over power bins instead of taking the max over bins, so
  one lucky bin no longer wins
absolute mode is preserved byte-for-byte for rollback.

A/B vs DrussGT, real server hit rate, 3 runs x 10 rounds per config, one frozen
binary:
  absolute+point  3.66 / 2.45 / 5.01   pooled 3.76%
  absolute+path   7.55 / 8.21 / 6.83   pooled 7.57%
  relative+point  7.66 / 6.18 / 5.79   pooled 6.59%
  relative+path   7.15 / 7.55 / 6.90   pooled 7.21%
absolute+point is SEPARATED from all three (p < 0.0001); the other three
OVERLAP each other (p = 0.18-0.64). So the METRIC is the dominant lever and
under path the two threshold models are statistically tied.

DEFAULT SET: metric = path, thresholds = relative. absolute+path was nominally
0.35pp higher but indistinguishable (p = 0.64); relative is the principled
scale-aware fix, is the only model that works under BOTH metrics, and prevents
the point-metric catastrophe if anyone switches back. Shipping absolute would
ship the accidental side-effect this work exists to remove.

STILL NOT SOLVED: the selector remains only a moderate ranker.
Spearman(virtual rank, real rank) is 0.52 for the winning config, 0.36 pooled
for path and 0.04 for point - and it is INCONSISTENT across run sets. The
metric switch won by de-selecting HeadOn, not by ranking guns better. That is
the next problem.

TASK B, report only: do NOT drive selection from raw real hit rates yet.
Only the selected gun fires, so unselected guns get near-zero real shots
(GuessFactor 20, Linear 24 vs HeadOn 733); noise is fatal (n=470 at p=10% gives
+/-2.8pp, most guns n<200 gives +/-5pp+ across a 3-15% spread); and real rate is
conditional on when the gun was selected. A blended signal with forced
exploration and shrinkage is defensible in principle but needs thousands of
shots per gun across many battles. Real rate is best used OFFLINE as the
evaluation metric - which is exactly what this A/B did.

RELATED BUG FLAGGED, not fixed: MinHitRate = 0.40 in bestPower is on the same
wrong scale - no bin ever clears 40%, so once every bin has data, power
selection falls back to bin 0 (power 1.0) late in a round.

Verified: 33/33 guard checks, 11/11 metric checks, tsetlin green, 12/12
offline==online acceptance under the shipped default, run_range rc=0 over 20
fixtures. Adds analyze_selector.nim to measure floor/tie/bestRate/HeadOn-share
per config on any fixture.
2026-09-21 04:33:52 +02:00
SirStone 3b5d70b7c3 feat(gun_harness): runtime metric switch + A/B proving the point metric mis-selects
Adds GUN_VBULLET_METRIC (point|path, default point = unchanged behaviour) so
the virtual-bullet hit model can be selected at runtime with no rebuild. Both
the live tracker and the offline replay read the same value, so the 12/12
offline==online acceptance holds under EITHER setting (verified for both).

A/B AGAINST THE LIVE BOSS, real server-side hit rate as ground truth, 5
battles x 12 rounds per metric on one frozen binary:
  point  4660 shots / 219 hits = 4.70%   (per-run 3.16-5.53)
  path   4834 shots / 359 hits = 7.43%   (per-run 6.55-8.24)
The distributions DO NOT OVERLAP: path's worst run beats point's best run.
+2.73pp, +58% relative, z = 5.56, p < 0.0001. Range distributions were
identical (~460-478 px), so this is not a range confound.

MECHANISM - and this is the important part. The gain is SELECTION, not better
gun learning. Under the point model every gun's virtual rate is compressed
into 0.6-4.4%, so HeadOn sits inside the 2pp tie margin and takes 72.6% of
selection ticks / 76.9% of shots - while HeadOn is 11th of 13 by REAL hit rate
(2.3%). The path model widens the band to 4.7-13.7% and ranks HeadOn 10th, so
its shot share falls to 35.9% and Pattern/Accel/WallBounce get picked instead.
Counterfactual: applying the point model's per-gun real rates to the path
model's shot mix yields 7.65%, i.e. essentially the whole observed gain.
So the selector, not the guns, is where the win lives.

PER-GUN REAL HIT RATE vs DrussGT (path mix, the answer to 'which guns are
worth keeping'): WallBounce 10.8, Pattern 10.5, Accel 10.0, Displace 9.3,
Circular 9.2, AvgLead 8.5, KNN 5.7, StopShot 5.2, GuessFactor 3.7,
Tsetlin 2.9. Per-gun N is small (hundreds of shots) so single-gun ordering is
indicative, not definitive.

TWO CAVEATS, recorded because they undercut a naive reading:
1. One adversary. DrussGT is a wave surfer and HeadOn is genuinely bad against
   surfers, so part of this may be matchup-specific.
2. The path model is NOT a better general ranker. Spearman(virtual rank, real
   rank) is 0.52 under point vs -0.04 under path. It wins by accidentally
   fixing HeadOn's mis-rank, not by ranking guns better. A more durable fix is
   to address the selection logic directly - which is the next job.

Also adds a focused guard test (test_vbullet_metric) covering parsing/default,
a receding-target point-miss/path-hit, a perpendicular-target path-miss, and
replay determinism.

Verified: 33 guard checks, 12/12 acceptance under both metrics, tsetlin tests
green, range 34.3% (point, unchanged) / 50.8% (path).
2026-09-21 03:58:27 +02:00
SirStone e2ca2fc7d8 fix(guns): recover the DrussGT regression with a radial-fraction range blend
The previous fix (learn the residual against a constant-velocity base) was
structurally right but cost us on real wave-surfing movement: GF 108 -> 55,
KNN 101 -> 74 on the classic DrussGT captures. Root cause: the linear base is
a poor model for a surfer, so the residual histogram is noisier than the old
total-lead histogram.

FIX: blend the RANGE between a radial-only forecast and the geometric one by
radialFrac (the fraction of recent per-tick motion that is radial), keeping
the constant-velocity bearing. dist = radialDist + rf*(linearDist - radialDist).
New VelocityTracker in common_libs/guns/lead_forecast.nim; the window default
is 32 and results were identical at 16 and 40, so it is not tightly tuned.

Nine candidate bases were measured and rejected WITH NUMBERS rather than by
argument, which is why I trust the winner:
  velocity scaling 0.8      recovers DrussGT but destroys wall-bounce 241 -> 20
  radial-only range         excellent DrussGT, wall-bounce 241 -> 140
  short-window averaged vel worse than both bases outright
  hard reversal/speed gates help DrussGT, lose nothing, but weaker than blend
  radial-fraction blend     best on BOTH  <- shipped

Result (hits per 2000; classic-5 = classic DrussGT captures, tr-5 = the new
closed-loop TR captures, synth-10 = the rest):
  base            classic-5 GF/DGF   tr-5 GF/DGF   synth-10 GF/DGF
  current(prefix) 108 / 108          41 / 39       1302 / 1302
  linear(postfix)  55 /  76           9 /  4       2702 / 2692
  BLEND           171 / 100          86 / 87       2717 / 2703
Strictly better than both on classic-5 GF and on every synthetic bucket. The
one figure below the old base is classic-5 DecayGF (108 -> 100, -8/2000,
within noise) and that is stated plainly rather than hidden.

TASK B - enemy energy in learners. KNN gains an 8th feature, enemyEnergy/100,
on a FIXED [0,1] scale (not min-max) because threshold behaviour keys off
absolute energy. Honest result: it is NEUTRAL on the target fixture (77 vs 77)
and roughly neutral in aggregate. The base change, not the feature, moved that
fixture. Tsetlin already encoded enemyEnergy and now scores 88/400 on
energy-threshold-turner against Linear's 43/400 - a 2x margin, which is the
'can a TM learn a high-level pattern' question answered in gun form.

TASK C - is the virtual-bullet metric itself faithful? Quantified: scoring the
bullet's PATH against BotRadius instead of the single point at aim distance
raises every gun by +31% (GF) to +86% (HeadOn), so the current model is
PESSIMISTIC, and it RE-RANKS materially: Linear 9th -> 6th, AvgLead 7th -> 3rd,
GuessFactor 4th -> 9th, DecayGF 6th -> 12th. The 12/12 offline==online
acceptance still holds under the path model (verified with a temporary env
hook driving both sides), so no red flag. VERDICT: do NOT switch. The point
model is the standard virtual-bullet PREDICTION-ACCURACY fitness - the bullet
must arrive at the predicted point at the right time - while the path model
measures hypothetical hit chance against a target that never dodges, and in
open-loop fixtures it over-credits directional guns (HeadOn 35% on DrussGT,
100% on constant-velocity) for exactly that reason. The models differ
materially but the current one is not shown to be unfaithful FOR ITS PURPOSE.
Because the metric drives gun SELECTION, this is now being A/B'd against real
hit rate versus the live DrussGT boss, which is the only ground truth we have.

Verified: 20 fixtures 35636/104000 (34.3%); 33 guard checks; 12/12 acceptance;
tsetlin tests green; live gauntlet 5/5.
2026-09-21 02:14:30 +02:00
SirStone 17c99f542f feat(fixtures): closed-loop DrussGT captures from real Tank Royale battles
The classic captures were OPEN-LOOP: replayed DrussGT never dodged OUR
bullets. These come from real TR battles through the working Java bridge, so
the recording contains genuine reactions to ModularBot's live fire. The
open-loop caveat is gone (perfect-information remains).

PRIMARY RESULT - the boss beats us badly. DrussGT 1447 - ModularBot 300 over
15 rounds, ModularBot winning only round 5 (DrussGT died at tick 1893). Rounds
are long, not truncated: mean 1335 ticks, ModularBot got off 1134 shots.
  ModularBot  1134 shots /  60 hits =  5.3% real hit rate
  DrussGT     1400 shots / 169 hits = 12.1% real hit rate
So DrussGT's gun is ~2.3x more accurate than our entire rack, on top of far
better movement. That is the number to move.

Also captured: shield-on variant (DrussGT 939-287, 9/10 - ModularBot takes
round 1 to the known shield warm-up), and vs SpinBot 1175-0, Crazy 1080-1,
Corners 1659-0. 20,026 + 12,629 + 10,824 + 11,507 + 2,575 ticks.

Movement statistics match the classic set within ~0.04 on the perpendicular
and radial fractions, so this is the same wave surfer in TR physics:
  TR vs modularbot: perp 0.967, radial 0.001, 52.8% at full speed,
                    reversing 46.3%, median range 464 px.

CLOSED LOOP PROVEN, not asserted. ModularBot's fire is a heat-limited near
metronome (median interval 14 ticks), which gives a usable exogenous clock:
  - event-locked |delta heading| oscillates 0.96 -> 2.69 deg about a 1.47 deg
    mean with the fire period, almost every lag outside the 95% band of a
    400-iteration phase-shuffled null;
  - cross-correlation of |delta heading| against the fire impulse peaks at
    r = +0.111, lag 12 ticks, permutation p = 0.005 (null peak mean +0.016);
  - OWN-FIRE CONTROL is flat, so the oscillation is enemy-driven rather than
    an internal cadence;
  - range response is weak (~3 px over 30 ticks, near noise) and is therefore
    NOT claimed, and per-bullet dodging is not claimed either because the
    bullet detector (shield) is off.

CAVEATS: still perfect-information (observer gives true positions every tick,
unlike the live bot's stale between-scan WorldState) so these remain optimistic
vs live play; and they are open-loop AT REPLAY TIME - 'closed_loop' describes
the capture, not a later replay. TR conversion residual is ~1.5 deg mean
because the TR server moves along the pre-turn heading, vs 0.000 deg for the
classic captures.

Adds analyze_closed_loop.py (PSTH event-locking, phase-shuffle permutation
null, cross-correlation, own-fire control) and per-round result sidecars.
2026-09-21 01:18:13 +02:00
SirStone 7f706e5b14 fix(guns): GF family aimed at the wrong RADIUS, not the wrong angle
The entire GuessFactor family scored 0% on clean circular and wall-bounce
trajectories. Two hypotheses were on the table and BOTH were wrong:

- MEA range too narrow / edge clamping: REFUTED. Measured 0 clamped shots
  out of 837/849/957, required offsets peak at ~33 deg against MEA
  28.1-46.7 deg, and the 8 in arcsin(8/bulletSpeed) is correct (it is the max
  robot SPEED, not the hit radius). Changing it to BotRadius=18 would have
  coarsened resolution for nothing.
- Peak selection: REFUTED. A sweep of every constant GF value showed the
  ORACLE-BEST constant offset on the original gun was only 6% circular,
  4% wall-bounce, 7.5% random-walk. No peak choice could have done better.
  The learning path was fine too: ~850-960 observations per fixture, 0
  starved waves, well-populated histograms.

REAL CAUSE: the GF family aimed at the FIRE-TIME distance. The virtual-bullet
metric resolves a bullet at the AIM-POINT distance and scores that single
point against the enemy's position on that tick, so with any radial target
motion the bullet stops at the wrong radius and misses even with a perfect
angle. Angle-only prediction is structurally unscoreable under this metric.

FIX: give the GF family a self-consistent constant-velocity forecast as its
base reference (new common_libs/guns/lead_forecast.nim, which iterates the
flight time to the same fixed point circular.nim uses), so the histogram
learns the RESIDUAL against that forecast and the aim point lands at the
right radius. Applied to guess_factor, decay_gf and knn_gun.

Same defect fixed in Linear: it did a one-shot dist/bulletSpeed extrapolation
and never iterated its flight time.

The oracle sweep proves the structural fix, independently of tuning: the best
achievable constant GF moved 6% -> 20% (circular), 4% -> 57% (wall-bounce),
7.5% -> 49% (random-walk).

MEASURED, all 15 fixtures: total 39.0% -> 44.4% (30399 -> 34654 hits).
  circular       GF 6 -> 23,   DecayGF 6 -> 21
  wall-bounce    GF 0 -> 60.2, DecayGF 0 -> 60.2
  constant-vel   GF 26 -> 100, DecayGF 26 -> 100, KNN 26 -> 100, Linear 87 -> 100
  random-walk    GF 0 -> 53,   DecayGF 0 -> 52,  Linear 24 -> 53
  StraightLine   GF 8 -> 77,   DecayGF 8 -> 77
Non-regression: 33 guard checks pass, the range's 12/12 offline==online
acceptance still PASSES, tsetlin tests green, live gauntlet 5/5.

HONEST TRADE-OFF, recorded rather than hidden: on the 5 real DrussGT
wave-surfing captures the GF family REGRESSES - GuessFactor 108 -> 55,
DecayGF 108 -> 76, KNN 101 -> 74 hits per 2000. The linear base is a poor
model for a surfer, so the residual histogram is noisier than the old
total-lead histogram. Linear itself improved there (95 -> 105). The synthetic
range and the live gauntlet both improved, and the structural bug is provably
fixed, so this was judged worth the cost - but recovering the DrussGT
regression is the next job, not something to wave away.
2026-09-21 00:56:02 +02:00
SirStone 8e2be6a4c6 feat(tools): the real DrussGT now plays and wins Tank Royale battles
The unmodified DrussGT.jar connects, wave-surfs, fires and beats every
adversary we have. 5 rounds each, all rounds won:
  SpinBot 542-16, Corners 823-4, Crazy 604-0, RamFire 900-0,
  ModularBot 506-72.

It is genuinely surfing, not drifting or stalling. Movement statistics
against the classic captures, same metrics, same analyzer:

  opp        perp TR/classic   reversing TR/cl   median range TR/cl
  SpinBot    0.967 / 0.962     0.466 / 0.464     462 / 402
  Corners    0.894 / 0.920     0.462 / 0.419     520 / 504
  Crazy      0.841 / 0.829     0.527 / 0.439     370 / 378
  RamFire    0.762 / 0.651     0.533 / 0.417     326 / 283
  ModularBot 0.956 / -         0.466 / -         453 / -

Every round starts moving within 3-11 ticks and runs at 75-82% full speed.

Implemented: the classic remaining-quantity motion model (delegated to the
TR Bot's own Nat-Pavasant model - identical constants: accel 1, decel -2,
max 8, body turn 10-0.75|v|, gun 20, radar 45 - so getDistanceRemaining and
getTurnRemaining are exactly self-consistent with what is emulated); event
synthesis with classic ordering; bullet identity via object identity; gun
heat; rounds; radar cadence; firing translation; and the ThreadManager
landmine is killed by installing a no-op IThreadManagerBase in
ContainerBase.instance (verified: without it 'RobotException: ThreadManager
cannot be null!' kills the bot thread; with it the write succeeds).

Also adds TrBattleCapture, an observer that dumps per-tick state in the SAME
JSONL fixture format as the classic capture, so legacy bots can be captured
from Tank Royale battles too.

HONEST DIVERGENCES (README section 5.9): the TR server moves along the
PRE-turn heading and then turns, while classic aligns displacement with the
POST-turn heading, so the analyzer's conversion error is ~1.5 deg rather than
0.000 deg; distanceRemaining decrements by target speed rather than actual
distance; collision clamping differs; BulletMissedEvent can fire less often
because age-expiry has no TR event; bulletId is a local temp id;
StatusEvent/onPaint/SkippedTurnEvent are never delivered. Physics fidelity
diverges by construction - expect to retune.

EnergyDomeWorker (the bullet shield) is off by default: its precise
bullet-detection warm-up makes round 1 up to 79% stationary vs 35% in
classic. Rounds 2+ match classic closely with it on, but the pure surfer is
the consistent path. DRUSSGT_SHIELD=1 re-enables it.

Jars remain out of git.
2026-09-21 00:29:37 +02:00
SirStone d5061ee215 test(range): restore the 12/12 offline==online proof; measure TM clause readability
Task 1 - the acceptance proof was unrunnable because RecordWorldState was a
compile-time const set to false. It is now a RUNTIME switch
(let RecordWorldState* = existsEnv("TR_RECORD_WORLDSTATE")), default OFF, so
ordinary runs write no fixture, and acceptance_offline_vs_online.nim enables
it for the battle it spawns and clears it afterwards. Restored and run twice:
12/12 deterministic guns match exactly (128-tick and 546-tick battles), with
Tsetlin reported separately as stochastic. Both nimble build variants clean.

Task 2 - does a compact encoding turn the TM's 99.35% into a READABLE rule?
Measured across window sizes (fixed seed, no tuning):

  frames  TEST acc  eff.lits/clause  firing clauses  counterfactual low/high/mean
  10      99.35%    152.8            37              100/24/62.4%
  3       95.94%    54.9             35              96/20/58.6%
  2       99.48%    39.6             38              95/25/60.7%
  1       98.30%    19.2             45              100/24/62.3%

So 2 frames is strictly better than 10 on BOTH axes: +0.13 accuracy for 4x
smaller clauses. The 3-frame dip is non-monotonic and left unexplained rather
than smoothed over.

A readable rule WAS partially recovered. Five clauses carry the exact Gray
form !g10 ^ !g9 ^ !g8; g10 is inert in this data, so the effective rule is the
2-literal proposition !g9 ^ !g8, i.e. energy < 25.6. That is a genuine
threshold in readable propositional form - but at 25.6, NOT the labelled 30,
because 256 is a power-of-two Gray boundary expressible in two literals while
300 needs a longer conjunction. The TM found the nearest SIMPLE threshold.

The honest caveat: that threshold is not the ensemble's decision mechanism.
The counterfactual follow rate (high 24%, mean 62.3%) is statistically
identical at 1, 2 and 10 frames, so compactness did not make the model read
energy - its vote is carried by co-occurring bearing/velocity/heading/wall
literals. Also identified: clauses containing all 11 Gray energy bits are
satisfied at exactly one raw value (50, the dataset floor), so they are
'energy has hit the floor' detectors, not thresholds.

Methodological fix worth keeping: the earlier single-frame counterfactual
wrote energy into all 10 frame slots including the zeroed ones, reviving dead
clauses and producing a spurious 2% high-follow rate. setEnergyFrames now
rewrites only the exposed frames; the corrected figure is 24%.
2026-09-21 00:14:46 +02:00
SirStone 89370008da fix(tsetlin): make the TM actually learn - saturation 714 -> 13.8 literals/clause
The gun has never contributed anything: Tsetlin.vHits was byte-for-byte
equal to Linear.vHits in every measured round of every run, because its
learned correction was always exactly 0.

Six diagnosed defects fixed, plus one that was required to make the first
one work:

1. Type I now conditions on the clause output. It previously rewarded
   included true literals unconditionally, omitting Granmo's (c=0, lk=1)
   -> toward Exclude counter-force, so true literals ratcheted toward
   Include forever. This was the root cause of the saturation.
2. Type II was unreachable dead code: its guard required cOut==1 AND
   lits[lit]==0 AND st>0 (included), but cOut==1 guarantees every included
   literal is 1. Its direction was wrong too - it should increment EXCLUDED
   false literals when the clause fires.
3. Resource allocation restored: Granmo's (T - clip(v,-T,T))/(2T) target
   replaces |error|/(2*RESID_MAX); TM_T was only an output normaliser.
4. Label baseline fixed - the factor-2 shrink. predX = linearX + cx, so the
   label was delta - cx while the learner's output IS cx, giving
   error = delta - 2cx and a fixed point of cx = delta/2: HALF the needed
   correction even with perfect feedback. TmTrace now stores linearX/linearY
   and training uses delta.
5. Hits no longer zero their label (a hit means |miss| < 18px, not 0).
6. The enemy-energy feature was duplicated - tmEncodeFrame passed
   state.selfEnergy with a stale comment claiming enemyEnergy was absent,
   while WorldState.enemyEnergy exists. Enemy-energy rules were literally
   unrepresentable.
7. REQUIRED EXTRA: tmEvalClause now implements Granmo Eq. 6 - an all-Exclude
   clause outputs 1 during learning and 0 during classification. Without it,
   fix #1 deadlocks every clause at empty.

MEASURED EFFECT (energy-threshold-turner fixture, seed 1):
  mean included literals per active clause   714.0 -> 13.8
  active clauses                             100/100 -> 53/100
  nonzero corrections                        8/764 -> 708/764
  Tsetlin virtual hits (Linear = 27/400)     27/400 -> 69/400

Divergence achieved: offline on 7/8 fixtures, and in a live gauntlet
(RandomMover: Tsetlin 199/1200 vs Linear 288/1200, vDropped=vStarved=0).
Tsetlin now LEARNS but is not yet competitive with Linear - the regression
head is untuned, flagged as follow-up rather than claimed as a win.

Also ignores compiled test harnesses that have no file extension, which the
existing '**/tests/test_*' rule misses.
2026-09-20 23:59:32 +02:00
SirStone e0bfa9b5e1 feat(tools): drive the real DrussGT jar outside the Robocode engine
The question was whether we can fight a genuine legacy leader bot in Tank
Royale. Answer: the API side is now PROVEN, not estimated.

KEY FINDING: the classic robocode.* API is a thin delegation layer over a
public seam. javap -c shows AdvancedRobot forwarding every call to
_RobotBase.peer (IBasicRobotPeer/IAdvancedRobotPeer), and _RobotBase.setPeer
is public final. So we do NOT need to reimplement the API: we reuse the
genuine robocode.jar and implement only the 75-method peer interface.

Consequences:
- DrussGT's 22 sources compile against the real API with ZERO unresolved
  symbols. (The literal 22-file javac fails only on two PRE-EXISTING
  duplicate classes - GFRange and Indice are declared both inline in
  DrussGunDC.java and as standalone files - and the 6 classes that ship
  without source. The jar supplies all of them, so a shim never cares.)
- Runtime smoke test PASSES: the unmodified DrussGT.jar is loaded through a
  child URLClassLoader and driven for 200 synthetic ticks, emitting movement
  intents every tick and 169 fire requests, with its thread surviving.

This is the cheap path to the 'final boss', and it also unlocks the whole
roborumble archive rather than one bot.

Traps found by measurement:
- ScannedRobotEvent's constructor order is (name, energy, bearing, distance,
  heading, velocity) - NOT heading-before-bearing. The wrong order silently
  yields distance=0, an immediate KD-tree insert and an NPE in
  EnemyMoves.predict; it hung the first smoke run.
- RobocodeFileOutputStream has a hard engine dependency (resolves
  IThreadManagerBase via ContainerBase) and throws 'ThreadManager cannot be
  null!' outside the engine, killing the bot thread. Reached from DrussGT's
  own contain() error logging, so it must be stubbed.
- robocode.RobotDeathEvent is required and was NOT in the predicted API list.
- Bullet.equals() is genuinely called for bullet identity, not just getters.

REMAINING WORK (README section 5): coordinate rotation DONE, execute()->tick
bridge DONE and proven, ThreadManager fix scoped. The main open item is the
classic motion model (setAhead distance semantics vs TR speed), estimated
1-3 days, plus event synthesis/ordering, gun heat, bullet identity, round
and radar cadence, and firing translation. NO API UNKNOWNS REMAIN.
Physics fidelity will still diverge from classic - expect to retune.

Jars stay out of git (blocked by tools/robocode_shim/.gitignore); the
genuine robocode.jar and DrussGT.jar are referenced from /tmp via
ROBOCODE_JAR / DRUSSGT_JAR.
2026-09-20 23:52:51 +02:00
SirStone 76ae6170f8 docs(research): portability audit of DrussGT 3.1.4159
Measured hazard counts for a Java->Nim port, with the correction that only
22 of 28 top-level classes ship source (6 do not: 5 in the gun package plus
dMove/Scan), so a full port would need a decompiler while a shim would not
care at all.

Real traps: 193 float / 35 casts / 147 literals / 60 float[] in the movement
closure (the danger histogram is float[171] -- porting 32-bit Java floats to
Nim's default float64 diverges silently); ~68 non-final statics; and 4 Java
single-& sites with side effects, which break under Nim's short-circuiting
'and'. Non-issues, correcting earlier assumptions: 0 sites of %-on-negative
(angle normalisation is floor-based) and FastTrig has no lookup tables, it is
7 coefficient-exact polynomials.

Movement scoping: 5007 LOC across 14 files, ~4.8-6.1k Nim LOC, 4-8 focused
agent-days to first-compiles. It can be ported WITHOUT the gun (data flows
movement->gun only), but the harness never routes HitByBullet into movement
modules, so the danger bins would never train -- that plumbing is the real
blocker, not the translation.
2026-09-20 23:44:42 +02:00
SirStone 4f18c8ce07 feat(tools): capture real DrussGT movement from classic Robocode as fixtures
There are no genuinely competitive adversaries for Tank Royale, and the
in-repo ones were broken until recently. Classic Robocode 1.9.5.5 is
obtainable (SourceForge, 20.4 MB) and its programmatic control API
(robocode.control.RobocodeEngine + BattleAdaptor.onTurnEnded) can run real
battles headless and expose per-turn robot state. So a legacy leader bot's
MOVEMENT can be captured and used as a gun-testing fixture with no port.

Captured unmodified DrussGT 3.1.4159 vs spinbot/ramfire/crazy/corners and a
mirror match: 28,797 ticks, plus two trivial-bot contrasts. Conversion to
the Tank Royale convention is validated to 0.000-0.001 deg by recomputing
the direction implied by (heading, speed) and comparing it against the
recorded per-tick displacement -- i.e. the data is proven to be genuine
recorded motion rather than a mangled export. (A first attempt treated the
snapshot API's headings as degrees; they are radians, ~95 deg off.)

The statistics confirm it is really a wave surfer: perpendicular to the
opponent 65-96% of ticks, radial ~0.001, 41-72% of ticks at full speed,
reversing on 42-46% of ticks, holding range at a 283-526 px median. The
straight-line contrast is radial-dominant (0.75) with ZERO reversals.

Discovery: DrussGT detects predictable guns and switches to a bullet-shield
stand-still mode, so captures against sample.Walls/TrackFire had to be
rejected as non-movement.

CAVEATS, recorded in DRUSSGT_FIXTURES.md: these are open-loop (replayed
DrussGT never dodges OUR bullets) and perfect-information (the observer
gives true positions every tick, unlike our stale live WorldState). Both
make our guns look better than in live play, so use them for RELATIVE gun
ranking, not absolute hit rates.

Jars stay out of git; capture tooling is reproducible via capture.sh.
2026-09-20 23:44:42 +02:00