The belief "BotDeathEvent never reaches ModularBot, so enemyTracker keeps dead
enemies alive forever" was written into a code comment and then believed twice.
It is FALSE. Measured in a 7-bot melee with a per-tick probe comparing
enemyTracker's alive count against the server's getEnemyCount():
metric 1.3.1 (20 rd) 0.35.5 (15 rd)
observed enemy deaths 83 68
...non-round-ending 83 (100%) 66 (97%)
ekBotDeath events DROPPED 0 0
max dispatch lag (turns behind) 1 1
phantom ticks 1 / 16,820 1 / 12,596
MAX CORPSE LIFETIME 0 ticks 0 ticks
victims still alive at round end 0 0
onBotDeath fires for every death, including non-round-ending ones. The
API-level event-drop mechanism IS real (test_event_drop_mechanism.nim proves
it: ekBotDeath is not in isCritical and MAX_EVENTS_AGE=2) - the bot simply
never falls far enough behind for it to trigger (max lag 1 turn).
Removed:
- reconcileWithServer + ReconcilePersistTicks/mismatchTicks/sawServerAlive
(uncommitted, and ON BY DEFAULT despite the premise being false). Its own
comment admitted a shorter window once KILLED A LIVE ENEMY ("it fired three
more times after the tracker marked it dead") - a latent mis-prune path
defending against a bug that does not exist.
- The radar's CorpseTicks=40 filter and the same-class age>60 filter in
recordRadarStats, both carrying the false comment. Removal changes no real
behaviour: buildState feeds the radar enemyTracker.allAlive(), so a dead
enemy never reaches computeScan.
Kept:
- The TR_TRACKER_PROBE instrument (default OFF), which produced the table above.
- test_event_drop_mechanism.nim - the drop mechanism is a genuine library
behaviour worth guarding.
- isAlive/aliveCount on the tracker.
Added: docs/tracker_death_events.md (the durable negative, so this is not
re-invented a third time) and test_enemy_tracker_death.nim (13 checks) in place
of the test for the deleted feature.
Guards: test_gun_harness 39/39, test_vbullet_metric 11, test_power_selection 3,
test_adaptive_radar 41/41, test_event_drop_mechanism 6, test_enemy_tracker_death
13, acceptance 12/12, ModularBot compiles.
Replaces melee_scan in the rack. melee_scan spun the radar at the 45 deg/tick
cap unconditionally, so a full 360 deg revolution took 8 ticks and every enemy
was scanned roughly every 8 ticks. The new module starts with the same full
spin, and once it is SURE it has covered every enemy it sweeps back and forth
over only the minimal covering arc of all enemy bearings.
MEASURED, real melees via the bridge, per-enemy onScannedBot counts:
3-bot melee (2 enemies): 29.6 -> 61.1 scans/100 melee-ticks (2.06x)
4-bot melee (3 enemies): 36.0 -> 75.7 scans/100 melee-ticks (2.10x)
Covering-arc widths observed: mostly <90 deg in the 2-enemy case, up to 240 deg
in the 3-enemy case, so the gain shrinks as the arc widens - and at the
ExitTrackWidthDeg=300 fallback it degenerates to exactly the old full spin, so
there is no loss when narrowing would not help.
TRADEOFF, recorded rather than hidden: a wider arc legitimately takes longer to
traverse, so the freshness window costs 5-7 points (fresh<=16: 93-95% vs
98-100%) and more at fresh<=8 (75-76% vs 97-100%). More scans per enemy, at
slightly staler individual fixes.
DESIGN: acquisition spins 360 until every live known enemy was seen within
FreshnessTicks=16 (two revolutions of slack), no new id appeared, and the live
count matches getEnemyCount(); that must hold FreshStreakTicks=3 consecutive
ticks. Tracking then bang-bang sweeps the wraparound-aware covering arc
(350+10 -> 20 through 0) widened by MarginDeg=20 each end, at up to 45 deg/tick.
Fallbacks return to acquisition: any stale enemy, any new id, or an arc >= 300
deg. Enter 270 / Exit 300 gives 30 deg of hysteresis so it cannot flap.
Adds EnemyInfo.lastSeenTick (additive) so coverage is judged on staleness, not
mere knowledge - without it an enemy that slipped behind the sweep would keep
contributing its own stale bearing, which is self-confirming. The offline range
now round-trips that field from the fixture 'lst'.
COMPANION FIX, and it matters: the radar-mode switch used the TRACKER's known
enemy count, so in melee the bot saw one enemy before scanning the second, locked
to 1v1, and the melee radar never ran at all. Now uses getEnemyCount() (server
truth), so melee mode persists until one enemy is genuinely left.
41 new unit checks (wraparound arcs, straddle at 0/360, single/empty enemies,
the 45 deg/tick cap, every phase transition and fallback). melee_scan is kept
but marked DEPRECATED; nothing in the rack imports it.
Non-regression: 33 gun-harness checks, vbullet metric, power selection, and
12/12 offline==online acceptance all pass.
The user spotted this from the game itself: the [config] line always said
gun=HeadOn while the in-game turret and bullet COLOURS varied. The colours are
set at the selection site, so they were truthful and the log was not.
MEASURED, one 2-round battle, same process:
[config] output : 6 lines, ALL gun=HeadOn
tracker selection : Displace 25.1%, HeadOn 20.8%, Pattern 15.5%,
KNN 14.1%, Tsetlin 6.3%, WallBounce 6.1%, ...
Root cause: printConfig did GunNames[bot.currentGun] but EVERY call site ran
before the tick's gun selection - onRoundStarted right after currentGun = 0
(so HeadOn by construction), and the target/radar-change prints. The selection
that sets currentGun is ~490 lines later in the same tick. radarMode and
currentTargetId ARE updated before those sites, which is exactly why the radar
and target columns looked plausible while the gun column did not.
WHERE IT CAME FROM: git history shows commit 2cc2a3b ('cleaner logging')
removed the original printConfig call at the selection site while leaving the
prevGun highlight logic in place. That removal is when the regression appeared -
before it, the bot logged on round start AND on gun switch.
FIX: one emission at the end of the tick, after selectShot has run, gated by a
cfgDirty flag set on gun switch / target change / radar change. The index is
guarded (currentGun may be -1), the round-start line omits the gun field rather
than inventing one, and the existing output contract is preserved (all white,
only changed fields green, enemies= and target= kept).
VERIFIED: before 6 lines all HeadOn; after 782 lines over 2118 ticks with 13
distinct guns and ZERO violations - no printed gun was one that had not been
selected. Counts differ from the selection totals because the log prints only
on change, which is the behaviour the user asked for.
Also removes the 'currentGun = 0' initialisation in onRoundStarted, keeping the
existing -1 sentinel so 'no gun chosen yet' is representable.
TASK 2 - THE FLAKY ACCEPTANCE TEST, root-caused. It was NOT a live/offline
boundary race as suspected. The replay spawned gun 13 (TMSelect) while live has
EnableTmSelector = false and never does. The shared VirtualTracker ring is
ORDER-SENSITIVE, so gun 13's extra 4 bullets/tick shift the ring head and
permute the per-tick RESOLUTION ORDER of every other gun. The learning guns
append observations in resolution order, so their predictions shifted and
produced small hit deltas that moved between runs.
Evidence: the first KNN divergence was at rtick=174 with the SAME resolution
set merely reordered (live ft133,138,139,142,148,150,151 vs offline
ft150,151,133,138,139,142,148); after closing gun 13's ready gate offline the
live and offline KNN traces became BYTE-IDENTICAL (diff empty, 904/904 lines).
Fix: mirror the live rack in the replay. No tick exclusion, no tolerance
loosening. Stability: 5/5 consecutive runs now report 12/12 exact, each with
enemyDied=true - the death boundary is included, not excluded. The proof is
now real rather than a lucky run.
TASK 1 - the tie-break was not random. randomize() was only reached
incidentally through initTsetlinGun(), so a rack without Tsetlin had a fixed
rand() stream and ties always resolved the same way across process restarts.
Added seedSelectorRng() after gun construction, honouring GUN_SELECTOR_SEED.
Evidence: unseeded, 6 separate processes gave different pick sequences; with
GUN_SELECTOR_SEED=42, 3 processes gave identical sequences.
TASK 3 - PRUNING DOES NOT HELP; keep the full rack. 15 PAIRED runs per variant
vs DrussGT, 8 rounds, identical seeds:
baseline 3238 shots 6.18% (events 6.16%) 200 dmg/run
Tsetlin disabled 3522 shots 5.76% (events 5.71%) 197 dmg/run
Tsetlin+Displace 3478 shots 5.46% (events 5.37%) 183 dmg/run
Paired permutation tests: -0.34pp p=0.57 and -0.70pp p=0.21. Per-run
distributions completely overlap (baseline range [2.68, 10.00]; 15/15 and 14/15
runs inside it). A Crazy control showed no separation either. So removing the
measured-worst real performers is neutral-to-slightly-negative, and with
sd ~1.8pp a definitive claim either way would need far more runs.
CORRECTION TO A CLAIM I MADE: the 'virtual metric is INVERTED' finding does NOT
reproduce. Job-24 measured Spearman -0.374; this job measures +0.335 over the
same 13 guns with a different but equally defensible aggregation. Two opposite
signs means the correlation is NOT robustly negative - it is WEAK AND
SIGN-UNSTABLE. The honest statement is that virtual hit rate is a poor ranker,
not an inverted one. The docs assert the inversion and need correcting.
Also adds per-process GUN_STATS_PATH/GUN_SHOTLOG_PATH so concurrent A/B runs do
not clobber each other, and an env-gated GUN_RACK_DISABLE for rack A/Bs. All
default behaviour is unchanged when the env vars are unset.
TASK 2 - power selection, a clear win. bestPower used an ABSOLUTE
MinHitRate = 0.40 bar. Measured per-bin virtual rates (rolling-100 fraction)
show no bin ever clears 40%, so 11 of 14 guns were stuck at bin 0 (power 1.0)
even where higher bins were comparable:
Linear p1.0 44% p1.5 39% p2.0 30% p3.0 29% old bin 0 -> new bin 3
Accel p1.0 44% p1.5 40% p2.0 26% p3.0 29% old bin 1 -> new bin 3
Pattern p1.0 50% p1.5 40% p2.0 27% p3.0 12% old bin 1 -> new bin 2
Replaced with a scale-aware PowerBarFrac = 0.50 (a dimensionless FRACTION of
the gun's own best bin rate). 13 of 14 selections now pick heavier bullets.
Real effect vs DrussGT (8 rounds x 3 runs): hit rate unchanged (7.56% ->
7.47%) but damage dealt +52% (157 -> 239 per run) and rounds end faster.
Same accuracy, half the shots, half again more damage.
TASK 1 - the TM pattern-classifier gun does NOT earn its slot. It was built as
a mixture of experts with a corrected-Granmo TM as a multi-class gate over
HeadOn/Linear/Circular/WallBounce/Accel, labelled by which expert's prediction
was closest to the actual enemy position (an exact, supervised, per-shot
label - no delayed credit). Offline it loses to the best of its OWN experts on
essentially every fixture, and against DrussGT it cost real performance:
baseline (path+relative) 7.56% real hit rate, damage 157
+ power fix 7.47%, damage 239
+ power fix + TM gun 5.59%, damage 133
The gun was selected on 806 ticks and fired 24 real shots at 4.2%.
So the tree ships with EnableTmSelector = false: code and wiring kept intact
for re-enabling, but it is not in the active rack.
Worth recording from the clause dump: the gate DOES latch onto meaningful
structure. On energy-threshold-turner, HeadOn's clauses key on the energy bits
(the rule's own driving variable) while Circular keys on distance/velocity. So
the TM is learning something real and interpretable - it simply cannot beat
'always pick the best expert'. Root cause (INFERRED): the closest-expert label
is noisy because several experts are near-tied, and under the path metric the
winner varies by power bin while the gate sees one shared per-tick input, so a
one-vs-rest gate over a saturated 870-bit clause space has no margin to exploit.
(Zero-padding the 2-frame window was tried first and saturated every clause at
256-755 included literals; alternating the two real frames fixed that.)
Also factors the corrected feedback into an exported tmLearnDir and exports the
encoding/TM primitives; the Tsetlin tests still reproduce the documented
mean=13.8 included literals, so the refactor is behaviour-preserving.
Verified: 33/33 guard checks, tsetlin tests green, metric checks green, new
power-selection guard green (13/14 selections change; relative bar still picks
bin 1 and not bin 3 for a [30,25,12,5]% profile), 12/12 offline==online
acceptance under the shipped default.
Task 1 - the acceptance proof was unrunnable because RecordWorldState was a
compile-time const set to false. It is now a RUNTIME switch
(let RecordWorldState* = existsEnv("TR_RECORD_WORLDSTATE")), default OFF, so
ordinary runs write no fixture, and acceptance_offline_vs_online.nim enables
it for the battle it spawns and clears it afterwards. Restored and run twice:
12/12 deterministic guns match exactly (128-tick and 546-tick battles), with
Tsetlin reported separately as stochastic. Both nimble build variants clean.
Task 2 - does a compact encoding turn the TM's 99.35% into a READABLE rule?
Measured across window sizes (fixed seed, no tuning):
frames TEST acc eff.lits/clause firing clauses counterfactual low/high/mean
10 99.35% 152.8 37 100/24/62.4%
3 95.94% 54.9 35 96/20/58.6%
2 99.48% 39.6 38 95/25/60.7%
1 98.30% 19.2 45 100/24/62.3%
So 2 frames is strictly better than 10 on BOTH axes: +0.13 accuracy for 4x
smaller clauses. The 3-frame dip is non-monotonic and left unexplained rather
than smoothed over.
A readable rule WAS partially recovered. Five clauses carry the exact Gray
form !g10 ^ !g9 ^ !g8; g10 is inert in this data, so the effective rule is the
2-literal proposition !g9 ^ !g8, i.e. energy < 25.6. That is a genuine
threshold in readable propositional form - but at 25.6, NOT the labelled 30,
because 256 is a power-of-two Gray boundary expressible in two literals while
300 needs a longer conjunction. The TM found the nearest SIMPLE threshold.
The honest caveat: that threshold is not the ensemble's decision mechanism.
The counterfactual follow rate (high 24%, mean 62.3%) is statistically
identical at 1, 2 and 10 frames, so compactness did not make the model read
energy - its vote is carried by co-occurring bearing/velocity/heading/wall
literals. Also identified: clauses containing all 11 Gray energy bits are
satisfied at exactly one raw value (50, the dataset floor), so they are
'energy has hit the floor' detectors, not thresholds.
Methodological fix worth keeping: the earlier single-frame counterfactual
wrote energy into all 10 frame slots including the zeroed ones, reviving dead
clauses and producing a spurious 2% high-follow rate. setEnergyFrames now
rewrites only the exposed frames; the corrected figure is 24%.
Gun evaluation previously required a full end-to-end battle (Java server +
battle runner + websocket IPC to 2 bot processes, 50 rounds, ~3.4 min) and
yielded only ~300-900 REAL shots across 13 guns -- far too few to rank
guns, which is why tuning needed many repetitions.
VirtualTracker is already a pure function of (WorldState stream, gun list);
the only reason it needed Java was where WorldState came from. So the range
replays a seq[WorldState] through the SAME tracker: offline and online
scores are the same metric by construction, not an approximation.
ACCEPTANCE TEST (the point of the whole thing): record one live round, replay
it offline, compare per-gun virtual hit rates. 12/12 deterministic guns match
EXACTLY, reproduced twice. Tsetlin is compared separately because tmLearnOne
calls rand(). Getting to 12/12 exposed two real ordering quirks in the live
loop: run() calls go() before the aim/fire block, so tickBullets resolves
against the NEXT tick's scan while the prediction used the previous one; and
if the target dies during that go() the final tick's spawn+resolution is
skipped entirely. The recorder emits an end marker for the second case.
The 5th (selected-gun) predict call was verified to be a no-op.
Measured cost: 8 fixtures (1770 ticks, ~92k virtual bullets, 13 guns) replay
in 2.9 s, ~32k virtual bullets/s -- roughly 70x faster and 100x more samples
than a live gauntlet.
Also adds a per-tick WorldState recorder behind const RecordWorldState
(default off, mirrors the ShotLog idiom) which records the state the bot
ACTUALLY builds, staleness included, rather than true positions -- recording
the latter would hand the guns perfect information and produce flattering
scores.
9 new guard checks (33 total, all passing), including fixture round-trip,
replay determinism, stationary->HeadOn 100%, constant-velocity->Linear>HeadOn,
and the energy-threshold turner crossing at t=41.
Measured, not assumed. With the gate temporarily opened to 20 deg, every
real shot was logged (tick, angle error at fire time, distance, power,
hit) across 3 gauntlets: 2611 shots, 57.3% aggregate. Findings:
- The geometric cone atan(BotRadius/d) is directionally confirmed but a
WEAK lever: even at 0.0-0.1 deg error the hit rate at 400-600px is only
~53-57%, because PREDICTION error dominates alignment error.
- Real effect of tightening the gate: 57.9% -> 68.0% aggregate hit rate
(fixed 0.1 deg), not the 76.9% previously reported -- that was a
high-variance draw (per-rep 62.8/66.4/77.2%).
- The shipped range-aware gate (SafetyFactor 0.6) does NOT beat the fixed
2.0 deg gate on hit rate (55.8% vs 57.9%, ~1.5 sigma, inside noise). It
fires 22-28% more shots and therefore lands more total hits (~509 vs
~434 per rep). No per-adversary score delta exceeded the 300-point
run-to-run noise band, so no config is demonstrably better on score.
Shipped anyway because it is strictly more expressive (a fixed threshold is
the special case), tunable from one const, and physically motivated, but
the honest verdict is recorded in-code: the gate is not the bottleneck.
AimThresholdDeg is removed; shouldFire now takes distPx. Degenerate or NaN
distance falls back to the ceiling rather than dividing by zero.
Also adds a per-shot logger to ModularBot behind 'const ShotLog' so the
measurement above is reproducible, and 10 new guard checks (24 total, all
passing) covering monotonicity, clamping, formula, perfect alignment,
gross misalignment and degenerate distance.
Cross-checked against the server source: the gun fires BEFORE the turn is
applied, so the logged angle error is the true departure error, and
fireAssist auto-aim is off (unset by the Nim API and forced false by
setAdjustRadarForGunTurn).
Wave queues (guess_factor, decay_gf, knn_gun): predict() stored ONE wave
per tick while onResult() popped one per resolved bullet (~4/tick), so the
queue drained to empty within a few dozen ticks, ~3 of every 4 resolutions
returned without learning, and the survivor paired with a same-tick wave
(bearingDelta ~= 0) pinning the histogram at centre. PROOF: GF.vHits ==
HeadOn.vHits and DecayGF.vHits == HeadOn.vHits byte-for-byte in every one
of 50 rounds — the guns had degenerated to HeadOn.
Now each gun keeps a per-bin FIFO with an O(1) head cursor. At most one
push per (tick, bin) so the fire site's 5th predict() call is a no-op, and
onResult pops the oldest wave of its OWN bin via e.bulletPower. Aiming
math untouched (it was already correct: 0 deg = East, CCW+).
maxBullets 2048 -> 8192: the rack spawns 52 bullets/tick so the ring wrapped
every ~39 ticks while a long power-3 shot needs ~90, silently discarding
unresolved bullets and biasing every measured hit rate by range. Added a
droppedBullets counter so a future overflow is measurable, and wavePushes/
waveStarved counters on the three guns. After the fix: vDropped = 0 and
vStarved = 0 across all 48 recorded rounds.
fitnessFor is now exported, deterministic (enemies iterated in ascending id
order) and shared by the selector and the stats dump, replacing a hand-rolled
merge in ModularBot that never advanced its window head.
Round lines gain additive keys: vDropped, vStarved.
Attribution is proven, not guessed: the server assigns a per-round-unique
bulletId (GunEngine.nextBulletId) and stamps the same id on BulletFired,
BulletHitBot, BulletHitWall and BulletHitBullet. Keep a FIFO of fired gun
ids, stamp bulletId -> gunId on onBulletFired, resolve through that map.
Hits are deferred when onBulletHit precedes onBulletFired in the same
turn (client dispatches priority 70 > 60), which recovered 14
unattributed hits. 99.9% of shots and 99.8% of hits attributed.
Stats lines now carry per-gun realShots/realHits/realHitRate; the old
keys and round-level totals are unchanged.
bestPower: a gun with zero observations in every bin previously returned
the HIGHEST bin (power 3.0) because an empty bin satisfied the
'count == 0' clause on the first countdown iteration. Cold guns now
return the lowest bin as the docstring always claimed. Warm-gun path
untouched.
- bestGun: replace first-index-wins argmax with random pick among guns
within TieMargin (2%) of best rate. HeadOn at index 0 was silently
winning every tie, starving Tsetlin/Linear/etc.
- MinObsBeforeCompete 15 -> 50 (Pattern entered competition on noise)
- add MinHitRateFloor 0.10: if no gun clears it, fall back to HeadOn
instead of selecting the best of a bad field
- ModularBot: remove AntiSurfer gun (0% virtual hit rate everywhere),
14 -> 13 guns, renumber ids and selection counters
- 30-tick cooldown after ghost-stuck/timeout ram exit prevents re-entry loop
- enemy_tracker.update() skips dead bots to prevent same-tick scan resurrection
- TFIL graphics cleared when ramming is active movement
- [config] logs: white base with green-highlighted changes only
- [ram:enter] logs trigger reason and key values on false→true transition
- [death] and [target-invalid] logs retained for diagnostics
PhantomMeteor:
- Ram finisher: charge at enemy when <200px and their energy <10
- Ram opportunity: charge when <60px and we have >20 energy advantage
- Integrated gunheat tracker for 1-2 tick earlier wave detection
- Distance control: smooth linear ramp toward preferred engagement distance
- Phantom range expanded 150→250px to catch closer threats
WaveSurfer:
- Wall-aware dodge bin selection: penalize bins leading off-arena
- Dodge timing: predict future position 15 ticks ahead for safety
- Distance control: radial blend when outside deadband (350±50px)
- Wall escape: invert strafe if pushing further into wall, blend toward center
ModularBot:
- Wired KNN gun (purple/magenta)
- Shadows tracked for movement (safer GF prediction)
- Bullet lifecycle management (onBulletFired/onBulletHitBot/onBulletHitWall)
- Unified phantom_meteor movement (wave_surfer unplugged)
- Config logging on round start + gun switch
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Implements VirtualBodyTracker (wave-based hit/miss scoring) instead of EMA damage accumulation. Movement switching now happens every tick, not every 3 rounds. Also refactors radar colors to dark teal (#004444/#0D4D4D) for faint visibility.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Adds WaveSurferModule as second movement option
- Tracks damage per round via onHitByBullet, EMA decay α=0.5
- Every round after round 2: switches to lower-damage movement
- Body color signals active mover (orange=phantom, blue=wave)
- vs SpinBot 10r: ModularBot 1412 vs SpinBot 650
- vs WaveSurfer 10r: ModularBot 1858 vs WaveSurfer 10
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>