07f6f3af3f22f6ec9e6d3ac0949dfa1c456395db
261 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
07f6f3af3f |
Tsetlin gun: NO configuration adapts faster than random feedback
The user's goal was "a TM gun that can learn fast and generalize better". Swept offline over the real DrussGT fixtures (no live battles) by coordinate descent, one lever at a time, with a SHUFFLED-FEEDBACK CONTROL - a TM trained on randomised targets. That control is what settles the question. Final confirmation, 4 seeds each (~74,600 first-100-tick bullets per config): config EARLY(first 100) OVERALL Shuf_w3 (RANDOM feedback) 23.9% 20.0% win3_s1.1 (best real TM found) 23.7% 20.2% Shuf_w10 (RANDOM feedback) 23.1% 20.0% win3_st100 (prior job's edit) 23.0% 20.1% win3_off (TM correction ~= 0) 22.7% 20.2% def_w10 (shipped default) 22.1% 20.3% Linear (deterministic reference) 34.0% 24.3% The best real config beats the default early (23.7% vs 22.1%, non-overlapping per-seed ranges, z=+7.34, p=2e-13) - but its own SHUFFLED control scores 23.9%, i.e. HIGHER, z=-0.91, p=0.37. Random targets do at least as well. So the early gain is not learning. Per-lever screens were flat: TM_N_CLAUSES 25/50/100/200 all 23.0% early, completely flat; TM_N_STATES 4/32/100 all ~22-23% (unstable across seeds); TM_S mildly monotonic (lower better early); TM_T flat; TM_WINDOW_SIZE 2/3/10 all within noise of each other and of the shuffled control. Two further findings: - The TM-off ablation (correction ~= 0) scores 22.7%/20.2%, essentially the same as TM-on. The TM's correction is near-zero-mean noise; the gun's one-shot internal linear baseline accounts for its accuracy. - The TM gun is 10.3 pp behind Linear early and 4.1 pp behind overall. That deficit is in the BASELINE MODEL (LinearGun iterates flight time; this gun does not), not in the TM hyper-parameters. Tuning knobs cannot close it. Conclusion: do not tune TM hyper-parameters further. Either the input representation or the prediction target is what needs to change - the shuffled control shows the TM is not extracting target information beyond its baseline. Defaults left UNCHANGED (window=10/states=32/S=1.5/T=25/clauses=50); an uncommitted prior edit (window=3/states=100) was reverted as unsupported. Hyper-parameters are now compile-time overridable (-d:TM_WINDOW_SIZE=3 etc.) so future sweeps need no gun edit. NOT MEASURED: real hit rate vs DrussGT (offline only by design). The repo's own docs/gun_rack_analysis.md 2 reports the virtual metric is a sign-unstable ranker of real hit rate, so the comparison against "Linear 10.7% real" is not direct - whether the TM is competitive live is INFERRED-unknown, not measured. Guards: test_gun_harness 39/39, test_vbullet_metric, test_power_selection, test_tsetlin_gun, test_tm_pattern_learning all green. |
||
|
|
fb36a0a685 |
tracker: corpses do not exist - revert the fix and retire the workaround
The belief "BotDeathEvent never reaches ModularBot, so enemyTracker keeps dead
enemies alive forever" was written into a code comment and then believed twice.
It is FALSE. Measured in a 7-bot melee with a per-tick probe comparing
enemyTracker's alive count against the server's getEnemyCount():
metric 1.3.1 (20 rd) 0.35.5 (15 rd)
observed enemy deaths 83 68
...non-round-ending 83 (100%) 66 (97%)
ekBotDeath events DROPPED 0 0
max dispatch lag (turns behind) 1 1
phantom ticks 1 / 16,820 1 / 12,596
MAX CORPSE LIFETIME 0 ticks 0 ticks
victims still alive at round end 0 0
onBotDeath fires for every death, including non-round-ending ones. The
API-level event-drop mechanism IS real (test_event_drop_mechanism.nim proves
it: ekBotDeath is not in isCritical and MAX_EVENTS_AGE=2) - the bot simply
never falls far enough behind for it to trigger (max lag 1 turn).
Removed:
- reconcileWithServer + ReconcilePersistTicks/mismatchTicks/sawServerAlive
(uncommitted, and ON BY DEFAULT despite the premise being false). Its own
comment admitted a shorter window once KILLED A LIVE ENEMY ("it fired three
more times after the tracker marked it dead") - a latent mis-prune path
defending against a bug that does not exist.
- The radar's CorpseTicks=40 filter and the same-class age>60 filter in
recordRadarStats, both carrying the false comment. Removal changes no real
behaviour: buildState feeds the radar enemyTracker.allAlive(), so a dead
enemy never reaches computeScan.
Kept:
- The TR_TRACKER_PROBE instrument (default OFF), which produced the table above.
- test_event_drop_mechanism.nim - the drop mechanism is a genuine library
behaviour worth guarding.
- isAlive/aliveCount on the tracker.
Added: docs/tracker_death_events.md (the durable negative, so this is not
re-invented a third time) and test_enemy_tracker_death.nim (13 checks) in place
of the test for the deleted feature.
Guards: test_gun_harness 39/39, test_vbullet_metric 11, test_power_selection 3,
test_adaptive_radar 41/41, test_event_drop_mechanism 6, test_enemy_tracker_death
13, acceptance 12/12, ModularBot compiles.
|
||
|
|
c091bf3c34 |
harness: upgrade to server 1.3.1, keep 0.35.5 selectable, re-baseline
All prior measurements ran on server 0.35.5. The default is now the current 1.3.1 jar, with the legacy jar kept and switchable via TR_SERVER_JAR (no code edit). test_gauntlet_5bots.nim no longer clobbers a caller's TR_SERVER_JAR - it used to putEnv() unconditionally, so an override was silently ignored. RE-BASELINE (controlled RulesProbe battle, stationary bot, powers 0.1/0.5/1/2/3): dimension 1.3.1 0.35.5 verdict bullet damage per hit 0.4/2/4/10/16 identical SAME bullet speed (20-3p) within noise within noise SAME post-fire gun heat (1+p/5) identical identical SAME cooling 0.1/tick 0.1/tick SAME bulletDamage SCORE exactly 100/round 104..113/round DIFFERENT bulletKillBonus (20%) 20/round 20..23/round DIFFERENT LOUD FINDING - a SCORING rule changed, physics did not: 0.35.5 credits OVERKILL to bulletDamage (the killing bullet's full damage even past 0 energy); 1.3.1 caps it at the energy actually removed. Every 0.35.5 score is therefore inflated ~5-6%, and bulletKillBonus inherits the inflation. Gauntlet totals shift accordingly (SittingDuck 1936 -> 1800, WaveSurfer 1886 -> 1669). Consequence: score-based numbers recorded on 0.35.5 are NOT comparable to 1.3.1. Our gun A/Bs used real HIT RATE, not score, so those conclusions stand. Runner 1.0.2 (unchanged, no newer one on the box) is measured compatible with the 1.3.1 server. Note TrBattleCapture uses the runner's EMBEDDED server, which is 1.0.2 - so the capture path still runs an older engine than the gauntlet. Also re-ran acceptance_offline_vs_online on the new default: 12/12. |
||
|
|
18f778056b |
gun selector: hysteresis measured NEGATIVE, shipped at the lightest setting
Hypothesis under test: the selector chatters (~54 switches/100 ticks) and that chatter suppresses firing, so committing to the virtual-best gun should raise real hit rate. MEASURED AGAINST THE REAL DRUSSGT: it does not. setting switches/100t real hit % dmg/run shots/run no hysteresis 0/0 54.29 7.02% 217 244.8 light 10/0.05 1.95 6.22% 191 235.3 moderate 30/0.15 1.33 5.10% 156 233.9 aggressive 60/0.30 - 5.72% 175 242.5 (16 runs x 7 rounds per config except aggressive = 8; server-side events sidecar; permutation test baseline-vs-moderate p=0.002, baseline-vs-light p=0.18.) Hysteresis cuts chatter 28-54x but every variant fires slightly FEWER shots and deals LESS damage than baseline. Mechanism [INFERRED, consistent with docs/gun_rack_analysis.md 2/4]: the per-tick random tie-break among the tied band is a hedge, and hysteresis destroys it by committing to the virtual-best gun - which is not the real-best, because the virtual metric is a weak, sign-unstable ranker. The chattering was load-bearing. Shipped: GunDwellTicks=10, GunSwitchMargin=0.05 (GUN_SELECTOR_DWELL / GUN_SELECTOR_MARGIN) - the only setting within the baseline's run-to-run spread. GUN_SELECTOR_DWELL=0 GUN_SELECTOR_MARGIN=0 reproduces the pre-change selector exactly. Seam: VirtualTracker, which already owns the other selection state (fitness, the relative floor's peakRateRef), so the bot needs no new fields. bestGun and chooseFromFit stay pure/memoryless, which is why the existing random-tiebreak test needed no change. Guards: test_gun_harness 39/39 (33 original + 6 new hysteresis checks), test_vbullet_metric, test_power_selection, acceptance_offline_vs_online 12/12. |
||
|
|
68e0375be2 |
feat(radar): adaptive melee radar sweeps only the arc containing all enemies
Replaces melee_scan in the rack. melee_scan spun the radar at the 45 deg/tick cap unconditionally, so a full 360 deg revolution took 8 ticks and every enemy was scanned roughly every 8 ticks. The new module starts with the same full spin, and once it is SURE it has covered every enemy it sweeps back and forth over only the minimal covering arc of all enemy bearings. MEASURED, real melees via the bridge, per-enemy onScannedBot counts: 3-bot melee (2 enemies): 29.6 -> 61.1 scans/100 melee-ticks (2.06x) 4-bot melee (3 enemies): 36.0 -> 75.7 scans/100 melee-ticks (2.10x) Covering-arc widths observed: mostly <90 deg in the 2-enemy case, up to 240 deg in the 3-enemy case, so the gain shrinks as the arc widens - and at the ExitTrackWidthDeg=300 fallback it degenerates to exactly the old full spin, so there is no loss when narrowing would not help. TRADEOFF, recorded rather than hidden: a wider arc legitimately takes longer to traverse, so the freshness window costs 5-7 points (fresh<=16: 93-95% vs 98-100%) and more at fresh<=8 (75-76% vs 97-100%). More scans per enemy, at slightly staler individual fixes. DESIGN: acquisition spins 360 until every live known enemy was seen within FreshnessTicks=16 (two revolutions of slack), no new id appeared, and the live count matches getEnemyCount(); that must hold FreshStreakTicks=3 consecutive ticks. Tracking then bang-bang sweeps the wraparound-aware covering arc (350+10 -> 20 through 0) widened by MarginDeg=20 each end, at up to 45 deg/tick. Fallbacks return to acquisition: any stale enemy, any new id, or an arc >= 300 deg. Enter 270 / Exit 300 gives 30 deg of hysteresis so it cannot flap. Adds EnemyInfo.lastSeenTick (additive) so coverage is judged on staleness, not mere knowledge - without it an enemy that slipped behind the sweep would keep contributing its own stale bearing, which is self-confirming. The offline range now round-trips that field from the fixture 'lst'. COMPANION FIX, and it matters: the radar-mode switch used the TRACKER's known enemy count, so in melee the bot saw one enemy before scanning the second, locked to 1v1, and the melee radar never ran at all. Now uses getEnemyCount() (server truth), so melee mode persists until one enemy is genuinely left. 41 new unit checks (wraparound arcs, straddle at 0/360, single/empty enemies, the 45 deg/tick cap, every phase transition and fallback). melee_scan is kept but marked DEPRECATED; nothing in the rack imports it. Non-regression: 33 gun-harness checks, vbullet metric, power selection, and 12/12 offline==online acceptance all pass. |
||
|
|
d3b3c28cdf |
fix(adversaries): migrate to bot-api 1.0.7 - kills an intermittent crash that corrupted measurements
Four of the five adversaries imported the OLD package (tankroyale_botapi 1.0.1); only SittingDuck used robocode_tankroyale_botapi 1.0.7, which is what the rest of the repo requires. A previous report claimed OscillatorBot was already on 1.0.7 - that was WRONG, and OscillatorBot turned out to crash the MOST (8 SIGSEGVs in the first reproduction, 15 in its historical /tmp logs). THE CRASH, reproduced with an identical stack in every case: botThreadEntry -> run -> adversary run -> go -> dispatchPendingEvents -> tankroyale_botapi-1.0.1/event_queue.nim(89) addEvent -> realloc/rawDealloc -> SIGSEGV Counts, old API: 60 melee battles x 8 rounds gave RandomMover 1, PatternMover 3, WaveSurfer 0, OscillatorBot 8; 6 battles x 6 rounds vs SittingDuck gave 4/2/0/3. ROOT CAUSE: the main->bot event hand-off. 1.0.1 passes a lock-protected seq[BotEvent] (signalTick writes gPendingEvents, dispatchPendingEvents copies it under lock). 1.0.7 uses a Channel[seq[BotEvent]] (send(move(pending)) / tryRecv). The old path copied string-bearing BotEvent payloads across threads every tick, churning ORC refcounts on the shared heap until the freelist was corrupted. 1.0.7's own source documents this as the gdb-confirmed fix. WHY IT MATTERED MORE THAN IT LOOKED: the crash silently corrupted measurements. Against a stationary duck, crash contamination inflated WaveSurfer's rest fraction from 12.4% (clean) to 20.7%; in a focused run the server logged 'Bot left: OscillatorBot' while the game continued and its score stopped growing. So every gauntlet run tonight was fighting adversaries that were partially dead - which is a second, independent reason the user's instinct that these bots were bugged was correct, and why they should not be used as a measurement baseline. (The per-gun REAL hit rates are unaffected: those came from DrussGT battles.) FIX: all four migrated to robocode_tankroyale_botapi 1.0.7. NO API adaptations were needed beyond the module rename - every symbol these bots use is identical in 1.0.7, verified by diffing the two packages (constants/utils/json_parse/ schemas semantically identical; the movement and intent procs in bot.nim are byte-identical). The .nimble files now require robocode_tankroyale_botapi. VERIFIED: 120 melee battles x 8 rounds plus 6x6 vs SittingDuck -> 0 SIGSEGV in all four stderr logs (0 bytes). Behaviour unchanged: sub-1% absolute drift in mean speed, rest fraction, reversal rate, mean range and perpendicular fraction, all within run-to-run spread; the one >=3-sigma flag (WaveSurfer perpendicular relative to DrussGT) was isolated against a stationary opponent and shown to be the chaotic closed loop, not the migration. test_wavesurfer_velocity passes 7/7. NOT migrated, reported only: GotoTest_garage, OscillatorBot_garage (archived copy), PPO_Bot_garage, QBot_garage, SAC_LSTM_Bot_garage - older experiment garages, left alone deliberately. |
||
|
|
ab35540035 |
fix(logging): [config] reported a stale gun - it printed before selection ran
The user spotted this from the game itself: the [config] line always said
gun=HeadOn while the in-game turret and bullet COLOURS varied. The colours are
set at the selection site, so they were truthful and the log was not.
MEASURED, one 2-round battle, same process:
[config] output : 6 lines, ALL gun=HeadOn
tracker selection : Displace 25.1%, HeadOn 20.8%, Pattern 15.5%,
KNN 14.1%, Tsetlin 6.3%, WallBounce 6.1%, ...
Root cause: printConfig did GunNames[bot.currentGun] but EVERY call site ran
before the tick's gun selection - onRoundStarted right after currentGun = 0
(so HeadOn by construction), and the target/radar-change prints. The selection
that sets currentGun is ~490 lines later in the same tick. radarMode and
currentTargetId ARE updated before those sites, which is exactly why the radar
and target columns looked plausible while the gun column did not.
WHERE IT CAME FROM: git history shows commit
|
||
|
|
1e8f0d342a |
fix(adversaries): the launchers ran STALE binaries - this is why the earlier fix never took effect
P0. Three of the five launchers ran ./<Bot> (a tracked binary at the bot root) while config.nims sets outdir=out and both the test framework's compileBots and a manual 'nim c src/<Bot>.nim' write to out/. SittingDuck and OscillatorBot correctly ran ./out/<Bot>; RandomMover, PatternMover and WaveSurfer did not. cmp confirms the root and out binaries differed for all three. Consequence: the previous session's adversary fixes were compiled into out/ and never executed. Every gauntlet and every capture ran the OLD code. This is almost certainly why the user's instinct that these bots were still bugged was correct while the code claimed otherwise. Fixed by pointing all five launchers at ./out/<Bot>, and by deleting the three stale root binaries so the trap cannot recur. Verified end to end through the booter: WaveSurfer went from standing still 96.2% of ticks with a 1398-tick longest standstill, to rest 12.3% / mean speed 6.69 / longest zero run 18 / perpendicular 0.845. Also honours GUN_STATS_PATH in test_gauntlet_5bots.nim (same knob ModularBot reads) so pooled gauntlet runs append to one file instead of clobbering the default. NOTE for a follow-up: the out/ binaries are still TRACKED build artifacts, which is the same class of hazard that caused this. Untracking them (as was done for ModularBot_garage/ModularBot) would remove the failure mode entirely. |
||
|
|
c214abcfa8 |
fix(adversaries): repair four of the five sparring bots
The user suspected these were bugged. They were, and the verdicts are not uniform - three genuinely broken, one merely sloppy, one fine: - WaveSurfer: GENUINELY BUGGED, worst of the five. (a) The enemy velocity decomposition was sin/cos SWAPPED - enemyVx used sin and enemyVy used cos, while Tank Royale is 0 deg = East, CCW+, so it must be cos for X and sin for Y. Its linear-prediction gun was aiming at a reflected position. (b) The wall escape flipped strafeDir on EVERY tick the bot was inside the wall margin, so instead of turning away it flip-flopped in place: measured standing still (speed < 0.5) for 96.2% of ticks with a longest continuous standstill of 1398 ticks. Fixed with a hysteretic wall-escape selection plus a corner escape, dead enemyLastDir removed, and per-round state reset. AFTER, measured through the booter: rest 12.3%, mean speed 6.69, full speed 79.7%, longest zero run 18, perpendicular 0.845 / radial 0.012 - it now actually strafes. Gun sanity: lead error 1.0 px vs 106 px for head-on on a constant-velocity target; lead gun 45.8% hits vs 29.3% for head-on. - PatternMover: GENUINELY BUGGED. Real deadlock - it decremented its step counter by the REQUESTED amount while issuing setTargetSpeed(8), so against a wall the counter never reached 0, advanceStep never ran and it was stuck forever (309-tick standstill). Now counts down by ACTUAL distance/turn with a STALL_LIMIT watchdog and steers toward the arena centre. Standstill 309 -> 19 ticks; full-speed ticks 10.0% -> 28.4%. - OscillatorBot: GENUINELY BUGGED, milder. No wall handling at all, so it ground along walls 53.4% of ticks and could pin in a corner. Added wall steering that preserves the fixed 25-tick reversal cadence. Wall-band 53.4% -> 18.6%, mean wall distance 72 -> 119. - RandomMover: merely sloppy, not broken. Its turn intent saturated against the speed-dependent limit (18.4% of moving ticks clamped) and the fire gate was a very loose 10 deg. Now clamps to calcMaxTurnRate and fires within 3 deg. Saturation 18.4% -> 3.9%. - SittingDuck: FINE. Speed 0 for 100% of ticks, zero shots. Left untouched - it is a duck by design. Adds test_wavesurfer_velocity.nim, a direct assertion that the decomposition is cos/sin and explicitly NOT the swapped form (7 cases). KNOWN ISSUE, not fixed: RandomMover/PatternMover/WaveSurfer import tankroyale_botapi 1.0.1 and intermittently SIGSEGV in tankroyale_botapi/event_queue.nim:89 addEvent, freezing the bot for the rest of the battle. It reproduces on old and new code and never occurs for SittingDuck/ OscillatorBot, which import robocode_tankroyale_botapi 1.0.7. Migrating the three to 1.0.7 would likely fix it and is worth doing - it is a real reliability risk for these as sparring partners. |
||
|
|
e6653199bb |
fix(shim): make the generated DrussGT bot dir runnable by hand
Running /tmp/tr_bots/DrussGT/DrussGT.sh manually failed with: BotException: Required bot property 'name' is missing. The Tank Royale Java bot API reads bot identity from BOT_* environment variables (EnvVars: BOT_NAME, BOT_VERSION, BOT_AUTHORS, BOT_DESCRIPTION, BOT_HOMEPAGE, BOT_COUNTRY_CODES, BOT_GAME_TYPES, BOT_PLATFORM, BOT_PROG_LANG, BOT_INITIAL_POS). When the TR booter launches the directory it supplies those, derived from the .json - which is why every headless battle worked - but launching the script directly supplies nothing, so the API rejects the handshake. make_botdir.sh now exports them in the generated <dir>.sh (kept in sync with the .json it writes), so the script works standalone as well as under the booter. Verified against a server started for the test: before, the exact error above; after, 'Connected to: ws://localhost:4599' and the server logs 'Bot joined: DrussGT 3.1.4159'. |
||
|
|
013b9fe01e |
docs: correct the overstated 'inverted metric' claim; record tonight's fixes
Three corrections, all prompted by later measurements: 1. The virtual-vs-real rank correlation is NOT robustly negative. Six independent Spearman measurements now exist (-0.374, +0.335, +0.522, -0.371, -0.073, -0.037) and the sign flips on large samples, so it is near zero on average. The honest headline is that virtual hit rate is a POOR RANKER, not an inverted one. The report said 'not weak - it is inverted' in six places; it now says so in none. The practical conclusion (do not trust it for ranking) is unchanged; the mechanism claimed was wrong. 2. The offline==online acceptance is FIXED, not flaky. Root cause was that the replay spawned gun 13 (TMSelect) while the live rack has it disabled, and the shared VirtualTracker ring is ORDER-SENSITIVE, so gun 13's extra 4 bullets/tick permuted the per-tick resolution order for every other gun and shifted the learning guns' observations. After closing gun 13's ready gate offline the live and offline KNN traces are byte-identical (904/904 lines, empty diff). 5/5 consecutive runs now report 12/12 exact with the death boundary included. Recorded with the lesson: a flaky proof was hiding a real bug. Also records the general A/B confound - disabling a gun removes its 4 spawns/tick from the shared ring, perturbing resolution order for the rest. 3. Pruning was tested and does NOT help, so the verdict for Tsetlin and Displace changes from an implied drop to BELOW OVERALL - KEEP. 15 paired runs: baseline 6.18%, Tsetlin-off 5.76%, Tsetlin+Displace-off 5.46%; paired permutation p=0.57 and p=0.21; distributions completely overlap; a non-surfer control showed no separation. Being below average does not justify removal. Also records the tie-break randomness fix, and quotes run counts with every rate (6.95% over 13 runs vs 6.18% over 15 runs, same binary) rather than presenting a single figure as definitive. |
||
|
|
4cd5618435 |
fix(test): repair the flaky offline==online acceptance; make the tie-break truly random
TASK 2 - THE FLAKY ACCEPTANCE TEST, root-caused. It was NOT a live/offline boundary race as suspected. The replay spawned gun 13 (TMSelect) while live has EnableTmSelector = false and never does. The shared VirtualTracker ring is ORDER-SENSITIVE, so gun 13's extra 4 bullets/tick shift the ring head and permute the per-tick RESOLUTION ORDER of every other gun. The learning guns append observations in resolution order, so their predictions shifted and produced small hit deltas that moved between runs. Evidence: the first KNN divergence was at rtick=174 with the SAME resolution set merely reordered (live ft133,138,139,142,148,150,151 vs offline ft150,151,133,138,139,142,148); after closing gun 13's ready gate offline the live and offline KNN traces became BYTE-IDENTICAL (diff empty, 904/904 lines). Fix: mirror the live rack in the replay. No tick exclusion, no tolerance loosening. Stability: 5/5 consecutive runs now report 12/12 exact, each with enemyDied=true - the death boundary is included, not excluded. The proof is now real rather than a lucky run. TASK 1 - the tie-break was not random. randomize() was only reached incidentally through initTsetlinGun(), so a rack without Tsetlin had a fixed rand() stream and ties always resolved the same way across process restarts. Added seedSelectorRng() after gun construction, honouring GUN_SELECTOR_SEED. Evidence: unseeded, 6 separate processes gave different pick sequences; with GUN_SELECTOR_SEED=42, 3 processes gave identical sequences. TASK 3 - PRUNING DOES NOT HELP; keep the full rack. 15 PAIRED runs per variant vs DrussGT, 8 rounds, identical seeds: baseline 3238 shots 6.18% (events 6.16%) 200 dmg/run Tsetlin disabled 3522 shots 5.76% (events 5.71%) 197 dmg/run Tsetlin+Displace 3478 shots 5.46% (events 5.37%) 183 dmg/run Paired permutation tests: -0.34pp p=0.57 and -0.70pp p=0.21. Per-run distributions completely overlap (baseline range [2.68, 10.00]; 15/15 and 14/15 runs inside it). A Crazy control showed no separation either. So removing the measured-worst real performers is neutral-to-slightly-negative, and with sd ~1.8pp a definitive claim either way would need far more runs. CORRECTION TO A CLAIM I MADE: the 'virtual metric is INVERTED' finding does NOT reproduce. Job-24 measured Spearman -0.374; this job measures +0.335 over the same 13 guns with a different but equally defensible aggregation. Two opposite signs means the correlation is NOT robustly negative - it is WEAK AND SIGN-UNSTABLE. The honest statement is that virtual hit rate is a poor ranker, not an inverted one. The docs assert the inversion and need correcting. Also adds per-process GUN_STATS_PATH/GUN_SHOTLOG_PATH so concurrent A/B runs do not clobber each other, and an env-gated GUN_RACK_DISABLE for rack A/Bs. All default behaviour is unchanged when the env vars are unset. |
||
|
|
19410164f1 |
docs: definitive gun-rack report on real measured numbers
Replaces the stale 2026-09-20 docs, which predated per-gun real attribution, the offline gun range and the DrussGT boss, and whose verdicts were built on virtual hit rates that turned out to be ANTI-correlated with reality. docs/gun_rack_analysis.md (751 lines) covers: the test infrastructure described honestly (offline range with its flaky-acceptance caveat, the 20 fixtures and what each set is good for, the live boss, and the A/B methodology of per-run server-side real hit rate with an explicit overlap test); the virtual-vs-real metric lesson with Spearman -0.374 and the point-vs-path A/B; per-gun real performance and the 16-rule ranking A/B; the offline per-fixture gun matrix; KEEP/MARGINAL/BELOW verdicts; and the five root-cause bugs with before/after numbers. docs/gun_rack_summary.md (58 lines) is the verdict table plus top actions. The '~230 point' score-noise band that has been steering methodology all night was re-derived from the artifacts rather than asserted: the 13 shipped-config run scores span 175-526, s.d. ~105, i.e. a ~210-point 2-s.d. band. Caveats recorded verbatim rather than softened: the offline==online acceptance is flaky (typically 11/12 on unmodified HEAD), fixtures are perfect-information and therefore optimistic vs live play, per-gun real N is small so single-gun ordering is indicative, the headline numbers come from ONE wave-surfer adversary, and HeadOn must stay despite being lowest because it is the floor fallback (disabling it: 5.08% / 175 dmg vs 6.95% / 251 dmg). |
||
|
|
2c94dc221a |
test(selector): 16 ranking rules A/B'd against the boss - none beat the shipped config
Added runtime-tunable ranking knobs to the selector, all defaulting to the shipped values so behaviour is byte-identical when unset: GUN_SELECTOR_WINDOW, MINOBS, TIE, FLOOR, POOL, RANK, SHRINK, SEED. rankScore supports mean, Wilson lower bound, UCB, Thompson and shrinkage. Also fixed hitRate's most-recent-N read for sub-WindowSize windows (windowHits). RESULT: NO candidate credibly beat the shipped config. 13 runs x 8 rounds vs DrussGT, 3612 shots, base 6.95% at 251 dmg/run; every candidate's per-run interval overlaps base, and the nominal 'winners' are <=0.6 SE apart on far fewer shots. Kept the shipped default. Valid outcome, recorded plainly. THE FINDING THAT MATTERS MORE: the virtual-bullet ranking is ANTI-correlated with real hit rate - Spearman ~ -0.37 for the shipped config. It is not merely weak, it is INVERTED. The guns with the highest VIRTUAL rates have among the lowest REAL rates: Tsetlin 12.9% virtual / 5.8% real, WallBounce 12.9 / 6.2, StopShot 12.6 / 6.1, AvgLead 12.3 / 7.0 - while Linear sits at 10.2 virtual / 10.7 real and KNN at 7.5 / 9.0. So what carries the selector is the floor/tie HEDGING, not the ranking: removing the floor drops us to 5.08% / 175 dmg. That also kills the 'exploration' hypothesis - every gun spawns virtual bullets every tick, so sampling is uniform and the bottleneck is SIGNAL QUALITY, not under-sampling. FINAL PER-GUN REAL HIT RATE vs DrussGT (13 runs, 3612 shots, overall 6.95%): Linear 10.7 | Circular 9.9 | KNN 9.0 | Pattern 8.6 | Accel 7.3 | AvgLead 7.0 GuessFactor 6.9 | DecayGF 6.4 | WallBounce 6.2 | StopShot 6.1 | Tsetlin 5.8 Displace 5.3 | HeadOn 5.2 Keep: Linear, Circular, KNN, Pattern, Accel, AvgLead. Marginal: GuessFactor, DecayGF, WallBounce, StopShot. Below overall: Tsetlin, Displace, HeadOn - but HeadOn must STAY as the floor fallback, since disabling the floor measurably hurt. CORRECTION TO A CLAIM I HAVE BEEN MAKING: the 12/12 offline==online acceptance is FLAKY. It fails 11/12 on the UNMODIFIED HEAD source (control: KNN 81 online vs 71 offline), and the mismatching gun moves between runs (KNN, then WallBounce) - a live/offline boundary race. So '12/12' was a lucky run, and that proof should be treated as strong-but-not-exact until the race is fixed. This diff does not touch replayFixture/spawnBullets/tickBullets and the selector is never called during replay, so it is pre-existing. SIDE FINDING, not fixed: the shipped live bot never calls randomize(), so the 'random tie-break' is a FIXED sequence across process restarts. Overfitting guard vs a non-surfer (SpinBot): inconclusive - ModularBot fires only 17-31 real shots/run against fast bots because the range-aware firing gate is strict at long range, so the guard has little power. Wilson looked better (18.5% vs 8.6%) but on 70-92 shots with a 5-33% spread. Not evidence either way. |
||
|
|
57b2ac3849 |
feat(guns): scale-aware power selection (+52% damage); TM classifier gun built, measured, DISABLED
TASK 2 - power selection, a clear win. bestPower used an ABSOLUTE MinHitRate = 0.40 bar. Measured per-bin virtual rates (rolling-100 fraction) show no bin ever clears 40%, so 11 of 14 guns were stuck at bin 0 (power 1.0) even where higher bins were comparable: Linear p1.0 44% p1.5 39% p2.0 30% p3.0 29% old bin 0 -> new bin 3 Accel p1.0 44% p1.5 40% p2.0 26% p3.0 29% old bin 1 -> new bin 3 Pattern p1.0 50% p1.5 40% p2.0 27% p3.0 12% old bin 1 -> new bin 2 Replaced with a scale-aware PowerBarFrac = 0.50 (a dimensionless FRACTION of the gun's own best bin rate). 13 of 14 selections now pick heavier bullets. Real effect vs DrussGT (8 rounds x 3 runs): hit rate unchanged (7.56% -> 7.47%) but damage dealt +52% (157 -> 239 per run) and rounds end faster. Same accuracy, half the shots, half again more damage. TASK 1 - the TM pattern-classifier gun does NOT earn its slot. It was built as a mixture of experts with a corrected-Granmo TM as a multi-class gate over HeadOn/Linear/Circular/WallBounce/Accel, labelled by which expert's prediction was closest to the actual enemy position (an exact, supervised, per-shot label - no delayed credit). Offline it loses to the best of its OWN experts on essentially every fixture, and against DrussGT it cost real performance: baseline (path+relative) 7.56% real hit rate, damage 157 + power fix 7.47%, damage 239 + power fix + TM gun 5.59%, damage 133 The gun was selected on 806 ticks and fired 24 real shots at 4.2%. So the tree ships with EnableTmSelector = false: code and wiring kept intact for re-enabling, but it is not in the active rack. Worth recording from the clause dump: the gate DOES latch onto meaningful structure. On energy-threshold-turner, HeadOn's clauses key on the energy bits (the rule's own driving variable) while Circular keys on distance/velocity. So the TM is learning something real and interpretable - it simply cannot beat 'always pick the best expert'. Root cause (INFERRED): the closest-expert label is noisy because several experts are near-tied, and under the path metric the winner varies by power bin while the gate sees one shared per-tick input, so a one-vs-rest gate over a saturated 870-bit clause space has no margin to exploit. (Zero-padding the 2-frame window was tried first and saturated every clause at 256-755 included literals; alternating the two real frames fixed that.) Also factors the corrected feedback into an exported tmLearnDir and exports the encoding/TM primitives; the Tsetlin tests still reproduce the documented mean=13.8 included literals, so the refactor is behaviour-preserving. Verified: 33/33 guard checks, tsetlin tests green, metric checks green, new power-selection guard green (13/14 selections change; relative bar still picks bin 1 and not bin 3 for a [30,25,12,5]% profile), 12/12 offline==online acceptance under the shipped default. |
||
|
|
dea4dcb574 |
feat(gun_harness): scale-aware selector thresholds; default = path + relative
The selection thresholds were calibrated for a rate scale that does not exist.
MEASURED on an exact offline replay of a fogged live WorldState vs DrussGT
(1397 selection ticks), the 0.10 absolute floor fires on 53.0% of point-metric
ticks and forces HeadOn, which has a REAL hit rate of 2.0-4.4% - worst or
near-worst of 13 guns. HeadOn's selection share: 69.1% (abs+point) -> 43.5%
(rel+point). My earlier claim that the floor fires ALWAYS is REFUTED - it is
53%, because bestRate is a max over gun x bin and a >=50-sample bin
occasionally clears 10%. The mechanism is confirmed; the literal statement was
not.
Scale-aware mode (GUN_SELECTOR_MODE, absolute|relative, default relative):
RelTieMargin = 0.20 dimensionless FRACTION of bestRate, replacing the
fixed 2pp band so the band scales with the metric
FloorPeakFrac = 0.25 the floor fires iff bestRate < 0.25 * peakRateRef,
SelectorWindow = 256 where peakRateRef is the field-best rate over the last
256 selection ticks - keeping the original 'don't trust
a collapsed field' purpose but only when the field is
bad RELATIVE TO ITS OWN RECENT BEST, and counting only
guns with >= MinObsBeforeCompete samples so cold-start
100% spikes cannot pin HeadOn
also pools the rate over power bins instead of taking the max over bins, so
one lucky bin no longer wins
absolute mode is preserved byte-for-byte for rollback.
A/B vs DrussGT, real server hit rate, 3 runs x 10 rounds per config, one frozen
binary:
absolute+point 3.66 / 2.45 / 5.01 pooled 3.76%
absolute+path 7.55 / 8.21 / 6.83 pooled 7.57%
relative+point 7.66 / 6.18 / 5.79 pooled 6.59%
relative+path 7.15 / 7.55 / 6.90 pooled 7.21%
absolute+point is SEPARATED from all three (p < 0.0001); the other three
OVERLAP each other (p = 0.18-0.64). So the METRIC is the dominant lever and
under path the two threshold models are statistically tied.
DEFAULT SET: metric = path, thresholds = relative. absolute+path was nominally
0.35pp higher but indistinguishable (p = 0.64); relative is the principled
scale-aware fix, is the only model that works under BOTH metrics, and prevents
the point-metric catastrophe if anyone switches back. Shipping absolute would
ship the accidental side-effect this work exists to remove.
STILL NOT SOLVED: the selector remains only a moderate ranker.
Spearman(virtual rank, real rank) is 0.52 for the winning config, 0.36 pooled
for path and 0.04 for point - and it is INCONSISTENT across run sets. The
metric switch won by de-selecting HeadOn, not by ranking guns better. That is
the next problem.
TASK B, report only: do NOT drive selection from raw real hit rates yet.
Only the selected gun fires, so unselected guns get near-zero real shots
(GuessFactor 20, Linear 24 vs HeadOn 733); noise is fatal (n=470 at p=10% gives
+/-2.8pp, most guns n<200 gives +/-5pp+ across a 3-15% spread); and real rate is
conditional on when the gun was selected. A blended signal with forced
exploration and shrinkage is defensible in principle but needs thousands of
shots per gun across many battles. Real rate is best used OFFLINE as the
evaluation metric - which is exactly what this A/B did.
RELATED BUG FLAGGED, not fixed: MinHitRate = 0.40 in bestPower is on the same
wrong scale - no bin ever clears 40%, so once every bin has data, power
selection falls back to bin 0 (power 1.0) late in a round.
Verified: 33/33 guard checks, 11/11 metric checks, tsetlin green, 12/12
offline==online acceptance under the shipped default, run_range rc=0 over 20
fixtures. Adds analyze_selector.nim to measure floor/tie/bestRate/HeadOn-share
per config on any fixture.
|
||
|
|
3b5d70b7c3 |
feat(gun_harness): runtime metric switch + A/B proving the point metric mis-selects
Adds GUN_VBULLET_METRIC (point|path, default point = unchanged behaviour) so the virtual-bullet hit model can be selected at runtime with no rebuild. Both the live tracker and the offline replay read the same value, so the 12/12 offline==online acceptance holds under EITHER setting (verified for both). A/B AGAINST THE LIVE BOSS, real server-side hit rate as ground truth, 5 battles x 12 rounds per metric on one frozen binary: point 4660 shots / 219 hits = 4.70% (per-run 3.16-5.53) path 4834 shots / 359 hits = 7.43% (per-run 6.55-8.24) The distributions DO NOT OVERLAP: path's worst run beats point's best run. +2.73pp, +58% relative, z = 5.56, p < 0.0001. Range distributions were identical (~460-478 px), so this is not a range confound. MECHANISM - and this is the important part. The gain is SELECTION, not better gun learning. Under the point model every gun's virtual rate is compressed into 0.6-4.4%, so HeadOn sits inside the 2pp tie margin and takes 72.6% of selection ticks / 76.9% of shots - while HeadOn is 11th of 13 by REAL hit rate (2.3%). The path model widens the band to 4.7-13.7% and ranks HeadOn 10th, so its shot share falls to 35.9% and Pattern/Accel/WallBounce get picked instead. Counterfactual: applying the point model's per-gun real rates to the path model's shot mix yields 7.65%, i.e. essentially the whole observed gain. So the selector, not the guns, is where the win lives. PER-GUN REAL HIT RATE vs DrussGT (path mix, the answer to 'which guns are worth keeping'): WallBounce 10.8, Pattern 10.5, Accel 10.0, Displace 9.3, Circular 9.2, AvgLead 8.5, KNN 5.7, StopShot 5.2, GuessFactor 3.7, Tsetlin 2.9. Per-gun N is small (hundreds of shots) so single-gun ordering is indicative, not definitive. TWO CAVEATS, recorded because they undercut a naive reading: 1. One adversary. DrussGT is a wave surfer and HeadOn is genuinely bad against surfers, so part of this may be matchup-specific. 2. The path model is NOT a better general ranker. Spearman(virtual rank, real rank) is 0.52 under point vs -0.04 under path. It wins by accidentally fixing HeadOn's mis-rank, not by ranking guns better. A more durable fix is to address the selection logic directly - which is the next job. Also adds a focused guard test (test_vbullet_metric) covering parsing/default, a receding-target point-miss/path-hit, a perpendicular-target path-miss, and replay determinism. Verified: 33 guard checks, 12/12 acceptance under both metrics, tsetlin tests green, range 34.3% (point, unchanged) / 50.8% (path). |
||
|
|
e2ca2fc7d8 |
fix(guns): recover the DrussGT regression with a radial-fraction range blend
The previous fix (learn the residual against a constant-velocity base) was structurally right but cost us on real wave-surfing movement: GF 108 -> 55, KNN 101 -> 74 on the classic DrussGT captures. Root cause: the linear base is a poor model for a surfer, so the residual histogram is noisier than the old total-lead histogram. FIX: blend the RANGE between a radial-only forecast and the geometric one by radialFrac (the fraction of recent per-tick motion that is radial), keeping the constant-velocity bearing. dist = radialDist + rf*(linearDist - radialDist). New VelocityTracker in common_libs/guns/lead_forecast.nim; the window default is 32 and results were identical at 16 and 40, so it is not tightly tuned. Nine candidate bases were measured and rejected WITH NUMBERS rather than by argument, which is why I trust the winner: velocity scaling 0.8 recovers DrussGT but destroys wall-bounce 241 -> 20 radial-only range excellent DrussGT, wall-bounce 241 -> 140 short-window averaged vel worse than both bases outright hard reversal/speed gates help DrussGT, lose nothing, but weaker than blend radial-fraction blend best on BOTH <- shipped Result (hits per 2000; classic-5 = classic DrussGT captures, tr-5 = the new closed-loop TR captures, synth-10 = the rest): base classic-5 GF/DGF tr-5 GF/DGF synth-10 GF/DGF current(prefix) 108 / 108 41 / 39 1302 / 1302 linear(postfix) 55 / 76 9 / 4 2702 / 2692 BLEND 171 / 100 86 / 87 2717 / 2703 Strictly better than both on classic-5 GF and on every synthetic bucket. The one figure below the old base is classic-5 DecayGF (108 -> 100, -8/2000, within noise) and that is stated plainly rather than hidden. TASK B - enemy energy in learners. KNN gains an 8th feature, enemyEnergy/100, on a FIXED [0,1] scale (not min-max) because threshold behaviour keys off absolute energy. Honest result: it is NEUTRAL on the target fixture (77 vs 77) and roughly neutral in aggregate. The base change, not the feature, moved that fixture. Tsetlin already encoded enemyEnergy and now scores 88/400 on energy-threshold-turner against Linear's 43/400 - a 2x margin, which is the 'can a TM learn a high-level pattern' question answered in gun form. TASK C - is the virtual-bullet metric itself faithful? Quantified: scoring the bullet's PATH against BotRadius instead of the single point at aim distance raises every gun by +31% (GF) to +86% (HeadOn), so the current model is PESSIMISTIC, and it RE-RANKS materially: Linear 9th -> 6th, AvgLead 7th -> 3rd, GuessFactor 4th -> 9th, DecayGF 6th -> 12th. The 12/12 offline==online acceptance still holds under the path model (verified with a temporary env hook driving both sides), so no red flag. VERDICT: do NOT switch. The point model is the standard virtual-bullet PREDICTION-ACCURACY fitness - the bullet must arrive at the predicted point at the right time - while the path model measures hypothetical hit chance against a target that never dodges, and in open-loop fixtures it over-credits directional guns (HeadOn 35% on DrussGT, 100% on constant-velocity) for exactly that reason. The models differ materially but the current one is not shown to be unfaithful FOR ITS PURPOSE. Because the metric drives gun SELECTION, this is now being A/B'd against real hit rate versus the live DrussGT boss, which is the only ground truth we have. Verified: 20 fixtures 35636/104000 (34.3%); 33 guard checks; 12/12 acceptance; tsetlin tests green; live gauntlet 5/5. |
||
|
|
17c99f542f |
feat(fixtures): closed-loop DrussGT captures from real Tank Royale battles
The classic captures were OPEN-LOOP: replayed DrussGT never dodged OUR
bullets. These come from real TR battles through the working Java bridge, so
the recording contains genuine reactions to ModularBot's live fire. The
open-loop caveat is gone (perfect-information remains).
PRIMARY RESULT - the boss beats us badly. DrussGT 1447 - ModularBot 300 over
15 rounds, ModularBot winning only round 5 (DrussGT died at tick 1893). Rounds
are long, not truncated: mean 1335 ticks, ModularBot got off 1134 shots.
ModularBot 1134 shots / 60 hits = 5.3% real hit rate
DrussGT 1400 shots / 169 hits = 12.1% real hit rate
So DrussGT's gun is ~2.3x more accurate than our entire rack, on top of far
better movement. That is the number to move.
Also captured: shield-on variant (DrussGT 939-287, 9/10 - ModularBot takes
round 1 to the known shield warm-up), and vs SpinBot 1175-0, Crazy 1080-1,
Corners 1659-0. 20,026 + 12,629 + 10,824 + 11,507 + 2,575 ticks.
Movement statistics match the classic set within ~0.04 on the perpendicular
and radial fractions, so this is the same wave surfer in TR physics:
TR vs modularbot: perp 0.967, radial 0.001, 52.8% at full speed,
reversing 46.3%, median range 464 px.
CLOSED LOOP PROVEN, not asserted. ModularBot's fire is a heat-limited near
metronome (median interval 14 ticks), which gives a usable exogenous clock:
- event-locked |delta heading| oscillates 0.96 -> 2.69 deg about a 1.47 deg
mean with the fire period, almost every lag outside the 95% band of a
400-iteration phase-shuffled null;
- cross-correlation of |delta heading| against the fire impulse peaks at
r = +0.111, lag 12 ticks, permutation p = 0.005 (null peak mean +0.016);
- OWN-FIRE CONTROL is flat, so the oscillation is enemy-driven rather than
an internal cadence;
- range response is weak (~3 px over 30 ticks, near noise) and is therefore
NOT claimed, and per-bullet dodging is not claimed either because the
bullet detector (shield) is off.
CAVEATS: still perfect-information (observer gives true positions every tick,
unlike the live bot's stale between-scan WorldState) so these remain optimistic
vs live play; and they are open-loop AT REPLAY TIME - 'closed_loop' describes
the capture, not a later replay. TR conversion residual is ~1.5 deg mean
because the TR server moves along the pre-turn heading, vs 0.000 deg for the
classic captures.
Adds analyze_closed_loop.py (PSTH event-locking, phase-shuffle permutation
null, cross-correlation, own-fire control) and per-round result sidecars.
|
||
|
|
7f706e5b14 |
fix(guns): GF family aimed at the wrong RADIUS, not the wrong angle
The entire GuessFactor family scored 0% on clean circular and wall-bounce trajectories. Two hypotheses were on the table and BOTH were wrong: - MEA range too narrow / edge clamping: REFUTED. Measured 0 clamped shots out of 837/849/957, required offsets peak at ~33 deg against MEA 28.1-46.7 deg, and the 8 in arcsin(8/bulletSpeed) is correct (it is the max robot SPEED, not the hit radius). Changing it to BotRadius=18 would have coarsened resolution for nothing. - Peak selection: REFUTED. A sweep of every constant GF value showed the ORACLE-BEST constant offset on the original gun was only 6% circular, 4% wall-bounce, 7.5% random-walk. No peak choice could have done better. The learning path was fine too: ~850-960 observations per fixture, 0 starved waves, well-populated histograms. REAL CAUSE: the GF family aimed at the FIRE-TIME distance. The virtual-bullet metric resolves a bullet at the AIM-POINT distance and scores that single point against the enemy's position on that tick, so with any radial target motion the bullet stops at the wrong radius and misses even with a perfect angle. Angle-only prediction is structurally unscoreable under this metric. FIX: give the GF family a self-consistent constant-velocity forecast as its base reference (new common_libs/guns/lead_forecast.nim, which iterates the flight time to the same fixed point circular.nim uses), so the histogram learns the RESIDUAL against that forecast and the aim point lands at the right radius. Applied to guess_factor, decay_gf and knn_gun. Same defect fixed in Linear: it did a one-shot dist/bulletSpeed extrapolation and never iterated its flight time. The oracle sweep proves the structural fix, independently of tuning: the best achievable constant GF moved 6% -> 20% (circular), 4% -> 57% (wall-bounce), 7.5% -> 49% (random-walk). MEASURED, all 15 fixtures: total 39.0% -> 44.4% (30399 -> 34654 hits). circular GF 6 -> 23, DecayGF 6 -> 21 wall-bounce GF 0 -> 60.2, DecayGF 0 -> 60.2 constant-vel GF 26 -> 100, DecayGF 26 -> 100, KNN 26 -> 100, Linear 87 -> 100 random-walk GF 0 -> 53, DecayGF 0 -> 52, Linear 24 -> 53 StraightLine GF 8 -> 77, DecayGF 8 -> 77 Non-regression: 33 guard checks pass, the range's 12/12 offline==online acceptance still PASSES, tsetlin tests green, live gauntlet 5/5. HONEST TRADE-OFF, recorded rather than hidden: on the 5 real DrussGT wave-surfing captures the GF family REGRESSES - GuessFactor 108 -> 55, DecayGF 108 -> 76, KNN 101 -> 74 hits per 2000. The linear base is a poor model for a surfer, so the residual histogram is noisier than the old total-lead histogram. Linear itself improved there (95 -> 105). The synthetic range and the live gauntlet both improved, and the structural bug is provably fixed, so this was judged worth the cost - but recovering the DrussGT regression is the next job, not something to wave away. |
||
|
|
8e2be6a4c6 |
feat(tools): the real DrussGT now plays and wins Tank Royale battles
The unmodified DrussGT.jar connects, wave-surfs, fires and beats every adversary we have. 5 rounds each, all rounds won: SpinBot 542-16, Corners 823-4, Crazy 604-0, RamFire 900-0, ModularBot 506-72. It is genuinely surfing, not drifting or stalling. Movement statistics against the classic captures, same metrics, same analyzer: opp perp TR/classic reversing TR/cl median range TR/cl SpinBot 0.967 / 0.962 0.466 / 0.464 462 / 402 Corners 0.894 / 0.920 0.462 / 0.419 520 / 504 Crazy 0.841 / 0.829 0.527 / 0.439 370 / 378 RamFire 0.762 / 0.651 0.533 / 0.417 326 / 283 ModularBot 0.956 / - 0.466 / - 453 / - Every round starts moving within 3-11 ticks and runs at 75-82% full speed. Implemented: the classic remaining-quantity motion model (delegated to the TR Bot's own Nat-Pavasant model - identical constants: accel 1, decel -2, max 8, body turn 10-0.75|v|, gun 20, radar 45 - so getDistanceRemaining and getTurnRemaining are exactly self-consistent with what is emulated); event synthesis with classic ordering; bullet identity via object identity; gun heat; rounds; radar cadence; firing translation; and the ThreadManager landmine is killed by installing a no-op IThreadManagerBase in ContainerBase.instance (verified: without it 'RobotException: ThreadManager cannot be null!' kills the bot thread; with it the write succeeds). Also adds TrBattleCapture, an observer that dumps per-tick state in the SAME JSONL fixture format as the classic capture, so legacy bots can be captured from Tank Royale battles too. HONEST DIVERGENCES (README section 5.9): the TR server moves along the PRE-turn heading and then turns, while classic aligns displacement with the POST-turn heading, so the analyzer's conversion error is ~1.5 deg rather than 0.000 deg; distanceRemaining decrements by target speed rather than actual distance; collision clamping differs; BulletMissedEvent can fire less often because age-expiry has no TR event; bulletId is a local temp id; StatusEvent/onPaint/SkippedTurnEvent are never delivered. Physics fidelity diverges by construction - expect to retune. EnergyDomeWorker (the bullet shield) is off by default: its precise bullet-detection warm-up makes round 1 up to 79% stationary vs 35% in classic. Rounds 2+ match classic closely with it on, but the pure surfer is the consistent path. DRUSSGT_SHIELD=1 re-enables it. Jars remain out of git. |
||
|
|
d5061ee215 |
test(range): restore the 12/12 offline==online proof; measure TM clause readability
Task 1 - the acceptance proof was unrunnable because RecordWorldState was a
compile-time const set to false. It is now a RUNTIME switch
(let RecordWorldState* = existsEnv("TR_RECORD_WORLDSTATE")), default OFF, so
ordinary runs write no fixture, and acceptance_offline_vs_online.nim enables
it for the battle it spawns and clears it afterwards. Restored and run twice:
12/12 deterministic guns match exactly (128-tick and 546-tick battles), with
Tsetlin reported separately as stochastic. Both nimble build variants clean.
Task 2 - does a compact encoding turn the TM's 99.35% into a READABLE rule?
Measured across window sizes (fixed seed, no tuning):
frames TEST acc eff.lits/clause firing clauses counterfactual low/high/mean
10 99.35% 152.8 37 100/24/62.4%
3 95.94% 54.9 35 96/20/58.6%
2 99.48% 39.6 38 95/25/60.7%
1 98.30% 19.2 45 100/24/62.3%
So 2 frames is strictly better than 10 on BOTH axes: +0.13 accuracy for 4x
smaller clauses. The 3-frame dip is non-monotonic and left unexplained rather
than smoothed over.
A readable rule WAS partially recovered. Five clauses carry the exact Gray
form !g10 ^ !g9 ^ !g8; g10 is inert in this data, so the effective rule is the
2-literal proposition !g9 ^ !g8, i.e. energy < 25.6. That is a genuine
threshold in readable propositional form - but at 25.6, NOT the labelled 30,
because 256 is a power-of-two Gray boundary expressible in two literals while
300 needs a longer conjunction. The TM found the nearest SIMPLE threshold.
The honest caveat: that threshold is not the ensemble's decision mechanism.
The counterfactual follow rate (high 24%, mean 62.3%) is statistically
identical at 1, 2 and 10 frames, so compactness did not make the model read
energy - its vote is carried by co-occurring bearing/velocity/heading/wall
literals. Also identified: clauses containing all 11 Gray energy bits are
satisfied at exactly one raw value (50, the dataset floor), so they are
'energy has hit the floor' detectors, not thresholds.
Methodological fix worth keeping: the earlier single-frame counterfactual
wrote energy into all 10 frame slots including the zeroed ones, reviving dead
clauses and producing a spurious 2% high-follow rate. setEnergyFrames now
rewrites only the exposed frames; the corrected figure is 24%.
|
||
|
|
89370008da |
fix(tsetlin): make the TM actually learn - saturation 714 -> 13.8 literals/clause
The gun has never contributed anything: Tsetlin.vHits was byte-for-byte equal to Linear.vHits in every measured round of every run, because its learned correction was always exactly 0. Six diagnosed defects fixed, plus one that was required to make the first one work: 1. Type I now conditions on the clause output. It previously rewarded included true literals unconditionally, omitting Granmo's (c=0, lk=1) -> toward Exclude counter-force, so true literals ratcheted toward Include forever. This was the root cause of the saturation. 2. Type II was unreachable dead code: its guard required cOut==1 AND lits[lit]==0 AND st>0 (included), but cOut==1 guarantees every included literal is 1. Its direction was wrong too - it should increment EXCLUDED false literals when the clause fires. 3. Resource allocation restored: Granmo's (T - clip(v,-T,T))/(2T) target replaces |error|/(2*RESID_MAX); TM_T was only an output normaliser. 4. Label baseline fixed - the factor-2 shrink. predX = linearX + cx, so the label was delta - cx while the learner's output IS cx, giving error = delta - 2cx and a fixed point of cx = delta/2: HALF the needed correction even with perfect feedback. TmTrace now stores linearX/linearY and training uses delta. 5. Hits no longer zero their label (a hit means |miss| < 18px, not 0). 6. The enemy-energy feature was duplicated - tmEncodeFrame passed state.selfEnergy with a stale comment claiming enemyEnergy was absent, while WorldState.enemyEnergy exists. Enemy-energy rules were literally unrepresentable. 7. REQUIRED EXTRA: tmEvalClause now implements Granmo Eq. 6 - an all-Exclude clause outputs 1 during learning and 0 during classification. Without it, fix #1 deadlocks every clause at empty. MEASURED EFFECT (energy-threshold-turner fixture, seed 1): mean included literals per active clause 714.0 -> 13.8 active clauses 100/100 -> 53/100 nonzero corrections 8/764 -> 708/764 Tsetlin virtual hits (Linear = 27/400) 27/400 -> 69/400 Divergence achieved: offline on 7/8 fixtures, and in a live gauntlet (RandomMover: Tsetlin 199/1200 vs Linear 288/1200, vDropped=vStarved=0). Tsetlin now LEARNS but is not yet competitive with Linear - the regression head is untuned, flagged as follow-up rather than claimed as a win. Also ignores compiled test harnesses that have no file extension, which the existing '**/tests/test_*' rule misses. |
||
|
|
e0bfa9b5e1 |
feat(tools): drive the real DrussGT jar outside the Robocode engine
The question was whether we can fight a genuine legacy leader bot in Tank Royale. Answer: the API side is now PROVEN, not estimated. KEY FINDING: the classic robocode.* API is a thin delegation layer over a public seam. javap -c shows AdvancedRobot forwarding every call to _RobotBase.peer (IBasicRobotPeer/IAdvancedRobotPeer), and _RobotBase.setPeer is public final. So we do NOT need to reimplement the API: we reuse the genuine robocode.jar and implement only the 75-method peer interface. Consequences: - DrussGT's 22 sources compile against the real API with ZERO unresolved symbols. (The literal 22-file javac fails only on two PRE-EXISTING duplicate classes - GFRange and Indice are declared both inline in DrussGunDC.java and as standalone files - and the 6 classes that ship without source. The jar supplies all of them, so a shim never cares.) - Runtime smoke test PASSES: the unmodified DrussGT.jar is loaded through a child URLClassLoader and driven for 200 synthetic ticks, emitting movement intents every tick and 169 fire requests, with its thread surviving. This is the cheap path to the 'final boss', and it also unlocks the whole roborumble archive rather than one bot. Traps found by measurement: - ScannedRobotEvent's constructor order is (name, energy, bearing, distance, heading, velocity) - NOT heading-before-bearing. The wrong order silently yields distance=0, an immediate KD-tree insert and an NPE in EnemyMoves.predict; it hung the first smoke run. - RobocodeFileOutputStream has a hard engine dependency (resolves IThreadManagerBase via ContainerBase) and throws 'ThreadManager cannot be null!' outside the engine, killing the bot thread. Reached from DrussGT's own contain() error logging, so it must be stubbed. - robocode.RobotDeathEvent is required and was NOT in the predicted API list. - Bullet.equals() is genuinely called for bullet identity, not just getters. REMAINING WORK (README section 5): coordinate rotation DONE, execute()->tick bridge DONE and proven, ThreadManager fix scoped. The main open item is the classic motion model (setAhead distance semantics vs TR speed), estimated 1-3 days, plus event synthesis/ordering, gun heat, bullet identity, round and radar cadence, and firing translation. NO API UNKNOWNS REMAIN. Physics fidelity will still diverge from classic - expect to retune. Jars stay out of git (blocked by tools/robocode_shim/.gitignore); the genuine robocode.jar and DrussGT.jar are referenced from /tmp via ROBOCODE_JAR / DRUSSGT_JAR. |
||
|
|
76ae6170f8 |
docs(research): portability audit of DrussGT 3.1.4159
Measured hazard counts for a Java->Nim port, with the correction that only 22 of 28 top-level classes ship source (6 do not: 5 in the gun package plus dMove/Scan), so a full port would need a decompiler while a shim would not care at all. Real traps: 193 float / 35 casts / 147 literals / 60 float[] in the movement closure (the danger histogram is float[171] -- porting 32-bit Java floats to Nim's default float64 diverges silently); ~68 non-final statics; and 4 Java single-& sites with side effects, which break under Nim's short-circuiting 'and'. Non-issues, correcting earlier assumptions: 0 sites of %-on-negative (angle normalisation is floor-based) and FastTrig has no lookup tables, it is 7 coefficient-exact polynomials. Movement scoping: 5007 LOC across 14 files, ~4.8-6.1k Nim LOC, 4-8 focused agent-days to first-compiles. It can be ported WITHOUT the gun (data flows movement->gun only), but the harness never routes HitByBullet into movement modules, so the danger bins would never train -- that plumbing is the real blocker, not the translation. |
||
|
|
4f18c8ce07 |
feat(tools): capture real DrussGT movement from classic Robocode as fixtures
There are no genuinely competitive adversaries for Tank Royale, and the in-repo ones were broken until recently. Classic Robocode 1.9.5.5 is obtainable (SourceForge, 20.4 MB) and its programmatic control API (robocode.control.RobocodeEngine + BattleAdaptor.onTurnEnded) can run real battles headless and expose per-turn robot state. So a legacy leader bot's MOVEMENT can be captured and used as a gun-testing fixture with no port. Captured unmodified DrussGT 3.1.4159 vs spinbot/ramfire/crazy/corners and a mirror match: 28,797 ticks, plus two trivial-bot contrasts. Conversion to the Tank Royale convention is validated to 0.000-0.001 deg by recomputing the direction implied by (heading, speed) and comparing it against the recorded per-tick displacement -- i.e. the data is proven to be genuine recorded motion rather than a mangled export. (A first attempt treated the snapshot API's headings as degrees; they are radians, ~95 deg off.) The statistics confirm it is really a wave surfer: perpendicular to the opponent 65-96% of ticks, radial ~0.001, 41-72% of ticks at full speed, reversing on 42-46% of ticks, holding range at a 283-526 px median. The straight-line contrast is radial-dominant (0.75) with ZERO reversals. Discovery: DrussGT detects predictable guns and switches to a bullet-shield stand-still mode, so captures against sample.Walls/TrackFire had to be rejected as non-movement. CAVEATS, recorded in DRUSSGT_FIXTURES.md: these are open-loop (replayed DrussGT never dodges OUR bullets) and perfect-information (the observer gives true positions every tick, unlike our stale live WorldState). Both make our guns look better than in live play, so use them for RELATIVE gun ranking, not absolute hit rates. Jars stay out of git; capture tooling is reproducible via capture.sh. |
||
|
|
974528d5cf |
feat(gun_harness): offline gun range, proven equivalent to live play
Gun evaluation previously required a full end-to-end battle (Java server + battle runner + websocket IPC to 2 bot processes, 50 rounds, ~3.4 min) and yielded only ~300-900 REAL shots across 13 guns -- far too few to rank guns, which is why tuning needed many repetitions. VirtualTracker is already a pure function of (WorldState stream, gun list); the only reason it needed Java was where WorldState came from. So the range replays a seq[WorldState] through the SAME tracker: offline and online scores are the same metric by construction, not an approximation. ACCEPTANCE TEST (the point of the whole thing): record one live round, replay it offline, compare per-gun virtual hit rates. 12/12 deterministic guns match EXACTLY, reproduced twice. Tsetlin is compared separately because tmLearnOne calls rand(). Getting to 12/12 exposed two real ordering quirks in the live loop: run() calls go() before the aim/fire block, so tickBullets resolves against the NEXT tick's scan while the prediction used the previous one; and if the target dies during that go() the final tick's spawn+resolution is skipped entirely. The recorder emits an end marker for the second case. The 5th (selected-gun) predict call was verified to be a no-op. Measured cost: 8 fixtures (1770 ticks, ~92k virtual bullets, 13 guns) replay in 2.9 s, ~32k virtual bullets/s -- roughly 70x faster and 100x more samples than a live gauntlet. Also adds a per-tick WorldState recorder behind const RecordWorldState (default off, mirrors the ShotLog idiom) which records the state the bot ACTUALLY builds, staleness included, rather than true positions -- recording the latter would hand the guns perfect information and produce flattering scores. 9 new guard checks (33 total, all passing), including fixture round-trip, replay determinism, stationary->HeadOn 100%, constant-velocity->Linear>HeadOn, and the energy-threshold turner crossing at t=41. |
||
|
|
3c90a5941d |
feat(selector): range-aware firing gate fitted to 2611 measured shots
Measured, not assumed. With the gate temporarily opened to 20 deg, every real shot was logged (tick, angle error at fire time, distance, power, hit) across 3 gauntlets: 2611 shots, 57.3% aggregate. Findings: - The geometric cone atan(BotRadius/d) is directionally confirmed but a WEAK lever: even at 0.0-0.1 deg error the hit rate at 400-600px is only ~53-57%, because PREDICTION error dominates alignment error. - Real effect of tightening the gate: 57.9% -> 68.0% aggregate hit rate (fixed 0.1 deg), not the 76.9% previously reported -- that was a high-variance draw (per-rep 62.8/66.4/77.2%). - The shipped range-aware gate (SafetyFactor 0.6) does NOT beat the fixed 2.0 deg gate on hit rate (55.8% vs 57.9%, ~1.5 sigma, inside noise). It fires 22-28% more shots and therefore lands more total hits (~509 vs ~434 per rep). No per-adversary score delta exceeded the 300-point run-to-run noise band, so no config is demonstrably better on score. Shipped anyway because it is strictly more expressive (a fixed threshold is the special case), tunable from one const, and physically motivated, but the honest verdict is recorded in-code: the gate is not the bottleneck. AimThresholdDeg is removed; shouldFire now takes distPx. Degenerate or NaN distance falls back to the ceiling rather than dividing by zero. Also adds a per-shot logger to ModularBot behind 'const ShotLog' so the measurement above is reproducible, and 10 new guard checks (24 total, all passing) covering monotonicity, clamping, formula, perfect alignment, gross misalignment and degenerate distance. Cross-checked against the server source: the gun fires BEFORE the turn is applied, so the logged angle error is the true departure error, and fireAssist auto-aim is off (unset by the Nim API and forced false by setAdjustRadarForGunTurn). |
||
|
|
54e9757567 |
docs(research): TM learning tracks + note that the TM-DEB source was deleted
tm-learning-tracks.md covers three things, all marked [FACT]/[INFERENCE]/ [UNKNOWN]: - Section A: the Tsetlin gun's label is measured against the wrong baseline. predX = linearX + cx, so rx = actual - predX = delta - cx, and inside tmLearnOne error = residual - predicted = (delta - cx) - cx = delta - 2cx. The fixed point is cx = delta/2 -- HALF the correction needed, even with perfect Granmo feedback. Fix: store linearX/linearY in TmTrace and train on delta. Also: hits zero the label instead of carrying their true residual, and the per-clause step is magnitude-blind. - Section B: what a TM is actually good at (AND-clauses over binary literals, readable output) and why this repo suits it -- the gun already builds an 83-bit x 10-frame Gray-coded window (870 bits). Includes a falsifiable known-rule benchmark proposal. - Section C: delayed-reward learning belongs to the MOVEMENT layer, not the gun. The gun's outcome is delayed but exactly pairable via (fireTick, powerBin), so its effective lambda is 1 and discounting would only destroy information. Also records that docs/papers/tm-deb-paper.pdf was deleted by the user as AI-generated and unverifiable, while the Granmo-based feedback diff in tm-deb-assessment.md stands on its own. |
||
|
|
669f9acd41 |
docs(research): assess TM-DEB paper against the Tsetlin gun's real failure
Verdict: the document (docs/papers/tm-deb-paper.pdf, 'Generated by Gemini Notebook') solves temporal credit assignment under delayed reward, which is not our problem. Our gun's failure is clause saturation (~131 of 1740 literals included per clause -> conjunction fires with probability ~2^-131 -> correction identically 0), and TM-DEB only scales update FREQUENCY by gamma^dt, so it would leave the fixed point untouched and additionally delete the long-range feedback we need: gamma^90 = 4.4e-7 at the paper's own gamma=0.85, and those are the shots whose lead matters most. Credibility signals recorded in the doc: reference [2] misattributes authors/venue/year; Algorithm Spec 2 steps 9-12 are literal '%' placeholders so the automata update is simply absent; Table 1 is titled 'Expected' and reports never-measured accuracies; Eq. 12 is not a faithful copy of Granmo's Lemma 2. The audit also produced the actionable result: a line-by-line diff of Granmo Table 2/3 feedback against tmLearnOne, identifying why the automata saturate - Type I never conditions on the clause output so it omits Granmo's c=0, lk=1 -> -1 w.p. 1/s counter-force (true literals ratchet toward Include), Type II's guard cOut==1 AND lits[lit]==0 is unsatisfiable and therefore dead code, and the (T - clip(v,-T,T))/(2T) resource allocation is missing entirely. |
||
|
|
e53690036b |
fix(guns): speed-sensitive caches, dead stop-shot branch, exact TM trace pairing
Four guns cached a whole prediction per tick while predict() is called once per power bin, so every bin after the first (and the real fired shot, which shares lastState) reused the power-1.0 lead. Fixed by caching only the speed-INDEPENDENT derived state and recomputing the lead per requested speed: - stop_shot: also fixes prevSpeed being written before it was read, which made abs(speed) < abs(prev) permanently false and the entire stop-prediction branch unreachable (it was just Linear). - displacement: the cache key included bulletSpeed, so the guard missed on all four bins and the 15-tick window advanced ~4x/tick, making the inferred velocity ~4x too small. - averaged_lead: tick cache removed outright. pattern_matcher: split into speed-independent match+path and per-call lead. FeedbackEvent gains fireTick/powerBin (additive; only virtual_bullets constructs one) so guns can pair feedback to the exact shot instead of guessing by coordinates. tsetlin uses it: traces are now keyed exactly by (fireTick, powerBin) with a 1024-slot ring, and the 10-frame window shifts at most once per tick (it was shifting ~4-5x/tick, so isWarmedUp tripped after ~2 ticks). KNOWN INCOMPLETE: tsetlin still does not diverge from Linear in battle. The two named bugs are fixed (a 600-tick sim shows trainedShots=2141, traceMisses=0, and a fixed-input probe converges to a 9.6px correction), but the TM's clause feedback itself is broken: ~131 of 1740 literals end up included per clause, so its conjunction never fires. Sweeping TM_S, TM_N_CLAUSES and a two-branch Type-I update did not change the correction from 0. Needs a real TM fix or removal, not another bug fix. First-ever guard tests for the gun selector: common_libs/tests/ test_gun_harness.nim (14 checks, headless, no Java). There were none before, which is how six broken guns survived a full analysis cycle. Against the previous HEAD, 5 of these checks FAIL - that is the regression guard. |
||
|
|
0cc682152d |
fix(guns): per-bin wave queues unbreak GF/DecayGF/KNN learning; fix vbullet drops
Wave queues (guess_factor, decay_gf, knn_gun): predict() stored ONE wave per tick while onResult() popped one per resolved bullet (~4/tick), so the queue drained to empty within a few dozen ticks, ~3 of every 4 resolutions returned without learning, and the survivor paired with a same-tick wave (bearingDelta ~= 0) pinning the histogram at centre. PROOF: GF.vHits == HeadOn.vHits and DecayGF.vHits == HeadOn.vHits byte-for-byte in every one of 50 rounds — the guns had degenerated to HeadOn. Now each gun keeps a per-bin FIFO with an O(1) head cursor. At most one push per (tick, bin) so the fire site's 5th predict() call is a no-op, and onResult pops the oldest wave of its OWN bin via e.bulletPower. Aiming math untouched (it was already correct: 0 deg = East, CCW+). maxBullets 2048 -> 8192: the rack spawns 52 bullets/tick so the ring wrapped every ~39 ticks while a long power-3 shot needs ~90, silently discarding unresolved bullets and biasing every measured hit rate by range. Added a droppedBullets counter so a future overflow is measurable, and wavePushes/ waveStarved counters on the three guns. After the fix: vDropped = 0 and vStarved = 0 across all 48 recorded rounds. fitnessFor is now exported, deterministic (enemies iterated in ascending id order) and shared by the selector and the stats dump, replacing a hand-rolled merge in ModularBot that never advanced its window head. Round lines gain additive keys: vDropped, vStarved. |
||
|
|
a90a0cc9b5 |
chore: untrack ModularBot_garage/ModularBot build artifact
nimble bin output lands at the garage root, so the existing '*_garage/out/' ignore rule never covered it and every build dirtied the tree with a 1.1 MB binary. File stays on disk; build regenerates it. |
||
|
|
26b66cbb24 |
feat(gun_harness): per-gun REAL hit attribution + bestPower cold-start fix
Attribution is proven, not guessed: the server assigns a per-round-unique bulletId (GunEngine.nextBulletId) and stamps the same id on BulletFired, BulletHitBot, BulletHitWall and BulletHitBullet. Keep a FIFO of fired gun ids, stamp bulletId -> gunId on onBulletFired, resolve through that map. Hits are deferred when onBulletHit precedes onBulletFired in the same turn (client dispatches priority 70 > 60), which recovered 14 unattributed hits. 99.9% of shots and 99.8% of hits attributed. Stats lines now carry per-gun realShots/realHits/realHitRate; the old keys and round-level totals are unchanged. bestPower: a gun with zero observations in every bin previously returned the HIGHEST bin (power 3.0) because an empty bin satisfied the 'count == 0' clause on the first countdown iteration. Cold guns now return the lowest bin as the docstring always claimed. Warm-gun path untouched. |
||
|
|
343e631633 |
fix(gun_harness): random tiebreak + drop AntiSurfer + raise MinObsBeforeCompete
- bestGun: replace first-index-wins argmax with random pick among guns within TieMargin (2%) of best rate. HeadOn at index 0 was silently winning every tie, starving Tsetlin/Linear/etc. - MinObsBeforeCompete 15 -> 50 (Pattern entered competition on noise) - add MinHitRateFloor 0.10: if no gun clears it, fall back to HeadOn instead of selecting the best of a bad field - ModularBot: remove AntiSurfer gun (0% virtual hit rate everywhere), 14 -> 13 guns, renumber ids and selection counters |
||
|
|
c034eb9d25 |
feat(testing): gun rack gauntlet + analysis reports
- fix(ModularBot): onBulletHitBot → onBulletHit (real hits were never tracked) - feat(ModularBot): per-round gun stats dump to /tmp/gun_stats.jsonl - feat(ModularBot): gun selection counter per round - fix(tests): adversary paths _garage suffix removed from 7 test files - feat(tests): test_gauntlet_5bots.nim — 10-round gauntlet vs all 5 adversaries - feat(tests): analyze_gun_stats.nim — JSONL parser for gun performance tables - docs: gun_rack_analysis.md — full per-gun performance report - docs: gun_rack_summary.md — TL;DR verdict table (keep/drop/tune) Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
2cc2a3bd87 |
fix(ModularBot): ram loop prevention, dead-target guards, cleaner logging
- 30-tick cooldown after ghost-stuck/timeout ram exit prevents re-entry loop - enemy_tracker.update() skips dead bots to prevent same-tick scan resurrection - TFIL graphics cleared when ramming is active movement - [config] logs: white base with green-highlighted changes only - [ram:enter] logs trigger reason and key values on false→true transition - [death] and [target-invalid] logs retained for diagnostics |
||
|
|
1f8574db1d | fix(rammer): rewrite heading logic — correct forward/backward steering toward enemy | ||
|
|
7657699531 | fix(movement): lock target during ram — no switching until target dies or ram exits | ||
|
|
c90874affd |
refactor(movement): extract ram to harness — phantom meteor dodge-only, rammer module via harness decision
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
ae9a5fec6d |
feat(movement): multi-enemy awareness — phantom meteor + minimum risk track all enemies
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
f06b263ae5 |
feat(phantom_meteor): aggressive ram — wider thresholds, desperation mode
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
ab473c6b68 |
feat(ModularBot): melee targeting — multi-enemy tracker, per-enemy gun fitness, radar auto-switch
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
30d00ba163 |
feat(movement): dodge timing, wall avoidance, distance control, ram finisher
PhantomMeteor: - Ram finisher: charge at enemy when <200px and their energy <10 - Ram opportunity: charge when <60px and we have >20 energy advantage - Integrated gunheat tracker for 1-2 tick earlier wave detection - Distance control: smooth linear ramp toward preferred engagement distance - Phantom range expanded 150→250px to catch closer threats WaveSurfer: - Wall-aware dodge bin selection: penalize bins leading off-arena - Dodge timing: predict future position 15 ticks ahead for safety - Distance control: radial blend when outside deadband (350±50px) - Wall escape: invert strafe if pushing further into wall, blend toward center ModularBot: - Wired KNN gun (purple/magenta) - Shadows tracked for movement (safer GF prediction) - Bullet lifecycle management (onBulletFired/onBulletHitBot/onBulletHitWall) - Unified phantom_meteor movement (wave_surfer unplugged) - Config logging on round start + gun switch Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
5bbc8cbda7 |
fix(gun_harness): min 15 obs before competing, window 50→100, KNN min k=5
Gun selector now gates competition: guns with <15 total observations across all power bins sit out until at least one gun qualifies. Falls back to ungated selection if no gun reaches threshold, preventing cold-start stalls. Sliding window increased from 50 to 100 ticks to reduce switching noise. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
651ce80620 |
feat(ModularBot): KNN gun, gunheat tracker, bullet shadows — inspired by DrussGT
New DrussGT-inspired modules: - KNNGun: K-nearest-neighbor statistical targeting using GF density peaks - GunheatTracker: dual-heat system (predicted + confirmed) for 1-2 tick lead - ShadowTracker: computes GF regions safe from in-flight bullets (enemy wave dodge) VirtualBodyTracker now integrates gunheat for earlier fire detection and shadows for safe-zone multiplier (90% reduction in danger zones). Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
6d16081fcc | style(ModularBot): ANSI colored log tags — gun yellow, move cyan | ||
|
|
63db0a4439 |
refactor: strip _garage suffix from adversarial test bots
Renames PatternMover_garage → PatternMover, RandomMover_garage → RandomMover, WaveSurfer_garage → WaveSurfer. Updates all .json, .sh, .nimble, and config.nims files to match TR Booter naming convention (directory name = bot name). Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
1eaaadf627 | refactor: move adversarial test bots to common_libs/test_framework/adversaries/ | ||
|
|
c0abd4c960 |
feat(ModularBot): virtual body movement selector — wave-based scoring replaces EMA damage
Implements VirtualBodyTracker (wave-based hit/miss scoring) instead of EMA damage accumulation. Movement switching now happens every tick, not every 3 rounds. Also refactors radar colors to dark teal (#004444/#0D4D4D) for faint visibility. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |