There are no genuinely competitive adversaries for Tank Royale, and the
in-repo ones were broken until recently. Classic Robocode 1.9.5.5 is
obtainable (SourceForge, 20.4 MB) and its programmatic control API
(robocode.control.RobocodeEngine + BattleAdaptor.onTurnEnded) can run real
battles headless and expose per-turn robot state. So a legacy leader bot's
MOVEMENT can be captured and used as a gun-testing fixture with no port.
Captured unmodified DrussGT 3.1.4159 vs spinbot/ramfire/crazy/corners and a
mirror match: 28,797 ticks, plus two trivial-bot contrasts. Conversion to
the Tank Royale convention is validated to 0.000-0.001 deg by recomputing
the direction implied by (heading, speed) and comparing it against the
recorded per-tick displacement -- i.e. the data is proven to be genuine
recorded motion rather than a mangled export. (A first attempt treated the
snapshot API's headings as degrees; they are radians, ~95 deg off.)
The statistics confirm it is really a wave surfer: perpendicular to the
opponent 65-96% of ticks, radial ~0.001, 41-72% of ticks at full speed,
reversing on 42-46% of ticks, holding range at a 283-526 px median. The
straight-line contrast is radial-dominant (0.75) with ZERO reversals.
Discovery: DrussGT detects predictable guns and switches to a bullet-shield
stand-still mode, so captures against sample.Walls/TrackFire had to be
rejected as non-movement.
CAVEATS, recorded in DRUSSGT_FIXTURES.md: these are open-loop (replayed
DrussGT never dodges OUR bullets) and perfect-information (the observer
gives true positions every tick, unlike our stale live WorldState). Both
make our guns look better than in live play, so use them for RELATIVE gun
ranking, not absolute hit rates.
Jars stay out of git; capture tooling is reproducible via capture.sh.
Gun evaluation previously required a full end-to-end battle (Java server +
battle runner + websocket IPC to 2 bot processes, 50 rounds, ~3.4 min) and
yielded only ~300-900 REAL shots across 13 guns -- far too few to rank
guns, which is why tuning needed many repetitions.
VirtualTracker is already a pure function of (WorldState stream, gun list);
the only reason it needed Java was where WorldState came from. So the range
replays a seq[WorldState] through the SAME tracker: offline and online
scores are the same metric by construction, not an approximation.
ACCEPTANCE TEST (the point of the whole thing): record one live round, replay
it offline, compare per-gun virtual hit rates. 12/12 deterministic guns match
EXACTLY, reproduced twice. Tsetlin is compared separately because tmLearnOne
calls rand(). Getting to 12/12 exposed two real ordering quirks in the live
loop: run() calls go() before the aim/fire block, so tickBullets resolves
against the NEXT tick's scan while the prediction used the previous one; and
if the target dies during that go() the final tick's spawn+resolution is
skipped entirely. The recorder emits an end marker for the second case.
The 5th (selected-gun) predict call was verified to be a no-op.
Measured cost: 8 fixtures (1770 ticks, ~92k virtual bullets, 13 guns) replay
in 2.9 s, ~32k virtual bullets/s -- roughly 70x faster and 100x more samples
than a live gauntlet.
Also adds a per-tick WorldState recorder behind const RecordWorldState
(default off, mirrors the ShotLog idiom) which records the state the bot
ACTUALLY builds, staleness included, rather than true positions -- recording
the latter would hand the guns perfect information and produce flattering
scores.
9 new guard checks (33 total, all passing), including fixture round-trip,
replay determinism, stationary->HeadOn 100%, constant-velocity->Linear>HeadOn,
and the energy-threshold turner crossing at t=41.
Measured, not assumed. With the gate temporarily opened to 20 deg, every
real shot was logged (tick, angle error at fire time, distance, power,
hit) across 3 gauntlets: 2611 shots, 57.3% aggregate. Findings:
- The geometric cone atan(BotRadius/d) is directionally confirmed but a
WEAK lever: even at 0.0-0.1 deg error the hit rate at 400-600px is only
~53-57%, because PREDICTION error dominates alignment error.
- Real effect of tightening the gate: 57.9% -> 68.0% aggregate hit rate
(fixed 0.1 deg), not the 76.9% previously reported -- that was a
high-variance draw (per-rep 62.8/66.4/77.2%).
- The shipped range-aware gate (SafetyFactor 0.6) does NOT beat the fixed
2.0 deg gate on hit rate (55.8% vs 57.9%, ~1.5 sigma, inside noise). It
fires 22-28% more shots and therefore lands more total hits (~509 vs
~434 per rep). No per-adversary score delta exceeded the 300-point
run-to-run noise band, so no config is demonstrably better on score.
Shipped anyway because it is strictly more expressive (a fixed threshold is
the special case), tunable from one const, and physically motivated, but
the honest verdict is recorded in-code: the gate is not the bottleneck.
AimThresholdDeg is removed; shouldFire now takes distPx. Degenerate or NaN
distance falls back to the ceiling rather than dividing by zero.
Also adds a per-shot logger to ModularBot behind 'const ShotLog' so the
measurement above is reproducible, and 10 new guard checks (24 total, all
passing) covering monotonicity, clamping, formula, perfect alignment,
gross misalignment and degenerate distance.
Cross-checked against the server source: the gun fires BEFORE the turn is
applied, so the logged angle error is the true departure error, and
fireAssist auto-aim is off (unset by the Nim API and forced false by
setAdjustRadarForGunTurn).
tm-learning-tracks.md covers three things, all marked [FACT]/[INFERENCE]/
[UNKNOWN]:
- Section A: the Tsetlin gun's label is measured against the wrong baseline.
predX = linearX + cx, so rx = actual - predX = delta - cx, and inside
tmLearnOne error = residual - predicted = (delta - cx) - cx = delta - 2cx.
The fixed point is cx = delta/2 -- HALF the correction needed, even with
perfect Granmo feedback. Fix: store linearX/linearY in TmTrace and train
on delta. Also: hits zero the label instead of carrying their true
residual, and the per-clause step is magnitude-blind.
- Section B: what a TM is actually good at (AND-clauses over binary
literals, readable output) and why this repo suits it -- the gun already
builds an 83-bit x 10-frame Gray-coded window (870 bits). Includes a
falsifiable known-rule benchmark proposal.
- Section C: delayed-reward learning belongs to the MOVEMENT layer, not the
gun. The gun's outcome is delayed but exactly pairable via
(fireTick, powerBin), so its effective lambda is 1 and discounting would
only destroy information.
Also records that docs/papers/tm-deb-paper.pdf was deleted by the user as
AI-generated and unverifiable, while the Granmo-based feedback diff in
tm-deb-assessment.md stands on its own.
Verdict: the document (docs/papers/tm-deb-paper.pdf, 'Generated by
Gemini Notebook') solves temporal credit assignment under delayed
reward, which is not our problem. Our gun's failure is clause
saturation (~131 of 1740 literals included per clause -> conjunction
fires with probability ~2^-131 -> correction identically 0), and
TM-DEB only scales update FREQUENCY by gamma^dt, so it would leave
the fixed point untouched and additionally delete the long-range
feedback we need: gamma^90 = 4.4e-7 at the paper's own gamma=0.85,
and those are the shots whose lead matters most.
Credibility signals recorded in the doc: reference [2] misattributes
authors/venue/year; Algorithm Spec 2 steps 9-12 are literal '%'
placeholders so the automata update is simply absent; Table 1 is
titled 'Expected' and reports never-measured accuracies; Eq. 12 is
not a faithful copy of Granmo's Lemma 2.
The audit also produced the actionable result: a line-by-line diff of
Granmo Table 2/3 feedback against tmLearnOne, identifying why the
automata saturate - Type I never conditions on the clause output so it
omits Granmo's c=0, lk=1 -> -1 w.p. 1/s counter-force (true literals
ratchet toward Include), Type II's guard cOut==1 AND lits[lit]==0 is
unsatisfiable and therefore dead code, and the (T - clip(v,-T,T))/(2T)
resource allocation is missing entirely.
Four guns cached a whole prediction per tick while predict() is called once
per power bin, so every bin after the first (and the real fired shot, which
shares lastState) reused the power-1.0 lead. Fixed by caching only the
speed-INDEPENDENT derived state and recomputing the lead per requested speed:
- stop_shot: also fixes prevSpeed being written before it was read, which
made abs(speed) < abs(prev) permanently false and the entire
stop-prediction branch unreachable (it was just Linear).
- displacement: the cache key included bulletSpeed, so the guard missed on
all four bins and the 15-tick window advanced ~4x/tick, making the
inferred velocity ~4x too small.
- averaged_lead: tick cache removed outright. pattern_matcher: split into
speed-independent match+path and per-call lead.
FeedbackEvent gains fireTick/powerBin (additive; only virtual_bullets
constructs one) so guns can pair feedback to the exact shot instead of
guessing by coordinates. tsetlin uses it: traces are now keyed exactly by
(fireTick, powerBin) with a 1024-slot ring, and the 10-frame window shifts
at most once per tick (it was shifting ~4-5x/tick, so isWarmedUp tripped
after ~2 ticks).
KNOWN INCOMPLETE: tsetlin still does not diverge from Linear in battle. The
two named bugs are fixed (a 600-tick sim shows trainedShots=2141,
traceMisses=0, and a fixed-input probe converges to a 9.6px correction), but
the TM's clause feedback itself is broken: ~131 of 1740 literals end up
included per clause, so its conjunction never fires. Sweeping TM_S,
TM_N_CLAUSES and a two-branch Type-I update did not change the correction
from 0. Needs a real TM fix or removal, not another bug fix.
First-ever guard tests for the gun selector: common_libs/tests/
test_gun_harness.nim (14 checks, headless, no Java). There were none before,
which is how six broken guns survived a full analysis cycle. Against the
previous HEAD, 5 of these checks FAIL - that is the regression guard.
Wave queues (guess_factor, decay_gf, knn_gun): predict() stored ONE wave
per tick while onResult() popped one per resolved bullet (~4/tick), so the
queue drained to empty within a few dozen ticks, ~3 of every 4 resolutions
returned without learning, and the survivor paired with a same-tick wave
(bearingDelta ~= 0) pinning the histogram at centre. PROOF: GF.vHits ==
HeadOn.vHits and DecayGF.vHits == HeadOn.vHits byte-for-byte in every one
of 50 rounds — the guns had degenerated to HeadOn.
Now each gun keeps a per-bin FIFO with an O(1) head cursor. At most one
push per (tick, bin) so the fire site's 5th predict() call is a no-op, and
onResult pops the oldest wave of its OWN bin via e.bulletPower. Aiming
math untouched (it was already correct: 0 deg = East, CCW+).
maxBullets 2048 -> 8192: the rack spawns 52 bullets/tick so the ring wrapped
every ~39 ticks while a long power-3 shot needs ~90, silently discarding
unresolved bullets and biasing every measured hit rate by range. Added a
droppedBullets counter so a future overflow is measurable, and wavePushes/
waveStarved counters on the three guns. After the fix: vDropped = 0 and
vStarved = 0 across all 48 recorded rounds.
fitnessFor is now exported, deterministic (enemies iterated in ascending id
order) and shared by the selector and the stats dump, replacing a hand-rolled
merge in ModularBot that never advanced its window head.
Round lines gain additive keys: vDropped, vStarved.
nimble bin output lands at the garage root, so the existing
'*_garage/out/' ignore rule never covered it and every build dirtied
the tree with a 1.1 MB binary. File stays on disk; build regenerates it.
Attribution is proven, not guessed: the server assigns a per-round-unique
bulletId (GunEngine.nextBulletId) and stamps the same id on BulletFired,
BulletHitBot, BulletHitWall and BulletHitBullet. Keep a FIFO of fired gun
ids, stamp bulletId -> gunId on onBulletFired, resolve through that map.
Hits are deferred when onBulletHit precedes onBulletFired in the same
turn (client dispatches priority 70 > 60), which recovered 14
unattributed hits. 99.9% of shots and 99.8% of hits attributed.
Stats lines now carry per-gun realShots/realHits/realHitRate; the old
keys and round-level totals are unchanged.
bestPower: a gun with zero observations in every bin previously returned
the HIGHEST bin (power 3.0) because an empty bin satisfied the
'count == 0' clause on the first countdown iteration. Cold guns now
return the lowest bin as the docstring always claimed. Warm-gun path
untouched.
- bestGun: replace first-index-wins argmax with random pick among guns
within TieMargin (2%) of best rate. HeadOn at index 0 was silently
winning every tie, starving Tsetlin/Linear/etc.
- MinObsBeforeCompete 15 -> 50 (Pattern entered competition on noise)
- add MinHitRateFloor 0.10: if no gun clears it, fall back to HeadOn
instead of selecting the best of a bad field
- ModularBot: remove AntiSurfer gun (0% virtual hit rate everywhere),
14 -> 13 guns, renumber ids and selection counters
- 30-tick cooldown after ghost-stuck/timeout ram exit prevents re-entry loop
- enemy_tracker.update() skips dead bots to prevent same-tick scan resurrection
- TFIL graphics cleared when ramming is active movement
- [config] logs: white base with green-highlighted changes only
- [ram:enter] logs trigger reason and key values on false→true transition
- [death] and [target-invalid] logs retained for diagnostics
PhantomMeteor:
- Ram finisher: charge at enemy when <200px and their energy <10
- Ram opportunity: charge when <60px and we have >20 energy advantage
- Integrated gunheat tracker for 1-2 tick earlier wave detection
- Distance control: smooth linear ramp toward preferred engagement distance
- Phantom range expanded 150→250px to catch closer threats
WaveSurfer:
- Wall-aware dodge bin selection: penalize bins leading off-arena
- Dodge timing: predict future position 15 ticks ahead for safety
- Distance control: radial blend when outside deadband (350±50px)
- Wall escape: invert strafe if pushing further into wall, blend toward center
ModularBot:
- Wired KNN gun (purple/magenta)
- Shadows tracked for movement (safer GF prediction)
- Bullet lifecycle management (onBulletFired/onBulletHitBot/onBulletHitWall)
- Unified phantom_meteor movement (wave_surfer unplugged)
- Config logging on round start + gun switch
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Gun selector now gates competition: guns with <15 total observations across
all power bins sit out until at least one gun qualifies. Falls back to ungated
selection if no gun reaches threshold, preventing cold-start stalls.
Sliding window increased from 50 to 100 ticks to reduce switching noise.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
New DrussGT-inspired modules:
- KNNGun: K-nearest-neighbor statistical targeting using GF density peaks
- GunheatTracker: dual-heat system (predicted + confirmed) for 1-2 tick lead
- ShadowTracker: computes GF regions safe from in-flight bullets (enemy wave dodge)
VirtualBodyTracker now integrates gunheat for earlier fire detection and shadows
for safe-zone multiplier (90% reduction in danger zones).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Renames PatternMover_garage → PatternMover, RandomMover_garage → RandomMover,
WaveSurfer_garage → WaveSurfer. Updates all .json, .sh, .nimble, and config.nims
files to match TR Booter naming convention (directory name = bot name).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Implements VirtualBodyTracker (wave-based hit/miss scoring) instead of EMA damage accumulation. Movement switching now happens every tick, not every 3 rounds. Also refactors radar colors to dark teal (#004444/#0D4D4D) for faint visibility.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Adds WaveSurferModule as second movement option
- Tracks damage per round via onHitByBullet, EMA decay α=0.5
- Every round after round 2: switches to lower-damage movement
- Body color signals active mover (orange=phantom, blue=wave)
- vs SpinBot 10r: ModularBot 1412 vs SpinBot 650
- vs WaveSurfer 10r: ModularBot 1858 vs WaveSurfer 10
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- GF gun: triangular head-on prior replaces flat bins (was always aiming at GF=-1)
- GF gun: recompute MEA per bullet power in onResult (was using stale first-bin value)
- Scores improved: SpinBot 1805-48, Crazy 1557-46, WaveSurfer 1868-10, PatternMover 1846-4
- All 6 guns active in selection across gauntlet
- Cold-start bug: uniform bins[0..30]=0.1 made peakBin() always return
0 (first-wins tie), giving GF=-1 (max CW escape) before any learning.
Fixed with a triangular head-on bump at bin 15 (GF=0) as the prior.
- onResult now recomputes mea from FeedbackEvent.bulletPower instead of
the stale first-bin mea cached by predict; correct per-power-bin GF.
- Add DebugGF const (default false) with [gf-dbg] echoes in predict/onResult.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Two bugs fixed:
1. Multi-tick gaps: turn rate assumed 1 tick between observations, but scans can be 5+ ticks apart. Now divides by actual tickDelta.
2. Per-power-bin state corruption: predict() called 4x per tick (per power bin). After first call, prevHeading was already updated, causing subsequent calls to compute 0° delta. Now captures oldHeading/oldTick before updating.
Verified: OscillatorBot at 4°/tick captured correctly; normalization [-180°,180°] works; tickDelta=1 typical; first bin gets delta, subsequent bins see 0 (expected).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace binary hit/miss reward with exponential decay based on miss distance:
- reward = 2.0 * exp(-missDistance / 36.0) - 1.0
- At 0px: +1.0 (perfect hit)
- At 36px: -0.26 (near miss, small penalty)
- At 100px: -0.87 (big miss, large penalty)
Maintains virtual hit/miss counters for display (threshold: 36px).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Swap sin/cos in virtual bullet position calc to match Tank Royale convention
- Change exploration from flipping each bit with 50% chance to flipping exactly 1 random bit
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>