5 Commits

Author SHA1 Message Date
SirStone e40c8493a6 Offline harness: audited, calibrated against live, and one real bug fixed
AUDIT (docs/offline_harness_trust.md, new):
- Re-ran acceptance_offline_vs_online myself TWICE: 12/12 deterministic guns
  exact both times (264 ticks/enemyId=1, 244 ticks/enemyId=2), death boundary
  included. The offline range reproduces the live bot's own per-gun virtual
  telemetry exactly.
- Re-verified the (fireTick, powerBin) wave-pairing fix: exact-key lookup,
  collisions counted not silently mislabelled; test_wave_pairing 17/17 PASS.
- The offline score is the live TELEMETRY (last-100 virtual hit rate) but NOT
  the live BATTLE score (damage/round wins). Two-level answer, documented.
- bmPoint scores up to one tick-step (~17px) PAST its documented aim distance,
  while the tie-break probe scores exactly the aim point. Real, low-impact,
  deliberately NOT fixed (point metric is non-default, measured negative, and
  the committed point baselines would silently change).
- bmPoint/bmPath, perfect-info captures, conditional-on-selection live rates,
  and hit-rate-as-objective-for-movement all catalogued as non-apples comparisons.

FIX (unambiguous, fail-before/pass-after):
- common_libs/tests/range_guns.nim: buildAllGunDrivers defaulted to
  enableTmSelector=true, so run_range / analyze_selector / test_power_selection /
  measure_power_policy spawned gun 13 (TMSelect) - a gun the shipped bot NEVER
  spawns. The shared VirtualTracker ring is order-sensitive, so those 4
  spawns/tick permuted the learning guns' resolution order (the exact confound
  4cd5618 fixed for the acceptance test, left broken for every default caller).
  Default is now false (mirror the shipped rack). Impact on
  tr_drussgt_vs_modularbot: Tsetlin 18.8->18.5%, KNN 7.5->7.2%, TMSelect 15.2->0.
- New guard common_libs/tests/test_range_rack_parity.nim (3 checks); proven to
  FAIL before and PASS after by stash-reverting the fix.

CALIBRATION (offline prediction vs live outcome, 9 usable arms):
- Direction agreement 3/9 = 33%. Split by domain: open-loop (single-tick
  prediction / metric / threshold) 3/3; closed-loop (adaptation / range /
  movement / selection) 0/6. Small, non-random, hand-assembled set - no
  correlation coefficient is claimed.
- The four motivating "offline wins" re-attributed: ring mover was NEVER
  offline (it is a live server-side hit rate, mislabelled "offline" in
  env_reference.md:342 and commit 7f6ccfb); TMHorizon window/NSTATES and the TM
  gun are the H3 classifier-accuracy harness (not hit rate); TFIL is the H2
  open-loop movement replay, whose mechanism prediction was right and whose
  outcome prediction was wrong.
- Open-loop hypothesis tested: TR_RACK_* knobs leave the offline range output
  BYTE-IDENTICAL (the replay never calls the selector), and the range has no
  driver for guns 14/15 (TMPATTERN/TMHORIZON). BUG vs LIMIT separated.

VERDICT: trust the harness for single-tick prediction quality only; never for
anything running through the closed loop. MEASURED vs INFERRED labelled.

Green counts unchanged: test_gun_harness 39, test_vbullet_metric 11,
test_power_selection 3, test_power_policy 58, test_adaptive_radar 41,
test_tfil_ring_weights 24, test_ram_decision 40, test_rack_membership 48,
test_selector_tiebreak 19, test_tm_pattern_registration 20,
test_vbullet_admit_gate 12, test_tm_horizon 104, test_tm_diag 48,
test_tm_automata_diag 55, test_tm_clause_shape 66, test_env_report 25,
test_tfil_commit_env 30 (as-is). New: test_range_rack_parity 3.
2026-09-24 21:40:38 +02:00
SirStone 4cd5618435 fix(test): repair the flaky offline==online acceptance; make the tie-break truly random
TASK 2 - THE FLAKY ACCEPTANCE TEST, root-caused. It was NOT a live/offline
boundary race as suspected. The replay spawned gun 13 (TMSelect) while live has
EnableTmSelector = false and never does. The shared VirtualTracker ring is
ORDER-SENSITIVE, so gun 13's extra 4 bullets/tick shift the ring head and
permute the per-tick RESOLUTION ORDER of every other gun. The learning guns
append observations in resolution order, so their predictions shifted and
produced small hit deltas that moved between runs.
Evidence: the first KNN divergence was at rtick=174 with the SAME resolution
set merely reordered (live ft133,138,139,142,148,150,151 vs offline
ft150,151,133,138,139,142,148); after closing gun 13's ready gate offline the
live and offline KNN traces became BYTE-IDENTICAL (diff empty, 904/904 lines).
Fix: mirror the live rack in the replay. No tick exclusion, no tolerance
loosening. Stability: 5/5 consecutive runs now report 12/12 exact, each with
enemyDied=true - the death boundary is included, not excluded. The proof is
now real rather than a lucky run.

TASK 1 - the tie-break was not random. randomize() was only reached
incidentally through initTsetlinGun(), so a rack without Tsetlin had a fixed
rand() stream and ties always resolved the same way across process restarts.
Added seedSelectorRng() after gun construction, honouring GUN_SELECTOR_SEED.
Evidence: unseeded, 6 separate processes gave different pick sequences; with
GUN_SELECTOR_SEED=42, 3 processes gave identical sequences.

TASK 3 - PRUNING DOES NOT HELP; keep the full rack. 15 PAIRED runs per variant
vs DrussGT, 8 rounds, identical seeds:
  baseline             3238 shots  6.18%  (events 6.16%)  200 dmg/run
  Tsetlin disabled     3522 shots  5.76%  (events 5.71%)  197 dmg/run
  Tsetlin+Displace     3478 shots  5.46%  (events 5.37%)  183 dmg/run
Paired permutation tests: -0.34pp p=0.57 and -0.70pp p=0.21. Per-run
distributions completely overlap (baseline range [2.68, 10.00]; 15/15 and 14/15
runs inside it). A Crazy control showed no separation either. So removing the
measured-worst real performers is neutral-to-slightly-negative, and with
sd ~1.8pp a definitive claim either way would need far more runs.

CORRECTION TO A CLAIM I MADE: the 'virtual metric is INVERTED' finding does NOT
reproduce. Job-24 measured Spearman -0.374; this job measures +0.335 over the
same 13 guns with a different but equally defensible aggregation. Two opposite
signs means the correlation is NOT robustly negative - it is WEAK AND
SIGN-UNSTABLE. The honest statement is that virtual hit rate is a poor ranker,
not an inverted one. The docs assert the inversion and need correcting.

Also adds per-process GUN_STATS_PATH/GUN_SHOTLOG_PATH so concurrent A/B runs do
not clobber each other, and an env-gated GUN_RACK_DISABLE for rack A/Bs. All
default behaviour is unchanged when the env vars are unset.
2026-09-21 06:56:44 +02:00
SirStone 57b2ac3849 feat(guns): scale-aware power selection (+52% damage); TM classifier gun built, measured, DISABLED
TASK 2 - power selection, a clear win. bestPower used an ABSOLUTE
MinHitRate = 0.40 bar. Measured per-bin virtual rates (rolling-100 fraction)
show no bin ever clears 40%, so 11 of 14 guns were stuck at bin 0 (power 1.0)
even where higher bins were comparable:
  Linear  p1.0 44% p1.5 39% p2.0 30% p3.0 29%   old bin 0 -> new bin 3
  Accel   p1.0 44% p1.5 40% p2.0 26% p3.0 29%   old bin 1 -> new bin 3
  Pattern p1.0 50% p1.5 40% p2.0 27% p3.0 12%   old bin 1 -> new bin 2
Replaced with a scale-aware PowerBarFrac = 0.50 (a dimensionless FRACTION of
the gun's own best bin rate). 13 of 14 selections now pick heavier bullets.
Real effect vs DrussGT (8 rounds x 3 runs): hit rate unchanged (7.56% ->
7.47%) but damage dealt +52% (157 -> 239 per run) and rounds end faster.
Same accuracy, half the shots, half again more damage.

TASK 1 - the TM pattern-classifier gun does NOT earn its slot. It was built as
a mixture of experts with a corrected-Granmo TM as a multi-class gate over
HeadOn/Linear/Circular/WallBounce/Accel, labelled by which expert's prediction
was closest to the actual enemy position (an exact, supervised, per-shot
label - no delayed credit). Offline it loses to the best of its OWN experts on
essentially every fixture, and against DrussGT it cost real performance:
  baseline (path+relative)  7.56% real hit rate, damage 157
  + power fix               7.47%,                 damage 239
  + power fix + TM gun      5.59%,                 damage 133
The gun was selected on 806 ticks and fired 24 real shots at 4.2%.
So the tree ships with EnableTmSelector = false: code and wiring kept intact
for re-enabling, but it is not in the active rack.

Worth recording from the clause dump: the gate DOES latch onto meaningful
structure. On energy-threshold-turner, HeadOn's clauses key on the energy bits
(the rule's own driving variable) while Circular keys on distance/velocity. So
the TM is learning something real and interpretable - it simply cannot beat
'always pick the best expert'. Root cause (INFERRED): the closest-expert label
is noisy because several experts are near-tied, and under the path metric the
winner varies by power bin while the gate sees one shared per-tick input, so a
one-vs-rest gate over a saturated 870-bit clause space has no margin to exploit.
(Zero-padding the 2-frame window was tried first and saturated every clause at
256-755 included literals; alternating the two real frames fixed that.)

Also factors the corrected feedback into an exported tmLearnDir and exports the
encoding/TM primitives; the Tsetlin tests still reproduce the documented
mean=13.8 included literals, so the refactor is behaviour-preserving.

Verified: 33/33 guard checks, tsetlin tests green, metric checks green, new
power-selection guard green (13/14 selections change; relative bar still picks
bin 1 and not bin 3 for a [30,25,12,5]% profile), 12/12 offline==online
acceptance under the shipped default.
2026-09-21 05:19:07 +02:00
SirStone 89370008da fix(tsetlin): make the TM actually learn - saturation 714 -> 13.8 literals/clause
The gun has never contributed anything: Tsetlin.vHits was byte-for-byte
equal to Linear.vHits in every measured round of every run, because its
learned correction was always exactly 0.

Six diagnosed defects fixed, plus one that was required to make the first
one work:

1. Type I now conditions on the clause output. It previously rewarded
   included true literals unconditionally, omitting Granmo's (c=0, lk=1)
   -> toward Exclude counter-force, so true literals ratcheted toward
   Include forever. This was the root cause of the saturation.
2. Type II was unreachable dead code: its guard required cOut==1 AND
   lits[lit]==0 AND st>0 (included), but cOut==1 guarantees every included
   literal is 1. Its direction was wrong too - it should increment EXCLUDED
   false literals when the clause fires.
3. Resource allocation restored: Granmo's (T - clip(v,-T,T))/(2T) target
   replaces |error|/(2*RESID_MAX); TM_T was only an output normaliser.
4. Label baseline fixed - the factor-2 shrink. predX = linearX + cx, so the
   label was delta - cx while the learner's output IS cx, giving
   error = delta - 2cx and a fixed point of cx = delta/2: HALF the needed
   correction even with perfect feedback. TmTrace now stores linearX/linearY
   and training uses delta.
5. Hits no longer zero their label (a hit means |miss| < 18px, not 0).
6. The enemy-energy feature was duplicated - tmEncodeFrame passed
   state.selfEnergy with a stale comment claiming enemyEnergy was absent,
   while WorldState.enemyEnergy exists. Enemy-energy rules were literally
   unrepresentable.
7. REQUIRED EXTRA: tmEvalClause now implements Granmo Eq. 6 - an all-Exclude
   clause outputs 1 during learning and 0 during classification. Without it,
   fix #1 deadlocks every clause at empty.

MEASURED EFFECT (energy-threshold-turner fixture, seed 1):
  mean included literals per active clause   714.0 -> 13.8
  active clauses                             100/100 -> 53/100
  nonzero corrections                        8/764 -> 708/764
  Tsetlin virtual hits (Linear = 27/400)     27/400 -> 69/400

Divergence achieved: offline on 7/8 fixtures, and in a live gauntlet
(RandomMover: Tsetlin 199/1200 vs Linear 288/1200, vDropped=vStarved=0).
Tsetlin now LEARNS but is not yet competitive with Linear - the regression
head is untuned, flagged as follow-up rather than claimed as a win.

Also ignores compiled test harnesses that have no file extension, which the
existing '**/tests/test_*' rule misses.
2026-09-20 23:59:32 +02:00
SirStone 974528d5cf feat(gun_harness): offline gun range, proven equivalent to live play
Gun evaluation previously required a full end-to-end battle (Java server +
battle runner + websocket IPC to 2 bot processes, 50 rounds, ~3.4 min) and
yielded only ~300-900 REAL shots across 13 guns -- far too few to rank
guns, which is why tuning needed many repetitions.

VirtualTracker is already a pure function of (WorldState stream, gun list);
the only reason it needed Java was where WorldState came from. So the range
replays a seq[WorldState] through the SAME tracker: offline and online
scores are the same metric by construction, not an approximation.

ACCEPTANCE TEST (the point of the whole thing): record one live round, replay
it offline, compare per-gun virtual hit rates. 12/12 deterministic guns match
EXACTLY, reproduced twice. Tsetlin is compared separately because tmLearnOne
calls rand(). Getting to 12/12 exposed two real ordering quirks in the live
loop: run() calls go() before the aim/fire block, so tickBullets resolves
against the NEXT tick's scan while the prediction used the previous one; and
if the target dies during that go() the final tick's spawn+resolution is
skipped entirely. The recorder emits an end marker for the second case.
The 5th (selected-gun) predict call was verified to be a no-op.

Measured cost: 8 fixtures (1770 ticks, ~92k virtual bullets, 13 guns) replay
in 2.9 s, ~32k virtual bullets/s -- roughly 70x faster and 100x more samples
than a live gauntlet.

Also adds a per-tick WorldState recorder behind const RecordWorldState
(default off, mirrors the ShotLog idiom) which records the state the bot
ACTUALLY builds, staleness included, rather than true positions -- recording
the latter would hand the guns perfect information and produce flattering
scores.

9 new guard checks (33 total, all passing), including fixture round-trip,
replay determinism, stationary->HeadOn 100%, constant-velocity->Linear>HeadOn,
and the energy-threshold turner crossing at t=41.
2026-09-20 23:44:42 +02:00