7632aaba06950bbc5dd33dcb021fcc0a2c85e5de
6 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
3142b70aa5 |
Gate virtual-bullet spawn on rack admission: +68% tick rate, selected gun unchanged
The default rack is now Pattern-only (
|
||
|
|
c091bf3c34 |
harness: upgrade to server 1.3.1, keep 0.35.5 selectable, re-baseline
All prior measurements ran on server 0.35.5. The default is now the current 1.3.1 jar, with the legacy jar kept and switchable via TR_SERVER_JAR (no code edit). test_gauntlet_5bots.nim no longer clobbers a caller's TR_SERVER_JAR - it used to putEnv() unconditionally, so an override was silently ignored. RE-BASELINE (controlled RulesProbe battle, stationary bot, powers 0.1/0.5/1/2/3): dimension 1.3.1 0.35.5 verdict bullet damage per hit 0.4/2/4/10/16 identical SAME bullet speed (20-3p) within noise within noise SAME post-fire gun heat (1+p/5) identical identical SAME cooling 0.1/tick 0.1/tick SAME bulletDamage SCORE exactly 100/round 104..113/round DIFFERENT bulletKillBonus (20%) 20/round 20..23/round DIFFERENT LOUD FINDING - a SCORING rule changed, physics did not: 0.35.5 credits OVERKILL to bulletDamage (the killing bullet's full damage even past 0 energy); 1.3.1 caps it at the energy actually removed. Every 0.35.5 score is therefore inflated ~5-6%, and bulletKillBonus inherits the inflation. Gauntlet totals shift accordingly (SittingDuck 1936 -> 1800, WaveSurfer 1886 -> 1669). Consequence: score-based numbers recorded on 0.35.5 are NOT comparable to 1.3.1. Our gun A/Bs used real HIT RATE, not score, so those conclusions stand. Runner 1.0.2 (unchanged, no newer one on the box) is measured compatible with the 1.3.1 server. Note TrBattleCapture uses the runner's EMBEDDED server, which is 1.0.2 - so the capture path still runs an older engine than the gauntlet. Also re-ran acceptance_offline_vs_online on the new default: 12/12. |
||
|
|
4cd5618435 |
fix(test): repair the flaky offline==online acceptance; make the tie-break truly random
TASK 2 - THE FLAKY ACCEPTANCE TEST, root-caused. It was NOT a live/offline boundary race as suspected. The replay spawned gun 13 (TMSelect) while live has EnableTmSelector = false and never does. The shared VirtualTracker ring is ORDER-SENSITIVE, so gun 13's extra 4 bullets/tick shift the ring head and permute the per-tick RESOLUTION ORDER of every other gun. The learning guns append observations in resolution order, so their predictions shifted and produced small hit deltas that moved between runs. Evidence: the first KNN divergence was at rtick=174 with the SAME resolution set merely reordered (live ft133,138,139,142,148,150,151 vs offline ft150,151,133,138,139,142,148); after closing gun 13's ready gate offline the live and offline KNN traces became BYTE-IDENTICAL (diff empty, 904/904 lines). Fix: mirror the live rack in the replay. No tick exclusion, no tolerance loosening. Stability: 5/5 consecutive runs now report 12/12 exact, each with enemyDied=true - the death boundary is included, not excluded. The proof is now real rather than a lucky run. TASK 1 - the tie-break was not random. randomize() was only reached incidentally through initTsetlinGun(), so a rack without Tsetlin had a fixed rand() stream and ties always resolved the same way across process restarts. Added seedSelectorRng() after gun construction, honouring GUN_SELECTOR_SEED. Evidence: unseeded, 6 separate processes gave different pick sequences; with GUN_SELECTOR_SEED=42, 3 processes gave identical sequences. TASK 3 - PRUNING DOES NOT HELP; keep the full rack. 15 PAIRED runs per variant vs DrussGT, 8 rounds, identical seeds: baseline 3238 shots 6.18% (events 6.16%) 200 dmg/run Tsetlin disabled 3522 shots 5.76% (events 5.71%) 197 dmg/run Tsetlin+Displace 3478 shots 5.46% (events 5.37%) 183 dmg/run Paired permutation tests: -0.34pp p=0.57 and -0.70pp p=0.21. Per-run distributions completely overlap (baseline range [2.68, 10.00]; 15/15 and 14/15 runs inside it). A Crazy control showed no separation either. So removing the measured-worst real performers is neutral-to-slightly-negative, and with sd ~1.8pp a definitive claim either way would need far more runs. CORRECTION TO A CLAIM I MADE: the 'virtual metric is INVERTED' finding does NOT reproduce. Job-24 measured Spearman -0.374; this job measures +0.335 over the same 13 guns with a different but equally defensible aggregation. Two opposite signs means the correlation is NOT robustly negative - it is WEAK AND SIGN-UNSTABLE. The honest statement is that virtual hit rate is a poor ranker, not an inverted one. The docs assert the inversion and need correcting. Also adds per-process GUN_STATS_PATH/GUN_SHOTLOG_PATH so concurrent A/B runs do not clobber each other, and an env-gated GUN_RACK_DISABLE for rack A/Bs. All default behaviour is unchanged when the env vars are unset. |
||
|
|
57b2ac3849 |
feat(guns): scale-aware power selection (+52% damage); TM classifier gun built, measured, DISABLED
TASK 2 - power selection, a clear win. bestPower used an ABSOLUTE MinHitRate = 0.40 bar. Measured per-bin virtual rates (rolling-100 fraction) show no bin ever clears 40%, so 11 of 14 guns were stuck at bin 0 (power 1.0) even where higher bins were comparable: Linear p1.0 44% p1.5 39% p2.0 30% p3.0 29% old bin 0 -> new bin 3 Accel p1.0 44% p1.5 40% p2.0 26% p3.0 29% old bin 1 -> new bin 3 Pattern p1.0 50% p1.5 40% p2.0 27% p3.0 12% old bin 1 -> new bin 2 Replaced with a scale-aware PowerBarFrac = 0.50 (a dimensionless FRACTION of the gun's own best bin rate). 13 of 14 selections now pick heavier bullets. Real effect vs DrussGT (8 rounds x 3 runs): hit rate unchanged (7.56% -> 7.47%) but damage dealt +52% (157 -> 239 per run) and rounds end faster. Same accuracy, half the shots, half again more damage. TASK 1 - the TM pattern-classifier gun does NOT earn its slot. It was built as a mixture of experts with a corrected-Granmo TM as a multi-class gate over HeadOn/Linear/Circular/WallBounce/Accel, labelled by which expert's prediction was closest to the actual enemy position (an exact, supervised, per-shot label - no delayed credit). Offline it loses to the best of its OWN experts on essentially every fixture, and against DrussGT it cost real performance: baseline (path+relative) 7.56% real hit rate, damage 157 + power fix 7.47%, damage 239 + power fix + TM gun 5.59%, damage 133 The gun was selected on 806 ticks and fired 24 real shots at 4.2%. So the tree ships with EnableTmSelector = false: code and wiring kept intact for re-enabling, but it is not in the active rack. Worth recording from the clause dump: the gate DOES latch onto meaningful structure. On energy-threshold-turner, HeadOn's clauses key on the energy bits (the rule's own driving variable) while Circular keys on distance/velocity. So the TM is learning something real and interpretable - it simply cannot beat 'always pick the best expert'. Root cause (INFERRED): the closest-expert label is noisy because several experts are near-tied, and under the path metric the winner varies by power bin while the gate sees one shared per-tick input, so a one-vs-rest gate over a saturated 870-bit clause space has no margin to exploit. (Zero-padding the 2-frame window was tried first and saturated every clause at 256-755 included literals; alternating the two real frames fixed that.) Also factors the corrected feedback into an exported tmLearnDir and exports the encoding/TM primitives; the Tsetlin tests still reproduce the documented mean=13.8 included literals, so the refactor is behaviour-preserving. Verified: 33/33 guard checks, tsetlin tests green, metric checks green, new power-selection guard green (13/14 selections change; relative bar still picks bin 1 and not bin 3 for a [30,25,12,5]% profile), 12/12 offline==online acceptance under the shipped default. |
||
|
|
d5061ee215 |
test(range): restore the 12/12 offline==online proof; measure TM clause readability
Task 1 - the acceptance proof was unrunnable because RecordWorldState was a
compile-time const set to false. It is now a RUNTIME switch
(let RecordWorldState* = existsEnv("TR_RECORD_WORLDSTATE")), default OFF, so
ordinary runs write no fixture, and acceptance_offline_vs_online.nim enables
it for the battle it spawns and clears it afterwards. Restored and run twice:
12/12 deterministic guns match exactly (128-tick and 546-tick battles), with
Tsetlin reported separately as stochastic. Both nimble build variants clean.
Task 2 - does a compact encoding turn the TM's 99.35% into a READABLE rule?
Measured across window sizes (fixed seed, no tuning):
frames TEST acc eff.lits/clause firing clauses counterfactual low/high/mean
10 99.35% 152.8 37 100/24/62.4%
3 95.94% 54.9 35 96/20/58.6%
2 99.48% 39.6 38 95/25/60.7%
1 98.30% 19.2 45 100/24/62.3%
So 2 frames is strictly better than 10 on BOTH axes: +0.13 accuracy for 4x
smaller clauses. The 3-frame dip is non-monotonic and left unexplained rather
than smoothed over.
A readable rule WAS partially recovered. Five clauses carry the exact Gray
form !g10 ^ !g9 ^ !g8; g10 is inert in this data, so the effective rule is the
2-literal proposition !g9 ^ !g8, i.e. energy < 25.6. That is a genuine
threshold in readable propositional form - but at 25.6, NOT the labelled 30,
because 256 is a power-of-two Gray boundary expressible in two literals while
300 needs a longer conjunction. The TM found the nearest SIMPLE threshold.
The honest caveat: that threshold is not the ensemble's decision mechanism.
The counterfactual follow rate (high 24%, mean 62.3%) is statistically
identical at 1, 2 and 10 frames, so compactness did not make the model read
energy - its vote is carried by co-occurring bearing/velocity/heading/wall
literals. Also identified: clauses containing all 11 Gray energy bits are
satisfied at exactly one raw value (50, the dataset floor), so they are
'energy has hit the floor' detectors, not thresholds.
Methodological fix worth keeping: the earlier single-frame counterfactual
wrote energy into all 10 frame slots including the zeroed ones, reviving dead
clauses and producing a spurious 2% high-follow rate. setEnergyFrames now
rewrites only the exposed frames; the corrected figure is 24%.
|
||
|
|
974528d5cf |
feat(gun_harness): offline gun range, proven equivalent to live play
Gun evaluation previously required a full end-to-end battle (Java server + battle runner + websocket IPC to 2 bot processes, 50 rounds, ~3.4 min) and yielded only ~300-900 REAL shots across 13 guns -- far too few to rank guns, which is why tuning needed many repetitions. VirtualTracker is already a pure function of (WorldState stream, gun list); the only reason it needed Java was where WorldState came from. So the range replays a seq[WorldState] through the SAME tracker: offline and online scores are the same metric by construction, not an approximation. ACCEPTANCE TEST (the point of the whole thing): record one live round, replay it offline, compare per-gun virtual hit rates. 12/12 deterministic guns match EXACTLY, reproduced twice. Tsetlin is compared separately because tmLearnOne calls rand(). Getting to 12/12 exposed two real ordering quirks in the live loop: run() calls go() before the aim/fire block, so tickBullets resolves against the NEXT tick's scan while the prediction used the previous one; and if the target dies during that go() the final tick's spawn+resolution is skipped entirely. The recorder emits an end marker for the second case. The 5th (selected-gun) predict call was verified to be a no-op. Measured cost: 8 fixtures (1770 ticks, ~92k virtual bullets, 13 guns) replay in 2.9 s, ~32k virtual bullets/s -- roughly 70x faster and 100x more samples than a live gauntlet. Also adds a per-tick WorldState recorder behind const RecordWorldState (default off, mirrors the ShotLog idiom) which records the state the bot ACTUALLY builds, staleness included, rather than true positions -- recording the latter would hand the guns perfect information and produce flattering scores. 9 new guard checks (33 total, all passing), including fixture round-trip, replay determinism, stationary->HeadOn 100%, constant-velocity->Linear>HeadOn, and the energy-threshold turner crossing at t=41. |