c305ef4212b641a44a26eba98b41f8bd1e6ace69
57 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
dea4dcb574 |
feat(gun_harness): scale-aware selector thresholds; default = path + relative
The selection thresholds were calibrated for a rate scale that does not exist.
MEASURED on an exact offline replay of a fogged live WorldState vs DrussGT
(1397 selection ticks), the 0.10 absolute floor fires on 53.0% of point-metric
ticks and forces HeadOn, which has a REAL hit rate of 2.0-4.4% - worst or
near-worst of 13 guns. HeadOn's selection share: 69.1% (abs+point) -> 43.5%
(rel+point). My earlier claim that the floor fires ALWAYS is REFUTED - it is
53%, because bestRate is a max over gun x bin and a >=50-sample bin
occasionally clears 10%. The mechanism is confirmed; the literal statement was
not.
Scale-aware mode (GUN_SELECTOR_MODE, absolute|relative, default relative):
RelTieMargin = 0.20 dimensionless FRACTION of bestRate, replacing the
fixed 2pp band so the band scales with the metric
FloorPeakFrac = 0.25 the floor fires iff bestRate < 0.25 * peakRateRef,
SelectorWindow = 256 where peakRateRef is the field-best rate over the last
256 selection ticks - keeping the original 'don't trust
a collapsed field' purpose but only when the field is
bad RELATIVE TO ITS OWN RECENT BEST, and counting only
guns with >= MinObsBeforeCompete samples so cold-start
100% spikes cannot pin HeadOn
also pools the rate over power bins instead of taking the max over bins, so
one lucky bin no longer wins
absolute mode is preserved byte-for-byte for rollback.
A/B vs DrussGT, real server hit rate, 3 runs x 10 rounds per config, one frozen
binary:
absolute+point 3.66 / 2.45 / 5.01 pooled 3.76%
absolute+path 7.55 / 8.21 / 6.83 pooled 7.57%
relative+point 7.66 / 6.18 / 5.79 pooled 6.59%
relative+path 7.15 / 7.55 / 6.90 pooled 7.21%
absolute+point is SEPARATED from all three (p < 0.0001); the other three
OVERLAP each other (p = 0.18-0.64). So the METRIC is the dominant lever and
under path the two threshold models are statistically tied.
DEFAULT SET: metric = path, thresholds = relative. absolute+path was nominally
0.35pp higher but indistinguishable (p = 0.64); relative is the principled
scale-aware fix, is the only model that works under BOTH metrics, and prevents
the point-metric catastrophe if anyone switches back. Shipping absolute would
ship the accidental side-effect this work exists to remove.
STILL NOT SOLVED: the selector remains only a moderate ranker.
Spearman(virtual rank, real rank) is 0.52 for the winning config, 0.36 pooled
for path and 0.04 for point - and it is INCONSISTENT across run sets. The
metric switch won by de-selecting HeadOn, not by ranking guns better. That is
the next problem.
TASK B, report only: do NOT drive selection from raw real hit rates yet.
Only the selected gun fires, so unselected guns get near-zero real shots
(GuessFactor 20, Linear 24 vs HeadOn 733); noise is fatal (n=470 at p=10% gives
+/-2.8pp, most guns n<200 gives +/-5pp+ across a 3-15% spread); and real rate is
conditional on when the gun was selected. A blended signal with forced
exploration and shrinkage is defensible in principle but needs thousands of
shots per gun across many battles. Real rate is best used OFFLINE as the
evaluation metric - which is exactly what this A/B did.
RELATED BUG FLAGGED, not fixed: MinHitRate = 0.40 in bestPower is on the same
wrong scale - no bin ever clears 40%, so once every bin has data, power
selection falls back to bin 0 (power 1.0) late in a round.
Verified: 33/33 guard checks, 11/11 metric checks, tsetlin green, 12/12
offline==online acceptance under the shipped default, run_range rc=0 over 20
fixtures. Adds analyze_selector.nim to measure floor/tie/bestRate/HeadOn-share
per config on any fixture.
|
||
|
|
3b5d70b7c3 |
feat(gun_harness): runtime metric switch + A/B proving the point metric mis-selects
Adds GUN_VBULLET_METRIC (point|path, default point = unchanged behaviour) so the virtual-bullet hit model can be selected at runtime with no rebuild. Both the live tracker and the offline replay read the same value, so the 12/12 offline==online acceptance holds under EITHER setting (verified for both). A/B AGAINST THE LIVE BOSS, real server-side hit rate as ground truth, 5 battles x 12 rounds per metric on one frozen binary: point 4660 shots / 219 hits = 4.70% (per-run 3.16-5.53) path 4834 shots / 359 hits = 7.43% (per-run 6.55-8.24) The distributions DO NOT OVERLAP: path's worst run beats point's best run. +2.73pp, +58% relative, z = 5.56, p < 0.0001. Range distributions were identical (~460-478 px), so this is not a range confound. MECHANISM - and this is the important part. The gain is SELECTION, not better gun learning. Under the point model every gun's virtual rate is compressed into 0.6-4.4%, so HeadOn sits inside the 2pp tie margin and takes 72.6% of selection ticks / 76.9% of shots - while HeadOn is 11th of 13 by REAL hit rate (2.3%). The path model widens the band to 4.7-13.7% and ranks HeadOn 10th, so its shot share falls to 35.9% and Pattern/Accel/WallBounce get picked instead. Counterfactual: applying the point model's per-gun real rates to the path model's shot mix yields 7.65%, i.e. essentially the whole observed gain. So the selector, not the guns, is where the win lives. PER-GUN REAL HIT RATE vs DrussGT (path mix, the answer to 'which guns are worth keeping'): WallBounce 10.8, Pattern 10.5, Accel 10.0, Displace 9.3, Circular 9.2, AvgLead 8.5, KNN 5.7, StopShot 5.2, GuessFactor 3.7, Tsetlin 2.9. Per-gun N is small (hundreds of shots) so single-gun ordering is indicative, not definitive. TWO CAVEATS, recorded because they undercut a naive reading: 1. One adversary. DrussGT is a wave surfer and HeadOn is genuinely bad against surfers, so part of this may be matchup-specific. 2. The path model is NOT a better general ranker. Spearman(virtual rank, real rank) is 0.52 under point vs -0.04 under path. It wins by accidentally fixing HeadOn's mis-rank, not by ranking guns better. A more durable fix is to address the selection logic directly - which is the next job. Also adds a focused guard test (test_vbullet_metric) covering parsing/default, a receding-target point-miss/path-hit, a perpendicular-target path-miss, and replay determinism. Verified: 33 guard checks, 12/12 acceptance under both metrics, tsetlin tests green, range 34.3% (point, unchanged) / 50.8% (path). |
||
|
|
d5061ee215 |
test(range): restore the 12/12 offline==online proof; measure TM clause readability
Task 1 - the acceptance proof was unrunnable because RecordWorldState was a
compile-time const set to false. It is now a RUNTIME switch
(let RecordWorldState* = existsEnv("TR_RECORD_WORLDSTATE")), default OFF, so
ordinary runs write no fixture, and acceptance_offline_vs_online.nim enables
it for the battle it spawns and clears it afterwards. Restored and run twice:
12/12 deterministic guns match exactly (128-tick and 546-tick battles), with
Tsetlin reported separately as stochastic. Both nimble build variants clean.
Task 2 - does a compact encoding turn the TM's 99.35% into a READABLE rule?
Measured across window sizes (fixed seed, no tuning):
frames TEST acc eff.lits/clause firing clauses counterfactual low/high/mean
10 99.35% 152.8 37 100/24/62.4%
3 95.94% 54.9 35 96/20/58.6%
2 99.48% 39.6 38 95/25/60.7%
1 98.30% 19.2 45 100/24/62.3%
So 2 frames is strictly better than 10 on BOTH axes: +0.13 accuracy for 4x
smaller clauses. The 3-frame dip is non-monotonic and left unexplained rather
than smoothed over.
A readable rule WAS partially recovered. Five clauses carry the exact Gray
form !g10 ^ !g9 ^ !g8; g10 is inert in this data, so the effective rule is the
2-literal proposition !g9 ^ !g8, i.e. energy < 25.6. That is a genuine
threshold in readable propositional form - but at 25.6, NOT the labelled 30,
because 256 is a power-of-two Gray boundary expressible in two literals while
300 needs a longer conjunction. The TM found the nearest SIMPLE threshold.
The honest caveat: that threshold is not the ensemble's decision mechanism.
The counterfactual follow rate (high 24%, mean 62.3%) is statistically
identical at 1, 2 and 10 frames, so compactness did not make the model read
energy - its vote is carried by co-occurring bearing/velocity/heading/wall
literals. Also identified: clauses containing all 11 Gray energy bits are
satisfied at exactly one raw value (50, the dataset floor), so they are
'energy has hit the floor' detectors, not thresholds.
Methodological fix worth keeping: the earlier single-frame counterfactual
wrote energy into all 10 frame slots including the zeroed ones, reviving dead
clauses and producing a spurious 2% high-follow rate. setEnergyFrames now
rewrites only the exposed frames; the corrected figure is 24%.
|
||
|
|
89370008da |
fix(tsetlin): make the TM actually learn - saturation 714 -> 13.8 literals/clause
The gun has never contributed anything: Tsetlin.vHits was byte-for-byte equal to Linear.vHits in every measured round of every run, because its learned correction was always exactly 0. Six diagnosed defects fixed, plus one that was required to make the first one work: 1. Type I now conditions on the clause output. It previously rewarded included true literals unconditionally, omitting Granmo's (c=0, lk=1) -> toward Exclude counter-force, so true literals ratcheted toward Include forever. This was the root cause of the saturation. 2. Type II was unreachable dead code: its guard required cOut==1 AND lits[lit]==0 AND st>0 (included), but cOut==1 guarantees every included literal is 1. Its direction was wrong too - it should increment EXCLUDED false literals when the clause fires. 3. Resource allocation restored: Granmo's (T - clip(v,-T,T))/(2T) target replaces |error|/(2*RESID_MAX); TM_T was only an output normaliser. 4. Label baseline fixed - the factor-2 shrink. predX = linearX + cx, so the label was delta - cx while the learner's output IS cx, giving error = delta - 2cx and a fixed point of cx = delta/2: HALF the needed correction even with perfect feedback. TmTrace now stores linearX/linearY and training uses delta. 5. Hits no longer zero their label (a hit means |miss| < 18px, not 0). 6. The enemy-energy feature was duplicated - tmEncodeFrame passed state.selfEnergy with a stale comment claiming enemyEnergy was absent, while WorldState.enemyEnergy exists. Enemy-energy rules were literally unrepresentable. 7. REQUIRED EXTRA: tmEvalClause now implements Granmo Eq. 6 - an all-Exclude clause outputs 1 during learning and 0 during classification. Without it, fix #1 deadlocks every clause at empty. MEASURED EFFECT (energy-threshold-turner fixture, seed 1): mean included literals per active clause 714.0 -> 13.8 active clauses 100/100 -> 53/100 nonzero corrections 8/764 -> 708/764 Tsetlin virtual hits (Linear = 27/400) 27/400 -> 69/400 Divergence achieved: offline on 7/8 fixtures, and in a live gauntlet (RandomMover: Tsetlin 199/1200 vs Linear 288/1200, vDropped=vStarved=0). Tsetlin now LEARNS but is not yet competitive with Linear - the regression head is untuned, flagged as follow-up rather than claimed as a win. Also ignores compiled test harnesses that have no file extension, which the existing '**/tests/test_*' rule misses. |
||
|
|
974528d5cf |
feat(gun_harness): offline gun range, proven equivalent to live play
Gun evaluation previously required a full end-to-end battle (Java server + battle runner + websocket IPC to 2 bot processes, 50 rounds, ~3.4 min) and yielded only ~300-900 REAL shots across 13 guns -- far too few to rank guns, which is why tuning needed many repetitions. VirtualTracker is already a pure function of (WorldState stream, gun list); the only reason it needed Java was where WorldState came from. So the range replays a seq[WorldState] through the SAME tracker: offline and online scores are the same metric by construction, not an approximation. ACCEPTANCE TEST (the point of the whole thing): record one live round, replay it offline, compare per-gun virtual hit rates. 12/12 deterministic guns match EXACTLY, reproduced twice. Tsetlin is compared separately because tmLearnOne calls rand(). Getting to 12/12 exposed two real ordering quirks in the live loop: run() calls go() before the aim/fire block, so tickBullets resolves against the NEXT tick's scan while the prediction used the previous one; and if the target dies during that go() the final tick's spawn+resolution is skipped entirely. The recorder emits an end marker for the second case. The 5th (selected-gun) predict call was verified to be a no-op. Measured cost: 8 fixtures (1770 ticks, ~92k virtual bullets, 13 guns) replay in 2.9 s, ~32k virtual bullets/s -- roughly 70x faster and 100x more samples than a live gauntlet. Also adds a per-tick WorldState recorder behind const RecordWorldState (default off, mirrors the ShotLog idiom) which records the state the bot ACTUALLY builds, staleness included, rather than true positions -- recording the latter would hand the guns perfect information and produce flattering scores. 9 new guard checks (33 total, all passing), including fixture round-trip, replay determinism, stationary->HeadOn 100%, constant-velocity->Linear>HeadOn, and the energy-threshold turner crossing at t=41. |
||
|
|
3c90a5941d |
feat(selector): range-aware firing gate fitted to 2611 measured shots
Measured, not assumed. With the gate temporarily opened to 20 deg, every real shot was logged (tick, angle error at fire time, distance, power, hit) across 3 gauntlets: 2611 shots, 57.3% aggregate. Findings: - The geometric cone atan(BotRadius/d) is directionally confirmed but a WEAK lever: even at 0.0-0.1 deg error the hit rate at 400-600px is only ~53-57%, because PREDICTION error dominates alignment error. - Real effect of tightening the gate: 57.9% -> 68.0% aggregate hit rate (fixed 0.1 deg), not the 76.9% previously reported -- that was a high-variance draw (per-rep 62.8/66.4/77.2%). - The shipped range-aware gate (SafetyFactor 0.6) does NOT beat the fixed 2.0 deg gate on hit rate (55.8% vs 57.9%, ~1.5 sigma, inside noise). It fires 22-28% more shots and therefore lands more total hits (~509 vs ~434 per rep). No per-adversary score delta exceeded the 300-point run-to-run noise band, so no config is demonstrably better on score. Shipped anyway because it is strictly more expressive (a fixed threshold is the special case), tunable from one const, and physically motivated, but the honest verdict is recorded in-code: the gate is not the bottleneck. AimThresholdDeg is removed; shouldFire now takes distPx. Degenerate or NaN distance falls back to the ceiling rather than dividing by zero. Also adds a per-shot logger to ModularBot behind 'const ShotLog' so the measurement above is reproducible, and 10 new guard checks (24 total, all passing) covering monotonicity, clamping, formula, perfect alignment, gross misalignment and degenerate distance. Cross-checked against the server source: the gun fires BEFORE the turn is applied, so the logged angle error is the true departure error, and fireAssist auto-aim is off (unset by the Nim API and forced false by setAdjustRadarForGunTurn). |
||
|
|
e53690036b |
fix(guns): speed-sensitive caches, dead stop-shot branch, exact TM trace pairing
Four guns cached a whole prediction per tick while predict() is called once per power bin, so every bin after the first (and the real fired shot, which shares lastState) reused the power-1.0 lead. Fixed by caching only the speed-INDEPENDENT derived state and recomputing the lead per requested speed: - stop_shot: also fixes prevSpeed being written before it was read, which made abs(speed) < abs(prev) permanently false and the entire stop-prediction branch unreachable (it was just Linear). - displacement: the cache key included bulletSpeed, so the guard missed on all four bins and the 15-tick window advanced ~4x/tick, making the inferred velocity ~4x too small. - averaged_lead: tick cache removed outright. pattern_matcher: split into speed-independent match+path and per-call lead. FeedbackEvent gains fireTick/powerBin (additive; only virtual_bullets constructs one) so guns can pair feedback to the exact shot instead of guessing by coordinates. tsetlin uses it: traces are now keyed exactly by (fireTick, powerBin) with a 1024-slot ring, and the 10-frame window shifts at most once per tick (it was shifting ~4-5x/tick, so isWarmedUp tripped after ~2 ticks). KNOWN INCOMPLETE: tsetlin still does not diverge from Linear in battle. The two named bugs are fixed (a 600-tick sim shows trainedShots=2141, traceMisses=0, and a fixed-input probe converges to a 9.6px correction), but the TM's clause feedback itself is broken: ~131 of 1740 literals end up included per clause, so its conjunction never fires. Sweeping TM_S, TM_N_CLAUSES and a two-branch Type-I update did not change the correction from 0. Needs a real TM fix or removal, not another bug fix. First-ever guard tests for the gun selector: common_libs/tests/ test_gun_harness.nim (14 checks, headless, no Java). There were none before, which is how six broken guns survived a full analysis cycle. Against the previous HEAD, 5 of these checks FAIL - that is the regression guard. |