7f6ccfb0151ea2df76a469f681dababfd1ff5d16
2 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
dea4dcb574 |
feat(gun_harness): scale-aware selector thresholds; default = path + relative
The selection thresholds were calibrated for a rate scale that does not exist.
MEASURED on an exact offline replay of a fogged live WorldState vs DrussGT
(1397 selection ticks), the 0.10 absolute floor fires on 53.0% of point-metric
ticks and forces HeadOn, which has a REAL hit rate of 2.0-4.4% - worst or
near-worst of 13 guns. HeadOn's selection share: 69.1% (abs+point) -> 43.5%
(rel+point). My earlier claim that the floor fires ALWAYS is REFUTED - it is
53%, because bestRate is a max over gun x bin and a >=50-sample bin
occasionally clears 10%. The mechanism is confirmed; the literal statement was
not.
Scale-aware mode (GUN_SELECTOR_MODE, absolute|relative, default relative):
RelTieMargin = 0.20 dimensionless FRACTION of bestRate, replacing the
fixed 2pp band so the band scales with the metric
FloorPeakFrac = 0.25 the floor fires iff bestRate < 0.25 * peakRateRef,
SelectorWindow = 256 where peakRateRef is the field-best rate over the last
256 selection ticks - keeping the original 'don't trust
a collapsed field' purpose but only when the field is
bad RELATIVE TO ITS OWN RECENT BEST, and counting only
guns with >= MinObsBeforeCompete samples so cold-start
100% spikes cannot pin HeadOn
also pools the rate over power bins instead of taking the max over bins, so
one lucky bin no longer wins
absolute mode is preserved byte-for-byte for rollback.
A/B vs DrussGT, real server hit rate, 3 runs x 10 rounds per config, one frozen
binary:
absolute+point 3.66 / 2.45 / 5.01 pooled 3.76%
absolute+path 7.55 / 8.21 / 6.83 pooled 7.57%
relative+point 7.66 / 6.18 / 5.79 pooled 6.59%
relative+path 7.15 / 7.55 / 6.90 pooled 7.21%
absolute+point is SEPARATED from all three (p < 0.0001); the other three
OVERLAP each other (p = 0.18-0.64). So the METRIC is the dominant lever and
under path the two threshold models are statistically tied.
DEFAULT SET: metric = path, thresholds = relative. absolute+path was nominally
0.35pp higher but indistinguishable (p = 0.64); relative is the principled
scale-aware fix, is the only model that works under BOTH metrics, and prevents
the point-metric catastrophe if anyone switches back. Shipping absolute would
ship the accidental side-effect this work exists to remove.
STILL NOT SOLVED: the selector remains only a moderate ranker.
Spearman(virtual rank, real rank) is 0.52 for the winning config, 0.36 pooled
for path and 0.04 for point - and it is INCONSISTENT across run sets. The
metric switch won by de-selecting HeadOn, not by ranking guns better. That is
the next problem.
TASK B, report only: do NOT drive selection from raw real hit rates yet.
Only the selected gun fires, so unselected guns get near-zero real shots
(GuessFactor 20, Linear 24 vs HeadOn 733); noise is fatal (n=470 at p=10% gives
+/-2.8pp, most guns n<200 gives +/-5pp+ across a 3-15% spread); and real rate is
conditional on when the gun was selected. A blended signal with forced
exploration and shrinkage is defensible in principle but needs thousands of
shots per gun across many battles. Real rate is best used OFFLINE as the
evaluation metric - which is exactly what this A/B did.
RELATED BUG FLAGGED, not fixed: MinHitRate = 0.40 in bestPower is on the same
wrong scale - no bin ever clears 40%, so once every bin has data, power
selection falls back to bin 0 (power 1.0) late in a round.
Verified: 33/33 guard checks, 11/11 metric checks, tsetlin green, 12/12
offline==online acceptance under the shipped default, run_range rc=0 over 20
fixtures. Adds analyze_selector.nim to measure floor/tie/bestRate/HeadOn-share
per config on any fixture.
|
||
|
|
89370008da |
fix(tsetlin): make the TM actually learn - saturation 714 -> 13.8 literals/clause
The gun has never contributed anything: Tsetlin.vHits was byte-for-byte equal to Linear.vHits in every measured round of every run, because its learned correction was always exactly 0. Six diagnosed defects fixed, plus one that was required to make the first one work: 1. Type I now conditions on the clause output. It previously rewarded included true literals unconditionally, omitting Granmo's (c=0, lk=1) -> toward Exclude counter-force, so true literals ratcheted toward Include forever. This was the root cause of the saturation. 2. Type II was unreachable dead code: its guard required cOut==1 AND lits[lit]==0 AND st>0 (included), but cOut==1 guarantees every included literal is 1. Its direction was wrong too - it should increment EXCLUDED false literals when the clause fires. 3. Resource allocation restored: Granmo's (T - clip(v,-T,T))/(2T) target replaces |error|/(2*RESID_MAX); TM_T was only an output normaliser. 4. Label baseline fixed - the factor-2 shrink. predX = linearX + cx, so the label was delta - cx while the learner's output IS cx, giving error = delta - 2cx and a fixed point of cx = delta/2: HALF the needed correction even with perfect feedback. TmTrace now stores linearX/linearY and training uses delta. 5. Hits no longer zero their label (a hit means |miss| < 18px, not 0). 6. The enemy-energy feature was duplicated - tmEncodeFrame passed state.selfEnergy with a stale comment claiming enemyEnergy was absent, while WorldState.enemyEnergy exists. Enemy-energy rules were literally unrepresentable. 7. REQUIRED EXTRA: tmEvalClause now implements Granmo Eq. 6 - an all-Exclude clause outputs 1 during learning and 0 during classification. Without it, fix #1 deadlocks every clause at empty. MEASURED EFFECT (energy-threshold-turner fixture, seed 1): mean included literals per active clause 714.0 -> 13.8 active clauses 100/100 -> 53/100 nonzero corrections 8/764 -> 708/764 Tsetlin virtual hits (Linear = 27/400) 27/400 -> 69/400 Divergence achieved: offline on 7/8 fixtures, and in a live gauntlet (RandomMover: Tsetlin 199/1200 vs Linear 288/1200, vDropped=vStarved=0). Tsetlin now LEARNS but is not yet competitive with Linear - the regression head is untuned, flagged as follow-up rather than claimed as a win. Also ignores compiled test harnesses that have no file extension, which the existing '**/tests/test_*' rule misses. |