feat(guns): scale-aware power selection (+52% damage); TM classifier gun built, measured, DISABLED

TASK 2 - power selection, a clear win. bestPower used an ABSOLUTE
MinHitRate = 0.40 bar. Measured per-bin virtual rates (rolling-100 fraction)
show no bin ever clears 40%, so 11 of 14 guns were stuck at bin 0 (power 1.0)
even where higher bins were comparable:
  Linear  p1.0 44% p1.5 39% p2.0 30% p3.0 29%   old bin 0 -> new bin 3
  Accel   p1.0 44% p1.5 40% p2.0 26% p3.0 29%   old bin 1 -> new bin 3
  Pattern p1.0 50% p1.5 40% p2.0 27% p3.0 12%   old bin 1 -> new bin 2
Replaced with a scale-aware PowerBarFrac = 0.50 (a dimensionless FRACTION of
the gun's own best bin rate). 13 of 14 selections now pick heavier bullets.
Real effect vs DrussGT (8 rounds x 3 runs): hit rate unchanged (7.56% ->
7.47%) but damage dealt +52% (157 -> 239 per run) and rounds end faster.
Same accuracy, half the shots, half again more damage.

TASK 1 - the TM pattern-classifier gun does NOT earn its slot. It was built as
a mixture of experts with a corrected-Granmo TM as a multi-class gate over
HeadOn/Linear/Circular/WallBounce/Accel, labelled by which expert's prediction
was closest to the actual enemy position (an exact, supervised, per-shot
label - no delayed credit). Offline it loses to the best of its OWN experts on
essentially every fixture, and against DrussGT it cost real performance:
  baseline (path+relative)  7.56% real hit rate, damage 157
  + power fix               7.47%,                 damage 239
  + power fix + TM gun      5.59%,                 damage 133
The gun was selected on 806 ticks and fired 24 real shots at 4.2%.
So the tree ships with EnableTmSelector = false: code and wiring kept intact
for re-enabling, but it is not in the active rack.

Worth recording from the clause dump: the gate DOES latch onto meaningful
structure. On energy-threshold-turner, HeadOn's clauses key on the energy bits
(the rule's own driving variable) while Circular keys on distance/velocity. So
the TM is learning something real and interpretable - it simply cannot beat
'always pick the best expert'. Root cause (INFERRED): the closest-expert label
is noisy because several experts are near-tied, and under the path metric the
winner varies by power bin while the gate sees one shared per-tick input, so a
one-vs-rest gate over a saturated 870-bit clause space has no margin to exploit.
(Zero-padding the 2-frame window was tried first and saturated every clause at
256-755 included literals; alternating the two real frames fixed that.)

Also factors the corrected feedback into an exported tmLearnDir and exports the
encoding/TM primitives; the Tsetlin tests still reproduce the documented
mean=13.8 included literals, so the refactor is behaviour-preserving.

Verified: 33/33 guard checks, tsetlin tests green, metric checks green, new
power-selection guard green (13/14 selections change; relative bar still picks
bin 1 and not bin 3 for a [30,25,12,5]% profile), 12/12 offline==online
acceptance under the shipped default.
This commit is contained in:
2026-09-21 05:19:07 +02:00
parent dea4dcb574
commit 57b2ac3849
8 changed files with 601 additions and 60 deletions
+24 -3
View File
@@ -17,12 +17,16 @@ import guns/displacement
import guns/averaged_lead
import guns/decay_gf
import guns/knn_gun
import guns/tm_selector
proc buildAllGunDrivers*(seed = -1): seq[GunDriver] =
## seed >= 0 re-seeds the global RNG after constructing Tsetlin so the
## stochastic gun's learning is reproducible for offline runs. (Its
## constructor calls randomize(); we override that seed afterwards.)
## seed >= 0 re-seeds the global RNG after constructing the stochastic guns
## (Tsetlin and the TM selector both call randomize() in their constructors),
## so their learning is reproducible for offline runs.
##
## Order matches ModularBot's gun ids exactly (TMSelect appended at 13).
var tsetlin = initTsetlinGun()
var tmSelector = initTmSelectorGun()
if seed >= 0:
randomize(seed)
result = @[
@@ -39,6 +43,7 @@ proc buildAllGunDrivers*(seed = -1): seq[GunDriver] =
makeDriver("AvgLead", initAveragedLeadGun()),
makeDriver("DecayGF", initDecayGFGun()),
makeDriver("KNN", initKNNGun()),
makeDriver("TMSelect", tmSelector),
]
proc makeTsetlinDriver*(seed = -1): tuple[driver: GunDriver, gun: ref TsetlinGun] =
@@ -57,3 +62,19 @@ proc makeTsetlinDriver*(seed = -1): tuple[driver: GunDriver, gun: ref TsetlinGun
resultCb: proc(e: FeedbackEvent) = g[].onResult(e),
readyCb: proc(): bool = g[].isWarmedUp(),
)
proc makeTmSelectorDriver*(seed = -1): tuple[driver: GunDriver, gun: ref TmSelectorGun] =
## Same as makeDriver("TMSelect", ...) but keeps a handle to the concrete gun
## so a test can inspect its votes / clause interpretability after a replay.
let g = new(TmSelectorGun)
g[] = initTmSelectorGun()
if seed >= 0:
randomize(seed)
result.gun = g
result.driver = GunDriver(
name: "TMSelect",
predictCb: proc(state: WorldState, bulletSpeed: float): GunPrediction =
g[].predict(state, bulletSpeed),
resultCb: proc(e: FeedbackEvent) = g[].onResult(e),
readyCb: proc(): bool = g[].isWarmedUp(),
)