feat(guns): scale-aware power selection (+52% damage); TM classifier gun built, measured, DISABLED
TASK 2 - power selection, a clear win. bestPower used an ABSOLUTE MinHitRate = 0.40 bar. Measured per-bin virtual rates (rolling-100 fraction) show no bin ever clears 40%, so 11 of 14 guns were stuck at bin 0 (power 1.0) even where higher bins were comparable: Linear p1.0 44% p1.5 39% p2.0 30% p3.0 29% old bin 0 -> new bin 3 Accel p1.0 44% p1.5 40% p2.0 26% p3.0 29% old bin 1 -> new bin 3 Pattern p1.0 50% p1.5 40% p2.0 27% p3.0 12% old bin 1 -> new bin 2 Replaced with a scale-aware PowerBarFrac = 0.50 (a dimensionless FRACTION of the gun's own best bin rate). 13 of 14 selections now pick heavier bullets. Real effect vs DrussGT (8 rounds x 3 runs): hit rate unchanged (7.56% -> 7.47%) but damage dealt +52% (157 -> 239 per run) and rounds end faster. Same accuracy, half the shots, half again more damage. TASK 1 - the TM pattern-classifier gun does NOT earn its slot. It was built as a mixture of experts with a corrected-Granmo TM as a multi-class gate over HeadOn/Linear/Circular/WallBounce/Accel, labelled by which expert's prediction was closest to the actual enemy position (an exact, supervised, per-shot label - no delayed credit). Offline it loses to the best of its OWN experts on essentially every fixture, and against DrussGT it cost real performance: baseline (path+relative) 7.56% real hit rate, damage 157 + power fix 7.47%, damage 239 + power fix + TM gun 5.59%, damage 133 The gun was selected on 806 ticks and fired 24 real shots at 4.2%. So the tree ships with EnableTmSelector = false: code and wiring kept intact for re-enabling, but it is not in the active rack. Worth recording from the clause dump: the gate DOES latch onto meaningful structure. On energy-threshold-turner, HeadOn's clauses key on the energy bits (the rule's own driving variable) while Circular keys on distance/velocity. So the TM is learning something real and interpretable - it simply cannot beat 'always pick the best expert'. Root cause (INFERRED): the closest-expert label is noisy because several experts are near-tied, and under the path metric the winner varies by power bin while the gate sees one shared per-tick input, so a one-vs-rest gate over a saturated 870-bit clause space has no margin to exploit. (Zero-padding the 2-frame window was tried first and saturated every clause at 256-755 included literals; alternating the two real frames fixed that.) Also factors the corrected feedback into an exported tmLearnDir and exports the encoding/TM primitives; the Tsetlin tests still reproduce the documented mean=13.8 included literals, so the refactor is behaviour-preserving. Verified: 33/33 guard checks, tsetlin tests green, metric checks green, new power-selection guard green (13/14 selections change; relative bar still picks bin 1 and not bin 3 for a [30,25,12,5]% profile), 12/12 offline==online acceptance under the shipped default.
This commit is contained in:
@@ -17,12 +17,16 @@ import guns/displacement
|
||||
import guns/averaged_lead
|
||||
import guns/decay_gf
|
||||
import guns/knn_gun
|
||||
import guns/tm_selector
|
||||
|
||||
proc buildAllGunDrivers*(seed = -1): seq[GunDriver] =
|
||||
## seed >= 0 re-seeds the global RNG after constructing Tsetlin so the
|
||||
## stochastic gun's learning is reproducible for offline runs. (Its
|
||||
## constructor calls randomize(); we override that seed afterwards.)
|
||||
## seed >= 0 re-seeds the global RNG after constructing the stochastic guns
|
||||
## (Tsetlin and the TM selector both call randomize() in their constructors),
|
||||
## so their learning is reproducible for offline runs.
|
||||
##
|
||||
## Order matches ModularBot's gun ids exactly (TMSelect appended at 13).
|
||||
var tsetlin = initTsetlinGun()
|
||||
var tmSelector = initTmSelectorGun()
|
||||
if seed >= 0:
|
||||
randomize(seed)
|
||||
result = @[
|
||||
@@ -39,6 +43,7 @@ proc buildAllGunDrivers*(seed = -1): seq[GunDriver] =
|
||||
makeDriver("AvgLead", initAveragedLeadGun()),
|
||||
makeDriver("DecayGF", initDecayGFGun()),
|
||||
makeDriver("KNN", initKNNGun()),
|
||||
makeDriver("TMSelect", tmSelector),
|
||||
]
|
||||
|
||||
proc makeTsetlinDriver*(seed = -1): tuple[driver: GunDriver, gun: ref TsetlinGun] =
|
||||
@@ -57,3 +62,19 @@ proc makeTsetlinDriver*(seed = -1): tuple[driver: GunDriver, gun: ref TsetlinGun
|
||||
resultCb: proc(e: FeedbackEvent) = g[].onResult(e),
|
||||
readyCb: proc(): bool = g[].isWarmedUp(),
|
||||
)
|
||||
|
||||
proc makeTmSelectorDriver*(seed = -1): tuple[driver: GunDriver, gun: ref TmSelectorGun] =
|
||||
## Same as makeDriver("TMSelect", ...) but keeps a handle to the concrete gun
|
||||
## so a test can inspect its votes / clause interpretability after a replay.
|
||||
let g = new(TmSelectorGun)
|
||||
g[] = initTmSelectorGun()
|
||||
if seed >= 0:
|
||||
randomize(seed)
|
||||
result.gun = g
|
||||
result.driver = GunDriver(
|
||||
name: "TMSelect",
|
||||
predictCb: proc(state: WorldState, bulletSpeed: float): GunPrediction =
|
||||
g[].predict(state, bulletSpeed),
|
||||
resultCb: proc(e: FeedbackEvent) = g[].onResult(e),
|
||||
readyCb: proc(): bool = g[].isWarmedUp(),
|
||||
)
|
||||
|
||||
Reference in New Issue
Block a user