feat(guns): scale-aware power selection (+52% damage); TM classifier gun built, measured, DISABLED
TASK 2 - power selection, a clear win. bestPower used an ABSOLUTE MinHitRate = 0.40 bar. Measured per-bin virtual rates (rolling-100 fraction) show no bin ever clears 40%, so 11 of 14 guns were stuck at bin 0 (power 1.0) even where higher bins were comparable: Linear p1.0 44% p1.5 39% p2.0 30% p3.0 29% old bin 0 -> new bin 3 Accel p1.0 44% p1.5 40% p2.0 26% p3.0 29% old bin 1 -> new bin 3 Pattern p1.0 50% p1.5 40% p2.0 27% p3.0 12% old bin 1 -> new bin 2 Replaced with a scale-aware PowerBarFrac = 0.50 (a dimensionless FRACTION of the gun's own best bin rate). 13 of 14 selections now pick heavier bullets. Real effect vs DrussGT (8 rounds x 3 runs): hit rate unchanged (7.56% -> 7.47%) but damage dealt +52% (157 -> 239 per run) and rounds end faster. Same accuracy, half the shots, half again more damage. TASK 1 - the TM pattern-classifier gun does NOT earn its slot. It was built as a mixture of experts with a corrected-Granmo TM as a multi-class gate over HeadOn/Linear/Circular/WallBounce/Accel, labelled by which expert's prediction was closest to the actual enemy position (an exact, supervised, per-shot label - no delayed credit). Offline it loses to the best of its OWN experts on essentially every fixture, and against DrussGT it cost real performance: baseline (path+relative) 7.56% real hit rate, damage 157 + power fix 7.47%, damage 239 + power fix + TM gun 5.59%, damage 133 The gun was selected on 806 ticks and fired 24 real shots at 4.2%. So the tree ships with EnableTmSelector = false: code and wiring kept intact for re-enabling, but it is not in the active rack. Worth recording from the clause dump: the gate DOES latch onto meaningful structure. On energy-threshold-turner, HeadOn's clauses key on the energy bits (the rule's own driving variable) while Circular keys on distance/velocity. So the TM is learning something real and interpretable - it simply cannot beat 'always pick the best expert'. Root cause (INFERRED): the closest-expert label is noisy because several experts are near-tied, and under the path metric the winner varies by power bin while the gate sees one shared per-tick input, so a one-vs-rest gate over a saturated 870-bit clause space has no margin to exploit. (Zero-padding the 2-frame window was tried first and saturated every clause at 256-755 included literals; alternating the two real frames fixed that.) Also factors the corrected feedback into an exported tmLearnDir and exports the encoding/TM primitives; the Tsetlin tests still reproduce the documented mean=13.8 included literals, so the refactor is behaviour-preserving. Verified: 33/33 guard checks, tsetlin tests green, metric checks green, new power-selection guard green (13/14 selections change; relative bar still picks bin 1 and not bin 3 for a [30,25,12,5]% profile), 12/12 offline==online acceptance under the shipped default.
This commit is contained in:
@@ -18,7 +18,13 @@ const
|
||||
## full-map long shot (~90 ticks) need ~4700 slots;
|
||||
## 8192 wraps only after ~157 ticks. Each VirtualBullet
|
||||
## is ~120 bytes, so this array costs ~960 KiB.
|
||||
MinHitRate* = 0.40 ## 40% threshold for acceptable power selection
|
||||
MinHitRate* = 0.40 ## LEGACY absolute bar; no longer used by bestPower
|
||||
## (no bin on the live path-metric scale cleared it, so
|
||||
## once every bin had data bestPower fell to power 1.0).
|
||||
PowerBarFrac* = 0.50 ## RELATIVE power bar (dimensionless): a bin is
|
||||
## acceptable when its virtual hit rate is at least this
|
||||
## FRACTION of the same gun's best bin rate. Scales with
|
||||
## the metric instead of assuming a ~40% hit rate.
|
||||
MinObsBeforeCompete* = 50 ## min observations before a gun×bin enters competition
|
||||
TieMargin* = 0.02 ## ABSOLUTE mode: guns within this hit-rate margin of best are tied
|
||||
MinHitRateFloor* = 0.10 ## ABSOLUTE mode: if best gun < this, fall back to gun 0 (HeadOn)
|
||||
@@ -418,26 +424,32 @@ proc fitnessFor*(t: VirtualTracker, targetId: int): seq[GunFitness] =
|
||||
result[gunId].bins[binIdx].record(src.hits[k])
|
||||
|
||||
proc bestPower*(t: VirtualTracker, gunId: GunId, targetId: int = -1): (int, float) =
|
||||
## Returns (binIdx, power) with highest power that has >= MinHitRate.
|
||||
## Falls back to lowest power bin if nothing qualifies yet.
|
||||
## Uses per-enemy fitness when targetId >= 0 and data exists; else aggregate.
|
||||
## Returns (binIdx, power). Prefers the HIGHEST power bin whose virtual hit
|
||||
## rate is acceptable, where "acceptable" is measured RELATIVE to the same
|
||||
## gun's best bin (`rate >= PowerBarFrac * bestBinRate`, dimensionless) — not
|
||||
## against the legacy absolute `MinHitRate`. On the live path-metric scale a
|
||||
## gun's rates sit around 3-40%, so the absolute 40% bar never fired once every
|
||||
## bin had data and bestPower silently collapsed to power 1.0; the relative bar
|
||||
## discriminates between bins at any scale.
|
||||
##
|
||||
## An EMPTY bin is still handed out (highest power first) so every bin keeps
|
||||
## getting sampled, and a fully cold gun (no data anywhere) returns the lowest
|
||||
## power bin. Uses per-enemy fitness when targetId >= 0 and data exists; else
|
||||
## the deterministic aggregate.
|
||||
let fit = t.fitnessFor(targetId)
|
||||
result = (0, PowerBins[0])
|
||||
# Cold gun (zero observations in every bin): fall back to the lowest power bin,
|
||||
# as documented. Without this the countdown loop below would hit the empty
|
||||
# highest bin first and wrongly return power 3.0.
|
||||
var anyObs = false
|
||||
var bestRate = 0.0
|
||||
for binIdx in 0..<len(PowerBins):
|
||||
if fit[gunId].bins[binIdx].count > 0:
|
||||
anyObs = true
|
||||
break
|
||||
bestRate = max(bestRate, fit[gunId].bins[binIdx].hitRate())
|
||||
if not anyObs:
|
||||
return (0, PowerBins[0])
|
||||
# Warm gun: unchanged — return the highest power bin clearing MinHitRate
|
||||
# (or an empty bin, which the existing logic treats as acceptable).
|
||||
let bar = PowerBarFrac * bestRate
|
||||
for binIdx in countdown(len(PowerBins) - 1, 0):
|
||||
let rate = fit[gunId].bins[binIdx].hitRate()
|
||||
if rate >= MinHitRate or fit[gunId].bins[binIdx].count == 0:
|
||||
let fw = fit[gunId].bins[binIdx]
|
||||
if fw.count == 0 or fw.hitRate() >= bar:
|
||||
return (binIdx, PowerBins[binIdx])
|
||||
|
||||
proc chooseFromFit*(fit: seq[GunFitness], diag: ptr SelectorDiag = nil,
|
||||
|
||||
Reference in New Issue
Block a user