feat(guns): scale-aware power selection (+52% damage); TM classifier gun built, measured, DISABLED
TASK 2 - power selection, a clear win. bestPower used an ABSOLUTE MinHitRate = 0.40 bar. Measured per-bin virtual rates (rolling-100 fraction) show no bin ever clears 40%, so 11 of 14 guns were stuck at bin 0 (power 1.0) even where higher bins were comparable: Linear p1.0 44% p1.5 39% p2.0 30% p3.0 29% old bin 0 -> new bin 3 Accel p1.0 44% p1.5 40% p2.0 26% p3.0 29% old bin 1 -> new bin 3 Pattern p1.0 50% p1.5 40% p2.0 27% p3.0 12% old bin 1 -> new bin 2 Replaced with a scale-aware PowerBarFrac = 0.50 (a dimensionless FRACTION of the gun's own best bin rate). 13 of 14 selections now pick heavier bullets. Real effect vs DrussGT (8 rounds x 3 runs): hit rate unchanged (7.56% -> 7.47%) but damage dealt +52% (157 -> 239 per run) and rounds end faster. Same accuracy, half the shots, half again more damage. TASK 1 - the TM pattern-classifier gun does NOT earn its slot. It was built as a mixture of experts with a corrected-Granmo TM as a multi-class gate over HeadOn/Linear/Circular/WallBounce/Accel, labelled by which expert's prediction was closest to the actual enemy position (an exact, supervised, per-shot label - no delayed credit). Offline it loses to the best of its OWN experts on essentially every fixture, and against DrussGT it cost real performance: baseline (path+relative) 7.56% real hit rate, damage 157 + power fix 7.47%, damage 239 + power fix + TM gun 5.59%, damage 133 The gun was selected on 806 ticks and fired 24 real shots at 4.2%. So the tree ships with EnableTmSelector = false: code and wiring kept intact for re-enabling, but it is not in the active rack. Worth recording from the clause dump: the gate DOES latch onto meaningful structure. On energy-threshold-turner, HeadOn's clauses key on the energy bits (the rule's own driving variable) while Circular keys on distance/velocity. So the TM is learning something real and interpretable - it simply cannot beat 'always pick the best expert'. Root cause (INFERRED): the closest-expert label is noisy because several experts are near-tied, and under the path metric the winner varies by power bin while the gate sees one shared per-tick input, so a one-vs-rest gate over a saturated 870-bit clause space has no margin to exploit. (Zero-padding the 2-frame window was tried first and saturated every clause at 256-755 included literals; alternating the two real frames fixed that.) Also factors the corrected feedback into an exported tmLearnDir and exports the encoding/TM primitives; the Tsetlin tests still reproduce the documented mean=13.8 included literals, so the refactor is behaviour-preserving. Verified: 33/33 guard checks, tsetlin tests green, metric checks green, new power-selection guard green (13/14 selections change; relative bar still picks bin 1 and not bin 3 for a [30,25,12,5]% profile), 12/12 offline==online acceptance under the shipped default.
This commit is contained in:
@@ -95,28 +95,28 @@ proc tmEncodeFullVector*(window: array[TM_WINDOW_SIZE, TmFrameEncoded],
|
||||
# ── Tsetlin Machine (adapted from BNNBot_garage/src/tsetlin_predictor.nim) ───
|
||||
|
||||
const
|
||||
TM_N_IN = TM_TOTAL_BITS # 870
|
||||
TM_N_OUT = 2 # cx, cy pixel corrections
|
||||
TM_N_LITERALS = TM_N_IN * 2 # 1740
|
||||
TM_N_CLAUSES = 50 # per output; issue #184 default
|
||||
TM_HALF = TM_N_CLAUSES div 2
|
||||
TM_N_STATES = 32 # automaton range [-32..32]
|
||||
TM_T = float(TM_HALF) # vote clamped to [-T, T]
|
||||
TM_S = 1.5 # specificity
|
||||
TM_N_IN* = TM_TOTAL_BITS # 870
|
||||
TM_N_OUT* = 2 # cx, cy pixel corrections
|
||||
TM_N_LITERALS* = TM_N_IN * 2 # 1740
|
||||
TM_N_CLAUSES* = 50 # per output; issue #184 default
|
||||
TM_HALF* = TM_N_CLAUSES div 2
|
||||
TM_N_STATES* = 32 # automaton range [-32..32]
|
||||
TM_T* = float(TM_HALF) # vote clamped to [-T, T]
|
||||
TM_S* = 1.5 # specificity
|
||||
TM_RESID_MAX = 80.0 # pixel correction range
|
||||
# ponytail: TM_N_STATES=32 needs int16 (int8 only fits ≤127, fine here); raise N_CLAUSES if underfitting
|
||||
|
||||
type
|
||||
TmClauseCache = array[TM_N_OUT * TM_N_CLAUSES, uint8]
|
||||
TmClauseCache* = array[TM_N_OUT * TM_N_CLAUSES, uint8]
|
||||
|
||||
TmNet = object
|
||||
states: array[TM_N_OUT * TM_N_CLAUSES * TM_N_LITERALS, int16]
|
||||
TmNet* = object
|
||||
states*: array[TM_N_OUT * TM_N_CLAUSES * TM_N_LITERALS, int16]
|
||||
# ponytail: int16 to safely hold [-32..32]; TM_N_STATES=32 fits int8 too but int16 is safer
|
||||
|
||||
proc tmStateIdx(outIdx, clause, lit: int): int {.inline.} =
|
||||
proc tmStateIdx*(outIdx, clause, lit: int): int {.inline.} =
|
||||
(outIdx * TM_N_CLAUSES + clause) * TM_N_LITERALS + lit
|
||||
|
||||
proc tmPolarity(clause: int): float {.inline.} =
|
||||
proc tmPolarity*(clause: int): float {.inline.} =
|
||||
if clause < TM_HALF: 1.0 else: -1.0
|
||||
|
||||
type
|
||||
@@ -134,12 +134,12 @@ type
|
||||
meanIncluded*: float ## mean include count over ACTIVE clauses
|
||||
meanIncludedAll*: float ## mean include count over ALL clauses (incl. empty)
|
||||
|
||||
proc tmMakeLiterals(input: TmBinaryVector): array[TM_N_LITERALS, uint8] =
|
||||
proc tmMakeLiterals*(input: TmBinaryVector): array[TM_N_LITERALS, uint8] =
|
||||
for i in 0..<TM_N_IN:
|
||||
result[i] = input[i]
|
||||
result[i + TM_N_IN] = 1'u8 - input[i]
|
||||
|
||||
proc tmEvalClause(net: TmNet, outIdx, clause: int,
|
||||
proc tmEvalClause*(net: TmNet, outIdx, clause: int,
|
||||
lits: array[TM_N_LITERALS, uint8],
|
||||
learning = false): uint8 =
|
||||
var hasIncluded = false
|
||||
@@ -157,7 +157,7 @@ proc tmEvalClause(net: TmNet, outIdx, clause: int,
|
||||
# clause at empty forever.
|
||||
return if learning: 1'u8 else: 0'u8
|
||||
|
||||
proc tmForwardWithCache(net: TmNet, input: TmBinaryVector,
|
||||
proc tmForwardWithCache*(net: TmNet, input: TmBinaryVector,
|
||||
cache: var TmClauseCache): (float, float, float, float) =
|
||||
## Returns (correctionX, correctionY, voteX, voteY). `cache` receives the
|
||||
## clause outputs under LEARNING semantics (empty clause = 1) for tmLearnOne;
|
||||
@@ -176,23 +176,19 @@ proc tmForwardWithCache(net: TmNet, input: TmBinaryVector,
|
||||
vy = clamp(vy, -TM_T, TM_T)
|
||||
(vx / TM_T * TM_RESID_MAX, vy / TM_T * TM_RESID_MAX, vx, vy)
|
||||
|
||||
proc tmLearnOne(net: var TmNet, outIdx: int, lits: array[TM_N_LITERALS, uint8],
|
||||
cache: TmClauseCache, vote: float, residual: float) =
|
||||
## One faithful Granmo Table 2/3 update against a continuous residual target.
|
||||
proc tmLearnDir*(net: var TmNet, outIdx: int, lits: array[TM_N_LITERALS, uint8],
|
||||
cache: TmClauseCache, vote: float, d: float) =
|
||||
## One faithful Granmo Table 2/3 clause update with an EXPLICIT desired vote
|
||||
## direction `d` in {-1, +1}. This is the exact corrected core the Tsetlin gun
|
||||
## uses; `tmLearnOne` is the regression wrapper that derives `d` from a
|
||||
## continuous residual, and the multi-class selector passes the class label
|
||||
## directly.
|
||||
##
|
||||
## `cache` holds clause outputs under LEARNING semantics (empty = 1, Eq. 6).
|
||||
## `vote` is the classification-semantics clause sum at prediction time,
|
||||
## already clamped to [-TM_T, TM_T]. `residual` is the correction target
|
||||
## delta = actual - linear baseline (see onResult), NOT actual - prediction.
|
||||
##
|
||||
## Regression adaptation: the "label" direction is d = sign(residual -
|
||||
## predicted), i.e. which way the correction must move. The resource
|
||||
## allocation of Granmo Eq. 8-11 collapses to a single probability
|
||||
## p = (T - d*clip(v,-T,T)) / (2T), applied both to the aligned clauses
|
||||
## (Type I) and the opposed ones (Type II), exactly as Algorithm 1 lines 11-22.
|
||||
let predicted = vote / TM_T * TM_RESID_MAX
|
||||
let error = residual - predicted
|
||||
let d = if error > 0.0: 1.0 elif error < 0.0: -1.0 else: return
|
||||
## already clamped to [-TM_T, TM_T]. The Granmo resource allocation collapses
|
||||
## to a single probability p = (T - d*clip(v,-T,T)) / (2T): it is high when the
|
||||
## vote opposes `d` and falls to 0 once the class is already won.
|
||||
let pFeedback = (TM_T - d * vote) / (2.0 * TM_T)
|
||||
if pFeedback <= 0.0: return
|
||||
|
||||
@@ -228,6 +224,25 @@ proc tmLearnOne(net: var TmNet, outIdx: int, lits: array[TM_N_LITERALS, uint8],
|
||||
if net.states[si] <= 0:
|
||||
net.states[si] = int16(min(int(net.states[si]) + 1, TM_N_STATES))
|
||||
|
||||
proc tmLearnOne*(net: var TmNet, outIdx: int, lits: array[TM_N_LITERALS, uint8],
|
||||
cache: TmClauseCache, vote: float, residual: float) =
|
||||
## One faithful Granmo Table 2/3 update against a continuous residual target.
|
||||
##
|
||||
## `cache` holds clause outputs under LEARNING semantics (empty = 1, Eq. 6).
|
||||
## `vote` is the classification-semantics clause sum at prediction time,
|
||||
## already clamped to [-TM_T, TM_T]. `residual` is the correction target
|
||||
## delta = actual - linear baseline (see onResult), NOT actual - prediction.
|
||||
##
|
||||
## Regression adaptation: the "label" direction is d = sign(residual -
|
||||
## predicted), i.e. which way the correction must move. The resource
|
||||
## allocation of Granmo Eq. 8-11 collapses to a single probability
|
||||
## p = (T - d*clip(v,-T,T)) / (2T), applied both to the aligned clauses
|
||||
## (Type I) and the opposed ones (Type II), exactly as Algorithm 1 lines 11-22.
|
||||
let predicted = vote / TM_T * TM_RESID_MAX
|
||||
let error = residual - predicted
|
||||
let d = if error > 0.0: 1.0 elif error < 0.0: -1.0 else: return
|
||||
net.tmLearnDir(outIdx, lits, cache, vote, d)
|
||||
|
||||
# ── TsetlinGun public type ────────────────────────────────────────────────────
|
||||
|
||||
const
|
||||
|
||||
Reference in New Issue
Block a user