Files
SirRoboGarage/common_libs/tests/test_tm_horizon.nim
T
SirStone 69debbe347 Re-adaptation: the sliding WINDOW is a control-validated +9.3pp on late accuracy
The user's first-hand diagnosis: "once the pattern is learnt enough we hit DrussGT,
but as soon as it adapts we are not fast enough to re-adapt again." The gun kept
EVERY sample for the whole battle (which the user explicitly asked for), so stale
evidence weighed the same as new evidence - accumulation without forgetting.

TASK 1 - THE CURVE, MEASURED FIRST (prequential side accuracy, 15 rounds x 4
horizons, deciles of the eval stream):
  arm                 early%  late%  decay
  frozen (early only)   64.9   61.4   -3.5
  accum (shipped)       76.4   75.3   -1.1
  window N=150          84.7   84.6   -0.0
  resetdrop 5pp         83.2   84.3   +1.0
**The decay IS real but modest and LOCALIZED IN THE LAST ~30% of the round**
(last-third vs middle-third: frozen -7.7pp, accum -3.7, window -1.3, resetdrop -1.3).

TASK 3 - WHICH FIX HELPS (late-half accuracy, shuffled control in parens):
  accum (shipped)                       75.3
  **window N=150                       84.6  (51)**   +9.3pp
  **resetdrop 5pp                      84.3  (51)**   +9.0pp
  rehearse-all (retrain, NO forgetting) 79.5
- **Sliding window: +9.3pp late (within-round), +9.7 (cross-round), +10.3 (shield).**
  The shuffled control stays ~50-51%, so it learns the ENEMY, not noise.
- **Forgetting is the essential ingredient**: the same periodic retrain WITHOUT
  forgetting reaches only ~79.5%, so roughly half the gain is the retraining
  mechanics and half is the forgetting.
- Change detection ties the window on late accuracy and gives the best decay, but
  at a 5pp threshold it fired 300-500 times in the offline stream - noisy.
- INERTIA IS REDUNDANT WITH THE WINDOW: lower inertia helps the keep-everything
  model a lot (accum late 75.3 -> 83.3 at N=8) but leaves the window FLAT
  (84.6-84.7 at every N). Low inertia and forgetting are SUBSTITUTES, and
  forgetting is the robust one - `TR_TMHORIZON_NSTATES=8` is NOT the primary fix.

VERDICT: SHIPPING CANDIDATE = `TR_TMHORIZON_WINDOW=150`, kept at default 0 until a
live A/B confirms.

HONEST CAVEATS THAT SET EXPECTATIONS:
- The fixtures are OPEN-LOOP (DrussGT does not react to our bullets), so a true
  mid-round adaptation is NOT present; the dominant measured effect is the LEVEL
  gap, not the decay magnitude. The user's "it adapts" magnitude is still INFERRED.
- The harness uses a STRAIGHT-LINE base while the live gun uses Pattern's
  prediction, and it ignores the h-tick label delay, so its absolute side accuracy
  (75-85%) is INFLATED: the same gun measured ~52% = chance live against DrussGT.
  So +9.3pp is a real ARM DELTA, not a promise that the gun now clears the ~80%
  accuracy wall that hits need. A live A/B must decide.

Adds `measure_tm_readapt.nim` (prequential harness with the mandatory shuffled
control, within-round + cross-round protocols, inertia sweep) and its captured
results. test_tm_horizon 79 -> 104 (five new groups). test_tm_diag 48,
test_tm_automata_diag 55, test_tm_clause_shape 66, test_rack_membership 48 pass.
ModularBot compiles.

.gitignore: switched from a broad `measure_*` pattern to EXPLICIT binary names.
The broad rule was too blunt - it also excluded the `.txt` results file, which made
`git add` refuse the whole commit twice. Sources stay tracked; binaries do not.
2026-09-23 00:29:37 +02:00

434 lines
20 KiB
Nim

## Pure unit guard for the horizon-based TM gun (common_libs/guns/tm_horizon).
##
## No Java, no server, no battle. Covers what can be tested without a battle:
## * the horizon-from-flight-time maths and its [10,50] clamp;
## * the 4-bit horizon bucket boundaries;
## * feature extraction shapes: 49 draft bits with exactly one-hot blocks, and
## the 53-bit literal vector with pos/neg complementary literals;
## * the applied-shift geometry (0 deg = identity, +90 = CCW);
## * label lookup: a sample resolves h ticks later, is DROPPED when the
## observation is stale, the last h ticks of a round never resolve, and a
## round boundary wipes the pending queue AND the observation ring (so a
## label can never be built from a position across a round boundary);
## * the reset SPLIT: a round boundary keeps the machines (learning
## accumulates), a game start wipes them, and a target change wipes them
## unless TR_TMHORIZON_RESET_ON_TARGET is off;
## * a full round trains the two binary heads.
##
## Run with plain:
## nim c -r common_libs/tests/test_tm_horizon.nim
import std/[math, os]
import gun_harness/gun_interface
import gun_harness/virtual_bullets
import gun_harness/offline_range
import guns/pattern_matcher
import guns/tm_horizon
var failures = 0
proc check(name: string, ok: bool) =
if ok: echo "PASS: ", name
else: echo "FAIL: ", name; inc failures
proc approx(a, b, tol: float): bool {.inline.} = abs(a - b) <= tol
proc sideTeamCopy(g: TmHorizonGun): seq[int16] =
## Snapshot of the learned clause states, to prove the machines survive or are
## wiped by a given reset.
g.sideClauseStates()
proc drive(g: var TmHorizonGun, fx: Fixture) =
for state in fx.states:
for b in 0..<len(PowerBins):
discard g.predict(state, bulletSpeed(PowerBins[b]))
# ── horizon maths ─────────────────────────────────────────────────────────────
proc testHorizonMaths() =
check "horizon: dist 220 / speed 11 -> 20 ticks",
tmhHorizonFor(220.0, 11.0) == 20
check "horizon: dist 100 / speed 17 -> clamped to H_MIN (10)",
tmhHorizonFor(100.0, 17.0) == TMH_H_MIN
check "horizon: dist 800 / speed 11 -> clamped to H_MAX (50)",
tmhHorizonFor(800.0, 11.0) == TMH_H_MAX
check "horizon: zero bullet speed falls back to H_MIN",
tmhHorizonFor(300.0, 0.0) == TMH_H_MIN
check "horizon: rounding is to the nearest tick",
tmhHorizonFor(250.0, 10.0) == 25
proc testHorizonBuckets() =
check "bucket: 10..19 -> 0",
tmhHorizonBucket(10) == 0 and tmhHorizonBucket(19) == 0
check "bucket: 20..29 -> 1",
tmhHorizonBucket(20) == 1 and tmhHorizonBucket(29) == 1
check "bucket: 30..39 -> 2",
tmhHorizonBucket(30) == 2 and tmhHorizonBucket(39) == 2
check "bucket: 40..50 -> 3",
tmhHorizonBucket(40) == 3 and tmhHorizonBucket(50) == 3
proc testQuadrantNames() =
check "quadrant: cold model names COLD", tmhQuadrantName(-1, -1) == "COLD"
check "quadrant: LARGE-LEFT", tmhQuadrantName(1, 1) == "LARGE-LEFT"
check "quadrant: SMALL-RIGHT", tmhQuadrantName(0, 0) == "SMALL-RIGHT"
# ── feature shapes ────────────────────────────────────────────────────────────
proc blockSum(b: array[TMH_N_BASE, uint8], lo, hi: int): int =
for i in lo..hi: result += int(b[i])
proc testFeatureShapes() =
var g = initTmHorizonGun()
# A straight, moving enemy so every history-dependent block is populated.
let fx = synthesizeConstantVelocity(ticks = 40, speed = 4.0)
for state in fx.states:
discard g.predict(state, bulletSpeed(PowerBins[0]))
let state = fx.states[^1]
let b = g.tmhBaseBits(state)
check "bits: the draft vector is exactly 49 bits", b.len == TMH_N_BASE
# Every one-hot block must carry exactly one set bit.
check "bits: dist-to-nearest-wall is one-hot", blockSum(b, 0, 3) == 1
check "bits: which-wall-nearest is one-hot", blockSum(b, 4, 7) == 1
check "bits: dist-from-us is one-hot", blockSum(b, 8, 13) == 1
check "bits: enemy-heading-vs-line-to-us is one-hot", blockSum(b, 14, 16) == 1
check "bits: ticks-since-reversal is one-hot", blockSum(b, 20, 24) == 1
check "bits: turn-consistency-10 is one-hot", blockSum(b, 25, 27) == 1
check "bits: distance-moved-10 is one-hot", blockSum(b, 28, 30) == 1
check "bits: speed-trend-10 is one-hot", blockSum(b, 31, 33) == 1
check "bits: turn-rate-change-5 is one-hot", blockSum(b, 34, 36) == 1
check "bits: time-until-bullet is one-hot", blockSum(b, 37, 41) == 1
check "bits: bullet-lateral-offset is one-hot", blockSum(b, 42, 48) == 1
proc testLiteralLayout() =
var base: array[TMH_N_BASE, uint8]
base[0] = 1'u8
base[8] = 1'u8
for bucket in 0..<TMH_NH:
let lits = tmhLits(base, bucket)
check "lits: length is 2 * 53", lits.len == TMH_NLITS
check "lits: horizon bucket " & $bucket & " sets exactly one of the 4 raw bits",
lits[TMH_N_BASE + bucket] == 1'u8 and
lits[TMH_N_BASE + ((bucket + 1) mod TMH_NH)] == 0'u8
var complement = true
for i in 0..<TMH_N_BITS:
if int(lits[i]) + int(lits[i + TMH_N_BITS]) != 1: complement = false
check "lits: every literal has its complementary negation (bucket " & $bucket & ")",
complement
# ── shift geometry ────────────────────────────────────────────────────────────
proc testShiftGeometry() =
let p0 = tmhApplyShift(100.0, 100.0, 200.0, 100.0, 0.0)
check "shift: 0 deg is the identity", approx(p0.x, 200.0, 1e-9) and approx(p0.y, 100.0, 1e-9)
# East point rotated +90 deg (CCW) -> north (+y in the math convention).
let p90 = tmhApplyShift(100.0, 100.0, 200.0, 100.0, 90.0)
check "shift: +90 deg rotates east to +y (CCW)",
approx(p90.x, 100.0, 1e-9) and approx(p90.y, 200.0, 1e-9)
let pm90 = tmhApplyShift(100.0, 100.0, 200.0, 100.0, -90.0)
check "shift: -90 deg rotates east to -y (CW)",
approx(pm90.x, 100.0, 1e-9) and approx(pm90.y, 0.0, 1e-9)
check "shift: the aim distance is preserved",
approx(hypot(p90.x - 100.0, p90.y - 100.0), 100.0, 1e-9)
# ── label resolution ──────────────────────────────────────────────────────────
proc testRoundTrains() =
var g = initTmHorizonGun()
g.setShift(0.0) # pure-predict arm: still trains
let fx = synthesizeCircular(ticks = 240)
drive(g, fx)
check "train: a full round resolves samples (trained > 0)", g.trained > 0
check "train: the model warmed past the cold gate", g.trained >= TMH_MIN_OBS
check "train: the last h ticks are still pending at round end",
g.pendingCount > 0
proc testRoundReset() =
var g = initTmHorizonGun()
g.setShift(0.0)
drive(g, synthesizeCircular(ticks = 240))
let t1 = g.trained
g.resetLearning()
check "reset: trained is wiped", g.trained == 0
check "reset: the pending queue is wiped (no cross-round labels)", g.pendingCount == 0
check "reset: side accuracy counters are wiped", g.sideTotal == 0
drive(g, synthesizeCircular(ticks = 240))
check "reset: the fresh round trains again", g.trained > 0 and t1 > 0
proc testRoundBoundaryKeepsMachines() =
## THE SPLIT: a plain round boundary (`resetRoundState`) must keep the learned
## machines and drop only the per-round observation/label state.
var g = initTmHorizonGun()
g.setShift(0.0)
drive(g, synthesizeCircular(ticks = 240))
let trainedBefore = g.trained
let sigBefore = sideTeamCopy(g)
check "round-boundary: samples trained before the boundary", trainedBefore > 0
check "round-boundary: unresolved labels exist at the boundary", g.pendingCount > 0
g.resetRoundState()
check "round-boundary: trained SURVIVES the boundary", g.trained == trainedBefore
check "round-boundary: the clause states SURVIVE the boundary",
sideTeamCopy(g) == sigBefore
check "round-boundary: the pending label queue is cleared", g.pendingCount == 0
check "round-boundary: the observation ring is cleared", g.ringValidCount() == 0
proc testLearningAccumulatesAcrossRounds() =
## The whole point of the fix: round 2 keeps round 1's learning and keeps
## training on top of it, so `trained` CLIMBS across the battle.
var g = initTmHorizonGun()
g.setShift(0.0)
drive(g, synthesizeCircular(ticks = 240))
let r1 = g.trained
g.resetRoundState() # exactly what onRoundStarted now does
drive(g, synthesizeCircular(ticks = 240))
check "accumulate: round 2 starts from round 1's total (trained climbs)",
g.trained > r1
proc testGameStartWipesMachines() =
## A new BATTLE (`resetLearning`) must wipe machines + stats + round state.
var g = initTmHorizonGun()
g.setShift(0.0)
drive(g, synthesizeCircular(ticks = 240))
check "game-start: samples trained before the wipe", g.trained > 0
g.resetLearning("game_start")
check "game-start: trained is wiped", g.trained == 0
check "game-start: side accuracy counters are wiped", g.sideTotal == 0
check "game-start: the pending label queue is cleared", g.pendingCount == 0
check "game-start: the observation ring is cleared", g.ringValidCount() == 0
check "game-start: every clause is back to the Exclude boundary",
g.sideClausesAllExclude()
proc testTargetChangeResetsMachines() =
## Reset on TARGET change to a different bot id, gated by the knob.
var g = initTmHorizonGun()
g.setShift(0.0)
g.setResetOnTarget(true)
discard g.targetChanged(7) # first acquisition: never wipes
drive(g, synthesizeCircular(ticks = 240))
check "target: trained before the change", g.trained > 0
discard g.targetChanged(7) # same enemy: no wipe
check "target: the same enemy does not wipe the machines", g.trained > 0
let wiped = g.targetChanged(9) # different enemy: wipe
check "target: a DIFFERENT enemy wipes the machines", wiped and g.trained == 0
# Knob OFF: a different enemy must NOT wipe the machines.
var h = initTmHorizonGun()
h.setShift(0.0)
h.setResetOnTarget(false)
discard h.targetChanged(7)
drive(h, synthesizeCircular(ticks = 240))
check "target: trained before the change (knob off)", h.trained > 0
let wipedOff = h.targetChanged(9)
check "target: knob off leaves the machines intact",
(not wipedOff) and h.trained > 0
proc testLabelCannotCrossBoundary() =
## The most dangerous interaction: a deferred label must never be resolved
## against a position from the previous round. At a round boundary BOTH the
## pending queue and the observation ring are cleared, so the lookup at
## `fireTick + h` cannot find an old position (and no old pending survives).
var g = initTmHorizonGun()
g.setShift(0.0)
let fx = synthesizeCircular(ticks = 240)
drive(g, fx)
let endTick = fx.states[^1].tick
check "cross-boundary: unresolved labels exist at round end", g.pendingCount > 0
check "cross-boundary: the round-end observation is in the ring", g.ringHas(endTick)
g.resetRoundState() # the round boundary
check "cross-boundary: the old observation is gone", not g.ringHas(endTick)
check "cross-boundary: the deferred label queue is gone", g.pendingCount == 0
check "cross-boundary: the ring is empty after the boundary", g.ringValidCount() == 0
proc testTickRegressionKeepsMachines() =
## The tick-regression self-reset (a missed onRoundStarted) must clear only the
## per-round state and must NOT wipe the machines.
var g = initTmHorizonGun()
g.setShift(0.0)
drive(g, synthesizeCircular(ticks = 120))
let t = g.trained
let sig = sideTeamCopy(g)
check "regression: samples trained before the new round", t > 0
# Simulate the server resetting the tick counter to 0 for a new round.
let fx = synthesizeCircular(ticks = 5)
discard g.predict(fx.states[0], bulletSpeed(PowerBins[0]))
check "regression: a tick regression does NOT wipe the machines", g.trained == t
check "regression: the clause states survive the regression",
sideTeamCopy(g) == sig
check "regression: the old observation ring is cleared (one fresh tick only)",
g.ringValidCount() == 1
proc testStaleObservationsDropped() =
## Build a fixture whose `lastSeenTick` is frozen far in the past: every
## resolved label must be dropped, never trained on.
var g = initTmHorizonGun()
g.setShift(0.0)
var states: seq[WorldState]
for t in 0..<120:
var s = WorldState(
enemyX: 400.0 + 3.0 * t.float, enemyY: 300.0,
enemyHeading: 0.0, enemySpeed: 3.0, enemyEnergy: 100.0,
selfX: 100.0, selfY: 300.0, selfEnergy: 100.0,
arenaWidth: 800.0, arenaHeight: 600.0, tick: t,
enemies: @[EnemyInfo(id: 1, x: 400.0 + 3.0 * t.float, y: 300.0,
heading: 0.0, speed: 3.0, energy: 100.0,
lastSeenTick: 0)]) # frozen -> always stale
states.add s
for state in states:
for b in 0..<len(PowerBins):
discard g.predict(state, bulletSpeed(PowerBins[b]))
check "stale: no sample with a stale observation is ever trained", g.trained == 0
check "stale: the dropped samples are counted", g.pendingDropped > 0
proc testColdModelEmitsNoShift() =
## A cold machine must return Pattern's prediction UNCHANGED. Constant-velocity
## motion is predicted perfectly by Pattern, so no sample trains and the model
## stays cold for the whole fixture.
var g = initTmHorizonGun()
g.setShift(3.0)
var pm = PatternMatcherGun()
let fx = synthesizeConstantVelocity(ticks = 60, speed = 4.0)
for state in fx.states:
for b in 0..<len(PowerBins):
let p = g.predict(state, bulletSpeed(PowerBins[b]))
let q = pm.predict(state, bulletSpeed(PowerBins[b]))
if not approx(p.x, q.x, 1e-9) or not approx(p.y, q.y, 1e-9):
check "cold: prediction must equal Pattern byte-for-byte", false
return
check "cold: prediction equals Pattern byte-for-byte while cold", true
proc testWarmShiftMovesAim() =
## Force the model warm with an all-Exclude head (votes 0/0 -> RIGHT) and check
## the correction actually rotates Pattern's base aim by the configured -3 deg.
var g = initTmHorizonGun()
g.setShift(3.0, 1.0)
g.trained = 100 # force warm; the fresh head votes 0/0 -> RIGHT
let fx = synthesizeCircular(ticks = 6)
let state = fx.states[2]
let speed = bulletSpeed(PowerBins[0])
let p = g.predict(state, speed)
# Pattern caches per tick, so this is the exact base point g used.
let q = g.pattern.predict(state, speed)
check "warm: the corrected aim differs from the Pattern base",
not (approx(p.x, q.x, 1e-9) and approx(p.y, q.y, 1e-9))
let b0 = arctan2(q.y - state.selfY, q.x - state.selfX)
let b1 = arctan2(p.y - state.selfY, p.x - state.selfX)
var d = radToDeg(b1 - b0)
while d > 180.0: d -= 360.0
while d < -180.0: d += 360.0
check "warm: the applied rotation is the configured -3.0 deg",
approx(d, -3.0, 1e-6)
proc testReadaptDefaultKnobs() =
## The shipped defaults must be EXACTLY today's behaviour: no buffering, no
## change detection, compile-time state count, rolling curve off.
var g = initTmHorizonGun()
check "readapt: WINDOW defaults to 0 (keep everything)", g.windowN == 0
check "readapt: RESET_DROP defaults to 0 (off)", g.resetDrop == 0.0
check "readapt: the sample buffer is disabled by default", g.bufferCapacity() == 0
check "readapt: NSTATES defaults to the compile-time TMH_NSTATES",
g.nStates == TMH_NSTATES and g.headStates() == TMH_NSTATES
check "readapt: the accuracy curve is off by default", not g.accurveEnabled
drive(g, synthesizeCircular(ticks = 240))
check "readapt: no reset fires when RESET_DROP is off", g.resetDrops == 0
check "readapt: rolling accuracy is still recorded (measurement only)",
g.rollingAcc(100) >= 0.0 and g.rollingAcc(100) <= 1.0
proc testRuntimeStatesKnob() =
## Inertia is sweepable WITHOUT a rebuild via TR_TMHORIZON_NSTATES.
putEnv("TR_TMHORIZON_NSTATES", "16")
var g = initTmHorizonGun()
delEnv("TR_TMHORIZON_NSTATES")
check "readapt: NSTATES env sets the runtime state count", g.nStates == 16
check "readapt: both heads get the runtime state count",
g.headStates() == 16
var d = initTmHorizonGun()
check "readapt: an unset env cannot move the state count",
d.nStates == TMH_NSTATES
proc testWindowBuffersAndRebuilds() =
## Sliding-window mode must allocate a bounded ring, keep buffering resolved
## samples, and rebuild both heads deterministically from that ring.
var g = initTmHorizonGun()
g.setShift(0.0)
g.setWindow(50)
g.setRetrainConfig(50, 1)
check "window: the ring is allocated to the window width", g.bufferCapacity() == 50
drive(g, synthesizeCircular(ticks = 240))
check "window: more than the window width of samples were buffered",
g.bufferedCount() > 50 and g.trained > 0
let before = sideTeamCopy(g)
g.retrainFromBuffer(50)
let a = sideTeamCopy(g)
g.retrainFromBuffer(50)
let b = sideTeamCopy(g)
check "window: a rebuild is deterministic from the buffer", a == b
check "window: a rebuild actually resets + retrains the machine", a != before
check "window: the resolved-sample count keeps climbing across rebuilds",
g.trained > 0
proc testWindowRetrainKeepsDeferredState() =
## A window rebuild / re-learn must NOT touch the deferred-label queue or the
## observation ring; only `resetRoundState`/`resetLearning` may clear those.
var g = initTmHorizonGun()
g.setShift(0.0)
g.setWindow(50)
g.setRetrainConfig(50, 1)
drive(g, synthesizeCircular(ticks = 240))
let pend = g.pendingCount
let ringN = g.ringValidCount()
check "window: deferred labels exist before the rebuild", pend > 0
g.retrainFromBuffer(50)
check "window: retrain keeps the deferred label queue", g.pendingCount == pend
check "window: retrain keeps the observation ring", g.ringValidCount() == ringN
g.resetLearning()
check "window: a battle reset drops the buffered samples", g.bufferedCount() == 0
check "window: a battle reset drops the rolling accuracy ring", g.accCount == 0
proc testResetDropTriggersRelearn() =
## Change detection: with the knob on, a rolling-accuracy drop below its own
## peak fires a re-learn; with it off, nothing fires. The circular fixture
## mixes a cold start with warm tracking, so the peak/current gap is real.
var off = initTmHorizonGun()
off.setShift(0.0)
drive(off, synthesizeCircular(ticks = 240))
check "resetdrop: off by default fires nothing", off.resetDrops == 0
var g = initTmHorizonGun()
g.setShift(0.0)
g.setResetDrop(0.1)
g.setRetrainConfig(50, 1)
check "resetdrop: enabling it allocates the re-learn ring",
g.bufferCapacity() >= TMH_RESET_WINDOW_DEF
drive(g, synthesizeCircular(ticks = 240))
check "resetdrop: a rolling-accuracy drop triggers at least one re-learn",
g.resetDrops >= 1
check "resetdrop: the accuracy ring is populated", g.accCount == g.sideTotal
check "resetdrop: rolling accuracy stays a valid fraction",
g.rollingAcc(100) >= 0.0 and g.rollingAcc(100) <= 1.0
when isMainModule:
testHorizonMaths()
testHorizonBuckets()
testQuadrantNames()
testFeatureShapes()
testLiteralLayout()
testShiftGeometry()
testRoundTrains()
testRoundReset()
testRoundBoundaryKeepsMachines()
testLearningAccumulatesAcrossRounds()
testGameStartWipesMachines()
testTargetChangeResetsMachines()
testLabelCannotCrossBoundary()
testTickRegressionKeepsMachines()
testStaleObservationsDropped()
testColdModelEmitsNoShift()
testWarmShiftMovesAim()
testReadaptDefaultKnobs()
testRuntimeStatesKnob()
testWindowBuffersAndRebuilds()
testWindowRetrainKeepsDeferredState()
testResetDropTriggersRelearn()
if failures > 0:
echo "\n", failures, " check(s) FAILED"
quit(1)
echo "\nAll tm-horizon checks passed."