The user's first-hand diagnosis: "once the pattern is learnt enough we hit DrussGT,
but as soon as it adapts we are not fast enough to re-adapt again." The gun kept
EVERY sample for the whole battle (which the user explicitly asked for), so stale
evidence weighed the same as new evidence - accumulation without forgetting.
TASK 1 - THE CURVE, MEASURED FIRST (prequential side accuracy, 15 rounds x 4
horizons, deciles of the eval stream):
arm early% late% decay
frozen (early only) 64.9 61.4 -3.5
accum (shipped) 76.4 75.3 -1.1
window N=150 84.7 84.6 -0.0
resetdrop 5pp 83.2 84.3 +1.0
**The decay IS real but modest and LOCALIZED IN THE LAST ~30% of the round**
(last-third vs middle-third: frozen -7.7pp, accum -3.7, window -1.3, resetdrop -1.3).
TASK 3 - WHICH FIX HELPS (late-half accuracy, shuffled control in parens):
accum (shipped) 75.3
**window N=150 84.6 (51)** +9.3pp
**resetdrop 5pp 84.3 (51)** +9.0pp
rehearse-all (retrain, NO forgetting) 79.5
- **Sliding window: +9.3pp late (within-round), +9.7 (cross-round), +10.3 (shield).**
The shuffled control stays ~50-51%, so it learns the ENEMY, not noise.
- **Forgetting is the essential ingredient**: the same periodic retrain WITHOUT
forgetting reaches only ~79.5%, so roughly half the gain is the retraining
mechanics and half is the forgetting.
- Change detection ties the window on late accuracy and gives the best decay, but
at a 5pp threshold it fired 300-500 times in the offline stream - noisy.
- INERTIA IS REDUNDANT WITH THE WINDOW: lower inertia helps the keep-everything
model a lot (accum late 75.3 -> 83.3 at N=8) but leaves the window FLAT
(84.6-84.7 at every N). Low inertia and forgetting are SUBSTITUTES, and
forgetting is the robust one - `TR_TMHORIZON_NSTATES=8` is NOT the primary fix.
VERDICT: SHIPPING CANDIDATE = `TR_TMHORIZON_WINDOW=150`, kept at default 0 until a
live A/B confirms.
HONEST CAVEATS THAT SET EXPECTATIONS:
- The fixtures are OPEN-LOOP (DrussGT does not react to our bullets), so a true
mid-round adaptation is NOT present; the dominant measured effect is the LEVEL
gap, not the decay magnitude. The user's "it adapts" magnitude is still INFERRED.
- The harness uses a STRAIGHT-LINE base while the live gun uses Pattern's
prediction, and it ignores the h-tick label delay, so its absolute side accuracy
(75-85%) is INFLATED: the same gun measured ~52% = chance live against DrussGT.
So +9.3pp is a real ARM DELTA, not a promise that the gun now clears the ~80%
accuracy wall that hits need. A live A/B must decide.
Adds `measure_tm_readapt.nim` (prequential harness with the mandatory shuffled
control, within-round + cross-round protocols, inertia sweep) and its captured
results. test_tm_horizon 79 -> 104 (five new groups). test_tm_diag 48,
test_tm_automata_diag 55, test_tm_clause_shape 66, test_rack_membership 48 pass.
ModularBot compiles.
.gitignore: switched from a broad `measure_*` pattern to EXPLICIT binary names.
The broad rule was too blunt - it also excluded the `.txt` results file, which made
`git add` refuse the whole commit twice. Sources stay tracked; binaries do not.
The user's request: "when our bot is low OR enemy is low, it is useless to use high
power instead low fast bullets have more chances to finish the enemy. Let's do a
math slope: starting from some health down, the power goes down with it."
1. ENERGY SLOPE (`TR_POWER_ENERGY_*`), replacing the old hard step at 50 energy:
cap = ENERGY_MAX at/above ENERGY_HI, ENERGY_MIN at/below ENERGY_LO, LINEAR in
power between, clamped. Defaults HI=80 LO=20 MIN=0.5 MAX=3.0, so no cap >=80,
0.5 at <=20, and e.g. E=65 -> 2.375, E=50 -> 1.75, E=35 -> 1.125.
Rationale: bullet speed is 20-3p, so lower power = FASTER bullet (less lead
error, higher hit chance), fires more often (10+2p) and drains slower (p/shot).
E[dE] = p(3P-1) => break-even hit probability is 1/3 INDEPENDENT of power, and
our measured rates are 5-27%, far below it.
2. FINISHING CAP (`TR_POWER_FINISH_KILL`, default ON): cap power at the SMALLEST
bullet that still removes the enemy's remaining energy -
`E<=4 -> p=E/4` (min 0.1), `4<E<=16 -> p=(E+2)/6`, `E>16 -> no cap`.
Rationale, and it makes the user's instinct stronger than a heuristic: server
1.3.1 caps the damage SCORE at the energy ACTUALLY REMOVED, so overkill is
WASTED damage AND ~6x the energy for ZERO extra score. Damage is 4p (p<=1) /
6p-2 (p>1).
Both are min-composed with the existing far/below-average caps, may only LOWER
power (exhaustively tested), and are exempt while ramming.
`TR_POWER_POLICY=0` still returns the uncapped control exactly.
MEASURED ENERGY SAVING (offline replay of the DrussGT fixtures, 28,797 ticks):
arm shots energy meanP E/1k ticks vs cliff
control(uncapped) 1913 4646 2.43 161.4 -90.2%
cliff (today) 2363 2443 1.03 84.8 0.0%
slope 2404 2178 0.91 75.6 ** 10.9% LESS **
slope+finish 2404 2167 0.90 75.3 ** 11.3% LESS **
So the slope spends ~11% less energy than the cliff AND fires slightly MORE shots
(2404 vs 2363) - both directions at once.
HONEST NOTE on the finishing rule's reach here: ticks where the enemy is low
(0 < E <= 16) are only 2252/28797 = 7.8% of these fixtures, so finishing adds just
~11 energy of saving against DrussGT. It matters in CLOSER fights, not this one.
Verification: test_power_policy 58 (was 26) in BOTH the default and TR_POWER_POLICY=0
control arms - slope at E=100/80/65/50/35/20/5, powerToKill across E=0.1..100, the
inverse-cover property for E<=16, monotonicity, ram exemption, and an exhaustive
sweep proving power <= preference. Guards: test_gun_harness 39, test_vbullet_metric
11, test_power_selection 3, test_adaptive_radar 41, test_tfil_ring_weights 24,
test_ram_decision 40, test_rack_membership 48, test_selector_tiebreak 19,
test_tm_pattern_registration 20, test_vbullet_admit_gate 12. acceptance
12/12 PASS. ModularBot compiles release.
Adds `common_libs/tests/measure_power_policy.nim` (the energy/histogram tool) and
updates docs/env_reference.md for the new `energySlope|finishKill` log reasons.
NOT MEASURED: the battle/hit-rate effect. The offline figures use the fixture
shooter's energy as a proxy, open-loop; the RELATIVE saving is the meaningful part.
The gun has never contributed anything: Tsetlin.vHits was byte-for-byte
equal to Linear.vHits in every measured round of every run, because its
learned correction was always exactly 0.
Six diagnosed defects fixed, plus one that was required to make the first
one work:
1. Type I now conditions on the clause output. It previously rewarded
included true literals unconditionally, omitting Granmo's (c=0, lk=1)
-> toward Exclude counter-force, so true literals ratcheted toward
Include forever. This was the root cause of the saturation.
2. Type II was unreachable dead code: its guard required cOut==1 AND
lits[lit]==0 AND st>0 (included), but cOut==1 guarantees every included
literal is 1. Its direction was wrong too - it should increment EXCLUDED
false literals when the clause fires.
3. Resource allocation restored: Granmo's (T - clip(v,-T,T))/(2T) target
replaces |error|/(2*RESID_MAX); TM_T was only an output normaliser.
4. Label baseline fixed - the factor-2 shrink. predX = linearX + cx, so the
label was delta - cx while the learner's output IS cx, giving
error = delta - 2cx and a fixed point of cx = delta/2: HALF the needed
correction even with perfect feedback. TmTrace now stores linearX/linearY
and training uses delta.
5. Hits no longer zero their label (a hit means |miss| < 18px, not 0).
6. The enemy-energy feature was duplicated - tmEncodeFrame passed
state.selfEnergy with a stale comment claiming enemyEnergy was absent,
while WorldState.enemyEnergy exists. Enemy-energy rules were literally
unrepresentable.
7. REQUIRED EXTRA: tmEvalClause now implements Granmo Eq. 6 - an all-Exclude
clause outputs 1 during learning and 0 during classification. Without it,
fix#1 deadlocks every clause at empty.
MEASURED EFFECT (energy-threshold-turner fixture, seed 1):
mean included literals per active clause 714.0 -> 13.8
active clauses 100/100 -> 53/100
nonzero corrections 8/764 -> 708/764
Tsetlin virtual hits (Linear = 27/400) 27/400 -> 69/400
Divergence achieved: offline on 7/8 fixtures, and in a live gauntlet
(RandomMover: Tsetlin 199/1200 vs Linear 288/1200, vDropped=vStarved=0).
Tsetlin now LEARNS but is not yet competitive with Linear - the regression
head is untuned, flagged as follow-up rather than claimed as a win.
Also ignores compiled test harnesses that have no file extension, which the
existing '**/tests/test_*' rule misses.
Gun evaluation previously required a full end-to-end battle (Java server +
battle runner + websocket IPC to 2 bot processes, 50 rounds, ~3.4 min) and
yielded only ~300-900 REAL shots across 13 guns -- far too few to rank
guns, which is why tuning needed many repetitions.
VirtualTracker is already a pure function of (WorldState stream, gun list);
the only reason it needed Java was where WorldState came from. So the range
replays a seq[WorldState] through the SAME tracker: offline and online
scores are the same metric by construction, not an approximation.
ACCEPTANCE TEST (the point of the whole thing): record one live round, replay
it offline, compare per-gun virtual hit rates. 12/12 deterministic guns match
EXACTLY, reproduced twice. Tsetlin is compared separately because tmLearnOne
calls rand(). Getting to 12/12 exposed two real ordering quirks in the live
loop: run() calls go() before the aim/fire block, so tickBullets resolves
against the NEXT tick's scan while the prediction used the previous one; and
if the target dies during that go() the final tick's spawn+resolution is
skipped entirely. The recorder emits an end marker for the second case.
The 5th (selected-gun) predict call was verified to be a no-op.
Measured cost: 8 fixtures (1770 ticks, ~92k virtual bullets, 13 guns) replay
in 2.9 s, ~32k virtual bullets/s -- roughly 70x faster and 100x more samples
than a live gauntlet.
Also adds a per-tick WorldState recorder behind const RecordWorldState
(default off, mirrors the ShotLog idiom) which records the state the bot
ACTUALLY builds, staleness included, rather than true positions -- recording
the latter would hand the guns perfect information and produce flattering
scores.
9 new guard checks (33 total, all passing), including fixture round-trip,
replay determinism, stationary->HeadOn 100%, constant-velocity->Linear>HeadOn,
and the energy-threshold turner crossing at t=41.
nimble bin output lands at the garage root, so the existing
'*_garage/out/' ignore rule never covered it and every build dirtied
the tree with a 1.1 MB binary. File stays on disk; build regenerates it.
- Removed non-existent exported/ directory from layout docs
- Replaced hardcoded bot binary entries with *_garage/out/ pattern for build directories
- Added all known bot binaries (GotoTest, OscillatorBot, PPO_Bot, QBot, SAC_LSTM_Bot)
to .gitignore with clarifying comment about Nim's compilation target structure
Nim places compiled binaries at bot root (no extension) and in out/ subdirs.
The out/ pattern catches all build outputs; specific bot binaries listed for
root-level executables since gitignore lacks a reliable "no-extension files" glob.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>