Commit Graph

330 Commits

Author SHA1 Message Date
SirStone 1ea84f5b6e STRAFE: draw the heat field, real bullet danger, corrected ring comment
Three defects the owner hit as "no heat tiles anymore" under TR_MOVEMENT=strafe.

1. The strafe overlay drew ONLY the tiles on its strafe line, so the computed
   heat field was essentially invisible. It now draws the WHOLE field exactly as
   TFIL does (every non-zero tile, yellow->orange->red ramp by field max, integer
   value label) behind the same debugGraphics flag, with TR_STRAFE_HEAT_GRID=0 to
   hide it. The strafe overlays draw on top, unchanged.

2. STRAFE carried the SHIPPED bullet constants (core 10 / aura 5), so a bullet's
   own heat sat exactly ON PathDangerThreshold (10.0) and a bullet was never
   dangerous on its own in this mover; it only ever bit through its corridor.
   Defaults are now the retune's 20/10, exposed as TR_STRAFE_BULLET_CORE /
   TR_STRAFE_BULLET_AURA.

3. The ring mover's header documented CorridorHeat 5.0 / WallHotness 10.0 while
   the code has always been 10.0 / 15.0. A job read the comment and handed out
   sub-threshold heat values, which emptied the field. The comment now states the
   real values and their actual behaviour; no code values changed.

Also sets strafe's heat defaults to the retune shape (bullet 20/10, corridor 10,
wall 15/5, pillar 0), documented with the reason.

Gate A re-run (j110, offline DrussGT fixture, measure_strafe_gates.nim):
  corrected DEFAULT : 24.6% of picks with ZERO safe tile, mean 11.17 safe
  j108 shipped field: 63.4% / 3.70   (reproduced exactly)
  j108 ring retune  :  8.1% / 18.41  (reproduced exactly)
  bullet isolated   : 11.4% / 17.07
The corrected default beats the shipped field but is WORSE than j108's retune
row: the bullet retune alone costs 8.1 -> 11.4, the corridor/wall retune accounts
for the rest. That is the deliberate price of making a bullet dangerous.

Guards green: test_env_report 24 PASS, test_tfil_commit_env 30 PASS (shipped TFIL
default untouched, byte-for-byte), test_tfil_ring_weights 24 PASS. The three new
knobs are registered in the boot env report so the tree-scan guard stays clean.
2026-09-25 23:25:34 +02:00
SirStone dc071b83f3 ModularBot: [result] round/battle outcome log lines (TR_RESULT_LOG, default on)
One greppable '[result]' line per round plus one at battle end, stdout only:
  [result] round 3/7  WE WON    (enemy destroyed)  | us 42.1 energy, them 0.0, 812 ticks | rounds won 3/7
  [result] battle END: rounds won 4/7

Outcome is authoritative from RoundEndedEventForBot.results.rank (1 = winner);
death observations (our onDeath, enemy onBotDeath) and onWonRound refine it
into WE WON (enemy destroyed) / WE DIED (killed) / BOTH DIED (score decided) /
TIMEOUT (score decided). Our own death is reported the instant it happens.
TR_RESULT_LOG registers in the boot env report; default on, only explicit
off-values disable it. No behaviour change - logging/state only.
2026-09-25 22:50:46 +02:00
SirStone a50c0125d5 STRAFE movement: body pinned perpendicular to the threat, reversals by sign flip
New engine movements/strafe.nim, selected by TR_MOVEMENT=strafe (default stays
tfil, byte-identical — test_tfil_commit_env.nim's 30 checks still pass).

Design (the owner's):
- AXIS = incoming bullet's direction when a bullet is in flight, else the
  perpendicular of the enemy bearing. The body heading is kept inside a band
  (TR_STRAFE_BAND, default 20 deg) around the perpendicular LINE; it turns only
  when outside the band, and never turns to face a movement target.
- Candidate tiles on the perpendicular line through our position, both forward
  and backward, within TR_STRAFE_REACH px, with a perpendicular jitter of
  +/- TR_STRAFE_SPREAD tiles. A tile is acceptable when its path max heat is
  <= PathDangerThreshold, the SAME safety rule TFIL uses.
- Move by SIGN only: setForward(+/-MaxSpeed>). Dwell is re-picked after a random
  number of ticks in [TR_STRAFE_DWELL_MIN, TR_STRAFE_DWELL_MAX], on arrival, or
  on a serious threat spike.
- Heat machinery is REUSED from the shipped mover, not re-implemented: the
  exported heatDecay()/bulletMagScale() (j105 time-indexed model) and the
  PillarHotness/PillarRadiance globals (j106 pillar-free default). The heat
  shape is overridable via TR_STRAFE_CORRIDOR_HEAT/WALL_HOTNESS/WALL_RADIANCE
  (defaults = the shipped TFIL field).
- GUI overlay: strafe line, threat axis, candidate tiles (safe/unsafe), chosen
  target, sign-coloured movement ray, and the heading band.

Gates (offline, recorded DrussGT fixture, 20026 ticks):
- A TILE AVAILABILITY: shipped heat field -> a safe tile exists on only 36.6%
  of picks (63.4% fall back to the least-hot tile); the ring retune
  (corridor 5, wall 10/5) raises it to 91.9%.
- B PREDICTABILITY: reversal-interval entropy 5.84 bits vs TFIL 5.09; direction
  entropy 1.00 both; long-lag autocorrelation ~0 for both (no periodic
  component). Fewer reversals (710 vs 1453) and more full-speed ticks.
  measurements: common_libs/tests/measure_strafe_gates.nim

Also registers TR_STRAFE_* in the boot env report (ModularBot_garage/src/
env_report.nim) and wires the engine into ModularBot.nim (hold -> strafe,
ram trigger -> rammer).
2026-09-25 22:38:52 +02:00
SirStone 99cf9e5c82 tfil heat/pillar A/B: record the owner's decision to keep the virtual pillar removed 2026-09-25 22:24:47 +02:00
SirStone 48f38b80e7 TFIL heat-time + virtual pillar: live A/B (6 arms x 70 rounds) - neither change beats the pre-change mover
Runs the pre-registered A/B for the two movement changes in HEAD: the
time-indexed bullet heat (TR_TFIL_HEAT_TIME, fca8993) and the removal of the
invented virtual centre pillar (d0750ab). One frozen binary from HEAD vs real
DrussGT: 6 arms x 10 runs x 7 rounds = 60 battles, 420 rounds, 0 failed.

Judged on damage/run and ROUND WINS only (hit rate and hits-taken are context):
hit rate would have inverted the verdict again - tau3 has the best pooled hit
rate of all arms (11.56%) and the fewest round wins (20/70).

RESULT (vs the reconstructed pre-change mover "old"):
  heat-time HURTS. tau3/tau5/tau9 lose 1.3-1.7 wins/run (p=0.0010-0.0125) and
  deal 22-38 less damage/run (p=0.004-0.047); tau15 is a wash on wins (p=0.64)
  and 22 damage/run lower (p=0.046). Nothing improves either metric.
  pillar removal does nothing measurable. old vs pillaoff: +5.7 damage/run
  (p=0.71), +0.5 wins/run (35 vs 30, p=0.43), 30.8 MORE damage taken/run
  without the pillar (p=0.040). The mechanism check proves the knob works
  (centre-box occupancy 0.09% -> 2.37%, p<0.0001; range 469 -> 443 px,
  p=0.0002), so this is a real behaviour change that buys nothing. At n=10 the
  pillar contrast is inside the MDE (33 damage/run, 1.2 wins/run), so this is
  not a proven regression.

Flags that the shipped default (pillar removed) should be reverted to the
TR_TFIL_PILLAR_ON behaviour; heat-time stays off.

Adds tools/ab/arms_heat_pillar.txt and tools/ab/ab_mechanism.py (per-tick
mechanism check: central-box occupancy, range distribution, live enemy-bullet
proximity) plus the captured summary/report fixtures.
2026-09-25 22:23:11 +02:00
SirStone f58d65d2e8 TMComposites gate: per-gun confidence faithful for 3 guns; no pair composes
Adds a per-sample intrinsic-confidence field (GunPrediction.confidence,
threaded through FeedbackEvent/VirtualBullet, populated by Pattern, DecayGF,
KNN, GuessFactor, Tsetlin, TMHorizon) and an offline recorder + analyzer that
reproduce the paper's Figure 2 per gun and its Eq-8 composite.

Measured on 3 held-out tr-bridge DrussGT battles (33k ticks, ~133k samples/gun):
- FAITHFUL: DecayGF (rho +0.133), KNN (+0.090), Pattern (+0.064, weak).
- GuessFactor is ANTI-faithful (rho -0.067); Tsetlin c_max is useless (0.001).
- No pair of guns specialises complementarily: the same gun dominates both
  high-confidence slices in every pair.
- Eq-8 alpha-normalised confidence-weighted composite: 18.41% vs Pattern
  20.45% (McNemar p=3.1e-126). Faithful-only variant 18.68%, still loses.
  Shuffle control passes weakly (composite > shuffle, p=4e-14) so ~0.7pp of
  competence is real but ~2pp short. Offline veto: design is dead.

See docs/tmcomposites_gate.md.
2026-09-25 22:04:39 +02:00
SirStone d0750ab020 TFIL: remove the virtual centre pillar from the shipped default; register j102 env reads
The default mover painted a 30/10 radiance blob on the arena centre even
though the arena has NO physical pillar there, creating a 4x4 tile
(144x144 px) exclusion zone over open centre floor. Set
PillarHotness/PillarRadiance to 0/0 in the shipped default (matching the
ring variant) and add TR_TFIL_PILLAR_ON=1 to restore the old 30/10 field
for A/B without a rebuild; registered in env_report.

Because the shipped default legitimately changed, the default-path parity
golden (fixtures/tfil_commit_default.golden) was regenerated from the NEW
default, with an explicit 'deliberate default change' note in the test so
a future failure is treated as a real regression.

Also register the three env reads job j102 added in common_libs/bitbrain
(TR_BITBRAIN_MODE / _DECAY_EVERY / _DECAY_SHIFT), which the env-report
guard was failing on.

Verification: test_env_report all green; test_tfil_commit_env 30/30.
2026-09-25 22:02:00 +02:00
SirStone fca899376e TFIL: time-indexed bullet heat (TR_TFIL_HEAT_TIME, DEFAULT OFF)
Make danger a function of time-to-arrival instead of flat distance. Bullet
core/aura/corridor heat becomes magnitude(power) * decay(dt), dt = along/speed:

  * decay(dt) = exp(-dt/tau) is a function of TIME; a fixed tau projects a
    pixel reach of speed*tau, so fast/weak bullets get a longer slope and slow
    ones a shorter one — derived from speed = 20 - 3*power, not hand-tuned.
    tau = TR_TFIL_HEAT_TAU.
  * magnitude(power) scales the near-end heat with power from DAMAGE
    (calcBulletDamage = 4p, linear in p; SCORE_PER_BULLET_DAMAGE = 1.0). Hit
    probability is FLAT across power (docs/env_reference.md), so risk does not
    justify power scaling — the cost of the hit does. Floored at 1.0 so a weak
    bullet's near end is never less dangerous than the flat model.
    Gain = TR_TFIL_HEAT_POWER_GAIN.

Every source is already f(dt), so the time-indexed planner (evaluate a cell at
the tick the bot would ARRIVE, i.e. heatDecay(dt - arrivalDelay)) is a one-line
change. It is intentionally NOT implemented here.

Default path is byte-identical: with TR_TFIL_HEAT_TIME unset both factors are
exactly 1.0 (IEEE x*1.0 is exact), and the committed golden replay in
common_libs/tests/test_tfil_commit_env.nim (20,026 ticks) still passes
byte-for-byte against the pre-change mover. The debug corridor outline is also
drawn only to the model's reach when enabled, so the GUI shows the shortening.

Offline field measurement (common_libs/tests/measure_tfil_heat_time.nim,
46,054 fixture ticks, tau=9/gain=1): corridor reach drops from 443px
wall-to-wall to 143px mean (32% retained); fraction of tiles > 10 goes
0.61 -> 0.57; largest contiguous safe region 118 -> 140 tiles; mean
distance-to-nearest-safe-tile 49 -> 42px. Saturation stays high because wall
radiance + pillar alone are 44% of tiles over threshold and are untouched.

Registers the three knobs in env_report (report + known-name set).
2026-09-25 21:47:54 +02:00
SirStone 4270136948 docs: correct counted-SBC decay amortised cost figure 2026-09-25 08:43:48 +02:00
SirStone d85ff53d34 State-window gate: a temporal window of wave-relative states does NOT beat a single state
Re-runs the SBC coincidence premise as a cheap veto test on the 70-battle
live-vs-DrussGT corpus (/tmp/tfil_ab2).  Defines the wave-relative state
(lat/vlat/toa/room/turn, 5/7.9/10 bits at Q=2/3/4), quantises the miss offset
at the bullet's arrival into 7 bins, and sweeps window length K in
{1,4,8,16,32,48} with an interpolated suffix-backoff model under a BY-BATTLE
70/30 split (3 seeds).

Result: NO.  On the pre-fire frame the window is worse than the single
fire-tick state at every K/Q/A (e.g. K=8 costs +0.35..+0.46 bits).  On the
during-flight frame the entire apparent gain is the later decision tick, not
the window; the single state alone drops 2.70 -> 1.31 bits as K goes 1 -> 32.
The shuffle-order control confirms recency matters but the windows do not:
by Q=4/K=8 they average ~1 observation and never recur.  The single state
survives as a strong predictor (log-loss 2.346 vs 2.698 majority; bin
accuracy 0.409 vs 0.235).

Gate only: no gun, no live-win claim.
2026-09-25 08:43:36 +02:00
SirStone 40ba96f649 BitBrain SBC: counted mode + global decay (forgetting, probabilities)
Adds an smCounted storage mode alongside the default smBitset. Each
(i,j,class) cell becomes a saturating uint8 counter; learn increments it and
a global fractional decay (c -= c shr decayShift every decayEvery learns)
makes forgetting possible. infer sums raw counters; new inferProb sums the
per-cell posterior P(class|cell) (scale-free, recommended readout).

Bitset path is the default and byte-for-byte unchanged: test_bitbrain 56/56
(was 32), and test_bitbrain_mnist reproduces 97.210% corrected / 96.540%
bug-compatible exactly.

Counted mode configurable at runtime (TR_BITBRAIN_MODE / TR_BITBRAIN_DECAY_*)
and compile time (-d:bitbrainDecay*). Measured: forgetting (86.2% vs 48.9% on
a permuted-label stream), probabilities (rare-class balanced 0.998 vs 0.500),
and the stationary cost (counted hurts MNIST; see docs/bitbrain_counted_sbc.md).

Harness: common_libs/tests/measure_counted_sbc.nim
2026-09-25 08:39:10 +02:00
SirStone 39e06719fb Campaign phase 2 ledger: live lead-gain sweep (gains above 1.0 do NOT beat Pattern)
6 arms x 7 runs x 7 rounds vs real DrussGT on one frozen binary (2747ebd).
Validity check PASSES: g100 (fixed gain 1.0) is statistically indistinguishable
from control (dmg p=0.65, wins p=0.62, ALL hit rate +0.04pp p=0.92; zero [bb]
lines = provably no correction). Fixed gains above 1.0 LOSE at 300+: gfix150
-94 dmg/run (p=0.0006), 4/49 vs 16/49 wins (p=0.009), -3.85pp at 300-450
(p=0.0006) and -2.25pp at 450+ (p=0.0023). The hypothesis arm ghi (learner
allowed above 1) is directionally positive but inside the MDE (+14.3 dmg/run
p=0.41; +0.83pp at 450+ p=0.35). Kill the gain axis in both directions.

Includes the mandatory correction notice: Phase 1's [1,1,1,0,0] is LIVE-REFUTED
by 140fe25, and the standing rule that offline is veto-only / live decides.
2026-09-25 00:24:45 +02:00
SirStone 2747ebd323 BitBrain: TR_BITBRAIN_GAINS env knob (candidate set + fixed-gain degenerate)
Task A of campaign phase 2: the lead-gain candidate set is now pure env, so the
live arms need no recompile.

- common_libs/guns/bitbrain_gun.nim: BB_GAINS_ENV (TR_BITBRAIN_GAINS); the
  candidate list is parsed once at gun construction into a dynamic seq, so the
  hit counts/hit rates are sized to it. Unset/unparsable -> the shipped
  BB_CAND set [0,0.25,0.5,0.75,1.0] (byte-identical behaviour). Exactly ONE
  candidate degenerates to a FIXED gain applied from the first shot (learning
  bypassed), still gated to the long bands. parseGains clamps to [0,8],
  de-dupes and sorts so the argmax tie rule is unchanged. The [bb] line now
  prints the APPLIED gain AND the resulting angular shift, so a run's
  correction is auditable from stdout.
- ModularBot_garage/src/env_report.nim: emit TR_BITBRAIN_GAINS (resolved
  candidate set) and add BB_GAINS_ENV to the known-name list.
- tools/ab/arms_leadgain.txt: the 6-arm phase-2 sweep definition.
2026-09-25 00:15:27 +02:00
SirStone c305ef4212 BitBrain campaign phase 1: the missing gain sweep and BitBrain as a lead-gain corrector
Task A — the gain region Phase 0 never covered (gain < 1). Extend the
prediction-quality ruler with gain 0.25/0.50/0.75 arms and a fixed causal
per-band arm. Full 70-run result: the hitProxy-argmax curve is
[1.00, 1.00, 1.00, 0.00, 0.00] — Pattern below 300 px, HeadOn above —
worth +0.13 pp at 300-450 and +2.16 pp at 450+ (0.0767 -> 0.0984). Lead
correlation is identical for every g>0 (Pearson is scale-invariant), so a
shrinking gain adds no lead information; and the least-squares optimum
[1,1,1,1,0.25,0.25] diverges from the hitProxy optimum because Pattern's
lead errors are bimodal.

Task B — rebuild guns/bitbrain_gun.nim as a lead-gain corrector:
aim = LOS + gain*(patternAim - LOS), gain learned online per range band by
ranking candidate gains on the hit-probability proxy (the observed lead
label via tmhObservedAt), gated to range >= 300 px. ADE+SBC output removed.
Offline (70 runs): Pattern below 300 px, +0.37 pp at 300-450, +1.90 pp at
450+ (hitProxy 0.0957 vs 0.0767), matching the fixed rule to within 0.27 pp
at 450+ and exceeding it at 300-450. ~0.0007 ms/tick marginal (old gun
~0.114 ms/tick). Default off; rack membership, env report and guard tests
(bitbrain 32, registration 13, rack 48, tm_pattern 20, env_report) unchanged
and green.

Ledger: docs/bitbrain_campaign.md Phase 1, including the three on-file
negatives and the causal-shippability note. No live claim.
2026-09-25 00:10:19 +02:00
SirStone 140fe2519a HeadOn (no-lead) vs Pattern LIVE at long range: clean negative, offline ruler killed
2 arms x 15 runs x 7 rounds, one frozen binary from HEAD a82c864, real DrussGT,
server-side events sidecar. Shipped rack is onlyPattern, so control=Pattern-only
and headon=HeadOn-only (TR_RACK_PATTERN=off TR_RACK_HEADON=both).

  arm      dmg/run  dmgtk/run  round wins  shots/run
  control      279        211     48/105       785
  headon        14        228      0/105       580

Round wins and dmg/run both separate at p<0.0001 (MC permutation, se 0.0000),
~7x the damage MDE (35.8). Per range band (pooled, 15 runs):
  300-450: Pattern 12.3% (4590 shots) vs HeadOn 0.6% (3701)  p<0.0001, MDE 2.0pp
  450+   : Pattern  9.2% (6671)       vs HeadOn 0.4% (4177)  p<0.0001, MDE 1.1pp
HeadOn loses EVERY long-range band by 20-23x, so the whole-battle loss is not a
close-range artefact.

The offline ruler (prediction_quality_results.txt) predicted the opposite: HeadOn
meanAbs 14.61 vs Pattern 17.53 at 300-450 and 12.33 vs 16.19 at 450+, hitProxy
.105/.104 and .098/.077 (+27%). That is an open-loop replay of a FIXED enemy
track, so it cannot see that a different bullet makes the surfer dodge
differently; live, the static gun does not lead at all.

TR_PATTERN_RAD_SCALE arms were skipped: applyRadial scales aim DISTANCE along an
unchanged bearing, so it cannot express 'less lead' (bearing is what firing uses).
HeadOn confirmed to ignore bulletSpeed (head_on.nim:9), liveness OK 15/15.

Adds the range-band analyzer tools/ab/ab_range_bands.py (reuses the lead-capture
Run alignment) and the captured fixtures. Does not touch bitbrain_gun.nim /
bitbrain_campaign.md (job-100).
2026-09-24 23:52:49 +02:00
SirStone a82c864c60 bitbrain campaign phase 0: offline prediction-quality ruler and the bar
New harness (common_libs/gun_harness/prediction_quality.nim +
common_libs/tests/run_prediction_quality.nim): per-gun single-tick aim error in
degrees against the true continuous interception point on the recorded
live-vs-real-DrussGT corpus (/tmp/tfil_ab2/out, 70 runs, 899607 ticks), per
range band, with the hit-probability proxy mean(|err|<=atan(18/range)).
Validated: recorded hits separate from misses 13.34x px (reference 11.59x),
perfect-oracle max |err| = 0, correct ordering on synthetic ground truth, two
full runs byte-identical. Fixed a wrap180 bug (Nim float mod keeps the dividend
sign) that inflated the negative error tail.

Bar (mean|err| deg [hitProxy] at 450+): Pattern 16.19 [0.077], naive-linear
22.86 [0.054], TMHorizon 16.20 [0.076], BitBrain 16.20 [0.077], static HeadOn
12.33 [0.098], oracle 0 [1.0]. Lead-gain sweep on Pattern is a dead end (1.0
wins every band). Naive-linear applies ~1.8x Pattern's lead but carries no more
lead information (corr 0.178 vs 0.165) and is strictly worse. Ledger:
docs/bitbrain_campaign.md. All verdicts remain live-only.
2026-09-24 23:40:56 +02:00
SirStone 32a5e72fac Gun mixing (TMHorizon+BitBrain) vs DrussGT: clean negative, no dodge disruption
4 arms x 7 runs vs real DrussGT. mix alternates the two guns 476 times/7 runs
(liveness OK) but our bullets are no more varied (power sd / aim-offset sd flat)
and DrussGT's dodge quality is unchanged (miss/tick mix-pat +0.03, p=0.66; MDE
3.8%). mix wins 24/49 = the 49% baseline; the user's 6/10 has P=0.353 at 49%.
New tools/ab/ab_dodge_analyze.py splits the validated per-shot dodge instrument
by arm and adds gun-switch/power/bearing liveness; fixtures committed.
2026-09-24 23:07:05 +02:00
SirStone d93ce444c0 BitBrain verdict: clean negative at 30 runs/arm; analyzer gets MC + Mann-Whitney + MDE
- docs/bitbrain_gun_verdict.md: control vs bb_decay (decay SBC memory) vs a
  provably-zero placebo, 30 runs/arm vs real DrussGT. Nothing separates
  (bb_decay +3.3 dmg/run, p=0.71; round wins 97/210 vs 97/210, p=1.00); the
  7-run shape does not replicate. TR_BITBRAIN_RANGE=0 is clamped to 1.0 deg
  (bitbrain_gun.nim:207) so it is NOT a zero-shift placebo; TR_BITBRAIN_MIN_OBS
  unreachable is used instead.
- tools/ab/ab_analyze.py: keep exact enumeration for C(n,na)<=20e6 (7v7), add
  a seeded Monte-Carlo permutation test (1e6 draws, 0x5eed5eed) with its
  standard error, a tie-corrected Mann-Whitney U cross-check, a minimum
  detectable effect line, all-pairs comparisons, and a [bb] shift check.
- tools/ab/README.md: document the new analyzer output.
2026-09-24 23:02:59 +02:00
SirStone f91e121965 lead capture by range: we under-lead (0.40->0.135) — a RANGE effect, not a power one
Measure, for every shot ModularBot fires at the real DrussGT, the lead we
actually applied vs the lead the enemy's motion required, from the recorded
live battles (/tmp/tfil_ab2, 70 battles / 490 rounds / 54926 shots, plus a
35-battle powtest replication of a different binary).

- requiredLead from an AIM-INDEPENDENT interception solve (bullet speed vs
  enemy truth), appliedLead from the server-recorded bullet bearing.
- capture = applied/required, guarded at 2px lateral lead (1.6% excluded);
  headline metric is the robust proportional slope.
- validation: hits 11.6px mean miss / 80.8% inside 18px, misses 134px,
  11.6x separation; 496/496 death + 70/70 owner attributions correct.

Direct answer: capture falls with RANGE (capSlp 0.401 -> 0.135, and
|err|/tolerance 1.27 -> 7.54) but is FLAT across fired POWER within a band
(450+, enemy alive: 0.154 / 0.127 / 0.127). The sub-0.5 long-range shots
(1.18% hit) are finishKill endgame shots at a near-dead DrussGT, not a
lead-capture failure. A naive linear predictor captures 0.29-0.60; we reach
46-67% of that, so the under-lead is real but capture=1.0 is unattainable
against a dodger (oracle required lead).
2026-09-24 22:47:18 +02:00
SirStone 795a0e59fe BitBrain gun (id 16): Pattern-relative ADE+SBC aim corrector, default off
Wire the verified common_libs/bitbrain ADE+SBC library into ModularBot as a
fine-grained angular corrector on top of Pattern's prediction, the shape the
offline gate test measured (argmax readout over N correction classes).

- common_libs/guns/bitbrain_gun.nim: new gun. Input = the existing TMHorizon
  53 bits (tmhBaseBits + tmhLits); output = argmax class centre over
  +-TR_BITBRAIN_RANGE, applied by rotating the Pattern point around the shooter
  exactly as tmhApplyShift does. Label = the +h-tick fact from TmHorizonGun's
  own observation ring (never across a round). Prequential (defer + resolve).
  AD layer synthesised online for our binary inputs (center=0): heuristic
  cold-start thresholds + running-histogram ~1% percentile init + the library's
  adaptThresholds. Memory modes perRound (default, measured best) / retained /
  decay (periodic partial SBC wipe). Lazy network build + local RNG, so the
  default path builds nothing and consumes no global randomness.
- tm_horizon.nim: export tmhUpdateHistory and add tmhObservedAt (label seam).
- selector.nim: register BITBRAIN at rack id 16, default rmOff, in the SAME
  commit as the id and the wiring (the aed579b admission bug is not repeated).
- ModularBot.nim: id 16 wired through predict/spawn/onResult/resets/colors,
  arrays grown 16->17, spawn gated on rack admission, per-round/per-battle/
  target reset hooks.
- env_report.nim: report every TR_BITBRAIN_* knob + add names to the known set.
- tests: update the rack length literals; new test_bitbrain_registration
  (default-parity: off, lazy, global-RNG clean).

Guard counts unchanged: rack 48, tm_pattern_registration 20, vbullet_admit 12,
env_report 25, and the rest of the suite green.
2026-09-24 22:25:04 +02:00
SirStone 1fec87537d gitignore: keep the BitBrain primary-source zip local (explicit path, not a pattern) 2026-09-24 22:09:36 +02:00
SirStone 17c50159ba tools/ab: reusable A/B runner + analyzer (frozen-HEAD build, exact permutation test, liveness check) 2026-09-24 22:08:59 +02:00
SirStone e670788eee env_reference: the ring mover was NEVER measured offline - correct a false label
The offline-harness audit (`e40c849`, `docs/offline_harness_trust.md`) found that the
claim "best offline hit rate of anything measured" for the ring mover was false.

The 20.28% figure is a LIVE number: `docs/feature_ab_results.md` and commit `bfdcdf8`
record 35 real-DrussGT bridge battles with a server-side event sidecar as the ground
truth, and the 6/49 round wins is likewise live. There is no offline measurement of
the ring mover anywhere - the offline harness scores GUNS, not movements, and has no
movement driver at all.

So this was NOT an offline-vs-live calibration failure, which is how it has been
described repeatedly (including by the orchestrator). It was a METRIC MISMATCH: a
movement arm judged on hit rate instead of damage/run and round wins - and hit rate is
precisely the metric that concealed its collapse.

The lesson previously attached to this result was therefore the wrong one. The
paragraph now says what actually happened and points at the audit.

Docs-only; no code touched.
2026-09-24 21:44:17 +02:00
SirStone e40c8493a6 Offline harness: audited, calibrated against live, and one real bug fixed
AUDIT (docs/offline_harness_trust.md, new):
- Re-ran acceptance_offline_vs_online myself TWICE: 12/12 deterministic guns
  exact both times (264 ticks/enemyId=1, 244 ticks/enemyId=2), death boundary
  included. The offline range reproduces the live bot's own per-gun virtual
  telemetry exactly.
- Re-verified the (fireTick, powerBin) wave-pairing fix: exact-key lookup,
  collisions counted not silently mislabelled; test_wave_pairing 17/17 PASS.
- The offline score is the live TELEMETRY (last-100 virtual hit rate) but NOT
  the live BATTLE score (damage/round wins). Two-level answer, documented.
- bmPoint scores up to one tick-step (~17px) PAST its documented aim distance,
  while the tie-break probe scores exactly the aim point. Real, low-impact,
  deliberately NOT fixed (point metric is non-default, measured negative, and
  the committed point baselines would silently change).
- bmPoint/bmPath, perfect-info captures, conditional-on-selection live rates,
  and hit-rate-as-objective-for-movement all catalogued as non-apples comparisons.

FIX (unambiguous, fail-before/pass-after):
- common_libs/tests/range_guns.nim: buildAllGunDrivers defaulted to
  enableTmSelector=true, so run_range / analyze_selector / test_power_selection /
  measure_power_policy spawned gun 13 (TMSelect) - a gun the shipped bot NEVER
  spawns. The shared VirtualTracker ring is order-sensitive, so those 4
  spawns/tick permuted the learning guns' resolution order (the exact confound
  4cd5618 fixed for the acceptance test, left broken for every default caller).
  Default is now false (mirror the shipped rack). Impact on
  tr_drussgt_vs_modularbot: Tsetlin 18.8->18.5%, KNN 7.5->7.2%, TMSelect 15.2->0.
- New guard common_libs/tests/test_range_rack_parity.nim (3 checks); proven to
  FAIL before and PASS after by stash-reverting the fix.

CALIBRATION (offline prediction vs live outcome, 9 usable arms):
- Direction agreement 3/9 = 33%. Split by domain: open-loop (single-tick
  prediction / metric / threshold) 3/3; closed-loop (adaptation / range /
  movement / selection) 0/6. Small, non-random, hand-assembled set - no
  correlation coefficient is claimed.
- The four motivating "offline wins" re-attributed: ring mover was NEVER
  offline (it is a live server-side hit rate, mislabelled "offline" in
  env_reference.md:342 and commit 7f6ccfb); TMHorizon window/NSTATES and the TM
  gun are the H3 classifier-accuracy harness (not hit rate); TFIL is the H2
  open-loop movement replay, whose mechanism prediction was right and whose
  outcome prediction was wrong.
- Open-loop hypothesis tested: TR_RACK_* knobs leave the offline range output
  BYTE-IDENTICAL (the replay never calls the selector), and the range has no
  driver for guns 14/15 (TMPATTERN/TMHORIZON). BUG vs LIMIT separated.

VERDICT: trust the harness for single-tick prediction quality only; never for
anything running through the closed loop. MEASURED vs INFERRED labelled.

Green counts unchanged: test_gun_harness 39, test_vbullet_metric 11,
test_power_selection 3, test_power_policy 58, test_adaptive_radar 41,
test_tfil_ring_weights 24, test_ram_decision 40, test_rack_membership 48,
test_selector_tiebreak 19, test_tm_pattern_registration 20,
test_vbullet_admit_gate 12, test_tm_horizon 104, test_tm_diag 48,
test_tm_automata_diag 55, test_tm_clause_shape 66, test_env_report 25,
test_tfil_commit_env 30 (as-is). New: test_range_rack_parity 3.
2026-09-24 21:40:38 +02:00
SirStone f41cd08718 BitBrain gate test: fine-grained aim correction vs naive + Pattern (offline) 2026-09-24 21:23:29 +02:00
SirStone 1adefaba26 DrussGT dodge vs fired power: no movement response once range is controlled
Answers the user's hypothesis that DrussGT dodges low-power shots better.
Measured on 70 live battles / 490 rounds / 54939 real shots vs real DrussGT
(/tmp/tfil_ab2) and replicated on 35 more battles / 24280 shots (/tmp/powtest).

Power is not randomly assigned - our policy caps it by RANGE
(TR_POWER_FAR_DIST=200 -> 1.0) and by OUR OWN ENERGY (the slope), so inside a
range band power is almost a deterministic function of our energy and a naive
low-vs-high comparison is secretly a losing-vs-healthy comparison. Everything
is stratified by range band and backed by a within-band shuffled-label null
(arrival re-derived, so the null keeps the kinematic channel), a round-cluster
bootstrap, and a within-shot CONTROL window 40 ticks later when the bullet is
long gone.

RESULT: no behavioural response. In band 450+ the raw miss distance at arrival
is +8.25 px [+5.39,+11.25] for HIGH power - but per flight tick it is 4.52 vs
4.51 px/tick (delta -0.01 [-0.12,+0.10]), i.e. entirely the 2.13-tick longer
flight window of the slower bullet. Fixed-12-tick lateral displacement is flat
(55.63 vs 55.47, -0.15 [-1.15,+0.84]) and turn rate / speed are flat. The whole
difference is already present 5 ticks after the trigger pull (+4.2 px) and is
just as large in the bullet-free control window (+5.6 px), so it is a property
of the low-energy situation, not of the shot. Hit rate is flat (0.10 vs 0.09).

Corpus/attribution notes: e*=DrussGT (subject), s*=ModularBot, per
TrBattleCapture.java; the Tank-Royale owner id is NOT stable across runs and is
recovered per battle from fire geometry + the energy decrement, cross-checked on
496/496 death events. Geometry validated on the server's own hits (mean miss
11.6 px, 80.6% inside the 18 px radius).
2026-09-24 21:22:38 +02:00
SirStone 77e6dace01 BitBrain: generic clean-room ADE + SBC library with MNIST acceptance
Implement the BitBrain (Address Decoder Element + Sparse Binary Coincidence)
classifier as a generic, deterministic Nim library under common_libs/bitbrain/,
written from the published algorithm (Front. Neuroinform. 17:1125844), not from
the GPL-3.0 reference C.

- ade.nim: signed thresholded random projection (scale 64 / centre 127 defaults
  reproduce the reference), multi-width ADs, optional deterministic homeostatic
  threshold adaptation. Hebbian longevity and Metropolis-Hastings sampling are
  described but not implemented.
- sbc.nim: packed class-bit coincidence memory; idempotent learn, counting
  inference.
- bitbrain.nim: container over several ADs and SBCs, online learn/infer, argmax
  readout, memory accounting.
- tests: 32 unit checks (idempotence, planted rule + monotone online curve,
  shuffled-label chance control, unseen input, homeostasis, memory).
- tests/test_bitbrain_mnist.nim: loads the reference pretrained ADs/thresholds
  and MNIST from /tmp, reproduces the reference exactly - 97.210% corrected and
  96.540% bug-compatible - confirming the port.

No gun/wiring integration yet; inputs and outputs to be agreed separately.
2026-09-24 20:52:45 +02:00
SirStone cc11ede824 TFIL: measure the commitment A/B on 490 live DrussGT rounds - no effect
Runs the five-arm commitment A/B that 19bf461 only implemented. One frozen
binary (git archive 19bf461, sha256 46e7ce19...) vs the real DrussGT through
tools/robocode_shim/run_bridge_battle.sh: 7 runs x 7 rounds per arm against
real DrussGT in the pre-registered block, plus an independent replication
block (runs 8-14) - 70 battles, 490 rounds, 5 arms in parallel.

  A control (shipped)          B TR_TFIL_TILE_REPLAN=off
  C B + TR_TFIL_NO_REV=1       D TR_TFIL_TILE_REPLAN=enemy
  E B + TR_TFIL_COMMIT_TICKS=30   (TR_TFIL_COMMIT_LOG=1 on every arm)

RESULT: no arm improves damage/run or round wins vs the shipped mover. Pooled
(14 runs/arm, exact two-sided permutation test over all C(28,14) relabellings):

  arm  damage/run (delta, p)        round wins (delta, p)
  A    274.44                       2.79  (39/98)
  B    274.07 (-0.37, p=0.97)       3.07  (+0.29, p=0.66)
  C    257.34 (-17.10, p=0.16)      2.43  (-0.36, p=0.47)
  D    286.61 (+12.17, p=0.36)      3.29  (+0.50, p=0.35)
  E    265.29 (-9.15, p=0.56)       2.93  (+0.14, p=0.89)

The replication is what settles it: block 1 alone showed D at +20.99 damage
(p=0.26); block 2 put D at +3.36. Nothing replicates.

The PREMISE fails. Honouring the commitment does not reduce reversals - it
raises the reversal-pick rate from 33.9% (control) to 50.9% (B) / 57.0% (E);
the no-reversal arm C only claws part of it back (42.4%). The post-reversal
speed dip (~4.9 -> ~4.0 px/tick at +1..+2 calls) is identical in every arm, so
it is a property of turning around, not of the tile replan, and arm A - the
arm carrying the bug - has the LOWEST mean abs(speed) of all five.

Both premise premises are also caught by the metric the report is careful
about: arm C takes significantly FEWER hits (87.2 vs 96.5/run, p=0.0022) while
dealing the least damage and winning the fewest rounds - a movement change
alters hits taken, shots fired and round length at once, which is exactly why
damage and round wins are the primary metrics and hit rate is not.

Treatment verified per arm from the per-tick TR_TFIL_COMMIT_LOG: the pick
interval moves off ~5 calls only where the arm says it should (A 5.05 with
96.6% tile_self replans; B 14.77 with 0%; C 14.79; D 5.29 with 95.7%
tile_enemy; E 26.62), matching the offline fixture replay (5.07 / 14.65 /
14.65 / 4.94 / 27.57). Every arm ran.

Round wins are the final RESULTS firstPlaces, cross-checked against each
round's BotDeathEvent (35/35 and 35/35 runs agree). Damage is the server's
BulletHitBotEvent damage, attributed by numeric owner/victim id and verified
against the capture's own event counters.

measure_tfil_commit_ab.nim reads the raw run artifacts (JSONL states, events
sidecar, capture stdout, commit log) and prints the whole deliverable:
arm table, per-run values, the exact per-run permutation test, the exact
round-level hypergeometric test (flagged anti-conservative - rounds cluster
within a run), the treatment diagnostics, and the direct answer. It reproduces
either report offline from the committed summary:

  nim c -r --path:common_libs common_libs/tests/measure_tfil_commit_ab.nim \
      --from-summary common_libs/tests/fixtures/tfil_commit_ab_results_runs14.json

The raw ~200 MB of battle artifacts are not committed; the committed summary
holds every per-run value plus the arm-level diagnostics.
2026-09-24 20:24:34 +02:00
SirStone c8b2a8c6c5 docs: reverse-engineer the BitBrain algorithm from its C source; assess as a gun
Reads docs/BitBrain_C_code.zip end to end (full_mnist_2048.c 568 lines +
bitarray.h), reconciles every weight/data file's byte size against its
loader, and answers two questions:

1. What the algorithm is: signed thresholded address decoders (random
   projections) -> write-once sparse binary coincidence bit tensors ->
   counting/argmax readout. Thresholds and ADs come from an unsupervised
   program that is NOT in the zip; only the SBC tensors are learned here.
   The supervised rule is genuinely online, single-pass, order-free, and
   has no learning rate - so it satisfies ModularBot's reset-on-enemy-change
   constraint. Found and isolated a real bug: read_from_sbc's uint8_t
   bit_test truncates the 32-bit bit test, so only ADEs with i%%32 < 8 are
   ever counted. Shipped reader: 96.540%% (reproduced exactly); with the bug
   fixed: 97.210%%.

2. Whether it can become a gun: not on this evidence. Measured full
   inference at ~0.56 ms/sample (cost is not the blocker), but the AD layer
   is unobtainable, and this repo has already shown the binding constraint
   is the target signal, not the learner - the TM head scored 2.08pp below
   its own majority class and the AD/SBC primitives are already in the
   16-family shootout as WiSARD/Bloom/SDM. Recommends a readout swap on the
   existing WiSARD feature extractor and a Gate-2b-style >=80%% side-accuracy
   gate before any port.

Also: recommends committing the 11 MB zip under its explicit name.
2026-09-24 20:17:08 +02:00
SirStone 19bf4610e4 TFIL: env-gated commitment arms (tile-replan cancel is a bug)
The tile-change replan cancels the 15-tick movement commitment whenever OUR
tile changes. With GridSize=36 and speed up to 8 px/tick that is every ~5
ticks, so the commitment is cancelled by the motion it commands (measured:
96.9% of picks were tile-change replans, 33.8% of picks reversed direction).

Adds four env knobs, every default reproducing the shipped mover
byte-for-byte:
  TR_TFIL_TILE_REPLAN  self (default) | off | enemy
  TR_TFIL_COMMIT_TICKS 15 (default)
  TR_TFIL_NO_REV       0 (default)
  TR_TFIL_COMMIT_LOG   off (default, JSONL per-tick diagnostics)

- `off` honours the commitment; the danger replan stays the safety valve.
- `enemy` keys the cancel to the TARGET's tile displacement (the intent the
  original comment claimed).
- `TR_TFIL_NO_REV` down-weights (never filters) tiles >90 deg from the travel
  direction; the pool can never be emptied.

Default-path parity is guarded by test_tfil_commit_env.nim, which replays
tools/fixtures/tr_drussgt_vs_modularbot.jsonl and diffs every move command
against a golden generated from the pre-change build (git archive f842ac0).
env_report known-name list updated for the four new names.
2026-09-23 23:28:45 +02:00
SirStone f842ac0f76 Env report Part 2: warn on misnamed env vars (unknown + compile-time defines)
The boot report printed the raw env and the resolved values, but it never
told the user when a name was WRONG - which is the failure mode that cost
real time: `TMH_NSTATES` was exported as an env var although it is a
compile-time `{.intdefine.}` (`guns/tm_horizon.nim:102`), so the export was
a silent no-op. Add the two warning paths, both boot-only and stdout:

  * unknown TR_*/GUN_* names are named explicitly. The known set is built
    from the modules' exported env-name constants; the remaining inline
    reads are listed once, and common_libs/tests/test_env_report.nim scans
    the tree and fails if a name read anywhere is missing.
  * any `{.intdefine.}`/`{.strdefine.}`/`{.booldefine.}` symbol present as
    an env var is flagged, with the real runtime equivalent when one exists
    (`TMH_NSTATES` -> `TR_TMHORIZON_NSTATES`) and an explicit "does not
    exist - this knob is compile-time only" when it does not. The map is
    derived by grepping the tree; the same test re-greps and fails on drift.

Warnings are emitted only when the env is dirty, so a correct run keeps the
documented A/B/build shape. TR_ENV_REPORT=0 still suppresses everything.

test_env_report.nim: 25 checks (known set, pure helpers, tree literal scan,
tree define scan). All 15 existing guard tests keep their exact counts.
2026-09-23 08:38:37 +02:00
SirStone 036979e78e env_reference: verify every knob against the code, fix the trap that cost real time
DOCS ONLY. The user set TMH_NSTATES, which is a compile-time {-d:intdefine.}
(-d:TMH_NSTATES=2, tm_horizon.nim:102), not an env var; the runtime var is
TR_TMHORIZON_NSTATES (tm_horizon.nim:139). This rewrites the reference so that
class of confusion cannot recur.

What was WRONG and is now fixed:
- "unparseable warns and falls back" was false for the numeric/bool knobs: only
  the enumerated string knobs warn; envInt/envFloat/envBool fall back silently.
- TR_POWER_LOG/TR_RAM_LOG/TR_MOVEMENT_LOG/TR_RECORD_WORLDSTATE/
  TR_RADAR_FORCE_SPIN/TR_RADAR_SCANLOG/TR_TRACKER_PROBE are read with existsEnv,
  so TR_POWER_LOG=0 turns the log ON. Documented per knob.
- TR_POWER_ENERGY_MIN is a CAP at low energy, not a minimum-power floor.
- TR_TMHORIZON_* knobs are inert unless TR_RACK_TMHORIZON=both; the doc implied
  they were live.
- GUN_SELECTOR_WINDOW is clamped 1..100; TR_MOVEMENT silently falls back to tfil
  for any value other than tfil_ring.
- The ring file's own header comment (corridor 5 / wall 10) is stale; the code
  defaults are 10.0/15.0 (commit 7f6ccfb).
- Compile-time section was incomplete and conflated the two TM modules:
  tm_pattern uses TM_NCLAUSES/TM_NSTATES, tsetlin uses TM_N_CLAUSES/TM_N_STATES,
  and -d:TM_S_DEF is defined in BOTH.

What was ADDED:
- "COMPILE-TIME vs RUNTIME: the two namespaces": the only define/env pair is
  TMH_NSTATES <-> TR_TMHORIZON_NSTATES; everything else is compile-time only.
- "Did my env vars actually reach the bot?": the /proc exec-time check, the note
  that grepping only TR_|GUN_ hides a wrongly-named var (grep -i tmh), and the
  boot report described as an interface with the two sections + build identity.
- "Measured verdicts" table: window 26.5% vs 49.0% p=0.036 (harmful live),
  TMHorizon N=2/8/64 42.9/42.9/53.1% (all p>=0.8), power floor 22/49 vs 21/49
  p=1.0, sub-1.0 accuracy 10.80% vs 10.35% p=0.37, shipped bot 49% vs DrussGT.
- "Names that look real but do nothing": TMH_NSTATES (env), TR_VBULLET_METRIC,
  TR_POWER_LOW_ENERGY, TR_TRACKER_RECONCILE.
- The adaptive-melee radar's compile-time constants (no env form).
2026-09-23 08:29:10 +02:00
SirStone 9bf3005850 Boot-time env report + fix the env reference
The bot is spawned by the server/GUI, so it inherits the SERVER's
environment. The user could not tell whether their exports reached the
bot, so print a one-shot greppable report at boot:

  grep '^\[env\]' /tmp/modularbot_stdout.log

Section A prints every TR_*/GUN_* this process actually received, the
count vs the total env size, a loud warning when nothing matched, and
the process identity (pid/ppid, cwd, self command line, and the PARENT
command line) so the spawn trap is obvious. Section B prints the
resolved effective value of every documented knob with its source
(env|default), including clamps and the rack's empty-set fallback.
Build identity (NimVersion, compile date/time, binary path/size/mtime)
pins the exact artifact. Suppress with TR_ENV_REPORT=0.

docs/env_reference.md: add the missing GUN_SHOTLOG_PATH,
GUN_SELECTOR_MINOBS/FLOOR/POOL/RANK/SHRINK/SEED, TR_ENV_REPORT and
-d:TM_NCLAUSES; record the measured TR_TMHORIZON_WINDOW verdict; and
add a prominent 'Did my env vars actually reach the bot?' section with
the boot report, the /proc/PID/environ no-code check, the correct GUI
launch recipe, and how to prove the trap deliberately.
2026-09-23 08:22:04 +02:00
SirStone 7f6ccfb015 Ring mover heat retune (author: the user) - dodge bullets, ignore walls/centre,
re-plan 3x more often

Uncommitted working-tree change in `the_floor_is_lava_ring.nim`, confirmed by the
user as theirs. Committing it so it stops appearing in every job's `git status`.

WHAT IT DOES - a coherent "be far more afraid of bullets" strategy:
- BULLETS 2x hotter: BulletCore 10 -> 20, BulletAura 5 -> 10.
- Walls weaker and thinner: WallRadiance 10 -> 5; the `WallHotness` default
  10 -> 15 (so the heat falls off over ~3 tiles instead of 1, with a higher peak).
- CENTRE PILLAR DISABLED: PillarHotness 30 -> 0, PillarRadiance 10 -> 0, so the
  bot may now use the middle of the arena instead of treating it as a no-go zone.
- MUCH more reactive: CommitTicks 15 -> 5 and MinCommitTicks 5 -> 0, i.e. the
  dodge target is re-chosen 3x more often and a replan is allowed immediately.
- `CorridorHeat` default 5 -> 10.

This applies ONLY to the ring mover, which is NOT the shipped movement (the default
is `tfil`), so the shipped bot is unaffected.

TWO FACTS RECORDED, not objections - it is the user's call:
1. `CorridorHeat = 10` now EQUALS `PathDangerThreshold = 10`. The earlier measured
   taming set it to 5 precisely so a single corridor could no longer poison a path
   on its own (a corridor at 10 puts every tile along it at/over the threshold).
   That is part of what lifted the "band-weightable" fraction from 10.8% to 26.5%.
   At 10 the range weighting gets less to work with.
2. UNMEASURED: this retune has not been A/B'd. The ring mover it modifies was
   itself measured as a glass cannon (best offline hit rate, but round wins
   16/49 -> 6/49, p=0.012). The two changes push in the same direction - hotter
   bullets and more frequent replanning mean MORE dodging and LESS time spent
   closing to the 100-200px band where our hit rate peaks (27% vs 5% at 450px).
   Worth an A/B before drawing conclusions, but it is an experiment, not a
   shipped change.
2026-09-23 00:30:52 +02:00
SirStone 69debbe347 Re-adaptation: the sliding WINDOW is a control-validated +9.3pp on late accuracy
The user's first-hand diagnosis: "once the pattern is learnt enough we hit DrussGT,
but as soon as it adapts we are not fast enough to re-adapt again." The gun kept
EVERY sample for the whole battle (which the user explicitly asked for), so stale
evidence weighed the same as new evidence - accumulation without forgetting.

TASK 1 - THE CURVE, MEASURED FIRST (prequential side accuracy, 15 rounds x 4
horizons, deciles of the eval stream):
  arm                 early%  late%  decay
  frozen (early only)   64.9   61.4   -3.5
  accum (shipped)       76.4   75.3   -1.1
  window N=150          84.7   84.6   -0.0
  resetdrop 5pp         83.2   84.3   +1.0
**The decay IS real but modest and LOCALIZED IN THE LAST ~30% of the round**
(last-third vs middle-third: frozen -7.7pp, accum -3.7, window -1.3, resetdrop -1.3).

TASK 3 - WHICH FIX HELPS (late-half accuracy, shuffled control in parens):
  accum (shipped)                       75.3
  **window N=150                       84.6  (51)**   +9.3pp
  **resetdrop 5pp                      84.3  (51)**   +9.0pp
  rehearse-all (retrain, NO forgetting) 79.5
- **Sliding window: +9.3pp late (within-round), +9.7 (cross-round), +10.3 (shield).**
  The shuffled control stays ~50-51%, so it learns the ENEMY, not noise.
- **Forgetting is the essential ingredient**: the same periodic retrain WITHOUT
  forgetting reaches only ~79.5%, so roughly half the gain is the retraining
  mechanics and half is the forgetting.
- Change detection ties the window on late accuracy and gives the best decay, but
  at a 5pp threshold it fired 300-500 times in the offline stream - noisy.
- INERTIA IS REDUNDANT WITH THE WINDOW: lower inertia helps the keep-everything
  model a lot (accum late 75.3 -> 83.3 at N=8) but leaves the window FLAT
  (84.6-84.7 at every N). Low inertia and forgetting are SUBSTITUTES, and
  forgetting is the robust one - `TR_TMHORIZON_NSTATES=8` is NOT the primary fix.

VERDICT: SHIPPING CANDIDATE = `TR_TMHORIZON_WINDOW=150`, kept at default 0 until a
live A/B confirms.

HONEST CAVEATS THAT SET EXPECTATIONS:
- The fixtures are OPEN-LOOP (DrussGT does not react to our bullets), so a true
  mid-round adaptation is NOT present; the dominant measured effect is the LEVEL
  gap, not the decay magnitude. The user's "it adapts" magnitude is still INFERRED.
- The harness uses a STRAIGHT-LINE base while the live gun uses Pattern's
  prediction, and it ignores the h-tick label delay, so its absolute side accuracy
  (75-85%) is INFLATED: the same gun measured ~52% = chance live against DrussGT.
  So +9.3pp is a real ARM DELTA, not a promise that the gun now clears the ~80%
  accuracy wall that hits need. A live A/B must decide.

Adds `measure_tm_readapt.nim` (prequential harness with the mandatory shuffled
control, within-round + cross-round protocols, inertia sweep) and its captured
results. test_tm_horizon 79 -> 104 (five new groups). test_tm_diag 48,
test_tm_automata_diag 55, test_tm_clause_shape 66, test_rack_membership 48 pass.
ModularBot compiles.

.gitignore: switched from a broad `measure_*` pattern to EXPLICIT binary names.
The broad rule was too blunt - it also excluded the `.txt` results file, which made
`git add` refuse the whole commit twice. Sources stay tracked; binaries do not.
2026-09-23 00:29:37 +02:00
SirStone b68707c867 Energy economy: the cliff becomes a SLOPE, plus a finishing cap. 11% less energy.
The user's request: "when our bot is low OR enemy is low, it is useless to use high
power instead low fast bullets have more chances to finish the enemy. Let's do a
math slope: starting from some health down, the power goes down with it."

1. ENERGY SLOPE (`TR_POWER_ENERGY_*`), replacing the old hard step at 50 energy:
   cap = ENERGY_MAX at/above ENERGY_HI, ENERGY_MIN at/below ENERGY_LO, LINEAR in
   power between, clamped. Defaults HI=80 LO=20 MIN=0.5 MAX=3.0, so no cap >=80,
   0.5 at <=20, and e.g. E=65 -> 2.375, E=50 -> 1.75, E=35 -> 1.125.
   Rationale: bullet speed is 20-3p, so lower power = FASTER bullet (less lead
   error, higher hit chance), fires more often (10+2p) and drains slower (p/shot).
   E[dE] = p(3P-1) => break-even hit probability is 1/3 INDEPENDENT of power, and
   our measured rates are 5-27%, far below it.

2. FINISHING CAP (`TR_POWER_FINISH_KILL`, default ON): cap power at the SMALLEST
   bullet that still removes the enemy's remaining energy -
   `E<=4 -> p=E/4` (min 0.1), `4<E<=16 -> p=(E+2)/6`, `E>16 -> no cap`.
   Rationale, and it makes the user's instinct stronger than a heuristic: server
   1.3.1 caps the damage SCORE at the energy ACTUALLY REMOVED, so overkill is
   WASTED damage AND ~6x the energy for ZERO extra score. Damage is 4p (p<=1) /
   6p-2 (p>1).

Both are min-composed with the existing far/below-average caps, may only LOWER
power (exhaustively tested), and are exempt while ramming.
`TR_POWER_POLICY=0` still returns the uncapped control exactly.

MEASURED ENERGY SAVING (offline replay of the DrussGT fixtures, 28,797 ticks):
  arm              shots  energy  meanP  E/1k ticks   vs cliff
  control(uncapped) 1913    4646   2.43    161.4      -90.2%
  cliff (today)     2363    2443   1.03     84.8       0.0%
  slope             2404    2178   0.91     75.6    ** 10.9% LESS **
  slope+finish      2404    2167   0.90     75.3    ** 11.3% LESS **
So the slope spends ~11% less energy than the cliff AND fires slightly MORE shots
(2404 vs 2363) - both directions at once.

HONEST NOTE on the finishing rule's reach here: ticks where the enemy is low
(0 < E <= 16) are only 2252/28797 = 7.8% of these fixtures, so finishing adds just
~11 energy of saving against DrussGT. It matters in CLOSER fights, not this one.

Verification: test_power_policy 58 (was 26) in BOTH the default and TR_POWER_POLICY=0
control arms - slope at E=100/80/65/50/35/20/5, powerToKill across E=0.1..100, the
inverse-cover property for E<=16, monotonicity, ram exemption, and an exhaustive
sweep proving power <= preference. Guards: test_gun_harness 39, test_vbullet_metric
11, test_power_selection 3, test_adaptive_radar 41, test_tfil_ring_weights 24,
test_ram_decision 40, test_rack_membership 48, test_selector_tiebreak 19,
test_tm_pattern_registration 20, test_vbullet_admit_gate 12. acceptance
12/12 PASS. ModularBot compiles release.

Adds `common_libs/tests/measure_power_policy.nim` (the energy/histogram tool) and
updates docs/env_reference.md for the new `energySlope|finishKill` log reasons.

NOT MEASURED: the battle/hit-rate effect. The offline figures use the fixture
shooter's energy as a proxy, open-loop; the RELATIVE saving is the meaningful part.
2026-09-23 00:12:04 +02:00
SirStone 81af5854df FIX an incomplete commit: TMHORIZON was admitted unconditionally at HEAD
MY ERROR. `aed579b` committed `ModularBot.nim` (which wires the new gun as id 15)
and `guns/tm_horizon.nim`, but I staged only three files and left job-72's RACK
REGISTRATION behind in the working tree. Consequence at HEAD:

- `common_libs/gun_harness/selector.nim` still had the 15-entry rack table
  (ids 0..14) with no TMHORIZON name, so
- `TR_RACK_TMHORIZON` was never read (a silent no-op), and
- id 15 is PAST the table, and `rackAdmitted` documents "an id past the
  membership table is admitted" -> **TMHORIZON was admitted unconditionally**.

So the committed default rack was `Pattern + TMHORIZON`, not the intended
`Pattern only`, and the gun could not be switched off by env at HEAD. This was
caught by a live A/B job that had to work around it with `GUN_RACK_DISABLE=15`.

FIX (this commit): the rack table goes to 16 entries with `TMHORIZON` appended at
id 15 defaulting to `rmOff`, plus the two tests whose length literals follow it
(`test_rack_membership` 15->16, `test_tm_pattern_registration` 15->16).
`DefaultRackMembership` is again Pattern-only with every other gun `off`.

NOTE ON SCOPE: `selector.nim` in the working tree also contains a CONCURRENT job's
power-policy threading (an `enemyEnergy` parameter). That work is still in flight
and is deliberately NOT included here - only the three rack hunks were staged.
The remainder stays unstaged for its own commit.

The lesson, recorded because it has bitten twice tonight in different forms: a
change is not committed until its registration/table counterpart is, and
`git add` of a hand-picked file list is exactly how a half-change ships.
2026-09-22 23:46:56 +02:00
SirStone 9ba932d1b1 docs: complete environment variable reference (every knob, its default, and why)
Single place to answer 'what env vars exist, what do they default to, and what are
they for'. Grouped by area, with the measured reason each knob exists recorded
next to it, since several defaults are counterintuitive:

- The shipped rack is PATTERN ONLY, with the revert one-liner included.
- GUN_SELECTOR_TIEBREAK defaults off because it measured NEGATIVE on real hit
  rate, and the per-tick random draw inside the tie band is load-bearing
  (commitment cost 7.02% -> 5.10%, p=0.002).
- The power policy is ON because turning it off measured 10.6% -> 7.9% real hit
  rate; TR_POWER_FINISH_KILL exists because server 1.3.1 caps the damage SCORE at
  the energy actually removed, so overkill scores nothing.
- The ring mover is NOT the default and must not be shipped: best offline hit
  rate of anything measured, but it halves survival (16/49 -> 6/49, p=0.012).
- The proactive ram is off because it converts 0/6 times.
- TR_PATTERN_RAD_* are kept but are structurally incapable of changing the shot
  (the aim is bearing-only, measured byte-identical on the path metric).
- TR_TMHORIZON_NSTATES is the automata inertia ('mood'): lower = adapts faster.
- TR_VBULLET_ADMIT_ONLY=1 gave +68% tick rate (87 -> 146 ticks/s).

Documents the read-once-at-start convention and the operational trap that the bot
is spawned by the server/GUI, so a variable exported in an unrelated terminal does
NOT reach it. Also lists the compile-time -d: knobs and the test-harness jars.
Snapshot of the source at this commit; f9f8d84-era knobs added by the in-flight
jobs are included where already present.
2026-09-22 23:45:53 +02:00
SirStone aed579b3af TM horizon: retain learning ACROSS ROUNDS, reset only when the ENEMY changes
The user's requirement: "every battle i means from round 1 to round end-battle, so
retain all learning until the enemy change." What was built wiped the Tsetlin
machines EVERY ROUND, in two places (`onRoundStarted` and the gun's own
tick-regression self-reset), so in a 7-round battle each round started cold,
trained ~360 samples and threw them away - discarding most of its one chance to
do what was asked: overfit the current enemy over the whole battle.

THE FIX - two kinds of state, two triggers:
- **`resetRoundState` (per ROUND)**: the observation ring, pending/deferred
  labels, the bullet proxy, motion history, per-tick caches, `roundStartTrained`.
  These MUST clear every round, because bots teleport back to the starting corners
  between rounds - an old position would build a garbage label. (That exact class
  of bug shipped 36-58% wrong labels in the old gun.)
- **`resetLearning` (per BATTLE / per ENEMY)**: both Tsetlin machines, `trained`,
  `sideCorrect/sideTotal`, all histograms, the magnitude median, `pendingDropped`,
  `observedTargetId`. These now SURVIVE round boundaries.
Triggers for the machine wipe: `onGameStarted` (primary) plus a redundant
`roundNumber <= 1` fallback in `onRoundStarted`; and a TARGET CHANGE
(`targetChanged`, knob `TR_TMHORIZON_RESET_ON_TARGET` default on - a no-op in 1v1,
fires on melee target switches; first acquisition never wipes). The
tick-regression self-reset now clears ONLY per-round state.
Still NO cross-battle persistence: grep for file I/O in the gun finds none.

PROOF IT WORKS (live 2-round battle, `TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both`):
  [tmh-reset] reason=game_start trained_was=0
  [tmh-reset] reason=round1     trained_was=0
  ...exactly TWO reset lines in the whole battle, both at battle start, and NONE
  at the round-2 boundary. And the per-round summaries:
  [tmh-round] trained=1249 thisRound=1249 ... sideAcc=711/1177  (60.4%)
  [tmh-round] trained=2226 thisRound=977  ... sideAcc=1361/2154 (63.2%)
`trained` CLIMBED 1249 -> 2226 across the boundary, and side accuracy rose
60.4% -> 63.2% in round 2 (one battle - suggestive, not proof).

Unit tests: `test_tm_horizon` 79 (was 54), including "trained SURVIVES the
boundary", "clause states SURVIVE", "trained climbs round1->round2", "game-start
wipes and all clauses end Exclude", "different enemy wipes / same enemy does not /
knob-off does not", and crucially "a label CANNOT be built across a round
boundary" (ringValidCount==0, ringHas(oldTick)==false, pendingCount==0) - the
single most dangerous interaction of this change.

Guards: test_tm_horizon 79, test_rack_membership 48, test_gun_harness 39,
test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41,
test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 40,
test_selector_tiebreak 19, test_tm_pattern_registration 20,
test_vbullet_admit_gate 12, test_tm_diag 48, test_tm_automata_diag 55,
test_tm_clause_shape 66. acceptance_offline_vs_online 12/12 VERDICT PASS.
Shipped rack unchanged: DefaultRackMembership is still Pattern-only, TMHORIZON off.

Residual (pre-existing, out of scope, stated): the internal base
`PatternMatcherGun` has a rolling move-history buffer that is NOT cleared at round
boundaries - it never was, and the SHIPPED Pattern gun carries history across
rounds too. It cannot affect label correctness (labels come from `g.ring`), only
base-prediction quality in a round's first ticks.
2026-09-22 23:24:11 +02:00
SirStone be74e369eb Gate 2b: the shift form WORKS - but break-even needs ~80% side accuracy and the
problem gives 60%. Thread closed with a number, not a shrug.

Gate 2 (fd2f7f6) said STOP because every correction lost hits. But the shifts it
applied were the conditional MEDIANS of the error (4-16 deg) while the gate metric
is HITS, and the naive guess is UNBIASED - so shifting an already-centred
distribution can only destroy near-target mass. That suggested the experiment was
MISCALIBRATED rather than the idea being dead. One cheap test settled it.

ARMS (pooled, both primary fixtures, 25 rounds, N~9433/h; deltas in pp vs naive):
  h    naive   TM-med   TM-hit  TM-x0.25   PERF-SIGN   shuffled
  15    39.3    29.0(-10.3)  35.1(-4.2)  37.9(-1.5)  **47.0 (+7.7)**  35.1(-4.2)
  20    28.2    19.2(-9.0)   22.8(-5.4)  25.6(-2.6)  **33.4 (+5.1)**  22.0(-6.2)
  25    20.9    13.1(-7.8)   15.4(-5.5)  18.8(-2.2)  **25.6 (+4.7)**  16.0(-4.9)
  30    16.4    10.5(-5.9)   11.4(-5.0)  14.1(-2.3)  **18.3 (+1.9)**  11.2(-5.2)

1. **THE FORM CAN BUY HITS.** Perfect side knowledge + the hit-optimal shift gains
   **+4.9 pp mean** (+1.9..+7.7), positive in ~19/25 rounds and in BOTH fixtures -
   and it improves the median AND p90 residual. So the premise is NOT structurally
   impossible.
2. **BUT THE TM'S 60% SIDE ACCURACY LOSES** (-4.2 pp quadrant, -3.3..-5.6 side).
   Shrinking the shift toward zero monotonically reduces the loss but never turns
   it positive.
3. **BREAK-EVEN IS ~80% SIDE ACCURACY** (synthetic sweep: 0.70 -> -1.6 pp, **0.80
   -> 0.0**, 1.00 -> +4.9 pp). The observed side signal tops out at **~60% (TM) /
   53-56% (turn-only rule)** against 50% chance, and the headroom study found it is
   essentially ONE WEAK FEATURE. **No realistic predictor of this problem clears
   the 80% wall.**

WHY IT IS SO SENSITIVE - the hits-vs-shift curve (h15, calibration): shift 0 ->
36.2%, +-2deg -> 27.4/26.8, +-4deg -> 14.2, +-6deg -> 11.7, +-10deg -> 8.3.
Shifting unconditionally is catastrophic. The whole value comes from shifting ONLY
when the side is known, because the signed-error distribution becomes ONE-SIDED
once you condition on the true side. And the hit-optimal PERF-SIGN shift is only
**+-2 to 3.5 deg** - an order of magnitude smaller than the 4-16 deg medians Gate 2
used, which is exactly why Gate 2's calibration could never win.

VERDICT: **MISCALIBRATED, but practically dead at achievable accuracy.**
Recommendation: stop the per-bucket-shift direction, and record the RIGHT reason -
an **accuracy wall at ~80%**, not an impossibility of the form. If ever revived,
the only viable path is a side predictor that materially exceeds 60%; a better
calibration cannot fix it (we already used the hit-optimal one).

METHOD NOTE: the hit-optimal shift is fitted on the calibration slice (the
pipeline's existing out-of-sample offset split) and applied on the later eval
slice - NOT fitted on the TM training slice, to avoid in-sample label leakage.
INTEGRITY: TM-med reproduces Gate 2's committed deltas to the decimal
(-10.3/-9.0/-7.8/-5.9); PS-shift0 and TM-x0.0 are exactly naive; the shuffled
control never improves. Guards all pass (test_tm_diag 48, test_tm_automata_diag 55,
test_tm_clause_shape 66, diag_synthetic, diag_automata_validation, test_gun_harness,
test_vbullet_metric, test_power_selection, test_adaptive_radar,
test_tfil_ring_weights, test_power_policy, test_ram_decision, test_rack_membership,
test_selector_tiebreak, test_tm_pattern_registration, test_vbullet_admit_gate).

Caveats: DrussGT-only; enemy movement is a closed-loop response to our CURRENT
movement; offline observation is perfect while live we see the enemy only on scans
-> all absolute hit fractions are optimistic upper bounds, only deltas are
meaningful.
2026-09-22 22:21:32 +02:00
SirStone fd2f7f608c GATE 2: STOP - the TM learns the side but the correction cannot buy hits
Offline four-arm pipeline: extracts the 49-bit draft spec + a 4-bit horizon
one-hot (53 bits) from the DrussGT fixtures, derives fact-based quadrant labels
(side vs the naive guess, magnitude vs the TRAIN median), trains the validated
`tm_diag/tm_core.nim` fresh PER ROUND with a fit / calibrate / eval within-round
split, applies each predicted quadrant's out-of-sample conditional-median offset
and measures the residual angular error. 50 clauses, N=64, s=3.0, 5 epochs
(verdict identical at 1 and 10). Estimated hit = |residual| < atan(18px/range).

POOLED (tr_drussgt_vs_modularbot + _shield, N~9.1-9.4k per horizon):
  h    arm        med|err|   p90|err|   hit%
  15   naive        4.08      14.24     39.3
  15   TM           4.81      13.53     29.0
  15   shuffled     4.45      14.58     34.3
  15   turn-only    4.13      14.10     37.1
  20   naive        6.81      21.60     28.2
  20   TM           7.61      20.85     19.2
  25   naive        9.71      29.13     20.9
  25   TM          10.33      28.63     13.1
  30   naive       12.62      36.12     16.4
  30   TM          13.09      35.17     10.5
Delta hits (pp): TM-naive **-10.3 / -9.0 / -7.8 / -5.9**;
**TM-turn -8.1 / -5.1 / -2.0 / -0.9** (h=15/20/25/30).
Median miss: TM is WORSE by +0.5..+0.8 deg. p90: marginally better by 0.5-1.3 deg.
So it pulls in the TAIL but not the typical miss.
Shuffled control: quadrant accuracy 22.9-24.8% ~= 25% chance, and dHit <= 0 at
every horizon -> NO LEAK. (The mild negative is the expected cost of applying a
noisy offset, not leakage.)

THE MECHANISM, and it is the important part: **the learning is REAL - the TM gets
the side right 59.6-61.0% vs 50% (quadrant 34-35% vs 25% chance; turn-only
53-56%) - but it does not translate into hits because the per-quadrant
conditional-median offsets (~4-16 deg for "big") are FAR LARGER than the body
half-angle (~2.6-3.4 deg at typical range).** The naive guess is essentially
UNBIASED (the median signed error is 0.00 deg at every horizon, from the headroom
study), so its error is centred on the target; shifting an already-centred
distribution away from zero DESTROYS near-target mass. That is why the turn-only
arm loses too (-2.2 to -5.8 pp): any constant shift hurts.

VERDICT AS IMPLEMENTED: **STOP. There is no case for building the new TM gun on
this evidence.**

NOT YET DISTINGUISHED, and worth one cheap test before the idea is declared dead:
whether this is a STRUCTURALLY dead application (no shift can help, because the
baseline is unbiased and the correction is coarser than the target) or merely a
MISCALIBRATED one (the offset was fitted to minimise the conditional MEDIAN of
the error, which is NOT the objective - hits are maximised by the shift that
maximises P(|error| < body), typically a SMALLER shift or none at all on a dense
near-zero distribution). That distinction decides whether the whole "TM predicts
the enemy's position" premise is dead or only this instantiation of it.

Caveats: DrussGT-only; the enemy's movement is a closed-loop response to our
CURRENT movement so the numbers are conditional on how we move now; offline
observation is perfect every tick while live we see the enemy only on scans, so
all of this is an UPPER BOUND. The bullet block was proxy-based (no gun heading or
power is recorded in the fixtures) - INFERRED, and stated as such rather than
silently dropped.
2026-09-22 22:06:01 +02:00
SirStone 9064377740 s settled from OUR code (higher = LONGER clauses), and it cannot rescue the gun
=== TASK 1: THE DIRECTION QUESTION, ANSWERED WITH A DEMONSTRATION ===
I told the user `s` controls clause length but refused to claim the DIRECTION,
because I had seen it described both ways. It is now read out of our own code -
one site per core, in the Type I branch of `tmLearnDir` (`guns/tm_pattern.nim:258`,
`guns/tsetlin.nim:224`, `tm_diag/tm_core.nim:110`):

  if pol * d > 0.0:
    if lits[lit] == 1:
      if cOut == 1:  if rand < (s-1)/s: st += 1   # toward Include, w.p. (s-1)/s
      else:          if rand < 1/s:     st -= 1   # toward Exclude, w.p. 1/s
    else:            if rand < 1/s:     st -= 1   # toward Exclude, w.p. 1/s
=> **HIGHER `s` GIVES LONGER CLAUSES.** The include step runs w.p. (s-1)/s
(rising with s); both exclude steps run w.p. 1/s (falling with s).

DEMONSTRATED (49-bit draft, planted 2-literal rule, 3000 train / 1500 eval):
  s=1.0 len 1.51 acc 100%   s=2.0 len 1.82 acc 100%   s=5.0 len 3.02 acc 100%
  s=1.5 len 1.42 acc 100%   s=3.0 len 2.19 acc 100%   s=10 len 4.04 acc 99.7%
                            s=20  len 5.17 acc 94.5%
WHY s=1.0 DEGENERATES: (s-1)/s = 0 so the include step NEVER fires while 1/s = 1
so BOTH exclude steps always fire - Type I can only remove literals, so a clause
can grow only through the Type II penalty. (On random labels that leaves 43/120
non-empty clauses vs 120/120 at s>=3.)
USABLE RANGE ~[1.5, 5]. tm_pattern uses 3.0; tsetlin uses 1.5.

=== TASK 4: WOULD `s` HELP THE SHIPPED GUN? NO - MEASURED ===
Recompiling the offline driver with -d:TM_S_DEF=<v> (source untouched) retrains
the gun end to end:
  s      mean len   verdict     warm acc   margin vs majority
  1.5     15.82     too long     32.22%      -2.03pp
  2.0     14.99     too long     34.03%      -0.22pp
  3.0*    19.17     too long     35.72%      +1.48pp   (*shipped)
  5.0     20.47     too long     34.72%      +0.47pp
Lowering `s` shrinks the clauses and makes accuracy WORSE; raising it pads them
and also loses. The shipped 3.0 is the best of the four, and **no value comes
near the healthy 3-8 band.** Combined with the settledness finding, the shape is
consistent with "NO CONSISTENT SHORT RULE EXISTS in this representation/target".
So the bottleneck is the SIGNAL - now confirmed from a THIRD independent angle
(settledness, churn trend, and clause shape). This is the measurement behind the
decision not to spend effort sweeping N or s.

=== TASK 3: AN HONEST CORRECTION TO MY OWN HYPOTHESIS ===
I predicted that random labels would produce `too long` clauses (the TM padding).
MEASURED: on this encoding noise reads as **short / `collapsed`** (mean 1.88,
median 2.0, acc 33.3%) - the TM FAILS TO COMMIT rather than padding. So "too
long" is not the noise signature, which means the shipped gun's 19.17 mean is not
explained by label noise. Worth knowing.

Adds diagnostic group 8: the clause-shape checker - full length distribution
(min/median/p10/p90/std), per-polarity and per-class breakdowns, a
`clauseShapeVerdict` against a parameterised healthy band (default 3-8),
per-BLOCK length contributions, and clause coverage (mean firing clauses,
effectiveClauses = participation ratio, top3Share). `healthLine` now appends
`shape=<mean> (<verdict>)`.
Validation: `test_tm_clause_shape` 66 checks. A planted 2-literal rule reads
`healthy` with the literals recovered exactly; per-block correctly names the
planted blocks (WALLS 43.0%, BULLETS 28.1%) and buries an irrelevant block (3.4%,
below its uniform 8.3% share); random labels read `collapsed`.
REAL READING, shipped gun: mean 19.17 / median 16.00 / p90 44.80 / max 57,
173 non-empty of 200, 27 empty => **`too long`**; coverage firing/sample 48.53
(24.3%), effectiveClauses 97.85/200, top3Share 5.3% (voting NOT concentrated);
per-block is diffuse with no dominator, EXCEPT **UNUSED 6.4%** - the always-true
negations of the never-written bits 38/39 acting as FREE PADDING, the same bug the
kit found earlier now visible as clause bloat.

Guards: test_tm_clause_shape 66 (new), test_tm_diag 48, test_tm_automata_diag 55,
diag_synthetic 17, diag_automata_validation 11, test_gun_harness 39,
test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41,
test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 40,
test_rack_membership 48, test_selector_tiebreak 19, test_tm_pattern_registration 20,
test_vbullet_admit_gate 12. acceptance_offline_vs_online not run (needs a live
battle; no tm_diag dependency).
2026-09-22 21:56:06 +02:00
SirStone e180626b50 Horizon headroom: QUALIFIED PASS at h=10..50, but the signal is one feature
Measures the learnable headroom of the "where will the enemy be in h ticks"
problem on the DrussGT fixtures, for h=1..50, using the ANGULAR error (the aim
only cares about the angle - a distance-only error cannot change the shot).
Observer = ModularBot at t; naive guess = straight-line extrapolation, never
bounced off walls. Round-bounded via the .rounds.json sidecars; the tail h ticks
of each round are dropped, never labelled with garbage. N = 32,630 (h=1) down to
31,405 (h=50).

WHAT THE NUMBERS SAY
- **The question is FAIR.** Sign balance is dead-centre at EVERY horizon
  (47.7-50.2% left) with zero systematic bias (median signed error = 0.00 deg at
  every h). No de-biasing needed - a healthy symmetric question, which is what
  the two-binary output shape needs.
- **The dumb guess is exact short, wrong long.** naiveMiss (aim lands outside the
  18px body): 0% at h<=3, 3.3% at 4, 11.8% at 5, 27.6% at 6, 48.6% at 10, 62.6%
  at 15, 73.5% at 20, 85.9% at 30, 93.7% at 50. Error magnitudes: median 2.0 deg
  (h10), 7.0 (h20), 12.8 (h30), 18.5 (h40), 24.5 (h50). Body half-angle at 300px
  is 3.43 deg - so at long horizons the error is 4-7x the body size.
- **No trivial rule solves it.** Turn-direction accuracy is 53-54% at h<=3, drops
  through 50% at h~7, and INVERTS to 38-45% at long h (flipped = 55-62%, the best
  single rule anywhere). Causal 1-step persistence peaks ~65% at h2-4 and decays
  to ~49% by h50. Majority is 50-51.5%.
- **REAL signal exists in exactly ONE feature: the enemy's current turn/reversal
  direction.** dTurn swings from +8.4pp (h2) through 0 (h7) to **-24.0pp at h50",
  all |z|>6. Every other planned feature is WEAK: walls <=2-4pp, speed <=5pp,
  bullets <=4pp, reversal <=6pp, closing <=3pp.

TWO FINDINGS I DID NOT EXPECT
1. **The dead zone is h=5..9.** The naive guess starts missing there (12-44%) but
   NO tested feature shifts the left/right split by >=5pp - the sign is a
   featureless coin flip in that band. So those horizons carry no learnable signal
   despite looking promising.
2. **The bullet block - which we designed with enthusiasm - shows <=4pp of shift.**
   The enemy's dodging reaction to our bullet is NOT a strong conditioning signal
   at these horizons in this measurement. CAVEAT: the "bullet in flight" split
   used an energy-drop proxy with ~370 false positives in 1504 positives, so this
   is a weak negative, not a settled one - the block should be measured properly
   before being cut.

RECOMMENDED RANGE: **h = 10..50** (41 horizons, contiguous). Criteria: (a) 42-58%
left, (b) max(turnAcc, pers1Acc) < 75%, (c) some feature |delta| >= 5pp with
|z| >= 3, (d) naiveMiss >= 15%. h=1-4 fail (d); h=5-9 fail (c); from h=10 all four
hold and strengthen with h.

VERDICT - QUALIFIED PASS, and the job's own calibration is worth quoting: there is
genuine, non-degenerate structure, so the horizon-input design is not obviously
wasted; BUT the per-sample signal is weak and concentrated almost entirely in one
feature which already captures most of the easy structure (~60% sign accuracy).
**"I would not treat this as a green light for a big build; I would first check
that a model can beat 60% sign accuracy on a held-out round at h~15-25."**

CAVEATS (from the tool): DrussGT-only, and the enemy's movement at capture time
was a RESPONSE to our current movement, so the headroom is conditional on how we
move now; offline observation is perfect every tick while live we see the enemy
only on radar scans, so these numbers are an UPPER BOUND; and the replay is
open-loop even though the capture was closed-loop.
2026-09-22 21:46:05 +02:00
SirStone ab8d383121 Automata metrics: settledness alone does NOT separate learning from fidgeting
Added the four automata-level metrics to the TM diagnostics kit (settledness,
clause diversity, churn, vote disagreement) plus a state histogram, a per-input
confidence table and a one-line health summary, and validated them on a
learnable-vs-noise pair.

STATE CONVENTIONS, read off OUR code rather than from memory:
  range [-nStates, nStates] as int16; nStates = 64 for tm_pattern, 32 for tsetlin
  initial value 0 = the Exclude boundary
  INCLUDE iff state > 0; EXCLUDE iff state <= 0
  flip boundary sits between state 0 and 1; commitment = abs(st)/nStates in [0,1]

=== THE GATE, AND A RESULT THAT MATTERS ===
Case A (learnable planted rule) vs Case B (shuffled labels), 49 bits, N=64:
  metric                    A (learnable)     B (shuffled)
  settledness mean              0.970            0.719
  churn flip/sample        0.000055 FALLING  0.000788 FLAT
  clause-change/sample       0.00263 falling   0.0595 flat
  diversity (Jaccard)           0.176            0.014
  disagreement                  0.003            0.298
  verdict                    settling        mixed (NOT settling)

**SETTLEDNESS ALONE DOES NOT WORK.** On noise the automata still COMMIT (0.719) -
they just commit to the wrong thing. The decisive separators are **churn TREND
(falling vs flat)** and **vote DISAGREEMENT (0.003 vs 0.298)**. Had we built only
the settledness metric - the one that seems most obvious - we would have been
misled. That is now recorded in the README.

INERTIA SWEEP: A vs B separate at N=16/32/64/128. **Raising N raises A's
commitment but does NOT reduce B's noise-fitting** - so more inertia does not
rescue a noise-fitting TM.

=== REAL READING ON THE SHIPPED GUN, AND THE INFERENCE IT SUPPORTS ===
tm_pattern GF head over the DrussGT fixtures: settledness 0.484 (settling),
diversity 0.267 (moderate), churn 0.094/100 FALLING, disagreement 0.145
(coherent). **VERDICT: SETTLING** - not fidgeting, not collapsed. Constant inputs
flagged: 38/39 (the known never-written bits) plus 19/36/37.
Context: pooled warm accuracy 35.72% vs 34.24% majority = +1.48pp.
So: **the old gun was NOT failing because of inertia or instability - it settled
properly and its settled rules still barely beat a lazy guess.** Its settledness
(0.484) is LOWER than both synthetic cases (0.97/0.72), which is the signature of
WEAK OR CONFLICTING SIGNAL rather than too much inertia.
CONCLUSION: **N and s are not the observed bottleneck. The target/representation
is.** That is exactly why the new design changes the target and the label
pipeline rather than sweeping knobs - and it means we should NOT spend effort on
an N/s sweep expecting it to fix anything.

Also adds `diag_automata_validation.nim` (Case A/B/C + inertia sweep) and
`test_tm_automata_diag.nim` (55 pure checks); `test_tm_diag` 48 and
`diag_synthetic` 17 still pass, plus all other guards. acceptance_offline_vs_online
was NOT run (it needs a live battle and there is no tm_diag dependency).

Caveat: churn on the real gun is a PROXY (a tm_core retrain over captured samples
in live order) because the live gun exposes no per-sample state trace; the other
metrics are read directly off the exported teams.
2026-09-22 21:40:27 +02:00
SirStone f9f8d84671 TM diagnostics kit: VALIDATED (finds a known dead input), and it found a real bug
Built `common_libs/tm_diag/` as a first-class offline diagnostics kit for Tsetlin
work, BEFORE writing the new gun - because we hit two data problems tonight that no
amount of reading the TM's clauses would have revealed (a 38.8% majority answer,
and 36-58% mislabelled training samples).

WHAT IT PROVIDES
- `feature_spec.nim`: a NAMED feature container, so a learned clause prints as a
  sentence (`IF near-wall AND bullet-dead-on AND turn-left(t-2) THEN class=3`)
  instead of "feature 17". Includes the 49-bit draft spec from the design session
  and the shipped 40-bit encoding.
- `tm_core.nim`: a compact deterministic Granmo multiclass TM with an
  INTROSPECTABLE clause layout (mirrors the tm_pattern core).
- `diagnostics.nim`, six groups: (1) pre-flight DATA checks + shuffled-label
  control, (2) clause introspection (readable dump, per-clause vote counts, empty
  and never-fired clauses, length distribution, per-class balance), (3)
  per-feature contribution with an explicit DEAD-INPUT LIST and a ranked
  most-valuable list, (4) accuracy vs the majority baseline with per-class
  precision/recall and pred-majority share, (5) learning curve, (6) ablation hooks
  (drop a block / scramble a bit).

=== TASK 3: THE VALIDATION THAT GATES EVERYTHING - PASSED WITH NUMBERS ===
A diagnostic we never checked is worthless, so the kit was tested on a synthetic
set with a PLANTED RULE (class2 = A and B, class1 = A and not B, class0 = not A),
a deliberately IRRELEVANT block (US, 9 bits) and a PURE-NOISE bit (17).
- majority baseline 60.63% (class0); over-30% correctly flagged
- **the planted rule is recovered EXACTLY** via `necessaryLiterals`:
    class0 IF NOT dist-wall<50 | class1 IF dist-wall<50 AND NOT lat DEAD-ON |
    class2 IF dist-wall<50 AND lat DEAD-ON
- **DEAD-INPUT LIST = all 9 US bits AND the noise bit 17**, while the planted bits
  0 and 45 are correctly NOT listed
- top contributors: bit0 w=1241.7, bit45 w=583.3, then 49.8 - a 12-25x gap, so the
  relevant bits are unmistakable
- **ABLATION: drop WALLS -39.47pp, drop BULLETS -19.33pp, drop US 0.00pp**,
  scramble A -42.00pp, scramble the noise bit 0.00pp
- shuffled-label control 60.40% vs majority 60.63% = -0.23pp -> no leak
So the kit reliably finds a known dead input and a known relevant one.

=== TASK 4: THE REAL READING, AND A BUG IN THE SHIPPED GUN ===
`tm_pattern` GF head, 6 DrussGT fixtures, pooled 250,745 samples:
- label balance c2 = **34.4%** (majority-heavy, flagged); accuracy **35.72%** vs
  majority **34.24%** -> margin **+1.48pp**. On `tr_drussgt_vs_crazy` it is BELOW
  majority (33.81% vs 37.72%, -3.92pp).
- 200 clauses: **27 empty, 45 never fired**, mean length 19.17, max 57. The
  majority class is starved (class2: 22 non-empty, 18 empty, only 2 positive
  fired). Class4 fires 11-24-literal clauses -> memorisation signature.
- **REPRESENTATION BUG FOUND (reported, not silently fixed):** `tmBuildBits` writes
  only 38 raw bits into `var bits: array[TM_NBITS=40, uint8]` - bits 38 and 39 are
  NEVER ASSIGNED, so they are always 0 and their negated literals are always 1.
  The kit's `constantInputs` confirms 38/39 are constant, and **`UNUSED-38`/`39`
  rank #6 and #8 in the most-valuable-inputs list** - i.e. the model's
  highest-usage inputs are information-free. That is a representation bug, not a
  display artefact, and it is a concrete mechanism for part of the poor learning.

DRAFT ENCODING CHECKED: the 49-bit draft is arithmetically consistent
(4+4=8 walls, 6+3=9 us, 3+5+3+3+3+3=20 motion, 5+7=12 bullets = 49). No draft
inconsistency.

Guards: test_tm_diag 48 (new), diag_synthetic 17 (new), test_gun_harness 39,
test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41,
test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 40,
test_rack_membership 48, test_selector_tiebreak 19, test_tm_pattern_registration 20,
test_vbullet_admit_gate 12, acceptance_offline_vs_online 12/12; tm_pattern_learning
passes. The tm_pattern hook is additive and default-OFF (no behaviour change).

NOT YET INCLUDED (the automata metrics discussed for the next step): per-clause
automata settledness (distance from the flip point), clause diversity (pairwise
overlap), literal-set churn over time, and cross-clause vote disagreement. The kit
has clause-level diagnostics but not the automata-state ones.
2026-09-22 21:24:27 +02:00
SirStone b0654d18eb TM verdict, settled: it loses LIVE and sits at/below its majority class - (c)
The user pushed back on "the TM can't be your best 1v1 gun", correctly, because two
decisive tests had never been run. Both are now run and they agree.

TASK 1 - THE GF HEAD vs ITS MAJORITY-CLASS BASELINE (offline, n=1,751,067):
  label histogram [254286, 284578, 678879, 297055, 236269]
  majority class = 2 (the CENTRE bucket) = 38.77%
  RAW head accuracy = 36.69%  ->  margin **-2.08 pp, BELOW majority**
  GATED head accuracy = 40.37% vs 38.75% majority -> +1.62 pp, BUT it predicts the
  majority class on 62.4% of ticks and its minority recall is 13.6% / 12.9% - a
  base-rate predictor wearing a classifier's clothes.
  Shuffled control sits at its own majority (20.04% vs 20.12%), confirming chance.
**THE OLD "46% vs 20% CHANCE" FIGURE I QUOTED WAS WRONG ON TWO COUNTS:** the
baseline is 38.8%, not 20%, and the 46% predated the deferred-label fix. Against
the correct baseline the head is BELOW it.

TASK 2 - THE FIRST-EVER LIVE A/B OF THE TM GUN (7 runs x 7 rounds per arm, one
frozen binary from git archive HEAD = eb74f9b2, sha256 cb66d66b..., real DrussGT,
every arm forced alone with TR_RACK_<GUN>=both and all 14 others off, liveness
confirmed per run):
  arm                     shots   real %   dmg/run   round wins
  onlyPattern              4610   10.74%     285      25/49
  onlyTMPATTERN (radial)   3374    3.50%      71       0/49
  onlyLinear               3218    3.23%      61       0/49
  Pattern vs TM:  +7.22 pp / +213.7 dmg, exact p=0.0006
  TM vs Linear:   +0.30 pp, p=0.659  (dmg p=0.438)
**The TM is statistically INDISTINGUISHABLE from its own Linear base live.** So it
is not "the TM works and we are aiming it wrong".

DIRECT ANSWER: **(c) It loses live AND sits at/below majority - the target carries
no learnable signal beyond the base rate, and that is the reason.** The reason is
not the machine, not the knobs, and not the application alone: the thing it was
asked to predict is dominated by the modal answer.

This closes the TM-as-gun thread. If a TM is wanted in the bot, a firing gate or a
movement decision is a better fit for a boolean-rule classifier than an aim point -
that is untested and is a different project.

A LIVE GF-MODE ARM WAS NOT RUN (stated as unmeasured): the task pinned one frozen
HEAD binary and HEAD registers the TM gun as radial only; Task 1 already makes GF
the unpromising candidate.

HARNESS FIX WORTH KEEPING: `tools/ab/which_gun_arm_env.sh` left the TARGET gun
unset, so with the now-Pattern-only default it silently fell back to the FULL rack
- an arm could appear to test a single gun while actually running the whole rack.
It now emits `TR_RACK_<GUN>=both` for the target and `=off` for all 14 others.
(Earlier which-gun results are unaffected: they ran before the Pattern-only default,
or - as in the melee/1v1 campaign - set the explicit `=both` themselves.)

tm_pattern.nim gains a per-class confusion matrix (warm samples only) to support the
majority baseline; no behaviour change. Adds Round 4 to
tm_pattern_sweep_results.md with both tasks and the interpretation rule.
2026-09-22 08:21:49 +02:00
SirStone eb74f9b2e3 Ram: finisher-only by default, and the bullet-rain abort now measures real energy
Follows the diagnosis that proactive straight-line ramming CANNOT work: both bots
have MAX_SPEED=8, so a pursuit cannot catch an evading equal-speed opponent.
Measured over 49 rounds per arm, opportunity -> contact was **0/6** (base), 0/40
(ring), 0/12 (ringhot). The only proactive conversion in the whole corpus came
from a FINISHER, and only because a <20-energy DrussGT stops fleeing (that episode
closed at 6-8 px/tick). Opportunity episodes never got below ~80px; one ran the
full 60-tick duration cap and closed only 198->171px; a perfectly aligned
full-speed one closed 195->114px then plateaued.

CHANGES
- **Finisher-only default.** `finisher` (<20 energy, dist<300, we are healthier)
  and the rare `desperation` (both <5, dist<150) are kept; `opportunity` and the
  speculative `plan` are OFF. Both are env-reenableable with no rebuild:
  `TR_RAM_OPPORTUNITY=1` (tune via TR_RAM_OPP_DIST/MARGIN) and `TR_RAM_PLAN=1`.
  Justification: it removes 100+ non-converting episodes per fixture at zero
  measured loss (oldram vs base was p=0.69, damage 279 vs 284, survival 17/49 vs
  16/49) - and each of those episodes spent up to 60 ticks driving STRAIGHT at
  the enemy, abandoning the mover's dodging and disrupting aim.
- **`desperation` KEPT** deliberately: it is cheap and rare, fires only when both
  bots are nearly dead at short range (a coin-flip where 0.6 contact can decide
  it), and it is not the refuted straight-line pursuit.
- **THE BULLET-RAIN ABORT WAS DEAD CODE AND IS NOW FIXED.** `onHitByBullet`
  accumulated raw bullet FIREPOWER while `TR_RAM_ABORT_DMG = 0.5` was documented
  as a DAMAGE rate - so the bar was implicitly "sum of power > 7.5 over 15 turns"
  and the maximum rate ever observed was 0.27. It now accumulates REAL ENERGY via
  a `bulletDamage(power)` helper matching the server's `4p` / `6p-2` formula, and
  `TR_RAM_ABORT_DMG` defaults to **2.0 energy/turn** (~30 HP over 15 turns):
  "abort an in-progress ram if we take > 2.0 energy per turn". Same effective bar
  for normal firepower, and it can now actually fire - the live run reports
  `dmgRate=1.07/turn` where the old units said 0.27.
- **`ramStuckTicks` REMOVED.** It required `dist < 5px`; contact occurs at ~36px
  (two 18px radii) and position rewind prevents getting closer, so it could never
  increment. Only the 60-tick duration cap can now self-end a ram.

LIVENESS (measured, default config, vs a charging Java RamFire, 3 rounds):
  default              -> `[ram] ON reason=finisher` x3, `reason=opportunity` x0
  TR_RAM_OPPORTUNITY=1 -> `reason=opportunity` x4, `reason=finisher` x2
So the opportunity states DID occur and are suppressed by the new default - the
removal is real, not an arm that never fires. A line also read
`[ram] OFF reason=duration dmgRate=1.07/turn`, confirming the new energy units.

Adds docs/ramming_negative_result.md (70 lines) recording the question, the five
diagnostic answers, the geometric reason, the finisher exception, the two dead
code paths, and an explicit "do not re-attempt a proactive straight-line ram; if
point-blank forcing is ever wanted it is an INTERCEPTION/cornering movement
problem" note - the same pattern that stopped the corpse bug recurring.

Guards: test_ram_decision 40 (was 28), test_gun_harness 39, test_vbullet_metric 11,
test_power_selection 3, test_adaptive_radar 41, test_tfil_ring_weights 24,
test_power_policy 26, test_rack_membership 48, test_selector_tiebreak 19,
test_tm_pattern_registration 20, test_vbullet_admit_gate 12,
acceptance_offline_vs_online 12/12. ModularBot compiles.

Honest note: the abort-threshold fix is a real (tiny) behaviour change, NOT
measured-neutral - it only bites while a finisher ram is under sustained fire,
which is exactly the user's stated wish. The finisher-only removal itself is
measured-neutral per the given A/B.
2026-09-22 08:14:03 +02:00
SirStone 99f9532f55 Melee vs 1v1 racks: mechanism built, but NO gun-level payoff - do not split
The user's plan was separate melee and 1v1 racks. The mechanism is built and
committed (TR_RACK_<GUN>=both|1v1|melee|off, mode from server truth). This job
produced the missing evidence: each candidate gun forced ALONE in BOTH modes,
one frozen binary, env-only arms, exact two-sided permutation tests.

1v1 vs the real DrussGT (6 arms x 5 runs x 7 rounds):
  Pattern      10.80%  292 dmg/run   11/35 round wins
  KNN           5.58%  128            1/35
  Linear        2.99%   59            0/35
  Circular      2.89%   60            0/35
  WallBounce    2.72%   57            0/35
  GuessFactor   2.14%   38            0/35
  -> every arm differs from Pattern at p=0.0079 (the 5v5 floor). Per-gun damage
     span 7.7x. In 1v1 the gun matters ENORMOUSLY.

Melee (+DrussGT, RandomMover, WaveSurfer, OscillatorBot; enemyCount 4):
  WallBounce   25.79%  643 dmg/run    7/25
  Linear       24.88%  624            7/25
  GuessFactor  24.48%  613            4/25
  Pattern      24.39%  599            5/25
  Circular     25.88%  596            5/25
  KNN          23.87%  539            3/25
  -> NO arm beats Pattern with significance (p=0.42-0.96, fully overlapping).
     WallBounce's nominal +7.3% is p=0.42. Per-gun damage span only 1.19x.

THE FINDING: in 1v1 the gun matters enormously (7.7x spread); in melee it barely
matters (1.19x). Melee is won by movement, survival and placement, not by which
gun you carry - every candidate lands in the same ~24-26% band. So a separate
melee rack has NO gun-level payoff and DefaultRackMembership stays Pattern-only.
The mechanism remains available if it is ever wanted.

Honest caveats: the melee comparison is UNDER-POWERED at 5 runs (detecting the
~44-damage WallBounce-Pattern gap would need ~2.5-3x the runs), so "no significant
difference" is NOT "no difference"; the only candidate worth re-testing is
WallBounce in melee, and it must not be shipped on this evidence.

SHIM LIMITATION FOUND, worth recording: run_bridge_battle.sh captures with
`--subject DrussGT`, and the shim DROPS every tick once the subject dies - losing
5-20% of ModularBot's shots in melee. The campaign was re-run with
`--subject ModularBot` so ModularBot's events are complete (fires match gun_stats
realShots exactly). Any future melee evidence through this shim must do the same
or it will silently under-count.

Adds docs/melee_vs_1v1_racks.md. No repository source changed.
2026-09-22 08:08:43 +02:00
SirStone bfdcdf8919 The three unmeasured features, A/B'd - and a methodology correction I had wrong
Five arms x 7 runs x 7 rounds (35 real-DrusGT battles, 8 concurrent), one frozen
binary from git archive HEAD at 185a32e (includes the shipped Pattern-only rack),
env knobs only, server-side event sidecar, exact two-sided permutation tests.

  arm                   real %   dmg/run   survival(rounds won)   p vs base
  base (shipped)         10.61    284.1     16/49 (32.7%)          --
  nopower (policy off)    7.88    247.1      9/49 (18.4%)          0.0012
  oldram (old gates)     10.78    279.4     17/49 (34.7%)          0.6888
  ring (tfil_ring)       20.28    267.4      6/49 (12.2%)          0.0006
  ringhot (orig heat)    11.73    285.3     15/49 (30.6%)          0.0303
Every round ends with exactly one death (0 timeouts), so survival = round win.

1. POWER POLICY HELPS - KEEP. Turning it off drops real hit rate 10.61 -> 7.88
   (p=0.0012), damage/run 284 -> 247, and wins FEWER rounds (16 -> 9). The cap
   trades per-shot damage for many more shots and a higher per-shot rate; that
   trade is a clear win. This was shipped on unit tests alone until now.

2. PROACTIVE RAMMING - INDISTINGUISHABLE. oldram vs base p=0.69, damage 279 vs
   284, survival 17/49 vs 16/49 (p=1.0), ram contacts 1 vs 2. The lever IS live
   (6 `opportunity` ON events vs 0 under the old 50px/+30 gates; base reached
   <40px on 10 ticks vs 0) but converts to essentially no extra collisions and no
   measurable outcome change. Safe to keep, but it is not earning its keep and
   reverting it is equally defensible.

3. RING MOVER HURTS THE OBJECTIVE - DO NOT SHIP. And this is the important one.

*** METHODOLOGY CORRECTION - I HAD THIS WRONG ALL NIGHT ***
The ring arm has the BEST hit rate of anything measured tonight: 20.28% vs 10.61%
(+9.67pp, p=0.0006, non-overlapping). Read alone it says "ship it immediately".
It is a CONFOUND. The ring halves engagement range (median 460 -> 240px), which
halves round length (1542 -> 636 ticks) and shots (4466 -> 1834). So damage/run is
FLAT (284 -> 267) while survival/round-win MORE THAN HALVES (16/49 -> 6/49,
p=0.0122). It is a GLASS CANNON: same damage dealt, twice as many deaths. The
hit-rate gain is a geometric artefact of fighting closer, not an improvement.
** For MOVEMENT arms, hit rate alone INVERTS the verdict. ** A movement change
alters range, shots fired and round length simultaneously, so the objective
metrics are DAMAGE/RUN and ROUND-WIN RATE (survival) - report all three. For
gun/selection arms, where range and round length are held fixed, real hit rate
remains the right ground truth. I had been enforcing the hit-rate-only rule
without qualification; it is now qualified.

HEAT TAMING is the knob that moves the tradeoff: tamed (corridor 5 / wall 10)
lets the range weighting pull to ~240px (20.28% / 6 wins); original (20/30) keeps
it at ~400px (11.73% / 15 wins). So heat taming buys hit rate at the cost of
survival - a knob to keep conservative, and the reason the ring is not the default.

Nothing to revert: the shipped movement is already `tfil` and the ring is opt-in.
Caveats recorded: only adversary is the real DrussGT jar (the shim hosts only
jk.mega.DrussGT), so the ring-vs-winning result needs re-checking elsewhere; and
base (10.61%) is consistent with the committed onlyPattern result (9.99%),
validating the frozen binary and pipeline.

Adds docs/feature_ab_results.md. No source files changed by this job.
2026-09-22 02:34:05 +02:00
SirStone 3142b70aa5 Gate virtual-bullet spawn on rack admission: +68% tick rate, selected gun unchanged
The default rack is now Pattern-only (31c7c01), but membership filters SELECTION,
not SPAWNING - so all 13 unselected guns still ran `predict` + `spawnBullets`
every tick to feed fitness tables nobody reads. Measured waste: Tsetlin alone
0.98 ms/tick, KNN 0.30, plus 10 more. This generalises the gate TMPATTERN already
had to every gun, behind `TR_VBULLET_ADMIT_ONLY` (default 1 = gate, 0 = old).

OFFLINE COST (release build, 400 ticks, min of 2 reps, 13-gun rack):
  OFF  2.14 ms/tick  (implied 467 ticks/s)
  ON   0.04 ms/tick  (implied 27667 ticks/s)
  -> reclaimed 2.10 ms/tick, ~98% of the virtual-bullet cost. Tsetlin's 0.98
     disappears, KNN's 0.30 disappears, only Pattern (0.017) survives.
Note the absolute scale is lower than an earlier unoptimised measurement (~6.1
ms/t) because this is a -d:release build; the ON-vs-OFF DELTA is the robust result.

LIVE TICK RATE (ONE frozen binary, 4 runs/arm x 4 rounds, all 8 concurrent, gate
varied by env only):
  ON   146.2 ticks/s  (142.3, 145.2, 147.9, 149.4)
  OFF   86.9 ticks/s  ( 95.0,  41.4,  99.1, 112.1)
  NO OVERLAP: ON min 142.3 > OFF max 112.1. Excluding a game-outcome outlier in
  the OFF arm, OFF max is still 112.1. Outside the noise.
**+68% tick rate**, which also makes every future A/B faster. Both arms still pay
the fixed 8192-slot ring scan in tickBullets.

SAFETY, verified not assumed: Pattern is admitted under the shipped rack, so its
own fitness keeps accumulating and the SELECTED gun is unchanged - Pattern 100% in
all 4 runs both arms, and Pattern was the ONLY gun with vShots>0 in the ON arm
while all 13 had vShots>0 in the OFF arm. So admission is the correct predicate.
Mid-round transitions are safe by construction: only predict/spawn are gated, while
tickBullets still resolves every active bullet and the feedback case still calls
the owning gun's onResult.

A TOOLING BUG THIS CAUGHT, and a correction to the task's assumption:
`acceptance_offline_vs_online.nim` IS affected (I had assumed it was not). It
compares the live per-gun vShots against an offline replay that always spawns all
guns, so under the default gate the online non-Pattern vShots are 0 while offline
is ~400 - a guaranteed mismatch. Fixed by pinning `TR_VBULLET_ADMIT_ONLY=0` inside
that test (same putEnv/defer pattern as TR_RECORD_WORLDSTATE), keeping 12/12. The
test is about offline/online METRIC parity, so it needs every gun spawning.
Unaffected (verified from source): audit_virtual_guns.nim,
measure_cornering_guns.nim, sweep_tm_pattern.nim - all offline, none read
GUN_STATS_PATH.

Guards: test_vbullet_admit_gate 12 (new, pure), test_gun_harness 39,
test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41,
test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 28,
test_rack_membership 48, test_selector_tiebreak 19, test_tm_pattern_registration 20,
test_tm_pattern_learning 3, acceptance_offline_vs_online 12/12. ModularBot compiles.
2026-09-22 02:24:18 +02:00