Commit Graph

14 Commits

Author SHA1 Message Date
SirStone 18f778056b gun selector: hysteresis measured NEGATIVE, shipped at the lightest setting
Hypothesis under test: the selector chatters (~54 switches/100 ticks) and that
chatter suppresses firing, so committing to the virtual-best gun should raise
real hit rate. MEASURED AGAINST THE REAL DRUSSGT: it does not.

  setting              switches/100t   real hit %   dmg/run   shots/run
  no hysteresis 0/0         54.29        7.02%        217       244.8
  light 10/0.05              1.95        6.22%        191       235.3
  moderate 30/0.15           1.33        5.10%        156       233.9
  aggressive 60/0.30           -         5.72%        175       242.5
(16 runs x 7 rounds per config except aggressive = 8; server-side events
sidecar; permutation test baseline-vs-moderate p=0.002, baseline-vs-light
p=0.18.)

Hysteresis cuts chatter 28-54x but every variant fires slightly FEWER shots and
deals LESS damage than baseline. Mechanism [INFERRED, consistent with
docs/gun_rack_analysis.md 2/4]: the per-tick random tie-break among the tied
band is a hedge, and hysteresis destroys it by committing to the virtual-best
gun - which is not the real-best, because the virtual metric is a weak,
sign-unstable ranker. The chattering was load-bearing.

Shipped: GunDwellTicks=10, GunSwitchMargin=0.05 (GUN_SELECTOR_DWELL /
GUN_SELECTOR_MARGIN) - the only setting within the baseline's run-to-run spread.
GUN_SELECTOR_DWELL=0 GUN_SELECTOR_MARGIN=0 reproduces the pre-change selector
exactly.

Seam: VirtualTracker, which already owns the other selection state (fitness,
the relative floor's peakRateRef), so the bot needs no new fields. bestGun and
chooseFromFit stay pure/memoryless, which is why the existing random-tiebreak
test needed no change.

Guards: test_gun_harness 39/39 (33 original + 6 new hysteresis checks),
test_vbullet_metric, test_power_selection, acceptance_offline_vs_online 12/12.
2026-09-21 22:41:01 +02:00
SirStone 2c94dc221a test(selector): 16 ranking rules A/B'd against the boss - none beat the shipped config
Added runtime-tunable ranking knobs to the selector, all defaulting to the
shipped values so behaviour is byte-identical when unset: GUN_SELECTOR_WINDOW,
MINOBS, TIE, FLOOR, POOL, RANK, SHRINK, SEED. rankScore supports mean, Wilson
lower bound, UCB, Thompson and shrinkage. Also fixed hitRate's most-recent-N
read for sub-WindowSize windows (windowHits).

RESULT: NO candidate credibly beat the shipped config. 13 runs x 8 rounds vs
DrussGT, 3612 shots, base 6.95% at 251 dmg/run; every candidate's per-run
interval overlaps base, and the nominal 'winners' are <=0.6 SE apart on far
fewer shots. Kept the shipped default. Valid outcome, recorded plainly.

THE FINDING THAT MATTERS MORE: the virtual-bullet ranking is ANTI-correlated
with real hit rate - Spearman ~ -0.37 for the shipped config. It is not merely
weak, it is INVERTED. The guns with the highest VIRTUAL rates have among the
lowest REAL rates: Tsetlin 12.9% virtual / 5.8% real, WallBounce 12.9 / 6.2,
StopShot 12.6 / 6.1, AvgLead 12.3 / 7.0 - while Linear sits at 10.2 virtual /
10.7 real and KNN at 7.5 / 9.0. So what carries the selector is the floor/tie
HEDGING, not the ranking: removing the floor drops us to 5.08% / 175 dmg.
That also kills the 'exploration' hypothesis - every gun spawns virtual bullets
every tick, so sampling is uniform and the bottleneck is SIGNAL QUALITY, not
under-sampling.

FINAL PER-GUN REAL HIT RATE vs DrussGT (13 runs, 3612 shots, overall 6.95%):
  Linear 10.7 | Circular 9.9 | KNN 9.0 | Pattern 8.6 | Accel 7.3 | AvgLead 7.0
  GuessFactor 6.9 | DecayGF 6.4 | WallBounce 6.2 | StopShot 6.1 | Tsetlin 5.8
  Displace 5.3 | HeadOn 5.2
Keep: Linear, Circular, KNN, Pattern, Accel, AvgLead. Marginal: GuessFactor,
DecayGF, WallBounce, StopShot. Below overall: Tsetlin, Displace, HeadOn - but
HeadOn must STAY as the floor fallback, since disabling the floor measurably
hurt.

CORRECTION TO A CLAIM I HAVE BEEN MAKING: the 12/12 offline==online acceptance
is FLAKY. It fails 11/12 on the UNMODIFIED HEAD source (control: KNN 81 online
vs 71 offline), and the mismatching gun moves between runs (KNN, then
WallBounce) - a live/offline boundary race. So '12/12' was a lucky run, and
that proof should be treated as strong-but-not-exact until the race is fixed.
This diff does not touch replayFixture/spawnBullets/tickBullets and the
selector is never called during replay, so it is pre-existing.

SIDE FINDING, not fixed: the shipped live bot never calls randomize(), so the
'random tie-break' is a FIXED sequence across process restarts.

Overfitting guard vs a non-surfer (SpinBot): inconclusive - ModularBot fires
only 17-31 real shots/run against fast bots because the range-aware firing gate
is strict at long range, so the guard has little power. Wilson looked better
(18.5% vs 8.6%) but on 70-92 shots with a 5-33% spread. Not evidence either way.
2026-09-21 06:31:00 +02:00
SirStone 57b2ac3849 feat(guns): scale-aware power selection (+52% damage); TM classifier gun built, measured, DISABLED
TASK 2 - power selection, a clear win. bestPower used an ABSOLUTE
MinHitRate = 0.40 bar. Measured per-bin virtual rates (rolling-100 fraction)
show no bin ever clears 40%, so 11 of 14 guns were stuck at bin 0 (power 1.0)
even where higher bins were comparable:
  Linear  p1.0 44% p1.5 39% p2.0 30% p3.0 29%   old bin 0 -> new bin 3
  Accel   p1.0 44% p1.5 40% p2.0 26% p3.0 29%   old bin 1 -> new bin 3
  Pattern p1.0 50% p1.5 40% p2.0 27% p3.0 12%   old bin 1 -> new bin 2
Replaced with a scale-aware PowerBarFrac = 0.50 (a dimensionless FRACTION of
the gun's own best bin rate). 13 of 14 selections now pick heavier bullets.
Real effect vs DrussGT (8 rounds x 3 runs): hit rate unchanged (7.56% ->
7.47%) but damage dealt +52% (157 -> 239 per run) and rounds end faster.
Same accuracy, half the shots, half again more damage.

TASK 1 - the TM pattern-classifier gun does NOT earn its slot. It was built as
a mixture of experts with a corrected-Granmo TM as a multi-class gate over
HeadOn/Linear/Circular/WallBounce/Accel, labelled by which expert's prediction
was closest to the actual enemy position (an exact, supervised, per-shot
label - no delayed credit). Offline it loses to the best of its OWN experts on
essentially every fixture, and against DrussGT it cost real performance:
  baseline (path+relative)  7.56% real hit rate, damage 157
  + power fix               7.47%,                 damage 239
  + power fix + TM gun      5.59%,                 damage 133
The gun was selected on 806 ticks and fired 24 real shots at 4.2%.
So the tree ships with EnableTmSelector = false: code and wiring kept intact
for re-enabling, but it is not in the active rack.

Worth recording from the clause dump: the gate DOES latch onto meaningful
structure. On energy-threshold-turner, HeadOn's clauses key on the energy bits
(the rule's own driving variable) while Circular keys on distance/velocity. So
the TM is learning something real and interpretable - it simply cannot beat
'always pick the best expert'. Root cause (INFERRED): the closest-expert label
is noisy because several experts are near-tied, and under the path metric the
winner varies by power bin while the gate sees one shared per-tick input, so a
one-vs-rest gate over a saturated 870-bit clause space has no margin to exploit.
(Zero-padding the 2-frame window was tried first and saturated every clause at
256-755 included literals; alternating the two real frames fixed that.)

Also factors the corrected feedback into an exported tmLearnDir and exports the
encoding/TM primitives; the Tsetlin tests still reproduce the documented
mean=13.8 included literals, so the refactor is behaviour-preserving.

Verified: 33/33 guard checks, tsetlin tests green, metric checks green, new
power-selection guard green (13/14 selections change; relative bar still picks
bin 1 and not bin 3 for a [30,25,12,5]% profile), 12/12 offline==online
acceptance under the shipped default.
2026-09-21 05:19:07 +02:00
SirStone dea4dcb574 feat(gun_harness): scale-aware selector thresholds; default = path + relative
The selection thresholds were calibrated for a rate scale that does not exist.
MEASURED on an exact offline replay of a fogged live WorldState vs DrussGT
(1397 selection ticks), the 0.10 absolute floor fires on 53.0% of point-metric
ticks and forces HeadOn, which has a REAL hit rate of 2.0-4.4% - worst or
near-worst of 13 guns. HeadOn's selection share: 69.1% (abs+point) -> 43.5%
(rel+point). My earlier claim that the floor fires ALWAYS is REFUTED - it is
53%, because bestRate is a max over gun x bin and a >=50-sample bin
occasionally clears 10%. The mechanism is confirmed; the literal statement was
not.

Scale-aware mode (GUN_SELECTOR_MODE, absolute|relative, default relative):
  RelTieMargin   = 0.20  dimensionless FRACTION of bestRate, replacing the
                         fixed 2pp band so the band scales with the metric
  FloorPeakFrac  = 0.25  the floor fires iff bestRate < 0.25 * peakRateRef,
  SelectorWindow = 256   where peakRateRef is the field-best rate over the last
                         256 selection ticks - keeping the original 'don't trust
                         a collapsed field' purpose but only when the field is
                         bad RELATIVE TO ITS OWN RECENT BEST, and counting only
                         guns with >= MinObsBeforeCompete samples so cold-start
                         100% spikes cannot pin HeadOn
  also pools the rate over power bins instead of taking the max over bins, so
  one lucky bin no longer wins
absolute mode is preserved byte-for-byte for rollback.

A/B vs DrussGT, real server hit rate, 3 runs x 10 rounds per config, one frozen
binary:
  absolute+point  3.66 / 2.45 / 5.01   pooled 3.76%
  absolute+path   7.55 / 8.21 / 6.83   pooled 7.57%
  relative+point  7.66 / 6.18 / 5.79   pooled 6.59%
  relative+path   7.15 / 7.55 / 6.90   pooled 7.21%
absolute+point is SEPARATED from all three (p < 0.0001); the other three
OVERLAP each other (p = 0.18-0.64). So the METRIC is the dominant lever and
under path the two threshold models are statistically tied.

DEFAULT SET: metric = path, thresholds = relative. absolute+path was nominally
0.35pp higher but indistinguishable (p = 0.64); relative is the principled
scale-aware fix, is the only model that works under BOTH metrics, and prevents
the point-metric catastrophe if anyone switches back. Shipping absolute would
ship the accidental side-effect this work exists to remove.

STILL NOT SOLVED: the selector remains only a moderate ranker.
Spearman(virtual rank, real rank) is 0.52 for the winning config, 0.36 pooled
for path and 0.04 for point - and it is INCONSISTENT across run sets. The
metric switch won by de-selecting HeadOn, not by ranking guns better. That is
the next problem.

TASK B, report only: do NOT drive selection from raw real hit rates yet.
Only the selected gun fires, so unselected guns get near-zero real shots
(GuessFactor 20, Linear 24 vs HeadOn 733); noise is fatal (n=470 at p=10% gives
+/-2.8pp, most guns n<200 gives +/-5pp+ across a 3-15% spread); and real rate is
conditional on when the gun was selected. A blended signal with forced
exploration and shrinkage is defensible in principle but needs thousands of
shots per gun across many battles. Real rate is best used OFFLINE as the
evaluation metric - which is exactly what this A/B did.

RELATED BUG FLAGGED, not fixed: MinHitRate = 0.40 in bestPower is on the same
wrong scale - no bin ever clears 40%, so once every bin has data, power
selection falls back to bin 0 (power 1.0) late in a round.

Verified: 33/33 guard checks, 11/11 metric checks, tsetlin green, 12/12
offline==online acceptance under the shipped default, run_range rc=0 over 20
fixtures. Adds analyze_selector.nim to measure floor/tie/bestRate/HeadOn-share
per config on any fixture.
2026-09-21 04:33:52 +02:00
SirStone 3b5d70b7c3 feat(gun_harness): runtime metric switch + A/B proving the point metric mis-selects
Adds GUN_VBULLET_METRIC (point|path, default point = unchanged behaviour) so
the virtual-bullet hit model can be selected at runtime with no rebuild. Both
the live tracker and the offline replay read the same value, so the 12/12
offline==online acceptance holds under EITHER setting (verified for both).

A/B AGAINST THE LIVE BOSS, real server-side hit rate as ground truth, 5
battles x 12 rounds per metric on one frozen binary:
  point  4660 shots / 219 hits = 4.70%   (per-run 3.16-5.53)
  path   4834 shots / 359 hits = 7.43%   (per-run 6.55-8.24)
The distributions DO NOT OVERLAP: path's worst run beats point's best run.
+2.73pp, +58% relative, z = 5.56, p < 0.0001. Range distributions were
identical (~460-478 px), so this is not a range confound.

MECHANISM - and this is the important part. The gain is SELECTION, not better
gun learning. Under the point model every gun's virtual rate is compressed
into 0.6-4.4%, so HeadOn sits inside the 2pp tie margin and takes 72.6% of
selection ticks / 76.9% of shots - while HeadOn is 11th of 13 by REAL hit rate
(2.3%). The path model widens the band to 4.7-13.7% and ranks HeadOn 10th, so
its shot share falls to 35.9% and Pattern/Accel/WallBounce get picked instead.
Counterfactual: applying the point model's per-gun real rates to the path
model's shot mix yields 7.65%, i.e. essentially the whole observed gain.
So the selector, not the guns, is where the win lives.

PER-GUN REAL HIT RATE vs DrussGT (path mix, the answer to 'which guns are
worth keeping'): WallBounce 10.8, Pattern 10.5, Accel 10.0, Displace 9.3,
Circular 9.2, AvgLead 8.5, KNN 5.7, StopShot 5.2, GuessFactor 3.7,
Tsetlin 2.9. Per-gun N is small (hundreds of shots) so single-gun ordering is
indicative, not definitive.

TWO CAVEATS, recorded because they undercut a naive reading:
1. One adversary. DrussGT is a wave surfer and HeadOn is genuinely bad against
   surfers, so part of this may be matchup-specific.
2. The path model is NOT a better general ranker. Spearman(virtual rank, real
   rank) is 0.52 under point vs -0.04 under path. It wins by accidentally
   fixing HeadOn's mis-rank, not by ranking guns better. A more durable fix is
   to address the selection logic directly - which is the next job.

Also adds a focused guard test (test_vbullet_metric) covering parsing/default,
a receding-target point-miss/path-hit, a perpendicular-target path-miss, and
replay determinism.

Verified: 33 guard checks, 12/12 acceptance under both metrics, tsetlin tests
green, range 34.3% (point, unchanged) / 50.8% (path).
2026-09-21 03:58:27 +02:00
SirStone e53690036b fix(guns): speed-sensitive caches, dead stop-shot branch, exact TM trace pairing
Four guns cached a whole prediction per tick while predict() is called once
per power bin, so every bin after the first (and the real fired shot, which
shares lastState) reused the power-1.0 lead. Fixed by caching only the
speed-INDEPENDENT derived state and recomputing the lead per requested speed:
- stop_shot: also fixes prevSpeed being written before it was read, which
  made abs(speed) < abs(prev) permanently false and the entire
  stop-prediction branch unreachable (it was just Linear).
- displacement: the cache key included bulletSpeed, so the guard missed on
  all four bins and the 15-tick window advanced ~4x/tick, making the
  inferred velocity ~4x too small.
- averaged_lead: tick cache removed outright. pattern_matcher: split into
  speed-independent match+path and per-call lead.

FeedbackEvent gains fireTick/powerBin (additive; only virtual_bullets
constructs one) so guns can pair feedback to the exact shot instead of
guessing by coordinates. tsetlin uses it: traces are now keyed exactly by
(fireTick, powerBin) with a 1024-slot ring, and the 10-frame window shifts
at most once per tick (it was shifting ~4-5x/tick, so isWarmedUp tripped
after ~2 ticks).

KNOWN INCOMPLETE: tsetlin still does not diverge from Linear in battle. The
two named bugs are fixed (a 600-tick sim shows trainedShots=2141,
traceMisses=0, and a fixed-input probe converges to a 9.6px correction), but
the TM's clause feedback itself is broken: ~131 of 1740 literals end up
included per clause, so its conjunction never fires. Sweeping TM_S,
TM_N_CLAUSES and a two-branch Type-I update did not change the correction
from 0. Needs a real TM fix or removal, not another bug fix.

First-ever guard tests for the gun selector: common_libs/tests/
test_gun_harness.nim (14 checks, headless, no Java). There were none before,
which is how six broken guns survived a full analysis cycle. Against the
previous HEAD, 5 of these checks FAIL - that is the regression guard.
2026-09-20 22:47:26 +02:00
SirStone 0cc682152d fix(guns): per-bin wave queues unbreak GF/DecayGF/KNN learning; fix vbullet drops
Wave queues (guess_factor, decay_gf, knn_gun): predict() stored ONE wave
per tick while onResult() popped one per resolved bullet (~4/tick), so the
queue drained to empty within a few dozen ticks, ~3 of every 4 resolutions
returned without learning, and the survivor paired with a same-tick wave
(bearingDelta ~= 0) pinning the histogram at centre. PROOF: GF.vHits ==
HeadOn.vHits and DecayGF.vHits == HeadOn.vHits byte-for-byte in every one
of 50 rounds — the guns had degenerated to HeadOn.

Now each gun keeps a per-bin FIFO with an O(1) head cursor. At most one
push per (tick, bin) so the fire site's 5th predict() call is a no-op, and
onResult pops the oldest wave of its OWN bin via e.bulletPower. Aiming
math untouched (it was already correct: 0 deg = East, CCW+).

maxBullets 2048 -> 8192: the rack spawns 52 bullets/tick so the ring wrapped
every ~39 ticks while a long power-3 shot needs ~90, silently discarding
unresolved bullets and biasing every measured hit rate by range. Added a
droppedBullets counter so a future overflow is measurable, and wavePushes/
waveStarved counters on the three guns. After the fix: vDropped = 0 and
vStarved = 0 across all 48 recorded rounds.

fitnessFor is now exported, deterministic (enemies iterated in ascending id
order) and shared by the selector and the stats dump, replacing a hand-rolled
merge in ModularBot that never advanced its window head.

Round lines gain additive keys: vDropped, vStarved.
2026-09-20 22:27:52 +02:00
SirStone 26b66cbb24 feat(gun_harness): per-gun REAL hit attribution + bestPower cold-start fix
Attribution is proven, not guessed: the server assigns a per-round-unique
bulletId (GunEngine.nextBulletId) and stamps the same id on BulletFired,
BulletHitBot, BulletHitWall and BulletHitBullet. Keep a FIFO of fired gun
ids, stamp bulletId -> gunId on onBulletFired, resolve through that map.

Hits are deferred when onBulletHit precedes onBulletFired in the same
turn (client dispatches priority 70 > 60), which recovered 14
unattributed hits. 99.9% of shots and 99.8% of hits attributed.

Stats lines now carry per-gun realShots/realHits/realHitRate; the old
keys and round-level totals are unchanged.

bestPower: a gun with zero observations in every bin previously returned
the HIGHEST bin (power 3.0) because an empty bin satisfied the
'count == 0' clause on the first countdown iteration. Cold guns now
return the lowest bin as the docstring always claimed. Warm-gun path
untouched.
2026-09-20 22:15:42 +02:00
SirStone 343e631633 fix(gun_harness): random tiebreak + drop AntiSurfer + raise MinObsBeforeCompete
- bestGun: replace first-index-wins argmax with random pick among guns
  within TieMargin (2%) of best rate. HeadOn at index 0 was silently
  winning every tie, starving Tsetlin/Linear/etc.
- MinObsBeforeCompete 15 -> 50 (Pattern entered competition on noise)
- add MinHitRateFloor 0.10: if no gun clears it, fall back to HeadOn
  instead of selecting the best of a bad field
- ModularBot: remove AntiSurfer gun (0% virtual hit rate everywhere),
  14 -> 13 guns, renumber ids and selection counters
2026-09-20 22:07:49 +02:00
SirStone ab473c6b68 feat(ModularBot): melee targeting — multi-enemy tracker, per-enemy gun fitness, radar auto-switch
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-09-20 12:25:10 +02:00
SirStone 5bbc8cbda7 fix(gun_harness): min 15 obs before competing, window 50→100, KNN min k=5
Gun selector now gates competition: guns with <15 total observations across
all power bins sit out until at least one gun qualifies. Falls back to ungated
selection if no gun reaches threshold, preventing cold-start stalls.

Sliding window increased from 50 to 100 ticks to reduce switching noise.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-09-20 11:59:00 +02:00
SirStone 3a1149359d fix(gun_harness): bump virtual bullet ring buffer 512→2048
With 12 guns × 4 power bins = 48 bullets/tick, 512 slots overflow
before slow bullets resolve (~36 ticks travel). 2048 gives headroom.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-09-20 01:49:21 +02:00
SirStone 1ed7797cb6 feat(ModularBot): 6 guns, pattern matcher, melee modules, adversarial bots
- New guns: guess-factor (GF histogram), pattern-matcher (movement tape replay)
- New modules: minimum-risk melee movement, spinning melee radar
- New test bots: PatternMover, RandomMover, WaveSurfer
- Fixed: FeedbackEvent now carries actualX/actualY for proper GF learning
- Fixed: TM gun warmup gating + directional residuals
- Fixed: circular gun integrated formula + multi-bin omega cache
- Fixed: oscillator wall-bounce lockout
- Fixed: phantom meteor perpendicular body orientation
- 6/6 battle wins across all enemy types
2026-09-20 00:59:53 +02:00
SirStone 254c7dc997 feat(ModularBot): pluggable bot with 4 guns, phantom meteor movement, radar harness
- Gun harness: virtual bullet tracker, rolling fitness, auto-selector
- Guns: head-on, linear (extrapolation), circular (integrated formula), tsetlin machine (learning)
- Movement: phantom meteor gravity engine (danger histograms, phantom bullets, fire detection)
- Radar: harness + radar_lock adapter
- Color-coded modules: turret/bullet color per gun, body per movement, scan per radar
- Beats Target, SpinBot, Crazy, TrackFire in 10-round battles
2026-09-20 00:37:10 +02:00