Commit Graph

407 Commits

Author SHA1 Message Date
SirStone 467e07a6a4 gun ledger: fix the arm count in the final answer (six configurations, not seven) 2026-09-26 03:55:32 +02:00
SirStone 2a5ea5c17b gun j121 Batch 2 outcome: the wider panel kills the damage hint (bitbrain wash, tmhorizon WORSE); onlyPattern confirmed; phase closed 2026-09-26 03:54:52 +02:00
SirStone c343c00aaf gun j121 Batch 2: correct the SpinBot path in the wider panel before any battle (it is not under /tmp/tr_bots) 2026-09-26 03:35:11 +02:00
SirStone 3fd6db97e7 movement ship: gate v2 fresh-data primary passed (sign-flip p=0.045, CI [+0.02,+0.58]); default TR_MOVEMENT flipped to strafe 2026-09-26 03:19:33 +02:00
SirStone 087e26955f gun j121 Batch 1 outcome: onlyPattern CONFIRMED across the 15-opponent panel; pre-register Batch 2 wider-panel calibration 2026-09-26 02:58:41 +02:00
SirStone 514674886d movement gate v2: pre-register the sign-flip primary test on fresh 300-battle data (before any battle) 2026-09-26 02:44:22 +02:00
SirStone feefc1912e movement final: confirmation result and ship decision (gate failed on the sign-test leg; default NOT flipped) 2026-09-26 02:41:08 +02:00
SirStone 1d8143a15e gun j121 Batch 1: pre-register the 6-arm Pattern-panel tournament and the decision rules 2026-09-26 02:22:38 +02:00
SirStone ff03e81591 movement final: pre-register the ship criterion before the confirmation battles 2026-09-26 02:21:09 +02:00
SirStone 47ce244740 movement j119 Batch 3+4 results: the reversal/dwell and heat-field axes are a clean negative; nothing beats strafe 2026-09-26 02:18:22 +02:00
SirStone 1256357f07 movement j119 Batch 3+4: pre-register the reversal/dwell and heat-field arms 2026-09-26 01:40:03 +02:00
SirStone 7311aaef5a movement j119 Task A: make the shipped TFIL heat shape env-overridable (default path byte-identical) 2026-09-26 01:38:51 +02:00
SirStone c07e6d9599 movement ledger: the post-hoc structure of the strafe win (60 points, both batches)
56/60 opponent-level win deltas are non-negative and no opponent family
regresses reproducibly (the one WallAvoider loss reverses in Batch 2).
corr(dwins, d_hit_rate) = -0.09, corr(dwins, d_damage) = +0.38: the aggregate
win is a survival effect, but a per-opponent hit-rate gain does not predict a
per-opponent win gain - so hit rate stays an explanation, never a proxy.
2026-09-26 01:33:24 +02:00
SirStone 44e2d191e3 movement ledger: fix two claims in the Batch-1/2 prose
- tfil is 4th of five on round wins, not last (ring is nominally 0.04 lower,
  ns) - the Batch-1 commit message overstates one word; the correction is
  recorded in the ledger rather than rewritten.
- the Batch-2 direct answer quoted three of four CIs excluding 0; it is four of
  four ([+0.04,+0.63], [+0.16,+0.60], [+0.27,+0.89], [+0.22,+0.72]).
- added the cleanest aggression isolation of Batch 1 (ring - ring_notemp, same
  engine and heat field, range weighting alone): +41.9 dmg/run, -0.11 wins/run,
  +12.8 pp incoming hit rate at 236 vs 395 px.
2026-09-26 01:32:28 +02:00
SirStone 7d3645da5a movement Batch 2: the strafe win over shipped tfil replicates; the range target decides nothing
Same frozen panel, same 3x3 design, new session on commit 8efa627 (no source file
changed since 1984a78, so the same code), 225 battles, 0 invalid runs. Arms:
tfil, strafe_notilt, strafe_325 + the tilt re-armed at 600px and 250px.

Paired vs tfil: strafe_325 +0.58 wins/run [CI +0.27,+0.89] 11/12 p=0.0063;
strafe_notilt +0.47 [+0.22,+0.72] 10/11 p=0.0117; tilt_600 +0.40 [+0.04,+0.76]
(sign test 8/11 p=0.23, sign-flip p=0.049); tilt_250 +0.38 [+0.07,+0.69] 10/12
p=0.039. Incoming hit rate -5.2..-6.9 pp with 0/15 opponents favouring tfil.

The range TARGET is not the lever: re-arming the tilt moved the achieved
distance from 459px (no steering) to 478px and 415px, and none of the three is
separable on wins. This overturns Batch 1's reading that the tilt costs wins -
the honest statement is that the tilt's win effect is below this design's
resolution. tfil reproduced to within 1.4 pp (40.7% -> 39.3% of rounds), so the
baseline itself is stable across sessions.

Ledger: Batch 2 section, the verbatim analyzer report, a data-driven
what-to-try-next, and the session log.
2026-09-26 01:31:08 +02:00
SirStone 07766303f5 movement Batch 1: pure strafe (range tilt OFF) beats the shipped tfil on round wins across a 15-opponent panel
225 battles, one frozen binary, five env-only arms, the frozen panel, 0 invalid
runs. Paired per opponent vs the shipped tfil:

  strafe_notilt  wins/run +0.38  [CI +0.16,+0.60]  9/9 opponents p=0.0039
                 dmg/run  -10.2  [CI -25.8,+5.5]   p=0.61, MDE 20.4 (not detectable)
                 incoming hit rate 12.24% vs 18.17%, dmg taken 150 vs 200
  strafe_325     wins/run +0.33  [CI +0.04,+0.63]  10/12 p=0.0386
  ring           dmg/run  +31.2  [CI +11.5,+50.9]  13/15 p=0.0074, wins/run -0.04 (ns)
                 but hit rate 29.4% at 236 px: a damage/survival trade, not a win
  ring_notemp    indistinguishable from tfil on both primaries

Round wins in this harness are survival wins (in 216/219 attributable runs the
win count equals the rounds the opponent died in), and the winner takes ~1/3
fewer hits while fighting ~74 px farther out. The shipped tfil is last of five
on wins: the DrussGT-only picture did not generalize.

Also: tournament_analyze.py now prints BOTH readings of the pre-registered
'while the other does not go down' clause (strict: nothing is better;
substantive: the two strafe arms and ring are better on one metric each).
2026-09-26 01:13:29 +02:00
SirStone 8efa627c05 gauntlet: BitBrain vs Pattern across 32 legacy opponents (does not generalize)
The repo's first multi-opponent gun measurement. Adds tools/ab/gauntlet_run.sh
(per-opponent A/B over the legacy roster, subject = frozen ModularBot),
tools/ab/gauntlet_analyze.py (paired per-opponent deltas, cross-opponent sign
test, style split, MDE) and the arm/opponent fixtures.

Result: BitBrain does NOT generalize beyond DrussGT. 32 opponents x 2 arms x
3 runs x 5 rounds = 192 battles / 960 rounds, 0 failed, 0 retries: damage/run
214.5 (pattern) vs 210.9 (bb), sign-flip p=0.53; round wins 237/480 vs 239/480,
p=0.91. Sign test: bb better on 13/32 opponents (damage). The DrussGT-only
penalty does not carry. The owner's 'killer vs regular movers' sub-claim is not
supported: regular bucket +1.3 dmg/run vs dodgers -0.2 (MW p=0.85), and the
measured movement predictability does not correlate with the delta.
2026-09-26 00:55:57 +02:00
SirStone 1984a780f4 melee A/B doc: correct the per-arm [bb-reset] counts (perRound 68.75, retained 68.19, decay 66.75) 2026-09-26 00:43:04 +02:00
SirStone da4a971ca9 MELEE A/B: BitBrain vs shipped Pattern rack — repo's first melee measurement (null)
Adds common_libs/tests/measure_melee_bitbrain_ab.nim (+ .sh driver, .py analyzer,
committed per-run fixtures) and docs/melee_bitbrain_ab.md.

Experiment: 4-bot Free-For-All (ModularBot + WaveSurfer + PatternMover +
RandomMover), 4 arms x 16 runs x 7 rounds, frozen ModularBot from git archive
HEAD (commit 0f5cfe3, binary 11bba27), shipped tfil movement in every run.
Arms differ only in the gun rack: pattern (shipped), bb_round, bb_ret, bb_learn.

Result: NOT DETECTABLE. Score (server round score = damage + survival bonus)
differs by -63..+33 pts (perm p=0.16-0.71) against an MDE of 151 (~5.1%).
Every arm finishes rank 1. Round wins hint BitBrain's way (112/112 and 111/112
vs 109/112) but p=0.225 (MW 0.080), half the 0.40-win MDE.

Liveness proven: rack boot lines flip (rack active melee = PATTERN / BITBRAIN),
every run faced 3 distinct targets and ~66-69 target changes, and the bb arms
logged one [bb-reset] reason=target_change per switch. The melee premise was
exercised; the fast adaptation bought no measurable score edge at this sample.
2026-09-26 00:35:44 +02:00
SirStone 02b691dd5f docs: add per-run damage/wins series and liveness line to surfer_wiring_ab 2026-09-26 00:29:34 +02:00
SirStone 6d6648ccf8 docs: wave surfer wiring A/B vs TFIL and STRAFE (surf does not beat the shipped default)
Live 3-arm x 15-run x 7-round A/B vs real DrussGT on commit 0f5cfe37.
Primary: surf ties strafe on round wins (37/105) and damage (255 vs 250/run),
both below tfil (45/105, 293/run; damage p=0.003). Incoming hit rate: surf
13.51% (worst) vs strafe 9.40% (best) and tfil 10.40%. So the plain surfer does
NOT dodge better and does NOT win more. Also records the j107 trap: strafe
dodges best yet wins fewer rounds than tfil. Next step: range/aggression A/B,
not a BitBrain upgrade.
2026-09-26 00:27:34 +02:00
SirStone ef16dd982d robocode_shim: generalize bridge to any legacy bot; validate 28 champions as sparring partners
Generalize the classic-Robocode shim so the hosted bot's main class, jar, extra
classpath and data directory are configurable (SHIM_BOT_CLASS / SHIM_BOT_JAR /
SHIM_EXTRA_CP / SHIM_DATA), with DrussGT kept as the default so every existing
script, fixture and generated bot dir behaves identically.

- LegacyBotBridge: generalized bridge (DrussGTBridge kept as an alias).
- BotHost: (mainClass, jar, extraJars, dataDir); disableShield -> generic
  disableStaticBoolean.
- ClassicPeer: implement ITeamRobotPeer; synthesize StatusEvent each turn;
  make move/turnBody/turnGun/turnRadar the immediate (turn-ending) variants;
  guard re-entrant execute() from event handlers; record delivered events.
- make_botdir.sh / run_bridge_battle.sh / run_smoke.sh generalized, legacy
  invocations unchanged; generated launcher uses a per-process mktemp data dir
  so concurrent bots/battles cannot clobber one classic data directory.

Validated 35 legacy bots against a real Tank Royale battle (2 rounds vs sample
SpinBot): 28 usable (27 effective + DrussGT), 5 weak-but-playing, 2 failing.
Adds robots.json (manifest) and LEGACY_BOTS.md (how-to, status, missing-API
costs). No third-party jar is committed.
2026-09-26 00:25:33 +02:00
SirStone 0f5cfe37b2 surf: wire the dormant wave surfer as TR_MOVEMENT=surf + fix 4 real defects
common_libs/movements/wave_surfer.nim was written in an early session and
never wired to the bot. This connects it exactly like tfil/strafe and fixes
the defects a full read found:

  1. the dodge direction was INVERTED: the perpendicular was built from the
     bot->enemy bearing while GF lives in the enemy->bot frame, so the bot
     moved toward MORE danger. Now built from the wave's origin->bot bearing:
     +90 provably increases GF.
  2. the danger histogram was never reset (resetRound cleared waves only), so
     it was a battle-long static average. Now reset to the uniform prior each
     round.
  3. fire detection tracked only the current target's energy via one scalar;
     now per-enemy (seq[(id,energy)]) so melee target switches cannot invent
     or hide waves.
  4. the wall penalty projected a point from the wave origin, not from the
     bot, making the wall test meaningless. Now projects the bot->candidate
     direction.

Wave speed uses the actual firepower (the one-tick energy drop IS the
firepower, so speed = 20 - 3*drop is exact). Adds TR_SURF_* knobs and
registers them in the env report. Shipped TR_MOVEMENT=tfil default untouched
(test_tfil_commit_env: 30/30 pass; test_env_report: pass).
2026-09-26 00:10:24 +02:00
SirStone 10723ab5f3 strafe: default range 200 -> 325 (mid of the owner's 300-350 band), with the j107 evidence noted 2026-09-25 23:56:19 +02:00
SirStone 4829f9ca13 BitBrain vs TMHorizon vs Pattern: live A/B on shipped TFIL (null result)
4 arms x 15 runs x 7 rounds (60 battles, 0 failed) vs real DrussGT on the
shipped TFIL default, frozen at ed25ce2. bb_id (gain 1.0 identity) is
statistically indistinguishable from shipped Pattern -> plumbing validity
check passes. No BitBrain arm beats TMHorizon or Pattern: bb_learn (the config
the owner likely ran) is the worst arm (276 dmg/run, 39/105 wins), the only
comparison at alpha=0.05 is Pattern beating it on damage. Learned gains
(>=1.0, gated >=300px) over-lead and lose 1.61pp of hit rate at 300-450px.
MDE 29.5 dmg/run, 1.235 wins/run; a 6-4-sized effect needs ~39 runs/arm.
2026-09-25 23:55:21 +02:00
SirStone ccff7e3e4a STRAFE: curved wings + guaranteed wall/corner escape; shipped TFIL default untouched
Task j112. Two changes to the TR_MOVEMENT=strafe engine, both OFF the shipped
tfil path; the binary default is still tfil.

CURVED WINGS: the candidate set was the straight 1-D line through the bot, which
a bounded segment always terminates at a wall. It is now an adaptive parabola
with the vertex on the bot:
  point(y) = bot + yhat*y + xhat*kappa(y)*y^2
xhat is the unit vector away from the nearest wall(s) (summed inward normals, so
a corner yields the diagonal). kappa grows as the wall approaches and saturates
at TR_STRAFE_KAPPA; every wing point is clamped inside TR_STRAFE_WALL_SAFE, so
the wing FLATTENS and runs parallel to the wall instead of touching it. In open
space kappa == 0 and the wing is exactly the old straight line. The wing chord
at the reach tilts the heading band toward the interior by atan(kappa*reach)
(capped by TR_STRAFE_WING_MAX); the body still only turns slowly to follow that
tangent, never to face the target, and reversals are still setForward sign flips.

GUARANTEED ESCAPE: with every candidate over threshold the old fallback minimised
pathMaxHeat, whose gradient points AT the wall (the shortest path has the least
wall exposure), so the least-hot tile was the adjacent one and led further along
the wall. Near a wall the picker now ranks by the DESTINATION (farthest from the
wall, then coolest tile) and commands the sign whose velocity has a positive
component along the wall-away normal. That sign is re-asserted EVERY tick, so
speed*heading . away >= 0 while escape is active: the clearance cannot fall.
mode=escape reaches the [strafe] log. A mild wall-margin bias
(TR_STRAFE_WALL_BIAS) prefers higher-clearance tiles when near a wall.

GUI/log: the curved wings are drawn as an orange polyline (candidates follow the
curve), a white ray + ESCAPE label marks the escape, and the [strafe] line now
carries wall=<dist> kappa=<..> mode=<pick|fallback|escape|radial>.

Gates (offline, kinematic replay of the DrussGT fixtures; see
common_libs/tests/measure_strafe_wings.nim, plus the reused j108/j111 gates):
wall occupancy (within 54 px) falls 25.6->4.3 / 18.3->4.1 / 23.8->4.3 / 27.8->4.7
percent and the longest continuous wall run 191->31 / 49->25 / 208->47 / 246->27
ticks; corner-region occupancy 4.7->0.0 percent with the longest corner run
71->7. The escape sweep (3520 start x heading x enemy runs, 110k escape ticks)
shows ZERO per-tick guarantee violations and a worst corner run of 21 ticks.
Open-space parity is bit-identical (kappa == 0), reversals are still sign flips
(0 non-sign commands), and mean turn/speed are unchanged (OFF 4.42 deg/tick,
31.3 percent no-turn vs ON 4.45 / 30.7; reversal-interval entropy 5.309 -> 5.311
bits). The fixtures are OPEN-LOOP, so these are veto-capable checks, not a live
win claim.
2026-09-25 23:51:12 +02:00
SirStone ed25ce29ad STRAFE: range control (tilt) + corner-stall escape; shipped TFIL default untouched
Task j111. Two changes to the TR_MOVEMENT=strafe engine, both OFF the shipped
tfil path; the binary default is still tfil.

RANGE CONTROL (a hypothesis under test, no default changed elsewhere):
the body is still pinned ~perpendicular to the threat, but the line is tilted
by the range error: lineAngle = threat + 90 + appliedTilt, with the tilt zero
inside +/-TR_STRAFE_RANGE_TOL around TR_STRAFE_RANGE (200 px, chosen because it
is exactly TR_POWER_FAR_DIST) and clamped to +/-TR_STRAFE_TILT_MAX. A tilt alone
cannot change range (the picker chooses both ends at random), so the picker also
PREFERS the end that reduces |distance - target| with a probability that grows
with |tilt|; both ends stay possible. The tilt sign is aligned to the ENEMY
bearing, since  is the bullet direction (roughly its opposite) when a
bullet is in flight. Knobs: TR_STRAFE_RANGE (200), TR_STRAFE_RANGE_TOL (25),
TR_STRAFE_TILT_MAX (15), TR_STRAFE_TILT_GAIN (0.10), all registered in
env_report.nim (emit + knownEnvNames). NOT claimed to be better: j107 measured
that drifting 25-30 px closer made damage/run and wins WORSE.

CORNER STALL (a real defect): a line whose in-arena candidate set was empty set
targetValid=false and kept driving on the last sign, so the bot could oscillate
inside a corner tile forever. Three defenses: (1) a deterministic corner guard
projects the outward component off the line whenever BOTH ends are outside, so
the line becomes wall-parallel and a candidate always exists; (2) a degenerate
line (<=1 candidate) falls back to a radial search for the coolest in-arena
tile and commits the sign; (3) a commanded move with no displacement for
StuckFlipTicks (5) ticks flips the sign. Both warnings now reach the [strafe]
log.

GUI/log: the tilt is drawn as the existing strafe line (it is lineForward), plus
a green/red ray toward the enemy (length = |distance-target|) and white text
d=.. tgt=.. tilt=..; a red disc marks a stuck tick. The existing overlays and
the j110 heat grid are unchanged.

Gates (offline, kinematic replay of the DrussGT fixtures; see
common_libs/tests/measure_strafe_range_stall.nim): on the j110 field all four
corners that were 100% confined inside 72 px / 22.6 px max before now escape
(<=2.6% confined, 209-741 px); the achieved |distance-200| falls on 3 of 4
fixtures (mean -16% to -30%); mean |turnRate| and the 8.00 px/tick speed are
essentially unchanged (no-turn property survives). Reversal-interval entropy
falls 5.85 -> 5.31 bits (still above TFIL's 5.09): the range bias costs some
reversal randomness while closing.
2026-09-25 23:39:28 +02:00
SirStone 1ea84f5b6e STRAFE: draw the heat field, real bullet danger, corrected ring comment
Three defects the owner hit as "no heat tiles anymore" under TR_MOVEMENT=strafe.

1. The strafe overlay drew ONLY the tiles on its strafe line, so the computed
   heat field was essentially invisible. It now draws the WHOLE field exactly as
   TFIL does (every non-zero tile, yellow->orange->red ramp by field max, integer
   value label) behind the same debugGraphics flag, with TR_STRAFE_HEAT_GRID=0 to
   hide it. The strafe overlays draw on top, unchanged.

2. STRAFE carried the SHIPPED bullet constants (core 10 / aura 5), so a bullet's
   own heat sat exactly ON PathDangerThreshold (10.0) and a bullet was never
   dangerous on its own in this mover; it only ever bit through its corridor.
   Defaults are now the retune's 20/10, exposed as TR_STRAFE_BULLET_CORE /
   TR_STRAFE_BULLET_AURA.

3. The ring mover's header documented CorridorHeat 5.0 / WallHotness 10.0 while
   the code has always been 10.0 / 15.0. A job read the comment and handed out
   sub-threshold heat values, which emptied the field. The comment now states the
   real values and their actual behaviour; no code values changed.

Also sets strafe's heat defaults to the retune shape (bullet 20/10, corridor 10,
wall 15/5, pillar 0), documented with the reason.

Gate A re-run (j110, offline DrussGT fixture, measure_strafe_gates.nim):
  corrected DEFAULT : 24.6% of picks with ZERO safe tile, mean 11.17 safe
  j108 shipped field: 63.4% / 3.70   (reproduced exactly)
  j108 ring retune  :  8.1% / 18.41  (reproduced exactly)
  bullet isolated   : 11.4% / 17.07
The corrected default beats the shipped field but is WORSE than j108's retune
row: the bullet retune alone costs 8.1 -> 11.4, the corridor/wall retune accounts
for the rest. That is the deliberate price of making a bullet dangerous.

Guards green: test_env_report 24 PASS, test_tfil_commit_env 30 PASS (shipped TFIL
default untouched, byte-for-byte), test_tfil_ring_weights 24 PASS. The three new
knobs are registered in the boot env report so the tree-scan guard stays clean.
2026-09-25 23:25:34 +02:00
SirStone dc071b83f3 ModularBot: [result] round/battle outcome log lines (TR_RESULT_LOG, default on)
One greppable '[result]' line per round plus one at battle end, stdout only:
  [result] round 3/7  WE WON    (enemy destroyed)  | us 42.1 energy, them 0.0, 812 ticks | rounds won 3/7
  [result] battle END: rounds won 4/7

Outcome is authoritative from RoundEndedEventForBot.results.rank (1 = winner);
death observations (our onDeath, enemy onBotDeath) and onWonRound refine it
into WE WON (enemy destroyed) / WE DIED (killed) / BOTH DIED (score decided) /
TIMEOUT (score decided). Our own death is reported the instant it happens.
TR_RESULT_LOG registers in the boot env report; default on, only explicit
off-values disable it. No behaviour change - logging/state only.
2026-09-25 22:50:46 +02:00
SirStone a50c0125d5 STRAFE movement: body pinned perpendicular to the threat, reversals by sign flip
New engine movements/strafe.nim, selected by TR_MOVEMENT=strafe (default stays
tfil, byte-identical — test_tfil_commit_env.nim's 30 checks still pass).

Design (the owner's):
- AXIS = incoming bullet's direction when a bullet is in flight, else the
  perpendicular of the enemy bearing. The body heading is kept inside a band
  (TR_STRAFE_BAND, default 20 deg) around the perpendicular LINE; it turns only
  when outside the band, and never turns to face a movement target.
- Candidate tiles on the perpendicular line through our position, both forward
  and backward, within TR_STRAFE_REACH px, with a perpendicular jitter of
  +/- TR_STRAFE_SPREAD tiles. A tile is acceptable when its path max heat is
  <= PathDangerThreshold, the SAME safety rule TFIL uses.
- Move by SIGN only: setForward(+/-MaxSpeed>). Dwell is re-picked after a random
  number of ticks in [TR_STRAFE_DWELL_MIN, TR_STRAFE_DWELL_MAX], on arrival, or
  on a serious threat spike.
- Heat machinery is REUSED from the shipped mover, not re-implemented: the
  exported heatDecay()/bulletMagScale() (j105 time-indexed model) and the
  PillarHotness/PillarRadiance globals (j106 pillar-free default). The heat
  shape is overridable via TR_STRAFE_CORRIDOR_HEAT/WALL_HOTNESS/WALL_RADIANCE
  (defaults = the shipped TFIL field).
- GUI overlay: strafe line, threat axis, candidate tiles (safe/unsafe), chosen
  target, sign-coloured movement ray, and the heading band.

Gates (offline, recorded DrussGT fixture, 20026 ticks):
- A TILE AVAILABILITY: shipped heat field -> a safe tile exists on only 36.6%
  of picks (63.4% fall back to the least-hot tile); the ring retune
  (corridor 5, wall 10/5) raises it to 91.9%.
- B PREDICTABILITY: reversal-interval entropy 5.84 bits vs TFIL 5.09; direction
  entropy 1.00 both; long-lag autocorrelation ~0 for both (no periodic
  component). Fewer reversals (710 vs 1453) and more full-speed ticks.
  measurements: common_libs/tests/measure_strafe_gates.nim

Also registers TR_STRAFE_* in the boot env report (ModularBot_garage/src/
env_report.nim) and wires the engine into ModularBot.nim (hold -> strafe,
ram trigger -> rammer).
2026-09-25 22:38:52 +02:00
SirStone 99cf9e5c82 tfil heat/pillar A/B: record the owner's decision to keep the virtual pillar removed 2026-09-25 22:24:47 +02:00
SirStone 48f38b80e7 TFIL heat-time + virtual pillar: live A/B (6 arms x 70 rounds) - neither change beats the pre-change mover
Runs the pre-registered A/B for the two movement changes in HEAD: the
time-indexed bullet heat (TR_TFIL_HEAT_TIME, fca8993) and the removal of the
invented virtual centre pillar (d0750ab). One frozen binary from HEAD vs real
DrussGT: 6 arms x 10 runs x 7 rounds = 60 battles, 420 rounds, 0 failed.

Judged on damage/run and ROUND WINS only (hit rate and hits-taken are context):
hit rate would have inverted the verdict again - tau3 has the best pooled hit
rate of all arms (11.56%) and the fewest round wins (20/70).

RESULT (vs the reconstructed pre-change mover "old"):
  heat-time HURTS. tau3/tau5/tau9 lose 1.3-1.7 wins/run (p=0.0010-0.0125) and
  deal 22-38 less damage/run (p=0.004-0.047); tau15 is a wash on wins (p=0.64)
  and 22 damage/run lower (p=0.046). Nothing improves either metric.
  pillar removal does nothing measurable. old vs pillaoff: +5.7 damage/run
  (p=0.71), +0.5 wins/run (35 vs 30, p=0.43), 30.8 MORE damage taken/run
  without the pillar (p=0.040). The mechanism check proves the knob works
  (centre-box occupancy 0.09% -> 2.37%, p<0.0001; range 469 -> 443 px,
  p=0.0002), so this is a real behaviour change that buys nothing. At n=10 the
  pillar contrast is inside the MDE (33 damage/run, 1.2 wins/run), so this is
  not a proven regression.

Flags that the shipped default (pillar removed) should be reverted to the
TR_TFIL_PILLAR_ON behaviour; heat-time stays off.

Adds tools/ab/arms_heat_pillar.txt and tools/ab/ab_mechanism.py (per-tick
mechanism check: central-box occupancy, range distribution, live enemy-bullet
proximity) plus the captured summary/report fixtures.
2026-09-25 22:23:11 +02:00
SirStone f58d65d2e8 TMComposites gate: per-gun confidence faithful for 3 guns; no pair composes
Adds a per-sample intrinsic-confidence field (GunPrediction.confidence,
threaded through FeedbackEvent/VirtualBullet, populated by Pattern, DecayGF,
KNN, GuessFactor, Tsetlin, TMHorizon) and an offline recorder + analyzer that
reproduce the paper's Figure 2 per gun and its Eq-8 composite.

Measured on 3 held-out tr-bridge DrussGT battles (33k ticks, ~133k samples/gun):
- FAITHFUL: DecayGF (rho +0.133), KNN (+0.090), Pattern (+0.064, weak).
- GuessFactor is ANTI-faithful (rho -0.067); Tsetlin c_max is useless (0.001).
- No pair of guns specialises complementarily: the same gun dominates both
  high-confidence slices in every pair.
- Eq-8 alpha-normalised confidence-weighted composite: 18.41% vs Pattern
  20.45% (McNemar p=3.1e-126). Faithful-only variant 18.68%, still loses.
  Shuffle control passes weakly (composite > shuffle, p=4e-14) so ~0.7pp of
  competence is real but ~2pp short. Offline veto: design is dead.

See docs/tmcomposites_gate.md.
2026-09-25 22:04:39 +02:00
SirStone d0750ab020 TFIL: remove the virtual centre pillar from the shipped default; register j102 env reads
The default mover painted a 30/10 radiance blob on the arena centre even
though the arena has NO physical pillar there, creating a 4x4 tile
(144x144 px) exclusion zone over open centre floor. Set
PillarHotness/PillarRadiance to 0/0 in the shipped default (matching the
ring variant) and add TR_TFIL_PILLAR_ON=1 to restore the old 30/10 field
for A/B without a rebuild; registered in env_report.

Because the shipped default legitimately changed, the default-path parity
golden (fixtures/tfil_commit_default.golden) was regenerated from the NEW
default, with an explicit 'deliberate default change' note in the test so
a future failure is treated as a real regression.

Also register the three env reads job j102 added in common_libs/bitbrain
(TR_BITBRAIN_MODE / _DECAY_EVERY / _DECAY_SHIFT), which the env-report
guard was failing on.

Verification: test_env_report all green; test_tfil_commit_env 30/30.
2026-09-25 22:02:00 +02:00
SirStone fca899376e TFIL: time-indexed bullet heat (TR_TFIL_HEAT_TIME, DEFAULT OFF)
Make danger a function of time-to-arrival instead of flat distance. Bullet
core/aura/corridor heat becomes magnitude(power) * decay(dt), dt = along/speed:

  * decay(dt) = exp(-dt/tau) is a function of TIME; a fixed tau projects a
    pixel reach of speed*tau, so fast/weak bullets get a longer slope and slow
    ones a shorter one — derived from speed = 20 - 3*power, not hand-tuned.
    tau = TR_TFIL_HEAT_TAU.
  * magnitude(power) scales the near-end heat with power from DAMAGE
    (calcBulletDamage = 4p, linear in p; SCORE_PER_BULLET_DAMAGE = 1.0). Hit
    probability is FLAT across power (docs/env_reference.md), so risk does not
    justify power scaling — the cost of the hit does. Floored at 1.0 so a weak
    bullet's near end is never less dangerous than the flat model.
    Gain = TR_TFIL_HEAT_POWER_GAIN.

Every source is already f(dt), so the time-indexed planner (evaluate a cell at
the tick the bot would ARRIVE, i.e. heatDecay(dt - arrivalDelay)) is a one-line
change. It is intentionally NOT implemented here.

Default path is byte-identical: with TR_TFIL_HEAT_TIME unset both factors are
exactly 1.0 (IEEE x*1.0 is exact), and the committed golden replay in
common_libs/tests/test_tfil_commit_env.nim (20,026 ticks) still passes
byte-for-byte against the pre-change mover. The debug corridor outline is also
drawn only to the model's reach when enabled, so the GUI shows the shortening.

Offline field measurement (common_libs/tests/measure_tfil_heat_time.nim,
46,054 fixture ticks, tau=9/gain=1): corridor reach drops from 443px
wall-to-wall to 143px mean (32% retained); fraction of tiles > 10 goes
0.61 -> 0.57; largest contiguous safe region 118 -> 140 tiles; mean
distance-to-nearest-safe-tile 49 -> 42px. Saturation stays high because wall
radiance + pillar alone are 44% of tiles over threshold and are untouched.

Registers the three knobs in env_report (report + known-name set).
2026-09-25 21:47:54 +02:00
SirStone 4270136948 docs: correct counted-SBC decay amortised cost figure 2026-09-25 08:43:48 +02:00
SirStone d85ff53d34 State-window gate: a temporal window of wave-relative states does NOT beat a single state
Re-runs the SBC coincidence premise as a cheap veto test on the 70-battle
live-vs-DrussGT corpus (/tmp/tfil_ab2).  Defines the wave-relative state
(lat/vlat/toa/room/turn, 5/7.9/10 bits at Q=2/3/4), quantises the miss offset
at the bullet's arrival into 7 bins, and sweeps window length K in
{1,4,8,16,32,48} with an interpolated suffix-backoff model under a BY-BATTLE
70/30 split (3 seeds).

Result: NO.  On the pre-fire frame the window is worse than the single
fire-tick state at every K/Q/A (e.g. K=8 costs +0.35..+0.46 bits).  On the
during-flight frame the entire apparent gain is the later decision tick, not
the window; the single state alone drops 2.70 -> 1.31 bits as K goes 1 -> 32.
The shuffle-order control confirms recency matters but the windows do not:
by Q=4/K=8 they average ~1 observation and never recur.  The single state
survives as a strong predictor (log-loss 2.346 vs 2.698 majority; bin
accuracy 0.409 vs 0.235).

Gate only: no gun, no live-win claim.
2026-09-25 08:43:36 +02:00
SirStone 40ba96f649 BitBrain SBC: counted mode + global decay (forgetting, probabilities)
Adds an smCounted storage mode alongside the default smBitset. Each
(i,j,class) cell becomes a saturating uint8 counter; learn increments it and
a global fractional decay (c -= c shr decayShift every decayEvery learns)
makes forgetting possible. infer sums raw counters; new inferProb sums the
per-cell posterior P(class|cell) (scale-free, recommended readout).

Bitset path is the default and byte-for-byte unchanged: test_bitbrain 56/56
(was 32), and test_bitbrain_mnist reproduces 97.210% corrected / 96.540%
bug-compatible exactly.

Counted mode configurable at runtime (TR_BITBRAIN_MODE / TR_BITBRAIN_DECAY_*)
and compile time (-d:bitbrainDecay*). Measured: forgetting (86.2% vs 48.9% on
a permuted-label stream), probabilities (rare-class balanced 0.998 vs 0.500),
and the stationary cost (counted hurts MNIST; see docs/bitbrain_counted_sbc.md).

Harness: common_libs/tests/measure_counted_sbc.nim
2026-09-25 08:39:10 +02:00
SirStone 39e06719fb Campaign phase 2 ledger: live lead-gain sweep (gains above 1.0 do NOT beat Pattern)
6 arms x 7 runs x 7 rounds vs real DrussGT on one frozen binary (2747ebd).
Validity check PASSES: g100 (fixed gain 1.0) is statistically indistinguishable
from control (dmg p=0.65, wins p=0.62, ALL hit rate +0.04pp p=0.92; zero [bb]
lines = provably no correction). Fixed gains above 1.0 LOSE at 300+: gfix150
-94 dmg/run (p=0.0006), 4/49 vs 16/49 wins (p=0.009), -3.85pp at 300-450
(p=0.0006) and -2.25pp at 450+ (p=0.0023). The hypothesis arm ghi (learner
allowed above 1) is directionally positive but inside the MDE (+14.3 dmg/run
p=0.41; +0.83pp at 450+ p=0.35). Kill the gain axis in both directions.

Includes the mandatory correction notice: Phase 1's [1,1,1,0,0] is LIVE-REFUTED
by 140fe25, and the standing rule that offline is veto-only / live decides.
2026-09-25 00:24:45 +02:00
SirStone 2747ebd323 BitBrain: TR_BITBRAIN_GAINS env knob (candidate set + fixed-gain degenerate)
Task A of campaign phase 2: the lead-gain candidate set is now pure env, so the
live arms need no recompile.

- common_libs/guns/bitbrain_gun.nim: BB_GAINS_ENV (TR_BITBRAIN_GAINS); the
  candidate list is parsed once at gun construction into a dynamic seq, so the
  hit counts/hit rates are sized to it. Unset/unparsable -> the shipped
  BB_CAND set [0,0.25,0.5,0.75,1.0] (byte-identical behaviour). Exactly ONE
  candidate degenerates to a FIXED gain applied from the first shot (learning
  bypassed), still gated to the long bands. parseGains clamps to [0,8],
  de-dupes and sorts so the argmax tie rule is unchanged. The [bb] line now
  prints the APPLIED gain AND the resulting angular shift, so a run's
  correction is auditable from stdout.
- ModularBot_garage/src/env_report.nim: emit TR_BITBRAIN_GAINS (resolved
  candidate set) and add BB_GAINS_ENV to the known-name list.
- tools/ab/arms_leadgain.txt: the 6-arm phase-2 sweep definition.
2026-09-25 00:15:27 +02:00
SirStone c305ef4212 BitBrain campaign phase 1: the missing gain sweep and BitBrain as a lead-gain corrector
Task A — the gain region Phase 0 never covered (gain < 1). Extend the
prediction-quality ruler with gain 0.25/0.50/0.75 arms and a fixed causal
per-band arm. Full 70-run result: the hitProxy-argmax curve is
[1.00, 1.00, 1.00, 0.00, 0.00] — Pattern below 300 px, HeadOn above —
worth +0.13 pp at 300-450 and +2.16 pp at 450+ (0.0767 -> 0.0984). Lead
correlation is identical for every g>0 (Pearson is scale-invariant), so a
shrinking gain adds no lead information; and the least-squares optimum
[1,1,1,1,0.25,0.25] diverges from the hitProxy optimum because Pattern's
lead errors are bimodal.

Task B — rebuild guns/bitbrain_gun.nim as a lead-gain corrector:
aim = LOS + gain*(patternAim - LOS), gain learned online per range band by
ranking candidate gains on the hit-probability proxy (the observed lead
label via tmhObservedAt), gated to range >= 300 px. ADE+SBC output removed.
Offline (70 runs): Pattern below 300 px, +0.37 pp at 300-450, +1.90 pp at
450+ (hitProxy 0.0957 vs 0.0767), matching the fixed rule to within 0.27 pp
at 450+ and exceeding it at 300-450. ~0.0007 ms/tick marginal (old gun
~0.114 ms/tick). Default off; rack membership, env report and guard tests
(bitbrain 32, registration 13, rack 48, tm_pattern 20, env_report) unchanged
and green.

Ledger: docs/bitbrain_campaign.md Phase 1, including the three on-file
negatives and the causal-shippability note. No live claim.
2026-09-25 00:10:19 +02:00
SirStone 140fe2519a HeadOn (no-lead) vs Pattern LIVE at long range: clean negative, offline ruler killed
2 arms x 15 runs x 7 rounds, one frozen binary from HEAD a82c864, real DrussGT,
server-side events sidecar. Shipped rack is onlyPattern, so control=Pattern-only
and headon=HeadOn-only (TR_RACK_PATTERN=off TR_RACK_HEADON=both).

  arm      dmg/run  dmgtk/run  round wins  shots/run
  control      279        211     48/105       785
  headon        14        228      0/105       580

Round wins and dmg/run both separate at p<0.0001 (MC permutation, se 0.0000),
~7x the damage MDE (35.8). Per range band (pooled, 15 runs):
  300-450: Pattern 12.3% (4590 shots) vs HeadOn 0.6% (3701)  p<0.0001, MDE 2.0pp
  450+   : Pattern  9.2% (6671)       vs HeadOn 0.4% (4177)  p<0.0001, MDE 1.1pp
HeadOn loses EVERY long-range band by 20-23x, so the whole-battle loss is not a
close-range artefact.

The offline ruler (prediction_quality_results.txt) predicted the opposite: HeadOn
meanAbs 14.61 vs Pattern 17.53 at 300-450 and 12.33 vs 16.19 at 450+, hitProxy
.105/.104 and .098/.077 (+27%). That is an open-loop replay of a FIXED enemy
track, so it cannot see that a different bullet makes the surfer dodge
differently; live, the static gun does not lead at all.

TR_PATTERN_RAD_SCALE arms were skipped: applyRadial scales aim DISTANCE along an
unchanged bearing, so it cannot express 'less lead' (bearing is what firing uses).
HeadOn confirmed to ignore bulletSpeed (head_on.nim:9), liveness OK 15/15.

Adds the range-band analyzer tools/ab/ab_range_bands.py (reuses the lead-capture
Run alignment) and the captured fixtures. Does not touch bitbrain_gun.nim /
bitbrain_campaign.md (job-100).
2026-09-24 23:52:49 +02:00
SirStone a82c864c60 bitbrain campaign phase 0: offline prediction-quality ruler and the bar
New harness (common_libs/gun_harness/prediction_quality.nim +
common_libs/tests/run_prediction_quality.nim): per-gun single-tick aim error in
degrees against the true continuous interception point on the recorded
live-vs-real-DrussGT corpus (/tmp/tfil_ab2/out, 70 runs, 899607 ticks), per
range band, with the hit-probability proxy mean(|err|<=atan(18/range)).
Validated: recorded hits separate from misses 13.34x px (reference 11.59x),
perfect-oracle max |err| = 0, correct ordering on synthetic ground truth, two
full runs byte-identical. Fixed a wrap180 bug (Nim float mod keeps the dividend
sign) that inflated the negative error tail.

Bar (mean|err| deg [hitProxy] at 450+): Pattern 16.19 [0.077], naive-linear
22.86 [0.054], TMHorizon 16.20 [0.076], BitBrain 16.20 [0.077], static HeadOn
12.33 [0.098], oracle 0 [1.0]. Lead-gain sweep on Pattern is a dead end (1.0
wins every band). Naive-linear applies ~1.8x Pattern's lead but carries no more
lead information (corr 0.178 vs 0.165) and is strictly worse. Ledger:
docs/bitbrain_campaign.md. All verdicts remain live-only.
2026-09-24 23:40:56 +02:00
SirStone 32a5e72fac Gun mixing (TMHorizon+BitBrain) vs DrussGT: clean negative, no dodge disruption
4 arms x 7 runs vs real DrussGT. mix alternates the two guns 476 times/7 runs
(liveness OK) but our bullets are no more varied (power sd / aim-offset sd flat)
and DrussGT's dodge quality is unchanged (miss/tick mix-pat +0.03, p=0.66; MDE
3.8%). mix wins 24/49 = the 49% baseline; the user's 6/10 has P=0.353 at 49%.
New tools/ab/ab_dodge_analyze.py splits the validated per-shot dodge instrument
by arm and adds gun-switch/power/bearing liveness; fixtures committed.
2026-09-24 23:07:05 +02:00
SirStone d93ce444c0 BitBrain verdict: clean negative at 30 runs/arm; analyzer gets MC + Mann-Whitney + MDE
- docs/bitbrain_gun_verdict.md: control vs bb_decay (decay SBC memory) vs a
  provably-zero placebo, 30 runs/arm vs real DrussGT. Nothing separates
  (bb_decay +3.3 dmg/run, p=0.71; round wins 97/210 vs 97/210, p=1.00); the
  7-run shape does not replicate. TR_BITBRAIN_RANGE=0 is clamped to 1.0 deg
  (bitbrain_gun.nim:207) so it is NOT a zero-shift placebo; TR_BITBRAIN_MIN_OBS
  unreachable is used instead.
- tools/ab/ab_analyze.py: keep exact enumeration for C(n,na)<=20e6 (7v7), add
  a seeded Monte-Carlo permutation test (1e6 draws, 0x5eed5eed) with its
  standard error, a tie-corrected Mann-Whitney U cross-check, a minimum
  detectable effect line, all-pairs comparisons, and a [bb] shift check.
- tools/ab/README.md: document the new analyzer output.
2026-09-24 23:02:59 +02:00
SirStone f91e121965 lead capture by range: we under-lead (0.40->0.135) — a RANGE effect, not a power one
Measure, for every shot ModularBot fires at the real DrussGT, the lead we
actually applied vs the lead the enemy's motion required, from the recorded
live battles (/tmp/tfil_ab2, 70 battles / 490 rounds / 54926 shots, plus a
35-battle powtest replication of a different binary).

- requiredLead from an AIM-INDEPENDENT interception solve (bullet speed vs
  enemy truth), appliedLead from the server-recorded bullet bearing.
- capture = applied/required, guarded at 2px lateral lead (1.6% excluded);
  headline metric is the robust proportional slope.
- validation: hits 11.6px mean miss / 80.8% inside 18px, misses 134px,
  11.6x separation; 496/496 death + 70/70 owner attributions correct.

Direct answer: capture falls with RANGE (capSlp 0.401 -> 0.135, and
|err|/tolerance 1.27 -> 7.54) but is FLAT across fired POWER within a band
(450+, enemy alive: 0.154 / 0.127 / 0.127). The sub-0.5 long-range shots
(1.18% hit) are finishKill endgame shots at a near-dead DrussGT, not a
lead-capture failure. A naive linear predictor captures 0.29-0.60; we reach
46-67% of that, so the under-lead is real but capture=1.0 is unattainable
against a dodger (oracle required lead).
2026-09-24 22:47:18 +02:00
SirStone 795a0e59fe BitBrain gun (id 16): Pattern-relative ADE+SBC aim corrector, default off
Wire the verified common_libs/bitbrain ADE+SBC library into ModularBot as a
fine-grained angular corrector on top of Pattern's prediction, the shape the
offline gate test measured (argmax readout over N correction classes).

- common_libs/guns/bitbrain_gun.nim: new gun. Input = the existing TMHorizon
  53 bits (tmhBaseBits + tmhLits); output = argmax class centre over
  +-TR_BITBRAIN_RANGE, applied by rotating the Pattern point around the shooter
  exactly as tmhApplyShift does. Label = the +h-tick fact from TmHorizonGun's
  own observation ring (never across a round). Prequential (defer + resolve).
  AD layer synthesised online for our binary inputs (center=0): heuristic
  cold-start thresholds + running-histogram ~1% percentile init + the library's
  adaptThresholds. Memory modes perRound (default, measured best) / retained /
  decay (periodic partial SBC wipe). Lazy network build + local RNG, so the
  default path builds nothing and consumes no global randomness.
- tm_horizon.nim: export tmhUpdateHistory and add tmhObservedAt (label seam).
- selector.nim: register BITBRAIN at rack id 16, default rmOff, in the SAME
  commit as the id and the wiring (the aed579b admission bug is not repeated).
- ModularBot.nim: id 16 wired through predict/spawn/onResult/resets/colors,
  arrays grown 16->17, spawn gated on rack admission, per-round/per-battle/
  target reset hooks.
- env_report.nim: report every TR_BITBRAIN_* knob + add names to the known set.
- tests: update the rack length literals; new test_bitbrain_registration
  (default-parity: off, lazy, global-RNG clean).

Guard counts unchanged: rack 48, tm_pattern_registration 20, vbullet_admit 12,
env_report 25, and the rest of the suite green.
2026-09-24 22:25:04 +02:00
SirStone 1fec87537d gitignore: keep the BitBrain primary-source zip local (explicit path, not a pattern) 2026-09-24 22:09:36 +02:00
SirStone 17c50159ba tools/ab: reusable A/B runner + analyzer (frozen-HEAD build, exact permutation test, liveness check) 2026-09-24 22:08:59 +02:00
SirStone e670788eee env_reference: the ring mover was NEVER measured offline - correct a false label
The offline-harness audit (`e40c849`, `docs/offline_harness_trust.md`) found that the
claim "best offline hit rate of anything measured" for the ring mover was false.

The 20.28% figure is a LIVE number: `docs/feature_ab_results.md` and commit `bfdcdf8`
record 35 real-DrussGT bridge battles with a server-side event sidecar as the ground
truth, and the 6/49 round wins is likewise live. There is no offline measurement of
the ring mover anywhere - the offline harness scores GUNS, not movements, and has no
movement driver at all.

So this was NOT an offline-vs-live calibration failure, which is how it has been
described repeatedly (including by the orchestrator). It was a METRIC MISMATCH: a
movement arm judged on hit rate instead of damage/run and round wins - and hit rate is
precisely the metric that concealed its collapse.

The lesson previously attached to this result was therefore the wrong one. The
paragraph now says what actually happened and points at the audit.

Docs-only; no code touched.
2026-09-24 21:44:17 +02:00