Commit Graph

29 Commits

Author SHA1 Message Date
SirStone d21f7ce5f5 j147: the 1-tick fire-detection lag is OURS — measure it, then back-date it (TR_FIRE_LAG)
MEASURED LIVE (common_libs/tests/measure_fire_ghost_lag.py, 4 sessions, 1777
matched ghost spawns, both movers): the server dispatches a turn's fire AFTER
our go() for that same turn, so a turn-T shot's energy drop first reaches our
scan at turn T+1 — and a bullet takes its FIRST step during the turn it is
fired, so the true bullet is already one whole bullet step (11-20 px) downrange.
Both movers place the ghost at the SCANNED enemy position (where the bullet was
born), so the whole ghost trajectory is the true one shifted one turn later and
the arrival deadline is a full tick late.

MEASURED: detection lag +1 tick on 100% of 1777 matched spawns; ghost-vs-
observer displacement 19.06 px mean / 22.00 p90 (tfil) and 16.08 / 21.81
(strafe); arrival-deadline error 0.99 / 0.77 ticks. NOT a rendering artefact:
the draw/advance order is correct (advanceBullets -> detectFires -> build).

THE FIX: TR_FIRE_LAG (int, default 0 = today byte-for-byte) in the shared
fire_tracker, applied by both movers at spawn: x = origin + dir*speed*lag.
The deadline needs no separate change — both movers derive it from the ghost's
own position, so a correct position gives a correct deadline.
WITH IT: displacement 19.06 -> 5.37 px mean (the residue is the enemy's own
<=8 px scan staleness) and the deadline error 0.99 -> 0.06 ticks.

Guards: test_tfil_commit_env 77 -> 87 checks (default golden parity, exact
n-step back-date, deadline shortens by exactly lag, junk/negative degrade to 0,
reaped exactly one tick earlier); test_env_report + test_env_dotenv green.
TR_FIRE_LAG registered in env_report + knownEnvNames + .env.example +
docs/env_reference.md. Live A/B pre-registered in docs/movement_campaign.md
(Batch 8) with its MDE stated up front; arms tools/ab/arms_fire_lag.txt.
TR_FIRE_DIAG gains a per-round ROUND line (the tick->getTurn anchor) and a
per-spawn SPAWN line (the ghost's drawn position).
2026-09-26 23:08:57 +02:00
SirStone de5d02ba3f j146 ledger: the field shape restores a real safe set offline (filter broken 63.5%->30.4%) and is a live null on damage/run and round wins; 375 battles, 5 arms 2026-09-26 22:35:17 +02:00
SirStone 298ea6d586 j146 TFIL: the bullet's own heat becomes tunable (TR_TFIL_BULLET_CORE/AURA, default 10/5 = today) and the field-shape A/B is pre-registered 2026-09-26 22:16:06 +02:00
SirStone 5dd921c26b j145 ledger: 300 battles, the turn tiebreak is a real but small mechanism with an under-powered outcome null 2026-09-26 22:13:03 +02:00
SirStone 39c90fd930 j145 TFIL: a turn-cost tiebreak among the SAFE tiles (default-off)
The picker scored candidates on pathMaxHeat alone and then drew uniformly
among the survivors, so a mirror-side tile was as likely as a straight-ahead
one. Added a continuous turn cost as a DRAW WEIGHT applied only after the
hard heat filter:

  w = max(1, round(1 + TR_TFIL_TURN_BIAS * (1 - max(0,|turn| - REF)/180)))

- TR_TFIL_TURN_BIAS (default 0) is the odds ratio straight-ahead vs 180 deg;
  TR_TFIL_TURN_REF_DEG (default 45) is where the penalty starts. Both
  default-off-effect: the default-path golden in test_tfil_commit_env.nim is
  unchanged and still passes.
- Turn cost is NEVER folded into the heat score. The filter stays hard.
- The draw stays random (j51 measured an argmin worse); every weight is
  floored at 1, so the pool can never be emptied and bias 0 is exactly the
  shipped uniform draw.
- |turn| now travels on the ScoredTile, and the commit log gained turn /
  minturn / promote so a caller can measure the regret of the draw.

Guards: 51 -> 66 checks (an absurd 99:1 bias never rescues an over-threshold
tile; mean |turn|, draw regret, >90 and mirror-side shares all fall; path
heat does not rise). env_report + .env.example updated.
2026-09-26 21:56:54 +02:00
SirStone 0e7e124c7f j144 ledger: the live A/B result, 600 battles, two independent blocks
Pooled verdict on the pre-registered rule is NOT DISTINGUISHABLE: the primary
cross-opponent sign test on round wins is 11/14, p=0.05737 for 'arrive', over the
0.05 line. Block 1 alone passed (p=0.01294); block 2 alone did not (p=0.0654).
Every delta is positive in both blocks for all three arms and damage/run is UP
on all three, so this reads as an under-powered null at the MDE boundary
(MDE 0.31 wins/run, observed 0.28) - but the verdict layer is not
reinterpreted, exactly as gate v1 was not.

The MECHANISM is established cleanly and replicates in both blocks: incoming hit
rate 18.07% -> 14.92%, sign-flip p=0.0013, CI [-5.29, -1.53] pp, damage taken
-24.57/run. That is precisely the pathology the owner watched.

Recommended for his own .env: TR_TFIL_COMMIT_ARRIVAL=1 and
TR_TFIL_NOREV_SPEED=4; leave TR_TFIL_COMMIT_MARGIN at 0 (weakest arm in both
blocks). Shipped default untouched: TR_MOVEMENT=strafe, all three new knobs
default off.
2026-09-26 21:43:03 +02:00
SirStone d2005abee9 j144 TFIL: arrival-based commitment + hysteresis + no mid-flight reversal
The owner's live-GUI report was correct on all four counts, and all four are
one bug: the commitment is cancelled by our own tile-boundary crossing
(96.1% of picks, 3793/3946, mean hold 5.06 ticks) while the bot is still
accelerating, and the picker is an unconstrained uniform draw over every
safe tile, so the new target can land in the mirror direction at |speed| < 4.

New knobs, all env-gated and default = today's behaviour (byte-for-byte
default parity guard re-run and green, 51 checks):
  TR_TFIL_COMMIT_ARRIVAL  hold the committed tile until we are ON it; the
                          tick knob becomes a MINIMUM dwell. 0 = shipped.
  TR_TFIL_COMMIT_MARGIN   leave only if the best alternative is at least
                          this much cooler on the same pathMaxHeat scale.
                          0 = shipped.
  TR_TFIL_NOREV_SPEED     while |speed| is below this, a mid-flight switch
                          may not take a tile >90 deg off the travel
                          direction. 0 = shipped. norevPool() never returns
                          an empty pool: with every candidate behind us it
                          takes the least-bad turn.

Offline gate (recorded DrussGT fixture, 20026 ticks): mean hold 4.1 -> 24.0
ticks, abandoned-before-arrival 92.8% -> 40.5%, committed tile actually
reached 3.3% -> 17.2%, opposite-direction slow mid-flight switches 394 -> 64
(-84%). 'TR_TFIL_TILE_REPLAN=off' alone - what cc11ede's arm B already tried -
only gets the hold to 13.6, which is why that A/B could not find this.

strafe is untouched: it imports only heatDecay/bulletMagScale/Pillar*, none
of which this touches. TR_MOVEMENT default stays strafe. Registered in
env_report.nim + knownEnvNames() + .env.example. Arms pre-registered in
docs/movement_campaign.md and tools/ab/arms_tfil_commit.txt.
2026-09-26 21:07:09 +02:00
SirStone 3903ec71b4 j134 ledger: note phantom_meteor is dead code (not selectable), out of scope 2026-09-26 13:01:36 +02:00
SirStone 3237b65f40 j134 ledger: propagated fire fix to all 5 movers via one shared fire_tracker (TR_FIRE_FIX), per-mover catch table 98.888->100, live tick alignment measured (correction belongs on event_tick+2) 2026-09-26 12:59:35 +02:00
SirStone 3bee4353b2 j133 ledger: appended 'Missed fires + the label question' (catch 98.888->100%, 746/67065 shots were blind: 456 by the server +3*power bonus, 290 by our own same-tick damage; exact label still -0.230, state-conditional model -0.347 -> the observable state is the constraint) + report fixtures 2026-09-26 12:22:31 +02:00
SirStone cc332138b3 j131 learned movement: real bullet-endpoint resolution (TR_LEARNED_REAL_EVENTS, default off) + exact-geometry Gate A/B (inversion NOT fixed; state still the constraint) 2026-09-26 11:54:11 +02:00
SirStone 8dd9b3b3b5 j130 learned movement outcome label: live panel results (no arm beats strafe; label swap is a dead heat; gap is information) + Gate A report 2026-09-26 11:35:47 +02:00
SirStone 61def1c3e9 j130 learned movement: outcome label (P(hit|state,g)) mode + Gate A; pre-registered outcome arms 2026-09-26 11:27:03 +02:00
SirStone a5c893ace6 learned movement: record the pre-registered prediction as partly wrong (right on wins, wrong on the hit-rate mechanism) 2026-09-26 10:56:02 +02:00
SirStone 4f332dcb95 learned movement (SBC): live panel results and the direct answer (module is a wash vs strafe on wins; state conditioning is a CI-separated hit-rate gain vs the same mover without it; the label is the wrong quantity) 2026-09-26 10:55:14 +02:00
SirStone ee98827a62 learned movement: danger-map alignment diagnostic (corr(P(arrival bin), P(hit|bin)) = -0.342) in the gate + ledger 2026-09-26 10:45:26 +02:00
SirStone a436e9f16a j128 learned movement: state-conditional counted-SBC wave danger (TR_MOVEMENT=learned, default-off), offline gate + pre-registered panel arms 2026-09-26 10:42:10 +02:00
SirStone 8e109bae0b closing j125: verify movement ship end-to-end (clean build SuccessX, full 16-test guard suite all counts match, real default+revert battles) and restructure gun/movement ledgers outcome-first 2026-09-26 05:06:49 +02:00
SirStone 3fd6db97e7 movement ship: gate v2 fresh-data primary passed (sign-flip p=0.045, CI [+0.02,+0.58]); default TR_MOVEMENT flipped to strafe 2026-09-26 03:19:33 +02:00
SirStone 514674886d movement gate v2: pre-register the sign-flip primary test on fresh 300-battle data (before any battle) 2026-09-26 02:44:22 +02:00
SirStone feefc1912e movement final: confirmation result and ship decision (gate failed on the sign-test leg; default NOT flipped) 2026-09-26 02:41:08 +02:00
SirStone ff03e81591 movement final: pre-register the ship criterion before the confirmation battles 2026-09-26 02:21:09 +02:00
SirStone 47ce244740 movement j119 Batch 3+4 results: the reversal/dwell and heat-field axes are a clean negative; nothing beats strafe 2026-09-26 02:18:22 +02:00
SirStone 1256357f07 movement j119 Batch 3+4: pre-register the reversal/dwell and heat-field arms 2026-09-26 01:40:03 +02:00
SirStone c07e6d9599 movement ledger: the post-hoc structure of the strafe win (60 points, both batches)
56/60 opponent-level win deltas are non-negative and no opponent family
regresses reproducibly (the one WallAvoider loss reverses in Batch 2).
corr(dwins, d_hit_rate) = -0.09, corr(dwins, d_damage) = +0.38: the aggregate
win is a survival effect, but a per-opponent hit-rate gain does not predict a
per-opponent win gain - so hit rate stays an explanation, never a proxy.
2026-09-26 01:33:24 +02:00
SirStone 44e2d191e3 movement ledger: fix two claims in the Batch-1/2 prose
- tfil is 4th of five on round wins, not last (ring is nominally 0.04 lower,
  ns) - the Batch-1 commit message overstates one word; the correction is
  recorded in the ledger rather than rewritten.
- the Batch-2 direct answer quoted three of four CIs excluding 0; it is four of
  four ([+0.04,+0.63], [+0.16,+0.60], [+0.27,+0.89], [+0.22,+0.72]).
- added the cleanest aggression isolation of Batch 1 (ring - ring_notemp, same
  engine and heat field, range weighting alone): +41.9 dmg/run, -0.11 wins/run,
  +12.8 pp incoming hit rate at 236 vs 395 px.
2026-09-26 01:32:28 +02:00
SirStone 7d3645da5a movement Batch 2: the strafe win over shipped tfil replicates; the range target decides nothing
Same frozen panel, same 3x3 design, new session on commit 8efa627 (no source file
changed since 1984a78, so the same code), 225 battles, 0 invalid runs. Arms:
tfil, strafe_notilt, strafe_325 + the tilt re-armed at 600px and 250px.

Paired vs tfil: strafe_325 +0.58 wins/run [CI +0.27,+0.89] 11/12 p=0.0063;
strafe_notilt +0.47 [+0.22,+0.72] 10/11 p=0.0117; tilt_600 +0.40 [+0.04,+0.76]
(sign test 8/11 p=0.23, sign-flip p=0.049); tilt_250 +0.38 [+0.07,+0.69] 10/12
p=0.039. Incoming hit rate -5.2..-6.9 pp with 0/15 opponents favouring tfil.

The range TARGET is not the lever: re-arming the tilt moved the achieved
distance from 459px (no steering) to 478px and 415px, and none of the three is
separable on wins. This overturns Batch 1's reading that the tilt costs wins -
the honest statement is that the tilt's win effect is below this design's
resolution. tfil reproduced to within 1.4 pp (40.7% -> 39.3% of rounds), so the
baseline itself is stable across sessions.

Ledger: Batch 2 section, the verbatim analyzer report, a data-driven
what-to-try-next, and the session log.
2026-09-26 01:31:08 +02:00
SirStone 07766303f5 movement Batch 1: pure strafe (range tilt OFF) beats the shipped tfil on round wins across a 15-opponent panel
225 battles, one frozen binary, five env-only arms, the frozen panel, 0 invalid
runs. Paired per opponent vs the shipped tfil:

  strafe_notilt  wins/run +0.38  [CI +0.16,+0.60]  9/9 opponents p=0.0039
                 dmg/run  -10.2  [CI -25.8,+5.5]   p=0.61, MDE 20.4 (not detectable)
                 incoming hit rate 12.24% vs 18.17%, dmg taken 150 vs 200
  strafe_325     wins/run +0.33  [CI +0.04,+0.63]  10/12 p=0.0386
  ring           dmg/run  +31.2  [CI +11.5,+50.9]  13/15 p=0.0074, wins/run -0.04 (ns)
                 but hit rate 29.4% at 236 px: a damage/survival trade, not a win
  ring_notemp    indistinguishable from tfil on both primaries

Round wins in this harness are survival wins (in 216/219 attributable runs the
win count equals the rounds the opponent died in), and the winner takes ~1/3
fewer hits while fighting ~74 px farther out. The shipped tfil is last of five
on wins: the DrussGT-only picture did not generalize.

Also: tournament_analyze.py now prints BOTH readings of the pre-registered
'while the other does not go down' clause (strict: nothing is better;
substantive: the two strafe arms and ring are better on one metric each).
2026-09-26 01:13:29 +02:00
SirStone 1984a780f4 melee A/B doc: correct the per-arm [bb-reset] counts (perRound 68.75, retained 68.19, decay 66.75) 2026-09-26 00:43:04 +02:00