feat(gun_harness): runtime metric switch + A/B proving the point metric mis-selects
Adds GUN_VBULLET_METRIC (point|path, default point = unchanged behaviour) so the virtual-bullet hit model can be selected at runtime with no rebuild. Both the live tracker and the offline replay read the same value, so the 12/12 offline==online acceptance holds under EITHER setting (verified for both). A/B AGAINST THE LIVE BOSS, real server-side hit rate as ground truth, 5 battles x 12 rounds per metric on one frozen binary: point 4660 shots / 219 hits = 4.70% (per-run 3.16-5.53) path 4834 shots / 359 hits = 7.43% (per-run 6.55-8.24) The distributions DO NOT OVERLAP: path's worst run beats point's best run. +2.73pp, +58% relative, z = 5.56, p < 0.0001. Range distributions were identical (~460-478 px), so this is not a range confound. MECHANISM - and this is the important part. The gain is SELECTION, not better gun learning. Under the point model every gun's virtual rate is compressed into 0.6-4.4%, so HeadOn sits inside the 2pp tie margin and takes 72.6% of selection ticks / 76.9% of shots - while HeadOn is 11th of 13 by REAL hit rate (2.3%). The path model widens the band to 4.7-13.7% and ranks HeadOn 10th, so its shot share falls to 35.9% and Pattern/Accel/WallBounce get picked instead. Counterfactual: applying the point model's per-gun real rates to the path model's shot mix yields 7.65%, i.e. essentially the whole observed gain. So the selector, not the guns, is where the win lives. PER-GUN REAL HIT RATE vs DrussGT (path mix, the answer to 'which guns are worth keeping'): WallBounce 10.8, Pattern 10.5, Accel 10.0, Displace 9.3, Circular 9.2, AvgLead 8.5, KNN 5.7, StopShot 5.2, GuessFactor 3.7, Tsetlin 2.9. Per-gun N is small (hundreds of shots) so single-gun ordering is indicative, not definitive. TWO CAVEATS, recorded because they undercut a naive reading: 1. One adversary. DrussGT is a wave surfer and HeadOn is genuinely bad against surfers, so part of this may be matchup-specific. 2. The path model is NOT a better general ranker. Spearman(virtual rank, real rank) is 0.52 under point vs -0.04 under path. It wins by accidentally fixing HeadOn's mis-rank, not by ranking guns better. A more durable fix is to address the selection logic directly - which is the next job. Also adds a focused guard test (test_vbullet_metric) covering parsing/default, a receding-target point-miss/path-hit, a perpendicular-target path-miss, and replay determinism. Verified: 33 guard checks, 12/12 acceptance under both metrics, tsetlin tests green, range 34.3% (point, unchanged) / 50.8% (path).
This commit is contained in:
@@ -218,7 +218,8 @@ proc reportFor(tracker: VirtualTracker, drivers: seq[GunDriver],
|
||||
result.add r
|
||||
|
||||
proc replayFixture*(fx: Fixture, drivers: seq[GunDriver],
|
||||
targetId = -1, liveActual = false): seq[GunReport] =
|
||||
targetId = -1, liveActual = false,
|
||||
metric = ActiveMetric): seq[GunReport] =
|
||||
## Drive a fresh `VirtualTracker` over the whole fixture, one tick at a time,
|
||||
## in the same order the live loop uses:
|
||||
## 1. predict(state, bulletSpeed(PowerBins[i])) for i = 0..3, per gun
|
||||
@@ -238,9 +239,12 @@ proc replayFixture*(fx: Fixture, drivers: seq[GunDriver],
|
||||
## resolution) is skipped entirely; the end marker tells us so and we drop
|
||||
## the final tick's resolutions. Synthetic fixtures leave liveActual false
|
||||
## (the state at the resolution tick is the ground truth).
|
||||
##
|
||||
## `metric` defaults to the process-wide `GUN_VBULLET_METRIC` switch read by
|
||||
## virtual_bullets; pass it explicitly only to force a model in one process.
|
||||
let tid = if targetId >= 0: targetId else: fx.enemyId
|
||||
let skipFinal = liveActual and fx.enemyDied
|
||||
var tracker = initTracker(drivers.len)
|
||||
var tracker = initTracker(drivers.len, metric)
|
||||
for si in 0..<fx.states.len:
|
||||
let state = fx.states[si]
|
||||
for gi in 0..<drivers.len:
|
||||
|
||||
Reference in New Issue
Block a user