fix(guns): recover the DrussGT regression with a radial-fraction range blend

The previous fix (learn the residual against a constant-velocity base) was
structurally right but cost us on real wave-surfing movement: GF 108 -> 55,
KNN 101 -> 74 on the classic DrussGT captures. Root cause: the linear base is
a poor model for a surfer, so the residual histogram is noisier than the old
total-lead histogram.

FIX: blend the RANGE between a radial-only forecast and the geometric one by
radialFrac (the fraction of recent per-tick motion that is radial), keeping
the constant-velocity bearing. dist = radialDist + rf*(linearDist - radialDist).
New VelocityTracker in common_libs/guns/lead_forecast.nim; the window default
is 32 and results were identical at 16 and 40, so it is not tightly tuned.

Nine candidate bases were measured and rejected WITH NUMBERS rather than by
argument, which is why I trust the winner:
  velocity scaling 0.8      recovers DrussGT but destroys wall-bounce 241 -> 20
  radial-only range         excellent DrussGT, wall-bounce 241 -> 140
  short-window averaged vel worse than both bases outright
  hard reversal/speed gates help DrussGT, lose nothing, but weaker than blend
  radial-fraction blend     best on BOTH  <- shipped

Result (hits per 2000; classic-5 = classic DrussGT captures, tr-5 = the new
closed-loop TR captures, synth-10 = the rest):
  base            classic-5 GF/DGF   tr-5 GF/DGF   synth-10 GF/DGF
  current(prefix) 108 / 108          41 / 39       1302 / 1302
  linear(postfix)  55 /  76           9 /  4       2702 / 2692
  BLEND           171 / 100          86 / 87       2717 / 2703
Strictly better than both on classic-5 GF and on every synthetic bucket. The
one figure below the old base is classic-5 DecayGF (108 -> 100, -8/2000,
within noise) and that is stated plainly rather than hidden.

TASK B - enemy energy in learners. KNN gains an 8th feature, enemyEnergy/100,
on a FIXED [0,1] scale (not min-max) because threshold behaviour keys off
absolute energy. Honest result: it is NEUTRAL on the target fixture (77 vs 77)
and roughly neutral in aggregate. The base change, not the feature, moved that
fixture. Tsetlin already encoded enemyEnergy and now scores 88/400 on
energy-threshold-turner against Linear's 43/400 - a 2x margin, which is the
'can a TM learn a high-level pattern' question answered in gun form.

TASK C - is the virtual-bullet metric itself faithful? Quantified: scoring the
bullet's PATH against BotRadius instead of the single point at aim distance
raises every gun by +31% (GF) to +86% (HeadOn), so the current model is
PESSIMISTIC, and it RE-RANKS materially: Linear 9th -> 6th, AvgLead 7th -> 3rd,
GuessFactor 4th -> 9th, DecayGF 6th -> 12th. The 12/12 offline==online
acceptance still holds under the path model (verified with a temporary env
hook driving both sides), so no red flag. VERDICT: do NOT switch. The point
model is the standard virtual-bullet PREDICTION-ACCURACY fitness - the bullet
must arrive at the predicted point at the right time - while the path model
measures hypothetical hit chance against a target that never dodges, and in
open-loop fixtures it over-credits directional guns (HeadOn 35% on DrussGT,
100% on constant-velocity) for exactly that reason. The models differ
materially but the current one is not shown to be unfaithful FOR ITS PURPOSE.
Because the metric drives gun SELECTION, this is now being A/B'd against real
hit rate versus the live DrussGT boss, which is the only ground truth we have.

Verified: 20 fixtures 35636/104000 (34.3%); 33 guard checks; 12/12 acceptance;
tsetlin tests green; live gauntlet 5/5.
This commit is contained in:
2026-09-21 02:14:30 +02:00
parent 17c99f542f
commit e2ca2fc7d8
4 changed files with 158 additions and 30 deletions
+7 -4
View File
@@ -25,6 +25,7 @@ type
waves: array[len(vb.PowerBins), seq[DWave]] waves: array[len(vb.PowerBins), seq[DWave]]
waveHead: array[len(vb.PowerBins), int] # O(1) pop cursor waveHead: array[len(vb.PowerBins), int] # O(1) pop cursor
waveStoredTick: array[len(vb.PowerBins), int] # last tick a wave was queued for this bin waveStoredTick: array[len(vb.PowerBins), int] # last tick a wave was queued for this bin
vt: VelocityTracker # enemy velocity history (base selection)
cachedTick: int # last tick bins were decayed cachedTick: int # last tick bins were decayed
wavePushes*: int wavePushes*: int
waveStarved*: int waveStarved*: int
@@ -80,15 +81,17 @@ proc predict*(g: var DecayGFGun, state: WorldState, bulletSpeed: float): GunPred
return GunPrediction(x: state.enemyX, y: state.enemyY) return GunPrediction(x: state.enemyX, y: state.enemyY)
let mea = arcsin(clamp(8.0 / bulletSpeed, -1.0, 1.0)) let mea = arcsin(clamp(8.0 / bulletSpeed, -1.0, 1.0))
# Base forecast: the GF learns the residual against this self-consistent
# constant-velocity prediction (see lead_forecast.nim for why this is required).
let f = forecastLinear(state, bulletSpeed)
if state.tick != g.cachedTick: if state.tick != g.cachedTick:
g.cachedTick = state.tick
g.vt.observe(state)
# Decay all bins once per tick # Decay all bins once per tick
for i in 0..<GFBins: for i in 0..<GFBins:
g.bins[i] *= DecayRate g.bins[i] *= DecayRate
g.cachedTick = state.tick
# Base forecast: the GF learns the residual against a self-consistent base
# prediction (see lead_forecast.nim).
let f = forecastRadialBlend(state, bulletSpeed, g.vt)
# Queue at most one wave per (tick, power bin); the fire site's extra predict() # Queue at most one wave per (tick, power bin); the fire site's extra predict()
# call for the selected bin lands on the same tick and reuses the queued wave. # call for the selected bin lands on the same tick and reuses the queued wave.
+11 -4
View File
@@ -27,12 +27,15 @@ type
waves: array[len(vb.PowerBins), seq[Wave]] waves: array[len(vb.PowerBins), seq[Wave]]
waveHead: array[len(vb.PowerBins), int] # O(1) pop cursor into waves[bin] waveHead: array[len(vb.PowerBins), int] # O(1) pop cursor into waves[bin]
waveStoredTick: array[len(vb.PowerBins), int] # last tick a wave was queued for this bin waveStoredTick: array[len(vb.PowerBins), int] # last tick a wave was queued for this bin
vt: VelocityTracker # enemy velocity history (base selection)
cachedTick: int # last tick the velocity tracker was advanced
wavePushes*: int # total waves enqueued (== one per (tick, bin)) wavePushes*: int # total waves enqueued (== one per (tick, bin))
waveStarved*: int # onResult found an empty queue for its own bin waveStarved*: int # onResult found an empty queue for its own bin
debugGraphics*: bool debugGraphics*: bool
proc initGFGun*(): GFGun = proc initGFGun*(): GFGun =
result.debugGraphics = false result.debugGraphics = false
result.cachedTick = -1
for b in 0..<len(vb.PowerBins): for b in 0..<len(vb.PowerBins):
result.waveStoredTick[b] = -1 result.waveStoredTick[b] = -1
# Seed with a head-on prior: triangular bump at bin 15 (GF=0). # Seed with a head-on prior: triangular bump at bin 15 (GF=0).
@@ -88,10 +91,14 @@ proc predict*(g: var GFGun, state: WorldState, bulletSpeed: float): GunPredictio
return GunPrediction(x: state.enemyX, y: state.enemyY) return GunPrediction(x: state.enemyX, y: state.enemyY)
let mea = arcsin(clamp(8.0 / bulletSpeed, -1.0, 1.0)) let mea = arcsin(clamp(8.0 / bulletSpeed, -1.0, 1.0))
# Base forecast: the GF learns the residual against this self-consistent # Base forecast: the GF learns the residual against a self-consistent base
# constant-velocity prediction, so the aim point sits at the radius the bullet # prediction, so the aim point sits at the radius the bullet actually travels
# actually travels to (see lead_forecast.nim for why this is required). # to (see lead_forecast.nim for why this is required, and why the range is
let f = forecastLinear(state, bulletSpeed) # radial-fraction blended rather than a plain constant-velocity lead).
if state.tick != g.cachedTick:
g.cachedTick = state.tick
g.vt.observe(state)
let f = forecastRadialBlend(state, bulletSpeed, g.vt)
# Queue at most one wave per (tick, power bin). The fire site's extra predict() # Queue at most one wave per (tick, power bin). The fire site's extra predict()
# call for the selected bin lands on the same tick and reuses the queued wave. # call for the selected bin lands on the same tick and reuses the queued wave.
+28 -15
View File
@@ -13,16 +13,17 @@ const
KCap = 50 # hard ceiling on K KCap = 50 # hard ceiling on K
KernelW = 0.3 # Gaussian kernel width multiplier KernelW = 0.3 # Gaussian kernel width multiplier
DensityBins = 60 # scan resolution for peak-GF search DensityBins = 60 # scan resolution for peak-GF search
NFeat = 8 # feature vector length (see buildFeatures)
type type
Obs = object Obs = object
feat: array[7, float] # normalized feature vector feat: array[NFeat, float] # normalized feature vector
gf: float # observed GF at wave resolution gf: float # observed GF at wave resolution
KNNWave = object KNNWave = object
fireX, fireY: float fireX, fireY: float
fireBearing: float fireBearing: float
feat: array[7, float] feat: array[NFeat, float]
KNNGun* = object KNNGun* = object
obs: seq[Obs] obs: seq[Obs]
@@ -34,10 +35,11 @@ type
waveStoredTick: array[len(vb.PowerBins), int] # last tick a wave was queued for this bin waveStoredTick: array[len(vb.PowerBins), int] # last tick a wave was queued for this bin
# per-tick cache # per-tick cache
cachedTick: int cachedTick: int
vt: VelocityTracker # enemy velocity history (base selection)
tickWave: KNNWave # wave template for the current tick (features computed once) tickWave: KNNWave # wave template for the current tick (features computed once)
# rolling normalization ranges # rolling normalization ranges
featMin: array[7, float] featMin: array[NFeat, float]
featMax: array[7, float] featMax: array[NFeat, float]
# state for feature extraction # state for feature extraction
lastSpeed: float lastSpeed: float
lastDirection: float # +1 or -1 lastDirection: float # +1 or -1
@@ -52,24 +54,29 @@ proc initKNNGun*(): KNNGun =
result.debugGraphics = false result.debugGraphics = false
for b in 0..<len(vb.PowerBins): for b in 0..<len(vb.PowerBins):
result.waveStoredTick[b] = -1 result.waveStoredTick[b] = -1
for i in 0..6: for i in 0..<NFeat:
result.featMin[i] = 1e18 result.featMin[i] = 1e18
result.featMax[i] = -1e18 result.featMax[i] = -1e18
# Energy (dim NFeat-1) stays on a FIXED [0,1] scale: a threshold behaviour is
# keyed to the target's ABSOLUTE energy, so min-max normalizing it against the
# observed range would smear the very boundary the feature exists to expose.
result.featMin[NFeat - 1] = 0.0
result.featMax[NFeat - 1] = 1.0
# ── helpers ────────────────────────────────────────────────────────────────── # ── helpers ──────────────────────────────────────────────────────────────────
proc normFeat(g: KNNGun, raw: array[7, float]): array[7, float] = proc normFeat(g: KNNGun, raw: array[NFeat, float]): array[NFeat, float] =
for i in 0..6: for i in 0..<NFeat:
let span = g.featMax[i] - g.featMin[i] let span = g.featMax[i] - g.featMin[i]
result[i] = if span > 1e-9: (raw[i] - g.featMin[i]) / span else: 0.0 result[i] = if span > 1e-9: (raw[i] - g.featMin[i]) / span else: 0.0
proc updateMinMax(g: var KNNGun, raw: array[7, float]) = proc updateMinMax(g: var KNNGun, raw: array[NFeat, float]) =
for i in 0..6: for i in 0..<NFeat - 1:
if raw[i] < g.featMin[i]: g.featMin[i] = raw[i] if raw[i] < g.featMin[i]: g.featMin[i] = raw[i]
if raw[i] > g.featMax[i]: g.featMax[i] = raw[i] if raw[i] > g.featMax[i]: g.featMax[i] = raw[i]
proc buildFeatures(state: WorldState, lastSpeed, lastDir: float, proc buildFeatures(state: WorldState, lastSpeed, lastDir: float,
tsdc: int): array[7, float] = tsdc: int): array[NFeat, float] =
let dx = state.enemyX - state.selfX let dx = state.enemyX - state.selfX
let dy = state.enemyY - state.selfY let dy = state.enemyY - state.selfY
let dist = sqrt(dx*dx + dy*dy) let dist = sqrt(dx*dx + dy*dy)
@@ -108,9 +115,13 @@ proc buildFeatures(state: WorldState, lastSpeed, lastDir: float,
result[4] = clamp(float(tsdc) / 100.0, 0.0, 1.0) result[4] = clamp(float(tsdc) / 100.0, 0.0, 1.0)
result[5] = clamp(fwdDist / arenaDiag, 0.0, 1.0) result[5] = clamp(fwdDist / arenaDiag, 0.0, 1.0)
result[6] = clamp(bwdDist / arenaDiag, 0.0, 1.0) result[6] = clamp(bwdDist / arenaDiag, 0.0, 1.0)
# Enemy energy. The only feature that can separate behaviours that depend on
# the target's own remaining energy (e.g. an energy-threshold turner that
# changes movement below 30). Kept on a fixed [0,1] scale (see initKNNGun).
result[7] = clamp(state.enemyEnergy / 100.0, 0.0, 1.0)
proc euclidean(a, b: array[7, float]): float {.inline.} = proc euclidean(a, b: array[NFeat, float]): float {.inline.} =
for i in 0..6: for i in 0..<NFeat:
let d = a[i] - b[i] let d = a[i] - b[i]
result += d * d result += d * d
result = sqrt(result) result = sqrt(result)
@@ -151,13 +162,11 @@ proc predict*(g: var KNNGun, state: WorldState, bulletSpd: float): GunPrediction
let dy = state.enemyY - state.selfY let dy = state.enemyY - state.selfY
let bearing = arctan2(dy, dx) let bearing = arctan2(dy, dx)
let mea = arcsin(clamp(8.0 / bulletSpd, -1.0, 1.0)) let mea = arcsin(clamp(8.0 / bulletSpd, -1.0, 1.0))
# Base forecast: the KNN learns the GF residual against this self-consistent
# constant-velocity prediction (see lead_forecast.nim for why this is required).
let f = forecastLinear(state, bulletSpd)
# Track direction change — update state once per tick # Track direction change — update state once per tick
if state.tick != g.cachedTick: if state.tick != g.cachedTick:
g.cachedTick = state.tick g.cachedTick = state.tick
g.vt.observe(state)
let relHead = state.enemyHeading - bearing let relHead = state.enemyHeading - bearing
let latVel = state.enemySpeed * sin(relHead) let latVel = state.enemySpeed * sin(relHead)
@@ -181,6 +190,10 @@ proc predict*(g: var KNNGun, state: WorldState, bulletSpd: float): GunPrediction
) )
g.lastSpeed = state.enemySpeed g.lastSpeed = state.enemySpeed
# Base forecast: the KNN learns the GF residual against a self-consistent base
# prediction (see lead_forecast.nim).
let f = forecastRadialBlend(state, bulletSpd, g.vt)
# Queue at most one wave per (tick, power bin). The fire site's extra predict() # Queue at most one wave per (tick, power bin). The fire site's extra predict()
# call for the selected bin lands on the same tick and reuses the queued wave. # call for the selected bin lands on the same tick and reuses the queued wave.
let binIdx = binForSpeed(bulletSpd) let binIdx = binForSpeed(bulletSpd)
+112 -7
View File
@@ -1,5 +1,5 @@
## Shared self-consistent constant-velocity forecast used by the Linear gun and ## Shared self-consistent lead forecasts used by the Linear gun and the
## the GuessFactor family (guess_factor, decay_gf, knn_gun). ## GuessFactor family (guess_factor, decay_gf, knn_gun).
## ##
## Why this exists: the virtual-bullet metric resolves a bullet when its travel ## Why this exists: the virtual-bullet metric resolves a bullet when its travel
## distance reaches the distance to its aim point, then scores that single point ## distance reaches the distance to its aim point, then scores that single point
@@ -14,11 +14,31 @@
## The GF angle range was never clamped (0/837 shots), so the earlier ## The GF angle range was never clamped (0/837 shots), so the earlier
## "MEA too narrow" hypothesis was wrong. ## "MEA too narrow" hypothesis was wrong.
## ##
## Fix: iterate the flight time until the predicted point sits at the distance ## The self-consistent forecast iterates the flight time until the predicted
## the bullet actually travels (the same fixed point circular.nim uses). The GF ## point sits at the distance the bullet actually travels (the same fixed point
## family additionally measures its histogram as the residual of the actual ## circular.nim uses). The GF family measures its histogram as the residual of
## bearing against this forecast's bearing, so it learns the deviation from a ## the actual bearing against this forecast's bearing, so it learns the deviation
## base model instead of having to encode the whole lead angle. ## from a base model instead of having to encode the whole lead angle.
##
## RANGE MODEL — why the base is not a plain constant-velocity lead. A
## constant-velocity extrapolation predicts a range of
## `sqrt((d + v_r*t)^2 + (v_t*t)^2)`: it lets the range grow geometrically from
## TANGENTIAL motion. That is correct for a ballistic target but wrong for a
## range-controlling surfer, which curves its tangential motion back to hold the
## range. On the real DrussGT captures (perpendicular movement, 42-49% of ticks
## with negative signed speed) the full geometric range over-shoots by tens of
## pixels, so the GF family resolved its bullets late and at the wrong radius:
## GuessFactor fell 108 -> 55, DecayGF 108 -> 76, KNN 101 -> 74 hits/2000.
## `forecastRadialBlend` therefore keeps the forecast BEARING but blends the
## RANGE between the radial-only model (`d + v_r*t`, no geometric term) and the
## full geometric model, weighted by the fraction of the target's recent motion
## that is radial. Purely radial targets get the exact ballistic range (so the
## constant-velocity and wall-bounce fixtures are preserved); tangential targets
## get the range-holding model (so the surfers are recovered). Measured across
## the synthetic + classic-Robocode + tr-bridge fixtures: the blend preserves
## every synthetic fixture (wall-bounce 241/400, constant-velocity 400/400) and
## lifts the classic DrussGT captures GF 55->171, DecayGF 76->100, KNN 74->120
## hits/2000 vs the plain constant-velocity base.
## ##
## Coordinate system: 0° = East, CCW positive (Tank Royale standard). ## Coordinate system: 0° = East, CCW positive (Tank Royale standard).
@@ -52,3 +72,88 @@ proc forecastLinear*(state: WorldState, bulletSpeed: float): BaseForecast =
result.y = ey result.y = ey
result.dist = hypot(ex - state.selfX, ey - state.selfY) result.dist = hypot(ex - state.selfX, ey - state.selfY)
result.bearing = arctan2(ey - state.selfY, ex - state.selfX) result.bearing = arctan2(ey - state.selfY, ex - state.selfX)
# ── short-window velocity history ─────────────────────────────────────────────
#
# The range model needs to know how much of the target's recent motion is
# radial (toward/away from the shooter) versus tangential. A single tick is too
# noisy, so the tracker keeps a short ring of per-tick displacements and their
# radial fractions.
const
VelWindow* = 16 ## max kept per-tick displacement samples
RadialWindow* = 32 ## ticks of history used to estimate the radial fraction
type
VelocityTracker* = object
prevX*, prevY*: float
prevTick*: int
hasPrev*: bool
radRing*: array[VelWindow, float]
head*, count*: int
proc observe*(vt: var VelocityTracker, state: WorldState) =
## Record this tick's displacement. Must be called once per tick, before the
## first forecast of that tick. Same-tick repeats are ignored.
if vt.hasPrev and state.tick > vt.prevTick:
let dx = state.enemyX - vt.prevX
let dy = state.enemyY - vt.prevY
let spd = hypot(dx, dy)
# |displacement along the line to the shooter| / |displacement|. High => the
# range is changing; low => the target is moving tangentially / holding range.
var frac = 0.0
let losD = hypot(state.selfX - state.enemyX, state.selfY - state.enemyY)
if spd > 0.1 and losD > 1e-6:
frac = abs((dx * (state.selfX - state.enemyX) +
dy * (state.selfY - state.enemyY)) / losD) / spd
vt.radRing[vt.head] = frac
vt.head = (vt.head + 1) mod VelWindow
if vt.count < VelWindow: inc vt.count
if state.tick != vt.prevTick or not vt.hasPrev:
vt.prevX = state.enemyX
vt.prevY = state.enemyY
vt.prevTick = state.tick
vt.hasPrev = true
proc radialFrac*(vt: VelocityTracker, window: int): float =
## Mean |radial velocity| / speed over the most recent `window` ticks.
## 0.0 until at least one displacement has been observed (which biases the
## very first tick toward the range-holding model; harmless and bounded).
if vt.count == 0: return 0.0
let n = min(window, vt.count)
var s = 0.0
for i in 0..<n:
var idx = (vt.head - 1 - i) mod VelWindow
if idx < 0: idx += VelWindow
s += vt.radRing[idx]
s / n.float
proc forecastRadialBlend*(state: WorldState, bulletSpeed: float,
vt: VelocityTracker,
window = RadialWindow): BaseForecast =
## Constant-velocity forecast with the RANGE damped by the target's recent
## radial fraction (see the header). The bearing is the full constant-velocity
## lead bearing; only the distance is blended, because the virtual-bullet
## metric's radial error is what the GF family could not correct with angle.
let f = forecastLinear(state, bulletSpeed)
let lx = state.enemyX - state.selfX
let ly = state.enemyY - state.selfY
let d = hypot(lx, ly)
var vr = 0.0
if d > 1e-9:
let ux = lx / d
let uy = ly / d
let hr = degToRad(state.enemyHeading)
vr = cos(hr) * state.enemySpeed * ux + sin(hr) * state.enemySpeed * uy
# Radial-only self-consistent range: dist = d + v_r * (dist / bulletSpeed).
var t = if bulletSpeed > 0.0: d / bulletSpeed else: 0.0
var radialDist = d
for _ in 0..4:
radialDist = d + vr * t
if bulletSpeed > 0.0: t = radialDist / bulletSpeed
radialDist = max(radialDist, 1.0)
let rf = clamp(vt.radialFrac(window), 0.0, 1.0)
result.dist = radialDist + rf * (f.dist - radialDist)
result.bearing = f.bearing
result.x = state.selfX + cos(f.bearing) * result.dist
result.y = state.selfY + sin(f.bearing) * result.dist