## Virtual bullet tracker. ## Spawns virtual bullets per gun×power bin every tick (no real firing). ## Resolves by travel distance. Rolling window fitness per gun×power. ## Calls onResult() on the owning gun when a bullet resolves. import std/math import std/tables import std/random import std/algorithm import std/os import std/strutils import gun_interface const PowerBins* = [1.0, 1.5, 2.0, 3.0] ## 4 bins; ponytail: fixed array, add runtime config if needed WindowSize* = 100 ## rolling window ticks for fitness MaxBullets* = 8192 ## hard cap; ring buffer. 52 spawns/tick and a ## full-map long shot (~90 ticks) need ~4700 slots; ## 8192 wraps only after ~157 ticks. Each VirtualBullet ## is ~120 bytes, so this array costs ~960 KiB. MinHitRate* = 0.40 ## LEGACY absolute bar; no longer used by bestPower ## (no bin on the live path-metric scale cleared it, so ## once every bin had data bestPower fell to power 1.0). PowerBarFrac* = 0.50 ## RELATIVE power bar (dimensionless): a bin is ## acceptable when its virtual hit rate is at least this ## FRACTION of the same gun's best bin rate. Scales with ## the metric instead of assuming a ~40% hit rate. MinObsBeforeCompete* = 50 ## min observations before a gun×bin enters competition TieMargin* = 0.02 ## ABSOLUTE mode: guns within this hit-rate margin of best are tied MinHitRateFloor* = 0.10 ## ABSOLUTE mode: if best gun < this, fall back to gun 0 (HeadOn) RelTieMargin* = 0.20 ## RELATIVE mode: tied if rate >= bestRate*(1-this). Dimensionless ## fraction of the best rate, so it scales with the metric. FloorPeakFrac* = 0.25 ## RELATIVE mode: floor fires if bestRate < this*peakRateRef. ## Dimensionless: only if the field collapsed vs its own recent best. SelectorWindow* = 256 ## ticks of per-tick bestRate kept for the RELATIVE floor reference MetricEnvVar* = "GUN_VBULLET_METRIC" ## RUNTIME switch selecting how a virtual bullet is scored. Read once per ## process at module init, so the SAME compiled binary can be A/B'd by ## exporting it — no rebuild needed. Both the live ModularBot tracker and ## the offline range replay call `initTracker`, so they always agree. type GunId* = int ## index into the guns seq BulletMetric* = enum bmPoint ## A bullet is scored at the single point it reaches at the ## fire-time aim distance. HIT iff that point is within BotRadius of ## the target on that tick. Measures prediction accuracy (does the ## bullet arrive at the predicted point at the right time). bmPath ## DEFAULT. The bullet flies along its straight ray until it leaves ## the arena. Each tick the swept segment (previous -> new position) ## is tested against the target's radius; HIT iff ANY segment came ## within BotRadius. Measures hypothetical hit chance against the ## target's real path. Chosen by the DrussGT A/B: 7.2-7.6% real hit ## rate vs 3.8% for point (p<0.0001). const DefaultMetric* = bmPath ## Shipped virtual-bullet scoring model. `GUN_VBULLET_METRIC` overrides it at ## runtime; an unset OR empty value means this default. proc parseMetric*(value: string): BulletMetric = ## Parse a `GUN_VBULLET_METRIC` value. Empty / unknown values fall back to ## the shipped `DefaultMetric` and emit a one-line warning on stderr, so a ## typo can never silently change the metric and a bad value can never take ## the bot down. case value.strip().toLowerAscii() of "", "default": DefaultMetric of "point", "points", "bmpoint": bmPoint of "path", "paths", "bmpath": bmPath else: stderr.writeLine("[gun_harness] unknown " & MetricEnvVar & "='" & value & "'; falling back to '" & $DefaultMetric & "' (valid: point|path)") DefaultMetric let ActiveMetric* = parseMetric(getEnv(MetricEnvVar, "")) ## The metric every tracker uses unless a caller overrides it explicitly in ## `initTracker`. Frozen at process start from the environment. const SelectorModeEnvVar* = "GUN_SELECTOR_MODE" ## RUNTIME switch selecting the selection-threshold model. Read once per ## process, so one binary can A/B both (§ virtual_bullets). type SelectorMode* = enum smAbsolute ## legacy: fixed 2pp tie band + 10% absolute floor. Correct only ## if the virtual hit-rate scale happens to land near 10%. smRelative ## scale-aware: tie band is a fraction of the best rate; the floor ## fires only when the field has collapsed vs its own recent peak. proc parseSelectorMode*(value: string): SelectorMode = ## Empty / unknown values fall back to the shipped `relative` model and warn. case value.strip().toLowerAscii() of "", "relative", "rel": smRelative of "absolute", "abs", "legacy": smAbsolute else: stderr.writeLine("[gun_harness] unknown " & SelectorModeEnvVar & "='" & value & "'; falling back to 'relative' (valid: absolute|relative)") smRelative let ActiveSelectorMode* = parseSelectorMode(getEnv(SelectorModeEnvVar, "relative")) # ── runtime ranking knobs (A/B without rebuilding) ──────────────────────────── # # Every knob below defaults to the SHIPPED constant, so an unset environment # reproduces the shipped behaviour byte-for-byte. They exist so one compiled # binary can be swept across candidate ranking rules. The measured A/B found no # candidate that credibly beats the shipped statistic: keep the defaults below # unless a new adversary/run set changes that. type RankStat* = enum rsMean ## plain window mean (SHIPPED) rsWilson ## Wilson lower confidence bound (z=1); penalises small n rsUCB ## mean + c*se (exploration bonus) rsThompson ## one Normal-approx Beta(h+1,n-h+1) sample per gun (Thompson) rsShrunk ## empirical-Bayes shrink toward the field mean proc envInt(name: string, default: int): int = let v = getEnv(name, "") if v.len == 0: return default try: parseInt(v.strip()) except ValueError: default proc envFloat(name: string, default: float): float = let v = getEnv(name, "") if v.len == 0: return default try: parseFloat(v.strip()) except ValueError: default proc envBool(name: string, default: bool): bool = case getEnv(name, "").strip().toLowerAscii() of "1", "true", "yes", "on": true of "0", "false", "no", "off": false else: default proc parseRank(value: string): RankStat = case value.strip().toLowerAscii() of "", "mean", "avg": rsMean of "wilson", "lcb": rsWilson of "ucb": rsUCB of "thompson", "ts": rsThompson of "shrunk", "shrink", "eb": rsShrunk else: stderr.writeLine("[gun_harness] unknown GUN_SELECTOR_RANK='" & value & "'; falling back to 'mean' (valid: mean|wilson|ucb|thompson|shrunk)") rsMean let ActiveWindow* = clamp(envInt("GUN_SELECTOR_WINDOW", WindowSize), 1, WindowSize) let ActiveMinObs* = max(1, envInt("GUN_SELECTOR_MINOBS", MinObsBeforeCompete)) let ActiveRelTie* = envFloat("GUN_SELECTOR_TIE", RelTieMargin) let ActiveFloorFrac* = envFloat("GUN_SELECTOR_FLOOR", FloorPeakFrac) let ActivePooled* = envBool("GUN_SELECTOR_POOL", true) let ActiveRank* = parseRank(getEnv("GUN_SELECTOR_RANK", "")) let ActiveShrink* = max(0.0, envFloat("GUN_SELECTOR_SHRINK", 20.0)) # Optional per-process seed so independent A/B runs use independent tie-breaks # (Nim's default rand() stream is identical in every process, which would make # "random" tie-breaks repeat across runs). Unset => leave the RNG untouched. block: let s = getEnv("GUN_SELECTOR_SEED", "") if s.len > 0: try: randomize(parseInt(s.strip())) except ValueError: discard type VirtualBullet* = object gunId*: GunId powerBin*: int ## index into PowerBins targetId*: int ## enemy bot ID this bullet was aimed at fireTick*: int ## tick this bullet was spawned; lets a gun pair its ## predict() trace with the exact resolution event fireX*, fireY*: float aimX*, aimY*: float ## predicted target (absolute) bulletSpeed*: float travelDist*: float ## accumulated px so far fireDist*: float ## distance to target at fire time active*: bool # --- path-metric bookkeeping (unused by the point metric) --- hitSeen*: bool ## a swept segment already touched the target bestMissDist*: float ## closest segment->target distance seen so far bestMissX*: float ## target position at that closest approach bestMissY*: float FitnessWindow* = object ## Ring buffer of hit booleans. hits*: array[WindowSize, bool] count*: int ## total samples so far (capped at WindowSize for rate) head*: int GunFitness* = object bins*: array[len(PowerBins), FitnessWindow] SelectorDiag* = object ## Optional observability for `chooseFromFit`/`bestGun`. Never needed by the ## bot; lets the offline range report WHY a gun was selected (floor vs tie). bestRate*: float ## max hit rate over eligible guns (the floor comparison value) floorFired*: bool ## bestRate below the active floor -> returned gun 0 tiedCount*: int ## eligible guns within the tie band of bestRate (0 if floor fired) anyQualifies*: bool ## at least one gun reached MinObsBeforeCompete floorRate*: float ## the floor actually applied this tick VirtualTracker* = object bullets*: array[MaxBullets, VirtualBullet] head*: int ## ring buffer head numGuns*: int metric*: BulletMetric ## scoring model (defaults to ActiveMetric) fitness*: Table[int, seq[GunFitness]] ## keyed by enemy bot ID, indexed by GunId droppedBullets*: int ## unresolved bullets clobbered by the ring buffer (should stay 0) # RELATIVE-mode floor reference: per-tick bestRate history + its running max. rateHist*: array[SelectorWindow, float] rateHistHead*: int rateHistCount*: int peakRateRef*: float ## max bestRate in rateHist; 0.0 = not enough history yet proc initTracker*(numGuns: int, metric = ActiveMetric): VirtualTracker = ## `metric` defaults to the process-wide `GUN_VBULLET_METRIC` switch; pass it ## explicitly only from tests that need both models in one process. result.numGuns = numGuns result.metric = metric proc windowHits(fw: FitnessWindow, want: int): int = ## Hits among the most recent `want` samples, in ring order. Reading the last ## `want` slots (head backwards) is what makes a runtime window shorter than ## `WindowSize` correct even after the ring has wrapped. let n = min(fw.count, min(want, WindowSize)) for k in 1..n: let idx = (fw.head - k + WindowSize) mod WindowSize if fw.hits[idx]: inc result proc hitRate*(fw: FitnessWindow): float = ## Returns fraction of hits in the rolling window. 0.0 when no data. ## Uses `ActiveWindow` (defaults to `WindowSize`). if fw.count == 0: return 0.0 let n = min(fw.count, min(ActiveWindow, WindowSize)) if n == 0: return 0.0 result = windowHits(fw, n).float / n.float proc record(fw: var FitnessWindow, hit: bool) = fw.hits[fw.head] = hit fw.head = (fw.head + 1) mod WindowSize inc fw.count proc spawnBullets*(t: var VirtualTracker, gunId: GunId, predictions: array[len(PowerBins), GunPrediction], state: WorldState, targetId: int) = ## Call once per gun per tick with predictions for all power bins. ## Lazily creates fitness entry for targetId on first spawn. if targetId notin t.fitness: t.fitness[targetId] = newSeq[GunFitness](t.numGuns) for binIdx in 0.. 1e-12: s = clamp(((px - ax)*abx + (py - ay)*aby) / abLen2, 0.0, 1.0) hypot(px - (ax + s*abx), py - (ay + s*aby)) # ── rate helpers (shared by the selector and the reference tracker) ────────── proc gunEligible*(fit: GunFitness, requireMin: bool): bool = ## True if a gun has >= MinObsBeforeCompete samples in at least one bin (or ## when `requireMin` is false, every gun is eligible). if not requireMin: return true for binIdx in 0..= ActiveMinObs: return true false proc gunCounts*(fit: GunFitness, pooled: bool): tuple[hits, n: int] = ## Sample counts behind `gunRate`. `pooled` sums all power bins; otherwise the ## single bin with the best rate (the "specialist" view). if pooled: for binIdx in 0.. best: best = r result = (h, m) proc gunRate*(fit: GunFitness, pooled: bool): float = ## A gun's hit rate. `pooled` sums hits/shots across all power bins (more ## samples, immune to one lucky bin); otherwise the max single-bin rate. let (h, n) = gunCounts(fit, pooled) result = if n > 0: h.float / n.float else: 0.0 proc rankScore(fit: GunFitness, pooled: bool, stat: RankStat, fieldRate, shrinkK: float): float = ## Ranking statistic over the gun's window. All are monotone-ish in the mean, ## but differ in how they trade mean against sample size/noise. `rsMean` is the ## shipped statistic. let (h, n) = gunCounts(fit, pooled) if n == 0: return 0.0 let p = h.float / n.float case stat of rsMean: p of rsWilson: let z = 1.0 let z2 = z * z let denom = 1.0 + z2 / n.float let centre = p + z2 / (2.0 * n.float) let margin = z * sqrt((p * (1.0 - p) + z2 / (4.0 * n.float)) / n.float) max(0.0, (centre - margin) / denom) of rsUCB: p + 0.5 * sqrt(p * (1.0 - p) / n.float) of rsThompson: let se = sqrt(max(1e-9, p * (1.0 - p) / n.float)) clamp(p + gauss(0.0, se), 0.0, 1.0) of rsShrunk: (h.float + shrinkK * fieldRate) / (n.float + shrinkK) proc tableBestRate(t: VirtualTracker, pooled: bool): float = ## Best eligible gun rate across every target's fitness (no merge/allocation). ## Only guns with >= MinObsBeforeCompete samples count: under-sampled bins ## produce 100%/-looking spikes that would inflate the floor reference and ## force HeadOn for the whole window. Zero when nothing is warmed up yet. result = 0.0 for _, perEnemy in t.fitness: for gunId in 0.. peak: peak = t.rateHist[i] t.peakRateRef = peak proc tickBullets*(t: var VirtualTracker, state: WorldState, enemies: Table[int, tuple[x, y: float, lastSeenTick: int, alive: bool]], onResolved: proc(gunId: GunId, binIdx: int, e: FeedbackEvent)) = ## Advance all active bullets one tick. ## ## The `bmPoint` branch (default) is unchanged: resolve when the bullet ## reaches the fire-time aim distance and score the single point it lands on. ## ## The `bmPath` branch flies the bullet along its ray until it leaves the ## arena and tests each tick's swept segment against the target's radius. It ## records exactly one outcome per bullet (at the wall), so every resolved ## bullet contributes exactly one fitness sample. A bullet that goes dead or ## stale is discarded without scoring, mirroring the point metric at ## resolution time. for i in 0.. StaleTicks: b.active = false continue ex = e.x; ey = e.y else: # No data for this target — fall back to selected enemy in state ex = state.enemyX; ey = state.enemyY let dx = b.aimX - b.fireX let dy = b.aimY - b.fireY let dist = hypot(dx, dy) let (bx, by) = if dist < 1e-6: (b.aimX, b.aimY) else: (b.fireX + dx / dist * b.travelDist, b.fireY + dy / dist * b.travelDist) let missDist = hypot(bx - ex, by - ey) let hit = missDist < BotRadius if b.targetId in t.fitness: t.fitness[b.targetId][b.gunId].bins[b.powerBin].record(hit) let fe = FeedbackEvent( prediction: GunPrediction(x: b.aimX, y: b.aimY), actualX: ex, actualY: ey, bulletPower: PowerBins[b.powerBin], fireTick: b.fireTick, powerBin: b.powerBin, missDistance: missDist, hit: hit, ) onResolved(b.gunId, b.powerBin, fe) b.active = false of bmPath: # Enemy pose for this tick. A dead/stale target abandons the bullet # without scoring, exactly as the point metric does at resolution time. var ex, ey: float if b.targetId in enemies: let e = enemies[b.targetId] if not e.alive or (state.tick - e.lastSeenTick) > StaleTicks: b.active = false continue ex = e.x; ey = e.y else: ex = state.enemyX; ey = state.enemyY let dx = b.aimX - b.fireX let dy = b.aimY - b.fireY let dist = hypot(dx, dy) var ux, uy: float if dist < 1e-6: ux = 0.0; uy = 0.0 else: ux = dx / dist; uy = dy / dist let prevD = max(0.0, b.travelDist - b.bulletSpeed) let ax = b.fireX + ux * prevD let ay = b.fireY + uy * prevD let bx = b.fireX + ux * b.travelDist let by = b.fireY + uy * b.travelDist let segMiss = distPointToSegment(ex, ey, ax, ay, bx, by) if not b.hitSeen: if segMiss < BotRadius: # First physical contact — freeze it so a later closer approach # cannot overwrite the contact position the guns learn from. b.hitSeen = true b.bestMissDist = segMiss b.bestMissX = ex b.bestMissY = ey elif segMiss < b.bestMissDist: b.bestMissDist = segMiss b.bestMissX = ex b.bestMissY = ey # Despawn only at a wall (a degenerate zero-length ray also ends here). let outside = dist < 1e-6 or bx < 0.0 or bx > state.arenaWidth or by < 0.0 or by > state.arenaHeight if outside: let missDist = if b.bestMissDist == Inf: segMiss else: b.bestMissDist let rx = if b.bestMissDist == Inf: ex else: b.bestMissX let ry = if b.bestMissDist == Inf: ey else: b.bestMissY if b.targetId in t.fitness: t.fitness[b.targetId][b.gunId].bins[b.powerBin].record(b.hitSeen) let fe = FeedbackEvent( prediction: GunPrediction(x: b.aimX, y: b.aimY), actualX: rx, actualY: ry, bulletPower: PowerBins[b.powerBin], fireTick: b.fireTick, powerBin: b.powerBin, missDistance: missDist, hit: b.hitSeen, ) onResolved(b.gunId, b.powerBin, fe) b.active = false if ActiveSelectorMode == smRelative: noteBestRate(t) proc fitnessFor*(t: VirtualTracker, targetId: int): seq[GunFitness] = ## Returns fitness seq for targetId, or merges all enemies as fallback. ## ## The fallback is a RECENCY-WEIGHTED AGGREGATE over the last WindowSize ## samples, NOT a pooled rate: each per-enemy window is replayed into one fresh ## window, so once the total exceeds WindowSize the earliest samples are ## overwritten by later ones. Enemies are visited in ascending target-id order ## so the result is identical on every run (std/tables iteration order is hash ## order and therefore nondeterministic). ## ponytail: merge is O(enemies*guns*bins*WindowSize), fine for small counts if targetId >= 0 and targetId in t.fitness: return t.fitness[targetId] # Aggregate across all enemies, deterministically ordered. result = newSeq[GunFitness](t.numGuns) var enemyIds: seq[int] for id in t.fitness.keys: enemyIds.add id enemyIds.sort() for id in enemyIds: let perEnemy = t.fitness[id] for gunId in 0..= PowerBarFrac * bestBinRate`, dimensionless) — not ## against the legacy absolute `MinHitRate`. On the live path-metric scale a ## gun's rates sit around 3-40%, so the absolute 40% bar never fired once every ## bin had data and bestPower silently collapsed to power 1.0; the relative bar ## discriminates between bins at any scale. ## ## An EMPTY bin is still handed out (highest power first) so every bin keeps ## getting sampled, and a fully cold gun (no data anywhere) returns the lowest ## power bin. Uses per-enemy fitness when targetId >= 0 and data exists; else ## the deterministic aggregate. let fit = t.fitnessFor(targetId) result = (0, PowerBins[0]) var anyObs = false var bestRate = 0.0 for binIdx in 0.. 0: anyObs = true bestRate = max(bestRate, fit[gunId].bins[binIdx].hitRate()) if not anyObs: return (0, PowerBins[0]) let bar = PowerBarFrac * bestRate for binIdx in countdown(len(PowerBins) - 1, 0): let fw = fit[gunId].bins[binIdx] if fw.count == 0 or fw.hitRate() >= bar: return (binIdx, PowerBins[binIdx]) proc chooseFromFit*(fit: seq[GunFitness], diag: ptr SelectorDiag = nil, mode: SelectorMode = smAbsolute, referenceRate = -1.0): GunId = ## Core gun ranking over an already-resolved fitness seq. Split out from ## `bestGun` so the offline range can rank without copying a VirtualTracker, ## and so callers can request `diag` for the selection internals. ## ## Guns with fewer than MinObsBeforeCompete observations are skipped unless ## every gun is below threshold (then fall back to best of all). ## ## `mode` chooses the threshold model: ## smAbsolute — legacy fixed TieMargin / MinHitRateFloor. ## smRelative — tie band = bestRate*RelTieMargin; floor = FloorPeakFrac ## * `referenceRate` (the recent field-best rate). Pooled over ## power bins, since one lucky bin is a poor ranker. ## `referenceRate` <= 0 disables the RELATIVE floor (no history yet). ## Ties (within the band) are broken randomly to avoid index-0 bias. let pooled = if mode == smRelative: ActivePooled else: false var anyQualifies = false for gunId in 0.. 0: fieldSum / fieldN.float else: 0.0 if diag != nil: diag[].bestRate = bestRate let floorRate = if mode == smAbsolute: MinHitRateFloor elif referenceRate > 0.0: ActiveFloorFrac * referenceRate else: 0.0 if diag != nil: diag[].floorRate = floorRate # No hit at all, or the field collapsed below its own recent peak: HeadOn. if bestRate <= 0.0 or bestRate < floorRate: if diag != nil: diag[].floorFired = true return 0 # Ranking scores (the active statistic) and the best of them. var scores = newSeq[float](fit.len) var bestScore = 0.0 for gunId in 0..= bestScore - tieBand: tied.add(gunId) if diag != nil: diag[].tiedCount = tied.len if tied.len == 0: return 0 result = tied[rand(tied.len - 1)] proc bestGun*(t: VirtualTracker, targetId: int = -1, diag: ptr SelectorDiag = nil): GunId = ## Pick gun with highest hit rate across all power bins. ## Uses per-enemy fitness when targetId >= 0 and data exists; else aggregate. ## `diag`, when non-nil, receives the selection internals (bestRate, floor, ## tie count) exactly as used by the decision. result = chooseFromFit(t.fitnessFor(targetId), diag, mode = ActiveSelectorMode, referenceRate = t.peakRateRef)