Files
SirRoboGarage/common_libs/gun_harness/selector.nim
T
SirStone 589a230106 TM radial gun: registered (default OFF) + label-bias fix that removes the bias but
retracts its own earlier learning claim

=== TASK 1: REGISTERED AS GUN 14, DEFAULT `off` ===
The radial TM gun is now a first-class rack member (`TMPATTERN`, id 14), forceable
alone with `TR_RACK_TMPATTERN=both` plus every other `TR_RACK_*=off`.
DEFAULT IS `off`, and the justification matters: `both` would let it compete for
selection AND (because the shared VirtualTracker ring is order-sensitive) shift
every other gun's learning order, so it CANNOT leave the default path unchanged.
With `off` its predict and spawnBullets are additionally GATED on rack admission
(the only gun wired that way), so the shipped default never spawns it at all:
zero cost, zero ring perturbation.
Live proof: 1-round battle with only TMPATTERN racked ->
  `gun 14 (TMPattern): vShots=400 selected=104 other-gun selections=0`.
Default-path-unchanged proof: parity checks that the 15-gun default bestGun/
selectGun equals the old 14-gun rack RNG-draw-for-RNG-draw, that gun 14 is never
selected by default, and acceptance 12/12.
Cost: 0.36 ms/tick (predict 0.30 + onResult 0.05) ~= 3% of the 13.16 ms budget.
Tsetlin in the same harness is 1.62 ms/tick, so the new gun is ~4.5x cheaper.

=== TASK 2: THE LABEL-BIAS FIX - AND A RETRACTION ===
Root cause confirmed: under bmPoint a SHORT radial correction resolves the virtual
bullet BEFORE the base arrival tick, so the label was dropped (labelMisses).
Fix: defer the label in a pending queue and flush it once the arrival tick is
recorded; labels still come from the BASE arrival tick.
  labelMisses        4,281,695  ->  0
  training samples   1,071,824  ->  5,345,847  (x5)
  radial head acc         48.8% ->  57.0%   (shuffled control 20.0%)
  bmPoint hit rate     9.4/5.8% ->  9.1/5.7%  (unchanged, within noise)
So the fix IMPROVES LEARNING but NOT the metric.

**RETRACTION OF THE PREVIOUS JOB'S CLAIM.** It reported the radial head's 48.8%
against a 36.7% majority baseline and concluded "conditional learning, not a
constant bias". With the bias removed, the correctly-measured majority baseline is
**58.2%** - so the head at 57.0% is AT/BELOW majority. The earlier apparent
conditional learning was PARTLY AN ARTEFACT OF THE BIASED SAMPLE. The bmPoint
metric win is real (TMRadial > Linear early 16/2 p=0.0013, overall 18/0 p<0.0001;
> shuffled 18/0 p<0.0001) but it comes from a NET-POSITIVE AVERAGE RADIAL SHIFT,
not from beating a majority classifier. Recorded plainly rather than left standing.

Guards: test_tm_pattern_registration 20 (new), test_tm_pattern_rack_live 4 (new),
test_gun_harness 39, test_vbullet_metric 11, test_power_selection 3 (the SIGSEGV is
gone - the knn_gun rewrite is now committed), test_adaptive_radar 41,
test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 28,
test_rack_membership 38, test_selector_tiebreak 19, test_tm_pattern_learning 3,
acceptance_offline_vs_online 12/12. ModularBot compiles (release).

Note: `common_libs/tests/range_guns.nim` still builds 14 offline drivers (the
offline sweep constructs TmPatternGun directly and acceptance only inspects ids
0..13), so nothing breaks - but a future job wanting it in the offline rack must
add a 15th driver and mirror the live admission gating. gun_stats.jsonl now emits
15 rows; downstream tooling should ignore id 14.
2026-09-22 01:58:33 +02:00

211 lines
11 KiB
Nim
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
## Gun selector — picks best gun×power, computes aim angle, gates firing.
## Fires highest power with acceptable hit rate when the gun is aimed within a
## range-dependent angular tolerance and gunHeat == 0.
import std/math
import std/os
import std/strutils
import gun_interface
import virtual_bullets
# ── rack membership (TR_RACK_*) ──────────────────────────────────────────────
#
# Per-gun rack membership, read ONCE at process start so a single frozen binary
# can be re-racked without a rebuild — the same runtime pattern as
# GUN_RACK_DISABLE. A gun's membership admits it into the 1v1 rack, the melee
# rack, both, or neither:
#
# TR_RACK_HEADON=both (shipped default for every gun)
# TR_RACK_TSETLIN=1v1 -> 1v1 rack only
# TR_RACK_DISPLACE=melee -> melee rack only
# TR_RACK_KNN=off -> removed from both racks
# TR_RACK_TMPATTERN=off (shipped default for the new TM pattern gun)
#
# The mode itself is derived from SERVER truth (`getEnemyCount()`), never from
# the tracker's known-enemy count, by `rackMode` in virtual_bullets — the same
# transition the radar uses. Every gun except TMPATTERN defaults to `both`, so an
# unset environment preserves the pre-change single-rack selection byte-for-byte;
# TMPATTERN defaults to `off` so it cannot alter that selection.
const
RackGunNames*: array[15, string] = [
"HEADON", "LINEAR", "TSETLIN", "CIRCULAR", "GUESSFACTOR", "PATTERN",
"WALLBOUNCE", "ACCEL", "STOPSHOT", "DISPLACE", "AVGLEAD", "DECAYGF",
"KNN", "TMSELECT", "TMPATTERN"]
RackEnvPrefix* = "TR_RACK_"
## Defaults are all-`both` EXCEPT the new TM pattern gun (id 14), which ships
## `off`: it is registered and forceable (`TR_RACK_TMPATTERN=both|1v1|melee`)
## but never spawns a virtual bullet unless explicitly enabled, so the shared
## VirtualTracker ring head — and therefore every other gun's learning order
## and the default selection sequence — is byte-for-byte unchanged. Defaulting
## it to `both` would let it compete for selection and change the default rack.
DefaultRackMembership*: array[15, RackMembership] = [
rmBoth, rmBoth, rmBoth, rmBoth, rmBoth, rmBoth, rmBoth,
rmBoth, rmBoth, rmBoth, rmBoth, rmBoth, rmBoth, rmBoth,
rmOff]
proc parseRackMembership*(value: string): RackMembership =
## Parse a `TR_RACK_<GUN>` value. Empty / unknown values fall back to the
## shipped `both` and warn on stderr, so a typo cannot silently move a gun and
## a bad value cannot take the bot down.
case value.strip().toLowerAscii()
of "", "both", "any": rmBoth
of "1v1", "only1v1", "1v1only", "single", "lock": rmOnly1v1
of "melee", "onlymelee", "multi": rmOnlyMelee
of "off", "none", "disabled", "disable": rmOff
else:
stderr.writeLine("[gun_harness] unknown " & RackEnvPrefix & "<GUN>='" & value &
"'; falling back to 'both' (valid: both|1v1|melee|off)")
rmBoth
proc loadRackMembership*(): array[len(RackGunNames), RackMembership] =
## Default table plus every `TR_RACK_<GUN>` override. A proc (not inlined into
## the `let`) so the unit test can exercise env parsing in-process.
result = DefaultRackMembership
for i in 0..<len(RackGunNames):
let key = RackEnvPrefix & RackGunNames[i]
let v = getEnv(key, "")
if v.len > 0:
result[i] = parseRackMembership(v)
let ActiveRackMembership* = loadRackMembership()
## Process-wide rack table, frozen at startup.
proc rackMembershipName*(m: RackMembership): string =
case m
of rmBoth: "both"
of rmOnly1v1: "1v1"
of rmOnlyMelee: "melee"
of rmOff: "off"
proc rackModeName*(m: RackMode): string =
case m
of rm1v1: "1v1"
of rmMelee: "melee"
proc rackOverrides*(membership: openArray[RackMembership]): string =
## Compact `GUN:mode,GUN:mode` list of entries that differ from the shipped
## all-`both` default. Empty when the rack is at its default.
for i in 0..<min(len(RackGunNames), membership.len):
if membership[i] != rmBoth:
if result.len > 0: result.add ","
result.add RackGunNames[i] & ":" & rackMembershipName(membership[i])
proc rackActive*(membership: openArray[RackMembership], mode: RackMode): string =
## Comma-separated gun names admitted in `mode` (empty set prints as
## `FULL` — the graceful-degradation fallback).
for i in 0..<min(len(RackGunNames), membership.len):
if membership[i].admits(mode):
if result.len > 0: result.add ","
result.add RackGunNames[i]
if result.len == 0: result = "FULL"
const
## ── Range-aware firing gate ────────────────────────────────────────────────
## A real shot departs with whatever misalignment the gun had at fire time,
## while a virtual bullet is spawned exactly on the prediction and carries zero
## aim error. At distance `d` the target subtends an angular half-width of
## `atan(BotRadius / d)`, so a fixed degree threshold is simultaneously too
## loose at long range (throws away shots that cannot hit) and too tight up
## close (holds fire when the bot is already inside the hit cone).
##
## We therefore derive the tolerance from the target's angular radius:
##
## tolDeg = radToDeg(arctan(BotRadius * SafetyFactor / distPx))
##
## clamped to [MinAimThresholdDeg, MaxAimThresholdDeg].
##
## SafetyFactor shrinks/expands the accepted cone: 1.0 == the full geometric
## half-width, < 1.0 is stricter. Fitted empirically from real-shot data
## (Task A, 2611 real shots behind a wide-open 20 deg measurement gate).
## The geometric model is only weakly identified: prediction error dominates
## the hit rate, and the measured 50%-hit knee is noisy (0.9-1.4x the
## geometric cone at 200-800 px; the 400-600 px bucket is ill-defined because
## its baseline hit rate is already ~50%). Simulating the gate directly on the
## measurement data showed 0.6 Pareto-dominates the old fixed 2.0 deg gate
## (61.4% vs 60.2% hit rate with MORE shots), and the live sweep confirms the
## observed preference for tighter gates. 0.6 is the shipped compromise:
## tighter than the raw geometry while still loosening close range.
SafetyFactor* = 0.6
## Floor: keeps the tolerance strictly positive so a perfectly aligned gun can
## always fire at any range, and guards the gate against collapsing to 0
## (a never-fire deadlock) at extreme distances.
MinAimThresholdDeg* = 0.05
## Ceiling: at point-blank range the geometric cone grows without bound; a
## >10 deg misalignment is a coin toss even at ~100 px, so cap it here.
MaxAimThresholdDeg* = 10.0
proc aimToleranceDeg*(distPx: float): float =
## Angular half-width (deg) the gun may be off by and still plausibly hit a
## target `distPx` px away, scaled by SafetyFactor and clamped.
##
## Degenerate distances (0 or unavailable) fall back to the ceiling rather than
## dividing by zero; NaN is treated the same way (the `not (distPx > 0.0)`
## test is false for NaN). +Inf falls through to arctan(0) == 0 and then the
## floor, which is correct: an infinitely distant target is a point.
if not (distPx > 0.0): return MaxAimThresholdDeg
result = radToDeg(arctan(BotRadius * SafetyFactor / distPx))
if result < MinAimThresholdDeg: result = MinAimThresholdDeg
elif result > MaxAimThresholdDeg: result = MaxAimThresholdDeg
proc aimAngle*(selfX, selfY, targetX, targetY: float): float =
## Absolute bearing in degrees (0=East, CCW+) toward (targetX, targetY).
result = radToDeg(arctan2(targetY - selfY, targetX - selfX))
proc shouldFire*(currentGunDir, targetAngle, gunHeat, distPx: float): bool =
## Returns true when the gun is within the range-aware angular tolerance and
## cool enough to fire. `distPx` is the distance (px) to the aim point.
var delta = (targetAngle - currentGunDir) mod 360.0
if delta > 180.0: delta -= 360.0
elif delta < -180.0: delta += 360.0
abs(delta) <= aimToleranceDeg(distPx) and gunHeat <= 0.0
proc selectShotPolicy*(t: var VirtualTracker, targetId = -1, tick = 0,
dist = 0.0, selfEnergy = 100.0,
ramming = false,
rackMode: RackMode = rm1v1,
membership: openArray[RackMembership] = []
): (GunId, int, float, PowerCap) =
## `selectShot` plus the energy-aware power-policy decision, so a caller can
## log the cap and its reason (see `applyPowerPolicy` in virtual_bullets).
##
## `dist` is the current distance (px) to the target and `selfEnergy` our own
## energy; `ramming` exempts the caps (the movement code's `shouldRam` is the
## single source of truth). The policy is applied identically wherever this is
## called, so live and any offline caller cannot diverge.
##
## `rackMode` is the server-truth enemy-count mode (`rackMode`); `membership`
## is the process-wide `TR_RACK_*` table, passed by the live bot. An empty
## membership admits every gun (the pre-change behaviour).
let gunId = t.selectGun(targetId, tick,
rackMode = rackMode, membership = membership)
let (prefBin, preferred) = t.bestPower(gunId, targetId)
# pEst / pRef mirror `bestPower`'s own fitness source (per-target when data
# exists, else the deterministic aggregate). An empty bin carries no rate of
# its own, so it borrows the gun's aggregate — the same "no data" case the
# policy documents.
let fit = t.fitnessFor(targetId)
let pRef = if PowerRefFixed > 0.0: PowerRefFixed
else: gunRate(fit[gunId], pooled = true)
let pEst =
if fit[gunId].bins[prefBin].count == 0: pRef
else: fit[gunId].bins[prefBin].hitRate()
let dec = applyPowerPolicy(preferred, dist, selfEnergy, pEst, pRef, ramming)
result = (gunId, binIndexForPower(dec.power), dec.power, dec)
proc selectShot*(t: var VirtualTracker, targetId = -1, tick = 0,
dist = 0.0, selfEnergy = 100.0,
ramming = false,
rackMode: RackMode = rm1v1,
membership: openArray[RackMembership] = []): (GunId, int, float) =
## Returns (gunId, powerBinIdx, power) — the shot to take this tick.
## Pass targetId to pick the best gun for that specific enemy. `tick` drives
## the minimum-dwell hysteresis (see `selectGun`). `dist`/`selfEnergy`/`ramming`
## feed the energy-aware power cap (`TR_POWER_POLICY`); defaults keep every
## existing caller compiling, and `TR_POWER_POLICY=0` reproduces the uncapped
## `bestPower` preference. Use `selectShotPolicy` when the cap/reason is needed.
let (gunId, binIdx, power, _) =
t.selectShotPolicy(targetId, tick, dist, selfEnergy, ramming,
rackMode = rackMode, membership = membership)
result = (gunId, binIdx, power)