Files
SirRoboGarage/common_libs/gun_harness/selector.nim
T
SirStone b68707c867 Energy economy: the cliff becomes a SLOPE, plus a finishing cap. 11% less energy.
The user's request: "when our bot is low OR enemy is low, it is useless to use high
power instead low fast bullets have more chances to finish the enemy. Let's do a
math slope: starting from some health down, the power goes down with it."

1. ENERGY SLOPE (`TR_POWER_ENERGY_*`), replacing the old hard step at 50 energy:
   cap = ENERGY_MAX at/above ENERGY_HI, ENERGY_MIN at/below ENERGY_LO, LINEAR in
   power between, clamped. Defaults HI=80 LO=20 MIN=0.5 MAX=3.0, so no cap >=80,
   0.5 at <=20, and e.g. E=65 -> 2.375, E=50 -> 1.75, E=35 -> 1.125.
   Rationale: bullet speed is 20-3p, so lower power = FASTER bullet (less lead
   error, higher hit chance), fires more often (10+2p) and drains slower (p/shot).
   E[dE] = p(3P-1) => break-even hit probability is 1/3 INDEPENDENT of power, and
   our measured rates are 5-27%, far below it.

2. FINISHING CAP (`TR_POWER_FINISH_KILL`, default ON): cap power at the SMALLEST
   bullet that still removes the enemy's remaining energy -
   `E<=4 -> p=E/4` (min 0.1), `4<E<=16 -> p=(E+2)/6`, `E>16 -> no cap`.
   Rationale, and it makes the user's instinct stronger than a heuristic: server
   1.3.1 caps the damage SCORE at the energy ACTUALLY REMOVED, so overkill is
   WASTED damage AND ~6x the energy for ZERO extra score. Damage is 4p (p<=1) /
   6p-2 (p>1).

Both are min-composed with the existing far/below-average caps, may only LOWER
power (exhaustively tested), and are exempt while ramming.
`TR_POWER_POLICY=0` still returns the uncapped control exactly.

MEASURED ENERGY SAVING (offline replay of the DrussGT fixtures, 28,797 ticks):
  arm              shots  energy  meanP  E/1k ticks   vs cliff
  control(uncapped) 1913    4646   2.43    161.4      -90.2%
  cliff (today)     2363    2443   1.03     84.8       0.0%
  slope             2404    2178   0.91     75.6    ** 10.9% LESS **
  slope+finish      2404    2167   0.90     75.3    ** 11.3% LESS **
So the slope spends ~11% less energy than the cliff AND fires slightly MORE shots
(2404 vs 2363) - both directions at once.

HONEST NOTE on the finishing rule's reach here: ticks where the enemy is low
(0 < E <= 16) are only 2252/28797 = 7.8% of these fixtures, so finishing adds just
~11 energy of saving against DrussGT. It matters in CLOSER fights, not this one.

Verification: test_power_policy 58 (was 26) in BOTH the default and TR_POWER_POLICY=0
control arms - slope at E=100/80/65/50/35/20/5, powerToKill across E=0.1..100, the
inverse-cover property for E<=16, monotonicity, ram exemption, and an exhaustive
sweep proving power <= preference. Guards: test_gun_harness 39, test_vbullet_metric
11, test_power_selection 3, test_adaptive_radar 41, test_tfil_ring_weights 24,
test_ram_decision 40, test_rack_membership 48, test_selector_tiebreak 19,
test_tm_pattern_registration 20, test_vbullet_admit_gate 12. acceptance
12/12 PASS. ModularBot compiles release.

Adds `common_libs/tests/measure_power_policy.nim` (the energy/histogram tool) and
updates docs/env_reference.md for the new `energySlope|finishKill` log reasons.

NOT MEASURED: the battle/hit-rate effect. The offline figures use the fixture
shooter's energy as a proxy, open-loop; the RELATIVE saving is the meaningful part.
2026-09-23 00:12:04 +02:00

248 lines
13 KiB
Nim
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
## Gun selector — picks best gun×power, computes aim angle, gates firing.
## Fires highest power with acceptable hit rate when the gun is aimed within a
## range-dependent angular tolerance and gunHeat == 0.
import std/math
import std/os
import std/strutils
import gun_interface
import virtual_bullets
# ── rack membership (TR_RACK_*) ──────────────────────────────────────────────
#
# Per-gun rack membership, read ONCE at process start so a single frozen binary
# can be re-racked without a rebuild — the same runtime pattern as
# GUN_RACK_DISABLE. A gun's membership admits it into the 1v1 rack, the melee
# rack, both, or neither:
#
# TR_RACK_PATTERN=both (shipped default: the ONLY admitted gun)
# TR_RACK_HEADON=both -> re-admit HeadOn (used to restore the old rack)
# TR_RACK_TSETLIN=1v1 -> 1v1 rack only
# TR_RACK_DISPLACE=melee -> melee rack only
# TR_RACK_KNN=off -> removed from both racks
# TR_RACK_TMPATTERN=off (shipped default for the new TM pattern gun)
#
# SHIPPED DEFAULT IS `onlyPattern`: Pattern (id 5) is admitted in both racks and
# every other gun is `off`. This is a deliberate, measured decision, not a
# pruning heuristic — the virtual-fitness selector was measured to be NEGATIVE
# value at every rack size tested (full, lean8, lean6, pairPC/PK/PL) and against
# 10/10 adversaries, while Pattern alone is the best single gun in general. See
# docs/selector_negative_value.md. The selector MECHANISM is retained in full
# (chooseFromFit, the floor/band logic, hysteresis, virtual fitness) — the rack
# merely has one member by default, so re-enabling any gun is a one-line env
# override with no rebuild:
#
# Revert to the old full rack (all guns `both`, TMPATTERN `off`):
# TR_RACK_PATTERN=both TR_RACK_HEADON=both TR_RACK_LINEAR=both \
# TR_RACK_TSETLIN=both TR_RACK_CIRCULAR=both TR_RACK_GUESSFACTOR=both \
# TR_RACK_WALLBOUNCE=both TR_RACK_ACCEL=both TR_RACK_STOPSHOT=both \
# TR_RACK_DISPLACE=both TR_RACK_AVGLEAD=both TR_RACK_DECAYGF=both \
# TR_RACK_KNN=both TR_RACK_TMSELECT=both ./ModularBot
#
# The mode itself is derived from SERVER truth (`getEnemyCount()`), never from
# the tracker's known-enemy count, by `rackMode` in virtual_bullets — the same
# transition the radar uses.
const
RackGunNames*: array[16, string] = [
"HEADON", "LINEAR", "TSETLIN", "CIRCULAR", "GUESSFACTOR", "PATTERN",
"WALLBOUNCE", "ACCEL", "STOPSHOT", "DISPLACE", "AVGLEAD", "DECAYGF",
"KNN", "TMSELECT", "TMPATTERN", "TMHORIZON"]
RackEnvPrefix* = "TR_RACK_"
## SHIPPED DEFAULT: `onlyPattern`. Pattern (id 5) is admitted in both racks;
## every other gun is `off`. The selection mechanism is untouched and remains
## fully functional — only the rack's membership changed. Re-enable any gun
## with `TR_RACK_<GUN>`, or restore the old full rack with the one-liner in the
## header comment. TMPATTERN (id 14) stays `off`: registered and forceable but
## it never spawns a virtual bullet unless explicitly enabled, so the shared
## VirtualTracker ring head — and every other gun's learning order — is
## unchanged.
DefaultRackMembership*: array[16, RackMembership] = [
rmOff, # 0 HEADON — off (measured: worst over-selected gun)
rmOff, # 1 LINEAR — off
rmOff, # 2 TSETLIN — off
rmOff, # 3 CIRCULAR — off
rmOff, # 4 GUESSFACTOR — off
rmBoth, # 5 PATTERN — the only admitted gun (best single gun in general)
rmOff, # 6 WALLBOUNCE — off
rmOff, # 7 ACCEL — off
rmOff, # 8 STOPSHOT — off
rmOff, # 9 DISPLACE — off
rmOff, # 10 AVGLEAD — off
rmOff, # 11 DECAYGF — off
rmOff, # 12 KNN — off
rmOff, # 13 TMSELECT — off
rmOff, # 14 TMPATTERN — off (already shipped off; TM pattern gun)
rmOff] # 15 TMHORIZON — off (horizon-based TM corrector; expected to lose)
proc parseRackMembership*(value: string): RackMembership =
## Parse a `TR_RACK_<GUN>` value. Empty / unknown values fall back to the
## shipped `both` and warn on stderr, so a typo cannot silently move a gun and
## a bad value cannot take the bot down.
case value.strip().toLowerAscii()
of "", "both", "any": rmBoth
of "1v1", "only1v1", "1v1only", "single", "lock": rmOnly1v1
of "melee", "onlymelee", "multi": rmOnlyMelee
of "off", "none", "disabled", "disable": rmOff
else:
stderr.writeLine("[gun_harness] unknown " & RackEnvPrefix & "<GUN>='" & value &
"'; falling back to 'both' (valid: both|1v1|melee|off)")
rmBoth
proc loadRackMembership*(): array[len(RackGunNames), RackMembership] =
## Default table plus every `TR_RACK_<GUN>` override. A proc (not inlined into
## the `let`) so the unit test can exercise env parsing in-process.
result = DefaultRackMembership
for i in 0..<len(RackGunNames):
let key = RackEnvPrefix & RackGunNames[i]
let v = getEnv(key, "")
if v.len > 0:
result[i] = parseRackMembership(v)
let ActiveRackMembership* = loadRackMembership()
## Process-wide rack table, frozen at startup.
proc rackMembershipName*(m: RackMembership): string =
case m
of rmBoth: "both"
of rmOnly1v1: "1v1"
of rmOnlyMelee: "melee"
of rmOff: "off"
proc rackModeName*(m: RackMode): string =
case m
of rm1v1: "1v1"
of rmMelee: "melee"
proc rackOverrides*(membership: openArray[RackMembership]): string =
## Compact `GUN:mode,GUN:mode` list of entries that differ from the shipped
## default table. Empty when the rack is at its default.
for i in 0..<min(len(RackGunNames), membership.len):
if membership[i] != rmBoth:
if result.len > 0: result.add ","
result.add RackGunNames[i] & ":" & rackMembershipName(membership[i])
proc rackActive*(membership: openArray[RackMembership], mode: RackMode): string =
## Comma-separated gun names admitted in `mode` (empty set prints as
## `FULL` — the graceful-degradation fallback).
for i in 0..<min(len(RackGunNames), membership.len):
if membership[i].admits(mode):
if result.len > 0: result.add ","
result.add RackGunNames[i]
if result.len == 0: result = "FULL"
const
## ── Range-aware firing gate ────────────────────────────────────────────────
## A real shot departs with whatever misalignment the gun had at fire time,
## while a virtual bullet is spawned exactly on the prediction and carries zero
## aim error. At distance `d` the target subtends an angular half-width of
## `atan(BotRadius / d)`, so a fixed degree threshold is simultaneously too
## loose at long range (throws away shots that cannot hit) and too tight up
## close (holds fire when the bot is already inside the hit cone).
##
## We therefore derive the tolerance from the target's angular radius:
##
## tolDeg = radToDeg(arctan(BotRadius * SafetyFactor / distPx))
##
## clamped to [MinAimThresholdDeg, MaxAimThresholdDeg].
##
## SafetyFactor shrinks/expands the accepted cone: 1.0 == the full geometric
## half-width, < 1.0 is stricter. Fitted empirically from real-shot data
## (Task A, 2611 real shots behind a wide-open 20 deg measurement gate).
## The geometric model is only weakly identified: prediction error dominates
## the hit rate, and the measured 50%-hit knee is noisy (0.9-1.4x the
## geometric cone at 200-800 px; the 400-600 px bucket is ill-defined because
## its baseline hit rate is already ~50%). Simulating the gate directly on the
## measurement data showed 0.6 Pareto-dominates the old fixed 2.0 deg gate
## (61.4% vs 60.2% hit rate with MORE shots), and the live sweep confirms the
## observed preference for tighter gates. 0.6 is the shipped compromise:
## tighter than the raw geometry while still loosening close range.
SafetyFactor* = 0.6
## Floor: keeps the tolerance strictly positive so a perfectly aligned gun can
## always fire at any range, and guards the gate against collapsing to 0
## (a never-fire deadlock) at extreme distances.
MinAimThresholdDeg* = 0.05
## Ceiling: at point-blank range the geometric cone grows without bound; a
## >10 deg misalignment is a coin toss even at ~100 px, so cap it here.
MaxAimThresholdDeg* = 10.0
proc aimToleranceDeg*(distPx: float): float =
## Angular half-width (deg) the gun may be off by and still plausibly hit a
## target `distPx` px away, scaled by SafetyFactor and clamped.
##
## Degenerate distances (0 or unavailable) fall back to the ceiling rather than
## dividing by zero; NaN is treated the same way (the `not (distPx > 0.0)`
## test is false for NaN). +Inf falls through to arctan(0) == 0 and then the
## floor, which is correct: an infinitely distant target is a point.
if not (distPx > 0.0): return MaxAimThresholdDeg
result = radToDeg(arctan(BotRadius * SafetyFactor / distPx))
if result < MinAimThresholdDeg: result = MinAimThresholdDeg
elif result > MaxAimThresholdDeg: result = MaxAimThresholdDeg
proc aimAngle*(selfX, selfY, targetX, targetY: float): float =
## Absolute bearing in degrees (0=East, CCW+) toward (targetX, targetY).
result = radToDeg(arctan2(targetY - selfY, targetX - selfX))
proc shouldFire*(currentGunDir, targetAngle, gunHeat, distPx: float): bool =
## Returns true when the gun is within the range-aware angular tolerance and
## cool enough to fire. `distPx` is the distance (px) to the aim point.
var delta = (targetAngle - currentGunDir) mod 360.0
if delta > 180.0: delta -= 360.0
elif delta < -180.0: delta += 360.0
abs(delta) <= aimToleranceDeg(distPx) and gunHeat <= 0.0
proc selectShotPolicy*(t: var VirtualTracker, targetId = -1, tick = 0,
dist = 0.0, selfEnergy = 100.0,
enemyEnergy = 100.0,
ramming = false,
rackMode: RackMode = rm1v1,
membership: openArray[RackMembership] = []
): (GunId, int, float, PowerCap) =
## `selectShot` plus the energy-aware power-policy decision, so a caller can
## log the cap and its reason (see `applyPowerPolicy` in virtual_bullets).
##
## `dist` is the current distance (px) to the target, `selfEnergy` our own
## energy and `enemyEnergy` the target's remaining energy (drives the
## finishing cap); `ramming` exempts the caps (the movement code's `shouldRam`
## is the single source of truth). The policy is applied identically wherever
## this is called, so live and any offline caller cannot diverge.
##
## `rackMode` is the server-truth enemy-count mode (`rackMode`); `membership`
## is the process-wide `TR_RACK_*` table, passed by the live bot. An empty
## membership admits every gun (the pre-change behaviour).
let gunId = t.selectGun(targetId, tick,
rackMode = rackMode, membership = membership)
let (prefBin, preferred) = t.bestPower(gunId, targetId)
# pEst / pRef mirror `bestPower`'s own fitness source (per-target when data
# exists, else the deterministic aggregate). An empty bin carries no rate of
# its own, so it borrows the gun's aggregate — the same "no data" case the
# policy documents.
let fit = t.fitnessFor(targetId)
let pRef = if PowerRefFixed > 0.0: PowerRefFixed
else: gunRate(fit[gunId], pooled = true)
let pEst =
if fit[gunId].bins[prefBin].count == 0: pRef
else: fit[gunId].bins[prefBin].hitRate()
let dec = applyPowerPolicy(preferred, dist, selfEnergy, pEst, pRef, ramming,
enemyEnergy = enemyEnergy)
result = (gunId, binIndexForPower(dec.power), dec.power, dec)
proc selectShot*(t: var VirtualTracker, targetId = -1, tick = 0,
dist = 0.0, selfEnergy = 100.0,
enemyEnergy = 100.0,
ramming = false,
rackMode: RackMode = rm1v1,
membership: openArray[RackMembership] = []): (GunId, int, float) =
## Returns (gunId, powerBinIdx, power) — the shot to take this tick.
## Pass targetId to pick the best gun for that specific enemy. `tick` drives
## the minimum-dwell hysteresis (see `selectGun`). `dist`/`selfEnergy`/
## `enemyEnergy`/`ramming` feed the energy-aware power cap (`TR_POWER_POLICY`);
## defaults keep every existing caller compiling, and `TR_POWER_POLICY=0`
## reproduces the uncapped `bestPower` preference. Use `selectShotPolicy` when
## the cap/reason is needed.
let (gunId, binIdx, power, _) =
t.selectShotPolicy(targetId, tick, dist, selfEnergy,
enemyEnergy = enemyEnergy, ramming = ramming,
rackMode = rackMode, membership = membership)
result = (gunId, binIdx, power)