power policy: cap power by range and energy, gate 3.0 on above-average chances
Implements the user's energy management request: "firing from more than 200px should be a 'not good chances zone' so faster bullets and more chances to hit matters more than single hit damage with low chances. When we are lower than 50 health, same thing. I would like to use 3.0 power only when the chances of hitting are higher than average." Design: a CAP on top of the existing `bestPower`, not a rewrite. `bestPower` still answers "which bin does this gun's own data prefer"; the policy caps it: ramming -> 3.0 (reason ram, exempt) dist > TR_POWER_FAR_DIST (200) -> 1.0 (far) elif selfEnergy < TR_POWER_LOW_ENERGY (50) -> 1.0 (lowEnergy) elif pEst <= pRef -> 2.0 (belowAvg) else -> 3.0 (full) power = min(gunPreferredBinPower, cap) # can only LOWER power p=1.0 is the right "low" tier on the measured mechanics: bullet speed 20-3p so p=1.0 gives speed 17 vs 11 at p=3.0 (55% faster = less lead error), fire interval 10+2p so 12 ticks vs 16 (33% more shots), and drain 0.083/turn vs 0.1875 (2.25x slower). All three things the user asked for at long range. pEst = the chosen bin's virtual rate (gun aggregate when the bin is empty); pRef = the gun's aggregate mean unless TR_POWER_REF > 0. No-data guns are vacuously below-average -> cap 2.0 (conservative, documented). Control arm: TR_POWER_POLICY=0 = uncapped = today's behaviour exactly. Knobs: TR_POWER_POLICY, TR_POWER_FAR_DIST, TR_POWER_LOW_ENERGY, TR_POWER_FAR_CAP, TR_POWER_MID_CAP, TR_POWER_REF, TR_POWER_LOG. TR_POWER_MID_CAP exists because the user did not specify the middle case (close + healthy + not-above-average); 2.0 is the default, flippable to 1.0. Seam: the cap lives in a pure `applyPowerPolicy` and is applied only in `selectShot` (the single place real shots are chosen), so the logic is testable without a battle. Ram is wired from `shouldRam` - the same value the movement dispatch uses for the (0,50) band. CORRECTION TO AN ASSUMPTION IN THE TASK: `offline_range.nim` does NOT call `bestPower`/`selectShot` - it only replays virtual-bullet spawn/resolve across all power bins, independent of the real shot's power. So there is no offline power-selection path that could diverge from the live one, and the acceptance test guards the metric, not the policy. Policy coverage therefore comes from the new unit test. Verification: test_power_policy 26/26 in BOTH modes (default and TR_POWER_POLICY=0 control arm); test_gun_harness 39, test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41, test_tfil_ring_weights 24; acceptance_offline_vs_online 12/12 VERDICT PASS (live battle). ModularBot compiles. UNVERIFIED: the live effect on damage/survival/score. No A/B has run.
This commit is contained in:
@@ -66,10 +66,40 @@ proc shouldFire*(currentGunDir, targetAngle, gunHeat, distPx: float): bool =
|
||||
elif delta < -180.0: delta += 360.0
|
||||
abs(delta) <= aimToleranceDeg(distPx) and gunHeat <= 0.0
|
||||
|
||||
proc selectShot*(t: var VirtualTracker, targetId: int = -1, tick = 0): (GunId, int, float) =
|
||||
proc selectShotPolicy*(t: var VirtualTracker, targetId = -1, tick = 0,
|
||||
dist = 0.0, selfEnergy = 100.0,
|
||||
ramming = false): (GunId, int, float, PowerCap) =
|
||||
## `selectShot` plus the energy-aware power-policy decision, so a caller can
|
||||
## log the cap and its reason (see `applyPowerPolicy` in virtual_bullets).
|
||||
##
|
||||
## `dist` is the current distance (px) to the target and `selfEnergy` our own
|
||||
## energy; `ramming` exempts the caps (the movement code's `shouldRam` is the
|
||||
## single source of truth). The policy is applied identically wherever this is
|
||||
## called, so live and any offline caller cannot diverge.
|
||||
let gunId = t.selectGun(targetId, tick)
|
||||
let (prefBin, preferred) = t.bestPower(gunId, targetId)
|
||||
# pEst / pRef mirror `bestPower`'s own fitness source (per-target when data
|
||||
# exists, else the deterministic aggregate). An empty bin carries no rate of
|
||||
# its own, so it borrows the gun's aggregate — the same "no data" case the
|
||||
# policy documents.
|
||||
let fit = t.fitnessFor(targetId)
|
||||
let pRef = if PowerRefFixed > 0.0: PowerRefFixed
|
||||
else: gunRate(fit[gunId], pooled = true)
|
||||
let pEst =
|
||||
if fit[gunId].bins[prefBin].count == 0: pRef
|
||||
else: fit[gunId].bins[prefBin].hitRate()
|
||||
let dec = applyPowerPolicy(preferred, dist, selfEnergy, pEst, pRef, ramming)
|
||||
result = (gunId, binIndexForPower(dec.power), dec.power, dec)
|
||||
|
||||
proc selectShot*(t: var VirtualTracker, targetId = -1, tick = 0,
|
||||
dist = 0.0, selfEnergy = 100.0,
|
||||
ramming = false): (GunId, int, float) =
|
||||
## Returns (gunId, powerBinIdx, power) — the shot to take this tick.
|
||||
## Pass targetId to pick the best gun for that specific enemy. `tick` drives
|
||||
## the minimum-dwell hysteresis (see `selectGun`).
|
||||
let gunId = t.selectGun(targetId, tick)
|
||||
let (binIdx, power) = t.bestPower(gunId, targetId)
|
||||
## the minimum-dwell hysteresis (see `selectGun`). `dist`/`selfEnergy`/`ramming`
|
||||
## feed the energy-aware power cap (`TR_POWER_POLICY`); defaults keep every
|
||||
## existing caller compiling, and `TR_POWER_POLICY=0` reproduces the uncapped
|
||||
## `bestPower` preference. Use `selectShotPolicy` when the cap/reason is needed.
|
||||
let (gunId, binIdx, power, _) =
|
||||
t.selectShotPolicy(targetId, tick, dist, selfEnergy, ramming)
|
||||
result = (gunId, binIdx, power)
|
||||
|
||||
@@ -599,6 +599,103 @@ proc bestPower*(t: VirtualTracker, gunId: GunId, targetId: int = -1): (int, floa
|
||||
if fw.count == 0 or fw.hitRate() >= bar:
|
||||
return (binIdx, PowerBins[binIdx])
|
||||
|
||||
# ── energy-aware power policy (TR_POWER_*) ───────────────────────────────────
|
||||
#
|
||||
# `bestPower` answers "which power bin does this gun's own virtual data prefer?"
|
||||
# and is deliberately left untouched. The policy below CAPS that preference using
|
||||
# only cheap, always-available state — range, our own energy, and the gun's own
|
||||
# rate — so a long-range or low-energy shot trades single-hit damage for a
|
||||
# faster bullet (speed = 20-3p, so LOW power is FASTER and needs less lead) and a
|
||||
# shorter fire interval (10+2p, so LOW power = MORE shots). It never RAISES
|
||||
# power, so the shipped behaviour is exactly the `cap = 3.0` case, which is also
|
||||
# the control arm (`TR_POWER_POLICY=0`).
|
||||
#
|
||||
# Measured basis (real shots vs DrussGT, 8-16 runs): hit rate 21.6% at 0-100px,
|
||||
# 27.1% at 100-200, then 19.3% at 200-300, 10.9% at 300-400, 6.8% at 400-600 and
|
||||
# 5.4% at 600-800. Energy math: E[dE] = p(3P-1), so the break-even hit
|
||||
# probability is 1/3 INDEPENDENT of power; at range/low energy the extra speed
|
||||
# and shots of p=1.0 dominate. Damage is 4p (p<=1) / 6p-2 (p>1).
|
||||
const
|
||||
PowerPolicyEnvVar* = "TR_POWER_POLICY" ## 0 = control arm (uncapped)
|
||||
PowerFarDistEnvVar* = "TR_POWER_FAR_DIST" ## px; beyond this = bad-chances zone
|
||||
PowerLowEnergyEnvVar* = "TR_POWER_LOW_ENERGY" ## self energy below this = conserve
|
||||
PowerFarCapEnvVar* = "TR_POWER_FAR_CAP" ## cap for far / low-energy
|
||||
PowerMidCapEnvVar* = "TR_POWER_MID_CAP" ## cap when close+healthy but not above avg
|
||||
PowerRefEnvVar* = "TR_POWER_REF" ## 0 = gun's own mean; >0 = fixed P_ref
|
||||
|
||||
let PowerPolicyEnabled* = envBool(PowerPolicyEnvVar, true)
|
||||
let PowerFarDist* = envFloat(PowerFarDistEnvVar, 200.0)
|
||||
let PowerLowEnergy* = envFloat(PowerLowEnergyEnvVar, 50.0)
|
||||
let PowerFarCap* = envFloat(PowerFarCapEnvVar, 1.0)
|
||||
let PowerMidCap* = envFloat(PowerMidCapEnvVar, 2.0)
|
||||
let PowerRefFixed* = envFloat(PowerRefEnvVar, 0.0)
|
||||
|
||||
type
|
||||
PowerReason* = enum
|
||||
prFull ## above-average chances, close, healthy -> full power
|
||||
prFar ## beyond TR_POWER_FAR_DIST -> bad-chances zone
|
||||
prLowEnergy ## self energy below TR_POWER_LOW_ENERGY -> conserve
|
||||
prBelowAvg ## chances not above the gun's own average -> no power 3.0
|
||||
prRam ## ramming: exempt (at contact P->1, so 3.0 is correct)
|
||||
|
||||
PowerCap* = object
|
||||
power*: float ## the power to fire (<= the gun's preference)
|
||||
cap*: float ## the cap applied (3.0 = uncapped)
|
||||
reason*: PowerReason ## why
|
||||
|
||||
proc powerReasonName*(r: PowerReason): string =
|
||||
case r
|
||||
of prFull: "full"
|
||||
of prFar: "far"
|
||||
of prLowEnergy: "lowEnergy"
|
||||
of prBelowAvg: "belowAvg"
|
||||
of prRam: "ram"
|
||||
|
||||
proc binIndexForPower*(power: float): int =
|
||||
## Index of `power` in `PowerBins`; if it is not an exact bin value, the
|
||||
## highest bin whose power does not exceed it (0 if none). Keeps the returned
|
||||
## bin index consistent with a capped power.
|
||||
result = 0
|
||||
for i in 0..<len(PowerBins):
|
||||
if abs(PowerBins[i] - power) < 1e-9: return i
|
||||
if PowerBins[i] <= power: result = i
|
||||
|
||||
proc applyPowerPolicy*(preferredPower, dist, selfEnergy, pEst, pRef: float,
|
||||
ramming: bool,
|
||||
enabled = PowerPolicyEnabled,
|
||||
farDist = PowerFarDist,
|
||||
lowEnergy = PowerLowEnergy,
|
||||
farCap = PowerFarCap,
|
||||
midCap = PowerMidCap): PowerCap =
|
||||
## PURE cap core — no tracker, no battle. `power = min(preferredPower, cap)`,
|
||||
## so the result can only ever LOWER the gun's own preference. Order of
|
||||
## precedence: ram (exempt) > far > low energy > below average > full.
|
||||
##
|
||||
## `pEst` is the gun's rate for the bin it chose (or its aggregate when that
|
||||
## bin is empty); `pRef` is the gun's aggregate mean (or the fixed
|
||||
## `TR_POWER_REF`). A cold gun has no data, so `pEst <= pRef` is vacuously
|
||||
## true and it gets the mid cap — deliberately conservative until it has
|
||||
## evidence its chances are above average.
|
||||
if ramming:
|
||||
return PowerCap(power: preferredPower, cap: 3.0, reason: prRam)
|
||||
if not enabled:
|
||||
return PowerCap(power: preferredPower, cap: 3.0, reason: prFull)
|
||||
var cap: float
|
||||
var reason: PowerReason
|
||||
if dist > farDist:
|
||||
cap = farCap
|
||||
reason = prFar
|
||||
elif selfEnergy < lowEnergy:
|
||||
cap = farCap
|
||||
reason = prLowEnergy
|
||||
elif pEst <= pRef:
|
||||
cap = midCap
|
||||
reason = prBelowAvg
|
||||
else:
|
||||
cap = 3.0
|
||||
reason = prFull
|
||||
PowerCap(power: min(preferredPower, cap), cap: cap, reason: reason)
|
||||
|
||||
proc chooseFromFit*(fit: seq[GunFitness], diag: ptr SelectorDiag = nil,
|
||||
mode: SelectorMode = smAbsolute,
|
||||
referenceRate = -1.0,
|
||||
|
||||
Reference in New Issue
Block a user