power policy: cap power by range and energy, gate 3.0 on above-average chances

Implements the user's energy management request: "firing from more than 200px
should be a 'not good chances zone' so faster bullets and more chances to hit
matters more than single hit damage with low chances. When we are lower than 50
health, same thing. I would like to use 3.0 power only when the chances of
hitting are higher than average."

Design: a CAP on top of the existing `bestPower`, not a rewrite. `bestPower`
still answers "which bin does this gun's own data prefer"; the policy caps it:

  ramming                                       -> 3.0  (reason ram, exempt)
  dist > TR_POWER_FAR_DIST (200)                -> 1.0  (far)
  elif selfEnergy < TR_POWER_LOW_ENERGY (50)    -> 1.0  (lowEnergy)
  elif pEst <= pRef                             -> 2.0  (belowAvg)
  else                                          -> 3.0  (full)
  power = min(gunPreferredBinPower, cap)   # can only LOWER power

p=1.0 is the right "low" tier on the measured mechanics: bullet speed 20-3p so
p=1.0 gives speed 17 vs 11 at p=3.0 (55% faster = less lead error), fire
interval 10+2p so 12 ticks vs 16 (33% more shots), and drain 0.083/turn vs
0.1875 (2.25x slower). All three things the user asked for at long range.
pEst = the chosen bin's virtual rate (gun aggregate when the bin is empty);
pRef = the gun's aggregate mean unless TR_POWER_REF > 0. No-data guns are
vacuously below-average -> cap 2.0 (conservative, documented).

Control arm: TR_POWER_POLICY=0 = uncapped = today's behaviour exactly.
Knobs: TR_POWER_POLICY, TR_POWER_FAR_DIST, TR_POWER_LOW_ENERGY,
TR_POWER_FAR_CAP, TR_POWER_MID_CAP, TR_POWER_REF, TR_POWER_LOG.
TR_POWER_MID_CAP exists because the user did not specify the middle case
(close + healthy + not-above-average); 2.0 is the default, flippable to 1.0.

Seam: the cap lives in a pure `applyPowerPolicy` and is applied only in
`selectShot` (the single place real shots are chosen), so the logic is testable
without a battle. Ram is wired from `shouldRam` - the same value the movement
dispatch uses for the (0,50) band.

CORRECTION TO AN ASSUMPTION IN THE TASK: `offline_range.nim` does NOT call
`bestPower`/`selectShot` - it only replays virtual-bullet spawn/resolve across
all power bins, independent of the real shot's power. So there is no offline
power-selection path that could diverge from the live one, and the acceptance
test guards the metric, not the policy. Policy coverage therefore comes from the
new unit test.

Verification: test_power_policy 26/26 in BOTH modes (default and TR_POWER_POLICY=0
control arm); test_gun_harness 39, test_vbullet_metric 11, test_power_selection 3,
test_adaptive_radar 41, test_tfil_ring_weights 24; acceptance_offline_vs_online
12/12 VERDICT PASS (live battle). ModularBot compiles.

UNVERIFIED: the live effect on damage/survival/score. No A/B has run.
This commit is contained in:
2026-09-21 23:59:56 +02:00
parent 9caf1d3728
commit c9825dfb0b
4 changed files with 356 additions and 5 deletions
+18 -1
View File
@@ -100,6 +100,11 @@ let MovementName* = getEnv("TR_MOVEMENT", "tfil").strip().toLowerAscii()
## Per-process output paths so concurrent A/B runs do not clobber each other.
let GunStatsPath = getEnv("GUN_STATS_PATH", "/tmp/gun_stats.jsonl")
let ShotLogPath = getEnv("GUN_SHOTLOG_PATH", "/tmp/shot_log.jsonl")
## Energy-aware power policy observability (TR_POWER_LOG=1): emit ONE line per
## CHANGE of the (power, cap, reason) decision — not per tick — so the GUI log
## shows why the cap moved. The policy itself lives in the shared gun harness
## (`applyPowerPolicy`), so both live and offline paths see the same rule.
let PowerLog = existsEnv("TR_POWER_LOG")
const GunNames = ["HeadOn", "Linear", "Tsetlin", "Circular", "GuessFactor", "Pattern", "WallBounce", "Accel", "StopShot", "Displace", "AvgLead", "DecayGF", "KNN", "TMSelect"]
const
@@ -175,6 +180,7 @@ type
bulletShot: Table[int, PendingShot] ## bulletId -> shot metadata (Task A shot log)
pendingHitBullets: HashSet[int] ## hit bulletIds seen before their onBulletFired stamp
gunSelectionCount: array[14, int]
lastPowerLogKey: string ## change detector for the TR_POWER_LOG line
lastKnownTargetId: int ## persists through death, used for round-end stats
# Radar measurement instrumentation (only touched when RadarScanLog is set).
radarScanCounts: Table[int, int] ## onScannedBot calls per enemy id
@@ -920,7 +926,18 @@ method run*(bot: ModularBot) =
)
# Gun selection + fire. `bot.tick` drives the selector's dwell window.
let (selectedGun, _, power) = selectShot(bot.tracker, tid, bot.tick)
# The energy-aware power cap uses the current target distance and our own
# energy; `shouldRam` (the movement code's decision) exempts it.
let (selectedGun, _, power, pdec) = selectShotPolicy(
bot.tracker, tid, bot.tick,
dist = ramDist, selfEnergy = ws.selfEnergy, ramming = shouldRam)
if PowerLog:
let pkey = fmt"{power:.1f}|{pdec.cap:.1f}|{pdec.reason}"
if pkey != bot.lastPowerLogKey:
bot.lastPowerLogKey = pkey
echo fmt"[power] p={power:.1f} cap={pdec.cap:.1f} " &
fmt"reason={powerReasonName(pdec.reason)} " &
fmt"dist={ramDist:.0f} selfE={ws.selfEnergy:.0f} gun={GunNames[selectedGun]}"
bot.gunSelectionCount[selectedGun] += 1
if selectedGun != bot.currentGun:
bot.currentGun = selectedGun