feat(BNNBot): shaped reward based on miss distance — every shot teaches something
Replace binary hit/miss reward with exponential decay based on miss distance: - reward = 2.0 * exp(-missDistance / 36.0) - 1.0 - At 0px: +1.0 (perfect hit) - At 36px: -0.26 (near miss, small penalty) - At 100px: -0.87 (big miss, large penalty) Maintains virtual hit/miss counters for display (threshold: 36px). Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -96,12 +96,17 @@ method onScannedBot*(bot: BNNBot, e: ScannedBotEvent) =
|
||||
if bulletDist >= b.fireDist or b.trace.age >= TRACE_MAX_AGE:
|
||||
let bulletX = b.fireX + cos(degToRad(b.aimAngleDeg)) * bulletDist
|
||||
let bulletY = b.fireY + sin(degToRad(b.aimAngleDeg)) * bulletDist
|
||||
let enemyDist = hypot(bulletX - bot.lastEnemyX, bulletY - bot.lastEnemyY)
|
||||
if enemyDist < 36.0:
|
||||
bot.net.learn(b.trace, HIT_REWARD)
|
||||
let missDistance = hypot(bulletX - bot.lastEnemyX, bulletY - bot.lastEnemyY)
|
||||
|
||||
# Shaped reward: +1.0 for perfect hit, decays toward -1.0 as miss distance grows
|
||||
# Using exponential decay: reward = 2.0 * exp(-missDistance / 36.0) - 1.0
|
||||
let reward = 2.0 * exp(-missDistance / 36.0) - 1.0
|
||||
bot.net.learn(b.trace, reward)
|
||||
|
||||
# Track hit/miss for display: hit if within 36px, miss otherwise
|
||||
if missDistance < 36.0:
|
||||
inc bot.virtualHits
|
||||
else:
|
||||
bot.net.learn(b.trace, MISS_PENALTY)
|
||||
inc bot.virtualMiss
|
||||
b.active = false
|
||||
|
||||
|
||||
Reference in New Issue
Block a user