j131 learned movement: real bullet-endpoint resolution (TR_LEARNED_REAL_EVENTS, default off) + exact-geometry Gate A/B (inversion NOT fixed; state still the constraint)
This commit is contained in:
@@ -2243,3 +2243,134 @@ pre-registered rule records as WORSE.
|
||||
**Status: the default is UNCHANGED (`TR_MOVEMENT=strafe`); the outcome mode is
|
||||
default-off behind `TR_MOVEMENT=learned TR_LEARNED_LABEL=outcome`.** Revert = do
|
||||
not set the env vars.
|
||||
|
||||
---
|
||||
|
||||
## Learned movement — real bullet endpoints (exact geometry)
|
||||
|
||||
**Job j131. The owner's request:** *"use real bullets: bullets that really hit
|
||||
me, bullets that hit the wall, both detectable. We ignore bullets that hit
|
||||
other bots, this movement is only for 1v1."* The task's premise was that
|
||||
`ModularBot.nim` already handles `onBulletHit`/`onBulletHitWall`, so the exact
|
||||
bullet line was available live and j130's rejection of the exact label ("needs
|
||||
bullet bodies the bot lacks") was wrong.
|
||||
|
||||
### THE PREMISE IS HALF WRONG — VERIFIED (MEASURED, not inferred)
|
||||
|
||||
The **fields** exist: `BulletState` has `x, y, direction, power, ownerId,
|
||||
bulletId`, and `BulletHitWallEvent`/`HitByBulletEvent` both expose
|
||||
`bullet: BulletState`. But **the events are not routed to the dodger**:
|
||||
|
||||
* `BulletHitWallEvent` is delivered **only to the bullet's owner**
|
||||
(`addPrivateBotEvent(outcome.bullet.botId, …)` — verified by decompiling the
|
||||
running server jar `robocode-tankroyale-server-0.35.5-all.jar`, and identical
|
||||
in the 1.1.0 source `CollisionDetector.applyBulletWallCollisions`). So an
|
||||
**enemy** bullet hitting a wall is **not observable** by us.
|
||||
* `TurnToTickEventForBotMapper` builds `bulletStates = turn.bullets.filter
|
||||
{ it.botId == bot.id }`, so `getBulletStates()` returns **only our own**
|
||||
bullets too.
|
||||
* The events the dodger **does** receive with a real enemy-bullet endpoint are:
|
||||
`onHitByBullet` (the bullet hit US — endpoint = our impact point) and a
|
||||
bullet-vs-bullet event where **our** bullet intercepted an enemy bullet
|
||||
(`e.hitBullet` is the enemy bullet, with its endpoint + heading).
|
||||
|
||||
**So the "exact straight line from a wall hit" cannot be built live.** In 1v1 a
|
||||
missed bullet does end on a wall, but the server keeps that observation private
|
||||
to the shooter. This is the second time the availability premise is the binding
|
||||
constraint, now for the exact label rather than the proxy.
|
||||
|
||||
### WHAT CHANGED (code)
|
||||
|
||||
* `common_libs/movements/learned_surfer.nim` — **default-off**
|
||||
`TR_LEARNED_REAL_EVENTS=1` (registered in `env_report.knownEnvNames()`). When
|
||||
on, a wave is resolved by the REAL event instead of the arrival deadline:
|
||||
the exact `origin → endpoint` straight line sets the label's GF bin, the real
|
||||
flight time `currentTick − fireTick` is recorded (`resolvedReal`, `lastFlightErr`
|
||||
— a cross-check on the energy-drop speed inference), and the wave is **dropped
|
||||
at once** (`resolveEnemyBullet`), so no ghost accumulates. A wave no event
|
||||
claims resolves `RealEventsGrace` ticks past nominal as a **wall MISS**. With
|
||||
the knob off the byte-for-byte j130 behaviour is preserved (tests pin it).
|
||||
* `ModularBot_garage/src/ModularBot.nim` — forwards `onHitByBullet` (hit on us),
|
||||
a bullet-vs-bullet intercept of an enemy bullet (`e.hitBullet`), and (guarded,
|
||||
dead on 0.35.5) an enemy `onBulletHitWall` to `learnedMover.resolveEnemyBullet`.
|
||||
* `ModularBot_garage/tests/test_learned_surfer.nim` — real-event unit checks
|
||||
(default-off parity, exact centre-bin resolution, ghost drop, wall-miss
|
||||
deadline). `common_libs/tests/exact_geometry_gate.py` — Gate A/B below.
|
||||
|
||||
### GATE A — danger-map alignment, ONE consistent computation (MEASURED)
|
||||
|
||||
`python3 common_libs/tests/exact_geometry_gate.py --corpus /tmp/tfil_ab2/out`
|
||||
(70 battles, 54 923 shots, the same extraction and the same
|
||||
`corr(danger(g), P(hit | b_our=g))` metric j128/j130 used):
|
||||
|
||||
| danger map | corr vs `P(hit\|b_our=g)` | corr vs `P(hit\|b_bullet=g)` |
|
||||
|---|---:|---:|
|
||||
| histogram P(arrival = g) (j128) | **−0.341** | −0.206 |
|
||||
| outcome proxy `P(hit & \|g−b_our\|≤w)` (j130 live) | **+0.566** | +0.604 |
|
||||
| **EXACT bullet line `P(\|g−b_bullet\|≤w)`** | **−0.230** | **+0.120** |
|
||||
| exact bullet line & hit | +0.465 | +0.684 |
|
||||
|
||||
**The exact-geometry label does NOT fix the inversion on the j128 metric** —
|
||||
−0.230 is still negative (minimising it still steers into where the observed
|
||||
hits happen). It is *less* negative than the histogram (−0.341) and turns
|
||||
weakly positive (+0.120) only when the target is conditioned on the bullet's
|
||||
own line `b_bullet`, while the +0.566 proxy is inflated by being conditioned on
|
||||
`b_our` (the realised arrival, i.e. where the recorded wave already was). Under
|
||||
the task's own gate, **the veto fires and the live batch is not run.**
|
||||
|
||||
### GATE B — state information under the EXACT label (MEASURED)
|
||||
|
||||
Held-out per-candidate log-loss of the exact label, split BY BATTLE, 3 seeds:
|
||||
|
||||
| model | log-loss (bits) |
|
||||
|---|---:|
|
||||
| state-free `P(label \| g)` | **0.1879** |
|
||||
| state-conditional `P(label \| state, g)` | **0.3747** |
|
||||
| Δ (state − state-free) | **+0.1868** |
|
||||
|
||||
state conditioning is better in **0/3** splits. This **replicates j130 almost
|
||||
exactly** (proxy: 0.3906 vs 0.1873, Δ +0.203, 0/3). Under the exact label the
|
||||
coarse four-field state is still *worse* than the state-free model: the state
|
||||
buys no held-out information, so it cannot be the thing the learned mover is
|
||||
missing — **the observable state is still the binding constraint.**
|
||||
|
||||
### GATE C — live panel (NOT RUN, by the pre-registered rule)
|
||||
|
||||
Gate A's veto fired (exact correlation negative), so no live battles were
|
||||
fought. Independently, the live batch would have been testing a label the module
|
||||
**cannot construct** in the miss case (enemy wall endpoints are owner-private),
|
||||
so a live "exact" arm would in practice be j130's proxy for ~90% of waves.
|
||||
|
||||
### Direct answer
|
||||
|
||||
**Does exact bullet geometry fix the label? NO — not on the measured metric and
|
||||
not live.** The physically-exact map reads −0.230 against the j128 target
|
||||
(still inverted; the proxy's +0.566 is the one that is inflated). And the
|
||||
geometric endpoint **is not observable** by the dodger on this server for the
|
||||
miss case: `BulletHitWallEvent` and `bulletStates` are owner-private, so the
|
||||
only real enemy-bullet endpoints we get are the ~13% that hit us (and the rare
|
||||
intercepts). The exact line therefore cannot be built live for the waves that
|
||||
matter.
|
||||
|
||||
**Is the binding constraint the STATE rather than the label or the learner?
|
||||
YES — the same answer as j130, now measured for the third label.** Under the
|
||||
exact label the state still loses to state-free on held-out log-loss (0.3747 vs
|
||||
0.1879, 0/3 splits). j128 (histogram), j130 (outcome proxy) and j131 (exact
|
||||
line) each change the label; none moves the live result and none makes the
|
||||
state informative. The wave-crossing signal a 1v1 dodger needs is simply not in
|
||||
the four-field observable state, and hand-tuned `strafe` remains hard to beat.
|
||||
|
||||
### MEASURED vs INFERRED
|
||||
|
||||
**MEASURED:** the event routing (decompiled the running 0.35.5 jar +
|
||||
`TurnToTickEventForBotMapper`); the three-way Gate A correlation and the
|
||||
exact-label Gate B log-loss on the recorded corpus; the module unit tests
|
||||
(24/24, including the real-event and default-off parity checks); the env-report
|
||||
guard (25/25); the clean-archive compile. **INFERRED:** that the offline
|
||||
alignment transfers live — it cannot (open-loop corpus, see
|
||||
`docs/offline_harness_trust.md`).
|
||||
|
||||
**Status: the default is UNCHANGED (`TR_MOVEMENT=strafe`).** The real-event
|
||||
resolution is default-off behind `TR_MOVEMENT=learned TR_LEARNED_REAL_EVENTS=1`
|
||||
(combined with `TR_LEARNED_LABEL=outcome` for the dense readout). Revert = do
|
||||
not set the env vars.
|
||||
|
||||
Reference in New Issue
Block a user