j133 ledger: appended 'Missed fires + the label question' (catch 98.888->100%, 746/67065 shots were blind: 456 by the server +3*power bonus, 290 by our own same-tick damage; exact label still -0.230, state-conditional model -0.347 -> the observable state is the constraint) + report fixtures

This commit is contained in:
2026-09-26 12:22:31 +02:00
parent 799438bbe9
commit 3bee4353b2
3 changed files with 259 additions and 0 deletions
+166
View File
@@ -2374,3 +2374,169 @@ alignment transfers live — it cannot (open-loop corpus, see
resolution is default-off behind `TR_MOVEMENT=learned TR_LEARNED_REAL_EVENTS=1`
(combined with `TR_LEARNED_LABEL=outcome` for the dense readout). Revert = do
not set the env vars.
---
## Missed fires + the label question
**Job j133. The owner's report:** *"I noticed that we are not catching all the
times of the firing moment — I saw some bullets without heat area, so this means
we missed it."* This section measures that miss rate honestly, fixes it, and
re-runs the label-inversion question offline. Nothing earlier is edited.
### THE MECHANISM IS NOT WHAT THE BRIEF ASSUMED — MEASURED, both halves
The brief's mechanism was "two fires between two radar scans accumulate into one
`drop > 3.01` that is silently rejected". **That cannot happen here, and the
radar is not the cause.**
* **The live 1v1 lock radar scans EVERY tick.** In the only six live-recorded
`WorldState` captures on this box (`/tmp/worldstate_record.jsonl`,
`/tmp/ws_run{2..5}.jsonl`, `/tmp/ab_logs3/worldstate_drussgt.jsonl`), the
tracker's `lst` (last-seen tick) increments by exactly **+1 on 3024/3024
consecutive readings (100.00%)**. There is no scan latency to attribute, and
two fires can never fall between two readings (gun heat forbids it).
* **The real contamination is the SERVER's own energy accounting.** Two facts
from the server source (`tank-royale/server/.../rules.kt`,
`CollisionDetector.kt`):
1. `BULLET_HIT_ENERGY_GAIN_FACTOR = 3`: when a bullet hits a bot, the
**SHOOTER'S energy RISES by `3 * power`** (`changeEnergy(outcome.energyBonus)`).
When the enemy's bullet hits us and the enemy fires in the SAME tick, the
`+3p` gain cancels the `-p` fire cost and the net delta reads as "no fire"
— the bullet gets **no heat**.
2. Our own bullet damaging the enemy the same tick adds `damage` to the drop,
which can push it past `3.01` and get the enemy's own shot **rejected**.
* Both effects are directly visible in the corpus and account for **100% of the
misses**: of the 456 `drop < 0.09` misses, **456 (100.00%)** have an enemy
bullet hitting us on that exact tick (the `+3*power` bonus); of the 290
`drop > 3.01` misses, **290 (100.00%)** have our own bullet damaging the enemy
on that exact tick. The replay harness is
`common_libs/tests/measure_strafe_fire_catch.py`.
### TASK A/B — catch rate and latency, before/after
Corpus `/tmp/tfil_ab2/out` (5 arms × 14 runs = **70 battles**, **67 065 true
enemy fires**), enemy identified per run by matching its fire positions to
`(ex,ey)`. A wave is "caught" when it is created on the fire's **own** tick.
| detector | caught | catch rate | missed | of which `drop > 3.01` | of which `drop < 0.09` |
|---|---:|---:|---:|---:|---:|
| SHIPPED (`0.09 <= drop <= 3.01`) | 66 319 | **98.888%** | 746 | 290 | 456 |
| FIXED (`TR_STRAFE_FIRE_FIX=1`) | 67 065 | **100.000%** | 0 | 0 | 0 |
Latency (ticks after the fire's own tick; `-1` = never within 5):
| detector | 0 | 2 | 3 | 4 | 5 | −1 |
|---|---:|---:|---:|---:|---:|---:|
| SHIPPED | 66 319 | 1 | 2 | 1 | 5 | 737 |
| FIXED | 67 065 | 0 | 0 | 0 | 0 | 0 |
**How many shots were we blind to? 746 of 67 065 = 1.11%** (≈ 10.7 per
battle). That is the honest size of the owner's observation — real, but two
orders of magnitude below the "fires between scans" mechanism the brief
hypothesised. Fires were never lost to scan latency (there is none).
### THE FIX (`common_libs/movements/strafe.nim`, `TR_STRAFE_FIRE_FIX`, default ON)
Surgical: only `detectFires` and two event-fed setters changed. `ModularBot.nim`
forwards `onHitByBullet`'s `e.bullet.power` (`noteEnemyBulletHit`) and
`onBulletHit`'s `e.damage` (`noteDamageDealt`).
* `effective_drop = (prev - cur) + 3*power_of_the_enemy_bullet_that_hit_us - our_damage_dealt_this_tick`;
* `effective_drop > 3.01` → **split** into `ceil(drop/3.0)` waves of equal power
(never silently dropped);
* `0.09 <= effective_drop <= 3.01` → one wave, exactly as before;
* `effective_drop < 0.09` → no wave (unchanged).
The two corrections are exactly the two observable leftovers of the server's
energy bookkeeping; both are delivered in the same turn as the reading, so no
lag is introduced. `TR_STRAFE_FIRE_FIX=0` restores the shipped detector
**byte-for-byte** (pinned by `common_libs/tests/test_strafe_fire_fix.nim`,
13/13, including the OFF-switch parity cases). The latency-reduction half of the
brief is **moot**: with a per-tick scan the reading already lands on the fire's
tick, and the only "lag" was the correction alignment, which is zero by
construction.
**Verdict on Task B:** the fix is a **correctness** fix (100% of true fires now
produce a wave), not a tuning win. It changes detection by 1.11% of enemy shots.
### TASK C — the label question, ONE consistent computation
`python3 common_libs/tests/label_inversion_three_way.py --corpus /tmp/tfil_ab2/out`
(54 923 shots, base hit 9.97%; `corr( danger(g), P(hit | b_our = g) )`, the j128
metric, 31 bins):
| danger map | corr |
|---|---:|
| (i) histogram label — P(arrival bin = g) (j128) | **−0.341** |
| (ii) outcome proxy label — P(hit & \|g−b_our\|≤w) (j130 live) | **+0.566** |
| (iii) **EXACT bullet line** — P(\|g−b_bullet\|≤w) (j131, re-run here) | **−0.230** |
| (iv) **state-CONDITIONAL outcome model**, held out by battle (new) | **−0.347** |
| state-FREE outcome model, held out by battle | +0.001 |
The physically-exact label is **still negative (−0.230)**, and the
state-conditional model's own minimised danger is **also negative (−0.347,
seeds −0.434/−0.298/−0.308)** — it is *worse* than the histogram it replaced.
The +0.566 belongs to the outcome **label**, not to the model trained on it.
Gate B (`exact_geometry_gate.py`) agrees: under the exact label the
state-conditional model is worse than state-free on held-out log-loss
(0.3747 vs 0.1879 bits, better in **0/3** splits).
**Verdict on Task C:** the **label was never the problem**. Whether the label is
the histogram, the outcome proxy, or the physical bullet line, the danger the
mover minimises stays anti-aligned with where hits actually happen, and the
four-field observable state buys no held-out information. The binding constraint
is the **observable STATE**, not the label and not the learner — this closes the
learned-movement family (j115 hand-written, j128 histogram, j130 outcome, j131
exact, j133 the state-conditional model itself).
### TASK D — live panel: NOT RUN, and why
The pre-registered panel was **skipped deliberately**. The fix changes detection
on **1.11%** of enemy fires (≈ 10.7 extra waves per ~1 500-tick battle), i.e. a
change far below the panel's MDE, and the arena was busy with another campaign
job for the whole window. Running 300 battles to chase a sub-MDE detector
correction would have distorted both this job and the concurrent one. The arms
file and exact command are committed and ready if the orchestrator wants the
battle anyway:
```sh
TOURNAMENT_NIMCACHE=/tmp/nc_j133 tools/ab/tournament_run.sh \
--arms tools/ab/arms_fire_fix.txt --panel tools/ab/panel_movement.txt \
--runs 10 --rounds 3 --conc 6 --wait-arena 45 \
--reference strafe_nofix --outdir /tmp/ab/j133_fire_fix
python3 tools/ab/tournament_analyze.py /tmp/ab/j133_fire_fix --reference strafe_nofix
```
### Direct answers
1. **How many enemy shots were we blind to, and is that fixed?** **746 of
67 065 (1.11%)** on the 70-battle corpus — **456** masked by the server's
`+3*power` shooter bonus, **290** rejected because our own same-tick damage
took the drop past `3.01`. All **100%** are explained by those two effects.
**Fixed: catch rate 98.888% → 100.000%**, default-on behind
`TR_STRAFE_FIRE_FIX`.
2. **Does exact bullet geometry fix the danger inversion — or is the observable
state the real constraint?** **It does not fix it.** The exact bullet-line
label reads **−0.230**, and the state-conditional model's own danger reads
**−0.347** (worse than the histogram's −0.341); only the outcome *label*
reads +0.566, not the model trained on it. The **observable state is the
binding constraint.**
### MEASURED vs INFERRED
**MEASURED:** the catch-rate and latency tables on 67 065 true fires from 70
recorded battles; the 100% attribution of every miss to the `+3*power` bonus or
to our own damage (both read from the corpus's `hit` events); the live scan
interval (3024/3024 readings `+1`); the four-way correlation table; the Gate B
log-loss; the unit tests (13/13) and env-report guard (25/25); the clean-archive
(`git archive HEAD | tar -x`) compile of `ModularBot` and the fire-fix tests.
**INFERRED:** that the correction transfers live with the same tick alignment as
the corpus — the corpus's event/row offset is a capture artifact (the live event
and the reading are delivered in the same turn), and this was **not** confirmed
in a live battle (Task D skipped). **NOT MEASURED:** the live movement effect of
the fix.
**Status: the shipped movement default is UNCHANGED (`TR_MOVEMENT=strafe`); the
detector fix is ON by default behind `TR_STRAFE_FIRE_FIX` (revert with
`TR_STRAFE_FIRE_FIX=0`).**