39 lines
2.5 KiB
Plaintext
39 lines
2.5 KiB
Plaintext
# ─────────────────────────────────────────────────────────────────────────────
|
||
# arms_movement_outcome.txt — the OUTCOME-LABELLED learned-movement arms
|
||
# (job j130), on the FROZEN panel tools/ab/panel_movement.txt.
|
||
#
|
||
# Pre-registered in docs/movement_campaign.md, "Learned movement — outcome
|
||
# label (P(hit))", BEFORE any battle. REFERENCE is `strafe` — the SHIPPED
|
||
# champion. Every delta is (arm − strafe).
|
||
#
|
||
# j128 measured `corr( P(arrival bin), P(hit | arrival bin) ) = -0.342` over the
|
||
# 31 bins: minimising the resolved-position histogram steers INTO the bullets.
|
||
# j130 replaces the label with the dense outcome
|
||
# hit(state, g) = hit and |g - b_our| <= window(wave)
|
||
# and learns P(hit | state, candidate g) with a counted 2-class SBC.
|
||
#
|
||
# Gate A (common_libs/tests/outcome_label_gate.py, corpus /tmp/tfil_ab2/out):
|
||
# the alignment correlation flips to +0.566 (histogram -0.341), so the veto
|
||
# does NOT fire; the state-conditional outcome model however is NOT better than
|
||
# the state-free one on held-out log-loss, and the open-loop decision
|
||
# counterfactual barely moves (3.53% -> 3.33%). The live panel decides.
|
||
#
|
||
# Format: name | ENV=value ENV=value | label
|
||
# ─────────────────────────────────────────────────────────────────────────────
|
||
|
||
# 1. THE CHAMPION — the arm a challenger has to beat (round wins + hit rate).
|
||
strafe | TR_MOVEMENT=strafe | champion/reference — shipped strafe defaults
|
||
|
||
# 2. THE OLD LABEL — j128's state-conditional counted SBC (arrival-bin label),
|
||
# so the new label is isolated against the old one on the same binary.
|
||
learned | TR_MOVEMENT=learned | j128 histogram label (arrival bin)
|
||
|
||
# 3. THE NEW LABEL — the same mover, same geometry, same counted+decayed SBC
|
||
# and penalties; only the training label changes (dense hit outcome).
|
||
learned_outcome | TR_MOVEMENT=learned TR_LEARNED_LABEL=outcome | outcome label P(hit | state, g)
|
||
|
||
# 4. INFORMATION CONTROL — the outcome label with the state forced to one cell
|
||
# (state-free outcome model). Isolates whether the state carries anything
|
||
# under the new label (Gate A says it does not).
|
||
learned_outcome_global | TR_MOVEMENT=learned TR_LEARNED_LABEL=outcome TR_LEARNED_GLOBAL=1 | outcome label, state conditioning OFF
|