Files
SirStone 2747ebd323 BitBrain: TR_BITBRAIN_GAINS env knob (candidate set + fixed-gain degenerate)
Task A of campaign phase 2: the lead-gain candidate set is now pure env, so the
live arms need no recompile.

- common_libs/guns/bitbrain_gun.nim: BB_GAINS_ENV (TR_BITBRAIN_GAINS); the
  candidate list is parsed once at gun construction into a dynamic seq, so the
  hit counts/hit rates are sized to it. Unset/unparsable -> the shipped
  BB_CAND set [0,0.25,0.5,0.75,1.0] (byte-identical behaviour). Exactly ONE
  candidate degenerates to a FIXED gain applied from the first shot (learning
  bypassed), still gated to the long bands. parseGains clamps to [0,8],
  de-dupes and sorts so the argmax tie rule is unchanged. The [bb] line now
  prints the APPLIED gain AND the resulting angular shift, so a run's
  correction is auditable from stdout.
- ModularBot_garage/src/env_report.nim: emit TR_BITBRAIN_GAINS (resolved
  candidate set) and add BB_GAINS_ENV to the known-name list.
- tools/ab/arms_leadgain.txt: the 6-arm phase-2 sweep definition.
2026-09-25 00:15:27 +02:00

23 lines
1.6 KiB
Plaintext

# Phase 2 live A/B: does a lead gain ABOVE 1.0 beat shipped Pattern?
#
# One frozen binary (git archive HEAD), real DrussGT, 6 arms x 7 runs x 7 rounds.
# Every BitBrain arm swaps the admitted rack gun (Pattern off, BitBrain on) so the
# ONLY thing that changes is the lead gain: BitBrain's base prediction IS Pattern
# (`tmh.pattern.predict`), scaled over the line of sight by `gain`. The correction
# is applied only in the long bands (range >= 300 px, BB_GAIN_BAND_MIN), i.e. the
# region never tested live. TR_BITBRAIN_LOG=1 puts the APPLIED gain on the [bb]
# line; liveness is read from the boot env report (raw `[env] VAR=VALUE`).
#
# control = shipped Pattern-only rack, no env (reference)
# g100 = single candidate 1.0 -> FIXED gain, identity: MUST match control
# glo = learner restricted to <= 1 (expected HARMFUL per the live HeadOn kill)
# ghi = learner allowed above 1 (THE HYPOTHESIS)
# gfix150 = fixed 1.5, no learning
# gfix125 = fixed 1.25, no learning
control |
g100 | TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0 TR_BITBRAIN_LOG=1 | fixed gain 1.0 (identity/validity)
glo | TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=0.25,0.5,0.75,1.0 TR_BITBRAIN_LOG=1 | learner restricted to <= 1
ghi | TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_LOG=1 | learner allowed above 1 (HYPOTHESIS)
gfix150 | TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.5 TR_BITBRAIN_LOG=1 | fixed gain 1.5, no learning
gfix125 | TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.25 TR_BITBRAIN_LOG=1 | fixed gain 1.25, no learning