298ea6d586a5db53a2643896cbefdeafb4492091
94 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f58d65d2e8 |
TMComposites gate: per-gun confidence faithful for 3 guns; no pair composes
Adds a per-sample intrinsic-confidence field (GunPrediction.confidence, threaded through FeedbackEvent/VirtualBullet, populated by Pattern, DecayGF, KNN, GuessFactor, Tsetlin, TMHorizon) and an offline recorder + analyzer that reproduce the paper's Figure 2 per gun and its Eq-8 composite. Measured on 3 held-out tr-bridge DrussGT battles (33k ticks, ~133k samples/gun): - FAITHFUL: DecayGF (rho +0.133), KNN (+0.090), Pattern (+0.064, weak). - GuessFactor is ANTI-faithful (rho -0.067); Tsetlin c_max is useless (0.001). - No pair of guns specialises complementarily: the same gun dominates both high-confidence slices in every pair. - Eq-8 alpha-normalised confidence-weighted composite: 18.41% vs Pattern 20.45% (McNemar p=3.1e-126). Faithful-only variant 18.68%, still loses. Shuffle control passes weakly (composite > shuffle, p=4e-14) so ~0.7pp of competence is real but ~2pp short. Offline veto: design is dead. See docs/tmcomposites_gate.md. |
||
|
|
4270136948 | docs: correct counted-SBC decay amortised cost figure | ||
|
|
d85ff53d34 |
State-window gate: a temporal window of wave-relative states does NOT beat a single state
Re-runs the SBC coincidence premise as a cheap veto test on the 70-battle
live-vs-DrussGT corpus (/tmp/tfil_ab2). Defines the wave-relative state
(lat/vlat/toa/room/turn, 5/7.9/10 bits at Q=2/3/4), quantises the miss offset
at the bullet's arrival into 7 bins, and sweeps window length K in
{1,4,8,16,32,48} with an interpolated suffix-backoff model under a BY-BATTLE
70/30 split (3 seeds).
Result: NO. On the pre-fire frame the window is worse than the single
fire-tick state at every K/Q/A (e.g. K=8 costs +0.35..+0.46 bits). On the
during-flight frame the entire apparent gain is the later decision tick, not
the window; the single state alone drops 2.70 -> 1.31 bits as K goes 1 -> 32.
The shuffle-order control confirms recency matters but the windows do not:
by Q=4/K=8 they average ~1 observation and never recur. The single state
survives as a strong predictor (log-loss 2.346 vs 2.698 majority; bin
accuracy 0.409 vs 0.235).
Gate only: no gun, no live-win claim.
|
||
|
|
40ba96f649 |
BitBrain SBC: counted mode + global decay (forgetting, probabilities)
Adds an smCounted storage mode alongside the default smBitset. Each (i,j,class) cell becomes a saturating uint8 counter; learn increments it and a global fractional decay (c -= c shr decayShift every decayEvery learns) makes forgetting possible. infer sums raw counters; new inferProb sums the per-cell posterior P(class|cell) (scale-free, recommended readout). Bitset path is the default and byte-for-byte unchanged: test_bitbrain 56/56 (was 32), and test_bitbrain_mnist reproduces 97.210% corrected / 96.540% bug-compatible exactly. Counted mode configurable at runtime (TR_BITBRAIN_MODE / TR_BITBRAIN_DECAY_*) and compile time (-d:bitbrainDecay*). Measured: forgetting (86.2% vs 48.9% on a permuted-label stream), probabilities (rare-class balanced 0.998 vs 0.500), and the stationary cost (counted hurts MNIST; see docs/bitbrain_counted_sbc.md). Harness: common_libs/tests/measure_counted_sbc.nim |
||
|
|
39e06719fb |
Campaign phase 2 ledger: live lead-gain sweep (gains above 1.0 do NOT beat Pattern)
6 arms x 7 runs x 7 rounds vs real DrussGT on one frozen binary ( |
||
|
|
c305ef4212 |
BitBrain campaign phase 1: the missing gain sweep and BitBrain as a lead-gain corrector
Task A — the gain region Phase 0 never covered (gain < 1). Extend the prediction-quality ruler with gain 0.25/0.50/0.75 arms and a fixed causal per-band arm. Full 70-run result: the hitProxy-argmax curve is [1.00, 1.00, 1.00, 0.00, 0.00] — Pattern below 300 px, HeadOn above — worth +0.13 pp at 300-450 and +2.16 pp at 450+ (0.0767 -> 0.0984). Lead correlation is identical for every g>0 (Pearson is scale-invariant), so a shrinking gain adds no lead information; and the least-squares optimum [1,1,1,1,0.25,0.25] diverges from the hitProxy optimum because Pattern's lead errors are bimodal. Task B — rebuild guns/bitbrain_gun.nim as a lead-gain corrector: aim = LOS + gain*(patternAim - LOS), gain learned online per range band by ranking candidate gains on the hit-probability proxy (the observed lead label via tmhObservedAt), gated to range >= 300 px. ADE+SBC output removed. Offline (70 runs): Pattern below 300 px, +0.37 pp at 300-450, +1.90 pp at 450+ (hitProxy 0.0957 vs 0.0767), matching the fixed rule to within 0.27 pp at 450+ and exceeding it at 300-450. ~0.0007 ms/tick marginal (old gun ~0.114 ms/tick). Default off; rack membership, env report and guard tests (bitbrain 32, registration 13, rack 48, tm_pattern 20, env_report) unchanged and green. Ledger: docs/bitbrain_campaign.md Phase 1, including the three on-file negatives and the causal-shippability note. No live claim. |
||
|
|
140fe2519a |
HeadOn (no-lead) vs Pattern LIVE at long range: clean negative, offline ruler killed
2 arms x 15 runs x 7 rounds, one frozen binary from HEAD
|
||
|
|
a82c864c60 |
bitbrain campaign phase 0: offline prediction-quality ruler and the bar
New harness (common_libs/gun_harness/prediction_quality.nim + common_libs/tests/run_prediction_quality.nim): per-gun single-tick aim error in degrees against the true continuous interception point on the recorded live-vs-real-DrussGT corpus (/tmp/tfil_ab2/out, 70 runs, 899607 ticks), per range band, with the hit-probability proxy mean(|err|<=atan(18/range)). Validated: recorded hits separate from misses 13.34x px (reference 11.59x), perfect-oracle max |err| = 0, correct ordering on synthetic ground truth, two full runs byte-identical. Fixed a wrap180 bug (Nim float mod keeps the dividend sign) that inflated the negative error tail. Bar (mean|err| deg [hitProxy] at 450+): Pattern 16.19 [0.077], naive-linear 22.86 [0.054], TMHorizon 16.20 [0.076], BitBrain 16.20 [0.077], static HeadOn 12.33 [0.098], oracle 0 [1.0]. Lead-gain sweep on Pattern is a dead end (1.0 wins every band). Naive-linear applies ~1.8x Pattern's lead but carries no more lead information (corr 0.178 vs 0.165) and is strictly worse. Ledger: docs/bitbrain_campaign.md. All verdicts remain live-only. |
||
|
|
32a5e72fac |
Gun mixing (TMHorizon+BitBrain) vs DrussGT: clean negative, no dodge disruption
4 arms x 7 runs vs real DrussGT. mix alternates the two guns 476 times/7 runs (liveness OK) but our bullets are no more varied (power sd / aim-offset sd flat) and DrussGT's dodge quality is unchanged (miss/tick mix-pat +0.03, p=0.66; MDE 3.8%). mix wins 24/49 = the 49% baseline; the user's 6/10 has P=0.353 at 49%. New tools/ab/ab_dodge_analyze.py splits the validated per-shot dodge instrument by arm and adds gun-switch/power/bearing liveness; fixtures committed. |
||
|
|
d93ce444c0 |
BitBrain verdict: clean negative at 30 runs/arm; analyzer gets MC + Mann-Whitney + MDE
- docs/bitbrain_gun_verdict.md: control vs bb_decay (decay SBC memory) vs a provably-zero placebo, 30 runs/arm vs real DrussGT. Nothing separates (bb_decay +3.3 dmg/run, p=0.71; round wins 97/210 vs 97/210, p=1.00); the 7-run shape does not replicate. TR_BITBRAIN_RANGE=0 is clamped to 1.0 deg (bitbrain_gun.nim:207) so it is NOT a zero-shift placebo; TR_BITBRAIN_MIN_OBS unreachable is used instead. - tools/ab/ab_analyze.py: keep exact enumeration for C(n,na)<=20e6 (7v7), add a seeded Monte-Carlo permutation test (1e6 draws, 0x5eed5eed) with its standard error, a tie-corrected Mann-Whitney U cross-check, a minimum detectable effect line, all-pairs comparisons, and a [bb] shift check. - tools/ab/README.md: document the new analyzer output. |
||
|
|
f91e121965 |
lead capture by range: we under-lead (0.40->0.135) — a RANGE effect, not a power one
Measure, for every shot ModularBot fires at the real DrussGT, the lead we actually applied vs the lead the enemy's motion required, from the recorded live battles (/tmp/tfil_ab2, 70 battles / 490 rounds / 54926 shots, plus a 35-battle powtest replication of a different binary). - requiredLead from an AIM-INDEPENDENT interception solve (bullet speed vs enemy truth), appliedLead from the server-recorded bullet bearing. - capture = applied/required, guarded at 2px lateral lead (1.6% excluded); headline metric is the robust proportional slope. - validation: hits 11.6px mean miss / 80.8% inside 18px, misses 134px, 11.6x separation; 496/496 death + 70/70 owner attributions correct. Direct answer: capture falls with RANGE (capSlp 0.401 -> 0.135, and |err|/tolerance 1.27 -> 7.54) but is FLAT across fired POWER within a band (450+, enemy alive: 0.154 / 0.127 / 0.127). The sub-0.5 long-range shots (1.18% hit) are finishKill endgame shots at a near-dead DrussGT, not a lead-capture failure. A naive linear predictor captures 0.29-0.60; we reach 46-67% of that, so the under-lead is real but capture=1.0 is unattainable against a dodger (oracle required lead). |
||
|
|
e670788eee |
env_reference: the ring mover was NEVER measured offline - correct a false label
The offline-harness audit (`e40c849`, `docs/offline_harness_trust.md`) found that the claim "best offline hit rate of anything measured" for the ring mover was false. The 20.28% figure is a LIVE number: `docs/feature_ab_results.md` and commit `bfdcdf8` record 35 real-DrussGT bridge battles with a server-side event sidecar as the ground truth, and the 6/49 round wins is likewise live. There is no offline measurement of the ring mover anywhere - the offline harness scores GUNS, not movements, and has no movement driver at all. So this was NOT an offline-vs-live calibration failure, which is how it has been described repeatedly (including by the orchestrator). It was a METRIC MISMATCH: a movement arm judged on hit rate instead of damage/run and round wins - and hit rate is precisely the metric that concealed its collapse. The lesson previously attached to this result was therefore the wrong one. The paragraph now says what actually happened and points at the audit. Docs-only; no code touched. |
||
|
|
e40c8493a6 |
Offline harness: audited, calibrated against live, and one real bug fixed
AUDIT (docs/offline_harness_trust.md, new): - Re-ran acceptance_offline_vs_online myself TWICE: 12/12 deterministic guns exact both times (264 ticks/enemyId=1, 244 ticks/enemyId=2), death boundary included. The offline range reproduces the live bot's own per-gun virtual telemetry exactly. - Re-verified the (fireTick, powerBin) wave-pairing fix: exact-key lookup, collisions counted not silently mislabelled; test_wave_pairing 17/17 PASS. - The offline score is the live TELEMETRY (last-100 virtual hit rate) but NOT the live BATTLE score (damage/round wins). Two-level answer, documented. - bmPoint scores up to one tick-step (~17px) PAST its documented aim distance, while the tie-break probe scores exactly the aim point. Real, low-impact, deliberately NOT fixed (point metric is non-default, measured negative, and the committed point baselines would silently change). - bmPoint/bmPath, perfect-info captures, conditional-on-selection live rates, and hit-rate-as-objective-for-movement all catalogued as non-apples comparisons. FIX (unambiguous, fail-before/pass-after): - common_libs/tests/range_guns.nim: buildAllGunDrivers defaulted to enableTmSelector=true, so run_range / analyze_selector / test_power_selection / measure_power_policy spawned gun 13 (TMSelect) - a gun the shipped bot NEVER spawns. The shared VirtualTracker ring is order-sensitive, so those 4 spawns/tick permuted the learning guns' resolution order (the exact confound |
||
|
|
f41cd08718 | BitBrain gate test: fine-grained aim correction vs naive + Pattern (offline) | ||
|
|
1adefaba26 |
DrussGT dodge vs fired power: no movement response once range is controlled
Answers the user's hypothesis that DrussGT dodges low-power shots better. Measured on 70 live battles / 490 rounds / 54939 real shots vs real DrussGT (/tmp/tfil_ab2) and replicated on 35 more battles / 24280 shots (/tmp/powtest). Power is not randomly assigned - our policy caps it by RANGE (TR_POWER_FAR_DIST=200 -> 1.0) and by OUR OWN ENERGY (the slope), so inside a range band power is almost a deterministic function of our energy and a naive low-vs-high comparison is secretly a losing-vs-healthy comparison. Everything is stratified by range band and backed by a within-band shuffled-label null (arrival re-derived, so the null keeps the kinematic channel), a round-cluster bootstrap, and a within-shot CONTROL window 40 ticks later when the bullet is long gone. RESULT: no behavioural response. In band 450+ the raw miss distance at arrival is +8.25 px [+5.39,+11.25] for HIGH power - but per flight tick it is 4.52 vs 4.51 px/tick (delta -0.01 [-0.12,+0.10]), i.e. entirely the 2.13-tick longer flight window of the slower bullet. Fixed-12-tick lateral displacement is flat (55.63 vs 55.47, -0.15 [-1.15,+0.84]) and turn rate / speed are flat. The whole difference is already present 5 ticks after the trigger pull (+4.2 px) and is just as large in the bullet-free control window (+5.6 px), so it is a property of the low-energy situation, not of the shot. Hit rate is flat (0.10 vs 0.09). Corpus/attribution notes: e*=DrussGT (subject), s*=ModularBot, per TrBattleCapture.java; the Tank-Royale owner id is NOT stable across runs and is recovered per battle from fire geometry + the energy decrement, cross-checked on 496/496 death events. Geometry validated on the server's own hits (mean miss 11.6 px, 80.6% inside the 18 px radius). |
||
|
|
c8b2a8c6c5 |
docs: reverse-engineer the BitBrain algorithm from its C source; assess as a gun
Reads docs/BitBrain_C_code.zip end to end (full_mnist_2048.c 568 lines + bitarray.h), reconciles every weight/data file's byte size against its loader, and answers two questions: 1. What the algorithm is: signed thresholded address decoders (random projections) -> write-once sparse binary coincidence bit tensors -> counting/argmax readout. Thresholds and ADs come from an unsupervised program that is NOT in the zip; only the SBC tensors are learned here. The supervised rule is genuinely online, single-pass, order-free, and has no learning rate - so it satisfies ModularBot's reset-on-enemy-change constraint. Found and isolated a real bug: read_from_sbc's uint8_t bit_test truncates the 32-bit bit test, so only ADEs with i%%32 < 8 are ever counted. Shipped reader: 96.540%% (reproduced exactly); with the bug fixed: 97.210%%. 2. Whether it can become a gun: not on this evidence. Measured full inference at ~0.56 ms/sample (cost is not the blocker), but the AD layer is unobtainable, and this repo has already shown the binding constraint is the target signal, not the learner - the TM head scored 2.08pp below its own majority class and the AD/SBC primitives are already in the 16-family shootout as WiSARD/Bloom/SDM. Recommends a readout swap on the existing WiSARD feature extractor and a Gate-2b-style >=80%% side-accuracy gate before any port. Also: recommends committing the 11 MB zip under its explicit name. |
||
|
|
036979e78e |
env_reference: verify every knob against the code, fix the trap that cost real time
DOCS ONLY. The user set TMH_NSTATES, which is a compile-time {-d:intdefine.}
(-d:TMH_NSTATES=2, tm_horizon.nim:102), not an env var; the runtime var is
TR_TMHORIZON_NSTATES (tm_horizon.nim:139). This rewrites the reference so that
class of confusion cannot recur.
What was WRONG and is now fixed:
- "unparseable warns and falls back" was false for the numeric/bool knobs: only
the enumerated string knobs warn; envInt/envFloat/envBool fall back silently.
- TR_POWER_LOG/TR_RAM_LOG/TR_MOVEMENT_LOG/TR_RECORD_WORLDSTATE/
TR_RADAR_FORCE_SPIN/TR_RADAR_SCANLOG/TR_TRACKER_PROBE are read with existsEnv,
so TR_POWER_LOG=0 turns the log ON. Documented per knob.
- TR_POWER_ENERGY_MIN is a CAP at low energy, not a minimum-power floor.
- TR_TMHORIZON_* knobs are inert unless TR_RACK_TMHORIZON=both; the doc implied
they were live.
- GUN_SELECTOR_WINDOW is clamped 1..100; TR_MOVEMENT silently falls back to tfil
for any value other than tfil_ring.
- The ring file's own header comment (corridor 5 / wall 10) is stale; the code
defaults are 10.0/15.0 (commit
|
||
|
|
9bf3005850 |
Boot-time env report + fix the env reference
The bot is spawned by the server/GUI, so it inherits the SERVER's environment. The user could not tell whether their exports reached the bot, so print a one-shot greppable report at boot: grep '^\[env\]' /tmp/modularbot_stdout.log Section A prints every TR_*/GUN_* this process actually received, the count vs the total env size, a loud warning when nothing matched, and the process identity (pid/ppid, cwd, self command line, and the PARENT command line) so the spawn trap is obvious. Section B prints the resolved effective value of every documented knob with its source (env|default), including clamps and the rack's empty-set fallback. Build identity (NimVersion, compile date/time, binary path/size/mtime) pins the exact artifact. Suppress with TR_ENV_REPORT=0. docs/env_reference.md: add the missing GUN_SHOTLOG_PATH, GUN_SELECTOR_MINOBS/FLOOR/POOL/RANK/SHRINK/SEED, TR_ENV_REPORT and -d:TM_NCLAUSES; record the measured TR_TMHORIZON_WINDOW verdict; and add a prominent 'Did my env vars actually reach the bot?' section with the boot report, the /proc/PID/environ no-code check, the correct GUI launch recipe, and how to prove the trap deliberately. |
||
|
|
b68707c867 |
Energy economy: the cliff becomes a SLOPE, plus a finishing cap. 11% less energy.
The user's request: "when our bot is low OR enemy is low, it is useless to use high power instead low fast bullets have more chances to finish the enemy. Let's do a math slope: starting from some health down, the power goes down with it." 1. ENERGY SLOPE (`TR_POWER_ENERGY_*`), replacing the old hard step at 50 energy: cap = ENERGY_MAX at/above ENERGY_HI, ENERGY_MIN at/below ENERGY_LO, LINEAR in power between, clamped. Defaults HI=80 LO=20 MIN=0.5 MAX=3.0, so no cap >=80, 0.5 at <=20, and e.g. E=65 -> 2.375, E=50 -> 1.75, E=35 -> 1.125. Rationale: bullet speed is 20-3p, so lower power = FASTER bullet (less lead error, higher hit chance), fires more often (10+2p) and drains slower (p/shot). E[dE] = p(3P-1) => break-even hit probability is 1/3 INDEPENDENT of power, and our measured rates are 5-27%, far below it. 2. FINISHING CAP (`TR_POWER_FINISH_KILL`, default ON): cap power at the SMALLEST bullet that still removes the enemy's remaining energy - `E<=4 -> p=E/4` (min 0.1), `4<E<=16 -> p=(E+2)/6`, `E>16 -> no cap`. Rationale, and it makes the user's instinct stronger than a heuristic: server 1.3.1 caps the damage SCORE at the energy ACTUALLY REMOVED, so overkill is WASTED damage AND ~6x the energy for ZERO extra score. Damage is 4p (p<=1) / 6p-2 (p>1). Both are min-composed with the existing far/below-average caps, may only LOWER power (exhaustively tested), and are exempt while ramming. `TR_POWER_POLICY=0` still returns the uncapped control exactly. MEASURED ENERGY SAVING (offline replay of the DrussGT fixtures, 28,797 ticks): arm shots energy meanP E/1k ticks vs cliff control(uncapped) 1913 4646 2.43 161.4 -90.2% cliff (today) 2363 2443 1.03 84.8 0.0% slope 2404 2178 0.91 75.6 ** 10.9% LESS ** slope+finish 2404 2167 0.90 75.3 ** 11.3% LESS ** So the slope spends ~11% less energy than the cliff AND fires slightly MORE shots (2404 vs 2363) - both directions at once. HONEST NOTE on the finishing rule's reach here: ticks where the enemy is low (0 < E <= 16) are only 2252/28797 = 7.8% of these fixtures, so finishing adds just ~11 energy of saving against DrussGT. It matters in CLOSER fights, not this one. Verification: test_power_policy 58 (was 26) in BOTH the default and TR_POWER_POLICY=0 control arms - slope at E=100/80/65/50/35/20/5, powerToKill across E=0.1..100, the inverse-cover property for E<=16, monotonicity, ram exemption, and an exhaustive sweep proving power <= preference. Guards: test_gun_harness 39, test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41, test_tfil_ring_weights 24, test_ram_decision 40, test_rack_membership 48, test_selector_tiebreak 19, test_tm_pattern_registration 20, test_vbullet_admit_gate 12. acceptance 12/12 PASS. ModularBot compiles release. Adds `common_libs/tests/measure_power_policy.nim` (the energy/histogram tool) and updates docs/env_reference.md for the new `energySlope|finishKill` log reasons. NOT MEASURED: the battle/hit-rate effect. The offline figures use the fixture shooter's energy as a proxy, open-loop; the RELATIVE saving is the meaningful part. |
||
|
|
9ba932d1b1 |
docs: complete environment variable reference (every knob, its default, and why)
Single place to answer 'what env vars exist, what do they default to, and what are
they for'. Grouped by area, with the measured reason each knob exists recorded
next to it, since several defaults are counterintuitive:
- The shipped rack is PATTERN ONLY, with the revert one-liner included.
- GUN_SELECTOR_TIEBREAK defaults off because it measured NEGATIVE on real hit
rate, and the per-tick random draw inside the tie band is load-bearing
(commitment cost 7.02% -> 5.10%, p=0.002).
- The power policy is ON because turning it off measured 10.6% -> 7.9% real hit
rate; TR_POWER_FINISH_KILL exists because server 1.3.1 caps the damage SCORE at
the energy actually removed, so overkill scores nothing.
- The ring mover is NOT the default and must not be shipped: best offline hit
rate of anything measured, but it halves survival (16/49 -> 6/49, p=0.012).
- The proactive ram is off because it converts 0/6 times.
- TR_PATTERN_RAD_* are kept but are structurally incapable of changing the shot
(the aim is bearing-only, measured byte-identical on the path metric).
- TR_TMHORIZON_NSTATES is the automata inertia ('mood'): lower = adapts faster.
- TR_VBULLET_ADMIT_ONLY=1 gave +68% tick rate (87 -> 146 ticks/s).
Documents the read-once-at-start convention and the operational trap that the bot
is spawned by the server/GUI, so a variable exported in an unrelated terminal does
NOT reach it. Also lists the compile-time -d: knobs and the test-harness jars.
Snapshot of the source at this commit; f9f8d84-era knobs added by the in-flight
jobs are included where already present.
|
||
|
|
eb74f9b2e3 |
Ram: finisher-only by default, and the bullet-rain abort now measures real energy
Follows the diagnosis that proactive straight-line ramming CANNOT work: both bots have MAX_SPEED=8, so a pursuit cannot catch an evading equal-speed opponent. Measured over 49 rounds per arm, opportunity -> contact was **0/6** (base), 0/40 (ring), 0/12 (ringhot). The only proactive conversion in the whole corpus came from a FINISHER, and only because a <20-energy DrussGT stops fleeing (that episode closed at 6-8 px/tick). Opportunity episodes never got below ~80px; one ran the full 60-tick duration cap and closed only 198->171px; a perfectly aligned full-speed one closed 195->114px then plateaued. CHANGES - **Finisher-only default.** `finisher` (<20 energy, dist<300, we are healthier) and the rare `desperation` (both <5, dist<150) are kept; `opportunity` and the speculative `plan` are OFF. Both are env-reenableable with no rebuild: `TR_RAM_OPPORTUNITY=1` (tune via TR_RAM_OPP_DIST/MARGIN) and `TR_RAM_PLAN=1`. Justification: it removes 100+ non-converting episodes per fixture at zero measured loss (oldram vs base was p=0.69, damage 279 vs 284, survival 17/49 vs 16/49) - and each of those episodes spent up to 60 ticks driving STRAIGHT at the enemy, abandoning the mover's dodging and disrupting aim. - **`desperation` KEPT** deliberately: it is cheap and rare, fires only when both bots are nearly dead at short range (a coin-flip where 0.6 contact can decide it), and it is not the refuted straight-line pursuit. - **THE BULLET-RAIN ABORT WAS DEAD CODE AND IS NOW FIXED.** `onHitByBullet` accumulated raw bullet FIREPOWER while `TR_RAM_ABORT_DMG = 0.5` was documented as a DAMAGE rate - so the bar was implicitly "sum of power > 7.5 over 15 turns" and the maximum rate ever observed was 0.27. It now accumulates REAL ENERGY via a `bulletDamage(power)` helper matching the server's `4p` / `6p-2` formula, and `TR_RAM_ABORT_DMG` defaults to **2.0 energy/turn** (~30 HP over 15 turns): "abort an in-progress ram if we take > 2.0 energy per turn". Same effective bar for normal firepower, and it can now actually fire - the live run reports `dmgRate=1.07/turn` where the old units said 0.27. - **`ramStuckTicks` REMOVED.** It required `dist < 5px`; contact occurs at ~36px (two 18px radii) and position rewind prevents getting closer, so it could never increment. Only the 60-tick duration cap can now self-end a ram. LIVENESS (measured, default config, vs a charging Java RamFire, 3 rounds): default -> `[ram] ON reason=finisher` x3, `reason=opportunity` x0 TR_RAM_OPPORTUNITY=1 -> `reason=opportunity` x4, `reason=finisher` x2 So the opportunity states DID occur and are suppressed by the new default - the removal is real, not an arm that never fires. A line also read `[ram] OFF reason=duration dmgRate=1.07/turn`, confirming the new energy units. Adds docs/ramming_negative_result.md (70 lines) recording the question, the five diagnostic answers, the geometric reason, the finisher exception, the two dead code paths, and an explicit "do not re-attempt a proactive straight-line ram; if point-blank forcing is ever wanted it is an INTERCEPTION/cornering movement problem" note - the same pattern that stopped the corpse bug recurring. Guards: test_ram_decision 40 (was 28), test_gun_harness 39, test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41, test_tfil_ring_weights 24, test_power_policy 26, test_rack_membership 48, test_selector_tiebreak 19, test_tm_pattern_registration 20, test_vbullet_admit_gate 12, acceptance_offline_vs_online 12/12. ModularBot compiles. Honest note: the abort-threshold fix is a real (tiny) behaviour change, NOT measured-neutral - it only bites while a finisher ram is under sustained fire, which is exactly the user's stated wish. The finisher-only removal itself is measured-neutral per the given A/B. |
||
|
|
99f9532f55 |
Melee vs 1v1 racks: mechanism built, but NO gun-level payoff - do not split
The user's plan was separate melee and 1v1 racks. The mechanism is built and
committed (TR_RACK_<GUN>=both|1v1|melee|off, mode from server truth). This job
produced the missing evidence: each candidate gun forced ALONE in BOTH modes,
one frozen binary, env-only arms, exact two-sided permutation tests.
1v1 vs the real DrussGT (6 arms x 5 runs x 7 rounds):
Pattern 10.80% 292 dmg/run 11/35 round wins
KNN 5.58% 128 1/35
Linear 2.99% 59 0/35
Circular 2.89% 60 0/35
WallBounce 2.72% 57 0/35
GuessFactor 2.14% 38 0/35
-> every arm differs from Pattern at p=0.0079 (the 5v5 floor). Per-gun damage
span 7.7x. In 1v1 the gun matters ENORMOUSLY.
Melee (+DrussGT, RandomMover, WaveSurfer, OscillatorBot; enemyCount 4):
WallBounce 25.79% 643 dmg/run 7/25
Linear 24.88% 624 7/25
GuessFactor 24.48% 613 4/25
Pattern 24.39% 599 5/25
Circular 25.88% 596 5/25
KNN 23.87% 539 3/25
-> NO arm beats Pattern with significance (p=0.42-0.96, fully overlapping).
WallBounce's nominal +7.3% is p=0.42. Per-gun damage span only 1.19x.
THE FINDING: in 1v1 the gun matters enormously (7.7x spread); in melee it barely
matters (1.19x). Melee is won by movement, survival and placement, not by which
gun you carry - every candidate lands in the same ~24-26% band. So a separate
melee rack has NO gun-level payoff and DefaultRackMembership stays Pattern-only.
The mechanism remains available if it is ever wanted.
Honest caveats: the melee comparison is UNDER-POWERED at 5 runs (detecting the
~44-damage WallBounce-Pattern gap would need ~2.5-3x the runs), so "no significant
difference" is NOT "no difference"; the only candidate worth re-testing is
WallBounce in melee, and it must not be shipped on this evidence.
SHIM LIMITATION FOUND, worth recording: run_bridge_battle.sh captures with
`--subject DrussGT`, and the shim DROPS every tick once the subject dies - losing
5-20% of ModularBot's shots in melee. The campaign was re-run with
`--subject ModularBot` so ModularBot's events are complete (fires match gun_stats
realShots exactly). Any future melee evidence through this shim must do the same
or it will silently under-count.
Adds docs/melee_vs_1v1_racks.md. No repository source changed.
|
||
|
|
bfdcdf8919 |
The three unmeasured features, A/B'd - and a methodology correction I had wrong
Five arms x 7 runs x 7 rounds (35 real-DrusGT battles, 8 concurrent), one frozen
binary from git archive HEAD at
|
||
|
|
31c7c01d28 |
SHIPPED: the default rack is now Pattern-only (+49% hit rate, +66% damage on the boss)
`DefaultRackMembership` now admits Pattern (id 5) and marks all 14 other guns `rmOff`. **The selector mechanism is untouched** - `chooseFromFit`, the floor/band logic, the hysteresis and the virtual-fitness plumbing are all intact and functional. Only the rack membership changed, so this is reverted by env alone. Evidence (measured, replicated three times, 10 adversaries): Pattern alone gives 10.36% real hit rate / 264 damage per run vs the full rack's 6.93% / 159. Pattern significantly wins on DrussGT, Corners, Crazy and PatternMover, ties on three, and the full rack never significantly beats it on ANY adversary. Mechanism: the virtual signal keeps ranking the wrong guns first (HeadOn 46% of ticks at 2.0% real; Linear 57.7% at 6.0% real while Pattern sits at 11.1%). **THIS CONTRADICTS THE USER'S STANDING DIRECTIVE** to keep virtual-fitness selection. Recorded plainly in docs/selector_negative_value.md with a SHIPPED DECISION banner rather than done quietly: the mechanism is retained and one env var away, because the measurement says it is negative value on every rack size tested and on 10/10 adversaries. Revert one-liner (no rebuild): TR_RACK_PATTERN=both TR_RACK_HEADON=both TR_RACK_LINEAR=both TR_RACK_TSETLIN=both \ TR_RACK_CIRCULAR=both TR_RACK_GUESSFACTOR=both TR_RACK_WALLBOUNCE=both \ TR_RACK_ACCEL=both TR_RACK_STOPSHOT=both TR_RACK_DISPLACE=both TR_RACK_AVGLEAD=both \ TR_RACK_DECAYGF=both TR_RACK_KNN=both TR_RACK_TMSELECT=both ./out/ModularBot The unit test `testRevertOverrideRestoresFullRack` exercises exactly this table. FLOOR PATH, verified not assumed: `chooseFromFit` already returns `admitted[0]` on the floor path, so it respects admission by construction. Cold field + shipped default -> floor returns Pattern (id 5), NOT HeadOn. With an explicit all-`both` membership the same cold field returns gun 0 (HeadOn) - the old behaviour. Four assertions in `testFloorRespectsAdmission`. LIVENESS: one 1-round battle with NO overrides -> Pattern selected 105/105 = 100%, every other gun 0 including TMPattern. Honesty caveat retained in the doc: 4 of the 10 opponents were Tank Royale sample-bot PORTS rather than the original classic jars (only DrussGT is a real classic jar through the shim). Guards: test_rack_membership 48 (was 38; new floor/revert/default checks), test_tm_pattern_registration 20 (5 checks hard-coded the old default and were updated to assert the new one, with the TMPATTERN parity proof moved onto an explicit old-rack table), test_gun_harness 39, test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41, test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 28, test_selector_tiebreak 19, test_tm_pattern_rack_live 4, test_tm_pattern_learning 3, acceptance_offline_vs_online 12/12 VERDICT PASS. ModularBot compiles. FOLLOW-ON THIS EXPOSED: membership filters SELECTION but not virtual-bullet SPAWNING, so under `onlyPattern` the 13 unselected guns still predict and spawn every tick. Tsetlin alone is ~5.3 ms/tick (~41% of the 13.16 ms per-tick budget), so we are still paying for it while never using it. Gating spawn on admission would reclaim that; it was deliberately NOT done here because it would alter the measurement protocol mid-A/B. |
||
|
|
394b3deeed |
SETTLED: Pattern alone is the best single gun IN GENERAL, not just vs DrussGT
Closes the one-adversary caveat that blocked shipping `onlyPattern`. Two arms
(`full` vs `onlyPattern`, rack knobs only), one frozen binary from clean HEAD
(
|
||
|
|
a54ae6a162 |
SETTLED: no small rack beats Pattern alone; the selector is negative value on a GOOD rack
Follow-up to |
||
|
|
e0666a562d |
The gun selector is NEGATIVE value: Pattern alone beats the full rack (p=0.0012)
MEASURED against the real DrussGT, one frozen binary built from clean HEAD, rack knobs only (no source edits), 5 arms x 7 runs x 7 rounds, 8 concurrent battles, judged ONLY on server-side real hit rate from the events sidecar, exact two-sided permutation test on per-run rates. arm runs shots hits real % dmg/run p vs full full (shipped) 7 3898 270 6.93 159 -- onlyPattern 7 4582 494 10.78 287 0.0012 <- BETTER onlyKNN 7 4033 207 5.13 119 0.1340 onlyLinear 7 3215 105 3.27 65 0.0082 onlyGF 7 3193 72 2.25 45 0.0012 Firing Pattern ALONE gives +3.85pp pooled hit rate and +80% damage per run, and it fires MORE shots (4582 vs 3898) - it dominates on rate and volume. This is not "any single gun wins" (full beats Linear, GF and KNN); it is specifically "Pattern alone beats the rack". WHY - the virtual fitness signal mis-ranks guns against real outcomes: - HeadOn is massively over-selected: 31.4% of ticks, the most real shots (1070), but only 4.5% REAL. It alone drags the rack down. - Pattern has the best virtual rank and near-best real rate (11.9%, rank 2), yet is selected only 22.6% of the time. - Linear's apparent strength was SELECTION BIAS: conditional on being selected it looked like 15.2% (n=33), but its UNCONDITIONAL rate (onlyLinear) is 3.27%. Every earlier per-gun "real rate" in this repo is conditional on selection and is therefore confounded. This experiment is the clean measurement. NOT YET SETTLED (do not overclaim): - ONE ADVERSARY. All of this is vs DrussGT. Pattern must be re-checked against other bots before it becomes the default on this evidence alone. - Whether a SMALL rack of good guns beats Pattern alone. The selector is negative value on the CURRENT bloated rack; that does not prove it is negative value on a rack of only good guns. That is the next experiment and it decides whether the selection apparatus is fixed or disabled. - The user's standing directive is to KEEP virtual-fitness selection. This measurement conflicts with it, so the next step tests the selector on a small good rack rather than assuming either answer. Context - three prior selection-side attempts all failed: hysteresis (7.02% -> 5.10%, p=0.002), commitment (7.17% -> 4.44%, p=0.0012), arrival-accuracy tie-break (7.08%, p=0.88 null). The per-tick random draw is load-bearing on three independent measurements. This experiment locates the real problem one level up: which guns are in the rack, and that the virtual signal ranks them wrongly. Preserves the reusable harness (tools/ab/which_gun_run_one.sh, which_gun_arm_env.sh, which_gun_analyze.py) and the full writeup (docs/selector_negative_value.md). |
||
|
|
0ede6d12ec |
selector: arrival-accuracy tie-break measured NEGATIVE; randomness is load-bearing
Hypothesis under test (from the gun audit, which named the tie-band as "the lever that matters most"): `bmPath` is deliberately generous (2.3-3.6x `bmPoint`), so a gun can sit in the tied band on a ray that sweeps the target's path while its bullets ARRIVE badly. So: keep the `path`-ranked band (path beat point on real hit rate 7.43% vs 4.70%, z=5.56), but narrow the random draw inside it using a parallel `point` (arrival-accuracy) window. RESULT: NO EFFECT. Real DrussGT, ONE frozen binary (/tmp/ModularBot_tieband, md5 2c0c56e6...), env knobs only, 7 runs x 7 rounds per arm, server-side events sidecar, exact two-sided permutation test on per-run rates. arm runs shots real % dmg/run d p tbbase (shipped) 7 4128 7.17 175 -- -- tbpt path-rank + point-narrow 7 3938 7.08 165 +0.14 0.88 tbpc =commit control 7 3759 4.44 98 +2.74 0.0012 tbpt25 point margin 0.25 7 3683 5.59 119 +1.65 0.20 tbtie05 / tbtie40 (band width) 7 3937/3917 5.84/6.28 133/144 1.49/1.00 0.11/0.25 tbwin50 (SelectorWindow=50) 7 3983 6.05 139 +1.20 0.11 tbfloor10 (FloorPeakFrac=0.10) 7 3829 5.33 118 +2.12 0.11 tbpt vs base: fully overlapping ranges, p=0.88. This is a REAL null, not a dead arm - the mechanism was live, and it visibly changed the selected-gun mix (Pattern 24%->16%, Accel 6%->16%, Tsetlin ~0%->13%). CONTROL VALIDATED, AND THIS IS THE THIRD TIME: removing the random draw inside the band is SIGNIFICANTLY WORSE (4.44%, p=0.0012). Combined with the earlier hysteresis A/B (7.02% -> 5.10% for commitment) and the light-hysteresis result, the selector's per-tick randomness is now load-bearing on three independent measurements. Narrowing the band on ANY second virtual statistic has not helped. Every knob swept (band width, floor, window) is nominally worse than shipped at n=7; that is "no credible win" rather than "proven harm" (sd ~1.8pp, ~1pp resolution, underpowered). Shipped default stays `GUN_SELECTOR_TIEBREAK=off`; the feature is opt-in, fully guarded, and costs zero extra work on the default path (point windows are scored only when the mode is on). Guards: test_selector_tiebreak 19 (new, pure), test_gun_harness 39, test_vbullet_metric 11, test_adaptive_radar 41, test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 28, test_rack_membership 38, acceptance_offline_vs_online 12/12 PASS (offline path calls neither chooseFromFit nor the tie-break). STRATEGIC CONCLUSION: three selection-side attempts have now failed (hysteresis, commitment, point tie-break). The selector is at a local optimum and the remaining lever is the QUALITY OF THE GUNS, not the selection among them. |
||
|
|
fb36a0a685 |
tracker: corpses do not exist - revert the fix and retire the workaround
The belief "BotDeathEvent never reaches ModularBot, so enemyTracker keeps dead
enemies alive forever" was written into a code comment and then believed twice.
It is FALSE. Measured in a 7-bot melee with a per-tick probe comparing
enemyTracker's alive count against the server's getEnemyCount():
metric 1.3.1 (20 rd) 0.35.5 (15 rd)
observed enemy deaths 83 68
...non-round-ending 83 (100%) 66 (97%)
ekBotDeath events DROPPED 0 0
max dispatch lag (turns behind) 1 1
phantom ticks 1 / 16,820 1 / 12,596
MAX CORPSE LIFETIME 0 ticks 0 ticks
victims still alive at round end 0 0
onBotDeath fires for every death, including non-round-ending ones. The
API-level event-drop mechanism IS real (test_event_drop_mechanism.nim proves
it: ekBotDeath is not in isCritical and MAX_EVENTS_AGE=2) - the bot simply
never falls far enough behind for it to trigger (max lag 1 turn).
Removed:
- reconcileWithServer + ReconcilePersistTicks/mismatchTicks/sawServerAlive
(uncommitted, and ON BY DEFAULT despite the premise being false). Its own
comment admitted a shorter window once KILLED A LIVE ENEMY ("it fired three
more times after the tracker marked it dead") - a latent mis-prune path
defending against a bug that does not exist.
- The radar's CorpseTicks=40 filter and the same-class age>60 filter in
recordRadarStats, both carrying the false comment. Removal changes no real
behaviour: buildState feeds the radar enemyTracker.allAlive(), so a dead
enemy never reaches computeScan.
Kept:
- The TR_TRACKER_PROBE instrument (default OFF), which produced the table above.
- test_event_drop_mechanism.nim - the drop mechanism is a genuine library
behaviour worth guarding.
- isAlive/aliveCount on the tracker.
Added: docs/tracker_death_events.md (the durable negative, so this is not
re-invented a third time) and test_enemy_tracker_death.nim (13 checks) in place
of the test for the deleted feature.
Guards: test_gun_harness 39/39, test_vbullet_metric 11, test_power_selection 3,
test_adaptive_radar 41/41, test_event_drop_mechanism 6, test_enemy_tracker_death
13, acceptance 12/12, ModularBot compiles.
|
||
|
|
013b9fe01e |
docs: correct the overstated 'inverted metric' claim; record tonight's fixes
Three corrections, all prompted by later measurements: 1. The virtual-vs-real rank correlation is NOT robustly negative. Six independent Spearman measurements now exist (-0.374, +0.335, +0.522, -0.371, -0.073, -0.037) and the sign flips on large samples, so it is near zero on average. The honest headline is that virtual hit rate is a POOR RANKER, not an inverted one. The report said 'not weak - it is inverted' in six places; it now says so in none. The practical conclusion (do not trust it for ranking) is unchanged; the mechanism claimed was wrong. 2. The offline==online acceptance is FIXED, not flaky. Root cause was that the replay spawned gun 13 (TMSelect) while the live rack has it disabled, and the shared VirtualTracker ring is ORDER-SENSITIVE, so gun 13's extra 4 bullets/tick permuted the per-tick resolution order for every other gun and shifted the learning guns' observations. After closing gun 13's ready gate offline the live and offline KNN traces are byte-identical (904/904 lines, empty diff). 5/5 consecutive runs now report 12/12 exact with the death boundary included. Recorded with the lesson: a flaky proof was hiding a real bug. Also records the general A/B confound - disabling a gun removes its 4 spawns/tick from the shared ring, perturbing resolution order for the rest. 3. Pruning was tested and does NOT help, so the verdict for Tsetlin and Displace changes from an implied drop to BELOW OVERALL - KEEP. 15 paired runs: baseline 6.18%, Tsetlin-off 5.76%, Tsetlin+Displace-off 5.46%; paired permutation p=0.57 and p=0.21; distributions completely overlap; a non-surfer control showed no separation. Being below average does not justify removal. Also records the tie-break randomness fix, and quotes run counts with every rate (6.95% over 13 runs vs 6.18% over 15 runs, same binary) rather than presenting a single figure as definitive. |
||
|
|
19410164f1 |
docs: definitive gun-rack report on real measured numbers
Replaces the stale 2026-09-20 docs, which predated per-gun real attribution, the offline gun range and the DrussGT boss, and whose verdicts were built on virtual hit rates that turned out to be ANTI-correlated with reality. docs/gun_rack_analysis.md (751 lines) covers: the test infrastructure described honestly (offline range with its flaky-acceptance caveat, the 20 fixtures and what each set is good for, the live boss, and the A/B methodology of per-run server-side real hit rate with an explicit overlap test); the virtual-vs-real metric lesson with Spearman -0.374 and the point-vs-path A/B; per-gun real performance and the 16-rule ranking A/B; the offline per-fixture gun matrix; KEEP/MARGINAL/BELOW verdicts; and the five root-cause bugs with before/after numbers. docs/gun_rack_summary.md (58 lines) is the verdict table plus top actions. The '~230 point' score-noise band that has been steering methodology all night was re-derived from the artifacts rather than asserted: the 13 shipped-config run scores span 175-526, s.d. ~105, i.e. a ~210-point 2-s.d. band. Caveats recorded verbatim rather than softened: the offline==online acceptance is flaky (typically 11/12 on unmodified HEAD), fixtures are perfect-information and therefore optimistic vs live play, per-gun real N is small so single-gun ordering is indicative, the headline numbers come from ONE wave-surfer adversary, and HeadOn must stay despite being lowest because it is the floor fallback (disabling it: 5.08% / 175 dmg vs 6.95% / 251 dmg). |
||
|
|
76ae6170f8 |
docs(research): portability audit of DrussGT 3.1.4159
Measured hazard counts for a Java->Nim port, with the correction that only 22 of 28 top-level classes ship source (6 do not: 5 in the gun package plus dMove/Scan), so a full port would need a decompiler while a shim would not care at all. Real traps: 193 float / 35 casts / 147 literals / 60 float[] in the movement closure (the danger histogram is float[171] -- porting 32-bit Java floats to Nim's default float64 diverges silently); ~68 non-final statics; and 4 Java single-& sites with side effects, which break under Nim's short-circuiting 'and'. Non-issues, correcting earlier assumptions: 0 sites of %-on-negative (angle normalisation is floor-based) and FastTrig has no lookup tables, it is 7 coefficient-exact polynomials. Movement scoping: 5007 LOC across 14 files, ~4.8-6.1k Nim LOC, 4-8 focused agent-days to first-compiles. It can be ported WITHOUT the gun (data flows movement->gun only), but the harness never routes HitByBullet into movement modules, so the danger bins would never train -- that plumbing is the real blocker, not the translation. |
||
|
|
54e9757567 |
docs(research): TM learning tracks + note that the TM-DEB source was deleted
tm-learning-tracks.md covers three things, all marked [FACT]/[INFERENCE]/ [UNKNOWN]: - Section A: the Tsetlin gun's label is measured against the wrong baseline. predX = linearX + cx, so rx = actual - predX = delta - cx, and inside tmLearnOne error = residual - predicted = (delta - cx) - cx = delta - 2cx. The fixed point is cx = delta/2 -- HALF the correction needed, even with perfect Granmo feedback. Fix: store linearX/linearY in TmTrace and train on delta. Also: hits zero the label instead of carrying their true residual, and the per-clause step is magnitude-blind. - Section B: what a TM is actually good at (AND-clauses over binary literals, readable output) and why this repo suits it -- the gun already builds an 83-bit x 10-frame Gray-coded window (870 bits). Includes a falsifiable known-rule benchmark proposal. - Section C: delayed-reward learning belongs to the MOVEMENT layer, not the gun. The gun's outcome is delayed but exactly pairable via (fireTick, powerBin), so its effective lambda is 1 and discounting would only destroy information. Also records that docs/papers/tm-deb-paper.pdf was deleted by the user as AI-generated and unverifiable, while the Granmo-based feedback diff in tm-deb-assessment.md stands on its own. |
||
|
|
669f9acd41 |
docs(research): assess TM-DEB paper against the Tsetlin gun's real failure
Verdict: the document (docs/papers/tm-deb-paper.pdf, 'Generated by Gemini Notebook') solves temporal credit assignment under delayed reward, which is not our problem. Our gun's failure is clause saturation (~131 of 1740 literals included per clause -> conjunction fires with probability ~2^-131 -> correction identically 0), and TM-DEB only scales update FREQUENCY by gamma^dt, so it would leave the fixed point untouched and additionally delete the long-range feedback we need: gamma^90 = 4.4e-7 at the paper's own gamma=0.85, and those are the shots whose lead matters most. Credibility signals recorded in the doc: reference [2] misattributes authors/venue/year; Algorithm Spec 2 steps 9-12 are literal '%' placeholders so the automata update is simply absent; Table 1 is titled 'Expected' and reports never-measured accuracies; Eq. 12 is not a faithful copy of Granmo's Lemma 2. The audit also produced the actionable result: a line-by-line diff of Granmo Table 2/3 feedback against tmLearnOne, identifying why the automata saturate - Type I never conditions on the clause output so it omits Granmo's c=0, lk=1 -> -1 w.p. 1/s counter-force (true literals ratchet toward Include), Type II's guard cOut==1 AND lits[lit]==0 is unsatisfiable and therefore dead code, and the (T - clip(v,-T,T))/(2T) resource allocation is missing entirely. |
||
|
|
c034eb9d25 |
feat(testing): gun rack gauntlet + analysis reports
- fix(ModularBot): onBulletHitBot → onBulletHit (real hits were never tracked) - feat(ModularBot): per-round gun stats dump to /tmp/gun_stats.jsonl - feat(ModularBot): gun selection counter per round - fix(tests): adversary paths _garage suffix removed from 7 test files - feat(tests): test_gauntlet_5bots.nim — 10-round gauntlet vs all 5 adversaries - feat(tests): analyze_gun_stats.nim — JSONL parser for gun performance tables - docs: gun_rack_analysis.md — full per-gun performance report - docs: gun_rack_summary.md — TL;DR verdict table (keep/drop/tune) Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
1ed7797cb6 |
feat(ModularBot): 6 guns, pattern matcher, melee modules, adversarial bots
- New guns: guess-factor (GF histogram), pattern-matcher (movement tape replay) - New modules: minimum-risk melee movement, spinning melee radar - New test bots: PatternMover, RandomMover, WaveSurfer - Fixed: FeedbackEvent now carries actualX/actualY for proper GF learning - Fixed: TM gun warmup gating + directional residuals - Fixed: circular gun integrated formula + multi-bin omega cache - Fixed: oscillator wall-bounce lockout - Fixed: phantom meteor perpendicular body orientation - 6/6 battle wins across all enemy types |
||
|
|
b509195ee9 |
chore: rename libs→common_libs, all bot dirs to _garage suffix, fix all path refs
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
8ba4bae21a | chore: rename CLAUDE.md → AGENTS.md, move research doc to docs/ | ||
|
|
d2d205e6f6 | Merge branch 'research/ga-parameters' into research/goto-controller | ||
|
|
6241c41e52 | Merge branch 'worktree-agent-a4c06ae3' into research/goto-controller | ||
|
|
cb1bbc35dc |
docs(adr): Evo_Bot neuroevolution gun architecture (#63)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
eae6fc15a2 |
research(Evo_Bot): GA/ES parameter recommendations for ~750-weight neuroevolution (#65)
Extracted concrete numbers from 13 papers in docs/papers/neuroevolution/. Key findings: use CMA-ES or mutation-only truncation GA, mutate ALL weights (not 5%), sigma=0.005-0.01, pop=64-200, no crossover, single elite. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
56e0b306c9 |
docs(research): goto controller algorithm for issue #20
Covers forward/reverse decision, proportional steering with speed-dependent turn rate clamping, and deceleration using the existing getNewTargetSpeed util. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
64f73dd413 |
chore: add agent skills configuration
Add CLAUDE.md and docs/agents/ with issue tracker (Gitea), triage labels, and domain doc consumer rules. |