Files
SirStone 1ed7797cb6 feat(ModularBot): 6 guns, pattern matcher, melee modules, adversarial bots
- New guns: guess-factor (GF histogram), pattern-matcher (movement tape replay)
- New modules: minimum-risk melee movement, spinning melee radar
- New test bots: PatternMover, RandomMover, WaveSurfer
- Fixed: FeedbackEvent now carries actualX/actualY for proper GF learning
- Fixed: TM gun warmup gating + directional residuals
- Fixed: circular gun integrated formula + multi-bin omega cache
- Fixed: oscillator wall-bounce lockout
- Fixed: phantom meteor perpendicular body orientation
- 6/6 battle wins across all enemy types
2026-09-20 00:59:53 +02:00

38 KiB
Raw Permalink Blame History

BNNBot Deep Learning: Gradient-Free Methods for Binary Networks

Context: 690-bit binary input, online battle learning, reward = miss distance from circular wave feedback, binary weights/activations preferred. No gradient descent, no STE, no supervised labels. Backprop of signals (not gradients) is OK.


Evaluation Scorecard Legend

Symbol Meaning
✅ Full support
⚠️ Partial / with caveats
❌ Fails this criterion

Criteria for each method:

  1. No GD — truly avoids gradient descent (no derivatives, no STE, no surrogate gradients)
  2. Binary — works with binary weights and/or activations
  3. RL — can learn from reward signal instead of labels
  4. Online — can update weights tick-by-tick, not batch-offline
  5. Depth — practically trains 3+ layers
  6. Cost — compute per weight update (relative: Low / Med / High)
  7. Finicky — sensitivity to hyperparameters (Low / Med / High)

Method 1: Contrastive Hebbian Learning (CHL) / Equilibrium Propagation

How It Works

CHL runs two phases using an energy-based (Hopfield-like) network:

  1. Free phase: Input clamped, hidden + output units relax to energy minimum → record activations s⁻
  2. Clamped phase: Input + output clamped to target, hidden units relax → record activations s⁺
  3. Weight update: ΔW_ij ∝ s⁺_i · s⁺_j − s⁻_i · s⁻_j

This is purely local — each weight only needs the activations of its two endpoints in both phases. No global error signal flows backward; activity differences drive learning.

Equilibrium Propagation (Scellier & Bengio 2017) is the modern formulation: the clamped phase is a "nudged" version where the output is soft-clamped with a small β factor. In the limit β→0, EqProp provably computes the same updates as backprop, but without passing gradients.

Binary variant (Laydevant et al., CVPRW 2021): EqProp was directly applied to dynamical binary neural networks. The free/clamped phases remain unchanged; binary activations are compatible because the update rule only compares activation products, not derivatives.

Can It Work Without Supervised Targets?

Critical limitation: The clamped phase requires knowing the target output. In standard CHL/EqProp the output neurons are pushed toward a labeled target. To use reward instead of labels, you need a "nudging" signal at the output layer.

Viable workaround for RL: Instead of nudging toward a ground-truth label, nudge the output in the direction that increases reward. If reward is scalar, define a "desired output shift":

  • Positive reward: nudge current output activations → reinforce current pattern
  • Negative reward: nudge away from current output

This converts CHL into a reward-modulated version. It is not a standard use case but has been explored in neo-Hebbian RL literature. The key: the "nudge" is a signal (direction to move), not a gradient. It fits the constraint.

Scorecard

Criterion Score Notes
No GD ✅ Pure activity differences, no derivatives
Binary ✅ Demonstrated in CVPRW 2021 on dynamic BNNs
RL ⚠️ Needs reward→nudge translation; not native but viable
Online ⚠️ Requires two settling phases per step; adds latency
Depth ✅ EqProp scales to multiple layers; credit is local
Cost Med Two full network relaxations per update
Finicky High β nudge strength, settling iterations, energy landscape

Fit Assessment

Strong theoretical foundation, but the two-phase settling requirement means you need the network to relax to equilibrium twice per firing tick. In Robocode's 1ms budget this is expensive. The reward→ nudge translation is non-trivial to implement correctly. Viable but complex.


Method 2: Forward-Forward Algorithm (FF) / Self-Contrastive FF

How It Works

Hinton (2022): Replace one forward + one backward pass with two forward passes:

  1. Positive pass: real/good data through the layer → maximize "goodness" G = Σ(y²)
  2. Negative pass: corrupted/bad data → minimize goodness
  3. Layer-local update for each layer independently: push goodness above threshold θ for positive, below for negative. No signal crosses layer boundaries.

The Self-Contrastive Forward-Forward (SCFF, 2024/2025, Nature Communications) eliminates the label requirement entirely: positive sample = [x_k, x_k] (same input concatenated), negative sample = [x_k, x_n] (two different inputs). The network learns to distinguish same-from-different without any labels.

Goodness Function

G(l) = (1/M) Σ_m y²_m(l) — sum of squared activations in layer l over M neurons.

Learning rule: ΔW ∝ (σ(G - θ) - target) · ∂G/∂W

Wait — does this have gradients? Yes, the standard FF uses a sigmoid-of-goodness loss with gradient descent on the goodness. However: (a) the update is strictly local to each layer, and (b) ∂G/∂W = 2y · ∂y/∂W. For binary activations with sign() this derivative is zero/undefined.

Binary compatibility issue: The goodness gradient requires ∂y/∂W. With hard binary activations (sign function), this derivative is zero everywhere except at the threshold, breaking the goodness update. The STE would be required to push gradients through — which is explicitly disallowed.

Alternative: Replace the goodness gradient with a Hebbian rule — if G > θ for positive data, do Hebbian update on active neurons; if G < θ for negative data, do anti-Hebbian. This is the "spirit" of FF without the gradient. It loses convergence guarantees but preserves the structure.

Adaptation for RL

Since SCFF already generates its own positive/negative pairs without labels, it operates as a representation learning system. For Robocode: the FF layers would learn good binary representations; then a reward-modulated Hebbian output layer would learn to map representations to predictions.

Scorecard

Criterion Score Notes
No GD ⚠️ Standard FF uses local GD on goodness; binary requires Hebbian approximation
Binary ⚠️ Needs Hebbian approximation of goodness; SCFF uses ReLU throughout
RL ⚠️ SCFF is self-supervised (representation learning); needs separate RL output layer
Online ✅ Layer-local, greedy, update after each sample
Depth ✅ By design — each layer trains independently
Cost Low Two forward passes per update, strictly local
Finicky Med Threshold θ, concatenation design choices

Fit Assessment

FF is excellent for building deep binary representations without labels. It cannot directly predict position with reward alone — you'd hybridize it: FF middle layers for feature extraction, reward- modulated Hebbian output layer for prediction. The goodness-gradient problem with hard binary activations is the main obstacle. Good as a representation sub-system; needs hybridization.


Method 3: Predictive Coding Networks (PCN)

How It Works

Each layer l predicts the activation of layer l-1 from above using top-down weights. The prediction error e_l = x_l − μ_l (actual minus predicted) propagates locally upward and downward.

Weight update: ΔW ∝ e_l · x_{l-1}^T (Hebbian on error × lower-layer activation)

This is self-supervised by construction: the "target" of each layer is the actual activation of the layer below it. No external labels needed. With temporal data, each layer can predict the next time frame — perfectly suited to 10-frame temporal windows.

Active Predictive Coding (APC, 2022): Combines predictive coding with RL for robotic sparse reward problems. The prediction error serves as an intrinsic reward signal. Demonstrated on robotic control tasks with sparse rewards.

PCN-TA (2025): Temporal amortization preserves latent states across frames, yielding 50% fewer inference steps and 10% fewer weight updates vs. backprop.

Gradient Status

Standard PCN inference minimizes free energy F = Σ e_l² by gradient descent on activations (not weights). This is allowed — you're adjusting activity, not weights, via derivatives. Weight updates are then Hebbian: ΔW ∝ e · x. No gradient of the loss w.r.t. weights is computed.

Technically pure: the weight update is ΔW_l = η · e_l · x_{l-1}^T, which is purely local and Hebbian-like. The inference phase uses a form of gradient descent on activations, which is a signal-flow operation (activations are signals), consistent with "backprop of signals allowed."

Binary Compatibility

PCN with continuous activations works cleanly. Binary activations are harder: the inference phase requires iterative adjustment of continuous latent variables before thresholding. A common approach: keep latent variables continuous during inference, threshold only for output/communication. This is the "semi-binary" regime used in neuromorphic PCN implementations.

Fit for Temporal Prediction

Our 10 temporal frames map directly to a hierarchical PCN structure:

  • Layer 1: current-frame features
  • Layer 2: short-term dynamics (frames 0-2)
  • Layer 3: trajectory trend (frames 0-9)
  • Each layer predicts the layer below across time

The reward signal (miss distance) can modulate the prediction error magnitude at the output layer, biasing learning toward accurate predictions.

Scorecard

Criterion Score Notes
No GD ✅ Weight update is Hebbian e·x; inference GD is on activations (signals)
Binary ⚠️ Latent variables need continuous inference; output can be binary
RL ✅ APC demonstrated; error = intrinsic reward; reward modulates output layer
Online ✅ PCN-TA designed for online per-frame learning
Depth ✅ Hierarchical by design; 3+ layers well studied
Cost Med Multiple inference iterations per step (reduced 50% in PCN-TA)
Finicky Med Inference step count, precision/ratio of prediction error weights

Fit Assessment

Excellent match for our temporal structure. The self-supervised nature, temporal hierarchy, and local Hebbian weight updates align with every constraint. Main cost: inference iterations. PCN-TA reduces this significantly. The semi-binary regime (continuous latent, binary output) is a realistic implementation path.


Method 4: Target Propagation / Difference Target Propagation (DTP)

How It Works

Instead of propagating gradients, TP learns an inverse model at each layer: a backward network g_l that maps activations of layer l+1 back to layer l. The target for layer l is:

h_l^target = g_l(h_{l+1}^target)

Difference Target Propagation (DTP): corrects for imperfect inverses by using:

h_l^target = h_l + g_l(h_{l+1}^target) − g_l(h_{l+1})

The target for the output layer still requires a supervised target — this is the fundamental limitation. DTP cannot generate its own target; it only propagates given targets layer-by-layer.

Does It Need Supervised Output?

Yes, inherently. The output-layer target must come from somewhere. In RL, you could define the output target as "the output that would have hit the enemy" — but this requires knowing the correct output, which is equivalent to supervised learning. The reward tells you how much you were off, not where you should have aimed.

Additionally: "the target for the final hidden layer is determined by a formula which relies on gradient descent" (from DTP theory paper). This means gradient descent leaks in even in DTP.

Binary Compatibility

DTP can work with binary units (Lee et al. showed this: "it can be applied even when units exchange stochastic bits rather than real numbers"). But non-invertible binary networks cause reconstruction errors to interfere with target propagation updates.

Scorecard

Criterion Score Notes
No GD ❌ Output-layer target computation involves GD; inverse model also trained with GD
Binary ⚠️ Stochastic binary supported but reconstruction error problem
RL ❌ Requires supervised output target; reward alone insufficient
Online ⚠️ Possible but inverse models add computational overhead
Depth ✅ Designed for depth; that's the whole point
Cost High Two networks (forward + inverse) per layer
Finicky High Inverse model training, reconstruction accuracy

Fit Assessment

Not suitable. Despite claiming "no backprop," DTP requires gradient descent for the inverse models and for the output target. Does not fit RL-only constraint. Skip.


Method 5: Reward-Modulated STDP (R-STDP) / Three-Factor Learning / E-prop

How It Works

Three-factor rule: ΔW_ij = η · e_ij · r(t) where:

  • e_ij = eligibility trace: local pre×post spike correlation (STDP-like)
  • r(t) = global reward signal (scalar, broadcast to all synapses)

Eligibility trace: de_ij/dt = −e_ij/τ + STDP(t_pre, t_post) — decaying memory of recent spike correlations.

This handles temporal credit assignment: the trace remembers which synapses fired recently when the (delayed) reward arrives.

E-prop (Bellec et al., Nature Comms 2020): Extends three-factor learning to deep and recurrent SNNs. Each layer l has an eligibility trace that accounts for the temporal dynamics of the LIF neuron. E-prop with reward-based RL (r-e-prop) was demonstrated winning Atari games.

Critical depth finding: "e-prop implemented in a single-layer recurrent SNN consistently outperforms a multi-layer variant." Adding more layers hurts because the eligibility trace at layer l only captures local pre/post correlations — it cannot account for multi-layer credit assignment without approximate gradients. E-prop for deep feedforward networks approximates backprop, not a clean departure from it.

Binary Compatibility

SNNs fire binary spikes by definition (0 or 1), so R-STDP is natively binary-compatible. Weights can also be binarized. The eligibility trace is continuous (a real-valued running average), but that is internal state, not a gradient propagated through the network.

Credit Assignment Through Depth

This is the weak point. R-STDP solves temporal credit assignment (delay between action and reward) but not spatial credit assignment (which layer contributed to the good/bad output). In a 3-layer network with R-STDP applied uniformly, all layers update on the same reward signal regardless of their contribution. This works empirically in shallow networks but degrades with depth.

Scorecard

Criterion Score Notes
No GD ✅ Local STDP trace × reward; no derivatives anywhere
Binary ✅ Native binary (spikes); weights can also be binary
RL ✅ Reward is the direct learning signal; no labels
Online ✅ Per-spike updates; inherently online
Depth ⚠️ Works but degrades with depth; single-layer LSNN beats multi-layer
Cost Low Eligibility trace update O(n_synapses) per tick
Finicky Med τ_eligibility, reward scaling, STDP window

Fit Assessment

Best fit for shallow networks. For a 2-layer binary SNN (one hidden layer), R-STDP is the cleanest solution: no gradients, native binary, direct reward learning, very cheap. For 3+ layers, credit assignment degrades. Consider using it for the output layer in combination with another method for hidden layers. Highly recommended for shallow or hybrid architectures.


Method 6: Node Perturbation / Weight Perturbation

How It Works

Node perturbation (NP): Add noise ξ to each layer's activations, measure reward R, update: ΔW_ij ∝ ξ_j · (R − R̄) where R̄ is baseline reward.

Weight perturbation (WP): Add noise directly to weights: ΔW_ij ∝ ε_ij · (R − R̄)

Key insight: This estimates the gradient without computing it. It is a Monte Carlo gradient estimate — unbiased but high variance.

Decorrelated NP (DNP, 2023): Applying input decorrelation at each layer "dramatically improves convergence by orders of magnitude." Makes NP practical for multi-layer networks.

Variance problem: The SNR of the NP gradient estimate scales as 1/(N·σ_noise²) where N = number of parameters. For a 690-bit → 256-neuron hidden layer, N ≈ 176,640 weights. The variance explosion makes learning essentially random without variance reduction.

For binary networks specifically: Weight perturbation on binary weights means randomly flipping bits and measuring reward change. This is exactly the "simulated annealing on neural net weights" approach. It is well-defined and requires no derivatives. The update becomes:

if flip W_ij causes ΔR > 0: keep flip; else: revert (with probability based on ΔR)

This is pure stochastic search — not gradient estimation. Works cleanly with binary weights.

Scorecard

Criterion Score Notes
No GD ✅ Statistical gradient estimation; no derivatives
Binary ✅ WP maps naturally to bit-flip search
RL ✅ Reward is the direct signal
Online ✅ Update per trial; eligible for online use
Depth ⚠️ DNP scales to 3-9 layers; variance still significant
Cost High Many perturbation samples needed for convergence
Finicky Med Perturbation magnitude σ, baseline estimator

Fit Assessment

The cleanest in principle, but expensive. Weight perturbation on a binary network is essentially a guided random walk through {-1,+1}^N. With N ≈ 176K weights in a hidden layer, random perturbation converges very slowly. Works best when: (a) network is small, (b) few samples needed per weight. For the output layer only (small, ~64 weights to prediction neurons), this is practical. Suitable as output-layer fine-tuner; too slow as sole learning method for deep networks.


Method 7: InfoMax / Information Maximization

How It Works

Linsker (1988), Bell & Sejnowski (1995): Maximize mutual information I(X; Y) between layer input X and output Y, subject to noise constraints.

Local learning rule (continuous case): ΔW_ij ∝ [φ'(y_i)/φ(y_i)] · x_j − W_ij^{-T}

For the binary/ICA case (Bell-Sejnowski): involves a nonlinear function of output activation and the input pattern. In practice, this maximizes entropy of the output distribution, preventing the network from collapsing all outputs to zero or all-ones.

Binary Compatibility

InfoMax with binary activations maximizes entropy H(Y) — binary outputs should have p(y=1) ≈ 0.5 across data. The update rule involves the "score function" of the output distribution, which for binary units is well-defined without derivatives through the activation function itself.

Self-Supervised Nature

InfoMax requires no labels. It maximizes information preserved from input to output. This makes it useful for representation learning in hidden layers — building diverse, non-collapsed binary features. It does not directly learn to predict enemy position.

Limitations for RL

InfoMax is a representation learning method. It maximizes informativeness of intermediate features but has no notion of task reward. You cannot directly optimize miss distance with InfoMax alone. It would serve as a pretraining / regularization layer to prevent dead neurons and maintain diverse binary representations.

Scorecard

Criterion Score Notes
No GD ✅ Local Hebbian-like rule with anti-Hebbian lateral term
Binary ✅ Entropy maximization is well-defined for binary units
RL ❌ Pure representation learning; no reward optimization
Online ✅ Fully online, local update
Depth ✅ Applied layer-by-layer independently
Cost Low Single forward pass per update
Finicky Low Few parameters; entropy maximization is self-stabilizing

Fit Assessment

Excellent regularizer / auxiliary learning rule. Run InfoMax as an auxiliary update on hidden layers to prevent representational collapse while the output layer learns from reward. Very cheap. Use as a complement to primary method, not standalone.


Method 8: Evolutionary / Genetic Algorithms for Binary BNNs

How It Works

Maintain a population of binary weight matrices. Each individual is a complete set of weights {W_l} with values in {-1, +1}. Evolution:

  1. Evaluate fitness (reward = hit rate, negative miss distance)
  2. Select top-K individuals (tournament or truncation selection)
  3. Crossover: combine weight sub-matrices from two parents
  4. Mutation: randomly flip bits with probability p_mut
  5. Repeat

Binary networks are particularly GA-friendly because:

  • The search space is discrete and finite
  • Crossover has a clean interpretation (taking different weight blocks)
  • No gradient computation at all
  • No issues with binary activations (trivially compatible)

Population-Based Training During a Battle

A battle has ~1000+ ticks. With a population of P=10-20 individuals, each tick evaluates the current "best" policy, and every N ticks (after a reward arrives from circular wave) the population updates. This is a form of online evolution.

Key problem: Each individual needs to be evaluated on the same input to compare fitnesses. During a battle, inputs change each tick — you cannot hold them constant while evaluating 20 models. Solutions:

  • Maintain an experience replay buffer; evaluate candidates on stored states
  • Use the single "best" candidate in battle, explore with perturbations (= weight perturbation)
  • Use island model: different individuals fight in different battles

Practical Issues

  • Population of 20 requires 20× memory for weights
  • Fitness estimates from sequential battle experience are noisy (different opponents, different positions each eval)
  • Crossover between weights that serve different layers may break learned structure
  • Convergence within a single 1000-tick battle is unlikely for deep networks

Scorecard

Criterion Score Notes
No GD ✅ Pure fitness-based search
Binary ✅ Perfect; binary is ideal for genetic operators
RL ✅ Fitness = battle reward; no labels
Online ❌ Population requires multiple evaluations; impossible in a single live battle
Depth ✅ Architecture-agnostic
Cost High O(P × network_size) per generation
Finicky Med Population size, mutation rate, selection pressure

Fit Assessment

Not suitable for within-battle online learning. The population-based evaluation requirement breaks against the single-agent, real-time constraint. Best applied between battles (offline evolution over many battles). If we're restricted to online learning, this is out. Viable only as inter-battle meta-learning.


Method 9: Feedback Alignment (FA) / Direct Feedback Alignment (DFA)

How It Works

Feedback Alignment (Lillicrap et al., 2016): Replace backprop's transposed weight matrices W^T in the backward pass with fixed random matrices B. The claim: the forward weights align to B during training, so the random backward signal still provides useful credit assignment.

Direct Feedback Alignment (DFA): Each layer receives credit directly from the output error multiplied by a random matrix B_l (layer-specific). No sequential backward pass needed.

The Gradient Problem

FA and DFA fundamentally still compute a gradient of the loss at the output layer and propagate that signal (or a random-projection of it) backward. The output gradient ∂L/∂y is a derivative. This violates the "no gradient descent" constraint. FA avoids the weight transport problem but does not avoid derivatives entirely.

Is the "signal" interpretation valid? You could argue: the output error (y - target) is a signal (not a derivative). If you replace this with a reward-modulated output unit error, you avoid computing a derivative. But this is a stretch — in practice, FA implementations use ∂L/∂y.

Binary compatibility: FA fails on binary networks because the backward signal through hard thresholds is zero everywhere (the same STE problem as standard backprop). Papers report "FA fails in deep networks, convolutional layers, or architectures with bottlenecks."

Scorecard

Criterion Score Notes
No GD ❌ Output gradient still computed; only backward weight symmetry is removed
Binary ❌ Fails with binary activations without STE
RL ⚠️ Could replace loss gradient with reward signal, but then it's R-STDP
Online ✅ Online update possible
Depth ⚠️ DFA works at depth but with performance degradation
Cost Low Same cost as forward pass + random projection
Finicky Low Random fixed matrices; no tuning needed

Fit Assessment

Not suitable as stated. FA/DFA require output-layer gradients and fail with binary activations. If you replace the output gradient with a reward signal and the backward pass with fixed random projections, you get something closer to random reward feedback — a variant of node perturbation, which is covered above. Skip in favor of R-STDP or node perturbation.


Method 10: Additional Methods Found

10a: Equilibrium Propagation + CHL as a Unified Framework

Recent work (arXiv 2206.02629) shows that CHL, EqProp, and predictive coding all converge to the same limit at infinitesimal inference steps — they are the same algorithm in different dynamical regimes. This means choosing among them is primarily about implementation tradeoffs (settling speed, binary compatibility) not fundamental differences.

10b: Counter-Current Learning (CCL, 2024)

A dual-network approach: one network runs forward (inference), the other runs backward (teaching). The backward network generates targets for the forward network without computing gradients. Used as a biologically plausible alternative to backprop. Requires paired network architecture, doubling memory. Interesting but adds complexity. Skip for now.

10c: E-prop with Reward (r-e-prop)

Specifically the reward-based variant of e-prop demonstrated on Atari. This is three-factor learning with eligibility traces designed for temporal credit assignment, formally shown to approximate policy gradients. Does use gradient approximations internally (via the eligibility trace derivation), but the weight update itself is ΔW = e_ij · r(t) — a local multiplication.

Whether this "counts" as gradient descent: The eligibility trace e_ij is derived from the neuron's dynamical equations and approximates ∂h_l/∂W. This is a gradient of activations w.r.t. weights (signal propagation), not a gradient of the loss. It satisfies "backprop of signals" while avoiding "gradient descent on loss." This is the cleanest theoretical fit.

10d: Noise-Based Reward-Modulated Learning (2025, arXiv 2503.23972)

Explicitly designed for spiking/binary units: natural noise in binary neurons drives stochastic exploration, reward signal modulates which noise patterns are reinforced. Weight update: ΔW ∝ ξ · r(t) where ξ is inherent neural noise (spike timing jitter). Online, local, binary-native. Very relevant — this is weight perturbation where perturbation = natural spike noise.

10e: Self-Contrastive Forward-Forward (SCFF, Nature Comms 2025)

Published result: MNIST 98.7%, CIFAR-10 80.75%, STL-10 77.3% without any labels. Layer-local goodness function. Greedy layer-wise training. The best current result for label-free local learning in deep networks. Needs Hebbian approximation for hard binary activations.


Comparative Scorecard Summary

Method No GD Binary RL Online Depth Cost Finicky Score
CHL / EqProp ✅ ✅ ⚠️ ⚠️ ✅ Med High 6/9
FF / SCFF ⚠️ ⚠️ ⚠️ ✅ ✅ Low Med 5.5/9
Predictive Coding ✅ ⚠️ ✅ ✅ ✅ Med Med 7.5/9
Target Propagation ❌ ⚠️ ❌ ⚠️ ✅ High High 2/9
R-STDP / 3-factor ✅ ✅ ✅ ✅ ⚠️ Low Med 7.5/9
Node/Weight Perturb ✅ ✅ ✅ ✅ ⚠️ High Med 6.5/9
InfoMax ✅ ✅ ❌ ✅ ✅ Low Low 5/9 (aux only)
Genetic / EA ✅ ✅ ✅ ❌ ✅ High Med 4/9
Feedback Alignment ❌ ❌ ⚠️ ✅ ⚠️ Low Low 2/9
r-e-prop ✅ ✅ ✅ ✅ ⚠️ Low Med 7.5/9

TOP 3 RECOMMENDATIONS

Rank 1: R-STDP / Three-Factor Learning (Shallow + Output Layer)

Why #1: Every criterion met cleanly. No derivatives anywhere. Binary spikes are native. Reward is the direct learning signal. Per-tick online updates. Cheap (O(n_synapses) per tick).

The constraint: poor depth scaling. Solution: limit depth or use shallow BNN with this method.

Architecture sketch (2-layer BNN):

Input (690 bits)
     │
     ▼ W1 ∈ {-1,+1}^{690×128}
Hidden Layer (128 binary neurons)
     │   ↑ InfoMax auxiliary update (prevents collapse)
     ▼ W2 ∈ {-1,+1}^{128×64}
Output Layer (64 binary neurons → [predX, predY] via dot-product readout)
     │
     ▼
Circular wave feedback → miss distance → reward r(t)
     │
     └──→ Eligibility trace e_ij = e_ij · τ_decay + STDP(pre_j, post_i)
          Weight update: ΔW_ij = η · e_ij · r(t)   [for all layers uniformly]

Signal flow for depth: Use a reward-gated reverse signal — after reward arrives, compute output layer error signal (not gradient, just "output was wrong by X direction") and multiply by a random fixed matrix B to project to hidden layer. This gives hidden layer a noisy credit signal. This is the "feedback alignment without gradients" version where the output error is a reward signal (binary direction: aim left or right), not a loss derivative.

Hypers: τ_eligibility ≈ 10 ticks (covers bullet travel time), η ≈ 0.01, STDP window ≈ 3 ticks.


Rank 2: Predictive Coding with Reward-Gated Output

Why #2: Best theoretical fit to our temporal 10-frame structure. Self-supervised hidden layers (no labels needed), reward modulates only the output layer. Proven online on edge robots (IROS 2025).

Architecture sketch (3-layer PCN with temporal hierarchy):

Frame buffer: [f0..f9] (10 temporal frames × 69 fields = 690 bits)

Layer 3 (top): trajectory-level representation, 64 neurons
  ↕ predicts/corrects ↕     W3 = Hebbian update: e3 · h2^T
Layer 2: motion dynamics, 128 neurons
  ↕ predicts/corrects ↕     W2 = e2 · h1^T
Layer 1: per-frame features, 256 neurons
  ↕ corrects ↕              W1 = e1 · x^T
Input: 690-bit frame vector

Signal flow:
  Forward: h_l = σ(W_l · h_{l-1})         [inference]
  Backward: μ_l = W_{l+1}^T · h_{l+1}    [top-down prediction, W^T of SAME weights]
  Error: e_l = h_l − μ_l                  [local prediction error]
  Weight: ΔW_l = η_rep · e_l · h_{l-1}^T [Hebbian on error × input]

Reward integration:
  Output layer e_out += η_reward · r(t) · sign(h_out − target_direction)
  where target_direction is estimated from miss distance (left/right of predicted pos)

Key: the weight update ΔW = e · h^T is Hebbian, not gradient descent. The inference phase (adjusting activations to minimize prediction error) uses signal flow, not backprop. Binarize outputs only; keep latent variables as integers 0-255 (8-bit) for inference, threshold for communication.

Hypers: inference iterations ≈ 3-5 per tick (PCN-TA reduces this), η_rep ≈ 0.001, η_reward ≈ 0.05.


Rank 3: Self-Contrastive FF (SCFF) for Deep Representations + R-STDP Output

Why #3: Best combination for a deeper (3+ layer) network when richer representations are needed. SCFF trains hidden layers purely from self-supervised data (no labels, no reward). R-STDP on the output layer uses reward directly. Two independent, specialized learning mechanisms.

Architecture sketch:

Input: 690-bit × 2 (concatenated for SCFF contrast)

Layer 1: 512 binary neurons
  SCFF update: Hebbian-approx goodness rule
  positive pair: [x_k, x_k] → increase goodness
  negative pair: [x_k, x_n] → decrease goodness
  ΔW ∝ (target_goodness − actual_goodness) · pre_activity   [no STE — approximate]

Layer 2: 256 binary neurons
  SCFF update: same as Layer 1

Layer 3: 128 binary neurons
  SCFF update: same

Output: 32 binary neurons → [predX, predY] decoding
  R-STDP update: ΔW ∝ e_ij · r(t)   [three-factor, reward = −miss_distance]
  Eligibility trace: STDP on output spikes × layer-3 spikes

Why this works: Hidden layers learn to preserve temporal motion patterns from the input buffer (self-supervised: same-frame vs different-frame contrast captures motion coherence). Output layer learns which patterns correlate with correct predictions (reward). The two learning rules are orthogonal and can run simultaneously without interference.

The SCFF-binary approximation: Replace ∂goodness/∂W with Hebbian on active neurons when goodness exceeds/falls below threshold θ. Loses convergence guarantees but works empirically for binary feature learning. θ ≈ 0.5 × expected_activation_rate.

Hypers: θ ≈ 0.4, τ_eligibility ≈ 10, η_scff ≈ 0.001 (slow), η_rstdp ≈ 0.01 (fast).


Final Notes and Red Flags

Things That Look Gradient-Free But Aren't

  1. Target Propagation: Output target still needs gradient descent. ❌
  2. Feedback Alignment: Random backward weights but still computes ∂L/∂y. ❌
  3. Standard Forward-Forward: Goodness gradient is a real derivative through relu. Needs approximation for binary. ⚠️
  4. E-prop (standard version): Eligibility trace approximates ∂h/∂W — it is a gradient of activations, which satisfies "backprop of signals" but is worth flagging. ✅ (barely)

The Depth-vs-Purity Trade-off

There is a fundamental tension: pure local learning rules (InfoMax, R-STDP) have no depth credit assignment. Any method that propagates credit through depth (CHL, EqProp, PCN, e-prop) uses some form of signal backpropagation. The question is whether those signals are gradients of a loss (forbidden) or prediction errors / activity differences (allowed as signals). All three recommendations above fall in the "allowed" category by that interpretation.

Practical Starting Point

Start with Rank 1 (R-STDP, 2 layers). It is the fastest to implement, cheapest to run, and cleanest theoretically. Add InfoMax as a free auxiliary update on the hidden layer to prevent dead neurons. Measure performance. If representation quality is bottlenecking predictions, move to Rank 2 (PCN temporal hierarchy). Only move to Rank 3 (SCFF+R-STDP) if deeper representations are demonstrably needed.


Sources