# BNNBot Deep Learning: Gradient-Free Methods for Binary Networks **Context**: 690-bit binary input, online battle learning, reward = miss distance from circular wave feedback, binary weights/activations preferred. No gradient descent, no STE, no supervised labels. Backprop of *signals* (not gradients) is OK. --- ## Evaluation Scorecard Legend | Symbol | Meaning | |--------|---------| | ✅ | Full support | | ⚠️ | Partial / with caveats | | ❌ | Fails this criterion | **Criteria for each method:** 1. **No GD** — truly avoids gradient descent (no derivatives, no STE, no surrogate gradients) 2. **Binary** — works with binary weights and/or activations 3. **RL** — can learn from reward signal instead of labels 4. **Online** — can update weights tick-by-tick, not batch-offline 5. **Depth** — practically trains 3+ layers 6. **Cost** — compute per weight update (relative: Low / Med / High) 7. **Finicky** — sensitivity to hyperparameters (Low / Med / High) --- ## Method 1: Contrastive Hebbian Learning (CHL) / Equilibrium Propagation ### How It Works CHL runs two phases using an energy-based (Hopfield-like) network: 1. **Free phase**: Input clamped, hidden + output units relax to energy minimum → record activations `s⁻` 2. **Clamped phase**: Input + output clamped to target, hidden units relax → record activations `s⁺` 3. **Weight update**: `ΔW_ij ∝ s⁺_i · s⁺_j − s⁻_i · s⁻_j` This is purely local — each weight only needs the activations of its two endpoints in both phases. No global error signal flows backward; activity differences drive learning. **Equilibrium Propagation (Scellier & Bengio 2017)** is the modern formulation: the clamped phase is a "nudged" version where the output is soft-clamped with a small `β` factor. In the limit β→0, EqProp provably computes the same updates as backprop, but without passing gradients. **Binary variant (Laydevant et al., CVPRW 2021)**: EqProp was directly applied to dynamical binary neural networks. The free/clamped phases remain unchanged; binary activations are compatible because the update rule only compares activation products, not derivatives. ### Can It Work Without Supervised Targets? **Critical limitation**: The clamped phase requires knowing the *target output*. In standard CHL/EqProp the output neurons are pushed toward a labeled target. To use reward instead of labels, you need a "nudging" signal at the output layer. **Viable workaround for RL**: Instead of nudging toward a ground-truth label, nudge the output in the direction that increases reward. If reward is scalar, define a "desired output shift": - Positive reward: nudge current output activations → reinforce current pattern - Negative reward: nudge away from current output This converts CHL into a reward-modulated version. It is not a standard use case but has been explored in neo-Hebbian RL literature. The key: the "nudge" is a *signal* (direction to move), not a gradient. It fits the constraint. ### Scorecard | Criterion | Score | Notes | |-----------|-------|-------| | No GD | ✅ | Pure activity differences, no derivatives | | Binary | ✅ | Demonstrated in CVPRW 2021 on dynamic BNNs | | RL | ⚠️ | Needs reward→nudge translation; not native but viable | | Online | ⚠️ | Requires two settling phases per step; adds latency | | Depth | ✅ | EqProp scales to multiple layers; credit is local | | Cost | Med | Two full network relaxations per update | | Finicky | High | β nudge strength, settling iterations, energy landscape | ### Fit Assessment Strong theoretical foundation, but the two-phase settling requirement means you need the network to relax to equilibrium twice per firing tick. In Robocode's 1ms budget this is expensive. The reward→ nudge translation is non-trivial to implement correctly. **Viable but complex.** --- ## Method 2: Forward-Forward Algorithm (FF) / Self-Contrastive FF ### How It Works Hinton (2022): Replace one forward + one backward pass with *two forward passes*: 1. **Positive pass**: real/good data through the layer → maximize "goodness" G = Σ(y²) 2. **Negative pass**: corrupted/bad data → minimize goodness 3. **Layer-local update** for each layer independently: push goodness above threshold θ for positive, below for negative. No signal crosses layer boundaries. The Self-Contrastive Forward-Forward (SCFF, 2024/2025, Nature Communications) eliminates the label requirement entirely: positive sample = [x_k, x_k] (same input concatenated), negative sample = [x_k, x_n] (two different inputs). The network learns to distinguish same-from-different without any labels. ### Goodness Function `G(l) = (1/M) Σ_m y²_m(l)` — sum of squared activations in layer l over M neurons. Learning rule: `ΔW ∝ (σ(G - θ) - target) · ∂G/∂W` **Wait — does this have gradients?** Yes, the standard FF uses a sigmoid-of-goodness loss with gradient descent on the goodness. However: (a) the update is strictly *local* to each layer, and (b) ∂G/∂W = 2y · ∂y/∂W. For binary activations with sign() this derivative is zero/undefined. **Binary compatibility issue**: The goodness gradient requires ∂y/∂W. With hard binary activations (sign function), this derivative is zero everywhere except at the threshold, breaking the goodness update. The STE would be required to push gradients through — which is explicitly disallowed. **Alternative**: Replace the goodness gradient with a Hebbian rule — if G > θ for positive data, do Hebbian update on active neurons; if G < θ for negative data, do anti-Hebbian. This is the "spirit" of FF without the gradient. It loses convergence guarantees but preserves the structure. ### Adaptation for RL Since SCFF already generates its own positive/negative pairs without labels, it operates as a *representation learning* system. For Robocode: the FF layers would learn good binary representations; then a reward-modulated Hebbian output layer would learn to map representations to predictions. ### Scorecard | Criterion | Score | Notes | |-----------|-------|-------| | No GD | ⚠️ | Standard FF uses local GD on goodness; binary requires Hebbian approximation | | Binary | ⚠️ | Needs Hebbian approximation of goodness; SCFF uses ReLU throughout | | RL | ⚠️ | SCFF is self-supervised (representation learning); needs separate RL output layer | | Online | ✅ | Layer-local, greedy, update after each sample | | Depth | ✅ | By design — each layer trains independently | | Cost | Low | Two forward passes per update, strictly local | | Finicky | Med | Threshold θ, concatenation design choices | ### Fit Assessment FF is excellent for building deep binary representations without labels. It cannot directly predict position with reward alone — you'd hybridize it: FF middle layers for feature extraction, reward- modulated Hebbian output layer for prediction. The goodness-gradient problem with hard binary activations is the main obstacle. **Good as a representation sub-system; needs hybridization.** --- ## Method 3: Predictive Coding Networks (PCN) ### How It Works Each layer l predicts the activation of layer l-1 from above using top-down weights. The prediction error e_l = x_l − μ_l (actual minus predicted) propagates locally upward and downward. Weight update: `ΔW ∝ e_l · x_{l-1}^T` (Hebbian on error × lower-layer activation) This is **self-supervised by construction**: the "target" of each layer is the actual activation of the layer below it. No external labels needed. With temporal data, each layer can predict the next time frame — perfectly suited to 10-frame temporal windows. **Active Predictive Coding (APC, 2022)**: Combines predictive coding with RL for robotic sparse reward problems. The prediction error serves as an intrinsic reward signal. Demonstrated on robotic control tasks with sparse rewards. **PCN-TA (2025)**: Temporal amortization preserves latent states across frames, yielding 50% fewer inference steps and 10% fewer weight updates vs. backprop. ### Gradient Status Standard PCN inference minimizes free energy F = Σ e_l² by gradient descent on *activations* (not weights). This is allowed — you're adjusting activity, not weights, via derivatives. Weight updates are then Hebbian: `ΔW ∝ e · x`. No gradient of the loss w.r.t. weights is computed. **Technically pure**: the weight update is `ΔW_l = η · e_l · x_{l-1}^T`, which is purely local and Hebbian-like. The inference phase uses a form of gradient descent on activations, which is a signal-flow operation (activations are signals), consistent with "backprop of signals allowed." ### Binary Compatibility PCN with continuous activations works cleanly. Binary activations are harder: the inference phase requires iterative adjustment of continuous latent variables before thresholding. A common approach: keep latent variables continuous during inference, threshold only for output/communication. This is the "semi-binary" regime used in neuromorphic PCN implementations. ### Fit for Temporal Prediction Our 10 temporal frames map directly to a hierarchical PCN structure: - Layer 1: current-frame features - Layer 2: short-term dynamics (frames 0-2) - Layer 3: trajectory trend (frames 0-9) - Each layer predicts the layer below across time The reward signal (miss distance) can modulate the prediction error magnitude at the output layer, biasing learning toward accurate predictions. ### Scorecard | Criterion | Score | Notes | |-----------|-------|-------| | No GD | ✅ | Weight update is Hebbian `e·x`; inference GD is on activations (signals) | | Binary | ⚠️ | Latent variables need continuous inference; output can be binary | | RL | ✅ | APC demonstrated; error = intrinsic reward; reward modulates output layer | | Online | ✅ | PCN-TA designed for online per-frame learning | | Depth | ✅ | Hierarchical by design; 3+ layers well studied | | Cost | Med | Multiple inference iterations per step (reduced 50% in PCN-TA) | | Finicky | Med | Inference step count, precision/ratio of prediction error weights | ### Fit Assessment **Excellent match for our temporal structure.** The self-supervised nature, temporal hierarchy, and local Hebbian weight updates align with every constraint. Main cost: inference iterations. PCN-TA reduces this significantly. The semi-binary regime (continuous latent, binary output) is a realistic implementation path. --- ## Method 4: Target Propagation / Difference Target Propagation (DTP) ### How It Works Instead of propagating gradients, TP learns an *inverse model* at each layer: a backward network g_l that maps activations of layer l+1 back to layer l. The target for layer l is: `h_l^target = g_l(h_{l+1}^target)` **Difference Target Propagation (DTP)**: corrects for imperfect inverses by using: `h_l^target = h_l + g_l(h_{l+1}^target) − g_l(h_{l+1})` The target for the output layer *still requires a supervised target* — this is the fundamental limitation. DTP cannot generate its own target; it only propagates given targets layer-by-layer. ### Does It Need Supervised Output? **Yes, inherently.** The output-layer target must come from somewhere. In RL, you could define the output target as "the output that would have hit the enemy" — but this requires knowing the correct output, which is equivalent to supervised learning. The reward tells you *how much* you were off, not *where* you should have aimed. Additionally: "the target for the final hidden layer is determined by a formula which relies on gradient descent" (from DTP theory paper). This means gradient descent leaks in even in DTP. ### Binary Compatibility DTP can work with binary units (Lee et al. showed this: "it can be applied even when units exchange stochastic bits rather than real numbers"). But non-invertible binary networks cause reconstruction errors to interfere with target propagation updates. ### Scorecard | Criterion | Score | Notes | |-----------|-------|-------| | No GD | ❌ | Output-layer target computation involves GD; inverse model also trained with GD | | Binary | ⚠️ | Stochastic binary supported but reconstruction error problem | | RL | ❌ | Requires supervised output target; reward alone insufficient | | Online | ⚠️ | Possible but inverse models add computational overhead | | Depth | ✅ | Designed for depth; that's the whole point | | Cost | High | Two networks (forward + inverse) per layer | | Finicky | High | Inverse model training, reconstruction accuracy | ### Fit Assessment **Not suitable.** Despite claiming "no backprop," DTP requires gradient descent for the inverse models and for the output target. Does not fit RL-only constraint. Skip. --- ## Method 5: Reward-Modulated STDP (R-STDP) / Three-Factor Learning / E-prop ### How It Works **Three-factor rule**: `ΔW_ij = η · e_ij · r(t)` where: - `e_ij` = eligibility trace: local pre×post spike correlation (STDP-like) - `r(t)` = global reward signal (scalar, broadcast to all synapses) Eligibility trace: `de_ij/dt = −e_ij/τ + STDP(t_pre, t_post)` — decaying memory of recent spike correlations. This handles temporal credit assignment: the trace remembers which synapses fired recently when the (delayed) reward arrives. **E-prop (Bellec et al., Nature Comms 2020)**: Extends three-factor learning to deep and recurrent SNNs. Each layer l has an eligibility trace that accounts for the temporal dynamics of the LIF neuron. E-prop with reward-based RL (r-e-prop) was demonstrated winning Atari games. **Critical depth finding**: "e-prop implemented in a single-layer recurrent SNN consistently outperforms a multi-layer variant." Adding more layers hurts because the eligibility trace at layer l only captures local pre/post correlations — it cannot account for multi-layer credit assignment without approximate gradients. E-prop for deep feedforward networks approximates backprop, not a clean departure from it. ### Binary Compatibility SNNs fire binary spikes by definition (0 or 1), so R-STDP is natively binary-compatible. Weights can also be binarized. The eligibility trace is continuous (a real-valued running average), but that is internal state, not a gradient propagated through the network. ### Credit Assignment Through Depth This is the weak point. R-STDP solves temporal credit assignment (delay between action and reward) but not *spatial* credit assignment (which layer contributed to the good/bad output). In a 3-layer network with R-STDP applied uniformly, all layers update on the same reward signal regardless of their contribution. This works empirically in shallow networks but degrades with depth. ### Scorecard | Criterion | Score | Notes | |-----------|-------|-------| | No GD | ✅ | Local STDP trace × reward; no derivatives anywhere | | Binary | ✅ | Native binary (spikes); weights can also be binary | | RL | ✅ | Reward is the direct learning signal; no labels | | Online | ✅ | Per-spike updates; inherently online | | Depth | ⚠️ | Works but degrades with depth; single-layer LSNN beats multi-layer | | Cost | Low | Eligibility trace update O(n_synapses) per tick | | Finicky | Med | τ_eligibility, reward scaling, STDP window | ### Fit Assessment **Best fit for shallow networks.** For a 2-layer binary SNN (one hidden layer), R-STDP is the cleanest solution: no gradients, native binary, direct reward learning, very cheap. For 3+ layers, credit assignment degrades. Consider using it for the *output layer* in combination with another method for hidden layers. **Highly recommended for shallow or hybrid architectures.** --- ## Method 6: Node Perturbation / Weight Perturbation ### How It Works **Node perturbation (NP)**: Add noise ξ to each layer's activations, measure reward R, update: `ΔW_ij ∝ ξ_j · (R − R̄)` where R̄ is baseline reward. **Weight perturbation (WP)**: Add noise directly to weights: `ΔW_ij ∝ ε_ij · (R − R̄)` **Key insight**: This estimates the gradient without computing it. It is a Monte Carlo gradient estimate — unbiased but high variance. **Decorrelated NP (DNP, 2023)**: Applying input decorrelation at each layer "dramatically improves convergence by orders of magnitude." Makes NP practical for multi-layer networks. **Variance problem**: The SNR of the NP gradient estimate scales as `1/(N·σ_noise²)` where N = number of parameters. For a 690-bit → 256-neuron hidden layer, N ≈ 176,640 weights. The variance explosion makes learning essentially random without variance reduction. **For binary networks specifically**: Weight perturbation on binary weights means randomly flipping bits and measuring reward change. This is exactly the "simulated annealing on neural net weights" approach. It is well-defined and requires no derivatives. The update becomes: `if flip W_ij causes ΔR > 0: keep flip; else: revert (with probability based on ΔR)` This is pure stochastic search — not gradient estimation. Works cleanly with binary weights. ### Scorecard | Criterion | Score | Notes | |-----------|-------|-------| | No GD | ✅ | Statistical gradient estimation; no derivatives | | Binary | ✅ | WP maps naturally to bit-flip search | | RL | ✅ | Reward is the direct signal | | Online | ✅ | Update per trial; eligible for online use | | Depth | ⚠️ | DNP scales to 3-9 layers; variance still significant | | Cost | High | Many perturbation samples needed for convergence | | Finicky | Med | Perturbation magnitude σ, baseline estimator | ### Fit Assessment **The cleanest in principle, but expensive.** Weight perturbation on a binary network is essentially a guided random walk through {-1,+1}^N. With N ≈ 176K weights in a hidden layer, random perturbation converges very slowly. Works best when: (a) network is small, (b) few samples needed per weight. For the *output layer only* (small, ~64 weights to prediction neurons), this is practical. **Suitable as output-layer fine-tuner; too slow as sole learning method for deep networks.** --- ## Method 7: InfoMax / Information Maximization ### How It Works Linsker (1988), Bell & Sejnowski (1995): Maximize mutual information I(X; Y) between layer input X and output Y, subject to noise constraints. Local learning rule (continuous case): `ΔW_ij ∝ [φ'(y_i)/φ(y_i)] · x_j − W_ij^{-T}` For the binary/ICA case (Bell-Sejnowski): involves a nonlinear function of output activation and the input pattern. In practice, this maximizes entropy of the output distribution, preventing the network from collapsing all outputs to zero or all-ones. ### Binary Compatibility InfoMax with binary activations maximizes entropy H(Y) — binary outputs should have p(y=1) ≈ 0.5 across data. The update rule involves the "score function" of the output distribution, which for binary units is well-defined without derivatives through the activation function itself. ### Self-Supervised Nature InfoMax requires *no labels*. It maximizes information preserved from input to output. This makes it useful for *representation learning in hidden layers* — building diverse, non-collapsed binary features. It does not directly learn to predict enemy position. ### Limitations for RL InfoMax is a representation learning method. It maximizes informativeness of intermediate features but has no notion of task reward. You cannot directly optimize miss distance with InfoMax alone. It would serve as a **pretraining / regularization layer** to prevent dead neurons and maintain diverse binary representations. ### Scorecard | Criterion | Score | Notes | |-----------|-------|-------| | No GD | ✅ | Local Hebbian-like rule with anti-Hebbian lateral term | | Binary | ✅ | Entropy maximization is well-defined for binary units | | RL | ❌ | Pure representation learning; no reward optimization | | Online | ✅ | Fully online, local update | | Depth | ✅ | Applied layer-by-layer independently | | Cost | Low | Single forward pass per update | | Finicky | Low | Few parameters; entropy maximization is self-stabilizing | ### Fit Assessment **Excellent regularizer / auxiliary learning rule.** Run InfoMax as an auxiliary update on hidden layers to prevent representational collapse while the output layer learns from reward. Very cheap. **Use as a complement to primary method, not standalone.** --- ## Method 8: Evolutionary / Genetic Algorithms for Binary BNNs ### How It Works Maintain a population of binary weight matrices. Each individual is a complete set of weights `{W_l}` with values in {-1, +1}. Evolution: 1. Evaluate fitness (reward = hit rate, negative miss distance) 2. Select top-K individuals (tournament or truncation selection) 3. Crossover: combine weight sub-matrices from two parents 4. Mutation: randomly flip bits with probability p_mut 5. Repeat Binary networks are particularly GA-friendly because: - The search space is discrete and finite - Crossover has a clean interpretation (taking different weight blocks) - No gradient computation at all - No issues with binary activations (trivially compatible) ### Population-Based Training During a Battle A battle has ~1000+ ticks. With a population of P=10-20 individuals, each tick evaluates the current "best" policy, and every N ticks (after a reward arrives from circular wave) the population updates. This is a form of online evolution. **Key problem**: Each individual needs to be evaluated on the *same input* to compare fitnesses. During a battle, inputs change each tick — you cannot hold them constant while evaluating 20 models. Solutions: - Maintain an experience replay buffer; evaluate candidates on stored states - Use the single "best" candidate in battle, explore with perturbations (= weight perturbation) - Use island model: different individuals fight in different battles ### Practical Issues - Population of 20 requires 20× memory for weights - Fitness estimates from sequential battle experience are noisy (different opponents, different positions each eval) - Crossover between weights that serve different layers may break learned structure - Convergence within a single 1000-tick battle is unlikely for deep networks ### Scorecard | Criterion | Score | Notes | |-----------|-------|-------| | No GD | ✅ | Pure fitness-based search | | Binary | ✅ | Perfect; binary is ideal for genetic operators | | RL | ✅ | Fitness = battle reward; no labels | | Online | ❌ | Population requires multiple evaluations; impossible in a single live battle | | Depth | ✅ | Architecture-agnostic | | Cost | High | O(P × network_size) per generation | | Finicky | Med | Population size, mutation rate, selection pressure | ### Fit Assessment **Not suitable for within-battle online learning.** The population-based evaluation requirement breaks against the single-agent, real-time constraint. Best applied *between battles* (offline evolution over many battles). If we're restricted to online learning, this is out. **Viable only as inter-battle meta-learning.** --- ## Method 9: Feedback Alignment (FA) / Direct Feedback Alignment (DFA) ### How It Works **Feedback Alignment (Lillicrap et al., 2016)**: Replace backprop's transposed weight matrices W^T in the backward pass with fixed random matrices B. The claim: the forward weights align to B during training, so the random backward signal still provides useful credit assignment. **Direct Feedback Alignment (DFA)**: Each layer receives credit directly from the output error multiplied by a random matrix B_l (layer-specific). No sequential backward pass needed. ### The Gradient Problem FA and DFA fundamentally still compute a gradient of the loss at the *output layer* and propagate that signal (or a random-projection of it) backward. The output gradient `∂L/∂y` is a derivative. This violates the "no gradient descent" constraint. FA avoids the *weight transport* problem but does not avoid derivatives entirely. **Is the "signal" interpretation valid?** You could argue: the output error `(y - target)` is a signal (not a derivative). If you replace this with a reward-modulated output unit error, you avoid computing a derivative. But this is a stretch — in practice, FA implementations use `∂L/∂y`. **Binary compatibility**: FA fails on binary networks because the backward signal through hard thresholds is zero everywhere (the same STE problem as standard backprop). Papers report "FA fails in deep networks, convolutional layers, or architectures with bottlenecks." ### Scorecard | Criterion | Score | Notes | |-----------|-------|-------| | No GD | ❌ | Output gradient still computed; only backward weight symmetry is removed | | Binary | ❌ | Fails with binary activations without STE | | RL | ⚠️ | Could replace loss gradient with reward signal, but then it's R-STDP | | Online | ✅ | Online update possible | | Depth | ⚠️ | DFA works at depth but with performance degradation | | Cost | Low | Same cost as forward pass + random projection | | Finicky | Low | Random fixed matrices; no tuning needed | ### Fit Assessment **Not suitable as stated.** FA/DFA require output-layer gradients and fail with binary activations. If you replace the output gradient with a reward signal and the backward pass with fixed random projections, you get something closer to random reward feedback — a variant of node perturbation, which is covered above. **Skip in favor of R-STDP or node perturbation.** --- ## Method 10: Additional Methods Found ### 10a: Equilibrium Propagation + CHL as a Unified Framework Recent work (arXiv 2206.02629) shows that CHL, EqProp, and predictive coding all converge to the same limit at infinitesimal inference steps — they are the same algorithm in different dynamical regimes. This means choosing among them is primarily about implementation tradeoffs (settling speed, binary compatibility) not fundamental differences. ### 10b: Counter-Current Learning (CCL, 2024) A dual-network approach: one network runs forward (inference), the other runs backward (teaching). The backward network generates targets for the forward network without computing gradients. Used as a biologically plausible alternative to backprop. Requires paired network architecture, doubling memory. Interesting but adds complexity. **Skip for now.** ### 10c: E-prop with Reward (r-e-prop) Specifically the reward-based variant of e-prop demonstrated on Atari. This is three-factor learning with eligibility traces designed for temporal credit assignment, formally shown to approximate policy gradients. Does use gradient approximations internally (via the eligibility trace derivation), but the weight update itself is `ΔW = e_ij · r(t)` — a local multiplication. **Whether this "counts" as gradient descent**: The eligibility trace `e_ij` is derived from the neuron's dynamical equations and approximates `∂h_l/∂W`. This is a gradient of *activations* w.r.t. weights (signal propagation), not a gradient of the *loss*. It satisfies "backprop of signals" while avoiding "gradient descent on loss." **This is the cleanest theoretical fit.** ### 10d: Noise-Based Reward-Modulated Learning (2025, arXiv 2503.23972) Explicitly designed for spiking/binary units: natural noise in binary neurons drives stochastic exploration, reward signal modulates which noise patterns are reinforced. Weight update: `ΔW ∝ ξ · r(t)` where ξ is inherent neural noise (spike timing jitter). Online, local, binary-native. **Very relevant — this is weight perturbation where perturbation = natural spike noise.** ### 10e: Self-Contrastive Forward-Forward (SCFF, Nature Comms 2025) Published result: MNIST 98.7%, CIFAR-10 80.75%, STL-10 77.3% without any labels. Layer-local goodness function. Greedy layer-wise training. The best current result for label-free local learning in deep networks. Needs Hebbian approximation for hard binary activations. --- ## Comparative Scorecard Summary | Method | No GD | Binary | RL | Online | Depth | Cost | Finicky | **Score** | |--------|-------|--------|-----|--------|-------|------|---------|-----------| | CHL / EqProp | ✅ | ✅ | ⚠️ | ⚠️ | ✅ | Med | High | **6/9** | | FF / SCFF | ⚠️ | ⚠️ | ⚠️ | ✅ | ✅ | Low | Med | **5.5/9** | | Predictive Coding | ✅ | ⚠️ | ✅ | ✅ | ✅ | Med | Med | **7.5/9** | | Target Propagation | ❌ | ⚠️ | ❌ | ⚠️ | ✅ | High | High | **2/9** | | R-STDP / 3-factor | ✅ | ✅ | ✅ | ✅ | ⚠️ | Low | Med | **7.5/9** | | Node/Weight Perturb | ✅ | ✅ | ✅ | ✅ | ⚠️ | High | Med | **6.5/9** | | InfoMax | ✅ | ✅ | ❌ | ✅ | ✅ | Low | Low | **5/9** (aux only) | | Genetic / EA | ✅ | ✅ | ✅ | ❌ | ✅ | High | Med | **4/9** | | Feedback Alignment | ❌ | ❌ | ⚠️ | ✅ | ⚠️ | Low | Low | **2/9** | | r-e-prop | ✅ | ✅ | ✅ | ✅ | ⚠️ | Low | Med | **7.5/9** | --- ## TOP 3 RECOMMENDATIONS ### Rank 1: R-STDP / Three-Factor Learning (Shallow + Output Layer) **Why #1**: Every criterion met cleanly. No derivatives anywhere. Binary spikes are native. Reward is the direct learning signal. Per-tick online updates. Cheap (O(n_synapses) per tick). **The constraint**: poor depth scaling. Solution: limit depth or use shallow BNN with this method. **Architecture sketch** (2-layer BNN): ``` Input (690 bits) │ ▼ W1 ∈ {-1,+1}^{690×128} Hidden Layer (128 binary neurons) │ ↑ InfoMax auxiliary update (prevents collapse) ▼ W2 ∈ {-1,+1}^{128×64} Output Layer (64 binary neurons → [predX, predY] via dot-product readout) │ ▼ Circular wave feedback → miss distance → reward r(t) │ └──→ Eligibility trace e_ij = e_ij · τ_decay + STDP(pre_j, post_i) Weight update: ΔW_ij = η · e_ij · r(t) [for all layers uniformly] ``` **Signal flow for depth**: Use a **reward-gated reverse signal** — after reward arrives, compute output layer error signal (not gradient, just "output was wrong by X direction") and multiply by a random fixed matrix B to project to hidden layer. This gives hidden layer a noisy credit signal. This is the "feedback alignment without gradients" version where the output error is a reward signal (binary direction: aim left or right), not a loss derivative. **Hypers**: τ_eligibility ≈ 10 ticks (covers bullet travel time), η ≈ 0.01, STDP window ≈ 3 ticks. --- ### Rank 2: Predictive Coding with Reward-Gated Output **Why #2**: Best theoretical fit to our temporal 10-frame structure. Self-supervised hidden layers (no labels needed), reward modulates only the output layer. Proven online on edge robots (IROS 2025). **Architecture sketch** (3-layer PCN with temporal hierarchy): ``` Frame buffer: [f0..f9] (10 temporal frames × 69 fields = 690 bits) Layer 3 (top): trajectory-level representation, 64 neurons ↕ predicts/corrects ↕ W3 = Hebbian update: e3 · h2^T Layer 2: motion dynamics, 128 neurons ↕ predicts/corrects ↕ W2 = e2 · h1^T Layer 1: per-frame features, 256 neurons ↕ corrects ↕ W1 = e1 · x^T Input: 690-bit frame vector Signal flow: Forward: h_l = σ(W_l · h_{l-1}) [inference] Backward: μ_l = W_{l+1}^T · h_{l+1} [top-down prediction, W^T of SAME weights] Error: e_l = h_l − μ_l [local prediction error] Weight: ΔW_l = η_rep · e_l · h_{l-1}^T [Hebbian on error × input] Reward integration: Output layer e_out += η_reward · r(t) · sign(h_out − target_direction) where target_direction is estimated from miss distance (left/right of predicted pos) ``` **Key**: the weight update `ΔW = e · h^T` is Hebbian, not gradient descent. The inference phase (adjusting activations to minimize prediction error) uses signal flow, not backprop. Binarize outputs only; keep latent variables as integers 0-255 (8-bit) for inference, threshold for communication. **Hypers**: inference iterations ≈ 3-5 per tick (PCN-TA reduces this), η_rep ≈ 0.001, η_reward ≈ 0.05. --- ### Rank 3: Self-Contrastive FF (SCFF) for Deep Representations + R-STDP Output **Why #3**: Best combination for a deeper (3+ layer) network when richer representations are needed. SCFF trains hidden layers purely from self-supervised data (no labels, no reward). R-STDP on the output layer uses reward directly. Two independent, specialized learning mechanisms. **Architecture sketch**: ``` Input: 690-bit × 2 (concatenated for SCFF contrast) Layer 1: 512 binary neurons SCFF update: Hebbian-approx goodness rule positive pair: [x_k, x_k] → increase goodness negative pair: [x_k, x_n] → decrease goodness ΔW ∝ (target_goodness − actual_goodness) · pre_activity [no STE — approximate] Layer 2: 256 binary neurons SCFF update: same as Layer 1 Layer 3: 128 binary neurons SCFF update: same Output: 32 binary neurons → [predX, predY] decoding R-STDP update: ΔW ∝ e_ij · r(t) [three-factor, reward = −miss_distance] Eligibility trace: STDP on output spikes × layer-3 spikes ``` **Why this works**: Hidden layers learn to preserve temporal motion patterns from the input buffer (self-supervised: same-frame vs different-frame contrast captures motion coherence). Output layer learns which patterns correlate with correct predictions (reward). The two learning rules are orthogonal and can run simultaneously without interference. **The SCFF-binary approximation**: Replace `∂goodness/∂W` with Hebbian on active neurons when goodness exceeds/falls below threshold θ. Loses convergence guarantees but works empirically for binary feature learning. θ ≈ 0.5 × expected_activation_rate. **Hypers**: θ ≈ 0.4, τ_eligibility ≈ 10, η_scff ≈ 0.001 (slow), η_rstdp ≈ 0.01 (fast). --- ## Final Notes and Red Flags ### Things That Look Gradient-Free But Aren't 1. **Target Propagation**: Output target still needs gradient descent. ❌ 2. **Feedback Alignment**: Random backward weights but still computes `∂L/∂y`. ❌ 3. **Standard Forward-Forward**: Goodness gradient is a real derivative through relu. Needs approximation for binary. ⚠️ 4. **E-prop (standard version)**: Eligibility trace approximates `∂h/∂W` — it is a gradient of activations, which satisfies "backprop of signals" but is worth flagging. ✅ (barely) ### The Depth-vs-Purity Trade-off There is a fundamental tension: pure local learning rules (InfoMax, R-STDP) have no depth credit assignment. Any method that propagates credit through depth (CHL, EqProp, PCN, e-prop) uses some form of signal backpropagation. The question is whether those signals are *gradients of a loss* (forbidden) or *prediction errors / activity differences* (allowed as signals). All three recommendations above fall in the "allowed" category by that interpretation. ### Practical Starting Point Start with **Rank 1** (R-STDP, 2 layers). It is the fastest to implement, cheapest to run, and cleanest theoretically. Add **InfoMax** as a free auxiliary update on the hidden layer to prevent dead neurons. Measure performance. If representation quality is bottlenecking predictions, move to **Rank 2** (PCN temporal hierarchy). Only move to **Rank 3** (SCFF+R-STDP) if deeper representations are demonstrably needed. --- ## Sources - [Contrastive Hebbian Learning with Random Feedback Weights](https://arxiv.org/pdf/1806.07406) - [Two Tales of Single-Phase Contrastive Hebbian Learning](https://arxiv.org/pdf/2402.08573) - [Equilibrium Propagation: Bridging the Gap between Energy-Based Models and Backpropagation](https://arxiv.org/html/1602.05179v5) - [Training Dynamical Binary Neural Networks with Equilibrium Propagation (CVPRW 2021)](https://openaccess.thecvf.com/content/CVPR2021W/BiVision/papers/Laydevant_Training_Dynamical_Binary_Neural_Networks_With_Equilibrium_Propagation_CVPRW_2021_paper.pdf) - [The Forward-Forward Algorithm: Some Preliminary Investigations (Hinton 2022)](https://www.cs.toronto.edu/~hinton/FFA13.pdf) - [Self-Contrastive Forward-Forward Algorithm (Nature Comms 2025)](https://www.nature.com/articles/s41467-025-61037-0) - [Self-Contrastive Forward-Forward Algorithm (arXiv 2409.11593)](https://arxiv.org/abs/2409.11593) - [Efficient Online Learning with Predictive Coding Networks: Exploiting Temporal Correlations (PCN-TA)](https://arxiv.org/abs/2510.25993) - [Introduction to Predictive Coding Networks for Machine Learning](https://arxiv.org/pdf/2506.06332) - [Active Predicting Coding: Brain-Inspired RL for Sparse Reward Robotic Control](https://arxiv.org/pdf/2209.09174) - [A Theoretical Framework for Target Propagation](https://arxiv.org/pdf/2006.14331) - [Towards Scaling Difference Target Propagation by Learning Backprop Targets](https://proceedings.mlr.press/v162/ernoult22a/ernoult22a.pdf) - [Three-factor learning in spiking neural networks (PMC 2025)](https://pmc.ncbi.nlm.nih.gov/articles/PMC12745983/) - [A solution to the learning dilemma for recurrent networks of spiking neurons (e-prop, Nature Comms 2020)](https://www.nature.com/articles/s41467-020-17236-y) - [Including STDP to eligibility propagation in multi-layer recurrent SNNs](https://arxiv.org/pdf/2201.07602) - [BioLCNet: Reward-Modulated Locally Connected SNNs](https://arxiv.org/pdf/2109.05539) - [On the Stability and Scalability of Node Perturbation Learning (NeurIPS 2022)](https://proceedings.neurips.cc/paper_files/paper/2022/file/cf38eb1549024cce4b3d2c1bb87a6c27-Paper-Conference.pdf) - [Effective Learning with Node Perturbation in Deep Neural Networks](https://arxiv.org/html/2310.00965v3) - [Noise-based reward-modulated learning (arXiv 2503.23972)](https://arxiv.org/pdf/2503.23972) - [Weight versus Node Perturbation Learning (Phys. Rev. X 2023)](https://link.aps.org/doi/10.1103/PhysRevX.13.021006) - [Infomax (Wikipedia)](https://en.wikipedia.org/wiki/Infomax) - [Local Synaptic Learning Rules Suffice to Maximize Mutual Information](https://scite.ai/reports/local-synaptic-learning-rules-suffice-4kGY1R) - [Are alternatives to backpropagation useful for training Binary Neural Networks? (ACM SAC 2023)](https://dl.acm.org/doi/10.1145/3555776.3577674) - [Learning Without Feedback: Fixed Random Learning Signals for Deep Networks (PMC)](https://pmc.ncbi.nlm.nih.gov/articles/PMC7902857/) - [Counter-Current Learning: A Biologically Plausible Dual Network Approach](https://arxiv.org/pdf/2409.19841) - [Towards Biologically Plausible Computing: A Comprehensive Comparison](https://arxiv.org/pdf/2406.16062)