SAC training module #47

Closed
opened 2026-08-20 23:28:46 +02:00 by SirStone · 1 comment
Owner

Parent

#37

What to build

The SAC-v2 training update logic. Given a batch of sequences from the replay buffer, computes losses and updates all networks.

Training step:

  1. Burn-in: forward pass through LSTM for burn-in steps (no loss, just warm up hidden state)
  2. Critic update: minimize TD error using twin Q-networks (take minimum), target networks for bootstrap
  3. Actor update: maximize expected Q-value minus alpha × log_prob
  4. Alpha update: adjust entropy temperature toward target entropy (configurable, default -4.0)
  5. Target network soft update: polyak averaging with tau (configurable, default 0.005)

Gradient clipping at max_norm=1.0 on critic gradients. All learning rates configurable via env vars.

Acceptance criteria

  • SAC update produces finite (non-NaN) losses for actor, critics, and alpha
  • Soft target update blends weights correctly at configured tau
  • Alpha stays positive throughout training
  • Burn-in steps update hidden state but produce no gradients
  • Gradient clipping applied to critic updates
  • test_training.nim passes

Blocked by

  • #41 (LSTM network module)
  • #44 (Reward module)
  • #45 (Replay buffer module)
  • #46 (Weight persistence module)
## Parent #37 ## What to build The SAC-v2 training update logic. Given a batch of sequences from the replay buffer, computes losses and updates all networks. Training step: 1. Burn-in: forward pass through LSTM for burn-in steps (no loss, just warm up hidden state) 2. Critic update: minimize TD error using twin Q-networks (take minimum), target networks for bootstrap 3. Actor update: maximize expected Q-value minus alpha × log_prob 4. Alpha update: adjust entropy temperature toward target entropy (configurable, default -4.0) 5. Target network soft update: polyak averaging with tau (configurable, default 0.005) Gradient clipping at max_norm=1.0 on critic gradients. All learning rates configurable via env vars. ## Acceptance criteria - [ ] SAC update produces finite (non-NaN) losses for actor, critics, and alpha - [ ] Soft target update blends weights correctly at configured tau - [ ] Alpha stays positive throughout training - [ ] Burn-in steps update hidden state but produce no gradients - [ ] Gradient clipping applied to critic updates - [ ] `test_training.nim` passes ## Blocked by - #41 (LSTM network module) - #44 (Reward module) - #45 (Replay buffer module) - #46 (Weight persistence module)
SirStone added the ready-for-agent label 2026-08-20 23:28:46 +02:00
Author
Owner

SAC training module complete — SAC-v2 with manual backprop, truncated BPTT (ponytail: full BPTT if gradient quality matters), twin critics, auto-alpha, soft target update. 11/11 training tests pass, full suite 80+ tests green. Post-review fixes applied: actor hidden state ordering corrected, redundant critic forwards eliminated (4→2), actor gradient clipping added.

SAC training module complete — SAC-v2 with manual backprop, truncated BPTT (ponytail: full BPTT if gradient quality matters), twin critics, auto-alpha, soft target update. 11/11 training tests pass, full suite 80+ tests green. Post-review fixes applied: actor hidden state ordering corrected, redundant critic forwards eliminated (4→2), actor gradient clipping added.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#47