LSTM network module #41

Closed
opened 2026-08-20 23:27:55 +02:00 by SirStone · 1 comment
Owner

Parent

#37

What to build

The neural network module containing the LSTM-based actor and dual critic networks for SAC-v2.

Actor: Linear(state_dim → 256) + ReLU → LSTM(256, 256) → Linear(256 → 128) + ReLU → two heads: μ (4) and log_σ (4, clamped [-5, 2]). Squashed Gaussian policy with reparameterization trick and tanh correction.

Critic (×2): Linear(state_dim + action_dim → 256) + ReLU → LSTM(256, 256) → Linear(256 → 128) + ReLU → Linear(128 → 1).

Hidden size is configurable via SACLSTM_HIDDEN_SIZE env var (default: 256). Weight init: He-style for hidden layers, smaller scale for output.

Supports deterministic mode (output μ directly, no sampling) via SACLSTM_EVAL_MODE env var.

Acceptance criteria

  • Actor forward pass: given state tensor + (h, c), returns (actions, log_prob, h', c') with correct shapes
  • Critic forward pass: given state+action tensor + (h, c), returns (Q-value scalar, h', c')
  • LSTM hidden state dimensions propagate correctly across sequential calls
  • Deterministic mode suppresses sampling noise and outputs μ directly
  • Hidden size responds to env var configuration
  • test_network.nim passes

Blocked by

  • #38 (SAC_LSTM_Bot: project scaffold)
## Parent #37 ## What to build The neural network module containing the LSTM-based actor and dual critic networks for SAC-v2. **Actor:** Linear(state_dim → 256) + ReLU → LSTM(256, 256) → Linear(256 → 128) + ReLU → two heads: μ (4) and log_σ (4, clamped [-5, 2]). Squashed Gaussian policy with reparameterization trick and tanh correction. **Critic (×2):** Linear(state_dim + action_dim → 256) + ReLU → LSTM(256, 256) → Linear(256 → 128) + ReLU → Linear(128 → 1). Hidden size is configurable via `SACLSTM_HIDDEN_SIZE` env var (default: 256). Weight init: He-style for hidden layers, smaller scale for output. Supports deterministic mode (output μ directly, no sampling) via `SACLSTM_EVAL_MODE` env var. ## Acceptance criteria - [ ] Actor forward pass: given state tensor + (h, c), returns (actions, log_prob, h', c') with correct shapes - [ ] Critic forward pass: given state+action tensor + (h, c), returns (Q-value scalar, h', c') - [ ] LSTM hidden state dimensions propagate correctly across sequential calls - [ ] Deterministic mode suppresses sampling noise and outputs μ directly - [ ] Hidden size responds to env var configuration - [ ] `test_network.nim` passes ## Blocked by - #38 (SAC_LSTM_Bot: project scaffold)
SirStone added the ready-for-agent label 2026-08-20 23:27:55 +02:00
Author
Owner

LSTM network module complete — actor + dual critics with proper LSTM cells, squashed Gaussian policy, deterministic mode. 9/9 tests pass.

LSTM network module complete — actor + dual critics with proper LSTM cells, squashed Gaussian policy, deterministic mode. 9/9 tests pass.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#41