Learning rate and optimizer hyperparameters #12
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Question
What optimizer, learning rate, and related hyperparameters for the PPO actor and critic networks?
Context: separate actor/critic networks (#9), 64→64 tanh, ~8K params total. 4 PPO epochs, mini-batch 64 (#8). Entropy coefficient 0.01 (#10).
Blocked by: nothing (all dependencies resolved)
Resolution
Optimizer: Adam. Two instances — one for actor, one for critic. Arraymancer has Adam built in (confirmed working in spike).
Learning rate: 3e-4, same for actor and critic. Fixed — no decay (the bot trains indefinitely across battles, no endpoint to decay toward).
Adam hyperparameters: Defaults — beta1=0.9, beta2=0.999, epsilon=1e-8.
Gradient clipping: Max global norm = 0.5. Standard PPO practice, prevents catastrophic weight updates from noisy advantages. May need hand-rolling if Arraymancer doesn't support it natively (a few lines).
Upgrade paths: