When Does Muon Help Agentic Reinforcement Learning?
Updated
Updated · arxiv.org · Jul 17
When Does Muon Help Agentic Reinforcement Learning?
1 articles · Updated · arxiv.org · Jul 17
Summary
A new study finds that the Muon optimizer can significantly improve agentic reinforcement learning (RL) performance compared to AdamW in certain settings.
In experiments on ALFWorld with the Qwen2.5-0.5B-Instruct model, Muon boosted validation success rates by up to 88% when paired with the GiGPO advantage estimator.
The effectiveness of Muon depends on the choice of advantage estimator and learning rate, highlighting the need for joint optimization of RL components.