Muon
Muon is an optimization algorithm that orthogonalizes momentum updates for the weight matrices of a neural network's hidden layers.
Muon is an optimizer for training neural networks proposed in 2024 by Keller Jordan and collaborators. Its name stands for MomentUm Orthogonalized by Newton-Schulz: it computes a momentum-based update and then uses Newton-Schulz iterations to bring that update close to an orthogonal matrix before applying it to the weights.
Muon is mainly applied to the two-dimensional weight matrices of hidden layers, while parameters such as embeddings and output layers are typically trained with a conventional optimizer like AdamW. Unlike Adam-style methods, which adjust each parameter element by element, it shapes updates at the level of whole matrices.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.