TY - UNPB
T1 - A unified convergence theory for adaptive first-order methods in the nonconvex case, including AdaNorm, full and diagonal AdaGrad, Shampoo and Muon
AU - Gratton, Serge
AU - Toint, Philippe
PY - 2026/4/17
Y1 - 2026/4/17
N2 - A unified framework for first-order optimization algorithms for nonconvex unconstrained optimization is proposed that uses adaptively preconditioned gradients and includes popular methods such as full and diagonal AdaGrad, AdaNorm, as well as adpative variants of Shampoo and Muon. This framework also allows combining heterogeneous geometries across different groups of variables while preserving a unified convergence analysis. A fully stochastic global rate-of-convergence analysis is conducted for all methods in the framework, with and without two types of momentum, using reasonable assumptions on the
variance of the gradient oracle and without assuming bounded stochastic gradients or small enough stepsize.
AB - A unified framework for first-order optimization algorithms for nonconvex unconstrained optimization is proposed that uses adaptively preconditioned gradients and includes popular methods such as full and diagonal AdaGrad, AdaNorm, as well as adpative variants of Shampoo and Muon. This framework also allows combining heterogeneous geometries across different groups of variables while preserving a unified convergence analysis. A fully stochastic global rate-of-convergence analysis is conducted for all methods in the framework, with and without two types of momentum, using reasonable assumptions on the
variance of the gradient oracle and without assuming bounded stochastic gradients or small enough stepsize.
M3 - Preprint
VL - 2604.17423
BT - A unified convergence theory for adaptive first-order methods in the nonconvex case, including AdaNorm, full and diagonal AdaGrad, Shampoo and Muon
PB - Arxiv
ER -