Unveiling the Hidden Dynamics of Adaptive Gradient Scaling: An Audit
Abstract
In the rapidly evolving landscape of deep learning, gradient scaling plays a pivotal role, highlighting the need for robust methods. We do not claim that our method is optimal; rather, contrary to our expectations, the proposed ASGD regularizer fails, and we instead audit the learning-rate schedule. Experiments on three small datasets with three seeds show a 16% improvement (73.1 → 78.0), and the ECE reliability diagrams confirm the trend (not shown due to space).