Methods from machine learning use only (possibly noisy) gradients, with per-coordinate adaptive step sizes:
- SGD with momentum: , .
- RMSProp: scales by a running root-mean-square of past gradients.
- Adam: exponential moving averages of the gradient () and its square (), with bias correction:
- NAdam adds Nesterov momentum to Adam; AdaGrad accumulates all past squared gradients.
In EIT. The objective is a sum over current patterns, . Using one pattern, or a random subset, per step gives an unbiased stochastic gradient at a fraction of the cost. This is the finite-sum setting of SGD. Stochastic objectives such as the RED-Diff regulariser also require these methods, since their gradients are random by construction.
Caveats. Adaptive methods do not use line searches or curvature information. On deterministic, ill-conditioned least-squares problems such as full-batch EIT, Gauss–Newton or L-BFGS typically converge much faster. The coordinate-wise scaling of Adam also depends on the discretisation.
References
- D. P. Kingma, J. Ba (2015). Adam: A Method for Stochastic Optimization. ICLR 2015. arXiv:1412.6980
- L. Bottou, F. E. Curtis, J. Nocedal (2018). Optimization Methods for Large-Scale Machine Learning. SIAM Review 60(2), 223–311. doi:10.1137/16M1080173