Methods from machine learning use only (possibly noisy) gradients, with per-coordinate adaptive step sizes:

  • SGD with momentum: , .
  • RMSProp: scales by a running root-mean-square of past gradients.
  • Adam: exponential moving averages of the gradient () and its square (), with bias correction:
  • NAdam adds Nesterov momentum to Adam; AdaGrad accumulates all past squared gradients.

In EIT. The objective is a sum over current patterns, . Using one pattern, or a random subset, per step gives an unbiased stochastic gradient at a fraction of the cost. This is the finite-sum setting of SGD. Stochastic objectives such as the RED-Diff regulariser also require these methods, since their gradients are random by construction.

Caveats. Adaptive methods do not use line searches or curvature information. On deterministic, ill-conditioned least-squares problems such as full-batch EIT, Gauss–Newton or L-BFGS typically converge much faster. The coordinate-wise scaling of Adam also depends on the discretisation.

References

  1. D. P. Kingma, J. Ba (2015). Adam: A Method for Stochastic Optimization. ICLR 2015. arXiv:1412.6980
  2. L. Bottou, F. E. Curtis, J. Nocedal (2018). Optimization Methods for Large-Scale Machine Learning. SIAM Review 60(2), 223–311. doi:10.1137/16M1080173