The overdamped Langevin SDE
has as its stationary distribution. Discretising with the Euler–Maruyama method (see Euler-Maruyama Method) gives the unadjusted Langevin algorithm (ULA):
For an energy-based model this is : gradient descent plus correctly scaled noise. The noise scales like the square root of the step size.
- For small the chain samples approximately from . The bias from discretisation is removed by a Metropolis correction (MALA).
- Stochastic gradient Langevin dynamics (SGLD) (Welling & Teh 2011) uses mini-batch gradients and decreasing step sizes, which turns an SGD optimiser into a posterior sampler.
- Posterior sampling for inverse problems: replace by , a data gradient plus a learned score. Annealed Langevin with decreasing noise levels was the first score-based generative sampler (Song & Ermon 2019).
References
- M. Welling, Y. W. Teh (2011). Bayesian Learning via Stochastic Gradient Langevin Dynamics. ICML 2011. stats.ox.ac.uk/~teh/research/compstats/WelTeh2011a.pdf
- G. O. Roberts, R. L. Tweedie (1996). Exponential convergence of Langevin distributions and their discrete approximations. Bernoulli 2(4), 341–363. doi:10.2307/3318418
- Y. Song, S. Ermon (2019). Generative Modeling by Estimating Gradients of the Data Distribution. NeurIPS 32. arXiv:1907.05600