The score of a density is the gradient of its log:
It points towards regions of higher probability. The normalising constant is irrelevant, since , so scores can be learned for unnormalised models such as Energy-Based Models.
Scores of the noisy marginals. Diffusion models need for the marginal of the forward process at every time . For the Gaussian transition kernel of the DDPM Forward Process,
using .
Noise prediction = score. A network trained to predict the added noise (see Denoising Score Matching) gives the score estimate
Note that estimates the conditional expectation . It therefore targets the score of the marginal , not of the conditional .
Why noise helps. Real data lie near a low-dimensional set (see Manifold Hypothesis). The score of is undefined or unreliable off that set. Adding noise fills the ambient space, which gives well-defined scores everywhere and a continuum of noise levels to anneal through.
References
- A. Hyvärinen (2005). Estimation of Non-Normalized Statistical Models by Score Matching. J. Mach. Learn. Res. 6, 695–709. jmlr.org/papers/v6/hyvarinen05a.html
- Y. Song, S. Ermon (2019). Generative Modeling by Estimating Gradients of the Data Distribution. NeurIPS 32. arXiv:1907.05600
- Y. Song et al. (2021). Score-Based Generative Modeling through Stochastic Differential Equations. ICLR 2021. arXiv:2011.13456