The marginal score is unknown, but it can be learned from samples without knowing it.

Denoising score matching (Vincent 2011). For any noise kernel ,

Regressing onto the conditional score, which is known in closed form, gives the marginal score at the optimum.

-prediction loss (DDPM). Inserting the Gaussian kernel of the DDPM Forward Process and parametrising gives the simple training objective

with , , and weight in DDPM. The network sees the noisy image and the time (via a Sinusoidal Time Embedding) and predicts the noise. Its optimum is .

Training loop. Sample a batch of images, times and noise; form ; take a gradient step on the squared error. No simulation of the SDE and no ODE solver are needed, which makes training cheap compared with Neural ODEs. Other parametrisations predict or ; they differ only in the implied weighting .

References

  1. P. Vincent (2011). A Connection Between Score Matching and Denoising Autoencoders. Neural Computation 23(7), 1661–1674. doi:10.1162/NECO_a_00142
  2. J. Ho, A. Jain, P. Abbeel (2020). Denoising Diffusion Probabilistic Models. NeurIPS 33. arXiv:2006.11239
  3. T. Salimans, J. Ho (2022). Progressive Distillation for Fast Sampling of Diffusion Models. ICLR 2022. arXiv:2202.00512