A diffusion denoiser must know the noise level. A scalar is fed in through a sinusoidal embedding into a -dimensional vector:
for . The frequencies are geometrically spaced, so the network can resolve both coarse and fine differences in . The embedding is typically passed through a small MLP and added to the feature maps of every residual block of the U-Net.
For continuous , is scaled, for example by 1000, before embedding. Random Fourier features ( drawn from a Gaussian) are an alternative.
References
- A. Vaswani et al. (2017). Attention Is All You Need. NeurIPS 30. arXiv:1706.03762
- J. Ho, A. Jain, P. Abbeel (2020). Denoising Diffusion Probabilistic Models. NeurIPS 33. arXiv:2006.11239
- M. Tancik et al. (2020). Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains. NeurIPS 33. arXiv:2006.10739