An energy-based model represents a probability density through a learned scalar energy :
Low energy means high probability. The normalising constant is intractable, but many tasks do not need it:
- the score is free of (see Score Function);
- sampling works with Langevin Dynamics;
- as a regulariser, plugs directly into a variational reconstruction (MAP estimate), and its gradient comes from autodiff.
Training.
- Maximum likelihood / contrastive divergence: . The negative phase needs samples from the model, usually short-run Langevin chains.
- Score matching / denoising score matching: fit to the score of noise-perturbed data (see Denoising Score Matching). This avoids sampling.
Why interesting for EIT. An explicit energy gives an objective that can be evaluated, a gradient, and a well-defined Proximal Operator. These plug into the same optimisers as the physics. Diffusion models only provide scores, not energies.
References
- Y. LeCun, S. Chopra, R. Hadsell, M. Ranzato, F. J. Huang (2006). A Tutorial on Energy-Based Learning. In: Predicting Structured Data, MIT Press. yann.lecun.com/exdb/publis/pdf/lecun-06.pdf
- Y. Du, I. Mordatch (2019). Implicit Generation and Generalization in Energy-Based Models. NeurIPS 32. arXiv:1903.08689
- Y. Song, D. P. Kingma (2021). How to Train Your Energy-Based Models. arXiv:2101.03288