An energy-based model represents a probability density through a learned scalar energy :

Low energy means high probability. The normalising constant is intractable, but many tasks do not need it:

Training.

  • Maximum likelihood / contrastive divergence: . The negative phase needs samples from the model, usually short-run Langevin chains.
  • Score matching / denoising score matching: fit to the score of noise-perturbed data (see Denoising Score Matching). This avoids sampling.

Why interesting for EIT. An explicit energy gives an objective that can be evaluated, a gradient, and a well-defined Proximal Operator. These plug into the same optimisers as the physics. Diffusion models only provide scores, not energies.

References

  1. Y. LeCun, S. Chopra, R. Hadsell, M. Ranzato, F. J. Huang (2006). A Tutorial on Energy-Based Learning. In: Predicting Structured Data, MIT Press. yann.lecun.com/exdb/publis/pdf/lecun-06.pdf
  2. Y. Du, I. Mordatch (2019). Implicit Generation and Generalization in Energy-Based Models. NeurIPS 32. arXiv:1903.08689
  3. Y. Song, D. P. Kingma (2021). How to Train Your Energy-Based Models. arXiv:2101.03288