Optimise-then-discretise (OTD). Derive the optimality system (state, adjoint, gradient) in function spaces, as in Lagrangian Formulation, then discretise each equation. The discrete gradient approximates the continuous gradient , but it is not necessarily the exact gradient of the discrete objective.
Discretise-then-optimise (DTO). Discretise the objective and PDE first: s.t. . Then differentiate the finite-dimensional problem exactly. The discrete adjoint is
(the last equality holds if the quadrature is exact for the integrand). This is what AD produces.
When do they coincide? For the conductivity equation with a Galerkin discretisation, the two agree whenever the gradient is assembled with the same quadrature as and represented in the dual (coefficient) sense. Differences appear when
- the gradient is interpolated or projected onto a space in which does not lie;
- different quadrature rules are used for assembly and gradient;
- solvers are stopped early (inexact states and adjoints);
- the misfit is measured in a norm other than the one used for discretisation.
Both gradients come from the same assembled quantity. The DTO gradient is the dual vector , and the -projected OTD gradient is times it (see Conductivity Tensor).
DTO gradients make line searches and quasi-Newton methods behave consistently, because they are exact derivatives of the function being minimised. OTD gradients are mesh-independent approximations of the true gradient (see Gradient Representation and the Riesz Map).
In ModularEIT.jl: CoefficientGradient, L2Gradient.
References
- M. D. Gunzburger (2002). Perspectives in Flow Control and Optimization. SIAM. doi:10.1137/1.9780898718720
- M. Hinze, R. Pinnau, M. Ulbrich, S. Ulbrich (2009). Optimization with PDE Constraints. Springer. doi:10.1007/978-1-4020-8839-1