Phiacta
ExplorePostDocsGuidesContributingAbout

Phiacta

The knowledge backend.

Contact Us
ExploreGeodesic Convexity Theory for the Induced-Metric Optimizer

Geodesic Convexity Theory for the Induced-Metric Optimizer

A complete pointwise theory of when the induced metric can convert a non-convex loss into a geodesically convex one. The Riemannian Hessian is (Hess_g L)_ij = (H_ij - C_ij)/(1 + ξ‖∇L‖²_γ), where C_ij = Γ^k_ij(γ)·∇_k L is a curvature correction depending on ∇γ and ∇L but not on H. Three-level hierarchy: (1) constant γ gives C ∝ H — eigenvalue signs are preserved, saddles cannot be fixed; (2) scalar γ = e^s I gives trace-constrained C (tr(C) = 0 in 2D), which can fix asymmetric saddles but not symmetric ones; (3) diagonal γ = diag(e^{s_i}) gives unrestricted C via exponential anisotropy ratios e^{s_i - s_j}, enabling sign-flipping for all N. The diagonal-N theorem is constructively proved via a Gershgorin argument and formally verified in Lean 4 (zero sorry). A universal critical-point obstruction (C ∝ ∇L vanishes at saddles) precludes global geodesic convexity for any function with stationary points, but the geodesic convexity basin around any minimum can be dramatically enlarged beyond the Euclidean convexity basin (demonstrated on the quartic well, four disconnected basins merge into one). The optimal correction has anti-correlation structure (∇s_i with sign opposite to H_ii) which motivates the curvature-aware diagonal variant.

ContentIssuesEditsHistoryFilesReferences5

Geodesic Convexity Theory for the Induced-Metric Optimizer

1. Setup

A function L:RN→RL: \mathbb{R}^N \to \mathbb{R}L:RN→R is geodesically convex with respect to a Riemannian metric ggg if its Riemannian Hessian is positive semidefinite at every point. A geodesic in this context is the shortest path between two points measured using ggg, equivalently a curve whose tangent vector is parallel-transported along itself. Geodesically convex functions are the Riemannian analogue of convex functions: they have no saddle points and no spurious local minima, and every local minimum is the global minimum. This entry asks when the induced metric from a loss embedding can make a non-convex loss geodesically convex.

The induced metric arising from the loss embedding ϕ(θ)=(θ,L(θ))\phi(\theta) = (\theta, L(\theta))ϕ(θ)=(θ,L(θ)) with ambient metric diag(γ,ξ)\mathrm{diag}(\gamma, \xi)diag(γ,ξ) has the form

gij(θ)=γij(θ)+ξ ℓi(θ) ℓj(θ),g_{ij}(\theta) = \gamma_{ij}(\theta) + \xi\, \ell_i(\theta)\, \ell_j(\theta),gij​(θ)=γij​(θ)+ξℓi​(θ)ℓj​(θ),

where ℓ=∇L\ell = \nabla Lℓ=∇L and γ\gammaγ is a positive-definite matrix field on parameter space. The Riemannian Hessian of LLL with respect to ggg has the closed form

(HessgL)ij=Hij−Cij1+ξ ∥ℓ∥γ2,(\mathrm{Hess}_g L)_{ij} = \frac{H_{ij} - C_{ij}}{1 + \xi\,\|\ell\|_\gamma^2},(Hessg​L)ij​=1+ξ∥ℓ∥γ2​Hij​−Cij​​,

where Hij=∂i∂jLH_{ij} = \partial_i \partial_j LHij​=∂i​∂j​L is the ordinary Euclidean Hessian, ∥ℓ∥γ2=ℓ⊤γ−1ℓ\|\ell\|_\gamma^2 = \ell^\top \gamma^{-1} \ell∥ℓ∥γ2​=ℓ⊤γ−1ℓ, and

Cij=Γijk(γ) ℓkC_{ij} = \Gamma^k_{ij}(\gamma)\,\ell_kCij​=Γijk​(γ)ℓk​

is what we call the curvature correction. The full derivation of this formula is in Appendix A, which also defines the Christoffel bracket and Christoffel symbols used throughout. The denominator is strictly positive, so the eigenvalue sign structure of HessgL\mathrm{Hess}_g LHessg​L is determined entirely by H−CH - CH−C.

The point of this entry is the structure of CCC. It depends on the spatial derivatives of γ\gammaγ and on the gradient ℓ\ellℓ, but not on HHH. The correction has its own eigenvalue structure, independent of the Hessian. If γ\gammaγ can be chosen so that CCC cancels the negative eigenvalues of HHH, the function becomes geodesically convex.

The three levels of the hierarchy below correspond to three families of induced-metric optimizers. The fixed variant has γ\gammaγ constant. The learnable scalar variant has γ=es(θ)I\gamma = e^{s(\theta)} Iγ=es(θ)I for a single scalar function sss. The learnable diagonal variant has γ=diag(es1(θ),…,esN(θ))\gamma = \mathrm{diag}(e^{s_1(\theta)}, \ldots, e^{s_N(\theta)})γ=diag(es1​(θ),…,esN​(θ)) for NNN independent scale functions. Each level adds degrees of freedom to γ\gammaγ, and we will see exactly what that unlocks.

Sign convention for sss. Throughout this entry, sss parameterises the metric directly: increasing sis_isi​ increases the metric weight in direction iii, and therefore decreases the parameter step in that direction. The learnable diagonal entry parameterises the inverse metric instead (γ−1=diag(esi)\gamma^{-1} = \mathrm{diag}(e^{s_i})γ−1=diag(esi​)), so its sss has the opposite sign: increasing sis_isi​ in that entry's convention means stronger preconditioning and a larger step in direction iii. The underlying object (the Newton preconditioner) is the same in both entries; the formulas look opposite because they parameterise opposite things. When transferring a formula across, substitute s→−ss \to -ss→−s.

One note on dimension. In neural-network applications, NNN (the number of parameters) is huge, typically in the millions or billions. The two-dimensional examples we use below are pedagogical, because two is the smallest dimension in which a saddle point exists and the algebra is fully transparent. The general-NNN statements appear afterwards.

The Riemannian Hessian is a (0, 2)-tensor and its eigenvalues with respect to ggg are coordinate-independent quantities, so the analysis below is genuinely geometric and not an artifact of coordinate choice.

2. Level 0: fixed metric

When γ\gammaγ is constant, all its derivatives vanish. The full induced metric ggg still has non-zero Christoffel symbols, because the ξℓℓ⊤\xi \ell \ell^\topξℓℓ⊤ piece depends on θ\thetaθ, but the calculation in Appendix A gives them in a special form,

Γijk(g)=ξ Hij ℓk1+ξ ∥ℓ∥2.\Gamma^k_{ij}(g) = \frac{\xi\,H_{ij}\,\ell^k}{1 + \xi\,\|\ell\|^2}.Γijk​(g)=1+ξ∥ℓ∥2ξHij​ℓk​.

Contracting with ℓk\ell_kℓk​,

Cij=ξ ∥ℓ∥21+ξ ∥ℓ∥2 Hij.C_{ij} = \frac{\xi\,\|\ell\|^2}{1 + \xi\,\|\ell\|^2}\,H_{ij}.Cij​=1+ξ∥ℓ∥2ξ∥ℓ∥2​Hij​.

The correction is proportional to HHH itself. Substituting into the Riemannian Hessian formula,

(HessgL)ij=11+ξ ∥ℓ∥2 Hij.(\mathrm{Hess}_g L)_{ij} = \frac{1}{1 + \xi\,\|\ell\|^2}\,H_{ij}.(Hessg​L)ij​=1+ξ∥ℓ∥21​Hij​.

The Riemannian Hessian equals the Euclidean Hessian multiplied by a positive scalar in (0,1](0, 1](0,1]. Concretely, if HHH has eigenvalues λ1,…,λN\lambda_1, \ldots, \lambda_Nλ1​,…,λN​ with eigenvectors v1,…,vNv_1, \ldots, v_Nv1​,…,vN​, then HessgL\mathrm{Hess}_g LHessg​L has the same eigenvectors with eigenvalues λi/(1+ξ∥ℓ∥2)\lambda_i / (1 + \xi \|\ell\|^2)λi​/(1+ξ∥ℓ∥2). The signs are preserved, the eigenvectors are preserved, and only the magnitudes are uniformly rescaled. A saddle remains a saddle and a bowl remains a bowl. The fixed induced metric provides adaptive curvature damping but zero curvature correction.

The same conclusion holds for the fixed symmetric off-diagonal extension of the ambient metric (constant coupling vectors a=ba = ba=b; see the off-diagonal metric variant). The Christoffel bracket factorises the same way and the Riemannian Hessian remains a scalar multiple of HHH. The asymmetric case (a≠ba \ne ba=b) does not yield a Riemannian metric in the first place and is treated separately in the off-diagonal entry as a non-symmetric preconditioner. The full derivation is in Appendix B.

3. Level 1: scalar learnable

Set γ=es(θ)I\gamma = e^{s(\theta)} Iγ=es(θ)I for a single scalar function s:RN→Rs: \mathbb{R}^N \to \mathbb{R}s:RN→R, the learnable scalar variant. The metric scales every direction by the same factor ese^ses, but the factor varies in space. This is a conformally flat metric. Its Christoffel symbols are the classical conformal expression,

Γijk(γ)=12(∂is δjk+∂js δik−∂ks δij).\Gamma^k_{ij}(\gamma) = \tfrac{1}{2}\bigl(\partial_i s\,\delta_{jk} + \partial_j s\,\delta_{ik} - \partial_k s\,\delta_{ij}\bigr).Γijk​(γ)=21​(∂i​sδjk​+∂j​sδik​−∂k​sδij​).

The derivation in our general framework is given in Appendix A. Contracting with ℓk\ell_kℓk​,

Cij=12(∂is ℓj+∂js ℓi)−12(∇s⋅ℓ) δij.C_{ij} = \tfrac{1}{2}\bigl(\partial_i s\,\ell_j + \partial_j s\,\ell_i\bigr) - \tfrac{1}{2}(\nabla s \cdot \ell)\,\delta_{ij}.Cij​=21​(∂i​sℓj​+∂j​sℓi​)−21​(∇s⋅ℓ)δij​.

The correction is a symmetric rank-2 outer product (the first two terms) plus a scalar multiple of the identity (the third term).

3.1 The trace identity

Take the trace of CCC, that is, sum the diagonal entries. The diagonal of CCC has Cii=12(∂is ℓi+∂is ℓi)−12(∇s⋅ℓ)=∂is ℓi−12(∇s⋅ℓ)C_{ii} = \tfrac{1}{2}(\partial_i s\,\ell_i + \partial_i s\,\ell_i) - \tfrac{1}{2}(\nabla s\cdot\ell) = \partial_i s\,\ell_i - \tfrac{1}{2}(\nabla s\cdot\ell)Cii​=21​(∂i​sℓi​+∂i​sℓi​)−21​(∇s⋅ℓ)=∂i​sℓi​−21​(∇s⋅ℓ). Summing over iii,

tr(C)=∑i∂is ℓi−N2(∇s⋅ℓ)=(∇s⋅ℓ)−N2(∇s⋅ℓ)=2−N2 (∇s⋅ℓ).\mathrm{tr}(C) = \sum_i \partial_i s\,\ell_i - \tfrac{N}{2}(\nabla s\cdot\ell) = (\nabla s\cdot\ell) - \tfrac{N}{2}(\nabla s\cdot\ell) = \frac{2 - N}{2}\,(\nabla s \cdot \ell).tr(C)=i∑​∂i​sℓi​−2N​(∇s⋅ℓ)=(∇s⋅ℓ)−2N​(∇s⋅ℓ)=22−N​(∇s⋅ℓ).

This is direct bookkeeping from the explicit formula for CCC, without further assumption. It is a standard consequence of conformal geometry: the trace of the correction is fixed by the dimension and by the inner product of ∇s\nabla s∇s with the gradient. In two dimensions (N=2N = 2N=2), the trace is identically zero, regardless of how cleverly sss varies in space. The conformal correction is trace-free in two dimensions. The metric can redistribute eigenvalues but cannot change their sum. We refer to this as the trace constraint, with the understanding that it is a hard algebraic identity in N=2N = 2N=2 and not an inequality.

3.2 Consequences in two dimensions

Consider two two-dimensional examples.

For the asymmetric saddle L=3x2−y2L = 3x^2 - y^2L=3x2−y2 at (1,1)(1, 1)(1,1), the Euclidean Hessian is H=diag(6,−2)H = \mathrm{diag}(6, -2)H=diag(6,−2) with tr(H)=4>0\mathrm{tr}(H) = 4 > 0tr(H)=4>0. Since tr(H−C)=tr(H)=4\mathrm{tr}(H - C) = \mathrm{tr}(H) = 4tr(H−C)=tr(H)=4, redistributing the eigenvalues to (2,2)(2, 2)(2,2) would give H−C=2IH - C = 2IH−C=2I, which is positive definite. Choosing ∇s=(6/5,2/5)\nabla s = (6/5, 2/5)∇s=(6/5,2/5) gives exactly C=diag(4,−4)C = \mathrm{diag}(4, -4)C=diag(4,−4), and the Riemannian Hessian is 2I2I2I times the positive scalar 1/(1+ξ∥ℓ∥γ2)1/(1 + \xi \|\ell\|^2_\gamma)1/(1+ξ∥ℓ∥γ2​). Geodesic convexity is achieved.

For the symmetric saddle L=x2−y2L = x^2 - y^2L=x2−y2 at (1,1)(1, 1)(1,1), the Euclidean Hessian is H=diag(2,−2)H = \mathrm{diag}(2, -2)H=diag(2,−2) with tr(H)=0\mathrm{tr}(H) = 0tr(H)=0. Since tr(H−C)=0\mathrm{tr}(H - C) = 0tr(H−C)=0 and a real symmetric matrix with zero trace must have at least one non-positive eigenvalue, strict geodesic convexity is impossible for the scalar metric here at any choice of ∇s\nabla s∇s.

So in two dimensions the scalar metric can fix saddles where positive curvature dominates (tr(H)>0\mathrm{tr}(H) > 0tr(H)>0) but cannot fix symmetric or negative-dominant saddles.

3.3 What about N>2N > 2N>2?

For N>2N > 2N>2 the trace identity becomes tr(C)=((2−N)/2)(∇s⋅ℓ)\mathrm{tr}(C) = ((2-N)/2)(\nabla s \cdot \ell)tr(C)=((2−N)/2)(∇s⋅ℓ), which is no longer identically zero (it vanishes only when ∇s⋅ℓ=0\nabla s \cdot \ell = 0∇s⋅ℓ=0). The hard trace obstruction is gone, and the scalar metric can in principle shift the sum of eigenvalues. The scalar metric is nonetheless still limited in higher dimensions, for a different reason.

The scalar metric has NNN free parameters (the components ∂ks\partial_k s∂k​s for k=1,…,Nk = 1, \ldots, Nk=1,…,N). The condition that H−CH - CH−C be positive definite involves controlling all N(N+1)/2N(N+1)/2N(N+1)/2 entries of a symmetric matrix. For N>2N > 2N>2, the scalar metric is under-determined: it has fewer knobs than constraints. This is not a proof that it always fails, but it does mean that the scalar class is highly restricted in high dimensions. It can fix specific easy Hessians, those that are a low-rank perturbation away from positive-definite, but not arbitrary ones.

A concrete example. With H=diag(1,−1,−1,−1,−1)H = \mathrm{diag}(1, -1, -1, -1, -1)H=diag(1,−1,−1,−1,−1) in N=5N = 5N=5, numerically optimizing λmin⁡(H−C)\lambda_{\min}(H - C)λmin​(H−C) over the components of ∇s\nabla s∇s using the SDP described in Section 4.4 returns a maximum value of approximately −0.8-0.8−0.8. Geodesic convexity is unattainable here, even though the trace constraint is formally lifted. The full numerical run is in the attached Mathematica notebook.

The general lesson is that the scalar learnable metric works when the Hessian is easy and the negative-eigenvalue subspace has small rank, but it lacks the degrees of freedom to handle arbitrary Hessians in high dimensions.

Since real neural-network applications have NNN in the millions, the two-dimensional-only trace constraint is not directly the binding obstruction in practice. The under-determination is the real practical limitation, and it gets worse with NNN.

4. Level 2: diagonal learnable

The learnable diagonal variant replaces the single scalar s(θ)s(\theta)s(θ) with NNN independent scale functions,

γ(θ)=diag(es1(θ), es2(θ), …, esN(θ)).\gamma(\theta) = \mathrm{diag}\bigl(e^{s_1(\theta)},\, e^{s_2(\theta)},\, \ldots,\, e^{s_N(\theta)}\bigr).γ(θ)=diag(es1​(θ),es2​(θ),…,esN​(θ)).

Each si:RN→Rs_i: \mathbb{R}^N \to \mathbb{R}si​:RN→R is its own function of all parameters. The correction CCC now depends on the partial derivatives ∂ksi\partial_k s_i∂k​si​, one for each pair (k,i)(k, i)(k,i). Collect these into a matrix,

σki:=∂ksi,σ∈RN×N.\sigma_{ki} := \partial_k s_i, \qquad \sigma \in \mathbb{R}^{N \times N}.σki​:=∂k​si​,σ∈RN×N.

This σ\sigmaσ is the Jacobian of the vector-valued function s=(s1,…,sN):RN→RNs = (s_1, \ldots, s_N): \mathbb{R}^N \to \mathbb{R}^Ns=(s1​,…,sN​):RN→RN. Each column iii of σ\sigmaσ is the gradient ∇si\nabla s_i∇si​. The diagonal entries σii\sigma_{ii}σii​ describe how the iii-th scale changes in its own coordinate direction, and the off-diagonal entries σki\sigma_{ki}σki​ for k≠ik \ne ik=i describe cross-coupling between scales and coordinates.

4.1 Christoffel structure

The Christoffel symbols of the diagonal metric, derived in Appendix C, are non-zero in three configurations:

(a) Γiii(γ)=12 ∂isi\Gamma^i_{ii}(\gamma) = \tfrac{1}{2}\,\partial_i s_iΓiii​(γ)=21​∂i​si​, the "same-index" piece, with no exponential factor.

(b) Γiik(γ)=−12 esi−sk ∂ksi\Gamma^k_{ii}(\gamma) = -\tfrac{1}{2}\,e^{s_i - s_k}\,\partial_k s_iΓiik​(γ)=−21​esi​−sk​∂k​si​ for k≠ik \ne ik=i, the "cross-index" piece, with exponential anisotropy factor esi−ske^{s_i - s_k}esi​−sk​.

(c) Γijj(γ)=12 ∂isj\Gamma^j_{ij}(\gamma) = \tfrac{1}{2}\,\partial_i s_jΓijj​(γ)=21​∂i​sj​ and Γiji(γ)=12 ∂jsi\Gamma^i_{ij}(\gamma) = \tfrac{1}{2}\,\partial_j s_iΓiji​(γ)=21​∂j​si​ for i≠ji \ne ji=j.

All other components vanish. The factors esi−ske^{s_i - s_k}esi​−sk​ are the key new ingredient. They did not appear in the conformal (scalar) case because there all sis_isi​ were equal, making every esi−sk=1e^{s_i - s_k} = 1esi​−sk​=1.

4.2 Why there is no trace constraint

Compute tr(C)=∑iCii\mathrm{tr}(C) = \sum_i C_{ii}tr(C)=∑i​Cii​ for the diagonal metric. Expanding Cii=∑kΓiik ℓkC_{ii} = \sum_k \Gamma^k_{ii}\,\ell_kCii​=∑k​Γiik​ℓk​ using the cases above,

tr(C)=∑i[12 ∂isi ℓi−12∑k≠i esi−sk ∂ksi ℓk].\mathrm{tr}(C) = \sum_i \Bigl[\tfrac{1}{2}\,\partial_i s_i \,\ell_i - \tfrac{1}{2}\sum_{k \ne i}\, e^{s_i - s_k}\,\partial_k s_i\,\ell_k\Bigr].tr(C)=i∑​[21​∂i​si​ℓi​−21​k=i∑​esi​−sk​∂k​si​ℓk​].

This is a sum of terms with different exponential factors esi−ske^{s_i - s_k}esi​−sk​. When all the sis_isi​ are equal (the scalar case), every esi−sk=1e^{s_i - s_k} = 1esi​−sk​=1 and the sum collapses to the clean identity ((2−N)/2)(∇s⋅ℓ)((2-N)/2)(\nabla s \cdot \ell)((2−N)/2)(∇s⋅ℓ) of Level 1. A sanity-check derivation of this reduction is given in Appendix C. When the sis_isi​ are free to differ, the exponential factors are independent positive numbers and no algebraic identity ties the trace to anything specific. The trace can take any real value.

So when we say "the diagonal metric breaks the trace constraint," we do not mean we have violated a real theorem. We mean that the constraint never existed for this metric class. The Level 1 trace identity is an artifact of insisting all sis_isi​ be equal, and relaxing that uniformity removes the identity.

4.3 The Jacobian as the design variable, and the role of the norm bound

Once we move to the diagonal metric, the object we have freedom to choose is the Jacobian σki=∂ksi\sigma_{ki} = \partial_k s_iσki​=∂k​si​. The correction CCC is determined as a linear function of σ\sigmaσ at any fixed θ\thetaθ (with ℓ\ellℓ and the sis_isi​ playing the role of constants in this fixed-point analysis). To talk about what is achievable, we need to specify what range of σ\sigmaσ we allow.

Pointwise feasibility versus a global metric field. The analysis in this section and the Gershgorin construction in §4.7 treat σ\sigmaσ as a free matrix at the point θ\thetaθ under consideration. For σ\sigmaσ to be the Jacobian of an actual s:RN→RNs: \mathbb R^N \to \mathbb R^Ns:RN→RN on a region, the integrability conditions

∂lσki=∂kσlifor all i,k,l\partial_l \sigma_{ki} = \partial_k \sigma_{li} \qquad \text{for all } i, k, l∂l​σki​=∂k​σli​for all i,k,l

(symmetry of the Hessian of each component sis_isi​) must hold. The constructions below establish pointwise first-jet feasibility: at each θ\thetaθ separately, there exists a σ\sigmaσ that achieves geodesic convexity. Whether those pointwise σ\sigmaσ's can be stitched together into a smooth global metric field γ(θ)=diag(esi(θ))\gamma(\theta) = \mathrm{diag}(e^{s_i(\theta)})γ(θ)=diag(esi​(θ)) is a separate (and nontrivial) question that is open and not addressed here. The online learnable-diagonal optimizer in practice uses an EMA-driven sss that is not required to be the gradient of any global potential; the pointwise feasibility analysis is the relevant theoretical object for its local behaviour.

We impose the Frobenius-norm bound

∥σ∥F:=∑k,iσki2≤R\|\sigma\|_F := \sqrt{\sum_{k,i}\sigma_{ki}^2} \le R∥σ∥F​:=k,i∑​σki2​​≤R

for a budget RRR. There are three reasons for this choice:

  1. Frobenius is the natural Hilbert-Schmidt norm on RN×N\mathbb{R}^{N \times N}RN×N. It treats every entry symmetrically and does not single out any coordinate direction.

  2. ∥σ∥F2=∑k,i(∂ksi)2\|\sigma\|_F^2 = \sum_{k,i}(\partial_k s_i)^2∥σ∥F2​=∑k,i​(∂k​si​)2 is the total Dirichlet energy of the vector field sss. It is the standard "rate of change" measure used in calculus of variations and in regularity theory for vector fields.

  3. The constraint ∥σ∥F≤R\|\sigma\|_F \le R∥σ∥F​≤R is convex and quadratic. Combined with the concave objective λmin⁡(H−C(σ))\lambda_{\min}(H - C(\sigma))λmin​(H−C(σ)) (concave because λmin⁡\lambda_{\min}λmin​ of an affine matrix function is concave), the result is a convex semidefinite program, solvable by standard interior-point methods. Uniqueness of the optimum is not guaranteed in general (active-set ambiguities at the norm-ball boundary, eigenvalue-multiplicity issues), but the optimal value is unique and the set of optimisers is convex.

Other matrix norms would also work and give qualitatively similar results. The operator norm ∥σ∥2\|\sigma\|_2∥σ∥2​ is the actual Lipschitz constant of sss; the bound ∥σ∥2≤∥σ∥F\|\sigma\|_2 \le \|\sigma\|_F∥σ∥2​≤∥σ∥F​ means a Frobenius bound implies a Lipschitz bound. The ℓ1\ell_1ℓ1​ and ℓ∞\ell_\inftyℓ∞​ norms would give similar sign-flip-vs-budget pictures with different quantitative thresholds. We picked Frobenius because it is the smoothest, the most rotation-invariant, and the cleanest for SDP.

The budget RRR is not a free knob in deployment. In the online learnable-diagonal optimizer, the metric-learning rate μ\muμ controls how aggressively the metric can change per step. Combined with the parameter learning rate η\etaη and the rate at which the optimizer moves through parameter space, μ\muμ sets the scale of σ\sigmaσ that the algorithm can produce. Larger μ\muμ allows larger ∥σ∥F\|\sigma\|_F∥σ∥F​. So the budget RRR stands in for "how much the online rule can move the metric within a single optimization step." The exact quantitative relationship between RRR and μ\muμ depends on the specific rule and on the trajectory, but the qualitative correspondence is direct: doubling μ\muμ roughly doubles the achievable RRR.

4.4 The sign-flip budget on the symmetric saddle

Return to L=x2−y2L = x^2 - y^2L=x2−y2 at (1,1)(1, 1)(1,1), with H=diag(2,−2)H = \mathrm{diag}(2, -2)H=diag(2,−2) and ℓ=(2,−2)\ell = (2, -2)ℓ=(2,−2), the case where the scalar metric provably failed. With the diagonal metric, the free variables are the four entries of σ∈R2×2\sigma \in \mathbb{R}^{2 \times 2}σ∈R2×2.

We numerically solve

max⁡∥σ∥F≤R  λmin⁡((H−C(σ,ℓ)) / (1+ξ ∥ℓ∥γ2)),\max_{\|\sigma\|_F \le R}\;\lambda_{\min}\bigl((H - C(\sigma, \ell))\,/\,(1 + \xi\,\|\ell\|^2_\gamma)\bigr),∥σ∥F​≤Rmax​λmin​((H−C(σ,ℓ))/(1+ξ∥ℓ∥γ2​)),

i.e., the smallest eigenvalue of the actual Riemannian Hessian rather than of the unnormalised H−CH - CH−C. With ξ=1\xi = 1ξ=1 and γ=I\gamma = Iγ=I at this evaluation point, ∥ℓ∥γ2=8\|\ell\|^2_\gamma = 8∥ℓ∥γ2​=8 and the denominator is 999. Since CCC is linear in σ\sigmaσ and the denominator is constant in σ\sigmaσ, the objective is a concave function of σ\sigmaσ. We use Mathematica's NMinimize with the SDP backend; the notebook is attached as notebooks/sign-flip-sdp.nb.

The numerical results (Riemannian-Hessian eigenvalue, after dividing by 999):

Budget RRRMax λmin⁡\lambda_{\min}λmin​ of Riemannian HessianSign flip
0−0.222-0.222−0.222 (=−2/9= -2/9=−2/9)No (matches Level 0)
1.5−0.009-0.009−0.009No
1.6≈0\approx 0≈0Threshold crossed
2+0.055+0.055+0.055Yes
5+0.360+0.360+0.360Yes
10+0.770+0.770+0.770Yes

At R≈1.6R \approx 1.6R≈1.6 the minimum eigenvalue crosses zero. With R=5R = 5R=5 the saddle is converted to a comfortable Riemannian bowl. This was impossible for the scalar metric at any budget.

Sign-flip budget vs eigenvalue on the symmetric saddle: \lambda_{\min}(H - C) crosses zero at R \approx 1.6, where the indefinite Hessian is converted into a positive Riemannian Hessian.

4.5 The anti-correlation structure of the optimal σ\sigmaσ

At the optimum, σ\sigmaσ has a characteristic shape: its diagonal entries are anti-correlated with sign(Hii)\mathrm{sign}(H_{ii})sign(Hii​). Concretely, σii<0\sigma_{ii} < 0σii​<0 in positive-curvature directions (shrink the metric where the loss curves up) and σii>0\sigma_{ii} > 0σii​>0 in negative-curvature directions (grow the metric where the loss curves down). The off-diagonal entries of σ\sigmaσ are small in this two-dimensional example. This anti-correlation is the geometric mechanism that produces non-zero trace in CCC. It also motivates the curvature-aware variant of the learnable diagonal optimizer, which adds a term to the online update rule that drives σii\sigma_{ii}σii​ in this anti-correlated direction.

4.6 Other saddles

The same numerical procedure works on every saddle we tested. For the asymmetric saddle L=x2−3y2L = x^2 - 3y^2L=x2−3y2, where the scalar metric struggles because tr(H)<0\mathrm{tr}(H) < 0tr(H)<0, the diagonal still flips signs at moderate RRR. For the monkey saddle L=x3−3xy2L = x^3 - 3xy^2L=x3−3xy2, with its position-dependent Hessian, sign-flipping still works. For the Rosenbrock function in regions of negative curvature, sign-flipping is achievable. The corresponding budget tables and optimal σ\sigmaσ values are in notebooks/sign-flip-sdp.nb.

4.7 Argument for general NNN: Gershgorin construction

For general dimension, we can give a constructive sign-flip argument. Choose σ\sigmaσ purely diagonal,

σii=−2M/ℓi,σki=0 for k≠i.\sigma_{ii} = -2M/\ell_i, \qquad \sigma_{ki} = 0 \text{ for } k \ne i.σii​=−2M/ℓi​,σki​=0 for k=i.

The algebra in Appendix C gives two consequences. Each diagonal entry of H−CH - CH−C shifts by +M+M+M. The off-diagonal entries of H−CH - CH−C are unchanged.

By Gershgorin's disc theorem, every eigenvalue of a matrix lies in a disc centred at one of its diagonal entries with radius equal to the corresponding off-diagonal row sum. If M>max⁡i(Ri−Hii)M > \max_i (R_i - H_{ii})M>maxi​(Ri​−Hii​), where Ri=∑j≠i∣Hij∣R_i = \sum_{j \ne i} |H_{ij}|Ri​=∑j=i​∣Hij​∣ is the iii-th off-diagonal row sum, every Gershgorin disc lies strictly in the positive half-plane. Then H−C≻0H - C \succ 0H−C≻0, and the sign flip is achieved for arbitrary NNN.

This argument is constructive as a pointwise existence statement (see §4.3 on pointwise feasibility versus global metric fields). It shows existence of a sign-flipping σ\sigmaσ at each θ\thetaθ separately, requiring ℓi≠0\ell_i \ne 0ℓi​=0 for every iii, and degenerating on coordinate axes where some ℓi=0\ell_i = 0ℓi​=0. The pointwise σ\sigmaσ produced by the construction does not in general satisfy the integrability conditions ∂lσki=∂kσli\partial_l \sigma_{ki} = \partial_k \sigma_{li}∂l​σki​=∂k​σli​, so it cannot in general be integrated to a global s(θ)s(\theta)s(θ); the construction proves feasibility of the first jet, not existence of a smooth metric field. The Gershgorin construction has ∥σ∥F2=4M2∑i1/ℓi2\|\sigma\|_F^2 = 4M^2 \sum_i 1/\ell_i^2∥σ∥F2​=4M2∑i​1/ℓi2​, which grows without bound as any ℓi→0\ell_i \to 0ℓi​→0. For an isotropic gradient with ℓi∼∥ℓ∥/N\ell_i \sim \|\ell\|/\sqrt Nℓi​∼∥ℓ∥/N​, the sum ∑i1/ℓi2∼N2/∥ℓ∥2\sum_i 1/\ell_i^2 \sim N^2/\|\ell\|^2∑i​1/ℓi2​∼N2/∥ℓ∥2, so the total Frobenius budget scales as ∥σ∥F∼2MN/∥ℓ∥\|\sigma\|_F \sim 2 M N/\|\ell\|∥σ∥F​∼2MN/∥ℓ∥, that is O(N/∥ℓ∥)O(N/\|\ell\|)O(N/∥ℓ∥) overall. The budget blows up when needed directions have small gradient components, which is when sign-flipping is structurally hardest.

4.8 Verification

The sign-flip results were verified numerically in Mathematica for N=3,5,10,50,100,1000N = 3, 5, 10, 50, 100, 1000N=3,5,10,50,100,1000 on diagonal Hessians with random off-diagonal perturbations. The notebook notebooks/sign-flip-sdp.nb contains the full code and the eigenvalue tables.

5. The hierarchy

LevelMetric γ\gammaγCorrection CCCTrace identityWhat it can fix
0 (constant)γ\gammaγ constantC∝HC \propto HC∝Htr(C)∝tr(H)\mathrm{tr}(C) \propto \mathrm{tr}(H)tr(C)∝tr(H)Nothing; eigenvalue signs preserved
0+ (fixed off-diag, a=ba = ba=b)const + fixed symmetric a=ba = ba=bC∝HC \propto HC∝Htr(C)∝tr(H)\mathrm{tr}(C) \propto \mathrm{tr}(H)tr(C)∝tr(H)Nothing (symmetric off-diagonal case only; the asymmetric a≠ba \ne ba=b case adds a non-scalar correction whose general analysis is open, see Appendix B)
1 (scalar)esIe^s IesIRank-2 + scalar ⋅\cdot⋅ IIItr(C)=2−N2(∇s⋅ℓ)\mathrm{tr}(C) = \frac{2-N}{2}(\nabla s \cdot \ell)tr(C)=22−N​(∇s⋅ℓ), identically zero in 2D2D: only saddles with tr(H)>0\mathrm{tr}(H) > 0tr(H)>0. High-NNN: under-determined
2 (diagonal)diag(esi)\mathrm{diag}(e^{s_i})diag(esi​)N2N^2N2 free entries (the Jacobian σ\sigmaσ); no algebraic identityUnrestrictedAll saddles in principle, for arbitrary NNN

Each level strictly contains the previous. Setting all sis_isi​ equal in the diagonal case recovers scalar, and setting sss constant recovers fixed. The hierarchy gives a geometric explanation for the empirical performance ordering of induced-metric variants: the variants with richer curvature-correction capability outperform those without, holding implementation cost constant.

6. The critical-point obstruction

Look again at Cij=Γijk(γ) ℓkC_{ij} = \Gamma^k_{ij}(\gamma)\,\ell_kCij​=Γijk​(γ)ℓk​. Every term contains a factor of ℓk\ell_kℓk​ for some kkk. At any critical point θ∗\theta^*θ∗ where ∇L(θ∗)=0\nabla L(\theta^*) = 0∇L(θ∗)=0, we have C(θ∗)≡0C(\theta^*) \equiv 0C(θ∗)≡0, so

HessgL(θ∗)=H(θ∗)\mathrm{Hess}_g L(\theta^*) = H(\theta^*)Hessg​L(θ∗)=H(θ∗)

regardless of the metric γ\gammaγ. The proof is one line: ℓ=0\ell = 0ℓ=0 kills the contraction, and the denominator equals 1.

6.1 What this rules out

Global geodesic convexity is impossible for any loss function with a saddle point. At the saddle, the Riemannian Hessian equals the indefinite Euclidean Hessian, no matter what metric is chosen. This is a fundamental obstruction. The curvature-correction mechanism is most powerful in the bulk and powerless exactly at saddles, where one might most want it to act.

6.2 What this does not rule out

Three things soften the obstruction.

First, optimizers rarely hit exact critical points. Iterates approach saddles asymptotically but generically never land on them, and stochastic noise further prevents it. The obstruction applies to a measure-zero set.

Second, the correction grows linearly with ∥ℓ∥\|\ell\|∥ℓ∥. Near a saddle θ∗\theta^*θ∗, ℓ≈Hδθ\ell \approx H \delta\thetaℓ≈Hδθ, so ∥C∥∝∥δθ∥\|C\| \propto \|\delta\theta\|∥C∥∝∥δθ∥. The correction weakens as the optimizer approaches the saddle, but the optimizer's velocity decreases in step. What matters is the ratio.

Third, the useful question becomes basin enlargement, not global convexity. The next section makes "the obstruction applies to a measure-zero set" quantitative: the failure region at correction budget RRR is contained in tubes of radius O(1/R)O(1/R)O(1/R) around each non-minimum critical point.

6.3 Quantitative obstruction: ε\varepsilonε-excised sign flip

Section 6.1 is a measure-zero negative statement. Its quantitative complement, which characterises how fast the obstruction decays as the optimizer moves away from a saddle, is the main result of this subsection. The failure region of the attainable sign flip shrinks linearly with the correction budget.

Define the attainable geodesic-convexity basin at budget RRR,

Bgopt(R):={θ:∃ σ with ∥σ∥F≤R and H(θ)−C(σ,ℓ(θ))≻0}.\mathcal{B}_g^{\mathrm{opt}}(R) := \bigl\{\theta : \exists\,\sigma\ \text{with}\ \|\sigma\|_F \le R\ \text{and}\ H(\theta) - C(\sigma,\ell(\theta)) \succ 0\bigr\}.Bgopt​(R):={θ:∃σ with ∥σ∥F​≤R and H(θ)−C(σ,ℓ(θ))≻0}.

At each critical point θk∗\theta^*_kθk∗​ of a Morse loss, let

μk∗:=∣λmin⁡(H(θk∗))∣,ρk∗:=max⁡i[∑j≠i∣Hij(θk∗)∣−Hii(θk∗)]\mu^*_k := |\lambda_{\min}(H(\theta^*_k))|, \qquad \rho^*_k := \max_i\Bigl[\sum_{j \ne i} |H_{ij}(\theta^*_k)| - H_{ii}(\theta^*_k)\Bigr]μk∗​:=∣λmin​(H(θk∗​))∣,ρk∗​:=imax​[j=i∑​∣Hij​(θk∗​)∣−Hii​(θk∗​)]

denote the spectral gap and the Gershgorin deficit at θk∗\theta^*_kθk∗​. At a minimum, ρk∗≤0\rho^*_k \le 0ρk∗​≤0; at a saddle or maximum, ρk∗>0\rho^*_k > 0ρk∗​>0.

Theorem 1ʹ (ε\varepsilonε-excised sign flip). Let L∈C3(Ω)L \in C^3(\Omega)L∈C3(Ω) be Morse on a bounded open set Ω⊂RN\Omega \subset \mathbb{R}^NΩ⊂RN with finite critical set Crit(L)\mathrm{Crit}(L)Crit(L). Then for every correction budget R>0R > 0R>0,

Bgopt(R)c∩Ω  ⊆  ⋃k : ρk∗>0B(θk∗, εk(R))‾,\mathcal{B}_g^{\mathrm{opt}}(R)^c \cap \Omega \;\subseteq\; \bigcup_{k\,:\,\rho^*_k > 0} \overline{B\bigl(\theta^*_k,\,\varepsilon_k(R)\bigr)},Bgopt​(R)c∩Ω⊆k:ρk∗​>0⋃​B(θk∗​,εk​(R))​,

with the explicit estimate

εk(R)  ≤  2Nρk∗μk∗ R  +  O(R−2)as R→∞.\varepsilon_k(R) \;\le\; \frac{2 N \rho^*_k}{\mu^*_k\,R} \;+\; O(R^{-2}) \qquad \text{as } R \to \infty.εk​(R)≤μk∗​R2Nρk∗​​+O(R−2)as R→∞.

In particular every εk(R)→0\varepsilon_k(R) \to 0εk​(R)→0 as R→∞R \to \inftyR→∞. The bound is sharp on the symmetric saddle L=x2−y2L = x^2 - y^2L=x2−y2 in N=2N = 2N=2, where ρ∗=μ∗=2\rho^* = \mu^* = 2ρ∗=μ∗=2 and ε(R)=4/R\varepsilon(R) = 4/Rε(R)=4/R.

The theorem is essentially the Gershgorin construction of Appendix C.2 wired through a Frobenius-budget bookkeeping and a Taylor expansion at the critical point.

Proof. Apply the diagonal-only choice σii=−2M/ℓi\sigma_{ii} = -2 M/\ell_iσii​=−2M/ℓi​ from Appendix C.2. The resulting correction is C=−M IC = -M\,IC=−MI, so H−C=H+M IH - C = H + M\,IH−C=H+MI has the same off-diagonals as HHH and diagonal entries Hii+MH_{ii} + MHii​+M. By Gershgorin's disc theorem, H−C≻0H - C \succ 0H−C≻0 whenever M>ρ(θ)M > \rho(\theta)M>ρ(θ), where

ρ(θ):=max⁡i[∑j≠i∣Hij(θ)∣−Hii(θ)]\rho(\theta) := \max_i\Bigl[\sum_{j \ne i} |H_{ij}(\theta)| - H_{ii}(\theta)\Bigr]ρ(θ):=imax​[j=i∑​∣Hij​(θ)∣−Hii​(θ)]

is the local Gershgorin deficit. The Frobenius cost of this choice is

∥σ∥F=2M(∑iℓi−2)1/2,\|\sigma\|_F = 2 M\Bigl(\sum_i \ell_i^{-2}\Bigr)^{1/2},∥σ∥F​=2M(i∑​ℓi−2​)1/2,

so the budget constraint ∥σ∥F≤R\|\sigma\|_F \le R∥σ∥F​≤R caps M≤R/(2∑iℓi−2)M \le R/(2\sqrt{\sum_i \ell_i^{-2}})M≤R/(2∑i​ℓi−2​​). The diagonal construction succeeds when both inequalities can be simultaneously satisfied, equivalently when

(∑iℓi−2(θ))1/2  <  R2 ρ(θ).(6.3.1)\Bigl(\sum_i \ell_i^{-2}(\theta)\Bigr)^{1/2} \;<\; \frac{R}{2\,\rho(\theta)}. \tag{6.3.1}(i∑​ℓi−2​(θ))1/2<2ρ(θ)R​.(6.3.1)

This is the budget-translated sign-flip criterion.

Near θk∗\theta^*_kθk∗​, change coordinates so Hk∗:=H(θk∗)=diag(λ1,…,λN)H^*_k := H(\theta^*_k) = \mathrm{diag}(\lambda_1, \ldots, \lambda_N)Hk∗​:=H(θk∗​)=diag(λ1​,…,λN​) with all λi≠0\lambda_i \ne 0λi​=0 (Morse). Let δ:=θ−θk∗\delta := \theta - \theta^*_kδ:=θ−θk∗​, r:=∥δ∥r := \|\delta\|r:=∥δ∥, and u^:=δ/r\hat u := \delta/ru^:=δ/r. Taylor expansion gives

ℓi(θ)=λiuir+O(r2),ρ(θ)=ρk∗+O(r),\ell_i(\theta) = \lambda_i u_i r + O(r^2), \qquad \rho(\theta) = \rho^*_k + O(r),ℓi​(θ)=λi​ui​r+O(r2),ρ(θ)=ρk∗​+O(r),

so ∑iℓi−2(θ)=r−2 ∑i(λiui)−2(1+O(r))\sum_i \ell_i^{-2}(\theta) = r^{-2}\,\sum_i (\lambda_i u_i)^{-2}(1 + O(r))∑i​ℓi−2​(θ)=r−2∑i​(λi​ui​)−2(1+O(r)). Substituting into (6.3.1) gives the per-direction bound

r  >  ε(R,u^)  :=  2ρk∗R ∑i(λiui)−2  +  O(R−2).(6.3.2)r \;>\; \varepsilon(R, \hat u) \;:=\; \frac{2 \rho^*_k}{R}\,\sqrt{\sum_i (\lambda_i u_i)^{-2}} \;+\; O(R^{-2}). \tag{6.3.2}r>ε(R,u^):=R2ρk∗​​i∑​(λi​ui​)−2​+O(R−2).(6.3.2)

The directional factor ∑i(λiui)−2\sqrt{\sum_i (\lambda_i u_i)^{-2}}∑i​(λi​ui​)−2​ diverges only on the coordinate-axis null-cones (when some ui→0u_i \to 0ui​→0), precisely the directions where the diagonal-only construction needs the cross-term lift. On those directions, replace the divergent self-derivative σii=−2M/ℓi\sigma_{ii} = -2 M/\ell_iσii​=−2M/ℓi​ with a cross-derivative σk(i), i=2M/ℓk(i)\sigma_{k(i),\,i} = 2 M/\ell_{k(i)}σk(i),i​=2M/ℓk(i)​ for a donor coordinate k(i)k(i)k(i) with ℓk(i)≠0\ell_{k(i)} \ne 0ℓk(i)​=0. The lifted construction (Section 4.7 / Appendix C.2 extension) achieves the same shift C=−M IC = -M\,IC=−MI and the lifted directional factor satisfies the uniform bound

∑i(λiui)lifted−2  ≤  Nμk∗\sqrt{\sum_i (\lambda_i u_i)^{-2}_{\text{lifted}}} \;\le\; \frac{N}{\mu^*_k}i∑​(λi​ui​)lifted−2​​≤μk∗​N​

since max⁡iui2≥1/N\max_i u_i^2 \ge 1/Nmaxi​ui2​≥1/N and ∣λi∣≥μk∗|\lambda_i| \ge \mu^*_k∣λi​∣≥μk∗​. Substituting into (6.3.2) yields εk(R)≤2Nρk∗/(μk∗ R)+O(R−2)\varepsilon_k(R) \le 2 N \rho^*_k/(\mu^*_k\,R) + O(R^{-2})εk​(R)≤2Nρk∗​/(μk∗​R)+O(R−2). □\square□

What the theorem buys

Three takeaways.

  1. The obstruction is a fence, not a wall. Larger budget shrinks the fence linearly. Doubling RRR halves the failure tube around each saddle. The R→∞R \to \inftyR→∞ limit recovers Section 6.1 as a measure-zero residue.

  2. The bound is structural, not directly computable on a real network. The constants ρk∗\rho^*_kρk∗​ and μk∗\mu^*_kμk∗​ are local to each critical point, and a real network has unknown saddle locations. What the theorem gives is the shape of the failure region (tubes of radius O(1/R)O(1/R)O(1/R) around the critical set), not its absolute size at a specific training run. Restricting to a bounded sublevel set {L≤L0}\{L \le L_0\}{L≤L0​} is needed to control the O(R−2)O(R^{-2})O(R−2) remainder (which absorbs a ∥D3L∥L∞\|D^3 L\|_{L^\infty}∥D3L∥L∞​ factor from the Taylor expansion).

  3. The theorem is about attainable basin coverage, not about online tracking. It says a good σ\sigmaσ exists outside the ε\varepsilonε-tube. Whether the EMA-based online estimator finds that σ\sigmaσ in finite time is a separate question that lives in Section 8.1.

The Morse assumption is load-bearing. Degenerate critical points (e.g., the monkey saddle L=x3−3xy2L = x^3 - 3 x y^2L=x3−3xy2, where H∗=0H^* = 0H∗=0 so μk∗=0\mu^*_k = 0μk∗​=0) make the bound diverge. The right ε\varepsilonε-scaling at a degenerate critical point uses the leading nonzero coefficient of the Taylor expansion of LLL and gives a slower decay than 1/R1/R1/R; that case is not treated here.

Numerical verification on the symmetric saddle, the asymmetric saddle L=x2−3y2L = x^2 - 3 y^2L=x2−3y2, and the quartic well of Section 7.1 is in notebooks/epsilon-excised-sign-flip.nb. On the symmetric saddle the predicted ε(R)=4/R\varepsilon(R) = 4/Rε(R)=4/R matches the observed failure radius along the diagonal direction to scan resolution (table reproduced below).

RRR4/R4/R4/R (predicted)observed ε\varepsilonε
2222.0002.0002.0002.0012.0012.001
5550.8000.8000.8000.8010.8010.801
1010100.4000.4000.4000.4010.4010.401
5050500.0800.0800.0800.0810.0810.081
1001001000.0400.0400.0400.0410.0410.041

7. Basin enlargement

Around a strict local minimum θ∗\theta^*θ∗ with H(θ∗)≻0H(\theta^*) \succ 0H(θ∗)≻0, define two basins:

Beucl={θ:H(θ)≻0},Bgopt(R)={θ:∃ σ with ∥σ∥F≤R and H(θ)−C(σ,ℓ(θ))≻0}.\mathcal{B}_{\mathrm{eucl}} = \{\theta : H(\theta) \succ 0\}, \qquad \mathcal{B}_g^{\mathrm{opt}}(R) = \{\theta : \exists\,\sigma\text{ with }\|\sigma\|_F \le R\text{ and } H(\theta) - C(\sigma, \ell(\theta)) \succ 0\}.Beucl​={θ:H(θ)≻0},Bgopt​(R)={θ:∃σ with ∥σ∥F​≤R and H(θ)−C(σ,ℓ(θ))≻0}.

Near the minimum, ℓ≈0\ell \approx 0ℓ≈0, so C≈0C \approx 0C≈0 for any σ\sigmaσ in the budget. In particular σ=0\sigma = 0σ=0 is feasible and gives H−C=H≻0H - C = H \succ 0H−C=H≻0. Therefore Beucl⊆Bgopt(R)\mathcal{B}_{\mathrm{eucl}} \subseteq \mathcal{B}_g^{\mathrm{opt}}(R)Beucl​⊆Bgopt​(R) for every R≥0R \ge 0R≥0: any point at which HHH is already positive-definite is in the attainable geodesic-convexity basin. This inclusion is about the attainable geodesic basin under optimal choice of σ\sigmaσ, not about the basin produced by any particular online rule (an online rule may apply σ≠0\sigma \ne 0σ=0 in the Euclidean basin and inadvertently push H−CH - CH−C out of positive-definiteness). The interesting region is outside Beucl\mathcal{B}_{\mathrm{eucl}}Beucl​, where HHH has negative eigenvalues but ℓ\ellℓ is non-zero. There CCC can compensate.

7.1 The quartic well

A vivid example. L=x4+y4−2x2−2y2L = x^4 + y^4 - 2x^2 - 2y^2L=x4+y4−2x2−2y2 has four minima at (±1,±1)(\pm 1, \pm 1)(±1,±1), saddle points at the origin and along the axes, and Hessian H=diag(12x2−4,12y2−4)H = \mathrm{diag}(12x^2 - 4, 12y^2 - 4)H=diag(12x2−4,12y2−4). The Euclidean convexity basin around each minimum extends only to ∣x∣,∣y∣>1/3≈0.58|x|, |y| > 1/\sqrt{3} \approx 0.58∣x∣,∣y∣>1/3​≈0.58. The four basins are disconnected square regions.

With the diagonal learnable metric at correction budget R=5R = 5R=5, the geodesic convexity basin extends almost to the axes, covering most of the domain. The four disconnected Euclidean basins merge into a single large basin that nearly fills the plane. Only small pockets of non-convexity remain near the coordinate axes, consistent with the critical-point obstruction.

The metric is redrawing the map. An optimizer using this metric can start much further from the minimum and still see a convex-looking landscape all the way down. The boundary of Bg\mathcal{B}_gBg​ forms a thin layer around each saddle, with thickness O(1/R)O(1/R)O(1/R) in the correction budget (formalised in Theorem 1ʹ). Larger budgets push the boundary closer to the saddle but never reach it.

Euclidean basins (left, four disconnected square regions, one per minimum) and the geodesic basin \mathcal{B}_g on the quartic well at R = 5 (right, single connected basin nearly filling the plane).

8. Trajectory-level analysis

The pointwise theory above asks whether at a given θ\thetaθ the metric can achieve geodesic convexity. An optimizer, however, moves. Three questions extend the theory to trajectories.

8.1 Does the optimal metric vary smoothly?

At each θ\thetaθ, the optimal σ∗\sigma^*σ∗ solves

σ∗(θ)=arg⁡max⁡∥σ∥F≤Rλmin⁡(H(θ)−C(σ,ℓ(θ))).\sigma^*(\theta) = \arg\max_{\|\sigma\|_F \le R} \lambda_{\min}\bigl(H(\theta) - C(\sigma, \ell(\theta))\bigr).σ∗(θ)=arg∥σ∥F​≤Rmax​λmin​(H(θ)−C(σ,ℓ(θ))).

This is the same concave semidefinite program as in Section 4.4. Solutions to smoothly parameterised concave SDPs vary upper-hemicontinuously with the parameters; full smoothness can fail not only at eigenvalue-degeneracy points but also when the active set on the norm-ball boundary switches or the optimum set is non-singleton. Adding a small Tikhonov regulariser to the objective (e.g., +ϵ∥σ∥F2+ \epsilon\|\sigma\|_F^2+ϵ∥σ∥F2​) restores smoothness by making the optimum unique. An exponential moving average (EMA)-smoothed online estimator of σ∗\sigma^*σ∗ can track the underlying σ∗\sigma^*σ∗ to within the smoothing window, but with no guarantee of pointwise tracking through active-set transitions.

8.2 Are the coupled dynamics stable?

At a strict local minimum θ∗\theta^*θ∗, the θ\thetaθ dynamics reduce to standard gradient descent (the denominator factor tends to 1), and the sss dynamics relax to s=0s = 0s=0 (the identity metric) under regularisation. The Jacobian of the joint update at (θ∗,s=0)(\theta^*, s = 0)(θ∗,s=0) is block-triangular, because the off-diagonal blocks vanish (ℓ=0\ell = 0ℓ=0 kills all cross-terms). Eigenvalues of a block-triangular matrix are the eigenvalues of the diagonal blocks. Both blocks contract under standard learning-rate conditions, giving local asymptotic stability of the joint system without any requirement that the metric converge faster than the parameters.

8.3 Where the corrections operate

Near a minimum, corrections are unnecessary because H≻0H \succ 0H≻0 already; the metric correctly relaxes to identity and composition is trivially coherent.

In the non-convex bulk, the gradient is large, CCC is strong, regularisation bounds the tracking error, and the denominator damps any aggressive deviation. Corrections compose coherently.

Near a saddle, the gradient is small, the correction weakens by the critical-point obstruction, and the geodesic convexity margin shrinks toward zero. The mechanism hands off to stochastic gradient noise for saddle escape.

The curvature correction is a bulk phenomenon, not a saddle-escape mechanism. It complements stochastic noise rather than replacing it: noise provides the perturbation to leave the saddle's neighbourhood, the metric provides the geometric correction that makes the surrounding bulk look convex.

9. The Gauss equation connection

The loss graph ϕ(θ)=(θ,L(θ))\phi(\theta) = (\theta, L(\theta))ϕ(θ)=(θ,L(θ)) is a submanifold of RN+1\mathbb{R}^{N+1}RN+1. As a submanifold it has two kinds of curvature: intrinsic curvature, detectable by inhabitants living on the surface, captured by the Riemann tensor of the induced metric; and extrinsic curvature, describing how the surface bends in the ambient space, captured by the second fundamental form.

For the loss embedding, the second fundamental form is proportional to HHH:

bij=ξ1+ξ∥ℓ∥γ2 Hij.b_{ij} = \sqrt{\frac{\xi}{1 + \xi \|\ell\|^2_\gamma}}\,H_{ij}.bij​=1+ξ∥ℓ∥γ2​ξ​​Hij​.

The Hessian of the loss is the shape operator of the loss surface. Via the Gauss equation, the intrinsic Riemann curvature encodes Hessian information through the second fundamental form. In two dimensions, the Gaussian curvature is

K=ξdet⁡(H)det⁡(γ) (1+ξ∥ℓ∥γ2)2.K = \frac{\xi \det(H)}{\det(\gamma)\,(1 + \xi \|\ell\|^2_\gamma)^2}.K=det(γ)(1+ξ∥ℓ∥γ2​)2ξdet(H)​.

Three cases: K>0K > 0K>0 if the Hessian has same-sign eigenvalues (bowl or dome), K<0K < 0K<0 if HHH is indefinite (saddle), K=0K = 0K=0 if HHH is singular. The qualitative structure of the Euclidean Hessian is faithfully reflected in the intrinsic geometry of the loss surface as a submanifold of the ambient space. This is why the framework feels canonical: the induced metric is not an arbitrary preconditioner choice but the natural geometric structure of the loss graph.

10. Empirical: the curvature mechanism is practically inactive (decoupled metrics)

The theory above says the diagonal metric can convert saddles into convex bulk when the curvature correction CCC is engaged. The natural way to engage it explicitly is to split the single metric vector into two, decoupling the role that sets the parameter step from the role that drives the correction:

  • sipreconds^{\mathrm{precond}}_isiprecond​ controls the parameter update (escape-focused),
  • sicurvs^{\mathrm{curv}}_isicurv​ controls the curvature correction CCC (convexity-focused).

The hope was that both could be pursued at once: saddle escape and geodesic convexity. The probe (saddle and weighted-saddle losses; raw runs in results/saddle/, results/weighted_saddle/, and results/decoupled_metrics_probe.json) showed no benefit. The decoupled variant produced essentially identical trajectories to the escape-only variant, for three reasons.

  1. The base gradient-magnitude drive (ξ esi gi2\xi\,e^{s_i}\,g_i^2ξesi​gi2​) dominates the curvature correction (βc sign(Hii) esi gi2\beta_c\,\mathrm{sign}(H_{ii})\,e^{s_i}\,g_i^2βc​sign(Hii​)esi​gi2​) when βc≪ξ\beta_c \ll \xiβc​≪ξ.
  2. Both spreconds^{\mathrm{precond}}sprecond and scurvs^{\mathrm{curv}}scurv saturate at the metric clip boundary (±4\pm 4±4) and converge to the same values.
  3. On the x2−y2x^2 - y^2x2−y2 saddle the two metrics stayed within ≈0.4{\approx}0.4≈0.4 of each other through step ∼440{\sim}440∼440 and diverged only in a brief transient (peak ≈4.9{\approx}4.9≈4.9 at step 480480480) before both collapsed to the same clipped boundary by step 500500500. The escape behaviour was identical to the escape-only variant (escape step 794794794 for both, identical final loss −6.069-6.069−6.069).

The fundamental issue is timescale separation: the two drives operate at timescales ∼1/(μξ){\sim}1/(\mu\xi)∼1/(μξ) and ∼1/(μβc){\sim}1/(\mu\beta_c)∼1/(μβc​). With βc/ξ≈0.1\beta_c/\xi \approx 0.1βc​/ξ≈0.1 the correction is a 10% perturbation on the base drive, and the metric clip forces both to the same saturated state. For decoupled metrics to engage, the correction must be comparable in magnitude to the base drive (βc∼ξ\beta_c \sim \xiβc​∼ξ), or the metric clip must be much larger, or the two metrics must run on genuinely different timescales — none of which were practical in the swept configurations.

This is the honest empirical status of the curvature mechanism. Theoretically the diagonal class flips eigenvalue signs in the bulk (Sections 4-7); at standard hyperparameter scales the correction term is swamped by the gradient-magnitude drive, and the mechanism is practically inactive as an independently controllable knob. It remains consistent with the implicit-sharpness behaviour documented in the learnable diagonal entry §9.2, which is the more durable practical signature of the same geometry.

Appendix A: Derivation of the Riemannian Hessian formula

The general Riemannian Hessian of a scalar function LLL in coordinates is the classical expression

(HessgL)ij=∂i∂jL−Γijk(g) ∂kL=Hij−Γijk(g) ℓk,(\mathrm{Hess}_g L)_{ij} = \partial_i \partial_j L - \Gamma^k_{ij}(g)\,\partial_k L = H_{ij} - \Gamma^k_{ij}(g)\,\ell_k,(Hessg​L)ij​=∂i​∂j​L−Γijk​(g)∂k​L=Hij​−Γijk​(g)ℓk​,

with ℓi=∂iL\ell_i = \partial_i Lℓi​=∂i​L and Hij=∂i∂jLH_{ij} = \partial_i \partial_j LHij​=∂i​∂j​L. The first term is the Euclidean Hessian; the second is the correction that accounts for the curvature of the coordinate system through the Christoffel symbols Γijk(g)\Gamma^k_{ij}(g)Γijk​(g), which encode how the basis vectors twist as you move on the manifold. The Christoffel symbols of a metric ggg are themselves built from a quantity we call the Christoffel bracket,

Bij,l(g):=∂igjl+∂jgil−∂lgij,Γijk(g)=12 gkl Bij,l(g).B_{ij,l}(g) := \partial_i g_{jl} + \partial_j g_{il} - \partial_l g_{ij}, \qquad \Gamma^k_{ij}(g) = \tfrac{1}{2}\,g^{kl}\,B_{ij,l}(g).Bij,l​(g):=∂i​gjl​+∂j​gil​−∂l​gij​,Γijk​(g)=21​gklBij,l​(g).

The bracket Bij,lB_{ij,l}Bij,l​ is a particular symmetric combination of first derivatives of ggg. The full Christoffel symbol is obtained by contracting it with the inverse metric. The bracket is the "raw" object: it captures all the information about how the metric varies in space, and the contraction with gklg^{kl}gkl is the rescaling needed to make Γijk\Gamma^k_{ij}Γijk​ transform correctly as a connection coefficient.

Applying the Riemannian Hessian formula to the induced metric gij=γij(θ)+ξ ℓi ℓjg_{ij} = \gamma_{ij}(\theta) + \xi\,\ell_i\,\ell_jgij​=γij​(θ)+ξℓi​ℓj​ requires computing Γijk(g)\Gamma^k_{ij}(g)Γijk​(g). The clean factorisation

(HessgL)ij=Hij−Cij1+ξ ∥ℓ∥γ2,Cij=Γijk(γ) ℓk(\mathrm{Hess}_g L)_{ij} = \frac{H_{ij} - C_{ij}}{1 + \xi\,\|\ell\|_\gamma^2}, \qquad C_{ij} = \Gamma^k_{ij}(\gamma)\,\ell_k(Hessg​L)ij​=1+ξ∥ℓ∥γ2​Hij​−Cij​​,Cij​=Γijk​(γ)ℓk​

is not obvious from the formula for Γijk(g)\Gamma^k_{ij}(g)Γijk​(g). Both the γ\gammaγ-part and the ξ\xiξ-part of the metric contribute to the bracket, and they combine through a Sherman-Morrison correction in the inverse metric. The derivation has two steps.

A.1 The Christoffel bracket splits cleanly

Split gijg_{ij}gij​ as γij+ξ ℓi ℓj\gamma_{ij} + \xi\,\ell_i\,\ell_jγij​+ξℓi​ℓj​ and use ∂iℓj=Hij\partial_i \ell_j = H_{ij}∂i​ℓj​=Hij​,

∂igjl=∂iγjl+ξ (Hij ℓl+ℓj Hil),\partial_i g_{jl} = \partial_i \gamma_{jl} + \xi\,(H_{ij}\,\ell_l + \ell_j\,H_{il}),∂i​gjl​=∂i​γjl​+ξ(Hij​ℓl​+ℓj​Hil​), ∂jgil=∂jγil+ξ (Hij ℓl+ℓi Hjl),\partial_j g_{il} = \partial_j \gamma_{il} + \xi\,(H_{ij}\,\ell_l + \ell_i\,H_{jl}),∂j​gil​=∂j​γil​+ξ(Hij​ℓl​+ℓi​Hjl​), ∂lgij=∂lγij+ξ (Hil ℓj+ℓi Hjl).\partial_l g_{ij} = \partial_l \gamma_{ij} + \xi\,(H_{il}\,\ell_j + \ell_i\,H_{jl}).∂l​gij​=∂l​γij​+ξ(Hil​ℓj​+ℓi​Hjl​).

Add the first two and subtract the third. The γ\gammaγ terms reproduce the bracket of γ\gammaγ alone,

Bij,l(γ):=∂iγjl+∂jγil−∂lγij.B_{ij,l}(\gamma) := \partial_i \gamma_{jl} + \partial_j \gamma_{il} - \partial_l \gamma_{ij}.Bij,l​(γ):=∂i​γjl​+∂j​γil​−∂l​γij​.

The ξ\xiξ terms simplify by pairwise cancellation. The term +ℓjHil+\ell_j H_{il}+ℓj​Hil​ from ∂igjl\partial_i g_{jl}∂i​gjl​ cancels −Hilℓj-H_{il}\ell_j−Hil​ℓj​ from −∂lgij-\partial_l g_{ij}−∂l​gij​, and +ℓiHjl+\ell_i H_{jl}+ℓi​Hjl​ from ∂jgil\partial_j g_{il}∂j​gil​ cancels −ℓiHjl-\ell_i H_{jl}−ℓi​Hjl​ from −∂lgij-\partial_l g_{ij}−∂l​gij​. What survives is 2ξ Hij ℓl2\xi\,H_{ij}\,\ell_l2ξHij​ℓl​. Therefore

Bij,l(g)=Bij,l(γ)+2ξ Hij ℓl.(A.1)B_{ij,l}(g) = B_{ij,l}(\gamma) + 2\xi\,H_{ij}\,\ell_l. \tag{A.1}Bij,l​(g)=Bij,l​(γ)+2ξHij​ℓl​.(A.1)

The bracket of ggg separates cleanly into a γ\gammaγ-part and an HHH-part with no cross terms. This is the first non-trivial fact.

The reason this works is the special structure of the ξ\xiξ-piece of ggg. The term ξℓiℓj\xi \ell_i \ell_jξℓi​ℓj​ depends on θ\thetaθ only through ℓ\ellℓ, and ∂iℓj=Hij\partial_i \ell_j = H_{ij}∂i​ℓj​=Hij​. So all the derivative terms that arise involve HijH_{ij}Hij​, and the cyclic combination forces them to combine into the simple product HijℓlH_{ij} \ell_lHij​ℓl​. A different non-loss-based embedding could give a more tangled bracket.

A.2 Contracting with the inverse metric and the gradient

The Christoffel symbol of ggg is Γijk(g)=12 gkl Bij,l(g)\Gamma^k_{ij}(g) = \tfrac{1}{2}\,g^{kl}\,B_{ij,l}(g)Γijk​(g)=21​gklBij,l​(g), and we need its contraction with ℓk\ell_kℓk​,

Γijk(g) ℓk=12 gkl Bij,l(g) ℓk.\Gamma^k_{ij}(g)\,\ell_k = \tfrac{1}{2}\,g^{kl}\,B_{ij,l}(g)\,\ell_k.Γijk​(g)ℓk​=21​gklBij,l​(g)ℓk​.

The inverse metric comes from Sherman-Morrison applied to g=γ+ξ ℓ ℓ⊤g = \gamma + \xi\,\ell\,\ell^\topg=γ+ξℓℓ⊤. This is applicable because ggg is a rank-1 perturbation of γ\gammaγ:

gkl=γkl−ξ (γ−1ℓ)k (γ−1ℓ)l1+ξ ∥ℓ∥γ2,∥ℓ∥γ2:=ℓ⊤γ−1ℓ.(A.2)g^{kl} = \gamma^{kl} - \frac{\xi\,(\gamma^{-1}\ell)^k\,(\gamma^{-1}\ell)^l}{1 + \xi\,\|\ell\|_\gamma^2}, \qquad \|\ell\|_\gamma^2 := \ell^\top \gamma^{-1} \ell. \tag{A.2}gkl=γkl−1+ξ∥ℓ∥γ2​ξ(γ−1ℓ)k(γ−1ℓ)l​,∥ℓ∥γ2​:=ℓ⊤γ−1ℓ.(A.2)

We will need one intermediate identity. Contracting gklg^{kl}gkl twice with ℓ\ellℓ,

ℓ⊤g−1ℓ=∥ℓ∥γ2−ξ ∥ℓ∥γ41+ξ ∥ℓ∥γ2=∥ℓ∥γ21+ξ ∥ℓ∥γ2.(A.3)\ell^\top g^{-1} \ell = \|\ell\|_\gamma^2 - \frac{\xi\,\|\ell\|_\gamma^4}{1 + \xi\,\|\ell\|_\gamma^2} = \frac{\|\ell\|_\gamma^2}{1 + \xi\,\|\ell\|_\gamma^2}. \tag{A.3}ℓ⊤g−1ℓ=∥ℓ∥γ2​−1+ξ∥ℓ∥γ2​ξ∥ℓ∥γ4​​=1+ξ∥ℓ∥γ2​∥ℓ∥γ2​​.(A.3)

Now contract (A.1) with 12 gkl ℓk\tfrac{1}{2}\,g^{kl}\,\ell_k21​gklℓk​ piece by piece.

From the ξ\xiξ-part of (A.1), using (A.3),

12 gkl (2ξ Hij ℓl) ℓk=ξ Hij (gkl ℓk ℓl)=ξ ∥ℓ∥γ21+ξ ∥ℓ∥γ2 Hij.(A.4)\tfrac{1}{2}\,g^{kl}\,(2\xi\,H_{ij}\,\ell_l)\,\ell_k = \xi\,H_{ij}\,(g^{kl}\,\ell_k\,\ell_l) = \frac{\xi\,\|\ell\|_\gamma^2}{1 + \xi\,\|\ell\|_\gamma^2}\,H_{ij}. \tag{A.4}21​gkl(2ξHij​ℓl​)ℓk​=ξHij​(gklℓk​ℓl​)=1+ξ∥ℓ∥γ2​ξ∥ℓ∥γ2​​Hij​.(A.4)

From the γ\gammaγ-part of (A.1), plug (A.2) into 12 gkl Bij,l(γ) ℓk\tfrac{1}{2}\,g^{kl}\,B_{ij,l}(\gamma)\,\ell_k21​gklBij,l​(γ)ℓk​,

12 γkl Bij,l(γ) ℓk−12⋅ξ1+ξ ∥ℓ∥γ2 (γ−1ℓ)k ℓk (γ−1ℓ)l Bij,l(γ).\tfrac{1}{2}\,\gamma^{kl}\,B_{ij,l}(\gamma)\,\ell_k - \tfrac{1}{2}\cdot \frac{\xi}{1 + \xi\,\|\ell\|_\gamma^2}\,(\gamma^{-1}\ell)^k\,\ell_k\,(\gamma^{-1}\ell)^l\,B_{ij,l}(\gamma).21​γklBij,l​(γ)ℓk​−21​⋅1+ξ∥ℓ∥γ2​ξ​(γ−1ℓ)kℓk​(γ−1ℓ)lBij,l​(γ).

The first piece is exactly Γijk(γ) ℓk=Cij\Gamma^k_{ij}(\gamma)\,\ell_k = C_{ij}Γijk​(γ)ℓk​=Cij​. In the second piece, (γ−1ℓ)k ℓk=∥ℓ∥γ2(\gamma^{-1}\ell)^k\,\ell_k = \|\ell\|_\gamma^2(γ−1ℓ)kℓk​=∥ℓ∥γ2​, and (γ−1ℓ)l Bij,l(γ)=γkl ℓk Bij,l(γ)=2 Cij(\gamma^{-1}\ell)^l\,B_{ij,l}(\gamma) = \gamma^{kl}\,\ell_k\,B_{ij,l}(\gamma) = 2\,C_{ij}(γ−1ℓ)lBij,l​(γ)=γklℓk​Bij,l​(γ)=2Cij​ (twice the definition of CCC, because the factor of 1/21/21/2 is absorbed). Combining,

Cij−ξ ∥ℓ∥γ21+ξ ∥ℓ∥γ2 Cij=Cij1+ξ ∥ℓ∥γ2.(A.5)C_{ij} - \frac{\xi\,\|\ell\|_\gamma^2}{1 + \xi\,\|\ell\|_\gamma^2}\,C_{ij} = \frac{C_{ij}}{1 + \xi\,\|\ell\|_\gamma^2}. \tag{A.5}Cij​−1+ξ∥ℓ∥γ2​ξ∥ℓ∥γ2​​Cij​=1+ξ∥ℓ∥γ2​Cij​​.(A.5)

Adding (A.4) and (A.5),

Γijk(g) ℓk=Cij1+ξ ∥ℓ∥γ2+ξ ∥ℓ∥γ21+ξ ∥ℓ∥γ2 Hij.\Gamma^k_{ij}(g)\,\ell_k = \frac{C_{ij}}{1 + \xi\,\|\ell\|_\gamma^2} + \frac{\xi\,\|\ell\|_\gamma^2}{1 + \xi\,\|\ell\|_\gamma^2}\,H_{ij}.Γijk​(g)ℓk​=1+ξ∥ℓ∥γ2​Cij​​+1+ξ∥ℓ∥γ2​ξ∥ℓ∥γ2​​Hij​.

Substituting into the Riemannian Hessian formula,

(HessgL)ij=Hij−Γijk(g) ℓk=Hij (1−ξ ∥ℓ∥γ21+ξ ∥ℓ∥γ2)−Cij1+ξ ∥ℓ∥γ2.(\mathrm{Hess}_g L)_{ij} = H_{ij} - \Gamma^k_{ij}(g)\,\ell_k = H_{ij}\,\bigl(1 - \tfrac{\xi\,\|\ell\|_\gamma^2}{1 + \xi\,\|\ell\|_\gamma^2}\bigr) - \frac{C_{ij}}{1 + \xi\,\|\ell\|_\gamma^2}.(Hessg​L)ij​=Hij​−Γijk​(g)ℓk​=Hij​(1−1+ξ∥ℓ∥γ2​ξ∥ℓ∥γ2​​)−1+ξ∥ℓ∥γ2​Cij​​.

The coefficient on HijH_{ij}Hij​ simplifies to 1/(1+ξ ∥ℓ∥γ2)1/(1 + \xi\,\|\ell\|_\gamma^2)1/(1+ξ∥ℓ∥γ2​), giving the boxed result

(HessgL)ij=Hij−Cij1+ξ ∥ℓ∥γ2.(\mathrm{Hess}_g L)_{ij} = \frac{H_{ij} - C_{ij}}{1 + \xi\,\|\ell\|_\gamma^2}.(Hessg​L)ij​=1+ξ∥ℓ∥γ2​Hij​−Cij​​.

In the fixed-γ\gammaγ case, Bij,l(γ)=0B_{ij,l}(\gamma) = 0Bij,l​(γ)=0 and so Cij=0C_{ij} = 0Cij​=0, recovering (HessgL)ij=Hij/(1+ξ∥ℓ∥γ2)(\mathrm{Hess}_g L)_{ij} = H_{ij}/(1 + \xi\|\ell\|^2_\gamma)(Hessg​L)ij​=Hij​/(1+ξ∥ℓ∥γ2​), a positive scalar times the Euclidean Hessian. This is the Level 0 result.

The single factor 1/(1+ξ∥ℓ∥γ2)1/(1 + \xi\|\ell\|_\gamma^2)1/(1+ξ∥ℓ∥γ2​) in the denominator comes from Sherman-Morrison; it is the same scalar that appears in the smooth-clipping update. The numerator splits cleanly because the ξ\xiξ-bracket has the special form Hij ℓlH_{ij}\,\ell_lHij​ℓl​, which is what allows the contraction with gkl ℓkg^{kl}\,\ell_kgklℓk​ to factor through the scalar ℓ⊤g−1ℓ\ell^\top g^{-1} \ellℓ⊤g−1ℓ. The γ\gammaγ-bracket contributes the new piece CijC_{ij}Cij​ that depends on ∇γ\nabla\gamma∇γ but not on HHH. That is the door to curvature correction.

Appendix B: Fixed off-diagonal extensions

Some variants of the induced metric add constant off-diagonal coupling vectors a,b∈RNa, b \in \mathbb{R}^Na,b∈RN to the ambient metric, giving the pullback

gij=γij+ξ ℓi ℓj+ξ (ai ℓj+ℓi bj),g_{ij} = \gamma_{ij} + \xi\,\ell_i\,\ell_j + \xi\,(a_i\,\ell_j + \ell_i\,b_j),gij​=γij​+ξℓi​ℓj​+ξ(ai​ℓj​+ℓi​bj​),

with γ,a,b\gamma, a, bγ,a,b all constant. This appendix carefully computes the bracket and shows that the sign-preservation conclusion holds only in the symmetric subset a=ba = ba=b. In the asymmetric case a≠ba \ne ba=b, an additional non-scalar correction appears, and the Riemannian Hessian is no longer forced to inherit the eigenstructure of HHH.

For the calculation to make sense as a Riemannian construction, ggg must be symmetric. Direct check: gij−gji=ξ[(ai−bi)ℓj−(aj−bj)ℓi]g_{ij} - g_{ji} = \xi[(a_i - b_i)\ell_j - (a_j - b_j)\ell_i]gij​−gji​=ξ[(ai​−bi​)ℓj​−(aj​−bj​)ℓi​], which is nonzero whenever a−ba - ba−b is not parallel to ℓ\ellℓ. The bracket-based machinery below applies only when a=ba = ba=b (or, more loosely, when the symmetric part of ggg is used). The non-symmetric case is treated in the off-diagonal metric entry, where it is acknowledged as a non-Riemannian preconditioner rather than a metric.

Bracket derivation for a=ba = ba=b. Using ∂iℓj=Hij\partial_i \ell_j = H_{ij}∂i​ℓj​=Hij​ and the constancy of a,ba, ba,b,

∂i(aj ℓl+ℓj bl)=aj Hil+Hij bl,\partial_i (a_j\,\ell_l + \ell_j\,b_l) = a_j\,H_{il} + H_{ij}\,b_l,∂i​(aj​ℓl​+ℓj​bl​)=aj​Hil​+Hij​bl​, ∂j(ai ℓl+ℓi bl)=ai Hjl+Hij bl,\partial_j (a_i\,\ell_l + \ell_i\,b_l) = a_i\,H_{jl} + H_{ij}\,b_l,∂j​(ai​ℓl​+ℓi​bl​)=ai​Hjl​+Hij​bl​, ∂l(ai ℓj+ℓi bj)=ai Hjl+Hil bj.\partial_l (a_i\,\ell_j + \ell_i\,b_j) = a_i\,H_{jl} + H_{il}\,b_j.∂l​(ai​ℓj​+ℓi​bj​)=ai​Hjl​+Hil​bj​.

The cyclic sum ∂igjl+∂jgil−∂lgij\partial_i g_{jl} + \partial_j g_{il} - \partial_l g_{ij}∂i​gjl​+∂j​gil​−∂l​gij​ from the new term is

ξ[ajHil+2Hijbl−Hilbj]=2ξ Hij bl+ξ Hil (aj−bj).\xi\bigl[a_j H_{il} + 2 H_{ij} b_l - H_{il} b_j\bigr] = 2\xi\,H_{ij}\,b_l + \xi\,H_{il}\,(a_j - b_j).ξ[aj​Hil​+2Hij​bl​−Hil​bj​]=2ξHij​bl​+ξHil​(aj​−bj​).

With a=ba = ba=b, the second term vanishes and the total bracket reduces to

Bij,l(g)=Bij,l(γ)+2ξ Hij (ℓl+bl).B_{ij,l}(g) = B_{ij,l}(\gamma) + 2\xi\,H_{ij}\,(\ell_l + b_l).Bij,l​(g)=Bij,l​(γ)+2ξHij​(ℓl​+bl​).

Define w:=ℓ+bw := \ell + bw:=ℓ+b. Then for constant γ\gammaγ the bracket is 2ξ Hij wl2\xi\,H_{ij}\,w_l2ξHij​wl​, the same structure as the rank-one loss embedding with ℓ\ellℓ replaced by www. Contracting with 12 gkl ℓk\tfrac{1}{2}\,g^{kl}\,\ell_k21​gklℓk​,

Γijk(g) ℓk=ξ Hij (gkl ℓk wl)=ξ β(θ) Hij,\Gamma^k_{ij}(g)\,\ell_k = \xi\,H_{ij}\,(g^{kl}\,\ell_k\,w_l) = \xi\,\beta(\theta)\,H_{ij},Γijk​(g)ℓk​=ξHij​(gklℓk​wl​)=ξβ(θ)Hij​,

giving (HessgL)ij=Hij (1−ξ β(θ))(\mathrm{Hess}_g L)_{ij} = H_{ij}\,(1 - \xi\,\beta(\theta))(Hessg​L)ij​=Hij​(1−ξβ(θ)). The Riemannian Hessian is HHH multiplied by a single scalar, so eigenvalue signs are preserved (or globally flipped, exchanging minima and maxima but not turning an indefinite Hessian into a definite one). Fixed off-diagonal couplings with a=ba = ba=b cannot enable sign flips.

The case a≠ba \ne ba=b. When a≠ba \ne ba=b, the ambient metric hhh is not symmetric, so the pullback is not a Riemannian metric and the Riemannian Hessian construction does not apply. The construction in this case should be analysed as a non-symmetric preconditioner (see the off-diagonal entry Section 4 for the honest framing). The bracket calculation above still produces meaningful algebraic content, however: the extra term ξ Hil (aj−bj)\xi\,H_{il}\,(a_j - b_j)ξHil​(aj​−bj​) in the cyclic sum is not of the form "HijH_{ij}Hij​ times a vector contracted on lll," so the analogue of the contraction Γijk ℓk\Gamma^k_{ij}\,\ell_kΓijk​ℓk​ in the symmetric case picks up a rank-one piece 12 ξ (aj−bj) (Hu)i\tfrac{1}{2}\,\xi\,(a_j - b_j)\,(H u)_i21​ξ(aj​−bj​)(Hu)i​ with ul=gkl ℓku^l = g^{kl}\,\ell_kul=gklℓk​. This term has separate row and column structure and is not proportional to HHH. A direct generalisation of "Riemannian Hessian" to this non-Riemannian setting is not available; whether some symmetrised or "preconditioned-Hessian" eigenvalue analysis recovers a useful sign-flip statement here is open.

The factor of HijH_{ij}Hij​ that comes out of the cyclic cancellation is the algebraic obstruction in the symmetric case: as long as the bracket has the form "HijH_{ij}Hij​ times a vector contracted later," the Riemannian Hessian is forced to inherit the eigenstructure of HHH unchanged. The asymmetric case breaks this factorisation but pays the price of no longer being a Riemannian construction.

Appendix C: Christoffel symbols of the diagonal metric

For the diagonal metric γmn(θ)=esm(θ) δmn\gamma_{mn}(\theta) = e^{s_m(\theta)}\,\delta_{mn}γmn​(θ)=esm​(θ)δmn​, the inverse metric is γmn=e−smδmn\gamma^{mn} = e^{-s_m}\delta^{mn}γmn=e−sm​δmn, and the Christoffel bracket of γ\gammaγ alone has components

Bij,l(γ)=∂iγjl+∂jγil−∂lγij.B_{ij,l}(\gamma) = \partial_i \gamma_{jl} + \partial_j \gamma_{il} - \partial_l \gamma_{ij}.Bij,l​(γ)=∂i​γjl​+∂j​γil​−∂l​γij​.

Using γjl=esjδjl\gamma_{jl} = e^{s_j}\delta_{jl}γjl​=esj​δjl​, so ∂iγjl=esj ∂isj δjl\partial_i \gamma_{jl} = e^{s_j}\,\partial_i s_j\,\delta_{jl}∂i​γjl​=esj​∂i​sj​δjl​, this expands to

Bij,l(γ)=esj ∂isj δjl+esi ∂jsi δil−esi ∂lsi δij.B_{ij,l}(\gamma) = e^{s_j}\,\partial_i s_j\,\delta_{jl} + e^{s_i}\,\partial_j s_i\,\delta_{il} - e^{s_i}\,\partial_l s_i\,\delta_{ij}.Bij,l​(γ)=esj​∂i​sj​δjl​+esi​∂j​si​δil​−esi​∂l​si​δij​.

Each term is supported only when specific indices coincide. Combining with γkl=e−skδkl\gamma^{kl} = e^{-s_k}\delta^{kl}γkl=e−sk​δkl, the Christoffel symbols Γijk(γ)=12γklBij,l(γ)\Gamma^k_{ij}(\gamma) = \tfrac{1}{2}\gamma^{kl}B_{ij,l}(\gamma)Γijk​(γ)=21​γklBij,l​(γ) are

Γijk(γ)=12 e−sk(esj ∂isj δjk+esi ∂jsi δik−esi ∂ksi δij).\Gamma^k_{ij}(\gamma) = \tfrac{1}{2}\,e^{-s_k}\bigl(e^{s_j}\,\partial_i s_j\,\delta_{jk} + e^{s_i}\,\partial_j s_i\,\delta_{ik} - e^{s_i}\,\partial_k s_i\,\delta_{ij}\bigr).Γijk​(γ)=21​e−sk​(esj​∂i​sj​δjk​+esi​∂j​si​δik​−esi​∂k​si​δij​).

Case-by-case:

(i) i=j=ki = j = ki=j=k. All three terms contribute and equal esi ∂isie^{s_i}\,\partial_i s_iesi​∂i​si​, giving Γiii=12 ∂isi\Gamma^i_{ii} = \tfrac{1}{2}\,\partial_i s_iΓiii​=21​∂i​si​.

(ii) i=j≠ki = j \ne ki=j=k. Only the third term survives; the δij\delta_{ij}δij​ matches, the others do not. Γiik=−12 esi−sk ∂ksi\Gamma^k_{ii} = -\tfrac{1}{2}\,e^{s_i - s_k}\,\partial_k s_iΓiik​=−21​esi​−sk​∂k​si​. This is where the exponential anisotropy esi−ske^{s_i - s_k}esi​−sk​ first appears.

(iii) i≠ji \ne ji=j, k=jk = jk=j. Only the first term survives. Γijj=12 ∂isj\Gamma^j_{ij} = \tfrac{1}{2}\,\partial_i s_jΓijj​=21​∂i​sj​.

(iv) i≠ji \ne ji=j, k=ik = ik=i. Only the second term survives. Γjii=12 ∂jsi\Gamma^i_{ji} = \tfrac{1}{2}\,\partial_j s_iΓjii​=21​∂j​si​ (note the symmetry Γiji=Γjii\Gamma^i_{ij} = \Gamma^i_{ji}Γiji​=Γjii​ in the lower indices).

(v) i≠ji \ne ji=j, k≠ik \ne ik=i, k≠jk \ne jk=j. All three terms vanish; the Kronecker deltas do not match. Γijk=0\Gamma^k_{ij} = 0Γijk​=0.

C.1 Trace of the diagonal correction

Compute Cii=∑kΓiik(γ) ℓkC_{ii} = \sum_k \Gamma^k_{ii}(\gamma)\,\ell_kCii​=∑k​Γiik​(γ)ℓk​ using the cases above,

Cii=12 ∂isi ℓi−12∑k≠iesi−sk ∂ksi ℓk.C_{ii} = \tfrac{1}{2}\,\partial_i s_i\,\ell_i - \tfrac{1}{2}\sum_{k \ne i} e^{s_i - s_k}\,\partial_k s_i\,\ell_k.Cii​=21​∂i​si​ℓi​−21​k=i∑​esi​−sk​∂k​si​ℓk​.

Summing over iii,

tr(C)=12∑i∂isi ℓi−12∑i∑k≠iesi−sk ∂ksi ℓk.(C.1)\mathrm{tr}(C) = \tfrac{1}{2}\sum_i \partial_i s_i\,\ell_i - \tfrac{1}{2}\sum_i \sum_{k \ne i} e^{s_i - s_k}\,\partial_k s_i\,\ell_k. \tag{C.1}tr(C)=21​i∑​∂i​si​ℓi​−21​i∑​k=i∑​esi​−sk​∂k​si​ℓk​.(C.1)

Sanity check: recover the scalar trace identity. Restrict to si=ss_i = ssi​=s for all iii (the conformal case). Then ∂ksi=∂ks\partial_k s_i = \partial_k s∂k​si​=∂k​s and every esi−sk=1e^{s_i - s_k} = 1esi​−sk​=1, so

tr(C)=12 (∇s⋅ℓ)−12 (N−1) (∇s⋅ℓ)=2−N2 (∇s⋅ℓ).\mathrm{tr}(C) = \tfrac{1}{2}\,(\nabla s \cdot \ell) - \tfrac{1}{2}\,(N - 1)\,(\nabla s \cdot \ell) = \tfrac{2 - N}{2}\,(\nabla s \cdot \ell).tr(C)=21​(∇s⋅ℓ)−21​(N−1)(∇s⋅ℓ)=22−N​(∇s⋅ℓ).

This matches Level 1.

Why no constraint without that restriction. In (C.1) without the conformal restriction, the exponential factors esi−ske^{s_i - s_k}esi​−sk​ for i≠ki \ne ki=k are independent positive numbers (one for each ordered pair). Combined with the freely chosen ∂ksi\partial_k s_i∂k​si​ (the entries of the Jacobian σ\sigmaσ), the right-hand side of (C.1) is a linear combination of ℓk\ell_kℓk​ with arbitrary positive-weighted coefficients. No algebraic identity forces it to vanish or to take a specific form. The trace is unconstrained.

C.2 The Gershgorin construction

With σki=∂ksi\sigma_{ki} = \partial_k s_iσki​=∂k​si​ chosen as σii=−2M/ℓi\sigma_{ii} = -2M/\ell_iσii​=−2M/ℓi​ and σki=0\sigma_{ki} = 0σki​=0 for k≠ik \ne ik=i, every Christoffel symbol that requires off-diagonal σ\sigmaσ vanishes. From case (ii), Γiik=−12esi−sk⋅0=0\Gamma^k_{ii} = -\tfrac{1}{2}e^{s_i - s_k}\cdot 0 = 0Γiik​=−21​esi​−sk​⋅0=0 for k≠ik \ne ik=i. From cases (iii) and (iv), Γijj=12 ∂isj=0\Gamma^j_{ij} = \tfrac{1}{2}\,\partial_i s_j = 0Γijj​=21​∂i​sj​=0 and Γjii=12 ∂jsi=0\Gamma^i_{ji} = \tfrac{1}{2}\,\partial_j s_i = 0Γjii​=21​∂j​si​=0 for i≠ji \ne ji=j, since ∂isj=σij=0\partial_i s_j = \sigma_{ij} = 0∂i​sj​=σij​=0 in this construction.

Only the same-index Γiii=12σii\Gamma^i_{ii} = \tfrac{1}{2}\sigma_{ii}Γiii​=21​σii​ survives. Therefore

Cii=Γiii ℓi=12σii ℓi=−M for every i,Cij=0 for i≠j.C_{ii} = \Gamma^i_{ii}\,\ell_i = \tfrac{1}{2}\sigma_{ii}\,\ell_i = -M \text{ for every } i, \qquad C_{ij} = 0 \text{ for } i \ne j.Cii​=Γiii​ℓi​=21​σii​ℓi​=−M for every i,Cij​=0 for i=j.

So C=−M IC = -M\,IC=−MI, and H−C=H+M IH - C = H + M\,IH−C=H+MI. Each diagonal entry of HHH shifts up by +M+M+M and off-diagonals are unchanged.

By Gershgorin's disc theorem, the eigenvalues of any matrix lie in discs centred at the diagonal entries with radii equal to the off-diagonal row sums. Shifting every disc centre up by MMM while leaving the radii unchanged, for sufficiently large M>max⁡i(Ri−Hii)M > \max_i (R_i - H_{ii})M>maxi​(Ri​−Hii​), every disc lies strictly in the positive half-plane. Then every eigenvalue of H−CH - CH−C is strictly positive and geodesic convexity is achieved.

The construction is not minimum-budget; it is a constructive proof of existence. The actual minimum-budget σ\sigmaσ is given by the SDP in Section 4.4.

Appendix D: Coordinate invariance

The Riemannian Hessian (HessgL)ij(\mathrm{Hess}_g L)_{ij}(Hessg​L)ij​ transforms as a (0, 2)-tensor under coordinate changes. Its eigenvalues with respect to the metric ggg (that is, the solutions of det⁡(HessgL−λ g)=0\det(\mathrm{Hess}_g L - \lambda\,g) = 0det(Hessg​L−λg)=0) are coordinate-invariant scalars. The sign-flip analysis above is therefore geometrically meaningful, not an artifact of the coordinate basis we used.

A practical implication: even though the diagonal-metric class γ=diag(esi)\gamma = \mathrm{diag}(e^{s_i})γ=diag(esi​) does single out the coordinate basis (it is diagonal in the parameter coordinates, not in some other basis), the questions it answers (can we achieve geodesic convexity? what is the budget threshold?) are coordinate-invariant. The choice of diagonal metric is a restriction on the class of γ\gammaγ we are willing to use, not on the geometry of the question. A different metric class would give a different feasibility region; the diagonal class is chosen for its computational tractability (O(N)O(N)O(N) storage and update cost) rather than for any geometric reason.

Appendix E: Attached materials

Notebooks (notebooks/):

  • notebooks/derivations.nb: symbolic verification of the Christoffel bracket factorisation, the Sherman-Morrison contraction, the conformal trace identity, the case-by-case Christoffel symbols of the diagonal metric, and the trace formula (C.1) including the sanity-check reduction to Level 1.
  • notebooks/sign-flip-sdp.nb: numerical SDP setup for max⁡λmin⁡(H−C)\max \lambda_{\min}(H - C)maxλmin​(H−C) over ∥σ∥F≤R\|\sigma\|_F \le R∥σ∥F​≤R. Budget-vs-eigenvalue tables for the symmetric saddle, asymmetric saddle, monkey saddle, and Rosenbrock function. The N=5N = 5N=5 scalar-failure check. Sign-flip verification on diagonal Hessians with random off-diagonal perturbations for N=3,5,10,50,100,1000N = 3, 5, 10, 50, 100, 1000N=3,5,10,50,100,1000.
  • notebooks/basin-enlargement.nb: computation and visualisation of Beucl\mathcal{B}_{\mathrm{eucl}}Beucl​ versus Bg\mathcal{B}_gBg​ on the quartic well, including the four-basins-merge-into-one figure.
  • notebooks/epsilon-excised-sign-flip.nb: symbolic derivation of Theorem 1ʹ (the budget-translated sign-flip criterion (6.3.1), the Taylor expansion at a Morse critical point, the cross-term lift and the uniform ε\varepsilonε-bound) together with numerical verification of the predicted ε(R)\varepsilon(R)ε(R) on the symmetric saddle (matches 4/R4/R4/R to scan resolution; sharp), the asymmetric saddle L=x2−3y2L = x^2 - 3 y^2L=x2−3y2, and the quartic well (failure region forms thin tubes around the 4 saddles and the maximum at the origin). The algebraic core is also formalised in Lean 4 in the project repository (zero sorry, six documented axioms).

Scripts (scripts/):

  • scripts/make_figures.py: reproduces every figure shown in this entry (PDF and PNG written to figures/).

Metadata

Type
theorem
Visibility
public
Published
May 17, 2026
Last updated
Jun 4, 2026

Tags

basin-enlargementchristoffel-symbolscritical-point-obstructioncurvature-correctiondifferential-geometrygeodesic-convexityinduced-metriclean-formalizationriemannian-hessiansign-flip