Prove2Me
Navigate
DiscoverFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

Linear algebra

50 missions · 32 completed

Missions

Open18Completed32All50
CombinatoricsGraph Theory·Captain: mikedeng1

On a Conjecture of Spectral Extremal Problems: If the Extremal Graphs for F Are Turán Graphs Plus a Fixed Number of Edges, Every F-Free Graph of Maximum Spectral Radius Is ExtremalResearch Paper

Motivation

Extremal graph theory asks how many edges a graph on nnn vertices can have without containing a fixed graph FFF. The answer, the Turán number ex(n,F)\mathrm{ex}(n,F)ex(n,F), and the graphs attaining it, the set Ex(n,F)\mathrm{Ex}(n,F)Ex(n,F) of extremal graphs, are known precisely only for special FFF; the Erdős–Stone–Simonovits theorem gives ex(n,F)=(1−1χ(F)−1+o(1))n22\mathrm{ex}(n,F) = (1 - \frac{1}{\chi(F)-1} + o(1))\frac{n^2}{2}ex(n,F)=(1−χ(F)−11​+o(1))2n2​, where χ(F)\chi(F)χ(F) is the chromatic number.

Spectral extremal graph theory asks the same question with the number of edges replaced by the spectral radius λ(G)\lambda(G)λ(G), the largest eigenvalue of the adjacency matrix. Since λ(G)≥2e(G)/n\lambda(G) \ge 2e(G)/nλ(G)≥2e(G)/n, a spectral bound implies an edge bound, and spectral extremal results are usually stronger than their edge versions. Nikiforov (Linear Algebra Appl. 427, 2007) showed that the Turán graph Tn,rT_{n,r}Tn,r​ maximises λ\lambdaλ among Kr+1K_{r+1}Kr+1​-free graphs, so for F=Kr+1F = K_{r+1}F=Kr+1​ the spectral and the edge extremal graphs coincide. Whether this happens for other FFF is the subject of the paper.

Timeline:

  • 1941 Turán: Tn,rT_{n,r}Tn,r​ is the unique extremal graph for Kr+1K_{r+1}Kr+1​.
  • 2003 Chen, Gould, Pfender, Wei (J. Combin. Theory Ser. B 89): ex(n,Fk,r+1)=e(Tn,r)+O(1)\mathrm{ex}(n, F_{k,r+1}) = e(T_{n,r}) + O(1)ex(n,Fk,r+1​)=e(Tn,r​)+O(1) for kkk copies of Kr+1K_{r+1}Kr+1​ sharing one vertex.
  • 2007 Nikiforov (Linear Algebra Appl. 427): spectral Turán theorem for Kr+1K_{r+1}Kr+1​. 2009 Nikiforov (J. Graph Theory 62): spectral stability for large forbidden subgraphs, the source of Lemma 2.5.
  • 2020 Cioabă, Feng, Tait, Zhang (Electron. J. Combin. 27): the spectral extremal graph for the friendship graph FkF_kFk​ lies in Ex(n,Fk)\mathrm{Ex}(n, F_k)Ex(n,Fk​).
  • 2022 Cioabă, Desai, Tait (European J. Combin. 99) conjecture: if the graphs in Ex(n,F)\mathrm{Ex}(n,F)Ex(n,F) are Turán graphs plus O(1)O(1)O(1) edges, then the spectral extremal graphs lie in Ex(n,F)\mathrm{Ex}(n,F)Ex(n,F) for large nnn. Before the general proof it was known for Kr+1K_{r+1}Kr+1​, the friendship graphs FkF_kFk​, the graphs Hs,kH_{s,k}Hs,k​ (Li, Peng) and the intersecting cliques Fk,rF_{k,r}Fk,r​ (Desai, Kang, Li, Ni, Tait, Wang, arXiv:2108.03587).
  • 2022/2023 Wang, Kang, Xue (arXiv:2203.10831; J. Combin. Theory Ser. B, 2023): the conjecture holds in general. This mission formalizes their Theorem 1.2.

Setting

All graphs are finite and simple. For a graph GGG on nnn vertices, A(G)A(G)A(G) is its 0/10/10/1 adjacency matrix and λ(G)\lambda(G)λ(G) is the largest eigenvalue of A(G)A(G)A(G). A graph GGG is FFF-free if no subgraph of GGG is isomorphic to FFF. The Turán number ex(n,F)\mathrm{ex}(n,F)ex(n,F) is the maximum number of edges e(G)e(G)e(G) over FFF-free graphs GGG on nnn vertices, and Ex(n,F)\mathrm{Ex}(n,F)Ex(n,F) is the set of FFF-free nnn-vertex graphs with ex(n,F)\mathrm{ex}(n,F)ex(n,F) edges. The Turán graph Tn,rT_{n,r}Tn,r​ is the complete rrr-partite graph on nnn vertices with parts of sizes ⌊n/r⌋\lfloor n/r\rfloor⌊n/r⌋ or ⌈n/r⌉\lceil n/r\rceil⌈n/r⌉.

The hypothesis on FFF is: for fixed integers r≥2r \ge 2r≥2 and a≥0a \ge 0a≥0 and all large nnn, ex(n,F)=e(Tn,r)+a\mathrm{ex}(n,F) = e(T_{n,r}) + aex(n,F)=e(Tn,r​)+a and every graph in Ex(n,F)\mathrm{Ex}(n,F)Ex(n,F) contains a spanning copy of Tn,rT_{n,r}Tn,r​, i.e. is Tn,rT_{n,r}Tn,r​ plus aaa edges. The paper states "adding O(1)O(1)O(1) edges" and fixes the constant at the start of Section 3 (p. 4): "We may assume that the graphs in Ex(n,F)\mathrm{Ex}(n,F)Ex(n,F) are obtained from Tn,rT_{n,r}Tn,r​ by adding aaa edges." The mission follows that reading. Examples: Kr+1K_{r+1}Kr+1​ with a=0a = 0a=0, the friendship graphs, and the intersecting cliques Fk,r+1F_{k,r+1}Fk,r+1​.

A graph GGG is spectral extremal for FFF if it is FFF-free and λ(G)≥λ(G′)\lambda(G) \ge \lambda(G')λ(G)≥λ(G′) for every FFF-free G′G'G′ on the same nnn vertices.

Formalization targets

Goal: Theorem 1.2

For r≥2r \ge 2r≥2, a≥0a \ge 0a≥0 and FFF satisfying the hypothesis, there is NNN such that for all n≥Nn \ge Nn≥N every spectral extremal graph GGG for FFF on nnn vertices satisfies

e(G)=ex(n,F),i.e.G∈Ex(n,F).e(G) = \mathrm{ex}(n,F), \qquad\text{i.e.}\qquad G \in \mathrm{Ex}(n,F).e(G)=ex(n,F),i.e.G∈Ex(n,F).

The statement contains no numerical constant, and NNN depends only on FFF, rrr, aaa.

Milestones

The milestones follow the proof, in order: strict monotonicity of λ\lambdaλ under proper subgraphs of a connected graph (Lemma 2.3); connectivity of GGG (Lemma 3.1); the bound λ(G)≥(1−1r)n−r4n+2an\lambda(G) \ge (1-\frac1r)n - \frac{r}{4n} + \frac{2a}{n}λ(G)≥(1−r1​)n−4nr​+n2a​ (Lemma 3.2); spectral stability for χ(F)=r+1\chi(F) = r+1χ(F)=r+1 (Corollary 2.6); for every maximum rrr-cut V1∪⋯∪VrV_1 \cup \dots \cup V_rV1​∪⋯∪Vr​, ∑ie(Vi)≤εn2\sum_i e(V_i) \le \varepsilon n^2∑i​e(Vi​)≤εn2 and ∣Vi∣=(1r±3ε)n|V_i| = (\frac1r \pm 3\sqrt\varepsilon)n∣Vi​∣=(r1​±3ε​)n (Lemma 3.3); a counting inequality for intersections (Lemma 2.8); e(G[Vi])≤ae(G[V_i]) \le ae(G[Vi​])≤a and minimum degree above (1−1r−3rε1/3)n(1 - \frac1r - 3r\varepsilon^{1/3})n(1−r1​−3rε1/3)n (Lemma 3.6); at most 2a2a2a vertices of each part have a neighbour in that part, and all others see every other part completely (Lemma 3.7); Perron entries xu≥1−20a2r2/nx_u \ge 1 - 20a^2r^2/nxu​≥1−20a2r2/n for a≥1a \ge 1a≥1 (Lemma 3.8); e(Gin)−e(Gout)≤ae(G_{in}) - e(G_{out}) \le ae(Gin​)−e(Gout​)≤a (Lemma 3.9); balancing two parts of a complete multipartite graph increases λ\lambdaλ (Lemma 2.7); and the maximum partition is balanced, ∣ni−nj∣≤1|n_i - n_j| \le 1∣ni​−nj​∣≤1 (Lemma 3.10).

Significance

The theorem settles the Cioabă–Desai–Tait conjecture: for every FFF whose extremal graphs are Turán graphs plus a bounded number of edges, the spectral extremal problem reduces to the edge extremal problem for large nnn. This recovers the earlier cases (friendship graphs, the graphs Hs,kH_{s,k}Hs,k​, intersecting cliques) at once, and it turns any future determination of Ex(n,F)\mathrm{Ex}(n,F)Ex(n,F) of this type into a spectral result with no further work.

The result is proved on paper; it has no machine-checked proof. Mathlib has Turán's theorem (extremalNumber_top, uniqueness of turanGraph) and the definition of extremalNumber, but no spectral extremal graph theory: no Perron–Frobenius theorem for graphs, no spectral Turán theorem, no stability theorem. The formalization would supply these, and the milestone statements are independently reusable: strict monotonicity of λ\lambdaλ (Lemma 2.3), the spectral comparison of complete multipartite graphs (Lemma 2.7) and spectral stability (Corollary 2.6) are standard tools of the area.

Difficulty

The natural first idea, comparing GGG with an extremal graph HHH by Rayleigh quotients, fails at the start: λ(G)≥λ(H)\lambda(G) \ge \lambda(H)λ(G)≥λ(H) gives only e(G)≥e(Tn,r)−o(n2)e(G) \ge e(T_{n,r}) - o(n^2)e(G)≥e(Tn,r​)−o(n2), far from ex(n,F)\mathrm{ex}(n,F)ex(n,F). Closing an additive gap of O(1)O(1)O(1) edges requires control of the Perron vector to within O(1/n)O(1/n)O(1/n) at every vertex and exact control of the part sizes. The second is the delicate step: an imbalance of one vertex between two parts costs Θ(1/n)\Theta(1/n)Θ(1/n) in λ\lambdaλ, while the aaa extra edges contribute only O(1/n2)O(1/n^2)O(1/n2) beyond the Turán graph, so the two effects must be compared at different scales. Further, the proof needs the deep spectral stability theorem of Nikiforov (Lemma 2.5), whose own proof is long.

Formalization scope

Graphs on nnn vertices are SimpleGraph (Fin n), matching Mathlib's extremalNumber n F; FFF is a graph on any finite type. λ(G)\lambda(G)λ(G) is the largest eigenvalue of G.adjMatrix ℝ (index 000 of Mathlib's decreasingly sorted eigenvalues₀), not an absolute value and not a norm. "FFF-free" is F.Free G (no copy of FFF, not necessarily induced). "Sufficiently large nnn" is ∃ N, ∀ n ≥ N with NNN chosen after F,r,aF, r, aF,r,a and before GGG; "sufficiently small ε\varepsilonε" is ∃ ε₀ > 0, ∀ ε ∈ (0, ε₀). Partitions V1∪⋯∪VrV_1 \cup \dots \cup V_rV1​∪⋯∪Vr​ are labellings Fin n → Fin r whose parts may be empty; the Section 3 lemmas hold for every partition maximising the number of crossing edges. Lemma 3.8 carries the extra hypothesis a≥1a \ge 1a≥1, because as printed it is false for a=0a = 0a=0 (for F=Kr+1F = K_{r+1}F=Kr+1​ and r∤nr \nmid nr∤n the Perron vector of Tn,rT_{n,r}Tn,r​ has entries below 111); the goal does not assume it.

Trivializing formalizations are ruled out: λ\lambdaλ is not defined from the edge count (which would make spectral and edge extremality the same), the hypothesis on FFF does not contain the conclusion and is satisfiable (a sorry-free check for F=Kr+1F = K_{r+1}F=Kr+1​, a=0a = 0a=0 was compiled), the conclusion is e(G)=ex(n,F)e(G) = \mathrm{ex}(n,F)e(G)=ex(n,F) and not a weaker bound, and the threshold is not chosen after GGG.

A complete development needs the Perron–Frobenius theorem for irreducible nonnegative symmetric matrices, the Rayleigh quotient characterisation of λ\lambdaλ, the spectrum of complete multipartite graphs, Nikiforov's spectral stability lemma, and max-cut partition arguments. All of these are reusable beyond this mission, and proofs of any milestone, of the cited Lemmas 2.1, 2.2 and 2.5, or of Nikiforov's spectral Turán theorem are welcome.

Selected references

  • J. Wang, L. Kang, Y. Xue, On a conjecture of spectral extremal problems, J. Combin. Theory Ser. B, 2023; arXiv:2203.10831v1 (2022). https://arxiv.org/abs/2203.10831
  • S. Cioabă, D. N. Desai, M. Tait, The spectral radius of graphs with no odd wheels, European J. Combin. 99 (2022) 103420.
  • S. Cioabă, L. H. Feng, M. Tait, X. D. Zhang, The maximum spectral radius of graphs without friendship subgraphs, Electron. J. Combin. 27(4) (2020) P4.22.
  • G. Chen, R. J. Gould, F. Pfender, B. Wei, Extremal graphs for intersecting cliques, J. Combin. Theory Ser. B 89 (2003) 159–171.
  • V. Nikiforov, Bounds on graph eigenvalues II, Linear Algebra Appl. 427 (2007) 183–189.
  • V. Nikiforov, Stability for large forbidden subgraphs, J. Graph Theory 62(4) (2009) 362–368.
  • D. N. Desai, L. Kang, Y. Li, Z. Ni, M. Tait, J. Wang, Spectral extremal graphs for intersecting cliques, arXiv:2108.03587v2 (2021). https://arxiv.org/abs/2108.03587
18 thms2 active usersReviewed
🏆Completed
Numerical AnalysisOperations ResearchOptimization·Captain: mikedeng1

Conjugate Gradient Methods with Inexact Searches: The Self-Scaled Direction Is a Multiple of Beale's Restart DirectionResearch Paper

Motivation

Conjugate gradient methods minimize a smooth function f:Rn→Rf:\mathbb R^n\to\mathbb Rf:Rn→R using only gradients and a handful of stored vectors. This makes them the standard choice when nnn is too large for Newton or quasi-Newton methods, which store an n×nn\times nn×n matrix. On a strictly convex quadratic with exact line searches the classical method of Hestenes and Stiefel terminates in at most nnn steps. On general functions, and with the inexact line searches used in practice, its behaviour is much less clear.

D. F. Shanno's 1978 paper in Mathematics of Operations Research (doi:10.1287/moor.3.3.244) links conjugate gradient methods to quasi-Newton methods. It writes the search direction as −H^g-\hat H g−H^g, where H^\hat HH^ is a positive definite approximation of the inverse Hessian that is never stored. The resulting "memoryless" BFGS directions give descent without exact line searches. The paper's new algorithm uses two BFGS updates: one from the last restart and one from the current step. Its first update is scaled by the Oren–Spedicato factor γt\gamma_tγt​. Shanno and Phua's CONMIN code implements the algorithm, and the memoryless BFGS direction is the one-pair case of the later limited-memory BFGS methods.

  • 1952: Hestenes and Stiefel, linear conjugate gradients.
  • 1964: Fletcher and Reeves, nonlinear conjugate gradients.
  • 1969: Polak and Ribière, a second nonlinear variant.
  • 1972: Beale, a restart procedure that keeps the computed direction dtd_tdt​.
  • 1977: Powell's restart criterion (Powell 1977).
  • 1978: Shanno's reformulation as memoryless and two-update quasi-Newton methods (this paper).

Setting

Vectors are columns in Rn\mathbb R^nRn. A prime denotes transpose: u′vu'vu′v is the inner product and uv′uv'uv′ the outer product. An iterative method produces points xkx_kxk​, steps pk=xk+1−xk=αkdkp_k = x_{k+1}-x_k = \alpha_k d_kpk​=xk+1​−xk​=αk​dk​ along search directions dkd_kdk​, gradients gk=∇f(xk)g_k = \nabla f(x_k)gk​=∇f(xk​), and gradient changes yk=gk+1−gky_k = g_{k+1}-g_kyk​=gk+1​−gk​. A line search is exact when pk′gk+1=0p_k'g_{k+1} = 0pk′​gk+1​=0.

The BFGS update of a matrix HHH with the pair (p,y)(p,y)(p,y) is

H+=H−p y′H+Hy p′p′y+(1+y′Hyp′y)pp′p′y.H^+ = H - \frac{p\,y'H + H y\,p'}{p'y} + \left(1+\frac{y'Hy}{p'y}\right)\frac{pp'}{p'y}.H+=H−p′ypy′H+Hyp′​+(1+p′yy′Hy​)p′ypp′​.

A restart cycle begins at iteration ttt. At a later iteration k>tk>tk>t, Shanno's self-scaled restart matrix is

H^k=γt(I−ptyt′+ytpt′pt′yt+yt′ytpt′ytptpt′pt′yt)+ptpt′pt′yt,γt=pt′ytyt′yt.\hat H_k = \gamma_t\left(I - \frac{p_ty_t'+y_tp_t'}{p_t'y_t} + \frac{y_t'y_t}{p_t'y_t}\frac{p_tp_t'}{p_t'y_t}\right) + \frac{p_tp_t'}{p_t'y_t}, \qquad \gamma_t = \frac{p_t'y_t}{y_t'y_t}.H^k​=γt​(I−pt′​yt​pt​yt′​+yt​pt′​​+pt′​yt​yt′​yt​​pt′​yt​pt​pt′​​)+pt′​yt​pt​pt′​​,γt​=yt′​yt​pt′​yt​​.

The matrix H^k+1\hat H_{k+1}H^k+1​ is its BFGS update with (pk,yk)(p_k,y_k)(pk​,yk​), and the self-scaled two-update direction is dk+1=−H^k+1gk+1d_{k+1} = -\hat H_{k+1}g_{k+1}dk+1​=−H^k+1​gk+1​. The unscaled variant uses the BFGS update of III in place of the first matrix.

Beale's direction is

dk+1=−gk+1+yk′gk+1dk′ykdk+yt′gk+1dt′ytdt.d_{k+1} = -g_{k+1} + \frac{y_k'g_{k+1}}{d_k'y_k}d_k + \frac{y_t'g_{k+1}}{d_t'y_t}d_t.dk+1​=−gk+1​+dk′​yk​yk′​gk+1​​dk​+dt′​yt​yt′​gk+1​​dt​.

The quadratic case has gradient g(x)=Ax+cg(x) = Ax + cg(x)=Ax+c with AAA symmetric positive definite.

Formalization targets

Goal: reduction of the self-scaled method to Beale's method

Let AAA be symmetric positive definite, gi=Axi+cg_i = Ax_i + cgi​=Axi​+c and t<kt<kt<k. Assume that for t≤i≤kt\le i\le kt≤i≤k we have xi+1=xi+pix_{i+1} = x_i + p_ixi+1​=xi​+pi​, pi=αidip_i = \alpha_i d_ipi​=αi​di​ and pi′gi+1=0p_i'g_{i+1}=0pi′​gi+1​=0, that pt′Api=0p_t'Ap_i = 0pt′​Api​=0 for t<i≤kt<i\le kt<i≤k, and that pt,pk≠0p_t, p_k \ne 0pt​,pk​=0. Then

−H^k+1gk+1=γt(−gk+1+yk′gk+1dk′ykdk+yt′gk+1dt′ytdt).-\hat H_{k+1}g_{k+1} = \gamma_t\left(-g_{k+1} + \frac{y_k'g_{k+1}}{d_k'y_k}d_k + \frac{y_t'g_{k+1}}{d_t'y_t}d_t\right).−H^k+1​gk+1​=γt​(−gk+1​+dk′​yk​yk′​gk+1​​dk​+dt′​yt​yt′​gk+1​​dt​).

This is the paper's claim that "for f(x)f(x)f(x) quadratic with exact searches each of the above methods reduces exactly to Beale's method defined by (28)", with the conclusion (44). The scale is exactly γt\gamma_tγt​.

Companion statements

  • The unscaled two-update direction equals Beale's direction exactly.
  • Both two-update directions are descent directions, gk+1′dk+1<0g_{k+1}'d_{k+1} < 0gk+1′​dk+1​<0, whenever pt′yt>0p_t'y_t > 0pt′​yt​>0 and pk′yk>0p_k'y_k > 0pk′​yk​>0. No exact search is needed.

Milestones on the path

  • (34): the expansion of −H^k+1gk+1-\hat H_{k+1}g_{k+1}−H^k+1​gk+1​.
  • (40): its form under an exact search.
  • (41): gradients along a run on a quadratic.
  • pt′gk+1=0p_t'g_{k+1} = 0pt′​gk+1​=0.
  • (38), corrected by a factor 2: the action of the self-scaled restart matrix.
  • (42): its form when pt′gk+1=0p_t'g_{k+1} = 0pt′​gk+1​=0.
  • (43): the direction after substitution.

Significance

The result. The reduction says that the new algorithm reproduces Beale's restarted conjugate gradient directions on a quadratic with exact line searches. So it keeps the finite-termination and rate-of-convergence properties behind Beale's restart. Away from that setting it behaves as a quasi-Newton method, whose directions are descent directions under any line search with p′y>0p'y>0p′y>0. The two regimes are what justify relaxing the line search, which the paper's computations exploit. The scale γt\gamma_tγt​ changes only the length of the step, not its direction.

Formalizing it. The claim is proved in the paper by a short computation, and no machine-checked version is known. This mission produces:

  • a checked statement of the claim with every hypothesis explicit, including the conjugacy the proof takes as known;
  • a corrected version of display (38), which is misprinted;
  • reusable definitions of the additive BFGS update and of Beale's direction.

Difficulty

The obvious attempt is to expand both BFGS updates symbolically and compare with Beale's formula. This fails without two facts that are not algebraic identities. The first is that the restart step stays orthogonal to every later gradient, pt′gk+1=0p_t'g_{k+1}=0pt′​gk+1​=0. It needs the affine gradient of a quadratic, the exact search at the restart step, and conjugacy of ptp_tpt​ with all later steps. The second is yk′pt=0y_k'p_t = 0yk′​pt​=0, which again comes from conjugacy. Beale's formula also has to be matched in its ddd-form: the coefficient y′gd′yd\frac{y'g}{d'y}dd′yy′g​d equals y′gp′yp\frac{y'g}{p'y}pp′yy′g​p only when the step length is nonzero. The descent statements need a different argument: the BFGS update of a positive definite matrix with p′y>0p'y>0p′y>0 must be shown to remain positive definite, and this has to be done twice.

Formalization scope

Vectors are Fin n → ℝ and matrices Matrix (Fin n) (Fin n) ℝ. The inner product u′vu'vu′v is u ⬝ᵥ v, the outer product uv′uv'uv′ is vecMulVec u v, and HvHvHv is H *ᵥ v. The quadratic enters only through its gradient A *ᵥ x + c with A.PosDef. The paper's (4) is the case c=−Ax^c = -A\hat xc=−Ax^. Iterates, steps, directions and gradients are sequences indexed by ℕ. Division is Lean's total division. Every statement that divides therefore carries hypotheses making its denominators nonzero: pt≠0p_t \ne 0pt​=0 and pk≠0p_k\ne0pk​=0 in the quadratic statements, and pt′yt≠0p_t'y_t\ne 0pt′​yt​=0 or p′y>0p'y>0p′y>0 in the generic ones.

The conjugacy pt′Api=0p_t'Ap_i=0pt′​Api​=0 for t<i≤kt<i\le kt<i≤k is a hypothesis, exactly as the paper's proof uses it. It is not derived from a full run of Beale's algorithm. The range k≤t+n−1k\le t+n-1k≤t+n−1 of Beale's formula is not assumed.

Several formalizations would make the claim easier than the paper's, and none of them is used:

  • defining the direction by the expanded formula (34), or the restart matrix by (38);
  • adding orthogonality or conjugacy hypotheses beyond those listed;
  • concluding only that the two directions are parallel;
  • dropping the restart term of Beale's direction;
  • allowing a zero denominator.

The generic milestones — (34), (40), (38), (42), (43) and the descent statements — are statements about arbitrary vectors and matrices and are reusable for any BFGS-based method. Proofs of any milestone, and alternative derivations of the goal, are welcome.

Selected references

  • D. F. Shanno, Conjugate Gradient Methods with Inexact Searches, Mathematics of Operations Research 3(3) (1978) 244–256. https://doi.org/10.1287/moor.3.3.244
  • E. M. L. Beale, A derivation of conjugate gradients, in F. A. Lootsma (ed.), Numerical Methods for Nonlinear Optimization, Academic Press, 1972, 39–43.
  • M. J. D. Powell, Restart procedures for the conjugate gradient method, Mathematical Programming 12 (1977) 241–254. https://doi.org/10.1007/BF01593790
  • M. R. Hestenes and E. Stiefel, Methods of conjugate gradients for solving linear systems, J. Res. Nat. Bur. Standards 49 (1952) 409–436. https://doi.org/10.6028/jres.049.044
  • S. S. Oren and E. Spedicato, Optimal conditioning of self-scaling variable metric algorithms, Mathematical Programming 10 (1976) 70–90. https://doi.org/10.1007/BF01580654
18 thms2 active usersReviewed
Control TheoryDynamical Systems·Captain: mikedeng1

Stabilization for a Perturbed Chain of Integrators in Prescribed Time I: Prescribed-Time Estimates for the Linear Time-Varying FeedbackResearch Paper

Motivation

Many control tasks have to settle at a given deadline, not merely eventually. Examples are missile guidance, where the interception time is fixed in advance, and rendezvous or consensus manoeuvres that must finish by a scheduled instant. Prescribed-time stabilization asks for a feedback that brings the state of a system to the origin at a time T>0T>0T>0 chosen by the designer, independently of the initial condition, and keeps this property under disturbances. It is stronger than finite-time stabilization, where the settling time depends on the initial state, and stronger than fixed-time stabilization, where it is only bounded uniformly.

Song, Wang, Holloway and Krstić (Automatica 2017) obtained prescribed-time regulation of normal-form systems with feedback gains that grow without bound as t→Tt\to Tt→T. The computations required explicit time derivatives of the gain function. Chitour, Ushirobira and Bouhemou (SIAM J. Control Optim. 2020) recast this construction as time-varying homogeneity: a time-dependent dilation of the state combined with a change of time turns the prescribed-time problem into a standard exponential-stabilization problem on an infinite horizon. The analysis then reduces to linear matrix inequalities. This mission formalizes the linear-feedback part of that paper (§3.1, pp. 1026–1030).

Timeline:

  • 2010: Chitour and Sigalotti (SIAM J. Control Optim. 48) prove a Lyapunov LMI for Jn−benKTJ_n - b e_n K^TJn​−ben​KT, uniform over a bounded range b‾≤b≤bˉ\underline b\le b\le\bar bb​≤b≤bˉ.
  • 2017: Song, Wang, Holloway and Krstić give prescribed-time regulation with a time-varying feedback.
  • 2020: Chitour, Ushirobira and Bouhemou remove the upper bound bˉ\bar bbˉ (Proposition 10) and derive the prescribed-time estimate by time-varying homogeneity (Corollary 14).

Setting

Fix a positive integer nnn and a prescribed time T>0T>0T>0. Let (ei)1≤i≤n(e_i)_{1\le i\le n}(ei​)1≤i≤n​ be the canonical basis of Rn\mathbb R^nRn, and let JnJ_nJn​ be the nnn-th Jordan block, Jnei=ei−1J_n e_i = e_{i-1}Jn​ei​=ei−1​ with e0=0e_0=0e0​=0. The perturbed chain of integrators is

x˙(t)=Jnx(t)+(d(t)+b(t)u(t))en,t∈[0,T),\dot x(t) = J_n x(t) + \big(d(t) + b(t)u(t)\big) e_n, \qquad t\in[0,T),x˙(t)=Jn​x(t)+(d(t)+b(t)u(t))en​,t∈[0,T),

with state x(t)∈Rnx(t)\in\mathbb R^nx(t)∈Rn and control u(t)∈Ru(t)\in\mathbb Ru(t)∈R. Here ddd is a measurable matched disturbance, and bbb is an uncertain control gain with b(t)≥b‾>0b(t)\ge\underline b>0b(t)≥b​>0 and no upper bound.

With weights ri=n−i+1r_i = n-i+1ri​=n−i+1, the dilation is Dμr=diag⁡(μr1,…,μrn)D^{\mathbf r}_\mu = \operatorname{diag}(\mu^{r_1},\dots,\mu^{r_n})Dμr​=diag(μr1​,…,μrn​) and Dr=diag⁡(r1,…,rn)D_{\mathbf r} = \operatorname{diag}(r_1,\dots,r_n)Dr​=diag(r1​,…,rn​). An admissible weight is a continuous function a≥0a\ge0a≥0 on [0,T][0,T][0,T] with ∫tTa>0\int_t^T a>0∫tT​a>0 for t<Tt<Tt<T. It defines

λ(t)=1∫tTa(ξ) dξ,s(t)=∫0tλ(ξ) dξ,0≤t<T.\lambda(t) = \frac{1}{\int_t^T a(\xi)\,d\xi},\qquad s(t) = \int_0^t\lambda(\xi)\,d\xi,\qquad 0\le t<T.λ(t)=∫tT​a(ξ)dξ1​,s(t)=∫0t​λ(ξ)dξ,0≤t<T.

The parameter λ\lambdaλ blows up at TTT, and sss maps [0,T)[0,T)[0,T) onto [0,∞)[0,\infty)[0,∞). In the coordinates y=Dλ(t)rxy = D^{\mathbf r}_{\lambda(t)}xy=Dλ(t)r​x and the time sss, the system becomes

y′=(aDr+Jn)y+(bu+d)en.y' = \big(a D_{\mathbf r} + J_n\big)y + (b u + d)e_n .y′=(aDr​+Jn​)y+(bu+d)en​.

The feedback studied is u=−KTDηryu = -K^T D^{\mathbf r}_\eta yu=−KTDηr​y, that is u(t)=−KTDηλ(t)rx(t)u(t) = -K^T D^{\mathbf r}_{\eta\lambda(t)}x(t)u(t)=−KTDηλ(t)r​x(t) in the original variables, where K∈RnK\in\mathbb R^nK∈Rn and η≥1\eta\ge1η≥1.

Formalization targets

Goal: Corollary 14 (corrected)

There is a rate μ>0\mu>0μ>0 depending only on nnn and b‾\underline bb​ such that, for every TTT and admissible aaa, there are KKK and C>0C>0C>0 with

∣xi(t)∣≤1(ηλ(t))n−i+1(Cηmax⁡(1,ηn−1)e−μηs(t)∥x(0)∥+Cmax⁡r∈[0,t]∣d(r)∣)|x_i(t)| \le \frac{1}{(\eta\lambda(t))^{n-i+1}}\Big(C\eta\max(1,\eta^{n-1})e^{-\mu\eta s(t)}\|x(0)\| + C\max_{r\in[0,t]}|d(r)|\Big)∣xi​(t)∣≤(ηλ(t))n−i+11​(Cηmax(1,ηn−1)e−μηs(t)∥x(0)∥+Cr∈[0,t]max​∣d(r)∣)

for every η≥1\eta\ge1η≥1, every b≥b‾b\ge\underline bb≥b​, every disturbance ddd, every closed-loop solution, every t∈[0,T)t\in[0,T)t∈[0,T) and every 1≤i≤n1\le i\le n1≤i≤n. Since λ(t)→∞\lambda(t)\to\inftyλ(t)→∞, every coordinate reaches 000 at the prescribed time TTT.

Milestones

  1. §3 (10): λ˙=aλ2\dot\lambda = a\lambda^2λ˙=aλ2, λ\lambdaλ is nondecreasing and tends to ∞\infty∞ at TTT, and sss is an increasing C1C^1C1 bijection [0,T)→[0,∞)[0,T)\to[0,\infty)[0,T)→[0,∞).
  2. §3 (7)–(11): the transformed dynamics.
  3. Proposition 10: ∃ρ,S,K\exists\rho,S,K∃ρ,S,K such that (Jn−benKT)TS+S(Jn−benKT)⪯−ρ Id(J_n - be_nK^T)^TS + S(J_n-be_nK^T)\preceq-\rho\,\mathrm{Id}(Jn​−ben​KT)TS+S(Jn​−ben​KT)⪯−ρId for all b≥b‾b\ge\underline bb≥b​.
  4. Proposition 11: the same LMI with aDraD_{\mathbf r}aDr​ added, uniformly in ∣a∣≤C0|a|\le C_0∣a∣≤C0​.
  5. Proposition 12 (corrected): the ISS estimate (17) in the time sss.
  6. Proposition 13: the η\etaη-scaled LMI (20).

Significance

Corollary 14 says that a linear feedback with a time-varying gain drives the chain of integrators to the origin at the designer's deadline TTT. The guarantee holds for every control gain b≥b‾b\ge\underline bb≥b​, however large, and gives an explicit input-to-state bound for the matched disturbance. The weight aaa is free, so the blow-up profile of the gain can be chosen, for example to converge faster than polynomially. Proposition 10 removes the upper bound on bbb from the classical LMI, and is of independent use for high-gain and persistently excited systems.

On formalization: the results are proved on paper; no formalization of them is known. The mission produces checked statements of the LMIs, the time change, and the prescribed-time estimate. Two printed estimates are false as stated, (17) and (23), and the mission states corrected versions (see Formalization scope). Formal proofs would certify the corrected constants.

Difficulty

The obvious approach is to take a Lyapunov function for the constant-coefficient closed loop and differentiate it along trajectories. This fails for two reasons. First, bbb ranges over an unbounded interval, so no compactness argument yields a common Lyapunov matrix, and a single SSS and KKK must work for all b≥b‾b\ge\underline bb≥b​ at once. Second, the transformed system carries the time-varying term a(s)Dra(s)D_{\mathbf r}a(s)Dr​, and in the original time the feedback gain λ(t)\lambda(t)λ(t) is unbounded. Estimates must survive both the rescaling by Dηλ(t)rD^{\mathbf r}_{\eta\lambda(t)}Dηλ(t)r​ and the change of time. Constants that look harmless in the new time, such as the condition number of SSS and the value λ(0)\lambda(0)λ(0), reappear in the original variables, and dropping them makes the printed estimates false.

Formalization scope

Lean representation: vectors of Rn\mathbb R^nRn are Fin n → ℝ, and coordinate i : Fin n is the paper's index i.val + 1, so rir_iri​ = n - i.val. Matrices are Matrix (Fin n) (Fin n) ℝ, and enKTe_nK^Ten​KT is vecMulVec (eN n) K. Matrix inequalities use the Loewner order (open scoped MatrixOrder), and "real symmetric positive definite" is PosDef. The norm ∥⋅∥\|\cdot\|∥⋅∥ is Euclidean, written out as eucNorm. Solutions are Carathéodory solutions in integral form: for every ttt in the interval, the right-hand side is integrable on [0,t][0,t][0,t] and the integral identity holds. This matches the paper's measurable bbb and ddd. The standing assumptions are n≥1n\ge1n≥1 (§3), T>0T>0T>0, aaa continuous and nonnegative on [0,T][0,T][0,T] with ∫tTa>0\int_t^Ta>0∫tT​a>0 on [0,T)[0,T)[0,T), and b≥b‾>0b\ge\underline b>0b≥b​>0. No upper bound on bbb is assumed anywhere. max⁡∣d∣\max|d|max∣d∣ is replaced by an arbitrary bound DDD of ∣d∣|d|∣d∣ on the interval.

Corrected statements, with milestone texts kept verbatim:

  • (17) as printed has no constant on the transient term and fails for n≥2n\ge2n≥2: with a≡1a\equiv1a≡1 and y(0)=e1y(0)=e_1y(0)=e1​, ∣y1∣|y_1|∣y1​∣ first increases. Proposition 12 is stated with CSC_SCS​ on that term.
  • (23) as printed fails at t=0t=0t=0 whenever ∫0Ta<1\int_0^Ta<1∫0T​a<1. The goal is stated with a constant CCC on the transient term, and KKK and CCC may depend on TTT and aaa. The rate μ\muμ and the disturbance gain CSC_SCS​ depend only on nnn and b‾\underline bb​, as on the page.
  • Proposition 13's undefined μ∗\mu_*μ∗​ is the ρ\rhoρ of its statement.
  • "λ\lambdaλ increasing" is read as nondecreasing.
  • Corollary 14's "t≥0t\ge0t≥0" is read as t∈[0,T)t\in[0,T)t∈[0,T).

Trivializing encodings are ruled out as follows. The integrability clause of IsIntegralSolution prevents a non-integrable right-hand side from integrating to 000. Every constant is quantified before the objects it must not depend on: the rate μ\muμ before TTT and aaa, KKK before η\etaη, and ρ,S,K\rho,S,Kρ,S,K before bbb. No statement weakens the estimates to mere convergence.

Needed infrastructure: Lyapunov estimates for absolutely continuous solutions, Loewner-order congruence, and interval-integral calculus for the time change. The LMI lemmas and the integral-form solution predicate are reusable in other control missions. Proofs of any milestone are welcome, and so are sharper constants.

Selected references

  • Y. Chitour, R. Ushirobira, H. Bouhemou, Stabilization for a Perturbed Chain of Integrators in Prescribed Time, SIAM J. Control Optim. 58(2):1022–1048, 2020. https://doi.org/10.1137/19M1285937
  • Y. Song, Y. Wang, J. Holloway, M. Krstić, Time-varying feedback for regulation of normal-form nonlinear systems in prescribed finite time, Automatica 83:243–251, 2017. https://doi.org/10.1016/j.automatica.2017.06.008
  • Y. Chitour, M. Sigalotti, On the stabilization of persistently excited linear systems, SIAM J. Control Optim. 48:4032–4055, 2010. https://doi.org/10.1137/080737812
10 thms2 active usersReviewed
🏆Completed
Numerical AnalysisProbabilityRandom Matrix Theory·Captain: mikedeng1

Randomized Algorithms for Estimating the Trace of an Implicit Symmetric Positive Semi-Definite Matrix V: Sample Bound for the Mixed Unit Vector Trace EstimatorResearch Paper

Motivation

Many computations in numerical linear algebra, statistics and computational physics need the trace of a matrix AAA that is never formed explicitly: AAA may be an inverse, a matrix function f(B)f(B)f(B), or a product of large operators, and the only affordable access is a routine that returns AvAvAv or vTAvv^TAvvTAv for a given vector vvv. Monte Carlo trace estimators handle this setting: draw random vectors zzz and average the quadratic forms zTAzz^TAzzTAz, each of which costs one matrix–vector product.

Avron and Toledo (J. ACM 2011) compare such estimators by the number of samples MMM that guarantee relative error ϵ\epsilonϵ with probability 1−δ1-\delta1−δ, and by the number of random bits each sample consumes. Hutchinson's estimator (Hutchinson 1990) and the Gaussian estimator need Ω(n)\Omega(n)Ω(n) random bits per sample. Section 8 of the paper studies two estimators that sample only from the nnn standard basis vectors and so need about log⁡2n\log_2 nlog2​n bits per sample, which allows the samples to be generated in advance. The plain version has a sample bound that depends on how uneven the diagonal of AAA is; the mixed version first multiplies AAA on both sides by a random orthogonal mixing matrix of the kind introduced by Ailon and Chazelle (2006) for the fast Johnson–Lindenstrauss transform and used by Avron, Maymounkov and Toledo (2010) in least-squares solvers. This mission formalizes the resulting sample bound, Theorem 8.4.

Setting

Let n≥1n \ge 1n≥1, let A∈Rn×nA \in \mathbb{R}^{n\times n}A∈Rn×n be symmetric positive semi-definite, and let e1,…,ene_1,\ldots,e_ne1​,…,en​ be the standard basis of Rn\mathbb{R}^nRn.

A random variable TTT is an (ϵ,δ)(\epsilon,\delta)(ϵ,δ)-approximator of trace(A)\mathrm{trace}(A)trace(A) if

Pr⁡(∣T−trace(A)∣≤ϵ trace(A))≥1−δ\Pr\bigl(|T-\mathrm{trace}(A)| \le \epsilon\,\mathrm{trace}(A)\bigr) \ge 1-\deltaPr(∣T−trace(A)∣≤ϵtrace(A))≥1−δ

(Definition 4.1).

The unit vector estimator with MMM samples is

UM=nM∑i=1MziTAzi,U_M = \frac{n}{M}\sum_{i=1}^M z_i^TAz_i,UM​=Mn​i=1∑M​ziT​Azi​,

where z1,…,zMz_1,\ldots,z_Mz1​,…,zM​ are independent uniform random samples from {e1,…,en}\{e_1,\ldots,e_n\}{e1​,…,en​} (Definition 3.4). Each term ziTAziz_i^TAz_iziT​Azi​ is a diagonal entry of AAA chosen uniformly at random. Its behaviour is governed by

rD(A)=n⋅max⁡iAiitrace(A),r_D(A) = \frac{n\cdot\max_i A_{ii}}{\mathrm{trace}(A)},rD​(A)=trace(A)n⋅maxi​Aii​​,

which lies between 111 and nnn.

A random mixing matrix is F=FD\mathcal F = FDF=FD, where the seed FFF is a fixed orthogonal n×nn\times nn×n matrix and DDD is diagonal with i.i.d. Rademacher entries, Pr⁡(Dii=±1)=1/2\Pr(D_{ii}=\pm1) = 1/2Pr(Dii​=±1)=1/2 (Definition 3.5). The seed enters through

η=max⁡i,j∣Fij∣2,\eta = \max_{i,j}|F_{ij}|^2,η=i,jmax​∣Fij​∣2,

which satisfies 1/n≤η≤11/n \le \eta \le 11/n≤η≤1; normalized DFT and Hadamard matrices attain η=1/n\eta = 1/nη=1/n, DCT and DHT matrices have η=2/n\eta = 2/nη=2/n (p. 8:5).

The mixed unit vector estimator is

TM=nM∑i=1MziTFAFTzi,T_M = \frac{n}{M}\sum_{i=1}^M z_i^T\mathcal F A\mathcal F^T z_i,TM​=Mn​i=1∑M​ziT​FAFTzi​,

with z1,…,zMz_1,\ldots,z_Mz1​,…,zM​ as above, independent of DDD (Definition 3.6). It is the unit vector estimator applied to FAFT\mathcal FA\mathcal F^TFAFT, whose trace equals trace(A)\mathrm{trace}(A)trace(A).

Formalization targets

Goal: Theorem 8.4

For every orthogonal seed FFF, every symmetric positive semi-definite AAA, every ϵ>0\epsilon > 0ϵ>0, δ∈(0,1)\delta \in (0,1)δ∈(0,1) and every M≥1M \ge 1M≥1,

M ≥ 2n2η2ϵ−2ln⁡(4/δ)ln⁡2(4n2/δ)⟹TM is an (ϵ,δ)-approximator of trace(A).M \ \ge\ 2n^2\eta^2\epsilon^{-2}\ln(4/\delta)\ln^2(4n^2/\delta) \quad\Longrightarrow\quad T_M \text{ is an } (\epsilon,\delta)\text{-approximator of } \mathrm{trace}(A).M ≥ 2n2η2ϵ−2ln(4/δ)ln2(4n2/δ)⟹TM​ is an (ϵ,δ)-approximator of trace(A).

Milestones

  1. Lemma 8.1. For symmetric AAA, E(U1)=trace(A)\mathrm{E}(U_1) = \mathrm{trace}(A)E(U1​)=trace(A) and Var(U1)=n∑iAii2−trace2(A)\mathrm{Var}(U_1) = n\sum_{i}A_{ii}^2 - \mathrm{trace}^2(A)Var(U1​)=n∑i​Aii2​−trace2(A).
  2. Theorem 8.2. UMU_MUM​ is an (ϵ,δ)(\epsilon,\delta)(ϵ,δ)-approximator of trace(A)\mathrm{trace}(A)trace(A) whenever
M≥12ϵ−2ln⁡(2/δ) rD2(A).M \ge \tfrac12\epsilon^{-2}\ln(2/\delta)\,r_D^2(A).M≥21​ϵ−2ln(2/δ)rD2​(A).
  1. Lemma 8.3. For U∈Rn×mU \in \mathbb{R}^{n\times m}U∈Rn×m with orthonormal columns and δ>0\delta > 0δ>0, with probability at least 1−δ1-\delta1−δ,
∣(FU)ij∣≤2ηln⁡(2mn/δ)for all i,j.|(\mathcal FU)_{ij}| \le \sqrt{2\eta\ln(2mn/\delta)} \quad\text{for all } i,j.∣(FU)ij​∣≤2ηln(2mn/δ)​for all i,j.
  1. Proof of Theorem 8.4, p. 8:13. With probability at least 1−δ/21-\delta/21−δ/2 over DDD, 0≤(FAFT)jj≤2ηln⁡(4n2/δ) trace(A)0 \le (\mathcal FA\mathcal F^T)_{jj} \le 2\eta\ln(4n^2/\delta)\,\mathrm{trace}(A)0≤(FAFT)jj​≤2ηln(4n2/δ)trace(A) for all jjj, and hence
rD(FAFT)≤2nηln⁡(4n2/δ).r_D(\mathcal FA\mathcal F^T) \le 2n\eta\ln(4n^2/\delta).rD​(FAFT)≤2nηln(4n2/δ).

Significance

Theorem 8.2 alone shows that the unit vector estimator can need order n2n^2n2 samples: when the trace is concentrated on one diagonal entry, rD(A)=nr_D(A) = nrD​(A)=n. Theorem 8.4 removes the dependence on AAA entirely. For a Fourier-type seed with η=Θ(1/n)\eta = \Theta(1/n)η=Θ(1/n) the bound becomes O(ϵ−2ln⁡(1/δ)ln⁡2(n/δ))O(\epsilon^{-2}\ln(1/\delta)\ln^2(n/\delta))O(ϵ−2ln(1/δ)ln2(n/δ)) samples for every positive semi-definite AAA, while each sample still costs about log⁡2n\log_2 nlog2​n random bits, and the nnn bits of DDD are drawn once. Among the estimators of the paper this is the only one with both an AAA-independent sample bound and logarithmic randomness per sample (Table I, p. 8:5). Lemma 8.3 is a standalone statement about randomized orthogonal transforms that is used well beyond trace estimation, in the analysis of subsampled randomized Hadamard transforms, sketching-based least squares, and fast Johnson–Lindenstrauss embeddings.

All results of the mission are proved in the literature: Lemma 8.3 in the cited works, the rest in the paper. As far as the platform's catalogue shows, none has a machine-checked proof. The mission produces checked statements of the paper's Section 8 with their exact constants, a probability model for random sign matrices and uniform basis-vector sampling that other randomized linear-algebra missions can reuse, and, once proved, a checked instance of Hoeffding's inequality applied to a concrete estimator.

Difficulty

The obvious argument for Theorem 8.4 applies Theorem 8.2 to FAFT\mathcal FA\mathcal F^TFAFT. That matrix is random, so Theorem 8.2, which is a statement about a fixed matrix, cannot be applied directly: the proof must condition on DDD, use that the samples ziz_izi​ are independent of DDD, and combine a failure event over DDD with a conditional failure event over the ziz_izi​, each with probability at most δ/2\delta/2δ/2. The second difficulty is Lemma 8.3: each entry (FU)ij=∑kFikDkkUkj(\mathcal FU)_{ij} = \sum_k F_{ik}D_{kk}U_{kj}(FU)ij​=∑k​Fik​Dkk​Ukj​ is a Rademacher sum whose coefficient vector has squared norm at most η\etaη, and the bound needs a sub-Gaussian tail for such sums together with a union bound over all mnmnmn entries. Bounding the diagonal of FAFT\mathcal FA\mathcal F^TFAFT through the diagonal of AAA alone does not work: each mixed diagonal entry depends on all entries of AAA, including the off-diagonal ones.

Formalization scope

Everything is over R\mathbb{R}R. The paper allows complex unitary seeds; since the estimator uses the transpose FT\mathcal F^TFT, the mission takes FFF real orthogonal (FTF=IF^TF = IFTF=I). Matrices are Matrix (Fin n) (Fin n) ℝ, "symmetric positive semi-definite" is Matrix.PosSemidef, and n≥1n \ge 1n≥1 is assumed throughout. The sample spaces are explicit product measures: indices k1,…,kMk_1,\ldots,k_Mk1​,…,kM​ uniform on Fin n with zi=ekiz_i = e_{k_i}zi​=eki​​, the diagonal of DDD with the nnn-fold Rademacher product law, and, for TMT_MTM​, the product of the two, which makes DDD and the ziz_izi​ independent as the paper assumes implicitly. Probabilities are Measure.real. η\etaη and max⁡iAii\max_i A_{ii}maxi​Aii​ are maxima over finite nonempty index sets; rDr_DrD​ uses real division, whose value at trace(A)=0\mathrm{trace}(A) = 0trace(A)=0 is irrelevant because a positive semi-definite matrix with zero trace is 000. The sample-count thresholds are exactly the paper's constants.

Deviations from the page, all recorded in the items' Formalization Notes: Definition 3.4 and Theorem 8.2 are stated for positive semi-definite rather than positive definite AAA (the proof uses only Aii≥0A_{ii} \ge 0Aii​≥0); Table I's entry 8ϵ−2ln⁡(4n2/δ)ln⁡(4/δ)8\epsilon^{-2}\ln(4n^2/\delta)\ln(4/\delta)8ϵ−2ln(4n2/δ)ln(4/δ) for the mixed estimator, which disagrees with Theorem 8.4, is not used; the proof of Theorem 8.4 prints the conditional failure probability as "≤1−δ/2\le 1-\delta/2≤1−δ/2" where δ/2\delta/2δ/2 is meant, and no statement copies it; Remark 8.5 ("for some small CCC") has no pinned constant and is not stated.

A formalization in which DDD is an arbitrary orthogonal diagonal matrix, the ziz_izi​ are correlated with DDD, or the law of the estimator is assumed rather than constructed would make the goal either false or a restatement of its hypotheses; the product-measure model rules this out.

Needed infrastructure: Hoeffding's inequality for bounded i.i.d. sums (in Mathlib as sub-Gaussian moment generating function bounds), a sub-Gaussian tail for Rademacher linear combinations, conditioning on one factor of a product measure, and the spectral theorem for real symmetric matrices. The Rademacher sign model and the random-mixing-matrix entry bound are reusable beyond this mission. Proofs of any milestone, and alternative arguments for Lemma 8.3, are welcome.

Selected references

  • H. Avron and S. Toledo, Randomized algorithms for estimating the trace of an implicit symmetric positive semi-definite matrix, J. ACM 58(2), Article 8, 2011. https://doi.org/10.1145/1944345.1944349
  • N. Ailon and B. Chazelle, Approximate nearest neighbors and the fast Johnson–Lindenstrauss transform, STOC 2006. https://doi.org/10.1145/1132516.1132597
  • H. Avron, P. Maymounkov and S. Toledo, Blendenpik: Supercharging LAPACK's least-squares solver, SIAM J. Sci. Comput. 32(3), 2010. https://doi.org/10.1137/090767911
  • M. F. Hutchinson, A stochastic estimator of the trace of the influence matrix for Laplacian smoothing splines, Comm. Statist. Simulation Comput. 19(2), 1990. https://doi.org/10.1080/03610919008812866
  • W. Hoeffding, Probability inequalities for sums of bounded random variables, J. Amer. Statist. Assoc. 58, 1963. https://doi.org/10.1080/01621459.1963.10500830
10 thms2 active usersReviewed
🏆Completed
Numerical AnalysisProbabilityRandom Matrix Theory·Captain: mikedeng1

Randomized Algorithms for Estimating the Trace of an Implicit Symmetric Positive Semi-Definite Matrix III: Sample Bound for Normalized Rayleigh-Quotient Trace EstimatorsResearch Paper

Motivation

Many computations in numerical linear algebra, statistics and computational physics need the trace of a matrix AAA that is never formed explicitly: AAA may be an inverse, a matrix function f(B)f(B)f(B), or a product of large operators, and the only affordable access is a routine that returns AvAvAv for a given vector vvv. Examples include log-determinants and the generalized cross-validation criterion in statistics, counting eigenvalues in an interval, and charge densities in electronic-structure computations. Monte Carlo trace estimators handle this setting: draw random vectors zzz, and average the quadratic forms zTAzz^TAzzTAz, each of which costs one matrix–vector product.

Hutchinson (1990) introduced the estimator with Rademacher vectors and computed its variance. Avron and Toledo (J. ACM 2011) replaced variance statements with sample bounds: how many samples MMM guarantee relative error ϵ\epsilonϵ with probability 1−δ1-\delta1−δ. Section 6 of their paper proves one such bound for an entire class of estimators at once, the normalized Rayleigh-quotient trace estimators, which contains Hutchinson's estimator and the unit vector estimator. This mission formalizes that bound (Theorem 6.1).

Setting

Let A∈Rn×nA \in \mathbb{R}^{n\times n}A∈Rn×n be symmetric positive semi-definite, with eigenvalues 0≤λ1≤⋯≤λn0 \le \lambda_1 \le \cdots \le \lambda_n0≤λ1​≤⋯≤λn​ and rank rank(A)\mathrm{rank}(A)rank(A). Write λn\lambda_nλn​ for the largest eigenvalue and

κf(A)=largest nonzero eigenvalue of Asmallest nonzero eigenvalue of A,\kappa_f(A) = \frac{\text{largest nonzero eigenvalue of }A}{\text{smallest nonzero eigenvalue of }A},κf​(A)=smallest nonzero eigenvalue of Alargest nonzero eigenvalue of A​,

defined for A≠0A \ne 0A=0; it is the condition number of AAA on its range.

A normalized Rayleigh-quotient trace estimator of AAA with MMM samples is

RM=1M∑i=1MziTAzi,R_M = \frac1M\sum_{i=1}^M z_i^TAz_i,RM​=M1​i=1∑M​ziT​Azi​,

where z1,…,zMz_1,\ldots,z_Mz1​,…,zM​ are independent random vectors in Rn\mathbb{R}^nRn with ziTzi=nz_i^Tz_i = nziT​zi​=n and E(ziTAzi)=trace(A)\mathrm{E}(z_i^TAz_i) = \mathrm{trace}(A)E(ziT​Azi​)=trace(A) for each iii (Definition 3.2). The vectors need not be identically distributed. Hutchinson's vectors (±1\pm1±1 entries, i.i.d. uniform) and the vectors n ek\sqrt n\,e_kn​ek​ with kkk uniform are instances.

A random variable TTT is an (ϵ,δ)(\epsilon,\delta)(ϵ,δ)-approximator of trace(A)\mathrm{trace}(A)trace(A) if

Pr⁡(∣T−trace(A)∣≤ϵ trace(A))≥1−δ\Pr\bigl(|T-\mathrm{trace}(A)| \le \epsilon\,\mathrm{trace}(A)\bigr) \ge 1-\deltaPr(∣T−trace(A)∣≤ϵtrace(A))≥1−δ

(Definition 4.1).

Formalization targets

Goal: Theorem 6.1, in the form its proof establishes

For every nonzero symmetric positive semi-definite AAA, every ϵ>0\epsilon>0ϵ>0, δ∈(0,1)\delta\in(0,1)δ∈(0,1), and every normalized Rayleigh-quotient estimator RMR_MRM​ of AAA,

M ≥ ln⁡(2/δ)⋅n2 κf2(A)2 rank2(A) ϵ2⟹RM is an (ϵ,δ)-approximator of trace(A).M \ \ge\ \frac{\ln(2/\delta)\cdot n^2\,\kappa_f^2(A)}{2\,\mathrm{rank}^2(A)\,\epsilon^2} \quad\Longrightarrow\quad R_M \text{ is an } (\epsilon,\delta)\text{-approximator of } \mathrm{trace}(A).M ≥ 2rank2(A)ϵ2ln(2/δ)⋅n2κf2​(A)​⟹RM​ is an (ϵ,δ)-approximator of trace(A).

Milestones (the displayed steps of the proof, p. 8:10)

  1. trace(A) κf(A)≥rank(A) λn\mathrm{trace}(A)\,\kappa_f(A) \ge \mathrm{rank}(A)\,\lambda_ntrace(A)κf​(A)≥rank(A)λn​.
  2. For every zzz with zTz=nz^Tz = nzTz=n:  0≤zTAz≤λnzTz=nλn≤nrank(A)trace(A) κf(A)\ 0 \le z^TAz \le \lambda_n z^Tz = n\lambda_n \le \frac{n}{\mathrm{rank}(A)}\mathrm{trace}(A)\,\kappa_f(A) 0≤zTAz≤λn​zTz=nλn​≤rank(A)n​trace(A)κf​(A).
  3. For every t>0t>0t>0:
Pr⁡(∣RM−trace(A)∣≥t)≤2exp⁡(−2M2rank2(A)t2Mn2trace2(A)κf2(A)).\Pr(|R_M-\mathrm{trace}(A)| \ge t) \le 2\exp\left(-\frac{2M^2\mathrm{rank}^2(A)t^2}{M n^2\mathrm{trace}^2(A)\kappa_f^2(A)}\right).Pr(∣RM​−trace(A)∣≥t)≤2exp(−Mn2trace2(A)κf2​(A)2M2rank2(A)t2​).
  1. For every ϵ>0\epsilon>0ϵ>0:
Pr⁡(∣RM−trace(A)∣≥ϵ trace(A))≤2exp⁡(−2Mrank2(A)ϵ2n2κf2(A)).\Pr(|R_M-\mathrm{trace}(A)| \ge \epsilon\,\mathrm{trace}(A)) \le 2\exp\left(-\frac{2M\mathrm{rank}^2(A)\epsilon^2}{n^2\kappa_f^2(A)}\right).Pr(∣RM​−trace(A)∣≥ϵtrace(A))≤2exp(−n2κf2​(A)2Mrank2(A)ϵ2​).

Significance

The result is distribution-free within the class: it needs only normalization and unbiasedness, so it covers Hutchinson's estimator, the unit vector estimator and any future normalized scheme with one argument. For well-conditioned matrices of full or nearly full rank the required number of samples is O(ϵ−2ln⁡(1/δ))O(\epsilon^{-2}\ln(1/\delta))O(ϵ−2ln(1/δ)), independent of nnn. For ill-conditioned matrices the bound degrades with κf2(A)\kappa_f^2(A)κf2​(A), which is the reason the paper proves sharper estimator-specific bounds in Sections 7 and 8; Theorem 6.1 is the baseline those results are compared against (Table I, p. 8:5).

The theorem is proved in the paper. No machine-checked version of it, or of any sample bound for trace estimators, is known to exist. The formalization adds a precise statement of the class of estimators on a general probability space, a corrected threshold (see below), and a Lean development that connects Mathlib's spectral theorem for symmetric matrices with its Hoeffding inequality for independent bounded variables.

Difficulty

Each step is short on paper; the work is in the interfaces. The eigenvalue inequality of milestone 1 requires relating the number of nonzero eigenvalues (with multiplicity) to rank(A)\mathrm{rank}(A)rank(A) and handling the maximum and minimum over the nonzero spectrum. Milestone 2 is the Rayleigh-quotient bound zTAz≤λnzTzz^TAz \le \lambda_n z^TzzTAz≤λn​zTz, which is a consequence of the spectral decomposition rather than a one-line identity. Milestone 3 applies Hoeffding's inequality to summands that are bounded only almost surely, are not identically distributed, and whose mean is fixed by hypothesis rather than computed; the two-sided bound must be assembled from two one-sided tails, and the event ∣RM−trace(A)∣≥t|R_M - \mathrm{trace}(A)| \ge t∣RM​−trace(A)∣≥t must be rescaled to a statement about the sum ∑iziTAzi\sum_i z_i^TAz_i∑i​ziT​Azi​. A naive attempt that fixes a particular distribution for the ziz_izi​ (Rademacher, say) proves a different, narrower theorem and does not settle the goal.

Formalization scope

  • Matrices and spectrum. AAA is Matrix (Fin n) (Fin n) ℝ with A.PosSemidef and A ≠ 0; eigenvalues are Mathlib's IsHermitian.eigenvalues. λn\lambda_nλn​ is lambdaMax (the maximum eigenvalue) and κf(A)\kappa_f(A)κf​(A) is kappaF (maximum over minimum of the finite set of nonzero eigenvalues); both are Finset.max'/min'/sup' of nonempty finite sets. κf(0)\kappa_f(0)κf​(0) is a placeholder, and every statement assumes A≠0A \ne 0A=0.
  • Probability model. A general probability space (Ω,P)(\Omega, P)(Ω,P) and random vectors z : Fin M → Ω → Fin n → ℝ satisfying IsNormalizedRayleighSample P A z: each ziz_izi​ measurable, the family mutually independent (iIndepFun), ziTzi=nz_i^Tz_i = nziT​zi​=n almost surely, and ∫ziTAzi dP=trace(A)\int z_i^TAz_i\,dP = \mathrm{trace}(A)∫ziT​Azi​dP=trace(A). The estimator is universally quantified over this class. Unbiasedness is required for the given AAA only, as on the page. Probabilities are P.real of events; M≥1M \ge 1M≥1 is a natural number and 1/M1/M1/M is (M : ℝ)⁻¹.
  • Correction of the printed statement. Theorem 6.1 and Table I print the threshold 12ϵ−2n−2rank2(A)ln⁡(2/δ)κf2(A)\tfrac12\epsilon^{-2}n^{-2}\mathrm{rank}^2(A)\ln(2/\delta)\kappa_f^2(A)21​ϵ−2n−2rank2(A)ln(2/δ)κf2​(A). The last display of the proof gives ln⁡(2/δ) n2κf2(A)/(2 rank2(A)ϵ2)\ln(2/\delta)\,n^2\kappa_f^2(A)/(2\,\mathrm{rank}^2(A)\epsilon^2)ln(2/δ)n2κf2​(A)/(2rank2(A)ϵ2), with the exponents of nnn and rank(A)\mathrm{rank}(A)rank(A) swapped. The printed version is false: for n=2n=2n=2, A=e1e1TA=e_1e_1^TA=e1​e1T​, z=2 ekz=\sqrt2\,e_kz=2​ek​ with kkk uniform and ϵ=δ=1/2\epsilon=\delta=1/2ϵ=δ=1/2 it admits M=1M=1M=1, while R1∈{0,2}R_1\in\{0,2\}R1​∈{0,2}. The goal states the proof's threshold.
  • Indexing slip. The proof writes 0=λ1=⋯=λk0=\lambda_1=\cdots=\lambda_k0=λ1​=⋯=λk​ with k=n−rank(A)+1k=n-\mathrm{rank}(A)+1k=n−rank(A)+1 and κf(A)=λn/λk\kappa_f(A)=\lambda_n/\lambda_kκf​(A)=λn​/λk​, which would make λk=0\lambda_k=0λk​=0; milestone 1 states the inequality with κf\kappa_fκf​ as defined, not the indexing.
  • Non-trivialization. The hypotheses of the class are satisfiable (by the unit vector and Hutchinson estimators), A≠0A\ne0A=0 excludes the degenerate κf(0)\kappa_f(0)κf​(0), and all divisions in the statements have positive denominators, so no statement holds vacuously or through a junk value.
  • Reusable parts. A proof of milestone 2 is a general Rayleigh-quotient bound for symmetric matrices; milestone 3 is a two-sided Hoeffding bound for independent, almost surely bounded, non-identically distributed summands, useful well beyond this mission. Contributions welcome: proofs of any milestone, and a sorry-free instance showing a concrete estimator (e.g. n ek\sqrt n\,e_kn​ek​) satisfies IsNormalizedRayleighSample.

Selected references

  • H. Avron and S. Toledo, Randomized algorithms for estimating the trace of an implicit symmetric positive semi-definite matrix, Journal of the ACM 58(2), Article 8, 2011. https://doi.org/10.1145/1944345.1944349
  • M. F. Hutchinson, A stochastic estimator of the trace of the influence matrix for Laplacian smoothing splines, Communications in Statistics – Simulation and Computation 19(2), 433–450, 1990. https://doi.org/10.1080/03610919008812866
  • W. Hoeffding, Probability inequalities for sums of bounded random variables, Journal of the American Statistical Association 58(301), 13–30, 1963. https://doi.org/10.1080/01621459.1963.10500830
8 thms2 active usersReviewed
🏆Completed
CombinatoricsGraph TheoryProbability+1·Captain: mikedeng1

Matching Is as Easy as Matrix Inversion: Steps 1–3 Find a Minimum Weight Perfect Matching with Probability at Least 1/2Research Paper

Motivation

Deciding whether a graph has a perfect matching, and finding one, are basic problems of combinatorial optimization; Edmonds' blossom algorithm solves them sequentially in polynomial time. The question behind this paper is whether they can also be solved in parallel, in polylogarithmic time on polynomially many processors (the class NC, or RNC when random bits are allowed).

The algebraic route to that question goes through the Tutte matrix. Tutte (1947) showed that a graph has a perfect matching if and only if its Tutte matrix, a skew-symmetric matrix of indeterminates, has a nonzero determinant. Substituting random numbers for the indeterminates turns this into a randomized parallel decision procedure, but it does not say which perfect matching exists, and a graph may have exponentially many.

Mulmuley, Vazirani and Vazirani (Combinatorica 7 (1987) 105–113) resolve this with the isolating lemma: random small integer weights make the minimum weight member of an arbitrary set family unique with probability at least one half. Once a single perfect matching is isolated, one determinant and one adjugate of an integer matrix reveal it. The isolating lemma has since become a standard tool in randomized algorithms and complexity theory, well beyond matchings.

Timeline:

  • 1947, Tutte: a graph has a perfect matching iff the determinant of its Tutte matrix is a nonzero polynomial (doi:10.1112/jlms/s1-22.2.107).
  • 1979, Lovász: random substitution into the Tutte matrix gives a randomized algorithm for deciding whether a perfect matching exists (Fundamentals of Computation Theory, LNCS 1979).
  • 1986, Karp, Upfal and Wigderson: the first RNC algorithm that finds a perfect matching, with RNC³ running time (Combinatorica 6 (1986) 35–48).
  • 1987, Mulmuley, Vazirani and Vazirani: the isolating lemma and an RNC² algorithm that inverts one integer matrix (this paper).
  • 2016–2017, Fenner, Gurjar and Thierauf (arXiv:1601.06319) for bipartite graphs, and Svensson and Tarnawski (arXiv:1704.01929) for general graphs, partially derandomize the isolation step and place perfect matching in quasi-NC. Whether perfect matching is in NC remains open.

Setting

A set system (S,F)(S, F)(S,F) is a finite set SSS of elements together with a family FFF of subsets of SSS. Given a weight wx∈Nw_x \in \mathbb{N}wx​∈N for each element xxx, the weight of T⊆ST \subseteq ST⊆S is w(T)=∑x∈Twxw(T) = \sum_{x \in T} w_xw(T)=∑x∈T​wx​, and FFF has a unique minimum weight set if one member of FFF is strictly lighter than every other member.

A graph GGG has vertices v1,…,vnv_1, \dots, v_nv1​,…,vn​ (in Lean, Fin n, in their natural order) and edge set EEE, with m=∣E∣m = |E|m=∣E∣. A perfect matching is a set M⊆EM \subseteq EM⊆E such that every vertex lies in exactly one edge of MMM. The edges and the perfect matchings of GGG form a set system.

Given edge weights wij∈Nw_{ij} \in \mathbb{N}wij​∈N, the integer matrix BBB is obtained from the Tutte matrix by substituting 2wij2^{w_{ij}}2wij​ for its indeterminates:

bij=2wij if (vi,vj)∈E, i<j;bij=−2wij if (vi,vj)∈E, i>j;bij=0 otherwise.b_{ij} = 2^{w_{ij}} \ \text{if } (v_i, v_j) \in E,\ i < j; \qquad b_{ij} = -2^{w_{ij}} \ \text{if } (v_i, v_j) \in E,\ i > j; \qquad b_{ij} = 0 \ \text{otherwise}.bij​=2wij​ if (vi​,vj​)∈E, i<j;bij​=−2wij​ if (vi​,vj​)∈E, i>j;bij​=0 otherwise.

∣B∣|B|∣B∣ is its determinant, BijB_{ij}Bij​ the submatrix with row iii and column jjj removed, and adj⁡(B)\operatorname{adj}(B)adj(B) its adjugate, whose (j,i)(j, i)(j,i) entry is ±∣Bij∣\pm|B_{ij}|±∣Bij​∣.

The algorithm of §4 is:

  1. Step 1. Compute ∣B∣|B|∣B∣ and obtain www, the exponent for which 22w2^{2w}22w is the highest power of 2 dividing ∣B∣|B|∣B∣.
  2. Step 2. Compute adj⁡(B)\operatorname{adj}(B)adj(B).
  3. Step 3. Output every edge (vi,vj)(v_i, v_j)(vi​,vj​) for which the integer ∣Bij∣ 2wij/22w|B_{ij}|\,2^{w_{ij}}/2^{2w}∣Bij​∣2wij​/22w is odd.

Formalization targets

Goal: Steps 1–3 find a minimum weight perfect matching with probability at least 1/2

For every graph GGG that has a perfect matching, with edge weights drawn uniformly and independently from {1,…,2m}\{1, \dots, 2m\}{1,…,2m},

Pr⁡[the output of Steps 1–3 is a perfect matching of G of minimum weight] ≥ 12.\Pr\bigl[\text{the output of Steps 1–3 is a perfect matching of } G \text{ of minimum weight}\bigr] \ \ge\ \tfrac12 .Pr[the output of Steps 1–3 is a perfect matching of G of minimum weight] ≥ 21​.

This is the correctness half of the paper's Theorem (p. 109). The probability is a fraction of the (2m)m(2m)^m(2m)m weight functions.

Milestones

  1. Lemma 1 (isolating lemma): for a nonempty family FFF over an nnn-element set, weights uniform in [1,2n][1, 2n][1,2n] give a unique minimum weight set with probability ≥1/2\ge 1/2≥1/2.
  2. Isolation for perfect matchings (§4): with edge weights uniform in [1,2m][1, 2m][1,2m], the minimum weight perfect matching is unique with probability ≥1/2\ge 1/2≥1/2.
  3. Odd-cycle cancellation (proof of Lemma 2): for a skew-symmetric integer matrix, only permutations all of whose cycles have even length contribute to the determinant.
  4. Lemma 2: if the minimum weight perfect matching is unique, of weight www, then ∣B∣≠0|B| \neq 0∣B∣=0 and 22w2^{2w}22w is the highest power of 2 dividing ∣B∣|B|∣B∣.
  5. Lemma 3: under the same hypothesis, (vi,vj)∈M(v_i, v_j) \in M(vi​,vj​)∈M iff ∣Bij∣ 2wij/22w|B_{ij}|\,2^{w_{ij}}/2^{2w}∣Bij​∣2wij​/22w is odd.
  6. Steps 1–3, deterministic core: under the same hypothesis, Step 1 obtains the weight of MMM and Steps 2–3 output exactly MMM.

Two companion items are included but are not on the goal's path: the maximum weight version of Lemma 1 (the remark after its proof, p. 107) and Lemma 4 (p. 110): the lexicographically largest matching set, for vertices sorted by decreasing weight, is a heaviest matching set.

Significance

The isolating lemma is a statement about arbitrary set families with no structure assumed, which is why it transfers: it is used for isolating satisfying assignments, for parallel algorithms for exact matching and minimum weight matchings with small weights, and in the derandomization program that led to the quasi-NC matching algorithms cited above. Lemmas 2 and 3 are the bridge from a combinatorial object (a unique minimum weight perfect matching) to arithmetic facts about one integer matrix (2-adic valuations of its determinant and adjugate entries), which is what makes the algorithm reducible to matrix inversion.

All results of this mission are proved in the paper. What the mission adds is machine-checked proofs: Mathlib at the pinned revision contains Tutte's barrier theorem but neither the isolating lemma nor the Tutte-matrix determinant arguments, and a search of Prove2Me (September 2026) found no formalization of them. A complete development yields a reusable isolating lemma for finite set systems and a reusable determinant expansion for skew-symmetric matrices.

Difficulty

The probabilistic step is a union bound over elements, but the event bounded for each element, "the element is ambiguous", is defined through a threshold that depends on all the other weights; the argument needs independence of that threshold from the element's own weight, which is a product-space (Fubini-type) counting statement rather than a one-line estimate. In a counting formalization over {1,…,2n}S\{1, \dots, 2n\}^S{1,…,2n}S, each fibre must be handled separately.

The determinant steps require a genuine combinatorial involution on permutations: reversing an odd cycle must be well defined (a canonical choice of cycle) and self-inverse, preserve the sign, negate the value, and in Lemma 3 also preserve the constraint σ(i)=j\sigma(i) = jσ(i)=j, which is where "since nnn is even, there are at least two odd cycles" enters. Relating a permutation with only even cycles to a pair of perfect matchings whose union is its trail is the second nontrivial bijection. Divisibility must be tracked exactly: 22w2^{2w}22w divides every term, and every term other than the one of MMM is divisible by 22w+12^{2w+1}22w+1.

Formalization scope

Vertices are Fin n and the graph is G : SimpleGraph (Fin n) with decidable adjacency. Edge weights are functions G.edgeSet → ℕ; perfect matchings are Finset G.edgeSet in which every vertex lies in exactly one edge. The matrix is weightedTutteMatrix G w : Matrix (Fin n) (Fin n) ℤ, with the positive entry above the diagonal. Probabilities are ratios of counts over Fintype.piFinset (fun _ => Finset.Icc 1 (2m)), stated without division as (2m)m≤2⋅#{… }(2m)^m \le 2 \cdot \#\{\dots\}(2m)m≤2⋅#{…}; the weight range is exactly [1,2m][1, 2m][1,2m] (resp. [1,2n][1, 2n][1,2n] in Lemma 1). "x/2kx/2^kx/2k is odd" means 2k∣x2^k \mid x2k∣x and x/2kx/2^kx/2k is an odd integer. The minor ∣Bij∣|B_{ij}|∣Bij​∣ is taken as Mathlib's signed cofactor adjugate B j i; parity and divisibility do not see the sign. Step 1's www is ⌊ν2(∣B∣)/2⌋\lfloor \nu_2(|B|)/2\rfloor⌊ν2​(∣B∣)/2⌋.

Added hypotheses: Lemma 1 and its maximum version assume FFF nonempty (the printed lemma omits it and is false for F=∅F = \emptysetF=∅); the goal and the isolation milestone assume GGG has a perfect matching, which is the paper's own input assumption. Lemmas 2 and 3 allow arbitrary natural weights, as printed.

The algorithm's output is defined from BBB, ∣B∣|B|∣B∣, adj⁡(B)\operatorname{adj}(B)adj(B), the 2-adic valuation and parity only; a definition of the output that refers to perfect matchings or to minimality would trivialize the goal and is ruled out. The complexity half of the Theorem (RNC², O(n3.5m)O(n^{3.5}m)O(n3.5m) processors), which rests on Pan's matrix-inversion algorithm, is not formalized, nor are §5a–b and §6.

Contributions welcome: proofs of the milestones in any order, general lemmas about the permutation expansion of skew-symmetric determinants, and a counting form of the union bound over product spaces, all of which are reusable outside this mission.

Selected references

  • K. Mulmuley, U. V. Vazirani, V. V. Vazirani, Matching is as easy as matrix inversion, Combinatorica 7(1) (1987) 105–113. https://doi.org/10.1007/BF02579206
  • W. T. Tutte, The factorization of linear graphs, J. London Math. Soc. 22 (1947) 107–111. https://doi.org/10.1112/jlms/s1-22.2.107
  • R. M. Karp, E. Upfal, A. Wigderson, Constructing a perfect matching is in random NC, Combinatorica 6(1) (1986) 35–48. https://doi.org/10.1007/BF02579407
  • L. Lovász, On determinants, matchings, and random algorithms, Fundamentals of Computation Theory (FCT '79), 1979, 565–574.
  • S. Fenner, R. Gurjar, T. Thierauf, Bipartite perfect matching is in quasi-NC, STOC 2016. https://arxiv.org/abs/1601.06319
  • O. Svensson, J. Tarnawski, The matching problem in general graphs is in quasi-NC, FOCS 2017. https://arxiv.org/abs/1704.01929
10 thms2 active usersReviewed
🏆Completed
Numerical AnalysisOptimizationTheoretical Computer Science·Captain: mikedeng1

Sparse Approximate Solutions to Linear Systems 1: The Column Bound for Greedy SelectionResearch Paper

Motivation

Many problems in scientific computing and statistics ask for a solution of a linear system Ax≈bAx\approx bAx≈b that uses as few unknowns as possible. In statistics this is subset selection (Golub and Van Loan, Matrix Computations, 1983). In coding theory over binary matrices it is the minimum weight solution problem (Gallager, 1968). Natarajan's own motivation was radial basis interpolation (Hardy, 1988). There the coefficients of the interpolant solve a square nonsingular linear system (Michelli, 1986). Few nonzero coefficients make the interpolant cheap to evaluate and, by Occam's razor, less prone to fitting noise.

Natarajan's paper (SIAM J. Comput. 24 (1995) 227–234) makes two contributions. First, finding the sparsest approximate solution over the reals is NP-hard (Theorem 1, the subject of the companion mission). Second, the obvious greedy heuristic, a QR factorization whose column pivots are chosen by their correlation with the right-hand side, is provably good (Theorem 2). This mission formalizes Theorem 2. The greedy method is known today as orthogonal least squares (OLS), a variant of orthogonal matching pursuit. Natarajan's bound is among the earliest worst-case guarantees for this family of algorithms and is widely cited in the sparse approximation and compressed sensing literature.

Setting

Let A∈Rm×nA\in\mathbb R^{m\times n}A∈Rm×n have columns a1,…,ana_1,\dots,a_na1​,…,an​, let b∈Rmb\in\mathbb R^mb∈Rm and ε>0\varepsilon>0ε>0. Write ∥⋅∥2\|\cdot\|_2∥⋅∥2​ for the Euclidean norm and ∥x∥0\|x\|_0∥x∥0​ for the number of nonzero entries of xxx. The sparse approximate solution problem asks for xxx with ∥Ax−b∥2≤ε\|Ax-b\|_2\le\varepsilon∥Ax−b∥2​≤ε and ∥x∥0\|x\|_0∥x∥0​ minimal. Define

Opt⁡(δ)=min⁡{∥x∥0:∥Ax−b∥2≤δ}.\operatorname{Opt}(\delta)=\min\{\|x\|_0 : \|Ax-b\|_2\le\delta\}.Opt(δ)=min{∥x∥0​:∥Ax−b∥2​≤δ}.

Let A\mathbf AA be AAA with every column divided by its Euclidean norm. Let A+\mathbf A^+A+ be its Moore–Penrose pseudo-inverse, the unique matrix PPP with APA=A\mathbf AP\mathbf A=\mathbf AAPA=A, PAP=PP\mathbf AP=PPAP=P and AP\mathbf APAP, PAP\mathbf APA symmetric. Let ∥A+∥2\|\mathbf A^+\|_2∥A+∥2​ be its spectral norm, the ℓ2→ℓ2\ell_2\to\ell_2ℓ2​→ℓ2​ operator norm.

Algorithm Greedy keeps a working matrix A(r)A^{(r)}A(r) with columns aj(r)a^{(r)}_jaj(r)​, a working vector b(r)b^{(r)}b(r) and a set τ\tauτ of chosen indices. It starts from A(0)=AA^{(0)}=\mathbf AA(0)=A, b(0)=bb^{(0)}=bb(0)=b, τ=∅\tau=\emptysetτ=∅. While ∥b(r)∥2>ε\|b^{(r)}\|_2>\varepsilon∥b(r)∥2​>ε, it chooses an index k∉τk\notin\tauk∈/τ that maximizes ∣ak(r)Tb(r)∣|a_k^{(r)T}b^{(r)}|∣ak(r)T​b(r)∣ and replaces b(r)b^{(r)}b(r) by its projection onto the orthogonal complement of ak(r)a^{(r)}_kak(r)​. It adds kkk to τ\tauτ and replaces every column outside τ\tauτ by its normalized projection onto that complement. If every correlation aj(r)Tb(r)a_j^{(r)T}b^{(r)}aj(r)T​b(r) vanishes, the algorithm stops ("no solution exists"). A final solution phase solves the linear system Bx=b(0)−b(r)Bx=b^{(0)}-b^{(r)}Bx=b(0)−b(r) in the chosen columns BBB of AAA. The number of nonzero entries of the output is therefore at most the number ttt of selection iterations.

Formalization targets

Goal: Theorem 2, for AAA with linearly independent columns

If the columns of AAA are linearly independent and some xxx satisfies ∥Ax−b∥2≤ε/2\|Ax-b\|_2\le\varepsilon/2∥Ax−b∥2​≤ε/2, then every run of the selection phase, with any tie-breaking, performs

t≤⌈18 Opt⁡(ε/2) ∥A+∥22 ln⁡∥b∥2ε⌉t\le\Big\lceil 18\,\operatorname{Opt}(\varepsilon/2)\,\|\mathbf A^+\|_2^2\,\ln\frac{\|b\|_2}{\varepsilon}\Big\rceilt≤⌈18Opt(ε/2)∥A+∥22​lnε∥b∥2​​⌉

iterations. The paper prints the theorem without the independence hypothesis. The hypothesis is needed (see Formalization scope).

Milestones

The proof on pp. 230–233 passes through the following statements, in order:

  1. (12): some column satisfies ∣aj(r)Tb(r)∣≥∥b(r)∥22/(2N(r)∥u(r)∥2)|a_j^{(r)T}b^{(r)}|\ge\|b^{(r)}\|_2^2/(2\sqrt{N^{(r)}}\|u^{(r)}\|_2)∣aj(r)T​b(r)∣≥∥b(r)∥22​/(2N(r)​∥u(r)∥2​). Here u(r)u^{(r)}u(r) is a sparsest vector with ∥A(r)u(r)−b(r)∥2≤ε/2\|A^{(r)}u^{(r)}-b^{(r)}\|_2\le\varepsilon/2∥A(r)u(r)−b(r)∥2​≤ε/2 and N(r)=∥u(r)∥0N^{(r)}=\|u^{(r)}\|_0N(r)=∥u(r)∥0​.
  2. (18): ∥b(r+1)∥22≤(1−1/ρ)∥b(r)∥22\|b^{(r+1)}\|_2^2\le(1-1/\rho)\|b^{(r)}\|_2^2∥b(r+1)∥22​≤(1−1/ρ)∥b(r)∥22​ whenever ρ≥4N(r)∥u(r)∥22/∥b(r)∥22\rho\ge 4N^{(r)}\|u^{(r)}\|_2^2/\|b^{(r)}\|_2^2ρ≥4N(r)∥u(r)∥22​/∥b(r)∥22​.
  3. Lemma 1: t≤⌈2ρln⁡(∥b∥2/ε)⌉t\le\lceil2\rho\ln(\|b\|_2/\varepsilon)\rceilt≤⌈2ρln(∥b∥2​/ε)⌉ for any such ρ\rhoρ valid at every iteration.
  4. Lemma 3: N(r+1)≤N(r)≤N(0)N^{(r+1)}\le N^{(r)}\le N^{(0)}N(r+1)≤N(r)≤N(0).
  5. N(0)=Opt⁡(ε/2)N^{(0)}=\operatorname{Opt}(\varepsilon/2)N(0)=Opt(ε/2).
  6. The columns of A\mathbf AA indexed by the support σ\sigmaσ of u(r)u^{(r)}u(r) and by the chosen set τ\tauτ are linearly independent, and σ∩τ=∅\sigma\cap\tau=\emptysetσ∩τ=∅.
  7. (31): ∥u(r)∥2≤32∥Z+∥2∥b(r)∥2\|u^{(r)}\|_2\le\frac32\|Z^+\|_2\|b^{(r)}\|_2∥u(r)∥2​≤23​∥Z+∥2​∥b(r)∥2​ for the matrix ZZZ of those columns.
  8. The singular-value comparison ∥Z+∥2≤∥M+∥2\|Z^+\|_2\le\|M^+\|_2∥Z+∥2​≤∥M+∥2​ for a column submatrix ZZZ of a matrix MMM with independent columns.
  9. Lemma 2: ∥u(r)∥2≤32∥A+∥2∥b(r)∥2\|u^{(r)}\|_2\le\frac32\|\mathbf A^+\|_2\|b^{(r)}\|_2∥u(r)∥2​≤23​∥A+∥2​∥b(r)∥2​, for AAA with independent columns.

Items 1–7 hold for every matrix AAA. Items 8, 9 and the goal carry the independence hypothesis.

Significance

Theorem 2 is a bicriteria approximation guarantee for an NP-hard problem. The greedy output meets the error ε\varepsilonε with at most a factor 18∥A+∥22ln⁡(∥b∥2/ε)18\|\mathbf A^+\|_2^2\ln(\|b\|_2/\varepsilon)18∥A+∥22​ln(∥b∥2​/ε) more nonzeros than the best solution at error ε/2\varepsilon/2ε/2. The factor depends only on the conditioning of the normalized matrix and logarithmically on the required accuracy. Its structure follows Johnson's analysis of the greedy set cover algorithm (1974): a potential decreases by a constant factor per step, which gives a logarithmic number of steps. The intermediate facts (12), (18) and Lemma 1 are the template of many later analyses of matching pursuit and OLS.

The result is proved on paper, with a gap. The last step of the proof of Lemma 2 compares singular values of a submatrix with those of A\mathbf AA, and this comparison holds only when A\mathbf AA has full column rank. For general AAA, Theorem 2 and Lemma 2 are false as printed. The formalization produces a machine-checked proof of the corrected theorem and pins down exactly where the hypothesis enters. The hypothesis-free statements (12), (18), Lemma 1, Lemma 3 and (31) form reusable infrastructure for greedy sparse approximation. No existing formalization of this algorithm or of its guarantee, in Lean or elsewhere, was found for this mission.

Difficulty

Each step of the proof is short, but the objects are defined by an iteration. The columns aj(r)a^{(r)}_jaj(r)​ are repeatedly projected and renormalized, and the columns already chosen are left untouched. Every claim about iteration rrr therefore needs invariants: chosen columns are orthonormal and orthogonal to b(r)b^{(r)}b(r), and the remaining columns are normalized projections of the original ones onto the orthogonal complement of the chosen ones. A proof has to establish these by induction before any lemma can be applied. The sparsest vector u(r)u^{(r)}u(r) is defined by minimality, so Lemma 3 and the linear-independence claim are exchange arguments on supports rather than computations. Finally, the passage from (31) to Lemma 2 needs a quantitative fact about pseudo-inverses of column submatrices. Mathlib has neither the Moore–Penrose inverse of a rectangular matrix nor its norm as a reciprocal singular value.

A naive attempt to bound ∥u(r)∥2\|u^{(r)}\|_2∥u(r)∥2​ directly by ∥A+∥2∥A(r)u(r)∥2\|\mathbf A^+\|_2\|A^{(r)}u^{(r)}\|_2∥A+∥2​∥A(r)u(r)∥2​ fails: u(r)u^{(r)}u(r) multiplies the projected columns A(r)A^{(r)}A(r), not A\mathbf AA, and different sparsest solutions can have different norms.

Formalization scope

Vectors live in EuclideanSpace ℝ (Fin m), so every ∥⋅∥2\|\cdot\|_2∥⋅∥2​ is the Euclidean norm. The only ∥⋅∥∞\|\cdot\|_\infty∥⋅∥∞​ is the maximum of ∣aj(r)Tb(r)∣|a_j^{(r)T}b^{(r)}|∣aj(r)T​b(r)∣, which is written out explicitly. The algorithm is a recursion greedyState A b k r in the sequence of choices k : ℕ → Fin n. A run of ttt iterations (IsGreedyRun) requires, at each r<tr<tr<t: the strict while-condition ∥b(r)∥2>ε\|b^{(r)}\|_2>\varepsilon∥b(r)∥2​>ε, an unchosen index, a nonzero correlation, and maximality over the unchosen columns. The residual and the columns are computed, never assumed. Normalization sends 000 to 000, so a column lying in the span of the chosen ones stays zero and is never chosen. Opt⁡\operatorname{Opt}Opt is an infimum over ℕ, and the goal assumes that some xxx has ∥Ax−b∥2≤ε/2\|Ax-b\|_2\le\varepsilon/2∥Ax−b∥2​≤ε/2, since otherwise the infimum would be 000. The ceiling is the natural-number ceiling. It agrees with the printed one whenever the loop runs at least once, because then ∥b∥2>ε\|b\|_2>\varepsilon∥b∥2​>ε. The pseudo-inverse is any matrix satisfying the four Penrose equations. It is never defined as (ATA)−1AT(\mathbf A^T\mathbf A)^{-1}\mathbf A^T(ATA)−1AT, which would hide the rank assumption.

Added hypothesis. The goal, Lemma 2 and the singular-value step assume that the columns of AAA are linearly independent, which forces n≤mn\le mn≤m. Without it, Theorem 2 fails. Take m=2m=2m=2, n=200n=200n=200, columns (cos⁡θj,sin⁡θj)(\cos\theta_j,\sin\theta_j)(cosθj​,sinθj​) and (sin⁡θj,cos⁡θj)(\sin\theta_j,\cos\theta_j)(sinθj​,cosθj​) for 100 distinct θj∈[0.001,0.01]\theta_j\in[0.001,0.01]θj​∈[0.001,0.01], b=2(1,1)b=\sqrt2(1,1)b=2​(1,1) and ε=1\varepsilon=1ε=1. Then Opt⁡(1/2)=2\operatorname{Opt}(1/2)=2Opt(1/2)=2 and the bound evaluates to 111, but Greedy selects two columns. Lemma 2 fails for A=[e1,e2,(e1+e2)/2]\mathbf A=[e_1,e_2,(e_1+e_2)/\sqrt2]A=[e1​,e2​,(e1​+e2​)/2​] and b=β(−1,1)/2b=\beta(-1,1)/\sqrt2b=β(−1,1)/2​. The paper's motivating interpolation systems are square and nonsingular, so they satisfy the hypothesis. A hypothesis-free goal would replace ∥A+∥2\|\mathbf A^+\|_2∥A+∥2​ by the largest ∥Z+∥2\|Z^+\|_2∥Z+∥2​ over linearly independent column subsets ZZZ of A\mathbf AA, which is what (31) gives. That quantity is not printed in the paper, so it is not the goal here.

A statement in which the iterates are free sequences constrained by hypotheses, the greedy choice is dropped, or Opt⁡\operatorname{Opt}Opt is taken over an empty set would be trivially true or would not describe this algorithm. The encoding above rules these out.

A complete development needs Gram–Schmidt-type invariants of the iteration, exchange arguments for sparsest solutions, and the Moore–Penrose inverse with its spectral norm. The last of these is reusable well beyond this mission. Contributions of any milestone, of the general Penrose-inverse facts, or of alternative proofs are welcome.

Selected references

  • B. K. Natarajan, Sparse Approximate Solutions to Linear Systems, SIAM J. Comput. 24(2):227–234, 1995. https://doi.org/10.1137/s0097539792240406
  • G. H. Golub and C. F. Van Loan, Matrix Computations, Johns Hopkins University Press, 1983.
  • D. S. Johnson, Approximation algorithms for combinatorial problems, J. Comput. System Sci. 9:256–278, 1974. https://doi.org/10.1016/S0022-0000(74)80044-9
  • R. Penrose, A generalized inverse for matrices, Proc. Cambridge Philos. Soc. 51:406–413, 1955. https://doi.org/10.1017/S0305004100030401
12 thms2 active usersReviewed
Control TheoryConvex OptimizationNumerical Analysis+2·Captain: mikedeng1

Robust Solutions to Least-Squares Problems with Uncertain Data IV: A Semidefinite Upper Bound on the Linear-Fractional Worst-Case Residual, Exact for Full PerturbationsResearch Paper

Motivation

Least-squares fitting is a standard tool in estimation, identification and data analysis, and its data AAA, bbb are rarely known exactly. El Ghaoui and Lebret (SIAM J. Matrix Anal. Appl. 18(4), 1997) proposed to choose xxx to minimize the worst-case residual over a set of admissible data perturbations. Earlier missions of this series treat unstructured perturbations of [A b][A\ b][A b] and perturbations affine in a parameter vector. §5 of the paper covers a more general model, taken from robust identification (Doyle et al.): the perturbed data depend on an uncertain matrix Δ\DeltaΔ through a linear-fractional transformation. This form covers rational dependence of the data on uncertain parameters, max-norm bounds on independent parameters, and data matrices with some columns known exactly (pp. 1046–1047).

In this generality, deciding whether the worst-case residual is finite is NP-complete, and computing it is NP-hard even when the dependence is affine (§5.3, Lemma 5.1). Theorem 5.2 gives the tractable replacement: a semidefinite program whose value bounds the worst-case residual from above, and equals it when the perturbation is unstructured. The main tool is a structured form of the S-procedure. Robust control uses the same tool, with the scalings SSS and GGG below, to bound the real structured singular value (Fan, Tits and Doyle, 1991).

Setting

Vectors carry the Euclidean norm ∥v∥\|v\|∥v∥. For a matrix XXX, ∥X∥\|X\|∥X∥ is its largest singular value (operator norm between Euclidean spaces). Let D\mathcal DD be a linear subspace of RN×N\mathbb R^{N\times N}RN×N (the perturbation structure), and fix A∈Rn×mA \in \mathbb R^{n\times m}A∈Rn×m, b∈Rnb \in \mathbb R^nb∈Rn, L∈Rn×NL \in \mathbb R^{n\times N}L∈Rn×N, RA∈RN×mR_A \in \mathbb R^{N\times m}RA​∈RN×m, Rb∈RNR_b \in \mathbb R^NRb​∈RN, D∈RN×ND \in \mathbb R^{N\times N}D∈RN×N. For Δ∈D\Delta \in \mathcal DΔ∈D with det⁡(I−DΔ)≠0\det(I - D\Delta) \ne 0det(I−DΔ)=0 the perturbed data are

A(Δ)=A+LΔ(I−DΔ)−1RA,b(Δ)=b+LΔ(I−DΔ)−1Rb.A(\Delta) = A + L\Delta(I - D\Delta)^{-1}R_A, \qquad b(\Delta) = b + L\Delta(I - D\Delta)^{-1}R_b .A(Δ)=A+LΔ(I−DΔ)−1RA​,b(Δ)=b+LΔ(I−DΔ)−1Rb​.

With the normalization ρ=1\rho = 1ρ=1 (the paper's, with no loss of generality), the worst-case residual of x∈Rmx \in \mathbb R^mx∈Rm is

rD(A,b,x)=max⁡Δ∈D, ∥Δ∥≤1∥A(Δ)x−b(Δ)∥r_{\mathcal D}(A,b,x) = \max_{\Delta \in \mathcal D,\ \|\Delta\| \le 1} \|A(\Delta)x - b(\Delta)\|rD​(A,b,x)=Δ∈D, ∥Δ∥≤1max​∥A(Δ)x−b(Δ)∥

if det⁡(I−DΔ)≠0\det(I - D\Delta) \ne 0det(I−DΔ)=0 for every such Δ\DeltaΔ, and +∞+\infty+∞ otherwise (35). The commutant scalings are S={S=ST:SΔ=ΔS ∀Δ∈D}\mathcal S = \{S = S^T : S\Delta = \Delta S\ \forall \Delta \in \mathcal D\}S={S=ST:SΔ=ΔS ∀Δ∈D} and G={G=−GT:GΔ=ΔG ∀Δ∈D}\mathcal G = \{G = -G^T : G\Delta = \Delta G\ \forall \Delta \in \mathcal D\}G={G=−GT:GΔ=ΔG ∀Δ∈D} (37). The SDP constraint is

F(λ,S,G,x)=[ΘAx−bRAx−Rb(Ax−b)T(RAx−Rb)Tλ]≻0,Θ=[λI−LSLT−LSDT+LG−DSLT+GTLTS+DG−GDT−DSDT].(38),(39)\mathcal F(\lambda,S,G,x) = \begin{bmatrix} \Theta & \begin{matrix} Ax - b \\ R_Ax - R_b\end{matrix} \\ \begin{matrix}(Ax-b)^T & (R_Ax - R_b)^T\end{matrix} & \lambda\end{bmatrix} \succ 0, \quad \Theta = \begin{bmatrix} \lambda I - LSL^T & -LSD^T + LG \\ -DSL^T + G^TL^T & S + DG - GD^T - DSD^T\end{bmatrix}. \qquad (38),(39)F(λ,S,G,x)=​Θ(Ax−b)T​(RA​x−Rb​)T​​Ax−bRA​x−Rb​​λ​​≻0,Θ=[λI−LSLT−DSLT+GTLT​−LSDT+LGS+DG−GDT−DSDT​].(38),(39)

Formalization targets

Goal: Theorem 5.2 (corrected)

For all xxx and λ\lambdaλ:

(a)S∈S, G∈G, S≻0, GΔ skew ∀Δ∈D, F(λ,S,G,x)≻0 ⟹ λ>rD(A,b,x);\text{(a)}\quad S \in \mathcal S,\ G \in \mathcal G,\ S \succ 0,\ G\Delta \text{ skew } \forall \Delta \in \mathcal D,\ \mathcal F(\lambda,S,G,x) \succ 0 \ \Longrightarrow\ \lambda > r_{\mathcal D}(A,b,x);(a)S∈S, G∈G, S≻0, GΔ skew ∀Δ∈D, F(λ,S,G,x)≻0 ⟹ λ>rD​(A,b,x); (b)D=RN×N, λ>rD(A,b,x) ⟹ ∃s>0: F(λ,sI,0,x)≻0.\text{(b)}\quad \mathcal D = \mathbb R^{N\times N},\ \lambda > r_{\mathcal D}(A,b,x) \ \Longrightarrow\ \exists s > 0:\ \mathcal F(\lambda, sI, 0, x) \succ 0 .(b)D=RN×N, λ>rD​(A,b,x) ⟹ ∃s>0: F(λ,sI,0,x)≻0.

Part (a) says the value of the SDP inf⁡{λ:(λ,S,G) feasible}\inf\{\lambda : (\lambda, S, G) \text{ feasible}\}inf{λ:(λ,S,G) feasible} (40) is an upper bound on rDr_{\mathcal D}rD​. Part (b) says this upper bound is exact for full perturbations, including the case rD=∞r_{\mathcal D} = \inftyrD​=∞, where (40) is infeasible.

Milestones

  1. Lemma 2.2, both directions: the full-block S-procedure. det⁡(I−T4Δ)≠0\det(I - T_4\Delta) \ne 0det(I−T4​Δ)=0 and T(Δ)⪰0T(\Delta) \succeq 0T(Δ)⪰0 for all ∥Δ∥≤1\|\Delta\| \le 1∥Δ∥≤1 if and only if ∥T4∥<1\|T_4\| < 1∥T4​∥<1 and a one-scalar LMI (10) holds (the "only if" under T2≠0T_2 \ne 0T2​=0 or T3=0T_3 = 0T3​=0).
  2. Lemma 2.3: sufficiency of the scaled LMI for a structured D\mathcal DD, and its strict necessity for D=RN×N\mathcal D = \mathbb R^{N\times N}D=RN×N.
  3. §5.4, p. 1047: λ>rD(A,b,x)\lambda > r_{\mathcal D}(A,b,x)λ>rD​(A,b,x) if and only if a linear-fractional matrix function of Δ\DeltaΔ is positive definite on the structured unit ball.
  4. §5.4, (38)–(39): the certificate (a) in the paper's own words.

Significance

The worst-case residual under linear-fractional uncertainty cannot be computed efficiently unless P = NP. Theorem 5.2 gives an SDP-computable upper bound with an explicit certificate (S,G)(S, G)(S,G). Since xxx enters (38) linearly, the same constraint can also be optimized over xxx (Theorem 5.3, not part of this mission). For D=RN×N\mathcal D = \mathbb R^{N\times N}D=RN×N the bound is exact, which covers the model [A(Δ) b(Δ)]=[A b]+LΔ[RA Rb][A(\Delta)\ b(\Delta)] = [A\ b] + L\Delta[R_A\ R_b][A(Δ) b(Δ)]=[A b]+LΔ[RA​ Rb​] and, as a special case, the unstructured problem of §3.

The results are proved in the paper (the proof of Theorem 5.2 is only indicated, through Appendix C). No machine-checked version of these statements, of Lemma 2.2 or of the structured S-procedure with commutant scalings is known. The formalization also fixes the statements. As printed, Lemma 2.2's "only if", Lemma 2.3 and the upper bound of Theorem 5.2 are each false in a boundary or structural case (see Formalization scope). The corrected forms stated here are the ones the paper's proofs support.

Difficulty

Part (a) reduces to robust positivity of a linear-fractional matrix function, and the difficulty is the inverse (I−DΔ)−1(I - D\Delta)^{-1}(I−DΔ)−1. The certificate is one LMI in which Δ\DeltaΔ does not appear, while the conclusion is about a rational function of Δ\DeltaΔ over a whole structured ball. The certificate also has to guarantee that I−DΔI - D\DeltaI−DΔ is invertible everywhere on that ball, and not only that the residual is small where it is defined. Evaluating F\mathcal FF at a single point does not show this. Part (b) needs a lossless S-procedure in its strict form. The standard (non-strict) S-lemma gives only ⪰\succeq⪰, and the gap between strict and non-strict inequalities is exactly where the printed statements fail. The degenerate case T2=0T_2 = 0T2​=0 is not covered by the S-lemma's regularity condition and has to be handled separately.

Formalization scope

  • Dimensions are Fin n, Fin m, Fin N; D\mathcal DD is a Submodule ℝ (Matrix (Fin N) (Fin N) ℝ), with D=RN×N\mathcal D = \mathbb R^{N\times N}D=RN×N as ⊤. The Euclidean norm is written out, because ‖·‖ on Fin n → ℝ is the sup norm. ∥Δ∥\|\Delta\|∥Δ∥ is the operator norm of Matrix.toEuclideanLin Δ, the largest singular value.
  • λ>rD(A,b,x)\lambda > r_{\mathcal D}(A,b,x)λ>rD​(A,b,x) is the predicate ResidualBelow: every Δ∈D\Delta \in \mathcal DΔ∈D with ∥Δ∥≤1\|\Delta\| \le 1∥Δ∥≤1 has det⁡(I−DΔ)≠0\det(I - D\Delta) \ne 0det(I−DΔ)=0 and residual <λ< \lambda<λ. It is false for every λ\lambdaλ when rD=∞r_{\mathcal D} = \inftyrD​=∞. No real-valued supremum is used, so the ∞\infty∞ branch of (35) cannot turn into a default 000. Matrix inverses are Mathlib's Matrix.inv, and every use carries the determinant condition.
  • ρ=1\rho = 1ρ=1 throughout, as in the paper; general ρ\rhoρ follows by scaling Δ\DeltaΔ.
  • Corrections of the printed statements. (i) (40) must require S≻0S \succ 0S≻0. Without it, N=n=m=1N = n = m = 1N=n=m=1, D=2D = 2D=2, L=1L = 1L=1, A=b=RA=Rb=0A = b = R_A = R_b = 0A=b=RA​=Rb​=0, x=0x = 0x=0, S=−1S = -1S=−1, G=0G = 0G=0 satisfy (38) for every λ>1/3\lambda > 1/3λ>1/3, while rD=∞r_{\mathcal D} = \inftyrD​=∞. (ii) GGG must make GΔG\DeltaGΔ skew-symmetric for every Δ∈D\Delta \in \mathcal DΔ∈D, which is the identity pTGq=0p^TGq = 0pTGq=0 used in the proof of Lemma 2.3. For D=span⁡{I,J}\mathcal D = \operatorname{span}\{I, J\}D=span{I,J}, J=[01−10]J = \begin{bmatrix}0&1\\-1&0\end{bmatrix}J=[0−1​10​], the printed bound certifies λ=3/2\lambda = 3/2λ=3/2 for an instance with worst-case residual 222. The added condition holds automatically when every element of D\mathcal DD is symmetric (e.g. the diagonal structures (36)) and when G=0G = 0G=0 (e.g. D=RN×N\mathcal D = \mathbb R^{N\times N}D=RN×N). (iii) Lemma 2.2's "only if" is stated under T2≠0T_2 \ne 0T2​=0 or T3=0T_3 = 0T3​=0. (iv) Lemma 2.3's necessity is stated in strict form, and its sufficiency concludes T(Δ)≻0T(\Delta) \succ 0T(Δ)≻0.
  • Not stated: "If Θ>0\Theta > 0Θ>0 at the optimum, the upper bound is also exact". The infimum over the strict LMI (38) is not attained, and the paper does not say which limit is meant. Theorem 5.3, Lemma 2.4 and Lemma 5.1 are also not stated.
  • Trivializing encodings ruled out: the goal is not a statement about the value of an infimum (which a junk value could satisfy), and the added hypotheses are satisfiable (for instance S=sIS = sIS=sI, G=0G = 0G=0 for full D\mathcal DD, which part (b) produces).
  • Infrastructure needed: the Schur complement for block matrices (in Mathlib), a lossless S-lemma for two homogeneous quadratic forms in strict and non-strict form (the platform has ConvexOptimization.s_procedure, in a different sign convention), square roots of positive definite matrices that commute with D\mathcal DD, and compactness of the structured unit ball. The S-procedure lemmas are reusable in robust control and trust-region analysis. Proofs of the milestones in any order are welcome.

Selected references

  • L. El Ghaoui and H. Lebret, Robust solutions to least-squares problems with uncertain data, SIAM J. Matrix Anal. Appl. 18(4):1035–1064, 1997. https://doi.org/10.1137/S0895479896298130
  • S. Boyd, L. El Ghaoui, E. Feron and V. Balakrishnan, Linear Matrix Inequalities in System and Control Theory, SIAM, 1994. https://doi.org/10.1137/1.9781611970777
  • M. K. H. Fan, A. L. Tits and J. C. Doyle, Robustness in the presence of mixed parametric uncertainty and unmodeled dynamics, IEEE Trans. Automat. Control 36(1):25–38, 1991. https://doi.org/10.1109/9.62265
  • I. Pólik and T. Terlaky, A survey of the S-lemma, SIAM Review 49(3):371–418, 2007. https://doi.org/10.1137/S003614450444614X
8 thms2 active usersReviewed
🏆Completed
Convex OptimizationNumerical AnalysisOperations Research+1·Captain: mikedeng1

Robust Solutions to Least-Squares Problems with Uncertain Data II: Robust Least Squares as Tikhonov RegularizationResearch Paper

Motivation

Least squares fits a linear model Ax≃bAx \simeq bAx≃b by minimizing ∥Ax−b∥\|Ax - b\|∥Ax−b∥, and its solution can be extremely sensitive to errors in the data (A,b)(A, b)(A,b) when AAA is ill-conditioned. The standard remedy is Tikhonov regularization (ridge regression): minimize ∥Ax−b∥2+μ∥x∥2\|Ax - b\|^2 + \mu\|x\|^2∥Ax−b∥2+μ∥x∥2, whose solution x=(A⊤A+μI)−1A⊤bx = (A^\top A + \mu I)^{-1}A^\top bx=(A⊤A+μI)−1A⊤b is stable but depends on a parameter μ>0\mu > 0μ>0 that must be chosen by some external rule.

El Ghaoui and Lebret (SIAM J. Matrix Anal. Appl. 18(4), 1997) proposed instead to take the uncertainty in (A,b)(A, b)(A,b) seriously: the robust least-squares (RLS) solution minimizes the worst-case residual over all perturbations [ΔA Δb][\Delta A\ \Delta b][ΔA Δb] of Frobenius norm at most ρ\rhoρ. Their Theorem 3.1 shows that for ρ=1\rho = 1ρ=1 this worst-case residual equals ∥Ax−b∥+∥x∥2+1\|Ax - b\| + \sqrt{\|x\|^2 + 1}∥Ax−b∥+∥x∥2+1​ and that its minimization is the second-order cone program (15). Theorem 3.2, the subject of this mission, reads off the optimal solution: it is a Tikhonov-regularized solution, and the regularization parameter is not a free choice but is fixed by the data. This gives a principled answer to the question of how to choose μ\muμ, and it is the reason the paper describes RLS as "a Tikhonov regularization procedure" with "a rigorous way to compute the regularization parameter" (abstract, p. 1035).

A closely related model for least squares with bounded data uncertainty was developed at the same time by Chandrasekaran, Golub, Gu and Sayed; the paper notes that their preliminary draft (its reference [5]) gives a solution to the unstructured RLS problem similar to that of §3.2 (pp. 1036–1037).

Setting

Throughout, A∈Rn×mA \in \mathbb R^{n\times m}A∈Rn×m, b∈Rnb \in \mathbb R^nb∈Rn, x∈Rmx \in \mathbb R^mx∈Rm, and every vector norm is Euclidean, ∥v∥=∑ivi2\|v\| = \sqrt{\sum_i v_i^2}∥v∥=∑i​vi2​​. For x∈Rmx \in \mathbb R^mx∈Rm, [x;1]∈Rm+1[x; 1] \in \mathbb R^{m+1}[x;1]∈Rm+1 is xxx with a coordinate 111 appended, so ∥[x;1]∥=∥x∥2+1\|[x;1]\| = \sqrt{\|x\|^2 + 1}∥[x;1]∥=∥x∥2+1​.

The SOCP (15) is the problem, in the variables x∈Rmx \in \mathbb R^mx∈Rm and λ,τ∈R\lambda, \tau \in \mathbb Rλ,τ∈R,

minimize λsubject to∥Ax−b∥≤λ−τ,∥[x;1]∥≤τ.\text{minimize } \lambda \quad\text{subject to}\quad \|Ax - b\| \le \lambda - \tau,\qquad \|[x;1]\| \le \tau.minimize λsubject to∥Ax−b∥≤λ−τ,∥[x;1]∥≤τ.

A triple (x,λ,τ)(x, \lambda, \tau)(x,λ,τ) is optimal for (15) if it is feasible and λ≤λ′\lambda \le \lambda'λ≤λ′ for every feasible (x′,λ′,τ′)(x', \lambda', \tau')(x′,λ′,τ′). Its dual, derived in the paper from the general second-order cone duality of §2.1, is the problem in z∈Rnz \in \mathbb R^nz∈Rn, u∈Rmu \in \mathbb R^mu∈Rm, v∈Rv \in \mathbb Rv∈R

maximize b⊤z−vsubject toA⊤z+u=0,∥z∥≤1,∥[u;v]∥≤1.\text{maximize } b^\top z - v \quad\text{subject to}\quad A^\top z + u = 0,\quad \|z\| \le 1,\quad \|[u; v]\| \le 1.maximize b⊤z−vsubject toA⊤z+u=0,∥z∥≤1,∥[u;v]∥≤1.

The minimum-norm solution of Ax=bAx = bAx=b is a solution xxx with ∥x∥≤∥y∥\|x\| \le \|y\|∥x∥≤∥y∥ for every other solution yyy; when Ax=bAx = bAx=b is consistent it is A†bA^\dagger bA†b, with A†A^\daggerA† the Moore–Penrose pseudoinverse.

In the Lean development these objects are IsSOCPFeasible, IsSOCPOptimal, IsDualFeasible, dualObjective, IsDualOptimal and IsMinNormSolution, in the namespace RobustLS.Tikhonov, with the Euclidean norm eucNorm.

Formalization targets

Goal: Theorem 3.2 with the identity for μ\muμ

Let (x,λ,τ)(x, \lambda, \tau)(x,λ,τ) be optimal for (15) and set μ=(λ−τ)/τ\mu = (\lambda - \tau)/\tauμ=(λ−τ)/τ. Then

x={(μI+A⊤A)−1A⊤bif μ>0,A†belse,andμ=∥Ax−b∥∥x∥2+1.x = \begin{cases} (\mu I + A^\top A)^{-1}A^\top b & \text{if } \mu > 0,\\ A^\dagger b & \text{else,}\end{cases}\qquad\text{and}\qquad \mu = \frac{\|Ax - b\|}{\sqrt{\|x\|^2 + 1}}.x={(μI+A⊤A)−1A⊤bA†b​if μ>0,else,​andμ=∥x∥2+1​∥Ax−b∥​.

By Theorem 3.1 (the subject of the companion mission I of this series), the xxx-part of an optimal point of (15) is the RLS solution for ρ=1\rho = 1ρ=1, so this is formula (17) of the paper. The identity for μ\muμ is the final display of the paper's proof and is the claim in the mission's title.

Milestones (in the order of the paper's proof, p. 1041)

  1. Both (15) and its dual have optimal points.
  2. If λ=τ\lambda = \tauλ=τ at the optimum, then Ax=bAx = bAx=b and λ=τ=∥x∥2+1\lambda = \tau = \sqrt{\|x\|^2 + 1}λ=τ=∥x∥2+1​.
  3. In that case xxx is the minimum-norm solution of Ax=bAx = bAx=b, x=A†bx = A^\dagger bx=A†b.
  4. Eq. (18): for λ>τ\lambda > \tauλ>τ, primal and dual optimal values coincide,
∥Ax−b∥+∥[x;1]∥=λ=b⊤z−v=−(Ax−b)⊤z−[x⊤ 1][−A⊤zv].\|Ax - b\| + \|[x;1]\| = \lambda = b^\top z - v = -(Ax-b)^\top z - [x^\top\ 1]\begin{bmatrix} -A^\top z\\ v\end{bmatrix}.∥Ax−b∥+∥[x;1]∥=λ=b⊤z−v=−(Ax−b)⊤z−[x⊤ 1][−A⊤zv​].
  1. The dual optimal point is z=−(Ax−b)/∥Ax−b∥z = -(Ax - b)/\|Ax - b\|z=−(Ax−b)/∥Ax−b∥, [u;v]=−[x;1]/∥x∥2+1[u; v] = -[x; 1]/\sqrt{\|x\|^2 + 1}[u;v]=−[x;1]/∥x∥2+1​.
  2. Substituting into A⊤z+u=0A^\top z + u = 0A⊤z+u=0: x=(A⊤A+μI)−1A⊤bx = (A^\top A + \mu I)^{-1}A^\top bx=(A⊤A+μI)−1A⊤b with μ=(λ−τ)/τ=∥Ax−b∥/∥x∥2+1\mu = (\lambda - \tau)/\tau = \|Ax - b\|/\sqrt{\|x\|^2 + 1}μ=(λ−τ)/τ=∥Ax−b∥/∥x∥2+1​.

A further item states Remark 3.1: for λ>τ\lambda > \tauλ>τ, xxx is the unique minimizer of the weighted residual ∥[A;I;0]y−[b;0;1]∥Θ\big\|[A; I; 0]y - [b; 0; 1]\big\|_\Theta​[A;I;0]y−[b;0;1]​Θ​ with Θ=diag((λ−τ)I,τI,τ)\Theta = \mathbf{diag}((\lambda-\tau)I, \tau I, \tau)Θ=diag((λ−τ)I,τI,τ) and ∥r∥Θ=∥Θ−1/2r∥\|r\|_\Theta = \|\Theta^{-1/2} r\|∥r∥Θ​=∥Θ−1/2r∥.

Significance

The result. Theorem 3.2 turns a robust optimization problem into a familiar linear-algebra object. It says that the robust solution always lies on the Tikhonov path {(A⊤A+μI)−1A⊤b:μ>0}\{(A^\top A + \mu I)^{-1}A^\top b : \mu > 0\}{(A⊤A+μI)−1A⊤b:μ>0} or at its endpoint A†bA^\dagger bA†b, and it identifies the point on the path through a fixed-point equation relating μ\muμ to the residual and the size of the solution. The paper builds on this in §3.3 (a one-dimensional search for μ\muμ via the SVD) and in §6 (continuity of the RLS solution in the data), and Remark 3.1 is the template for the weighted least-squares interpretation of the structured and linear-fractional problems in §5.

Formalizing it. The theorem is proved in the paper; to our knowledge it has no machine-checked proof. The mission produces a formal account of second-order cone duality for a concrete program, the characterization of the optimal dual point by equality in the Cauchy–Schwarz inequality, and the minimum-norm characterization of A†bA^\dagger bA†b, all in terms of explicit Euclidean norms on Fin k → ℝ.

Difficulty

The paper's proof rests on strong duality for (15) ("both primal and dual problems are strictly feasible"), which it cites from the SOCP literature rather than proving; Mathlib has no second-order cone duality, so this step is the main gap. The degenerate case λ=τ\lambda = \tauλ=τ also needs care: there ∥Ax−b∥=0\|Ax - b\| = 0∥Ax−b∥=0, the residual term is not differentiable at the optimum, and the conclusion changes from a regularized inverse to a pseudoinverse. A statement that only handles the case Ax≠bAx \ne bAx=b, or that assumes the matrix A⊤A+μIA^\top A + \mu IA⊤A+μI invertible without deriving it from μ>0\mu > 0μ>0, misses part of the theorem.

Formalization scope

  • Normalization. The paper states Theorem 3.2 for ρ=1\rho = 1ρ=1 ("we take ρ=1\rho = 1ρ=1 in what follows", p. 1039) and obtains general ρ\rhoρ by the scaling φ(A,b,ρ)=ρ φ(A/ρ,b/ρ,1)\varphi(A, b, \rho) = \rho\,\varphi(A/\rho, b/\rho, 1)φ(A,b,ρ)=ρφ(A/ρ,b/ρ,1). Only the ρ=1\rho = 1ρ=1 statement is formalized.
  • The RLS solution. The perturbation model is not used here: all statements are about optimal points of (15). That the xxx-part of such a point is the RLS solution is Theorem 3.1 (mission I), and it is recalled in prose only.
  • Norms. Vectors are Fin k → ℝ; the Euclidean norm is the explicit eucNorm v = √(∑ vᵢ²) (Mathlib's ‖·‖ on Fin k → ℝ is the sup norm). Stacked vectors [x;1][x;1][x;1] and [u;v][u;v][u;v] are indexed by Fin m ⊕ Unit.
  • Optimality. "Optimal point" means feasible with objective no worse than every feasible point; the minimum and maximum are therefore attained by definition, and milestone 1 guarantees they exist.
  • Pseudoinverse. Mathlib has no matrix pseudoinverse, so A†bA^\dagger bA†b is stated as the minimum-norm solution of Ax=bAx = bAx=b, which is how the proof uses it. The branch "else" is ¬(μ>0)\neg(\mu > 0)¬(μ>0).
  • Inverse. (μI+A⊤A)−1(\mu I + A^\top A)^{-1}(μI+A⊤A)−1 is Mathlib's Matrix.inv; it is used only where μ>0\mu > 0μ>0, where the matrix is positive definite. τ≥1\tau \ge 1τ≥1 at every feasible point, so μ\muμ is well defined without an extra hypothesis.
  • No trivialization. The goal quantifies over optimal points of (15) over the whole feasible set, not over feasible points, and milestone 1 shows the hypothesis is satisfiable for every (A,b)(A, b)(A,b), including n=0n = 0n=0 or m=0m = 0m=0.
  • Weighted norm. For Remark 3.1, ∥r∥Θ\|r\|_\Theta∥r∥Θ​ for the diagonal Θ\ThetaΘ is written as ∑iri2/θi\sqrt{\sum_i r_i^2/\theta_i}∑i​ri2​/θi​​, which equals ∥Θ−1/2r∥\|\Theta^{-1/2}r\|∥Θ−1/2r∥ for positive weights.

Contributions welcome: second-order cone (or general conic) weak and strong duality for finite-dimensional programs, the equality case of Cauchy–Schwarz in the explicit-norm form used here, and a Moore–Penrose pseudoinverse for real matrices with its minimum-norm property. The platform's ConvexOptimization.conic_slater_strong_duality may help with the duality step.

Selected references

  • L. El Ghaoui and H. Lebret, Robust Solutions to Least-Squares Problems with Uncertain Data, SIAM J. Matrix Anal. Appl. 18(4):1035–1064, 1997. https://doi.org/10.1137/S0895479896298130
  • S. Chandrasekaran, G. H. Golub, M. Gu and A. H. Sayed, A new linear least-squares type model for parameter estimation in the presence of data uncertainties, cited as submitted to SIAM J. Matrix Anal. Appl. (reference [5] of the paper).
  • A. N. Tikhonov and V. Y. Arsenin, Solutions of Ill-Posed Problems, Wiley, New York, 1977 (reference [43] of the paper).
  • Y. Nesterov and A. Nemirovskii, Interior-Point Polynomial Algorithms in Convex Programming, SIAM, 1994. https://doi.org/10.1137/1.9781611970791
  • M. S. Lobo, L. Vandenberghe, S. Boyd and H. Lebret, Applications of Second-Order Cone Programming, Linear Algebra Appl. 284:193–228, 1998. https://doi.org/10.1016/S0024-3795(98)10032-0
9 thms2 active usersReviewed
🏆Completed
Convex OptimizationOptimization·Captain: mikedeng1

A Singular Value Thresholding Algorithm for Matrix Completion 3: Convergence to the Minimum Nuclear Norm SolutionResearch Paper

Motivation

Nuclear norm minimization is the standard convex surrogate for rank minimization: to recover a low-rank matrix from a few linear measurements, or from a subset of its entries, one minimizes the sum of the singular values subject to the data constraints. For matrix completion, Candès and Recht (Found. Comput. Math. 2009) showed that this convex program recovers a low-rank matrix exactly from sufficiently many random entries. Solving it at scale is another matter: interior-point methods for the equivalent semidefinite program become impractical beyond matrices of a few hundred rows and columns.

Cai, Candès and Shen (SIAM J. Optim. 2010) proposed the singular value thresholding (SVT) algorithm, whose iterates are cheap and typically of low rank. SVT does not solve the nuclear norm problem itself. It solves a proximal problem, in which the nuclear norm is replaced by τ∥X∥∗+12∥X∥F2\tau\|X\|_* + \tfrac12\|X\|_F^2τ∥X∥∗​+21​∥X∥F2​ for a fixed parameter τ>0\tau>0τ>0. Section 3.4 of the paper justifies this substitution: as τ→∞\tau\to\inftyτ→∞, the solutions of the proximal problem converge to a specific solution of the nuclear norm problem, the one of least Frobenius norm. This mission formalizes that result, Theorem 3.1 of the paper, under general convex constraints.

Setting

Let n1,n2n_1, n_2n1​,n2​ be natural numbers and Rn1×n2\mathbb R^{n_1\times n_2}Rn1​×n2​ the space of real n1×n2n_1\times n_2n1​×n2​ matrices, with the inner product ⟨X,Y⟩=trace⁡(X∗Y)=∑i,jXijYij\langle X, Y\rangle = \operatorname{trace}(X^*Y) = \sum_{i,j}X_{ij}Y_{ij}⟨X,Y⟩=trace(X∗Y)=∑i,j​Xij​Yij​. Three functions of a matrix XXX are used:

  • the Frobenius norm ∥X∥F=⟨X,X⟩\|X\|_F = \sqrt{\langle X, X\rangle}∥X∥F​=⟨X,X⟩​;
  • the nuclear norm ∥X∥∗\|X\|_*∥X∥∗​, the sum of the singular values of XXX;
  • for a parameter τ\tauτ, the proximal objective fτ(X)=τ∥X∥∗+12∥X∥F2f_\tau(X) = \tau\|X\|_* + \tfrac12\|X\|_F^2fτ​(X)=τ∥X∥∗​+21​∥X∥F2​.

Let f1,…,fm:Rn1×n2→Rf_1,\dots,f_m:\mathbb R^{n_1\times n_2}\to\mathbb Rf1​,…,fm​:Rn1​×n2​→R be constraint functions and C={X:fi(X)≤0, i=1,…,m}\mathcal C = \{X : f_i(X)\le 0,\ i = 1,\dots,m\}C={X:fi​(X)≤0, i=1,…,m} the feasible set. The nuclear norm problem is

(1.6)minimize ∥X∥∗subject to fi(X)≤0, i=1,…,m,\text{(1.6)}\qquad \text{minimize } \|X\|_* \quad \text{subject to } f_i(X)\le 0,\ i=1,\dots,m,(1.6)minimize ∥X∥∗​subject to fi​(X)≤0, i=1,…,m,

and, for τ>0\tau>0τ>0, the proximal problem is

(3.4)minimize fτ(X)subject to fi(X)≤0, i=1,…,m.\text{(3.4)}\qquad \text{minimize } f_\tau(X) \quad \text{subject to } f_i(X)\le 0,\ i=1,\dots,m.(3.4)minimize fτ​(X)subject to fi​(X)≤0, i=1,…,m.

When the fif_ifi​ are convex and C\mathcal CC is nonempty, (3.4) has exactly one solution, written Xτ⋆X^\star_\tauXτ⋆​, because fτf_\taufτ​ is strongly convex. Problem (1.6) may have many solutions. Among them, the paper singles out the minimum Frobenius norm solution

(3.14)X∞:=arg⁡min⁡X{∥X∥F2:X is a solution of (1.6)}.\text{(3.14)}\qquad X_\infty := \arg\min_X\{\|X\|_F^2 : X \text{ is a solution of (1.6)}\}.(3.14)X∞​:=argXmin​{∥X∥F2​:X is a solution of (1.6)}.

Linear equality constraints, and in particular the matrix completion constraints Xij=MijX_{ij} = M_{ij}Xij​=Mij​ for sampled entries (i,j)(i,j)(i,j), are covered by taking pairs of affine functionals.

Formalization targets

Goal: Theorem 3.1

Assume that the fif_ifi​ are convex and lower semicontinuous. Then

(3.15)lim⁡τ→∞∥Xτ⋆−X∞∥F=0.\text{(3.15)}\qquad \lim_{\tau\to\infty}\|X^\star_\tau - X_\infty\|_F = 0.(3.15)τ→∞lim​∥Xτ⋆​−X∞​∥F​=0.

Milestones

In the order in which the paper's proof uses them (all on p. 1967):

  1. Eq. (3.16), for every τ>0\tau>0τ>0:
∥Xτ⋆∥∗+12τ∥Xτ⋆∥F2≤∥X∞∥∗+12τ∥X∞∥F2and∥X∞∥∗≤∥Xτ⋆∥∗.\|X^\star_\tau\|_* + \frac{1}{2\tau}\|X^\star_\tau\|_F^2 \le \|X_\infty\|_* + \frac{1}{2\tau}\|X_\infty\|_F^2 \quad\text{and}\quad \|X_\infty\|_*\le\|X^\star_\tau\|_*.∥Xτ⋆​∥∗​+2τ1​∥Xτ⋆​∥F2​≤∥X∞​∥∗​+2τ1​∥X∞​∥F2​and∥X∞​∥∗​≤∥Xτ⋆​∥∗​.
  1. Eq. (3.17), for every τ>0\tau>0τ>0: ∥Xτ⋆∥F2≤∥X∞∥F2\|X^\star_\tau\|_F^2 \le \|X_\infty\|_F^2∥Xτ⋆​∥F2​≤∥X∞​∥F2​.
  2. Convergence of the nuclear norms: lim⁡τ→∞∥Xτ⋆∥∗=∥X∞∥∗\lim_{\tau\to\infty}\|X^\star_\tau\|_* = \|X_\infty\|_*limτ→∞​∥Xτ⋆​∥∗​=∥X∞​∥∗​.
  3. Uniqueness of X∞X_\inftyX∞​: two minimum Frobenius norm solutions of (1.6) coincide when the fif_ifi​ are convex.
  4. Cluster points: if τk→∞\tau_k\to\inftyτk​→∞ and Xτk⋆→XcX^\star_{\tau_k}\to X_cXτk​⋆​→Xc​, then Xc=X∞X_c = X_\inftyXc​=X∞​.

Significance

The result itself. Theorem 3.1 is the link between the problem SVT actually solves and the problem one wants solved. The companion missions of this series prove that the SVT iteration, and its variant for general convex constraints, converges to Xτ⋆X^\star_\tauXτ⋆​. Theorem 3.1 says what Xτ⋆X^\star_\tauXτ⋆​ is worth: for large τ\tauτ it is close to a nuclear norm minimizer, and the minimizer it approaches is identified exactly, namely the one of least Frobenius norm. The statement is not specific to matrix completion. It covers every finite family of convex, lower semicontinuous constraints, and hence noisy variants such as the inequality-constrained problems of §3.3 of the paper.

Formalizing it. The theorem is proved in the paper, in about half a page. It has not, to our knowledge, been machine-checked. A formal proof pins down the hypotheses: the argument needs the minimizers to exist, and it uses continuity and convexity of the nuclear norm, closedness of the feasible set, and uniqueness of X∞X_\inftyX∞​. It also produces a reusable fact about the nuclear norm in Lean, namely that the sum of singular values is a continuous convex function of the matrix.

Difficulty

The first steps are elementary consequences of the definitions of Xτ⋆X^\star_\tauXτ⋆​ and X∞X_\inftyX∞​: (3.16) compares objective values, and (3.17) and the convergence of the nuclear norms follow by algebra and a squeeze. The difficulty lies elsewhere.

  • Identifying the limit. Boundedness gives cluster points of Xτ⋆X^\star_\tauXτ⋆​, not convergence. Each cluster point must be shown to be feasible, to be optimal for (1.6), and to have the least Frobenius norm among the optimal points. Feasibility uses lower semicontinuity of the constraints. Optimality uses continuity of the nuclear norm. Minimality uses (3.17) passed to the limit.
  • Uniqueness of X∞X_\inftyX∞​. The last step concludes Xc=X∞X_c = X_\inftyXc​=X∞​ from ∥Xc∥F=∥X∞∥F\|X_c\|_F = \|X_\infty\|_F∥Xc​∥F​=∥X∞​∥F​, which needs uniqueness of the minimum Frobenius norm solution. That in turn needs convexity of the solution set of (1.6), hence convexity of the nuclear norm, together with strict convexity of ∥⋅∥F2\|\cdot\|_F^2∥⋅∥F2​.
  • Nuclear norm in Lean. The nuclear norm is defined from singular values, and its convexity (the triangle inequality for the sum of singular values) and continuity are not currently available as ready-made statements. They are the main groundwork.

A tempting shortcut, reading the family Xτ⋆X^\star_\tauXτ⋆​ as a sequence indexed by integers, proves a weaker statement: the limit in (3.15) is over real τ→∞\tau\to\inftyτ→∞.

Formalization scope

  • Matrices. Matrices are Matrix (Fin n₁) (Fin n₂) ℝ, abbreviated Mat n₁ n₂, over the reals as in the paper. ⟨X,Y⟩=∑i,jXijYij\langle X,Y\rangle = \sum_{i,j}X_{ij}Y_{ij}⟨X,Y⟩=∑i,j​Xij​Yij​ and ∥X∥F=⟨X,X⟩\|X\|_F = \sqrt{\langle X,X\rangle}∥X∥F​=⟨X,X⟩​.
  • Nuclear norm. ∥X∥∗\|X\|_*∥X∥∗​ is the sum of Mathlib's LinearMap.singularValues of Matrix.toEuclideanLin X. It is the genuine sum of singular values, not an abstract norm or the Frobenius norm.
  • Constraints. The constraints are a family f : Fin m → Mat n₁ n₂ → ℝ of real-valued functions. m=0m = 0m=0 (no constraints) is allowed.
  • Hypotheses of Theorem 3.1. The hypotheses are ConvexOn ℝ Set.univ (f i) and LowerSemicontinuous (f i) for every iii. Lower semicontinuity is redundant for real-valued convex functions on a finite-dimensional space, but it is kept because the theorem states it.
  • Xτ⋆X^\star_\tauXτ⋆​ and X∞X_\inftyX∞​. Xτ⋆X^\star_\tauXτ⋆​ is a family Xτ : ℝ → Mat n₁ n₂ assumed to solve (3.4) for every τ>0\tau>0τ>0, and its values at τ≤0\tau\le 0τ≤0 play no role. X∞X_\inftyX∞​ is a matrix assumed to satisfy the defining property (3.14): it solves (1.6) and has the least ∥⋅∥F2\|\cdot\|_F^2∥⋅∥F2​ among its solutions. Uniqueness of X∞X_\inftyX∞​ is a milestone to prove, not an assumption.
  • Vacuous case. These hypotheses presuppose, as the paper does, that (1.6) has a solution. They can be met exactly when the feasible set is nonempty. When it is empty the statement is vacuous, which matches the paper, where X∞X_\inftyX∞​ is then undefined.
  • Limits and topology. Limits in τ\tauτ are along Filter.atTop on R\mathbb RR. Convergence of matrices uses Mathlib's entrywise topology, which is the topology of ∥⋅∥F\|\cdot\|_F∥⋅∥F​. The goal states (3.15) literally, with the Frobenius norm of the difference tending to 000.
  • Excluded shortcuts. A formalization that replaces the nuclear norm by the Frobenius norm or by an arbitrary norm, indexes τ\tauτ by N\mathbb NN, or assumes uniqueness or convergence as a hypothesis would not be Theorem 3.1. It is ruled out.

Infrastructure. The needed facts, all reusable beyond this mission:

  • nonnegativity, convexity and continuity of the nuclear norm on real matrices;
  • closedness and convexity of sublevel sets of convex lower semicontinuous functions;
  • uniqueness of the minimizer of a strictly convex function over a convex set;
  • a cluster-point argument for bounded families in finite-dimensional spaces.

Contributions of these general lemmas as separate theorems are welcome.

Selected references

  • J.-F. Cai, E. J. Candès, Z. Shen, A Singular Value Thresholding Algorithm for Matrix Completion, SIAM J. Optim. 20(4):1956–1982, 2010. https://doi.org/10.1137/080738970
  • E. J. Candès, B. Recht, Exact Matrix Completion via Convex Optimization, Found. Comput. Math. 9:717–772, 2009. https://doi.org/10.1007/s10208-009-9045-5
  • B. Recht, M. Fazel, P. A. Parrilo, Guaranteed Minimum-Rank Solutions of Linear Matrix Equations via Nuclear Norm Minimization, SIAM Rev. 52(3):471–501, 2010. https://doi.org/10.1137/070697835
8 thms2 active usersReviewed
🏆Completed
Convex OptimizationNumerical AnalysisOptimization·Captain: mikedeng1

A Singular Value Thresholding Algorithm for Matrix Completion 2: Convergence of the SVT Iteration under General Convex ConstraintsResearch Paper

Motivation

Singular value thresholding (SVT) is a first-order method introduced by Cai, Candès and Shen (SIAM J. Optim. 20 (2010)) for recovering a low-rank matrix from incomplete or indirect information. Its basic form, for matrix completion, alternates a soft-thresholding of singular values with a gradient step on a dual variable, and needs only one sparse singular value decomposition per iteration. That is what made nuclear-norm heuristics usable on matrices with tens of thousands of rows and columns, where interior-point methods for the equivalent semidefinite program do not fit in memory.

Matrix completion is only one constraint set. In applications the data are noisy linear measurements b=A(M)+zb = \mathcal A(M) + zb=A(M)+z, and the constraint takes the form of componentwise error bounds or norm balls around the data (§3.3 of the paper). Section 3.2 of the paper extends the method to a general finite family of convex constraints, and §4.2 proves that the extended iteration converges. This mission formalizes that extension and its convergence theorem, Theorem 4.4.

Setting

Let n1,n2,mn_1, n_2, mn1​,n2​,m be natural numbers and Rn1×n2\mathbb R^{n_1\times n_2}Rn1​×n2​ the real n1×n2n_1\times n_2n1​×n2​ matrices, with the Frobenius inner product ⟨X,Y⟩=∑i,jXijYij\langle X, Y\rangle = \sum_{i,j} X_{ij}Y_{ij}⟨X,Y⟩=∑i,j​Xij​Yij​ and norm ∥X∥F=⟨X,X⟩\|X\|_F = \sqrt{\langle X, X\rangle}∥X∥F​=⟨X,X⟩​. The nuclear norm ∥X∥∗\|X\|_*∥X∥∗​ is the sum of the singular values of XXX. For a fixed τ>0\tau > 0τ>0 the objective is

fτ(X)=τ∥X∥∗+12∥X∥F2.f_\tau(X) = \tau\|X\|_* + \tfrac12\|X\|_F^2 .fτ​(X)=τ∥X∥∗​+21​∥X∥F2​.

A matrix ZZZ is a subgradient of a function ggg at X0X_0X0​, written Z∈∂g(X0)Z\in\partial g(X_0)Z∈∂g(X0​), if g(X)≥g(X0)+⟨Z,X−X0⟩g(X)\ge g(X_0) + \langle Z, X - X_0\rangleg(X)≥g(X0​)+⟨Z,X−X0​⟩ for all XXX.

Let f1,…,fm:Rn1×n2→Rf_1,\dots,f_m:\mathbb R^{n_1\times n_2}\to\mathbb Rf1​,…,fm​:Rn1​×n2​→R be convex and put F(X)=(f1(X),…,fm(X))∈Rm\mathcal F(X) = (f_1(X),\dots,f_m(X))\in\mathbb R^mF(X)=(f1​(X),…,fm​(X))∈Rm. On Rm\mathbb R^mRm, ⟨u,v⟩=∑iuivi\langle u, v\rangle = \sum_i u_iv_i⟨u,v⟩=∑i​ui​vi​ and ∥v∥\|v\|∥v∥ is the Euclidean norm. The constrained problem is

(3.4)minimize fτ(X)subject to fi(X)≤0, i=1,…,m,\text{(3.4)}\qquad \text{minimize } f_\tau(X)\quad\text{subject to } f_i(X)\le 0,\ i=1,\dots,m,(3.4)minimize fτ​(X)subject to fi​(X)≤0, i=1,…,m,

with Lagrangian L(X,y)=fτ(X)+⟨y,F(X)⟩\mathcal L(X, y) = f_\tau(X) + \langle y, \mathcal F(X)\rangleL(X,y)=fτ​(X)+⟨y,F(X)⟩ for y≥0y\ge 0y≥0. A pair (X⋆,y⋆)(X^\star, y^\star)(X⋆,y⋆) with y⋆≥0y^\star\ge0y⋆≥0 is primal-dual optimal if it is a saddle point:

L(X⋆,y)≤L(X⋆,y⋆)≤L(X,y⋆)for all y≥0, X.\mathcal L(X^\star, y)\le \mathcal L(X^\star, y^\star)\le \mathcal L(X, y^\star)\qquad\text{for all } y\ge 0,\ X .L(X⋆,y)≤L(X⋆,y⋆)≤L(X,y⋆)for all y≥0, X.

The paper's standing assumption "strong duality holds" is the existence of such a pair.

The iteration (3.5) starts from y0=0y^0 = 0y0=0 and, for step sizes δk\delta_kδk​, sets for k=1,2,…k = 1, 2, \dotsk=1,2,…

Xk=arg⁡min⁡X{fτ(X)+⟨yk−1,F(X)⟩},yk=[ yk−1+δkF(Xk) ]+,X^k = \arg\min_X\{f_\tau(X) + \langle y^{k-1}, \mathcal F(X)\rangle\},\qquad y^k = [\,y^{k-1} + \delta_k\mathcal F(X^k)\,]_+ ,Xk=argXmin​{fτ​(X)+⟨yk−1,F(X)⟩},yk=[yk−1+δk​F(Xk)]+​,

where x+x_+x+​ has entries max⁡(xi,0)\max(x_i, 0)max(xi​,0). It is Uzawa's method for (3.4): an exact minimization in the primal variable followed by a projected ascent step on the dual. When F(X)=b−A(X)\mathcal F(X) = b - \mathcal A(X)F(X)=b−A(X) is affine, the minimization is a singular value thresholding step, which gives the algorithm its name.

The analysis of §4.2 assumes F\mathcal FF is Lipschitz in the sense

(4.2)∥F(X)−F(Y)∥≤L ∥X−Y∥Ffor all X,Y,\text{(4.2)}\qquad \|\mathcal F(X) - \mathcal F(Y)\|\le L\,\|X - Y\|_F\quad\text{for all } X, Y,(4.2)∥F(X)−F(Y)∥≤L∥X−Y∥F​for all X,Y,

for a constant L≥0L\ge 0L≥0.

Formalization targets

Goal: Theorem 4.4 (p. 1969)

If 0<inf⁡kδk≤sup⁡kδk<2/L20 < \inf_k\delta_k\le\sup_k\delta_k < 2/L^20<infk​δk​≤supk​δk​<2/L2 and strong duality holds, then the sequence XkX^kXk of (3.5) converges to the unique solution of (3.4):

∃! X⋆ solving (3.4),lim⁡k→∞Xk=X⋆.\exists!\,X^\star\ \text{solving (3.4)},\qquad \lim_{k\to\infty} X^k = X^\star .∃!X⋆ solving (3.4),k→∞lim​Xk=X⋆.

Milestones, in the order the proof uses them

  • Lemma 4.1 (p. 1968): ⟨Z−Z′,X−X′⟩≥∥X−X′∥F2\langle Z - Z', X - X'\rangle\ge\|X - X'\|_F^2⟨Z−Z′,X−X′⟩≥∥X−X′∥F2​ for Z∈∂fτ(X)Z\in\partial f_\tau(X)Z∈∂fτ​(X), Z′∈∂fτ(X′)Z'\in\partial f_\tau(X')Z′∈∂fτ​(X′).
  • Lemma 4.3 (p. 1969): for a primal-dual optimal pair and each δ>0\delta > 0δ>0, y⋆=[y⋆+δF(X⋆)]+y^\star = [y^\star + \delta\mathcal F(X^\star)]_+y⋆=[y⋆+δF(X⋆)]+​.
  • Eq. (4.4) (p. 1969): there are Zk∈∂fτ(Xk)Z^k\in\partial f_\tau(X^k)Zk∈∂fτ​(Xk) and Z⋆∈∂fτ(X⋆)Z^\star\in\partial f_\tau(X^\star)Z⋆∈∂fτ​(X⋆) with ⟨Zk,X−Xk⟩+⟨yk−1,F(X)−F(Xk)⟩≥0\langle Z^k, X - X^k\rangle + \langle y^{k-1}, \mathcal F(X) - \mathcal F(X^k)\rangle\ge 0⟨Zk,X−Xk⟩+⟨yk−1,F(X)−F(Xk)⟩≥0 and ⟨Z⋆,X−X⋆⟩+⟨y⋆,F(X)−F(X⋆)⟩≥0\langle Z^\star, X - X^\star\rangle + \langle y^\star, \mathcal F(X) - \mathcal F(X^\star)\rangle\ge 0⟨Z⋆,X−X⋆⟩+⟨y⋆,F(X)−F(X⋆)⟩≥0 for all XXX.
  • Eq. (4.5) (p. 1969): ⟨yk−1−y⋆,F(Xk)−F(X⋆)⟩≤−∥Xk−X⋆∥F2\langle y^{k-1} - y^\star, \mathcal F(X^k) - \mathcal F(X^\star)\rangle\le -\|X^k - X^\star\|_F^2⟨yk−1−y⋆,F(Xk)−F(X⋆)⟩≤−∥Xk−X⋆∥F2​.
  • Contraction step (p. 1969): ∥yk−y⋆∥≤∥yk−1−y⋆+δk(F(Xk)−F(X⋆))∥\|y^k - y^\star\|\le\|y^{k-1} - y^\star + \delta_k(\mathcal F(X^k) - \mathcal F(X^\star))\|∥yk−y⋆∥≤∥yk−1−y⋆+δk​(F(Xk)−F(X⋆))∥.
  • Eq. (4.6) (p. 1970): if 2δk−δk2L2≥β>02\delta_k - \delta_k^2L^2\ge\beta > 02δk​−δk2​L2≥β>0 for k≥1k\ge1k≥1, then ∥yk−y⋆∥2≤∥yk−1−y⋆∥2−β∥Xk−X⋆∥F2\|y^k - y^\star\|^2\le\|y^{k-1} - y^\star\|^2 - \beta\|X^k - X^\star\|_F^2∥yk−y⋆∥2≤∥yk−1−y⋆∥2−β∥Xk−X⋆∥F2​.

Significance

The result. Theorem 4.4 is the convergence guarantee for SVT beyond matrix completion. The componentwise error bounds of (3.8), whose SVT iteration is (3.9), are finitely many affine constraints and fall under it directly, as does any finite family of Lipschitz convex constraints, for instance a Frobenius-norm ball around the data. The conic variants of §3.3 ((3.11)–(3.13)) project the dual variable onto a cone rather than onto the nonnegative orthant and are not covered by the theorem as stated. Together with Theorem 3.1 of the same paper, which says that the solution of (3.4) tends to the minimum-nuclear-norm solution as τ→∞\tau\to\inftyτ→∞, it justifies using SVT as a solver for nuclear-norm minimization under general convex constraints.

Formalizing it. The theorem is proved in the paper, with two steps delegated to the literature: Lemma 4.3 cites [31], and the concluding step reads "the conclusion is as before". Its proof is short but relies on convex-analytic facts that are standard on paper and missing, in this form, from Mathlib: subgradients of the nuclear norm, the subdifferential sum rule for finite convex functions, and nonexpansiveness of the projection onto the nonnegative orthant. No machine-checked proof of this theorem or of Uzawa-type convergence for nuclear-norm objectives is known to exist. The mission produces a complete, checked version of the argument, including the omitted closing step.

Difficulty

The obvious approach is to view (3.5) as projected gradient ascent on the dual function g(y)=min⁡XL(X,y)g(y) = \min_X\mathcal L(X, y)g(y)=minX​L(X,y) and quote the standard convergence theorem for gradient methods with Lipschitz gradients. That does not apply directly: for general convex fif_ifi​ the dual function need not be differentiable, F(Xk)\mathcal F(X^k)F(Xk) is only a supergradient, and the Lipschitz hypothesis (4.2) is on F\mathcal FF, not on a dual gradient. The proof instead works with the primal-dual pair: it needs first-order optimality conditions (4.4), which require a subdifferential sum rule for fτ+∑iyifif_\tau + \sum_i y_i f_ifτ​+∑i​yi​fi​ with nonsmooth fif_ifi​, and it needs the strong monotonicity of ∂fτ\partial f_\tau∂fτ​ (Lemma 4.1), which depends on the description of subgradients of the nuclear norm. A second subtlety is that the theorem asserts convergence of the whole primal sequence to the unique solution, not to some solution along a subsequence, while nothing is claimed about convergence of the dual sequence.

Formalization scope

Matrices are Matrix (Fin n₁) (Fin n₂) ℝ, vectors in Rm\mathbb R^mRm are Fin m → ℝ, and convergence of matrices is in Mathlib's product topology, which coincides with the Frobenius topology. The nuclear norm is the sum of Mathlib's LinearMap.singularValues of the matrix viewed as a map between Euclidean spaces. Each fif_ifi​ is a real-valued function with ConvexOn ℝ Set.univ. The iteration is a predicate on sequences indexed by ℕ: the paper's step kkk produces X (k+1) and y (k+1) from y k with step size δ (k+1), and y 0 = 0. XkX^kXk is required to minimize L(⋅,yk−1)\mathcal L(\cdot, y^{k-1})L(⋅,yk−1); for τ>0\tau>0τ>0 and convex fif_ifi​ this minimizer exists and is unique, so the predicate is satisfiable and determines the sequence. The step-size condition is stated as a≤δk≤Ca\le\delta_k\le Ca≤δk​≤C for k≥1k\ge1k≥1 with a>0a>0a>0 and CL2<2C L^2 < 2CL2<2, which avoids the division 2/L22/L^22/L2 (evaluated as 000 in Lean when L=0L=0L=0); for L=0L=0L=0 it requires only bounded steps, matching the convention 2/0=∞2/0 = \infty2/0=∞. Strong duality is the hypothesis that a saddle point exists; Slater's condition is not assumed. The paper's standing assumptions (τ>0\tau>0τ>0, convex fif_ifi​, and (4.2) where LLL enters) appear as explicit hypotheses in every statement.

A formalization that assumes convergence or boundedness of the dual iterates, replaces the primal minimization by a closed-form thresholding step (valid only for affine F\mathcal FF), or states only subsequential convergence would not be this theorem; each of these is excluded by the statements above.

A complete development needs: subgradients of the nuclear norm and strong monotonicity of ∂fτ\partial f_\tau∂fτ​; existence and characterization of minimizers of strongly convex continuous functions on a finite-dimensional space; the subdifferential sum rule for finite convex functions; complementary slackness from the saddle-point inequalities; and nonexpansiveness of the entrywise positive part. These are reusable beyond this mission, especially for other Uzawa and augmented Lagrangian analyses. Contributions of any of these pieces as separate lemmas are welcome.

Selected references

  • J.-F. Cai, E. J. Candès, Z. Shen, A Singular Value Thresholding Algorithm for Matrix Completion, SIAM J. Optim. 20(4):1956–1982, 2010. https://doi.org/10.1137/080738970
  • E. J. Candès, B. Recht, Exact Matrix Completion via Convex Optimization, Found. Comput. Math. 9:717–772, 2009. https://doi.org/10.1007/s10208-009-9045-5
  • K. J. Arrow, L. Hurwicz, H. Uzawa, Studies in Linear and Nonlinear Programming, Stanford University Press, 1958.
  • S. Boyd, L. Vandenberghe, Convex Optimization, Cambridge University Press, 2004. https://doi.org/10.1017/CBO9780511804441
9 thms2 active usersReviewed
🏆Completed
Convex OptimizationNumerical AnalysisOptimization·Captain: mikedeng1

A Singular Value Thresholding Algorithm for Matrix Completion 1: The SVT Iteration Converges to the Unique Solution of the Proximal ProblemResearch Paper

Motivation

Matrix completion asks to recover an n1×n2n_1\times n_2n1​×n2​ matrix MMM from a subset Ω\OmegaΩ of its entries. When MMM has low rank, a standard convex surrogate is to minimize the nuclear norm ∥X∥∗\|X\|_*∥X∥∗​ (the sum of the singular values) subject to agreeing with MMM on Ω\OmegaΩ; Candès and Recht showed that this recovers MMM exactly under incoherence and sampling conditions (Candès–Recht 2009). Generic interior-point solvers for this semidefinite program do not scale beyond matrices of a few hundred rows.

Cai, Candès and Shen (SIAM J. Optim. 2010) proposed the singular value thresholding (SVT) algorithm: a first-order iteration whose only nonlinear step is a soft-thresholding of singular values, and whose other iterate is a sparse matrix supported on Ω\OmegaΩ. The algorithm has become a standard baseline in low-rank matrix recovery and a model example of dual (Uzawa-type) methods for nuclear-norm problems. This mission formalizes its convergence theorem.

Setting

All matrices are real. For X,Y∈Rn1×n2X,Y\in\mathbb R^{n_1\times n_2}X,Y∈Rn1​×n2​ write ⟨X,Y⟩=trace⁡(X∗Y)=∑i,jXijYij\langle X,Y\rangle=\operatorname{trace}(X^*Y)=\sum_{i,j}X_{ij}Y_{ij}⟨X,Y⟩=trace(X∗Y)=∑i,j​Xij​Yij​ and ∥X∥F2=⟨X,X⟩\|X\|_F^2=\langle X,X\rangle∥X∥F2​=⟨X,X⟩. The nuclear norm ∥X∥∗\|X\|_*∥X∥∗​ is the sum of the singular values of XXX.

For an index set Ω\OmegaΩ, the sampling projector PΩP_\OmegaPΩ​ keeps the entries with indices in Ω\OmegaΩ and sets the others to zero.

A reduced singular value decomposition of a matrix YYY of rank rrr is Y=UΣV∗Y=U\Sigma V^*Y=UΣV∗ with UUU (n1×rn_1\times rn1​×r) and VVV (n2×rn_2\times rn2​×r) having orthonormal columns and Σ=diag⁡(σ1,…,σr)\Sigma=\operatorname{diag}(\sigma_1,\dots,\sigma_r)Σ=diag(σ1​,…,σr​) with σi>0\sigma_i>0σi​>0. For τ≥0\tau\ge0τ≥0 the singular value shrinkage operator is

Dτ(Y)=Udiag⁡((σi−τ)+)V∗,t+=max⁡(0,t).\mathcal D_\tau(Y)=U\operatorname{diag}\big((\sigma_i-\tau)_+\big)V^*,\qquad t_+=\max(0,t).Dτ​(Y)=Udiag((σi​−τ)+​)V∗,t+​=max(0,t).

Fix τ>0\tau>0τ>0, a sequence of step sizes {δk}k≥1\{\delta_k\}_{k\ge1}{δk​}k≥1​ and data MMM. The SVT iteration (2.7) starts from Y0=0Y^0=0Y0=0 and sets, for k=1,2,…k=1,2,\dotsk=1,2,…,

Xk=Dτ(Yk−1),Yk=Yk−1+δkPΩ(M−Xk).X^k=\mathcal D_\tau(Y^{k-1}),\qquad Y^k=Y^{k-1}+\delta_k P_\Omega(M-X^k).Xk=Dτ​(Yk−1),Yk=Yk−1+δk​PΩ​(M−Xk).

The proximal problem (2.8) is

minimize  fτ(X)=τ∥X∥∗+12∥X∥F2subject to  PΩ(X)=PΩ(M).\text{minimize}\ \ f_\tau(X)=\tau\|X\|_*+\tfrac12\|X\|_F^2\quad\text{subject to}\ \ P_\Omega(X)=P_\Omega(M).minimize  fτ​(X)=τ∥X∥∗​+21​∥X∥F2​subject to  PΩ​(X)=PΩ​(M).

More generally, for a linear map A:Rn1×n2→Rm\mathcal A:\mathbb R^{n_1\times n_2}\to\mathbb R^mA:Rn1​×n2​→Rm with adjoint A∗\mathcal A^*A∗ and spectral norm ∥A∥=sup⁡{∥A(X)∥ℓ2:∥X∥F=1}\|\mathcal A\|=\sup\{\|\mathcal A(X)\|_{\ell_2}:\|X\|_F=1\}∥A∥=sup{∥A(X)∥ℓ2​​:∥X∥F​=1}, and b∈Rmb\in\mathbb R^mb∈Rm, problem (3.1) is to minimize fτ(X)f_\tau(X)fτ​(X) subject to A(X)=b\mathcal A(X)=bA(X)=b, and Uzawa's iteration (3.3) starts from y0=0y^0=0y0=0 and sets Xk=Dτ(A∗(yk−1))X^k=\mathcal D_\tau(\mathcal A^*(y^{k-1}))Xk=Dτ​(A∗(yk−1)), yk=yk−1+δk(b−A(Xk))y^k=y^{k-1}+\delta_k(b-\mathcal A(X^k))yk=yk−1+δk​(b−A(Xk)).

Formalization targets

Goal: Theorem 4.2, second sentence (p. 1968)

If 0<inf⁡kδk≤sup⁡kδk<20<\inf_k\delta_k\le\sup_k\delta_k<20<infk​δk​≤supk​δk​<2, then (2.8) has a unique solution X⋆X^\starX⋆ and the SVT iterates satisfy

lim⁡k→∞Xk=X⋆.\lim_{k\to\infty}X^k=X^\star .k→∞lim​Xk=X⋆.

Theorem 4.2, first sentence (p. 1968)

If (3.1) is feasible and 0<inf⁡kδk≤sup⁡kδk<2/∥A∥20<\inf_k\delta_k\le\sup_k\delta_k<2/\|\mathcal A\|^20<infk​δk​≤supk​δk​<2/∥A∥2, then (3.1) has a unique solution and the iterates XkX^kXk of (3.3) converge to it.

Supporting results (milestones, in attack order)

  1. Well-definedness of Dτ\mathcal D_\tauDτ​ (§2.1, p. 1960): the output does not depend on the chosen SVD.
  2. Theorem 2.1 (p. 1960): Dτ(Y)=arg⁡min⁡X12∥X−Y∥F2+τ∥X∥∗\mathcal D_\tau(Y)=\arg\min_X \tfrac12\|X-Y\|_F^2+\tau\|X\|_*Dτ​(Y)=argminX​21​∥X−Y∥F2​+τ∥X∥∗​.
  3. Sparsity of the iterates (§2.2, p. 1961): since Y0=0Y^0=0Y0=0, every YkY^kYk vanishes outside Ω\OmegaΩ.
  4. Eq. (2.14) (p. 1964): the minimizers of the Lagrangian fτ(X)+⟨Y,PΩ(M−X)⟩f_\tau(X)+\langle Y,P_\Omega(M-X)\ranglefτ​(X)+⟨Y,PΩ​(M−X)⟩ are those of τ∥X∥∗+12∥X−PΩY∥F2\tau\|X\|_*+\tfrac12\|X-P_\Omega Y\|_F^2τ∥X∥∗​+21​∥X−PΩ​Y∥F2​.
  5. Lemma 4.1 (p. 1968): for Z∈∂fτ(X)Z\in\partial f_\tau(X)Z∈∂fτ​(X), Z′∈∂fτ(X′)Z'\in\partial f_\tau(X')Z′∈∂fτ​(X′), ⟨Z−Z′,X−X′⟩≥∥X−X′∥F2\langle Z-Z',X-X'\rangle\ge\|X-X'\|_F^2⟨Z−Z′,X−X′⟩≥∥X−X′∥F2​.
  6. The §3.1 reduction (p. 1964): for a sampling operator, A∗A=PΩ\mathcal A^*\mathcal A=P_\OmegaA∗A=PΩ​ and (3.3) becomes (2.7) under Yk=A∗(yk)Y^k=\mathcal A^*(y^k)Yk=A∗(yk).
  7. Theorem 4.2, first sentence, as above.

Significance

The theorem certifies that SVT, run with any step sizes in a fixed interval (0,2)(0,2)(0,2), computes the unique minimizer of the strongly convex surrogate (2.8). Together with the separate fact that the solution of (2.8) tends to the minimum nuclear norm completion as τ→∞\tau\to\inftyτ→∞ (the paper's Theorem 3.1, a companion mission), this is what justifies using SVT as a solver for nuclear-norm matrix completion. Theorem 2.1, the proximal characterization of singular value soft-thresholding, is used throughout the literature on proximal methods for low-rank problems.

The paper's proof of Theorem 4.2 consists of the reduction to Uzawa's method and a citation of a general convergence theorem for projected gradient methods on the dual. The formalization produces a self-contained, machine-checked chain: the proximal characterization of Dτ\mathcal D_\tauDτ​, the Lagrangian identity, strong monotonicity of ∂fτ\partial f_\tau∂fτ​, and the convergence argument itself. To our knowledge none of these results has a machine-checked proof; Mathlib at the pinned revision has singular values of linear maps but no SVD structure, no nuclear norm and no subgradient calculus.

Difficulty

Nothing in the iteration is a gradient step of a smooth function in XXX: the XXX-update is a nonsmooth proximal map, and the convergence of XkX^kXk is not visible from the recursion itself. The paper's argument cites a general theorem on projected gradient methods ([25, Theorem 2.1]) and takes for granted that "strong duality holds" for (2.8) (p. 1963), so the existence of a Lagrange multiplier is part of what must be formalized. Theorem 2.1 depends on the subdifferential of the nuclear norm, which Mathlib does not provide, and therefore on the singular value decomposition and the duality between the nuclear and spectral norms. Convergence of objective values or of a subsequence would not suffice: the target is convergence of the whole sequence XkX^kXk to the unique solution.

Formalization scope

Matrices are Matrix (Fin n₁) (Fin n₂) ℝ; convergence is Mathlib's topology on matrices, which coincides with the Frobenius-norm topology. ⟨⋅,⋅⟩\langle\cdot,\cdot\rangle⟨⋅,⋅⟩ and ∥⋅∥F\|\cdot\|_F∥⋅∥F​ are defined entrywise; ∥X∥∗\|X\|_*∥X∥∗​ is the sum of Mathlib's LinearMap.singularValues of XXX viewed as a map Rn2→Rn1\mathbb R^{n_2}\to\mathbb R^{n_1}Rn2​→Rn1​. The shrinkage operator is a relation IsShrink τ Y X defined, as in (2.1)–(2.2), through some reduced SVD of YYY; well-definedness is a milestone. It is not defined as the minimizer of (2.3), which would make Theorem 2.1 definitional. Linear maps A\mathcal AA are given by matrices A1,…,AmA_1,\dots,A_mA1​,…,Am​ with A(X)i=⟨Ai,X⟩\mathcal A(X)_i=\langle A_i,X\rangleA(X)i​=⟨Ai​,X⟩, and a sampling operator by an injective enumeration of Ω\OmegaΩ. Subgradients are those of (2.4).

Sequences are indexed by N\mathbb NN: Lean's step k+1k+1k+1 is the paper's step kkk, so X0X^0X0 and δ0\delta_0δ0​ are unused. Committed conventions:

  • Y0=0Y^0=0Y0=0 and y0=0y^0=0y0=0 are hypotheses; with a start that is nonzero outside Ω\OmegaΩ the iterates converge to a different matrix.
  • The standing τ>0\tau>0τ>0 is kept, except in Theorem 2.1 and the well-definedness statement, which are printed for τ≥0\tau\ge0τ≥0.
  • The step-size conditions are explicit bounds a>0a>0a>0, CCC with a≤δk≤Ca\le\delta_k\le Ca≤δk​≤C for k≥1k\ge1k≥1, together with C<2C<2C<2, respectively C∥A∥2<2C\|\mathcal A\|^2<2C∥A∥2<2. The multiplicative form avoids Lean's x/0=0x/0=0x/0=0: for A=0\mathcal A=0A=0 the condition does not become unsatisfiable.
  • Feasibility of (3.1) is an added hypothesis of Theorem 4.2's first sentence, since "the unique solution" presupposes it.
  • "Converges to the unique solution" is stated as existence and uniqueness of the solution together with convergence of the whole sequence to it.

A small unprinted helper, ∥A∥≤1\|\mathcal A\|\le1∥A∥≤1 for sampling operators, is included to pass from the first sentence of Theorem 4.2 to the second; it is not a milestone. The subdifferential formula (2.6) of the nuclear norm and the Fejér-type condition of §5.1.2 are not stated. Reusable infrastructure welcome from solvers: existence and uniqueness properties of the reduced SVD, the nuclear/spectral norm duality, the subdifferential of the nuclear norm, and a general convergence theorem for Uzawa's method with a strongly convex objective.

Selected references

  • J.-F. Cai, E. J. Candès, Z. Shen, A Singular Value Thresholding Algorithm for Matrix Completion, SIAM J. Optim. 20(4):1956–1982, 2010. https://doi.org/10.1137/080738970
  • E. J. Candès, B. Recht, Exact Matrix Completion via Convex Optimization, Found. Comput. Math. 9:717–772, 2009. https://doi.org/10.1007/s10208-009-9045-5
  • K. J. Arrow, L. Hurwicz, H. Uzawa, Studies in Linear and Non-Linear Programming, Stanford University Press, 1958.
12 thms2 active usersReviewed
Convex OptimizationLinear OptimizationOperations Research+1·Captain: mikedeng1

Path-Finding Methods for Linear Programming II: Properties of the Regularized D-Optimal-Design Weight FunctionResearch Paper

Motivation

Interior point methods for a linear program min⁡{c⊤x:Ax≥b}\min\{c^\top x : Ax\ge b\}min{c⊤x:Ax≥b} with A∈Rm×nA\in\mathbb R^{m\times n}A∈Rm×n follow the central path of the logarithmic barrier −∑ilog⁡si-\sum_i\log s_i−∑i​logsi​, where s=Ax−bs=Ax-bs=Ax−b is the slack vector. Renegar's path-following analysis (1988) gives O(m L)O(\sqrt m\,L)O(m​L) iterations, and for decades this was the best bound for methods whose iterations cost a linear system solve. Vaidya's volumetric barrier −log⁡det⁡(A⊤S−2A)-\log\det(A^\top S^{-2}A)−logdet(A⊤S−2A) and the hybrid volumetric barriers of Vaidya and of Anstreicher (references [45] and [2] of the paper) reached O((m rank(A))1/4L)O((m\,\mathrm{rank}(A))^{1/4}L)O((mrank(A))1/4L) iterations at the price of more expensive linear algebra. Nesterov and Nemirovski showed that a universal barrier gives O(n L)O(\sqrt n\,L)O(n​L) iterations, but that barrier cannot be evaluated efficiently.

Lee and Sidford (FOCS 2014; full version arXiv:1312.6677) obtained O~(rank(A) L)\tilde O(\sqrt{\mathrm{rank}(A)}\,L)O~(rank(A)​L) iterations, each costing O~(1)\tilde O(1)O~(1) linear system solves, by following a weighted central path whose weights are recomputed from the slacks. The weights come from a weight function ggg, defined as the minimizer of a regularized D-optimal-design problem. This mission is about that weight function and the theorem (Theorem 1 of the paper) certifying its properties. The companion mission, Path-Finding Methods for Linear Programming I, formalizes the path-following framework (Theorem 5 of §IV.C) that consumes these properties.

Setting

Fix A∈Rm×nA\in\mathbb R^{m\times n}A∈Rm×n with full column rank, rank(A)=n\mathrm{rank}(A)=nrank(A)=n, and 1≤n<m1\le n<m1≤n<m. For vectors s,w∈R>0ms,w\in\mathbb R^m_{>0}s,w∈R>0m​ write S=diag(s)S=\mathrm{diag}(s)S=diag(s), W=diag(w)W=\mathrm{diag}(w)W=diag(w), Wα=diag(wiα)W^\alpha=\mathrm{diag}(w_i^\alpha)Wα=diag(wiα​), and As=S−1AA_s=S^{-1}AAs​=S−1A. For a matrix MMM let ∥v∥M=v⊤Mv\|v\|_M=\sqrt{v^\top Mv}∥v∥M​=v⊤Mv​.

Projection matrix and slack sensitivity (Definition 2, p. 428). The projection matrix is PS−1A(w)=W1/2S−1A (A⊤S−1WS−1A)−1A⊤S−1W1/2P_{S^{-1}A}(w)=W^{1/2}S^{-1}A\,(A^\top S^{-1}WS^{-1}A)^{-1}A^\top S^{-1}W^{1/2}PS−1A​(w)=W1/2S−1A(A⊤S−1WS−1A)−1A⊤S−1W1/2, and the slack sensitivity is

γ(s,w)=max⁡i∈[m]∥W−1/21i∥PS−1A(w).\gamma(s,w)=\max_{i\in[m]}\big\|W^{-1/2}\mathbb 1_i\big\|_{P_{S^{-1}A}(w)} .γ(s,w)=i∈[m]max​​W−1/21i​​PS−1A​(w)​.

Weight function (Definition 4, p. 428). A map g:R>0m→R>0mg:\mathbb R^m_{>0}\to\mathbb R^m_{>0}g:R>0m​→R>0m​ is a weight function with constants c1,cγ,crc_1,c_\gamma,c_rc1​,cγ​,cr​ if it is differentiable and, for every s>0s>0s>0, with G(s)=diag(g(s))G(s)=\mathrm{diag}(g(s))G(s)=diag(g(s)), G′(s)G'(s)G′(s) the Jacobian of ggg at sss, and ∥y∥G(s)=∑igi(s)yi2\|y\|_{G(s)}=\sqrt{\sum_ig_i(s)y_i^2}∥y∥G(s)​=∑i​gi​(s)yi2​​:

  1. Size: ∥g(s)∥1≤c1\|g(s)\|_1\le c_1∥g(s)∥1​≤c1​;
  2. Slack sensitivity: cγ≥1c_\gamma\ge1cγ​≥1 and γ(s,g(s))≤cγ\gamma(s,g(s))\le c_\gammaγ(s,g(s))≤cγ​;
  3. Step consistency: cr≥1c_r\ge1cr​≥1 and for all r≥crr\ge c_rr≥cr​, y∈Rmy\in\mathbb R^my∈Rm: ∥(I+r−1G−1G′S)y∥G(s)≤∥y∥G(s)\|(I+r^{-1}G^{-1}G'S)y\|_{G(s)}\le\|y\|_{G(s)}∥(I+r−1G−1G′S)y∥G(s)​≤∥y∥G(s)​ and ∥y+r−1G−1G′Sy∥∞≤∥y∥∞+cr∥y∥G(s)\|y+r^{-1}G^{-1}G'Sy\|_\infty\le\|y\|_\infty+c_r\|y\|_{G(s)}∥y+r−1G−1G′Sy∥∞​≤∥y∥∞​+cr​∥y∥G(s)​;
  4. Uniformity: ∥g(s)∥∞≤2\|g(s)\|_\infty\le2∥g(s)∥∞​≤2.

The regularized objective (6), p. 429. For α,β∈R\alpha,\beta\in\mathbb Rα,β∈R,

f^(s,w)=1⊤w−1αlog⁡det⁡(As⊤WαAs)−β∑i∈[m]log⁡wi,g(s)=arg⁡min⁡w∈R>0mf^(s,w).\hat f(s,w)=\mathbb 1^\top w-\frac1\alpha\log\det\big(A_s^\top W^\alpha A_s\big)-\beta\sum_{i\in[m]}\log w_i ,\qquad g(s)=\arg\min_{w\in\mathbb R^m_{>0}}\hat f(s,w).f^​(s,w)=1⊤w−α1​logdet(As⊤​WαAs​)−βi∈[m]∑​logwi​,g(s)=argw∈R>0m​min​f^​(s,w).

At α=1,β=0\alpha=1,\beta=0α=1,β=0 this is the D-optimal design problem, dual to computing the John ellipsoid of the polytope {y:∣[A(y−x)]i∣≤si}\{y:|[A(y-x)]_i|\le s_i\}{y:∣[A(y−x)]i​∣≤si​} (§V.B).

Formalization targets

Goal: Theorem 1 (Properties of Weight Function), §V.A, p. 429

With

α=1−(log⁡22mrank(A))−1,β=rank(A)2m,\alpha=1-\Big(\log_2\frac{2m}{\mathrm{rank}(A)}\Big)^{-1},\qquad \beta=\frac{\mathrm{rank}(A)}{2m},α=1−(log2​rank(A)2m​)−1,β=2mrank(A)​,

the objective f^(s,⋅)\hat f(s,\cdot)f^​(s,⋅) has a unique minimizer over R>0m\mathbb R^m_{>0}R>0m​ for every s>0s>0s>0, and the resulting ggg is a weight function with

c1(g)=2 rank(A),cγ(g)=2,cr(g)=2log⁡22mrank(A).c_1(g)=2\,\mathrm{rank}(A),\qquad c_\gamma(g)=2,\qquad c_r(g)=2\log_2\frac{2m}{\mathrm{rank}(A)} .c1​(g)=2rank(A),cγ​(g)=2,cr​(g)=2log2​rank(A)2m​.

Milestones: the three bullets of Theorem 1

  • Size: every minimizer www of f^(s,⋅)\hat f(s,\cdot)f^​(s,⋅) satisfies ∥w∥1≤2 rank(A)\|w\|_1\le2\,\mathrm{rank}(A)∥w∥1​≤2rank(A).
  • Slack sensitivity: every minimizer www satisfies γ(s,w)≤2\gamma(s,w)\le2γ(s,w)≤2.
  • Step consistency: any map ggg selecting a minimizer at every s>0s>0s>0 is differentiable on R>0m\mathbb R^m_{>0}R>0m​ and satisfies the two step-consistency inequalities for every r≥2log⁡22mrank(A)r\ge2\log_2\frac{2m}{\mathrm{rank}(A)}r≥2log2​rank(A)2m​.

A supporting (non-milestone) item states the existence and uniqueness of the minimizer on its own.

Significance

The result. Theorem 1 is the input that turns the weighted path-following framework into an O~(rank(A) L)\tilde O(\sqrt{\mathrm{rank}(A)}\,L)O~(rank(A)​L)-iteration method: the framework needs O(cγ−1cr−3c1−1/2)O(c_\gamma^{-1}c_r^{-3}c_1^{-1/2})O(cγ−1​cr−3​c1−1/2​)-sized steps in ttt (p. 428), and Theorem 1 makes that Ω~(1/rank(A))\tilde\Omega(1/\sqrt{\mathrm{rank}(A)})Ω~(1/rank(A)​). The step consistency bound is what allows the weights to be recomputed after each Newton step without losing centrality. The same construction underlies later work on Lewis-weight barriers and on fast approximate John ellipsoids and maximum flow (§VIII of the paper).

Formalizing it. The theorem is proved in the full version of the paper (arXiv:1312.6677); the FOCS extended abstract contains no proofs. No part of it has a machine-checked proof. A complete formalization would give a verified account of leverage-score calculus (sums of leverage scores equal the rank; derivatives of projection matrices), of the convexity of w↦−log⁡det⁡(A⊤WαA)w\mapsto-\log\det(A^\top W^\alpha A)w↦−logdet(A⊤WαA) for α∈(0,1)\alpha\in(0,1)α∈(0,1), and of differentiability of an argmin via the implicit function theorem, none of which is currently packaged in Mathlib in this form.

Difficulty

Size and slack sensitivity are statements about the minimizer, which is only characterized implicitly; they require precise matrix calculus for log⁡det⁡(As⊤WαAs)\log\det(A_s^\top W^\alpha A_s)logdet(As⊤​WαAs​) and a comparison between the matrices A⊤WAA^\top WAA⊤WA (which defines γ\gammaγ) and A⊤WαAA^\top W^\alpha AA⊤WαA (which defines ggg). The specific values of α\alphaα and β\betaβ matter here: the unregularized choice α=1\alpha=1α=1, β=0\beta=0β=0 makes the problem degenerate (p. 429).

The hard part is step consistency. The Jacobian G′G'G′ of an argmin is available only implicitly, as the solution of a linear system obtained by differentiating the optimality condition. A bound on ∥G′∥\|G'\|∥G′∥ that depends on mmm is easy to get and useless: the theorem needs the operator norm of I+r−1G−1G′SI+r^{-1}G^{-1}G'SI+r−1G−1G′S in the G(s)G(s)G(s)-norm to be at most 111 as soon as rrr exceeds 2log⁡2(2m/rank(A))2\log_2(2m/\mathrm{rank}(A))2log2​(2m/rank(A)), and an ℓ∞\ell_\inftyℓ∞​ bound with only an additive cr∥y∥G(s)c_r\|y\|_{G(s)}cr​∥y∥G(s)​ loss.

Existence and differentiability of the minimizer are conclusions, not hypotheses. The minimization is over an open orthant on which the objective is not obviously coercive or strictly convex for α<1\alpha<1α<1, and differentiability of ggg requires the Hessian of f^\hat ff^​ at the minimizer to be invertible.

Formalization scope

Vectors are Fin m → ℝ, matrices Matrix (Fin m) (Fin n) ℝ; inverses are Matrix.inv, log⁡det⁡\log\detlogdet is Real.log (Matrix.det …), wiαw_i^\alphawiα​ is Real.rpow, log⁡2\log_2log2​ is Real.logb 2, the Jacobian is fderiv ℝ g s, and ∥⋅∥∞\|\cdot\|_\infty∥⋅∥∞​ is Mathlib's sup norm on Fin m → ℝ.

Conventions and pinned hypotheses:

  • Full column rank A.rank = n is assumed in every theorem. The paper never states it, but without it As⊤WαAsA_s^\top W^\alpha A_sAs⊤​WαAs​ is singular and every formula is undefined (in Lean, Matrix.inv and Real.log would return junk 000).
  • 1≤n<m1\le n<m1≤n<m. β=rank(A)/(2m)\beta=\mathrm{rank}(A)/(2m)β=rank(A)/(2m) and log⁡2(2m/rank(A))\log_2(2m/\mathrm{rank}(A))log2​(2m/rank(A)) need rank(A)≥1\mathrm{rank}(A)\ge1rank(A)≥1; at m=rank(A)m=\mathrm{rank}(A)m=rank(A) the page's α\alphaα is 000 and 1/α1/\alpha1/α in (6) is undefined.
  • Reading of α\alphaα: the exponent −1-1−1 is the reciprocal of log⁡22mrank(A)\log_2\frac{2m}{\mathrm{rank}(A)}log2​rank(A)2m​, giving α∈(0,1)\alpha\in(0,1)α∈(0,1).
  • Size is an upper bound ∥g(s)∥1≤c1\|g(s)\|_1\le c_1∥g(s)∥1​≤c1​ (the paper's weight function has ∥g(s)∥1=32rank(A)\|g(s)\|_1=\tfrac32\mathrm{rank}(A)∥g(s)∥1​=23​rank(A), while Theorem 1 reports c1=2 rank(A)c_1=2\,\mathrm{rank}(A)c1​=2rank(A)).
  • The first step-consistency bullet (an operator-norm bound) is stated for every vector yyy.
  • ggg is any map Rm→Rm\mathbb R^m\to\mathbb R^mRm→Rm whose value at each positive sss minimizes f^(s,⋅)\hat f(s,\cdot)f^​(s,⋅) over R>0m\mathbb R^m_{>0}R>0m​. Only its values on the open orthant matter. The goal also asserts that such minimizers exist and are unique, so it is not vacuous.

Ruling out trivializations: the goal does not assume ggg to be a weight function or to be differentiable, and it does not replace ggg by an arbitrary weight function; differentiability is a conclusion (a predicate using fderiv without it would make step consistency hold vacuously wherever ggg fails to be differentiable).

Useful infrastructure, reusable beyond this mission: leverage scores and their sum; derivatives of w↦log⁡det⁡(A⊤WA)w\mapsto\log\det(A^\top WA)w↦logdet(A⊤WA) and of projection matrices; convexity of −log⁡det⁡(A⊤WαA)-\log\det(A^\top W^\alpha A)−logdet(A⊤WαA) in www (related to the published ConvexOptimization.log_det_concaveOn); differentiability of the argmin of a strictly convex smooth function. Contributions of these as separate theorems are welcome, as is a proof of any single bullet of Theorem 1.

Selected references

  • Y. T. Lee, A. Sidford, Path Finding Methods for Linear Programming: Solving Linear Programs in Õ(√rank) Iterations and Faster Algorithms for Maximum Flow, FOCS 2014, pp. 424–433. https://doi.org/10.1109/FOCS.2014.52
  • Y. T. Lee, A. Sidford, Path Finding I: Solving Linear Programs with Õ(√rank) Linear System Solves, arXiv, 2013. https://arxiv.org/abs/1312.6677
  • J. Renegar, A polynomial-time algorithm, based on Newton's method, for linear programming, Mathematical Programming 40 (1988). https://doi.org/10.1007/BF01580724
7 thms2 active usersReviewed
🏆Completed
Mathematical Physics·Captain: ShapeZero

The passivity-admissible couplings have dimension n² = dim u(n)Textbook

Motivation

The Shape Zero model (Shape Zero LLC, unpublished) aims to obtain all three factors of the Standard Model gauge structure — U(1)U(1)U(1), SU(2)SU(2)SU(2) and SU(3)SU(3)SU(3) — from a single linear-algebra statement applied at three node sizes. A node carrying nnn oscillator pairs has a 2n2n2n-dimensional real state space with a complex structure JJJ. Within the model, the couplings that do no net work (passive couplings) are the symmetric matrices, and those that respect JJJ are the ones commuting with it. The claim is that the couplings satisfying both conditions form a real vector space of dimension exactly n2n^2n2, which is the dimension of the unitary Lie algebra u(n)\mathfrak{u}(n)u(n) (Unitary group):

node size nnnadmissible dimensionLie algebra
11u(1)\mathfrak{u}(1)u(1)
24u(2)=u(1)⊕su(2)\mathfrak{u}(2) = \mathfrak{u}(1) \oplus \mathfrak{su}(2)u(2)=u(1)⊕su(2)
39u(3)=u(1)⊕su(3)\mathfrak{u}(3) = \mathfrak{u}(1) \oplus \mathfrak{su}(3)u(3)=u(1)⊕su(3)

The dimension count has so far been checked numerically for n=1,…,5n = 1, \dots, 5n=1,…,5 only, giving 1,4,9,16,251, 4, 9, 16, 251,4,9,16,25. A machine-checked proof covers every nnn, and it is the claim a physicist examining the model would check first.

What this mission does NOT prove. This mission proves the linear algebra: symmetric matrices commuting with JJJ form a space of dimension n2n^2n2. It does not prove the physics step that passivity forces a coupling to be symmetric. That step is a separate premise of the model, and a completed mission must not be read as establishing it.

Setting

Fix a natural number nnn. Consider real 2n×2n2n \times 2n2n×2n matrices, with rows and columns indexed by two copies of {0,…,n−1}\{0, \dots, n-1\}{0,…,n−1}, so that each matrix is written in 2×22 \times 22×2 block form with n×nn \times nn×n blocks:

W=(ABCD).W = \begin{pmatrix} A & B \\ C & D \end{pmatrix}.W=(AC​BD​).

Let III be the n×nn \times nn×n identity matrix and define the standard complex structure

J=(0−II0),J = \begin{pmatrix} 0 & -I \\ I & 0 \end{pmatrix},J=(0I​−I0​),

which satisfies J2=−1J^2 = -1J2=−1. In the Lean development this matrix is PassivityUn.stdJ n, and the index type is PassivityUn.Blk n, the disjoint union of two copies of Fin n.

Two real linear subspaces of the 4n24n^24n2-dimensional space of such matrices are defined:

  1. symm n: the symmetric matrices, WT=WW^{\mathsf T} = WWT=W;
  2. commJ n: the matrices commuting with JJJ, WJ=JWWJ = JWWJ=JW.

The admissible coupling class admissible n is their intersection:

An={ W∈R2n×2n:WT=W and WJ=JW }.\mathcal{A}_n = \{\, W \in \mathbb{R}^{2n \times 2n} : W^{\mathsf T} = W \ \text{and}\ WJ = JW \,\}.An​={W∈R2n×2n:WT=W and WJ=JW}.

Formalization targets

Goal: the admissible class has dimension n2n^2n2

dim⁡RAn=n2for every n∈N.\dim_{\mathbb{R}} \mathcal{A}_n = n^2 \qquad \text{for every } n \in \mathbb{N}.dimR​An​=n2for every n∈N.

This is PassivityUn.admissible_finrank. It asserts the exact dimension for all nnn at once, not for a particular node size.

Milestones

  1. M1. J⋅J=−1J \cdot J = -1J⋅J=−1, so JJJ is a complex structure.
  2. M2. A block matrix (ABCD)\begin{pmatrix} A & B \\ C & D \end{pmatrix}(AC​BD​) commutes with JJJ exactly when D=AD = AD=A and B=−CB = -CB=−C.
  3. M3. A block matrix (A−BBA)\begin{pmatrix} A & -B \\ B & A \end{pmatrix}(AB​−BA​) is symmetric exactly when AT=AA^{\mathsf T} = AAT=A and BT=−BB^{\mathsf T} = -BBT=−B.
  4. M4. The dimension counts of symmetric and antisymmetric n×nn \times nn×n matrices add to n2n^2n2: n(n+1)2+n(n−1)2=n2\tfrac{n(n+1)}{2} + \tfrac{n(n-1)}{2} = n^22n(n+1)​+2n(n−1)​=n2.

Significance

The result itself. The goal identifies the admissible coupling class with the real form of the n×nn \times nn×n Hermitian matrices, the space whose dimension is that of u(n)\mathfrak{u}(n)u(n) (Hermitian matrix). Applied at n=1,2,3n = 1, 2, 3n=1,2,3 it gives the dimensions 111, 444 and 999, which the Shape Zero model matches with u(1)\mathfrak{u}(1)u(1), u(2)=u(1)⊕su(2)\mathfrak{u}(2) = \mathfrak{u}(1) \oplus \mathfrak{su}(2)u(2)=u(1)⊕su(2) and u(3)=u(1)⊕su(3)\mathfrak{u}(3) = \mathfrak{u}(1) \oplus \mathfrak{su}(3)u(3)=u(1)⊕su(3). Without a proof for general nnn, the model rests on a finite numerical check.

Formalizing it. The underlying fact is standard linear algebra, a special case of the correspondence between real matrices commuting with a complex structure and complex-linear maps (Linear complex structure). No machine-checked statement of this exact dimension count was found in the platform library. The mission produces a verified, general-nnn statement whose hypotheses are fully explicit, and it separates the verified linear algebra from the unverified physical premise.

Difficulty

Every step is standard linear algebra, so the difficulty is in the formal bookkeeping rather than the mathematics. The dimension of a subspace defined by equations is not computed by any Mathlib tactic. It has to be obtained by exhibiting an explicit linear equivalence with spaces of known dimension, and that requires moving between the 2n×2n2n \times 2n2n×2n matrix indexed by a disjoint union and its four n×nn \times nn×n blocks. Checking small cases numerically, as has already been done, does not extend to a statement about every nnn.

Formalization scope

  • Matrices are real (ℝ), not complex. The complex structure enters only through the fixed real matrix JJJ.
  • Matrices are indexed by Fin n ⊕ Fin n (PassivityUn.Blk n) rather than Fin (2n), so that block decomposition via Mathlib's Matrix.fromBlocks is direct. The first copy of Fin n indexes the upper/left blocks.
  • Each condition is defined as the kernel of a linear map: symmetry as the kernel of W↦WT−WW \mapsto W^{\mathsf T} - WW↦WT−W, and commuting with JJJ as the kernel of W↦WJ−JWW \mapsto WJ - JWW↦WJ−JW. Both are therefore subspaces by construction, with no hand-written closure proofs.
  • Dimension is Module.finrank ℝ. The ambient space is finite-dimensional, so the convention that finrank of an infinite-dimensional space is 000 never applies.
  • The case n=0n = 0n=0 is included, and there the statement reads 0=00 = 00=0. This is not a trivialization: the claim is quantified over all nnn, and every n≥1n \ge 1n≥1 is a nontrivial instance.
  • M4 is stated with natural-number division and truncated subtraction. Both are exact here, because n(n±1)n(n \pm 1)n(n±1) is always even and n⋅(n−1)=0n \cdot (n - 1) = 0n⋅(n−1)=0 when n=0n = 0n=0.

The development needs only Mathlib's block matrices, transpose, linear maps and finrank. The block-characterization lemmas (M2, M3) are reusable for any statement relating real matrices commuting with JJJ to complex matrices. Contributions are welcome on the milestones and on the assembling isomorphism.

Selected references

  • B. C. Hall, Lie Groups, Lie Algebras, and Representations: An Elementary Introduction, 2nd ed., Graduate Texts in Mathematics 222, Springer, 2015. https://doi.org/10.1007/978-3-319-13467-3
  • Wikipedia, Unitary group. https://en.wikipedia.org/wiki/Unitary_group
  • Wikipedia, Linear complex structure. https://en.wikipedia.org/wiki/Linear_complex_structure
  • Wikipedia, Hermitian matrix. https://en.wikipedia.org/wiki/Hermitian_matrix
  • Mathlib, Mathlib.Data.Matrix.Block (block matrices, Matrix.fromBlocks). https://leanprover-community.github.io/mathlib4_docs/Mathlib/Data/Matrix/Block.html
7 thms2 active usersReviewed
CombinatoricsGraph TheoryTheoretical Computer Science·Captain: mikedeng1

Explicit Expanders of Every Degree and Size 3: Deleting Far-Apart Tree-Like Vertices of a Near-Ramanujan Graph and Matching Their Neighbours Keeps λ ≤ 2√(d−1) + εResearch Paper

Motivation

Sparse graphs with small nontrivial eigenvalues, expanders, are basic objects in combinatorics and theoretical computer science. They are used in error-correcting codes, derandomization, sorting networks, and the analysis of random walks. The Alon–Boppana bound says that a ddd-regular graph on nnn vertices has a nontrivial eigenvalue of absolute value at least 2d−1−o(1)2\sqrt{d-1}-o(1)2d−1​−o(1) (Alon 1986; Nilli 1991). Graphs that reach 2d−12\sqrt{d-1}2d−1​ are Ramanujan graphs.

The classical explicit Ramanujan graphs of Lubotzky, Phillips and Sarnak (1988) and Margulis exist only for degrees d=p+1d = p+1d=p+1 with ppp prime, and only for very sparse sequences of vertex counts. Constructions with λ≤2d−1+ε\lambda\le 2\sqrt{d-1}+\varepsilonλ≤2d−1​+ε for every degree came from Mohanty, O'Donnell and Paredes (STOC 2020, arXiv:1909.06988), but their graphs also do not have every number of vertices. Alon (arXiv:2003.11673, Combinatorica 41, 2021) asked for near-Ramanujan graphs of every degree and every large size. Theorem 1.3 of that paper answers this up to ε\varepsilonε: for every ddd, every ε>0\varepsilon>0ε>0 and every large nnn with ndndnd even there is an explicit (n,d,λ)(n,d,\lambda)(n,d,λ)-graph with λ≤2d−1+ε\lambda\le 2\sqrt{d-1}+\varepsilonλ≤2d−1​+ε.

Setting

A (n,d,λ)(n,d,\lambda)(n,d,λ)-graph is a ddd-regular simple graph on nnn vertices in which every nontrivial eigenvalue of the adjacency matrix AAA has absolute value at most λ\lambdaλ. The trivial eigenvalue is ddd, with the constant eigenvector 1\mathbf 11. Equivalently, every eigenvalue μ\muμ of AAA with an eigenvector f≠0f\ne0f=0, ∑vf(v)=0\sum_v f(v)=0∑v​f(v)=0, satisfies ∣μ∣≤λ|\mu|\le\lambda∣μ∣≤λ.

Distances dist⁡(v,w)\operatorname{dist}(v,w)dist(v,w) are graph distances, and they are ∞\infty∞ between components. The kkk-neighbourhood of a vertex vvv is B(v,k)={w:dist⁡(v,w)≤k}B(v,k)=\{w:\operatorname{dist}(v,w)\le k\}B(v,k)={w:dist(v,w)≤k}. The kkk-neighbourhood of an edge uvuvuv is B(u,k)∪B(v,k)B(u,k)\cup B(v,k)B(u,k)∪B(v,k), and NiN_iNi​ is the set of vertices at distance exactly iii from {u,v}\{u,v\}{u,v}. A set contains no cycle if the subgraph induced on it is a forest. A ball contains at most one cycle if its induced subgraph has at most as many edges as vertices.

The construction starts from a ddd-regular graph HHH on a vertex set VVV and a set U⊆VU\subseteq VU⊆V. Write N(U)N(U)N(U) for the set of neighbours of UUU, and let mmm be a perfect matching on N(U)N(U)N(U). Then H′H'H′ is the subgraph induced on V∖UV\setminus UV∖U, MMM is the graph of matching edges {x,m(x)}\{x,m(x)\}{x,m(x)}, and G=H′∪MG=H'\cup MG=H′∪M.

Formalization targets

Goal: Theorem 1.3, relative to the input graph

Let d≥3d\ge3d≥3, ε>0\varepsilon>0ε>0, r=⌈2/ε⌉r=\lceil 2/\varepsilon\rceilr=⌈2/ε⌉. Suppose HHH is an (N,d,2d−1+ε/2)(N,d,2\sqrt{d-1}+\varepsilon/2)(N,d,2d−1​+ε/2)-graph in which the (2r+4)(2r+4)(2r+4)-neighbourhood of every vertex contains at most one cycle, and r≤log⁡d−1Nr\le\log_{d-1}Nr≤logd−1​N. Then for every uuu with ududud even and u≤N/(2d2r+3)u\le N/(2d^{2r+3})u≤N/(2d2r+3),

∃ G on N−u vertices:G is an (N−u, d, 2d−1+ε)-graph.\exists\, G \text{ on } N-u \text{ vertices}:\quad G \text{ is an } \bigl(N-u,\ d,\ 2\sqrt{d-1}+\varepsilon\bigr)\text{-graph}.∃G on N−u vertices:G is an (N−u, d, 2d−1​+ε)-graph.

The hypotheses on HHH are what Theorem 3.3 (Mohanty–O'Donnell–Paredes) supplies, and that theorem is not formalized.

Milestones

  • Lemma 3.1 (p. 10). A ddd-regular graph whose (2r+4)(2r+4)(2r+4)-balls contain at most one cycle has a set UUU with ∣U∣≥n/(2d2r+3)|U|\ge n/(2d^{2r+3})∣U∣≥n/(2d2r+3), cycle-free (r+1)(r+1)(r+1)-balls, and pairwise distances ≥2r+3\ge 2r+3≥2r+3.
  • Lemma 3.2 (p. 11). If the rrr-neighbourhood of an edge uvuvuv contains no cycle and Af=μfA f=\mu fAf=μf with μ≥2d−1\mu\ge2\sqrt{d-1}μ≥2d−1​, then
∑w∈Nif2(w) ≥ ∑w∈Ni−1f2(w),1≤i≤r.\sum_{w\in N_i}f^2(w)\ \ge\ \sum_{w\in N_{i-1}}f^2(w),\qquad 1\le i\le r .w∈Ni​∑​f2(w) ≥ w∈Ni−1​∑​f2(w),1≤i≤r.
  • The variational characterization of nontrivial eigenvalues (§2.4, p. 8).
  • In G=H′∪MG=H'\cup MG=H′∪M: GGG is ddd-regular on ∣V∣−∣U∣|V|-|U|∣V∣−∣U∣ vertices and AG=AH′+AMA_G=A_{H'}+A_MAG​=AH′​+AM​. Matching edges have cycle-free (r−1)(r-1)(r−1)-neighbourhoods and pairwise disjoint rrr-neighbourhoods.
  • Inequalities (9), (10), (11) (p. 13), and the spectral step: for every admissible UUU and mmm, GGG is an (N−∣U∣,d,2d−1+ε)(N-|U|,d,2\sqrt{d-1}+\varepsilon)(N−∣U∣,d,2d−1​+ε)-graph.

Significance

Theorem 1.3 shows that the size restrictions of algebraic Ramanujan constructions cost nothing spectrally: up to an arbitrarily small ε\varepsilonε, the Alon–Boppana bound is attained by explicit graphs on every admissible vertex count. The deletion method is local. It turns any near-Ramanujan graph whose short cycles are sparse into graphs of all nearby sizes, so it applies to future constructions as well. Lemma 3.2 is a self-contained delocalization statement in the tradition of Kahale 1995: eigenvectors of eigenvalues at least 2d−12\sqrt{d-1}2d−1​ in absolute value cannot concentrate near tree-like edges.

The result is proved on paper. To our knowledge none of it is formalized; Mathlib has adjacency matrices, extended graph distance and acyclicity, but no theory of expanders. A complete development would give machine-checked versions of a delocalization lemma, of the greedy selection of far-apart vertices away from short cycles, and of the variational eigenvalue bound for induced subgraphs. It would also check two points the paper passes over. The proof of Theorem 1.3 treats only positive eigenvalues λ≥2d−1\lambda\ge2\sqrt{d-1}λ≥2d−1​. And its claim that the rrr-neighbourhood of a matching edge is cycle-free fails when two deleted vertices are at distance exactly 2r+32r+32r+3. This mission states the corrected forms (see Formalization scope).

Difficulty

The spectral bound for GGG does not follow from interlacing alone. Deleting vertices is harmless, since by (9) the quadratic form of H′H'H′ is controlled by HHH. But the added matching contributes up to ∑x∈N(U)f(x)2\sum_{x\in N(U)}f(x)^2∑x∈N(U)​f(x)2 to ftAGff^tA_GfftAG​f, which can be as large as ∥f∥2\|f\|^2∥f∥2 for an eigenvector concentrated on N(U)N(U)N(U). The obvious estimate therefore gives only λ≤2d−1+1+ε/2\lambda\le 2\sqrt{d-1}+1+\varepsilon/2λ≤2d−1​+1+ε/2. Closing the gap requires showing that an eigenvector of a large eigenvalue spreads its mass over the rrr layers around each matching edge (Lemma 3.2). That in turn needs those neighbourhoods to be trees in GGG and pairwise disjoint, which is where Lemma 3.1's choice of UUU is used. The combinatorial part, tracking distances and cycles in GGG when GGG mixes edges of HHH with matching edges, is the main formalization burden.

Formalization scope

  • Representation. Vertex sets are finite types. Graphs are Mathlib SimpleGraphs with real adjacency matrices adjMatrix ℝ. The (n, d, λ) predicate requires IsRegularOfDegree d, symmetry, row sums ddd, and ∣μ∣≤λ|\mu|\le\lambda∣μ∣≤λ for every eigenpair (μ,f)(\mu,f)(μ,f) with f≠0f\ne0f=0, ∑f=0\sum f=0∑f=0. Distances use the extended SimpleGraph.edist, never dist (which is 000 across components). Cycle conditions are on induced subgraphs, and "at most one cycle" on a ball is ∣E∣≤∣V∣|E|\le|V|∣E∣≤∣V∣. Deleted vertices are a Finset U; the new graph lives on the subtype {v // v ∉ U}. The matching is a fixed-point-free involution of N(U)N(U)N(U).
  • Explicit quantities replacing the paper's asymptotics. The paper writes "sufficiently large nnn" and u=o(n)u=o(n)u=o(n). The goal instead takes any u≤N/(2d2r+3)u\le N/(2d^{2r+3})u≤N/(2d2r+3) (the size Lemma 3.1 guarantees) with ududud even, plus Lemma 3.1's side condition r≤log⁡d−1Nr\le\log_{d-1}Nr≤logd−1​N. The equality r=⌈2/ε⌉r=\lceil2/\varepsilon\rceilr=⌈2/ε⌉ is used as ⌈2/ε⌉+∈N\lceil2/\varepsilon\rceil_+\in\mathbb N⌈2/ε⌉+​∈N. The paper's "every degree ddd" becomes d≥3d\ge3d≥3, the range of its proof.
  • Corrections. (11) and Lemma 3.2's companion are stated for ∣μ∣≥2d−1|\mu|\ge2\sqrt{d-1}∣μ∣≥2d−1​, both signs. The matching-edge note is stated for the (r−1)(r-1)(r−1)-neighbourhood, which still yields the factor 1/r1/r1/r in (11). Lemma 3.2 itself is stated as printed.
  • Out of scope. Theorem 3.3 ([18]) is a cited input: its graph is the hypothesis HHH. All claims of explicitness and polynomial running time are out of scope, as is §4's remark on applying the method to LPS graphs directly.
  • No trivialization. The input hypotheses are exactly Theorem 3.3's conclusions plus Lemma 3.1's side condition, and they are met by high-girth Ramanujan graphs. No hypothesis mentions the spectrum or Rayleigh quotients of the constructed graph, and the goal's graph must be ddd-regular on exactly N−uN-uN−u vertices.
  • Reusable infrastructure. Welcome contributions include the variational characterization for symmetric matrices with constant row sums, a forest edge-count lemma for balls, BFS-layer structure of cycle-free balls in regular graphs, and the edge-disjoint decomposition AG=AH′+AMA_{G}=A_{H'}+A_MAG​=AH′​+AM​. Each is useful beyond this mission.

Selected references

  • N. Alon, Explicit expanders of every degree and size, arXiv:2003.11673v1, 2020; Combinatorica 41 (2021). https://arxiv.org/abs/2003.11673
  • S. Mohanty, R. O'Donnell, P. Paredes, Explicit near-Ramanujan graphs of every degree, STOC 2020. https://arxiv.org/abs/1909.06988
  • A. Lubotzky, R. Phillips, P. Sarnak, Ramanujan graphs, Combinatorica 8 (1988). https://doi.org/10.1007/BF02126799
  • N. Alon, Eigenvalues and expanders, Combinatorica 6 (1986). https://doi.org/10.1007/BF02579166
  • A. Nilli, On the second eigenvalue of a graph, Discrete Mathematics 91 (1991). https://doi.org/10.1016/0012-365X(91)90112-F
  • N. Kahale, Eigenvalues and expansion of regular graphs, J. ACM 42 (1995). https://doi.org/10.1145/210118.210136
13 thms1 active userReviewed
CombinatoricsGraph TheoryTheoretical Computer Science·Captain: mikedeng1

Explicit Expanders of Every Degree and Size 2: Attaching New Vertices to a (p+1)-Regular Ramanujan Graph and Adding Loops Keeps Every Nontrivial Eigenvalue at Most √(2(p+1)) + √p + o(1)Research Paper

Motivation

Sparse graphs whose adjacency spectrum is concentrated near zero, expanders, are used throughout theoretical computer science: in error-correcting codes, derandomization, sorting and routing networks, and the construction of pseudorandom objects (Hoory, Linial and Wigderson, survey). The best possible spectral expansion for a ddd-regular graph is governed by the Alon–Boppana bound 2d−12\sqrt{d-1}2d−1​, and graphs attaining it, Ramanujan graphs, were constructed explicitly by Lubotzky, Phillips and Sarnak (LPS 1988) and by Margulis. These constructions exist only for special degrees (d=p+1d = p+1d=p+1 with ppp prime) and special numbers of vertices (orders of PSL(2,Fq)PSL(2,\mathbb F_q)PSL(2,Fq​) or SL(2,Fq)SL(2,\mathbb F_q)SL(2,Fq​)). Applications often need a graph of a prescribed size nnn.

N. Alon's paper Explicit expanders of every degree and size (arXiv:2003.11673v1; Combinatorica 41, 2021) shows how to obtain explicit near-Ramanujan graphs on exactly nnn vertices. This mission formalizes the spectral core of its Theorem 1.2: a Ramanujan graph on mmm vertices can be enlarged to n=m+rn = m + rn=m+r vertices, with degree raised by one, while the nontrivial eigenvalues stay within a constant factor of optimal.

Setting

Let VVV be a finite set of m≥1m \ge 1m≥1 vertices. A (n,d,λ)(n,d,\lambda)(n,d,λ)-graph is a ddd-regular graph on nnn vertices whose adjacency matrix AAA satisfies ∣μ∣≤λ|\mu| \le \lambda∣μ∣≤λ for every nontrivial eigenvalue μ\muμ, that is, every eigenvalue other than the top eigenvalue ddd of the constant vector 1\mathbf 11. For a symmetric AAA with A1=d 1A\mathbf 1 = d\,\mathbf 1A1=d1, the nontrivial eigenvalues are those with an eigenvector f≠0f \ne 0f=0 satisfying ∑vf(v)=0\sum_v f(v) = 0∑v​f(v)=0. Graphs may carry loops, at most one per vertex, and a loop adds one to the degree: it is a diagonal entry 111 of AAA.

Fix an integer p≥0p \ge 0p≥0 and let HHH be an (m,p+1,2p)(m, p+1, 2\sqrt p)(m,p+1,2p​)-graph on VVV, a (p+1)(p+1)(p+1)-regular Ramanujan graph. Let R={u1,…,ur}R = \{u_1, \dots, u_r\}R={u1​,…,ur​} be rrr new vertices and let W1,…,Wr⊆VW_1, \dots, W_r \subseteq VW1​,…,Wr​⊆V be pairwise disjoint sets of p+2p+2p+2 vertices each. Put W=⋃iWiW = \bigcup_i W_iW=⋃i​Wi​ and L=V∖WL = V \setminus WL=V∖W. The graph GGG on U=V∪RU = V \cup RU=V∪R is obtained from HHH by joining each uiu_iui​ to every vertex of WiW_iWi​ and adding one loop at each vertex of LLL. Its adjacency matrix is

AG=AH+AR+AL,A_G = A_H + A_R + A_L,AG​=AH​+AR​+AL​,

where AHA_HAH​ is the adjacency matrix of HHH (zero on RRR), ARA_RAR​ that of the stars joining uiu_iui​ to WiW_iWi​, and ALA_LAL​ the diagonal matrix of the loops. Every vertex of GGG has degree p+2p+2p+2.

Formalization targets

Goal: Theorem 1.2, spectral core

AG is an (m+r,  p+2,  2(p+1)+p+(p+1) rm) matrix.A_G \text{ is an } \Big(m+r,\; p+2,\; \sqrt{2(p+1)} + \sqrt p + \frac{(p+1)\,r}{m}\Big)\text{ matrix.}AG​ is an (m+r,p+2,2(p+1)​+p​+m(p+1)r​) matrix.

The paper states λ≤2(d−1)+d−1+o(1)\lambda \le \sqrt{2(d-1)} + \sqrt{d-1} + o(1)λ≤2(d−1)​+d−1​+o(1) for d=p+2d = p+2d=p+2; its proof gives 2(p+1)+p+o(1)\sqrt{2(p+1)} + \sqrt p + o(1)2(p+1)​+p​+o(1), which is stronger, and the error term it produces is (p+1)r/m(p+1)r/m(p+1)r/m. The goal is parametrised by HHH, rrr and the sets WiW_iWi​, so it does not depend on how mmm and rrr are chosen.

Milestones

  1. The variational characterization of the nontrivial eigenvalues: for a symmetric matrix with constant row sums and λ≥0\lambda \ge 0λ≥0, ∣μ∣≤λ|\mu| \le \lambda∣μ∣≤λ for every nontrivial eigenvalue if and only if ∣ftAf∣≤λ∥f∥2|f^tAf| \le \lambda\|f\|^2∣ftAf∣≤λ∥f∥2 whenever ∑f=0\sum f = 0∑f=0.
  2. The Cauchy–Schwarz display: ∑Uf=0\sum_U f = 0∑U​f=0 implies ∣∑Vf∣2=∣∑Rf∣2≤∣R∣∑Rf2|\sum_V f|^2 = |\sum_R f|^2 \le |R| \sum_R f^2∣∑V​f∣2=∣∑R​f∣2≤∣R∣∑R​f2.
  3. Inequality (3): ∣ftAHf∣≤b2(p+1)+c2 2p|f^tA_Hf| \le b^2(p+1) + c^2\, 2\sqrt p∣ftAH​f∣≤b2(p+1)+c22p​ with b2=(∑Vf)2/mb^2 = (\sum_V f)^2/mb2=(∑V​f)2/m and c2=∑Vf2−b2c^2 = \sum_V f^2 - b^2c2=∑V​f2−b2.
  4. Display (4): ftALf=∑v∈Lf2(v)f^tA_Lf = \sum_{v\in L} f^2(v)ftAL​f=∑v∈L​f2(v).
  5. Inequality (5): ∣ftARf∣≤p+2x∑Rf2+x∑Wf2|f^tA_Rf| \le \frac{p+2}{x}\sum_R f^2 + x\sum_W f^2∣ftAR​f∣≤xp+2​∑R​f2+x∑W​f2 for every x>0x > 0x>0.
  6. Inequality (6): for ∑Uf=0\sum_U f = 0∑U​f=0 and x>0x > 0x>0,
∣ftAGf∣≤(2p+1)∑Lf2+(2p+x)∑Wf2+p+2x∑Rf2+(p+1)rm∑Rf2.|f^tA_Gf| \le (2\sqrt p+1)\sum_L f^2 + (2\sqrt p+x)\sum_W f^2 + \frac{p+2}{x}\sum_R f^2 + (p+1)\frac rm \sum_R f^2.∣ftAG​f∣≤(2p​+1)L∑​f2+(2p​+x)W∑​f2+xp+2​R∑​f2+(p+1)mr​R∑​f2.

Significance

With HHH the Lubotzky–Phillips–Sarnak graph on m=∣SL(2,Fq)∣m = |SL(2,\mathbb F_q)|m=∣SL(2,Fq​)∣ vertices for the largest suitable prime qqq with m≤nm \le nm≤n, and r=n−mr = n - mr=n−m, the distribution of primes in arithmetic progressions gives r=o(m)r = o(m)r=o(m), and the goal yields an explicit (n,p+2,λ)(n, p+2, \lambda)(n,p+2,λ)-graph with λ≤(1+2)d−1+o(1)\lambda \le (1+\sqrt2)\sqrt{d-1} + o(1)λ≤(1+2​)d−1​+o(1) for every sufficiently large nnn. This is within a factor of about 1.211.211.21 of the Ramanujan bound 2d−12\sqrt{d-1}2d−1​, for every number of vertices, by an elementary modification of an existing graph. The statement is useful independently of LPS: any Ramanujan graph, or any graph with a bound on its nontrivial eigenvalues, can be padded to a nearby size in the same way.

The result is proved in the paper. No formalization of it, of the (n,d,λ)(n,d,\lambda)(n,d,λ) notion, or of the variational characterization of nontrivial eigenvalues for regular graphs exists on the platform. The mission produces a checked version of the spectral argument, and the variational characterization (milestone 1) is a general fact about symmetric matrices with constant row sums that applies to any spectral expander argument.

Difficulty

The vertices of WWW and LLL lie in the old graph HHH, whose spectrum is controlled, but the new vertices of RRR are not; and a vector orthogonal to 1\mathbf 11 on UUU need not be orthogonal to the constant vector on VVV. Bounding ftAGff^tA_GfftAG​f by applying the Ramanujan bound for HHH to fff restricted to VVV therefore fails: the restriction has a component along the trivial eigenvector of HHH, whose eigenvalue p+1p+1p+1 is large. The argument must show that this component is small, of order r/mr/mr/m, and must balance the star edges between RRR and WWW against the loops on LLL so that every vertex class gets the same coefficient. The naive bound ∣ftARf∣≤∥AR∥ ∥f∥2=p+2 ∥f∥2|f^tA_Rf| \le \|A_R\|\,\|f\|^2 = \sqrt{p+2}\,\|f\|^2∣ftAR​f∣≤∥AR​∥∥f∥2=p+2​∥f∥2 added to 2p2\sqrt p2p​ for HHH and 111 for LLL gives a constant larger than 2(p+1)+p\sqrt{2(p+1)}+\sqrt p2(p+1)​+p​; the stated constant needs the weighted estimate.

On the Lean side, milestone 1 concerns the spectrum of a symmetric matrix on the invariant subspace 1⊥\mathbf 1^\perp1⊥, while Mathlib states the spectral theorem for the whole space.

Formalization scope

  • Vertices of GGG are the disjoint union V⊕Fin rV \oplus \mathrm{Fin}\, rV⊕Finr. GGG is represented by its real adjacency matrix, since it has loops; HHH is a Mathlib SimpleGraph with adjMatrix.
  • The (n,d,λ)(n,d,\lambda)(n,d,λ) predicate is stated for matrices: ∣V∣=n|V| = n∣V∣=n, symmetry, A1=d 1A\mathbf 1 = d\,\mathbf 1A1=d1, and ∣μ∣≤λ|\mu| \le \lambda∣μ∣≤λ for every eigenpair (μ,f)(\mu, f)(μ,f) with f≠0f \ne 0f=0 and ∑f=0\sum f = 0∑f=0. For simple graphs, ddd-regularity is added.
  • The paper's o(1)o(1)o(1) terms are replaced by the explicit quantities its proof produces: (p+1)r/m(p+1)r/m(p+1)r/m in the goal, and (p+1)rm∑Rf2(p+1)\frac rm\sum_R f^2(p+1)mr​∑R​f2 in (6). Inequality (3) is stated with the corrected relation b2+c2=∑Vf2b^2 + c^2 = \sum_V f^2b2+c2=∑V​f2; the paper's "b2+c2=1b^2 + c^2 = 1b2+c2=1" holds only for unit restrictions.
  • The bound uses p=d−2\sqrt p = \sqrt{d-2}p​=d−2​, as in the proof and the abstract, which implies the printed d−1\sqrt{d-1}d−1​.
  • ppp is any natural number. The hypothesis "ppp prime, p≡1(mod4)p \equiv 1 \pmod 4p≡1(mod4)" serves only to obtain HHH from LPS, and HHH is a hypothesis here. The sets WiW_iWi​ are arbitrary pairwise disjoint sets of size p+2p+2p+2, not the paper's consecutive blocks of a numbering of SL(2,Fq)SL(2,\mathbb F_q)SL(2,Fq​).
  • Out of scope: the existence of the prime qqq and the estimate n−m=o(m)n - m = o(m)n−m=o(m); the numbering of SL(2,Fq)SL(2,\mathbb F_q)SL(2,Fq​); the "strongly explicit" and polynomial-time claims; the LPS construction (Theorem 2.1, cited); the variant that replaces loops by a matching for even nnn.
  • The goal cannot be satisfied trivially: dropping the condition ∑f=0\sum f = 0∑f=0 makes it false, since p+2p+2p+2 is always an eigenvalue, and for p≥2p \ge 2p≥2 and small r/mr/mr/m the bound is below p+2p + 2p+2 (for p=5p = 5p=5 it is about 5.70+6r/m5.70 + 6r/m5.70+6r/m). At p=1p = 1p=1 the bound 3+2r/m3 + 2r/m3+2r/m is at least the degree 333, so that case holds trivially; it is the paper's statement there as well.
  • Contributions welcome: proofs of each milestone, especially the variational characterization, which is reusable for mission 3 of this series and for any regular-graph spectral argument.

Selected references

  • N. Alon, Explicit expanders of every degree and size, arXiv:2003.11673v1, 2020; Combinatorica 41 (2021). https://arxiv.org/abs/2003.11673
  • A. Lubotzky, R. Phillips, P. Sarnak, Ramanujan graphs, Combinatorica 8 (1988) 261–277. https://doi.org/10.1007/BF02126799
  • S. Hoory, N. Linial, A. Wigderson, Expander graphs and their applications, Bull. AMS 43 (2006) 439–561. https://doi.org/10.1090/S0273-0979-06-01126-8
9 thms1 active userReviewed
CombinatoricsOperations Research·Captain: mikedeng1

On the Abstract Properties of Linear Dependence 5: The Seven-Element Fano Matroid Corresponds to No Real MatrixResearch Paper

Motivation

Whitney's 1935 paper On the Abstract Properties of Linear Dependence introduced matroids: finite sets equipped with a rank function, or equivalently a family of independent sets, obeying a few postulates abstracted from the linear dependence of the columns of a matrix. The obvious first question about such an abstraction is whether it is genuinely more general than its model, that is, whether there are matroids that do not arise from any matrix. Section 16 of the paper answers it with a seven-element example, now called the Fano matroid F7F_7F7​, and proves that no real matrix corresponds to it.

The question has had a long life. Representability of matroids over a given field is a central theme of matroid theory: Tutte (1958) characterized the matroids representable over the field with two elements by a single excluded minor, the four-point line U2,4U_{2,4}U2,4​, and the regular matroids by three excluded minors, U2,4U_{2,4}U2,4​, F7F_7F7​ and its dual; and Seymour's decomposition of regular matroids (1980) rests on the same objects. Whitney's §16 is the starting point of this line: the first proof that the abstract postulates admit matroids outside linear algebra over R\mathbb RR.

Timeline:

  • 1935. Whitney defines matroids, the circuit matrix of a matrix, and proves (§16) that the seven-element matroid M′M'M′ corresponds to no real matrix; in a footnote he credits Saunders MacLane with finding that M′M'M′ corresponds to no matrix and identifying it with a finite projective geometry. On p. 533 he exhibits a matrix of integers mod 2 for M′M'M′.
  • 1958. Tutte characterizes binary and regular matroids by excluded minors; F7F_7F7​ appears as an excluded minor for regularity (Tutte 1958).

Setting

Let M=(aij)\mathbf M=(a_{ij})M=(aij​) be an m×nm\times nm×n matrix with columns C1,…,CnC_1,\dots,C_nC1​,…,Cn​. For a set NNN of columns, let r(N)r(N)r(N) be the rank of the submatrix formed by those columns. Regarding the columns as abstract elements gives a matroid MMM on {C1,…,Cn}\{C_1,\dots,C_n\}{C1​,…,Cn​} with rank function rrr: the matroid of M\mathbf MM. A matroid corresponds to M\mathbf MM if it is the matroid of M\mathbf MM, with elements matched to columns.

A circuit of a matroid is a minimal dependent set. For a circuit P={i1,…,ip}P=\{i_1,\dots,i_p\}P={i1​,…,ip​} of the matroid of M\mathbf MM, there are numbers b1,…,bnb_1,\dots,b_nb1​,…,bn​ with ∑jaijbj=0\sum_j a_{ij}b_j=0∑j​aij​bj​=0 for every row iii, and bj≠0b_j\neq 0bj​=0 exactly for j∈Pj\in Pj∈P; the set of such vectors is written Zi1⋯ipZ_{i_1\cdots i_p}Zi1​⋯ip​​ when only the support condition is meant. Stacking one such row per circuit gives the circuit matrix M′\mathbf M'M′ of M\mathbf MM, determined up to nonzero factors on its rows.

A fundamental set of circuits of a matroid MMM with nullity n(M)=ρ(M)−r(M)n(M)=\rho(M)-r(M)n(M)=ρ(M)−r(M) (ρ\rhoρ the number of elements) is a family of circuits P1,…,PqP_1,\dots,P_qP1​,…,Pq​ with q=n(M)q=n(M)q=n(M) such that the elements can be ordered e1,…,ene_1,\dots,e_ne1​,…,en​ with en−q+i∈Pie_{n-q+i}\in P_ien−q+i​∈Pi​ and en−q+j∉Pie_{n-q+j}\notin P_ien−q+j​∈/Pi​ for j>ij>ij>i; it is strict if en−q+j∉Pie_{n-q+j}\notin P_ien−q+j​∈/Pi​ for every j≠ij\neq ij=i.

The matroid M′M'M′ of §16 has elements 1,…,71,\dots,71,…,7; its bases (maximal independent sets) are all three-element sets except

124,135,167,236,257,347,456.(16.1)124,\quad 135,\quad 167,\quad 236,\quad 257,\quad 347,\quad 456. \qquad (16.1)124,135,167,236,257,347,456.(16.1)

Formalization targets

Goal: §16, pp. 529–530

∃ M′and∀m ∀ M∈Rm×7: M′ is not the matroid of M.\exists\,M' \quad\text{and}\quad \forall m\ \forall\,\mathbf M\in\mathbb R^{m\times 7}:\ M' \text{ is not the matroid of } \mathbf M .∃M′and∀m ∀M∈Rm×7: M′ is not the matroid of M.

The number of rows is arbitrary; the existence clause makes the non-existence statement non-vacuous.

Milestones

  1. §12. Every real matrix has a matroid: the ranks of column submatrices satisfy the rank postulates.
  2. §14, (14.1). Every real matrix has a circuit matrix.
  3. Theorem 29. The rows of a fundamental set of circuits form a base for the rows of the circuit matrix, so r(M′)=q=n(M)r(\mathbf M')=q=n(\mathbf M)r(M′)=q=n(M).
  4. Lemma 10. The support of a vector in the row space HHH of a circuit matrix is a union of circuits.
  5. Lemma 11. Two vectors of HHH with the same circuit as support are proportional.
  6. Theorem 32. For a circuit matrix normalised along a strict fundamental set, a minor DDD vanishes iff an associated q×qq\times qq×q minor D′D'D′ vanishes, iff some circuit avoids a prescribed set of columns.
  7. §16, rank of M′M'M′. The rank of a kkk-set is kkk for k≤2k\le 2k≤2, 333 for k≥4k\ge 4k≥4, and for k=3k=3k=3 it is 222 on (16.1) and 333 otherwise.
  8. p. 533. M′M'M′ is the matroid of an explicit 3×73\times 73×7 matrix of integers mod 2.

Significance

The result. The theorem separates the abstract notion of matroid from linear dependence over R\mathbb RR: some matroids are not real-representable. It also exhibits that representability depends on the field, because the same matroid is the matroid of a matrix over the integers mod 2 (milestone 8). Everything later written about representability over particular fields, excluded-minor characterizations, and the gap between abstract and linear matroids starts from this distinction. Theorem 32 is of independent interest: it translates statements about circuits of a represented matroid into the vanishing of minors of a normalised circuit matrix.

Formalizing it. The result is classical and its proof is short on paper, but it is not formalized in Mathlib, which has matroids (Matroid, circuits, ranks) but no column matroid of a matrix with a rank-of-submatrix characterization, no circuit matrix, and no Fano matroid. The mission produces those objects and the bridge lemmas (Theorem 29, Lemmas 10–11, Theorem 32) that connect matroid circuits with linear algebra of the circuit matrix. No machine-checked proof of the non-representability of the Fano matroid over R\mathbb RR in Lean is known to the curators.

Difficulty

The obvious attempt is a direct search: suppose a real m×7m\times 7m×7 matrix has M′M'M′ as its matroid and derive a contradiction from the seven dependent triples. This does not work as stated. Each rank condition is a determinantal (nonlinear) condition on the entries, the number of rows mmm is unbounded, and a representation is determined only up to row operations and column scalings, so there is no finite case check and no single linear computation that settles the question. The contradiction has to come from an argument that is invariant under these symmetries, and the milestones (circuit vectors determined up to scaling, fundamental sets spanning, circuits detected by minors) are what such an argument needs to be stated in. The field also matters: the argument must use that 2≠02\neq 02=0 in R\mathbb RR, since over a field of characteristic 2 the statement is false (milestone 8).

Formalization scope

  • Elements and matrices. Matroids are Mathlib Matroids whose ground set is the whole (finite) type. The Fano matroid lives on Fin 7, Whitney's element kkk being k - 1; the seven triples are written out literally. Matrices are Matrix (Fin m) ι K; "the matroid of M\mathbf MM" means: ground set everything, and the rank M.eRk N of every finite set NNN of columns equals Matrix.rank of the column submatrix.
  • Field. The goal and Lemmas 10–11, Theorems 29 and 32 are stated over R\mathbb RR, as in the paper; the predicate "matroid of a matrix" is stated over any field so that the mod-2 milestone uses the same notion.
  • Circuit matrix. Rows are determined up to nonzero factors, so "circuit matrix" is a predicate on a matrix together with a bijection between its rows and the circuits; every theorem holds for every such choice.
  • Nullity and indices. q=n(M)q=n(M)q=n(M) is written q+r(M)=ρ(M)q+r(M)=\rho(M)q+r(M)=ρ(M) in extended naturals, with no truncated subtraction. In Theorem 32, n=p+qn=p+qn=p+q, the complement of i1,…,isi_1,\dots,i_si1​,…,is​ is given as an order embedding of Fin t with s+t=qs+t=qs+t=q, and determinants are of square submatrices in the paper's row and column order.
  • Ruling out trivial readings. The goal includes the existence of M′M'M′; without it "every matroid with these bases has no real matrix" could hold vacuously. The goal quantifies over every number of rows; fixing m=3m=3m=3 would be a weaker statement.

Reusable beyond this mission: the matroid of a matrix over a field, the circuit matrix, fundamental sets of circuits, and the Fano matroid. Contributions welcome: proofs of the milestones, and a proof of the goal by any route, including one that does not go through Theorem 32.

Selected references

  • H. Whitney, On the Abstract Properties of Linear Dependence, American Journal of Mathematics 57 (1935), 509–533. https://doi.org/10.2307/2371182
  • W. T. Tutte, A homotopy theorem for matroids, I, II, Transactions of the American Mathematical Society 88 (1958), 144–174. https://doi.org/10.2307/1993244
  • J. Oxley, Matroid Theory, 2nd ed., Oxford University Press, 2011. https://doi.org/10.1093/acprof:oso/9780198566946.001.0001
  • O. Veblen and J. W. Young, Projective Geometry, Vol. I, Ginn, 1910 (cited by Whitney for the finite projective geometry).
13 thms1 active userReviewed
CombinatoricsOperations Research·Captain: mikedeng1

On the Abstract Properties of Linear Dependence 6: Every Matroid Satisfying (C*) Is Represented by a Matrix of Integers Mod 2Research Paper

Motivation

Whitney's 1935 paper introduced matroids as an abstraction of linear dependence among the columns of a matrix. Most of the paper works over the real numbers; its appendix asks which matroids arise from matrices of integers mod 2, that is, matrices with entries 0 and 1 in which rank and dependence are computed over the two-element field. These are today's binary matroids. They include the cycle matroids of graphs (Whitney closes the paper by noting that graphs correspond to mod-2 matrices with exactly two ones in each column) and they are the setting of several later structure theorems: Tutte's excluded-minor characterization of binary matroids (Tutte 1958), Seymour's decomposition of regular matroids (Seymour 1980) and Seymour's theory of binary clutters and max-flow min-cut (Seymour 1977), which underlies parts of combinatorial optimization.

Whitney's answer is an intrinsic postulate, (C*), on the circuits of the matroid, stated without reference to any matrix, and a constructive representation theorem (Theorem 37): a matroid satisfying (C*) is the matroid of a mod-2 matrix, and the matrix is unique once the columns of one base are fixed.

Setting

A matroid MMM on elements e1,…,ene_1, \dots, e_ne1​,…,en​ is given by its independent sets; its circuits are its minimal dependent sets, its rank r(M)r(M)r(M) is the size of a base, and its nullity is n(M)=n−r(M)n(M) = n - r(M)n(M)=n−r(M). Here MMM is a Mathlib Matroid (Fin n) whose ground set is all of Fin n.

Subsets of the elements are added mod 2: a sum of finitely many sets is the set of elements lying in an odd number of them (for two sets, the symmetric difference). A cycle is a sum mod 2 of circuits; the empty sum is the null cycle ∅\emptyset∅. A set is a true sum of sets that have no common elements and whose union it is. Postulate (C*) requires that each cycle be a true sum of circuits.

With n=r+qn = r + qn=r+q, a family P1,…,PqP_1, \dots, P_qP1​,…,Pq​ is a strict fundamental set of circuits with respect to en−q+1,…,ene_{n-q+1}, \dots, e_nen−q+1​,…,en​ if q=n(M)q = n(M)q=n(M), each PiP_iPi​ is a circuit, and PiP_iPi​ contains en−q+ie_{n-q+i}en−q+i​ but no other en−q+je_{n-q+j}en−q+j​.

For a matrix M\mathbf MM over the integers mod 2 with columns C1,…,CnC_1, \dots, C_nC1​,…,Cn​, columns are independent (mod 2) if no non-null subset of them sums to the zero column. The matroid corresponding to M\mathbf MM has the column indices as elements and these independent sets.

Formalization targets

Goal: Theorem 37 (p. 533)

Let MMM satisfy (C*), with elements e1,…,ene_1, \dots, e_ne1​,…,en​ and base {e1,…,en−q}\{e_1, \dots, e_{n-q}\}{e1​,…,en−q​}. For every matrix M1\mathbf M_1M1​ mod 2 (any number of rows) whose n−qn - qn−q columns are independent mod 2,

∃! M=(M1∣Cn−q+1⋯Cn)  whose corresponding matroid is M.\exists!\ \mathbf M = (\mathbf M_1 \mid C_{n-q+1} \cdots C_n) \ \text{ whose corresponding matroid is } M.∃! M=(M1​∣Cn−q+1​⋯Cn​)  whose corresponding matroid is M.

Milestones

  1. Theorem 9 (p. 517): if e1,…,en−qe_1, \dots, e_{n-q}e1​,…,en−q​ is a base, there is a unique strict fundamental set of circuits with respect to en−q+1,…,ene_{n-q+1}, \dots, e_nen−q+1​,…,en​.
  2. Appendix, p. 531: (C*) implies the circuit postulate (C₂), for any family of sets.
  3. Theorem 33: under (C*), the circuits are exactly the minimal non-null cycles.
  4. Theorem 34: under (C*), the cycles are exactly the 2q2^q2q sums mod 2 of a strict fundamental set.
  5. Theorem 35: two (C*)-matroids with a common strict fundamental set have the same circuits.
  6. Theorem 36: any P1,…,PqP_1, \dots, P_qP1​,…,Pq​ with en−q+i∈Pi⊆{e1,…,en−q,en−q+i}e_{n-q+i} \in P_i \subseteq \{e_1, \dots, e_{n-q}, e_{n-q+i}\}en−q+i​∈Pi​⊆{e1​,…,en−q​,en−q+i​} is the strict fundamental set of exactly one (C*)-matroid.
  7. Appendix, p. 532: the matroid of a matrix mod 2 exists, satisfies (C*), and its cycles are the supports of the mod-2 dependencies among the columns.

Milestone 7 and the goal together characterize binary matroids as the matroids satisfying (C*).

Significance

The result. Theorem 37 and the p. 532 claim give an intrinsic, matrix-free description of the matroids representable over the two-element field, and Theorem 36 parametrizes all of them by qqq arbitrary subsets of a base. Uniqueness in Theorem 37 says that a binary representation is determined by the columns of one base; in modern terms, binary matroids are uniquely representable over GF(2) up to row operations. Every later theory of binary matroids, including graphic and cographic matroids, Tutte's excluded-minor theorem and Seymour's decomposition, starts from this equivalence.

Formalizing it. The results are proved in the paper and in textbooks (e.g. Oxley, Matroid Theory, Ch. 9) but, at the Mathlib revision used here, there is no notion of a matroid represented by a matrix over a field, and no binary-matroid theory. On Prove2Me, the existing binary objects (SeymourMFMC.Binary.*) are binary clutters defined through blockers, not matroids represented by mod-2 matrices. This mission produces the representation predicate for mod-2 matrices, the cycle space of a matroid, and the equivalence between (C*) and binary representability.

Difficulty

Writing down candidate columns is not the hard part; showing that the matroid of the completed matrix is MMM itself, and not merely a matroid sharing some of its circuits, is. Whitney's example at the end of §9 exhibits two different matroids with a common strict fundamental set, so agreement on fundamental circuits does not by itself identify a matroid; any argument must use (C*) on both the given matroid and the matroid of the matrix. A naive comparison of independent sets column by column does not close this gap. Uniqueness likewise depends on the independence mod 2 of the prescribed columns: without it, different completions can give the same matroid.

Formalization scope

  • Matroids are Mathlib Matroid (Fin (r + q)) with ground set Set.univ; Whitney's eke_kek​ is k - 1, his e1,…,en−qe_1, \dots, e_{n-q}e1​,…,en−q​ is the range of Fin.castAdd q, and en−q+ie_{n-q+i}en−q+i​ is Fin.natAdd r (i - 1). Writing n=r+qn = r + qn=r+q removes natural-number subtraction; qqq is not a free parameter, since {e1,…,er}\{e_1, \dots, e_r\}{e1​,…,er​} is required to be a base.
  • Sums mod 2 count parity of membership (sumMod2); cycles are sums over finite sets of circuits; true sums are unions over finite pairwise-disjoint sets of circuits; (C*) is SatisfiesCStar on the circuit family {C | M.IsCircuit C}. These definitions take the circuit family as a parameter, so that the (C₂) milestone is posed for an arbitrary family of sets, as Whitney poses it.
  • A strict fundamental set includes the nullity condition r(M)+q=ρ(M)r(M) + q = \rho(M)r(M)+q=ρ(M), stated in N∞\mathbb N_\inftyN∞​.
  • Matrices are Matrix (Fin m) (Fin n) (ZMod 2) with any mmm; independence mod 2 of columns is LinearIndepOn (ZMod 2) of the columns (the rows of the transpose). IsMatroidOf M A compares all independent sets, not only bases.
  • Ruled out: the goal is not satisfied by any statement that compares only the bases of one size, by an existence-only statement without uniqueness, or by real (instead of mod-2) independence.
  • Tacit hypotheses made explicit: the matroid's ground set is exactly e1,…,ene_1, \dots, e_ne1​,…,en​ (ρ(M)=n\rho(M) = nρ(M)=n); the elements and matroids are finite.

Contributions welcome: the general fact that the matroid of a vector family over a field exists (a reusable Matroid.ofFun-style construction over any field), the cycle-space lemmas, and proofs of the milestones in any order.

Selected references

  • H. Whitney, On the Abstract Properties of Linear Dependence, American Journal of Mathematics 57 (1935), 509–533. https://doi.org/10.2307/2371182
  • W. T. Tutte, A homotopy theorem for matroids, I, II, Transactions of the AMS 88 (1958), 144–174. https://doi.org/10.2307/1993244
  • P. D. Seymour, The matroids with the max-flow min-cut property, Journal of Combinatorial Theory Ser. B 23 (1977), 189–222. https://doi.org/10.1016/0095-8956(77)90031-4
  • P. D. Seymour, Decomposition of regular matroids, Journal of Combinatorial Theory Ser. B 28 (1980), 305–359. https://doi.org/10.1016/0095-8956(80)90075-1
  • J. Oxley, Matroid Theory, 2nd ed., Oxford University Press, 2011. https://doi.org/10.1093/acprof:oso/9780198566946.001.0001
11 thms1 active userReviewed
Numerical AnalysisOptimization·Captain: mikedeng1

Projection-like Retractions on Matrix Manifolds IV: Within σ_r(X̄)/2 of a Rank-r Matrix, the Truncated SVD Is the Unique Nearest Matrix of Rank Exactly rResearch Paper

Motivation

Optimization over sets of matrices of fixed rank comes up in low-rank matrix completion, model reduction, and the low-rank approximation of solutions of large matrix equations. A standard approach treats the constraint set as a Riemannian manifold and runs gradient or Newton-type methods on it (Absil, Mahony & Sepulchre 2008). Each iteration takes a step in the tangent space and then has to return to the manifold. A map that does this to first order is called a retraction, and its cost is often what decides whether a manifold method is practical.

Absil and Malick (2012) study retractions defined by projection: step to X+ZX+ZX+Z in the ambient space, then take the nearest point of the manifold. Their Proposition 3.2 shows that for any CkC^kCk submanifold (k≥2k\ge2k≥2) this projective retraction is a retraction. Section 3.2 makes it explicit for the manifold of fixed-rank matrices: near any matrix of rank rrr, the nearest matrix of rank exactly rrr is the truncated singular value decomposition (Proposition 3.3). Section 4.4 also gives a closed form for a second retraction on the same manifold, the orthographic retraction (Proposition 4.11).

Setting

Fix natural numbers nnn, mmm and r≥1r\ge1r≥1. The space Rn×m\mathbb R^{n\times m}Rn×m of real n×mn\times mn×m matrices carries the Frobenius norm ∥X∥2=∑i,jXij2=trace⁡(X⊤X)\|X\|^2=\sum_{i,j}X_{ij}^2=\operatorname{trace}(X^\top X)∥X∥2=∑i,j​Xij2​=trace(X⊤X) (3.6). The fixed-rank set is

Rr={X∈Rn×m: rank⁡(X)=r},\mathcal R_r=\{X\in\mathbb R^{n\times m}:\ \operatorname{rank}(X)=r\},Rr​={X∈Rn×m: rank(X)=r},

a smooth submanifold of Rn×m\mathbb R^{n\times m}Rn×m. It is not closed: its closure is the set of matrices of rank at most rrr.

A singular value decomposition (3.5) of XXX is a factorization X=UΣV⊤X=U\Sigma V^\topX=UΣV⊤ in which U=[u1,…,un]∈Rn×nU=[u_1,\dots,u_n]\in\mathbb R^{n\times n}U=[u1​,…,un​]∈Rn×n and V=[v1,…,vm]∈Rm×mV=[v_1,\dots,v_m]\in\mathbb R^{m\times m}V=[v1​,…,vm​]∈Rm×m are orthogonal and Σ∈Rn×m\Sigma\in\mathbb R^{n\times m}Σ∈Rn×m is zero off its diagonal. The diagonal of Σ\SigmaΣ holds the singular values of XXX in nonincreasing order,

σ1(X)≥σ2(X)≥⋯≥σmin⁡{n,m}(X)≥0,\sigma_1(X)\ge\sigma_2(X)\ge\cdots\ge\sigma_{\min\{n,m\}}(X)\ge0,σ1​(X)≥σ2​(X)≥⋯≥σmin{n,m}​(X)≥0,

and σi(X)=0\sigma_i(X)=0σi​(X)=0 for i>min⁡{n,m}i>\min\{n,m\}i>min{n,m}. A matrix has rank rrr exactly when σr(X)>0=σr+1(X)\sigma_r(X)>0=\sigma_{r+1}(X)σr​(X)>0=σr+1​(X). The truncated SVD of rank rrr is X^=∑i=1rσi(X)uivi⊤\hat X=\sum_{i=1}^r\sigma_i(X)u_iv_i^\topX^=∑i=1r​σi​(X)ui​vi⊤​ (3.7).

For a set QQQ and a point XXX, the projection PQ(X)P_Q(X)PQ​(X) is the set of nearest points: the Y∈QY\in QY∈Q with ∥X−Y∥≤∥X−W∥\|X-Y\|\le\|X-W\|∥X−Y∥≤∥X−W∥ for all W∈QW\in QW∈Q. For a set that is not closed, PQ(X)P_Q(X)PQ​(X) may be empty or contain several points.

Formalization targets

Goal: Proposition 3.3

Let Xˉ∈Rr\bar X\in\mathcal R_rXˉ∈Rr​. For every XXX with ∥X−Xˉ∥<σr(Xˉ)/2\|X-\bar X\|<\sigma_r(\bar X)/2∥X−Xˉ∥<σr​(Xˉ)/2 and every singular value decomposition X=UΣV⊤X=U\Sigma V^\topX=UΣV⊤,

PRr(X)={∑i=1rσi(X) uivi⊤}.P_{\mathcal R_r}(X)=\Bigl\{\sum_{i=1}^r\sigma_i(X)\,u_iv_i^\top\Bigr\}.PRr​​(X)={i=1∑r​σi​(X)ui​vi⊤​}.

The projection exists, is unique, and is the truncated SVD. The radius σr(Xˉ)/2\sigma_r(\bar X)/2σr​(Xˉ)/2 and the strict inequality are those of the paper.

Milestones

The paper's proof goes through four claims, which are the milestones in attack order:

  1. Weyl's bound (§3.2, proof of Proposition 3.3, citing Horn–Johnson 7.3.8): ∣σi(Xˉ)−σi(X)∣≤∥X−Xˉ∥|\sigma_i(\bar X)-\sigma_i(X)|\le\|X-\bar X\|∣σi​(Xˉ)−σi​(X)∣≤∥X−Xˉ∥ for every i≥1i\ge1i≥1.
  2. Eckart–Young (3.7): for every singular value decomposition of XXX, X^\hat XX^ is a nearest matrix to XXX of rank at most rrr.
  3. The gap (3.8): if rank⁡Xˉ=r\operatorname{rank}\bar X=rrankXˉ=r and ∥X−Xˉ∥<σr(Xˉ)/2\|X-\bar X\|<\sigma_r(\bar X)/2∥X−Xˉ∥<σr​(Xˉ)/2, then σr+1(X)<σr(Xˉ)/2<σr(X)\sigma_{r+1}(X)<\sigma_r(\bar X)/2<\sigma_r(X)σr+1​(X)<σr​(Xˉ)/2<σr​(X).
  4. Uniqueness under a gap (§3.2, proof of Proposition 3.3, citing Helmke–Moore): if σr(X)>σr+1(X)\sigma_r(X)>\sigma_{r+1}(X)σr​(X)>σr+1​(X), then X^\hat XX^ is the only nearest matrix of rank at most rrr.

Further result: Proposition 4.11

Write X=U[Σ0000]V⊤X=U\left[\begin{smallmatrix}\Sigma_0&0\\0&0\end{smallmatrix}\right]V^\topX=U[Σ0​0​00​]V⊤ (4.7), with Σ0\Sigma_0Σ0​ the positive diagonal of nonzero singular values, and a tangent vector Z=U[ACB0]V⊤Z=U\left[\begin{smallmatrix}A&C\\B&0\end{smallmatrix}\right]V^\topZ=U[AB​C0​]V⊤ (4.8). If Σ0+A\Sigma_0+AΣ0​+A is invertible, the orthographic retraction is

R(X,Z)=U[Σ0+ACBB(Σ0+A)−1C]V⊤.R(X,Z)=U\begin{bmatrix}\Sigma_0+A&C\\B&B(\Sigma_0+A)^{-1}C\end{bmatrix}V^\top .R(X,Z)=U[Σ0​+AB​CB(Σ0​+A)−1C​]V⊤.

Here R(X,Z)R(X,Z)R(X,Z) is the nearest point to X+ZX+ZX+Z of Rr∩(X+Z+NRr(X))\mathcal R_r\cap(X+Z+N_{\mathcal R_r}(X))Rr​∩(X+Z+NRr​​(X)).

Significance

The result. Proposition 3.3 reduces the projective retraction on Rr\mathcal R_rRr​ to one truncated singular value decomposition. With Proposition 3.2 this gives an explicit, computable retraction for Riemannian optimization on fixed-rank matrices. It is used, for instance, in low-rank matrix completion algorithms. The point is local: Rr\mathcal R_rRr​ is not closed, so far from Rr\mathcal R_rRr​ the nearest matrix of rank exactly rrr need not exist, and the proposition gives an explicit radius on which it does exist and is unique. Proposition 4.11 gives a second retraction that needs only products of matrices and one r×rr\times rr×r inverse.

Formalizing it. All results here are proved, on paper. To the best of a search of Mathlib and the Prove2Me catalog, none is machine-checked. Mathlib has singular values of linear maps between finite-dimensional inner product spaces (LinearMap.singularValues), but no singular value decomposition in matrix form, no Weyl perturbation inequality for singular values, and no Eckart–Young theorem. A complete development would add these, and the Eckart–Young theorem with its uniqueness case is a standard result of numerical linear algebra in its own right.

Difficulty

The proof in the paper is short only because it cites three facts, and each of them is a real piece of matrix analysis.

  • Weyl's bound needs the variational (min–max) description of singular values, which Mathlib does not have for singular values.
  • Eckart–Young requires comparing ∥X−Y∥\|X-Y\|∥X−Y∥ with the singular values of XXX for every YYY of rank at most rrr, not only for those diagonal in the same bases. The obvious approach, writing YYY in the singular bases of XXX, fails because YYY need not be diagonal there.
  • Uniqueness under the gap requires showing that every minimizer is diagonal in some singular bases of XXX and then using the gap to fix its support. Without the gap uniqueness fails: for X=I2X=I_2X=I2​ and r=1r=1r=1 every uu⊤uu^\topuu⊤ with ∥u∥=1\|u\|=1∥u∥=1 is a nearest point.

A further subtlety: Proposition 3.3 must hold for any singular value decomposition of XXX. Singular vectors are not unique, so the proof has to show that the truncation does not depend on the choice under the gap (3.8).

Formalization scope

Matrices are Matrix (Fin n) (Fin m) ℝ with Mathlib's Frobenius norm (open scoped Matrix.Norms.Frobenius). The projection is the published platform predicate RandomGradFree.Nonsmooth.IsMetricProjection, and PRr(X)P_{\mathcal R_r}(X)PRr​​(X) is the set {Y | IsMetricProjection (rankSet r) X Y}. The set equality with a singleton states existence, uniqueness and the formula together.

Conventions committed to:

  • sv X i is σi(X)\sigma_i(X)σi​(X), 1-based as on the page. It is Mathlib's 0-based singularValues of Matrix.toEuclideanLin X at i - 1, so the index 000 is meaningless and every statement uses indices ≥1\ge1≥1.
  • IsSVD X U S V is the predicate of (3.5). The singular value decomposition is a hypothesis of the theorems, never a choice made inside them, so the theorems hold for every singular value decomposition.
  • truncSVD r U S V is ∑i≤rΣiiuivi⊤\sum_{i\le r}\Sigma_{ii}u_iv_i^\top∑i≤r​Σii​ui​vi⊤​, written with the diagonal of S. That these entries are the singular values σi(X)\sigma_i(X)σi​(X) is a fact to be proved, not part of the definition.
  • The hypothesis r≥1r\ge1r≥1 is added, because σr\sigma_rσr​ needs it (and R0={0}\mathcal R_0=\{0\}R0​={0}). It is the only hypothesis of Proposition 3.3 not printed on the page.
  • In Proposition 4.11, matrices use the block index types Fin r ⊕ Fin p and Fin r ⊕ Fin q. The tangent vector and the normal space are taken in the form the page gives, and invertibility of Σ0+A\Sigma_0+AΣ0​+A is an explicit hypothesis. In the paper's proof it comes from "ZZZ in a neighborhood of the origin".

A formalization in which the radius is replaced by a smaller one, the projection is taken onto the matrices of rank at most rrr, or the singular value decomposition is chosen inside the statement is a different theorem and is ruled out.

Wanted contributions, all reusable beyond this mission:

  • existence of a singular value decomposition in matrix form, and the identification of its diagonal with singularValues;
  • Weyl's inequality for singular values;
  • the Eckart–Young theorem and its uniqueness case;
  • the Schur-complement rank formula for 2×22\times22×2 block matrices, used for Proposition 4.11.

Selected references

  • P.-A. Absil and J. Malick, Projection-like retractions on matrix manifolds, SIAM J. Optim. 22(1):135–158, 2012. https://doi.org/10.1137/100802529 (authors' version HAL hal-00651608v2: https://hal.science/hal-00651608v2)
  • P.-A. Absil, R. Mahony and R. Sepulchre, Optimization Algorithms on Matrix Manifolds, Princeton University Press, 2008. https://press.princeton.edu/absil
  • C. Eckart and G. Young, The approximation of one matrix by another of lower rank, Psychometrika 1(3):211–218, 1936. https://doi.org/10.1007/BF02288367
  • R. A. Horn and C. R. Johnson, Matrix Analysis, Cambridge University Press, 1985 (§7.3, §7.4). https://doi.org/10.1017/CBO9780511810817
  • U. Helmke and J. B. Moore, Optimization and Dynamical Systems, Springer, 1994 (Ch. 5). https://doi.org/10.1007/978-1-4471-3467-1
8 thms1 active userReviewed
Differential GeometryOptimization·Captain: mikedeng1

Projection-like Retractions on Matrix Manifolds III: Near a Locally Symmetric Point, the Projection onto a Spectral Manifold Is U Diag(P_M(λ(X))) UᵀResearch Paper

Motivation

Riemannian optimization algorithms move along a manifold by taking a step in a tangent direction and then coming back to the manifold. The map that does the coming back is a retraction, and the most natural candidate on a submanifold of a Euclidean space is the projective retraction: add the tangent step, then take the nearest point of the manifold. Its practical value depends on whether that nearest point can be computed.

For many manifolds of symmetric matrices the defining property is a property of the eigenvalues: matrices whose largest eigenvalue has multiplicity ppp, matrices with prescribed spectrum, and similar sets that arise in eigenvalue optimization and in alternating projection methods (Lewis and Malick 2008). Daniilidis, Malick and Sendov (reference [8] of the paper) showed that such a spectral set is a smooth manifold when the underlying set of eigenvalue vectors is a smooth, locally symmetric manifold. P.-A. Absil and J. Malick (SIAM J. Optim. 2012; authors' version hal-00651608v2) then showed that, close to such a manifold, the metric projection onto it has a closed form: one eigendecomposition and one projection in Rn\mathbb R^nRn. This mission formalizes that result, Theorem 3.9 of the paper.

Setting

Write Sn\mathbf S_nSn​ for the real symmetric n×nn \times nn×n matrices with the Frobenius norm ∥X∥2=∑i,jXij2\|X\|^2 = \sum_{i,j} X_{ij}^2∥X∥2=∑i,j​Xij2​, On\mathbf O_nOn​ for the orthogonal matrices, and Σn\mathbf \Sigma_nΣn​ for the permutation matrices, which act on Rn\mathbb R^nRn (with the Euclidean norm) by permuting coordinates. Let

R↓n={x∈Rn:x1≥x2≥⋯≥xn}.\mathbb R^n_\downarrow = \{x \in \mathbb R^n : x_1 \ge x_2 \ge \cdots \ge x_n\}.R↓n​={x∈Rn:x1​≥x2​≥⋯≥xn​}.

For X∈SnX \in \mathbf S_nX∈Sn​, λ(X)∈R↓n\lambda(X) \in \mathbb R^n_\downarrowλ(X)∈R↓n​ is the vector of its eigenvalues, with multiplicity, in nonincreasing order; every XXX has an eigendecomposition X=UDiag⁡(λ(X))U⊤X = U \operatorname{Diag}(\lambda(X)) U^\topX=UDiag(λ(X))U⊤ with U∈OnU \in \mathbf O_nU∈On​. For M⊆RnM \subseteq \mathbb R^nM⊆Rn the spectral set of MMM is

λ−1(M)={X∈Sn:λ(X)∈M}.\lambda^{-1}(M) = \{X \in \mathbf S_n : \lambda(X) \in M\}.λ−1(M)={X∈Sn​:λ(X)∈M}.

For a set QQQ and a point yyy, PQ(y)P_Q(y)PQ​(y) is the set of nearest points of QQQ to yyy (it may be empty or contain several points).

Let M\mathcal MM be a C2C^2C2 submanifold of Rn\mathbb R^nRn: around each of its points it is a coordinate slice of a C2C^2C2 chart with C2C^2C2 inverse. Let S=λ−1(M∩R↓n)\mathcal S = \lambda^{-1}(\mathcal M \cap \mathbb R^n_\downarrow)S=λ−1(M∩R↓n​), let Xˉ∈S\bar X \in \mathcal SXˉ∈S and xˉ=λ(Xˉ)\bar x = \lambda(\bar X)xˉ=λ(Xˉ). The set M∩B(xˉ,δ)\mathcal M \cap B(\bar x, \delta)M∩B(xˉ,δ) (open ball) is strongly locally symmetric if for every x∈M∩B(xˉ,δ)x \in \mathcal M \cap B(\bar x, \delta)x∈M∩B(xˉ,δ) and every P∈ΣnP \in \mathbf \Sigma_nP∈Σn​ with Px=xPx = xPx=x,

P(M∩B(xˉ,δ))=M∩B(xˉ,δ).(3.15)P\big(\mathcal M \cap B(\bar x, \delta)\big) = \mathcal M \cap B(\bar x, \delta). \qquad (3.15)P(M∩B(xˉ,δ))=M∩B(xˉ,δ).(3.15)

Formalization targets

Goal: Theorem 3.9 (projection onto spectral manifolds)

Assume M\mathcal MM is a C2C^2C2 submanifold of Rn\mathbb R^nRn, Xˉ∈S\bar X \in \mathcal SXˉ∈S, δ>0\delta > 0δ>0 and (3.15). Then there is δ0∈(0,δ]\delta_0 \in (0, \delta]δ0​∈(0,δ] such that for every X∈SnX \in \mathbf S_nX∈Sn​ with ∥X−Xˉ∥≤δ0/2\|X - \bar X\| \le \delta_0/2∥X−Xˉ∥≤δ0​/2, the set PM(λ(X))P_{\mathcal M}(\lambda(X))PM​(λ(X)) is a single point ppp, and for every U∈OnU \in \mathbf O_nU∈On​ with X=UDiag⁡(λ(X))U⊤X = U \operatorname{Diag}(\lambda(X)) U^\topX=UDiag(λ(X))U⊤,

PS(X)={ UDiag⁡(p) U⊤ }.P_{\mathcal S}(X) = \{\, U \operatorname{Diag}(p)\, U^\top \,\}.PS​(X)={UDiag(p)U⊤}.

The goal fixes no constant: δ0\delta_0δ0​ is only asserted to exist, as the paper's proof restricts δ\deltaδ to make the local uniqueness of PMP_{\mathcal M}PM​ and Lemma 3.8 apply.

Milestones

  1. (3.11): ∥λ(X)−λ(Y)∥≤∥X−Y∥\|\lambda(X) - \lambda(Y)\| \le \|X - Y\|∥λ(X)−λ(Y)∥≤∥X−Y∥ for X,Y∈SnX, Y \in \mathbf S_nX,Y∈Sn​.
  2. Lemma 3.7: for closed M⊆R↓nM \subseteq \mathbb R^n_\downarrowM⊆R↓n​, an eigendecomposition X=UDiag⁡(λ(X))U⊤X = U\operatorname{Diag}(\lambda(X))U^\topX=UDiag(λ(X))U⊤ and sorted zzz, UDiag⁡(z)U⊤∈Pλ−1(M)(X)  ⟺  z∈PM(λ(X))U \operatorname{Diag}(z) U^\top \in P_{\lambda^{-1}(M)}(X) \iff z \in P_M(\lambda(X))UDiag(z)U⊤∈Pλ−1(M)​(X)⟺z∈PM​(λ(X)).
  3. Lemma 3.8: for xˉ∈R↓n\bar x \in \mathbb R^n_\downarrowxˉ∈R↓n​ and all small δ>0\delta > 0δ>0, for y∈B(xˉ,δ)y \in B(\bar x, \delta)y∈B(xˉ,δ) and sorted x∈B(xˉ,δ)x \in B(\bar x, \delta)x∈B(xˉ,δ), the maximum of x⊤Pyx^\top P yx⊤Py over permutations fixing xˉ\bar xxˉ is attained at some PPP with PyPyPy sorted.
  4. (3.20): under (3.15) and for small δ\deltaδ, the distance from a sorted x∈B(xˉ,δ)x \in B(\bar x, \delta)x∈B(xˉ,δ) to M∩B(xˉ,δ)\mathcal M \cap B(\bar x, \delta)M∩B(xˉ,δ) is attained up to equality on sorted points.

Significance

The theorem turns the projective retraction on a spectral manifold, an optimization problem over n×nn \times nn×n matrices, into a projection in Rn\mathbb R^nRn onto M\mathcal MM plus one eigendecomposition. For the matrices whose largest eigenvalue has multiplicity ppp, the projection onto Mp\mathcal M_pMp​ is an explicit averaging of the top ppp eigenvalues (Example 3.10 of the paper), which completes a partial result of Oustry (reference [29, Th. 13] of the paper). Together with Theorem 3.5, it is what makes the projective retraction a practical alternative on manifolds where the Riemannian exponential has no known efficient formula.

The result is proved in the paper; to our knowledge none of it is formalized. The mission produces a machine-checked version of the theorem and of the spectral-set toolkit beneath it: the Lipschitz property of sorted eigenvalues, the reduction of projections onto spectral sets to projections onto sets of vectors, and the permutation rearrangement lemma. The formalization also corrects a misstatement: Lemma 3.7 as printed quantifies over all z∈Rnz \in \mathbb R^nz∈Rn, and its direction "⇒\Rightarrow⇒" fails for unsorted zzz (for n=2n = 2n=2, M={(2,0)}M = \{(2, 0)\}M={(2,0)}, X=U=IX = U = IX=U=I, z=(0,2)z = (0, 2)z=(0,2)). The mission states it for z∈R↓nz \in \mathbb R^n_\downarrowz∈R↓n​, which is the only case Theorem 3.9 uses.

Difficulty

The obvious argument shows only half of the goal. Lemma 3.7 characterizes the nearest points of S\mathcal SS that share the eigenvectors UUU of XXX; it does not exclude a nearest point with other eigenvectors, and the goal asserts that PS(X)P_{\mathcal S}(X)PS​(X) is a single point. Excluding the others needs the equality case of the trace inequality trace⁡(XY)≤λ(X)⊤λ(Y)\operatorname{trace}(XY) \le \lambda(X)^\top \lambda(Y)trace(XY)≤λ(X)⊤λ(Y) (3.10), which the paper quotes from the literature without proof, and the fact that the projection ppp inherits the ties of λ(X)\lambda(X)λ(X), which comes from strong local symmetry and uniqueness of PMP_{\mathcal M}PM​.

The local uniqueness of PMP_{\mathcal M}PM​ near a point of a C2C^2C2 submanifold (Lemma 3.1 of the paper) is itself a tubular-neighbourhood argument through the inverse function theorem. Mathlib has no metric projection onto embedded submanifolds, no von Neumann trace inequality for sorted eigenvalues, and no Hoffman–Wielandt inequality. Finally, the radii interact: the restriction δ0\delta_0δ0​ must make Lemma 3.1, Lemma 3.8 and the ball inclusion PM(x)∈B(xˉ,δ)P_{\mathcal M}(x) \in B(\bar x, \delta)PM​(x)∈B(xˉ,δ) all hold at once.

Formalization scope

  • Vectors are EuclideanSpace ℝ (Fin n), with the Euclidean norm (not the sup norm of Fin n → ℝ). Matrices are Matrix (Fin n) (Fin n) ℝ with Mathlib's Frobenius norm (open scoped Matrix.Norms.Frobenius); Sn\mathbf S_nSn​ is IsHermitian (symmetry over R\mathbb RR) and On\mathbf O_nOn​ is Matrix.orthogonalGroup.
  • λ(X)\lambda(X)λ(X) is eig X, built from Mathlib's sorted IsHermitian.eigenvalues₀; on a non-symmetric matrix it returns 000, and every statement assumes symmetry. R↓n\mathbb R^n_\downarrowR↓n​ is sortedDesc n, spectral sets are specSet, permutations act by permAct (an isometry), and (3.15) is IsStronglyLocallySymmetric on the open ball.
  • Nearest points are the platform predicate RandomGradFree.Nonsmooth.IsMetricProjection; PQ(y)P_Q(y)PQ​(y) is {z | IsMetricProjection Q y z}.
  • The submanifold hypothesis is a local slice-chart predicate IsSubmanifold 2 d M. The paper allows k=2k = 2k=2 or ∞\infty∞; k=2k = 2k=2 covers both. Only the hypotheses of Theorem 3.5 (cited from Daniilidis–Malick–Sendov without proof) are assumed, not its conclusion that S\mathcal SS is a manifold.
  • "For δ\deltaδ small enough" is ∃δ1>0,∀δ∈(0,δ1]\exists \delta_1 > 0, \forall \delta \in (0, \delta_1]∃δ1​>0,∀δ∈(0,δ1​]; the radius of Theorem 3.9 is an existential δ0≤δ\delta_0 \le \deltaδ0​≤δ, with the paper's non-strict ∥X−Xˉ∥≤δ0/2\|X - \bar X\| \le \delta_0/2∥X−Xˉ∥≤δ0​/2.
  • Both singleton claims of Theorem 3.9 are part of the conclusion. A version proving only ⊆\subseteq⊆ would be satisfied by the empty set, and a version that assumes PM(λ(X))P_{\mathcal M}(\lambda(X))PM​(λ(X)) is a singleton would delete the theorem's local-uniqueness content; neither is the target.

Infrastructure that a full development needs and that is reusable beyond this mission: the von Neumann/Fan trace inequality and its equality case for real symmetric matrices, the Hoffman–Wielandt inequality (3.11), the rearrangement inequality over permutations fixing a vector, and the local existence and uniqueness of metric projections onto C2C^2C2 submanifolds. Useful platform items: RHLinalg.vonNeumann_trace_ineq (the trace inequality for Hermitian matrices), RHLinalg.bilinear_doublyStochastic_le_of_monovary (the rearrangement step), and Bhatia.trace_mul_perm_bounds (trace pairings between permutation pairings of unsorted spectra). Contributions of these lemmas, of alternative proofs of (3.11), and of the equality case of (3.10) are welcome.

Selected references

  • P.-A. Absil and J. Malick, Projection-like retractions on matrix manifolds, SIAM J. Optim. 22(1):135–158, 2012. https://doi.org/10.1137/100802529; authors' version https://hal.science/hal-00651608v2
  • A. Daniilidis, J. Malick and H. Sendov, Locally symmetric submanifolds lift up to spectral manifolds, preprint, 2009 (reference [8] of the paper; the source of Theorem 3.5).
  • A. S. Lewis and J. Malick, Alternating projections on manifolds, Math. Oper. Res. 33(1):216–234, 2008. https://doi.org/10.1287/moor.1070.0291
  • F. Oustry, A second-order bundle method to minimize the maximum eigenvalue function, Math. Program. 89:1–34, 2000 (reference [29] of the paper).
  • A. S. Lewis, Convex analysis on the Hermitian matrices, SIAM J. Optim. 6(1):164–177, 1996. https://doi.org/10.1137/0806009
8 thms1 active userReviewed
Control TheoryDynamic ProgrammingMarkov Chain+2·Captain: mikedeng1

Bellman's Dynamic Programming IX: Markovian Decision Processes and the Maximal Perron RootTextbook

Motivation

Chapter XI of Richard Bellman's Dynamic Programming (Princeton University Press, 1957; DOI 10.2307/j.ctv1nxcw0f) studies decision processes whose state is a vector of nonnegative quantities, for example the probabilities that a system is in each of NNN states, or the stocks of NNN commodities, and whose transitions are linear maps chosen stage by stage by a controller. Maximizing a linear functional of the state at every stage leads to the nonlinear difference equation

xi(n+1)=max⁡q∑j=1Naij(q) xj(n),xi(0)=ci,x_i(n+1) = \max_q \sum_{j=1}^N a_{ij}(q)\, x_j(n), \qquad x_i(0) = c_i,xi​(n+1)=qmax​j=1∑N​aij​(q)xj​(n),xi​(0)=ci​,

and, in the limit of small time steps, to differential equations of the form dx/dt=max⁡q[A(q,t)x+b(q,t)]dx/dt = \max_q [A(q,t)x + b(q,t)]dx/dt=maxq​[A(q,t)x+b(q,t)] and, when two opposing controllers act, dx/dt=max⁡pmin⁡q[… ]dx/dt = \max_p \min_q[\dots]dx/dt=maxp​minq​[…].

These equations are the multiplicative counterpart of the additive Bellman equation. Their growth rate is the natural object for controlled population models, controlled Markov chains observed through their unnormalized state vectors, and economic growth models with a choice of technology. Bellman announced the discrete results in "A Markovian decision process" (J. Math. Mech. 6, 1957) the same year as the book, and R. A. Howard's Dynamic Programming and Markov Processes (MIT Press, 1960) developed policy iteration for the related average-reward problem. The central discrete result of the chapter, Theorem 2, is an early instance of what is now called nonlinear Perron–Frobenius theory (Lemmens and Nussbaum, 2012).

Setting

Fix N≥1N \ge 1N≥1. Row iii of the matrix carries its own control qiq_iqi​, ranging over a set SiS_iSi​; the joint control is q=(q1,…,qN)q = (q_1, \dots, q_N)q=(q1​,…,qN​) in S=S1×⋯×SNS = S_1 \times \dots \times S_NS=S1​×⋯×SN​, and A(q)=(aij(qi))A(q) = (a_{ij}(q_i))A(q)=(aij​(qi​)). Bellman insists on this row-wise structure (§ 3): "the set of q's for each row is distinct from the corresponding set for any other row ... so that there is no interaction between the various maximizations". The maximum of a vector over qqq is then taken row by row.

The Perron root φ(q)\varphi(q)φ(q) is the characteristic root of A(q)A(q)A(q) of largest absolute value, the spectral radius of A(q)A(q)A(q) as a complex matrix. The conditions (10.3) of the chapter are:

  1. for every yyy and every row the maximum of ∑jaij(qi)yj\sum_j a_{ij}(q_i) y_j∑j​aij​(qi​)yj​ over SiS_iSi​ is attained;
  2. 0<aij(q)≤m<∞0 < a_{ij}(q) \le m < \infty0<aij​(q)≤m<∞ on SSS;
  3. φ\varphiφ attains its maximum on SSS.

For the continuous processes, ∥x∥=∑i∣xi∣\|x\| = \sum_i |x_i|∥x∥=∑i​∣xi​∣ and ∥A∥=∑i,j∣aij∣\|A\| = \sum_{i,j}|a_{ij}|∥A∥=∑i,j​∣aij​∣, and a solution of dx/dt=F(t,x)dx/dt = F(t,x)dx/dt=F(t,x), x(0)=cx(0)=cx(0)=c, on [0,T][0,T][0,T] is a continuous xxx with x(t)=c+∫0tF(s,x(s)) dsx(t) = c + \int_0^t F(s, x(s))\,dsx(t)=c+∫0t​F(s,x(s))ds, which is the book's "satisfying the equation almost everywhere". The successive approximations are x0=cx_0 = cx0​=c, xn+1(t)=c+∫0tF(s,xn(s)) dsx_{n+1}(t) = c + \int_0^t F(s, x_n(s))\,dsxn+1​(t)=c+∫0t​F(s,xn​(s))ds.

Formalization targets

Goal: Chapter XI, Theorem 2

Under (10.3) there is exactly one λ>0\lambda > 0λ>0 for which

λyi=max⁡q∑j=1Naij(q) yj,i=1,…,N,\lambda y_i = \max_q \sum_{j=1}^N a_{ij}(q)\, y_j, \qquad i = 1,\dots,N,λyi​=qmax​j=1∑N​aij​(q)yj​,i=1,…,N,

has a solution with all yi>0y_i > 0yi​>0. That solution is unique up to a positive factor, and

λ=max⁡q∈Sφ(q).\lambda = \max_{q \in S} \varphi(q).λ=q∈Smax​φ(q).

Milestones

  1. § 4, Lemma. For row-wise maximized operators T1(x)=max⁡q[b1(q,t)+∫0tA(q,s)x ds]T_1(x) = \max_q[b_1(q,t) + \int_0^t A(q,s)x\,ds]T1​(x)=maxq​[b1​(q,t)+∫0t​A(q,s)xds] and T2(y)T_2(y)T2​(y) likewise, ∥T1(x)−T2(y)∥≤max⁡q[∥b1−b2∥+∫0t∥A(q,s)∥ ∥x−y∥ ds]\|T_1(x) - T_2(y)\| \le \max_q[\|b_1 - b_2\| + \int_0^t \|A(q,s)\|\,\|x-y\|\,ds]∥T1​(x)−T2​(y)∥≤maxq​[∥b1​−b2​∥+∫0t​∥A(q,s)∥∥x−y∥ds].
  2. Theorem 1. If ∥A(q,t)∥,∥b(q,t)∥≤f(t)\|A(q,t)\|, \|b(q,t)\| \le f(t)∥A(q,t)∥,∥b(q,t)∥≤f(t) with fff locally integrable and the maximum is attained, then dx/dt=max⁡q[A(q,t)x+b(q,t)]dx/dt = \max_q[A(q,t)x + b(q,t)]dx/dt=maxq​[A(q,t)x+b(q,t)], x(0)=cx(0) = cx(0)=c, has a unique solution, the uniform limit of the successive approximations.
  3. Theorem 3 (corrected). If moreover φ\varphiφ has a unique maximizer on SSS and c≥0c \ge 0c≥0, c≠0c \ne 0c=0, then the recurrence satisfies xi(n)∼a yi λnx_i(n) \sim a\,y_i\,\lambda^nxi​(n)∼ayi​λn with a=a(c)>0a = a(c) > 0a=a(c)>0.
  4. Theorem 4. The same well-posedness for dx/dt=max⁡pmin⁡q[A(p,q,t)x+b(p,q,t)]=min⁡qmax⁡p[… ]dx/dt = \max_p\min_q[A(p,q,t)x + b(p,q,t)] = \min_q\max_p[\dots]dx/dt=maxp​minq​[A(p,q,t)x+b(p,q,t)]=minq​maxp​[…] on [0,T][0,T][0,T].
  5. Theorem 5. If (Bp,q)≥d>0(Bp,q) \ge d > 0(Bp,q)≥d>0 on probability vectors, the solution of du/dt=max⁡pmin⁡q[(Ap,q)−(Bp,q)u]du/dt = \max_p\min_q[(Ap,q) - (Bp,q)u]du/dt=maxp​minq​[(Ap,q)−(Bp,q)u] satisfies
lim⁡t→∞u(t)=max⁡pmin⁡q(Ap,q)(Bp,q)=min⁡qmax⁡p(Ap,q)(Bp,q).\lim_{t\to\infty} u(t) = \max_p \min_q \frac{(Ap,q)}{(Bp,q)} = \min_q \max_p \frac{(Ap,q)}{(Bp,q)} .t→∞lim​u(t)=pmax​qmin​(Bp,q)(Ap,q)​=qmin​pmax​(Bp,q)(Ap,q)​.

Significance

Theorem 2 identifies the optimal long-run growth rate of a controlled multiplicative process with the largest Perron root among the admissible matrices, and shows that the optimal process has a single positive stationary direction. Theorem 3 turns this into the asymptotics of the value iteration x(n+1)=max⁡qA(q)x(n)x(n+1) = \max_q A(q)x(n)x(n+1)=maxq​A(q)x(n): after normalization by λn\lambda^nλn the iterates converge to a multiple of the eigenvector. Theorems 1 and 4 are the existence and uniqueness results that justify defining continuous-time controlled processes and differential games by these equations. Theorem 5 recovers the min-max theorem for ratios of bilinear forms (Chapter X) as the long-run limit of a scalar differential game.

The results are classical, and none of them is formalized. Mathlib has the spectral radius and irreducible matrices but no Perron–Frobenius theorem and no Brouwer fixed point theorem; the platform has a statement of the Perron theorem for a single positive matrix (ClassicalGaps.perron_positive_matrix). A formal proof of the goal therefore also produces a reusable monotone, positively homogeneous eigenvector theorem on the positive orthant.

Difficulty

The map y↦max⁡qA(q)yy \mapsto \max_q A(q)yy↦maxq​A(q)y is not linear, so the linear-algebra proof of the Perron theorem through the characteristic polynomial does not apply. Existence of a positive eigenvector needs a fixed point argument for a nonlinear map of the simplex (Bellman uses Brouwer's theorem). The identification λ=max⁡qφ(q)\lambda = \max_q \varphi(q)λ=maxq​φ(q) must connect the nonlinear eigenvalue with the spectra of the individual matrices, which requires the Perron theory of each A(q)A(q)A(q), including the fact that the Perron root dominates every complex eigenvalue in modulus. For Theorem 3, the iterates may switch controls infinitely often when SSS is infinite, so an argument that the optimal control is eventually constant does not settle convergence. For Theorems 1 and 4, the right-hand side is only measurable in ttt and Lipschitz in xxx with an integrable constant, so the classical Picard–Lindelöf theorem with a continuous right-hand side does not apply directly.

Formalization scope

Everything lives in the namespace BellmanDP.Markovian. Vectors are Fin N → ℝ and matrices are Matrix (Fin N) (Fin N) ℝ. Row iii's control type is Q i with admissible set S i, and the joint admissible set is Set.pi Set.univ S. The Perron root is (spectralRadius ℂ (A.map (algebraMap ℝ ℂ))).toReal, the largest modulus of a complex eigenvalue; it is not defined as a positive eigenvalue with a positive eigenvector, which would make the Perron–Frobenius content of the goal definitional. The maximized eigen-equation is stated with IsGreatest, so the maxima are attained. The goal and Theorem 3 assume N≥1N \ge 1N≥1; for N=0N = 0N=0 every λ\lambdaλ would qualify.

Conventions and repairs:

  • Theorem 3 prints "a unique q for which the maximum value of q is assumed". A control has no maximum value; the proof uses "q∗q^*q∗ ... the value of qqq for which λ=φ(q∗)\lambda = \varphi(q^*)λ=φ(q∗)", so the hypothesis is uniqueness of the maximizer of φ\varphiφ. For c=0c = 0c=0 the iterates vanish and xi(n)∼ayiλnx_i(n) \sim a y_i\lambda^nxi​(n)∼ayi​λn fails, so c≠0c \ne 0c=0 is assumed (the proof takes c>0c > 0c>0 "without loss of generality"). The asymptotic is stated as xi(n)/λn→ayix_i(n)/\lambda^n \to a y_ixi​(n)/λn→ayi​ with a>0a > 0a>0.
  • Theorems 1 and 4: the book's controls are functions of ttt with the maximum outside the integral; since the maximization is pointwise (§ 4), the statements use pointwise sets and the integral of the pointwise maximum. Measurability of t↦F(t,x)t \mapsto F(t,x)t↦F(t,x) is not stated in the book and is assumed. In Theorem 4 the max-min is taken row by row, and (2a) is encoded as the existence of a saddle point in each row.
  • § 4 Lemma: "≤max⁡q[… ]\le \max_q[\dots]≤maxq​[…]" is stated as "≤[… ]\le [\dots]≤[…] at some admissible joint qqq".
  • Theorem 5: the right-hand side is the max-min form; the equality of the two ratio values is part of the conclusion.

Degenerate readings are ruled out: the maxima are attained or taken over nonempty compact sets, never Lean's junk sSup of an unbounded set, and the Perron root is spectral rather than defined through the conclusion. Contributions welcome: a proof of the single-matrix Perron theorem in the form needed here, a Brouwer or Kakutani fixed point theorem for the simplex, and a Carathéodory existence theorem for dx/dt=F(t,x)dx/dt = F(t,x)dx/dt=F(t,x) with an integrable Lipschitz constant, each reusable well beyond this mission.

Selected references

  • R. Bellman, Dynamic Programming, Princeton University Press, 1957; Princeton Landmarks in Mathematics ed., 2010, Chapter XI. https://doi.org/10.2307/j.ctv1nxcw0f
  • R. Bellman, "A Markovian decision process", Journal of Mathematics and Mechanics 6 (1957), 679–684.
  • R. A. Howard, Dynamic Programming and Markov Processes, MIT Press, 1960.
  • O. Perron, "Zur Theorie der Matrices", Mathematische Annalen 64 (1907), 248–263. https://doi.org/10.1007/BF01449896
  • B. Lemmens and R. Nussbaum, Nonlinear Perron–Frobenius Theory, Cambridge University Press, 2012. https://doi.org/10.1017/CBO9781139026079
9 thms1 active userReviewed
🏆Completed
Machine LearningProbability·Captain: Minghui

Fine-Tuning Can Distort Pretrained Features: Perfect-Feature LP-FT SeparationResearch Paper

Why initialization matters for transfer learning

Transfer learning starts with a representation learned on an earlier task and adapts it to a new one. Two common choices are linear probing, which changes only the final linear predictor, and fine-tuning, which changes the representation as well. These procedures optimize related training objectives, but their behavior away from the training data can differ. Kumar and coauthors study this distinction through two-layer linear networks, alongside experiments with nonlinear networks. This mission formalizes their perfect-feature LP-FT result, rather than the empirical claims or the general imperfect-feature comparison. See Section 3.4, Proposition 3.7, PDF p. 10.

LP-FT first learns a head by linear probing and then uses that head to initialize full fine-tuning. The perfect-feature setting isolates the effect of head initialization: the representation already contains exactly the features needed to predict the labels, but the head initially need not use them correctly. The mathematical question is whether joint training preserves or loses the representation's ability to predict outside the observed training subspace.

Linear predictors, training data, and OOD loss

An input is a vector x∈Rdx\in\mathbb R^dx∈Rd. A feature extractor is a matrix B∈Rk×dB\in\mathbb R^{k\times d}B∈Rk×d, and a head is a vector v∈Rkv\in\mathbb R^kv∈Rk. Together they predict v⊤Bxv^\top Bxv⊤Bx, with effective weight vector B⊤vB^\top vB⊤v. The fixed matrix X∈Rn×dX\in\mathbb R^{n\times d}X∈Rn×d contains the nnn training inputs as rows. Their span is S=rowspace⁡(X)S=\operatorname{rowspace}(X)S=rowspace(X), with dimension mmm.

The ground truth has orthonormal-row features B⋆B_\starB⋆​ and a nonzero head v⋆v_\starv⋆​. Write w⋆=B⋆⊤v⋆w_\star=B_\star^\top v_\starw⋆​=B⋆⊤​v⋆​ and Y=Xw⋆Y=Xw_\starY=Xw⋆​. Perfect pretrained features mean B0=UB⋆B_0=UB_\starB0​=UB⋆​ for an orthogonal matrix UUU. The corresponding aligned head is u=Uv⋆u=Uv_\staru=Uv⋆​. The dimensions satisfy 1≤k≤m1\le k\le m1≤k≤m and m+k<dm+k<dm+k<d.

The two geometric assumptions require the orthogonal projections from R0=rowspace⁡(B0)R_0=\operatorname{rowspace}(B_0)R0​=rowspace(B0​) into SSS and into S⊥S^\perpS⊥ to be injective. In this dimension regime these are exactly the positive largest-principal-angle cosine conditions used by the paper. They demand more than two subspaces having some nonorthogonal directions. The Lean definition spells out injectivity of v↦ΠTB0⊤vv\mapsto\Pi_T B_0^\top vv↦ΠT​B0⊤​v for each T∈{S,S⊥}T\in\{S,S^\perp\}T∈{S,S⊥}. See Definition 3.2 and Appendix A.1, PDF pp. 7 and 22-23.

An out-of-distribution law μ\muμ is any probability measure on Rd\mathbb R^dRd with a finite second moment and positive-definite uncentered second-moment matrix Σ=Eμ[xx⊤]\Sigma=\mathbb E_\mu[xx^\top]Σ=Eμ​[xx⊤]. Its mean need not be zero. Define

LOOD(v,B)=Ex∼μ[(v⊤Bx−w⋆⊤x)2].L_{\rm OOD}(v,B)=\mathbb E_{x\sim\mu} [(v^\top Bx-w_\star^\top x)^2].LOOD​(v,B)=Ex∼μ​[(v⊤Bx−w⋆⊤​x)2].

Both training methods use the unnormalized loss L^(v,B)=∥XB⊤v−Y∥22\widehat L(v,B)=\|XB^\top v-Y\|_2^2L(v,B)=∥XB⊤v−Y∥22​. Fine-tuning follows its gradient flow in both parameters; linear probing keeps B=B0B=B_0B=B0​. Time is real and nonnegative. These are the paper's equations (3.2)-(3.3), PDF p. 6.

Formalization targets

The goal is Proposition 3.7 in an explicit nonzero-signal regime. For σ>0\sigma>0σ>0, initialize an FT head with independent Gaussian coordinates, v0∼N(0,σ2Ik)v_0\sim\mathcal N(0,\sigma^2I_k)v0​∼N(0,σ2Ik​). Establish

Pr⁡ ⁣[∀t≥0,LOOD(vFT(t),BFT(t))>0]=1.\Pr\!\left[\forall t\ge0,\quad L_{\rm OOD}(v_{\rm FT}(t),B_{\rm FT}(t))>0\right]=1.Pr[∀t≥0,LOOD​(vFT​(t),BFT​(t))>0]=1.

Linear probing, from any initial head, must converge to uuu. Fine-tuning initialized at its limit must satisfy

vLP(t)⟶u,∀t≥0,LOOD(vLP-FT(t),BLP-FT(t))=0.v_{\rm LP}(t)\longrightarrow u,\qquad \forall t\ge0,\quad L_{\rm OOD}(v_{\rm LP\text{-}FT}(t),B_{\rm LP\text{-}FT}(t))=0.vLP​(t)⟶u,∀t≥0,LOOD​(vLP-FT​(t),BLP-FT​(t))=0.

The probability-one event applies to all times simultaneously. The goal also asserts existence of the relevant global flows; a conditional claim about a possibly nonexistent trajectory would not suffice. The statement does not assert a numerical error lower bound or a positive time-infimum.

Seven milestones supply the supporting results: global flow existence and FT uniqueness; unchanged features orthogonal to the training span; the balancedness invariant; the second-moment identity for OOD risk; almost-sure Gaussian head misalignment; exact LP recovery; and stationarity after LP initialization. The principal source is Appendices A.2 and A.7, PDF pp. 23-31 and 45-47.

What the result establishes

The result distinguishes two initializations of the same joint-training procedure. In this idealized setting, a head obtained by linear probing gives zero OOD loss throughout subsequent fine-tuning, while a Gaussian head almost surely has positive OOD loss at every finite time. The conclusion concerns population squared prediction error, not classification accuracy or a finite test-set estimate.

The paper establishes the mathematical claim; this mission asks for a Lean proof of the stated model and result. The scope is deliberately limited to perfect pretrained features. It does not claim an LP-FT upper bound for imperfect features, which the authors identify as a further challenge in Section 3.4, PDF p. 10. A completed development would also provide reusable components for finite dimensional gradient flows, factorized linear models, and population risk.

Why the proof needs the training dynamics

The training loss alone does not select a unique effective predictor in an overparameterized problem. Knowing that a predictor fits the observed examples therefore does not determine its OOD loss. Formalization must track the head and feature extractor together, and it must distinguish parameter stationarity from a claim that a derivative happens to vanish at one time. The Gaussian conclusion also requires one event controlling an uncountable set of times; separate probability-one statements for individual times would be weaker.

Formalization scope and conventions

Vectors use Mathlib's finite dimensional real Euclidean spaces. Matrices are represented as continuous linear maps, with Euclidean adjoints and operator norms. The feature update is written explicitly as the Frobenius-gradient equation; it is not a gradient with respect to the operator norm. Differentiability is imposed within [0,∞)[0,\infty)[0,∞), including the right derivative at zero.

The dimensions, nonzero target, positive Gaussian scale, finite second moments, and projection injectivity are explicit. The nonzero target restricts the formalization to the regime of the Gaussian alignment argument in Lemma A.12; k≤mk\le mk≤m makes the identifiability condition used in Proposition A.20 precise. The random-head law is the scaled standard Gaussian measure. No randomness of the fixed training matrix or independence from an additional data draw is assumed.

The model contains no assumed convergence, invariant, or desired risk bound. Each of those is a theorem obligation. The well-posedness milestone makes explicit an analytic prerequisite of the source's flow notation. The risk milestone uses the identity in (A.29)-(A.32), avoiding the reversed inequality printed in (A.28). The quantitative constant in Theorem 3.3 is outside this mission. Source-aligned proofs and the supporting analysis infrastructure are welcome; changing the learning rule or assuming a milestone inside the model would change the task.

Selected references

  • Ananya Kumar, Aditi Raghunathan, Robbie Jones, Tengyu Ma, and Percy Liang, Fine-Tuning can Distort Pretrained Features and Underperform Out-of-Distribution, ICLR 2022, arXiv:2202.10054v1. Main target: Section 3.4, Proposition 3.7, PDF p. 10, equations (3.10)-(3.11); proof: Appendix A.7, PDF pp. 45-47, Proposition A.20 and (A.208)-(A.218). Supporting invariants: Appendix A.2, PDF p. 24, Lemmas A.3-A.4, equations (A.15)-(A.20). Gaussian alignment: Appendix A.3, PDF pp. 34-35, Lemmas A.11-A.12.
9 thms1 active userReviewed
🏆Completed
Mathematical Physics·Captain: ShapeZero

The spectrum of the octonionic two-generator flowTextbook

Motivation

This mission is pure mathematics about objects it defines itself: an octonion multiplication built from the Fano plane of the companion mission The role postulates force exactly seven points, and the matrix of the linear map

p  ↦  p a+b pp \;\mapsto\; p\,a + b\,pp↦pa+bp

for two unit imaginary octonions a,ba, ba,b. It is motivated by the two-generator D8 flow ψ˙=ψa+bψ\dot\psi = \psi a + b\psiψ˙​=ψa+bψ in the Shape Zero derivation (00_START_HERE/MODEL_SPEC.md §1b; 02_synthesis/D8_SYNTHESIS.md). The statements below do not depend on that motivation.

Result. For a=e1a = e_1a=e1​ and b=c e1+s e2b = c\,e_1 + s\,e_2b=ce1​+se2​ with c2+s2=1c^2 + s^2 = 1c2+s2=1, the characteristic polynomial of M=Ra+LbM = R_a + L_bM=Ra​+Lb​ is

X2 (X2+4) (X2+(2−2c))2.X^2\,(X^2 + 4)\,\bigl(X^2 + (2 - 2c)\bigr)^2 .X2(X2+4)(X2+(2−2c))2.

With c=cos⁡θc = \cos\thetac=cosθ, 2−2c=4sin⁡2(θ/2)2 - 2c = 4\sin^2(\theta/2)2−2c=4sin2(θ/2), so the eigenvalues are 0,0,±2i0, 0, \pm 2i0,0,±2i and ±2isin⁡(θ/2)\pm 2i\sin(\theta/2)±2isin(θ/2) (each twice), and the ratio of the two nonzero frequencies is 1/sin⁡(θ/2)1/\sin(\theta/2)1/sin(θ/2).

What this mission does NOT prove.

  • Only the canonical pair. It proves the result for a=e1a = e_1a=e1​, b=cos⁡θ e1+sin⁡θ e2b = \cos\theta\, e_1 + \sin\theta\, e_2b=cosθe1​+sinθe2​. That every pair of unit imaginary octonions at angle θ\thetaθ gives the same spectrum follows from the classical fact that G2G_2G2​ acts transitively on such pairs — not formalized here.
  • Nothing about any model. It says nothing about which angle, if any, a physical model selects, or whether the D8 flow plays a role in one.
  • One orientation. The octonion table uses one orientation of the Fano lines; all valid orientations give isomorphic algebras, and the spectrum is basis-independent, but only this table is formalized.

Setting

Write e0=1,e1,…,e7e_0 = 1, e_1, \dots, e_7e0​=1,e1​,…,e7​ for the standard basis of R8\mathbb{R}^8R8. The imaginary unit eue_ueu​ (u=1,…,7u = 1, \dots, 7u=1,…,7) is labelled by the Fano point u−1u - 1u−1. For each line {l,l+1,l+3}\{l, l+1, l+3\}{l,l+1,l+3} (mod 777) of the companion mission's Fano plane RolesForceSeven.fanoLine, the product is oriented cyclically along the ordered triple (l,l+1,l+3)(l, l+1, l+3)(l,l+1,l+3):

el+1 el+2=el+4,el+2 el+4=el+1,el+4 el+1=el+2e_{l+1}\,e_{l+2} = e_{l+4}, \qquad e_{l+2}\,e_{l+4} = e_{l+1}, \qquad e_{l+4}\,e_{l+1} = e_{l+2}el+1​el+2​=el+4​,el+2​el+4​=el+1​,el+4​el+1​=el+2​

(unit indices mod 777, shifted into 1,…,71, \dots, 71,…,7), with the reversed products negative, eu2=−1e_u^2 = -1eu2​=−1, and e0e_0e0​ the identity. This defines OctonionD8.octTable and the bilinear product OctonionD8.omul on R8\mathbb{R}^8R8.

For a,b∈R8a, b \in \mathbb{R}^8a,b∈R8, Rmat a is the matrix of p↦p ap \mapsto p\,ap↦pa and Lmat b the matrix of p↦b pp \mapsto b\,pp↦bp (column jjj is the image of eje_jej​). The flow matrix is

M=flowMat  c  s=Re1+Lce1+se2.M = \texttt{flowMat}\;c\;s = R_{e_1} + L_{c e_1 + s e_2}.M=flowMatcs=Re1​​+Lce1​+se2​​.

The Fano plane is imported from the companion mission, not restated. The reordering blockEquiv and the two blocks blockA c s, blockB c s are defined explicitly for the milestones.

Formalization targets

Goal: the characteristic polynomial

c2+s2=1  ⟹  χM(X)=X2 (X2+4) (X2+(2−2c))2.c^2 + s^2 = 1 \;\Longrightarrow\; \chi_M(X) = X^2\,(X^2 + 4)\,\bigl(X^2 + (2 - 2c)\bigr)^2 .c2+s2=1⟹χM​(X)=X2(X2+4)(X2+(2−2c))2.

This is OctonionD8.flow_charpoly. The hypothesis c2+s2=1c^2 + s^2 = 1c2+s2=1 is the only one, and it is needed: without it the characteristic polynomial differs.

Milestones — the proof outline

The proof goes through an invariant-subspace decomposition. Reorder the basis as (e0,e1,e2,e4∣e3,e5,e6,e7)(e_0, e_1, e_2, e_4 \mid e_3, e_5, e_6, e_7)(e0​,e1​,e2​,e4​∣e3​,e5​,e6​,e7​) (OctonionD8.blockEquiv).

  1. M1 (block-diagonal form). For all real c,sc, sc,s, in the reordered basis MMM is block diagonal: M=(A00B)M = \begin{pmatrix} A & 0 \\ 0 & B \end{pmatrix}M=(A0​0B​), with explicit 4×44\times44×4 blocks AAA on (e0,e1,e2,e4)(e_0, e_1, e_2, e_4)(e0​,e1​,e2​,e4​) and BBB on (e3,e5,e6,e7)(e_3, e_5, e_6, e_7)(e3​,e5​,e6​,e7​) (OctonionD8.blockA, OctonionD8.blockB).
  2. M2 (block AAA). If c2+s2=1c^2 + s^2 = 1c2+s2=1, then χA(X)=X2 (X2+4)\chi_A(X) = X^2\,(X^2 + 4)χA​(X)=X2(X2+4).
  3. M3 (block BBB). If c2+s2=1c^2 + s^2 = 1c2+s2=1, then χB(X)=(X2+(2−2c))2\chi_B(X) = \bigl(X^2 + (2 - 2c)\bigr)^2χB​(X)=(X2+(2−2c))2.

The goal follows: the characteristic polynomial is unchanged by reordering the basis, and that of a block-diagonal matrix is the product of the blocks' characteristic polynomials.

Further results (after the goal)

  • It is the octonions: the norm is multiplicative. For all p,q∈R8p, q \in \mathbb{R}^8p,q∈R8, ∑k(pq)k2=(∑ipi2)(∑jqj2)\sum_k (pq)_k^2 = \bigl(\sum_i p_i^2\bigr)\bigl(\sum_j q_j^2\bigr)∑k​(pq)k2​=(∑i​pi2​)(∑j​qj2​) — the eight-square identity, which certifies that the table defines a normed (composition) algebra, the octonions, rather than some other algebra.
  • The flow conserves the norm. MT=−MM^{\mathsf T} = -MMT=−M.
  • An annihilating polynomial. If c2+s2=1c^2 + s^2 = 1c2+s2=1, then M (M2+4) (M2+(2−2c))=0M\,(M^2 + 4)\,\bigl(M^2 + (2 - 2c)\bigr) = 0M(M2+4)(M2+(2−2c))=0.

Corollary: the frequencies

For 0<θ<π0 < \theta < \pi0<θ<π, with c=cos⁡θc = \cos\thetac=cosθ and s=sin⁡θs = \sin\thetas=sinθ, the roots over C\mathbb{C}C of the characteristic polynomial, with multiplicity, are exactly

0, 0, ±2i, ±2isin⁡(θ/2), ±2isin⁡(θ/2),0,\ 0,\ \pm 2i,\ \pm 2i\sin(\theta/2),\ \pm 2i\sin(\theta/2),0, 0, ±2i, ±2isin(θ/2), ±2isin(θ/2),

and 0<sin⁡(θ/2)<10 < \sin(\theta/2) < 10<sin(θ/2)<1. So the two nonzero frequencies 222 and 2sin⁡(θ/2)2\sin(\theta/2)2sin(θ/2) are distinct, and their ratio is 1/sin⁡(θ/2)1/\sin(\theta/2)1/sin(θ/2). Uses 2−2cos⁡θ=4sin⁡2(θ/2)2 - 2\cos\theta = 4\sin^2(\theta/2)2−2cosθ=4sin2(θ/2).

Significance

The result itself. It gives the spectrum of the two-generator flow for the canonical pair in closed form, for every angle, from an octonion table built on the formalized Fano plane.

Formalizing it. The spectrum had been checked numerically and symbolically only.

Numerical and symbolic cross-check (independent code)

checkresult
table from the companion mission's Fano lines is a normed algebra (200 random pairs)true
eigenvalue magnitudes {0×2, 2sin⁡(θ/2)×4, 2×2}\{0 \times 2,\ 2\sin(\theta/2) \times 4,\ 2 \times 2\}{0×2, 2sin(θ/2)×4, 2×2} at 5°, 30°, 60°, 76.3°, 90°, 120°, 150°max error 1.8×10−151.8\times10^{-15}1.8×10−15
symbolic characteristic polynomial (sympy, s2=1−c2s^2 = 1 - c^2s2=1−c2)X2(X2+4)(X2+2−2c)2X^2(X^2 + 4)(X^2 + 2 - 2c)^2X2(X2+4)(X2+2−2c)2

These agree with the repository's shape_zero_tests/d8_closed_form.py (frequencies {0,2sin⁡(θ/2),2}\{0, 2\sin(\theta/2), 2\}{0,2sin(θ/2),2} to 5×10−115\times10^{-11}5×10−11), which uses a different octonion table.

Difficulty

Moderate. The goal needs the characteristic polynomial of an 8×88\times88×8 matrix with symbolic entries; a direct determinant expansion is expensive. The milestones take the invariant-subspace route: the block-diagonal form (M1) is an entrywise computation from the octonion table, and each 4×44\times44×4 block's characteristic polynomial (M2, M3) is a small determinant, reduced with c2+s2=1c^2 + s^2 = 1c2+s2=1. The further results are large but mechanical polynomial identities (the eight-square identity; the annihilating polynomial), where the risk is performance, not ideas.

Formalization scope

  • Vectors are Fin 8 → ℝ; matrices are Matrix (Fin 8) (Fin 8) ℝ, with column jjj the image of the basis vector eje_jej​.
  • The characteristic polynomial is Mathlib's Matrix.charpoly over R\mathbb{R}R; the corollary maps it to C\mathbb{C}C and uses Polynomial.roots, a multiset, so multiplicities are part of the statement.
  • The table is one fixed orientation of the Fano lines; G2G_2G2​-invariance and other orientations are out of scope.

Selected references

  • Wikipedia, Octonion. https://en.wikipedia.org/wiki/Octonion
  • Wikipedia, Fano plane. https://en.wikipedia.org/wiki/Fano_plane
  • Shape Zero repository (motivation only). https://github.com/ShapeZeroSZ/shape-zero
11 thms1 active userReviewed
🏆Completed
Mathematical Physics·Captain: ShapeZero

Zero net power forces every link coupling to be symmetric, in any dimensionTextbook

Motivation

The Shape Zero model (Shape Zero LLC, unpublished) couples neighbouring nodes through a velocity-dependent force, one coupling matrix per link. The companion mission Zero net power forces every link coupling to be symmetric (rings of ≥ 3 sites) proves that a coupling doing no net work (a passive coupling) must have every link matrix symmetric — but only on a one-dimensional ring. The model's base is three-dimensional. This mission extends the result to a periodic cubic lattice with any number q of spatial axes, so it covers q = 3 directly and reduces exactly to the ring result at q = 1.

What this mission does NOT prove. As with the companion missions, commuting with the complex structure JJJ is not derived here; it remains a modelling premise. This mission shows only that passivity forces symmetry on every link in every direction.

Setting

Fix natural numbers qqq (number of axes), LLL (sites along each axis) and ddd. A site is a point x∈(Z/LZ)qx \in (\mathbb{Z}/L\mathbb{Z})^qx∈(Z/LZ)q of the periodic cubic lattice; x±eax \pm e_ax±ea​ is the site one step forward or back along axis aaa, wrapping around. Each link — a site xxx together with an axis aaa — carries its own real d×dd\times dd×d matrix W(x,a)W(x,a)W(x,a), and each site carries a velocity v(x)∈Rdv(x) \in \mathbb{R}^dv(x)∈Rd. The force on site xxx is

F(x)=∑a(W(x,a) v(x+ea)−W(x−ea,a) v(x−ea)),F(x) = \sum_{a} \Bigl( W(x,a)\, v(x+e_a) - W(x-e_a,a)\, v(x-e_a) \Bigr),F(x)=a∑​(W(x,a)v(x+ea​)−W(x−ea​,a)v(x−ea​)),

where the second term uses the matrix of the link behind xxx. The total power is

PW(v)=∑xv(x)⋅F(x).P_W(v) = \sum_{x} v(x) \cdot F(x).PW​(v)=x∑​v(x)⋅F(x).

In Lean, sites are Site q L = Fin q → Fin L, the step is shift x a s (update coordinate aaa by sss in Fin L), and the power is PassivityTorus.power q L d W v. The coupling is passive when PW(v)=0P_W(v) = 0PW​(v)=0 for every vvv.

Formalization targets

Goal: passive if and only if every link is symmetric (L≥3L \ge 3L≥3)

L≥3  ⟹  (∀v,  PW(v)=0)  ⟺  (∀x,a,  W(x,a)T=W(x,a)).L \ge 3 \;\Longrightarrow\; \Bigl(\forall v,\; P_W(v) = 0\Bigr) \iff \Bigl(\forall x, a,\; W(x,a)^{\mathsf T} = W(x,a)\Bigr).L≥3⟹(∀v,PW​(v)=0)⟺(∀x,a,W(x,a)T=W(x,a)).

This is PassivityTorus.passive_iff_symm, for every qqq and ddd.

Milestones

  1. M1. Stepping back then forward along an axis returns to the start.
  2. M2 (power identity, every qqq, LLL). PW(v)=∑x∑av(x)⋅((W(x,a)−W(x,a)T) v(x+ea))P_W(v) = \sum_x \sum_a v(x) \cdot \bigl((W(x,a) - W(x,a)^{\mathsf T})\, v(x+e_a)\bigr)PW​(v)=∑x​∑a​v(x)⋅((W(x,a)−W(x,a)T)v(x+ea​)).
  3. M3 (symmetric links are passive, every LLL).
  4. M4 (passive forces symmetric, L≥3L \ge 3L≥3).

Significance

The result itself. It is the dimension-independent form of the passivity theorem: whatever the number of spatial axes, a per-link coupling does no net work for every motion exactly when every link matrix is symmetric. It removes the one-dimensional restriction from the step that justifies restricting the model's coupling class to symmetric matrices.

Formalizing it. The ring result has been machine-checked; the lattice extension had been checked numerically only (identity error ≤3.6×10−14\le 3.6\times10^{-14}≤3.6×10−14 at (q,L)=(1,3),(2,3),(3,3),(3,4)(q,L) = (1,3), (2,3), (3,3), (3,4)(q,L)=(1,3),(2,3),(3,3),(3,4)). A proof covers every qqq, L≥3L \ge 3L≥3 and ddd.

Difficulty

The algebra is the ring argument; the difficulty is bookkeeping. The incoming term must be reindexed by the one-step shift, which is a bijection of the torus, and the converse needs a test motion supported on two neighbouring sites xxx and x+eax+e_ax+ea​ together with a proof that no other link joins them.

The hypothesis L≥3L \ge 3L≥3 is necessary. At L=2L = 2L=2 the step forward and the step back along an axis reach the same site. This is proved (in Lean, locally): for every q≥1q \ge 1q≥1, putting the same non-symmetric matrix (0100)\begin{pmatrix} 0 & 1 \\ 0 & 0 \end{pmatrix}(00​10​) on every link of the L=2L = 2L=2 lattice gives exactly zero power for every velocity field, so passivity does not force symmetry there. Weakening the hypothesis makes the goal false.

Formalization scope

  • The lattice is periodic in every axis, with the same LLL along each; [NeZero L] is needed for the literal 111 in Fin L and is implied by L≥3L \ge 3L≥3 in the goal.
  • There is one matrix per site per axis, with no relation assumed between links.
  • Passivity is quantified over all velocity fields; everything is real.
  • q=0q = 0q=0 (a single site, no links) and d=0d = 0d=0 are included; both sides of the goal hold trivially there, and this is not a trivialization since the statement is quantified over all qqq and ddd.

Selected references

  • Wikipedia, Symmetric matrix. https://en.wikipedia.org/wiki/Symmetric_matrix
  • Wikipedia, Passivity (engineering). https://en.wikipedia.org/wiki/Passivity_(engineering)
  • Mathlib, Mathlib.Data.Matrix.Mul (dot product ⬝ᵥ, matrix–vector product *ᵥ). https://leanprover-community.github.io/mathlib4_docs/Mathlib/Data/Matrix/Mul.html
6 thms1 active userReviewed
🏆Completed
Mathematical Physics·Captain: ShapeZero

Zero net power forces every link coupling to be symmetric (rings of ≥ 3 sites)Textbook

Motivation

The Shape Zero model (Shape Zero LLC, unpublished) couples neighbouring nodes on a ring through a velocity-dependent force, with one coupling matrix per link. A companion mission, The passivity-admissible couplings have dimension n² = dim u(n), shows that real matrices which are symmetric and commute with a complex structure JJJ form a space of dimension n2=dim⁡u(n)n^2 = \dim \mathfrak{u}(n)n2=dimu(n). That mission deliberately takes symmetry as given. This mission supplies the reason for it: a coupling that never does net work (a passive coupling) must have every link matrix symmetric, provided the ring has at least three sites.

Together the two missions give the chain

passive  ⟹  every Wi symmetric(this mission),\text{passive} \;\Longrightarrow\; \text{every } W_i \text{ symmetric} \qquad\text{(this mission)},passive⟹every Wi​ symmetric(this mission), symmetric+commutes with J  ⟹  dim⁡=n2(companion mission).\text{symmetric} + \text{commutes with } J \;\Longrightarrow\; \dim = n^2 \qquad\text{(companion mission)}.symmetric+commutes with J⟹dim=n2(companion mission).

What these two missions do NOT prove. Commuting with JJJ is not derived here. It is the requirement that the coupling respect each node's complex structure, and in the model it is imposed, not forced by passivity. After both missions the proved content is: passivity forces symmetry, and symmetric JJJ-compatible couplings have the dimension of u(n)\mathfrak{u}(n)u(n). The JJJ-compatibility step remains a modelling premise. In addition, the coupling space consists of the Hermitian matrices, which match u(n)\mathfrak{u}(n)u(n) in dimension; the Lie algebra u(n)\mathfrak{u}(n)u(n) itself appears only after multiplying by iii (Unitary group).

Setting

Let NNN and ddd be natural numbers with N≥1N \ge 1N≥1. A ring has NNN sites labelled by Z/NZ={0,…,N−1}\mathbb{Z}/N\mathbb{Z} = \{0, \dots, N-1\}Z/NZ={0,…,N−1}, so site N−1N-1N−1 is next to site 000; indices i+1i + 1i+1 and i−1i - 1i−1 are taken modulo NNN. Each site carries a velocity vector vi∈Rdv_i \in \mathbb{R}^dvi​∈Rd, and each link i→i+1i \to i+1i→i+1 carries a real d×dd \times dd×d link matrix WiW_iWi​. The force on site iii is

Fi=Wi vi+1−Wi−1 vi−1,F_i = W_i\, v_{i+1} - W_{i-1}\, v_{i-1},Fi​=Wi​vi+1​−Wi−1​vi−1​,

where the second term uses the previous link's matrix Wi−1W_{i-1}Wi−1​. The total power delivered by the coupling is

PW(v)=∑i∈Z/NZvi⋅(Wi vi+1−Wi−1 vi−1).P_W(v) = \sum_{i \in \mathbb{Z}/N\mathbb{Z}} v_i \cdot \bigl( W_i\, v_{i+1} - W_{i-1}\, v_{i-1} \bigr).PW​(v)=i∈Z/NZ∑​vi​⋅(Wi​vi+1​−Wi−1​vi−1​).

In the Lean development this is PassivityRing.power N d W v, with sites indexed by Fin N (whose addition and subtraction wrap around modulo NNN), vectors in Fin d → ℝ, and ⋅\cdot⋅ the dot product ⬝ᵥ. The coupling is passive when PW(v)=0P_W(v) = 0PW​(v)=0 for every choice of velocities vvv.

Formalization targets

Goal: passive if and only if every link is symmetric (rings of at least 3 sites)

N≥3  ⟹  (∀v,  PW(v)=0)  ⟺  (∀i,  WiT=Wi).N \ge 3 \;\Longrightarrow\; \Bigl( \forall v,\; P_W(v) = 0 \Bigr) \iff \Bigl( \forall i,\; W_i^{\mathsf T} = W_i \Bigr).N≥3⟹(∀v,PW​(v)=0)⟺(∀i,WiT​=Wi​).

This is PassivityRing.passive_iff_symm. It holds for every ddd and every family of link matrices, with no assumption relating different links.

Milestones

  1. M1 (power identity, every NNN). PW(v)=∑ivi⋅((Wi−WiT) vi+1)P_W(v) = \sum_i v_i \cdot \bigl((W_i - W_i^{\mathsf T})\, v_{i+1}\bigr)PW​(v)=∑i​vi​⋅((Wi​−WiT​)vi+1​).
  2. M2 (symmetric links are passive, every NNN). If every WiW_iWi​ is symmetric, then PW(v)=0P_W(v) = 0PW​(v)=0 for all vvv.
  3. M3 (passive forces symmetric, N≥3N \ge 3N≥3). If PW(v)=0P_W(v) = 0PW​(v)=0 for all vvv, then every WiW_iWi​ is symmetric.

Significance

The result itself. The goal turns a physical requirement — no net work for any motion — into an exact algebraic condition on each link separately. It is the step that justifies restricting the companion mission's coupling class to symmetric matrices, so that the dimension count n2n^2n2 applies to the passive couplings of the model rather than to an assumed class.

Formalizing it. The model's claim has so far been checked numerically at ring sizes 1,2,3,4,71, 2, 3, 4, 71,2,3,4,7. A machine-checked proof covers every N≥3N \ge 3N≥3 and every ddd, makes the role of the hypothesis N≥3N \ge 3N≥3 explicit, and separates what is proved (passivity forces symmetry) from what remains a modelling premise (JJJ-compatibility).

Difficulty

The forward direction and the power identity are finite-sum bookkeeping. The central difficulty is the converse: the hypothesis constrains a single scalar quantity summed around the ring, and it must be shown to constrain each link matrix individually. On small rings this fails because distinct terms of the sum coincide, so the argument depends on the ring being large enough that neighbouring sites are genuinely distinct, and in Lean that is modular arithmetic on Fin N.

The hypothesis N≥3N \ge 3N≥3 is necessary, not a convenience:

  • N=1N = 1N=1: the site is its own neighbour on both sides, so the force is W0v0−W0v0=0W_0 v_0 - W_0 v_0 = 0W0​v0​−W0​v0​=0 and the power vanishes for every W0W_0W0​, symmetric or not.
  • N=2N = 2N=2: the power reduces to v0⋅(D−DT) v1v_0 \cdot (D - D^{\mathsf T})\, v_1v0​⋅(D−DT)v1​ with D=W0−W1D = W_0 - W_1D=W0​−W1​, so it depends only on the difference of the two link matrices. Taking both links equal to the same non-symmetric matrix, for example (0100)\begin{pmatrix} 0 & 1 \\ 0 & 0 \end{pmatrix}(00​10​), gives zero power for all velocities.

Weakening the hypothesis to 2≤N2 \le N2≤N, or dropping it, makes the goal false.

Formalization scope

  • Sites are Fin N with its wrap-around arithmetic, so i+1i + 1i+1 and i−1i - 1i−1 go around the ring; [NeZero N] is required for the literal 111 in Fin N and is implied by N≥3N \ge 3N≥3 in the goal.
  • There is one matrix per link, W : Fin N → Matrix (Fin d) (Fin d) ℝ, not a single shared matrix, and no relation between different links is assumed.
  • Passivity is quantified over all velocity configurations v : Fin N → (Fin d → ℝ).
  • Everything is real (ℝ); symmetry is (W i)ᵀ = W i.
  • d=0d = 0d=0 is included; there both sides of the goal hold trivially. This is not a trivialization, since the statement is quantified over all ddd.

M1 and M2 hold for every N≥1N \ge 1N≥1 and carry no ring-size hypothesis; M3 and the goal require N≥3N \ge 3N≥3. The development needs only Mathlib's finite sums, dot products, matrix–vector products and Fin arithmetic.

Selected references

  • B. C. Hall, Lie Groups, Lie Algebras, and Representations: An Elementary Introduction, 2nd ed., Graduate Texts in Mathematics 222, Springer, 2015. https://doi.org/10.1007/978-3-319-13467-3
  • Wikipedia, Unitary group. https://en.wikipedia.org/wiki/Unitary_group
  • Wikipedia, Symmetric matrix. https://en.wikipedia.org/wiki/Symmetric_matrix
  • Mathlib, Mathlib.Data.Matrix.Mul (dot product ⬝ᵥ, matrix–vector product *ᵥ). https://leanprover-community.github.io/mathlib4_docs/Mathlib/Data/Matrix/Mul.html
5 thms1 active userReviewed
PreviousPage 2 of 2Next

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me