Prove2Me
Navigate
DiscoverFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

Optimization

633 missions · 406 completed

Missions

Open227Completed406All633
🏆Completed
CombinatoricsLinear OptimizationOperations Research·Captain: mikedeng1

Proximity Results and Faster Algorithms for Integer Programming Using the Steinitz Lemma: ℓ1-Proximity of Integer and LP OptimaResearch Paper

Motivation

Integer programs are routinely solved by first solving their linear programming (LP) relaxation and then searching for an integer optimum near the fractional one. How near an integer optimum must be is the subject of proximity theorems. They bound the search region of branch-and-bound and of dynamic programming, and they turn a fractional optimum into a starting point for exact algorithms.

The classical bound is due to Cook, Gerards, Schrijver and Tardos (Math. Programming 34, 1986): for an integer program in inequality form max⁡{cTx:Ax≤b, x∈Zn}\max\{c^Tx : Ax\le b,\ x\in\mathbb Z^n\}max{cTx:Ax≤b, x∈Zn} that is feasible and bounded, every optimal LP solution x∗x^*x∗ has an optimal integer solution z∗z^*z∗ with ∥x∗−z∗∥∞≤n⋅δ\|x^*-z^*\|_\infty\le n\cdot\delta∥x∗−z∗∥∞​≤n⋅δ, where δ\deltaδ is the largest absolute value of a subdeterminant of AAA. For programs in standard form Ax=bAx=bAx=b with mmm rows this gives, via the Hadamard bound, ∥z∗−x∗∥1≤n2⋅mm/2Δm\|z^*-x^*\|_1\le n^2\cdot m^{m/2}\Delta^m∥z∗−x∗∥1​≤n2⋅mm/2Δm, which grows with the number of variables nnn.

Eisenbrand and Weismantel (ACM Trans. Algorithms 16(1), Article 5, 2019; conference version SODA 2018) removed the dependence on nnn altogether, using the Steinitz lemma on rearranging vectors so that all partial sums stay short. Their bound depends only on mmm and on the largest absolute value Δ\DeltaΔ of an entry of AAA, and it is the basis of their faster algorithms for integer programs with few constraints.

Setting

Fix natural numbers mmm (rows) and nnn (variables). The data are a matrix A∈Zm×nA\in\mathbb Z^{m\times n}A∈Zm×n, a right-hand side b∈Zmb\in\mathbb Z^mb∈Zm, an objective c∈Znc\in\mathbb Z^nc∈Zn and upper bounds u∈Nnu\in\mathbb N^nu∈Nn. A natural number Δ\DeltaΔ bounds the entries: ∣aij∣≤Δ|a_{ij}|\le\Delta∣aij​∣≤Δ for all i,ji,ji,j. The integer program (10) is

max⁡{cTx:Ax=b, 0≤x≤u, x∈Zn},\max\{c^Tx : Ax=b,\ 0\le x\le u,\ x\in\mathbb Z^n\},max{cTx:Ax=b, 0≤x≤u, x∈Zn},

and its LP relaxation is the same problem over x∈Rnx\in\mathbb R^nx∈Rn. Its feasible region P={x∈Rn:Ax=b, 0≤x≤u}P=\{x\in\mathbb R^n: Ax=b,\ 0\le x\le u\}P={x∈Rn:Ax=b, 0≤x≤u} is a polytope, lpPolytope A b u. An optimal vertex solution is an optimal solution of the LP relaxation (IsLPOptimal) that is an extreme point of PPP. An optimal integer solution is IsIPOptimal. Both are maxima.

Distances are measured in the ℓ1\ell_1ℓ1​-norm ∥z−x∥1=∑i∣zi−xi∣\|z-x\|_1=\sum_i|z_i-x_i|∥z−x∥1​=∑i​∣zi​−xi​∣.

A vector y∈Zny\in\mathbb Z^ny∈Zn is a cycle of z∗−x∗z^*-x^*z∗−x∗ (Eq. (14)) if Ay=0Ay=0Ay=0 and, for every iii, ∣yi∣≤∣(z∗−x∗)i∣|y_i|\le|(z^*-x^*)_i|∣yi​∣≤∣(z∗−x∗)i​∣ and yi(z∗−x∗)i≥0y_i(z^*-x^*)_i\ge0yi​(z∗−x∗)i​≥0: an integer kernel vector that is sign-compatible with z∗−x∗z^*-x^*z∗−x∗ and dominated by it (IsCycle).

The Steinitz lemma (Theorem 1.1) concerns vectors x1,…,xnx_1,\dots,x_nx1​,…,xn​ in an mmm-dimensional normed space with ∑ixi=0\sum_i x_i=0∑i​xi​=0 and ∥xi∥≤1\|x_i\|\le1∥xi​∥≤1. It asserts a permutation π\piπ with ∥∑j≤kxπ(j)∥≤c(m)\|\sum_{j\le k}x_{\pi(j)}\|\le c(m)∥∑j≤k​xπ(j)​∥≤c(m) for all kkk, and the paper uses Sevast'anov's constant c(m)=mc(m)=mc(m)=m.

Formalization targets

Goal: Theorem 3.3 (p. 5:8)

If (10) has an integer feasible point and x∗x^*x∗ is an optimal vertex solution of its LP relaxation, then there is an optimal solution z∗z^*z∗ of (10) with

∥z∗−x∗∥1 ≤ m⋅(2mΔ+1)m.\|z^*-x^*\|_1\ \le\ m\cdot(2m\Delta+1)^m .∥z∗−x∗∥1​ ≤ m⋅(2mΔ+1)m.

The constant is the paper's. The goal holds for all mmm, nnn, bbb, ccc and uuu; only mmm and Δ\DeltaΔ enter the bound.

Milestones, in the order the proof uses them

  1. Lemma 3.1 (p. 5:8): for an LP optimum x∗x^*x∗, an integer optimum z∗z^*z∗ and a cycle yyy of z∗−x∗z^*-x^*z∗−x∗, the vector z∗−yz^*-yz∗−y is integer feasible, x∗+yx^*+yx∗+y is LP feasible, and cTy≤0c^Ty\le0cTy≤0.
  2. Lemma 3.2 (p. 5:8): if z∗z^*z∗ minimizes ∥z∗−x∗∥1\|z^*-x^*\|_1∥z∗−x∗∥1​ among the optimal integer solutions, then z∗−x∗z^*-x^*z∗−x∗ has no nonzero cycle.
  3. Theorem 1.1 with c(m)=mc(m)=mc(m)=m (p. 5:4): the Steinitz lemma in any mmm-dimensional real normed space.
  4. Proof of Theorem 3.3 (pp. 5:8–5:9): round a vertex x∗x^*x∗ towards an integer vector and write {x∗}\{x^*\}{x∗} for the remainder. Then ∥−A{x∗}∥∞≤Δm\|-A\{x^*\}\|_\infty\le\Delta m∥−A{x∗}∥∞​≤Δm and −A{x∗}=w1+⋯+wm-A\{x^*\}=w_1+\dots+w_m−A{x∗}=w1​+⋯+wm​ with integer wjw_jwj​, ∥wj∥∞≤Δ\|w_j\|_\infty\le\Delta∥wj​∥∞​≤Δ.
  5. Proof of Theorem 3.3, Eq. (20) (p. 5:9): a sequence of integer vectors of ℓ∞\ell_\inftyℓ∞​-norm at most mΔm\DeltamΔ in which no value repeats m+1m+1m+1 times has length at most m(2mΔ+1)mm(2m\Delta+1)^mm(2mΔ+1)m.
  6. Eq. (21) (p. 5:9), a consequence: cT(x∗−z∗)≤∥c∥∞⋅m(2mΔ+1)mc^T(x^*-z^*)\le\|c\|_\infty\cdot m(2m\Delta+1)^mcT(x∗−z∗)≤∥c∥∞​⋅m(2mΔ+1)m for every optimal integer solution z∗z^*z∗.

Significance

The bound is independent of the number of variables. Combined with the paper's dynamic program, it gives the paper's running-time results for integer programs with upper bounds: an optimal LP vertex is computed, and the integer optimum is searched for within an ℓ1\ell_1ℓ1​-ball of radius m(2mΔ+1)mm(2m\Delta+1)^mm(2mΔ+1)m around it. Eq. (21) bounds the absolute integrality gap by the same quantity, scaled by ∥c∥∞\|c\|_\infty∥c∥∞​. The Steinitz lemma with constant mmm is a general tool in discrepancy theory and in scheduling algorithms.

All of these results have published proofs. No machine-checked proof of Theorem 3.3 or of the Steinitz lemma is known to this mission, and Mathlib has no Steinitz lemma. The mission asks for complete Lean proofs of the milestones and of the goal. A proof of the Steinitz lemma with constant mmm for arbitrary norms is reusable well beyond integer programming.

Difficulty

Lemmas 3.1 and 3.2 and the counting step are elementary. The substance lies in two places. The first is the Steinitz lemma with the linear constant mmm for an arbitrary norm: the bound must hold uniformly in the number nnn of vectors, and the constant must be exactly mmm, because the goal's constant (2mΔ+1)m(2m\Delta+1)^m(2mΔ+1)m counts integer points of ℓ∞\ell_\inftyℓ∞​-norm at most mΔm\DeltamΔ. The second is the passage from a vertex to at most mmm fractional coordinates. The paper argues this in one sentence ("x∗x^*x∗ has at most mmm positive entries"), which is not literally true for (10) with upper bounds: coordinates at their upper bound ui>0u_i>0ui​>0 are positive. The correct fact concerns coordinates strictly between 000 and uiu_iui​, and it has to be derived from the extreme-point property of PPP.

Formalization scope

  • All declarations live in the namespace IPProximity.Eisenbrand. The data are integral: A : Matrix (Fin m) (Fin n) ℤ, b : Fin m → ℤ, c : Fin n → ℤ, u : Fin n → ℕ (entries ui=0u_i=0ui​=0 allowed), Δ : ℕ. They are cast to ℝ once, inside the LP definitions. m=0m=0m=0 and n=0n=0n=0 are allowed.
  • "Vertex" is Mathlib's Set.extremePoints ℝ (lpPolytope A b u). It is not defined through bases or by counting fractional coordinates.
  • The ℓ1\ell_1ℓ1​-distance is the explicit sum ∑ i, |(z i : ℝ) - x i|. Mathlib's norm on Fin n → ℝ is the sup norm, and it is used only where the paper has ∥⋅∥∞\|\cdot\|_\infty∥⋅∥∞​ (the ∥c∥∞\|c\|_\infty∥c∥∞​ of Eq. (21)).
  • The goal adds one hypothesis the paper leaves implicit: (10) has an integer feasible point. The paper's proof begins with "Let z∗z^*z∗ be an optimal integer solution"; without this hypothesis the conclusion is false.
  • Eq. (14) is formalized literally, so y=0y=0y=0 is a cycle, and Lemma 3.2 is stated for nonzero cycles, which is what its proof establishes. Dropping the vertex hypothesis would make the goal false, so the goal keeps it. The constant is exactly m(2mΔ+1)mm(2m\Delta+1)^mm(2mΔ+1)m, with no hidden existential constant.
  • The Steinitz milestone is stated for any finite-dimensional real normed space of dimension mmm with the explicit constant mmm. The goal needs only the ℓ∞\ell_\inftyℓ∞​ case on Rm\mathbb R^mRm.
  • Out of scope: the dynamic program and the running-time theorems of Sections 2 and 4, and the refinement ∥z∗−x∗∥1≤2Δ\|z^*-x^*\|_1\le2\Delta∥z∗−x∗∥1​≤2Δ for m=1m=1m=1.

Contributions welcome: proofs of any milestone, in particular the Steinitz lemma, and a proof of the goal from the milestones.

Selected references

  • F. Eisenbrand, R. Weismantel, Proximity Results and Faster Algorithms for Integer Programming Using the Steinitz Lemma, ACM Transactions on Algorithms 16(1), Article 5, 2019. https://doi.org/10.1145/3340322
  • W. Cook, A. M. H. Gerards, A. Schrijver, É. Tardos, Sensitivity theorems in integer linear programming, Mathematical Programming 34, 251–264, 1986. https://doi.org/10.1007/BF01582230
  • E. Steinitz, Bedingt konvergente Reihen und konvexe Systeme, Journal für die reine und angewandte Mathematik 143, 128–176, 1913. https://doi.org/10.1515/crll.1913.143.128
  • S. Sevast'janov, Approximate solution of some problems of scheduling theory (in Russian), Metody Diskretnogo Analiza 32, 66–75, 1978 (reference [31] of the paper).
  • V. S. Grinberg, S. V. Sevast'yanov, Value of the Steinitz constant, Functional Analysis and Its Applications 14(2), 125–126, 1980 (reference [16] of the paper).
12 thms3 active usersReviewed
🏆Completed
CombinatoricsGraph TheoryOperations Research·Captain: mikedeng1

Theoretical Improvements in Algorithmic Efficiency for Network Flow Problems 2: The Augmentation Bound for Maximum-Augmentation PathsResearch Paper

Motivation

The maximum flow problem asks how much of a commodity can be sent from a source to a sink through a network whose arcs have capacities. It underlies bipartite matching, transportation, scheduling and many reductions in combinatorial optimization. The classical method for it, the labeling method of Ford and Fulkerson (Flows in Networks, 1962), repeatedly finds an augmenting path and pushes flow along it. With integer capacities it terminates, but the number of augmentations can be as large as the maximum flow value itself, and Edmonds and Karp exhibit a four-node network on which this happens (p. 250). With irrational capacities the method need not terminate at all.

Edmonds and Karp, Theoretical Improvements in Algorithmic Efficiency for Network Flow Problems, J. ACM 19(2):248–264, 1972 (doi:10.1145/321694.321699), showed that two simple rules for choosing the augmenting path repair this. The first, augmenting along a path with fewest arcs, is the subject of mission 1 of this series. This mission covers the second (§1.3): augment along a path that gives the largest possible augmentation. For integer capacities the number of augmentations then grows only logarithmically in the maximum flow value.

Setting

A network NNN has a finite set VVV of nodes, a source sss and a sink t≠st \neq st=s, and a set of arcs, ordered pairs (u,v)(u,v)(u,v) with u≠vu \neq vu=v, at most one from each node to another. One arc is the return arc (t,s)(t,s)(t,s); the other arcs form the set AAA, and each (u,v)∈A(u,v) \in A(u,v)∈A has a capacity c(u,v)>0c(u,v) > 0c(u,v)>0. A flow is a nonnegative function fff on the arcs of NNN with f(u,v)≤c(u,v)f(u,v) \le c(u,v)f(u,v)≤c(u,v) on AAA and flow conservation at every node, sss and ttt included. Its value is f(t,s)f(t,s)f(t,s), the flow returned along the return arc; a maximum flow has the largest value among all flows, and f∗(t,s)f^*(t,s)f∗(t,s) denotes that value.

The residual network NfN^fNf has an arc (u,v)(u,v)(u,v) whenever (u,v)∈A(u,v) \in A(u,v)∈A and c(u,v)−f(u,v)>0c(u,v) - f(u,v) > 0c(u,v)−f(u,v)>0, or (v,u)∈A(v,u) \in A(v,u)∈A and f(v,u)>0f(v,u) > 0f(v,u)>0. An augmenting path is a directed path s=u1,…,up=ts = u_1, \dots, u_p = ts=u1​,…,up​=t of distinct nodes in NfN^fNf. Each of its arcs (u,v)(u,v)(u,v) has a residual amount e(u,v)e(u,v)e(u,v), equal to c(u,v)−f(u,v)c(u,v) - f(u,v)c(u,v)−f(u,v), f(v,u)f(v,u)f(v,u), or c(u,v)−f(u,v)+f(v,u)c(u,v) - f(u,v) + f(v,u)c(u,v)−f(u,v)+f(v,u) according to which of (u,v)(u,v)(u,v), (v,u)(v,u)(v,u) lie in AAA, and the path's augmentation is ε=min⁡e(ui,ui+1)\varepsilon = \min e(u_i, u_{i+1})ε=mine(ui​,ui+1​). Augmenting increases f(t,s)f(t,s)f(t,s) by ε\varepsilonε and changes the flow on the arcs of the path accordingly, with the paper's own rule when both (u,v)(u,v)(u,v) and (v,u)(v,u)(v,u) are arcs. The labeling method produces flows f0,f1,…f^0, f^1, \dotsf0,f1,… by augmenting along a path relative to fkf^kfk as long as one exists.

The rule studied here chooses, at every step, an augmenting path whose ε\varepsilonε is at least that of every other augmenting path relative to the current flow. The bound involves an integer M>1M > 1M>1 such that every partition of the nodes into X∋sX \ni sX∋s and Xˉ∋t\bar X \ni tXˉ∋t has at most MMM arcs of NNN with one end on each side.

Formalization targets

Goal: Theorem 2 (p. 253)

For a network with integer capacities, MMM as above, and a run f0,…,fKf^0, \dots, f^Kf0,…,fK of the labeling method with maximum augmentations started from an integer-valued flow,

K  ≤  1+log⁡M/(M−1)f∗(t,s),K \;\le\; 1 + \log_{M/(M-1)} f^*(t,s),K≤1+logM/(M−1)​f∗(t,s),

and if no augmenting path relative to fKf^KfK exists, then fKf^KfK is a maximum flow.

Milestones

The milestone list follows the paper's argument:

  1. augmentation produces a flow of value f(t,s)+εf(t,s) + \varepsilonf(t,s)+ε (§1.1, p. 249);
  2. a flow is maximum if and only if it has no augmenting path (§1.1, pp. 249–250);
  3. with integer capacities, ε\varepsilonε is a positive integer and the flows of the method stay integer-valued (§1.1, p. 250);
  4. the cut inequality c(X,Xˉ)≥f(X,Xˉ)−f(Xˉ,X)=f(t,s)c(X,\bar X) \ge f(X,\bar X) - f(\bar X,X) = f(t,s)c(X,Xˉ)≥f(X,Xˉ)−f(Xˉ,X)=f(t,s) (p. 254);
  5. f∗(t,s)−fk(t,s)≤εkMf^*(t,s) - f^k(t,s) \le \varepsilon^k Mf∗(t,s)−fk(t,s)≤εkM, where εk=fk+1(t,s)−fk(t,s)\varepsilon^k = f^{k+1}(t,s) - f^k(t,s)εk=fk+1(t,s)−fk(t,s) (p. 254);
  6. f∗(t,s)−fk+1(t,s)≤[f∗(t,s)−fk(t,s)](1−M−1)f^*(t,s) - f^{k+1}(t,s) \le [f^*(t,s) - f^k(t,s)](1 - M^{-1})f∗(t,s)−fk+1(t,s)≤[f∗(t,s)−fk(t,s)](1−M−1) (p. 254);
  7. f∗(t,s)−fk(t,s)≤f∗(t,s)(1−M−1)kf^*(t,s) - f^k(t,s) \le f^*(t,s)(1 - M^{-1})^kf∗(t,s)−fk(t,s)≤f∗(t,s)(1−M−1)k (p. 254).

Significance

Theorem 2 was among the first bounds showing that a maximum flow algorithm can be made polynomial in the size of the numbers rather than in their values: since M≤n2/2M \le n^2/2M≤n2/2 and f∗(t,s)f^*(t,s)f∗(t,s) is at most n2n^2n2 times the average capacity, the bound is O(n2log⁡(n2cˉ))O(n^2 \log(n^2 \bar c))O(n2log(n2cˉ)) in terms of the number of nodes nnn and the average capacity cˉ\bar ccˉ (p. 254). The largest-augmentation rule, often called the fattest-path or maximum-capacity augmenting path rule, is a standard textbook variant, and its geometric-decrease argument is the model for later capacity-scaling methods, including the scaling algorithm for the Hitchcock problem in §2 of the same paper (mission 3 of this series).

The theorem has been proved since 1972 and appears in standard texts. As far as a platform search shows (2026-09-26), no machine-checked proof of it exists on Prove2Me. The platform does contain LinearOptimization.max_flow_min_cut and LinearOptimization.max_flow_ford_fulkerson_integer_termination, which state max-flow min-cut and termination of the generic method in a different network model (parallel arcs, extended nonnegative capacities, no return arc); they give no count of augmentations and are related work only. This mission would contribute a formal proof of the counting bound together with the general labeling-method facts (milestones 1–3), which mission 1 needs as well.

Difficulty

The obvious argument, that each augmentation raises the value by at least 1, gives only the bound f∗(t,s)f^*(t,s)f∗(t,s), and on the four-node example of p. 250 that bound is attained by an arbitrary choice of paths. The logarithmic bound needs a lower bound on the size of the largest augmentation in terms of the remaining gap f∗(t,s)−fk(t,s)f^*(t,s) - f^k(t,s)f∗(t,s)−fk(t,s). The largest augmentation is defined by comparison with all augmenting paths relative to the current flow, while the gap is a global quantity of the network, and neither integrality nor the maximum-augmentation rule alone controls it. Milestone 2's converse, that a non-maximum flow always admits an augmenting path, is itself the max-flow min-cut theorem in this model, and the formal proof has to establish it for the paper's return-arc model rather than import it from a different one.

Formalization scope

  • Nodes form a finite type V with decidable equality. A : Finset (V × V) contains no loops and not (t,s)(t,s)(t,s). Capacities are real, c : V → V → ℝ, positive on A. Integrality is the hypothesis IntegralCaps N, and for the initial flow IsIntegralOn N (f 0) (integer values on the arcs of NNN, the return arc included).
  • Flows are functions V → V → ℝ constrained only on the arcs of NNN. A maximum flow is the predicate IsMaxFlow, comparing f(t,s)f(t,s)f(t,s) with every flow, not a supremum. The goal takes a maximum flow g as a hypothesis and sets f∗(t,s)=g(t,s)f^*(t,s) = g(t,s)f∗(t,s)=g(t,s); every network has one.
  • Augmenting paths are duplicate-free node lists whose consecutive pairs are arcs of NfN^fNf. The page prints Case (b) of the definition of εi\varepsilon_iεi​ with the same hypothesis as Case (c); the corrected Case (b), (u,v)∉A(u,v) \notin A(u,v)∈/A and (v,u)∈A(v,u) \in A(v,u)∈A, is used, as the definition of NfN^fNf (p. 251) and the list for e(u,v)e(u,v)e(u,v) (p. 253) confirm.
  • A run is IsMaxAugRun N K f P. Its initial flow is arbitrary except for integrality, and each later flow is the augmentation of the previous one along a path of maximum ε\varepsilonε among all augmenting paths.
  • The crossing bound CrossArcsBounded N M counts the arcs of NNN, return arc included, with one end on each side of every sss–ttt partition. This is the literal reading of p. 253.
  • Explicit constants. The bound is exactly 1+log⁡M/(M−1)f∗(t,s)1 + \log_{M/(M-1)} f^*(t,s)1+logM/(M−1)​f∗(t,s), written (K : ℝ) ≤ 1 + Real.logb ((M : ℝ) / ((M : ℝ) - 1)) (g N.t N.s) with M>1M > 1M>1 a natural number. When f∗(t,s)=0f^*(t,s) = 0f∗(t,s)=0, Real.logb gives 000 and the bound reads K≤1K \le 1K≤1. The contraction factor is 1 - (M : ℝ)⁻¹.
  • A statement that bounds only runs of an unsatisfiable step predicate, drops the integrality of f0f^0f0 or of the capacities (the bound is false without them), or compares ε\varepsilonε only among paths of some restricted class does not formalize Theorem 2. A sorry-free check exhibits a four-node network with integer capacities and a valid maximum-augmentation step.
  • Reusable beyond this mission: the return-arc network model, the augmentation step with the paper's opposite-arc rule, the integrality lemma, and the cut inequality. Proofs of any milestone are welcome, as are proofs of the converse in milestone 2 that could later be shared with mission 1.

Selected references

  • J. Edmonds, R. M. Karp, Theoretical Improvements in Algorithmic Efficiency for Network Flow Problems, Journal of the ACM 19(2):248–264, 1972. https://doi.org/10.1145/321694.321699
  • L. R. Ford, D. R. Fulkerson, Flows in Networks, RAND report R-375-PR, 1962; Princeton University Press, 1962. https://www.rand.org/pubs/reports/R375.html
11 thms3 active usersReviewed
🏆Completed
Control TheoryDynamic ProgrammingOperations Research·Captain: mikedeng1

Monotone Mappings with Application in Dynamic Programming II: Convergence of the DP Algorithm under Uniform DecreaseResearch Paper

Motivation

Infinite-horizon sequential decision problems (deterministic optimal control, Markov decision processes, minimax control) share one computational question: does the dynamic programming (DP) algorithm, which starts from a terminal cost and repeatedly applies the Bellman operator, converge to the optimal cost? For discounted problems with bounded costs the answer is yes, by the contraction mapping theorem (Blackwell 1965; Denardo 1967). Without discounting and boundedness the answer depends on the sign structure of the problem. Strauch's negative programming model (Strauch 1966) and Blackwell's positive programming model behave differently, and in the former the DP algorithm can fail to converge to the optimal cost even for simple deterministic problems.

Bertsekas (1977) recast these models in one abstract framework: a monotone mapping HHH that encodes the one-stage problem, with no probabilistic or additive structure assumed. Two sign conditions organise the theory: uniform increase (Assumption I, containing Strauch's model) and uniform decrease (Assumption D, containing the deterministic version of Blackwell's positive model, e.g. deterministic problems with nonpositive stage costs). This mission formalizes the uniform-decrease half of Section 5: under D, the finite-horizon problems are solved by the DP algorithm, J∗J^*J∗ is the limit of the finite-horizon values, Bellman's equation holds, and the DP algorithm converges to J∗J^*J∗. The same framework became the basis of Bertsekas–Shreve's Stochastic Optimal Control: The Discrete-Time Case (1978) and of Bertsekas's Abstract Dynamic Programming (2013, 3rd ed. 2022).

Setting

States, controls, policies. SSS (nonempty) and CCC are sets. Each x∈Sx\in Sx∈S has a nonempty constraint set U(x)⊆CU(x)\subseteq CU(x)⊆C. MMM is the set of selectors μ:S→C\mu:S\to Cμ:S→C with μ(x)∈U(x)\mu(x)\in U(x)μ(x)∈U(x) for all xxx, and a policy is a sequence π={μ0,μ1,… }\pi=\{\mu_0,\mu_1,\dots\}π={μ0​,μ1​,…} of selectors. The policy is stationary if μk=μ\mu_k=\muμk​=μ for all kkk.

Functions and the mapping HHH. FFF is the set of functions J:S→[−∞,∞]J:S\to[-\infty,\infty]J:S→[−∞,∞], ordered pointwise, and eee is the constant function 111. A mapping H:S×C×F→[−∞,∞]H:S\times C\times F\to[-\infty,\infty]H:S×C×F→[−∞,∞] is given, and it is monotone: J≤J′J\le J'J≤J′ implies H(x,u,J)≤H(x,u,J′)H(x,u,J)\le H(x,u,J')H(x,u,J)≤H(x,u,J′) for every xxx and u∈U(x)u\in U(x)u∈U(x). It defines

Tμ(J)(x)=H(x,μ(x),J),T(J)(x)=inf⁡u∈U(x)H(x,u,J).T_\mu(J)(x)=H(x,\mu(x),J),\qquad T(J)(x)=\inf_{u\in U(x)}H(x,u,J).Tμ​(J)(x)=H(x,μ(x),J),T(J)(x)=u∈U(x)inf​H(x,u,J).

TkT^kTk is the kkk-fold composition, with T0T^0T0 the identity, and (Tμ0⋯TμN−1)(T_{\mu_0}\cdots T_{\mu_{N-1}})(Tμ0​​⋯TμN−1​​) applies TμN−1T_{\mu_{N-1}}TμN−1​​ first.

Costs. A terminal function Jˉ∈F\bar J\in FJˉ∈F with Jˉ(x)>−∞\bar J(x)>-\inftyJˉ(x)>−∞ is given. The cost of a policy, the optimal cost, the NNN-stage optimal cost and the limit of the DP algorithm are

Jπ=lim⁡N→∞(Tμ0⋯TμN−1)(Jˉ),J∗=inf⁡πJπ,JN=inf⁡π(Tμ0⋯TμN−1)(Jˉ),J∞=lim⁡N→∞TN(Jˉ),J_\pi=\lim_{N\to\infty}(T_{\mu_0}\cdots T_{\mu_{N-1}})(\bar J),\quad J^*=\inf_{\pi}J_\pi,\quad J_N=\inf_{\pi}(T_{\mu_0}\cdots T_{\mu_{N-1}})(\bar J),\quad J_\infty=\lim_{N\to\infty}T^N(\bar J),Jπ​=N→∞lim​(Tμ0​​⋯TμN−1​​)(Jˉ),J∗=πinf​Jπ​,JN​=πinf​(Tμ0​​⋯TμN−1​​)(Jˉ),J∞​=N→∞lim​TN(Jˉ),

all pointwise. JμJ_\muJμ​ denotes the cost of the stationary policy {μ,μ,… }\{\mu,\mu,\dots\}{μ,μ,…}.

Assumptions. D: H(x,u,Jˉ)≤Jˉ(x)H(x,u,\bar J)\le\bar J(x)H(x,u,Jˉ)≤Jˉ(x) for all xxx, u∈U(x)u\in U(x)u∈U(x). Under D every sequence above is nonincreasing, so the limits exist in [−∞,∞][-\infty,\infty][−∞,∞]. D.1: for every sequence with Jk+1≤Jk≤JˉJ_{k+1}\le J_k\le\bar JJk+1​≤Jk​≤Jˉ, lim⁡kH(x,u,Jk)=H(x,u,lim⁡kJk)\lim_k H(x,u,J_k)=H(x,u,\lim_k J_k)limk​H(x,u,Jk​)=H(x,u,limk​Jk​). D.2: there is α>0\alpha>0α>0 such that H(x,u,J)−αr≤H(x,u,J−re)≤H(x,u,J)H(x,u,J)-\alpha r\le H(x,u,J-re)\le H(x,u,J)H(x,u,J)−αr≤H(x,u,J−re)≤H(x,u,J) for all r>0r>0r>0 and J≤JˉJ\le\bar JJ≤Jˉ.

Formalization targets

Goal: convergence of the DP algorithm (Proposition 9)

If D holds, and either D.1 holds or JN=TN(Jˉ)J_N=T^N(\bar J)JN​=TN(Jˉ) for every N≥1N\ge1N≥1, then

J∞=J∗.J_\infty=J^*.J∞​=J∗.

Milestones

  1. Lemma 1. Under D, J∗(x)=lim⁡N→∞JN(x)J^*(x)=\lim_{N\to\infty}J_N(x)J∗(x)=limN→∞​JN​(x) for every xxx.
  2. Proposition 3. Under D, and either D.1 or (D.2 and TN(Jˉ)>−∞T^N(\bar J)>-\inftyTN(Jˉ)>−∞ everywhere), JN=TN(Jˉ)J_N=T^N(\bar J)JN​=TN(Jˉ) for a given N≥1N\ge1N≥1.
  3. Proposition 6. Under D and D.1, J∗=T(J∗)J^*=T(J^*)J∗=T(J∗), and every J′≤JˉJ'\le\bar JJ′≤Jˉ with J′≤T(J′)J'\le T(J')J′≤T(J′) satisfies J′≤J∗J'\le J^*J′≤J∗.
  4. Corollary 6.2. Under D and D.1, Jμ=Tμ(Jμ)J_\mu=T_\mu(J_\mu)Jμ​=Tμ​(Jμ​) for every stationary policy, and every J′≤JˉJ'\le\bar JJ′≤Jˉ with J′≤Tμ(J′)J'\le T_\mu(J')J′≤Tμ​(J′) satisfies J′≤JμJ'\le J_\muJ′≤Jμ​.
  5. Proposition 8. Under D and D.1, a stationary policy {μ∗,μ∗,… }\{\mu^*,\mu^*,\dots\}{μ∗,μ∗,…} is optimal if and only if Tμ∗(Jμ∗)=T(Jμ∗)T_{\mu^*}(J_{\mu^*})=T(J_{\mu^*})Tμ∗​(Jμ∗​)=T(Jμ∗​).

The goal is the paper's answer, in the uniform-decrease case, to the question it poses in the introduction: when is lim⁡NTN(Jˉ)=J∗\lim_N T^N(\bar J)=J^*limN​TN(Jˉ)=J∗?

Significance

The result. Proposition 9 justifies value iteration from Jˉ\bar JJˉ for every problem that fits Assumption D, including deterministic and stochastic control with nonpositive costs (reward maximization with nonnegative rewards) and minimax problems satisfying D.1. Propositions 6 and 8 characterise J∗J^*J∗ as the largest solution of Bellman's equation below Jˉ\bar JJˉ and give a verification test for stationary policies. The hypotheses are sharp in the sense the paper documents: its Counterexamples 2 and 3 show JN≠TN(Jˉ)J_N\ne T^N(\bar J)JN​=TN(Jˉ) when D.1 is dropped together with D.2 or with the finiteness condition TN(Jˉ)>−∞T^N(\bar J)>-\inftyTN(Jˉ)>−∞. Under the mirror assumption I, J∞=J∗J_\infty=J^*J∞​=J∗ can fail, so the asymmetry between the two sign conditions is part of the content.

Formalizing it. All results are proved in the 1977 paper and reappear in later monographs. No machine-checked version of this abstract framework is known. The platform's existing dynamic programming items are finite-state, real-valued and contraction-based, so this mission would add the first formal treatment of extended-real-valued, non-contractive dynamic programming, and a model definition that other results of the same theory can reuse.

Difficulty

The obvious argument for Proposition 9, "JN=TN(Jˉ)J_N=T^N(\bar J)JN​=TN(Jˉ) and JN→J∗J_N\to J^*JN​→J∗", hides two separate interchanges of limits and infima. Lemma 1 interchanges inf⁡π\inf_\piinfπ​ with lim⁡N\lim_NlimN​, which works only because every sequence is monotone in the right direction under D. Proposition 3 is where the work is: the NNN-stage infimum over policies must be matched by the iterated infimum TNT^NTN, which requires building near-optimal selectors stage by stage and passing a limit through HHH NNN times, using D.1, or controlling accumulated errors through D.2. The latter breaks down when values reach −∞-\infty−∞, which is why that branch needs TN(Jˉ)>−∞T^N(\bar J)>-\inftyTN(Jˉ)>−∞. All arithmetic is in [−∞,∞][-\infty,\infty][−∞,∞], where expressions such as ∞−∞\infty-\infty∞−∞ are not defined, and J∗J^*J∗, JNJ_NJN​, TN(Jˉ)T^N(\bar J)TN(Jˉ) may equal −∞-\infty−∞ even though Jˉ\bar JJˉ does not.

Formalization scope

The model is a Lean structure MonotoneDP.Decrease.Model S C with fields U, U_nonempty, H, mono, Jbar, Jbar_ne_bot and S_nonempty; FFF is S → EReal. Policies are ℕ → Selector, where a selector is a function with values in the constraint sets. TTT is an infimum over U x only, and J∗J^*J∗, JNJ_NJN​ are infima over admissible policies. JπJ_\piJπ​ and J∞J_\inftyJ∞​ are limUnder atTop; every theorem assumes D, under which both sequences are nonincreasing and converge, so these are the paper's limits. In D.1 both limits are limUnder. D.2 carries its scalar as a parameter, and "D.2 holds" is ∃ α, AssumptionD2 α. Only real scalars are ever subtracted from extended reals.

JNJ_NJN​ is defined for every NNN, and Propositions 3 and 9 quantify over N≥1N\ge1N≥1 as the paper does. In Proposition 3 the condition TN(Jˉ)>−∞T^N(\bar J)>-\inftyTN(Jˉ)>−∞ belongs to the D.2 branch only. No hypothesis beyond the page is added. Nonempty constraint sets and Jˉ>−∞\bar J>-\inftyJˉ>−∞ are the paper's standing assumptions, stated in the model, not in the theorems. Without nonempty constraint sets there would be no policies, J∗J^*J∗ and JNJ_NJN​ would be +∞+\infty+∞, and several statements would hold trivially; the model rules this out.

Useful contributions: general lemmas about monotone sequences in EReal (interchanging ⨅ and limits), the monotonicity facts (25) and TN+1(Jˉ)≤TN(Jˉ)T^{N+1}(\bar J)\le T^N(\bar J)TN+1(Jˉ)≤TN(Jˉ) under D, and reusable constructions of near-optimal selectors. Corollary 6.1 (the finite-state D.2 variant) is not included.

Selected references

  • D. P. Bertsekas, Monotone mappings with application in dynamic programming, SIAM J. Control Optim. 15(3), 438–464, 1977. https://doi.org/10.1137/0315031
  • E. V. Denardo, Contraction mappings in the theory underlying dynamic programming, SIAM Review 9(2), 165–177, 1967. https://doi.org/10.1137/1009030
  • R. E. Strauch, Negative dynamic programming, Ann. Math. Statist. 37(4), 871–890, 1966. https://doi.org/10.1214/aoms/1177699147
  • D. Blackwell, Discounted dynamic programming, Ann. Math. Statist. 36(1), 226–235, 1965. https://doi.org/10.1214/aoms/1177700285
  • D. P. Bertsekas, Abstract Dynamic Programming, 3rd ed., Athena Scientific, 2022. https://www.mit.edu/~dimitrib/abstractdp_MIT.html
8 thms3 active usersReviewed
🏆Completed
Convex OptimizationFunctional AnalysisOperations Research·Captain: mikedeng1

A Three-Operator Splitting Scheme and its Optimization Applications 3: Accelerated Convergence under Strong MonotonicityResearch Paper

Motivation

Many problems in convex optimization, variational inequalities and signal processing reduce to finding a zero of a sum of three monotone operators, one of which is single-valued and smooth. Davis and Yin (Set-Valued Var. Anal. 25, 2017) introduced a splitting scheme that evaluates each of the three operators separately: the two set-valued ones through their resolvents, the single-valued one through a forward step. With a fixed stepsize, their Algorithm 1 converges weakly but can be slow: the paper's Section 3.4 constructs examples where the squared distance of the iterates to the solution decays no faster than (k+1)−(1+ϵ)(k+1)^{-(1+\epsilon)}(k+1)−(1+ϵ) for every ϵ>0\epsilon > 0ϵ>0.

When one of the operators is strongly monotone (for example the subdifferential of a strongly convex function), first-order splitting methods can be accelerated by letting the stepsize shrink like 1/k1/k1/k; the paper relates its stepsizes to those of Chambolle and Pock's accelerated primal–dual method (J. Math. Imaging Vis. 40, 2011, Algorithm 2) and of Boţ, Csetnek, Heinrich and Hendrich (Math. Program. 150, 2015, Algorithm 5). Section 3.3 of Davis–Yin carries this device over to three-operator splitting and obtains an O(1/(k+1)2)O(1/(k+1)^2)O(1/(k+1)2) rate for the squared distance. This mission formalizes that result.

Setting

Let HHH be a real Hilbert space. A set-valued operator A:H→2HA : H \to 2^HA:H→2H is monotone if ⟨x−y,u−v⟩≥0\langle x - y, u - v\rangle \ge 0⟨x−y,u−v⟩≥0 for all u∈Axu \in Axu∈Ax, v∈Ayv \in Ayv∈Ay, and maximal monotone if its graph is not properly contained in the graph of another monotone operator. It is μ\muμ-strongly monotone if ⟨x−y,u−v⟩≥μ∥x−y∥2\langle x - y, u - v\rangle \ge \mu\|x-y\|^2⟨x−y,u−v⟩≥μ∥x−y∥2 for all such pairs. A single-valued C:H→HC : H \to HC:H→H is β\betaβ-cocoercive if β∥Cx−Cy∥2≤⟨Cx−Cy,x−y⟩\beta\|Cx - Cy\|^2 \le \langle Cx - Cy, x - y\rangleβ∥Cx−Cy∥2≤⟨Cx−Cy,x−y⟩, and LCL_CLC​-Lipschitz if ∥Cx−Cy∥≤LC∥x−y∥\|Cx - Cy\| \le L_C\|x - y\|∥Cx−Cy∥≤LC​∥x−y∥.

The problem is to find x∗∈zer⁡(A+B+C)x^* \in \operatorname{zer}(A + B + C)x∗∈zer(A+B+C), that is, 0∈Ax∗+Bx∗+Cx∗0 \in Ax^* + Bx^* + Cx^*0∈Ax∗+Bx∗+Cx∗, where AAA, BBB are maximal monotone and CCC is monotone and single-valued. For γ>0\gamma > 0γ>0 the resolvent JγA=(I+γA)−1J_{\gamma A} = (I + \gamma A)^{-1}JγA​=(I+γA)−1 is the map with x∈JγAx+γA(JγAx)x \in J_{\gamma A}x + \gamma A(J_{\gamma A}x)x∈JγA​x+γA(JγA​x).

Algorithm 3 fixes stepsizes (γk)k≥0⊆(0,∞)(\gamma_k)_{k\ge 0} \subseteq (0,\infty)(γk​)k≥0​⊆(0,∞) and an initial point xA0∈Hx_A^0 \in HxA0​∈H, sets xB0=Jγ0B(xA0)x_B^0 = J_{\gamma_0 B}(x_A^0)xB0​=Jγ0​B​(xA0​), uB0=γ0−1(xA0−xB0)u_B^0 = \gamma_0^{-1}(x_A^0 - x_B^0)uB0​=γ0−1​(xA0​−xB0​), and iterates for k≥0k \ge 0k≥0

xBk+1=JγkB(xAk+γkuBk),uBk+1=1γk(xAk+γkuBk−xBk+1),xAk+1=Jγk+1A(xBk+1−γk+1uBk+1−γk+1CxBk+1).x_B^{k+1} = J_{\gamma_k B}(x_A^k + \gamma_k u_B^k),\quad u_B^{k+1} = \tfrac{1}{\gamma_k}(x_A^k + \gamma_k u_B^k - x_B^{k+1}),\quad x_A^{k+1} = J_{\gamma_{k+1}A}(x_B^{k+1} - \gamma_{k+1}u_B^{k+1} - \gamma_{k+1}Cx_B^{k+1}).xBk+1​=Jγk​B​(xAk​+γk​uBk​),uBk+1​=γk​1​(xAk​+γk​uBk​−xBk+1​),xAk+1​=Jγk+1​A​(xBk+1​−γk+1​uBk+1​−γk+1​CxBk+1​).

The stepsize changes in the middle of an iteration. Two stepsize rules are considered, each defined recursively from γ0\gamma_0γ0​:

(3.6)γk+1=−2γk2μCη+(2γk2μCη)2+4(1+2γkμB)γk22(1+2γkμB),(3.7)γk+1=γk1+2γk(μB−γkLC2/2).\text{(3.6)}\quad \gamma_{k+1} = \frac{-2\gamma_k^2\mu_C\eta + \sqrt{(2\gamma_k^2\mu_C\eta)^2 + 4(1+2\gamma_k\mu_B)\gamma_k^2}}{2(1+2\gamma_k\mu_B)}, \qquad \text{(3.7)}\quad \gamma_{k+1} = \frac{\gamma_k}{\sqrt{1 + 2\gamma_k(\mu_B - \gamma_kL_C^2/2)}}.(3.6)γk+1​=2(1+2γk​μB​)−2γk2​μC​η+(2γk2​μC​η)2+4(1+2γk​μB​)γk2​​​,(3.7)γk+1​=1+2γk​(μB​−γk​LC2​/2)​γk​​.

Formalization targets

Goal: Theorem 3.3, both parts

Let BBB be μB\mu_BμB​-strongly monotone with μB≥0\mu_B \ge 0μB​≥0.

  1. If CCC is β\betaβ-cocoercive and μC\mu_CμC​-strongly monotone (μC>0\mu_C > 0μC​>0), η∈(0,1)\eta \in (0,1)η∈(0,1), γ0∈(0,2β(1−η))\gamma_0 \in (0, 2\beta(1-\eta))γ0​∈(0,2β(1−η)) and the stepsizes follow (3.6), then for every x∗∈zer⁡(A+B+C)x^* \in \operatorname{zer}(A+B+C)x∗∈zer(A+B+C)
∃K ∀k≥0:∥xBk−x∗∥2≤K(k+1)2.\exists K\ \forall k \ge 0:\quad \|x_B^k - x^*\|^2 \le \frac{K}{(k+1)^2}.∃K ∀k≥0:∥xBk​−x∗∥2≤(k+1)2K​.
  1. If CCC is LCL_CLC​-Lipschitz, μB>0\mu_B > 0μB​>0, γ0∈(0,2μB/LC2)\gamma_0 \in (0, 2\mu_B/L_C^2)γ0​∈(0,2μB​/LC2​) and the stepsizes follow (3.7), the same conclusion holds.

The goal asserts the shape of the rate only; the constant KKK is not fixed.

Milestones

  • Proposition 3.1, Parts 1 and 2: the one-step inequalities (3.9) and (3.10) for Algorithm 3 with arbitrary admissible stepsizes.
  • Stepsize facts from the proof of Theorem 3.3: the identities that make (3.9) and (3.10) telescope, the monotonicity of the stepsizes (3.6), and the limits (k+1)γk→1/(μCη+μB)(k+1)\gamma_k \to 1/(\mu_C\eta + \mu_B)(k+1)γk​→1/(μC​η+μB​) for (3.6) and (k+1)γk→1/μB(k+1)\gamma_k \to 1/\mu_B(k+1)γk​→1/μB​ for (3.7).

Significance

The theorem shows that strong monotonicity of BBB or CCC can be converted into a quadratically decaying distance bound without knowledge of the solution, with stepsizes that are computable from the strong monotonicity and cocoercivity (or Lipschitz) constants alone. Since the rate is established for xBkx_B^kxBk​, it applies directly to splitting schemes for strongly convex composite problems min⁡f+g+h\min f + g + hminf+g+h with hhh smooth, where xBkx_B^kxBk​ is the proximal point of ggg.

The result is proved in the paper; no machine-checked version is known. The formalization adds a precise statement of the admissible parameter ranges, a check of the index conventions of a scheme whose stepsize changes mid-iteration, and a correction of the one-step inequalities at the first iteration (see Formalization scope). The stepsize limits are statements about explicit real recursions and are of independent use for other accelerated schemes.

Difficulty

The one-step inequalities (3.9) and (3.10) are long but elementary chains of inner-product identities and Young's inequality; the work lies in bookkeeping two stepsizes per iteration. The rate itself does not follow from the one-step inequality alone: telescoping gives a bound of the form ∥xBk−x∗∥2≲γk2\|x_B^{k}-x^*\|^2 \lesssim \gamma_k^2∥xBk​−x∗∥2≲γk2​, and one must then show γk\gamma_kγk​ decays exactly like 1/k1/k1/k. The rules (3.6) and (3.7) are nonlinear recursions without closed form, so their asymptotics require a Stolz–Cesàro type argument, which is not available in Mathlib under that name. Choosing a stepsize sequence of the form c/kc/kc/k instead is a different algorithm and not covered by the theorem.

Formalization scope

  • HHH is an arbitrary real Hilbert space (InnerProductSpace ℝ H, CompleteSpace H), not a Euclidean space.
  • Resolvents are not constructed. They are families JA JB : ℝ → H → H required to satisfy the resolvent inclusion γ−1(x−J(γ)x)∈A(J(γ)x)\gamma^{-1}(x - J(\gamma)x) \in A(J(\gamma)x)γ−1(x−J(γ)x)∈A(J(γ)x) for every γ>0\gamma > 0γ>0; for maximal monotone operators such maps exist and are unique, so nothing is lost.
  • Algorithm 3 is a single recursive definition of the triple (xAk,xBk,uBk)(x_A^k, x_B^k, u_B^k)(xAk​,xBk​,uBk​) from xA0x_A^0xA0​, the stepsizes, the resolvent families and CCC; the paper's loop index k=1,2,…k = 1, 2, \dotsk=1,2,… matches recursion (3.8) shifted by one.
  • The stepsize rules (3.6) and (3.7) are recursive real sequences, used verbatim; each theorem assumes the paper's parameter ranges.
  • O(1/(k+1)2)O(1/(k+1)^2)O(1/(k+1)2) is rendered as ∃K ∀k, ∥xBk−x∗∥2≤K/(k+1)2\exists K\,\forall k,\ \|x_B^k - x^*\|^2 \le K/(k+1)^2∃K∀k, ∥xBk​−x∗∥2≤K/(k+1)2, with KKK chosen after all data (initial point, operators, constants, γ0\gamma_0γ0​, x∗x^*x∗) and before kkk. No explicit constant is stated.
  • Strong monotonicity of CCC means μC>0\mu_C > 0μC​>0; only μB=0\mu_B = 0μB​=0 is allowed, as on the page. With μB=μC=0\mu_B = \mu_C = 0μB​=μC​=0 rule (3.6) keeps γk\gamma_kγk​ constant and the rate fails, so a formalization allowing μC=0\mu_C = 0μC​=0 would be false. In Part 2, LC>0L_C > 0LC​>0 is assumed so that the stepsize interval is meaningful, and CCC is assumed monotone, as in problem (1.1) and as used in the paper's proof of (3.10).
  • The paper states (3.9) and (3.10) for all k≥0k \ge 0k≥0; at k=0k = 0k=0 the initial point xA0x_A^0xA0​ is not a resolvent output, and both inequalities fail in general. The milestones state them for k≥1k \ge 1k≥1. Theorem 3.3 is unaffected, since finitely many initial terms do not change an O(⋅)O(\cdot)O(⋅) bound.
  • The display γk2−γk+12=γkγk+1(2γkμB+2γk+1μCη)\gamma_k^2 - \gamma_{k+1}^2 = \gamma_k\gamma_{k+1}(2\gamma_k\mu_B + 2\gamma_{k+1}\mu_C\eta)γk2​−γk+12​=γk​γk+1​(2γk​μB​+2γk+1​μC​η) on p. 845 has γk\gamma_kγk​ and γk+1\gamma_{k+1}γk+1​ swapped inside the bracket; the milestone states the corrected identity γkγk+1(2γk+1μB+2γkμCη)\gamma_k\gamma_{k+1}(2\gamma_{k+1}\mu_B + 2\gamma_k\mu_C\eta)γk​γk+1​(2γk+1​μB​+2γk​μC​η).
  • A trivializing formalization, such as one in which the resolvent hypothesis is unsatisfiable, the stepsize interval is empty, or the rate constant may depend on kkk, is ruled out: the hypotheses are met by A=0A = 0A=0, B=μBIB = \mu_B IB=μB​I (with resolvents JγA=IJ_{\gamma A} = IJγA​=I, JγB=(1+γμB)−1IJ_{\gamma B} = (1+\gamma\mu_B)^{-1}IJγB​=(1+γμB​)−1I) and C=cIC = cIC=cI with c>0c > 0c>0, and KKK is quantified before kkk.

Contributions are welcome at every level: proofs of the real-sequence milestones (a general Stolz–Cesàro lemma would be reusable well beyond this mission), of the two one-step inequalities, and of the telescoping argument that assembles the goal.

Selected references

  • D. Davis and W. Yin, A Three-Operator Splitting Scheme and its Optimization Applications, Set-Valued and Variational Analysis 25 (2017), 829–858. https://doi.org/10.1007/s11228-017-0421-z (preprint: https://arxiv.org/abs/1504.01032)
  • R. I. Boţ, E. R. Csetnek, A. Heinrich and C. Hendrich, On the convergence rate improvement of a primal-dual splitting algorithm for solving monotone inclusion problems, Mathematical Programming 150 (2015), 251–279. https://doi.org/10.1007/s10107-014-0766-0
  • A. Chambolle and T. Pock, A First-Order Primal-Dual Algorithm for Convex Problems with Applications to Imaging, Journal of Mathematical Imaging and Vision 40 (2011), 120–145. https://doi.org/10.1007/s10851-010-0251-1
  • H. H. Bauschke and P. L. Combettes, Convex Analysis and Monotone Operator Theory in Hilbert Spaces, 2nd ed., Springer, 2017. https://doi.org/10.1007/978-3-319-48311-5
11 thms3 active usersReviewed
🏆Completed
Convex OptimizationOperations Research·Captain: mikedeng1

On Polyhedral Approximations of the Second-Order Cone III: Closeness of the Relaxed Feasible SetResearch Paper

Motivation

Conic quadratic problems (also called second-order cone programs) arise directly in applications such as contact problems with Coulomb friction, and a wide range of nonlinear convex problems can be rewritten in this form (Lobo, Vandenberghe, Boyd and Lebret 1998). Interior-point methods solve them in polynomial time, but around 2000 the available software for conic quadratic problems handled far fewer variables than linear programming software. Ben-Tal and Nemirovski (2001) therefore asked whether a conic quadratic problem can be replaced by a linear program of comparable size. Their construction replaces each second-order cone by a polyhedral cone that is exact up to a factor 1+ε1+\varepsilon1+ε. The feasible set of the resulting linear program, projected back to the original variables, lies between the feasible set of the original problem and that of its ε\varepsilonε-relaxation.

This sandwich is only useful if the relaxed problem is close to the original one, and in general it is not: the paper notes that (CQP) can be infeasible while every relaxation with ε>0\varepsilon>0ε>0 is feasible. Proposition 4.1 of the paper, the target of this mission, gives a sufficient condition under which the two feasible sets are O(ε)O(\varepsilon)O(ε)-close.

Setting

For y∈Rky\in\mathbb R^ky∈Rk let ∥y∥2=yTy\|y\|_2=\sqrt{y^Ty}∥y∥2​=yTy​ be the Euclidean norm. A conic quadratic problem in the variable x∈Rnx\in\mathbb R^nx∈Rn is

(CQP)min⁡x{eTx∣Ax≥b, ∥Aℓx−bℓ∥2≤cℓTx−dℓ, ℓ=1,…,m},\text{(CQP)}\qquad \min_x\bigl\{e^Tx \bigm| Ax\ge b,\ \|A_\ell x-b_\ell\|_2\le c_\ell^Tx-d_\ell,\ \ell=1,\dots,m\bigr\},(CQP)xmin​{eTx​Ax≥b, ∥Aℓ​x−bℓ​∥2​≤cℓT​x−dℓ​, ℓ=1,…,m},

where AAA is a k0×nk_0\times nk0​×n matrix and b∈Rk0b\in\mathbb R^{k_0}b∈Rk0​ (the inequality Ax≥bAx\ge bAx≥b is componentwise), and for each ℓ\ellℓ the matrix AℓA_\ellAℓ​ is kℓ×nk_\ell\times nkℓ​×n, bℓ∈Rkℓb_\ell\in\mathbb R^{k_\ell}bℓ​∈Rkℓ​, cℓ∈Rnc_\ell\in\mathbb R^ncℓ​∈Rn and dℓ∈Rd_\ell\in\mathbb Rdℓ​∈R. For ε>0\varepsilon>0ε>0 the ε\varepsilonε-relaxation is

(CQPε)min⁡x{eTx∣Ax≥b, ∥Aℓx−bℓ∥2≤(1+ε)[cℓTx−dℓ], ℓ=1,…,m}.\text{(CQP}_\varepsilon)\qquad \min_x\bigl\{e^Tx \bigm| Ax\ge b,\ \|A_\ell x-b_\ell\|_2\le (1+\varepsilon)\bigl[c_\ell^Tx-d_\ell\bigr],\ \ell=1,\dots,m\bigr\}.(CQPε​)xmin​{eTx​Ax≥b, ∥Aℓ​x−bℓ​∥2​≤(1+ε)[cℓT​x−dℓ​], ℓ=1,…,m}.

Feas(P)\mathrm{Feas}(P)Feas(P) denotes the feasible set of a problem (P)(P)(P); in Lean these are feas P and feasRelaxed P ε, subsets of Fin n → ℝ, for a problem datum P : CQP n k₀ m.

Two conditions on (CQP) are used.

  1. Strict feasibility: there are xˉ\bar xxˉ and r>0r>0r>0 with Axˉ≥bA\bar x\ge bAxˉ≥b and ∥Aℓxˉ−bℓ∥2≤[cℓTxˉ−dℓ]−r\|A_\ell\bar x-b_\ell\|_2\le[c_\ell^T\bar x-d_\ell]-r∥Aℓ​xˉ−bℓ​∥2​≤[cℓT​xˉ−dℓ​]−r for every ℓ\ellℓ (IsStrictlyFeasible P x̄ r).
  2. Semiboundedness: there is RRR such that every feasible xxx of (CQP) satisfies cℓTx−dℓ≤Rc_\ell^Tx-d_\ell\le RcℓT​x−dℓ​≤R for every ℓ\ellℓ (IsSemibounded P R).

Put γ(ε)=Rε/r\gamma(\varepsilon)=R\varepsilon/rγ(ε)=Rε/r.

Formalization targets

Goal: Proposition 4.1

If (CQP) has m≥1m\ge1m≥1 conic constraints and is strictly feasible and semibounded, then for every ε>0\varepsilon>0ε>0 with γ(ε)<1\gamma(\varepsilon)<1γ(ε)<1,

γ(ε)xˉ+(1−γ(ε)) Feas(CQPε) ⊆ Feas(CQP) ⊆ Feas(CQPε).(14)\gamma(\varepsilon)\bar x+(1-\gamma(\varepsilon))\,\mathrm{Feas}(\mathrm{CQP}_\varepsilon)\ \subseteq\ \mathrm{Feas}(\mathrm{CQP})\ \subseteq\ \mathrm{Feas}(\mathrm{CQP}_\varepsilon). \tag{14}γ(ε)xˉ+(1−γ(ε))Feas(CQPε​) ⊆ Feas(CQP) ⊆ Feas(CQPε​).(14)

The left-hand side is the image of Feas(CQPε)\mathrm{Feas}(\mathrm{CQP}_\varepsilon)Feas(CQPε​) under y↦γ(ε)xˉ+(1−γ(ε))yy\mapsto\gamma(\varepsilon)\bar x+(1-\gamma(\varepsilon))yy↦γ(ε)xˉ+(1−γ(ε))y, not a Minkowski sum.

Milestones

The milestones follow the paper's proof in order.

  1. The right inclusion Feas(CQP)⊆Feas(CQPε)\mathrm{Feas}(\mathrm{CQP})\subseteq\mathrm{Feas}(\mathrm{CQP}_\varepsilon)Feas(CQP)⊆Feas(CQPε​) for ε>0\varepsilon>0ε>0.
  2. For y∈Feas(CQPε)y\in\mathrm{Feas}(\mathrm{CQP}_\varepsilon)y∈Feas(CQPε​) and tℓ=cℓTy−dℓt_\ell=c_\ell^Ty-d_\elltℓ​=cℓT​y−dℓ​, every δ∈[0,1]\delta\in[0,1]δ∈[0,1] with δ≥εtℓ/(r+εtℓ)\delta\ge\varepsilon t_\ell/(r+\varepsilon t_\ell)δ≥εtℓ​/(r+εtℓ​) for all ℓ\ellℓ makes xδ=(1−δ)y+δxˉx_\delta=(1-\delta)y+\delta\bar xxδ​=(1−δ)y+δxˉ feasible for (CQP).
  3. Under semiboundedness, the same δ\deltaδ satisfies (1−δ)tℓ≤R(1-\delta)t_\ell\le R(1−δ)tℓ​≤R for all ℓ\ellℓ.
  4. If δ=εt/(r+εt)\delta=\varepsilon t/(r+\varepsilon t)δ=εt/(r+εt) with t≥0t\ge0t≥0, (1−δ)t≤R(1-\delta)t\le R(1−δ)t≤R and γ(ε)<1\gamma(\varepsilon)<1γ(ε)<1, then t≤R/(1−γ(ε))t\le R/(1-\gamma(\varepsilon))t≤R/(1−γ(ε)) and δ≤γ(ε)\delta\le\gamma(\varepsilon)δ≤γ(ε).

Significance

The result. Proposition 4.1 turns the qualitative sandwich "exact ⊆ polyhedral ⊆ relaxed" into a quantitative statement. When a problem is strictly feasible with margin rrr and its conic right-hand sides are bounded by RRR on the feasible set, the relaxed feasible set, shrunk towards xˉ\bar xxˉ by 1−γ(ε)1-\gamma(\varepsilon)1−γ(ε), lies inside the exact one. The error of the relaxation is thus controlled by γ(ε)=Rε/r\gamma(\varepsilon)=R\varepsilon/rγ(ε)=Rε/r, which is linear in ε\varepsilonε. Together with the paper's main theorem, that a polyhedral ε\varepsilonε-approximation of the Lorentz cone with O(kln⁡(1/ε))O(k\ln(1/\varepsilon))O(kln(1/ε)) variables and inequalities exists, this measures how well a linear program of moderate size approximates the conic problem. The paper uses it this way for the examples in its introduction.

The formalization. The proposition is proved in the paper; no machine-checked version is known. This mission produces a Lean formalization of conic quadratic problems and their relaxations with the Euclidean norm, together with the strict feasibility and semiboundedness conditions and the proof. The Lorentz-cone approximation results of the same paper are the subject of the companion missions I and II of this series.

Difficulty

The right inclusion is immediate. The left inclusion does not follow from convexity alone. A relaxed-feasible point yyy may violate every conic constraint of (CQP), and nothing about yyy bounds how far it is from Feas(CQP)\mathrm{Feas}(\mathrm{CQP})Feas(CQP). The needed information comes from semiboundedness, which constrains only feasible points of (CQP). That hypothesis therefore cannot be applied to yyy itself, and the shrink factor γ(ε)\gamma(\varepsilon)γ(ε) must be obtained without any bound on cℓTy−dℓc_\ell^Ty-d_\ellcℓT​y−dℓ​ given in advance. The obvious attempt, bounding the violation at yyy by εR\varepsilon RεR, fails for exactly this reason.

Formalization scope

  • Vectors of Rn\mathbb R^nRn are Fin n → ℝ; the mmm conic constraints are indexed by Fin m (0-based) with a dependent family of matrices (ℓ : Fin m) → Matrix (Fin (k ℓ)) (Fin n) ℝ, so the row sizes kℓk_\ellkℓ​ may differ. The norm is written out as eucNorm y = √(∑ i, y i ^ 2); Mathlib's norm on Fin k → ℝ is the sup norm and is not used.
  • Only feasible sets are compared; the objective eee is carried as data but plays no role.
  • Correction 1. In hypothesis (i) the page prints [cℓTx−dℓ]−r[c_\ell^Tx-d_\ell]-r[cℓT​x−dℓ​]−r without the bar over xxx. The proof uses cℓTxˉ−dℓ−rc_\ell^T\bar x-d_\ell-rcℓT​xˉ−dℓ​−r, which is what IsStrictlyFeasible states.
  • Correction 2. The goal assumes m≥1m\ge1m≥1, which the paper leaves implicit. With m=0m=0m=0, semiboundedness is vacuous and RRR may be negative, so γ(ε)<0\gamma(\varepsilon)<0γ(ε)<0. Then the map y↦γxˉ+(1−γ)yy\mapsto\gamma\bar x+(1-\gamma)yy↦γxˉ+(1−γ)y extrapolates beyond yyy and can leave {Ax≥b}\{Ax\ge b\}{Ax≥b}. An example is n=1n=1n=1, A=[1]A=[1]A=[1], b=0b=0b=0, xˉ=1\bar x=1xˉ=1, y=0y=0y=0, R=−1R=-1R=−1, r=ε=1r=\varepsilon=1r=ε=1. For m≥1m\ge1m≥1 the hypotheses force R≥r>0R\ge r>0R≥r>0.
  • ε\varepsilonε ranges over all ε>0\varepsilon>0ε>0 with γ(ε)<1\gamma(\varepsilon)<1γ(ε)<1, as in the paper; it is not restricted to (0,1](0,1](0,1].
  • The second milestone is stated for every δ∈[0,1]\delta\in[0,1]δ∈[0,1] that dominates all ratios εtℓ/(r+εtℓ)\varepsilon t_\ell/(r+\varepsilon t_\ell)εtℓ​/(r+εtℓ​), rather than only for the paper's δ=max⁡ℓ\delta=\max_\ellδ=maxℓ​. This includes the paper's case.
  • The goal cannot be satisfied trivially. The strict feasibility and semiboundedness hypotheses are jointly satisfiable (for example n=m=1n=m=1n=m=1, the constraint ∣x∣≤1|x|\le 1∣x∣≤1 written as ∥x∥2≤1\|x\|_2\le 1∥x∥2​≤1, xˉ=0\bar x=0xˉ=0, r=1r=1r=1, R=1R=1R=1), and the conclusion is the full two-sided inclusion with the paper's γ(ε)\gamma(\varepsilon)γ(ε), not the existence of some contraction factor.
  • Needed infrastructure: Euclidean-norm convexity (the triangle inequality and homogeneity for eucNorm, or a transfer to EuclideanSpace ℝ (Fin k)) and linearity of Matrix.mulVec and dotProduct. A convexity lemma for feas P would be reusable beyond this mission, and contributions of it are welcome.

Selected references

  • A. Ben-Tal and A. Nemirovski, On Polyhedral Approximations of the Second-Order Cone, Mathematics of Operations Research 26(2):193–205, 2001. https://doi.org/10.1287/moor.26.2.193.10561
  • M. S. Lobo, L. Vandenberghe, S. Boyd and H. Lebret, Applications of Second-Order Cone Programming, Linear Algebra and its Applications 284:193–228, 1998. https://doi.org/10.1016/S0024-3795(98)10032-0
  • Yu. Nesterov and A. Nemirovski, Interior-Point Polynomial Algorithms in Convex Programming, SIAM, 1994. https://doi.org/10.1137/1.9781611970791
7 thms3 active usersReviewed
🏆Completed
Operations ResearchProbability·Captain: mikedeng1

On Minimizing a Convex Function Subject to Linear Inequalities III: The Expected Cost of a Linear Program with Random Coefficients Is ConvexResearch Paper

Motivation

A linear program is solved with known data, but in planning problems the data are often only known in distribution when the main decision is taken: demands, yields and requirements are revealed later, and a corrective action is taken after they are. E. M. L. Beale's 1955 paper On Minimizing a Convex Function Subject to Linear Inequalities formulates this situation in its §5, "Linear Programming with Random Coefficients", as what is now called a two-stage stochastic linear program with recourse. Beale's motivating example is the transportation problem of Hitchcock (1941) with random requirements at the destinations, where every unit of shortage or excess incurs a loss. The same model was put forward in the same year by Dantzig, Linear Programming under Uncertainty (Management Science, 1955), as the paper's note added in proof acknowledges.

Timeline:

  • 1955. Beale (§5, Theorems 2 and 3) and Dantzig independently introduce two-stage linear programs with random data; Beale proves that the expected cost is convex in the first-stage decision, and that the cost is convex in the random data for fixed decision.
  • 1967. Walkup and Wets, Stochastic Programs with Recourse, study the domain of the expected recourse function and its properties under fixed recourse.
  • 1974. Wets, Stochastic Programs with Fixed Recourse: The Equivalent Deterministic Program, gives the systematic treatment of convexity, finiteness and polyhedrality of the expected recourse function, now textbook material (Birge and Louveaux, Introduction to Stochastic Programming, Ch. 3).

Setting

Constants c∈Rnc\in\mathbb R^nc∈Rn, f∈Rpf\in\mathbb R^pf∈Rp and an m×pm\times pm×p matrix D=(dik)D=(d_{ik})D=(dik​) are given. The data A=(αij)A=(\alpha_{ij})A=(αij​), an m×nm\times nm×n matrix, and β∈Rm\beta\in\mathbb R^mβ∈Rm are random variables on a probability space (Ω,P)(\Omega,P)(Ω,P): their distribution is known when the first-stage decision x∈Rnx\in\mathbb R^nx∈Rn, x≥0x\ge0x≥0, is chosen, and their values are known when the second-stage decision y∈Rpy\in\mathbb R^py∈Rp, y≥0y\ge0y≥0, is chosen. The cost is

C=c′x+f′y,Ax+Dy=β.(5.3),(5.4)C=c'x+f'y,\qquad Ax+Dy=\beta. \qquad(5.3),(5.4)C=c′x+f′y,Ax+Dy=β.(5.3),(5.4)

For a right-hand side b∈Rmb\in\mathbb R^mb∈Rm the second-stage value is

Q(b)=min⁡{f′y:y≥0, Dy=b},Q(b)=\min\{f'y : y\ge0,\ Dy=b\},Q(b)=min{f′y:y≥0, Dy=b},

and for fixed data the cost of a first-stage decision is C(x)=c′x+Q(β−Ax)C(x)=c'x+Q(\beta-Ax)C(x)=c′x+Q(β−Ax). The expected cost is

E(C)(x)=∫Ω(c′x+Q(β(ω)−A(ω)x)) dP(ω).E(C)(x)=\int_\Omega \bigl(c'x+Q(\beta(\omega)-A(\omega)x)\bigr)\,dP(\omega).E(C)(x)=∫Ω​(c′x+Q(β(ω)−A(ω)x))dP(ω).

The problem is to choose x≥0x\ge0x≥0 minimising E(C)E(C)E(C). In Lean the value is secondStageValue D f b, the cost is cost c f D A β x, and the expected cost is expectedCost P c f D A β x, all in the namespace BealeConvexMin.RandomLP.

Formalization targets

Goal: Theorem 2 (p. 182)

Assume that for every x≥0x\ge0x≥0 the second-stage minimum is attained for almost every outcome and that ω↦C(x,ω)\omega\mapsto C(x,\omega)ω↦C(x,ω) is integrable. Then

E(C)(λ1x1+λ2x2)≤λ1E(C)(x1)+λ2E(C)(x2)(x1,x2≥0, λ1,λ2≥0, λ1+λ2=1),E(C)(\lambda_1x_1+\lambda_2x_2)\le\lambda_1E(C)(x_1)+\lambda_2E(C)(x_2)\qquad(x_1,x_2\ge0,\ \lambda_1,\lambda_2\ge0,\ \lambda_1+\lambda_2=1),E(C)(λ1​x1​+λ2​x2​)≤λ1​E(C)(x1​)+λ2​E(C)(x2​)(x1​,x2​≥0, λ1​,λ2​≥0, λ1​+λ2​=1),

that is, E(C)E(C)E(C) is convex on the non-negative orthant. The statement fixes no distribution class: it is claimed for any known distribution of (A,β)(A,\beta)(A,β).

Milestones

  1. Pointwise convexity (last display of the proof of Theorem 2, p. 182): for fixed data (A,β)(A,\beta)(A,β), with the minimum attained at every x≥0x\ge0x≥0,
C(λ1x1+λ2x2)≤λ1C(x1)+λ2C(x2).C(\lambda_1x_1+\lambda_2x_2)\le\lambda_1C(x_1)+\lambda_2C(x_2).C(λ1​x1​+λ2​x2​)≤λ1​C(x1​)+λ2​C(x2​).
  1. Theorem 3 (p. 182): for fixed xxx, the cost (A,β)↦c′x+Q(β−Ax)(A,\beta)\mapsto c'x+Q(\beta-Ax)(A,β)↦c′x+Q(β−Ax) is jointly convex on every convex set of data on which the second-stage minimum is attained.
  2. Eqs. (5.5)–(5.6) (p. 182): for a finitely supported distribution, A=ArA=A_rA=Ar​ and β=βr\beta=\beta_rβ=βr​ with probability prp_rpr​, the value E(C)(x)E(C)(x)E(C)(x) is the minimum of c′x+∑rprf′yrc'x+\sum_r p_r f'y_rc′x+∑r​pr​f′yr​ over non-negative yry_ryr​ with Arx+Dyr=βrA_rx+Dy_r=\beta_rAr​x+Dyr​=βr​ for all rrr; minimising E(C)E(C)E(C) is then a linear program.

Significance

The result. Theorem 2 is the basic structural fact of two-stage stochastic linear programming: the first-stage problem is a convex program in xxx, whatever the distribution of the data. It is what makes local optimality global for the first-stage problem, what justifies cutting-plane and decomposition methods that approximate E(C)E(C)E(C) from below by supporting hyperplanes, and what makes sample-average approximations convex programs. Theorem 3, joint convexity in the data, gives through Jensen's inequality the comparison between the stochastic problem and its mean-value problem that Beale draws on p. 182. The discrete reformulation (5.5)–(5.6) is the deterministic-equivalent linear program used for finitely many scenarios.

Formalizing it. The theorems are proved in the paper, and their content is classical. The mission produces machine-checked statements of the model with its implicit hypotheses made explicit (attainment of the second stage, integrability of the cost), and proofs of the three results in Lean. The platform already has related statements in other models (finite scenario sets with extended-real recourse, and a complete-recourse, finite-second-moment version); none has Beale's hypotheses, and none states convexity of c′x+E Qc'x+E\,Qc′x+EQ for an arbitrary distribution.

Difficulty

The mathematics is short; the difficulty is in the encoding. The second-stage value is a minimum that may fail to exist: the second stage may be infeasible for some xxx and some outcomes, or unbounded below. A real-valued infimum then takes an arbitrary default value, and convexity would become a statement about that default. Similarly, the mean value only exists when the cost is integrable. A faithful statement has to carry attainment and integrability exactly where the paper tacitly assumes them, on the domain x≥0x\ge0x≥0 the paper uses, and no stronger condition (such as complete recourse or moment bounds) that the paper does not make. In the discrete reformulation, the minimum over the whole family (yr)r(y_r)_r(yr​)r​ has to be matched with the probability-weighted sum of per-scenario minima.

Formalization scope

  • Vectors are Fin n → ℝ, matrices Matrix (Fin m) (Fin n) ℝ, inner products dotProduct, and y≥0y\ge0y≥0 is the componentwise order. The random data are functions A : Ω → Matrix (Fin m) (Fin n) ℝ and β : Ω → Fin m → ℝ on a measurable space with a probability measure P; no measurability of the data is assumed beyond integrability of the cost.
  • The second-stage value is the real infimum of f′yf'yf′y over the feasible set. It equals 000 on an infeasible or unbounded-below second stage, so each theorem assumes attainment of the minimum where it is evaluated (the paper's "value of yyy that minimizes CCC"). The goal assumes attainment for almost every outcome at every x≥0x\ge0x≥0.
  • E(C)E(C)E(C) is the Bochner integral, which is 000 for a non-integrable integrand, so the goal assumes integrability of C(x,⋅)C(x,\cdot)C(x,⋅) at every x≥0x\ge0x≥0 (the paper's "mean value E(C)E(C)E(C)").
  • Convexity is claimed on {x:x≥0}\{x : x\ge0\}{x:x≥0}, the paper's domain, not on all of Rn\mathbb R^nRn. Theorem 3 is stated for fixed non-negative xxx (the model's first-stage domain) and on every convex set of data on which the minimum is attained, since the paper names no domain.
  • A formalization in which the value is an unconstrained infimum without attainment, or the expectation is taken without integrability, is trivially convex on the region where the default values apply and does not state Beale's theorem; such variants are ruled out.
  • Reusable beyond this mission: basic facts on the optimal value of a parametric linear program in its right-hand side and cost data, and convexity of integrals of pointwise-convex integrands. Proofs of the milestones and of the goal, and alternative formulations in extended reals, are welcome.

Selected references

  • E. M. L. Beale, On Minimizing a Convex Function Subject to Linear Inequalities, Journal of the Royal Statistical Society, Series B 17(2):173–184, 1955. https://doi.org/10.1111/j.2517-6161.1955.tb00191.x
  • G. B. Dantzig, Linear Programming under Uncertainty, Management Science 1(3–4):197–206, 1955. https://doi.org/10.1287/mnsc.1.3-4.197
  • F. L. Hitchcock, The Distribution of a Product from Several Sources to Numerous Localities, Journal of Mathematics and Physics 20:224–230, 1941. https://doi.org/10.1002/sapm1941201224
  • D. W. Walkup and R. J.-B. Wets, Stochastic Programs with Recourse, SIAM Journal on Applied Mathematics 15(5):1299–1314, 1967. https://doi.org/10.1137/0115113
  • R. J.-B. Wets, Stochastic Programs with Fixed Recourse: The Equivalent Deterministic Program, SIAM Review 16(3):309–339, 1974. https://doi.org/10.1137/1016053
  • J. R. Birge and F. Louveaux, Introduction to Stochastic Programming, 2nd ed., Springer, 2011. https://doi.org/10.1007/978-1-4614-0237-4
6 thms3 active usersReviewed
Convex OptimizationOperations Research·Captain: mikedeng1

On Minimizing a Convex Function Subject to Linear Inequalities I: Beale's Simplex Method for a Convex Quadratic Function TerminatesResearch Paper

Motivation

Quadratic programming, the minimization of a convex quadratic function subject to linear constraints, is the simplest nonlinear extension of linear programming. It arises in least-squares estimation with sign constraints, in portfolio selection, and as the subproblem solved at each iteration of Newton-type methods for general smooth convex programs. E. M. L. Beale's 1955 paper (DOI 10.1111/j.2517-6161.1955.tb00191.x) gave one of the first finite algorithms for it by extending Dantzig's simplex method: the method keeps the simplex tableau and adds free variables, linear functions of the original variables with no sign restriction, along which the quadratic stops decreasing.

Timeline:

  • 1951: Dantzig publishes the simplex method for linear programming.
  • 1952: Charnes introduces ε-perturbations to resolve degeneracy in the simplex method.
  • 1955: Beale extends the simplex method to convex quadratic objectives and proves that the iteration terminates (§3 of the paper; the result formalized here).
  • 1959: Beale's "On quadratic programming" (Naval Research Logistics Quarterly 6) develops the method further; Wolfe's simplex method for quadratic programming (Econometrica 27) appears the same year.

Setting

There are nnn restricted variables xj≥0x_j \ge 0xj​≥0 satisfying mmm linearly independent linear equations, and a convex quadratic objective CCC. The iteration keeps N=n−mN = n - mN=n−m nonbasic variables z1,…,zNz_1, \dots, z_Nz1​,…,zN​, each either a restricted variable or a free variable, and writes every restricted variable as an affine function of them:

xh=ah0+∑l=1Nahlzl.(2.3)x_h = a_{h0} + \sum_{l=1}^{N} a_{hl} z_l. \qquad (2.3)xh​=ah0​+l=1∑N​ahl​zl​.(2.3)

A restricted variable that is not nonbasic is basic. The associated solution sets every zl=0z_l = 0zl​=0, so xh=ah0x_h = a_{h0}xh​=ah0​. The objective is written as

C=∑k=0N∑l=0Ncklzkzl,z0=1,(3.1)C = \sum_{k=0}^{N} \sum_{l=0}^{N} c_{kl} z_k z_l, \qquad z_0 = 1, \qquad (3.1)C=k=0∑N​l=0∑N​ckl​zk​zl​,z0​=1,(3.1)

with (ckl)(c_{kl})(ckl​) symmetric. Thus c00c_{00}c00​ is the value of CCC at the associated solution and 2ck02c_{k0}2ck0​ is its linear coefficient in zkz_kzk​. The number of nonbasic free variables is sss.

One step chooses a nonbasic zpz_pzp​ that can profitably be altered: a free one with cp0≠0c_{p0} \ne 0cp0​=0 if there is one, otherwise a restricted one with cp0<0c_{p0} < 0cp0​<0. It orients zpz_pzp​ so that it is to be increased, and increases it from 000. It stops at the first of two events. Either a basic variable xqx_qxq​ reaches 000 (the ratio test (2.4)), and then xqx_qxq​ becomes nonbasic in place of zpz_pzp​. Or CCC stops decreasing where the free variable ur=cp0+∑lcplzlu_r = c_{p0} + \sum_l c_{pl} z_lur​=cp0​+∑l​cpl​zl​ vanishes (3.2), and then uru_rur​ becomes nonbasic in place of zpz_pzp​. The coefficients are then transformed by substituting for zpz_pzp​ (eqs. (3.4)–(3.6)). CCC is in standard form when it has no linear term in any free variable.

Formalization targets

Goal: the iteration terminates

From a tableau with symmetric (ckl)(c_{kl})(ckl​), positive semidefinite quadratic block (ckl)k,l≥1(c_{kl})_{k,l \ge 1}(ckl​)k,l≥1​ and consistent labels, there is no infinite run

T0→T1→T2→⋯T_0 \to T_1 \to T_2 \to \cdotsT0​→T1​→T2​→⋯

of steps along which every basic variable stays strictly positive in the associated solution. No bound on the number of steps is claimed, as in the paper.

Milestones

  1. Eq. (3.7): the closed form of the transformed matrix, its symmetry, and the invariance ∑cklzkzl=∑ckl′′zk′zl′\sum c_{kl} z_k z_l = \sum c''_{kl} z'_k z'_l∑ckl​zk​zl​=∑ckl′′​zk′​zl′​.
  2. Lemma 1: when a free variable enters, its row and column vanish off the diagonal, the index 000 included.
  3. Lemma 2: a slot whose row and column vanish off the diagonal keeps this property when another free variable enters.
  4. The optimality criterion (p. 175): if no nonbasic variable can profitably be altered and CCC is convex, then c00c_{00}c00​ is the minimum over the feasible region.
  5. In standard form, c00≤C(z)c_{00} \le C(z)c00​≤C(z) for every zzz with the restricted nonbasic variables at 000.
  6. CCC decreases at every step: c00′<c00c'_{00} < c_{00}c00′​<c00​.
  7. A finite run never returns to a standard form with the same set of restricted nonbasic variables.
  8. If CCC is not in standard form and s=s0s = s_0s=s0​, then within s0s_0s0​ steps either standard form is reached or sss drops, and sss never exceeds s0s_0s0​ on the way.

Significance

The theorem makes Beale's method an algorithm: a finite procedure that ends either at an optimal tableau (milestone 4) or with a ray along which CCC decreases without bound. This finiteness is what later active-set methods for quadratic programming inherit.

The result was proved in 1955. What remains is to formalize it: a machine-checked account of a simplex-type method whose state includes variables that are created during the run and later discarded. Mathlib has no simplex-type algorithm for quadratic programming, and no machine-checked proof of this theorem is known. The pivot algebra (3.4)–(3.7) and the tableau model are reusable for other pivoting methods for quadratic programs.

Difficulty

The argument for linear programming does not carry over. There, the objective strictly decreases and a basis is a subset of a finite set of columns, so no basis repeats. Here each step may create a new free variable, and nothing bounds the number of distinct free variables that can occur. Tableaux are therefore not drawn from a finite set, and a strictly decreasing objective alone does not give termination. The paper states this itself: "there is no obvious limit to the number of free variables that may be involved". The difficulty is to bound the number of steps between returns to a well-behaved tableau, and this depends both on the rule that free variables are chosen first and on how the coefficient matrix evolves under repeated pivots.

Formalization scope

  • Representation. The nonbasic variables occupy fixed slots Fin (N+1). Slot 0 is z0=1z_0 = 1z0​=1; the nonbasic slot k : Fin N is index k.succ. A pivot stores the new nonbasic variable in the slot of the variable it replaces, so the paper's index qqq in (3.4)–(3.7) is that slot. The tableau holds the labels (restricted xjx_jxj​ or free), the rows of all nnn restricted variables (a nonbasic one has the unit row), and (ckl)(c_{kl})(ckl​). Free variables carry no row, as in the paper.
  • The pivot. pivotC is computed literally from (3.5) and then (3.6). Rows are transformed by the same substitution, as the paper states.
  • The step. The step is a relation. It allows any profitable choice of zpz_pzp​ subject to the free-first rule, and at a tie either outcome. No pricing rule is fixed, since the paper fixes none.
  • Convexity. Convexity of CCC is the symmetry of (ckl)(c_{kl})(ckl​) plus positive semidefiniteness of the block (ckl)k,l≥1(c_{kl})_{k,l \ge 1}(ckl​)k,l≥1​, assumed on the initial tableau.
  • Added hypothesis. The one hypothesis not on the page is that every basic restricted variable is strictly positive in the associated solution of every tableau of the run. It replaces Charnes's ε-perturbations, by which the paper ensures "the ah0a_{h0}ah0​ are always positive, and not zero". Positivity is required of basic variables only; nonbasic variables are 000 in the associated solution.
  • Out of scope. The link to the original equations (2.1) and phase 1 (artificial variables, the M-method) are not formalized: the iteration starts from a tableau already in the form (2.3).
  • Ruling out a trivial goal. A step relation that never fires, or a positivity hypothesis that no tableau can meet after a step, would make the goal trivially true. A sorry-free check exhibits a convex instance with consistent labels, a step, and positive basic variables before and after it.

Contributions are welcome on every milestone. The algebraic milestones 1–3 are self-contained.

Selected references

  • E. M. L. Beale, On Minimizing a Convex Function Subject to Linear Inequalities, Journal of the Royal Statistical Society, Series B 17(2):173–184, 1955. https://doi.org/10.1111/j.2517-6161.1955.tb00191.x
  • A. Charnes, Optimality and Degeneracy in Linear Programming, Econometrica 20(2):160–170, 1952. https://doi.org/10.2307/1907845
  • G. B. Dantzig, Maximization of a Linear Function of Variables Subject to Linear Inequalities, in T. C. Koopmans (ed.), Activity Analysis of Production and Allocation, Wiley, 1951, pp. 339–347.
  • E. M. L. Beale, On Quadratic Programming, Naval Research Logistics Quarterly 6(3):227–243, 1959. https://doi.org/10.1002/nav.3800060305
  • P. Wolfe, The Simplex Method for Quadratic Programming, Econometrica 27(3):382–398, 1959. https://doi.org/10.2307/1909468
12 thms3 active usersReviewed
🏆Completed
Algorithmic Game TheoryOperations ResearchProbability·Captain: mikedeng1

Optimal Pricing of Seasonal Products in the Presence of Forward-Looking Consumers 2: A Threshold Nash Equilibrium under Announced Fixed-Discount PricingResearch Paper

Motivation

Retailers of fashion and seasonal goods sell at a premium price early in the season and mark down later. When customers anticipate the markdown, some of them wait, and the seller's pricing problem becomes a game between the seller and a population of forward-looking (strategic) customers. Aviv and Pazgal (MSOM 10(3), 2008) study this game in a model with limited inventory, stochastic arrivals and valuations that decline over the season, under two classes of seller policies: contingent pricing, where the discount depends on the inventory left, and announced fixed-discount pricing, where the seller commits to both prices upfront. Their numerical study (§7.3) compares the two classes and finds that precommitment can raise expected revenue by up to about 8%.

That comparison needs, for every announced price path, the customers' equilibrium response. Theorem 2 of the paper (p. 348) supplies it: a threshold purchasing policy, pinned down by a scalar fixed-point equation for the probability that a waiting customer is served. This mission formalizes Theorem 2. A companion mission of the same series formalizes Theorem 1, the contingent-pricing counterpart.

Setting

A seller has Q≥1Q \ge 1Q≥1 units to sell over a season [0,H][0, H][0,H], split at a fixed time TTT with 0<T≤H0 < T \le H0<T≤H. Customers arrive by a Poisson process with rate λ>0\lambda > 0λ>0. Customer jjj has a base valuation VjV_jVj​ drawn from a continuous distribution FFF (tail Fˉ=1−F\bar F = 1 - FFˉ=1−F), and at time ttt values the product at Vj(t)=Vje−αtV_j(t) = V_j e^{-\alpha t}Vj​(t)=Vj​e−αt, where the decline factor α≥0\alpha \ge 0α≥0 is common to all customers.

Under an announced price path the seller commits to a premium price p1p_1p1​ on [0,T)[0, T)[0,T) and a discount price p2≤p1p_2 \le p_1p2​≤p1​ from TTT on; p2p_2p2​ does not depend on the remaining inventory. Customers know the initial inventory but not the current one.

A customer arriving at t<Tt < Tt<T buys immediately if and only if (i) the current surplus V(t)−p1V(t) - p_1V(t)−p1​ is nonnegative and (ii) it is at least the expected surplus of waiting,

ω⋅max⁡{V(T)−p2,0},\omega\cdot\max\{V(T) - p_2, 0\},ω⋅max{V(T)−p2​,0},

where ω\omegaω is the probability that a unit will be allocated to the customer at time TTT. Units left at TTT are rationed at random among the customers who request one.

For a threshold function ψ\psiψ on [0,T)[0, T)[0,T) the paper defines three segment rates: ΛI(ψ)\Lambda_I(\psi)ΛI​(ψ), the expected number of customers who buy at p1p_1p1​; ΛS(ψ,p1,p2)\Lambda_S(\psi, p_1, p_2)ΛS​(ψ,p1​,p2​), those who could buy at p1p_1p1​ but wait and want to buy at p2p_2p2​; and ΛW(p1,p2)\Lambda_W(p_1, p_2)ΛW​(p1​,p2​), those whose valuation was below p1p_1p1​ and who want to buy at p2p_2p2​. Each is λ\lambdaλ times an integral over [0,T][0, T][0,T] of Fˉ\bar FFˉ at scaled prices (p. 345). With P(x∣Λ)P(x \mid \Lambda)P(x∣Λ) the Poisson probabilities, the allocation probability of qqq units is

A(q∣Λ)=∑y=0∞qmax⁡{1+y,q} P(y∣Λ).A(q \mid \Lambda) = \sum_{y=0}^{\infty} \frac{q}{\max\{1+y, q\}}\,P(y \mid \Lambda).A(q∣Λ)=y=0∑∞​max{1+y,q}q​P(y∣Λ).

Formalization targets

Goal: Theorem 2 (p. 348)

For w∈[0,1]w \in [0,1]w∈[0,1] let

ψA(t)=max⁡{p1,p1−wp21−we−α(T−t)},0≤t<T,(7)\psi_A(t) = \max\left\{p_1, \frac{p_1 - wp_2}{1 - we^{-\alpha(T-t)}}\right\},\qquad 0 \le t < T, \tag{7}ψA​(t)=max{p1​,1−we−α(T−t)p1​−wp2​​},0≤t<T,(7)

and suppose www solves

w=∑x=0Q−1P(x∣ΛI(ψA))⋅A(Q−x∣ΛS(ψA,p1,p2)+ΛW(p1,p2)).(8)w = \sum_{x=0}^{Q-1} P\big(x \mid \Lambda_I(\psi_A)\big)\cdot A\big(Q-x \mid \Lambda_S(\psi_A, p_1, p_2) + \Lambda_W(p_1, p_2)\big). \tag{8}w=x=0∑Q−1​P(x∣ΛI​(ψA​))⋅A(Q−x∣ΛS​(ψA​,p1​,p2​)+ΛW​(p1​,p2​)).(8)

Then, when all other customers use ψA\psi_AψA​ (so that a waiting customer is served with the probability on the right of (8)), every customer arriving at t∈[0,T)t \in [0, T)t∈[0,T) buys immediately if and only if V(t)≥ψA(t)V(t) \ge \psi_A(t)V(t)≥ψA​(t): the symmetric threshold profile is a Nash equilibrium.

Milestones: the two cases of the proof (p. 358)

  1. If e−α(T−t)≤p2/p1e^{-\alpha(T-t)} \le p_2/p_1e−α(T−t)≤p2​/p1​, the threshold is p1p_1p1​.
  2. If e−α(T−t)>p2/p1e^{-\alpha(T-t)} > p_2/p_1e−α(T−t)>p2​/p1​, the threshold is (p1−wp2)/(1−we−α(T−t))≥p1(p_1 - wp_2)/(1 - we^{-\alpha(T-t)}) \ge p_1(p1​−wp2​)/(1−we−α(T−t))≥p1​.

Significance

Theorem 2 reduces the customers' equilibrium under an announced path to a single scalar www. Everything downstream in §5 and §7 rests on it: the seller's expected revenue πA/S(p1,p2)\pi_{A/S}(p_1, p_2)πA/S​(p1​,p2​) (p. 348) is written in terms of ψA\psi_AψA​, the seller's optimal announced path maximizes it, and the comparison between announced and contingent pricing uses the resulting value πA/S∗\pi^*_{A/S}πA/S∗​. The theorem also explains the qualitative prediction of the model: the threshold exceeds p1p_1p1​ exactly when the announced discount is deep relative to the decline of valuations, and it rises with the perceived availability www.

The result is proved in the paper; to the best of our search it has no machine-checked proof. A formal development contributes the model objects (segment rates for threshold policies, the allocation probability for random rationing among Poisson requesters) in a form reusable by the rest of the series and by other strategic-customer pricing models, and a checked proof of the equilibrium property. The existence of a solution to (8) is not proved in the paper and is a natural further target.

Difficulty

The best-response part of the argument is elementary once the availability is known. The substance of the statement lies in the availability itself: the probability that a waiting customer is served is not a free parameter but the one generated, through (8), by the other customers' use of the same threshold. A formalization must connect the segment rates, the Poisson counts and random rationing into one expression and keep the fixed-point coupling between www and ψA\psi_AψA​ intact; dropping it turns the theorem into a one-line inequality about an arbitrary www. The division by 1−we−α(T−t)1 - we^{-\alpha(T-t)}1−we−α(T−t) also degenerates when w=1w = 1w=1 and α=0\alpha = 0α=0, and has to be excluded explicitly.

Formalization scope

The Lean development lives in namespace SeasonalPricing.Announced. Conventions:

  • Time is real; base valuations have law μ : Measure ℝ with IsProbabilityMeasure μ, FFF = ProbabilityTheory.cdf μ, and continuity of FFF (the paper's "continuous distribution") is a hypothesis of the goal. No support condition on [0,∞)[0,\infty)[0,∞) is imposed; the statement quantifies over every real base valuation VVV.
  • ΛI,ΛS,ΛW\Lambda_I, \Lambda_S, \Lambda_WΛI​,ΛS​,ΛW​ are interval integrals over [0,T][0, T][0,T] exactly as printed. P(x∣Λ)=e−ΛΛx/x!P(x \mid \Lambda) = e^{-\Lambda}\Lambda^x/x!P(x∣Λ)=e−ΛΛx/x! is written out; A(q∣Λ)A(q\mid\Lambda)A(q∣Λ) is the infinite series (tsum) as printed, not its closed form.
  • availability is the right-hand side of (8), with ψA\psi_AψA​ built from www by (7).

Readings of the paper's informal words:

  • "Nash equilibrium" is read as the best-response property the paper's proof checks: against the availability generated by (8), the immediate-purchase rule of p. 344 coincides with the threshold ψA\psi_AψA​ at every t∈[0,T)t \in [0, T)t∈[0,T) and every valuation. The paper defines no strategy space beyond threshold rules.
  • "www is a solution to (8)": the theorem is conditional on a solution; its existence is neither assumed elsewhere nor claimed. The conditional statement has content only when (8) has a solution, which the paper does not prove.
  • www as a likelihood: 0≤w≤10 \le w \le 10≤w≤1 is a hypothesis (it also follows from (8)).
  • Added hypothesis: α>0\alpha > 0α>0 or w<1w < 1w<1, which keeps 1−we−α(T−t)>01 - we^{-\alpha(T-t)} > 01−we−α(T−t)>0 for t<Tt < Tt<T; the paper's formula is undefined when it fails. In the milestones the same condition appears as we−α(T−t)<1we^{-\alpha(T-t)} < 1we−α(T−t)<1, and 0<p10 < p_10<p1​ is added so that p2/p1p_2/p_1p2​/p1​ is meaningful.
  • The rule on [T,H][T, H][T,H] (buy at TTT iff V(T)>p2V(T) > p_2V(T)>p2​) is part of the model and is not restated; HHH does not enter the statements.

A formalization in which www is an arbitrary number in [0,1][0,1][0,1], not tied to (8), is ruled out: it is the best-response lemma alone, not Theorem 2. Contributions welcome: proofs of the two milestones and the goal; lemmas such as 0≤A(q∣Λ)≤10 \le A(q\mid\Lambda) \le 10≤A(q∣Λ)≤1 and summability of its series; the closed form of A(q∣Λ)A(q \mid \Lambda)A(q∣Λ) printed on p. 346; and an existence result for (8).

Selected references

  • Y. Aviv and A. Pazgal, Optimal Pricing of Seasonal Products in the Presence of Forward-Looking Consumers, Manufacturing & Service Operations Management 10(3):339–359, 2008. https://doi.org/10.1287/msom.1070.0183
  • G. Gallego and G. van Ryzin, Optimal Dynamic Pricing of Inventories with Stochastic Demand over Finite Horizons, Management Science 40(8):999–1020, 1994. https://doi.org/10.1287/mnsc.40.8.999
  • X. Su, Intertemporal Pricing with Strategic Customer Behavior, Management Science 53(5):726–741, 2007. https://doi.org/10.1287/mnsc.1060.0667
5 thms3 active usersReviewed
🏆Completed
Machine LearningProbabilityStatistics·Captain: mikedeng1

Simultaneous Analysis of Lasso and Dantzig Selector III: A Sparsity Oracle Inequality for the LassoResearch Paper

Motivation

In high-dimensional regression the number of candidate predictors MMM can far exceed the number of observations nnn. A regression function can then be estimated only if it is well approximated by a combination of a few elements of a large dictionary. The Lasso is the most widely used estimator in this regime. The question this mission formalizes is how well the Lasso predicts when the truth is not assumed to be sparse, or even to lie in the span of the dictionary.

A sparsity oracle inequality answers it. It bounds the prediction error of the estimator by the error of the best sparse approximation of the truth, which only an oracle knowing the truth could compute, plus a remainder proportional to the sparsity of that approximation times log⁡M/n\log M/nlogM/n. Bickel, Ritov and Tsybakov (arXiv:0801.1095, Ann. Statist. 37(4), 2009) proved such an inequality for the Lasso under their restricted eigenvalue (RE) condition. Earlier oracle inequalities for Lasso-type estimators in fixed design (Bunea, Tsybakov and Wegkamp, 2006–2007) required the Gram matrix to be positive definite or to satisfy a mutual-coherence condition. The RE condition is weaker and allows M≫nM\gg nM≫n, and it is now the standard hypothesis in this literature.

Setting

A dictionary f1,…,fMf_1,\dots,f_Mf1​,…,fM​ is evaluated at fixed points Z1,…,ZnZ_1,\dots,Z_nZ1​,…,Zn​. This gives the design matrix X=(fj(Zi))∈Rn×MX=(f_j(Z_i))\in\mathbb R^{n\times M}X=(fj​(Zi​))∈Rn×M and, for an unknown regression function fff, the vector f=(f(Z1),…,f(Zn))⊤f=(f(Z_1),\dots,f(Z_n))^\topf=(f(Z1​),…,f(Zn​))⊤. The observations are

y=f+W,W1,…,Wn independent N(0,σ2), σ>0.y=f+W,\qquad W_1,\dots,W_n\ \text{independent}\ \mathcal N(0,\sigma^2),\ \sigma>0 .y=f+W,W1​,…,Wn​ independent N(0,σ2), σ>0.

Nothing is assumed about fff. For v∈Rnv\in\mathbb R^nv∈Rn the empirical norm is ∥v∥n=(1n∑ivi2)1/2\|v\|_n=(\frac1n\sum_iv_i^2)^{1/2}∥v∥n​=(n1​∑i​vi2​)1/2, and for β∈RM\beta\in\mathbb R^Mβ∈RM we write fβ=Xβf_\beta=X\betafβ​=Xβ. The column norms ∥fj∥n\|f_j\|_n∥fj​∥n​ are assumed nonzero, with fmax⁡=max⁡j∥fj∥nf_{\max}=\max_j\|f_j\|_nfmax​=maxj​∥fj​∥n​ and fmin⁡=min⁡j∥fj∥nf_{\min}=\min_j\|f_j\|_nfmin​=minj​∥fj​∥n​. The support of β\betaβ is J(β)={j:βj≠0}J(\beta)=\{j:\beta_j\neq0\}J(β)={j:βj​=0} and its sparsity is M(β)=∣J(β)∣\mathcal M(\beta)=|J(\beta)|M(β)=∣J(β)∣.

The Lasso β^L\hat\beta_Lβ^​L​ is any minimiser of

1n∑i=1n(yi−(Xβ)i)2+2r∑j=1M∥fj∥n∣βj∣,r=Aσlog⁡Mn, A>22,\frac1n\sum_{i=1}^n\big(y_i-(X\beta)_i\big)^2+2r\sum_{j=1}^M\|f_j\|_n|\beta_j|,\qquad r=A\sigma\sqrt{\frac{\log M}{n}},\ A>2\sqrt2,n1​i=1∑n​(yi​−(Xβ)i​)2+2rj=1∑M​∥fj​∥n​∣βj​∣,r=AσnlogM​​, A>22​,

and f^L=Xβ^L\hat f_L=X\hat\beta_Lf^​L​=Xβ^​L​.

Assumption RE(s,c0)(s,c_0)(s,c0​) holds with constant κ>0\kappa>0κ>0 if, for every J0⊆{1,…,M}J_0\subseteq\{1,\dots,M\}J0​⊆{1,…,M} with ∣J0∣≤s|J_0|\le s∣J0​∣≤s and every δ≠0\delta\neq0δ=0 with ∣δJ0c∣1≤c0∣δJ0∣1|\delta_{J_0^c}|_1\le c_0|\delta_{J_0}|_1∣δJ0c​​∣1​≤c0​∣δJ0​​∣1​,

κn ∣δJ0∣2≤∣Xδ∣2.\kappa\sqrt n\,|\delta_{J_0}|_2\le|X\delta|_2 .κn​∣δJ0​​∣2​≤∣Xδ∣2​.

The paper's κ(s,c0)\kappa(s,c_0)κ(s,c0​) is the largest such constant.

Formalization targets

Goal: Theorem 6.1

Fix ε>0\varepsilon>0ε>0, n≥1n\ge1n≥1, M≥2M\ge2M≥2, 1≤s≤M1\le s\le M1≤s≤M, and let RE(s,(3+4/ε)fmax⁡/fmin⁡)(s,(3+4/\varepsilon)f_{\max}/f_{\min})(s,(3+4/ε)fmax​/fmin​) hold with constant κ\kappaκ. With probability at least 1−M1−A2/81-M^{1-A^2/8}1−M1−A2/8, every Lasso solution satisfies, simultaneously for all β\betaβ with M(β)≤s\mathcal M(\beta)\le sM(β)≤s,

∥f^L−f∥n2≤(1+ε){∥fβ−f∥n2+C(ε)fmax⁡2A2σ2κ2 M(β)log⁡Mn},C(ε)=4(2+ε)2ε(1+ε).\|\hat f_L-f\|_n^2\le(1+\varepsilon)\Big\{\|f_\beta-f\|_n^2+C(\varepsilon)\frac{f_{\max}^2A^2\sigma^2}{\kappa^2}\,\frac{\mathcal M(\beta)\log M}{n}\Big\},\qquad C(\varepsilon)=\frac{4(2+\varepsilon)^2}{\varepsilon(1+\varepsilon)} .∥f^​L​−f∥n2​≤(1+ε){∥fβ​−f∥n2​+C(ε)κ2fmax2​A2σ2​nM(β)logM​},C(ε)=ε(1+ε)4(2+ε)2​.

Milestones

  1. (B.4): the noise event A=⋂j{2∣Vj∣≤r∥fj∥n}\mathcal A=\bigcap_j\{2|V_j|\le r\|f_j\|_n\}A=⋂j​{2∣Vj​∣≤r∥fj​∥n​}, with Vj=n−1∑iXijWiV_j=n^{-1}\sum_iX_{ij}W_iVj​=n−1∑i​Xij​Wi​, satisfies P(Ac)≤M1−A2/8P(\mathcal A^c)\le M^{1-A^2/8}P(Ac)≤M1−A2/8.
  2. (B.1) on A\mathcal AA: for every Lasso solution and every β\betaβ,
∥f^L−f∥n2+r∑j∥fj∥n∣β^j−βj∣≤∥fβ−f∥n2+4r∑j∈J(β)∥fj∥n∣β^j−βj∣.\|\hat f_L-f\|_n^2+r\sum_j\|f_j\|_n|\hat\beta_j-\beta_j|\le\|f_\beta-f\|_n^2+4r\sum_{j\in J(\beta)}\|f_j\|_n|\hat\beta_j-\beta_j| .∥f^​L​−f∥n2​+rj∑​∥fj​∥n​∣β^​j​−βj​∣≤∥fβ​−f∥n2​+4rj∈J(β)∑​∥fj​∥n​∣β^​j​−βj​∣.
  1. Lemma B.1: the same inequality with probability at least 1−M1−A2/81-M^{1-A^2/8}1−M1−A2/8.
  2. Cone step: in the case ε∥fβ−f∥n2<4r∑J(β)∥fj∥n∣β^j−βj∣\varepsilon\|f_\beta-f\|_n^2<4r\sum_{J(\beta)}\|f_j\|_n|\hat\beta_j-\beta_j|ε∥fβ​−f∥n2​<4r∑J(β)​∥fj​∥n​∣β^​j​−βj​∣, the difference β^L−β\hat\beta_L-\betaβ^​L​−β lies in the cone with constant (3+4/ε)fmax⁡/fmin⁡(3+4/\varepsilon)f_{\max}/f_{\min}(3+4/ε)fmax​/fmin​ at J(β)J(\beta)J(β).
  3. Inequality before decoupling: ∥f^L−f∥n2≤∥fβ−f∥n2+4rfmax⁡κ−1M(β) (∥f^L−f∥n+∥fβ−f∥n)\|\hat f_L-f\|_n^2\le\|f_\beta-f\|_n^2+4rf_{\max}\kappa^{-1}\sqrt{\mathcal M(\beta)}\,(\|\hat f_L-f\|_n+\|f_\beta-f\|_n)∥f^​L​−f∥n2​≤∥fβ​−f∥n2​+4rfmax​κ−1M(β)​(∥f^​L​−f∥n​+∥fβ​−f∥n​).
  4. Decoupled bound: ∥f^L−f∥n2≤b+1b−1∥fβ−f∥n2+8b2fmax⁡2(b−1)κ2r2M(β)\|\hat f_L-f\|_n^2\le\frac{b+1}{b-1}\|f_\beta-f\|_n^2+\frac{8b^2f_{\max}^2}{(b-1)\kappa^2}r^2\mathcal M(\beta)∥f^​L​−f∥n2​≤b−1b+1​∥fβ​−f∥n2​+(b−1)κ28b2fmax2​​r2M(β) for all b>1b>1b>1.
  5. Corollary 6.2: the same oracle inequality with γ\gammaγ in place of κ\kappaκ and no global RE assumption. The infimum runs over those β\betaβ with M(β)≤s\mathcal M(\beta)\le sM(β)≤s whose support alone satisfies the restricted eigenvalue inequality with constant γ\gammaγ.

Significance

The theorem says that, up to the factor 1+ε1+\varepsilon1+ε and a remainder of order M(β)log⁡M/n\mathcal M(\beta)\log M/nM(β)logM/n, the Lasso predicts as well as the best sss-sparse linear combination of the dictionary. This is the case even when fff is not sparse and not in the span of the dictionary. The remainder is the parametric rate for M(β)\mathcal M(\beta)M(β) parameters, inflated by log⁡M\log MlogM and by the ill-posedness factor fmax⁡2/κ2f_{\max}^2/\kappa^2fmax2​/κ2. Together with Theorem 5.1 of the same paper (mission II of this series), it shows that the Lasso and the Dantzig selector are within the same distance of the sparse oracle. The oracle inequality is used in aggregation, in model selection, and as a black box in later sparse-estimation papers.

The result is proved in the paper. It has not been formalized: at the time of writing, no Lasso oracle inequality and no probabilistic Lasso bound exist on Prove2Me or in Mathlib. What this mission contributes is a machine-checked proof of the paper's Theorem 6.1 with an explicit constant C(ε)C(\varepsilon)C(ε). The paper leaves C(ε)C(\varepsilon)C(ε) unspecified, and its proof fixes the value used here. The mission also formalizes the Gaussian-tail step (B.4) and the deterministic basic inequality (B.1), both of which are shared with the paper's other Lasso results.

Difficulty

There is no sparse truth, so the usual argument does not apply. That argument places the error β^L−β∗\hat\beta_L-\beta^*β^​L​−β∗ in the RE cone and reads off a rate. Here the competitor β\betaβ is arbitrary, and the approximation error ∥fβ−f∥n\|f_\beta-f\|_n∥fβ​−f∥n​ can dominate the penalty terms, in which case the error is not in the cone. The RE assumption can be used only where the error does lie in a cone, and the cone constant available there depends on ε\varepsilonε and on the column-norm ratio fmax⁡/fmin⁡f_{\max}/f_{\min}fmax​/fmin​, because the penalty is weighted while RE is stated for unweighted vectors. What RE then yields is an inequality quadratic in ∥f^L−f∥n\|\hat f_L-f\|_n∥f^​L​−f∥n​ with a cross term, not the (1+ε)(1+\varepsilon)(1+ε) form directly, and the constant C(ε)C(\varepsilon)C(ε) is determined by how that cross term is absorbed. On the probabilistic side, the whole argument must run on one event of probability at least 1−M1−A2/81-M^{1-A^2/8}1−M1−A2/8. That event may depend neither on β\betaβ nor on the choice of minimiser. The Lasso need not have a unique solution.

Formalization scope

  • The dictionary enters only through X∈Rn×MX\in\mathbb R^{n\times M}X∈Rn×M (Matrix (Fin n) (Fin M) ℝ) and the target only through f∈Rnf\in\mathbb R^nf∈Rn, which is arbitrary. The noise is a family W : Fin n → Ω → ℝ of measurable, independent random variables, each with law gaussianReal 0 σ², and σ>0\sigma>0σ>0.
  • The Lasso is an argmin predicate, and every statement is made for every minimiser. "With probability at least ppp" means a measurable event EEE with P(E)≥pP(E)\ge pP(E)≥p, chosen before the competitor β\betaβ and the minimiser.
  • RE is stated through a witness κ>0\kappa>0κ>0. Since κ(s,c0)\kappa(s,c_0)κ(s,c0​) is attained and every bound decreases in κ\kappaκ, this is equivalent to the paper's form, and it avoids a real infimum over an empty set.
  • The infimum over {β:M(β)≤s}\{\beta:\mathcal M(\beta)\le s\}{β:M(β)≤s} is written as "for every such β\betaβ". This is equivalent, because the set contains β=0\beta=0β=0 and the bracket is nonnegative.
  • Correction/strengthening. The printed theorem has an unspecified C(ε)>0C(\varepsilon)>0C(ε)>0. The goal instead uses the value C(ε)=4(2+ε)2/(ε(1+ε))C(\varepsilon)=4(2+\varepsilon)^2/(\varepsilon(1+\varepsilon))C(ε)=4(2+ε)2/(ε(1+ε)) that the proof yields with b=1+2/εb=1+2/\varepsilonb=1+2/ε, and this implies the printed statement. Corollary 6.2 uses the same explicit constant.
  • The standing assumptions of Section 2 (M≥2M\ge2M≥2 and every ∥fj∥n≠0\|f_j\|_n\neq0∥fj​∥n​=0) are hypotheses of every theorem.
  • Some formalizations would make the result trivial, and they are excluded here. The noise must be exactly i.i.d. N(0,σ2)\mathcal N(0,\sigma^2)N(0,σ2) with σ>0\sigma>0σ>0 and must enter only through y=f+Wy=f+Wy=f+W. The target fff must not be restricted to Xβ∗X\beta^*Xβ∗. The event must be measurable. The constant must depend on ε\varepsilonε alone.
  • A single definition file provides the empirical norms, fmax⁡f_{\max}fmax​, fmin⁡f_{\min}fmin​, support and sparsity, the weighted Lasso, RE and its single-set version (the family Λs,γ,c0\Lambda_{s,\gamma,c_0}Λs,γ,c0​​ of Corollary 6.2), the Gaussian noise model and the event A\mathcal AA. The same objects appear in the other missions of this series. Gaussian-tail and union-bound lemmas proved along the way are reusable, and contributions of such lemmas are welcome.

Selected references

  • P. J. Bickel, Y. Ritov, A. B. Tsybakov, Simultaneous analysis of Lasso and Dantzig selector, Ann. Statist. 37(4), 1705–1732, 2009. Cited version: arXiv:0801.1095v3; DOI 10.1214/08-AOS620.
  • F. Bunea, A. B. Tsybakov, M. H. Wegkamp, Sparsity oracle inequalities for the Lasso, Electron. J. Statist. 1, 169–194, 2007. DOI 10.1214/07-EJS008.
  • F. Bunea, A. B. Tsybakov, M. H. Wegkamp, Aggregation for Gaussian regression, Ann. Statist. 35(4), 1674–1697, 2007. DOI 10.1214/009053606000001587.
  • R. Tibshirani, Regression shrinkage and selection via the lasso, J. R. Stat. Soc. B 58(1), 267–288, 1996. DOI 10.1111/j.2517-6161.1996.tb02080.x.
9 thms3 active usersReviewed
🏆Completed
Linear algebraMachine LearningStatistics·Captain: mikedeng1

Simultaneous Analysis of Lasso and Dantzig Selector I: Sparse Eigenvalue and Correlation Conditions Imply the Restricted Eigenvalue ConditionResearch Paper

Motivation

In high-dimensional linear regression one observes y=Xβ∗+w∈Rny = X\beta^* + w \in \mathbb R^ny=Xβ∗+w∈Rn with a design matrix X∈Rn×MX\in\mathbb R^{n\times M}X∈Rn×M whose number of columns MMM may far exceed the sample size nnn. The two standard estimators of a sparse β∗\beta^*β∗, the Lasso (Tibshirani, 1996) and the Dantzig selector (Candès and Tao, 2007), both come with error bounds of order slog⁡M/ns\log M/nslogM/n for an sss-sparse β∗\beta^*β∗, but only under a condition on XXX: since XXX has a non-trivial kernel when M>nM>nM>n, some form of restricted invertibility is unavoidable.

Bickel, Ritov and Tsybakov (arXiv:0801.1095, Ann. Statist. 2009) introduced the restricted eigenvalue (RE) condition, which asks for invertibility of XXX only on a cone of approximately sparse vectors. It is weaker than the conditions used before it and has since become the default assumption in the sparse-estimation literature. Section 4 of the paper relates RE to the earlier conditions:

  • 2005–2007: Candès and Tao (arXiv:math/0506081) analyse the Dantzig selector under a uniform uncertainty principle involving restricted eigenvalues and restricted correlations of XXX; the condition ϕmin⁡(2s)>θs,2s\phi_{\min}(2s)>\theta_{s,2s}ϕmin​(2s)>θs,2s​ is Assumption 1 below with c0=1c_0=1c0​=1.
  • 2006–2009: Meinshausen and Yu (arXiv:math/0605584) analyse the Lasso under a lower bound on sparse eigenvalues of order slog⁡ns\log nslogn.
  • 2006: Donoho, Elad and Temlyakov (doi:10.1109/TIT.2005.860430) use mutual coherence for sparse recovery; 2007: Bunea, Tsybakov and Wegkamp (doi:10.1214/07-EJS008) use coherence-type conditions for the Lasso.
  • 2009: Bickel, Ritov and Tsybakov show (Lemma 4.1 and Section 4) that each of these conditions implies RE.

This mission formalizes those implications.

Setting

Fix integers n≥1n\ge1n≥1 and M≥2M\ge2M≥2 and a matrix X∈Rn×MX\in\mathbb R^{n\times M}X∈Rn×M with columns x1,…,xMx_1,\dots,x_Mx1​,…,xM​. The Gram matrix is Ψn=XTX/n\Psi_n = X^TX/nΨn​=XTX/n. For δ∈RM\delta\in\mathbb R^Mδ∈RM and J⊆{1,…,M}J\subseteq\{1,\dots,M\}J⊆{1,…,M}, δJ\delta_JδJ​ is the vector equal to δ\deltaδ on JJJ and 000 off JJJ; ∣⋅∣1|\cdot|_1∣⋅∣1​, ∣⋅∣2|\cdot|_2∣⋅∣2​ are the ℓ1\ell_1ℓ1​ and Euclidean norms; M(δ)\mathcal M(\delta)M(δ) is the number of non-zero coordinates of δ\deltaδ; J0cJ_0^cJ0c​ is the complement of J0J_0J0​.

The cone condition for J0J_0J0​ and c0>0c_0>0c0​>0 is

∣δJ0c∣1≤c0 ∣δJ0∣1.(4.1)|\delta_{J_0^c}|_1\le c_0\,|\delta_{J_0}|_1. \tag{4.1}∣δJ0c​​∣1​≤c0​∣δJ0​​∣1​.(4.1)

Assumption RE(s,c0)(s,c_0)(s,c0​) holds with constant κ>0\kappa>0κ>0 if ∣Xδ∣2≥κn ∣δJ0∣2|X\delta|_2\ge\kappa\sqrt n\,|\delta_{J_0}|_2∣Xδ∣2​≥κn​∣δJ0​​∣2​ for every J0J_0J0​ with ∣J0∣≤s|J_0|\le s∣J0​∣≤s and every δ≠0\delta\ne0δ=0 satisfying (4.1). For m≥sm\ge sm≥s, let J1J_1J1​ be a set of mmm indices outside J0J_0J0​ carrying the mmm largest ∣δj∣|\delta_j|∣δj​∣, and J01=J0∪J1J_{01}=J_0\cup J_1J01​=J0​∪J1​; Assumption RE(s,m,c0)(s,m,c_0)(s,m,c0​) replaces ∣δJ0∣2|\delta_{J_0}|_2∣δJ0​​∣2​ by ∣δJ01∣2|\delta_{J_{01}}|_2∣δJ01​​∣2​.

The restricted eigenvalues are ϕmin⁡(u)\phi_{\min}(u)ϕmin​(u) and ϕmax⁡(u)\phi_{\max}(u)ϕmax​(u), the minimum and maximum of xTΨnx/∣x∣22x^T\Psi_nx/|x|_2^2xTΨn​x/∣x∣22​ over xxx with 1≤M(x)≤u1\le\mathcal M(x)\le u1≤M(x)≤u. The restricted correlations θm1,m2\theta_{m_1,m_2}θm1​,m2​​ are the maximum of c1TXI1TXI2c2/(n∣c1∣2∣c2∣2)c_1^TX_{I_1}^TX_{I_2}c_2/(n|c_1|_2|c_2|_2)c1T​XI1​T​XI2​​c2​/(n∣c1​∣2​∣c2​∣2​) over disjoint index sets I1,I2I_1,I_2I1​,I2​ with ∣Ii∣≤mi|I_i|\le m_i∣Ii​∣≤mi​ and non-zero ci∈RIic_i\in\mathbb R^{I_i}ci​∈RIi​. Two constants are attached to them:

κ1(s,c0)=ϕmin⁡(2s)(1−c0θs,2sϕmin⁡(2s)),κ2(s,m,c0)=ϕmin⁡(s+m)(1−c0s ϕmax⁡(m)m ϕmin⁡(s+m)).\kappa_1(s,c_0)=\sqrt{\phi_{\min}(2s)}\Big(1-\frac{c_0\theta_{s,2s}}{\phi_{\min}(2s)}\Big),\qquad \kappa_2(s,m,c_0)=\sqrt{\phi_{\min}(s+m)}\Big(1-c_0\sqrt{\tfrac{s\,\phi_{\max}(m)}{m\,\phi_{\min}(s+m)}}\Big).κ1​(s,c0​)=ϕmin​(2s)​(1−ϕmin​(2s)c0​θs,2s​​),κ2​(s,m,c0​)=ϕmin​(s+m)​(1−c0​mϕmin​(s+m)sϕmax​(m)​​).

P01P_{01}P01​ is the orthogonal projector in Rn\mathbb R^nRn onto the span of the columns xjx_jxj​, j∈J01j\in J_{01}j∈J01​.

Formalization targets

Goal: Lemma 4.1 (ii)

For integers 1≤s≤M/21\le s\le M/21≤s≤M/2, m≥sm\ge sm≥s, s+m≤Ms+m\le Ms+m≤M and c0>0c_0>0c0​>0, if Assumption 2 m ϕmin⁡(s+m)>c02 s ϕmax⁡(m)m\,\phi_{\min}(s+m)>c_0^2\,s\,\phi_{\max}(m)mϕmin​(s+m)>c02​sϕmax​(m) holds, then κ2(s,m,c0)>0\kappa_2(s,m,c_0)>0κ2​(s,m,c0​)>0, RE(s,c0)(s,c_0)(s,c0​) and RE(s,m,c0)(s,m,c_0)(s,m,c0​) hold with constant κ2(s,m,c0)\kappa_2(s,m,c_0)κ2​(s,m,c0​), and for every J0J_0J0​ with ∣J0∣≤s|J_0|\le s∣J0​∣≤s and every δ\deltaδ satisfying (4.1)

1n∣P01Xδ∣2 ≥ κ2(s,m,c0) ∣δJ01∣2.\frac1{\sqrt n}|P_{01}X\delta|_2\ \ge\ \kappa_2(s,m,c_0)\,|\delta_{J_{01}}|_2 .n​1​∣P01​Xδ∣2​ ≥ κ2​(s,m,c0​)∣δJ01​​∣2​.

Assumption 2 involves no correlations, only extreme eigenvalues of small principal submatrices of Ψn\Psi_nΨn​.

Lemma 4.1 (i)

For 1≤s≤M/21\le s\le M/21≤s≤M/2 and c0>0c_0>0c0​>0, Assumption 1 ϕmin⁡(2s)>c0θs,2s\phi_{\min}(2s)>c_0\theta_{s,2s}ϕmin​(2s)>c0​θs,2s​ implies the same conclusions with m=sm=sm=s and constant κ1(s,c0)\kappa_1(s,c_0)κ1​(s,c0​).

Coherence-type conditions (Section 4)

For 1≤s≤M1\le s\le M1≤s≤M and c0>0c_0>0c0​>0, each of

ϕmin⁡(s)>2c0θs,1s,ϕmin⁡(s)>2c0θ1,1s,diag⁡Ψn=1 and θ1,1<1(1+2c0)s\phi_{\min}(s)>2c_0\theta_{s,1}\sqrt s,\qquad \phi_{\min}(s)>2c_0\theta_{1,1}s,\qquad \operatorname{diag}\Psi_n=1\ \text{and}\ \theta_{1,1}<\frac1{(1+2c_0)s}ϕmin​(s)>2c0​θs,1​s​,ϕmin​(s)>2c0​θ1,1​s,diagΨn​=1 and θ1,1​<(1+2c0​)s1​

(Assumptions 3, 4, 5) implies RE(s,c0)(s,c_0)(s,c0​), with the constants κ2=ϕmin⁡(s)−2c0θs,1s\kappa^2=\phi_{\min}(s)-2c_0\theta_{s,1}\sqrt sκ2=ϕmin​(s)−2c0​θs,1​s​, ϕmin⁡(s)−2c0θ1,1s\phi_{\min}(s)-2c_0\theta_{1,1}sϕmin​(s)−2c0​θ1,1​s and 1−(1+2c0)θ1,1s1-(1+2c_0)\theta_{1,1}s1−(1+2c0​)θ1,1​s respectively.

The milestones are the steps of the proof in Appendix A — the projection inequality (A.1), the block bound (A.2), the shelling bound (A.3), the Candès–Tao correlation bound used for part (i) — followed by part (i) and the three coherence-type implications.

Significance

RE(s,c0)(s,c_0)(s,c0​) with c0=3c_0=3c0​=3 and c0=1c_0=1c0​=1 is the hypothesis of the paper's prediction and ℓ1\ell_1ℓ1​ bounds for the Lasso and the Dantzig selector (Theorems 5.1, 6.1, 7.1, 7.2), and RE(s,m,c0)(s,m,c_0)(s,m,c0​) is the hypothesis of its ℓp\ell_pℓp​ bounds. Assumptions 1–5 are stated through quantities that are standard in compressed sensing and random matrix theory, so known bounds for ϕmin⁡\phi_{\min}ϕmin​, ϕmax⁡\phi_{\max}ϕmax​ and θ\thetaθ of random designs transfer, through this mission's theorems, to every result stated under RE. Lemma 4.1 also shows that RE is weaker than the Candès–Tao condition used for the Dantzig selector.

The results are proved in the paper; parts of Lemma 4.1's proof (the correlation bound for part (i)) are cited from Candès and Tao without proof. None of these results is formalized: the platform has pairwise-incoherence and restricted-nullspace statements from Wainwright's textbook (a different conclusion and normalization) and restricted isometry definitions, but neither restricted eigenvalues ϕmin⁡(u),ϕmax⁡(u)\phi_{\min}(u),\phi_{\max}(u)ϕmin​(u),ϕmax​(u), restricted correlations θm1,m2\theta_{m_1,m_2}θm1​,m2​​, nor the RE condition in this form.

Difficulty

The naive attempt to bound ∣Xδ∣2|X\delta|_2∣Xδ∣2​ from below splits δ=δJ0+δJ0c\delta=\delta_{J_0}+\delta_{J_0^c}δ=δJ0​​+δJ0c​​ and applies an eigenvalue bound to each part. This fails: δJ0c\delta_{J_0^c}δJ0c​​ can have up to M−sM-sM−s non-zero coordinates, and no condition on sss- or 2s2s2s-sparse submatrices controls ∣XδJ0c∣2|X\delta_{J_0^c}|_2∣XδJ0c​​∣2​ directly. The cone condition bounds only the ℓ1\ell_1ℓ1​ norm of δJ0c\delta_{J_0^c}δJ0c​​, while eigenvalue conditions speak about ℓ2\ell_2ℓ2​ norms of sparse vectors; bridging the two with the right constant s/m\sqrt{s/m}s/m​, and keeping track of how the leading block J01J_{01}J01​ interacts with the rest through the projector P01P_{01}P01​, is where the work lies. For part (i), the interaction between disjoint sparse blocks has to be controlled by θs,2s\theta_{s,2s}θs,2s​ rather than by ϕmax⁡\phi_{\max}ϕmax​.

Formalization scope

  • Representation. XXX is Matrix (Fin n) (Fin M) ℝ; vectors are Fin M → ℝ and Fin n → ℝ; ∣Xδ∣2=(∑i(Xδ)i2)1/2|X\delta|_2=(\sum_i (X\delta)_i^2)^{1/2}∣Xδ∣2​=(∑i​(Xδ)i2​)1/2. The projector P01P_{01}P01​ is Mathlib's orthogonal projection on EuclideanSpace ℝ (Fin n) onto the span of the columns indexed by J01J_{01}J01​.
  • RE through a witness. RE X s c0 κ asserts the RE inequality with constant κ\kappaκ for all admissible J0J_0J0​ and δ\deltaδ. The paper's κ(s,c0)\kappa(s,c_0)κ(s,c0​) is the largest such κ\kappaκ (the minimum is attained), so "RE holds with κ(s,c0)≥κ2\kappa(s,c_0)\ge\kappa_2κ(s,c0​)≥κ2​" is exactly "κ2>0\kappa_2>0κ2​>0 is a witness". This avoids a real infimum over an empty set when J0=∅J_0=\emptysetJ0​=∅.
  • Ties. Every admissible choice of J1J_1J1​ (the mmm largest ∣δj∣|\delta_j|∣δj​∣ outside J0J_0J0​) is quantified over.
  • Restricted eigenvalues and correlations are sInf/sSup over nonempty bounded sets (a basis vector for ϕ\phiϕ; two disjoint singletons for θ\thetaθ, since M≥2M\ge2M≥2), so they equal the paper's attained min/max. uuu, sss, mmm are natural numbers; s≤M/2s\le M/2s≤M/2 is written 2s≤M2s\le M2s≤M.
  • Corrections of the printed statement. (1) Lemma 4.1 says the RE assumptions "hold with κ(s,c0)=κ(s,m,c0)=κ2(s,m,c0)\kappa(s,c_0)=\kappa(s,m,c_0)=\kappa_2(s,m,c_0)κ(s,c0​)=κ(s,m,c0​)=κ2​(s,m,c0​)" (and likewise with κ1\kappa_1κ1​); the proof gives only the lower bound, and the lower bound is what is stated. (2) The paper calls P01P_{01}P01​ "the projector in RM\mathbb R^MRM"; it acts on Rn\mathbb R^nRn. (3) The Section 4 claims "Assumption 3/4/5 implies RE(s,c0)(s,c_0)(s,c0​)" are stated with the explicit constant produced by the displayed argument, a labelled strengthening. (4) The Candès–Tao bound is stated with the hypotheses the proof uses: the blocks are disjoint, of sizes at most sss and 2s2s2s, and ϕmin⁡(2s)>0\phi_{\min}(2s)>0ϕmin​(2s)>0.
  • Ruling out trivializations. RE quantifies over all J0J_0J0​ with ∣J0∣≤s|J_0|\le s∣J0​∣≤s and all non-zero δ\deltaδ in the cone, and bounds the full ∣Xδ∣2|X\delta|_2∣Xδ∣2​, not ∣XδJ0∣2|X\delta_{J_0}|_2∣XδJ0​​∣2​; no hypothesis restricts XXX beyond the stated assumptions. The hypotheses are satisfiable: for n=M=4n=M=4n=M=4, X=2IX=2IX=2I (so Ψn=I\Psi_n=IΨn​=I), s=1s=1s=1, m=2m=2m=2, c0=1c_0=1c0​=1, Assumption 2 reads 2>12>12>1.
  • Infrastructure. A sparse-vector library (restriction, support, sorting coordinates into blocks) and facts about orthogonal projections onto column spans are needed; both are reusable for the other missions of this series and for compressed-sensing results. Proofs of any milestone, and alternative arguments, are welcome.

Selected references

  • P. J. Bickel, Y. Ritov, A. B. Tsybakov, Simultaneous analysis of Lasso and Dantzig selector, Ann. Statist. 37(4), 1705–1732, 2009. arXiv:0801.1095v3. https://arxiv.org/abs/0801.1095
  • E. Candès, T. Tao, The Dantzig selector: statistical estimation when p is much larger than n, Ann. Statist. 35(6), 2313–2351, 2007. https://arxiv.org/abs/math/0506081
  • N. Meinshausen, B. Yu, Lasso-type recovery of sparse representations for high-dimensional data, Ann. Statist. 37(1), 246–270, 2009. https://arxiv.org/abs/math/0605584
  • F. Bunea, A. B. Tsybakov, M. H. Wegkamp, Sparsity oracle inequalities for the Lasso, Electron. J. Statist. 1, 169–194, 2007. https://doi.org/10.1214/07-EJS008
  • D. L. Donoho, M. Elad, V. N. Temlyakov, Stable recovery of sparse overcomplete representations in the presence of noise, IEEE Trans. Inform. Theory 52(1), 6–18, 2006. https://doi.org/10.1109/TIT.2005.860430
11 thms3 active usersReviewed
🏆Completed
AnalysisNumerical AnalysisOperations Research·Captain: mikedeng1

Analysis of Generalized Pattern Searches: Nonnegative Clarke Derivatives at Limits of Refining SubsequencesResearch Paper

Motivation

Generalized pattern search (GPS) is a class of derivative-free methods for minimizing a function that can only be evaluated, not differentiated. Such objectives arise in engineering design, where one evaluation is an expensive simulation that may fail and return no value at all. The helicopter rotor design problem of Booker et al. is one example: no value was returned for roughly 66% of the trial points (Booker et al., 1999). A method for such problems has to tolerate objectives that are discontinuous or take the value +∞+\infty+∞.

Earlier convergence theory for GPS assumed continuous differentiability of the objective on a neighbourhood of the level set. Torczon established it for unconstrained problems (SIAM J. Optim. 7, 1997), and Lewis and Torczon extended it to bound constraints (1999) and to finitely many linear constraints (SIAM J. Optim. 10, 2000). Audet and Dennis (SIAM J. Optim. 13, 2003) replaced these analyses with a single argument. Its conclusions are local and are graded by the smoothness of the objective at the limit point only, through Clarke's generalized directional derivative. That paper is the source of this mission. Its analysis is the basis of the later mesh adaptive direct search (MADS) theory (Audet, Dennis, SIAM J. Optim. 17, 2006).

Setting

The problem is

min⁡x∈Ωf(x),f:Rn→R∪{+∞},Ω={x∈Rn:ℓ≤Ax≤u},\min_{x\in\Omega} f(x),\qquad f:\mathbb R^n\to\mathbb R\cup\{+\infty\},\qquad \Omega=\{x\in\mathbb R^n:\ell\le Ax\le u\},x∈Ωmin​f(x),f:Rn→R∪{+∞},Ω={x∈Rn:ℓ≤Ax≤u},

with A∈Rm×nA\in\mathbb R^{m\times n}A∈Rm×n and ℓ≤u\ell\le uℓ≤u in (R∪{±∞})m(\mathbb R\cup\{\pm\infty\})^m(R∪{±∞})m. The algorithm works with the barrier function fΩf_\OmegafΩ​, equal to fff on Ω\OmegaΩ and to +∞+\infty+∞ elsewhere.

The algorithm uses a finite set of directions D=GZˉD=G\bar ZD=GZˉ, the columns dj=Gzˉjd_j=G\bar z_jdj​=Gzˉj​ of the product of a nonsingular G∈Rn×nG\in\mathbb R^{n\times n}G∈Rn×n and an integer matrix Zˉ∈Zn×p\bar Z\in\mathbb Z^{n\times p}Zˉ∈Zn×p. The directions form a positive spanning set: their nonnegative combinations give all of Rn\mathbb R^nRn. At iteration kkk, with iterate xkx_kxk​ and mesh size parameter Δk>0\Delta_k>0Δk​>0, the mesh is Mk={xk+ΔkDz:z∈Z+p}M_k=\{x_k+\Delta_k Dz: z\in\mathbb Z_+^{p}\}Mk​={xk​+Δk​Dz:z∈Z+p​}. A poll set {xk+Δkd:d∈Dk}\{x_k+\Delta_k d: d\in D_k\}{xk​+Δk​d:d∈Dk​} is drawn from a positive spanning subset Dk⊆DD_k\subseteq DDk​⊆D. Each iteration ends in one of two ways:

  1. Improved mesh point. Some xk+1∈Mk∩Ωx_{k+1}\in M_k\cap\Omegaxk+1​∈Mk​∩Ω with fΩ(xk+1)<fΩ(xk)f_\Omega(x_{k+1})<f_\Omega(x_k)fΩ​(xk+1​)<fΩ​(xk​) was found, by the free SEARCH step or by the poll. Then Δk+1=τwkΔk\Delta_{k+1}=\tau^{w_k}\Delta_kΔk+1​=τwk​Δk​ with 0≤wk≤w+0\le w_k\le w^+0≤wk​≤w+.
  2. Mesh local optimizer. fΩ(xk)≤fΩ(xk+Δkd)f_\Omega(x_k)\le f_\Omega(x_k+\Delta_k d)fΩ​(xk​)≤fΩ​(xk​+Δk​d) for every d∈Dkd\in D_kd∈Dk​. Then xk+1=xkx_{k+1}=x_kxk+1​=xk​ and Δk+1=τwkΔk\Delta_{k+1}=\tau^{w_k}\Delta_kΔk+1​=τwk​Δk​ with w−≤wk≤−1w^-\le w_k\le-1w−≤wk​≤−1.

Here τ>1\tau>1τ>1 is rational and w−≤−1≤0≤w+w^-\le-1\le 0\le w^+w−≤−1≤0≤w+ are integers. The assumptions are A1 fΩ(x0)<∞f_\Omega(x_0)<\inftyfΩ​(x0​)<∞, A2 AAA is rational, and A3 all iterates lie in a compact set. A refining subsequence is an infinite set of mesh local optimizers {xk}k∈K\{x_k\}_{k\in K}{xk​}k∈K​ along which Δk→0\Delta_k\to 0Δk​→0 (Definition 3.5). For fff Lipschitz near x^\hat xx^, Clarke's derivative is

f∘(x^;d)=lim sup⁡y→x^, t↓0f(y+td)−f(y)t.f^\circ(\hat x;d)=\limsup_{y\to\hat x,\ t\downarrow 0}\frac{f(y+td)-f(y)}{t}.f∘(x^;d)=y→x^, t↓0limsup​tf(y+td)−f(y)​.

Formalization targets

Goal: Theorem 3.7

Assume A1–A3. Let x^\hat xx^ be the limit of a refining subsequence, and let d∈Dd\in Dd∈D be a direction polled at a feasible point xk+Δkdx_k+\Delta_k dxk​+Δk​d for infinitely many kkk in the subsequence. If fff is Lipschitz near x^\hat xx^, then

f∘(x^;d) ≥ 0.f^\circ(\hat x;d)\ \ge\ 0 .f∘(x^;d) ≥ 0.

Milestones on the way

  • Theorem 3.1: the iterates have a limit point, lim⁡kf(xk)\lim_k f(x_k)limk​f(xk​) exists and dominates fff at lower semicontinuity limit points, and all continuity limit points share one value.
  • Lemma 3.2: min⁡u≠v∈Mk∥u−v∥≥Δk/∥G−1∥\min_{u\ne v\in M_k}\|u-v\|\ge\Delta_k/\|G^{-1}\|minu=v∈Mk​​∥u−v∥≥Δk​/∥G−1∥ for every norm giving nonzero integer vectors norm at least 111.
  • Lemma 3.3: Δk≤Δ0τr+\Delta_k\le\Delta_0\tau^{r^+}Δk​≤Δ0​τr+ for some positive integer r+r^+r+.
  • Proposition 3.4: lim inf⁡k→∞Δk=0\liminf_{k\to\infty}\Delta_k=0liminfk→∞​Δk​=0.
  • Theorem 3.6: a convergent refining subsequence exists.

Corollaries

  • Theorem 3.9: if Ω=Rn\Omega=\mathbb R^nΩ=Rn and fff is strictly differentiable at x^\hat xx^, then ∇f(x^)=0\nabla f(\hat x)=0∇f(x^)=0.
  • Theorem 3.14: if the poll sets conform to the boundary of Ω\OmegaΩ (Definition 3.13) and fff is strictly differentiable at x^\hat xx^, then ∇f(x^)Tw≥0\nabla f(\hat x)^Tw\ge 0∇f(x^)Tw≥0 on the tangent cone TΩ(x^)T_\Omega(\hat x)TΩ​(x^) and −∇f(x^)∈NΩ(x^)-\nabla f(\hat x)\in N_\Omega(\hat x)−∇f(x^)∈NΩ​(x^). So x^\hat xx^ is a KKT point.

Significance

Theorem 3.7 gives a first-order conclusion at a limit point from a local hypothesis at that point alone. It does not require smoothness elsewhere, finiteness of fff elsewhere, or continuity. It turns the heuristic "the method stopped improving on ever finer meshes" into a statement about generalized derivatives. The unconstrained stationarity result (Theorem 3.9) and the linearly constrained KKT result (Theorem 3.14) follow from it, and they recover the Torczon and Lewis–Torczon theorems under weaker smoothness assumptions. The chain Lemma 3.2 → Lemma 3.3 → Proposition 3.4 → Theorem 3.6 shows that the goal's hypothesis is always met. Every run satisfying A1 and A3 has a refining subsequence, which rests on the rationality of τ\tauτ and on the integer structure of DDD.

All results in this mission are proved in the source paper. None of them has, to the best of our knowledge, a machine-checked proof. The mission contributes a formal model of the GPS algorithm class as a class of runs, a formal Clarke directional derivative, and checked proofs of the mesh-refinement chain and the main theorem.

Difficulty

Given a refining subsequence, the goal is a comparison of limsups: the poll inequalities give nonnegative difference quotients at the points (xk,Δk)(x_k,\Delta_k)(xk​,Δk​), which converge to (x^,0+)(\hat x,0^+)(x^,0+). The difficulty lies in two places. First, the objective is extended-valued, and the barrier hides fff at infeasible poll points, where the poll inequality fΩ(xk)≤+∞f_\Omega(x_k)\le+\inftyfΩ​(xk​)≤+∞ says nothing. The hypothesis on ddd has to supply feasibility, and the Lipschitz hypothesis has to supply finiteness near x^\hat xx^. Second, the existence of refining subsequences is not a compactness argument alone. Coarsening is allowed, so Δk\Delta_kΔk​ need not decrease, and with an irrational τ\tauτ or a direction set that is not an integer lattice image (for instance D=[−1,+π]D=[-1,+\pi]D=[−1,+π] in R\mathbb RR) the meshes can be dense and lim inf⁡Δk\liminf\Delta_kliminfΔk​ can be positive. The lattice argument behind Proposition 3.4 is where the integrality hypotheses are used.

Formalization scope

Points of Rn\mathbb R^nRn are Fin n → ℝ, fff takes values in WithTop ℝ, and the bounds ℓ,u\ell,uℓ,u are EReal-valued, so m=0m=0m=0 gives Ω=Rn\Omega=\mathbb R^nΩ=Rn. The barrier is defined by cases, never by extended addition. Directions are the columns of G * Zbar indexed by Fin p, and DkD_kDk​ is a Finset (Fin p). A GPS run is a structure of sequences xk,Δk,Dk,wkx_k,\Delta_k,D_k,w_kxk​,Δk​,Dk​,wk​ and a per-iteration predicate "mesh local optimizer", subject to exactly the two update rules above, Δ0>0\Delta_0>0Δ0​>0, rational τ>1\tau>1τ>1 and the exponent bounds. The SEARCH step, the choice of DkD_kDk​ and the exponents are left free, since the paper allows any strategy. A subsequence is a strictly increasing map K:N→NK:\mathbb N\to\mathbb NK:N→N. The Clarke derivative of a real function is an EReal-valued limit superior along y→x^y\to\hat xy→x^, t→0+t\to 0^+t→0+. "fff Lipschitz near x^\hat xx^" means that fff agrees near x^\hat xx^ with a real function Lipschitz there, and the conclusions are stated for every such function. Strict differentiability is the directional notion of Section 3.4 of the paper.

The goal is not trivialized by an empty run class: Theorem 3.6, on the same class, asserts that refining subsequences exist. The mesh-local-optimizer branch requires the complete poll inequality over DkD_kDk​. The Clarke limit superior cannot take a default value. The direction ddd must be polled at feasible points infinitely often, which is the paper's "fff was evaluated".

Contributions welcome: proofs of any milestone, and reusable lemmas on positive spanning sets, lattice points in compact sets, and the Clarke derivative (for instance, that it equals ∇f(x^)Td\nabla f(\hat x)^Td∇f(x^)Td under strict differentiability).

Selected references

  • C. Audet, J. E. Dennis Jr., Analysis of Generalized Pattern Searches, SIAM J. Optim. 13(3):889–903, 2003. https://doi.org/10.1137/S1052623400378742
  • V. Torczon, On the Convergence of Pattern Search Algorithms, SIAM J. Optim. 7(1):1–25, 1997. https://doi.org/10.1137/S1052623493250780
  • R. M. Lewis, V. Torczon, Pattern Search Methods for Linearly Constrained Minimization, SIAM J. Optim. 10(3):917–941, 2000. https://doi.org/10.1137/S1052623497331373
  • F. H. Clarke, Optimization and Nonsmooth Analysis, Wiley, 1983; reprinted SIAM Classics in Applied Mathematics 5, 1990. https://doi.org/10.1137/1.9781611971309
  • A. J. Booker, J. E. Dennis Jr., P. D. Frank, D. B. Serafini, V. Torczon, M. W. Trosset, A rigorous framework for optimization of expensive functions by surrogates, Structural Optimization 17:1–13, 1999. https://doi.org/10.1007/BF01197559
  • C. Audet, J. E. Dennis Jr., Mesh Adaptive Direct Search Algorithms for Constrained Optimization, SIAM J. Optim. 17(1):188–217, 2006. https://doi.org/10.1137/040603371
13 thms3 active usersReviewed
🏆Completed
CombinatoricsOperations ResearchProbability+1·Captain: mikedeng1

Maximizing Non-Monotone Submodular Functions II: A Nonadaptive Algorithm Achieves 1/3 of the OptimumResearch Paper

Motivation

Maximizing a submodular set function without constraints contains Max Cut, Max Directed Cut, maximum facility location and several graph and hypergraph cut problems as special cases, and it appears in operations research wherever a value exhibits diminishing returns but is not monotone (profit that combines coverage with a cost, for example). These problems are NP-hard, so the question is which fraction of the optimum an efficient algorithm can guarantee when the function is accessible only through a value oracle that returns f(S)f(S)f(S) for a queried set SSS.

Feige, Mirrokni and Vondrák (SIAM J. Comput. 40(4), 2011) gave the first constant-factor approximation algorithms for maximizing a general nonnegative submodular function. The simplest of them returns a uniformly random set and achieves 1/41/41/4 of the optimum; this mission is about the next one, a nonadaptive algorithm: it decides all of its oracle queries before seeing any answer, then computes a set from the answers. Such an algorithm can be run in one round of parallel queries. The paper shows that this restricted access already beats 1/41/41/4 and reaches 1/31/31/3.

Timeline. For Max Directed Cut, a random cut achieves 1/41/41/4. Feige, Mirrokni and Vondrák (FOCS 2007; journal version 2011) proved 1/41/41/4 for a random set and 1/31/31/3 nonadaptively for general nonnegative submodular functions, 1/31/31/3 and 2/52/52/5 by adaptive local search, and that 1/21/21/2 requires exponentially many queries. Buchbinder, Feldman, Naor and Schwartz (FOCS 2012, SIAM J. Comput. 2015) later reached the optimal 1/21/21/2 with a randomized double-greedy algorithm.

Setting

Let XXX be a finite ground set with n=∣X∣≥1n = |X| \ge 1n=∣X∣≥1 elements. A function f:2X→Rf : 2^X \to \mathbb{R}f:2X→R is submodular (Definition 1.1) if

f(S∪T)+f(S∩T)≤f(S)+f(T)for all S,T⊆X.f(S \cup T) + f(S \cap T) \le f(S) + f(T) \qquad \text{for all } S, T \subseteq X .f(S∪T)+f(S∩T)≤f(S)+f(T)for all S,T⊆X.

Throughout, fff is nonnegative, the paper's standing assumption, and OPT=max⁡S⊆Xf(S)OPT = \max_{S \subseteq X} f(S)OPT=maxS⊆X​f(S).

For p∈[0,1]p \in [0,1]p∈[0,1], X(p)X(p)X(p) denotes the random subset of XXX containing each element independently with probability ppp; R=X(1/2)R = X(1/2)R=X(1/2) is a uniformly random subset. For a set A⊆XA \subseteq XA⊆X, A(p)A(p)A(p) is the analogous random subset of AAA. The averaged marginal value of an element (Definition 2.4) is

ω(x)=E[f(R∪{x})−f(R∖{x})],R=X(1/2).\omega(x) = \mathbf{E}\big[f(R \cup \{x\}) - f(R \setminus \{x\})\big], \qquad R = X(1/2).ω(x)=E[f(R∪{x})−f(R∖{x})],R=X(1/2).

Algorithm NA (p. 1139):

  1. by random sampling, compute estimates ω~(x)\tilde\omega(x)ω~(x) with ∣ω~(x)−ω(x)∣<OPT/n2|\tilde\omega(x) - \omega(x)| < OPT/n^2∣ω~(x)−ω(x)∣<OPT/n2 for all xxx, with high probability;
  2. independently, sample R=X(1/2)R = X(1/2)R=X(1/2);
  3. with probability 8/98/98/9 return RRR;
  4. with probability 1/91/91/9 return A={x∈X:ω~(x)>0}A = \{x \in X : \tilde\omega(x) > 0\}A={x∈X:ω~(x)>0}.

Given the estimates, the expected value NA returns is 89 E[f(X(1/2))]+19f(A)\tfrac89\,\mathbf{E}[f(X(1/2))] + \tfrac19 f(A)98​E[f(X(1/2))]+91​f(A).

Formalization targets

Goal: Theorem 2.6 in the explicit form of its proof

For every nonnegative submodular fff and every estimate ω~\tilde\omegaω~ with ∣ω~(x)−ω(x)∣<OPT/n2|\tilde\omega(x) - \omega(x)| < OPT/n^2∣ω~(x)−ω(x)∣<OPT/n2 for all xxx,

89 E[f(X(1/2))]+19 f({x:ω~(x)>0}) ≥ (13−49n) OPT.\frac89\,\mathbf{E}[f(X(1/2))] + \frac19\, f\big(\{x : \tilde\omega(x) > 0\}\big) \ \ge\ \Big(\frac13 - \frac{4}{9n}\Big)\, OPT .98​E[f(X(1/2))]+91​f({x:ω~(x)>0}) ≥ (31​−9n4​)OPT.

The printed theorem says "at least (1/3−o(1)) OPT(1/3 - o(1))\,OPT(1/3−o(1))OPT"; the term 4/(9n)4/(9n)4/(9n) is what the proof establishes (p. 1140, last display).

Milestones

  1. Lemma 2.2: E[g(A(p))]≥(1−p) g(∅)+p g(A)\mathbf{E}[g(A(p))] \ge (1-p)\,g(\emptyset) + p\,g(A)E[g(A(p))]≥(1−p)g(∅)+pg(A) for submodular ggg.
  2. Lemma 2.3: E[f(A(p)∪B(q))]≥(1−p)(1−q)f(∅)+p(1−q)f(A)+(1−p)qf(B)+pqf(A∪B)\mathbf{E}[f(A(p) \cup B(q))] \ge (1-p)(1-q) f(\emptyset) + p(1-q) f(A) + (1-p)q f(B) + pq f(A \cup B)E[f(A(p)∪B(q))]≥(1−p)(1−q)f(∅)+p(1−q)f(A)+(1−p)qf(B)+pqf(A∪B) for independently sampled, possibly overlapping A,BA, BA,B.
  3. For B=X∖AB = X \setminus AB=X∖A and any CCC: f(A)+f(B∩C)+f(B∪C)≥f(C)f(A) + f(B \cap C) + f(B \cup C) \ge f(C)f(A)+f(B∩C)+f(B∪C)≥f(C).
  4. If ω≤OPT/n2\omega \le OPT/n^2ω≤OPT/n2 on BBB: E[f(R∪(B∩C))]≤E[f(R)]+OPT/(2n)\mathbf{E}[f(R \cup (B \cap C))] \le \mathbf{E}[f(R)] + OPT/(2n)E[f(R∪(B∩C))]≤E[f(R)]+OPT/(2n).
  5. E[f(R∪(B∩C))]≥14f(B∩C)+14f(C)\mathbf{E}[f(R \cup (B \cap C))] \ge \tfrac14 f(B \cap C) + \tfrac14 f(C)E[f(R∪(B∩C))]≥41​f(B∩C)+41​f(C).
  6. If ω≥−OPT/n2\omega \ge -OPT/n^2ω≥−OPT/n2 on AAA and B=X∖AB = X \setminus AB=X∖A: E[f(R)]≥E[f(R∩(B∪C))]−OPT/(2n)\mathbf{E}[f(R)] \ge \mathbf{E}[f(R \cap (B \cup C))] - OPT/(2n)E[f(R)]≥E[f(R∩(B∪C))]−OPT/(2n).
  7. E[f(R∩(B∪C))]≥14f(C)+14f(B∪C)\mathbf{E}[f(R \cap (B \cup C))] \ge \tfrac14 f(C) + \tfrac14 f(B \cup C)E[f(R∩(B∪C))]≥41​f(C)+41​f(B∪C).

Milestones 3–7 are the displayed steps of the proof of Theorem 2.6, stated for arbitrary sets where the page's argument does not use the optimality of CCC.

Significance

The theorem shows that nonadaptive access, a fixed batch of polynomially many value queries followed by a computation, suffices for a 1/31/31/3-approximation of unconstrained nonnegative submodular maximization, strictly better than the 1/41/41/4 of any algorithm that must return one of its queried sets (the paper shows 1/41/41/4 is optimal in that class, §4.2). The quantity ω\omegaω generalizes the in-degree/out-degree test for Max Directed Cut to arbitrary submodular functions, and Lemmas 2.2 and 2.3 are general sampling inequalities for submodular functions that the paper reuses for its adaptive smooth local search.

Formalizing it produces machine-checked versions of Lemmas 2.2 and 2.3 as statements about exact finite averages, a reusable expectation operator on product-distributed random subsets, and a checked version of the 1/31/31/3 argument with its explicit error term. The result is proved in the paper; to our knowledge none of it has been formalized in a proof assistant.

Difficulty

The two regimes the proof separates, "AAA is already good" and "one of f(B∩C)f(B \cap C)f(B∩C), f(B∪C)f(B \cup C)f(B∪C) is large", must be tied to the value of a uniformly random set, whereas the elements of AAA and BBB are chosen from estimated averages, not from the optimal set CCC. The natural attempt, comparing f(R)f(R)f(R) with f(C)f(C)f(C) element by element, fails because fff is not monotone: adding elements of CCC to RRR can decrease the value. The accuracy OPT/n2OPT/n^2OPT/n2 of the estimates must also be propagated through a sum over up to nnn elements, which is where the error term 4/(9n)4/(9n)4/(9n) comes from. The sampling lemmas require handling expectations over pairs of independent random subsets of possibly overlapping sets.

Formalization scope

  • The ground set is a Fintype X with DecidableEq, assumed Nonempty, so n=∣X∣≥1n = |X| \ge 1n=∣X∣≥1 and the divisions by nnn and n2n^2n2 are genuine; sets are Finset X; fff is real valued with nonnegativity ∀S, 0≤f(S)\forall S,\ 0 \le f(S)∀S, 0≤f(S) as an explicit hypothesis. Lemmas 2.2 and 2.3 are stated for real fff with no sign condition, as printed.
  • OPTOPTOPT is Finset.univ.sup' _ f, the true maximum over all subsets.
  • Every expectation over an independently sampled random set is the exact finite sum F(x)=∑Sf(S)∏i∈Sxi∏i∉S(1−xi)F(x) = \sum_{S} f(S)\prod_{i \in S} x_i \prod_{i \notin S}(1 - x_i)F(x)=∑S​f(S)∏i∈S​xi​∏i∈/S​(1−xi​); X(1/2)X(1/2)X(1/2) is x≡1/2x \equiv 1/2x≡1/2. Expectations over two independent samples (Lemma 2.3) are the corresponding iterated sums. Sampling probabilities carry the hypotheses 0≤p,q≤10 \le p, q \le 10≤p,q≤1.
  • The goal quantifies over every estimate ω~\tilde\omegaω~ satisfying the printed accuracy ∣ω~(x)−ω(x)∣<OPT/n2|\tilde\omega(x) - \omega(x)| < OPT/n^2∣ω~(x)−ω(x)∣<OPT/n2 (strict), with A={x:ω~(x)>0}A = \{x : \tilde\omega(x) > 0\}A={x:ω~(x)>0} (strict). The "with high probability" of NA's first step is this hypothesis; the sampling estimate that makes it likely (Lemma 2.5, a Chernoff-bound argument) is not part of the goal. When OPT=0OPT = 0OPT=0 the hypothesis is unsatisfiable, but then f≡0f \equiv 0f≡0 and nothing is lost.
  • The left-hand side is exactly the mixture 89 E[f(X(1/2))]+19f(A)\tfrac89\,\mathbf{E}[f(X(1/2))] + \tfrac19 f(A)98​E[f(X(1/2))]+91​f(A). A statement with the maximum of the two terms, with exact values ω~=ω\tilde\omega = \omegaω~=ω, or with the o(1)o(1)o(1) replaced by an existential constant or a limit, is a different (and weaker or stronger) theorem and does not close this mission.
  • Printed slip corrected: in the second display on p. 1140, the "===" before −∣A∖C∣ OPT/(2n2)-|A \setminus C|\,OPT/(2n^2)−∣A∖C∣OPT/(2n2) should be "≥\ge≥"; milestone 6 states the inequality.

Welcome contributions: proofs of Lemmas 2.2 and 2.3 (reusable for mission IV of this series), the identity E[f(R∪{x})−f(R)]=12ω(x)\mathbf{E}[f(R \cup \{x\}) - f(R)] = \tfrac12\omega(x)E[f(R∪{x})−f(R)]=21​ω(x), and general lemmas about the operator FFF (splitting a uniform random set along a partition).

Selected references

  • U. Feige, V. S. Mirrokni, J. Vondrák, Maximizing Non-Monotone Submodular Functions, SIAM J. Comput. 40(4):1133–1153, 2011. https://doi.org/10.1137/090779346
  • N. Buchbinder, M. Feldman, J. Naor, R. Schwartz, A Tight Linear Time (1/2)-Approximation for Unconstrained Submodular Maximization, SIAM J. Comput. 44(5):1384–1402, 2015. https://doi.org/10.1137/130929205
12 thms3 active usersReviewed
🏆Completed
Operations ResearchProbability·Captain: mikedeng1

On Properties of Stochastic Inventory Systems IV: The (Q, r) Cost Is Flatter in the Order Quantity than the EOQ CostResearch Paper

Motivation

The continuous-review (Q,r)(Q, r)(Q,r) policy is the standard replenishment rule of inventory theory: whenever the inventory position (stock on hand plus on order minus backorders) drops to the reorder point rrr, order a fixed order quantity QQQ. It is used in practice and taught in every operations management course, usually after the deterministic economic order quantity (EOQ) model, which is the same system with a constant demand stream.

Practitioners and textbooks rely on a robustness property of the EOQ: its cost is very insensitive to the choice of order quantity. If the order quantity is off by a factor α\alphaα, the cost rises only by the factor 12(α+1/α)\tfrac12(\alpha + 1/\alpha)21​(α+1/α); ordering 50% too much costs about 8% extra. The insensitivity of the stochastic (Q,r)(Q, r)(Q,r) system to its control parameters had been observed numerically (Wagner, O'Hagan and Lundh 1965; Naddor 1975; Archibald and Silver 1978), but, as Zheng notes, no analytical result on it was known.

Timeline:

  • 1963: Hadley and Whitin derive the (Q,r)(Q, r)(Q,r) cost for Poisson demand.
  • 1986: Zipkin proves that the average backorders of a (Q,r)(Q,r)(Q,r) policy are jointly convex in (Q,r)(Q, r)(Q,r) under continuous demand (Zipkin 1986).
  • 1992: Zheng derives simple optimality conditions for the continuous (Q,r)(Q, r)(Q,r) model and compares it with the EOQ model under the same cost structure. One of the results is that the stochastic cost curve is flatter in the order quantity than the EOQ curve (Zheng 1992). This mission formalizes that result.

Setting

Demands arrive at rate λ>0\lambda>0λ>0; orders arrive after a fixed leadtime L>0L>0L>0; all stockouts are backordered. Each order costs K>0K>0K>0; holding costs accrue at rate h>0h>0h>0 per unit in stock and penalty costs at rate p>0p>0p>0 per unit backordered. The leadtime demand D≥0D\ge 0D≥0 has distribution μ\muμ with finite mean E(D)=λLE(D) = \lambda LE(D)=λL.

The inventory cost rate at inventory position yyy is

G(y)=E[h(y−D)++p(D−y)+],G(y) = E\big[h(y-D)^+ + p(D-y)^+\big],G(y)=E[h(y−D)++p(D−y)+],

assumed to attain its minimum at a unique point y0y^0y0. The long-run average cost of the policy (Q,r)(Q, r)(Q,r) is

c(Q,r)=λK+∫rr+QG(y) dyQ,Q>0.c(Q, r) = \frac{\lambda K + \int_r^{r+Q} G(y)\,dy}{Q}, \qquad Q>0.c(Q,r)=QλK+∫rr+Q​G(y)dy​,Q>0.

For fixed Q>0Q>0Q>0 let r(Q)r(Q)r(Q) be a reorder point minimizing c(Q,⋅)c(Q,\cdot)c(Q,⋅), and let

C(Q)=c(Q,r(Q)),H(Q)=G(r(Q)) (Q>0),H(0)=G(y0).C(Q) = c(Q, r(Q)), \qquad H(Q) = G(r(Q))\ (Q>0), \quad H(0) = G(y^0).C(Q)=c(Q,r(Q)),H(Q)=G(r(Q)) (Q>0),H(0)=G(y0).

CCC is the cost of the order quantity QQQ when the reorder point is always chosen optimally for it. An optimal order quantity Q∗Q^*Q∗ minimizes CCC over Q>0Q>0Q>0, and C∗=C(Q∗)C^* = C(Q^*)C∗=C(Q∗).

The EOQ model is the same system with the constant leadtime demand λL\lambda LλL. Its cost rate is Gd(y)=h(y−λL)++p(λL−y)+G_d(y) = h(y-\lambda L)^+ + p(\lambda L-y)^+Gd​(y)=h(y−λL)++p(λL−y)+, and rdr_drd​, HdH_dHd​, CdC_dCd​ are the objects above at GdG_dGd​, with optimum Qd∗Q^*_dQd∗​ and Cd∗C^*_dCd∗​.

Formalization targets

Goal: Theorem 4

C(αQ∗)C∗≤12(α+1α)∀α>0.\frac{C(\alpha Q^*)}{C^*} \le \frac12\left(\alpha + \frac1\alpha\right) \qquad \forall \alpha>0.C∗C(αQ∗)​≤21​(α+α1​)∀α>0.

The goal holds for every demand distribution satisfying the standing assumptions and every optimal Q∗Q^*Q∗. Both regimes, α<1\alpha<1α<1 and α>1\alpha>1α>1, are included.

Milestones

In the order the proof uses them:

  1. Eq. (7): ∫r(Q)r(Q)+QG=∫0QH\int_{r(Q)}^{r(Q)+Q} G = \int_0^Q H∫r(Q)r(Q)+Q​G=∫0Q​H, hence C(Q)=(λK+∫0QH(y)dy)/QC(Q) = \big(\lambda K + \int_0^Q H(y)dy\big)/QC(Q)=(λK+∫0Q​H(y)dy)/Q for Q>0Q>0Q>0.
  2. Lemma 4: HHH is increasing and convex on [0,∞)[0,\infty)[0,∞) with asymptotic slope hp/(h+p)hp/(h+p)hp/(h+p).
  3. Eq. (8): an optimal Q∗Q^*Q∗ exists, and Q>0Q>0Q>0 is optimal iff H(Q)=C(Q)H(Q) = C(Q)H(Q)=C(Q).
  4. Eq. (18): Hd(Q)=hph+pQH_d(Q) = \frac{hp}{h+p}QHd​(Q)=h+php​Q, with rd(Q)=λL−hh+pQr_d(Q) = \lambda L - \frac{h}{h+p}Qrd​(Q)=λL−h+ph​Q.
  5. Lemma 7: H0(Q)≤Hd(Q)≤H(Q)H_0(Q) \le H_d(Q) \le H(Q)H0​(Q)≤Hd​(Q)≤H(Q) and A(Q)≤Ad(Q)A(Q)\le A_d(Q)A(Q)≤Ad​(Q), where H0=H−G(y0)H_0 = H - G(y^0)H0​=H−G(y0) and A(Q)=QH(Q)−∫0QHA(Q) = QH(Q) - \int_0^Q HA(Q)=QH(Q)−∫0Q​H.
  6. Eqs. (26)–(27): H(αQ)≤αH(Q)H(\alpha Q)\le \alpha H(Q)H(αQ)≤αH(Q) for α>1\alpha>1α>1 and H(αQ)≥αH(Q)H(\alpha Q)\ge\alpha H(Q)H(αQ)≥αH(Q) for 0<α<10<\alpha<10<α<1.
  7. Lemma 9: ∫QαQH(y) dy≤α2−12 QH(Q)\int_Q^{\alpha Q} H(y)\,dy \le \frac{\alpha^2-1}{2}\,Q H(Q)∫QαQ​H(y)dy≤2α2−1​QH(Q) for all α>0\alpha>0α>0, Q>0Q>0Q>0.

Significance

In the EOQ model the relative cost of a scaled order quantity is exactly Cd(αQd∗)/Cd∗=12(α+1/α)C_d(\alpha Q^*_d)/C^*_d = \tfrac12(\alpha + 1/\alpha)Cd​(αQd∗​)/Cd∗​=21​(α+1/α) (Eq. (25) of the paper). Theorem 4 shows that the stochastic system is at least as forgiving. The bound holds for every leadtime-demand distribution with a unique newsvendor minimizer, and it does not depend on the parameters KKK, hhh, ppp, λ\lambdaλ or LLL. Because the reorder point is re-optimized for each quantity, the bound applies to the practical question of how much a misestimated lot size costs when the safety stock is set correctly.

Together with the other results of the paper (the 1/81/81/8 bound for the EOQ heuristic and the bounds between Q∗Q^*Q∗ and Qd∗Q^*_dQd∗​, which are separate missions of this series), it gives a closed-form account of why the EOQ is a good heuristic for stochastic systems.

The result has a complete published proof. It has not been machine-checked. The work that remains is a formal proof for general distributions: the paper differentiates GGG and r(Q)r(Q)r(Q) twice, and a formal proof has to replace those derivatives with arguments that need no density.

Difficulty

C(Q)C(Q)C(Q) is defined through an inner minimization over the reorder point, so its shape in QQQ is controlled by the implicitly defined function H(Q)=G(r(Q))H(Q) = G(r(Q))H(Q)=G(r(Q)) rather than by GGG directly. The obvious approach would bound C(αQ∗)C(\alpha Q^*)C(αQ∗) with the reorder point fixed at r(Q∗)r(Q^*)r(Q∗). That approach is the wrong comparison: it bounds a larger quantity, and the resulting bound depends on the distribution.

The paper's proof uses three properties of HHH: that it is convex, that its slope never exceeds the EOQ slope hp/(h+p)hp/(h+p)hp/(h+p), and that it dominates HdH_dHd​. The paper obtains these from the derivatives r′(Q)r'(Q)r′(Q) and H′(Q)H'(Q)H′(Q) under a smooth demand distribution. Without a density, r(Q)r(Q)r(Q) is only an argmin and HHH need not be differentiable, so none of these three properties can be read off a derivative formula; the asymptotic slope in particular depends on the finite mean E(D)=λLE(D) = \lambda LE(D)=λL and on the behaviour of GGG at ±∞\pm\infty±∞.

Formalization scope

The mission is set in Lean 4 with Mathlib. All objects are real valued.

  • Model. The structure QRModel bundles λ,L,K,h,p>0\lambda, L, K, h, p>0λ,L,K,h,p>0, a probability measure μ\muμ on R\mathbb{R}R with integrable identity, ∫x dμ=λL\int x\,d\mu = \lambda L∫xdμ=λL, D≥0D\ge 0D≥0 almost surely, and the unique-minimizer hypothesis on GGG. K>0K>0K>0 is implicit in the paper and made explicit here. No density is assumed; deterministic and discrete demands are allowed, and the paper's own numerical study uses Poisson demand.
  • Generic machinery. ccc, r(Q)r(Q)r(Q), y0y^0y0, HHH, H0H_0H0​, CCC and AAA are defined for an arbitrary cost rate and instantiated at GGG and at GdG_dGd​. r(Q)r(Q)r(Q) and y0y^0y0 are chosen minimizers; they are never defined by the equation G(r)=G(r+Q)G(r) = G(r+Q)G(r)=G(r+Q), which is a lemma of the paper. H(0)=G(y0)H(0) = G(y^0)H(0)=G(y0). Values at Q<0Q<0Q<0 (and of ccc, CCC at Q≤0Q\le 0Q≤0) are junk, and every statement restricts to Q>0Q>0Q>0 or Q≥0Q\ge 0Q≥0.
  • Readings of informal words. "Increasing" in Lemma 4 is strict on [0,∞)[0,\infty)[0,∞), since the proof shows H′>0H'>0H′>0. "Asymptotic slope hp/(h+p)hp/(h+p)hp/(h+p)" is stated as H(Q)/Q→hp/(h+p)H(Q)/Q\to hp/(h+p)H(Q)/Q→hp/(h+p) together with the chord bound H(Q2)−H(Q1)≤hph+p(Q2−Q1)H(Q_2)-H(Q_1)\le \frac{hp}{h+p}(Q_2-Q_1)H(Q2​)−H(Q1​)≤h+php​(Q2​−Q1​) for 0≤Q1≤Q20\le Q_1\le Q_20≤Q1​≤Q2​. The chord bound is the derivative-free form of H′≤hp/(h+p)H'\le hp/(h+p)H′≤hp/(h+p) that the proofs of Lemmas 7–9 use. "The optimal order quantity" is IsOptQty Q, meaning Q>0Q>0Q>0 and C(Q)≤C(Q′)C(Q)\le C(Q')C(Q)≤C(Q′) for all Q′>0Q'>0Q′>0. Its existence is asserted in the Eq. (8) milestone, so the goal is not vacuous. "∀α>0\forall\alpha>0∀α>0" is a real α>0\alpha>0α>0 with real division 1/α1/\alpha1/α. In Lemma 9 the integral ∫QαQ\int_Q^{\alpha Q}∫QαQ​ is oriented, as on the page.
  • Ruling out trivializations. C(αQ∗)C(\alpha Q^*)C(αQ∗) re-optimizes the reorder point for αQ∗\alpha Q^*αQ∗; holding it at r(Q∗)r(Q^*)r(Q∗) would be a different theorem. C∗>0C^*>0C∗>0 is a consequence of the model, not a hypothesis.

A complete development needs the following:

  • integrability and continuity of GGG;
  • existence of the optimal reorder point;
  • convexity of HHH;
  • the asymptotics G−Gd→0G - G_d\to 0G−Gd​→0 at ±∞\pm\infty±∞;
  • Jensen's inequality Gd≤GG_d\le GGd​≤G (Eq. (22));
  • existence of Q∗Q^*Q∗.

These facts about newsvendor cost functions are reusable in the other missions of this series. Contributions of any of them as separate lemmas are welcome.

Selected references

  • Y.-S. Zheng, On Properties of Stochastic Inventory Systems, Management Science 38(1):87–103, 1992. https://doi.org/10.1287/mnsc.38.1.87
  • P. H. Zipkin, Inventory Service-Level Measures: Convexity and Approximation, Management Science 32(8):975–981, 1986. https://doi.org/10.1287/mnsc.32.8.975
  • G. Hadley and T. M. Whitin, Analysis of Inventory Systems, Prentice-Hall, 1963.
  • H. M. Wagner, M. O'Hagan and B. Lundh, An Empirical Study of Exactly and Approximately Optimal Inventory Policies, Management Science 11(7):690–723, 1965. https://doi.org/10.1287/mnsc.11.7.690
  • A. Federgruen and Y.-S. Zheng, An Efficient Algorithm for Computing an Optimal (r, Q) Policy in Continuous Review Stochastic Inventory Systems, Operations Research 40(4):808–813, 1992. https://doi.org/10.1287/opre.40.4.808
11 thms3 active usersReviewed
🏆Completed
Operations ResearchProbability·Captain: mikedeng1

On Properties of Stochastic Inventory Systems II: The Optimal Order Quantity of the Stochastic (Q, r) Model Exceeds the EOQ by a Bounded GapResearch Paper

Motivation

The continuous-review (Q,r)(Q, r)(Q,r) policy is the standard control rule for a single stocked item with random demand: whenever the inventory position falls to the reorder point rrr, an order of fixed size QQQ is placed. It is implemented in a large share of commercial inventory systems. Choosing the two parameters jointly has traditionally required numerical search (Hadley and Whitin, 1963; Federgruen and Zheng, 1992). In practice the order quantity is therefore often taken from the deterministic economic order quantity (EOQ) formula with backorders, and the reorder point is then set for the random demand.

Zheng (1992) turned this practice into a question with an exact answer: how does the optimal order quantity Q∗Q^*Q∗ of the stochastic model compare with the EOQ quantity Qd∗Q^*_dQd∗​ computed from the same cost data and the same mean demand? Its Theorem 2 answers it with a two-sided bound. This mission formalizes that theorem. Companion missions of the same series formalize the paper's cost bounds (Theorem 3), the flatness of the cost curve (Theorem 4) and the 1/81/81/8 bound on the cost of using the EOQ quantity (Theorem 5).

Setting

Demand arrives at rate λ>0\lambda > 0λ>0 and replenishment orders arrive after a fixed leadtime L>0L > 0L>0. Shortages are backordered. Holding costs accrue at rate h>0h > 0h>0 per unit held, backorder penalties at rate p>0p > 0p>0 per unit short, and every order costs K>0K > 0K>0. The leadtime demand DDD is a nonnegative random variable with law μ\muμ and mean E(D)=λL\mathbb{E}(D) = \lambda LE(D)=λL. The expected inventory cost rate at inventory position yyy is the newsvendor cost

G(y)=E[h(y−D)++p(D−y)+],G(y) = \mathbb{E}\big[h(y - D)^+ + p(D - y)^+\big],G(y)=E[h(y−D)++p(D−y)+],

assumed, as in the paper, to attain its minimum at a unique point y0y^0y0. The long-run average cost of the policy (Q,r)(Q, r)(Q,r) is

c(Q,r)=λK+∫rr+QG(y) dyQ.c(Q, r) = \frac{\lambda K + \int_r^{r+Q} G(y)\,dy}{Q}.c(Q,r)=QλK+∫rr+Q​G(y)dy​.

For each Q>0Q > 0Q>0, let r(Q)r(Q)r(Q) be an optimal reorder point, i.e. a minimizer of c(Q,⋅)c(Q, \cdot)c(Q,⋅). The analysis runs through the curves

H(Q)=G(r(Q)) (Q>0),H(0)=G(y0),H0(Q)=H(Q)−G(y0),A(Q)=QH(Q)−∫0QH(y) dy,H(Q) = G(r(Q))\ (Q > 0),\quad H(0) = G(y^0),\qquad H_0(Q) = H(Q) - G(y^0),\qquad A(Q) = QH(Q) - \int_0^Q H(y)\,dy,H(Q)=G(r(Q)) (Q>0),H(0)=G(y0),H0​(Q)=H(Q)−G(y0),A(Q)=QH(Q)−∫0Q​H(y)dy,

and through the cost C(Q)=c(Q,r(Q))C(Q) = c(Q, r(Q))C(Q)=c(Q,r(Q)) of order quantity QQQ with the reorder point set optimally. The optimal order quantity Q∗Q^*Q∗ is the minimizer of CCC over Q>0Q > 0Q>0.

The EOQ model is the case of a constant leadtime demand λL\lambda LλL. Its cost rate is Gd(y)=h(y−λL)++p(λL−y)+G_d(y) = h(y - \lambda L)^+ + p(\lambda L - y)^+Gd​(y)=h(y−λL)++p(λL−y)+, and the same construction gives rdr_drd​, HdH_dHd​, AdA_dAd​ and the optimal quantity

Qd∗=2λK(h+p)hp.Q^*_d = \sqrt{\frac{2\lambda K(h+p)}{hp}}.Qd∗​=hp2λK(h+p)​​.

Formalization targets

Goal: Theorem 2 (p. 96)

For K>0K > 0K>0, let Qˉ\bar QQˉ​, Qˉ1\bar Q_1Qˉ​1​, Qˉ2\bar Q_2Qˉ​2​ be the positive solutions of

QH0(Q)=2λK,H0(Q)=Hd(Qd∗),∫0QH0(y) dy=λK.Q H_0(Q) = 2\lambda K,\qquad H_0(Q) = H_d(Q^*_d),\qquad \int_0^Q H_0(y)\,dy = \lambda K.QH0​(Q)=2λK,H0​(Q)=Hd​(Qd∗​),∫0Q​H0​(y)dy=λK.

Each has exactly one positive solution, and

Qd∗≤Q∗≤Qˉ,Qˉ≤Qˉ1,Qˉ≤Qˉ2.Q^*_d \le Q^* \le \bar Q,\qquad \bar Q \le \bar Q_1,\qquad \bar Q \le \bar Q_2.Qd∗​≤Q∗≤Qˉ​,Qˉ​≤Qˉ​1​,Qˉ​≤Qˉ​2​.

Moreover, with λ,L,h,p\lambda, L, h, pλ,L,h,p and the demand law fixed, K↦Qˉ1(K)−Qd∗(K)K \mapsto \bar Q_1(K) - Q^*_d(K)K↦Qˉ​1​(K)−Qd∗​(K) is nondecreasing on (0,∞)(0, \infty)(0,∞) and converges to a finite constant as K→∞K \to \inftyK→∞.

Milestones

The milestones are the paper's own numbered results that feed Theorem 2, listed in the order the argument uses them:

  1. Lemma 2 (p. 90): for Q>0Q > 0Q>0, rrr is optimal iff G(r)=G(r+Q)G(r) = G(r + Q)G(r)=G(r+Q).
  2. Eq. (7) (p. 91): C(Q)=(λK+∫0QH(y) dy)/QC(Q) = (\lambda K + \int_0^Q H(y)\,dy)/QC(Q)=(λK+∫0Q​H(y)dy)/Q.
  3. Lemma 4 (p. 91): HHH is increasing and convex with asymptotic slope hp/(h+p)hp/(h+p)hp/(h+p).
  4. Lemma 6 (p. 92): AAA is increasing and convex, and Q=Q∗Q = Q^*Q=Q∗ iff A(Q)=λKA(Q) = \lambda KA(Q)=λK.
  5. Eqs. (18), (20) (p. 94): Hd(Q)=hph+pQH_d(Q) = \frac{hp}{h+p}QHd​(Q)=h+php​Q, and Qd∗Q^*_dQd∗​ is optimal for the EOQ model.
  6. Lemma 7 (p. 95): H0≤Hd≤HH_0 \le H_d \le HH0​≤Hd​≤H and A≤AdA \le A_dA≤Ad​.
  7. Lemma 8 (p. 95): ∫0QH≥12QH(Q)≥A(Q)≥12QH0(Q)≥∫0QH0\int_0^Q H \ge \tfrac12 QH(Q) \ge A(Q) \ge \tfrac12 QH_0(Q) \ge \int_0^Q H_0∫0Q​H≥21​QH(Q)≥A(Q)≥21​QH0​(Q)≥∫0Q​H0​, with equalities for deterministic demand.

Significance

The result. Theorem 2 says that the EOQ formula always underestimates the optimal order quantity when leadtime demand is random. The underestimate is bounded by Qˉ1−Qd∗\bar Q_1 - Q^*_dQˉ​1​−Qd∗​, a quantity that stays bounded however large the ordering cost is. So the relative error of the EOQ quantity vanishes as KKK grows. The first inequality, Qd∗≤Q∗Q^*_d \le Q^*Qd∗​≤Q∗, is also an ingredient of the paper's Theorem 3 (cost bounds) and Theorem 5 (the EOQ quantity raises costs by at most 1/81/81/8). The explicit bounds Qˉ\bar QQˉ​, Qˉ1\bar Q_1Qˉ​1​, Qˉ2\bar Q_2Qˉ​2​ bracket Q∗Q^*Q∗ and give a search interval for it.

Formalizing it. The theorem has been proved on paper since 1992. No machine-checked version of it, or of the continuous-review (Q,r)(Q, r)(Q,r) cost of Eq. (1), exists on this platform. The inventory items already here treat the discrete cost with integer order quantities, a normally distributed demand, or the EOQ without backorders. This mission provides a machine-checked version of the paper's optimality conditions for a general demand distribution. The paper's argument differentiates GGG twice, i.e. it tacitly assumes a density. The formal statements do not, so a formal proof must redo those steps with one-sided (convexity) arguments. The printed argument for the limit in part (b) shows only that a derivative tends to zero. A complete proof of convergence is part of the work.

Difficulty

The obvious route to Qd∗≤Q∗Q^*_d \le Q^*Qd∗​≤Q∗ compares the two cost curves CCC and CdC_dCd​ directly. It fails because C≥CdC \ge C_dC≥Cd​ pointwise, and a pointwise inequality between two convex functions says nothing about the order of their minimizers. The stochastic curve HHH is defined only implicitly, as GGG evaluated at a minimizer of a parametric integral, so its growth relative to the linear HdH_dHd​ has to be established before any comparison of order quantities. For part (b), a vanishing derivative does not imply convergence (log⁡K\log KlogK also has a vanishing derivative), so the printed proof of the limit does not go through as written.

Without a density, r(Q)r(Q)r(Q) need not be differentiable. Every derivative in the paper's proofs (of rrr, HHH and AAA) must be replaced by monotonicity or chord arguments.

Formalization scope

The Lean development uses the namespace ZhengQR.OrderQty. Its conventions:

  • Parameters. λ,L,K,h,p\lambda, L, K, h, pλ,L,K,h,p are reals, all assumed strictly positive. K>0K > 0K>0 is implicit in the paper; at K=0K = 0K=0 the optimal quantity degenerates.
  • Demand. The law μ\muμ of DDD is a probability measure on R\mathbb{R}R that is integrable, has mean λL\lambda LλL and is carried by [0,∞)[0, \infty)[0,∞). No density is assumed, so discrete laws such as the Poisson of the paper's §4 are allowed.
  • Standing assumption. GGG has a unique global minimizer (p. 90). It is a hypothesis of every statement about the stochastic model.
  • Generic machinery. ccc, r(Q)r(Q)r(Q), y0y^0y0, HHH, CCC, AAA, H0H_0H0​ and optimality of QQQ are defined for an arbitrary cost rate GGG and applied to both the newsvendor cost and GdG_dGd​. So Eqs. (18) and (20) are theorems, not definitions. r(Q)r(Q)r(Q) and y0y^0y0 are chosen minimizers, never solutions of Lemma 2's equation. r(Q)r(Q)r(Q) minimizes ∫rr+QG\int_r^{r+Q}G∫rr+Q​G, which for Q>0Q > 0Q>0 has the same minimizers as c(Q,⋅)c(Q, \cdot)c(Q,⋅), so HHH, H0H_0H0​ and AAA do not depend on KKK.
  • Domains. HHH, H0H_0H0​ and AAA are used on [0,∞)[0, \infty)[0,∞), ccc and CCC for Q>0Q > 0Q>0 only, and Q∗Q^*Q∗ is a Q>0Q > 0Q>0 minimizing CCC over (0,∞)(0, \infty)(0,∞).
  • Readings of informal words.
    • Lemma 4's "increasing" and Lemma 6's "increasing/decreasing" mean strictly.
    • Lemma 4's "asymptotic slope hp/(h+p)hp/(h+p)hp/(h+p)" means H(Q)/Q→hp/(h+p)H(Q)/Q \to hp/(h+p)H(Q)/Q→hp/(h+p) together with the chord bound H(Q′)−H(Q)≤hph+p(Q′−Q)H(Q') - H(Q) \le \frac{hp}{h+p}(Q' - Q)H(Q′)−H(Q)≤h+php​(Q′−Q) for 0≤Q<Q′0 \le Q < Q'0≤Q<Q′.
    • "Qˉ=def{Q:… }\bar Q \overset{\text{def}}{=} \{Q : \dots\}Qˉ​=def{Q:…}" means the unique positive solution. The goal quantifies over every positive solution and separately asserts that exactly one exists.
    • Theorem 2's "increasing function of KKK" means nondecreasing, which is what the paper's proof establishes (a nonnegative derivative).
    • "Converges to a constant" means a finite real limit.
    • Lemma 8's "the leadtime demand is deterministic" means the EOQ model with cost rate GdG_dGd​.
  • Ruling out trivial readings. The goal's hypotheses are satisfiable (for example by a deterministic leadtime demand). Existence of Q∗Q^*Q∗ (Lemma 6) and of Qˉ\bar QQˉ​, Qˉ1\bar Q_1Qˉ​1​, Qˉ2\bar Q_2Qˉ​2​ (the goal itself) is asserted, so neither the bounds nor the limit hold vacuously.

Infrastructure needed includes the following. Much of it is reusable for any single-item inventory model:

  • differentiation under the expectation, or one-sided substitutes, for GGG;
  • convexity of HHH as the inverse of the width of the sublevel sets of GGG;
  • the envelope identity behind Eq. (7);
  • elementary convex-analysis facts about chords.

Contributions welcome: proofs of the milestones in any order, general lemmas on the newsvendor cost, and a complete convergence argument for part (b).

Selected references

  • Y.-S. Zheng, On Properties of Stochastic Inventory Systems, Management Science 38(1):87–103, 1992. https://doi.org/10.1287/mnsc.38.1.87
  • A. Federgruen, Y.-S. Zheng, An Efficient Algorithm for Computing an Optimal (r, Q) Policy in Continuous Review Stochastic Inventory Systems, Operations Research 40(4):808–813, 1992. https://doi.org/10.1287/opre.40.4.808
10 thms3 active usersReviewed
Convex OptimizationOperations Research·Captain: mikedeng1

An Exact Duality Theory for Semidefinite Programming and Its Complexity Implications: The Extended Lagrange–Slater Dual Has Zero Duality Gap and Attains Its OptimumResearch Paper

Motivation

Semidefinite programming (SDP) optimizes a linear function over the intersection of the cone of positive semidefinite matrices with an affine subspace. It contains linear programming as the diagonal case and is the computational core of relaxations in combinatorial optimization, control theory and polynomial optimization. Its standard duality theory, however, is weaker than that of linear programming. The Lagrangian dual of an SDP can have a strictly positive duality gap, can fail to attain its optimal value, and an infeasible semidefinite system need not have a certificate of infeasibility of the naive Farkas form. All the classical strong duality theorems for SDP therefore assume a constraint qualification such as Slater's condition (a strictly feasible point).

M. V. Ramana (1997, Math. Program. 77, 129–162) constructed a dual, the Extended Lagrange–Slater Dual (ELSD), whose size is polynomial in the data and which enjoys every property of linear programming duality for every SDP, with no constraint qualification. The same construction yields an exact theorem of the alternative for semidefinite feasibility and the complexity consequence that semidefinite feasibility lies in NP if and only if it lies in co-NP in the Turing model.

Timeline:

  • 1980s–1990s: Lagrangian (Slater-type) duality for SDP, with strong duality under strict feasibility (see e.g. the surveys of Vandenberghe and Boyd, SIAM Rev. 38 (1996)).
  • 1981: Borwein and Wolkowicz, facial reduction for general convex programs, which regularizes a problem by passing to the minimal face containing the feasible set; not of polynomial size in the SDP data (J. Math. Anal. Appl. 83 (1981)).
  • 1997: Ramana, the ELSD, an explicit polynomial-size dual with zero gap and dual attainment for every SDP.
  • 1997: Ramana, Tunçel and Wolkowicz relate the ELSD to facial reduction (SIAM J. Optim. 7 (1997)).

Setting

Let n,mn, mn,m be natural numbers, Mn\mathcal M_nMn​ the space of real n×nn\times nn×n matrices, and Sn⊆Mn\mathcal S_n\subseteq\mathcal M_nSn​⊆Mn​ the symmetric ones. On Mn\mathcal M_nMn​ the inner product is A∙B=∑i,jAijBijA\bullet B = \sum_{i,j}A_{ij}B_{ij}A∙B=∑i,j​Aij​Bij​. For symmetric AAA, A⪰0A\succeq 0A⪰0 means AAA is positive semidefinite. The data are symmetric Q0,Q1,…,Qm∈SnQ_0, Q_1,\dots,Q_m\in\mathcal S_nQ0​,Q1​,…,Qm​∈Sn​ and c∈Rmc\in\mathbb R^mc∈Rm. The primal SDP is

(P)sup⁡ cTxs.t.Q(x):=Q0−∑i=1mxiQi⪰0,(\mathrm P)\qquad \sup\ c^{\mathsf T}x\quad\text{s.t.}\quad Q(x) := Q_0-\sum_{i=1}^m x_iQ_i\succeq 0 ,(P)sup cTxs.t.Q(x):=Q0​−i=1∑m​xi​Qi​⪰0,

with feasible region G={x∣Q(x)⪰0}G = \{x\mid Q(x)\succeq 0\}G={x∣Q(x)⪰0}, a spectrahedron. Define Q∗:Mn→RmQ^*:\mathcal M_n\to\mathbb R^mQ∗:Mn​→Rm by Q∗(U)=(U∙Qi)i=1mQ^*(U) = (U\bullet Q_i)_{i=1}^mQ∗(U)=(U∙Qi​)i=1m​ and write Q#(U)=0Q^\#(U) = 0Q#(U)=0 for "Q0∙U=0Q_0\bullet U = 0Q0​∙U=0 and Q∗(U)=0Q^*(U) = 0Q∗(U)=0".

For k≥1k\ge 1k≥1 let Ck\mathcal C_kCk​ be the set of tuples (Ui,Wi)i=1k(U_i, W_i)_{i=1}^k(Ui​,Wi​)i=1k​ of real n×nn\times nn×n matrices with W0=0W_0 = 0W0​=0 and, for i=1,…,ki = 1,\dots,ki=1,…,k,

Q#(Ui+Wi−1)=0,Ui⪰WiWiT.Q^\#(U_i+W_{i-1}) = 0,\qquad U_i\succeq W_iW_i^{\mathsf T}.Q#(Ui​+Wi−1​)=0,Ui​⪰Wi​WiT​.

The WiW_iWi​ need not be symmetric. Uk\mathcal U_kUk​ and Wk\mathcal W_kWk​ are the sets of last components UkU_kUk​ and WkW_kWk​; W0={0}\mathcal W_0 = \{0\}W0​={0}. The ELSD is

inf⁡ (U+W)∙Q0s.t.Q∗(U+W)=c,W∈Wm,U⪰0,\inf\ (U+W)\bullet Q_0\quad\text{s.t.}\quad Q^*(U+W) = c,\quad W\in\mathcal W_m,\quad U\succeq 0,inf (U+W)∙Q0​s.t.Q∗(U+W)=c,W∈Wm​,U⪰0,

and Weak-ELSD is the same program with Wm−1\mathcal W_{m-1}Wm−1​. For the milestones: the polar G∘={y∣xTy≤1 ∀x∈G}G^\circ = \{y\mid x^{\mathsf T}y\le 1\ \forall x\in G\}G∘={y∣xTy≤1 ∀x∈G}, the algebraic polar G∗={Q∗(U)∣U∙Q0≤1, U⪰0}G^* = \{Q^*(U)\mid U\bullet Q_0\le 1,\ U\succeq 0\}G∗={Q∗(U)∣U∙Q0​≤1, U⪰0}, and Sk=Q∗(Wk)S_k = Q^*(\mathcal W_k)Sk​=Q∗(Wk​).

Formalization targets

Goal: Theorem 6 (Duality Theorem)

For all data (Q0,…,Qm,c)(Q_0,\dots,Q_m,c)(Q0​,…,Qm​,c):

  1. weak duality: cTx≤(U+W)∙Q0c^{\mathsf T}x\le (U+W)\bullet Q_0cTx≤(U+W)∙Q0​ for x∈Gx\in Gx∈G and (U,W)(U,W)(U,W) feasible for ELSD or Weak-ELSD;
  2. if G≠∅G\neq\emptysetG=∅, then sup⁡x∈GcTx<∞\sup_{x\in G}c^{\mathsf T}x<\inftysupx∈G​cTx<∞ iff ELSD is feasible, iff Weak-ELSD is feasible;
  3. if G≠∅G\ne\emptysetG=∅ and ELSD (or Weak-ELSD) is feasible, there is v∈Rv\in\mathbb Rv∈R with
v=sup⁡x∈GcTx=inf⁡ELSD(U+W)∙Q0=inf⁡Weak-ELSD(U+W)∙Q0;v = \sup_{x\in G}c^{\mathsf T}x = \inf_{\mathrm{ELSD}}(U+W)\bullet Q_0 = \inf_{\mathrm{Weak\text{-}ELSD}}(U+W)\bullet Q_0;v=x∈Gsup​cTx=ELSDinf​(U+W)∙Q0​=Weak-ELSDinf​(U+W)∙Q0​;
  1. if G≠∅G\ne\emptysetG=∅ and the primal is bounded, ELSD attains vvv.

Milestones

Propositions 7(vi) and 7(vii) (facts on PSD matrices), Lemma 9 (annihilation Q(x)U=Q(x)W=0Q(x)U = Q(x)W = 0Q(x)U=Q(x)W=0), weak duality over every Wk\mathcal W_kWk​, Lemma 10 (nested subspaces), Lemma 13 (G∘=Cl(G∗)G^\circ = \mathrm{Cl}(G^*)G∘=Cl(G∗)), Corollary 14, Claims 17 and 16, the central Theorem 12,

G∘={Q∗(U+W)∣W∈Wk, U⪰0, U∙Q0≤1}(0∈G, k≥m−1),G^\circ = \{Q^*(U+W)\mid W\in\mathcal W_k,\ U\succeq 0,\ U\bullet Q_0\le 1\}\qquad(0\in G,\ k\ge m-1),G∘={Q∗(U+W)∣W∈Wk​, U⪰0, U∙Q0​≤1}(0∈G, k≥m−1),

the translation invariance of Ck,Uk,Wk\mathcal C_k,\mathcal U_k,\mathcal W_kCk​,Uk​,Wk​ (§2.5), and system (14) (dual attainment at value 0). Theorems 19–21 (Farkas lemma for SDP, optimality condition, primal attainment) are further items stated on the same definitions.

Significance

The Duality Theorem gives SDP a dual with the full strength of linear programming duality for every instance, at polynomial size. Consequences in the paper: an exact theorem of the alternative for semidefinite feasibility (Theorem 19); semidefinite characterizations of optimality of a given point and of primal attainment (Theorems 20, 21); and the complexity results that semidefinite feasibility is in NP iff it is in co-NP in the Turing model and in NP ∩ co-NP in the Blum–Shub–Smale model (Theorem 25, not part of this mission). Theorem 12 separately gives an exact semidefinite description of the polar of any spectrahedron containing the origin.

The results are proved on paper and are classical. No machine-checked version is known to exist; the platform's existing SDP duality theorem assumes Slater's condition. A formalization would provide the first constraint-qualification-free SDP duality in Lean, together with reusable infrastructure on PSD matrices (range inclusion, A∙B=0⇒AB=0A\bullet B = 0\Rightarrow AB = 0A∙B=0⇒AB=0) and on polars of convex sets.

Difficulty

The obvious route to SDP strong duality separates the primal's value from the image of the PSD cone under a linear map and invokes a closed-cone Farkas lemma. That step fails: the linear image of the PSD cone need not be closed, which is exactly why Lagrangian duality has gaps. In this mission the obstruction reappears as the non-closedness of the algebraic polar G∗G^*G∗ (Lemma 13 only gives G∘=Cl(G∗)G^\circ = \mathrm{Cl}(G^*)G∘=Cl(G∗)). The difficulty is to show that finitely many, and at most m−1m-1m−1, corrections by the sets SkS_kSk​ close G∗+SkG^*+S_kG∗+Sk​ (Claims 16, 17), and to control dimensions in doing so. A proof by assuming closedness, strict feasibility or a Slater point is a different theorem.

Formalization scope

Everything lives in the namespace ExactSDPDuality.ELSD, in one definition file. Matrices are Matrix (Fin n) (Fin n) ℝ, vectors Fin m → ℝ; "⪰0\succeq 0⪰0" is Mathlib's PosSemidef (which over ℝ includes symmetry); A∙BA\bullet BA∙B is the entrywise sum on all of Mn\mathcal M_nMn​; cTxc^{\mathsf T}xcTx is the dot product. The data Q0,…,QmQ_0,\dots,Q_mQ0​,…,Qm​ carry symmetry hypotheses in every statement, as the paper assumes throughout. Ck\mathcal C_kCk​ is encoded by sequences U,W:N→MnU, W:\mathbb N\to\mathcal M_nU,W:N→Mn​ with U0=W0=0U_0 = W_0 = 0U0​=W0​=0, so U0=W0={0}\mathcal U_0 = \mathcal W_0 = \{0\}U0​=W0​={0}; for m=0m = 0m=0 the index m−1m-1m−1 is 000. Optimal values are least upper and greatest lower bounds of the value sets, never real sSup/sInf. The polar is the one-sided polar. In §2.4 statements the standing assumption 0∈G0\in G0∈G is a hypothesis. In Claim 16 the index satisfies k+1≤mk+1\le mk+1≤m, the range where Sk+1S_{k+1}Sk+1​ is introduced, and dim⁡Sk\dim S_kdimSk​ is the rank of the span of SkS_kSk​.

Theorems 20 and 21 are printed with Q∗(U+W)=0Q^*(U+W) = 0Q∗(U+W)=0; both are false as printed (counterexamples in the items) and are stated with the corrected Q∗(U+W)=cQ^*(U+W) = cQ∗(U+W)=c that the paper's derivation from Theorem 6 gives.

Trivializing formalizations are ruled out: no Slater or other constraint qualification appears; the dual is the ELSD built from the recursively defined Wm\mathcal W_mWm​, not the Lagrangian dual or an arbitrary subspace; the WiW_iWi​ range over all of Mn\mathcal M_nMn​, not only symmetric matrices (the paper's Example 4 needs a nonsymmetric W2W_2W2​).

Needed infrastructure: PSD matrix facts (Proposition 7), bipolar theorem for closed convex sets containing the origin (Proposition 11), closedness arguments for linear images of cones, and dimension counting of subspaces of Rm\mathbb R^mRm. Contributions of any milestone, of these general lemmas, and of alternative proofs (for instance via facial reduction) are welcome.

Selected references

  • M. V. Ramana, An exact duality theory for semidefinite programming and its complexity implications, Mathematical Programming 77 (1997) 129–162. https://doi.org/10.1007/BF02614433
  • M. V. Ramana, L. Tunçel, H. Wolkowicz, Strong duality for semidefinite programming, SIAM Journal on Optimization 7 (1997) 641–662. https://doi.org/10.1137/S1052623495288350
  • J. M. Borwein, H. Wolkowicz, Regularizing the abstract convex program, Journal of Mathematical Analysis and Applications 83 (1981) 495–530. https://doi.org/10.1016/0022-247X(81)90138-4
  • L. Vandenberghe, S. Boyd, Semidefinite programming, SIAM Review 38 (1996) 49–95. https://doi.org/10.1137/1038003
  • R. T. Rockafellar, Convex Analysis, Princeton University Press, 1970. https://doi.org/10.1515/9781400873173
14 thms3 active usersReviewed
🏆Completed
AnalysisNumerical AnalysisOperations Research·Captain: mikedeng1

A Nonsmooth Version of Newton's Method I: local superlinear convergence of the generalized-Jacobian Newton method at a semismooth regular rootResearch Paper

Motivation

Many problems in optimization and equilibrium modelling reduce to a system of equations F(x)=0F(x) = 0F(x)=0 whose map F:Rn→RnF : \mathbb R^n \to \mathbb R^nF:Rn→Rn is Lipschitz but not differentiable: reformulations of nonlinear complementarity problems through the componentwise minimum or the Fischer–Burmeister function, Karush–Kuhn–Tucker systems of constrained programs, and gradients of augmented Lagrangians all have kinks. Newton's method, xk+1=xk−F′(xk)−1F(xk)x^{k+1} = x^k - F'(x^k)^{-1}F(x^k)xk+1=xk−F′(xk)−1F(xk), is the standard fast local solver for smooth systems, but it needs a derivative at every iterate.

Qi and Sun (Math. Programming 58, 1993) replaced the Jacobian by an arbitrary element of Clarke's generalized Jacobian and showed that the resulting method converges locally superlinearly under a regularity condition they called semismoothness, extending Mifflin's notion for functionals (Mifflin, SIAM J. Control Optim. 15, 1977) to vector-valued maps. This theorem is the foundation of the family of semismooth Newton methods used in complementarity, variational inequalities and PDE-constrained optimization.

Timeline. Robinson (1988) and Pang (Math. OR 15, 1990) studied Newton methods built on B-derivatives, with convergence proved under a strong Fréchet derivative at the solution; Kummer (1988) gave an abstract framework for Newton methods for nonsmooth equations; Qi and Sun (1993) proved local superlinear convergence for the generalized-Jacobian iteration under semismoothness and nonsingularity of ∂F(x∗)\partial F(x^*)∂F(x∗), with order 1+p1+p1+p under ppp-order semismoothness.

Setting

Let F:Rn→RmF : \mathbb R^n \to \mathbb R^mF:Rn→Rm be locally Lipschitz. By Rademacher's theorem FFF is differentiable on a set DFD_FDF​ of full measure; write JF(y)JF(y)JF(y) for the Jacobian at y∈DFy \in D_Fy∈DF​. The generalized Jacobian is

∂F(x)=co{lim⁡i→∞JF(xi):xi→x, xi∈DF},\partial F(x) = \mathrm{co}\Big\{\lim_{i\to\infty} JF(x_i) : x_i \to x,\ x_i \in D_F\Big\},∂F(x)=co{i→∞lim​JF(xi​):xi​→x, xi​∈DF​},

the convex hull of all limits of Jacobians along sequences of differentiability points converging to xxx. The one-sided directional derivative is F′(x;h)=lim⁡t↓0(F(x+th)−F(x))/tF'(x;h) = \lim_{t\downarrow 0}(F(x+th)-F(x))/tF′(x;h)=limt↓0​(F(x+th)−F(x))/t.

FFF is semismooth at xxx if it is Lipschitz near xxx and, for every hhh, the limit of Vh′Vh'Vh′ over V∈∂F(x+th′)V \in \partial F(x+th')V∈∂F(x+th′), h′→hh' \to hh′→h, t↓0t \downarrow 0t↓0 exists. For 0<p≤10 < p \le 10<p≤1, FFF is ppp-order semismooth at xxx if in addition Vh−F′(x;h)=O(∥h∥1+p)Vh - F'(x;h) = O(\|h\|^{1+p})Vh−F′(x;h)=O(∥h∥1+p) for V∈∂F(x+h)V \in \partial F(x+h)V∈∂F(x+h), h→0h \to 0h→0.

For m=nm = nm=n, the nonsmooth Newton method is

xk+1=xk−Vk−1F(xk),Vk∈∂F(xk),(3.2)x^{k+1} = x^k - V_k^{-1}F(x^k), \qquad V_k \in \partial F(x^k), \tag{3.2}xk+1=xk−Vk−1​F(xk),Vk​∈∂F(xk),(3.2)

where any element of ∂F(xk)\partial F(x^k)∂F(xk) may be chosen at each step. A run is a pair of sequences (xk)(x^k)(xk), (Vk)(V_k)(Vk​) with Vk∈∂F(xk)V_k \in \partial F(x^k)Vk​∈∂F(xk) and Vk(xk+1−xk)=−F(xk)V_k(x^{k+1}-x^k) = -F(x^k)Vk​(xk+1−xk)=−F(xk) for all kkk. A root x∗x^*x∗ (F(x∗)=0F(x^*) = 0F(x∗)=0) is regular when every V∈∂F(x∗)V \in \partial F(x^*)V∈∂F(x∗) is nonsingular.

Formalization targets

Goal: Theorem 3.2, local superlinear convergence

Let FFF be locally Lipschitz, F(x∗)=0F(x^*) = 0F(x∗)=0, FFF semismooth at x∗x^*x∗, and every V∈∂F(x∗)V \in \partial F(x^*)V∈∂F(x∗) nonsingular. Then there is δ>0\delta > 0δ>0 such that every V∈∂F(y)V \in \partial F(y)V∈∂F(y) with ∥y−x∗∥<δ\|y - x^*\| < \delta∥y−x∗∥<δ is nonsingular, a Newton step from such a yyy stays within δ\deltaδ of x∗x^*x∗, and every run with ∥x0−x∗∥<δ\|x^0 - x^*\| < \delta∥x0−x∗∥<δ satisfies

xk→x∗,∥xk+1−x∗∥=o(∥xk−x∗∥).x^k \to x^*, \qquad \|x^{k+1} - x^*\| = o(\|x^k - x^*\|).xk→x∗,∥xk+1−x∗∥=o(∥xk−x∗∥).

The goal asserts only the shape of the convergence (superlinear) and fixes no constants.

Stronger: Theorem 3.2, order 1+p1 + p1+p

If moreover FFF is ppp-order semismooth at x∗x^*x∗, 0<p≤10 < p \le 10<p≤1, there are δ>0\delta > 0δ>0 and CCC with

∥xk+1−x∗∥≤C∥xk−x∗∥1+p\|x^{k+1} - x^*\| \le C\|x^k - x^*\|^{1+p}∥xk+1−x∗∥≤C∥xk−x∗∥1+p

for every run started within δ\deltaδ of x∗x^*x∗.

Milestones

The milestones follow the paper's route: Proposition 2.1 (the limit in the definition of semismoothness is the directional derivative), Lemma 2.2 (Lipschitz continuity of F′(x;⋅)F'(x;\cdot)F′(x;⋅) and its realisation by an element of ∂F(x)\partial F(x)∂F(x)), Theorem 2.3 (semismoothness is equivalent to Vh−F′(x;h)=o(∥h∥)Vh - F'(x;h) = o(\|h\|)Vh−F′(x;h)=o(∥h∥) and to the corresponding condition at differentiability points), the Remark's expansion (2.17), Proposition 3.1 (uniform invertibility near a regular point), the order-(1+p)(1+p)(1+p) sentence of Theorem 3.2, and Corollary 2.5 (strong Fréchet differentiability implies semismoothness).

Significance

The theorem gives a locally superlinearly convergent method for Lipschitz equations with no smoothness beyond semismoothness at the root. Convex, smooth and subsmooth functions are semismooth, as are sums and scalar products of semismooth functions (the paper, citing Mifflin), and later work showed that the complementarity and KKT reformulations on which semismooth Newton solvers are built are semismooth as well; the order-(1+p)(1+p)(1+p) variant gives local quadratic convergence for strongly semismooth maps. Mission II of this series treats the paper's global convergence theorem on a ball, and Mission III the semismoothness of augmented Lagrangian gradients, which supplies the application.

The results are proved in the paper. No machine-checked version of the generalized Jacobian, of semismoothness or of the nonsmooth Newton method is known to exist in Mathlib or on this platform; the platform's formalized Newton results concern one-dimensional C2C^2C2 functions (MetodosNumericos.newton_local_convergence) and smooth convex minimization. A complete development would provide the first formal library for Clarke's generalized Jacobian and semismooth maps.

Difficulty

The classical Newton proof compares F(xk)F(x^k)F(xk) with its linearization JF(x∗)(xk−x∗)JF(x^*)(x^k - x^*)JF(x∗)(xk−x∗) and uses continuity of the Jacobian at x∗x^*x∗. Here neither is available: FFF need not be differentiable at x∗x^*x∗ or at any iterate, the element VkV_kVk​ is chosen arbitrarily from a set, and VkV_kVk​ need not be close to any fixed linear map. The comparison has to go through the directional derivative F′(x∗;⋅)F'(x^*; \cdot)F′(x∗;⋅), which is only positively homogeneous, not linear. The analytic content therefore sits in Section 2: showing that semismoothness, defined through a limit over a set-valued map, controls Vh−F′(x;h)Vh - F'(x;h)Vh−F′(x;h) uniformly in the direction, and that F(x+h)−F(x)−F′(x;h)F(x+h) - F(x) - F'(x;h)F(x+h)−F(x)−F′(x;h) is small. Both rest on Clarke's mean-value inclusion and on compactness and upper semicontinuity of ∂F\partial F∂F, none of which is in Mathlib. The superlinear rate also requires a uniform bound on ∥V−1∥\|V^{-1}\|∥V−1∥ in a whole neighbourhood, not just at x∗x^*x∗.

Formalization scope

Everything lives in the namespace NonsmoothNewton.Local. Section 2 results are stated for maps between finite-dimensional real normed spaces E→GE \to GE→G (the paper's Rn→Rm\mathbb R^n \to \mathbb R^mRn→Rm is the Euclidean instance); Section 3 results use EuclideanSpace ℝ (Fin n). Conventions fixed by the Lean statements:

  • JFJFJF is fderiv; the generalized Jacobian is the convex hull (no closure) of limits of fderiv along sequences xi→xx_i \to xxi​→x of differentiability points.
  • F′(x;h)F'(x;h)F′(x;h) is the one-sided limit over t↓0t \downarrow 0t↓0, never the two-sided lineDeriv; its value is a limUnder, used only where existence is a hypothesis or a consequence.
  • Nonsingular means IsUnit in the ring of continuous linear endomorphisms; ∥V−1∥≤C\|V^{-1}\| \le C∥V−1∥≤C is a two-sided inverse of operator norm at most CCC.
  • A run of (3.2) is encoded by the linear equation Vk(xk+1−xk)=−F(xk)V_k(x^{k+1} - x^k) = -F(x^k)Vk​(xk+1−xk)=−F(xk) with Vk∈∂F(xk)V_k \in \partial F(x^k)Vk​∈∂F(xk); all choices of VkV_kVk​ are quantified, and δ\deltaδ is chosen before the run.
  • Pinned asymptotics. The goal's rate is the proof's display (3.3), stated as IsLittleO along atTop; the printed Theorem 3.2 states only well-definedness and convergence. "Order 1+p1+p1+p" is pinned as ∥xk+1−x∗∥≤C∥xk−x∗∥1+p\|x^{k+1}-x^*\| \le C\|x^k-x^*\|^{1+p}∥xk+1−x∗∥≤C∥xk−x∗∥1+p with δ\deltaδ and CCC uniform over runs. Every o(∥h∥)o(\|h\|)o(∥h∥) in (2.8), (2.9) and (2.17) is its ε\varepsilonε–δ\deltaδ form with a non-strict inequality ≤ε∥h∥\le \varepsilon\|h\|≤ε∥h∥, and every O(∥h∥1+p)O(\|h\|^{1+p})O(∥h∥1+p) is an explicit constant and radius.
  • The standing assumptions "FFF locally Lipschitzian" of Sections 2 and 3 are hypotheses of every statement.
  • The strong Fréchet derivative of Corollary 2.5 is Mathlib's HasStrictFDerivAt, which corrects the misprint F(x)F(x)F(x) for F(z)F(z)F(z) in the paper's display (2.16).

A trivializing formalization is ruled out: the update is not written with a junk inverse (which would make a singular step "well defined"), the generalized Jacobian is the paper's nonempty set rather than one that could be empty, and the theorem quantifies over every run rather than asserting that some run converges.

A complete development needs Clarke's mean-value inclusion (2.2), compactness and upper semicontinuity of ∂F\partial F∂F for locally Lipschitz maps (via Rademacher's theorem, available in Mathlib), and perturbation bounds for inverses of linear maps. The generalized-Jacobian and semismoothness layer is reusable beyond this mission, in particular for Missions II and III of this series. Contributions of proofs of any milestone, and of general lemmas about ∂F\partial F∂F, are welcome.

Selected references

  • L. Qi, J. Sun, A nonsmooth version of Newton's method, Mathematical Programming 58 (1993) 353–367. https://doi.org/10.1007/BF01581275
  • F. H. Clarke, Optimization and Nonsmooth Analysis, Wiley, 1983 (SIAM reprint 1990). https://doi.org/10.1137/1.9781611971309
  • R. Mifflin, Semismooth and semiconvex functions in constrained optimization, SIAM Journal on Control and Optimization 15 (1977) 959–972. https://doi.org/10.1137/0315061
  • J.-S. Pang, Newton's method for B-differentiable equations, Mathematics of Operations Research 15 (1990) 311–341. https://doi.org/10.1287/moor.15.2.311
  • J. M. Ortega, W. C. Rheinboldt, Iterative Solution of Nonlinear Equations in Several Variables, Academic Press, 1970 (SIAM reprint 2000). https://doi.org/10.1137/1.9780898719468
14 thms3 active usersReviewed
🏆Completed
CombinatoricsGraph TheoryOperations Research·Captain: mikedeng1

Shortest Connection Networks And Some Generalizations: Construction Principles P1 and P2 Yield a Shortest Spanning Subtree of Every Connected Labelled GraphResearch Paper

Motivation

Connecting a set of terminals by a network of direct links of least total length is one of the oldest problems of combinatorial optimization. R. C. Prim's 1957 paper in the Bell System Technical Journal (DOI) was motivated by the rate structure for Bell System leased-line services, in which the charge for connecting a set of terminals depends on the length of a shortest network connecting them. The paper states two local construction principles, P1 and P2, and shows that any sequence of their applications produces a shortest network, first for points in the plane and then for arbitrary connected labelled graphs with arbitrary real edge lengths. The paper's §V specialization of the principles, growing a single fragment, is what is now called Prim's algorithm, and its §IV statement is the form of the minimum spanning tree theorem used throughout network design, clustering and approximation algorithms.

Timeline. O. Borůvka (1926) solved the problem for an electrical network in Moravia; V. Jarník (1930) gave the single-fragment procedure; J. B. Kruskal (1956, Proc. AMS 7, 48–50) proved that adding globally shortest links avoiding cycles yields a shortest spanning tree; Prim (1957) gave the more permissive principles P1 and P2, which contain both the Jarník procedure and Kruskal's rule as special orders of application; E. W. Dijkstra (1959) rediscovered the single-fragment procedure.

Setting

Let VVV be a finite set of NNN terminals and GGG a simple graph on VVV, the labelled graph whose edges are the possible links. Each edge eee carries a real length w(e)w(e)w(e); lengths may be negative, zero, or tie. For a finite set FFF of links, H(F)H(F)H(F) denotes the graph on VVV whose edges are the links of FFF.

  • A spanning subtree of GGG is a set FFF of edges of GGG such that H(F)H(F)H(F) is a tree on VVV. Its length is ℓw(F)=∑e∈Fw(e)\ell_w(F) = \sum_{e \in F} w(e)ℓw​(F)=∑e∈F​w(e).
  • A shortest spanning subtree (SSS) is a spanning subtree of least length among all spanning subtrees of GGG. Prim's dictionary is "shortest connection network (SCN) ↔ shortest spanning subtree (SSS)". L(G,w)L(G,w)L(G,w) denotes that least length.
  • Given the links FFF made so far, the connected components of H(F)H(F)H(F) are the isolated terminals (one terminal) and isolated fragments (two or more terminals).
  • Principle 1: any isolated terminal ttt can be connected to a nearest neighbor, a GGG-neighbor nnn with w({t,n})≤w({t,m})w(\{t,n\}) \le w(\{t,m\})w({t,n})≤w({t,m}) for all GGG-neighbors mmm of ttt.
  • Principle 2: any isolated fragment CCC can be connected to a nearest neighbor n∉Cn \notin Cn∈/C by a shortest available link {u,n}\{u,n\}{u,n}, u∈Cu \in Cu∈C; equivalently {u,n}\{u,n\}{u,n} is a shortest edge of GGG with one end in CCC and the other outside.
  • A construction is a sequence of links e0,e1,…e_0, e_1, \dotse0​,e1​,…, each an application of P1 or P2 with respect to the links before it. It is complete when it has N−1N-1N−1 links.

Only edges of GGG are possible links; in Prim's distance table a missing edge has length ∞\infty∞.

Formalization targets

Goal (§IV, p. 1396)

For every finite connected graph GGG and every www,

(∃ a complete construction) ∧ (∀ complete constructions e0,…,eN−2: {e0,…,eN−2} is a SSS of G).\Bigl(\exists\ \text{a complete construction}\Bigr) \ \wedge\ \Bigl(\forall\ \text{complete constructions } e_0,\dots,e_{N-2}:\ \{e_0,\dots,e_{N-2}\} \text{ is a SSS of } G\Bigr).(∃ a complete construction) ∧ (∀ complete constructions e0​,…,eN−2​: {e0​,…,eN−2​} is a SSS of G).

This is the sentence "P1 and P2 will provide a SSS for any connected labelled graph with any set of real edge lengths." It fixes nothing about the order of applications, the component chosen, or the tie-breaking.

Milestones

  1. Counting (§II, p. 1392): after any construction with kkk links, H(F)H(F)H(F) is acyclic with N−kN-kN−k components; a complete construction is a spanning subtree; a construction with fewer than N−1N-1N−1 links can be extended.
  2. Necessary Condition 1 (p. 1392): every terminal of a SSS is linked in it to at least one nearest neighbor.
  3. Necessary Condition 2 (p. 1392): every fragment SSS of a SSS, ∅≠S≠V\emptyset \ne S \ne V∅=S=V, is linked in it to a nearest neighbor by a shortest available link.
  4. Distinct lengths (§III, p. 1393): if the edge lengths are pairwise distinct, every link of every construction belongs to every SSS.
  5. Continuity (§III, p. 1394): w↦L(G,w)w \mapsto L(G,w)w↦L(G,w) is continuous.

Significance

The goal is the correctness theorem of a whole family of greedy minimum spanning tree procedures at once: Jarník–Prim (one growing fragment), Kruskal (globally shortest link first) and Borůvka-style interleavings all produce sequences of P1/P2 applications. Because lengths are arbitrary reals, it also covers maximum spanning trees by a sign change (p. 1397) and graphs that are not complete.

The result is classical and fully proved in the literature. What this mission adds is a machine-checked statement in exactly Prim's generality. Mathlib has spanning trees of connected graphs (SimpleGraph.Connected.exists_isTree_le) and the edge count of trees, but no minimum spanning tree theory. Existing Prove2Me items on minimum spanning trees are either restricted to complete graphs with distance matrices or state a cut property in existence form at a single vertex; none states Prim's principles or his necessary conditions.

Difficulty

The obvious argument, "each link P1 or P2 adds belongs to the shortest network", uses a unique shortest network, and that fails with ties: when two links tie, a P1/P2 link need not lie in a given SSS. Prim's own treatment of ties (§III) is an informal perturbation argument; the formal statement must hold for every tie-breaking choice made during a construction, not only for a generic perturbed instance. Negative lengths remove the easy reading "shortest connected spanning subgraph": the minimum must range over trees only. The statements also involve the component structure of H(F)H(F)H(F) as it changes during a construction, and tree paths in an arbitrary, not necessarily complete, graph.

Formalization scope

Namespace ShortestConnection.Principles, Mathlib SimpleGraph. Conventions:

  • VVV is a Fintype with decidable equality; GGG is a SimpleGraph V (at most one link per pair, no loops, which is Prim's setting). Lengths are w : Sym2 V → ℝ; only values on edges of GGG matter.
  • Link sets are Finset (Sym2 V); linkGraph F is SimpleGraph.fromEdgeSet F. A spanning subtree requires ↑F ⊆ G.edgeSet and (linkGraph F).IsTree.
  • An isolated fragment is a whole connected component of linkGraph F; the P2 condition is a single inequality against every GGG-edge leaving it, which is equivalent to "nearest neighbor and shortest link" in Prim's sense.
  • A construction is a List (Sym2 V) checked entrywise against l.take i; complete means length Fintype.card V - 1 (natural subtraction, used only for nonempty VVV).
  • LLL is sInf of the lengths of spanning subtrees; continuity is in the product topology.

Implicit hypotheses made explicit: GGG connected (hence V≠∅V \ne \emptysetV=∅) wherever an SSS or a complete construction is involved; at least two terminals for Necessary Condition 1; SSS nonempty and S≠VS \ne VS=V for Necessary Condition 2; pairwise distinct edge lengths only in milestone 4, as in the paper's temporary assumption.

The goal's existence clause rules out a vacuous formalization in which no complete construction exists; the step predicates are defined from lengths and components only, never through shortest spanning subtrees, and they are not restricted to one growing fragment or to the globally shortest link.

Needed infrastructure: tree exchange (adding an edge to a spanning tree creates one cycle; removing any other cycle edge yields a spanning tree), component counts under edge addition, and minima of finitely many continuous functions. The exchange and counting lemmas are reusable for any matroid-greedy or spanning-tree mission. Contributions of intermediate lemmas, and proofs of the milestones in any order, are welcome.

Selected references

  • R. C. Prim, Shortest Connection Networks And Some Generalizations, Bell System Technical Journal 36 (1957), 1389–1401. https://doi.org/10.1002/j.1538-7305.1957.tb01515.x
  • J. B. Kruskal, On the shortest spanning subtree of a graph and the traveling salesman problem, Proceedings of the AMS 7 (1956), 48–50. https://doi.org/10.1090/S0002-9939-1956-0078686-7
  • V. Jarník, O jistém problému minimálním, Práce Moravské Přírodovědecké Společnosti 6 (1930), 57–63.
  • O. Borůvka, O jistém problému minimálním, Práce Moravské Přírodovědecké Společnosti 3 (1926), 37–58.
  • R. L. Graham, P. Hell, On the history of the minimum spanning tree problem, Annals of the History of Computing 7 (1985), 43–57. https://doi.org/10.1109/MAHC.1985.10011
10 thms3 active usersReviewed
🏆Completed
Linear algebraNumerical AnalysisOperations Research·Captain: mikedeng1

Updating Quasi-Newton Matrices with Limited Storage: The Limited-Storage BFGS Method Reaches the Minimizer of a Strictly Convex Quadratic in at Most n StepsResearch Paper

Motivation

Quasi-Newton methods minimize a smooth function fff on Rn\mathbb{R}^nRn by moving along dk=−Hkgkd_k = -H_k g_kdk​=−Hk​gk​, where gkg_kgk​ is the gradient and HkH_kHk​ is an approximation of the inverse Hessian built from observed gradient differences. The BFGS update is the most widely used way of building HkH_kHk​, but it stores a dense n×nn \times nn×n matrix, which is prohibitive for large nnn.

Nocedal's 1980 paper (Math. Comp. 35, 773–782) proposed keeping only the last mmm correction pairs and rebuilding the matrix from a simple initial matrix H0H_0H0​ at every step. The resulting method, called SQN in the paper, is now known as L-BFGS, and it is the default large-scale unconstrained optimizer in many numerical libraries and in machine learning. The paper's main theoretical claim is that this truncation does not destroy the finite termination of BFGS on quadratics.

Timeline:

  • 1970: Broyden, Fletcher, Goldfarb and Shanno introduce the BFGS update (references [1] and [5] of the paper).
  • 1977: Nazareth relates BFGS to conjugate gradients (Argonne Tech. Memo 282, reference [7]); his form of preconditioned conjugate gradients is the one the paper uses.
  • 1977–1978: Shanno studies the memoryless BFGS update, the case m=1m = 1m=1 (reference [11]; journal version Math. Oper. Res. 3, 1978).
  • 1980: Nocedal defines the special BFGS matrices and the SQN method and states that on quadratics with exact line searches SQN is identical to preconditioned conjugate gradients, hence has quadratic termination.
  • 1989: Liu and Nocedal (Math. Programming 45) study the method, now called L-BFGS, for large-scale problems.
  • 1998: Kolda, O'Leary and Nazareth (SIAM J. Optim. 8) treat limited-memory and update-skipping BFGS variants with exact line searches on quadratics.

Setting

Let AAA be a symmetric positive definite n×nn \times nn×n matrix and b∈Rnb \in \mathbb{R}^nb∈Rn, and let f(x)=12xTAx+bTxf(x) = \tfrac12 x^T A x + b^T xf(x)=21​xTAx+bTx, a strictly convex quadratic with gradient g(x)=Ax+bg(x) = Ax + bg(x)=Ax+b and unique minimizer x∗=−A−1bx^\ast = -A^{-1} bx∗=−A−1b.

Exact line search. Along a direction d≠0d \neq 0d=0 from xxx, the step α=−g(x)Td/dTAd\alpha = -g(x)^T d / d^T A dα=−g(x)Td/dTAd minimizes f(x+αd)f(x + \alpha d)f(x+αd).

BFGS update. For a pair (s,y)(s, y)(s,y) with ρ=1/yTs\rho = 1/y^T sρ=1/yTs and v=I−ρysTv = I - \rho y s^Tv=I−ρysT, the BFGS update of HHH is

Hˉ=vTHv+ρssT.\bar H = v^T H v + \rho s s^T .Hˉ=vTHv+ρssT.

Special BFGS matrices. Fix H0H_0H0​ symmetric positive definite and a number m≥1m \ge 1m≥1 of stored corrections. Given pairs (sj,yj)(s_j, y_j)(sj​,yj​), the special matrix HKH_KHK​ is H0H_0H0​ updated by the pairs j=K−min⁡(K,m),…,K−1j = K - \min(K, m), \dots, K-1j=K−min(K,m),…,K−1, oldest first (the paper's (4)–(5)). Only the mmm most recent pairs enter, and the matrix is rebuilt from H0H_0H0​.

SQN. Starting from x0x_0x0​, with gi=g(xi)g_i = g(x_i)gi​=g(xi​):

di=−Higi,xi+1=xi+αidi,si=xi+1−xi,yi=gi+1−gi,d_i = -H_i g_i, \qquad x_{i+1} = x_i + \alpha_i d_i, \qquad s_i = x_{i+1} - x_i,\quad y_i = g_{i+1} - g_i,di​=−Hi​gi​,xi+1​=xi​+αi​di​,si​=xi+1​−xi​,yi​=gi+1​−gi​,

with αi\alpha_iαi​ the exact step and Hi+1H_{i+1}Hi+1​ the special matrix built from the last min⁡(i+1,m)\min(i+1, m)min(i+1,m) pairs.

PCG with fixed preconditioner H0H_0H0​. d0=−H0g0d_0 = -H_0 g_0d0​=−H0​g0​, xi+1=xi+αidix_{i+1} = x_i + \alpha_i d_ixi+1​=xi​+αi​di​, di+1=−H0gi+1+βi+1did_{i+1} = -H_0 g_{i+1} + \beta_{i+1} d_idi+1​=−H0​gi+1​+βi+1​di​ with βi+1=yiTH0gi+1/yiTdi\beta_{i+1} = y_i^T H_0 g_{i+1} / y_i^T d_iβi+1​=yiT​H0​gi+1​/yiT​di​.

Formalization targets

Goal: quadratic termination of SQN

For every nnn, every symmetric positive definite AAA and H0H_0H0​, every bbb, x0x_0x0​ and every m≥1m \ge 1m≥1,

∃ k≤n:Axk+b=0,\exists\, k \le n : \quad A x_k + b = 0 ,∃k≤n:Axk​+b=0,

where xkx_kxk​ are the SQN iterates. The statement fixes no constant beyond the dimension bound nnn.

Milestones

  1. Property (a): the special matrices are positive definite whenever yiTsi>0y_i^T s_i > 0yiT​si​>0 for all iii.
  2. Eq. (7): along conjugate steps, viyi=0v_i y_i = 0vi​yi​=0 and viyj=yjv_i y_j = y_jvi​yj​=yj​ for i>ji > ji>j.
  3. Eq. (6): along conjugate steps, Hkyj=sjH_k y_j = s_jHk​yj​=sj​ for the mmm most recent jjj (when k>mk > mk>m).
  4. Eq. (10): the special matrix equals mmm sum-form BFGS corrections applied to H0H_0H0​.
  5. Eq. (15): the PCG directions satisfy diTyj=0d_i^T y_j = 0diT​yj​=0 for i≠ji \neq ji=j.
  6. Eq. (16): giTH0gj=0g_i^T H_0 g_j = 0giT​H0​gj​=0 for i≠ji \neq ji=j and giTdj=0g_i^T d_j = 0giT​dj​=0 for j<ij < ij<i.
  7. The PCG with fixed preconditioner H0H_0H0​ reaches the minimizer in at most nnn steps.
  8. SQN and this PCG produce identical iterates and directions at every step.

Significance

The result shows that storing only mmm correction pairs costs nothing on quadratics: for any m≥1m \ge 1m≥1, SQN terminates in at most nnn steps, like full BFGS and conjugate gradients. It explains why L-BFGS with small mmm is competitive, and it is the model case for later analyses of limited-memory methods (their linear convergence on uniformly convex functions, and their relation to Krylov methods). Property (b) is the reason one expects efficiency to grow with mmm: the matrix satisfies the secant equation on the mmm most recent directions.

The claims are classical and generally accepted, but the paper argues them in a few lines ("it is straightforward to show"), deferring the PCG facts (15)–(16) to a reference. No machine-checked proof of the termination of BFGS, L-BFGS or preconditioned conjugate gradients is known to this mission. A formalization would provide a verified model of L-BFGS on quadratics and a reusable development of conjugate-direction methods with a preconditioner.

Difficulty

The obvious route, "SQN is BFGS and BFGS terminates", fails: SQN discards old corrections, so the classical BFGS argument (hereditary secant conditions on all past directions) does not apply once more than mmm steps have been taken. The paper asserts the identity of SQN with preconditioned conjugate gradients in one sentence ("using a similar argument as for the SCG"), and the PCG relations it relies on are quoted from a technical report. The other difficulty is bookkeeping: the window of stored pairs shifts, the matrix is a nested product, and the runs must remain meaningful after the minimizer is reached.

Formalization scope

Vectors are Fin n → ℝ, matrices Matrix (Fin n) (Fin n) ℝ, xTyx^T yxTy is dotProduct, and syTs y^TsyT is Matrix.vecMulVec. Symmetric positive definiteness is Matrix.PosDef. Indices are 0-based, as in the paper. The exact line search is the closed-form step −gTd/dTAd-g^T d / d^T A d−gTd/dTAd. The iterations have no stopping rule: once the gradient vanishes the direction and step are zero and the iterate stays at the minimizer (Lean's 0/0=00/0 = 00/0=0). Past that point the zero pair stored by SQN leaves the BFGS step unchanged. The hypotheses are exactly the paper's: A≻0A \succ 0A≻0, H0≻0H_0 \succ 0H0​≻0, m≥1m \ge 1m≥1 and exact line searches. H0H_0H0​ need not be diagonal.

Two misprints are corrected and flagged in the items: the denominator of β\betaβ in (13) is yi−1Tdi−1y_{i-1}^T d_{i-1}yi−1T​di−1​ (as in (12) and p. 778), and the second relation of (16) is stated for j<ij < ij<i (as used on p. 778), since it fails for i<ji < ji<j.

Ruled out: SQN is defined through its own matrices (4)–(5), rebuilt from H0H_0H0​ and the last mmm pairs. It is not defined through the PCG recurrence, not by one BFGS update of the previous matrix, and not with a stop rule that returns −A−1b-A^{-1}b−A−1b. The standing assumption ykTsk>0y_k^T s_k > 0ykT​sk​>0 is not a hypothesis of any statement about a run (it fails after termination and would make the goal vacuous). With m=0m = 0m=0 SQN is steepest descent and the goal is false, so m≥1m \ge 1m≥1 is required.

Needed infrastructure: algebra of rank-one updates and of Matrix.PosDef under congruence, conjugate-direction lemmas for quadratics, and the fact that n+1n+1n+1 mutually H0H_0H0​-orthogonal vectors in Rn\mathbb{R}^nRn include a zero vector. The PCG results (milestones 5–7) are reusable beyond this mission. Proofs of any milestone, or of the goal directly, are welcome.

Selected references

  • J. Nocedal, Updating Quasi-Newton Matrices with Limited Storage, Mathematics of Computation 35(151), 1980, 773–782. https://doi.org/10.1090/s0025-5718-1980-0572855-7
  • D. F. Shanno, Conjugate gradient methods with inexact searches, Mathematics of Operations Research 3(3), 1978, 244–256. https://doi.org/10.1287/moor.3.3.244
  • L. Nazareth, A Relationship Between the BFGS and Conjugate Gradient Algorithms, ANL-AMD Tech. Memo 282 (rev.), Argonne National Laboratory, 1977 (reference [7] of Nocedal 1980; no online copy located).
  • T. G. Kolda, D. P. O'Leary, L. Nazareth, BFGS with update skipping and varying memory, SIAM Journal on Optimization 8(4), 1998, 1060–1083. https://doi.org/10.1137/S1052623496306450
  • D. C. Liu, J. Nocedal, On the limited memory BFGS method for large scale optimization, Mathematical Programming 45, 1989, 503–528. https://doi.org/10.1007/BF01589116
16 thms3 active usersReviewed
🏆Completed
Operations Research·Captain: mikedeng1

Cubic Regularization of Newton Method and Its Global Performance I: Global Rate of Convergence to Second-Order Stationary PointsResearch Paper

Motivation

Newton's method is the standard second-order algorithm for unconstrained minimization, but without safeguards it has no global guarantee: far from a minimizer the Newton step can increase the objective, and at a point where the Hessian is indefinite the step can head for a saddle point or a maximum. The usual repairs (line search, trust regions, Levenberg–Marquardt damping) come with convergence proofs, but for nonconvex objectives those proofs typically give no rate at all, or only the rate of the gradient method.

Nesterov and Polyak (Math. Program. 108 (2006) 177–205) proposed to regularize the second-order Taylor model of the objective with a cubic term and to take as the next iterate a global minimizer of the regularized model. They showed that the resulting method has a global worst-case rate of convergence to points satisfying the second-order necessary conditions, for every objective with a Lipschitz continuous Hessian and without any convexity. That rate, O(k−2/3)O(k^{-2/3})O(k−2/3) for the gradient norm, is better than the O(k−1/2)O(k^{-1/2})O(k−1/2) of the gradient method. It became the reference point for the complexity theory of nonconvex second-order optimization: adaptive variants (Cartis, Gould and Toint, Math. Program. 127 (2011) 245–295) and lower bounds showing that O(ϵ−3/2)O(\epsilon^{-3/2})O(ϵ−3/2) iterations are optimal among second-order methods (Carmon, Duchi, Hinder and Sidford, Math. Program. 184 (2020) 71–120) are stated against it.

This mission formalizes the general convergence result of that paper, Theorem 1 of Section 3, together with the properties of the cubic step from Section 2 on which it rests.

Setting

Let F⊆RnF \subseteq \mathbb{R}^nF⊆Rn be a closed convex set with nonempty interior, and let fff be twice differentiable on FFF with gradient f′(x)f'(x)f′(x) and Hessian f′′(x)f''(x)f′′(x). A starting point x0∈int⁡Fx_0 \in \operatorname{int} Fx0​∈intF is fixed, and FFF is assumed to contain the level set L(f(x0))={x∈Rn:f(x)≤f(x0)}\mathcal{L}(f(x_0)) = \{x \in \mathbb{R}^n : f(x) \le f(x_0)\}L(f(x0​))={x∈Rn:f(x)≤f(x0​)} in its interior. Assumption 1: the Hessian is Lipschitz continuous on FFF in the spectral norm, ∥f′′(x)−f′′(y)∥≤L∥x−y∥\|f''(x) - f''(y)\| \le L\|x - y\|∥f′′(x)−f′′(y)∥≤L∥x−y∥ for all x,y∈Fx, y \in Fx,y∈F, with L>0L > 0L>0.

For a parameter M>0M > 0M>0 the cubic model of fff at xxx is

mM,x(y)=⟨f′(x),y−x⟩+12⟨f′′(x)(y−x),y−x⟩+M6∥y−x∥3.m_{M,x}(y) = \langle f'(x), y - x\rangle + \tfrac12 \langle f''(x)(y - x), y - x\rangle + \tfrac{M}{6}\|y - x\|^3 .mM,x​(y)=⟨f′(x),y−x⟩+21​⟨f′′(x)(y−x),y−x⟩+6M​∥y−x∥3.

The cubic-regularized Newton step TM(x)T_M(x)TM​(x) is any global minimizer of mM,xm_{M,x}mM,x​ over Rn\mathbb{R}^nRn; it exists because the model is continuous and coercive. Write rM(x)=∥x−TM(x)∥r_M(x) = \|x - T_M(x)\|rM​(x)=∥x−TM​(x)∥ and fˉM(x)=f(x)+min⁡ymM,x(y)\bar f_M(x) = f(x) + \min_y m_{M,x}(y)fˉ​M​(x)=f(x)+miny​mM,x​(y).

The cubic regularization of Newton method (3.3) fixes L0∈(0,L]L_0 \in (0, L]L0​∈(0,L], starts at x0x_0x0​ and, for k≥0k \ge 0k≥0, chooses Mk∈[L0,2L]M_k \in [L_0, 2L]Mk​∈[L0​,2L] such that f(TMk(xk))≤fˉMk(xk)f(T_{M_k}(x_k)) \le \bar f_{M_k}(x_k)f(TMk​​(xk​))≤fˉ​Mk​​(xk​), then sets xk+1=TMk(xk)x_{k+1} = T_{M_k}(x_k)xk+1​=TMk​​(xk​). The choice Mk=LM_k = LMk​=L always passes the test.

Write λn(A)\lambda_n(A)λn​(A) for the smallest eigenvalue of a symmetric matrix AAA. The measure of local optimality is

μM(x)=max⁡{2L+M ∥f′(x)∥, −22L+M λn(f′′(x))}.\mu_M(x) = \max\Big\{ \sqrt{\tfrac{2}{L + M}\,\|f'(x)\|},\ -\tfrac{2}{2L + M}\,\lambda_n(f''(x)) \Big\}.μM​(x)=max{L+M2​∥f′(x)∥​, −2L+M2​λn​(f′′(x))}.

It is nonnegative and vanishes exactly when f′(x)=0f'(x) = 0f′(x)=0 and f′′(x)⪰0f''(x) \succeq 0f′′(x)⪰0.

Formalization targets

Goal: Theorem 1, inequality (3.4)

If f(x)≥f∗f(x) \ge f^*f(x)≥f∗ for all x∈Fx \in Fx∈F, then every run of method (3.3) satisfies, for every k≥1k \ge 1k≥1,

min⁡1≤i≤kμL(xi)≤83⋅(3 (f(x0)−f∗)2k⋅L0)1/3.\min_{1 \le i \le k} \mu_L(x_i) \le \frac{8}{3}\cdot\left(\frac{3\,(f(x_0) - f^*)}{2k\cdot L_0}\right)^{1/3}.1≤i≤kmin​μL​(xi​)≤38​⋅(2k⋅L0​3(f(x0​)−f∗)​)1/3.

The constant 8/38/38/3 and the exponent 1/31/31/3 are the paper's; the statement holds for every admissible choice of the parameters MkM_kMk​ and of the global minimizers xk+1x_{k+1}xk+1​.

Milestones, in attack order

  1. Lemma 1 (2.2): ∥f′(y)−f′(x)−f′′(x)(y−x)∥≤12L∥y−x∥2\|f'(y) - f'(x) - f''(x)(y - x)\| \le \tfrac12 L\|y - x\|^2∥f′(y)−f′(x)−f′′(x)(y−x)∥≤21​L∥y−x∥2 on FFF.
  2. Eq. (2.5): f′(x)+f′′(x)(T−x)+12M∥T−x∥(T−x)=0f'(x) + f''(x)(T - x) + \tfrac12 M\|T - x\|(T - x) = 0f′(x)+f′′(x)(T−x)+21​M∥T−x∥(T−x)=0 for T=TM(x)T = T_M(x)T=TM​(x).
  3. Proposition 1 (2.7): f′′(x)+12MrM(x)I⪰0f''(x) + \tfrac12 M r_M(x) I \succeq 0f′′(x)+21​MrM​(x)I⪰0.
  4. Lemma 2 (2.8): ⟨f′(x),x−TM(x)⟩≥0\langle f'(x), x - T_M(x)\rangle \ge 0⟨f′(x),x−TM​(x)⟩≥0 when f(x)≤f(x0)f(x) \le f(x_0)f(x)≤f(x0​).
  5. Lemma 4 (2.11): f(x)−fˉM(x)≥M12rM(x)3f(x) - \bar f_M(x) \ge \tfrac{M}{12} r_M(x)^3f(x)−fˉ​M​(x)≥12M​rM​(x)3.
  6. Lemma 4 (2.12): for M≥LM \ge LM≥L, TM(x)∈FT_M(x) \in FTM​(x)∈F and f(TM(x))≤fˉM(x)f(T_M(x)) \le \bar f_M(x)f(TM​(x))≤fˉ​M​(x).
  7. Lemma 3 (2.9): ∥f′(TM(x))∥≤12(L+M)rM(x)2\|f'(T_M(x))\| \le \tfrac12(L + M) r_M(x)^2∥f′(TM​(x))∥≤21​(L+M)rM​(x)2 when TM(x)∈FT_M(x) \in FTM​(x)∈F.
  8. Lemma 5: μM(TM(x))≤rM(x)\mu_M(T_M(x)) \le r_M(x)μM​(TM​(x))≤rM​(x).
  9. Theorem 1, first claim: ∑i≥0rMi(xi)3≤12L0(f(x0)−f∗)\sum_{i \ge 0} r_{M_i}(x_i)^3 \le \tfrac{12}{L_0}(f(x_0) - f^*)∑i≥0​rMi​​(xi​)3≤L0​12​(f(x0​)−f∗).
  10. Theorem 1, second claim: lim⁡i→∞μL(xi)=0\lim_{i\to\infty} \mu_L(x_i) = 0limi→∞​μL​(xi​)=0.

Significance

Inequality (3.4) is a global, dimension-free complexity bound for reaching approximate second-order stationarity. It controls both the gradient norm, min⁡1≤i≤k∥f′(xi)∥=O(k−2/3)\min_{1\le i\le k}\|f'(x_i)\| = O(k^{-2/3})min1≤i≤k​∥f′(xi​)∥=O(k−2/3), and the most negative curvature, max⁡{0,−λn(f′′(xi))}=O(k−1/3)\max\{0, -\lambda_n(f''(x_i))\} = O(k^{-1/3})max{0,−λn​(f′′(xi​))}=O(k−1/3), along the best iterate, from a single scalar potential f(x0)−f∗f(x_0) - f^*f(x0​)−f∗. The second claim of Theorem 1 gives the asymptotic counterpart: every limit point satisfies the second-order necessary conditions. Section 4 of the paper derives its faster rates for star-convex and gradient-dominated functions from the same Section 2 lemmas.

The result is proved on paper and widely cited; to our knowledge no machine-checked proof of it or of the Section 2 lemmas exists. A formalization adds a checked statement of the method with its exact constants, and reusable facts about global minimizers of cubic models (Proposition 1 in particular) that the companion missions on star-convex, gradient-dominated and locally quadratic convergence also rely on.

Difficulty

Most steps are short inequalities, but two are not. Proposition 1 is a statement about a global minimizer of a nonconvex function: the first- and second-order conditions of a local minimizer give only f′′(x)+12MrI+M2r(T−x)(T−x)⊤⪰0f''(x) + \tfrac12 M r I + \tfrac{M}{2r}(T - x)(T - x)^\top \succeq 0f′′(x)+21​MrI+2rM​(T−x)(T−x)⊤⪰0, which is weaker. The natural first attempt, "take the second-order optimality condition of the model at TTT", therefore fails. The paper proves it in Section 5.1 through a one-dimensional dual characterization of the minimizer.

The second is Lemma 2's second claim, used for (2.12): showing that TM(x)T_M(x)TM​(x) stays in FFF requires a boundary argument along the segment from xxx to TM(x)T_M(x)TM​(x), since the Taylor bounds are only available inside FFF. The remaining work is calculus in Rn\mathbb{R}^nRn: the integral form of Taylor's theorem for the gradient under a Lipschitz Hessian, and eigenvalue perturbation for the second entry of μ\muμ.

Formalization scope

The space is EuclideanSpace ℝ (Fin n) for arbitrary n : ℕ. The gradient and Hessian are maps g and H with HasGradientAt f (g x) x and HasFDerivAt g (H x) x at every x ∈ F. At boundary points of FFF this asks for two-sided derivatives, a mild strengthening of "twice differentiable on FFF". The Lipschitz condition uses the operator norm, which is the spectral norm. TM(x)T_M(x)TM​(x) is represented by the predicate IsCubicStep (global minimizer of cubicModel), and every lemma is stated for every such minimizer. The run predicate IsCubicNewtonRun is 0-based. It writes fˉMk(xk)\bar f_{M_k}(x_k)fˉ​Mk​​(xk​) as f(xk)f(x_k)f(xk​) plus the model value at xk+1x_{k+1}xk+1​, which is the minimum because xk+1x_{k+1}xk+1​ attains it. λn\lambda_nλn​ is lamMin, the Rayleigh-quotient infimum over the unit sphere, which equals the smallest eigenvalue for the (symmetric) Hessian. The lower bound f∗f^*f∗ is required on FFF only. The minimum over 1≤i≤k1 \le i \le k1≤i≤k is written as the existence of an index attaining the bound.

A stationary point of the cubic model is not an admissible step, and the run must keep the test Mk∈[L0,2L]M_k \in [L_0, 2L]Mk​∈[L0​,2L] and the acceptance test. Replacing the step by any point with f(xk+1)≤f(xk)f(x_{k+1}) \le f(x_k)f(xk+1​)≤f(xk​) makes the goal false, and dropping the square root in μM\mu_MμM​ makes Lemma 5 false. The statements rule out all three. Lemma 5 carries the hypothesis TM(x)∈FT_M(x) \in FTM​(x)∈F, which its printed proof uses and which holds at every iterate.

A complete development needs the Taylor bounds (2.2)–(2.3) for vector-valued derivatives on convex sets, and first- and second-order optimality for the cubic model. It also needs a proof of Proposition 1 (Section 5.1 or any other correct argument) and eigenvalue perturbation via Rayleigh quotients. The cubic-model lemmas and Proposition 1 are reusable across the whole series. Proofs of any milestone, alternative proofs of Proposition 1, and general Mathlib-level lemmas about Rayleigh quotients are welcome.

Selected references

  • Yu. Nesterov and B. T. Polyak, Cubic regularization of Newton method and its global performance, Mathematical Programming, Ser. A 108 (2006) 177–205. https://doi.org/10.1007/s10107-006-0706-8
  • C. Cartis, N. I. M. Gould and Ph. L. Toint, Adaptive cubic regularisation methods for unconstrained optimization. Part I: motivation, convergence and numerical results, Mathematical Programming 127 (2011) 245–295. https://doi.org/10.1007/s10107-009-0286-5
  • Y. Carmon, J. C. Duchi, O. Hinder and A. Sidford, Lower bounds for finding stationary points I, Mathematical Programming 184 (2020) 71–120. https://doi.org/10.1007/s10107-019-01406-y
  • Yu. Nesterov, Introductory Lectures on Convex Optimization: A Basic Course, Kluwer, 2004. https://doi.org/10.1007/978-1-4419-8853-9
16 thms3 active usersReviewed
CombinatoricsOperations ResearchTheoretical Computer Science·Captain: mikedeng1

Worst-Case Performance Bounds for Simple One-Dimensional Packing Algorithms 4: First-Fit Decreasing Uses at Most 71/60 L* + 5 Bins When No Item Exceeds 1/2Research Paper

Motivation

Bin packing asks how to place a list of items with sizes in (0,1](0,1](0,1] into as few unit-capacity bins as possible. It models the cutting of stock material, the packing of files onto tracks of a disc and the assignment of jobs to machines with a common deadline. Deciding the optimum is NP-hard, so in practice simple rules are used, and the question is how far they can stray from the optimum in the worst case.

Johnson, Demers, Ullman, Garey and Graham (SIAM J. Comput. 3(4), 1974) gave the first sharp worst-case bounds for the four classical rules. For First-Fit Decreasing (FFD), the rule that sorts the items into nonincreasing order and then places each into the first bin with room, they announced the bound FFD(L)≤119L∗+4FFD(L)\le\frac{11}{9}L^*+4FFD(L)≤911​L∗+4, whose full proof in Johnson's thesis exceeds 75 pages. To show the method, Section 4 of the paper proves a simpler bound in detail: when no item exceeds 1/21/21/2, FFD uses at most 7160L∗+5\frac{71}{60}L^*+56071​L∗+5 bins. That result is the subject of this mission.

Timeline:

  • 1973: D. S. Johnson's MIT thesis, Near-optimal bin packing algorithms, contains the complete proofs of the 11/911/911/9 and 71/6071/6071/60 bounds.
  • 1974: Johnson, Demers, Ullman, Garey and Graham publish the 71/6071/6071/60 bound for lists in (0,1/2](0,1/2](0,1/2] (Theorem 4.1) with a proof that is complete except for parts of two lemmas, and show by example that 71/6071/6071/60 cannot be lowered.
  • 1985: B. S. Baker gives a shorter proof of the 11/911/911/9 bound for FFD (J. Algorithms 6, 1985).
  • 2007: G. Dósa determines the tight additive constant 6/96/96/9 in the 11/911/911/9 bound (ESCAPE 2007, LNCS 4614).

Setting

A list is a finite sequence L=(a1,…,an)L=(a_1,\dots,a_n)L=(a1​,…,an​) of real numbers in (0,1](0,1](0,1]; values may repeat. A bin has capacity 111; its level is the sum of the numbers in it. The optimum L∗L^*L∗ is the least number of bins into which the elements of LLL can be placed with no bin level exceeding 111.

First-Fit places a1,a2,…a_1,a_2,\dotsa1​,a2​,… in order into bins B1,B2,…B_1,B_2,\dotsB1​,B2​,…, each initially at level 000: aia_iai​ goes into the bin of least index whose level β\betaβ satisfies β≤1−ai\beta\le 1-a_iβ≤1−ai​. First-Fit Decreasing first arranges LLL into nonincreasing order and then runs First-Fit. FFD(L)FFD(L)FFD(L) is the number of bins it uses.

The proof uses a weight WWW on finite sets of elements. For an integer k≥1k\ge1k≥1, xxx is a kkk-piece if x∈(1k+1,1k]x\in(\frac1{k+1},\frac1k]x∈(k+11​,k1​], and a kkk-bin is a bin whose largest element is a kkk-piece. Set w1(x)=⌊1/x⌋−1w_1(x)=\lfloor 1/x\rfloor^{-1}w1​(x)=⌊1/x⌋−1. A pair (x,y)(x,y)(x,y) obeys relation kkk if xxx is a kkk-piece and kx+y≤1kx+y\le1kx+y≤1; then w2(x,y)=w1(x)+k−1kw1(y)w_2(x,y)=w_1(x)+\frac{k-1}{k}w_1(y)w2​(x,y)=w1​(x)+kk−1​w1​(y), and otherwise w2(x,y)=w1(x)+w1(y)w_2(x,y)=w_1(x)+w_1(y)w2​(x,y)=w1​(x)+w1​(y). For a partition π\piπ of XXX into one- and two-element sets, with each pair ordered (earlier, later) in the nonincreasing order,

w12(π)=∑{x}∈πw1(x)+∑(x,y)∈πw2(x,y),W(X)=min⁡πw12(π).w_{12}(\pi)=\sum_{\{x\}\in\pi}w_1(x)+\sum_{(x,y)\in\pi}w_2(x,y),\qquad W(X)=\min_\pi w_{12}(\pi).w12​(π)={x}∈π∑​w1​(x)+(x,y)∈π∑​w2​(x,y),W(X)=πmin​w12​(π).

BASIC is the set of elements of LLL that are kkk-pieces lying in a kkk-bin of the FFD packing of LLL, for some kkk; SURPLUS is the rest of LLL.

Formalization targets

Goal: Theorem 4.1

for every list L⊆(0,12]:FFD(L)≤7160L∗+5.\text{for every list } L\subseteq(0,\tfrac12]:\qquad FFD(L)\le\frac{71}{60}L^*+5 .for every list L⊆(0,21​]:FFD(L)≤6071​L∗+5.

The constants are those printed in the paper. The multiplicative constant 71/6071/6071/60 is best possible.

Milestones

  1. Lemma 3.3 (FFD part): if FFD(L)>rL∗+dFFD(L)>rL^*+dFFD(L)>rL∗+d with r,d≥1r,d\ge1r,d≥1, the list L′L'L′ of the elements of LLL exceeding (r−1)/r(r-1)/r(r−1)/r also has FFD(L′)>rL′∗+dFFD(L')>rL'^*+dFFD(L′)>rL′∗+d.
  2. Claim 4.2.1: for N≥4N\ge4N≥4 and L⊆(1N,12]L\subseteq(\frac1N,\frac12]L⊆(N1​,21​], ∑x∈BASICw1(x)≥FFD(L)−∑j=2N−1j−1j\sum_{x\in\mathrm{BASIC}}w_1(x)\ge FFD(L)-\sum_{j=2}^{N-1}\frac{j-1}{j}∑x∈BASIC​w1​(x)≥FFD(L)−∑j=2N−1​jj−1​.
  3. Claim 4.2.2: for N≥4N\ge4N≥4, L⊆(1N,12]L\subseteq(\frac1N,\frac12]L⊆(N1​,21​] and every partition π\piπ of LLL into one- and two-element sets, w12(π)≥w1(BASIC)−∑j=3N−11jw_{12}(\pi)\ge w_1(\mathrm{BASIC})-\sum_{j=3}^{N-1}\frac1jw12​(π)≥w1​(BASIC)−∑j=3N−1​j1​.
  4. Lemma 4.2: for N≥4N\ge4N≥4 and L⊆(1N,12]L\subseteq(\frac1N,\frac12]L⊆(N1​,21​], W(L)≥FFD(L)−N+2W(L)\ge FFD(L)-N+2W(L)≥FFD(L)−N+2.
  5. Subadditivity: W(X1∪⋯∪Xk)≤∑iW(Xi)W(X_1\cup\dots\cup X_k)\le\sum_i W(X_i)W(X1​∪⋯∪Xk​)≤∑i​W(Xi​).
  6. Lemma 4.3: if X⊆(17,12]X\subseteq(\frac17,\frac12]X⊆(71​,21​] and ∑x∈Xx≤1\sum_{x\in X}x\le1∑x∈X​x≤1, then W(X)≤7160W(X)\le\frac{71}{60}W(X)≤6071​.

A companion item states the Remark after Theorem 4.1: for every N≥1N\ge1N≥1 there is a list with all elements below 1/31/31/3, L∗=60NL^*=60NL∗=60N and FFD(L)=71NFFD(L)=71NFFD(L)=71N.

Significance

Theorem 4.1 shows the weighting-function method in its simplest nontrivial form: a weight whose total is within a constant of the algorithm's bin count, and which no feasible bin can exceed by more than the target ratio. The same method, with more elaborate weights, gives the 11/911/911/9 bound for FFD, and it is the model for later worst-case analyses of packing heuristics. The Remark shows that 71/6071/6071/60 is exact for items in (0,1/2](0,1/2](0,1/2], and the Corollary on p. 322 extends the analysis to the asymptotic ratio RFFDαR^\alpha_{FFD}RFFDα​ when items are bounded by α∈(8/29,1/2]\alpha\in(8/29,1/2]α∈(8/29,1/2].

The source proof is partial. The billing argument behind Claim 4.2.2 is given only when two auxiliary conditions (G1) and (G2) hold ("The more intricate argument here omitted", p. 321), and Lemma 4.3 is checked in four of about seventy-four cases ("leaving the remaining 70-odd, more or less routine, cases to the ambitious reader", p. 321). Complete details are in Johnson's thesis. The theorem itself is established. A formalization therefore gives the first complete, checked proof in a single place. The finite case analysis of Lemma 4.3 is well suited to machine checking. No machine-checked proof of any FFD bound is known to exist.

Difficulty

The obvious weight w1w_1w1​ alone fails. Claim 4.2.1 shows that w1(BASIC)w_1(\mathrm{BASIC})w1​(BASIC) covers the FFD bins, but many sets XXX of elements with sum at most 111 have w1(X)>71/60w_1(X)>71/60w1​(X)>71/60, for example two 222-pieces, a 555-piece and a 666-piece. The pair discounts of w2w_2w2​ repair Lemma 4.3, but they must then be paid for in Lemma 4.2, for every partition. That is Claim 4.2.2: a charge from each discounted pair to distinct SURPLUS elements that are no larger. The charge is straightforward only when no member of a pair obeying relation kkk lies in a bin of type k′<kk'<kk′<k. In general a pair's larger element may already have been charged by a smaller relation, and the paper omits the argument that handles this. Lemma 4.3 is elementary but has many cases, each determined by the piece types in XXX and the relations they obey.

Formalization scope

A list is L : List ℝ with IsList L (0<a≤10<a\le10<a≤1 for each element) in every statement. L∗L^*L∗ is optBins L, the least bbb such that some map from positions to Fin b has every bin sum at most 111. The First-Fit run keeps the nonempty bins as a List (List ℝ), and opens a new bin at the end exactly when no existing bin fits, which is the paper's "least jjj". The fit test is β+a≤1\beta+a\le1β+a≤1. FFD is First-Fit on sortDesc L, the mergeSort into nonincreasing order; ties do not affect the bin count. Indices are 000-based.

W(X)W(X)W(X) sorts XXX into nonincreasing order and minimises w12w_{12}w12​ over the involutions of its positions: fixed points are singletons, and a pair i<σ(i)i<\sigma(i)i<σ(i) is oriented (larger, smaller). The minimum is over a finite nonempty set, so it is attained. BASIC is a set of positions of sortDesc L, and each position's bin is its bin in the final FFD packing. In w2w_2w2​, k=⌊1/x⌋k=\lfloor1/x\rfloork=⌊1/x⌋ is the piece type of the first element. Sums ∑j=2N−1\sum_{j=2}^{N-1}∑j=2N−1​ are over Finset.Icc 2 (N - 1) with N≥4N\ge4N≥4.

The goal's range is (0,1/2](0,1/2](0,1/2]. The restriction to (1/7,1/2](1/7,1/2](1/7,1/2] belongs only to the proof, through Lemma 3.3. Stating the goal for (1/7,1/2](1/7,1/2](1/7,1/2], weakening 71/6071/6071/60 or 555, or making WWW an unattained infimum would each change the theorem. Only the FFD half of Lemma 3.3 is stated. Claim 4.2.1 is stated with Lemma 4.2's standing hypothesis N≥4N\ge4N≥4. The Remark's printed range 0<ε≤5/870<\varepsilon\le5/870<ε≤5/87 is a misprint: its FFD packing needs ε<1/174\varepsilon<1/174ε<1/174, and the companion item states only the existence claim.

Infrastructure needed: a usable API for the First-Fit run (the invariants of the fold, bin levels, the order of bins), a lemma that FFD bins receive items in nonincreasing order, and a decision procedure for Lemma 4.3's case analysis over piece types. The model file and the weight file are reusable for the 11/911/911/9 bound (mission 3 of this series) and for the bounded-α\alphaα corollaries. Contributions of proofs of Lemma 4.3 by computer-checked case enumeration, and of the missing general case of Claim 4.2.2, are especially welcome.

Selected references

  • D. S. Johnson, A. Demers, J. D. Ullman, M. R. Garey, R. L. Graham, Worst-Case Performance Bounds for Simple One-Dimensional Packing Algorithms, SIAM J. Comput. 3(4):299–325, 1974. https://doi.org/10.1137/0203025
  • D. S. Johnson, Near-Optimal Bin Packing Algorithms, Ph.D. thesis, Massachusetts Institute of Technology, 1973 (reference [8] of the paper).
  • B. S. Baker, A new proof for the first-fit decreasing bin-packing algorithm, J. Algorithms 6, 1985.
  • G. Dósa, The tight bound of first fit decreasing bin-packing algorithm is FFD(I) ≤ 11/9 OPT(I) + 6/9, ESCAPE 2007, Lecture Notes in Computer Science 4614, 2007.
9 thms3 active usersReviewed
🏆Completed
Linear algebraNumerical AnalysisOperations Research·Captain: mikedeng1

Methods of Conjugate Gradients for Solving Linear Systems II: Each Conjugate Gradient Step Shortens the Error VectorResearch Paper

Motivation

The conjugate gradient method (cg-method) of Hestenes and Stiefel is the standard iterative solver for linear systems Ax=kAx=kAx=k with a symmetric positive definite matrix AAA. It is used for the large sparse systems of finite-element and finite-difference discretizations, as the inner solver of Newton-type and interior-point methods in optimization, and as the prototype of the Krylov subspace methods. Its original 1952 paper (Hestenes and Stiefel, J. Res. NBS 49(6), 1952) already presented it as two things at once: a direct method that reaches the exact solution in at most nnn steps, and a method of successive approximations whose intermediate estimates are useful in their own right.

The second view needs a guarantee that the intermediate estimates actually approach the solution. The method is built to decrease the AAA-weighted error f(x)=(h−x,A(h−x))f(x)=(h-x,A(h-x))f(x)=(h−x,A(h−x)), and the residual ∣k−Axi∣|k-Ax_i|∣k−Axi​∣ need not decrease (Section 18 of the paper, p. 432, notes that it can increase at every step). Theorem 6:3 of the paper supplies the guarantee in the plain Euclidean length: the distance ∣h−xi∣|h-x_i|∣h−xi​∣ from the estimate to the solution decreases strictly at every step, by an exactly computable amount. This mission formalizes that theorem together with the relations from Sections 5 and 6 of the paper on which its proof rests.

Timeline:

  • 1952: Hestenes and Stiefel introduce the method and prove, in one paper, finite termination (Theorems 4:2 and 5:2), the monotone decrease of the error function fff (Theorem 6:1), and the monotone decrease of the Euclidean error (Theorem 6:3). The later literature on cg as an iterative method for large sparse systems takes these properties as its starting point.

Setting

Let AAA be a real n×nn\times nn×n matrix that is symmetric and positive definite, let k∈Rnk\in\mathbb{R}^nk∈Rn, and let hhh be the solution of Ah=kAh=kAh=k. Write (x,y)=x1y1+⋯+xnyn(x,y)=x_1y_1+\cdots+x_ny_n(x,y)=x1​y1​+⋯+xn​yn​ and ∣x∣2=(x,x)|x|^2=(x,x)∣x∣2=(x,x). From an arbitrary starting point x0x_0x0​, the cg-method (5:1) computes estimates xix_ixi​, residuals rir_iri​ and direction vectors pip_ipi​ by

p0=r0=k−Ax0,ai=∣ri∣2(pi,Api),xi+1=xi+aipi,ri+1=ri−aiApi,bi=∣ri+1∣2∣ri∣2,pi+1=ri+1+bipi.p_0=r_0=k-Ax_0,\quad a_i=\frac{|r_i|^2}{(p_i,Ap_i)},\quad x_{i+1}=x_i+a_ip_i,\quad r_{i+1}=r_i-a_iAp_i,\quad b_i=\frac{|r_{i+1}|^2}{|r_i|^2},\quad p_{i+1}=r_{i+1}+b_ip_i .p0​=r0​=k−Ax0​,ai​=(pi​,Api​)∣ri​∣2​,xi+1​=xi​+ai​pi​,ri+1​=ri​−ai​Api​,bi​=∣ri​∣2∣ri+1​∣2​,pi+1​=ri+1​+bi​pi​.

The error vector of xix_ixi​ is yi=h−xiy_i=h-x_iyi​=h−xi​. The error function (4:5) is f(x)=(h−x,A(h−x))f(x)=(h-x,A(h-x))f(x)=(h−x,A(h−x)), which is nonnegative and vanishes only at x=hx=hx=h. The Rayleigh quotient (4:12) of a vector z≠0z\neq 0z=0 is μ(z)=(z,Az)/∣z∣2\mu(z)=(z,Az)/|z|^2μ(z)=(z,Az)/∣z∣2. The Lean development names these cgIter A k x₀ i (with fields .x, .r, .p), cgAlpha for aia_iai​, errorFun A h x and rayleigh A z.

Formalization targets

Goal: Theorem 6:3

For every step that the method performs, that is, every iii with ri≠0r_i\neq 0ri​=0,

∣yi∣2−∣yi+1∣2=f(xi+1)+f(xi)μ(pi)and∣yi+1∣<∣yi∣.|y_i|^2-|y_{i+1}|^2=\frac{f(x_{i+1})+f(x_i)}{\mu(p_i)}\qquad\text{and}\qquad |y_{i+1}|<|y_i| .∣yi​∣2−∣yi+1​∣2=μ(pi​)f(xi+1​)+f(xi​)​and∣yi+1​∣<∣yi​∣.

The paper writes the step from xi−1x_{i-1}xi−1​ to xix_ixi​; the Lean statement shifts the index by one. The goal holds for every dimension nnn, every symmetric positive definite AAA, every kkk and every x0x_0x0​.

Milestones

  1. Theorems 4:2 and 5:2: some m≤nm\le nm≤n has xm=hx_m=hxm​=h.
  2. Theorem 5:3, (5:6a): (pi,pj)=∣rj∣2∣pi∣2/∣ri∣2(p_i,p_j)=|r_j|^2|p_i|^2/|r_i|^2(pi​,pj​)=∣rj​∣2∣pi​∣2/∣ri​∣2 for i≤ji\le ji≤j.
  3. Theorem 6:1, (6:1): f(xi)−f(xi+1)=ai∣ri∣2=μ(pi)∣xi−xi+1∣2f(x_i)-f(x_{i+1})=a_i|r_i|^2=\mu(p_i)|x_i-x_{i+1}|^2f(xi​)−f(xi+1​)=ai​∣ri​∣2=μ(pi​)∣xi​−xi+1​∣2.
  4. Theorem 6:1, (6:2): f(xi)−f(xj)=∑l=ij−1al∣rl∣2f(x_i)-f(x_j)=\sum_{l=i}^{j-1}a_l|r_l|^2f(xi​)−f(xj​)=∑l=ij−1​al​∣rl​∣2 for i<ji<ji<j.
  5. Section 6, (6:6): (yi+1,xi+1−xi)=f(xi+1)/μ(pi)(y_{i+1},x_{i+1}-x_i)=f(x_{i+1})/\mu(p_i)(yi+1​,xi+1​−xi​)=f(xi+1​)/μ(pi​).

Significance

The theorem is what makes an early stop of the cg-method safe in the norm a user usually cares about. Every intermediate estimate is closer to the solution, in Euclidean distance, than the previous one, and the identity (6:5) states by how much. It also separates the cg-method from methods that minimize the residual: the AAA-norm error, the Euclidean error and the residual behave differently, and only the first two are monotone along cg.

The results are proved in the 1952 paper. They have not been formalized: the Prove2Me library has no statement of the conjugate gradient recursion (5:1), and Mathlib has none either. What this mission adds is a machine-checked version of the paper's Section 6 argument for the recursion exactly as printed, including the case analysis at termination that the paper leaves implicit, and a reusable Lean definition of the cg iteration with its basic identities.

Difficulty

The obvious argument does not reach the conclusion. The method decreases f(x)=(y,Ay)f(x)=(y,Ay)f(x)=(y,Ay) at every step, but a decrease in this AAA-weighted norm does not imply a decrease in the Euclidean norm: for a single step along an arbitrary direction, even the best step for fff can lengthen the Euclidean error. So the theorem cannot be proved one step at a time from the local minimization property. It depends on how the current direction relates to all the later directions of the same run, and those relations in turn rest on the mutual orthogonality of the residuals and the conjugacy of the directions, which are established by an induction over the whole run.

A second difficulty is bookkeeping at the end of the run. The recursion divides by ∣ri∣2|r_i|^2∣ri​∣2 and by (pi,Api)(p_i,Ap_i)(pi​,Api​), which vanish after termination. Every milestone has to hold, or be guarded, past that point, and the goal needs the hypothesis ri≠0r_i\neq 0ri​=0 exactly because the strict inequality fails once xi=hx_i=hxi​=h.

Formalization scope

Vectors are Fin n → ℝ, the scalar product is dotProduct (⬝ᵥ), AxAxAx is Matrix.mulVec (*ᵥ), and the standing assumption is A.PosDef, which in Mathlib includes symmetry. The solution hhh is a variable with the hypothesis A *ᵥ h = k. Indices are 0-based. The cg recursion is the definition cgIter, which computes (5:1b)–(5:1f) literally and in order; it has no stopping rule, and Lean's convention t/0=0t/0=0t/0=0 makes it stay at hhh with ri=pi=0r_i=p_i=0ri​=pi​=0 once rm=0r_m=0rm​=0. Lengths appear squared, as (y,y)(y,y)(y,y). The milestones are stated for every index without a termination guard, because both sides of each identity vanish after termination; only the goal carries ri≠0r_i\neq 0ri​=0.

Two formalizations would trivialize the goal and are ruled out. The goal does not assume termination (xm=hx_m=hxm​=h) or any bound on iii: it quantifies over every cg run and every step that takes place. And it is about the Euclidean length ∣h−xi∣|h-x_i|∣h−xi​∣, not the error function fff (that is Theorem 6:1, a different and weaker statement) and not the residual.

A complete development needs Theorem 5:1 (orthogonality of residuals, conjugacy of directions) for the literal recursion, the identities (5:2) and (5:3c), and the positivity of (p,Ap)(p,Ap)(p,Ap) for p≠0p\neq 0p=0. These are reusable for any further work on the cg-method, including the sister mission on finite termination. Proofs of any milestone, and alternative proofs of Theorem 6:3 through the Krylov-subspace characterization, are welcome.

Selected references

  • M. R. Hestenes and E. Stiefel, Methods of Conjugate Gradients for Solving Linear Systems, J. Res. Natl. Bur. Stand. 49(6), 409–436, 1952. https://doi.org/10.6028/jres.049.044 (publisher's scan: https://nvlpubs.nist.gov/nistpubs/jres/049/jresv49n6p409_A1b.pdf)
9 thms3 active usersReviewed
🏆Completed
Linear algebraNumerical AnalysisOperations Research·Captain: mikedeng1

Methods of Conjugate Gradients for Solving Linear Systems I: Finite Termination of the Conjugate Gradient MethodResearch Paper

Motivation

Solving a linear system Ax=kAx = kAx=k with a large symmetric positive definite matrix AAA is a basic task of scientific computing: it arises from discretized elliptic equations, least-squares problems and the Newton steps of optimization methods. In 1952 Magnus Hestenes and Eduard Stiefel published the conjugate gradient method (cg-method) (J. Res. Natl. Bur. Stand. 49 (1952) 409–436). The method uses AAA only through matrix–vector products and stores a few vectors. The paper's abstract states its central property in one sentence: "The solution is given in nnn steps."

The paper obtains this property from a more general scheme, the method of conjugate directions (cd-method), which also contains Gaussian elimination as a special case. Its argument has two parts. Every cd-method with nonzero directions reaches the solution within nnn steps (Theorem 4:2). The cg-method is a cd-method (Theorem 5:2), which follows from the orthogonality and conjugacy relations of Theorem 5:1. This mission formalizes that chain.

Setting

Vectors are real nnn-tuples, with scalar product (x,y)=x1y1+⋯+xnyn(x, y) = x_1y_1 + \cdots + x_ny_n(x,y)=x1​y1​+⋯+xn​yn​ and squared length ∣x∣2=(x,x)|x|^2 = (x, x)∣x∣2=(x,x). The matrix AAA is real, n×nn \times nn×n, symmetric and positive definite, which is the paper's standing assumption (p. 410). The solution hhh satisfies Ah=kAh = kAh=k. The residual of an estimate xxx is r=k−Axr = k - Axr=k−Ax. Two vectors x,yx, yx,y are conjugate when (x,Ay)=0(x, Ay) = 0(x,Ay)=0.

The cg-method (eq. (3:1), p. 411) starts from an arbitrary estimate x0x_0x0​ and sets p0=r0=k−Ax0p_0 = r_0 = k - Ax_0p0​=r0​=k−Ax0​. Given xix_ixi​, rir_iri​, pip_ipi​, it computes

ai=∣ri∣2(pi,Api),xi+1=xi+aipi,ri+1=ri−aiApi,bi=∣ri+1∣2∣ri∣2,pi+1=ri+1+bipi.a_i = \frac{|r_i|^2}{(p_i, Ap_i)},\quad x_{i+1} = x_i + a_i p_i,\quad r_{i+1} = r_i - a_i Ap_i,\quad b_i = \frac{|r_{i+1}|^2}{|r_i|^2},\quad p_{i+1} = r_{i+1} + b_i p_i.ai​=(pi​,Api​)∣ri​∣2​,xi+1​=xi​+ai​pi​,ri+1​=ri​−ai​Api​,bi​=∣ri​∣2∣ri+1​∣2​,pi+1​=ri+1​+bi​pi​.

In Lean the iterates are cgIter A k x₀ i, a structure with fields x, r, p. The step length is cgA.

The cd-method (Section 4, p. 412) chooses an arbitrary first direction p0p_0p0​ and then sets xi+1=xi+aipix_{i+1} = x_i + a_i p_ixi+1​=xi​+ai​pi​ with ai=(pi,ri)/(pi,Api)a_i = (p_i, r_i)/(p_i, Ap_i)ai​=(pi​,ri​)/(pi​,Api​) and ri=k−Axir_i = k - Ax_iri​=k−Axi​. Each new direction pi+1p_{i+1}pi+1​ may be any vector conjugate to p0,…,pip_0, \dots, p_ip0​,…,pi​. Because the directions are free, a cd-run is a property of sequences: IsCDRun A k x r p.

Formalization targets

Goal: finite termination of the cg-method

For every nnn, every symmetric positive definite AAA, every kkk and hhh with Ah=kAh = kAh=k, and every initial estimate x0x_0x0​,

∃ m≤n:xm=h,\exists\, m \le n:\quad x_m = h,∃m≤n:xm​=h,

where xmx_mxm​ is the mmm-th cg iterate. This is the statement of the abstract and of Section 3 (p. 410): "one will reach an estimate xmx_mxm​ (m≤nm \le nm≤n) at which rm=0r_m = 0rm​=0. This estimate is the desired solution hhh."

Milestones

  1. Theorem 4:1 (p. 412). For every cd-run, the directions are mutually conjugate (4:3a). The residual rir_iri​ is orthogonal to p0,…,pi−1p_0, \dots, p_{i-1}p0​,…,pi−1​ (4:3b). The products (pi,rj)(p_i, r_j)(pi​,rj​) are the same for all j≤ij \le ij≤i (4:3c). Hence ai=(pi,r0)/(pi,Api)a_i = (p_i, r_0)/(p_i, Ap_i)ai​=(pi​,r0​)/(pi​,Api​) (4:4).
  2. Theorem 4:2 (p. 412). Every cd-run whose directions p0,…,pn−1p_0, \dots, p_{n-1}p0​,…,pn−1​ are nonzero has xm=hx_m = hxm​=h for some m≤nm \le nm≤n.
  3. Theorem 5:1 (p. 414), in four items. For the cg-method:
    • (5:3a) (ri,rj)=0(r_i, r_j) = 0(ri​,rj​)=0 for i≠ji \ne ji=j;
    • (5:3b) (pi,Apj)=0(p_i, Ap_j) = 0(pi​,Apj​)=0 for i≠ji \ne ji=j;
    • (5:3c) (pi,rj)=0(p_i, r_j) = 0(pi​,rj​)=0 for i<ji < ji<j and (pi,rj)=∣ri∣2(p_i, r_j) = |r_i|^2(pi​,rj​)=∣ri​∣2 for i≥ji \ge ji≥j;
    • (5:3d) (ri,Api)=(pi,Api)(r_i, Ap_i) = (p_i, Ap_i)(ri​,Api​)=(pi​,Api​), and (ri,Apj)=0(r_i, Ap_j) = 0(ri​,Apj​)=0 for i≠j,j+1i \ne j, j + 1i=j,j+1.
  4. Theorem 5:5, eq. (5:10) (p. 416):
ai=∣ri∣2(pi,Api)=(pi,ri)(pi,Api)=(pi,r0)(pi,Api).a_i = \frac{|r_i|^2}{(p_i, Ap_i)} = \frac{(p_i, r_i)}{(p_i, Ap_i)} = \frac{(p_i, r_0)}{(p_i, Ap_i)}.ai​=(pi​,Api​)∣ri​∣2​=(pi​,Api​)(pi​,ri​)​=(pi​,Api​)(pi​,r0​)​.
  1. Theorem 5:2, first sentence (p. 415). The cg-method is a cd-method: its iterates satisfy IsCDRun.

Significance

Finite termination is the property that distinguishes the conjugate gradient method from stationary iterations such as Jacobi or Gauss–Seidel. It explains why the method can be used as a direct solver in exact arithmetic. It is the starting point for the later theory of Krylov subspace methods. The relations of Theorem 5:1 are the checks the paper recommends for monitoring a computation (eq. (3:3)). They are also the input to the paper's further results on the monotone decrease of the error (Section 6) and on the connection with orthogonal polynomials (Sections 14–18).

The result has been proved since 1952 and appears in every numerical linear algebra textbook. As far as a search of the Prove2Me catalog shows (September 2026), it has no machine-checked proof there. The only conjugate gradient material on the platform is a Hilbert-space convergence result for a different recurrence. This mission adds a faithful formal version of the original recursion (3:1) and of the cd-method. It also adds the complete termination argument, organized as in the paper. The definitions and the Theorem 5:1 relations are reusable by any later formalization of Krylov methods, including the second mission of this series on the decrease of the error ∣h−xi∣|h - x_i|∣h−xi​∣.

Difficulty

The obvious argument says that the residuals are mutually orthogonal, so at most nnn of them are nonzero. That argument is only as good as the orthogonality, and Theorem 5:1 must be established by a simultaneous induction over four families of relations. The recursion defines ri+1r_{i+1}ri+1​ by an update, not as k−Axi+1k - Ax_{i+1}k−Axi+1​, so even ri=k−Axir_i = k - Ax_iri​=k−Axi​ needs a proof. Orthogonality of ri+1r_{i+1}ri+1​ to the earlier residuals needs the conjugacy of the earlier directions, and conjugacy of pi+1p_{i+1}pi+1​ needs the orthogonality of the earlier residuals. Neither family can be proved first.

A second obstacle is the passage from orthogonality to termination. A cd-run with a zero direction stalls, so Theorem 4:2 needs the directions p0,…,pn−1p_0, \dots, p_{n-1}p0​,…,pn−1​ to be nonzero. The cg directions become zero exactly when the solution is reached, so applying Theorem 4:2 to cg needs a case split at the first vanishing residual.

Formalization scope

Vectors are Fin n → ℝ, the scalar product is dotProduct (⬝ᵥ), and AxAxAx is Matrix.mulVec (*ᵥ). The hypothesis on AAA is Mathlib's Matrix.PosDef, which includes symmetry. The solution enters only through the hypothesis A *ᵥ h = k. Indices are 0-based, as in the paper.

The cg iteration is total and has no stopping test. Once rm=0r_m = 0rm​=0, Lean's convention x/0=0x/0 = 0x/0=0 gives pm=0p_m = 0pm​=0 and am=0a_m = 0am​=0. From then on the iteration stays at xmx_mxm​, with zero residuals and directions. For this reason the relations of Theorem 5:1 are stated for all indices without guards: after termination they hold trivially.

Two formulations would make the goal trivial, and neither is used. One is a stopping test or step that refers to hhh or to A−1kA^{-1}kA−1k. The other replaces (3:1b), (3:1d) or (3:1e) by the equivalent formulas (3:2a), (3:2b) or ri+1=k−Axi+1r_{i+1} = k - Ax_{i+1}ri+1​=k−Axi+1​, which would move Theorem 5:5 and part of Theorem 5:2 into the definition. The iteration is (3:1) literally.

The cd-method's implicit hypothesis, that the directions p0,…,pn−1p_0, \dots, p_{n-1}p0​,…,pn−1​ are nonzero, is an explicit binder of Theorem 4:2. Without it the statement fails (take p0=0p_0 = 0p0​=0). The hypothesis is satisfiable: the cg run with nonzero residuals is one example.

A complete development needs the standard facts that mutually conjugate nonzero vectors are linearly independent and that nnn independent vectors span Rn\mathbb{R}^nRn. It also needs the induction behind Theorem 5:1. Contributions welcome beyond the milestones include the converse half of Theorem 5:2, the relation (5:2) expressing pkp_kpk​ through r0,…,rkr_0, \dots, r_kr0​,…,rk​, and Theorem 4:5 (the cd-method computes A−1A^{-1}A−1).

Selected references

  • M. R. Hestenes and E. Stiefel, Methods of Conjugate Gradients for Solving Linear Systems, Journal of Research of the National Bureau of Standards 49(6), 409–436, 1952. https://doi.org/10.6028/jres.049.044
  • L. Fox, H. D. Huskey and J. H. Wilkinson, Notes on the solution of algebraic linear simultaneous equations, Quarterly Journal of Mechanics and Applied Mathematics 1(1), 149–173, 1948 (the cd-method from a different point of view; cited in the paper's Section 4 footnote). https://doi.org/10.1093/qjmam/1.1.149
  • G. H. Golub and C. F. Van Loan, Matrix Computations, 4th ed., Johns Hopkins University Press, 2013, §11.3 (textbook account of the method).
11 thms3 active usersReviewed
🏆Completed
Convex OptimizationMachine LearningOperations Research·Captain: mikedeng1

Logarithmic Regret Algorithms for Online Convex Optimization 4: Logarithmic Regret of Exponentially Weighted Online OptimizationResearch Paper

Motivation

In online convex optimization a player repeatedly chooses a point xtx_txt​ from a convex set P⊆RnP \subseteq \mathbb{R}^nP⊆Rn, after which an adversary reveals a convex cost function ftf_tft​ and the player pays ft(xt)f_t(x_t)ft​(xt​). The player's regret after TTT rounds is its total cost minus the cost of the best fixed point in hindsight. The model covers online portfolio selection, online regression and prediction with expert advice, and it underlies the analysis of stochastic and adaptive optimization methods (Zinkevich 2003; Cesa-Bianchi and Lugosi 2006).

For general convex costs the best achievable regret is of order T\sqrt{T}T​. Hazan, Agarwal and Kale (Mach Learn 69, 2007) showed that a curvature condition, α\alphaα-exp-concavity, brings the regret down to order log⁡T\log TlogT, and gave several algorithms that achieve it. This mission concerns the simplest of them, Exponentially Weighted Online Optimization (EWOO), which needs nothing beyond exp-concavity: no bound on gradients and no bound on the diameter of PPP.

Timeline.

  • 1991: Cover's universal portfolio algorithm attains regret O(nlog⁡T)O(n \log T)O(nlogT) for online portfolio selection, whose log-loss is 111-exp-concave (Cover 1991).
  • 1997: Blum and Kalai give a short analysis of the universal portfolio with transaction costs, using a shrinking argument around the best portfolio (Blum and Kalai 1997/1999).
  • 2003: Kalai and Vempala give a polynomial-time randomized implementation of Cover's algorithm via random walks (JMLR 3, 2003).
  • 2007: Hazan, Agarwal and Kale state EWOO for general α\alphaα-exp-concave costs and prove the regret bound of Theorem 7, alongside the Online Newton Step and Follow the Approximate Leader.

Setting

Fix n≥0n \ge 0n≥0 and a set P⊆RnP \subseteq \mathbb{R}^nP⊆Rn that is nonempty, closed, bounded and convex, with positive Lebesgue volume vol(P)\mathrm{vol}(P)vol(P). Fix α>0\alpha > 0α>0. The cost functions are f1,f2,⋯:Rn→Rf_1, f_2, \dots : \mathbb{R}^n \to \mathbb{R}f1​,f2​,⋯:Rn→R, each continuous on PPP and α\alphaα-exp-concave on PPP: the function ht(x)=e−αft(x)h_t(x) = e^{-\alpha f_t(x)}ht​(x)=e−αft​(x) is concave on PPP (LogRegretOCO.EWOO.IsExpConcave).

EWOO keeps the weights

wt(x)=exp⁡(−α∑τ=1t−1fτ(x))=∏τ=1t−1hτ(x),w_t(x) = \exp\Bigl(-\alpha \sum_{\tau=1}^{t-1} f_\tau(x)\Bigr) = \prod_{\tau=1}^{t-1} h_\tau(x),wt​(x)=exp(−ατ=1∑t−1​fτ​(x))=τ=1∏t−1​hτ​(x),

and on round ttt plays the wtw_twt​-weighted mean of PPP,

xt=∫Px wt(x) dx∫Pwt(x) dxx_t = \frac{\int_P x\, w_t(x)\, dx}{\int_P w_t(x)\, dx}xt​=∫P​wt​(x)dx∫P​xwt​(x)dx​

(LogRegretOCO.EWOO.ewooPoint). In particular x1x_1x1​ is the centroid of PPP, and each xtx_txt​ depends only on f1,…,ft−1f_1, \dots, f_{t-1}f1​,…,ft−1​. The regret against a comparator u∈Pu \in Pu∈P is ∑t=1T(ft(xt)−ft(u))\sum_{t=1}^T \bigl(f_t(x_t) - f_t(u)\bigr)∑t=1T​(ft​(xt​)−ft​(u)).

Formalization targets

Goal: Theorem 7

For every T≥1T \ge 1T≥1 and every u∈Pu \in Pu∈P,

∑t=1Tft(xt)−∑t=1Tft(u)  ≤  1α n (1+log⁡(T+1)).\sum_{t=1}^{T} f_t(x_t) - \sum_{t=1}^{T} f_t(u) \;\le\; \frac{1}{\alpha}\, n\, \bigl(1 + \log(T+1)\bigr).t=1∑T​ft​(xt​)−t=1∑T​ft​(u)≤α1​n(1+log(T+1)).

This is the paper's printed constant. The paper's proof yields the slightly sharper 1α(1+nlog⁡(T+1))\frac{1}{\alpha}\bigl(1 + n\log(T+1)\bigr)α1​(1+nlog(T+1)); the printed form is the goal.

Milestones

The proof in §3.4 (p. 187) passes through five displays, each a milestone:

  1. Jensen for the weighted mean (first display on p. 187): ht(xt)≥∫Pht wt /∫Pwth_t(x_t) \ge \int_P h_t\, w_t \,/ \int_P w_tht​(xt​)≥∫P​ht​wt​/∫P​wt​.
  2. Eq. (18): ∏τ=1thτ(xτ)≥∫P∏τ=1thτ / vol(P)\prod_{\tau=1}^t h_\tau(x_\tau) \ge \int_P \prod_{\tau=1}^t h_\tau \,/\, \mathrm{vol}(P)∏τ=1t​hτ​(xτ​)≥∫P​∏τ=1t​hτ​/vol(P).
  3. The nearby set S={TT+1x∗+1T+1y:y∈P}S = \{\frac{T}{T+1}x^* + \frac{1}{T+1}y : y \in P\}S={T+1T​x∗+T+11​y:y∈P}: for x∈Sx \in Sx∈S, ht(x)≥TT+1ht(x∗)h_t(x) \ge \frac{T}{T+1}h_t(x^*)ht​(x)≥T+1T​ht​(x∗) and ∏τ=1Thτ(x)≥1e∏τ=1Thτ(x∗)\prod_{\tau=1}^T h_\tau(x) \ge \frac1e \prod_{\tau=1}^T h_\tau(x^*)∏τ=1T​hτ​(x)≥e1​∏τ=1T​hτ​(x∗).
  4. Volume of SSS: vol(S)=vol(P)/(T+1)n\mathrm{vol}(S) = \mathrm{vol}(P)/(T+1)^nvol(S)=vol(P)/(T+1)n.
  5. Multiplicative regret bound (last display on p. 187): ∏τ=1Thτ(xτ)≥1e(T+1)n∏τ=1Thτ(x∗)\prod_{\tau=1}^T h_\tau(x_\tau) \ge \frac{1}{e(T+1)^n}\prod_{\tau=1}^T h_\tau(x^*)∏τ=1T​hτ​(xτ​)≥e(T+1)n1​∏τ=1T​hτ​(x∗).

Significance

The result. Theorem 7 shows that exp-concavity alone suffices for logarithmic regret, with a constant n/αn/\alphan/α that does not depend on the size of PPP or on the gradients of the costs. Specialised to the log-loss ft(x)=−log⁡(rt⊤x)f_t(x) = -\log(r_t^\top x)ft​(x)=−log(rt⊤​x) on the simplex, where α=1\alpha = 1α=1, it recovers the O(nlog⁡T)O(n\log T)O(nlogT) regret of Cover's universal portfolio. The bound is the benchmark against which the computationally cheaper second-order methods of the same paper (Online Newton Step, Follow the Approximate Leader) are compared: those need a gradient bound GGG and diameter DDD and pay a factor (1/α+GD)(1/\alpha + GD)(1/α+GD).

Formalizing it. The theorem is proved in the paper, and a textbook version with a different constant, (n/α)log⁡T+2/α(n/\alpha)\log T + 2/\alpha(n/α)logT+2/α, appears in Hazan's Introduction to Online Convex Optimization (Theorem 4.4). No machine-checked proof of either is known. The work here is to formalize the paper's proof: Jensen's inequality for a weighted Lebesgue average in Rn\mathbb{R}^nRn, the change of volume under homothety, and the elementary inequality (1+1/T)T≤e(1 + 1/T)^T \le e(1+1/T)T≤e. A companion draft of the textbook version exists on the platform as a private item (OnlineConvexOpt.SecondOrder.ewoo_regret) with another constant; it is not reused.

Difficulty

The pieces are classical, but they have to be assembled in measure-theoretic form. The point xtx_txt​ is a Bochner integral of a vector-valued function over PPP, and its membership in PPP and the Jensen inequality both require the normalised weight wt dx/∫Pwtw_t\,dx/\int_P w_twt​dx/∫P​wt​ to be a genuine probability measure on PPP, with every integrand integrable. The obvious one-dimensional intuition — "the weighted mean of a convex set lies in the set" — hides the requirement that PPP be closed and have positive volume.

The second obstacle is that Eq. (18) compares the algorithm with an average of the product ∏hτ\prod h_\tau∏hτ​ over all of PPP, while the regret compares it with a single point. The natural attempt, bounding the average below by the value at the comparator, fails: the average can be far smaller than the maximum, and a lower bound that loses more than a factor polynomial in TTT destroys the logarithmic rate. Controlling this loss in nnn dimensions, with a constant independent of the shape and size of PPP, is the heart of the argument.

Formalization scope

Points live in EuclideanSpace ℝ (Fin n) with its Lebesgue (Haar) measure volume. Rounds are numbered from 111: the weights sum over Finset.Ico 1 t, the regret over Finset.Icc 1 T. Cost functions are defined on all of Rn\mathbb{R}^nRn; only their values on PPP enter. The algorithm is the total function ewooPoint P α f t, and the goal is stated for xtx_txt​ equal to it — not for an arbitrary sequence satisfying a Jensen-type inequality.

Conventions and corrections relative to the printed text:

  • Regret against every comparator. The regret is stated as ∑t(ft(xt)−ft(u))≤\sum_t (f_t(x_t) - f_t(u)) \le∑t​(ft​(xt​)−ft​(u))≤ bound for every u∈Pu \in Pu∈P, never through a real-valued ⨅ or sInf over PPP, which in Lean would return a junk value off its intended domain and trivialize the statement.
  • Positive volume volume P ≠ 0 is added: the algorithm divides by ∫Pwt\int_P w_t∫P​wt​, which the paper leaves implicit. Without it Lean's convention 0−1=00^{-1} = 00−1=0 would set xt=0x_t = 0xt​=0.
  • Continuity of each ftf_tft​ on PPP is the paper's standing assumption (§2.2: costs twice differentiable and convex) weakened to what the argument uses; it makes every integral in the development an integral of an integrable function.
  • Typos. Theorem 7's "ft:P→Rnf_t : P \to \mathbb{R}^nft​:P→Rn" is read as real-valued, and its "exp⁡(−αf(x))\exp(-\alpha f(x))exp(−αf(x))" as exp⁡(−αft(x))\exp(-\alpha f_t(x))exp(−αft​(x)). The set-builder "S={x∈S∣… }S = \{x \in S \mid \dots\}S={x∈S∣…}" defines SSS in terms of itself and is read as the set of all TT+1x∗+1T+1y\frac{T}{T+1}x^* + \frac{1}{T+1}yT+1T​x∗+T+11​y, y∈Py \in Py∈P; the printed "S=x∗+1T+1PS = x^* + \frac{1}{T+1}PS=x∗+T+11​P" is a translate of that set with the same volume.
  • Comparator. The paper's x∗x^*x∗ is a minimizer of ∑tft\sum_t f_t∑t​ft​; milestones 3 and 5 are stated for every x∗∈Px^* \in Px∗∈P, which implies the minimizer case.
  • Constant. The printed 1αn(1+log⁡(T+1))\frac{1}{\alpha}n(1+\log(T+1))α1​n(1+log(T+1)) is stated, although the proof gives the sharper 1α(1+nlog⁡(T+1))\frac{1}{\alpha}(1 + n\log(T+1))α1​(1+nlog(T+1)).
  • Not in scope. The randomized variant (sampling xtx_txt​ with density proportional to wtw_twt​, "in expectation") and the running-time discussion of §3.4.1 have no separate proof in the paper.

Infrastructure that a complete development needs, and that is reusable beyond this mission: Jensen's inequality for concave functions under a probability measure with a continuous density on a compact convex set (Mathlib has ConcaveOn.le_map_integral and Convex.integral_mem); the scaling identity for Haar measure (MeasureTheory.Measure.addHaar_smul); and the elementary bound (T/(T+1))T≥1/e(T/(T+1))^T \ge 1/e(T/(T+1))T≥1/e. Proofs of any milestone, and a general weighted-Jensen lemma usable across the milestones, are welcome.

Selected references

  • E. Hazan, A. Agarwal, S. Kale, Logarithmic regret algorithms for online convex optimization, Machine Learning 69 (2007), 169–192. https://doi.org/10.1007/s10994-007-5016-8
  • T. M. Cover, Universal portfolios, Mathematical Finance 1 (1991), 1–29. https://doi.org/10.1111/j.1467-9965.1991.tb00002.x
  • A. Blum, A. Kalai, Universal portfolios with and without transaction costs, Machine Learning 35 (1999), 193–205 (COLT 1997). https://doi.org/10.1023/A:1007530728748
  • A. Kalai, S. Vempala, Efficient algorithms for universal portfolios, Journal of Machine Learning Research 3 (2003), 423–440. https://www.jmlr.org/papers/v3/kalai02a.html
  • M. Zinkevich, Online convex programming and generalized infinitesimal gradient ascent, ICML 2003. https://dl.acm.org/doi/10.5555/3041838.3041955
  • E. Hazan, Introduction to Online Convex Optimization, 2nd ed., MIT Press 2022; arXiv:1909.05207, Theorem 4.4. https://arxiv.org/abs/1909.05207
9 thms3 active usersReviewed
🏆Completed
Convex OptimizationMachine LearningOperations Research·Captain: mikedeng1

Logarithmic Regret Algorithms for Online Convex Optimization 3: Logarithmic Regret of Follow the Approximate LeaderResearch Paper

Motivation

Online convex optimization models repeated decision making against an adversary: in each round a player chooses a point of a convex set, and only then learns the convex cost of that round. It covers online portfolio selection, online regression and routing, and it is the standard lens for analysing learning algorithms that must commit before seeing data. The figure of merit is regret, the player's total cost minus the cost of the best fixed decision in hindsight. For general convex costs regret Θ(T)\Theta(\sqrt T)Θ(T​) over TTT rounds is optimal; for costs with curvature it can be logarithmic.

Hazan, Agarwal and Kale (Mach Learn 69 (2007) 169–192) gave several algorithms with O(log⁡T)O(\log T)O(logT) regret for α\alphaα-exp-concave costs, the class that contains the log-loss of portfolio selection. This mission formalizes one of them, Follow the Approximate Leader (FTAL). It connects to the oldest online algorithm, Follow the Leader (FTL), which plays the minimiser of all past costs: FTAL is FTL run on quadratic lower models of the costs, and the paper's analysis shows that FTL itself has logarithmic regret on a class of curved costs.

Timeline. Zinkevich (2003) proved O(T)O(\sqrt T)O(T​) regret for online gradient descent on convex costs. Cover (1991) gave a universal portfolio with logarithmic regret for the log-loss, at a running time exponential in the dimension. Kalai and Vempala (2005) analysed perturbed Follow the Leader through the "be the leader" argument. Hazan, Agarwal and Kale (2007) gave efficient algorithms (Online Newton Step, FTAL, EWOO) with O(nlog⁡T)O(n \log T)O(nlogT) regret for exp-concave costs.

Setting

The decision set P⊆RnP \subseteq \mathbb{R}^nP⊆Rn is nonempty, convex, closed and bounded, and DDD bounds its diameter: ∥y−z∥2≤D\|y - z\|_2 \le D∥y−z∥2​≤D for y,z∈Py, z \in Py,z∈P. In rounds t=1,2,…t = 1, 2, \dotst=1,2,… the player picks xt∈Px_t \in Pxt​∈P and then pays ft(xt)f_t(x_t)ft​(xt​), where ftf_tft​ is a cost function differentiable at the points of PPP with gradient norm ∥∇ft(x)∥≤G\|\nabla f_t(x)\| \le G∥∇ft​(x)∥≤G on PPP. The cost ftf_tft​ is α\alphaα-exp-concave (α>0\alpha > 0α>0) if x↦exp⁡(−αft(x))x \mapsto \exp(-\alpha f_t(x))x↦exp(−αft​(x)) is concave on PPP. The regret over TTT rounds against a comparator u∈Pu \in Pu∈P is ∑t=1T(ft(xt)−ft(u))\sum_{t=1}^T \bigl(f_t(x_t) - f_t(u)\bigr)∑t=1T​(ft​(xt​)−ft​(u)).

Follow the Leader plays xt∈arg⁡min⁡x∈P∑τ=1t−1fτ(x)x_t \in \arg\min_{x \in P} \sum_{\tau=1}^{t-1} f_\tau(x)xt​∈argminx∈P​∑τ=1t−1​fτ​(x) (any point of PPP in round 1). Follow the Approximate Leader (version 1 of the paper's Fig. 3) with parameter β\betaβ plays FTL on the approximate costs

f~τ(x)=fτ(xτ)+∇τ⊤(x−xτ)+β2(x−xτ)⊤∇τ∇τ⊤(x−xτ),∇τ=∇fτ(xτ).\tilde f_\tau(x) = f_\tau(x_\tau) + \nabla_\tau^\top(x - x_\tau) + \frac{\beta}{2}(x - x_\tau)^\top \nabla_\tau\nabla_\tau^\top (x - x_\tau), \qquad \nabla_\tau = \nabla f_\tau(x_\tau).f~​τ​(x)=fτ​(xτ​)+∇τ⊤​(x−xτ​)+2β​(x−xτ​)⊤∇τ​∇τ⊤​(x−xτ​),∇τ​=∇fτ​(xτ​).

In the Lean development these are IsFTLRun P f x and IsFTALRun P β f x, predicates on a whole trajectory xxx.

Formalization targets

Goal: Theorem 6

With β=12min⁡{1/(4GD),α}\beta = \tfrac12 \min\{1/(4GD), \alpha\}β=21​min{1/(4GD),α}, every FTAL run on α\alphaα-exp-concave costs satisfies, for every T≥1T \ge 1T≥1 and u∈Pu \in Pu∈P,

∑t=1T(ft(xt)−ft(u))≤64(1α+GD)n (log⁡T+1).\sum_{t=1}^T \bigl(f_t(x_t) - f_t(u)\bigr) \le 64\left(\frac1\alpha + GD\right) n\,(\log T + 1).t=1∑T​(ft​(xt​)−ft​(u))≤64(α1​+GD)n(logT+1).

This is the paper's statement with its constant, stated for the algorithm as defined, and for every adversarial sequence of costs.

Milestones

  1. Lemma 3: an α\alphaα-exp-concave cost with gradients bounded by GGG lies above the paraboloid f(y)+∇f(y)⊤(x−y)+β2(∇f(y)⊤(x−y))2f(y) + \nabla f(y)^\top(x-y) + \frac\beta2 (\nabla f(y)^\top (x - y))^2f(y)+∇f(y)⊤(x−y)+2β​(∇f(y)⊤(x−y))2 on PPP.
  2. Lemma 9: regret on lower surrogates that touch the costs at the played points dominates the true regret.
  3. Lemma 10: ∑tft(xt+1)≤∑tft(u)\sum_t f_t(x_{t+1}) \le \sum_t f_t(u)∑t​ft​(xt+1​)≤∑t​ft​(u) for an FTL run ("be the leader").
  4. Lemma 12: A−1∙(A−B)≤log⁡(∣A∣/∣B∣)A^{-1} \bullet (A - B) \le \log(|A|/|B|)A−1∙(A−B)≤log(∣A∣/∣B∣) for A⪰B≻0A \succeq B \succ 0A⪰B≻0.
  5. Lemma 11: ∑t=1Tut⊤Vt−1ut≤nlog⁡(r2T/ε+1)\sum_{t=1}^T u_t^\top V_t^{-1} u_t \le n\log(r^2T/\varepsilon + 1)∑t=1T​ut⊤​Vt−1​ut​≤nlog(r2T/ε+1) with Vt=∑τ≤tuτuτ⊤+εIV_t = \sum_{\tau \le t} u_\tau u_\tau^\top + \varepsilon IVt​=∑τ≤t​uτ​uτ⊤​+εI.
  6. Theorem 5 (corrected constant): FTL on costs gt(vt⊤x)g_t(v_t^\top x)gt​(vt⊤​x) with ∥vt∥≤R\|v_t\| \le R∥vt​∥≤R, ∣gt′∣≤b|g_t'| \le b∣gt′​∣≤b, gt′′≥ag_t'' \ge agt′′​≥a has regret at most nb2alog⁡(a2D2R2T2b2+1)+b2a\frac{nb^2}{a}\log\bigl(\frac{a^2D^2R^2T^2}{b^2} + 1\bigr) + \frac{b^2}{a}anb2​log(b2a2D2R2T2​+1)+ab2​.

Significance

Theorem 6 shows that a simple rule, re-solving a convex quadratic program over all past linearized costs, achieves O(nlog⁡T)O(n\log T)O(nlogT) regret on exp-concave costs, matching the Online Newton Step up to constants. Theorem 5 is of independent interest: it shows that unmodified Follow the Leader, which has linear regret on linear costs, has logarithmic regret whenever each cost is a strongly curved function of one linear form. Portfolio selection is such a case. The appendix lemmas (log-determinant potential, elliptical potential) are standard tools reused throughout the bandit and online-learning literature.

On formalization: the results are proved on paper; none is formalized. The Lean development provides a reusable encoding of Follow the Leader as a trajectory predicate, the "be the leader" reduction, the surrogate reduction for regret, and the matrix potential inequalities, which the Online Newton Step analysis also needs. The paper's printed statements of Theorem 5 and Lemma 10 contain errors (see below); this mission states corrected versions that suffice for the goal.

Difficulty

The obvious attempt, bounding each term ft(xt)−ft(xt+1)f_t(x_t) - f_t(x_{t+1})ft​(xt​)−ft​(xt+1​) by how far the leader moves, requires knowing how far the minimiser of a constrained problem moves when one cost is added. For unconstrained strongly convex quadratics this is an explicit Newton step, but here the minimiser lies in a general convex set and each cost contributes curvature in only one direction, so the accumulated curvature can be singular for many rounds and no per-round strong convexity is available. Turning the per-round movement into a sum that grows only like log⁡T\log TlogT, with the paper's explicit constant, is the core of the work; the printed Theorem 5 bound is negative for small TTT, so the constants must be tracked exactly rather than asymptotically.

Formalization scope

Points are in EuclideanSpace ℝ (Fin n) so that ∥⋅∥\|\cdot\|∥⋅∥ is the Euclidean norm; cost functions are functions on all of Rn\mathbb{R}^nRn, differentiable at the points of PPP, with Mathlib's gradient. Rounds are 111-based; x0x_0x0​ and f0f_0f0​ are unused. DDD is any upper bound on pairwise distances in PPP. Exp-concavity is ConcaveOn ℝ P (fun x => Real.exp (-α * f t x)). Algorithms are predicates on the trajectory, required at every round, so every tie-breaking rule is covered and adaptive adversaries are included.

Regret is always stated against every comparator u∈Pu \in Pu∈P. A formalization with a real-valued ⨅/sInf over PPP, or one that bounds the regret of an arbitrary sequence of points rather than of an FTAL run with the paper's β\betaβ, would be trivial or false, and is excluded: the goal carries IsFTALRun with β=12min⁡{1/(4GD),α}\beta = \frac12\min\{1/(4GD),\alpha\}β=21​min{1/(4GD),α}.

Corrections and conventions relative to the printed paper:

  • Theorem 5: the printed bound 2nb2a[log⁡(DRaT/b)+1]\frac{2nb^2}{a}[\log(DRaT/b) + 1]a2nb2​[log(DRaT/b)+1] is false when DRaT/b<1/eDRaT/b < 1/eDRaT/b<1/e. The milestone states the bound the paper's proof gives, nb2alog⁡(a2D2R2T2b2+1)+b2a\frac{nb^2}{a}\log(\frac{a^2D^2R^2T^2}{b^2} + 1) + \frac{b^2}{a}anb2​log(b2a2D2R2T2​+1)+ab2​, which implies the printed one when DRaT≥bDRaT \ge bDRaT≥b. Derivatives are deriv with explicit differentiability at the points vt⊤xv_t^\top xvt⊤​x, x∈Px \in Px∈P.
  • Lemma 10: printed with xt=arg⁡min⁡∑τ=1tfτx_t = \arg\min \sum_{\tau=1}^{t} f_\tauxt​=argmin∑τ=1t​fτ​, under which it is false at T=1T = 1T=1; the proof and its use require the FTL index ∑τ=1t−1\sum_{\tau=1}^{t-1}∑τ=1t−1​, which is stated.
  • Lemma 11: the typo ∑τutut⊤\sum_\tau u_t u_t^\top∑τ​ut​ut⊤​ is read as ∑τuτuτ⊤\sum_\tau u_\tau u_\tau^\top∑τ​uτ​uτ⊤​, and ε>0\varepsilon > 0ε>0 is stated.
  • Lemma 3: β>0\beta > 0β>0 is added (the proof divides by β\betaβ), and G,D>0G, D > 0G,D>0 so that 1/(4GD)1/(4GD)1/(4GD) is meaningful.
  • Theorem 6: "ft:P→Rnf_t : P \to \mathbb{R}^nft​:P→Rn" is read as R\mathbb{R}R-valued; only first-order differentiability is assumed; G,D>0G, D > 0G,D>0. The theorem is true as printed, although the paper's route through the printed Theorem 5 is invalid for T<16T < 16T<16.
  • Only version 1 of FTAL is formalized; Lemma 4 (equivalence with the pseudoinverse form) is out of scope.

Needed infrastructure: first-order optimality for convex minimisation over a convex set, a mean-value theorem along segments, determinants and eigenvalues of symmetric positive definite matrices (Mathlib has most of this), and the matrix inequality ∣A∣≤(tr⁡A/n)n|A| \le (\operatorname{tr} A/n)^n∣A∣≤(trA/n)n. Contributions welcome: proofs of any milestone, and a proof of the goal from the milestones.

Selected references

  • E. Hazan, A. Agarwal, S. Kale, Logarithmic regret algorithms for online convex optimization, Machine Learning 69 (2007), 169–192. https://doi.org/10.1007/s10994-007-5016-8
  • M. Zinkevich, Online convex programming and generalized infinitesimal gradient ascent, ICML 2003. https://dl.acm.org/doi/10.5555/3041838.3041955
  • T. M. Cover, Universal portfolios, Mathematical Finance 1 (1991), 1–29. https://doi.org/10.1111/j.1467-9965.1991.tb00002.x
  • A. Kalai, S. Vempala, Efficient algorithms for online decision problems, J. Comput. System Sci. 71 (2005), 291–307. https://doi.org/10.1016/j.jcss.2004.10.016
  • E. Hazan, Introduction to Online Convex Optimization, Foundations and Trends in Optimization 2 (2016). https://arxiv.org/abs/1909.05207
9 thms3 active usersReviewed
🏆Completed
Convex OptimizationMachine LearningOperations Research·Captain: mikedeng1

Logarithmic Regret Algorithms for Online Convex Optimization 2: Logarithmic Regret of the Online Newton StepResearch Paper

Motivation

Online convex optimization models repeated decision making against an unknown, possibly adversarial environment: in each round t=1,…,Tt=1,\dots,Tt=1,…,T a player picks a point xtx_txt​ of a convex set P⊆Rn\mathcal P\subseteq\mathbb R^nP⊆Rn, and only then learns a convex cost function ftf_tft​ and pays ft(xt)f_t(x_t)ft​(xt​). Performance is measured by regret, the excess of the total cost over that of the best fixed point in hindsight. Zinkevich (ICML 2003) showed that online gradient descent has regret O(T)O(\sqrt T)O(T​) for arbitrary convex costs with bounded gradients, and this rate cannot be improved in general.

Many costs met in practice have more curvature than bare convexity. The log-loss f(x)=−log⁡(x⊤a)f(x)=-\log(x^\top a)f(x)=−log(x⊤a) of universal portfolio management (Cover, Math. Finance 1991) is not strongly convex, but it is exp-concave. Hazan, Agarwal and Kale (Mach Learn 69, 2007) gave the first efficient algorithms with regret logarithmic in TTT for exp-concave costs. This mission formalizes the second of their algorithms, the Online Newton Step (ONS), and its regret bound (Theorem 2 of the paper). ONS is the basis of later second-order online methods and appears as a standard algorithm in textbooks on online learning.

Timeline:

  • 2003 — Zinkevich: O(T)O(\sqrt T)O(T​) regret for general convex costs by online gradient descent.
  • 2006–2007 — Hazan, Agarwal, Kale (COLT 2006; Mach Learn 2007): O(log⁡T)O(\log T)O(logT) regret for strongly convex costs by gradient descent, and O(nlog⁡T)O(n\log T)O(nlogT) regret for exp-concave costs by ONS, Follow the Approximate Leader, and exponentially weighted online optimization.
  • 2016 — Hazan, Introduction to Online Convex Optimization (Found. Trends Optim., arXiv:1909.05207): textbook treatment of ONS with modified parameters.

Setting

The decision set P⊆Rn\mathcal P\subseteq\mathbb R^nP⊆Rn is nonempty, closed, bounded and convex, and DDD bounds its diameter: ∥x−y∥≤D\|x-y\|\le D∥x−y∥≤D for all x,y∈Px,y\in\mathcal Px,y∈P, with the Euclidean norm. The costs f1,f2,…f_1,f_2,\dotsf1​,f2​,… are real functions, differentiable at every point of P\mathcal PP, with gradient bound ∥∇ft(x)∥≤G\|\nabla f_t(x)\|\le G∥∇ft​(x)∥≤G on P\mathcal PP. A cost is α\alphaα-exp-concave (α>0\alpha>0α>0) if x↦exp⁡(−αft(x))x\mapsto\exp(-\alpha f_t(x))x↦exp(−αft​(x)) is concave on P\mathcal PP.

For a matrix AAA, the generalized projection ΠPA(y)\Pi^A_{\mathcal P}(y)ΠPA​(y) is a point of P\mathcal PP minimising (y−x)⊤A(y−x)(y-x)^\top A(y-x)(y−x)⊤A(y−x) over x∈Px\in\mathcal Px∈P.

The Online Newton Step fixes

β=12min⁡{14GD,α},ε=1β2D2,\beta=\tfrac12\min\Big\{\frac1{4GD},\alpha\Big\},\qquad \varepsilon=\frac1{\beta^2D^2},β=21​min{4GD1​,α},ε=β2D21​,

writes ∇t=∇ft(xt)\nabla_t=\nabla f_t(x_t)∇t​=∇ft​(xt​) and At=∑i=1t∇i∇i⊤+εInA_t=\sum_{i=1}^t\nabla_i\nabla_i^\top+\varepsilon I_nAt​=∑i=1t​∇i​∇i⊤​+εIn​, plays an arbitrary x1∈Px_1\in\mathcal Px1​∈P, and then

xt+1=ΠPAt(xt−1βAt−1∇t).x_{t+1}=\Pi^{A_t}_{\mathcal P}\Big(x_t-\frac1\beta A_t^{-1}\nabla_t\Big).xt+1​=ΠPAt​​(xt​−β1​At−1​∇t​).

The regret after TTT rounds against a comparator u∈Pu\in\mathcal Pu∈P is ∑t=1T(ft(xt)−ft(u))\sum_{t=1}^T\big(f_t(x_t)-f_t(u)\big)∑t=1T​(ft​(xt​)−ft​(u)); the paper's regret is its maximum over u∈Pu\in\mathcal Pu∈P.

In Lean the objects are LogRegretOCO.ONS.onsBeta, onsEps, onsMatrix, IsGenProj and IsONSRun, with the regularised Gram matrix regGram and the quadratic form quadForm.

Formalization targets

Goal: Theorem 2 with nlog⁡T≥4n\log T\ge4nlogT≥4

For every run of ONS, every horizon TTT with nlog⁡T≥4n\log T\ge 4nlogT≥4, and every u∈Pu\in\mathcal Pu∈P,

∑t=1T(ft(xt)−ft(u))≤5(1α+GD) nlog⁡T.\sum_{t=1}^T\big(f_t(x_t)-f_t(u)\big)\le 5\Big(\frac1\alpha+GD\Big)\,n\log T.t=1∑T​(ft​(xt​)−ft​(u))≤5(α1​+GD)nlogT.

The added condition nlog⁡T≥4n\log T\ge4nlogT≥4 is what makes the printed constant correct (see Formalization scope).

Milestones

  1. Lemma 3 (p. 177): for 0<β≤12min⁡{1/(4GD),α}0<\beta\le\frac12\min\{1/(4GD),\alpha\}0<β≤21​min{1/(4GD),α} and x,y∈Px,y\in\mathcal Px,y∈P,
f(x)≥f(y)+∇f(y)⊤(x−y)+β2(∇f(y)⊤(x−y))2.f(x)\ge f(y)+\nabla f(y)^\top(x-y)+\tfrac\beta2\big(\nabla f(y)^\top(x-y)\big)^2 .f(x)≥f(y)+∇f(y)⊤(x−y)+2β​(∇f(y)⊤(x−y))2.
  1. Lemma 8 (p. 188): for convex P\mathcal PP, A⪰0A\succeq0A⪰0, z=ΠPA(y)z=\Pi^A_{\mathcal P}(y)z=ΠPA​(y) and a∈Pa\in\mathcal Pa∈P: (y−a)⊤A(y−a)≥(z−a)⊤A(z−a)(y-a)^\top A(y-a)\ge(z-a)^\top A(z-a)(y−a)⊤A(y−a)≥(z−a)⊤A(z−a).
  2. The display on p. 178: for every run of ONS and u∈Pu\in\mathcal Pu∈P,
∑t=1T(ft(xt)−ft(u))≤12β∑t=1T∇t⊤At−1∇t+12β.\sum_{t=1}^T\big(f_t(x_t)-f_t(u)\big)\le\frac1{2\beta}\sum_{t=1}^T\nabla_t^\top A_t^{-1}\nabla_t+\frac1{2\beta}.t=1∑T​(ft​(xt​)−ft​(u))≤2β1​t=1∑T​∇t⊤​At−1​∇t​+2β1​.
  1. Lemma 12 (p. 191): for A⪰B≻0A\succeq B\succ0A⪰B≻0, A−1∙(A−B)≤log⁡(∣A∣/∣B∣)A^{-1}\bullet(A-B)\le\log(|A|/|B|)A−1∙(A−B)≤log(∣A∣/∣B∣).
  2. Lemma 11 (p. 190): if ∥ut∥≤r\|u_t\|\le r∥ut​∥≤r, ε>0\varepsilon>0ε>0 and Vt=∑τ≤tuτuτ⊤+εInV_t=\sum_{\tau\le t}u_\tau u_\tau^\top+\varepsilon I_nVt​=∑τ≤t​uτ​uτ⊤​+εIn​, then ∑t=1Tut⊤Vt−1ut≤nlog⁡(r2T/ε+1)\sum_{t=1}^Tu_t^\top V_t^{-1}u_t\le n\log(r^2T/\varepsilon+1)∑t=1T​ut⊤​Vt−1​ut​≤nlog(r2T/ε+1).

Significance

Theorem 2 shows that exp-concavity alone, without strong convexity, suffices for regret logarithmic in TTT, at a per-round cost of one rank-one matrix update and one generalized projection. Its consequences include logarithmic regret for universal portfolio selection with a polynomial-time algorithm, and, by online-to-batch conversion, fast rates for stochastic exp-concave optimization. Lemma 11 (the elliptical potential bound) is used well beyond this paper, in linear bandits and online regression.

The result has been proved on paper since 2007. The remaining work is its machine-checked proof: the potential argument, the log-determinant inequality and the generalized-projection inequality for positive semidefinite matrices. As far as is known, none of these results is formalized in Mathlib. Prove2Me holds a related elliptical potential lemma for linear bandits (BanditAlgorithm.elliptical_potential_lemma, with Vt−1V_{t-1}Vt−1​ and a min⁡(1,⋅)\min(1,\cdot)min(1,⋅), a different statement) and the Euclidean case A=IA=IA=I of Lemma 8 (UnderstandingML.projection_lemma). The textbook version of ONS (OnlineConvexOpt.SecondOrder.online_newton_step_regret, with γ=12min⁡{1/(GD),α}\gamma=\frac12\min\{1/(GD),\alpha\}γ=21​min{1/(GD),α} and bound 2(1/α+GD)nlog⁡T2(1/\alpha+GD)n\log T2(1/α+GD)nlogT) is an open private draft with different parameters.

Difficulty

The obvious route to logarithmic regret, the gradient-descent argument of Theorem 1 with step sizes 1/(Ht)1/(Ht)1/(Ht), needs a uniform lower bound H>0H>0H>0 on the Hessians. Exp-concave costs such as the log-loss have no such bound: their curvature vanishes in directions orthogonal to the gradients seen so far. The analysis therefore has to track curvature only along the observed gradient directions. This requires a matrix-valued potential ∑t∇t⊤At−1∇t\sum_t\nabla_t^\top A_t^{-1}\nabla_t∑t​∇t⊤​At−1​∇t​ and a projection in the norm of AtA_tAt​ rather than the Euclidean norm. The Euclidean projection inequality does not transfer to this norm, which changes from round to round. Bounding the potential requires determinant inequalities for positive definite matrices. The analytic facts are elementary, but their Lean statements involve the interaction of EuclideanSpace, Matrix.mulVec, Matrix.inv and Matrix.det.

Formalization scope

Points live in EuclideanSpace ℝ (Fin n), so all norms are Euclidean; matrices are Matrix (Fin n) (Fin n) ℝ acting on coordinate vectors. Rounds are 1-based: sums run over Finset.Icc 1 T and the index 000 is unused. Cost functions are ambient functions Rn→R\mathbb R^n\to\mathbb RRn→R, differentiable at the points of P\mathcal PP, with ∇ft\nabla f_t∇ft​ given by Mathlib's gradient. The paper's standing assumptions of convexity and twice differentiability are not needed and are omitted. DDD enters only as an upper bound on distances in P\mathcal PP. The generalized projection is a predicate that every minimiser satisfies, and ONS is the predicate IsONSRun on the whole trajectory, so the goal covers every tie-break and every adaptive adversary.

Corrections and added hypotheses:

  • Theorem 2 is false as printed at T=1T=1T=1. Take n=1n=1n=1, P=[−1,1]\mathcal P=[-1,1]P=[−1,1], f1(x)=x2f_1(x)=x^2f1​(x)=x2, α=12\alpha=\frac12α=21​, G=D=2G=D=2G=D=2 and x1=1x_1=1x1​=1: the regret is 111 and the bound is 000. The paper's proof gives 4(1/α+GD)(nlog⁡T+1)4(1/\alpha+GD)(n\log T+1)4(1/α+GD)(nlogT+1) for T≥2T\ge2T≥2; the final sentence drops the additive 1/(2β)1/(2\beta)1/(2β) of the p. 178 display. The goal adds nlog⁡T≥4n\log T\ge4nlogT≥4, under which the printed constant 555 follows.
  • G,D,α>0G,D,\alpha>0G,D,α>0 are assumed wherever β\betaβ or ε\varepsilonε appear: they are the non-degeneracy the formulas presuppose (in Lean, 1/0=01/0=01/0=0).
  • Lemma 3 adds 0<β0<\beta0<β; the proof divides by β\betaβ.
  • Lemma 11 adds ε>0\varepsilon>0ε>0 and reads the printed ∑τ=1tutut⊤\sum_{\tau=1}^tu_tu_t^\top∑τ=1t​ut​ut⊤​ as ∑τ=1tuτuτ⊤\sum_{\tau=1}^tu_\tau u_\tau^\top∑τ=1t​uτ​uτ⊤​.
  • Lemma 12's product ∙\bullet∙ is the entrywise inner product ∑i,jCijEij\sum_{i,j}C_{ij}E_{ij}∑i,j​Cij​Eij​, written out as a double sum.
  • The printed "ft:P→Rnf_t:\mathcal P\to\mathbb R^nft​:P→Rn" is read as ft:P→Rf_t:\mathcal P\to\mathbb Rft​:P→R, and "ΠSnAt\Pi^{A_t}_{S_n}ΠSn​At​​" on p. 177 as ΠPAt\Pi^{A_t}_{\mathcal P}ΠPAt​​.

Regret is stated against every comparator u∈Pu\in\mathcal Pu∈P, never as a real infimum ⨅ over P\mathcal PP, which is junk-valued in Lean on unbounded or empty sets. The goal is a statement about runs of the paper's algorithm with the paper's β\betaβ, ε\varepsilonε and AtA_tAt​. A bound for an arbitrary sequence satisfying the p. 178 display would be a milestone, not Theorem 2. The hypotheses are jointly satisfiable: the closed unit ball with ft(x)=∥x∥2/2f_t(x)=\|x\|^2/2ft​(x)=∥x∥2/2, α=1\alpha=1α=1, G=1G=1G=1, D=2D=2D=2 is a model.

A complete development needs: first-order conditions for concave functions on convex sets at boundary points; the optimality condition for minimising a convex quadratic over a convex set; spectral facts about symmetric positive definite matrices (square roots, eigenvalues, tr⁡\operatorname{tr}tr and det⁡\detdet); and the telescoping of log-determinants. Lemmas 8, 11 and 12 are reusable beyond this mission, in the sibling missions of this series (Follow the Approximate Leader) and in linear-bandit analyses. Proofs of any milestone are welcome, as are alternative proofs of Lemma 12 through concavity of log⁡det⁡\log\detlogdet.

Selected references

  • E. Hazan, A. Agarwal, S. Kale, Logarithmic regret algorithms for online convex optimization, Machine Learning 69 (2007), 169–192. https://doi.org/10.1007/s10994-007-5016-8
  • M. Zinkevich, Online convex programming and generalized infinitesimal gradient ascent, ICML 2003. https://dl.acm.org/doi/10.5555/3041838.3041955
  • T. M. Cover, Universal portfolios, Mathematical Finance 1 (1991), 1–29. https://doi.org/10.1111/j.1467-9965.1991.tb00002.x
  • E. Hazan, Introduction to Online Convex Optimization, Foundations and Trends in Optimization 2 (2016); 2nd ed. arXiv:1909.05207. https://arxiv.org/abs/1909.05207
8 thms3 active usersReviewed
PreviousPage 8 of 26Next

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me