Prove2Me
Navigate
DiscoverFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

Operations Research

889 missions · 513 completed

The discipline of applying mathematical analysis to complex decision problems in operations: allocating scarce resources, scheduling, routing, inventory, and the design of service and production systems. Drawing on mathematical programming, stochastic modeling, queueing, simulation, and game-theoretic reasoning, it seeks policies that perform provably well in systems shaped by constraints, congestion, and uncertainty.

Missions

Open376Completed513All889
🏆Completed
Linear OptimizationOptimization·Captain: Shuze Chen

Disjunctive Programming III: Projecting Polyhedra and the Convex Hull via PolarityTextbook

Motivation

Theorem 2.1 (the previous mission in this series) shows that the closed convex hull of a union of polyhedra has a compact description after lifting to a higher-dimensional space. That description comes in two dual flavors: a primal one, as the projection of an explicit lifted polyhedron, and a polar one, characterizing the hull's facets directly via a cone built from the disjuncts' own data. Both flavors matter in practice: a cutting-plane algorithm needs to know exactly which inequalities are facet-defining (so as not to waste effort generating redundant cuts), and the two routes — projection and polarity — offer complementary tools for deciding this. This mission formalizes both routes and the machinery connecting them, closing out Chapter 2 of Balas, Disjunctive Programming (Springer, 2018).

The projection route (§2.2–2.3) develops general facts about projecting an arbitrary polyhedron that predate and underlie the disjunctive-programming application: the classical projection formula via extreme rays of a projection cone, how dimension and facet structure behave under projection, and a refinement (via a coordinate transformation) that eliminates the redundant inequalities the plain projection formula can produce. The polarity route (§2.4) develops the reverse polar, an object introduced by Balas specifically for this purpose, whose iterated application recovers the closed convex hull of a disjunctive set directly, culminating in an exact characterization of when an inequality is facet-defining purely in terms of extreme rays of an explicit cone W0W_0W0​.

Setting

For a matrix system (A,B,b)(A,B,b)(A,B,b) with mmm rows, let Q:={(u,x)∈Rp×Rq:Au+Bx≤b}Q := \{(u,x) \in \mathbb{R}^p \times \mathbb{R}^q : Au+Bx \le b\}Q:={(u,x)∈Rp×Rq:Au+Bx≤b}, and let Projx(Q):={x:∃ u, (u,x)∈Q}\mathrm{Proj}_x(Q) := \{x : \exists\, u,\ (u,x) \in Q\}Projx​(Q):={x:∃u, (u,x)∈Q} be its projection onto the xxx-space. The projection cone is W:={v:vA=0, v≥0}W := \{v : vA=0,\ v \ge 0\}W:={v:vA=0, v≥0}. A vector vvv is an extreme ray of a cone WWW if v≠0v \ne 0v=0, v∈Wv \in Wv∈W, and the ray it generates is an extreme subset of WWW. The dimension dim⁡(P)\dim(P)dim(P) of a polyhedron is the dimension of its affine hull, and a set FFF is a facet of PPP if it is a proper face of PPP of dimension dim⁡(P)−1\dim(P)-1dim(P)−1. Partitioning (A,B,b)(A,B,b)(A,B,b)'s rows into those tight throughout QQQ (the equality subsystem) and the rest, rrr and r∗r^*r∗ denote the rank of the tight rows' combined and AAA-only submatrices, respectively.

For S⊆RnS \subseteq \mathbb{R}^nS⊆Rn, the polar is S0:={x:xy≤1 ∀y∈S}S^0 := \{x : xy \le 1\ \forall y \in S\}S0:={x:xy≤1 ∀y∈S} and the reverse polar is S#:={x:xy≥1 ∀y∈S}S^\# := \{x : xy \ge 1\ \forall y \in S\}S#:={x:xy≥1 ∀y∈S}; more generally the scaled polar at level α0\alpha_0α0​ is F(α0):={y:xy≥α0 ∀x∈F}F_{(\alpha_0)} := \{y : xy \ge \alpha_0\ \forall x \in F\}F(α0​)​:={y:xy≥α0​ ∀x∈F}. For a disjunctive set F=⋃h∈QPhF = \bigcup_{h \in Q} P_hF=⋃h∈Q​Ph​ with Ph:={x:Ahx≥bh}P_h := \{x : A_h x \ge b_h\}Ph​:={x:Ah​x≥bh​} and Q∗:={h:Ph≠∅}Q^* := \{h : P_h \ne \emptyset\}Q∗:={h:Ph​=∅}, the cone W0:={(α,α0):∃ (uh)h∈Q∗, ∀h, uhAh=α, α0≤uhbh, uh≥0}W_0 := \{(\alpha,\alpha_0) : \exists\, (u_h)_{h \in Q^*},\ \forall h,\ u_h A_h = \alpha,\ \alpha_0 \le u_h b_h,\ u_h \ge 0\}W0​:={(α,α0​):∃(uh​)h∈Q∗​, ∀h, uh​Ah​=α, α0​≤uh​bh​, uh​≥0}.

Formalization targets

Theorem 2.18 (goal) — facet characterization via polarity

For a full-dimensional disjunctive set FFF (dim⁡(F)=n\dim(F)=ndim(F)=n) and α0≠0\alpha_0 \ne 0α0​=0:

αx≥α0 defines a facet of cl conv(F)  ⟺  (α,α0) is an extreme ray of W0.\alpha x \ge \alpha_0 \text{ defines a facet of } \mathrm{cl\,conv}(F) \iff (\alpha,\alpha_0) \text{ is an extreme ray of } W_0.αx≥α0​ defines a facet of clconv(F)⟺(α,α0​) is an extreme ray of W0​.

The polarity chain feeding the goal

Proposition 2.13 (0∈cl conv(S)  ⟺  S#=∅  ⟺  S#0 \in \mathrm{cl\,conv}(S) \iff S^\# = \emptyset \iff S^\#0∈clconv(S)⟺S#=∅⟺S# bounded), Theorem 2.14 (S##=cl conv(S)+cl cone(S)S^{\#\#} = \mathrm{cl\,conv}(S) + \mathrm{cl\,cone}(S)S##=clconv(S)+clcone(S) when 0∉cl conv(S)0 \notin \mathrm{cl\,conv}(S)0∈/clconv(S)), Corollary 2.15 (cl conv(S)=S00∩S##\mathrm{cl\,conv}(S) = S^{00} \cap S^{\#\#}clconv(S)=S00∩S##), Theorem 2.16 (the scaled polar stabilizes: F(α0)###=F(α0)#F_{(\alpha_0)}^{\#\#\#} = F_{(\alpha_0)}^{\#}F(α0​)###​=F(α0​)#​), and Corollary 2.17 (F(α0)={α:(α,α0)∈W0}F_{(\alpha_0)} = \{\alpha : (\alpha,\alpha_0) \in W_0\}F(α0​)​={α:(α,α0​)∈W0​}) — each the weakest statement needed for the next.

The projection track (independent of the goal's direct proof, sharing its definitions)

Theorem 2.5 (Projx(Q)={x:(vB)x≤vb, v∈extr(W)}\mathrm{Proj}_x(Q) = \{x : (vB)x \le vb,\ v \in \mathrm{extr}(W)\}Projx​(Q)={x:(vB)x≤vb, v∈extr(W)}), Proposition 2.6 (projection preserves integrality), Theorem 2.7 (dim⁡(Projx(Q))=dim⁡(Q)−p+r∗\dim(\mathrm{Proj}_x(Q)) = \dim(Q)-p+r^*dim(Projx​(Q))=dim(Q)−p+r∗), Corollaries 2.8–2.10 (facet/face behavior under projection), and Proposition 2.11 / Corollary 2.12 (sharper facet characterizations via a coordinate-transformed projection cone).

Significance

The results themselves. Theorem 2.18 is the practical payoff of the entire polarity apparatus: it turns "is this inequality facet-defining for the convex hull of a union of polyhedra" from a geometric question into an algebraic one about extreme rays of an explicit, finitely-generated cone built directly from the disjuncts' own constraint data — exactly the kind of question a cutting-plane algorithm needs answered to avoid generating redundant cuts. The projection track is foundational general polyhedral theory in its own right (Theorem 2.5's formula underlies Benders decomposition and classical Fourier-Motzkin elimination as special cases, per the book's own remarks), independently useful beyond the disjunctive setting.

Formalizing it. No object in this mission — polars, reverse polars, projection cones, extreme rays of a cone, or the dimension/facet apparatus of a polyhedron — exists on the platform prior to this mission or in Mathlib (a q=polar search returns only an unrelated cyclic-polytope construction from the Hirsch-conjecture series, with different conventions and object). This mission restates the disjunctive-set vocabulary of the companion ConvexHull mission locally (per the series convention that a draft mission cannot import another draft mission's definitions) and builds the polarity apparatus from scratch on top of it.

Difficulty

The natural first attempt at Theorem 2.18 tries to characterize facets of cl conv(F)\mathrm{cl\,conv}(F)clconv(F) directly from the lifted-polyhedron representation of Theorem 2.1, projecting facet by facet. This misses the point of the polarity route entirely: Theorem 2.18's proof instead goes through F(α0)F_{(\alpha_0)}F(α0​)​, showing a vertex of F(α0)F_{(\alpha_0)}F(α0​)​ corresponds to a nonhomogeneous subset of rank nnn of F(α0)F_{(\alpha_0)}F(α0​)​'s own defining system being tight — algebra entirely in the dual space of multipliers, never touching the lifted polyhedron's facets directly. The two obstacles Theorem 2.14 and Proposition 2.13 exist to clear are, respectively: reverse polars do not satisfy the ordinary polar's clean involution property (an extra cl cone(S)\mathrm{cl\,cone}(S)clcone(S) summand appears, capturing recession directions the reverse-polar construction alone cannot see), and reverse polars are either empty or automatically unbounded (never merely "small"), which is why the apparatus needs the normalization 0∉cl conv(F)0 \notin \mathrm{cl\,conv}(F)0∈/clconv(F) throughout.

Formalization scope

All results are stated over finite index sets and matrices Matrix (Fin (m h)) (Fin n) ℝ (disjunctive-set data, m : Q → ℕ dependent) or Matrix (Fin m) (Fin p) ℝ / Matrix (Fin m) (Fin q) ℝ (projection-track data). PolyDim and IsFacet are stated generically over any real vector space (via Module.finrank of vectorSpan and Mathlib's IsExtreme), so the same definitions serve both Poly2-shaped pairs and cl conv F ⊆ Fin n → ℝ directly in Theorem 2.18. IsExtremeRay is likewise stated generically, reused for cones in plain vector space, (v,v0)-space, and the triple (v,w,v0)-space Proposition 2.11 needs.

Two results (Proposition 2.11, Corollary 2.12) build on a coordinate-transformed polyhedron Q̃/cone W̃ that the book itself only cites from [14] rather than constructing; consistent with the book's own treatment, this mission takes W̃ (or its (v,v0)-projection) as given data together with its defining relationship to Proj_x(Q), rather than re-deriving the transformation — a choice recorded in MODERATION_NOTES.md, not a weakening of either statement's content. Proposition 2.11's complexity remark ("O(max{m,q}³)") is a proof aside about the transformation's cost, not part of either result's mathematical claim, and is out of scope per the book-wide disposition (triage.json).

A trivializing formalization is ruled out explicitly: the projection-track results are stated for generic m, p, q, never fixed at small values, and Theorem 2.18 is stated for a generic finite disjunctive index set Q, not specialized to |Q| = 1 (which would collapse W_0 to ordinary LP polarity and prove nothing about unions).

Selected references

  • E. Balas, Disjunctive Programming, Springer, 2018. DOI: 10.1007/978-3-030-00148-3, Chapter 2, §2.2–2.4.
  • E. Balas, Disjunctive programming: Properties of the convex hull of feasible points, Discrete Applied Mathematics 89 (1998), 3–44 (cited in the text as [6], the origin of the reverse-polar apparatus alongside [10]).
  • Balas, Pordli (cited as [14] in the text) — the coordinate-transformation construction behind Proposition 2.11 and Corollary 2.12.
  • Balas, Portugal (cited as [30] in the text) — the source of the dimensional results of §2.2.2.
18 thms3 active usersReviewed
🏆Completed
CombinatoricsGraph TheoryLinear Optimization·Captain: mikedeng1

Cones of Matrices and Set-Functions and 0–1 Optimization II: One Round of N on the Stable Set Polytope Gives Exactly the Odd Hole ConstraintsResearch Paper

Motivation

The stable set problem (vertex packing) asks for a largest set of pairwise non-adjacent nodes of a graph. It is NP-hard, and its polyhedral study, the description of the stable set polytope STAB(G)\mathrm{STAB}(G)STAB(G) by linear inequalities, is one of the most studied topics of polyhedral combinatorics. Classes of valid inequalities (clique, odd hole, odd antihole, wheel constraints) and the graph classes they describe exactly (perfect, ttt-perfect, hhh-perfect graphs) organize much of that literature; see Grötschel, Lovász and Schrijver, Geometric Algorithms and Combinatorial Optimization (Springer, 1988).

Lovász and Schrijver (SIAM J. Optim. 1(2), 1991) introduced a general lift-and-project procedure for 0–1 programs: lift a relaxation KKK into a space of matrices, impose linear conditions that every 0–1 point satisfies, and project back. One round of their operator NNN gives a tighter relaxation N(K)N(K)N(K) that still contains every 0–1 point of KKK; nnn rounds give the 0–1 hull. The procedure is an ancestor of the Sherali–Adams and Lasserre hierarchies, and the stable set problem is its first test case. This mission formalizes the paper's exact description of what one round of NNN does to the fractional stable set polytope: it adds precisely the odd hole constraints.

Setting

Let G=(V,E)G = (V, E)G=(V,E) be a finite graph with no isolated nodes, n=∣V∣n = |V|n=∣V∣. Vectors of RV∪{0}\mathbb{R}^{V \cup \{0\}}RV∪{0} have a distinguished coordinate x0x_0x0​; RV\mathbb{R}^VRV sits inside as the hyperplane H0={x0=1}H_0 = \{x_0 = 1\}H0​={x0​=1}, via x↦(1,x)x \mapsto (1, x)x↦(1,x).

  • FRAC(G)⊆RV\mathrm{FRAC}(G) \subseteq \mathbb{R}^VFRAC(G)⊆RV is the solution set of the nonnegativity constraints xi≥0x_i \ge 0xi​≥0 (i∈Vi \in Vi∈V) and the edge constraints xi+xj≤1x_i + x_j \le 1xi​+xj​≤1 (ij∈Eij \in Eij∈E).
  • FR(G)⊆RV∪{0}\mathrm{FR}(G) \subseteq \mathbb{R}^{V\cup\{0\}}FR(G)⊆RV∪{0} is the cone given by xi≥0x_i \ge 0xi​≥0 and xi+xj≤x0x_i + x_j \le x_0xi​+xj​≤x0​; it is the cone spanned by the vectors (1,x)(1, x)(1,x) with x∈FRAC(G)x \in \mathrm{FRAC}(G)x∈FRAC(G).
  • QQQ is the cone spanned by the 0–1 vectors with x0=1x_0 = 1x0​=1. For a convex cone KKK, its polar cone is K∗={u:uTx≥0 ∀x∈K}K^* = \{u : u^{\mathsf T}x \ge 0 \ \forall x \in K\}K∗={u:uTx≥0 ∀x∈K}.
  • M(K)=M(K,Q)M(K) = M(K, Q)M(K)=M(K,Q) is the set of (n+1)×(n+1)(n+1)\times(n+1)(n+1)×(n+1) matrices Y=(yij)Y = (y_{ij})Y=(yij​) that are symmetric, satisfy yii=y0iy_{ii} = y_{0i}yii​=y0i​ for i∈Vi \in Vi∈V, and satisfy uTYv≥0u^{\mathsf T} Y v \ge 0uTYv≥0 for all u∈K∗u \in K^*u∈K∗, v∈Q∗v \in Q^*v∈Q∗.
  • N(K)={Ye0:Y∈M(K)}N(K) = \{Y e_0 : Y \in M(K)\}N(K)={Ye0​:Y∈M(K)}, and N(G)={x∈RV:(1,x)∈N(FR(G))}N(G) = \{x \in \mathbb{R}^V : (1, x) \in N(\mathrm{FR}(G))\}N(G)={x∈RV:(1,x)∈N(FR(G))}.
  • A set C⊆VC \subseteq VC⊆V is an odd hole if it induces a chordless cycle of odd length ∣C∣≥3|C| \ge 3∣C∣≥3 (triangles included). Its odd hole constraint is ∑i∈Cxi≤12(∣C∣−1)\sum_{i \in C} x_i \le \frac12(|C| - 1)∑i∈C​xi​≤21​(∣C∣−1).

Formalization targets

Goal: Theorem 2.3 (p. 178)

For every finite graph GGG without isolated nodes,

N(G)={x∈RV:xi≥0 (i∈V),  xi+xj≤1 (ij∈E),  ∑i∈Cxi≤12(∣C∣−1) (C an odd hole)}.N(G) = \Big\{x \in \mathbb{R}^V : x_i \ge 0\ (i \in V),\ \ x_i + x_j \le 1\ (ij \in E),\ \ \sum_{i \in C} x_i \le \tfrac12(|C|-1)\ (C \text{ an odd hole})\Big\}.N(G)={x∈RV:xi​≥0 (i∈V),  xi​+xj​≤1 (ij∈E),  i∈C∑​xi​≤21​(∣C∣−1) (C an odd hole)}.

Milestones, in the order the proof uses them

  1. Lemma 1.3 (p. 171): for a convex cone K⊆QK \subseteq QK⊆Q and i∈Vi \in Vi∈V, N(K)⊆(K∩Hi)+(K∩Gi)N(K) \subseteq (K \cap H_i) + (K \cap G_i)N(K)⊆(K∩Hi​)+(K∩Gi​), with Hi={xi=0}H_i = \{x_i = 0\}Hi​={xi​=0}, Gi={xi=x0}G_i = \{x_i = x_0\}Gi​={xi​=x0​}.
  2. Lemma 2.2 (p. 178): if both the deletion and the contraction of some node vvv give inequalities valid for KKK, then aTx≤ba^{\mathsf T}x \le baTx≤b is valid for N(K)N(K)N(K).
  3. Part (1) of the proof of Theorem 2.3 (p. 178): for an odd hole CCC and i∈Ci \in Ci∈C, the deletion and contraction of iii in the odd hole constraint are valid for FRAC(G)\mathrm{FRAC}(G)FRAC(G).
  4. Observation of Section 2.b (p. 177): every Y∈M(FR(G))Y \in M(\mathrm{FR}(G))Y∈M(FR(G)) has yij=0y_{ij} = 0yij​=0 for ij∈Eij \in Eij∈E.
  5. Part (2) of the proof of Theorem 2.3 (p. 178): x∈N(G)x \in N(G)x∈N(G) if and only if some nonnegative symmetric YYY with y00=1y_{00} = 1y00​=1, yi0=yii=xiy_{i0} = y_{ii} = x_iyi0​=yii​=xi​ satisfies xi+xj+xk−1≤yik+yjk≤xkx_i + x_j + x_k - 1 \le y_{ik} + y_{jk} \le x_kxi​+xj​+xk​−1≤yik​+yjk​≤xk​ for all i,j,ki, j, ki,j,k with ij∈Eij \in Eij∈E.
  6. Lemma 2.4 (p. 178): a system a(ij)≤yi+yj≤b(ij)a(ij) \le y_i + y_j \le b(ij)a(ij)≤yi​+yj​≤b(ij), y≥0y \ge 0y≥0, y∣U=0y|_U = 0y∣U​=0 on a graph is infeasible if and only if a walk with a negative alternating sum of one of four types exists.

Significance

Theorem 2.3 gives a complete description of one round of NNN on the stable set problem: the only new constraints are the odd hole constraints. Consequences:

  • For ttt-perfect graphs (those for which nonnegativity, edge and odd hole constraints describe STAB(G)\mathrm{STAB}(G)STAB(G)), N(G)=STAB(G)N(G) = \mathrm{STAB}(G)N(G)=STAB(G).
  • It is the base case for the paper's bounds on the NNN-index of stable set inequalities (Theorem 2.13), and it contrasts with the semidefinite operator N+N_+N+​, which after one round already satisfies clique, odd antihole and wheel constraints.
  • Lemma 2.4 is a combinatorial feasibility criterion for systems with two variables per inequality, useful beyond this paper.

The result has been proved since 1991. At the time of drafting, Prove2Me holds no formalization of it or of any part of the Lovász–Schrijver construction, and Mathlib has none. The mission produces a formal account of the NNN operator on the stable set polytope and a formal proof of the walk criterion for two-variable systems.

Difficulty

The inclusion of N(G)N(G)N(G) in the odd hole system is a short argument once Lemma 1.3 is available. The reverse inclusion is the substance: given xxx satisfying all odd hole constraints, one must exhibit a lifted matrix YYY. A direct appeal to Farkas' lemma yields a certificate with no visible relation to odd cycles; the difficulty is to show that every obstruction to solvability of the matrix system forces a violated odd hole constraint, which is what Lemma 2.4 and the analysis of its four walk types accomplish. Case (d) of that analysis needs the odd hole constraints; the other cases need only the edge constraints. Lemma 2.4 itself is called folklore on the page and is stated without proof there.

A further point: Lemma 2.4 is stated for lower bounds 0≤a0 \le a0≤a, while the lower bounds that arise from the matrix system, xi+xj+xk−1x_i + x_j + x_k - 1xi​+xj​+xk​−1, can be negative.

Formalization scope

  • Coordinates of RV∪{0}\mathbb{R}^{V\cup\{0\}}RV∪{0} are indexed by Option V, with none the coordinate x0x_0x0​. Graphs are Mathlib SimpleGraphs on a finite type VVV with decidable adjacency. Every statement about a graph carries the paper's standing assumption that GGG has no isolated nodes (∀ v, ∃ w, G.Adj v w).
  • MMM is defined by condition (iii), never by its rewritings. Lemma 1.3 and Lemma 2.2 take the cone KKK closed, a hypothesis the paper leaves tacit (its cones are polyhedral); for a non-closed KKK Lemma 1.3 is false. FR(G)\mathrm{FR}(G)FR(G) is polyhedral, so the goal needs no such hypothesis.
  • FR(G)\mathrm{FR}(G)FR(G) is defined by its constraints; this agrees with the cone over FRAC(G)\mathrm{FRAC}(G)FRAC(G) because GGG has no isolated nodes.
  • Lemma 2.2 is stated in cone form: KKK is any closed convex cone inside FR(G)\mathrm{FR}(G)FR(G), and validity is read on the slice x0=1x_0 = 1x0​=1. The paper's extra hypothesis STAB(G)⊆K\mathrm{STAB}(G) \subseteq KSTAB(G)⊆K is dropped, which strengthens the lemma.
  • Deletion and contraction of a node are coefficient vectors on the same graph (coefficients set to 000), not inequalities on the subgraphs G−vG - vG−v and G−Γ(v)−vG - \Gamma(v) - vG−Γ(v)−v.
  • Odd holes are chordless odd cycles including triangles; triangles are needed, as 121\tfrac12\mathbf 121​1 satisfies all other constraints on a triangle.
  • The matrix system of part (2) is stated as an equivalence; the page uses one direction.
  • Lemma 2.4 uses edge values on unordered pairs and strict inequalities, exactly as printed.

A trivializing formalization is ruled out: the goal is the set equality for every graph without isolated nodes, not the existence of a lifted matrix and not a single graph.

Not formalized here: the semidefinite operator N+N_+N+​, the operator N^\hat NN^, algorithmic statements (Theorems 1.6, 2.1, Corollary 2.5), and the set-function results of Section 3.

Reusable beyond this mission: the matrix cone layer (QQQ, MMM, NNN), the stable-set cones, and the two-variable feasibility criterion of Lemma 2.4. Contributions of any of the milestones, and of general facts about polar cones of polyhedral cones in this setting, are welcome.

Selected references

  • L. Lovász and A. Schrijver, Cones of matrices and set-functions and 0–1 optimization, SIAM Journal on Optimization 1(2) (1991) 166–190. https://doi.org/10.1137/0801013
  • M. Grötschel, L. Lovász and A. Schrijver, Geometric Algorithms and Combinatorial Optimization, Springer, 1988. https://doi.org/10.1007/978-3-642-97881-4
  • H. D. Sherali and W. P. Adams, A hierarchy of relaxations between the continuous and convex hull representations for zero-one programming problems, SIAM Journal on Discrete Mathematics 3(3) (1990) 411–430. https://doi.org/10.1137/0403036
13 thms3 active usersReviewed
🏆Completed
Convex OptimizationLinear OptimizationOptimization·Captain: mikedeng1

Cones of Matrices and Set-Functions and 0–1 Optimization I: n Rounds of the Lovász–Schrijver N Operator Give the 0–1 HullResearch Paper

Motivation

A 0–1 integer program asks for the best 0–1 vector satisfying a system of linear inequalities. Its linear relaxation is easy to optimize over, but the relaxation is usually much larger than the convex hull of the 0–1 solutions. Lift-and-project methods close this gap systematically: they lift the relaxation to a higher-dimensional space, add constraints that every 0–1 point satisfies there, and project back, obtaining a tighter relaxation that still contains every 0–1 solution.

L. Lovász and A. Schrijver introduced one of the two standard lift-and-project hierarchies in Cones of matrices and set-functions and 0–1 optimization (SIAM J. Optim., 1991). Their operators NNN and N+N_+N+​ represent a 0–1 point xxx by the matrix xxTxx^{\mathsf T}xxT, impose linear (and for N+N_+N+​ semidefinite) constraints on such matrices, and project back to Rn+1\mathbb R^{n+1}Rn+1. The same paper applies the operators to the stable set polytope, where one round already produces the odd hole, odd wheel, clique and odd antihole constraints. The Lovász–Schrijver hierarchy, the Sherali–Adams hierarchy (1990) and Lasserre's semidefinite hierarchy (2001) are the three reference lift-and-project methods; their rank lower bounds are a standard tool for proving that a relaxation cannot solve a combinatorial problem in few rounds.

This mission formalizes the first structural fact about the operator NNN: iterating it nnn times on any relaxation in nnn variables yields exactly the 0–1 hull (Theorem 1.4 of the paper).

Setting

Vectors live in Rn+1\mathbb R^{n+1}Rn+1 with coordinates x0,x1,…,xnx_0, x_1, \dots, x_nx0​,x1​,…,xn​; the space Rn\mathbb R^nRn of the original problem is the hyperplane x0=1x_0 = 1x0​=1, and polytopes are replaced by the convex cones they generate.

  • A convex cone is a nonempty set closed under addition and nonnegative scaling. For a set SSS, cone⁡(S)\operatorname{cone}(S)cone(S) is the set of nonnegative combinations of finitely many vectors of SSS.
  • The polar cone of KKK is K∗={u:uTx≥0 for all x∈K}K^* = \{u : u^{\mathsf T}x \ge 0 \text{ for all } x \in K\}K∗={u:uTx≥0 for all x∈K}.
  • A 0–1 vector has every coordinate, x0x_0x0​ included, equal to 000 or 111. The cube cone QQQ is the cone spanned by the 0–1 vectors with x0=1x_0 = 1x0​=1; it is the cone over the unit cube.
  • For a convex cone KKK, K∘K^\circK∘ is the cone spanned by the 0–1 vectors in KKK. For K⊆QK \subseteq QK⊆Q this is the cone over the convex hull of the 0–1 points of the relaxation.

For convex cones K1,K2⊆QK_1, K_2 \subseteq QK1​,K2​⊆Q, the matrix cone M(K1,K2)M(K_1, K_2)M(K1​,K2​) consists of the (n+1)×(n+1)(n+1)\times(n+1)(n+1)×(n+1) real matrices Y=(yij)Y = (y_{ij})Y=(yij​) such that

  1. YYY is symmetric;
  2. yii=y0iy_{ii} = y_{0i}yii​=y0i​ for 1≤i≤n1 \le i \le n1≤i≤n (the diagonal equals the 0th column);
  3. uTYv≥0u^{\mathsf T} Y v \ge 0uTYv≥0 for every u∈K1∗u \in K_1^*u∈K1∗​ and v∈K2∗v \in K_2^*v∈K2∗​.

M+(K1,K2)M_+(K_1, K_2)M+​(K1​,K2​) adds the condition that YYY is positive semidefinite. The projections are N(K1,K2)={Ye0:Y∈M(K1,K2)}N(K_1, K_2) = \{Ye_0 : Y \in M(K_1, K_2)\}N(K1​,K2​)={Ye0​:Y∈M(K1​,K2​)} and N+(K1,K2)={Ye0:Y∈M+(K1,K2)}N_+(K_1, K_2) = \{Ye_0 : Y \in M_+(K_1, K_2)\}N+​(K1​,K2​)={Ye0​:Y∈M+​(K1​,K2​)}, where e0e_0e0​ is the 0th unit vector. The cut operator is N(K)=N(K,Q)N(K) = N(K, Q)N(K)=N(K,Q), and its iterates are N0(K)=KN^0(K) = KN0(K)=K, Nt(K)=N(Nt−1(K))N^t(K) = N(N^{t-1}(K))Nt(K)=N(Nt−1(K)).

Two families of hyperplanes appear in the proofs: Hi={x:xi=0}H_i = \{x : x_i = 0\}Hi​={x:xi​=0} and Gi={x:xi=x0}G_i = \{x : x_i = x_0\}Gi​={x:xi​=x0​}, the hyperplanes through the two opposite facets of QQQ in direction iii.

Formalization targets

Goal: Theorem 1.4

For every closed convex cone K⊆QK \subseteq QK⊆Q,

Nn(K)=K∘.N^n(K) = K^\circ .Nn(K)=K∘.

The statement is uniform in nnn and in KKK: no polyhedrality, no bound on the number of constraints, and no assumption that KKK contains a 0–1 point.

Milestones

  1. Condition (iii″). For a closed convex cone K⊆QK \subseteq QK⊆Q and a symmetric YYY with yii=y0iy_{ii} = y_{0i}yii​=y0i​: Y∈M(K,Q)Y \in M(K, Q)Y∈M(K,Q) if and only if every column of YYY is in KKK and the difference of the first column and any other column is in KKK.
  2. Lemma 1.1. For closed convex cones K1,K2⊆QK_1, K_2 \subseteq QK1​,K2​⊆Q,
(K1∩K2)∘⊆N+(K1,K2)⊆N(K1,K2)⊆K1∩K2.(K_1 \cap K_2)^\circ \subseteq N_+(K_1, K_2) \subseteq N(K_1, K_2) \subseteq K_1 \cap K_2 .(K1​∩K2​)∘⊆N+​(K1​,K2​)⊆N(K1​,K2​)⊆K1​∩K2​.
  1. Lemma 1.3. For a closed convex cone K⊆QK \subseteq QK⊆Q and every 1≤i≤n1 \le i \le n1≤i≤n,
N(K)⊆(K∩Hi)+(K∩Gi).N(K) \subseteq (K \cap H_i) + (K \cap G_i).N(K)⊆(K∩Hi​)+(K∩Gi​).
  1. Claim (4) in the proof of Theorem 1.4. For every set TTT of t≥1t \ge 1t≥1 coordinates, with Fˉ\bar FFˉ the union of the faces of the unit cube that fix the coordinates in TTT to 000 or 111,
Nt(K)⊆cone⁡(K∩Fˉ).N^t(K) \subseteq \operatorname{cone}(K \cap \bar F).Nt(K)⊆cone(K∩Fˉ).
  1. The remark after Lemma 1.1. N(K1∩K2,K1∩K2)⊆N(K1,K2)⊆N(K1∩K2,Q)N(K_1 \cap K_2, K_1 \cap K_2) \subseteq N(K_1, K_2) \subseteq N(K_1 \cap K_2, Q)N(K1​∩K2​,K1​∩K2​)⊆N(K1​,K2​)⊆N(K1​∩K2​,Q).

Significance

Theorem 1.4 is what makes NNN a hierarchy rather than a single cut: the relaxations K⊇N(K)⊇N2(K)⊇…K \supseteq N(K) \supseteq N^2(K) \supseteq \dotsK⊇N(K)⊇N2(K)⊇… reach the 0–1 hull after at most nnn rounds, so the NNN-rank of a valid inequality (the least ttt with the inequality valid for Nt(K)N^t(K)Nt(K)) is a well-defined number between 000 and nnn. The rest of the paper measures combinatorial constraints by this rank: odd hole constraints have rank one on the stable set polytope, and the rank of a stable set inequality is bounded by its defect. Rank lower bounds for lift-and-project hierarchies, in the literature that followed, all presuppose this finite convergence.

The theorem is proved in the paper; to the best of available knowledge none of the Lovász–Schrijver operators has been formalized in a proof assistant. A formalization provides machine-checked definitions of the matrix cones and the cut operators that later missions in this series (odd holes, the defect bound, the N+N_+N+​ constraints) state their results against, and a checked proof of the column characterization (iii″) that all of those proofs use.

Difficulty

The inclusion K∘⊆Nn(K)K^\circ \subseteq N^n(K)K∘⊆Nn(K) follows from Lemma 1.1 once each Nt(K)N^t(K)Nt(K) is known to be a convex cone. The reverse inclusion is the content. A first attempt shows that one round of NNN forces one coordinate to be integral, and then iterates; but N(K)N(K)N(K) is not contained in the union of K∩HiK \cap H_iK∩Hi​ and K∩GiK \cap G_iK∩Gi​, only in their Minkowski sum (Lemma 1.3), so a point of N(K)N(K)N(K) is not itself integral in any coordinate. The induction must carry a statement about cones spanned by intersections with unions of cube faces, and it needs each iterate Nt(K)N^t(K)Nt(K) to again be a closed convex cone inside QQQ so that Lemma 1.3 can be reapplied. Closedness of the projection N(K)N(K)N(K) is not automatic: a linear image of a closed cone need not be closed.

Formalization scope

  • Coordinates of Rn+1\mathbb R^{n+1}Rn+1 are indexed by Option ι for a finite type ι; none is x0x_0x0​ and some i is xix_ixi​, and nnn is the cardinality of ι, which may be 000.
  • cone⁡(S)\operatorname{cone}(S)cone(S) is Mathlib's PointedCone.hull ℝ S; QQQ and K∘K^\circK∘ are defined as spans of 0–1 vectors, as on the page, not by the inequality description 0≤xi≤x00 \le x_i \le x_00≤xi​≤x0​.
  • M(K1,K2)M(K_1, K_2)M(K1​,K2​) is defined by condition (iii) through the polar cones; the column form (iii″) is a milestone, not the definition.
  • The operators NNN, N+N_+N+​ and the iterates are defined on arbitrary sets; the hypotheses (convex cone, contained in QQQ, closed) are carried by the theorems.
  • Closedness. The paper tacitly takes its cones closed (they are polyhedral in all its applications), and the rewriting (iii′) on p. 169 needs it. Every statement here assumes the cones closed. Without this the goal is false: for K={x:0<x1<x0}∪{0}K = \{x : 0 < x_1 < x_0\} \cup \{0\}K={x:0<x1​<x0​}∪{0} in R2\mathbb R^2R2, K∘={0}K^\circ = \{0\}K∘={0} while N(K)=QN(K) = QN(K)=Q.
  • In the proof of Theorem 1.4 the page places the cube Q′Q'Q′ in the hyperplane "x0=0x_0 = 0x0​=0"; this is a misprint for x0=1x_0 = 1x0​=1, and claim (4) is formalized with x0=1x_0 = 1x0​=1.
  • Not formalized in this mission: Lemma 1.2 (the dual description of N(K)∗N(K)^*N(K)∗), Lemma 1.5 (the N+N_+N+​ analogue of Lemma 1.3, part of a later mission), and the algorithmic results of Section 1.c.

Contributions welcome: proofs that N(K)N(K)N(K) is a closed convex cone contained in QQQ whenever KKK is, a proof of Q∗=cone⁡{ei,e0−ei}Q^* = \operatorname{cone}\{e_i, e_0 - e_i\}Q∗=cone{ei​,e0​−ei​}, and lemmas on cones spanned by the intersection of a generating set with a supporting hyperplane; these are reusable by the other missions of the series.

Selected references

  • L. Lovász and A. Schrijver, Cones of matrices and set-functions and 0–1 optimization, SIAM Journal on Optimization 1(2) (1991) 166–190. https://doi.org/10.1137/0801013
  • H. D. Sherali and W. P. Adams, A hierarchy of relaxations between the continuous and convex hull representations for zero-one programming problems, SIAM Journal on Discrete Mathematics 3(3) (1990) 411–430. https://doi.org/10.1137/0403036
  • J. B. Lasserre, Global optimization with polynomials and the problem of moments, SIAM Journal on Optimization 11(3) (2001) 796–817. https://doi.org/10.1137/S1052623400366802
  • M. Laurent, A comparison of the Sherali–Adams, Lovász–Schrijver, and Lasserre relaxations for 0–1 programming, Mathematics of Operations Research 28(3) (2003) 470–496. https://doi.org/10.1287/moor.28.3.470.16391
8 thms3 active usersReviewed
🏆Completed
Control TheoryTheoretical Computer Science·Captain: mikedeng1

Supervisory Control of a Class of Discrete Event Processes I: Minimally Restrictive Supervisors Exist iff the Supremal Controllable Legal Sublanguage Contains the Minimal Acceptable LanguageResearch Paper

Motivation

Manufacturing cells, communication protocols, traffic systems and database transaction managers are naturally described not by differential equations but by sequences of discrete events: a machine starts, a part arrives, a message is lost. Ramadge and Wonham's 1987 paper (SIAM J. Control Optim. 25(1)) set up a control theory for such systems in which the plant is an automaton, the controller is another automaton that may disable some events, and specifications are formal languages. The framework, now called supervisory control theory or the Ramadge–Wonham framework, is the standard model for the logical control of discrete event systems and underlies the textbook treatment in Cassandras and Lafortune (Introduction to Discrete Event Systems, 2008) and Wonham and Cai (Supervisory Control of Discrete-Event Systems, 2019).

The question the paper answers is the basic synthesis question of the theory: given a plant, a set of legal behaviours and a set of minimally acceptable behaviours, when does a controller exist that keeps the plant legal, achieves at least the acceptable behaviour, and never deadlocks, and what is the least restrictive such controller?

Setting

A generator is G=(Q,Σ,δ,q0,Qm)\mathcal G = (Q, \Sigma, \delta, q_0, Q_m)G=(Q,Σ,δ,q0​,Qm​) with a state set QQQ, a finite alphabet Σ\SigmaΣ of events, a partial transition function δ:Σ×Q→Q\delta : \Sigma \times Q \to Qδ:Σ×Q→Q, an initial state q0q_0q0​ and marker states Qm⊆QQ_m \subseteq QQm​⊆Q. Extending δ\deltaδ to strings, the generated language L(G)L(\mathcal G)L(G) is the set of strings www for which δ(w,q0)\delta(w, q_0)δ(w,q0​) is defined, and the marked language Lm(G)L_m(\mathcal G)Lm​(G) is the subset of those that end in QmQ_mQm​. The closure Kˉ\bar KKˉ of a language KKK is its set of prefixes; KKK is closed if K=KˉK = \bar KK=Kˉ. Throughout, G\mathcal GG is assumed trim in the sense L(G)=Lˉm(G)L(\mathcal G) = \bar L_m(\mathcal G)L(G)=Lˉm​(G): every generated string can be completed to a marked one.

The alphabet is split into controllable events Σc\Sigma_cΣc​ and uncontrollable events Σu=Σ−Σc\Sigma_u = \Sigma - \Sigma_cΣu​=Σ−Σc​. A supervisor S=(S,ϕ)\mathcal S = (S, \phi)S=(S,ϕ) consists of a deterministic, accessible automaton S=(X,Σ,ξ,x0,Xm)S = (X, \Sigma, \xi, x_0, X_m)S=(X,Σ,ξ,x0​,Xm​), whose state set XXX may be infinite, and a map ϕ\phiϕ assigning to each state xxx the set of controllable events it enables; uncontrollable events are always enabled. The closed loop S/G\mathcal S/\mathcal GS/G runs SSS and G\mathcal GG in lockstep: an event occurs when the plant can execute it, the supervisor enables it, and the supervisor's automaton can follow it. This defines the languages L(S/G)L(\mathcal S/\mathcal G)L(S/G), Lm(S/G)L_m(\mathcal S/\mathcal G)Lm​(S/G) and the controlled language Lc(S/G)=L(S/G)∩Lm(G)L_c(\mathcal S/\mathcal G) = L(\mathcal S/\mathcal G) \cap L_m(\mathcal G)Lc​(S/G)=L(S/G)∩Lm​(G).

S\mathcal SS is complete if its automaton never refuses an event that the plant can execute and ϕ\phiϕ enables; it is proper if it is complete and

Lˉm(S/G)=Lˉc(S/G)=L(S/G),\bar L_m(\mathcal S/\mathcal G) = \bar L_c(\mathcal S/\mathcal G) = L(\mathcal S/\mathcal G),Lˉm​(S/G)=Lˉc​(S/G)=L(S/G),

that is, every closed-loop string can be completed to a marked task. A language KKK is controllable if K⊆L(G)K \subseteq L(\mathcal G)K⊆L(G) and KˉΣu∩L(G)⊆Kˉ\bar K \Sigma_u \cap L(\mathcal G) \subseteq \bar KKˉΣu​∩L(G)⊆Kˉ. For L⊆L(G)L \subseteq L(\mathcal G)L⊆L(G), CG(L)\mathbf C_{\mathcal G}(L)CG​(L) is the class of controllable sublanguages of LLL and FG(L)\mathbf F_{\mathcal G}(L)FG​(L) the class of sublanguages KKK of LLL with K=Kˉ∩Lm(G)K = \bar K \cap L_m(\mathcal G)K=Kˉ∩Lm​(G).

Given ∅≠La⊆Lg⊆Lm(G)\emptyset \neq L_a \subseteq L_g \subseteq L_m(\mathcal G)∅=La​⊆Lg​⊆Lm​(G), the Supervisory Marking Problem (SMP) asks for a proper S\mathcal SS with La⊆Lm(S/G)⊆LgL_a \subseteq L_m(\mathcal S/\mathcal G) \subseteq L_gLa​⊆Lm​(S/G)⊆Lg​, and the Supervisory Control Problem (SCP) for a proper S\mathcal SS with La⊆Lc(S/G)⊆LgL_a \subseteq L_c(\mathcal S/\mathcal G) \subseteq L_gLa​⊆Lc​(S/G)⊆Lg​.

Formalization targets

Goal: Theorem 7.1 (pp. 218–219)

SMP solvable  ⟺  sup⁡CG(Lg)⊇La,SCP solvable  ⟺  sup⁡{CG(Lg)∩FG(Lg)}⊇La,\text{SMP solvable} \iff \sup \mathbf C_{\mathcal G}(L_g) \supseteq L_a, \qquad \text{SCP solvable} \iff \sup\{\mathbf C_{\mathcal G}(L_g) \cap \mathbf F_{\mathcal G}(L_g)\} \supseteq L_a,SMP solvable⟺supCG​(Lg​)⊇La​,SCP solvable⟺sup{CG​(Lg​)∩FG​(Lg​)}⊇La​,

and in each case the solving supervisor can be taken minimally restrictive: its marked (respectively controlled) language equals the supremal element and contains that of every proper supervisor whose language lies in LgL_gLg​.

Milestones

  1. Proposition 4.1 (i), (ii): marking is independent of control; every K⊆Lm(G)K \subseteq L_m(\mathcal G)K⊆Lm​(G), or K⊆L∩Lm(G)K \subseteq L \cap L_m(\mathcal G)K⊆L∩Lm​(G) for an achievable closed LLL, is the marked language of a complete supervisor.
  2. Proposition 5.1: a complete supervisor realizes (Lm,Lc,L)=(K1,K2,K3)(L_m, L_c, L) = (K_1, K_2, K_3)(Lm​,Lc​,L)=(K1​,K2​,K3​) iff K1⊆K2K_1 \subseteq K_2K1​⊆K2​, K2=K3∩Lm(G)K_2 = K_3 \cap L_m(\mathcal G)K2​=K3​∩Lm​(G) and K3K_3K3​ is closed and controllable.
  3. Theorem 6.1 (i), (ii): a proper supervisor with Lm(S/G)=KL_m(\mathcal S/\mathcal G) = KLm​(S/G)=K exists iff KKK is controllable; one with Lc(S/G)=KL_c(\mathcal S/\mathcal G) = KLc​(S/G)=K exists iff KKK is controllable and Lm(G)L_m(\mathcal G)Lm​(G)-closed.
  4. Proposition 7.1: CG(L)\mathbf C_{\mathcal G}(L)CG​(L) and FG(L)\mathbf F_{\mathcal G}(L)FG​(L) contain ∅\emptyset∅ and are closed under arbitrary unions.
  5. The supremal elements (p. 218): sup⁡CG(L)\sup \mathbf C_{\mathcal G}(L)supCG​(L), sup⁡FG(L)\sup \mathbf F_{\mathcal G}(L)supFG​(L) and sup⁡{CG(L)∩FG(L)}\sup\{\mathbf C_{\mathcal G}(L) \cap \mathbf F_{\mathcal G}(L)\}sup{CG​(L)∩FG​(L)} belong to their classes.

Significance

Theorem 7.1 reduces the existence of a correct, non-blocking controller to a single language inclusion, and identifies the supremal controllable sublanguage as the behaviour of the least restrictive solution. That object is the backbone of the later theory: modular and decentralized control, control under partial observation, and the computational results that sup⁡CG(L)\sup \mathbf C_{\mathcal G}(L)supCG​(L) is regular and computable when G\mathcal GG is finite and LLL regular all start from it.

The result is proved in the paper and has been taught for decades; it is not open. Mathlib has no model of generators with partial transitions, supervisors or controllability (its DFA has a total transition function and a single accepted language), and no formalization of this theory exists on the platform. This mission provides one: a reusable Lean model of generators with partial transitions, supervisors with possibly infinite state, closed loops and controllability, together with the paper's existence theorems stated against it. A second mission on the same paper (quotients of supervisors, Theorem 10.1) builds on the same objects.

Difficulty

The combinatorial core of Proposition 7.1 is a short prefix-closure computation. The work lies in the constructions: to show existence, a supervisor must be built for an arbitrary controllable language, which in general is not regular, so the supervisor needs an infinite state set (for example strings of the target language) together with a proof that the closed loop generates exactly the intended language, is complete, and is proper. The converse directions require relating the closed-loop run to separate runs of the plant and of the supervisor's automaton. The naive shortcut of reading the "sup" as an arbitrary member of CG(Lg)\mathbf C_{\mathcal G}(L_g)CG​(Lg​) containing LaL_aLa​ skips the content of the supremal-element milestone; the goal is stated with the actual union.

Formalization scope

  • The alphabet is a type α with [Fintype α]; strings are List α, the empty string is [], and sσs\sigmasσ is s ++ [σ]. Languages are Set (List α); the closure is pre K = {s | ∃ t, s ++ t ∈ K}.
  • A generator is a structure with a state type Q : Type, a partial transition δ : α → Q → Option Q, an initial state and a marker set. No finiteness is assumed on states, of the plant or of the supervisor. Supervisors and generators live in Type 1.
  • ϕ\phiϕ maps states to Ec → Bool; an event outside Σc\Sigma_cΣc​ is enabled by definition, so uncontrollable events cannot be disabled.
  • The closed loop is the product run from (x0,q0)(x_0, q_0)(x0​,q0​); the accessible part in the paper's display (2.1) changes no language and is not built.
  • Standing assumptions carried by every theorem: Σ\SigmaΣ finite, L(G)=Lˉm(G)L(\mathcal G) = \bar L_m(\mathcal G)L(G)=Lˉm​(G), supervisor automata accessible (as a hypothesis on every input supervisor and a conjunct of every "there exists a supervisor").
  • sup⁡\supsup is sSup in the complete lattice Set (List α).
  • A formalization in which the closed loop ignores ϕ\phiϕ or the plant, in which completeness is dropped from properness, or in which the supervisor is restricted to finitely many states, is not the paper's theorem and is ruled out by the definitions above.

Welcome contributions: the basic run lemmas (closed-loop runs project to plant and supervisor runs), the string-state supervisor construction, and proofs of the milestones in the given order.

Selected references

  • P. J. Ramadge and W. M. Wonham, Supervisory Control of a Class of Discrete Event Processes, SIAM J. Control Optim. 25(1):206–230, 1987. https://doi.org/10.1137/0325013
  • W. M. Wonham and P. J. Ramadge, On the Supremal Controllable Sublanguage of a Given Language, SIAM J. Control Optim. 25(3):637–659, 1987. https://doi.org/10.1137/0325036
  • C. G. Cassandras and S. Lafortune, Introduction to Discrete Event Systems, 2nd ed., Springer, 2008. https://doi.org/10.1007/978-0-387-68612-7
  • W. M. Wonham and K. Cai, Supervisory Control of Discrete-Event Systems, Springer, 2019. https://doi.org/10.1007/978-3-319-77452-7
12 thms3 active usersReviewed
CombinatoricsGraph TheoryLinear algebra+1·Captain: mikedeng1

Approximating Clique-Width and Branch-Width: Well-Linked Sets Certify Clique-WidthResearch Paper

Motivation

Clique-width is a graph parameter introduced by Courcelle and Olariu (Discrete Appl. Math. 101 (2000)) that measures how far a graph is from being built by a few labelled operations. Every problem expressible in monadic second-order logic with quantification over vertices and vertex sets (MSO1_11​) can be solved in linear time on graphs given together with a decomposition of bounded clique-width (Courcelle, Makowsky and Rotics, Theory Comput. Syst. 33 (2000)). Bounded clique-width is more general than bounded tree-width: complete graphs have unbounded tree-width but clique-width 222.

For fixed kkk there was, before this paper, no polynomial-time algorithm that either decides that a graph has clique-width at least k+1k+1k+1 or outputs a decomposition of clique-width bounded by a function of kkk; the best known algorithm, by Johansson (2001), gave width 2klog⁡n2k\log n2klogn. Oum and Seymour (J. Combin. Theory Ser. B 96 (2006)) closed this gap with approximation 23k+2−12^{3k+2}-123k+2−1, through rank-width and a factor-3 approximation for the branch-width of symmetric submodular functions.

Timeline:

  • 1991: Robertson and Seymour introduce branch-width of graphs and hypergraphs (J. Combin. Theory Ser. B 52).
  • 2000: Courcelle and Olariu define clique-width; Courcelle, Makowsky and Rotics solve MSO1_11​ problems on graphs given with a kkk-expression.
  • 2001: Johansson gives a 2klog⁡n2k\log n2klogn approximation.
  • 2006: Oum and Seymour define rank-width, prove rwd(G)≤cwd(G)≤2rwd(G)+1−1\mathrm{rwd}(G) \le \mathrm{cwd}(G) \le 2^{\mathrm{rwd}(G)+1}-1rwd(G)≤cwd(G)≤2rwd(G)+1−1, and give an O(n9log⁡n)O(n^9 \log n)O(n9logn) algorithm that outputs a (23k+2−1)(2^{3k+2}-1)(23k+2−1)-expression or certifies clique-width above kkk.

Setting

All graphs are finite and simple. For a finite set VVV, a function f:2V→Zf : 2^V \to \mathbb{Z}f:2V→Z is submodular if f(X)+f(Y)≥f(X∩Y)+f(X∪Y)f(X)+f(Y) \ge f(X\cap Y)+f(X\cup Y)f(X)+f(Y)≥f(X∩Y)+f(X∪Y) and symmetric if f(X)=f(V∖X)f(X) = f(V\setminus X)f(X)=f(V∖X).

A branch-decomposition of fff is a pair (T,L)(T, L)(T,L) where TTT is a tree with at least two vertices and all degrees at most 333, and LLL is a bijection from VVV onto the leaves of TTT. Removing an edge eee of TTT splits the leaves in two; the width of eee is fff of the set of elements of VVV on one side. The width of (T,L)(T, L)(T,L) is the largest edge width, and the branch-width bw(f)\mathrm{bw}(f)bw(f) is the least width of a branch-decomposition, with bw(f)=f(∅)\mathrm{bw}(f) = f(\emptyset)bw(f)=f(∅) when ∣V∣≤1|V| \le 1∣V∣≤1.

A set W⊆VW \subseteq VW⊆V is well-linked with respect to fff if for every partition (X,Y)(X, Y)(X,Y) of WWW and every ZZZ with X⊆Z⊆V∖YX \subseteq Z \subseteq V\setminus YX⊆Z⊆V∖Y, f(Z)≥min⁡(∣X∣,∣Y∣)f(Z) \ge \min(|X|, |Y|)f(Z)≥min(∣X∣,∣Y∣).

Let A(G)A(G)A(G) be the adjacency matrix of GGG over GF(2)\mathrm{GF}(2)GF(2). For disjoint X,Y⊆V(G)X, Y \subseteq V(G)X,Y⊆V(G), cutrkG∗(X,Y)\mathrm{cutrk}^*_G(X, Y)cutrkG∗​(X,Y) is the rank of the submatrix of A(G)A(G)A(G) with rows XXX and columns YYY, and the cut-rank function is cutrkG(X)=cutrkG∗(X,V(G)∖X)\mathrm{cutrk}_G(X) = \mathrm{cutrk}^*_G(X, V(G)\setminus X)cutrkG​(X)=cutrkG∗​(X,V(G)∖X). The rank-width rwd(G)\mathrm{rwd}(G)rwd(G) is bw(cutrkG)\mathrm{bw}(\mathrm{cutrk}_G)bw(cutrkG​).

A kkk-expression is a term built from constants ⋅i\cdot_i⋅i​ (a vertex with label i∈{1,…,k}i \in \{1,\dots,k\}i∈{1,…,k}), the operators ηi,j\eta_{i,j}ηi,j​ (i≠ji \ne ji=j; add all edges between labels iii and jjj), ρi→j\rho_{i\to j}ρi→j​ (relabel iii into jjj) and disjoint union ⊕\oplus⊕. Its value is the labelled graph it produces; GGG has clique-width cwd(G)≤k\mathrm{cwd}(G) \le kcwd(G)≤k if some kkk-expression has value isomorphic to GGG.

An interpolation of fff is a function f∗f^*f∗ on disjoint pairs (X,Y)(X, Y)(X,Y) that agrees with fff on (X,V∖X)(X, V\setminus X)(X,V∖X), is monotone, submodular in the sense f∗(A,B)+f∗(C,D)≥f∗(A∩C,B∪D)+f∗(A∪C,B∩D)f^*(A,B)+f^*(C,D) \ge f^*(A\cap C, B\cup D) + f^*(A\cup C, B\cap D)f∗(A,B)+f∗(C,D)≥f∗(A∩C,B∪D)+f∗(A∪C,B∩D), and has f∗(∅,∅)=f(∅)f^*(\emptyset,\emptyset)=f(\emptyset)f∗(∅,∅)=f(∅).

Formalization targets

Goal: Theorem 1.1, certificate form

For a graph GGG with at least one vertex and an integer k≥1k \ge 1k≥1:

∃ W, ∣W∣=3k+1, W well-linked for cutrkG  ⟹  cwd(G)≥k+1,\exists\, W,\ |W| = 3k+1,\ W \text{ well-linked for } \mathrm{cutrk}_G \;\Longrightarrow\; \mathrm{cwd}(G) \ge k+1,∃W, ∣W∣=3k+1, W well-linked for cutrkG​⟹cwd(G)≥k+1, ∄ W, ∣W∣=3k+1, W well-linked for cutrkG  ⟹  cwd(G)≤23k+2−1.\nexists\, W,\ |W| = 3k+1,\ W \text{ well-linked for } \mathrm{cutrk}_G \;\Longrightarrow\; \mathrm{cwd}(G) \le 2^{3k+2}-1.∄W, ∣W∣=3k+1, W well-linked for cutrkG​⟹cwd(G)≤23k+2−1.

The same explicit condition decides which side of the approximation holds; this is what the paper's algorithm certifies.

Milestones

  1. Proposition 4.1: properties of an interpolation, including that X↦f∗(X,B)−f(∅)X \mapsto f^*(X, B) - f(\emptyset)X↦f∗(X,B)−f(∅) is a matroid rank function on V∖BV\setminus BV∖B when f({v})−f(∅)≤1f(\{v\}) - f(\emptyset) \le 1f({v})−f(∅)≤1.
  2. Proposition 4.2: fmin⁡(X,Y)=min⁡X⊆Z⊆V∖Yf(Z)f_{\min}(X,Y) = \min_{X\subseteq Z\subseteq V\setminus Y} f(Z)fmin​(X,Y)=minX⊆Z⊆V∖Y​f(Z) is an interpolation.
  3. Theorem 5.1: a well-linked set of size kkk forces bw(f)≥k/3\mathrm{bw}(f) \ge k/3bw(f)≥k/3 (for k≠1k \ne 1k=1).
  4. Theorem 5.2: no well-linked set of size kkk implies bw(f)≤k\mathrm{bw}(f) \le kbw(f)≤k, when f({v})≤1f(\{v\}) \le 1f({v})≤1.
  5. Proposition 6.1: rk M[X1,Y1]+rk M[X2,Y2]≥rk M[X1∪X2,Y1∩Y2]+rk M[X1∩X2,Y1∪Y2]\mathrm{rk}\,M[X_1,Y_1] + \mathrm{rk}\,M[X_2,Y_2] \ge \mathrm{rk}\,M[X_1\cup X_2, Y_1\cap Y_2] + \mathrm{rk}\,M[X_1\cap X_2, Y_1\cup Y_2]rkM[X1​,Y1​]+rkM[X2​,Y2​]≥rkM[X1​∪X2​,Y1​∩Y2​]+rkM[X1​∩X2​,Y1​∪Y2​].
  6. Corollary 6.2: submodularity of cutrkG∗\mathrm{cutrk}^*_GcutrkG∗​ and cutrkG\mathrm{cutrk}_GcutrkG​.
  7. Section 6 claim: cutrkG\mathrm{cutrk}_GcutrkG​ is symmetric submodular and cutrkG∗\mathrm{cutrk}^*_GcutrkG∗​ interpolates it.
  8. Proposition 6.3: rwd(G)≤cwd(G)≤2rwd(G)+1−1\mathrm{rwd}(G) \le \mathrm{cwd}(G) \le 2^{\mathrm{rwd}(G)+1}-1rwd(G)≤cwd(G)≤2rwd(G)+1−1.

Significance

The dichotomy turns clique-width, for which no exact polynomial algorithm is known even for fixed kkk, into a parameter that can be approximated with an explicit witness in each direction. Downstream, every algorithm for graphs of bounded clique-width that needs a kkk-expression as input becomes applicable to graphs given without one, at the cost of an exponential blow-up of the width.

The result is proved in the literature; this mission formalizes it. To our knowledge none of the objects involved — branch-width of set functions, rank-width, cut-rank, kkk-expressions, clique-width — has been formalized in Mathlib, and the submodularity of submatrix rank (Proposition 6.1) is absent from Mathlib's Matrix.rank API. The formal development would give reusable definitions of branch-decompositions of arbitrary integer set functions, of cut-rank, and of clique-width, and a machine-checked link between the combinatorial and the linear-algebraic width parameters.

Difficulty

The upper bound in Theorem 5.2 is the core. The natural approach, growing a branch-decomposition one leaf split at a time while keeping the width at most kkk, gets stuck at a leaf carrying a set BBB with f(B)=kf(B) = kf(B)=k: a split of BBB into two parts of fff-value below kkk has to be found, and it must be found from the failure of well-linkedness of a set that is not obviously related to BBB. The paper's device is the interpolation f∗f^*f∗, which attaches a matroid to BBB whose base has exactly f(B)f(B)f(B) elements. Formalizing this requires handling partial branch-decompositions, their extensions, and a maximality argument over trees, none of which exists in Mathlib.

Proposition 6.3's upper bound is a second, independent difficulty: a rank-decomposition must be converted into a kkk-expression by an induction over a rooted binary tree, with a relabelling argument bounding the number of labels by the number of distinct nonzero rows of a GF(2)\mathrm{GF}(2)GF(2) matrix of rank kkk. Its lower bound needs the tree structure of a kkk-expression to be read as a branch-decomposition.

Formalization scope

The ground set is a Fintype V with DecidableEq V; subsets are Finset V; set functions are Finset V → ℤ, as in the paper. A branch-decomposition is a tree T : SimpleGraph (Fin n) with n≥2n \ge 2n≥2, all neighbour sets of size at most 333, and an injective map LLL from VVV onto the vertices of degree 111; the side of an edge uwuwuw is found by reachability from uuu after deleting uwuwuw. Branch-width, rank-width and clique-width are never computed as minima: "bw(f)≤k\mathrm{bw}(f) \le kbw(f)≤k" is the predicate "∣V∣≤1|V| \le 1∣V∣≤1 and f(∅)≤kf(\emptyset) \le kf(∅)≤k, or a branch-decomposition of width at most kkk exists", lower bounds say that every branch-decomposition has a wide edge, and "cwd(G)≤k\mathrm{cwd}(G) \le kcwd(G)≤k" is "GGG has a kkk-expression". Labels {1,…,k}\{1,\dots,k\}{1,…,k} are Fin k. The value of a kkk-expression has as vertex type the occurrences of constants (a nested sum type), and ηi,j\eta_{i,j}ηi,j​ requires i≠ji \ne ji=j. Cut-rank uses Matrix.rank over ZMod 2 of submatrices of SimpleGraph.adjMatrix. An interpolation is a function on all pairs of subsets whose axioms are imposed on disjoint pairs only.

Running time is not formalized. The paper's Theorem 1.1 asserts an O(n9log⁡n)O(n^9\log n)O(n9logn) algorithm; there is no cost model on the page, and the goal states the certificate the algorithm returns instead. Without the running time, "cwd(G)≥k+1\mathrm{cwd}(G) \ge k+1cwd(G)≥k+1 or cwd(G)≤23k+2−1\mathrm{cwd}(G) \le 2^{3k+2}-1cwd(G)≤23k+2−1" holds for every graph, so that reading is ruled out as a formalization of the goal; so are well-linkedness with respect to anything other than cutrkG\mathrm{cutrk}_GcutrkG​, widths defined by an unguarded infimum (which is 000 on an empty family), kkk-expressions whose value is not the graph up to isomorphism or whose η\etaη may join equal labels, and Theorem 5.1 stated for k=1k = 1k=1.

Correction of Theorem 5.1. As printed, Theorem 5.1 fails for k=1k = 1k=1: a singleton is always well-linked, but the edgeless graph on two vertices has cut-rank identically 000 and branch-width 0<1/30 < 1/30<1/3. The milestone carries the hypothesis k≠1k \ne 1k=1; the goal uses the theorem only at size 3k+1≥43k+1 \ge 43k+1≥4.

The graph with no vertex is excluded from the goal and from the upper bound of Proposition 6.3, since it has no kkk-expression for any kkk. Contributions welcome: proofs of the milestones, lemmas on branch-decompositions (suppressing degree-2 vertices, extending partial decompositions), and submatrix-rank submodularity, which is reusable beyond this mission.

Selected references

  • S. Oum and P. Seymour, Approximating clique-width and branch-width, J. Combin. Theory Ser. B 96 (2006) 514–528. https://doi.org/10.1016/j.jctb.2005.10.006
  • B. Courcelle and S. Olariu, Upper bounds to the clique width of graphs, Discrete Appl. Math. 101 (2000) 77–114. https://doi.org/10.1016/S0166-218X(99)00184-5
  • B. Courcelle, J. A. Makowsky and U. Rotics, Linear time solvable optimization problems on graphs of bounded clique-width, Theory Comput. Syst. 33 (2000) 125–150. https://doi.org/10.1007/s002249910009
  • N. Robertson and P. D. Seymour, Graph minors. X. Obstructions to tree-decomposition, J. Combin. Theory Ser. B 52 (1991) 153–190. https://doi.org/10.1016/0095-8956(91)90061-N
14 thms3 active usersReviewed
🏆Completed
CombinatoricsOptimizationTheoretical Computer Science·Captain: mikedeng1

Approximation Algorithms for Combinatorial Problems V: The Overlap-Ratio Greedy C2 Is Within 1 + ln k of the Least-Overlap Cover on EC(k)Research Paper

Motivation

David S. Johnson's 1974 paper Approximation Algorithms for Combinatorial Problems (J. Comput. System Sci. 9 (1974) 256–278) was one of the first systematic worst-case analyses of polynomial-time heuristics for NP-complete optimization problems. Its Section 5 proves the harmonic bound ∑j=1k1/j\sum_{j=1}^k 1/j∑j=1k​1/j for the greedy algorithm on minimum-cardinality set cover, a result that still underlies the standard ln⁡n\ln nlnn approximation guarantee.

Section 6, the subject of this mission, asks what happens when the cost of a cover is its total size rather than its number of sets. This problem, SET COVERING II (EC), is the optimization version of the EXACT COVER recognition problem of Karp's list (Karp 1972): a family has a disjoint subcover exactly when the optimum equals the number of covered points. Johnson shows that the change of measure breaks the cardinality greedy but that a greedy rule based on an overlap ratio recovers essentially the same guarantee. The same accounting (paying for each newly covered point) later became the standard analysis of greedy weighted set cover (Chvátal 1979).

Setting

An input is a finite family F={S1,…,Sp}F = \{S_1, \dots, S_p\}F={S1​,…,Sp​} of finite sets. Its covered set is T=⋃S∈FST = \bigcup_{S \in F} ST=⋃S∈F​S. A subcover is a subfamily F′⊆FF' \subseteq FF′⊆F with ⋃S∈F′S=T\bigcup_{S\in F'} S = T⋃S∈F′​S=T, and its measure is

mEC(F′)=∑S∈F′∣S∣.m_{EC}(F') = \sum_{S \in F'} |S|.mEC​(F′)=S∈F′∑​∣S∣.

The optimum F∗F^*F∗ is the least measure of a subcover; since every subcover has measure at least ∣T∣|T|∣T∣, an optimal subcover is one with the least possible overlapping. The subproblem EC(k)(k)(k) admits only families in which every set has at most kkk points.

Algorithm C2 keeps a subfamily SUB (initially empty), the unused sets LEFT (initially FFF) and the uncovered points UNCOV (initially TTT). While UNCOV is nonempty it chooses S′∈S' \inS′∈ LEFT minimizing

Ratio(S)=∣S−UNCOV∣∣S∩UNCOV∣,\mathrm{Ratio}(S) = \frac{|S - \mathrm{UNCOV}|}{|S \cap \mathrm{UNCOV}|},Ratio(S)=∣S∩UNCOV∣∣S−UNCOV∣​,

the number of already-covered points of SSS per newly covered point, and moves S′S'S′ from LEFT to SUB, removing its points from UNCOV. When several sets tie, any of them may be chosen; a subcover is choosable by C2 if some sequence of admissible choices returns it.

The overlap of a chosen set is ∣S′−UNCOV∣|S' - \mathrm{UNCOV}|∣S′−UNCOV∣ at the moment it is chosen, and the cumulative overlap OV(F1)\mathrm{OV}(F_1)OV(F1​) of a run returning F1F_1F1​ is the sum of these overlaps.

Formalization targets

Goal: Theorem 6 (p. 271)

For all k≥1k \ge 1k≥1 and n>0n > 0n>0,

R[C2,EC(k)](n)≤1+ln⁡(k)≤∑j=1k1j+12,R[C2, EC(k)](n) \le 1 + \ln(k) \le \sum_{j=1}^k \frac1j + \frac12,R[C2,EC(k)](n)≤1+ln(k)≤j=1∑k​j1​+21​,

and for all sufficiently large nnn, R[C2,EC(k)](n)≥∑j=1k(1/j)R[C2, EC(k)](n) \ge \sum_{j=1}^k (1/j)R[C2,EC(k)](n)≥∑j=1k​(1/j). In the size-free form used here, for every k≥1k \ge 1k≥1:

  1. every subcover MMM choosable by C2 on an input of EC(k)(k)(k) satisfies mEC(M)≤(1+ln⁡k) F∗m_{EC}(M) \le (1 + \ln k)\,F^*mEC​(M)≤(1+lnk)F∗;
  2. 1+ln⁡k≤∑j=1k1/j+1/21 + \ln k \le \sum_{j=1}^k 1/j + 1/21+lnk≤∑j=1k​1/j+1/2;
  3. some input of EC(k)(k)(k) with F∗>0F^* > 0F∗>0 has a choosable subcover with mEC(M)≥(∑j=1k1/j)F∗m_{EC}(M) \ge \big(\sum_{j=1}^k 1/j\big) F^*mEC​(M)≥(∑j=1k​1/j)F∗.

Milestones (proof of Theorem 6, pp. 271–272)

  • the measure of the output is ∣T∣+OV(F1)|T| + \mathrm{OV}(F_1)∣T∣+OV(F1​);
  • if C2 may choose a set with Ratio(S′)≥y\mathrm{Ratio}(S') \ge yRatio(S′)≥y, then (y+1) ∣UNCOV∣≤F∗(y+1)\,|\mathrm{UNCOV}| \le F^*(y+1)∣UNCOV∣≤F∗;
  • with a=F∗/∣T∣a = F^*/|T|a=F∗/∣T∣ and x=∣T−UNCOV∣/∣T∣x = |T - \mathrm{UNCOV}|/|T|x=∣T−UNCOV∣/∣T∣, the next chosen set has Ratio(S′)≤a/(1−x)−1\mathrm{Ratio}(S') \le a/(1-x) - 1Ratio(S′)≤a/(1−x)−1;
  • on EC(k)(k)(k), OV(F1)≤∣T∣ (a[ln⁡(k)+1]−1)\mathrm{OV}(F_1) \le |T|\,(a[\ln(k) + 1] - 1)OV(F1​)≤∣T∣(a[ln(k)+1]−1);
  • the analytic inequality 1+ln⁡(k)≤∑j=1k1/j+1/21 + \ln(k) \le \sum_{j=1}^k 1/j + 1/21+ln(k)≤∑j=1k​1/j+1/2;
  • the lower-bound input (Fig. 1 of the paper with every set of F1F_1F1​ filled out to exactly kkk points) on which C2 may pay ∑j=1k1/j\sum_{j=1}^k 1/j∑j=1k​1/j times the optimum.

Significance

The result. The measure ∑∣S∣\sum|S|∑∣S∣ penalizes overlap, and the paper notes (without proof, p. 270) that an algorithm returning an optimal cover for the cardinality measure can be a factor kkk from optimal for this one. Theorem 6 shows that the ratio rule C2 is within 1+ln⁡k1 + \ln k1+lnk of the least-overlap cover, and the lower bound shows that no analysis of C2 can beat ∑j=1k1/j\sum_{j=1}^k 1/j∑j=1k​1/j. The two bounds differ by less than 1/21/21/2 for every kkk. The theorem was an early instance of a logarithmic guarantee for a weighted covering problem, where each set's cost is its size.

Formalizing it. The theorem has been proved since 1974; no machine-checked proof is known to exist. The mission produces a formal model of the EC problem and of C2 as a nondeterministic process, the overlap identity, and the discrete form of the paper's area-under-a-curve estimate. The last is the part the paper argues informally, through a step function and an integral.

Difficulty

The cardinality argument for C1 counts the sets chosen; here the sets have different sizes, so it does not apply. The overlap C2 pays per newly covered point is not bounded by a constant: early choices can be free and late ones cost up to k−1k-1k−1 per point, and the bound on the cumulative overlap must hold against the whole run, for every sequence of tie-breaks.

In a formal proof the integral must be replaced by a sum. Covered points arrive in blocks (one block per chosen set), the charge is constant on a block but the bound depends on the covered fraction at the start of the block, and the sum has to be compared with a logarithm. The terms aln⁡aa\ln aalna and a/ka/ka/k that the paper drops using 1≤a≤k1 \le a \le k1≤a≤k must be controlled as well, and the relation 1≤a≤k1 \le a \le k1≤a≤k must itself be proved from optimality. The lower bound needs an explicit run of C2 through ties on an input with k⋅k!k \cdot k!k⋅k! points, checking at every stage that the intended set is a ratio minimizer.

Formalization scope

  • Inputs. A family is p : ℕ with S : Fin p → Finset α (0-based, repetitions allowed; a repeated set counts twice in the measure if both copies are chosen, which C2 never does). Subcovers and SUB, LEFT are index sets. F∗F^*F∗ is a minimum over the finite, nonempty set of subcovers (Finset.inf'), never a junk value.
  • Algorithm. C2 is a step relation on states (SUB, LEFT, UNCOV). The choice at Step 3 is existential over all minimizers, so every result quantifies over every choosable output (the paper's WORST). Ratio(S)\mathrm{Ratio}(S)Ratio(S) is +∞+\infty+∞ when S∩UNCOV=∅S \cap \mathrm{UNCOV} = \emptysetS∩UNCOV=∅; the formal rule requires the chosen set to meet UNCOV and compares ratios by cross-multiplication, with no division.
  • Overlap. The cumulative overlap depends on the run, not on the output alone, so it is carried by an inductive run relation RunOV.
  • No problem size. The paper's R[A,P](n)R[A, P](n)R[A,P](n) maximizes over inputs of size at most nnn in an unspecified encoding. Upper bounds are stated for every input and every choosable output; the lower bound exhibits one input and one choosable output. Given monotonicity of RRR in nnn, these are equivalent to the paper's claims. Ratios are stated multiplicatively, so F∗=0F^* = 0F∗=0 does not create a vacuous bound.
  • Numbers. Measures are natural numbers cast to R\mathbb RR; ln⁡\lnln is Real.log; ∑j=1k1/j\sum_{j=1}^k 1/j∑j=1k​1/j is Mathlib's harmonic k.
  • Ruled out. A deterministic tie-break would prove a weaker upper bound and could not realize the lower-bound run, and a ratio with x/0=0x/0 = 0x/0=0 would make disjoint-from-UNCOV sets the most attractive choice. The formalization uses neither.

The overlap identity and the discrete integral comparison are reusable for any greedy covering analysis that charges cost per newly covered point. Contributions welcome include proofs of the milestones, the invariants of the C2 run relation (SUB and LEFT partition the indices; UNCOV =T−⋃= T - \bigcup=T−⋃ SUB), and the explicit lower-bound run.

Selected references

  • D. S. Johnson, Approximation algorithms for combinatorial problems, J. Comput. System Sci. 9 (1974), 256–278. https://doi.org/10.1016/S0022-0000(74)80044-9
  • R. M. Karp, Reducibility among combinatorial problems, in Complexity of Computer Computations, Plenum, 1972, 85–103. https://doi.org/10.1007/978-1-4684-2001-2_9
  • V. Chvátal, A greedy heuristic for the set-covering problem, Math. Oper. Res. 4 (1979), 233–235. https://doi.org/10.1287/moor.4.3.233
10 thms3 active usersReviewed
🏆Completed
CombinatoricsOptimization·Captain: mikedeng1

Scenario Reduction Algorithms in Stochastic Programming III: The Minimal Reduction Distance of a Regular Ternary Scenario TreeResearch Paper

Motivation

Multistage stochastic programs are solved on a finite scenario tree: a discrete probability distribution whose support points are paths of a random process. Realistic trees have far too many scenarios for the resulting optimization problem, so practitioners reduce the tree, keeping nnn of its NNN scenarios and redistributing the probability of the deleted ones. The reduction should keep the reduced distribution as close as possible to the original one in a probability metric that controls the optimal value of the stochastic program (Dupačová, Gröwe-Kuska, Römisch, Math. Program. 95 (2003)).

Choosing the best nnn scenarios is a set-covering problem and NP-hard, and the algorithms used in practice (backward reduction, fast forward selection) are heuristics without error guarantees. Heitsch and Römisch (2003) therefore derived test instances with an exactly known optimum: regular binary and ternary scenario trees, for which the minimal reduction distance has a closed form once nnn is not too small. This mission formalizes the ternary case, Proposition 3.2 of that paper. The binary case (Proposition 3.1) is a separate mission of the same series.

Setting

Fix a depth K∈NK \in \mathbb{N}K∈N and branch widths δ1,…,δK≥0\delta^1, \dots, \delta^K \ge 0δ1,…,δK≥0, with δ0=0\delta^0 = 0δ0=0. A regular ternary scenario tree has N=3KN = 3^KN=3K scenarios, one for each index tuple (i1,…,iK)∈{1,2,3}K(i_1, \dots, i_K) \in \{1, 2, 3\}^K(i1​,…,iK​)∈{1,2,3}K, where iki_kik​ is the successor chosen at level kkk. Choosing successor iki_kik​ adds the increment δikk=(ik−2) δk∈{−δk,0,δk}\delta^k_{i_k} = (i_k - 2)\,\delta^k \in \{-\delta^k, 0, \delta^k\}δik​k​=(ik​−2)δk∈{−δk,0,δk}, and scenario iii is the vector ωi=(ωi0,…,ωiK)∈RK+1\omega_i = (\omega_i^0, \dots, \omega_i^K) \in \mathbb{R}^{K+1}ωi​=(ωi0​,…,ωiK​)∈RK+1 with

ωik=∑j=0kδijj,k=0,…,K(eq. (19)).\omega_i^k = \sum_{j=0}^{k} \delta^j_{i_j}, \qquad k = 0, \dots, K \quad \text{(eq. (19))}.ωik​=j=0∑k​δij​j​,k=0,…,K(eq. (19)).

All scenarios have probability pi=1/Np_i = 1/Npi​=1/N. The distance between scenarios is the maximum norm c(ωi,ωj)=∥ωi−ωj∥∞=max⁡0≤k≤K∣ωik−ωjk∣c(\omega_i, \omega_j) = \|\omega_i - \omega_j\|_\infty = \max_{0 \le k \le K} |\omega_i^k - \omega_j^k|c(ωi​,ωj​)=∥ωi​−ωj​∥∞​=max0≤k≤K​∣ωik​−ωjk​∣.

Deleting the scenarios of an index set JJJ and moving each deleted scenario's probability to a nearest kept scenario costs the reduction distance

DJ=∑i∈Jpimin⁡j∉J∥ωi−ωj∥∞(eq. (8)),D_J = \sum_{i \in J} p_i \min_{j \notin J} \|\omega_i - \omega_j\|_\infty \quad \text{(eq. (8))},DJ​=i∈J∑​pi​j∈/Jmin​∥ωi​−ωj​∥∞​(eq. (8)),

which by Theorem 2.1 of the paper is the optimal transport-type distance between the original distribution and the best distribution supported on the kept scenarios. The minimal reduction distance to nnn scenarios is Dnmin=min⁡{DJ:#J=N−n}D^{min}_n = \min\{D_J : \#J = N - n\}Dnmin​=min{DJ​:#J=N−n}.

Formalization targets

Goal: Proposition 3.2 (7/9-solution)

Let K≥3K \ge 3K≥3 and let k0∈arg⁡min⁡1≤k≤Kδkk_0 \in \arg\min_{1 \le k \le K} \delta^kk0​∈argmin1≤k≤K​δk with k0≤K−2k_0 \le K - 2k0​≤K−2 and max⁡{δk0+1,δk0+2}≤2δk0\max\{\delta^{k_0+1}, \delta^{k_0+2}\} \le 2\delta^{k_0}max{δk0​+1,δk0​+2}≤2δk0​. Then any two distinct scenarios are at distance at least δk0\delta^{k_0}δk0​; there is a set of 79N\tfrac79 N97​N scenarios each paired with a scenario outside it at distance exactly δk0\delta^{k_0}δk0​; and for each n∈Nn \in \mathbb{N}n∈N with 29N≤n<N\tfrac29 N \le n < N92​N≤n<N,

Dnmin=min⁡{DJ:#J=N−n}=N−nN δk0(eq. (21)),D^{min}_n = \min\{D_J : \#J = N - n\} = \frac{N - n}{N}\,\delta^{k_0} \quad \text{(eq. (21))},Dnmin​=min{DJ​:#J=N−n}=NN−n​δk0​(eq. (21)),

with the minimum attained.

Milestones

  1. Distinct scenarios satisfy ∥ωi−ωj∥∞≥δk0\|\omega_i - \omega_j\|_\infty \ge \delta^{k_0}∥ωi​−ωj​∥∞​≥δk0​.
  2. Every JJJ with #J=N−n\#J = N - n#J=N−n has DJ≥N−nNδk0D_J \ge \frac{N-n}{N}\delta^{k_0}DJ​≥NN−n​δk0​.
  3. The index set I∗∗I_{**}I∗∗​ of the proof has #I∗∗=29N\#I_{**} = \tfrac29 N#I∗∗​=92​N, and its complement J∗∗J_{**}J∗∗​ has 79N\tfrac79 N97​N elements.
  4. Every j∈J∗∗j \in J_{**}j∈J∗∗​ has a partner i∈I∗∗i \in I_{**}i∈I∗∗​ with ∥ωi−ωj∥∞=δk0\|\omega_i - \omega_j\|_\infty = \delta^{k_0}∥ωi​−ωj​∥∞​=δk0​.
  5. Example 4.2: for K=6K = 6K=6 and (δ1,…,δ6)=(0.7,0.9,1.2,1.5,2.6,3.3)(\delta^1, \dots, \delta^6) = (0.7, 0.9, 1.2, 1.5, 2.6, 3.3)(δ1,…,δ6)=(0.7,0.9,1.2,1.5,2.6,3.3), Dnmin=0.7 N−nND^{min}_n = 0.7\,\frac{N-n}{N}Dnmin​=0.7NN−n​ for 162≤n<729162 \le n < 729162≤n<729.

Significance

The result gives an exact optimal value for an NP-hard reduction problem on an infinite family of instances. Heitsch and Römisch use it in their numerical section to measure how far the heuristics' reduced trees are from optimal (Examples 4.1 and 4.2 are the binary and ternary test trees of that study). A closed form of this kind is also the only way to certify that a heuristic is exactly optimal on some instances rather than only competitive with other heuristics.

The proposition is proved in the paper, but the published proof of its central step is one sentence: "Similarly as in Proposition 3.1 it can be shown that there exists an index i∈I∗∗i \in I_{**}i∈I∗∗​ for each j∈J∗∗j \in J_{**}j∈J∗∗​ …". A formal proof supplies that case analysis, which is absent from the literature. The formalization also settles two points the printed statement leaves loose (see Formalization scope): the count of pairs at distance δk0\delta^{k_0}δk0​, and the definition of I∗∗I_{**}I∗∗​ when some widths vanish. No machine-checked version of this result or of the reduction distance DJD_JDJ​ is known to exist.

Difficulty

The lower bound is routine: two distinct scenarios first differ at some level lll, where their coordinates differ by δl\delta^lδl or 2δl2\delta^l2δl. The substance is attainment: one must exhibit, for every n≥29Nn \ge \tfrac29 Nn≥92​N, a kept set of size nnn whose every deleted scenario lies at distance exactly δk0\delta^{k_0}δk0​ from some kept one. The obvious candidate, keeping the scenarios that take the middle branch at level k0k_0k0​, has every other scenario at distance exactly δk0\delta^{k_0}δk0​ from a kept one, but it keeps 13N\tfrac13 N31​N scenarios and so covers only n≥13Nn \ge \tfrac13 Nn≥31​N. Going down to 29N\tfrac29 N92​N kept scenarios forces a deleted scenario and its partner to differ at more than one level, and since coordinates are running sums the differences at levels k0+1k_0+1k0​+1 and k0+2k_0+2k0​+2 accumulate on top of the one at level k0k_0k0​. The paper's proof of this step is not written out, and it depends on the widths of the two levels below k0k_0k0​: the hypothesis max⁡{δk0+1,δk0+2}≤2δk0\max\{\delta^{k_0+1}, \delta^{k_0+2}\} \le 2\delta^{k_0}max{δk0​+1,δk0​+2}≤2δk0​ is essential, and the result is false without it (for K=3K = 3K=3 and (δ1,δ2,δ3)=(1,3,3)(\delta^1, \delta^2, \delta^3) = (1, 3, 3)(δ1,δ2,δ3)=(1,3,3) one has D6min=31/27D^{min}_6 = 31/27D6min​=31/27, not 7/97/97/9).

Formalization scope

  • A scenario is an index tuple σ:Fin K→Fin 3\sigma : \mathrm{Fin}\,K \to \mathrm{Fin}\,3σ:FinK→Fin3; σ(r)\sigma(r)σ(r) is the successor at paper level r+1r + 1r+1, with Fin 3\mathrm{Fin}\,3Fin3 values 0,1,20, 1, 20,1,2 standing for the paper's i=1,2,3i = 1, 2, 3i=1,2,3. The widths are δ:N→R\delta : \mathbb{N} \to \mathbb{R}δ:N→R, of which only δ(1),…,δ(K)\delta(1), \dots, \delta(K)δ(1),…,δ(K) are used; the standing assumption δk∈R+\delta^k \in \mathbb{R}_+δk∈R+​ (p. 196) is the hypothesis δ(k)≥0\delta(k) \ge 0δ(k)≥0 for 1≤k≤K1 \le k \le K1≤k≤K. Scenarios live in Fin(K+1)→R\mathrm{Fin}(K+1) \to \mathbb{R}Fin(K+1)→R, whose Mathlib norm is the maximum norm. Probabilities are uniform, 1/3K1/3^K1/3K.
  • DJD_JDJ​ is defined for a general finite index set, probabilities and cost, with the inner minimum a Finset.inf' over the complement of JJJ; the complement must be nonempty, so no default value arises. DnminD^{min}_nDnmin​ is stated as IsLeast of the set of all values DJD_JDJ​ with #J=N−n\#J = N - n#J=N−n: the goal asserts both the lower bound for every JJJ and attainment by some JJJ. A formalization that exhibits a single JJJ with DJ=N−nNδk0D_J = \frac{N-n}{N}\delta^{k_0}DJ​=NN−n​δk0​, or states an infimum without attainment, or drops any hypothesis on k0k_0k0​, is a different (and in the last case false) statement.
  • 29N≤n\tfrac29 N \le n92​N≤n is written 2⋅3K≤9n2 \cdot 3^K \le 9n2⋅3K≤9n, and 79N\tfrac79 N97​N as 7⋅3K−27 \cdot 3^{K-2}7⋅3K−2.
  • Pairs. As printed, "there are 79N\tfrac79 N97​N distinct pairs of scenarios such that the distance between the members of each pair is exactly δk0\delta^{k_0}δk0​" is false as an exact count (for K=3K = 3K=3 and δ=(1,1,1)\delta = (1,1,1)δ=(1,1,1) there are 130 such pairs, not 21). The goal states what the proof constructs: a set J∗∗J_{**}J∗∗​ of 79N\tfrac79 N97​N scenarios, each paired with a scenario outside J∗∗J_{**}J∗∗​ at distance exactly δk0\delta^{k_0}δk0​.
  • I∗∗I_{**}I∗∗​. The paper defines I∗∗I_{**}I∗∗​ by testing whether the increments δikk\delta^k_{i_k}δik​k​ vanish. When one of δk0,δk0+1,δk0+2\delta^{k_0}, \delta^{k_0+1}, \delta^{k_0+2}δk0​,δk0​+1,δk0​+2 is 000 this no longer identifies the middle branch and the count 29N\tfrac29 N92​N fails, although the proposition remains true. The formalization defines I∗∗I_{**}I∗∗​ by branch indices (middle branch versus outer branches), which agrees with the paper whenever these three widths are positive. No positivity hypothesis is added to the goal.
  • Welcome contributions: a reusable library for regular scenario trees (first differing level, distance of paths), and the case analysis of milestone 4. The binary mission of this series needs the same lower-bound argument with the constant 2δk02\delta^{k_0}2δk0​.

Selected references

  • H. Heitsch, W. Römisch, Scenario Reduction Algorithms in Stochastic Programming, Computational Optimization and Applications 24 (2003), 187–206. https://doi.org/10.1023/A:1021805924152
  • J. Dupačová, N. Gröwe-Kuska, W. Römisch, Scenario reduction in stochastic programming: An approach using probability metrics, Mathematical Programming 95 (2003), 493–511. https://doi.org/10.1007/s10107-002-0331-0
9 thms3 active usersReviewed
🏆Completed
CombinatoricsOptimization·Captain: mikedeng1

Scenario Reduction Algorithms in Stochastic Programming II: The Minimal Reduction Distance of a Regular Binary Scenario TreeResearch Paper

Why exact reduction distances matter

Multistage stochastic programs are solved on a finite scenario tree, a discrete probability measure whose atoms are paths of a stochastic process. The size of the deterministic equivalent grows with the number of scenarios, so practitioners replace the original measure P=∑i=1NpiδωiP=\sum_{i=1}^N p_i\delta_{\omega_i}P=∑i=1N​pi​δωi​​ by a measure supported on n<Nn<Nn<N of its scenarios. Stability theory for stochastic programs (Dupačová, Gröwe-Kuska and Römisch, Math. Program. 95 (2003), doi:10.1007/s10107-002-0331-0) bounds the change of the optimal value by a probability metric between the two measures, which leads to the optimal scenario reduction problem: choose which N−nN-nN−n scenarios to delete so that this distance is smallest.

That problem is a set-covering problem and is NP-hard, and the algorithms that Heitsch and Römisch study in the same paper (backward reduction, fast forward selection) are heuristics without error guarantees. To test them one needs original measures whose optimal reduction distance is known exactly. Section 3 of Heitsch and Römisch, Scenario Reduction Algorithms in Stochastic Programming, Comput. Optim. Appl. 24 (2003) (doi:10.1023/A:1021805924152) supplies such instances: regular binary and ternary scenario trees, for which the minimal distance to any reduced tree with at least a fixed fraction of the scenarios is an explicit formula. This mission formalizes the binary case, Proposition 3.1.

Setting

Fix a horizon K∈NK\in\mathbb NK∈N and level parameters δ1,…,δK≥0\delta^1,\dots,\delta^K\ge0δ1,…,δK≥0, with δ0=0\delta^0=0δ0=0. A regular binary scenario tree has N=2KN=2^KN=2K scenarios. A scenario is determined by a branch ik∈{1,2}i_k\in\{1,2\}ik​∈{1,2} at every level k=1,…,Kk=1,\dots,Kk=1,…,K, and is the vector ωi=(ωi0,…,ωiK)∈RK+1\omega_i=(\omega_i^0,\dots,\omega_i^K)\in\mathbb R^{K+1}ωi​=(ωi0​,…,ωiK​)∈RK+1 with

ωik=∑j=0kδijj,δijj=(2ij−3) δj∈{−δj,+δj}.(19)\omega_i^k=\sum_{j=0}^k\delta^j_{i_j},\qquad \delta^j_{i_j}=(2i_j-3)\,\delta^j\in\{-\delta^j,+\delta^j\}.\tag{19}ωik​=j=0∑k​δij​j​,δij​j​=(2ij​−3)δj∈{−δj,+δj}.(19)

All scenarios start at the root ωi0=0\omega_i^0=0ωi0​=0 and carry probability pi=1/Np_i=1/Npi​=1/N. Scenarios are compared in the maximum norm ∥ω−ω~∥∞=max⁡k=0,…,K∣ωk−ω~k∣\|\omega-\tilde\omega\|_\infty=\max_{k=0,\dots,K}|\omega^k-\tilde\omega^k|∥ω−ω~∥∞​=maxk=0,…,K​∣ωk−ω~k∣.

Deleting the scenarios with indices in J⊂{1,…,N}J\subset\{1,\dots,N\}J⊂{1,…,N} and moving each deleted scenario's probability to a nearest kept scenario costs the reduction cost

DJ=∑i∈Jpimin⁡j∉J∥ωi−ωj∥∞,(8)D_J=\sum_{i\in J}p_i\min_{j\notin J}\|\omega_i-\omega_j\|_\infty,\tag{8}DJ​=i∈J∑​pi​j∈/Jmin​∥ωi​−ωj​∥∞​,(8)

which by Theorem 2.1 of the paper is the minimal Kantorovich-type distance between PPP and a measure supported on the kept scenarios. The minimal reduction distance for nnn kept scenarios is Dnmin=min⁡{DJ:#J=N−n}D^{min}_n=\min\{D_J:\#J=N-n\}Dnmin​=min{DJ​:#J=N−n}.

In the Lean development, scenarios are indexed by σ : Fin K → Fin 2 (Fin-index rrr is tree level r+1r+1r+1, value 000 is the branch −δ-\delta−δ, value 111 is +δ+\delta+δ), lev σ k is the branch at level kkk, scenario δ σ : Fin (K+1) → ℝ is ωσ\omega_\sigmaωσ​, and redCost δ J hJ is DJD_JDJ​.

Formalization targets

Goal: Proposition 3.1 (3/4-solution)

Let K≥3K\ge3K≥3, k0∈arg⁡min⁡1≤k≤Kδkk_0\in\arg\min_{1\le k\le K}\delta^kk0​∈argmin1≤k≤K​δk, k0≤K−2k_0\le K-2k0​≤K−2 and max⁡{δk0+1,δk0+2}≤2δk0\max\{\delta^{k_0+1},\delta^{k_0+2}\}\le2\delta^{k_0}max{δk0​+1,δk0​+2}≤2δk0​. Then any two distinct scenarios are at distance at least 2δk02\delta^{k_0}2δk0​; there is a set J∗J_*J∗​ of 34N\frac34N43​N scenarios each of which has a partner outside J∗J_*J∗​ at distance exactly 2δk02\delta^{k_0}2δk0​; and for every n∈Nn\in\mathbb Nn∈N with N4≤n<N\frac N4\le n<N4N​≤n<N

Dnmin=min⁡{DJ:#J=N−n}=N−nN 2δk0.(20)D^{min}_n=\min\{D_J:\#J=N-n\}=\frac{N-n}{N}\,2\delta^{k_0}.\tag{20}Dnmin​=min{DJ​:#J=N−n}=NN−n​2δk0​.(20)

Milestones

  1. Two scenarios that first differ at level lll are at distance ≥2δl≥2δk0\ge2\delta^l\ge2\delta^{k_0}≥2δl≥2δk0​.
  2. DJ≥N−nN2δk0D_J\ge\frac{N-n}{N}2\delta^{k_0}DJ​≥NN−n​2δk0​ for every JJJ with #J=N−n\#J=N-n#J=N−n.
  3. The index set I∗I_*I∗​ (branch at k0k_0k0​ opposite to the common branch at k0+1,k0+2k_0+1,k_0+2k0​+1,k0​+2) has #I∗=N/4\#I_*=N/4#I∗​=N/4, so #J∗=34N\#J_*=\frac34N#J∗​=43​N.
  4. Every j∈J∗j\in J_*j∈J∗​ has a partner i∈I∗i\in I_*i∈I∗​ with ∥ωi−ωj∥∞=2δk0\|\omega_i-\omega_j\|_\infty=2\delta^{k_0}∥ωi​−ωj​∥∞​=2δk0​.
  5. Example 4.1: for K=10K=10K=10 and the paper's parameters, Proposition 3.1 applies with k0=1k_0=1k0​=1 and Dnmin=N−nND^{min}_n=\frac{N-n}{N}Dnmin​=NN−n​ for 256≤n<1024256\le n<1024256≤n<1024.

Significance

The result. Proposition 3.1 gives the exact optimum of an NP-hard combinatorial problem on an explicit, parametrized family of instances of every size N=2KN=2^KN=2K. Section 4 of the paper uses it (Example 4.1, N=1024N=1024N=1024) as ground truth for the relative accuracy of backward reduction and fast forward selection. Without it, the quality of a heuristic reduction on a large tree could only be compared with other heuristics or with lower bounds.

The formalization. The proposition is proved in the paper; to our knowledge no machine-checked version exists. A formal proof certifies the benchmark values, and the statement also corrects the printed text in two places. First, "there are 34N\frac34N43​N distinct pairs … at distance exactly 2δk02\delta^{k_0}2δk0​" is false as an exact count (for K=3K=3K=3, δ=(1,1,1)\delta=(1,1,1)δ=(1,1,1) there are twenty such pairs, not six), so the goal states "at least", in the form the proof exhibits. Second, the paper's sign-based definition of I∗I_*I∗​ degenerates when some δk=0\delta^k=0δk=0, although the proposition still holds; the mission defines I∗I_*I∗​ by branch indices.

Difficulty

The lower bound is a direct computation. The content is the matching upper bound: a set JJJ of the prescribed size for which every deleted scenario has a kept scenario at the minimal possible distance. A natural first attempt pairs scenarios that differ only at level k0k_0k0​. That handles only half of the scenarios with a single partner each, and it cannot reach 34N\frac34N43​N deleted scenarios. Once scenarios differ at more than one level, their maximum-norm distance is a maximum of several partial sums, and keeping all of them at most 2δk02\delta^{k_0}2δk0​ is exactly where the hypothesis max⁡{δk0+1,δk0+2}≤2δk0\max\{\delta^{k_0+1},\delta^{k_0+2}\}\le2\delta^{k_0}max{δk0​+1,δk0​+2}≤2δk0​ enters. Without it, eq. (20) fails: for K=3K=3K=3, δ=(1,3,3)\delta=(1,3,3)δ=(1,3,3) and n=2n=2n=2 the true minimum is 52\frac5225​, not 32\frac3223​. The passage from n=N/4n=N/4n=N/4 to general n≥N/4n\ge N/4n≥N/4 also needs care: the deleted set must shrink while each remaining deleted scenario keeps its partner among the kept ones.

Formalization scope

  • The index type is Fin K → Fin 2, which has exactly 2K2^K2K elements. The paper's (K+1)(K+1)(K+1)-tuple has a level-0 entry with no choice, so it is dropped, and the vector ω\omegaω keeps its K+1K+1K+1 coordinates with ω0=0\omega^0=0ω0=0.
  • The parameters are δ : ℕ → ℝ, with δk≥0\delta^k\ge0δk≥0 and the arg min stated for k=1,…,Kk=1,\dots,Kk=1,…,K only. δk>0\delta^k>0δk>0 is not assumed, because the paper allows δk∈R+\delta^k\in\mathbb R_+δk∈R+​ and the proposition holds with zeros.
  • The cost is Mathlib's norm on Fin (K+1) → ℝ, which is the maximum norm, and pi=1/2Kp_i=1/2^Kpi​=1/2K is written out.
  • DJD_JDJ​ requires a nonempty set of kept scenarios, and its inner minimum is a finite Finset.inf'.
  • DnminD^{min}_nDnmin​ is stated with IsLeast over the set of attained values DJD_JDJ​, #J=N−n\#J=N-n#J=N−n, so both the lower bound and attainment are part of the goal. A proof of DJ≤N−nN2δk0D_{J}\le\frac{N-n}{N}2\delta^{k_0}DJ​≤NN−n​2δk0​ for a single exhibited JJJ, or a real infimum without attainment, does not prove the goal.
  • "N4≤n\frac N4\le n4N​≤n" is written 2K≤4n2^K\le4n2K≤4n, and 34N\frac34N43​N is 3⋅2K−23\cdot2^{K-2}3⋅2K−2.
  • The pairs claim is "at least 34N\frac34N43​N pairs", expressed as a set J∗J_*J∗​ of that size with a partner outside J∗J_*J∗​ for every member.
  • All hypotheses on k0k_0k0​ appear in the goal; dropping any of them makes (20) false.

Useful infrastructure: sup-norm lemmas for Fin n → ℝ (pi_norm_le_iff_of_nonneg, norm_le_pi_norm), Finset.inf' lemmas, and counting functions Fin K → Fin 2 with prescribed values (Fintype.card_fun, Fintype.card_pi). The tree and reduction-cost definitions are shared in spirit with the ternary-tree mission of this series (Proposition 3.2), and a proof whose structure transfers to d=3d=3d=3 is welcome. Contributions of proofs of the milestones individually, in any order, are welcome.

Selected references

  • H. Heitsch, W. Römisch, Scenario Reduction Algorithms in Stochastic Programming, Computational Optimization and Applications 24 (2003), 187–206. doi:10.1023/A:1021805924152
  • J. Dupačová, N. Gröwe-Kuska, W. Römisch, Scenario reduction in stochastic programming: An approach using probability metrics, Mathematical Programming 95 (2003), 493–511. doi:10.1007/s10107-002-0331-0
9 thms3 active usersReviewed
🏆Completed
Optimal TransportOptimization·Captain: mikedeng1

Scenario Reduction Algorithms in Stochastic Programming I: Fast Forward Selection Realizes the Forward Selection PrincipleResearch Paper

Why reduce scenarios

Multistage and two-stage stochastic programs are solved numerically on a discrete probability distribution: a finite set of scenarios ω1,…,ωN\omega_1,\dots,\omega_Nω1​,…,ωN​ with probabilities p1,…,pNp_1,\dots,p_Np1​,…,pN​. The size of the resulting optimization problem grows with NNN, and scenario sets produced by sampling or by historical data are often far too large to be solved directly. Scenario reduction replaces the original distribution by one supported on a small subset of the scenarios, chosen so that the optimal value and solutions of the stochastic program change as little as possible.

Stability theory for stochastic programs (Rachev and Römisch, 2002) shows that this change is controlled by a probability distance of Fortet–Mourier type, which for discrete measures is bounded by the value of a transportation problem. Dupačová, Gröwe-Kuska and Römisch (2003) turned this into a combinatorial problem and proposed greedy backward and forward heuristics. Heitsch and Römisch (2003) gave faster versions of both heuristics; the forward version, fast forward selection, is the subject of this mission. Implementations of these reduction heuristics are distributed with the GAMS modelling system (SCENRED) and are used in energy and finance applications of stochastic programming.

Setting

Let EEE be a finite-dimensional real vector space with a norm ∥⋅∥\|\cdot\|∥⋅∥, let ω0∈E\omega_0\in Eω0​∈E, and let h:[0,∞)→[0,∞)h:[0,\infty)\to[0,\infty)h:[0,∞)→[0,∞) be continuous and nondecreasing with h(0)=0h(0)=0h(0)=0. The cost between two points of EEE is

c(ω,ω~)=max⁡{1, h(∥ω−ω0∥), h(∥ω~−ω0∥)} ∥ω−ω~∥.c(\omega,\tilde\omega)=\max\bigl\{1,\,h(\|\omega-\omega_0\|),\,h(\|\tilde\omega-\omega_0\|)\bigr\}\,\|\omega-\tilde\omega\| .c(ω,ω~)=max{1,h(∥ω−ω0​∥),h(∥ω~−ω0​∥)}∥ω−ω~∥.

It is nonnegative, symmetric, and zero on the diagonal.

The original distribution is P=∑i=1NpiδωiP=\sum_{i=1}^N p_i\delta_{\omega_i}P=∑i=1N​pi​δωi​​ with pi>0p_i>0pi​>0 and ∑ipi=1\sum_i p_i=1∑i​pi​=1. Deleting the scenarios in a set J⊂{1,…,N}J\subset\{1,\dots,N\}J⊂{1,…,N} and assigning new weights qj≥0q_j\ge 0qj​≥0, ∑j∉Jqj=1\sum_{j\notin J}q_j=1∑j∈/J​qj​=1, to the kept ones gives Q=∑j∉JqjδωjQ=\sum_{j\notin J}q_j\delta_{\omega_j}Q=∑j∈/J​qj​δωj​​. The distance D(J;q)D(J;q)D(J;q) between PPP and QQQ is the optimal value of the transportation problem

D(J;q)=min⁡{∑i=1N∑j∉Jc(ωi,ωj)ηij: ηij≥0, ∑iηij=qj, ∑j∉Jηij=pi}.D(J;q)=\min\Bigl\{\sum_{i=1}^N\sum_{j\notin J}c(\omega_i,\omega_j)\eta_{ij}:\ \eta_{ij}\ge 0,\ \sum_{i}\eta_{ij}=q_j,\ \sum_{j\notin J}\eta_{ij}=p_i\Bigr\}.D(J;q)=min{i=1∑N​j∈/J∑​c(ωi​,ωj​)ηij​: ηij​≥0, i∑​ηij​=qj​, j∈/J∑​ηij​=pi​}.

The reduction cost of deleting JJJ is

DJ=∑i∈Jpimin⁡j∉Jc(ωi,ωj),D_J=\sum_{i\in J}p_i\min_{j\notin J}c(\omega_i,\omega_j),DJ​=i∈J∑​pi​j∈/Jmin​c(ωi​,ωj​),

and the optimal reduction problem (8) minimizes DJD_JDJ​ over all JJJ with #J=N−n\#J=N-n#J=N−n, where nnn is the number of scenarios to keep.

Forward selection builds the kept set greedily. With J[0]={1,…,N}J^{[0]}=\{1,\dots,N\}J[0]={1,…,N} and J[i]={1,…,N}∖{u1,…,ui}J^{[i]}=\{1,\dots,N\}\setminus\{u_1,\dots,u_i\}J[i]={1,…,N}∖{u1​,…,ui​}, it chooses

ui∈arg⁡min⁡u∈J[i−1]DJ[i−1]∖{u},i=1,…,n.(16)u_i\in\arg\min_{u\in J^{[i-1]}}D_{J^{[i-1]}\setminus\{u\}},\qquad i=1,\dots,n. \tag{16}ui​∈argu∈J[i−1]min​DJ[i−1]∖{u}​,i=1,…,n.(16)

Fast forward selection (Algorithm 2.4) computes the same choices through an updated cost matrix: cku[1]=c(ωk,ωu)c^{[1]}_{ku}=c(\omega_k,\omega_u)cku[1]​=c(ωk​,ωu​), cku[i]=min⁡{cku[i−1],ckui−1[i−1]}c^{[i]}_{ku}=\min\{c^{[i-1]}_{ku},c^{[i-1]}_{ku_{i-1}}\}cku[i]​=min{cku[i−1]​,ckui−1​[i−1]​}, zu[i]=∑k∈J[i−1]∖{u}pkcku[i]z^{[i]}_u=\sum_{k\in J^{[i-1]}\setminus\{u\}}p_kc^{[i]}_{ku}zu[i]​=∑k∈J[i−1]∖{u}​pk​cku[i]​, and ui∈arg⁡min⁡u∈J[i−1]zu[i]u_i\in\arg\min_{u\in J^{[i-1]}}z^{[i]}_uui​∈argminu∈J[i−1]​zu[i]​.

Formalization targets

Goal: Theorem 2.5

For 1≤n≤N1\le n\le N1≤n≤N and every run u1,…,unu_1,\dots,u_nu1​,…,un​ of Algorithm 2.4, with any tie-breaking in the arg min,

ui satisfies (16)andzui[i]=DJ[i](i=1,…,n).u_i\ \text{satisfies (16)}\quad\text{and}\quad z^{[i]}_{u_i}=D_{J^{[i]}}\qquad(i=1,\dots,n).ui​ satisfies (16)andzui​[i]​=DJ[i]​(i=1,…,n).

Milestones

  1. Theorem 2.1 (redistribution). For JJJ with at least one kept scenario, DJ=min⁡qD(J;q)D_J=\min_q D(J;q)DJ​=minq​D(J;q), and the minimum is attained at qˉj=pj+∑i∈J, j(i)=jpi\bar q_j=p_j+\sum_{i\in J,\,j(i)=j}p_iqˉ​j​=pj​+∑i∈J,j(i)=j​pi​ for every choice of nearest kept scenarios j(i)j(i)j(i).
  2. Eq. (10). D{1,…,N}∖{u}=∑i=1Npic(ωi,ωu)D_{\{1,\dots,N\}\setminus\{u\}}=\sum_{i=1}^Np_ic(\omega_i,\omega_u)D{1,…,N}∖{u}​=∑i=1N​pi​c(ωi​,ωu​), so (8) with #J=N−1\#J=N-1#J=N−1 is problem (10).
  3. Eq. (12). The sum lblblb of the N−nN-nN−n smallest single-deletion costs plmin⁡j≠lc(ωl,ωj)p_l\min_{j\neq l}c(\omega_l,\omega_j)pl​minj=l​c(ωl​,ωj​), taken in the greedy order (11), is at most DJD_JDJ​ for every JJJ with #J=N−n\#J=N-n#J=N−n.
  4. Optimality condition (p. 191). If each lil_ili​ has a nearest other scenario outside {l1,…,lN−n}∖{li}\{l_1,\dots,l_{N-n}\}\setminus\{l_i\}{l1​,…,lN−n​}∖{li​}, then {l1,…,lN−n}\{l_1,\dots,l_{N-n}\}{l1​,…,lN−n​} solves (8).
  5. Eq. (17), unrolled recursion. For any index sequence, cku[i]=min⁡j∉J[i−1]∖{u}c(ωk,ωj)c^{[i]}_{ku}=\min_{j\notin J^{[i-1]}\setminus\{u\}}c(\omega_k,\omega_j)cku[i]​=minj∈/J[i−1]∖{u}​c(ωk​,ωj​) for u∈J[i−1]u\in J^{[i-1]}u∈J[i−1].
  6. Eq. (17), conclusion. For any index sequence, zu[i]=DJ[i−1]∖{u}z^{[i]}_u=D_{J^{[i-1]}\setminus\{u\}}zu[i]​=DJ[i−1]∖{u}​ for u∈J[i−1]u\in J^{[i-1]}u∈J[i−1].

Significance

Theorem 2.5 certifies that the cheap update of Algorithm 2.4 (one pairwise minimum per matrix entry and step) produces exactly the greedy forward selection defined through the reduction costs, and that the running objective zui[i]z^{[i]}_{u_i}zui​[i]​ is the reduction cost of the scenarios deleted so far. Combined with Theorem 2.1, zui[i]z^{[i]}_{u_i}zui​[i]​ is the optimal transportation distance between PPP and the best measure on the kept scenarios, which is the quantity practitioners monitor to decide how many scenarios to keep. The lower bound (12) and the optimality condition give a posteriori quality certificates for any reduced set.

All results of this mission are proved in the paper or in the works it cites (Dupačová et al., 2003); none is open. To the best of a platform search, none has been machine-checked. The mission provides a verified specification of a widely deployed algorithm, a formal link between a combinatorial set-covering objective and a finite transportation problem, and definitions (reduction cost, transportation plans with a partially free target marginal, greedy runs with arbitrary tie-breaking) reusable by the regular-tree missions of this series and by later scenario-tree construction papers.

Difficulty

The mathematics is elementary; the difficulty is bookkeeping. The recursion for c[i]c^{[i]}c[i] refers to the previous step's column ui−1u_{i-1}ui−1​, which itself was updated, so unrolling it to a minimum over {u,u1,…,ui−1}\{u,u_1,\dots,u_{i-1}\}{u,u1​,…,ui−1​} is an induction on iii in which the index sets J[i]J^{[i]}J[i], the 1-based step counter and the complement structure all move together. The natural first attempt, identifying cku[i]c^{[i]}_{ku}cku[i]​ with the minimum over the complement of J[i]J^{[i]}J[i], is off by one step: the correct set is the complement of J[i−1]∖{u}J^{[i-1]}\setminus\{u\}J[i−1]∖{u}, which contains uuu itself. For Theorem 2.1 the lower bound requires using that c(ωi,ωi)=0c(\omega_i,\omega_i)=0c(ωi​,ωi​)=0 for kept scenarios and that every plan ships all of pip_ipi​ somewhere outside JJJ; the attainment part requires constructing the plan explicitly from the choice j(⋅)j(\cdot)j(⋅), including scenarios for which several kept scenarios are equally near.

Formalization scope

  • Scenarios are ω : Fin N → E with E a finite-dimensional real normed space; the paper's closed set Ω⊂Rs\Omega\subset\mathbb R^sΩ⊂Rs plays no role beyond containing the scenarios and is omitted. Scenarios need not be distinct.
  • hhh is a function ℝ → ℝ with the paper's assumptions imposed on [0,∞)[0,\infty)[0,∞) (IsGrowthFunction); every theorem carries them, together with pi>0p_i>0pi​>0 and ∑ipi=1\sum_ip_i=1∑i​pi​=1.
  • The functions f0f_0f0​, ggg and the stochastic program (1)–(2) that motivate ccc appear in no statement.
  • D(J;q)D(J;q)D(J;q) is the paper's finite transportation problem (p. 188), not the Kantorovich functional on measures. Weights qqq and plans η\etaη are indexed by all of {1,…,N}\{1,\dots,N\}{1,…,N} with entries at deleted indices fixed to 000.
  • DJD_JDJ​ requires a proof that the complement of JJJ is nonempty; minima are Finset.inf', never a real infimum with a default value.
  • Algorithm 2.4 is a relation on sequences u : ℕ → Fin N with 1-based steps. c[i]c^{[i]}c[i] is the printed recursion, extended to all indices; runs are any sequences satisfying the arg-min conditions, so every tie-breaking rule is covered.
  • The paper's standing restriction n<Nn<Nn<N is relaxed to n≤Nn\le Nn≤N in Theorem 2.5; the statement remains true at n=Nn=Nn=N.
  • A trivializing formalization is ruled out: defining c[i]c^{[i]}c[i] or z[i]z^{[i]}z[i] directly as the minimum over the selected set or as DJ[i−1]∖{u}D_{J^{[i-1]}\setminus\{u\}}DJ[i−1]∖{u}​ would make Theorem 2.5 hold by definition, and proving it for one fixed tie-breaking rule would prove less than the paper; neither is done.
  • Proofs of the milestones, alternative proofs of Theorem 2.1 via LP duality, and a verified executable implementation of Algorithm 2.4 are all welcome.

Selected references

  • H. Heitsch, W. Römisch, Scenario Reduction Algorithms in Stochastic Programming, Computational Optimization and Applications 24 (2003), 187–206. https://doi.org/10.1023/A:1021805924152
  • J. Dupačová, N. Gröwe-Kuska, W. Römisch, Scenario reduction in stochastic programming: an approach using probability metrics, Mathematical Programming 95 (2003), 493–511. https://doi.org/10.1007/s10107-002-0331-0
  • S. T. Rachev, W. Römisch, Quantitative stability in stochastic programming: the method of probability metrics, Mathematics of Operations Research 27 (2002), 792–818. https://doi.org/10.1287/moor.27.4.792.304
12 thms3 active usersReviewed
🏆Completed
Convex OptimizationNumerical AnalysisOptimization·Captain: mikedeng1

Golden Ratio Algorithms for Variational Inequalities I: The Golden Ratio Algorithm with a Fixed Step Converges to a Solution of a Monotone Variational InequalityResearch Paper

Motivation

A monotone variational inequality asks for a point at which a monotone operator and a convex function are in equilibrium. It unifies convex minimization (where FFF is a gradient), convex–concave saddle-point problems (where FFF is the skew gradient of a Lagrangian), Nash equilibria of monotone games, and complementarity problems in economics and traffic assignment. In operations research, first-order methods for such problems are the workhorse behind large-scale saddle-point formulations of linear and conic programs, where only one operator evaluation and one projection or proximal step per iteration are affordable.

The classical method for Lipschitz monotone operators is Korpelevich's extragradient method (1976) and its proximal variant, Tseng's forward–backward–forward method (2000); both need two evaluations of FFF per iteration. The reflected projected gradient method of Malitsky (SIAM J. Optim., 2015) uses one evaluation of FFF but evaluates it at 2zk−zk−12z^k-z^{k-1}2zk−zk−1, a point that may lie outside the domain of ggg. Malitsky's Golden Ratio Algorithm (GRAAL), introduced in Golden Ratio Algorithms for Variational Inequalities (preprint 2018; published in Mathematical Programming, doi:10.1007/s10107-019-01416-w), uses one evaluation of FFF, always at a feasible point, and one proximal step per iteration. Its fixed-step version, Theorem 1 of that paper, is the subject of this mission; the explicit, adaptive-step version (Theorem 2) is a separate mission of this series.

Setting

Let E\mathcal EE be a finite-dimensional real inner product space with inner product ⟨⋅,⋅⟩\langle\cdot,\cdot\rangle⟨⋅,⋅⟩ and norm ∥⋅∥=⟨⋅,⋅⟩\|\cdot\| = \sqrt{\langle\cdot,\cdot\rangle}∥⋅∥=⟨⋅,⋅⟩​. Let g:E→(−∞,+∞]g:\mathcal E\to(-\infty,+\infty]g:E→(−∞,+∞] and write dom⁡g={x:g(x)<+∞}\operatorname{dom} g = \{x : g(x)<+\infty\}domg={x:g(x)<+∞}. Let F:dom⁡g→EF:\operatorname{dom} g\to\mathcal EF:domg→E. The variational inequality is

find z∗∈Esuch that⟨F(z∗),z−z∗⟩+g(z)−g(z∗) ≥ 0∀z∈E.(1)\text{find } z^*\in\mathcal E \quad\text{such that}\quad \langle F(z^*), z-z^*\rangle + g(z)-g(z^*)\ \ge\ 0\qquad \forall z\in\mathcal E. \tag{1}find z∗∈Esuch that⟨F(z∗),z−z∗⟩+g(z)−g(z∗) ≥ 0∀z∈E.(1)

The standing assumptions are:

  • (C1) the solution set SSS of (1) is nonempty;
  • (C2) ggg is proper (never −∞-\infty−∞, finite somewhere), convex, and lower semicontinuous;
  • (C3) FFF is monotone: ⟨F(u)−F(v),u−v⟩≥0\langle F(u)-F(v),u-v\rangle\ge0⟨F(u)−F(v),u−v⟩≥0 for all u,v∈dom⁡gu,v\in\operatorname{dom} gu,v∈domg.

The proximal operator of ggg is prox⁡g(z)=argmin⁡x{g(x)+12∥x−z∥2}\operatorname{prox}_g(z) = \operatorname{argmin}_x\{g(x)+\tfrac12\|x-z\|^2\}proxg​(z)=argminx​{g(x)+21​∥x−z∥2}. Let φ=5+12\varphi = \frac{\sqrt5+1}{2}φ=25​+1​ be the golden ratio, so that φ2=1+φ\varphi^2 = 1+\varphiφ2=1+φ. For a step λ>0\lambda>0λ>0 and arbitrary starting points z1,zˉ0∈Ez^1,\bar z^0\in\mathcal Ez1,zˉ0∈E, the Golden Ratio Algorithm generates, for k≥1k\ge1k≥1,

zˉk=(φ−1)zk+zˉk−1φ,zk+1=prox⁡λg(zˉk−λF(zk)).(6)\bar z^k = \frac{(\varphi-1)z^k + \bar z^{k-1}}{\varphi},\qquad z^{k+1} = \operatorname{prox}_{\lambda g}\big(\bar z^k - \lambda F(z^k)\big). \tag{6}zˉk=φ(φ−1)zk+zˉk−1​,zk+1=proxλg​(zˉk−λF(zk)).(6)

The first line is a convex combination of the newest iterate and the previous average; the second is a forward–backward step taken from the average rather than from zkz^kzk.

Formalization targets

Goal: Theorem 1

If FFF is LLL-Lipschitz on dom⁡g\operatorname{dom} gdomg (L>0L>0L>0), (C1)–(C3) hold, and λ∈(0,φ2L]\lambda\in\big(0,\frac{\varphi}{2L}\big]λ∈(0,2Lφ​], then there is z∗∈Sz^*\in Sz∗∈S with

zk→z∗andzˉk→z∗(k→∞).z^k\to z^*\qquad\text{and}\qquad \bar z^k\to z^*\qquad(k\to\infty).zk→z∗andzˉk→z∗(k→∞).

Both sequences converge, to one and the same solution. The goal is stated with the paper's exact step range; no rate is claimed, as the paper claims none.

Milestones

  1. Eq. (4), the prox-inequality: for proper convex lsc ggg,
xˉ=prox⁡gz  ⟺  ⟨xˉ−z,x−xˉ⟩≥g(xˉ)−g(x)∀x∈E.\bar x = \operatorname{prox}_g z \iff \langle\bar x - z, x-\bar x\rangle\ge g(\bar x)-g(x)\quad\forall x\in\mathcal E.xˉ=proxg​z⟺⟨xˉ−z,x−xˉ⟩≥g(xˉ)−g(x)∀x∈E.
  1. Eq. (12), an identity using only the averaging step of (6): for every point z∗z^*z∗,
∥zk+1−z∗∥2=(1+φ)∥zˉk+1−z∗∥2−φ∥zˉk−z∗∥2+1φ∥zk+1−zˉk∥2.\|z^{k+1}-z^*\|^2 = (1+\varphi)\|\bar z^{k+1}-z^*\|^2-\varphi\|\bar z^k-z^*\|^2+\tfrac1\varphi\|z^{k+1}-\bar z^k\|^2 .∥zk+1−z∗∥2=(1+φ)∥zˉk+1−z∗∥2−φ∥zˉk−z∗∥2+φ1​∥zk+1−zˉk∥2.
  1. Eq. (14), the energy inequality: for z∗∈Sz^*\in Sz∗∈S and k≥2k\ge2k≥2,
(1+φ)∥zˉk+1−z∗∥2+φ2∥zk+1−zk∥2≤(1+φ)∥zˉk−z∗∥2+φ2∥zk−zk−1∥2−φ∥zk−zˉk∥2.(1+\varphi)\|\bar z^{k+1}-z^*\|^2+\tfrac\varphi2\|z^{k+1}-z^k\|^2\le(1+\varphi)\|\bar z^k-z^*\|^2+\tfrac\varphi2\|z^k-z^{k-1}\|^2-\varphi\|z^k-\bar z^k\|^2 .(1+φ)∥zˉk+1−z∗∥2+2φ​∥zk+1−zk∥2≤(1+φ)∥zˉk−z∗∥2+2φ​∥zk−zk−1∥2−φ∥zk−zˉk∥2.
  1. Lemma 1 (Bauschke–Combettes, Theorem 5.5): a sequence that is Fejér monotone with respect to a nonempty set CCC and whose cluster points all lie in CCC converges to a point of CCC.

Significance

The result. Theorem 1 shows that monotone variational inequalities with a Lipschitz operator can be solved with one operator evaluation and one proximal step per iteration, with FFF evaluated only at points of dom⁡g\operatorname{dom} gdomg, where it is defined. This matters when FFF is expensive (a large matrix–vector product, a simulation) or undefined outside the feasible set (for instance an operator involving log⁡x\log xlogx on the positive orthant). The analysis also explains the constant: the averaging weight φ\varphiφ is the largest ccc with 1/c≥c−11/c\ge c-11/c≥c−1, and the step bound φ/(2L)\varphi/(2L)φ/(2L) follows from it. The fixed-step analysis is the template for the explicit, adaptive-step EGRAAL of the same paper (Theorem 2), which needs only local Lipschitz continuity of FFF.

The formalization. The theorem has a published proof, and no machine-checked version of it or of GRAAL is known. Mathlib contains the golden ratio, Lipschitz conditions, lower semicontinuity and cluster points, but no proximal operator of an extended-valued function, no prox-inequality and no Fejér-monotonicity convergence lemma. This mission produces those pieces and a complete convergence proof for a first-order VI method, which are reusable for projected gradient, forward–backward, extragradient and reflected-gradient analyses.

Difficulty

The naive approach, to show that ∥zk−z∗∥\|z^k-z^*\|∥zk−z∗∥ decreases, fails: GRAAL is not Fejér monotone in zkz^kzk, because the forward step is taken from the average zˉk\bar z^kzˉk and uses F(zk)F(z^k)F(zk) rather than FFF at the new point. The quantity that decreases is an energy mixing ∥zˉk−z∗∥2\|\bar z^k-z^*\|^2∥zˉk−z∗∥2 with the successive difference ∥zk−zk−1∥2\|z^k-z^{k-1}\|^2∥zk−zk−1∥2, and both the averaging identity and the Lipschitz estimate must produce matching coefficients for the cross terms to cancel. The energy inequality alone gives only boundedness and vanishing successive differences; convergence of the whole sequence, and the fact that the limit solves (1) when ggg is merely lower semicontinuous and extended-valued, is a separate step. On the formal side, ggg takes the value +∞+\infty+∞, so the prox-inequality and the variational inequality must be handled in extended arithmetic without letting ∞−∞\infty-\infty∞−∞ decide anything.

Formalization scope

  • E\mathcal EE is a type E with [NormedAddCommGroup E] [InnerProductSpace ℝ E] [FiniteDimensional ℝ E].
  • ggg is E → EReal. (C2) is IsProperConvexLSC g: never ⊥\bot⊥, somewhere finite, convex epigraph {(x,t)∈E×R:g(x)≤t}\{(x,t)\in E\times\mathbb R: g(x)\le t\}{(x,t)∈E×R:g(x)≤t}, and LowerSemicontinuous g on all of E. dom⁡g\operatorname{dom} gdomg is effDom g = {x | g x ≠ ⊤}.
  • FFF is a total function E → E; monotonicity and the Lipschitz bound ∥F(u)−F(v)∥≤L∥u−v∥\|F(u)-F(v)\|\le L\|u-v\|∥F(u)−F(v)∥≤L∥u−v∥ are required on effDom g only. The step range is 0 < λ, λ ≤ φ / (2 * L) with 0 < L and φ = Real.goldenRatio.
  • SSS is solutionSet g F: points of effDom g satisfying (1) for every z∈Ez\in Ez∈E, evaluated in EReal.
  • The proximal step is the argmin predicate IsProxPoint (fun x => λ * g x) w z⁺, not a choice function, so no junk value is involved. A run of (6) is IsGRAALRun g F λ z zbar on sequences ℕ → E indexed as in the paper: z1z^1z1 and zˉ0\bar z^0zˉ0 are free and the entry z0z^0z0 is unused.
  • The conclusion is ∃ zs ∈ solutionSet g F, Tendsto z atTop (𝓝 zs) ∧ Tendsto zbar atTop (𝓝 zs).

The hypotheses of the goal are jointly satisfiable, so the theorem is not vacuous: for g≡0g\equiv0g≡0 and F≡0F\equiv0F≡0 every point is a solution and constant sequences form a run of (6); a formalization under which IsGRAALRun has no instances, or in which SSS may be empty, is ruled out. Two hypotheses are added to printed statements and flagged in their notes: C≠∅C\neq\emptysetC=∅ in Lemma 1, which is false without it, and z1∈dom⁡gz^1\in\operatorname{dom} gz1∈domg in Eq. (14), needed at k=2k=2k=2 because the paper's FFF is only defined on dom⁡g\operatorname{dom} gdomg.

Welcome contributions: existence and uniqueness of the proximal point of a proper convex lsc function in finite dimensions; the prox-inequality; Fejér-monotonicity lemmas; the energy inequality; and the final convergence argument. The prox and Fejér infrastructure is independent of the golden ratio and is shared with the second mission of this series.

Selected references

  • Y. Malitsky, Golden Ratio Algorithms for Variational Inequalities, preprint, Optimization Online 6598, 2018. https://optimization-online.org/wp-content/uploads/2018/05/6598.pdf ; published in Mathematical Programming. https://doi.org/10.1007/s10107-019-01416-w
  • H. H. Bauschke, P. L. Combettes, Convex Analysis and Monotone Operator Theory in Hilbert Spaces, Springer, 2011 (2nd ed. 2017). https://doi.org/10.1007/978-3-319-48311-5
  • G. M. Korpelevich, The extragradient method for finding saddle points and other problems, Ekonomika i Matematicheskie Metody 12 (1976) 747–756.
  • P. Tseng, A modified forward–backward splitting method for maximal monotone mappings, SIAM J. Control Optim. 38 (2000) 431–446. https://doi.org/10.1137/S0363012998338806
  • Y. Malitsky, Projected reflected gradient methods for monotone variational inequalities, SIAM J. Optim. 25 (2015) 502–520. https://doi.org/10.1137/14097238X
8 thms3 active usersReviewed
OptimizationProbability·Captain: mikedeng1

Advance Demand Information, Price Discrimination, and Preorder Strategies: When Preorder Profit Increases with Demand CorrelationResearch Paper

Motivation

Firms that sell new products (consoles, books, films) often take preorders before release. A preorder serves two purposes at once. It lets the firm charge early adopters a different price from later buyers, a form of price discrimination, and the number of preorders is advance demand information: an early signal of how large regular-season demand will be. Li and Zhang (MSOM 15(1), 2013) ask whether better advance information always helps a seller who runs a preorder, when consumers are strategic and anticipate the seller's stocking decision. Their answer is no. The second-period stocking decision improves, but the preorder price can fall. Whether the net effect is positive depends on the low-type margin and the size of the early-adopter segment.

This mission formalizes that answer, §4 of the paper: the preorder profit as a function of the demand correlation ρ\rhoρ.

Setting

A seller sells a perishable product over two periods. High-type consumers, with valuation vHv_HvH​, arrive in the first period and may preorder at price p1p_1p1​. Low-type consumers, with valuation vL<vHv_L < v_HvL​<vH​, arrive in the second period, when the price is p2p_2p2​. A high type who waits values the product at δvH\delta v_HδvH​, where δ≤1\delta \le 1δ≤1 and δvH>vL\delta v_H > v_LδvH​>vL​. The unit cost is ccc, with 0<c<vL0 < c < v_L0<c<vL​. Unsold units have no salvage value and unmet demand carries no penalty. Write Δ=δvH−vL\Delta = \delta v_H - v_LΔ=δvH​−vL​.

The demands are jointly normal with correlation ρ∈[0,1)\rho \in [0,1)ρ∈[0,1). The high-type demand has mean μH\mu_HμH​, and XXX denotes its standardization, a standard normal variable. The low-type demand has mean μL\mu_LμL​, standard deviation σL\sigma_LσL​ and λL=μL/σL\lambda_L = \mu_L/\sigma_LλL​=μL​/σL​. Given X=xX = xX=x, the updated low-type demand X~L(x)\tilde X_L(x)X~L​(x) is normal with mean μ~L(x)=μL+ρσLx\tilde\mu_L(x) = \mu_L + \rho\sigma_L xμ~​L​(x)=μL​+ρσL​x and standard deviation σ~L=σL1−ρ2\tilde\sigma_L = \sigma_L\sqrt{1-\rho^2}σ~L​=σL​1−ρ2​. Φ\PhiΦ and ϕ\phiϕ are the standard normal distribution function and density, and zLz_LzL​ solves Φ(zL)=(vL−c)/vL\Phi(z_L) = (v_L - c)/v_LΦ(zL​)=(vL​−c)/vL​.

In the second period the seller charges p2=vLp_2 = v_Lp2​=vL​ and solves a newsvendor problem. It orders

Q(x)=μ~L(x)+zLσ~L(1)Q(x) = \tilde\mu_L(x) + z_L\tilde\sigma_L \qquad (1)Q(x)=μ~​L​(x)+zL​σ~L​(1)

and earns ΠL(x)=(vL−c)(μL+ρσLx)−vLϕ(zL)σL1−ρ2\Pi_L(x) = (v_L - c)(\mu_L + \rho\sigma_L x) - v_L\phi(z_L)\sigma_L\sqrt{1-\rho^2}ΠL​(x)=(vL​−c)(μL​+ρσL​x)−vL​ϕ(zL​)σL​1−ρ2​ (2). A waiting high type believes half of the remaining consumers will be served before her, so her belief of product availability is

ξ(ρ)=E[Pr⁡(X~L(X)2<Q(X))].(3)\xi(\rho) = \mathbb E\Big[\Pr\Big(\tfrac{\tilde X_L(X)}{2} < Q(X)\Big)\Big]. \qquad (3)ξ(ρ)=E[Pr(2X~L​(X)​<Q(X))].(3)

In the rational-expectations equilibrium every high type preorders at p1=vH−Δξp_1 = v_H - \Delta\xip1​=vH​−Δξ. The preorder profit is

Πp(ρ)=(vH−Δξ(ρ)−c)μH+ΠL(0).(4)\Pi^p(\rho) = (v_H - \Delta\xi(\rho) - c)\mu_H + \Pi_L(0). \qquad (4)Πp(ρ)=(vH​−Δξ(ρ)−c)μH​+ΠL​(0).(4)

The threshold is

μ~(ρ)=−vLϕ(zL)σL2ΔzL ϕ(2zL1−ρ2+λL).(6)\tilde\mu(\rho) = -\frac{v_L\phi(z_L)\sigma_L}{2\Delta z_L\,\phi\big(2z_L\sqrt{1-\rho^2} + \lambda_L\big)}. \qquad (6)μ~​(ρ)=−2ΔzL​ϕ(2zL​1−ρ2​+λL​)vL​ϕ(zL​)σL​​.(6)

Formalization targets

Goal: PROPOSITION 2

(i) c<vL<2c: Πp strictly increases on every interval where μH<μ~(ρ),and strictly decreases on every interval where μH≥μ~(ρ);(ii) vL≥2c: Πp strictly increases on [0,1).\begin{aligned} &\text{(i) } c < v_L < 2c:\ \Pi^p \text{ strictly increases on every interval where } \mu_H < \tilde\mu(\rho),\\ &\qquad\text{and strictly decreases on every interval where } \mu_H \ge \tilde\mu(\rho);\\ &\text{(ii) } v_L \ge 2c:\ \Pi^p \text{ strictly increases on } [0,1). \end{aligned}​(i) c<vL​<2c: Πp strictly increases on every interval where μH​<μ~​(ρ),and strictly decreases on every interval where μH​≥μ~​(ρ);(ii) vL​≥2c: Πp strictly increases on [0,1).​

Milestones

  1. (1)–(2): Q(x)Q(x)Q(x) is the unique maximizer of E[vLmin⁡(Q,X~L(x))−cQ]\mathbb E[v_L\min(Q,\tilde X_L(x)) - cQ]E[vL​min(Q,X~L​(x))−cQ], and its value is ΠL(x)\Pi_L(x)ΠL​(x).
  2. (3): ξ(ρ)=E[Φ((λL+ρX)/1−ρ2+2zL)]\xi(\rho) = \mathbb E\big[\Phi\big((\lambda_L + \rho X)/\sqrt{1-\rho^2} + 2z_L\big)\big]ξ(ρ)=E[Φ((λL​+ρX)/1−ρ2​+2zL​)].
  3. LEMMA 1(i): if c<vL<2cc < v_L < 2cc<vL​<2c, then zL<0z_L < 0zL​<0 and ξ\xiξ is strictly increasing in ρ\rhoρ.
  4. LEMMA 1(ii): if vL≥2cv_L \ge 2cvL​≥2c, then zL≥0z_L \ge 0zL​≥0 and ξ\xiξ is non-increasing in ρ\rhoρ, strictly so when vL>2cv_L > 2cvL​>2c.
  5. After (5): ddρΠL(0)=vLϕ(zL)σL ρ/1−ρ2>0\dfrac{d}{d\rho}\Pi_L(0) = v_L\phi(z_L)\sigma_L\,\rho/\sqrt{1-\rho^2} > 0dρd​ΠL​(0)=vL​ϕ(zL​)σL​ρ/1−ρ2​>0 for ρ∈(0,1)\rho \in (0,1)ρ∈(0,1).
  6. PROPOSITION 2(i), pointwise: for ρ∈(0,1)\rho \in (0,1)ρ∈(0,1), dΠp/dρ>0  ⟺  μH<μ~(ρ)d\Pi^p/d\rho > 0 \iff \mu_H < \tilde\mu(\rho)dΠp/dρ>0⟺μH​<μ~​(ρ) and dΠp/dρ<0  ⟺  μH>μ~(ρ)d\Pi^p/d\rho < 0 \iff \mu_H > \tilde\mu(\rho)dΠp/dρ<0⟺μH​>μ~​(ρ).
  7. LEMMA 2(i): if −λL/2<zL<0-\lambda_L/2 < z_L < 0−λL​/2<zL​<0, then μ~\tilde\muμ~​ is strictly increasing in ρ\rhoρ.
  8. LEMMA 2(ii): if zL≤−λL/2z_L \le -\lambda_L/2zL​≤−λL​/2, then μ~\tilde\muμ~​ is quasi-convex in ρ\rhoρ.

Significance

The proposition splits the value of advance demand information into two effects with opposite signs. Better information always raises the second-period profit (milestone 5). Its effect on the preorder price depends on the margin. When vL<2cv_L < 2cvL​<2c, the seller stocks below the conditional mean. A more precise forecast then raises the stock and the availability ξ\xiξ, and waiting becomes more attractive, which lowers the preorder price. The threshold μ~(ρ)\tilde\mu(\rho)μ~​(ρ) says which effect wins, and LEMMA 2 describes its shape. This is the basis for the paper's later comparisons of preorder, price-guarantee and no-preorder strategies.

The paper states these results and leaves the proofs to an online appendix. None of them has a machine-checked proof. A complete development would also give reusable facts about normal laws: the expectation E[Φ(a+bX)]\mathbb E[\Phi(a + bX)]E[Φ(a+bX)] for standard normal XXX, the normal newsvendor solution with its closed-form optimal profit, and derivatives of Gaussian integrals with respect to a correlation parameter.

Difficulty

The model is explicit, so the difficulty is not in modelling. It is in turning the two expectations into closed forms and differentiating them. The availability ξ\xiξ is an integral over XXX of a normal probability whose mean and variance both move with ρ\rhoρ. The expected newsvendor profit involves E[min⁡(Q,Y)]\mathbb E[\min(Q, Y)]E[min(Q,Y)] for a normal YYY. Neither closed form is in Mathlib, and neither is a derivative in ρ\rhoρ of a Gaussian integral. The obvious route, differentiating under the integral sign in (3), needs domination estimates that are uniform in ρ\rhoρ near each point, and these degenerate as ρ→1\rho \to 1ρ→1.

Signs are the second difficulty. zLz_LzL​ changes sign at vL=2cv_L = 2cvL​=2c, μ~\tilde\muμ~​ carries a leading minus and divides by zLz_LzL​, and the derivative of Πp\Pi^pΠp vanishes at ρ=0\rho = 0ρ=0 and wherever μ~(ρ)=μH\tilde\mu(\rho) = \mu_Hμ~​(ρ)=μH​. Strict monotonicity on an interval has to be recovered from a derivative that is positive except at finitely many points.

Formalization scope

Everything lives in the namespace PreorderADI.Correlation. A structure Params holds vH,vL,c,δ,μH,μL,σL,zLv_H, v_L, c, \delta, \mu_H, \mu_L, \sigma_L, z_LvH​,vL​,c,δ,μH​,μL​,σL​,zL​. The predicate Params.Standing records the model's assumptions: vH>vLv_H > v_LvH​>vL​, c<vLc < v_Lc<vL​, δ≤1\delta \le 1δ≤1 and δvH>vL\delta v_H > v_LδvH​>vL​, together with Φ(zL)=(vL−c)/vL\Phi(z_L) = (v_L-c)/v_LΦ(zL​)=(vL​−c)/vL​. σH\sigma_HσH​ is omitted because nothing in §4 uses it after standardization. Φ\PhiΦ is ProbabilityTheory.cdf (gaussianReal 0 1) and ϕ\phiϕ is gaussianPDFReal 0 1. The normal law with mean mmm and standard deviation sss is gaussianReal m (s^2).

The formalization commits to the following conventions and additions:

  • Added hypotheses. Four hypotheses are added to the page's assumptions:
    • c>0c > 0c>0, so that zLz_LzL​ exists;
    • σL>0\sigma_L > 0σL​>0, so that the conditional law is a genuine normal law;
    • μL>0\mu_L > 0μL​>0, so that λL>0\lambda_L > 0λL​>0, as LEMMA 2 presupposes;
    • μH>0\mu_H > 0μH​>0, which PROPOSITION 2(ii) needs.
  • zLz_LzL​. zLz_LzL​ is a parameter pinned down by Φ(zL)=(vL−c)/vL\Phi(z_L) = (v_L-c)/v_LΦ(zL​)=(vL​−c)/vL​, not an inverse function with junk values.
  • Range of ρ\rhoρ. ρ\rhoρ ranges over [0,1)[0,1)[0,1), the paper's "we will focus on ρ≥0\rho \ge 0ρ≥0" together with ρ<1\rho < 1ρ<1. Derivative statements use (0,1)(0,1)(0,1).
  • ξ\xiξ and Πp\Pi^pΠp. ξ\xiξ is defined as the first expression of (3), the expected probability, so that no closed form is assumed. Πp\Pi^pΠp is defined by (4). PROPOSITION 1, which derives (4) as the unique equilibrium profit, is not formalized: the paper defines the equilibrium only through conditions that already assume every high type preorders.
  • Monotonicity. "Increases in ρ\rhoρ when μH<μ~(ρ)\mu_H < \tilde\mu(\rho)μH​<μ~​(ρ)" is formalized as strict monotonicity on every order-connected I⊆[0,1)I \subseteq [0,1)I⊆[0,1) on which the condition holds. "Decreasing" in LEMMA 1(ii) is non-strict, since ξ\xiξ is constant when vL=2cv_L = 2cvL​=2c. Quasi-convexity is Mathlib's QuasiconvexOn.
  • The ε_v sentence is omitted. PROPOSITION 2(i) also claims that some εv>0\varepsilon_v > 0εv​>0 makes Πp\Pi^pΠp always decrease when vL<c+εvv_L < c + \varepsilon_vvL​<c+εv​. That claim is false in the paper's own model. As vL↓cv_L \downarrow cvL​↓c, μ~(ρ)→∞\tilde\mu(\rho) \to \inftyμ~​(ρ)→∞ for ρ\rhoρ below roughly 3/2\sqrt 3/23​/2, so Πp\Pi^pΠp increases there for every fixed μH\mu_HμH​. The paper's Figure 1 shows the same behaviour.
  • Out of scope. §5 rests on an approximation that treats normal demands as nonnegative, so it is excluded. §§6–7 depend on models given only in the online appendix and are excluded too.

A statement about the closed form Φ(λL+2zL1−ρ2)\Phi(\lambda_L + 2z_L\sqrt{1-\rho^2})Φ(λL​+2zL​1−ρ2​) in place of ξ\xiξ would make LEMMA 1 a one-line monotonicity fact. So would a definition of Πp\Pi^pΠp that bypasses (3). The definitions rule both out.

Welcome contributions include proofs of the milestones, general Mathlib-style lemmas on Gaussian expectations of Φ\PhiΦ and of min⁡(Q,Y)\min(Q, Y)min(Q,Y), and a formalization of PROPOSITION 1 from conditions (i)–(v).

Selected references

  • C. Li and F. Zhang, Advance Demand Information, Price Discrimination, and Preorder Strategies, Manufacturing & Service Operations Management 15(1):57–71, 2013. https://doi.org/10.1287/msom.1120.0398
  • G. P. Cachon and R. Swinney, Purchasing, Pricing, and Quick Response in the Presence of Strategic Consumers, Management Science 55(3):497–511, 2009. https://doi.org/10.1287/mnsc.1080.0948
  • X. Su and F. Zhang, On the Value of Commitment and Availability Guarantees When Selling to Strategic Consumers, Management Science 55(5):713–726, 2009. https://doi.org/10.1287/mnsc.1080.0967
10 thms3 active usersReviewed
🏆Completed
Graph TheoryOptimizationTheoretical Computer Science·Captain: mikedeng1

A New Approach to the Maximum-Flow Problem 2: The Nonsaturating-Push Bound for FIFO Push-RelabelResearch Paper

Motivation

The maximum-flow problem asks how much of a commodity can be sent from a source to a sink through a network whose edges carry capacities. It is a basic model of operations research. Transportation, scheduling, bipartite matching and image segmentation reduce to it, and it is the inner step of many combinatorial algorithms.

Goldberg and Tarjan introduced the push–relabel (preflow) method in A New Approach to the Maximum-Flow Problem (J. ACM 35(4), 1988). Ford–Fulkerson-type algorithms augment along whole source–sink paths. The push–relabel method instead moves excess flow across single edges, guided by integer distance labels on the vertices. Whatever order its local operations are applied in, it is correct and performs O(n2m)O(n^2 m)O(n2m) of them (§3 of the paper). Section 4 shows that one particular order, processing the active vertices first-in, first-out, cuts the dominant term, the number of nonsaturating pushes, to O(n3)O(n^3)O(n3). The method and its FIFO and highest-label variants are the standard practical maximum-flow codes.

Timeline:

  • 1956: Ford and Fulkerson, augmenting paths and max-flow min-cut.
  • 1970–72: Dinic, and Edmonds and Karp, give polynomial augmenting-path bounds.
  • 1974: Karzanov introduces preflows and obtains O(n3)O(n^3)O(n3).
  • 1982: Shiloach and Vishkin give a parallel O(n2log⁡n)O(n^2 \log n)O(n2logn) preflow algorithm with a first-in, first-out flavour.
  • 1988: Goldberg and Tarjan, the generic push–relabel method, the FIFO bound of this mission, and O(nmlog⁡(n2/m))O(nm \log(n^2/m))O(nmlog(n2/m)) with dynamic trees.

Setting

A flow network has a finite vertex set VVV with n=∣V∣n = |V|n=∣V∣, a capacity c(v,w)≥0c(v,w) \ge 0c(v,w)≥0 on every ordered pair, a source sss and a sink t≠st \neq st=s. The edges are the pairs with c(v,w)>0c(v,w) > 0c(v,w)>0, and there are no loops. A preflow is a function fff on vertex pairs with f(v,w)≤c(v,w)f(v,w) \le c(v,w)f(v,w)≤c(v,w) and f(v,w)=−f(w,v)f(v,w) = -f(w,v)f(v,w)=−f(w,v). Its excess e(v)=∑uf(u,v)e(v) = \sum_u f(u,v)e(v)=∑u​f(u,v) must be nonnegative at every v≠sv \neq sv=s. The residual capacity is rf(v,w)=c(v,w)−f(v,w)r_f(v,w) = c(v,w) - f(v,w)rf​(v,w)=c(v,w)−f(v,w). A labeling ddd assigns each vertex a value in N∪{∞}\mathbb{N} \cup \{\infty\}N∪{∞}. A vertex v∉{s,t}v \notin \{s,t\}v∈/{s,t} is active if d(v)<∞d(v) < \inftyd(v)<∞ and e(v)>0e(v) > 0e(v)>0.

The two basic operations (Fig. 1 of the paper) are:

  • push(v,w)(v,w)(v,w), applicable when vvv is active, rf(v,w)>0r_f(v,w) > 0rf​(v,w)>0 and d(v)=d(w)+1d(v) = d(w)+1d(v)=d(w)+1. It sends δ=min⁡(e(v),rf(v,w))\delta = \min(e(v), r_f(v,w))δ=min(e(v),rf​(v,w)) from vvv to www. The push is saturating if rf(v,w)=0r_f(v,w) = 0rf​(v,w)=0 afterwards and nonsaturating otherwise.
  • relabel(v)(v)(v), applicable when vvv is active and d(v)≤d(w)d(v) \le d(w)d(v)≤d(w) for every residual edge (v,w)(v,w)(v,w). It sets d(v)←min⁡{d(w)+1:rf(v,w)>0}d(v) \leftarrow \min\{d(w)+1 : r_f(v,w) > 0\}d(v)←min{d(w)+1:rf​(v,w)>0}.

The algorithm starts by saturating every edge leaving sss, with d(s)=nd(s) = nd(s)=n and d(v)=0d(v) = 0d(v)=0 for v≠sv \neq sv=s.

In the first-in, first-out algorithm (§4), each vertex vvv scans a fixed list L(v)L(v)L(v) of its neighbours through a current edge. The push/relabel(v)(v)(v) operation pushes through the current edge if possible. Otherwise it advances the current edge, or, at the end of the list, returns to the first edge and relabels vvv. Active vertices wait in a queue QQQ, initially {v∈V−{s,t}:c(s,v)>0}\{v \in V - \{s,t\} : c(s,v) > 0\}{v∈V−{s,t}:c(s,v)>0}. The discharge operation removes the front vertex vvv and repeats push/relabel(v)(v)(v) until e(v)=0e(v) = 0e(v)=0 or d(v)d(v)d(v) increases. Every vertex that becomes active meanwhile is appended to QQQ, and vvv is appended too if it is still active. Passes over the queue are defined inductively. Pass 1 consists of the discharges of the initially queued vertices. Pass i+1i+1i+1 consists of the discharges of vertices added during pass iii.

Formalization targets

Goal: Corollary 4.4 (p. 931)

For every network, every edge-list order, every initial queue order, and every run of the FIFO algorithm,

#{nonsaturating pushes}≤4n3.\#\{\text{nonsaturating pushes}\} \le 4n^3 .#{nonsaturating pushes}≤4n3.

The constant is the printed one.

Milestones

  • Lemma 4.1 (p. 929): the push/relabel operation relabels only when relabeling is applicable.
  • Lemma 3.5 (p. 926): from any vertex with positive excess, the source is reachable in the residual graph.
  • Lemma 3.7 (p. 927): at any time, d(v)≤2n−1d(v) \le 2n-1d(v)≤2n−1 for every vertex.
  • Lemma 3.8 (p. 927): at most 2n−12n-12n−1 relabelings per vertex and at most (2n−1)(n−2)<2n2(2n-1)(n-2) < 2n^2(2n−1)(n−2)<2n2 in total.
  • Lemma 4.3 (p. 930): at most 4n24n^24n2 passes over the queue.

Significance

Corollary 4.4 is the combinatorial core of Theorem 4.5, which states that the FIFO algorithm runs in O(n3)O(n^3)O(n3) time. Theorem 4.2 shows that the remaining work of the implementation is O(nm)O(nm)O(nm) plus constant time per nonsaturating push. The bound of Corollary 4.4 is therefore what separates the O(n3)O(n^3)O(n3) FIFO method from the O(n2m)O(n^2 m)O(n2m) bound of the generic method, which matters on dense networks. The same pass-counting argument is reused for the parallel algorithm of §6 and underlies later analyses of highest-label and wave variants.

The results are proved in the paper. Formalizing them adds an analysis of a push–relabel algorithm, which the platform does not yet have. Its existing network-flow material states max-flow min-cut and Ford–Fulkerson termination in an arc-based model with nonnegative flows (the Introduction to Linear Optimization missions). The mission builds a precise operational model of the FIFO implementation, with edge lists, current edges and a queue carrying pass numbers, and states an explicit operation count for it. A companion mission in this series treats the generic algorithm's correctness and its (2n−1)(n−2)+2nm+4n2m(2n-1)(n-2) + 2nm + 4n^2m(2n−1)(n−2)+2nm+4n2m operation bound.

Difficulty

The obvious argument is the potential-function count of §3, over the sum of the labels of active vertices. It yields only 4n2m4n^2 m4n2m and does not use the queue discipline at all. The 4n34n^34n3 bound has to charge nonsaturating pushes to passes over the queue, and then bound the number of passes by the total growth of the labels. Neither step is visible in the generic algorithm, because both depend on the order in which vertices are processed.

Making this rigorous requires invariants of the implementation that the paper uses silently:

  • a vertex is in the queue exactly when it is active, and at most once;
  • pass numbers are nondecreasing along the queue;
  • current edges only move forward between relabelings.

Lemma 4.1 in particular depends on the current-edge scan: an edge passed over earlier stays inadmissible until vvv is relabeled.

Formalization scope

The Lean development works in namespace GoldbergTarjan.FIFO. Vertices form a type V with [Fintype V] [DecidableEq V], and nnn is Fintype.card V. Capacities are c : V → V → ℝ with c ≥ 0 and c v v = 0. Flows are antisymmetric real functions on all ordered pairs, not nonnegative arc flows. Excess is computed from the flow. Labels are in ℕ∞, and the empty minimum in relabel is ⊤.

The state of the algorithm consists of the flow, the labels, the current-edge index cur v into the edge list L v, and the queue Q : List (V × ℕ), each entry tagged with its pass number. Push/relabel (Fig. 3) is a total function, and a discharge (Fig. 4) is a relation carrying the number of push/relabel operations it performs. A run consists of the states S 0, …, S K with S 0 the initial state and consecutive states related by one discharge. The printed variant of Fig. 4, which stops as soon as vvv is relabeled, is the one formalized. Counts are natural numbers over all push/relabel operations of all discharges. The number of passes is the largest pass tag of a discharged entry.

All constants are explicit, exactly as printed:

  • 2n−12n-12n−1 (Lemmas 3.7, 3.8);
  • (2n−1)(n−2)(2n-1)(n-2)(2n−1)(n−2) and 2n22n^22n2 (Lemma 3.8);
  • 4n24n^24n2 (Lemma 4.3);
  • 4n34n^34n3 (Corollary 4.4).

No asymptotic notation is used, and no m≥n−1m \ge n-1m≥n−1 assumption is made.

A model without current edges, where relabeling happens whenever no push applies, would make Lemma 4.1 vacuous and change the algorithm. Pass numbers that are not propagated by the "added during pass iii" rule would make the pass count arbitrary. Both are ruled out by the definitions. A sorry-free check, outside the proposal, exhibits a three-vertex network with two legal discharges, two passes and no nonsaturating push, so the run hypotheses are satisfiable.

Reusable beyond this mission are the network, preflow, push and relabel definitions and Lemma 3.5, which is about an arbitrary preflow. Contributions welcome: invariants of FIFO runs (preflow, valid labeling, queue = active set, cur within bounds), proofs of the milestones, and the reduction of Corollary 4.4 to Lemma 4.3.

Selected references

  • A. V. Goldberg, R. E. Tarjan, A New Approach to the Maximum-Flow Problem, Journal of the ACM 35(4):921–940, 1988. https://doi.org/10.1145/48014.61051
  • A. V. Karzanov, Determining the maximal flow in a network by the method of preflows, Soviet Math. Doklady 15:434–437, 1974.
  • Y. Shiloach, U. Vishkin, An O(n² log n) parallel max-flow algorithm, Journal of Algorithms 3(2):128–146, 1982. https://doi.org/10.1016/0196-6774(82)90013-X
  • L. R. Ford, D. R. Fulkerson, Maximal flow through a network, Canadian Journal of Mathematics 8:399–404, 1956. https://doi.org/10.4153/CJM-1956-045-5
  • J. Edmonds, R. M. Karp, Theoretical improvements in algorithmic efficiency for network flow problems, Journal of the ACM 19(2):248–264, 1972. https://doi.org/10.1145/321694.321699
10 thms3 active usersReviewed
🏆Completed
Graph TheoryOptimizationTheoretical Computer Science·Captain: mikedeng1

A New Approach to the Maximum-Flow Problem 1: The Generic Push-Relabel Algorithm and Its Operation BoundResearch Paper

Motivation

The maximum-flow problem asks how much of a commodity can be sent from a source to a sink through a network whose edges have capacities. It is a basic model in operations research (transportation, scheduling, bipartite matching) and a standard subroutine in combinatorial optimization.

Classical algorithms, from Ford and Fulkerson (1956) through Edmonds–Karp and Dinic (1970–1972) and Karzanov (1974), increase a feasible flow along augmenting paths or blocking flows. Goldberg and Tarjan, A New Approach to the Maximum-Flow Problem (J. ACM 35(4), 1988, doi:10.1145/48014.61051), replaced this global view by a local one: the push-relabel method maintains a preflow, which may violate conservation at intermediate vertices, and moves excess along edges toward vertices with smaller distance labels. The generic method, with the basic operations applied in any order, is the starting point of the FIFO, highest-label and dynamic-tree implementations analysed later in the same paper, and push-relabel codes remain among the fastest practical maximum-flow solvers.

This mission formalizes §2–§3 of the paper: the generic algorithm is correct, and it stops after a number of basic operations bounded by an explicit polynomial in the numbers of vertices and edges, whatever order of operations is chosen.

Setting

A flow network has a finite vertex set VVV with n=∣V∣n = |V|n=∣V∣, a source sss and a sink t≠st \ne st=s, and a capacity c(v,w)≥0c(v,w) \ge 0c(v,w)≥0 for every ordered pair of vertices, positive exactly on the edges E={(v,w):c(v,w)>0}E = \{(v,w) : c(v,w) > 0\}E={(v,w):c(v,w)>0}; m=∣E∣m = |E|m=∣E∣, and there are no loops, c(v,v)=0c(v,v) = 0c(v,v)=0.

Flows are real functions on all vertex pairs. A function fff satisfies the capacity constraint if f(v,w)≤c(v,w)f(v,w) \le c(v,w)f(v,w)≤c(v,w) and antisymmetry if f(v,w)=−f(w,v)f(v,w) = -f(w,v)f(v,w)=−f(w,v) for all pairs. The excess of vvv is e(v)=∑uf(u,v)e(v) = \sum_{u} f(u,v)e(v)=∑u​f(u,v). A flow also has e(v)=0e(v) = 0e(v)=0 for v∉{s,t}v \notin \{s,t\}v∈/{s,t}; a preflow only e(v)≥0e(v) \ge 0e(v)≥0 for v≠sv \ne sv=s. The value of a flow is ∣f∣=∑vf(v,t)|f| = \sum_v f(v,t)∣f∣=∑v​f(v,t), and a maximum flow is a flow of maximum value.

The residual capacity is rf(v,w)=c(v,w)−f(v,w)r_f(v,w) = c(v,w) - f(v,w)rf​(v,w)=c(v,w)−f(v,w); pairs with rf(v,w)>0r_f(v,w) > 0rf​(v,w)>0 are the edges of the residual graph GfG_fGf​. A valid labeling is d:V→N∪{∞}d : V \to \mathbb{N} \cup \{\infty\}d:V→N∪{∞} with d(s)=nd(s) = nd(s)=n, d(t)=0d(t) = 0d(t)=0 and d(v)≤d(w)+1d(v) \le d(w) + 1d(v)≤d(w)+1 on every residual edge. A vertex vvv is active if v∉{s,t}v \notin \{s,t\}v∈/{s,t}, d(v)<∞d(v) < \inftyd(v)<∞ and e(v)>0e(v) > 0e(v)>0.

The two basic operations (Fig. 1 of the paper) are:

  • Push(v,w)(v,w)(v,w), applicable when vvv is active, rf(v,w)>0r_f(v,w) > 0rf​(v,w)>0 and d(v)=d(w)+1d(v) = d(w)+1d(v)=d(w)+1: send δ=min⁡(e(v),rf(v,w))\delta = \min(e(v), r_f(v,w))δ=min(e(v),rf​(v,w)), i.e. f(v,w)+=δf(v,w) \mathrel{+}= \deltaf(v,w)+=δ, f(w,v)−=δf(w,v) \mathrel{-}= \deltaf(w,v)−=δ. It is saturating if rf(v,w)=0r_f(v,w) = 0rf​(v,w)=0 afterwards and nonsaturating otherwise.
  • Relabel(v)(v)(v), applicable when vvv is active and d(v)≤d(w)d(v) \le d(w)d(v)≤d(w) for every residual edge (v,w)(v,w)(v,w): set d(v)←min⁡{d(w)+1:(v,w)∈Ef}d(v) \leftarrow \min\{d(w)+1 : (v,w) \in E_f\}d(v)←min{d(w)+1:(v,w)∈Ef​} (∞\infty∞ if there is none).

The generic algorithm (Fig. 2) starts from the preflow that saturates every edge leaving sss and is zero elsewhere, with the simple labeling d(s)=nd(s) = nd(s)=n, d(v)=0d(v) = 0d(v)=0 otherwise, and applies applicable basic operations in any order while one exists. An execution with KKK basic operations is a sequence of states (f0,d0),…,(fK,dK)(f_0,d_0),\dots,(f_K,d_K)(f0​,d0​),…,(fK​,dK​) from the initial state, each obtained from the previous one by one applicable operation.

Formalization targets

Goal: Theorems 3.11 and 3.4

Assume the paper's standing assumption m≥n−1m \ge n-1m≥n−1. For every execution with KKK basic operations,

K≤(2n−1)(n−2)+2nm+4n2m,K \le (2n-1)(n-2) + 2nm + 4n^2 m,K≤(2n−1)(n−2)+2nm+4n2m,

and if no basic operation applies in the final state, then fKf_KfK​ is a maximum flow. The paper states the bound as O(n2m)O(n^2m)O(n2m) and proves it as "immediate from Lemmas 3.8, 3.9, and 3.10"; the goal states the sum of those three printed bounds. Since every execution is this short, no order of operations runs forever.

Milestones

In the order the proof uses them: Lemma 2.1 (at an active vertex a push or a relabel applies); Lemma 3.1 (the labeling stays valid); Theorem 3.2 (Ford–Fulkerson: a flow is maximum iff ttt is unreachable from sss in GfG_fGf​); Lemma 3.3 (under a valid labeling ttt is unreachable from sss); Lemma 3.5 (from any vertex with positive excess, sss is reachable); Lemma 3.6 (labels never decrease; a relabeling increases the label); Lemma 3.7 (d(v)≤2n−1d(v) \le 2n-1d(v)≤2n−1 throughout); Theorem 3.4 (termination with finite labels gives a maximum flow); Lemma 3.8 (≤2n−1\le 2n-1≤2n−1 relabelings per vertex, ≤(2n−1)(n−2)<2n2\le (2n-1)(n-2) < 2n^2≤(2n−1)(n−2)<2n2 in total); Lemma 3.9 (≤2nm\le 2nm≤2nm saturating pushes); Lemma 3.10 (≤4n2m\le 4n^2m≤4n2m nonsaturating pushes, under m≥n−1m \ge n-1m≥n−1). A further, non-milestone item states the unnumbered invariant that every fkf_kfk​ is a preflow.

Significance

The generic bound shows that push-relabel terminates in a polynomial number of steps without any rule for choosing the next operation; the specific orderings of §4–§5 of the paper (first-in first-out, O(n3)O(n^3)O(n3); dynamic trees, O(nmlog⁡(n2/m))O(nm\log(n^2/m))O(nmlog(n2/m))) refine only the count of nonsaturating pushes, and reuse Lemmas 3.1–3.9 unchanged. The correctness argument, a valid labeling excludes augmenting paths, is the template for the push-relabel minimum-cost flow and assignment algorithms that followed.

These results are proved in the paper and are textbook material. Their machine-checked counterparts are, as far as is known here, not on the Prove2Me platform: the platform's network-flow statements (from Introduction to Linear Optimization, e.g. LinearOptimization.max_flow_min_cut) use a different model, with arc-indexed nonnegative flows and extended-real capacities, and contain nothing about preflows, labels or operation counts. This mission produces a formal account of the antisymmetric-flow model, of Ford–Fulkerson in that model, and of the amortized counting arguments, with the constants the paper prints.

Difficulty

The correctness half is short once the invariants are in place; the difficulty is in the counting. The label bound (Lemma 3.7) is a statement about the whole execution, and it depends on a structural fact about preflows (Lemma 3.5) whose truth rests on antisymmetry and on the nonnegativity of excesses. The obvious first idea for the push counts, bounding pushes per edge or per vertex locally, fails for nonsaturating pushes: flow pushed across a pair can be pushed back later, and nothing local limits how often this happens, so Lemma 3.10 holds only as an amortized statement over the entire execution and depends on both earlier counts. Saturating pushes on a pair can also recur, in both directions, and Lemma 3.9 has to control the interaction between the two directions.

Formally, all of this is reasoning about arbitrary interleavings of operations, with labels in N∪{∞}\mathbb{N} \cup \{\infty\}N∪{∞} and real-valued flows.

Formalization scope

  • Vertices form a finite type with decidable equality; nnn is its cardinality, s≠ts \ne ts=t, so n≥2n \ge 2n≥2 and the natural-number subtractions 2n−12n-12n−1 and n−2n-2n−2 are exact. Capacities are a real function on all pairs, nonnegative, zero on the diagonal; EEE is its support and mmm its cardinality.
  • Flows and preflows are antisymmetric real functions on all pairs (not nonnegative arc flows); the excess is computed from fff, never stored. A maximum flow is a flow whose value is at least that of every flow.
  • Labels live in ℕ∞, with ∞+1=∞\infty + 1 = \infty∞+1=∞; the relabel value is an infimum, which is ∞\infty∞ on the empty set.
  • An execution is a sequence of states σ : ℕ → State V with a length KKK, starting at the Fig. 2 state with the simple labeling (the paper's own assumption for its proofs), each step an applicable push or relabel. "Terminates" means that no basic operation applies, the loop guard of Fig. 2. The three counts are cardinalities of the sets of step indices of each kind.
  • Explicit constants: 2n−12n-12n−1 per-vertex relabelings, (2n−1)(n−2)<2n2(2n-1)(n-2) < 2n^2(2n−1)(n−2)<2n2 total relabelings, 2nm2nm2nm saturating pushes, 4n2m4n^2m4n2m nonsaturating pushes, label bound 2n−12n-12n−1, and the total (2n−1)(n−2)+2nm+4n2m(2n-1)(n-2)+2nm+4n^2m(2n−1)(n−2)+2nm+4n2m. The standing assumption m≥n−1m \ge n-1m≥n−1 appears only on Lemma 3.10 and the goal.
  • A trivializing formalization is ruled out: the step relation fixes the pushed amount δ=min⁡(e(v),rf(v,w))\delta = \min(e(v), r_f(v,w))δ=min(e(v),rf​(v,w)) and the new label exactly as in Fig. 1, termination is the loop guard rather than "the result is a flow", and a sorry-free check exhibits a concrete network s→a→ts \to a \to ts→a→t with a two-step execution (relabel aaa, then push (a,t)(a,t)(a,t)), so the run hypotheses are satisfiable.

Welcome contributions: proofs of the invariants (preflow, valid labeling, label monotonicity), of Ford–Fulkerson for antisymmetric flows (reusable beyond this mission), and of the counting lemmas. The FIFO bound of §4 is the subject of a companion mission.

Selected references

  • A. V. Goldberg, R. E. Tarjan, A New Approach to the Maximum-Flow Problem, Journal of the ACM 35(4):921–940, 1988. doi:10.1145/48014.61051
  • L. R. Ford, D. R. Fulkerson, Flows in Networks, Princeton University Press, 1962.
  • J. Edmonds, R. M. Karp, Theoretical improvements in algorithmic efficiency for network flow problems, Journal of the ACM 19(2):248–264, 1972. doi:10.1145/321694.321699
  • R. K. Ahuja, T. L. Magnanti, J. B. Orlin, Network Flows: Theory, Algorithms, and Applications, Prentice Hall, 1993.
17 thms3 active usersReviewed
Algorithmic Game TheoryOptimizationProbability·Captain: mikedeng1

A Supply Chain Theory of Factoring and Reverse Factoring 2: The Retailer's Optimal Reverse Factoring Payment ExtensionResearch Paper

Motivation

Large retailers pay their suppliers weeks or months after delivery, and small suppliers fill the gap with short-term finance. In factoring the supplier sells the receivable to a factor for immediate cash; in reverse factoring the retailer arranges the program with a bank, which pays the supplier early at a rate priced on the retailer's credit rating. Retailers commonly attach a condition: the supplier must accept a longer payment term. Wuttke et al. (Journal of Operations Management, 2019) report that buyers extended payment terms by 54 days on average on adopting reverse factoring and that many suppliers delayed adoption; Corsten (2010) reports suppliers resisting a program because of the demanded payment delay (both as cited by Kouvelis and Xu, pp. 6082–6083). How long an extension a retailer can demand, and what it gains by demanding it, is therefore a practical design question.

Kouvelis and Xu (Management Science 67(10), 2021) answer it inside a Stackelberg supply chain model with credit and liquidity risk. This mission formalizes their answer, Proposition 6 of §5.3: the retailer's optimal payment extension when she keeps the existing wholesale price.

Setting

Demand D≥0D\ge0D≥0 has density fff, distribution function FFF and Fˉ=1−F\bar F=1-FFˉ=1−F; f>0f>0f>0 on [0,Z][0,\mathbb Z][0,Z] with Z≤+∞\mathbb Z\le+\inftyZ≤+∞ the upper end of the support, fff is continuous there, the mean is finite, and the failure rate z(ξ)=f(ξ)/Fˉ(ξ)z(\xi)=f(\xi)/\bar F(\xi)z(ξ)=f(ξ)/Fˉ(ξ) is strictly increasing. Write S(q)=∫0qFˉ(ξ) dξS(q)=\int_0^q\bar F(\xi)\,d\xiS(q)=∫0q​Fˉ(ξ)dξ for expected sales and k(q)=S(q)/Fˉ(q)k(q)=S(q)/\bar F(q)k(q)=S(q)/Fˉ(q).

A retailer (the leader) sets a wholesale price www, and a capital-constrained supplier (the follower) chooses a production quantity q≥0q\ge0q≥0; the retail price ppp exceeds the unit cost ccc. Each firm j∈{s,r}j\in\{s,r\}j∈{s,r} has a credit rating Cj∈(Cmin⁡,Cmax⁡)C_j\in(C_{\min},C_{\max})Cj​∈(Cmin​,Cmax​), a default probability ρj=ρ(Cj)∈[0,1]\rho_j=\rho(C_j)\in[0,1]ρj​=ρ(Cj​)∈[0,1] with ρ\rhoρ strictly decreasing, and an interest premium ηj=η(Cj)>0\eta_j=\eta(C_j)>0ηj​=η(Cj​)>0 with η\etaη decreasing. The lead time is t1t_1t1​, the payment term t2t_2t2​, and λs,λr≥0\lambda_s,\lambda_r\ge0λs​,λr​≥0 are the liquidity risks.

Under a post-shipment scheme with coefficient Λ\LambdaΛ the supplier earns

π(q;w)=(1−ρs)(Λe−λst1wS(q)−c q eηst1),\pi(q;w)=(1-\rho_s)\bigl(\Lambda e^{-\lambda_s t_1}wS(q)-c\,q\,e^{\eta_s t_1}\bigr),π(q;w)=(1−ρs​)(Λe−λs​t1​wS(q)−cqeηs​t1​),

with ΛF=(1−ρr)+(1−ρs)−eηst2\Lambda_{\mathcal F}=(1-\rho_r)+(1-\rho_s)-e^{\eta_s t_2}ΛF​=(1−ρr​)+(1−ρs​)−eηs​t2​ (recourse factoring), ΛN=e−ηrt2(1−ρr)\Lambda_{\mathcal N}=e^{-\eta_r t_2}(1-\rho_r)ΛN​=e−ηr​t2​(1−ρr​) (non-recourse factoring) and ΛR=e−ηr(t2+τ)\Lambda_{\mathcal R}=e^{-\eta_r(t_2+\tau)}ΛR​=e−ηr​(t2​+τ) (reverse factoring with payment extension τ≥0\tau\ge0τ≥0). The retailer earns Π=e−λst1(1−ρr)(p−w)S(q)\Pi=e^{-\lambda_s t_1}(1-\rho_r)(p-w)S(q)Π=e−λs​t1​(1−ρr​)(p−w)S(q) under factoring and

ΠR(w,τ)=e−λst1(1−ρr)(2−e−λrτ)(p−w)S(qR)\Pi_{\mathcal R}(w,\tau)=e^{-\lambda_s t_1}(1-\rho_r)(2-e^{-\lambda_r\tau})(p-w)S(q_{\mathcal R})ΠR​(w,τ)=e−λs​t1​(1−ρr​)(2−e−λr​τ)(p−w)S(qR​)

under reverse factoring, where qRq_{\mathcal R}qR​ is the supplier's best response. Its first-order condition is wFˉ(qR)=cR(τ)=c e(ηs+λs)t1+ηr(t2+τ)w\bar F(q_{\mathcal R})=c_{\mathcal R}(\tau)=c\,e^{(\eta_s+\lambda_s)t_1+\eta_r(t_2+\tau)}wFˉ(qR​)=cR​(τ)=ce(ηs​+λs​)t1​+ηr​(t2​+τ) (Eq. (12)).

Before reverse factoring, the supplier uses the better of the two factoring schemes. By Proposition 4 this is non-recourse, with equilibrium (wN∗,qN∗)(w^*_{\mathcal N},q^*_{\mathcal N})(wN∗​,qN∗​), when CN<Cs≤C1\mathbb C_{\mathcal N}<C_s\le\mathbb C_1CN​<Cs​≤C1​, and recourse, with (wF∗,qF∗)(w^*_{\mathcal F},q^*_{\mathcal F})(wF∗​,qF∗​), when Cs>CF∨C1C_s>\mathbb C_{\mathcal F}\vee\mathbb C_1Cs​>CF​∨C1​. The retailer keeps the existing wholesale price wsw_sws​ and solves problem (13): maximize ΠR(ws,τ)\Pi_{\mathcal R}(w_s,\tau)ΠR​(ws​,τ) over τ≥0\tau\ge0τ≥0, subject to the supplier's acceptance (his reverse factoring profit is at least his existing one). CRmax⁡\mathbb C^{\max}_{\mathcal R}CRmax​ is the rating at which ΛF=e−ηrt2\Lambda_{\mathcal F}=e^{-\eta_r t_2}ΛF​=e−ηr​t2​, and Ξ[0,z](x)=max⁡{0,min⁡{z,x}}\Xi_{[0,z]}(x)=\max\{0,\min\{z,x\}\}Ξ[0,z]​(x)=max{0,min{z,x}}.

Formalization targets

Goal: Proposition 6

(i) If Cs≥CRmax⁡C_s\ge\mathbb C^{\max}_{\mathcal R}Cs​≥CRmax​, reverse factoring is dominated by recourse factoring. (ii) If CN<Cs<CRmax⁡\mathbb C_{\mathcal N}<C_s<\mathbb C^{\max}_{\mathcal R}CN​<Cs​<CRmax​, reverse factoring should be offered with

τR∗=Ξ[0,τs](τ0∗),λrk(q)z(q)+ηr=2ηreλrτ0∗,wsFˉ(q)=cR(τ0∗),\tau^*_{\mathcal R}=\Xi_{[0,\tau_s]}(\tau^*_0),\qquad \lambda_r k(q)z(q)+\eta_r=2\eta_r e^{\lambda_r\tau^*_0},\quad w_s\bar F(q)=c_{\mathcal R}(\tau^*_0),τR∗​=Ξ[0,τs​]​(τ0∗​),λr​k(q)z(q)+ηr​=2ηr​eλr​τ0∗​,ws​Fˉ(q)=cR​(τ0∗​),

where τs=−ηr−1ln⁡(1−ρr)\tau_s=-\eta_r^{-1}\ln(1-\rho_r)τs​=−ηr−1​ln(1−ρr​) with ws=wN∗w_s=w^*_{\mathcal N}ws​=wN∗​ in the non-recourse case, and τs=−ηr−1ln⁡[(1−ρr)+(1−ρs)−eηst2]−t2\tau_s=-\eta_r^{-1}\ln[(1-\rho_r)+(1-\rho_s)-e^{\eta_s t_2}]-t_2τs​=−ηr−1​ln[(1−ρr​)+(1−ρs​)−eηs​t2​]−t2​ with ws=wF∗w_s=w^*_{\mathcal F}ws​=wF∗​ in the recourse case.

Milestones, in attack order

  1. Eq. (12): the supplier's best response under reverse factoring.
  2. Proposition 4: which factoring scheme is in force before reverse factoring.
  3. §5.3, τs\tau_sτs​: acceptance holds exactly on [0,τs][0,\tau_s][0,τs​].
  4. §5.3, τ0∗\tau^*_0τ0∗​: the retailer's unconstrained profit is unimodal around τ0∗\tau^*_0τ0∗​.

A follow-on item states Corollary 3(ii): the retailer's profit strictly increases, and the supplier's profit is unchanged when τ0∗≥τs\tau^*_0\ge\tau_sτ0∗​≥τs​.

Significance

Proposition 6 is the paper's prescription for program design. It says which suppliers should be offered reverse factoring: every supplier below the indifference rating CRmax⁡\mathbb C^{\max}_{\mathcal R}CRmax​ and above the non-recourse feasibility threshold. It also gives the extension in closed form, the unconstrained optimum clipped to the supplier's acceptance limit. Two consequences are drawn in the paper: non-recourse factoring is dominated once the extension is optimized, and reverse factoring may leave the supplier exactly as well off as before, so it is not necessarily a win-win (Corollary 3).

The proofs are in the paper's Online Appendix B and have not been machine-checked. A formal proof here produces a checked derivation of the projection formula from the model's primitives. It covers the strict-IFR analysis of the follower's response, the reduction of the acceptance constraint to an interval, and the unimodality of the retailer's objective. The same analysis of the pull game with an effective unit cost recurs across the supply chain finance literature.

Difficulty

The retailer's objective depends on τ\tauτ through two opposing channels: the liquidity factor 2−e−λrτ2-e^{-\lambda_r\tau}2−e−λr​τ increases, while expected sales S(qR(τ))S(q_{\mathcal R}(\tau))S(qR​(τ)) decrease because the supplier's effective cost rises. Neither factor is concave in τ\tauτ, and the objective need not be concave. The natural move, to set the derivative to zero and call the root a maximum, proves nothing without a sign analysis. That analysis needs the monotonicity of k⋅zk\cdot zk⋅z along the implicitly defined response qR(τ)q_{\mathcal R}(\tau)qR​(τ), which is where strict IFR enters. The acceptance constraint compares the supplier's profits in two different games (reverse factoring at τ\tauτ against the existing equilibrium). Reducing it to τ≤τs\tau\le\tau_sτ≤τs​ requires the supplier's best-response profit as an explicit increasing function of his quantity. Identifying the existing equilibrium requires Proposition 4, whose "adopted" compares equilibrium profits of two Stackelberg games.

Formalization scope

The model is a single Lean structure SupplyChainFactoring.Extension.Model. Demand is a probability measure on R\mathbb RR with a density fff, and Z\mathbb ZZ is an extended real. "Continuous p.d.f. with f>0f>0f>0 in [0,Z][0,\mathbb Z][0,Z]" is read as continuity on [0,Z][0,\mathbb Z][0,Z], with f=0f=0f=0 outside the support. Credit functions ρ,η\rho,\etaρ,η are real functions constrained on (Cmin⁡,Cmax⁡)(C_{\min},C_{\max})(Cmin​,Cmax​). The finance derivations behind the profit functions (Eqs. (1), (5), (6), Lemma 1) are not formalized: the profit functions are the model.

Readings of informal words, each also recorded in the item's Formalization Note:

  • Best response: a maximizer of the supplier's profit over q≥0q\ge0q≥0; equilibrium: a best response pair from which no nonnegative wholesale price with a best response gives the retailer more. Neither is defined through first-order conditions.
  • Feasible: some w≥0w\ge0w≥0 with a best response gives the retailer positive profit; adopted (Proposition 4): feasible, with equilibrium supplier profit at least (non-recourse) or strictly above (recourse) the other feasible scheme's.
  • Thresholds "the unique value of CsC_sCs​ that satisfies …" are hypotheses in exactly that form; cN=pc_{\mathcal N}=pcN​=p and cF=pc_{\mathcal F}=pcF​=p are cross-multiplied because ΛF\Lambda_{\mathcal F}ΛF​ can be ≤0\le0≤0.
  • In (13) www is fixed at wsw_sws​ (§5.3's first sentence, footnote 23). πR∗\pi^*_{\mathcal R}πR∗​ is the supplier's best-response profit under reverse factoring at (ws,τ)(w_s,\tau)(ws​,τ), and max⁡{πF∗,πN∗}\max\{\pi^*_{\mathcal F},\pi^*_{\mathcal N}\}max{πF∗​,πN∗​} is his profit in the existing equilibrium.
  • Dominated (Proposition 6(i)): at every τ≥0\tau\ge0τ≥0 and every www, the supplier's reverse factoring best-response profit is at most his recourse one. Should be offered (6(ii)): τR∗\tau^*_{\mathcal R}τR∗​ solves (13) and the retailer's profit is at least her existing equilibrium profit.
  • τ0∗\tau^*_0τ0∗​ is a hypothesis: it and some q∈(0,Z)q\in(0,\mathbb Z)q∈(0,Z) solve the paper's two equations (the paper does not argue existence). Its optimality "without the nonnegativity constraint" is stated as unimodality of ΠR\Pi_{\mathcal R}ΠR​ on the set of real τ\tauτ with cR(τ)<wc_{\mathcal R}(\tau)<wcR​(τ)<w.
  • Always increases (Corollary 3(ii)) is strict; may remain unchanged when τ0∗≥τs\tau^*_0\ge\tau_sτ0∗​≥τs​ is read as "is unchanged whenever τ0∗≥τs\tau^*_0\ge\tau_sτ0∗​≥τs​".

Three misprints of the paper are corrected: Ξ[0,z](x)=0\Xi_{[0,z]}(x)=0Ξ[0,z]​(x)=0 "if x<zx<zx<z" is read as "if x<0x<0x<0"; "the retailer's maximization problem in (16)" refers to (13); the middle line of the ΠR\Pi_{\mathcal R}ΠR​ display on p. 6082 carries a stray factor www, and the last line is used.

The hypotheses on τ0∗\tau^*_0τ0∗​ cannot be met when λr=0\lambda_r=0λr​=0, and the goal then says nothing about the case, as in the paper. The existing equilibrium, τs\tau_sτs​ and τ0∗\tau^*_0τ0∗​ are never free parameters: τs\tau_sτs​ is the paper's explicit formula, and the reduction of acceptance to τ≤τs\tau\le\tau_sτ≤τs​ is a milestone to be proved, not an assumption. Every logarithm is applied to a quantity the hypotheses force positive. A formalization that assumed acceptance equivalent to τ≤τs\tau\le\tau_sτ≤τs​ or assumed unimodality would be trivial and is excluded.

The pull game with an effective cost has the same structure as Cachon's pull contract without salvage value (platform items CachonPushPull.*), but those items assume IGFR demand with a salvage value, so they are not reused. Reusable infrastructure welcome: the strict-IFR lemmas (kkk, k⋅zk\cdot zk⋅z and k(q)−qk(q)-qk(q)−q increasing) and the explicit best response of a newsvendor-type follower.

Selected references

  • P. Kouvelis, F. Xu, A Supply Chain Theory of Factoring and Reverse Factoring, Management Science 67(10):6071–6088, 2021. https://doi.org/10.1287/mnsc.2020.3788
  • G. P. Cachon, The Allocation of Inventory Risk in a Supply Chain: Push, Pull, and Advance-Purchase Discount Contracts, Management Science 50(2):222–238, 2004. https://doi.org/10.1287/mnsc.1030.0190
  • D. A. Wuttke, E. S. Rosenzweig, H. S. Heese, An Empirical Analysis of Supply Chain Finance Adoption, Journal of Operations Management 65(3):242–261, 2019. https://doi.org/10.1002/joom.1023
6 thms3 active usersReviewed
🏆Completed
ProbabilityStatistics·Captain: mikedeng1

Are Call Center and Hospital Arrivals Well Modeled by Nonhomogeneous Poisson Processes?: Combining k Equal Subintervals of a Linear Arrival Rate Bounds the Degree of Nonhomogeneity by C/kResearch Paper

Motivation

Arrival processes to call centers and hospital emergency departments are routinely modeled as nonhomogeneous Poisson processes (NHPPs): Poisson processes whose arrival rate varies over the day. Staffing and queueing models built on this assumption are only as good as the assumption itself, so practitioners test it on data. The standard test, going back to Brown et al. (2005, doi:10.1198/016214504000001808), divides the day into short subintervals, treats the rate as constant on each, rescales the arrival times within each subinterval to [0,1][0,1][0,1], combines all the rescaled data, and applies a Kolmogorov–Smirnov (KS) test of uniformity.

Kim and Whitt (2014, doi:10.1287/msom.2014.0490) ask when this piecewise-constant approximation is justified. If the true rate is not constant on a subinterval, the rescaled arrival times are not uniform, and with enough data the KS test rejects the Poisson hypothesis even when the process really is an NHPP. Section 3 of the paper quantifies this effect through a single number, the degree of nonhomogeneity, and shows how it behaves when the interval is cut into kkk equal pieces. This mission formalizes that section's exact computations for a linear arrival rate.

Setting

An arrival rate function λ\lambdaλ on an interval [0,T][0,T][0,T], T>0T > 0T>0, is nonnegative, integrable, and strictly positive except at finitely many points. Its cumulative arrival rate is

Λ(t)=∫0tλ(s) ds.\Lambda(t) = \int_0^t \lambda(s)\,ds .Λ(t)=∫0t​λ(s)ds.

Conditionally on nnn arrivals in [0,T][0,T][0,T], the arrival times of an NHPP with rate λ\lambdaλ, divided by TTT, are distributed as the order statistics of nnn independent random variables on [0,1][0,1][0,1] with the conditional cdf

F(t)=Λ(tT)Λ(T),0≤t≤1.F(t) = \frac{\Lambda(tT)}{\Lambda(T)}, \qquad 0 \le t \le 1 .F(t)=Λ(T)Λ(tT)​,0≤t≤1.

The degree of nonhomogeneity is the Kolmogorov distance of FFF from the uniform cdf,

D=sup⁡0≤t≤1∣F(t)−t∣.D = \sup_{0 \le t \le 1} |F(t) - t| .D=0≤t≤1sup​∣F(t)−t∣.

It is zero exactly when λ\lambdaλ is constant, and it is the limit of the KS test statistic as the amount of data grows.

For k≥1k \ge 1k≥1, divide [0,T][0,T][0,T] into kkk subintervals of length T/kT/kT/k. For 1≤j≤k1 \le j \le k1≤j≤k the jjj-th subinterval has cumulative rate Λj(t)=Λ((j−1)T/k+t)−Λ((j−1)T/k)\Lambda_j(t) = \Lambda((j-1)T/k + t) - \Lambda((j-1)T/k)Λj​(t)=Λ((j−1)T/k+t)−Λ((j−1)T/k), conditional cdf Fj(t)=Λj(tT/k)/Λj(T/k)F_j(t) = \Lambda_j(tT/k)/\Lambda_j(T/k)Fj​(t)=Λj​(tT/k)/Λj​(T/k), and share of arrivals pj=(Λ(jT/k)−Λ((j−1)T/k))/Λ(T)p_j = (\Lambda(jT/k) - \Lambda((j-1)T/k))/\Lambda(T)pj​=(Λ(jT/k)−Λ((j−1)T/k))/Λ(T). The data of all subintervals, each rescaled to [0,1][0,1][0,1] and combined, have the conditional cdf F=∑j=1kpjFjF = \sum_{j=1}^k p_j F_jF=∑j=1k​pj​Fj​ (LEMMA 1).

The linear arrival rate is λ(t)=a+bt\lambda(t) = a + btλ(t)=a+bt with b≥0b \ge 0b≥0 and a≥0a \ge 0a≥0, not identically zero. When a>0a > 0a>0 its relative slope is r=b/ar = b/ar=b/a; on the jjj-th subinterval the relative slope is rj=b/λ((j−1)T/k)r_j = b/\lambda((j-1)T/k)rj​=b/λ((j−1)T/k).

In the Lean development these are cumRate, condCdf, degree, subCum, subCdf, weight, mixCdf, linRate and subSlope, in the namespace NHPPArrivals.LinearRate.

Formalization targets

Goal: THEOREM 5, combining equally spaced subintervals

For the linear rate, there is a constant CCC such that for every k≥1k \ge 1k≥1

D=sup⁡0≤t≤1∣F(t)−t∣=∑j=1kpjDj=∑j=1kpjsup⁡0≤t≤1∣Fj(t)−t∣,(20)D = \sup_{0 \le t \le 1}|F(t) - t| = \sum_{j=1}^k p_j D_j = \sum_{j=1}^k p_j \sup_{0 \le t \le 1}|F_j(t) - t|, \tag{20}D=0≤t≤1sup​∣F(t)−t∣=j=1∑k​pj​Dj​=j=1∑k​pj​0≤t≤1sup​∣Fj​(t)−t∣,(20)

with, if a>0a > 0a>0,

D=∑j=1kpj rjT/k8+4rjT/k,(21)D = \sum_{j=1}^k \frac{p_j\, r_j T/k}{8 + 4 r_j T/k}, \tag{21}D=j=1∑k​8+4rj​T/kpj​rj​T/k​,(21)

and, if a=0a = 0a=0,

D=p14+∑j=2kpj/(j−1)8+4/(j−1),(22)D = \frac{p_1}{4} + \sum_{j=2}^k \frac{p_j/(j-1)}{8 + 4/(j-1)}, \tag{22}D=4p1​​+j=2∑k​8+4/(j−1)pj​/(j−1)​,(22)

and in both cases D≤C/kD \le C/kD≤C/k. The constant CCC may depend on aaa, bbb and TTT, but not on kkk; its value is left open, as in the paper.

Milestones

  1. LEMMA 1, (17): for a general rate, the rescaled combined data have cdf ∑jpjFj\sum_j p_j F_j∑j​pj​Fj​, and the pjp_jpj​ form a probability vector.
  2. THEOREM 4, a>0a > 0a>0, (14), (16): F(t)=(tT+r(tT)2/2)/(T+rT2/2)F(t) = (tT + r(tT)^2/2)/(T + rT^2/2)F(t)=(tT+r(tT)2/2)/(T+rT2/2) and D=∣F(1/2)−1/2∣=rT/(8+4rT)D = |F(1/2) - 1/2| = rT/(8 + 4rT)D=∣F(1/2)−1/2∣=rT/(8+4rT).
  3. THEOREM 4, a=0a = 0a=0, (15): F(t)=t2F(t) = t^2F(t)=t2 and D=1/4D = 1/4D=1/4.
  4. LEMMA 1, (18): closed forms of Λj\Lambda_jΛj​, FjF_jFj​, pjp_jpj​, rjr_jrj​ when a>0a > 0a>0.
  5. LEMMA 1, (19): closed forms of Λj\Lambda_jΛj​, FjF_jFj​, pjp_jpj​, rjr_jrj​ when a=0a = 0a=0.
  6. THEOREM 5, (20): D=∑jpjDjD = \sum_j p_j D_jD=∑j​pj​Dj​ for one fixed kkk.

Significance

The result gives a quantitative criterion for the piecewise-constant approximation: for a linear rate, cutting the interval into kkk equal pieces reduces the degree of nonhomogeneity of the combined data by a factor of order 1/k1/k1/k. Since the KS critical value at sample size nnn is of order 1/n1/\sqrt n1/n​, this tells a practitioner how fine the subintervals must be, relative to the amount of data, before a KS test of the Poisson hypothesis stops rejecting merely because the rate varies within subintervals. The paper's later THEOREM 6 and its practical guidelines (§3.4, §3.6) rest on these formulas.

The results are proved in the paper by direct calculation; none of them has a machine-checked proof. The mission produces a verified library of the conditional-cdf calculus for NHPPs on an interval (the conditional cdf, its degree of nonhomogeneity, the subinterval decomposition) and the exact linear-rate formulas that the testing literature cites.

Difficulty

The computations are elementary, but two steps are not immediate. First, the supremum of ∣F(t)−t∣|F(t) - t|∣F(t)−t∣ over [0,1][0,1][0,1] is a supremum of a nonsmooth function; showing that it is attained at t=1/2t = 1/2t=1/2 requires knowing the sign of F(t)−tF(t) - tF(t)−t on the whole interval, and for the combined cdf it requires that all the pieces FjF_jFj​ attain their maximal deviation at the same point, which is special to linear rates. For a general rate the naive identity D=∑jpjDjD = \sum_j p_j D_jD=∑j​pj​Dj​ fails: the sup of a sum is at most the sum of the sups, with equality only when the maximizers coincide. Second, LEMMA 1 is a statement about the law of a rescaled random variable (the fractional part of kX/TkX/TkX/T), which requires splitting a measure along the kkk subintervals and handling their boundary points.

Formalization scope

Rates are real functions λ:R→R\lambda : \mathbb R \to \mathbb Rλ:R→R; only their values on [0,T][0,T][0,T] enter. Λ\LambdaΛ is an interval integral, subintervals are indexed by j∈{1,…,k}j \in \{1, \dots, k\}j∈{1,…,k} with k,jk, jk,j natural numbers cast to reals, and (j−1)(j-1)(j−1) is computed in R\mathbb RR. All quotients are real divisions; the hypotheses of every statement (T>0T > 0T>0, k≥1k \ge 1k≥1, b≥0b \ge 0b≥0, and a>0a > 0a>0 or b>0b > 0b>0 for the linear rate; integrability, nonnegativity and a finite zero set for a general rate) make every denominator Λ(T)\Lambda(T)Λ(T) and Λj(T/k)\Lambda_j(T/k)Λj​(T/k) positive. b≥0b \ge 0b≥0 is the paper's standing assumption of §3.3; excluding a=b=0a = b = 0a=b=0 is §3.2's requirement that the rate be positive except at finitely many points. The degree of nonhomogeneity is sSup of the image of [0,1][0,1][0,1], and every statement that uses it also asserts that the supremum is attained, so no default value of sSup can make a statement true. The constant CCC of THEOREM 5 is quantified before kkk; choosing it after kkk would make the bound empty. The statements are about the general definitions of (17) applied to λ(t)=a+bt\lambda(t) = a + btλ(t)=a+bt, not about the closed forms (18)–(19), which are separate milestones. The formula for rjr_jrj​ in (19) is stated for 2≤j≤k2 \le j \le k2≤j≤k only: r1=b/λ(0)r_1 = b/\lambda(0)r1​=b/λ(0) is undefined when a=0a = 0a=0.

The Poisson process itself is not formalized. LEMMA 1's "i.i.d. random variables" is the paper's THEOREM 1 (the conditioning property) applied to each arrival; LEMMA 1 is stated for the law of one arrival time, the probability measure with density λ/Λ(T)\lambda/\Lambda(T)λ/Λ(T) on [0,T][0,T][0,T]. THEOREM 1, THEOREMS 2–3 and COROLLARY 1 (limits of the empirical cdf and of the KS statistic) are out of scope: they need a point-process layer, the Glivenko–Cantelli theorem and KS critical values, none of which exists in Mathlib. THEOREM 6 is out of scope because the paper gives only a sketch comparing DDD with the KS critical value.

Contributions welcome: proofs of the milestones, general lemmas on sups of ∣F(t)−t∣|F(t) - t|∣F(t)−t∣ for convex cdfs, and the measure-splitting argument of LEMMA 1, which is reusable for any subinterval-based test of the Poisson hypothesis.

Selected references

  • S.-H. Kim and W. Whitt, Are call center and hospital arrivals well modeled by nonhomogeneous Poisson processes?, Manufacturing & Service Operations Management 16(3):464–480, 2014. doi:10.1287/msom.2014.0490
  • L. Brown, N. Gans, A. Mandelbaum, A. Sakov, H. Shen, S. Zeltyn, L. Zhao, Statistical analysis of a telephone call center: a queueing-science perspective, Journal of the American Statistical Association 100(469):36–50, 2005. doi:10.1198/016214504000001808
  • F. J. Massey, The Kolmogorov–Smirnov test for goodness of fit, Journal of the American Statistical Association 46(253):68–78, 1951. doi:10.1080/01621459.1951.10500769
8 thms3 active usersReviewed
🏆Completed
Dynamic ProgrammingOptimizationTopology·Captain: mikedeng1

The Structure of Dynamic Programing Models: A Solution of the Principle of Optimality with Vanishing Tail Is the Optimal ReturnResearch Paper

Motivation

Dynamic programming, as introduced by Bellman in the early 1950s, solves sequential decision problems through a functional equation: the value of a problem started in a given state equals the best one-stage return plus the value of the problem started in the state that decision leads to. In practice the argument usually runs backwards. One writes down the functional equation, finds or characterizes a solution, and reads off the structure of optimal decisions from that solution. This is legitimate only if two things hold: an optimal policy exists at all, and the solution of the functional equation that was found is the optimal value, not some other solution of the same equation.

Samuel Karlin's 1955 paper The Structure of Dynamic Programing Models (Naval Research Logistics Quarterly 2(4):285–294) gives an abstract deterministic model in which both questions can be posed precisely. It proves existence of optimal strategies by a compactness argument (Theorem 1), derives the functional equation, which it calls the Principle of Optimality, and identifies the condition under which a solution of that equation is the optimal return: a tail term must vanish. Later treatments of dynamic programming on general state spaces, such as Blackwell's discounted and positive programming (1965–1967) and the monographs of Bertsekas and Shreve, state their verification theorems in the same form, with a solution of the optimality equation plus a condition at infinity.

Setting

The model has a state space Ω\OmegaΩ, a Hausdorff topological space, and a decision space DDD, a nonempty compact Hausdorff space. A strategy is a sequence s=(δ1,δ2,… )s = (\delta_1, \delta_2, \dots)s=(δ1​,δ2​,…) of decisions, one per stage. The strategy space S=D×D×⋯S = D \times D \times \cdotsS=D×D×⋯ carries the product topology and is compact by Tychonoff's theorem.

The data are:

  • a return function L:Ω×D→RL : \Omega \times D \to \mathbb{R}L:Ω×D→R, continuous and non-negative, where L(ω,δ)L(\omega, \delta)L(ω,δ) is the return for taking decision δ\deltaδ in state ω\omegaω;
  • a transition (δ,ω)↦Tδ ω∈Ω(\delta, \omega) \mapsto T_\delta\,\omega \in \Omega(δ,ω)↦Tδ​ω∈Ω, the state faced at the next stage after decision δ\deltaδ in state ω\omegaω;
  • a normalization factor P:D→RP : D \to \mathbb{R}P:D→R, continuous and positive.

From an initial state ω\omegaω, a strategy sss generates the trajectory ω1=ω\omega_1 = \omegaω1​=ω, ωn=Tδn−1 ωn−1\omega_n = T_{\delta_{n-1}}\,\omega_{n-1}ωn​=Tδn−1​​ωn−1​, and the weights Pn(s)=∏i=1n−1P(δi)P_n(s) = \prod_{i=1}^{n-1} P(\delta_i)Pn​(s)=∏i=1n−1​P(δi​) with P1(s)=1P_1(s) = 1P1​(s)=1. The total yield is

Φ(ω,s)=∑n=1∞L(ωn,δn) Pn(s),\Phi(\omega, s) = \sum_{n=1}^{\infty} L(\omega_n, \delta_n)\, P_n(s),Φ(ω,s)=n=1∑∞​L(ωn​,δn​)Pn​(s),

and the optimal return is K(ω)=max⁡s∈SΦ(ω,s)K(\omega) = \max_{s \in S} \Phi(\omega, s)K(ω)=maxs∈S​Φ(ω,s). The standing assumption of the paper, display (1), is that the partial sums ∑n=1kL(ωn,δn)Pn(s)\sum_{n=1}^{k} L(\omega_n, \delta_n) P_n(s)∑n=1k​L(ωn​,δn​)Pn​(s) converge uniformly in s∈Ss \in Ss∈S for each ω\omegaω.

Formalization targets

Goal: uniqueness of solutions with vanishing tail

Let M:Ω→RM : \Omega \to \mathbb{R}M:Ω→R solve the functional equation

M(ω)=max⁡δ∈D{L(ω,δ)+P(δ) M(Tδ ω)}for all ω,M(\omega) = \max_{\delta \in D} \bigl\{ L(\omega, \delta) + P(\delta)\, M(T_\delta\,\omega) \bigr\} \quad \text{for all } \omega,M(ω)=δ∈Dmax​{L(ω,δ)+P(δ)M(Tδ​ω)}for all ω,

with the maximum attained, and suppose that for every ω\omegaω

lim⁡n→∞ sup⁡δ1,…,δn∣M(ωn)∣∏i=1n−1P(δi)=0.\lim_{n \to \infty} \ \sup_{\delta_1, \dots, \delta_n} |M(\omega_n)| \prod_{i=1}^{n-1} P(\delta_i) = 0.n→∞lim​ δ1​,…,δn​sup​∣M(ωn​)∣i=1∏n−1​P(δi​)=0.

Then M(ω)=max⁡s∈SΦ(ω,s)M(\omega) = \max_{s \in S} \Phi(\omega, s)M(ω)=maxs∈S​Φ(ω,s) for every ω\omegaω, and the maximum is attained (pp. 290–291, §Uniqueness).

Milestones

  1. Theorem 1 (p. 287). If the series (1) converges uniformly in SSS, an optimal strategy s∗s^*s∗ exists: Φ(ω,s∗)=max⁡SΦ(ω,s)\Phi(\omega, s^*) = \max_S \Phi(\omega, s)Φ(ω,s∗)=maxS​Φ(ω,s).
  2. Shift identity (p. 290, first display). For a strategy sss with convergent yield series and the shift s′=(δ2,δ3,… )s' = (\delta_2, \delta_3, \dots)s′=(δ2​,δ3​,…),
Φ(ω,s)=L(ω,δ1)+P(δ1) Φ(Tδ1 ω,s′).\Phi(\omega, s) = L(\omega, \delta_1) + P(\delta_1)\, \Phi(T_{\delta_1}\,\omega, s').Φ(ω,s)=L(ω,δ1​)+P(δ1​)Φ(Tδ1​​ω,s′).
  1. Principle of Optimality, eq. (2) (p. 290). K(ω)=max⁡δ1{L(ω,δ1)+P(δ1)K(Tδ1 ω)}K(\omega) = \max_{\delta_1} \{ L(\omega, \delta_1) + P(\delta_1) K(T_{\delta_1}\,\omega) \}K(ω)=maxδ1​​{L(ω,δ1​)+P(δ1​)K(Tδ1​​ω)}.
  2. n-step expansion (p. 291, first display). A solution MMM of (2) satisfies, for every nnn,
M(ω)=max⁡δ1,…,δn{∑m=1nL(ωm,δm)Pm(s)+M(ωn+1)Pn+1(s)}.M(\omega) = \max_{\delta_1, \dots, \delta_n} \Bigl\{ \sum_{m=1}^{n} L(\omega_m, \delta_m) P_m(s) + M(\omega_{n+1}) P_{n+1}(s) \Bigr\}.M(ω)=δ1​,…,δn​max​{m=1∑n​L(ωm​,δm​)Pm​(s)+M(ωn+1​)Pn+1​(s)}.

Significance

Milestone 3 says that the optimal return solves the functional equation. The goal gives the converse on a class of candidate solutions: any solution with a vanishing tail is the optimal return. Together they justify solving a dynamic program by solving its functional equation. The paper's two examples, a two-operation allocation problem with discounting and a resource allocation model driving the state to the origin, obtain uniqueness among bounded solutions and among continuous solutions vanishing at the origin respectively, by checking the tail condition. Without the tail condition the conclusion fails; the paper notes that the limit term "need not be true in general for any solution to the functional equation".

All four milestones and the goal are classical results with published proofs. None of them has a machine-checked proof on this platform: its existing Bellman-equation theorems concern finite-state stochastic models with a constant discount factor, and this model has neither restriction. The mission produces a formal version of the general deterministic model on topological state spaces, with optimality characterized by a verification theorem, which later missions on the paper's examples can reuse.

Difficulty

Existence rests on continuity of s↦Φ(ω,s)s \mapsto \Phi(\omega, s)s↦Φ(ω,s) on the product space. Each term L(ωn,δn)Pn(s)L(\omega_n, \delta_n) P_n(s)L(ωn​,δn​)Pn​(s) depends on the first nnn decisions through the composite map ωn=Tδn−1∘⋯∘Tδ1 ω\omega_n = T_{\delta_{n-1}} \circ \cdots \circ T_{\delta_1}\,\omegaωn​=Tδn−1​​∘⋯∘Tδ1​​ω, and continuity of that composite in all decisions at once does not follow from separate continuity of Tδ ωT_\delta\,\omegaTδ​ω in δ\deltaδ and in ω\omegaω. The limit of the series is continuous only because the convergence is uniform.

For uniqueness, the paper's display "M(ω)=max⁡SΦ(ω,s)+lim⁡nmax⁡M(ωn)∏P(δi)M(\omega) = \max_S \Phi(\omega, s) + \lim_n \max M(\omega_n) \prod P(\delta_i)M(ω)=maxS​Φ(ω,s)+limn​maxM(ωn​)∏P(δi​)" is not an identity: a maximum of a sum is not the sum of the maxima. A proof has to bound MMM from above by the yield of every strategy and from below by the yield of one particular strategy, and the lower bound fails if the tail term is controlled only from above. Attainment of the maximum in the conclusion needs a strategy to be exhibited, not only a supremum computed.

Formalization scope

A strategy is a function s : ℕ → D, with the product topology. Lean's 0-based index kkk is the paper's stage k+1k+1k+1: s 0 is δ1\delta_1δ1​, trajectory T ω s 0 is ω1=ω\omega_1 = \omegaω1​=ω, and weight P s k is Pk+1(s)P_{k+1}(s)Pk+1​(s), so weight P s 0 = 1. The total yield is a tsum and the optimal return a supremum over all strategies. "Maximum" is encoded as IsGreatest of a range, so every stated maximum is attained. Uniform convergence of (1) is TendstoUniformly of the partial sums to Φ(ω,⋅)\Phi(\omega, \cdot)Φ(ω,⋅) along atTop.

The formalization commits to the following, relative to the page:

  • The model's assumptions (1)–(3) of p. 286, non-negativity of LLL, and positivity and continuity of PPP appear as hypotheses of every statement.
  • A single return function LLL is used, not stage-dependent LnL_nLn​, as in display (1) and as the functional equation (2) requires. PnP_nPn​ has the product form, which the paper adopts "unless stated to the contrary".
  • Assumption (4), separate continuity of Tδ ωT_\delta\,\omegaTδ​ω in δ\deltaδ and in ω\omegaω, is strengthened to joint continuity of (δ,ω)↦Tδ ω(\delta, \omega) \mapsto T_\delta\,\omega(δ,ω)↦Tδ​ω. This supports the paper's assertion (p. 287) that each term is a continuous function of sss, which separate continuity does not give.
  • DDD is assumed nonempty. With DDD empty there is no strategy and Theorem 1 is false.
  • The paper leaves the class of admissible solutions open ("an appropriate class of M's for which the lim = 0"). The goal fixes it as the two-sided condition: for every ω\omegaω and ε>0\varepsilon > 0ε>0 there is NNN with ∣M(ωn)∣ Pn(s)≤ε|M(\omega_n)|\,P_n(s) \le \varepsilon∣M(ωn​)∣Pn​(s)≤ε for all n≥Nn \ge Nn≥N and all sss. Both of the paper's examples verify this form.

Lean's tsum of a non-summable series is 000, and a supremum of an unbounded family is 000. Neither default can make a statement trivially true. The uniform-convergence hypothesis forces the series to converge, and under Theorem 1's hypotheses Φ(ω,⋅)\Phi(\omega, \cdot)Φ(ω,⋅) is continuous on a compact space, so the supremum is a maximum. The goal's conclusion is stated without a supremum. A sorry-free check confirms that all hypotheses of the goal hold on a concrete instance with non-zero return: two decisions, constant return 111, P≡1/2P \equiv 1/2P≡1/2, and M≡2M \equiv 2M≡2.

Needed infrastructure: Tychonoff's theorem, continuity of uniform limits, and attainment of maxima on compact spaces, all in Mathlib. The mission's own definitions are the trajectory, weights, partial and total yield, and the optimal return. Contributions of intermediate lemmas are welcome, for example continuity of s↦ωns \mapsto \omega_ns↦ωn​, summability from uniform convergence, and the upper and lower tail estimates for MMM. So are formalizations of the paper's Remarks 1 and 3 (Dini's theorem and convergence of the kkk-stage optimal returns).

Selected references

  • S. Karlin, The Structure of Dynamic Programing Models, Naval Research Logistics Quarterly 2(4):285–294, 1955. https://doi.org/10.1002/nav.3800020408
  • R. Bellman, Dynamic Programming, Princeton University Press, 1957. https://press.princeton.edu/books/paperback/9780691146683/dynamic-programming
  • D. Blackwell, Discounted Dynamic Programming, Annals of Mathematical Statistics 36(1):226–235, 1965. https://doi.org/10.1214/aoms/1177700285
  • D. P. Bertsekas and S. E. Shreve, Stochastic Optimal Control: The Discrete-Time Case, Academic Press, 1978. https://web.mit.edu/dimitrib/www/soc.html
6 thms3 active usersReviewed
Discrete GeometryNumber Theory·Captain: mikedeng1

Minkowski's Convex Body Theorem and Integer Programming: Lattice-Free Convex Bodies Meet Few Translates of an Integral SubspaceResearch Paper

Motivation

Integer programming asks whether a system of linear inequalities Ax≤bAx\le bAx≤b has a solution x∈Znx\in\mathbb Z^nx∈Zn. In fixed dimension nnn it is solvable in polynomial time: Lenstra (1983) proved this by showing that a convex body without integer points is "flat" in some integral direction, so that the search splits into few lower-dimensional subproblems. Kannan's 1987 paper in Mathematics of Operations Research sharpened this approach. It computes a Korkine–Zolotarev ("reduced") basis of a lattice, solves the shortest and closest vector problems exactly in nO(n)n^{O(n)}nO(n) operations, and runs integer programming in O(n9n/2s)O(n^{9n/2}s)O(n9n/2s) arithmetic operations. Underneath the algorithm sits a purely geometric statement, Theorem (5.5): a lattice-free convex body meets only boundedly many integer translates of some integral subspace.

Timeline.

  • Korkine and Zolotareff (1873): the reduced bases used here.
  • Minkowski (1896): a symmetric convex body of volume greater than 2n2^n2n contains a nonzero integer point.
  • Khinchine (1948): lattice-free convex bodies have lattice width bounded by a function of nnn alone (the flatness theorem).
  • Lenstra (1983): integer programming in fixed dimension is polynomial, via a flat direction.
  • Kannan (1987, this paper): Theorem (5.5), with subspaces VVV of any dimension between 111 and n−1n-1n−1 and an explicit bound n2(n−dim⁡V)n^{2(n-\dim V)}n2(n−dimV).
  • Kannan and Lovász (1988), Banaszczyk et al. (1999), and later work: polynomial bounds on the flatness constant.

Setting

Rn\mathcal R^nRn is Euclidean space with dot product (a,b)(a,b)(a,b) and length ∣a∣|a|∣a∣, and Zn\mathbb Z^nZn is the set of integer vectors. For linearly independent b1,…,bm∈Rkb_1,\dots,b_m\in\mathcal R^kb1​,…,bm​∈Rk, the lattice L(b1,…,bm)L(b_1,\dots,b_m)L(b1​,…,bm​) is the set of integer combinations ∑jλjbj\sum_j\lambda_jb_j∑j​λj​bj​, λj∈Z\lambda_j\in\mathbb Zλj​∈Z, and b1,…,bmb_1,\dots,b_mb1​,…,bm​ is a basis. Gram–Schmidt orthogonalisation gives b1∗,…,bm∗b_1^*,\dots,b_m^*b1∗​,…,bm∗​ and unit vectors uj=bj∗/∣bj∗∣u_j=b_j^*/|b_j^*|uj​=bj∗​/∣bj∗​∣, and bi(j)=(bi,uj)b_i(j)=(b_i,u_j)bi​(j)=(bi​,uj​), so bi=∑jbi(j)ujb_i=\sum_jb_i(j)u_jbi​=∑j​bi​(j)uj​ and bj(j)=∣bj∗∣b_j(j)=|b_j^*|bj​(j)=∣bj∗​∣. The determinant is d(L)=∏j∣bj∗∣d(L)=\prod_j|b_j^*|d(L)=∏j​∣bj∗​∣. Λ1(L)\Lambda_1(L)Λ1​(L) is the length of a shortest nonzero vector of LLL. The projected lattice Lj(b1,…,bm)L_j(b_1,\dots,b_m)Lj​(b1​,…,bm​) is the image of LLL under orthogonal projection onto the complement of span⁡(b1,…,bj−1)\operatorname{span}(b_1,\dots,b_{j-1})span(b1​,…,bj−1​). A basis is reduced (Definition 2.6) if bj(j)=Λ1(Lj)b_j(j)=\Lambda_1(L_j)bj​(j)=Λ1​(Lj​) for every jjj and ∣bi(j)∣≤bj(j)/2|b_i(j)|\le b_j(j)/2∣bi​(j)∣≤bj​(j)/2 for i>ji>ji>j.

A convex body is a convex set of positive volume, which for a convex set means nonempty interior. A subspace VVV has a basis of integer vectors if it is the real span of integer vectors. Its integer translates are the sets z+Vz+Vz+V with z∈Znz\in\mathbb Z^nz∈Zn.

In Lean, Rk\mathcal R^kRk is EuclideanSpace ℝ (Fin k), a basis is b : Fin m → EuclideanSpace ℝ (Fin k), the lattice is lattice b = Submodule.span ℤ (Set.range b), ∣bj∗∣|b_j^*|∣bj∗​∣ is gsLen b j, bi(j)b_i(j)bi​(j) is gsCoeff b i j, d(L)d(L)d(L) is latticeDet b, Λ1\Lambda_1Λ1​ is lambdaOne, Lj+1L_{j+1}Lj+1​ is projLattice b j, and a reduced basis is IsReduced b.

Formalization targets

Goal: Theorem (5.5), corrected reading

For n≥2n\ge2n≥2 and every bounded convex set K⊆RnK\subseteq\mathcal R^nK⊆Rn with nonempty interior and K∩Zn=∅K\cap\mathbb Z^n=\emptysetK∩Zn=∅ there is a subspace VVV spanned by integer vectors with 1≤dim⁡V≤n−11\le\dim V\le n-11≤dimV≤n−1 and

#{ z+V:z∈Zn, (z+V)∩K≠∅ } ≤ n2(n−dim⁡V).\#\{\,z+V : z\in\mathbb Z^n,\ (z+V)\cap K\ne\emptyset\,\}\ \le\ n^{2(n-\dim V)}.#{z+V:z∈Zn, (z+V)∩K=∅} ≤ n2(n−dimV).

The printed theorem allows "an iii dimensional space VVV" with 1≤i≤n1\le i\le n1≤i≤n and bound n2(n−i+1)n^{2(n-i+1)}n2(n−i+1). Taken literally that is trivial (V=RnV=\mathcal R^nV=Rn, one translate). The proof on the same page takes V=span⁡(b1,…,bi−1)V=\operatorname{span}(b_1,\dots,b_{i-1})V=span(b1​,…,bi−1​), of dimension i−1i-1i−1, and remarks that this "ensures that the subspace VVV is always of dimension at least 1". The goal states that reading.

Milestones

  1. Theorem (1.11), Minkowski's convex body theorem (referenced from the platform, in Mathlib's general form).
  2. Theorem (1.12): every mmm-dimensional lattice has a nonzero vector with ∣v∣≤m d(L)1/m|v|\le\sqrt m\,d(L)^{1/m}∣v∣≤m​d(L)1/m.
  3. Proposition 1.9: a primitive lattice vector belongs to some basis.
  4. Proposition 2.16, existence form: every lattice has a reduced basis.
  5. Proposition 4.2: for any b0b_0b0​ with projection bˉ0\bar b_0bˉ0​ onto the span, some b∈Lb\in Lb∈L has ∣b−bˉ0∣≤12(∑jbj(j)2)1/2≤m2max⁡jbj(j)|b-\bar b_0|\le\frac12(\sum_jb_j(j)^2)^{1/2}\le\frac{\sqrt m}2\max_jb_j(j)∣b−bˉ0​∣≤21​(∑j​bj​(j)2)1/2≤2m​​maxj​bj​(j).
  6. Proposition 4.3: for a reduced basis and iii maximising bi(i)b_i(i)bi​(i), the tail (λi,…,λm)(\lambda_i,\dots,\lambda_m)(λi​,…,λm​) of every closest lattice point to b0b_0b0​ lies in an explicit set of at most mm−i+1m^{m-i+1}mm−i+1 integer vectors.

Significance

Theorem (5.5) is a structural form of the flatness theorem. For dim⁡V=n−1\dim V=n-1dimV=n−1 it says that a lattice-free convex body meets fewer than n2n^2n2 consecutive integer hyperplanes of some integral direction. For smaller dim⁡V\dim VdimV it gives a finer decomposition of Zn\mathbb Z^nZn into translates, each a lower-dimensional integer program. This is the recursion behind fixed-dimension integer programming, and statements of this form are used in lattice-point enumeration, in the geometry of numbers (covering minima), and in cutting-plane theory (lattice-free bodies define split and intersection cuts). Propositions 4.2 and 4.3 are the correctness core of exact closest-vector enumeration.

The results are proved in the literature, though Theorem (5.5) is proved "albeit sketchily" in the paper itself. As far as is known, none of them is formalized: Mathlib has Minkowski's convex body theorem and the ZLattice API, but not Gram–Schmidt lattice invariants, Korkine–Zolotarev bases, Hermite-type bounds, nearest-plane rounding, or any flatness theorem. This mission produces the first machine-checked versions. It also corrects three statements that are wrong as printed (below), so the formal statements are the ones that can be relied on.

Difficulty

The naive route to (5.5) is to take a flat direction directly: bound the lattice width of KKK and count hyperplanes. That needs a flatness theorem with an explicit bound below n2n^2n2, which is itself the hard part. The paper's argument instead needs John's theorem (every convex body lies between an ellipsoid and its nnn-fold dilation), a reduced basis of the transformed lattice, and a counting argument across projected lattices that combines Minkowski's bound on each LiL_iLi​ with the covering estimate of Proposition 4.2. None of John's theorem, reduced bases or the projected-lattice counting is in Mathlib.

Proposition 4.3 is also delicate as printed: the per-coordinate count on p. 24 undercounts the integers in a closed interval, so the printed arithmetic cannot be transcribed as it stands. Proposition 2.16 in the paper is the correctness of the algorithm SHORTEST. Here only the existence of a reduced basis is needed, which requires attainment of Λ1\Lambda_1Λ1​ on every projected lattice and a lifting argument (Proposition 1.9).

Formalization scope

Conventions: indices are 0-based (Fin m), so the paper's LjL_jLj​ is projLattice b (j-1) and its bound nn−i+1n^{n-i+1}nn−i+1 is m ^ (m - i). Gram–Schmidt is Mathlib's unnormalised gramSchmidt. Lattices are Submodule ℤs of a real Euclidean space generated by a linearly independent family, and m≤km\le km≤k is allowed, because (1.12) and 4.2 are applied to projected lattices. The goal counts translates as sets with Set.encard, so the bound includes finiteness. KKK is assumed convex, bounded and with nonempty interior, but not closed.

Three printed statements are corrected, and the corrections are recorded in each item's Formalization Note.

  • (1.12)'s constant 12n\frac12\sqrt n21​n​ is false for n≤7n\le7n≤7 (for example L=ZL=\mathbb ZL=Z, or the hexagonal lattice) and is replaced by n\sqrt nn​, the constant the paper's own later proofs use.
  • Proposition 4.2's second sentence is stated for bˉ0\bar b_0bˉ0​ instead of b0b_0b0​.
  • Proposition 4.3 fails at n=1n=1n=1 and is stated for m≥2m\ge2m≥2 with the proof's explicit candidate set TTT, since an existential TTT is satisfied by the set of tails of closest points and says nothing.

Trivializing formalizations are ruled out: the goal forbids dim⁡V=n\dim V=ndimV=n, which gives one translate, and dim⁡V=0\dim V=0dimV=0, where no translate meets KKK. It requires nonempty interior (the empty set would satisfy everything) and counts with encard (an infinite count cannot become 000).

Out of scope: the paper's algorithms (SHORTEST, SELECT-BASIS, ENUMERATE, CLP, CLP′, ILP) and their operation and bit counts (Theorems 2.17, 3.9, 4.5, 5.4), because Mathlib has no cost model. Also out of scope is §6 (NP-completeness of the L2L_2L2​ closest vector problem and Cook reductions), because Mathlib has no complexity classes. The definitions of this mission (lattice, gsLen, gsCoeff, latticeDet, lambdaOne, projLattice, IsReduced) are reusable for any later work on lattice reduction. Contributions are welcome at every level: John's theorem, Hermite-type bounds, Korkine–Zolotarev existence, and the counting lemmas.

Selected references

  • R. Kannan, Minkowski's Convex Body Theorem and Integer Programming, Mathematics of Operations Research 12(3):415–440, 1987. https://doi.org/10.1287/moor.12.3.415
  • H. W. Lenstra Jr., Integer programming with a fixed number of variables, Mathematics of Operations Research 8(4):538–548, 1983. https://doi.org/10.1287/moor.8.4.538
  • R. Kannan, L. Lovász, Covering minima and lattice-point-free convex bodies, Annals of Mathematics 128(3):577–602, 1988. https://doi.org/10.2307/1971436
  • A. K. Lenstra, H. W. Lenstra Jr., L. Lovász, Factoring polynomials with rational coefficients, Mathematische Annalen 261:515–534, 1982. https://doi.org/10.1007/BF01457454
  • F. John, Extremum problems with inequalities as subsidiary conditions, Studies and Essays presented to R. Courant, 1948, 187–204.
9 thms3 active usersReviewed
🏆Completed
Convex OptimizationOptimization·Captain: mikedeng1

Robust Solutions to Uncertain Semidefinite Programs IV: Closed-Form Robust Counterparts under Unstructured PerturbationsResearch Paper

Motivation

A semidefinite program (SDP) minimizes a linear objective cTxc^TxcTx subject to a linear matrix inequality (LMI) F(x)=F0+∑i=1mxiFi⪰0F(x) = F_0 + \sum_{i=1}^m x_i F_i \succeq 0F(x)=F0​+∑i=1m​xi​Fi​⪰0. In applications the coefficient matrices FiF_iFi​ are measured, estimated or rounded. A solution that is feasible for the nominal data can become infeasible for data that differ from it by an arbitrarily small amount.

El Ghaoui, Oustry and Lebret (SIAM J. Optim. 9(1), 1998) introduced robust semidefinite programs (RSDPs): the constraint must hold for every admissible perturbation of the data, and the robust solution is the best point that survives all of them. Their §5 works out the examples in which the robust counterpart has a closed form. The simplest and most widely quoted is the case where every coefficient matrix is perturbed independently and without structure (§5.1): the robust LMI becomes the single convex constraint F(x)⪰2ρ∥x∥2+1 IF(x) \succeq 2\rho\sqrt{\|x\|^2+1}\,IF(x)⪰2ρ∥x∥2+1​I. The same computation gives closed-form robust versions of linear programs (§5.3), of largest-eigenvalue minimization (§5.4) and of matrix-norm minimization (§5.6), each of which is the nominal problem plus a Tikhonov-type term ρ∥x∥2+1\rho\sqrt{\|x\|^2+1}ρ∥x∥2+1​. Robust linear programming under ellipsoidal uncertainty was developed at the same time by Ben-Tal and Nemirovski (Math. Oper. Res., 1998); robust least squares, the prototype of §5.6, by El Ghaoui and Lebret (SIAM J. Matrix Anal. Appl., 1997).

Setting

Fix m,n∈Nm, n \in \mathbb{N}m,n∈N, a level ρ>0\rho > 0ρ>0, and symmetric matrices F0,…,Fm∈Rn×nF_0, \dots, F_m \in \mathbb{R}^{n\times n}F0​,…,Fm​∈Rn×n. For x∈Rmx \in \mathbb{R}^mx∈Rm write F(x)=F0+∑i=1mxiFiF(x) = F_0 + \sum_{i=1}^m x_i F_iF(x)=F0​+∑i=1m​xi​Fi​ and ∥x∥2=∑i=1mxi2\|x\|^2 = \sum_{i=1}^m x_i^2∥x∥2=∑i=1m​xi2​ (the Euclidean norm). For a matrix MMM, ∥M∥\|M\|∥M∥ is its spectral norm, the largest singular value, and X⪰0X \succeq 0X⪰0 means that XXX is symmetric positive semidefinite.

An unstructured perturbation is a block row Δ=[Δ0 ⋯ Δm]\Delta = [\Delta_0 \ \cdots \ \Delta_m]Δ=[Δ0​ ⋯ Δm​] of n×nn\times nn×n blocks, viewed as one n×n(m+1)n \times n(m+1)n×n(m+1) matrix. It perturbs each coefficient independently:

F(x,Δ)=F(x)+Δ0+Δ0T+∑i=1mxi(Δi+ΔiT).\mathbf{F}(x,\Delta) = F(x) + \Delta_0 + \Delta_0^T + \sum_{i=1}^m x_i(\Delta_i + \Delta_i^T).F(x,Δ)=F(x)+Δ0​+Δ0T​+i=1∑m​xi​(Δi​+ΔiT​).

The robust feasible set is

Xρ={x∈Rm:F(x,Δ)⪰0 for every Δ with ∥Δ∥≤ρ},\mathcal{X}_\rho = \{x \in \mathbb{R}^m : \mathbf{F}(x,\Delta) \succeq 0 \text{ for every } \Delta \text{ with } \|\Delta\| \le \rho\},Xρ​={x∈Rm:F(x,Δ)⪰0 for every Δ with ∥Δ∥≤ρ},

and the RSDP is: minimize cTxc^TxcTx over Xρ\mathcal{X}_\rhoXρ​. With R(x)=[1; x]⊗IR(x) = [1;\,x]\otimes IR(x)=[1;x]⊗I, the n(m+1)×nn(m+1)\times nn(m+1)×n matrix whose iii-th block is x~iI\tilde x_i Ix~i​I for x~=(1,x1,…,xm)\tilde x = (1, x_1, \dots, x_m)x~=(1,x1​,…,xm​), the perturbation reads F(x,Δ)=F(x)+ΔR(x)+R(x)TΔT\mathbf{F}(x,\Delta) = F(x) + \Delta R(x) + R(x)^T\Delta^TF(x,Δ)=F(x)+ΔR(x)+R(x)TΔT (the paper's (19)).

Three further models use the same pattern. In a robust LP, the data [aiT bi]T[a_i^T\ b_i]^T[aiT​ bi​]T of each constraint aiTx≥bia_i^Tx \ge b_iaiT​x≥bi​ are shifted by an independent δi∈Rm+1\delta_i \in \mathbb{R}^{m+1}δi​∈Rm+1 with ∥δi∥2≤ρ\|\delta_i\|_2 \le \rho∥δi​∥2​≤ρ. In robust eigenvalue minimization one minimizes the worst case over ∥Δ∥≤ρ\|\Delta\|\le\rho∥Δ∥≤ρ of λmax⁡(F(x,Δ))\lambda_{\max}(\mathbf{F}(x,\Delta))λmax​(F(x,Δ)). In robust maximum-norm minimization, H(x)=H0+∑ixiHiH(x) = H_0 + \sum_i x_i H_iH(x)=H0​+∑i​xi​Hi​ with Hi∈Rp×qH_i \in \mathbb{R}^{p\times q}Hi​∈Rp×q, H(x,Δ)=H0+Δ0+∑ixi(Hi+Δi)\mathbf{H}(x,\Delta) = H_0 + \Delta_0 + \sum_i x_i(H_i + \Delta_i)H(x,Δ)=H0​+Δ0​+∑i​xi​(Hi​+Δi​), and one minimizes max⁡∥Δ∥≤ρ∥H(x,Δ)∥\max_{\|\Delta\|\le\rho}\|\mathbf{H}(x,\Delta)\|max∥Δ∥≤ρ​∥H(x,Δ)∥.

Formalization targets

Goal: Theorem 5.1 (first sentence)

For every x∈Rmx \in \mathbb{R}^mx∈Rm,

x∈Xρ  ⟺  F(x)⪰2ρ∥x∥2+1  I.x \in \mathcal{X}_\rho \iff F(x) \succeq 2\rho\sqrt{\|x\|^2+1}\; I .x∈Xρ​⟺F(x)⪰2ρ∥x∥2+1​I.

The RSDP and problem (21), "minimize cTxc^TxcTx subject to F(x)⪰2ρ∥x∥2+1 IF(x) \succeq 2\rho\sqrt{\|x\|^2+1}\,IF(x)⪰2ρ∥x∥2+1​I", therefore have the same feasible set, optimal value and solutions. The goal fixes no numerical data: F0,…,FmF_0, \dots, F_mF0​,…,Fm​, mmm, nnn and ρ>0\rho > 0ρ>0 are arbitrary.

Milestones on the way (§5.1)

  1. (19)–(20): x∈Xρx \in \mathcal{X}_\rhox∈Xρ​ iff there is τ∈R\tau \in \mathbb{R}τ∈R with [F(x)−τIρR(x)TρR(x)τI]⪰0\begin{bmatrix} F(x) - \tau I & \rho R(x)^T \\ \rho R(x) & \tau I\end{bmatrix} \succeq 0[F(x)−τIρR(x)​ρR(x)TτI​]⪰0.
  2. Positivity of τ\tauτ and the Schur form (for n≥1n \ge 1n≥1): that block matrix is ⪰0\succeq 0⪰0 iff τ>0\tau > 0τ>0 and F(x)⪰(τ+ρ2(1+∥x∥2)/τ)IF(x) \succeq \bigl(\tau + \rho^2(1+\|x\|^2)/\tau\bigr) IF(x)⪰(τ+ρ2(1+∥x∥2)/τ)I.
  3. (21): some τ>0\tau > 0τ>0 satisfies the Schur form iff F(x)⪰2ρ∥x∥2+1 IF(x) \succeq 2\rho\sqrt{\|x\|^2+1}\, IF(x)⪰2ρ∥x∥2+1​I.

Further milestones: the value halves of Theorems 5.2–5.4

  • Theorem 5.2: the robust LP constraints hold iff aiTx−ρ∥x∥22+1≥bia_i^Tx - \rho\sqrt{\|x\|_2^2+1} \ge b_iaiT​x−ρ∥x∥22​+1​≥bi​ for all iii (problem (23)).
  • Theorem 5.3: for every ttt, tI⪰F(x,Δ)tI \succeq \mathbf{F}(x,\Delta)tI⪰F(x,Δ) for all ∥Δ∥≤ρ\|\Delta\| \le \rho∥Δ∥≤ρ iff (t−2ρ∥x∥2+1)I⪰F(x)\bigl(t - 2\rho\sqrt{\|x\|^2+1}\bigr) I \succeq F(x)(t−2ρ∥x∥2+1​)I⪰F(x); that is, the worst-case largest eigenvalue is λmax⁡(F(x))+2ρ∥x∥2+1\lambda_{\max}(F(x)) + 2\rho\sqrt{\|x\|^2+1}λmax​(F(x))+2ρ∥x∥2+1​ (problem (25)).
  • Theorem 5.4: for p,q≥1p, q \ge 1p,q≥1, max⁡∥Δ∥≤ρ∥H(x,Δ)∥=∥H(x)∥+ρ∥x∥2+1\max_{\|\Delta\|\le\rho}\|\mathbf{H}(x,\Delta)\| = \|H(x)\| + \rho\sqrt{\|x\|^2+1}max∥Δ∥≤ρ​∥H(x,Δ)∥=∥H(x)∥+ρ∥x∥2+1​, and the maximum is attained (problem (29)).

Significance

The goal shows that robustness against unstructured perturbations costs no more than the nominal problem: the robust counterpart is an LMI of the same size n×nn\times nn×n, with a right-hand side that is a convex function of xxx and grows like 2ρ∥x∥2\rho\|x\|2ρ∥x∥. The sets Xρ\mathcal{X}_\rhoXρ​ have no flat faces, which the paper's §5.2 uses to define the robust center of an LMI and which underlies the uniqueness and continuity of the robust solution (the second sentences of Theorems 5.1–5.4, from §4 under hypotheses H1–H3). Theorems 5.3 and 5.4 exhibit robustification as a Tikhonov regularization with parameter 2ρ2\rho2ρ or ρ\rhoρ, and Theorem 5.2 turns a robust LP into a second-order cone program.

All four closed forms are proved in the paper, partly by appeal to the general SDP reformulation of its §3. No machine-checked version of any of them exists, to our knowledge. The mission produces the robust counterparts as identities of feasible sets, stated for every xxx, together with the three intermediate steps of §5.1, so that later missions on the uniqueness and stability halves can import them.

Difficulty

The goal is an exchange of a universal quantifier over an infinite family of matrices with a single matrix inequality. The inequality F(x,Δ)⪰F(x)−2ρ∥x∥2+1 I\mathbf{F}(x,\Delta) \succeq F(x) - 2\rho\sqrt{\|x\|^2+1}\,IF(x,Δ)⪰F(x)−2ρ∥x∥2+1​I bounds each perturbation, but the converse needs, for each failing direction, one admissible perturbation that attains the bound; the constant 222 comes from the two copies ΔR(x)\Delta R(x)ΔR(x) and R(x)TΔTR(x)^T\Delta^TR(x)TΔT, and the constant ∥x∥2+1\sqrt{\|x\|^2+1}∥x∥2+1​ is the spectral norm of R(x)R(x)R(x), which holds only because Δ\DeltaΔ is normed as one block row. Normed block by block, the worst case and the constant change. In the milestone route, the positivity of τ\tauτ needs a separate argument before any Schur complement can be taken, since the Schur complement with respect to τI\tau IτI is undefined at τ=0\tau = 0τ=0, and the elimination of τ\tauτ needs the attainment of min⁡τ>0τ+a/τ\min_{\tau>0} \tau + a/\tauminτ>0​τ+a/τ. For Theorem 5.4 the difficulty is the attainment: an upper bound on the maximum is immediate, while the lower bound requires exhibiting an admissible perturbation that attains it.

Formalization scope

Matrices are Matrix (Fin r) (Fin c) ℝ. Coefficients are indexed by Fin (m + 1) with index 0 the constant term. A block row Δ\DeltaΔ is one matrix with columns indexed by pairs (i, b) : Fin (m + 1) × Fin n (or Fin q), and ∥Δ∥\|\Delta\|∥Δ∥ is Mathlib's ℓ2\ell^2ℓ2 operator norm (open scoped Matrix.Norms.L2Operator), the largest singular value, never the default entrywise norm. The vector norm ∥x∥2\|x\|^2∥x∥2 is written as ∑ixi2\sum_i x_i^2∑i​xi2​, never as Mathlib's sup norm on Fin m → ℝ. A⪰BA \succeq BA⪰B is (A - B).PosSemidef. Standing assumptions made explicit: F0,…,FmF_0, \dots, F_mF0​,…,Fm​ symmetric; ρ>0\rho > 0ρ>0 (§3, p. 36); n≥1n \ge 1n≥1 in milestone 2 (at n=0n = 0n=0 every τ\tauτ is feasible); p,q≥1p, q \ge 1p,q≥1 in Theorem 5.4 (empty matrices have norm 000).

Readings and corrections of the printed text:

  1. "The optimal value of the RSDP can be computed by solving (21)" is stated as the identity of the two feasible sets for every xxx, which implies equality of values and of solutions. Theorems 5.2 and 5.4 are stated the same way (5.4 through the pointwise worst-case value, with attainment), and Theorem 5.3 in epigraph form, λmax⁡(M)≤t  ⟺  tI−M⪰0\lambda_{\max}(M) \le t \iff tI - M \succeq 0λmax​(M)≤t⟺tI−M⪰0.
  2. Only the first sentence of each theorem is in scope. Uniqueness, regularity, Lipschitz stability and the limit ρ→0\rho \to 0ρ→0 rest on Theorem 4.3 and on external results ([31], [3]) and are not stated.
  3. In (19) the paper writes D=Rn×nm\mathcal D = \mathbb R^{n\times nm}D=Rn×nm and "the representation in section 5"; Δ\DeltaΔ has m+1m+1m+1 blocks, so D=Rn×n(m+1)\mathcal D = \mathbb R^{n\times n(m+1)}D=Rn×n(m+1), and the representation is that of §2.2.
  4. The paper derives (20) from Lemma 3.2 and (29) from Theorem 3.2, which give only sufficient conditions; the exact equivalences are the full-perturbation Lemma 3.1 / Theorem 3.1.
  5. Before (21) the paper says "the scalar in the left-hand side" (it is on the right) and "the RSDP (1)" (it means the RSDP (4)). Theorem 5.3's "min-max problem (24)" is the robust version of the nominal problem (24).

A formalization in which ∥Δ∥\|\Delta\|∥Δ∥ is an entrywise or blockwise norm, ∥x∥\|x\|∥x∥ is the sup norm, or the robust set quantifies over a single block, changes the constant 2ρ∥x∥2+12\rho\sqrt{\|x\|^2+1}2ρ∥x∥2+1​ and is not this theorem; the statements here rule these out by construction.

Useful, reusable infrastructure: the spectral norm of [1; x]⊗I[1;\,x] \otimes I[1;x]⊗I, Schur complements for positive semidefinite block matrices, and spectral norms of rank-one matrices. Proofs of the three §5.1 milestones and direct proofs of the goal are both welcome.

Selected references

  • L. El Ghaoui, F. Oustry, H. Lebret, Robust Solutions to Uncertain Semidefinite Programs, SIAM J. Optim. 9(1):33–52, 1998. https://doi.org/10.1137/S1052623496305717
  • L. El Ghaoui, H. Lebret, Robust Solutions to Least-Squares Problems with Uncertain Data, SIAM J. Matrix Anal. Appl. 18(4):1035–1064, 1997. https://doi.org/10.1137/S0895479896298130
  • A. Ben-Tal, A. Nemirovski, Robust Convex Optimization, Math. Oper. Res. 23(4):769–805, 1998. https://doi.org/10.1287/moor.23.4.769
8 thms3 active usersReviewed
🏆Completed
Control TheoryConvex OptimizationOptimization·Captain: mikedeng1

Robust Solutions to Uncertain Semidefinite Programs I: Exact SDP Reformulation of the Robust LMI under Full Linear-Fractional PerturbationsResearch Paper

Motivation

A semidefinite program (SDP) minimizes a linear objective cTxc^TxcTx subject to a linear matrix inequality (LMI) F(x)=F0+∑i=1mxiFi⪰0F(x) = F_0 + \sum_{i=1}^m x_iF_i \succeq 0F(x)=F0​+∑i=1m​xi​Fi​⪰0. SDPs model problems in control, combinatorial optimization, statistics and engineering design, and they are solved efficiently by interior-point methods. In applications the data F0,…,FmF_0,\dots,F_mF0​,…,Fm​ are rarely known exactly: they come from measurements, from linearized models, or from rounding. A solution that is optimal for the nominal data may violate the constraint for data that differ only slightly.

El Ghaoui, Oustry and Lebret (SIAM J. Optim. 9(1), 1998) asked for robust solutions: points xxx that satisfy the constraint for every admissible value of an unknown but bounded perturbation, and among them one that minimizes cTxc^TxcTx. Their paper, together with the contemporaneous work of Ben-Tal and Nemirovski on robust convex optimization (Math. Oper. Res. 23(4), 1998), founded robust semidefinite programming. The perturbation model they use, the linear-fractional representation (LFR), is the standard uncertainty model of robust control, where the same exact reformulation appears as the multiplier characterization of quadratic stability under norm-bounded uncertainty.

This mission formalizes the first main result of the paper: when the perturbation is full (an arbitrary matrix of bounded spectral norm), the robust problem is exactly an SDP with one extra scalar variable.

Setting

Fix natural numbers m,n,p,qm, n, p, qm,n,p,q and a decision vector x∈Rmx \in \mathbb{R}^mx∈Rm. The data are:

  • symmetric matrices F0,…,Fm∈Rn×nF_0,\dots,F_m \in \mathbb{R}^{n\times n}F0​,…,Fm​∈Rn×n, defining the affine map F(x)=F0+∑ixiFiF(x) = F_0 + \sum_i x_iF_iF(x)=F0​+∑i​xi​Fi​;
  • matrices R0,…,Rm∈Rq×nR_0,\dots,R_m \in \mathbb{R}^{q\times n}R0​,…,Rm​∈Rq×n, defining R(x)=R0+∑ixiRiR(x) = R_0 + \sum_i x_iR_iR(x)=R0​+∑i​xi​Ri​;
  • fixed matrices L∈Rn×pL \in \mathbb{R}^{n\times p}L∈Rn×p and D∈Rq×pD \in \mathbb{R}^{q\times p}D∈Rq×p;
  • a level ρ>0\rho > 0ρ>0.

For a matrix XXX, ∥X∥\|X\|∥X∥ denotes its largest singular value (the spectral norm), and X⪰0X \succeq 0X⪰0 means that XXX is symmetric positive semidefinite. A perturbation is a matrix Δ∈Rp×q\Delta \in \mathbb{R}^{p\times q}Δ∈Rp×q. The perturbed constraint matrix is the LFR (5)

F(x,Δ)=F(x)+LΔ(I−DΔ)−1R(x)+R(x)T(I−ΔTDT)−1ΔTLT,\mathbf{F}(x,\Delta) = F(x) + L\Delta(I - D\Delta)^{-1}R(x) + R(x)^T(I - \Delta^TD^T)^{-1}\Delta^TL^T,F(x,Δ)=F(x)+LΔ(I−DΔ)−1R(x)+R(x)T(I−ΔTDT)−1ΔTLT,

which is well defined exactly when det⁡(I−DΔ)≠0\det(I - D\Delta) \neq 0det(I−DΔ)=0. For a linear subspace D\mathcal{D}D of Rp×q\mathbb{R}^{p\times q}Rp×q, the robust feasible set (2) is

Xρ={x∈Rm:for every Δ∈D with ∥Δ∥≤ρ, F(x,Δ) is well defined and F(x,Δ)⪰0},\mathcal{X}_\rho = \bigl\{x \in \mathbb{R}^m : \text{for every } \Delta \in \mathcal{D} \text{ with } \|\Delta\| \le \rho,\ \mathbf{F}(x,\Delta) \text{ is well defined and } \mathbf{F}(x,\Delta) \succeq 0\bigr\},Xρ​={x∈Rm:for every Δ∈D with ∥Δ∥≤ρ, F(x,Δ) is well defined and F(x,Δ)⪰0},

and the robust SDP (4) is: minimize cTxc^TxcTx subject to x∈Xρx \in \mathcal{X}_\rhox∈Xρ​, for a given c∈Rm∖{0}c \in \mathbb{R}^m \setminus \{0\}c∈Rm∖{0}. In this mission D=Rp×q\mathcal{D} = \mathbb{R}^{p\times q}D=Rp×q, the full perturbation case, and the paper's standing assumption of §3.1 is ∥D∥<ρ−1\|D\| < \rho^{-1}∥D∥<ρ−1.

Formalization targets

Goal: Theorem 3.1 (p. 36), as a set identity

Under ρ>0\rho > 0ρ>0, ∥D∥<ρ−1\|D\| < \rho^{-1}∥D∥<ρ−1, q≥1q \ge 1q≥1 and L≠0L \ne 0L=0, for every x∈Rmx \in \mathbb{R}^mx∈Rm,

x∈Xρ  ⟺  ∃ τ∈R: [F(x)−τLLTR(x)T−τLDTR(x)−τDLTτ(ρ−2I−DDT)]⪰0.(10)x \in \mathcal{X}_\rho \iff \exists\,\tau \in \mathbb{R}:\ \begin{bmatrix} F(x) - \tau LL^T & R(x)^T - \tau LD^T \\ R(x) - \tau DL^T & \tau(\rho^{-2}I - DD^T)\end{bmatrix} \succeq 0. \qquad (10)x∈Xρ​⟺∃τ∈R: [F(x)−τLLTR(x)−τDLT​R(x)T−τLDTτ(ρ−2I−DDT)​]⪰0.(10)

The paper states that the robust SDP and a corresponding solution can be computed by solving the SDP "minimize cTxc^TxcTx subject to (10)" in the variables (x,τ)(x, \tau)(x,τ). Both problems have the objective cTxc^TxcTx, so the identity above, between Xρ\mathcal{X}_\rhoXρ​ and the xxx-projection of the feasible set of (10), is the content of that sentence. A companion item states the solution correspondence explicitly: xxx is optimal for the robust SDP if and only if (x,τ)(x,\tau)(x,τ) is optimal for (10) for some τ\tauτ.

Milestones

  1. Well-posedness (§3.1, p. 36). For ρ>0\rho > 0ρ>0: det⁡(I−DΔ)≠0\det(I - D\Delta) \ne 0det(I−DΔ)=0 for every Δ\DeltaΔ with ∥Δ∥≤ρ\|\Delta\| \le \rho∥Δ∥≤ρ if and only if ∥D∥<ρ−1\|D\| < \rho^{-1}∥D∥<ρ−1.
  2. Lemma 3.1 (p. 36). For F=FTF = F^TF=FT, q≥1q \ge 1q≥1 and L≠0L \ne 0L=0: det⁡(I−DΔ)≠0\det(I - D\Delta) \ne 0det(I−DΔ)=0 and F+LΔ(I−DΔ)−1R+RT(I−DΔ)−TΔTLT⪰0F + L\Delta(I - D\Delta)^{-1}R + R^T(I - D\Delta)^{-T}\Delta^TL^T \succeq 0F+LΔ(I−DΔ)−1R+RT(I−DΔ)−TΔTLT⪰0 for every ∥Δ∥≤1\|\Delta\| \le 1∥Δ∥≤1 if and only if ∥D∥<1\|D\| < 1∥D∥<1 and some scalar τ\tauτ satisfies
[F−τLLTRT−τLDTR−τDLTτ(I−DDT)]⪰0.\begin{bmatrix} F - \tau LL^T & R^T - \tau LD^T \\ R - \tau DL^T & \tau(I - DD^T)\end{bmatrix} \succeq 0.[F−τLLTR−τDLT​RT−τLDTτ(I−DDT)​]⪰0.

The paper cites the S-procedure as the classical result behind Lemma 3.1; it is already proved on the platform (ConvexOptimization.s_procedure) and is included as a reference item.

Significance

The robust feasible set is defined by infinitely many matrix inequalities, one per perturbation, each rational in Δ\DeltaΔ; in general such a set is convex but has no tractable description, and the paper notes that the structured version of the problem is NP-hard. Theorem 3.1 shows that for full perturbations nothing is lost by replacing that semi-infinite constraint with a single LMI of size n+qn + qn+q in one extra variable. Consequences: the robust problem is solved by a standard SDP solver; the largest admissible perturbation level is a generalized eigenvalue problem; and the exact result is the benchmark against which the paper's sufficient conditions for structured perturbations (Theorem 3.2) and its closed-form counterparts for unstructured perturbations (Theorem 5.1) are measured.

The result is proved in the paper (from the S-procedure, with the details deferred to a cited report). To the best of available knowledge it has no machine-checked proof. The mission produces a formal statement of the LFR model and of the robust feasible set that later missions on robust SDPs can reuse, a formal proof of the well-posedness condition, and a formal proof of the exact reformulation built on the platform's S-procedure. Formalizing it also records two points the printed statement leaves implicit: the result needs L≠0L \ne 0L=0 and a nonempty perturbation output dimension q≥1q \ge 1q≥1.

Difficulty

The direction from the LMI to robust feasibility is elementary. The converse is the substance: robust feasibility is a statement about a continuum of perturbations, each entering rationally, and testing the LMI against finitely many extreme perturbations does not produce a multiplier τ\tauτ. The exactness of the reformulation rests on a lossless certificate for an implication between quadratic inequalities, which holds only under a strict feasibility condition; that condition is where L≠0L \ne 0L=0 enters, and without it the lemma is false. The well-posedness milestone requires showing that ∥D∥<ρ−1\|D\| < \rho^{-1}∥D∥<ρ−1 is also necessary, which is not a norm estimate but needs a perturbation that makes I−DΔI - D\DeltaI−DΔ singular.

Formalization scope

Matrices are Mathlib Matrix (Fin a) (Fin b) ℝ. The affine maps are given by coefficient lists indexed by Fin (m + 1), the constant term first. The norm on matrices is the ℓ2\ell^2ℓ2 operator norm, opened with open scoped Matrix.Norms.L2Operator; it is the largest singular value, and no other matrix norm is used. X⪰0X \succeq 0X⪰0 is Matrix.PosSemidef, which includes symmetry. Block matrices are Matrix.fromBlocks over the index type Fin n ⊕ Fin q, with R(x)T−τLDTR(x)^T - \tau LD^TR(x)T−τLDT top-right and R(x)−τDLTR(x) - \tau DL^TR(x)−τDLT bottom-left. Mathlib's matrix inverse returns 000 at a singular matrix, so the condition det⁡(I−DΔ)≠0\det(I - D\Delta) \ne 0det(I−DΔ)=0 appears in the robust feasible set in the same universally quantified clause as positive semidefiniteness, as the paper's "well defined" requires; dropping it, or using an entrywise matrix norm, would change the set and is excluded.

Readings and corrections of the printed statements:

  • "The RSDP (4) and a corresponding solution xxx can be computed by solving the SDP" is read as the identity of Xρ\mathcal{X}_\rhoXρ​ with the xxx-projection of the feasible set of (10), for every xxx, together with the solution correspondence item. A statement of equal optimal values alone would be weaker and is not used.
  • Correction: L≠0L \ne 0L=0 is added to Lemma 3.1 and Theorem 3.1. The printed statements fail for L=0L = 0L=0: with n=p=q=1n = p = q = 1n=p=q=1, F=0F = 0F=0, L=0L = 0L=0, D=0D = 0D=0, R=1R = 1R=1, the perturbation does not enter, so the robust condition holds, while the LMI reads [011⋅]⪰0\begin{bmatrix}0 & 1\\1 & \cdot\end{bmatrix} \succeq 0[01​1⋅​]⪰0, which is infeasible.
  • q≥1q \ge 1q≥1 makes "matrices of appropriate size" explicit; for q=0q = 0q=0 the lower-right block is empty and the equivalence fails.
  • The standing assumptions ρ>0\rho > 0ρ>0 (§3) and ∥D∥<ρ−1\|D\| < \rho^{-1}∥D∥<ρ−1 (§3.1) are hypotheses of the goal. In Lemma 3.1, ∥D∥<1\|D\| < 1∥D∥<1 is part of the conclusion, as printed, and τ\tauτ carries no sign constraint, as printed.
  • The paper's standing assumption that the nominal problem is feasible (X0≠∅\mathcal{X}_0 \ne \emptysetX0​=∅) is not needed for the identity and is not added.

Welcome contributions: proofs of the well-posedness milestone (a spectral-norm and singular-vector argument, reusable wherever I−DΔI - D\DeltaI−DΔ must be invertible); of Lemma 3.1 from the S-procedure (the reachability lemma for norm-bounded perturbations is reusable in robust control); of the goal from Lemma 3.1 by rescaling; and general lemmas on the spectral norm of rank-one matrices and on Schur complements of block matrices.

Selected references

  • L. El Ghaoui, F. Oustry and H. Lebret, Robust Solutions to Uncertain Semidefinite Programs, SIAM J. Optim. 9(1), 33–52, 1998. https://doi.org/10.1137/S1052623496305717
  • A. Ben-Tal and A. Nemirovski, Robust Convex Optimization, Math. Oper. Res. 23(4), 769–805, 1998. https://doi.org/10.1287/moor.23.4.769
  • S. Boyd, L. El Ghaoui, E. Feron and V. Balakrishnan, Linear Matrix Inequalities in System and Control Theory, SIAM, 1994. https://doi.org/10.1137/1.9781611970777
  • S. Boyd and L. Vandenberghe, Convex Optimization, Cambridge University Press, 2004, Appendix B.2 (the S-procedure). https://web.stanford.edu/~boyd/cvxbook/
5 thms3 active usersReviewed
🏆Completed
CombinatoricsLinear OptimizationOptimization·Captain: mikedeng1

Validation of Subgradient Optimization II: A Unique Optimal Assignment Makes the Dual Optimal Set Full-DimensionalResearch Paper

Why the assignment dual matters

The subgradient method maximizes a concave, piecewise-linear function w(π)=min⁡k{ck+π⋅vk}w(\pi)=\min_k\{c_k+\pi\cdot v_k\}w(π)=mink​{ck​+π⋅vk​} by moving along a subgradient vkv_kvk​ of an active piece with a prescribed step. Held, Wolfe and Crowder's 1974 paper Validation of subgradient optimization tested the method on three families of Lagrangean duals from combinatorial optimization — the assignment problem, a relaxation of the travelling salesman problem in the style of Held and Karp, and multicommodity flows — and gave the first systematic account of when the method works in practice.

On randomly generated assignment problems of order n≤30n\le 30n≤30 the authors observed that the method usually did not merely converge: it stopped, after finitely many steps, at an iterate whose subgradient was exactly zero. Their explanation is a structural fact about the assignment dual, Theorem 3.1 of the paper: when the optimal assignment is unique — the typical case for random integer costs — the set of optimal dual prices has full dimension nnn, so a sequence of steps of decreasing length can land inside it. This mission formalizes that theorem and the steps of its proof.

Setting

There are nnn men and nnn jobs, and a real n×nn\times nn×n cost matrix A=(air)A=(a_{ir})A=(air​): aira_{ir}air​ is the cost for which man iii does job rrr. A one-to-one assignment is a permutation σ\sigmaσ of {1,…,n}\{1,\dots,n\}{1,…,n}, where σ(r)\sigma(r)σ(r) is the man doing job rrr; its cost is ∑raσ(r) r\sum_r a_{\sigma(r)\,r}∑r​aσ(r)r​. The assignment problem (3.1) asks for a permutation of minimal cost; the assignment is unique if exactly one permutation attains that minimum.

The linear relaxation of (3.1), over doubly stochastic matrices x=(xir)x=(x_{ir})x=(xir​), has the dual linear program (3.2), max⁡{∑iπi+∑rρr:πi+ρr≤air}\max\{\sum_i\pi_i+\sum_r\rho_r : \pi_i+\rho_r\le a_{ir}\}max{∑i​πi​+∑r​ρr​:πi​+ρr​≤air​}. For fixed prices π∈Rn\pi\in\mathbb R^nπ∈Rn on the men the best ρ\rhoρ is ρr=min⁡s[asr−πs]\rho_r=\min_s[a_{sr}-\pi_s]ρr​=mins​[asr​−πs​], which leaves the dual function (3.3)

w(π)=∑i=1nπi+∑r=1nmin⁡s [asr−πs],w(\pi)=\sum_{i=1}^n\pi_i+\sum_{r=1}^n\min_s\,[a_{sr}-\pi_s],w(π)=i=1∑n​πi​+r=1∑n​smin​[asr​−πs​],

the inner minimum being over the men sss for each job rrr. The optimal set is Ω={π:w(π′)≤w(π) for all π′}\Omega=\{\pi : w(\pi')\le w(\pi)\ \text{for all }\pi'\}Ω={π:w(π′)≤w(π) for all π′}.

To put www in the form min⁡k{ck+π⋅vk}\min_k\{c_k+\pi\cdot v_k\}mink​{ck​+π⋅vk​} the paper uses assignments in a weaker sense: arbitrary functions A:{1,…,n}→{1,…,n}A:\{1,\dots,n\}\to\{1,\dots,n\}A:{1,…,n}→{1,…,n}, nnn^nnn of them, with cost cA=∑raA(r) rc_A=\sum_r a_{A(r)\,r}cA​=∑r​aA(r)r​ and vector (vA)i=1−#{r:A(r)=i}(v_A)_i=1-\#\{r:A(r)=i\}(vA​)i​=1−#{r:A(r)=i} (3.4). The subgradient step raises the price of a man assigned no job and lowers the price of a man assigned several; vA=0v_A=0vA​=0 exactly when AAA is a permutation.

In the Lean development these are assignCost, assignVec, IsOptimalAssignment, w and optSet in the namespace HeldWolfeCrowder.Assignment.

Formalization targets

Goal: Theorem 3.1 (p. 70)

If the assignment problem has a unique optimal permutation, then

dim⁡aff⁡ Ω=n.\dim\operatorname{aff}\,\Omega=n .dimaffΩ=n.

The hypothesis is uniqueness among permutations; the conclusion is the dimension of the affine hull of the optimal set.

Milestones, in the order the proof uses them

  1. Eq. (3.4): w(π)=min⁡A{cA+∑iπi(vA)i}w(\pi)=\min_A\{c_A+\sum_i\pi_i(v_A)_i\}w(π)=minA​{cA​+∑i​πi​(vA​)i​} over all nnn^nnn assignments AAA.
  2. §3, Eqs. (3.1)–(3.3): www attains its maximum, and max⁡w\max wmaxw equals the cost of an optimal permutation.
  3. Eq. (3.5): if σ\sigmaσ is the unique optimal permutation, some maximizer πˉ\bar\piπˉ of www has, for every job rrr, the minimum min⁡s[asr−πˉs]\min_s[a_{sr}-\bar\pi_s]mins​[asr​−πˉs​] attained only at s=σ(r)s=\sigma(r)s=σ(r).
  4. Eq. (3.6): for an optimal permutation σ\sigmaσ, the set Π={π:air−πi>aσ(r) r−πσ(r) for all r, i≠σ(r)}\Pi=\{\pi : a_{ir}-\pi_i>a_{\sigma(r)\,r}-\pi_{\sigma(r)}\ \text{for all } r,\ i\ne\sigma(r)\}Π={π:air​−πi​>aσ(r)r​−πσ(r)​ for all r, i=σ(r)} is convex and open, v=0v=0v=0 on it, and Π⊆Ω\Pi\subseteq\OmegaΠ⊆Ω.

Significance

The theorem turns an empirical observation into a statement about the problem: finite termination of the subgradient method on assignment problems is a property of the dual, not luck. Since www is unchanged by adding the same constant to every price, Ω\OmegaΩ always contains a line; Theorem 3.1 says that, under uniqueness, it is as large as it can be. The paper (p. 70) cites the argument of its Section 2 that, with a full-dimensional optimal set, termination of the method is "nearly certain".

The result is proved in the paper; none of it is known to be machine-checked. What the formalization adds is a checked link between three classical ingredients: the integrality of the assignment polytope (Birkhoff–von Neumann, which Mathlib has as doublyStochastic_eq_convexHull_permMatrix), linear-programming duality, and strict complementary slackness, which neither Mathlib nor the platform has in the form needed. The piecewise-linear representation (3.4) is reusable wherever the assignment dual appears as a Lagrangean subproblem.

Difficulty

The inclusion Π⊆Ω\Pi\subseteq\OmegaΠ⊆Ω is elementary; the substance is that Π\PiΠ is nonempty. The obvious candidate — any optimal dual solution — fails: an optimal π\piπ may leave ties asr−πs=aσ(r) r−πσ(r)a_{sr}-\pi_s=a_{\sigma(r)\,r}-\pi_{\sigma(r)}asr​−πs​=aσ(r)r​−πσ(r)​ for some s≠σ(r)s\ne\sigma(r)s=σ(r), so it sits on the boundary of Ω\OmegaΩ and shows nothing about dimension. What is needed is an optimal price vector with all these inequalities strict at once, and uniqueness of the optimal permutation is a statement about the primal side only; transferring it to the dual side goes through the linear relaxation (3.1), whose uniqueness is not the hypothesis, and through a strict complementarity property that is not available in Mathlib or on the platform.

Formalization scope

Men and jobs are both Fin n; the costs are a : Matrix (Fin n) (Fin n) ℝ with a i r the cost of man i on job r; prices are π : Fin n → ℝ (no inner product or norm is needed, so no EuclideanSpace). The inner minimum of (3.3) is Finset.univ.inf' over the men, well defined for every n. One-to-one assignments are Equiv.Perm (Fin n) with σ r the man doing job r, so the orientation of the matrix matches (3.3); arbitrary assignments are functions Fin n → Fin n. "Of dimension nnn" is Module.finrank ℝ (vectorSpan ℝ (optSet a)) = n. The page prints the index condition of (3.6) as "i≠ri\ne ri=r"; the formalization uses i≠σ(r)i\ne\sigma(r)i=σ(r), which is what the argument requires. The case n=0n=0n=0 is allowed and trivial.

A statement asserting only that Ω\OmegaΩ is nonempty, or that it has dimension at least one, is not this theorem: both hold for every cost matrix, the second because Ω\OmegaΩ is invariant under adding a constant to all prices. The goal requires the full value nnn, and its hypothesis is uniqueness of the optimal permutation, not of the optimal linear-programming solution.

A complete development needs: the assignment linear program and its integrality (Mathlib's Birkhoff–von Neumann theorem), weak and strong duality between (3.1) and (3.2) or directly max⁡w=min⁡σcσ\max w=\min_\sigma c_\sigmamaxw=minσ​cσ​, and a strict complementarity statement for this primal–dual pair; the last two are reusable beyond this mission. Contributions of any of these, and of alternative arguments for (3.5) that avoid strict complementary slackness, are welcome.

Selected references

  • M. Held, P. Wolfe, H. P. Crowder, Validation of subgradient optimization, Mathematical Programming 6 (1974) 62–88. https://doi.org/10.1007/BF01580223
  • M. Held, R. M. Karp, The traveling-salesman problem and minimum spanning trees: Part II, Mathematical Programming 1 (1971) 6–25. https://doi.org/10.1007/BF01584070
  • H. W. Kuhn, The Hungarian method for the assignment problem, Naval Research Logistics Quarterly 2 (1955) 83–97. https://doi.org/10.1002/nav.3800020109
  • A. J. Goldman, A. W. Tucker, Theory of linear programming, in H. W. Kuhn, A. W. Tucker (eds.), Linear Inequalities and Related Systems, Annals of Mathematics Studies 38, Princeton University Press, 1956, 53–97.
  • Mathlib, Mathlib/Analysis/Convex/Birkhoff.lean (Birkhoff–von Neumann theorem, doublyStochastic_eq_convexHull_permMatrix). https://github.com/leanprover-community/mathlib4/blob/master/Mathlib/Analysis/Convex/Birkhoff.lean
6 thms3 active usersReviewed
Dynamic ProgrammingOptimization·Captain: mikedeng1

Integrating Replenishment Decisions with Advance Demand Information I: The Myopic Order-Up-To Level Is Optimal When Observed Demand Beyond the Protection Period Is LargeResearch Paper

Motivation

Many firms learn part of future demand before it has to be served: customers place orders days or weeks ahead of the date they need the goods, or commit to delivery dates in contracts. Advance demand information of this kind reduces the uncertainty the inventory manager has to protect against, and the question is how replenishment decisions should use it. Gallego and Özer (Management Science 47(10), 2001) model orders placed up to NNN periods ahead and a supply lead time LLL. They show that, with a fixed ordering cost, the classical (s,S)(s,S)(s,S) structure survives with parameters that depend on the observed future demand, and they identify when that dependence disappears.

The classical theory this builds on is the finite-horizon inventory model with a set-up cost. Scarf (1960) introduced KKK-convexity to prove that (s,S)(s,S)(s,S) policies are optimal there. Veinott (1966) and Iglehart (1963) bounded the optimal policy parameters by myopic quantities. Gallego and Özer extend both results to a state that carries a vector of observed demands.

Setting

Periods are t=1,…,Tt = 1, \dots, Tt=1,…,T. In period ttt customers place orders Dt=(Dt,t,…,Dt,t+N)D_t = (D_{t,t}, \dots, D_{t,t+N})Dt​=(Dt,t​,…,Dt,t+N​) for periods t,…,t+Nt, \dots, t+Nt,…,t+N; DtD_tDt​ is a random vector with law μt\mu_tμt​ and nonnegative components. The lead time LLL and the information horizon NNN satisfy N>L+1N > L+1N>L+1; write M=N−L−1≥1M = N - L - 1 \ge 1M=N−L−1≥1. The state at the start of period ttt is a pair (xt,ot)(x_t, o_t)(xt​,ot​): the modified inventory position xt∈Rx_t \in \mathbb{R}xt​∈R, and the vector

ot=(ot,t+L+1,…,ot,t+N−1)∈RMo_t = (o_{t,t+L+1}, \dots, o_{t,t+N-1}) \in \mathbb{R}^Mot​=(ot,t+L+1​,…,ot,t+N−1​)∈RM

of demands already observed for the periods beyond the protection period t,…,t+Lt, \dots, t+Lt,…,t+L.

The manager raises xtx_txt​ to an order-up-to level y≥xty \ge x_ty≥xt​, paying a set-up cost Kt>0K_t > 0Kt​>0 if y>xty > x_ty>xt​. After DtD_tDt​ is observed, the state moves to

xt+1=y−∑s=tt+L+1Dt,s−ot,t+L+1,ot+1,s=ot,s+Dt,s  (s=t+L+2,…,t+N),x_{t+1} = y - \sum_{s=t}^{t+L+1} D_{t,s} - o_{t,t+L+1},\qquad o_{t+1,s} = o_{t,s} + D_{t,s}\ \ (s = t+L+2,\dots,t+N),xt+1​=y−s=t∑t+L+1​Dt,s​−ot,t+L+1​,ot+1,s​=ot,s​+Dt,s​  (s=t+L+2,…,t+N),

with ot,t+N=0o_{t,t+N} = 0ot,t+N​=0. With a convex single-period cost GtG_tGt​, discount factors αt>0\alpha_t > 0αt​>0 and δ(z)=1{z>0}\delta(z) = \mathbf 1\{z > 0\}δ(z)=1{z>0}, the optimal cost satisfies JT+1≡0J_{T+1} \equiv 0JT+1​≡0 and

Jt(x,o)=min⁡y≥x{Ktδ(y−x)+Vt(y,o)},Vt(y,o)=Gt(y)+αt+1 E Jt+1(xt+1,ot+1).J_t(x, o) = \min_{y \ge x}\{K_t\delta(y - x) + V_t(y, o)\},\qquad V_t(y, o) = G_t(y) + \alpha_{t+1}\,\mathbb E\,J_{t+1}(x_{t+1}, o_{t+1}).Jt​(x,o)=y≥xmin​{Kt​δ(y−x)+Vt​(y,o)},Vt​(y,o)=Gt​(y)+αt+1​EJt+1​(xt+1​,ot+1​).

Let Ht(x,o)=Kt+min⁡y≥xVt(y,o)−Vt(x,o)H_t(x, o) = K_t + \min_{y\ge x}V_t(y,o) - V_t(x,o)Ht​(x,o)=Kt​+miny≥x​Vt​(y,o)−Vt​(x,o). The order-up-to level St(o)S_t(o)St​(o) is the least minimizer of Vt(⋅,o)V_t(\cdot, o)Vt​(⋅,o), and the reorder point is st(o)=max⁡{x:Ht(x,o)≤0}s_t(o) = \max\{x : H_t(x,o) \le 0\}st​(o)=max{x:Ht​(x,o)≤0}.

A function ggg is (a,b)(a,b)(a,b)-convex, g∈C(a,b)g \in C(a,b)g∈C(a,b), if g(θx1+(1−θ)x2)≤θ(a+g(x1))+(1−θ)(b+g(x2))g(\theta x_1 + (1-\theta)x_2) \le \theta(a + g(x_1)) + (1-\theta)(b + g(x_2))g(θx1​+(1−θ)x2​)≤θ(a+g(x1​))+(1−θ)(b+g(x2​)) for all x1≤x2x_1 \le x_2x1​≤x2​ and θ∈[0,1]\theta \in [0,1]θ∈[0,1]; C(0,K)C(0,K)C(0,K) is Scarf's KKK-convexity. In the stationary problem Gt=GG_t = GGt​=G, Kt=KK_t = KKt​=K, αt=α\alpha_t = \alphaαt​=α and μt=ν\mu_t = \nuμt​=ν, and the myopic levels are

Sm=min⁡{y:G(y)≤G(x) ∀x},sm=max⁡{y≤Sm:G(y)≥K+G(Sm)},S‾=inf⁡{y>Sm:G(y)>G(Sm)+αK}.S^m = \min\{y : G(y) \le G(x)\ \forall x\},\quad s^m = \max\{y \le S^m : G(y) \ge K + G(S^m)\},\quad \overline S = \inf\{y > S^m : G(y) > G(S^m) + \alpha K\}.Sm=min{y:G(y)≤G(x) ∀x},sm=max{y≤Sm:G(y)≥K+G(Sm)},S=inf{y>Sm:G(y)>G(Sm)+αK}.

Formalization targets

Goal: Theorem 2 (p. 1350)

For the stationary problem, every 1≤t≤T1 \le t \le T1≤t≤T and every observed-demand vector ot≥0o_t \ge 0ot​≥0,

ot,t+L+1≥S‾−sm  ⟹  St(ot)=Sm.o_{t,t+L+1} \ge \overline S - s^m \implies S_t(o_t) = S^m.ot,t+L+1​≥S−sm⟹St​(ot​)=Sm.

Only the observed demand for the first period beyond the protection period is compared with the threshold; the other components of oto_tot​ are free.

Milestones

  • Lemma 1, Parts 1, 2, 4, 5 (p. 1349): inclusion, positive combinations, expectations, and g(max⁡(x,s))g(\max(x,s))g(max(x,s)) for (a,b)(a,b)(a,b)-convex functions.
  • Lemma 2 and Corollary 1 (p. 1350): for V∈C(0,K)V \in C(0,K)V∈C(0,K) with a minimizer SSS, HHH changes sign once from −-− to +++, and (for continuous VVV) J(x)=V(max⁡(s,x))J(x) = V(\max(s,x))J(x)=V(max(s,x)).
  • Theorem 1, Parts 1–3 (p. 1350): Vt(⋅,ot)∈C(0,Kt)V_t(\cdot, o_t) \in C(0, K_t)Vt​(⋅,ot​)∈C(0,Kt​) is coercive, a state-dependent (st(ot),St(ot))(s_t(o_t), S_t(o_t))(st​(ot​),St​(ot​)) policy is optimal, and Jt(⋅,ot)∈C(0,Kt)J_t(\cdot, o_t) \in C(0, K_t)Jt​(⋅,ot​)∈C(0,Kt​) with its limits at ±∞\pm\infty±∞.
  • Lemma 3 (p. 1351): Sm≤St(ot)≤S‾S^m \le S_t(o_t) \le \overline SSm≤St​(ot​)≤S and sm≤st(ot)s^m \le s_t(o_t)sm≤st​(ot​).

Significance

Theorem 1 says that advance demand information does not destroy the (s,S)(s,S)(s,S) structure: the optimal policy is still a reorder point and an order-up-to level, now functions of oto_tot​. Theorem 2 is a horizon result. Once enough demand is already booked for period t+L+1t+L+1t+L+1, the order-up-to level is the myopic one, computed from GGG alone, and the rest of the information vector can be ignored. The threshold S‾−sm\overline S - s^mS−sm grows with the set-up cost. In practice the manager then orders only to cover demand up to period t+Lt+Lt+L, knowing another order will be placed in period t+1t+1t+1, and the search for state-dependent policies is confined to states with little booked demand.

The results are proved in the paper, with some steps argued informally (the unit forward difference in Lemma 3, a minimum over an open set in the definition of S‾\overline SS). As far as the platform's catalog shows, no (s,S)(s,S)(s,S) optimality theorem of this form has a machine-checked proof: the platform has KKK-convexity lemmas for a single constant, not for (a,b)(a,b)(a,b)-convexity or a dynamic program with a vector state. A formal proof would check the infinite-state induction and the conditions under which the expectation in (9) is finite. The (a,b)(a,b)(a,b)-convexity layer and the one-period results (Lemma 2, Corollary 1) are reusable for any set-up cost model.

Difficulty

The obvious induction proves KKK-convexity of VtV_tVt​ and then applies Scarf's argument. The vector state makes each step conditional on facts the scalar case gets for free. The expectation in (9) mixes a random shift of xxx with a random update of ooo, so preserving KKK-convexity needs Lemma 1, Part 4 rather than the scalar version. Coercivity and continuity of Vt(⋅,o)V_t(\cdot, o)Vt​(⋅,o), and measurability and integrability of D↦Jt+1(xt+1,ot+1)D \mapsto J_{t+1}(x_{t+1}, o_{t+1})D↦Jt+1​(xt+1​,ot+1​), must be carried through the induction jointly in (x,o)(x, o)(x,o). They cannot be assumed.

For Theorem 2, VtV_tVt​ is not a function of GGG alone, and a comparison of St(ot)S_t(o_t)St​(ot​) with SmS^mSm has to control EJt+1\mathbb E J_{t+1}EJt+1​. That requires the lower bound sm≤st+1(ot+1)s^m \le s_{t+1}(o_{t+1})sm≤st+1​(ot+1​) at the next period, which holds only on the nonnegative state space.

Formalization scope

Everything is stated about the functional equation (8)–(9). The paper derives it in Appendix A from a control problem over history-dependent policies, citing Özer (2000); that reduction is out of scope. The demand vector is Fin (L + M + 2) → ℝ and ooo is Fin M → ℝ, with component jjj equal to ot,t+L+1+jo_{t,t+L+1+j}ot,t+L+1+j​. JJJ is defined by backward recursion, minima over y≥xy \ge xy≥x are infima over {y:x≤y}\{y : x \le y\}{y:x≤y}, and expectations are Bochner integrals. GtG_tGt​ is an abstract primitive, not built from holding and penalty costs.

The paper's hypotheses are GtG_tGt​ convex with Gt(y)→∞G_t(y) \to \inftyGt​(y)→∞ as ∣y∣→∞|y| \to \infty∣y∣→∞ (stated for G~t\widetilde G_tGt​, used for GtG_tGt​), Kt>0K_t > 0Kt​>0 (Section 4), and αt+1Kt+1≤Kt\alpha_{t+1}K_{t+1} \le K_tαt+1​Kt+1​≤Kt​. The formalization adds hypotheses the paper uses without stating:

  • positive discount factors;
  • demands nonnegative almost surely;
  • every demand component has a finite mean, and ∣Gt(y)∣≤at+bt∣y∣|G_t(y)| \le a_t + b_t|y|∣Gt​(y)∣≤at​+bt​∣y∣ (so the expectation in (9) is finite);
  • continuity of VVV in the abstract Corollary 1;
  • nonnegative observed demands oto_tot​ in Lemma 3 and Theorem 2, where the lower bound fails for negative ot,t+L+1o_{t,t+L+1}ot,t+L+1​.

All hypotheses are placed on the primitives; nothing is assumed about the derived VtV_tVt​ or JtJ_tJt​. A Solutions/verification instance (G(y)=∣y∣G(y) = |y|G(y)=∣y∣, K=1K = 1K=1, α=1/2\alpha = 1/2α=1/2, zero demand) shows these hypotheses can be met together.

The definition of S‾\overline SS is read as an infimum: the paper prints a minimum over an open set. St(ot)S_t(o_t)St​(ot​) and st(ot)s_t(o_t)st​(ot​) appear through IsLeast and IsGreatest, never as sInf/sSup values. A default value therefore cannot satisfy a conclusion, and vacuous readings (an empty minimizer set, an unbounded reorder set) are excluded. The infinite-horizon results (Lemma 4, Theorem 3, Corollary 2) and the zero set-up cost case are not part of this mission. Contributions are welcome at every level, and most of all a reusable library for (a,b)(a,b)(a,b)-convex functions.

Selected references

  • G. Gallego, Ö. Özer, Integrating Replenishment Decisions with Advance Demand Information, Management Science 47(10):1344–1360, 2001. https://doi.org/10.1287/mnsc.47.10.1344.10261
  • H. Scarf, The Optimality of (S, s) Policies in the Dynamic Inventory Problem, in Mathematical Methods in the Social Sciences, Stanford University Press, 1960.
  • A. F. Veinott, On the Optimality of (s,S) Inventory Policies: New Conditions and a New Proof, SIAM Journal on Applied Mathematics 14(5):1067–1083, 1966. https://doi.org/10.1137/0114086
  • D. L. Iglehart, Optimality of (s, S) Policies in the Infinite Horizon Dynamic Inventory Problem, Management Science 9(2):259–267, 1963. https://doi.org/10.1287/mnsc.9.2.259
15 thms3 active usersReviewed
🏆Completed
Convex OptimizationOptimization·Captain: mikedeng1

Lifts of Convex Sets and Cone Factorizations I: A Proper K-Lift of a Convex Body Yields a K-Factorization of Its Slack Operator, and a K-Factorization Yields a K-LiftResearch Paper

Motivation

Many convex sets that appear in optimization have complicated descriptions in their own space but simple descriptions as projections of higher-dimensional sets. A polytope with exponentially many facets can be the shadow of a polyhedron with polynomially many; the unit disk is the projection of a slice of the cone of 2×22\times 22×2 positive semidefinite matrices. Such a representation, a lift, turns linear optimization over the original set into a linear or semidefinite program over the lifted one, so the size of the smallest lift measures how hard the set is for conic optimization.

For polytopes and polyhedral lifts, Yannakakis (Yannakakis 1991) showed that the minimal size of a lift equals the nonnegative rank of the polytope's slack matrix. This turned questions about extended formulations into questions about matrix factorizations, and it is the basis of the lower bounds of Fiorini, Massar, Pokutta, Tiwary and de Wolf (2012) for the cut, stable set and traveling salesman polytopes. Lift-and-project hierarchies (Sherali–Adams, Lovász–Schrijver, Lasserre) all produce lifts to nonnegative orthants or positive semidefinite cones, so a criterion for the existence of a lift is also a criterion for when such a hierarchy can succeed.

Gouveia, Parrilo and Thomas (arXiv:1111.3164, Mathematics of Operations Research 38(2), 2013) extended Yannakakis' theorem from polytopes and polyhedral cones to arbitrary convex bodies and arbitrary closed convex cones. Their Theorem 2.4 is the target of this mission.

Timeline:

  • 1991, Yannakakis: polytopes, polyhedral lifts, nonnegative factorizations of the slack matrix.
  • 2012, Fiorini, Massar, Pokutta, Tiwary, de Wolf: superpolynomial lower bounds on polyhedral lifts via nonnegative rank; a positive semidefinite analogue for polytopes.
  • 2011/2013, Gouveia, Parrilo, Thomas: convex bodies and general closed convex cones (Theorem 2.4), with psd rank as the semidefinite analogue of nonnegative rank.

Setting

Throughout, Rk\mathbb R^kRk carries the Euclidean inner product ⟨⋅,⋅⟩\langle\cdot,\cdot\rangle⟨⋅,⋅⟩.

A convex body is a set C⊆RnC \subseteq \mathbb R^nC⊆Rn that is convex, compact, and contains the origin in its interior. Its polar is

C∘={ y∈Rn:⟨x,y⟩≤1 for all x∈C }.C^\circ = \{\, y \in \mathbb R^n : \langle x, y\rangle \le 1 \text{ for all } x \in C \,\}.C∘={y∈Rn:⟨x,y⟩≤1 for all x∈C}.

A point p∈Cp \in Cp∈C is an extreme point if p=(p1+p2)/2p = (p_1+p_2)/2p=(p1​+p2​)/2 with p1,p2∈Cp_1,p_2\in Cp1​,p2​∈C forces p1=p2=pp_1 = p_2 = pp1​=p2​=p; ext⁡(C)\operatorname{ext}(C)ext(C) is the set of extreme points. The slack operator of CCC is

SC:ext⁡(C)×ext⁡(C∘)→R,SC(x,y)=1−⟨x,y⟩.S_C : \operatorname{ext}(C)\times\operatorname{ext}(C^\circ) \to \mathbb R, \qquad S_C(x,y) = 1 - \langle x,y\rangle .SC​:ext(C)×ext(C∘)→R,SC​(x,y)=1−⟨x,y⟩.

It is nonnegative, and for a polytope it is the slack matrix: rows indexed by vertices, columns by facet normals.

Let K⊆RmK \subseteq \mathbb R^mK⊆Rm be a full-dimensional closed convex cone: closed, convex, closed under nonnegative scaling, with nonempty interior. Its dual is K∗={y:⟨x,y⟩≥0 ∀x∈K}K^* = \{y : \langle x,y\rangle \ge 0 \ \forall x\in K\}K∗={y:⟨x,y⟩≥0 ∀x∈K}.

  • A KKK-lift of CCC is Q=K∩LQ = K\cap LQ=K∩L, where L⊆RmL\subseteq\mathbb R^mL⊆Rm is an affine subspace and π:Rm→Rn\pi:\mathbb R^m\to\mathbb R^nπ:Rm→Rn is a linear map with C=π(K∩L)C = \pi(K\cap L)C=π(K∩L). The lift is proper if LLL meets the interior of KKK (Definition 2.1).
  • SCS_CSC​ is KKK-factorizable if there are maps, not necessarily linear, A:ext⁡(C)→KA:\operatorname{ext}(C)\to KA:ext(C)→K and B:ext⁡(C∘)→K∗B:\operatorname{ext}(C^\circ)\to K^*B:ext(C∘)→K∗ with SC(x,y)=⟨A(x),B(y)⟩S_C(x,y) = \langle A(x), B(y)\rangleSC​(x,y)=⟨A(x),B(y)⟩ for all (x,y)(x,y)(x,y) (Definition 2.2).

In Lean these are IsConvexBody, IsClosedConvexCone, HasLift, HasProperLift and SlackFactorizable in the namespace ConeLifts.Factorization, together with the series' shared ConeLifts.Shared.polar and ConeLifts.Shared.dualCone.

Formalization targets

Goal: Theorem 2.4

For n≥1n \ge 1n≥1, a convex body C⊆RnC\subseteq\mathbb R^nC⊆Rn and a full-dimensional closed convex cone K⊆RmK\subseteq\mathbb R^mK⊆Rm:

(C has a proper K-lift⇒SC is K-factorizable)  ∧  (SC is K-factorizable⇒C has a K-lift).\bigl(C \text{ has a proper } K\text{-lift} \Rightarrow S_C \text{ is } K\text{-factorizable}\bigr) \;\wedge\; \bigl(S_C \text{ is } K\text{-factorizable} \Rightarrow C \text{ has a } K\text{-lift}\bigr).(C has a proper K-lift⇒SC​ is K-factorizable)∧(SC​ is K-factorizable⇒C has a K-lift).

The two implications are not an equivalence: the forward one assumes properness, and the lift produced by the converse may be improper.

Milestones

In the order the paper's proof uses them:

  1. (§2, p. 3) C=conv⁡(ext⁡C)C = \operatorname{conv}(\operatorname{ext} C)C=conv(extC) and C∘=conv⁡(ext⁡C∘)C^\circ = \operatorname{conv}(\operatorname{ext} C^\circ)C∘=conv(extC∘).
  2. (proof, p. 4) For every c∈ext⁡(C∘)c\in\operatorname{ext}(C^\circ)c∈ext(C∘), max⁡{⟨c,x⟩:x∈C}=1\max\{\langle c,x\rangle : x\in C\} = 1max{⟨c,x⟩:x∈C}=1, attained.
  3. (proof, p. 4) If C=π(K∩L)C = \pi(K\cap L)C=π(K∩L), L=w0+L0L = w_0 + L_0L=w0​+L0​ and w0∈int⁡Kw_0\in\operatorname{int}Kw0​∈intK, then for c∈ext⁡(C∘)c \in \operatorname{ext}(C^\circ)c∈ext(C∘)
1=min⁡{⟨w0,z⟩:z−π∗(c)∈K∗, z∈L0⊥},1 = \min\{\langle w_0, z\rangle : z - \pi^*(c)\in K^*,\ z\in L_0^\perp\},1=min{⟨w0​,z⟩:z−π∗(c)∈K∗, z∈L0⊥​},

with the minimum attained. 4. (proof, p. 5) For L={(x,z):1−⟨x,y⟩=⟨z,B(y)⟩ ∀y∈ext⁡(C∘)}L = \{(x,z) : 1-\langle x,y\rangle = \langle z, B(y)\rangle\ \forall y\in\operatorname{ext}(C^\circ)\}L={(x,z):1−⟨x,y⟩=⟨z,B(y)⟩ ∀y∈ext(C∘)} and its projection LKL_KLK​ to Rm\mathbb R^mRm: 0∉LK0\notin L_K0∈/LK​. 5. (proof, p. 5) If BBB maps into K∗K^*K∗, z∈Kz\in Kz∈K and (x,z)∈L(x,z)\in L(x,z)∈L, then x∈Cx\in Cx∈C. 6. (proof, p. 5) For each z∈K∩LKz\in K\cap L_Kz∈K∩LK​ there is a unique xzx_zxz​ with (xz,z)∈L(x_z,z)\in L(xz​,z)∈L.

Significance

The result. Theorem 2.4 makes the existence of a lift of a convex body to a given cone a purely algebraic question about its slack operator. Every lower bound on lift size in the paper and its successors goes through it: the nonnegative-rank bounds for polytopes (Section 4 of the paper), the proof that the stable set polytope of an nnn-vertex graph has no lift to S+n\mathcal S^n_+S+n​ (Section 5), and the later psd-rank literature. It also puts Yannakakis' theorem and its semidefinite analogue under a single statement.

Formalizing it. The theorem is proved on paper; no machine-checked version is known to exist. Formalizing it requires conic strong duality with dual attainment under a Slater condition, which Mathlib does not have, and finite-dimensional Krein–Milman for the polar body. The companion missions of this series (nonnegative-rank lower bounds; stable set polytopes and psd lifts) use the correspondence as their entry point.

Difficulty

The converse half is elementary once the extreme points of C∘C^\circC∘ are known to generate it. The forward half is not: B(c)B(c)B(c) must be an element of K∗K^*K∗ that certifies ⟨c,x⟩≤1\langle c, x\rangle \le 1⟨c,x⟩≤1 on CCC through the lift. A separating functional gives this certificate on π(K∩L)\pi(K\cap L)π(K∩L), but writing it as z−π∗(c)z - \pi^*(c)z−π∗(c) with z⊥L0z \perp L_0z⊥L0​, z−π∗(c)∈K∗z - \pi^*(c)\in K^*z−π∗(c)∈K∗ and ⟨w0,z⟩=1\langle w_0,z\rangle = 1⟨w0​,z⟩=1 exactly is conic duality with a zero gap and an attained dual optimum. For closed convex cones the gap can be positive or the dual unattained unless a constraint qualification holds; this is why properness is assumed. Weak duality alone gives only ≥1\ge 1≥1, and a dual sequence approaching 111 does not yield a factor. The paper notes (p. 5) that, since the proof uses strong duality, it is not obvious how to remove properness for a general closed convex cone.

Formalization scope

Conventions fixed by the Lean statements:

  • Rk\mathbb R^kRk is EuclideanSpace ℝ (Fin k); every pairing, in SSS, in K∗K^*K∗ and in the factorization, is its inner product.
  • The polar is one-sided, ⟨x,y⟩≤1\langle x,y\rangle\le 1⟨x,y⟩≤1; Mathlib's absolute polar is not used.
  • A convex body is compact, convex, with 000 in its interior. The paper's "full-dimensional convex body in Rn\mathbb R^nRn" is read as including n≥1n\ge 1n≥1: for n=0n = 0n=0, C={0}C = \{0\}C={0} has the proper Rm\mathbb R^mRm-lift {0}\{0\}{0} while SC(0,0)=1S_C(0,0) = 1SC​(0,0)=1 cannot factor through K∗={0}K^* = \{0\}K∗={0}, so the forward half is false there. The goal and milestones 2–3 assume 1≤n1\le n1≤n.
  • KKK is closed, convex, contains 000 and is closed under nonnegative scaling; full-dimensionality is (interior K).Nonempty. Pointedness is not assumed.
  • LLL is a Mathlib AffineSubspace and π\piπ a linear map; the lift condition is the set equality C=π(K∩L)C = \pi(K\cap L)C=π(K∩L).
  • A,BA, BA,B are total functions Rn→Rm\mathbb R^n\to\mathbb R^mRn→Rm constrained only on ext⁡(C)\operatorname{ext}(C)ext(C), resp. ext⁡(C∘)\operatorname{ext}(C^\circ)ext(C∘), which is equivalent to maps out of the extreme points. They are not required to be linear or continuous.
  • Milestone 3 is the second, substituted form of the paper's dual (z=MTyz = M^{\mathsf T}yz=MTy), stated with L.directionᗮ and LinearMap.adjoint π; minima and maxima are stated with IsLeast/IsGreatest, so attainment is part of every claim.

Trivializing readings are excluded: π\piπ is linear, not an arbitrary function (with an arbitrary function every set is a "lift"); LLL is an affine subspace, not an arbitrary set; and BBB takes values in K∗K^*K∗, not KKK, which for a cone that is not self-dual would be a different and generally false statement.

Needed infrastructure: finite-dimensional Krein–Milman in the form C=conv⁡(ext⁡C)C = \operatorname{conv}(\operatorname{ext} C)C=conv(extC) for compact convex sets (Mathlib has the closure form); the bipolar theorem (C∘)∘=C(C^\circ)^\circ = C(C∘)∘=C for closed convex C∋0C\ni 0C∋0 with the one-sided polar; compactness of C∘C^\circC∘ when 0∈int⁡C0\in\operatorname{int} C0∈intC; and conic linear programming duality with a Slater point, including dual attainment. The last two are reusable well beyond this mission. Proofs of individual milestones, and of these general facts as separate lemmas, are welcome.

Selected references

  • J. Gouveia, P. A. Parrilo, R. R. Thomas, Lifts of Convex Sets and Cone Factorizations, Mathematics of Operations Research 38(2):248–264, 2013. arXiv:1111.3164v2, doi:10.1287/moor.1120.0575
  • M. Yannakakis, Expressing combinatorial optimization problems by linear programs, Journal of Computer and System Sciences 43(3):441–466, 1991. doi:10.1016/0022-0000(91)90024-Y
  • S. Fiorini, S. Massar, S. Pokutta, H. R. Tiwary, R. de Wolf, Linear vs. semidefinite extended formulations: exponential separation and strong lower bounds, STOC 2012. arXiv:1111.0837
14 thms3 active usersReviewed
🏆Completed
Convex OptimizationLinear algebraNumerical Analysis+1·Captain: mikedeng1

Robust Solutions to Least-Squares Problems with Uncertain Data III: Structured Robust Least Squares Is Solved Exactly by a Semidefinite ProgramResearch Paper

Motivation

Least squares fits a model Ax≈bAx \approx bAx≈b as if the data (A,b)(A, b)(A,b) were exact. In practice they are measured, rounded or estimated, and the least-squares solution can be very sensitive to such errors. El Ghaoui and Lebret (SIAM J. Matrix Anal. Appl. 18(4), 1997) proposed to treat the errors as deterministic, unknown but bounded, and to choose xxx minimizing the worst-case residual over all admissible data. For unstructured perturbations of [A b][A\ b][A b] bounded in Frobenius norm this leads to a second-order cone program (missions I and II of this series).

In many applications the perturbations have a known structure: a Toeplitz matrix stays Toeplitz, a parameter enters several entries at once, or only some entries are uncertain. An unstructured bound then over-estimates the worst case. The paper's §4 treats perturbations that are affine in a parameter vector δ\deltaδ bounded in Euclidean norm, and shows that the resulting structured robust least-squares (SRLS) problem is still solved exactly, now by a semidefinite program (SDP). This model of uncertainty (an ellipsoid of affinely parametrized data) is the one later adopted as the basic uncertainty set of robust optimization; see Ben-Tal and Nemirovski, Math. Oper. Res. 23(4), 1998.

Setting

Vectors carry the Euclidean norm ∥v∥=vTv\|v\| = \sqrt{v^Tv}∥v∥=vTv​. Given matrices A0,A1,…,Ap∈Rn×mA_0, A_1, \dots, A_p \in \mathbb{R}^{n\times m}A0​,A1​,…,Ap​∈Rn×m and vectors b0,b1,…,bp∈Rnb_0, b_1, \dots, b_p \in \mathbb{R}^nb0​,b1​,…,bp​∈Rn, define for every δ∈Rp\delta \in \mathbb{R}^pδ∈Rp

A(δ)=A0+∑i=1pδiAi,b(δ)=b0+∑i=1pδibi.\mathbf A(\delta) = A_0 + \sum_{i=1}^p \delta_i A_i, \qquad \mathbf b(\delta) = b_0 + \sum_{i=1}^p \delta_i b_i .A(δ)=A0​+i=1∑p​δi​Ai​,b(δ)=b0​+i=1∑p​δi​bi​.

For ρ≥0\rho \ge 0ρ≥0 and x∈Rmx \in \mathbb{R}^mx∈Rm the structured worst-case residual is

rS(A,b,ρ,x)=max⁡∥δ∥≤ρ∥A(δ)x−b(δ)∥,r_S(\mathbf A, \mathbf b, \rho, x) = \max_{\|\delta\| \le \rho} \|\mathbf A(\delta)x - \mathbf b(\delta)\|,rS​(A,b,ρ,x)=∥δ∥≤ρmax​∥A(δ)x−b(δ)∥,

and xxx is an SRLS solution if it minimizes rS(A,b,ρ,⋅)r_S(\mathbf A, \mathbf b, \rho, \cdot)rS​(A,b,ρ,⋅) over Rm\mathbb{R}^mRm. The paper takes ρ=1\rho = 1ρ=1 throughout §4 and writes rS(A,b,x)r_S(\mathbf A, \mathbf b, x)rS​(A,b,x).

For fixed xxx let M(x)=[A1x−b1 ⋯ Apx−bp]∈Rn×pM(x) = [A_1x - b_1\ \cdots\ A_px - b_p] \in \mathbb{R}^{n\times p}M(x)=[A1​x−b1​ ⋯ Ap​x−bp​]∈Rn×p and

F=M(x)TM(x),g=M(x)T(A0x−b0),h=∥A0x−b0∥2.F = M(x)^TM(x), \qquad g = M(x)^T(A_0x - b_0), \qquad h = \|A_0x - b_0\|^2 .F=M(x)TM(x),g=M(x)T(A0​x−b0​),h=∥A0​x−b0​∥2.

Since A(δ)x−b(δ)=(A0x−b0)+M(x)δ\mathbf A(\delta)x - \mathbf b(\delta) = (A_0x - b_0) + M(x)\deltaA(δ)x−b(δ)=(A0​x−b0​)+M(x)δ, the squared residual at δ\deltaδ is the quadratic function h+2gTδ+δTFδh + 2g^T\delta + \delta^TF\deltah+2gTδ+δTFδ. Finally, for scalars λ,τ\lambda, \tauλ,τ,

F(λ,τ)=[λ−τ−h−gT−gτI−F].\mathcal F(\lambda, \tau) = \begin{bmatrix} \lambda - \tau - h & -g^T \\ -g & \tau I - F \end{bmatrix}.F(λ,τ)=[λ−τ−h−g​−gTτI−F​].

Formalization targets

Goal: Theorem 4.2

With p≥1p \ge 1p≥1 and ρ=1\rho = 1ρ=1, consider the SDP in (λ,τ,x)(\lambda, \tau, x)(λ,τ,x)

minimize λsubject to[λ−τ0(A0x−b0)T0τIM(x)TA0x−b0M(x)I]⪰0.(32)\text{minimize } \lambda \quad \text{subject to} \quad \begin{bmatrix} \lambda - \tau & 0 & (A_0x - b_0)^T \\ 0 & \tau I & M(x)^T \\ A_0x - b_0 & M(x) & I \end{bmatrix} \succeq 0. \tag{32}minimize λsubject to​λ−τ0A0​x−b0​​0τIM(x)​(A0​x−b0​)TM(x)TI​​⪰0.(32)

The goal states that (a) for all xxx and λ\lambdaλ, some τ\tauτ makes (λ,τ,x)(\lambda, \tau, x)(λ,τ,x) feasible if and only if rS(A,b,x)2≤λr_S(\mathbf A, \mathbf b, x)^2 \le \lambdarS​(A,b,x)2≤λ; and (b) (λ,τ,x)(\lambda, \tau, x)(λ,τ,x) is optimal for (32) if and only if xxx is an SRLS solution, λ=rS(A,b,x)2\lambda = r_S(\mathbf A, \mathbf b, x)^2λ=rS​(A,b,x)2, and (λ,τ,x)(\lambda, \tau, x)(λ,τ,x) is feasible. This is the precise content of the paper's "the SRLS can be solved by computing an optimal solution of (32)".

Milestones

  1. Lemma 2.1 (S-procedure), in two items: the multiplier condition is sufficient for every ppp; for p=1p = 1p=1 it is also necessary when F1(ζ0)>0F_1(\zeta_0) > 0F1​(ζ0​)>0 for some ζ0\zeta_0ζ0​.
  2. Eq. (28): rS(A,b,x)2=max⁡δTδ≤1[1;δ]T[hgTgF][1;δ]r_S(\mathbf A, \mathbf b, x)^2 = \max_{\delta^T\delta \le 1} [1;\delta]^T \begin{bmatrix} h & g^T \\ g & F\end{bmatrix} [1;\delta]rS​(A,b,x)2=maxδTδ≤1​[1;δ]T[hg​gTF​][1;δ].
  3. Eq. (29): for λ≥0\lambda \ge 0λ≥0, that quadratic form is ≤λ\le \lambda≤λ on the unit ball if and only if F(λ,τ)⪰0\mathcal F(\lambda, \tau) \succeq 0F(λ,τ)⪰0 for some τ\tauτ.
  4. Theorem 4.1, first assertion: rS(A,b,x)2=min⁡{λ:∃τ, F(λ,τ)⪰0}r_S(\mathbf A, \mathbf b, x)^2 = \min\{\lambda : \exists \tau,\ \mathcal F(\lambda, \tau) \succeq 0\}rS​(A,b,x)2=min{λ:∃τ, F(λ,τ)⪰0}, the minimum attained.
  5. §4.2, Schur-complement step: the matrix of (32) is positive semidefinite if and only if F(λ,τ)\mathcal F(\lambda, \tau)F(λ,τ) is.

Significance

The result shows that a min–max problem over a nonconvex worst case (the inner problem maximizes a convex quadratic over a ball) is equivalent to a single convex SDP whose size is linear in nnn, mmm and ppp, and hence solvable in polynomial time by interior-point methods. It covers as special cases the unstructured problem of §3, least squares with uncertainty in selected entries, and Toeplitz or otherwise patterned perturbations. The exactness contrasts with the next section of the paper, where the linear-fractional and ℓ∞\ell_\inftyℓ∞​-bounded versions are in general only bounded from above, or shown NP-hard.

The result is proved in the paper; to the best of current knowledge it has not been formalized. The platform already has the one-constraint S-procedure (ConvexOptimization.s_procedure, proved, in a different sign and block convention); this mission adds the robust least-squares objects, the reduction to the S-procedure, the Schur-complement step, and the optimal-solution correspondence of Theorem 4.2. The worst-case residual and SDP (32) definitions are reusable by later robust-regression missions.

Difficulty

The obvious approach is to compute the inner maximum directly. The function δ↦h+2gTδ+δTFδ\delta \mapsto h + 2g^T\delta + \delta^TF\deltaδ↦h+2gTδ+δTFδ is convex, so its maximum over the unit ball is attained on the boundary, but it is not given by any closed-form expression in general, and maximizing a convex function is not a convex problem. Exactness therefore rests on the lossless S-procedure for one quadratic constraint, a nonconvex duality statement that fails for two or more constraints; the sufficient direction alone only yields an upper bound.

A second point is passing from "for fixed xxx" (Theorem 4.1) to "optimal over xxx" (Theorem 4.2): F(λ,τ)\mathcal F(\lambda, \tau)F(λ,τ) is quadratic in xxx, and only the Schur-complement lift (32) is jointly affine in (λ,τ,x)(\lambda, \tau, x)(λ,τ,x). The correspondence of optimal solutions must then be checked in both directions, including that the optimal λ\lambdaλ is the squared residual and not the residual.

Formalization scope

  • Data are A0 : Matrix (Fin n) (Fin m) ℝ, A : Fin p → Matrix (Fin n) (Fin m) ℝ, b0 : Fin n → ℝ, b : Fin p → Fin n → ℝ; A i is the paper's Ai+1A_{i+1}Ai+1​ (0-based index). Vectors live in Fin k → ℝ with the Euclidean norm written out as ∑ivi2\sqrt{\sum_i v_i^2}∑i​vi2​​, never Mathlib's sup norm.
  • The maximum defining rSr_SrS​ is sSup of the set of attained residuals over the closed ball; for ρ≥0\rho \ge 0ρ≥0 this set is nonempty and bounded, so sSup is the true maximum. The theorems use ρ=1\rho = 1ρ=1, as the paper does; the paper derives general ρ\rhoρ by scaling and that is not stated here.
  • Block matrices are Matrix.fromBlocks in the printed order (scalar block first: Unit ⊕ Fin p; for (32), (Unit ⊕ Fin p) ⊕ Fin n). "⪰0\succeq 0⪰0" is Mathlib's PosSemidef, which includes symmetry; all matrices here are symmetric by construction.
  • p≥1p \ge 1p≥1 is assumed in (29), Theorem 4.1 and Theorem 4.2, although the paper does not state it: for p=0p = 0p=0 the block τI\tau IτI is empty, τ\tauτ is unconstrained, every λ\lambdaλ is feasible and both SDPs lose their meaning. Eq. (28), Lemma 2.1 and the Schur-complement step hold for every ppp and are stated without it.
  • Optimality in (32) is stated as feasibility plus λ≤λ′\lambda \le \lambda'λ≤λ′ for every feasible (λ′,τ′,x′)(\lambda', \tau', x')(λ′,τ′,x′). A formalization that only proves existence of some feasible τ\tauτ, or only an inequality between the optimal values, is weaker than Theorem 4.2 and does not close the goal.
  • Theorem 4.1's second and third assertions (the one-dimensional reformulation (30)–(31) and the worst-case perturbation) are not included: they use the notion "(F,g)(F, g)(F,g)-controllable", which the paper does not define.
  • Useful infrastructure: Mathlib's Schur-complement lemmas (Matrix.PosSemidef.fromBlocks₂₂ and relatives in LinearAlgebra.Matrix.SchurComplement); the platform's ConvexOptimization.s_procedure and ConvexOptimization.single_constraint_quadratic_strong_duality with their definitions ConvexOptimization_quadraticForms, included as reference items. A bridge lemma between the platform's block convention and this mission's is a welcome contribution, as is a general-ρ\rhoρ version.

Selected references

  • L. El Ghaoui and H. Lebret, Robust Solutions to Least-Squares Problems with Uncertain Data, SIAM J. Matrix Anal. Appl. 18(4):1035–1064, 1997. https://doi.org/10.1137/S0895479896298130
  • S. Boyd, L. El Ghaoui, E. Feron and V. Balakrishnan, Linear Matrix Inequalities in System and Control Theory, SIAM, 1994 (the S-procedure, p. 24). https://doi.org/10.1137/1.9781611970777
  • A. Ben-Tal and A. Nemirovski, Robust Convex Optimization, Math. Oper. Res. 23(4):769–805, 1998. https://doi.org/10.1287/moor.23.4.769
  • I. Pólik and T. Terlaky, A Survey of the S-Lemma, SIAM Review 49(3):371–418, 2007. https://doi.org/10.1137/S003614450444614X
11 thms3 active usersReviewed
🏆Completed
Convex OptimizationLinear algebraNumerical Analysis+1·Captain: mikedeng1

Robust Solutions to Least-Squares Problems with Uncertain Data I: The Worst-Case Residual and Its Unique MinimizerResearch Paper

Motivation

The least-squares (LS) problem min⁡x∥Ax−b∥\min_x \|Ax - b\|minx​∥Ax−b∥ assumes that the data A∈Rn×mA \in \mathbb{R}^{n\times m}A∈Rn×m, b∈Rnb \in \mathbb{R}^nb∈Rn are exact. In applications they rarely are: they come from measurements, from linearizations, or from models with neglected dynamics. A classical response is sensitivity analysis or regularization (Tikhonov), where a weight trades the size of the solution against the fit, and the choice of that weight is left to the user. El Ghaoui and Lebret (SIAM J. Matrix Anal. Appl. 18(4), 1997) take a deterministic view instead: the true data lie in a known ball around (A,b)(A, b)(A,b), and the solution should minimize the residual it can be forced to have in the worst case over that ball. The paper shows that this robust least-squares (RLS) problem is solvable exactly, in the unstructured case by a second-order cone program (SOCP). The same worst-case idea, applied to regression, underlies the later equivalence between robustness and regularization (Xu, Caramanis and Mannor, 2009) and is a standard entry point to robust optimization (Ben-Tal, El Ghaoui and Nemirovski, Robust Optimization, 2009).

This mission formalizes the first main result of the paper, Theorem 3.1: the worst-case residual has a closed form, its minimizer is unique, and minimizing it is an SOCP.

Setting

Vectors carry the Euclidean norm ∥v∥=(∑ivi2)1/2\|v\| = (\sum_i v_i^2)^{1/2}∥v∥=(∑i​vi2​)1/2. For a matrix XXX, ∥X∥F=(∑i,jXij2)1/2\|X\|_F = (\sum_{i,j} X_{ij}^2)^{1/2}∥X∥F​=(∑i,j​Xij2​)1/2 is the Frobenius norm and ∥X∥\|X\|∥X∥ the largest singular value, i.e. the smallest c≥0c \ge 0c≥0 with ∥Xv∥≤c∥v∥\|Xv\| \le c\|v\|∥Xv∥≤c∥v∥ for all vvv.

Fix A∈Rn×mA \in \mathbb{R}^{n\times m}A∈Rn×m and b∈Rnb \in \mathbb{R}^nb∈Rn. A perturbation is a pair ΔA∈Rn×m\Delta A \in \mathbb{R}^{n\times m}ΔA∈Rn×m, Δb∈Rn\Delta b \in \mathbb{R}^nΔb∈Rn, collected in the augmented matrix Δ=[ΔA Δb]∈Rn×(m+1)\Delta = [\Delta A\ \Delta b] \in \mathbb{R}^{n\times(m+1)}Δ=[ΔA Δb]∈Rn×(m+1). For a bound ρ≥0\rho \ge 0ρ≥0 and x∈Rmx \in \mathbb{R}^mx∈Rm, the worst-case residual is (paper, eq. (1))

r(A,b,ρ,x)=max⁡∥[ΔA Δb]∥F≤ρ∥(A+ΔA)x−(b+Δb)∥,r(A,b,\rho,x) = \max_{\|[\Delta A\ \Delta b]\|_F \le \rho} \|(A+\Delta A)x - (b+\Delta b)\|,r(A,b,ρ,x)=∥[ΔA Δb]∥F​≤ρmax​∥(A+ΔA)x−(b+Δb)∥,

and xxx is an RLS solution if it minimizes r(A,b,ρ,⋅)r(A,b,\rho,\cdot)r(A,b,ρ,⋅). The bound constrains the augmented matrix jointly, not ΔA\Delta AΔA and Δb\Delta bΔb separately. The paper normalizes ρ=1\rho = 1ρ=1 and writes r(A,b,x)=r(A,b,1,x)r(A,b,x) = r(A,b,1,x)r(A,b,x)=r(A,b,1,x). Finally, [x;1]∈Rm+1[x;1] \in \mathbb{R}^{m+1}[x;1]∈Rm+1 denotes xxx stacked over 111. In the Lean development these are RobustLS.Unstructured.eucNorm, frobNorm, specNorm, augment, stackOne, worstCaseResidual A b ρ x, its largest-singular-value variant worstCaseResidualSpec, and the SOCP constraint predicate SocpFeasible A b x λ τ.

Formalization targets

Goal: Theorem 3.1 (p. 1040)

For n≥1n \ge 1n≥1, every AAA, bbb:

r(A,b,x)=∥Ax−b∥+∥x∥2+1for all x∈Rm,r(A,b,x) = \|Ax-b\| + \sqrt{\|x\|^2+1} \quad \text{for all } x \in \mathbb{R}^m,r(A,b,x)=∥Ax−b∥+∥x∥2+1​for all x∈Rm,

the problem min⁡x∈Rmr(A,b,x)\min_{x \in \mathbb{R}^m} r(A,b,x)minx∈Rm​r(A,b,x) has exactly one solution xRLSx_{\mathrm{RLS}}xRLS​, and it is the SOCP

minimize λsubject to∥Ax−b∥≤λ−τ,∥[x;1]∥≤τ,(15)\text{minimize } \lambda \quad\text{subject to}\quad \|Ax-b\| \le \lambda-\tau,\quad \|[x;1]\| \le \tau, \tag{15}minimize λsubject to∥Ax−b∥≤λ−τ,∥[x;1]∥≤τ,(15)

in the sense that r(A,b,x)r(A,b,x)r(A,b,x) is the least λ\lambdaλ for which some τ\tauτ makes (x,λ,τ)(x,\lambda,\tau)(x,λ,τ) feasible.

Milestones

  1. Eq. (16). Every perturbation with ∥[ΔA Δb]∥F≤1\|[\Delta A\ \Delta b]\|_F \le 1∥[ΔA Δb]∥F​≤1 has residual at most ∥Ax−b∥+∥x∥2+1\|Ax-b\| + \sqrt{\|x\|^2+1}∥Ax−b∥+∥x∥2+1​.
  2. The worst-case perturbation. For a unit vector uuu aligned with Ax−bAx - bAx−b (arbitrary if Ax=bAx = bAx=b), the rank-one matrix Δ=u[xT −1]/∥x∥2+1\Delta = u[x^T\ {-1}]/\sqrt{\|x\|^2+1}Δ=u[xT −1]/∥x∥2+1​ has ∥Δ∥F=∥Δ∥=1\|\Delta\|_F = \|\Delta\| = 1∥Δ∥F​=∥Δ∥=1 and attains the bound.
  3. Spectral norm. The worst case over the larger ball ∥[ΔA Δb]∥≤1\|[\Delta A\ \Delta b]\| \le 1∥[ΔA Δb]∥≤1 is the same value.
  4. Strict convexity. x↦r(A,b,x)x \mapsto r(A,b,x)x↦r(A,b,x) is strictly convex on Rm\mathbb{R}^mRm.
  5. The SOCP (15). For every xxx, r(A,b,x)r(A,b,x)r(A,b,x) is the optimal λ\lambdaλ of (15) with xxx fixed, and xxx is an RLS solution exactly when it is the xxx-part of an optimal solution of (15).

Significance

The closed form replaces a maximization over a matrix ball of dimension n(m+1)n(m+1)n(m+1) by two Euclidean norms. It shows that the RLS objective is the LS residual plus a penalty ∥x∥2+1\sqrt{\|x\|^2+1}∥x∥2+1​ that does not depend on AAA or bbb, which is the starting point for the paper's Theorem 3.2 (the RLS solution is a Tikhonov-regularized LS solution with a data-dependent weight) and its analysis of continuity and conditioning. The SOCP formulation places the problem in the class solved by interior-point methods, at a cost the paper compares with one singular value decomposition of AAA. The spectral-norm statement says the worst case does not depend on which of the two standard matrix norms bounds the perturbation.

The result is proved in the paper; the proof is short. To the best of the planning survey (September 2026), no machine-checked proof exists, and Prove2Me has no statement about worst-case residuals or robust least squares. The mission produces a verified closed form that later missions of this series (Tikhonov form of the solution, structured and linear-fractional perturbations) and any formalization of robust regression can import.

Difficulty

The upper bound alone does not give the theorem: the statement is an equality, and the equality needs an explicit maximizer. The paper's printed maximizer is wrong by a sign: with [xT 1][x^T\ 1][xT 1] in place of [xT −1][x^T\ {-1}][xT −1] the perturbation does not attain the bound (for A=0A = 0A=0, x=0x = 0x=0, b=e1b = e_1b=e1​ it gives residual 000 instead of 222), so a transcription of the printed proof fails. Two further points are silent in the paper. The operator norm of a rank-one matrix has to be computed from the definition of the largest singular value. Uniqueness of the minimizer needs existence first, which follows from growth of rrr at infinity and is not stated. Working with the sSup definition of the worst case requires showing the set of residuals is bounded, which is milestone 1.

Formalization scope

  • Dimensions are Fin n, Fin m; AAA is Matrix (Fin n) (Fin m) ℝ, bbb and xxx are functions Fin n → ℝ, Fin m → ℝ. The augmented matrix [ΔA Δb][\Delta A\ \Delta b][ΔA Δb] is indexed by Fin m ⊕ Unit, and so is [x;1][x;1][x;1].
  • Vector norms are the Euclidean norm written as ∑ivi2\sqrt{\sum_i v_i^2}∑i​vi2​​ (eucNorm), never Mathlib's ‖·‖ on Fin n → ℝ, which is the sup norm. The Frobenius norm and the largest singular value are explicit definitions (frobNorm, specNorm); specNorm is the infimum of admissible operator constants.
  • The maximum in (1) is sSup of the set of attained residuals. For ρ≥0\rho \ge 0ρ≥0 the set is nonempty and bounded, so this is the true maximum; milestones 1 and 2 state the bound and the attaining perturbation directly, so no statement relies on the value of sSup on an unbounded set.
  • The paper's normalization ρ=1\rho = 1ρ=1 is kept; general ρ>0\rho > 0ρ>0 follows from the scaling ϕ(A,b,ρ)=ρ ϕ(A/ρ,b/ρ,1)\phi(A,b,\rho) = \rho\,\phi(A/\rho,b/\rho,1)ϕ(A,b,ρ)=ρϕ(A/ρ,b/ρ,1) the paper records on p. 1039 and is not a target.
  • The goal assumes n≥1n \ge 1n≥1. For n=0n = 0n=0 the only perturbation is the empty matrix, the worst case is 000, and the closed form fails; the paper's setting (Ax≃bAx \simeq bAx≃b with data b∈Rnb \in \mathbb{R}^nb∈Rn) has n≥1n \ge 1n≥1. Milestones 3–5 carry the same hypothesis.
  • Milestone 2 states the corrected perturbation [xT −1][x^T\ {-1}][xT −1]; the printed [xT 1][x^T\ 1][xT 1] is false.
  • A trivializing formalization — an upper bound in place of the equality, a worst case over ΔA\Delta AΔA and Δb\Delta bΔb bounded separately, or uniqueness among critical points only — is ruled out: the goal is the equality for the jointly bounded augmented matrix and ∃! of a global minimizer over all of Rm\mathbb{R}^mRm.

Contributions welcome: lemmas on Frobenius and operator norms of rank-one matrices, the inequality ∥Mz∥≤∥M∥F∥z∥\|Mz\| \le \|M\|_F\|z\|∥Mz∥≤∥M∥F​∥z∥ in this explicit setting, and strict convexity of x↦∥x∥2+1x \mapsto \sqrt{\|x\|^2+1}x↦∥x∥2+1​; these are reusable beyond the mission.

Selected references

  • L. El Ghaoui and H. Lebret, Robust Solutions to Least-Squares Problems with Uncertain Data, SIAM Journal on Matrix Analysis and Applications 18(4):1035–1064, 1997. https://doi.org/10.1137/S0895479896298130
  • A. Ben-Tal, L. El Ghaoui and A. Nemirovski, Robust Optimization, Princeton University Press, 2009. https://doi.org/10.1515/9781400831050
  • H. Xu, C. Caramanis and S. Mannor, Robust Regression and Lasso, Journal of Machine Learning Research 10:1485–1510, 2009 (IEEE Trans. Inf. Theory 56(7), 2010). https://jmlr.org/papers/v10/xu09b.html
  • M. S. Lobo, L. Vandenberghe, S. Boyd and H. Lebret, Applications of Second-Order Cone Programming, Linear Algebra and its Applications 284:193–228, 1998. https://doi.org/10.1016/S0024-3795(98)10032-0
7 thms3 active usersReviewed
🏆Completed
ProbabilityTheoretical Computer Science·Captain: mikedeng1

Competitive Paging Algorithms III: No Randomized Paging Algorithm Is Better than H_k-CompetitiveResearch Paper

Motivation

Paging is the problem of managing a two-level memory: a cache holds kkk of the nnn pages a program uses, every request must find its page in the cache, and a request to a page outside the cache (a page fault) forces the algorithm to bring the page in and evict another. An on-line algorithm chooses what to evict without seeing future requests. Sleator and Tarjan (CACM 1985) measured on-line paging algorithms against the optimal off-line algorithm, which knows the whole request sequence, and showed that no deterministic on-line algorithm can be within a factor smaller than kkk of it.

Randomization changes that picture. Fiat, Karp, Luby, McGeoch, Sleator and Young (J. Algorithms 1991; arXiv:cs/0205038) gave a randomized algorithm, the marking algorithm, whose expected number of faults is within 2Hk2H_k2Hk​ of the optimum, where Hk=1+12+⋯+1k≈ln⁡kH_k = 1 + \tfrac12 + \dots + \tfrac1k \approx \ln kHk​=1+21​+⋯+k1​≈lnk. This mission formalizes the other half of their paper's picture: no randomized paging algorithm can do better than HkH_kHk​. The bound says that the logarithmic behaviour is not an artefact of one algorithm but a property of the problem.

Timeline:

  • 1985 — Sleator and Tarjan: deterministic paging algorithms have competitive factor at least kkk; LRU and FIFO achieve kkk.
  • 1988 — Karlin, Manasse, Rudolph and Sleator (Algorithmica 3, 1988) introduce the term competitive; Manasse, McGeoch and Sleator (STOC 1988; J. Algorithms 1990) extend it to randomized algorithms and pose the kkk-server problem, of which paging is the uniform-metric case.
  • 1991 — Fiat et al.: the marking algorithm is 2Hk2H_k2Hk​-competitive, and no randomized algorithm is better than HkH_kHk​-competitive (Theorem 4 and Corollary 5 of the paper). Raghavan gave an alternative proof of the lower bound through Yao's minimax principle.
  • 1991 — McGeoch and Sleator give an HkH_kHk​-competitive randomized paging algorithm (Algorithmica 6, 1991), so the lower bound is tight.

Setting

Let MMM be a set of nnn vertices with the uniform metric: any two distinct vertices are at distance 111. A configuration of kkk servers is a map C:{1,…,k}→MC : \{1,\dots,k\} \to MC:{1,…,k}→M; server sss sits at C(s)C(s)C(s), and a vertex is covered when some server sits on it. A request sequence σ\sigmaσ is a finite list of vertices. A deterministic on-line algorithm assigns to every prefix of requests the configuration after serving it, in such a way that the vertex just requested is covered; its cost on σ\sigmaσ is the total distance travelled by its servers, which on the uniform metric is the number of server moves. Paging with kkk cache slots and nnn pages is exactly this kkk-server problem on nnn uniform vertices.

The optimal off-line cost OPTC0(σ)\mathrm{OPT}_{C_0}(\sigma)OPTC0​​(σ) is the least cost of any schedule of configurations that starts at C0C_0C0​ and covers each request of σ\sigmaσ in turn.

A randomized on-line algorithm AAA is a probability space (Ω,μ)(\Omega,\mu)(Ω,μ) of coin outcomes together with a deterministic on-line algorithm AωA_\omegaAω​ for each outcome ω\omegaω. Its expected cost CA(σ)C_A(\sigma)CA​(σ) is the average of the cost of AωA_\omegaAω​ on σ\sigmaσ over ω\omegaω. The request sequence is fixed in advance and does not depend on the coins (an oblivious adversary). Following the paper, AAA is ccc-competitive from the initial configuration C0C_0C0​ if there is a constant aaa such that

CA(σ)  ≤  c⋅OPTC0(σ)+afor every request sequence σ.C_A(\sigma) \;\le\; c \cdot \mathrm{OPT}_{C_0}(\sigma) + a \qquad \text{for every request sequence } \sigma .CA​(σ)≤c⋅OPTC0​​(σ)+afor every request sequence σ.

For the lower-bound argument, the probability vector p=(pi)i∈Mp=(p_i)_{i\in M}p=(pi​)i∈M​ after a prefix σ\sigmaσ has pip_ipi​ equal to the probability, over ω\omegaω, that vertex iii is not covered by AωA_\omegaAω​ after serving σ\sigmaσ. A set SSS of marked vertices and the number u=n−∣S∣u = n - |S|u=n−∣S∣ of unmarked vertices are bookkeeping of the adversary, updated as the marking algorithm would update them.

Formalization targets

Goal: Corollary 5

For 1≤k≤n−11 \le k \le n-11≤k≤n−1, every randomized on-line algorithm AAA with kkk servers on nnn uniform vertices, every initial configuration C0C_0C0​ and every real ccc,

c<Hk  ⟹  A is not c-competitive from C0.c < H_k \;\Longrightarrow\; A \text{ is not } c\text{-competitive from } C_0 .c<Hk​⟹A is not c-competitive from C0​.

Theorem 4 (milestone)

The case k=n−1k = n-1k=n−1: no randomized algorithm for the uniform (n−1)(n-1)(n−1)-server problem on nnn vertices is ccc-competitive with c<Hn−1c < H_{n-1}c<Hn−1​.

Claims of the proof of Theorem 4 (milestones)

With ppp the probability vector, SSS the marked set, P=∑i∈SpiP = \sum_{i\in S} p_iP=∑i∈S​pi​ and u=n−∣S∣u = n - |S|u=n−∣S∣:

∑ipi=1(servers on distinct vertices),CA(σ i)≥CA(σ)+pi,\sum_i p_i = 1 \quad(\text{servers on distinct vertices}),\qquad C_A(\sigma\,i) \ge C_A(\sigma) + p_i,i∑​pi​=1(servers on distinct vertices),CA​(σi)≥CA​(σ)+pi​, P=0⇒∃ i∉S, pi≥1u,P>ϵ>0⇒max⁡j∈Spj≥ϵ∣S∣>0,P = 0 \Rightarrow \exists\, i\notin S,\ p_i \ge \tfrac1u, \qquad P > \epsilon > 0 \Rightarrow \max_{j\in S} p_j \ge \tfrac{\epsilon}{|S|} > 0,P=0⇒∃i∈/S, pi​≥u1​,P>ϵ>0⇒j∈Smax​pj​≥∣S∣ϵ​>0, pj=max⁡j′∉Spj′⇒pj≥1−Pu,P≤ϵ⇒ϵ+pj≥ϵ+1−Pu≥ϵ+1−ϵu≥1u.p_j = \max_{j'\notin S} p_{j'} \Rightarrow p_j \ge \tfrac{1-P}{u}, \qquad P \le \epsilon \Rightarrow \epsilon + p_j \ge \epsilon + \tfrac{1-P}{u} \ge \epsilon + \tfrac{1-\epsilon}{u} \ge \tfrac1u .pj​=j′∈/Smax​pj′​⇒pj​≥u1−P​,P≤ϵ⇒ϵ+pj​≥ϵ+u1−P​≥ϵ+u1−ϵ​≥u1​.

Significance

The result. Together with the marking algorithm's 2Hk2H_k2Hk​ upper bound, the corollary pins the randomized competitive ratio of paging to Θ(log⁡k)\Theta(\log k)Θ(logk), an exponential improvement over the deterministic ratio kkk that no randomized algorithm can push below HkH_kHk​. For k=n−1k = n-1k=n−1 the marking algorithm itself is Hn−1H_{n-1}Hn−1​-competitive, so Theorem 4 makes it optimal there. The HkH_kHk​ bound is the benchmark every later randomized paging algorithm is measured against, including the HkH_kHk​-competitive algorithm of McGeoch and Sleator, and it is the uniform-metric base case of the randomized kkk-server conjecture.

Formalizing it. The theorem is proved and classical; no machine-checked proof is known to exist. The platform already has the deterministic bound (KServer.uniform_not_competitive_below_k, ratio kkk) and a formal Yao averaging principle for randomized kkk-server algorithms (KServer.randomized_yao_averaging), but no randomized paging lower bound. This mission produces the first formal HkH_kHk​ lower bound, stated against the published randomized kkk-server model, and a formal version of the paper's adversary argument. Either route — the paper's adaptive construction of a nemesis sequence from the probability vector, or Raghavan's distributional argument through Yao's principle — is welcome.

Difficulty

The adversary may not look at the coins, yet it must build one fixed sequence against which the expected cost is high in every phase. Requesting an uncovered vertex is not available, since which vertex is uncovered depends on the coins; requesting the vertex with the largest uncovered probability gives only 1/n1/n1/n per request and loses the harmonic sum. Lifting the per-phase bound to the asymptotic statement also requires handling the additive constant aaa, the initial configuration of the off-line algorithm, and, for Corollary 5, the reduction from nnn vertices to k+1k+1k+1 of them for an algorithm that may still place servers on the others.

Formalization scope

The Lean development reuses the published definitions KServer_model (configurations Fin k → M, deterministic on-line algorithms as functions of the request prefix, offlineCost) and KServer_randomized (RandomizedAlgorithm: a probability measure on coin outcomes, a deterministic algorithm per outcome, measurable costs; expCost as a lower Lebesgue integral in [0,∞][0,\infty][0,∞]; IsCompetitiveFrom C₀ c: every drawn algorithm starts at C0C_0C0​ and there is one constant aaa, fixed before the sequence, with expCost σ ≤ ENNReal.ofReal (c * offlineCost C₀ σ + a)). The clamp at 000 in ENNReal.ofReal only weakens the property the goal refutes. The vertex set is an abstract type MMM with an equivalence Fin n ≃ M and the uniform metric as a hypothesis, never the line metric of Fin n. HkH_kHk​ is Mathlib's harmonic k cast to R\mathbb RR. The goal quantifies over every algorithm and every initial configuration, with no laziness or distinct-positions assumption, and over every real c<Hkc < H_kc<Hk​, including c≤0c \le 0c≤0.

A formalization in which competitiveness is vacuous (a model with no algorithms, or a cost that is always infinite), in which the adversary may choose the sequence after seeing the coins, or which fixes ccc or the additive constant, would be a different statement and is ruled out by the published definitions used here.

The probability vector is the one new definition, uncoveredProb A σ i. Milestones about it assume the uncovered events measurable, the standing convention that pip_ipi​ is a probability; the model itself only guarantees measurable costs. The milestone ∑ipi=1\sum_i p_i = 1∑i​pi​=1 assumes the n−1n-1n−1 servers occupy distinct vertices, as in the paper; in general ∑ipi≥1\sum_i p_i \ge 1∑i​pi​≥1. The arithmetic milestones are stated for an arbitrary probability vector on a finite set. Reusable pieces: the probability vector and the cost lemma apply to any randomized kkk-server algorithm on a uniform metric, and a restriction lemma (from nnn vertices to k+1k+1k+1) would serve other paging lower bounds.

Selected references

  • A. Fiat, R. M. Karp, M. Luby, L. A. McGeoch, D. D. Sleator, N. E. Young, Competitive Paging Algorithms, J. Algorithms 12(4):685–699, 1991. https://doi.org/10.1016/0196-6774(91)90041-V ; preprint arXiv:cs/0205038v1 (cited version). https://arxiv.org/abs/cs/0205038
  • D. D. Sleator, R. E. Tarjan, Amortized Efficiency of List Update and Paging Rules, Comm. ACM 28(2):202–208, 1985. https://doi.org/10.1145/2786.2793
  • M. S. Manasse, L. A. McGeoch, D. D. Sleator, Competitive Algorithms for Server Problems, J. Algorithms 11(2):208–230, 1990. https://doi.org/10.1016/0196-6774(90)90003-W
  • L. A. McGeoch, D. D. Sleator, A Strongly Competitive Randomized Paging Algorithm, Algorithmica 6:816–825, 1991.
  • P. Raghavan, Lecture Notes on Randomized Algorithms, IBM Research Report, Yorktown Heights, 1990 (the alternative proof of the lower bound, pp. 118–119).
11 thms3 active usersReviewed
PreviousPage 10 of 36Next

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me