Prove2Me
Navigate
DiscoverFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

Convex Optimization

253 missions · 152 completed

Missions

Open101Completed152All253
Operations ResearchOptimizationProbability·Captain: mikedeng1

Minimization Methods for Non-Differentiable Functions VI: Almost-Sure Convergence of the Stochastic Subgradient MethodTextbook

Motivation

Many optimization problems in operations research are posed on an expectation: a two-stage or multistage stochastic program minimizes f(x)=E F(x,ξ)f(x) = E\,F(x,\xi)f(x)=EF(x,ξ), where F(⋅,ξ)F(\cdot,\xi)F(⋅,ξ) is convex but nonsmooth and the expectation cannot be computed exactly. What can be computed is a stochastic subgradient, a random vector whose mean is a subgradient of fff. The stochastic subgradient method replaces the exact subgradient in the classical method by such a random vector. It was introduced by Yu. M. Ermoliev and N. Z. Shor in 1968 and developed by Ermoliev, Nurminski and others into a standard tool of stochastic programming; the same scheme, under the name stochastic (sub)gradient descent, underlies most of large-scale machine learning.

This mission formalizes Section 2.6 of N. Z. Shor, Minimization Methods for Non-Differentiable Functions (Springer 1985): the almost-sure convergence theorem for the stochastic subgradient method (Theorem 2.19), together with two deterministic results of the same section on perturbed and restarted variants of the subgradient method (Theorems 2.18 and 2.20).

Timeline, as recorded in the book:

  • 1968, Ermoliev and Shor: the notion of a stochastic subgradient, introduced for a random search method for two-stage stochastic programs; the convergence theorem reproduced as Theorem 2.19.
  • 1972, Bazhenov: convergence of a subgradient method with restarts for almost differentiable (in general nonconvex) functions, Theorem 2.18.
  • 1976, Shepilov: stability of the subgradient method with respect to errors in the point where the subgradient is computed, Theorem 2.20.

Setting

EnE_nEn​ is nnn-dimensional Euclidean space with inner product (x,y)(x,y)(x,y). A vector ggg is a subgradient of f:En→Rf : E_n \to \mathbb{R}f:En​→R at x0x_0x0​ if f(x)−f(x0)≥(g,x−x0)f(x) - f(x_0) \ge (g, x - x_0)f(x)−f(x0​)≥(g,x−x0​) for all xxx; M∗M^*M∗ is the set of minimum points of fff.

Stochastic subgradient method. Fix a probability space (Ω,F,P)(\Omega, \mathcal F, P)(Ω,F,P) with a filtration (Fk)k≥0(\mathcal F_k)_{k \ge 0}(Fk​)k≥0​, a deterministic starting point x0x_0x0​, stepsize rules hk:En→Rh_k : E_n \to \mathbb{R}hk​:En​→R and random vectors gk:Ω→Eng_k : \Omega \to E_ngk​:Ω→En​. The iterates are

xk+1=xk−hk(xk) gk,k=0,1,…x_{k+1} = x_k - h_k(x_k)\, g_k, \qquad k = 0,1,\dotsxk+1​=xk​−hk​(xk​)gk​,k=0,1,…

In the book's notation gk=gω(xk)g_k = g_\omega(x_k)gk​=gω​(xk​): a random vector whose expectation, given the state at step kkk, is a subgradient of fff at xkx_kxk​. In the Lean development the iterates are stochIter h G x₀ k ω.

Perturbed subgradient method (Shepilov). Given a subgradient selection gfg_fgf​, points x~k\tilde x_kx~k​ with ∥x~k−xk∥≤δk\|\tilde x_k - x_k\| \le \delta_k∥x~k​−xk​∥≤δk​, and steps hk>0h_k > 0hk​>0: xk+1=xk−hk gf(x~k)/∥gf(x~k)∥x_{k+1} = x_k - h_k\, g_f(\tilde x_k)/\|g_f(\tilde x_k)\|xk+1​=xk​−hk​gf​(x~k​)/∥gf​(x~k​)∥.

Restarted method (Bazhenov). For a function fff that is almost differentiable (Lipschitz on bounded sets, differentiable almost everywhere, with gradient continuous where it exists) and a selection gf(x)g_f(x)gf​(x) of almost-gradients (limit points of gradients at nearby points of differentiability), with Sr={x:∥x−x∗∥≤r}S_r = \{x : \|x - x^*\| \le r\}Sr​={x:∥x−x∗∥≤r}: take the normalized step xˉk+1=xk−hk gf(xk)/∥gf(xk)∥\bar x_{k+1} = x_k - h_k\, g_f(x_k)/\|g_f(x_k)\|xˉk+1​=xk​−hk​gf​(xk​)/∥gf​(xk​)∥ and restart from x0x_0x0​ whenever xˉk+1\bar x_{k+1}xˉk+1​ leaves SrS_rSr​ (resetIter).

Formalization targets

Goal: Theorem 2.19 (p. 46)

Let fff be convex with a unique minimum point x∗x^*x∗. Suppose E{gk∣Fk}E\{g_k \mid \mathcal F_k\}E{gk​∣Fk​} is a subgradient of fff at xkx_kxk​, E{∥gk∥2∣Fk}≤cE\{\|g_k\|^2 \mid \mathcal F_k\} \le cE{∥gk​∥2∣Fk​}≤c, and almost surely hk(xk)>0h_k(x_k) > 0hk​(xk​)>0, ∑khk(xk)=+∞\sum_k h_k(x_k) = +\infty∑k​hk​(xk​)=+∞, ∑khk2(xk)<∞\sum_k h_k^2(x_k) < \infty∑k​hk2​(xk​)<∞. Then

P(lim⁡k→∞∥xk−x∗∥=0)=1.P\Big(\lim_{k\to\infty} \|x_k - x^*\| = 0\Big) = 1 .P(k→∞lim​∥xk​−x∗∥=0)=1.

Milestones

  1. Eq. (2.42), the conditional one-step inequality
E{∥xk+1−x∗∥2∣Fk}≤∥xk−x∗∥2+c hk2(xk).E\{\|x_{k+1} - x^*\|^2 \mid \mathcal F_k\} \le \|x_k - x^*\|^2 + c\,h_k^2(x_k).E{∥xk+1​−x∗∥2∣Fk​}≤∥xk​−x∗∥2+chk2​(xk​).
  1. Proof of Theorem 2.19, pp. 46–47: with probability one ∥xk−x∗∥2\|x_k - x^*\|^2∥xk​−x∗∥2 converges to a finite limit (no divergence condition on the steps).
  2. Theorem 2.20 (Shepilov): under δk→0\delta_k \to 0δk​→0, ∑hkδk<∞\sum h_k\delta_k < \infty∑hk​δk​<∞, ∑hk2<∞\sum h_k^2 < \infty∑hk2​<∞, ∑hk=∞\sum h_k = \infty∑hk​=∞, the perturbed method converges to a point of M∗M^*M∗.
  3. Theorem 2.18 (Bazhenov): if f(x∗)=min⁡Srff(x^*) = \min_{S_r} ff(x∗)=minSr​​f and inf⁡Sr∖Sε(gf(x),x−x∗)>0\inf_{S_r\setminus S_\varepsilon} (g_f(x), x - x^*) > 0infSr​∖Sε​​(gf​(x),x−x∗)>0 for every 0<ε<r0 < \varepsilon < r0<ε<r, the restarted method with hk→0h_k \to 0hk​→0, ∑hk=∞\sum h_k = \infty∑hk​=∞ converges to x∗x^*x∗ from any x0∈Srx_0 \in S_rx0​∈Sr​.

Significance

Theorem 2.19 is the basic justification of stochastic subgradient methods: without computing fff or any exact subgradient, the method reaches the minimizer with probability one, under stepsize conditions that are met by hk=1/(k+1)h_k = 1/(k+1)hk​=1/(k+1). It is the nonsmooth convex counterpart of the Robbins–Monro theorem and the prototype of the almost-sure convergence results for stochastic quasi-gradient methods used in stochastic programming. Theorem 2.20 shows that the deterministic method tolerates summable errors in the point where the subgradient is evaluated, which is what allows subgradients to be approximated by finite differences (Section 1.3). Theorem 2.18 extends the convergence of the normalized method to local minima of a class of nonconvex functions.

All four results are proved in the literature. To the best of the platform search (September 2026), none is machine-checked: the platform has almost-sure convergence theorems for smooth stochastic approximation under ODE-type hypotheses (Borkar–Meyn) and in-expectation bounds for stochastic gradient descent, neither of which covers this recursion. A formal proof of the goal would give a reusable almost-sure convergence argument for nonsmooth stochastic methods on top of Mathlib's martingale theory.

Difficulty

The deterministic proof of convergence of the subgradient method compares ∥xk+1−x∗∥2\|x_{k+1}-x^*\|^2∥xk+1​−x∗∥2 with ∥xk−x∗∥2\|x_k - x^*\|^2∥xk​−x∗∥2 along the whole trajectory. With random directions this comparison holds only in conditional expectation, and the term hk(gk−E{gk∣Fk},xk−x∗)h_k(g_k - E\{g_k\mid\mathcal F_k\}, x_k - x^*)hk​(gk​−E{gk​∣Fk​},xk​−x∗) is not controlled pathwise. Taking expectations of the one-step inequality and summing gives only bounds on E∥xk−x∗∥2E\|x_k - x^*\|^2E∥xk​−x∗∥2, which do not yield almost-sure convergence. Moreover the stepsize hk(xk)h_k(x_k)hk​(xk​) depends on the random iterate, so the conditions ∑hk2(xk)<∞\sum h_k^2(x_k) < \infty∑hk2​(xk​)<∞ and ∑hk(xk)=∞\sum h_k(x_k) = \infty∑hk​(xk​)=∞ hold only almost surely, not uniformly, and the iterates need not be square-integrable. Identifying the almost-sure limit as 000 requires using the uniqueness of the minimizer to bound (E{gk∣Fk},xk−x∗)(E\{g_k\mid\mathcal F_k\}, x_k - x^*)(E{gk​∣Fk​},xk​−x∗) away from zero outside a neighbourhood of x∗x^*x∗.

In Theorems 2.18 and 2.20 the difficulty is that the distance to x∗x^*x∗ is not monotone: steps taken near the solution, or with a perturbed subgradient, can increase it, and a restart can move the iterate far away.

Formalization scope

  • EnE_nEn​ is EuclideanSpace ℝ (Fin n); fff is real-valued (finite everywhere); convexity is ConvexOn ℝ Set.univ f; uniqueness of x∗x^*x∗ is a separate hypothesis.
  • Probabilistic model. The book assumes the distribution of gω(xk)g_\omega(x_k)gω​(xk​) is determined by xkx_kxk​ and independent of the past, and remarks this is inessential. The formalization uses a filtration: gkg_kgk​ is Fk+1\mathcal F_{k+1}Fk+1​-measurable, each hkh_khk​ is Borel measurable, x0x_0x0​ is deterministic, and the hypotheses are on conditional expectations given Fk\mathcal F_kFk​. This contains the book's model.
  • Condition (iii) is printed as E∥gω(xk)∥2≤cE\|g_\omega(x_k)\|^2 \le cE∥gω​(xk​)∥2≤c; the proof uses the conditional bound in (2.42), and the formalization assumes the conditional bound E{∥gk∥2∣Fk}≤cE\{\|g_k\|^2\mid\mathcal F_k\} \le cE{∥gk​∥2∣Fk​}≤c almost surely.
  • Every expectation carries an integrability hypothesis (gkg_kgk​ and ∥gk∥2\|g_k\|^2∥gk​∥2 integrable), so no conditional expectation defaults to Lean's junk value 000. The one-step milestone assumes ∥xk−x∗∥2\|x_k - x^*\|^2∥xk​−x∗∥2 integrable and a bounded stepsize rule at that step, and concludes integrability of ∥xk+1−x∗∥2\|x_{k+1}-x^*\|^2∥xk+1​−x∗∥2.
  • Conditions (i)–(ii) on the random stepsizes are required almost surely. "With probability one lim⁡∥xk−x∗∥=0\lim\|x_k - x^*\| = 0lim∥xk​−x∗∥=0" is ∀ᵐ ω ∂μ, Tendsto (fun k => ‖x k ω - x*‖) atTop (𝓝 0).
  • Division by zero. In Theorems 2.18 and 2.20 the normalized step is undefined when the subgradient vanishes; the formalization skips the step (the iterate is repeated) by an explicit branch, not through Lean's convention x/0=0x/0 = 0x/0=0. When the subgradient never vanishes the sequences are exactly the book's.
  • The printed display (2.42) has xkx_kxk​ where xk+1x_{k+1}xk+1​ is meant on its left-hand side; the corrected inequality is stated.
  • A trivializing formalization, for instance dropping the integrability hypotheses so that the conditional expectations vanish, or quantifying the stepsize conditions so that they cannot hold, is excluded by the hypotheses above; the hypotheses are satisfiable (deterministic subgradients of f(x)=∥x∥f(x) = \|x\|f(x)=∥x∥ with hk=1/(k+1)h_k = 1/(k+1)hk​=1/(k+1)).
  • Mathlib supplies conditional expectation (MeasureTheory.condExp), filtrations, and almost-sure convergence of L1L^1L1-bounded (sub/super)martingales; the supermartingale convergence theorem the book cites from Doob is used from Mathlib, not restated. A Robbins–Siegmund-type lemma for nonnegative almost-supermartingales would be the natural reusable contribution. The almost-differentiability and subgradient definitions duplicate drafts of other missions in this series.

Selected references

  • N. Z. Shor, Minimization Methods for Non-Differentiable Functions, Springer Series in Computational Mathematics 3, Springer, 1985, Section 2.6, pp. 44–47. https://doi.org/10.1007/978-3-642-82118-9
  • Yu. M. Ermoliev and N. Z. Shor, A random search method for two-stage problems of stochastic programming and its generalization, Kibernetika (Kiev), no. 1, 90–92, 1968.
  • L. G. Bazhenov, On the conditions for convergence of methods for minimizing almost differentiable functions, Kibernetika (Kiev), no. 4, 71–72, 1972.
  • M. A. Shepilov, On a method of generalized gradient for finding the absolute minimum of a convex function, Kibernetika (Kiev), no. 4, 52–57, 1976.
  • Yu. M. Ermoliev, Methods of Stochastic Programming, Nauka, Moscow, 1976.
  • H. Robbins and D. Siegmund, A convergence theorem for non negative almost supermartingales and some applications, in Optimizing Methods in Statistics, Academic Press, 1971, pp. 233–257. https://doi.org/10.1016/B978-0-12-604550-5.50015-8
  • J. L. Doob, Stochastic Processes, Wiley, New York, 1953 (supermartingale convergence theorem).
11 thms1 active userReviewed
Operations ResearchOptimization·Captain: mikedeng1

Minimization Methods for Non-Differentiable Functions V: Polyak's Stepsize and Fejér-Type ApproximationsTextbook

Motivation

The subgradient method for a convex function fff moves from xkx_kxk​ against a subgradient gf(xk)g_f(x_k)gf​(xk​), and everything hinges on the step length. Divergent-series stepsizes guarantee convergence but are slow and need no information about fff. When the optimal value f∗f^*f∗, or any level ccc that is known to be attainable, is available, B. T. Polyak proposed in 1969 the step

xk+1=xk−γ [f(xk)−c]∥gf(xk)∥2 gf(xk),x_{k+1} = x_k - \frac{\gamma\,[f(x_k) - c]}{\|g_f(x_k)\|^2}\, g_f(x_k),xk+1​=xk​−∥gf​(xk​)∥2γ[f(xk​)−c]​gf​(xk​),

which uses the current gap f(xk)−cf(x_k) - cf(xk​)−c to scale the move. This Polyak stepsize is still the reference adaptive rule in nonsmooth convex optimization, in the solution of convex feasibility problems, and in the Lagrangian relaxation heuristics of integer programming (Held–Wolfe–Crowder, Camerini–Fratta–Maffioli), where it is known under the name "relaxation step".

Section 2.4 of N. Z. Shor's Minimization Methods for Non-Differentiable Functions (Springer 1985) places Polyak's rule in the framework of Fejér-type approximations developed by I. I. Eremin: an iteration whose map strictly decreases the distance to every point of a target set. It then proves convergence of the rule, linear rates under growth conditions, its behaviour when the level is set too low, and a property of the conjugate-subgradient direction of Camerini, Fratta and Maffioli (1975).

Timeline:

  • 1965–1969: Eremin introduces Fejér mappings for systems of convex inequalities.
  • 1969: Polyak, Minimization of unsmooth functionals, proposes the step with the known optimal value and proves convergence and a linear rate under a sharp-minimum condition.
  • 1975: Camerini, Fratta and Maffioli combine the Polyak step with a conjugate direction for Lagrangian relaxation.
  • 1985: Shor's book collects these results in Section 2.4 (Theorems 2.10–2.16).

Setting

EnE_nEn​ is the nnn-dimensional Euclidean space with inner product (x,y)(x, y)(x,y) and norm ∥x∥\|x\|∥x∥. A vector ggg is a subgradient of f:En→Rf : E_n \to \mathbb{R}f:En​→R at x0x_0x0​ if f(x)−f(x0)≥(g,x−x0)f(x) - f(x_0) \ge (g, x - x_0)f(x)−f(x0​)≥(g,x−x0​) for all xxx. A subgradient selection is a map gfg_fgf​ with gf(x)g_f(x)gf​(x) a subgradient at every xxx; nothing else is assumed about it, in particular not continuity.

For a nonempty set M⊆EnM \subseteq E_nM⊆En​, a map φ:En→En\varphi : E_n \to E_nφ:En​→En​ is MMM-Fejér if φ(y)=y\varphi(y) = yφ(y)=y and ∥φ(x)−y∥<∥x−y∥\|\varphi(x) - y\| < \|x - y\|∥φ(x)−y∥<∥x−y∥ for all y∈My \in My∈M and x∉Mx \notin Mx∈/M.

For a convex fff with f∗=inf⁡ff^* = \inf ff∗=inff and a level c≥f∗c \ge f^*c≥f∗, let M(c)={x:f(x)≤c}M(c) = \{x : f(x) \le c\}M(c)={x:f(x)≤c}. Polyak's method (2.32) is the iteration xk+1=φc(xk)x_{k+1} = \varphi_c(x_k)xk+1​=φc​(xk​) with the map displayed above for x∉M(c)x \notin M(c)x∈/M(c) and φc(y)=y\varphi_c(y) = yφc​(y)=y on M(c)M(c)M(c); the factor γ\gammaγ is fixed in (0,2)(0, 2)(0,2).

The conjugate-subgradient procedure (2.38), for a convex fff with minimum point x∗x^*x∗ and f∗=f(x∗)f^* = f(x^*)f∗=f(x∗), is

xk+1=xk−hksk,hk=[f(xk)−f∗]γk∥sk∥2,s0=gf(x0),sk=gf(xk)+βksk−1.x_{k+1} = x_k - h_k s_k, \quad h_k = \frac{[f(x_k) - f^*]\gamma_k}{\|s_k\|^2}, \qquad s_0 = g_f(x_0),\quad s_k = g_f(x_k) + \beta_k s_{k-1}.xk+1​=xk​−hk​sk​,hk​=∥sk​∥2[f(xk​)−f∗]γk​​,s0​=gf​(x0​),sk​=gf​(xk​)+βk​sk−1​.

Formalization targets

Goal: Theorem 2.11

If 0<γ<20 < \gamma < 20<γ<2 and M(c)≠∅M(c) \neq \emptysetM(c)=∅, then for any x0∈Enx_0 \in E_nx0​∈En​

∃k∗:xk∗∈M(c)orlim⁡k→∞xk exists and lies in M(c).\exists k^* : x_{k^*} \in M(c) \qquad \text{or} \qquad \lim_{k \to \infty} x_k \text{ exists and lies in } M(c).∃k∗:xk∗​∈M(c)ork→∞lim​xk​ exists and lies in M(c).

The goal fixes no constant and no rate; it asserts only that the method finds a point of the level set, in finite time or in the limit.

Milestones

  1. Theorem 2.10. Iterates of a continuous MMM-Fejér map converge to a point of MMM.
  2. Inequality (2.33). For xk∉M(c)x_k \notin M(c)xk​∈/M(c) and y∈M(c)y \in M(c)y∈M(c),
∥xk+1−y∥2≤∥xk−y∥2−γ(2−γ)[f(xk)−c]2∥gf(xk)∥2<∥xk−y∥2.\|x_{k+1} - y\|^2 \le \|x_k - y\|^2 - \gamma(2-\gamma)\frac{[f(x_k) - c]^2}{\|g_f(x_k)\|^2} < \|x_k - y\|^2 .∥xk+1​−y∥2≤∥xk​−y∥2−γ(2−γ)∥gf​(xk​)∥2[f(xk​)−c]2​<∥xk​−y∥2.
  1. Theorem 2.12. Under f(x)−f∗≥m∥x−x∗∥2f(x) - f^* \ge m\|x - x^*\|^2f(x)−f∗≥m∥x−x∗∥2 and an LLL-Lipschitz gradient near x∗x^*x∗, with c=f∗c = f^*c=f∗: ∥xk−x∗∥≤qk∥x0−x∗∥\|x_k - x^*\| \le q^k \|x_0 - x^*\|∥xk​−x∗∥≤qk∥x0​−x∗∥, q=(1−γ(2−γ)m2/L2)1/2<1q = (1 - \gamma(2-\gamma)m^2/L^2)^{1/2} < 1q=(1−γ(2−γ)m2/L2)1/2<1.
  2. Theorem 2.13. Under the sharp-minimum condition f(x)−f(x∗)≥m∥x−x∗∥f(x) - f(x^*) \ge m\|x - x^*\|f(x)−f(x∗)≥m∥x−x∗∥ and subgradients bounded by LLL near x∗x^*x∗, with c=f(x∗)c = f(x^*)c=f(x∗): ∥xk+1−x∗∥≤q∥xk−x∗∥\|x_{k+1} - x^*\| \le q\|x_k - x^*\|∥xk+1​−x∗∥≤q∥xk​−x∗∥.
  3. Theorem 2.14. If min⁡ψ=d>0\min \psi = d > 0minψ=d>0 and the method runs with c=0c = 0c=0, then lim⁡kmin⁡0≤i≤kψ(xi)≤2d/(2−γ)\lim_k \min_{0 \le i \le k} \psi(x_i) \le 2d/(2-\gamma)limk​min0≤i≤k​ψ(xi​)≤2d/(2−γ).
  4. Theorem 2.15. For (2.38) with 0<γk≤10 < \gamma_k \le 10<γk​≤1, βk≥0\beta_k \ge 0βk​≥0: (xk−x∗,sk)≥(xk−x∗,gf(xk))(x_k - x^*, s_k) \ge (x_k - x^*, g_f(x_k))(xk​−x∗,sk​)≥(xk​−x∗,gf​(xk​)).
  5. Theorem 2.16. With the Camerini–Fratta–Maffioli coefficient βk\beta_kβk​ and 0≤αk≤20 \le \alpha_k \le 20≤αk​≤2: (xk−x∗,sk)/∥sk∥≥(xk−x∗,gf(xk))/∥gf(xk)∥(x_k - x^*, s_k)/\|s_k\| \ge (x_k - x^*, g_f(x_k))/\|g_f(x_k)\|(xk​−x∗,sk​)/∥sk​∥≥(xk​−x∗,gf​(xk​))/∥gf​(xk​)∥.

Significance

Theorem 2.11 is the convergence guarantee of the most widely used adaptive step rule for nonsmooth convex problems. With c=f∗c = f^*c=f∗ it yields a minimizer; with c>f∗c > f^*c>f∗ it solves the convex inequality f(x)≤cf(x) \le cf(x)≤c, and applied to ψ=max⁡ifi+\psi = \max_i f_i^+ψ=maxi​fi+​ it solves consistent systems of convex inequalities. The linear rates of Theorems 2.12–2.13 are the prototype of the "sharpness implies linear convergence" results of modern first-order methods, and Theorem 2.14 quantifies the loss when the level is underestimated, which is the situation of every practical variant that estimates f∗f^*f∗ on the fly. Theorems 2.15–2.16 are the justification of the conjugate-subgradient directions used in Lagrangian relaxation.

All results are classical and proved on paper. None of them is formalized on Prove2Me: the platform has a smooth, strongly convex Polyak gradient-descent bound (a different theorem) and Fejér-monotonicity statements for polyhedral relaxation methods, but no Polyak subgradient step, no MMM-Fejér map and no conjugate-subgradient procedure. The mission produces machine-checked versions of the whole section, with the page's misprints corrected where the proof and the statement disagree.

Difficulty

The obvious argument for the goal is to observe that φc\varphi_cφc​ is M(c)M(c)M(c)-Fejér, by (2.33), and invoke Theorem 2.10. That argument fails: Theorem 2.10 needs a continuous map, and φc\varphi_cφc​ depends on an arbitrary subgradient selection, which is discontinuous wherever fff is not differentiable. The book says so explicitly. Fejér monotonicity gives boundedness and a limit of each distance ∥xk−y∥\|x_k - y\|∥xk​−y∥, but convergence of the whole sequence to a single point of M(c)M(c)M(c), and the fact that an accumulation point cannot lie outside M(c)M(c)M(c), have to be obtained without continuity of the map.

For Theorem 2.14 the level c=0c = 0c=0 lies strictly below the minimum, so M(0)=∅M(0) = \emptysetM(0)=∅, the target set of the iteration as run is empty, no Fejér property is available for it, and the theorem controls only the best value found, not the iterates.

Formalization scope

  • EnE_nEn​ is EuclideanSpace ℝ (Fin n); fff is real-valued on all of EnE_nEn​ and ConvexOn ℝ Set.univ f.
  • The subgradient selection is universally quantified; no theorem assumes continuity of it.
  • Iterations are sequences x : ℕ → E_n with the recursion as a hypothesis; the first term is arbitrary.
  • Polyak's step map is defined piecewise: it returns xxx on M(c)M(c)M(c) (as the book sets φc(y)=y\varphi_c(y) = yφc​(y)=y) and at gf(x)=0g_f(x) = 0gf​(x)=0. No statement relies on Lean's convention x/0=0x/0 = 0x/0=0. The same holds for hkh_khk​ when sk=0s_k = 0sk​=0.
  • M(c)≠∅M(c) \neq \emptysetM(c)=∅ is a hypothesis of the goal: the book's proof picks y∈M(c)y \in M(c)y∈M(c), and for c=f∗c = f^*c=f∗ not attained the conclusion is false (for f=exf = e^xf=ex, c=0c = 0c=0, the method moves by γ\gammaγ each step and diverges).
  • "lim⁡xk∈M(c)\lim x_k \in M(c)limxk​∈M(c)" is the existence of a limit in M(c)M(c)M(c), not a statement about cluster points.
  • γ∈(0,2)\gamma \in (0, 2)γ∈(0,2) is stated in every theorem on Polyak's method; the book fixes this range at Theorem 2.11.
  • Theorem 2.12's "strongly convex" is used through its displayed growth condition only; the statement is made for convex fff satisfying it, with L>0L > 0L>0 and qqq computed with Real.sqrt.
  • Theorem 2.13 assumes the bound ∥g∥≤L\|g\| \le L∥g∥≤L on subgradients in the ball, which is what the proof uses; a Lipschitz constant on the closed ball alone does not give it, and the printed statement fails without it.
  • Theorem 2.14 is stated with "≤2d/(2−γ)\le 2d/(2-\gamma)≤2d/(2−γ)"; the printed "===" is false in general.
  • Theorem 2.16's inequality is stated at indices where sk≠0s_k \neq 0sk​=0 and gf(xk)≠0g_f(x_k) \neq 0gf​(xk​)=0.

A formalization in which M(c)M(c)M(c) may be empty, the selection is assumed continuous, or the step divides by zero through Lean's conventions would be a different theorem; these are ruled out above.

Needed infrastructure: Fejér monotone sequences in finite dimensions (bounded, with convergent distances), the subgradient inequality, and the fact that a zero subgradient characterizes a minimum. These are reusable for every subgradient-type method. Contributions of general lemmas on Fejér-monotone sequences are welcome.

Selected references

  • N. Z. Shor, Minimization Methods for Non-Differentiable Functions, Springer Series in Computational Mathematics 3, Springer, 1985, §2.4, pp. 36–42. https://doi.org/10.1007/978-3-642-82118-9
  • B. T. Polyak, Minimization of unsmooth functionals, USSR Computational Mathematics and Mathematical Physics 9(3), 1969, 14–29. https://doi.org/10.1016/0041-5553(69)90061-5
  • I. I. Eremin, The relaxation method of solving systems of inequalities with convex functions on the left-hand side, Soviet Mathematics Doklady 6, 1965, 219–222.
  • P. M. Camerini, L. Fratta, F. Maffioli, On improving relaxation methods by modified gradient techniques, Mathematical Programming Study 3, 1975, 26–34.
10 thms1 active userReviewed
Operations ResearchOptimization·Captain: mikedeng1

Minimization Methods for Non-Differentiable Functions IV: Linear Convergence of the Subgradient Method under Level-Set Shape ConditionsTextbook

Motivation

The subgradient method minimizes a convex function fff on Rn\mathbb{R}^nRn that need not be differentiable, by stepping against an arbitrary subgradient. With stepsizes hk→0h_k \to 0hk​→0, ∑hk=∞\sum h_k = \infty∑hk​=∞ it converges (Shor, Theorem 2.2), but in general only slowly: no stepsize rule that ignores the structure of fff gives a geometric rate. Section 2.3 of N. Z. Shor's Minimization Methods for Non-Differentiable Functions (Springer 1985) identifies geometric conditions on fff under which a simple geometric stepsize rule does give linear convergence — convergence with the speed of a geometric progression — and computes the rate explicitly.

The results are the origin of what is now studied as sharpness or error-bound conditions for nonsmooth optimization. Their practical content is that nonsmooth problems whose level sets are not too elongated near the minimum (piecewise-linear functions, maxima of finitely many well-conditioned pieces, positive definite quadratics) can be solved by the subgradient method at a linear rate, with a stepsize rule that needs only one or two scalar parameters.

Timeline. Shor proposed the subgradient method in 1962. The book presents the geometric stepsize rule under an angle condition (Theorem 2.7) and its level-surface form (Theorem 2.8), and attributes the block-halving rule of Theorem 2.9 to its reference [94]. Goffin (Math. Programming 13, 1977) gave the sharp rate in terms of a condition number of the level sets. The book compares the quadratic case with L. V. Kantorovich's rate for steepest descent.

Setting

Let EnE_nEn​ be the nnn-dimensional Euclidean space with inner product (x,y)(x, y)(x,y) and f:En→Rf : E_n \to \mathbb{R}f:En​→R convex. A vector ggg is a subgradient of fff at xxx if f(y)−f(x)≥(g,y−x)f(y) - f(x) \ge (g, y - x)f(y)−f(x)≥(g,y−x) for all yyy; gf(x)g_f(x)gf​(x) denotes an arbitrary subgradient at xxx, chosen once for each xxx. Let M∗M^*M∗ be the set of minimum points of fff, assumed nonempty; for x∈Enx \in E_nx∈En​, x∗(x)x^*(x)x∗(x) is the point of M∗M^*M∗ nearest to xxx.

Given a starting point x0x_0x0​ and positive stepsizes h1,h2,…h_1, h_2, \dotsh1​,h2​,…, the normalized subgradient method is

xk+1=xk−hk+1gf(xk)∥gf(xk)∥,k=0,1,2,…,x_{k+1} = x_k - h_{k+1}\frac{g_f(x_k)}{\|g_f(x_k)\|}, \qquad k = 0, 1, 2, \dots,xk+1​=xk​−hk+1​∥gf​(xk​)∥gf​(xk​)​,k=0,1,2,…,

stopped when gf(xk)=0g_f(x_k) = 0gf​(xk​)=0 (then xk∈M∗x_k \in M^*xk​∈M∗).

Two shape conditions are used. The angle condition (2.12) with angle 0≤φ<π/20 \le \varphi < \pi/20≤φ<π/2 asks that every subgradient make an angle at most φ\varphiφ with the direction to the nearest minimum point:

(gf(x),x−x∗(x))≥cos⁡φ ∥gf(x)∥ ∥x−x∗(x)∥.(g_f(x), x - x^*(x)) \ge \cos\varphi\,\|g_f(x)\|\,\|x - x^*(x)\|.(gf​(x),x−x∗(x))≥cosφ∥gf​(x)∥∥x−x∗(x)∥.

The level-surface ratio condition (2.20), for a function with unique minimum point x∗x^*x∗, asks that on a ball YYY around x∗x^*x∗ any two points x,zx, zx,z on a common level surface f(x)=f(z)≠f(x∗)f(x) = f(z) \ne f(x^*)f(x)=f(z)=f(x∗) satisfy ∥x−x∗∥≤σ∥z−x∗∥\|x - x^*\| \le \sigma\|z - x^*\|∥x−x∗∥≤σ∥z−x∗∥.

Formalization targets

Goal: Theorem 2.8 (p. 32)

If fff has a unique minimum point x∗x^*x∗, σ≥2\sigma \ge \sqrt2σ≥2​, h1≥∥x0−x∗∥/σh_1 \ge \|x_0 - x^*\|/\sigmah1​≥∥x0​−x∗∥/σ, and (2.20) holds on Y={y:∥y−x∗∥≤σh1}Y = \{y : \|y - x^*\| \le \sigma h_1\}Y={y:∥y−x∗∥≤σh1​}, then with hk+1=hkσ2−1/σh_{k+1} = h_k\sqrt{\sigma^2-1}/\sigmahk+1​=hk​σ2−1​/σ

∥xk−x∗∥≤hk+1 σ,k=0,1,2,…\|x_k - x^*\| \le h_{k+1}\,\sigma, \qquad k = 0, 1, 2, \dots∥xk​−x∗∥≤hk+1​σ,k=0,1,2,…

Milestones

  1. Theorem 2.7 (pp. 30–31): under (2.12) and the geometric rule hk+1=hkr(φ)h_{k+1} = h_k r(\varphi)hk+1​=hk​r(φ) with r(φ)=sin⁡φr(\varphi) = \sin\varphir(φ)=sinφ for φ≥π/4\varphi \ge \pi/4φ≥π/4 and r(φ)=1/(2cos⁡φ)r(\varphi) = 1/(2\cos\varphi)r(φ)=1/(2cosφ) for φ<π/4\varphi < \pi/4φ<π/4,
∥xk−x∗(xk)∥≤hk+1/cos⁡φresp.2hk+1cos⁡φ.\|x_k - x^*(x_k)\| \le h_{k+1}/\cos\varphi \quad\text{resp.}\quad 2h_{k+1}\cos\varphi .∥xk​−x∗(xk​)∥≤hk+1​/cosφresp.2hk+1​cosφ.
  1. Remark after Theorem 2.7 (p. 32): the same conclusion when (2.12) holds only at the iterates.
  2. Inequality (2.22) (p. 33): (2.20) on YYY implies (g,x−x∗)≥σ−1∥g∥ ∥x−x∗∥(g, x - x^*) \ge \sigma^{-1}\|g\|\,\|x - x^*\|(g,x−x∗)≥σ−1∥g∥∥x−x∗∥ for every x∈Yx \in Yx∈Y and every subgradient ggg at xxx.
  3. Example (p. 33): for AAA symmetric positive definite with extreme eigenvalues λ≤μ\lambda \le \muλ≤μ,
min⁡x≠0(Ax,x)∥Ax∥ ∥x∥=2λμλ+μ,\min_{x \ne 0}\frac{(Ax, x)}{\|Ax\|\,\|x\|} = \frac{2\sqrt{\lambda\mu}}{\lambda + \mu},x=0min​∥Ax∥∥x∥(Ax,x)​=λ+μ2λμ​​,

attained at x=μ/(λ+μ) s1+λ/(λ+μ) s2x = \sqrt{\mu/(\lambda+\mu)}\,s_1 + \sqrt{\lambda/(\lambda+\mu)}\,s_2x=μ/(λ+μ)​s1​+λ/(λ+μ)​s2​. 5. Theorem 2.9 (p. 34): under the assumptions of Theorem 2.8 with σ≥2\sigma \ge 2σ≥2, the rule hk+1=h02−[(k+1)/N]h_{k+1} = h_0 2^{-[(k+1)/N]}hk+1​=h0​2−[(k+1)/N] with N≥3σ2+1N \ge 3\sigma^2 + 1N≥3σ2+1 gives ∥xk−x∗∥≤2σhk+1\|x_k - x^*\| \le 2\sigma h_{k+1}∥xk​−x∗∥≤2σhk+1​.

Significance

The results. Theorem 2.7 shows that the subgradient method, often dismissed as sublinear, converges linearly once the geometry of fff is controlled and the stepsizes decrease geometrically at the right ratio; the rate r(φ)r(\varphi)r(φ) depends only on the angle. Theorem 2.8 restates the hypothesis in terms of the shape of level surfaces, a condition that can be checked for concrete functions, and gives rate σ2−1/σ\sqrt{\sigma^2-1}/\sigmaσ2−1​/σ. The Example computes the angle for positive definite quadratics, yielding rate (ϱ−1)/(ϱ+1)(\varrho - 1)/(\varrho + 1)(ϱ−1)/(ϱ+1) with ϱ=μ/λ\varrho = \mu/\lambdaϱ=μ/λ the condition number — the same rate as steepest descent with exact line search in Kantorovich's analysis, obtained with less storage. Theorem 2.9 removes the need to know σ\sigmaσ exactly in the stepsize ratio.

Formalizing them. All five results are proved in the book; none is formalized, and no linear-rate result for a nonsmooth first-order method is on the platform. The formalization pins the constants (2.13)–(2.19), the stepsize indexing, and the treatment of the stopped iteration; the Example is a Kantorovich-type inequality for symmetric operators that is reusable beyond this mission.

Difficulty

The obvious one-step estimate ∥xk+1−x∗∥2=∥xk−x∗∥2−2hk+1(g,xk−x∗)/∥g∥+hk+12\|x_{k+1} - x^*\|^2 = \|x_k - x^*\|^2 - 2h_{k+1}(g, x_k - x^*)/\|g\| + h_{k+1}^2∥xk+1​−x∗∥2=∥xk​−x∗∥2−2hk+1​(g,xk​−x∗)/∥g∥+hk+12​ alone does not contract: the step length is fixed in advance and does not shrink with the distance, so a step may overshoot the minimum. The rate argument has to track the ratio between the current distance and the current stepsize, and the admissible ratio of stepsizes is dictated by the worst case of this quadratic in the distance; in the two regimes φ≥π/4\varphi \ge \pi/4φ≥π/4 and φ<π/4\varphi < \pi/4φ<π/4 the worst case sits at different ends. For Theorem 2.8, the shape condition is only assumed on the ball YYY, so the iterates must be shown to stay in YYY, and the passage from level surfaces to subgradients needs the distance from x∗x^*x∗ to a level surface, which is not a quantity the iteration computes. Theorem 2.9's constant stepsize blocks are not monotone in distance at all, and the count 3σ2+13\sigma^2 + 13σ2+1 must be matched against a worst-case phase.

Formalization scope

The space EnE_nEn​ is EuclideanSpace ℝ (Fin n); fff is real-valued and ConvexOn ℝ Set.univ. A subgradient selection is an arbitrary function g with g x a subgradient at every x, quantified universally. Stepsizes are h : ℕ → ℝ with h (k+1) used at step k; the recursions hk+1=hkrh_{k+1} = h_k rhk+1​=hk​r are imposed for k≥1k \ge 1k≥1, h1h_1h1​ being the chosen initial step (the book's "k=0,1,2,…k = 0, 1, 2, \dotsk=0,1,2,…" in Theorem 2.8 is read this way). The iteration stops at gf(xk)=0g_f(x_k) = 0gf​(xk​)=0 by an explicit branch that repeats xkx_kxk​, never through x/0=0x/0 = 0x/0=0; the bounds are asserted for every kkk, which implies the book's "either the method stops or …" form. x∗(x)x^*(x)x∗(x) is a definition (the nearest point of the set of minima), and the set of minima is assumed nonempty. Unique minimum is stated as M∗={x∗}M^* = \{x^*\}M∗={x∗}. The ball YYY is closed, and (2.20) is assumed only for pairs in YYY with a common value different from f(x∗)f(x^*)f(x∗). In (2.16) the ratio is 1/(2cos⁡φ)1/(2\cos\varphi)1/(2cosφ), as the proof requires, where the page prints "1/2 cos φ". In Theorem 2.9, "the assumptions of Theorem 2.8" are taken with h1=h0h_1 = h_0h1​=h0​, which fixes h0≥∥x0−x∗∥/σh_0 \ge \|x_0 - x^*\|/\sigmah0​≥∥x0​−x∗∥/σ and YYY of radius σh0\sigma h_0σh0​. In the Example, the extreme eigenvalues are pinned by λ∥x∥2≤(Ax,x)≤μ∥x∥2\lambda\|x\|^2 \le (Ax,x) \le \mu\|x\|^2λ∥x∥2≤(Ax,x)≤μ∥x∥2 together with unit eigenvectors, and the minimum is stated with IsLeast.

A statement in which the stepsizes or the bound constants could be chosen after the iterates, or in which the shape condition quantified over an empty set of pairs, would be trivially true; here every constant is fixed by the hypotheses before the sequence is generated, and the conditions are the book's.

Needed infrastructure: the subgradient inequality, continuity of convex functions on Rn\mathbb{R}^nRn, nearest points of closed convex sets, and elementary trigonometry. The nearest-point map and the stepsize ratio are defined within the mission; the subgradient inequality, the set of minima and the iteration come from the series' shared definitions. Contributions welcome: proofs of the milestones, and a reusable Kantorovich-type cosine bound for symmetric positive definite operators.

Selected references

  • N. Z. Shor, Minimization Methods for Non-Differentiable Functions, Springer Series in Computational Mathematics 3, Springer 1985, §2.3, pp. 30–36. https://doi.org/10.1007/978-3-642-82118-9
  • J.-L. Goffin, On convergence rates of subgradient optimization methods, Mathematical Programming 13 (1977) 329–347. https://doi.org/10.1007/BF01584346
  • L. V. Kantorovich, Functional analysis and applied mathematics, Uspekhi Mat. Nauk 3 (1948) 89–185 (steepest descent rate for quadratics, cited by Shor as [45]).
9 thms1 active userReviewed
Operations ResearchOptimizationProbability·Captain: mikedeng1

Numerical Techniques for Stochastic Optimization VI: Adaptive Stepsizes and Cesàro Convergence of Stochastic Quasigradient MethodsTextbook

Motivation

Stochastic quasigradient (SQG) methods minimize an expectation F(x)=Eωf(x,ω)F(x)=E_\omega f(x,\omega)F(x)=Eω​f(x,ω) over a constraint set X⊆RnX\subseteq\mathbb R^nX⊆Rn when neither FFF nor its gradient can be computed, only random vectors whose conditional mean is (close to) a subgradient. They are the workhorse of stochastic programming and, under the name stochastic gradient descent, of large-scale statistical learning. The classical convergence theory, going back to Robbins and Monro (1951) and to Ermoliev's quasi-Féjer analysis, asks the stepsizes to be chosen in advance with ρs→0\rho_s\to0ρs​→0, ∑ρs=∞\sum\rho_s=\infty∑ρs​=∞, ∑ρs2<∞\sum\rho_s^2<\infty∑ρs2​<∞. Uryasev, in Chapter 18 of Numerical Techniques for Stochastic Optimization (Ermoliev and Wets, eds., 1988), points out that such programmed rules are slow in practice, and that practitioners want adaptive stepsizes computed on line from the observed directions.

Timeline:

  • 1951: Robbins and Monro, stochastic approximation with programmed steps.
  • 1976: Ermoliev, Methods of Stochastic Programming: the SQG projection method and its a.s. convergence through stochastic quasi-Féjer sequences.
  • 1983: Mirzoakhmedov and Uryasev (Zh. Vychisl. Mat. i Mat. Fiz., cited as [7] in Ch. 18 and [14] in Ch. 17): Cesàro convergence of the weighted mean with ρs→0\rho_s\to0ρs​→0 and ∑ρs=∞\sum\rho_s=\infty∑ρs​=∞ only, under the two measurability regimes. Chapter 17 states it as Theorem (ii); Chapter 18 as Theorem 1.
  • 1988: Uryasev, Ch. 18, applies it to the adaptive rule (18.5) (Theorem 2).
  • 1992: Polyak and Juditsky, averaging of iterates for smooth stochastic approximation, with optimal asymptotic variance.

Setting

Let X⊆RnX\subseteq\mathbb R^nX⊆Rn be nonempty, convex and compact, C1=max⁡x,y∈X∥x−y∥C_1=\max_{x,y\in X}\|x-y\|C1​=maxx,y∈X​∥x−y∥ its diameter, and FFF convex on an open convex set U⊇XU\supseteq XU⊇X, with subdifferential ∂F(x)\partial F(x)∂F(x). The projection πX(y)\pi_X(y)πX​(y) is the point of XXX nearest to yyy. On a probability space, the SQG method generates

xs+1=πX(xs−ρsξs),s=0,1,…(18.2)x^{s+1}=\pi_X(x^s-\rho_s\xi^s),\qquad s=0,1,\dots\qquad(18.2)xs+1=πX​(xs−ρs​ξs),s=0,1,…(18.2)

from x0∈Xx^0\in Xx0∈X, where the direction ξs\xi^sξs is a stochastic quasigradient: E(ξs∣Bs)=Fx(xs)+bsE(\xi^s\mid B_s)=F_x(x^s)+b^sE(ξs∣Bs​)=Fx​(xs)+bs with Fx(xs)∈∂F(xs)F_x(x^s)\in\partial F(x^s)Fx​(xs)∈∂F(xs), a bias bsb^sbs, and BsB_sBs​ the σ\sigmaσ-algebra induced by (x0,…,xs,ξ0,…,ξs−1)(x^0,\dots,x^s,\xi^0,\dots,\xi^{s-1})(x0,…,xs,ξ0,…,ξs−1).

The adaptive stepsize rule of the chapter is, for fixed a>1a>1a>1, δ>0\delta>0δ>0 and ρ0>0\rho_0>0ρ0​>0,

ρs+1=ρs a⟨ξs+1, xs−xs+1⟩−δρs(18.5).\rho_{s+1}=\rho_s\,a^{\langle\xi^{s+1},\,x^s-x^{s+1}\rangle-\delta\rho_s}\qquad(18.5).ρs+1​=ρs​a⟨ξs+1,xs−xs+1⟩−δρs​(18.5).

The step grows when consecutive moves point the same way and shrinks otherwise. The weighted (Cesàro) averages are

xˉs=∑ℓ=0sρℓxℓ/∑ℓ=0sρℓ(18.6).\bar x^s=\sum_{\ell=0}^s\rho_\ell x^\ell\Big/\sum_{\ell=0}^s\rho_\ell\qquad(18.6).xˉs=ℓ=0∑s​ρℓ​xℓ/ℓ=0∑s​ρℓ​(18.6).

The sequence xsx^sxs is Cesàro convergent when xˉs\bar x^sxˉs converges to the solution set.

Chapter 17 (Pflug) uses the same method for f(x)=EP q(x,ξ)f(x)=E_P\,q(x,\xi)f(x)=EP​q(x,ξ) over a closed convex S⊆RkS\subseteq\mathbb R^kS⊆Rk, with Y=∇q(Xn,ξn)Y=\nabla q(X_n,\xi_n)Y=∇q(Xn​,ξn​) from i.i.d. ξn\xi_nξn​ and stepsizes adapted to σ(ξ0,…,ξn−1)\sigma(\xi_0,\dots,\xi_{n-1})σ(ξ0​,…,ξn−1​).

Formalization targets

Goal: Theorem 2 of Chapter 18

Under sup⁡s∥ξs∥<C2\sup_s\|\xi^s\|<C_2sups​∥ξs∥<C2​ (18.15), lim sup⁡∥bs∥≤bˉ\limsup\|b^s\|\le\bar blimsup∥bs∥≤bˉ (18.16) and δ>C2lim sup⁡sinf⁡h∈∂F(xs)∥ξs−h∥\delta>C_2\limsup_s\inf_{h\in\partial F(x^s)}\|\xi^s-h\|δ>C2​limsups​infh∈∂F(xs)​∥ξs−h∥ (18.17), almost surely,

lim sup⁡s→∞(F(xˉs)−min⁡x∈XF(x))≤bˉ C1,\limsup_{s\to\infty}\Big(F(\bar x^s)-\min_{x\in X}F(x)\Big)\le\bar b\,C_1,s→∞limsup​(F(xˉs)−x∈Xmin​F(x))≤bˉC1​,

and if bs→0b^s\to0bs→0 a.s., then F(xˉs)→min⁡XFF(\bar x^s)\to\min_XFF(xˉs)→minX​F and all accumulation points of xˉs\bar x^sxˉs are minimizers, almost surely.

Milestones

  1. Chapter 17, Theorem (i): ∑ρn=∞\sum\rho_n=\infty∑ρn​=∞ and ∑ρn2<∞\sum\rho_n^2<\infty∑ρn2​<∞ a.s. imply Xn→x∗X_n\to x^*Xn​→x∗ a.s.
  2. Chapter 17, Theorem (ii): for convex fff and bounded SSS, ρn→0\rho_n\to0ρn​→0 and ∑ρn=∞\sum\rho_n=\infty∑ρn​=∞ a.s. imply Xˉn→x∗\bar X_n\to x^*Xˉn​→x∗ a.s.
  3. Chapter 18, Theorem 1: for any stepsizes with ρs>0\rho_s>0ρs​>0, Eρs2<∞E\rho_s^2<\inftyEρs2​<∞, ρs→0\rho_s\to0ρs​→0, ∑ρs=∞\sum\rho_s=\infty∑ρs​=∞ and measurability condition (1) or (2), lim sup⁡F(xˉs)−F(x∗)≤bˉC1\limsup F(\bar x^s)-F(x^*)\le\bar bC_1limsupF(xˉs)−F(x∗)≤bˉC1​ a.s.
  4. Chapter 18, Corollary: with bs→0b^s\to0bs→0, the accumulation points of xˉs\bar x^sxˉs are solutions.
  5. Eq. (18.18): ∥xs+1−xs∥≤∥ρsξs∥≤ρsC2\|x^{s+1}-x^s\|\le\|\rho_s\xi^s\|\le\rho_sC_2∥xs+1−xs∥≤∥ρs​ξs∥≤ρs​C2​.
  6. Proof of Theorem 2, step 1: the adaptive steps satisfy ∑ρs=∞\sum\rho_s=\infty∑ρs​=∞.
  7. Proof of Theorem 2, step 2: under (18.17), ρs→0\rho_s\to0ρs​→0.
  8. End of step 2: ρs→0\rho_s\to0ρs​→0 implies ρs+1/ρs→1\rho_{s+1}/\rho_s\to1ρs+1​/ρs​→1.

Significance

Theorem 2 is a convergence guarantee for a stepsize rule that is computed from the run itself. It needs no square summability of the steps, and it tolerates a nonvanishing bias at a cost linear in the bias. This is the regime of practical SQG codes; §18.4–18.5 of the chapter discuss implementation and numerical experiments. Theorem 1 isolates the reason: Cesàro convergence needs only ρs→0\rho_s\to0ρs​→0 and ∑ρs=∞\sum\rho_s=\infty∑ρs​=∞. It also allows a stepsize that depends on the current direction, provided consecutive steps have ratio tending to 111.

The volume proves none of the probabilistic results in full. Theorem 1 of Chapter 18 is cited from Uryasev's earlier report. Theorem 2 has an outline proof that reduces it to Theorem 1. Chapter 17 gives a sketch through the Robbins–Siegmund lemma. None of these results is formalized. The mission produces machine-checked statements of all of them, with the misprints of the page resolved, and it separates the pathwise part of the Theorem 2 argument (steps 1 and 2, which are deterministic) from the martingale part (Theorem 1).

Difficulty

The obvious route to a.s. convergence is the quasi-Féjer or Robbins–Siegmund argument. It controls ∥xs−x∗∥2\|x^s-x^*\|^2∥xs−x∗∥2 and needs ∑ρs2∥ξs∥2<∞\sum\rho_s^2\|\xi^s\|^2<\infty∑ρs2​∥ξs∥2<∞, which is exactly what is not available here. The averaged analysis has to show that the martingale term ∑ℓρℓ⟨ξℓ−E(ξℓ∣Bℓ),x∗−xℓ⟩\sum_\ell\rho_\ell\langle\xi^\ell-E(\xi^\ell\mid B_\ell),x^*-x^\ell\rangle∑ℓ​ρℓ​⟨ξℓ−E(ξℓ∣Bℓ​),x∗−xℓ⟩ is o(∑ℓρℓ)o(\sum_\ell\rho_\ell)o(∑ℓ​ρℓ​) almost surely, and that ∑ℓρℓ2∥ξℓ∥2\sum_\ell\rho_\ell^2\|\xi^\ell\|^2∑ℓ​ρℓ2​∥ξℓ∥2 is o(∑ℓρℓ)o(\sum_\ell\rho_\ell)o(∑ℓ​ρℓ​), when the stepsizes are themselves random. Under condition (2) of Theorem 1, ρs\rho_sρs​ is not even measurable with respect to the σ\sigmaσ-algebra of the conditional expectation. So E(ρsξs∣Bs)≠ρsE(ξs∣Bs)E(\rho_s\xi^s\mid B_s)\ne\rho_sE(\xi^s\mid B_s)E(ρs​ξs∣Bs​)=ρs​E(ξs∣Bs​), and the standard decomposition breaks. For the adaptive rule, the stepsizes are coupled to the iterates through the exponent. Neither ∑ρs=∞\sum\rho_s=\infty∑ρs​=∞ nor ρs→0\rho_s\to0ρs​→0 is given, and both must be derived path by path.

Formalization scope

Rn\mathbb R^nRn is EuclideanSpace ℝ (Fin n). Sequences are indexed from 000. Chapter 17 is shifted by one against the page: its Xn,ξn,FnX_n,\xi_n,\mathcal F_nXn​,ξn​,Fn​, n≥1n\ge1n≥1, become indices n−1n-1n−1. Conditional expectations are Mathlib's condExp with respect to the history σ\sigmaσ-algebras of the definition file. Every lim sup⁡\limsuplimsup bound is written out as "for every ε>0\varepsilon>0ε>0, eventually ⋯≤⋯+ε\dots\le\dots+\varepsilon⋯≤⋯+ε", or in (18.17) as a bound LLL with C2L<δC_2L<\deltaC2​L<δ. Expectations of squared norms are lower Lebesgue integrals. The deterministic proof steps (items 5 to 8) are stated for one sample path.

Readings of the page, each recorded in the item's Formalization Note:

  • (18.17) prints C1C_1C1​. The proof's estimate gives (C2Cs−δ)ρs(C_2C_s-\delta)\rho_s(C2​Cs​−δ)ρs​, and only C2C_2C2​ is invariant under rescaling of Rn\mathbb R^nRn, so C2C_2C2​ is stated.
  • (18.5) has two forms that agree only without projection. The proof uses the second, a⟨ξs+1,xs−xs+1⟩−δρsa^{\langle\xi^{s+1},x^s-x^{s+1}\rangle-\delta\rho_s}a⟨ξs+1,xs−xs+1⟩−δρs​, which is stated.
  • (18.8) prints Fs(xs)F_s(x^s)Fs​(xs) for Fx(xs)F_x(x^s)Fx​(xs). (18.11) prints EρssE\rho_s^sEρss​, read as Eρs2<∞E\rho_s^2<\inftyEρs2​<∞.
  • Theorem 2's "F(xs)−min⁡z∈XF(x)→0F(x^s)-\min z\in XF(x)\to0F(xs)−minz∈XF(x)→0" is read as F(xˉs)−min⁡XF→0F(\bar x^s)-\min_XF\to0F(xˉs)−minX​F→0.
  • The end of step 2 prints ρs+1/ρs→0\rho_{s+1}/\rho_s\to0ρs+1​/ρs​→0, read as →1\to1→1.
  • Chapter 17, assumption (ii) prints ∥∇f(x)∥≤A+B∥x−x∗∥2\|\nabla f(x)\|\le A+B\|x-x^*\|^2∥∇f(x)∥≤A+B∥x−x∗∥2. The proof uses ∥∇f(x)∥2\|\nabla f(x)\|^2∥∇f(x)∥2, and the printed form makes part (i) false, so the squared form is stated. Var(Yx)≤C\mathrm{Var}(Y_x)\le CVar(Yx​)≤C is read as E∥Yx−EYx∥2≤CE\|Y_x-EY_x\|^2\le CE∥Yx​−EYx​∥2≤C.
  • The Corollary adds lower semicontinuity of FFF on XXX, without which it fails.
  • x0∈Xx^0\in Xx0∈X is assumed, and ρ0\rho_0ρ0​ in Theorem 2 is a fixed positive number.

No explicit constants replace an O(·) or an unspecified "C": every constant appears in the book's statements.

A trivializing formalization states Theorem 2 for arbitrary stepsizes satisfying (18.10)–(18.13), which is Theorem 1 again. Here the stepsizes are tied to the iterates by (18.5), and the δ\deltaδ of (18.17) is the δ\deltaδ of the rule.

Needed infrastructure: a Robbins–Siegmund almost-supermartingale lemma, which Mathlib does not have; a strong law for martingale differences with random weights (Kronecker's lemma in its stochastic form); nonexpansiveness of the projection onto a closed convex set; and nonemptiness of the subdifferential of a finite convex function on an open set. The first two are reusable across stochastic approximation. Contributions of any of the milestones, or of these lemmas as separate theorems, are welcome.

Selected references

  • G. Ch. Pflug, Stepsize Rules, Stopping Times and their Implementation in Stochastic Quasigradient Algorithms, in Yu. Ermoliev and R. J-B Wets (eds.), Numerical Techniques for Stochastic Optimization, Springer 1988, Ch. 17. https://doi.org/10.1007/978-3-642-61370-8
  • S. Uryasev, Adaptive Stochastic Quasigradient Procedures, ibid., Ch. 18. https://doi.org/10.1007/978-3-642-61370-8
  • Yu. Ermoliev, Stochastic Quasigradient Methods, ibid., Ch. 6. https://doi.org/10.1007/978-3-642-61370-8
  • F. Mirzoakhmedov and S. P. Uryasev, Adaptive step size control for stochastic optimization algorithm, Zh. Vychisl. Mat. i Mat. Fiz. 23(6) (1983) 1314–1325 (in Russian); cited in the volume above, no online copy linked.
  • H. Robbins and S. Monro, A Stochastic Approximation Method, Ann. Math. Statist. 22 (1951) 400–407. https://doi.org/10.1214/aoms/1177729586
  • H. Robbins and D. Siegmund, A convergence theorem for non negative almost supermartingales and some applications, in Optimizing Methods in Statistics, Academic Press 1971, 233–257. https://doi.org/10.1016/B978-0-12-604550-5.50015-8
  • B. T. Polyak and A. B. Juditsky, Acceleration of Stochastic Approximation by Averaging, SIAM J. Control Optim. 30 (1992) 838–855. https://doi.org/10.1137/0330046
11 thms1 active userReviewed
Operations ResearchOptimizationProbability·Captain: mikedeng1

Introduction to the Scenario Approach II: Violation Guarantees after Discarding k ConstraintsTextbook

Motivation

Decisions under uncertainty are often required to satisfy a constraint θ∈Θδ\theta \in \Theta_\deltaθ∈Θδ​ that depends on a random parameter δ\deltaδ, and requiring it for every possible δ\deltaδ is usually too conservative or infeasible. The scenario approach replaces the unknown distribution of δ\deltaδ by NNN independent samples (scenarios) and enforces only the sampled constraints; its generalization theorem (Campi and Garatti, 2008) bounds the probability that the resulting decision violates a fresh constraint.

Enforcing all NNN sampled constraints can still be costly: a few unusual scenarios may dominate the solution. A practitioner therefore often discards kkk of the sampled constraints, optimally, greedily or at random, and re-solves. The question is what guarantee survives: the removed constraints were chosen by looking at the data, so the solution is biased towards points of higher risk. Campi and Garatti (2011) answered it with a bound that holds for every removal procedure. This mission formalizes that answer as it is presented in Chapter 3, Section 3.3 and Chapter 5, Section 5.3 of the textbook Introduction to the Scenario Approach (Campi and Garatti, SIAM/MOS 2018), together with its explicit corollary, Theorem 1.2. Applications include chance-constrained control, portfolio selection and prediction, where discarding scenarios trades a controlled amount of risk for a better cost.

Setting

A decision θ\thetaθ ranges over Rd\mathbb R^dRd (in Lean, EuclideanSpace ℝ (Fin d)), with a closed convex domain Θ\ThetaΘ and a linear cost cTθc^{\mathsf T}\thetacTθ. An uncertain parameter δ\deltaδ takes values in a measurable space Δ\DeltaΔ with probability P\mathbb PP, and each δ\deltaδ determines a closed convex constraint set Θδ\Theta_\deltaΘδ​. The violation probability of a decision is

V(θ)=P{δ∈Δ:θ∉Θδ}.V(\theta) = \mathbb P\{\delta \in \Delta : \theta \notin \Theta_\delta\}.V(θ)=P{δ∈Δ:θ∈/Θδ​}.

Given independent samples δ1,…,δN\delta_1,\dots,\delta_Nδ1​,…,δN​ with joint law PN\mathbb P^NPN, the scenario program minimizes cTθc^{\mathsf T}\thetacTθ over θ∈Θ∩⋂i=1NΘδi\theta \in \Theta \cap \bigcap_{i=1}^N \Theta_{\delta_i}θ∈Θ∩⋂i=1N​Θδi​​. For a set III of indexes, the program without the constraints in III minimizes the same cost over Θ∩⋂i∉IΘδi\Theta \cap \bigcap_{i \notin I} \Theta_{\delta_i}Θ∩⋂i∈/I​Θδi​​; its solution is written θI∗\theta^*_IθI∗​. A removal procedure selects, as a function of the whole sample, a set of kkk indexes, and θk∗\theta^*_kθk∗​ denotes the solution of the program without them. The procedure is required to output a solution that violates exactly the kkk removed constraints (with probability one): a removed constraint that turns out to be satisfied is reinstated and another is removed. Two standing assumptions are used throughout: Assumption 3.4, that Θ\ThetaΘ and every Θδ\Theta_\deltaΘδ​ are convex and closed, and Assumption 3.6, that for every sample size mmm and every sample the scenario program has exactly one solution.

Formalization targets

Goal: Theorem 3.9

For N≥dN \ge dN≥d, under Assumptions 3.4 and 3.6, for every removal procedure and every ε∈[0,1]\varepsilon \in [0,1]ε∈[0,1],

PN{V(θk∗)>ε}≤(k+d−1k)∑i=0k+d−1(Ni)εi(1−ε)N−i.\mathbb P^N\{V(\theta^*_k) > \varepsilon\} \le \binom{k+d-1}{k} \sum_{i=0}^{k+d-1} \binom Ni \varepsilon^i (1-\varepsilon)^{N-i}.PN{V(θk∗​)>ε}≤(kk+d−1​)i=0∑k+d−1​(iN​)εi(1−ε)N−i.

The bound depends on the problem only through ddd, and on the removal procedure not at all. For k=0k = 0k=0 it is Theorem 3.7.

Milestones

  1. Theorem 3.7 (no removal): PN{V(θ∗)>ε}≤∑i=0d−1(Ni)εi(1−ε)N−i\mathbb P^N\{V(\theta^*) > \varepsilon\} \le \sum_{i=0}^{d-1}\binom Ni\varepsilon^i(1-\varepsilon)^{N-i}PN{V(θ∗)>ε}≤∑i=0d−1​(iN​)εi(1−ε)N−i, used for the program with the N−kN-kN−k kept constraints.
  2. Eq. (5.11): up to a zero probability set, the event {V(θk∗)>ε}\{V(\theta^*_k) > \varepsilon\}{V(θk∗​)>ε} is contained in the union over all kkk-element index sets III of the events "θI∗\theta^*_IθI∗​ violates all constraints in III and V(θI∗)>εV(\theta^*_I) > \varepsilonV(θI∗​)>ε".
  3. Eq. (5.13): for a fixed III, the probability of that event equals ∫(ε,1]αkFV(dα)\int_{(\varepsilon,1]} \alpha^k F_V(d\alpha)∫(ε,1]​αkFV​(dα), where FVF_VFV​ is the law of V(θI∗)V(\theta^*_I)V(θI∗​).
  4. Eq. (5.14) and Theorem 3.9 for d=2d = 2d=2: the book's complete proof in the plane.
  5. Eqs. (3.15)–(3.17) and the conclusion of Section 3.3.1: with the explicit level εk\varepsilon_kεk​ of (1.9), the right-hand side of (3.13) is at most β\betaβ.
  6. Theorem 1.2: with probability at least 1−β1-\beta1−β, V(θk∗)≤εkV(\theta^*_k) \le \varepsilon_kV(θk∗​)≤εk​, where
εk=kN+[kN+k+1N((d−1)ln⁡(k+d−1)+d−1k+ln⁡1β)].\varepsilon_k = \frac{k}{N} + \left[\frac{\sqrt k}{N} + \frac{\sqrt k+1}{N}\left((d-1)\ln(k+d-1) + \frac{d-1}{\sqrt k} + \ln\frac1\beta\right)\right].εk​=Nk​+[Nk​​+Nk​+1​((d−1)ln(k+d−1)+k​d−1​+lnβ1​)].

Significance

Theorem 3.9 certifies every constraint-removal heuristic at once. Since the guarantee is the same for optimal, greedy and random removal, a user may pick the removal strategy purely for cost, and may inspect several values of kkk before choosing, paying only a union bound over the values tried (Section 3.3). Theorem 1.2 turns the bound into an explicit rate: when k/Nk/Nk/N is held fixed, the violation exceeds the empirical risk k/Nk/Nk/N by a margin of order ln⁡N/N\ln N/\sqrt NlnN/N​, only slightly worse than the 1/N1/\sqrt N1/N​ rate for estimating the probability of a fixed event. The result also shows that the violation after removal concentrates around the target level, which is the basis of the book's comparison between sampling-and-discarding and simply using fewer scenarios (Example 3.10).

Theorem 3.9 is proved in the literature for general ddd (Campi and Garatti, 2011); the textbook proves it for d=2d = 2d=2. To our knowledge no part of the scenario approach has a machine-checked proof. A formal development would supply the first verified version of the removal bound, a Lean treatment of solution maps of random convex programs, and reusable combinatorial and binomial-tail estimates.

Difficulty

The removed set is chosen after seeing the data, so the kept constraints are not an independent sample and Theorem 3.7 cannot be applied to θk∗\theta^*_kθk∗​ directly. The argument must pass through all (Nk)\binom Nk(kN​) fixed index sets and account for the event that the removed constraints are violated; a plain union bound that ignores this event loses a factor (Nk)\binom Nk(kN​) and does not give (3.13). For a fixed index set, the probability that the kkk removed scenarios are all violated involves the distribution of V(θI∗)V(\theta^*_I)V(θI∗​), which is only known to be dominated by a Beta law, so a stochastic-domination argument for the increasing function α↦αk\alpha \mapsto \alpha^kα↦αk is needed. In general dimension the combinatorial constant (k+d−1k)\binom{k+d-1}{k}(kk+d−1​) comes from a sharper counting than the two-dimensional computation of Section 5.3, and that argument is in the cited paper rather than in the book.

Formalization scope

Decisions live in EuclideanSpace ℝ (Fin d), samples of size mmm are maps Fin m → Δ with law Measure.pi (fun _ => P) for a probability measure P, and the violation is the real number (P {δ | θ ∉ Θδ δ}).toReal. Events over samples are compared in ℝ≥0∞ with ENNReal.ofReal of the book's right-hand side. The removal procedure is an arbitrary map I : (Fin N → Δ) → Finset (Fin N) with (I ω).card = k, and θk is a map that, for every sample, solves the program without the constraints in I ω, and violates each of them with probability one. The following implicit hypotheses of the book are written as binders:

  • d≥1d \ge 1d≥1, d≤Nd \le Nd≤N, k≤Nk \le Nk≤N and ε∈[0,1]\varepsilon \in [0,1]ε∈[0,1];
  • Assumption 3.6 for every mmm, including m=0m = 0m=0, and for every sample (not almost every);
  • the removed constraints are violated with probability one (∀ᵐ ω ∂ℙ^N), the book's own hypothesis on p. 65, so (5.11) is an inclusion up to a null set as on the page; requiring the violation for every sample would be unsatisfiable for 1≤k<N1 \le k < N1≤k<N (on a sample with all δi\delta_iδi​ equal a kept constraint coincides with a removed one) and would make the results vacuous;
  • measurability, which the book glosses over (p. 33): the constraint relation {(θ,δ):θ∈Θδ}\{(\theta,\delta) : \theta \in \Theta_\delta\}{(θ,δ):θ∈Θδ​} is jointly measurable, the solution map of the scenario program with mmm constraints is measurable for every mmm, and θk∗\theta^*_kθk∗​ is measurable;
  • for Theorem 1.2 and Section 3.3.1: k≥1k \ge 1k≥1 (formula (1.9) divides by k\sqrt kk​), N≥1N \ge 1N≥1, β∈(0,1)\beta \in (0,1)β∈(0,1); Section 3.3.1 additionally assumes εk≤1\varepsilon_k \le 1εk​≤1, the range in which its chain of inequalities holds.

Theorem 1.2 is stated in the constraint formulation of Chapter 3, to which the book says it "straightforwardly generalizes" (p. 20), with the hypotheses of Theorem 3.9 from which Section 3.3.1 derives it. Eq. (5.14) and the closing display of Section 5.3 are stated for d=2d = 2d=2 only, as in the book.

A trivializing formalization is excluded: the removal procedure is universally quantified, the solutions are exact minimizers rather than arbitrary feasible points, and the event is the strict V(θk∗)>εV(\theta^*_k) > \varepsilonV(θk∗​)>ε; a statement for one fixed rule, or with θk∗\theta^*_kθk∗​ unconstrained, would be a different theorem.

A complete development needs: product measures and Fubini over Fin N → Δ, reindexing of the kept constraints as a sample of size N−kN-kN−k, the Beta form of the binomial tail (the platform's binomial_upper_tail_eq_incomplete_beta is available), and stochastic domination for monotone integrands. Solution-map and violation infrastructure is shared with the sibling missions of this series. Contributions on any milestone, including the general-ddd counting argument of the cited paper, are welcome.

Selected references

  • M. C. Campi and S. Garatti, Introduction to the Scenario Approach, MOS-SIAM Series on Optimization 26, SIAM/MOS, 2018. https://doi.org/10.1137/1.9781611975444
  • M. C. Campi and S. Garatti, A sampling-and-discarding approach to chance-constrained optimization: feasibility and optimality, Journal of Optimization Theory and Applications 148(2), 257–280, 2011. https://doi.org/10.1007/s10957-010-9754-6
  • M. C. Campi and S. Garatti, The exact feasibility of randomized solutions of uncertain convex programs, SIAM Journal on Optimization 19(3), 1211–1230, 2008. https://doi.org/10.1137/07069821X
  • G. C. Calafiore and M. C. Campi, The scenario approach to robust control design, IEEE Transactions on Automatic Control 51(5), 742–753, 2006. https://doi.org/10.1109/TAC.2006.875041
13 thms1 active userReviewed
Operations ResearchOptimizationProbability·Captain: mikedeng1

Introduction to the Scenario Approach IV: The FAST Algorithm Keeps the Beta Bound and Adds a Factor (1−ε)^{N₂}Textbook

Motivation

The scenario approach turns an optimization problem under uncertainty into a finite, data-driven program: sample NNN instances of the uncertain parameter, optimize against all of them, and certify how often the resulting design fails on a new instance. Its main guarantee (Theorem 3.7 of Campi and Garatti's Introduction to the Scenario Approach) bounds the probability of failure by a binomial tail in NNN and in the number ddd of optimization variables. To reach a failure level ε\varepsilonε with confidence 1−β1-\beta1−β, the number of scenarios grows roughly like 2ε(ln⁡1β+d−1)\frac{2}{\varepsilon}\big(\ln\frac1\beta+d-1\big)ε2​(lnβ1​+d−1) (Theorem 1.1 of the book). The product of ddd and 1/ε1/\varepsilon1/ε is what makes medium- and large-scale designs expensive: each scenario is one more constraint in the program that has to be solved.

FAST (Fast Algorithm for the Scenario Technique), introduced by Carè, Garatti and Campi in Operations Research 62 (2014), removes that product. It solves the program with a moderate number N1N_1N1​ of scenarios and then, instead of re-optimizing, raises the returned cost level until it covers N2N_2N2​ further scenarios. The book presents the algorithm and its guarantee, Theorem 8.5, in §8.3, and refers to the paper for the proof. This mission formalizes that guarantee.

Timeline:

  • 2006, Calafiore and Campi: violation bounds for the solution of convex scenario programs.
  • 2008, Campi and Garatti: the exact binomial bound, tight for fully supported problems (Theorem 3.7 of the book).
  • 2014, Carè, Garatti and Campi: FAST and its two-stage bound, Eq. (8.5).
  • 2018, Campi and Garatti's textbook, §8.3, the source of this mission.

Setting

Let Δ\DeltaΔ be a measurable space carrying a probability measure P\mathbb PP, and let ℓ(ν,δ)\ell(\nu,\delta)ℓ(ν,δ) be a real loss of a decision ν∈Rd−1\nu\in\mathbb R^{d-1}ν∈Rd−1 under the uncertain parameter δ∈Δ\delta\in\Deltaδ∈Δ. As a standing assumption of the book, ℓ(⋅,δ)\ell(\cdot,\delta)ℓ(⋅,δ) is convex for every δ\deltaδ.

Given scenarios δ1,…,δm\delta_1,\dots,\delta_mδ1​,…,δm​ drawn independently from P\mathbb PP, the scenario program (1.4) is

min⁡ν∈Rd−1 [max⁡i=1,…,m ℓ(ν,δi)].\min_{\nu\in\mathbb R^{d-1}}\ \Big[\max_{i=1,\dots,m}\ \ell(\nu,\delta_i)\Big].ν∈Rd−1min​ [i=1,…,mmax​ ℓ(ν,δi​)].

Its solution is ν∗\nu^*ν∗ and its optimal value ℓ∗\ell^*ℓ∗. Assumption 3.6 requires that for every mmm and every sample the solution exist and be unique. The pair (ν,ℓ)(\nu,\ell)(ν,ℓ) has ddd components, and ddd is the number that enters every bound.

The risk (Definition 8.2) of a decision ν\nuν with cost level ℓ\ellℓ is

R(ν,ℓ)=P{δ∈Δ: ℓ(ν,δ)>ℓ},R(\nu,\ell)=\mathbb P\{\delta\in\Delta:\ \ell(\nu,\delta)>\ell\},R(ν,ℓ)=P{δ∈Δ: ℓ(ν,δ)>ℓ},

the probability that a new instance costs more than promised. It is the violation V(ν,ℓ)V(\nu,\ell)V(ν,ℓ) of the epigraphic constraint ℓ≥ℓ(ν,δ)\ell\ge\ell(\nu,\delta)ℓ≥ℓ(ν,δ).

FAST takes N1+N2N_1+N_2N1​+N2​ independent scenarios. It solves (1.4) with the first N1N_1N1​ of them, obtaining νN1∗\nu^*_{N_1}νN1​∗​. In the detuning step it then sets

ℓF∗=max⁡i=1,…,N1+N2 ℓ(νN1∗,δi),\ell^*_F=\max_{i=1,\dots,N_1+N_2}\ \ell(\nu^*_{N_1},\delta_i),ℓF∗​=i=1,…,N1​+N2​max​ ℓ(νN1​∗​,δi​),

the smallest level that covers every scenario seen. The output is (νF∗,ℓF∗)(\nu^*_F,\ell^*_F)(νF∗​,ℓF∗​) with νF∗=νN1∗\nu^*_F=\nu^*_{N_1}νF∗​=νN1​∗​.

Formalization targets

Goal: Theorem 8.5, Eq. (8.5)

For every ε∈[0,1]\varepsilon\in[0,1]ε∈[0,1],

PN1+N2{V(νF∗,ℓF∗)>ε} ≤ (1−ε)N2∑i=0d−1(N1i)εi(1−ε)N1−i.\mathbb P^{N_1+N_2}\{V(\nu^*_F,\ell^*_F)>\varepsilon\}\ \le\ (1-\varepsilon)^{N_2}\sum_{i=0}^{d-1}\binom{N_1}{i}\varepsilon^i(1-\varepsilon)^{N_1-i}.PN1​+N2​{V(νF∗​,ℓF∗​)>ε} ≤ (1−ε)N2​i=0∑d−1​(iN1​​)εi(1−ε)N1​−i.

No relation between N1N_1N1​ and ddd is required. When N1<dN_1<dN1​<d the sum equals 111 and the bound reads (1−ε)N2(1-\varepsilon)^{N_2}(1−ε)N2​.

Milestone: Theorem 3.7 for program (1.4)

The first stage is an ordinary scenario program with N1N_1N1​ scenarios. For N≥dN\ge dN≥d,

PN{R(ν∗,ℓ∗)>ε} ≤ ∑i=0d−1(Ni)εi(1−ε)N−i,\mathbb P^N\{R(\nu^*,\ell^*)>\varepsilon\}\ \le\ \sum_{i=0}^{d-1}\binom{N}{i}\varepsilon^i(1-\varepsilon)^{N-i},PN{R(ν∗,ℓ∗)>ε} ≤ i=0∑d−1​(iN​)εi(1−ε)N−i,

that is, R(ν∗,ℓ∗)R(\nu^*,\ell^*)R(ν∗,ℓ∗) is dominated by a B(d,N−d+1)B(d,N-d+1)B(d,N−d+1) distribution (recalled on p. 90).

Milestone: the N2N_2N2​ rule

For ε,β∈(0,1)\varepsilon,\beta\in(0,1)ε,β∈(0,1), N2≥1εln⁡1βN_2\ge\frac1\varepsilon\ln\frac1\betaN2​≥ε1​lnβ1​ makes the right-hand side of (8.5) at most β\betaβ (p. 95).

Significance

The result. Theorem 8.5 makes the guarantee of the scenario approach cheap to obtain. With N1=KdN_1=KdN1​=Kd (the book suggests K≈20K\approx20K≈20) and N2≥1εln⁡1βN_2\ge\frac1\varepsilon\ln\frac1\betaN2​≥ε1​lnβ1​, the total number of scenarios is Kd+1εln⁡1βKd+\frac1\varepsilon\ln\frac1\betaKd+ε1​lnβ1​. This is additive in ddd and 1/ε1/\varepsilon1/ε rather than multiplicative, and the added N2N_2N2​ scenarios cost only function evaluations, not a larger optimization. The price is suboptimality: ℓF∗\ell^*_FℓF∗​ is in general higher than the value a classical scenario program with the same confidence would return.

Formalizing it. The result is proved on paper, in the cited 2014 article; the book states it without proof. No part of the scenario theory has been machine-checked on this platform, as far as a search of the catalog shows. The mission produces a checked two-stage bound whose first stage is the loss-function form of Theorem 3.7, which is reusable by every mission of the series that works with program (1.4). The N2N_2N2​ rule is an elementary but explicit sample-size certificate.

Difficulty

The obvious route treats the detuning step as a fresh scenario program with N1+N2N_1+N_2N1​+N2​ scenarios and applies Theorem 3.7 to it. That fails: νF∗\nu^*_FνF∗​ is not the solution of that program, and Theorem 3.7 with N1+N2N_1+N_2N1​+N2​ scenarios gives a bound that is not of the product form (8.5). The level ℓF∗\ell^*_FℓF∗​ depends on all N1+N2N_1+N_2N1​+N2​ scenarios at once, including those that determined νN1∗\nu^*_{N_1}νN1​∗​, and the map c↦R(ν,c)c\mapsto R(\nu,c)c↦R(ν,c) is monotone but need not be continuous, so the event V(νF∗,ℓF∗)>εV(\nu^*_F,\ell^*_F)>\varepsilonV(νF∗​,ℓF∗​)>ε is not a simple event about the new scenarios. Underneath the goal sits Theorem 3.7 itself, which is the main theorem of the book and whose proof occupies Chapter 5.

Formalization scope

Lean representation:

  • The decision space Rd−1\mathbb R^{d-1}Rd−1 is EuclideanSpace ℝ (Fin n); the book's ddd is written n+1n+1n+1, never with natural-number subtraction.
  • A sample of size mmm is ω : Fin m → Δ with law Measure.pi (fun _ => P), and the same P\mathbb PP defines the risk. FAST draws one sample ω : Fin (N₁ + N₂) → Δ; its first stage is ω ∘ Fin.castAdd N₂.
  • The maximum in (1.4) and in ℓF∗\ell^*_FℓF∗​ is Finset.sup' over a nonempty index set. ℓF∗\ell^*_FℓF∗​ runs over all N1+N2N_1+N_2N1​+N2​ scenarios, not over the N2N_2N2​ new ones only.
  • The risk is (P {δ | c < ℓ ν δ}).toReal, with the strict inequality of Definition 8.2 and the strict event V>εV>\varepsilonV>ε of (8.5). Probabilities of sample events are compared in ℝ≥0∞ through ENNReal.ofReal.
  • The first-stage solution is a map νstar from samples to decisions, with the hypothesis that νstar ω₁ solves the program for every sample ω₁.

Hypotheses the book leaves implicit, stated explicitly:

  1. ℓ(⋅,δ)\ell(\cdot,\delta)ℓ(⋅,δ) is convex for every δ\deltaδ (standing assumption, p. 6).
  2. Existence and uniqueness of the solution (Assumption 3.6) for every m≥1m\ge1m≥1 and every sample. The program with no scenario has no minimum, so m=0m=0m=0 is excluded.
  3. N1≥1N_1\ge1N1​≥1, since the first stage needs a scenario.
  4. ε∈[0,1]\varepsilon\in[0,1]ε∈[0,1]; for ε>1\varepsilon>1ε>1 the factor (1−ε)N2(1-\varepsilon)^{N_2}(1−ε)N2​ can be negative.
  5. The loss is jointly measurable in (ν,δ)(\nu,\delta)(ν,δ) and the first-stage solution map is measurable. The book glosses over measurability (p. 6, footnote 1; p. 33).

A formalization that bounds only the N2N_2N2​ new scenarios is ruled out, because ℓF∗\ell^*_FℓF∗​ is defined as a maximum over all N1+N2N_1+N_2N1​+N2​ scenarios. So is one that takes ℓF∗\ell^*_FℓF∗​ as a free variable or drops Assumption 3.6: the goal is stated for the output of FAST as the book defines it.

A complete development needs the scenario program in loss form, product-measure conditioning on ΔN1×ΔN2\Delta^{N_1}\times\Delta^{N_2}ΔN1​×ΔN2​, and Theorem 3.7. The loss-form Theorem 3.7 is the reusable piece. Proofs of the milestones and of intermediate conditioning lemmas are welcome.

Selected references

  • M. C. Campi, S. Garatti, Introduction to the Scenario Approach, MOS-SIAM Series on Optimization 26, SIAM, 2018, §8.3 and Theorem 3.7. https://doi.org/10.1137/1.9781611975444
  • A. Carè, S. Garatti, M. C. Campi, FAST—Fast Algorithm for the Scenario Technique, Operations Research 62(3):662–671, 2014. https://doi.org/10.1287/opre.2014.1257
  • M. C. Campi, S. Garatti, The exact feasibility of randomized solutions of uncertain convex programs, SIAM Journal on Optimization 19(3):1211–1230, 2008. https://doi.org/10.1137/07069821X
  • G. C. Calafiore, M. C. Campi, The scenario approach to robust control design, IEEE Transactions on Automatic Control 51(5):742–753, 2006. https://doi.org/10.1109/TAC.2006.875041
6 thms1 active userReviewed
Linear OptimizationOperations ResearchOptimization·Captain: mikedeng1

Linear Programming: Foundations and Extensions V: Convergence Rates of the Path-Following MethodTextbook

Motivation

Interior-point methods are, together with the simplex method, the standard algorithms for linear programming, and the primal–dual path-following method is the form in which they are implemented in most solvers. Unlike the simplex method, it is a one-phase method: it can start from any point whose primal and dual variables are strictly positive, feasible or not, and drives infeasibility and complementarity to zero simultaneously. The question every user of such a method eventually asks is how fast these three measures of non-optimality decrease.

Chapter 18 of R. J. Vanderbei, Linear Programming: Foundations and Extensions (4th ed., Springer 2014, DOI 10.1007/978-1-4614-7630-6) defines the method from scratch (Fig. 18.1, p. 273) and proves a rate statement, Theorem 18.1 (pp. 277–279): as long as the step lengths stay bounded below and the iterates stay bounded, the primal and dual infeasibilities decay geometrically, and so does the complementarity, at a slower rate. This mission formalizes that theorem and the one-step identities and estimates it is built from. It is the fifth mission of a series on the book; the missions are independent of each other.

Setting

Let AAA be a real m×nm \times nm×n matrix, b∈Rmb \in \mathbb{R}^mb∈Rm, c∈Rnc \in \mathbb{R}^nc∈Rn. The primal problem is to maximize cTxc^TxcTx subject to Ax+w=bAx + w = bAx+w=b, x,w≥0x, w \ge 0x,w≥0; the dual is to minimize bTyb^TybTy subject to ATy−z=cA^Ty - z = cATy−z=c, y,z≥0y, z \ge 0y,z≥0. A primal–dual point is a quadruple (x,w,y,z)(x, w, y, z)(x,w,y,z) with x,z∈Rnx, z \in \mathbb{R}^nx,z∈Rn, w,y∈Rmw, y \in \mathbb{R}^mw,y∈Rm; it is strictly positive, (x,w,y,z)>0(x, w, y, z) > 0(x,w,y,z)>0, if every component is. Write X,W,Y,ZX, W, Y, ZX,W,Y,Z for the diagonal matrices of x,w,y,zx, w, y, zx,w,y,z and eee for the all-ones vector. The norms are ∥v∥1=∑j∣vj∣\|v\|_1 = \sum_j |v_j|∥v∥1​=∑j​∣vj​∣ and ∥v∥∞=max⁡j∣vj∣\|v\|_\infty = \max_j |v_j|∥v∥∞​=maxj​∣vj​∣.

At a point (x,w,y,z)(x, w, y, z)(x,w,y,z) the three measures of progress are the primal infeasibility ρ=b−Ax−w\rho = b - Ax - wρ=b−Ax−w, the dual infeasibility σ=c−ATy+z\sigma = c - A^Ty + zσ=c−ATy+z, and the complementarity γ=zTx+yTw\gamma = z^Tx + y^Twγ=zTx+yTw. Fix parameters 0<δ<10 < \delta < 10<δ<1 and 0<r<10 < r < 10<r<1. One iteration of the method, from a strictly positive point, sets μ=δγ/(n+m)\mu = \delta\gamma/(n+m)μ=δγ/(n+m), takes any solution (Δx,Δw,Δy,Δz)(\Delta x, \Delta w, \Delta y, \Delta z)(Δx,Δw,Δy,Δz) of the Newton system

AΔx+Δw=ρ,ATΔy−Δz=σ,ZΔx+XΔz=μe−XZe,WΔy+YΔw=μe−YWe,A\Delta x + \Delta w = \rho, \quad A^T\Delta y - \Delta z = \sigma, \quad Z\Delta x + X\Delta z = \mu e - XZe, \quad W\Delta y + Y\Delta w = \mu e - YWe,AΔx+Δw=ρ,ATΔy−Δz=σ,ZΔx+XΔz=μe−XZe,WΔy+YΔw=μe−YWe,

computes the step length

θ=r(max⁡i,j{∣Δxjxj∣,∣Δwiwi∣,∣Δyiyi∣,∣Δzjzj∣})−1∧1(18.7)\theta = r\left(\max_{i,j}\left\{\left|\tfrac{\Delta x_j}{x_j}\right|, \left|\tfrac{\Delta w_i}{w_i}\right|, \left|\tfrac{\Delta y_i}{y_i}\right|, \left|\tfrac{\Delta z_j}{z_j}\right|\right\}\right)^{-1} \wedge 1 \qquad (18.7)θ=r(i,jmax​{​xj​Δxj​​​,​wi​Δwi​​​,​yi​Δyi​​​,​zj​Δzj​​​})−1∧1(18.7)

(with θ=1\theta = 1θ=1 when all ratios vanish), and moves to (x+θΔx,w+θΔw,y+θΔy,z+θΔz)(x + \theta\Delta x, w + \theta\Delta w, y + \theta\Delta y, z + \theta\Delta z)(x+θΔx,w+θΔw,y+θΔy,z+θΔz). This is Fig. 18.1 with the shorter step (18.7) the book adopts for its analysis. Along a sequence of iterates, superscripts (k)^{(k)}(k) denote the quantities at the kkk-th iterate, and θ(k)\theta^{(k)}θ(k) is the step length computed there.

Formalization targets

Goal: Theorem 18.1 with the explicit constant

If t>0t > 0t>0, MMM is real, and for all k≤Kk \le Kk≤K one has θ(k)≥t\theta^{(k)} \ge tθ(k)≥t, ∥x(k)∥∞≤M\|x^{(k)}\|_\infty \le M∥x(k)∥∞​≤M, ∥y(k)∥∞≤M\|y^{(k)}\|_\infty \le M∥y(k)∥∞​≤M, then for all k≤Kk \le Kk≤K, with t~=t(1−δ)\tilde t = t(1-\delta)t~=t(1−δ),

∥ρ(k)∥1≤(1−t)k∥ρ(0)∥1,∥σ(k)∥1≤(1−t)k∥σ(0)∥1,γ(k)≤(1−t~)k(γ(0)+M(∥ρ(0)∥1+∥σ(0)∥1)δt).\|\rho^{(k)}\|_1 \le (1-t)^k\|\rho^{(0)}\|_1, \qquad \|\sigma^{(k)}\|_1 \le (1-t)^k\|\sigma^{(0)}\|_1, \qquad \gamma^{(k)} \le (1-\tilde t)^k \left(\gamma^{(0)} + \frac{M(\|\rho^{(0)}\|_1 + \|\sigma^{(0)}\|_1)}{\delta t}\right).∥ρ(k)∥1​≤(1−t)k∥ρ(0)∥1​,∥σ(k)∥1​≤(1−t)k∥σ(0)∥1​,γ(k)≤(1−t~)k(γ(0)+δtM(∥ρ(0)∥1​+∥σ(0)∥1​)​).

Milestones

The one-step identities for the infeasibilities, ρ~=(1−θ)ρ\tilde\rho = (1-\theta)\rhoρ~​=(1−θ)ρ (18.8) and σ~=(1−θ)σ\tilde\sigma = (1-\theta)\sigmaσ~=(1−θ)σ (18.9); the one-step complementarity estimate

γ~≤(1−(1−δ)θ)γ+M∥ρ∥1+M∥σ∥1(18.10)\tilde\gamma \le (1 - (1-\delta)\theta)\gamma + M\|\rho\|_1 + M\|\sigma\|_1 \qquad (18.10)γ~​≤(1−(1−δ)θ)γ+M∥ρ∥1​+M∥σ∥1​(18.10)

under ∥x∥∞,∥y∥∞≤M\|x\|_\infty, \|y\|_\infty \le M∥x∥∞​,∥y∥∞​≤M; and the recursion γ(k)≤(1−t~)γ(k−1)+M(1−t)k−1(∥ρ(0)∥1+∥σ(0)∥1)\gamma^{(k)} \le (1-\tilde t)\gamma^{(k-1)} + M(1-t)^{k-1}(\|\rho^{(0)}\|_1 + \|\sigma^{(0)}\|_1)γ(k)≤(1−t~)γ(k−1)+M(1−t)k−1(∥ρ(0)∥1​+∥σ(0)∥1​) (18.11). Two unnumbered statements complete the picture: every iteration has 0<θ≤10 < \theta \le 10<θ≤1 and keeps the point strictly positive, and at any strictly positive point the duality gap satisfies ∣bTy−cTx∣≤γ+∥σ∥1∥x∥∞+∥ρ∥1∥y∥∞|b^Ty - c^Tx| \le \gamma + \|\sigma\|_1\|x\|_\infty + \|\rho\|_1\|y\|_\infty∣bTy−cTx∣≤γ+∥σ∥1​∥x∥∞​+∥ρ∥1​∥y∥∞​ (§18.5.3).

Significance

Theorem 18.1 separates the convergence question for the path-following method into two parts: a rate statement that holds whenever steps stay long and iterates stay bounded, and the remaining question of when those two conditions hold. It also explains an effect seen in practice: the infeasibilities fall by the factor 1−t1 - t1−t per iteration while the complementarity, and hence (by the duality-gap estimate) the gap bTy−cTxb^Ty - c^TxbTy−cTx, falls only by 1−t~1 - \tilde t1−t~. The book stresses that the result is partial, because it does not show that the step lengths remain bounded away from zero; that requires modifications of the method and of the starting point that the book does not carry out.

All statements here are proved in the book. The mission's contribution is a machine-checked version, with the constant of the complementarity bound made explicit. Neither Mathlib nor the platform contains a formal proof of this theorem or a formalization of the infeasible-start primal–dual iteration it concerns; the platform's existing path-following result concerns a different, feasible-start short-step method in equality form.

Difficulty

The infeasibility identities are linear and follow from the first two Newton equations. The complementarity is where the Newton system linearizes a bilinear equation, so the new complementarity contains a second-order term θ2(ΔyTρ−σTΔx)\theta^2(\Delta y^T\rho - \sigma^T\Delta x)θ2(ΔyTρ−σTΔx) that has no sign. Bounding it requires relating the size of the step θΔ\theta\DeltaθΔ to the size of the current iterate through the specific form of the step-length rule (18.7); the rule (18.6) of Fig. 18.1, with signed ratios, does not give such a bound. The multi-step estimate then couples two geometric sequences with different rates, and keeping the constant independent of the horizon KKK is what makes the statement non-trivial.

Formalization scope

Vectors are Fin n → ℝ and Fin m → ℝ, AAA is a Matrix (Fin m) (Fin n) ℝ, and points and step directions are a structure PDPoint m n with fields x w y z. The sup-norm is ⨆ j, |v j| (the maximum; 0 for an empty vector). The step length is written with the explicit case θ=1\theta = 1θ=1 when all ratios vanish, since Lean's r / 0 = 0 would otherwise give θ=0\theta = 0θ=0. An iteration is a relation between the current point, a step direction and the next point: the current point is strictly positive, the direction is some solution of the Newton system (uniqueness, which the book asserts under a full-rank assumption, is not assumed), and the next point is current + θ⋅+\ \theta \cdot+ θ⋅ direction. The hypotheses 0<δ<10 < \delta < 10<δ<1, 0<r<10 < r < 10<r<1 (pp. 272–273) are stated in every theorem; MMM is an arbitrary real number and KKK a natural number. As in the book, the hypotheses of Theorem 18.1 range over k≤Kk \le Kk≤K, so the iteration from index KKK is part of the data.

Explicit constants. The book's Theorem 18.1 asserts only "there exists a constant Mˉ<∞\bar M < \inftyMˉ<∞". Because KKK is fixed, that existential is satisfied trivially by max⁡k≤Kγ(k)/(1−t~)k\max_{k \le K}\gamma^{(k)}/(1-\tilde t)^kmaxk≤K​γ(k)/(1−t~)k, and a statement with ∃Mˉ\exists \bar M∃Mˉ would be empty. The goal therefore uses the constant the book's proof establishes (p. 279, last display): Mˉ=γ(0)+M(∥ρ(0)∥1+∥σ(0)∥1)/(δt)\bar M = \gamma^{(0)} + M(\|\rho^{(0)}\|_1 + \|\sigma^{(0)}\|_1)/(\delta t)Mˉ=γ(0)+M(∥ρ(0)∥1​+∥σ(0)∥1​)/(δt). Eq. (18.11) is stated with the book's M~=M(∥ρ(0)∥1+∥σ(0)∥1)\tilde M = M(\|\rho^{(0)}\|_1 + \|\sigma^{(0)}\|_1)M~=M(∥ρ(0)∥1​+∥σ(0)∥1​) written out.

The formalization needs only finite sums, dot products and matrix–vector products from Mathlib; the definitions of the iteration are reusable for other analyses of the same method (Chapters 19–22 of the book). Contributions are welcome for each milestone separately.

Selected references

  • R. J. Vanderbei, Linear Programming: Foundations and Extensions, 4th ed., International Series in Operations Research & Management Science 196, Springer, 2014, Chapter 18, pp. 269–283. https://doi.org/10.1007/978-1-4614-7630-6
  • S. J. Wright, Primal-Dual Interior-Point Methods, SIAM, 1997. https://doi.org/10.1137/1.9781611971453
6 thms1 active userReviewed
Linear OptimizationOperations ResearchOptimization·Captain: mikedeng1

Linear Programming: Foundations and Extensions IV: Existence of the Central PathTextbook

Motivation

Interior-point methods solve linear programs by moving through the interior of the feasible region instead of along its edges, as the simplex method does. The methods used in practice are path-following methods: they track a curve, the central path, that runs through the interior of the feasible region and ends at an optimal solution. Before any such method can be analysed, the curve has to exist. This mission formalizes Chapter 17 of R. J. Vanderbei, Linear Programming: Foundations and Extensions (4th ed., Springer 2014), which defines the central path through the logarithmic barrier problem and proves that it exists exactly when the primal and the dual problem both have strictly positive feasible points.

The chapter's results have a short history. Barrier methods for nonlinear programming go back to Fiacco and McCormick (1968). Interest in interior-point methods for linear programming began with Karmarkar (1984), whose projective algorithm does not mention a central path; the connection between Karmarkar's method and the primal–dual central path was found by Megiddo (1989), with central points traced back to Huard (1967) and an extended study of the path by Bayer and Lagarias (1989). The chapter is the textbook entry point to this line of work and the foundation for the path-following algorithm of Chapter 18.

Setting

Let AAA be a real m×nm \times nm×n matrix, b∈Rmb \in \mathbb{R}^mb∈Rm and c∈Rnc \in \mathbb{R}^nc∈Rn. The primal linear program is to maximize cTxc^T xcTx subject to Ax≤bAx \le bAx≤b, x≥0x \ge 0x≥0; its dual is to minimize bTyb^T ybTy subject to ATy≥cA^T y \ge cATy≥c, y≥0y \ge 0y≥0. With slack variables w∈Rmw \in \mathbb{R}^mw∈Rm and z∈Rnz \in \mathbb{R}^nz∈Rn they read (17.1)

Ax+w=b, x,w≥0andATy−z=c, y,z≥0.Ax + w = b,\ x, w \ge 0 \qquad\text{and}\qquad A^T y - z = c,\ y, z \ge 0.Ax+w=b, x,w≥0andATy−z=c, y,z≥0.

For a vector ξ\xiξ, ξ>0\xi > 0ξ>0 means that every component is strictly positive. The primal feasible region has nonempty interior when some (xˉ,wˉ)(\bar x, \bar w)(xˉ,wˉ) satisfies Axˉ+wˉ=bA\bar x + \bar w = bAxˉ+wˉ=b with xˉ>0\bar x > 0xˉ>0, wˉ>0\bar w > 0wˉ>0; the dual feasible region has nonempty interior when some (yˉ,zˉ)(\bar y, \bar z)(yˉ​,zˉ) satisfies ATyˉ−zˉ=cA^T \bar y - \bar z = cATyˉ​−zˉ=c with yˉ>0\bar y > 0yˉ​>0, zˉ>0\bar z > 0zˉ>0.

For a parameter μ>0\mu > 0μ>0, the barrier function (17.7) is

f(x,w)=cTx+μ∑j=1nlog⁡xj+μ∑i=1mlog⁡wi,f(x, w) = c^T x + \mu \sum_{j=1}^n \log x_j + \mu \sum_{i=1}^m \log w_i ,f(x,w)=cTx+μj=1∑n​logxj​+μi=1∑m​logwi​,

and the barrier problem (17.2) is to maximize f(x,w)f(x, w)f(x,w) subject to Ax+w=bAx + w = bAx+w=b, over the domain x>0x > 0x>0, w>0w > 0w>0 where the logarithms are finite. A solution of the barrier problem is a point of that domain at which fff attains its maximum over the domain.

Writing X,Z,Y,WX, Z, Y, WX,Z,Y,W for the diagonal matrices carrying x,z,y,wx, z, y, wx,z,y,w and eee for the all-ones vector, the primal–dual central-path system (17.6) is

Ax+w=b,ATy−z=c,XZe=μe,YWe=μe,Ax + w = b, \qquad A^T y - z = c, \qquad XZe = \mu e, \qquad YWe = \mu e,Ax+w=b,ATy−z=c,XZe=μe,YWe=μe,

with x,w,y,z>0x, w, y, z > 0x,w,y,z>0. The last two equations say xjzj=μx_j z_j = \muxj​zj​=μ and yiwi=μy_i w_i = \muyi​wi​=μ for all jjj and iii. The set of its solutions (xμ,wμ,yμ,zμ)(x_\mu, w_\mu, y_\mu, z_\mu)(xμ​,wμ​,yμ​,zμ​), μ>0\mu > 0μ>0, is the primal–dual central path.

The chapter also uses one fact from nonlinear programming: for the problem "maximize f(x)f(x)f(x) subject to gi(x)=0g_i(x) = 0gi​(x)=0, i=1,…,mi = 1, \dots, mi=1,…,m", a critical point is a feasible x∗x^*x∗ with ∇f(x∗)=∑iyi∇gi(x∗)\nabla f(x^*) = \sum_i y_i \nabla g_i(x^*)∇f(x∗)=∑i​yi​∇gi​(x∗) for some Lagrange multipliers yiy_iyi​ (17.3), and Hf(x∗)H_f(x^*)Hf​(x∗) is the Hessian of fff at x∗x^*x∗.

Formalization targets

Goal: Theorem 17.2 (p. 265)

For each fixed μ>0\mu > 0μ>0,

∃ (x,w) solving the barrier problem  ⟺  (∃ xˉ,wˉ>0:Axˉ+wˉ=b)∧(∃ yˉ,zˉ>0:ATyˉ−zˉ=c).\exists\, (x, w) \text{ solving the barrier problem} \iff \big(\exists\, \bar x, \bar w > 0 : A\bar x + \bar w = b\big) \wedge \big(\exists\, \bar y, \bar z > 0 : A^T\bar y - \bar z = c\big).∃(x,w) solving the barrier problem⟺(∃xˉ,wˉ>0:Axˉ+wˉ=b)∧(∃yˉ​,zˉ>0:ATyˉ​−zˉ=c).

Both directions are part of the goal. The statement fixes no constants and no rate; it asserts only when the barrier problem is solvable.

Milestones

  1. Theorem 17.1 (p. 261), second-order sufficiency under linear constraints: if the constraints are linear, a critical point x∗x^*x∗ with ξTHf(x∗)ξ<0\xi^T H_f(x^*)\xi < 0ξTHf​(x∗)ξ<0 for every ξ≠0\xi \ne 0ξ=0 satisfying ξT∇gi(x∗)=0\xi^T \nabla g_i(x^*) = 0ξT∇gi​(x∗)=0 for all iii is a local maximum on the feasible set.
  2. Exercise 10.7 (p. 150): if the primal is feasible and its feasible set {x:Ax≤b, x≥0}\{x : Ax \le b,\ x \ge 0\}{x:Ax≤b, x≥0} is bounded, then there are y>0y > 0y>0, z>0z > 0z>0 with ATy−z=cA^T y - z = cATy−z=c.
  3. Corollary 17.3 (p. 266): if the primal feasible set (or the dual feasible set) has nonempty interior and is bounded, then for each μ>0\mu > 0μ>0 the system (17.6) has exactly one solution with x,w,y,z>0x, w, y, z > 0x,w,y,z>0.

The corollary is stronger than the goal in one direction (it adds uniqueness and the dual variables) and weaker in another (it assumes boundedness).

Significance

The result itself. Theorem 17.2 gives an exact criterion for the barrier problem to be solvable for a fixed μ\muμ, and Corollary 17.3 turns it into the statement that the central path is a well-defined curve μ↦(xμ,wμ,yμ,zμ)\mu \mapsto (x_\mu, w_\mu, y_\mu, z_\mu)μ↦(xμ​,wμ​,yμ​,zμ​) for all μ>0\mu > 0μ>0. Every path-following method, including the one analysed in Chapter 18 of the same book, targets points on this curve; without existence and uniqueness the "target" of an iteration is undefined. The system (17.6) is also the starting point of the primal–dual Newton step.

Formalizing it. These results are classical and proved in the book; none is formalized in the Ax≤bAx \le bAx≤b, x≥0x \ge 0x≥0 form used here. The platform has the converse fact in the standard form Ax=bAx = bAx=b, x≥0x \ge 0x≥0 (a solution of the central-path conditions minimizes the barrier, Introduction to Linear Optimization), but not existence. A formal proof of Theorem 17.2 and Corollary 17.3 produces a reusable existence theorem for the central path that downstream missions on path-following and self-dual methods can import.

Difficulty

The "if" direction of Theorem 17.2 is an existence claim on a set that is neither closed nor bounded: the domain x>0x > 0x>0, w>0w > 0w>0 is open, and the feasible region itself may be unbounded, so the obvious appeal to "a continuous function on a compact set attains its maximum" does not apply directly. The example "maximize 000 subject to x≥0x \ge 0x≥0" (p. 264), whose barrier μlog⁡x\mu \log xμlogx has no maximum, shows that the dual hypothesis cannot be dropped. The "only if" direction, which the book calls trivial and does not prove, needs first-order conditions at a maximizer over a relatively open set.

Theorem 17.1 needs a second-order Taylor expansion with a remainder that is o(∥ξ∥2)o(\|\xi\|^2)o(∥ξ∥2) uniformly along the constraint subspace, not along individual lines. Exercise 10.7 is a theorem of the alternative and is not a consequence of weak duality alone. Uniqueness in Corollary 17.3 requires the positivity of the solution: the equations xjzj=μx_j z_j = \muxj​zj​=μ, yiwi=μy_i w_i = \muyi​wi​=μ admit sign-flipped solutions.

Formalization scope

Vectors are Fin n → ℝ and Fin m → ℝ; AAA is a Matrix (Fin m) (Fin n) ℝ. The book's primal–dual pair in Ax≤bAx \le bAx≤b, x≥0x \ge 0x≥0 form with slacks is used throughout; there are no explicit constants in this chapter.

Conventions committed to:

  • "Nonempty interior" means a feasible point with every component strictly positive, as the proof of Theorem 17.2 says. The topological interior of {(x,w):Ax+w=b, x,w≥0}\{(x, w) : Ax + w = b,\ x, w \ge 0\}{(x,w):Ax+w=b, x,w≥0} in Rn+m\mathbb{R}^{n+m}Rn+m is empty whenever m≥1m \ge 1m≥1; reading the theorem that way would make its right-hand side always false for m≥1m \ge 1m≥1, and that reading is ruled out.
  • The barrier problem is posed over x>0x > 0x>0, w>0w > 0w>0 explicitly; Real.log returns 000 at nonpositive arguments and is never evaluated there.
  • Solutions of (17.6) are required to be strictly positive, as in Exercise 17.3 (p. 267).
  • "Bounded" is Bornology.IsBounded of the feasible set in Rn\mathbb{R}^nRn (resp. Rm\mathbb{R}^mRm).
  • In Theorem 17.1 the constraints are Gx=βGx = \betaGx=β; fff is differentiable near x∗x^*x∗ with derivative differentiable at x∗x^*x∗, and ξTHf(x∗)ξ\xi^T H_f(x^*) \xiξTHf​(x∗)ξ is the second Fréchet derivative applied to (ξ,ξ)(\xi, \xi)(ξ,ξ). The local maximum is relative to the feasible set.

Infrastructure a complete development needs: attainment of maxima on compact sets (IsCompact.exists_isMaxOn in Mathlib), first-order conditions on relatively open sets, a theorem of the alternative for Exercise 10.7, and concavity facts about the logarithm. A second-order sufficient condition under affine constraints is not in Mathlib and is reusable beyond linear programming. Proofs of any milestone, including the "only if" half of the goal separately, are welcome contributions.

Selected references

  • R. J. Vanderbei, Linear Programming: Foundations and Extensions, 4th ed., International Series in Operations Research & Management Science 196, Springer, 2014, Chapter 17 and Exercise 10.7. https://doi.org/10.1007/978-1-4614-7630-6
  • A. V. Fiacco and G. P. McCormick, Nonlinear Programming: Sequential Unconstrained Minimization Techniques, Wiley, 1968. https://doi.org/10.1137/1.9781611971316
  • N. Karmarkar, A new polynomial-time algorithm for linear programming, Combinatorica 4 (1984), 373–395. https://doi.org/10.1007/BF02579150
  • N. Megiddo, Pathways to the optimal set in linear programming, in Progress in Mathematical Programming, Springer, 1989, 131–158. https://doi.org/10.1007/978-1-4613-9617-8_8
  • D. A. Bayer and J. C. Lagarias, The nonlinear geometry of linear programming I, II, Transactions of the AMS 314 (1989), 499–526 and 527–581. https://doi.org/10.1090/S0002-9947-1989-1005525-6
5 thms1 active userReviewed
Operations ResearchOptimizationProbability·Captain: mikedeng1

Introduction to the Scenario Approach III: The Risks of the Empirical Costs Follow an Ordered Dirichlet DistributionTextbook

Motivation

A scenario program replaces an uncertain optimization problem by its worst case over finitely many sampled instances. In its simplest form it reads

min⁡ν∈Rd−1[max⁡i=1,…,Nℓ(ν,δi)],\min_{\nu\in\mathbb R^{d-1}}\Big[\max_{i=1,\dots,N}\ell(\nu,\delta_i)\Big],ν∈Rd−1min​[i=1,…,Nmax​ℓ(ν,δi​)],

where ℓ(ν,δ)\ell(\nu,\delta)ℓ(ν,δ) is the cost of a decision ν\nuν when the uncertain parameter takes the value δ\deltaδ, and δ1,…,δN\delta_1,\dots,\delta_Nδ1​,…,δN​ are independent draws from an unknown probability P\mathbb PP. The classical guarantee of the scenario approach (Campi and Garatti, 2008) bounds the probability that a new instance produces a cost above the optimal value ℓ∗\ell^*ℓ∗, and it does so without any knowledge of P\mathbb PP.

That guarantee concerns a single number, ℓ∗\ell^*ℓ∗. Two scenario programs with the same NNN and the same optimal value can look very different at the solution: in one, most sampled costs lie just below ℓ∗\ell^*ℓ∗; in the other, they are widely scattered. The costs that do not determine the solution still carry information about how the cost of the chosen decision is distributed on future instances. Carè, Garatti and Campi (2015) showed that this information can be extracted with the same distribution-free character as the classical result, which is the subject of this mission. It is Chapter 8, §8.1 ("Probability box") of Campi and Garatti, Introduction to the Scenario Approach (SIAM/MOS 2018), the third mission of the series formalizing that book.

Timeline:

  • 2008: Campi and Garatti prove that the violation of the scenario solution is dominated by a beta distribution B(d,N−d+1)B(d,N-d+1)B(d,N−d+1), with equality for fully supported problems (doi:10.1137/07069821X).
  • 2015: Carè, Garatti and Campi prove that the risks of all empirical costs from index ddd on have a joint ordered Dirichlet law (doi:10.1137/130928546).
  • 2018: the book states the result as Theorem 8.4 and draws the probability box from it.

Setting

Let Δ\DeltaΔ be a measurable space with a probability P\mathbb PP, and ℓ:Rd−1×Δ→R\ell:\mathbb R^{d-1}\times\Delta\to\mathbb Rℓ:Rd−1×Δ→R a cost that is convex in ν\nuν for every δ\deltaδ (a standing assumption of the book). For a sample (δ1,…,δN)(\delta_1,\dots,\delta_N)(δ1​,…,δN​) of independent draws from P\mathbb PP, let ν∗\nu^*ν∗ be the solution of the program above and ℓ∗=max⁡iℓ(ν∗,δi)\ell^*=\max_i\ell(\nu^*,\delta_i)ℓ∗=maxi​ℓ(ν∗,δi​) its optimal value.

Empirical costs (Definition 8.1). Sort the costs of the solution on the sampled scenarios in decreasing order, ℓ1∗≥ℓ2∗≥⋯≥ℓN∗\ell^*_1\ge\ell^*_2\ge\dots\ge\ell^*_Nℓ1∗​≥ℓ2∗​≥⋯≥ℓN∗​; so ℓ1∗=ℓ∗\ell^*_1=\ell^*ℓ1∗​=ℓ∗.

Risk (Definition 8.2). For a decision ν\nuν and a level ℓ\ellℓ, R(ν,ℓ)=P{δ:ℓ(ν,δ)>ℓ}R(\nu,\ell)=\mathbb P\{\delta:\ell(\nu,\delta)>\ell\}R(ν,ℓ)=P{δ:ℓ(ν,δ)>ℓ}. The risk of the kkk-th empirical cost is Rk=R(ν∗,ℓk∗)R_k=R(\nu^*,\ell^*_k)Rk​=R(ν∗,ℓk∗​), and R1≤R2≤⋯≤RNR_1\le R_2\le\dots\le R_NR1​≤R2​≤⋯≤RN​.

Nondegeneracy (Definition 8.3). For every N≥dN\ge dN≥d, with probability 111, ℓd∗≠ℓd+1∗≠…≠ℓN∗\ell^*_d\ne\ell^*_{d+1}\ne\dots\ne\ell^*_Nℓd∗​=ℓd+1∗​=…=ℓN∗​. Costs with index below ddd are excluded because several scenarios typically attain the maximum at ν∗\nu^*ν∗.

Support constraints and full support (Definitions 5.1 and 5.4). In epigraph form, min⁡t\min tmint subject to t≥ℓ(ν,δi)t\ge\ell(\nu,\delta_i)t≥ℓ(ν,δi​), the constraint of scenario iii is a support constraint if removing it lowers the optimal value; the problem is fully supported if for every m≥dm\ge dm≥d the program with mmm scenarios has exactly ddd support constraints with probability 111.

The ordered Dirichlet distribution with parameters (d,1,…,1)(d,1,\dots,1)(d,1,…,1) is the law on {0≤αd≤⋯≤αN≤1}\{0\le\alpha_d\le\dots\le\alpha_N\le1\}{0≤αd​≤⋯≤αN​≤1} with density N!(d−1)!αdd−1\frac{N!}{(d-1)!}\alpha_d^{d-1}(d−1)!N!​αdd−1​.

Formalization targets

Goal: Theorem 8.4

Under nondegeneracy, for N≥dN\ge dN≥d and all εd,…,εN\varepsilon_d,\dots,\varepsilon_Nεd​,…,εN​,

PN{Rd≤εd,…,RN≤εN}=N!(d−1)!∫0εdαdd−1∫0εd+1 ⁣ ⁣⋯∫0εN1{0≤αd≤⋯≤αN≤1} dαN⋯dαd.\mathbb P^N\{R_d\le\varepsilon_d,\dots,R_N\le\varepsilon_N\}=\frac{N!}{(d-1)!}\int_0^{\varepsilon_d}\alpha_d^{d-1}\int_0^{\varepsilon_{d+1}}\!\!\cdots\int_0^{\varepsilon_N}\mathbf 1_{\{0\le\alpha_d\le\dots\le\alpha_N\le1\}}\,\mathrm d\alpha_N\cdots\mathrm d\alpha_d .PN{Rd​≤εd​,…,RN​≤εN​}=(d−1)!N!​∫0εd​​αdd−1​∫0εd+1​​⋯∫0εN​​1{0≤αd​≤⋯≤αN​≤1}​dαN​⋯dαd​.

This is an identity of joint distribution functions, not a bound, and it does not depend on ℓ\ellℓ or P\mathbb PP.

Milestones

  1. Theorem 3.7 for the min-max program: PN{R(ν∗,ℓ∗)>ε}≤∑i=0d−1(Ni)εi(1−ε)N−i\mathbb P^N\{R(\nu^*,\ell^*)>\varepsilon\}\le\sum_{i=0}^{d-1}\binom Ni\varepsilon^i(1-\varepsilon)^{N-i}PN{R(ν∗,ℓ∗)>ε}≤∑i=0d−1​(iN​)εi(1−ε)N−i for ε∈[0,1]\varepsilon\in[0,1]ε∈[0,1].
  2. Fully supported problems: ℓ∗=ℓd∗\ell^*=\ell^*_dℓ∗=ℓd∗​ with probability 111.
  3. Marginal of RdR_dRd​ (a corollary of the goal): PN{Rd≤ε}=1−∑i=0d−1(Ni)εi(1−ε)N−i\mathbb P^N\{R_d\le\varepsilon\}=1-\sum_{i=0}^{d-1}\binom Ni\varepsilon^i(1-\varepsilon)^{N-i}PN{Rd​≤ε}=1−∑i=0d−1​(iN​)εi(1−ε)N−i, the beta law B(d,N−d+1)B(d,N-d+1)B(d,N−d+1).

Significance

The result. The theorem controls the whole distribution function of the cost ℓ(ν∗,δ)\ell(\nu^*,\delta)ℓ(ν∗,δ) of the scenario solution on a new instance, not just one quantile of it. Discarding the extreme tails of the laws of Rd,…,RNR_d,\dots,R_NRd​,…,RN​ yields, with a prescribed confidence 1−β1-\beta1−β, a region (the book's "probability box", Figure 8.2) that contains the entire cumulative distribution function of ℓ(ν∗,δ)\ell(\nu^*,\delta)ℓ(ν∗,δ), computed from the sample alone. The first marginal recovers the classical Theorem 3.7, since ℓ∗≥ℓd∗\ell^*\ge\ell^*_dℓ∗≥ℓd∗​ makes the risk of ℓ∗\ell^*ℓ∗ at most RdR_dRd​.

Formalizing it. The theorem is proved in Carè, Garatti and Campi (2015); the book states it and gives no proof. No machine-checked version of the scenario approach, of its generalization theorem, or of ordered Dirichlet laws of risks is known to exist. A formalization would produce a checked proof of the distribution-free identity together with the combinatorial and measure-theoretic infrastructure (order statistics of sampled costs, laws of random risks) that the rest of scenario theory reuses. The milestones separate the classical beta bound, which is also the goal of the first mission of this series, from the new exact joint law.

Difficulty

The obvious attempt treats Rd,…,RNR_d,\dots,R_NRd​,…,RN​ as the order statistics of the uniform variables 1−F(ℓ(ν∗,δi))1-F(\ell(\nu^*,\delta_i))1−F(ℓ(ν∗,δi​)). That works only for d=1d=1d=1, when the decision space is a point and the costs are independent. For d≥2d\ge2d≥2 the decision ν∗\nu^*ν∗ is itself a function of the whole sample, so the sampled costs at ν∗\nu^*ν∗ are neither independent nor identically distributed, and the ddd-th cost is tied to the scenarios that determine the solution. The factor αdd−1\alpha_d^{d-1}αdd−1​ and the constant N!/(d−1)!N!/(d-1)!N!/(d−1)! encode exactly this dependence. Any argument must account for which scenarios are active at ν∗\nu^*ν∗ without assuming full support, since the theorem holds whether or not ℓ∗=ℓd∗\ell^*=\ell^*_dℓ∗=ℓd∗​.

Formalization scope

The decision space is EuclideanSpace ℝ (Fin n) and the book's ddd is n+1n+1n+1; the sample is ω : Fin N → Δ with law Measure.pi (fun _ => P). Empirical costs are read from Tuple.sort with kkk counted from 111; risks are real numbers (P {δ | c < ℓ ν δ}).toReal. The right-hand side of the goal is a Lebesgue integral over the box ∏k[0,εk]\prod_k[0,\varepsilon_k]∏k​[0,εk​] intersected with the ordered simplex, stated for all real εk\varepsilon_kεk​. The solution map ω↦ν∗\omega\mapsto\nu^*ω↦ν∗ is a hypothesis-constrained function, never an arbitrary map.

Implicit hypotheses of the book pinned down in the binders:

  • ℓ(⋅,δ)\ell(\cdot,\delta)ℓ(⋅,δ) is convex for every δ\deltaδ (p. 6).
  • Existence and uniqueness of the solution of the program for every sample size m≥1m\ge1m≥1 and every sample; the book's Assumption 3.6 says "every mmm", but the program with no scenario has no solution.
  • Nondegeneracy for every sample size m≥dm\ge dm≥d, not only for the NNN of the theorem, as Definition 8.3 is written.
  • Joint measurability of (ν,δ)↦ℓ(ν,δ)(\nu,\delta)\mapsto\ell(\nu,\delta)(ν,δ)↦ℓ(ν,δ) and measurability of the solution map (measurability is glossed over in the book, p. 6 footnote 1 and p. 33).
  • N≥dN\ge dN≥d, and ε∈[0,1]\varepsilon\in[0,1]ε∈[0,1] in the binomial-form statements.

A statement in which the solution is an arbitrary measurable map, or in which the nondegeneracy or existence hypothesis is unsatisfiable, would make the goal vacuous; the hypotheses here are met, for example, by ℓ(ν,δ)=∥ν−δ∥2\ell(\nu,\delta)=\|\nu-\delta\|^2ℓ(ν,δ)=∥ν−δ∥2 with a continuous law on Rn\mathbb R^{n}Rn, and for d=1d=1d=1 by any cost independent of ν\nuν with an atomless law.

A complete development needs: the scenario approach generalization theorem (reusable across this series), laws of order statistics of i.i.d. uniform variables, and the combinatorics of support sets of convex min-max programs. Proofs of the milestones, of the d=1d=1d=1 case of the goal, and of auxiliary facts about kthLargest are all welcome.

Selected references

  • M. C. Campi, S. Garatti, Introduction to the Scenario Approach, MOS-SIAM Series on Optimization 26, SIAM, 2018. doi:10.1137/1.9781611975444
  • A. Carè, S. Garatti, M. C. Campi, Scenario min-max optimization and the risk of empirical costs, SIAM J. Optim. 25(4):2061–2080, 2015. doi:10.1137/130928546
  • M. C. Campi, S. Garatti, The exact feasibility of randomized solutions of uncertain convex programs, SIAM J. Optim. 19:1211–1230, 2008. doi:10.1137/07069821X
10 thms1 active userReviewed
🏆Completed
Machine LearningProbability·Captain: Minghui

A Field Guide to Federated Optimization: Convex FedAvg ConvergenceResearch Paper

Why local training needs a convergence guarantee

Federated optimization studies learning when data and computation are spread across clients. Communicating after every stochastic gradient step can be costly, so clients often take several steps before averaging their models. The difficulty is that different clients can optimize different objective functions. Their models then move apart between communication rounds. A convergence guarantee must account for both stochastic gradient noise and this disagreement. Wang et al. give an explicit analysis of this tradeoff in A Field Guide to Federated Optimization, Section 6.1.

The relevant historical sequence is the introduction of FedAvg by McMahan et al. in 2017, the development of local-SGD convergence analyses reviewed by Wang et al., and the unified illustrative analysis in the 2021 field guide. The present target is that guide's known convex convergence theorem, rather than a new conjecture about arbitrary federated learning. The remaining open task is a Lean proof of the stated result. The mission is categorized as ResearchPaper: it formalizes a specific known result rather than proposing a new mathematical conjecture.

Setting: full participation with uniform weights

There are M≥1M\ge1M≥1 clients and a parameter vector in Rd\mathbb R^dRd. Client iii has a differentiable convex objective FiF_iFi​. The global objective is F(x)=M−1∑i=1MFi(x)F(x)=M^{-1}\sum_{i=1}^M F_i(x)F(x)=M−1∑i=1M​Fi​(x). Every local gradient is LLL-Lipschitz for the same L>0L>0L>0. Fix a global minimizer x⋆x^\starx⋆ of FFF and a deterministic initial model x0x_0x0​, and write D=∥x0−x⋆∥D=\|x_0-x^\star\|D=∥x0​−x⋆∥.

Every client participates in every round. Each of T≥1T\ge1T≥1 rounds consists of τ≥1\tau\ge1τ≥1 local steps with constant learning rate η\etaη. If xit,kx_i^{t,k}xit,k​ is client iii's state after kkk local steps of round ttt, its next state is xit,k+1=xit,k−ηgit,kx_i^{t,k+1}=x_i^{t,k}-\eta g_i^{t,k}xit,k+1​=xit,k​−ηgit,k​. At the next round all clients restart from the average of the preceding round's terminal states. The shadow iterate is xˉt,k=M−1∑ixit,k\bar x^{t,k}=M^{-1}\sum_i x_i^{t,k}xˉt,k=M−1∑i​xit,k​; it is defined even at local steps where clients do not communicate.

On a probability space (Ω,A,P)(\Omega,\mathcal A,\mathbb P)(Ω,A,P), the history before a step contains all past oracle draws. The stochastic gradients are conditionally unbiased, their conditional squared errors have expectation at most σ2\sigma^2σ2, and the different clients' current gradients are conditionally independent. The heterogeneity bound is ∥∇Fi(x)−∇F(x)∥≤ζ\|\nabla F_i(x)-\nabla F(x)\|\le\zeta∥∇Fi​(x)−∇F(x)∥≤ζ for every client and every point, where σ,ζ≥0\sigma,\zeta\ge0σ,ζ≥0. These are the assumptions of Section 6.1.1, PDF p. 40, equations (11)–(14), with the history and independence convention used explicitly in Appendix D.1 immediately after equation (27), PDF p. 87.

Formalization targets

The principal goal is Theorem 1, equation (15). For 0<η≤1/(4L)0<\eta\le1/(4L)0<η≤1/(4L), establish

E ⁣[1τT∑t=0T−1∑k=1τ(F(xˉt,k)−F(x⋆))]≤D22ητT+ησ2M+4τη2Lσ2+18τ2η2Lζ2.\mathbb E\!\left[\frac1{\tau T}\sum_{t=0}^{T-1}\sum_{k=1}^{\tau} \bigl(F(\bar x^{t,k})-F(x^\star)\bigr)\right] \le \frac{D^2}{2\eta\tau T}+\frac{\eta\sigma^2}{M} +4\tau\eta^2L\sigma^2+18\tau^2\eta^2L\zeta^2.E[τT1​t=0∑T−1​k=1∑τ​(F(xˉt,k)−F(x⋆))]≤2ητTD2​+Mησ2​+4τη2Lσ2+18τ2η2Lζ2.

Two source milestones describe the intermediate results. Lemma 1 bounds the conditional average loss within a round by the decrease of squared distance to the minimizer, plus a noise term and a sum of client disagreements. Lemma 2 bounds each conditional squared disagreement by 18τ2η2ζ2+4τη2σ218\tau^2\eta^2\zeta^2+4\tau\eta^2\sigma^218τ2η2ζ2+4τη2σ2. Both are stated on PDF p. 41, Section 6.1.2, with proofs in Appendix D, PDF pp. 86–88. They remain genuine open proof obligations; the algorithm model does not assume either estimate.

An additional milestone records the same theorem's tuned-step consequence, equations (16)–(17), in the regime D,σ,ζ>0D,\sigma,\zeta>0D,σ,ζ>0. With

η=min⁡{14L,MDτTσ,D2/3τ2/3T1/3L1/3σ2/3,D2/3τT1/3L1/3ζ2/3},\eta=\min\left\{\frac1{4L}, \frac{\sqrt M D}{\sqrt\tau\sqrt T\sigma}, \frac{D^{2/3}}{\tau^{2/3}T^{1/3}L^{1/3}\sigma^{2/3}}, \frac{D^{2/3}}{\tau T^{1/3}L^{1/3}\zeta^{2/3}}\right\},η=min{4L1​,τ​T​σM​D​,τ2/3T1/3L1/3σ2/3D2/3​,τT1/3L1/3ζ2/3D2/3​},

the same expected loss is at most

2LD2τT+2σDMτT+5L1/3σ2/3D4/3τ1/3T2/3+19L1/3ζ2/3D4/3T2/3.\frac{2LD^2}{\tau T}+\frac{2\sigma D}{\sqrt{M\tau T}} +\frac{5L^{1/3}\sigma^{2/3}D^{4/3}}{\tau^{1/3}T^{2/3}} +\frac{19L^{1/3}\zeta^{2/3}D^{4/3}}{T^{2/3}}.τT2LD2​+MτT​2σD​+τ1/3T2/35L1/3σ2/3D4/3​+T2/319L1/3ζ2/3D4/3​.

The principal goal includes zero-noise and zero-heterogeneity cases. The extra positivity conditions apply only to the printed tuned-step formula, whose denominators otherwise require separate conventions.

What this establishes

The bound separates an initial-distance term, a noise term improved by the number of clients, and two costs of local updates. It quantifies how local work interacts with stochastic noise and differing client objectives. Its conclusion concerns the average objective gap along the post-update shadow sequence; it does not assert the same bound for every last iterate or for the average of client losses. These distinctions follow directly from the quantity defined in equation (14).

A completed formalization would provide reusable checked components for stochastic optimization: finite client averages, history-conditioned oracle assumptions, per-round potential estimates, and disagreement bounds. The paper supplies the mathematical proof; this draft supplies checked statements and definitions. No convergence proof is claimed by creating or compiling the proposal.

Where the difficulty lies

The average update evaluates each gradient at its own client's state. It therefore does not directly equal a centralized stochastic-gradient step at the shadow iterate. A proof must control that discrepancy quantitatively, preserve the conditioning on the round's starting history, and justify the 1/M1/M1/M noise improvement using independent client sampling. Ignoring the sampling relationship can invalidate the advertised bound even for scalar quadratic objectives.

Formalization scope

The model uses finite-dimensional real Euclidean space, including the harmless zero-dimensional case, and a standard Borel probability space. A filtration indexed by tτ+kt\tau+ktτ+k records the full past. Local states and gradients carry explicit measurability and finite-second-moment conditions. These probability conventions support actual Bochner and conditional expectations; integrals are not treated as arbitrary total functions without analytic obligations.

The local objectives, global objective, gradients, iterates, and shadow averages are concrete functions. The model assumes neither a drift bound nor a progress bound. Source hypotheses are uniform in the model point, and the minimizer is an actual minimizer of the averaged objective. Unequal weighting, partial client participation, nonconvex objectives, adaptive step sizes, and privacy mechanisms are outside this particular theorem. Contributions should prove the named source lemmas or their necessary analytic infrastructure while retaining these statements.

Selected references

  • Jianyu Wang et al., A Field Guide to Federated Optimization, 2021, arXiv:2107.06917v1, Section 6.1.1–6.1.2, PDF pp. 40–41; Appendix D, PDF pp. 86–88.
  • H. Brendan McMahan et al., Communication-Efficient Learning of Deep Networks from Decentralized Data, AISTATS 2017, arXiv:1602.05629, the FedAvg algorithm cited by the field guide. This is historical context, not an additional target.
5 thms1 active userReviewed
🏆Completed
Functional AnalysisOptimization·Captain: Shuze Chen

Vector Space Methods IV: Hahn–Banach and Minimum Norm DualityTextbook

Motivation

Chapter 5 of Luenberger's Optimization by Vector Space Methods (Wiley, 1969) carries the minimum norm theory of Chapter 3 (Mission I of this series) from Hilbert space to arbitrary real normed spaces. The inner product is gone, so orthogonal projection is no longer available; its role is taken over by the Hahn–Banach theorem, in two classical forms. The extension form generalizes the projection theorem and yields a duality principle equating a minimum norm problem in a space XXX with a maximization problem in its dual X∗X^*X∗; the geometric form (separating hyperplanes) extends that duality from subspaces to convex sets. These duality theorems are the backbone of the optimization theory in the remainder of the book — conjugate functionals (Ch. 7) and Lagrange duality (Ch. 8) both trace back to them.

Setting

Throughout, XXX is a real normed linear space. A linear functional fff on XXX is bounded if ∣f(x)∣≤M∥x∥|f(x)| \le M\|x\|∣f(x)∣≤M∥x∥ for some constant MMM and all xxx; the least such MMM is the norm ∥f∥\|f\|∥f∥. The (normed) dual X∗X^*X∗ is the space of bounded (equivalently, continuous) linear functionals with this norm; ⟨x,x∗⟩\langle x, x^*\rangle⟨x,x∗⟩ denotes x∗(x)x^*(x)x∗(x). A functional p:X→Rp : X \to \mathbb{R}p:X→R is sublinear when p(x+y)≤p(x)+p(y)p(x+y) \le p(x) + p(y)p(x+y)≤p(x)+p(y) and p(αx)=α p(x)p(\alpha x) = \alpha\, p(x)p(αx)=αp(x) for α>0\alpha > 0α>0. Vectors x∈Xx \in Xx∈X and x∗∈X∗x^* \in X^*x∗∈X∗ are aligned when ⟨x,x∗⟩=∥x∗∥ ∥x∥\langle x, x^*\rangle = \|x^*\|\,\|x\|⟨x,x∗⟩=∥x∗∥∥x∥, and orthogonal when ⟨x,x∗⟩=0\langle x, x^*\rangle = 0⟨x,x∗⟩=0; for S⊆XS \subseteq XS⊆X, the complement S⊥⊆X∗S^\perp \subseteq X^*S⊥⊆X∗ consists of the functionals vanishing on SSS, and for U⊆X∗U \subseteq X^*U⊆X∗, ⊥U⊆X{}^\perp U \subseteq X⊥U⊆X consists of the vectors annihilated by every member of UUU. A hyperplane is a maximal proper linear variety; closed hyperplanes are the level sets {x:⟨x,x∗⟩=c}\{x : \langle x, x^*\rangle = c\}{x:⟨x,x∗⟩=c} of nonzero bounded functionals. The support functional of a convex set KKK is h(x∗)=sup⁡k∈K ⟨k,x∗⟩h(x^*) = \sup_{k \in K}\, \langle k, x^*\rangleh(x∗)=supk∈K​⟨k,x∗⟩.

Formalization targets

The goal is §5.13 Theorem 1 (Minimum Norm Duality): if x1∈Xx_1 \in Xx1​∈X has distance d>0d > 0d>0 from a convex set KKK with support functional hhh, then

d  =  inf⁡x∈K∥x−x1∥  =  max⁡∥x∗∥≤1 [⟨x1,x∗⟩−h(x∗)],d \;=\; \inf_{x \in K} \|x - x_1\| \;=\; \max_{\|x^*\| \le 1}\ \big[\langle x_1, x^*\rangle - h(x^*)\big],d=x∈Kinf​∥x−x1​∥=∥x∗∥≤1max​ [⟨x1​,x∗⟩−h(x∗)],

the maximum on the right being achieved by some x0∗x_0^*x0∗​; and if the infimum is achieved by x0∈Kx_0 \in Kx0​∈K, then −x0∗-x_0^*−x0∗​ is aligned with x0−x1x_0 - x_1x0​−x1​.

The milestones trace the chapter's route there: boundedness ⇔\Leftrightarrow⇔ continuity (§5.2); the Hahn–Banach theorem in sublinear form (§5.4 Theorem 1) with its norm-preserving extension and norming-functional corollaries; the annihilator identity ⊥(M⊥)=M{}^\perp(M^\perp) = M⊥(M⊥)=M for closed subspaces (§5.7 Theorem 1); the two subspace duality theorems and the alignment characterization of best approximations (§5.8 — the chapter's principal results); and the geometric form: Mazur's separation theorem, the support theorem, and Eidelheit's separation theorem (§5.12).

Significance

The §5.8 duality theorems are the exact normed-space analogue of the projection theorem: existence transfers to the dual problem (minimum norm problems should be formulated in a dual space to guarantee solutions — the chapter's methodological moral), orthogonality becomes alignment, and infinite-dimensional problems with finitely many constraints reduce to finite-dimensional dual problems. The geometric form underpins all of convex duality.

All results are classical and proved in the source. Mathlib contains the Hahn–Banach extension theorem and point/convex separation theorems, so several milestones are exercises in connecting Luenberger's formulations to existing library lemmas; the two §5.8 duality theorems, the alignment corollary, and the §5.13 convex duality theorem have no direct Mathlib counterpart and are the mission's genuinely new content.

Difficulty

Degenerate cases are the trap throughout. In §5.8 Corollary 1 the "only if" direction fails literally when MMM is dense and x∈Mx \in Mx∈M (then M⊥={0}M^\perp = \{0\}M⊥={0} and no nonzero aligned functional exists); the formalization therefore carries the hypothesis x∉M‾x \notin \overline{M}x∈/M. In the separation theorems the strict inequality holds only on the interior of the convex set — on the set itself only ≤\le≤ survives — and nonemptiness hypotheses (of the interior, of K2K_2K2​, of the variety) are what make the "nonzero functional" claims true; dropping any of them creates false statements in trivial spaces. In §5.13 the support functional may take the value +∞+\infty+∞, so the dual maximum is formalized by two quantified inequalities (the witness achieves ddd; no admissible functional exceeds ddd) rather than by a real-valued supremum. The infimum in the primal problems need not be attained — attainment appears only as a hypothesis in the alignment clauses.

Formalization scope

Real scalars throughout. The dual space is represented concretely as continuous linear maps X →L[ℝ] ℝ, and annihilators are written as explicit quantified conditions rather than named subspaces. Five notions the chapter needs and Mathlib lacks are published as definitions and used by the statements rather than inlined: alignment (⟨x,x∗⟩=∥x∗∥ ∥x∥\langle x, x^*\rangle = \|x^*\|\,\|x\|⟨x,x∗⟩=∥x∗∥∥x∥), the support functional (h(x∗)=sup⁡k∈K⟨k,x∗⟩h(x^*) = \sup_{k \in K} \langle k, x^*\rangleh(x∗)=supk∈K​⟨k,x∗⟩, valued in the extended reals since it may be infinite), the total variation of a function on an interval, the normalized space NBV[a,b]NBV[a,b]NBV[a,b], and the Riemann–Stieltjes integral (defined relationally, so that no existence claim is built into the definition). The Minkowski functional needed for Mazur's theorem is Mathlib's gauge. Minimum distances are infima ⨅ over coerced sets or submodules; in §5.8 Theorem 2 the dual-side supremum is a real sSup over {⟨x,x∗⟩:x∈M, ∥x∥≤1}\{\langle x, x^*\rangle : x \in M,\ \|x\| \le 1\}{⟨x,x∗⟩:x∈M, ∥x∥≤1}, which is nonempty and bounded. Sublinearity in §5.4 is hypothesized exactly as in the source (subadditivity plus positive homogeneity plus continuity). Linear varieties are parametrized as x0+Mx_0 + Mx0​+M with MMM a Submodule ℝ X. No completeness of XXX is assumed anywhere — the chapter's results are genuinely about normed spaces, and Hahn–Banach needs no completeness. The concrete dual of C[a,b]C[a,b]C[a,b] (§5.5) is in scope, and carries most of the mission's new infrastructure: Mathlib has the property of bounded variation (eVariationOn) but no total-variation norm, no normalized space NBV[a,b]NBV[a,b]NBV[a,b], and no Riemann–Stieltjes integral — its StieltjesFunction is the different object of a monotone right-continuous function inducing a Borel measure, and its Riesz–Markov–Kakutani development represents positive functionals on Cc(X)C_c(X)Cc​(X) by measures, not bounded functionals on C[a,b]C[a,b]C[a,b] by functions of bounded variation. This mission therefore publishes those notions as definitions and states the representation theorem in both directions. §5.3 (the Riesz–Fréchet theorem, i.e. self-duality of Hilbert space) is the one omission: Mathlib's InnerProductSpace.toDual already provides it. §5.6 (second dual, reflexivity) is definitional and likewise present in Mathlib.

Selected references

  • David G. Luenberger, Optimization by Vector Space Methods, John Wiley & Sons, 1969. Chapter 5, pp. 103–142. ISBN 0-471-55359-X.
  • H. Hahn, Über lineare Gleichungssysteme in linearen Räumen, J. Reine Angew. Math. 157 (1927), 214–229; S. Banach, Sur les fonctionnelles linéaires II, Studia Math. 1 (1929), 223–239.
  • S. Mazur, Über konvexe Mengen in linearen normierten Räumen, Studia Math. 4 (1933), 70–84.
17 thms1 active userReviewed
Operations ResearchOptimizationProbability·Captain: mikedeng1

Optimality and Duality Theory for Stochastic Optimization Problems with Nonlinear Dominance Constraints 2: With Finite Scenarios and Slater's Condition, Piecewise-Linear Utilities Are MultipliersResearch Paper

Motivation

Second-order stochastic dominance constraints let a decision maker require that a random outcome of a decision be preferred to a fixed benchmark outcome by every risk-averse expected-utility maximizer, without choosing a utility function in advance. Dentcheva and Ruszczyński introduced optimization under such constraints in Optimization with stochastic dominance constraints (SIAM J. Optim., 2003), for the case where the decision enters the outcome linearly (the pure-dominance case). Their follow-up paper, Optimality and duality theory for stochastic optimization problems with nonlinear dominance constraints (Math. Program., 2004), allows the decision to affect many random outcomes in a nonlinear, concave way, and derives optimality and duality theory in which the Lagrange multipliers of the dominance constraints are utility functions.

In applications (portfolio selection against a benchmark index is the paper's own example in §6) the probability space is a finite set of scenarios. Section 5 of the paper specialises the theory to that case. This mission formalizes that section: the reduction of the dominance constraints to finitely many inequalities, and the optimality and duality theorems (Theorems 6 and 7) in which the multipliers become piecewise-linear concave utilities.

Setting

There are nnn scenarios ω1,…,ωn\omega_1,\dots,\omega_nω1​,…,ωn​ with probabilities pj≥0p_j \ge 0pj​≥0, ∑jpj=1\sum_j p_j = 1∑j​pj​=1, and mmm benchmark constraints, indexed by i∈I={1,…,m}i \in I = \{1,\dots,m\}i∈I={1,…,m}; J={1,…,n}J = \{1,\dots,n\}J={1,…,n}. A decision zzz ranges over a convex set Z⊆RNZ \subseteq \mathbb R^NZ⊆RN. For each scenario jjj, hj:RN→Rh_j:\mathbb R^N\to\mathbb Rhj​:RN→R is the objective contribution and gij:RN→Rg_{ij}:\mathbb R^N\to\mathbb Rgij​:RN→R the iiith outcome, all concave. The benchmark YiY_iYi​ has realizations yijy_{ij}yij​. Write (t)+=max⁡(t,0)(t)_+=\max(t,0)(t)+​=max(t,0).

The second-order dominance of a finitely distributed XiX_iXi​ (realizations xijx_{ij}xij​) over YiY_iYi​ on an interval [ai,bi][a_i,b_i][ai​,bi​] reads

∑jpj(η−xij)+≤∑jpj(η−yij)+for all η∈[ai,bi].(36)\sum_{j} p_j(\eta - x_{ij})_+ \le \sum_j p_j(\eta-y_{ij})_+ \quad\text{for all } \eta\in[a_i,b_i]. \tag{36}j∑​pj​(η−xij​)+​≤j∑​pj​(η−yij​)+​for all η∈[ai​,bi​].(36)

The split-variable problem (38)–(41) is

max⁡∑j=1npjhj(z)s.t.∑jpj(yik−xij)+≤∑jpj(yik−yij)+,xik≤gik(z),z∈Z,\max \sum_{j=1}^n p_j h_j(z)\quad\text{s.t.}\quad \sum_{j} p_j(y_{ik}-x_{ij})_+ \le \sum_j p_j(y_{ik}-y_{ij})_+,\quad x_{ik}\le g_{ik}(z),\quad z\in Z,maxj=1∑n​pj​hj​(z)s.t.j∑​pj​(yik​−xij​)+​≤j∑​pj​(yik​−yij​)+​,xik​≤gik​(z),z∈Z,

for all i∈Ii\in Ii∈I, k∈Jk\in Jk∈J, over zzz and X=(xij)∈RmnX=(x_{ij})\in\mathbb R^{mn}X=(xij​)∈Rmn. The Slater condition asks for z~∈relint⁡Z\tilde z \in \operatorname{relint} Zz~∈relintZ and X~\tilde XX~ satisfying the dominance constraints (39) with x~ik<gik(z~)\tilde x_{ik} < g_{ik}(\tilde z)x~ik​<gik​(z~) for all i,ki,ki,k.

The utility set ViV_iVi​ consists of the functions u:R→Ru:\mathbb R\to\mathbb Ru:R→R that are concave, nondecreasing, piecewise linear with break points only at the yiky_{ik}yik​, and zero on [max⁡kyik,∞)[\max_k y_{ik},\infty)[maxk​yik​,∞). With θij≥0\theta_{ij}\ge 0θij​≥0 multipliers for the splitting constraints xij≤gij(z)x_{ij}\le g_{ij}(z)xij​≤gij​(z), the Lagrangian is

L(z,X,u,θ)=∑j=1npj[hj(z)+∑i=1mθijgij(z)]+∑i=1m∑j=1npj[ui(xij)−ui(yij)−θijxij].(42)L(z,X,u,\theta) = \sum_{j=1}^n p_j\Big[h_j(z)+\sum_{i=1}^m\theta_{ij}g_{ij}(z)\Big]+\sum_{i=1}^m\sum_{j=1}^n p_j\big[u_i(x_{ij})-u_i(y_{ij})-\theta_{ij}x_{ij}\big]. \tag{42}L(z,X,u,θ)=j=1∑n​pj​[hj​(z)+i=1∑m​θij​gij​(z)]+i=1∑m​j=1∑n​pj​[ui​(xij​)−ui​(yij​)−θij​xij​].(42)

Multipliers μik\mu_{ik}μik​ of the inequalities (39) generate the utility ui(t)=−∑kμik(yik−t)+u_i(t)=-\sum_k\mu_{ik}(y_{ik}-t)_+ui​(t)=−∑k​μik​(yik​−t)+​ (46). The dual functional is D(u,θ)=sup⁡z∈Z, XL(z,X,u,θ)D(u,\theta)=\sup_{z\in Z,\,X}L(z,X,u,\theta)D(u,θ)=supz∈Z,X​L(z,X,u,θ) (47).

Formalization targets

Goal: Theorem 6

Under the Slater condition, (z^,X^)(\hat z,\hat X)(z^,X^) optimal for (38)–(41) implies that there are u^i∈Vi\hat u_i\in V_iu^i​∈Vi​ and θ^≥0\hat\theta\ge 0θ^≥0 with

L(z^,X^,u^,θ^)=max⁡(z,X)∈Z×RmnL(z,X,u^,θ^),∑jpj[u^i(x^ij)−u^i(yij)]=0,θ^ij(x^ij−gij(z^))=0;L(\hat z,\hat X,\hat u,\hat\theta)=\max_{(z,X)\in Z\times\mathbb R^{mn}}L(z,X,\hat u,\hat\theta),\qquad \sum_j p_j[\hat u_i(\hat x_{ij})-\hat u_i(y_{ij})]=0,\qquad \hat\theta_{ij}(\hat x_{ij}-g_{ij}(\hat z))=0;L(z^,X^,u^,θ^)=(z,X)∈Z×Rmnmax​L(z,X,u^,θ^),j∑​pj​[u^i​(x^ij​)−u^i​(yij​)]=0,θ^ij​(x^ij​−gij​(z^))=0;

conversely, these conditions together with feasibility imply optimality.

Milestones

  1. Lemma 2 (p. 15): if ai≤yij≤bia_i\le y_{ij}\le b_iai​≤yij​≤bi​, then (36) is equivalent to the mnmnmn inequalities (37) at the realizations η=yik\eta=y_{ik}η=yik​, and also to (36) on the whole line.
  2. Eq. (46) (p. 17): for any μ\muμ, the standard Lagrangian Λ(z,X,μ,θ)\Lambda(z,X,\mu,\theta)Λ(z,X,μ,θ) equals L(z,X,u,θ)L(z,X,u,\theta)L(z,X,u,θ) with uuu given by (46).
  3. p. 18: for μi≥0\mu_i\ge0μi​≥0, the utility (46) lies in ViV_iVi​.
  4. pp. 16–17: under Slater, an optimal solution admits Kuhn–Tucker multipliers μ≥0\mu\ge0μ≥0, θ≥0\theta\ge0θ≥0 for (38)–(41) with complementarity.
  5. p. 18: every v∈Viv\in V_iv∈Vi​ is of the form (46) with μi≥0\mu_i\ge0μi​≥0.
  6. Theorem 7 (p. 18), after the goal: the dual problem min⁡{D(u,θ):u∈V1×⋯×Vm, θ≥0}\min\{D(u,\theta): u\in V_1\times\dots\times V_m,\ \theta\ge0\}min{D(u,θ):u∈V1​×⋯×Vm​, θ≥0} has a solution and no duality gap.

Significance

Theorem 6 says that, for finitely many scenarios, the infinite-dimensional multiplier of the general theory (a concave utility in a cone of functions, Theorem 2 of the paper) can always be taken piecewise linear with kinks exactly at the benchmark's realizations. The multiplier space becomes finite-dimensional, ViV_iVi​ is a polyhedral cone, and the dual problem of Theorem 7 is a finite-dimensional convex program. The paper's decomposition (49)–(51) of the dual functional and its numerical method in §6 rest on this. Lemma 2 is the standard reduction that makes dominance against a finitely distributed benchmark a finite set of polyhedral constraints, used throughout the later literature on dominance-constrained portfolio optimization.

The results are proved in the paper; none of them is formalized. The mission produces machine-checked statements of the finite-scenario theory, a Lean model of the utility set ViV_iVi​ and of the correspondence between nonnegative multipliers and piecewise-linear utilities, and a Kuhn–Tucker theorem for concave programs with polyhedral constraints and a relative-interior Slater point.

Difficulty

The obvious route to Theorem 6 is to invoke a Kuhn–Tucker theorem. The available formal versions require every inequality constraint to hold strictly at the Slater point and range over all of RN\mathbb R^NRN. Neither fits: the dominance constraint at the smallest realization yi,[1]y_{i,[1]}yi,[1]​ has right-hand side 000 and a nonnegative left-hand side, so it can never hold strictly, and ZZZ may be lower-dimensional (a simplex), so only its relative interior is available. The polyhedral structure of (39) must be used, as in Rockafellar's Theorem 28.2. The second obstacle is the converse direction of the multiplier–utility correspondence: a utility in ViV_iVi​ must be written as a nonnegative combination of the kinks (yik−t)+(y_{ik}-t)_+(yik​−t)+​, which requires handling repeated realizations and the one-sided slopes at each break point.

Formalization scope

  • RN\mathbb R^NRN is Fin N → ℝ; XXX, θ\thetaθ, μ\muμ are Fin m → Fin n → ℝ; expectations are finite sums and positive parts are max t 0. No measure theory is used.
  • Probabilities satisfy pj≥0p_j\ge0pj​≥0, ∑jpj=1\sum_jp_j=1∑j​pj​=1; pj=0p_j=0pj​=0 is allowed, as on the page.
  • Standing assumptions of p. 2 are explicit hypotheses: ZZZ convex and hjh_jhj​, gijg_{ij}gij​ concave on RN\mathbb R^NRN. Continuity is not stated, since finite concave functions on RN\mathbb R^NRN are continuous.
  • The relative interior is intrinsicInterior ℝ Z, not the topological interior. In the Slater condition only the splitting constraints are strict; the dominance constraints hold non-strictly.
  • ViV_iVi​ is defined by concavity, monotonicity, affinity on every interval whose interior contains no yiky_{ik}yik​, and u=0u=0u=0 on [max⁡kyik,∞)[\max_ky_{ik},\infty)[maxk​yik​,∞). This last clause is the page's u(yi,[n])=0u(y_{i,[n]})=0u(yi,[n]​)=0 combined with Vi⊂U1([ai,bi])V_i\subset\mathcal U_1([a_i,b_i])Vi​⊂U1​([ai​,bi​]). No positive slope is required, because the printed "c>0c>0c>0" in U1\mathcal U_1U1​ is a misprint for c≥0c\ge0c≥0.
  • "max" in (43) is an attained maximum over all of Z×RmnZ\times\mathbb R^{mn}Z×Rmn, with no constraints on XXX. The dual functional (47) is an EReal supremum.
  • Theorem 6 keeps the Slater condition as a hypothesis of the whole statement, as printed, although its converse part does not use it.
  • A trivializing formalization is ruled out. ViV_iVi​ is not defined as the set of functions of the form (46), which would make milestones 3 and 5 true by definition. Slater does not require strict dominance constraints, which would make it unsatisfiable. A sorry-free check confirms that the goal's hypotheses hold on an instance (n=2n=2n=2, Z=[0,1]Z=[0,1]Z=[0,1]).
  • Reusable beyond this mission: the Kuhn–Tucker theorem with polyhedral constraints and relative-interior Slater point (milestone 4), and Lemma 2. Proofs of any item, and alternative proofs of the goal that avoid milestone 4, are welcome.
  • The pure-dominance case is the earlier paper of Dentcheva–Ruszczyński (2003). The function F2F_2F2​ and its expected-shortfall form are due to Ogryczak–Ruszczyński. The general Lagrange duality on the platform (ConvexOptimization.slater_strong_duality, Boyd–Vandenberghe §5.3.2) assumes a strict Slater point for every constraint and no set constraint, so it does not cover milestone 4.

Selected references

  • D. Dentcheva, A. Ruszczyński, Optimality and duality theory for stochastic optimization problems with nonlinear dominance constraints, Math. Program., 2004 (cited here from the authors' revised manuscript, April 2003). https://doi.org/10.1007/s10107-003-0453-z
  • D. Dentcheva, A. Ruszczyński, Optimization with stochastic dominance constraints, SIAM J. Optim. 14 (2003) 548–566. https://doi.org/10.1137/S1052623402420528
  • W. Ogryczak, A. Ruszczyński, Dual stochastic dominance and related mean-risk models, SIAM J. Optim. 13 (2002) 60–78. https://doi.org/10.1137/S1052623400375075
  • R. T. Rockafellar, Convex Analysis, Princeton University Press, 1970, §28. https://doi.org/10.1515/9781400873173
8 thms0 active usersReviewed
Operations ResearchOptimizationProbability·Captain: mikedeng1

Optimality and Duality Theory for Stochastic Optimization Problems with Nonlinear Dominance Constraints 1: Under Uniform Dominance, Optimal Solutions Have Concave Utility and L∞ MultipliersResearch Paper

Motivation

Stochastic programs often optimize a decision that changes several random outcomes at once. A reference outcome may be acceptable even when no fixed threshold captures its risk: one wants the new outcome to be preferable under every increasing concave assessment of gains. Second order stochastic dominance expresses that comparison. Dentcheva and Ruszczyński study optimization with several such constraints, each imposed on a nonlinear outcome operator, and show how the constraint multipliers can be represented by utility functions rather than scalar penalties (Dentcheva–Ruszczyński, 2004). Their earlier paper, Optimization with stochastic dominance constraints, treats the pure dominance case without the nonlinear decision map; the present result adds decision dependent outcomes, multiple constraints, and split variables. Ogryczak and Ruszczyński's second performance function supplies the stochastic order used here (Ogryczak–Ruszczyński, 2002).

The utility interpretation matters when a modeler wants a certificate explaining why a solution satisfies a risk preference expressed by dominance. The theorem identifies a concave utility for each binding dominance constraint and an essentially bounded multiplier for each comparison between the split outcome and the outcome produced by the decision. The source is a revised April 2003 author manuscript, later published in Mathematical Programming in 2004; the page and equation numbers below follow that manuscript (author manuscript).

Setting

Work on a probability space (Ω,F,P)(\Omega,\mathcal F,P)(Ω,F,P). An integrable random outcome is a measurable real function with finite expected absolute value; L1\mathcal L^1L1 denotes these outcomes, and L∞\mathcal L^\inftyL∞ denotes essentially bounded ones. The decisions lie in a convex set ZZZ inside a separable locally convex Hausdorff real vector space Z\mathcal ZZ. An integrable objective outcome H(z)H(z)H(z) and integrable constraint outcomes Gi(z)G_i(z)Gi​(z) depend continuously in the L1\mathcal L^1L1 norm on zzz. Almost every realized map z↦H(z)(ω)z\mapsto H(z)(\omega)z↦H(z)(ω) and z↦Gi(z)(ω)z\mapsto G_i(z)(\omega)z↦Gi​(z)(ω) is concave and continuous on all of Z\mathcal ZZ. Fixed integrable outcomes YiY_iYi​ serve as references; the iiith comparison is required over a bounded interval [ai,bi][a_i,b_i][ai​,bi​].

For an outcome XXX, its second performance function is the area below its distribution function:

F2(X;η)=∫−∞ηP{X≤ξ} dξ.F_2(X;\eta)=\int_{-\infty}^{\eta}P\{X\le\xi\}\,d\xi.F2​(X;η)=∫−∞η​P{X≤ξ}dξ.

The split program (11)–(14) chooses z∈Zz\in Zz∈Z and X=(X1,…,Xm)∈(L1)mX=(X_1,\ldots,X_m)\in(\mathcal L^1)^mX=(X1​,…,Xm​)∈(L1)m to maximize EH(z)\mathbb E H(z)EH(z), subject to F2(Xi;η)≤F2(Yi;η)F_2(X_i;\eta)\le F_2(Y_i;\eta)F2​(Xi​;η)≤F2​(Yi​;η) for every η∈[ai,bi]\eta\in[a_i,b_i]η∈[ai​,bi​], and Xi≤Gi(z)X_i\le G_i(z)Xi​≤Gi​(z) almost surely. Larger outcomes are preferred, so a dominating XiX_iXi​ has the smaller F2F_2F2​ curve. The split variables expose the dominance and decision coupling as separate constraints (manuscript, pp. 3–4).

The utility cone U1([a,b])\mathcal U_1([a,b])U1​([a,b]) consists of concave nondecreasing functions u:R→Ru:\mathbb R\to\mathbb Ru:R→R that vanish for t≥bt\ge bt≥b and are affine with a nonnegative slope for t≤at\le at≤a. Given uiu_iui​ in these cones and θi∈L∞\theta_i\in\mathcal L^\inftyθi​∈L∞, the Lagrangian is

L(z,X,u,θ)=E ⁣[H(z)+∑i=1m(ui(Xi)−ui(Yi)+θi(Gi(z)−Xi))].L(z,X,u,\theta)=\mathbb E\!\left[H(z)+\sum_{i=1}^m\bigl(u_i(X_i)-u_i(Y_i)+\theta_i(G_i(z)-X_i)\bigr)\right].L(z,X,u,θ)=E[H(z)+i=1∑m​(ui​(Xi​)−ui​(Yi​)+θi​(Gi​(z)−Xi​))].

Uniform dominance means one decision z~∈Z\tilde z\in Zz~∈Z makes every dominance inequality uniformly strict on its interval: for each iii, F2(Yi;η)−F2(Gi(z~);η)F_2(Y_i;\eta)-F_2(G_i(\tilde z);\eta)F2​(Yi​;η)−F2​(Gi​(z~);η) has a positive lower bound over [ai,bi][a_i,b_i][ai​,bi​] (Definition 1, p. 7).

Formalization targets

Utility and bounded multiplier characterization

Theorem 2 is the goal. Under uniform dominance, every optimum (z^,X^)(\hat z,\hat X)(z^,X^) of the split program admits u^i∈U1([ai,bi])\hat u_i\in\mathcal U_1([a_i,b_i])u^i​∈U1​([ai​,bi​]) and nonnegative θ^i∈L∞\hat\theta_i\in\mathcal L^\inftyθ^i​∈L∞ with

L(z^,X^,u^,θ^)=max⁡z∈Z, X∈(L1)mL(z,X,u^,θ^),L(\hat z,\hat X,\hat u,\hat\theta)=\max_{z\in Z,\,X\in(\mathcal L^1)^m}L(z,X,\hat u,\hat\theta),L(z^,X^,u^,θ^)=z∈Z,X∈(L1)mmax​L(z,X,u^,θ^), Eu^i(X^i)=Eu^i(Yi),θ^i(X^i−Gi(z^))=0almost surely.\mathbb E\hat u_i(\hat X_i)=\mathbb E\hat u_i(Y_i),\qquad \hat\theta_i\bigl(\hat X_i-G_i(\hat z)\bigr)=0\quad\text{almost surely}.Eu^i​(X^i​)=Eu^i​(Yi​),θ^i​(X^i​−Gi​(z^))=0almost surely.

Conversely, an attained Lagrangian maximum satisfying the split constraints and these complementarity equations is a primal optimum. The milestone list follows the source's measure multiplier equations (23)–(24), the measure to utility identity (25), Theorem 1's expected concave subgradient characterization, and the converse's weak duality inequality (manuscript, pp. 5, 8–10).

Significance

The result gives a concrete optimality certificate in a program whose constraints compare entire outcome distributions. Each utility multiplier represents the active part of one dominance constraint. Each θi\theta_iθi​ accounts for the almost sure inequality linking a split outcome to the decision. The equalities show exactly where those constraints are complementary, while the Lagrangian maximum compares the proposed solution with all integrable split outcomes. The paper derives a dual problem from the same Lagrangian in its following section (manuscript, p. 11).

The mathematical theorem is proved in the paper. This mission seeks a machine checked version of its definitions, measure identity, subgradient statement, and both directions of Theorem 2. The published second performance definition is reused as a reference; the nonlinear split program and its utility and measure Lagrangians require a development specific to this paper. The 2003 pure dominance mission contains related local drafts, but those items are not published and cannot currently be imported as platform theorems.

Difficulty

The dominance inequality contains a continuum of thresholds for each outcome. A scalar multiplier at one threshold cannot capture the whole constraint, while the dual object for continuous functions on [ai,bi][a_i,b_i][ai​,bi​] is a measure. The split inequality lives in L1\mathcal L^1L1, where the nonnegative cone has empty interior, so an ordinary interior point argument applied to all constraints at once does not match the paper's setting. The source also needs a subgradient of expected concave utility represented by an almost surely selected, essentially bounded random vector; the conclusion is stronger than merely knowing that the expected objective has a deterministic supporting functional (manuscript, pp. 5–9).

Formalization scope

The Lean development keeps the general separable locally convex Hausdorff decision space, the convex set ZZZ, and a finite index type for the mmm dominance constraints. Operators are function representatives with explicit integrability, continuity in L1\mathcal L^1L1, and samplewise concavity and continuity. Almost sure comparisons use the probability measure PPP; the null set for each realization condition precedes the quantifier over decisions. Split outcomes range only over integrable functions, and utility multipliers range over the exact cone U1([ai,bi])\mathcal U_1([a_i,b_i])U1​([ai​,bi​]). The L∞\mathcal L^\inftyL∞ condition includes almost sure strong measurability and essential boundedness. Maxima in Theorems 1 and 2 are attained maxima, expressed by membership and comparison against every competitor, never a real supremum with a default value.

The source prints a strictly positive affine slope in its definition of U1\mathcal U_1U1​, but immediately calls this class a cone and later uses the zero measure. The formalization uses c≥0c\ge0c≥0; with c>0c>0c>0, Theorem 2 is false for a slack dominance constraint. Uniform dominance is expressed as a positive lower bound rather than a real infimum. The measure milestone uses finite nonnegative measures supported on closed intervals, including endpoint atoms. These conditions exclude default zero integrals, an empty interval disguised by an infimum, and a vacuous utility class. Contributions to the measure to utility correspondence, integration identities, and expected concave subgradient infrastructure can be reused beyond this program.

Selected references

  • D. Dentcheva and A. Ruszczyński, Optimality and duality theory for stochastic optimization problems with nonlinear dominance constraints, Mathematical Programming (2004), DOI; revised author manuscript, April 2003.
  • D. Dentcheva and A. Ruszczyński, Optimization with stochastic dominance constraints, manuscript submitted for publication (2002), cited as reference [6] in the 2003 author manuscript.
  • W. Ogryczak and A. Ruszczyński, Dual stochastic dominance and related mean risk models, SIAM Journal on Optimization 13 (2002), DOI.
7 thms0 active usersReviewed
Operations ResearchOptimizationProbability·Captain: mikedeng1

The Exact Feasibility of Randomized Solutions of Uncertain Convex Programs: Fully-Supported Problems Attain the Binomial Violation Tail ExactlyResearch Paper

Motivation

Many design problems in control, finance and engineering are convex programs whose constraints depend on an uncertain parameter δ\deltaδ: a solution must satisfy x∈Xδx\in\mathcal X_\deltax∈Xδ​ for every δ\deltaδ in a possibly infinite set Δ\DeltaΔ. Enforcing all constraints (robust optimization) is often intractable or overly conservative. The scenario approach draws NNN independent samples of δ\deltaδ, solves the convex program with those NNN constraints only, and asks how likely it is that the resulting solution violates a fresh constraint. The question matters wherever a randomized design is certified by a confidence statement, from robust control to chance-constrained portfolio selection.

Timeline.

  • Calafiore and Campi (Math. Program. 2005; IEEE TAC 2006) introduced the method and bounded the probability that the violation exceeds ε\varepsilonε by a quantity of order (Nd)(1−ε)N−d\binom Nd(1-\varepsilon)^{N-d}(dN​)(1−ε)N−d. The bound is valid but loose.
  • Campi and Garatti (SIAM J. Optim. 2008, this mission's source) proved the bound ∑i=0d−1(Ni)εi(1−ε)N−i\sum_{i=0}^{d-1}\binom Ni\varepsilon^i(1-\varepsilon)^{N-i}∑i=0d−1​(iN​)εi(1−ε)N−i for every convex problem satisfying existence and uniqueness of solutions. They showed it is attained with equality by every fully-supported problem, so it cannot be improved without further assumptions.
  • Later work extended the result to non-unique solutions, constraint removal, and non-convex decisions (Campi and Garatti, Introduction to the Scenario Approach, SIAM 2018).

Setting

Let (Δ,D,P)(\Delta,\mathcal D,\mathbb P)(Δ,D,P) be a probability space, c∈Rdc\in\mathbb R^dc∈Rd with d≥1d\ge1d≥1, and let X⊆Rd\mathcal X\subseteq\mathbb R^dX⊆Rd and Xδ⊆Rd\mathcal X_\delta\subseteq\mathbb R^dXδ​⊆Rd (δ∈Δ\delta\in\Deltaδ∈Δ) be convex closed sets. The violation probability of a point xxx is

V(x)=P{δ∈Δ: x∉Xδ}.V(x)=\mathbb P\{\delta\in\Delta:\ x\notin\mathcal X_\delta\}.V(x)=P{δ∈Δ: x∈/Xδ​}.

For a multi-extraction (δ(1),…,δ(m))∈Δm(\delta^{(1)},\dots,\delta^{(m)})\in\Delta^m(δ(1),…,δ(m))∈Δm, the program PmP_mPm​ minimises c⊤xc^\top xc⊤x over x∈X∩⋂i=1mXδ(i)x\in\mathcal X\cap\bigcap_{i=1}^m\mathcal X_{\delta^{(i)}}x∈X∩⋂i=1m​Xδ(i)​. It is assumed that every PmP_mPm​ has a unique solution xm∗x^*_mxm∗​. A constraint δ(r)\delta^{(r)}δ(r) is a support constraint of PmP_mPm​ if its removal changes the solution. A convex PmP_mPm​ has at most ddd support constraints (Proposition 2.2). The problem is fully-supported if, for every m≥dm\ge dm≥d, the program PmP_mPm​ built from mmm independent samples has exactly ddd support constraints with Pm\mathbb P^mPm-probability one.

Two further objects carry the argument. For I⊆{1,…,m}\mathcal I\subseteq\{1,\dots,m\}I⊆{1,…,m} of cardinality ddd, SIS_{\mathcal I}SI​ is the set of multi-extractions whose support constraints have exactly the indexes in I\mathcal II. The violation law is

F(α)=Pd{V(xd∗)≤α},F(\alpha)=\mathbb P^d\{V(x^*_d)\le\alpha\},F(α)=Pd{V(xd∗​)≤α},

the distribution of the violation of the solution built from ddd samples.

Formalization targets

Goal: Theorem 2.4, equation (2.3)

For a fully-supported problem, every N≥dN\ge dN≥d and every ε∈[0,1]\varepsilon\in[0,1]ε∈[0,1],

PN{V(xN∗)>ε}=∑i=0d−1(Ni)εi(1−ε)N−i.\mathbb P^N\{V(x^*_N)>\varepsilon\}=\sum_{i=0}^{d-1}\binom Ni\varepsilon^i(1-\varepsilon)^{N-i}.PN{V(xN∗​)>ε}=i=0∑d−1​(iN​)εi(1−ε)N−i.

Milestones (PART 1 of §3)

  • Proposition 2.2: at most ddd support constraints.
  • SIˉ⊆S~IˉS_{\bar{\mathcal I}}\subseteq\widetilde S_{\bar{\mathcal I}}SIˉ​⊆SIˉ​ for Iˉ={1,…,d}\bar{\mathcal I}=\{1,\dots,d\}Iˉ={1,…,d}, where S~Iˉ\widetilde S_{\bar{\mathcal I}}SIˉ​ is the set where δ(d+1),…,δ(m)\delta^{(d+1)},\dots,\delta^{(m)}δ(d+1),…,δ(m) are not violated by the solution generated by δ(1),…,δ(d)\delta^{(1)},\dots,\delta^{(d)}δ(1),…,δ(d); and S~Iˉ⊆SIˉ\widetilde S_{\bar{\mathcal I}}\subseteq S_{\bar{\mathcal I}}SIˉ​⊆SIˉ​ up to a probability-zero set.
  • (3.3): Pm{SI}=∫01(1−α)m−dF(dα)\mathbb P^m\{S_{\mathcal I}\}=\int_0^1(1-\alpha)^{m-d}F(\mathrm d\alpha)Pm{SI​}=∫01​(1−α)m−dF(dα) for every I\mathcal II of cardinality ddd.
  • (3.4): (md)∫01(1−α)m−dF(dα)=1\binom md\int_0^1(1-\alpha)^{m-d}F(\mathrm d\alpha)=1(dm​)∫01​(1−α)m−dF(dα)=1 for all m≥dm\ge dm≥d.
  • Moment uniqueness: F(α)=αdF(\alpha)=\alpha^dF(α)=αd is the only distribution on [0,1][0,1][0,1] satisfying (3.4).
  • (3.2): F(α)=αdF(\alpha)=\alpha^dF(α)=αd.
  • Partition chain: PN{V(xN∗)>ε}=(Nd)∫(ε,1](1−α)N−dF(dα)\mathbb P^N\{V(x^*_N)>\varepsilon\}=\binom Nd\int_{(\varepsilon,1]}(1-\alpha)^{N-d}F(\mathrm d\alpha)PN{V(xN∗​)>ε}=(dN​)∫(ε,1]​(1−α)N−dF(dα).
  • Integration by parts: (Nd)∫ε1(1−α)N−d d αd−1 dα=∑i=0d−1(Ni)εi(1−ε)N−i\binom Nd\int_\varepsilon^1(1-\alpha)^{N-d}\,d\,\alpha^{d-1}\,\mathrm d\alpha=\sum_{i=0}^{d-1}\binom Ni\varepsilon^i(1-\varepsilon)^{N-i}(dN​)∫ε1​(1−α)N−ddαd−1dα=∑i=0d−1​(iN​)εi(1−ε)N−i.

Significance

The result. Equation (2.3) shows that the scenario bound (2.2) is tight: no bound that depends only on NNN, ddd and ε\varepsilonε can be smaller, because a fully-supported problem attains it. The distribution of V(xN∗)V(x^*_N)V(xN∗​) is then a Beta law, PN{V(xN∗)≤ε}\mathbb P^N\{V(x^*_N)\le\varepsilon\}PN{V(xN∗​)≤ε} being the probability that a Binomial(N,ε)\mathrm{Binomial}(N,\varepsilon)Binomial(N,ε) variable is at least ddd, the same for every fully-supported problem. This is what fixes the sample sizes used in practice: NNN is chosen so that the binomial tail is below a confidence level β\betaβ. Fact (3.2), that V(xd∗)V(x^*_d)V(xd∗​) has distribution function αd\alpha^dαd whatever the problem, is a distribution-free statement of independent interest.

Formalizing it. The result is proved in the source. As far as is known it has no machine-checked proof. The goal statement is already posed on the platform, and this mission supplies the paper's proof structure as milestones. Two milestones are reusable outside the scenario approach: the uniqueness of a distribution on [0,1][0,1][0,1] given the moments ∫(1−α)k dF=1/(d+kd)\int(1-\alpha)^k\,\mathrm dF=1/\binom{d+k}d∫(1−α)kdF=1/(dd+k​), and the incomplete-beta identity for binomial tails.

Difficulty

The obvious route would compute the law of V(xN∗)V(x^*_N)V(xN∗​) directly, but it depends on the geometry of the constraints. The paper never computes it. It obtains the law of V(xd∗)V(x^*_d)V(xd∗​) only implicitly, through the infinite family of identities (3.4), and recovers it by a uniqueness theorem for moment problems. Two points need care. First, full support holds only almost surely: duplicated samples, for instance, produce programs with fewer than ddd support constraints, so every set identity holds only up to null sets. Second, the claim that removing a non-support constraint keeps the first ddd constraints as the only support constraints uses Proposition 2.2. Two identical non-support constraints show that a constraint can become a support constraint after another is removed, unless the count is bounded by ddd.

Formalization scope

Goal. The goal is the already-posed platform statement ScenarioApproach.Generalization.violation_tail_eq_binomial_sum_of_fullySupported (theorem id cffaa932-832c-42ca-9e81-1848ffab7e34), referenced as it stands and not restated. Proposition 2.2 is the platform statement card_support_constraints_le_dim (f70e8aa3-…). This mission adds the PART 1 steps as milestones under ScenarioExact.PartOne.

Representation. Decisions are vectors in EuclideanSpace ℝ (Fin d). A multi-extraction is ω : Fin m → Δ, with 0-based indexes, so Iˉ\bar{\mathcal I}Iˉ is {i:i<d}\{i : i<d\}{i:i<d} and "δ(d+1),…,δ(m)\delta^{(d+1)},\dots,\delta^{(m)}δ(d+1),…,δ(m)" are the indexes j≥dj\ge dj≥d. Pm\mathbb P^mPm is Measure.pi (fun _ : Fin m => P). VVV, the feasible set, solutions, support constraints and full support are the published definitions violation, feasibleSet, IsSolution, IsSupportConstraint and FullySupported. A support constraint is one whose removal admits a feasible point of strictly smaller cost, which under uniqueness is the paper's "its removal changes the solution". Full support is almost sure, not pointwise.

Hypotheses made explicit. Assumption 1 is entered as existence and uniqueness of the solution for every number of constraints and every sample, together with a family of solution maps θs k, each assumed to solve PkP_kPk​ and to be measurable. Under uniqueness, θs N is the goal's solution map. The paper's "measurability ... is assumed for granted" (p. 4) is replaced by joint measurability of {(x,δ):x∈Xδ}\{(x,\delta):x\in\mathcal X_\delta\}{(x,δ):x∈Xδ​} and measurability of the solution maps, the same two hypotheses as the goal. No set SIS_{\mathcal I}SI​ is assumed measurable. The nonempty-interior clause of Assumption 1 is unused in PART 1 and is not assumed, so the milestones compose with the goal.

Conventions. FFF is the push-forward measure violationLaw on R\mathbb RR, with F(α)F(\alpha)F(α) = violationLaw … (Set.Iic α). Integrals against FFF are lower Lebesgue integrals of nonnegative integrands, as extended nonnegative reals: over [0,1][0,1][0,1] for ∫01\int_0^1∫01​, and over (ε,1](\varepsilon,1](ε,1] for ∫ε1\int_\varepsilon^1∫ε1​ in the partition chain, since that integral comes from the event V>εV>\varepsilonV>ε. The integration-by-parts identity is a real interval integral. Ranges are 1≤d1\le d1≤d, d≤md\le md≤m, d≤Nd\le Nd≤N and 0≤ε≤10\le\varepsilon\le10≤ε≤1.

Ruled out. A pointwise "exactly ddd support constraints for every sample" would be unsatisfiable for many problems (repeated samples) and would trivialise the probabilistic content, so it is not used. Assuming measurability of the event {V(xN∗)>ε}\{V(x^*_N)>\varepsilon\}{V(xN∗​)>ε} or of SIS_{\mathcal I}SI​, or the identity Pm{SI}=Pm{S~I}\mathbb P^m\{S_{\mathcal I}\}=\mathbb P^m\{\widetilde S_{\mathcal I}\}Pm{SI​}=Pm{SI​}, as a hypothesis would assume part of the conclusion, so none of these is a hypothesis.

Infrastructure. A complete development needs: the support-constraint count (Proposition 2.2, a Helly-type argument), invariance of product measures under coordinate permutations, the change-of-variables formula for push-forward measures, the Hausdorff moment uniqueness theorem on [0,1][0,1][0,1], and the binomial–incomplete-beta identity. The last two are general results, and contributions of them are welcome independently.

Selected references

  • M. C. Campi, S. Garatti, The exact feasibility of randomized solutions of uncertain convex programs, SIAM J. Optim. 19(3) (2008) 1211–1230. https://doi.org/10.1137/07069821X
  • G. Calafiore, M. C. Campi, Uncertain convex programs: randomized solutions and confidence levels, Math. Program. 102 (2005) 25–46. https://doi.org/10.1007/s10107-003-0499-y
  • G. Calafiore, M. C. Campi, The scenario approach to robust control design, IEEE Trans. Automat. Control 51(5) (2006) 742–753. https://doi.org/10.1109/TAC.2006.875041
  • M. C. Campi, S. Garatti, Introduction to the Scenario Approach, SIAM, 2018. https://doi.org/10.1137/1.9781611975444
  • A. N. Shiryaev, Probability, 2nd ed., Springer, 1996, Chapter II, §12. https://doi.org/10.1007/978-1-4757-2539-1
14 thms0 active usersReviewed
CombinatoricsOptimizationProbability·Captain: mikedeng1

Convex Optimization: Algorithms and Complexity XVII: Goemans–Williamson Rounding of the MAXCUT SDP Relaxation Has Expected Value at Least 0.878 Times the Maximum CutTextbook

Motivation

MAXCUT asks for a partition of the vertices of a weighted graph into two sets that maximizes the total weight of the edges between them. It is one of Karp's original NP-hard problems, so no polynomial-time exact algorithm is expected, and the natural question is how close a polynomial-time algorithm can come to the optimum. Sampling a uniformly random partition already achieves, in expectation, half of the optimal value. For two decades this factor 1/21/21/2 was essentially the best known.

Goemans and Williamson (J. ACM 42(6), 1995) replaced the combinatorial problem by a semidefinite relaxation, solvable in polynomial time by interior point methods, and rounded its solution with a random Gaussian hyperplane. They proved that the resulting cut has expected weight at least 0.8780.8780.878 times the maximum. The technique founded the use of semidefinite programming in approximation algorithms. Khot, Kindler, Mossel and O'Donnell (SIAM J. Comput. 37(1), 2007) showed that, assuming the Unique Games Conjecture, no polynomial-time algorithm achieves a better constant. Nesterov (Optim. Methods Softw. 9, 1998) extended the rounding analysis to maximizing any positive semidefinite quadratic form over the hypercube, with the constant 2/π2/\pi2/π.

This mission formalizes the presentation of these results in §6.6 of S. Bubeck, Convex Optimization: Algorithms and Complexity (arXiv:1405.4980v2), pp. 343–347.

Setting

Let n≥0n\ge 0n≥0 and let A∈Rn×nA\in\mathbb R^{n\times n}A∈Rn×n be a symmetric matrix with non-negative entries; Ai,jA_{i,j}Ai,j​ is the weight between points iii and jjj. The graph Laplacian is L=D−AL=D-AL=D−A, where DDD is the diagonal matrix with entries ∑j=1nAi,j\sum_{j=1}^n A_{i,j}∑j=1n​Ai,j​. For x∈{−1,1}nx\in\{-1,1\}^nx∈{−1,1}n the vector xxx encodes a partition, and MAXCUT is (6.7)

max⁡x∈{−1,1}nx⊤Lx.\max_{x\in\{-1,1\}^n} x^\top L x .x∈{−1,1}nmax​x⊤Lx.

Write ⟨M,X⟩=Tr⁡(M⊤X)\langle M,X\rangle=\operatorname{Tr}(M^\top X)⟨M,X⟩=Tr(M⊤X) for the Frobenius inner product and S+n\mathbb S^n_+S+n​ for the symmetric positive semidefinite matrices. Since x⊤Lx=⟨L,xx⊤⟩x^\top Lx=\langle L,xx^\top\ranglex⊤Lx=⟨L,xx⊤⟩ and xx⊤∈S+nxx^\top\in\mathbb S^n_+xx⊤∈S+n​ has unit diagonal, MAXCUT is bounded above by the SDP relaxation

max⁡{⟨L,X⟩:X∈S+n, Xi,i=1, i∈[n]}.\max\bigl\{\langle L,X\rangle : X\in\mathbb S^n_+,\ X_{i,i}=1,\ i\in[n]\bigr\}.max{⟨L,X⟩:X∈S+n​, Xi,i​=1, i∈[n]}.

A solution Σ\SigmaΣ of the relaxation is any feasible matrix attaining this maximum. The rounding draws ξ∼N(0,Σ)\xi\sim\mathcal N(0,\Sigma)ξ∼N(0,Σ), a centered Gaussian vector with covariance Σ\SigmaΣ, and outputs ζ=sign⁡(ξ)∈{−1,1}n\zeta=\operatorname{sign}(\xi)\in\{-1,1\}^nζ=sign(ξ)∈{−1,1}n coordinatewise.

Formalization targets

Goal: Theorem 6.11 (Goemans–Williamson)

For AAA symmetric with non-negative entries, L=D−AL=D-AL=D−A, Σ\SigmaΣ any solution of the relaxation, ξ∼N(0,Σ)\xi\sim\mathcal N(0,\Sigma)ξ∼N(0,Σ) and ζ=sign⁡(ξ)\zeta=\operatorname{sign}(\xi)ζ=sign(ξ):

E ζ⊤Lζ ≥ 0.878max⁡x∈{−1,1}nx⊤Lx.\mathbb E\,\zeta^\top L\zeta\ \ge\ 0.878\max_{x\in\{-1,1\}^n}x^\top Lx.Eζ⊤Lζ ≥ 0.878x∈{−1,1}nmax​x⊤Lx.

Milestones

  1. Bounded entries. If Σ∈S+n\Sigma\in\mathbb S^n_+Σ∈S+n​ and Σi,i=1\Sigma_{i,i}=1Σi,i​=1, then ∣Σi,j∣≤1|\Sigma_{i,j}|\le 1∣Σi,j​∣≤1 (remark in the proof of Lemma 6.12).
  2. Lemma 6.12 (Sheppard's formula). If ξ∼N(0,Σ)\xi\sim\mathcal N(0,\Sigma)ξ∼N(0,Σ) with Σi,i=1\Sigma_{i,i}=1Σi,i​=1 and ζ=sign⁡(ξ)\zeta=\operatorname{sign}(\xi)ζ=sign(ξ), then E ζiζj=2πarcsin⁡(Σi,j)\mathbb E\,\zeta_i\zeta_j=\frac{2}{\pi}\arcsin(\Sigma_{i,j})Eζi​ζj​=π2​arcsin(Σi,j​).
  3. Inequality (6.8). 1−2πarcsin⁡(t)≥0.878(1−t)1-\frac{2}{\pi}\arcsin(t)\ge 0.878(1-t)1−π2​arcsin(t)≥0.878(1−t) for all t∈[−1,1]t\in[-1,1]t∈[−1,1].
  4. Relaxation inequality. max⁡xx⊤Lx=max⁡x⟨L,xx⊤⟩≤⟨L,Σ⟩\max_{x}x^\top Lx=\max_x\langle L,xx^\top\rangle\le\langle L,\Sigma\ranglemaxx​x⊤Lx=maxx​⟨L,xx⊤⟩≤⟨L,Σ⟩ for every solution Σ\SigmaΣ.

The separately stated Laplacian identity on p. 346 is also included as a theorem item: if Xi,i=1X_{i,i}=1Xi,i​=1 for all iii, then ⟨L,X⟩=∑i,jAi,j(1−Xi,j)\langle L,X\rangle=\sum_{i,j}A_{i,j}(1-X_{i,j})⟨L,X⟩=∑i,j​Ai,j​(1−Xi,j​); for x∈{−1,1}nx\in\{-1,1\}^nx∈{−1,1}n, x⊤Lx=∑i,jAi,j(1−xixj)x^\top Lx=\sum_{i,j}A_{i,j}(1-x_ix_j)x⊤Lx=∑i,j​Ai,j​(1−xi​xj​).

Companion: Theorem 6.13 (Nesterov)

For B∈S+nB\in\mathbb S^n_+B∈S+n​, Σ\SigmaΣ a solution of max⁡{⟨B,X⟩:X∈S+n, Xi,i=1}\max\{\langle B,X\rangle : X\in\mathbb S^n_+,\ X_{i,i}=1\}max{⟨B,X⟩:X∈S+n​, Xi,i​=1}, ξ∼N(0,Σ)\xi\sim\mathcal N(0,\Sigma)ξ∼N(0,Σ) and ζ=sign⁡(ξ)\zeta=\operatorname{sign}(\xi)ζ=sign(ξ):

E ζ⊤Bζ ≥ 2πmax⁡x∈{−1,1}nx⊤Bx.\mathbb E\,\zeta^\top B\zeta\ \ge\ \frac{2}{\pi}\max_{x\in\{-1,1\}^n}x^\top Bx.Eζ⊤Bζ ≥ π2​x∈{−1,1}nmax​x⊤Bx.

Significance

The result. Theorem 6.11 is a polynomial-time randomized 0.8780.8780.878-approximation for MAXCUT: the relaxation is a semidefinite program, and sampling a Gaussian vector and taking signs is cheap. Repeated sampling turns the bound in expectation into a cut of value close to 0.8780.8780.878 times the optimum with high probability. The same scheme of relaxation followed by randomized rounding underlies approximation algorithms for MAX-2SAT, correlation clustering and quadratic programs over the hypercube, and Nesterov's Theorem 6.13 is the version for an arbitrary positive semidefinite objective.

Formalizing it. Both theorems were proved long ago. To our knowledge neither has a machine-checked proof in Mathlib. The platform has related statements from other books, in different forms: Grothendieck's identity for a standard Gaussian and two unit vectors, and the relaxation guarantee with a Grothendieck constant. This mission states the textbook's results for a Gaussian with a possibly singular covariance matrix, which is the form the rounding uses. A complete development needs Sheppard's formula for a degenerate bivariate Gaussian, an elementary but careful real-variable inequality, and a link between Mathlib's multivariate Gaussian and Gram factorizations of Σ\SigmaΣ. All three are reusable.

Difficulty

The algebra (the Laplacian identity and milestone 4) is routine. The probabilistic core is Lemma 6.12. The textbook argument reduces it to the probability that a uniformly random direction separates two unit vectors, which is "a quick picture" on paper. In Lean this requires showing that the pair (ξi,ξj)(\xi_i,\xi_j)(ξi​,ξj​) has the law of (⟨Vi,ε⟩,⟨Vj,ε⟩)(\langle V_i,\varepsilon\rangle,\langle V_j,\varepsilon\rangle)(⟨Vi​,ε⟩,⟨Vj​,ε⟩) for a standard Gaussian ε\varepsilonε, and then computing an angular measure in the plane, including the degenerate cases Σi,j=±1\Sigma_{i,j}=\pm1Σi,j​=±1, where the pair is supported on a line. A density-based argument fails there, because N(0,Σ)\mathcal N(0,\Sigma)N(0,Σ) has no density when Σ\SigmaΣ is singular, and singular solutions of the relaxation occur (for instance Σ=xx⊤\Sigma=xx^\topΣ=xx⊤). Inequality (6.8) is a statement about a transcendental function on a closed interval with a tight constant (0.8780.8780.878 against the true minimum ≈0.87856\approx0.87856≈0.87856), so crude estimates do not suffice near the minimizer t≈−0.689t\approx-0.689t≈−0.689.

Formalization scope

  • Matrices are Matrix (Fin n) (Fin n) ℝ, vectors Fin n → ℝ. S+n\mathbb S^n_+S+n​ is Matrix.PosSemidef, which includes symmetry, and ⟨M,X⟩\langle M,X\rangle⟨M,X⟩ is trace (Mᵀ * X).
  • N(0,Σ)\mathcal N(0,\Sigma)N(0,Σ) is Mathlib's ProbabilityTheory.multivariateGaussian 0 Σ on EuclideanSpace ℝ (Fin n), defined for every positive semidefinite Σ\SigmaΣ, singular ones included. Expectations are Bochner integrals against it, and each theorem also asserts integrability of its (bounded) integrand.
  • The sign is {−1,1}\{-1,1\}{−1,1}-valued: sign⁡(r)=1\operatorname{sign}(r)=1sign(r)=1 for r≥0r\ge0r≥0 and −1-1−1 for r<0r<0r<0. Mathlib's Real.sign would give sign⁡(0)=0\operatorname{sign}(0)=0sign(0)=0, which takes ζ\zetaζ out of {−1,1}n\{-1,1\}^n{−1,1}n; the two agree almost surely because Σi,i=1\Sigma_{i,i}=1Σi,i​=1.
  • The maximum over the hypercube is a finite maximum (Finset.sup') over the 2n2^n2n Boolean vectors read as ±1\pm1±1 vectors, so it is never a junk value. "The solution" of the relaxation means any maximizer, and maximizers exist since the feasible set is compact and contains the identity.
  • Standing hypotheses: in Theorem 6.11, AAA symmetric with non-negative entries (the book's MAXCUT setting); in Lemma 6.12, Σ\SigmaΣ positive semidefinite (implicit in "ξ∼N(0,Σ)\xi\sim\mathcal N(0,\Sigma)ξ∼N(0,Σ)"); in Theorem 6.13, BBB positive semidefinite. The identities of milestones 4 and 5 hold for every real matrix AAA and are stated without hypotheses on AAA.
  • Ruled out: tying ξ\xiξ's law to anything other than Σ\SigmaΣ, or dropping optimality of Σ\SigmaΣ, would make the goal false or vacuous; here the law is exactly N(0,Σ)\mathcal N(0,\Sigma)N(0,Σ) and Σ\SigmaΣ is a maximizer.
  • Welcome contributions: Sheppard's formula in Mathlib's multivariate Gaussian language, a proof of (6.8), and the Schur product theorem (A,B⪰0⇒A∘B⪰0A,B\succeq0\Rightarrow A\circ B\succeq0A,B⪰0⇒A∘B⪰0) used in Theorem 6.13.

Selected references

  • S. Bubeck, Convex Optimization: Algorithms and Complexity, Foundations and Trends in Machine Learning 8(3–4):231–358, 2015. arXiv:1405.4980v2
  • M. X. Goemans, D. P. Williamson, Improved approximation algorithms for maximum cut and satisfiability problems using semidefinite programming, J. ACM 42(6):1115–1145, 1995. doi:10.1145/227683.227684
  • Yu. Nesterov, Semidefinite relaxation and nonconvex quadratic optimization, Optim. Methods Softw. 9(1–3):141–160, 1998. doi:10.1080/10556789808805690
  • S. Khot, G. Kindler, E. Mossel, R. O'Donnell, Optimal inapproximability results for MAX-CUT and other 2-variable CSPs?, SIAM J. Comput. 37(1):319–357, 2007. doi:10.1137/S0097539705447372
  • W. F. Sheppard, On the application of the theory of error to cases of normal distribution and normal correlation, Phil. Trans. R. Soc. A 192:101–167, 1899. doi:10.1098/rsta.1899.0003
6 thms0 active usersReviewed
Machine LearningOptimizationProbability·Captain: mikedeng1

Convex Optimization: Algorithms and Complexity XVI: Random Coordinate Descent RCD(γ) on a Strongly Convex Coordinate-Smooth Function Has Rate (1 − 1/κ_γ)^tTextbook

Motivation

When a problem has millions of variables, even one full gradient can be too expensive to compute, while a single partial derivative ∂f/∂xi\partial f/\partial x_i∂f/∂xi​ is often cheap: in regularized regression, support vector machines and many structured problems, updating one coordinate costs a small fraction of a full gradient step. Coordinate descent methods exploit this by moving along one coordinate at a time. They are among the oldest optimization schemes and were for a long time analysed only for cyclic orders and only asymptotically.

Nesterov (2012) showed that choosing the coordinate at random, with probabilities depending on the coordinate-wise smoothness constants, gives global, non-asymptotic rates that can beat full gradient descent in total work. This mission formalizes that analysis as presented in §6.4 of S. Bubeck, Convex Optimization: Algorithms and Complexity (arXiv:1405.4980v2), pp. 338–342, and in particular its linear rate for strongly convex functions (Theorem 6.8).

Setting

Let f:Rn→Rf:\mathbb R^n\to\mathbb Rf:Rn→R be differentiable, write ∇if(x)=∂f∂xi(x)\nabla_i f(x)=\frac{\partial f}{\partial x_i}(x)∇i​f(x)=∂xi​∂f​(x) and let eie_iei​ be the iii-th standard basis vector. The function is directionally smooth with constants β1,…,βn>0\beta_1,\dots,\beta_n>0β1​,…,βn​>0 if

∣∇if(x+uei)−∇if(x)∣≤βi∣u∣for all i∈[n], x∈Rn, u∈R,|\nabla_i f(x+ue_i)-\nabla_i f(x)|\le\beta_i|u|\qquad\text{for all } i\in[n],\ x\in\mathbb R^n,\ u\in\mathbb R,∣∇i​f(x+uei​)−∇i​f(x)∣≤βi​∣u∣for all i∈[n], x∈Rn, u∈R,

equivalently, each one-variable restriction u↦f(x+uei)u\mapsto f(x+ue_i)u↦f(x+uei​) is βi\beta_iβi​-smooth.

For a real exponent ccc, the weighted norms are

∥x∥[c]=∑iβicxi2,∥x∥[c]∗=∑iβi−cxi2.\|x\|_{[c]}=\sqrt{\textstyle\sum_{i}\beta_i^{c}x_i^2},\qquad \|x\|^*_{[c]}=\sqrt{\textstyle\sum_{i}\beta_i^{-c}x_i^2}.∥x∥[c]​=∑i​βic​xi2​​,∥x∥[c]∗​=∑i​βi−c​xi2​​.

For α>0\alpha>0α>0, fff is α\alphaα-strongly convex w.r.t. a norm ∥⋅∥\|\cdot\|∥⋅∥ if f(x)−f(y)≤∇f(x)⊤(x−y)−α2∥x−y∥2f(x)-f(y)\le\nabla f(x)^\top(x-y)-\frac{\alpha}{2}\|x-y\|^2f(x)−f(y)≤∇f(x)⊤(x−y)−2α​∥x−y∥2 for all x,yx,yx,y. The point x∗x^*x∗ is a minimizer of fff.

For γ≥0\gamma\ge0γ≥0, RCD(γ\gammaγ) starts at x1∈Rnx_1\in\mathbb R^nx1​∈Rn and iterates

xs+1=xs−1βis∇isf(xs) eis,x_{s+1}=x_s-\frac{1}{\beta_{i_s}}\nabla_{i_s}f(x_s)\,e_{i_s},xs+1​=xs​−βis​​1​∇is​​f(xs​)eis​​,

where i1,i2,…i_1,i_2,\dotsi1​,i2​,… are drawn independently from pγ(i)=βiγ/∑jβjγp_\gamma(i)=\beta_i^\gamma/\sum_{j}\beta_j^\gammapγ​(i)=βiγ​/∑j​βjγ​. The case γ=0\gamma=0γ=0 is uniform sampling; γ=1\gamma=1γ=1 samples proportionally to βi\beta_iβi​.

Formalization targets

Goal: Theorem 6.8 (p. 341)

Let γ≥0\gamma\ge0γ≥0, let fff be α\alphaα-strongly convex w.r.t. ∥⋅∥[1−γ]\|\cdot\|_{[1-\gamma]}∥⋅∥[1−γ]​ and directionally smooth with constants βi\beta_iβi​, and let κγ=∑iβiγ/α\kappa_\gamma=\sum_i\beta_i^\gamma/\alphaκγ​=∑i​βiγ​/α. Then for every t≥0t\ge0t≥0

Ef(xt+1)−f(x∗)≤(1−1κγ)t(f(x1)−f(x∗)).\mathbb E f(x_{t+1})-f(x^*)\le\Big(1-\frac{1}{\kappa_\gamma}\Big)^t\big(f(x_1)-f(x^*)\big).Ef(xt+1​)−f(x∗)≤(1−κγ​1​)t(f(x1​)−f(x∗)).

Milestones

  1. Lemma 6.9 (p. 341): for fff α\alphaα-strongly convex w.r.t. any norm, f(x)−f(x∗)≤12α∥∇f(x)∥∗2f(x)-f(x^*)\le\frac{1}{2\alpha}\|\nabla f(x)\|_*^2f(x)−f(x∗)≤2α1​∥∇f(x)∥∗2​.
  2. One coordinate step (p. 340): f(x−1βi∇if(x)ei)−f(x)≤−12βi(∇if(x))2f\big(x-\frac{1}{\beta_i}\nabla_i f(x)e_i\big)-f(x)\le-\frac{1}{2\beta_i}(\nabla_i f(x))^2f(x−βi​1​∇i​f(x)ei​)−f(x)≤−2βi​1​(∇i​f(x))2.
  3. Expected decrease (p. 340): Eisf(xs+1)−f(xs)≤−12∑iβiγ(∥∇f(xs)∥[1−γ]∗)2\mathbb E_{i_s}f(x_{s+1})-f(x_s)\le-\frac{1}{2\sum_i\beta_i^\gamma}\big(\|\nabla f(x_s)\|^*_{[1-\gamma]}\big)^2Eis​​f(xs+1​)−f(xs​)≤−2∑i​βiγ​1​(∥∇f(xs​)∥[1−γ]∗​)2.
  4. Lemma 6.9 in the weighted norm (p. 342): (∥∇f(x)∥[1−γ]∗)2≥2α(f(x)−f(x∗))\big(\|\nabla f(x)\|^*_{[1-\gamma]}\big)^2\ge2\alpha(f(x)-f(x^*))(∥∇f(x)∥[1−γ]∗​)2≥2α(f(x)−f(x∗)).
  5. Contraction (pp. 341–342): one step multiplies the expected gap by at most 1−1/κγ1-1/\kappa_\gamma1−1/κγ​.

Companion: Theorem 6.7 (pp. 339–340)

For fff convex and directionally smooth, and t≥2t\ge2t≥2,

Ef(xt)−f(x∗)≤2R1−γ2(x1)∑iβiγt−1,R1−γ(x1)=sup⁡f(x)≤f(x1)∥x−x∗∥[1−γ].\mathbb E f(x_t)-f(x^*)\le\frac{2R_{1-\gamma}^2(x_1)\sum_i\beta_i^\gamma}{t-1},\qquad R_{1-\gamma}(x_1)=\sup_{f(x)\le f(x_1)}\|x-x^*\|_{[1-\gamma]}.Ef(xt​)−f(x∗)≤t−12R1−γ2​(x1​)∑i​βiγ​​,R1−γ​(x1​)=f(x)≤f(x1​)sup​∥x−x∗∥[1−γ]​.

Significance

Theorem 6.8 says random coordinate descent converges linearly, with a rate governed by ∑iβiγ/α\sum_i\beta_i^\gamma/\alpha∑i​βiγ​/α instead of the global smoothness constant. For γ=1\gamma=1γ=1, directional smoothness implies fff is β\betaβ-smooth with β≤∑iβi\beta\le\sum_i\beta_iβ≤∑i​βi​, so for functions whose global smoothness constant is of the order of ∑iβi\sum_i\beta_i∑i​βi​, RCD(1) attains the accuracy of gradient descent after the same number of iterations (book, p. 340, comparing Theorem 6.7 with Theorem 3.3), while each iteration touches a single coordinate. The same per-step inequalities underlie later accelerated and parallel coordinate methods.

These results are proved in the literature (Nesterov 2012; Bubeck 2015). Their contribution here is a machine-checked version. As far as a search of the Prove2Me catalogue shows, no coordinate descent rate of this kind has been formalized there; a Euclidean-norm special case of Lemma 6.9 exists on the platform as a separate result, but not the arbitrary-norm lemma or the weighted-norm instance used here.

Difficulty

The main obstacle is bookkeeping of the randomness: the per-step inequality holds for each fixed iterate, while the theorem is about the expectation over the whole sequence of draws i1,…,iti_1,\dots,i_ti1​,…,it​, so the pointwise contraction has to be passed through the tower of conditional expectations. In the strongly convex case this is linear and exact; for Theorem 6.7 the recursion on δs=Ef(xs)−f(x∗)\delta_s=\mathbb Ef(x_s)-f(x^*)δs​=Ef(xs​)−f(x∗) is quadratic, and since δs\delta_sδs​ is an expectation while the gradient norm at xsx_sxs​ is random, the pointwise inequality does not transfer to δs\delta_sδs​ verbatim. A second point is geometric: strong convexity, the dual norm and the sampling distribution must use matching weights (βi1−γ\beta_i^{1-\gamma}βi1−γ​ against βiγ\beta_i^{\gamma}βiγ​), and a mismatch silently changes the constant.

Formalization scope

  • Rn\mathbb R^nRn is EuclideanSpace ℝ (Fin n). The gradient is an explicit map ggg with HasGradientAt f (g x) x, so ∇if(x)=g(x)i\nabla_i f(x)=g(x)_i∇i​f(x)=g(x)i​. Lemma 6.9 is stated for a finite-dimensional real normed space with the Fréchet derivative and the operator norm as the dual norm.
  • Powers βic\beta_i^cβic​ are real powers. The theorems assume n≥1n\ge1n≥1, α>0\alpha>0α>0 and βi>0\beta_i>0βi​>0, which the book uses implicitly; γ≥0\gamma\ge0γ≥0 is the book's.
  • RCD(γ) is a deterministic function of the drawn coordinates, and the expectation over ttt independent draws from pγp_\gammapγ​ is the finite sum ∑(i1,…,it)∈[n]t∏spγ(is) F(i1,…,it)\sum_{(i_1,\dots,i_t)\in[n]^t}\prod_s p_\gamma(i_s)\,F(i_1,\dots,i_t)∑(i1​,…,it​)∈[n]t​∏s​pγ​(is​)F(i1​,…,it​). No measure theory or integrability conventions are involved.
  • The minimizer x∗x^*x∗ is assumed to exist, as the book does throughout; its uniqueness, which the book assumes "only for sake of notation", is not used.
  • In Theorem 6.7 the supremum R1−γ(x1)R_{1-\gamma}(x_1)R1−γ​(x1​) is passed as any real upper bound RRR on the sublevel set, which is equivalent when the supremum is finite and avoids Lean's value 000 for an unbounded supremum.
  • Directional smoothness is required at every xxx and uuu, and pγp_\gammapγ​ is fixed by the βi\beta_iβi​; neither is weakened to hold only along the iterates, which would change the theorem.

A complete development needs the one-dimensional descent lemma (3.5), weighted Cauchy–Schwarz for the dual pair ∥⋅∥[c],∥⋅∥[c]∗\|\cdot\|_{[c]},\|\cdot\|^*_{[c]}∥⋅∥[c]​,∥⋅∥[c]∗​, and a decomposition of the finite expectation over [n]t+1[n]^{t+1}[n]t+1 into the last draw and the first ttt. The weighted-norm and finite-expectation lemmas are reusable for other randomized coordinate and sampling methods. Proofs of the milestones and of either theorem are welcome.

Selected references

  • S. Bubeck, Convex Optimization: Algorithms and Complexity, Foundations and Trends in Machine Learning 8(3–4):231–358, 2015. arXiv:1405.4980v2, §6.4, pp. 338–342.
  • Yu. Nesterov, Efficiency of coordinate descent methods on huge-scale optimization problems, SIAM Journal on Optimization 22(2):341–362, 2012. doi:10.1137/100802001
  • P. Richtárik and M. Takáč, Parallel coordinate descent methods for big data optimization, Mathematical Programming 156:433–484, 2016. arXiv:1212.0873
7 thms0 active usersReviewed
Machine LearningOptimizationProbability·Captain: mikedeng1

Convex Optimization: Algorithms and Complexity XV: SVRG with η = 1/(10β) and k = 20κ Contracts the Expected Optimality Gap by 0.9 per EpochTextbook

Motivation

Many optimization problems in machine learning minimize an average of losses, one loss for each observation. A full gradient step examines every observation, while a stochastic gradient step examines one. The latter is cheaper per step, but its sampled gradient can remain noisy even near the optimum. Section 6.3 of Bubeck's monograph studies stochastic variance reduced gradient descent (SVRG), which periodically computes a full gradient at an anchor point and uses it to correct subsequent sampled gradients. The question for this mission is whether that correction gives a geometric reduction of the expected objective gap at the constants printed in Theorem 6.5.

Bubeck places this method alongside full gradient descent and stochastic gradient descent for finite sums. The section records that earlier stochastic average gradient and dual coordinate ascent methods attain a gradient-computation cost of order (m+κ)log⁡(1/ε)(m+\kappa)\log(1/\varepsilon)(m+κ)log(1/ε) for the same regime, where mmm is the number of components and κ\kappaκ is a condition number. The target here is the precise SVRG convergence statement in the book, rather than a comparison of implementation costs. The source's discussion on pp. 334–336 gives the context and the algorithm.

Setting

Let f1,…,fm:Rn→Rf_1,\ldots,f_m:\mathbb R^n\to\mathbb Rf1​,…,fm​:Rn→R be differentiable convex functions, with m≥1m\ge1m≥1, and define the finite-sum objective and its gradient by

f(x)=1m∑i=1mfi(x),G(x)=1m∑i=1m∇fi(x).f(x)=\frac1m\sum_{i=1}^m f_i(x),\qquad G(x)=\frac1m\sum_{i=1}^m \nabla f_i(x).f(x)=m1​i=1∑m​fi​(x),G(x)=m1​i=1∑m​∇fi​(x).

Each component is β\betaβ-smooth when its gradient is β\betaβ-Lipschitz in the Euclidean norm: ∥∇fi(x)−∇fi(z)∥2≤β∥x−z∥2\|\nabla f_i(x)-\nabla f_i(z)\|_2\le\beta\|x-z\|_2∥∇fi​(x)−∇fi​(z)∥2​≤β∥x−z∥2​ for all x,zx,zx,z. The average fff is α\alphaα-strongly convex, meaning that for all x,zx,zx,z it lies at least α2∥z−x∥22\frac\alpha2\|z-x\|_2^22α​∥z−x∥22​ above its first-order affine approximation at xxx. The constants α\alphaα and β\betaβ are positive, x∗x^*x∗ minimizes fff over Rn\mathbb R^nRn, and κ=β/α\kappa=\beta/\alphaκ=β/α.

An epoch begins at an anchor yyy. Its first inner iterate is x1=yx_1=yx1​=y. For t=1,…,kt=1,\ldots,kt=1,…,k, draw iti_tit​ uniformly from {1,…,m}\{1,\ldots,m\}{1,…,m}, independently across steps and epochs, and update

xt+1=xt−η(∇fit(xt)−∇fit(y)+G(y)).x_{t+1}=x_t-\eta\bigl(\nabla f_{i_t}(x_t)-\nabla f_{i_t}(y)+G(y)\bigr).xt+1​=xt​−η(∇fit​​(xt​)−∇fit​​(y)+G(y)).

The next anchor is the average y+=k−1∑t=1kxty^+=k^{-1}\sum_{t=1}^k x_ty+=k−1∑t=1k​xt​. In particular, this average uses x1x_1x1​ through xkx_kxk​, while the last updated point xk+1x_{k+1}xk+1​ is excluded. Starting from an arbitrary y(1)y^{(1)}y(1) and repeating the epoch produces y(s+1)y^{(s+1)}y(s+1). The expectation of f(y(s+1))f(y^{(s+1)})f(y(s+1)) is over all sksksk sampled indices in the first sss epochs.

Formalization targets

Goal: geometric contraction across epochs

Theorem 6.5 sets η=1/(10β)\eta=1/(10\beta)η=1/(10β) and k=20κk=20\kappak=20κ and asserts, for every s≥1s\ge1s≥1,

Ef(y(s+1))−f(x∗)≤0.9s(f(y(1))−f(x∗)).\mathbb E f(y^{(s+1)})-f(x^*) \le 0.9^s\bigl(f(y^{(1)})-f(x^*)\bigr).Ef(y(s+1))−f(x∗)≤0.9s(f(y(1))−f(x∗)).

The epoch length is a count, so the statement takes k∈Nk\in\mathbb Nk∈N and explicitly requires k=20β/αk=20\beta/\alphak=20β/α. The goal uses exactly the book's step size, epoch length, and contraction factor.

Milestones: second moments and a single epoch

Lemma 6.4 bounds Ei∥∇fi(x)−∇fi(x∗)∥22\mathbb E_i\|\nabla f_i(x)-\nabla f_i(x^*)\|_2^2Ei​∥∇fi​(x)−∇fi​(x∗)∥22​ by 2β(f(x)−f(x∗))2\beta(f(x)-f(x^*))2β(f(x)−f(x∗)). Equation (6.3) bounds the second moment of the corrected sampled direction by the objective gaps at the current point and the anchor. Equation (6.2), the unbiased-direction display, and the one-step display express how that direction changes squared distance to x∗x^*x∗. The later display on p. 338 bounds one epoch for any positive step size with 2βη<12\beta\eta<12βη<1. Finally, equation (6.1) substitutes the stated constants to obtain the factor 0.90.90.9 for one epoch. These seven source claims form the milestone list in reading order.

Significance

The theorem gives an explicit accuracy guarantee after a specified number of epochs: an initial gap DDD falls below 0.9sD0.9^sD0.9sD in expectation. Because each epoch uses a full gradient at its anchor as well as sampled component gradients, the result makes clear which quantity contracts and which operations are counted. It is a concrete linear-rate statement for a method whose individual stochastic gradients need not approach zero at the optimum. Bubeck, §6.3 discusses this issue when introducing the correction term.

The mathematical result is already proved in the monograph. The remaining task is to produce machine-checked proofs of its precise finite-sum model, the single-index estimates, the epoch inequality, and the full repeated-epoch guarantee. The mission drafts those statements and definitions; no proof is claimed for the open theorem items. The finite uniform-average representation and the separation between a conditional one-step average and the full multi-epoch average can be reused in other finite-sum stochastic algorithms.

Difficulty

The sampled component gradient ∇fit(xt)\nabla f_{i_t}(x_t)∇fit​​(xt​) need not be small when xtx_txt​ is near x∗x^*x∗, so a bound using only its norm does not yield the desired fixed-step contraction. The correction −∇fit(y)+G(y)-\nabla f_{i_t}(y)+G(y)−∇fit​​(y)+G(y) has mean zero relative to the full gradient at the current iterate, but its second moment still depends on both xtx_txt​ and yyy. The proof must control those two gaps while respecting the fact that xtx_txt​ depends on earlier samples. A single-index estimate with xtx_txt​ held fixed and an expectation over complete sample histories are different statements; confusing them would make the goal weaker or false.

Formalization scope

The carrier is EuclideanSpace ℝ (Fin n) with its usual inner product and norm. The Fin m components and every sample array are finite. A real-valued uniform average is an ordinary finite sum divided by the number of arrays, and m≥1m\ge1m≥1 and k≥1k\ge1k≥1 prevent an empty average. Independent uniform sampling is represented by averaging over every function from step positions to component indices. The multi-epoch sample space has one such block for every epoch. There are no integrals or measurability side conditions.

The component assumptions include differentiability with an explicit gradient map, convexity on all of Rn\mathbb R^nRn, and the book's gradient-Lipschitz version of smoothness. Strong convexity is imposed on the average objective alone, using the published OnlineConvexOpt.ConvexBasics.StronglyConvexOn definition on the whole space. The book's standing notation assumes a minimizing x∗x^*x∗ exists; this is explicit. Positivity of α\alphaα and β\betaβ, and integrality of 20β/α20\beta/\alpha20β/α, make the displayed divisions and epoch length meaningful. The general epoch bound also requires 0<η0<\eta0<η and 2βη<12\beta\eta<12βη<1. Dimension zero is allowed: the theorem remains a statement about the unique point of R0\mathbb R^0R0 and its zero objective gap.

The direction always contains the sampled difference ∇fit(xt)−∇fit(y)\nabla f_{i_t}(x_t)-\nabla f_{i_t}(y)∇fit​​(xt​)−∇fit​​(y) and the full anchor gradient G(y)G(y)G(y). Replacing that direction with G(xt)G(x_t)G(xt​) would define gradient descent and would not satisfy this mission's algorithm. Contributions are welcome for the finite averaging identities, the component-gradient estimate, the conditional one-step calculation, the epoch inequality, and the induction across epochs.

Selected references

  • Sébastien Bubeck, Convex Optimization: Algorithms and Complexity, Foundations and Trends in Machine Learning 8(3–4), 2015, pp. 231–358. arXiv:1405.4980v2
  • Rie Johnson and Tong Zhang, Accelerating Stochastic Gradient Descent using Predictive Variance Reduction, Advances in Neural Information Processing Systems 26 (NIPS), 2013 (the origin of SVRG, cited by Bubeck on p. 335). https://proceedings.neurips.cc/paper/2013/hash/ac1dd209cbcc5e5d1c6e28598e8cbbe8-Abstract.html
10 thms0 active usersReviewed
Machine LearningOptimizationProbability·Captain: mikedeng1

Convex Optimization: Algorithms and Complexity XIV: Stochastic Mirror Descent on a β-Smooth Function with Noise σ Has Rate Rσ√(2/t) + βR²/tTextbook

Motivation

Many optimization problems in statistics and machine learning ask to minimize an expected loss f(x)=Eξ ℓ(x,ξ)f(x)=\mathbb E_\xi\,\ell(x,\xi)f(x)=Eξ​ℓ(x,ξ), or an average f(x)=1m∑i=1mfi(x)f(x)=\frac1m\sum_{i=1}^m f_i(x)f(x)=m1​∑i=1m​fi​(x) over a large data set. Exact gradients of such an fff are unavailable or too expensive, but unbiased random estimates are cheap: the gradient of the loss at one sample, or of one randomly chosen summand. The observation that first-order methods still make progress when the gradients are only correct on average goes back to Robbins and Monro (1951) and underlies stochastic gradient descent.

Chapter 6 of S. Bubeck, Convex Optimization: Algorithms and Complexity (2015), studies this setting through stochastic mirror descent (S-MD). Its Section 6.1 shows that in the non-smooth case a noisy oracle costs nothing in rate. Section 6.2 asks what smoothness buys: for a general stochastic oracle it cannot buy acceleration, but Theorem 6.3, whose proof the book takes from Dekel, Gilad-Bachrach, Shamir and Xiao (2012), shows that the rate splits into a noise term of order 1/t1/\sqrt t1/t​ and a smoothness term of order 1/t1/t1/t. The book uses it to justify mini-batch SGD. This mission is the fourteenth of a series that formalizes the section capstones of the book.

Setting

Let EEE be a finite-dimensional real vector space with an arbitrary norm ∥⋅∥\|\cdot\|∥⋅∥. Gradients are linear forms ggg on EEE, the value of ggg at vvv is written g⊤vg^\top vg⊤v, and the dual norm is ∥g∥∗=sup⁡∥v∥≤1g⊤v\|g\|_*=\sup_{\|v\|\le1}g^\top v∥g∥∗​=sup∥v∥≤1​g⊤v. Let X⊆E\mathcal X\subseteq EX⊆E be compact and convex.

A mirror map is a function Φ\PhiΦ on an open convex set D\mathcal DD with X⊆D‾\mathcal X\subseteq\overline{\mathcal D}X⊆D and X∩D≠∅\mathcal X\cap\mathcal D\ne\emptysetX∩D=∅. It is strictly convex and differentiable on D\mathcal DD, its gradient ∇Φ\nabla\Phi∇Φ takes every value, and ∥∇Φ(x)∥∗→∞\|\nabla\Phi(x)\|_*\to\infty∥∇Φ(x)∥∗​→∞ as xxx approaches the boundary of D\mathcal DD. Its Bregman divergence is DΦ(x,y)=Φ(x)−Φ(y)−∇Φ(y)⊤(x−y)D_\Phi(x,y)=\Phi(x)-\Phi(y)-\nabla\Phi(y)^\top(x-y)DΦ​(x,y)=Φ(x)−Φ(y)−∇Φ(y)⊤(x−y). The map is 1-strongly convex on X∩D\mathcal X\cap\mathcal DX∩D if DΦ(y,x)≥12∥x−y∥2D_\Phi(y,x)\ge\frac12\|x-y\|^2DΦ​(y,x)≥21​∥x−y∥2 there. A function fff is β\betaβ-smooth on X\mathcal XX if ∥∇f(x)−∇f(y)∥∗≤β∥x−y∥\|\nabla f(x)-\nabla f(y)\|_*\le\beta\|x-y\|∥∇f(x)−∇f(y)∥∗​≤β∥x−y∥ for x,y∈Xx,y\in\mathcal Xx,y∈X.

A stochastic oracle returns, at a query point xxx, a random linear form g~(x)\tilde g(x)g~​(x). When the query point is itself random, the book requires the conditional expectation given the query point, E(g~(x)∣x)\mathbb E(\tilde g(x)\mid x)E(g~​(x)∣x), to be a subgradient of fff at xxx. In the smooth case it requires E(g~(x)∣x)=∇f(x)\mathbb E(\tilde g(x)\mid x)=\nabla f(x)E(g~​(x)∣x)=∇f(x) together with the variance bound E(∥g~(x)−∇f(x)∥∗2∣x)≤σ2\mathbb E(\|\tilde g(x)-\nabla f(x)\|_*^2\mid x)\le\sigma^2E(∥g~​(x)−∇f(x)∥∗2​∣x)≤σ2.

S-MD with step γ\gammaγ starts at x1∈argmin⁡X∩DΦx_1\in\operatorname{argmin}_{\mathcal X\cap\mathcal D}\Phix1​∈argminX∩D​Φ and, writing g~s=g~(xs)\tilde g_s=\tilde g(x_s)g~​s​=g~​(xs​), iterates

xs+1∈argmin⁡x∈X∩D γ g~s⊤x+DΦ(x,xs).x_{s+1}\in\operatorname*{argmin}_{x\in\mathcal X\cap\mathcal D}\ \gamma\,\tilde g_s^\top x+D_\Phi(x,x_s).xs+1​∈x∈X∩Dargmin​ γg~​s⊤​x+DΦ​(x,xs​).

Let R2≥sup⁡x∈X∩DΦ(x)−Φ(x1)R^2\ge\sup_{x\in\mathcal X\cap\mathcal D}\Phi(x)-\Phi(x_1)R2≥supx∈X∩D​Φ(x)−Φ(x1​), and let x∗x^*x∗ minimize fff on X\mathcal XX.

Formalization targets

Goal: Theorem 6.3

Let fff be convex and β\betaβ-smooth, and let the oracle have variance at most σ2\sigma^2σ2. Then for every t≥1t\ge1t≥1, S-MD with step 1/(β+1/η)1/(\beta+1/\eta)1/(β+1/η) and η=Rσ2/t\eta=\frac R\sigma\sqrt{2/t}η=σR​2/t​ satisfies

E f(1t∑s=1txs+1)−f(x∗)≤Rσ2t+βR2t.\mathbb E\,f\Big(\frac1t\sum_{s=1}^t x_{s+1}\Big)-f(x^*)\le R\sigma\sqrt{\frac2t}+\frac{\beta R^2}{t}.Ef(t1​s=1∑t​xs+1​)−f(x∗)≤Rσt2​​+tβR2​.

Milestones (the proof's four displays)

For points xs,xs+1∈X∩Dx_s,x_{s+1}\in\mathcal X\cap\mathcal Dxs​,xs+1​∈X∩D and η>0\eta>0η>0, the smoothness step is

f(xs+1)−f(xs)≤g~s⊤(xs+1−xs)+η2∥∇f(xs)−g~s∥∗2+(β+1/η)DΦ(xs+1,xs).f(x_{s+1})-f(x_s)\le\tilde g_s^\top(x_{s+1}-x_s)+\tfrac\eta2\|\nabla f(x_s)-\tilde g_s\|_*^2+(\beta+1/\eta)D_\Phi(x_{s+1},x_s).f(xs+1​)−f(xs​)≤g~​s⊤​(xs+1​−xs​)+2η​∥∇f(xs​)−g~​s​∥∗2​+(β+1/η)DΦ​(xs+1​,xs​).

If xs+1x_{s+1}xs+1​ is the S-MD step, the mirror step is

1β+1/ηg~s⊤(xs+1−x∗)≤DΦ(x∗,xs)−DΦ(x∗,xs+1)−DΦ(xs+1,xs).\tfrac{1}{\beta+1/\eta}\tilde g_s^\top(x_{s+1}-x^*)\le D_\Phi(x^*,x_s)-D_\Phi(x^*,x_{s+1})-D_\Phi(x_{s+1},x_s).β+1/η1​g~​s⊤​(xs+1​−x∗)≤DΦ​(x∗,xs​)−DΦ​(x∗,xs+1​)−DΦ​(xs+1​,xs​).

Combining the two gives a pathwise bound on f(xs+1)f(x_{s+1})f(xs+1​) with the cross term (g~s−∇f(xs))⊤(x∗−xs)(\tilde g_s-\nabla f(x_s))^\top(x^*-x_s)(g~​s​−∇f(xs​))⊤(x∗−xs​). Taking expectations gives the expected one-step bound

Ef(xs+1)−f(x∗)≤(β+1/η) E(DΦ(x∗,xs)−DΦ(x∗,xs+1))+ησ22.\mathbb Ef(x_{s+1})-f(x^*)\le(\beta+1/\eta)\,\mathbb E\big(D_\Phi(x^*,x_s)-D_\Phi(x^*,x_{s+1})\big)+\frac{\eta\sigma^2}{2}.Ef(xs+1​)−f(x∗)≤(β+1/η)E(DΦ​(x∗,xs​)−DΦ​(x∗,xs+1​))+2ησ2​.

Companion: Theorem 6.1 and (4.10)

For a convex fff with E(∥g~(x)∥∗2∣x)≤B2\mathbb E(\|\tilde g(x)\|_*^2\mid x)\le B^2E(∥g~​(x)∥∗2​∣x)≤B2, S-MD with η=RB2/t\eta=\frac RB\sqrt{2/t}η=BR​2/t​ satisfies

E f(1t∑s=1txs)−min⁡Xf≤RB2/t.\mathbb E\,f\Big(\frac1t\sum_{s=1}^tx_s\Big)-\min_{\mathcal X}f\le RB\sqrt{2/t}.Ef(t1​s=1∑t​xs​)−Xmin​f≤RB2/t​.

This rests on the deterministic regret bound (4.10) of mirror descent along arbitrary vectors gsg_sgs​:

∑s≤tgs⊤(xs−x)≤R2η+η2ρ∑s≤t∥gs∥∗2.\sum_{s\le t}g_s^\top(x_s-x)\le\frac{R^2}{\eta}+\frac{\eta}{2\rho}\sum_{s\le t}\|g_s\|_*^2.s≤t∑​gs⊤​(xs​−x)≤ηR2​+2ρη​s≤t∑​∥gs​∥∗2​.

Significance

Theorem 6.3 says exactly how much smoothness helps under noise. As σ→0\sigma\to0σ→0 it recovers the βR2/t\beta R^2/tβR2/t rate of deterministic smooth optimization. For large ttt the noise term Rσ2/tR\sigma\sqrt{2/t}Rσ2/t​ dominates; the book notes, citing Tsybakov (2003), that smoothness brings no acceleration for a general stochastic oracle. Averaging mmm independent oracle answers divides the variance by mmm, so the theorem quantifies the benefit of mini-batches: the noise term shrinks by m\sqrt mm​ while the smoothness term is unchanged. Theorem 6.1 is the matching non-smooth statement and the template for stochastic subgradient methods in any norm.

These are classical, proved results. None of them is known to be formalized in Lean, and the platform has no stochastic mirror descent statement. Its stochastic gradient items cover the Euclidean strongly convex case and the non-convex gradient-norm case. This mission adds a reusable stochastic-oracle layer in an arbitrary norm, with conditional expectations given random query points, on top of the mirror-map layer of Chapter 4.

Difficulty

The deterministic steps are short manipulations of Bregman divergences. The difficulty is in the passage to expectations. The query point xsx_sxs​ is random, so unbiasedness enters only through the conditional expectation given xsx_sxs​. Making the cross term vanish requires pulling the σ(xs)\sigma(x_s)σ(xs​)-measurable vector x∗−xsx^*-x_sx∗−xs​ out of a conditional expectation of a dual-valued random variable. Every expectation also has to exist. When ∇Φ\nabla\Phi∇Φ blows up at the boundary of D\mathcal DD, the Bregman terms DΦ(x∗,xs)D_\Phi(x^*,x_s)DΦ​(x∗,xs​) are not bounded a priori, and their integrability has to be derived from the recursion. A further obstacle is that the minimizer x∗x^*x∗ may lie on the boundary of D\mathcal DD, where Φ\PhiΦ is not part of the book's data. Treating E\mathbb EE informally, or assuming x∗∈Dx^*\in\mathcal Dx∗∈D, skips exactly these points.

Formalization scope

  • Spaces and gradients. EEE is a finite-dimensional real normed space. Gradients are explicit maps Φ' f' : E → (E →L[ℝ] ℝ), g⊤vg^\top vg⊤v is g v, and ∥⋅∥∗\|\cdot\|_*∥⋅∥∗​ is the operator norm. β\betaβ-smoothness is stated with derivatives relative to X\mathcal XX. Φ\PhiΦ is a total function, constrained only by the mirror-map axioms on D\mathcal DD.
  • Runs and oracle. S-MD is a run predicate. For every outcome, x1x_1x1​ minimizes Φ\PhiΦ on X∩D\mathcal X\cap\mathcal DX∩D, and xs+1x_{s+1}xs+1​ is some minimizer of the step objective. The oracle is a predicate on the random sequences (xs,g~s)(x_s,\tilde g_s)(xs​,g~​s​): each xsx_sxs​ is measurable, and the conditional expectations are taken given σ(xs)\sigma(x_s)σ(xs​). Every conditioned quantity is integrable.
  • Conclusions. Every bound on an expectation also asserts integrability. Without it, the Lean integral of a non-integrable function is 000 and the bound could hold trivially.
  • Standing assumptions. The book's R2=sup⁡(Φ−Φ(x1))R^2=\sup(\Phi-\Phi(x_1))R2=sup(Φ−Φ(x1​)) is replaced by any upper bound R2R^2R2. The minimizer x∗∈Xx^*\in\mathcal Xx∗∈X exists (p. 242). X\mathcal XX is compact and convex (Chapter 4), and convex functions are closed (p. 236).
  • Positivity side conditions. R,σ,B>0R,\sigma,B>0R,σ,B>0 and t≥1t\ge1t≥1 make the step sizes and bounds defined, and β≥0\beta\ge0β≥0.

A variance hypothesis stated only at deterministic points would not control the random iterates, and is not used. Run predicates that let xs+1x_{s+1}xs+1​ be an arbitrary point of X∩D\mathcal X\cap\mathcal DX∩D would make the theorems false, and are not used either.

A complete development needs: first-order optimality over a convex set, the three-point identity of Bregman divergences, the descent lemma in an arbitrary norm, and continuity of the gradient of a differentiable convex function. On the probability side it needs pull-out and conditional Jensen properties for dual-valued conditional expectations. The probability layer is reusable for every stochastic first-order method in the book, including SVRG and random coordinate descent. Proofs of the milestones are welcome, and so are general lemmas about conditional expectations of continuous-linear-map-valued random variables.

Selected references

  • S. Bubeck, Convex Optimization: Algorithms and Complexity, Foundations and Trends in Machine Learning 8(3–4):231–358, 2015. arXiv:1405.4980v2, https://arxiv.org/abs/1405.4980 (Chapter 6, pp. 329–333; Chapter 4, pp. 297–307).
  • O. Dekel, R. Gilad-Bachrach, O. Shamir, L. Xiao, Optimal distributed online prediction using mini-batches, Journal of Machine Learning Research 13:165–202, 2012. https://jmlr.org/papers/v13/dekel12a.html
  • H. Robbins, S. Monro, A stochastic approximation method, Annals of Mathematical Statistics 22(3):400–407, 1951. https://doi.org/10.1214/aoms/1177729586
  • A. Beck, M. Teboulle, Mirror descent and nonlinear projected subgradient methods for convex optimization, Operations Research Letters 31(3):167–175, 2003. https://doi.org/10.1016/S0167-6377(02)00231-6
  • A. Nemirovski, A. Juditsky, G. Lan, A. Shapiro, Robust stochastic approximation approach to stochastic programming, SIAM Journal on Optimization 19(4):1574–1609, 2009. https://doi.org/10.1137/070704277
6 thms0 active usersReviewed
Numerical AnalysisOptimization·Captain: mikedeng1

Convex Optimization: Algorithms and Complexity XIII: Newton's Method Converges Quadratically, ‖x_{k+1} − x*‖ ≤ (M/μ)‖x_k − x*‖², from ‖x₀ − x*‖ ≤ μ/(2M)Textbook

Motivation

Newton's method is the basic second-order method of continuous optimization: at the current point it replaces the objective by its second-order Taylor model and jumps to the stationary point of that model. Its defining property is speed near a nondegenerate minimum, where the error is squared at every step, so that the number of correct digits roughly doubles per iteration. This local behaviour is what makes Newton's method the inner engine of interior point methods, the polynomial-time algorithms for linear, conic and general convex programming (Nesterov and Nemirovski, 1994). In S. Bubeck's monograph Convex Optimization: Algorithms and Complexity (Foundations and Trends in Machine Learning, 2015; arXiv:1405.4980v2), §5.3.2 recalls the traditional local analysis of Newton's method, Theorem 5.3, before turning to the affine-invariant self-concordance analysis used for interior point methods. This mission formalizes that theorem and the four steps of its proof.

Setting

Let Rn\mathbb R^nRn carry the Euclidean norm ∥⋅∥\|\cdot\|∥⋅∥, and write ∥A∥\|A\|∥A∥ for the operator norm of a linear map A:Rn→RnA:\mathbb R^n\to\mathbb R^nA:Rn→Rn, so that ∥Ax∥≤∥A∥ ∥x∥\|Ax\|\le\|A\|\,\|x\|∥Ax∥≤∥A∥∥x∥. Let f:Rn→Rf:\mathbb R^n\to\mathbb Rf:Rn→R be a C2C^2C2 function, with gradient ∇f(x)∈Rn\nabla f(x)\in\mathbb R^n∇f(x)∈Rn and Hessian ∇2f(x)\nabla^2 f(x)∇2f(x), a linear map Rn→Rn\mathbb R^n\to\mathbb R^nRn→Rn (the derivative of the gradient map). For a real number ccc, A⪰cInA\succeq cI_nA⪰cIn​ means ⟨Av,v⟩≥c∥v∥2\langle Av,v\rangle\ge c\|v\|^2⟨Av,v⟩≥c∥v∥2 for all v∈Rnv\in\mathbb R^nv∈Rn.

The Hessian is MMM-Lipschitz if ∥∇2f(x)−∇2f(y)∥≤M∥x−y∥\|\nabla^2 f(x)-\nabla^2 f(y)\|\le M\|x-y\|∥∇2f(x)−∇2f(y)∥≤M∥x−y∥ for all x,y∈Rnx,y\in\mathbb R^nx,y∈Rn.

Newton's method starts at x0∈Rnx_0\in\mathbb R^nx0​∈Rn and iterates, for k≥0k\ge0k≥0,

xk+1=xk−[∇2f(xk)]−1∇f(xk).x_{k+1}=x_k-[\nabla^2 f(x_k)]^{-1}\nabla f(x_k).xk+1​=xk​−[∇2f(xk​)]−1∇f(xk​).

A point x∗x^*x∗ is a local minimum of fff if f(x∗)≤f(x)f(x^*)\le f(x)f(x∗)≤f(x) for all xxx in a neighbourhood of x∗x^*x∗; it has strictly positive Hessian if ∇2f(x∗)⪰μIn\nabla^2 f(x^*)\succeq\mu I_n∇2f(x∗)⪰μIn​ for some μ>0\mu>0μ>0.

Formalization targets

Goal: Theorem 5.3 (p. 320)

Assume the Hessian of fff is MMM-Lipschitz, M>0M>0M>0, and x∗x^*x∗ is a local minimum with ∇2f(x∗)⪰μIn\nabla^2 f(x^*)\succeq\mu I_n∇2f(x∗)⪰μIn​, μ>0\mu>0μ>0. If ∥x0−x∗∥≤μ/(2M)\|x_0-x^*\|\le\mu/(2M)∥x0​−x∗∥≤μ/(2M), then Newton's method from x0x_0x0​ is well defined (every Hessian along the iterates is invertible, so the sequence exists and is unique) and

∥xk+1−x∗∥≤Mμ ∥xk−x∗∥2(k≥0),xk→x∗.\|x_{k+1}-x^*\|\le\frac M\mu\,\|x_k-x^*\|^2\quad(k\ge0),\qquad x_k\to x^*.∥xk+1​−x∗∥≤μM​∥xk​−x∗∥2(k≥0),xk​→x∗.

Milestones (p. 321, the steps of the proof)

  1. The integral formula ∫01∇2f(x+sh) h ds=∇f(x+h)−∇f(x)\int_0^1\nabla^2 f(x+sh)\,h\,ds=\nabla f(x+h)-\nabla f(x)∫01​∇2f(x+sh)hds=∇f(x+h)−∇f(x).
  2. The error representation of one Newton step, xk+1−x∗=[∇2f(xk)]−1∫01[∇2f(xk)−∇2f(x∗+s(xk−x∗))](xk−x∗) dsx_{k+1}-x^*=[\nabla^2 f(x_k)]^{-1}\int_0^1[\nabla^2 f(x_k)-\nabla^2 f(x^*+s(x_k-x^*))](x_k-x^*)\,dsxk+1​−x∗=[∇2f(xk​)]−1∫01​[∇2f(xk​)−∇2f(x∗+s(xk​−x∗))](xk​−x∗)ds.
  3. The Lipschitz bound ∫01∥∇2f(xk)−∇2f(x∗+s(xk−x∗))∥ ds≤M2∥xk−x∗∥\int_0^1\|\nabla^2 f(x_k)-\nabla^2 f(x^*+s(x_k-x^*))\|\,ds\le\frac M2\|x_k-x^*\|∫01​∥∇2f(xk​)−∇2f(x∗+s(xk​−x∗))∥ds≤2M​∥xk​−x∗∥.
  4. The Hessian lower bound ∇2f(xk)⪰(μ−M∥xk−x∗∥)In⪰μ2In\nabla^2 f(x_k)\succeq(\mu-M\|x_k-x^*\|)I_n\succeq\frac\mu2I_n∇2f(xk​)⪰(μ−M∥xk​−x∗∥)In​⪰2μ​In​ when ∥xk−x∗∥≤μ/(2M)\|x_k-x^*\|\le\mu/(2M)∥xk​−x∗∥≤μ/(2M).

Significance

The theorem gives a quantitative basin of quadratic convergence: an explicit radius μ/(2M)\mu/(2M)μ/(2M), depending only on the curvature at the minimum and the Lipschitz constant of the Hessian, inside which Newton's method needs only O(log⁡log⁡(1/ε))O(\log\log(1/\varepsilon))O(loglog(1/ε)) iterations to reach accuracy ε\varepsilonε. It is the classical statement whose shortcomings (dependence on a choice of norm, constants that change under linear changes of variables) motivate the self-concordance theory of the following subsections, and it is the local convergence result invoked whenever a damped or globalized Newton scheme is shown to enter its quadratic phase.

On the formal side, Mathlib has the calculus this needs (Fréchet derivatives, interval integrals of vector-valued maps, operator norms) but no convergence theorem for multivariate Newton's method for minimization. A formal proof produces reusable pieces: the integral form of the mean value theorem for gradients, the stability of a positive-definite lower bound under Lipschitz perturbations, and an inverse-operator norm bound from a quadratic-form lower bound. The result itself is classical and fully proved in the literature; what is open here is its machine-checked proof in this form.

Difficulty

The individual inequalities are short, but the argument is an induction in which well-definedness and the rate are proved together: the Hessian at xkx_kxk​ is invertible only because xkx_kxk​ is still in the ball of radius μ/(2M)\mu/(2M)μ/(2M), and xk+1x_{k+1}xk+1​ stays in that ball only because of the rate. A proof that first assumes the sequence exists and then bounds it is circular. The proof also passes between two kinds of control on the Hessian, a lower bound on its quadratic form and an operator-norm bound on its inverse, and the second is only meaningful once invertibility is established. Finally, the integral manipulations need integrability of the maps s↦∇2f(x∗+s(xk−x∗))(xk−x∗)s\mapsto\nabla^2 f(x^*+s(x_k-x^*))(x_k-x^*)s↦∇2f(x∗+s(xk​−x∗))(xk​−x∗), which comes from the continuity of the Hessian.

Formalization scope

Rn\mathbb R^nRn is EuclideanSpace ℝ (Fin n). The gradient and Hessian are explicit maps g:Rn→Rng:\mathbb R^n\to\mathbb R^ng:Rn→Rn and H:Rn→(Rn→LRn)H:\mathbb R^n\to(\mathbb R^n\to_L\mathbb R^n)H:Rn→(Rn→L​Rn) with ContDiff ℝ 2 f, HasGradientAt f (g x) x and HasFDerivAt g (H x) x at every point; the norm on H(x)H(x)H(x) is Mathlib's operator norm, as on the page. A⪰cInA\succeq cI_nA⪰cIn​ is the quadratic-form inequality. A Newton run is a sequence x:N→Rnx:\mathbb N\to\mathbb R^nx:N→Rn indexed from 000 satisfying the linear system ∇2f(xk)(xk−xk+1)=∇f(xk)\nabla^2 f(x_k)(x_k-x_{k+1})=\nabla f(x_k)∇2f(xk​)(xk​−xk+1​)=∇f(xk​); no inverse of a possibly singular operator appears in any hypothesis, and "well defined" is a conclusion: a unique run exists from x0x_0x0​ and every Hessian along it is bijective. The rate and xk→x∗x_k\to x^*xk​→x∗ are asserted for every run. The error representation is stated with both sides multiplied by ∇2f(xk)\nabla^2 f(x_k)∇2f(xk​), which is equivalent to the printed form once the Hessian is invertible. Milestones 3 and 4 use only the Lipschitz property and are stated for any Lipschitz map HHH.

Added hypothesis: M>0M>0M>0 (the radius μ/(2M)\mu/(2M)μ/(2M) divides by MMM; with M=0M=0M=0, Lean's convention μ/0=0\mu/0=0μ/0=0 would collapse the hypothesis to x0=x∗x_0=x^*x0​=x∗). Convexity of fff is not assumed, as on the page; x∗x^*x∗ is a local minimum and ∇f(x∗)=0\nabla f(x^*)=0∇f(x∗)=0 is derived, not assumed. Encoding the Newton step with Lean's inverse (which returns 000 on singular maps), or replacing ∇2f(x∗)⪰μIn\nabla^2 f(x^*)\succeq\mu I_n∇2f(x∗)⪰μIn​ by mere invertibility, would change the theorem and is ruled out.

A complete development needs the fundamental theorem of calculus for C1C^1C1 vector-valued maps along segments, Hessian-based quadratic-form estimates, and operator-norm bounds for inverses; all are reusable for the analysis of damped Newton, cubic regularization and interior point methods. Proofs of the milestones independently of the goal are welcome.

Selected references

  • S. Bubeck, Convex Optimization: Algorithms and Complexity, Foundations and Trends in Machine Learning 8(3–4):231–358, 2015. arXiv:1405.4980v2, §5.3.2, Theorem 5.3, pp. 320–321.
  • Yu. Nesterov and A. Nemirovski, Interior-Point Polynomial Algorithms in Convex Programming, SIAM Studies in Applied Mathematics 13, 1994. doi:10.1137/1.9781611970791
  • Yu. Nesterov, Introductory Lectures on Convex Optimization: A Basic Course, Kluwer, 2004, Theorem 1.2.5. doi:10.1007/978-1-4419-8853-9
6 thms0 active usersReviewed
Machine LearningOptimization·Captain: mikedeng1

Convex Optimization: Algorithms and Complexity XII: Mirror Prox on a Convex β-Smooth Function Has Rate βR²/(ρt)Textbook

Motivation

First-order methods for constrained convex optimization are usually analysed in the Euclidean norm, where gradient descent and its projected and accelerated variants have well-understood rates. Many constraint sets of practical interest, such as the probability simplex, the ℓ1\ell_1ℓ1​ ball and the spectrahedron, are poorly adapted to the Euclidean geometry: the Euclidean diameter and the Euclidean Lipschitz or smoothness constants grow with the dimension. Mirror descent, due to Nemirovski and Yudin, replaces the Euclidean step by a step taken in the dual space through a mirror map Φ\PhiΦ, so that the rate depends on constants measured in a norm matched to the set. For non-smooth Lipschitz functions it attains the rate 1/t1/\sqrt t1/t​ (Theorem 4.2 of the source, the subject of the previous mission of this series).

For smooth functions, a rate of order 1/t1/t1/t is available in any geometry by an extragradient-type modification. Mirror prox was introduced by Nemirovski in 2004 (Prox-method with rate of convergence O(1/t)) for variational inequalities with Lipschitz monotone operators and convex–concave saddle-point problems. It is the basis of smoothing approaches to non-smooth optimization and of stochastic saddle-point methods. This mission formalizes the version for minimizing a single smooth convex function, Theorem 4.4 of S. Bubeck, Convex Optimization: Algorithms and Complexity (Foundations and Trends in Machine Learning, 2015), §4.5.

Setting

Fix an arbitrary norm ∥⋅∥\|\cdot\|∥⋅∥ on a finite-dimensional real space EEE. Gradients are linear functionals on EEE; for a functional ggg write g⊤vg^\top vg⊤v for its value at vvv, and ∥g∥∗=sup⁡∥v∥≤1g⊤v\|g\|_*=\sup_{\|v\|\le1}g^\top v∥g∥∗​=sup∥v∥≤1​g⊤v for the dual norm. Let X⊆E\mathcal X\subseteq EX⊆E be compact and convex.

  • The Bregman divergence of a differentiable Φ\PhiΦ is DΦ(x,y)=Φ(x)−Φ(y)−∇Φ(y)⊤(x−y)D_\Phi(x,y)=\Phi(x)-\Phi(y)-\nabla\Phi(y)^\top(x-y)DΦ​(x,y)=Φ(x)−Φ(y)−∇Φ(y)⊤(x−y).
  • Let D\mathcal DD be a convex open set with X⊆D‾\mathcal X\subseteq\overline{\mathcal D}X⊆D and X∩D≠∅\mathcal X\cap\mathcal D\ne\emptysetX∩D=∅. A mirror map on D\mathcal DD is a function Φ\PhiΦ that is strictly convex and differentiable on D\mathcal DD, whose gradient takes every value (∇Φ(D)\nabla\Phi(\mathcal D)∇Φ(D) is the whole dual space), and whose gradient norm tends to +∞+\infty+∞ at the boundary of D\mathcal DD.
  • The Bregman projection of y∈Dy\in\mathcal Dy∈D is ΠXΦ(y)=argmin⁡x∈X∩DDΦ(x,y)\Pi^\Phi_{\mathcal X}(y)=\operatorname{argmin}_{x\in\mathcal X\cap\mathcal D}D_\Phi(x,y)ΠXΦ​(y)=argminx∈X∩D​DΦ​(x,y).
  • Φ\PhiΦ is ρ\rhoρ-strongly convex on X∩D\mathcal X\cap\mathcal DX∩D if Φ(x)−Φ(y)≤∇Φ(x)⊤(x−y)−ρ2∥x−y∥2\Phi(x)-\Phi(y)\le\nabla\Phi(x)^\top(x-y)-\frac\rho2\|x-y\|^2Φ(x)−Φ(y)≤∇Φ(x)⊤(x−y)−2ρ​∥x−y∥2 there; fff is β\betaβ-smooth on X\mathcal XX if ∥∇f(x)−∇f(y)∥∗≤β∥x−y∥\|\nabla f(x)-\nabla f(y)\|_*\le\beta\|x-y\|∥∇f(x)−∇f(y)∥∗​≤β∥x−y∥ for x,y∈Xx,y\in\mathcal Xx,y∈X.

Mirror prox with step size η\etaη generates, from xtx_txt​,

∇Φ(yt+1′)=∇Φ(xt)−η∇f(xt),yt+1∈argmin⁡x∈X∩DDΦ(x,yt+1′),\nabla\Phi(y'_{t+1})=\nabla\Phi(x_t)-\eta\nabla f(x_t),\qquad y_{t+1}\in\operatorname*{argmin}_{x\in\mathcal X\cap\mathcal D}D_\Phi(x,y'_{t+1}),∇Φ(yt+1′​)=∇Φ(xt​)−η∇f(xt​),yt+1​∈x∈X∩Dargmin​DΦ​(x,yt+1′​), ∇Φ(xt+1′)=∇Φ(xt)−η∇f(yt+1),xt+1∈argmin⁡x∈X∩DDΦ(x,xt+1′).\nabla\Phi(x'_{t+1})=\nabla\Phi(x_t)-\eta\nabla f(y_{t+1}),\qquad x_{t+1}\in\operatorname*{argmin}_{x\in\mathcal X\cap\mathcal D}D_\Phi(x,x'_{t+1}).∇Φ(xt+1′​)=∇Φ(xt​)−η∇f(yt+1​),xt+1​∈x∈X∩Dargmin​DΦ​(x,xt+1′​).

The first half is a mirror descent step from xtx_txt​ to yt+1y_{t+1}yt+1​; the second restarts from xtx_txt​ with the gradient evaluated at yt+1y_{t+1}yt+1​.

Formalization targets

Goal: Theorem 4.4

With η=ρ/β\eta=\rho/\betaη=ρ/β, x1∈argmin⁡X∩DΦx_1\in\operatorname{argmin}_{\mathcal X\cap\mathcal D}\Phix1​∈argminX∩D​Φ, R2≥sup⁡x∈X∩DΦ(x)−Φ(x1)R^2\ge\sup_{x\in\mathcal X\cap\mathcal D}\Phi(x)-\Phi(x_1)R2≥supx∈X∩D​Φ(x)−Φ(x1​), fff convex and β\betaβ-smooth and x∗x^*x∗ a minimizer of fff on X\mathcal XX, for every t≥1t\ge1t≥1

f(1t∑s=1tys+1)−f(x∗)≤βR2ρt.f\Bigl(\frac1t\sum_{s=1}^t y_{s+1}\Bigr)-f(x^*)\le\frac{\beta R^2}{\rho t}.f(t1​s=1∑t​ys+1​)−f(x∗)≤ρtβR2​.

Milestones

  1. Lemma 4.1 (Bregman projections): (∇Φ(ΠXΦ(y))−∇Φ(y))⊤(ΠXΦ(y)−x)≤0(\nabla\Phi(\Pi^\Phi_{\mathcal X}(y))-\nabla\Phi(y))^\top(\Pi^\Phi_{\mathcal X}(y)-x)\le0(∇Φ(ΠXΦ​(y))−∇Φ(y))⊤(ΠXΦ​(y)−x)≤0 and DΦ(x,ΠXΦ(y))+DΦ(ΠXΦ(y),y)≤DΦ(x,y)D_\Phi(x,\Pi^\Phi_{\mathcal X}(y))+D_\Phi(\Pi^\Phi_{\mathcal X}(y),y)\le D_\Phi(x,y)DΦ​(x,ΠXΦ​(y))+DΦ​(ΠXΦ​(y),y)≤DΦ​(x,y) for x∈X∩Dx\in\mathcal X\cap\mathcal Dx∈X∩D, y∈Dy\in\mathcal Dy∈D.
  2. First term: η∇f(yt+1)⊤(xt+1−x)≤DΦ(x,xt)−DΦ(x,xt+1)−DΦ(xt+1,xt)\eta\nabla f(y_{t+1})^\top(x_{t+1}-x)\le D_\Phi(x,x_t)-D_\Phi(x,x_{t+1})-D_\Phi(x_{t+1},x_t)η∇f(yt+1​)⊤(xt+1​−x)≤DΦ​(x,xt​)−DΦ​(x,xt+1​)−DΦ​(xt+1​,xt​).
  3. Second term, (4.9): η∇f(xt)⊤(yt+1−xt+1)≤DΦ(xt+1,xt)−ρ2∥xt+1−yt+1∥2−ρ2∥yt+1−xt∥2\eta\nabla f(x_t)^\top(y_{t+1}-x_{t+1})\le D_\Phi(x_{t+1},x_t)-\frac\rho2\|x_{t+1}-y_{t+1}\|^2-\frac\rho2\|y_{t+1}-x_t\|^2η∇f(xt​)⊤(yt+1​−xt+1​)≤DΦ​(xt+1​,xt​)−2ρ​∥xt+1​−yt+1​∥2−2ρ​∥yt+1​−xt​∥2.
  4. Third term: (∇f(yt+1)−∇f(xt))⊤(yt+1−xt+1)≤β2∥yt+1−xt∥2+β2∥yt+1−xt+1∥2(\nabla f(y_{t+1})-\nabla f(x_t))^\top(y_{t+1}-x_{t+1})\le\frac\beta2\|y_{t+1}-x_t\|^2+\frac\beta2\|y_{t+1}-x_{t+1}\|^2(∇f(yt+1​)−∇f(xt​))⊤(yt+1​−xt+1​)≤2β​∥yt+1​−xt​∥2+2β​∥yt+1​−xt+1​∥2.
  5. Per-step bound: f(yt+1)−f(x)≤(DΦ(x,xt)−DΦ(x,xt+1))/ηf(y_{t+1})-f(x)\le\bigl(D_\Phi(x,x_t)-D_\Phi(x,x_{t+1})\bigr)/\etaf(yt+1​)−f(x)≤(DΦ​(x,xt​)−DΦ​(x,xt+1​))/η for x∈X∩Dx\in\mathcal X\cap\mathcal Dx∈X∩D.

The three-point identity (4.1), (∇f(x)−∇f(y))⊤(x−z)=Df(x,y)+Df(z,x)−Df(z,y)(\nabla f(x)-\nabla f(y))^\top(x-z)=D_f(x,y)+D_f(z,x)-D_f(z,y)(∇f(x)−∇f(y))⊤(x−z)=Df​(x,y)+Df​(z,x)−Df​(z,y), is already posed on the platform as BeckTeboulleMD.EMDA.lemma_4_1 and is included by reference.

Significance

Theorem 4.4 shows that the 1/t1/t1/t rate of gradient descent on smooth functions survives in non-Euclidean geometries, with the dimension entering only through R2/ρR^2/\rhoR2/ρ. With the negative entropy on the simplex, R2/ρR^2/\rhoR2/ρ is of order log⁡n\log nlogn for the ℓ1\ell_1ℓ1​ norm, against a polynomial dependence for Euclidean methods. Beyond this single-function version, the same three-term argument is the core of the analysis of mirror prox for monotone variational inequalities and saddle-point problems (§4.6 and §5.2 of the source), where it gives the 1/t1/t1/t rate for smooth convex–concave games.

The result is proved in the source and in Nemirovski's paper. As far as this mission is aware it has no machine-checked proof; Mathlib has no Bregman divergences, mirror maps or mirror-descent-type algorithms. A formal proof would supply reusable infrastructure: the Bregman projection lemma (Lemma 4.1) and the three-point identity are used by every mirror-descent analysis, including the stochastic and online variants later in the source.

Difficulty

The obvious argument for mirror descent bounds f(xt)−f(x)f(x_t)-f(x)f(xt​)−f(x) by ∇f(xt)⊤(xt−x)\nabla f(x_t)^\top(x_t-x)∇f(xt​)⊤(xt​−x) and leaves a stability term of size η2∥∇f∥∗2\eta^2\|\nabla f\|_*^2η2∥∇f∥∗2​, which yields only 1/t1/\sqrt t1/t​. Smoothness has to be used to cancel that term, and a single gradient step does not do so in a general norm. Mirror prox evaluates the gradient at the extrapolated point yt+1y_{t+1}yt+1​; the analysis must show that the error created by using ∇f(xt)\nabla f(x_t)∇f(xt​) for the first step is paid for by the strong convexity of Φ\PhiΦ, which requires splitting ∇f(yt+1)⊤(yt+1−x)\nabla f(y_{t+1})^\top(y_{t+1}-x)∇f(yt+1​)⊤(yt+1​−x) into three terms and bounding each with matching constants.

In Lean the difficulty is in the infrastructure: first-order optimality of a Bregman projection over X∩D\mathcal X\cap\mathcal DX∩D, which need not be closed; the passage from x∈X∩Dx\in\mathcal X\cap\mathcal Dx∈X∩D to a minimizer x∗x^*x∗ that may lie on the boundary of D\mathcal DD; and Jensen's inequality for the average of the ys+1y_{s+1}ys+1​.

Formalization scope

  • EEE is a finite-dimensional real normed space. Gradients are continuous linear functionals E →L[ℝ] ℝ given by explicit maps Φ' and f'; the dual norm is the operator norm. Φ' is the Fréchet derivative of Φ\PhiΦ at each point of D\mathcal DD; f' is the derivative of fff relative to X\mathcal XX at each point of X\mathcal XX.
  • The mirror map definition keeps all three properties of §4.1, including surjectivity of the gradient and divergence at each frontier point of D\mathcal DD.
  • Projections and runs are relations: every admissible argmin choice is covered; no choice function is used. The run is indexed from t=1t=1t=1; index 000 is unused.
  • Standing and implicit hypotheses: X\mathcal XX compact convex; a minimizer x∗∈Xx^*\in\mathcal Xx∗∈X exists (the source's standing assumption); ρ>0\rho>0ρ>0 and β>0\beta>0β>0 (so that η=ρ/β\eta=\rho/\betaη=ρ/β and the division by ρt\rho tρt are meaningful); x1∈argmin⁡X∩DΦx_1\in\operatorname{argmin}_{\mathcal X\cap\mathcal D}\Phix1​∈argminX∩D​Φ (unspecified in §4.5, taken as in mirror descent, §4.2).
  • RRR is any real with Φ−Φ(x1)≤R2\Phi-\Phi(x_1)\le R^2Φ−Φ(x1​)≤R2 on X∩D\mathcal X\cap\mathcal DX∩D. The source's R2=sup⁡R^2=\supR2=sup is the case where the supremum is finite; when it is infinite, the source's bound is void.
  • A run in which the second step started from yt+1y_{t+1}yt+1​, or projected yt+1′y'_{t+1}yt+1′​ again, would be mirror descent, for which the claimed rate is false; the formal run follows the four equations of the source exactly. The average is over the extrapolated points ys+1y_{s+1}ys+1​, s=1,…,ts=1,\dots,ts=1,…,t.

Contributions welcome: proofs of Lemma 4.1 and the three-term bounds, which are reusable for the stochastic mirror descent and saddle-point missions of this series, and the final telescoping and limit argument.

Selected references

  • S. Bubeck, Convex Optimization: Algorithms and Complexity, Foundations and Trends in Machine Learning 8(3–4):231–358, 2015. https://arxiv.org/abs/1405.4980v2
  • A. Nemirovski, Prox-method with rate of convergence O(1/t) for variational inequalities with Lipschitz continuous monotone operators and smooth convex-concave saddle point problems, SIAM Journal on Optimization 15(1):229–251, 2004. https://doi.org/10.1137/S1052623403425629
  • A. Beck and M. Teboulle, Mirror descent and nonlinear projected subgradient methods for convex optimization, Operations Research Letters 31(3):167–175, 2003. https://doi.org/10.1016/S0167-6377(02)00231-6
  • A. Nemirovski and D. Yudin, Problem Complexity and Method Efficiency in Optimization, Wiley, 1983.
8 thms0 active usersReviewed
Machine LearningOptimization·Captain: mikedeng1

Convex Optimization: Algorithms and Complexity XI: Mirror Descent with a ρ-Strongly Convex Mirror Map Has Rate RL√(2/(ρt))Textbook

Motivation

The projected subgradient method of Chapter 3 measures distances in the Euclidean norm, and its rate RL/tRL/\sqrt tRL/t​ depends on the geometry only through a Euclidean radius RRR and a Euclidean bound LLL on the subgradients. For many constraint sets that occur in practice this is the wrong geometry. On the probability simplex Δn\Delta_nΔn​ with subgradients bounded in ℓ∞\ell_\inftyℓ∞​, as for linear losses with bounded coefficients, the Euclidean constants make RLRLRL grow like n\sqrt nn​, while the problem itself is nearly dimension-free.

Mirror descent, introduced by Nemirovski and Yudin (1983), repairs this. The gradient step is taken in the dual space, after mapping the current point through the gradient of a strictly convex mirror map Φ\PhiΦ, and feasibility is restored by a projection in the Bregman divergence of Φ\PhiΦ rather than in the Euclidean distance. Beck and Teboulle (2003, doi:10.1016/S0167-6377(02)00231-6) recast it as a nonlinear projected subgradient method and showed that with the negative entropy on the simplex its rate is O(Llog⁡n/t)O(L\sqrt{\log n/t})O(Llogn/t​). Nesterov (2009, doi:10.1007/s10107-007-0149-x) introduced dual averaging, a lazy variant that averages the subgradients in the dual space and maps back only when a primal point is needed. Both methods underlie online learning (exponential weights is mirror descent with the entropy), stochastic optimization, and saddle-point methods such as mirror prox.

This mission is the eleventh of a series formalizing S. Bubeck, Convex Optimization: Algorithms and Complexity (Foundations and Trends in Machine Learning, 2015, arXiv:1405.4980). It covers the preamble of Chapter 4, Sections 4.1 and 4.2, and Section 4.4.

Setting

Let EEE be a finite-dimensional real vector space (the book's Rn\mathbb R^nRn) with an arbitrary norm ∥⋅∥\|\cdot\|∥⋅∥. A linear functional ggg on EEE acts on xxx by g⊤xg^\top xg⊤x and has dual norm ∥g∥∗=sup⁡∥x∥≤1g⊤x\|g\|_*=\sup_{\|x\|\le1}g^\top x∥g∥∗​=sup∥x∥≤1​g⊤x. Let X⊆E\mathcal X\subseteq EX⊆E be compact and convex.

Let D⊆E\mathcal D\subseteq ED⊆E be a convex open set with X⊆D‾\mathcal X\subseteq\overline{\mathcal D}X⊆D and X∩D≠∅\mathcal X\cap\mathcal D\neq\emptysetX∩D=∅. A function Φ:D→R\Phi:\mathcal D\to\mathbb RΦ:D→R is a mirror map if it is strictly convex and differentiable, its gradient takes every value in the dual space, and ∥∇Φ(x)∥→+∞\|\nabla\Phi(x)\|\to+\infty∥∇Φ(x)∥→+∞ as xxx tends to the boundary of D\mathcal DD. Its Bregman divergence is

DΦ(x,y)=Φ(x)−Φ(y)−∇Φ(y)⊤(x−y),D_\Phi(x,y)=\Phi(x)-\Phi(y)-\nabla\Phi(y)^\top(x-y),DΦ​(x,y)=Φ(x)−Φ(y)−∇Φ(y)⊤(x−y),

and the Bregman projection of y∈Dy\in\mathcal Dy∈D is ΠXΦ(y)=argmin⁡x∈X∩DDΦ(x,y)\Pi^\Phi_{\mathcal X}(y)=\operatorname{argmin}_{x\in\mathcal X\cap\mathcal D}D_\Phi(x,y)ΠXΦ​(y)=argminx∈X∩D​DΦ​(x,y). The mirror map is ρ\rhoρ-strongly convex on X∩D\mathcal X\cap\mathcal DX∩D if Φ(x)−Φ(y)≤∇Φ(x)⊤(x−y)−ρ2∥x−y∥2\Phi(x)-\Phi(y)\le\nabla\Phi(x)^\top(x-y)-\frac\rho2\|x-y\|^2Φ(x)−Φ(y)≤∇Φ(x)⊤(x−y)−2ρ​∥x−y∥2 for all x,y∈X∩Dx,y\in\mathcal X\cap\mathcal Dx,y∈X∩D.

Let fff be convex on X\mathcal XX with a minimizer x∗∈Xx^*\in\mathcal Xx∗∈X. A linear functional ggg is a subgradient of fff at xxx if f(x)−f(y)≤g⊤(x−y)f(x)-f(y)\le g^\top(x-y)f(x)−f(y)≤g⊤(x−y) for all y∈Xy\in\mathcal Xy∈X, and fff is LLL-Lipschitz if the subgradients satisfy ∥g∥∗≤L\|g\|_*\le L∥g∥∗​≤L.

Mirror descent with step η\etaη starts at x1∈argmin⁡X∩DΦx_1\in\operatorname{argmin}_{\mathcal X\cap\mathcal D}\Phix1​∈argminX∩D​Φ and, for t≥1t\ge1t≥1, picks gt∈∂f(xt)g_t\in\partial f(x_t)gt​∈∂f(xt​) and

∇Φ(yt+1)=∇Φ(xt)−ηgt,yt+1∈D,xt+1=ΠXΦ(yt+1).(4.2–4.3)\nabla\Phi(y_{t+1})=\nabla\Phi(x_t)-\eta g_t,\quad y_{t+1}\in\mathcal D,\qquad x_{t+1}=\Pi^\Phi_{\mathcal X}(y_{t+1}).\tag{4.2–4.3}∇Φ(yt+1​)=∇Φ(xt​)−ηgt​,yt+1​∈D,xt+1​=ΠXΦ​(yt+1​).(4.2–4.3)

Dual averaging instead sets xt∈argmin⁡x∈X∩Dη∑s=1t−1gs⊤x+Φ(x)x_t\in\operatorname{argmin}_{x\in\mathcal X\cap\mathcal D}\eta\sum_{s=1}^{t-1}g_s^\top x+\Phi(x)xt​∈argminx∈X∩D​η∑s=1t−1​gs⊤​x+Φ(x) (4.6). Let R2R^2R2 bound Φ(x)−Φ(x1)\Phi(x)-\Phi(x_1)Φ(x)−Φ(x1​) over X∩D\mathcal X\cap\mathcal DX∩D.

Formalization targets

Goal: Theorem 4.2

With η=RL2ρt\eta=\frac RL\sqrt{\frac{2\rho}{t}}η=LR​t2ρ​​, every run of mirror descent satisfies

f(1t∑s=1txs)−f(x∗)≤RL2ρt.f\Big(\frac1t\sum_{s=1}^t x_s\Big)-f(x^*)\le RL\sqrt{\frac{2}{\rho t}}.f(t1​s=1∑t​xs​)−f(x∗)≤RLρt2​​.

Milestones

  1. The three-point identity (4.1): (∇f(x)−∇f(y))⊤(x−z)=Df(x,y)+Df(z,x)−Df(z,y)(\nabla f(x)-\nabla f(y))^\top(x-z)=D_f(x,y)+D_f(z,x)-D_f(z,y)(∇f(x)−∇f(y))⊤(x−z)=Df​(x,y)+Df​(z,x)−Df​(z,y).
  2. Lemma 4.1: for x∈X∩Dx\in\mathcal X\cap\mathcal Dx∈X∩D and y∈Dy\in\mathcal Dy∈D, DΦ(x,ΠXΦ(y))+DΦ(ΠXΦ(y),y)≤DΦ(x,y)D_\Phi(x,\Pi^\Phi_{\mathcal X}(y))+D_\Phi(\Pi^\Phi_{\mathcal X}(y),y)\le D_\Phi(x,y)DΦ​(x,ΠXΦ​(y))+DΦ​(ΠXΦ​(y),y)≤DΦ​(x,y), together with the first-order inequality that implies it.
  3. The per-step inequality: f(xs)−f(x)≤1η(DΦ(x,xs)+DΦ(xs,ys+1)−DΦ(x,xs+1)−DΦ(xs+1,ys+1))f(x_s)-f(x)\le\frac1\eta\big(D_\Phi(x,x_s)+D_\Phi(x_s,y_{s+1})-D_\Phi(x,x_{s+1})-D_\Phi(x_{s+1},y_{s+1})\big)f(xs​)−f(x)≤η1​(DΦ​(x,xs​)+DΦ​(xs​,ys+1​)−DΦ​(x,xs+1​)−DΦ​(xs+1​,ys+1​)).
  4. The stability term: DΦ(xs,ys+1)−DΦ(xs+1,ys+1)≤(ηL)2/(2ρ)D_\Phi(x_s,y_{s+1})-D_\Phi(x_{s+1},y_{s+1})\le(\eta L)^2/(2\rho)DΦ​(xs​,ys+1​)−DΦ​(xs+1​,ys+1​)≤(ηL)2/(2ρ).
  5. The summed bound: ∑s=1t(f(xs)−f(x))≤DΦ(x,x1)/η+ηL2t/(2ρ)\sum_{s=1}^t(f(x_s)-f(x))\le D_\Phi(x,x_1)/\eta+\eta L^2t/(2\rho)∑s=1t​(f(xs​)−f(x))≤DΦ​(x,x1​)/η+ηL2t/(2ρ).

Companion: Theorem 4.3

With η=RLρ2t\eta=\frac RL\sqrt{\frac{\rho}{2t}}η=LR​2tρ​​, dual averaging satisfies f(1t∑s=1txs)−f(x∗)≤2RL2ρtf\big(\frac1t\sum_{s=1}^tx_s\big)-f(x^*)\le2RL\sqrt{\frac2{\rho t}}f(t1​∑s=1t​xs​)−f(x∗)≤2RLρt2​​, with the two displays (4.7) and (4.8) of its proof as supporting items.

Significance

Theorem 4.2 is the basic rate of non-Euclidean first-order optimization. With the Euclidean mirror map Φ=12∥⋅∥22\Phi=\frac12\|\cdot\|_2^2Φ=21​∥⋅∥22​ it recovers the projected subgradient rate of Chapter 3; with the negative entropy on the simplex, which is 111-strongly convex for ℓ1\ell_1ℓ1​ by Pinsker's inequality and has R2=log⁡nR^2=\log nR2=logn, it gives L∞2log⁡n/tL_\infty\sqrt{2\log n/t}L∞​2logn/t​, so the dependence on the dimension becomes logarithmic. The same analysis, with fff replaced by a sequence of losses, is the regret bound of online mirror descent, and its per-step and stability inequalities are reused for stochastic mirror descent and mirror prox in later chapters of the book.

The results are classical and fully proved in the book. What this mission adds is a machine-checked version in a general finite-dimensional normed space, with gradients as linear functionals and the dual norm as the operator norm, for an arbitrary mirror map on an arbitrary open domain. On Prove2Me the three-point identity (4.1) is already posed, as Lemma 4.1 of the Beck–Teboulle formalization, and is reused here; a regret bound for online mirror descent with Legendre functions is proved in another library, but it is not this theorem and does not cover a domain D\mathcal DD distinct from the whole space or a Bregman projection onto X∩D\mathcal X\cap\mathcal DX∩D.

Difficulty

The algebra of the proof is short. The difficulty lies at the boundary of D\mathcal DD. The minimizer x∗x^*x∗ may lie on ∂D\partial\mathcal D∂D, where Φ\PhiΦ and ∇Φ\nabla\Phi∇Φ are not defined; this is the typical case for the entropy on the simplex, where x∗x^*x∗ is often a vertex. The per-step bound holds only for comparison points x∈X∩Dx\in\mathcal X\cap\mathcal Dx∈X∩D, so the bound at x∗x^*x∗ has to be recovered from points of X∩D\mathcal X\cap\mathcal DX∩D, where no regularity of fff beyond convexity on X\mathcal XX is available. Lemma 4.1 needs the first-order optimality condition of the Bregman projection over the convex set X∩D\mathcal X\cap\mathcal DX∩D, which is not closed. In the non-Euclidean setting the gradient step cannot be written in the primal space at all: the update (4.2) is an equation between linear functionals, and the norm and dual norm must be kept apart throughout.

Formalization scope

  • The space is a finite-dimensional real normed space E with an arbitrary norm; gradients and subgradients are elements of E →L[ℝ] ℝ, the dual norm is the operator norm, and ∇Φ\nabla\Phi∇Φ is an explicit map Φ' with HasFDerivAt Φ (Φ' x) x on D\mathcal DD.
  • The mirror map carries all three properties (i)–(iii) of the book; the standing setting (X\mathcal XX compact convex, X⊆D‾\mathcal X\subseteq\overline{\mathcal D}X⊆D, X∩D≠∅\mathcal X\cap\mathcal D\neq\emptysetX∩D=∅) is a single predicate.
  • Runs are predicates over the first ttt steps, quantified universally: any subgradient, any yt+1y_{t+1}yt+1​ solving (4.2), any minimizer. A run whose next iterate were an arbitrary point of X∩D\mathcal X\cap\mathcal DX∩D, rather than the Bregman projection, would make the theorem false; the projection is part of the run.
  • Sequences are indexed from 111; t≥1t\ge1t≥1, ρ>0\rho>0ρ>0, L>0L>0L>0, R>0R>0R>0 and η>0\eta>0η>0 are explicit, since the step and the bound divide by them.
  • R2R^2R2 is any upper bound of sup⁡X∩DΦ−Φ(x1)\sup_{\mathcal X\cap\mathcal D}\Phi-\Phi(x_1)supX∩D​Φ−Φ(x1​) (the book's RRR is the least one); when the supremum is infinite the hypotheses cannot be met, as on the page.
  • "fff is LLL-Lipschitz" is assumed for the subgradients the run uses. Subgradients relative to X\mathcal XX are unbounded at boundary points of X\mathcal XX, so the literal "every subgradient at every point" would make the hypothesis unsatisfiable.
  • The definitions (Bregman divergence, mirror map, Bregman projection, the two runs) are reusable by the later chapters on mirror prox and stochastic mirror descent. Proofs of any milestone, and of the general facts they need (nonnegativity of Bregman divergences, first-order optimality over a convex set that is not closed), are welcome.

Selected references

  • S. Bubeck, Convex Optimization: Algorithms and Complexity, Foundations and Trends in Machine Learning 8(3–4):231–358, 2015. arXiv:1405.4980
  • A. Nemirovski and D. Yudin, Problem Complexity and Method Efficiency in Optimization, Wiley, 1983 (book; no DOI).
  • A. Beck and M. Teboulle, Mirror descent and nonlinear projected subgradient methods for convex optimization, Operations Research Letters 31(3):167–175, 2003. doi:10.1016/S0167-6377(02)00231-6
  • Y. Nesterov, Primal-dual subgradient methods for convex problems, Mathematical Programming 120(1):221–259, 2009. doi:10.1007/s10107-007-0149-x
7 thms0 active usersReviewed
Machine LearningOptimization·Captain: mikedeng1

Convex Optimization: Algorithms and Complexity X: Nesterov's Accelerated Gradient Descent on a β-Smooth α-Strongly Convex Function Has Rate ((α + β)/2)‖x₁ − x*‖² exp(−(t − 1)/√κ)Textbook

Why accelerated rates matter

First-order methods, which query only function values and gradients, are the workhorse of large-scale optimization in machine learning, signal processing and operations research, because each step costs little more than one gradient evaluation. For a function that is both strongly convex and smooth, plain gradient descent converges geometrically, but the number of steps needed to reach accuracy ε\varepsilonε scales with the condition number κ\kappaκ of the problem. In 1983 Nesterov showed that a gradient method with a carefully chosen momentum term needs a number of steps proportional to κ\sqrt\kappaκ​ instead, and that this is optimal for black-box first-order methods. On ill-conditioned problems, where κ\kappaκ is in the thousands or millions, the difference between κ\kappaκ and κ\sqrt\kappaκ​ is the difference between practical and impractical.

This mission is the tenth of a series formalizing S. Bubeck's monograph Convex Optimization: Algorithms and Complexity (Foundations and Trends in Machine Learning, 2015; arXiv:1405.4980v2). It covers §3.7.1, the smooth and strongly convex case of Nesterov's accelerated gradient descent, and its main result, Theorem 3.18.

Setting

Work in Rn\mathbb R^nRn with the Euclidean inner product x⊤yx^\top yx⊤y and norm ∥⋅∥\|\cdot\|∥⋅∥. Let f:Rn→Rf:\mathbb R^n\to\mathbb Rf:Rn→R be differentiable with gradient ∇f\nabla f∇f.

  • fff is β\betaβ-smooth if its gradient is β\betaβ-Lipschitz: ∥∇f(x)−∇f(y)∥≤β∥x−y∥\|\nabla f(x)-\nabla f(y)\|\le\beta\|x-y\|∥∇f(x)−∇f(y)∥≤β∥x−y∥ for all x,yx,yx,y.
  • fff is α\alphaα-strongly convex (α>0\alpha>0α>0) if for all x,yx,yx,y
f(y)≥f(x)+∇f(x)⊤(y−x)+α2∥y−x∥2.f(y)\ge f(x)+\nabla f(x)^\top(y-x)+\frac\alpha2\|y-x\|^2 .f(y)≥f(x)+∇f(x)⊤(y−x)+2α​∥y−x∥2.
  • The condition number is κ=β/α\kappa=\beta/\alphaκ=β/α; for n≥1n\ge1n≥1 one always has κ≥1\kappa\ge1κ≥1.
  • x∗x^*x∗ denotes a minimizer of fff on Rn\mathbb R^nRn.

Nesterov's accelerated gradient descent starts at an arbitrary point x1=y1x_1=y_1x1​=y1​ and iterates, for t≥1t\ge1t≥1,

yt+1=xt−1β∇f(xt),xt+1=(1+κ−1κ+1)yt+1−κ−1κ+1 yt.y_{t+1}=x_t-\frac1\beta\nabla f(x_t),\qquad x_{t+1}=\Big(1+\frac{\sqrt\kappa-1}{\sqrt\kappa+1}\Big)y_{t+1}-\frac{\sqrt\kappa-1}{\sqrt\kappa+1}\,y_t .yt+1​=xt​−β1​∇f(xt​),xt+1​=(1+κ​+1κ​−1​)yt+1​−κ​+1κ​−1​yt​.

The point yt+1y_{t+1}yt+1​ is a gradient step from xtx_txt​, and xt+1x_{t+1}xt+1​ moves beyond yt+1y_{t+1}yt+1​ in the direction yt+1−yty_{t+1}-y_tyt+1​−yt​ by the fixed momentum factor (κ−1)/(κ+1)(\sqrt\kappa-1)/(\sqrt\kappa+1)(κ​−1)/(κ​+1).

The analysis in the book uses auxiliary quadratic functions Φs\Phi_sΦs​ (an estimate sequence), defined from the points xsx_sxs​ by

Φ1(x)=f(x1)+α2∥x−x1∥2,Φs+1(x)=(1−1κ)Φs(x)+1κ(f(xs)+∇f(xs)⊤(x−xs)+α2∥x−xs∥2),\Phi_1(x)=f(x_1)+\frac\alpha2\|x-x_1\|^2,\qquad \Phi_{s+1}(x)=\Big(1-\frac1{\sqrt\kappa}\Big)\Phi_s(x)+\frac1{\sqrt\kappa}\Big(f(x_s)+\nabla f(x_s)^\top(x-x_s)+\frac\alpha2\|x-x_s\|^2\Big),Φ1​(x)=f(x1​)+2α​∥x−x1​∥2,Φs+1​(x)=(1−κ​1​)Φs​(x)+κ​1​(f(xs​)+∇f(xs​)⊤(x−xs​)+2α​∥x−xs​∥2),

together with their centres vsv_svs​ (with v1=x1v_1=x_1v1​=x1​ and the recursion (3.21) of the book) and their minimum values Φs∗\Phi^*_sΦs∗​.

Formalization targets

Goal: Theorem 3.18

For every run of the method and every t≥1t\ge1t≥1,

f(yt)−f(x∗)≤α+β2 ∥x1−x∗∥2exp⁡(−t−1κ).f(y_t)-f(x^*)\le\frac{\alpha+\beta}2\,\|x_1-x^*\|^2\exp\Big(-\frac{t-1}{\sqrt\kappa}\Big).f(yt​)−f(x∗)≤2α+β​∥x1​−x∗∥2exp(−κ​t−1​).

Milestones, from the book's proof

  1. (3.18): Φs+1(x)≤f(x)+(1−1/κ)s(Φ1(x)−f(x))\Phi_{s+1}(x)\le f(x)+(1-1/\sqrt\kappa)^s(\Phi_1(x)-f(x))Φs+1​(x)≤f(x)+(1−1/κ​)s(Φ1​(x)−f(x)) for all xxx.
  2. (3.19): f(ys)≤min⁡x∈RnΦs(x)f(y_s)\le\min_{x\in\mathbb R^n}\Phi_s(x)f(ys​)≤minx∈Rn​Φs​(x).
  3. The geometric rate: f(yt)−f(x∗)≤α+β2∥x1−x∗∥2(1−1/κ)t−1f(y_t)-f(x^*)\le\frac{\alpha+\beta}2\|x_1-x^*\|^2(1-1/\sqrt\kappa)^{t-1}f(yt​)−f(x∗)≤2α+β​∥x1​−x∗∥2(1−1/κ​)t−1.
  4. The form Φs(x)=Φs∗+α2∥x−vs∥2\Phi_s(x)=\Phi^*_s+\frac\alpha2\|x-v_s\|^2Φs​(x)=Φs∗​+2α​∥x−vs​∥2 with vsv_svs​ given by (3.21).
  5. The identity (3.22) for Φs+1∗\Phi^*_{s+1}Φs+1∗​.
  6. The inequality (3.20), the inductive step of (3.19).
  7. The coupling vs−xs=κ (xs−ys)v_s-x_s=\sqrt\kappa\,(x_s-y_s)vs​−xs​=κ​(xs​−ys​).

The geometric form in milestone 3 is slightly stronger than the goal, which follows from 1−u≤e−u1-u\le e^{-u}1−u≤e−u.

Significance

Theorem 3.18 gives ε\varepsilonε-accuracy after O(κlog⁡(1/ε))O(\sqrt\kappa\log(1/\varepsilon))O(κ​log(1/ε)) gradient evaluations. Projected gradient descent with step 1/β1/\beta1/β on the same class contracts only at the rate exp⁡(−t/κ)\exp(-t/\kappa)exp(−t/κ) (Theorem 3.10 of the book). The lower bound of Theorem 3.15 shows that no black-box first-order method can do better than ((κ−1)/(κ+1))2(t−1)((\sqrt\kappa-1)/(\sqrt\kappa+1))^{2(t-1)}((κ​−1)/(κ​+1))2(t−1), so the accelerated rate is optimal up to constants. The estimate-sequence argument is the template for many later accelerated methods: proximal, stochastic and variance-reduced variants such as Katyusha, and accelerated coordinate descent.

The result is classical and fully proved on paper. No machine-checked proof of the accelerated rate for strongly convex smooth functions is known to exist in Lean's Mathlib. This mission produces one, with the estimate sequence Φs\Phi_sΦs​, its centres and its minimum values as reusable objects, and with every algebraic identity of the book's proof stated separately.

Difficulty

The algorithm is two lines, but its analysis is not a one-step contraction: neither ∥xt−x∗∥\|x_t-x^*\|∥xt​−x∗∥ nor f(yt)−f(x∗)f(y_t)-f(x^*)f(yt​)−f(x∗) decreases by the factor 1−1/κ1-1/\sqrt\kappa1−1/κ​ at every step. A Lyapunov argument for gradient descent, applied directly to yty_tyt​, gives only the rate 1−1/κ1-1/\kappa1−1/κ. The book obtains the rate through the auxiliary functions Φs\Phi_sΦs​. The inequality (3.18) is easy, but (3.19), that the minimum of Φs\Phi_sΦs​ never drops below f(ys)f(y_s)f(ys​), depends on the exact choice of the momentum factor. It holds only through the identity vs−xs=κ(xs−ys)v_s-x_s=\sqrt\kappa(x_s-y_s)vs​−xs​=κ​(xs​−ys​), which ties the centre of Φs\Phi_sΦs​ to the iterates. Formally, the obstacles are the bookkeeping of the recursive quadratics on Rn\mathbb R^nRn and the algebra in κ\sqrt\kappaκ​, 1/κ1/\sqrt\kappa1/κ​ and 1/(ακ)=κ/β1/(\alpha\sqrt\kappa)=\sqrt\kappa/\beta1/(ακ​)=κ​/β.

Formalization scope

  • Rn\mathbb R^nRn is EuclideanSpace ℝ (Fin n). The gradient is an explicit map g with HasGradientAt f (g x) x at every point, which is part of the smoothness predicate IsBetaSmooth f g β. Strong convexity is the published definition OnlineConvexOpt.ConvexBasics.StronglyConvexOn Set.univ f g α, which is the book's (3.13).
  • A run of the method is a predicate IsNesterovSCRun g α β x y on two sequences indexed from 111, with x1=y1x_1=y_1x1​=y1​ arbitrary. Every theorem holds for every run, that is, every starting point.
  • κ\kappaκ is kappa α β = β / α. All theorems assume α>0\alpha>0α>0 and β>0\beta>0β>0. The second is implied by the other hypotheses for n≥1n\ge1n≥1; no hypothesis α≤β\alpha\le\betaα≤β is added.
  • The existence of a minimizer x∗x^*x∗ is the book's standing assumption, written as a hypothesis.
  • Φs\Phi_sΦs​, vsv_svs​ and Φs∗=Φs(vs)\Phi^*_s=\Phi_s(v_s)Φs∗​=Φs​(vs​) are explicit recursive definitions. The book's Φs∗=min⁡Φs\Phi^*_s=\min\Phi_sΦs∗​=minΦs​ is recovered by milestone 4, and no real infimum is used. The minimum in (3.19) is stated as f(ys)≤Φs(x)f(y_s)\le\Phi_s(x)f(ys​)≤Φs​(x) for every xxx.
  • The identities of milestones 4, 5 and 7 are algebraic and are stated without convexity or smoothness, for arbitrary sequences or runs.
  • Ruled out as trivializing: a run predicate that drops x1=y1x_1=y_1x1​=y1​ breaks (3.19) at s=1s=1s=1 and is not used. A minimum value Φs∗\Phi^*_sΦs∗​ defined through (3.22) would make that identity a tautology, so Φs∗\Phi^*_sΦs∗​ is defined as a value of Φs\Phi_sΦs​.
  • The definitions are local to the namespace ConvexOptAlg.NesterovStrong. β-smoothness duplicates the predicate of other missions of the series and will be merged afterwards. Contributions welcome: proofs of the milestones, and general lemmas on quadratics z↦c+α2∥z−v∥2z\mapsto c+\frac\alpha2\|z-v\|^2z↦c+2α​∥z−v∥2 that the algebraic milestones need.

Selected references

  • S. Bubeck, Convex Optimization: Algorithms and Complexity, Foundations and Trends in Machine Learning 8(3–4):231–358, 2015. arXiv:1405.4980v2, §3.7.1, Theorem 3.18, pp. 290–293.
  • Y. Nesterov, A method of solving a convex programming problem with convergence rate O(1/k²), Soviet Mathematics Doklady 27:372–376, 1983.
  • Y. Nesterov, Introductory Lectures on Convex Optimization: A Basic Course, Kluwer, 2004. doi:10.1007/978-1-4419-8853-9
10 thms0 active usersReviewed
Machine LearningOptimization·Captain: mikedeng1

Convex Optimization: Algorithms and Complexity IX: Nesterov's Accelerated Gradient Descent on a Convex β-Smooth Function Has Rate 2β‖x₁ − x*‖²/t²Textbook

Motivation

Minimizing a convex function f:Rn→Rf:\mathbb R^n\to\mathbb Rf:Rn→R whose gradient is Lipschitz is the basic problem of first-order optimization, and it is the regime in which most large-scale methods of machine learning, signal processing and operations research are analysed. The natural method, gradient descent, reaches accuracy ε\varepsilonε after O(1/ε)O(1/\varepsilon)O(1/ε) gradient evaluations. In 1983 Nesterov showed that a method using the same oracle, but combining the current and the previous iterate, reaches accuracy ε\varepsilonε after only O(1/ε)O(1/\sqrt\varepsilon)O(1/ε​) evaluations, and that no method using only gradient information can do better by more than a constant factor. This accelerated gradient descent and its proximal variants (FISTA) are now the default fast first-order methods for smooth and composite convex problems.

This mission formalizes the accelerated rate as presented in S. Bubeck, Convex Optimization: Algorithms and Complexity (Foundations and Trends in Machine Learning, 2015; arXiv:1405.4980v2), §3.7.2, Theorem 3.19, whose proof follows Beck and Teboulle (2009).

Timeline.

  • 1983: Nesterov introduces the accelerated method with rate O(1/t2)O(1/t^2)O(1/t2) for convex functions with Lipschitz gradient (Soviet Math. Dokl. 27).
  • 1983: Nemirovski and Yudin's black-box lower bounds show that Ω(1/t2)\Omega(1/t^2)Ω(1/t2) is the best possible rate for this class (Theorem 3.14 of the book gives 3β∥x1−x∗∥2/(32(t+1)2)3\beta\|x_1-x^*\|^2/(32(t+1)^2)3β∥x1​−x∗∥2/(32(t+1)2)).
  • 2009: Beck and Teboulle's FISTA extends the method, with the same step sequence λt\lambda_tλt​, to composite problems with a simple nonsmooth term (SIAM J. Imaging Sci. 2(1)).
  • 2008: Tseng gives a unified treatment with simpler step sizes (manuscript).

Setting

Let Rn\mathbb R^nRn carry the Euclidean inner product x⊤yx^\top yx⊤y and norm ∥⋅∥\|\cdot\|∥⋅∥. A differentiable function f:Rn→Rf:\mathbb R^n\to\mathbb Rf:Rn→R is β-smooth if its gradient is β\betaβ-Lipschitz:

∥∇f(x)−∇f(y)∥≤β∥x−y∥(x,y∈Rn).\|\nabla f(x)-\nabla f(y)\|\le\beta\|x-y\|\qquad(x,y\in\mathbb R^n).∥∇f(x)−∇f(y)∥≤β∥x−y∥(x,y∈Rn).

Assume fff is convex and β\betaβ-smooth with β>0\beta>0β>0, and let x∗x^*x∗ be a minimizer of fff.

Define the step sequences λ0=0\lambda_0=0λ0​=0,

λt=1+1+4λt−122(t≥1),γt=1−λtλt+1.\lambda_t=\frac{1+\sqrt{1+4\lambda_{t-1}^2}}{2}\quad(t\ge1),\qquad\gamma_t=\frac{1-\lambda_t}{\lambda_{t+1}}.λt​=21+1+4λt−12​​​(t≥1),γt​=λt+1​1−λt​​.

Then λ1=1\lambda_1=1λ1​=1, λt≥1\lambda_t\ge1λt​≥1 for t≥1t\ge1t≥1, and γt≤0\gamma_t\le0γt​≤0. Nesterov's accelerated gradient descent for the smooth case starts from an arbitrary point x1=y1x_1=y_1x1​=y1​ and sets, for t≥1t\ge1t≥1,

yt+1=xt−1β∇f(xt),xt+1=(1−γt) yt+1+γt yt.y_{t+1}=x_t-\frac1\beta\nabla f(x_t),\qquad x_{t+1}=(1-\gamma_t)\,y_{t+1}+\gamma_t\,y_t.yt+1​=xt​−β1​∇f(xt​),xt+1​=(1−γt​)yt+1​+γt​yt​.

The primary sequence (yt)(y_t)(yt​) consists of gradient steps of length 1/β1/\beta1/β; the sequence (xt)(x_t)(xt​), at which the gradient is queried, moves past yt+1y_{t+1}yt+1​ away from yty_tyt​ (since γt≤0\gamma_t\le0γt​≤0). Write δt=f(yt)−f(x∗)\delta_t=f(y_t)-f(x^*)δt​=f(yt​)−f(x∗) for the optimality gap.

Formalization targets

Goal: Theorem 3.19

For every run of the method and every t≥1t\ge1t≥1,

f(yt)−f(x∗)≤2β∥x1−x∗∥2t2.f(y_t)-f(x^*)\le\frac{2\beta\|x_1-x^*\|^2}{t^2}.f(yt​)−f(x∗)≤t22β∥x1​−x∗∥2​.

The constant 222 and the exponent 222 are those printed in the book.

Milestones

In the order of the book's proof (pp. 294–295):

  1. Lemma 3.6, unconstrained (p. 270): f(x−1β∇f(x))−f(y)≤∇f(x)⊤(x−y)−12β∥∇f(x)∥2f(x-\tfrac1\beta\nabla f(x))-f(y)\le\nabla f(x)^\top(x-y)-\tfrac1{2\beta}\|\nabla f(x)\|^2f(x−β1​∇f(x))−f(y)≤∇f(x)⊤(x−y)−2β1​∥∇f(x)∥2 for all x,yx,yx,y.
  2. (3.23): f(ys+1)−f(ys)≤β(xs−ys+1)⊤(xs−ys)−β2∥xs−ys+1∥2f(y_{s+1})-f(y_s)\le\beta(x_s-y_{s+1})^\top(x_s-y_s)-\tfrac\beta2\|x_s-y_{s+1}\|^2f(ys+1​)−f(ys​)≤β(xs​−ys+1​)⊤(xs​−ys​)−2β​∥xs​−ys+1​∥2.
  3. (3.24): f(ys+1)−f(x∗)≤β(xs−ys+1)⊤(xs−x∗)−β2∥xs−ys+1∥2f(y_{s+1})-f(x^*)\le\beta(x_s-y_{s+1})^\top(x_s-x^*)-\tfrac\beta2\|x_s-y_{s+1}\|^2f(ys+1​)−f(x∗)≤β(xs​−ys+1​)⊤(xs​−x∗)−2β​∥xs​−ys+1​∥2.
  4. The λ identity: λs−12=λs2−λs\lambda_{s-1}^2=\lambda_s^2-\lambda_sλs−12​=λs2​−λs​ for s≥1s\ge1s≥1.
  5. (3.25): λs2δs+1−λs−12δs≤β2(∥λsxs−(λs−1)ys−x∗∥2−∥λsys+1−(λs−1)ys−x∗∥2)\lambda_s^2\delta_{s+1}-\lambda_{s-1}^2\delta_s\le\tfrac\beta2\bigl(\|\lambda_sx_s-(\lambda_s-1)y_s-x^*\|^2-\|\lambda_sy_{s+1}-(\lambda_s-1)y_s-x^*\|^2\bigr)λs2​δs+1​−λs−12​δs​≤2β​(∥λs​xs​−(λs​−1)ys​−x∗∥2−∥λs​ys+1​−(λs​−1)ys​−x∗∥2).
  6. (3.26): λs+1xs+1−(λs+1−1)ys+1=λsys+1−(λs−1)ys\lambda_{s+1}x_{s+1}-(\lambda_{s+1}-1)y_{s+1}=\lambda_sy_{s+1}-(\lambda_s-1)y_sλs+1​xs+1​−(λs+1​−1)ys+1​=λs​ys+1​−(λs​−1)ys​.
  7. Telescoped bound: δt≤β2λt−12∥u1∥2\delta_t\le\frac{\beta}{2\lambda_{t-1}^2}\|u_1\|^2δt​≤2λt−12​β​∥u1​∥2 for t≥2t\ge2t≥2, with u1=λ1x1−(λ1−1)y1−x∗u_1=\lambda_1x_1-(\lambda_1-1)y_1-x^*u1​=λ1​x1​−(λ1​−1)y1​−x∗.
  8. Growth of λ: λt−1≥t/2\lambda_{t-1}\ge t/2λt−1​≥t/2 for t≥2t\ge2t≥2.

Significance

The result. Theorem 3.19 is the upper half of the statement that first-order methods on smooth convex functions have complexity Θ(β∥x1−x∗∥2/ε)\Theta(\sqrt{\beta\|x_1-x^*\|^2/\varepsilon})Θ(β∥x1​−x∗∥2/ε​): together with the black-box lower bound of Theorem 3.14 it shows that the accelerated method is optimal up to a constant factor, while plain gradient descent (Theorem 3.3, rate 2β∥x1−x∗∥2/(t−1)2\beta\|x_1-x^*\|^2/(t-1)2β∥x1​−x∗∥2/(t−1)) is not. The same potential-function argument, with the gradient step replaced by a proximal step, yields the rate of FISTA for composite objectives, and the identity λs−12=λs2−λs\lambda_{s-1}^2=\lambda_s^2-\lambda_sλs−12​=λs2​−λs​ is the algebraic core of most later analyses of accelerated and momentum methods.

Formalizing it. The theorem is classical and fully proved on paper. What this mission adds is a machine-checked proof of the O(1/t2)O(1/t^2)O(1/t2) rate for the exact step sequence of the book, with every intermediate inequality of the proof stated separately so that each can be closed and reused. A Lean development of accelerated gradient descent with this rate is not, to our knowledge, in Mathlib.

Difficulty

The rate does not follow from monotone decrease of f(yt)f(y_t)f(yt​), which is the engine of the analysis of plain gradient descent: the accelerated iterates are not monotone, and no single-step inequality on δt\delta_tδt​ alone gives 1/t21/t^21/t2. The proof instead tracks a weighted potential λs−12δs+β2∥us∥2\lambda_{s-1}^2\delta_s+\frac\beta2\|u_s\|^2λs−12​δs​+2β​∥us​∥2 and needs three exact algebraic coincidences: the identity λs−12=λs2−λs\lambda_{s-1}^2=\lambda_s^2-\lambda_sλs−12​=λs2​−λs​, the completion of a square 2a⊤b−∥a∥2=∥b∥2−∥b−a∥22a^\top b-\|a\|^2=\|b\|^2-\|b-a\|^22a⊤b−∥a∥2=∥b∥2−∥b−a∥2, and the rewriting (3.26) of the update rule, which makes consecutive potentials match exactly. In Lean the work is in this bookkeeping (vector identities in an inner-product space, a real recursion with square roots, and a telescoping sum indexed from 111), not in any deep analytic fact.

Formalization scope

  • Rn\mathbb R^nRn is EuclideanSpace ℝ (Fin n); x⊤yx^\top yx⊤y is ⟪x, y⟫_ℝ.
  • The gradient is an explicit map ggg with HasGradientAt f (g x) x for every xxx; β-smoothness is the Lipschitz bound on ggg (not the quadratic upper bound (3.4), which is a consequence). Convexity is ConvexOn ℝ Set.univ f.
  • λ\lambdaλ is a real sequence defined by recursion on N\mathbb NN with λ0=0\lambda_0=0λ0​=0; γt=(1−λt)/λt+1\gamma_t=(1-\lambda_t)/\lambda_{t+1}γt​=(1−λt​)/λt+1​ never divides by zero because λt+1≥1\lambda_{t+1}\ge1λt+1​≥1.
  • A run is a predicate on two sequences x,y:N→Rnx,y:\mathbb N\to\mathbb R^nx,y:N→Rn, indexed from 111 as in the book; index 000 is unconstrained. Every theorem quantifies over all runs.
  • Standing assumptions: x∗x^*x∗ is a minimizer of fff (book, p. 242); β>0\beta>0β>0 (implicit in the step 1/β1/\beta1/β) is a stated hypothesis.
  • The page prints the update as xt+1=(1−γs)yt+1+γtytx_{t+1}=(1-\gamma_s)y_{t+1}+\gamma_ty_txt+1​=(1−γs​)yt+1​+γt​yt​; the formalization uses γt\gamma_tγt​ in both places, as the proof's (3.26) requires. A run predicate with a fixed coefficient γs\gamma_sγs​ would describe a different (non-accelerated) method and is ruled out.
  • The goal is stated for all t≥1t\ge1t≥1, including t=1t=1t=1, which the printed proof does not cover but which follows from (3.4). The growth bound λt−1≥t/2\lambda_{t-1}\ge t/2λt−1​≥t/2 is stated for t≥2t\ge2t≥2, since λ0=0\lambda_0=0λ0​=0. The telescoping display prints δs2\delta_s^2δs2​ for δs\delta_sδs​; the formalization uses δs\delta_sδs​.

Contributions of every kind are welcome: proofs of the λ facts and of (3.26) are pure algebra; (3.23)–(3.25) need the descent lemma for β-smooth functions and first-order convexity, both of which are reusable well beyond this mission.

Selected references

  • S. Bubeck, Convex Optimization: Algorithms and Complexity, Foundations and Trends in Machine Learning 8(3–4):231–358, 2015. arXiv:1405.4980
  • Y. Nesterov, A method of solving a convex programming problem with convergence rate O(1/k²), Soviet Mathematics Doklady 27:372–376, 1983. mathnet
  • A. Beck, M. Teboulle, A fast iterative shrinkage-thresholding algorithm for linear inverse problems, SIAM J. Imaging Sciences 2(1):183–202, 2009. doi:10.1137/080716542
  • P. Tseng, On accelerated proximal gradient methods for convex-concave optimization, manuscript, 2008. pdf
  • A. Nemirovski, D. Yudin, Problem Complexity and Method Efficiency in Optimization, Wiley, 1983.
10 thms0 active usersReviewed
Machine LearningOptimization·Captain: mikedeng1

Convex Optimization: Algorithms and Complexity VIII: No Black-Box Method Beats 3β‖x₁ − x*‖²/(32(t + 1)²) on β-Smooth Convex FunctionsTextbook

Why lower bounds for first-order methods

Upper bounds for an optimization method say how fast it converges; oracle complexity lower bounds say how fast any method of a given kind can possibly converge. Chapter 3 of S. Bubeck, Convex Optimization: Algorithms and Complexity (Foundations and Trends in Machine Learning 8(3–4), 2015, arXiv:1405.4980) proves upper bounds for subgradient descent on Lipschitz functions and for gradient methods on smooth functions. Section 3.5 (Lower bounds, pp. 279–283) shows that these rates cannot be improved by more than a numerical constant, as long as the number of queries is smaller than the dimension. For smooth convex functions the matching lower bound is what identifies Nesterov's accelerated gradient descent, with its 1/t21/t^21/t2 rate, as an optimal method.

Timeline. The lower bounds first appeared in A. Nemirovski and D. Yudin, Problem Complexity and Method Efficiency in Optimization (Wiley, 1983). The presentation followed by the book, with an explicit tridiagonal quadratic as the hard instance and the "span of past gradients" restriction on the method, is that of Y. Nesterov, Introductory Lectures on Convex Optimization (Kluwer, 2004), §2.1.2. Nesterov's accelerated method (1983) attains the matching upper bound.

Setting

Work in Rn\mathbb R^nRn with the Euclidean inner product x⊤yx^\top yx⊤y, coordinates x(1),…,x(n)x(1),\dots,x(n)x(1),…,x(n), canonical basis e1,…,ene_1,\dots,e_ne1​,…,en​ and balls B2(R)={x:∥x∥≤R}\mathrm B_2(R)=\{x:\|x\|\le R\}B2​(R)={x:∥x∥≤R}. A first-order oracle for fff answers a query xxx with a subgradient g∈∂f(x)g\in\partial f(x)g∈∂f(x) (the gradient when fff is differentiable). A black-box procedure maps the history (x1,g1,…,xt,gt)(x_1,g_1,\dots,x_t,g_t)(x1​,g1​,…,xt​,gt​) to the next query xt+1x_{t+1}xt+1​. Section 3.5 restricts attention to procedures with

x1=0,xt+1∈Span(g1,…,gt)(t≥0),(3.15)x_1=0,\qquad x_{t+1}\in\mathrm{Span}(g_1,\dots,g_t)\quad(t\ge0), \tag{3.15}x1​=0,xt+1​∈Span(g1​,…,gt​)(t≥0),(3.15)

which covers gradient descent, its accelerated variants and conjugate gradient. A function is β\betaβ-smooth if its gradient is β\betaβ-Lipschitz; LLL-Lipschitz on X\mathcal XX if every subgradient at every point of X\mathcal XX has norm at most LLL; α\alphaα-strongly convex if x↦f(x)−α2∥x∥2x\mapsto f(x)-\frac\alpha2\|x\|^2x↦f(x)−2α​∥x∥2 is convex.

The hard smooth instance uses, for k≤nk\le nk≤n, the symmetric tridiagonal matrix AkA_kAk​ with entries 222 on the first kkk diagonal positions and −1-1−1 on the neighbouring off-diagonal positions of the leading k×kk\times kk×k block, zero elsewhere, and the quadratics

fk(x)=β8x⊤Akx−β4x⊤e1,fk∗=inf⁡x∈Rnfk(x).f_k(x)=\frac\beta8x^\top A_kx-\frac\beta4x^\top e_1 ,\qquad f_k^*=\inf_{x\in\mathbb R^n}f_k(x).fk​(x)=8β​x⊤Ak​x−4β​x⊤e1​,fk∗​=x∈Rninf​fk​(x).

Formalization targets

Goal: Theorem 3.14 (p. 282)

For 1≤t≤n−121\le t\le\frac{n-1}21≤t≤2n−1​ and β>0\beta>0β>0 there are a β\betaβ-smooth convex fff and a minimizer x∗x^*x∗ such that every procedure satisfying (3.15) has

min⁡1≤s≤tf(xs)−f(x∗) ≥ 3β32 ∥x1−x∗∥2(t+1)2.\min_{1\le s\le t}f(x_s)-f(x^*)\ \ge\ \frac{3\beta}{32}\,\frac{\|x_1-x^*\|^2}{(t+1)^2}.1≤s≤tmin​f(xs​)−f(x∗) ≥ 323β​(t+1)2∥x1​−x∗∥2​.

The constant 3/323/323/32 is the book's.

Milestones (proof of Theorem 3.14, pp. 282–283)

  1. 0⪯Ak⪯4In0\preceq A_k\preceq4I_n0⪯Ak​⪯4In​, through x⊤Akx=x(1)2+x(k)2+∑i=1k−1(x(i)−x(i+1))2x^\top A_kx=x(1)^2+x(k)^2+\sum_{i=1}^{k-1}(x(i)-x(i+1))^2x⊤Ak​x=x(1)2+x(k)2+∑i=1k−1​(x(i)−x(i+1))2.
  2. For f=f2t+1f=f_{2t+1}f=f2t+1​ and any procedure satisfying (3.15), xs∈Span(e1,…,es−1)x_s\in\mathrm{Span}(e_1,\dots,e_{s-1})xs​∈Span(e1​,…,es−1​); hence f(xs)=fs(xs)f(x_s)=f_s(x_s)f(xs​)=fs​(xs​) for s≤ts\le ts≤t.
  3. xk∗(i)=1−ik+1x_k^*(i)=1-\frac i{k+1}xk∗​(i)=1−k+1i​ solves Akx=e1A_kx=e_1Ak​x=e1​, minimizes fkf_kfk​, and fk∗=−β8(1−1k+1)f_k^*=-\frac\beta8\bigl(1-\frac1{k+1}\bigr)fk∗​=−8β​(1−k+11​).
  4. ∥xk∗∥2≤k+13\|x_k^*\|^2\le\frac{k+1}3∥xk∗​∥2≤3k+1​.
  5. ft∗−f2t+1∗=β8(1t+1−12t+2)≥3β32∥x2t+1∗∥2(t+1)2f_t^*-f_{2t+1}^*=\frac\beta8\bigl(\frac1{t+1}-\frac1{2t+2}\bigr)\ge\frac{3\beta}{32}\frac{\|x^*_{2t+1}\|^2}{(t+1)^2}ft∗​−f2t+1∗​=8β​(t+11​−2t+21​)≥323β​(t+1)2∥x2t+1∗​∥2​.

Companion: Theorem 3.13 (p. 280)

For 1≤t≤n1\le t\le n1≤t≤n and L,R>0L,R>0L,R>0 there are a convex fff, LLL-Lipschitz on B2(R)\mathrm B_2(R)B2​(R), and a first-order oracle for it such that every procedure satisfying (3.15) has min⁡s≤tf(xs)−min⁡B2(R)f≥RL2(1+t)\min_{s\le t}f(x_s)-\min_{\mathrm B_2(R)}f\ge\frac{RL}{2(1+\sqrt t)}mins≤t​f(xs​)−minB2​(R)​f≥2(1+t​)RL​; and for α>0\alpha>0α>0 there are an α\alphaα-strongly convex fff, LLL-Lipschitz on B2(L2α)\mathrm B_2(\frac L{2\alpha})B2​(2αL​), and an oracle with gap at least L28αt\frac{L^2}{8\alpha t}8αtL2​ over that ball.

Significance

The upper bounds of Chapter 3 (projected subgradient descent at rate RL/tRL/\sqrt tRL/t​, accelerated gradient descent at rate β∥x1−x∗∥2/t2\beta\|x_1-x^*\|^2/t^2β∥x1​−x∗∥2/t2) become optimal statements only through these lower bounds: no method in the class (3.15) can be faster by more than a constant factor while ttt is below the dimension. The restriction to t≲nt\lesssim nt≲n is necessary, since Chapter 2's cutting-plane methods converge exponentially once the number of queries exceeds the dimension.

The results are classical and proved. Formalizing them yields machine-checked versions of the quadratic-form computation for the tridiagonal matrix, of the Krylov-type support argument under (3.15), and of the explicit minimizer of fkf_kfk​, each reusable in other lower-bound arguments (Theorem 3.15 in ℓ2\ell_2ℓ2​, lower bounds for strongly convex smooth functions, conjugate gradient analyses). The platform held no formal statement of these oracle lower bounds when this mission was drafted.

Difficulty

Each analytic step is elementary; the difficulty is in the bookkeeping. The span argument is an induction that must track, at each step, that the gradient of a tridiagonal quadratic at a vector supported on the first s−1s-1s−1 coordinates is supported on the first sss, and that the span hypothesis transfers this to the next query. The minimizer computation requires solving Akx=e1A_kx=e_1Ak​x=e1​ on the leading block and showing that the coordinates beyond kkk do not affect fkf_kfk​. A natural first attempt, choosing the hard function after seeing the procedure, proves a much weaker statement and is excluded by the quantifier order: the function is fixed first and must defeat every procedure.

Formalization scope

Rn\mathbb R^nRn is EuclideanSpace ℝ (Fin n). Coordinates in Lean are 0-based; the definitions provide the book's 1-based coordinate coord x i and basis vector basisVec n i, and the matrix tridiag n k translates book index iii to Fin n index i−1i-1i−1. The query sequence starts at index 111. The oracle is a fixed map ggg, so (3.15) reads xt+1∈Span(g(x1),…,g(xt))x_{t+1}\in\mathrm{Span}(g(x_1),\dots,g(x_t))xt+1​∈Span(g(x1​),…,g(xt​)) with x1=0x_1=0x1​=0. In Theorem 3.14 the oracle is the gradient, given as a map with HasGradientAt everywhere; β\betaβ-smoothness is the Lipschitz bound on that map. In Theorem 3.13 the oracle is part of what is constructed, because the book's proof uses a specific "resisting" subgradient selection and the claim fails for an arbitrary one.

Committed conventions, each stated in the item's Formalization Note: the minimum over 1≤s≤t1\le s\le t1≤s≤t is the bound for every such sss, and t≥1t\ge1t≥1 is required; t≤(n−1)/2t\le(n-1)/2t≤(n−1)/2 is 2t+1≤n2t+1\le n2t+1≤n; the minimizer x∗x^*x∗ is existentially chosen together with fff (the hard function has many minimizers when 2t+1<n2t+1<n2t+1<n, and the bound is false for some of them); the minimum over a ball is the bound against every point of the ball; fk∗f_k^*fk∗​ is the real infimum, asserted to be attained.

A formalization that let the function depend on the procedure, dropped x1=0x_1=0x1​=0, or took the span over gradients at points other than the queries would state a different and weaker theorem; the statements here keep fff (and the oracle) before the universally quantified procedure.

Infrastructure needed: quadratic forms of explicit matrices on EuclideanSpace, gradients of quadratics, and span/support lemmas for EuclideanSpace.single-type vectors. Contributions of proofs for any milestone, and of the strongly convex ℓ2\ell_2ℓ2​ lower bound (Theorem 3.15, not included here), are welcome.

Selected references

  • S. Bubeck, Convex Optimization: Algorithms and Complexity, Foundations and Trends in Machine Learning 8(3–4):231–358, 2015. https://arxiv.org/abs/1405.4980
  • A. Nemirovski and D. Yudin, Problem Complexity and Method Efficiency in Optimization, Wiley, 1983.
  • Y. Nesterov, Introductory Lectures on Convex Optimization: A Basic Course, Kluwer, 2004. https://doi.org/10.1007/978-1-4419-8853-9
8 thms0 active usersReviewed
Machine LearningOptimization·Captain: mikedeng1

Convex Optimization: Algorithms and Complexity VII: Conditional Gradient Descent (Frank–Wolfe) with γ_s = 2/(s + 1) Has Rate 2βR²/(t + 1) in Any NormTextbook

Motivation

Many constrained optimization problems in machine learning and statistics have a feasible set X\mathcal XX over which a linear function is cheap to minimize but a Euclidean projection is expensive: the ℓ1\ell_1ℓ1​-ball, the simplex, the nuclear-norm ball, the convex hull of a combinatorial family. Projected gradient descent needs a projection at every step. Conditional gradient descent, introduced by Frank and Wolfe in 1956 for quadratic programming, replaces the projection by a call to a linear minimization oracle over X\mathcal XX, and its iterates are convex combinations of oracle outputs, which makes them sparse when X\mathcal XX is a polytope.

This mission is the seventh of a series formalizing S. Bubeck's monograph Convex Optimization: Algorithms and Complexity (Foundations and Trends in Machine Learning, 2015, arXiv:1405.4980v2). It covers Section 3.3, whose main result, Theorem 3.8, is the O(1/t)O(1/t)O(1/t) rate of the method in the form given by Jaggi (2013), going back to Dunn and Harshbarger (1978).

Setting

Let EEE be a finite-dimensional real vector space with an arbitrary norm ∥⋅∥\|\cdot\|∥⋅∥. For a linear form ggg on EEE, written v↦g⊤vv\mapsto g^\top vv↦g⊤v, the dual norm is ∥g∥∗=sup⁡∥v∥≤1g⊤v\|g\|_*=\sup_{\|v\|\le1}g^\top v∥g∥∗​=sup∥v∥≤1​g⊤v. Let X⊆E\mathcal X\subseteq EX⊆E be nonempty, compact and convex, with diameter R=sup⁡x,y∈X∥x−y∥R=\sup_{x,y\in\mathcal X}\|x-y\|R=supx,y∈X​∥x−y∥.

Let f:E→Rf:E\to\mathbb Rf:E→R be differentiable with gradient ∇f(x)\nabla f(x)∇f(x), a linear form on EEE. For β≥0\beta\ge0β≥0, fff is β-smooth with respect to ∥⋅∥\|\cdot\|∥⋅∥ on X\mathcal XX if

∥∇f(x)−∇f(y)∥∗≤β∥x−y∥(x,y∈X).\|\nabla f(x)-\nabla f(y)\|_*\le\beta\|x-y\|\qquad(x,y\in\mathcal X).∥∇f(x)−∇f(y)∥∗​≤β∥x−y∥(x,y∈X).

A point x∗∈Xx^*\in\mathcal Xx∗∈X with f(x∗)=min⁡x∈Xf(x)f(x^*)=\min_{x\in\mathcal X}f(x)f(x∗)=minx∈X​f(x) is fixed throughout, and δt=f(xt)−f(x∗)\delta_t=f(x_t)-f(x^*)δt​=f(xt​)−f(x∗).

Given step sizes (γs)s≥1(\gamma_s)_{s\ge1}(γs​)s≥1​, a run of conditional gradient descent is a pair of sequences with x1∈Xx_1\in\mathcal Xx1​∈X and, for every t≥1t\ge1t≥1,

yt∈argmin⁡y∈X∇f(xt)⊤y(3.8),xt+1=(1−γt)xt+γtyt(3.9).y_t\in\operatorname*{argmin}_{y\in\mathcal X}\nabla f(x_t)^\top y\quad(3.8),\qquad x_{t+1}=(1-\gamma_t)x_t+\gamma_ty_t\quad(3.9).yt​∈y∈Xargmin​∇f(xt​)⊤y(3.8),xt+1​=(1−γt​)xt​+γt​yt​(3.9).

The minimizer yty_tyt​ need not be unique; any choice is allowed.

Formalization targets

Goal: Theorem 3.8 (p. 272)

If fff is convex and β\betaβ-smooth with respect to ∥⋅∥\|\cdot\|∥⋅∥ and γs=2s+1\gamma_s=\frac{2}{s+1}γs​=s+12​ for s≥1s\ge1s≥1, then every run satisfies, for every t≥2t\ge2t≥2,

f(xt)−f(x∗)≤2βR2t+1.f(x_t)-f(x^*)\le\frac{2\beta R^2}{t+1}.f(xt​)−f(x∗)≤t+12βR2​.

Milestones

  1. Inequality (3.4) in an arbitrary norm (p. 267, used on p. 272): for x,y∈Xx,y\in\mathcal Xx,y∈X, 0≤f(x)−f(y)−∇f(y)⊤(x−y)≤β2∥x−y∥20\le f(x)-f(y)-\nabla f(y)^\top(x-y)\le\frac{\beta}{2}\|x-y\|^20≤f(x)−f(y)−∇f(y)⊤(x−y)≤2β​∥x−y∥2.
  2. The one-step recursion (pp. 272–273): for any run with γs∈[0,1]\gamma_s\in[0,1]γs​∈[0,1],
δs+1≤(1−γs)δs+β2γs2R2.\delta_{s+1}\le(1-\gamma_s)\delta_s+\frac{\beta}{2}\gamma_s^2R^2 .δs+1​≤(1−γs​)δs​+2β​γs2​R2.
  1. Initialization (p. 273): if γ1=1\gamma_1=1γ1​=1, then δ2≤β2R2\delta_2\le\frac{\beta}{2}R^2δ2​≤2β​R2.
  2. The induction (p. 273): a real sequence with δ2≤β2R2\delta_2\le\frac{\beta}{2}R^2δ2​≤2β​R2 and the recursion of milestone 2 with γs=2s+1\gamma_s=\frac{2}{s+1}γs​=s+12​ for s≥2s\ge2s≥2 satisfies δt≤2βR2t+1\delta_t\le\frac{2\beta R^2}{t+1}δt​≤t+12βR2​ for t≥2t\ge2t≥2.

Significance

Theorem 3.8 is the basic guarantee for projection-free first-order optimization. Its rate does not depend on the dimension, and it depends on the geometry only through the product βR2\beta R^2βR2, both measured in the same norm, which may be chosen to fit X\mathcal XX: for the ℓ1\ell_1ℓ1​-ball, smoothness in ∥⋅∥1\|\cdot\|_1∥⋅∥1​ with dual norm ∥⋅∥∞\|\cdot\|_\infty∥⋅∥∞​ gives much better constants than the Euclidean analysis. The book applies it this way to a LASSO-type problem (Section 3.3, pp. 274–276), and the same bound underlies the sparse-approximation corollary on the simplex (p. 273) and the many later variants of the method (away steps, stochastic and online conditional gradient).

The result is classical and proved in the book. What this mission adds is a machine-checked statement and proof in an arbitrary finite-dimensional normed space, with the dual norm as the operator norm on linear forms, and a reusable encoding of norm-smoothness and of conditional gradient runs. On Prove2Me a related result is already proved: Lan's Theorem 7.1 (First-order and Stochastic Optimization Methods), which bounds f(yk)−f∗f(y_k)-f^*f(yk​)−f∗ by 2Lk(k+1)∑i≤k∥xi−yi−1∥2\frac{2L}{k(k+1)}\sum_{i\le k}\|x_i-y_{i-1}\|^2k(k+1)2L​∑i≤k​∥xi​−yi−1​∥2 with a different indexing; after the diameter bound it yields 2βR2/t2\beta R^2/t2βR2/t at Bubeck's iterate xtx_txt​, which is weaker than Theorem 3.8 by one step.

Difficulty

The difficulty is in the bookkeeping of norms and indices, not in a deep idea. Inequality (3.4) is proved in the book only for the Euclidean norm, where ∇f(x)∈Rn\nabla f(x)\in\mathbb R^n∇f(x)∈Rn and the Cauchy–Schwarz inequality is used; in a general norm it needs the pairing between a linear form and a vector and the bound ∣g⊤v∣≤∥g∥∗∥v∥|g^\top v|\le\|g\|_*\|v\|∣g⊤v∣≤∥g∥∗​∥v∥. The rate 2βR2/(t+1)2\beta R^2/(t+1)2βR2/(t+1) is attained only by starting the induction at t=2t=2t=2, where the step γ1=1\gamma_1=1γ1​=1 erases the dependence on the starting point; at t=1t=1t=1 the bound can fail, since δ1\delta_1δ1​ is arbitrary. A first attempt that runs the induction from t=1t=1t=1 with an arbitrary δ1\delta_1δ1​ does not give the stated constant.

Formalization scope

EEE is a type with [NormedAddCommGroup E] [NormedSpace ℝ E] [FiniteDimensional ℝ E]; nothing is specialised to the Euclidean norm. The gradient is an explicit derivative map f' : E → (E →L[ℝ] ℝ) with HasFDerivAt f (f' x) x for every x; ∇f(x)⊤v\nabla f(x)^\top v∇f(x)⊤v is f' x v, and ∥⋅∥∗\|\cdot\|_*∥⋅∥∗​ is the operator norm, which equals sup⁡∥v∥≤1g⊤v\sup_{\|v\|\le1}g^\top vsup∥v∥≤1​g⊤v. RRR is Metric.diam X, which equals the supremum of ∥x−y∥\|x-y\|∥x−y∥ over X\mathcal XX because X\mathcal XX is compact. Sequences are indexed by ℕ with the first iterate at index 1.

Committed conventions: X\mathcal XX compact, convex and containing x∗x^*x∗ (hence nonempty); convexity of fff and the Lipschitz bound on the gradient are assumed on X\mathcal XX only, which is weaker than the book's global assumptions; β≥0\beta\ge0β≥0; the existence of the minimizer x∗x^*x∗ is the book's standing assumption (p. 242); the conclusion is stated for t≥2t\ge2t≥2, as in the book. Runs are predicates: yty_tyt​ is any minimizer of the linear form over X\mathcal XX and yt∈Xy_t\in\mathcal Xyt​∈X is required, so the goal quantifies over every run with γs=2/(s+1)\gamma_s=2/(s+1)γs​=2/(s+1). A formalization in which yty_tyt​ need not lie in X\mathcal XX, or in which smoothness is assumed only along the iterates, states a different theorem and is excluded.

A complete development needs the descent inequality (3.4) for Fréchet derivatives in a normed space (reusable for every smooth method in the series and beyond), the fact that the iterates stay in X\mathcal XX, and a scalar induction. Proofs of the milestones are welcome independently; milestone 4 is a statement about real sequences only.

Selected references

  • S. Bubeck, Convex Optimization: Algorithms and Complexity, Foundations and Trends in Machine Learning 8(3–4):231–358, 2015. https://arxiv.org/abs/1405.4980
  • M. Frank and P. Wolfe, An algorithm for quadratic programming, Naval Research Logistics Quarterly 3(1–2):95–110, 1956. https://doi.org/10.1002/nav.3800030109
  • J. C. Dunn and S. Harshbarger, Conditional gradient algorithms with open loop step size rules, Journal of Mathematical Analysis and Applications 62(2):432–444, 1978. https://doi.org/10.1016/0022-247X(78)90137-3
  • M. Jaggi, Revisiting Frank–Wolfe: projection-free sparse convex optimization, Proceedings of ICML 2013, PMLR 28(1):427–435. https://proceedings.mlr.press/v28/jaggi13.html
  • G. Lan, First-order and Stochastic Optimization Methods for Machine Learning, Springer, 2020, Theorem 7.1. https://doi.org/10.1007/978-3-030-39568-1
6 thms0 active usersReviewed
PreviousPage 10 of 11Next

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me