Prove2Me
Navigate
DiscoverFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

Optimization

661 missions · 438 completed

Missions

Open223Completed438All661
Convex OptimizationMachine LearningProbability·Captain: mikedeng1

Convex Optimization: Algorithms and Complexity XV: SVRG with η = 1/(10β) and k = 20κ Contracts the Expected Optimality Gap by 0.9 per EpochTextbook

Motivation

Many optimization problems in machine learning minimize an average of losses, one loss for each observation. A full gradient step examines every observation, while a stochastic gradient step examines one. The latter is cheaper per step, but its sampled gradient can remain noisy even near the optimum. Section 6.3 of Bubeck's monograph studies stochastic variance reduced gradient descent (SVRG), which periodically computes a full gradient at an anchor point and uses it to correct subsequent sampled gradients. The question for this mission is whether that correction gives a geometric reduction of the expected objective gap at the constants printed in Theorem 6.5.

Bubeck places this method alongside full gradient descent and stochastic gradient descent for finite sums. The section records that earlier stochastic average gradient and dual coordinate ascent methods attain a gradient-computation cost of order (m+κ)log⁡(1/ε)(m+\kappa)\log(1/\varepsilon)(m+κ)log(1/ε) for the same regime, where mmm is the number of components and κ\kappaκ is a condition number. The target here is the precise SVRG convergence statement in the book, rather than a comparison of implementation costs. The source's discussion on pp. 334–336 gives the context and the algorithm.

Setting

Let f1,…,fm:Rn→Rf_1,\ldots,f_m:\mathbb R^n\to\mathbb Rf1​,…,fm​:Rn→R be differentiable convex functions, with m≥1m\ge1m≥1, and define the finite-sum objective and its gradient by

f(x)=1m∑i=1mfi(x),G(x)=1m∑i=1m∇fi(x).f(x)=\frac1m\sum_{i=1}^m f_i(x),\qquad G(x)=\frac1m\sum_{i=1}^m \nabla f_i(x).f(x)=m1​i=1∑m​fi​(x),G(x)=m1​i=1∑m​∇fi​(x).

Each component is β\betaβ-smooth when its gradient is β\betaβ-Lipschitz in the Euclidean norm: ∥∇fi(x)−∇fi(z)∥2≤β∥x−z∥2\|\nabla f_i(x)-\nabla f_i(z)\|_2\le\beta\|x-z\|_2∥∇fi​(x)−∇fi​(z)∥2​≤β∥x−z∥2​ for all x,zx,zx,z. The average fff is α\alphaα-strongly convex, meaning that for all x,zx,zx,z it lies at least α2∥z−x∥22\frac\alpha2\|z-x\|_2^22α​∥z−x∥22​ above its first-order affine approximation at xxx. The constants α\alphaα and β\betaβ are positive, x∗x^*x∗ minimizes fff over Rn\mathbb R^nRn, and κ=β/α\kappa=\beta/\alphaκ=β/α.

An epoch begins at an anchor yyy. Its first inner iterate is x1=yx_1=yx1​=y. For t=1,…,kt=1,\ldots,kt=1,…,k, draw iti_tit​ uniformly from {1,…,m}\{1,\ldots,m\}{1,…,m}, independently across steps and epochs, and update

xt+1=xt−η(∇fit(xt)−∇fit(y)+G(y)).x_{t+1}=x_t-\eta\bigl(\nabla f_{i_t}(x_t)-\nabla f_{i_t}(y)+G(y)\bigr).xt+1​=xt​−η(∇fit​​(xt​)−∇fit​​(y)+G(y)).

The next anchor is the average y+=k−1∑t=1kxty^+=k^{-1}\sum_{t=1}^k x_ty+=k−1∑t=1k​xt​. In particular, this average uses x1x_1x1​ through xkx_kxk​, while the last updated point xk+1x_{k+1}xk+1​ is excluded. Starting from an arbitrary y(1)y^{(1)}y(1) and repeating the epoch produces y(s+1)y^{(s+1)}y(s+1). The expectation of f(y(s+1))f(y^{(s+1)})f(y(s+1)) is over all sksksk sampled indices in the first sss epochs.

Formalization targets

Goal: geometric contraction across epochs

Theorem 6.5 sets η=1/(10β)\eta=1/(10\beta)η=1/(10β) and k=20κk=20\kappak=20κ and asserts, for every s≥1s\ge1s≥1,

Ef(y(s+1))−f(x∗)≤0.9s(f(y(1))−f(x∗)).\mathbb E f(y^{(s+1)})-f(x^*) \le 0.9^s\bigl(f(y^{(1)})-f(x^*)\bigr).Ef(y(s+1))−f(x∗)≤0.9s(f(y(1))−f(x∗)).

The epoch length is a count, so the statement takes k∈Nk\in\mathbb Nk∈N and explicitly requires k=20β/αk=20\beta/\alphak=20β/α. The goal uses exactly the book's step size, epoch length, and contraction factor.

Milestones: second moments and a single epoch

Lemma 6.4 bounds Ei∥∇fi(x)−∇fi(x∗)∥22\mathbb E_i\|\nabla f_i(x)-\nabla f_i(x^*)\|_2^2Ei​∥∇fi​(x)−∇fi​(x∗)∥22​ by 2β(f(x)−f(x∗))2\beta(f(x)-f(x^*))2β(f(x)−f(x∗)). Equation (6.3) bounds the second moment of the corrected sampled direction by the objective gaps at the current point and the anchor. Equation (6.2), the unbiased-direction display, and the one-step display express how that direction changes squared distance to x∗x^*x∗. The later display on p. 338 bounds one epoch for any positive step size with 2βη<12\beta\eta<12βη<1. Finally, equation (6.1) substitutes the stated constants to obtain the factor 0.90.90.9 for one epoch. These seven source claims form the milestone list in reading order.

Significance

The theorem gives an explicit accuracy guarantee after a specified number of epochs: an initial gap DDD falls below 0.9sD0.9^sD0.9sD in expectation. Because each epoch uses a full gradient at its anchor as well as sampled component gradients, the result makes clear which quantity contracts and which operations are counted. It is a concrete linear-rate statement for a method whose individual stochastic gradients need not approach zero at the optimum. Bubeck, §6.3 discusses this issue when introducing the correction term.

The mathematical result is already proved in the monograph. The remaining task is to produce machine-checked proofs of its precise finite-sum model, the single-index estimates, the epoch inequality, and the full repeated-epoch guarantee. The mission drafts those statements and definitions; no proof is claimed for the open theorem items. The finite uniform-average representation and the separation between a conditional one-step average and the full multi-epoch average can be reused in other finite-sum stochastic algorithms.

Difficulty

The sampled component gradient ∇fit(xt)\nabla f_{i_t}(x_t)∇fit​​(xt​) need not be small when xtx_txt​ is near x∗x^*x∗, so a bound using only its norm does not yield the desired fixed-step contraction. The correction −∇fit(y)+G(y)-\nabla f_{i_t}(y)+G(y)−∇fit​​(y)+G(y) has mean zero relative to the full gradient at the current iterate, but its second moment still depends on both xtx_txt​ and yyy. The proof must control those two gaps while respecting the fact that xtx_txt​ depends on earlier samples. A single-index estimate with xtx_txt​ held fixed and an expectation over complete sample histories are different statements; confusing them would make the goal weaker or false.

Formalization scope

The carrier is EuclideanSpace ℝ (Fin n) with its usual inner product and norm. The Fin m components and every sample array are finite. A real-valued uniform average is an ordinary finite sum divided by the number of arrays, and m≥1m\ge1m≥1 and k≥1k\ge1k≥1 prevent an empty average. Independent uniform sampling is represented by averaging over every function from step positions to component indices. The multi-epoch sample space has one such block for every epoch. There are no integrals or measurability side conditions.

The component assumptions include differentiability with an explicit gradient map, convexity on all of Rn\mathbb R^nRn, and the book's gradient-Lipschitz version of smoothness. Strong convexity is imposed on the average objective alone, using the published OnlineConvexOpt.ConvexBasics.StronglyConvexOn definition on the whole space. The book's standing notation assumes a minimizing x∗x^*x∗ exists; this is explicit. Positivity of α\alphaα and β\betaβ, and integrality of 20β/α20\beta/\alpha20β/α, make the displayed divisions and epoch length meaningful. The general epoch bound also requires 0<η0<\eta0<η and 2βη<12\beta\eta<12βη<1. Dimension zero is allowed: the theorem remains a statement about the unique point of R0\mathbb R^0R0 and its zero objective gap.

The direction always contains the sampled difference ∇fit(xt)−∇fit(y)\nabla f_{i_t}(x_t)-\nabla f_{i_t}(y)∇fit​​(xt​)−∇fit​​(y) and the full anchor gradient G(y)G(y)G(y). Replacing that direction with G(xt)G(x_t)G(xt​) would define gradient descent and would not satisfy this mission's algorithm. Contributions are welcome for the finite averaging identities, the component-gradient estimate, the conditional one-step calculation, the epoch inequality, and the induction across epochs.

Selected references

  • Sébastien Bubeck, Convex Optimization: Algorithms and Complexity, Foundations and Trends in Machine Learning 8(3–4), 2015, pp. 231–358. arXiv:1405.4980v2
  • Rie Johnson and Tong Zhang, Accelerating Stochastic Gradient Descent using Predictive Variance Reduction, Advances in Neural Information Processing Systems 26 (NIPS), 2013 (the origin of SVRG, cited by Bubeck on p. 335). https://proceedings.neurips.cc/paper/2013/hash/ac1dd209cbcc5e5d1c6e28598e8cbbe8-Abstract.html
10 thms1 active userReviewed
Convex OptimizationMachine LearningProbability·Captain: mikedeng1

Convex Optimization: Algorithms and Complexity XIV: Stochastic Mirror Descent on a β-Smooth Function with Noise σ Has Rate Rσ√(2/t) + βR²/tTextbook

Motivation

Many optimization problems in statistics and machine learning ask to minimize an expected loss f(x)=Eξ ℓ(x,ξ)f(x)=\mathbb E_\xi\,\ell(x,\xi)f(x)=Eξ​ℓ(x,ξ), or an average f(x)=1m∑i=1mfi(x)f(x)=\frac1m\sum_{i=1}^m f_i(x)f(x)=m1​∑i=1m​fi​(x) over a large data set. Exact gradients of such an fff are unavailable or too expensive, but unbiased random estimates are cheap: the gradient of the loss at one sample, or of one randomly chosen summand. The observation that first-order methods still make progress when the gradients are only correct on average goes back to Robbins and Monro (1951) and underlies stochastic gradient descent.

Chapter 6 of S. Bubeck, Convex Optimization: Algorithms and Complexity (2015), studies this setting through stochastic mirror descent (S-MD). Its Section 6.1 shows that in the non-smooth case a noisy oracle costs nothing in rate. Section 6.2 asks what smoothness buys: for a general stochastic oracle it cannot buy acceleration, but Theorem 6.3, whose proof the book takes from Dekel, Gilad-Bachrach, Shamir and Xiao (2012), shows that the rate splits into a noise term of order 1/t1/\sqrt t1/t​ and a smoothness term of order 1/t1/t1/t. The book uses it to justify mini-batch SGD. This mission is the fourteenth of a series that formalizes the section capstones of the book.

Setting

Let EEE be a finite-dimensional real vector space with an arbitrary norm ∥⋅∥\|\cdot\|∥⋅∥. Gradients are linear forms ggg on EEE, the value of ggg at vvv is written g⊤vg^\top vg⊤v, and the dual norm is ∥g∥∗=sup⁡∥v∥≤1g⊤v\|g\|_*=\sup_{\|v\|\le1}g^\top v∥g∥∗​=sup∥v∥≤1​g⊤v. Let X⊆E\mathcal X\subseteq EX⊆E be compact and convex.

A mirror map is a function Φ\PhiΦ on an open convex set D\mathcal DD with X⊆D‾\mathcal X\subseteq\overline{\mathcal D}X⊆D and X∩D≠∅\mathcal X\cap\mathcal D\ne\emptysetX∩D=∅. It is strictly convex and differentiable on D\mathcal DD, its gradient ∇Φ\nabla\Phi∇Φ takes every value, and ∥∇Φ(x)∥∗→∞\|\nabla\Phi(x)\|_*\to\infty∥∇Φ(x)∥∗​→∞ as xxx approaches the boundary of D\mathcal DD. Its Bregman divergence is DΦ(x,y)=Φ(x)−Φ(y)−∇Φ(y)⊤(x−y)D_\Phi(x,y)=\Phi(x)-\Phi(y)-\nabla\Phi(y)^\top(x-y)DΦ​(x,y)=Φ(x)−Φ(y)−∇Φ(y)⊤(x−y). The map is 1-strongly convex on X∩D\mathcal X\cap\mathcal DX∩D if DΦ(y,x)≥12∥x−y∥2D_\Phi(y,x)\ge\frac12\|x-y\|^2DΦ​(y,x)≥21​∥x−y∥2 there. A function fff is β\betaβ-smooth on X\mathcal XX if ∥∇f(x)−∇f(y)∥∗≤β∥x−y∥\|\nabla f(x)-\nabla f(y)\|_*\le\beta\|x-y\|∥∇f(x)−∇f(y)∥∗​≤β∥x−y∥ for x,y∈Xx,y\in\mathcal Xx,y∈X.

A stochastic oracle returns, at a query point xxx, a random linear form g~(x)\tilde g(x)g~​(x). When the query point is itself random, the book requires the conditional expectation given the query point, E(g~(x)∣x)\mathbb E(\tilde g(x)\mid x)E(g~​(x)∣x), to be a subgradient of fff at xxx. In the smooth case it requires E(g~(x)∣x)=∇f(x)\mathbb E(\tilde g(x)\mid x)=\nabla f(x)E(g~​(x)∣x)=∇f(x) together with the variance bound E(∥g~(x)−∇f(x)∥∗2∣x)≤σ2\mathbb E(\|\tilde g(x)-\nabla f(x)\|_*^2\mid x)\le\sigma^2E(∥g~​(x)−∇f(x)∥∗2​∣x)≤σ2.

S-MD with step γ\gammaγ starts at x1∈argmin⁡X∩DΦx_1\in\operatorname{argmin}_{\mathcal X\cap\mathcal D}\Phix1​∈argminX∩D​Φ and, writing g~s=g~(xs)\tilde g_s=\tilde g(x_s)g~​s​=g~​(xs​), iterates

xs+1∈argmin⁡x∈X∩D γ g~s⊤x+DΦ(x,xs).x_{s+1}\in\operatorname*{argmin}_{x\in\mathcal X\cap\mathcal D}\ \gamma\,\tilde g_s^\top x+D_\Phi(x,x_s).xs+1​∈x∈X∩Dargmin​ γg~​s⊤​x+DΦ​(x,xs​).

Let R2≥sup⁡x∈X∩DΦ(x)−Φ(x1)R^2\ge\sup_{x\in\mathcal X\cap\mathcal D}\Phi(x)-\Phi(x_1)R2≥supx∈X∩D​Φ(x)−Φ(x1​), and let x∗x^*x∗ minimize fff on X\mathcal XX.

Formalization targets

Goal: Theorem 6.3

Let fff be convex and β\betaβ-smooth, and let the oracle have variance at most σ2\sigma^2σ2. Then for every t≥1t\ge1t≥1, S-MD with step 1/(β+1/η)1/(\beta+1/\eta)1/(β+1/η) and η=Rσ2/t\eta=\frac R\sigma\sqrt{2/t}η=σR​2/t​ satisfies

E f(1t∑s=1txs+1)−f(x∗)≤Rσ2t+βR2t.\mathbb E\,f\Big(\frac1t\sum_{s=1}^t x_{s+1}\Big)-f(x^*)\le R\sigma\sqrt{\frac2t}+\frac{\beta R^2}{t}.Ef(t1​s=1∑t​xs+1​)−f(x∗)≤Rσt2​​+tβR2​.

Milestones (the proof's four displays)

For points xs,xs+1∈X∩Dx_s,x_{s+1}\in\mathcal X\cap\mathcal Dxs​,xs+1​∈X∩D and η>0\eta>0η>0, the smoothness step is

f(xs+1)−f(xs)≤g~s⊤(xs+1−xs)+η2∥∇f(xs)−g~s∥∗2+(β+1/η)DΦ(xs+1,xs).f(x_{s+1})-f(x_s)\le\tilde g_s^\top(x_{s+1}-x_s)+\tfrac\eta2\|\nabla f(x_s)-\tilde g_s\|_*^2+(\beta+1/\eta)D_\Phi(x_{s+1},x_s).f(xs+1​)−f(xs​)≤g~​s⊤​(xs+1​−xs​)+2η​∥∇f(xs​)−g~​s​∥∗2​+(β+1/η)DΦ​(xs+1​,xs​).

If xs+1x_{s+1}xs+1​ is the S-MD step, the mirror step is

1β+1/ηg~s⊤(xs+1−x∗)≤DΦ(x∗,xs)−DΦ(x∗,xs+1)−DΦ(xs+1,xs).\tfrac{1}{\beta+1/\eta}\tilde g_s^\top(x_{s+1}-x^*)\le D_\Phi(x^*,x_s)-D_\Phi(x^*,x_{s+1})-D_\Phi(x_{s+1},x_s).β+1/η1​g~​s⊤​(xs+1​−x∗)≤DΦ​(x∗,xs​)−DΦ​(x∗,xs+1​)−DΦ​(xs+1​,xs​).

Combining the two gives a pathwise bound on f(xs+1)f(x_{s+1})f(xs+1​) with the cross term (g~s−∇f(xs))⊤(x∗−xs)(\tilde g_s-\nabla f(x_s))^\top(x^*-x_s)(g~​s​−∇f(xs​))⊤(x∗−xs​). Taking expectations gives the expected one-step bound

Ef(xs+1)−f(x∗)≤(β+1/η) E(DΦ(x∗,xs)−DΦ(x∗,xs+1))+ησ22.\mathbb Ef(x_{s+1})-f(x^*)\le(\beta+1/\eta)\,\mathbb E\big(D_\Phi(x^*,x_s)-D_\Phi(x^*,x_{s+1})\big)+\frac{\eta\sigma^2}{2}.Ef(xs+1​)−f(x∗)≤(β+1/η)E(DΦ​(x∗,xs​)−DΦ​(x∗,xs+1​))+2ησ2​.

Companion: Theorem 6.1 and (4.10)

For a convex fff with E(∥g~(x)∥∗2∣x)≤B2\mathbb E(\|\tilde g(x)\|_*^2\mid x)\le B^2E(∥g~​(x)∥∗2​∣x)≤B2, S-MD with η=RB2/t\eta=\frac RB\sqrt{2/t}η=BR​2/t​ satisfies

E f(1t∑s=1txs)−min⁡Xf≤RB2/t.\mathbb E\,f\Big(\frac1t\sum_{s=1}^tx_s\Big)-\min_{\mathcal X}f\le RB\sqrt{2/t}.Ef(t1​s=1∑t​xs​)−Xmin​f≤RB2/t​.

This rests on the deterministic regret bound (4.10) of mirror descent along arbitrary vectors gsg_sgs​:

∑s≤tgs⊤(xs−x)≤R2η+η2ρ∑s≤t∥gs∥∗2.\sum_{s\le t}g_s^\top(x_s-x)\le\frac{R^2}{\eta}+\frac{\eta}{2\rho}\sum_{s\le t}\|g_s\|_*^2.s≤t∑​gs⊤​(xs​−x)≤ηR2​+2ρη​s≤t∑​∥gs​∥∗2​.

Significance

Theorem 6.3 says exactly how much smoothness helps under noise. As σ→0\sigma\to0σ→0 it recovers the βR2/t\beta R^2/tβR2/t rate of deterministic smooth optimization. For large ttt the noise term Rσ2/tR\sigma\sqrt{2/t}Rσ2/t​ dominates; the book notes, citing Tsybakov (2003), that smoothness brings no acceleration for a general stochastic oracle. Averaging mmm independent oracle answers divides the variance by mmm, so the theorem quantifies the benefit of mini-batches: the noise term shrinks by m\sqrt mm​ while the smoothness term is unchanged. Theorem 6.1 is the matching non-smooth statement and the template for stochastic subgradient methods in any norm.

These are classical, proved results. None of them is known to be formalized in Lean, and the platform has no stochastic mirror descent statement. Its stochastic gradient items cover the Euclidean strongly convex case and the non-convex gradient-norm case. This mission adds a reusable stochastic-oracle layer in an arbitrary norm, with conditional expectations given random query points, on top of the mirror-map layer of Chapter 4.

Difficulty

The deterministic steps are short manipulations of Bregman divergences. The difficulty is in the passage to expectations. The query point xsx_sxs​ is random, so unbiasedness enters only through the conditional expectation given xsx_sxs​. Making the cross term vanish requires pulling the σ(xs)\sigma(x_s)σ(xs​)-measurable vector x∗−xsx^*-x_sx∗−xs​ out of a conditional expectation of a dual-valued random variable. Every expectation also has to exist. When ∇Φ\nabla\Phi∇Φ blows up at the boundary of D\mathcal DD, the Bregman terms DΦ(x∗,xs)D_\Phi(x^*,x_s)DΦ​(x∗,xs​) are not bounded a priori, and their integrability has to be derived from the recursion. A further obstacle is that the minimizer x∗x^*x∗ may lie on the boundary of D\mathcal DD, where Φ\PhiΦ is not part of the book's data. Treating E\mathbb EE informally, or assuming x∗∈Dx^*\in\mathcal Dx∗∈D, skips exactly these points.

Formalization scope

  • Spaces and gradients. EEE is a finite-dimensional real normed space. Gradients are explicit maps Φ' f' : E → (E →L[ℝ] ℝ), g⊤vg^\top vg⊤v is g v, and ∥⋅∥∗\|\cdot\|_*∥⋅∥∗​ is the operator norm. β\betaβ-smoothness is stated with derivatives relative to X\mathcal XX. Φ\PhiΦ is a total function, constrained only by the mirror-map axioms on D\mathcal DD.
  • Runs and oracle. S-MD is a run predicate. For every outcome, x1x_1x1​ minimizes Φ\PhiΦ on X∩D\mathcal X\cap\mathcal DX∩D, and xs+1x_{s+1}xs+1​ is some minimizer of the step objective. The oracle is a predicate on the random sequences (xs,g~s)(x_s,\tilde g_s)(xs​,g~​s​): each xsx_sxs​ is measurable, and the conditional expectations are taken given σ(xs)\sigma(x_s)σ(xs​). Every conditioned quantity is integrable.
  • Conclusions. Every bound on an expectation also asserts integrability. Without it, the Lean integral of a non-integrable function is 000 and the bound could hold trivially.
  • Standing assumptions. The book's R2=sup⁡(Φ−Φ(x1))R^2=\sup(\Phi-\Phi(x_1))R2=sup(Φ−Φ(x1​)) is replaced by any upper bound R2R^2R2. The minimizer x∗∈Xx^*\in\mathcal Xx∗∈X exists (p. 242). X\mathcal XX is compact and convex (Chapter 4), and convex functions are closed (p. 236).
  • Positivity side conditions. R,σ,B>0R,\sigma,B>0R,σ,B>0 and t≥1t\ge1t≥1 make the step sizes and bounds defined, and β≥0\beta\ge0β≥0.

A variance hypothesis stated only at deterministic points would not control the random iterates, and is not used. Run predicates that let xs+1x_{s+1}xs+1​ be an arbitrary point of X∩D\mathcal X\cap\mathcal DX∩D would make the theorems false, and are not used either.

A complete development needs: first-order optimality over a convex set, the three-point identity of Bregman divergences, the descent lemma in an arbitrary norm, and continuity of the gradient of a differentiable convex function. On the probability side it needs pull-out and conditional Jensen properties for dual-valued conditional expectations. The probability layer is reusable for every stochastic first-order method in the book, including SVRG and random coordinate descent. Proofs of the milestones are welcome, and so are general lemmas about conditional expectations of continuous-linear-map-valued random variables.

Selected references

  • S. Bubeck, Convex Optimization: Algorithms and Complexity, Foundations and Trends in Machine Learning 8(3–4):231–358, 2015. arXiv:1405.4980v2, https://arxiv.org/abs/1405.4980 (Chapter 6, pp. 329–333; Chapter 4, pp. 297–307).
  • O. Dekel, R. Gilad-Bachrach, O. Shamir, L. Xiao, Optimal distributed online prediction using mini-batches, Journal of Machine Learning Research 13:165–202, 2012. https://jmlr.org/papers/v13/dekel12a.html
  • H. Robbins, S. Monro, A stochastic approximation method, Annals of Mathematical Statistics 22(3):400–407, 1951. https://doi.org/10.1214/aoms/1177729586
  • A. Beck, M. Teboulle, Mirror descent and nonlinear projected subgradient methods for convex optimization, Operations Research Letters 31(3):167–175, 2003. https://doi.org/10.1016/S0167-6377(02)00231-6
  • A. Nemirovski, A. Juditsky, G. Lan, A. Shapiro, Robust stochastic approximation approach to stochastic programming, SIAM Journal on Optimization 19(4):1574–1609, 2009. https://doi.org/10.1137/070704277
6 thms1 active userReviewed
Convex OptimizationNumerical Analysis·Captain: mikedeng1

Convex Optimization: Algorithms and Complexity XIII: Newton's Method Converges Quadratically, ‖x_{k+1} − x*‖ ≤ (M/μ)‖x_k − x*‖², from ‖x₀ − x*‖ ≤ μ/(2M)Textbook

Motivation

Newton's method is the basic second-order method of continuous optimization: at the current point it replaces the objective by its second-order Taylor model and jumps to the stationary point of that model. Its defining property is speed near a nondegenerate minimum, where the error is squared at every step, so that the number of correct digits roughly doubles per iteration. This local behaviour is what makes Newton's method the inner engine of interior point methods, the polynomial-time algorithms for linear, conic and general convex programming (Nesterov and Nemirovski, 1994). In S. Bubeck's monograph Convex Optimization: Algorithms and Complexity (Foundations and Trends in Machine Learning, 2015; arXiv:1405.4980v2), §5.3.2 recalls the traditional local analysis of Newton's method, Theorem 5.3, before turning to the affine-invariant self-concordance analysis used for interior point methods. This mission formalizes that theorem and the four steps of its proof.

Setting

Let Rn\mathbb R^nRn carry the Euclidean norm ∥⋅∥\|\cdot\|∥⋅∥, and write ∥A∥\|A\|∥A∥ for the operator norm of a linear map A:Rn→RnA:\mathbb R^n\to\mathbb R^nA:Rn→Rn, so that ∥Ax∥≤∥A∥ ∥x∥\|Ax\|\le\|A\|\,\|x\|∥Ax∥≤∥A∥∥x∥. Let f:Rn→Rf:\mathbb R^n\to\mathbb Rf:Rn→R be a C2C^2C2 function, with gradient ∇f(x)∈Rn\nabla f(x)\in\mathbb R^n∇f(x)∈Rn and Hessian ∇2f(x)\nabla^2 f(x)∇2f(x), a linear map Rn→Rn\mathbb R^n\to\mathbb R^nRn→Rn (the derivative of the gradient map). For a real number ccc, A⪰cInA\succeq cI_nA⪰cIn​ means ⟨Av,v⟩≥c∥v∥2\langle Av,v\rangle\ge c\|v\|^2⟨Av,v⟩≥c∥v∥2 for all v∈Rnv\in\mathbb R^nv∈Rn.

The Hessian is MMM-Lipschitz if ∥∇2f(x)−∇2f(y)∥≤M∥x−y∥\|\nabla^2 f(x)-\nabla^2 f(y)\|\le M\|x-y\|∥∇2f(x)−∇2f(y)∥≤M∥x−y∥ for all x,y∈Rnx,y\in\mathbb R^nx,y∈Rn.

Newton's method starts at x0∈Rnx_0\in\mathbb R^nx0​∈Rn and iterates, for k≥0k\ge0k≥0,

xk+1=xk−[∇2f(xk)]−1∇f(xk).x_{k+1}=x_k-[\nabla^2 f(x_k)]^{-1}\nabla f(x_k).xk+1​=xk​−[∇2f(xk​)]−1∇f(xk​).

A point x∗x^*x∗ is a local minimum of fff if f(x∗)≤f(x)f(x^*)\le f(x)f(x∗)≤f(x) for all xxx in a neighbourhood of x∗x^*x∗; it has strictly positive Hessian if ∇2f(x∗)⪰μIn\nabla^2 f(x^*)\succeq\mu I_n∇2f(x∗)⪰μIn​ for some μ>0\mu>0μ>0.

Formalization targets

Goal: Theorem 5.3 (p. 320)

Assume the Hessian of fff is MMM-Lipschitz, M>0M>0M>0, and x∗x^*x∗ is a local minimum with ∇2f(x∗)⪰μIn\nabla^2 f(x^*)\succeq\mu I_n∇2f(x∗)⪰μIn​, μ>0\mu>0μ>0. If ∥x0−x∗∥≤μ/(2M)\|x_0-x^*\|\le\mu/(2M)∥x0​−x∗∥≤μ/(2M), then Newton's method from x0x_0x0​ is well defined (every Hessian along the iterates is invertible, so the sequence exists and is unique) and

∥xk+1−x∗∥≤Mμ ∥xk−x∗∥2(k≥0),xk→x∗.\|x_{k+1}-x^*\|\le\frac M\mu\,\|x_k-x^*\|^2\quad(k\ge0),\qquad x_k\to x^*.∥xk+1​−x∗∥≤μM​∥xk​−x∗∥2(k≥0),xk​→x∗.

Milestones (p. 321, the steps of the proof)

  1. The integral formula ∫01∇2f(x+sh) h ds=∇f(x+h)−∇f(x)\int_0^1\nabla^2 f(x+sh)\,h\,ds=\nabla f(x+h)-\nabla f(x)∫01​∇2f(x+sh)hds=∇f(x+h)−∇f(x).
  2. The error representation of one Newton step, xk+1−x∗=[∇2f(xk)]−1∫01[∇2f(xk)−∇2f(x∗+s(xk−x∗))](xk−x∗) dsx_{k+1}-x^*=[\nabla^2 f(x_k)]^{-1}\int_0^1[\nabla^2 f(x_k)-\nabla^2 f(x^*+s(x_k-x^*))](x_k-x^*)\,dsxk+1​−x∗=[∇2f(xk​)]−1∫01​[∇2f(xk​)−∇2f(x∗+s(xk​−x∗))](xk​−x∗)ds.
  3. The Lipschitz bound ∫01∥∇2f(xk)−∇2f(x∗+s(xk−x∗))∥ ds≤M2∥xk−x∗∥\int_0^1\|\nabla^2 f(x_k)-\nabla^2 f(x^*+s(x_k-x^*))\|\,ds\le\frac M2\|x_k-x^*\|∫01​∥∇2f(xk​)−∇2f(x∗+s(xk​−x∗))∥ds≤2M​∥xk​−x∗∥.
  4. The Hessian lower bound ∇2f(xk)⪰(μ−M∥xk−x∗∥)In⪰μ2In\nabla^2 f(x_k)\succeq(\mu-M\|x_k-x^*\|)I_n\succeq\frac\mu2I_n∇2f(xk​)⪰(μ−M∥xk​−x∗∥)In​⪰2μ​In​ when ∥xk−x∗∥≤μ/(2M)\|x_k-x^*\|\le\mu/(2M)∥xk​−x∗∥≤μ/(2M).

Significance

The theorem gives a quantitative basin of quadratic convergence: an explicit radius μ/(2M)\mu/(2M)μ/(2M), depending only on the curvature at the minimum and the Lipschitz constant of the Hessian, inside which Newton's method needs only O(log⁡log⁡(1/ε))O(\log\log(1/\varepsilon))O(loglog(1/ε)) iterations to reach accuracy ε\varepsilonε. It is the classical statement whose shortcomings (dependence on a choice of norm, constants that change under linear changes of variables) motivate the self-concordance theory of the following subsections, and it is the local convergence result invoked whenever a damped or globalized Newton scheme is shown to enter its quadratic phase.

On the formal side, Mathlib has the calculus this needs (Fréchet derivatives, interval integrals of vector-valued maps, operator norms) but no convergence theorem for multivariate Newton's method for minimization. A formal proof produces reusable pieces: the integral form of the mean value theorem for gradients, the stability of a positive-definite lower bound under Lipschitz perturbations, and an inverse-operator norm bound from a quadratic-form lower bound. The result itself is classical and fully proved in the literature; what is open here is its machine-checked proof in this form.

Difficulty

The individual inequalities are short, but the argument is an induction in which well-definedness and the rate are proved together: the Hessian at xkx_kxk​ is invertible only because xkx_kxk​ is still in the ball of radius μ/(2M)\mu/(2M)μ/(2M), and xk+1x_{k+1}xk+1​ stays in that ball only because of the rate. A proof that first assumes the sequence exists and then bounds it is circular. The proof also passes between two kinds of control on the Hessian, a lower bound on its quadratic form and an operator-norm bound on its inverse, and the second is only meaningful once invertibility is established. Finally, the integral manipulations need integrability of the maps s↦∇2f(x∗+s(xk−x∗))(xk−x∗)s\mapsto\nabla^2 f(x^*+s(x_k-x^*))(x_k-x^*)s↦∇2f(x∗+s(xk​−x∗))(xk​−x∗), which comes from the continuity of the Hessian.

Formalization scope

Rn\mathbb R^nRn is EuclideanSpace ℝ (Fin n). The gradient and Hessian are explicit maps g:Rn→Rng:\mathbb R^n\to\mathbb R^ng:Rn→Rn and H:Rn→(Rn→LRn)H:\mathbb R^n\to(\mathbb R^n\to_L\mathbb R^n)H:Rn→(Rn→L​Rn) with ContDiff ℝ 2 f, HasGradientAt f (g x) x and HasFDerivAt g (H x) x at every point; the norm on H(x)H(x)H(x) is Mathlib's operator norm, as on the page. A⪰cInA\succeq cI_nA⪰cIn​ is the quadratic-form inequality. A Newton run is a sequence x:N→Rnx:\mathbb N\to\mathbb R^nx:N→Rn indexed from 000 satisfying the linear system ∇2f(xk)(xk−xk+1)=∇f(xk)\nabla^2 f(x_k)(x_k-x_{k+1})=\nabla f(x_k)∇2f(xk​)(xk​−xk+1​)=∇f(xk​); no inverse of a possibly singular operator appears in any hypothesis, and "well defined" is a conclusion: a unique run exists from x0x_0x0​ and every Hessian along it is bijective. The rate and xk→x∗x_k\to x^*xk​→x∗ are asserted for every run. The error representation is stated with both sides multiplied by ∇2f(xk)\nabla^2 f(x_k)∇2f(xk​), which is equivalent to the printed form once the Hessian is invertible. Milestones 3 and 4 use only the Lipschitz property and are stated for any Lipschitz map HHH.

Added hypothesis: M>0M>0M>0 (the radius μ/(2M)\mu/(2M)μ/(2M) divides by MMM; with M=0M=0M=0, Lean's convention μ/0=0\mu/0=0μ/0=0 would collapse the hypothesis to x0=x∗x_0=x^*x0​=x∗). Convexity of fff is not assumed, as on the page; x∗x^*x∗ is a local minimum and ∇f(x∗)=0\nabla f(x^*)=0∇f(x∗)=0 is derived, not assumed. Encoding the Newton step with Lean's inverse (which returns 000 on singular maps), or replacing ∇2f(x∗)⪰μIn\nabla^2 f(x^*)\succeq\mu I_n∇2f(x∗)⪰μIn​ by mere invertibility, would change the theorem and is ruled out.

A complete development needs the fundamental theorem of calculus for C1C^1C1 vector-valued maps along segments, Hessian-based quadratic-form estimates, and operator-norm bounds for inverses; all are reusable for the analysis of damped Newton, cubic regularization and interior point methods. Proofs of the milestones independently of the goal are welcome.

Selected references

  • S. Bubeck, Convex Optimization: Algorithms and Complexity, Foundations and Trends in Machine Learning 8(3–4):231–358, 2015. arXiv:1405.4980v2, §5.3.2, Theorem 5.3, pp. 320–321.
  • Yu. Nesterov and A. Nemirovski, Interior-Point Polynomial Algorithms in Convex Programming, SIAM Studies in Applied Mathematics 13, 1994. doi:10.1137/1.9781611970791
  • Yu. Nesterov, Introductory Lectures on Convex Optimization: A Basic Course, Kluwer, 2004, Theorem 1.2.5. doi:10.1007/978-1-4419-8853-9
6 thms1 active userReviewed
Convex OptimizationMachine Learning·Captain: mikedeng1

Convex Optimization: Algorithms and Complexity X: Nesterov's Accelerated Gradient Descent on a β-Smooth α-Strongly Convex Function Has Rate ((α + β)/2)‖x₁ − x*‖² exp(−(t − 1)/√κ)Textbook

Why accelerated rates matter

First-order methods, which query only function values and gradients, are the workhorse of large-scale optimization in machine learning, signal processing and operations research, because each step costs little more than one gradient evaluation. For a function that is both strongly convex and smooth, plain gradient descent converges geometrically, but the number of steps needed to reach accuracy ε\varepsilonε scales with the condition number κ\kappaκ of the problem. In 1983 Nesterov showed that a gradient method with a carefully chosen momentum term needs a number of steps proportional to κ\sqrt\kappaκ​ instead, and that this is optimal for black-box first-order methods. On ill-conditioned problems, where κ\kappaκ is in the thousands or millions, the difference between κ\kappaκ and κ\sqrt\kappaκ​ is the difference between practical and impractical.

This mission is the tenth of a series formalizing S. Bubeck's monograph Convex Optimization: Algorithms and Complexity (Foundations and Trends in Machine Learning, 2015; arXiv:1405.4980v2). It covers §3.7.1, the smooth and strongly convex case of Nesterov's accelerated gradient descent, and its main result, Theorem 3.18.

Setting

Work in Rn\mathbb R^nRn with the Euclidean inner product x⊤yx^\top yx⊤y and norm ∥⋅∥\|\cdot\|∥⋅∥. Let f:Rn→Rf:\mathbb R^n\to\mathbb Rf:Rn→R be differentiable with gradient ∇f\nabla f∇f.

  • fff is β\betaβ-smooth if its gradient is β\betaβ-Lipschitz: ∥∇f(x)−∇f(y)∥≤β∥x−y∥\|\nabla f(x)-\nabla f(y)\|\le\beta\|x-y\|∥∇f(x)−∇f(y)∥≤β∥x−y∥ for all x,yx,yx,y.
  • fff is α\alphaα-strongly convex (α>0\alpha>0α>0) if for all x,yx,yx,y
f(y)≥f(x)+∇f(x)⊤(y−x)+α2∥y−x∥2.f(y)\ge f(x)+\nabla f(x)^\top(y-x)+\frac\alpha2\|y-x\|^2 .f(y)≥f(x)+∇f(x)⊤(y−x)+2α​∥y−x∥2.
  • The condition number is κ=β/α\kappa=\beta/\alphaκ=β/α; for n≥1n\ge1n≥1 one always has κ≥1\kappa\ge1κ≥1.
  • x∗x^*x∗ denotes a minimizer of fff on Rn\mathbb R^nRn.

Nesterov's accelerated gradient descent starts at an arbitrary point x1=y1x_1=y_1x1​=y1​ and iterates, for t≥1t\ge1t≥1,

yt+1=xt−1β∇f(xt),xt+1=(1+κ−1κ+1)yt+1−κ−1κ+1 yt.y_{t+1}=x_t-\frac1\beta\nabla f(x_t),\qquad x_{t+1}=\Big(1+\frac{\sqrt\kappa-1}{\sqrt\kappa+1}\Big)y_{t+1}-\frac{\sqrt\kappa-1}{\sqrt\kappa+1}\,y_t .yt+1​=xt​−β1​∇f(xt​),xt+1​=(1+κ​+1κ​−1​)yt+1​−κ​+1κ​−1​yt​.

The point yt+1y_{t+1}yt+1​ is a gradient step from xtx_txt​, and xt+1x_{t+1}xt+1​ moves beyond yt+1y_{t+1}yt+1​ in the direction yt+1−yty_{t+1}-y_tyt+1​−yt​ by the fixed momentum factor (κ−1)/(κ+1)(\sqrt\kappa-1)/(\sqrt\kappa+1)(κ​−1)/(κ​+1).

The analysis in the book uses auxiliary quadratic functions Φs\Phi_sΦs​ (an estimate sequence), defined from the points xsx_sxs​ by

Φ1(x)=f(x1)+α2∥x−x1∥2,Φs+1(x)=(1−1κ)Φs(x)+1κ(f(xs)+∇f(xs)⊤(x−xs)+α2∥x−xs∥2),\Phi_1(x)=f(x_1)+\frac\alpha2\|x-x_1\|^2,\qquad \Phi_{s+1}(x)=\Big(1-\frac1{\sqrt\kappa}\Big)\Phi_s(x)+\frac1{\sqrt\kappa}\Big(f(x_s)+\nabla f(x_s)^\top(x-x_s)+\frac\alpha2\|x-x_s\|^2\Big),Φ1​(x)=f(x1​)+2α​∥x−x1​∥2,Φs+1​(x)=(1−κ​1​)Φs​(x)+κ​1​(f(xs​)+∇f(xs​)⊤(x−xs​)+2α​∥x−xs​∥2),

together with their centres vsv_svs​ (with v1=x1v_1=x_1v1​=x1​ and the recursion (3.21) of the book) and their minimum values Φs∗\Phi^*_sΦs∗​.

Formalization targets

Goal: Theorem 3.18

For every run of the method and every t≥1t\ge1t≥1,

f(yt)−f(x∗)≤α+β2 ∥x1−x∗∥2exp⁡(−t−1κ).f(y_t)-f(x^*)\le\frac{\alpha+\beta}2\,\|x_1-x^*\|^2\exp\Big(-\frac{t-1}{\sqrt\kappa}\Big).f(yt​)−f(x∗)≤2α+β​∥x1​−x∗∥2exp(−κ​t−1​).

Milestones, from the book's proof

  1. (3.18): Φs+1(x)≤f(x)+(1−1/κ)s(Φ1(x)−f(x))\Phi_{s+1}(x)\le f(x)+(1-1/\sqrt\kappa)^s(\Phi_1(x)-f(x))Φs+1​(x)≤f(x)+(1−1/κ​)s(Φ1​(x)−f(x)) for all xxx.
  2. (3.19): f(ys)≤min⁡x∈RnΦs(x)f(y_s)\le\min_{x\in\mathbb R^n}\Phi_s(x)f(ys​)≤minx∈Rn​Φs​(x).
  3. The geometric rate: f(yt)−f(x∗)≤α+β2∥x1−x∗∥2(1−1/κ)t−1f(y_t)-f(x^*)\le\frac{\alpha+\beta}2\|x_1-x^*\|^2(1-1/\sqrt\kappa)^{t-1}f(yt​)−f(x∗)≤2α+β​∥x1​−x∗∥2(1−1/κ​)t−1.
  4. The form Φs(x)=Φs∗+α2∥x−vs∥2\Phi_s(x)=\Phi^*_s+\frac\alpha2\|x-v_s\|^2Φs​(x)=Φs∗​+2α​∥x−vs​∥2 with vsv_svs​ given by (3.21).
  5. The identity (3.22) for Φs+1∗\Phi^*_{s+1}Φs+1∗​.
  6. The inequality (3.20), the inductive step of (3.19).
  7. The coupling vs−xs=κ (xs−ys)v_s-x_s=\sqrt\kappa\,(x_s-y_s)vs​−xs​=κ​(xs​−ys​).

The geometric form in milestone 3 is slightly stronger than the goal, which follows from 1−u≤e−u1-u\le e^{-u}1−u≤e−u.

Significance

Theorem 3.18 gives ε\varepsilonε-accuracy after O(κlog⁡(1/ε))O(\sqrt\kappa\log(1/\varepsilon))O(κ​log(1/ε)) gradient evaluations. Projected gradient descent with step 1/β1/\beta1/β on the same class contracts only at the rate exp⁡(−t/κ)\exp(-t/\kappa)exp(−t/κ) (Theorem 3.10 of the book). The lower bound of Theorem 3.15 shows that no black-box first-order method can do better than ((κ−1)/(κ+1))2(t−1)((\sqrt\kappa-1)/(\sqrt\kappa+1))^{2(t-1)}((κ​−1)/(κ​+1))2(t−1), so the accelerated rate is optimal up to constants. The estimate-sequence argument is the template for many later accelerated methods: proximal, stochastic and variance-reduced variants such as Katyusha, and accelerated coordinate descent.

The result is classical and fully proved on paper. No machine-checked proof of the accelerated rate for strongly convex smooth functions is known to exist in Lean's Mathlib. This mission produces one, with the estimate sequence Φs\Phi_sΦs​, its centres and its minimum values as reusable objects, and with every algebraic identity of the book's proof stated separately.

Difficulty

The algorithm is two lines, but its analysis is not a one-step contraction: neither ∥xt−x∗∥\|x_t-x^*\|∥xt​−x∗∥ nor f(yt)−f(x∗)f(y_t)-f(x^*)f(yt​)−f(x∗) decreases by the factor 1−1/κ1-1/\sqrt\kappa1−1/κ​ at every step. A Lyapunov argument for gradient descent, applied directly to yty_tyt​, gives only the rate 1−1/κ1-1/\kappa1−1/κ. The book obtains the rate through the auxiliary functions Φs\Phi_sΦs​. The inequality (3.18) is easy, but (3.19), that the minimum of Φs\Phi_sΦs​ never drops below f(ys)f(y_s)f(ys​), depends on the exact choice of the momentum factor. It holds only through the identity vs−xs=κ(xs−ys)v_s-x_s=\sqrt\kappa(x_s-y_s)vs​−xs​=κ​(xs​−ys​), which ties the centre of Φs\Phi_sΦs​ to the iterates. Formally, the obstacles are the bookkeeping of the recursive quadratics on Rn\mathbb R^nRn and the algebra in κ\sqrt\kappaκ​, 1/κ1/\sqrt\kappa1/κ​ and 1/(ακ)=κ/β1/(\alpha\sqrt\kappa)=\sqrt\kappa/\beta1/(ακ​)=κ​/β.

Formalization scope

  • Rn\mathbb R^nRn is EuclideanSpace ℝ (Fin n). The gradient is an explicit map g with HasGradientAt f (g x) x at every point, which is part of the smoothness predicate IsBetaSmooth f g β. Strong convexity is the published definition OnlineConvexOpt.ConvexBasics.StronglyConvexOn Set.univ f g α, which is the book's (3.13).
  • A run of the method is a predicate IsNesterovSCRun g α β x y on two sequences indexed from 111, with x1=y1x_1=y_1x1​=y1​ arbitrary. Every theorem holds for every run, that is, every starting point.
  • κ\kappaκ is kappa α β = β / α. All theorems assume α>0\alpha>0α>0 and β>0\beta>0β>0. The second is implied by the other hypotheses for n≥1n\ge1n≥1; no hypothesis α≤β\alpha\le\betaα≤β is added.
  • The existence of a minimizer x∗x^*x∗ is the book's standing assumption, written as a hypothesis.
  • Φs\Phi_sΦs​, vsv_svs​ and Φs∗=Φs(vs)\Phi^*_s=\Phi_s(v_s)Φs∗​=Φs​(vs​) are explicit recursive definitions. The book's Φs∗=min⁡Φs\Phi^*_s=\min\Phi_sΦs∗​=minΦs​ is recovered by milestone 4, and no real infimum is used. The minimum in (3.19) is stated as f(ys)≤Φs(x)f(y_s)\le\Phi_s(x)f(ys​)≤Φs​(x) for every xxx.
  • The identities of milestones 4, 5 and 7 are algebraic and are stated without convexity or smoothness, for arbitrary sequences or runs.
  • Ruled out as trivializing: a run predicate that drops x1=y1x_1=y_1x1​=y1​ breaks (3.19) at s=1s=1s=1 and is not used. A minimum value Φs∗\Phi^*_sΦs∗​ defined through (3.22) would make that identity a tautology, so Φs∗\Phi^*_sΦs∗​ is defined as a value of Φs\Phi_sΦs​.
  • The definitions are local to the namespace ConvexOptAlg.NesterovStrong. β-smoothness duplicates the predicate of other missions of the series and will be merged afterwards. Contributions welcome: proofs of the milestones, and general lemmas on quadratics z↦c+α2∥z−v∥2z\mapsto c+\frac\alpha2\|z-v\|^2z↦c+2α​∥z−v∥2 that the algebraic milestones need.

Selected references

  • S. Bubeck, Convex Optimization: Algorithms and Complexity, Foundations and Trends in Machine Learning 8(3–4):231–358, 2015. arXiv:1405.4980v2, §3.7.1, Theorem 3.18, pp. 290–293.
  • Y. Nesterov, A method of solving a convex programming problem with convergence rate O(1/k²), Soviet Mathematics Doklady 27:372–376, 1983.
  • Y. Nesterov, Introductory Lectures on Convex Optimization: A Basic Course, Kluwer, 2004. doi:10.1007/978-1-4419-8853-9
10 thms1 active userReviewed
Convex OptimizationMachine Learning·Captain: mikedeng1

Convex Optimization: Algorithms and Complexity IX: Nesterov's Accelerated Gradient Descent on a Convex β-Smooth Function Has Rate 2β‖x₁ − x*‖²/t²Textbook

Motivation

Minimizing a convex function f:Rn→Rf:\mathbb R^n\to\mathbb Rf:Rn→R whose gradient is Lipschitz is the basic problem of first-order optimization, and it is the regime in which most large-scale methods of machine learning, signal processing and operations research are analysed. The natural method, gradient descent, reaches accuracy ε\varepsilonε after O(1/ε)O(1/\varepsilon)O(1/ε) gradient evaluations. In 1983 Nesterov showed that a method using the same oracle, but combining the current and the previous iterate, reaches accuracy ε\varepsilonε after only O(1/ε)O(1/\sqrt\varepsilon)O(1/ε​) evaluations, and that no method using only gradient information can do better by more than a constant factor. This accelerated gradient descent and its proximal variants (FISTA) are now the default fast first-order methods for smooth and composite convex problems.

This mission formalizes the accelerated rate as presented in S. Bubeck, Convex Optimization: Algorithms and Complexity (Foundations and Trends in Machine Learning, 2015; arXiv:1405.4980v2), §3.7.2, Theorem 3.19, whose proof follows Beck and Teboulle (2009).

Timeline.

  • 1983: Nesterov introduces the accelerated method with rate O(1/t2)O(1/t^2)O(1/t2) for convex functions with Lipschitz gradient (Soviet Math. Dokl. 27).
  • 1983: Nemirovski and Yudin's black-box lower bounds show that Ω(1/t2)\Omega(1/t^2)Ω(1/t2) is the best possible rate for this class (Theorem 3.14 of the book gives 3β∥x1−x∗∥2/(32(t+1)2)3\beta\|x_1-x^*\|^2/(32(t+1)^2)3β∥x1​−x∗∥2/(32(t+1)2)).
  • 2009: Beck and Teboulle's FISTA extends the method, with the same step sequence λt\lambda_tλt​, to composite problems with a simple nonsmooth term (SIAM J. Imaging Sci. 2(1)).
  • 2008: Tseng gives a unified treatment with simpler step sizes (manuscript).

Setting

Let Rn\mathbb R^nRn carry the Euclidean inner product x⊤yx^\top yx⊤y and norm ∥⋅∥\|\cdot\|∥⋅∥. A differentiable function f:Rn→Rf:\mathbb R^n\to\mathbb Rf:Rn→R is β-smooth if its gradient is β\betaβ-Lipschitz:

∥∇f(x)−∇f(y)∥≤β∥x−y∥(x,y∈Rn).\|\nabla f(x)-\nabla f(y)\|\le\beta\|x-y\|\qquad(x,y\in\mathbb R^n).∥∇f(x)−∇f(y)∥≤β∥x−y∥(x,y∈Rn).

Assume fff is convex and β\betaβ-smooth with β>0\beta>0β>0, and let x∗x^*x∗ be a minimizer of fff.

Define the step sequences λ0=0\lambda_0=0λ0​=0,

λt=1+1+4λt−122(t≥1),γt=1−λtλt+1.\lambda_t=\frac{1+\sqrt{1+4\lambda_{t-1}^2}}{2}\quad(t\ge1),\qquad\gamma_t=\frac{1-\lambda_t}{\lambda_{t+1}}.λt​=21+1+4λt−12​​​(t≥1),γt​=λt+1​1−λt​​.

Then λ1=1\lambda_1=1λ1​=1, λt≥1\lambda_t\ge1λt​≥1 for t≥1t\ge1t≥1, and γt≤0\gamma_t\le0γt​≤0. Nesterov's accelerated gradient descent for the smooth case starts from an arbitrary point x1=y1x_1=y_1x1​=y1​ and sets, for t≥1t\ge1t≥1,

yt+1=xt−1β∇f(xt),xt+1=(1−γt) yt+1+γt yt.y_{t+1}=x_t-\frac1\beta\nabla f(x_t),\qquad x_{t+1}=(1-\gamma_t)\,y_{t+1}+\gamma_t\,y_t.yt+1​=xt​−β1​∇f(xt​),xt+1​=(1−γt​)yt+1​+γt​yt​.

The primary sequence (yt)(y_t)(yt​) consists of gradient steps of length 1/β1/\beta1/β; the sequence (xt)(x_t)(xt​), at which the gradient is queried, moves past yt+1y_{t+1}yt+1​ away from yty_tyt​ (since γt≤0\gamma_t\le0γt​≤0). Write δt=f(yt)−f(x∗)\delta_t=f(y_t)-f(x^*)δt​=f(yt​)−f(x∗) for the optimality gap.

Formalization targets

Goal: Theorem 3.19

For every run of the method and every t≥1t\ge1t≥1,

f(yt)−f(x∗)≤2β∥x1−x∗∥2t2.f(y_t)-f(x^*)\le\frac{2\beta\|x_1-x^*\|^2}{t^2}.f(yt​)−f(x∗)≤t22β∥x1​−x∗∥2​.

The constant 222 and the exponent 222 are those printed in the book.

Milestones

In the order of the book's proof (pp. 294–295):

  1. Lemma 3.6, unconstrained (p. 270): f(x−1β∇f(x))−f(y)≤∇f(x)⊤(x−y)−12β∥∇f(x)∥2f(x-\tfrac1\beta\nabla f(x))-f(y)\le\nabla f(x)^\top(x-y)-\tfrac1{2\beta}\|\nabla f(x)\|^2f(x−β1​∇f(x))−f(y)≤∇f(x)⊤(x−y)−2β1​∥∇f(x)∥2 for all x,yx,yx,y.
  2. (3.23): f(ys+1)−f(ys)≤β(xs−ys+1)⊤(xs−ys)−β2∥xs−ys+1∥2f(y_{s+1})-f(y_s)\le\beta(x_s-y_{s+1})^\top(x_s-y_s)-\tfrac\beta2\|x_s-y_{s+1}\|^2f(ys+1​)−f(ys​)≤β(xs​−ys+1​)⊤(xs​−ys​)−2β​∥xs​−ys+1​∥2.
  3. (3.24): f(ys+1)−f(x∗)≤β(xs−ys+1)⊤(xs−x∗)−β2∥xs−ys+1∥2f(y_{s+1})-f(x^*)\le\beta(x_s-y_{s+1})^\top(x_s-x^*)-\tfrac\beta2\|x_s-y_{s+1}\|^2f(ys+1​)−f(x∗)≤β(xs​−ys+1​)⊤(xs​−x∗)−2β​∥xs​−ys+1​∥2.
  4. The λ identity: λs−12=λs2−λs\lambda_{s-1}^2=\lambda_s^2-\lambda_sλs−12​=λs2​−λs​ for s≥1s\ge1s≥1.
  5. (3.25): λs2δs+1−λs−12δs≤β2(∥λsxs−(λs−1)ys−x∗∥2−∥λsys+1−(λs−1)ys−x∗∥2)\lambda_s^2\delta_{s+1}-\lambda_{s-1}^2\delta_s\le\tfrac\beta2\bigl(\|\lambda_sx_s-(\lambda_s-1)y_s-x^*\|^2-\|\lambda_sy_{s+1}-(\lambda_s-1)y_s-x^*\|^2\bigr)λs2​δs+1​−λs−12​δs​≤2β​(∥λs​xs​−(λs​−1)ys​−x∗∥2−∥λs​ys+1​−(λs​−1)ys​−x∗∥2).
  6. (3.26): λs+1xs+1−(λs+1−1)ys+1=λsys+1−(λs−1)ys\lambda_{s+1}x_{s+1}-(\lambda_{s+1}-1)y_{s+1}=\lambda_sy_{s+1}-(\lambda_s-1)y_sλs+1​xs+1​−(λs+1​−1)ys+1​=λs​ys+1​−(λs​−1)ys​.
  7. Telescoped bound: δt≤β2λt−12∥u1∥2\delta_t\le\frac{\beta}{2\lambda_{t-1}^2}\|u_1\|^2δt​≤2λt−12​β​∥u1​∥2 for t≥2t\ge2t≥2, with u1=λ1x1−(λ1−1)y1−x∗u_1=\lambda_1x_1-(\lambda_1-1)y_1-x^*u1​=λ1​x1​−(λ1​−1)y1​−x∗.
  8. Growth of λ: λt−1≥t/2\lambda_{t-1}\ge t/2λt−1​≥t/2 for t≥2t\ge2t≥2.

Significance

The result. Theorem 3.19 is the upper half of the statement that first-order methods on smooth convex functions have complexity Θ(β∥x1−x∗∥2/ε)\Theta(\sqrt{\beta\|x_1-x^*\|^2/\varepsilon})Θ(β∥x1​−x∗∥2/ε​): together with the black-box lower bound of Theorem 3.14 it shows that the accelerated method is optimal up to a constant factor, while plain gradient descent (Theorem 3.3, rate 2β∥x1−x∗∥2/(t−1)2\beta\|x_1-x^*\|^2/(t-1)2β∥x1​−x∗∥2/(t−1)) is not. The same potential-function argument, with the gradient step replaced by a proximal step, yields the rate of FISTA for composite objectives, and the identity λs−12=λs2−λs\lambda_{s-1}^2=\lambda_s^2-\lambda_sλs−12​=λs2​−λs​ is the algebraic core of most later analyses of accelerated and momentum methods.

Formalizing it. The theorem is classical and fully proved on paper. What this mission adds is a machine-checked proof of the O(1/t2)O(1/t^2)O(1/t2) rate for the exact step sequence of the book, with every intermediate inequality of the proof stated separately so that each can be closed and reused. A Lean development of accelerated gradient descent with this rate is not, to our knowledge, in Mathlib.

Difficulty

The rate does not follow from monotone decrease of f(yt)f(y_t)f(yt​), which is the engine of the analysis of plain gradient descent: the accelerated iterates are not monotone, and no single-step inequality on δt\delta_tδt​ alone gives 1/t21/t^21/t2. The proof instead tracks a weighted potential λs−12δs+β2∥us∥2\lambda_{s-1}^2\delta_s+\frac\beta2\|u_s\|^2λs−12​δs​+2β​∥us​∥2 and needs three exact algebraic coincidences: the identity λs−12=λs2−λs\lambda_{s-1}^2=\lambda_s^2-\lambda_sλs−12​=λs2​−λs​, the completion of a square 2a⊤b−∥a∥2=∥b∥2−∥b−a∥22a^\top b-\|a\|^2=\|b\|^2-\|b-a\|^22a⊤b−∥a∥2=∥b∥2−∥b−a∥2, and the rewriting (3.26) of the update rule, which makes consecutive potentials match exactly. In Lean the work is in this bookkeeping (vector identities in an inner-product space, a real recursion with square roots, and a telescoping sum indexed from 111), not in any deep analytic fact.

Formalization scope

  • Rn\mathbb R^nRn is EuclideanSpace ℝ (Fin n); x⊤yx^\top yx⊤y is ⟪x, y⟫_ℝ.
  • The gradient is an explicit map ggg with HasGradientAt f (g x) x for every xxx; β-smoothness is the Lipschitz bound on ggg (not the quadratic upper bound (3.4), which is a consequence). Convexity is ConvexOn ℝ Set.univ f.
  • λ\lambdaλ is a real sequence defined by recursion on N\mathbb NN with λ0=0\lambda_0=0λ0​=0; γt=(1−λt)/λt+1\gamma_t=(1-\lambda_t)/\lambda_{t+1}γt​=(1−λt​)/λt+1​ never divides by zero because λt+1≥1\lambda_{t+1}\ge1λt+1​≥1.
  • A run is a predicate on two sequences x,y:N→Rnx,y:\mathbb N\to\mathbb R^nx,y:N→Rn, indexed from 111 as in the book; index 000 is unconstrained. Every theorem quantifies over all runs.
  • Standing assumptions: x∗x^*x∗ is a minimizer of fff (book, p. 242); β>0\beta>0β>0 (implicit in the step 1/β1/\beta1/β) is a stated hypothesis.
  • The page prints the update as xt+1=(1−γs)yt+1+γtytx_{t+1}=(1-\gamma_s)y_{t+1}+\gamma_ty_txt+1​=(1−γs​)yt+1​+γt​yt​; the formalization uses γt\gamma_tγt​ in both places, as the proof's (3.26) requires. A run predicate with a fixed coefficient γs\gamma_sγs​ would describe a different (non-accelerated) method and is ruled out.
  • The goal is stated for all t≥1t\ge1t≥1, including t=1t=1t=1, which the printed proof does not cover but which follows from (3.4). The growth bound λt−1≥t/2\lambda_{t-1}\ge t/2λt−1​≥t/2 is stated for t≥2t\ge2t≥2, since λ0=0\lambda_0=0λ0​=0. The telescoping display prints δs2\delta_s^2δs2​ for δs\delta_sδs​; the formalization uses δs\delta_sδs​.

Contributions of every kind are welcome: proofs of the λ facts and of (3.26) are pure algebra; (3.23)–(3.25) need the descent lemma for β-smooth functions and first-order convexity, both of which are reusable well beyond this mission.

Selected references

  • S. Bubeck, Convex Optimization: Algorithms and Complexity, Foundations and Trends in Machine Learning 8(3–4):231–358, 2015. arXiv:1405.4980
  • Y. Nesterov, A method of solving a convex programming problem with convergence rate O(1/k²), Soviet Mathematics Doklady 27:372–376, 1983. mathnet
  • A. Beck, M. Teboulle, A fast iterative shrinkage-thresholding algorithm for linear inverse problems, SIAM J. Imaging Sciences 2(1):183–202, 2009. doi:10.1137/080716542
  • P. Tseng, On accelerated proximal gradient methods for convex-concave optimization, manuscript, 2008. pdf
  • A. Nemirovski, D. Yudin, Problem Complexity and Method Efficiency in Optimization, Wiley, 1983.
10 thms1 active userReviewed
Convex OptimizationMachine Learning·Captain: mikedeng1

Convex Optimization: Algorithms and Complexity VIII: No Black-Box Method Beats 3β‖x₁ − x*‖²/(32(t + 1)²) on β-Smooth Convex FunctionsTextbook

Why lower bounds for first-order methods

Upper bounds for an optimization method say how fast it converges; oracle complexity lower bounds say how fast any method of a given kind can possibly converge. Chapter 3 of S. Bubeck, Convex Optimization: Algorithms and Complexity (Foundations and Trends in Machine Learning 8(3–4), 2015, arXiv:1405.4980) proves upper bounds for subgradient descent on Lipschitz functions and for gradient methods on smooth functions. Section 3.5 (Lower bounds, pp. 279–283) shows that these rates cannot be improved by more than a numerical constant, as long as the number of queries is smaller than the dimension. For smooth convex functions the matching lower bound is what identifies Nesterov's accelerated gradient descent, with its 1/t21/t^21/t2 rate, as an optimal method.

Timeline. The lower bounds first appeared in A. Nemirovski and D. Yudin, Problem Complexity and Method Efficiency in Optimization (Wiley, 1983). The presentation followed by the book, with an explicit tridiagonal quadratic as the hard instance and the "span of past gradients" restriction on the method, is that of Y. Nesterov, Introductory Lectures on Convex Optimization (Kluwer, 2004), §2.1.2. Nesterov's accelerated method (1983) attains the matching upper bound.

Setting

Work in Rn\mathbb R^nRn with the Euclidean inner product x⊤yx^\top yx⊤y, coordinates x(1),…,x(n)x(1),\dots,x(n)x(1),…,x(n), canonical basis e1,…,ene_1,\dots,e_ne1​,…,en​ and balls B2(R)={x:∥x∥≤R}\mathrm B_2(R)=\{x:\|x\|\le R\}B2​(R)={x:∥x∥≤R}. A first-order oracle for fff answers a query xxx with a subgradient g∈∂f(x)g\in\partial f(x)g∈∂f(x) (the gradient when fff is differentiable). A black-box procedure maps the history (x1,g1,…,xt,gt)(x_1,g_1,\dots,x_t,g_t)(x1​,g1​,…,xt​,gt​) to the next query xt+1x_{t+1}xt+1​. Section 3.5 restricts attention to procedures with

x1=0,xt+1∈Span(g1,…,gt)(t≥0),(3.15)x_1=0,\qquad x_{t+1}\in\mathrm{Span}(g_1,\dots,g_t)\quad(t\ge0), \tag{3.15}x1​=0,xt+1​∈Span(g1​,…,gt​)(t≥0),(3.15)

which covers gradient descent, its accelerated variants and conjugate gradient. A function is β\betaβ-smooth if its gradient is β\betaβ-Lipschitz; LLL-Lipschitz on X\mathcal XX if every subgradient at every point of X\mathcal XX has norm at most LLL; α\alphaα-strongly convex if x↦f(x)−α2∥x∥2x\mapsto f(x)-\frac\alpha2\|x\|^2x↦f(x)−2α​∥x∥2 is convex.

The hard smooth instance uses, for k≤nk\le nk≤n, the symmetric tridiagonal matrix AkA_kAk​ with entries 222 on the first kkk diagonal positions and −1-1−1 on the neighbouring off-diagonal positions of the leading k×kk\times kk×k block, zero elsewhere, and the quadratics

fk(x)=β8x⊤Akx−β4x⊤e1,fk∗=inf⁡x∈Rnfk(x).f_k(x)=\frac\beta8x^\top A_kx-\frac\beta4x^\top e_1 ,\qquad f_k^*=\inf_{x\in\mathbb R^n}f_k(x).fk​(x)=8β​x⊤Ak​x−4β​x⊤e1​,fk∗​=x∈Rninf​fk​(x).

Formalization targets

Goal: Theorem 3.14 (p. 282)

For 1≤t≤n−121\le t\le\frac{n-1}21≤t≤2n−1​ and β>0\beta>0β>0 there are a β\betaβ-smooth convex fff and a minimizer x∗x^*x∗ such that every procedure satisfying (3.15) has

min⁡1≤s≤tf(xs)−f(x∗) ≥ 3β32 ∥x1−x∗∥2(t+1)2.\min_{1\le s\le t}f(x_s)-f(x^*)\ \ge\ \frac{3\beta}{32}\,\frac{\|x_1-x^*\|^2}{(t+1)^2}.1≤s≤tmin​f(xs​)−f(x∗) ≥ 323β​(t+1)2∥x1​−x∗∥2​.

The constant 3/323/323/32 is the book's.

Milestones (proof of Theorem 3.14, pp. 282–283)

  1. 0⪯Ak⪯4In0\preceq A_k\preceq4I_n0⪯Ak​⪯4In​, through x⊤Akx=x(1)2+x(k)2+∑i=1k−1(x(i)−x(i+1))2x^\top A_kx=x(1)^2+x(k)^2+\sum_{i=1}^{k-1}(x(i)-x(i+1))^2x⊤Ak​x=x(1)2+x(k)2+∑i=1k−1​(x(i)−x(i+1))2.
  2. For f=f2t+1f=f_{2t+1}f=f2t+1​ and any procedure satisfying (3.15), xs∈Span(e1,…,es−1)x_s\in\mathrm{Span}(e_1,\dots,e_{s-1})xs​∈Span(e1​,…,es−1​); hence f(xs)=fs(xs)f(x_s)=f_s(x_s)f(xs​)=fs​(xs​) for s≤ts\le ts≤t.
  3. xk∗(i)=1−ik+1x_k^*(i)=1-\frac i{k+1}xk∗​(i)=1−k+1i​ solves Akx=e1A_kx=e_1Ak​x=e1​, minimizes fkf_kfk​, and fk∗=−β8(1−1k+1)f_k^*=-\frac\beta8\bigl(1-\frac1{k+1}\bigr)fk∗​=−8β​(1−k+11​).
  4. ∥xk∗∥2≤k+13\|x_k^*\|^2\le\frac{k+1}3∥xk∗​∥2≤3k+1​.
  5. ft∗−f2t+1∗=β8(1t+1−12t+2)≥3β32∥x2t+1∗∥2(t+1)2f_t^*-f_{2t+1}^*=\frac\beta8\bigl(\frac1{t+1}-\frac1{2t+2}\bigr)\ge\frac{3\beta}{32}\frac{\|x^*_{2t+1}\|^2}{(t+1)^2}ft∗​−f2t+1∗​=8β​(t+11​−2t+21​)≥323β​(t+1)2∥x2t+1∗​∥2​.

Companion: Theorem 3.13 (p. 280)

For 1≤t≤n1\le t\le n1≤t≤n and L,R>0L,R>0L,R>0 there are a convex fff, LLL-Lipschitz on B2(R)\mathrm B_2(R)B2​(R), and a first-order oracle for it such that every procedure satisfying (3.15) has min⁡s≤tf(xs)−min⁡B2(R)f≥RL2(1+t)\min_{s\le t}f(x_s)-\min_{\mathrm B_2(R)}f\ge\frac{RL}{2(1+\sqrt t)}mins≤t​f(xs​)−minB2​(R)​f≥2(1+t​)RL​; and for α>0\alpha>0α>0 there are an α\alphaα-strongly convex fff, LLL-Lipschitz on B2(L2α)\mathrm B_2(\frac L{2\alpha})B2​(2αL​), and an oracle with gap at least L28αt\frac{L^2}{8\alpha t}8αtL2​ over that ball.

Significance

The upper bounds of Chapter 3 (projected subgradient descent at rate RL/tRL/\sqrt tRL/t​, accelerated gradient descent at rate β∥x1−x∗∥2/t2\beta\|x_1-x^*\|^2/t^2β∥x1​−x∗∥2/t2) become optimal statements only through these lower bounds: no method in the class (3.15) can be faster by more than a constant factor while ttt is below the dimension. The restriction to t≲nt\lesssim nt≲n is necessary, since Chapter 2's cutting-plane methods converge exponentially once the number of queries exceeds the dimension.

The results are classical and proved. Formalizing them yields machine-checked versions of the quadratic-form computation for the tridiagonal matrix, of the Krylov-type support argument under (3.15), and of the explicit minimizer of fkf_kfk​, each reusable in other lower-bound arguments (Theorem 3.15 in ℓ2\ell_2ℓ2​, lower bounds for strongly convex smooth functions, conjugate gradient analyses). The platform held no formal statement of these oracle lower bounds when this mission was drafted.

Difficulty

Each analytic step is elementary; the difficulty is in the bookkeeping. The span argument is an induction that must track, at each step, that the gradient of a tridiagonal quadratic at a vector supported on the first s−1s-1s−1 coordinates is supported on the first sss, and that the span hypothesis transfers this to the next query. The minimizer computation requires solving Akx=e1A_kx=e_1Ak​x=e1​ on the leading block and showing that the coordinates beyond kkk do not affect fkf_kfk​. A natural first attempt, choosing the hard function after seeing the procedure, proves a much weaker statement and is excluded by the quantifier order: the function is fixed first and must defeat every procedure.

Formalization scope

Rn\mathbb R^nRn is EuclideanSpace ℝ (Fin n). Coordinates in Lean are 0-based; the definitions provide the book's 1-based coordinate coord x i and basis vector basisVec n i, and the matrix tridiag n k translates book index iii to Fin n index i−1i-1i−1. The query sequence starts at index 111. The oracle is a fixed map ggg, so (3.15) reads xt+1∈Span(g(x1),…,g(xt))x_{t+1}\in\mathrm{Span}(g(x_1),\dots,g(x_t))xt+1​∈Span(g(x1​),…,g(xt​)) with x1=0x_1=0x1​=0. In Theorem 3.14 the oracle is the gradient, given as a map with HasGradientAt everywhere; β\betaβ-smoothness is the Lipschitz bound on that map. In Theorem 3.13 the oracle is part of what is constructed, because the book's proof uses a specific "resisting" subgradient selection and the claim fails for an arbitrary one.

Committed conventions, each stated in the item's Formalization Note: the minimum over 1≤s≤t1\le s\le t1≤s≤t is the bound for every such sss, and t≥1t\ge1t≥1 is required; t≤(n−1)/2t\le(n-1)/2t≤(n−1)/2 is 2t+1≤n2t+1\le n2t+1≤n; the minimizer x∗x^*x∗ is existentially chosen together with fff (the hard function has many minimizers when 2t+1<n2t+1<n2t+1<n, and the bound is false for some of them); the minimum over a ball is the bound against every point of the ball; fk∗f_k^*fk∗​ is the real infimum, asserted to be attained.

A formalization that let the function depend on the procedure, dropped x1=0x_1=0x1​=0, or took the span over gradients at points other than the queries would state a different and weaker theorem; the statements here keep fff (and the oracle) before the universally quantified procedure.

Infrastructure needed: quadratic forms of explicit matrices on EuclideanSpace, gradients of quadratics, and span/support lemmas for EuclideanSpace.single-type vectors. Contributions of proofs for any milestone, and of the strongly convex ℓ2\ell_2ℓ2​ lower bound (Theorem 3.15, not included here), are welcome.

Selected references

  • S. Bubeck, Convex Optimization: Algorithms and Complexity, Foundations and Trends in Machine Learning 8(3–4):231–358, 2015. https://arxiv.org/abs/1405.4980
  • A. Nemirovski and D. Yudin, Problem Complexity and Method Efficiency in Optimization, Wiley, 1983.
  • Y. Nesterov, Introductory Lectures on Convex Optimization: A Basic Course, Kluwer, 2004. https://doi.org/10.1007/978-1-4419-8853-9
8 thms1 active userReviewed
Operations ResearchOptimal TransportProbability·Captain: mikedeng1

Quantifying Distributional Model Risk via Optimal Transport 1: Strong Duality — the Worst-Case Expectation over an Optimal-Transport Ball on a Polish Space Equals Its Dual over (λ, φ)Research Paper

Motivation

A probability model μ\muμ for a random element XXX is rarely known exactly. Distributionally robust performance analysis replaces the single expectation Eμ[f(X)]E_\mu[f(X)]Eμ​[f(X)] by its worst case over all models within a prescribed distance of μ\muμ. When the distance is an optimal-transport cost, the neighbourhood contains models whose support differs from that of μ\muμ. That matters in stochastic-process applications such as ruin probabilities for insurance reserves, where the natural alternatives (a compensated Poisson process against a Brownian motion) are mutually singular and likelihood-based divergences such as Kullback–Leibler are infinite.

Blanchet and Murthy (arXiv:1604.01446, Math. Oper. Res. 2019) prove that the worst-case expectation over an optimal-transport ball equals a one-dimensional dual problem. They assume only that the underlying space is Polish, the cost lower semicontinuous and the performance function upper semicontinuous and integrable.

Timeline. Esfahani and Kuhn (arXiv:1505.05116, 2015/2018) obtained a dual reformulation for Wasserstein balls around empirical measures on Rd\mathbb R^dRd. Gao and Kleywegt (arXiv:1604.02199, 2016) proved a general duality whose proof, as Blanchet and Murthy note, uses the local compactness of the space. Blanchet and Murthy (2016, v2 2017) removed local compactness and continuity of the cost. This covers path spaces such as C[0,T]C[0,T]C[0,T] and D[0,T]D[0,T]D[0,T].

Setting

Let SSS be a Polish space with Borel σ-algebra B(S)\mathcal B(S)B(S), and let μ\muμ be a probability measure on SSS (the baseline model).

  • Cost (A1). c:S×S→[0,∞)c : S\times S\to[0,\infty)c:S×S→[0,∞) is lower semicontinuous, and c(x,y)=0c(x,y)=0c(x,y)=0 if and only if x=yx=yx=y.
  • Performance function (A2). f:S→Rf : S\to\mathbb Rf:S→R is upper semicontinuous and μ\muμ-integrable.
  • Budget. δ>0\delta>0δ>0.

The primal feasible set Φμ,δ\Phi_{\mu,\delta}Φμ,δ​ consists of the probability measures π\piπ on S×SS\times SS×S whose first marginal is μ\muμ and whose transport cost satisfies ∫c dπ≤δ\int c\,d\pi\le\delta∫cdπ≤δ. The second marginal of π\piπ is the alternative model. The primal objective is I(π)=∫f(y) dπ(x,y)I(\pi)=\int f(y)\,d\pi(x,y)I(π)=∫f(y)dπ(x,y), and the primal value is

I=sup⁡{I(π):π∈Φμ,δ}.I=\sup\{I(\pi):\pi\in\Phi_{\mu,\delta}\}.I=sup{I(π):π∈Φμ,δ​}.

The universal σ-algebra U(S)\mathcal U(S)U(S) is the intersection of the completions of B(S)\mathcal B(S)B(S) under all probability measures. Write mU(S;Rˉ)m\mathcal U(S;\bar{\mathbb R})mU(S;Rˉ) for the U(S)\mathcal U(S)U(S)-measurable functions S→[−∞,∞]S\to[-\infty,\infty]S→[−∞,∞]. The dual feasible set Λc,f\Lambda_{c,f}Λc,f​ consists of the pairs (λ,φ)(\lambda,\varphi)(λ,φ) with λ≥0\lambda\ge0λ≥0, φ∈mU(S;Rˉ)\varphi\in m\mathcal U(S;\bar{\mathbb R})φ∈mU(S;Rˉ) and φ(x)+λc(x,y)≥f(y)\varphi(x)+\lambda c(x,y)\ge f(y)φ(x)+λc(x,y)≥f(y) for all x,yx,yx,y. The dual objective is J(λ,φ)=λδ+∫φ dμJ(\lambda,\varphi)=\lambda\delta+\int\varphi\,d\muJ(λ,φ)=λδ+∫φdμ, and the dual value is J=inf⁡{J(λ,φ):(λ,φ)∈Λc,f}J=\inf\{J(\lambda,\varphi):(\lambda,\varphi)\in\Lambda_{c,f}\}J=inf{J(λ,φ):(λ,φ)∈Λc,f​}. Finally,

φλ(x)=sup⁡y∈S{f(y)−λc(x,y)}∈R∪{∞}.\varphi_\lambda(x)=\sup_{y\in S}\{f(y)-\lambda c(x,y)\}\in\mathbb R\cup\{\infty\}.φλ​(x)=y∈Ssup​{f(y)−λc(x,y)}∈R∪{∞}.

Formalization targets

Goal: Theorem 1

Under (A1) and (A2):

  1. strong duality,
sup⁡{I(π):π∈Φμ,δ}=inf⁡{J(λ,φ):(λ,φ)∈Λc,f};\sup\{I(\pi):\pi\in\Phi_{\mu,\delta}\}=\inf\{J(\lambda,\varphi):(\lambda,\varphi)\in\Lambda_{c,f}\};sup{I(π):π∈Φμ,δ​}=inf{J(λ,φ):(λ,φ)∈Λc,f​};
  1. there is λ∗≥0\lambda^*\ge0λ∗≥0 such that (λ∗,φλ∗)(\lambda^*,\varphi_{\lambda^*})(λ∗,φλ∗​) is a dual optimizer;
  2. a feasible π∗\pi^*π∗ and a feasible (λ∗,φλ∗)(\lambda^*,\varphi_{\lambda^*})(λ∗,φλ∗​) with finite J(λ∗,φλ∗)J(\lambda^*,\varphi_{\lambda^*})J(λ∗,φλ∗​) are optimal with I(π∗)=J(λ∗,φλ∗)I(\pi^*)=J(\lambda^*,\varphi_{\lambda^*})I(π∗)=J(λ∗,φλ∗​) if and only if the complementary slackness conditions hold:
f(y)−λ∗c(x,y)=φλ∗(x)  π∗-a.s.,λ∗(∫c dπ∗−δ)=0.f(y)-\lambda^*c(x,y)=\varphi_{\lambda^*}(x)\ \ \pi^*\text{-a.s.},\qquad \lambda^*\Big(\int c\,d\pi^*-\delta\Big)=0.f(y)−λ∗c(x,y)=φλ∗​(x)  π∗-a.s.,λ∗(∫cdπ∗−δ)=0.

The "if" direction is stated without the finiteness assumption.

Milestones

Weak duality I≤JI\le JI≤J (5). Lemma 15. Strong duality with a primal optimizer on compact SSS, first for continuous costs (Proposition 5), then for lower semicontinuous ones (Proposition 6). Universal measurability of φλ\varphi_\lambdaφλ​ (§4.2). Lemma 16. The restricted dual bound of Proposition 7. Lemma 8. The univariate formula (9):

I=inf⁡λ≥0{λδ+Eμ[sup⁡y∈S{f(y)−λc(X,y)}]}.I=\inf_{\lambda\ge0}\Big\{\lambda\delta+E_\mu\Big[\sup_{y\in S}\{f(y)-\lambda c(X,y)\}\Big]\Big\}.I=λ≥0inf​{λδ+Eμ​[y∈Ssup​{f(y)−λc(X,y)}]}.

Significance

The result. Formula (9) turns an infinite-dimensional optimization over probability measures into a one-dimensional convex minimization that involves only the baseline μ\muμ. A modeller can therefore evaluate it by sampling from μ\muμ. Theorem 1 is the input for the worst-case probability formula for closed sets (Theorem 3 of the paper) and for the existence of worst-case transport plans (Corollary 1). Its complementary slackness conditions describe the structure of every worst-case plan: mass is moved from xxx to maximizers of f(z)−λ∗c(x,z)f(z)-\lambda^*c(x,z)f(z)−λ∗c(x,z), and the budget is exhausted whenever λ∗>0\lambda^*>0λ∗>0.

Formalizing it. The result is proved on paper. To our knowledge it has no machine-checked proof. The only related statement on Prove2Me is a special case (empirical baseline, bounded continuous loss, power-of-norm cost on Rm\mathbb R^mRm). A complete development would contain duality on compact spaces via Fenchel duality, the extension to σ-compact supports, and measurable-selection arguments for universally measurable functions. The measurable-selection part reuses Bertsekas–Shreve's analytic-set theory, which is already posed on the platform.

Difficulty

The obvious route copies Kantorovich duality. That route fails here because the feasible set Φμ,δ\Phi_{\mu,\delta}Φμ,δ​ fixes only one marginal, so it is not tight on a non-compact space. Prokhorov compactness is available only on compact pieces Sn×SnS_n\times S_nSn​×Sn​. The duality must then be transported to the whole space by a limiting argument that keeps control of the dual multipliers.

A second obstacle is measurability. For a merely lower semicontinuous cost on a non-locally-compact space, φλ\varphi_\lambdaφλ​ need not be Borel measurable, so the dual must range over universally measurable functions. Removing the restriction y∈Sπy\in S_\piy∈Sπ​ from the envelope (Lemma 8) needs a measurable selection theorem. Arguments that assume closed balls are compact do not apply in the target spaces C[0,T]C[0,T]C[0,T] and D[0,T]D[0,T]D[0,T].

Formalization scope

  • Space and costs. S carries [TopologicalSpace S] [PolishSpace S] [MeasurableSpace S] [BorelSpace S]. The cost is a real-valued curried function c : S → S → ℝ; (A1) is the structure AssumptionA1; (A2) is UpperSemicontinuous f together with Integrable f μ; and 0 < δ is assumed throughout.
  • Extended reals. III, JJJ, I(π)I(\pi)I(π), J(λ,φ)J(\lambda,\varphi)J(λ,φ) and φλ\varphi_\lambdaφλ​ live in EReal. The integral of an extended-real function is ∫φ+−∫φ−\int\varphi^+-\int\varphi^-∫φ+−∫φ− with lower Lebesgue integrals, and ∞−∞\infty-\infty∞−∞ evaluates to −∞-\infty−∞. A coupling with ∫f− dπ=∞\int f^-\,d\pi=\infty∫f−dπ=∞ therefore never raises III, which is the paper's reading in footnote 2.
  • Measurability and integrals. Universal measurability is the published BertsekasShreve.AnalyticSelection.IsUniversallyMeasurable. For such φ\varphiφ the lower integral equals the integral against the completion of μ\muμ.
  • Variants. The dual feasible set takes a set KKK: with K=SK=SK=S it is (6b), and with K=SπK=S_\piK=Sπ​ it is (29).
  • Hidden hypothesis. The "only if" part of Theorem 1(b) carries the hypothesis J(λ∗,φλ∗)<∞J(\lambda^*,\varphi_{\lambda^*})<\inftyJ(λ∗,φλ∗​)<∞. Without it the equivalence fails when I=J=∞I=J=\inftyI=J=∞.
  • Ruled-out trivializations. A primal that fixes both marginals (or neither), a dual over Borel-measurable φ\varphiφ, and a Bochner integral for ∫f dπ\int f\,d\pi∫fdπ (which is 000 off L1(π)L^1(\pi)L1(π)) all describe different problems and are ruled out by the definitions.
  • Infrastructure and contributions. Needed: Fenchel duality on Cb(S×S)C_b(S\times S)Cb​(S×S) and its dual M(S×S)M(S\times S)M(S×S) (Riesz–Markov–Kakutani), Prokhorov's theorem, Sion's minimax theorem, and Jankov–von Neumann selection. Several are on the platform or in Mathlib, and all are reusable beyond this mission. Proofs of the milestones in any order, and of the posed Bertsekas–Shreve tools, are welcome.

Selected references

  • J. Blanchet and K. Murthy, Quantifying Distributional Model Risk via Optimal Transport, Math. Oper. Res. 44(2):565–600, 2019. arXiv:1604.01446v2, doi:10.1287/moor.2018.0936
  • R. Gao and A. Kleywegt, Distributionally Robust Stochastic Optimization with Wasserstein Distance, 2016. arXiv:1604.02199
  • P. Mohajerin Esfahani and D. Kuhn, Data-driven distributionally robust optimization using the Wasserstein metric, Math. Program. 171:115–166, 2018. arXiv:1505.05116
  • D. Bertsekas and S. Shreve, Stochastic Optimal Control: The Discrete-Time Case, Academic Press, 1978, Chapter 7. MIT open copy
  • C. Villani, Optimal Transport: Old and New, Springer, 2008. doi:10.1007/978-3-540-71050-9
19 thms1 active userReviewed
Operations ResearchProbability·Captain: mikedeng1

On the Power of Robust Solutions in Two-Stage Stochastic and Adaptive Optimization Problems 1: For Symmetric Right-Hand-Side Uncertainty, the Robust Optimum Is at Most Twice the Stochastic OptimumResearch Paper

Motivation

Many planning problems are made in two stages: a first decision xxx (capacity, inventory, a network design) is fixed before an uncertain demand is revealed, and a second decision yyy (recourse, routing, overtime) is taken afterwards. Two-stage stochastic optimization models the demand as random and minimizes expected cost; its second stage is a whole policy ω↦y(ω)\omega\mapsto y(\omega)ω↦y(ω), and the problem is intractable in general, especially with integer variables (Dyer and Stougie, 2006). Robust optimization instead picks one static pair (x,y)(x,y)(x,y) that is feasible for every possible demand and minimizes its worst-case cost; it is a single deterministic mixed-integer program and needs no knowledge of the distribution (Ben-Tal and Nemirovski, 2002; Bertsimas and Sim, 2004).

The question this mission addresses is how much is lost by solving the robust problem in place of the stochastic one. Bertsimas and Goyal (Math. Oper. Res. 2010) show that when only the right-hand side is uncertain, the uncertainty set is symmetric and the distribution is centred at its point of symmetry, the loss is at most a factor of two, and that this factor is tight.

Setting

Fix A∈Rm×n1A\in\mathbb R^{m\times n_1}A∈Rm×n1​, B∈Rm×n2B\in\mathbb R^{m\times n_2}B∈Rm×n2​ and nonnegative costs c∈R+n1c\in\mathbb R^{n_1}_+c∈R+n1​​, d∈R+n2d\in\mathbb R^{n_2}_+d∈R+n2​​. A set Ω\OmegaΩ of scenarios carries a probability measure μ\muμ, and each scenario ω\omegaω has a right-hand side b(ω)∈R+mb(\omega)\in\mathbb R^m_+b(ω)∈R+m​. The uncertainty set is Ib(Ω)={b(ω):ω∈Ω}\mathcal I_b(\Omega)=\{b(\omega):\omega\in\Omega\}Ib​(Ω)={b(ω):ω∈Ω}. First-stage variables are nonnegative, with integer values on a designated set of coordinates; second-stage variables are nonnegative reals (p2=0p_2=0p2​=0).

The stochastic problem ΠStoch(b)\Pi_{\mathrm{Stoch}}(b)ΠStoch​(b), (1.1), chooses xxx and a policy y(⋅)y(\cdot)y(⋅):

zStoch(b)=inf⁡ cTx+Eμ[dTy(ω)]s.t.Ax+By(ω)≥b(ω)  ∀ω∈Ω.z_{\mathrm{Stoch}}(b)=\inf\ c^Tx+\mathbb E_\mu[d^Ty(\omega)]\quad\text{s.t.}\quad Ax+By(\omega)\ge b(\omega)\ \ \forall\omega\in\Omega .zStoch​(b)=inf cTx+Eμ​[dTy(ω)]s.t.Ax+By(ω)≥b(ω)  ∀ω∈Ω.

The robust problem ΠRob(b)\Pi_{\mathrm{Rob}}(b)ΠRob​(b), (1.2), chooses one yyy for all scenarios:

zRob(b)=inf⁡ cTx+dTys.t.Ax+By≥b(ω)  ∀ω∈Ω.z_{\mathrm{Rob}}(b)=\inf\ c^Tx+d^Ty\quad\text{s.t.}\quad Ax+By\ge b(\omega)\ \ \forall\omega\in\Omega .zRob​(b)=inf cTx+dTys.t.Ax+By≥b(ω)  ∀ω∈Ω.

A set PPP is symmetric (Definition 1.2) if there is u0∈Pu^0\in Pu0∈P with u0+z∈P  ⟺  u0−z∈Pu^0+z\in P\iff u^0-z\in Pu0+z∈P⟺u0−z∈P for all zzz; u0u^0u0 is its point of symmetry. Hypercubes, ellipsoids and norm balls are symmetric. A probability measure on a symmetric set is symmetric (Definition 1.4) if it gives a set and its reflection {2u0−x}\{2u^0-x\}{2u0−x} the same mass.

Formalization targets

Goal: Theorem 2.1 (p. 10)

If Ib(Ω)\mathcal I_b(\Omega)Ib​(Ω) is symmetric with point of symmetry b(ω0)b(\omega^0)b(ω0), p2=0p_2=0p2​=0, and μ\muμ satisfies

Eμ[b(ω)] ≥ b(ω0)(2.1)\mathbb E_\mu[b(\omega)]\ \ge\ b(\omega^0)\qquad(2.1)Eμ​[b(ω)] ≥ b(ω0)(2.1)

then

zRob(b) ≤ 2⋅zStoch(b).z_{\mathrm{Rob}}(b)\ \le\ 2\cdot z_{\mathrm{Stoch}}(b).zRob​(b) ≤ 2⋅zStoch​(b).

Milestones on the way

  • Lemma 2.2 (p. 12): the coordinatewise bounding box HHH of a symmetric set SSS is the smallest hypercube containing SSS.
  • Lemma 2.3 (p. 12): the centre x0x^0x0 of HHH is the point of symmetry of SSS, and x≤2x0x\le 2x^0x≤2x0 on SSS when S⊆R+nS\subseteq\mathbb R^n_+S⊆R+n​.
  • Eqs. (2.9)–(2.10) (p. 13): if (x,y)(x,y)(x,y) covers b(ω0)b(\omega^0)b(ω0) then (2x,2y)(2x,2y)(2x,2y) covers every b(ω)b(\omega)b(ω), so it is robust feasible.
  • p. 14 display: under (2.1), the mean second-stage decision Eμ[y(ω)]\mathbb E_\mu[y(\omega)]Eμ​[y(ω)] covers b(ω0)b(\omega^0)b(ω0).
  • Lemma 2.1 (p. 11): a symmetric probability measure has mean u0u^0u0, so it satisfies (2.1).
  • Theorem 2.7 (p. 21): the same bound zRob(b)≤2 zStoch(b)z_{\mathrm{Rob}}(b)\le 2\,z_{\mathrm{Stoch}}(b)zRob​(b)≤2zStoch​(b) when Ib(Ω)\mathcal I_b(\Omega)Ib​(Ω) is convex and positive (contained in a symmetric subset of R+m\mathbb R^m_+R+m​ whose centre lies in Ib(Ω)\mathcal I_b(\Omega)Ib​(Ω)).

Significance

The result. The robust problem is one mixed-integer program, independent of μ\muμ; the stochastic problem optimizes over policies and requires the distribution. Theorem 2.1 says that under symmetry the static robust solution (x,y)(x,y)(x,y) used in every scenario is a 2-approximation of the optimal expected cost, for every centred distribution at once. The companion results of the paper show the hypotheses matter: the bound is tight for symmetric sets, the gap is unbounded (at least n+1n+1n+1) on the non-symmetric simplex (Theorem 2.6), and unbounded when costs are uncertain as well (Theorem 3.1). The theorem also underlies later work on the power of static and affine policies in adaptive optimization (Bertsimas and Goyal, 2012).

Formalizing it. The theorem and its proof are published; nothing in this mission is open mathematics. To our knowledge none of these statements has a machine-checked proof. The mission produces a reusable Lean model of two-stage stochastic and robust mixed-integer covering problems with arbitrary scenario spaces, extended-real optimal values and genuine expectations, together with the elementary geometry of point-symmetric sets. The same objects are used by the other missions of this series (the simplex and cost-uncertainty gaps, and the adaptability gap).

Difficulty

Each step of the published argument is short; the difficulty is in stating it at the right generality. The paper begins "consider an optimal solution" of ΠStoch(b)\Pi_{\mathrm{Stoch}}(b)ΠStoch​(b); optimal policies need not exist for an arbitrary scenario space, so the statement is about infima and every step must work for an arbitrary feasible pair. Passing from "Ax+By(ω)≥b(ω)Ax+By(\omega)\ge b(\omega)Ax+By(ω)≥b(ω) for all ω\omegaω" to "Ax+B Eμ[y]≥Eμ[b]Ax+B\,\mathbb E_\mu[y]\ge\mathbb E_\mu[b]Ax+BEμ​[y]≥Eμ​[b]" needs integrability of the policy and of bbb and linearity of the Bochner integral through a matrix. The bound b(ω)≤2b(ω0)b(\omega)\le 2b(\omega^0)b(ω)≤2b(ω0) uses symmetry together with nonnegativity of the uncertainty set; symmetry alone does not give it. Integrality of the second stage breaks the argument, since Eμ[y(ω)]\mathbb E_\mu[y(\omega)]Eμ​[y(ω)] need not be integral.

Formalization scope

  • Vectors are Fin k → ℝ with the componentwise order, products A *ᵥ x and inner products c ⬝ᵥ x. The mixed-integer domain is "nonnegative with integer values on a set III of coordinates", which is the paper's R+n−p×Z+p\mathbb R^{n-p}_+\times\mathbb Z^p_+R+n−p​×Z+p​ up to relabelling.
  • Ω\OmegaΩ is an arbitrary measurable space with a probability measure; Ib(Ω)\mathcal I_b(\Omega)Ib​(Ω) is Set.range b. Constraints hold for every scenario, not almost surely.
  • Second-stage policies are μ\muμ-integrable, and bbb is μ\muμ-integrable in every statement that uses (2.1). Without these, Lean's integral of a non-integrable function is 000 and (2.1) would degenerate.
  • zStochz_{\mathrm{Stoch}}zStoch​ and zRobz_{\mathrm{Rob}}zRob​ are infima in EReal, equal to +∞+\infty+∞ when infeasible; no attainment is assumed. A real-valued infimum would return 000 on an infeasible robust problem and make the goal trivial; that formalization is ruled out.
  • The bounding box of (2.5)–(2.7) uses suprema and infima, with boundedness assumed where needed.
  • Corrections to the page: Lemma 2.3's inequality x≤2x0x\le 2x^0x≤2x0 is stated under S⊆R+nS\subseteq\mathbb R^n_+S⊆R+n​, which its proof uses and which holds in every application; Lemma 2.1 assumes the measure has a mean; Theorem 2.7 carries the standing assumption p2=0p_2=0p2​=0 of §2.

Contributions welcome: proofs of the milestones and the goal, and general lemmas on point-symmetric sets and on interchanging Bochner integrals with matrix–vector products, both reusable outside this mission.

Selected references

  • D. Bertsimas, V. Goyal, On the power of robust solutions in two-stage stochastic and adaptive optimization problems, Mathematics of Operations Research 35(2), 2010. https://doi.org/10.1287/moor.1090.0440 (cited from the authors' manuscript, MIT DSpace)
  • A. Ben-Tal, A. Nemirovski, Robust optimization — methodology and applications, Mathematical Programming 92, 2002. https://doi.org/10.1007/s101070100286
  • D. Bertsimas, M. Sim, The price of robustness, Operations Research 52(1), 2004. https://doi.org/10.1287/opre.1030.0065
  • M. Dyer, L. Stougie, Computational complexity of stochastic programming problems, Mathematical Programming 106, 2006. https://doi.org/10.1007/s10107-005-0578-0
  • D. Bertsimas, V. Goyal, On the power and limitations of affine policies in two-stage adaptive optimization, Mathematical Programming 134, 2012. https://doi.org/10.1007/s10107-011-0444-4
10 thms1 active userReviewed
Linear OptimizationOperations Research·Captain: mikedeng1

A Robust Optimization Approach to Inventory Theory: The Optimal Robust Policy Is the Optimal Nominal Policy for an Explicit Modified Demand, at Extra Cost (2ph/(p+h))·ΣA_kResearch Paper

Motivation

Classical inventory theory chooses order quantities against a probability distribution of demand. The resulting dynamic programs are optimal in expectation but need the distribution, and they become intractable once several installations, capacities or fixed costs interact. Robust optimization replaces the distribution by an uncertainty set and asks for the order sequence whose worst-case cost over that set is smallest. Bertsimas and Thiele (Operations Research 54(1), 2006) applied the budget-of-uncertainty approach of Bertsimas and Sim (The Price of Robustness, Operations Research 52(1), 2004) to finite-horizon inventory control. Their main structural result says that robustness does not destroy the structure of the classical problem. The robust problem is a deterministic (nominal) inventory problem with an explicitly modified demand, and the price of robustness is an explicit constant.

Setting

A single item is ordered at a single installation over periods k=0,…,T−1k = 0, \dots, T-1k=0,…,T−1. The stock at the beginning of the horizon is x0x_0x0​. Orders uk≥0u_k \ge 0uk​≥0 arrive immediately, demand wkw_kwk​ is subtracted, and excess demand is backlogged, so the stock at the end of period kkk is

xk+1=x0+∑i=0k(ui−wi).x_{k+1} = x_0 + \sum_{i=0}^{k} (u_i - w_i).xk+1​=x0​+i=0∑k​(ui​−wi​).

The demand of period kkk is uncertain: wk=wˉk+w^kzkw_k = \bar w_k + \hat w_k z_kwk​=wˉk​+w^k​zk​ with a nominal demand wˉk\bar w_kwˉk​, a maximal deviation w^k≥0\hat w_k \ge 0w^k​≥0 and a scaled deviation zk∈[−1,1]z_k \in [-1, 1]zk​∈[−1,1]. A budget of uncertainty Γk\Gamma_kΓk​ limits the total scaled deviation up to period kkk. The budgets satisfy 0≤Γ00 \le \Gamma_00≤Γ0​ and Γk≤Γk+1≤Γk+1\Gamma_k \le \Gamma_{k+1} \le \Gamma_k + 1Γk​≤Γk+1​≤Γk​+1.

Each period costs C(uk)+R(xk+1)C(u_k) + R(x_{k+1})C(uk​)+R(xk+1​). The purchasing cost is C(u)=K+cuC(u) = K + cuC(u)=K+cu for u>0u > 0u>0 and C(0)=0C(0) = 0C(0)=0, with c>0c > 0c>0 and K≥0K \ge 0K≥0. The holding/shortage cost is R(x)=max⁡(hx,−px)R(x) = \max(hx, -px)R(x)=max(hx,−px), with h≥0h \ge 0h≥0 and p>cp > cp>c. The nominal problem with demand www minimizes ∑k<T(C(uk)+R(xk+1))\sum_{k<T} (C(u_k) + R(x_{k+1}))∑k<T​(C(uk​)+R(xk+1​)) over u≥0u \ge 0u≥0 for a fixed demand sequence www.

For each kkk, AkA_kAk​ is the optimal value of the linear program

Ak=max⁡{∑i=0kw^izi  :  ∑i=0kzi≤Γk, 0≤zi≤1}(13)A_k = \max\Big\{\sum_{i=0}^{k} \hat w_i z_i \;:\; \sum_{i=0}^{k} z_i \le \Gamma_k,\ 0 \le z_i \le 1\Big\} \qquad (13)Ak​=max{i=0∑k​w^i​zi​:i=0∑k​zi​≤Γk​, 0≤zi​≤1}(13)

It is the worst-case deviation of the cumulative demand up to kkk from its nominal value, with A−1=0A_{-1} = 0A−1​=0. Write xˉk+1=x0+∑i≤k(ui−wˉi)\bar x_{k+1} = x_0 + \sum_{i\le k}(u_i - \bar w_i)xˉk+1​=x0​+∑i≤k​(ui​−wˉi​) for the nominal stock. The robust formulation (14) minimizes ∑k<T(C(uk)+yk)\sum_{k<T} (C(u_k) + y_k)∑k<T​(C(uk​)+yk​) over (u,y,q,r)(u, y, q, r)(u,y,q,r) subject to the following constraints for every k<Tk < Tk<T:

  • uk≥0u_k \ge 0uk​≥0, qk≥0q_k \ge 0qk​≥0, and rik≥0r_{ik} \ge 0rik​≥0, qk+rik≥w^iq_k + r_{ik} \ge \hat w_iqk​+rik​≥w^i​ for i≤ki \le ki≤k;
  • yk≥h(xˉk+1+qkΓk+∑i≤krik)y_k \ge h(\bar x_{k+1} + q_k\Gamma_k + \sum_{i\le k} r_{ik})yk​≥h(xˉk+1​+qk​Γk​+∑i≤k​rik​);
  • yk≥p(−xˉk+1+qkΓk+∑i≤krik)y_k \ge p(-\bar x_{k+1} + q_k\Gamma_k + \sum_{i\le k} r_{ik})yk​≥p(−xˉk+1​+qk​Γk​+∑i≤k​rik​).

The variables q,rq, rq,r are the dual of (13). Formulation (14) is equivalent to requiring the holding and shortage constraints of period kkk for every demand whose scaled deviations satisfy ∣zi∣≤1|z_i| \le 1∣zi​∣≤1 and ∑i≤k∣zi∣≤Γk\sum_{i \le k}|z_i| \le \Gamma_k∑i≤k​∣zi​∣≤Γk​.

Formalization targets

Goal: Theorem 3.2 (a), (b), (d)

Let the modified demand be

wk′=wˉk+p−hp+h (Ak−Ak−1).(20)w'_k = \bar w_k + \frac{p-h}{p+h}\,(A_k - A_{k-1}). \qquad (20)wk′​=wˉk​+p+hp−h​(Ak​−Ak−1​).(20)

Write Nw′(u)N_{w'}(u)Nw′​(u) for the nominal cost of uuu under demand w′w'w′. Then:

  1. For every u≥0u \ge 0u≥0, the minimum of the objective of (14) over the feasible (y,q,r)(y, q, r)(y,q,r) is attained and equals
Nw′(u)+2php+h∑k=0T−1Ak.N_{w'}(u) + \frac{2ph}{p+h}\sum_{k=0}^{T-1} A_k.Nw′​(u)+p+h2ph​k=0∑T−1​Ak​.
  1. uuu is the order part of an optimal solution of (14) if and only if uuu is optimal for the nominal problem with demand w′w'w′.
  2. The optimal cost of (14) is the optimal nominal cost under w′w'w′ plus 2php+h∑kAk\frac{2ph}{p+h}\sum_k A_kp+h2ph​∑k​Ak​.
  3. If K=0K = 0K=0 and wk′≥0w'_k \ge 0wk′​≥0, the order-up-to policy with levels Sk=wk′S_k = w'_kSk​=wk′​ is robust-optimal.

Milestones

The milestones are the steps of the paper's proof, in order:

  • LP (13) and its dual are attained with the common value AkA_kAk​.
  • The constraints of (14) are the robust counterpart of the kkk-th holding/shortage pair (10)–(11).
  • For fixed orders, the value of (14) is the sum (21).
  • The modified stock (22) satisfies xk+1′=xˉk+1−p−hp+hAkx'_{k+1} = \bar x_{k+1} - \frac{p-h}{p+h}A_kxk+1′​=xˉk+1​−p+hp−h​Ak​.
  • The max identity (23): max⁡(h(xˉ+A),p(−xˉ+A))=max⁡(hx′,−px′)+2php+hA\max(h(\bar x+A), p(-\bar x+A)) = \max(hx', -px') + \frac{2ph}{p+h}Amax(h(xˉ+A),p(−xˉ+A))=max(hx′,−px′)+p+h2ph​A.
  • Lemma 3.1(b): for nonnegative demand, the nominal problem without fixed cost is solved by ordering up to Sk=wkS_k = w_kSk​=wk​.
  • Remark 1: Ak−1≤AkA_{k-1} \le A_kAk−1​≤Ak​, so wk′≥wˉkw'_k \ge \bar w_kwk′​≥wˉk​ when p≥hp \ge hp≥h.
  • Remark 3: under i.i.d. demand, Ak=w^ΓkA_k = \hat w\Gamma_kAk​=w^Γk​, which gives the closed-form thresholds.

Significance

The theorem reduces robust inventory control to nominal inventory control. Every structural fact known for the deterministic problem then transfers to the robust one. These include the optimality of base-stock policies without fixed cost and the threshold structure with a fixed cost. The robust base-stock levels are explicit: they shift the nominal levels by p−hp+h(Ak−Ak−1)\frac{p-h}{p+h}(A_k - A_{k-1})p+hp−h​(Ak​−Ak−1​), upward when shortage is more expensive than holding. The extra cost 2php+h∑kAk\frac{2ph}{p+h}\sum_k A_kp+h2ph​∑k​Ak​ quantifies the price of protection as a function of the budgets. The paper uses the same reduction for capacitated orders (Theorem 3.3) and for supply networks (§4).

The result has been proved since 2006, and no machine-checked proof of it, or of any budgeted robust counterpart, is known to exist. This mission provides several formalizations for reuse:

  • the budgeted robust counterpart of a pair of piecewise-linear constraints;
  • the duality of the fractional knapsack LP (13);
  • the optimality of base-stock orders for a deterministic backlogged inventory problem.

Difficulty

The algebraic core, identity (23), is elementary. The work is in the reductions around it. The first is that (14) really is the worst case of (10)–(11): this needs strong duality for (13), together with attainment on both sides, and the observation that the minimizing and maximizing deviations of a constraint pair differ. The second is that the minimum of (14) over the auxiliary variables, for fixed orders, is (21). This requires the dual optimum to be attained with the value of (13), and h,p≥0h, p \ge 0h,p≥0 so that the cost is monotone in AkA_kAk​. The third is the base-stock part, which needs Lemma 3.1(b), a global optimality statement for a TTT-period problem with backlogging. The paper proves that lemma by an explicit dual certificate. An argument through first-order conditions in each period is not enough, because orders in one period affect every later stock level.

Formalization scope

All data are real numbers and sequences are ℕ → ℝ; only indices k<Tk < Tk<T matter. stock w u k denotes xk+1x_{k+1}xk+1​, the stock at the end of period kkk. The standing assumptions of §3.1 are fields of the model:

  • c>0c > 0c>0, K≥0K \ge 0K≥0, h≥0h \ge 0h≥0, p>cp > cp>c;
  • w^k≥0\hat w_k \ge 0w^k​≥0;
  • Γ0≥0\Gamma_0 \ge 0Γ0​≥0 and Γk≤Γk+1≤Γk+1\Gamma_k \le \Gamma_{k+1} \le \Gamma_k + 1Γk​≤Γk+1​≤Γk​+1.

The conventions and corrections are:

  • Fixed cost. The paper writes it with binary variables and a big-MMM constraint. Here it is the indicator C(uk)C(u_k)C(uk​) in the objective, as in the paper's own (21).
  • "Optimal". It always means minimality over all feasible points.
  • The policy. It is the order sequence chosen at time 0.
  • AkA_kAk​. It is the value of (13), as in Remark 1 after the theorem, not "the optimal q∗,r∗q^*, r^*q∗,r∗ of (14)", which need not be unique.
  • The robust formulation. (14) is formalized as printed: the kkk-th constraint pair is protected by the budget Γk\Gamma_kΓk​ alone, not by the intersection of all budgets up to kkk.
  • Sign slip. The page's xk+1=xˉk+1+∑w^izix_{k+1} = \bar x_{k+1} + \sum \hat w_i z_ixk+1​=xˉk+1​+∑w^i​zi​ is a sign slip for xˉk+1−∑w^izi\bar x_{k+1} - \sum \hat w_i z_ixˉk+1​−∑w^i​zi​. The formal statements use the correct sign; the result is unaffected because the deviation set is symmetric.
  • Part (b). It is stated under wk′≥0w'_k \ge 0wk′​≥0. As printed it fails when p<hp < hp<h makes w′w'w′ negative, for example T=2T = 2T=2, wˉ=w^=(10,0)\bar w = \hat w = (10, 0)wˉ=w^=(10,0), Γ=(0,1)\Gamma = (0, 1)Γ=(0,1), c=1c = 1c=1, h=4h = 4h=4, p=2p = 2p=2, x0=0x_0 = 0x0​=0. Lemma 3.1(b) carries the matching hypothesis of nonnegative demand.
  • Remark 3. It adds Γ0≤1\Gamma_0 \le 1Γ0​≤1.
  • Remark 1. Its inequalities are weak.
  • Not stated. The (s, S) clause of (a) and part (c) are excluded. They rest on a stochastic theorem cited from Bertsekas (1995) and on thresholds stated through the optimal ordering times.

A trivializing formalization is ruled out: the robust cost is formulation (14) with its variables y,q,ry, q, ry,q,r, and AkA_kAk​ is the value of LP (13). Neither is the closed-form objective (21) nor an arbitrary sequence. Contributions are welcome on the LP duality of (13) (a fractional knapsack), on the robust counterpart milestone, and on Lemma 3.1(b), each of which is independent of the others.

Selected references

  • D. Bertsimas, A. Thiele, A Robust Optimization Approach to Inventory Theory, Operations Research 54(1):150–168, 2006. https://doi.org/10.1287/opre.1050.0238
  • D. Bertsimas, M. Sim, The Price of Robustness, Operations Research 52(1):35–53, 2004. https://doi.org/10.1287/opre.1030.0065
  • A. Ben-Tal, A. Nemirovski, Robust solutions of uncertain linear programs, Operations Research Letters 25(1):1–13, 1999. https://doi.org/10.1016/S0167-6377(99)00016-4
  • D. P. Bertsekas, Dynamic Programming and Optimal Control, Vol. 1, Athena Scientific, 1995.
12 thms1 active userReviewed
🏆Completed
Convex OptimizationMachine Learning·Captain: mikedeng1

Convex Optimization: Algorithms and Complexity VII: Conditional Gradient Descent (Frank–Wolfe) with γ_s = 2/(s + 1) Has Rate 2βR²/(t + 1) in Any NormTextbook

Motivation

Many constrained optimization problems in machine learning and statistics have a feasible set X\mathcal XX over which a linear function is cheap to minimize but a Euclidean projection is expensive: the ℓ1\ell_1ℓ1​-ball, the simplex, the nuclear-norm ball, the convex hull of a combinatorial family. Projected gradient descent needs a projection at every step. Conditional gradient descent, introduced by Frank and Wolfe in 1956 for quadratic programming, replaces the projection by a call to a linear minimization oracle over X\mathcal XX, and its iterates are convex combinations of oracle outputs, which makes them sparse when X\mathcal XX is a polytope.

This mission is the seventh of a series formalizing S. Bubeck's monograph Convex Optimization: Algorithms and Complexity (Foundations and Trends in Machine Learning, 2015, arXiv:1405.4980v2). It covers Section 3.3, whose main result, Theorem 3.8, is the O(1/t)O(1/t)O(1/t) rate of the method in the form given by Jaggi (2013), going back to Dunn and Harshbarger (1978).

Setting

Let EEE be a finite-dimensional real vector space with an arbitrary norm ∥⋅∥\|\cdot\|∥⋅∥. For a linear form ggg on EEE, written v↦g⊤vv\mapsto g^\top vv↦g⊤v, the dual norm is ∥g∥∗=sup⁡∥v∥≤1g⊤v\|g\|_*=\sup_{\|v\|\le1}g^\top v∥g∥∗​=sup∥v∥≤1​g⊤v. Let X⊆E\mathcal X\subseteq EX⊆E be nonempty, compact and convex, with diameter R=sup⁡x,y∈X∥x−y∥R=\sup_{x,y\in\mathcal X}\|x-y\|R=supx,y∈X​∥x−y∥.

Let f:E→Rf:E\to\mathbb Rf:E→R be differentiable with gradient ∇f(x)\nabla f(x)∇f(x), a linear form on EEE. For β≥0\beta\ge0β≥0, fff is β-smooth with respect to ∥⋅∥\|\cdot\|∥⋅∥ on X\mathcal XX if

∥∇f(x)−∇f(y)∥∗≤β∥x−y∥(x,y∈X).\|\nabla f(x)-\nabla f(y)\|_*\le\beta\|x-y\|\qquad(x,y\in\mathcal X).∥∇f(x)−∇f(y)∥∗​≤β∥x−y∥(x,y∈X).

A point x∗∈Xx^*\in\mathcal Xx∗∈X with f(x∗)=min⁡x∈Xf(x)f(x^*)=\min_{x\in\mathcal X}f(x)f(x∗)=minx∈X​f(x) is fixed throughout, and δt=f(xt)−f(x∗)\delta_t=f(x_t)-f(x^*)δt​=f(xt​)−f(x∗).

Given step sizes (γs)s≥1(\gamma_s)_{s\ge1}(γs​)s≥1​, a run of conditional gradient descent is a pair of sequences with x1∈Xx_1\in\mathcal Xx1​∈X and, for every t≥1t\ge1t≥1,

yt∈argmin⁡y∈X∇f(xt)⊤y(3.8),xt+1=(1−γt)xt+γtyt(3.9).y_t\in\operatorname*{argmin}_{y\in\mathcal X}\nabla f(x_t)^\top y\quad(3.8),\qquad x_{t+1}=(1-\gamma_t)x_t+\gamma_ty_t\quad(3.9).yt​∈y∈Xargmin​∇f(xt​)⊤y(3.8),xt+1​=(1−γt​)xt​+γt​yt​(3.9).

The minimizer yty_tyt​ need not be unique; any choice is allowed.

Formalization targets

Goal: Theorem 3.8 (p. 272)

If fff is convex and β\betaβ-smooth with respect to ∥⋅∥\|\cdot\|∥⋅∥ and γs=2s+1\gamma_s=\frac{2}{s+1}γs​=s+12​ for s≥1s\ge1s≥1, then every run satisfies, for every t≥2t\ge2t≥2,

f(xt)−f(x∗)≤2βR2t+1.f(x_t)-f(x^*)\le\frac{2\beta R^2}{t+1}.f(xt​)−f(x∗)≤t+12βR2​.

Milestones

  1. Inequality (3.4) in an arbitrary norm (p. 267, used on p. 272): for x,y∈Xx,y\in\mathcal Xx,y∈X, 0≤f(x)−f(y)−∇f(y)⊤(x−y)≤β2∥x−y∥20\le f(x)-f(y)-\nabla f(y)^\top(x-y)\le\frac{\beta}{2}\|x-y\|^20≤f(x)−f(y)−∇f(y)⊤(x−y)≤2β​∥x−y∥2.
  2. The one-step recursion (pp. 272–273): for any run with γs∈[0,1]\gamma_s\in[0,1]γs​∈[0,1],
δs+1≤(1−γs)δs+β2γs2R2.\delta_{s+1}\le(1-\gamma_s)\delta_s+\frac{\beta}{2}\gamma_s^2R^2 .δs+1​≤(1−γs​)δs​+2β​γs2​R2.
  1. Initialization (p. 273): if γ1=1\gamma_1=1γ1​=1, then δ2≤β2R2\delta_2\le\frac{\beta}{2}R^2δ2​≤2β​R2.
  2. The induction (p. 273): a real sequence with δ2≤β2R2\delta_2\le\frac{\beta}{2}R^2δ2​≤2β​R2 and the recursion of milestone 2 with γs=2s+1\gamma_s=\frac{2}{s+1}γs​=s+12​ for s≥2s\ge2s≥2 satisfies δt≤2βR2t+1\delta_t\le\frac{2\beta R^2}{t+1}δt​≤t+12βR2​ for t≥2t\ge2t≥2.

Significance

Theorem 3.8 is the basic guarantee for projection-free first-order optimization. Its rate does not depend on the dimension, and it depends on the geometry only through the product βR2\beta R^2βR2, both measured in the same norm, which may be chosen to fit X\mathcal XX: for the ℓ1\ell_1ℓ1​-ball, smoothness in ∥⋅∥1\|\cdot\|_1∥⋅∥1​ with dual norm ∥⋅∥∞\|\cdot\|_\infty∥⋅∥∞​ gives much better constants than the Euclidean analysis. The book applies it this way to a LASSO-type problem (Section 3.3, pp. 274–276), and the same bound underlies the sparse-approximation corollary on the simplex (p. 273) and the many later variants of the method (away steps, stochastic and online conditional gradient).

The result is classical and proved in the book. What this mission adds is a machine-checked statement and proof in an arbitrary finite-dimensional normed space, with the dual norm as the operator norm on linear forms, and a reusable encoding of norm-smoothness and of conditional gradient runs. On Prove2Me a related result is already proved: Lan's Theorem 7.1 (First-order and Stochastic Optimization Methods), which bounds f(yk)−f∗f(y_k)-f^*f(yk​)−f∗ by 2Lk(k+1)∑i≤k∥xi−yi−1∥2\frac{2L}{k(k+1)}\sum_{i\le k}\|x_i-y_{i-1}\|^2k(k+1)2L​∑i≤k​∥xi​−yi−1​∥2 with a different indexing; after the diameter bound it yields 2βR2/t2\beta R^2/t2βR2/t at Bubeck's iterate xtx_txt​, which is weaker than Theorem 3.8 by one step.

Difficulty

The difficulty is in the bookkeeping of norms and indices, not in a deep idea. Inequality (3.4) is proved in the book only for the Euclidean norm, where ∇f(x)∈Rn\nabla f(x)\in\mathbb R^n∇f(x)∈Rn and the Cauchy–Schwarz inequality is used; in a general norm it needs the pairing between a linear form and a vector and the bound ∣g⊤v∣≤∥g∥∗∥v∥|g^\top v|\le\|g\|_*\|v\|∣g⊤v∣≤∥g∥∗​∥v∥. The rate 2βR2/(t+1)2\beta R^2/(t+1)2βR2/(t+1) is attained only by starting the induction at t=2t=2t=2, where the step γ1=1\gamma_1=1γ1​=1 erases the dependence on the starting point; at t=1t=1t=1 the bound can fail, since δ1\delta_1δ1​ is arbitrary. A first attempt that runs the induction from t=1t=1t=1 with an arbitrary δ1\delta_1δ1​ does not give the stated constant.

Formalization scope

EEE is a type with [NormedAddCommGroup E] [NormedSpace ℝ E] [FiniteDimensional ℝ E]; nothing is specialised to the Euclidean norm. The gradient is an explicit derivative map f' : E → (E →L[ℝ] ℝ) with HasFDerivAt f (f' x) x for every x; ∇f(x)⊤v\nabla f(x)^\top v∇f(x)⊤v is f' x v, and ∥⋅∥∗\|\cdot\|_*∥⋅∥∗​ is the operator norm, which equals sup⁡∥v∥≤1g⊤v\sup_{\|v\|\le1}g^\top vsup∥v∥≤1​g⊤v. RRR is Metric.diam X, which equals the supremum of ∥x−y∥\|x-y\|∥x−y∥ over X\mathcal XX because X\mathcal XX is compact. Sequences are indexed by ℕ with the first iterate at index 1.

Committed conventions: X\mathcal XX compact, convex and containing x∗x^*x∗ (hence nonempty); convexity of fff and the Lipschitz bound on the gradient are assumed on X\mathcal XX only, which is weaker than the book's global assumptions; β≥0\beta\ge0β≥0; the existence of the minimizer x∗x^*x∗ is the book's standing assumption (p. 242); the conclusion is stated for t≥2t\ge2t≥2, as in the book. Runs are predicates: yty_tyt​ is any minimizer of the linear form over X\mathcal XX and yt∈Xy_t\in\mathcal Xyt​∈X is required, so the goal quantifies over every run with γs=2/(s+1)\gamma_s=2/(s+1)γs​=2/(s+1). A formalization in which yty_tyt​ need not lie in X\mathcal XX, or in which smoothness is assumed only along the iterates, states a different theorem and is excluded.

A complete development needs the descent inequality (3.4) for Fréchet derivatives in a normed space (reusable for every smooth method in the series and beyond), the fact that the iterates stay in X\mathcal XX, and a scalar induction. Proofs of the milestones are welcome independently; milestone 4 is a statement about real sequences only.

Selected references

  • S. Bubeck, Convex Optimization: Algorithms and Complexity, Foundations and Trends in Machine Learning 8(3–4):231–358, 2015. https://arxiv.org/abs/1405.4980
  • M. Frank and P. Wolfe, An algorithm for quadratic programming, Naval Research Logistics Quarterly 3(1–2):95–110, 1956. https://doi.org/10.1002/nav.3800030109
  • J. C. Dunn and S. Harshbarger, Conditional gradient algorithms with open loop step size rules, Journal of Mathematical Analysis and Applications 62(2):432–444, 1978. https://doi.org/10.1016/0022-247X(78)90137-3
  • M. Jaggi, Revisiting Frank–Wolfe: projection-free sparse convex optimization, Proceedings of ICML 2013, PMLR 28(1):427–435. https://proceedings.mlr.press/v28/jaggi13.html
  • G. Lan, First-order and Stochastic Optimization Methods for Machine Learning, Springer, 2020, Theorem 7.1. https://doi.org/10.1007/978-3-030-39568-1
6 thms1 active userReviewed
Convex OptimizationMachine Learning·Captain: mikedeng1

Convex Optimization: Algorithms and Complexity VI: Gradient Descent with η = 2/(α + β) on a β-Smooth α-Strongly Convex Function Has Rate (β/2)exp(−4t/(κ + 1))‖x₁ − x*‖²Textbook

Motivation

Gradient descent is a basic method for minimizing a differentiable function when evaluating its gradient is practical but solving the optimization problem directly is not. The rate at which its iterates approach an optimizer depends on the assumptions about the function. For a convex function with a Lipschitz gradient, the value error decreases at a sublinear rate. Adding strong convexity changes the behavior: the distance from the optimizer contracts at each step, giving an exponential bound on the value error. This section of Bubeck's monograph identifies a fixed step size that uses both the smoothness and curvature constants and gives the corresponding rate.

The result matters when a high-accuracy answer is needed. A sublinear bound makes each extra digit progressively more expensive; an exponential bound says that a fixed number of additional gradient evaluations reduces the error by a fixed factor. The theorem is a textbook result, already proved mathematically. This mission asks for its precise machine-checked statement and the source's supporting inequalities, rather than for a new optimization method.

Setting

Work in Euclidean space Rn\mathbb R^nRn with n≥1n\ge1n≥1, equipped with its usual inner product and norm. A differentiable function f:Rn→Rf:\mathbb R^n\to\mathbb Rf:Rn→R has gradient g(x)=∇f(x)g(x)=\nabla f(x)g(x)=∇f(x). It is β\betaβ-smooth when its gradient is β\betaβ-Lipschitz: ∥g(x)−g(y)∥≤β∥x−y∥\|g(x)-g(y)\|\le\beta\|x-y\|∥g(x)−g(y)∥≤β∥x−y∥ for every x,yx,yx,y. It is α\alphaα-strongly convex when, for every x,yx,yx,y,

f(y)≥f(x)+⟨g(x),y−x⟩+α2∥y−x∥2.f(y)\ge f(x)+\langle g(x),y-x\rangle+\frac\alpha2\|y-x\|^2.f(y)≥f(x)+⟨g(x),y−x⟩+2α​∥y−x∥2.

The first condition limits how rapidly the gradient changes. The second gives a quadratic lower bound on the function around any point. Here α>0\alpha>0α>0 and β≥0\beta\ge0β≥0. In positive dimension, the two conditions together entail β≥α\beta\ge\alphaβ≥α, so the condition number κ=β/α\kappa=\beta/\alphaκ=β/α is at least one. The case α=β\alpha=\betaα=β remains part of the target.

A point x∗x^*x∗ is a global minimizer when f(x∗)≤f(y)f(x^*)\le f(y)f(x∗)≤f(y) for every yyy. The book assumes such a point exists as a standing convention. A gradient descent run is a sequence (xt)t≥1(x_t)_{t\ge1}(xt​)t≥1​ satisfying xt+1=xt−ηg(xt)x_{t+1}=x_t-\eta g(x_t)xt+1​=xt​−ηg(xt​) at each positive index. Its first iterate x1x_1x1​ is arbitrary. The step size in this mission is fixed at η=2/(α+β)\eta=2/(\alpha+\beta)η=2/(α+β), rather than chosen by line search or adapted along the run.

Formalization targets

The central target is Theorem 3.12 of Bubeck, p. 279. For every integer t≥0t\ge0t≥0, the gradient descent run satisfies

f(xt+1)−f(x∗)≤β2exp⁡ ⁣(−4tκ+1)∥x1−x∗∥2.f(x_{t+1})-f(x^*)\le \frac\beta2\exp\!\left(-\frac{4t}{\kappa+1}\right)\|x_1-x^*\|^2.f(xt+1​)−f(x∗)≤2β​exp(−κ+14t​)∥x1​−x∗∥2.

At t=0t=0t=0 this is a smoothness bound on the initial value gap. For subsequent iterations it gives a linear convergence rate with the explicit exponential factor stated in the book. No initial-radius bound or bounded domain is imposed: the actual squared distance ∥x1−x∗∥2\|x_1-x^*\|^2∥x1​−x∗∥2 appears in the conclusion.

The milestones trace the mathematical claims stated in the source. Equation (3.6) is the co-coercivity inequality for gradients of convex smooth functions. The proof of Lemma 3.11 introduces ϕ(z)=f(z)−(α/2)∥z∥2\phi(z)=f(z)-(\alpha/2)\|z\|^2ϕ(z)=f(z)−(α/2)∥z∥2, and identifies it as convex and (β−α)(\beta-\alpha)(β−α)-smooth. Lemma 3.11 combines curvature and smoothness into a sharper inequality for two gradients. The proof of Theorem 3.12 then gives a value-gap bound, a one-step distance contraction, and its iterated exponential form. These statements are separately useful: the co-coercivity and contraction bounds can be reused in analyses of related first-order methods.

Significance

The theorem states a complete guarantee for the algorithm: an explicit rule, the hypotheses on the objective, and a bound valid for every iteration count. It makes the role of κ\kappaκ visible. When κ\kappaκ is close to one, the contraction is strong; when the smoothness constant is much larger than the curvature constant, more iterations are needed for the same error reduction. The stated dependence supports comparisons with projected and accelerated gradient methods elsewhere in the same monograph.

Formalizing the result requires a common interface for actual gradients, smoothness, strong convexity, and algorithm runs. The strong-convexity predicate is an existing published definition, while the local smoothness and run definitions use Bubeck's conventions. Once these interfaces and the inequalities are proved, later missions can use the resulting declarations to compare rates without translating between informal meanings of “smooth” or changing the iterate index. The source provides a mathematical proof; these draft Lean theorems carry sorry and do not yet constitute machine-checked proofs.

Difficulty

The main issue is getting the sharp contraction factor from two assumptions that control different parts of the gradient step. A direct Lipschitz estimate on the update map does not by itself express the mixed inner-product term with the constants needed for the stated factor. Lemma 3.11 is the source's precise bridge between the gradient difference, the point displacement, and their inner product. The case α=β\alpha=\betaα=β also needs to remain valid: a proof route that divides by β−α\beta-\alphaβ−α cannot cover that boundary by the same calculation.

The last display in the proof of Theorem 3.12 joins a one-step inequality involving xtx_txt​ to an exponential inequality involving x1x_1x1​. The latter is the cumulative statement after ttt steps. Keeping these as separate milestones makes each quantified claim explicit while preserving the theorem's bound.

Formalization scope

Lean represents Rn\mathbb R^nRn as EuclideanSpace ℝ (Fin n), with n>0n>0n>0. The gradient is an explicit map ggg required to be the actual gradient of fff at every point. Smoothness is the gradient Lipschitz condition, not a quadratic upper bound used as a definition. Strong convexity uses the published OnlineConvexOpt.ConvexBasics.StronglyConvexOn predicate on the whole space; its formula is the book's (3.13). Iterates are indexed from one, and index zero imposes no condition. The norm, inner product, constants, and real exponential follow the printed formulas.

The added explicit conditions are n>0n>0n>0, α>0\alpha>0α>0, and β>0\beta>0β>0 where Equation (3.6) divides by β\betaβ. Positive dimension excludes a degenerate space where curvature imposes no restriction on smoothness. The positivity of α\alphaα makes κ\kappaκ meaningful; the source treats it as a positive strong-convexity parameter. The minimizer hypothesis is the book's standing convention. There is no assumption that g(x∗)=0g(x^*)=0g(x∗)=0: that property follows from global minimality and differentiability. A gradient map unrelated to fff would trivialize the model, so the smoothness definition includes the gradient identity.

The local development needs Euclidean inner-product identities, convexity, differentiability, Lipschitz gradient bounds, and real exponential estimates. The auxiliary function and the two co-coercivity inequalities are reusable outside this chapter. Contributions should prove the exact milestone statements and the final theorem, including t=0t=0t=0 and α=β\alpha=\betaα=β, without weakening constants or substituting another gradient descent step.

Selected references

  • Sébastien Bubeck, Convex Optimization: Algorithms and Complexity, Foundations and Trends in Machine Learning 8(3–4):231–358, 2015. arXiv:1405.4980v2; DOI:10.1561/2200000050.
9 thms1 active userReviewed
Dynamic ProgrammingMarkov ChainOperations Research·Captain: mikedeng1

An Analysis of Stochastic Shortest Path Problems: If Every Improper Policy Has Infinite Cost, the Optimal Cost Is the Unique Fixed Point of Bellman's Operator and Value Iteration Converges to ItResearch Paper

Motivation

A shortest path problem asks how to reach a destination at minimum cost. In a stochastic shortest path problem, a decision at a state selects a probability distribution over successor states, so both the route and its total cost are random. Costs may have either sign. This makes the problem relevant to finite-state control models where rewards and expenses occur before eventual termination. Bertsekas and Tsitsiklis analyze this setting without requiring all one-stage costs to be nonnegative or all to be nonpositive. Their condition instead rules out an improper stationary policy whose costs stay finite from every initial state. Bertsekas and Tsitsiklis (1991), pp. 580–583.

Earlier treatments established Bellman-equation and algorithmic conclusions under positive or nonnegative costs. The 1991 paper traces the finite-control development from Eaton and Zadeh and the compact-control extension from Kushner, then removes the sign restriction while retaining finite state space. It also explains why the Bellman mapping need not contract when an improper policy is available. Bertsekas and Tsitsiklis (1991), pp. 581, 585. A separate, proved Prove2Me theorem from Bertsekas's textbook treats the special case in which every policy is proper and controls are finite; the result here permits improper policies and compact control spaces.

Setting

There are n≥1n\ge1n≥1 states. State 111 is the destination. At state iii, a control u∈U(i)u\in U(i)u∈U(i) incurs a real cost ci(u)c_i(u)ci​(u) and moves the process to state jjj with probability pij(u)p_{ij}(u)pij​(u). A selector μ\muμ chooses one control μ(i)\mu(i)μ(i) at every state. A policy π=(μ0,μ1,…)\pi=(\mu_0,\mu_1,\ldots)π=(μ0​,μ1​,…) may change selectors over time; a stationary policy repeats one selector. The matrix P(μ)P(\mu)P(μ) has entries pij(μ(i))p_{ij}(\mu(i))pij​(μ(i)), and c(μ)c(\mu)c(μ) is the vector of one-stage costs. Bertsekas and Tsitsiklis (1991), p. 582.

The cost vector x(π)x(\pi)x(π) is the coordinatewise limit inferior of expected partial costs, including the possibility of infinite values. The optimal cost xi∗x_i^*xi∗​ is the infimum of xi(π)x_i(\pi)xi​(π) over all policies, including nonstationary ones. Optimality of a policy means it attains that infimum from every initial state. The fixed-selector operator is Tμ(x)=c(μ)+P(μ)xT_\mu(x)=c(\mu)+P(\mu)xTμ​(x)=c(μ)+P(μ)x; the Bellman operator TTT takes the coordinatewise infimum of these vectors over selectors. Bertsekas and Tsitsiklis (1991), pp. 582–583, equations (2)–(6).

A stationary policy is proper when its probability of reaching state 111 tends to one from every starting state. Assumption 1 says state 111 is absorbing and cost-free, at least one proper stationary policy exists, and every improper stationary policy has a partial-cost coordinate tending to +∞+\infty+∞. Assumption 2 makes each control space compact, each cost function lower semicontinuous, and each transition-probability coordinate continuous. Work takes place in X={x∈Rn:x1=0}X=\{x\in\mathbb R^n:x_1=0\}X={x∈Rn:x1​=0}. Bertsekas and Tsitsiklis (1991), pp. 583–584.

Formalization targets

Proposition 2: Bellman's equation and value iteration

Under Assumptions 1 and 2, the optimal cost is finite and is the unique fixed point of TTT in XXX. Every initial x∈Xx\in Xx∈X has

lim⁡t→∞Tt(x)=x∗.\lim_{t\to\infty}T^t(x)=x^*.t→∞lim​Tt(x)=x∗.

A stationary selector μ\muμ is optimal exactly when Tμ(x∗)=T(x∗)T_\mu(x^*)=T(x^*)Tμ​(x∗)=T(x∗); an optimal proper stationary selector exists. These are all clauses of Proposition 2, rather than separate targets selected from it. Bertsekas and Tsitsiklis (1991), p. 586, Proposition 2.

Supporting results

The milestone list follows the source's Proposition 1, Lemmas 1–3, and the numbered equations used by their proof. Proposition 1 supplies a weighted maximum-norm contraction when every stationary policy is proper. Lemma 1 describes fixed costs of a proper policy and characterizes properness through a Bellman inequality. Lemma 2 gives continuity of TTT. Lemma 3 controls limits of proper policies; its second part detects a limit that becomes improper through diverging costs. Appendix equations (22) and (24), and the policy-improvement equation (15), provide the paper's intermediate targets. Bertsekas and Tsitsiklis (1991), pp. 585–587, 591–592.

Significance

Proposition 2 identifies the cost of the best policy by a finite-dimensional Bellman equation even though the definition of optimal cost ranges over all, possibly nonstationary, policies. Its convergence clause justifies value iteration from any vector whose destination coordinate is zero. Its policy criterion and existence clause connect a fixed point to an implementable stationary decision rule. The paper applies these conclusions to successive approximation and policy iteration in §4. Bertsekas and Tsitsiklis (1991), pp. 586, 590–591.

The paper proves these statements. This mission seeks machine-checked proofs of their exact finite-state formulation and of the listed intermediate results. The resulting finite stochastic-matrix, hitting, and policy-cost infrastructure can also support other undiscounted control problems. The related proved Prove2Me result for the all-proper finite-control case does not settle this mission's compact-control or improper-policy cases.

Difficulty

When every stationary policy is proper, a common weighted maximum norm makes the Bellman mapping contract. An improper policy can destroy this route: Figure 2 has an absorbing destination and satisfies both assumptions, yet T(0,x2)=(0,min⁡{1+x2,2})T(0,x_2)=(0,\min\{1+x_2,2\})T(0,x2​)=(0,min{1+x2​,2}) is not a contraction in any norm on XXX. The main result therefore needs a way to retain fixed-point uniqueness and convergence without a uniform contraction rate. Compactness matters because a sequence of proper selectors may converge to an improper selector; Lemma 3 explains the associated cost behavior. Bertsekas and Tsitsiklis (1991), pp. 585–586, 591–595.

Formalization scope

States are Fin n, with the paper's state 111 represented by 0; the model requires n≥1n\ge1n≥1. A control set U(i)U(i)U(i) is a type with a metric, with compactness asserted for its whole carrier. Transition rows explicitly have nonnegative entries summing to one. These are the probability-vector conditions implicit in the word “probability.” Assumption 1 supplies a selector and therefore nonempty control sets. The paper writes T:Rn→RnT:\mathbb R^n\to\mathbb R^nT:Rn→Rn; where a theorem uses only Assumption 1, finite real Bellman infima are made explicit through TRealValued. Under Assumption 2, compactness and lower semicontinuity give real attained minima. Bertsekas and Tsitsiklis (1991), pp. 582–584.

Policy costs and their infimum use extended reals so that an improper policy's +∞+\infty+∞ cost is represented without a default finite value. Proposition 2 concludes, rather than assumes, that x∗x^*x∗ has real coordinates. Its policy infimum ranges over every sequence of selectors. Matrix products use the identity at time zero; T0T^0T0 is the identity. The destination-zero restriction is retained in every fixed-point and iteration claim. Defining optimal cost as a Bellman fixed point, or restricting the infimum to stationary policies, would remove the result the paper proves. Solvers can contribute the finite-chain, matrix-inverse, semicontinuity, and convergence arguments needed by these targets.

Selected references

  • Dimitri P. Bertsekas and John N. Tsitsiklis, An Analysis of Stochastic Shortest Path Problems, Mathematics of Operations Research 16(3), 580–595, 1991. DOI.
11 thms1 active userReviewed
🏆Completed
Convex OptimizationMachine Learning·Captain: mikedeng1

Convex Optimization: Algorithms and Complexity IV: Gradient Descent on a Convex β-Smooth Function Has Rate 2β‖x₁ − x*‖²/(t − 1)Textbook

Motivation

Gradient descent goes back to Cauchy (1847). It is the simplest method for minimizing a differentiable function, and most of the first-order methods in large-scale optimization and machine learning are variants of it. Its appeal in high dimension is that its oracle complexity, the number of gradient evaluations needed to reach a given accuracy, can be bounded independently of the dimension. For a merely Lipschitz convex function the projected subgradient method needs on the order of 1/ε21/\varepsilon^21/ε2 steps to reach accuracy ε\varepsilonε (Theorem 3.2 of the book). Under a smoothness assumption gradient descent does much better, because the gradients shrink near the optimum and the steps adapt automatically.

This mission formalizes the basic result of that kind: Theorem 3.3 of S. Bubeck, Convex Optimization: Algorithms and Complexity (Foundations and Trends in Machine Learning 8(3–4), 2015; arXiv:1405.4980v2). It states that gradient descent with step size 1/β1/\beta1/β on a convex β\betaβ-smooth function on Rn\mathbb R^nRn has optimality gap O(1/t)O(1/t)O(1/t) after ttt steps. The result and its proof are standard; versions appear in Nesterov's Introductory Lectures on Convex Optimization (2004, §2.1.5). It is the fourth mission in a series that formalizes the capstone results of Bubeck's monograph.

Setting

Write Rn\mathbb R^nRn for Euclidean space with inner product x⊤yx^\top yx⊤y and norm ∥⋅∥\|\cdot\|∥⋅∥. Let f:Rn→Rf:\mathbb R^n\to\mathbb Rf:Rn→R be differentiable with gradient ∇f\nabla f∇f.

  • fff is convex if f((1−λ)x+λy)≤(1−λ)f(x)+λf(y)f((1-\lambda)x+\lambda y)\le(1-\lambda)f(x)+\lambda f(y)f((1−λ)x+λy)≤(1−λ)f(x)+λf(y) for all x,yx,yx,y and λ∈[0,1]\lambda\in[0,1]λ∈[0,1].
  • For β≥0\beta\ge0β≥0, fff is β\betaβ-smooth if its gradient is β\betaβ-Lipschitz:
∥∇f(x)−∇f(y)∥≤β∥x−y∥for all x,y∈Rn.\|\nabla f(x)-\nabla f(y)\|\le\beta\|x-y\|\qquad\text{for all }x,y\in\mathbb R^n.∥∇f(x)−∇f(y)∥≤β∥x−y∥for all x,y∈Rn.
  • A minimizer is a point x∗x^*x∗ with f(x∗)≤f(y)f(x^*)\le f(y)f(x∗)≤f(y) for every yyy. Throughout the book a minimizer is assumed to exist.
  • Gradient descent with step size η>0\eta>0η>0, started at x1∈Rnx_1\in\mathbb R^nx1​∈Rn, is the sequence
xt+1=xt−η∇f(xt),t≥1.(3.1)x_{t+1}=x_t-\eta\nabla f(x_t),\qquad t\ge1. \tag{3.1}xt+1​=xt​−η∇f(xt​),t≥1.(3.1)

The optimality gaps are δs=f(xs)−f(x∗)≥0\delta_s=f(x_s)-f(x^*)\ge0δs​=f(xs​)−f(x∗)≥0.

Formalization targets

Goal: Theorem 3.3 (p. 267)

If fff is convex and β\betaβ-smooth with β>0\beta>0β>0, x∗x^*x∗ is a minimizer, and (xt)(x_t)(xt​) is gradient descent with η=1/β\eta=1/\betaη=1/β, then for every t≥2t\ge2t≥2

f(xt)−f(x∗)≤2β∥x1−x∗∥2t−1.f(x_t)-f(x^*)\le\frac{2\beta\|x_1-x^*\|^2}{t-1}.f(xt​)−f(x∗)≤t−12β∥x1​−x∗∥2​.

Milestones

These are the statements the book's proof uses, in the book's order:

  1. Lemma 3.4 (p. 267). For any β\betaβ-smooth fff, with no convexity: ∣f(x)−f(y)−∇f(y)⊤(x−y)∣≤β2∥x−y∥2|f(x)-f(y)-\nabla f(y)^\top(x-y)|\le\frac\beta2\|x-y\|^2∣f(x)−f(y)−∇f(y)⊤(x−y)∣≤2β​∥x−y∥2.
  2. (3.4) (p. 267). For convex β\betaβ-smooth fff: 0≤f(x)−f(y)−∇f(y)⊤(x−y)≤β2∥x−y∥20\le f(x)-f(y)-\nabla f(y)^\top(x-y)\le\frac\beta2\|x-y\|^20≤f(x)−f(y)−∇f(y)⊤(x−y)≤2β​∥x−y∥2.
  3. (3.5) (p. 267). For convex fff, one step of length 1/β1/\beta1/β decreases fff by at least 12β∥∇f(x)∥2\frac1{2\beta}\|\nabla f(x)\|^22β1​∥∇f(x)∥2.
  4. Lemma 3.5 (p. 268). If (3.4) holds, then f(x)−f(y)≤∇f(x)⊤(x−y)−12β∥∇f(x)−∇f(y)∥2f(x)-f(y)\le\nabla f(x)^\top(x-y)-\frac1{2\beta}\|\nabla f(x)-\nabla f(y)\|^2f(x)−f(y)≤∇f(x)⊤(x−y)−2β1​∥∇f(x)−∇f(y)∥2.
  5. (3.6) (p. 269). Co-coercivity: (∇f(x)−∇f(y))⊤(x−y)≥1β∥∇f(x)−∇f(y)∥2(\nabla f(x)-\nabla f(y))^\top(x-y)\ge\frac1\beta\|\nabla f(x)-\nabla f(y)\|^2(∇f(x)−∇f(y))⊤(x−y)≥β1​∥∇f(x)−∇f(y)∥2.
  6. Distances decrease (proof of Theorem 3.3, p. 269). ∥xs+1−x∗∥≤∥xs−x∗∥\|x_{s+1}-x^*\|\le\|x_s-x^*\|∥xs+1​−x∗∥≤∥xs​−x∗∥ for every s≥1s\ge1s≥1.
  7. The recursion (p. 268). δs+1≤δs−12β∥x1−x∗∥2δs2\delta_{s+1}\le\delta_s-\frac{1}{2\beta\|x_1-x^*\|^2}\delta_s^2δs+1​≤δs​−2β∥x1​−x∗∥21​δs2​.
  8. From the recursion to the rate (p. 269). For ω>0\omega>0ω>0 and non-negative reals, ωδs2+δs+1≤δs\omega\delta_s^2+\delta_{s+1}\le\delta_sωδs2​+δs+1​≤δs​ for all s≥1s\ge 1s≥1 implies 1/δt≥ω(t−1)1/\delta_t\ge\omega(t-1)1/δt​≥ω(t−1).

Stronger companion: footnote 4 (p. 269)

Under the same hypotheses, f(xt)−f(x∗)≤2β∥x1−x∗∥2/(t+3)f(x_t)-f(x^*)\le 2\beta\|x_1-x^*\|^2/(t+3)f(xt​)−f(x∗)≤2β∥x1​−x∗∥2/(t+3) for every t≥1t\ge1t≥1.

Significance

Theorem 3.3 is the reference rate for first-order methods on smooth convex problems. Several later results in the book are measured against it. Nesterov's accelerated gradient descent (§3.7) improves 1/t1/t1/t to 1/t21/t^21/t2, the lower bounds of §3.5 show that 1/t21/t^21/t2 cannot be beaten by any black-box first-order method, and adding strong convexity (§3.4) upgrades 1/t1/t1/t to a linear rate. Its ingredients are reused throughout the book and the optimization literature: the descent lemma (Lemma 3.4), the one-step improvement (3.5) and co-coercivity (3.6). Co-coercivity is the finite-dimensional case of the Baillon–Haddad theorem.

The result is classical and fully proved on paper. To our knowledge Mathlib does not contain this rate for gradient descent on convex smooth functions, and no published Prove2Me theorem states it. The mission provides it, together with Lemma 3.4 and co-coercivity as reusable statements on EuclideanSpace ℝ (Fin n). These are the facts that later missions of the series on projected gradient descent, strong convexity and acceleration need. Formalizing the improved constant of footnote 4 is a welcome addition.

Difficulty

Two steps are not routine. The first is Lemma 3.4, where the book integrates the gradient along a segment. In Lean this needs a mean-value or fundamental-theorem-of-calculus argument for a function on a Euclidean space, together with a Cauchy–Schwarz estimate, and Mathlib's HasGradientAt interface has to be connected to one-variable derivatives along lines.

The second is the monotonicity of ∥xs−x∗∥\|x_s-x^*\|∥xs​−x∗∥. The natural first attempt is to telescope the one-step improvement (3.5) together with the convexity bound δs≤∥xs−x∗∥∥∇f(xs)∥\delta_s\le\|x_s-x^*\|\|\nabla f(x_s)\|δs​≤∥xs​−x∗∥∥∇f(xs​)∥. That gives a recursion involving ∥xs−x∗∥\|x_s-x^*\|∥xs​−x∗∥, which is not controlled by ∥x1−x∗∥\|x_1-x^*\|∥x1​−x∗∥ without further work. Making the recursion uniform requires co-coercivity (3.6), which comes from Lemma 3.5. That lemma's hypothesis is the two-sided inequality (3.4), not smoothness directly. The final numerical step divides by the gaps δs\delta_sδs​, so zero gaps need separate handling.

Formalization scope

  • Rn\mathbb R^nRn is EuclideanSpace ℝ (Fin n), and x⊤yx^\top yx⊤y is ⟪x, y⟫_ℝ.
  • The gradient is an explicit map g:Rn→Rng:\mathbb R^n\to\mathbb R^ng:Rn→Rn with the hypothesis ∀ x, HasGradientAt f (g x) x.
  • β\betaβ-smoothness (IsBetaSmooth f g β) is the book's definition: β≥0\beta\ge0β≥0, the gradient-existence hypothesis, and the Lipschitz bound ∥g(x)−g(y)∥≤β∥x−y∥\|g(x)-g(y)\|\le\beta\|x-y\|∥g(x)−g(y)∥≤β∥x−y∥. The book's "continuously differentiable" follows from it.
  • Convexity is Mathlib's ConvexOn ℝ Set.univ f. A minimizer is a point xstar with ∀ y, f xstar ≤ f y; its existence is the book's standing assumption, and ∇f(x∗)=0\nabla f(x^*)=0∇f(x∗)=0 is derived, not assumed.
  • A gradient-descent run (IsGDRun g η x) is a sequence x : ℕ → ℝⁿ with η>0\eta>0η>0 and xt+1=xt−ηg(xt)x_{t+1}=x_t-\eta g(x_t)xt+1​=xt​−ηg(xt​) for every t≥1t\ge1t≥1. The book's x1x_1x1​ is x 1, and index 000 is unused. Every theorem quantifies over all runs with η=1/β\eta=1/\betaη=1/β.

Disclosed side conditions:

  • β>0\beta>0β>0 wherever the page divides by β\betaβ. Convexity is retained for (3.5), as in its source context.
  • t≥2t\ge2t≥2 in the goal, where the page's bound has denominator t−1t-1t−1.
  • Statements in which the page divides by δs\delta_sδs​ or by ∥x1−x∗∥2\|x_1-x^*\|^2∥x1​−x∗∥2 are multiplied through, so that they stay true and meaningful when those quantities vanish.

A trivializing formalization is ruled out. Smoothness is the Lipschitz condition on the gradient, so Lemma 3.4 and (3.4) are not restatements of the definition, as they would be if the quadratic upper bound (3.4) were taken as the definition of smoothness. The run predicate also fixes the step size 1/β1/\beta1/β.

A complete development needs:

  • the segment integral or mean-value estimate for HasGradientAt functions;
  • first-order characterizations of convexity for differentiable functions;
  • elementary inner-product algebra.

Lemma 3.4, Lemma 3.5 and (3.6) are reusable well beyond this mission. Proofs of any milestone are welcome independently.

Selected references

  • S. Bubeck, Convex Optimization: Algorithms and Complexity, Foundations and Trends in Machine Learning 8(3–4):231–357, 2015. arXiv:1405.4980v2; §3.2, pp. 266–269.
  • A. Cauchy, Méthode générale pour la résolution des systèmes d'équations simultanées, C. R. Acad. Sci. Paris 25:536–538, 1847.
  • Y. Nesterov, Introductory Lectures on Convex Optimization: A Basic Course, Kluwer, 2004. doi:10.1007/978-1-4419-8853-9
  • J.-B. Baillon and G. Haddad, Quelques propriétés des opérateurs angle-bornés et n-cycliquement monotones, Israel J. Math. 26:137–150, 1977. doi:10.1007/BF03007664
10 thms1 active userReviewed
Operations ResearchProbability·Captain: mikedeng1

Airline Seat Allocation with Multiple Nested Fare Classes 1: Protection Levels Solving f₁Pr[X₁ > p₁ ∩ … ∩ X₁ + … + X_k > p_k] = f_{k+1} Maximize Expected RevenueResearch Paper

Motivation

An airline sells the seats of one flight leg at several fares. Cheaper fares are booked earlier, so the airline must decide, while low-fare requests arrive, how many seats to hold back for later and more valuable passengers. In nested booking control a seat that could be sold at a low fare is always available to a higher fare. The airline therefore chooses protection levels: pkp_kpk​ seats are reserved for the kkk most expensive classes together, and a request of class k+1k+1k+1 is accepted only while more than pkp_kpk​ seats remain.

For two classes the optimal protection level was found by Littlewood (1972): protect p1p_1p1​ seats, where f1Pr⁡[X1>p1]=f2f_1 \Pr[X_1 > p_1] = f_2f1​Pr[X1​>p1​]=f2​. For more classes the industry used the EMSRa heuristic of Belobaba (1987, 1989), which applies Littlewood's rule to each pair of classes separately and adds the results. Brumelle and McGill (1993) gave the exact optimality conditions for any number of nested classes and showed that EMSRa is in general not optimal. Their conditions are part of the standard theory of single-leg revenue management, as presented in Talluri and van Ryzin (2004).

Setting

There are fare classes k=1,2,…k = 1, 2, \dotsk=1,2,…, numbered from the highest fare. Class kkk has fare fkf_kfk​ and random demand Xk≥0X_k \ge 0Xk​≥0. The standing assumptions (pp. 128–129) are: the demands are mutually independent random variables on a probability space (Ω,F,P)(\Omega, \mathcal F, P)(Ω,F,P), and the fares are strictly decreasing, f1>f2>⋯f_1 > f_2 > \cdotsf1​>f2​>⋯. Demands arrive in order of increasing fare: all of class k+1k+1k+1 before any of class kkk. There are no cancellations or no-shows, and the decision to close a class depends only on the number of current bookings.

A protection-level policy is a vector p=(p1,p2,… )p = (p_1, p_2, \dots)p=(p1​,p2​,…) with pk≥0p_k \ge 0pk​≥0; the dummy p0=0p_0 = 0p0​=0. The revenue Rk[s;p;x]R_k[s; p; x]Rk​[s;p;x] of the kkk highest classes with sss seats available and demand vector xxx is defined recursively by (8)–(9), p. 130:

R1[s;p;x]=f1min⁡(s,x1),R_1[s; p; x] = f_1 \min(s, x_1),R1​[s;p;x]=f1​min(s,x1​), Rk+1[s;p;x]={Rk[s;p;x]0≤s<pk,(s−pk)fk+1+Rk[pk;p;x]pk≤s<pk+xk+1,xk+1fk+1+Rk[s−xk+1;p;x]pk+xk+1≤s.R_{k+1}[s; p; x] = \begin{cases} R_k[s; p; x] & 0 \le s < p_k, \\ (s - p_k) f_{k+1} + R_k[p_k; p; x] & p_k \le s < p_k + x_{k+1}, \\ x_{k+1} f_{k+1} + R_k[s - x_{k+1}; p; x] & p_k + x_{k+1} \le s. \end{cases}Rk+1​[s;p;x]=⎩⎨⎧​Rk​[s;p;x](s−pk​)fk+1​+Rk​[pk​;p;x]xk+1​fk+1​+Rk​[s−xk+1​;p;x]​0≤s<pk​,pk​≤s<pk​+xk+1​,pk​+xk+1​≤s.​

The expected revenue is ERk[s;p;X]=E Rk[s;p;X]ER_k[s; p; X] = E\,R_k[s; p; X]ERk​[s;p;X]=ERk​[s;p;X]. A policy ppp is optimal if ERk[s;q;X]≤ERk[s;p;X]ER_k[s; q; X] \le ER_k[s; p; X]ERk​[s;q;X]≤ERk​[s;p;X] for every policy qqq, every k≥1k \ge 1k≥1 and every s≥0s \ge 0s≥0.

For g:R→Rg : \mathbb R \to \mathbb Rg:R→R, δ+g[s]\delta_+ g[s]δ+​g[s] and δ−g[s]\delta_- g[s]δ−​g[s] denote the right and left derivatives, and the subdifferential δg[s]\delta g[s]δg[s] is the interval [δ+g[s],δ−g[s]][\delta_+ g[s], \delta_- g[s]][δ+​g[s],δ−​g[s]], with δ−g[0]=+∞\delta_- g[0] = +\inftyδ−​g[0]=+∞ (p. 131).

Formalization targets

Goal: Theorem 3 (p. 134)

If the protection levels satisfy

f1Pr⁡[X1>p1∩X1+X2>p2∩⋯∩X1+⋯+Xk>pk]=fk+1for all k≥1,(31)f_1 \Pr[X_1 > p_1 \cap X_1 + X_2 > p_2 \cap \dots \cap X_1 + \dots + X_k > p_k] = f_{k+1} \quad \text{for all } k \ge 1, \tag{31}f1​Pr[X1​>p1​∩X1​+X2​>p2​∩⋯∩X1​+⋯+Xk​>pk​]=fk+1​for all k≥1,(31)

then ppp is optimal.

Milestones

  1. (27), p. 132: ER1ER_1ER1​ is concave, and δER1[s;p;X]=[f1Pr⁡[X1>s],f1Pr⁡[X1≥s]]\delta ER_1[s; p; X] = [f_1 \Pr[X_1 > s], f_1 \Pr[X_1 \ge s]]δER1​[s;p;X]=[f1​Pr[X1​>s],f1​Pr[X1​≥s]].
  2. Lemma 1, p. 131: if ERk[ ⋅ ;p;X]ER_k[\,\cdot\,; p; X]ERk​[⋅;p;X] is concave on s≥0s \ge 0s≥0 and fk+1∈δERk[pk;p;X]f_{k+1} \in \delta ER_k[p_k; p; X]fk+1​∈δERk​[pk​;p;X], then E{Rk+1[s;p;X]∣Xk+1}E\{R_{k+1}[s; p; X] \mid X_{k+1}\}E{Rk+1​[s;p;X]∣Xk+1​} is concave in sss.
  3. Corollary 1, p. 131: under the same conditions ERk+1[ ⋅ ;p;X]ER_{k+1}[\,\cdot\,; p; X]ERk+1​[⋅;p;X] is concave on s≥0s \ge 0s≥0.
  4. Theorem 1, p. 131: if fk+1∈δERk[pk;p;X]f_{k+1} \in \delta ER_k[p_k; p; X]fk+1​∈δERk​[pk​;p;X] for every kkk (condition (20)), then ppp is optimal.
  5. Lemma 2, p. 134: under (31), for s≥pks \ge p_ks≥pk​,
δ+E{Rk+1[s;p;X]∣Xk+1}=f1Pr⁡[X1>p1∩⋯∩X1+⋯+Xk>pk∩X1+⋯+Xk+1>s∣Xk+1].\delta_+ E\{R_{k+1}[s; p; X] \mid X_{k+1}\} = f_1 \Pr[X_1 > p_1 \cap \dots \cap X_1 + \dots + X_k > p_k \cap X_1 + \dots + X_{k+1} > s \mid X_{k+1}].δ+​E{Rk+1​[s;p;X]∣Xk+1​}=f1​Pr[X1​>p1​∩⋯∩X1​+⋯+Xk​>pk​∩X1​+⋯+Xk+1​>s∣Xk+1​].
  1. Corollary 2, p. 134: the unconditional version (37) of Lemma 2 for δ+ERk+1[s;p;X]\delta_+ ER_{k+1}[s; p; X]δ+​ERk+1​[s;p;X].

Significance

Theorem 3 turns the optimal nested protection levels into a sequence of equations in the joint distribution of the cumulative demands X1+⋯+XjX_1 + \dots + X_jX1​+⋯+Xj​. For k=1k = 1k=1 it is Littlewood's rule. For k≥2k \ge 2k≥2 it identifies exactly what EMSRa approximates: EMSRa replaces the joint event in (31) by separate pairwise comparisons, and the paper shows (§4) that EMSRa can both over- and underestimate the optimal protection levels. The conditions are also the input of numerical methods: given demand forecasts, the levels p1,p2,…p_1, p_2, \dotsp1​,p2​,… are found one after another by solving (31), and §3.3 notes that a continuous joint demand distribution guarantees a solution exists.

The results are proved in the paper. As far as is known they have no machine-checked proof. Related platform items cover the two-class, integer-seat case from Belobaba (1987) (SeatInventory.Nested.emsr_protection_level_optimal) and the integer marginal-seat-revenue analogue of (27). They use a different model: two classes, natural-number seats and first differences. This mission formalizes the multi-class statement with real-valued seats and one-sided derivatives. A sister mission of the series proves the existence of optimal integer policies for integer-valued demand (Theorem 2).

Difficulty

The expected revenue is not differentiable: for discrete demand it is piecewise linear, so first-order conditions must be stated with one-sided derivatives and subdifferentials. The natural approach, to optimize each protection level separately with the others fixed, fails without concavity, and concavity of ERk+1ER_{k+1}ERk+1​ in sss is not automatic. It holds only when the lower protection levels already satisfy the first-order conditions. Concavity and optimality must therefore be carried through one joint induction over the classes. Passing from (31) to (20) requires computing the right derivative of the expected revenue in closed form for every s≥pks \ge p_ks≥pk​. This involves exchanging differentiation with expectation and conditioning on one class's demand at a time.

Formalization scope

  • Classes are indexed by N\mathbb NN from 111; fares, demands and protection levels are sequences N→R\mathbb N \to \mathbb RN→R, with no bound on the number of classes. Seats and protection levels are real numbers.
  • Expectation is the Bochner integral on a probability space. The standing assumptions are a single predicate: probability measure, measurable nonnegative demands, mutual independence (iIndepFun), strictly decreasing fares.
  • E{⋅∣Xk}E\{\cdot \mid X_k\}E{⋅∣Xk​} evaluated at Xk=yX_k = yXk​=y is the integral with the kkk-th demand frozen at yyy. Because the demands are independent this is a version of the conditional expectation, and "with probability 1" becomes "for every y≥0y \ge 0y≥0", which is stronger.
  • One-sided derivatives are HasDerivWithinAt on half-lines and must exist; derivWithin, which returns 000 where no derivative exists, is not used. δ−g[0]=+∞\delta_- g[0] = +\inftyδ−​g[0]=+∞ is encoded as a disjunct.
  • Optimality is global: ppp beats every policy qqq at every level kkk and every s≥0s \ge 0s≥0. The page's proof of Theorem 1 shows coordinatewise optimality of pkp_kpk​, and the global form follows by induction on kkk.
  • Fares are not assumed positive in the model: under (20) or (31) with strictly decreasing fares, f1>0f_1 > 0f1​>0 follows. The milestone (27), stated with only the hypotheses on X1X_1X1​ that it needs, assumes X1≥0X_1 \ge 0X1​≥0 and f1≥0f_1 \ge 0f1​≥0, without which ER1ER_1ER1​ is not concave.
  • No continuity of the demand distribution is assumed. Theorem 3 is conditional on a solution of (31).
  • The page's hypothesis of Lemma 1 has the misprint "(p0,…,pk+1)(p_0, \dots, p_{k+1})(p0​,…,pk+1​)" for (p0,…,pk−1)(p_0, \dots, p_{k-1})(p0​,…,pk−1​). The formal statement uses the latter.

The goal assumes only the standing assumptions, p≥0p \ge 0p≥0, and (31). It does not assume concavity, condition (20) or any derivative formula: those are milestones. A formalization that quantified optimality over one level, one value of sss, or policies differing from ppp in one coordinate would be weaker than the paper and is excluded.

A complete development needs one-sided derivatives of integrals of piecewise-linear functions (dominated convergence for difference quotients), concavity of piecewise functions glued at points where the slopes decrease, and the independence calculus that turns E[E{⋅∣Xk+1}]E[E\{\cdot \mid X_{k+1}\}]E[E{⋅∣Xk+1​}] into an iterated integral. These pieces are reusable for other newsvendor-type and revenue-management models. Proofs of any milestone, and alternative arguments for Theorem 1, are welcome.

Selected references

  • S. L. Brumelle and J. I. McGill, Airline Seat Allocation with Multiple Nested Fare Classes, Operations Research 41(1), 127–137, 1993. https://doi.org/10.1287/opre.41.1.127
  • K. Littlewood, Forecasting and Control of Passenger Bookings, AGIFORS Symposium Proceedings 12, 95–117, 1972; reprinted in Journal of Revenue and Pricing Management 4(2), 2005. https://doi.org/10.1057/palgrave.rpm.5170134
  • P. P. Belobaba, Air Travel Demand and Airline Seat Inventory Management, PhD thesis, MIT, 1987. http://hdl.handle.net/1721.1/68077
  • P. P. Belobaba, Application of a Probabilistic Decision Model to Airline Seat Inventory Control, Operations Research 37(2), 183–197, 1989. https://doi.org/10.1287/opre.37.2.183
  • K. T. Talluri and G. J. van Ryzin, The Theory and Practice of Revenue Management, Springer, 2004. https://doi.org/10.1007/b139000
8 thms1 active userReviewed
Control TheoryOperations Research·Captain: mikedeng1

Optimal Dynamic Pricing of Inventories with Stochastic Demand over Finite Horizons 2: The Optimal Revenue Is Strictly Concave in Stock and Time; the Optimal Price Falls with Stock, Rises with TimeResearch Paper

Why the shape of the optimal pricing policy matters

A retailer holding a fixed stock of a perishable or seasonal good (fashion items, airline seats, hotel rooms, concert tickets) must sell it before a deadline, after which unsold units are worthless. Demand is random and depends on the posted price, and the firm may change its price at any time. Dynamic pricing asks how the price should depend on the remaining stock and the remaining time.

Gallego and van Ryzin (Management Science 40(8), 1994) posed this problem as a continuous-time intensity control problem and established its basic structure. Their Theorem 1 says that the optimal expected revenue is strictly increasing and strictly concave in both the stock and the time remaining, and that the optimal price falls as stock grows and rises with the time left to sell. The paper is a standard reference of revenue management; the structural result is the continuous-time counterpart of the monotonicity of marginal values in discrete-time models (Talluri and van Ryzin, The Theory and Practice of Revenue Management, 2004, Proposition 5.2), and it is what makes the optimal policy computable by restricting attention to monotone policies. The paper credits a slightly weaker version to Kincaid and Darling (1963). The source formalized here is the published 1994 article.

The model and the Hamilton–Jacobi system

The firm chooses a demand rate λ\lambdaλ from a set Λ⊆[0,∞)\Lambda \subseteq [0,\infty)Λ⊆[0,∞) of allowable rates, an interval containing 000; the market then sets the price p(λ)p(\lambda)p(λ), where ppp is the inverse demand function, strictly decreasing and nonnegative on the positive rates. The rate 000 corresponds to the null price at which nothing sells. The revenue rate is

r(λ)=λ p(λ),r(0)=0.r(\lambda) = \lambda\,p(\lambda), \qquad r(0) = 0.r(λ)=λp(λ),r(0)=0.

The demand function is regular when rrr is continuous, bounded and concave on Λ\LambdaΛ and has a least maximizer λ∗=min⁡{λ:r(λ)=max⁡μ∈Λr(μ)}\lambda^* = \min\{\lambda : r(\lambda) = \max_{\mu\in\Lambda} r(\mu)\}λ∗=min{λ:r(λ)=maxμ∈Λ​r(μ)}. The exponential demand λ(p)=ae−p\lambda(p) = ae^{-p}λ(p)=ae−p, with Λ=[0,a]\Lambda=[0,a]Λ=[0,a], p(λ)=log⁡(a/λ)p(\lambda)=\log(a/\lambda)p(λ)=log(a/λ) and λ∗=a/e\lambda^*=a/eλ∗=a/e, is the running example.

With nnn units in stock and time remaining ttt, write J(n,t)J(n,t)J(n,t) for the optimal expected revenue. The paper derives the Hamilton–Jacobi system

∂J(n,t)∂t=sup⁡λ∈Λ[r(λ)−λ(J(n,t)−J(n−1,t))],n≥1, t>0,(8)\frac{\partial J(n,t)}{\partial t} = \sup_{\lambda\in\Lambda}\big[r(\lambda) - \lambda\big(J(n,t)-J(n-1,t)\big)\big], \qquad n\ge1,\ t>0, \tag{8}∂t∂J(n,t)​=λ∈Λsup​[r(λ)−λ(J(n,t)−J(n−1,t))],n≥1, t>0,(8)

with J(n,0)=0J(n,0)=0J(n,0)=0 and J(0,t)=0J(0,t)=0J(0,t)=0. The difference J(n,t)−J(n−1,t)J(n,t)-J(n-1,t)J(n,t)−J(n−1,t) is the marginal value of an item; a rate attaining the supremum is an optimal intensity λ∗(n,t)\lambda^*(n,t)λ∗(n,t), and p(λ∗(n,t))p(\lambda^*(n,t))p(λ∗(n,t)) is the optimal price p∗(n,t)p^*(n,t)p∗(n,t).

Formalization targets

Goal: Theorem 1 (p. 1005)

For the solution JJJ of (8), with rrr strictly concave, differentiable on the interior of Λ\LambdaΛ, and λ∗\lambda^*λ∗ interior:

J(n,t) strictly increasing in n (t>0) and in t (n≥1);J(n+1,t)−J(n,t)<J(n,t)−J(n−1,t);J(n,t)\ \text{strictly increasing in } n\ (t>0)\ \text{and in } t\ (n\ge1);\qquad J(n+1,t)-J(n,t) < J(n,t)-J(n-1,t);J(n,t) strictly increasing in n (t>0) and in t (n≥1);J(n+1,t)−J(n,t)<J(n,t)−J(n−1,t); t↦J(n,t) strictly concave;∃ λ∗(n,t): λ∗ ⁣↑n, λ∗ ⁣↓t,p∗ ⁣↓n, p∗ ⁣↑t (strictly).t\mapsto J(n,t)\ \text{strictly concave};\qquad \exists\,\lambda^*(n,t):\ \lambda^*\!\uparrow_n,\ \lambda^*\!\downarrow_t,\quad p^*\!\downarrow_n,\ p^*\!\uparrow_t\ \text{(strictly)}.t↦J(n,t) strictly concave;∃λ∗(n,t): λ∗↑n​, λ∗↓t​,p∗↓n​, p∗↑t​ (strictly).

Milestones

  1. The supremum in (8) is a maximum over [0,λ∗][0,\lambda^*][0,λ∗] whenever the marginal value is nonnegative (proof of Proposition 1).
  2. Proposition 1: (8) has a unique solution, and λ∗(n,s)≤λ∗\lambda^*(n,s)\le\lambda^*λ∗(n,s)≤λ∗.
  3. Eq. (26): J(n,t)−J(n−1,t)=r′(λ∗(n,t))>0J(n,t)-J(n-1,t) = r'(\lambda^*(n,t)) > 0J(n,t)−J(n−1,t)=r′(λ∗(n,t))>0 for t>0t>0t>0.
  4. The case n=1n=1n=1 of Theorem 1: λ∗(1,t)\lambda^*(1,t)λ∗(1,t) strictly decreasing and J(1,t)J(1,t)J(1,t) strictly concave in ttt.
  5. λ∗(n,0+)=λ∗\lambda^*(n,0^+) = \lambda^*λ∗(n,0+)=λ∗.
  6. Eqs. (9)–(10), exponential demand: J(n,t)=log⁡∑i=0n(λ∗t)i/i!J(n,t) = \log\sum_{i=0}^n(\lambda^*t)^i/i!J(n,t)=log∑i=0n​(λ∗t)i/i! and p∗(n,t)=J(n,t)−J(n−1,t)+1p^*(n,t) = J(n,t)-J(n-1,t)+1p∗(n,t)=J(n,t)−J(n−1,t)+1.
  7. Proposition 3, exponential demand: λ∗(n,t)≤λD(n,t)=min⁡{λ∗,n/t}\lambda^*(n,t)\le\lambda^D(n,t)=\min\{\lambda^*,n/t\}λ∗(n,t)≤λD(n,t)=min{λ∗,n/t} and p∗(n,t)≥p(λD(n,t))p^*(n,t)\ge p(\lambda^D(n,t))p∗(n,t)≥p(λD(n,t)).

Significance

Theorem 1 is the qualitative backbone of single-product dynamic pricing. Concavity of JJJ in nnn means that each additional unit is worth less than the previous one, which is the basis of bid-price and marginal-value reasoning in revenue management; the monotone price path justifies markdown practice as the deadline approaches and reduces the policy search to monotone policies. Proposition 3 answers, for exponential demand, a question raised by Mills (1959): the stochastic optimal price is never below the deterministic one. The closed form (9)–(10) is one of the few exactly solvable intensity control problems in pricing.

The results are proved in the paper; none of them has a machine-checked proof. Formalizing them requires the comparison and monotonicity theory of a countable system of coupled ordinary differential equations whose right-hand side is a convex conjugate, a theory that Mathlib does not package. A complete development would also certify the corrected hypotheses of Theorem 1 described below.

Difficulty

The value functions are defined only implicitly by (8), a triangular infinite system of ODEs in which each J(n,⋅)J(n,\cdot)J(n,⋅) is driven by J(n−1,⋅)J(n-1,\cdot)J(n−1,⋅) through the nonsmooth map Δ↦sup⁡λ[r(λ)−λΔ]\Delta\mapsto\sup_\lambda[r(\lambda)-\lambda\Delta]Δ↦supλ​[r(λ)−λΔ]. Monotonicity of the optimal intensity in ttt is a statement about the time derivative of a marginal value, and the paper establishes it by an induction on nnn combined with an argument by contradiction on the first interval where monotonicity could fail. The obvious approach of differentiating (8) twice in ttt needs second derivatives of rrr and of JJJ that the hypotheses do not provide, and the paper's own proof of Proposition 1 assumes that JJJ is nondecreasing in nnn, which is only established in Theorem 1; a rigorous development must break this circularity.

Formalization scope

A regular demand function is a Lean structure (GVRPricing.Structure.Model) holding Λ\LambdaΛ, ppp and λ∗\lambda^*λ∗ with the paper's standing assumptions of §2.1: 0∈Λ⊆[0,∞)0\in\Lambda\subseteq[0,\infty)0∈Λ⊆[0,∞) an interval, ppp strictly decreasing and nonnegative on Λ∖{0}\Lambda\setminus\{0\}Λ∖{0}, r(λ)=λp(λ)r(\lambda)=\lambda p(\lambda)r(λ)=λp(λ) continuous, concave and bounded on Λ\LambdaΛ, and λ∗\lambda^*λ∗ the least maximizer. The definition IsHJBSolution encodes (8) for J:N→R→RJ:\mathbb N\to\mathbb R\to\mathbb RJ:N→R→R, with the second argument the time remaining, the two-sided derivative at each t>0t>0t>0, continuity on [0,∞)[0,\infty)[0,∞), the boundary conditions, and the requirement that the set inside the supremum be bounded above, so that the real supremum is never a default value. An optimal intensity at (n,t)(n,t)(n,t) is any ℓ∈Λ\ell\in\Lambdaℓ∈Λ maximizing λ↦r(λ)−λ(J(n,t)−J(n−1,t))\lambda\mapsto r(\lambda)-\lambda(J(n,t)-J(n-1,t))λ↦r(λ)−λ(J(n,t)−J(n−1,t)) over Λ\LambdaΛ.

All theorems are about solutions of (8), on which the paper's proofs operate. The identification of the solution of (8) with the supremum of expected revenue over non-anticipating pricing policies is Brémaud's verification theorem, which the paper cites and does not prove; it is not part of this mission.

Added hypotheses and corrected statements.

  • As printed, Theorem 1 assumes only a regular demand function and is false: for r(λ)=λr(\lambda)=\sqrt\lambdar(λ)=λ​ on [0,1][0,1][0,1] and r=1r=1r=1 beyond, λ∗(1,t)=1\lambda^*(1,t)=1λ∗(1,t)=1 for all t∈(0,ln⁡2]t\in(0,\ln2]t∈(0,ln2]; for r(λ)=λ−λ2/4r(\lambda)=\lambda-\lambda^2/4r(λ)=λ−λ2/4 on Λ=[0,1]\Lambda=[0,1]Λ=[0,1], λ∗=1\lambda^*=1λ∗=1 is on the boundary and λ∗(1,t)=1\lambda^*(1,t)=1λ∗(1,t)=1 for small ttt. The goal, eq. (26) and the case n=1n=1n=1 therefore assume that rrr is strictly concave, differentiable on the interior of Λ\LambdaΛ, and that λ∗\lambda^*λ∗ is interior; the appendix proof uses all three. Proposition 1, the restriction lemma and λ∗(n,0+)=λ∗\lambda^*(n,0^+)=\lambda^*λ∗(n,0+)=λ∗ use only the printed assumptions.
  • Strict claims in nnn are made for t>0t>0t>0, since J(n,0)=0J(n,0)=0J(n,0)=0 for all nnn; optimal intensities are considered for n≥1n\ge1n≥1, t>0t>0t>0.
  • Proposition 1's bound "λ∗(n,s)≤λ∗\lambda^*(n,s)\le\lambda^*λ∗(n,s)≤λ∗ for 0≤s0\le s0≤s" is read at s=0s=0s=0 as "λ∗\lambda^*λ∗ is optimal", since every maximizer of rrr is optimal there.
  • Proposition 3 is stated for n≥1n\ge1n≥1, t>0t>0t>0 (the page says n≥0n\ge0n≥0, t≥0t\ge0t≥0, where n/tn/tn/t or the optimal intensity is undefined).
  • The exponential results use the paper's normalization α=1\alpha=1α=1 of λ(p)=ae−αp\lambda(p)=ae^{-\alpha p}λ(p)=ae−αp.
  • The case n=1n=1n=1 omits the displayed identities involving r′′r''r′′ and λ∗′\lambda^{*\prime}λ∗′, which presuppose second derivatives; its conclusions are stated.

A formalization that defines JJJ by a formula, or postulates a monotone function as the optimal intensity, would trivialize the goal: the goal quantifies over every solution of (8), and the optimal intensity it asserts must maximize the right-hand side of (8) at every (n,t)(n,t)(n,t).

A complete development needs comparison principles for scalar ODEs with Lipschitz right-hand sides, properties of the concave conjugate Δ↦sup⁡λ[r(λ)−λΔ]\Delta\mapsto\sup_\lambda[r(\lambda)-\lambda\Delta]Δ↦supλ​[r(λ)−λΔ] (monotonicity, Lipschitz continuity, envelope theorem), and monotone comparative statics of maximizers. These are reusable well beyond this mission; contributions of any of them, or of the milestones in any order, are welcome.

Selected references

  • G. Gallego, G. van Ryzin, Optimal Dynamic Pricing of Inventories with Stochastic Demand over Finite Horizons, Management Science 40(8), 999–1020, 1994. https://doi.org/10.1287/mnsc.40.8.999
  • P. Brémaud, Point Processes and Queues: Martingale Dynamics, Springer, 1981. https://doi.org/10.1007/978-1-4684-9477-8
  • W. M. Kincaid, D. A. Darling, An Inventory Pricing Problem, Journal of Mathematical Analysis and Applications 7, 183–208, 1963. https://doi.org/10.1016/0022-247X(63)90047-7
  • K. T. Talluri, G. J. van Ryzin, The Theory and Practice of Revenue Management, Springer, 2004. https://doi.org/10.1007/b139000
10 thms1 active userReviewed
🏆Completed
Linear OptimizationOperations Research·Captain: mikedeng1

Assortment Optimization under Variants of the Nested Logit Model 1: If the Restricted LP Optimum Scaled by α Is Feasible for the Full LP, Its Assortment Earns Within a Factor α of the Optimal RevenueResearch Paper

Motivation

A retailer that sells products in several categories, channels or stores has to decide which products to offer in each. Customers substitute: a product left out of the assortment sends some of its demand to other products, and some of it away. The nested logit model (McFadden 1974, 1981) is the standard choice model for this situation. It groups products into nests, so that substitution within a nest differs from substitution across nests. Assortment optimization under this model asks which products to offer in each nest so as to maximize expected revenue.

Davis, Gallego and Topaloglu (Oper. Res. 62(2), 2014) split the problem into four cases: dissimilarity parameters at most one or unrestricted, and nests that are fully or only partially captured. The problem is polynomially solvable in the first case and NP-hard in the other three. Every approximation guarantee in the paper for the hard cases (Theorems 7, 10, 11, 12) comes from one general framework, set up in §2: a linear program equivalent to the assortment problem, and Theorem 1, which turns a feasibility certificate for that linear program into a performance guarantee. This mission formalizes that framework.

Setting

There are mmm nests MMM and nnn products N={1,…,n}N = \{1, \dots, n\}N={1,…,n} in each nest. Product jjj of nest iii has revenue rij≥0r_{ij} \ge 0rij​≥0 and preference weight vij>0v_{ij} > 0vij​>0, and the products in each nest are ordered so that ri1≥⋯≥rinr_{i1} \ge \dots \ge r_{in}ri1​≥⋯≥rin​. Nest iii has a no-purchase weight vi0≥0v_{i0} \ge 0vi0​≥0 and a dissimilarity parameter γi>0\gamma_i > 0γi​>0. The weight of choosing no nest at all is v0≥0v_0 \ge 0v0​≥0. If the assortment Si⊆NS_i \subseteq NSi​⊆N is offered in nest iii, write

Vi(Si)=vi0+∑j∈Sivij,Ri(Si)=∑j∈SirijvijVi(Si),Ri(∅)=0.V_i(S_i) = v_{i0} + \sum_{j \in S_i} v_{ij}, \qquad R_i(S_i) = \frac{\sum_{j\in S_i} r_{ij} v_{ij}}{V_i(S_i)}, \quad R_i(\emptyset) = 0 .Vi​(Si​)=vi0​+j∈Si​∑​vij​,Ri​(Si​)=Vi​(Si​)∑j∈Si​​rij​vij​​,Ri​(∅)=0.

A customer picks nest iii with probability Qi=Vi(Si)γi/(v0+∑l∈MVl(Sl)γl)Q_i = V_i(S_i)^{\gamma_i} / (v_0 + \sum_{l\in M} V_l(S_l)^{\gamma_l})Qi​=Vi​(Si​)γi​/(v0​+∑l∈M​Vl​(Sl​)γl​), and then a product of that nest by the multinomial logit model. The expected revenue is

Π(S1,…,Sm)=∑i∈MQi(S1,…,Sm) Ri(Si),\Pi(S_1, \dots, S_m) = \sum_{i \in M} Q_i(S_1, \dots, S_m)\, R_i(S_i),Π(S1​,…,Sm​)=i∈M∑​Qi​(S1​,…,Sm​)Ri​(Si​),

and problem (2) is Z∗=max⁡Si⊆NΠ(S1,…,Sm)Z^* = \max_{S_i \subseteq N} \Pi(S_1, \dots, S_m)Z∗=maxSi​⊆N​Π(S1​,…,Sm​).

The linear program (3) in the variables (x,y1,…,ym)(x, y_1, \dots, y_m)(x,y1​,…,ym​) minimizes xxx subject to

v0x≥∑i∈Myi,yi≥Vi(Si)γi(Ri(Si)−x)∀Si⊆N, i∈M.v_0 x \ge \sum_{i\in M} y_i, \qquad y_i \ge V_i(S_i)^{\gamma_i}\big(R_i(S_i) - x\big) \quad \forall S_i \subseteq N,\ i \in M.v0​x≥i∈M∑​yi​,yi​≥Vi​(Si​)γi​(Ri​(Si​)−x)∀Si​⊆N, i∈M.

Given candidate collections {Ait:t∈Ti}\{A_{it} : t \in \mathcal T_i\}{Ait​:t∈Ti​} of assortments for each nest, the linear program (4) is (3) with the second family of constraints imposed only for SiS_iSi​ in the collection of nest iii.

Formalization targets

Goal: Theorem 1 (p. 13)

Let (x^,y^)(\hat x, \hat y)(x^,y^​) be an optimal solution of (4), and let S^i\hat S_iS^i​ solve max⁡Si∈{Ait}Vi(Si)γi(Ri(Si)−x^)\max_{S_i \in \{A_{it}\}} V_i(S_i)^{\gamma_i}(R_i(S_i) - \hat x)maxSi​∈{Ait​}​Vi​(Si​)γi​(Ri​(Si​)−x^), problem (5), in every nest. If (αx^,βy^)(\alpha \hat x, \beta \hat y)(αx^,βy^​) is feasible for (3) for some α,β\alpha, \betaα,β, then, with Z^=Π(S^1,…,S^m)\hat Z = \Pi(\hat S_1, \dots, \hat S_m)Z^=Π(S^1​,…,S^m​),

αZ^ ≥ Z∗ ≥ Z^.\alpha \hat Z \ \ge\ Z^* \ \ge\ \hat Z .αZ^ ≥ Z∗ ≥ Z^.

The theorem fixes no candidate collection and no value of α\alphaα. Each later section of the paper instantiates it with its own collection and its own factor, so a formal proof applies to all of them.

Milestones (§2, pp. 11–12)

  1. Problem (2) is equivalent to (3): Z∗Z^*Z∗ is the least xxx for which some yyy makes (x,y)(x, y)(x,y) feasible for (3).
  2. At an optimal solution of (4), the first constraint binds at the maximizers S^i\hat S_iS^i​ of (5), and x^=Π(S^1,…,S^m)\hat x = \Pi(\hat S_1, \dots, \hat S_m)x^=Π(S^1​,…,S^m​).
  3. Problem (4) relaxes (3), so x^≤Z∗\hat x \le Z^*x^≤Z∗.

Companions (§7, pp. 29–30)

  • The tighter program (16), which lets each nest's assortment be a fractional vector zi∈[0,1]nz_i \in [0,1]^nzi​∈[0,1]n, has every feasible xxx above Z∗Z^*Z∗.
  • Proposition 13: F^i(x)=max⁡zi∈[0,1]nFi(zi∣x)\hat F_i(x) = \max_{z_i \in [0,1]^n} F_i(z_i \mid x)F^i​(x)=maxzi​∈[0,1]n​Fi​(zi​∣x) is convex, with subgradient −(vi0+∑jvijz^ij(x))γi-(v_{i0} + \sum_j v_{ij}\hat z_{ij}(x))^{\gamma_i}−(vi0​+∑j​vij​z^ij​(x))γi​ at xxx.

Significance

Theorem 1 is the common step behind the paper's four approximation guarantees: the factor ρ\rhoρ or 2κ2\kappa2κ of Theorem 7, the factor 2 of Theorem 10, the factor of Theorem 11, and the δ2γˉ+1\delta^{2\bar\gamma+1}δ2γˉ​+1 of Theorem 12. Each of these reduces to checking that a scaled optimum of a small linear program is feasible for (3). With Theorem 1 formalized, those guarantees reduce to inequalities about candidate collections, which are the subject of the sister missions of this series. The upper bound (16) and Proposition 13 give the instance-specific bound that the paper uses to assess its assortments numerically.

The results are proved in the paper. To our knowledge none of them has a machine-checked proof. The formal work adds two things: the statements below are made exact at the degenerate inputs the prose passes over (an empty assortment, v0=0v_0 = 0v0​=0), and a formal proof certifies the framework once for every later instantiation.

Difficulty

The equivalence of (2) and (3) rests on decomposing a maximum over joint assortments into a sum of per-nest maxima, and on reading the fractional objective Π≤x\Pi \le xΠ≤x as a linear constraint. Both steps need care where a denominator v0+∑iVi(Si)γiv_0 + \sum_i V_i(S_i)^{\gamma_i}v0​+∑i​Vi​(Si​)γi​ can vanish. The binding argument for (4) is a perturbation argument: lowering x^\hat xx^ must keep every constraint satisfiable, which needs a continuity and monotonicity property of the right-hand side in xxx. The obvious one-line reading of Theorem 1, "x^=Z^\hat x = \hat Zx^=Z^ and αx^≥Z∗\alpha\hat x \ge Z^*αx^≥Z∗", is correct only once both of these facts are established with their hypotheses. In particular, it is false when v0=0v_0 = 0v0​=0 (see below). Proposition 13 requires that the supremum over the box be finite, which comes from the boundedness of FiF_iFi​ on [0,1]n[0,1]^n[0,1]n.

Formalization scope

Nests are a finite type ι and products are Fin n, indexed 0,…,n−10, \dots, n-10,…,n−1. An assortment is a finite set of products per nest, and a candidate collection is a set of such finite sets. Powers are real powers, and Lean's x/0=0x / 0 = 0x/0=0 gives Ri(∅)=0R_i(\emptyset) = 0Ri​(∅)=0. Z∗Z^*Z∗ is Π(S∗)\Pi(S^*)Π(S∗) for an arbitrary optimal assortment S∗S^*S∗; no supremum over assortments is taken. "Optimal solution of (4)" means feasible with minimal xxx, and "S^i\hat S_iS^i​ solves (5)" means S^i\hat S_iS^i​ belongs to the collection of nest iii and maximizes the objective of (5) over it at x^\hat xx^.

Standing assumptions and pins. These are v0,vi0≥0v_0, v_{i0} \ge 0v0​,vi0​≥0 and ordered revenues, together with vij>0v_{ij} > 0vij​>0, rij≥0r_{ij} \ge 0rij​≥0 and γi>0\gamma_i > 0γi​>0. The paper allows zero-weight padding products and γi=0\gamma_i = 0γi​=0, but its own conventions fail there. Theorem 1 and the binding milestone add v0>0v_0 > 0v0​>0. The page allows v0=0v_0 = 0v0​=0, but then Theorem 1 is false: take one nest with v10=0v_{10} = 0v10​=0, γ1=1\gamma_1 = 1γ1​=1, r11=v11=1r_{11} = v_{11} = 1r11​=v11​=1 and candidates {∅,{1}}\{\emptyset, \{1\}\}{∅,{1}}. Then x^=1\hat x = 1x^=1 and S^1=∅\hat S_1 = \emptysetS^1​=∅ meet every hypothesis with α=β=1\alpha = \beta = 1α=β=1, yet Z^=0<Z∗=1\hat Z = 0 < Z^* = 1Z^=0<Z∗=1. The equivalence of (2) and (3) and the bound from (16) keep v0≥0v_0 \ge 0v0​≥0, as the page does, and assume at least one nest and one product: with neither and v0=0v_0 = 0v0​=0, every xxx is feasible for (3).

A formalization that assumes the binding equality, the identity x^=Z^\hat x = \hat Zx^=Z^, or the inequality x^≤Z∗\hat x \le Z^*x^≤Z∗ in the goal would trivialize it. Those facts appear only as milestones. Likewise, reading "optimal solution of (4)" as mere feasibility would make the goal false rather than easier.

The development needs only finite sums, real powers and elementary order reasoning. Proposition 13 also needs the boundedness of a continuous function on a box and the convexity of a pointwise supremum of affine functions. Welcome contributions include proofs of the milestones, the goal from them, and reusable lemmas on the per-nest decomposition of maxima, which the sister missions of this series use as well.

Selected references

  • J. M. Davis, G. Gallego, H. Topaloglu, Assortment optimization under variants of the nested logit model, Operations Research 62(2), 2014 (revised manuscript of June 18, 2013, cited here). https://doi.org/10.1287/opre.2014.1256
  • D. McFadden, Econometric models of probabilistic choice, in C. Manski, D. McFadden (eds.), Structural Analysis of Discrete Data with Econometric Applications, MIT Press, 1981. https://eml.berkeley.edu/~mcfadden/discrete.html
  • P. Rusmevichientong, D. Shmoys, H. Topaloglu, Assortment optimization with mixtures of logits, Technical report, Cornell University, 2010. https://people.orie.cornell.edu/huseyin/publications/publications.html
  • M. S. Bazaraa, H. D. Sherali, C. M. Shetty, Nonlinear Programming: Theory and Algorithms, 2nd ed., Wiley, 1993. https://doi.org/10.1002/0471787779
6 thms1 active userReviewed
Control TheoryDynamic ProgrammingOperations Research+2·Captain: mikedeng1

Stochastic Optimal Control: The Discrete-Time Case IX: Imperfect State Information — Reduction to a Perfect-Information Model through a Statistic Sufficient for ControlTextbook

Motivation

In most control problems the controller does not see the state of the system. It sees noisy observations, remembers its past controls, and must act on that record. Inventory systems with delayed or inaccurate counts, maintenance of machines whose wear is only inspected, target tracking, and medical treatment planned from test results all have this form. The standard device for such problems is to replace the hidden state by a summary of the record, most often the conditional distribution of the state given the observations, and to solve a dynamic program whose state is that summary.

For finite or countable spaces this reduction goes back to Åström (1965) and Striebel (1965), who introduced the conditional distribution of the state as a "sufficient statistic" for control. Chapter 10 of Bertsekas and Shreve, Stochastic Optimal Control: The Discrete-Time Case (Academic Press 1978; Athena Scientific 1996) carries it out for Borel state, control and observation spaces, with universally measurable policies and costs that are only lower semianalytic. In that generality the measurability of the reduced model is the whole difficulty, and the chapter isolates exactly what a summary must satisfy for the reduction to be exact.

Setting

The imperfect state information model (ISI) of Definition 10.3 has a nonempty Borel state space SSS, control space CCC and observation space ZZZ; a discount factor α>0\alpha>0α>0; a lower semianalytic cost g:SC→R∗=[−∞,∞]g:SC\to R^*=[-\infty,\infty]g:SC→R∗=[−∞,∞]; a Borel state transition kernel t(dx′∣x,u)t(dx'\mid x,u)t(dx′∣x,u); Borel observation kernels s0(dz∣x)s_0(dz\mid x)s0​(dz∣x) and s(dz∣u,x)s(dz\mid u,x)s(dz∣u,x); and a horizon NNN. The initial state x0x_0x0​ has distribution p∈P(S)p\in P(S)p∈P(S), z0∼s0(⋅∣x0)z_0\sim s_0(\cdot\mid x_0)z0​∼s0​(⋅∣x0​), and then xk+1∼t(⋅∣xk,uk)x_{k+1}\sim t(\cdot\mid x_k,u_k)xk+1​∼t(⋅∣xk​,uk​), zk+1∼s(⋅∣uk,xk+1)z_{k+1}\sim s(\cdot\mid u_k,x_{k+1})zk+1​∼s(⋅∣uk​,xk+1​). The controller knows the information vector ik=(z0,u0,…,uk−1,zk)∈Iki_k=(z_0,u_0,\dots,u_{k-1},z_k)\in I_kik​=(z0​,u0​,…,uk−1​,zk​)∈Ik​ and must choose uk∈Uk(ik)u_k\in U_k(i_k)uk​∈Uk​(ik​), where the constraint set Γk={(ik,u)∣u∈Uk(ik)}\Gamma_k=\{(i_k,u)\mid u\in U_k(i_k)\}Γk​={(ik​,u)∣u∈Uk​(ik​)} is analytic.

A policy π=(μ0,…,μN−1)\pi=(\mu_0,\dots,\mu_{N-1})π=(μ0​,…,μN−1​) consists of universally measurable stochastic kernels μk(duk∣p;ik)\mu_k(du_k\mid p;i_k)μk​(duk​∣p;ik​) that respect the constraints (Definition 10.4). Together with ppp it determines probability measures Pk(π,p)P_k(\pi,p)Pk​(π,p) on the histories (x0,z0,u0,…,xk,zk,uk)(x_0,z_0,u_0,\dots,x_k,z_k,u_k)(x0​,z0​,u0​,…,xk​,zk​,uk​), the cost

JN,π(p)=∫[∑k=0N−1αkg(xk,uk)]dPN−1(π,p),J_{N,\pi}(p)=\int\Big[\sum_{k=0}^{N-1}\alpha^k g(x_k,u_k)\Big]dP_{N-1}(\pi,p),JN,π​(p)=∫[k=0∑N−1​αkg(xk​,uk​)]dPN−1​(π,p),

and the optimal cost JN∗(p)=inf⁡πJN,π(p)J^*_N(p)=\inf_\pi J_{N,\pi}(p)JN∗​(p)=infπ​JN,π​(p) (Definition 10.5). Assumption (F+)(F^+)(F+) asks that the expected discounted negative part of the cost be finite for every policy and initial distribution; (F−)(F^-)(F−) asks the same of the positive part.

A statistic is a sequence of Borel maps ηk:P(S)Ik→Yk\eta_k:P(S)I_k\to Y_kηk​:P(S)Ik​→Yk​ into nonempty Borel spaces. It is sufficient for control (Definition 10.6) if (a) the constraints can be read off from it, Γk={(ik,u)∣(ηk(p;ik),u)∈Γ^k}\Gamma_k=\{(i_k,u)\mid(\eta_k(p;i_k),u)\in\hat\Gamma_k\}Γk​={(ik​,u)∣(ηk​(p;ik​),u)∈Γ^k​} with Γ^k\hat\Gamma_kΓ^k​ analytic; (b) the conditional law of ηk+1\eta_{k+1}ηk+1​ given (ηk,uk)(\eta_k,u_k)(ηk​,uk​) is a Borel kernel t^k(dyk+1∣yk,uk)\hat t_k(dy_{k+1}\mid y_k,u_k)t^k​(dyk+1​∣yk​,uk​), for every ppp and every policy; and (c) the conditional expectation of g(xk,uk)g(x_k,u_k)g(xk​,uk​) given (ηk,uk)(\eta_k,u_k)(ηk​,uk​) is a lower semianalytic function g^k(yk,uk)\hat g_k(y_k,u_k)g^​k​(yk​,uk​). The perfect state information model (PSI) of Definition 10.7 has states yk∈Yky_k\in Y_kyk​∈Yk​, constraints U^k(yk)=(Γ^k)yk\hat U_k(y_k)=(\hat\Gamma_k)_{y_k}U^k​(yk​)=(Γ^k​)yk​​, costs g^k\hat g_kg^​k​ and transitions t^k\hat t_kt^k​; its cost and optimal cost at y∈Y0y\in Y_0y∈Y0​ are J^N,π^(y)\hat J_{N,\hat\pi}(y)J^N,π^​(y) and J^N∗(y)\hat J^*_N(y)J^N∗​(y). The initial distribution of y0y_0y0​ is

φ(p)(Y‾0)=∫Ss0({z0∣η0(p;z0)∈Y‾0}∣x0) p(dx0).\varphi(p)(\underline Y_0)=\int_S s_0(\{z_0\mid\eta_0(p;z_0)\in\underline Y_0\}\mid x_0)\,p(dx_0).φ(p)(Y​0​)=∫S​s0​({z0​∣η0​(p;z0​)∈Y​0​}∣x0​)p(dx0​).

A Markov (PSI) policy μ^k(du∣yk)\hat\mu_k(du\mid y_k)μ^​k​(du∣yk​) acts in (ISI) through μk(du∣p;ik)=μ^k(du∣ηk(p;ik))\mu_k(du\mid p;i_k)=\hat\mu_k(du\mid\eta_k(p;i_k))μk​(du∣p;ik​)=μ^​k​(du∣ηk​(p;ik​)).

Formalization targets

Goal: Proposition 10.3

Under (F+,F^+)(F^+,\hat F^+)(F+,F^+) or (F−,F^−)(F^-,\hat F^-)(F−,F^−),

JN∗(p)=∫Y0J^N∗(y0) φ(p)(dy0)∀p∈P(S),J^*_N(p)=\int_{Y_0}\hat J^*_N(y_0)\,\varphi(p)(dy_0)\qquad\forall p\in P(S),JN∗​(p)=∫Y0​​J^N∗​(y0​)φ(p)(dy0​)∀p∈P(S),

and a Markov (PSI) policy that is optimal, φ(p)\varphi(p)φ(p)-optimal or weakly φ(p)\varphi(p)φ(p)-ε\varepsilonε-optimal for (PSI) is respectively optimal, optimal at ppp, or ε\varepsilonε-optimal at ppp for (ISI); under (F+,F^+)(F^+,\hat F^+)(F+,F^+) an ε\varepsilonε-optimal (PSI) policy is ε\varepsilonε-optimal for (ISI). Here π^\hat\piπ^ is weakly qqq-ε\varepsilonε-optimal if ∫J^N,π^ dq≤∫J^N∗ dq+ε\int\hat J_{N,\hat\pi}\,dq\le\int\hat J^*_N\,dq+\varepsilon∫J^N,π^​dq≤∫J^N∗​dq+ε when ∫J^N∗ dq>−∞\int\hat J^*_N\,dq>-\infty∫J^N∗​dq>−∞ and ∫J^N,π^ dq≤−1/ε\int\hat J_{N,\hat\pi}\,dq\le-1/\varepsilon∫J^N,π^​dq≤−1/ε otherwise, and qqq-optimal if q({y0∣J^N,π^(y0)=J^N∗(y0)})=1q(\{y_0\mid\hat J_{N,\hat\pi}(y_0)=\hat J^*_N(y_0)\})=1q({y0​∣J^N,π^​(y0​)=J^N∗​(y0​)})=1 (Definition 10.8).

Milestones

  1. Lemma 10.1: the process (η0,u0,…,ηk,uk)(\eta_0,u_0,\dots,\eta_k,u_k)(η0​,u0​,…,ηk​,uk​) generated in (ISI) by a Markov (PSI) policy has the law P^k[π^,φ(p)]\hat P_k[\hat\pi,\varphi(p)]P^k​[π^,φ(p)].
  2. Proposition 10.2: JN,π^(p)=∫J^N,π^ dφ(p)J_{N,\hat\pi}(p)=\int\hat J_{N,\hat\pi}\,d\varphi(p)JN,π^​(p)=∫J^N,π^​dφ(p) for Markov π^\hat\piπ^.
  3. Corollary 10.2.1: JN∗(p)≤∫J^N∗ dφ(p)J^*_N(p)\le\int\hat J^*_N\,d\varphi(p)JN∗​(p)≤∫J^N∗​dφ(p).
  4. Lemma 10.2: every (ISI) policy is matched in cost by some Markov (PSI) policy.
  5. Proposition 10.4: ε\varepsilonε-optimal nonrandomized (ISI) policies that depend on iki_kik​ only through ηk(p;ik)\eta_k(p;i_k)ηk​(p;ik​).
  6. Proposition 10.6: the identity maps on P(S)IkP(S)I_kP(S)Ik​ form a statistic sufficient for control.

Significance

Proposition 10.3 says that an imperfect-information problem loses nothing by being solved in the reduced model: the optimal cost is the φ(p)\varphi(p)φ(p)-average of the reduced optimal cost, and good reduced policies are good original policies. Combined with Proposition 10.6, every (ISI) model has such a reduction, so the finite-horizon dynamic programming theory of Chapter 8 (existence of ε\varepsilonε-optimal policies, the dynamic programming algorithm) transfers to partially observed problems on Borel spaces. Proposition 10.4 turns this into a structural statement about the original problem: nearly optimal controllers need to retain only the statistic.

These results are proved in the book. None of them is formalized: the platform's related results (Bäuerle–Rieder's partially observable models with observation densities, and the linear-quadratic-Gaussian separation theorem) work in different models and do not cover universally measurable policies, analytic constraints, or lower semianalytic costs. A machine-checked version makes the conditional-expectation bookkeeping of the reduction explicit, and the definitions of this mission (universal measurability, lower semianalytic functions, the book's extended integral, history measures built from universally measurable kernels) are reusable by every other chapter of the book.

Difficulty

The obvious argument says: replace the state by the statistic, observe that costs and transitions depend only on the statistic, and conclude. In the Borel setting each step is a measurability claim that the naive argument does not supply. The conditions of Definition 10.6 are almost-everywhere statements about conditional distributions under every pair (p,π)(p,\pi)(p,π), while the reduced model needs genuine kernels; the policies are only universally measurable, so integrals and compositions must be taken with respect to completions; the costs take the values ±∞\pm\infty±∞, so interchanging sums and integrals requires the finiteness assumptions (F±)(F^\pm)(F±) and (F^±)(\hat F^\pm)(F^±); and the inequality JN∗≥∫J^N∗ dφ(p)J^*_N\ge\int\hat J^*_N\,d\varphi(p)JN∗​≥∫J^N∗​dφ(p) requires producing, from an arbitrary history-dependent (ISI) policy, a Markov (PSI) policy with the same cost, which the naive argument does not do.

Formalization scope

  • Horizon. Only finite horizons N≥1N\ge1N≥1 are covered, hence only the cases (F+,F^+)(F^+,\hat F^+)(F+,F^+) and (F−,F^−)(F^-,\hat F^-)(F−,F^−) of the book's statements; the infinite-horizon cases (P,P^)(P,\hat P)(P,P^), (N,N^)(N,\hat N)(N,N^), (D,D^)(D,\hat D)(D,D^) are out of scope.
  • Extended reals. Costs live in EReal with the book's convention ∞−∞=+∞\infty-\infty=+\infty∞−∞=+∞ written out explicitly (badd, bsum, extIntegral); Mathlib's EReal subtraction (⊤−⊤=⊥\top-\top=\bot⊤−⊤=⊥) is never used where both terms can be infinite.
  • Spaces and measures. SSS, CCC, ZZZ, YkY_kYk​ are Borel spaces in the sense of Definition 7.7 with their Borel σ\sigmaσ-algebras; P(S)P(S)P(S) carries the weak topology and the Giry σ\sigmaσ-algebra. Policies are families of maps into ProbabilityMeasure C that are measurable for the completion of every probability measure. History measures are characterized by their values on rectangles. Families indexed by the stage are indexed by all of N\mathbb NN; only stages k<Nk<Nk<N are constrained.
  • Conditional statements. Conditions (22) and (23) are stated through the defining relations of conditional probability and expectation, for every ppp and every policy, with (23) required when g(xk,uk)g(x_k,u_k)g(xk​,uk​) is quasi-integrable.
  • Policies in Proposition 10.3. The (PSI) policies in the optimality transfers are Markov, as in Proposition 10.2.
  • No trivialization. Definition 10.6 is the full definition: analytic Γ^k\hat\Gamma_kΓ^k​ with full projection, Borel kernels t^k\hat t_kt^k​ satisfying (22) for every ppp and policy, and lower semianalytic g^k\hat g_kg^​k​ satisfying (23); a weaker notion would make Proposition 10.6 empty.

Contributions are welcome on any milestone. Basic facts that a full development needs, such as composition of universally measurable maps (Proposition 7.44), measurability of integrals against universally measurable kernels (Proposition 7.46), and existence of the history measures (Proposition 7.45), can be posed and proved as supporting lemmas; they are reusable across the book.

Selected references

  • D. P. Bertsekas and S. E. Shreve, Stochastic Optimal Control: The Discrete-Time Case, Academic Press, 1978; Athena Scientific, 1996, Chapter 10. https://web.mit.edu/dimitrib/www/soc.html
  • K. J. Åström, Optimal control of Markov processes with incomplete state information, Journal of Mathematical Analysis and Applications 10 (1965) 174–205. https://doi.org/10.1016/0022-247X(65)90154-X
  • C. Striebel, Sufficient statistics in the optimum control of stochastic systems, Journal of Mathematical Analysis and Applications 12 (1965) 576–592. https://doi.org/10.1016/0022-247X(65)90027-2
  • N. Bäuerle and U. Rieder, Markov Decision Processes with Applications to Finance, Springer, 2011, Chapter 5. https://doi.org/10.1007/978-3-642-18324-9
12 thms1 active userReviewed
Dynamic ProgrammingOperations Research·Captain: mikedeng1

Stochastic Optimal Control: The Discrete-Time Case IV: The Generalized Abstract Model — Restricted Policy Classes under ContractionTextbook

Why restricted policy classes

Abstract dynamic programming, in the form developed by Denardo (1967) and Bertsekas (1977), studies sequential decision problems through a single monotone mapping H(x,u,J)H(x,u,J)H(x,u,J): the cost of using control uuu at state xxx when the future is valued by the function JJJ. Chapters 2–5 of Bertsekas and Shreve, Stochastic Optimal Control: The Discrete-Time Case (1978; Athena Scientific reprint 1996), analyze this model when policies are arbitrary selectors μ:S→C\mu:S\to Cμ:S→C and HHH is defined on all extended-real functions on SSS.

That generality breaks down as soon as the state and control spaces are uncountable. A stochastic control problem on Borel spaces needs measurable policies, so that the expected cost is an integral rather than an outer integral, and the functions on which HHH acts must be measurable for the same reason. Chapter 6 of the book introduces a generalized abstract model in which the policies are drawn from a prescribed class M~\tilde MM~ and HHH is only defined on a prescribed class F~\tilde FF~ of functions. The examples on p. 94 are the models of Part II: universally measurable policies with lower semianalytic costs (Chapters 8–9), analytically measurable policies (Section 11.2), and the semicontinuous models of Definitions 8.7–8.8. Chapter 6 is the bridge that lets the abstract results of Part I be invoked for these models.

Setting

The data are a state space SSS, a control space CCC, nonempty constraint sets U(x)⊆CU(x)\subseteq CU(x)⊆C, and three restricted classes: sets of functions F∗⊂F~⊂FF^*\subset\tilde F\subset FF∗⊂F~⊂F, where FFF is the set of all functions S→[−∞,∞]S\to[-\infty,\infty]S→[−∞,∞], and a set M~\tilde MM~ of selectors μ:S→C\mu:S\to Cμ:S→C with μ(x)∈U(x)\mu(x)\in U(x)μ(x)∈U(x). The mapping H:S×C×F~→[−∞,∞]H:S\times C\times\tilde F\to[-\infty,\infty]H:S×C×F~→[−∞,∞] is monotone: J≤J′J\le J'J≤J′ in F~\tilde FF~ implies H(x,u,J)≤H(x,u,J′)H(x,u,J)\le H(x,u,J')H(x,u,J)≤H(x,u,J′). For μ∈M~\mu\in\tilde Mμ∈M~ and J∈F~J\in\tilde FJ∈F~,

Tμ(J)(x)=H[x,μ(x),J],T(J)(x)=inf⁡u∈U(x)H(x,u,J).T_\mu(J)(x)=H[x,\mu(x),J],\qquad T(J)(x)=\inf_{u\in U(x)}H(x,u,J).Tμ​(J)(x)=H[x,μ(x),J],T(J)(x)=u∈U(x)inf​H(x,u,J).

A policy is a sequence π=(μ0,μ1,… )\pi=(\mu_0,\mu_1,\dots)π=(μ0​,μ1​,…) with every μk∈M~\mu_k\in\tilde Mμk​∈M~; their set is Π~\tilde\PiΠ~. Given J0∈F∗J_0\in F^*J0​∈F∗ with J0>−∞J_0>-\inftyJ0​>−∞, the NNN-stage and infinite-horizon costs are

JN,π=(Tμ0⋯TμN−1)(J0),Jπ(x)=lim⁡N→∞JN,π(x),J_{N,\pi}=(T_{\mu_0}\cdots T_{\mu_{N-1}})(J_0),\qquad J_\pi(x)=\lim_{N\to\infty}J_{N,\pi}(x),JN,π​=(Tμ0​​⋯TμN−1​​)(J0​),Jπ​(x)=N→∞lim​JN,π​(x),

and the optimal costs are JN∗=inf⁡π∈Π~JN,πJ^*_N=\inf_{\pi\in\tilde\Pi}J_{N,\pi}JN∗​=infπ∈Π~​JN,π​ and J∗=inf⁡π∈Π~JπJ^*=\inf_{\pi\in\tilde\Pi}J_\piJ∗=infπ∈Π~​Jπ​. For a stationary policy (μ,μ,… )(\mu,\mu,\dots)(μ,μ,…) write JμJ_\muJμ​.

Five standing conditions tie the classes together: A.1 (every control u∈U(x)u\in U(x)u∈U(x) is the value μ(x)\mu(x)μ(x) of some μ∈M~\mu\in\tilde Mμ∈M~), A.2 (F∗F^*F∗ is closed under TTT and under adding constants), A.3 (F~\tilde FF~ is closed under every TμT_\muTμ​, μ∈M~\mu\in\tilde Mμ∈M~, and under adding constants), A.4 (ε\varepsilonε-minimizing selectors for T(J)T(J)T(J), J∈F∗J\in F^*J∈F∗, exist in M~\tilde MM~), and A.5 (F~\tilde FF~ and F∗F^*F∗ are closed under pointwise limits). Assumption C~\tilde CC~ asks for a closed subset Bˉ\bar BBˉ of the space BBB of bounded real functions with the sup norm ∥⋅∥\|\cdot\|∥⋅∥, containing J0J_0J0​ and invariant under TTT on Bˉ∩F∗\bar B\cap F^*Bˉ∩F∗ and under TμT_\muTμ​ on Bˉ∩F~\bar B\cap\tilde FBˉ∩F~, such that every JπJ_\piJπ​ exists and is real, each TμT_\muTμ​ is α\alphaα-Lipschitz on B∩F~B\cap\tilde FB∩F~, and every mmm-fold composition Tμ0⋯Tμm−1T_{\mu_0}\cdots T_{\mu_{m-1}}Tμ0​​⋯Tμm−1​​ is a ρ\rhoρ-contraction on Bˉ∩F~\bar B\cap\tilde FBˉ∩F~ for some ρ<1\rho<1ρ<1.

Formalization targets

Goal: Proposition 6.4 (p. 97)

Under A.1–A.5 and C~\tilde CC~: J∗∈Bˉ∩F∗J^*\in\bar B\cap F^*J∗∈Bˉ∩F∗ is the unique fixed point of TTT in Bˉ∩F∗\bar B\cap F^*Bˉ∩F∗, with T(J′)≤J′⇒J∗≤J′T(J')\le J'\Rightarrow J^*\le J'T(J′)≤J′⇒J∗≤J′ and J′≤T(J′)⇒J′≤J∗J'\le T(J')\Rightarrow J'\le J^*J′≤T(J′)⇒J′≤J∗; each JμJ_\muJμ​, μ∈M~\mu\in\tilde Mμ∈M~, is the unique fixed point of TμT_\muTμ​ in Bˉ∩F~\bar B\cap\tilde FBˉ∩F~;

lim⁡N→∞∥TN(J)−J∗∥=0  (J∈Bˉ∩F∗),lim⁡N→∞∥TμN(J)−Jμ∥=0  (J∈Bˉ∩F~);\lim_{N\to\infty}\|T^N(J)-J^*\|=0\ \ (J\in\bar B\cap F^*),\qquad\lim_{N\to\infty}\|T_\mu^N(J)-J_\mu\|=0\ \ (J\in\bar B\cap\tilde F);N→∞lim​∥TN(J)−J∗∥=0  (J∈Bˉ∩F∗),N→∞lim​∥TμN​(J)−Jμ​∥=0  (J∈Bˉ∩F~);

a stationary (μ∗,μ∗,… )∈Π~(\mu^*,\mu^*,\dots)\in\tilde\Pi(μ∗,μ∗,…)∈Π~ is optimal iff Tμ∗(J∗)=T(J∗)T_{\mu^*}(J^*)=T(J^*)Tμ∗​(J∗)=T(J∗); and for every ε>0\varepsilon>0ε>0 some stationary policy in Π~\tilde\PiΠ~ satisfies ∥J∗−Jμε∥≤ε\|J^*-J_{\mu_\varepsilon}\|\le\varepsilon∥J∗−Jμε​​∥≤ε.

Milestones

In attack order:

  1. Proposition 6.3(a) (p. 96) — under A.1–A.4 and the exact selection assumption, a uniformly NNN-stage optimal policy exists iff the infimum in Tk+1(J0)(x)=inf⁡u∈U(x)H[x,u,Tk(J0)]T^{k+1}(J_0)(x)=\inf_{u\in U(x)}H[x,u,T^k(J_0)]Tk+1(J0​)(x)=infu∈U(x)​H[x,u,Tk(J0​)] is attained for each x∈Sx\in Sx∈S and k<Nk<Nk<N.
  2. Proposition 6.5(a) (p. 97) — under A.1–A.5, C~\tilde CC~ and exact selection: if for each xxx some policy in Π~\tilde\PiΠ~ is optimal at xxx, then an optimal stationary policy exists in Π~\tilde\PiΠ~.

Further results of the chapter

The other results of Sections 6.2–6.3 are posed in the mission as separate theorems:

  • Proposition 6.2 — π∗\pi^*π∗ is uniformly NNN-stage optimal iff (Tμk∗TN−k−1)(J0)=TN−k(J0)(T_{\mu_k^*}T^{N-k-1})(J_0)=T^{N-k}(J_0)(Tμk∗​​TN−k−1)(J0​)=TN−k(J0​) for k<Nk<Nk<N; such a policy forces JN∗=TN(J0)J^*_N=T^N(J_0)JN∗​=TN(J0​).
  • Proposition 6.1(a) — under Assumption F~.2\tilde F.2F~.2 and Jk∗>−∞J^*_k>-\inftyJk∗​>−∞: JN∗=TN(J0)J^*_N=T^N(J_0)JN∗​=TN(J0​) and NNN-stage ε\varepsilonε-optimal policies exist in Π~\tilde\PiΠ~.
  • Proposition 6.1(b) — under Assumption F~.3\tilde F.3F~.3 and Jk,π<∞J_{k,\pi}<\inftyJk,π​<∞: JN∗=TN(J0)J^*_N=T^N(J_0)JN∗​=TN(J0​) and {εn}\{\varepsilon_n\}{εn​}-dominated convergence to optimality.
  • Proposition 6.3(b) — compact level sets Uk(x,λ)U_k(x,\lambda)Uk​(x,λ) give both JN∗=TN(J0)J^*_N=T^N(J_0)JN∗​=TN(J0​) and a uniformly NNN-stage optimal policy.
  • Proposition 6.5(b) — compact level sets of the iterates Tk(J)T^k(J)Tk(J), k≥kˉk\ge\bar kk≥kˉ, give an optimal stationary policy.

Significance

Proposition 6.4 is the statement that makes value iteration, Bellman's equation and stationary ε\varepsilonε-optimal policies available for discounted problems whose admissible policies are restricted, for instance to measurable ones. Without it, each measurable model would need its own fixed-point argument. The finite-horizon Propositions 6.1–6.3 play the same role for the dynamic programming algorithm JN∗=TN(J0)J^*_N=T^N(J_0)JN∗​=TN(J0​), and their hypotheses (F~.3\tilde F.3F~.3, exact selection) are exactly what Chapters 7–8 verify for universally measurable policies.

The book states Propositions 6.4 and 6.5 without proof (p. 97), referring to the proofs of Chapter 4; Propositions 6.1–6.3 are justified by "nearly verbatim repetition" of Chapter 3. A formalization therefore supplies proofs that are only indicated in print, and checks that A.1–A.5 really suffice for each step of the Chapter 3–4 arguments. None of these results has a machine-checked proof that we know of; the companion missions of this series formalize the unrestricted special case (F∗=F~=FF^*=\tilde F=FF∗=F~=F, M~=M\tilde M=MM~=M) of Chapters 3 and 4.

Difficulty

The Chapter 4 proof of Proposition 4.2 applies the contraction mapping theorem to TTT on Bˉ\bar BBˉ. Here the obvious transcription fails at two points. First, TTT maps Bˉ∩F∗\bar B\cap F^*Bˉ∩F∗ into itself but TμT_\muTμ​ only maps Bˉ∩F~\bar B\cap\tilde FBˉ∩F~ into itself, so the fixed-point theorem must be applied on two different sets, and these are closed only because of A.5. Second, every argument that picks a near-minimizing selector at each state must produce a selector in M~\tilde MM~: pointwise choices are no longer allowed, and A.1, A.4 and the exact selection assumption are the only sources of admissible selectors. Proofs of Chapter 3–4 that build a policy state by state cannot be copied.

Formalization scope

Functions on SSS are S → EReal. HHH is a total Lean function, but monotonicity is assumed only on F~\tilde FF~ and every statement evaluates HHH only at functions of F~\tilde FF~. JπJ_\piJπ​ is limUnder; Assumption C~\tilde CC~ makes the limit exist. BBB is Mathlib's ℓ∞(S,R)\ell^\infty(S,\mathbb R)ℓ∞(S,R); a bound ∥G−G′∥≤c\|G-G'\|\le c∥G−G′∥≤c between extended-real functions means both are real everywhere and ∣G(x)−G′(x)∣≤c|G(x)-G'(x)|\le c∣G(x)−G′(x)∣≤c, which is how the book's convention ∞−∞=∞\infty-\infty=\infty∞−∞=∞ reads a norm of a difference. No statement adds values of opposite infinite sign, so Mathlib's EReal addition agrees with the book's wherever it is used. JN∗J^*_NJN∗​ and J∗J^*J∗ are infima over Π~\tilde\PiΠ~ only, the ε\varepsilonε-optimality notions keep the book's two-case form at −∞-\infty−∞, and NNN is a positive integer.

The chapter collapses to Chapters 3–4 if F∗=F~=FF^*=\tilde F=FF∗=F~=F or M~=M\tilde M=MM~=M is built in; here F∗F^*F∗, F~\tilde FF~ and M~\tilde MM~ are arbitrary and constrained only by A.1–A.5, and J∗J^*J∗ is never defined as a fixed point.

A complete development needs the mmm-step contraction mapping theorem on a closed subset of ℓ∞\ell^\inftyℓ∞, monotonicity lemmas for TTT and TμT_\muTμ​, and the restricted-class versions of Propositions 3.1–3.4 and 4.1–4.4. These are reusable for the Borel models of Chapters 8–9. Proofs of any milestone, and sorry-free lemmas about the Assumption C~\tilde CC~ contraction, are welcome.

Selected references

  • D. P. Bertsekas and S. E. Shreve, Stochastic Optimal Control: The Discrete-Time Case, Academic Press, 1978; Athena Scientific, 1996, Chapter 6. https://web.mit.edu/dimitrib/www/soc.html
  • D. P. Bertsekas, Monotone mappings with application in dynamic programming, SIAM J. Control and Optimization 15(3), 1977, 438–464. https://doi.org/10.1137/0315031
  • E. V. Denardo, Contraction mappings in the theory underlying dynamic programming, SIAM Review 9(2), 1967, 165–177. https://doi.org/10.1137/1009030
7 thms1 active userReviewed
CombinatoricsGraph TheoryOperations Research+1·Captain: mikedeng1

Maximal Flow Through a Network II: In an ab-Planar Network Some Chain from Source to Sink Meets Every Cut Exactly OnceResearch Paper

Motivation

The maximum flow problem asks how much of a commodity can be shipped from a source to a sink through a network whose arcs have limited capacities. L. R. Ford, Jr. and D. R. Fulkerson's 1956 paper Maximal Flow Through a Network proved the minimal cut theorem: the largest flow value equals the smallest total capacity of a set of arcs that separates source from sink. That theorem is formalized in the companion mission Maximal Flow Through a Network I.

The second section of the same paper treats a special class of networks, those that remain planar after an arc from source to sink is added. For these networks the paper shows that one particular source–sink chain crosses every minimal separating set exactly once. This structural fact turns the minimal cut theorem into a simple computing procedure: repeatedly push as much flow as possible along such a chain and delete the arcs it saturates. The paper notes that G. Dantzig had conjectured, before the minimal cut theorem was proved, that this procedure yields a maximal flow on planar networks. The same "uppermost path" idea underlies later algorithms for maximum flow in planar graphs with source and sink on a common face (Itai and Shiloach, 1979).

The statement is short and purely combinatorial in its conclusion, but its hypothesis is topological. This mission isolates that theorem.

Setting

A network NNN has a finite set VVV of vertices and a finite set EEE of arcs. Each arc eee joins two distinct end vertices, written tail(e)\mathrm{tail}(e)tail(e) and head(e)\mathrm{head}(e)head(e); arcs carry no direction, and two arcs may join the same pair of vertices. Two distinct vertices are distinguished, the source aaa and the sink bbb, and each arc carries a positive capacity (capacities play no role in the target below).

A chain joining uuu and www is a set CCC of distinct arcs that can be arranged as α1(v0v1),α2(v1v2),…,αk(vk−1vk)\alpha_1(v_0v_1), \alpha_2(v_1v_2), \dots, \alpha_k(v_{k-1}v_k)α1​(v0​v1​),α2​(v1​v2​),…,αk​(vk−1​vk​) with v0=uv_0 = uv0​=u, vk=wv_k = wvk​=w, and the vertices v0,…,vkv_0, \dots, v_kv0​,…,vk​ pairwise distinct; each arc may be traversed in either direction. The empty set is the null chain from uuu to uuu.

A set DDD of arcs is a disconnecting set if every chain joining aaa and bbb contains an arc of DDD. A disconnecting set none of whose proper subsets is disconnecting is a cut.

The network is ab-planar if the graph of NNN, together with one additional arc joining aaa and bbb, can be drawn in the plane without crossings: vertices go to distinct points of R2\mathbb R^2R2; each arc, including the added arc ababab, goes to an injective continuous path between the points of its end vertices; no arc passes through a vertex other than its ends; and two distinct arcs meet only at endpoints of both. In Lean the drawing is the structure ABPlaneDrawing N, and NNN is ab-planar when Nonempty (ABPlaneDrawing N). The section's standing assumption is that no arc of NNN already joins aaa and bbb.

Formalization targets

Goal: Theorem 2 (p. 403)

If NNN is ab-planar, no arc of NNN joins aaa and bbb, and some chain joins aaa and bbb, then

∃ T a chain joining a and b  such that  ∣T∩D∣=1  for every cut D of N.\exists\, T \text{ a chain joining } a \text{ and } b \ \text{ such that }\ |T \cap D| = 1 \ \text{ for every cut } D \text{ of } N.∃T a chain joining a and b  such that  ∣T∩D∣=1  for every cut D of N.

This is FordFulkerson56.Planar.ab_planar_exists_chain_meeting_each_cut_once. "Precisely once" is exact cardinality one, neither "at least once" (true of every chain) nor "at most once".

Milestone: a chain meeting a cut in one prescribed arc (proof of Theorem 2, p. 403)

For every network NNN, every cut DDD and every arc α∈D\alpha \in Dα∈D, there is a chain CCC joining aaa and bbb with C∩D={α}C \cap D = \{\alpha\}C∩D={α}. No planarity is involved; the statement is what the minimality of a cut provides to the proof.

Further item: the Fig. 2 example (p. 403)

In the "gas, water, electricity" graph K3,3K_{3,3}K3,3​ with the arc ababab removed, every chain joining aaa and bbb meets some cut in three arcs. This network is not ab-planar, so the example shows that the planarity hypothesis of Theorem 2 cannot be dropped.

Significance

Theorem 2 and the minimal cut theorem together give the paper's procedure for planar networks: if TTT meets every cut once, then imposing a flow kkk on TTT lowers the value of every cut by exactly kkk, so the minimal cut value, and hence the maximal flow value, drops by kkk. Saturated arcs can then be deleted and the step repeated. Without the "exactly once" property the reduction could overshoot the cut structure, and the greedy step would not be justified. The theorem is also one of the earliest instances of the link between planarity and cut structure that later underlies planar duality arguments for minimum cuts.

The result has been known since 1956 and is not open. No machine-checked version is recorded on the platform, and Mathlib, at the pinned revision, has neither planar graphs nor the Jordan curve theorem. A formal proof would be the first formalized statement about source–sink planar networks in this library, and the counterexample item records, as a checkable fact, that the hypothesis is necessary.

Difficulty

The conclusion is combinatorial while the hypothesis is a drawing in R2\mathbb R^2R2. The paper's proof normalises the drawing (the added arc ababab on the outer boundary, the graph in a vertical strip with aaa on the left line and bbb on the right), selects the "top-most" chain from aaa to bbb, and argues that a chain meeting a cut below the top-most chain must cross another such chain. Each of these steps rests on plane topology: the existence of the outer region, the meaning of "top-most", and the fact that two chains with interleaved endpoints on a boundary must intersect, which is a form of the Jordan curve theorem.

The naive purely combinatorial route fails: the analogous statement for arbitrary networks is false (Fig. 2), so any argument has to use the drawing somewhere. Replacing the drawing by a combinatorial embedding (rotation systems, faces) is possible but then requires proving that the two notions agree, which is again Jordan-curve territory.

Formalization scope

Conventions committed to in the Lean statements:

  • Vertices and arcs are finite types V, E with decidable equality. Arcs are undirected, may be parallel, and have two distinct end vertices. Source and sink are distinct, capacities are positive (structure Network).
  • A chain is a Finset E that is the arc set of some arrangement as a simple path (IsChainWalk, IsChain); the null chain is allowed.
  • IsDisconnecting and IsCut quantify over all chains joining source and sink; a cut is a disconnecting set no proper subset of which is disconnecting.
  • ab-planarity is a plane drawing of the graph with the extra arc indexed by none : Option E, with injective Paths in ℝ × ℝ as arcs.

Hypotheses of the goal: hno_ab, the standing assumption of §2 (no arc joins aaa and bbb, p. 403); hconn, that some chain joins aaa and bbb. The second is not stated in the paper; its proof starts from "the chain joining a and b which is top-most", which presupposes one, and without it the statement is false (if aaa and bbb are disconnected, the empty set is a cut and no chain exists).

The drawing structure is satisfiable (a three-vertex path network has an explicit drawing), so the planarity hypothesis is not vacuous; and it covers the added arc ababab and all crossings, so K3,3K_{3,3}K3,3​ minus ababab is not ab-planar and the goal is not refuted by the paper's own example. A formalization that dropped the arc ababab from the drawing, or quantified over disconnecting sets instead of cuts, would state a false theorem and is ruled out.

A complete development needs basic plane topology for paths in R2\mathbb R^2R2 (a Jordan-curve-type separation lemma for simple closed curves, or an equivalent statement about crossing paths in a strip), together with combinatorial lemmas about chains (concatenation and shortcutting of chains at a common vertex). The topological lemmas are reusable well beyond this mission. Proofs through a combinatorial embedding are welcome, provided the equivalence with ABPlaneDrawing is proved.

Selected references

  • L. R. Ford, Jr. and D. R. Fulkerson, Maximal Flow Through a Network, Canadian Journal of Mathematics 8 (1956), 399–404. https://doi.org/10.4153/CJM-1956-045-5
  • H. Whitney, Non-separable and planar graphs, Transactions of the American Mathematical Society 34 (1932), 339–362. https://doi.org/10.1090/S0002-9947-1932-1501641-2
  • A. Itai and Y. Shiloach, Maximum flow in planar networks, SIAM Journal on Computing 8 (1979), 135–150. https://doi.org/10.1137/0208012
  • H. Whitney, Planar graphs, Fundamenta Mathematicae 21 (1933), 73–84. https://doi.org/10.4064/fm-21-1-73-84
7 thms1 active userReviewed
Machine LearningOperations ResearchReinforcement Learning·Captain: mikedeng1

Twice Regularized MDPs and the Equivalence Between Robustness and Regularization 2: The Greedy Policy of the R2 Optimal Value Is the Unique Optimal R2 PolicyResearch Paper

Motivation

A robust Markov decision process (robust MDP) evaluates a policy against the worst transition kernel and reward in an uncertainty set around a nominal model (P0,r0)(P_0, r_0)(P0​,r0​). It is the standard model for planning when the dynamics are estimated from data (Iyengar 2005; Nilim and El Ghaoui 2005; Wiesemann, Kuhn and Rustem 2013). Its Bellman update contains an inner optimization over the uncertainty set at every state, which makes robust planning costly when the sets are not (s,a)(s,a)(s,a)-rectangular.

Derman, Geist and Mannor (arXiv:2110.06267, NeurIPS 2021) show that, for sss-rectangular ball uncertainty sets, this inner optimization can be replaced by an explicit penalty that depends both on the policy and on the value function. The resulting twice regularized (R²) MDPs have Bellman operators with no inner optimization over models. The first mission of this series formalizes the robust–regularized equivalence (Theorem 4.1 of the paper). This mission formalizes Section 5: the R² Bellman operators are monotone and contracting under a bound on the transition radius, and the greedy policy of the R² optimal value is optimal.

Setting

Let S\mathcal SS and A\mathcal AA be finite nonempty sets of states and actions, γ∈(0,1)\gamma\in(0,1)γ∈(0,1) a discount factor, P0(s′∣s,a)P_0(s'\mid s,a)P0​(s′∣s,a) a transition kernel and r0(s,a)r_0(s,a)r0​(s,a) a reward. A policy π∈ΔAS\pi\in\Delta_{\mathcal A}^{\mathcal S}π∈ΔAS​ assigns to each state a probability distribution πs\pi_sπs​ on A\mathcal AA. For v∈RSv\in\mathbb R^{\mathcal S}v∈RS write qs(a)=r0(s,a)+γ∑s′P0(s′∣s,a)v(s′)q_s(a)=r_0(s,a)+\gamma\sum_{s'}P_0(s'\mid s,a)v(s')qs​(a)=r0​(s,a)+γ∑s′​P0​(s′∣s,a)v(s′) and

[T(P0,r0)πv](s)=∑aπs(a) qs(a).[T^\pi_{(P_0,r_0)}v](s)=\sum_a\pi_s(a)\,q_s(a).[T(P0​,r0​)π​v](s)=a∑​πs​(a)qs​(a).

All norms ∥⋅∥\|\cdot\|∥⋅∥ below are ℓ2\ell_2ℓ2​-norms, ∥a∥=(∑za(z)2)1/2\|a\|=\big(\sum_z a(z)^2\big)^{1/2}∥a∥=(∑z​a(z)2)1/2; ∥⋅∥∞\|\cdot\|_\infty∥⋅∥∞​ is the sup norm.

Fix nonnegative radii αsr,αsP\alpha^r_s,\alpha^P_sαsr​,αsP​ for each state. The R² regularizer is Ωv,R2(πs)=∥πs∥ (αsr+αsPγ∥v∥)\Omega_{v,\mathrm R^2}(\pi_s)=\|\pi_s\|\,(\alpha^r_s+\alpha^P_s\gamma\|v\|)Ωv,R2​(πs​)=∥πs​∥(αsr​+αsP​γ∥v∥), and the R² Bellman operators are

[Tπ,R2v](s)=[T(P0,r0)πv](s)−Ωv,R2(πs),[T∗,R2v](s)=max⁡π∈ΔAS[Tπ,R2v](s).[T^{\pi,\mathrm R^2}v](s)=[T^\pi_{(P_0,r_0)}v](s)-\Omega_{v,\mathrm R^2}(\pi_s),\qquad [T^{*,\mathrm R^2}v](s)=\max_{\pi\in\Delta^{\mathcal S}_{\mathcal A}}[T^{\pi,\mathrm R^2}v](s).[Tπ,R2v](s)=[T(P0​,r0​)π​v](s)−Ωv,R2​(πs​),[T∗,R2v](s)=π∈ΔAS​max​[Tπ,R2v](s).

A policy π\piπ is greedy for vvv when Tπ,R2v=T∗,R2vT^{\pi,\mathrm R^2}v=T^{*,\mathrm R^2}vTπ,R2v=T∗,R2v.

Assumption 5.1 (bounded radius). For each sss there is ϵs>0\epsilon_s>0ϵs​>0 with

αsP≤min⁡(1−γ−ϵsγ∣S∣ ; min⁡u∈R+A,∥u∥=1, w∈R+S,∥w∥=1 ∑a,s′u(a)P0(s′∣s,a)w(s′)),\alpha^P_s\le\min\Big(\frac{1-\gamma-\epsilon_s}{\gamma\sqrt{|\mathcal S|}}\ ;\ \min_{u\in\mathbb R^{\mathcal A}_+,\|u\|=1,\ w\in\mathbb R^{\mathcal S}_+,\|w\|=1}\ \sum_{a,s'}u(a)P_0(s'\mid s,a)w(s')\Big),αsP​≤min(γ∣S∣​1−γ−ϵs​​ ; u∈R+A​,∥u∥=1, w∈R+S​,∥w∥=1min​ a,s′∑​u(a)P0​(s′∣s,a)w(s′)),

and ϵ∗=min⁡sϵs\epsilon_*=\min_s\epsilon_sϵ∗​=mins​ϵs​. The R² value function vπ,R2v^{\pi,\mathrm R^2}vπ,R2 of a policy and the R² optimal value v∗,R2v^{*,\mathrm R^2}v∗,R2 are the fixed points of Tπ,R2T^{\pi,\mathrm R^2}Tπ,R2 and T∗,R2T^{*,\mathrm R^2}T∗,R2.

Formalization targets

Goal: Theorem 5.1 (p. 8)

Under Assumption 5.1, T∗,R2T^{*,\mathrm R^2}T∗,R2 and every Tπ,R2T^{\pi,\mathrm R^2}Tπ,R2 have unique fixed points; a greedy policy π∗,R2\pi^{*,\mathrm R^2}π∗,R2 for v∗,R2v^{*,\mathrm R^2}v∗,R2 exists, and every such policy satisfies

vπ∗,R2,R2=v∗,R2 ≥ vπ,R2for all π∈ΔAS;v^{\pi^{*,\mathrm R^2},\mathrm R^2}=v^{*,\mathrm R^2}\ \ge\ v^{\pi,\mathrm R^2}\qquad\text{for all }\pi\in\Delta^{\mathcal S}_{\mathcal A};vπ∗,R2,R2=v∗,R2 ≥ vπ,R2for all π∈ΔAS​;

every optimal policy is greedy; and when αsr>0\alpha^r_s>0αsr​>0 for all sss the greedy policy is unique, hence the unique optimal R² policy.

Milestones

  1. Proposition 2.1 (p. 3): for Ω\OmegaΩ strongly convex on the simplex, Ω∗(y)=max⁡a∈Δ⟨a,y⟩−Ω(a)\Omega^*(y)=\max_{a\in\Delta}\langle a,y\rangle-\Omega(a)Ω∗(y)=maxa∈Δ​⟨a,y⟩−Ω(a) is differentiable with Lipschitz gradient equal to the unique maximizer, satisfies Ω∗(y+c1)=Ω∗(y)+c\Omega^*(y+c\mathbb 1)=\Omega^*(y)+cΩ∗(y+c1)=Ω∗(y)+c, and is non-decreasing.
  2. Proposition 5.1 (i) (p. 8): v1≤v2v_1\le v_2v1​≤v2​ implies Tπ,R2v1≤Tπ,R2v2T^{\pi,\mathrm R^2}v_1\le T^{\pi,\mathrm R^2}v_2Tπ,R2v1​≤Tπ,R2v2​ and T∗,R2v1≤T∗,R2v2T^{*,\mathrm R^2}v_1\le T^{*,\mathrm R^2}v_2T∗,R2v1​≤T∗,R2v2​.
  3. Proposition 5.1 (iii) (p. 8):
∥Tπ,R2v1−Tπ,R2v2∥∞≤(1−ϵ∗)∥v1−v2∥∞,∥T∗,R2v1−T∗,R2v2∥∞≤(1−ϵ∗)∥v1−v2∥∞.\|T^{\pi,\mathrm R^2}v_1-T^{\pi,\mathrm R^2}v_2\|_\infty\le(1-\epsilon_*)\|v_1-v_2\|_\infty,\qquad \|T^{*,\mathrm R^2}v_1-T^{*,\mathrm R^2}v_2\|_\infty\le(1-\epsilon_*)\|v_1-v_2\|_\infty.∥Tπ,R2v1​−Tπ,R2v2​∥∞​≤(1−ϵ∗​)∥v1​−v2​∥∞​,∥T∗,R2v1​−T∗,R2v2​∥∞​≤(1−ϵ∗​)∥v1​−v2​∥∞​.

Significance

Theorem 5.1 is the R² counterpart of the fundamental theorem of discounted dynamic programming: optimal R² values are achieved by stationary policies obtained by a single greedy step. Together with the contraction of Proposition 5.1 (iii) it justifies the R² modified policy iteration algorithm of the paper, whose greedy step is a projection onto the simplex rather than a robust max–min problem. Combined with the first mission of the series, which identifies the robust value of an sss-rectangular ball-constrained MDP with the optimum of an R²-regularized program, it gives a route to robust planning at the cost of regularized planning.

The results are proved in the paper (App. C), partly by reference to Geist, Scherrer and Pietquin (2019) for the optimality operator. No machine-checked proof of any of them exists; this mission produces the first. Prop. 2.1 is a general fact of convex analysis (Danskin-type smoothness of a conjugate on the simplex) that is reusable for any regularized MDP or entropy-regularized game.

Difficulty

The R² evaluation operator is not affine: the value regularizer −αsPγ∥πs∥ ∥v∥-\alpha^P_s\gamma\|\pi_s\|\,\|v\|−αsP​γ∥πs​∥∥v∥ is concave in vvv and decreases as ∥v∥\|v\|∥v∥ grows. Monotonicity therefore does not follow from the positivity of P0P_0P0​ as in the standard case; it requires the second bound of Assumption 5.1, which compares the ℓ2\ell_2ℓ2​ variation of ∥v∥\|v\|∥v∥ with the minimal nonnegative bilinear form of P0(⋅∣s,⋅)P_0(\cdot\mid s,\cdot)P0​(⋅∣s,⋅). Likewise the contraction modulus is not γ\gammaγ but 1−ϵ∗1-\epsilon_*1−ϵ∗​, because the regularizer is ∣S∣\sqrt{|\mathcal S|}∣S∣​-Lipschitz between the ℓ2\ell_2ℓ2​ and sup norms. The optimality step of the classical proof uses linearity of TπT^\piTπ when comparing values of policies; here only monotonicity and contraction are available. Uniqueness of the greedy policy rests on strict concavity on the simplex, which holds only when the regularization weight is positive.

Formalization scope

States and actions are finite nonempty types; transitions are arrays P₀ : S → A → S → ℝ with the published predicate IsTransitionKernel; value functions are S → ℝ with the pointwise order. The ℓ2\ell_2ℓ2​-norm is an explicit l2norm (Mathlib's norm on S → ℝ is the sup norm, used only for ∥⋅∥∞\|\cdot\|_\infty∥⋅∥∞​). T∗,R2v(s)T^{*,\mathrm R^2}v(s)T∗,R2v(s) is the real supremum over the simplex ΔA\Delta_{\mathcal A}ΔA​ (attained), and the inner minimum of Assumption 5.1 is the real infimum over nonnegative ℓ2\ell_2ℓ2​-unit vectors; the witnesses ϵs\epsilon_sϵs​ are explicit. Greedy policies are a predicate, never a function, and the R² value functions are not defined by choice: the goal asserts their existence and uniqueness and speaks about the fixed points.

Disclosed deviations from the page. Assumption 5.1 is a hypothesis of Theorem 5.1 (its proof assumes it). The uniqueness clause of Theorem 5.1 additionally assumes αsr>0\alpha^r_s>0αsr​>0 for all sss: with one state, two actions, zero reward and zero radii every policy is greedy and optimal. Proposition 2.1 assumes Ω\OmegaΩ continuous on the simplex, without which the maximum need not be attained, and strong convexity is Mathlib's StrongConvexOn for some modulus (norm-independent in finite dimension). Proposition 5.1 (ii) is false as printed and is not drafted: with one state, one action, P0=1P_0=1P0​=1, r0=0r_0=0r0​=0, γ=1/2\gamma=1/2γ=1/2, αr=0\alpha^r=0αr=0, αP=1/2\alpha^P=1/2αP=1/2, ϵ=1/4\epsilon=1/4ϵ=1/4, one has Tv=v/2−∣v∣/4Tv=v/2-|v|/4Tv=v/2−∣v∣/4, and v1=−1v_1=-1v1​=−1, c=1c=1c=1 give T(v1+c)=0>−1/4=Tv1+γcT(v_1+c)=0>-1/4=Tv_1+\gamma cT(v1​+c)=0>−1/4=Tv1​+γc. Remark 5.1, Algorithm 1 and the ℓp\ell_pℓp​ variant of App. C.1 are out of scope. The inner minimum of Assumption 5.1 is 000 whenever some P0(s′∣s,a)=0P_0(s'\mid s,a)=0P0​(s′∣s,a)=0, forcing αsP=0\alpha^P_s=0αsP​=0; this is the assumption as printed.

A formalization in which ∥⋅∥\|\cdot\|∥⋅∥ is the sup norm, the inner minimum ranges over all unit vectors (making the assumption unsatisfiable), or the value functions are postulated rather than shown to exist would be trivial or wrong; the drafted statements avoid all three. Contributions welcome: Prop. 2.1 as a general convex-analysis lemma, Banach fixed-point plumbing for S → ℝ with the sup norm, and the strict concavity of p↦⟨p,q⟩−c∥p∥p\mapsto\langle p,q\rangle-c\|p\|p↦⟨p,q⟩−c∥p∥ on the simplex.

Selected references

  • E. Derman, M. Geist, S. Mannor, Twice regularized MDPs and the equivalence between robustness and regularization, NeurIPS 2021. arXiv:2110.06267v1
  • M. Geist, B. Scherrer, O. Pietquin, A theory of regularized Markov decision processes, ICML 2019. arXiv:1901.11275
  • A. Nilim, L. El Ghaoui, Robust control of Markov decision processes with uncertain transition matrices, Operations Research 53(5), 2005. doi:10.1287/opre.1050.0216
  • G. N. Iyengar, Robust dynamic programming, Mathematics of Operations Research 30(2), 2005. doi:10.1287/moor.1040.0129
  • W. Wiesemann, D. Kuhn, B. Rustem, Robust Markov decision processes, Mathematics of Operations Research 38(1), 2013. doi:10.1287/moor.1120.0566
  • A. Mensch, M. Blondel, Differentiable dynamic programming for structured prediction and attention, ICML 2018. arXiv:1802.03676
9 thms1 active userReviewed
AnalysisDifferential Geometry·Captain: mikedeng1

Projection-like Retractions on Matrix Manifolds II: The Metric Projection onto a C^k Submanifold Is Locally Unique and C^(k−1), so Projecting a Tangent Step Is a RetractionResearch Paper

Motivation

Many optimization problems in statistics, signal processing and control are posed over sets of matrices with a constraint that makes them curved: matrices of fixed rank, matrices with orthonormal columns, symmetric matrices with a prescribed spectrum. Algorithms on such sets ("optimization on manifolds") compute a step in the tangent space at the current point, as in a vector space, and then need a rule that brings the point x+ux+ux+u back to the set. A retraction is such a rule; the notion was introduced by Adler, Dedieu, Margulies, Martens and Shub for Newton's method on Riemannian manifolds (IMA J. Numer. Anal. 2002) and is the basic building block of the algorithms in Absil, Mahony and Sepulchre's monograph (Princeton, 2008). Any retraction preserves the local convergence of Newton's method, so the choice among retractions is about cost and convenience.

The most natural candidate is to project x+ux+ux+u back onto the manifold: take the nearest point. Absil and Malick (SIAM J. Optim. 2012; preprint HAL hal-00651608) show in §3.1 that this projective retraction is always a valid retraction of maximal smoothness, and then compute it for fixed-rank, spectral and Stiefel manifolds. The underlying fact, that the nearest-point map onto a CkC^kCk submanifold is locally single-valued and Ck−1C^{k-1}Ck−1, is classical (the paper cites Lewis and Malick, Math. Oper. Res. 2008); the paper gives a short proof through the inverse function theorem on the normal bundle. This mission formalizes §3.1: that lemma and the resulting retraction.

Setting

Let E\mathcal EE be a Euclidean space, a finite-dimensional real inner product space, of dimension nnn (in the paper's examples, Rn×m\mathbb R^{n\times m}Rn×m with the Frobenius inner product). A set M⊆E\mathcal M\subseteq\mathcal EM⊆E is a submanifold of class CkC^kCk and dimension ddd around xˉ\bar xxˉ if xˉ∈M\bar x\in\mathcal Mxˉ∈M and there are an open neighbourhood UEU_{\mathcal E}UE​ of xˉ\bar xxˉ and a CkC^kCk diffeomorphism φ\varphiφ from UEU_{\mathcal E}UE​ onto an open subset of Rn\mathbb R^nRn with

M∩UE={x∈UE: φd+1(x)=⋯=φn(x)=0}.\mathcal M\cap U_{\mathcal E}=\{x\in U_{\mathcal E}:\ \varphi_{d+1}(x)=\cdots=\varphi_n(x)=0\}.M∩UE​={x∈UE​: φd+1​(x)=⋯=φn​(x)=0}.

The tangent space TM(x)T_{\mathcal M}(x)TM​(x) is the linear subspace of E\mathcal EE spanned by the tangent cone of M\mathcal MM at xxx, and the normal space is NM(x)=TM(x)⊥N_{\mathcal M}(x)=T_{\mathcal M}(x)^\perpNM​(x)=TM​(x)⊥. The tangent bundle and normal bundle are TM={(x,u):x∈M, u∈TM(x)}T\mathcal M=\{(x,u):x\in\mathcal M,\ u\in T_{\mathcal M}(x)\}TM={(x,u):x∈M, u∈TM​(x)} and NM={(x,v):x∈M, v∈NM(x)}N\mathcal M=\{(x,v):x\in\mathcal M,\ v\in N_{\mathcal M}(x)\}NM={(x,v):x∈M, v∈NM​(x)}, subsets of E×E\mathcal E\times\mathcal EE×E. PTM(x)P_{T_{\mathcal M}(x)}PTM​(x)​ denotes the orthogonal projector onto TM(x)T_{\mathcal M}(x)TM​(x).

The projection of x∈Ex\in\mathcal Ex∈E onto M\mathcal MM is the set of nearest points,

PM(x)=argmin⁡{∥x−y∥: y∈M},P_{\mathcal M}(x)=\operatorname{argmin}\{\|x-y\|:\ y\in\mathcal M\},PM​(x)=argmin{∥x−y∥: y∈M},

which may be empty (if M\mathcal MM is not closed) or contain several points (if M\mathcal MM is not convex).

A map RRR from TMT\mathcal MTM to M\mathcal MM is a retraction around xˉ\bar xxˉ (Definition 2.1) if on some neighbourhood U\mathcal UU of (xˉ,0)(\bar x,0)(xˉ,0) in TMT\mathcal MTM it is of class Ck−1C^{k-1}Ck−1, satisfies R(x,0)=xR(x,0)=xR(x,0)=x, and DR(x,⋅)(0)=idTM(x)\mathrm DR(x,\cdot)(0)=\mathrm{id}_{T_{\mathcal M}(x)}DR(x,⋅)(0)=idTM​(x)​ for (x,0)∈U(x,0)\in\mathcal U(x,0)∈U.

The formal retraction predicate includes k≥2k\ge2k≥2 and a local submanifold chart at xˉ\bar xxˉ; this ensures that its base point lies on M\mathcal MM.

Formalization targets

Goal: Proposition 3.2 (projective retraction)

For M\mathcal MM a CkC^kCk submanifold (k≥2k\ge2k≥2) around xˉ\bar xxˉ, the map

R(x,u)=PM(x+u),(x,u)∈TM,R(x,u)=P_{\mathcal M}(x+u),\qquad (x,u)\in T\mathcal M,R(x,u)=PM​(x+u),(x,u)∈TM,

is single-valued near (xˉ,0)(\bar x,0)(xˉ,0) in TMT\mathcal MTM, and this single value is a retraction around xˉ\bar xxˉ.

Milestones

  1. (3.2) If p∈PM(x)p\in P_{\mathcal M}(x)p∈PM​(x) and M\mathcal MM is a CkC^kCk submanifold around ppp, then p∈Mp\in\mathcal Mp∈M and x−p∈NM(p)x-p\in N_{\mathcal M}(p)x−p∈NM​(p).
  2. (3.3) TNM(xˉ,0)=TM(xˉ)×NM(xˉ)T_{N\mathcal M}(\bar x,0)=T_{\mathcal M}(\bar x)\times N_{\mathcal M}(\bar x)TNM​(xˉ,0)=TM​(xˉ)×NM​(xˉ).
  3. Lemma 3.1 There is δ>0\delta>0δ>0 such that on B(xˉ,δ)B(\bar x,\delta)B(xˉ,δ) the projection PMP_{\mathcal M}PM​ is a single point P(x)P(x)P(x), the map PPP is Ck−1C^{k-1}Ck−1, and
DPM(xˉ)=PTM(xˉ).\mathrm DP_{\mathcal M}(\bar x)=P_{T_{\mathcal M}(\bar x)}.DPM​(xˉ)=PTM​(xˉ)​.

Significance

The projective retraction is the reference retraction on an embedded submanifold: it exists for every CkC^kCk submanifold, it has the maximal smoothness Ck−1C^{k-1}Ck−1 allowed by the tangent bundle, and it is the one practitioners compute first (truncated SVD for fixed-rank matrices, polar factor for the Stiefel manifold). Section 4 of the paper shows that it is moreover second order and generalizes it to projections along arbitrary smooth fields of transverse subspaces; Lemma 3.1 is the model of that argument. Lemma 3.1 on its own is a basic tool well beyond retractions: local single-valuedness and smoothness of the nearest-point map underlies the local convergence analysis of alternating projections on manifolds and the theory of prox-regular sets.

All statements here are proved results. No machine-checked proof of them is known to exist: Mathlib has the inverse function theorem, tangent cones and orthogonal projections onto subspaces, but no embedded submanifolds of a Euclidean space with their normal bundle and no nearest-point map onto non-convex sets. The mission produces these statements in Lean and invites proofs of them.

Difficulty

The obvious argument writes the nearest point as a critical point of y↦∥x−y∥2y\mapsto\|x-y\|^2y↦∥x−y∥2 on M\mathcal MM and applies the implicit function theorem. Two things break. First, existence: M\mathcal MM is not assumed closed, so a nearest point exists only because M\mathcal MM is locally closed near xˉ\bar xxˉ and points of M\mathcal MM far from xˉ\bar xxˉ are farther from xxx than xˉ\bar xxˉ is. Second, uniqueness: critical points are not unique in general, and the implicit function theorem only describes critical points near a given one. Uniqueness needs a quantitative argument that every nearest point of xxx lies in the region where (p,v)↦p+v(p,v)\mapsto p+v(p,v)↦p+v on the normal bundle is injective. Finally, the normal bundle is itself only a Ck−1C^{k-1}Ck−1 manifold, whose tangent space at (xˉ,0)(\bar x,0)(xˉ,0) must be identified before the inverse function theorem applies; this is milestone (3.3).

Formalization scope

The ambient space is any type E with [NormedAddCommGroup E] [InnerProductSpace ℝ E] [FiniteDimensional ℝ E]; nnn is Module.finrank ℝ E, and k,dk,dk,d are natural numbers with k≥2k\ge2k≥2 stated as a hypothesis (so that k−1k-1k−1 in ℕ is honest). The paper's standing assumption "M\mathcal MM is a submanifold of class CkC^kCk (k≥2k\ge2k≥2) and dimension ddd" enters only as the local hypothesis IsSubmanifoldAt k d M xbar, exactly as Lemma 3.1 and Proposition 3.2 state it ("around xˉ\bar xxˉ"). The chart is an OpenPartialHomeomorph onto EuclideanSpace ℝ (Fin n), CkC^kCk in both directions, with 0-based coordinates. No closedness of M\mathcal MM is assumed, because the paper does not assume it.

The projection is the platform predicate IsMetricProjection M x z (z∈Mz\in\mathcal Mz∈M and ∥x−z∥≤∥x−w∥\|x-z\|\le\|x-w\|∥x−z∥≤∥x−w∥ for all w∈Mw\in\mathcal Mw∈M); single-valuedness is stated as equality of the set of such zzz with a singleton, which asserts existence and uniqueness. A formalization that only states "some selection is Ck−1C^{k-1}Ck−1", or uses ⊆\subseteq⊆ (satisfied by the empty set), would drop the main claim and is ruled out. The tangent space is the span of Mathlib's tangentConeAt; "class Ck−1C^{k-1}Ck−1 on a neighbourhood in TMT\mathcal MTM" is ContDiffOn on O∩TMO\cap T\mathcal MO∩TM with OOO open; DR(x,⋅)(0)=id\mathrm DR(x,\cdot)(0)=\mathrm{id}DR(x,⋅)(0)=id is a HasFDerivAt statement on the normed space TM(x)T_{\mathcal M}(x)TM​(x); PTM(xˉ)P_{T_{\mathcal M}(\bar x)}PTM​(xˉ)​ is Submodule.starProjection.

A complete development needs: the tangent space of a slice submanifold equals the image of the chart's derivative, the normal bundle as a Ck−1C^{k-1}Ck−1 manifold, an inverse function theorem on it, and compactness of M∩Bˉ(xˉ,r)\mathcal M\cap\bar B(\bar x,r)M∩Bˉ(xˉ,r) for small rrr. These are reusable for the other missions of this series (fixed-rank, spectral and Stiefel manifolds, and the retractor construction of Section 4), and contributions of such general lemmas as intermediate theorems are welcome.

Selected references

  • P.-A. Absil and J. Malick, Projection-like retractions on matrix manifolds, SIAM J. Optim. 22(1):135–158, 2012. https://doi.org/10.1137/100802529 (preprint: https://hal.science/hal-00651608, version 2, the basis of the statement indices here)
  • R. L. Adler, J.-P. Dedieu, J. Y. Margulies, M. Martens and M. Shub, Newton's method on Riemannian manifolds and a geometric model for the human spine, IMA J. Numer. Anal. 22:359–390, 2002. https://doi.org/10.1093/imanum/22.3.359
  • P.-A. Absil, R. Mahony and R. Sepulchre, Optimization Algorithms on Matrix Manifolds, Princeton University Press, 2008. https://doi.org/10.1515/9781400830244
  • A. S. Lewis and J. Malick, Alternating projections on manifolds, Math. Oper. Res. 33(1):216–234, 2008. https://doi.org/10.1287/moor.1070.0291
8 thms1 active userReviewed
Operations ResearchProbabilityStochastic Systems·Captain: mikedeng1

Dimensioning Large Call Centers IV: Asymptotically Optimal Staffing under a Waiting-Cost ConstraintResearch Paper

Motivation

A call center has to decide how many agents to staff. In practice the decision is often posed as a service-level constraint rather than a cost trade-off: use the fewest agents for which the expected waiting cost, or the fraction of customers who wait, stays below a target. Borst, Mandelbaum and Reiman (CWI Report PNA-R0015, 2000; journal version in Operations Research 52(1), 2004, doi:10.1287/opre.1030.0081) treat this constraint problem in Section 8 of their paper, alongside the cost-minimization problem of Sections 5–7, and show that a simple square-root staffing rule solves it asymptotically as the arrival rate grows.

The rule matters because it is what practitioners use. Under the classical Erlang-C model, the exact optimum requires evaluating the Erlang-C formula over many staffing levels. The asymptotic rule replaces this with a single equation in the Halfin–Whitt function PPP: when the target is a delay probability ε\varepsilonε (Example 8.5 of the paper), it reduces to staffing λ/μ+P−1(ε)λ/μ\lambda/\mu + P^{-1}(\varepsilon)\sqrt{\lambda/\mu}λ/μ+P−1(ε)λ/μ​ servers.

Timeline. Erlang's formula for the M/M/N delay probability dates from 1917. Halfin and Whitt (Operations Research 29, 1981) identified the limit P(x)P(x)P(x) of the delay probability under square-root staffing N=λ/μ+xλ/μN = \lambda/\mu + x\sqrt{\lambda/\mu}N=λ/μ+xλ/μ​ with integer NNN. Jagers and Van Doorn (Operations Research Letters 5, 1986; SIAM Review 33, 1991) studied the continued Erlang loss and delay functions at non-integer numbers of servers, including their convexity, which is what lets the staffing problem be relaxed to a continuous one. Borst, Mandelbaum and Reiman (2000/2004) used these to prove asymptotic optimality of square-root rules for both the cost and the constraint formulations.

Setting

Customers arrive at rate λ\lambdaλ to NNN identical servers, each with service rate μ>0\mu > 0μ>0; μ\muμ is fixed while λ→∞\lambda \to \inftyλ→∞. Stability requires N>λ/μN > \lambda/\muN>λ/μ. A customer who waits ttt time units costs Dλ(t)D_\lambda(t)Dλ​(t), where Dλ(0)=0D_\lambda(0) = 0Dλ​(0)=0, DλD_\lambdaDλ​ is strictly increasing on [0,∞)[0,\infty)[0,∞) and ∫0∞Dλ(t)e−θt dt<∞\int_0^\infty D_\lambda(t)e^{-\theta t}\,dt < \infty∫0∞​Dλ​(t)e−θtdt<∞ for all θ>0\theta > 0θ>0.

The Erlang-C probability of waiting is

π(N,ν)=νNN!{(1−ν/N)∑n=0N−1νnn!+νNN!}−1,\pi(N,\nu) = \frac{\nu^N}{N!}\Big\{(1-\nu/N)\sum_{n=0}^{N-1}\frac{\nu^n}{n!} + \frac{\nu^N}{N!}\Big\}^{-1},π(N,ν)=N!νN​{(1−ν/N)n=0∑N−1​n!νn​+N!νN​}−1,

and the conditional waiting cost is G(N,λ)=(Nμ−λ)∫0∞Dλ(t)e−(Nμ−λ)t dtG(N,\lambda) = (N\mu-\lambda)\int_0^\infty D_\lambda(t)e^{-(N\mu-\lambda)t}\,dtG(N,λ)=(Nμ−λ)∫0∞​Dλ​(t)e−(Nμ−λ)tdt. The waiting cost per unit time with NNN servers is

K(N,λ)=λ π(N,λ/μ) G(N,λ).K(N,\lambda) = \lambda\,\pi(N,\lambda/\mu)\,G(N,\lambda).K(N,λ)=λπ(N,λ/μ)G(N,λ).

Given a target Mλ>0M_\lambda > 0Mλ​>0, the optimal staffing level is the least integer N>λ/μN > \lambda/\muN>λ/μ with K(N,λ)≤MλK(N,\lambda) \le M_\lambdaK(N,λ)≤Mλ​; call it Nλ∗N^*_\lambdaNλ∗​.

In the continuous parametrization Nλ(x)=λ/μ+xλ/μN_\lambda(x) = \lambda/\mu + x\sqrt{\lambda/\mu}Nλ​(x)=λ/μ+xλ/μ​, define Gλ(x)=λG(Nλ(x),λ)G_\lambda(x) = \lambda G(N_\lambda(x),\lambda)Gλ​(x)=λG(Nλ​(x),λ), the continuous Erlang-C function πλ(x)=H(Nλ(x),λ/μ)\pi_\lambda(x) = H(N_\lambda(x),\lambda/\mu)πλ​(x)=H(Nλ​(x),λ/μ) with H(M,α)={α∫0∞e−αtt(1+t)M−1dt}−1H(M,\alpha) = \{\alpha\int_0^\infty e^{-\alpha t}t(1+t)^{M-1}dt\}^{-1}H(M,α)={α∫0∞​e−αtt(1+t)M−1dt}−1, and Kλ(x)=πλ(x)Gλ(x)K_\lambda(x) = \pi_\lambda(x)G_\lambda(x)Kλ​(x)=πλ​(x)Gλ​(x). The Halfin–Whitt function is P(x)=1/(1+x/h(−x))P(x) = 1/(1 + x/h(-x))P(x)=1/(1+x/h(−x)) with h=ϕ/(1−Φ)h = \phi/(1-\Phi)h=ϕ/(1−Φ) the standard normal hazard rate. A staffing function xλ>0x_\lambda > 0xλ​>0 is judged by the rounding gap

Tλ(x)=min⁡{∣K(⌊Nλ(x)⌋,λ)−Mλ∣, ∣K(⌈Nλ(x)⌉,λ)−Mλ∣, ∣K(⌈Nλ(x)⌉,λ)−K(Nλ∗,λ)∣}.T_\lambda(x) = \min\big\{|K(\lfloor N_\lambda(x)\rfloor,\lambda) - M_\lambda|,\ |K(\lceil N_\lambda(x)\rceil,\lambda) - M_\lambda|,\ |K(\lceil N_\lambda(x)\rceil,\lambda) - K(N^*_\lambda,\lambda)|\big\}.Tλ​(x)=min{∣K(⌊Nλ​(x)⌋,λ)−Mλ​∣, ∣K(⌈Nλ​(x)⌉,λ)−Mλ​∣, ∣K(⌈Nλ​(x)⌉,λ)−K(Nλ∗​,λ)∣}.

It is asymptotically optimal when Tλ(xλ)/Mλ→0T_\lambda(x_\lambda)/M_\lambda \to 0Tλ​(xλ​)/Mλ​→0 as λ→∞\lambda\to\inftyλ→∞.

Formalization targets

Goal: Theorem 8.2 (rationalized regime)

Suppose that for some κ>0\kappa > 0κ>0 and γ∈(0,∞)\gamma \in (0,\infty)γ∈(0,∞), Gλ(κ)/Mλ→γG_\lambda(\kappa)/M_\lambda \to \gammaGλ​(κ)/Mλ​→γ, i.e. the waiting cost is comparable to the target. Let yλ∗>0y^*_\lambda > 0yλ∗​>0 solve P(y)Gλ(y)=MλP(y)G_\lambda(y) = M_\lambdaP(y)Gλ​(y)=Mλ​. Then

lim⁡λ→∞Tλ(yλ∗)Mλ=0.\lim_{\lambda\to\infty}\frac{T_\lambda(y^*_\lambda)}{M_\lambda} = 0.λ→∞lim​Mλ​Tλ​(yλ∗​)​=0.

Supporting milestones

  • Lemma C.1: GλG_\lambdaGλ​ is strictly convex and decreasing on (0,∞)(0,\infty)(0,∞).
  • Section 3: πλ(x)=π(Nλ(x),λ/μ)\pi_\lambda(x) = \pi(N_\lambda(x),\lambda/\mu)πλ​(x)=π(Nλ​(x),λ/μ) when Nλ(x)N_\lambda(x)Nλ​(x) is an integer.
  • Lemma 8.1: if zλ∗>0z^*_\lambda > 0zλ∗​>0 solves π^λ(z)G^λ(z)=Mλ\hat\pi_\lambda(z)\hat G_\lambda(z) = M_\lambdaπ^λ​(z)G^λ​(z)=Mλ​ and Kλ(zλ∗)/(π^λG^λ)(zλ∗)→1K_\lambda(z^*_\lambda)/(\hat\pi_\lambda\hat G_\lambda)(z^*_\lambda) \to 1Kλ​(zλ∗​)/(π^λ​G^λ​)(zλ∗​)→1, then Tλ(zλ∗)/Mλ→0T_\lambda(z^*_\lambda)/M_\lambda \to 0Tλ​(zλ∗​)/Mλ​→0.
  • Lemma B.1: PPP is strictly convex and decreasing on (0,∞)(0,\infty)(0,∞).
  • Eq. (17): lim sup⁡aλ/b=∞\limsup a_\lambda/b = \inftylimsupaλ​/b=∞ implies lim inf⁡P(aλ)/P(b)=0\liminf P(a_\lambda)/P(b) = 0liminfP(aλ​)/P(b)=0 and lim inf⁡πλ(aλ)/πλ(b)=0\liminf \pi_\lambda(a_\lambda)/\pi_\lambda(b) = 0liminfπλ​(aλ​)/πλ​(b)=0.
  • Lemma 4.1 (Halfin–Whitt): for bounded xλ>0x_\lambda > 0xλ​>0, πλ(xλ)/P(xλ)→1\pi_\lambda(x_\lambda)/P(x_\lambda) \to 1πλ​(xλ​)/P(xλ​)→1; with xλ→xx_\lambda \to xxλ​→x, πλ(xλ)/P(x)→1\pi_\lambda(x_\lambda)/P(x)\to 1πλ​(xλ​)/P(x)→1.

Further target: Theorem 8.6 (efficiency-driven regime)

If Gλ(κ)/Mλ→0G_\lambda(\kappa)/M_\lambda \to 0Gλ​(κ)/Mλ​→0 for every κ>0\kappa > 0κ>0 and yλ∗>0y^*_\lambda > 0yλ∗​>0 solves Gλ(y)=MλG_\lambda(y) = M_\lambdaGλ​(y)=Mλ​, then Tλ(yλ∗)/Mλ→0T_\lambda(y^*_\lambda)/M_\lambda \to 0Tλ​(yλ∗​)/Mλ​→0.

Significance

The theorem certifies the staffing rule used in workforce-management practice: the excess staffing is determined by one scalar equation involving the Gaussian function PPP and the scaled waiting cost, and rounding the resulting staffing level misses the constraint by a vanishing fraction of the target. Lemma 8.1 is a reusable framework: any approximation π^λG^λ\hat\pi_\lambda\hat G_\lambdaπ^λ​G^λ​ that is asymptotically exact at the proposed staffing level yields an asymptotically optimal rule, and the paper instantiates it in three regimes (Theorems 8.2, 8.6, 8.9).

The results are proved on paper. To the best of current knowledge none of them, nor the Halfin–Whitt limit for the continuous Erlang-C extension, has a machine-checked proof. A formalization would produce the first verified heavy-traffic limit of the Erlang-C delay probability, a verified continuous Erlang-C extension with its integer identity, and the convexity facts about PPP and GλG_\lambdaGλ​ that many staffing papers cite without proof.

Difficulty

The obvious argument is to quote Halfin and Whitt: the delay probability converges to P(x)P(x)P(x) under square-root staffing, so PPP can replace the Erlang-C formula. That limit, as published in 1981, is about integer server counts along sequences with a convergent excess-staffing parameter. The paper needs it for the continuous function HHH at non-integer server counts and for staffing functions that are merely bounded, and it also needs the identity H(N,ν)=π(N,ν)H(N,\nu) = \pi(N,\nu)H(N,ν)=π(N,ν) at integers and the monotonicity of πλ\pi_\lambdaπλ​ in xxx, both cited from Jagers and Van Doorn rather than proved. None of these is in Mathlib. A second obstacle is that the staffing function yλ∗y^*_\lambdayλ∗​ is defined only implicitly by an equation involving GλG_\lambdaGλ​, which depends on the arbitrary cost functions DλD_\lambdaDλ​; nothing a priori prevents it from escaping to infinity, outside the range where the Halfin–Whitt approximation applies. Finally, TλT_\lambdaTλ​ compares integer-level costs given by the Erlang-C formula with a continuous approximation, so both representations of the delay probability are in play at once.

Formalization scope

The queue itself is not formalized: there is no Markov chain and no waiting-time distribution. Every statement is about the closed-form waiting cost K(N,λ)K(N,\lambda)K(N,λ) with π\piπ given by the Erlang-C formula, exactly as the paper's analysis is. Conventions, all in the namespace DimCallCenters.Constraint:

  • lam : ℝ is the arrival rate (λ is a Lean keyword); limits are Filter.atTop in lam, with μ fixed. Objects indexed by λ (MλM_\lambdaMλ​, Nλ∗N^*_\lambdaNλ∗​, yλ∗y^*_\lambdayλ∗​) are functions of lam constrained only for lam > 0.
  • WaitModel packages μ > 0 and DλD_\lambdaDλ​ with Dλ(0)=0D_\lambda(0) = 0Dλ​(0)=0, strict monotonicity on [0,∞)[0,\infty)[0,∞), and integrability of Dλ(t)e−θtD_\lambda(t)e^{-\theta t}Dλ​(t)e−θt on (0,∞)(0,\infty)(0,∞) for θ > 0 (the paper's finiteness of GGG; integrability is required because Lean's integral of a non-integrable function is 0).
  • Nλ∗N^*_\lambdaNλ∗​ is a function Nstar : ℝ → ℕ given with its two defining properties (feasible; below every feasible integer level above λ/μ). yλ∗y^*_\lambdayλ∗​ and zλ∗z^*_\lambdazλ∗​ are any positive solutions of their equations; existence and uniqueness are not hypotheses.
  • In TλT_\lambdaTλ​ the round-down term is dropped when ⌊Nλ(x)⌋≤λ/μ\lfloor N_\lambda(x)\rfloor \le \lambda/\mu⌊Nλ​(x)⌋≤λ/μ (an unstable level where KKK is undefined). This can only enlarge TλT_\lambdaTλ​.
  • Asymptotic relations are limits of ratios. lim sup⁡=∞\limsup = \inftylimsup=∞ and lim inf⁡=0\liminf = 0liminf=0 are stated with ∃ᶠ ("frequently"), lim sup⁡<∞\limsup < \inftylimsup<∞ as eventual boundedness.
  • PPP is defined through explicit ϕ\phiϕ, Φ\PhiΦ, hhh; the formula also gives P(0)=1P(0) = 1P(0)=1, used in Lemma 4.1(2) at x=0x = 0x=0.
  • No hypothesis lim⁡N↓λ/μG(N,λ)=∞\lim_{N\downarrow\lambda/\mu}G(N,\lambda) = \inftylimN↓λ/μ​G(N,λ)=∞ is added: it is not needed for the statements here.

A trivializing formalization is ruled out: TλT_\lambdaTλ​ keeps all of the paper's terms and is never replaced by a smaller quantity, and the hypotheses are jointly satisfiable — Dλ(t)=aλ/μ tD_\lambda(t) = a\sqrt{\lambda/\mu}\,tDλ​(t)=aλ/μ​t with Mλ=MλM_\lambda = M\lambdaMλ​=Mλ satisfies (33) for every κ\kappaκ with γ=a/(μκM)\gamma = a/(\mu\kappa M)γ=a/(μκM).

Infrastructure needed: the continuous Erlang-C function and its integer identity; the Halfin–Whitt limit (a Gaussian approximation of Poisson/gamma tails); calculus facts about the normal hazard rate. These are reusable beyond this mission, notably by the sibling missions on the cost-minimization problem. Example 8.5 (delay-probability target with Dλ=1t>0D_\lambda = 1_{t>0}Dλ​=1t>0​) motivates the rule but violates the strict monotonicity of DλD_\lambdaDλ​, so it is not an instance of the theorem as stated. Contributions on any milestone, and on Theorem 8.9 (quality-driven regime, which needs Lemma 4.2), are welcome.

Selected references

  • S. Borst, A. Mandelbaum, M. I. Reiman, Dimensioning Large Call Centers, CWI Report PNA-R0015, 2000; Operations Research 52(1):17–34, 2004. https://doi.org/10.1287/opre.1030.0081
  • S. Halfin, W. Whitt, Heavy-Traffic Limits for Queues with Many Exponential Servers, Operations Research 29(3):567–588, 1981. https://doi.org/10.1287/opre.29.3.567
  • A. A. Jagers, E. A. Van Doorn, On the Continued Erlang Loss Function, Operations Research Letters 5:43–46, 1986.
  • A. A. Jagers, E. A. Van Doorn, Convexity of Functions which are Generalizations of the Erlang Loss Function and the Erlang Delay Function, SIAM Review 33:281–282, 1991.
18 thms1 active userReviewed
AnalysisDifferential Geometry·Captain: mikedeng1

Projection-like Retractions on Matrix Manifolds I: Coming Back to a Submanifold Along a Smooth Field of Transverse Subspaces Defines a RetractionResearch Paper

Motivation

Iterative methods for optimization and equation solving on a smooth constraint set M\mathcal MM (orthogonal matrices, fixed-rank matrices, spheres, Stiefel and Grassmann manifolds) compute an update vector uuu in the tangent space at the current iterate xxx and then have to return to M\mathcal MM. The Riemannian exponential map does this along geodesics, but computing it means solving an ordinary differential equation. The notion of retraction, introduced by Adler, Dedieu, Margulies, Martens and Shub (IMA J. Numer. Anal., 2002) and developed in the book of Absil, Mahony and Sepulchre (Princeton, 2008), captures what such a return map needs for Newton's method to keep its local quadratic convergence and for gradient methods to converge: smoothness, R(x,0)=xR(x,0)=xR(x,0)=x, and first-order agreement with the exponential.

Absil and Malick (SIAM J. Optim., 2012; HAL hal-00651608v2) give a general recipe for building retractions on submanifolds of a Euclidean space: move tangentially from xxx to x+ux+ux+u, then come back to M\mathcal MM along a prescribed family of admissible directions. This mission formalizes that recipe, Theorem 4.2 of the paper ("retractors give retractions"), together with the two lemmas its proof rests on.

Setting

Let E\mathcal EE be a Euclidean space of dimension nnn (in the paper's examples, Rn×m\mathbb R^{n\times m}Rn×m with the Frobenius inner product). A set M⊆E\mathcal M\subseteq\mathcal EM⊆E is a CkC^kCk submanifold of dimension ddd if around every xˉ∈M\bar x\in\mathcal Mxˉ∈M it is a coordinate slice: there are an open neighbourhood U\mathcal UU of xˉ\bar xxˉ and a CkC^kCk diffeomorphism ϕ\phiϕ of U\mathcal UU onto an open subset of Rn\mathbb R^nRn with M∩U={x∈U:ϕd+1(x)=⋯=ϕn(x)=0}\mathcal M\cap\mathcal U=\{x\in\mathcal U:\phi_{d+1}(x)=\dots=\phi_n(x)=0\}M∩U={x∈U:ϕd+1​(x)=⋯=ϕn​(x)=0}. Throughout, k≥2k\ge2k≥2.

The tangent space TM(x)\mathrm T_{\mathcal M}(x)TM​(x) is the linear subspace of E\mathcal EE of tangent directions of M\mathcal MM at xxx, the normal space NM(x)\mathrm N_{\mathcal M}(x)NM​(x) is its orthogonal complement, and the tangent bundle is TM={(x,u):x∈M, u∈TM(x)}\mathrm T\mathcal M=\{(x,u):x\in\mathcal M,\ u\in\mathrm T_{\mathcal M}(x)\}TM={(x,u):x∈M, u∈TM​(x)}.

A map RRR from TM\mathrm T\mathcal MTM to M\mathcal MM is a retraction around xˉ\bar xxˉ (Definition 2.1) if, on a neighbourhood U\mathcal UU of (xˉ,0)(\bar x,0)(xˉ,0) in TM\mathrm T\mathcal MTM, it is of class Ck−1C^{k-1}Ck−1, satisfies R(x,0)=xR(x,0)=xR(x,0)=x, and u↦R(x,u)u\mapsto R(x,u)u↦R(x,u) has derivative idTM(x)\mathrm{id}_{\mathrm T_{\mathcal M}(x)}idTM​(x)​ at u=0u=0u=0. It is a retraction on M\mathcal MM if this holds around every point.

A retractor (Definition 4.1) is a Ck−1C^{k-1}Ck−1 map DDD, defined on a neighbourhood of the zero section of TM\mathrm T\mathcal MTM, with values in the Grassmann manifold Gr(n−d,E)\mathrm{Gr}(n-d,\mathcal E)Gr(n−d,E) of (n−d)(n-d)(n−d)-dimensional linear subspaces, such that D(x,0)∩TM(x)={0}D(x,0)\cap\mathrm T_{\mathcal M}(x)=\{0\}D(x,0)∩TM​(x)={0} for every x∈Mx\in\mathcal Mx∈M. Given DDD, set D(x,u)=x+u+D(x,u)\mathcal D(x,u)=x+u+D(x,u)D(x,u)=x+u+D(x,u) and let

R(x,u)={points of M∩D(x,u) nearest to x+u}.R(x,u)=\{\text{points of }\mathcal M\cap\mathcal D(x,u)\text{ nearest to }x+u\}.R(x,u)={points of M∩D(x,u) nearest to x+u}.

Formalization targets

Goal: Theorem 4.2 (retractors give retractions)

∀xˉ∈M  ∃r:R(x,u)={r(x,u)}  for (x,u)∈TM near (xˉ,0),and r is a retraction around xˉ.\forall\bar x\in\mathcal M\ \ \exists r:\quad R(x,u)=\{r(x,u)\}\ \text{ for }(x,u)\in\mathrm T\mathcal M\text{ near }(\bar x,0),\quad\text{and } r \text{ is a retraction around } \bar x.∀xˉ∈M  ∃r:R(x,u)={r(x,u)}  for (x,u)∈TM near (xˉ,0),and r is a retraction around xˉ.

The theorem asserts both that the point-to-set map RRR is single-valued near the zero section and that it is a retraction there.

Milestone: Lemma 4.7 (the normal case D(x,u)=NM(x)D(x,u)=\mathrm N_{\mathcal M}(x)D(x,u)=NM​(x))

Near (xˉ,0)(\bar x,0)(xˉ,0) there is one and only one smallest v(x,u)∈NM(x)v(x,u)\in\mathrm N_{\mathcal M}(x)v(x,u)∈NM​(x) with x+u+v(x,u)∈Mx+u+v(x,u)\in\mathcal Mx+u+v(x,u)∈M; Duv(x,0)=0\mathrm D_u v(x,0)=0Du​v(x,0)=0; and R(x,u)=x+u+v(x,u)R(x,u)=x+u+v(x,u)R(x,u)=x+u+v(x,u) is a retraction around xˉ\bar xxˉ, hence on M\mathcal MM.

Milestone: Lemma 4.8 (straightening up)

On a neighbourhood of the zero section, D(x,u)={v+A(x,u)v: v∈NM(x)}D(x,u)=\{v+A(x,u)v:\ v\in\mathrm N_{\mathcal M}(x)\}D(x,u)={v+A(x,u)v: v∈NM​(x)} for a unique linear A(x,u):NM(x)→TM(x)A(x,u):\mathrm N_{\mathcal M}(x)\to\mathrm T_{\mathcal M}(x)A(x,u):NM​(x)→TM​(x) depending Ck−1C^{k-1}Ck−1 on (x,u)(x,u)(x,u).

Further: Theorem 4.9 (second order)

If k≥3k\ge3k≥3 and D(x,0)=NM(x)D(x,0)=\mathrm N_{\mathcal M}(x)D(x,0)=NM​(x) for all x∈Mx\in\mathcal Mx∈M, then d2dt2R(x,tu)∣t=0∈NM(x)\frac{\mathrm d^2}{\mathrm dt^2}R(x,tu)|_{t=0}\in\mathrm N_{\mathcal M}(x)dt2d2​R(x,tu)∣t=0​∈NM​(x) for all (x,u)∈TM(x,u)\in\mathrm T\mathcal M(x,u)∈TM.

Significance

Theorem 4.2 reduces the construction of a retraction to the choice of a smooth field of subspaces transverse to the tangent space at u=0u=0u=0. The orthographic retraction (D=NM(x)D=\mathrm N_{\mathcal M}(x)D=NM​(x)) and the projective retraction R(x,u)=PM(x+u)R(x,u)=P_{\mathcal M}(x+u)R(x,u)=PM​(x+u) (D=NM(PM(x+u))D=\mathrm N_{\mathcal M}(P_{\mathcal M}(x+u))D=NM​(PM​(x+u))) are both instances, as are the gnomonic, orthographic and stereographic projections on the sphere. Theorem 4.9 then certifies, by a check at u=0u=0u=0 only, that a retraction agrees with the exponential to second order, which matters for the superlinear convergence of Riemannian trust-region and Newton methods. Lemma 4.7 goes beyond an earlier result on tangential parameterizations (reference [28, Th. 3.4] of the paper) by giving Ck−1C^{k-1}Ck−1 regularity jointly in (x,u)(x,u)(x,u), not only in uuu (Remark 4.4 of the paper).

The results are proved in the paper. None of them is machine-checked: Mathlib has the implicit function theorem and smooth manifolds, but no embedded submanifolds of a Euclidean space with their tangent and normal bundles, no retractions, and no smooth Grassmannian-valued maps. The work here is to formalize the known proof and to build this layer, which every later mission of this series (projective, spectral, fixed-rank and Stiefel retractions) also needs.

Difficulty

The statement is an implicit-function argument, but the obvious one does not apply directly. The unknown vvv lives in the normal space NM(x)\mathrm N_{\mathcal M}(x)NM​(x), which moves with xxx, and (x,u)(x,u)(x,u) ranges over the tangent bundle, a Ck−1C^{k-1}Ck−1 submanifold of E×E\mathcal E\times\mathcal EE×E rather than an open set of a vector space. The equation has to be read in charts of the bundle, which costs one derivative; the uniqueness given by the implicit function theorem holds only locally in (x,u,v)(x,u,v)(x,u,v), and turning it into "the smallest vvv", or "the nearest point of M∩D(x,u)\mathcal M\cap\mathcal D(x,u)M∩D(x,u)", requires excluding far-away intersection points. For a general retractor, the subspace D(x,u)D(x,u)D(x,u) also moves, and has to be written as a graph over the normal space with a Ck−1C^{k-1}Ck−1 dependence before a second implicit-function argument applies.

Formalization scope

E\mathcal EE is a finite-dimensional real inner product space E, nnn is Module.finrank ℝ E, and k,dk,dk,d are natural numbers with 2≤k2\le k2≤k and d≤nd\le nd≤n (the latter inside the submanifold definition). The coordinate slice uses an open partial homeomorphism E → EuclideanSpace ℝ (Fin n), CkC^kCk in both directions, with 0-based coordinates. TM(x)\mathrm T_{\mathcal M}(x)TM​(x) is the span of Mathlib's tangent cone (chart-free), NM(x)\mathrm N_{\mathcal M}(x)NM​(x) its orthogonal complement. Retractions and retractors are total maps on E×E\mathcal E\times\mathcal EE×E constrained only on O∩TMO\cap\mathrm T\mathcal MO∩TM with OOO open; "of class Ck−1C^{k-1}Ck−1 on a subset of TM\mathrm T\mathcal MTM" is ContDiffOn on that set. A Grassmannian-valued map is Ck−1C^{k-1}Ck−1 when its orthogonal projector PD(x,u)P_{D(x,u)}PD(x,u)​ is, and the dimension n−dn-dn−d of D(x,u)D(x,u)D(x,u) is part of the definition. The set-valued RRR uses the published nearest-point predicate RandomGradFree.Nonsmooth.IsMetricProjection. In Lemma 4.8, A(x,u)A(x,u)A(x,u) is extended by 000 on TM(x)\mathrm T_{\mathcal M}(x)TM​(x) so that it is an operator on E\mathcal EE. Theorem 4.9 assumes k≥3k\ge3k≥3, the standing assumption of the definition of second-order retractions (display (2.3)), and reads the page's "D(xˉ,0)=NM(x)D(\bar x,0)=\mathrm N_{\mathcal M}(x)D(xˉ,0)=NM​(x)" as D(x,0)=NM(x)D(x,0)=\mathrm N_{\mathcal M}(x)D(x,0)=NM​(x) for all x∈Mx\in\mathcal Mx∈M.

The goal is not satisfied by exhibiting some retraction: it requires the nearest-point set R(x,u)R(x,u)R(x,u) to equal a singleton {r(x,u)}\{r(x,u)\}{r(x,u)} (so it is nonempty) and that very rrr to be a retraction. Theorem 4.9 requires the curve t↦R(x,tu)t\mapsto R(x,tu)t↦R(x,tu) to be twice differentiable, so a junk second derivative of 000 does not satisfy it.

A complete development needs: finite-rank facts for tangent spaces of slices (dim⁡TM(x)=d\dim\mathrm T_{\mathcal M}(x)=ddimTM​(x)=d), smoothness of x↦PNM(x)x\mapsto P_{\mathrm N_{\mathcal M}(x)}x↦PNM​(x)​, charts of TM\mathrm T\mathcal MTM and of the Whitney sum TM⊕NM\mathrm T\mathcal M\oplus\mathrm N\mathcal MTM⊕NM, ContDiffOn on submanifolds, and local orthonormal frames of a smooth field of subspaces. All of these are reusable beyond this mission. Contributions to any of them, and proofs of the two lemmas, are welcome.

Selected references

  • P.-A. Absil, J. Malick, Projection-like retractions on matrix manifolds, SIAM J. Optim. 22(1):135–158, 2012. https://doi.org/10.1137/100802529 (authors' version: https://hal.science/hal-00651608)
  • P.-A. Absil, R. Mahony, R. Sepulchre, Optimization Algorithms on Matrix Manifolds, Princeton University Press, 2008. https://press.princeton.edu/absil
  • R. L. Adler, J.-P. Dedieu, J. Y. Margulies, M. Martens, M. Shub, Newton's method on Riemannian manifolds and a geometric model for the human spine, IMA J. Numer. Anal. 22(3):359–390, 2002. https://doi.org/10.1093/imanum/22.3.359
10 thms1 active userReviewed
AnalysisDynamic ProgrammingOperations Research+2·Captain: mikedeng1

Stochastic Optimal Control: The Discrete-Time Case VI: Lower Semianalytic Functions — Analytically Measurable ε-Optimal Selectors (Jankov–von Neumann)Textbook

Motivation

Dynamic programming over uncountable state and control spaces needs two things at every stage: the optimal cost-to-go, obtained by minimizing over the control, must be a function that can be integrated against the next stage's transition probabilities, and a policy that nearly attains the minimum must be measurable, so that it defines a stochastic process. With Borel-measurable costs and Borel-measurable policies both requirements fail. Minimizing a Borel function of (x,y)(x,y)(x,y) over yyy produces a function whose level sets are projections of Borel sets, and such projections need not be Borel (Suslin, 1917). The repair, developed by Blackwell, Freedman and Orkin (1974), Shreve and Bertsekas, and set out in Chapter 7 of Bertsekas and Shreve's Stochastic Optimal Control: The Discrete-Time Case (1978), is to enlarge the class of costs to the lower semianalytic functions and the class of policies to the analytically or universally measurable ones. Sections 7.6–7.7 of the book establish that this class is closed under partial minimization and admits measurable ε-optimal selectors. Chapters 8–10 of the book, and much of the later literature on Borel-space Markov decision processes (Hernández-Lerma and Lasserre; Feinberg and coauthors), build on these results.

Timeline:

  • 1917: Suslin shows that projections of Borel sets need not be Borel and introduces analytic sets; Lusin proves that analytic sets are universally measurable.
  • 1941–1949: Jankov and von Neumann independently prove that an analytic subset of a product admits a selector measurable with respect to the σ-algebra generated by analytic sets.
  • 1974: Blackwell, Freedman and Orkin use analytic sets to construct ε-optimal policies in Borel dynamic programming.
  • 1978: Bertsekas and Shreve give the treatment used here (§7.6–7.7), including the selection theorem for lower semianalytic functions, Proposition 7.50.

Setting

A Borel space is a topological space homeomorphic to a Borel subset of a complete separable metric space (Definition 7.7); its Borel σ-algebra is BX\mathscr B_XBX​. The Baire space is N=NN\mathscr N=\mathbb N^{\mathbb N}N=NN with the product topology. A set A⊆XA\subseteq XA⊆X is analytic if it is empty or the image of N\mathscr NN under a continuous map; by Proposition 7.41 this is the book's Definition 7.16 (the Suslin operation applied to closed sets). Every Borel set is analytic, and the converse fails when XXX is uncountable.

Three σ-algebras on XXX are in play. The analytic σ-algebra AX\mathscr A_XAX​ is generated by the analytic sets (Definition 7.19). The universal σ-algebra is UX=⋂pBX(p)\mathscr U_X=\bigcap_{p}\mathscr B_X(p)UX​=⋂p​BX​(p), the intersection over all probability measures ppp on (X,BX)(X,\mathscr B_X)(X,BX​) of the ppp-completions of BX\mathscr B_XBX​ (Definition 7.18). For a function fff from D⊆XD\subseteq XD⊆X into a Borel space YYY, fff is analytically measurable if D∈AXD\in\mathscr A_XD∈AX​ and f−1(B)∈AXf^{-1}(B)\in\mathscr A_Xf−1(B)∈AX​ for every B∈BYB\in\mathscr B_YB∈BY​, and universally measurable if the same holds with UX\mathscr U_XUX​ (Definition 7.20).

Let R∗=[−∞,∞]R^*=[-\infty,\infty]R∗=[−∞,∞]. A function f:D→R∗f:D\to R^*f:D→R∗ is lower semianalytic if DDD is analytic and {x∈D∣f(x)<c}\{x\in D\mid f(x)<c\}{x∈D∣f(x)<c} is analytic for every real ccc (Definition 7.21). For D⊆X×YD\subseteq X\times YD⊆X×Y write Dx={y∣(x,y)∈D}D_x=\{y\mid (x,y)\in D\}Dx​={y∣(x,y)∈D}, projX(D)={x∣Dx≠∅}\mathrm{proj}_X(D)=\{x\mid D_x\neq\emptyset\}projX​(D)={x∣Dx​=∅}, and define the partial infimum

f∗(x)=inf⁡y∈Dxf(x,y),x∈projX(D).f^*(x)=\inf_{y\in D_x}f(x,y),\qquad x\in\mathrm{proj}_X(D).f∗(x)=y∈Dx​inf​f(x,y),x∈projX​(D).

A selector is a function φ:projX(D)→Y\varphi:\mathrm{proj}_X(D)\to Yφ:projX​(D)→Y whose graph Gr(φ)\mathrm{Gr}(\varphi)Gr(φ) lies in DDD.

Formalization targets

Goal: Proposition 7.50

Let X,YX,YX,Y be Borel spaces, D⊆X×YD\subseteq X\times YD⊆X×Y analytic, and f:D→R∗f:D\to R^*f:D→R∗ lower semianalytic.

(a) For every ε>0\varepsilon>0ε>0 there is an analytically measurable selector φ\varphiφ with

f[x,φ(x)]≤{f∗(x)+εif f∗(x)>−∞,−1/εif f∗(x)=−∞.f[x,\varphi(x)]\le\begin{cases}f^*(x)+\varepsilon&\text{if }f^*(x)>-\infty,\\-1/\varepsilon&\text{if }f^*(x)=-\infty.\end{cases}f[x,φ(x)]≤{f∗(x)+ε−1/ε​if f∗(x)>−∞,if f∗(x)=−∞.​

(b) The set III of points where the infimum is attained is universally measurable, and for every ε>0\varepsilon>0ε>0 there is a universally measurable selector φ\varphiφ with f[x,φ(x)]=f∗(x)f[x,\varphi(x)]=f^*(x)f[x,φ(x)]=f∗(x) on III and the bounds of (a) off III.

The goal fixes no constant beyond the book's ε\varepsilonε and −1/ε-1/\varepsilon−1/ε.

Milestones

In attack order: Proposition 7.40 (Borel images and preimages of analytic sets are analytic), Corollary 7.42.1 (AX⊆UX\mathscr A_X\subseteq\mathscr U_XAX​⊆UX​), Corollary 7.44.2 (composites of analytically measurable maps are universally measurable), and Proposition 7.49, the Jankov–von Neumann theorem:

A⊆X×Y analytic ⟹ ∃ φ:projX(A)→Y analytically measurable, Gr(φ)⊆A.A\subseteq X\times Y\text{ analytic}\ \Longrightarrow\ \exists\,\varphi:\mathrm{proj}_X(A)\to Y\ \text{analytically measurable},\ \mathrm{Gr}(\varphi)\subseteq A.A⊆X×Y analytic ⟹ ∃φ:projX​(A)→Y analytically measurable, Gr(φ)⊆A.

Further items of the mission, on the same definitions: Proposition 7.39 (projections of analytic sets are analytic, and every analytic set is a projection of a Borel set), Lemma 7.30(1) (strict and non-strict, real and extended level sets give the same class) and Proposition 7.47 (lower semianalytic functions are exactly partial infima of Borel functions).

Significance

Proposition 7.50 is the selection theorem behind the existence of ε-optimal policies in Borel-space dynamic programming. In the finite-horizon model of Chapter 8 the optimal cost-to-go at each stage is lower semianalytic, by Propositions 7.47 and 7.48. Proposition 7.50 then turns the one-stage minimization into a measurable policy, analytically measurable when only ε-optimality is required and universally measurable when the minimum is attained. Chapters 8–9 of the book (the finite-horizon recursion JK∗=TK(J0)J^*_K=T^K(J_0)JK∗​=TK(J0​) and the optimality equation under (P), (N), (D)) use it at every step. Downstream catalog papers on average-cost and stochastic shortest-path problems over Borel spaces cite these results.

All results here are proved in the book and in the descriptive set theory literature (Kechris, Classical Descriptive Set Theory, §18 and §29). None is formalized on Prove2Me. Mathlib has analytic sets in Polish-type settings, the Lusin separation theorem and Suslin's theorem, but it has no universal σ-algebra, no analytic σ-algebra, no lower semianalytic functions and no Jankov–von Neumann uniformization. The definitions in this mission are reusable by the later missions of the series (Chapters 8–10), which restate them locally until these are published.

Difficulty

The obvious route to a selector is to choose, for each xxx, a minimizing or near-minimizing yyy. The axiom of choice provides such a function, but nothing makes it measurable, and the conclusion of the theorem is exactly that measurability. The Borel route fails too: the set {x∣f∗(x)<c}\{x\mid f^*(x)<c\}{x∣f∗(x)<c} is a projection of a Borel set, which is analytic but in general not Borel, so no Borel-measurable selector exists in general. The Jankov–von Neumann theorem needs a lexicographically least branch of a continuous parametrization of AAA by N\mathscr NN, and an argument that the resulting map is measurable with respect to AX\mathscr A_XAX​, which is generated by sets that are not closed under complementation. Part (b) adds a further obstacle: the composite of two analytically measurable maps need not be analytically measurable, so the exact selector is only universally measurable. Proving that requires Lusin's theorem that analytic sets are measurable for every completed probability measure.

Formalization scope

  • A Borel space is a type with a topology satisfying the class IsBorelSpace (Definition 7.7, the ambient complete separable metric space taken in the same universe), together with Mathlib's [MeasurableSpace X] [BorelSpace X], so measurable sets are exactly the Borel sets. On X×YX\times YX×Y the product σ-algebra is used; it coincides with BX×Y\mathscr B_{X\times Y}BX×Y​ for separable metrizable spaces (Proposition 7.13).
  • Analytic sets are Mathlib's MeasureTheory.AnalyticSet (empty or a continuous image of ℕ → ℕ).
  • R∗R^*R∗ is EReal. The book uses ∞−∞=∞\infty-\infty=\infty∞−∞=∞, and Mathlib's EReal uses ⊥+⊤=⊥\bot+\top=\bot⊥+⊤=⊥. No statement of this mission adds infinities of opposite sign; f∗(x)+εf^*(x)+\varepsilonf∗(x)+ε adds a real number.
  • Functions on DDD and on projX(D)\mathrm{proj}_X(D)projX​(D) are functions on subtypes. The graph condition Gr(φ)⊆D\mathrm{Gr}(\varphi)\subseteq DGr(φ)⊆D is part of every selector statement.
  • Universally measurable means NullMeasurableSet E p for every probability measure p.
  • "Analytically measurable" refers to the σ-algebra generated by analytic sets. Replacing it by the power set, dropping the graph condition, or dropping the −1/ε-1/\varepsilon−1/ε case would make the selection theorems a consequence of the axiom of choice. The statements rule all three out.

Not included: Lusin's theorem in Suslin-scheme form (Proposition 7.42, which needs the Suslin operation as a definition), Proposition 7.43 on P(X)P(X)P(X), the integration results of Propositions 7.46 and 7.48, and Lemma 7.30(2)–(4). None is used in the proof of the goal. Contributions welcome: the bridge between IsBorelSpace and Mathlib's StandardBorelSpace, the universal σ-algebra API, and the Jankov–von Neumann theorem itself.

Selected references

  • D. P. Bertsekas and S. E. Shreve, Stochastic Optimal Control: The Discrete-Time Case, Academic Press 1978; Athena Scientific 1996, §7.6–7.7. https://web.mit.edu/dimitrib/www/soc.html
  • D. Blackwell, D. Freedman and M. Orkin, The optimal reward operator in dynamic programming, Annals of Probability 2 (1974) 926–941. https://doi.org/10.1214/aop/1176996558
  • A. S. Kechris, Classical Descriptive Set Theory, Graduate Texts in Mathematics 156, Springer 1995, §18 (Jankov–von Neumann uniformization), §29 (measurability of analytic sets). https://doi.org/10.1007/978-1-4612-4190-4
  • S. E. Shreve and D. P. Bertsekas, Universally measurable policies in dynamic programming, Mathematics of Operations Research 4 (1979) 15–30. https://doi.org/10.1287/moor.4.1.15
8 thms1 active userReviewed
AnalysisDynamic ProgrammingOperations Research+2·Captain: mikedeng1

Stochastic Optimal Control: The Discrete-Time Case V: Semicontinuous Functions — a Borel-Measurable Minimizing Selector for Lower Semicontinuous CostsTextbook

Motivation

Every step of the dynamic programming algorithm on a general state space does three things: it takes a conditional expectation of the cost-to-go under a transition kernel, it minimizes the resulting function of state and control over the control, and, if a policy is to be produced, it picks a control for each state that attains or nearly attains that minimum. On a finite or countable state space all three are harmless. On an uncountable state space each can destroy the measurability needed to take the next expectation: the infimum over an uncountable family of measurable functions need not be measurable, and a minimizer chosen state by state need not be a measurable function of the state, so it does not define a policy at all.

Section 7.5 of Bertsekas and Shreve, Stochastic Optimal Control: The Discrete-Time Case (1978; Athena Scientific reprint 1996), settles the three operations for semicontinuous costs and continuous kernels. The results are the topological half of the book's measurability theory; the descriptive set theory half (lower semianalytic functions and analytically measurable selectors, §7.6–7.7) is a separate mission in this series. The semicontinuous results are what Propositions 8.6–8.7 and Corollaries 9.17.2–9.17.3 of the book use to obtain Borel-measurable optimal policies for finite-horizon and infinite-horizon models with lower semicontinuous costs and compact control sets.

Timeline. The exact selection theorem for lower semicontinuous functions (Proposition 7.33 below) is credited by the book's notes to Dubins and Savage, How to Gamble If You Must (1965). The Hausdorff metric on closed sets goes back to Hausdorff's Set Theory. Measurable selection in the closed-valued setting was later systematized by Kuratowski and Ryll-Nardzewski (1965), whose theorem gives a different route to results of this kind.

Setting

Throughout, R∗=[−∞,+∞]R^*=[-\infty,+\infty]R∗=[−∞,+∞] is the extended real line. A function f:X→R∗f:X\to R^*f:X→R∗ on a metrizable space XXX is lower semicontinuous if every sublevel set {x∣f(x)≤c}\{x\mid f(x)\le c\}{x∣f(x)≤c}, c∈Rc\in\mathbb Rc∈R, is closed, and upper semicontinuous if every superlevel set {x∣f(x)≥c}\{x\mid f(x)\ge c\}{x∣f(x)≥c} is closed (Definition 7.13). C(X)C(X)C(X) is the space of bounded continuous real-valued functions on XXX.

For a separable metrizable space YYY, P(Y)P(Y)P(Y) is the set of Borel probability measures on YYY with the weak topology (convergence of integrals of functions in C(Y)C(Y)C(Y)). A stochastic kernel q(dy∣x)q(dy\mid x)q(dy∣x) on YYY given XXX is a map x↦q(dy∣x)x\mapsto q(dy\mid x)x↦q(dy∣x) from XXX to P(Y)P(Y)P(Y), and it is continuous if this map is continuous (Definition 7.12). The integral of a Borel-measurable f:Y→R∗f:Y\to R^*f:Y→R∗ is ∫f dp=∫f+dp−∫f−dp\int f\,dp=\int f^+dp-\int f^-dp∫fdp=∫f+dp−∫f−dp with the convention −∞+∞=+∞−∞=+∞-\infty+\infty=+\infty-\infty=+\infty−∞+∞=+∞−∞=+∞ (Eq. (43) of Chapter 7).

For a compact metric space YYY, 2Y2^Y2Y is the collection of closed subsets of YYY with the topology of the Hausdorff metric (Appendix C). For D⊆X×YD\subseteq X\times YD⊆X×Y, the section at xxx is Dx={y∣(x,y)∈D}D_x=\{y\mid (x,y)\in D\}Dx​={y∣(x,y)∈D}, the projection is projX(D)={x∣Dx≠∅}\mathrm{proj}_X(D)=\{x\mid D_x\neq\emptyset\}projX​(D)={x∣Dx​=∅}, and a function φ:projX(D)→Y\varphi:\mathrm{proj}_X(D)\to Yφ:projX​(D)→Y has its graph in DDD if (x,φ(x))∈D(x,\varphi(x))\in D(x,φ(x))∈D for every x∈projX(D)x\in\mathrm{proj}_X(D)x∈projX​(D). "Borel-measurable" refers to the Borel σ-algebras of the topologies in question; on projX(D)\mathrm{proj}_X(D)projX​(D) this is the Borel σ-algebra of the subspace topology.

Formalization targets

Goal: Proposition 7.33

Let XXX be metrizable, YYY compact metrizable, D⊆X×YD\subseteq X\times YD⊆X×Y closed, and f:D→R∗f:D\to R^*f:D→R∗ lower semicontinuous. Put

f∗(x)=min⁡y∈Dxf(x,y),x∈projX(D).f^*(x)=\min_{y\in D_x}f(x,y),\qquad x\in\mathrm{proj}_X(D).f∗(x)=y∈Dx​min​f(x,y),x∈projX​(D).

Then projX(D)\mathrm{proj}_X(D)projX​(D) is closed, f∗f^*f∗ is lower semicontinuous, and there is a Borel-measurable φ:projX(D)→Y\varphi:\mathrm{proj}_X(D)\to Yφ:projX​(D)→Y with graph in DDD and

f(x,φ(x))=f∗(x)∀x∈projX(D).f\bigl(x,\varphi(x)\bigr)=f^*(x)\qquad\forall x\in\mathrm{proj}_X(D).f(x,φ(x))=f∗(x)∀x∈projX​(D).

Milestones

  • Proposition 7.32: for f∗(x)=inf⁡y∈Yf(x,y)f^*(x)=\inf_{y\in Y}f(x,y)f∗(x)=infy∈Y​f(x,y), lower semicontinuity of fff and compactness of YYY give lower semicontinuity of f∗f^*f∗ and attainment; upper semicontinuity of fff gives upper semicontinuity of f∗f^*f∗.
  • Lemma 7.18: there is a Borel-measurable σ:2Y−{∅}→Y\sigma:2^Y-\{\emptyset\}\to Yσ:2Y−{∅}→Y with σ(A)∈A\sigma(A)\in Aσ(A)∈A.
  • Lemma 7.20: for lower semicontinuous fff on a nonempty compact YYY, the argmin map x↦{y∣f(x,y)≤f∗(x)}x\mapsto\{y\mid f(x,y)\le f^*(x)\}x↦{y∣f(x,y)≤f∗(x)} is Borel-measurable into 2Y2^Y2Y.
  • Lemma 7.14: fff is lower semicontinuous and bounded below iff fn↑ff_n\uparrow ffn​↑f for some fn∈C(X)f_n\in C(X)fn​∈C(X) (and dually).
  • Proposition 7.30: x↦∫f(x,y) q(dy∣x)x\mapsto\int f(x,y)\,q(dy\mid x)x↦∫f(x,y)q(dy∣x) is continuous for f∈C(X×Y)f\in C(X\times Y)f∈C(X×Y) and continuous qqq.
  • Proposition 7.31: the same map is lower (upper) semicontinuous and bounded below (above) when fff is.
  • Lemma 7.21: an open G⊆X×YG\subseteq X\times YG⊆X×Y, YYY separable, has open projection and a Borel-measurable selector with graph in GGG.
  • Proposition 7.34: for open DDD and upper semicontinuous fff, projX(D)\mathrm{proj}_X(D)projX​(D) is open, f∗=inf⁡Dxff^*=\inf_{D_x}ff∗=infDx​​f is upper semicontinuous, and for each ε>0\varepsilon>0ε>0 there is a Borel-measurable φε\varphi_\varepsilonφε​ with graph in DDD and
f(x,φε(x))≤{f∗(x)+εif f∗(x)>−∞,−1/εif f∗(x)=−∞.f\bigl(x,\varphi_\varepsilon(x)\bigr)\le\begin{cases}f^*(x)+\varepsilon&\text{if }f^*(x)>-\infty,\\-1/\varepsilon&\text{if }f^*(x)=-\infty.\end{cases}f(x,φε​(x))≤{f∗(x)+ε−1/ε​if f∗(x)>−∞,if f∗(x)=−∞.​

Significance

The results. Propositions 7.31–7.33 are the closure properties that make the dynamic programming recursion stay inside the class of lower semicontinuous functions bounded below: the expectation step preserves the class (7.31), the minimization step preserves it (7.32, 7.33), and the minimization admits a Borel-measurable exact minimizer (7.33). This is why, in semicontinuous models, the optimal cost functions are lower semicontinuous and optimal policies can be taken Borel-measurable and nonrandomized. Proposition 7.34 gives the weaker, ε\varepsilonε-optimal counterpart for upper semicontinuous costs, where the infimum need not be attained.

Formalizing them. All of these results are proved in the book; none is open. As far as is known, none has a machine-checked proof: Mathlib has semicontinuity, the Hausdorff extended metric on closed and on nonempty compact sets, and the weak topology on probability measures, but no theorem combining them into a measurable selection result of this kind. A formal development would supply measurable selectors for semicontinuous minimization in Lean and the Borel-measurability of set-valued maps into the hyperspace of closed sets, both reusable well beyond dynamic programming.

Difficulty

The obvious attempt at the goal is to pick, for each xxx, some minimizer yyy of f(x,⋅)f(x,\cdot)f(x,⋅) over the compact section DxD_xDx​. The minimizer exists by compactness and lower semicontinuity, but the choice is made pointwise and gives no control on measurability: a minimizer chosen by the axiom of choice need not be Borel-measurable. The argmin sets F∗(x)F^*(x)F∗(x) vary with xxx only semicontinuously: they can jump from a single point to a large set, so a continuous selection generally does not exist, and continuity arguments cannot replace measurability. Lemma 7.18 isolates the hardest part: a choice of a point of each nonempty closed set that is measurable as a function of the set itself.

A second difficulty is bookkeeping at infinity. Values ±∞\pm\infty±∞ are allowed throughout, so sublevel sets, minima, integrals and ε\varepsilonε-bounds must all be handled in R∗R^*R∗; the integral in Proposition 7.31 uses the convention ∞−∞=+∞\infty-\infty=+\infty∞−∞=+∞, which is not Mathlib's.

Formalization scope

  • Extended reals. Values are in EReal. The only place where values of opposite infinite sign are combined is the integral, which is the published definition DupacovaWets.Consistency.expect (reused, not restated): ∫f+−∫f−\int f^+-\int f^-∫f+−∫f− with an explicit case returning +∞+\infty+∞ when ∫f+=∞\int f^+=\infty∫f+=∞, exactly the book's convention (42). The ε\varepsilonε-bound of Proposition 7.34 adds a real ε\varepsilonε to a value different from −∞-\infty−∞, which is safe in EReal.
  • Semicontinuity is Mathlib's LowerSemicontinuous/UpperSemicontinuous, equivalent to Definition 7.13 for EReal-valued functions. Lemma 7.13 of the book (the sequential characterization) is Mathlib's lowerSemicontinuous_iff_le_liminf together with first countability of metrizable spaces, and is not restated here.
  • Functions on DDD. Functions "on DDD" are functions on X×YX\times YX×Y with LowerSemicontinuousOn f D (resp. UpperSemicontinuousOn); values off DDD play no role. projX(D)\mathrm{proj}_X(D)projX​(D) is Prod.fst '' D, selectors are functions on that subtype, and its σ-algebra is the Borel σ-algebra of the subspace topology.
  • Hyperspace. 2Y2^Y2Y is Closeds Y, and 2Y−{∅}2^Y-\{\emptyset\}2Y−{∅} for compact YYY is NonemptyCompacts Y, each with the Hausdorff extended metric and the Borel σ-algebra of its topology. This topology agrees with the book's (the exponential topology of Appendix C, independent of the metric).
  • Boundedness. "Bounded below/above" is by a real constant. BddBelow in EReal would be vacuous and is not used.
  • Edge cases. Proposition 7.32(a)'s attainment clause is stated for nonempty YYY, since for Y=∅Y=\emptysetY=∅ the infimum is +∞+\infty+∞ and nothing attains it.
  • Argmin minimum. Lemma 7.20 assumes nonempty YYY because its defining formula uses a minimum; for empty YYY there is no minimizer.
  • Ruling out trivial readings. The graph condition (x,φ(x))∈D(x,\varphi(x))\in D(x,φ(x))∈D is part of every selection statement; without it the goal would follow from the unconstrained case. The selector must be Borel-measurable on projX(D)\mathrm{proj}_X(D)projX​(D) and must attain the minimum exactly, not up to ε\varepsilonε.

A complete development needs the Borel structure of the hyperspace (measurability of maps into Closeds Y from upper semicontinuity in the sense of Kuratowski, Proposition C.4 of the book), the construction of a measurable choice function on NonemptyCompacts Y, and approximation of semicontinuous functions by monotone sequences in C(X)C(X)C(X). Each of these is reusable on its own; proofs of individual milestones by any route are welcome.

Selected references

  • D. P. Bertsekas and S. E. Shreve, Stochastic Optimal Control: The Discrete-Time Case, Academic Press, 1978; Athena Scientific reprint, 1996, Section 7.5 and Appendix C. https://web.mit.edu/dimitrib/www/soc.html
  • L. E. Dubins and L. J. Savage, How to Gamble If You Must: Inequalities for Stochastic Processes, McGraw-Hill, 1965.
  • K. Kuratowski and C. Ryll-Nardzewski, "A general theorem on selectors," Bull. Acad. Polon. Sci. 13 (1965), 397–403.
  • F. Hausdorff, Set Theory, Chelsea, New York, 1957.
10 thms1 active userReviewed
PreviousPage 24 of 27Next

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me