Prove2Me
Navigate
DiscoverFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

Probability

550 missions · 266 completed

Missions

Open284Completed266All550
🏆Completed
Machine LearningOperations ResearchOptimization+1·Captain: mikedeng1

A Distributional Interpretation of Robust Optimization II: Box-Robust Sample Average Optimization Is ConsistentResearch Paper

Why robustify a sampled stochastic program

Many decision problems under uncertainty take the form of a stochastic program: choose a decision vvv from a feasible set F\mathcal FF to maximise the expected utility Ex∼μ[f(v,x)]\mathbb E_{x\sim\mu}[f(v,x)]Ex∼μ​[f(v,x)], where the distribution μ\muμ of the uncertain parameter x∈Rmx\in\mathbb R^mx∈Rm is known only through i.i.d. samples x1,…,xnx_1,\dots,x_nx1​,…,xn​. The standard remedy, sample average approximation, maximises 1n∑if(v,xi)\frac1n\sum_i f(v,x_i)n1​∑i​f(v,xi​) instead. Its consistency (convergence of the optimal expected utility of its solutions to the true optimum) is classical, but it needs regularity assumptions of its own, for example those of King and Wets (Stochastics and Stochastic Reports, 1991), cited on p. 98 of the paper; the paper presents its construction as a route to consistency under weaker conditions.

Robust optimization (RO) takes a different route: it protects each sample by an uncertainty set and optimises against the worst point in it. Xu, Caramanis and Mannor (Math. Oper. Res. 2012) show that RO over several overlapping uncertainty sets is equivalent to a distributionally robust stochastic program (their Theorem 2.1, the subject of mission I of this series). Section 3 of the paper uses that equivalence to show that a specific robustification of the sampled problem, with ℓ∞\ell_\inftyℓ∞​ boxes of shrinking radius around each sample, is consistent under only boundedness and equicontinuity of fff. This mission formalizes that result, Theorem 3.1.

Setting

Equip Rm\mathbb R^mRm with the sup norm ∥z∥∞=max⁡k∣zk∣\|z\|_\infty=\max_k|z_k|∥z∥∞​=maxk​∣zk​∣, its Borel σ\sigmaσ-algebra and Lebesgue measure dxdxdx. The data are:

  • a set of decisions VVV and a nonempty feasible set F⊆V\mathcal F\subseteq VF⊆V;
  • a utility f:V×Rm→Rf:V\times\mathbb R^m\to\mathbb Rf:V×Rm→R, Borel measurable in xxx for each vvv;
  • a true density h∗h^*h∗ on Rm\mathbb R^mRm (nonnegative, ∫h∗ dx=1\int h^*\,dx=1∫h∗dx=1) and i.i.d. samples x1,x2,…x_1,x_2,\dotsx1​,x2​,… with distribution h∗(x) dxh^*(x)\,dxh∗(x)dx;
  • radii ϵ(n)>0\epsilon(n)>0ϵ(n)>0.

For a sample x1,…,xnx_1,\dots,x_nx1​,…,xn​ the boxes are Zi={xi+δ∣∥δ∥∞≤ϵ(n)}\mathcal Z_i=\{x_i+\delta\mid\|\delta\|_\infty\le\epsilon(n)\}Zi​={xi​+δ∣∥δ∥∞​≤ϵ(n)}, and the box-robust sample objective is

Jn(v)=1n∑i=1n inf⁡∥δi∥∞≤ϵ(n)f(v,xi+δi)=∑i=1n1ninf⁡xi′∈Zif(v,xi′).J_n(v)=\frac1n\sum_{i=1}^n\ \inf_{\|\delta_i\|_\infty\le\epsilon(n)}f(v,x_i+\delta_i)=\sum_{i=1}^n\frac1n\inf_{x_i'\in\mathcal Z_i}f(v,x_i').Jn​(v)=n1​i=1∑n​ ∥δi​∥∞​≤ϵ(n)inf​f(v,xi​+δi​)=i=1∑n​n1​xi′​∈Zi​inf​f(v,xi′​).

The RO solution v(n)v(n)v(n) is a maximiser of JnJ_nJn​ over F\mathcal FF. The equicontinuity modulus of fff is

d(ϵ)=sup⁡v, x, ∥δ∥∞≤ϵ∣f(v,x)−f(v,x+δ)∣.d(\epsilon)=\sup_{v,\,x,\ \|\delta\|_\infty\le\epsilon}|f(v,x)-f(v,x+\delta)|.d(ϵ)=v,x, ∥δ∥∞​≤ϵsup​∣f(v,x)−f(v,x+δ)∣.

The proof works with the distribution set Pn\mathcal P_nPn​ of probability measures μ\muμ with μ(⋃i∈SZi)≥∣S∣/n\mu(\bigcup_{i\in S}\mathcal Z_i)\ge|S|/nμ(⋃i∈S​Zi​)≥∣S∣/n for every S⊆{1,…,n}S\subseteq\{1,\dots,n\}S⊆{1,…,n}, and with the uniform box kernel density estimator

hn(x)=(nϵ(n)m)−1∑i=1nK(x−xiϵ(n)),K(z)=1(∥z∥∞≤1)2m.h_n(x)=(n\epsilon(n)^m)^{-1}\sum_{i=1}^nK\Big(\frac{x-x_i}{\epsilon(n)}\Big),\qquad K(z)=\frac{\mathbf 1(\|z\|_\infty\le1)}{2^m}.hn​(x)=(nϵ(n)m)−1i=1∑n​K(ϵ(n)x−xi​​),K(z)=2m1(∥z∥∞​≤1)​.

Formalization targets

Goal: Theorem 3.1 (p. 98)

Assume ∣f(v,x)∣≤C|f(v,x)|\le C∣f(v,x)∣≤C for all v,xv,xv,x; d(ϵ)→0d(\epsilon)\to0d(ϵ)→0 as ϵ↓0\epsilon\downarrow0ϵ↓0; ϵ(n)↓0\epsilon(n)\downarrow0ϵ(n)↓0 and nϵ(n)m↑∞n\epsilon(n)^m\uparrow\inftynϵ(n)m↑∞. Then for every choice of maximisers v(n)v(n)v(n), with probability one,

lim⁡n→∞∫Rmf(v(n),x) h∗(x) dx=sup⁡v∈F∫Rmf(v,x) h∗(x) dx.\lim_{n\to\infty}\int_{\mathbb R^m}f(v(n),x)\,h^*(x)\,dx=\sup_{v\in\mathcal F}\int_{\mathbb R^m}f(v,x)\,h^*(x)\,dx .n→∞lim​∫Rm​f(v(n),x)h∗(x)dx=v∈Fsup​∫Rm​f(v,x)h∗(x)dx.

Milestones (proof of Theorem 3.1, p. 99)

  1. hnh_nhn​ is the density of a probability measure in Pn\mathcal P_nPn​.
  2. Jn(v)≤∫f(v,x) hn(x) dxJ_n(v)\le\int f(v,x)\,h_n(x)\,dxJn​(v)≤∫f(v,x)hn​(x)dx for every vvv.
  3. Oscillation over a box: sup⁡Zif(v,⋅)−inf⁡Zif(v,⋅)≤d(2ϵ(n))\sup_{\mathcal Z_i}f(v,\cdot)-\inf_{\mathcal Z_i}f(v,\cdot)\le d(2\epsilon(n))supZi​​f(v,⋅)−infZi​​f(v,⋅)≤d(2ϵ(n)).
  4. Eq. (7): with Mn=C∫∣hn−h∗∣ dxM_n=C\int|h_n-h^*|\,dxMn​=C∫∣hn​−h∗∣dx, for every vvv,
Jn(v)−Mn≤∫f(v,x)h∗(x) dx≤Jn(v)+Mn+d(2ϵ(n)).J_n(v)-M_n\le\int f(v,x)h^*(x)\,dx\le J_n(v)+M_n+d(2\epsilon(n)).Jn​(v)−Mn​≤∫f(v,x)h∗(x)dx≤Jn​(v)+Mn​+d(2ϵ(n)).
  1. Strong L1L^1L1 consistency of the box kernel density estimator: if ϵ(n)→0\epsilon(n)\to0ϵ(n)→0 and nϵ(n)m→∞n\epsilon(n)^m\to\inftynϵ(n)m→∞, then ∫∣hn−h∗∣ dx→0\int|h_n-h^*|\,dx\to0∫∣hn​−h∗∣dx→0 almost surely.

Milestones 1–4 are deterministic statements about a fixed sample; milestone 5 is the only probabilistic input.

Significance

Theorem 3.1 gives consistency of a tractable robust reformulation of a sampled stochastic program under conditions the paper notes are weaker than those of King and Wets for sampled stochastic programs: fff need only be bounded and equicontinuous in xxx, uniformly in vvv, and the true distribution need only have a density. It also gives an explicit schedule for the size of the uncertainty set, ϵ(n)→0\epsilon(n)\to0ϵ(n)→0 with nϵ(n)m→∞n\epsilon(n)^m\to\inftynϵ(n)m→∞, the bandwidth condition of kernel density estimation. Section 4 of the paper applies the same distributional interpretation to regularised learning methods such as the support vector machine and the Lasso.

The result is proved in the paper, with the L1L^1L1 consistency of kernel density estimators (Devroye 1983; Devroye and Györfi 1985) cited rather than proved. No part of it is formalized in Lean or on this platform as far as a search of the platform found. A complete development would produce, besides Theorem 3.1, a machine-checked strong L1L^1L1 consistency theorem for kernel density estimators, which is a basic result of nonparametric statistics in its own right.

Difficulty

The deterministic part (milestones 1–4) is measure-theoretic bookkeeping: the kernel integrates to one only because the box is a sup-norm ball of volume (2ϵ)m(2\epsilon)^m(2ϵ)m, and every infimum and supremum must be handled with care, since fff need not attain them.

The obstacle is milestone 5. Almost-sure L1L^1L1 convergence of hnh_nhn​ to an arbitrary density h∗h^*h∗, with no continuity or support assumption, does not follow from the strong law of large numbers applied pointwise: hn(x)h_n(x)hn​(x) is an average of nnn terms whose law changes with nnn through ϵ(n)\epsilon(n)ϵ(n), and almost-sure convergence at each fixed xxx does not give convergence of the integral along a single sample path. The theorem needs both a bias estimate valid for every integrable density and a concentration estimate for the random L1L^1L1 error. Mathlib has Lebesgue differentiation and the strong law, but no kernel density estimator and no such concentration result.

Formalization scope

  • Rm\mathbb R^mRm is Fin m → ℝ, whose Mathlib norm is the sup norm; boxes are Metric.closedBall. The integrals ∫f(v,x)h∗(x) dx\int f(v,x)h^*(x)\,dx∫f(v,x)h∗(x)dx are Bochner integrals against Lebesgue measure of integrable integrands.
  • The samples are a sequence X : ℕ → Ω → Fin m → ℝ on a probability space, independent (iIndepFun) and each with law volume.withDensity h*; x1,x2,…x_1,x_2,\dotsx1​,x2​,… become X 0, X 1, …, and the nnn-th problem uses the first nnn. "With probability one" is ∀ᵐ ω ∂P.
  • The goal quantifies over every selection v(n)v(n)v(n) of maximisers, with no measurability assumed; a version with one chosen maximiser would be weaker and is ruled out.
  • Readings and corrections of the printed text:
    • the kernel argument printed (x−xi)/ϵ(x-x_i)/\epsilon(x−xi​)/ϵ on p. 98 is read as (x−xi)/ϵ(n)(x-x_i)/\epsilon(n)(x−xi​)/ϵ(n), as the proof on p. 99 writes it;
    • "max⁡v,x∣f(v,x)∣≤C\max_{v,x}|f(v,x)|\le Cmaxv,x​∣f(v,x)∣≤C" is read as the uniform bound ∣f∣≤C|f|\le C∣f∣≤C and the "max" in d(ϵ)d(\epsilon)d(ϵ) as a supremum;
    • "d(ϵ)↓0d(\epsilon)\downarrow0d(ϵ)↓0" is read as d(ϵ)→0d(\epsilon)\to0d(ϵ)→0 as ϵ↓0\epsilon\downarrow0ϵ↓0;
    • implicit hypotheses made explicit: F≠∅\mathcal F\ne\emptysetF=∅, ϵ(n)>0\epsilon(n)>0ϵ(n)>0, measurability of f(v,⋅)f(v,\cdot)f(v,⋅), h∗h^*h∗ a Lebesgue density;
    • the monotonicity in "ϵ(n)↓0\epsilon(n)\downarrow0ϵ(n)↓0, nϵ(n)m↑∞n\epsilon(n)^m\uparrow\inftynϵ(n)m↑∞" is kept in the goal; milestone 5 uses the limits only, as the paper states it;
    • the paper's MnM_nMn​ ("there exists {Mn}→0\{M_n\}\to0{Mn​}→0") is made explicit as Mn=C∫∣hn−h∗∣M_n=C\int|h_n-h^*|Mn​=C∫∣hn​−h∗∣, so Eq. (7) is stated for every sample.
  • Remark 3.2 and Appendix B (an integrable envelope in place of boundedness) are not part of this mission.
  • Every real infimum and supremum ranges over a nonempty set of values bounded by CCC in absolute value, so no statement holds through a junk value; a formalization in which the supremum over F\mathcal FF or the box infimum could be vacuous is excluded.
  • The definitions (boxes, Pn\mathcal P_nPn​, the kernel, the estimator, JnJ_nJn​, ddd) live in one definition file. Pn\mathcal P_nPn​ duplicates, with weights 1/n1/n1/n, the distribution set of mission I; the duplication is deliberate because draft missions cannot import each other.
  • Welcome contributions: the kernel density estimator and its strong L1L^1L1 consistency as reusable infrastructure, and any of the deterministic milestones.

Selected references

  • H. Xu, C. Caramanis, S. Mannor, A Distributional Interpretation of Robust Optimization, Mathematics of Operations Research 37(1):95–110, 2012. https://doi.org/10.1287/moor.1110.0531
  • L. Devroye, The equivalence of weak, strong and complete convergence in L1L_1L1​ for kernel density estimates, Annals of Statistics 11(3):896–904, 1983.
  • L. Devroye, L. Györfi, Nonparametric Density Estimation: The L1L_1L1​ View, Wiley, 1985.
  • A. J. King, R. J.-B. Wets, Epi-consistency of convex stochastic programs, Stochastics and Stochastic Reports 34(1), 1991 (reference [22] of the paper).
7 thms4 active usersReviewed
🏆Completed
Linear OptimizationOperations ResearchOptimization·Captain: mikedeng1

Robust solutions of Linear Programming problems contaminated with uncertain data: A Violation-Probability Bound for the Robust CounterpartResearch Paper

Motivation

Linear programs solved in practice carry data that are measured, estimated or rounded. Ben-Tal and Nemirovski (Math. Program. 88, 2000) examined the NETLIB collection of real-world LPs and found that in 13 of them a relative perturbation of only 0.01% in the "ugly" coefficients of the inequality constraints can make the nominal optimal solution more than 50% infeasible (§2.3). Their remedy is the robust counterpart methodology: replace the nominal problem by a deterministic problem whose feasible solutions remain nearly feasible for every, or for all but a small probability of, data realizations.

The paper made the approach concrete for entry-wise uncertainty and gave the probabilistic guarantee that became a standard tool of robust and chance-constrained optimization. The main steps of the history are:

  • 1973, A. L. Soyster: the interval (worst-case, "box") counterpart, here called (IRC).
  • 1998–1999, Ben-Tal and Nemirovski (Math. Oper. Res. 23; Oper. Res. Lett. 25), and independently El Ghaoui and co-authors: robust optimization with ellipsoidal uncertainty sets.
  • 2000, this paper: the ellipsoid-plus-box counterpart (RC[ε, δ, Ω]) and Proposition 1, which bounds each constraint's violation probability by exp⁡{−Ω2/2}\exp\{-\Omega^2/2\}exp{−Ω2/2} under independent symmetric perturbations.
  • 2004, Bertsimas and Sim (Oper. Res. 52): the budgeted counterpart, with an analogous probability bound.

Setting

An uncertain linear program is

minimize cTxs.t.Ex=e,Ax≤b,ℓ≤x≤u,(LP)\text{minimize } c^Tx\quad\text{s.t.}\quad Ex=e,\qquad Ax\le b,\qquad \ell\le x\le u,\tag{LP}minimize cTxs.t.Ex=e,Ax≤b,ℓ≤x≤u,(LP)

with x∈Rnx\in\mathbb{R}^nx∈Rn, E∈Rp×nE\in\mathbb{R}^{p\times n}E∈Rp×n, A=(aij)∈Rm×nA=(a_{ij})\in\mathbb{R}^{m\times n}A=(aij​)∈Rm×n, and bounds ℓj∈R∪{−∞}\ell_j\in\mathbb{R}\cup\{-\infty\}ℓj​∈R∪{−∞}, uj∈R∪{+∞}u_j\in\mathbb{R}\cup\{+\infty\}uj​∈R∪{+∞}. For each inequality row iii a set Ji⊆{1,…,n}J_i\subseteq\{1,\dots,n\}Ji​⊆{1,…,n} lists the uncertain entries aija_{ij}aij​, j∈Jij\in J_ij∈Ji​. Only these entries are uncertain; E,e,b,ℓ,u,cE,e,b,\ell,u,cE,e,b,ℓ,u,c are exact.

Given an uncertainty level ϵ>0\epsilon>0ϵ>0 and a feasibility tolerance δ>0\delta>0δ>0, write bi+=bi+δmax⁡[1,∣bi∣]b_i^+=b_i+\delta\max[1,|b_i|]bi+​=bi​+δmax[1,∣bi​∣].

  • xxx is reliable if it is feasible for (LP) and ∑j∉Jiaijxj+∑j∈Jia~ijxj≤bi+\sum_{j\notin J_i}a_{ij}x_j+\sum_{j\in J_i}\tilde a_{ij}x_j\le b_i^+∑j∈/Ji​​aij​xj​+∑j∈Ji​​a~ij​xj​≤bi+​ for every iii and every choice of a~ij\tilde a_{ij}a~ij​ with ∣a~ij−aij∣≤ϵ∣aij∣|\tilde a_{ij}-a_{ij}|\le\epsilon|a_{ij}|∣a~ij​−aij​∣≤ϵ∣aij​∣.
  • In the random symmetric uncertainty model, the true coefficients are a~ij=(1+ϵξij)aij\tilde a_{ij}=(1+\epsilon\xi_{ij})a_{ij}a~ij​=(1+ϵξij​)aij​, where ξij=0\xi_{ij}=0ξij​=0 for j∉Jij\notin J_ij∈/Ji​ and, for each row iii, {ξij}j∈Ji\{\xi_{ij}\}_{j\in J_i}{ξij​}j∈Ji​​ are independent random variables, each symmetrically distributed in [−1,1][-1,1][−1,1].
  • xxx is almost reliable with level κ\kappaκ if it is feasible for (LP) and Pr⁡{∑ja~ijxj>bi+}≤κ\Pr\{\sum_j\tilde a_{ij}x_j>b_i^+\}\le\kappaPr{∑j​a~ij​xj​>bi+​}≤κ for every iii.

The robust counterpart (RC[ε, δ, Ω]), with a safety parameter Ω>0\Omega>0Ω>0, has variables xjx_jxj​, yijy_{ij}yij​, zijz_{ij}zij​ and constraints Ex=eEx=eEx=e, Ax≤bAx\le bAx≤b, ℓ≤x≤u\ell\le x\le uℓ≤x≤u, −yij≤xj−zij≤yij-y_{ij}\le x_j-z_{ij}\le y_{ij}−yij​≤xj​−zij​≤yij​ for all i,ji,ji,j, and

∑jaijxj+ϵ[∑j∈Ji∣aij∣yij+Ω∑j∈Jiaij2zij2]≤bi+∀i.\sum_j a_{ij}x_j+\epsilon\Big[\sum_{j\in J_i}|a_{ij}|y_{ij}+\Omega\sqrt{\sum_{j\in J_i}a_{ij}^2z_{ij}^2}\Big]\le b_i^+\qquad\forall i.j∑​aij​xj​+ϵ[j∈Ji​∑​∣aij​∣yij​+Ωj∈Ji​∑​aij2​zij2​​]≤bi+​∀i.

The interval robust counterpart (IRC[ε, δ]) has variables xj,yjx_j,y_jxj​,yj​ and the constraint ∑jaijxj+ϵ∑j∈Ji∣aij∣yj≤bi+\sum_ja_{ij}x_j+\epsilon\sum_{j\in J_i}|a_{ij}|y_j\le b_i^+∑j​aij​xj​+ϵ∑j∈Ji​​∣aij​∣yj​≤bi+​ with −yj≤xj≤yj-y_j\le x_j\le y_j−yj​≤xj​≤yj​, besides the nominal ones. Problem (∗) is the same with yjy_jyj​ replaced by ∣xj∣|x_j|∣xj​∣.

Formalization targets

Goal: Proposition 1 (pp. 418–419)

If xxx extends to a feasible solution (x,y,z)(x,y,z)(x,y,z) of (RC[ε, δ, Ω]), then xxx is feasible for (LP) and, for every iii,

Pr⁡{∑j(1+ϵξij)aijxj>bi+δmax⁡[1,∣bi∣]}≤exp⁡{−Ω2/2}.\Pr\Big\{\sum_j(1+\epsilon\xi_{ij})a_{ij}x_j>b_i+\delta\max[1,|b_i|]\Big\}\le\exp\{-\Omega^2/2\}.Pr{j∑​(1+ϵξij​)aij​xj​>bi​+δmax[1,∣bi​∣]}≤exp{−Ω2/2}.

Milestones

  1. The reduction in the proof of Proposition 1 (p. 419), in corrected pointwise form: a violation of row iii forces ∑j∈Jiξijaijzij>Ω∑j∈Jiaij2zij2\sum_{j\in J_i}\xi_{ij}a_{ij}z_{ij}>\Omega\sqrt{\sum_{j\in J_i}a_{ij}^2z_{ij}^2}∑j∈Ji​​ξij​aij​zij​>Ω∑j∈Ji​​aij2​zij2​​.
  2. Eq. (1), p. 419: for independent symmetric ηj∈[−1,1]\eta_j\in[-1,1]ηj​∈[−1,1] and reals pjp_jpj​,
Pr⁡{∑jηjpj>Ω∑jpj2}≤exp⁡{−Ω2/2}.\Pr\Big\{\sum_j\eta_jp_j>\Omega\sqrt{\textstyle\sum_jp_j^2}\Big\}\le\exp\{-\Omega^2/2\}.Pr{j∑​ηj​pj​>Ω∑j​pj2​​}≤exp{−Ω2/2}.
  1. xxx is reliable iff it is feasible for (∗) (p. 417).
  2. (∗) is equivalent to (IRC[ε, δ]) (pp. 417–418).
  3. Every feasible solution of (IRC) yields one of (RC) with yij=yjy_{ij}=y_jyij​=yj​, zij=0z_{ij}=0zij​=0 (p. 420).
  4. Feasibility for (LP) together with ∑jaijxj+ϵβi(x)≤bi+\sum_ja_{ij}x_j+\epsilon\beta_i(x)\le b_i^+∑j​aij​xj​+ϵβi​(x)≤bi+​, βi(x)=Ω∑j∈Jiaij2xj2\beta_i(x)=\Omega\sqrt{\sum_{j\in J_i}a_{ij}^2x_j^2}βi​(x)=Ω∑j∈Ji​​aij2​xj2​​, suffices to extend xxx to (RC) (p. 420).
  5. The ratio αi(x)/βi(x)\alpha_i(x)/\beta_i(x)αi​(x)/βi​(x), αi(x)=∑j∈Ji∣aij∣∣xj∣\alpha_i(x)=\sum_{j\in J_i}|a_{ij}||x_j|αi​(x)=∑j∈Ji​​∣aij​∣∣xj​∣, is at most card(Ji)/Ω\sqrt{\mathrm{card}(J_i)}/\Omegacard(Ji​)​/Ω, with equality attained (p. 420, corrected).

Significance

Proposition 1 turns a probabilistic requirement, which is hard to handle directly, into a single convex (second-order-cone) program. The bound exp⁡{−Ω2/2}\exp\{-\Omega^2/2\}exp{−Ω2/2} does not depend on the dimension, on the number of uncertain entries, or on which symmetric distributions the perturbations follow, so Ω\OmegaΩ can be chosen from the desired reliability level alone. Together with milestones 3–6, the mission certifies the whole chain: the worst-case notion of reliability is exactly Soyster's linear program (IRC), and (RC) is never more conservative than (IRC), with an advantage that can reach the factor card(Ji)/Ω\sqrt{\mathrm{card}(J_i)}/\Omegacard(Ji​)​/Ω.

The results are proved in the paper; to our knowledge none of them is machine-checked. A formal development produces a reusable model of entry-wise uncertain LPs, the counterparts (∗), (IRC) and (RC) as Lean predicates, and a Hoeffding-type bound for weighted sums of symmetric bounded variables in the exact form (1). The platform's HighDimProb.Concentration.hoeffding_rademacher covers the Rademacher special case only.

Difficulty

The deterministic parts (milestones 1, 3–7) are elementary: worst cases of interval perturbations, and the Cauchy–Schwarz inequality. The obstacle lies in the probabilistic step. The printed proof passes from ξijaij\xi_{ij}a_{ij}ξij​aij​ to ξij∣aij∣\xi_{ij}|a_{ij}|ξij​∣aij​∣ with an equality that holds only in distribution, and contains index misprints, so it cannot be transcribed line by line; the reduction has to be restated pointwise. Eq. (1) is a tail bound for general symmetric variables in [−1,1][-1,1][−1,1], not only for random signs; the step (c) of the printed proof of (1) is written as an equality that holds only for random signs, so that proof too needs repair. The degenerate case ∑jpj2=0\sum_jp_j^2=0∑j​pj2​=0 must be handled rather than assumed away.

Formalization scope

  • Data are a structure UncertainLP n p m over Fin indices (0-based), with A : Matrix (Fin m) (Fin n) ℝ, J : Fin m → Finset (Fin n) arbitrary, and EReal bounds so that infinite bounds are expressible. The objective ccc is omitted: no statement involves it.
  • The probability space is (S,P)(S,\mathbb P)(S,P) with IsProbabilityMeasure; the name SSS avoids a clash with the safety parameter Ω\OmegaΩ. Symmetry is equality of the laws of ξij\xi_{ij}ξij​ and −ξij-\xi_{ij}−ξij​; values lie in [−1,1][-1,1][−1,1] at every outcome; independence is required within each row only, with no identical distribution (§3.1 says only "independent", which is weaker than the "iid" of §2.2). Probabilities are P.real.
  • The hypotheses ϵ>0\epsilon>0ϵ>0, δ>0\delta>0δ>0, Ω>0\Omega>0Ω>0 are the paper's standing assumptions and are carried by every theorem that mentions the parameter.
  • Corrections of the printed text: aijxi→aijxja_{ij}x_i\to a_{ij}x_jaij​xi​→aij​xj​ in (IRC); ∑j∈J→∑j∈Ji\sum_{j\in J}\to\sum_{j\in J_i}∑j∈J​→∑j∈Ji​​ in (RC); the reduction of milestone 1 is stated with aija_{ij}aij​ and zijz_{ij}zij​ in place of the printed ∣aij∣|a_{ij}|∣aij​∣, xi−yijx_i-y_{ij}xi​−yij​ and yjy_jyj​, yijy_{ij}yij​; and the ratio of milestone 7 carries the factor 1/Ω1/\Omega1/Ω that the printed "card(Ji)\sqrt{\mathrm{card}(J_i)}card(Ji​)​" omits.
  • Ruling out trivializations: the violation event uses the signed multiplicative model (1+ϵξij)aij(1+\epsilon\xi_{ij})a_{ij}(1+ϵξij​)aij​, never ∣aij∣|a_{ij}|∣aij​∣ or an additive perturbation; the goal concludes both nominal feasibility (i) and the probability bound (ii′) for every row; no hypothesis excludes the degenerate case ∑j∈Jiaij2zij2=0\sum_{j\in J_i}a_{ij}^2z_{ij}^2=0∑j∈Ji​​aij2​zij2​=0; and the probability model is satisfiable (e.g. by ξ≡0\xi\equiv0ξ≡0 or by Rademacher signs), so the goal is not vacuous.
  • The numerical remarks of the paper (0.92, 5.24, 10−610^{-6}10−6, "at least 30") and the NETLIB case study are not formalized.
  • Reusable beyond this mission: the uncertain-LP model and the three counterparts, and the tail bound (1). Contributions of general lemmas about symmetric bounded random variables are welcome.

Selected references

  • A. Ben-Tal, A. Nemirovski, Robust solutions of Linear Programming problems contaminated with uncertain data, Math. Program. Ser. A 88 (2000) 411–424. https://doi.org/10.1007/s101070000163
  • A. L. Soyster, Convex programming with set-inclusive constraints and applications to inexact linear programming, Oper. Res. 21 (1973) 1154–1157. https://doi.org/10.1287/opre.21.5.1154
  • A. Ben-Tal, A. Nemirovski, Robust convex optimization, Math. Oper. Res. 23 (1998) 769–805. https://doi.org/10.1287/moor.23.4.769
  • A. Ben-Tal, A. Nemirovski, Robust solutions of uncertain linear programs, Oper. Res. Lett. 25 (1999) 1–13. https://doi.org/10.1016/S0167-6377(99)00016-4
  • D. Bertsimas, M. Sim, The price of robustness, Oper. Res. 52 (2004) 35–53. https://doi.org/10.1287/opre.1030.0065
  • W. Hoeffding, Probability inequalities for sums of bounded random variables, J. Amer. Statist. Assoc. 58 (1963) 13–30. https://doi.org/10.1080/01621459.1963.10500830
12 thms4 active usersReviewed
🏆Completed
Machine LearningStatistics·Captain: mikedeng1

Learnability, Stability and Uniform Convergence III: For an ERM, Leave-One-Out Stability, Universal Consistency and Universal Generalization Are EquivalentResearch Paper

Motivation

Algorithmic stability asks how much the output of a learning algorithm changes when its training sample is perturbed. Since Devroye and Wagner (IEEE Trans. Inf. Theory 1979) it has served as a route to generalization bounds that does not go through the complexity of the hypothesis class. Bousquet and Elisseeff (JMLR 2002) popularised uniform stability, and Mukherjee, Niyogi, Poggio and Rifkin (Adv. Comput. Math. 2006) showed that for empirical risk minimisation in supervised learning, a leave-one-out type of stability is necessary and sufficient for consistency.

Shalev-Shwartz, Shamir, Srebro and Sridharan (JMLR 11, 2010) study stability in Vapnik's General Learning Setting, where uniform convergence can fail even though the problem is learnable. In Appendix A.2 they compare replace-one and leave-one-out (LOO) stability. For an empirical risk minimiser they prove that LOO stability is equivalent to consistency and to generalization, provided each property holds with one rate for all distributions (Theorem 31, p. 2667). This mission formalizes that theorem and the lemmas of Section 5.3 on which its proof rests.

Timeline:

  • 1979, Devroye–Wagner: leave-one-out estimates for local rules.
  • 2002, Bousquet–Elisseeff: uniform stability implies generalization.
  • 2002, Kutin–Niyogi (UAI 2002): a taxonomy of stability notions.
  • 2006, Mukherjee et al.: LOO stability characterises consistency of ERM in supervised learning.
  • 2010, Shalev-Shwartz et al.: in the General Learning Setting, for ERMs, LOO stability, universal consistency and universal generalization are equivalent (Theorem 31). Universally consistent AERMs need not be LOO stable (Example 6).

Setting

A learning problem consists of an instance space Z\mathcal ZZ with a σ\sigmaσ-algebra, a nonempty hypothesis class H\mathcal HH, and an objective f:H×Z→Rf:\mathcal H\times\mathcal Z\to\mathbb Rf:H×Z→R with ∣f(h;z)∣≤B|f(h;z)|\le B∣f(h;z)∣≤B for all h,zh,zh,z. For a probability measure D\mathcal DD on Z\mathcal ZZ:

  • the risk is F(h)=Ez∼D[f(h;z)]F(h)=\mathbb E_{z\sim\mathcal D}[f(h;z)]F(h)=Ez∼D​[f(h;z)] and the optimal risk is F∗=inf⁡hF(h)F^*=\inf_{h}F(h)F∗=infh​F(h);
  • for a sample S=(z1,…,zm)∼DmS=(z_1,\dots,z_m)\sim\mathcal D^mS=(z1​,…,zm​)∼Dm of mmm i.i.d. draws, the empirical risk is FS(h)=1m∑if(h;zi)F_S(h)=\frac1m\sum_{i}f(h;z_i)FS​(h)=m1​∑i​f(h;zi​), and FS(h^S)=inf⁡hFS(h)F_S(\hat h_S)=\inf_hF_S(h)FS​(h^S​)=infh​FS​(h) denotes the minimal empirical risk;
  • a learning rule AAA maps each sample of size m≥1m\ge1m≥1 to a hypothesis A(S)A(S)A(S). It is an ERM if FS(A(S))=FS(h^S)F_S(A(S))=F_S(\hat h_S)FS​(A(S))=FS​(h^S​) for every sample. It is an AERM with rate εerm\varepsilon_{\mathrm{erm}}εerm​ if E[FS(A(S))−FS(h^S)]≤εerm(m)\mathbb E[F_S(A(S))-F_S(\hat h_S)]\le\varepsilon_{\mathrm{erm}}(m)E[FS​(A(S))−FS​(h^S​)]≤εerm​(m);
  • AAA is consistent with rate ε\varepsilonε if ES∼Dm[F(A(S))−F∗]≤ε(m)\mathbb E_{S\sim\mathcal D^m}[F(A(S))-F^*]\le\varepsilon(m)ES∼Dm​[F(A(S))−F∗]≤ε(m). It generalizes with rate ε\varepsilonε if E[∣F(A(S))−FS(A(S))∣]≤ε(m)\mathbb E[|F(A(S))-F_S(A(S))|]\le\varepsilon(m)E[∣F(A(S))−FS​(A(S))∣]≤ε(m), and it on-average generalizes if ∣E[F(A(S))−FS(A(S))]∣≤ε(m)|\mathbb E[F(A(S))-F_S(A(S))]|\le\varepsilon(m)∣E[F(A(S))−FS​(A(S))]∣≤ε(m);
  • writing S∖iS^{\setminus i}S∖i for SSS with ziz_izi​ removed, AAA is LOO stable with rate ε\varepsilonε (Definition 29) if
1m∑i=1mES∼Dm[∣f(A(S∖i);zi)−f(A(S);zi)∣]≤ε(m).\frac1m\sum_{i=1}^m\mathbb E_{S\sim\mathcal D^m}\Big[\big|f(A(S^{\setminus i});z_i)-f(A(S);z_i)\big|\Big]\le\varepsilon(m).m1​i=1∑m​ES∼Dm​[​f(A(S∖i);zi​)−f(A(S);zi​)​]≤ε(m).

A rate is a sequence ε(m)\varepsilon(m)ε(m) that is non-increasing and tends to 000. A property holds universally if it holds under every D\mathcal DD with one and the same rate.

Formalization targets

Goal: Theorem 31

For an ERM AAA:

A universally LOO stable  ⟺  A universally consistent  ⟺  A universally generalizes.A\ \text{universally LOO stable}\iff A\ \text{universally consistent}\iff A\ \text{universally generalizes}.A universally LOO stable⟺A universally consistent⟺A universally generalizes.

The statement has no rates. Each side asserts that some rate exists and serves every distribution.

Milestones

In the order the proof uses them:

  1. Utility Lemma 12: E∣X−EX∣≤B/m\mathbb E|X-\mathbb EX|\le B/\sqrt mE∣X−EX∣≤B/m​ for the mean XXX of mmm i.i.d. variables bounded by BBB.
  2. Utility Lemma 13: X≤YX\le YX≤Y a.s. implies E∣X∣≤∣EX∣+2E∣Y∣\mathbb E|X|\le|\mathbb EX|+2\mathbb E|Y|E∣X∣≤∣EX∣+2E∣Y∣.
  3. Lemma 14: an AERM that on-average generalizes with rate εoag\varepsilon_{\mathrm{oag}}εoag​ generalizes with rate εoag+2εerm+2B/m\varepsilon_{\mathrm{oag}}+2\varepsilon_{\mathrm{erm}}+2B/\sqrt mεoag​+2εerm​+2B/m​.
  4. Lemma 15: under the same hypotheses the rule is consistent with rate εoag+εerm\varepsilon_{\mathrm{oag}}+\varepsilon_{\mathrm{erm}}εoag​+εerm​.
  5. Lemma 16 (Main Converse Lemma): in a learnable problem, E∣FS(h^S)−F∗∣≤2εcons(m′)+2B/m+2Bm′2/m\mathbb E|F_S(\hat h_S)-F^*|\le2\varepsilon_{\mathrm{cons}}(m')+2B/\sqrt m+2Bm'^2/mE∣FS​(h^S​)−F∗∣≤2εcons​(m′)+2B/m​+2Bm′2/m for 2≤m′≤m/22\le m'\le m/22≤m′≤m/2.
  6. Lemma 17: Eq. (12), together with an AERM that is consistent, gives generalization with rate εemp+εerm+εcons\varepsilon_{\mathrm{emp}}+\varepsilon_{\mathrm{erm}}+\varepsilon_{\mathrm{cons}}εemp​+εerm​+εcons​.
  7. First display of the proof of Theorem 31: a generalizing ERM is LOO stable with rate εgen(m−1)\varepsilon_{\mathrm{gen}}(m-1)εgen​(m−1).
  8. Second display: a LOO stable ERM on-average generalizes on samples of size m−1m-1m−1 with rate εstable(m)+2B/m\varepsilon_{\mathrm{stable}}(m)+2B/mεstable​(m)+2B/m.

Significance

Theorem 31 shows that for exact ERMs, LOO stability is not only sufficient but necessary for consistency. It transfers the supervised-learning characterisation of Mukherjee et al. to the General Learning Setting, where uniform convergence is no longer available as an intermediate. The hypothesis is sharp in one direction: Example 6 of the paper gives a universally consistent AERM that is not LOO stable. The equivalence therefore depends on exact minimisation, and an asymptotic minimiser does not suffice.

The lemmas are useful on their own. Lemmas 14 and 15 are the standard bridges between on-average generalization, generalization and consistency. Lemma 16 says that the minimal empirical risk estimates F∗F^*F∗ consistently in every learnable problem, even when no ERM learns. Lemma 16 also underlies Theorem 7, the paper's main characterisation of learnability, which is the subject of mission I of this series.

The results are proved in the paper. As far as is known they have not been formalized in any proof assistant. The platform has the textbook side of this framework (Shalev-Shwartz and Ben-David, Understanding Machine Learning, Chapter 13). Those statements are in Rd\mathbb R^dRd and use a replace-one stability notion; they do not cover leave-one-out stability or the General Learning Setting.

Difficulty

The implications between stability and generalization change the sample size: S∖iS^{\setminus i}S∖i has m−1m-1m−1 points, so a statement about the rule at size mmm has to be compared with the rule at size m−1m-1m−1 under the marginal law of the reduced sample. The step from universal consistency to generalization needs Lemma 16. That lemma estimates F∗F^*F∗ from a sample on which the ERM itself may be inconsistent, and it is the only place where universality of the consistency rate is used. Per-distribution consistency of an ERM does not imply generalization (Example 1 of the paper). An argument that fixes D\mathcal DD throughout therefore cannot succeed.

Formalization scope

Samples are tuples S:Fin m→ZS:\mathrm{Fin}\,m\to\mathcal ZS:Finm→Z with law Measure.pi (the i.i.d. product), and S∖iS^{\setminus i}S∖i is Fin.removeNth i S. A learning rule is a family Am:Zm→HA_m:\mathcal Z^m\to\mathcal HAm​:Zm→H. Its value at m=0m=0m=0 is never used, and LOO stability is required only for m≥2m\ge2m≥2. FS(h^S)F_S(\hat h_S)FS​(h^S​) is the infimum inf⁡hFS(h)\inf_hF_S(h)infh​FS​(h) and no minimiser is chosen. An ERM is a rule attaining this infimum at every sample. A rate is non-increasing on m≥1m\ge1m≥1 and tends to 000. Universal properties are stated as "there exists ε\varepsilonε with IsRate ε such that for every probability measure D\mathcal DD …", with the rate chosen before the distribution. The bound BBB is any bound on ∣f∣|f|∣f∣; the paper's BBB is sup⁡∣f∣\sup|f|sup∣f∣, and all its rates increase with BBB.

Measurability is not discussed in the paper. The formalization assumes that each f(h;⋅)f(h;\cdot)f(h;⋅) is measurable and that the rule is measurable in the sense that (S,z)↦f(Am(S);z)(S,z)\mapsto f(A_m(S);z)(S,z)↦f(Am​(S);z) is jointly measurable. Lemmas 14–17 also assume that S↦inf⁡hFS(h)S\mapsto\inf_hF_S(h)S↦infh​FS​(h) is measurable. For a measurable ERM this holds automatically, so Theorem 31 makes no such assumption. Without these assumptions Lean's integral of a non-measurable function is 000 and every rate bound would hold trivially. For the same reason Utility Lemma 13 assumes X,YX,YX,Y integrable. The ERM hypothesis of the goal must not be weakened to an AERM: the statement would then be false (Example 6).

One statement corrects the printed text. In the second display of the proof of Theorem 31 (p. 2668), the chain adds 2B/m2B/m2B/m and then drops it. The milestone states the bound the argument proves, εstable(m)+2B/m\varepsilon_{\mathrm{stable}}(m)+2B/mεstable​(m)+2B/m. The rate-free Theorem 31 is unaffected.

The development needs product measures, independence and variance bounds, all available in Mathlib, and the marginals of Measure.pi under removal of a coordinate. The definitions of risks, rules and stability notions can be reused by other stability results. Proofs of any milestone are welcome, and so are alternative proofs of the goal.

Selected references

  • S. Shalev-Shwartz, O. Shamir, N. Srebro, K. Sridharan, Learnability, Stability and Uniform Convergence, Journal of Machine Learning Research 11 (2010) 2635–2670. https://jmlr.org/papers/v11/shalev-shwartz10a.html
  • O. Bousquet, A. Elisseeff, Stability and Generalization, Journal of Machine Learning Research 2 (2002) 499–526. https://www.jmlr.org/papers/v2/bousquet02a.html
  • S. Mukherjee, P. Niyogi, T. Poggio, R. Rifkin, Learning theory: stability is sufficient for generalization and necessary and sufficient for consistency of empirical risk minimization, Advances in Computational Mathematics 25 (2006) 161–193. https://doi.org/10.1007/s10444-004-7634-z
  • S. Kutin, P. Niyogi, Almost-everywhere algorithmic stability and generalization error, UAI 2002. https://arxiv.org/abs/1301.0579
  • L. Devroye, T. Wagner, Distribution-free performance bounds for potential function rules, IEEE Transactions on Information Theory 25(5) (1979) 601–604. https://doi.org/10.1109/TIT.1979.1056087
11 thms4 active usersReviewed
🏆Completed
Convex OptimizationMachine LearningOptimization+1·Captain: mikedeng1

Learnability, Stability and Uniform Convergence II: Tikhonov-Regularized ERM Learns Convex Lipschitz Stochastic Optimization in Hilbert Space with High ProbabilityResearch Paper

Motivation

Statistical learning theory asks when a rule that sees only an i.i.d. sample z1,…,zmz_1,\dots,z_mz1​,…,zm​ from an unknown distribution DDD can return a hypothesis whose expected loss is close to the best possible. In supervised classification the classical answer is uniform convergence: learnability holds exactly when empirical risks converge to expected risks uniformly over the hypothesis class, and then empirical risk minimization (ERM) learns. Shalev-Shwartz, Shamir, Srebro and Sridharan (JMLR 11, 2010) showed that in Vapnik's broader General Learning Setting this picture breaks down. Their motivating example is stochastic convex optimization in a Hilbert space: minimizing an expected convex, Lipschitz objective over a bounded convex set from samples. This problem underlies regularized linear prediction, kernel methods and online-to-batch conversions, and the paper shows (§4.1) that in infinite dimension uniform convergence can fail and the plain empirical minimizer can fail to converge, while the problem is still learnable.

This mission formalizes the positive half of that example: Tikhonov-regularized ERM learns every such problem, with an explicit bound holding with probability 1−δ1-\delta1−δ (Theorem 3, p. 2644), through the stability of strongly convex empirical minimization (Theorem 2).

Setting

Let ZZZ be a measurable space of instances and EEE a real Hilbert space. A stochastic convex optimization problem consists of a nonempty, closed, convex, bounded set H⊆E\mathcal H\subseteq EH⊆E and an objective f:E×Z→Rf:E\times Z\to\mathbb Rf:E×Z→R such that for every zzz the map h↦f(h;z)h\mapsto f(h;z)h↦f(h;z) is convex and LLL-Lipschitz on H\mathcal HH, each f(h;⋅)f(h;\cdot)f(h;⋅) is measurable, and ∣f(h;z)∣≤C|f(h;z)|\le C∣f(h;z)∣≤C on H×Z\mathcal H\times ZH×Z. For a distribution DDD on ZZZ define the risk and optimal risk

F(h)=Ez∼D[f(h;z)],F∗=inf⁡h∈HF(h),F(h)=\mathbb E_{z\sim D}[f(h;z)],\qquad F^*=\inf_{h\in\mathcal H}F(h),F(h)=Ez∼D​[f(h;z)],F∗=h∈Hinf​F(h),

and for a sample S=(z1,…,zm)∼DmS=(z_1,\dots,z_m)\sim D^mS=(z1​,…,zm​)∼Dm the empirical risk FS(h)=1m∑i=1mf(h;zi)F_S(h)=\frac1m\sum_{i=1}^m f(h;z_i)FS​(h)=m1​∑i=1m​f(h;zi​). A function ggg is λ\lambdaλ-strongly convex on H\mathcal HH if g−λ2∥⋅∥2g-\frac\lambda2\|\cdot\|^2g−2λ​∥⋅∥2 is convex there. The regularized empirical minimizer is

h^λ∈arg min⁡h∈H(FS(h)+λ2∥h∥2).(5)\hat h_\lambda\in\operatorname*{arg\,min}_{h\in\mathcal H}\Big(F_S(h)+\frac\lambda2\|h\|^2\Big).\tag{5}h^λ​∈h∈Hargmin​(FS​(h)+2λ​∥h∥2).(5)

For the general part, a learning rule AAA maps samples to hypotheses; it is an AERM with rate εerm\varepsilon_{\mathrm{erm}}εerm​ if E[FS(A(S))−inf⁡hFS(h)]≤εerm(m)\mathbb E[F_S(A(S))-\inf_hF_S(h)]\le\varepsilon_{\mathrm{erm}}(m)E[FS​(A(S))−infh​FS​(h)]≤εerm​(m), consistent with rate εcons\varepsilon_{\mathrm{cons}}εcons​ if E[F(A(S))−F∗]≤εcons(m)\mathbb E[F(A(S))-F^*]\le\varepsilon_{\mathrm{cons}}(m)E[F(A(S))−F∗]≤εcons​(m), and uniform-RO stable with rate εstable\varepsilon_{\mathrm{stable}}εstable​ if replacing any one sample point changes the loss at any test point by at most εstable(m)\varepsilon_{\mathrm{stable}}(m)εstable​(m) on average over the replaced index (Definition 4).

Formalization targets

Goal: Theorem 3

If ∥h∥≤B\|h\|\le B∥h∥≤B on H\mathcal HH, L,B>0L,B>0L,B>0, δ∈(0,1)\delta\in(0,1)δ∈(0,1), m≥1m\ge1m≥1 and λ=16L2/(δB2m)\lambda=\sqrt{16L^2/(\delta B^2m)}λ=16L2/(δB2m)​, then with probability at least 1−δ1-\delta1−δ over S∼DmS\sim D^mS∼Dm

F(h^λ)−F∗ ≤ 4L2B2δm(1+8δm).F(\hat h_\lambda)-F^*\ \le\ 4\sqrt{\frac{L^2B^2}{\delta m}}\Big(1+\frac8{\delta m}\Big).F(h^λ​)−F∗ ≤ 4δmL2B2​​(1+δm8​).

The constants are the paper's.

Milestones, in the order the proof uses them

  1. Quadratic growth at a minimizer of a λ\lambdaλ-strongly convex ggg: g(h′)−g(h)≥λ2∥h′−h∥2g(h')-g(h)\ge\frac\lambda2\|h'-h\|^2g(h′)−g(h)≥2λ​∥h′−h∥2 (§4.2, p. 2644).
  2. Eq. (6): if f(⋅;z)f(\cdot;z)f(⋅;z) is λ\lambdaλ-strongly convex and LLL-Lipschitz, empirical minimizers of SSS and of S(i)S^{(i)}S(i) satisfy ∣f(h^S,z)−f(h^S(i),z)∣≤4L2/(λm)|f(\hat h_S,z)-f(\hat h_S^{(i)},z)|\le 4L^2/(\lambda m)∣f(h^S​,z)−f(h^S(i)​,z)∣≤4L2/(λm) for all zzz (p. 2645).
  3. Theorem 8: a uniform- or average-RO stable AERM is consistent with rate εstable+εerm\varepsilon_{\mathrm{stable}}+\varepsilon_{\mathrm{erm}}εstable​+εerm​ and generalizes with rate εstable+2εerm+2C/m\varepsilon_{\mathrm{stable}}+2\varepsilon_{\mathrm{erm}}+2C/\sqrt mεstable​+2εerm​+2C/m​ (p. 2649).
  4. ES∼Dm[F(h^S)−F∗]≤4L2/(λm)\mathbb E_{S\sim D^m}[F(\hat h_S)-F^*]\le 4L^2/(\lambda m)ES∼Dm​[F(h^S​)−F∗]≤4L2/(λm) for the strongly convex empirical minimizer (p. 2645).
  5. Theorem 2: with probability 1−δ1-\delta1−δ, F(h^S)−F∗≤4L2/(δλm)F(\hat h_S)-F^*\le 4L^2/(\delta\lambda m)F(h^S​)−F∗≤4L2/(δλm) (p. 2644).
  6. Theorem 2 applied to r(h;z)=λ2∥h∥2+f(h;z)r(h;z)=\frac\lambda2\|h\|^2+f(h;z)r(h;z)=2λ​∥h∥2+f(h;z): with probability 1−δ1-\delta1−δ, λ2∥h^λ∥2+F(h^λ)≤inf⁡h(λ2∥h∥2+F(h))+4(L+λB)2/(δλm)\frac\lambda2\|\hat h_\lambda\|^2+F(\hat h_\lambda)\le\inf_h\big(\frac\lambda2\|h\|^2+F(h)\big)+4(L+\lambda B)^2/(\delta\lambda m)2λ​∥h^λ​∥2+F(h^λ​)≤infh​(2λ​∥h∥2+F(h))+4(L+λB)2/(δλm) (p. 2645).

Significance

The result. Theorem 3 shows that every convex, Lipschitz, bounded stochastic optimization problem over a bounded subset of a Hilbert space is learnable at rate O(LB/δm)O(LB/\sqrt{\delta m})O(LB/δm​), with no dimension dependence and no uniform convergence. Together with the counterexamples of §4.1 it separates learnability from uniform convergence and from ERM, and it motivates the paper's general characterization: a problem is learnable if and only if it admits a uniform-RO stable asymptotic empirical risk minimizer (Theorem 7). Theorem 8 is the sufficiency half of that characterization and is reused wherever stability arguments give generalization bounds.

Formalizing it. The results are proved in the paper; to our knowledge none has a machine-checked proof. The closest platform material is the textbook treatment in Understanding Machine Learning, chapter 13 (Shalev-Shwartz and Ben-David): Corollary 13.9 (UnderstandingML.convex_lipschitz_bounded_learnable), Corollary 13.6 (rlm_lipschitz_stable) and Lemma 13.5 (strongly_convex_lemma). Those are stated in Rd\mathbb R^dRd, bound the risk in expectation, use the regularizer λ∥w∥2\lambda\|w\|^2λ∥w∥2 over all of Rd\mathbb R^dRd, and have different constants; the present mission works in an arbitrary Hilbert space, over a constraint set H\mathcal HH, with high-probability bounds and the paper's constants. Its definitions of learning rules, AERM, consistency and replace-one stability in the General Learning Setting are reusable by the other missions of this series.

Difficulty

The obvious route, bounding sup⁡h∈H∣F(h)−FS(h)∣\sup_{h\in\mathcal H}|F(h)-F_S(h)|suph∈H​∣F(h)−FS​(h)∣, is unavailable: §4.1 exhibits problems of exactly this type in which that supremum stays bounded away from zero for every sample size. Any successful argument therefore has to rely on a property of the learning rule rather than of the class H\mathcal HH, and the plain empirical minimizer does not have it: §4.1 shows it can stay a constant away from F∗F^*F∗ at every sample size. A second difficulty is purely formal: the regularization parameter λ\lambdaλ depends on δ\deltaδ and mmm, so the regularized minimizer changes with them, and all expectations involve a data-dependent hypothesis in a possibly non-separable Hilbert space, where measurability is not automatic.

Formalization scope

Lean conventions, fixed for every item:

  • EEE is a real inner product space with CompleteSpace E, never assumed finite-dimensional; H\mathcal HH is Hset : Set E, and all infima, suprema, strong convexity and Lipschitz conditions are taken on Hset only. F∗F^*F∗ is ⨅ h : Hset, F h.
  • Samples are Fin m → Z, DmD^mDm is Measure.pi, S(i)S^{(i)}S(i) is Function.update S i z', and m≥1m\ge1m≥1 throughout.
  • The paper's standing loss bound ∣f∣≤B|f|\le B∣f∣≤B (p. 2637) is named CCC, because Theorem 3 uses BBB for the norm bound ∥h∥≤B\|h\|\le B∥h∥≤B. L>0L>0L>0 and B>0B>0B>0 are implicit in Theorem 3's choice of λ\lambdaλ and are stated.
  • Strong convexity is Mathlib's StrongConvexOn Hset λ, which is the paper's definition.
  • Minimizers are selections S↦h^S∈HS\mapsto\hat h_S\in\mathcal HS↦h^S​∈H satisfying the minimization property; the theorems hold for every such selection, hence for the minimizer, which is unique by strong convexity.
  • Measurability, not discussed in the paper, is the series' single standing convention: each f(h;⋅)f(h;\cdot)f(h;⋅) is measurable and the selection makes (S,z)↦f(h^S;z)(S,z)\mapsto f(\hat h_S;z)(S,z)↦f(h^S​;z) jointly measurable; for Theorem 8, the rule is measurable in the same sense and S↦inf⁡hFS(h)S\mapsto\inf_hF_S(h)S↦infh​FS​(h) is measurable.
  • "With probability at least 1−δ1-\delta1−δ" is the bound Dm{failure}≤δD^m\{\text{failure}\}\le\deltaDm{failure}≤δ with 0<δ<10<\delta<10<δ<1.

No statement of the paper is corrected: all printed constants were checked against the proofs and are reproduced exactly.

A formalization in which the expected excess risk is a Bochner integral of a non-measurable or non-integrable function, or in which F∗F^*F∗ is an infimum over all of EEE or over an unbounded family, would make the bounds trivially true; the measurability hypotheses, the bound ∣f∣≤C|f|\le C∣f∣≤C and the infimum over the nonempty set H\mathcal HH rule this out.

Contributions welcome: proofs of the milestones in order, and in particular a reusable replace-one identity E[FS(A(S))]=1m∑iE[f(A(S(i));zi′)]\mathbb E[F_S(A(S))]=\frac1m\sum_i\mathbb E[f(A(S^{(i)});z'_i)]E[FS​(A(S))]=m1​∑i​E[f(A(S(i));zi′​)] under Measure.pi, and Markov's inequality in the form used for high-probability bounds.

Selected references

  • S. Shalev-Shwartz, O. Shamir, N. Srebro, K. Sridharan, Learnability, Stability and Uniform Convergence, Journal of Machine Learning Research 11 (2010) 2635–2670. https://jmlr.org/papers/v11/shalev-shwartz10a.html
  • S. Shalev-Shwartz, O. Shamir, N. Srebro, K. Sridharan, Stochastic Convex Optimization, COLT 2009. https://www.cs.mcgill.ca/~colt2009/papers/018.pdf
  • O. Bousquet, A. Elisseeff, Stability and Generalization, Journal of Machine Learning Research 2 (2002) 499–526. https://jmlr.org/papers/v2/bousquet02a.html
  • S. Shalev-Shwartz, S. Ben-David, Understanding Machine Learning: From Theory to Algorithms, Cambridge University Press, 2014, chapter 13. https://doi.org/10.1017/CBO9781107298019
  • V. N. Vapnik, Statistical Learning Theory, Wiley, 1998.
9 thms4 active usersReviewed
🏆Completed
Operations ResearchTheoretical Computer Science·Captain: mikedeng1

Competitive Paging Algorithms I: The Marking Algorithm Is 2H_k-CompetitiveResearch Paper

Motivation

Paging is the problem of managing a two-level memory: a fast cache holds kkk pages out of an address space of nnn pages, requests to pages arrive one at a time, and a request to a page outside the cache (a page fault) forces the algorithm to bring that page in and, when the cache is full, to evict another. The cost is the number of faults. An on-line algorithm decides which page to evict without knowing future requests. The comparison of paging policies with the optimal off-line policy is where competitive analysis began.

Sleator and Tarjan showed that LRU and FIFO are within a factor kkk of the off-line optimum and that no deterministic on-line algorithm does better than kkk (Sleator–Tarjan 1985). Randomization changes the picture: Fiat, Karp, Luby, McGeoch, Sleator and Young introduced the marking algorithm and proved that its expected cost is within a factor 2Hk2H_k2Hk​ of the optimum, where Hk=1+12+⋯+1k≈ln⁡kH_k = 1 + \frac12 + \dots + \frac1k \approx \ln kHk​=1+21​+⋯+k1​≈lnk (arXiv:cs/0205038).

Timeline.

  • 1985: Sleator and Tarjan: LRU and FIFO are kkk-competitive; no deterministic algorithm beats kkk.
  • 1988: Karlin, Manasse, Rudolph and Sleator coin "competitive" and analyse flush-when-full (Algorithmica 3).
  • 1990: Manasse, McGeoch and Sleator introduce the kkk-server problem and define competitiveness for randomized algorithms (J. Algorithms 11).
  • 1991: Fiat et al.: the marking algorithm is 2Hk2H_k2Hk​-competitive, and Hn−1H_{n-1}Hn−1​-competitive when k=n−1k = n-1k=n−1; no randomized paging algorithm beats HkH_kHk​.
  • 1991: McGeoch and Sleator give an HkH_kHk​-competitive randomized paging algorithm (Algorithmica 6).
  • 2000: Achlioptas, Chrobak and Noga determine the exact competitive ratio of the marking algorithm, 2Hk−12H_k - 12Hk​−1 (Theoret. Comput. Sci. 234).

Setting

The paper works in the uniform kkk-server problem, which is isomorphic to paging. There is a set MMM of nnn vertices, enumerated e(0),…,e(n−1)e(0), \dots, e(n-1)e(0),…,e(n−1), and moving a server between two distinct vertices costs 111. There are kkk servers, 1≤k≤n1 \le k \le n1≤k≤n. A request is a vertex, and after each request some server must be on it. Cached pages are covered vertices; a fault is a server move.

The marking algorithm starts with its servers on e(0),…,e(k−1)e(0), \dots, e(k-1)e(0),…,e(k−1) and keeps a set of marked vertices, initially the covered ones. On a request to rrr:

  1. Marking. rrr is marked; the moment k+1k+1k+1 vertices are marked, all marks except the one on rrr are erased.
  2. Serving. If rrr is covered, nothing moves. Otherwise a server is chosen uniformly at random among the covered unmarked vertices and moved to rrr.

The marks are updated before the server is chosen. For a finite request sequence σ\sigmaσ, CM(σ)C_M(\sigma)CM​(σ) is the algorithm's expected number of server moves. OPT(σ)\mathrm{OPT}(\sigma)OPT(σ) is the least number of moves with which kkk servers, starting from the same configuration C0C_0C0​ and knowing σ\sigmaσ in advance, can serve σ\sigmaσ.

A randomized algorithm is ccc-competitive if there is a constant aaa such that CM(σ)≤c⋅CB(σ)+aC_M(\sigma) \le c \cdot C_B(\sigma) + aCM​(σ)≤c⋅CB​(σ)+a for every request sequence σ\sigmaσ and every algorithm BBB.

The marks divide σ\sigmaσ into phases. A new phase begins at the request that would make k+1k+1k+1 vertices marked. A vertex is clean in a phase if it was not requested in the previous phase and not yet in this one, and stale if it was requested in the previous phase but not yet in this one.

Formalization targets

Goal: Theorem 1

∃ a∈R  ∀σ:CM(σ)  ≤  2Hk⋅OPT(σ)+a.\exists\, a \in \mathbb R\ \ \forall \sigma:\qquad C_M(\sigma) \;\le\; 2H_k \cdot \mathrm{OPT}(\sigma) + a .∃a∈R  ∀σ:CM​(σ)≤2Hk​⋅OPT(σ)+a.

The constant aaa may depend on nnn, kkk and the enumeration, never on σ\sigmaσ.

Milestones (proof of Theorem 1, pp. 4–5)

  1. Without loss of generality the adversary is lazy: no move on a covered request, exactly one move otherwise (reference item, already proved on the platform).
  2. At the start of every phase the marked vertices are exactly the covered ones, and the first request of a phase is unmarked.
  3. In a phase with lll clean requests, a lazy adversary pays CA≥l−dC_A \ge l - dCA​≥l−d, where ddd counts its servers off the marking algorithm's servers at the start of the phase.
  4. It also pays CA≥d′C_A \ge d'CA​≥d′, where d′d'd′ counts its servers off the final marked set at the end of the phase.
  5. Hence CA≥max⁡(l−d,d′)≥12(l−d+d′)C_A \ge \max(l-d, d') \ge \tfrac12(l - d + d')CA​≥max(l−d,d′)≥21​(l−d+d′).
  6. A request to a stale vertex is a fault with probability c/sc/sc/s (ccc clean vertices requested so far, sss stale vertices left).
  7. The marking algorithm's expected cost in a phase is at most l(Hk−Hl+1)≤lHkl(H_k - H_l + 1) \le lH_kl(Hk​−Hl​+1)≤lHk​.

Companions

  • Theorem 2: for k=n−1k = n-1k=n−1, CM(σ)≤Hn−1⋅OPT(σ)+aC_M(\sigma) \le H_{n-1} \cdot \mathrm{OPT}(\sigma) + aCM​(σ)≤Hn−1​⋅OPT(σ)+a.
  • Tightness remark (pp. 5–6): for k=2k = 2k=2, n=4n = 4n=4 there is no aaa with CM(σ)≤H2⋅OPT(σ)+aC_M(\sigma) \le H_2 \cdot \mathrm{OPT}(\sigma) + aCM​(σ)≤H2​⋅OPT(σ)+a for all σ\sigmaσ.

Significance

The result. Theorem 1 was the first proof that randomization beats the deterministic barrier kkk for paging, bringing the ratio down to O(log⁡k)O(\log k)O(logk). Together with the paper's lower bound HkH_kHk​ for every randomized algorithm, it determines the randomized competitive ratio of paging up to a factor 222. Its phase and clean/stale accounting is reused throughout the analysis of randomized caching.

Formalizing it. The theorem is proved (1991). As far as is known it has no machine-checked proof. Formalizing it requires a probabilistic model of a randomized on-line algorithm, an off-line optimum, and a phase decomposition with an exchangeability argument, and these are the first such objects in this library. Theorem 2 and the k=2k = 2k=2, n=4n = 4n=4 example use the same definitions and also check that the formal algorithm is the paper's. The sharp ratio 2Hk−12H_k - 12Hk​−1 is a natural follow-up.

Difficulty

The comparison is between a random process and a deterministic adversary, and each side has its own obstacle.

On the algorithm's side, the configuration inside a phase is random, and the fault probability of a stale request depends on the whole history of the phase. The claim that the ccc uncovered stale vertices form a uniformly random subset of the sss stale ones is an exchangeability property of the process, and must be established from the step-by-step uniform choice. The worst-case ordering of the requests within a phase then has to be justified as a bound, not assumed.

On the adversary's side, the per-phase bound max⁡(l−d,d′)\max(l-d, d')max(l−d,d′) does not sum directly. The ddd and d′d'd′ terms telescope across phases only because the configuration of the marking algorithm at each phase boundary is deterministic. The first phase, which begins after an initial run of requests to e(0),…,e(k−1)e(0), \dots, e(k-1)e(0),…,e(k−1), and the last, incomplete phase have to be absorbed into the additive constant.

Formalization scope

The vertex set is an abstract metric space MMM with e:Fin n≃Me : \mathrm{Fin}\,n \simeq Me:Finn≃M and dist(x,y)=1\mathrm{dist}(x,y) = 1dist(x,y)=1 for x≠yx \ne yx=y. The natural metric ∣i−j∣|i - j|∣i−j∣ on Fin n\mathrm{Fin}\,nFinn is deliberately not used. The configurations and the off-line optimum OPT\mathrm{OPT}OPT are the published KServer definitions (KServer.Config, KServer.offlineCost), with OPT\mathrm{OPT}OPT taken from the marking algorithm's initial configuration. An off-line algorithm starting elsewhere changes the cost by at most kkk, which is absorbed into aaa.

The marking algorithm is a Markov chain on pairs (covered set, marked set). Each step is a PMF, with the eviction drawn by PMF.uniformOfFinset from the covered unmarked vertices. The expected cost is the sum over requests of the probability that the request is not covered, which is exact because the algorithm moves exactly one server per fault. Harmonic numbers are Mathlib's harmonic, cast to R\mathbb RR. Phases, clean counts and lazy off-line schedules are defined once, in the mission's definition file, and all milestones use them.

A trivializing formalization is ruled out as follows. The additive constant is quantified before σ\sigmaσ, so a per-sequence constant cannot be used. The comparison is with the optimum over all off-line schedules, not a particular one. The random choice is among the covered unmarked vertices, with marks updated first. The hypothesis 1≤k≤n1 \le k \le n1≤k≤n excludes the degenerate case k=0k = 0k=0, where H0=0H_0 = 0H0​=0.

Proofs of individual milestones are welcome. The laziness reduction for off-line schedules, the exchangeability lemma for the uniform eviction process, and the harmonic-sum identity ∑j=l+1kl/j=l(Hk−Hl)\sum_{j=l+1}^{k} l/j = l(H_k - H_l)∑j=l+1k​l/j=l(Hk​−Hl​) are reusable beyond this mission.

Selected references

  • A. Fiat, R. M. Karp, M. Luby, L. A. McGeoch, D. D. Sleator, N. E. Young, Competitive Paging Algorithms, J. Algorithms 12(4):685–699, 1991; arXiv:cs/0205038v1. https://arxiv.org/abs/cs/0205038
  • D. D. Sleator, R. E. Tarjan, Amortized Efficiency of List Update and Paging Rules, Comm. ACM 28(2):202–208, 1985. https://doi.org/10.1145/2786.2793
  • A. R. Karlin, M. S. Manasse, L. Rudolph, D. D. Sleator, Competitive Snoopy Caching, Algorithmica 3:79–119, 1988. https://doi.org/10.1007/BF01762111
  • M. S. Manasse, L. A. McGeoch, D. D. Sleator, Competitive Algorithms for Server Problems, J. Algorithms 11(2):208–230, 1990. https://doi.org/10.1016/0196-6774(90)90003-W
  • L. A. McGeoch, D. D. Sleator, A Strongly Competitive Randomized Paging Algorithm, Algorithmica 6:816–825, 1991. https://doi.org/10.1007/BF01759073
  • D. Achlioptas, M. Chrobak, J. Noga, Competitive Analysis of Randomized Paging Algorithms, Theoret. Comput. Sci. 234:203–218, 2000. https://doi.org/10.1016/S0304-3975(98)00116-9
10 thms4 active usersReviewed
🏆Completed
Machine LearningOptimizationStatistics·Captain: mikedeng1

Simultaneous Analysis of Lasso and Dantzig Selector V: Estimation, Prediction and Sparsity Bounds for the LassoResearch Paper

Motivation

In a linear regression with many more candidate variables than observations, least squares is not defined uniquely and does not estimate anything useful. The Lasso (Tibshirani, 1996) replaces it by an ℓ1\ell_1ℓ1​-penalised least-squares problem, which is convex, can be solved at scale, and returns sparse coefficient vectors. The question that a statistician, a signal-processing engineer or an operations researcher fitting a sparse model must answer before trusting it is quantitative: how far is the Lasso estimate from the true coefficient vector, how well does it predict, and how many variables does it select, as functions of the sample size nnn, the number of variables MMM and the sparsity sss of the truth?

Bickel, Ritov and Tsybakov (arXiv:0801.1095, Ann. Statist. 37(4), 2009) answered this under the restricted eigenvalue (RE) condition, which they introduced, with explicit constants and an explicit failure probability. Their Theorem 7.2, the goal of this mission, is a standard reference result of high-dimensional statistics and a model for the Lasso analyses in the textbooks of Bühlmann and van de Geer (2011) and Wainwright (2019).

Timeline, restricted to what each work proved:

  • 2007: Candès and Tao (arXiv:math/0506081) prove ℓ2\ell_2ℓ2​ bounds for the Dantzig selector under a uniform uncertainty principle.
  • 2007: Bunea, Tsybakov and Wegkamp (doi:10.1214/07-EJS008) prove sparsity oracle inequalities for the Lasso under mutual-coherence conditions; Lemma B.1 of the present paper is essentially their Lemma 1.
  • 2008/2009: Bickel, Ritov and Tsybakov prove Theorem 7.2 under RE(s,3)(s,3)(s,3) and RE(s,m,3)(s,m,3)(s,m,3), conditions weaker than those of the previous works.

Setting

A deterministic design matrix X∈Rn×MX\in\mathbb R^{n\times M}X∈Rn×M with columns x(1),…,x(M)x_{(1)},\dots,x_{(M)}x(1)​,…,x(M)​ is observed together with

y=Xβ∗+w,y=X\beta^*+w,y=Xβ∗+w,

where β∗∈RM\beta^*\in\mathbb R^Mβ∗∈RM is unknown and w=(W1,…,Wn)w=(W_1,\dots,W_n)w=(W1​,…,Wn​) has independent N(0,σ2)\mathcal N(0,\sigma^2)N(0,σ2) entries with σ>0\sigma>0σ>0. Throughout, n≥1n\ge1n≥1, M≥2M\ge2M≥2, and every diagonal entry of the Gram matrix Ψn=X⊤X/n\Psi_n=X^\top X/nΨn​=X⊤X/n equals 111.

For δ∈RM\delta\in\mathbb R^Mδ∈RM write ∣δ∣p=(∑j∣δj∣p)1/p|\delta|_p=(\sum_j|\delta_j|^p)^{1/p}∣δ∣p​=(∑j​∣δj​∣p)1/p, J(δ)={j:δj≠0}J(\delta)=\{j:\delta_j\ne0\}J(δ)={j:δj​=0} for the support, M(δ)=∣J(δ)∣\mathcal M(\delta)=|J(\delta)|M(δ)=∣J(δ)∣ for the sparsity, and δJ\delta_JδJ​ for the vector that agrees with δ\deltaδ on JJJ and vanishes off JJJ. The largest eigenvalue of Ψn\Psi_nΨn​ is ϕmax⁡\phi_{\max}ϕmax​.

The Lasso estimator with tuning parameter r>0r>0r>0 is any minimiser

β^L∈arg⁡min⁡β∈RM{1n∣y−Xβ∣22+2r∣β∣1}.\hat\beta_L\in\arg\min_{\beta\in\mathbb R^M}\Big\{\frac1n|y-X\beta|_2^2+2r|\beta|_1\Big\}.β^​L​∈argβ∈RMmin​{n1​∣y−Xβ∣22​+2r∣β∣1​}.

Minimisers exist but need not be unique.

Assumption RE(s,c0)(s,c_0)(s,c0​) (1≤s≤M1\le s\le M1≤s≤M, c0>0c_0>0c0​>0) asks that

κ(s,c0)=min⁡∣J0∣≤s min⁡δ≠0, ∣δJ0c∣1≤c0∣δJ0∣1∣Xδ∣2n ∣δJ0∣2>0.\kappa(s,c_0)=\min_{|J_0|\le s}\ \min_{\delta\ne0,\ |\delta_{J_0^c}|_1\le c_0|\delta_{J_0}|_1}\frac{|X\delta|_2}{\sqrt n\,|\delta_{J_0}|_2}>0 .κ(s,c0​)=∣J0​∣≤smin​ δ=0, ∣δJ0c​​∣1​≤c0​∣δJ0​​∣1​min​n​∣δJ0​​∣2​∣Xδ∣2​​>0.

Assumption RE(s,m,c0)(s,m,c_0)(s,m,c0​) (1≤s≤M/21\le s\le M/21≤s≤M/2, m≥sm\ge sm≥s, s+m≤Ms+m\le Ms+m≤M) is the same with ∣δJ01∣2|\delta_{J_{01}}|_2∣δJ01​​∣2​ in the denominator, where J01=J0∪J1J_{01}=J_0\cup J_1J01​=J0​∪J1​ and J1J_1J1​ collects the mmm largest in absolute value coordinates of δ\deltaδ outside J0J_0J0​.

Formalization targets

Goal: Theorem 7.2

Let M(β∗)≤s\mathcal M(\beta^*)\le sM(β∗)≤s, let RE(s,3)(s,3)(s,3) hold, and let r=Aσlog⁡M/nr=A\sigma\sqrt{\log M/n}r=AσlogM/n​ with A>22A>2\sqrt2A>22​. With probability at least 1−M1−A2/81-M^{1-A^2/8}1−M1−A2/8, every Lasso solution satisfies

∣β^L−β∗∣1≤16Aκ2(s,3)σslog⁡Mn,∣X(β^L−β∗)∣22≤16A2κ2(s,3)σ2slog⁡M,M(β^L)≤64ϕmax⁡κ2(s,3)s,|\hat\beta_L-\beta^*|_1\le\frac{16A}{\kappa^2(s,3)}\sigma s\sqrt{\frac{\log M}{n}},\qquad |X(\hat\beta_L-\beta^*)|_2^2\le\frac{16A^2}{\kappa^2(s,3)}\sigma^2s\log M,\qquad \mathcal M(\hat\beta_L)\le\frac{64\phi_{\max}}{\kappa^2(s,3)}s,∣β^​L​−β∗∣1​≤κ2(s,3)16A​σsnlogM​​,∣X(β^​L​−β∗)∣22​≤κ2(s,3)16A2​σ2slogM,M(β^​L​)≤κ2(s,3)64ϕmax​​s,

and, if RE(s,m,3)(s,m,3)(s,m,3) holds, on the same event and for all 1<p≤21<p\le21<p≤2,

∣β^L−β∗∣pp≤16{1+3sm}2(p−1)s(Aσκ2(s,m,3)log⁡Mn)p.|\hat\beta_L-\beta^*|_p^p\le16\Big\{1+3\sqrt{\tfrac sm}\Big\}^{2(p-1)}s\Big(\frac{A\sigma}{\kappa^2(s,m,3)}\sqrt{\frac{\log M}{n}}\Big)^p .∣β^​L​−β∗∣pp​≤16{1+3ms​​}2(p−1)s(κ2(s,m,3)Aσ​nlogM​​)p.

Milestones

In the order in which the paper's proof uses them:

  1. (B.4): the noise event A=⋂j{2∣1nx(j)⊤w∣≤r}\mathcal A=\bigcap_j\{2|\tfrac1n x_{(j)}^\top w|\le r\}A=⋂j​{2∣n1​x(j)⊤​w∣≤r} has P(Ac)≤M1−A2/8\mathbb P(\mathcal A^c)\le M^{1-A^2/8}P(Ac)≤M1−A2/8.
  2. (B.6): the optimality conditions of the Lasso.
  3. Lemma B.1 (Section 7 case): the basic inequality (B.1) for all β\betaβ, the residual bound (B.2) and the sparsity bound M(β^L)≤4ϕmax⁡∥fβ^L−f∥n2/r2\mathcal M(\hat\beta_L)\le4\phi_{\max}\|f_{\hat\beta_L}-f\|_n^2/r^2M(β^​L​)≤4ϕmax​∥fβ^​L​​−f∥n2​/r2 (B.3).
  4. Corollary B.2: the error δ=β^L−β\delta=\hat\beta_L-\betaδ=β^​L​−β lies in the cone ∣δJ0c∣1≤3∣δJ0∣1|\delta_{J_0^c}|_1\le3|\delta_{J_0}|_1∣δJ0c​​∣1​≤3∣δJ0​​∣1​.
  5. (B.30)–(B.31): on A\mathcal AA, 1n∣Xδ∣22≤16r2s/κ2\frac1n|X\delta|_2^2\le16r^2s/\kappa^2n1​∣Xδ∣22​≤16r2s/κ2 and ∣δJ0∣2≤4rs/κ2|\delta_{J_0}|_2\le4r\sqrt s/\kappa^2∣δJ0​​∣2​≤4rs​/κ2.
  6. (B.27) and (B.28) with c0=3c_0=3c0​=3: ℓ1\ell_1ℓ1​ and ℓ2\ell_2ℓ2​ norms of a cone vector.
  7. The ℓp\ell_pℓp​ interpolation ∑ajp≤b12−pb2p−1\sum a_j^p\le b_1^{2-p}b_2^{p-1}∑ajp​≤b12−p​b2p−1​.

Significance

The result. Theorem 7.2 gives, for fixed nnn and MMM rather than asymptotically, the rate slog⁡M/ns\log M/nslogM/n for the prediction loss and slog⁡M/ns\sqrt{\log M/n}slogM/n​ for the ℓ1\ell_1ℓ1​ loss of the Lasso, under a condition on the design only (RE), with no assumption on how MMM compares with nnn. The dependence on MMM is only logarithmic, which is what makes the Lasso usable when M≫nM\gg nM≫n. Bound (7.9) shows that the Lasso selects at most a constant multiple of sss variables, and (7.10) covers every ℓp\ell_pℓp​ loss between ℓ1\ell_1ℓ1​ and ℓ2\ell_2ℓ2​. Together with Theorem 7.1 for the Dantzig selector, the result shows that the two estimators have the same rates.

Formalizing it. The theorem is proved on paper. As far as a search of the platform shows, there is no machine-checked proof of a probabilistic Lasso rate. The closest platform statement, HighDimStat.SparseLinear.lasso_l2_error_bound (Wainwright, Theorem 7.13(a)), is deterministic, assumes a lower bound on the regularisation parameter in place of Gaussian noise, uses a restricted eigenvalue condition over the cone of one fixed support, and concludes an ℓ2\ell_2ℓ2​ bound with a different constant. A formal proof of Theorem 7.2 would supply the Gaussian maximal inequality, the Lasso optimality conditions, and the cone and interpolation inequalities as reusable lemmas.

Difficulty

Each step is short on paper, and none of the steps is deep. The main work is in three places. First, the probability: the event on which the deterministic argument runs involves all MMM correlations 1nx(j)⊤w\frac1n x_{(j)}^\top wn1​x(j)⊤​w at once, and its probability must be bounded by exactly M1−A2/8M^{1-A^2/8}M1−A2/8, which requires the law of a linear combination of independent Gaussians and a sharp Gaussian tail estimate, not a generic concentration bound with unspecified constants. Second, the Lasso is defined only through its minimising property, while the sparsity bound (7.9) is a statement about the number of non-zero coordinates of a minimiser of a non-differentiable objective; the characterisation (B.6) of minimisers is not in Mathlib. Third, (7.10) involves two restricted eigenvalue constants, a ranking of coordinates with possible ties, and real exponents, and every constant has to come out exactly.

The obvious idea of proving (7.7)–(7.10) for one fixed minimiser does not suffice: the statement quantifies over every minimiser on a single event.

Formalization scope

The design XXX is a Matrix (Fin n) (Fin M) ℝ; vectors are functions Fin M → ℝ. The noise is a family W : Fin n → Ω → ℝ of independent, measurable random variables with law gaussianReal 0 σ² on a probability space, and y(ω)=Xβ∗+W(ω)y(\omega)=X\beta^*+W(\omega)y(ω)=Xβ∗+W(ω). The probabilistic conclusion is one measurable event EEE with P(E)≥1−M1−A2/8\mathbb P(E)\ge1-M^{1-A^2/8}P(E)≥1−M1−A2/8 on which every minimiser of (7.2) satisfies all bounds; the event does not depend on the minimiser, on mmm or on ppp. log⁡\loglog is the natural logarithm.

RE(s,3)(s,3)(s,3) and RE(s,m,3)(s,m,3)(s,m,3) are stated through witnesses: a predicate "κn∣δJ0∣2≤∣Xδ∣2\kappa\sqrt n|\delta_{J_0}|_2\le|X\delta|_2κn​∣δJ0​​∣2​≤∣Xδ∣2​ for every admissible J0J_0J0​ and δ\deltaδ", and the theorem holds for every witness κ>0\kappa>0κ>0. Because the paper's κ(s,c0)\kappa(s,c_0)κ(s,c0​) is an attained minimum, every witness is at most it and the bounds decrease in κ\kappaκ, so this is equivalent to the printed statement. The assumption quantifies over every J0J_0J0​ with ∣J0∣≤s|J_0|\le s∣J0​∣≤s, as on page 7, not only over the support of β∗\beta^*β∗. ϕmax⁡\phi_{\max}ϕmax​ is the supremum of 1n∣Xx∣22\frac1n|Xx|_2^2n1​∣Xx∣22​ over unit vectors xxx. Lemma B.1 is stated in its Section-7 specialisation (unit column norms, f=Xβ∗f=X\beta^*f=Xβ∗), the form used in the proof of Theorem 7.2; (B.28) is stated for every c0>0c_0>0c0​>0 and (B.27) likewise, since the paper writes them with c0=1c_0=1c0​=1 and invokes them with c0=3c_0=3c0​=3. The printed Theorem 7.2 needs no correction; all four constants were checked against the proof.

A formalization in which the noise is not Gaussian, the Lasso predicate can be vacuous, the RE condition is imposed only on the support of β∗\beta^*β∗, or the probability is that of a non-measurable set, is not this theorem and is ruled out by the statement.

Needed infrastructure: Gaussian tail bounds and the law of a linear combination of independent Gaussians (largely in Mathlib), subdifferential calculus for ℓ1\ell_1ℓ1​-penalised least squares, and elementary finite-sum inequalities. The cone inequalities (B.27)–(B.28), the interpolation inequality and the optimality conditions (B.6) are reusable in other sparse-estimation missions; contributions to any milestone are welcome.

Selected references

  • P. J. Bickel, Y. Ritov, A. B. Tsybakov, Simultaneous analysis of Lasso and Dantzig selector, Ann. Statist. 37(4), 1705–1732, 2009. arXiv:0801.1095v3, https://arxiv.org/abs/0801.1095 ; https://doi.org/10.1214/08-AOS620
  • F. Bunea, A. B. Tsybakov, M. H. Wegkamp, Sparsity oracle inequalities for the Lasso, Electron. J. Statist. 1, 169–194, 2007. https://doi.org/10.1214/07-EJS008
  • E. Candès, T. Tao, The Dantzig selector: statistical estimation when p is much larger than n, Ann. Statist. 35(6), 2313–2351, 2007. https://arxiv.org/abs/math/0506081
  • R. Tibshirani, Regression shrinkage and selection via the lasso, J. R. Statist. Soc. B 58(1), 267–288, 1996. https://doi.org/10.1111/j.2517-6161.1996.tb02080.x
  • M. J. Wainwright, High-Dimensional Statistics: A Non-Asymptotic Viewpoint, Cambridge University Press, 2019, Chapter 7. https://doi.org/10.1017/9781108627771
10 thms4 active usersReviewed
🏆Completed
Machine LearningOptimizationStatistics·Captain: mikedeng1

Simultaneous Analysis of Lasso and Dantzig Selector IV: Estimation and Prediction Error Bounds for the Dantzig SelectorResearch Paper

Motivation

In high-dimensional linear regression the number of unknown coefficients MMM may be much larger than the number of observations nnn, and the coefficient vector can only be recovered because it is assumed to be sparse: few of its entries are non-zero. Two convex estimators dominate this setting: the Lasso of Tibshirani (1996), an ℓ1\ell_1ℓ1​-penalized least-squares estimator, and the Dantzig selector of Candès and Tao (2007), which minimizes the ℓ1\ell_1ℓ1​ norm subject to a bound on the correlation between the residual and the columns of the design. Both are used routinely in statistics, signal processing and machine learning, and their rates of convergence determine how many observations suffice to estimate a sparse vector.

Bickel, Ritov and Tsybakov (arXiv:0801.1095; Ann. Statist. 37(4), 2009) analysed the two estimators side by side under a single, weak condition on the design, the restricted eigenvalue (RE) assumption. This mission formalizes their rates for the Dantzig selector, Theorem 7.1 of the paper.

Timeline. Candès and Tao (Ann. Statist. 35, 2007) introduced the Dantzig selector and bounded its ℓ2\ell_2ℓ2​ error under a uniform uncertainty principle on the design. Bickel, Ritov and Tsybakov (2009) replaced that condition by the RE assumptions, which are implied by it (their Lemma 4.1), and obtained ℓp\ell_pℓp​ bounds for every 1≤p≤21\le p\le21≤p≤2 and a prediction bound, with explicit constants. Later work (van de Geer and Bühlmann, EJS 2009) compared RE with the compatibility condition and other design conditions.

Setting

Observations follow the linear model

y=Xβ∗+w,y=X\beta^*+w,y=Xβ∗+w,

where X∈Rn×MX\in\mathbb R^{n\times M}X∈Rn×M is a deterministic design matrix, n≥1n\ge1n≥1, M≥2M\ge2M≥2, β∗∈RM\beta^*\in\mathbb R^Mβ∗∈RM is unknown, and w=(W1,…,Wn)w=(W_1,\dots,W_n)w=(W1​,…,Wn​) has independent N(0,σ2)\mathcal N(0,\sigma^2)N(0,σ2) coordinates with σ>0\sigma>0σ>0. The columns are normalized: every diagonal element of the Gram matrix XTX/nX^TX/nXTX/n equals 1.

For β∈RM\beta\in\mathbb R^Mβ∈RM, J(β)={j:βj≠0}J(\beta)=\{j:\beta_j\ne0\}J(β)={j:βj​=0} is its support and M(β)=∣J(β)∣\mathcal M(\beta)=|J(\beta)|M(β)=∣J(β)∣ its sparsity; β∗\beta^*β∗ satisfies M(β∗)≤s\mathcal M(\beta^*)\le sM(β∗)≤s for an integer 1≤s≤M1\le s\le M1≤s≤M. Norms are ∣δ∣p=(∑j∣δj∣p)1/p|\delta|_p=(\sum_j|\delta_j|^p)^{1/p}∣δ∣p​=(∑j​∣δj​∣p)1/p and ∣v∣22=∑ivi2|v|_2^2=\sum_iv_i^2∣v∣22​=∑i​vi2​; for an index set JJJ, δJ\delta_JδJ​ keeps the coordinates of δ\deltaδ in JJJ and sets the others to 0, and JcJ^cJc is the complement of JJJ.

With a tuning level r=Aσlog⁡M/nr=A\sigma\sqrt{\log M/n}r=AσlogM/n​, A>2A>\sqrt2A>2​, the Dantzig selector is any minimizer

β^D∈arg⁡min⁡β∈Λ∣β∣1,Λ={β∈RM: ∣1nXT(y−Xβ)∣∞≤r}.\hat\beta_D\in\arg\min_{\beta\in\Lambda}|\beta|_1,\qquad \Lambda=\Big\{\beta\in\mathbb R^M:\ \Big|\tfrac1nX^T(y-X\beta)\Big|_\infty\le r\Big\}.β^​D​∈argβ∈Λmin​∣β∣1​,Λ={β∈RM: ​n1​XT(y−Xβ)​∞​≤r}.

The cone condition at an index set J0J_0J0​ with constant c0>0c_0>0c0​>0 is ∣δJ0c∣1≤c0∣δJ0∣1|\delta_{J_0^c}|_1\le c_0|\delta_{J_0}|_1∣δJ0c​​∣1​≤c0​∣δJ0​​∣1​. Assumption RE(s,c0)(s,c_0)(s,c0​) asks that

κ(s,c0)=min⁡∣J0∣≤s min⁡δ≠0, ∣δJ0c∣1≤c0∣δJ0∣1∣Xδ∣2n ∣δJ0∣2>0.\kappa(s,c_0)=\min_{|J_0|\le s}\ \min_{\delta\ne0,\ |\delta_{J_0^c}|_1\le c_0|\delta_{J_0}|_1}\frac{|X\delta|_2}{\sqrt n\,|\delta_{J_0}|_2}>0 .κ(s,c0​)=∣J0​∣≤smin​ δ=0, ∣δJ0c​​∣1​≤c0​∣δJ0​​∣1​min​n​∣δJ0​​∣2​∣Xδ∣2​​>0.

Assumption RE(s,m,c0)(s,m,c_0)(s,m,c0​) is the same with ∣δJ01∣2|\delta_{J_{01}}|_2∣δJ01​​∣2​ in the denominator, where J01=J0∪J1J_{01}=J_0\cup J_1J01​=J0​∪J1​ and J1J_1J1​ collects the mmm largest ∣δj∣|\delta_j|∣δj​∣ outside J0J_0J0​; it is used for s≤ms\le ms≤m, s+m≤Ms+m\le Ms+m≤M.

Formalization targets

Goal: Theorem 7.1

With probability at least 1−M1−A2/21-M^{1-A^2/2}1−M1−A2/2, every Dantzig selector satisfies

∣β^D−β∗∣1≤8Aκ2(s,1) σslog⁡Mn,∣X(β^D−β∗)∣22≤16A2κ2(s,1) σ2slog⁡M,|\hat\beta_D-\beta^*|_1\le\frac{8A}{\kappa^2(s,1)}\,\sigma s\sqrt{\frac{\log M}{n}},\qquad |X(\hat\beta_D-\beta^*)|_2^2\le\frac{16A^2}{\kappa^2(s,1)}\,\sigma^2s\log M,∣β^​D​−β∗∣1​≤κ2(s,1)8A​σsnlogM​​,∣X(β^​D​−β∗)∣22​≤κ2(s,1)16A2​σ2slogM,

and, on the same event, if RE(s,m,1)(s,m,1)(s,m,1) holds, simultaneously for all 1<p≤21<p\le21<p≤2,

∣β^D−β∗∣pp≤2p−1 8{1+sm}2(p−1)s(Aσκ2(s,m,1)log⁡Mn)p.|\hat\beta_D-\beta^*|_p^p\le2^{p-1}\,8\Big\{1+\sqrt{\tfrac sm}\Big\}^{2(p-1)}s\Big(\frac{A\sigma}{\kappa^2(s,m,1)}\sqrt{\frac{\log M}{n}}\Big)^p .∣β^​D​−β∗∣pp​≤2p−18{1+ms​​}2(p−1)s(κ2(s,m,1)Aσ​nlogM​​)p.

Milestones

In the order the proof of the paper uses them:

  1. The noise event B=⋂j{∣1n∑iXijWi∣≤r∥fj∥n}\mathcal B=\bigcap_j\{|\frac1n\sum_iX_{ij}W_i|\le r\|f_j\|_n\}B=⋂j​{∣n1​∑i​Xij​Wi​∣≤r∥fj​∥n​} has P{Bc}≤M1−A2/2\mathbb P\{\mathcal B^c\}\le M^{1-A^2/2}P{Bc}≤M1−A2/2 (proof of Lemma B.3).
  2. Lemma B.3, (B.9): for any β\betaβ satisfying the Dantzig constraint, δ=β^D−β\delta=\hat\beta_D-\betaδ=β^​D​−β satisfies the cone condition at J(β)J(\beta)J(β) with c0=1c_0=1c0​=1.
  3. (B.25): on B\mathcal BB, β∗∈Λ\beta^*\in\Lambdaβ∗∈Λ, 1n∣XTXδ∣∞≤2r\frac1n|X^TX\delta|_\infty\le2rn1​∣XTXδ∣∞​≤2r, and 1n∣Xδ∣22≤4rs ∣δJ0∣2\frac1n|X\delta|_2^2\le4r\sqrt s\,|\delta_{J_0}|_2n1​∣Xδ∣22​≤4rs​∣δJ0​​∣2​.
  4. (B.26): under RE(s,1)(s,1)(s,1), 1n∣Xδ∣22≤16r2s/κ2\frac1n|X\delta|_2^2\le16r^2s/\kappa^2n1​∣Xδ∣22​≤16r2s/κ2 and ∣δJ0∣2≤4rs/κ2|\delta_{J_0}|_2\le4r\sqrt s/\kappa^2∣δJ0​​∣2​≤4rs​/κ2.
  5. (B.27): on the cone, ∣δ∣1≤(1+c0)s ∣δJ0∣2|\delta|_1\le(1+c_0)\sqrt s\,|\delta_{J_0}|_2∣δ∣1​≤(1+c0​)s​∣δJ0​​∣2​.
  6. (B.28): on the cone, ∣δ∣2≤(1+c0s/m) ∣δJ01∣2|\delta|_2\le(1+c_0\sqrt{s/m})\,|\delta_{J_{01}}|_2∣δ∣2​≤(1+c0​s/m​)∣δJ01​​∣2​.
  7. (B.29): under RE(s,m,1)(s,m,1)(s,m,1), ∣δ∣22≤16(1+s/m)2(rs/κ2)2|\delta|_2^2\le16(1+\sqrt{s/m})^2(r\sqrt s/\kappa^2)^2∣δ∣22​≤16(1+s/m​)2(rs​/κ2)2.
  8. Interpolation: ∑jaj≤b1\sum_ja_j\le b_1∑j​aj​≤b1​, ∑jaj2≤b2\sum_ja_j^2\le b_2∑j​aj2​≤b2​, aj≥0a_j\ge0aj​≥0 imply ∑jajp≤b12−pb2p−1\sum_ja_j^p\le b_1^{2-p}b_2^{p-1}∑j​ajp​≤b12−p​b2p−1​ for 1<p≤21<p\le21<p≤2.

Significance

The result. Theorem 7.1 shows that, up to the factor log⁡M\log MlogM, the Dantzig selector estimates an sss-sparse vector as well as least squares would if the support were known: the prediction error 1n∣X(β^D−β∗)∣22\frac1n|X(\hat\beta_D-\beta^*)|_2^2n1​∣X(β^​D​−β∗)∣22​ is of order σ2slog⁡M/n\sigma^2s\log M/nσ2slogM/n, and the ℓp\ell_pℓp​ errors are of order s1/pσlog⁡M/ns^{1/p}\sigma\sqrt{\log M/n}s1/pσlogM/n​. The bounds hold for any MMM, including M≫nM\gg nM≫n, provided only that RE holds, and every constant is explicit. The paper's Theorem 7.2 gives the same rates for the Lasso; comparing the two is the paper's main message.

Formalizing it. The theorem is proved in the paper; to the best of current knowledge it has not been machine-checked. A complete formal proof would provide: a verified Gaussian maximal inequality for the noise event, the deterministic cone and RE arithmetic that underlies essentially all ℓ1\ell_1ℓ1​-regularized estimation theory, and a reusable ℓ1\ell_1ℓ1​–ℓ2\ell_2ℓ2​ interpolation lemma. Most milestones are deterministic and independent of the probability layer.

Difficulty

The obvious argument — compare β^D\hat\beta_Dβ^​D​ with β∗\beta^*β∗ in Euclidean norm using the smallest eigenvalue of XTX/nX^TX/nXTX/n — fails because that eigenvalue is 0 whenever M>nM>nM>n. The proof must instead show that the error vector lies in a cone on which XXX is injective in a quantitative sense, and this uses the optimality of β^D\hat\beta_Dβ^​D​ (not just feasibility) together with the event B\mathcal BB on which β∗\beta^*β∗ itself is feasible. The ℓp\ell_pℓp​ bound needs a second, stronger condition RE(s,m,1)(s,m,1)(s,m,1) and a control of the tail of the error outside the mmm largest coordinates. On the formal side, handling the non-uniqueness of the minimizer, real powers with exponent p−1p-1p−1 or 2−p2-p2−p, and the union over MMM Gaussian tails with the exact constant M1−A2/2M^{1-A^2/2}M1−A2/2 all need care.

Formalization scope

Vectors are functions Fin M → ℝ, the design is Matrix (Fin n) (Fin M) ℝ; the paper's dictionary of functions enters only through XXX. The unit diagonal of XTX/nX^TX/nXTX/n is a hypothesis, not a normalization performed in the proof. The noise is W : Fin n → Ω → ℝ on a probability space, measurable, mutually independent, each of law N(0,σ2)\mathcal N(0,\sigma^2)N(0,σ2); log⁡\loglog is the natural logarithm. The Dantzig selector is a predicate (feasible and of minimal ℓ1\ell_1ℓ1​ norm among feasible vectors), and every result is stated for every minimizer. RE(s,c0)(s,c_0)(s,c0​) and RE(s,m,c0)(s,m,c_0)(s,m,c0​) are stated through a witness κ\kappaκ (a number with the defining lower-bound property); κ(s,c0)\kappa(s,c_0)κ(s,c0​) is the largest witness and the bounds decrease in κ\kappaκ, so the statements are equivalent to the paper's while avoiding the value of a real infimum over an empty set. Two witnesses are kept apart: κ\kappaκ for RE(s,1)(s,1)(s,1) in (7.4)–(7.5), κ′\kappa'κ′ for RE(s,m,1)(s,m,1)(s,m,1) in (7.6). Ties in the choice of the mmm largest coordinates are handled by quantifying over every admissible J1J_1J1​. The probability statement asserts one measurable event EEE with P(E)≥1−M1−A2/2\mathbb P(E)\ge1-M^{1-A^2/2}P(E)≥1−M1−A2/2 on which all three bounds hold for every minimizer, every admissible mmm, every witness κ′\kappa'κ′ and every ppp.

The event EEE is fixed before the minimizer is quantified, so a formalization in which the event depends on β^D\hat\beta_Dβ^​D​, or in which RE is a hypothesis about the random error vector rather than the design, would be a different (weaker) statement and is not accepted. The deterministic milestones (B.26)–(B.29) take the conclusion of (B.25) as a hypothesis; they are true for every vector satisfying their hypotheses and are not restricted to the event.

A complete development needs Gaussian tail bounds and a union bound (Mathlib's gaussianReal), finite Hölder-type inequalities for real exponents, and elementary sorting arguments for the tail outside J01J_{01}J01​. The cone, RE and interpolation lemmas are reusable for the Lasso (Theorem 7.2, a sister mission) and beyond. Proofs of any milestone are welcome independently.

Selected references

  • P. J. Bickel, Y. Ritov, A. B. Tsybakov, Simultaneous analysis of Lasso and Dantzig selector, Ann. Statist. 37(4), 1705–1732, 2009. arXiv:0801.1095v3: https://arxiv.org/abs/0801.1095
  • E. Candès, T. Tao, The Dantzig selector: statistical estimation when p is much larger than n, Ann. Statist. 35(6), 2313–2351, 2007. https://doi.org/10.1214/009053606000001523
  • R. Tibshirani, Regression shrinkage and selection via the lasso, J. R. Stat. Soc. B 58(1), 267–288, 1996. https://doi.org/10.1111/j.2517-6161.1996.tb02080.x
  • S. van de Geer, P. Bühlmann, On the conditions used to prove oracle results for the Lasso, Electron. J. Statist. 3, 1360–1392, 2009. https://doi.org/10.1214/09-EJS506
10 thms4 active usersReviewed
🏆Completed
Machine LearningOptimizationStatistics·Captain: mikedeng1

Simultaneous Analysis of Lasso and Dantzig Selector II: Approximate Equivalence of the Lasso and Dantzig Prediction LossesResearch Paper

Motivation

Two estimators dominate sparse high-dimensional regression, where the number MMM of candidate regressors can far exceed the sample size nnn. The Lasso (Tibshirani, 1996) minimises a least-squares criterion plus an ℓ1\ell_1ℓ1​ penalty. The Dantzig selector (Candès and Tao, 2007) minimises the ℓ1\ell_1ℓ1​ norm of the coefficients subject to a bound on the correlation between the residual and the regressors, and is computed by a linear program. They were proposed independently and first analysed under different assumptions: sparsity oracle inequalities for the Lasso (Bunea, Tsybakov and Wegkamp, 2007) and ℓ2\ell_2ℓ2​ bounds for the Dantzig selector under a uniform uncertainty principle (Candès and Tao, 2007). A practitioner choosing between them needs to know whether guarantees for one say anything about the other.

Bickel, Ritov and Tsybakov (arXiv:0801.1095; Ann. Statist. 37(4), 2009, doi:10.1214/08-AOS620) analyse both estimators in parallel under one assumption on the design, the restricted eigenvalue condition. Their main message is that, under sparsity, the two estimators "exhibit similar behavior" (p. 2). Section 5 makes this precise: the prediction losses of the two estimators are close. The result holds in a nonparametric model: the regression function need not be a combination of the regressors. This mission formalizes that comparison, Theorem 5.1 of the paper.

Setting

Let f1,…,fMf_1,\dots,f_Mf1​,…,fM​ be real functions (the dictionary) on a set Z\mathcal ZZ, and Z1,…,Zn∈ZZ_1,\dots,Z_n\in\mathcal ZZ1​,…,Zn​∈Z fixed design points, with n≥1n\ge1n≥1 and M≥2M\ge2M≥2. The design matrix is X=(fj(Zi))∈Rn×MX=(f_j(Z_i))\in\mathbb R^{n\times M}X=(fj​(Zi​))∈Rn×M. Observations are Yi=f(Zi)+WiY_i=f(Z_i)+W_iYi​=f(Zi​)+Wi​, where fff is an unknown function and W1,…,WnW_1,\dots,W_nW1​,…,Wn​ are independent N(0,σ2)\mathcal N(0,\sigma^2)N(0,σ2) with σ>0\sigma>0σ>0. Write y=(Yi)y=(Y_i)y=(Yi​), f=(f(Zi))\boldsymbol f=(f(Z_i))f=(f(Zi​)) and w=(Wi)w=(W_i)w=(Wi​), so y=f+wy=\boldsymbol f+wy=f+w.

The empirical norm of ggg is ∥g∥n=(1n∑ig(Zi)2)1/2\|g\|_n=(\tfrac1n\sum_i g(Z_i)^2)^{1/2}∥g∥n​=(n1​∑i​g(Zi​)2)1/2. Every column has ∥fj∥n≠0\|f_j\|_n\neq0∥fj​∥n​=0, and fmax⁡=max⁡j∥fj∥nf_{\max}=\max_j\|f_j\|_nfmax​=maxj​∥fj​∥n​. For β∈RM\beta\in\mathbb R^Mβ∈RM, fβ=∑jβjfjf_\beta=\sum_j\beta_jf_jfβ​=∑j​βj​fj​ has value vector XβX\betaXβ. The support is J(β)={j:βj≠0}J(\beta)=\{j:\beta_j\ne0\}J(β)={j:βj​=0} and the sparsity is M(β)=∣J(β)∣\mathcal M(\beta)=|J(\beta)|M(β)=∣J(β)∣. For J⊆{1,…,M}J\subseteq\{1,\dots,M\}J⊆{1,…,M}, δJ\delta_JδJ​ agrees with δ\deltaδ on JJJ and vanishes elsewhere.

Fix r>0r>0r>0. The Lasso β^L\hat\beta_Lβ^​L​ is any minimiser of

1n∑i=1n(Yi−fβ(Zi))2+2r∑j=1M∥fj∥n∣βj∣.\frac1n\sum_{i=1}^n\big(Y_i-f_\beta(Z_i)\big)^2+2r\sum_{j=1}^M\|f_j\|_n|\beta_j| .n1​i=1∑n​(Yi​−fβ​(Zi​))2+2rj=1∑M​∥fj​∥n​∣βj​∣.

With D=diag(∥f1∥n2,…,∥fM∥n2)D=\mathrm{diag}(\|f_1\|_n^2,\dots,\|f_M\|_n^2)D=diag(∥f1​∥n2​,…,∥fM​∥n2​), the Dantzig constraint is ∣1nD−1/2X⊤(y−Xβ)∣∞≤r|\tfrac1nD^{-1/2}X^\top(y-X\beta)|_\infty\le r∣n1​D−1/2X⊤(y−Xβ)∣∞​≤r. The Dantzig selector β^D\hat\beta_Dβ^​D​ is any vector of smallest ∣β∣1=∑j∣βj∣|\beta|_1=\sum_j|\beta_j|∣β∣1​=∑j​∣βj​∣ that satisfies it. The estimators are f^L=fβ^L\hat f_L=f_{\hat\beta_L}f^​L​=fβ^​L​​ and f^D=fβ^D\hat f_D=f_{\hat\beta_D}f^​D​=fβ^​D​​.

Assumption RE(s,c0)(s,c_0)(s,c0​) with 1≤s≤M1\le s\le M1≤s≤M, c0>0c_0>0c0​>0 asks that

κ(s,c0)=min⁡∣J0∣≤s min⁡δ≠0, ∣δJ0c∣1≤c0∣δJ0∣1 ∣Xδ∣2n ∣δJ0∣2>0.\kappa(s,c_0)=\min_{|J_0|\le s}\ \min_{\delta\ne0,\ |\delta_{J_0^c}|_1\le c_0|\delta_{J_0}|_1}\ \frac{|X\delta|_2}{\sqrt n\,|\delta_{J_0}|_2}>0 .κ(s,c0​)=∣J0​∣≤smin​ δ=0, ∣δJ0c​​∣1​≤c0​∣δJ0​​∣1​min​ n​∣δJ0​​∣2​∣Xδ∣2​​>0.

Throughout, r=Aσlog⁡M/nr=A\sigma\sqrt{\log M/n}r=AσlogM/n​, where log⁡\loglog is the natural logarithm.

Formalization targets

Goal: Theorem 5.1

Assume RE(s,1)(s,1)(s,1) with 1≤s≤M1\le s\le M1≤s≤M, and let A>22A>2\sqrt2A>22​. With probability at least 1−M1−A2/81-M^{1-A^2/8}1−M1−A2/8, every Lasso solution with M(β^L)≤s\mathcal M(\hat\beta_L)\le sM(β^​L​)≤s and every Dantzig selector satisfy

∣ ∥f^D−f∥n2−∥f^L−f∥n2 ∣≤16A2 M(β^L)σ2n fmax⁡2κ2(s,1) log⁡M.\Big|\,\|\hat f_D-f\|_n^2-\|\hat f_L-f\|_n^2\,\Big|\le16A^2\,\frac{\mathcal M(\hat\beta_L)\sigma^2}{n}\,\frac{f_{\max}^2}{\kappa^2(s,1)}\,\log M .​∥f^​D​−f∥n2​−∥f^​L​−f∥n2​​≤16A2nM(β^​L​)σ2​κ2(s,1)fmax2​​logM.

Milestones

The proof uses one probabilistic event and two one-sided deterministic inequalities.

  1. The Lasso satisfies the Dantzig constraint (2.3).
  2. The noise event A=⋂j{2∣1n∑iXijWi∣≤r∥fj∥n}\mathcal A=\bigcap_j\{2|\tfrac1n\sum_iX_{ij}W_i|\le r\|f_j\|_n\}A=⋂j​{2∣n1​∑i​Xij​Wi​∣≤r∥fj​∥n​} has P(Ac)≤M1−A2/8\mathbb P(\mathcal A^c)\le M^{1-A^2/8}P(Ac)≤M1−A2/8 (B.4).
  3. On A\mathcal AA, ∣1nX⊤(f−Xβ^L)∣∞≤3rfmax⁡/2|\tfrac1nX^\top(\boldsymbol f-X\hat\beta_L)|_\infty\le 3rf_{\max}/2∣n1​X⊤(f−Xβ^​L​)∣∞​≤3rfmax​/2 (Lemma B.1, (B.2)).
  4. The Dantzig error lies in the cone ∣δJ0c∣1≤∣δJ0∣1|\delta_{J_0^c}|_1\le|\delta_{J_0}|_1∣δJ0c​​∣1​≤∣δJ0​​∣1​ (Lemma B.3, (B.9)).
  5. On the larger event B⊇A\mathcal B\supseteq\mathcal AB⊇A, ∣1nX⊤(f−Xβ^D)∣∞≤2rfmax⁡|\tfrac1nX^\top(\boldsymbol f-X\hat\beta_D)|_\infty\le 2rf_{\max}∣n1​X⊤(f−Xβ^​D​)∣∞​≤2rfmax​ (Lemma B.3, (B.10)).
  6. ∥f^D−f∥n2≤∥f^L−f∥n2+16fmax⁡2r2M(β^L)/κ2\|\hat f_D-f\|_n^2\le\|\hat f_L-f\|_n^2+16f_{\max}^2r^2\mathcal M(\hat\beta_L)/\kappa^2∥f^​D​−f∥n2​≤∥f^​L​−f∥n2​+16fmax2​r2M(β^​L​)/κ2 on B\mathcal BB (B.15).
  7. ∥f^L−f∥n2≤∥f^D−f∥n2+9fmax⁡2r2M(β^L)/κ2\|\hat f_L-f\|_n^2\le\|\hat f_D-f\|_n^2+9f_{\max}^2r^2\mathcal M(\hat\beta_L)/\kappa^2∥f^​L​−f∥n2​≤∥f^​D​−f∥n2​+9fmax2​r2M(β^​L​)/κ2 on A\mathcal AA (B.17).

Further result: Theorem 5.2

Assume ∥fj∥n=1\|f_j\|_n=1∥fj​∥n​=1 for all jjj and RE(s,5)(s,5)(s,5). With probability at least 1−M1−A2/81-M^{1-A^2/8}1−M1−A2/8, whenever M(β^D)≤s\mathcal M(\hat\beta_D)\le sM(β^​D​)≤s,

∥f^L−f∥n2≤10∥f^D−f∥n2+81A2 M(β^D)σ2n log⁡Mκ2(s,5).\|\hat f_L-f\|_n^2\le10\|\hat f_D-f\|_n^2+81A^2\,\frac{\mathcal M(\hat\beta_D)\sigma^2}{n}\,\frac{\log M}{\kappa^2(s,5)} .∥f^​L​−f∥n2​≤10∥f^​D​−f∥n2​+81A2nM(β^​D​)σ2​κ2(s,5)logM​.

Significance

The result. Theorem 5.1 bounds the gap between the two prediction losses by the rate M(β^L)σ2log⁡M/n\mathcal M(\hat\beta_L)\sigma^2\log M/nM(β^​L​)σ2logM/n of a sparse regression with M(β^L)\mathcal M(\hat\beta_L)M(β^​L​) parameters. The bound carries a factor fmax⁡2/κ2(s,1)f^2_{\max}/\kappa^2(s,1)fmax2​/κ2(s,1) that measures how ill-conditioned the Gram matrix is on sparse vectors. A prediction bound for one estimator therefore transfers to the other at this cost. The paper uses this transfer in Proposition 6.3, which combines Theorem 5.1 with the Lasso oracle inequality of Section 6 to derive an oracle inequality for the Dantzig selector. The theorem requires no assumption relating fff to the dictionary.

Formalizing it. The result has been proved since 2009, and this mission formalizes that proof. None of the objects involved exists on Prove2Me yet: the weighted Lasso, the Dantzig selector and the Gaussian noise events. The Wainwright series on the platform defines a differently normalised Lasso with an unweighted penalty, a single fixed support and a different restricted eigenvalue condition, so it cannot be reused here. A machine-checked proof would also confirm the paper's constants, 16A216A^216A2 and the thresholds 222\sqrt222​ and M1−A2/8M^{1-A^2/8}M1−A2/8, which appear in all later analyses.

Difficulty

Each estimator is defined only implicitly, as the solution of an optimisation problem, and neither need be unique. Comparing their losses directly gives ±2nδ⊤X⊤(f−Xβ^)\pm\tfrac2n\delta^\top X^\top(\boldsymbol f-X\hat\beta)±n2​δ⊤X⊤(f−Xβ^​) plus 1n∣Xδ∣22\tfrac1n|X\delta|_2^2n1​∣Xδ∣22​ with δ=β^L−β^D\delta=\hat\beta_L-\hat\beta_Dδ=β^​L​−β^​D​. A crude bound on the cross term, ∣δ∣1⋅∣X⊤(⋅)∣∞|\delta|_1\cdot|X^\top(\cdot)|_\infty∣δ∣1​⋅∣X⊤(⋅)∣∞​, yields an error proportional to ∣δ∣1|\delta|_1∣δ∣1​. This does not produce the sparse rate unless ∣δ∣1|\delta|_1∣δ∣1​ is controlled by ∣Xδ∣2|X\delta|_2∣Xδ∣2​. That control needs δ\deltaδ to lie in the restricted eigenvalue cone at the support of the random, data-dependent vector β^L\hat\beta_Lβ^​L​. The restricted eigenvalue condition must therefore hold uniformly over supports of size at most sss; a condition for one fixed support does not suffice. The probabilistic part is a union bound over MMM Gaussian coordinates, and it must be arranged so that a single event serves every minimiser of both programs.

Formalization scope

The dictionary and the design points enter every statement only through XXX and f\boldsymbol ff, so the Lean statements take X : Matrix (Fin n) (Fin M) ℝ and f : Fin n → ℝ directly, and fff is arbitrary. The noise is a family W : Fin n → Ω → ℝ of measurable, mutually independent random variables with law gaussianReal 0 σ², and y=f+W(ω)y=f+W(\omega)y=f+W(ω). The Lasso and the Dantzig selector are predicates (IsLasso, IsDantzig), and every theorem is stated for every solution. The Dantzig constraint is written coordinatewise as ∣1n∑iXij(yi−(Xβ)i)∣≤r∥fj∥n|\tfrac1n\sum_iX_{ij}(y_i-(X\beta)_i)|\le r\|f_j\|_n∣n1​∑i​Xij​(yi​−(Xβ)i​)∣≤r∥fj​∥n​. The Lasso penalty and the Dantzig constraint are weighted by ∥fj∥n\|f_j\|_n∥fj​∥n​, and the Dantzig objective ∣β∣1|\beta|_1∣β∣1​ is unweighted, exactly as in the paper. Theorem 5.1 does not normalise the columns.

RE(s,c0)(s,c_0)(s,c0​) is stated through a witness: a real κ>0\kappa>0κ>0 with κn∣δJ0∣2≤∣Xδ∣2\kappa\sqrt n|\delta_{J_0}|_2\le|X\delta|_2κn​∣δJ0​​∣2​≤∣Xδ∣2​ on the cone, for all ∣J0∣≤s|J_0|\le s∣J0​∣≤s. The paper's κ(s,c0)\kappa(s,c_0)κ(s,c0​) is attained, so it is the largest witness. Every bound decreases in κ\kappaκ, so this reading is equivalent to the paper's and avoids a real infimum over an empty set. "With probability at least ppp" becomes the existence of a measurable event EEE with P(E)≥p\mathbb P(E)\ge pP(E)≥p on which the conclusion holds for every Lasso solution and every Dantzig selector. The condition M(β^L)≤s\mathcal M(\hat\beta_L)\le sM(β^​L​)≤s is imposed inside the event, per realisation. The milestones (B.2), (B.10), (B.15) and (B.17) are stated deterministically, on the noise events A\mathcal AA and B\mathcal BB as predicates on the noise vector; this is how the proof uses them. (B.4) is stated for every A>0A>0A>0, which is stronger than the paper's A>22A>2\sqrt2A>22​ and still true.

The goal cannot be made vacuous. For A>22A>2\sqrt2A>22​ the probability bound 1−M1−A2/81-M^{1-A^2/8}1−M1−A2/8 is positive. Lasso solutions exist because r>0r>0r>0 and every ∥fj∥n>0\|f_j\|_n>0∥fj​∥n​>0, and Dantzig selectors exist because the Lasso is feasible. RE(s,1)(s,1)(s,1) with κ=1\kappa=1κ=1 holds for X=n IX=\sqrt n\,IX=n​I.

A complete development needs:

  • subgradient optimality for the weighted Lasso;
  • a Gaussian tail bound P(∣η∣≥t)≤e−t2/2\mathbb P(|\eta|\ge t)\le e^{-t^2/2}P(∣η∣≥t)≤e−t2/2 together with the law of a weighted sum of independent Gaussians;
  • Cauchy–Schwarz on supports;
  • the quadratic bound bx−x2≤b2/4bx-x^2\le b^2/4bx−x2≤b2/4.

The noise-event lemmas and the Lasso optimality condition can be reused by the companion missions on this paper. Proofs of any milestone are welcome, including proofs that route (B.4) through Mathlib's sub-Gaussian API.

Selected references

  • P. J. Bickel, Y. Ritov, A. B. Tsybakov, Simultaneous analysis of Lasso and Dantzig selector, Ann. Statist. 37(4), 1705–1732, 2009. arXiv:0801.1095v3: https://arxiv.org/abs/0801.1095 ; doi:10.1214/08-AOS620
  • E. Candès, T. Tao, The Dantzig selector: statistical estimation when p is much larger than n, Ann. Statist. 35(6), 2313–2351, 2007. https://doi.org/10.1214/009053606000001523
  • R. Tibshirani, Regression shrinkage and selection via the lasso, J. R. Stat. Soc. B 58(1), 267–288, 1996. https://doi.org/10.1111/j.2517-6161.1996.tb02080.x
  • F. Bunea, A. B. Tsybakov, M. H. Wegkamp, Sparsity oracle inequalities for the Lasso, Electron. J. Stat. 1, 169–194, 2007. https://doi.org/10.1214/07-EJS008
9 thms4 active usersReviewed
🏆Completed
Bandit AlgorithmsMachine LearningOperations Research·Captain: mikedeng1

Stochastic Linear Optimization under Bandit Feedback 1: The Regret Bound of ConfidenceBallResearch Paper

Motivation

In stochastic linear optimization under bandit feedback a learner repeatedly chooses a decision xtx_txt​ from a fixed set D⊆RnD\subseteq\mathbb R^nD⊆Rn and observes only the noisy cost of that one decision, whose expectation is a fixed but unknown linear function μ⊤xt\mu^\top x_tμ⊤xt​. The model covers online routing, ad and product selection with feature vectors, and any sequential decision problem whose decision set is too large to enumerate but whose expected cost is linear in a known representation. The multi-armed bandit is the special case where DDD is the set of standard basis vectors.

Dani, Hayes and Kakade (COLT 2008) analysed the algorithm ConfidenceBall₂, a generalization of Auer's LinRel (JMLR 2002), and proved that its regret is O∗(nT)O^*(n\sqrt T)O∗(nT​) with high probability for an arbitrary compact decision set, and that this is optimal up to logarithmic factors. Their confidence-ellipsoid construction became the template for later linear bandit algorithms (OFUL, LinUCB), and the ellipsoid-plus-potential analysis is the standard argument of the field (Lattimore and Szepesvári, Bandit Algorithms, Chapters 19–20).

Timeline: Auer (2002) introduced LinRel for finite decision sets; Dani, Hayes and Kakade (2008) extended it to arbitrary compact sets with the O∗(nT)O^*(n\sqrt T)O∗(nT​) bound and a matching Ω(nT)\Omega(n\sqrt T)Ω(nT​) lower bound; Rusmevichientong and Tsitsiklis (Math. Oper. Res. 2010) studied linearly parameterized bandits with dimension-dependent regret bounds; Abbasi-Yadkori, Pál and Szepesvári (NeurIPS 2011) sharpened the confidence ellipsoids with self-normalized martingale bounds.

Setting

Fix n≥1n\ge1n≥1 and a compact decision set D⊆RnD\subseteq\mathbb R^nD⊆Rn whose standard basis e1,…,ene_1,\dots,e_ne1​,…,en​ is a barycentric spanner: each ei∈De_i\in Dei​∈D and every x∈Dx\in Dx∈D lies in the cube [−1,1]n[-1,1]^n[−1,1]n. The paper's Section 5 adopts these coordinates without loss of generality. An unknown vector μ∈Rn\mu\in\mathbb R^nμ∈Rn satisfies ∣μ⊤x∣≤1|\mu^\top x|\le1∣μ⊤x∣≤1 for x∈Dx\in Dx∈D, and x∗∈Dx^*\in Dx∗∈D minimises μ⊤x\mu^\top xμ⊤x.

On round t=1,2,…t=1,2,\dotst=1,2,… the learner plays xt∈Dx_t\in Dxt​∈D, measurable with respect to the information Ft\mathcal F_tFt​ before round ttt, and observes a loss ℓt∈[−1,1]\ell_t\in[-1,1]ℓt​∈[−1,1] with E[ℓt∣Ft]=μ⊤xt\mathbb E[\ell_t\mid\mathcal F_t]=\mu^\top x_tE[ℓt​∣Ft​]=μ⊤xt​. The regret after TTT rounds is

RT=∑t=1T(μ⊤xt−μ⊤x∗).R_T=\sum_{t=1}^T\big(\mu^\top x_t-\mu^\top x^*\big).RT​=t=1∑T​(μ⊤xt​−μ⊤x∗).

ConfidenceBall₂(D,δ)(D,\delta)(D,δ) maintains the design matrix At=I+∑τ<txτxτ⊤A_t=I+\sum_{\tau<t}x_\tau x_\tau^\topAt​=I+∑τ<t​xτ​xτ⊤​, the least-squares estimate μ^t=At−1∑τ<tℓτxτ\hat\mu_t=A_t^{-1}\sum_{\tau<t}\ell_\tau x_\tauμ^​t​=At−1​∑τ<t​ℓτ​xτ​, the radius

βt=max⁡(128 nln⁡tln⁡(t2/δ), (83ln⁡(t2/δ))2),\beta_t=\max\Big(128\,n\ln t\ln(t^2/\delta),\ \big(\tfrac83\ln(t^2/\delta)\big)^2\Big),βt​=max(128nlntln(t2/δ), (38​ln(t2/δ))2),

and the confidence ellipsoid Bt2={ν:(ν−μ^t)⊤At(ν−μ^t)≤βt}B^2_t=\{\nu:(\nu-\hat\mu_t)^\top A_t(\nu-\hat\mu_t)\le\beta_t\}Bt2​={ν:(ν−μ^​t​)⊤At​(ν−μ^​t​)≤βt​}. It plays the optimistic decision xt∈argmin⁡x∈Dmin⁡ν∈Bt2ν⊤xx_t\in\operatorname{argmin}_{x\in D}\min_{\nu\in B^2_t}\nu^\top xxt​∈argminx∈D​minν∈Bt2​​ν⊤x. The analysis uses the width wt=xt⊤At−1xtw_t=\sqrt{x_t^\top A_t^{-1}x_t}wt​=xt⊤​At−1​xt​​, the error Zt=(μ^t−μ)⊤At(μ^t−μ)Z_t=(\hat\mu_t-\mu)^\top A_t(\hat\mu_t-\mu)Zt​=(μ^​t​−μ)⊤At​(μ^​t​−μ), and the noise ηt=ℓt−μ⊤xt\eta_t=\ell_t-\mu^\top x_tηt​=ℓt​−μ⊤xt​.

Formalization targets

Goal: Theorem 2, ConfidenceBall₂ bullet (corrected)

For 0<δ<10<\delta<10<δ<1 with n≤β1n\le\beta_1n≤β1​ and noise ∣ηt∣≤1|\eta_t|\le1∣ηt​∣≤1,

Pr⁡(∀T≥1, RT≤8nTβTln⁡(T+1))≥1−δ.\Pr\Big(\forall T\ge1,\ R_T\le\sqrt{8nT\beta_T\ln(T+1)}\Big)\ge1-\delta.Pr(∀T≥1, RT​≤8nTβT​ln(T+1)​)≥1−δ.

A single event covers every horizon, so the bound is anytime.

Milestones

  • Lemma 8. If μ∈Bt2\mu\in B^2_tμ∈Bt2​ then rt≤2min⁡(βtwt,1)r_t\le2\min(\sqrt{\beta_t}w_t,1)rt​≤2min(βt​​wt​,1).
  • Lemma 10. det⁡At+1=∏τ=1t(1+wτ2)\det A_{t+1}=\prod_{\tau=1}^t(1+w_\tau^2)detAt+1​=∏τ=1t​(1+wτ2​).
  • Lemma 9 (corrected). ∑τ=1tmin⁡(wτ2,1)≤2nln⁡(t+1)\sum_{\tau=1}^t\min(w_\tau^2,1)\le2n\ln(t+1)∑τ=1t​min(wτ2​,1)≤2nln(t+1).
  • Theorem 6 (corrected). If μ∈Bt2\mu\in B^2_tμ∈Bt2​ for all t≤Tt\le Tt≤T, then ∑t≤Trt2≤8nβTln⁡(T+1)\sum_{t\le T}r_t^2\le8n\beta_T\ln(T+1)∑t≤T​rt2​≤8nβT​ln(T+1).
  • Theorem 4 (Freedman). Pr⁡(∑Xi≥a, V≤v)≤exp⁡(−a2/(2v+2ab/3))\Pr(\sum X_i\ge a,\ V\le v)\le\exp(-a^2/(2v+2ab/3))Pr(∑Xi​≥a, V≤v)≤exp(−a2/(2v+2ab/3)) for martingale differences bounded above by bbb.
  • Lemma 12. Zt≤n+2∑τ<tητxτ⊤(μ^τ−μ)1+wτ2+∑τ<tητ2wτ21+wτ2Z_t\le n+2\sum_{\tau<t}\eta_\tau\frac{x_\tau^\top(\hat\mu_\tau-\mu)}{1+w_\tau^2}+\sum_{\tau<t}\eta_\tau^2\frac{w_\tau^2}{1+w_\tau^2}Zt​≤n+2∑τ<t​ητ​1+wτ2​xτ⊤​(μ^​τ​−μ)​+∑τ<t​ητ2​1+wτ2​wτ2​​.
  • Lemma 14. Pr⁡(∀t, ∑τ<tMτ≤βt/2)≥1−δ\Pr(\forall t,\ \sum_{\tau<t}M_\tau\le\beta_t/2)\ge1-\deltaPr(∀t, ∑τ<t​Mτ​≤βt​/2)≥1−δ.
  • Theorem 5 (Confidence). Pr⁡(∀t, μ∈Bt2)≥1−δ\Pr(\forall t,\ \mu\in B^2_t)\ge1-\deltaPr(∀t, μ∈Bt2​)≥1−δ.

Significance

The result gives a regret bound for linear bandits over an arbitrary compact decision set that depends on the dimension nnn rather than on ∣D∣|D|∣D∣, holds uniformly over horizons, and requires no gap between the best and second-best decision. With the paper's lower bound it shows that the price of bandit feedback, compared with full information, is a factor Θ∗(n)\Theta^*(\sqrt n)Θ∗(n​). The two components, a confidence theorem for a least-squares ellipsoid under martingale noise and a deterministic potential argument on log⁡det⁡At\log\det A_tlogdetAt​, are reused in the analysis of most optimistic linear and generalized-linear bandit algorithms.

The theorem has a published proof but, to our knowledge, no machine-checked one. The platform already has the elliptical potential lemma (BanditAlgorithm.elliptical_potential_lemma, Lattimore–Szepesvári Lemma 19.4), the matrix determinant lemma and the Woodbury identity, which cover the linear-algebra layer; the LinUCB regret theorem there (BanditAlgorithm.linear_bandit_linucb_regret_bound) is pathwise given confidence, for a different algorithm, so the probabilistic half is new. Formalization also settles the printed constants, three of which need correction (below).

Difficulty

The deterministic half is linear algebra. The difficulty is the confidence theorem. Hoeffding–Azuma applied to ∑τMτ\sum_\tau M_\tau∑τ​Mτ​ would need a deterministic step bound, and the natural one gives only a T3/4T^{3/4}T3/4 regret. The step sizes of MtM_tMt​ are bounded in terms of the random widths wtw_twt​, so the argument must control the conditional variances pathwise and apply Freedman's inequality, whose event {V≤v}\{V\le v\}{V≤v} is random. The escape indicator EtE_tEt​, which switches the martingale off after the first failure of confidence, is what makes the variance bound hold on every path, and the induction that turns Lemma 14 into Theorem 5 must be carried out on a single event for all ttt simultaneously. Freedman's inequality itself is not in Mathlib.

Formalization scope

Vectors are Fin n → ℝ, matrices Matrix (Fin n) (Fin n) ℝ, rounds are indexed t=1,2,…t=1,2,\dotst=1,2,… in ℕ. The spanner is the standard basis, as in Section 5 of the paper (the algorithm is equivariant under the linear change of coordinates). The probability model is a probability space with a general filtration (Ft)(\mathcal F_t)(Ft​); xtx_txt​ is Ft\mathcal F_tFt​-measurable and ℓt\ell_tℓt​ is Ft+1\mathcal F_{t+1}Ft+1​-measurable. The argmin is encoded as a joint minimiser over D×Bt2D\times B^2_tD×Bt2​, which admits every tie-break and is required on every outcome; measurability of xtx_txt​ is a hypothesis, not derived from the selection. The optimum x∗x^*x∗ is a hypothesis (x∗∈Dx^*\in Dx∗∈D, minimising), not an sInf.

Corrections of the printed statements, each labelled in the item's Formalization Note:

  1. ln⁡(T+1)\ln(T+1)ln(T+1) for ln⁡T\ln TlnT in Lemma 9, Theorem 6 and Theorem 2. The printed bounds are false at T=1T=1T=1 (with n=1n=1n=1, D=[−1,1]D=[-1,1]D=[−1,1], μ>0\mu>0μ>0, the tie-break x1=1x_1=1x1​=1 gives R1=2μ>0R_1=2\mu>0R1​=2μ>0 against a bound of 000); the proof of Lemma 9 gives 2ln⁡det⁡At+1≤2nln⁡(t+1)2\ln\det A_{t+1}\le2n\ln(t+1)2lndetAt+1​≤2nln(t+1).
  2. n≤β1=(83ln⁡(1/δ))2n\le\beta_1=(\tfrac83\ln(1/\delta))^2n≤β1​=(38​ln(1/δ))2 is added to Theorems 5 and 2: the proof of Theorem 5 claims Z1≤n<β1Z_1\le n<\beta_1Z1​≤n<β1​, which fails for δ\deltaδ near 111. Theorem 6 takes the proof's "1<β11<\beta_11<β1​" as the hypothesis β1≥1\beta_1\ge1β1​≥1.
  3. ∣ℓt−μ⊤xt∣≤1|\ell_t-\mu^\top x_t|\le1∣ℓt​−μ⊤xt​∣≤1 is added to Lemma 14, Theorems 5 and 2: Section 5.2 uses ∣ηt∣≤1|\eta_t|\le1∣ηt​∣≤1, while the model gives only ∣ηt∣≤2|\eta_t|\le2∣ηt​∣≤2. It holds when costs lie in [0,1][0,1][0,1].
  4. Theorem 5's "δ>0\delta>0δ>0" is stated with 0<δ<10<\delta<10<δ<1; Lemma 10's index typo (wtw_twt​ for wτw_\tauwτ​) and Theorem 4's ∑i=1n\sum_{i=1}^n∑i=1n​ (for TTT) are corrected.

A regret bound for an arbitrary decision sequence under the assumption μ∈Bt2\mu\in B^2_tμ∈Bt2​ for all ttt is Theorem 6, not the goal; the goal carries the ConfidenceBall₂ selection rule, the conditional-mean and measurability hypotheses, and δ\deltaδ as the algorithm's own parameter. The hypotheses are jointly satisfiable, for example by a finite DDD with a fixed tie-break and i.i.d. costs in [0,1][0,1][0,1].

Needed infrastructure: Freedman's inequality for a filtration (reusable across all of bandit theory), the potential lemma (available), and measurability of the algorithm's statistics. Contributions of alternative proofs of Theorem 5, for example via self-normalized bounds, are welcome.

Selected references

  • V. Dani, T. P. Hayes, S. M. Kakade, Stochastic Linear Optimization under Bandit Feedback, COLT 2008. http://colt2008.cs.helsinki.fi/papers/80-Dani.pdf
  • P. Auer, Using Confidence Bounds for Exploitation-Exploration Trade-offs, JMLR 3, 2002. https://www.jmlr.org/papers/v3/auer02a.html
  • D. A. Freedman, On Tail Probabilities for Martingales, Annals of Probability 3(1), 1975. https://doi.org/10.1214/aop/1176996452
  • B. Awerbuch, R. Kleinberg, Adaptive Routing with End-to-End Feedback, STOC 2004. https://doi.org/10.1145/1007352.1007367
  • P. Rusmevichientong, J. N. Tsitsiklis, Linearly Parameterized Bandits, Mathematics of Operations Research 35(2), 2010. https://doi.org/10.1287/moor.1100.0446
  • Y. Abbasi-Yadkori, D. Pál, Cs. Szepesvári, Improved Algorithms for Linear Stochastic Bandits, NeurIPS 2011. https://arxiv.org/abs/1102.2670
  • T. Lattimore, Cs. Szepesvári, Bandit Algorithms, Cambridge University Press, 2020. https://doi.org/10.1017/9781108571401
12 thms4 active usersReviewed
🏆Completed
Machine LearningStatistics·Captain: naimengye

Understanding Machine Learning XXV: PAC-BayesTextbook

Motivation

The MDL and Occam principles of Chapter 7 rank hypotheses by description length and pay for a hypothesis according to its rank. Chapter 31 of Shalev-Shwartz and Ben-David, Understanding Machine Learning: From Theory to Algorithms (doi:10.1017/CBO9781107298019), generalizes this to the PAC-Bayesian approach of McAllester: prior knowledge is a prior distribution PPP over the class, the learner outputs a posterior QQQ, read as the randomized predictor that draws h∼Qh \sim Qh∼Q, and the price of QQQ is its Kullback–Leibler divergence from PPP. The PAC-Bayes theorem (Theorem 31.1) says that with probability 1−δ1 - \delta1−δ, simultaneously for every posterior, the generalization loss exceeds the training loss by at most (D(Q∥P)+ln⁡(m/δ))/(2(m−1))\sqrt{(D(Q\|P) + \ln(m/\delta))/(2(m-1))}(D(Q∥P)+ln(m/δ))/(2(m−1))​. Its proof is a compact and elegant argument: Markov's inequality for an exponential moment, a change of measure from QQQ to PPP by Jensen's inequality, an exchange of expectations that is possible because the prior does not depend on the sample, and a moment bound for the deviation of a single hypothesis. The bound suggests the learning rule of Remark 31.1, minimize LS(Q)L_S(Q)LS​(Q) plus the divergence penalty, which is regularized risk minimization in disguise, and for a finite class with a uniform prior it recovers an Occam-type bound (Exercise 2).

Setting

HHH is a measurable space of hypotheses, ℓ:H×Z→[0,1]\ell : H \times Z \to [0,1]ℓ:H×Z→[0,1] a jointly measurable loss, DDD a distribution over ZZZ, and PPP a prior probability measure on HHH. For a posterior QQQ, ℓ(Q,z)=Eh∼Q[ℓ(h,z)]\ell(Q, z) = \mathbb{E}_{h \sim Q}[\ell(h, z)]ℓ(Q,z)=Eh∼Q​[ℓ(h,z)], LD(Q)=Eh∼Q[LD(h)]L_D(Q) = \mathbb{E}_{h \sim Q}[L_D(h)]LD​(Q)=Eh∼Q​[LD​(h)] and LS(Q)=Eh∼Q[LS(h)]L_S(Q) = \mathbb{E}_{h \sim Q}[L_S(h)]LS​(Q)=Eh∼Q​[LS​(h)], and D(Q∥P)=Eh∼Q[ln⁡(dQ/dP)]D(Q\|P) = \mathbb{E}_{h \sim Q}[\ln(dQ/dP)]D(Q∥P)=Eh∼Q​[ln(dQ/dP)] is the Kullback–Leibler divergence, a real number when Q≪PQ \ll PQ≪P and the log-density is QQQ-integrable.

Formalization targets

Goal: Theorem 31.1

For m≥2m \ge 2m≥2 and δ∈(0,1)\delta \in (0,1)δ∈(0,1), with probability at least 1−δ1 - \delta1−δ over S∼DmS \sim D^mS∼Dm, every probability measure Q≪PQ \ll PQ≪P with finite divergence satisfies

LD(Q)≤LS(Q)+D(Q∥P)+ln⁡(m/δ)2(m−1).L_D(Q) \le L_S(Q) + \sqrt{\frac{D(Q\|P) + \ln(m/\delta)}{2(m-1)}}.LD​(Q)≤LS​(Q)+2(m−1)D(Q∥P)+ln(m/δ)​​.

Milestones

The identity Ez∼D[ℓ(Q,z)]=LD(Q)\mathbb{E}_{z \sim D}[\ell(Q, z)] = L_D(Q)Ez∼D​[ℓ(Q,z)]=LD​(Q) (§31.1); the moment bound ES[e2(m−1)Δ(h)2]≤m\mathbb{E}_S[e^{2(m-1)\Delta(h)^2}] \le mES​[e2(m−1)Δ(h)2]≤m of the proof (p. 417); Exercise 2, the bound for a finite class with the uniform prior.

Significance

PAC-Bayes bounds are among the tightest generalization bounds known in practice, and the reason is visible in Theorem 31.1: the complexity term is not a property of the class but of the posterior actually chosen, measured against a prior, so a learner that stays close to its prior generalizes even in a huge class. The theorem is the ancestor of a large literature (Seeger, Langford, Catoni, Maurer) and of modern nonvacuous bounds for neural networks. Formally it is a pleasant target: the change-of-measure inequality Eh∼Q[f(h)]−D(Q∥P)≤ln⁡Eh∼P[ef(h)]\mathbb{E}_{h \sim Q}[f(h)] - D(Q\|P) \le \ln\mathbb{E}_{h \sim P}[e^{f(h)}]Eh∼Q​[f(h)]−D(Q∥P)≤lnEh∼P​[ef(h)] is the Donsker–Varadhan inequality, of independent value, and the moment bound is a sharp sub-Gaussian fact. On the platform, the mission introduces Gibbs risks and the Kullback–Leibler divergence between measures on a class, ending the book's series with its last learning principle.

Difficulty

The Gibbs risk identity is Fubini for a bounded jointly measurable function. The moment bound is the delicate step: the book derives it from Hoeffding's tail bound through Exercise 1, whose one-sided hypothesis is not enough (a constant negative variable satisfies it with an unbounded moment), and even the two-sided tail integrates only to 2m−12m - 12m−1; the claim ≤m\le m≤m is nevertheless true, for instance by writing eaΔ2=Eg[e2a gΔ]e^{a\Delta^2} = \mathbb{E}_g[e^{\sqrt{2a}\,g\Delta}]eaΔ2=Eg​[e2a​gΔ] for a standard Gaussian ggg and applying Hoeffding's lemma to the sample mean, which gives ES[e2(m−1)Δ2]≤(1−(m−1)/m)−1/2=m\mathbb{E}_S[e^{2(m-1)\Delta^2}] \le (1 - (m-1)/m)^{-1/2} = \sqrt mES​[e2(m−1)Δ2]≤(1−(m−1)/m)−1/2=m​. Theorem 31.1 then follows the book: Markov's inequality on ef(S)e^{f(S)}ef(S) with f(S)=sup⁡Q(2(m−1)Eh∼QΔ(h)2−D(Q∥P))f(S) = \sup_Q(2(m-1)\mathbb{E}_{h \sim Q}\Delta(h)^2 - D(Q\|P))f(S)=supQ​(2(m−1)Eh∼Q​Δ(h)2−D(Q∥P)), the change of measure (31.2) by Jensen's inequality for ln⁡\lnln applied to the density dQ/dPdQ/dPdQ/dP (the Donsker–Varadhan inequality, which needs Q≪PQ \ll PQ≪P and an integrable log-density), the exchange of expectations (31.4) by Fubini, the moment bound, and finally Jensen for x2x^2x2 in (31.6); formally the supremum over all posteriors is handled by proving the bound for each QQQ on the event {S:Eh∼P[e2(m−1)Δ(h)2]≤m/δ}\{S : \mathbb{E}_{h \sim P}[e^{2(m-1)\Delta(h)^2}] \le m/\delta\}{S:Eh∼P​[e2(m−1)Δ(h)2]≤m/δ}, whose complement has probability at most δ\deltaδ by Markov, which sidesteps any measurability question about fff. Exercise 2 is the theorem with Q=δhQ = \delta_hQ=δh​, for which D(Q∥P)=ln⁡∣H∣D(Q\|P) = \ln|H|D(Q∥P)=ln∣H∣.

Formalization scope

The class is an arbitrary measurable space, priors and posteriors are probability measures on it, and Q(h)/P(h)Q(h)/P(h)Q(h)/P(h) is Mathlib's Radon–Nikodym derivative; the divergence is a Bochner integral, so the theorem quantifies over posteriors Q≪PQ \ll PQ≪P whose log-density is QQQ-integrable, exactly the posteriors with a finite divergence, for which the bound has content, and no junk value can make a case false. Posteriors may depend on the sample: the statement bounds the outer measure of the set of samples for which some admissible QQQ violates the bound. The loss is jointly measurable so that h↦LD(h)h \mapsto L_D(h)h↦LD​(h) and (S,h)↦LS(h)(S, h) \mapsto L_S(h)(S,h)↦LS​(h) are measurable and the Gibbs risks are genuine integrals; m≥2m \ge 2m≥2 is forced by the denominator 2(m−1)2(m-1)2(m−1). The moment bound is stated as the claim the proof needs rather than as Exercise 1, whose printed hypothesis is insufficient; the item text records this. Exercise 2 is stated as a simultaneous bound over the finite class, the form in which the theorem delivers it.

Not stated: Remark 31.1 (a learning rule, not a theorem), Exercise 1 as printed, and the second part of Exercise 2.

Selected references

  • S. Shalev-Shwartz, S. Ben-David, Understanding Machine Learning: From Theory to Algorithms, Cambridge University Press, 2014, Chapter 31. doi:10.1017/CBO9781107298019
  • D. A. McAllester, Some PAC-Bayesian theorems, Machine Learning 37, 1999. doi:10.1023/A:1007618624809
  • D. A. McAllester, PAC-Bayesian stochastic model selection, Machine Learning 51, 2003. doi:10.1023/A:1021840411064
  • A. Maurer, A note on the PAC Bayesian theorem, arXiv:cs/0411099, 2004.
  • M. Seeger, PAC-Bayesian generalisation error bounds for Gaussian process classification, Journal of Machine Learning Research 3, 2002.
  • J. Langford, J. Shawe-Taylor, PAC-Bayes and margins, NIPS 2002.
6 thms4 active usersReviewed
🏆Completed
Machine Learning·Captain: Lucas

The Principles of Deep Learning Theory I: Wick's Theorem for Gaussian IntegralsTextbook

Motivation

The book The Principles of Deep Learning Theory by D. A. Roberts and S. Yaida (arXiv:2106.10165) develops a perturbative, physics-style theory of wide deep neural networks. Every later computation in the book, from the statistics of preactivations at initialization to the neural tangent kernel, reduces to Gaussian expectations of polynomials. The single tool that evaluates these expectations is Wick's theorem (eq. 1.45), which the authors ask the reader to "put a box around". This mission formalizes Chapter 1, §1.1 (Gaussian Integrals), culminating in that theorem. It is the first mission of a series formalizing the book, all in the Lean namespace DeepLearningTheory.

Setting

Fix N≥0N \ge 0N≥0 and a symmetric positive-definite real N×NN\times NN×N matrix K=(Kμν)K = (K_{\mu\nu})K=(Kμν​), the covariance. Write Kμν=(K−1)μνK^{\mu\nu} = (K^{-1})_{\mu\nu}Kμν=(K−1)μν​ for its inverse. The zero-mean multivariable Gaussian distribution with covariance KKK (eq. 1.31) has density on RN\mathbb{R}^NRN

p(z)=1∣2πK∣exp⁡(−12∑μ,ν=1NzμKμνzν),∣2πK∣=det⁡(2πK)=(2π)Ndet⁡K,p(z) = \frac{1}{\sqrt{|2\pi K|}} \exp\Big(-\tfrac12 \sum_{\mu,\nu=1}^N z_\mu K^{\mu\nu} z_\nu\Big), \qquad |2\pi K| = \det(2\pi K) = (2\pi)^N \det K,p(z)=∣2πK∣​1​exp(−21​μ,ν=1∑N​zμ​Kμνzν​),∣2πK∣=det(2πK)=(2π)NdetK,

and the expectation of an observable FFF is E[F(z)]=∫dNz p(z)F(z)\mathbb{E}[F(z)] = \int d^N z\, p(z) F(z)E[F(z)]=∫dNzp(z)F(z) (eqs. 1.36, 1.46), an integral against Lebesgue measure. A pairing of the labels 1,…,2m1,\dots,2m1,…,2m is a partition into mmm unordered pairs; equivalently, a fixed-point-free involution σ\sigmaσ of {1,…,2m}\{1,\dots,2m\}{1,…,2m} pairing aaa with σ(a)\sigma(a)σ(a). There are (2m−1)!!(2m-1)!!(2m−1)!! pairings.

Formalization targets

Goal: Wick's theorem (eq. 1.45)

For all indices μ1,…,μ2m∈{1,…,N}\mu_1,\dots,\mu_{2m} \in \{1,\dots,N\}μ1​,…,μ2m​∈{1,…,N} (repetitions allowed),

E[zμ1⋯zμ2m]=∑pairings σ ∏a<σ(a)Kμaμσ(a).\mathbb{E}[z_{\mu_1}\cdots z_{\mu_{2m}}] = \sum_{\text{pairings } \sigma}\ \prod_{a<\sigma(a)} K_{\mu_a \mu_{\sigma(a)}}.E[zμ1​​⋯zμ2m​​]=pairings σ∑​ a<σ(a)∏​Kμa​μσ(a)​​.

Milestones

  1. Eq. (1.30) — the normalization ∫dNz e−12z⊤K−1z=∣2πK∣\int d^N z\, e^{-\frac12 z^\top K^{-1} z} = \sqrt{|2\pi K|}∫dNze−21​z⊤K−1z=∣2πK∣​.
  2. Eq. (1.41) — the generating function ∫dNz e−12z⊤K−1z+J⋅z=∣2πK∣ e12J⊤KJ\int d^N z\, e^{-\frac12 z^\top K^{-1} z + J\cdot z} = \sqrt{|2\pi K|}\, e^{\frac12 J^\top K J}∫dNze−21​z⊤K−1z+J⋅z=∣2πK∣​e21​J⊤KJ for every source J∈RNJ \in \mathbb{R}^NJ∈RN.
  3. Odd moments vanish (p. 22, discussion after eq. 1.42).
  4. Eq. (1.43) — E[zμ1zμ2]=Kμ1μ2\mathbb{E}[z_{\mu_1} z_{\mu_2}] = K_{\mu_1\mu_2}E[zμ1​​zμ2​​]=Kμ1​μ2​​.
  5. Eq. (1.44) — E[zμ1zμ2zμ3zμ4]=Kμ1μ2Kμ3μ4+Kμ1μ3Kμ2μ4+Kμ1μ4Kμ2μ3\mathbb{E}[z_{\mu_1}z_{\mu_2}z_{\mu_3}z_{\mu_4}] = K_{\mu_1\mu_2}K_{\mu_3\mu_4} + K_{\mu_1\mu_3}K_{\mu_2\mu_4} + K_{\mu_1\mu_4}K_{\mu_2\mu_3}E[zμ1​​zμ2​​zμ3​​zμ4​​]=Kμ1​μ2​​Kμ3​μ4​​+Kμ1​μ3​​Kμ2​μ4​​+Kμ1​μ4​​Kμ2​μ3​​.

Significance

Wick's theorem (also known as Isserlis' theorem, 1918) is used throughout the book: the Gaussianity of the first layer (Chapter 4), the four-point correlators of deep linear networks (Chapter 3), and all perturbative expansions around the infinite-width limit are evaluated by Wick contractions. A machine-checked version stated in the book's own density-based conventions gives later missions of this series a dependable foundation. The result is classical; the work remaining is its formalization in exactly this form. A related statement for general (possibly degenerate) Gaussian measures already exists on the platform as FeynmanWick.wick_theorem_gaussian; the present mission keeps the book's formulation via the explicit density and a positive-definite covariance.

Difficulty

The book's derivation differentiates the generating function (1.41) 2m2m2m times at J=0J=0J=0, which requires justifying differentiation under the integral sign and organizing the combinatorics of the product rule into pairings. Both the analytic step (dominated convergence for Gaussian tails, in NNN dimensions) and the combinatorial step (a bijection between surviving terms and pairings, with the 1/(2mm!)1/(2^m m!)1/(2mm!) normalization) are routine on paper but need care in Lean. Computing the normalization (1.30) itself requires diagonalizing KKK and a linear change of variables.

Formalization scope

  • Vectors are functions Fin N → ℝ with Lebesgue measure; KKK is a Matrix (Fin N) (Fin N) ℝ with K.PosDef (which includes symmetry). The inverse is Mathlib's K⁻¹.
  • gaussianDensity K z and gaussExpect K F are exactly the book's p(z)p(z)p(z) and E[F]\mathbb{E}[F]E[F]; gaussQuadForm K z is ∑zμKμνzν\sum z_\mu K^{\mu\nu} z_\nu∑zμ​Kμνzν​.
  • pairings (2m) is the finite set of fixed-point-free involutions of Fin (2m); the product is over labels aaa with a<σ(a)a<\sigma(a)a<σ(a), so each pair contributes once.
  • Positive definiteness is assumed exactly as in the book; no statement is vacuous since, e.g., K=IK = IK=I satisfies every hypothesis.

Selected references

  • D. A. Roberts, S. Yaida (with B. Hanin), The Principles of Deep Learning Theory, Cambridge University Press, 2022. arXiv:2106.10165
  • L. Isserlis, On a formula for the product-moment coefficient of any order of a normal frequency distribution in any number of variables, Biometrika 12 (1918). doi:10.1093/biomet/12.1-2.134
7 thms4 active usersReviewed
🏆Completed
CombinatoricsMachine LearningStatistics·Captain: naimengye

An Introduction to Computational Learning Theory III: The Vapnik-Chervonenkis Dimension, ε-Nets and Sample ComplexityTextbook

Motivation

Chapter 3 of Kearns and Vazirani, An Introduction to Computational Learning Theory (MIT Press, 1994, doi:10.7551/mitpress/3897.001.0001), asks how many random examples suffice to learn a concept from an infinite class. The cardinality bound of Occam's Razor is useless there, yet the rectangle game of Chapter 1 shows that some infinite classes are learnable from a finite sample. The answer is the Vapnik–Chervonenkis dimension: the size of the largest set on which the class realizes every labeling. Sauer's lemma says that a class of VC dimension ddd realizes only Φd(m)=∑i≤d(mi)≤(em/d)d\Phi_d(m) = \sum_{i \le d}\binom{m}{i} \le (em/d)^dΦd​(m)=∑i≤d​(im​)≤(em/d)d labelings on any mmm points, polynomially many rather than 2m2^m2m, and the ε-net theorem of Blumer, Ehrenfeucht, Haussler and Warmuth turns this into a sample bound: a consistent hypothesis from a class of VC dimension ddd is probably approximately correct once mmm is of order (1/ϵ)(log⁡(1/δ)+dlog⁡(1/ϵ))(1/\epsilon)(\log(1/\delta) + d\log(1/\epsilon))(1/ϵ)(log(1/δ)+dlog(1/ϵ)). A matching lower bound shows that Ω(d/ϵ)\Omega(d/\epsilon)Ω(d/ϵ) examples are necessary. The chapter thus gives a single combinatorial parameter that characterizes, up to a logarithmic factor, the sample complexity of learning any class in the distribution-free model.

Setting

For a class CCC of concepts X→{0,1}X \to \{0,1\}X→{0,1} and a finite S⊆XS \subseteq XS⊆X, ΠC(S)\Pi_C(S)ΠC​(S) is the set of dichotomies of SSS realized by CCC; SSS is shattered if all 2∣S∣2^{|S|}2∣S∣ are realized; VCD(C)\mathrm{VCD}(C)VCD(C) is the supremum of the sizes of shattered sets, possibly ∞\infty∞; ΠC(m)\Pi_C(m)ΠC​(m) is the largest ∣ΠC(S)∣|\Pi_C(S)|∣ΠC​(S)∣ over ∣S∣=m|S| = m∣S∣=m; and Φd(m)\Phi_d(m)Φd​(m) is defined by Φd(m)=Φd(m−1)+Φd−1(m−1)\Phi_d(m) = \Phi_d(m-1) + \Phi_{d-1}(m-1)Φd​(m)=Φd​(m−1)+Φd−1​(m−1), Φd(0)=Φ0(m)=1\Phi_d(0) = \Phi_0(m) = 1Φd​(0)=Φ0​(m)=1. For a target ccc the error regions are c Δ hc \,\Delta\, hcΔh for hhh in the hypothesis class, and a set of points is an ε-net if it meets every error region of weight at least ϵ\epsilonϵ under the target distribution DDD. Samples, their product law, consistency and the error of a hypothesis are those of Mission I.

Formalization targets

Goal: Theorems 3.3 and 3.4

Let HHH be a class of VC dimension at most ddd, well-behaved for the target ccc (the double-sample event of the proof is null-measurable), and m≥8/ϵm \ge 8/\epsilonm≥8/ϵ. The points of a random sample of mmm examples of a target ccc fail to be an ε-net for the error regions {c Δ h:h∈H}\{c \,\Delta\, h : h \in H\}{cΔh:h∈H} with probability at most

2 Φd(2m) 2−ϵm/2,2\,\Phi_d(2m)\,2^{-\epsilon m/2},2Φd​(2m)2−ϵm/2,

so any algorithm that outputs a hypothesis in HHH consistent with its sample has error greater than ϵ\epsilonϵ with at most that probability; with m≥(4/ϵ)log⁡2(2/δ)m \ge (4/\epsilon)\log_2(2/\delta)m≥(4/ϵ)log2​(2/δ) and m≥(8d/ϵ)log⁡2(13/ϵ)m \ge (8d/\epsilon)\log_2(13/\epsilon)m≥(8d/ϵ)log2​(13/ϵ) the probability is at most δ\deltaδ; and, if HHH is nonempty, every class contained in HHH for whose targets HHH is well-behaved is PAC learnable using HHH.

Milestones

Lemma 3.1 (Sauer's lemma, ΠC(m)≤Φd(m)\Pi_C(m) \le \Phi_d(m)ΠC​(m)≤Φd​(m)); Lemma 3.2 (Φd(m)=∑i≤d(mi)\Phi_d(m) = \sum_{i \le d}\binom{m}{i}Φd​(m)=∑i≤d​(im​)); the polynomial bound Φd(m)≤(em/d)d\Phi_d(m) \le (em/d)^dΦd​(m)≤(em/d)d of p. 57; Theorem 3.5 (the Ω(d/ϵ)\Omega(d/\epsilon)Ω(d/ϵ) lower bound, in the two explicit forms of its proof).

Significance

Theorem 3.3 is the fundamental theorem of PAC learning: it replaces log⁡∣H∣\log|H|log∣H∣ in Occam's Razor by the VC dimension and thereby covers rectangles, halfspaces, polygons, neural networks with a fixed architecture, and every class whose dichotomies grow polynomially. Its proof, the double sample and random partition argument, is the origin of symmetrization in empirical process theory. Sauer's lemma is a cornerstone of extremal combinatorics with independent proofs by Sauer, Shelah and Vapnik–Chervonenkis, and the lower bound of Theorem 3.5 shows that the upper bound is tight to within log⁡(1/ϵ)\log(1/\epsilon)log(1/ϵ), so the VC dimension genuinely characterizes sample complexity. None of these is machine-checked. Formalizing them puts on the platform the VC dimension, the growth function and the ε-net theorem with explicit constants, stated on the same sample law as the rest of this series, and the first information-theoretic lower bound for learning.

Difficulty

Sauer's lemma is a double induction on ddd and mmm through the auxiliary class C′C'C′ of dichotomies whose two extensions to a distinguished point are both realized, which needs care with the identification of dichotomies of SSS and of S∖{x}S \setminus \{x\}S∖{x}. The ε-net theorem needs: the reduction Pr⁡[A]≤2Pr⁡[B]\Pr[A] \le 2\Pr[B]Pr[A]≤2Pr[B] from a failed ε-net on the first half to a region hit at least ϵm/2\epsilon m/2ϵm/2 times by the second half, which is a Chebyshev bound on a binomial variable and is where m≥8/ϵm \ge 8/\epsilonm≥8/ϵ enters; the exchangeability of the 2m2m2m draws with a random partition into two halves; the counting bound (mℓ)/(2mℓ)≤2−ℓ\binom{m}{\ell}/\binom{2m}{\ell} \le 2^{-\ell}(ℓm​)/(ℓ2m​)≤2−ℓ; and Sauer's lemma applied to the error regions, whose growth function equals that of HHH. The explicit constants require the numerical inequality 2(2em/d)d2−ϵm/2≤δ2(2em/d)^d 2^{-\epsilon m/2} \le \delta2(2em/d)d2−ϵm/2≤δ under the two stated conditions. The lower bound is a probabilistic argument with a random target: conditional on the sample, the labels of unseen points are fair coins, so the number of errors on them is binomial and exceeds half its range with probability at least 1/21/21/2; the refined bound scales this construction to a region of weight 16ϵ16\epsilon16ϵ and uses Markov's inequality to bound the number of draws landing in it. Measurability of the failure sets is avoided by stating outer-measure bounds, except for the double-sample event, which the proof integrates: it is assumed null-measurable (the well-behavedness of Blumer et al., without which the theorem is false for a class of VC dimension 111 on ω1\omega_1ω1​). For the lower bound it is avoided by working over a finitely supported distribution on a space with measurable singletons.

Formalization scope

The VC dimension is a supremum in N∪{∞}\mathbb{N} \cup \{\infty\}N∪{∞}, the growth function a supremum in N\mathbb{N}N (bounded by 2m2^m2m), and Φd\Phi_dΦd​ the book's recurrence, with its closed form and polynomial bound stated as separate theorems. The goal is stated for a hypothesis class HHH (Theorem 3.4), Theorem 3.3 being the case C=HC = HC=H; it carries the exact bound of the proof, the explicit constants of Blumer et al. in place of the book's c0c_0c0​, the requirement m≥8/ϵm \ge 8/\epsilonm≥8/ϵ of the proof's Chebyshev step, measurability of the hypotheses and the target, well-behavedness of HHH for the target, and 0<ϵ,δ<10 < \epsilon, \delta < 10<ϵ,δ<1. PAC learnability of the subclasses of HHH needs HHH nonempty, since no algorithm outputs hypotheses in the empty class. The lower bound is stated for every deterministic learning function, on an instance space with measurable singletons, for a class shattering some set of d≥1d \ge 1d≥1 points, with the explicit constants derived in the proof sketch (m≤d/2m \le d/2m≤d/2: error ≥1/8\ge 1/8≥1/8 with probability ≥1/2\ge 1/2≥1/2; ϵ≤1/16\epsilon \le 1/16ϵ≤1/16 and m≤(d−1)/(64ϵ)m \le (d-1)/(64\epsilon)m≤(d−1)/(64ϵ): error >ϵ> \epsilon>ϵ with probability ≥1/4\ge 1/4≥1/4). Running time is not modelled. The composition bound for layered networks (Theorems 3.6 and 3.7) is not stated.

Trivializing readings are excluded: the ε-net event ranges over every hypothesis of HHH, the failure bounds are uniform over all consistent learners, and the lower bound holds for every learning function. Welcome contributions: Sauer's lemma, the closed form and the (em/d)d(em/d)^d(em/d)d bound, the random-partition counting lemma, and the binomial median inequality used in the lower bound.

Selected references

  • M. J. Kearns, U. V. Vazirani, An Introduction to Computational Learning Theory, MIT Press, 1994, Chapter 3. doi:10.7551/mitpress/3897.001.0001
  • A. Blumer, A. Ehrenfeucht, D. Haussler, M. K. Warmuth, Learnability and the Vapnik–Chervonenkis dimension, Journal of the ACM 36(4), 1989. doi:10.1145/76359.76371
  • V. N. Vapnik, A. Ya. Chervonenkis, On the uniform convergence of relative frequencies of events to their probabilities, Theory of Probability and its Applications 16(2), 1971. doi:10.1137/1116025
  • N. Sauer, On the density of families of sets, Journal of Combinatorial Theory, Series A 13(1), 1972. doi:10.1016/0097-3165(72)90019-2
  • A. Ehrenfeucht, D. Haussler, M. Kearns, L. Valiant, A general lower bound on the number of examples needed for learning, Information and Computation 82(3), 1989. doi:10.1016/0890-5401(89)90002-3
8 thms4 active usersReviewed
🏆Completed
Bandit AlgorithmsOperations ResearchOptimization·Captain: naimengye

Multi-armed Bandit Allocation Indices I: The Gittins Index, Optimal Stopping and MonotonicityTextbook

Motivation

A decision-maker has nnn projects, each a Markov reward process, and at every decision time may advance exactly one of them; the others stay frozen. Which project to advance so as to maximize the expected total discounted reward? Posed as a dynamic program the problem has a state space that is the product of the nnn state spaces, and the size of that product defeats every general method. The index theorem of Gittins and Jones (1974) says the dynamic program is solved exactly by an index policy: there is a real number ν(B,x)\nu(B, x)ν(B,x), computable for each bandit process BBB from its own data and its own current state xxx, such that always advancing a process of greatest index is optimal. Chapter 2 of Gittins, Glazebrook and Weber, Multi-armed Bandit Allocation Indices (2nd ed., doi:10.1002/9780470980033), introduces the index, proves the theorem three times, and works out the properties of the index that the rest of the book, from jobs and superprocesses to restless bandits, is built on: which stopping time attains it, how it is computed, how it moves with the discount factor, and when it collapses to the myopic rule.

The index theorem itself is already on the platform, proved, as BanditAlgorithm.gittins_index_theorem in the Bandit Algorithms series (Lattimore and Szepesvári, Theorem 35.9). This mission cites it and formalizes what Chapter 2 establishes around it.

Setting

A bandit process BBB (§2.3–2.4) is a Markov reward process on a countable state space EEE: transition probabilities P(y∣x)P(y \mid x)P(y∣x), a bounded reward r(x)r(x)r(x) received each time the continuation control is applied in state xxx, and a discount factor a∈(0,1)a \in (0, 1)a∈(0,1); the freeze control leaves the state unchanged and yields nothing. The law of the process started at xxx is Px\mathbb{P}_xPx​ and x(t)x(t)x(t) is its state at process time t=0,1,2,…t = 0, 1, 2, \dotst=0,1,2,…. A stopping time τ\tauτ is a past-measurable rule for switching from continuation to freezing, taking values in {1,2,… }∪{∞}\{1, 2, \dots\} \cup \{\infty\}{1,2,…}∪{∞}. For such τ\tauτ, Rτ(B,x)=Ex[∑t<τatr(x(t))]R_\tau(B, x) = \mathbb{E}_x[\sum_{t < \tau} a^t r(x(t))]Rτ​(B,x)=Ex​[∑t<τ​atr(x(t))] is the expected discounted reward and Wτ(B,x)=Ex[∑t<τat]W_\tau(B, x) = \mathbb{E}_x[\sum_{t < \tau} a^t]Wτ​(B,x)=Ex​[∑t<τ​at] the expected discounted time; their ratio ντ(B,x)\nu_\tau(B, x)ντ​(B,x) (2.7) is the equivalent constant reward rate of that portion of BBB. The Gittins index is

ν(B,x)=sup⁡τ>0Rτ(B,x)Wτ(B,x)(2.6)\nu(B, x) = \sup_{\tau > 0} \frac{R_\tau(B, x)}{W_\tau(B, x)} \tag{2.6}ν(B,x)=τ>0sup​Wτ​(B,x)Rτ​(B,x)​(2.6)

and, equivalently, the fair charge (2.5): the greatest rent λ\lambdaλ per period for which continuing BBB for one or more periods, paying λ\lambdaλ each period, can be done without expected loss. A simple family of alternative bandit processes (SFABP) is nnn such processes with a common discount factor, one of which is continued at each decision time; an index policy continues a process of greatest index. In the Lean development the single-arm model is the platform's (GittinsIndex): the chain law is built by the Ionescu–Tulcea construction, stopping times are adapted N∪{∞}\mathbb{N}\cup\{\infty\}N∪{∞}-valued maps on trajectories, and gittinsIndex P r a x is (2.6).

Formalization targets

Goal: Lemma 2.2, the optimal stopping set

The supremum in (2.6) is attained. For the initial state ξ\xiξ, the attaining stopping rule may be taken to be "stop at the first time t≥1t \ge 1t≥1 at which the state lies in Σ0\Sigma_0Σ0​" for any set Σ0\Sigma_0Σ0​ with

{x:ν(B,x)<ν(B,ξ)}⊆Σ0⊆{x:ν(B,x)≤ν(B,ξ)},\{x : \nu(B, x) < \nu(B, \xi)\} \subseteq \Sigma_0 \subseteq \{x : \nu(B, x) \le \nu(B, \xi)\},{x:ν(B,x)<ν(B,ξ)}⊆Σ0​⊆{x:ν(B,x)≤ν(B,ξ)},

and every such rule has ντ(B,ξ)=ν(B,ξ)\nu_\tau(B, \xi) = \nu(B, \xi)ντ​(B,ξ)=ν(B,ξ).

Milestones

Theorem 2.1 as a reference to the proved platform theorem; Eq. (2.5), the fair-charge characterization of the index; the restart-in-state formulation of §2.6.4 as the convergence of the Katehakis–Veinott value iteration; Theorem 2.3, monotonicity in the discount factor; Lemma 2.4, the interchange of two bandit portions; Propositions 2.5–2.8, the monotone-index cases in which the index is the immediate reward, or is attained only after the first step, or only at τ=∞\tau = \inftyτ=∞.

Significance

Lemma 2.2 is the working form of the index: it turns the supremum over all stopping times into a specific rule, the first time the index falls below its starting value, and it is what the interchange proof of §2.7, the modified-forwards-induction policies of §2.6.6, the monotone-index propositions of §2.11 and the treatment of jobs in Chapter 3 all use. The fair-charge form (2.5) is the prevailing-charge proof of the theorem (Weber 1992) and the interpretation that carries over to superprocesses and restless bandits. The restart formulation is how indices are computed in practice (Katehakis and Veinott 1987), and Theorem 2.3 is the first of the comparative statics used throughout Chapters 7 and 8. Lemma 2.4 is the elementary inequality behind the original proof of Gittins and Jones.

Of these, only the index theorem has a machine-checked proof today. Formalizing the rest gives the platform the index as a usable object: a characterization of the optimal stopping rule, a convergent algorithm for it, and the monotonicity facts, all stated against the existing model so that every later mission of this series and every future use of the L&S model can build on them.

Difficulty

The obvious first move for Lemma 2.2, "take the stopping time that achieves the supremum", is what has to be proved: the supremum is over an uncountable family, and attainment comes from the optimal stopping problem with charge λ=ν(B,ξ)\lambda = \nu(B, \xi)λ=ν(B,ξ), whose value function satisfies φ(x)=max⁡{0,r(x)−λ+aE[φ(x(1))∣x(0)=x]}\varphi(x) = \max\{0, r(x) - \lambda + a\mathbb{E}[\varphi(x(1)) \mid x(0) = x]\}φ(x)=max{0,r(x)−λ+aE[φ(x(1))∣x(0)=x]}, together with the fact that its optimal stopping set is characterized by the strict and non-strict inequalities λ>ν(B,x)\lambda > \nu(B, x)λ>ν(B,x) and λ≥ν(B,x)\lambda \ge \nu(B, x)λ≥ν(B,x). That last step is the content: it identifies the local decision "stop or continue" with a comparison of indices, which is why any set between the two level sets works. On the platform's model this requires the dynamic-programming theory of discounted optimal stopping on a countable space (Theorem 2.10 of the book), the identification of fairChargeProfit with that value function, and the strong Markov property of markovChainMeasure at a trajectory stopping time. Theorem 2.3 needs randomized stopping times (a geometric kill) and the fact that they do not enlarge the supremum. The restart iteration is monotone and bounded but its operator is not a contraction in the restart value, so its limit has to be identified with the restart problem's value directly; that value is max⁡(0,ν/(1−a))\max(0, \nu/(1-a))max(0,ν/(1−a)), not ν/(1−a)\nu/(1-a)ν/(1−a), because restarting forever is free.

Formalization scope

The state space is a countable type with measurable singletons, so every subset is measurable; rewards are bounded; a∈(0,1)a \in (0, 1)a∈(0,1). The chain law, stopping times, the discounted stopped sums and the index are the platform's, unchanged. Expectations are Bochner integrals; with bounded rewards they are finite and no total-function default value enters. IsPositiveStoppingTime fixes τ≥1\tau \ge 1τ≥1 everywhere, so Wτ≥1W_\tau \ge 1Wτ​≥1 and the ratio (2.7) is a genuine quotient. The stopping rule of Lemma 2.2 is the hitting time from time 111 of a set, with ∞\infty∞ when the set is never hit. The fair-charge profit is a real supremum over the nonempty bounded family of positive stopping times, and (2.5) is stated with the outer supremum over {λ:profit(λ)≥0}\{\lambda : \text{profit}(\lambda) \ge 0\}{λ:profit(λ)≥0}, a nonempty set bounded above. The restart iteration is stated as a limit, with the value max⁡(0,ν(B,ξ)/(1−a))\max(0, \nu(B, \xi)/(1-a))max(0,ν(B,ξ)/(1−a)): for a nonnegative index this is the book's ν/(1−a)\nu/(1-a)ν/(1−a), and the maximum is forced by a one-state example with negative reward. The propositions' hypotheses are almost-sure events under Px\mathbb{P}_xPx​, written as events of measure one.

Two trivializing readings are excluded: the index is never taken over all N∪{∞}\mathbb{N}\cup\{\infty\}N∪{∞}-valued maps but over adapted stopping times, and the stopping set of the goal is not restricted to the two extreme level sets. Contributions welcome: the optimal-stopping dynamic program on markovChainMeasure (value iteration, the strong Markov property at a stopping time), the equivalence of randomized and non-randomized stopping times for the supremum, and the monotone convergence of the restart iteration.

Selected references

  • J. Gittins, K. Glazebrook, R. Weber, Multi-armed Bandit Allocation Indices, 2nd ed., Wiley, 2011, Chapter 2. doi:10.1002/9780470980033
  • J. C. Gittins, D. M. Jones, A dynamic allocation index for the sequential design of experiments, in Progress in Statistics (Gani, ed.), North-Holland, 1974.
  • R. Weber, On the Gittins index for multiarmed bandits, Annals of Applied Probability 2(4), 1992. doi:10.1214/aoap/1177005588
  • M. N. Katehakis, A. F. Veinott, The multi-armed bandit problem: decomposition and computation, Mathematics of Operations Research 12(2), 1987. doi:10.1287/moor.12.2.262
  • T. Lattimore, C. Szepesvári, Bandit Algorithms, Cambridge University Press, 2020, Chapter 35. doi:10.1017/9781108571401
12 thms4 active usersReviewed
🏆Completed
Machine LearningRandom Matrix TheoryStatistics·Captain: mikedeng1

High-Dimensional Probability VI: The Hanson-Wright InequalityTextbook

Motivation

Sums of independent random variables are well understood: Bernstein's inequality and its relatives give sharp, non-asymptotic tail bounds for ∑iaiXi\sum_i a_i X_i∑i​ai​Xi​ whenever the XiX_iXi​ are independent and light-tailed. Many quantities that arise in high-dimensional statistics and random matrix theory, however, are not linear but quadratic in an independent sample — the squared norm of a random vector after a linear transformation, a quadratic-form test statistic, the diagonal of a sample covariance matrix, or the number of edges cut by a random partition in a random graph. A quadratic form X⊤AX=∑i,jAijXiXjX^\top A X = \sum_{i,j} A_{ij} X_i X_jX⊤AX=∑i,j​Aij​Xi​Xj​ is a sum with dependent terms: XiXjX_iX_jXi​Xj​ and XiXkX_iX_kXi​Xk​ share the factor XiX_iXi​, so classical sum-of-independent-variables tools do not apply directly.

The Hanson-Wright inequality, first obtained by Hanson and Wright (1971) for sub-gaussian variables and later sharpened and popularized in this form by Rudelson and Vershynin (2013, "Hanson-Wright inequality and sub-gaussian concentration," Electronic Communications in Probability), closes this gap: it gives a concentration inequality for X⊤AXX^\top A XX⊤AX around its mean with the same two-regime (sub-gaussian near the center, sub-exponential in the tail) shape as Bernstein's inequality for linear sums. It is now a standard tool wherever quadratic statistics of independent data are analyzed: covariance estimation, compressed sensing, randomized numerical linear algebra, and the analysis of random matrices more broadly draw on it routinely.

Setting

Fix a probability space and let X=(X1,…,Xn)X = (X_1, \dots, X_n)X=(X1​,…,Xn​) be a random vector whose coordinates X1,…,XnX_1, \dots, X_nX1​,…,Xn​ are independent, mean zero, and sub-gaussian: each XiX_iXi​ has a finite sub-gaussian (Orlicz ψ2\psi_2ψ2​) norm ∥Xi∥ψ2\|X_i\|_{\psi_2}∥Xi​∥ψ2​​, the smallest t>0t > 0t>0 with Eexp⁡(Xi2/t2)≤2\mathbb E \exp(X_i^2/t^2) \le 2Eexp(Xi2​/t2)≤2. Write K=max⁡i∥Xi∥ψ2K = \max_i \|X_i\|_{\psi_2}K=maxi​∥Xi​∥ψ2​​.

Let A=(Aij)i,j=1nA = (A_{ij})_{i,j=1}^nA=(Aij​)i,j=1n​ be an n×nn \times nn×n real matrix, with no constraint on its diagonal, and form the quadratic form

X⊤AX=∑i,j=1nAijXiXj.X^\top A X = \sum_{i,j=1}^n A_{ij} X_i X_j.X⊤AX=i,j=1∑n​Aij​Xi​Xj​.

Two matrix norms measure the size of AAA: the Frobenius norm ∥A∥F=(∑i,jAij2)1/2\|A\|_F = \bigl(\sum_{i,j} A_{ij}^2\bigr)^{1/2}∥A∥F​=(∑i,j​Aij2​)1/2 (the Euclidean norm of AAA's entries) and the operator (spectral) norm ∥A∥=sup⁡∥x∥2=1∥Ax∥2\|A\| = \sup_{\|x\|_2=1} \|Ax\|_2∥A∥=sup∥x∥2​=1​∥Ax∥2​ (the largest singular value of AAA). Always ∥A∥≤∥A∥F≤n ∥A∥\|A\| \le \|A\|_F \le \sqrt{n}\,\|A\|∥A∥≤∥A∥F​≤n​∥A∥, so the two norms can differ by a factor as large as n\sqrt nn​ — the gap between them is exactly what produces the inequality's two regimes below.

Formalization targets

Goal — Theorem 6.2.1 (Hanson-Wright inequality)

P{ ∣X⊤AX−E X⊤AX∣≥t }  ≤  2exp⁡ ⁣[−cmin⁡ ⁣(t2K4∥A∥F2, tK2∥A∥)]for every t≥0,P\bigl\{\, |X^\top A X - \mathbb E\, X^\top A X| \ge t \,\bigr\} \;\le\; 2 \exp\!\left[-c \min\!\left(\frac{t^2}{K^4 \|A\|_F^2},\ \frac{t}{K^2 \|A\|}\right)\right] \qquad \text{for every } t \ge 0,P{∣X⊤AX−EX⊤AX∣≥t}≤2exp[−cmin(K4∥A∥F2​t2​, K2∥A∥t​)]for every t≥0,

where c>0c > 0c>0 is an absolute constant, not depending on nnn, XXX, AAA, or ttt. Stating the constant only as "some absolute ccc" (rather than pinning it to a numeral) is deliberate: the book's own proof does not track a sharp value, and a goal that only asserts the shape of the bound survives any later improvement to ccc.

Significance

The result itself. Hanson-Wright turns a two-dimensional (in i,ji,ji,j) dependency structure into a one-dimensional concentration statement controlled by two scalar quantities, ∥A∥F\|A\|_F∥A∥F​ and ∥A∥\|A\|∥A∥. This is what makes it usable: a practitioner bounding a quadratic statistic need only compute these two norms, not analyze the joint dependency structure of {XiXj}\{X_iX_j\}{Xi​Xj​} directly. It specializes to Bernstein's inequality (Chapter 2 of this book) when AAA is diagonal, and it underlies non-asymptotic guarantees for covariance estimation, the Johnson-Lindenstrauss lemma via a different route, and the concentration of Lipschitz functions of sub-gaussian vectors.

Formalizing it. The published proof of Hanson-Wright is not a single argument but a chain of four steps: a decoupling reduction (Section 6.1), a direct computation for Gaussian chaos (Lemma 6.2.2), a comparison lemma extending the Gaussian bound to general sub-gaussian vectors via a replacement trick (Lemma 6.2.3), and a final assembly that separates the diagonal part (handled by Bernstein's inequality) from the off-diagonal part (handled by decoupling and comparison). This mission formalizes the goal theorem's statement and the first, most reusable link in that chain — the decoupling machinery of Section 6.1, which reduces the analysis of the dependent chaos X⊤AXX^\top A XX⊤AX to the independent-once-conditioned bilinear form X⊤AX′X^\top A X'X⊤AX′ — together with the chapter's separate contraction principle (Section 6.7), a general comparison tool for Rademacher-weighted sums used repeatedly in the book's later chaining chapters. The Gaussian MGF computation and the replacement-trick comparison lemma (Lemmas 6.2.2–6.2.3) are left as future milestones on top of this mission: they require Gaussian rotation invariance and the singular value decomposition of AAA, substantially more machinery than the milestones included here.

Difficulty

The obvious first idea — treat X⊤AX=∑i,jAijXiXjX^\top A X = \sum_{i,j} A_{ij}X_iX_jX⊤AX=∑i,j​Aij​Xi​Xj​ as if it were a sum of independent terms and apply Bernstein's inequality termwise — fails immediately: the terms AijXiXjA_{ij}X_iX_jAij​Xi​Xj​ for fixed iii are not independent across jjj, since they all share the factor XiX_iXi​. Decoupling (Theorem 6.1.1) is the non-obvious fix: it replaces the off-diagonal chaos by a bilinear form X⊤AX′X^\top A X'X⊤AX′ in an independent copy X′X'X′, which genuinely does become a sum of independent terms once one of the two vectors is conditioned on. The price is a universal constant factor of 444 and the restriction to diagonal-free matrices, which is exactly why the full Hanson-Wright proof must separate the diagonal contribution to E X⊤AX\mathbb E\,X^\top A XEX⊤AX (handled directly by Bernstein's inequality, Chapter 2) before decoupling can be applied to what remains.

Formalization scope

Random variables and vectors are real-valued on an explicit probability space (Ω,F,P)(\Omega, \mathcal F, P)(Ω,F,P). The sub-gaussian norm is HighDimProb.Concentration.subgaussianNorm, the Orlicz-ψ2\psi_2ψ2​-norm definition already published for this series (01-concentration), reused here as a reference item rather than redefined. K=max⁡i∥Xi∥ψ2K = \max_i \|X_i\|_{\psi_2}K=maxi​∥Xi​∥ψ2​​ is written as a finite supremum over the coordinate index, ⨆ i, subgaussianNorm P (X i); because the index type is always a Fintype (Fin n), this supremum is well-defined and, at the degenerate index n=0n=0n=0, reduces to a true (if content-free) instance of the inequality rather than a vacuous or false one. The Frobenius and operator norms of AAA are this mission's own frobeniusNorm and opNorm, stated directly from their defining formulas rather than through Mathlib's scoped matrix-norm typeclass instances, which are deliberately not global defaults (to avoid a diamond between the two norms) and so are unsuitable for a statement that needs both simultaneously. Every place the goal or a milestone integrates a quantity, that quantity is required Integrable, guarding against Mathlib's convention of returning 0 for the Bochner integral of a non-integrable function — without these hypotheses, a mean-zero or expectation hypothesis could hold vacuously, or a conclusion could hold trivially, for reasons having nothing to do with the book's mathematics.

The formalization deliberately does not restrict AAA's diagonal in the goal theorem: doing so would collapse Hanson-Wright to a restatement of Bernstein's inequality for the special case of a diagonal matrix, discarding the chapter's actual content, which is handling the off-diagonal, genuinely quadratic dependence between coordinates. The diagonal-free restriction does appear, correctly, in the Decoupling theorem (6.1.1), whose proof needs it.

Reusable beyond this mission: frobeniusNorm and opNorm are needed by any future chapter using matrix norms (Chapter 4's random matrix norms, Chapter 9's matrix deviation inequality); the decoupling theorem and convex decoupling lemma are the standard entry point for any later formalization of chaos concentration; the contraction principle is reused throughout the book's chaining chapters (7 and 8). Welcome contributions include the Gaussian MGF and comparison lemmas (6.2.2–6.2.3) needed to complete a full proof of the goal theorem, and the two-sided version of Bernstein's inequality needed for the diagonal part of that proof.

Selected references

  • R. Vershynin, High-Dimensional Probability: An Introduction with Applications in Data Science, Cambridge University Press, 2018. DOI: 10.1017/9781108231596.
  • D. L. Hanson, F. T. Wright, "A bound on tail probabilities for quadratic forms in independent random variables," Annals of Mathematical Statistics 42 (1971), 1079–1083.
  • M. Rudelson, R. Vershynin, "Hanson-Wright inequality and sub-gaussian concentration," Electronic Communications in Probability 18 (2013), no. 82, 1–9. https://arxiv.org/abs/1306.2872
7 thms4 active usersReviewed
🏆Completed
Machine LearningStatistics·Captain: mikedeng1

High-Dimensional Statistics I: Gaussian Concentration of Lipschitz FunctionsTextbook

Motivation

A recurring question in high-dimensional statistics is how tightly a scalar quantity built from many random inputs concentrates around its mean, even as the number of inputs grows without bound. Two classical answers organize the whole toolkit: martingale methods, which control a sum of dependent increments one conditional step at a time, and Gaussian-specific isoperimetry, which shows that essentially any regular (Lipschitz) function of a high-dimensional Gaussian vector concentrates as tightly as a single Gaussian coordinate, regardless of dimension. This mission formalizes one representative theorem from each line: the general martingale Bernstein bound (Wainwright, High-Dimensional Statistics, 2019, Theorem 2.19) and the Gaussian concentration of Lipschitz functions (Theorem 2.26), following Chapter 2 of the same book.

Setting

A random variable XXX with mean μ=E[X]\mu=\mathbb E[X]μ=E[X] is sub-Gaussian with parameter σ\sigmaσ (Definition 2.2) if E[eλ(X−μ)]≤eσ2λ2/2\mathbb E[e^{\lambda(X-\mu)}]\le e^{\sigma^2\lambda^2/2}E[eλ(X−μ)]≤eσ2λ2/2 for all λ∈R\lambda\in\mathbb Rλ∈R; it is sub-exponential with parameters (ν,α)(\nu,\alpha)(ν,α) (Definition 2.7, a strictly milder condition) if the same bound holds only for ∣λ∣<1/α|\lambda|<1/\alpha∣λ∣<1/α, with the convention 1/0=+∞1/0=+\infty1/0=+∞ so that α=0\alpha=0α=0 recovers the sub-Gaussian case exactly.

A sequence {Dk}k≥1\{D_k\}_{k\ge1}{Dk​}k≥1​, adapted to a filtration {Fk}\{\mathcal F_k\}{Fk​}, is a martingale difference sequence if each DkD_kDk​ is Fk\mathcal F_kFk​-measurable and E[Dk∣Fk−1]=0\mathbb E[D_k\mid\mathcal F_{k-1}]=0E[Dk​∣Fk−1​]=0. Such sequences arise throughout statistics via the Doob martingale construction: given a function fff of independent variables X1,…,XnX_1,\dots,X_nX1​,…,Xn​, setting Dk:=E[f(X)∣X1,…,Xk]−E[f(X)∣X1,…,Xk−1]D_k:=\mathbb E[f(X)\mid X_1,\dots,X_k]-\mathbb E[f(X)\mid X_1,\dots,X_{k-1}]Dk​:=E[f(X)∣X1​,…,Xk​]−E[f(X)∣X1​,…,Xk−1​] telescopes to f(X)−E[f(X)]=∑kDkf(X)-\mathbb E[f(X)]=\sum_k D_kf(X)−E[f(X)]=∑k​Dk​, converting a deviation question about f(X)f(X)f(X) into a martingale concentration question.

A function f:Rn→Rf:\mathbb R^n\to\mathbb Rf:Rn→R is LLL-Lipschitz with respect to the Euclidean norm if ∣f(x)−f(y)∣≤L∥x−y∥2|f(x)-f(y)|\le L\|x-y\|_2∣f(x)−f(y)∣≤L∥x−y∥2​ for all x,yx,yx,y (Eq. (2.38)).

Formalization targets

Goal — Theorem 2.26 (Gaussian concentration of Lipschitz functions)

Let (X1,…,Xn)(X_1,\dots,X_n)(X1​,…,Xn​) be i.i.d. standard Gaussian and fff be LLL-Lipschitz with respect to the Euclidean norm. Then f(X)−E[f(X)]f(X)-\mathbb E[f(X)]f(X)−E[f(X)] is sub-Gaussian with parameter at most LLL, and hence

P[∣f(X)−E[f(X)]∣≥t]  ≤  2e−t2/2L2for all t≥0.\mathbb P[|f(X)-\mathbb E[f(X)]|\ge t] \;\le\; 2e^{-t^2/2L^2} \qquad \text{for all } t\ge 0.P[∣f(X)−E[f(X)]∣≥t]≤2e−t2/2L2for all t≥0.

The bound is dimension-free: it depends on nnn only through fff's Lipschitz constant, not the ambient dimension itself.

Milestone — Lemma 2.27 (Gaussian interpolation identity)

For any differentiable fff and convex φ\varphiφ, E[φ(f(X)−E[f(X)])]≤E[φ(π2⟨∇f(X),Y⟩)]\mathbb E[\varphi(f(X)-\mathbb E[f(X)])] \le \mathbb E[\varphi(\tfrac\pi2\langle\nabla f(X),Y\rangle)]E[φ(f(X)−E[f(X)])]≤E[φ(2π​⟨∇f(X),Y⟩)] for X,Y∼N(0,In)X,Y\sim N(0,I_n)X,Y∼N(0,In​) independent — the interpolation identity Theorem 2.26's proof is built on.

Milestone — Theorem 2.19 (martingale Bernstein bound)

Given a martingale difference sequence with a per-index sub-exponential conditional moment-generating-function bound E[eλDk∣Fk−1]≤eλ2νk2/2\mathbb E[e^{\lambda D_k}\mid\mathcal F_{k-1}]\le e^{\lambda^2\nu_k^2/2}E[eλDk​∣Fk−1​]≤eλ2νk2​/2 for ∣λ∣<1/αk|\lambda|<1/\alpha_k∣λ∣<1/αk​, the sum ∑kDk\sum_k D_k∑k​Dk​ is itself sub-exponential with parameters (∑kνk2, max⁡kαk)\big(\sqrt{\sum_k\nu_k^2},\ \max_k\alpha_k\big)(∑k​νk2​​, maxk​αk​), and satisfies the two-regime concentration inequality of Eq. (2.28): sub-Gaussian for small deviations, sub-exponential for large ones. This is the chapter's central general-purpose martingale concentration tool.

Significance

Theorem 2.19 is the source of two of the most-cited concentration inequalities in the field — the Azuma–Hoeffding inequality (Corollary 2.20) and the bounded-differences/McDiarmid inequality (Corollary 2.21), both already faithfully covered elsewhere on the platform (azuma_hoeffding_two_sided, bounded_diff_martingale_two_sided) and included here as kind: reference milestones rather than redrafted. Theorem 2.26's Gaussian Lipschitz concentration is separately significant: it is the tool behind dimension-free operator-norm bounds for random matrices, concentration of the empirical spectral distribution, and much of the machinery of Chapters 5 and 6 of the same book.

Formalizing it. No faithful prior art exists on the platform for either the martingale Bernstein bound or Lipschitz-Gaussian concentration itself (a fresh search for "martingale Bernstein," "sub-exponential martingale," "Gaussian interpolation," and "Lipschitz concentration" returned no hits; the existing Vershynin-book item HighDimProb.Isoperimetry.lipschitz_concentration_sphere concentrates a Lipschitz function on the sphere, a different underlying space and a different proof from Theorem 2.26's Gaussian vector). Both goal-adjacent theorems and the Gaussian interpolation lemma are drafted here as open goals (:= by sorry); the two Azuma–Hoeffding/bounded-differences corollaries are reused from the platform's existing, already-proved formalizations.

Difficulty

The naive approach to Theorem 2.26 — try to bound f(X)−E[f(X)]f(X)-\mathbb E[f(X)]f(X)−E[f(X)] directly via a Lipschitz-type argument in Rn\mathbb R^nRn — has no obvious route to a dimension-free bound, since a union bound over coordinates (or over an ε\varepsilonε-net of the domain) picks up a factor that grows with nnn. The resolution, Lemma 2.27's interpolation identity, instead exploits a special structural fact about the Gaussian distribution — its rotation invariance — to replace the nonlinear quantity f(X)−E[f(X)]f(X)-\mathbb E[f(X)]f(X)−E[f(X)] with the linear, and hence exactly computable, Gaussian quantity ⟨∇f(X),Y⟩\langle\nabla f(X),Y\rangle⟨∇f(X),Y⟩, at the mild cost of a non-optimal constant. Theorem 2.19's difficulty is bookkeeping rather than a conceptual obstruction: the recursive conditioning step (Eq. (2.29)) must be iterated exactly nnn times while keeping track of the interplay between the two parameters νk,αk\nu_k,\alpha_kνk​,αk​ per difference, and Proposition 2.9's two-regime tail bound (small-deviation sub-Gaussian behavior, large-deviation sub-exponential behavior) must be carried through unchanged into the final statement — dropping either regime understates what the theorem proves.

Formalization scope

Expectations are Bochner integrals against an explicit probability measure, with integrability required as an explicit hypothesis in IsSubGaussian and IsSubExponential (Mathlib's Bochner integral silently returns 000 for a non-integrable function, which this mission's definitions rule out as a trivializing formalization). The sub-exponential condition's domain restriction |λ| < 1/α is realized as the disjunction α = 0 ∨ |λ| < 1/α, since Lean's real division convention 1/0 = 0 is exactly backwards from the book's own stated 1/0 = +\infty convention for the degenerate sub-Gaussian case.

"X,Y∼N(0,In)X,Y\sim N(0,I_n)X,Y∼N(0,In​) independent" (Lemma 2.27, Theorem 2.26) is formalized via Mathlib's HasGaussianLaw predicate together with explicit coordinatewise mean-zero and identity-covariance hypotheses, which together pin down the standard multivariate normal law, plus IndepFun. The inner product ⟨∇f(X),Y⟩\langle\nabla f(X),Y\rangle⟨∇f(X),Y⟩ is realized as fderiv ℝ f (X ω) (Y ω), the Fréchet derivative applied to Y(ω)Y(\omega)Y(ω) — equal to ⟨∇f(X(ω)),Y(ω)⟩\langle\nabla f(X(\omega)),Y(\omega)\rangle⟨∇f(X(ω)),Y(ω)⟩ by the Riesz representation of the gradient on a Hilbert space, avoiding the need to separately construct a gradient vector field.

Theorem 2.19's printed parameter pair for part (a), "(∑kνk2,α∗)(\sum_k\nu_k^2,\alpha_*)(∑k​νk2​,α∗​)," is formalized as (∑kνk2,α∗)(\sqrt{\sum_k\nu_k^2},\alpha_*)(∑k​νk2​​,α∗​): Definition 2.7 parametrizes the sub-exponential MGF bound by ν\nuν (with ν2\nu^2ν2 appearing in the exponent), so a literal transcription of the printed pair's first entry would silently square the effective parameter and make part (a), read literally, inconsistent with part (b)'s own tail-bound formula (which the book derives from part (a) via the general sub-exponential tail bound, Proposition 2.9). The corrected pairing is the one the book's own proof actually establishes; see MODERATION_NOTES.md for the full derivation.

Out of scope for this mission: Proposition 2.5 (plain Hoeffding for a sum of independent sub-Gaussians), used in the book only as background for Theorem 2.26's proof and not redrafted, since the goal theorem's own statement does not depend on it once Lemma 2.27 is in hand.

Selected references

  • M. J. Wainwright, High-Dimensional Statistics: A Non-Asymptotic Viewpoint, Cambridge University Press, 2019. DOI: 10.1017/9781108627771. Chapter 2.
  • K. Azuma, "Weighted sums of certain dependent random variables," Tôhoku Mathematical Journal, 19:357–367, 1967.
  • W. Hoeffding, "Probability inequalities for sums of bounded random variables," Journal of the American Statistical Association, 58:13–30, 1963.
8 thms4 active usersReviewed
🏆Completed
Convex OptimizationLinear OptimizationOperations Research+1·Captain: mikedeng1

Introduction to Stochastic Programming I: Convexity, Attainment and Optimality of the Two-Stage Recourse ProblemTextbook

Motivation

Two-stage stochastic linear programming with recourse models a decision made before uncertainty resolves (the first-stage variables xxx) followed by a corrective decision made after (the second-stage, or recourse, variables yyy). Solving such a program means minimizing cTx+Q(x)c^{\mathsf T}x + Q(x)cTx+Q(x), where Q(x)Q(x)Q(x) is the expected cost of the best recourse action given xxx -- an object defined only implicitly, as the value of an embedded linear program that must be solved (or bounded) for every realization of the uncertain data. Before any algorithm for this problem can be justified -- the L-shaped method, stochastic decomposition, scenario decomposition, all developed in later chapters of Birge & Louveaux, Introduction to Stochastic Programming (Springer, 2011) -- one needs to know that QQQ is well-behaved enough to optimize over at all: that the feasible region is closed and convex, that QQQ itself is a finite, Lipschitz, convex function on it, that an optimal solution is actually attained rather than only approached in the limit, and finally what an optimality condition for the resulting nonsmooth convex program even looks like. This mission formalizes exactly that foundational layer, Chapter 3, Section 3.1 of the book.

Setting

Fix natural numbers n1,n2,m1,m2n_1, n_2, m_1, m_2n1​,n2​,m1​,m2​ and a finite scenario count KKK. A two-stage recourse instance consists of first-stage data A∈Rm1×n1A \in \mathbb{R}^{m_1 \times n_1}A∈Rm1​×n1​, b∈Rm1b \in \mathbb{R}^{m_1}b∈Rm1​, c∈Rn1c \in \mathbb{R}^{n_1}c∈Rn1​, a fixed recourse matrix W∈Rm2×n2W \in \mathbb{R}^{m_2 \times n_2}W∈Rm2​×n2​, and, for each scenario k=1,…,Kk = 1,\dots,Kk=1,…,K, a cost vector qk∈Rn2q_k \in \mathbb{R}^{n_2}qk​∈Rn2​, a right-hand side hk∈Rm2h_k \in \mathbb{R}^{m_2}hk​∈Rm2​, a technology matrix Tk∈Rm2×n1T_k \in \mathbb{R}^{m_2 \times n_1}Tk​∈Rm2​×n1​, and a probability pk≥0p_k \ge 0pk​≥0 with ∑kpk=1\sum_k p_k = 1∑k​pk​=1 (Eq. (1.1)). The first-stage feasible region is K1={x∣Ax=b, x≥0}K_1 = \{x \mid Ax = b,\ x \ge 0\}K1​={x∣Ax=b, x≥0}.

For a fixed xxx and scenario kkk, the second-stage value is

Q(x,ξk)=min⁡y{qkTy∣Wy=hk−Tkx, y≥0}Q(x,\xi_k) = \min_{y}\{q_k^{\mathsf T}y \mid Wy = h_k - T_k x,\ y \ge 0\}Q(x,ξk​)=ymin​{qkT​y∣Wy=hk​−Tk​x, y≥0}

(Eq. (1.6)), taken as an extended real: +∞+\infty+∞ if no feasible yyy exists, −∞-\infty−∞ if the inner program is unbounded below. The expected recourse value is Q(x)=∑kpk Q(x,ξk)Q(x) = \sum_k p_k\, Q(x,\xi_k)Q(x)=∑k​pk​Q(x,ξk​) (Eq. (1.3)), combined so that +∞+(−∞)=+∞+\infty + (-\infty) = +\infty+∞+(−∞)=+∞ -- the book's own convention (p. 109): infeasibility in one scenario is treated as fatal even if another scenario is unboundedly favorable. The second-stage feasibility set is K2={x∣Q(x)<∞}K_2 = \{x \mid Q(x) < \infty\}K2​={x∣Q(x)<∞}, and the deterministic-equivalent objective is z(x)=cTx+Q(x)z(x) = c^{\mathsf T}x + Q(x)z(x)=cTx+Q(x) (Eq. (1.2)). For xxx with Q(x)Q(x)Q(x) finite, the subdifferential ∂Q(x)\partial Q(x)∂Q(x) is the set of η\etaη satisfying Q(x)+ηT(y−x)≤Q(y)Q(x) + \eta^{\mathsf T}(y-x) \le Q(y)Q(x)+ηT(y−x)≤Q(y) for every yyy (p. 115).

A simple-recourse instance is the special case W=[I,−I]W = [I,-I]W=[I,−I]: the recourse cost splits as q=(q+,q−)q = (q^+,q^-)q=(q+,q−), and Q(x)Q(x)Q(x) decomposes componentwise via the closed form of Eq. (1.9)-(1.10) using the (left- and right-limit) distribution functions Fi−,Fi+F_i^-, F_i^+Fi−​,Fi+​ of each hih_ihi​.

Formalization targets

Goal -- Chapter 3, Theorem 9 (p. 116)

x∗∈K1 is optimal in (1.2)  ⟺  ∃ λ∗∈Rm1, μ∗∈R≥0n1, (μ∗)Tx∗=0,  s.t. −c+ATλ∗+μ∗∈∂Q(x∗),x^* \in K_1 \text{ is optimal in (1.2)} \iff \exists\, \lambda^* \in \mathbb{R}^{m_1},\ \mu^* \in \mathbb{R}^{n_1}_{\ge 0},\ (\mu^*)^{\mathsf T}x^* = 0,\ \text{ s.t. } -c + A^{\mathsf T}\lambda^* + \mu^* \in \partial Q(x^*),x∗∈K1​ is optimal in (1.2)⟺∃λ∗∈Rm1​, μ∗∈R≥0n1​​, (μ∗)Tx∗=0,  s.t. −c+ATλ∗+μ∗∈∂Q(x∗),

given that (1.2) has a finite optimal value. This is the KKT-style necessary and sufficient optimality condition for the two-stage recourse LP, and the weakest of the mission's targets in the sense that everything else supports it: convexity and finiteness of QQQ (Theorem 6) are what make the left-to-right implication meaningful, closedness/convexity of K2K_2K2​ (Theorem 5) makes the feasible region well-posed, and attainment (Theorem 8) is what makes "x∗x^*x∗ is optimal" a statement about a point that exists rather than an infimum that may not be reached.

Supporting milestones

  • Theorem 5(a) (p. 111): K2K_2K2​ is closed and convex.
  • Theorem 6(a) (p. 112): QQQ is finite on K2K_2K2​, and Lipschitzian and convex there.
  • Theorem 8 (p. 115): under boundedness of K1∩K2K_1 \cap K_2K1​∩K2​ or eventual linearity of QQQ along recession directions, a finite optimal value is attained.
  • Corollary 10 (p. 116): Theorem 9 specialized to simple recourse, with ∂Q(x∗)\partial Q(x^*)∂Q(x∗) replaced by its explicit componentwise description.

Significance

Theorem 9 is the hinge on which the rest of the book's algorithmic chapters turn. The L-shaped method (Chapter 5) is a cutting-plane scheme whose cuts are literally elements of ∂Q(x)\partial Q(x)∂Q(x); stochastic decomposition and sampling-based methods use the same subdifferential structure with estimated cuts; the differentiable specialization (Eq. (1.14), c+∇Q(x∗)=ATλ∗+μ∗c + \nabla Q(x^*) = A^{\mathsf T}\lambda^* + \mu^*c+∇Q(x∗)=ATλ∗+μ∗) underlies nonlinear-programming approaches to the smooth case. None of this is meaningful without first knowing QQQ is convex, finite where it needs to be, and that a minimizer exists to characterize. Formalizing this mission's four milestones from the actual definition of QQQ as an embedded linear program's value -- rather than assuming these properties -- is exactly the content the book itself proves (or, for Theorem 6, explicitly cites to Wets [1972] and Kall [1976] rather than proving); this mission asks for genuine Lean proofs of Theorems 5, 8, 9 and Corollary 10 from the LP structure of QQQ, and records Theorem 6 as a stated (not re-derived) input, matching the book's own presentation.

Difficulty

The obvious shortcut is to treat QQQ as an opaque convex function and apply a textbook convex-KKT theorem off the shelf. This fails to capture what Theorem 9 actually is: a statement about the specific function Q(x)=∑kpkmin⁡y{qkTy∣Wy=hk−Tkx, y≥0}Q(x) = \sum_k p_k \min_y\{q_k^{\mathsf T}y \mid Wy = h_k - T_k x,\ y \ge 0\}Q(x)=∑k​pk​miny​{qkT​y∣Wy=hk​−Tk​x, y≥0}, built from finitely many parametric linear programs, each of which can be infeasible (Q(x,ξk)=+∞Q(x,\xi_k) = +\inftyQ(x,ξk​)=+∞) or unbounded (Q(x,ξk)=−∞Q(x,\xi_k) = -\inftyQ(x,ξk​)=−∞) depending on xxx. Convexity of QQQ must come from convexity of the value function of a parametric LP in its right-hand side (the book's Theorem 2 argument: a convex combination of optimal solutions at two right-hand sides is feasible, hence suboptimal, at the combined right-hand side) -- not from an assumed hypothesis. Handling ±∞\pm\infty±∞ correctly is a second, easy-to-miss source of error: the book fixes an explicit, non-default convention (+∞+\infty+∞ dominates −∞-\infty−∞) for combining per-scenario values, the opposite of the convention Mathlib's own extended-real arithmetic uses, so any formalization that reaches for EReal's built-in addition to aggregate QQQ silently states a different theorem. Theorem 8's attainment condition is a genuine existence result, not an automatic consequence of convexity: continuity alone does not give attainment on an unbounded feasible region, and the book's own counterexample (Eq. (1.11), a negative-exponential tail with infimum 000 attained by no finite xxx) shows the boundedness/recession hypotheses are load-bearing.

Formalization scope

The scenario set is modeled as Fin K, a finite discrete random variable, matching Section 3.1b's development; under this model "ξ\xiξ has finite second moments" (the standing hypothesis of Theorems 4-11 in the general, possibly-continuous case) holds automatically and so does not appear as a separate hypothesis anywhere in this mission. Q(x,\xi_k) is defined as an EReal via sInf of the second-stage LP's feasible objective values -- sInf of the empty set is ⊤, and of a set unbounded below is ⊥ -- and is genuinely derived from that inner minimization rather than assumed convex; this rules out the chapter's trivializing formalization, which the paper-level triage explicitly warns against: taking Q(x) as an opaque convex-function hypothesis instead of deriving its properties from the inner LP's structure. Aggregating the KKK per-scenario values into Q(x)Q(x)Q(x) uses a bespoke bookAdd operation implementing the book's stated convention +∞+(−∞)=+∞+\infty+(-\infty)=+\infty+∞+(−∞)=+∞, since Mathlib's EReal addition is defined with the opposite convention (⊥+⊤=⊤+⊥=⊥\bot+\top=\top+\bot=\bot⊥+⊤=⊤+⊥=⊥). ∂Q(x)\partial Q(x)∂Q(x) is the ordinary subgradient-inequality set for this extended-real-valued function.

Theorem 8's condition (b) is stated with the book's own quantifier structure: the threshold λˉ\bar\lambdaλˉ and the recession value depend on the point xxx and direction vvv exactly as written, with no strengthening. Theorem 6(a)'s Lipschitz bound is stated, not derived -- the book itself cites it to Wets [1972] and Kall [1976] without proof -- so a faithful Lean proof of that milestone is expected to remain out of scope for this mission. Corollary 10 similarly takes the closed form of ∂Qi(x)\partial Q_i(x)∂Qi​(x) from Eq. (1.10) as a hypothesis on an abstract QQQ, matching how the book itself uses (1.10) as an already-established fact rather than re-deriving it from the second-stage LP in the corollary's own proof. Theorem 11's subdifferential-decomposition result (∂Q(x)=Eω[∂Q(x,ξ(ω))]+N(K2,x)\partial Q(x) = E_\omega[\partial Q(x,\xi(\omega))] + N(K_2,x)∂Q(x)=Eω​[∂Q(x,ξ(ω))]+N(K2​,x)) is deliberately left out of this mission's scope: it is not needed by Theorem 9's own proof, and its normal-cone term would require relatively-complete-recourse machinery this mission does not otherwise need. No prior-art match was found on the platform: VectorSpaceOpt.fenchel_duality and the Luenberger-derived VectorSpaceOpt.generalized_kuhn_tucker / kkt_complementary_slackness family use a differentiable (Gateaux-derivative) or conjugate-function KKT model over general normed spaces, not this chapter's finite-dimensional, possibly-nondifferentiable subgradient formulation over the specific polyhedral set K1K_1K1​, so none is a faithful match and all items here are original drafts.

Selected references

  • J.R. Birge and F. Louveaux, Introduction to Stochastic Programming, 2nd ed., Springer Series in Operations Research and Financial Engineering, Springer, 2011. https://doi.org/10.1007/978-1-4614-0237-4
  • R.J-B. Wets, "Programming Under Uncertainty: The Equivalent Convex Program," SIAM Journal on Applied Mathematics 14 (1966), 89-105 (Lipschitz continuity of the recourse function, cited by the book as Wets [1972] for the closely related result used in Theorem 6). https://doi.org/10.1137/0114008
  • D.P. Walkup and R.J-B. Wets, "Stochastic Programs with Recourse," SIAM Journal on Applied Mathematics 15 (1967), 1299-1314 (finiteness of the recourse function and coincidence of the possibility and expectation feasibility sets, underlying Proposition 3 and Theorem 4). https://doi.org/10.1137/0115113
8 thms4 active usersReviewed
🏆Completed
Functional Analysis·Captain: Lucas

An Introduction to Stochastic PDEs I: The Cameron–Martin TheoremTextbook

Motivation

A stochastic partial differential equation is driven by noise that lives on an infinite-dimensional function space, and the first object one has to control is the law of that noise: a Gaussian measure on a separable Banach space. Martin Hairer's lecture notes An Introduction to Stochastic PDEs (arXiv:0907.4178) devote their first technical chapter (Section 4) to exactly this, because every later construction — stochastic convolutions, invariant measures for semilinear equations, the ergodic theory of the stochastic Navier–Stokes equations — is phrased against it.

The single structural fact that chapter produces is the Cameron–Martin theorem (Theorem 4.44, p. 31). It answers the question: in which directions may one translate an infinite-dimensional Gaussian measure without destroying its null sets? In finite dimensions the answer is "all of them", because Lebesgue measure is translation invariant. In infinite dimensions the admissible directions form a proper, and typically much smaller, Hilbert subspace Hμ⊂BH_\mu \subset BHμ​⊂B — the Cameron–Martin space — and the translated measure is either equivalent to μ\muμ or mutually singular with it, with nothing in between. This dichotomy is the reason Girsanov-type changes of measure, Schilder-type large deviation rate functions, support theorems and Malliavin calculus all take the form they do.

Setting

Throughout, BBB is a separable Banach space, B∗B^{*}B∗ its topological dual, and μ\muμ a Borel probability measure on BBB.

μ\muμ is Gaussian (Definition 4.4, p. 19) if for every continuous linear functional ℓ∈B∗\ell \in B^{*}ℓ∈B∗ the push-forward ℓ∗μ\ell_{*}\muℓ∗​μ is a Gaussian measure on R\mathbb RR in the sense of Definition 4.1, i.e. has characteristic function exp⁡(−σ2ℓ2+iℓm)\exp(-\tfrac{\sigma}{2}\ell^{2} + i\ell m)exp(−2σ​ℓ2+iℓm); the degenerate case σ=0\sigma = 0σ=0, a Dirac mass, is included. It is centred if all these one-dimensional laws have mean zero, which is expressed here as ∫Bx μ(dx)=0\int_B x \, \mu(dx) = 0∫B​xμ(dx)=0.

For a centred Gaussian μ\muμ the covariance form (4.2, p. 20) is

Cμ(ℓ,ℓ′)  =  ∫Bℓ(x) ℓ′(x) μ(dx),ℓ,ℓ′∈B∗.C_\mu(\ell, \ell') \;=\; \int_B \ell(x)\,\ell'(x)\, \mu(dx), \qquad \ell, \ell' \in B^{*} .Cμ​(ℓ,ℓ′)=∫B​ℓ(x)ℓ′(x)μ(dx),ℓ,ℓ′∈B∗.

It is well defined because ∥x∥2\|x\|^{2}∥x∥2 is μ\muμ-integrable, and it is a bounded bilinear form (Corollary 4.14, p. 22).

The Cameron–Martin space (Definition 4.26, p. 27) is classically built as the completion of

H˚μ  =  { h∈B:∃ h∗∈B∗ with Cμ(h∗,ℓ)=ℓ(h)  ∀ℓ∈B∗ }\mathring H_\mu \;=\; \{\, h \in B : \exists\, h^{*} \in B^{*} \text{ with } C_\mu(h^{*}, \ell) = \ell(h) \ \ \forall \ell \in B^{*} \,\}H˚μ​={h∈B:∃h∗∈B∗ with Cμ​(h∗,ℓ)=ℓ(h)  ∀ℓ∈B∗}

under ∥h∥μ2=Cμ(h∗,h∗)\|h\|_\mu^{2} = C_\mu(h^{*}, h^{*})∥h∥μ2​=Cμ​(h∗,h∗). This mission uses the equivalent intrinsic description of Exercise 4.38 (p. 29), which avoids the completion:

∥h∥μ  =  sup⁡{ ℓ(h)  :  ℓ∈B∗, Cμ(ℓ,ℓ)≤1 },Hμ={ h∈B:∥h∥μ<∞ }.\|h\|_\mu \;=\; \sup\{\, \ell(h) \;:\; \ell \in B^{*},\ C_\mu(\ell, \ell) \le 1 \,\}, \qquad H_\mu = \{\, h \in B : \|h\|_\mu < \infty \,\} .∥h∥μ​=sup{ℓ(h):ℓ∈B∗, Cμ​(ℓ,ℓ)≤1},Hμ​={h∈B:∥h∥μ​<∞}.

The supremum is taken in [0,∞][0, \infty][0,∞]; since −ℓ-\ell−ℓ is admissible whenever ℓ\ellℓ is, it equals sup⁡∣ℓ(h)∣\sup |\ell(h)|sup∣ℓ(h)∣ over the same set. For h∈Bh \in Bh∈B write Th:B→BT_h : B \to BTh​:B→B, Th(x)=x+hT_h(x) = x + hTh​(x)=x+h.

Formalization targets

Goal — Theorem 4.44 (Cameron–Martin)

For a centred Gaussian measure μ\muμ on a separable Banach space BBB and h∈Bh \in Bh∈B,

(Th)∗μ ≪ μ⟺h∈Hμ.(T_h)_{*}\mu \ \ll \ \mu \qquad \Longleftrightarrow \qquad h \in H_\mu .(Th​)∗​μ ≪ μ⟺h∈Hμ​.

Both implications are asserted: translation along a Cameron–Martin direction produces an absolutely continuous measure, and translation along any other direction does not (in fact it produces a mutually singular measure).

Milestones

The milestone list follows the route of Section 4.2:

  1. Exercise 4.38 — the supremum description agrees with Definition 4.26 on H˚μ\mathring H_\muH˚μ​.
  2. Proposition 4.32 — Hμ⊂BH_\mu \subset BHμ​⊂B with ∥h∥2≤∥Cμ∥ ∥h∥μ2\|h\|^{2} \le \|C_\mu\| \, \|h\|_\mu^{2}∥h∥2≤∥Cμ​∥∥h∥μ2​.
  3. Proposition 4.40 — every L2(μ)L^{2}(\mu)L2(μ)-limit of elements of B∗B^{*}B∗ has a centred Gaussian law whose variance is its own L2L^2L2 norm squared.
  4. Equation (4.14) — the explicit density Dh(x)=exp⁡(h∗(x)−12∥h∥μ2)D_h(x) = \exp(h^{*}(x) - \tfrac12\|h\|_\mu^{2})Dh​(x)=exp(h∗(x)−21​∥h∥μ2​) of the shifted measure, for h∈H˚μh \in \mathring H_\muh∈H˚μ​.
  5. The total-variation separation bound ∥N(0,1)−N(m,1)∥TV≥2−2e−m2/8\|\mathcal N(0,1) - \mathcal N(m,1)\|_{\mathrm{TV}} \ge 2 - 2e^{-m^{2}/8}∥N(0,1)−N(m,1)∥TV​≥2−2e−m2/8 used in the converse half of Theorem 4.44.
  6. Proposition 4.45 — HμH_\muHμ​ is exactly the intersection of all measurable linear subspaces of full measure.

Significance

The Cameron–Martin theorem is what makes the Cameron–Martin space a canonical object rather than a formal construction: HμH_\muHμ​ is simultaneously the set of admissible shifts, the intersection of all full-measure linear subspaces (Proposition 4.45), and the space whose unit ball governs Gaussian isoperimetry (Theorem 4.53, Borell–Sudakov–Cirel'son). Downstream in the notes it is used to identify invariant measures of linear SPDEs and to compare them; outside the notes it is the starting point of Malliavin calculus and of large deviation theory for Gaussian measures.

Status, precisely. The Mathlib library pinned by this mission already contains a substantial part of Section 4: the predicate IsGaussian (Definition 4.4), uniqueness of measures with equal characteristic functionals on a separable Banach space (Propositions 4.8 and 4.11), invariance of μ⊗μ\mu \otimes \muμ⊗μ under rotations (Proposition 4.12), Fernique's theorem (Theorem 4.13), and the covariance form of (4.2) together with its boundedness (Corollary 4.14) as a continuous bilinear form on the dual. Those results are therefore not milestones here; they are the assumed foundation. What is absent, and what this mission asks for, is everything from the Cameron–Martin space onwards: its definition, its elementary properties, and Theorem 4.44 itself. No machine-checked proof of the infinite-dimensional Cameron–Martin theorem is known to the captain in any Lean library.

Difficulty

The naive route — write down the two densities and take their ratio — is unavailable: there is no translation-invariant reference measure on an infinite-dimensional Banach space, so "the density of μ\muμ" does not exist and the Radon–Nikodym derivative of (Th)∗μ(T_h)_{*}\mu(Th​)∗​μ with respect to μ\muμ must be produced directly, as the exponential of a random variable.

That random variable is the obstruction. For h∈H˚μh \in \mathring H_\muh∈H˚μ​ the functional h∗h^{*}h∗ is continuous and the computation is a characteristic-function identity. But H˚μ\mathring H_\muH˚μ​ is in general strictly smaller than HμH_\muHμ​: a general h∈Hμh \in H_\muh∈Hμ​ has an associated h∗h^{*}h∗ that exists only as an L2(μ)L^{2}(\mu)L2(μ)-limit of continuous functionals, defined μ\muμ-almost everywhere and linear only on a measurable subspace of full measure (Propositions 4.34 and 4.39). Establishing that these limits are Gaussian with the expected variance (Proposition 4.40) is the technical bridge, and it is why milestone 3 is stated as a statement about L2L^{2}L2-limits rather than about elements of B∗B^{*}B∗.

The converse half has a different shape. One must produce, for h∉Hμh \notin H_\muh∈/Hμ​, a single one-dimensional projection that separates μ\muμ from (Th)∗μ(T_h)_{*}\mu(Th​)∗​μ arbitrarily well; unboundedness of ℓ(h)\ell(h)ℓ(h) over the covariance unit ball supplies ℓ\ellℓ with Cμ(ℓ,ℓ)=1C_\mu(\ell,\ell) = 1Cμ​(ℓ,ℓ)=1 and ℓ(h)\ell(h)ℓ(h) as large as desired, and the quantitative Gaussian total-variation bound of milestone 5 converts this into total variation distance 222, i.e. mutual singularity.

Formalization scope

The development is in Lean 4 with Mathlib, in the namespace HairerSPDE, shared by the whole series drawn from these notes. The conventions it commits to:

  1. BBB carries NormedAddCommGroup, NormedSpace ℝ, its Borel σ-algebra, CompleteSpace and SecondCountableTopology — the last two encode "separable Banach space".
  2. Gaussianity is Mathlib's IsGaussian, which is Definition 4.4 verbatim; centredness is the extra hypothesis ∫x dμ=0\int x \, d\mu = 0∫xdμ=0, needed because IsGaussian permits a non-zero mean.
  3. The covariance form is Mathlib's covarianceBilinDual, which equals (4.2) for centred measures with finite second moment and is set to zero otherwise; Fernique's theorem rules the degenerate branch out for Gaussian measures.
  4. The Cameron–Martin norm is the [0,∞][0,\infty][0,∞]-valued supremum above, so membership in HμH_\muHμ​ is finiteness of that supremum; this is the only new definition the mission publishes.
  5. Translation is fun x ↦ x + h, absolute continuity is Mathlib's ≪, and "measurable linear subspace" is a Submodule ℝ B whose carrier is a measurable set.

The goal is an equivalence, so neither half can be discharged vacuously: the direction h∈Hμ⇒(Th)∗μ≪μh \in H_\mu \Rightarrow (T_h)_*\mu \ll \muh∈Hμ​⇒(Th​)∗​μ≪μ is non-trivial already for h≠0h \ne 0h=0 in finite dimensions, and the converse has content precisely when Hμ≠BH_\mu \ne BHμ​=B. Note that ∥0∥μ=0\|0\|_\mu = 0∥0∥μ​=0 always, and that for μ\muμ a Dirac mass the covariance form vanishes and Hμ={0}H_\mu = \{0\}Hμ​={0}; both degenerate cases are inside the statement rather than excluded by hypothesis.

Contributions welcome beyond the milestones: the reproducing kernel space RμR_\muRμ​ and the isomorphism of Proposition 4.34, measurable linear extensions (Proposition 4.39), the dilation singularity of Proposition 4.43, and μ(Hμ)=0\mu(H_\mu) = 0μ(Hμ​)=0 in the infinite-dimensional case (second half of Proposition 4.45). All of these are reusable outside this mission.

Selected references

  • M. Hairer, An Introduction to Stochastic PDEs, lecture notes, 2009/2023. arXiv:0907.4178
  • V. I. Bogachev, Gaussian Measures, Mathematical Surveys and Monographs 62, American Mathematical Society, 1998. DOI:10.1090/surv/062
  • X. Fernique, Intégrabilité des vecteurs gaussiens, C. R. Acad. Sci. Paris Sér. A-B 270 (1970), A1698–A1699.
  • G. Da Prato, J. Zabczyk, Stochastic Equations in Infinite Dimensions, Cambridge University Press, 1992. DOI:10.1017/CBO9780511666223
21 thms4 active usersReviewed
🏆Completed
Control TheoryDynamic ProgrammingOperations Research·Captain: Shuze Chen

Dynamic Programming and Optimal Control V: LQG and Certainty EquivalenceTextbook

Motivation

The separation theorem — certainty equivalence for linear-quadratic control with imperfect state information — is one of the celebrated structural results of stochastic control: the optimal controller splits into a least-squares estimator and the deterministic LQR actuator, designed independently. It underlies every LQG autopilot and Kalman-filter-based regulator. Section 5.2 of Bertsekas, Dynamic Programming and Optimal Control, Vol. I (3rd ed., 2005) proves it from the DP algorithm over information vectors, with Lemma 5.2.1 supplying the key fact that the estimation error is beyond the controller's influence. No formal analogue exists in Mathlib.

Setting

Linear dynamics and measurements

xk+1=Akxk+Bkuk+wk,zk=Ckxk+vk,x_{k+1} = A_k x_k + B_k u_k + w_k, \qquad z_k = C_k x_k + v_k,xk+1​=Ak​xk​+Bk​uk​+wk​,zk​=Ck​xk​+vk​,

with quadratic cost E[xN⊤QNxN+∑k<N(xk⊤Qkxk+uk⊤Rkuk)]\mathbb{E}\big[x_N^\top Q_N x_N + \sum_{k<N}(x_k^\top Q_k x_k + u_k^\top R_k u_k)\big]E[xN⊤​QN​xN​+∑k<N​(xk⊤​Qk​xk​+uk⊤​Rk​uk​)], Qk⪰0Q_k \succeq 0Qk​⪰0, Rk≻0R_k \succ 0Rk​≻0. The initial state and the zero-mean disturbances/noises are independent with finite ranges; independence is structural — the sample space is the product of an initial-state coordinate and per-stage noise coordinates (BertsekasLQGModel, BertsekasLQGSample, BertsekasLQGProb). A policy maps the realized measurement history (z0,…,zk)(z_0,\dots,z_k)(z0​,…,zk​) to uku_kuk​; the closed-loop process is BertsekasLQGTraj, the expected cost BertsekasLQGCost. The estimator E[xk∣Ik]\mathbb{E}[x_k \mid I_k]E[xk​∣Ik​] is an explicit conditional average (BertsekasCondExpVec, BertsekasLQGEstimate); the gains LkL_kLk​ come from the time-varying Riccati recursion (BertsekasLQGRiccati, BertsekasLQGGain).

Target

π∗(Ik)=Lk E[xk∣Ik]  along its own trajectories⟹J(π∗)≤J(π)  ∀π,\pi^*(I_k) = L_k\, \mathbb{E}[x_k \mid I_k] \ \text{ along its own trajectories} \quad\Longrightarrow\quad J(\pi^*) \le J(\pi)\ \ \forall \pi,π∗(Ik​)=Lk​E[xk​∣Ik​]  along its own trajectories⟹J(π∗)≤J(π)  ∀π,

— BertsekasDP.lqg_certainty_equivalence (goal). Milestone: Lemma 5.2.1 in pointwise form — the error xk−E[xk∣Ik]x_k - \mathbb{E}[x_k \mid I_k]xk​−E[xk​∣Ik​] is the same under any two policies, outcome by outcome (lqg_estimation_error_policy_independent).

Significance

This is the theorem that justifies designing estimator and controller separately — remove it and the entire LQG methodology loses its warrant. The formalization also yields the first machine-checked instance of the informational decomposition (control-dependent part + policy-independent error) that recurs throughout imperfect-information control. Notably the result needs no Gaussian assumption — only zero mean and independence — and the finite-support model makes that generality exact. The result is classical (Joseph–Tou 1961, Gunckel–Franklin 1963; the book's §5.2); the formal proof is new.

Difficulty

The heart is Lemma 5.2.1: showing the estimation error coincides, sample by sample, with the error of the control-free system — which requires proving that the observation-history σ-events under any policy coincide with those of the control-free system (controls are determined by the history, so they shift observations by a known amount). Then the DP argument over information histories must carry the quadratic decomposition through the backward recursion. Bookkeeping over histories-as-lists is the main formal burden; probability theory stays finite.

Formalization scope

Finite-support randomness (all expectations are finite sums); conditional expectation with the explicit junk value 0 on zero-probability events — the goal's hypothesis is accordingly restricted to outcomes of positive probability. Policies are functions of the measurement list only (equivalent to the book's information vector for deterministic policies, since past controls are recoverable from past measurements). Matrices are time-varying; positive definiteness of RkR_kRk​ makes every matrix inverse in the gains genuine. Measurement noise covariance is not assumed positive definite — the estimator is the abstract conditional expectation, not the Kalman filter (whose recursive form, §5.2.1, would be a natural follow-up mission).

Selected references

  • D. P. Bertsekas, Dynamic Programming and Optimal Control, Vol. I, 3rd ed., Athena Scientific, 2005. (§5.2, Lemma 5.2.1.) http://www.athenasc.com/dpbook.html
  • P. D. Joseph, J. T. Tou, On linear control theory, Trans. AIEE 80 (1961), 193–196. https://doi.org/10.1109/TAI.1961.6371743
  • T. L. Gunckel, G. F. Franklin, A general solution for linear sampled-data control, J. Basic Eng. 85 (1963), 197–201. https://doi.org/10.1115/1.3656559
5 thms4 active usersReviewed
🏆Completed
Markov ChainStochastic Systems·Captain: Shuze Chen

Markov Chains and Mixing Times VIII: Path Coupling and Approximate CountingTextbook

Motivation

The coupling method of Mission III asks for a coupling of two copies of a chain from every pair of starting states — often painful to construct globally. Chapter 14 of Levin–Peres–Wilmer, Markov Chains and Mixing Times (AMS, 2009) replaces that global demand by a local one. The path coupling technique of Bubley and Dyer says: put a connected graph structure on the state space, and couple one step of the chain only across edges; if each edge contracts in expectation, contraction propagates automatically along paths to arbitrary pairs of distributions. The bookkeeping runs through the transportation metric (Kantorovich distance) between distributions, whose theory — attainment by an optimal coupling, the triangle inequality — is developed on the way. The chapter's payoff is the sharpest elementary bound for sampling proper colorings (Theorem 14.8: the Glauber dynamics mixes in O(nlog⁡n)O(n\log n)O(nlogn) steps once q>2Δq>2\Deltaq>2Δ), and, through the sampling-to-counting reduction of Jerrum–Valiant–Vazirani, a polynomial-time approximation algorithm for counting colorings — the paradigm of the Markov chain Monte Carlo method as an algorithmic tool.

Setting

All chains live on a finite state space VVV with a transition matrix PPP; Pt(x,⋅)P^t(x,\cdot)Pt(x,⋅) is the time-ttt distribution from xxx, ∥μ−ν∥TV=max⁡A⊆V∣μ(A)−ν(A)∣\|\mu-\nu\|_{TV}=\max_{A\subseteq V}|\mu(A)-\nu(A)|∥μ−ν∥TV​=maxA⊆V​∣μ(A)−ν(A)∣ the total variation distance, d(t)=max⁡x∥Pt(x,⋅)−π∥TVd(t)=\max_x\|P^t(x,\cdot)-\pi\|_{TV}d(t)=maxx​∥Pt(x,⋅)−π∥TV​ the worst-case distance to the stationary distribution π\piπ, and tmix(ε)=min⁡{t:d(t)≤ε}t_{\mathrm{mix}}(\varepsilon)=\min\{t:d(t)\le\varepsilon\}tmix​(ε)=min{t:d(t)≤ε} the mixing time. A coupling of distributions μ,ν\mu,\nuμ,ν is a distribution qqq on V×VV\times VV×V with marginals μ\muμ and ν\nuν.

Given a metric-like cost ρ\rhoρ on pairs of states, the transportation metric between two distributions is the cheapest expected cost of moving one onto the other:

ρK(μ,ν)=min⁡{∑x,yρ(x,y) q(x,y)  :  q a coupling of μ,ν}.\rho_K(\mu,\nu)=\min\Bigl\{\sum_{x,y}\rho(x,y)\,q(x,y)\;:\;q\ \text{a coupling of}\ \mu,\nu\Bigr\}.ρK​(μ,ν)=min{x,y∑​ρ(x,y)q(x,y):q a coupling of μ,ν}.

Given a connected graph structure GGG on the state space with edge lengths ℓ≥1\ell\ge1ℓ≥1, the path metric ρ(x,y)\rho(x,y)ρ(x,y) is the least total length of a GGG-path from xxx to yyy.

For the colorings application: a qqq-coloring of the vertices of a graph is proper when adjacent vertices receive distinct colors, and the Glauber dynamics on proper colorings picks a uniform vertex and re-samples its color uniformly among the colors legal there; its stationary distribution is uniform on the proper colorings. Throughout, nnn is the number of vertices and Δ\DeltaΔ the maximum degree of the graph being colored.

Formalization targets

Goal

Theorem 14.8, the capstone of Chapter 14: for the Glauber dynamics on proper qqq-colorings, if q>2Δq>2\Deltaq>2Δ then

tmix(ε)  ≤  ⌈q−Δq−2Δ  n (log⁡n−log⁡ε)⌉.t_{\mathrm{mix}}(\varepsilon)\;\le\;\Bigl\lceil\frac{q-\Delta}{q-2\Delta}\;n\,\bigl(\log n-\log\varepsilon\bigr)\Bigr\rceil.tmix​(ε)≤⌈q−2Δq−Δ​n(logn−logε)⌉.

Milestones

  • Lemma 14.3 and Remark 14.2 — the transportation distance is attained by an optimal coupling, and satisfies the triangle inequality (so it is a genuine metric on distributions).
  • Theorem 14.6, path coupling (Bubley–Dyer) — if for every edge {x,y}\{x,y\}{x,y} of a connected graph structure there is a coupling of the one-step distributions P(x,⋅),P(y,⋅)P(x,\cdot),P(y,\cdot)P(x,⋅),P(y,⋅) contracting the path metric by e−αe^{-\alpha}e−α in expectation, then one step of the chain contracts the transportation metric of arbitrary distribution pairs by e−αe^{-\alpha}e−α.
  • Corollary 14.7 — under the same hypotheses, d(t)≤e−αt diam(V)d(t)\le e^{-\alpha t}\,\mathrm{diam}(V)d(t)≤e−αtdiam(V) and tmix(ε)≤⌈(log⁡diam(V)−log⁡ε)/α⌉t_{\mathrm{mix}}(\varepsilon)\le\lceil(\log\mathrm{diam}(V)-\log\varepsilon)/\alpha\rceiltmix​(ε)≤⌈(logdiam(V)−logε)/α⌉, where diam(V)\mathrm{diam}(V)diam(V) is the largest path-metric distance between two states.
  • Theorem 14.12, approximate counting — for q>2Δq>2\Deltaq>2Δ there is a randomized estimator, computed from an explicit polynomial number of independent uniform random seeds, which with probability at least 1−η1-\eta1−η estimates the number of proper qqq-colorings within a (1±ε)(1\pm\varepsilon)(1±ε) factor: rapid sampling yields rapid approximate counting.

Significance

The results. Path coupling converted the coupling method from an art into a calculus: one bounds a single-edge contraction constant, and the machinery does the rest. It is the standard tool for Glauber dynamics on colorings, independent sets, and other constraint-satisfaction models, and the q>2Δq>2\Deltaq>2Δ colorings bound is its flagship application. Theorem 14.12 is the discrete embodiment of the Jerrum–Valiant–Vazirani equivalence between approximate counting and sampling — the conceptual foundation of the entire MCMC approach to #P\#\mathrm P#P-hard counting problems.

Formalizing them. Mathlib has no transportation/Kantorovich metric in the finite setting, no path coupling, and nothing on approximate counting. The transportation-metric layer (optimal couplings, triangle inequality) is reusable far beyond this mission — it is the finite Wasserstein distance. The path-coupling theorem feeds directly into Mission IX (Ising) and is quoted throughout modern mixing literature.

Difficulty

The transportation metric asks for minimization over the (compact) polytope of couplings: attainment is a finite-dimensional compactness argument, and the triangle inequality requires gluing two optimal couplings along their common marginal — the classic construction that must be carried out with explicit finite sums here. Path coupling itself is an induction along geodesics of the path metric, with the subtlety that the composite coupling produced along a path need not be optimal, only admissible; the bookkeeping of the contraction constant through the induction is exactly the kind of argument Lean keeps honest. Theorem 14.8 instantiates the machinery: the single-edge coupling for colorings needs a careful case analysis of the proposed recolorings at the two endpoints (matching legal colors bijectively), and the contraction constant (q−2Δ)/(q−Δ)(q-2\Delta)/(q-\Delta)(q−2Δ)/(q−Δ) emerges from counting disagreeing proposals. Theorem 14.12 layers a probabilistic-amplification argument (medians of means over independent runs) on top of the mixing bound; its combinatorial core — expressing ∣Ω∣−1|\Omega|^{-1}∣Ω∣−1 as a telescoping product of marginal probabilities — is elementary but notation-heavy, and the formal statement quantifies over explicit seed spaces, so the whole estimator is a finite object.

Formalization scope

The transportation metric is an sInf over coupling costs (the coupling polytope is nonempty for genuine distributions, and attainment is part of the milestone, so the junk value never propagates); the path metric is an sInf over walk lengths in a connected graph. The Glauber dynamics on colorings is the restriction to proper colorings of the single-site heat-bath chain of Mission II, matching §3.3 of the book; its state space is the subtype of proper colorings, nonempty whenever q>2Δq>2\Deltaq>2Δ (a fact the hypotheses of the goal supply). Mixing-time upper bounds are stated with the book's explicit ceilings, so no rounding slack is hidden. In Theorem 14.12 the estimator is presented concretely as a function of finitely many uniform seeds, and "with probability ≥1−η\ge1-\eta≥1−η" is a counting inequality over the seed space — no measure theory enters. Contributions of intermediate lemmas (optimal-coupling gluing, geodesic decompositions, colorings edge-coupling) are welcome and will be reused by Mission IX.

Selected references

  • D. A. Levin, Y. Peres, E. L. Wilmer, Markov Chains and Mixing Times, American Mathematical Society, 2009. https://documents.epfl.ch/groups/i/ip/ipg/www/2013-2014/Random_Walks/markovmixing.pdf
  • R. Bubley, M. Dyer, Path coupling: a technique for proving rapid mixing in Markov chains, FOCS 1997. https://doi.org/10.1109/SFCS.1997.646111
  • M. Jerrum, A very simple algorithm for estimating the number of k-colorings of a low-degree graph, Random Structures Algorithms 7 (1995). https://doi.org/10.1002/rsa.3240070205
  • M. Jerrum, L. Valiant, V. Vazirani, Random generation of combinatorial structures from a uniform distribution, Theoret. Comput. Sci. 43 (1986). https://doi.org/10.1016/0304-3975(86)90174-X
9 thms4 active usersReviewed
🏆Completed
Markov ChainStochastic Systems·Captain: Shuze Chen

Markov Chains and Mixing Times VI: Networks, Hitting Times, and Cover TimesTextbook

Motivation

A reversible Markov chain is an electrical network: states are nodes, and the conductance c(x,y)=π(x)P(x,y)c(x,y)=\pi(x)P(x,y)c(x,y)=π(x)P(x,y) turns hitting probabilities into voltages and hitting times into resistances. This dictionary, going back to Kakutani and popularized by Doyle and Snell, converts probabilistic estimates into the physical laws of circuits — series/parallel reduction, energy minimization, monotonicity under edge removal. Chapters 9–11 of Levin–Peres–Wilmer, Markov Chains and Mixing Times (AMS, 2009) develop the dictionary and its two crown results: the commute time identity of Chandra–Raghavan–Ruzzo–Smolensky–Tiwari, Ea(τb)+Eb(τa)=cG R(a↔b)\mathbb E_a(\tau_b)+\mathbb E_b(\tau_a)=c_G\,R(a\leftrightarrow b)Ea​(τb​)+Eb​(τa​)=cG​R(a↔b), and the Matthews method bounding cover times by hitting times with harmonic-number precision.

Setting

A network is a symmetric nonnegative conductance function ccc on pairs of vertices; the associated walk moves with probabilities

P(x,y)=c(x,y)/c(x)P(x,y)=c(x,y)/c(x)P(x,y)=c(x,y)/c(x)

where c(x)=∑yc(x,y)c(x)=\sum_y c(x,y)c(x)=∑y​c(x,y), and cG=∑xc(x)c_G=\sum_x c(x)cG​=∑x​c(x). A function hhh is harmonic at xxx if h(x)=∑yP(x,y)h(y)h(x)=\sum_y P(x,y)h(y)h(x)=∑y​P(x,y)h(y). The voltage with boundary values 111 at aaa and 000 at zzz is W(x)=Px{τa<τz}W(x)=\mathbb P_x\{\tau_a<\tau_z\}W(x)=Px​{τa​<τz​}, the harmonic extension of its boundary data; the current flowing out of aaa has strength ∥I∥=∑yc(a,y) [W(a)−W(y)]\|I\|=\sum_y c(a,y)\,[W(a)-W(y)]∥I∥=∑y​c(a,y)[W(a)−W(y)], and the effective resistance is R(a↔z)=∥I∥−1R(a\leftrightarrow z)=\|I\|^{-1}R(a↔z)=∥I∥−1. A flow from aaa to zzz is an antisymmetric edge function obeying the node law off {a,z}\{a,z\}{a,z}; its energy is E(θ)=∑eθ(e)2/c(e)\mathcal E(\theta)=\sum_e\theta(e)^2/c(e)E(θ)=∑e​θ(e)2/c(e). Hitting times τS=min⁡{t≥0:Xt∈S}\tau_S=\min\{t\ge0:X_t\in S\}τS​=min{t≥0:Xt​∈S}, their expectations, the Green's function Gτz(a,x)G_{\tau_z}(a,x)Gτz​​(a,x), the maximal hitting time thitt_{\mathrm{hit}}thit​, and the cover time tcovt_{\mathrm{cov}}tcov​ (expected time to visit every state, maximized over starts) all use the trajectory calculus of Mission I.

Formalization targets

Goal

Ea(τb)+Eb(τa)  =  cG R(a↔b).\mathbb E_a(\tau_b)+\mathbb E_b(\tau_a)\;=\;c_G\,R(a\leftrightarrow b).Ea​(τb​)+Eb​(τa​)=cG​R(a↔b).

This is Proposition 10.6, the commute time identity — the exact bridge between the probabilistic and electrical sides, and the engine of the transience/recurrence theory of Mission XII.

Milestones

Reversibility and stationarity of the network walk with π(x)=c(x)/cG\pi(x)=c(x)/c_Gπ(x)=c(x)/cG​ (§9.1); Proposition 9.1 (existence and uniqueness of harmonic extensions with given boundary values, h(x)=Exf(XτB)h(x)=\mathbb E_x f(X_{\tau_B})h(x)=Ex​f(XτB​​)); Lemma 9.6 (the Green's function identity Gτz(a,a)=c(a)R(a↔z)G_{\tau_z}(a,a)=c(a)R(a\leftrightarrow z)Gτz​​(a,a)=c(a)R(a↔z)); Theorem 9.10 (Thomson's principle: R(a↔z)R(a\leftrightarrow z)R(a↔z) is the minimal energy of a unit flow, attained); Theorem 9.12 (Rayleigh monotonicity: lowering conductances raises effective resistance); Lemma 10.1 (the random target lemma: ∑yEa(τy)π(y)\sum_y\mathbb E_a(\tau_y)\pi(y)∑y​Ea​(τy​)π(y) does not depend on aaa); Corollary 10.8 (the resistance triangle inequality); Theorem 11.2 (Matthews: tcov≤thit (1+12+⋯+1n)t_{\mathrm{cov}}\le t_{\mathrm{hit}}\,(1+\tfrac12+\dots+\tfrac1n)tcov​≤thit​(1+21​+⋯+n1​)); Proposition 11.4 (the matching Matthews lower bound over subsets).

Significance

The results. The commute time identity computes hitting times from circuit reductions — this is how hitting times on trees, tori and glued graphs are actually evaluated — and, through Thomson and Rayleigh, makes them monotone under graph operations, something invisible probabilistically. The Matthews bounds pin cover times up to a log⁡n\log nlogn factor in complete generality; they are the tool behind cover-time results for lamplighter groups in Mission XI's sequel. Green's function identities feed Mission XII's recurrence theory, where R(a↔∞)R(a\leftrightarrow\infty)R(a↔∞) decides transience.

Formalizing them. Mathlib has graph Laplacians but no electrical network theory: no effective resistance, no flows, no energy, no Thomson/Rayleigh, no hitting or cover times. This mission publishes that layer over the trajectory calculus of Mission I. It is the most reusable single block of the series outside Missions I–II: effective resistance on finite networks is of independent interest to combinatorics (spanning trees, spectral sparsification) well beyond mixing times.

Difficulty

The identity chain behind the goal runs: Green's function of the stopped walk →\to→ escape probability Pa{τz<τa+}=(c(a)R(a↔z))−1\mathbb P_a\{\tau_z<\tau_a^+\}=\bigl(c(a)R(a\leftrightarrow z)\bigr)^{-1}Pa​{τz​<τa+​}=(c(a)R(a↔z))−1 (via harmonic uniqueness) →\to→ the Aldous–Fill occupation identity Gτ(a,x)=Ea(τ)π(x)G_\tau(a,x)=\mathbb E_a(\tau)\pi(x)Gτ​(a,x)=Ea​(τ)π(x) for stopping times with Xτ=aX_\tau=aXτ​=a — each step is a manipulation of infinite series of trajectory sums whose exchange steps (splitting a path at its first visit, last-exit decompositions) need summability from Mission I's Lemma 1.13. Thomson's principle is a finite-dimensional convex minimization: existence of the minimizer needs a compactness or completing-the-square argument, and the identification of the minimizer with the current flow needs the cycle law; the naive "differentiate the energy" route must be made exact. Matthews' method is a clean but genuinely clever argument — a uniformly random ordering of targets and the harmonic-number telescoping; the formal cost is the exchangeability of the randomized order against the chain, handled combinatorially.

Formalization scope

Networks are functions c:V×V→Rc:V\times V\to\mathbb Rc:V×V→R with a symmetry-and-nonnegativity predicate; loops are permitted; connectivity enters as irreducibility of the induced walk. The voltage is defined probabilistically as Px{τa<τz}\mathbb P_x\{\tau_a<\tau_z\}Px​{τa​<τz​} (the book's harmonic characterization is Proposition 9.1); R(a↔z)R(a\leftrightarrow z)R(a↔z) is the reciprocal of the explicit current strength, with total division junk when a,za,za,z are disconnected — statements carry irreducibility so this does not arise. Energy counts each undirected edge once, formalized as half the ordered double sum, and 02/0=00^2/0=002/0=0 handles absent edges. Cover times are tail sums of the explicit event "some state unvisited". The Matthews lower bound is stated with an arbitrary lower bound mmm for the pairwise hitting times of the subset AAA — equivalent to the book's min over pairs and easier to instantiate.

Welcome contributions: series/parallel reduction laws, the cycle and node law API for flows, escape-probability lemmas — all reused in Mission XII's infinite-network arguments.

Selected references

  • D. A. Levin, Y. Peres, E. L. Wilmer, Markov Chains and Mixing Times, American Mathematical Society, 2009. https://documents.epfl.ch/groups/i/ip/ipg/www/2013-2014/Random_Walks/markovmixing.pdf
  • P. G. Doyle, J. L. Snell, Random Walks and Electric Networks, MAA, 1984. https://arxiv.org/abs/math/0001057
  • A. K. Chandra, P. Raghavan, W. L. Ruzzo, R. Smolensky, P. Tiwari, The electrical resistance of a graph captures its commute and cover times, STOC 1989. https://doi.org/10.1145/73007.73062
  • P. Matthews, Covering problems for Markov chains, Ann. Probab. 16 (1988). https://doi.org/10.1214/aop/1176991894
15 thms4 active usersReviewed
🏆Completed
Experimental DesignOperations ResearchReinforcement Learning+1·Captain: Shuze Chen

Treatment Locality in A/B TestingResearch Paper

Modern A/B tests must infer lifetime treatment effects — e.g. customer lifetime value under a new feature — from short-horizon experiment data. Chen, Simchi-Levi and Wang (arXiv:2407.19618) model the experiment as a Markov decision process and exploit a structural fact of many practical interventions: the treatment is local, modifying the system at a single crucial state only. This mission formalizes the core asymptotic theory of the paper: for any differentiable estimator built from the experiment's transition and reward statistics, information sharing — pooling across test arms the samples collected away from the treated state — keeps the estimator asymptotically normal with the same asymptotic bias and never increases its asymptotic variance (Theorem 9), and is asymptotically efficient among unbiased estimators (Theorem 5). The route runs through a Markov chain central limit theorem with the asymptotic variance identified as the autocovariance series, and the linearization/delta method for functionals of chain statistics.

42 thms4 active users
🏆Completed
Markov ChainStochastic Systems·Captain: Shuze Chen

Markov Chains and Mixing Times I: Existence and Uniqueness of the Stationary DistributionTextbook

Markov Chains and Mixing Times I: Existence and Uniqueness of the Stationary Distribution

Motivation

Finite Markov chains are the basic model for memoryless random dynamics: card shuffles, random walks on graphs and groups, Monte Carlo samplers, and queueing systems are all chains on a finite state space. The single most used fact about them is that an irreducible chain has exactly one stationary distribution — a probability vector π\piπ with π=πP\pi = \pi Pπ=πP — and that this π\piπ is strictly positive and encodes the long-run behaviour of the chain through the return-time identity π(x)=1/Ex(τx+)\pi(x) = 1/\mathbb{E}_x(\tau_x^+)π(x)=1/Ex​(τx+​). Every later result in the theory of mixing times (convergence theorems, coupling bounds, spectral methods, cutoff) is a statement about the distance of the chain from this π\piπ, so nothing in the subject can be formalized before this mission is.

This mission is the first in a series formalizing D. A. Levin, Y. Peres and E. L. Wilmer, Markov Chains and Mixing Times (AMS, 2009), covering Chapters 1–2: the basic vocabulary of finite chains (stochastic matrices, irreducibility, period, reversibility, time reversal, random walks on graphs and groups) and the classical examples of Chapter 2 (gambler's ruin, coupon collecting, the reflection principle for simple random walk on Z\mathbb{Z}Z). Later missions in the series build on the definitions published here.

Setting

A chain on a finite state space Ω\OmegaΩ is presented by its transition matrix, a matrix P∈RΩ×ΩP \in \mathbb{R}^{\Omega\times\Omega}P∈RΩ×Ω with nonnegative entries whose rows sum to 111. A distribution is a row vector μ\muμ with nonnegative entries summing to 111; one step of the chain carries μ\muμ to μP\mu PμP, and the ttt-step transition probabilities are the entries of the matrix power PtP^tPt.

The chain is irreducible if for all states x,yx, yx,y there is a ttt with Pt(x,y)>0P^t(x,y) > 0Pt(x,y)>0. The period of a state xxx is gcd⁡ T(x)\gcd\,\mathcal{T}(x)gcdT(x) where T(x)={t≥1:Pt(x,x)>0}\mathcal{T}(x) = \{t \ge 1 : P^t(x,x) > 0\}T(x)={t≥1:Pt(x,x)>0}, and the chain is aperiodic if every state has period 111. A distribution π\piπ is stationary if πP=π\pi P = \piπP=π, and π\piπ and PPP are in detailed balance (the chain is reversible) if π(x)P(x,y)=π(y)P(y,x)\pi(x)P(x,y) = \pi(y)P(y,x)π(x)P(x,y)=π(y)P(y,x) for all x,yx,yx,y.

Trajectory events over a finite horizon are finite sums of path weights: a length-ttt trajectory is a function ω:{0,…,t}→Ω\omega : \{0,\dots,t\} \to \Omegaω:{0,…,t}→Ω, with weight ∏i<tP(ωi,ωi+1)\prod_{i<t} P(\omega_i, \omega_{i+1})∏i<t​P(ωi​,ωi+1​) conditional on its starting state. The tail probability Px{τz+>t}\mathbb{P}_x\{\tau_z^+ > t\}Px​{τz+​>t} of the first hitting time τz+=min⁡{t≥1:Xt=z}\tau_z^+ = \min\{t \ge 1 : X_t = z\}τz+​=min{t≥1:Xt​=z} is the sum of the weights of the trajectories from xxx that avoid zzz at times 1,…,t1,\dots,t1,…,t, and expectations of hitting times are recovered by the tail-sum formula E(Y)=∑t≥0P{Y>t}\mathbb{E}(Y) = \sum_{t\ge0} \mathbb{P}\{Y > t\}E(Y)=∑t≥0​P{Y>t}, formalized as a tsum over ttt.

Formalization targets

Goal

P stochastic and irreducible on a finite nonempty Ω  ⟹  ∃! π, π=πP.\text{$P$ stochastic and irreducible on a finite nonempty $\Omega$} \;\Longrightarrow\; \exists!\, \pi,\ \pi = \pi P .P stochastic and irreducible on a finite nonempty Ω⟹∃!π, π=πP.

This is Corollary 1.17 of the book. It asserts only existence and uniqueness, leaving the finer structure of π\piπ to the milestones; it is the weakest statement on which the rest of the series can stand, which is why it is the goal.

Milestones toward and around the goal

The milestone list follows the book's own route: well-definedness of the period (Lemma 1.6), positivity of some matrix power for irreducible aperiodic chains (Proposition 1.7), finiteness of expected hitting times (Lemma 1.13), existence of a positive stationary distribution together with

π(x) Ex(τx+)=1\pi(x)\,\mathbb{E}_x(\tau_x^+) = 1π(x)Ex​(τx+​)=1

(Proposition 1.14), constancy of harmonic functions (Lemma 1.16), stationarity from detailed balance (Proposition 1.19), the stationary and reversible measure π(x)=deg⁡(x)/2∣E∣\pi(x) = \deg(x)/2|E|π(x)=deg(x)/2∣E∣ of simple random walk on a graph (Examples 1.12 and 1.20), the time reversal P^\hat PP^ and its path-reversal identity (Proposition 1.22), and the random walks on finite groups of Section 2.6 (Propositions 2.12–2.14). From Chapter 2 the list adds the gambler's ruin formulas Pk{Xτ=n}=k/n\mathbb{P}_k\{X_\tau = n\} = k/nPk​{Xτ​=n}=k/n and Ek(τ)=k(n−k)\mathbb{E}_k(\tau) = k(n-k)Ek​(τ)=k(n−k) (Proposition 2.1), the coupon collector expectation n∑k≤n1/kn\sum_{k\le n} 1/kn∑k≤n​1/k and tail bound e−ce^{-c}e−c (Propositions 2.3 and 2.4), and the reflection principle and the bound Pk{τ0>r}≤12k/r\mathbb{P}_k\{\tau_0 > r\} \le 12k/\sqrt{r}Pk​{τ0​>r}≤12k/r​ for simple random walk on Z\mathbb{Z}Z (Lemma 2.18 and Theorem 2.17).

Significance

The result itself. Existence and uniqueness of π\piπ is the pivot on which the entire quantitative theory turns: it defines the target of convergence, and the identity π(x) Ex(τx+)=1\pi(x)\,\mathbb{E}_x(\tau_x^+) = 1π(x)Ex​(τx+​)=1 ties the stationary measure to return times, which later missions use for hitting-time and cover-time results. Detailed balance is the practical tool by which stationary measures of graph and group walks are computed, and the Chapter 2 examples (gambler's ruin, coupon collecting, reflection) are the standard building blocks reused throughout the book — the coupon collector bound, for instance, is exactly the estimate behind the nlog⁡n+cnn \log n + cnnlogn+cn analysis of the top-to-random shuffle in a later mission of this series.

Formalizing it. Mathlib currently has no theory of finite Markov chains: no stochastic-matrix predicate, no stationary distribution, no periodicity, no hitting times. Everything proved in this mission is new formal mathematics, and the definition layer published here (mm_basic, mm_path, mm_classical) is the shared foundation that all twelve subsequent missions of the series import. All results are classical and have textbook proofs; none has a machine-checked proof.

Difficulty

The delicate point is the existence proof. The natural first idea — extract π\piπ from an eigenvector of PTP^{\mathsf T}PT for eigenvalue 111, or invoke a fixed-point theorem — either does not give positivity and nonnegativity without further work, or uses compactness machinery (Brouwer) that is unavailable. The book's proof instead builds π~(y)=Ez(visits to y before τz+)\tilde\pi(y) = \mathbb{E}_z(\text{visits to } y \text{ before } \tau_z^+)π~(y)=Ez​(visits to y before τz+​) and verifies π~P=π~\tilde\pi P = \tilde\piπ~P=π~ by reindexing trajectory sums; formalizing it requires managing infinite series of path sums (summability from the geometric tail bound of Lemma 1.13, exchanging tsum with finite sums, splitting a trajectory at its last step). The uniqueness half is linear algebra via constancy of harmonic functions (Lemma 1.16), which is elementary but requires a maximum-principle argument over a finite state space. The reflection principle and Theorem 2.17 are finite combinatorics on ±1\pm 1±1 paths — the bijection is easy to describe and fiddly to implement.

Formalization scope

States form a Fintype with decidable equality; chains are Matrix V V ℝ with the row-stochasticity predicate IsStochastic; distributions are functions V → ℝ with the predicate IsDist. Everything is distribution-side: no probability space or measure theory is used. The period is formalized as sup⁡{d:d∣t for all t∈T(x)}\sup\{d : d \mid t \text{ for all } t \in \mathcal{T}(x)\}sup{d:d∣t for all t∈T(x)}, which equals gcd⁡T(x)\gcd \mathcal{T}(x)gcdT(x) when T(x)≠∅\mathcal{T}(x) \neq \varnothingT(x)=∅ and takes the junk value 000 otherwise. Expectations of hitting times are tsums of tail probabilities, with the usual junk value 000 for non-summable families — the statements are arranged (e.g. multiplicatively, π(x)⋅Ex(τx+)=1\pi(x)\cdot\mathbb{E}_x(\tau_x^+) = 1π(x)⋅Ex​(τx+​)=1) so that junk values cannot make them vacuously true. Existence statements carry a Nonempty V hypothesis; irreducibility on the empty space is vacuous, and without nonemptiness the goal would be false, not trivial. The coupon collector and the walk on Z\mathbb{Z}Z are presented directly by their driving randomness (uniform draws Fin t → Fin n, uniform sign strings Fin r → Bool), so those probabilities are elementary counting; in particular the reflection principle is stated as an equality of cardinalities of sets of sign strings — this is equivalent to the probabilistic statement because all 2r2^r2r strings are equally likely.

Contributions welcome beyond the milestone list: simp lemmas for the definition layer, the taboo-matrix representation of avoidance probabilities (useful for Lemma 1.13), and any interface lemmas connecting pathWeight sums to matrix powers — these will be reused by every later mission in the series.

Selected references

  • D. A. Levin, Y. Peres, E. L. Wilmer, Markov Chains and Mixing Times, American Mathematical Society, 2009. https://documents.epfl.ch/groups/i/ip/ipg/www/2013-2014/Random_Walks/markovmixing.pdf
  • J. R. Norris, Markov Chains, Cambridge University Press, 1998. https://doi.org/10.1017/CBO9780511810633
  • D. Aldous, J. A. Fill, Reversible Markov Chains and Random Walks on Graphs, 2002 (unfinished monograph). https://www.stat.berkeley.edu/~aldous/RWG/book.html
20 thms4 active usersReviewed
🏆Completed
Markov ChainOperations ResearchOptimization+1·Captain: mikedeng1

Reversibility and Stochastic Networks V: Optimal Capacity Allocation in a Network of QueuesTextbook

Motivation

Chapter 4 of F. P. Kelly's Reversibility and Stochastic Networks (Wiley, 1979) applies the product-form theory of Chapter 3 to concrete systems. Two of its results have numbered statements, and they answer two practical questions.

The first comes from the design of store-and-forward communication networks (telegraph and packet-switched data networks). Messages queue at channels; the designer chooses each channel's capacity subject to a budget, and wants to minimize the delay messages suffer. Because the equilibrium law at every channel is geometric under several different modelling assumptions (§4.1, pp. 95–96), the mean delay has a closed form, and the budget allocation problem becomes a small convex program with an explicit solution. This square-root capacity assignment goes back to L. Kleinrock's work on communication nets (Communication Nets, McGraw-Hill, 1964) and remains the textbook example of optimal design for a network of queues.

The second comes from compartmental models in biology, birth–illness–death processes and manpower planning (§4.5, pp. 113–115). Individuals enter a system as a Poisson stream and move through it independently. Equilibrium results follow from Chapter 3; Theorem 4.2 describes the transient behaviour exactly, starting from an empty system.

Setting

Capacity allocation (§4.1). A network has J≥1J \ge 1J≥1 channels. Channel jjj receives traffic at average rate aj>0a_j > 0aj​>0 and is given capacity ϕj\phi_jϕj​. In equilibrium the number njn_jnj​ of messages at channel jjj has the geometric law (4.1),

P(nj=n)=(1−ajϕj)(ajϕj)n,n=0,1,2,…,P(n_j = n) = \Big(1 - \frac{a_j}{\phi_j}\Big)\Big(\frac{a_j}{\phi_j}\Big)^n, \qquad n = 0, 1, 2, \dots,P(nj​=n)=(1−ϕj​aj​​)(ϕj​aj​​)n,n=0,1,2,…,

which requires ϕj>aj\phi_j > a_jϕj​>aj​. Its mean is aj/(ϕj−aj)a_j/(\phi_j - a_j)aj​/(ϕj​−aj​). Capacity on channel jjj costs fj>0f_j > 0fj​>0 per unit and the total budget is FFF, giving the cost constraint (4.2)

∑jfjϕj=F.\sum_j f_j\phi_j = F.j∑​fj​ϕj​=F.

The mean number of customers in the network is

N(ϕ)=∑jajϕj−aj,N(\phi) = \sum_j \frac{a_j}{\phi_j - a_j},N(ϕ)=j∑​ϕj​−aj​aj​​,

and the feasible set is the set of ϕ∈RJ\phi \in \mathbb{R}^Jϕ∈RJ with ϕj>aj\phi_j > a_jϕj​>aj​ for every jjj that satisfy (4.2). In Lean these are meanNumberInNetwork a φ and FeasibleCapacities a f F. The proof works with the Lagrangian lagrangian a f F y φ =N(ϕ)+y(∑jfjϕj−F)= N(\phi) + y(\sum_j f_j\phi_j - F)=N(ϕ)+y(∑j​fj​ϕj​−F).

Compartmental model (§4.5). Individuals arrive in a Poisson stream of rate ν>0\nu > 0ν>0 at a system of JJJ compartments that is empty at time 000. Let pj(s)p_j(s)pj​(s) be the probability that an individual is in compartment jjj a time sss after its arrival; pj(s)≥0p_j(s) \ge 0pj​(s)≥0 and ∑jpj(s)≤1\sum_j p_j(s) \le 1∑j​pj​(s)≤1, since individuals may leave. Let nj(t)n_j(t)nj​(t) be the number of individuals in compartment jjj at time t>0t > 0t>0, and

αj(t)=∫0tpj(u) du.\alpha_j(t) = \int_0^t p_j(u)\,du.αj​(t)=∫0t​pj​(u)du.

In Lean the model is the predicate IsCompartmentModel P ν t p M T Loc, the counts are compartmentCount M Loc j, and αj(t)\alpha_j(t)αj​(t) is alpha p j t.

Formalization targets

Goal: Theorem 4.1 (p. 97)

If J≥1J \ge 1J≥1, aj>0a_j > 0aj​>0, fj>0f_j > 0fj​>0 and F>∑kakfkF > \sum_k a_k f_kF>∑k​ak​fk​, then

ϕj∗=aj+ajfj∑kakfk⋅F−∑kakfkfj\phi^*_j = a_j + \frac{\sqrt{a_j f_j}}{\sum_k \sqrt{a_k f_k}}\cdot\frac{F - \sum_k a_k f_k}{f_j}ϕj∗​=aj​+∑k​ak​fk​​aj​fj​​​⋅fj​F−∑k​ak​fk​​

is feasible and minimizes NNN over the feasible set, and every other feasible ϕ\phiϕ has N(ϕ)>N(ϕ∗)N(\phi) > N(\phi^*)N(ϕ)>N(ϕ∗).

Milestones toward the goal

  1. The mean of the geometric law (4.1) is aj/(ϕj−aj)a_j/(\phi_j - a_j)aj​/(ϕj​−aj​) (p. 97).
  2. For y>0y > 0y>0 the Lagrangian is minimized over {ϕj>aj}\{\phi_j > a_j\}{ϕj​>aj​} by ϕj=aj+aj/(yfj)\phi_j = a_j + \sqrt{a_j/(y f_j)}ϕj​=aj​+aj​/(yfj​)​ (proof of Theorem 4.1).
  3. The choice 1/y=(F−∑kakfk)/∑kakfk1/\sqrt y = (F - \sum_k a_k f_k)/\sum_k\sqrt{a_k f_k}1/y​=(F−∑k​ak​fk​)/∑k​ak​fk​​ makes that minimizer equal to ϕ∗\phi^*ϕ∗ and feasible for (4.2).

Second result: Theorem 4.2 (pp. 114–115)

The proof's generating-function identity, for zj∈[0,1]z_j \in [0,1]zj​∈[0,1],

E(z1n1(t)⋯zJnJ(t))=∏j=1Jexp⁡[−(1−zj)ναj(t)],E\big(z_1^{n_1(t)}\cdots z_J^{n_J(t)}\big) = \prod_{j=1}^J \exp\big[-(1 - z_j)\nu\alpha_j(t)\big],E(z1n1​(t)​⋯zJnJ​(t)​)=j=1∏J​exp[−(1−zj​)ναj​(t)],

and the theorem itself: n1(t),…,nJ(t)n_1(t), \dots, n_J(t)n1​(t),…,nJ​(t) are independent and nj(t)n_j(t)nj​(t) is Poisson with mean ναj(t)\nu\alpha_j(t)ναj​(t).

Significance

Theorem 4.1 is a closed-form design rule. Every channel first receives the capacity aja_jaj​ needed to carry its traffic; the remaining budget is shared in proportion to ajfj\sqrt{a_j f_j}aj​fj​​, not to the traffic aja_jaj​. Sizing capacity in proportion to the load, which is the obvious rule, is therefore not optimal. The same calculation applies to any network whose stations have the geometric law (4.1), for example a manufacturing job shop (p. 97). The mean number in the network and the mean time a customer spends in it are minimized together.

Theorem 4.2 is the exact transient law of a network of infinite-server queues started empty. It holds however complicated the motion of an individual is, provided individuals move independently, and letting t→∞t \to \inftyt→∞ it recovers the equilibrium Poisson law of §4.5.

Both results are classical and proved in the book. Neither is formalized on Prove2Me. The platform has the geometric equilibrium law of the M/M/1 queue (KellyStochasticNetworks.mm1_equilibrium) but not its mean, and no result on Poisson thinning or marking by independent random locations. The formal work for Theorem 4.1 is a strict-convexity and Lagrangian-sufficiency argument in RJ\mathbb{R}^JRJ. For Theorem 4.2 it is a Poisson marking theorem in measure-theoretic probability. Both pieces are reusable.

Difficulty

For Theorem 4.1, the book's proof sets the partial derivatives of the Lagrangian to zero. A stationary point is not a global minimizer in general, so the formal proof must show that LLL is (strictly) convex on the open region ϕj>aj\phi_j > a_jϕj​>aj​, and it must use Lagrangian sufficiency, not first-order conditions alone. The region is open and the objective is unbounded near its boundary. Feasibility of ϕ∗\phi^*ϕ∗ needs F>∑kakfkF > \sum_k a_k f_kF>∑k​ak​fk​, which the book leaves implicit.

For Theorem 4.2, the steps of the proof that read "conditional on MMM" have to be carried out with measure-theoretic independence. One step averages a product over MMM independent uniform instants. Another sums the Poisson mixture into an exponential. The last turns a factorized generating function into mutual independence of JJJ counts with Poisson marginals. Mathlib has the Poisson distribution (ProbabilityTheory.poissonMeasure) but no marking or thinning theorem, and no uniqueness theorem for multivariate probability generating functions.

Formalization scope

Channels and compartments are indexed by Fin J; all rates, costs and capacities are real numbers.

Theorem 4.1. The statement carries the book's implicit hypotheses explicitly: J≥1J \ge 1J≥1, aj>0a_j > 0aj​>0, fj>0f_j > 0fj​>0, F>∑kakfkF > \sum_k a_k f_kF>∑k​ak​fk​. Stability ϕj>aj\phi_j > a_jϕj​>aj​ is part of the feasible set. The conclusion is global optimality over the feasible set (IsMinOn) together with feasibility of ϕ∗\phi^*ϕ∗, plus strict optimality against every other feasible point. Uniqueness is a slight strengthening of the book's "the optimal allocation is", and it holds by strict convexity. A statement that ϕ∗\phi^*ϕ∗ satisfies (4.2), or that it is a stationary point of the Lagrangian, is not the theorem: those are one-line computations or the proof method, and the goal is stated as global optimality to rule them out.

Theorem 4.2. The model is pinned down as in the proof on p. 115:

  • the number MMM of arrivals in (0,t)(0,t)(0,t) is Poisson with mean νt\nu tνt;
  • an i.i.d. sequence of (arrival instant, location at time ttt) pairs is independent of MMM, and only its first MMM entries are used;
  • each instant is uniform on (0,t)(0,t)(0,t), and an individual arriving at uuu is in compartment jjj at time ttt with probability pj(t−u)p_j(t-u)pj​(t−u), or has left.

"Individuals move independently" is formalized as this conditional independence. The pjp_jpj​ are measurable sub-probabilities, not assumed to sum to one. The conclusion is mutual independence of the JJJ counts (iIndepFun) together with the Poisson probability mass function of each. Infinite time horizons and the point-process description of the system are out of scope.

Useful contributions: a reusable Lagrangian-sufficiency lemma for separable convex objectives under one linear constraint; a Poisson marking (colouring) theorem for finitely many colours; the multivariate generating-function uniqueness lemma for NJ\mathbb{N}^JNJ-valued random vectors.

Selected references

  • F. P. Kelly, Reversibility and Stochastic Networks, John Wiley & Sons, 1979, Chapter 4 (§4.1, pp. 95–97; §4.5, pp. 113–115).
  • L. Kleinrock, Communication Nets: Stochastic Message Flow and Delay, McGraw-Hill, 1964 (reprinted Dover, 1972).
  • J. F. C. Kingman, Poisson Processes, Oxford University Press, 1993 (colouring and marking theorems).
8 thms3 active usersReviewed
🏆Completed
Algorithmic Game TheoryOperations Research·Captain: mikedeng1

Correlated Equilibrium as an Expression of Bayesian Rationality I: Bayes Rationality at Every State Yields Exactly the Correlated Equilibrium DistributionsResearch Paper

Motivation

Nash equilibrium is the standard solution concept for strategic games, but it is usually justified by appeal to what players "would" do once they somehow coordinate on a profile. Correlated equilibrium, introduced by Aumann (J. Math. Econ. 1, 1974), enlarges the set of outcomes by letting players condition their actions on correlated private signals. In Correlated Equilibrium as an Expression of Bayesian Rationality (Econometrica 55, 1987), Aumann gave the concept a decision-theoretic foundation: if the players share a common prior over the states of the world and each player maximizes expected utility given his information at every state, then what they play is a correlated equilibrium — and every correlated equilibrium arises this way. The result is a standard entry point to the epistemic foundations of game theory, and correlated equilibria are central in algorithmic game theory because no-swap-regret learning dynamics converge to them (Foster–Vohra 1997; Hart–Mas-Colell 2000).

Timeline:

  • 1974, Aumann: correlated equilibrium defined, in a measure-theoretic model with subjective probabilities and information σ-fields.
  • 1987, Aumann: the Main Theorem (Bayes rationality at every state under a common prior implies correlated equilibrium play) and its converse, in a finite model with information partitions.

Setting

A game GGG in strategic form has players iii, action sets SiS^iSi, action nnn-tuples s∈S=S1×⋯×Sns\in S=S^1\times\dots\times S^ns∈S=S1×⋯×Sn, and payoffs hi(s)∈Rh^i(s)\in\mathbb Rhi(s)∈R.

A correlated strategy nnn-tuple is a function f:Γ→Sf:\Gamma\to Sf:Γ→S on a finite probability space (Γ,q)(\Gamma,q)(Γ,q), with q≥0q\ge 0q≥0 and ∑γq(γ)=1\sum_\gamma q(\gamma)=1∑γ​q(γ)=1. Its expected payoff is Ehi(f)=∑γq(γ)hi(f(γ))Eh^i(f)=\sum_\gamma q(\gamma)h^i(f(\gamma))Ehi(f)=∑γ​q(γ)hi(f(γ)). For gi:Γ→Sig^i:\Gamma\to S^igi:Γ→Si, the profile (f−i,gi)(f^{-i},g^i)(f−i,gi) replaces player iii's coordinate of fff by gig^igi. The function fff is a correlated equilibrium (Definition 2.1) if

Ehi(f) ≥ Ehi(f−i,gi)(2.2)Eh^i(f)\ \ge\ Eh^i(f^{-i},g^i)\qquad(2.2)Ehi(f) ≥ Ehi(f−i,gi)(2.2)

for every player iii and every gig^igi that is a function of fif^ifi (i.e. gi=φ∘fig^i=\varphi\circ f^igi=φ∘fi). The distribution of fff assigns to each s∈Ss\in Ss∈S the number q{f−1(s)}q\{f^{-1}(s)\}q{f−1(s)}, and a correlated equilibrium distribution (c.e.d.) is the distribution of some correlated equilibrium.

An information system consists of a finite set Ω\OmegaΩ of states of the world, a common prior ppp on Ω\OmegaΩ, an information partition Pi\mathcal P^iPi of Ω\OmegaΩ for each player, and action functions si:Ω→Si\mathbf s^i:\Omega\to S^isi:Ω→Si with s=(s1,…,sn)\mathbf s=(\mathbf s^1,\dots,\mathbf s^n)s=(s1,…,sn), each si\mathbf s^isi constant on the elements of Pi\mathcal P^iPi (each player knows his own action). For a random variable xxx, E(x∣Pi)(ω)E(x\mid\mathcal P^i)(\omega)E(x∣Pi)(ω) is the ppp-average of xxx over the element of Pi\mathcal P^iPi containing ω\omegaω. Player iii is Bayes rational at ω\omegaω if

E(hi(s)∣Pi)(ω) ≥ E(hi(s−i,a)∣Pi)(ω)for every a∈Si.E\big(h^i(\mathbf s)\mid\mathcal P^i\big)(\omega)\ \ge\ E\big(h^i(\mathbf s^{-i},a)\mid\mathcal P^i\big)(\omega)\quad\text{for every }a\in S^i .E(hi(s)∣Pi)(ω) ≥ E(hi(s−i,a)∣Pi)(ω)for every a∈Si.

Formalization targets

Goal: Main Theorem with its converse

Q is a c.e.d. of G  ⟺  ∃ information system (Ω,p,(Pi),s): every player is Bayes rational at every state, and Q(a)=p{s=a} ∀a.Q\ \text{is a c.e.d. of }G\iff\exists\ \text{information system }(\Omega,p,(\mathcal P^i),\mathbf s):\ \text{every player is Bayes rational at every state, and } Q(a)=p\{\mathbf s=a\}\ \forall a.Q is a c.e.d. of G⟺∃ information system (Ω,p,(Pi),s): every player is Bayes rational at every state, and Q(a)=p{s=a} ∀a.

This is the two-sided statement the paper announces in its introduction (p. 2) and closes in Sect. 4d (p. 11): "under Bayesian rationality, the set of all information systems corresponds precisely to the set of all correlated equilibria."

Milestones

  1. Main Theorem, proof: summing the cell-wise inequalities over the partition, Ehi(s−i,gi)≤Ehi(s)Eh^i(\mathbf s^{-i},g^i)\le Eh^i(\mathbf s)Ehi(s−i,gi)≤Ehi(s) for gig^igi constant on the cells of Pi\mathcal P^iPi.
  2. Main Theorem, proof: s\mathbf ss itself is a correlated equilibrium on (Ω,p)(\Omega,p)(Ω,p).
  3. Main Theorem (p. 7): the distribution of s\mathbf ss is a c.e.d.
  4. Sect. 4d: in the system generated by fff (partitions generated by fif^ifi), Bayes rationality everywhere is equivalent to (2.2).
  5. Sect. 4d: every correlated equilibrium is realized by a Bayes-rational information system with the same distribution.

The mission also states, as a supporting lemma without a milestone, the cell-wise step of the proof of the Main Theorem: Bayes rationality at every state gives E(hi(s−i,gi)∣P)≤E(hi(s)∣P)E(h^i(\mathbf s^{-i},g^i)\mid P)\le E(h^i(\mathbf s)\mid P)E(hi(s−i,gi)∣P)≤E(hi(s)∣P) on every cell PPP for gig^igi constant on cells.

Significance

The theorem identifies correlated equilibrium as the outcome of individual Bayesian decision making under a common prior, without any assumption that players randomize or that their choices are independent. It shows that the Nash equilibrium's independence requirement is not implied by rationality alone, and it is the template for later epistemic characterizations of solution concepts. On the computational side, correlated equilibrium distributions form a polytope described by linear inequalities (the companion mission of this series), which is why they are the tractable equilibrium notion in algorithmic game theory.

The result is proved in the paper; it has no machine-checked formalization known to this mission. Formalizing it pins down the exact role of the standing assumptions — finiteness, a common prior, measurability of each player's action with respect to his own partition, rationality at every state — and of the conditional expectation on cells of probability zero. The definitions (information systems, Bayes rationality, correlated equilibria as functions on a finite probability space) are reusable for later formalizations of the paper's Sect. 5 (subjective correlated equilibrium) and of other epistemic results.

Difficulty

The mathematics is short; the difficulty lies in the bookkeeping that the paper's notation hides. The hypothesis is interim (a conditional inequality at each state), while Definition 2.1 is ex ante (an unconditional inequality). Passing between them requires the law of total expectation over a partition whose cells may have probability zero, where conditional expectations are undefined. Deviations in Definition 2.1 are functions of fif^ifi, not arbitrary maps, and must be shown constant on cells, which uses measurability of si\mathbf s^isi. A tempting first idea — that rationality against every fixed action already gives rationality against every deviation — fails without measurability: a player who does not know his own action could be rational at each state against constant deviations while a deviation φ∘si\varphi\circ\mathbf s^iφ∘si varies inside his cells. The converse direction needs an information system that satisfies every axiom of the goal's right-hand side, not just one that is Bayes rational.

Formalization scope

  • Players form a finite type ι with decidable equality; action sets S i are arbitrary types (the paper's finiteness of SiS^iSi is not used); payoffs are h : ι → (∀ i, S i) → ℝ.
  • Probability spaces are finite types with a real weight vector satisfying AGT.IsLottery (from the published definition agt_games). State spaces and the witnessing probability spaces of a c.e.d. range over Type; for finite sets this loses nothing.
  • Partitions are Setoids. The Common Prior Assumption is built into the information system, which has a single prior. Measurability of each action function is a field of the structure.
  • Conditional expectation on a cell is a ratio of finite sums; on a cell of probability zero Lean returns 000, so Bayes rationality at such states holds vacuously. The prior is not required to have full support. This matches the paper, whose argument multiplies each cell inequality by the cell's probability.
  • Deviations in Definition 2.1 are exactly the compositions φ∘fi\varphi\circ f^iφ∘fi; neither all maps nor only constant maps. A c.e.d. requires a genuine probability vector, ruling out the trivializing reading in which the zero weight function witnesses every QQQ.
  • Contributions welcome: proofs of the milestones, a general law-of-total-expectation lemma for finite partitions, and lemmas relating distr to sums over action profiles.

Selected references

  • R. J. Aumann, Correlated Equilibrium as an Expression of Bayesian Rationality, Econometrica 55 (1987), no. 1, 1–18. https://doi.org/10.2307/1911154
  • R. J. Aumann, Subjectivity and Correlation in Randomized Strategies, Journal of Mathematical Economics 1 (1974), 67–96. https://doi.org/10.1016/0304-4068(74)90037-8
  • D. P. Foster and R. V. Vohra, Calibrated Learning and Correlated Equilibrium, Games and Economic Behavior 21 (1997), 40–55. https://doi.org/10.1006/game.1997.0595
  • S. Hart and A. Mas-Colell, A Simple Adaptive Procedure Leading to Correlated Equilibrium, Econometrica 68 (2000), 1127–1150. https://doi.org/10.1111/1468-0262.00153
9 thms3 active usersReviewed
🏆Completed
Analysis·Captain: mikedeng1

Stochastic Dynamic Programming and the Control of Queueing Systems XII: A Tauberian Theorem Linking Abel and Cesàro Means of Nonnegative SeriesTextbook

Motivation

In the theory of Markov decision chains two cost criteria dominate: the discounted cost, in which a cost incurred at time nnn is weighted by αn\alpha^nαn for a discount factor α∈(0,1)\alpha \in (0,1)α∈(0,1), and the average cost, the long-run cost per unit time. A standard route to average cost optimal policies is to solve the discounted problem and let α→1−\alpha \to 1^-α→1−. That route needs a precise link between the two criteria for a single sequence of expected costs u0,u1,u2,⋯≥0u_0, u_1, u_2, \dots \ge 0u0​,u1​,u2​,⋯≥0: the discounted cost multiplied by 1−α1-\alpha1−α on one side, the running average on the other. Appendix A.4 of Sennott, Stochastic Dynamic Programming and the Control of Queueing Systems (Wiley, 1999) supplies this link as Theorem A.4.2, and the book invokes it whenever it passes from discounted to average cost (for instance in Chapter 6, where it is applied to un=Eθ[C(Xn,An)]u_n = E_\theta[C(X_n, A_n)]un​=Eθ​[C(Xn​,An​)]).

The result belongs to classical summability theory.

  • Abelian direction. Convergence of averages implies convergence of the power-series (Abel) means: a consequence of Abel's and Frobenius's theorems on power series (19th century).
  • Tauber (1897). The first converse, under a growth condition on the terms.
  • Hardy and Littlewood (1914). The converse for Cesàro means of nonnegative terms, the "Hardy–Littlewood Tauberian theorem".
  • Karamata (1930). A short proof of that converse by polynomial approximation, reproduced in Titchmarsh, The Theory of Functions (1939, pp. 227–229). The book follows this proof.
  • Widder (1941). The continuous-time (Laplace transform) version.
  • Sennott (1986b). A proof of the discrete statement, cited by the book as its source for Theorem A.4.2.

Setting

Let u0,u1,u2,…u_0, u_1, u_2, \dotsu0​,u1​,u2​,… be nonnegative terms un∈[0,∞]u_n \in [0, \infty]un​∈[0,∞] with u0<∞u_0 < \inftyu0​<∞. For α∈[0,∞)\alpha \in [0, \infty)α∈[0,∞) the power series

U(α)=∑n=0∞αnun∈[0,∞]U(\alpha) = \sum_{n=0}^{\infty} \alpha^n u_n \in [0, \infty]U(α)=n=0∑∞​αnun​∈[0,∞]

always has a sum, possibly +∞+\infty+∞, because its partial sums increase. Its radius of convergence is R=(lim sup⁡nun1/n)−1∈[0,∞]R = (\limsup_{n} u_n^{1/n})^{-1} \in [0,\infty]R=(limsupn​un1/n​)−1∈[0,∞]. The partial sums are wn=∑k=0n−1ukw_n = \sum_{k=0}^{n-1} u_kwn​=∑k=0n−1​uk​ for n≥1n \ge 1n≥1. Two ways of averaging the sequence are compared:

  • the Cesàro means wn/nw_n / nwn​/n, as n→∞n \to \inftyn→∞;
  • the Abel means (1−α)U(α)(1-\alpha) U(\alpha)(1−α)U(α), as α→1−\alpha \to 1^-α→1− (from below).

For the constant sequence un≡Bu_n \equiv Bun​≡B both equal BBB. In the Lean development these are U, radius, w, cesaroMean and abelMean in the namespace SennottDP.Tauberian, all valued in ℝ≥0∞. The proof also uses the function rrr of Fig. A.1, equal to 000 on (0,e−1)(0, e^{-1})(0,e−1) and to 1/x1/x1/x on [e−1,1)[e^{-1}, 1)[e−1,1).

Formalization targets

Goal: Theorem A.4.2

lim inf⁡n→∞wnn≤lim inf⁡α→1−(1−α)U(α)≤lim sup⁡α→1−(1−α)U(α)≤lim sup⁡n→∞wnn,(A.28)\liminf_{n\to\infty} \frac{w_n}{n} \le \liminf_{\alpha\to 1^-} (1-\alpha)U(\alpha) \le \limsup_{\alpha\to 1^-} (1-\alpha)U(\alpha) \le \limsup_{n\to\infty} \frac{w_n}{n}, \tag{A.28}n→∞liminf​nwn​​≤α→1−liminf​(1−α)U(α)≤α→1−limsup​(1−α)U(α)≤n→∞limsup​nwn​​,(A.28)

and the following are equivalent: (i) all terms in (A.28) are equal and finite; (ii) lim⁡nwn/n\lim_n w_n/nlimn​wn​/n exists and is finite; (iii) lim⁡α→1−(1−α)U(α)\lim_{\alpha \to 1^-} (1-\alpha)U(\alpha)limα→1−​(1−α)U(α) exists and is finite.

No growth or boundedness condition on unu_nun​ is assumed: nonnegativity is the only Tauberian condition.

Milestones

  1. Remark A.3.3: the geometric series. R=1R = 1R=1, U(α)=B/(1−α)U(\alpha) = B/(1-\alpha)U(α)=B/(1−α) and B∑nnαn=Bα/(1−α)2B\sum_n n\alpha^n = B\alpha/(1-\alpha)^2B∑n​nαn=Bα/(1−α)2.
  2. Remark A.3.1: inside the radius of convergence, UUU is differentiable term by term, and the derived series has the same radius.
  3. Eq. (A.31): (∑αn)U(α)=∑nαnwn+1=U(α)/(1−α)\big(\sum \alpha^n\big) U(\alpha) = \sum_n \alpha^n w_{n+1} = U(\alpha)/(1-\alpha)(∑αn)U(α)=∑n​αnwn+1​=U(α)/(1−α).
  4. Eq. (A.28): the Abelian inequalities on their own.
  5. Eq. (A.26): ∫01r=1\int_0^1 r = 1∫01​r=1.
  6. Lemma A.4.1: continuous s∗≤r≤ss^* \le r \le ss∗≤r≤s with integrals within ε\varepsilonε of 111.
  7. Eqs. (A.35)–(A.36): if the Abel means tend to L<∞L < \inftyL<∞, then (1−α)∑nαnunf(αn)→L∫01f(1-\alpha)\sum_n \alpha^n u_n f(\alpha^n) \to L\int_0^1 f(1−α)∑n​αnun​f(αn)→L∫01​f for f(x)=xkf(x) = x^kf(x)=xk.
  8. The same statement for every fff continuous on [0,1][0,1][0,1].
  9. The same statement for f=rf = rf=r.
  10. Example A.5.1: a 0/10/10/1 block sequence with lim inf⁡wn/n=1/2\liminf w_n/n = 1/2liminfwn​/n=1/2 and lim sup⁡wn/n=2/3\limsup w_n/n = 2/3limsupwn​/n=2/3 (Choice One) or 111 (Choice Two), for which the middle inequality of (A.28) is strict.

Significance

The result itself. Theorem A.4.2 lets every statement about lim⁡α→1−(1−α)Vα\lim_{\alpha\to1^-}(1-\alpha)V_\alphalimα→1−​(1−α)Vα​ for a discounted value be read as a statement about long-run average cost, and conversely. The strict cases of Example A.5.1 show why the book must work with lim inf⁡\liminfliminf and lim sup⁡\limsuplimsup rather than limits: average costs of policies need not exist as limits, even for bounded costs.

Formalizing it. Mathlib contains Abel's theorem for convergent series (Mathlib/Analysis/Complex/AbelLimit.lean) and Cesàro convergence of convergent sequences, but no Tauberian theorem. The pinned checkout has no file mentioning "Tauberian", and neither does the platform. The mathematics has been proved since 1914. Formalizing it here adds:

  • a machine-checked Hardy–Littlewood–Karamata theorem for nonnegative [0,∞][0,\infty][0,∞]-valued sequences;
  • the lim inf⁡\liminfliminf/lim sup⁡\limsuplimsup comparison (A.28) with infinite values allowed;
  • a reusable component for the average cost chapters of this series.

Difficulty

The Abelian inequalities (A.28) follow from rearranging the series. The converse (iii) ⇒ (ii) does not: the first idea, recovering wn/nw_n/nwn​/n from (1−α)U(α)(1-\alpha)U(\alpha)(1−α)U(α) at α=1−1/n\alpha = 1 - 1/nα=1−1/n, fails, because knowing UUU near 111 controls only weighted averages of all the uku_kuk​, not a sharp truncation. Some positivity argument is unavoidable, since the converse fails for signed sequences (for example un=(−1)nnu_n = (-1)^n nun​=(−1)nn). The sharp truncation at n≈(−ln⁡α)−1n \approx (-\ln\alpha)^{-1}n≈(−lnα)−1 corresponds to the discontinuous weight rrr. Passing from polynomial weights to the jump function rrr requires uniform approximation together with the nonnegativity of the unu_nun​ at every step.

The infinite values add bookkeeping. A single un0=∞u_{n_0} = \inftyun0​​=∞ makes every term of (A.28) infinite, and R<1R < 1R<1 forces U(α)=∞U(\alpha) = \inftyU(α)=∞ on (R,1)(R,1)(R,1).

Formalization scope

Conventions the statements commit to:

  • Values. Terms u:N→[0,∞]u : \mathbb{N} \to [0,\infty]u:N→[0,∞] (ℝ≥0∞) with u0≠∞u_0 \ne \inftyu0​=∞. UUU, the Abel means, wnw_nwn​ and the Cesàro means are ℝ≥0∞-valued, with ℝ≥0∞ sums (no summability side conditions).
  • Limits. lim inf⁡\liminfliminf and lim sup⁡\limsuplimsup are those of the complete lattice [0,∞][0,\infty][0,∞], never real liminf. "Exists and is finite" means convergence in [0,∞][0,\infty][0,∞] to some L≠∞L \ne \inftyL=∞.
  • Discount factor. α∈R≥0\alpha \in \mathbb{R}_{\ge 0}α∈R≥0​, and α→1−\alpha \to 1^-α→1− is the filter 𝓝[<] 1.
  • Index n=0n = 0n=0. n→∞n \to \inftyn→∞ is atTop on N\mathbb{N}N; the value w0/0=0w_0/0 = 0w0​/0=0 is irrelevant.
  • The equivalence is List.TFAE.
  • The function rrr. Its value at the jump is r(e−1)=er(e^{-1}) = er(e−1)=e (Fig. A.1).
  • Continuity. In Lemma A.4.1 and in the continuous-weight step, continuity is required on the closed interval [0,1][0,1][0,1], which is what the Weierstrass approximation step uses. The book writes "(0,1)(0,1)(0,1)".
  • Finite terms. Remark A.3.1 is stated for finite terms, which the book's discussion of the radius presumes.

A trivializing formalization is ruled out. Real-valued lim inf⁡\liminfliminf/lim sup⁡\limsuplimsup would make (A.28) hold by junk values on unbounded sequences; hypotheses such as boundedness of unu_nun​ would replace the theorem by an easier special case; and a definition of "exists and is finite" allowing L=∞L = \inftyL=∞ would make (ii) hold for un=nu_n = nun​=n. None of these is used. As sanity checks: for un≡1u_n \equiv 1un​≡1 all four terms of (A.28) equal 111, and for un=nu_n = nun​=n all equal ∞\infty∞ and (i)–(iii) all fail.

A complete development needs:

  • Cauchy products of [0,∞][0,\infty][0,∞]-valued power series;
  • the Weierstrass approximation theorem, which Mathlib has (polynomialFunctions.topologicalClosure);
  • interval integrals of step-like functions;
  • lim inf⁡\liminfliminf/lim sup⁡\limsuplimsup manipulation along 𝓝[<] 1.

Reusable beyond this mission: the ℝ≥0∞ power-series toolkit and the Karamata argument for general continuous weights (milestone 8). Contributions welcome: proofs of any milestone, and alternative proofs of (iii) ⇒ (ii).

Selected references

  • L. I. Sennott, Stochastic Dynamic Programming and the Control of Queueing Systems, Wiley, 1999, Appendix A.3–A.5, pp. 279–287. https://doi.org/10.1002/9780470317037
  • E. C. Titchmarsh, The Theory of Functions, 2nd ed., Oxford University Press, 1939, pp. 227–229 (Karamata's proof).
  • J. Karamata, "Über die Hardy–Littlewoodschen Umkehrungen des Abelschen Stetigkeitssatzes", Mathematische Zeitschrift 32 (1930), 319–320.
  • G. H. Hardy and J. E. Littlewood, "Tauberian theorems concerning power series and Dirichlet's series whose coefficients are positive", Proc. London Math. Soc. (2) 13 (1914), 174–191.
  • D. V. Widder, The Laplace Transform, Princeton University Press, 1941.
  • L. I. Sennott, "A new condition for the existence of optimum stationary policies in average cost Markov decision processes — unbounded cost case", Proc. 25th IEEE Conference on Decision and Control, Athens, 1986, pp. 1719–1721 (cited by the book as Sennott 1986b).
  • T. M. Liggett and S. A. Lippman, "Short notes: Stochastic games with perfect information and time average payoff", SIAM Review 11 (1969), 604–607. https://doi.org/10.1137/1011093
14 thms3 active usersReviewed
PreviousPage 2 of 11Next

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me