Prove2Me
Navigate
DiscoverFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

Statistics

144 missions · 100 completed

The mathematical discipline of drawing inferences from data under uncertainty: estimation, hypothesis testing, prediction, and the quantification of confidence. Grounded in probability, it spans classical and Bayesian inference, experimental design, and modern high-dimensional and nonparametric theory, asking what data can reveal and with what guarantees.

Missions

Open44Completed100All144
Machine LearningOptimizationProbability·Captain: mikedeng1

Variance-based Regularization with Convex Objectives III: Localized-Rademacher Risk Bounds for the Robust MinimizerResearch Paper

Why variance-regularized risk bounds

In statistical learning, one picks a function fff from a class F\mathcal FF to make the population risk E[f]\mathbb E[f]E[f] small, with access only to an i.i.d. sample x1,…,xnx_1,\dots,x_nx1​,…,xn​ from an unknown distribution PPP. Empirical risk minimization replaces E[f]\mathbb E[f]E[f] by the empirical mean EP^n[f]\mathbb E_{\widehat P_n}[f]EPn​​[f], and its classical guarantees decay like 1/n1/\sqrt n1/n​ regardless of how concentrated fff is. Bernstein-type inequalities show that the deviation of EP^n[f]\mathbb E_{\widehat P_n}[f]EPn​​[f] from E[f]\mathbb E[f]E[f] scales with the standard deviation of fff, so a procedure that minimizes "empirical risk plus a standard-deviation penalty" can, in principle, achieve faster rates when the variance at the optimum is small (Maurer and Pontil, 2009). The penalized objective is non-convex even when every fff is convex in its parameters, which makes it hard to optimize.

J. C. Duchi and H. Namkoong (arXiv:1610.02581v3, 2017) replace the penalty by a distributionally robust objective: the worst-case risk over all reweightings of the sample within a χ2\chi^2χ2-divergence ball. This objective is convex whenever the losses are, and (Theorem 1 of the paper) it equals the empirical mean plus a standard-deviation penalty up to an error of order 1/n1/n1/n. This mission formalizes the paper's guarantee for the minimizer of that robust objective in terms of localized Rademacher complexities (Section 3.2, Theorem 4), the sharpest of the paper's three generalization analyses. It is the third of four missions on the paper.

Setting

Let PPP be a probability measure on a measurable space X\mathcal XX and x1,…,xnx_1,\dots,x_nx1​,…,xn​, n≥1n\ge1n≥1, an i.i.d. sample from PPP with empirical distribution P^n\widehat P_nPn​. Let M≥1M\ge1M≥1 and let F\mathcal FF be a collection of measurable functions f:X→[0,M]f:\mathcal X\to[0,M]f:X→[0,M] (losses).

  • The χ2\chi^2χ2 ball of radius ρ≥0\rho\ge0ρ≥0 is the set Pn\mathcal P_nPn​ of weight vectors p∈Rnp\in\mathbb R^np∈Rn with pi≥0p_i\ge0pi​≥0, ∑ipi=1\sum_ip_i=1∑i​pi​=1 and 12∑i(npi−1)2≤ρ\frac12\sum_i(np_i-1)^2\le\rho21​∑i​(npi​−1)2≤ρ; equivalently, the distributions PPP on the sample with Dϕ(P∥P^n)≤ρ/nD_\phi(P\|\widehat P_n)\le\rho/nDϕ​(P∥Pn​)≤ρ/n for ϕ(t)=12(t−1)2\phi(t)=\frac12(t-1)^2ϕ(t)=21​(t−1)2.
  • The robust risk of fff is sup⁡P: Dϕ(P∥P^n)≤ρ/nEP[f]=sup⁡p∈Pn∑ipif(xi)\sup_{P:\,D_\phi(P\|\widehat P_n)\le\rho/n}\mathbb E_P[f]=\sup_{p\in\mathcal P_n}\sum_ip_if(x_i)supP:Dϕ​(P∥Pn​)≤ρ/n​EP​[f]=supp∈Pn​​∑i​pi​f(xi​), and a robust minimizer f^\widehat ff​ minimizes it over F\mathcal FF.
  • The empirical Rademacher complexity is Rn(F)=Eε[sup⁡f∈F1n∑iεif(xi)]\mathfrak R_n(\mathcal F)=\mathbb E_\varepsilon\big[\sup_{f\in\mathcal F}\frac1n\sum_i\varepsilon_if(x_i)\big]Rn​(F)=Eε​[supf∈F​n1​∑i​εi​f(xi​)] with i.i.d. uniform signs εi∈{−1,1}\varepsilon_i\in\{-1,1\}εi​∈{−1,1}, and E[Rn(F)]\mathbb E[\mathfrak R_n(\mathcal F)]E[Rn​(F)] averages it over the sample.
  • A function ψ:R+→R+\psi:\mathbb R_+\to\mathbb R_+ψ:R+​→R+​ is sub-root if it is nonnegative, nondecreasing, and r↦ψ(r)/rr\mapsto\psi(r)/\sqrt rr↦ψ(r)/r​ is nonincreasing on r>0r>0r>0.
  • The localization inequality (20) asks that, for all r≥0r\ge0r≥0,
ψn(r) ≥ E[Rn({cf:f∈F, c∈[0,1], E[c2f2]≤r})],\psi_n(r)\ \ge\ \mathbb E\big[\mathfrak R_n(\{cf : f\in\mathcal F,\ c\in[0,1],\ \mathbb E[c^2f^2]\le r\})\big],ψn​(r) ≥ E[Rn​({cf:f∈F, c∈[0,1], E[c2f2]≤r})],

with ψn\psi_nψn​ sub-root, and rn⋆>0r_n^\star>0rn⋆​>0 is a point with rn⋆≥ψn(rn⋆)r_n^\star\ge\psi_n(r_n^\star)rn⋆​≥ψn​(rn⋆​).

Formalization targets

Goal: Theorem 4, inequality (23), as its proof establishes it

Let 0<t<n0<t<n0<t<n and let ρ\rhoρ satisfy (21): ρn≥8(45Mn(t+log⁡⌈log⁡nt⌉)+18rn⋆)\frac\rho n\ge8\big(\frac{45M}n\big(t+\log\lceil\log\frac nt\rceil\big)+18r_n^\star\big)nρ​≥8(n45M​(t+log⌈logtn​⌉)+18rn⋆​). With probability at least 1−4e−t1-4e^{-t}1−4e−t, every robust minimizer f^\widehat ff​ satisfies

E[f^] ≤ (1+22ρn)inf⁡f∈F(E[f]+182ρ45nVar(f))+(14+62ρn)M(3ρ+t)n.\mathbb E[\widehat f]\ \le\ \Big(1+2\sqrt{\tfrac{2\rho}n}\Big)\inf_{f\in\mathcal F}\Big(\mathbb E[f]+\sqrt{\tfrac{182\rho}{45n}\mathrm{Var}(f)}\Big)+\Big(14+6\sqrt{\tfrac{2\rho}n}\Big)\frac{M(3\rho+t)}n .E[f​] ≤ (1+2n2ρ​​)f∈Finf​(E[f]+45n182ρ​Var(f)​)+(14+6n2ρ​​)nM(3ρ+t)​.

Milestones

In attack order: Bousquet's form of Talagrand's inequality (Lemma B.2); the elementary root bound (Lemma D.4); the contraction principle (Lemma D.5, a published theorem); the uniform Bernstein inequality with Rademacher complexity (Lemma D.1); its localized version in terms of rn⋆r_n^\starrn⋆​ (Lemma D.2); localized second-moment bounds (Lemma D.3); the deterministic expansion (10) of Theorem 1,

(2ρnsn2−2Mρn)+≤sup⁡PEP[Z]−EP^n[Z]≤2ρnsn2;\Big(\sqrt{\tfrac{2\rho}n s_n^2}-\tfrac{2M\rho}n\Big)_+\le\sup_{P}\mathbb E_P[Z]-\mathbb E_{\widehat P_n}[Z]\le\sqrt{\tfrac{2\rho}ns_n^2};(n2ρ​sn2​​−n2Mρ​)+​≤Psup​EP​[Z]−EPn​​[Z]≤n2ρ​sn2​​;

and the uniform bound (22): with probability at least 1−2e−t1-2e^{-t}1−2e−t, for all f∈Ff\in\mathcal Ff∈F,

E[f]≤(1+22ρn)sup⁡P: Dϕ(P∥P^n)≤ρ/nEP[f]+(13+42ρn)Mρn.\mathbb E[f]\le\Big(1+2\sqrt{\tfrac{2\rho}n}\Big)\sup_{P:\,D_\phi(P\|\widehat P_n)\le\rho/n}\mathbb E_P[f]+\Big(13+4\sqrt{\tfrac{2\rho}n}\Big)\frac{M\rho}n .E[f]≤(1+2n2ρ​​)P:Dϕ​(P∥Pn​)≤ρ/nsup​EP​[f]+(13+4n2ρ​​)nMρ​.

Significance

The bound (23) says that the robust minimizer competes with the best trade-off between risk and standard deviation in the class, and that the complexity of the class enters only through the fixed point rn⋆r_n^\starrn⋆​ of a localized complexity bound. For bounded VC classes rn⋆r_n^\starrn⋆​ is of order dlog⁡(n/d)n\frac{d\log(n/d)}nndlog(n/d)​ (Bartlett, Bousquet and Mendelson, 2005, Corollary 3.7), so when the optimal function has small variance the excess risk is of order ρ/n\rho/nρ/n, faster than the 1/n1/\sqrt n1/n​ of uniform covering arguments; and localized complexities apply to classes, such as balls of reproducing kernel Hilbert spaces, whose covering numbers are too large for the covering-number analysis of the paper's Theorem 3 (mission II of this series).

The paper's result is proved, not open. No part of it, and none of the localized-complexity machinery of Bartlett, Bousquet and Mendelson, is formalized in Lean or Mathlib to our knowledge. The mission produces a checked version of the theorem with every constant explicit and, along the way, the localization lemmas D.1–D.3, which are reusable for any localized-complexity analysis. Reading the proof also exposed three arithmetic slips in the printed statements; the mission states what the proof establishes (see Formalization scope).

Difficulty

The obvious route applies a uniform concentration inequality to F\mathcal FF and then a Bernstein bound to each fff. Talagrand's inequality applied to the whole class gives a deviation governed by the largest variance in the class and by the global complexity E[Rn(F)]\mathbb E[\mathfrak R_n(\mathcal F)]E[Rn​(F)], which yields only 1/n1/\sqrt n1/n​ rates. Obtaining a deviation that scales with each function's own second moment requires peeling the class into shells of comparable second moment and a fixed-point argument on the sub-root bound, with a union bound whose cost appears as log⁡⌈log⁡nt⌉\log\lceil\log\frac nt\rceillog⌈logtn​⌉. The two directions of the localized inequalities (population to sample, and sample to population for second moments) must then be combined with the deterministic expansion (10) while keeping the constants explicit. A further subtlety is the self-normalized rescaling f↦r/(E[f2]∨r) ff\mapsto\sqrt{r/(\mathbb E[f^2]\vee r)}\,ff↦r/(E[f2]∨r)​f, which differs from the variance normalization of Bartlett et al. and is what makes the bound compatible with the robust objective.

Formalization scope

Lean conventions. The sample is the coordinate map of the product measure PnP^nPn on Fin n → X. Distributions on the sample are weight vectors in the χ2\chi^2χ2 ball; the robust risk is the real supremum over that ball (attained, since the ball is nonempty and compact for n≥1n\ge1n≥1, ρ≥0\rho\ge0ρ≥0). Population means and variances are ∫ x, f x ∂P and ProbabilityTheory.variance f P for measurable bounded fff; empirical means and variances are normalized by 1/n1/n1/n. The empirical Rademacher complexity is the published UnderstandingML_Rademacher definition evaluated on {(f(x1),…,f(xn))}\{(f(x_1),\dots,f(x_n))\}{(f(x1​),…,f(xn​))}. Its expectation is a Bochner integral, and every hypothesis that bounds it also asserts that the integrand is integrable: otherwise the integral is 000, (20) would hold for free, and the theorem would be false. Probability bounds are stated for the failure event under PnP^nPn (an outer measure when the event is not measurable). The goal speaks about every minimizer of the robust risk, so it is not vacuous when the set of minimizers is empty. The condition rn⋆>0r_n^\star>0rn⋆​>0 is part of the page's "root" (and the proof divides by rn⋆\sqrt{r_n^\star}rn⋆​​); with rn⋆=0r_n^\star=0rn⋆​=0 allowed, ψ(r)=r\psi(r)=\sqrt rψ(r)=r​ would remove the complexity term from (21). The condition t<nt<nt<n makes log⁡⌈log⁡nt⌉\log\lceil\log\frac nt\rceillog⌈logtn​⌉ defined.

Corrections of printed statements, each recorded in the item's docstring and Formalization Note (the milestone texts stay verbatim):

  • (22) is stated with probability 1−2e−t1-2e^{-t}1−2e−t; the paper prints 1−e−t1-e^{-t}1−e−t, and its proof (p. 41) concludes 1−2e−t1-2e^{-t}1−2e−t.
  • (23) is stated with probability 1−4e−t1-4e^{-t}1−4e−t (printed 1−3e−t1-3e^{-t}1−3e−t; the proof adds two fixed-fff events to the two of (22)) and with 182ρ45n\frac{182\rho}{45n}45n182ρ​ (printed 91ρ45n\frac{91\rho}{45n}45n91ρ​; the proof's step ρ+t≤91ρ/45\sqrt\rho+\sqrt t\le\sqrt{91\rho/45}ρ​+t​≤91ρ/45​ multiplies 2Var(f)/n\sqrt{2\mathrm{Var}(f)/n}2Var(f)/n​).
  • Lemma D.3 is stated with the additive term 72M2(1+η)rn⋆+(4(1+η)+143)M2tn72M^2(1+\eta)r_n^\star+(4(1+\eta)+\frac{14}3)\frac{M^2t}n72M2(1+η)rn⋆​+(4(1+η)+314​)nM2t​ and, in the reversed direction, the coefficient 1+11+η1+\frac1{1+\eta}1+1+η1​, as its proof yields (printed: Mtn(4+73M)\frac{Mt}n(4+\frac73M)nMt​(4+37​M) and 1+η1+η1+\frac\eta{1+\eta}1+1+ηη​), under Theorem 4's standing hypothesis M≥1M\ge1M≥1.
  • Lemma D.5 is linked to the published contraction lemma UnderstandingML.contraction_lemma, which states it at a fixed sample for nonempty bounded classes and allows a different Lipschitz map per coordinate.

Contributions welcome: proofs of the milestones in any order; Lemma B.2 (Bousquet's inequality) is the deepest single ingredient and is reusable well beyond this mission, as are the peeling Lemma D.1 and the sub-root fixed-point Lemma D.2.

Selected references

  • J. C. Duchi and H. Namkoong, Variance-based regularization with convex objectives, arXiv:1610.02581v3, 2017. https://arxiv.org/abs/1610.02581
  • P. L. Bartlett, O. Bousquet and S. Mendelson, Local Rademacher complexities, Annals of Statistics 33(4), 2005. https://doi.org/10.1214/009053605000000282
  • O. Bousquet, A Bennett concentration inequality and its application to suprema of empirical processes, Comptes Rendus Mathématique 334(6), 2002. https://doi.org/10.1016/S1631-073X(02)02292-6
  • A. Maurer and M. Pontil, Empirical Bernstein bounds and sample variance penalization, COLT 2009. https://arxiv.org/abs/0907.3740
  • M. Ledoux and M. Talagrand, Probability in Banach Spaces, Springer, 1991. https://doi.org/10.1007/978-3-642-20212-4
14 thms4 active usersReviewed
Operations ResearchProbabilityStochastic Systems·Captain: mikedeng1

Weak Convergence and Optimal Scaling of Random Walk Metropolis Algorithms: Langevin Diffusion Limit of the First CoordinateResearch Paper

Motivation

The random walk Metropolis algorithm is one of the most widely used Markov chain Monte Carlo methods for sampling from a density known up to a constant. Its one tuning parameter is the variance of the Gaussian proposal. If the variance is too small, the chain accepts almost every move but barely moves. If it is too large, it proposes long jumps that are almost always rejected. Practitioners need a rule for choosing it, and the rule has to work in high dimension, where both failure modes are severe.

Roberts, Gelman and Gilks (Ann. Appl. Probab. 7(1), 1997) gave the first rigorous answer for product targets. As the dimension grows, one coordinate of the suitably speeded-up chain converges to a Langevin diffusion. The speed of that diffusion is an explicit function of the proposal scale, and maximising it gives the rule "tune the proposal so that about 23% of proposals are accepted". This rule, and the 2.38/√I scaling behind it, is now standard advice in applied Bayesian statistics.

Setting

Let f:R→Rf:\mathbb R\to\mathbb Rf:R→R be a target density: positive, C2C^2C2, integrating to one, with f′/ff'/ff′/f Lipschitz, and satisfying the moment conditions (A1) Ef[(f′/f)8]<∞\mathbb E_f[(f'/f)^8]<\inftyEf​[(f′/f)8]<∞ and (A2) Ef[(f′′/f)4]<∞\mathbb E_f[(f''/f)^4]<\inftyEf​[(f′′/f)4]<∞. Here Ef[g(X)]=∫g(x)f(x) dx\mathbb E_f[g(X)]=\int g(x)f(x)\,dxEf​[g(X)]=∫g(x)f(x)dx. In dimension n≥2n\ge2n≥2 the target is the product density πn(x)=∏i=1nf(xi)\pi_n(x)=\prod_{i=1}^n f(x_i)πn​(x)=∏i=1n​f(xi​) on Rn\mathbb R^nRn.

Fix a scale l>0l>0l>0 and set σn2=l2/(n−1)\sigma_n^2=l^2/(n-1)σn2​=l2/(n−1). The random walk Metropolis chain Xn=(X0n,X1n,… )X^n=(X^n_0,X^n_1,\dots)Xn=(X0n​,X1n​,…) moves as follows. From Xm−1nX^n_{m-1}Xm−1n​ it proposes Y∼N(Xm−1n,σn2In)Y\sim N(X^n_{m-1},\sigma_n^2I_n)Y∼N(Xm−1n​,σn2​In​). It sets Xmn=YX^n_m=YXmn​=Y with probability α(Xm−1n,Y)=1∧πn(Y)/πn(Xm−1n)\alpha(X^n_{m-1},Y)=1\wedge\pi_n(Y)/\pi_n(X^n_{m-1})α(Xm−1n​,Y)=1∧πn​(Y)/πn​(Xm−1n​), and Xmn=Xm−1nX^n_m=X^n_{m-1}Xmn​=Xm−1n​ otherwise. The chain starts from πn\pi_nπn​, which is stationary for it. The speeded-up first coordinate is Utn=X⌊nt⌋,1nU^n_t=X^n_{\lfloor nt\rfloor,1}Utn​=X⌊nt⌋,1n​ for t≥0t\ge0t≥0.

Let Φ\PhiΦ be the standard normal distribution function, and define the roughness I=Ef[(f′(X)/f(X))2]I=\mathbb E_f[(f'(X)/f(X))^2]I=Ef​[(f′(X)/f(X))2]. The speed and the limiting acceptance rate are

h(l)=2l2 Φ ⁣(−lI2),a(l)=2 Φ ⁣(−lI2).h(l)=2l^2\,\Phi\!\Big(-\frac{l\sqrt I}{2}\Big),\qquad a(l)=2\,\Phi\!\Big(-\frac{l\sqrt I}{2}\Big).h(l)=2l2Φ(−2lI​​),a(l)=2Φ(−2lI​​).

The Langevin generator is GV(x)=h(l)[12V′′(x)+12(log⁡f)′(x)V′(x)]GV(x)=h(l)\big[\tfrac12V''(x)+\tfrac12(\log f)'(x)V'(x)\big]GV(x)=h(l)[21​V′′(x)+21​(logf)′(x)V′(x)]. It generates the Langevin diffusion dUt=h(l)1/2dBt+h(l)f′(Ut)2f(Ut)dtdU_t=h(l)^{1/2}dB_t+h(l)\frac{f'(U_t)}{2f(U_t)}dtdUt​=h(l)1/2dBt​+h(l)2f(Ut​)f′(Ut​)​dt.

Formalization targets

Goal: Theorem 1.1

As n→∞n\to\inftyn→∞,

Un⇒U,U^n\Rightarrow U,Un⇒U,

where ⇒\Rightarrow⇒ denotes weak convergence in the Skorokhod topology, U0U_0U0​ has density fff, and UUU is the Langevin diffusion with speed h(l)h(l)h(l). The limit is asserted to exist. No constants appear in the statement beyond those the model defines.

Milestones: the proof

  1. Lemma 2.1. The stationary chain stays in the sets Fn={∣Rn−I∣<n−1/8}∩{∣Sn−I∣<n−1/8}F_n=\{|R_n-I|<n^{-1/8}\}\cap\{|S_n-I|<n^{-1/8}\}Fn​={∣Rn​−I∣<n−1/8}∩{∣Sn​−I∣<n−1/8} up to time ttt with probability tending to one. Here RnR_nRn​ and SnS_nSn​ are the empirical averages of ((log⁡f)′)2((\log f)')^2((logf)′)2 and −(log⁡f)′′-(\log f)''−(logf)′′ over coordinates 2,…,n2,\dots,n2,…,n.
  2. Proposition 2.2. ∣1∧ex−1∧ey∣≤∣x−y∣|1\wedge e^x-1\wedge e^y|\le|x-y|∣1∧ex−1∧ey∣≤∣x−y∣.
  3. Lemma 2.3. sup⁡x∈FnE∣Wn∣→0\sup_{x\in F_n}\mathbb E|W_n|\to0supx∈Fn​​E∣Wn​∣→0, where WnW_nWn​ is the second-order part of the log acceptance ratio.
  4. Proposition 2.4. E[1∧eA]=Φ(μ/σ)+eμ+σ2/2Φ(−σ−μ/σ)\mathbb E[1\wedge e^A]=\Phi(\mu/\sigma)+e^{\mu+\sigma^2/2}\Phi(-\sigma-\mu/\sigma)E[1∧eA]=Φ(μ/σ)+eμ+σ2/2Φ(−σ−μ/σ) for A∼N(μ,σ2)A\sim N(\mu,\sigma^2)A∼N(μ,σ2).
  5. Lemma 2.5. lim sup⁡nsup⁡x1n∣E[V(Y1)−V(x1)]∣<∞\limsup_n\sup_{x_1}n|\mathbb E[V(Y_1)-V(x_1)]|<\inftylimsupn​supx1​​n∣E[V(Y1​)−V(x1​)]∣<∞ for V∈Cc∞V\in C_c^\inftyV∈Cc∞​.
  6. Lemma 2.6. The discrete generator GnV(x)=n E[(V(Y)−V(x))α(x,Y)]G_nV(x)=n\,\mathbb E[(V(Y)-V(x))\alpha(x,Y)]Gn​V(x)=nE[(V(Y)−V(x))α(x,Y)] converges to GVGVGV uniformly on FnF_nFn​, for V∈Cc∞V\in C_c^\inftyV∈Cc∞​ a function of the first coordinate (stated with bounded (log⁡f)′′′(\log f)'''(logf)′′′, the assumption its proof uses).

Milestones: the optimal-scaling corollary

  1. Corollary 1.2 (i). an(l)=∬πn(x)α(x,y)qn(x,y) dx dy→a(l)a_n(l)=\iint\pi_n(x)\alpha(x,y)q_n(x,y)\,dx\,dy\to a(l)an​(l)=∬πn​(x)α(x,y)qn​(x,y)dxdy→a(l).
  2. Corollary 1.2 (ii). hhh is maximised at l^=2.38/I\hat l=2.38/\sqrt Il^=2.38/I​, with a(l^)=0.23a(\hat l)=0.23a(l^)=0.23 and h(l^)=1.3/Ih(\hat l)=1.3/Ih(l^)=1.3/I, to the printed precision.

Significance

Theorem 1.1 shows that, run for nnn times as many steps, the chain in dimension nnn looks like a fixed one-dimensional diffusion. The algorithm's cost therefore grows linearly in dimension, and its efficiency is measured by the single number h(l)h(l)h(l). Corollary 1.2 turns this into the 0.234 acceptance-rate heuristic and the 2.38/I2.38/\sqrt I2.38/I​ scaling. The same diffusion-limit method has since been applied to the Metropolis-adjusted Langevin algorithm, to Hamiltonian Monte Carlo and to non-product targets.

The theorem is proved on paper; it has no machine-checked proof. A formal development would give the first verified diffusion limit of an MCMC algorithm. It would also yield reusable components: the Metropolis chain on Rn\mathbb R^nRn as a measurable random mapping, a martingale-problem characterisation of one-dimensional diffusions, and a Gaussian computation (Proposition 2.4) that recurs throughout the optimal-scaling literature.

Difficulty

The obvious approach, a Taylor expansion of the log acceptance ratio, gives a sum of n−1n-1n−1 terms of size 1/n1/n1/n. That sum does not concentrate uniformly over the state space, since the coordinates 2,…,n2,\dots,n2,…,n are arbitrary. The expansion is controlled only on the sets FnF_nFn​, where the empirical averages RnR_nRn​ and SnS_nSn​ are close to III. The limit therefore holds only after showing that the chain rarely leaves FnF_nFn​ over a time horizon of ntntnt steps. A pointwise law of large numbers is not enough for that, because the bound has to survive a union over ntntnt steps. Passing from generator convergence on a set of high probability to weak convergence of processes requires the Ethier–Kurtz convergence theory: a core for the limit generator, and convergence of processes that are not themselves Markov. None of this theory is in Mathlib.

Formalization scope

All declarations live in the namespace Roberts1997.RWM. The following conventions are fixed.

  • Vectors are Fin n → ℝ, and the paper's first coordinate x1x_1x1​ is index 0. Its coordinates 2,…,n2,\dots,n2,…,n are the indices i ≠ 0.
  • σn2=l2/(n−1)\sigma_n^2=l^2/(n-1)σn2​=l2/(n−1) is computed in R\mathbb RR. All statements concern n≥2n\ge2n≥2 or large nnn.
  • l>0l>0l>0 is assumed. The paper leaves it implicit, but h(−l)≠h(l)h(-l)\ne h(l)h(−l)=h(l).
  • "fff is a density" is read as ∫f=1\int f=1∫f=1. The moment conditions are read as integrability of (f′/f)8f(f'/f)^8f(f′/f)8f and (f′′/f)4f(f''/f)^4f(f′′/f)4f. The standing assumption "f′/ff'/ff′/f is Lipschitz" (p. 111) is carried by every statement.
  • The chain is built as a random mapping on an explicit probability space: x0∼πnx_0\sim\pi_nx0​∼πn​, with i.i.d. standard normal innovations and uniform acceptance variables. Theorem 1.1's initial condition (components i.i.d. fff, shared across dimensions) is read as "the nnn-th chain starts from πn\pi_nπn​", since weak convergence depends only on the law of each UnU^nUn.
  • "UUU satisfies the Langevin SDE" is read as "the law of UUU solves the martingale problem for GGG on Cc∞C_c^\inftyCc∞​, with continuous paths and initial law f(x) dxf(x)\,dxf(x)dx". This is equivalent by Ethier–Kurtz (1986), Ch. 5, Prop. 3.1 and Thm 3.3, and follows the platform definition EthierKurtz_IsContinuousDiffusionLaw.
  • "Un⇒UU^n\Rightarrow UUn⇒U" is read as the existence of an almost-sure coupling in which càdlàg copies of the UnU^nUn converge to a continuous Langevin path uniformly on compact time intervals. For a continuous limit this is equivalent to weak convergence in DR[0,∞)D_{\mathbb R}[0,\infty)DR​[0,∞), by Skorokhod's representation theorem and Ethier–Kurtz Ch. 3, Thm 1.8, Prop. 5.3 and Prop. 7.1. It follows the platform encoding of Ethier–Kurtz Theorem 7.4.1.
  • "sup⁡→0\sup\to0sup→0" and "lim sup⁡sup⁡<∞\limsup\sup<\inftylimsupsup<∞" are stated as eventual uniform bounds. This avoids real suprema, whose value on an unbounded set is a default.
  • In Lemma 2.6, "as d→∞d\to\inftyd→∞" is a misprint for n→∞n\to\inftyn→∞, and "2f(Ut)2f(Ut)2f(Ut)" in (1.2) is read as 2f(Ut)2f(U_t)2f(Ut​).
  • Corollary 1.2 (ii) is stated for an arbitrary constant I>0I>0I>0. "To two decimal places" is read as explicit rounding intervals: 1.31.31.3 is read to one decimal, and all maximisers over l>0l>0l>0 are covered.

The goal cannot be satisfied trivially. The limit law QQQ must exist, and it must be a probability measure whose initial marginal is f(x) dxf(x)\,dxf(x)dx, so the zero measure is excluded. The coupled copies must carry exactly the laws of the paths UnU^nUn, not an arbitrary process with the same one-time marginals.

The statements carry the paper's hypotheses, with one exception. The printed proof of Lemma 2.6 bounds sup⁡z∣(log⁡f)′′′(z)∣\sup_z|(\log f)'''(z)|supz​∣(logf)′′′(z)∣, which Theorem 1.1 does not assume, and under C2C^2C2 alone the uniform convergence over FnF_nFn​ claimed by Lemma 2.6 fails (narrow spikes of (log⁡f)′′(\log f)''(logf)′′ far out let a positive fraction of the coordinates shift the log acceptance ratio by a constant while RnR_nRn​ and SnS_nSn​ stay close to III). Lemma 2.6 is therefore stated with the proof's own assumption, f∈C3f\in C^3f∈C3 with (log⁡f)′′′(\log f)'''(logf)′′′ bounded, named as an addition. Theorem 1.1 and the other results keep the paper's hypotheses.

A complete development needs several pieces not yet available: path spaces and the Skorokhod topology (or the coupling reading), the martingale problem and its well-posedness for Lipschitz drift, and the Ethier–Kurtz theorem on convergence of generators on sets of high probability. Proofs of the Gaussian milestones (Propositions 2.2 and 2.4, Lemma 2.5) and of Corollary 1.2 (ii) are independent of this infrastructure and are welcome contributions.

Selected references

  • G. O. Roberts, A. Gelman, W. R. Gilks, Weak convergence and optimal scaling of random walk Metropolis algorithms, Ann. Appl. Probab. 7(1), 110–120, 1997. https://doi.org/10.1214/aoap/1034625254
  • S. N. Ethier, T. G. Kurtz, Markov Processes: Characterization and Convergence, Wiley, 1986. https://doi.org/10.1002/9780470316658
  • A. Gelman, G. O. Roberts, W. R. Gilks, Efficient Metropolis jumping rules, Bayesian Statistics 5, Oxford University Press, 599–607, 1996.
  • G. O. Roberts, J. S. Rosenthal, Optimal scaling for various Metropolis–Hastings algorithms, Statistical Science 16(4), 351–367, 2001. https://doi.org/10.1214/ss/1015346320
15 thms4 active usersReviewed
Machine LearningOperations ResearchOptimization·Captain: mikedeng1

Generalization Bounds in the Predict-then-Optimize Framework I: Natarajan-Dimension Generalization Bound for Polyhedral Feasible RegionsResearch Paper

Motivation

In many operational problems (shortest paths, assignment, planning) the decision solves a linear program whose cost vector is unknown at decision time and is predicted from contextual features. The predict-then-optimize pipeline fits a model fff that maps a feature vector xxx to a predicted cost vector c^=f(x)\hat c=f(x)c^=f(x), and then acts on the decision that is optimal for c^\hat cc^. Elmachtoub and Grigas (Smart "Predict, then Optimize", Management Science 2022) proposed to measure the quality of such a model not by the prediction error but by the SPO loss (Smart Predict-then-Optimize loss): the excess true cost of the decision induced by the prediction over the best decision in hindsight.

The question is whether a small SPO loss on the training sample implies a small SPO loss on new data, uniformly over the models a training procedure may return. The SPO loss is neither convex nor continuous in the prediction, so the standard Lipschitz-based bounds for regression do not apply. El Balghiti, Elmachtoub, Grigas and Tewari (arXiv:1905.11488v3, Mathematics of Operations Research 2023; a preliminary version appeared at NeurIPS 2019) give the first such generalization bounds. This mission formalizes their first one, for polyhedral feasible regions, which treats every vertex of the feasible region as a class label of a multiclass classification problem. The bound has since been used by later work, e.g. Hu, Kallus and Mao (Fast rates for contextual linear optimization, Management Science 2022), who sharpen it by a log⁡n\sqrt{\log n}logn​ factor (as noted on p. 4 of the paper).

Setting

A feasible region S⊆RdS\subseteq\mathbb R^dS⊆Rd is nonempty, compact and convex. For a cost vector c∈Rdc\in\mathbb R^dc∈Rd the nominal problem is min⁡w∈Sc⊤w\min_{w\in S}c^\top wminw∈S​c⊤w. An optimization oracle is a fixed map w∗:Rd→Sw^*:\mathbb R^d\to Sw∗:Rd→S with w∗(c)∈arg⁡min⁡w∈Sc⊤ww^*(c)\in\arg\min_{w\in S}c^\top ww∗(c)∈argminw∈S​c⊤w for every ccc; nothing is assumed about how it breaks ties. The SPO loss of a prediction c^\hat cc^ when the realized cost is ccc is

ℓSPO(c^,c)=c⊤w∗(c^)−c⊤w∗(c) ≥0.\ell_{\rm SPO}(\hat c,c)=c^\top w^*(\hat c)-c^\top w^*(c)\ \ge 0 .ℓSPO​(c^,c)=c⊤w∗(c^)−c⊤w∗(c) ≥0.

The linear optimization gap is ωS(c)=max⁡w∈Sc⊤w−min⁡w∈Sc⊤w\omega_S(c)=\max_{w\in S}c^\top w-\min_{w\in S}c^\top wωS​(c)=maxw∈S​c⊤w−minw∈S​c⊤w, and for a set C\mathcal CC of cost vectors ωS(C)=sup⁡c∈CωS(c)\omega_S(\mathcal C)=\sup_{c\in\mathcal C}\omega_S(c)ωS​(C)=supc∈C​ωS​(c); the SPO loss of a cost in C\mathcal CC lies in [0,ωS(C)][0,\omega_S(\mathcal C)][0,ωS​(C)].

Data are pairs (x,c)(x,c)(x,c) drawn from a distribution D\mathcal DD on X×C\mathcal X\times\mathcal CX×C. A hypothesis class H\mathcal HH is a family of predictors f:X→Rdf:\mathcal X\to\mathbb R^df:X→Rd. The SPO risk is RSPO(f)=ED[ℓSPO(f(x),c)]R_{\rm SPO}(f)=\mathbb E_{\mathcal D}[\ell_{\rm SPO}(f(x),c)]RSPO​(f)=ED​[ℓSPO​(f(x),c)], and on an i.i.d. sample (x1,c1),…,(xn,cn)(x_1,c_1),\dots,(x_n,c_n)(x1​,c1​),…,(xn​,cn​) the empirical SPO risk is R^SPO(f)=1n∑iℓSPO(f(xi),ci)\hat R_{\rm SPO}(f)=\frac1n\sum_i\ell_{\rm SPO}(f(x_i),c_i)R^SPO​(f)=n1​∑i​ℓSPO​(f(xi​),ci​). The empirical Rademacher complexity with respect to the SPO loss is

R^SPOn(H)=Eσ[sup⁡f∈H1n∑i=1nσi ℓSPO(f(xi),ci)]\hat{\mathfrak R}^n_{\rm SPO}(\mathcal H)=\mathbb E_\sigma\Big[\sup_{f\in\mathcal H}\frac1n\sum_{i=1}^n\sigma_i\,\ell_{\rm SPO}(f(x_i),c_i)\Big]R^SPOn​(H)=Eσ​[f∈Hsup​n1​i=1∑n​σi​ℓSPO​(f(xi​),ci​)]

with independent uniform signs σi∈{±1}\sigma_i\in\{\pm1\}σi​∈{±1}, and RSPOn(H)\mathfrak R^n_{\rm SPO}(\mathcal H)RSPOn​(H) is its expectation over the sample.

The decisions induced by H\mathcal HH form the class w∗(H)={x↦w∗(f(x)):f∈H}w^*(\mathcal H)=\{x\mapsto w^*(f(x)):f\in\mathcal H\}w∗(H)={x↦w∗(f(x)):f∈H}. A class F\mathcal FF N-shatters a finite set X⊆X\mathbb X\subseteq\mathcal XX⊆X if there are two labelings g1,g2g_1,g_2g1​,g2​ that differ at every point of X\mathbb XX such that every mixture of them (follow g1g_1g1​ on a subset TTT, g2g_2g2​ on the rest) is realized by some member of F\mathcal FF. The Natarajan dimension dN(F)d_N(\mathcal F)dN​(F) is the largest size of an N-shattered set. When SSS is a polyhedron, S\mathfrak SS denotes its finite set of extreme points.

Formalization targets

Goal: Theorem 2, second display (p. 11)

For a polyhedral SSS and every δ>0\delta>0δ>0, with probability at least 1−δ1-\delta1−δ over an i.i.d. sample of size nnn, every f∈Hf\in\mathcal Hf∈H satisfies

RSPO(f)≤R^SPO(f)+2 ωS(C)2dN(w∗(H))log⁡(n∣S∣2)n+ωS(C)log⁡(1/δ)2n.R_{\rm SPO}(f)\le\hat R_{\rm SPO}(f)+2\,\omega_S(\mathcal C)\sqrt{\frac{2d_N(w^*(\mathcal H))\log(n|\mathfrak S|^2)}{n}}+\omega_S(\mathcal C)\sqrt{\frac{\log(1/\delta)}{2n}} .RSPO​(f)≤R^SPO​(f)+2ωS​(C)n2dN​(w∗(H))log(n∣S∣2)​​+ωS​(C)2nlog(1/δ)​​.

Milestones, in attack order

  1. Theorem 1 (p. 9): with probability 1−δ1-\delta1−δ, RSPO(f)≤R^SPO(f)+2RSPOn(H)+ωS(C)log⁡(1/δ)/(2n)R_{\rm SPO}(f)\le\hat R_{\rm SPO}(f)+2\mathfrak R^n_{\rm SPO}(\mathcal H)+\omega_S(\mathcal C)\sqrt{\log(1/\delta)/(2n)}RSPO​(f)≤R^SPO​(f)+2RSPOn​(H)+ωS​(C)log(1/δ)/(2n)​ for all f∈Hf\in\mathcal Hf∈H.
  2. Massart step (Appendix B.1, p. 31): for a fixed sample with costs in C\mathcal CC, R^SPOn(H)≤ωS(C)2log⁡∣F∣X∣/n\hat{\mathfrak R}^n_{\rm SPO}(\mathcal H)\le\omega_S(\mathcal C)\sqrt{2\log|\mathfrak F_{|\mathbb X}|/n}R^SPOn​(H)≤ωS​(C)2log∣F∣X​∣/n​, where F∣X\mathfrak F_{|\mathbb X}F∣X​ is the set of decision vectors (w∗(f(x1)),…,w∗(f(xn)))(w^*(f(x_1)),\dots,w^*(f(x_n)))(w∗(f(x1​)),…,w∗(f(xn​))).
  3. Natarajan lemma (cited on p. 31; proved on the platform as UnderstandingML.natarajan_lemma): a class from an mmm-point set to kkk labels with Natarajan dimension ddd has at most mdk2dm^d k^{2d}mdk2d members.
  4. Empirical bound (Appendix B.1, p. 31): for a fixed sample, R^SPOn(H)≤ωS(C)2dN(w∗(H))log⁡(n∣S∣2)/n\hat{\mathfrak R}^n_{\rm SPO}(\mathcal H)\le\omega_S(\mathcal C)\sqrt{2d_N(w^*(\mathcal H))\log(n|\mathfrak S|^2)/n}R^SPOn​(H)≤ωS​(C)2dN​(w∗(H))log(n∣S∣2)/n​.
  5. Theorem 2, first display (p. 11): the same bound for the expected complexity RSPOn(H)\mathfrak R^n_{\rm SPO}(\mathcal H)RSPOn​(H).

Significance

The bound controls the out-of-sample decision cost of every predictor in the class, not only of an empirical risk minimizer, so it applies to any training procedure (SPO+ surrogate minimization, decision trees, heuristics) that returns a member of H\mathcal HH. Its dependence on the feasible region is only through ωS(C)\omega_S(\mathcal C)ωS​(C) and log⁡∣S∣\log|\mathfrak S|log∣S∣: the number of vertices of a combinatorial polytope is typically exponential in ddd, and enters only logarithmically. For linear predictors x↦Bxx\mapsto Bxx↦Bx the paper's Corollary 2 bounds dN(w∗(Hlin))d_N(w^*(\mathcal H_{\rm lin}))dN​(w∗(Hlin​)) by dpdpdp, giving a rate of order dplog⁡(n∣S∣)/n\sqrt{dp\log(n|\mathfrak S|)/n}dplog(n∣S∣)/n​.

The results are proved in the paper; this mission formalizes them. No statement about predict-then-optimize or the SPO loss is known to have a machine-checked proof. The platform already has the Natarajan lemma (proved) and several Massart-type lemmas for generic classes; this mission connects that multiclass machinery to decision losses, and its Theorem 1 is a reusable Rademacher generalization bound for a loss with range [0,ω][0,\omega][0,ω].

Difficulty

The obvious route through Lipschitz contraction fails: the SPO loss jumps when the prediction crosses a point where the optimum is not unique, so the Rademacher complexity of the composed class cannot be bounded by that of H\mathcal HH times a Lipschitz constant. Any argument through the finitely many vertices of SSS needs the decisions w∗(f(xi))w^*(f(x_i))w∗(f(xi​)) to take finitely many values on a sample, i.e. the oracle to return vertices; for an oracle that returns a non-vertex optimal point under ties, w∗(H)w^*(\mathcal H)w∗(H) may take infinitely many values on a sample. On the probabilistic side, the passage from the empirical to the expected complexity and the McDiarmid concentration step (Theorem 1) require the suprema over an uncountable class to be measurable, which the paper does not discuss.

Formalization scope

Lean works in Rd\mathbb R^dRd = EuclideanSpace ℝ (Fin d); cost vectors and decisions live in the same space and c⊤wc^\top wc⊤w is the inner product. The standing assumptions of §2 are hypotheses of every theorem: SSS nonempty, compact and convex; w∗w^*w∗ an arbitrary oracle (a hypothesis IsOracle S w, never a specific selection); C\mathcal CC nonempty and bounded, with the cost component of D\mathcal DD in C\mathcal CC almost surely (or, for fixed-sample statements, every ci∈Cc_i\in\mathcal Cci​∈C); n≥1n\ge1n≥1. "Polyhedron" means the solution set of finitely many linear inequalities; with compactness it is a polytope, and ∣S∣|\mathfrak S|∣S∣ is the cardinality of Set.extremePoints ℝ S. The expectation over signs is the exact average over the 2n2^n2n sign vectors; RSPOR_{\rm SPO}RSPO​ and RSPOn\mathfrak R^n_{\rm SPO}RSPOn​ are Bochner integrals. "With probability at least 1−δ1-\delta1−δ" is stated as: the product measure of the set of samples on which some f∈Hf\in\mathcal Hf∈H violates the bound is at most δ\deltaδ.

The formalization commits to the following disclosed additions:

  • In the empirical bound, Theorem 2 and its first display, the oracle returns extreme points of SSS. This is the proof's own "w.l.o.g." (p. 31), made explicit because p. 10 allows non-vertex outputs under ties. The hypothesis is needed: on the unit square, an oracle that returns distinct interior points of an edge under ties can have dN(w∗(H))=1d_N(w^*(\mathcal H))=1dN​(w∗(H))=1 and empirical complexity near 12\frac1221​, which exceeds the printed bound for large nnn.
  • The Natarajan dimension is not defined as a number (a supremum in N\mathbb NN would silently be 000 for unboundedly large shattered sets). Statements carry a natural number kkk bounding the size of every N-shattered set, in place of dN(w∗(H))d_N(w^*(\mathcal H))dN​(w∗(H)). This is equivalent when dNd_NdN​ is finite; the printed bound is vacuous otherwise.
  • Theorem 1 and the goal carry three measurability hypotheses: each loss function z↦ℓSPO(f(z1),z2)z\mapsto\ell_{\rm SPO}(f(z_1),z_2)z↦ℓSPO​(f(z1​),z2​) is measurable; the uniform deviation sup⁡f(RSPO(f)−R^SPO(f))\sup_f(R_{\rm SPO}(f)-\hat R_{\rm SPO}(f))supf​(RSPO​(f)−R^SPO​(f)) and, for each sign vector, the signed supremum sup⁡f1n∑iσiℓSPO(f(xi),ci)\sup_f\frac1n\sum_i\sigma_i\ell_{\rm SPO}(f(x_i),c_i)supf​n1​∑i​σi​ℓSPO​(f(xi​),ci​) are almost-everywhere measurable functions of the sample. Without them the integral defining RSPOn\mathfrak R^n_{\rm SPO}RSPOn​ could default to 000.
  • The Massart step assumes F∣X\mathfrak F_{|\mathbb X}F∣X​ finite, the case in which its printed right-hand side is finite.

A formalization that let dNd_NdN​ be an sSup in N\mathbb NN, or chose a specific tie-breaking oracle inside the definitions, would prove a different and in part trivial statement; both are excluded.

Infrastructure needed: McDiarmid's bounded-differences inequality and symmetrization for the product measure; Massart's finite-class lemma (a proved version is on the platform as RademacherMassart.rad_le_massart, with its own normalization); the Natarajan lemma (proved, UnderstandingML.natarajan_lemma, stated with its own but identical notion of N-shattering over finite types); finiteness and nonemptiness of the extreme points of a nonempty polytope. Theorem 1 and the Massart step do not use polyhedrality and are reusable for any bounded decision loss. Contributions of these infrastructure lemmas are welcome.

Selected references

  • O. El Balghiti, A. N. Elmachtoub, P. Grigas, A. Tewari, Generalization Bounds in the Predict-then-Optimize Framework, arXiv:1905.11488v3, 2022; Mathematics of Operations Research, 2023. https://arxiv.org/abs/1905.11488
  • A. N. Elmachtoub, P. Grigas, Smart "Predict, then Optimize", Management Science 68(1), 9–26, 2022. https://arxiv.org/abs/1710.08005
  • P. L. Bartlett, S. Mendelson, Rademacher and Gaussian complexities: risk bounds and structural results, Journal of Machine Learning Research 3, 463–482, 2002. https://www.jmlr.org/papers/v3/bartlett02a.html
  • B. K. Natarajan, On learning sets and functions, Machine Learning 4(1), 67–97, 1989. https://doi.org/10.1007/BF00114804
  • S. Shalev-Shwartz, S. Ben-David, Understanding Machine Learning: From Theory to Algorithms, Cambridge University Press, 2014 (Lemma 29.4). https://doi.org/10.1017/CBO9781107298019
  • M. Mohri, A. Rostamizadeh, A. Talwalkar, Foundations of Machine Learning, 2nd ed., MIT Press, 2018 (Theorem 3.3, Corollary 3.8). https://cs.nyu.edu/~mohri/mlbook/
9 thms4 active usersReviewed
Linear OptimizationMachine LearningProbability·Captain: mikedeng1

The Dantzig Selector: Statistical Estimation When p Is Much Larger than n 2: Oracle Inequality within a Logarithmic Factor of the Ideal Mean Squared ErrorResearch Paper

Motivation

Many regression problems have far more unknown coefficients ppp than observations nnn: gene expression studies with tens of samples and thousands of genes, imaging from few measurements, nonparametric curve recovery from a finite number of noisy samples. Estimation is hopeless in general, but becomes possible when the parameter is sparse, that is, has few nonzero entries. Candès and Tao (arXiv:math/0506081; Ann. Statist. 35(6), 2007, doi:10.1214/009053606000001523) proposed the Dantzig selector, an estimator computed by a single linear program, and showed that its squared error is within a logarithmic factor of what an oracle that knew which coefficients matter could achieve.

The estimator became one of the two standard ℓ1\ell_1ℓ1​ methods for high-dimensional regression, alongside the Lasso; the comparison of the two by Bickel, Ritov and Tsybakov (arXiv:0801.1095, 2009) is built on it. This mission targets the paper's main result, the oracle inequality (Theorem 1.2). A companion mission covers the simpler ℓ2\ell_2ℓ2​ bound for sparse parameters (Theorem 1.1).

Setting

Observations follow the linear model

y=Xβ+z,y = X\beta + z,y=Xβ+z,

where X∈Rn×pX\in\mathbb R^{n\times p}X∈Rn×p is a deterministic design matrix with columns X1,…,XpX_1,\dots,X_pX1​,…,Xp​, each of Euclidean norm ∥Xj∥ℓ2=1\|X_j\|_{\ell_2}=1∥Xj​∥ℓ2​​=1; β∈Rp\beta\in\mathbb R^pβ∈Rp is an unknown deterministic parameter; and z=(z1,…,zn)z=(z_1,\dots,z_n)z=(z1​,…,zn​) has independent N(0,σ2)N(0,\sigma^2)N(0,σ2) coordinates, σ>0\sigma>0σ>0. The vector β\betaβ is SSS-sparse if at most SSS of its entries are nonzero.

Two constants of XXX measure how close sparse sets of columns are to being orthonormal. The restricted isometry constant δS\delta_SδS​ is the smallest δ≥0\delta\ge0δ≥0 such that (1−δ)∥c∥ℓ22≤∥Xc∥ℓ22≤(1+δ)∥c∥ℓ22(1-\delta)\|c\|_{\ell_2}^2\le\|Xc\|_{\ell_2}^2\le(1+\delta)\|c\|_{\ell_2}^2(1−δ)∥c∥ℓ2​2​≤∥Xc∥ℓ2​2​≤(1+δ)∥c∥ℓ2​2​ for every ccc supported on at most SSS indices. The restricted orthogonality constant θS,S′\theta_{S,S'}θS,S′​ (defined for S+S′≤pS+S'\le pS+S′≤p) is the smallest θ≥0\theta\ge0θ≥0 with ∣⟨Xc,Xc′⟩∣≤θ∥c∥ℓ2∥c′∥ℓ2|\langle Xc,Xc'\rangle|\le\theta\|c\|_{\ell_2}\|c'\|_{\ell_2}∣⟨Xc,Xc′⟩∣≤θ∥c∥ℓ2​​∥c′∥ℓ2​​ whenever c,c′c,c'c,c′ are supported on disjoint sets of sizes at most SSS and S′S'S′. Below δ:=δ2S\delta:=\delta_{2S}δ:=δ2S​ and θ:=θS,2S\theta:=\theta_{S,2S}θ:=θS,2S​.

For a tuning level λp>0\lambda_p>0λp​>0, a Dantzig selector β^\hat\betaβ^​ is any solution of

min⁡β~∈Rp∥β~∥ℓ1subject to∥X∗(y−Xβ~)∥ℓ∞=sup⁡1≤j≤p∣⟨y−Xβ~,Xj⟩∣≤λpσ.\min_{\tilde\beta\in\mathbb R^p}\|\tilde\beta\|_{\ell_1}\quad\text{subject to}\quad\|X^*(y-X\tilde\beta)\|_{\ell_\infty}=\sup_{1\le j\le p}|\langle y-X\tilde\beta,X_j\rangle|\le\lambda_p\sigma .β~​∈Rpmin​∥β~​∥ℓ1​​subject to∥X∗(y−Xβ~​)∥ℓ∞​​=1≤j≤psup​∣⟨y−Xβ~​,Xj​⟩∣≤λp​σ.

The ideal mean squared error is ∑i=1pmin⁡(βi2,σ2)\sum_{i=1}^p\min(\beta_i^2,\sigma^2)∑i=1p​min(βi2​,σ2): the risk of an oracle that keeps exactly the coordinates above the noise level.

Formalization targets

Goal: Theorem 1.2 (pp. 8–9)

Let t>0t>0t>0, a≥0a\ge0a≥0, and λp:=(1+a+t−1)2log⁡p\lambda_p:=(\sqrt{1+a}+t^{-1})\sqrt{2\log p}λp​:=(1+a​+t−1)2logp​. If β\betaβ is SSS-sparse and δ2S+θS,2S<1−t\delta_{2S}+\theta_{S,2S}<1-tδ2S​+θS,2S​<1−t, then with probability exceeding 1−(πlog⁡p⋅pa)−11-(\sqrt{\pi\log p}\cdot p^a)^{-1}1−(πlogp​⋅pa)−1 every Dantzig selector obeys

∥β^−β∥ℓ22≤C22⋅λp2⋅(σ2+∑i=1pmin⁡(βi2,σ2)),\|\hat\beta-\beta\|_{\ell_2}^2\le C_2^2\cdot\lambda_p^2\cdot\Big(\sigma^2+\sum_{i=1}^p\min(\beta_i^2,\sigma^2)\Big),∥β^​−β∥ℓ2​2​≤C22​⋅λp2​⋅(σ2+i=1∑p​min(βi2​,σ2)),

with the explicit constant (1.14)

C2=2C01−δ−θ+2θ(1+δ)(1−δ−θ)2+1+δ1−δ−θ,C0=22(1+1−δ21−δ−θ)+(1+12)(1+δ)21−δ−θ.C_2=\frac{2C_0}{1-\delta-\theta}+\frac{2\theta(1+\delta)}{(1-\delta-\theta)^2}+\frac{1+\delta}{1-\delta-\theta},\qquad C_0=2\sqrt2\Big(1+\frac{1-\delta^2}{1-\delta-\theta}\Big)+\Big(1+\frac1{\sqrt2}\Big)\frac{(1+\delta)^2}{1-\delta-\theta}.C2​=1−δ−θ2C0​​+(1−δ−θ)22θ(1+δ)​+1−δ−θ1+δ​,C0​=22​(1+1−δ−θ1−δ2​)+(1+2​1​)1−δ−θ(1+δ)2​.

Milestones

  1. Lemma 3.2: ∥Xβ∥ℓ2≤1+δ (∥β∥ℓ2+(2S)−1/2∥β∥ℓ1)\|X\beta\|_{\ell_2}\le\sqrt{1+\delta}\,(\|\beta\|_{\ell_2}+(2S)^{-1/2}\|\beta\|_{\ell_1})∥Xβ∥ℓ2​​≤1+δ​(∥β∥ℓ2​​+(2S)−1/2∥β∥ℓ1​​) for every β\betaβ.
  2. Lemma A.1 (dual sparse reconstruction, ℓ2\ell_2ℓ2​ version): for ccc supported on ∣T∣≤2S|T|\le2S∣T∣≤2S, a vector β\betaβ on TTT whose correlations ⟨Xβ,Xj⟩\langle X\beta,X_j\rangle⟨Xβ,Xj​⟩ equal cjc_jcj​ on TTT and are small off TTT except on an exceptional set of size at most SSS, with bounds (6.1)–(6.6).
  3. Corollary A.2 (ℓ∞\ell_\inftyℓ∞​ version): the same without exceptional set, constants 1/(1−δ−θ)1/(1-\delta-\theta)1/(1−δ−θ).
  4. Corollary A.3 (constrained thresholding): an SSS-sparse β\betaβ with ∥β∥ℓ2<λS\|\beta\|_{\ell_2}<\lambda\sqrt S∥β∥ℓ2​​<λS​ splits as β′+β′′\beta'+\beta''β′+β′′ with β′\beta'β′ small in ℓ2\ell_2ℓ2​ and ℓ1\ell_1ℓ1​ and ∥X∗Xβ′′∥ℓ∞<1−δ21−δ−θλ\|X^*X\beta''\|_{\ell_\infty}<\frac{1-\delta^2}{1-\delta-\theta}\lambda∥X∗Xβ′′∥ℓ∞​​<1−δ−θ1−δ2​λ.
  5. Gaussian tail bound (Section 3, p. 15): P(sup⁡j∣⟨z,Xj⟩∣>u)≤2p φ(u)/uP(\sup_j|\langle z,X_j\rangle|>u)\le2p\,\varphi(u)/uP(supj​∣⟨z,Xj​⟩∣>u)≤2pφ(u)/u for standard Gaussian noise.
  6. Lemma 3.1: the ℓ2\ell_2ℓ2​ mass of hhh on T0T_0T0​ and its top SSS positions outside T0T_0T0​ is controlled by ∥XT01TXh∥ℓ2\|X^T_{T_{01}}Xh\|_{\ell_2}∥XT01​T​Xh∥ℓ2​​ and ∥h∥ℓ1(T0c)\|h\|_{\ell_1(T_0^c)}∥h∥ℓ1​(T0c​)​.

Significance

The result. Theorem 1.2 says that a single linear program, which knows neither the support of β\betaβ nor which coefficients exceed the noise, matches the oracle risk ∑imin⁡(βi2,σ2)\sum_i\min(\beta_i^2,\sigma^2)∑i​min(βi2​,σ2) up to a factor O(log⁡p)O(\log p)O(logp), uniformly over SSS-sparse parameters and with explicit, nonasymptotic constants. For coefficients well below the noise level it is far sharper than the σ2Slog⁡p\sigma^2 S\log pσ2Slogp bound of Theorem 1.1. It is the template for later oracle inequalities for ℓ1\ell_1ℓ1​-penalized estimators under restricted isometry or restricted eigenvalue conditions.

Formalizing it. The theorem is proved in the paper, but parts of the argument are only sketched: Corollary A.2 refers to the 2005 Decoding by Linear Programming paper for its convergence argument, and Corollary A.3's ℓ1\ell_1ℓ1​ bound is printed with a constant its own proof does not deliver. A machine-checked proof settles these steps. The restricted isometry and orthogonality constants used here are already published on the platform from the decoding series; the appendix lemmas on dual vectors are reusable for any compressed-sensing result in that framework. No formalization of the Dantzig selector's oracle inequality is known to us.

Difficulty

The natural proof compares β^\hat\betaβ^​ with the hard-thresholded parameter β(1)\beta^{(1)}β(1) that keeps only the large coefficients: if β(1)\beta^{(1)}β(1) were feasible for the Dantzig constraint, the analysis of Theorem 1.1 would apply directly. It is not feasible in general, because the small coefficients β(2)\beta^{(2)}β(2), though individually below the noise level, can add up to a large correlation X∗Xβ(2)X^*X\beta^{(2)}X∗Xβ(2). The central difficulty is to split β(2)\beta^{(2)}β(2) into a part with controlled ℓ1\ell_1ℓ1​ and ℓ2\ell_2ℓ2​ norm and a part invisible to the constraint; this requires constructing dual vectors with prescribed correlations (Lemma A.1, Corollary A.2), via an iterative, geometrically convergent correction. The probabilistic part is a Gaussian tail estimate plus a union bound, and the bookkeeping of constants must be carried through exactly.

Formalization scope

Vectors are functions Fin p → ℝ, the design is Matrix (Fin n) (Fin p) ℝ, and the noise is a family z : Fin n → Ω → ℝ of mutually independent random variables (iIndepFun) each with law gaussianReal 0 σ². δ2S\delta_{2S}δ2S​ and θS,2S\theta_{S,2S}θS,2S​ are the published CandesTao.Decoding.restrictedIsometryConst X (2*S) and restrictedOrthogonalityConst X S (2*S) (infima, absolute value in the orthogonality condition). Domain: S≥1S\ge1S≥1 and 3S≤p3S\le p3S≤p (the paper defines θS,S′\theta_{S,S'}θS,S′​ for S+S′≤pS+S'\le pS+S′≤p), which forces p≥3p\ge3p≥3 and log⁡p>0\log p>0logp>0. A Dantzig selector is any ℓ1\ell_1ℓ1​ minimizer over the feasible set; the ℓ∞\ell_\inftyℓ∞​ constraint is a bound on every coordinate.

The goal bounds from above the (outer) probability of the bad event "no Dantzig selector exists, or some Dantzig selector violates (1.13)". Because the event includes non-existence, a definition no vector satisfies cannot make the theorem vacuous; and the constant C2C_2C2​ is the printed (1.14), evaluated at δ2S\delta_{2S}δ2S​, θS,2S\theta_{S,2S}θS,2S​ of XXX, not a free constant chosen after the fact.

Corrected constant: Corollary A.3 is stated with ∥β′∥ℓ1≤21+δ1−δ−θ∥β∥ℓ22/λ\|\beta'\|_{\ell_1}\le2\frac{1+\delta}{1-\delta-\theta}\|\beta\|_{\ell_2}^2/\lambda∥β′∥ℓ1​​≤21−δ−θ1+δ​∥β∥ℓ2​2​/λ, the bound its proof gives once Corollary A.2 is applied at an integer sparsity level; the printed statement omits the factor 222. Corollary A.2 carries Lemma A.1's standing hypothesis δ+θ<1\delta+\theta<1δ+θ<1. The deterministic lemmas (3.1, 3.2, A.1–A.3) assume nothing about column norms, since their statements do not need it.

Useful infrastructure: monotonicity of δS\delta_SδS​ and θS,S′\theta_{S,S'}θS,S′​ in their indices (the proof applies the lemmas at a smaller sparsity level), the Gaussian tail bound 1−Φ(u)<φ(u)/u1-\Phi(u)<\varphi(u)/u1−Φ(u)<φ(u)/u, existence of minimizers of the Dantzig linear program, and a sorting/blocking toolkit for "the SSS largest positions". Proofs of individual milestones are welcome independently.

Selected references

  • E. Candès and T. Tao, The Dantzig selector: statistical estimation when p is much larger than n, Ann. Statist. 35(6):2313–2351, 2007. arXiv:math/0506081v3, doi:10.1214/009053606000001523
  • E. Candès and T. Tao, Decoding by linear programming, IEEE Trans. Inform. Theory 51(12):4203–4215, 2005. arXiv:math/0502327, doi:10.1109/TIT.2005.858979
  • P. Bickel, Y. Ritov and A. Tsybakov, Simultaneous analysis of Lasso and Dantzig selector, Ann. Statist. 37(4):1705–1732, 2009. arXiv:0801.1095, doi:10.1214/08-AOS620
  • D. Donoho and I. Johnstone, Ideal spatial adaptation by wavelet shrinkage, Biometrika 81(3):425–455, 1994. doi:10.1093/biomet/81.3.425
11 thms3 active usersReviewed
Information TheoryQuantum Information·Captain: mikedeng1

Shadow Tomography of Quantum States 2: Even the Classical Special Case Needs Ω(min{D, log M}/ε²) CopiesResearch Paper

Why the copy count matters

Shadow tomography asks for predictions of many specified measurements of an unknown quantum state, while using as few prepared copies of that state as possible. The requested output is a list of acceptance probabilities, not a full description of the state. A procedure might exploit the fact that these are only MMM numbers, even when the state has dimension DDD. The natural question is how far this saving can go. Aaronson's paper gives upper bounds and separates two sources of difficulty: one already present for ordinary probability distributions, and one arising from noncommuting quantum measurements.

This mission concerns the first source. It formalizes Theorem 16, which says that even when the state and every requested measurement are diagonal in the same basis, the number of copies must grow with min⁡{D,log⁡M}/ε2\min\{D,\log M\}/\varepsilon^2min{D,logM}/ε2 in the relevant asymptotic regime. In that special case, the unknown state is an ordinary distribution on DDD outcomes. The theorem therefore puts a limit on any proposed improvement to shadow tomography that would promise fewer copies in all instances.

States, measurements, and estimates

A mixed state ρ\rhoρ on a DDD-dimensional system is a positive semidefinite D×DD\times DD×D complex matrix with trace one. A two-outcome measurement is represented by an effect EEE satisfying 0⪯E⪯I0\preceq E\preceq I0⪯E⪯I; it accepts ρ\rhoρ with probability Tr⁡(Eρ)\operatorname{Tr}(E\rho)Tr(Eρ). When ρ\rhoρ is diagonal, its diagonal entries are the probabilities of the DDD basis outcomes. When EEE is diagonal too, it specifies a randomized yes-or-no test on those outcomes. These conventions are stated in Section 3 of the paper.

Given effects E1,…,EME_1,\ldots,E_ME1​,…,EM​, a shadow-tomography strategy measures kkk independent copies ρ⊗k\rho^{\otimes k}ρ⊗k and outputs estimates b1,…,bMb_1,\ldots,b_Mb1​,…,bM​. It succeeds on ρ\rhoρ when every estimate differs from Tr⁡(Eiρ)\operatorname{Tr}(E_i\rho)Tr(Ei​ρ) by at most ε\varepsilonε. The lower bound requires success probability at least 2/32/32/3 for every diagonal mixed state. The strategy may make a joint quantum measurement on all copies and may choose its estimates from its observed outcome. This is the same measurement model used in Problem 1.

Formalization targets

Classical special-case lower bound

The goal is the classical clause of Theorem 16. There are absolute constants c>0c>0c>0 and N0N_0N0​ such that, for D≥N0D\ge N_0D≥N0​, log⁡2M≥N0\log_2 M\ge N_0log2​M≥N0​, and 0<ε≤1/60<\varepsilon\le1/60<ε≤1/6, there are MMM diagonal effects with 0/1 entries, corresponding to the known Boolean functions in Section 6.1, for which every strategy successful on all diagonal states must use

k≥c min⁡{D,log⁡2M}ε2.k\ge c\,\frac{\min\{D,\log_2 M\}}{\varepsilon^2}.k≥cε2min{D,log2​M}​.

The hard measurements are chosen before the strategy is quantified. The statement therefore also rules out a strategy with a smaller copy count that works uniformly for all quantum states and measurements. Its constants and threshold express the Ω\OmegaΩ notation in Theorem 16, rather than specifying a numerical optimum.

Supporting targets

Four milestones come from the proof on pages 20–21: the high-probability overlap bound for independently chosen half-size subsets (Eq. (1)); the acceptance probability of a subset under its associated biased distribution; an upper bound on the mutual information between the hidden subset index and the observed samples; and the exact entropy formula with a quadratic entropy deficit. These statements expose the combinatorial and information-theoretic parts of the lower bound while leaving the goal as the paper's copy-complexity result.

What the result supplies

Theorem 16 sets a floor for shadow tomography that survives even when all operators commute. Any uniform copy bound for the full quantum task must respect this floor. The result also distinguishes the difficulty of predicting many properties of a distribution from the extra difficulty possible for noncommuting states and measurements, which the paper treats in a separate lower bound. Section 6 presents both bounds.

A complete formalization would give machine-checked statements and proofs for the finite subset construction, the entropy calculation, the information inequality, and the reduction from a successful quantum measurement procedure on diagonal states to a lower bound on kkk. The theorem is proved on paper; these draft statements are open Lean goals and do not claim that its proof has been machine checked. The finite-distribution and information-theory infrastructure is reusable for other lower bounds based on hidden-index families.

Where the argument is delicate

Counting how many possible measurements there are does not by itself show that samples reveal enough about which distribution generated them. The lower bound needs a quantitative relation between estimation accuracy and information about a hidden index, while each individual sample carries limited information. The paper's printed overlap condition (1) is too weak for the next displayed ε/2\varepsilon/2ε/2 estimate: at its boundary it gives ε\varepsilonε. The milestone preserves Eq. (1) as printed; closing the goal requires the correspondingly sharper overlap fact with N/24N/24N/24, which follows from the same type of concentration statement after adjusting its constant. The printed assertion that learning the index requires mutual information at least log⁡2K\log_2 Klog2​K is also imprecise at success probability 2/32/32/3; a quantitative decoding inequality is needed. Neither incorrect display is a draft milestone.

Formalization scope

Matrices are indexed by Fin D. WildeQIT.IsDensityOperator supplies the mixed-state predicate ρ⪰0\rho\succeq0ρ⪰0 and Tr⁡(ρ)=1\operatorname{Tr}(\rho)=1Tr(ρ)=1. Diagonal states and effects use the standard matrix diagonal predicate. An effect is positive semidefinite together with its complement. The tensor power uses functions Fin k → Fin D as basis indices; at k=0k=0k=0 it is a one-by-one identity matrix. A strategy is a finite-outcome POVM on that tensor power, followed by a real estimate vector for each outcome. The output values are not restricted to [0,1][0,1][0,1]: clipping them to this interval cannot worsen an estimate of a probability. No restriction to classical estimators is placed in the goal; that would change the allowed strategies before the theorem has been proved.

The asymptotic threshold excludes the one-dimensional and single-measurement corners where the claimed rate does not describe the problem. The bound ε≤1/6\varepsilon\le1/6ε≤1/6 keeps the biased distributions used on page 20 nonnegative; ε≥1/2\varepsilon\ge1/2ε≥1/2 would permit a zero-copy constant estimate. Subset milestones require even NNN or explicitly require a half-size subset, so N/2N/2N/2 has its intended meaning. The natural logarithm appears nowhere in the lower-bound rate; entropy, mutual information, and log⁡2M\log_2 Mlog2​M use base two. At zero probability the entropy convention is 0log⁡0=00\log 0=00log0=0.

The finite distributions, entropy, conditional entropy, and mutual information reuse the published WildeQIT definitions. Contributions that prove the four source milestones, establish the sharper overlap fact, or supply the quantitative decoding step are welcome. The goal must retain its order of quantifiers: one hard measurement family, then every strategy, with success demanded on every diagonal state.

Selected references

  • Scott Aaronson, Shadow Tomography of Quantum States, arXiv preprint arXiv:1711.01053v2, 2018. Preprint.
14 thms3 active usersReviewed
Linear OptimizationMachine LearningProbability·Captain: mikedeng1

The Dantzig Selector: Statistical Estimation When p Is Much Larger than n 1: ℓ2 Error Bound for Sparse Parameters under the Uniform Uncertainty PrincipleResearch Paper

Motivation

In many statistical applications the number of unknown parameters ppp is far larger than the number of observations nnn: gene-expression studies with tens of samples and thousands of genes, imaging problems with fewer measurements than pixels, and nonparametric curve estimation from finitely many noisy samples. Least squares is useless in this regime, since the system Xβ=yX\beta=yXβ=y is underdetermined. If the parameter is sparse (only a few of its entries are nonzero), estimation becomes possible, and the question is how accurate a computationally tractable estimator can be.

Candès and Tao (arXiv:math/0506081; Ann. Statist. 35(6), 2007, doi:10.1214/009053606000001523) introduced the Dantzig selector, an estimator computed by a linear program, and proved that its squared error is within a factor of order log⁡p\log plogp of the error of an oracle that knows where the nonzero entries are. The paper, with its discussion in the same issue, is one of the founding results of high-dimensional sparse regression, alongside the Lasso analysis of Bickel, Ritov and Tsybakov (arXiv:0801.1095).

Timeline. Candès and Tao (2005, arXiv:math/0502327) showed that ℓ1\ell_1ℓ1​ minimization recovers a sparse vector exactly from noiseless data when the restricted isometry constants of the design satisfy δS+θS,S+θS,2S<1\delta_S+\theta_{S,S}+\theta_{S,2S}<1δS​+θS,S​+θS,2S​<1. The Dantzig selector paper (first posted 2005, published 2007) carried this to Gaussian noise, with the ℓ2\ell_2ℓ2​ error bound formalized here (Theorem 1.1) and an oracle inequality (Theorem 1.2). Bickel, Ritov and Tsybakov (2009) replaced the restricted isometry hypothesis by weaker restricted eigenvalue conditions and showed that the Lasso and the Dantzig selector behave alike.

Setting

Observe y∈Rny\in\mathbb R^ny∈Rn from the linear model

y=Xβ+z,y=X\beta+z ,y=Xβ+z,

where X∈Rn×pX\in\mathbb R^{n\times p}X∈Rn×p is a deterministic design matrix with columns X1,…,XpX_1,\dots,X_pX1​,…,Xp​, each of Euclidean norm ∥Xj∥ℓ2=1\|X_j\|_{\ell_2}=1∥Xj​∥ℓ2​​=1; β∈Rp\beta\in\mathbb R^pβ∈Rp is an unknown deterministic parameter; and z=(z1,…,zn)z=(z_1,\dots,z_n)z=(z1​,…,zn​) is a vector of independent N(0,σ2)N(0,\sigma^2)N(0,σ2) random variables with σ>0\sigma>0σ>0. The vector β\betaβ is SSS-sparse if at most SSS of its entries are nonzero.

For T⊆{1,…,p}T\subseteq\{1,\dots,p\}T⊆{1,…,p} let XTX_TXT​ be the submatrix of the columns indexed by TTT. The restricted isometry constant δS\delta_SδS​ is the smallest δ≥0\delta\ge0δ≥0 with

(1−δ)∥c∥ℓ22≤∥XTc∥ℓ22≤(1+δ)∥c∥ℓ22(1-\delta)\|c\|_{\ell_2}^2\le\|X_Tc\|_{\ell_2}^2\le(1+\delta)\|c\|_{\ell_2}^2(1−δ)∥c∥ℓ2​2​≤∥XT​c∥ℓ2​2​≤(1+δ)∥c∥ℓ2​2​

for all ∣T∣≤S|T|\le S∣T∣≤S and all coefficient vectors ccc; the restricted orthogonality constant θS,S′\theta_{S,S'}θS,S′​ (for S+S′≤pS+S'\le pS+S′≤p) is the smallest θ≥0\theta\ge0θ≥0 with ∣⟨XTc,XT′c′⟩∣≤θ∥c∥ℓ2∥c′∥ℓ2|\langle X_Tc,X_{T'}c'\rangle|\le\theta\|c\|_{\ell_2}\|c'\|_{\ell_2}∣⟨XT​c,XT′​c′⟩∣≤θ∥c∥ℓ2​​∥c′∥ℓ2​​ for all disjoint T,T′T,T'T,T′ with ∣T∣≤S|T|\le S∣T∣≤S, ∣T′∣≤S′|T'|\le S'∣T′∣≤S′.

Given a tuning parameter λp>0\lambda_p>0λp​>0, the Dantzig selector β^\hat\betaβ^​ is any solution of

min⁡β~∈Rp∥β~∥ℓ1subject to∥X∗(y−Xβ~)∥ℓ∞=max⁡1≤j≤p∣⟨y−Xβ~,Xj⟩∣≤λp⋅σ.\min_{\tilde\beta\in\mathbb R^p}\|\tilde\beta\|_{\ell_1}\quad\text{subject to}\quad\|X^*(y-X\tilde\beta)\|_{\ell_\infty}=\max_{1\le j\le p}|\langle y-X\tilde\beta,X_j\rangle|\le\lambda_p\cdot\sigma .β~​∈Rpmin​∥β~​∥ℓ1​​subject to∥X∗(y−Xβ~​)∥ℓ∞​​=1≤j≤pmax​∣⟨y−Xβ~​,Xj​⟩∣≤λp​⋅σ.

Formalization targets

Goal: Theorem 1.1

Let S≥1S\ge1S≥1, 3S≤p3S\le p3S≤p, β\betaβ SSS-sparse, and δ2S+θS,2S<1\delta_{2S}+\theta_{S,2S}<1δ2S​+θS,2S​<1. For every a≥0a\ge0a≥0, with λp=2(1+a)log⁡p\lambda_p=\sqrt{2(1+a)\log p}λp​=2(1+a)logp​, with probability exceeding 1−(πlog⁡p⋅pa)−11-(\sqrt{\pi\log p}\cdot p^a)^{-1}1−(πlogp​⋅pa)−1 the program has a solution and every solution satisfies

∥β^−β∥ℓ22≤C12⋅λp2⋅S⋅σ2,C1=41−δ2S−θS,2S.\|\hat\beta-\beta\|_{\ell_2}^2\le C_1^2\cdot\lambda_p^2\cdot S\cdot\sigma^2,\qquad C_1=\frac{4}{1-\delta_{2S}-\theta_{S,2S}} .∥β^​−β∥ℓ2​2​≤C12​⋅λp2​⋅S⋅σ2,C1​=1−δ2S​−θS,2S​4​.

For a=0a=0a=0 this is ∥β^−β∥ℓ22≤C12⋅(2log⁡p)⋅S⋅σ2\|\hat\beta-\beta\|_{\ell_2}^2\le C_1^2\cdot(2\log p)\cdot S\cdot\sigma^2∥β^​−β∥ℓ2​2​≤C12​⋅(2logp)⋅S⋅σ2, display (1.10) of the paper. The constant is the one the paper's proof establishes (see Formalization scope).

Milestones

  1. The cone constraint (3.2): if ∥β+h∥ℓ1≤∥β∥ℓ1\|\beta+h\|_{\ell_1}\le\|\beta\|_{\ell_1}∥β+h∥ℓ1​​≤∥β∥ℓ1​​ and β\betaβ vanishes off T0T_0T0​, then ∥hT0c∥ℓ1≤∥hT0∥ℓ1\|h_{T_0^c}\|_{\ell_1}\le\|h_{T_0}\|_{\ell_1}∥hT0c​​∥ℓ1​​≤∥hT0​​∥ℓ1​​.
  2. The tube constraint (3.3): with unit-normed columns, if ∣⟨z,Xj⟩∣≤λp|\langle z,X_j\rangle|\le\lambda_p∣⟨z,Xj​⟩∣≤λp​ for all jjj and β^\hat\betaβ^​ is feasible, then ∥X∗X(β^−β)∥ℓ∞≤2λp\|X^*X(\hat\beta-\beta)\|_{\ell_\infty}\le2\lambda_p∥X∗X(β^​−β)∥ℓ∞​​≤2λp​.
  3. Lemma 3.1 (under the section’s unit-column assumption): an ℓ2\ell_2ℓ2​ bound on hhh over T0∪T1T_0\cup T_1T0​∪T1​ (T1T_1T1​ the SSS largest entries of hhh off T0T_0T0​) in terms of ∥XT01TXh∥ℓ2\|X_{T_{01}}^TXh\|_{\ell_2}∥XT01​T​Xh∥ℓ2​​ and ∥h∥ℓ1(T0c)\|h\|_{\ell_1(T_0^c)}∥h∥ℓ1​(T0c​)​, and ∥h∥ℓ22≤∥h∥ℓ2(T01)2+S−1∥h∥ℓ1(T0c)2\|h\|_{\ell_2}^2\le\|h\|_{\ell_2(T_{01})}^2+S^{-1}\|h\|_{\ell_1(T_0^c)}^2∥h∥ℓ2​2​≤∥h∥ℓ2​(T01​)2​+S−1∥h∥ℓ1​(T0c​)2​.
  4. The deterministic core: with σ=1\sigma=1σ=1, on the event ∣⟨z,Xj⟩∣≤λp|\langle z,X_j\rangle|\le\lambda_p∣⟨z,Xj​⟩∣≤λp​ for all jjj, every Dantzig selector satisfies ∥β^−β∥ℓ22≤C12λp2S\|\hat\beta-\beta\|_{\ell_2}^2\le C_1^2\lambda_p^2S∥β^​−β∥ℓ2​2​≤C12​λp2​S.
  5. The Gaussian tail bound: for standard normal zzz and Zj=⟨z,Xj⟩Z_j=\langle z,X_j\rangleZj​=⟨z,Xj​⟩, P(sup⁡j∣Zj∣>u)≤2p φ(u)/u\mathbb P(\sup_j|Z_j|>u)\le2p\,\varphi(u)/uP(supj​∣Zj​∣>u)≤2pφ(u)/u with φ(u)=(2π)−1/2e−u2/2\varphi(u)=(2\pi)^{-1/2}e^{-u^2/2}φ(u)=(2π)−1/2e−u2/2.

Significance

The result. Theorem 1.1 shows that an estimator computable by linear programming reaches, up to the factor 2log⁡p2\log p2logp and the constant C12C_1^2C12​, the squared error Sσ2S\sigma^2Sσ2 that least squares would attain if the support of β\betaβ were known in advance, even when p≫np\gg np≫n. The factor log⁡p\log plogp is the price of not knowing the support; the paper argues (p. 5) that, apart from this factor, (1.10) is unimprovable in general. The bound is non-asymptotic, with an explicit constant and an explicit failure probability, and it holds for every SSS-sparse β\betaβ simultaneously in the sense that the good event (the noise being nearly orthogonal to every column) does not depend on β\betaβ. Its deterministic part, Lemma 3.1, is reused verbatim in the proof of the paper's oracle inequality (Theorem 1.2) and became a standard tool in compressed sensing.

Formalizing it. The result is proved, and to our knowledge no machine-checked proof exists. A formalization produces a checked version of the cone-and-tube argument behind most ℓ1\ell_1ℓ1​-recovery guarantees, a Lean statement of the restricted isometry machinery for noisy data, and a checked Gaussian maximal inequality usable for other high-dimensional estimators. It also settles the exact constant: the paper prints C1=4/(1−δS−θS,2S)C_1=4/(1-\delta_S-\theta_{S,2S})C1​=4/(1−δS​−θS,2S​), while its proof gives δ2S\delta_{2S}δ2S​ in place of δS\delta_SδS​.

Difficulty

Lemma 3.1 is the main obstacle. The obvious approach bounds ∥h∥ℓ2\|h\|_{\ell_2}∥h∥ℓ2​​ directly through restricted isometry, and it fails because the error hhh is not sparse: it spreads over all ppp coordinates, and restricted isometry controls XXX only on vectors with at most 2S2S2S nonzero entries. The two constraints (3.2) and (3.3) only say that hhh is concentrated in ℓ1\ell_1ℓ1​ on the SSS coordinates of T0T_0T0​ and that X∗XhX^*XhX∗Xh is small coordinatewise, and turning that into an ℓ2\ell_2ℓ2​ bound on all of hhh is where the work lies. In Lean this requires bookkeeping that is routine on paper: ordering the coordinates of hhh off T0T_0T0​ by magnitude, with ties and a possibly incomplete last group of coordinates, and working with the span of a selected set of columns. On the probabilistic side, the tail bound needs the law of ⟨z,Xj⟩\langle z,X_j\rangle⟨z,Xj​⟩ (a weighted sum of independent Gaussians), a sharp Gaussian tail estimate of Mills-ratio type, and a union over ppp events. A cruder sub-Gaussian bound 2e−u2/22e^{-u^2/2}2e−u2/2 would not give the stated failure probability.

Formalization scope

Indices are Fin n and Fin p; vectors are functions into ℝ. The norms, the column XjX_jXj​ and the constants δS\delta_SδS​, θS,S′\theta_{S,S'}θS,S′​ are the published definitions CandesTao_Decoding_Norms and CandesTao_Decoding_RestrictedIsometry (the smallest admissible constants, via sInf), from the formalization of Candès and Tao's Decoding by Linear Programming. The noise is a family z : Fin n → Ω → ℝ on a probability space, mutually independent (iIndepFun), each coordinate with law gaussianReal 0 σ². The ℓ∞\ell_\inftyℓ∞​ constraint is coordinatewise. A Dantzig selector is any minimizer; uniqueness is not assumed. Section 3 works with σ=1\sigma=1σ=1; the goal is stated for general σ>0\sigma>0σ>0.

Committed conventions and corrections:

  • Corrected constant. Theorem 1.1 is printed with C1=4/(1−δS−θS,2S)C_1=4/(1-\delta_S-\theta_{S,2S})C1​=4/(1−δS​−θS,2S​), but the proof (pp. 18–19) applies Lemma 3.1, whose δ\deltaδ is δ2S\delta_{2S}δ2S​. Since δS≤δ2S\delta_S\le\delta_{2S}δS​≤δ2S​, the printed constant is stronger than what is proved. The goal and the deterministic core are stated with C1=4/(1−δ2S−θS,2S)C_1=4/(1-\delta_{2S}-\theta_{S,2S})C1​=4/(1−δ2S​−θS,2S​).
  • Domain. 1≤S1\le S1≤S and 3S≤p3S\le p3S≤p, because θS,2S\theta_{S,2S}θS,2S​ is defined only for S+2S≤pS+2S\le pS+2S≤p. This forces p≥3p\ge3p≥3 and log⁡p>0\log p>0logp>0.
  • Failure event. The probability bounded is that of the set where no Dantzig selector exists or some Dantzig selector violates the bound. A version that only constrains existing solutions, or that assumes the feasible set is nonempty, would be weaker. The bound is strict, as in the paper's "exceeding", and is on the outer measure, so no measurability of the event is assumed.
  • Standing assumptions are binders: unit-normed columns, independent Gaussian noise, deterministic XXX and β\betaβ.

A trivializing formalization is excluded: the hypothesis δ2S+θS,2S<1\delta_{2S}+\theta_{S,2S}<1δ2S​+θS,2S​<1 is on the actual least constants of XXX, not on free parameters, and it is satisfiable (for instance by X=IpX=I_pX=Ip​, where both constants vanish).

Needed infrastructure: sums of independent real Gaussians (Mathlib has gaussianReal and its convolution), a Mills-ratio tail bound, a sorting-based block decomposition of a Finset, and orthogonal projection onto the span of finitely many columns. The block decomposition and the tail bound are reusable beyond this mission. Proofs of any milestone are welcome, as are alternative proofs of Lemma 3.1.

Selected references

  • E. Candès and T. Tao, The Dantzig selector: statistical estimation when p is much larger than n, Ann. Statist. 35(6) (2007), 2313–2351. arXiv:math/0506081, doi:10.1214/009053606000001523
  • E. Candès and T. Tao, Decoding by linear programming, IEEE Trans. Inform. Theory 51(12) (2005), 4203–4215. arXiv:math/0502327
  • P. Bickel, Y. Ritov and A. Tsybakov, Simultaneous analysis of Lasso and Dantzig selector, Ann. Statist. 37(4) (2009), 1705–1732. arXiv:0801.1095
9 thms3 active usersReviewed
Machine LearningProbability·Captain: mikedeng1

On the Uniform Convergence of Relative Frequencies of Events to Their Probabilities III: Uniform Convergence and the Entropy per ObservationResearch Paper

Motivation

Estimating a probability by the relative frequency of the event in an independent sample is justified for one event by the law of large numbers. Statistics and learning theory need more: the frequencies of a whole class of events SSS must approach their probabilities simultaneously, so that a quantity chosen after looking at the data (the empirical risk minimizer, the empirical distribution function) is still close to its expectation. Glivenko's theorem on the empirical distribution function is the classical instance; empirical risk minimization rests on the same property for the class of loss sets of a model.

Vapnik and Chervonenkis, On the Uniform Convergence of Relative Frequencies of Events to Their Probabilities, Theory Probab. Appl. 16 (1971), treat this question in two parts. The first gives a distribution-free sufficient condition through the growth function (Theorems 1–3). The second, which this mission formalizes, gives a condition that is necessary and sufficient for a fixed distribution: Theorem 4, the entropy criterion.

Timeline. 1933: Glivenko and Cantelli prove uniform convergence for the class of rays {x≤a}\{x \le a\}{x≤a} on the line. 1968: Vapnik and Chervonenkis announce the results in Dokl. Akad. Nauk SSSR 181. 1971: the full paper appears, with the growth-function bound and the entropy criterion. Later work (Talagrand 1987; Dudley, Giné and Zinn 1991) recasts such criteria as the theory of Glivenko–Cantelli classes.

Setting

Let XXX be a set carrying a probability measure PPP, and SSS a collection of measurable subsets of XXX (events). A sample of size lll is a sequence x1,…,xlx_1, \dots, x_lx1​,…,xl​ of independent draws from PPP; repetitions are allowed. For A∈SA \in SA∈S the relative frequency νA(l)\nu_A^{(l)}νA(l)​ is the fraction of sample terms lying in AAA, and PA=P(A)P_A = P(A)PA​=P(A). The maximal deviation is

π(l)(x1,…,xl)=sup⁡A∈S∣νA(l)−PA∣.\pi^{(l)}(x_1, \dots, x_l) = \sup_{A \in S} \bigl|\nu_A^{(l)} - P_A\bigr| .π(l)(x1​,…,xl​)=A∈Ssup​​νA(l)​−PA​​.

The relative frequencies converge in probability to the probabilities uniformly over SSS when P{π(l)>ε}→0\mathbf{P}\{\pi^{(l)} > \varepsilon\} \to 0P{π(l)>ε}→0 as l→∞l \to \inftyl→∞ for every ε>0\varepsilon > 0ε>0.

Each A∈SA \in SA∈S induces in a sample the subsample of terms lying in AAA. The index ΔS(x1,…,xl)\Delta^S(x_1, \dots, x_l)ΔS(x1​,…,xl​) is the number of different subsamples induced by the sets of SSS; it lies between 000 and 2l2^l2l. The entropy of SSS in samples of size lll is

HS(l)=Elog⁡2ΔS(x1,…,xl).H^S(l) = \mathbf{E} \log_2 \Delta^S(x_1, \dots, x_l) .HS(l)=Elog2​ΔS(x1​,…,xl​).

For a sample of size 2l2l2l, split into halves x1,…,xlx_1, \dots, x_lx1​,…,xl​ and xl+1,…,x2lx_{l+1}, \dots, x_{2l}xl+1​,…,x2l​ with relative frequencies νA′\nu'_AνA′​ and νA′′\nu''_AνA′′​, the semi-sample deviation is ρ(l)=sup⁡A∈S∣νA′−νA′′∣\rho^{(l)} = \sup_{A \in S} |\nu'_A - \nu''_A|ρ(l)=supA∈S​∣νA′​−νA′′​∣. Finally Φ(n,r)\Phi(n, r)Φ(n,r) is defined by the recurrence Φ(n,r)=Φ(n,r−1)+Φ(n−1,r−1)\Phi(n, r) = \Phi(n, r-1) + \Phi(n-1, r-1)Φ(n,r)=Φ(n,r−1)+Φ(n−1,r−1), Φ(0,r)=Φ(n,0)=1\Phi(0, r) = \Phi(n, 0) = 1Φ(0,r)=Φ(n,0)=1.

Formalization targets

Goal: Theorem 4 (p. 275)

(∀ε>0: lim⁡l→∞P{π(l)>ε}=0)  ⟺  lim⁡l→∞HS(l)l=0.\Bigl(\forall \varepsilon > 0:\ \lim_{l\to\infty} \mathbf{P}\{\pi^{(l)} > \varepsilon\} = 0\Bigr) \iff \lim_{l \to \infty} \frac{H^S(l)}{l} = 0 .(∀ε>0: l→∞lim​P{π(l)>ε}=0)⟺l→∞lim​lHS(l)​=0.

Milestones

  1. Entropy rate. (12) ΔS(x1,…,xl)≤ΔS(x1,…,xk)ΔS(xk+1,…,xl)\Delta^S(x_1, \dots, x_l) \le \Delta^S(x_1, \dots, x_k)\Delta^S(x_{k+1}, \dots, x_l)ΔS(x1​,…,xl​)≤ΔS(x1​,…,xk​)ΔS(xk+1​,…,xl​); the subadditivity HS(l1+l2)≤HS(l1)+HS(l2)H^S(l_1 + l_2) \le H^S(l_1) + H^S(l_2)HS(l1​+l2​)≤HS(l1​)+HS(l2​); Lemma 3, HS(l)/l→c∈[0,1]H^S(l)/l \to c \in [0, 1]HS(l)/l→c∈[0,1]; Lemma 4, P(∣l−1log⁡2ΔS−c∣>ε)→0\mathbf{P}(|l^{-1}\log_2 \Delta^S - c| > \varepsilon) \to 0P(∣l−1log2​ΔS−c∣>ε)→0.
  2. Sufficiency. Lemma 2, P{π(l)>ε}≤2 P{ρ(l)≥ε/2}\mathbf{P}\{\pi^{(l)} > \varepsilon\} \le 2\,\mathbf{P}\{\rho^{(l)} \ge \varepsilon/2\}P{π(l)>ε}≤2P{ρ(l)≥ε/2} for l≥2/ε2l \ge 2/\varepsilon^2l≥2/ε2; the per-sample permutation bound 2ΔS(x1,…,x2l)e−ε2l/82\Delta^S(x_1, \dots, x_{2l}) e^{-\varepsilon^2 l/8}2ΔS(x1​,…,x2l​)e−ε2l/8; and
P{ρ(l)≥ε2}≤2(2e)ε2l/8+P{12llog⁡2ΔS(x1,…,x2l)>ε216}.\mathbf{P}\{\rho^{(l)} \ge \tfrac{\varepsilon}{2}\} \le 2\Bigl(\frac{2}{e}\Bigr)^{\varepsilon^2 l/8} + \mathbf{P}\Bigl\{\tfrac{1}{2l}\log_2 \Delta^S(x_1, \dots, x_{2l}) > \tfrac{\varepsilon^2}{16}\Bigr\} .P{ρ(l)≥2ε​}≤2(e2​)ε2l/8+P{2l1​log2​ΔS(x1​,…,x2l​)>16ε2​}.
  1. Necessity. Lemma 1 (Sauer–Shelah in sequence form); step 1°, 1−P(C′)≥(1−P(Q))21 - \mathbf{P}(C') \ge (1 - \mathbf{P}(Q))^21−P(C′)≥(1−P(Q))2 with C′={ρ(l)>2ε}C' = \{\rho^{(l)} > 2\varepsilon\}C′={ρ(l)>2ε}; (26), P{ΔS>Φ([ql],l)}→1\mathbf{P}\{\Delta^S > \Phi([ql], l)\} \to 1P{ΔS>Φ([ql],l)}→1 when 0<q<140 < q < \frac140<q<41​ and qlog⁡2(2e/q)<cq\log_2(2e/q) < cqlog2​(2e/q)<c; and (29), P{π(l)>ε}→1\mathbf{P}\{\pi^{(l)} > \varepsilon\} \to 1P{π(l)>ε}→1 when moreover 0<ε<q/70 < \varepsilon < q/70<ε<q/7.

Significance

Theorem 4 characterizes uniform convergence for a given distribution exactly, with no gap between the necessary and the sufficient condition. It separates the cases the growth-function bound cannot: a class may have mS(l)=2lm^S(l) = 2^lmS(l)=2l for every lll (all open subsets of [0,1][0,1][0,1]) and still satisfy HS(l)/l→0H^S(l)/l \to 0HS(l)/l→0 under a particular PPP, or fail it. The entropy HS(l)H^S(l)HS(l) is the distribution-dependent quantity from which later work on Glivenko–Cantelli classes and on consistency of empirical risk minimization proceeds; the 1981 paper of the same authors extends the criterion to classes of functions. The quantitative form (29) states more than the negation of convergence: when the entropy rate is positive, the maximal deviation stays above a fixed ε\varepsilonε with probability tending to one.

The result has been proved since 1971; it has not been formalized. The platform holds Sauer–Shelah variants over sets of distinct points and PAC bounds with other constants, but no statement of the VC entropy or of Theorem 4. The mission produces machine-checked statements of the entropy criterion and of its supporting lemmas with the paper's own constants (l≥2/ε2l \ge 2/\varepsilon^2l≥2/ε2, 2e−ε2l/82e^{-\varepsilon^2 l/8}2e−ε2l/8, δ=ε2/16\delta = \varepsilon^2/16δ=ε2/16, ε<q/7\varepsilon < q/7ε<q/7).

Difficulty

The sufficiency half is a variant of the proof of the growth-function bound; its new ingredient is the concentration of l−1log⁡2ΔSl^{-1} \log_2 \Delta^Sl−1log2​ΔS (Lemma 4), which needs subadditivity and a law of large numbers over independent blocks of the sample rather than a single mean estimate. The hypergeometric tail estimate behind the permutation bound is omitted in the paper ("a simple but long computation").

Necessity is harder. The obvious attempt, bounding P{π(l)>ε}\mathbf{P}\{\pi^{(l)} > \varepsilon\}P{π(l)>ε} from below by exhibiting a single bad event, fails: SSS may be uncountable and no single AAA deviates with non-vanishing probability. A positive entropy rate has to be converted into a combinatorial statement about typical samples ((26) combines Lemma 4 with an estimate of Φ([ql],l)\Phi([ql], l)Φ([ql],l)), and that statement back into a lower bound on a probability over the product measure; the constants q<14q < \frac14q<41​ and ε<q/7\varepsilon < q/7ε<q/7 must be tracked through both conversions, and the conclusion lim⁡P{π(l)>ε}=1\lim \mathbf{P}\{\pi^{(l)} > \varepsilon\} = 1limP{π(l)>ε}=1 needs the unweakened inequality of step 1°.

Formalization scope

A sample of size lll is a function Fin l → X (positions 0,…,l−10, \dots, l-10,…,l−1) and its law is the product measure Measure.pi (fun _ => P), with P a probability measure. A subsample is a set of positions, so the index counts distinct Finset (Fin l) of the form {i:xi∈A}\{i : x_i \in A\}{i:xi​∈A}. The halves of x : Fin (l + l) → X are x ∘ Fin.castAdd l and x ∘ Fin.natAdd l. PAP_APA​ is P.real A; the suprema π(l)\pi^{(l)}π(l) and ρ(l)\rho^{(l)}ρ(l) are real suprema over the subtype of SSS (values in [0,1][0,1][0,1]; 000 for S=∅S = \emptysetS=∅). HS(l)H^S(l)HS(l) is a Bochner integral of Real.logb 2 of the index, and [ql][ql][ql] is ⌊q * l⌋₊. Probabilities are values in [0,∞][0, \infty][0,∞], except in the inequalities between probabilities (step 1°, the sufficiency estimate), which use Measure.real.

Measurability. The paper assumes, and the statements carry as hypotheses, that the events of SSS are measurable (p. 264), that π(l)\pi^{(l)}π(l) is a random variable (p. 265), that ρ(l)\rho^{(l)}ρ(l) is measurable (p. 268), and that the index is measurable in the sample (p. 273). Each statement carries the ones its proof uses. Without them the Bochner integral defining HS(l)H^S(l)HS(l) can be the junk value 000 and the equivalence can fail; replacing them by "SSS countable" would weaken the theorem. The goal is not trivialized by degenerate cases: the equivalence is not vacuous for any class, and S=∅S = \emptysetS=∅ gives the true instance HS=0H^S = 0HS=0, π(l)=0\pi^{(l)} = 0π(l)=0.

Corrections of the printed text. Lemma 2 is printed for l>2/ε2l > 2/\varepsilon^2l>2/ε2; its proof gives l≥2/ε2l \ge 2/\varepsilon^2l≥2/ε2, which is stated. On p. 275 Lemma 2 is recalled as "2P(C)≥12P(Q)2\mathbf{P}(C) \ge \frac12 P(Q)2P(C)≥21​P(Q)", meaning P(C)≥12P(Q)\mathbf{P}(C) \ge \frac12\mathbf{P}(Q)P(C)≥21​P(Q). On p. 276 the first display carries a stray upper limit "4" on the integral, and the region of integration is printed {log⁡2ΔS≤2δ}\{\log_2 \Delta^S \le 2\delta\}{log2​ΔS≤2δ} where {log⁡2ΔS≤2δl}\{\log_2\Delta^S \le 2\delta l\}{log2​ΔS≤2δl} is meant. The event C′C'C′ is defined with ">2ε> 2\varepsilon>2ε" (p. 276) but integrated in step 3° as θ(⋅−2ε)\theta(\cdot - 2\varepsilon)θ(⋅−2ε), which counts "≥2ε\ge 2\varepsilon≥2ε"; the strict form is stated, and the estimate of step 3° is itself strict. Step 1° is stated unweakened. Milestone texts are verbatim.

Contributions welcome: a reusable development of the index and its submultiplicativity, the hypergeometric tail bound for sampling without replacement, a block law of large numbers for subadditive functionals of i.i.d. samples, and the permutation-invariance argument for product measures on Fin (l + l) → X.

Selected references

  • V. N. Vapnik and A. Ya. Chervonenkis, On the uniform convergence of relative frequencies of events to their probabilities, Theory of Probability and Its Applications 16(2) (1971), 264–280. https://doi.org/10.1137/1116025
  • V. N. Vapnik and A. Ya. Chervonenkis, Necessary and sufficient conditions for the uniform convergence of means to their expectations, Theory of Probability and Its Applications 26(3) (1981), 532–553. https://doi.org/10.1137/1126059
  • M. Talagrand, The Glivenko–Cantelli problem, Annals of Probability 15(3) (1987), 837–870. https://doi.org/10.1214/aop/1176992069
  • R. M. Dudley, E. Giné and J. Zinn, Uniform and universal Glivenko–Cantelli classes, Journal of Theoretical Probability 4(3) (1991), 485–510. https://doi.org/10.1007/BF01210321
16 thms3 active usersReviewed
Algorithmic Game TheoryMachine LearningProbability·Captain: mikedeng1

Calibrated Learning and Correlated Equilibrium III: A Randomized Forecast Calibrated against Every OpponentResearch Paper

Motivation

A forecaster who announces "70% chance of rain" is calibrated if, among the days on which that number was announced, it rained on about 70% of them. Dawid (The well-calibrated Bayesian, JASA 1982) proposed calibration as the minimal requirement of an honest probability forecaster. Oakes (Self-calibrating priors do not exist, JASA 1985) showed that no deterministic forecasting rule can be calibrated against every sequence of outcomes: an adversary who knows the rule can always choose the outcome that contradicts the forecast.

Foster and Vohra (Calibrated learning and correlated equilibrium, Games Econ. Behav. 21 (1997) 40–55) use calibration as the bridge between learning and equilibrium in repeated games. Their Theorem 1 says that if each player best-responds to calibrated forecasts of the opponent, the empirical distribution of play converges to the set of correlated equilibria. That theorem is only useful if calibrated forecasts can actually be produced whatever the opponent does. Theorem 3 of the paper, credited to an unpublished 1991 manuscript of the same authors and proved in the paper's Appendix, says they can, provided the forecaster randomizes.

Timeline. Dawid (1982) defines calibration. Oakes (1985) rules out deterministic calibrated forecasting against arbitrary sequences. Foster and Vohra (1991 manuscript; 1997 paper, Theorem 3 and Appendix) give a randomized forecaster calibrated against any opponent, through a pairwise ("internal") no-regret property. The full argument appeared in Foster and Vohra, Asymptotic calibration, Biometrika 85 (1998). Hart and Mas-Colell (A simple adaptive procedure leading to correlated equilibrium, Econometrica 2000) later made internal regret the standard route to correlated equilibrium.

Setting

Player 2 has n≥1n\ge 1n≥1 pure strategies j∈{0,…,n−1}j\in\{0,\dots,n-1\}j∈{0,…,n−1}. In every round, player 1 announces a forecast p∈Rnp\in\mathbb R^np∈Rn, a probability vector (pj≥0p_j\ge 0pj​≥0, ∑jpj=1\sum_j p_j = 1∑j​pj​=1), and player 2 plays a strategy jjj. The two moves are simultaneous: player 2 does not see the current forecast.

A history hhh of length ttt is the list of the ttt pairs (forecast, play), oldest first. For a forecast vector ppp and a strategy jjj:

  • N(p,t)N(p,t)N(p,t) is the number of rounds of hhh in which ppp was forecast;
  • ρ(p,j,t)\rho(p,j,t)ρ(p,j,t) is the fraction of those rounds in which player 2 played jjj (and 000 if N(p,t)=0N(p,t)=0N(p,t)=0);
  • the calibration score (Eq. (1), p. 49) is
Ct=∑p∑j∣ρ(p,j,t)−pj∣ N(p,t)t.C_t = \sum_p\sum_j \bigl|\rho(p,j,t) - p_j\bigr|\,\frac{N(p,t)}{t}.Ct​=p∑​j∑​​ρ(p,j,t)−pj​​tN(p,t)​.

A randomized forecaster FFF maps each history to a probability distribution on forecasts. A learning rule AAA of player 2 maps each history to a probability distribution on strategies. In round t+1t+1t+1 the forecast is drawn from F(h)F(h)F(h) and the play from A(h)A(h)A(h), independently given the history hhh of the first ttt rounds. This defines the law PF,A\mathbb P_{F,A}PF,A​ of the first ttt rounds (histLaw F A t).

For the Appendix: with kkk forecasts, losses LtiL_t^iLti​ and mixing weights wtiw_t^iwti​, the pairwise regret of replacing forecast iii by forecast jjj is

RTi→j=max⁡{0, ∑t=1Twti (Lti−Ltj)}.R_T^{i\to j} = \max\Bigl\{0,\ \sum_{t=1}^T w_t^i\,(L_t^i - L_t^j)\Bigr\}.RTi→j​=max{0, t=1∑T​wti​(Lti​−Ltj​)}.

Formalization targets

Goal: Theorem 3 (p. 49)

There is a forecaster FFF, with probability-vector forecasts, such that for every learning rule AAA of player 2 and every ε>0\varepsilon>0ε>0,

lim⁡t→∞PF,A(Ct<ε)=1.\lim_{t\to\infty}\mathbb P_{F,A}\bigl(C_t<\varepsilon\bigr) = 1.t→∞lim​PF,A​(Ct​<ε)=1.

The forecaster is fixed before the opponent; no rate is claimed, and the rate may depend on AAA.

Milestones, in the order the Appendix uses them

  1. Flow conservation is solvable (p. 52). For every nonnegative k×kk\times kk×k matrix RRR, k≥1k\ge 1k≥1, there is a probability vector www with wi∑jRi→j=∑jwjRj→iw^i\sum_j R^{i\to j} = \sum_j w^j R^{j\to i}wi∑j​Ri→j=∑j​wjRj→i for all iii.
  2. Lemma 1 (No-Regret) (p. 52). With losses in [0,1][0,1][0,1] and weights solving flow conservation for the previous regrets,
RTi→j≤2kTfor all i,j,T.R_T^{i\to j}\le\sqrt{2kT}\quad\text{for all } i,j,T.RTi→j​≤2kT​for all i,j,T.
  1. Regrets sandwich L-2 calibration (p. 54). For a grid p1,…,pkp^1,\dots,p^kp1,…,pk that is ε\varepsilonε-dense in squared distance and losses Lti=∣Xt−pi∣2L_t^i = |X_t - p^i|^2Lti​=∣Xt​−pi∣2,
∑imax⁡jRTi→jT ≤ C2,w(T) ≤ ε+∑imax⁡jRTi→jT,\sum_i\max_j \frac{R_T^{i\to j}}{T}\ \le\ C_{2,w}(T)\ \le\ \varepsilon + \sum_i\max_j\frac{R_T^{i\to j}}{T},i∑​jmax​TRTi→j​​ ≤ C2,w​(T) ≤ ε+i∑​jmax​TRTi→j​​,

with C2,wC_{2,w}C2,w​ the fractional L-2 calibration score. 4. L-1 versus L-2 (p. 54). For each jjj, ∑p∣ρ(p,j,t)−pj∣ N(p,t)/t≤∑p(ρ(p,j,t)−pj)2N(p,t)/t\sum_p|\rho(p,j,t)-p_j|\,N(p,t)/t \le \sqrt{\sum_p(\rho(p,j,t)-p_j)^2 N(p,t)/t}∑p​∣ρ(p,j,t)−pj​∣N(p,t)/t≤∑p​(ρ(p,j,t)−pj​)2N(p,t)/t​.

Significance

The result. Theorem 3 makes the hypothesis of Theorem 1 attainable: combined, they show that there are learning procedures under which play converges in probability to the set of correlated equilibria of any finite game (the paper's Corollary, p. 49). The intermediate Lemma 1 is an early internal-regret bound; internal (swap) regret minimization later became the standard algorithmic route to correlated equilibria and to calibrated prediction in online learning.

Formalizing it. The result is proved, in the 1997 Appendix in telegraphic form and in full in Foster and Vohra (1998). The Appendix leaves several steps informal (see Formalization scope), so a machine-checked proof must supply them. No formalization of calibration or of internal regret was found on Prove2Me as of 2026-09-26. The mission produces a formal model of randomized forecasters against adaptive opponents, a checked internal-regret bound with an explicit constant, and the passage from regret to calibration.

Difficulty

The obvious approach is to pick, at each round, a forecast that corrects the current miscalibration. This is a deterministic rule, and by Oakes' theorem an opponent can defeat it. Randomization alone does not help either: the forecaster must randomize in a way that controls every pairwise regret Ri→jR^{i\to j}Ri→j at once, not only the regret against the best fixed forecast. External no-regret does not imply calibration.

Two further gaps separate Lemma 1 from Theorem 3. First, the Appendix controls a fractional score in which the event "forecast pip^ipi was issued" is replaced by its probability wtiw_t^iwti​. The realized calibration score involves the random choices, so a concentration argument is needed against an adaptive opponent. Second, a fixed grid gives calibration only up to its mesh ε\varepsilonε. Exact convergence requires letting the grid size kkk grow and ε\varepsilonε shrink over time, and the scores of the different phases must be combined.

Formalization scope

  • Model. Strategies are Fin n; forecasts are vectors Fin n → ℝ that are probability vectors (IsDist). Histories are Lean lists of (forecast, play) pairs, oldest first. Forecaster and opponent are maps from histories to Mathlib PMFs. The history law is built with PMF.bind/PMF.map, with the two draws independent given the past. P(Ct<ε)\mathbb P(C_t<\varepsilon)P(Ct​<ε) is the toOuterMeasure of the law, in ℝ≥0∞; no σ-algebra on histories is used.
  • Opponent. The opponent may be randomized and may depend on all past forecasts and plays; fixed sequences and deterministic rules are special cases. The opponent never sees the current forecast. Letting it see the current forecast would make the goal false, and restricting to fixed sequences would make it weaker than the paper.
  • Quantifiers. The forecaster is chosen before the opponent (∃ F, ∀ A). The reverse order is trivial, since one can forecast AAA's next play.
  • Conventions. Rounds are counted from 000, so the paper's rounds 1,…,t1,\dots,t1,…,t are the first ttt list entries. The paper's Rt−1R_{t-1}Rt−1​ is the regret over the 000-based rounds before ttt. Forecasts are grouped by exact equality of real vectors. Sums over ppp run over the forecasts that occur, and every other term vanishes. Scores divide by the history length, and are 000 for the empty history. "Converges in probability" is the lim⁡P(Ct<ε)=1\lim\mathbb P(C_t<\varepsilon)=1limP(Ct​<ε)=1 form the page states. The grid of milestone 3 is indexed 1,…,k1,\dots,k1,…,k (the page writes i=0,…,ki = 0,\dots,ki=0,…,k on p. 53 and 1,…,k1,\dots,k1,…,k in Lemma 1), and "within ε\varepsilonε" is read in squared Euclidean distance.
  • Pinned reading. The page writes the middle term of milestone 3 as E(C2(t))E(C_2(t))E(C2​(t)) with a garbled formula. The mission states the inequality for the fractional score C2,wC_{2,w}C2,w​, for which it holds.
  • Steps the paper leaves informal, not stated as milestones. (a) "E(C2(t))≤ε+O(k/2)E(C_2(t))\le\varepsilon + O(k/\sqrt2)E(C2​(t))≤ε+O(k/2​)" when the weights solve flow conservation; the OOO-term is garbled and should decay in ttt. (b) "if we let kkk grow slowly and ε\varepsilonε go slowly to zero … C2(t)→0C_2(t)\to 0C2​(t)→0 in expectation which implies C2(t)→0C_2(t)\to0C2​(t)→0 in probability by Jensen's inequality", together with the passage from the fractional score to the realized one. Solvers will have to formalize these steps on the way to the goal.
  • Not included. The Corollary on p. 49 (convergence in probability of play to the correlated equilibria when both players use the scheme). It needs the game layer and a quantitative form of Theorem 1, and the page gives no proof.
  • Infrastructure welcome. Finite Markov chain stationary distributions (for milestone 1), for which Prove2Me has MarkovChain.exists_isStationary for row-stochastic matrices. Also useful: martingale concentration for PMF-built processes, and Cauchy–Schwarz with weights. The history-law construction is reusable for any repeated forecasting game.

Selected references

  • D. P. Foster and R. V. Vohra, Calibrated learning and correlated equilibrium, Games and Economic Behavior 21 (1997) 40–55. https://doi.org/10.1006/game.1997.0595
  • D. P. Foster and R. V. Vohra, Asymptotic calibration, Biometrika 85 (1998) 379–390. https://doi.org/10.1093/biomet/85.2.379
  • A. P. Dawid, The well-calibrated Bayesian, J. Amer. Statist. Assoc. 77 (1982) 605–613. https://doi.org/10.1080/01621459.1982.10477856
  • D. Oakes, Self-calibrating priors do not exist, J. Amer. Statist. Assoc. 80 (1985) 339. https://doi.org/10.1080/01621459.1985.10478117
  • S. Hart and A. Mas-Colell, A simple adaptive procedure leading to correlated equilibrium, Econometrica 68 (2000) 1127–1150. https://doi.org/10.1111/1468-0262.00153
10 thms3 active usersReviewed
Dynamic ProgrammingOperations ResearchProbability·Captain: mikedeng1

Markovian Decision Processes with Uncertain Transition Probabilities II: Max-Max and Max-Min Optimal Returns Bound the Bayesian Optimal ReturnResearch Paper

Motivation

A Markovian decision process (Howard, 1960) models a controller who, in each of finitely many states, picks a decision, earns a reward and moves to a random next state according to known transition probabilities. In applications (inventory control, equipment replacement, quality control) those probabilities are estimated, not known. Satia and Lave (Operations Research 21(3), 1973) treat the uncertainty in two ways: a game-theoretic formulation, in which each unknown row only lies in a given set, and a Bayesian formulation, going back to Silver (1963) and Martin (1967), in which the controller holds a prior on the unknown matrix and learns from observed transitions.

The Bayesian problem is the natural one but its state includes the whole prior, so it cannot be solved exactly beyond small cases. The paper's contribution in the Bayesian part is a pair of computable bounds on the Bayesian optimal return in terms of the two game-theoretic values (max-max and max-min). This mission formalizes those bounds and the chain of facts they rest on.

Setting

There are NNN states iii and, in state iii, a finite nonempty set KiK_iKi​ of decisions. A transition i→ji \to ji→j under decision kkk earns rijkr^k_{ij}rijk​ and rewards are discounted by β\betaβ, 0≤β<10 \le \beta < 10≤β<1. The row pik=(pijk)jp_i^k = (p^k_{ij})_jpik​=(pijk​)j​ of transition probabilities is unknown; it is known to lie in a closed convex nonempty set SikS_i^kSik​ of probability vectors, and S={P:pik∈Sik for all i,k}S = \{P : p_i^k \in S_i^k \text{ for all } i, k\}S={P:pik​∈Sik​ for all i,k}.

A prior ggg is a probability distribution on matrices P=(pik)P = (p_i^k)P=(pik​) whose rows are all probability vectors. Its means are pˉijk=E(pijk)\bar p^k_{ij} = E(p^k_{ij})pˉ​ijk​=E(pijk​). After a transition l→jl \to jl→j under decision mmm the prior is replaced by the Bayes transformation Tljmg(P)=C pljm g(P)T^m_{lj} g(P) = C\,p^m_{lj}\,g(P)Tljm​g(P)=Cpljm​g(P) (Eq. (8)), with CCC the normalizing constant. The Bayesian optimal return f(i,g)f(i,g)f(i,g) solves the recursion

f(i,g)=max⁡k∈Ki{∑jpˉijkrijk+β∑jpˉijkf(j,Tijkg)}.(10)f(i, g) = \max_{k \in K_i} \Big\{ \sum_j \bar p^k_{ij} r^k_{ij} + \beta \sum_j \bar p^k_{ij} f(j, T^k_{ij} g) \Big\}. \qquad (10)f(i,g)=k∈Ki​max​{j∑​pˉ​ijk​rijk​+βj∑​pˉ​ijk​f(j,Tijk​g)}.(10)

The max-max and max-min values V+V^+V+, V−V^-V− solve

Vi±=max⁡k∈Kimax/min⁡pik∈Sik{∑jpijkrijk+β∑jpijkVj±},V_i^\pm = \max_{k \in K_i} \operatorname*{max/min}_{p_i^k \in S_i^k} \Big\{ \sum_j p^k_{ij} r^k_{ij} + \beta \sum_j p^k_{ij} V_j^\pm \Big\},Vi±​=k∈Ki​max​pik​∈Sik​max/min​{j∑​pijk​rijk​+βj∑​pijk​Vj±​},

with max for V+V^+V+ and min for V−V^-V−. Finally α=prob⁡(P∈S∣g)\alpha = \operatorname{prob}(P \in S \mid g)α=prob(P∈S∣g), the prior probability that the true matrix lies in SSS.

The Lean development lives in the namespace SatiaLave.Bayes: UncertainMDP, IsPrior, pbar, bayes, SolvesEq10, SolvesVplus, SolvesVminus, alpha, rmax, rmin, policyValue.

Formalization targets

Goal: Propositions 9 and 10

For every bounded solution fff of (10), all solutions V+V^+V+, V−V^-V−, every prior ggg and every state iii,

αVi−+(1−α)min⁡i,j,krijk1−β  ≤  f(i,g)  ≤  αVi++(1−α)max⁡i,j,krijk1−β,\alpha V_i^- + (1-\alpha)\min_{i,j,k}\frac{r^k_{ij}}{1-\beta} \;\le\; f(i,g) \;\le\; \alpha V_i^+ + (1-\alpha)\max_{i,j,k}\frac{r^k_{ij}}{1-\beta},αVi−​+(1−α)i,j,kmin​1−βrijk​​≤f(i,g)≤αVi+​+(1−α)i,j,kmax​1−βrijk​​,

together with the existence of fff, V+V^+V+ and V−V^-V−. Both halves are the paper's printed statements.

Milestones

  1. Proposition 6 (Martin): (9)/(10) has a unique bounded solution (unique at priors).
  2. No learning (p. 733): at a point-mass prior δP\delta_PδP​, f(⋅,δP)f(\cdot,\delta_P)f(⋅,δP​) solves the optimality equations of the process with known PPP.
  3. Proposition 8: f(i,g)f(i,g)f(i,g) is convex in ggg.
  4. Jensen step (proof of Proposition 9): f(i,g)≤∫f(i,δP) dg(P)f(i,g) \le \int f(i,\delta_P)\,dg(P)f(i,g)≤∫f(i,δP​)dg(P).
  5. Policy step (proof of Proposition 10): f(i,g)≥∫[q+βPAq+β2[PA]2q+⋯ ]i dg(P)f(i,g) \ge \int [q + \beta P^A q + \beta^2 [P^A]^2 q + \cdots]_i\,dg(P)f(i,g)≥∫[q+βPAq+β2[PA]2q+⋯]i​dg(P) for every pure stationary policy AAA.

Significance

The result. The bounds sandwich an intractable quantity between two quantities computable by finite algorithms (the max-max and max-min policy-iteration procedures of the same paper), weighted by a single prior probability α\alphaα. When the prior concentrates on SSS (α→1\alpha \to 1α→1) the bounds become Vi−≤f(i,g)≤Vi+V_i^- \le f(i,g) \le V_i^+Vi−​≤f(i,g)≤Vi+​: the Bayesian return lies between the pessimistic and optimistic robust values. They are the upper and lower bounds on the return that the paper's implicit-enumeration method (the decision tree of its Fig. 2 and Proposition 12) uses to compare decisions. The Jensen step is a value-of-information inequality (Bayesian optimal return is at most the expected full-information optimal return), which recurs throughout Bayesian control and bandit theory.

Formalizing it. The results are proved on paper (Propositions 6 and 8 by reference to Martin's book and Satia's thesis, Propositions 9 and 10 in the text); none is machine-checked. The mission produces a Lean model of Bayes-adaptive Markov decision processes with priors as measures, the Bayes transformation and its fixed-point recursion, and the link between the Bayesian and the robust (rectangular) formulations. Martin's existence-uniqueness theorem and the convexity of the Bayesian value are reusable for any Bayes-adaptive model.

Difficulty

The prior space is infinite-dimensional and not a vector space, so the recursion (10) lives on a space of measures, and the usual finite-state arguments do not apply verbatim. Proposition 8 gives convexity only along finite mixtures, while the proof of Proposition 9 applies Jensen's inequality to the integral mixture g=∫δP dg(P)g = \int \delta_P\,dg(P)g=∫δP​dg(P) of point masses; bridging the two, or proving the value-of-information inequality directly, is the central step. The paper also restricts the point masses to xik∈Sikx_i^k \in S_i^kxik​∈Sik​, which cannot represent a prior with mass outside SSS; the formal statement integrates over every transition matrix, as the next line of the paper's display requires. Measurability of P↦f(i,δP)P \mapsto f(i,\delta_P)P↦f(i,δP​) is not automatic, since fff is only characterized by a functional equation.

Formalization scope

  • States are a nonempty Fintype S; decisions a dependent family D i of nonempty finite types. A matrix is P : (i : S) → D i → S → ℝ with the product Borel σ\sigmaσ-algebra.
  • Priors are measures: a probability measure giving full mass to matrices whose rows are probability vectors. This generalizes the paper's densities g(P)g(P)g(P) and includes the point masses axa_xax​ its proof uses.
  • Bayes transformation at pˉ=0\bar p = 0pˉ​=0: the normalizing constant does not exist; bayes then returns ggg. That posterior is always multiplied by pˉ=0\bar p = 0pˉ​=0 in (10), so the choice is immaterial.
  • Readings of informal words. "The problem reduces to a Markovian decision process" = at a point-mass prior, fixed by every Bayes transformation, fff solves the known-PPP optimality equations. "Convex in ggg" = convex along mixtures of priors. "Unique set of bounded functions" = two bounded solutions agree at every prior (values at non-priors are unconstrained). "Satisfy (9)" is formalized as (10), which the paper derives from (9) by linearity of EEE. max⁡P∈S\max_{P\in S}maxP∈S​/min⁡P∈S\min_{P\in S}minP∈S​ in V±V^\pmV± is taken over the row pik∈Sikp_i^k \in S_i^kpik​∈Sik​ (the only row that enters; SSS is a product), as ⨆/⨅ over a nonempty bounded set. "Obviously f(i,g)≥ViAf(i,g)\ge V_i^Af(i,g)≥ViA​" is stated for every pure stationary policy AAA, not only a max-min optimal one. The policy return is the componentwise series ∑nβn(PA)nq\sum_n \beta^n (P^A)^n q∑n​βn(PA)nq.
  • Added hypotheses, not printed: 0≤β<10 \le \beta < 10≤β<1; Sik≠∅S_i^k \ne \emptysetSik​=∅; N≥1N \ge 1N≥1. Printed and kept: SikS_i^kSik​ closed and convex.
  • fff, V+V^+V+, V−V^-V− are quantified as solutions of their equations; α\alphaα is computed from ggg, never a free parameter; max⁡i,j,k[rijk/(1−β)]\max_{i,j,k}[r^k_{ij}/(1-\beta)]maxi,j,k​[rijk​/(1−β)] ranges over all states i,ji,ji,j and k∈Kik \in K_ik∈Ki​. Integrability of the integrands in milestones 4 and 5 is part of their conclusions.
  • Trivializations ruled out. A free α∈[0,1]\alpha \in [0,1]α∈[0,1], or fff defined off priors, would make the goal false or vacuous; the goal also asserts that bounded fff and V±V^\pmV± exist, so its universal part is not vacuous.
  • Not in scope: Proposition 7 (matrix-beta conjugacy, which needs a Dirichlet distribution), Propositions 11–13 and the numerical example.

Welcome contributions: the Banach fixed-point argument for (10) on bounded functions of priors; lemmas that bayes maps priors to priors and that point masses are fixed; continuity of the known-PPP optimal value in PPP; a general Jensen inequality for functions convex along mixtures of probability measures.

Selected references

  • J. K. Satia and R. E. Lave, Jr., Markovian Decision Processes with Uncertain Transition Probabilities, Operations Research 21(3), 728–740, 1973. https://doi.org/10.1287/opre.21.3.728
  • J. J. Martin, Bayesian Decision Problems and Markov Chains, Wiley, New York, 1967.
  • E. A. Silver, Markovian Decision Processes with Uncertain Transition Probabilities or Rewards, Interim Technical Report No. 1, Operations Research Center, Massachusetts Institute of Technology, August 1963.
  • R. A. Howard, Dynamic Programming and Markov Processes, MIT Press, 1960.
  • J. K. Satia, Markovian Decision Process with Uncertain Transition Matrices or/and Probabilistic Observation of States, Ph.D. dissertation, Stanford University, 1968.
7 thms3 active usersReviewed
Machine LearningProbability·Captain: mikedeng1

The Power of Convex Relaxation: Near-Optimal Matrix Completion III: No Method Recovers Incoherent Rank-r Matrices below the Sampling Rate (I.20)Research Paper

Motivation

Matrix completion asks to recover a matrix from a small random subset of its entries. It models collaborative filtering (a ratings table with most entries missing), sensor-network localisation from partial distance data, and system identification. With no structure the task is hopeless, so one assumes the matrix has low rank rrr and that its information is not concentrated in a few entries (incoherence).

Candès and Recht (Found. Comput. Math., 2009) showed that nuclear-norm minimisation recovers such a matrix from about n6/5rlog⁡nn^{6/5} r\log nn6/5rlogn random entries. Candès and Tao (IEEE Trans. Inf. Theory, 2010) lowered this to nr polylog(n)n r\,\mathrm{polylog}(n)nrpolylog(n). The same paper also asks how few entries any method could possibly use, and answers it with a lower bound, Theorem 1.7: below about μ0nrlog⁡n\mu_0 n r\log nμ0​nrlogn observed entries, no algorithm can succeed. This mission formalizes that lower bound. The two upper bounds of the same paper are separate missions of this series.

Setting

Work with real n×nn\times nn×n matrices. For a matrix MMM, let U⊆RnU \subseteq \mathbb{R}^nU⊆Rn be its column space and V⊆RnV\subseteq \mathbb{R}^nV⊆Rn its row space, and let PUP_UPU​, PVP_VPV​ be the orthogonal projections onto them. Let eae_aea​ be the aaa-th standard basis vector.

Fix an integer rrr and a real μ0\mu_0μ0​. A matrix MMM has rank at most rrr and obeys the incoherence property with parameter μ0\mu_0μ0​ (the paper's (I.18)) if rank⁡M≤r\operatorname{rank}M \le rrankM≤r and

∥PUea∥2≤μ0rn,∥PVeb∥2≤μ0rnfor all a,b∈[n].\|P_U e_a\|^2 \le \frac{\mu_0 r}{n},\qquad \|P_V e_b\|^2 \le \frac{\mu_0 r}{n}\qquad\text{for all } a,b\in[n].∥PU​ea​∥2≤nμ0​r​,∥PV​eb​∥2≤nμ0​r​for all a,b∈[n].

Since ∑a∥PUea∥2=dim⁡U\sum_a \|P_U e_a\|^2 = \dim U∑a​∥PU​ea​∥2=dimU, a matrix of rank exactly rrr can satisfy this only when μ0≥1\mu_0\ge 1μ0​≥1; the smallest possible value μ0=1\mu_0 = 1μ0​=1 means the column and row spaces are spread evenly over the coordinates.

Bernoulli sampling. Fix m≥1m \ge 1m≥1 and set p=m/n2p = m/n^2p=m/n2. The observed set Ω⊆[n]×[n]\Omega\subseteq[n]\times[n]Ω⊆[n]×[n] contains each entry independently with probability ppp, so mmm is the expected number of observed entries. The sampling operator PΩ\mathcal{P}_\OmegaPΩ​ keeps the entries of a matrix that lie in Ω\OmegaΩ and sets the others to 000. A recovery method sees only PΩ(M)\mathcal{P}_\Omega(M)PΩ​(M).

The sampling conditions are, with the natural logarithm,

m≥n2(1−e−μ0rnlog⁡(n2δ))(I.20)m \ge n^2\left(1 - e^{-\frac{\mu_0 r}{n}\log\left(\frac{n}{2\delta}\right)}\right) \qquad \text{(I.20)}m≥n2(1−e−nμ0​r​log(2δn​))(I.20) m≥(1−ϵ) μ0nrlog⁡(n2δ),ϵ:=12μ0rnlog⁡(n2δ).(I.21)m \ge (1-\epsilon)\,\mu_0 n r\log\left(\frac{n}{2\delta}\right),\qquad \epsilon := \frac12\frac{\mu_0 r}{n}\log\left(\frac{n}{2\delta}\right). \qquad \text{(I.21)}m≥(1−ϵ)μ0​nrlog(2δn​),ϵ:=21​nμ0​r​log(2δn​).(I.21)

Formalization targets

Goal: Theorem 1.7 (p. 2058)

Fix 1≤m1 \le m1≤m, 1≤r≤n1 \le r \le n1≤r≤n, μ0≥1\mu_0\ge 1μ0​≥1 and 0<δ<1/20<\delta<1/20<δ<1/2, with ℓ:=n/(μ0r)\ell := n/(\mu_0 r)ℓ:=n/(μ0​r) an integer. If (I.20) fails, or (I.21) fails, then

PΩ(there are infinitely many pairs M≠M′ of rank≤r, incoherent with parameter μ0, with PΩ(M)=PΩ(M′)) ≥ δ.\mathbb{P}_\Omega\Bigl(\text{there are infinitely many pairs } M\ne M' \text{ of rank} \le r, \text{ incoherent with parameter } \mu_0, \text{ with } \mathcal{P}_\Omega(M)=\mathcal{P}_\Omega(M')\Bigr) \ \ge\ \delta .PΩ​(there are infinitely many pairs M=M′ of rank≤r, incoherent with parameter μ0​, with PΩ​(M)=PΩ​(M′)) ≥ δ.

On that event, the observations cannot tell MMM from M′M'M′, so no method can recover every such matrix with probability greater than 1−δ1-\delta1−δ. The statement fixes no constant beyond those the paper prints.

Milestones (Section II)

  1. For pairwise disjoint sets of entries S1,…,SnS_1,\dots,S_nS1​,…,Sn​ of size ℓ\ellℓ, P(every Sa is sampled)=(1−(1−p)ℓ)n\mathbb{P}(\text{every } S_a \text{ is sampled}) = (1-(1-p)^\ell)^nP(every Sa​ is sampled)=(1−(1−p)ℓ)n.
  2. For n≥1n\ge1n≥1, π∈[0,1]\pi\in[0,1]π∈[0,1] and 0<δ<1/20<\delta<1/20<δ<1/2: (1−π)n≥1−δ(1-\pi)^n \ge 1-\delta(1−π)n≥1−δ implies π≤2δ/n\pi \le 2\delta/nπ≤2δ/n.
  3. With p=m/n2p = m/n^2p=m/n2 and the theorem's parameters: (1−p)ℓ≤2δ/n(1-p)^\ell \le 2\delta/n(1−p)ℓ≤2δ/n implies (I.20).
  4. 1−e−x>x−x2/21-e^{-x} > x - x^2/21−e−x>x−x2/2 for every x>0x>0x>0 (the paper prints x≥0x\ge0x≥0; see Formalization scope).
  5. The second part of Theorem 1.7: for the theorem's parameters, (I.20) implies (I.21).

Significance

The result. Theorem 1.7 shows that the sample complexity nr polylog(n)n r\,\mathrm{polylog}(n)nrpolylog(n) of the paper's upper bounds is close to optimal: about μ0nrlog⁡n\mu_0 n r\log nμ0​nrlogn entries are necessary, however the matrix is reconstructed. The count exceeds the 2nr−r22nr - r^22nr−r2 degrees of freedom of a rank-rrr matrix by the factor μ0log⁡n\mu_0\log nμ0​logn. The logarithm is a coupon-collector effect: every row has to be sampled. The factor μ0\mu_0μ0​ shows that the oversampling grows in proportion to the coherence. The bound is information-theoretic, and it holds even when the rank bound and the coherence are known in advance.

Formalizing it. The theorem and its proof in Section II are published. No machine-checked version is known to exist, and the platform has no lower bound for matrix completion. A formal proof would also check the printed argument, whose steps are compressed. It would yield reusable pieces: a Lean predicate for incoherence of matrices of bounded rank built on Mathlib's orthogonal projections, the independence computation for Bernoulli sampling over disjoint entry sets, and the elementary estimates that turn a success probability into a sampling rate.

Difficulty

The probabilistic and analytic parts are elementary. The difficulty is in building the hard instances as matrices and certifying them. For each observation set one has to exhibit, on an event of probability at least δ\deltaδ, an infinite family of distinct pairs that agree on Ω\OmegaΩ. Every member must have rank at most rrr and meet both incoherence bounds, measured through projections onto its column and row spaces, and the pairs must stay distinct across the family. The paper describes the instances only informally. They have to be pinned down so that whatever distinguishes MMM from M′M'M′ is really invisible on Ω\OmegaΩ, while the incoherence bounds still hold for every admissible μ0≥1\mu_0\ge1μ0​≥1 and r≤nr\le nr≤n. Computing the column space and the projection norms of an explicit matrix in Lean is the main infrastructure cost.

Formalization scope

Matrices are Matrix (Fin n) (Fin n) ℝ (the platform's MatrixCompletion.RealMatrix n n). The observation model is the platform's bernoulliEventProb with rate m/n2m/n^2m/n2, the sum over all Ω\OmegaΩ of p∣Ω∣(1−p)n2−∣Ω∣p^{|\Omega|}(1-p)^{n^2-|\Omega|}p∣Ω∣(1−p)n2−∣Ω∣. The sampling operator is the platform's samplingProjection. Logarithms and exponentials are Real.log, Real.exp. The mission's own definitions are IncoherentRankAtMost r μ₀ M (rank at most rrr, with the projection bounds computed from Mathlib's Submodule.starProjection onto the ranges of MMM and M⊤M^\topM⊤ in EuclideanSpace ℝ (Fin n)) and SamplingConditionI20, SamplingConditionI21.

The formalization commits to four readings:

  • Order of quantifiers. The event is "the set of bad pairs is infinite", evaluated for each Ω\OmegaΩ, so the pairs may depend on Ω\OmegaΩ. This is what Section II establishes and what the sentence after the theorem uses. The reading "fixed M≠M′M\ne M'M=M′ with P(PΩ(M)=PΩ(M′))≥δ\mathbb{P}(\mathcal{P}_\Omega(M) = \mathcal{P}_\Omega(M'))\ge\deltaP(PΩ​(M)=PΩ​(M′))≥δ" is a different statement and is not the goal.
  • Integrality of ℓ\ellℓ. The hypothesis that ℓ=n/(μ0r)\ell = n/(\mu_0 r)ℓ=n/(μ0​r) is an integer is the paper's own "without loss of generality" of Section II, and it is stated explicitly. It forces μ0r≤n\mu_0 r \le nμ0​r≤n.
  • "Fix 1≤m,r≤n1\le m, r\le n1≤m,r≤n" is read as 1≤m1\le m1≤m and 1≤r≤n1\le r\le n1≤r≤n. No upper bound on mmm is imposed, since the failure of (I.20) already gives m<n2m<n^2m<n2.
  • Standing assumptions. Section I-H assumes m≥2nrm\ge 2nrm≥2nr and nnn larger than an absolute constant for the rest of the paper. Those assumptions serve the upper bounds. Theorem 1.7 lists its own ranges, and only those are imposed.

The hypothesis "(I.20) fails or (I.21) fails" covers both parts of the theorem. The last sentence of Section II proves the second part from 1−e−x>x−x2/21 - e^{-x} > x - x^2/21−e−x>x−x2/2, which the paper states "whenever x≥0x \ge 0x≥0". At x=0x=0x=0 the two sides are equal, so the strict inequality is false there; milestone 4 states the corrected range x>0x>0x>0, which is all the paper uses, since its x=μ0rnlog⁡n2δx = \frac{\mu_0 r}{n}\log\frac{n}{2\delta}x=nμ0​r​log2δn​ is positive.

Trivializing formalizations are ruled out. The set of pairs requires M≠M′M\ne M'M=M′, so the diagonal pairs (M,M)(M,M)(M,M) do not count. "Infinitely many" is Set.Infinite of a set of pairs, not "at least one". Incoherence uses the theorem's rrr and the actual column and row spaces, so the class is the paper's. The hypotheses are satisfiable, for example n=4n=4n=4, r=1r=1r=1, μ0=1\mu_0=1μ0​=1, ℓ=4\ell=4ℓ=4, δ=0.1\delta=0.1δ=0.1, m=1m=1m=1.

Contributions are welcome on every milestone. The Bernoulli independence computation and the incoherence predicate can be reused in the other two missions of this series and in any lower bound for sampling problems.

Selected references

  • E. J. Candès and T. Tao, The Power of Convex Relaxation: Near-Optimal Matrix Completion, IEEE Transactions on Information Theory 56(5):2053–2080, 2010. https://doi.org/10.1109/TIT.2010.2044061
  • E. J. Candès and B. Recht, Exact Matrix Completion via Convex Optimization, Foundations of Computational Mathematics 9(6):717–772, 2009. https://doi.org/10.1007/s10208-009-9045-5
12 thms3 active usersReviewed
Machine LearningOperations ResearchOptimization·Captain: mikedeng1

Generalization Bounds in the Predict-then-Optimize Framework II: Margin-Based Generalization Bound for the SPO Loss under the Strength PropertyResearch Paper

Motivation

In the predict-then-optimize paradigm a model first predicts the cost vector of a linear optimization problem from contextual features, and the prediction is then fed to an optimization solver that returns a decision. Examples include routing with predicted travel times and portfolio choice with predicted returns. The quality of a prediction is judged by the decision it produces. The Smart Predict-then-Optimize (SPO) loss of Elmachtoub and Grigas (Management Science 2022) measures exactly that: the excess cost of acting on the prediction instead of on the true cost vector.

El Balghiti, Elmachtoub, Grigas and Tewari (arXiv:1905.11488v3) ask when a model with small empirical SPO loss also has small expected SPO loss. The SPO loss is non-convex and discontinuous, so standard Lipschitz-contraction arguments do not apply to it directly. Their Section 4 introduces a margin version of the SPO loss, in the spirit of the margin theory of Koltchinskii and Panchenko (Ann. Statist. 2002) for classification. They show that it is Lipschitz under a geometric condition on the feasible region, and derive a generalization bound in terms of the multivariate Rademacher complexity of the hypothesis class. This mission formalizes that bound.

Setting

Decisions live in Rd\mathbb R^dRd with a norm ∥⋅∥\|\cdot\|∥⋅∥; cost vectors are linear functionals with the dual norm ∥c∥∗=max⁡∥w∥≤1c⊤w\|c\|_*=\max_{\|w\|\le1}c^\top w∥c∥∗​=max∥w∥≤1​c⊤w. The feasible region S⊆RdS\subseteq\mathbb R^dS⊆Rd is nonempty, compact and convex, and throughout Section 4 it is not a singleton. An optimization oracle w∗w^*w∗ maps each cost vector ccc to some minimizer w∗(c)∈arg⁡min⁡w∈Sc⊤ww^*(c)\in\arg\min_{w\in S}c^\top ww∗(c)∈argminw∈S​c⊤w. The SPO loss of a prediction c^\hat cc^ against the realized cost ccc is

ℓSPO(c^,c)=c⊤w∗(c^)−c⊤w∗(c),\ell_{\rm SPO}(\hat c,c)=c^\top w^*(\hat c)-c^\top w^*(c),ℓSPO​(c^,c)=c⊤w∗(c^)−c⊤w∗(c),

and the linear optimization gap is ωS(c)=max⁡w∈Sc⊤w−min⁡w∈Sc⊤w\omega_S(c)=\max_{w\in S}c^\top w-\min_{w\in S}c^\top wωS​(c)=maxw∈S​c⊤w−minw∈S​c⊤w, with ωS(C)=sup⁡c∈CωS(c)\omega_S(\mathcal C)=\sup_{c\in\mathcal C}\omega_S(c)ωS​(C)=supc∈C​ωS​(c) and ρ2(C)=sup⁡c∈C∥c∥2\rho_2(\mathcal C)=\sup_{c\in\mathcal C}\|c\|_2ρ2​(C)=supc∈C​∥c∥2​ for the set C\mathcal CC of possible true costs.

A cost vector is degenerate if min⁡w∈Sc^⊤w\min_{w\in S}\hat c^\top wminw∈S​c^⊤w has more than one optimal solution; C∘\mathcal C^\circC∘ is the set of degenerate costs. The distance to degeneracy is νS(c^)=inf⁡c∈C∘∥c−c^∥∗\nu_S(\hat c)=\inf_{c\in\mathcal C^\circ}\|c-\hat c\|_*νS​(c^)=infc∈C∘​∥c−c^∥∗​. The region SSS has the strength property with parameter μ>0\mu>0μ>0 if

c^⊤(w−w∗(c^))≥μ νS(c^)2 ∥w−w∗(c^)∥2for all w∈S and all c^.\hat c^\top\big(w-w^*(\hat c)\big)\ge\frac{\mu\,\nu_S(\hat c)}{2}\,\|w-w^*(\hat c)\|^2\qquad\text{for all }w\in S\text{ and all }\hat c .c^⊤(w−w∗(c^))≥2μνS​(c^)​∥w−w∗(c^)∥2for all w∈S and all c^.

For γ>0\gamma>0γ>0 the γ\gammaγ-margin SPO loss ℓSPOγ(c^,c)\ell^\gamma_{\rm SPO}(\hat c,c)ℓSPOγ​(c^,c) equals ℓSPO(c^,c)\ell_{\rm SPO}(\hat c,c)ℓSPO​(c^,c) when νS(c^)>γ\nu_S(\hat c)>\gammaνS​(c^)>γ and νS(c^)γℓSPO(c^,c)+(1−νS(c^)γ)ωS(c)\frac{\nu_S(\hat c)}{\gamma}\ell_{\rm SPO}(\hat c,c)+\big(1-\frac{\nu_S(\hat c)}{\gamma}\big)\omega_S(c)γνS​(c^)​ℓSPO​(c^,c)+(1−γνS​(c^)​)ωS​(c) otherwise. It dominates the SPO loss.

Data (x,c)(x,c)(x,c) are drawn from a distribution D\mathcal DD on features X\mathcal XX and costs in C\mathcal CC, and H\mathcal HH is a class of prediction functions f:X→Rdf:\mathcal X\to\mathbb R^df:X→Rd. The SPO risk is RSPO(f)=ED[ℓSPO(f(x),c)]R_{\rm SPO}(f)=\mathbb E_{\mathcal D}[\ell_{\rm SPO}(f(x),c)]RSPO​(f)=ED​[ℓSPO​(f(x),c)] and the empirical margin risk is R^SPOγ(f)=1n∑iℓSPOγ(f(xi),ci)\hat R^\gamma_{\rm SPO}(f)=\frac1n\sum_i\ell^\gamma_{\rm SPO}(f(x_i),c_i)R^SPOγ​(f)=n1​∑i​ℓSPOγ​(f(xi​),ci​). The multivariate empirical Rademacher complexity is R^n(H)=Eσ[sup⁡f∈H1n∑iσi⊤f(xi)]\hat{\mathfrak R}^n(\mathcal H)=\mathbb E_{\boldsymbol\sigma}\big[\sup_{f\in\mathcal H}\frac1n\sum_i\boldsymbol\sigma_i^\top f(x_i)\big]R^n(H)=Eσ​[supf∈H​n1​∑i​σi⊤​f(xi​)] with i.i.d. Rademacher vectors σi∈{±1}d\boldsymbol\sigma_i\in\{\pm1\}^dσi​∈{±1}d, and Rn(H)\mathfrak R^n(\mathcal H)Rn(H) is its expectation over the sample.

Formalization targets

Goal: Theorem 4, second display (pp. 19–20)

In the ℓ2\ell_2ℓ2​ set-up, under the strength property with μ>0\mu>0μ>0 and for fixed γ>0\gamma>0γ>0, for every δ>0\delta>0δ>0, with probability at least 1−δ1-\delta1−δ over an i.i.d. sample of size nnn, for all f∈Hf\in\mathcal Hf∈H:

RSPO(f)≤R^SPOγ(f)+(22ρ2(C)+22μ ωS(C)γμ)Rn(H)+ωS(C)log⁡(1/δ)2n.R_{\rm SPO}(f)\le\hat R^\gamma_{\rm SPO}(f)+\Big(\frac{2\sqrt2\rho_2(\mathcal C)+2\sqrt2\mu\,\omega_S(\mathcal C)}{\gamma\mu}\Big)\mathfrak R^n(\mathcal H)+\omega_S(\mathcal C)\sqrt{\frac{\log(1/\delta)}{2n}} .RSPO​(f)≤R^SPOγ​(f)+(γμ22​ρ2​(C)+22​μωS​(C)​)Rn(H)+ωS​(C)2nlog(1/δ)​​.

Milestones

  1. Theorem 3(a): ∥w∗(c^1)−w∗(c^2)∥≤∥c^1−c^2∥∗μmin⁡{νS(c^1),νS(c^2)}\|w^*(\hat c_1)-w^*(\hat c_2)\|\le\frac{\|\hat c_1-\hat c_2\|_*}{\mu\min\{\nu_S(\hat c_1),\nu_S(\hat c_2)\}}∥w∗(c^1​)−w∗(c^2​)∥≤μmin{νS​(c^1​),νS​(c^2​)}∥c^1​−c^2​∥∗​​.
  2. Theorem 3(b): the same Lipschitz-like bound for ℓSPO(⋅,c)\ell_{\rm SPO}(\cdot,c)ℓSPO​(⋅,c), with an extra factor ∥c∥∗\|c\|_*∥c∥∗​.
  3. Theorem 3(c): ℓSPOγ(⋅,c)\ell^\gamma_{\rm SPO}(\cdot,c)ℓSPOγ​(⋅,c) is ∥c∥∗+μ ωS(c)γμ\frac{\|c\|_*+\mu\,\omega_S(c)}{\gamma\mu}γμ∥c∥∗​+μωS​(c)​-Lipschitz for the dual norm.
  4. Eq. (7) with C=2C=\sqrt2C=2​ (Maurer's vector contraction inequality): for LLL-Lipschitz Φi\Phi_iΦi​ on Euclidean Rd\mathbb R^dRd,
Eσ[sup⁡f∈H1n∑iσiΦi(f(xi))]≤2L R^n(H).\mathbb E_\sigma\Big[\sup_{f\in\mathcal H}\frac1n\sum_i\sigma_i\Phi_i(f(x_i))\Big]\le\sqrt2L\,\hat{\mathfrak R}^n(\mathcal H).Eσ​[f∈Hsup​n1​i∑​σi​Φi​(f(xi​))]≤2​LR^n(H).
  1. Theorem 4, first display: for any fixed sample with costs in C\mathcal CC,
R^γSPOn(H)≤(2ρ2(C)+2μ ωS(C)γμ)R^n(H).\hat{\mathfrak R}^n_{\gamma\rm SPO}(\mathcal H)\le\Big(\frac{\sqrt2\rho_2(\mathcal C)+\sqrt2\mu\,\omega_S(\mathcal C)}{\gamma\mu}\Big)\hat{\mathfrak R}^n(\mathcal H).R^γSPOn​(H)≤(γμ2​ρ2​(C)+2​μωS​(C)​)R^n(H).

Theorem 3 is stated for a general norm, as in the paper. Eq. (7), Theorem 4 and the goal are Euclidean. The paper's Theorem 5 (p. 20), a version of the goal uniform over γ∈(0,γˉ]\gamma\in(0,\bar\gamma]γ∈(0,γˉ​], is not part of this mission.

Significance

The bound replaces the loss-class complexity of the SPO loss, which is controlled only through combinatorial dimensions (Natarajan dimension in the polyhedral case, Section 3 of the paper), by the multivariate Rademacher complexity of H\mathcal HH itself. For norm-bounded linear hypothesis classes this complexity has mild, even logarithmic, dependence on the dimensions ppp and ddd (Section 4.4). The result applies to every feasible region with the strength property. By Section 5 of the paper these include strongly convex sets and polytopes, where νS\nu_SνS​ can also be computed. When most predictions stay far from degeneracy, R^SPOγ≈R^SPO\hat R^\gamma_{\rm SPO}\approx\hat R_{\rm SPO}R^SPOγ​≈R^SPO​ and the bound is much sharper than the combinatorial one. It is also a strict generalization of margin bounds for binary classification (Example 7).

The theorem is proved in the paper, which imports two external tools without proof: the Rademacher generalization bound of Bartlett and Mendelson, applied to the margin loss, and Maurer's inequality. To our knowledge none of these results has a machine-checked proof. The mission produces a checked proof of the margin bound and a Lean statement of Maurer's inequality. It also formalizes the strength property and the Lipschitz estimates of Theorem 3, which the companion missions on strongly convex sets and polytopes rely on.

Difficulty

The SPO loss is discontinuous in c^\hat cc^ at degenerate predictions. The standard route, scalar Ledoux–Talagrand contraction applied to the loss class, therefore fails at the first step. It would fail even for a Lipschitz loss, because it relates the loss class only to a scalar class, and H\mathcal HH is vector valued. Lipschitz continuity of the margin loss needs the oracle to be stable away from C∘\mathcal C^\circC∘. Convexity and compactness of SSS alone do not give that: for an ℓp\ell_pℓp​ ball with 2<p<∞2<p<\infty2<p<∞ the strength property fails for every μ>0\mu>0μ>0 (p. 14). The vector contraction inequality of Maurer (2016) is a nontrivial probabilistic inequality, and its constant 2\sqrt22​ must not depend on the dimension ddd. The final concentration step is McDiarmid's inequality for a supremum over a possibly uncountable class, which in a formal proof needs measurability of that supremum.

Formalization scope

The decision space is a finite-dimensional real normed space E. Cost vectors and predictions are continuous linear functionals, StrongDual ℝ E, whose operator norm is the paper's dual norm. In the ℓ2\ell_2ℓ2​ statements E = EuclideanSpace ℝ (Fin d), where the dual norm is Euclidean. Every statement carries the standing assumptions: SSS nonempty, compact, convex and not a singleton, an arbitrary oracle (no tie-breaking rule), and μ>0\mu>0μ>0, γ>0\gamma>0γ>0. The Lipschitz-like bounds of Theorem 3(a)–(b) are stated multiplied out, because the paper reads 1/01/01/0 as +∞+\infty+∞. Expectations over signs are finite averages over sign patterns. ωS(C)\omega_S(\mathcal C)ωS​(C) and ρ2(C)\rho_2(\mathcal C)ρ2​(C) are suprema over a nonempty bounded C\mathcal CC containing the cost almost surely. "With probability at least 1−δ1-\delta1−δ" is the statement that the outer Dn\mathcal D^nDn-measure of the failure event is at most δ\deltaδ.

Added hypotheses, all disclosed in the statements: the multivariate Rademacher sums are bounded above (almost surely in the goal) and R^n(H)\hat{\mathfrak R}^n(\mathcal H)R^n(H) is integrable, since otherwise Lean's junk value 000 would replace an infinite complexity and make the bound false rather than vacuous. Hypotheses fff and ℓSPO(f(x),c)\ell_{\rm SPO}(f(x),c)ℓSPO​(f(x),c) measurable, and the uniform deviation and margin Rademacher suprema a.e.-measurable, are also added; the paper is silent on measurability. A singleton SSS would make C∘\mathcal C^\circC∘ empty and the strength property hold for free; this is excluded explicitly, so the strength property is not vacuous.

A complete development needs the Bartlett–Mendelson symmetrization bound for bounded losses, McDiarmid's inequality, Maurer's inequality, and the Lipschitz and distance-to-degeneracy facts of Section 4.1. Maurer's inequality and the multivariate Rademacher complexity are reusable across vector-valued learning theory. Proofs of any milestone, and of Maurer's inequality in particular, are welcome.

Selected references

  • O. El Balghiti, A. N. Elmachtoub, P. Grigas, A. Tewari, Generalization Bounds in the Predict-then-Optimize Framework, Mathematics of Operations Research, 2023; preprint arXiv:1905.11488v3, 2022. https://arxiv.org/abs/1905.11488
  • A. N. Elmachtoub, P. Grigas, Smart "Predict, then Optimize", Management Science 68(1), 2022. https://doi.org/10.1287/mnsc.2020.3922
  • A. Maurer, A Vector-Contraction Inequality for Rademacher Complexities, Algorithmic Learning Theory (ALT), 2016. https://arxiv.org/abs/1605.00251
  • P. L. Bartlett, S. Mendelson, Rademacher and Gaussian Complexities: Risk Bounds and Structural Results, Journal of Machine Learning Research 3, 2002. https://www.jmlr.org/papers/v3/bartlett02a.html
  • V. Koltchinskii, D. Panchenko, Empirical Margin Distributions and Bounding the Generalization Error of Combined Classifiers, Annals of Statistics 30(1), 2002. https://doi.org/10.1214/aos/1015362183
9 thms3 active usersReviewed
Machine LearningProbability·Captain: mikedeng1

Support Vector Machines VI: An Oracle Inequality for Classifying with Support Vector MachinesTextbook

Motivation

A support vector machine for classification is trained by minimizing a regularized hinge-loss objective — never the classification loss itself, which is non-convex and computationally intractable to minimize. Every earlier mission in this series supplies one piece of the argument that this substitution is nonetheless justified: 01-loss-functions shows the excess hinge risk controls the excess classification risk (Zhang's inequality); 05-concentration supplies a Hilbert-space concentration inequality; and Chapter 6 of the book (not itself a mission in this series, but cited here) combines concentration with a stability argument to bound how far the empirical SVM solution's regularized hinge risk can be from its population minimum. Steinwart & Christmann, Support Vector Machines (Springer 2008, Information Science and Statistics), Chapter 8, assembles exactly these three pieces into Theorem 8.1: an explicit, finite-sample, non-asymptotic bound on how close an SVM classifier's classification risk gets to the Bayes risk — the payoff result the whole apparatus of Chapters 2, 5 and 6 was built to deliver.

Setting

Fix a measurable space XXX and Y:={−1,1}Y := \{-1,1\}Y:={−1,1}. A loss L:X×Y×R→[0,∞)L : X \times Y \times \mathbb R \to [0,\infty)L:X×Y×R→[0,∞), a distribution PPP on X×YX \times YX×Y, the LLL-risk RL,P(f):=∫L(x,y,f(x)) dP(x,y)R_{L,P}(f) := \int L(x,y,f(x)) \,dP(x,y)RL,P​(f):=∫L(x,y,f(x))dP(x,y), and the Bayes risk RL,P∗:=inf⁡fRL,P(f)R^*_{L,P} := \inf_f R_{L,P}(f)RL,P∗​:=inff​RL,P​(f) are exactly as in 01-loss-functions, restated locally here. The hinge loss is Lhinge(y,t):=max⁡{0,1−yt}L_{\mathrm{hinge}}(y,t) := \max\{0,1-yt\}Lhinge​(y,t):=max{0,1−yt} and the classification loss is Lclass(y,t):=1(−∞,0](y⋅sgn⁡t)L_{\mathrm{class}}(y,t) := \mathbf 1_{(-\infty,0]}(y\cdot\operatorname{sgn} t)Lclass​(y,t):=1(−∞,0]​(y⋅sgnt).

Let HHH be a reproducing kernel Hilbert space (RKHS) of a kernel kkk over XXX, i.e. a Hilbert space of functions X→RX \to \mathbb RX→R in which point evaluation is represented by an inner product against a feature map x↦kx∈Hx \mapsto k_x \in Hx↦kx​∈H with k(x,x′)=⟨kx,kx′⟩Hk(x,x') = \langle k_x,k_{x'}\rangle_Hk(x,x′)=⟨kx​,kx′​⟩H​. Write ∥k∥∞:=sup⁡xk(x,x)\|k\|_\infty := \sup_x \sqrt{k(x,x)}∥k∥∞​:=supx​k(x,x)​ for the kernel's sup-bound. For a sample D:=((x1,y1),…,(xn,yn))∈(X×Y)nD := ((x_1,y_1),\dots,(x_n,y_n)) \in (X\times Y)^nD:=((x1​,y1​),…,(xn​,yn​))∈(X×Y)n, the empirical risk is RL,D(f):=1n∑iL(xi,yi,f(xi))R_{L,D}(f) := \tfrac1n \sum_i L(x_i,y_i,f(x_i))RL,D​(f):=n1​∑i​L(xi​,yi​,f(xi​)), and the SVM decision function fD,λf_{D,\lambda}fD,λ​ is the minimizer over HHH of g↦λ∥g∥H2+RL,D(g)g \mapsto \lambda\|g\|_H^2 + R_{L,D}(g)g↦λ∥g∥H2​+RL,D​(g) — the regularized empirical risk minimizer a practical SVM solver computes. The restricted Bayes risk on HHH is RL,P,H∗:=inf⁡f∈HRL,P(f)R^*_{L,P,H} := \inf_{f\in H} R_{L,P}(f)RL,P,H∗​:=inff∈H​RL,P​(f), and the approximation error function is A2(λ):=inf⁡f∈Hλ∥f∥H2+RL,P(f)−RL,P,H∗A_2(\lambda) := \inf_{f\in H} \lambda\|f\|_H^2 + R_{L,P}(f) - R^*_{L,P,H}A2​(λ):=inff∈H​λ∥f∥H2​+RL,P​(f)−RL,P,H∗​: the price, in excess risk, of restricting attention to HHH at regularization strength λ\lambdaλ.

Formalization targets

Goal: Theorem 8.1 — oracle inequality for classifying with SVMs

RLclass,P(fD,λ)−RLclass,P∗<A2(λ)+λ−1(8τn+4n+8τ3n)R_{L_{\mathrm{class}},P}(f_{D,\lambda}) - R^*_{L_{\mathrm{class}},P} < A_2(\lambda) + \lambda^{-1}\left(\sqrt{\tfrac{8\tau}{n}} + \sqrt{\tfrac{4}{n} + \tfrac{8\tau}{3n}}\right)RLclass​,P​(fD,λ​)−RLclass​,P∗​<A2​(λ)+λ−1(n8τ​​+n4​+3n8τ​​)

with PnP^nPn-probability at least 1−e−τ1-e^{-\tau}1−e−τ, for the hinge loss, HHH a separable RKHS with ∥k∥∞≤1\|k\|_\infty \le 1∥k∥∞​≤1, and PPP such that HHH is dense in L1(PX)L^1(P_X)L1(PX​). The bound is finite-sample (valid for every fixed nnn, not just asymptotically) and fully explicit: no unspecified constants beyond A2(λ)A_2(\lambda)A2​(λ) itself, which is a genuine, computable-in-principle quantity depending on HHH, PPP and λ\lambdaλ, not a placeholder. Making the right-hand side small — e.g. letting λ→0\lambda \to 0λ→0 slowly as n→∞n\to\inftyn→∞ — is exactly what proves an SVM classifier consistent for the classification risk, even though it never optimizes that risk directly.

Three milestones, each the specific instance of an earlier chapter's result that this proof invokes (attack order):

  1. Theorem 6.24 instance (hinge loss): λ∥fD,λ∥H2+RL,P(fD,λ)−RL,P,H∗<A2(λ)+λ−1(8τ/n+4/n+8τ/(3n))\lambda\|f_{D,\lambda}\|_H^2 + R_{L,P}(f_{D,\lambda}) - R^*_{L,P,H} < A_2(\lambda) + \lambda^{-1}(\sqrt{8\tau/n}+\sqrt{4/n+8\tau/(3n)})λ∥fD,λ​∥H2​+RL,P​(fD,λ​)−RL,P,H∗​<A2​(λ)+λ−1(8τ/n​+4/n+8τ/(3n)​) with PnP^nPn-probability at least 1−e−τ1-e^{-\tau}1−e−τ — the general oracle inequality for regularized SVMs (Chapter 6, not itself a mission of this series), specialized to the hinge loss, whose global Lipschitz constant 111 collapses the general theorem's Lipschitz-constant factor away.
  2. Theorem 5.31 instance: RLhinge,P,H∗=RLhinge,P∗R^*_{L_{\mathrm{hinge}},P,H} = R^*_{L_{\mathrm{hinge}},P}RLhinge​,P,H∗​=RLhinge​,P∗​ — the RKHS's restricted Bayes hinge risk equals the unrestricted one, using HHH's density in L1(PX)L^1(P_X)L1(PX​) and the fact (Lemma 2.25 v)) that the hinge loss is automatically a PPP-integrable Nemitski loss.
  3. Theorem 2.31 instance (Zhang's inequality, second clause): RLclass,P(f)−RLclass,P∗≤RLhinge,P(f)−RLhinge,P∗R_{L_{\mathrm{class}},P}(f) - R^*_{L_{\mathrm{class}},P} \le R_{L_{\mathrm{hinge}},P}(f) - R^*_{L_{\mathrm{hinge}},P}RLclass​,P​(f)−RLclass​,P∗​≤RLhinge​,P​(f)−RLhinge​,P∗​ for every measurable fff with finite hinge and classification risk — this series' own 01-loss-functions mission's zhang_inequality, second assertion, restated locally.

Chaining these three (with milestone 2 used to rewrite milestone 1's RL,P,H∗R^*_{L,P,H}RL,P,H∗​ as RL,P∗R^*_{L,P}RL,P∗​, then milestone 3 applied to f=fD,λf=f_{D,\lambda}f=fD,λ​) is exactly the book's four-line proof of Theorem 8.1.

Significance

Theorem 8.1 is this series' capstone: every other chapter's result (loss calibration, RKHS theory, representer theorem, Hilbert-space concentration, the general SVM oracle inequality) is a prerequisite this theorem consumes, and nothing later in the book depends on formalizing it further to be meaningful in its own right — it is already a complete, citable, explicit consistency-and-rate statement for SVM classification. It is also the first result in this series whose statement combines three distinct chapters' machinery into a single inequality, making the "restate the specific instance, not the general machinery" discipline (Hard Rule 9) most visibly load-bearing here: none of Theorem 6.24, Theorem 5.31 or Theorem 2.31 in their full generality is needed, only the narrow slice each contributes to this one proof.

No machine-checked formalization of an SVM classification oracle inequality of this kind is known to exist in a public Lean/Mathlib development (see prior-art search below): statistical learning theory results of this shape (finite-sample high-probability bounds combining regularization, approximation error and concentration) are largely unformalized outside isolated concentration inequalities.

Difficulty

The difficulty here is compositional rather than computational: each of the three milestones is, in its own chapter, a short consequence of substantial earlier machinery (Theorem 6.24 rests on a stability argument plus Hilbert-space Hoeffding; Theorem 5.31 rests on continuity of the risk functional on LpL^pLp; Theorem 2.31 rests on a pointwise case analysis), but none of that earlier machinery is re-derived here — only the specific numerical instance each milestone hands to Theorem 8.1's proof. Getting the three instances to compose correctly (in particular, making sure milestone 1's restrictedBayesRisk and milestone 2's equality target the identical quantity, so the substitution the book's proof performs is literally available) is the main formalization risk, not any single proof step.

The probabilistic statement itself is genuinely over the product measure PnP^nPn on samples of size nnn, not an expectation or almost-sure claim, and the bounded-kernel hypothesis ∥k∥∞≤1\|k\|_\infty\le1∥k∥∞​≤1 is load-bearing (it is what fixes the "888" and "444" constants exactly, not just up to a normalization).

Formalization scope

XXX is an arbitrary measurable space; HHH is a general real Hilbert space (NormedAddCommGroup H, InnerProductSpace ℝ H, CompleteSpace H), not specialized to a concrete function space, matching the book's own generality. IsRKHSOfKernel, risk/bayesRisk, classLoss/hingeLoss and empiricalRisk are restated locally in this mission's own Classification sub-namespace — per Hard Rule 9, a draft mission cannot import another draft's definitions, so these duplicate (with identical mathematical content) definitions already drafted in 01-loss-functions and 04-representer. IsSVMSolution encodes "fD,λf_{D,\lambda}fD,λ​ minimizes the regularized empirical risk over HHH" directly as a hypothesis rather than re-deriving existence and uniqueness (04-representer's territory). DenseInL1 renders "HHH dense in L1(PX)L^1(P_X)L1(PX​)" as an ε\varepsilonε-approximation property in the L1L^1L1 seminorm rather than via the Lp subtype, to keep the statement self-contained without importing Chapter 5's own Lp-space apparatus. ∥k∥∞≤1\|k\|_\infty \le 1∥k∥∞​≤1 is ∀ x, k x x ≤ 1 (since ∥k∥∞:=sup⁡xk(x,x)\|k\|_\infty := \sup_x\sqrt{k(x,x)}∥k∥∞​:=supx​k(x,x)​, Eq. (4.15)). "With PnP^nPn-probability at least 1−e−τ1-e^{-\tau}1−e−τ" is stated as a lower bound on (Measure.pi (fun _ : Fin n => P)).real {D | ...}, the nnn-fold product measure of the event.

Theorem 8.2 (Classification with benign kernels), the polynomially-decaying-entropy-number specialization of Theorem 8.1 stated immediately after it in the book, is deliberately out of scope for this mission: it requires entropy-number and covering-number machinery (dyadic entropy numbers ei(id:H→C(X))e_i(\mathrm{id}: H\to C(X))ei​(id:H→C(X)), Lemma 6.21's covering-number bound) that none of this mission's three milestones need, and formalizing it faithfully would roughly double the mission's scope for a result that is a refinement, not a prerequisite, of Theorem 8.1. A trivializing formalization of the goal would state the conclusion for an unconstrained fSVM : (Fin n → X × ℝ) → H with no connection to L, D or λ (making the bound a tautology about whatever function is supplied, independent of what an SVM actually computes); this is ruled out here by requiring hfSVM : ∀ D, IsSVMSolution H toFun hingeLoss lam n D (fSVM D), which pins fSVM D to be an actual minimizer of the regularized empirical hinge risk for that specific sample D.

Selected references

  • I. Steinwart & A. Christmann, Support Vector Machines, Springer, Information Science and Statistics, 2008. https://doi.org/10.1007/978-0-387-77242-4 (Chapter 8, §8.1, pp. 287-291; Chapter 6, §6.4, pp. 223-225; Chapter 5, §5.4-5.5, pp. 179, 190-191; Chapter 2, §2.3, p. 37).
  • T. Zhang, "Statistical behavior and consistency of classification methods based on convex risk minimization," Annals of Statistics 32(1), 2004, pp. 56-85. https://doi.org/10.1214/aos/1079120130
  • This series' 01-loss-functions mission (Theorem 2.31, full statement and proof) and 04-representer mission (Chapter 5's RKHS and SVM-solution machinery, in full generality).
7 thms3 active usersReviewed
Convex OptimizationMachine LearningOptimization+1·Captain: mikedeng1

Variance-based Regularization with Convex Objectives IV: Fast Rates for Approximate Robust Minimizers under a Growth ConditionResearch Paper

Motivation

In stochastic optimization and statistical learning one chooses a parameter θ\thetaθ from a set Θ⊆Rd\Theta\subseteq\mathbb R^dΘ⊆Rd to make the risk R(θ)=EP[ℓ(θ;X)]R(\theta)=\mathbb E_P[\ell(\theta;X)]R(θ)=EP​[ℓ(θ;X)] small, having seen only a sample X1,…,XnX_1,\dots,X_nX1​,…,Xn​ from PPP. Generalization bounds suggest trading empirical risk against its standard deviation, but the variance-penalized objective is non-convex even for convex losses. Duchi and Namkoong (arXiv:1610.02581v3) replace it by the robustly regularized risk, the worst-case expected loss over a χ2\chi^2χ2-divergence ball around the empirical distribution. This objective is convex whenever ℓ\ellℓ is, and it agrees with the variance-penalized objective up to a small error.

When the risk has curvature near its minimizers, empirical risk minimization attains rates faster than 1/n1/\sqrt n1/n​ (Bartlett, Bousquet and Mendelson 2005; Shapiro, Dentcheva and Ruszczyński 2009). Section 4.1 of the paper asks whether minimizers of the robust risk, which carry an extra variance-dependent penalty of order ρ/n\sqrt{\rho/n}ρ/n​, keep these fast rates. Its Theorem 5 answers yes, and does so for approximate minimizers, which is what iterative solvers return.

Setting

A loss ℓ:Rd×X→R\ell:\mathbb R^d\times\mathcal X\to\mathbb Rℓ:Rd×X→R is fixed, with ℓ(⋅;x)\ell(\cdot;x)ℓ(⋅;x) convex and LLL-Lipschitz on a convex set Θ\ThetaΘ for every xxx, and ℓ(θ;⋅)\ell(\theta;\cdot)ℓ(θ;⋅) integrable. The risk is R(θ)=EP[ℓ(θ;X)]R(\theta)=\mathbb E_P[\ell(\theta;X)]R(θ)=EP​[ℓ(θ;X)].

For a radius ρ≥0\rho\ge0ρ≥0, the χ2\chi^2χ2 ball around the empirical distribution P^n\widehat P_nPn​ is the set of weight vectors

Pn={p∈R+n:12∥np−1∥22≤ρ, ⟨1,p⟩=1},\mathcal P_n=\Big\{p\in\mathbb R^n_+:\tfrac12\|np-\mathbf 1\|_2^2\le\rho,\ \langle\mathbf 1,p\rangle=1\Big\},Pn​={p∈R+n​:21​∥np−1∥22​≤ρ, ⟨1,p⟩=1},

and the robust risk is Rn(θ,Pn)=sup⁡p∈Pn∑ipi ℓ(θ;Xi)R_n(\theta,\mathcal P_n)=\sup_{p\in\mathcal P_n}\sum_i p_i\,\ell(\theta;X_i)Rn​(θ,Pn​)=supp∈Pn​​∑i​pi​ℓ(θ;Xi​).

For ϵ≥0\epsilon\ge0ϵ≥0 the ϵ\epsilonϵ-suboptimal sets of the risk and of the robust risk are

S⋆ϵ={θ∈Θ:R(θ)≤inf⁡ΘR+ϵ},S^⋆ϵ={θ∈Θ:Rn(θ,Pn)≤inf⁡ΘRn(⋅,Pn)+ϵ},S_\star^\epsilon=\{\theta\in\Theta:R(\theta)\le\inf_\Theta R+\epsilon\},\qquad\widehat S_\star^\epsilon=\{\theta\in\Theta:R_n(\theta,\mathcal P_n)\le\inf_\Theta R_n(\cdot,\mathcal P_n)+\epsilon\},S⋆ϵ​={θ∈Θ:R(θ)≤Θinf​R+ϵ},S⋆ϵ​={θ∈Θ:Rn​(θ,Pn​)≤Θinf​Rn​(⋅,Pn​)+ϵ},

with S⋆=S⋆0S_\star=S_\star^0S⋆​=S⋆0​ the solution set and πS⋆\pi_{S_\star}πS⋆​​ the Euclidean projection onto it. The risk satisfies a growth condition of order γ>1\gamma>1γ>1 if, for some λ>0\lambda>0λ>0 and r>0r>0r>0,

R(θ)−inf⁡ΘR ≥ λ dist(θ,S⋆)γwhenever dist(θ,S⋆)≤r.(26)R(\theta)-\inf_\Theta R\ \ge\ \lambda\,\mathrm{dist}(\theta,S_\star)^\gamma\quad\text{whenever }\mathrm{dist}(\theta,S_\star)\le r.\tag{26}R(θ)−Θinf​R ≥ λdist(θ,S⋆​)γwhenever dist(θ,S⋆​)≤r.(26)

The complexity of the problem enters through the localized class {x↦ℓ(θ;x)−ℓ(πS⋆(θ);x):θ∈A}\{x\mapsto\ell(\theta;x)-\ell(\pi_{S_\star}(\theta);x):\theta\in A\}{x↦ℓ(θ;x)−ℓ(πS⋆​​(θ);x):θ∈A} and its empirical Rademacher complexity Rn(A)=Eε[sup⁡θ∈A1n∑iεi(ℓ(θ;Xi)−ℓ(πS⋆(θ);Xi))]\mathfrak R_n(A)=\mathbb E_\varepsilon\big[\sup_{\theta\in A}\frac1n\sum_i\varepsilon_i(\ell(\theta;X_i)-\ell(\pi_{S_\star}(\theta);X_i))\big]Rn​(A)=Eε​[supθ∈A​n1​∑i​εi​(ℓ(θ;Xi​)−ℓ(πS⋆​​(θ);Xi​))], with independent uniform signs εi∈{±1}\varepsilon_i\in\{\pm1\}εi​∈{±1}.

Formalization targets

Goal: Theorem 5 (p. 19)

For t>0t>0t>0, ρ≥0\rho\ge0ρ≥0, and 0<ϵ≤12λrγ0<\epsilon\le\frac12\lambda r^\gamma0<ϵ≤21​λrγ satisfying

ϵ≥(28γLγλ)1γ−1(ρn)γ2(γ−1)andϵ2≥2 E[Rn(S⋆2ϵ)]+L(2ϵλ)1γ2tn,(27)\epsilon\ge\Big(2\frac{8^\gamma L^\gamma}{\lambda}\Big)^{\frac1{\gamma-1}}\Big(\frac\rho n\Big)^{\frac\gamma{2(\gamma-1)}}\quad\text{and}\quad\frac\epsilon2\ge2\,\mathbb E[\mathfrak R_n(S_\star^{2\epsilon})]+L\Big(\frac{2\epsilon}\lambda\Big)^{\frac1\gamma}\sqrt{\frac{2t}n},\tag{27}ϵ≥(2λ8γLγ​)γ−11​(nρ​)2(γ−1)γ​and2ϵ​≥2E[Rn​(S⋆2ϵ​)]+L(λ2ϵ​)γ1​n2t​​,(27) P(S^⋆ϵ⊂S⋆2ϵ) ≥ 1−e−t.\mathbb P\big(\widehat S_\star^\epsilon\subset S_\star^{2\epsilon}\big)\ \ge\ 1-e^{-t}.P(S⋆ϵ​⊂S⋆2ϵ​) ≥ 1−e−t.

Milestones, in attack order

  1. Localization (p. 44). Under (26), S⋆2ϵS_\star^{2\epsilon}S⋆2ϵ​ lies in {θ∈Θ:dist(θ,S⋆)≤(2ϵ/λ)1/γ}\{\theta\in\Theta:\mathrm{dist}(\theta,S_\star)\le(2\epsilon/\lambda)^{1/\gamma}\}{θ∈Θ:dist(θ,S⋆​)≤(2ϵ/λ)1/γ}.
  2. Theorem 1, upper half of (10) (p. 7). sup⁡p∈Pn⟨p,z⟩−zˉ≤2ρsn2/n\sup_{p\in\mathcal P_n}\langle p,z\rangle-\bar z\le\sqrt{2\rho s_n^2/n}supp∈Pn​​⟨p,z⟩−zˉ≤2ρsn2​/n​ for every z∈Rnz\in\mathbb R^nz∈Rn.
  3. Claim E.1 (p. 44). If S^⋆ϵ⊄S⋆2ϵ\widehat S_\star^\epsilon\not\subset S_\star^{2\epsilon}S⋆ϵ​⊂S⋆2ϵ​, the localized deviation Δn\Delta_nΔn​ plus a variance term reaches ϵ\epsilonϵ somewhere on S⋆2ϵS_\star^{2\epsilon}S⋆2ϵ​.
  4. Display (43) (p. 45). P(S^⋆ϵ⊄S⋆2ϵ)≤P(sup⁡S⋆2ϵΔn≥ϵ/2)\mathbb P(\widehat S_\star^\epsilon\not\subset S_\star^{2\epsilon})\le\mathbb P(\sup_{S_\star^{2\epsilon}}\Delta_n\ge\epsilon/2)P(S⋆ϵ​⊂S⋆2ϵ​)≤P(supS⋆2ϵ​​Δn​≥ϵ/2).
  5. Concentration (p. 45). P(sup⁡S⋆2ϵΔn≥2E[Rn(S⋆2ϵ)]+u)≤exp⁡(−nu22L2(λ2ϵ)2/γ)\mathbb P(\sup_{S_\star^{2\epsilon}}\Delta_n\ge2\mathbb E[\mathfrak R_n(S_\star^{2\epsilon})]+u)\le\exp(-\frac{nu^2}{2L^2}(\frac\lambda{2\epsilon})^{2/\gamma})P(supS⋆2ϵ​​Δn​≥2E[Rn​(S⋆2ϵ​)]+u)≤exp(−2L2nu2​(2ϵλ​)2/γ).

Significance

The theorem says that the variance penalty implicit in the robust objective does not cost the fast rates available under curvature. The ρ\rhoρ-dependent condition in (27) is of order (ρ/n)γ/(2(γ−1))(\rho/n)^{\gamma/(2(\gamma-1))}(ρ/n)γ/(2(γ−1)), which for quadratic growth (γ=2\gamma=2γ=2) is ρ/n\rho/nρ/n, the same order as the localized complexity term in typical parametric problems. Corollary 4.1 of the paper derives explicit rates of order dnlog⁡nd+tn+ρn\frac dn\log\frac nd+\frac tn+\frac\rho nnd​logdn​+nt​+nρ​ from it for a unique minimizer. The result applies to ϵ\epsilonϵ-approximate minimizers, so it covers the output of the stochastic-gradient methods used to solve the robust problem.

The result is proved in the paper (Appendix E). None of it is formalized: no statement about growth conditions, localized deviations of a robust objective, or fast rates for robust minimizers is on Prove2Me. A formal proof would check the printed constants, settle the boundary case ϵ=0\epsilon=0ϵ=0 (see below), and produce a localization lemma and a reduction from approximate robust minimizers to empirical processes that apply to other estimators.

Difficulty

The obvious argument fails at two points. First, a uniform deviation bound over all of Θ\ThetaΘ gives only the 1/n1/\sqrt n1/n​ rate: the speed-up comes from localizing to S⋆2ϵS_\star^{2\epsilon}S⋆2ϵ​, which requires transferring the growth condition, assumed only within distance rrr of S⋆S_\starS⋆​, to every 2ϵ2\epsilon2ϵ-suboptimal point by convexity. Second, the robust risk is not an empirical average, so standard comparisons between empirical and population minimizers do not apply. Claim E.1 handles this by moving along the segment from a bad approximate minimizer to its projection, which needs the projection to be preserved along that segment (a normal-cone property of πS⋆\pi_{S_\star}πS⋆​​) and the risk to be continuous there. The robust–empirical gap is then controlled by the variance expansion of Theorem 1. The concentration step needs a bounded-differences inequality for a supremum over an uncountable class, together with symmetrization; neither is in Mathlib in this form.

Formalization scope

Parameters live in EuclideanSpace ℝ (Fin d), so norms, distances and projections are Euclidean. The sample is the coordinate process of the product measure P⊗nP^{\otimes n}P⊗n on Fin n → X, n≥1n\ge1n≥1, and probabilities are measures of sample sets (the outer measure for a set that is not measurable). The χ2\chi^2χ2 ball is the weight-vector form (8). The suboptimal sets are written without infima (R(θ)≤R(θ′)+ϵR(\theta)\le R(\theta')+\epsilonR(θ)≤R(θ′)+ϵ for all θ′∈Θ\theta'\in\Thetaθ′∈Θ). Each supremum "sup⁡≥c\sup\ge csup≥c" is written as "for every δ>0\delta>0δ>0 some θ\thetaθ reaches c−δc-\deltac−δ", so no statement relies on the default value of a real supremum. The Rademacher complexity is the published UnderstandingML.rademacher, and its expectation over the sample is assumed integrable, so that it is the true expectation and not the default value 000 of a Bochner integral. Lipschitz continuity is required on Θ\ThetaΘ, as printed.

Corrections and presuppositions:

  • ϵ>0\epsilon>0ϵ>0. The paper prints 0≤ϵ0\le\epsilon0≤ϵ. At ϵ=0\epsilon=0ϵ=0, ρ=0\rho=0ρ=0, both conditions of (27) hold, yet for ℓ(θ;x)=12(θ−x)2\ell(\theta;x)=\frac12(\theta-x)^2ℓ(θ;x)=21​(θ−x)2 on Θ=[−1,1]\Theta=[-1,1]Θ=[−1,1] with XXX uniform on [−12,12][-\frac12,\frac12][−21​,21​] the robust minimizer is the sample mean, which is almost surely not in S⋆={0}S_\star=\{0\}S⋆​={0}. The proof divides by ϵ\epsilonϵ (p. 45). The goal is stated for ϵ>0\epsilon>0ϵ>0.
  • S⋆S_\starS⋆​ nonempty and closed are assumed. The projection πS⋆\pi_{S_\star}πS⋆​​ presupposes them, and Appendix E calls S⋆S_\starS⋆​ closed.
  • Only the upper half of Theorem 1's (10) is stated; it needs no boundedness of the values.

The constant (2⋅8γLγ/λ)1/(γ−1)\big(2\cdot8^\gamma L^\gamma/\lambda\big)^{1/(\gamma-1)}(2⋅8γLγ/λ)1/(γ−1) is the printed one; the proof uses a smaller one, which the printed condition implies. The hypotheses ϵ>0\epsilon>0ϵ>0, γ>1\gamma>1γ>1 and λ>0\lambda>0λ>0 make every power well defined. A formalization that assumed (26) vacuously, took ϵ=0\epsilon=0ϵ=0, or let the Rademacher term be a non-integrable Bochner integral would trivialize the goal; the statements rule these out.

Infrastructure: Euclidean projection onto closed convex sets and its normal-cone characterization (partly in Mathlib), convexity of integral functionals, McDiarmid's bounded-differences inequality, and symmetrization for suprema of empirical processes. The concentration tools and the localization lemma can be reused beyond this mission. Contributions toward McDiarmid's inequality and symmetrization are especially welcome.

Selected references

  • J. C. Duchi and H. Namkoong, Variance-based regularization with convex objectives, arXiv:1610.02581v3, 2017. https://arxiv.org/abs/1610.02581
  • P. L. Bartlett, O. Bousquet and S. Mendelson, Local Rademacher complexities, Annals of Statistics 33(4), 2005. https://doi.org/10.1214/009053605000000282
  • S. Boucheron, G. Lugosi and P. Massart, Concentration Inequalities: A Nonasymptotic Theory of Independence, Oxford University Press, 2013. https://doi.org/10.1093/acprof:oso/9780199535255.001.0001
  • A. Shapiro, D. Dentcheva and A. Ruszczyński, Lectures on Stochastic Programming: Modeling and Theory, SIAM, 2009. https://doi.org/10.1137/1.9780898718751
  • A. Maurer and M. Pontil, Empirical Bernstein bounds and sample variance penalization, COLT, 2009. https://arxiv.org/abs/0907.3740
12 thms2 active usersReviewed
Operations ResearchProbability·Captain: mikedeng1

Conditional Logit Analysis of Qualitative Choice Behavior 5: The Maximum Likelihood Estimator Exists with Probability Tending to One and Is Consistent and Asymptotically NormalResearch Paper

Motivation

The conditional logit model is the workhorse of discrete choice analysis in econometrics, transportation planning, marketing and revenue management. An individual facing a finite set of alternatives picks alternative iii with probability proportional to eziθe^{z_i\theta}ezi​θ, where ziz_izi​ is a vector of observed attributes and θ\thetaθ an unknown parameter vector. Daniel McFadden's 1974 chapter, Conditional Logit Analysis of Qualitative Choice Behavior, derived this model from a theory of random utility maximization and set out how to estimate θ\thetaθ by maximum likelihood. McFadden received the 2000 Nobel Memorial Prize in Economic Sciences for his development of theory and methods for analyzing discrete choice.

Every confidence interval and hypothesis test computed from a fitted logit model rests on the large-sample theory in §III of that chapter: the maximum likelihood estimator exists with probability tending to one, converges to the true parameter, and is approximately normal with covariance given by the inverse information matrix. This mission formalizes that theory, Lemmas 5 and 6 of the paper, as proved in its Appendix.

Setting

Observations are indexed serially, m=0,1,2,…m = 0, 1, 2, \dotsm=0,1,2,…, as in the paper's Appendix ("Let m be a serial index of trials and repetitions"). Observation mmm offers Jm≥1J_m \ge 1Jm​≥1 alternatives, and alternative iii carries a vector zim∈RKz_{im} \in \mathbb R^Kzim​∈RK of independent variables. For a parameter θ∈RK\theta \in \mathbb R^Kθ∈RK the selection probabilities are

Pim(θ)=ezimθ∑j=1Jmezjmθ,zˉm(θ)=∑iPim(θ) zim.P_{im}(\theta) = \frac{e^{z_{im}\theta}}{\sum_{j=1}^{J_m} e^{z_{jm}\theta}}, \qquad \bar z_m(\theta) = \sum_{i} P_{im}(\theta)\, z_{im}.Pim​(θ)=∑j=1Jm​​ezjm​θezim​θ​,zˉm​(θ)=i∑​Pim​(θ)zim​.

The data are generated at a true parameter θ0\theta^0θ0: the chosen alternatives Y0,Y1,…Y_0, Y_1, \dotsY0​,Y1​,… are independent random variables with Pr⁡(Ym=i)=Pim(θ0)\Pr(Y_m = i) = P_{im}(\theta^0)Pr(Ym​=i)=Pim​(θ0). The log-likelihood of the first qqq observations is Lq(θ)=∑m<qlog⁡PYmm(θ)L^q(\theta) = \sum_{m<q}\log P_{Y_m m}(\theta)Lq(θ)=∑m<q​logPYm​m​(θ). The moment matrix of observation mmm is

Ωm=∑iPim(θ0) (zim−zˉm)(zim−zˉm)′,zˉm=zˉm(θ0).\Omega_m = \sum_{i} P_{im}(\theta^0)\,(z_{im}-\bar z_m)(z_{im}-\bar z_m)', \qquad \bar z_m = \bar z_m(\theta^0).Ωm​=i∑​Pim​(θ0)(zim​−zˉm​)(zim​−zˉm​)′,zˉm​=zˉm​(θ0).

Axiom 7 asks that Jm≤J∗J_m \le J_*Jm​≤J∗​ and ∣zim∣≤M|z_{im}| \le M∣zim​∣≤M uniformly, and that 1q∑m<qΩm\frac1q\sum_{m<q}\Omega_mq1​∑m<q​Ωm​ converge to a positive definite matrix Ω\OmegaΩ. Axiom 6, for a given sample, asks that no nonzero γ\gammaγ satisfy (zjm−zYmm)γ≤0(z_{jm} - z_{Y_m m})\gamma \le 0(zjm​−zYm​m​)γ≤0 for all observed mmm and all jjj. A maximum likelihood estimator θ^q\hat\theta^qθ^q is a measurable choice of a maximizer of LqL^qLq, wherever one exists.

Formalization targets

Goal: Lemma 6

θ^q→Pr⁡θ0andq Ω1/2(θ^q−θ0)→dN(0,IK)(q→∞).\hat\theta^q \xrightarrow{\Pr} \theta^0 \quad\text{and}\quad \sqrt q\,\Omega^{1/2}(\hat\theta^q - \theta^0) \xrightarrow{d} N(0, I_K) \qquad (q \to \infty).θ^qPr​θ0andq​Ω1/2(θ^q−θ0)d​N(0,IK​)(q→∞).

Milestones

  1. Axiom 7 implies Axiom 5 (the full-rank condition) in all sufficiently large samples.
  2. Equation (42): Pim(θ)≥1/(J∗e2M∣θ∣)P_{im}(\theta) \ge 1/(J_* e^{2M|\theta|})Pim​(θ)≥1/(J∗​e2M∣θ∣).
  3. Lemma 5: Pr⁡(Axiom 6 holds and Lq attains its maximum)→1\Pr(\text{Axiom 6 holds and } L^q \text{ attains its maximum}) \to 1Pr(Axiom 6 holds and Lq attains its maximum)→1.
  4. Equation (43): the first three derivatives of log⁡Pim\log P_{im}logPim​ are bounded by 2M2M2M, 4M24M^24M2, 8M38M^38M3.
  5. Equation (46): each score ∇log⁡PYmm(θ0)\nabla\log P_{Y_m m}(\theta^0)∇logPYm​m​(θ0) has mean zero.
  6. Equation (47): each expected Hessian equals −Ωm-\Omega_m−Ωm​.
  7. Consistency of θ^q\hat\theta^qθ^q.
  8. Equation (58): q−1/2 Ω−1/2∑m<q∇log⁡PYmm(θ0)→dN(0,IK)q^{-1/2}\,\Omega^{-1/2}\sum_{m<q}\nabla\log P_{Y_m m}(\theta^0) \xrightarrow{d} N(0, I_K)q−1/2Ω−1/2∑m<q​∇logPYm​m​(θ0)d​N(0,IK​).

Significance

The result. Lemma 6 is what licenses reading θ^q\hat\theta^qθ^q as approximately N(θ0,q−1Ω−1)N(\theta^0, q^{-1}\Omega^{-1})N(θ0,q−1Ω−1), so that the diagonal of the inverse information matrix estimates the sampling variances and q(θ^q−θ0)′Ω(θ^q−θ0)q(\hat\theta^q-\theta^0)'\Omega(\hat\theta^q-\theta^0)q(θ^q−θ0)′Ω(θ^q−θ0) is asymptotically χK2\chi^2_KχK2​. Lemma 5 complements it: in finite samples the likelihood can fail to have a maximum (the observations are then "explained" by a direction γ\gammaγ of Axiom 6), and the lemma shows this failure is asymptotically negligible. The data are not identically distributed (each observation has its own alternatives), so the result is not an instance of the textbook i.i.d. maximum likelihood theorem.

Formalizing it. The results are proved in the paper, in outline. A machine-checked version adds: a complete proof of the existence part (Lemma 5), whose published argument is a sketch by induction over an infinite index set; a precise treatment of the estimator where no maximizer exists; the correction of two misprints in the published proof (the normalization 1/q1/q1/q in (58), which must be 1/q1/\sqrt q1/q​, and a constant in (51)); and a multivariate Lindeberg–Feller central limit theorem for bounded, independent, non-identically distributed vectors, which the proof invokes and which is reusable well beyond this paper. No machine-checked proof of these results is known.

Difficulty

The obvious route, "the log-likelihood is concave, so its maximizer converges", needs a maximizer to exist, and in a finite sample it may not; the estimator is defined only on an event whose probability must first be shown to tend to one. Consistency then needs a uniform law of large numbers for the gradient on a sphere around θ0\theta^0θ0, controlled by the third-derivative bound (43). Asymptotic normality needs a central limit theorem for independent but not identically distributed score vectors, with covariances Ωm\Omega_mΩm​ that converge only on average; the i.i.d. central limit theorem does not apply. Finally the random Hessian at an intermediate point must be shown to converge in probability, which ties the consistency result into the normality argument.

Formalization scope

  • Vectors live in EuclideanSpace ℝ (Fin K); zθz\thetazθ is the inner product, and all norms are Euclidean (footnote 11's sum-of-absolute-values norm is equivalent and gives the same qualitative axiom); derivative bounds use operator norms.
  • The paper's NNN trials with RnR_nRn​ repetitions are the special case of the serial indexing in which consecutive observations repeat their data; the sample size ∑nRn\sum_n R_n∑n​Rn​ is qqq.
  • Axiom 7's limit (27) is taken in its serial form (48), with PPP evaluated at θ0\theta^0θ0.
  • The estimator is any measurable selection that maximizes LqL^qLq whenever LqL^qLq has a maximum, and is unconstrained otherwise. Requiring a maximizer for every sample would be unsatisfiable, since Axiom 6 fails with positive probability, and would make the goal vacuous; this convention rules that out.
  • Consistency is TendstoInMeasure. Asymptotic normality is TendstoInDistribution to a random vector whose law is stdGaussian. Ω1/2\Omega^{1/2}Ω1/2 is the positive semidefinite square root CFC.sqrt.
  • Needed infrastructure: derivatives of log-sum-exp, a law of large numbers for bounded independent vectors, and a multivariate Lindeberg–Feller theorem. Mathlib provides the one-dimensional i.i.d. central limit theorem only. Contributions of these general results as separate theorems are welcome.

Selected references

  • D. McFadden, Conditional logit analysis of qualitative choice behavior, in P. Zarembka (ed.), Frontiers in Econometrics, Academic Press, New York, 1974, pp. 105–142.
  • W. Feller, An Introduction to Probability Theory and Its Applications, Vol. II, Wiley, 1966 (Lindeberg–Feller theorem, pp. 256–258).
  • C. R. Rao, Linear Statistical Inference and Its Applications, Wiley (cited by McFadden as Rao (1968), pp. 347–351, for the asymptotic χ2\chi^2χ2 test).
10 thms2 active usersReviewed
Machine LearningProbability·Captain: mikedeng1

Certified Adversarial Robustness via Randomized Smoothing 2: The Certified ℓ2 Radius Cannot Be EnlargedResearch Paper

Motivation

Neural-network classifiers can be made to change their output by perturbations of the input that are imperceptible to a person. A certified defense is a classifier together with a proof that its prediction at a point xxx does not change for any perturbation δ\deltaδ in a stated set, typically an ℓ2\ell_2ℓ2​ ball ∥δ∥2<R\|\delta\|_2<R∥δ∥2​<R. Randomized smoothing turns an arbitrary base classifier into one with such a certificate by classifying Gaussian-noised copies of the input and returning the most likely class. Cohen, Rosenfeld and Kolter (arXiv:1902.02918v2, ICML 2019) gave the certified radius R=σ2(Φ−1(pA‾)−Φ−1(pB‾))R=\frac{\sigma}{2}\big(\Phi^{-1}(\underline{p_A})-\Phi^{-1}(\overline{p_B})\big)R=2σ​(Φ−1(pA​​)−Φ−1(pB​​)) (their Theorem 1) and showed, in their Theorem 2, that this radius cannot be enlarged when only the two class-probability bounds are known about the base classifier. This mission formalizes Theorem 2. Theorem 1 is the subject of the companion mission of this series.

Earlier certificates for the same smoothed classifier, by Lecuyer et al. (2019) via differential privacy and Li et al. (2018) via Rényi divergence, gave smaller radii. Theorem 2 shows that no further analysis that uses only the class-probability bounds can improve on Theorem 1.

Setting

Inputs live in Rd\mathbb R^dRd with the Euclidean norm ∥⋅∥2\|\cdot\|_2∥⋅∥2​; classes form a set Y\mathcal YY. A base classifier is a map f:Rd→Yf:\mathbb R^d\to\mathcal Yf:Rd→Y with Borel decision regions. For a noise level σ>0\sigma>0σ>0, write N(x,σ2I)\mathcal N(x,\sigma^2I)N(x,σ2I) for the isotropic Gaussian law of x+εx+\varepsilonx+ε with ε∼N(0,σ2I)\varepsilon\sim\mathcal N(0,\sigma^2I)ε∼N(0,σ2I). The class probability of ccc at xxx is P(f(x+ε)=c)\mathbb P(f(x+\varepsilon)=c)P(f(x+ε)=c), and the smoothed classifier is

g(x)=arg⁡max⁡c∈Y P(f(x+ε)=c).g(x)=\arg\max_{c\in\mathcal Y}\ \mathbb P(f(x+\varepsilon)=c).g(x)=argc∈Ymax​ P(f(x+ε)=c).

Let Φ\PhiΦ be the standard Gaussian CDF and Φ−1\Phi^{-1}Φ−1 its inverse on (0,1)(0,1)(0,1). A classifier fff is consistent with the observed class probabilities (6) for a top class cAc_AcA​ and numbers pA‾≥pB‾\underline{p_A}\ge\overline{p_B}pA​​≥pB​​ if

P(f(x+ε)=cA) ≥ pA‾ ≥ pB‾ ≥ max⁡c≠cAP(f(x+ε)=c).\mathbb P(f(x+\varepsilon)=c_A)\ \ge\ \underline{p_A}\ \ge\ \overline{p_B}\ \ge\ \max_{c\ne c_A}\mathbb P(f(x+\varepsilon)=c).P(f(x+ε)=cA​) ≥ pA​​ ≥ pB​​ ≥ c=cA​max​P(f(x+ε)=c).

The certified radius is R=σ2(Φ−1(pA‾)−Φ−1(pB‾))R=\frac{\sigma}{2}\big(\Phi^{-1}(\underline{p_A})-\Phi^{-1}(\overline{p_B})\big)R=2σ​(Φ−1(pA​​)−Φ−1(pB​​)). In Lean these are gaussNoise x σ, classProb f σ x c, IsConsistent f σ x cA pA pB and radius σ pA pB in the namespace Cohen2019.Tight, with Phi and PhiInvReal from the series' shared module Cohen2019.Robust; the half-spaces A={z:δT(z−x)≤σ∥δ∥Φ−1(pA‾)}A=\{z:\delta^T(z-x)\le\sigma\|\delta\|\Phi^{-1}(\underline{p_A})\}A={z:δT(z−x)≤σ∥δ∥Φ−1(pA​​)} and B={z:δT(z−x)≥σ∥δ∥Φ−1(1−pB‾)}B=\{z:\delta^T(z-x)\ge\sigma\|\delta\|\Phi^{-1}(1-\overline{p_B})\}B={z:δT(z−x)≥σ∥δ∥Φ−1(1−pB​​)} of the paper's Appendix A are setA and setB.

Quotations write the paper's underlined lower bound as p̲A and its overlined upper bound as p̄B. The PDF has no printed page numbers; every page cited is the PDF page of arXiv:1902.02918v2.

Formalization targets

Goal: Theorem 2 (corrected)

Assume 0<pB‾≤pA‾<10<\overline{p_B}\le\underline{p_A}<10<pB​​≤pA​​<1, pA‾+pB‾≤1\underline{p_A}+\overline{p_B}\le1pA​​+pB​​≤1, and that some finite set sss of classes other than cAc_AcA​ satisfies 1≤pA‾+∣s∣ pB‾1\le\underline{p_A}+|s|\,\overline{p_B}1≤pA​​+∣s∣pB​​. Then for every δ\deltaδ with ∥δ∥2>R\|\delta\|_2>R∥δ∥2​>R there is a base classifier f∗f^*f∗ consistent with (6) and a class c≠cAc\ne c_Ac=cA​ with

P(f∗(x+δ+ε)=cA) < P(f∗(x+δ+ε)=c),\mathbb P(f^*(x+\delta+\varepsilon)=c_A)\ <\ \mathbb P(f^*(x+\delta+\varepsilon)=c),P(f∗(x+δ+ε)=cA​) < P(f∗(x+δ+ε)=c),

so that g(x+δ)≠cAg(x+\delta)\ne c_Ag(x+δ)=cA​ under any tie-breaking. The classifier may depend on δ\deltaδ.

The class-capacity hypothesis is a correction. As printed, with only pA‾+pB‾≤1\underline{p_A}+\overline{p_B}\le1pA​​+pB​​≤1, the theorem fails for two classes: with Y={cA,cB}\mathcal Y=\{c_A,c_B\}Y={cA​,cB​}, pA‾=0.6\underline{p_A}=0.6pA​​=0.6, pB‾=0.1\overline{p_B}=0.1pB​​=0.1 and σ=∥δ∥2=1\sigma=\|\delta\|_2=1σ=∥δ∥2​=1, one has R≈0.767<1R\approx0.767<1R≈0.767<1, yet every consistent fff gives cAc_AcA​ probability at least 0.90.90.9, and Theorem 1 then certifies radius Φ−1(0.9)≈1.28\Phi^{-1}(0.9)\approx1.28Φ−1(0.9)≈1.28.

Milestones

The milestones are the steps the paper itself states, in its order: the Claims P(X∈A)=pA‾\mathbb P(X\in A)=\underline{p_A}P(X∈A)=pA​​ and P(X∈B)=pB‾\mathbb P(X\in B)=\overline{p_B}P(X∈B)=pB​​ for X∼N(x,σ2I)X\sim\mathcal N(x,\sigma^2I)X∼N(x,σ2I); the disjointness of AAA and BBB (corrected to "null" when pA‾+pB‾=1\underline{p_A}+\overline{p_B}=1pA​​+pB​​=1); equations (13) and (14) for Y∼N(x+δ,σ2I)Y\sim\mathcal N(x+\delta,\sigma^2I)Y∼N(x+δ,σ2I),

P(Y∈A)=Φ(Φ−1(pA‾)−∥δ∥σ),P(Y∈B)=Φ(Φ−1(pB‾)+∥δ∥σ);\mathbb P(Y\in A)=\Phi\Big(\Phi^{-1}(\underline{p_A})-\tfrac{\|\delta\|}{\sigma}\Big),\qquad \mathbb P(Y\in B)=\Phi\Big(\Phi^{-1}(\overline{p_B})+\tfrac{\|\delta\|}{\sigma}\Big);P(Y∈A)=Φ(Φ−1(pA​​)−σ∥δ∥​),P(Y∈B)=Φ(Φ−1(pB​​)+σ∥δ∥​);

the equivalence P(Y∈A)<P(Y∈B)  ⟺  ∥δ∥2>R\mathbb P(Y\in A)<\mathbb P(Y\in B)\iff\|\delta\|_2>RP(Y∈A)<P(Y∈B)⟺∥δ∥2​>R; and the existence of the worst-case classifier f∗f^*f∗ satisfying (6) with equalities.

Significance

Theorem 2 makes the guarantee of Theorem 1 exact: when only (6) is known about fff, the set of perturbations under which the Gaussian-smoothed prediction is provably constant is exactly the open ℓ2\ell_2ℓ2​ ball of radius RRR. It settles that improvements to Gaussian-smoothing certificates must use more information about the base classifier than the two bounds, as later work on higher-order and Lipschitz-based certificates does.

The paper's proof is complete in its main lines and has two gaps that this mission records and repairs: the printed statement omits a condition on the number of classes, and the claim A∩B=∅A\cap B=\emptysetA∩B=∅ fails at pA‾+pB‾=1\underline{p_A}+\overline{p_B}=1pA​​+pB​​=1. To our knowledge neither Theorem 1 nor Theorem 2 has a machine-checked proof. Mathlib at the pinned revision has the multivariate standard Gaussian but no normal quantile function and no Gaussian half-space lemma; this mission adds statements for both kinds of fact.

Difficulty

Each step is elementary on paper but rests on facts about Gaussians that Mathlib does not package: the image of the standard Gaussian on Rd\mathbb R^dRd under a linear functional z↦δTzz\mapsto\delta^T zz↦δTz is the one-dimensional Gaussian with variance ∥δ∥2\|\delta\|^2∥δ∥2, and Φ\PhiΦ is a continuous strictly increasing bijection R→(0,1)\mathbb R\to(0,1)R→(0,1) with Φ−1(1−p)=−Φ−1(p)\Phi^{-1}(1-p)=-\Phi^{-1}(p)Φ−1(1−p)=−Φ−1(p). The construction of f∗f^*f∗ has a further step the paper leaves informal: the region between AAA and BBB, of mass 1−pA‾−pB‾1-\underline{p_A}-\overline{p_B}1−pA​​−pB​​, must be shared among "other classes" with none exceeding pB‾\overline{p_B}pB​​, which is where the capacity hypothesis enters. Measurability of the constructed decision regions must be carried along.

Formalization scope

Rd\mathbb R^dRd is EuclideanSpace ℝ (Fin d). N(x,σ2I)\mathcal N(x,\sigma^2I)N(x,σ2I) is the pushforward of Mathlib's stdGaussian under z↦x+σzz\mapsto x+\sigma zz↦x+σz, with σ>0\sigma>0σ>0 a binder. Φ\PhiΦ is cdf (gaussianReal 0 1); Φ−1(p)\Phi^{-1}(p)Φ−1(p) is the generalized inverse inf⁡{t:p≤Φ(t)}\inf\{t:p\le\Phi(t)\}inf{t:p≤Φ(t)}, which is the true inverse on (0,1)(0,1)(0,1) and the junk value 000 at the endpoints, so every statement that evaluates it assumes 0<p<10<p<10<p<1; at pB‾=0\overline{p_B}=0pB​​=0 or pA‾=1\underline{p_A}=1pA​​=1 the paper's radius is infinite and Theorem 2 is vacuous. Class probabilities are real numbers. The base classifier in the conclusion is deterministic with Borel decision regions, which is the stronger existence statement. The conclusion is the strict inequality between class probabilities, not merely the failure of cAc_AcA​ to be a strict unique argmax.

A formalization in which the junk endpoint value of Φ−1\Phi^{-1}Φ−1 makes RRR negative, or in which the classifier's decision regions are non-measurable so that its class probabilities are default values, would make the goal trivial; the hypotheses above exclude both.

Reusable beyond this mission: the Gaussian half-space probabilities and the normal quantile on (0,1)(0,1)(0,1). Contributions of general Mathlib-style lemmas (the law of δTX\delta^T XδTX for X∼N(x,σ2I)X\sim\mathcal N(x,\sigma^2I)X∼N(x,σ2I), properties of Φ−1\Phi^{-1}Φ−1) are welcome.

Selected references

  • J. M. Cohen, E. Rosenfeld, J. Z. Kolter, Certified Adversarial Robustness via Randomized Smoothing, ICML 2019; arXiv:1902.02918v2. https://arxiv.org/abs/1902.02918v2
  • M. Lecuyer, V. Atlidakis, R. Geambasu, D. Hsu, S. Jana, Certified Robustness to Adversarial Examples with Differential Privacy, IEEE S&P 2019. https://arxiv.org/abs/1802.03471
  • B. Li, C. Chen, W. Wang, L. Carin, Certified Adversarial Robustness with Additive Noise, NeurIPS 2019. https://arxiv.org/abs/1809.03113
  • J. Neyman, E. S. Pearson, On the Problem of the Most Efficient Tests of Statistical Hypotheses, Phil. Trans. R. Soc. A 231, 1933. https://doi.org/10.1098/rsta.1933.0009
11 thms2 active usersReviewed
Probability·Captain: mikedeng1

Weighted Sums of Certain Dependent Random Variables 3: Reversed Weighted Sums of Bounded Martingale Differences Obey a Strong LawResearch Paper

Motivation

A martingale difference sequence is the standard model of a "fair" sequence of observations whose terms may depend on the past: each new term has conditional mean zero given everything observed before it. Laws of large numbers for such sequences underlie the analysis of stochastic approximation, sequential estimation and online learning, where the noise terms are dependent but conditionally centred.

Classical strong laws concern averages in which every observation keeps the same weight as the sample grows. Kazuoki Azuma's 1967 paper Weighted sums of certain dependent random variables (Tôhoku Math. J. 19) studies weighted sums of dependent variables, and is best known for the exponential moment bound that is now called the Azuma inequality (its display (2.4) together with Remark 1). Its Theorem 3 uses that bound to prove a strong law for weighted averages in which the weights are applied in reverse order, so that the oldest observation always receives the newest, largest weight.

Timeline:

  • 1960s: Y. S. Chow (Ann. Math. Statist. 37, 1966) introduces a conditional exponential-moment condition close to Azuma's property [G] in a convergence theorem for independent variables.
  • 1967: Azuma proves the moment bound (2.4) for conditionally sub-Gaussian martingale differences, a law of the iterated logarithm for direct weighted sums (Theorem 2), and the strong law for reversed weighted sums (Theorem 3), the subject of this mission.

Setting

Let (Ω,A,P)(\Omega,\mathfrak A,P)(Ω,A,P) be a probability space and (An)n≥0(\mathfrak A_n)_{n\ge0}(An​)n≥0​ an increasing family of sub-σ\sigmaσ-fields of A\mathfrak AA. A sequence (xn)n≥1(x_n)_{n\ge1}(xn​)n≥1​ of real random variables is a sequence of martingale differences if, for every n≥1n\ge1n≥1, xnx_nxn​ is An\mathfrak A_nAn​-measurable and integrable and E{xn∣An−1}=0E\{x_n\mid\mathfrak A_{n-1}\}=0E{xn​∣An−1​}=0 almost surely. Theorem 3 assumes moreover ∣xn∣≤1|x_n|\le1∣xn​∣≤1 almost surely for every nnn.

Let (an)n≥1(a_n)_{n\ge1}(an​)n≥1​ be positive increasing weights: an>0a_n>0an​>0 and an≤an+1a_n\le a_{n+1}an​≤an+1​. Put

An=a1+a2+⋯+an,Sˉn=anx1+an−1x2+⋯+a1xn=∑j=1nan−j+1xj.A_n = a_1+a_2+\dots+a_n,\qquad \bar S_n = a_nx_1 + a_{n-1}x_2+\dots+a_1x_n=\sum_{j=1}^n a_{n-j+1}x_j .An​=a1​+a2​+⋯+an​,Sˉn​=an​x1​+an−1​x2​+⋯+a1​xn​=j=1∑n​an−j+1​xj​.

The sums Sˉn\bar S_nSˉn​ are the reversed weighted sums. In passing from Sˉn\bar S_nSˉn​ to Sˉn+1\bar S_{n+1}Sˉn+1​ every existing term changes its weight, so (Sˉn)(\bar S_n)(Sˉn​) is in general not a martingale. In the Lean development, AnA_nAn​ is A a n, Sˉn(ω)\bar S_n(\omega)Sˉn​(ω) is Sbar a x n ω, and the martingale-difference property is IsMartingaleDiff μ ℱ x.

Formalization targets

Goal: Theorem 3, (4.9)–(4.10)

If (xn)(x_n)(xn​) is a sequence of martingale differences with ∣xn∣≤1|x_n|\le1∣xn​∣≤1 a.s., (an)(a_n)(an​) is positive and nondecreasing, and

anAn=o(1log⁡log⁡An)(n→∞),(4.9)\frac{a_n}{A_n}=o\Big(\frac{1}{\log\log A_n}\Big)\qquad(n\to\infty),\tag{4.9}An​an​​=o(loglogAn​1​)(n→∞),(4.9)

then

SˉnAn⟶0almost surely.(4.10)\frac{\bar S_n}{A_n}\longrightarrow0\quad\text{almost surely}.\tag{4.10}An​Sˉn​​⟶0almost surely.(4.10)

Milestones

The milestones follow the paper's proof, in attack order.

  1. Remark 1 (p. 358): if ∣xn∣≤Kn|x_n|\le K_n∣xn​∣≤Kn​ a.s., then E{exp⁡(txn)∣An−1}≤cosh⁡(tKn)≤exp⁡(t2Kn2/2)E\{\exp(tx_n)\mid\mathfrak A_{n-1}\}\le\cosh(tK_n)\le\exp(t^2K_n^2/2)E{exp(txn​)∣An−1​}≤cosh(tKn​)≤exp(t2Kn2​/2) a.s.
  2. The tail step of (4.16) (p. 366): for ∣xn∣≤1|x_n|\le1∣xn​∣≤1, real c1,…,cNc_1,\dots,c_Nc1​,…,cN​ with ∑cj2>0\sum c_j^2>0∑cj2​>0 and λ≥0\lambda\ge0λ≥0,
P{∑j=1Ncjxj>λ}≤exp⁡(−λ22∑j=1Ncj2).P\Big\{\sum_{j=1}^Nc_jx_j>\lambda\Big\}\le\exp\Big(-\frac{\lambda^2}{2\sum_{j=1}^Nc_j^2}\Big).P{j=1∑N​cj​xj​>λ}≤exp(−2∑j=1N​cj2​λ2​).
  1. The blocks (4.11)–(4.14) (pp. 364–365): for every ε>0\varepsilon>0ε>0 there are indices n1<n2<⋯n_1<n_2<\cdotsn1​<n2​<⋯ with An1>2(3+ε)/(6+ε)A_{n_1}>2(3+\varepsilon)/(6+\varepsilon)An1​​>2(3+ε)/(6+ε), an/An<ε/(6+ε)a_n/A_n<\varepsilon/(6+\varepsilon)an​/An​<ε/(6+ε) and anlog⁡log⁡An/An<ε2/64a_n\log\log A_n/A_n<\varepsilon^2/64an​loglogAn​/An​<ε2/64 for n>n1n>n_1n>n1​, and Ank−1<Ank≤(1+ε/3)Ank−1<Ank+1A_{n_{k-1}}<A_{n_k}\le(1+\varepsilon/3)A_{n_{k-1}}<A_{n_k+1}Ank−1​​<Ank​​≤(1+ε/3)Ank−1​​<Ank​+1​.
  2. The maximal inequality (4.15) (p. 365): if AN1≤(1+ε/3)AN0A_{N_1}\le(1+\varepsilon/3)A_{N_0}AN1​​≤(1+ε/3)AN0​​ with 1≤N0<N11\le N_0<N_11≤N0​<N1​, then
2P{SˉN1>(ε/2)AN0}≥P{max⁡N0<n≤N1Sˉn>εAN0}.2P\{\bar S_{N_1}>(\varepsilon/2)A_{N_0}\}\ge P\Big\{\max_{N_0<n\le N_1}\bar S_n>\varepsilon A_{N_0}\Big\}.2P{SˉN1​​>(ε/2)AN0​​}≥P{N0​<n≤N1​max​Sˉn​>εAN0​​}.
  1. Block growth (p. 366): (4.11), (4.12) and (4.14) give Ank>(2(3+ε)/(6+ε))k−1A_{n_k}>(2(3+\varepsilon)/(6+\varepsilon))^{k-1}Ank​​>(2(3+ε)/(6+ε))k−1.

Significance

The result. Theorem 3 shows that reversed weighting does not destroy the strong law, under a growth condition on the weights that is strictly weaker than the condition an2/∑j≤naj2→0a_n^2/\sum_{j\le n}a_j^2\to0an2​/∑j≤n​aj2​→0 of the paper's Theorem 2: the paper reproduces an example of T. Tsuchikura satisfying (4.9) but not that condition. The condition allows rapidly growing weights, provided no single weight carries more than an o(1/log⁡log⁡An)o(1/\log\log A_n)o(1/loglogAn​) share of the total. Milestone 2 is the one-sided Azuma inequality in the paper's conditional-expectation form, a tool used throughout probability, combinatorics and learning theory. Milestone 4 is a maximal inequality for a process that is not a martingale, where Doob's inequality cannot be used.

Formalizing it. The theorem is proved in the paper; this mission produces a machine-checked proof. Mathlib contains sub-Gaussian moment-generating-function bounds for martingale differences in a kernel formulation (HasCondSubgaussianMGF, which assumes a standard Borel space), and the Prove2Me platform has two-sided Azuma–Hoeffding inequalities under the same assumption. Neither states Remark 1 or the one-sided tail bound in the paper's conditional-expectation form on an arbitrary probability space, and no strong law for reversed weighted sums is formalized.

Difficulty

The obvious route to a strong law for a martingale, Doob's maximal inequality applied along a geometric subsequence, fails at the first step: (Sˉn)(\bar S_n)(Sˉn​) is not a martingale, because each new step reweights all earlier terms. The maximum of Sˉn\bar S_nSˉn​ over a block of indices therefore needs a separate maximal inequality, and it is there that the monotonicity of the weights is indispensable. A second difficulty is quantitative: the exponential tail bound must be summable over blocks whose growth is controlled only through (4.9), which is weaker than the variance-type condition of Theorem 2, so the block sizes and the constants ε/(6+ε)\varepsilon/(6+\varepsilon)ε/(6+ε), ε2/64\varepsilon^2/64ε2/64 and 1+ε/31+\varepsilon/31+ε/3 have to be chosen to fit together.

Formalization scope

Conventions committed to in Lean:

  • Indices start at 111: sums run over Finset.Icc 1 n, and a0a_0a0​, x0x_0x0​ are never used. Every hypothesis on aaa and xxx is quantified over n≥1n\ge1n≥1.
  • The filtration is a Mathlib Filtration ℕ. Its first σ\sigmaσ-field plays the role of A0\mathfrak A_0A0​ and is arbitrary rather than trivial; the paper's A0={∅,Ω}\mathfrak A_0=\{\emptyset,\Omega\}A0​={∅,Ω} is a special case, so the formal statements are at least as general.
  • "Positive increasing" is read as an>0a_n>0an​>0 and an≤an+1a_n\le a_{n+1}an​≤an+1​ (nondecreasing), the weaker hypothesis.
  • (4.9) is stated literally as a little-ooo relation (IsLittleO along atTop). An→∞A_n\to\inftyAn​→∞ is not a hypothesis, since it follows from positivity and monotonicity.
  • The conclusion is convergence of the real sequence Sˉn(ω)/An\bar S_n(\omega)/A_nSˉn​(ω)/An​ to 000 for almost every ω\omegaω, which contains both the upper and the lower tail; a statement giving only lim sup⁡≤0\limsup\le0limsup≤0 is not the goal. No real-valued limsup is used anywhere.
  • Probabilities are real-valued (μ.real); conditional expectations are Mathlib's μ[f | ℱ n].

A formalization of the goal that replaces Sˉn\bar S_nSˉn​ by the direct sums a1x1+⋯+anxna_1x_1+\dots+a_nx_na1​x1​+⋯+an​xn​, drops the monotonicity of the weights, or strengthens (4.9) to an/An=o(1/log⁡An)a_n/A_n=o(1/\log A_n)an​/An​=o(1/logAn​) or to an2/∑j≤naj2→0a_n^2/\sum_{j\le n}a_j^2\to0an2​/∑j≤n​aj2​→0 states a different theorem and does not count.

A complete development needs the Azuma moment bound in conditional-expectation form, conditional Chebyshev arguments on events, the Borel–Cantelli lemma (in Mathlib) and elementary real analysis of the blocks. The one-sided Azuma inequality and the maximal inequality (4.15) are reusable beyond this mission. Proofs of any milestone are welcome, as are alternative proofs of the goal.

Selected references

  • K. Azuma, Weighted sums of certain dependent random variables, Tôhoku Mathematical Journal 19 (1967), 357–367. https://doi.org/10.2748/tmj/1178243286
  • Y. S. Chow, Some convergence theorems for independent random variables, Annals of Mathematical Statistics 37 (1966), 1482–1493.
  • J. L. Doob, Stochastic Processes, Wiley, New York, 1953.
8 thms2 active usersReviewed
Probability·Captain: mikedeng1

Weighted Sums of Certain Dependent Random Variables 1: Weighted Sums of Bounded Multiplicative Systems Grow at Most Like √(2Bₙ² log n)Research Paper

Motivation

Weighted sums of dependent random variables occur when the weights change with the observation horizon. Even if each random variable is bounded and centered, allowing the weights in row nnn to be chosen anew makes an almost-sure statement about all large nnn different from a bound on a single finite sum. Kazuo Azuma's 1967 paper treats this situation under a finite-product moment condition called class [M]. Its first main theorem bounds every row of an arbitrary real triangular array at the scale given by that row's Euclidean norm and log⁡n\log nlogn.

The condition is useful because it allows dependence. The paper notes that bounded martingale differences provide examples, but Theorem 1 is stated directly for class [M], without introducing a filtration in the result. The conclusion therefore records the property of the random variables actually used by this part of the paper, rather than restricting the mission to one familiar source of examples. Azuma, §1 and Theorem 1.

Setting

Fix a probability space (Ω,A,P)(\Omega,\mathcal A,P)(Ω,A,P). Let x1,x2,…x_1,x_2,\ldotsx1​,x2​,… be real measurable random variables. They form a bounded multiplicative system with unit bounds if ∣xk∣≤1|x_k|\le1∣xk​∣≤1 almost surely for every k≥1k\ge1k≥1, and

E ⁣[∏k∈Sxk]=0for every nonempty finite S⊆{1,2,…}.E\!\left[\prod_{k\in S}x_k\right]=0 \quad\text{for every nonempty finite }S\subseteq\{1,2,\ldots\}.E[k∈S∏​xk​]=0for every nonempty finite S⊆{1,2,…}.

The indices in SSS are distinct. Taking a singleton shows E[xk]=0E[x_k]=0E[xk​]=0; taking sets of two, three, or more indices imposes the full condition used in the paper. Pairwise zero correlations by themselves do not state class [M]. All bounds and moment conditions are for the positive indices, so x0x_0x0​ is outside the mathematical sequence. Azuma, p. 357, property [M].

For each n≥1n\ge1n≥1, choose real coefficients an1,…,anna_{n1},\ldots,a_{nn}an1​,…,ann​. There is no relation required between different rows. Define the weighted sum TnT_nTn​ and its weight norm BnB_nBn​ by

Tn=∑k=1nankxk,Bn=(∑k=1nank2)1/2.T_n=\sum_{k=1}^{n}a_{nk}x_k, \qquad B_n=\left(\sum_{k=1}^{n}a_{nk}^{2}\right)^{1/2}.Tn​=k=1∑n​ank​xk​,Bn​=(k=1∑n​ank2​)1/2.

Both definitions use the source's 1-based indices. A row of zero weights has Bn=0B_n=0Bn​=0 and Tn=0T_n=0Tn​=0 almost surely; such rows remain within the theorem. Azuma, p. 359, §3.

Formalization targets

The goal is Theorem 1, display (3.1): for every bounded multiplicative system and every real triangular array,

lim sup⁡n→∞∣Tn∣2Bn2log⁡n≤1P-almost surely.\limsup_{n\to\infty} \frac{|T_n|}{\sqrt{2B_n^2\log n}}\le1 \qquad P\text{-almost surely}.n→∞limsup​2Bn2​logn​∣Tn​∣​≤1P-almost surely.

The normalizing constant is exactly 222, and the upper bound is exactly 111. The mission does not assume that BnB_nBn​ grows, converges, or stays positive. In Lean, the target is stated as the equivalent operational bound: for each δ>0\delta>0δ>0, almost every outcome eventually satisfies ∣Tn∣≤(1+δ)2Bn2log⁡n|T_n|\le(1+\delta)\sqrt{2B_n^2\log n}∣Tn​∣≤(1+δ)2Bn2​logn​. This includes zero-weight rows without assigning a meaning to a real quotient 0/00/00/0.

The milestone list follows results displayed in the paper: the corrected convexity inequality (2.2), Lemma 1's exponential moment estimate (2.1), the exponential estimate in the proof of Theorem 1 with its factor 222, and the almost-sure finite exponential series on the next page. Lemma 1 gives, for arbitrary real b1,…,bnb_1,\ldots,b_nb1​,…,bn​ and t∈Rt\in\mathbb Rt∈R,

Eexp⁡ ⁣(t∑k=1nbkxk)≤exp⁡ ⁣(t22∑k=1nbk2).E\exp\!\left(t\sum_{k=1}^{n}b_kx_k\right) \le \exp\!\left(\frac{t^2}{2}\sum_{k=1}^{n}b_k^2\right).Eexp(tk=1∑n​bk​xk​)≤exp(2t2​k=1∑n​bk2​).

The paper prints (2.2) with a missing factor bnkb_{nk}bnk​ in its linear term. The mission records the printed text as provenance and states the corrected inequality in Lean; the printed version fails already when bnk=2b_{nk}=2bnk​=2 and xk=t=1x_k=t=1xk​=t=1. Azuma, pp. 357–360.

Significance

Theorem 1 turns an exponential moment bound for each finite weighted sum into a single almost-sure assertion along an entire triangular array. It gives a scale that adapts to the actual coefficients in each row: two arrays with different row norms receive different bounds, while no regularity across rows is required. The result is also the starting point for the weighted strong-law corollaries that follow it in the paper. Azuma, Theorem 1 and Corollary 1.

The mathematical theorem has been proved since 1967. This mission's remaining work is a machine-checked Lean proof of the exact theorem and its listed intermediate statements. A complete development would add reusable formal statements for bounded multiplicative systems and for their finite exponential moments. Those objects could support later work on dependent sums without importing a filtration or a stronger independence assumption. The proposal statements compile as open goals; compilation alone does not supply proofs.

Difficulty

The usual first step for independent bounded variables is to factor the exponential moment into one-variable expectations. Class [M] does not assume independence, so that factorization is unavailable. The condition controls every product with distinct indices, while allowing other dependence. The almost-sure conclusion must also hold when the coefficients change arbitrarily with nnn: bounds that depend on one fixed row do not by themselves settle what happens for all sufficiently large rows. Finally, rows with Bn=0B_n=0Bn​=0 require a statement that preserves the theorem rather than excluding them by an added positivity hypothesis.

Formalization scope

The Lean development represents the probability law by a measure μ\muμ with IsProbabilityMeasure μ, and a random sequence by x:N→Ω→Rx:\mathbb N\to\Omega\to\mathbb Rx:N→Ω→R. It uses measurable variables and states the unit bound almost surely at every positive index. The class [M] predicate quantifies over every nonempty finite set of positive indices; no conditional expectations, filtration, symmetry, or independence hypotheses enter Theorem 1. Finite products and finite weighted sums use ordinary real multiplication and Finset.Icc 1 n. The triangular weights have type N→N→R\mathbb N\to\mathbb N\to\mathbb RN→N→R, and only entries with 1≤k≤n1\le k\le n1≤k≤n contribute.

The norm BnB_nBn​ is the nonnegative real square root of the sum of squared weights. Real.log is zero at n=0n=0n=0 and n=1n=1n=1 in Lean, but the target is eventually quantified, so its asymptotic content concerns large nnn. The exponential estimate is stated for n≥1n\ge1n≥1. If Bn=0B_n=0Bn​=0, Lean's total division returns zero in the exponent's quotient; all weights in that row are zero, making this extension valid. In the almost-sure series, the term at index zero is set to zero. The source's limsup is represented by eventual inequalities for every positive excess, avoiding a real-valued limsup default on unbounded sequences.

The definition of class [M] includes every finite product, including singletons; replacing it by pairwise orthogonality would change the theorem. Measurability and the almost-sure unit bounds ensure that the finite products and the exponential functions in Lemma 1 are integrable, so their Lean integrals represent expectations. Contributions toward proofs of the corrected convexity bound, the moment estimate, the exponential series, and the final almost-sure step are all within scope. The finite-product predicate and the exponential estimate are reusable beyond this mission.

Selected references

  • Kazuo Azuma, Weighted sums of certain dependent random variables, Tôhoku Mathematical Journal 19 (1967), 357–367. DOI: 10.2748/tmj/1178243286.
7 thms2 active usersReviewed
Convex OptimizationMachine LearningProbability·Captain: mikedeng1

Stability and Generalization 4: Relative-Entropy Regularization of Mixtures Has Uniform Stability M²/(λm)Research Paper

Motivation

A learning algorithm generalizes when its error on fresh data is close to its error on the training sample. Bousquet and Elisseeff (JMLR 2, 2002) showed that a single property of the algorithm, uniform stability, controls this gap with exponential concentration: if removing any one example from a training set of size mmm changes the loss of the output at every point by at most β\betaβ, the generalization error exceeds the empirical error by roughly 2β+(4mβ+M)ln⁡(1/δ)/(2m)2\beta + (4m\beta + M)\sqrt{\ln(1/\delta)/(2m)}2β+(4mβ+M)ln(1/δ)/(2m)​ with probability 1−δ1-\delta1−δ (their Theorem 12). The bound is useful only when β=O(1/m)\beta = O(1/m)β=O(1/m), and the second half of the paper identifies algorithms with that rate: Tikhonov regularization in a reproducing kernel Hilbert space (Theorem 22), and relative-entropy regularization of mixtures (Theorem 24), the subject of this mission.

Mixtures arise whenever a learner outputs a distribution over a parametric base class instead of a single hypothesis: Bayesian posterior averaging, Gibbs and randomized classifiers, exponential weights. Regularizing by the relative entropy to a prior is the maximum-a-posteriori reading of these procedures, and Theorem 24 is one of the earliest results showing that such posteriors are uniformly stable with rate 1/(λm)1/(\lambda m)1/(λm). The same mechanism (entropic regularization, stability through Pinsker's inequality) reappears in PAC-Bayesian analysis and in the stability of exponential-weights methods.

Setting

Let Θ\ThetaΘ be a measurable space with a reference measure ν\nuν, and write dθd\thetadθ for integration against ν\nuν. A base class H={hθ:θ∈Θ}\mathcal H = \{h_\theta : \theta \in \Theta\}H={hθ​:θ∈Θ} is indexed by Θ\ThetaΘ, and r(hθ,z)∈[0,M]r(h_\theta, z) \in [0, M]r(hθ​,z)∈[0,M] is the loss of the base hypothesis hθh_\thetahθ​ at an example z∈Zz \in Zz∈Z.

The algorithm outputs a density ggg with respect to ν\nuν: a measurable, nonnegative, integrable g:Θ→Rg : \Theta \to \mathbb Rg:Θ→R with ∫Θg dθ=1\int_\Theta g\,d\theta = 1∫Θ​gdθ=1. FFF denotes the set of all densities. A density is scored by the averaged loss

ℓ(g,z)=∫Θr(hθ,z) g(θ) dθ(28),\ell(g, z) = \int_\Theta r(h_\theta, z)\, g(\theta)\, d\theta \qquad (28),ℓ(g,z)=∫Θ​r(hθ​,z)g(θ)dθ(28),

the expected loss of a randomized predictor that draws hθh_\thetahθ​ from ggg. The relative entropy of ggg to g′g'g′ is

K(g,g′)=∫Θg(θ)ln⁡g(θ)g′(θ) dθ∈[0,∞],K(g, g') = \int_\Theta g(\theta) \ln \frac{g(\theta)}{g'(\theta)}\, d\theta \in [0, \infty],K(g,g′)=∫Θ​g(θ)lng′(θ)g(θ)​dθ∈[0,∞],

with K(g,g′)=+∞K(g, g') = +\inftyK(g,g′)=+∞ when g νg\,\nugν is not absolutely continuous with respect to g′ νg'\,\nug′ν or the integrand is not integrable.

Fix a prior f0∈Ff_0 \in Ff0​∈F, a parameter λ>0\lambda > 0λ>0, and a training set S=(z1,…,zm)S = (z_1, \dots, z_m)S=(z1​,…,zm​). The algorithm returns a minimizer over FFF of

Rr(g)=1m∑j=1mℓ(g,zj)+λK(g,f0)(29).R_r(g) = \frac1m \sum_{j=1}^m \ell(g, z_j) + \lambda K(g, f_0) \qquad (29).Rr​(g)=m1​j=1∑m​ℓ(g,zj​)+λK(g,f0​)(29).

For an index iii, the truncated objective is Rr∖i(g)=1m∑j≠iℓ(g,zj)+λK(g,f0)R_r^{\setminus i}(g) = \frac1m \sum_{j \ne i} \ell(g, z_j) + \lambda K(g, f_0)Rr∖i​(g)=m1​∑j=i​ℓ(g,zj​)+λK(g,f0​), and f∖if^{\setminus i}f∖i denotes one of its minimizers over FFF.

Formalization targets

Goal: Theorem 24

For every minimizer fff of (29), every minimizer f∖if^{\setminus i}f∖i of the truncated objective, and every example zzz,

∣ℓ(f,z)−ℓ(f∖i,z)∣≤M2λm.|\ell(f, z) - \ell(f^{\setminus i}, z)| \le \frac{M^2}{\lambda m}.∣ℓ(f,z)−ℓ(f∖i,z)∣≤λmM2​.

Milestones

  1. MMM-admissibility of (28) (§5.2.3, p. 518): ∣ℓ(g,z)−ℓ(g′,z)∣≤M∫Θ∣g−g′∣ dθ|\ell(g,z) - \ell(g',z)| \le M \int_\Theta |g - g'|\,d\theta∣ℓ(g,z)−ℓ(g′,z)∣≤M∫Θ​∣g−g′∣dθ.
  2. Pinsker's inequality, L1L^1L1 form (proof of Theorem 24): 12(∫Θ∣g−g′∣ dθ)2≤K(g,g′)\tfrac12 \bigl(\int_\Theta |g - g'|\,d\theta\bigr)^2 \le K(g, g')21​(∫Θ​∣g−g′∣dθ)2≤K(g,g′) for densities g,g′g, g'g,g′.
  3. Lemma 21 (p. 513): for a differentiable convex regularizer NNN on a vector space and a σ\sigmaσ-admissible loss,
dN(f,f∖i)+dN(f∖i,f)≤1λm(ℓ(f∖i,zi)−ℓ(f,zi)−dℓ(⋅,zi)(f∖i,f))≤σλm∣Δf(xi)∣.d_N(f, f^{\setminus i}) + d_N(f^{\setminus i}, f) \le \frac{1}{\lambda m}\Bigl(\ell(f^{\setminus i}, z_i) - \ell(f, z_i) - d_{\ell(\cdot, z_i)}(f^{\setminus i}, f)\Bigr) \le \frac{\sigma}{\lambda m}|\Delta f(x_i)|.dN​(f,f∖i)+dN​(f∖i,f)≤λm1​(ℓ(f∖i,zi​)−ℓ(f,zi​)−dℓ(⋅,zi​)​(f∖i,f))≤λmσ​∣Δf(xi​)∣.
  1. Bregman divergence of the relative entropy (proof of Theorem 24): dK(⋅,f0)(g,g′)=K(g,g′)d_{K(\cdot, f_0)}(g, g') = K(g, g')dK(⋅,f0​)​(g,g′)=K(g,g′).
  2. L1L^1L1 displacement bound (proof of Theorem 24):
∫Θ∣f−f∖i∣ dθ≤Mλm.\int_\Theta |f - f^{\setminus i}|\,d\theta \le \frac{M}{\lambda m}.∫Θ​∣f−f∖i∣dθ≤λmM​.

Significance

Theorem 24 places entropy-regularized posteriors among the algorithms to which the paper's exponential generalization bound applies: combined with Theorem 12 it gives, for the averaged loss, a deviation of order M2/(λm)+(M2/λ+M)ln⁡(1/δ)/mM^2/(\lambda m) + (M^2/\lambda + M)\sqrt{\ln(1/\delta)/m}M2/(λm)+(M2/λ+M)ln(1/δ)/m​. The proof also yields the L1L^1L1 bound ∫∣f−f∖i∣≤M/(λm)\int |f - f^{\setminus i}| \le M/(\lambda m)∫∣f−f∖i∣≤M/(λm), which by itself gives classification stability M/(λm)M/(\lambda m)M/(λm) for base hypotheses with values in {−1,1}\{-1, 1\}{−1,1} (remark after Theorem 24, p. 518).

The result is proved in the paper; no machine-checked proof is known to exist. A formalization produces reusable pieces that Mathlib does not have: Pinsker's inequality for densities in L1L^1L1 form (Mathlib has the Kullback–Leibler divergence InformationTheory.klDiv, but not Pinsker), the Bregman identity for the relative entropy, and a stability statement for minimizers over a space of probability densities.

Difficulty

The paper derives Theorem 24 from Lemma 21, which is stated for a regularizer that is defined and differentiable on a vector space. The relative entropy K(⋅,f0)K(\cdot, f_0)K(⋅,f0​) is defined only on the convex set of densities and is not differentiable at densities that vanish on a set of positive measure, so the general lemma does not literally apply, and the identity dK(⋅,f0)=Kd_{K(\cdot,f_0)} = KdK(⋅,f0​)​=K needs integrability conditions that the page does not state. A complete proof of the goal must either justify that application on the set of densities, or work directly with the minimizers, which requires identifying them and handling the +∞+\infty+∞ values of KKK. Pinsker's inequality itself requires a separate argument at the level of general measures.

Formalization scope

  • Densities are IsDensity ν g: measurable, nonnegative, integrable, total mass one, with respect to a σ-finite reference measure ν. The integral dθd\thetadθ is always against ν, never Lebesgue measure.
  • The base loss is r : Θ → Z → ℝ, measurable in θ, with 0 ≤ r ≤ M; the paper's costs are nonnegative (p. 502).
  • KKK is InformationTheory.klDiv of the measures g · ν and g' · ν, in ℝ≥0∞. The objectives (29) and its truncation take values in ℝ≥0∞. A formalization that converts KKK to a real number with toReal would send K=+∞K = +\inftyK=+∞ to 000 and make the worst densities minimizers; that reading is excluded.
  • The minimizers are given as hypotheses: f minimizes (29) and f' minimizes the truncated objective over all densities, for the given S : Fin m → Z and i : Fin m.
  • Corrected reading of the algorithm on S∖iS^{\setminus i}S∖i. The goal is stated in the pairwise form of the paper's proof: f∖if^{\setminus i}f∖i minimizes the truncated objective with factor 1/m1/m1/m, the analogue of (20), not (29) run on the m−1m-1m−1 points of S∖iS^{\setminus i}S∖i with factor 1/(m−1)1/(m-1)1/(m−1).
  • Corrected display. The objective displayed before Theorem 24 has ℓ(g,z)\ell(g, z)ℓ(g,z) inside the sum; (29) has ℓ(g,zi)\ell(g, z_i)ℓ(g,zi​), which is used.
  • Lemma 21 is stated as printed, in its differentiable case, on a real normed space whose elements act as functions on XXX through a linear map; the goal does not instantiate it. The Bregman identity is stated with the explicit gradient ln⁡(g′/f0)+1\ln(g'/f_0) + 1ln(g′/f0​)+1, for f0,g′>0f_0, g' > 0f0​,g′>0, finite K(g,f0)K(g, f_0)K(g,f0​), K(g′,f0)K(g', f_0)K(g′,f0​), and integrable gln⁡(g′/f0)g \ln(g'/f_0)gln(g′/f0​).

Contributions welcome: proofs of Pinsker's inequality for klDiv (reusable far beyond this mission), of the Bregman identity, of Lemma 21, and of the goal by any route.

Selected references

  • O. Bousquet and A. Elisseeff, Stability and Generalization, Journal of Machine Learning Research 2 (2002) 499–526. https://www.jmlr.org/papers/v2/bousquet02a.html
  • T. M. Cover and J. A. Thomas, Elements of Information Theory, Wiley, 1991 (Pinsker's inequality). https://doi.org/10.1002/0471200611
  • R. T. Rockafellar, Convex Analysis, Princeton University Press, 1970 (Bregman divergences, Appendix C of the paper). https://doi.org/10.1515/9781400873173
11 thms2 active usersReviewed
Machine LearningProbability·Captain: mikedeng1

Stability and Generalization 2: Exponential Generalization Bounds for Uniformly Stable AlgorithmsResearch Paper

Motivation

A learning algorithm is judged by its generalization error: the expected loss of the hypothesis it outputs on a fresh example. That quantity depends on an unknown distribution, so it is estimated from the training data, by the empirical error (the average loss on the training set) or the leave-one-out error (the average loss on each training point of the hypothesis trained without it). Classical learning theory controls the gap between these estimates and the true error uniformly over a hypothesis class, through its VC dimension or covering numbers. Such bounds say nothing useful about algorithms that search very large or infinite-dimensional spaces, such as support vector machines and regularization networks in a reproducing kernel Hilbert space.

Bousquet and Elisseeff (JMLR 2 (2002) 499–526) replaced the capacity of the class by a property of the algorithm, its stability: how much its output changes when one training example is removed. Their exponential bound for uniformly stable algorithms is the starting point of the stability approach to generalization, which was later used for stochastic gradient descent (Hardt, Recht and Singer, 2016) and differential privacy, and sharpened by Feldman and Vondrák (2019) and Bousquet, Klochkov and Zhivotovskiy (2020).

Timeline. Rogers and Wagner (1978) and Devroye and Wagner (1979) bounded the leave-one-out error of local rules such as k-nearest neighbours through their stability. McDiarmid (1989) proved the bounded-differences inequality. Lugosi and Pawlak (1994) combined it with smoothed error estimates. Kearns and Ron (1999) named hypothesis and error stability and related them to the VC dimension. Bousquet and Elisseeff (2002) introduced uniform stability and proved the exponential bounds this mission formalizes.

Setting

Let Z=X×YZ = X \times YZ=X×Y be a measurable space of labelled examples with an unknown probability distribution DDD. A training set S=(z1,…,zm)S = (z_1, \dots, z_m)S=(z1​,…,zm​) is drawn from DmD^mDm. A learning algorithm AAA maps a training set to a hypothesis AS:X→Y′A_S : X \to Y'AS​:X→Y′. It is deterministic and symmetric: it depends on the training set only as a multiset, so it is a function Multiset (X × Y) → (X → Y'), defined for training sets of every size. For a cost ccc, the loss of a hypothesis fff at z=(x,y)z = (x, y)z=(x,y) is ℓ(f,z)=c(f(x),y)\ell(f, z) = c(f(x), y)ℓ(f,z)=c(f(x),y).

Given SSS, write S∖iS^{\setminus i}S∖i for SSS with its iii-th example removed, and SiS^iSi for SSS with ziz_izi​ replaced by an independent draw zi′∼Dz_i' \sim Dzi′​∼D. The three errors are

R=Ez∼D[ℓ(AS,z)],Remp=1m∑i=1mℓ(AS,zi),Rloo=1m∑i=1mℓ(AS∖i,zi).R = \mathbb E_{z \sim D}[\ell(A_S, z)], \qquad R_{\mathrm{emp}} = \frac1m \sum_{i=1}^m \ell(A_S, z_i), \qquad R_{\mathrm{loo}} = \frac1m \sum_{i=1}^m \ell(A_{S^{\setminus i}}, z_i).R=Ez∼D​[ℓ(AS​,z)],Remp​=m1​i=1∑m​ℓ(AS​,zi​),Rloo​=m1​i=1∑m​ℓ(AS∖i​,zi​).

An algorithm has uniform stability β\betaβ at sample size mmm (Definition 6) if for every S∈ZmS \in Z^mS∈Zm, every iii and every z∈Zz \in Zz∈Z,

∣ℓ(AS,z)−ℓ(AS∖i,z)∣≤β.|\ell(A_S, z) - \ell(A_{S^{\setminus i}}, z)| \le \beta .∣ℓ(AS​,z)−ℓ(AS∖i​,z)∣≤β.

As a function of the sample size this constant is written βm\beta_mβm​.

Formalization targets

Goal: Theorem 12

If AAA has uniform stability β\betaβ and 0≤ℓ(AS,z)≤M0 \le \ell(A_S, z) \le M0≤ℓ(AS​,z)≤M for all zzz and all training sets SSS, then for every m≥1m \ge 1m≥1 and δ∈(0,1)\delta \in (0,1)δ∈(0,1), each of the following holds, separately, with probability at least 1−δ1 - \delta1−δ over S∼DmS \sim D^mS∼Dm:

R≤Remp+2β+(4mβ+M)ln⁡(1/δ)2m,(11)R \le R_{\mathrm{emp}} + 2\beta + (4m\beta + M)\sqrt{\frac{\ln(1/\delta)}{2m}}, \qquad (11)R≤Remp​+2β+(4mβ+M)2mln(1/δ)​​,(11) R≤Rloo+β+(4mβ+M)ln⁡(1/δ)2m.(12)R \le R_{\mathrm{loo}} + \beta + (4m\beta + M)\sqrt{\frac{\ln(1/\delta)}{2m}}. \qquad (12)R≤Rloo​+β+(4mβ+M)2mln(1/δ)​​.(12)

Milestones

  1. McDiarmid's inequality (Theorem 2): for measurable F:Zm→RF : Z^m \to \mathbb RF:Zm→R with ∣F(S)−F(Si)∣≤ci|F(S) - F(S^i)| \le c_i∣F(S)−F(Si)∣≤ci​, PS[F−ESF≥ϵ]≤e−2ϵ2/∑ici2P_S[F - \mathbb E_S F \ge \epsilon] \le e^{-2\epsilon^2/\sum_i c_i^2}PS​[F−ES​F≥ϵ]≤e−2ϵ2/∑i​ci2​.
  2. Uniform stability β\betaβ implies ∣ℓ(AS,z)−ℓ(ASi,z)∣≤2β|\ell(A_S, z) - \ell(A_{S^i}, z)| \le 2\beta∣ℓ(AS​,z)−ℓ(ASi​,z)∣≤2β (p. 504).
  3. Lemma 7: the bias identities for ES[R−Remp]\mathbb E_S[R - R_{\mathrm{emp}}]ES​[R−Remp​], ES[R(A,S∖i)−Rloo]\mathbb E_S[R(A,S^{\setminus i}) - R_{\mathrm{loo}}]ES​[R(A,S∖i)−Rloo​] and ES[R−Rloo]\mathbb E_S[R - R_{\mathrm{loo}}]ES​[R−Rloo​].
  4. R−RempR - R_{\mathrm{emp}}R−Remp​ and R−RlooR - R_{\mathrm{loo}}R−Rloo​ have bounded differences ci=4β+M/mc_i = 4\beta + M/mci​=4β+M/m.
  5. ES[R−Remp]≤2β\mathbb E_S[R - R_{\mathrm{emp}}] \le 2\betaES​[R−Remp​]≤2β and ES[R−Rloo]≤β\mathbb E_S[R - R_{\mathrm{loo}}] \le \betaES​[R−Rloo​]≤β.
  6. The tail bounds PS[R−Remp>ϵ+2β]≤exp⁡(−2mϵ2/(4mβ+M)2)P_S[R - R_{\mathrm{emp}} > \epsilon + 2\beta] \le \exp(-2m\epsilon^2/(4m\beta+M)^2)PS​[R−Remp​>ϵ+2β]≤exp(−2mϵ2/(4mβ+M)2) and the leave-one-out analogue.

Significance

When β=O(1/m)\beta = O(1/m)β=O(1/m) both bounds are O(1/m)O(1/\sqrt m)O(1/m​), with constants that do not depend on any capacity of the hypothesis space. Later sections of the paper show that Tikhonov regularization in a reproducing kernel Hilbert space has β=O(1/(λm))\beta = O(1/(\lambda m))β=O(1/(λm)), so the theorem gives generalization bounds for support vector regression, kernel ridge regression and, through a smoothed loss, soft-margin classification. The theorem is also the template for later stability bounds: the decomposition into a bias term controlled by stability and a deviation term controlled by a concentration inequality recurs throughout the literature.

The result has been proved since 2002, and replace-one variants appear in textbooks (Mohri, Rostamizadeh and Talwalkar, Foundations of Machine Learning, Theorem 14.2; Shalev-Shwartz and Ben-David, Chapter 13). On Prove2Me the replace-one textbook version is not formalized, and Mathlib at the platform's environment has no McDiarmid inequality. This mission asks for a machine-checked proof of the paper's remove-one version with its exact constants, and a reusable McDiarmid inequality with per-coordinate constants.

Difficulty

The deterministic steps (the bias identity and the bounded-differences estimates) are short on paper. The central difficulty is McDiarmid's inequality itself: it needs a martingale argument along the coordinates of a product measure, or an equivalent tensorization of conditional sub-Gaussian bounds, with the Doob martingale E[F∣z1,…,zk]\mathbb E[F \mid z_1, \dots, z_k]E[F∣z1​,…,zk​] expressed through partial integration over Measure.pi. Hoeffding's inequality for sums, which Mathlib has, does not apply directly: R−RempR - R_{\mathrm{emp}}R−Remp​ is not a sum of independent terms. A second, bookkeeping difficulty is Lemma 7: the identities rest on exchanging ziz_izi​ with zi′z_i'zi′​ and on the symmetry of AAA, which in Lean means measure-preserving coordinate permutations of Dm⊗DD^m \otimes DDm⊗D and multiset equalities such as Si ∖i=S∖iS^{i\,\setminus i} = S^{\setminus i}Si∖i=S∖i.

Formalization scope

Conventions committed to by the Lean statements:

  • An algorithm is a function of a multiset; this is how symmetry is encoded. Samples are Fin m → X × Y, and SiS^iSi is Function.update.
  • The law of SSS is Measure.pi (fun _ => D) with D a probability measure; zi′z_i'zi′​ and zzz are independent draws, integrated against the product (Measure.pi fun _ => D).prod D.
  • The paper's standing assumption that all functions are measurable is one hypothesis: (S,z)↦ℓ(AS,z)(S, z) \mapsto \ell(A_S, z)(S,z)↦ℓ(AS​,z) is measurable for every sample size. With the bound 0≤ℓ(AT,z)≤M0 \le \ell(A_T, z) \le M0≤ℓ(AT​,z)≤M for training sets TTT of every size, every expectation is a genuine integral, so no bound can hold because a non-integrable expectation defaults to 000.
  • Uniform stability quantifies over every sample, every index and every point, not almost every one.
  • "With probability at least 1−δ1 - \delta1−δ" means the DmD^mDm-measure of the failure set is at most δ\deltaδ. The two bounds (11) and (12) are separate statements, joined by a conjunction; they are not claimed for one joint event.
  • The paper assumes βm\beta_mβm​ is non-increasing in mmm and bounds βm−1\beta_{m-1}βm−1​ by βm\beta_mβm​ (p. 504). The leave-one-out bound (12) and its milestones carry the explicit hypothesis of uniform stability β\betaβ at size m−1m-1m−1; the empirical bound (11) does not.
  • McDiarmid's inequality sums ci2c_i^2ci2​ over i=1,…,mi = 1, \dots, mi=1,…,m; the paper's printed upper index nnn is a slip.
  • When a displayed tail bound has a zero denominator, its formal statement uses the limiting bound 000. In McDiarmid's inequality this is the constant-function case; in the stability tails the loss is identically zero.

The stability notion is the paper's remove-one notion. A formalization with replace-one stability would prove a different theorem with different constants, and the published FoundationsML_Stability_UniformlyStable (replace-one) is therefore not used.

The development needs the published loss, empirical-error and generalization-error definitions from Foundations of Machine Learning, a McDiarmid inequality on product measures (reusable for any bounded-differences argument), and the coordinate-exchange lemmas for Measure.pi behind Lemma 7. Contributions of a general McDiarmid inequality, of exchangeability lemmas for product measures, and of proofs of any milestone are welcome.

Selected references

  • O. Bousquet and A. Elisseeff, Stability and Generalization, Journal of Machine Learning Research 2 (2002) 499–526. https://www.jmlr.org/papers/v2/bousquet02a.html
  • C. McDiarmid, On the method of bounded differences, Surveys in Combinatorics, LMS Lecture Note Series 141 (1989) 148–188. https://doi.org/10.1017/CBO9781107359949.008
  • L. Devroye and T. Wagner, Distribution-free performance bounds for potential function rules, IEEE Trans. Inform. Theory 25 (1979) 601–604. https://doi.org/10.1109/TIT.1979.1056087
  • M. Kearns and D. Ron, Algorithmic stability and sanity-check bounds for leave-one-out cross-validation, Neural Computation 11 (1999) 1427–1453. https://doi.org/10.1162/089976699300016304
  • M. Hardt, B. Recht and Y. Singer, Train faster, generalize better: stability of stochastic gradient descent, ICML 2016. https://arxiv.org/abs/1509.01240
  • V. Feldman and J. Vondrák, High probability generalization bounds for uniformly stable algorithms with nearly optimal rate, COLT 2019. https://arxiv.org/abs/1902.10710
  • O. Bousquet, Y. Klochkov and N. Zhivotovskiy, Sharper bounds for uniformly stable algorithms, COLT 2020. https://arxiv.org/abs/1910.07833
  • M. Mohri, A. Rostamizadeh and A. Talwalkar, Foundations of Machine Learning, 2nd ed., MIT Press, 2018, Chapter 14.
13 thms2 active usersReviewed
Bandit AlgorithmsMachine LearningOperations Research·Captain: mikedeng1

Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems III: Contextual Bandits and the Banditron Mistake BoundTextbook

Motivation

In many sequential decision problems the learner sees side information before acting. A news site chooses an article for a visitor whose history and location it knows; an ad server chooses an advertisement for a query. Only the reward of the chosen action is observed. These are contextual bandit problems, and Chapter 4 of Bubeck and Cesa-Bianchi's monograph Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems (arXiv:1204.5721v2) surveys several of their formal versions. In a contextual problem the learner is compared with the best policy, a map from contexts to arms, rather than with the best single arm.

This mission covers three of the chapter's models. The first marks each round with a context from a finite set. In the second, NNN experts give advice, as in prediction with expert advice. The third is the bandit multiclass problem: a linear classifier predicts one of KKK labels and then learns only whether its prediction was right. The goal is the mistake bound of the Banditron (Kakade, Shalev-Shwartz and Tewari, ICML 2008). The bound shows that one bit of feedback per round suffices to compete with every linear classifier, at regret O(n2/3)O(n^{2/3})O(n2/3).

Setting

There are K≥2K \ge 2K≥2 arms (or labels) {1,…,K}\{1,\dots,K\}{1,…,K} and rounds t=1,…,nt = 1, \dots, nt=1,…,n.

Adversarial losses. At round ttt an adversary assigns losses ℓi,t∈[0,1]\ell_{i,t} \in [0,1]ℓi,t​∈[0,1] to the arms and may adapt to the forecaster's past plays I1,…,It−1I_1, \dots, I_{t-1}I1​,…,It−1​. The forecaster draws ItI_tIt​ at random from a distribution ptp_tpt​ that depends on what it has observed, and it observes only ℓIt,t\ell_{I_t,t}ℓIt​,t​. Expectations E\mathbb EE are over the forecaster's draws.

Side information. Each round carries a context sts_tst​ from a finite set S\mathcal SS, and the sequence s1,s2,…s_1, s_2, \dotss1​,s2​,… is fixed in advance. The pseudo-regret against context-to-arm maps is

R‾nS=max⁡g:S→{1,…,K}E[∑t=1nℓIt,t−∑t=1nℓg(st),t].\overline R^{\mathcal S}_n = \max_{g:\mathcal S\to\{1,\dots,K\}} \mathbb E\Big[\sum_{t=1}^n \ell_{I_t,t} - \sum_{t=1}^n \ell_{g(s_t),t}\Big].RnS​=g:S→{1,…,K}max​E[t=1∑n​ℓIt​,t​−t=1∑n​ℓg(st​),t​].

The S-Exp3 forecaster runs one instance of Exp3 (Section 3.1 of the book) on each context.

Expert advice. At each round each of NNN experts jjj proposes a distribution ξtj\xi^j_tξtj​ over arms, which may depend on the forecaster's past plays. The contextual pseudo-regret is

R‾nctx=max⁡k=1,…,NE[∑t=1nℓIt,t−∑t=1nEi∼ξtkℓi,t].\overline R^{\mathrm{ctx}}_n = \max_{k=1,\dots,N}\mathbb E\Big[\sum_{t=1}^n \ell_{I_t,t} - \sum_{t=1}^n \mathbb E_{i\sim\xi^k_t}\ell_{i,t}\Big].Rnctx​=k=1,…,Nmax​E[t=1∑n​ℓIt​,t​−t=1∑n​Ei∼ξtk​​ℓi,t​].

Exp4 (Fig. 4.1) runs exponential weights over the experts with importance-weighted loss estimates.

Bandit multiclass. The examples (xt,yt)∈Rd×{1,…,K}(x_t, y_t) \in \mathbb R^d \times \{1,\dots,K\}(xt​,yt​)∈Rd×{1,…,K} are fixed in advance, with ∥xt∥=1\|x_t\| = 1∥xt​∥=1 (Euclidean). A K×dK\times dK×d matrix UUU classifies xxx by arg⁡max⁡i(Ux)i\arg\max_i (Ux)_iargmaxi​(Ux)i​. Its multiclass hinge loss on round ttt is ℓt(U)=[1−(Uxt)yt+max⁡i≠yt(Uxt)i]+\ell_t(U) = [1 - (Ux_t)_{y_t} + \max_{i\neq y_t}(Ux_t)_i]_+ℓt​(U)=[1−(Uxt​)yt​​+maxi=yt​​(Uxt​)i​]+​. Write Ln(U)=∑t≤nℓt(U)L_n(U) = \sum_{t\le n}\ell_t(U)Ln​(U)=∑t≤n​ℓt​(U) for the cumulative hinge loss, Lˉn(U)=Ln(U)/n\bar L_n(U) = L_n(U)/nLˉn​(U)=Ln​(U)/n for its average, and ∥U∥\|U\|∥U∥ for the Frobenius norm. The multiclass Perceptron predicts y^t=arg⁡max⁡i(Wtxt)i\hat y_t = \arg\max_i (W_tx_t)_iy^​t​=argmaxi​(Wt​xt​)i​ and, after seeing yty_tyt​, adds xtx_txt​ to row yty_tyt​ and subtracts it from row y^t\hat y_ty^​t​. The Banditron (p. 58) predicts YtY_tYt​ from pi,t=(1−γ)1y^t=i+γ/Kp_{i,t} = (1-\gamma)\mathbb 1_{\hat y_t = i} + \gamma/Kpi,t​=(1−γ)1y^​t​=i​+γ/K. It observes only 1Yt=yt\mathbb 1_{Y_t = y_t}1Yt​=yt​​ and updates Wt+1=Wt+X~tW_{t+1} = W_t + \widetilde X_tWt+1​=Wt​+Xt​, where (X~t)i,j=xt,j(1Yt=yt1Yt=i/pi,t−1y^t=i)(\widetilde X_t)_{i,j} = x_{t,j}\big(\mathbb 1_{Y_t=y_t}\mathbb 1_{Y_t=i}/p_{i,t} - \mathbb 1_{\hat y_t=i}\big)(Xt​)i,j​=xt,j​(1Yt​=yt​​1Yt​=i​/pi,t​−1y^​t​=i​). Its number of mistakes is Mn=∑t≤n1Yt≠ytM_n = \sum_{t\le n}\mathbb 1_{Y_t\neq y_t}Mn​=∑t≤n​1Yt​=yt​​.

Formalization targets

Goal: Theorem 4.7 (Banditron)

For n≥8Kn \ge 8Kn≥8K, γ=(K/n)1/3\gamma = (K/n)^{1/3}γ=(K/n)1/3, every example sequence as above and every K×dK\times dK×d matrix UUU,

E Mn≤Ln(U)+(1+∥U∥2Lˉn(U))K1/3n2/3+2∥U∥2K2/3n1/3+2 ∥U∥K1/6n1/3.\mathbb E\,M_n \le L_n(U) + \Big(1 + \|U\|\sqrt{2\bar L_n(U)}\Big)K^{1/3}n^{2/3} + 2\|U\|^2K^{2/3}n^{1/3} + \sqrt2\,\|U\|K^{1/6}n^{1/3}.EMn​≤Ln​(U)+(1+∥U∥2Lˉn​(U)​)K1/3n2/3+2∥U∥2K2/3n1/3+2​∥U∥K1/6n1/3.

Milestones

  1. Multiclass Perceptron bound (Section 4.4, p. 57). For every n≥1n \ge 1n≥1 and UUU, ∑t≤n1y^t≠yt≤Ln(U)+2∥U∥2+∥U∥2nLˉn(U)\sum_{t\le n}\mathbb 1_{\hat y_t\ne y_t} \le L_n(U) + 2\|U\|^2 + \|U\|\sqrt{2n\bar L_n(U)}∑t≤n​1y^​t​=yt​​≤Ln​(U)+2∥U∥2+∥U∥2nLˉn​(U)​.
  2. Theorem 4.1 (p. 44). S-Exp3 satisfies R‾nS≤2n∣S∣Kln⁡K\overline R^{\mathcal S}_n \le \sqrt{2n|\mathcal S|K\ln K}RnS​≤2n∣S∣KlnK​.
  3. Theorem 4.2 (p. 46), with corrected constants. Exp4 without mixing satisfies R‾nctx≤2nKln⁡N\overline R^{\mathrm{ctx}}_n \le \sqrt{2nK\ln N}Rnctx​≤2nKlnN​ for ηt=2ln⁡N/(nK)\eta_t = \sqrt{2\ln N/(nK)}ηt​=2lnN/(nK)​, and R‾nctx≤2nKln⁡N\overline R^{\mathrm{ctx}}_n \le 2\sqrt{nK\ln N}Rnctx​≤2nKlnN​ for ηt=ln⁡N/(tK)\eta_t = \sqrt{\ln N/(tK)}ηt​=lnN/(tK)​.
  4. Theorem 4.3 (p. 50), with corrected learning rate. Let the plays be drawn from distributions qtq_tqt​ with qi,t≥ε>0q_{i,t}\ge\varepsilon > 0qi,t​≥ε>0, and let Exp3 run on the estimates ℓi,t1It=i/qi,t\ell_{i,t}\mathbb 1_{I_t=i}/q_{i,t}ℓi,t​1It​=i​/qi,t​ with η=2εln⁡K/n\eta = \sqrt{2\varepsilon\ln K/n}η=2εlnK/n​. Then max⁡kE[∑tEi∼ptℓi,t−∑tℓk,t]≤(2n/ε)ln⁡K\max_k \mathbb E\big[\sum_t \mathbb E_{i\sim p_t}\ell_{i,t} - \sum_t\ell_{k,t}\big] \le \sqrt{(2n/\varepsilon)\ln K}maxk​E[∑t​Ei∼pt​​ℓi,t​−∑t​ℓk,t​]≤(2n/ε)lnK​.

Significance

Theorem 4.7 shows that, on any sequence of examples, the bandit version of online multiclass classification costs at most O(K1/3n2/3)O(K^{1/3}n^{2/3})O(K1/3n2/3) mistakes beyond the hinge loss of the best linear classifier. The full-information Perceptron, by comparison, pays O(n)O(\sqrt n)O(n​). The bound has no stochastic assumption and has explicit constants. Theorems 4.1–4.3 are the basic regret guarantees for side information and expert advice. Theorem 4.3 in particular lets learning algorithms serve as experts inside Exp4, which is the construction behind Theorem 4.5.

The mission produces machine-checked statements, and eventually proofs, of these results with fully explicit constants and an explicit model of adaptive adversaries and adaptive advice. To the curators' knowledge none of the Banditron, the multiclass Perceptron bound, S-Exp3 or Theorem 4.3 is formalized anywhere. The platform's Bandit Algorithms series has a proved Exp4 bound, but only for advice and rewards fixed in advance. The book proves all four milestones and the goal; two printed statements (4.2 and 4.3) contain misprints that this mission corrects.

Difficulty

The Banditron bound concerns a randomized process whose weight matrix depends on all earlier random predictions. The Perceptron argument tracks ⟨U,Wn+1⟩\langle U, W_{n+1}\rangle⟨U,Wn+1​⟩ and ∥Wn+1∥2\|W_{n+1}\|^2∥Wn+1​∥2. It carries over only in conditional expectation, and the second moment of the importance-weighted update is of order K/γK/\gammaK/γ on rounds where y^t≠yt\hat y_t \neq y_ty^​t​=yt​ and of order γ\gammaγ otherwise. Combining these into one inequality for ∑tP(y^t≠yt)\sum_t\mathbb P(\hat y_t\neq y_t)∑t​P(y^​t​=yt​) and then for EMn\mathbb E M_nEMn​ requires solving a quadratic inequality in the presence of expectations, and the constants must come out as printed. For the Exp3/Exp4 results, the obstacle is that losses and advice adapt to past plays. The standard potential argument has to be run conditionally on the history, and a version that fixes the losses in advance proves a weaker theorem.

Formalization scope

  • Rounds and laws. Rounds are numbered from 000 in Lean (Lean round ttt is the book's round t+1t+1t+1). Every forecaster is a sampling rule from past plays to weights on Fin K. The law of the first nnn plays is the product ∏tpt(ωt∣ω<t)\prod_t p_t(\omega_t\mid\omega_{<t})∏t​pt​(ωt​∣ω<t​) over sequences ω:Fin n→Fin K\omega : \mathrm{Fin}\,n\to\mathrm{Fin}\,Kω:Finn→FinK, and expectations are finite sums against it. The adversary and the experts are deterministic functions of past plays; an independent randomized adversary is a mixture of these. The examples of the Banditron are fixed.
  • Argmax. y^t\hat y_ty^​t​ uses any argmax selector; all tie-breaking rules are covered.
  • Norms. ∥xt∥=1\|x_t\| = 1∥xt​∥=1 is the Euclidean condition ∑jxt,j2=1\sum_j x_{t,j}^2 = 1∑j​xt,j2​=1; ∥U∥\|U\|∥U∥ is the Frobenius norm written out explicitly.
  • Infima and maxima. Each "inf⁡U\inf_UinfU​" and "max⁡k\max_kmaxk​" of the book is stated as "for every UUU" or "for every kkk", which is equivalent.
  • Explicit constants. Every bound is the one printed or, for the corrected items, the one the proof yields. No O(⋅)O(\cdot)O(⋅) appears.
  • Corrected misprints. Theorem 4.7 prints the examples in Rd×{−1,+1}\mathbb R^d\times\{-1,+1\}Rd×{−1,+1}; labels are in {1,…,K}\{1,\dots,K\}{1,…,K}. Theorem 4.2 prints 2nNln⁡K\sqrt{2nN\ln K}2nNlnK​ and 2nNln⁡K2\sqrt{nN\ln K}2nNlnK​; the proof gives 2nKln⁡N\sqrt{2nK\ln N}2nKlnN​ and 2nKln⁡N2\sqrt{nK\ln N}2nKlnN​. Theorem 4.3 prints η=2ln⁡K/(nK)\eta = \sqrt{2\ln K/(nK)}η=2lnK/(nK)​; (4.7) follows from the proof with η=2εln⁡K/n\eta = \sqrt{2\varepsilon\ln K/n}η=2εlnK/n​.
  • Parameter range. At n=8Kn = 8Kn=8K the Banditron's γ\gammaγ equals 1/21/21/2, outside the box's open interval (0,1/2)(0,1/2)(0,1/2). The proof uses only γ≤1/2\gamma\le 1/2γ≤1/2, so n=8Kn = 8Kn=8K is included.
  • Ruling out trivial forms. Theorem 4.1 is stated for the explicit S-Exp3 forecaster, not as an existence claim, so no forecaster tuned to the losses can witness it. The losses and the advice are allowed to adapt, so a proof for oblivious sequences does not suffice.
  • Left out. Theorem 4.4 (Exp4 with mixing) is proved in the book only by reference. The argument that reference suggests yields 32γn+Kln⁡N/γ\tfrac32\gamma n + K\ln N/\gamma23​γn+KlnN/γ, not the printed γn/2+Kln⁡N/γ\gamma n/2 + K\ln N/\gammaγn/2+KlnN/γ. Theorem 4.5 is stated with O(⋅)O(\cdot)O(⋅), Theorem 4.6 "for some constant ccc", and Eq. (4.8) is left to the reader.

Useful reusable infrastructure: the path-law expectation for history-dependent sampling, the exponential-weights potential argument under adaptive losses, and Perceptron-type inner-product arguments for matrices. Proofs of any milestone and of the goal are welcome.

Selected references

  • S. Bubeck, N. Cesa-Bianchi, Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems, Foundations and Trends in Machine Learning 5(1), 2012. arXiv:1204.5721v2. https://arxiv.org/abs/1204.5721 ; https://doi.org/10.1561/2200000024
  • S. M. Kakade, S. Shalev-Shwartz, A. Tewari, Efficient Bandit Algorithms for Online Multiclass Prediction, ICML 2008. https://doi.org/10.1145/1390156.1390212
  • P. Auer, N. Cesa-Bianchi, Y. Freund, R. E. Schapire, The Nonstochastic Multiarmed Bandit Problem, SIAM Journal on Computing 32(1), 2002. https://doi.org/10.1137/S0097539701398375
  • O.-A. Maillard, R. Munos, Adaptive Bandits: Towards the Best History-Dependent Strategy, AISTATS 2011. https://proceedings.mlr.press/v15/maillard11a.html
11 thms2 active usersReviewed
Machine LearningProbability·Captain: mikedeng1

Certified Adversarial Robustness via Randomized Smoothing 1: The Gaussian-Smoothed Classifier Is Constant on the ℓ2 Ball of Radius (σ/2)(Φ⁻¹(p_A) − Φ⁻¹(p_B))Research Paper

Motivation

Classifiers trained on images, speech and text can be made to change their prediction by perturbations of the input that are tiny in norm (Szegedy et al., 2014). Empirical defences against such adversarial examples have repeatedly been broken by stronger attacks (Athalye, Carlini, Wagner, 2018), which motivates certified defences: classifiers that come with a proof that their prediction at a given input cannot change inside a stated ball around it.

Randomized smoothing turns any classifier, however large or opaque, into one with such a certificate in the ℓ2\ell_2ℓ2​ norm. It was introduced with weaker radii by Lecuyer et al. (2019) and Li et al. (2018). Cohen, Rosenfeld and Kolter (arXiv:1902.02918v2, ICML 2019) proved the radius that is now standard, and showed it cannot be enlarged. Their guarantee underlies most later work on certified ℓ2\ell_2ℓ2​ robustness, including Salman et al. (2019).

This mission formalizes the robustness guarantee, Theorem 1 of that paper. Page numbers below are PDF pages of the arXiv v2 preprint, which has no printed page numbers.

Setting

Inputs are points of Rd\mathbb R^dRd with the Euclidean norm ∥⋅∥\|\cdot\|∥⋅∥ and inner product δ⊤z\delta^\top zδ⊤z. Classes form a set Y\mathcal YY. A base classifier is a deterministic or random function f:Rd→Yf : \mathbb R^d \to \mathcal Yf:Rd→Y. A random fff is described by the probabilities P(f(z)=c)\mathbb P(f(z) = c)P(f(z)=c), which for each zzz form a probability distribution on Y\mathcal YY.

Fix a noise level σ>0\sigma > 0σ>0 and let ε∼N(0,σ2I)\varepsilon \sim \mathcal N(0, \sigma^2 I)ε∼N(0,σ2I) be isotropic Gaussian noise. The class probabilities at xxx are P(f(x+ε)=c)\mathbb P(f(x + \varepsilon) = c)P(f(x+ε)=c), and the smoothed classifier is

g(x)=arg⁡max⁡c∈YP(f(x+ε)=c).(1)g(x) = \arg\max_{c \in \mathcal Y} \mathbb P\big(f(x + \varepsilon) = c\big). \qquad (1)g(x)=argc∈Ymax​P(f(x+ε)=c).(1)

The paper leaves g(x)g(x)g(x) undefined when the maximizer is not unique. "g(x)=cg(x) = cg(x)=c" therefore means that every class other than ccc has strictly smaller probability.

Write Φ\PhiΦ for the standard Gaussian cumulative distribution function and Φ−1\Phi^{-1}Φ−1 for its inverse. Φ−1\Phi^{-1}Φ−1 is a real number on (0,1)(0,1)(0,1), and Φ−1(0)=−∞\Phi^{-1}(0) = -\inftyΦ−1(0)=−∞, Φ−1(1)=+∞\Phi^{-1}(1) = +\inftyΦ−1(1)=+∞.

Formalization targets

Goal: Theorem 1 (p. 4; restated p. 13)

Suppose that at a specific xxx there are a class cAc_AcA​ and numbers pA‾,pB‾∈[0,1]\underline{p_A}, \overline{p_B} \in [0,1]pA​​,pB​​∈[0,1] with

P(f(x+ε)=cA) ≥ pA‾ ≥ pB‾ ≥ max⁡c≠cAP(f(x+ε)=c).(6)\mathbb P\big(f(x + \varepsilon) = c_A\big) \ \ge\ \underline{p_A} \ \ge\ \overline{p_B} \ \ge\ \max_{c \ne c_A} \mathbb P\big(f(x + \varepsilon) = c\big). \qquad (6)P(f(x+ε)=cA​) ≥ pA​​ ≥ pB​​ ≥ c=cA​max​P(f(x+ε)=c).(6)

Then g(x+δ)=cAg(x + \delta) = c_Ag(x+δ)=cA​ for every δ\deltaδ with ∥δ∥2<R\|\delta\|_2 < R∥δ∥2​<R, where

R=σ2(Φ−1(pA‾)−Φ−1(pB‾)).(7)R = \frac{\sigma}{2}\Big(\Phi^{-1}(\underline{p_A}) - \Phi^{-1}(\overline{p_B})\Big). \qquad (7)R=2σ​(Φ−1(pA​​)−Φ−1(pB​​)).(7)

The statement covers every base classifier and every set of classes. The radius is infinite when pA‾=1>pB‾\underline{p_A} = 1 > \overline{p_B}pA​​=1>pB​​ or pA‾>0=pB‾\underline{p_A} > 0 = \overline{p_B}pA​​>0=pB​​.

Milestones

The milestones are the paper's own steps, in attack order:

  • Lemma 3 (p. 12): the Neyman–Pearson lemma in both directions, for densities μX\mu_XμX​, μY\mu_YμY​ on Rd\mathbb R^dRd and the likelihood-ratio sets {μY≤tμX}\{\mu_Y \le t\mu_X\}{μY​≤tμX​} and {μY≥tμX}\{\mu_Y \ge t\mu_X\}{μY​≥tμX​}.
  • The likelihood ratio and (5) (proof of Lemma 4, p. 13): for X∼N(x,σ2I)X \sim \mathcal N(x,\sigma^2 I)X∼N(x,σ2I) and Y∼N(x+δ,σ2I)Y \sim \mathcal N(x+\delta,\sigma^2 I)Y∼N(x+δ,σ2I), μY/μX=exp⁡(aδ⊤z+b)\mu_Y/\mu_X = \exp(a\delta^\top z + b)μY​/μX​=exp(aδ⊤z+b), so half-spaces orthogonal to δ\deltaδ are likelihood-ratio sets.
  • Lemma 4 (pp. 12–13): Neyman–Pearson for these two Gaussians and the half-spaces {δ⊤z≤β}\{\delta^\top z \le \beta\}{δ⊤z≤β}, {δ⊤z≥β}\{\delta^\top z \ge \beta\}{δ⊤z≥β}.
  • The four Claims of Appendix A.0.1 (pp. 15–16, with (13) and (14) of p. 14): the probabilities of the half-spaces A={z:δ⊤(z−x)≤σ∥δ∥Φ−1(pA‾)}A = \{z : \delta^\top(z-x) \le \sigma\|\delta\|\Phi^{-1}(\underline{p_A})\}A={z:δ⊤(z−x)≤σ∥δ∥Φ−1(pA​​)} and B={z:δ⊤(z−x)≥σ∥δ∥Φ−1(1−pB‾)}B = \{z : \delta^\top(z-x) \ge \sigma\|\delta\|\Phi^{-1}(1-\overline{p_B})\}B={z:δ⊤(z−x)≥σ∥δ∥Φ−1(1−pB​​)} under XXX and YYY:
P(X∈A)=pA‾,P(X∈B)=pB‾,P(Y∈A)=Φ(Φ−1(pA‾)−∥δ∥σ),P(Y∈B)=Φ(Φ−1(pB‾)+∥δ∥σ).\mathbb P(X \in A) = \underline{p_A},\quad \mathbb P(X \in B) = \overline{p_B},\quad \mathbb P(Y \in A) = \Phi\Big(\Phi^{-1}(\underline{p_A}) - \tfrac{\|\delta\|}{\sigma}\Big),\quad \mathbb P(Y \in B) = \Phi\Big(\Phi^{-1}(\overline{p_B}) + \tfrac{\|\delta\|}{\sigma}\Big).P(X∈A)=pA​​,P(X∈B)=pB​​,P(Y∈A)=Φ(Φ−1(pA​​)−σ∥δ∥​),P(Y∈B)=Φ(Φ−1(pB​​)+σ∥δ∥​).
  • (15) (p. 14): P(Y∈A)>P(Y∈B)\mathbb P(Y \in A) > \mathbb P(Y \in B)P(Y∈A)>P(Y∈B) if and only if ∥δ∥<R\|\delta\| < R∥δ∥<R.

Significance

Theorem 1 turns three numbers at one input into a guarantee over a whole ball: a lower bound on the top-class probability, an upper bound on the other classes, and the noise level. These bounds can be estimated by sampling and certified with confidence intervals (the paper's CERTIFY procedure). That is what lets randomized smoothing certify ImageNet-scale networks, where exact verification methods do not scale. The companion result (Theorem 2, a separate mission of this series) shows that no larger ℓ2\ell_2ℓ2​ ball can be certified from the same information.

The theorem has a short pen-and-paper proof; no machine-checked proof of it is recorded on Prove2Me. A formal development adds:

  • a checked Neyman–Pearson lemma for randomized tests with densities on Rd\mathbb R^dRd, which Mathlib does not have;
  • the Gaussian likelihood-ratio and half-space computations;
  • an explicit treatment of the endpoint cases pA‾=1\underline{p_A} = 1pA​​=1 and pB‾=0\overline{p_B} = 0pB​​=0, where the radius is infinite.

Difficulty

Every step is classical, so the difficulty lies in the missing infrastructure, not in the idea.

  • Densities. Mathlib's multivariate Gaussian is defined as a pushforward of a product measure, not by a density. Identifying it with the density (2πσ2)−d/2e−∥z−x∥2/(2σ2)(2\pi\sigma^2)^{-d/2}e^{-\|z-x\|^2/(2\sigma^2)}(2πσ2)−d/2e−∥z−x∥2/(2σ2), which Lemma 4 needs, is not available.
  • Projections. The Claims need the law of δ⊤X\delta^\top Xδ⊤X for X∼N(x,σ2I)X \sim \mathcal N(x,\sigma^2 I)X∼N(x,σ2I) in closed form, namely a one-dimensional Gaussian with mean δ⊤x\delta^\top xδ⊤x and variance σ2∥δ∥2\sigma^2\|\delta\|^2σ2∥δ∥2.
  • The inverse CDF. Mathlib has no normal quantile, so Φ−1\Phi^{-1}Φ−1 is defined here as an infimum. The identities Φ(Φ−1(p))=p\Phi(\Phi^{-1}(p)) = pΦ(Φ−1(p))=p and Φ−1(1−p)=−Φ−1(p)\Phi^{-1}(1-p) = -\Phi^{-1}(p)Φ−1(1−p)=−Φ−1(p) must be derived.
  • Degenerate cases. The obvious argument through the worst-case half-spaces breaks down when δ=0\delta = 0δ=0 or when pA‾\underline{p_A}pA​​ or pB‾\overline{p_B}pB​​ is 000 or 111. There the half-spaces are empty or everything and Φ−1\Phi^{-1}Φ−1 is infinite, so these cases need a separate argument.

Formalization scope

  • Space and noise. Rd\mathbb R^dRd is EuclideanSpace ℝ (Fin d). N(x,σ2I)\mathcal N(x,\sigma^2 I)N(x,σ2I) is the pushforward of Mathlib's stdGaussian under z↦x+σzz \mapsto x + \sigma zz↦x+σz, with σ>0\sigma > 0σ>0 a hypothesis.
  • Classifiers. A random classifier is a map f : ℝᵈ → PMF 𝒴 with measurable class probabilities; deterministic classifiers are point masses. Y\mathcal YY is an arbitrary type, with no finiteness assumed. The class probability is the published Gaussian smoothing RandomGradFree.Shared.smoothing, applied to z↦P(f(z)=c)z \mapsto \mathbb P(f(z) = c)z↦P(f(z)=c).
  • The prediction. "g(x)=cg(x) = cg(x)=c" is the strict unique-maximizer predicate, and ggg itself is not defined. Defining ggg by an arbitrary choice at ties would make the theorem depend on the tie-break.
  • The radius. RRR is computed in the extended reals with Φ−1(0)=−∞\Phi^{-1}(0) = -\inftyΦ−1(0)=−∞ and Φ−1(1)=+∞\Phi^{-1}(1) = +\inftyΦ−1(1)=+∞. A real-valued Φ−1\Phi^{-1}Φ−1 with junk value 000 at the endpoints would assign a finite, wrong radius there, so it is used only in milestones that assume 0<p<10 < p < 10<p<1. In the two corners pA‾=pB‾∈{0,1}\underline{p_A} = \overline{p_B} \in \{0,1\}pA​​=pB​​∈{0,1}, where (7) reads ∞−∞\infty - \infty∞−∞, the radius is −∞-\infty−∞ and the goal is vacuous, as in the paper.
  • Neyman–Pearson. Random tests are [0,1][0,1][0,1]-valued measurable functions, so the lemmas apply with h(z)=P(f(z)=c)h(z) = \mathbb P(f(z) = c)h(z)=P(f(z)=c) for random fff. The likelihood-ratio sets are written multiplied out (μY≤tμX\mu_Y \le t\mu_XμY​≤tμX​), which avoids division by zero where μX\mu_XμX​ vanishes.
  • Added hypotheses. The milestones about AAA and BBB assume δ≠0\delta \ne 0δ=0 and 0<p<10 < p < 10<p<1, which the paper's computation uses implicitly.

Contributions are welcome at every level: proofs of the milestones, and general lemmas such as the Gaussian density, the law of linear functionals of a Gaussian vector and properties of the normal quantile. These lemmas are reusable well beyond this mission.

Selected references

  • J. M. Cohen, E. Rosenfeld, J. Z. Kolter, Certified Adversarial Robustness via Randomized Smoothing, ICML 2019; arXiv:1902.02918v2. https://arxiv.org/abs/1902.02918
  • J. Neyman, E. S. Pearson, On the Problem of the Most Efficient Tests of Statistical Hypotheses, Phil. Trans. R. Soc. A 231, 1933. https://doi.org/10.1098/rsta.1933.0009
  • M. Lecuyer, V. Atlidakis, R. Geambasu, D. Hsu, S. Jana, Certified Robustness to Adversarial Examples with Differential Privacy, IEEE S&P 2019. https://arxiv.org/abs/1802.03471
  • B. Li, C. Chen, W. Wang, L. Carin, Certified Adversarial Robustness with Additive Noise, NeurIPS 2019. https://arxiv.org/abs/1809.03113
  • H. Salman et al., Provably Robust Deep Learning via Adversarially Trained Smoothed Classifiers, NeurIPS 2019. https://arxiv.org/abs/1906.04584
  • A. Athalye, N. Carlini, D. Wagner, Obfuscated Gradients Give a False Sense of Security, ICML 2018. https://arxiv.org/abs/1802.00420
  • C. Szegedy et al., Intriguing Properties of Neural Networks, ICLR 2014. https://arxiv.org/abs/1312.6199
15 thms2 active usersReviewed
Machine LearningProbability·Captain: mikedeng1

The Sample Complexity of Pattern Classification with Neural Networks: The Size of the Weights is More Important than the Size of the Network I: Fat-Shattering Margin Bound with d = fat_H(γ/16)Research Paper

Motivation

Classical generalization bounds for classifiers, built on the VC dimension, grow with the number of adjustable parameters. For neural networks this is at odds with practice: networks with many more weights than training examples often generalize well. Bartlett's 1998 paper (IEEE Trans. Inform. Theory 44(2), 525–536) explains part of this by measuring a real-valued classifier's confidence. If a hypothesis classifies most training examples correctly with a margin γ\gammaγ, its misclassification probability is controlled by a scale-sensitive dimension of the class at scale proportional to γ\gammaγ, not by its VC dimension. Later in the paper this yields bounds for networks with small weights that do not depend on the number of weights.

This mission formalizes the first of the paper's two main technical results, the margin bound of Theorem 2 (p. 527), together with the steps of its proof on pp. 527–528.

The fat-shattering dimension was introduced by Kearns and Schapire (JCSS 1994). Alon, Ben-David, Cesa-Bianchi and Haussler (J. ACM 1997) proved the scale-sensitive Sauer-type covering bound used here (Theorem 5 of the paper). Shawe-Taylor, Bartlett, Williamson and Anthony (IEEE Trans. Inform. Theory 1998) proved the zero-training-error version (Theorem 1 of the paper). Theorem 2 extends it to hypotheses that make margin errors on the training data.

Setting

Let XXX be a set and PPP a probability distribution on X×{−1,1}X\times\{-1,1\}X×{−1,1}. The threshold function is sgn⁡(α)=−1\operatorname{sgn}(\alpha)=-1sgn(α)=−1 for α<0\alpha<0α<0 and sgn⁡(α)=1\operatorname{sgn}(\alpha)=1sgn(α)=1 for α≥0\alpha\ge0α≥0. For a real-valued hypothesis hhh on XXX, the misclassification probability is er⁡P(h)=P{sgn⁡(h(x))≠y}\operatorname{er}_P(h)=P\{\operatorname{sgn}(h(x))\ne y\}erP​(h)=P{sgn(h(x))=y}. For a sample z=((x1,y1),…,(xm,ym))z=((x_1,y_1),\dots,(x_m,y_m))z=((x1​,y1​),…,(xm​,ym​)) drawn independently from PPP and γ>0\gamma>0γ>0, the margin error estimate is

er⁡^zγ(h)=1m ∣{i:yih(xi)<γ}∣.\widehat{\operatorname{er}}{}^{\gamma}_z(h)=\tfrac1m\,|\{i : y_ih(x_i)<\gamma\}|.erzγ​(h)=m1​∣{i:yi​h(xi​)<γ}∣.

Let HHH be a class of real functions on XXX. Points x1,…,xmx_1,\dots,x_mx1​,…,xm​ are γ\gammaγ-shattered by HHH if some r∈Rmr\in\mathbb R^mr∈Rm has the following property: for every sign vector b∈{−1,1}mb\in\{-1,1\}^mb∈{−1,1}m, some h∈Hh\in Hh∈H satisfies (h(xi)−ri)bi≥γ(h(x_i)-r_i)b_i\ge\gamma(h(xi​)−ri​)bi​≥γ for all iii. The fat-shattering dimension fat⁡H(γ)\operatorname{fat}_H(\gamma)fatH​(γ) is the largest such mmm, possibly ∞\infty∞.

The proof uses the following objects:

  • the squashing function πγ(α)=max⁡(−γ,min⁡(γ,α))\pi_\gamma(\alpha)=\max(-\gamma,\min(\gamma,\alpha))πγ​(α)=max(−γ,min(γ,α)) and the class πγ(H)={πγ∘h:h∈H}\pi_\gamma(H)=\{\pi_\gamma\circ h:h\in H\}πγ​(H)={πγ​∘h:h∈H};
  • the sample ℓ∞\ell_\inftyℓ∞​ pseudometric dℓ∞(x)(f,g)=max⁡i∣f(xi)−g(xi)∣d_{\ell_\infty(x)}(f,g)=\max_i|f(x_i)-g(x_i)|dℓ∞​(x)​(f,g)=maxi​∣f(xi​)−g(xi​)∣;
  • the covering number N∞(F,ϵ,m)\mathcal N_\infty(F,\epsilon,m)N∞​(F,ϵ,m), the largest over x∈Xmx\in X^mx∈Xm of the size of the smallest ϵ\epsilonϵ-cover (Definition 3), and the corresponding packing number M∞(F,α,m)\mathcal M_\infty(F,\alpha,m)M∞​(F,α,m);
  • the quantization Qα(x)=⌈(x−α/2)/α⌉αQ_\alpha(x)=\lceil (x-\alpha/2)/\alpha\rceil\alphaQα​(x)=⌈(x−α/2)/α⌉α.

Formalization targets

Goal: Theorem 2

Assume 0<δ<1/20<\delta<1/20<δ<1/2, 0<γ<10<\gamma<10<γ<1, m≥1m\ge1m≥1, and d=fat⁡H(γ/16)d=\operatorname{fat}_H(\gamma/16)d=fatH​(γ/16) finite with d≤34md\le 34md≤34m. With probability at least 1−δ1-\delta1−δ over zzz, every h∈Hh\in Hh∈H satisfies

er⁡P(h)<er⁡^zγ(h)+2m(dln⁡34emdlog⁡2(578m)+ln⁡4δ).\operatorname{er}_P(h)<\widehat{\operatorname{er}}{}^{\gamma}_z(h)+\sqrt{\frac2m\Bigl(d\ln\frac{34em}{d}\log_2(578m)+\ln\frac4\delta\Bigr)} .erP​(h)<erzγ​(h)+m2​(dlnd34em​log2​(578m)+lnδ4​)​.

Milestones, in the order the proof uses them

  1. Lemma 4. er⁡P(h)<er⁡^zγ(h)+(2/m)ln⁡(2N∞(πγ(H),γ/2,2m)/δ)\operatorname{er}_P(h)<\widehat{\operatorname{er}}{}^{\gamma}_z(h)+\sqrt{(2/m)\ln(2\mathcal N_\infty(\pi_\gamma(H),\gamma/2,2m)/\delta)}erP​(h)<erzγ​(h)+(2/m)ln(2N∞​(πγ​(H),γ/2,2m)/δ)​ uniformly over HHH, with probability at least 1−δ1-\delta1−δ.
  2. Theorem 5 (Alon et al.). If F:{1,…,n}→{1,…,b}F:\{1,\dots,n\}\to\{1,\dots,b\}F:{1,…,n}→{1,…,b} and fat⁡F(1)≤d\operatorname{fat}_F(1)\le dfatF​(1)≤d, then log⁡2N∞(F,2,n)<1+log⁡2(nb2)log⁡2∑i≤d(ni)bi\log_2\mathcal N_\infty(F,2,n)<1+\log_2(nb^2)\log_2\sum_{i\le d}\binom ni b^ilog2​N∞​(F,2,n)<1+log2​(nb2)log2​∑i≤d​(in​)bi, provided nnn is large enough.
  3. Writing F=Qγ/8(πγ(H))F=Q_{\gamma/8}(\pi_\gamma(H))F=Qγ/8​(πγ​(H)): fat⁡F(γ/8)≤fat⁡πγ(H)(γ/16)\operatorname{fat}_F(\gamma/8)\le\operatorname{fat}_{\pi_\gamma(H)}(\gamma/16)fatF​(γ/8)≤fatπγ​(H)​(γ/16).
  4. M∞(πγ(H),γ/2,2m)≤M∞(F,γ/2,2m)\mathcal M_\infty(\pi_\gamma(H),\gamma/2,2m)\le\mathcal M_\infty(F,\gamma/2,2m)M∞​(πγ​(H),γ/2,2m)≤M∞​(F,γ/2,2m).
  5. N∞(πγ(H),γ/2,2m)≤N∞(F,γ/4,2m)\mathcal N_\infty(\pi_\gamma(H),\gamma/2,2m)\le\mathcal N_\infty(F,\gamma/4,2m)N∞​(πγ​(H),γ/2,2m)≤N∞​(F,γ/4,2m).
  6. log⁡2N∞(πγ(H),γ/2,2m)<1+dlog⁡2(34em/d)log⁡2(578m)\log_2\mathcal N_\infty(\pi_\gamma(H),\gamma/2,2m)<1+d\log_2(34em/d)\log_2(578m)log2​N∞​(πγ​(H),γ/2,2m)<1+dlog2​(34em/d)log2​(578m) when 1≤d≤2m1\le d\le 2m1≤d≤2m and m≥dlog⁡2(34em/d)+1m\ge d\log_2(34em/d)+1m≥dlog2​(34em/d)+1.
  7. fat⁡πγ(H)(γ/16)≤fat⁡H(γ/16)\operatorname{fat}_{\pi_\gamma(H)}(\gamma/16)\le\operatorname{fat}_H(\gamma/16)fatπγ​(H)​(γ/16)≤fatH​(γ/16).

A further item, Proposition 8 (p. 529), is the probabilistic device the paper uses to make such bounds uniform over γ\gammaγ.

Significance

Theorem 2 is the bound behind the paper's main message. Corollary 9 makes it uniform over γ\gammaγ, and Theorem 28 combines it with fat-shattering estimates for networks with bounded weights. Together they show that a network classifying the training data with a large margin generalizes at a rate governed by the size of its weights, not by its number of weights. The same template, a margin error plus a capacity term at scale γ\gammaγ, underlies later margin analyses of support vector machines and boosting.

All results in this mission are proved in the literature; none is open. None is machine-checked on this platform: the platform has Rademacher-complexity margin bounds, but no statement about fat-shattering dimension or ℓ∞\ell_\inftyℓ∞​ sample covering numbers of real-valued classes. A complete formalization would provide a reusable library of these objects, with their basic inequalities between squashing, quantization, packing and covering. It would also give a checked version of the explicit constants 34em/d34em/d34em/d and 578m578m578m, which differ from those in later textbook treatments.

Difficulty

The bound is uniform over a possibly uncountable class HHH, so a union bound over hypotheses does not apply. The obvious replacement is a union bound over a cover of HHH. Two steps make it hard:

  • Lemma 4. It needs a ghost-sample symmetrization and a random-swap argument, carried out with an ℓ∞\ell_\inftyℓ∞​ cover of the squashed class on the double sample, so the cover depends on the data.
  • Theorem 5. Bounding that covering number by the fat-shattering dimension is a combinatorial counting argument about strongly shattered pairs. It is the scale-sensitive analogue of the Sauer–Shelah lemma, and here the bookkeeping of constants is exact.

The quantization steps look routine but carry the factor-of-two losses that produce the constants γ/16\gamma/16γ/16, 171717 and 578578578.

Formalization scope

The model is in the namespace BartlettNN.Margin.

  • Labels and samples. Labels are Bool, read as ±1\pm1±1 through pm (true is +1+1+1). sgn⁡(0)=1\operatorname{sgn}(0)=1sgn(0)=1. Samples are functions Fin m → X × Bool, indexed from 000, with law Measure.pi (fun _ => P). The margin estimate uses the strict inequality yih(xi)<γy_ih(x_i)<\gammayi​h(xi​)<γ, and shattering uses ≥γ\ge\gamma≥γ.
  • Fat-shattering dimension. fat⁡\operatorname{fat}fat is valued in ℕ∞. A ℕ-valued supremum would be 000 on an unbounded set, so the goal assumes fat H (γ/16) = d with d : ℕ.
  • Covering and packing numbers. Covers are finite and external (centres are arbitrary functions), the cover inequality is strict, and covering numbers are ⊤ when no finite cover exists. N∞\mathcal N_\inftyN∞​ and M∞\mathcal M_\inftyM∞​ are suprema over all samples, with repetitions allowed. "α\alphaα-separated", which the paper leaves undefined, is read as distance ≥α\ge\alpha≥α.
  • Logarithms. ln⁡\lnln is Real.log, log⁡2\log_2log2​ is Real.logb 2, and eee is Real.exp 1.
  • High probability. "With probability at least 1−δ1-\delta1−δ, every hhh" bounds the measure of the event that some h∈Hh\in Hh∈H violates the inequality. It is not a per-hypothesis statement.

Measurability. The paper states "we ignore issues of measurability, and assume that all sets considered are measurable" (p. 526). This is made explicit, not removed, through three hypotheses:

  • every h∈Hh\in Hh∈H is measurable;
  • the bad events {z:∃h∈H, er⁡P(h)≥er⁡^zγ(h)+ϵ}\{z:\exists h\in H,\ \operatorname{er}_P(h)\ge\widehat{\operatorname{er}}{}^{\gamma}_z(h)+\epsilon\}{z:∃h∈H, erP​(h)≥erzγ​(h)+ϵ} are measurable;
  • the double-sample events of display (1) are measurable.

Replacing these by countability of HHH would weaken the theorem.

Corrections of the printed text.

  • Theorem 2. The goal adds d≤34md\le 34md≤34m. Beyond 34m34m34m the term dln⁡(34em/d)d\ln(34em/d)dln(34em/d) decreases, vanishes at d=34emd=34emd=34em and then turns negative, and the printed statement fails for rich classes. Within this range nothing is lost: the proof covers d≤2md\le2md≤2m, and for 2m<d≤34m2m<d\le34m2m<d≤34m the bound exceeds 111.
  • Milestone 6. It carries the hypothesis d≤2md\le 2md≤2m, the range of the binomial estimate behind 34em/d34em/d34em/d.
  • Milestone 3. Its printed justification ∣Qγ/8(a)−Qγ/8(b)∣<∣a−b∣+γ/16|Q_{\gamma/8}(a)-Q_{\gamma/8}(b)|<|a-b|+\gamma/16∣Qγ/8​(a)−Qγ/8​(b)∣<∣a−b∣+γ/16 is false; the correct term is γ/8\gamma/8γ/8. The milestone's conclusion is true as printed, and only the conclusion is formalized.

Trivializing formalizations, ruled out. The following would each make the statements empty or different, and none is used:

  • a ℕ-valued fat dimension or covering number;
  • Real.sign in place of sgn⁡\operatorname{sgn}sgn;
  • a per-hypothesis probability bound;
  • an unrestricted ddd, which makes ⋅\sqrt{\cdot}⋅​ of a negative number equal to 000;
  • a covering number that is 000 on classes without finite covers.

Infrastructure that a complete development needs, and contributions that are welcome:

  • product measures and Hoeffding's inequality, which Mathlib has;
  • a symmetrization (ghost-sample) lemma for margin events;
  • the combinatorics of Theorem 5;
  • the elementary inequalities between packing and covering numbers.

The covering/packing and fat-shattering lemmas apply beyond this mission. Proofs of individual milestones, or of Theorem 5 in the generality of Alon et al., are useful contributions in their own right.

Selected references

  • P. L. Bartlett, The Sample Complexity of Pattern Classification with Neural Networks: The Size of the Weights is More Important than the Size of the Network, IEEE Trans. Inform. Theory 44(2), 525–536, 1998. https://doi.org/10.1109/18.661502
  • N. Alon, S. Ben-David, N. Cesa-Bianchi, D. Haussler, Scale-sensitive dimensions, uniform convergence, and learnability, J. ACM 44(4), 615–631, 1997. https://doi.org/10.1145/263867.263927
  • J. Shawe-Taylor, P. L. Bartlett, R. C. Williamson, M. Anthony, Structural risk minimization over data-dependent hierarchies, IEEE Trans. Inform. Theory 44(5), 1926–1940, 1998. https://doi.org/10.1109/18.705570
  • M. J. Kearns, R. E. Schapire, Efficient distribution-free learning of probabilistic concepts, J. Comput. Syst. Sci. 48(3), 464–497, 1994. https://doi.org/10.1016/S0022-0000(05)80062-5
  • V. N. Vapnik, A. Ya. Chervonenkis, On the uniform convergence of relative frequencies of events to their probabilities, Theory Probab. Appl. 16(2), 264–280, 1971. https://doi.org/10.1137/1116025
12 thms2 active usersReviewed
Convex OptimizationMachine LearningProbability+1·Captain: mikedeng1

The Power of Convex Relaxation: Near-Optimal Matrix Completion II: Exact Nuclear-Norm Recovery from Nearly Minimally Many EntriesResearch Paper

Motivation

Many data sets are large matrices of which only a small fraction of the entries is observed, and of which the underlying object is believed to have low rank: user–item rating tables in collaborative filtering, distance matrices in sensor-network localization, and measurement matrices in structure-from-motion. Matrix completion asks when the missing entries can be recovered exactly. Rank minimization subject to the observed entries is intractable in general. Its convex relaxation, nuclear-norm minimization, is a semidefinite program, and the question is how many randomly placed entries it needs.

Timeline:

  • 2008–2009. Candès and Recht (arXiv:0805.4471) proved that nuclear-norm minimization recovers an incoherent n×nn\times nn×n matrix of rank rrr from about μ0n6/5rlog⁡n\mu_0 n^{6/5} r\log nμ0​n6/5rlogn uniformly sampled entries, and from n5/4n^{5/4}n5/4 in the low-rank regime. They also showed that about μ0nrlog⁡n\mu_0 nr\log nμ0​nrlogn entries are necessary for any method.
  • 2010. Candès and Tao (doi:10.1109/TIT.2010.2044061), the source of this mission, closed most of the gap. Under a strong incoherence assumption, Cμ2nrlog⁡6nC\mu^2 nr\log^6 nCμ2nrlog6n entries suffice (Theorem 1.2), within a polylogarithmic factor of the information-theoretic limit, which the same paper sharpens (Theorem 1.7).
  • 2009–2011. Keshavan, Montanari and Oh (arXiv:0901.3150) obtained comparable bounds for a non-convex method. Gross (arXiv:0910.1879) and Recht (arXiv:0910.0651) later gave much shorter proofs of an O(μ0nrlog⁡2n)O(\mu_0 nr\log^2 n)O(μ0​nrlog2n) bound under a different incoherence condition, using matrix Bernstein inequalities and a "golfing" construction of the dual certificate.

Setting

Fix M∈Rn×nM \in \mathbb{R}^{n\times n}M∈Rn×n of rank rrr with singular value decomposition M=∑k=1rσkukvk∗M = \sum_{k=1}^r\sigma_k u_kv_k^*M=∑k=1r​σk​uk​vk∗​, where σk>0\sigma_k>0σk​>0 and {uk}\{u_k\}{uk​}, {vk}\{v_k\}{vk​} are orthonormal. Let PU=∑kukuk∗P_U = \sum_k u_ku_k^*PU​=∑k​uk​uk∗​, PV=∑kvkvk∗P_V = \sum_k v_kv_k^*PV​=∑k​vk​vk∗​, and let E=∑kukvk∗E = \sum_k u_kv_k^*E=∑k​uk​vk∗​ be the sign matrix. The tangent space TTT at MMM is the image of the projection

PT(X)=PUX+XPV−PUXPV,\mathcal{P}_T(X) = P_UX + XP_V - P_UXP_V,PT​(X)=PU​X+XPV​−PU​XPV​,

and PT⊥=I−PT\mathcal{P}_{T^\perp} = \mathcal{I} - \mathcal{P}_TPT⊥​=I−PT​.

MMM obeys the strong incoherence property with parameter μ\muμ if every entry of PUP_UPU​ and PVP_VPV​ is within μr/n\mu\sqrt r/nμr​/n of the corresponding entry of (r/n)I(r/n)I(r/n)I, and every entry of EEE is at most μr/n\mu\sqrt r/nμr​/n in absolute value.

An observation set Ω⊆[n]×[n]\Omega \subseteq [n]\times[n]Ω⊆[n]×[n] is either a uniformly random mmm-subset (the uniform model) or contains each entry independently with probability p=m/n2p = m/n^2p=m/n2 (the Bernoulli model). PΩ\mathcal{P}_\OmegaPΩ​ keeps the entries in Ω\OmegaΩ and zeroes the rest. The program is

minimize ∥X∥∗ subject to PΩ(X)=PΩ(M),(I.3)\text{minimize } \|X\|_* \text{ subject to } \mathcal{P}_\Omega(X) = \mathcal{P}_\Omega(M), \qquad \text{(I.3)}minimize ∥X∥∗​ subject to PΩ​(X)=PΩ​(M),(I.3)

where ∥X∥∗\|X\|_*∥X∥∗​ is the sum of the singular values.

The analysis uses the centered operators QΩ=p−1PΩ−I\mathcal{Q}_\Omega = p^{-1}\mathcal{P}_\Omega - \mathcal{I}QΩ​=p−1PΩ​−I and QT=PT−ρ′I\mathcal{Q}_T = \mathcal{P}_T - \rho'\mathcal{I}QT​=PT​−ρ′I, where ρ=r/n\rho = r/nρ=r/n and ρ′=2ρ−ρ2\rho' = 2\rho-\rho^2ρ′=2ρ−ρ2. It also uses the random matrices (QΩQT)kQΩ(E)(\mathcal{Q}_\Omega\mathcal{Q}_T)^k\mathcal{Q}_\Omega(E)(QΩ​QT​)kQΩ​(E), where the operator is applied to EEE from the right. ∥⋅∥\|\cdot\|∥⋅∥ denotes the spectral norm.

Formalization targets

Goal: Theorem 1.2 (Matrix Completion II)

There is an absolute constant C>0C>0C>0 such that, for every fixed MMM as above and m≤n2m \le n^2m≤n2 uniformly sampled entries,

m≥Cμ2nrlog⁡6n  ⟹  Pr⁡[M is the unique solution of (I.3)]≥1−n−3.m \ge C\mu^2 nr\log^6 n \implies \Pr\bigl[M \text{ is the unique solution of (I.3)}\bigr] \ge 1 - n^{-3}.m≥Cμ2nrlog6n⟹Pr[M is the unique solution of (I.3)]≥1−n−3.

The constant CCC is not fixed; the goal asserts only its existence.

Milestones (in attack order)

  1. Lemma 3.1. A dual certificate YYY with PΩ(Y)=Y\mathcal{P}_\Omega(Y)=YPΩ​(Y)=Y, PT(Y)=E\mathcal{P}_T(Y)=EPT​(Y)=E, ∥PT⊥(Y)∥<1\|\mathcal{P}_{T^\perp}(Y)\|<1∥PT⊥​(Y)∥<1, together with injectivity of PΩ\mathcal{P}_\OmegaPΩ​ on TTT, implies unique recovery. This is already proved on the platform.
  2. Theorem 3.2 (Rudelson selection estimate). With probability at least 1−3n−β1-3n^{-\beta}1−3n−β,
p−1∥PTPΩPT−pPT∥≤CRμ0nrβlog⁡n/m,p^{-1}\|\mathcal{P}_T\mathcal{P}_\Omega\mathcal{P}_T - p\mathcal{P}_T\| \le C_R\sqrt{\mu_0nr\beta\log n/m},p−1∥PT​PΩ​PT​−pPT​∥≤CR​μ0​nrβlogn/m​,

provided the right-hand side is below 111. 3. Lemma 8.1. An exact expansion of (QΩPT)kQΩ(\mathcal{Q}_\Omega\mathcal{P}_T)^k\mathcal{Q}_\Omega(QΩ​PT​)kQΩ​ in powers of QΩQT\mathcal{Q}_\Omega\mathcal{Q}_TQΩ​QT​ with explicit recursive coefficients. 4. Lemma 8.2. The coefficients are at most λ⌈(k−j)/2⌉4k\lambda^{\lceil (k-j)/2\rceil}4^kλ⌈(k−j)/2⌉4k, with λ=ρ′/p\lambda = \rho'/pλ=ρ′/p. 5. Lemma 3.3. On the event ∥(QΩQT)kQΩ(E)∥≤σ(k+1)/2\|(\mathcal{Q}_\Omega\mathcal{Q}_T)^k\mathcal{Q}_\Omega(E)\| \le \sigma^{(k+1)/2}∥(QΩ​QT​)kQΩ​(E)∥≤σ(k+1)/2, the same terms with PT\mathcal{P}_TPT​ obey the bound with an extra factor 1+4k+11+4^{k+1}1+4k+1. 6. Theorem 3.6 (Moment bound II). Let A=(QΩQT)kQΩ(E)A = (\mathcal{Q}_\Omega\mathcal{Q}_T)^k\mathcal{Q}_\Omega(E)A=(QΩ​QT​)kQΩ​(E) and rμ=μ2rr_\mu = \mu^2 rrμ​=μ2r. Then

Etrace⁡((A∗A)j)≤n(C(j(k+1))6nrμ/m)j(k+1).\mathbb{E}\operatorname{trace}\bigl((A^*A)^j\bigr) \le n\bigl(C(j(k+1))^6nr_\mu/m\bigr)^{j(k+1)}.Etrace((A∗A)j)≤n(C(j(k+1))6nrμ​/m)j(k+1).
  1. Corollary 3.7. Under (I.12), with probability at least 1−n−31-n^{-3}1−n−3 the certificate (III.10) exists and has ∥PT⊥(Y)∥≤1/2\|\mathcal{P}_{T^\perp}(Y)\|\le 1/2∥PT⊥​(Y)∥≤1/2.

Significance

Theorem 1.2 shows that a polynomial-time convex program recovers an incoherent low-rank matrix from a number of entries that is linear in nrnrnr and within a polylogarithmic factor of what any method requires. It turned nuclear-norm minimization from a heuristic into a method with near-optimal guarantees, and much of the later work on low-rank recovery, robust PCA and phase retrieval uses its framework of dual certificates, tangent spaces and incoherence.

The theorem is proved; formalizing it is the remaining work here. None of these results has a machine-checked proof. The platform already has the Candès–Recht definitions (nuclear norm, SVD data, Bernoulli model, tangent projection), the deterministic Lemma 3.1, and the Bernoulli-to-uniform transfer. This mission adds:

  • the trace-moment bound, which is the combinatorial core of the paper;
  • the deterministic operator algebra of Appendix A;
  • the assembly into the main theorem.

Shorter later proofs (Gross, Recht) use a different incoherence condition. A formal proof of the goal along either route is welcome, provided it proves the statement as given.

Difficulty

The obvious approach bounds each term ∥(QΩPT)kQΩ(E)∥\|(\mathcal{Q}_\Omega\mathcal{P}_T)^k\mathcal{Q}_\Omega(E)\|∥(QΩ​PT​)kQΩ​(E)∥ of the Neumann series for the certificate separately, using noncommutative Khintchine inequalities and decoupling. This is what Candès and Recht did, and it fails beyond small kkk: the entries of these matrices are coupled through the same random indicators, and the bounds degrade with kkk. That is where their n6/5n^{6/5}n6/5 comes from.

The moment method avoids this but has its own obstruction. Taking absolute values inside the expansion of Etrace⁡(A∗A)j\mathbb{E}\operatorname{trace}(A^*A)^jEtrace(A∗A)j loses a factor of rrr, which gives the quadratic dependence of Theorem 1.1. The linear bound needs sign cancellations among the coefficients of QT\mathcal{Q}_TQT​ to be tracked through a nested induction over "generalized spider" configurations (Section VI). Replacing PT\mathcal{P}_TPT​ by QT\mathcal{Q}_TQT​ (Lemma 3.3) is necessary for those cancellations. Without it the diagonal coefficients are of size r/nr/nr/n instead of r/n\sqrt r/nr​/n.

Formalization scope

  • Objects. Matrices are Matrix (Fin n) (Fin n) ℝ (MatrixCompletion.RealMatrix). The SVD is the platform structure SVD M r. Logarithms are natural. Probabilities are the platform's finite sums: successProb (uniform mmm-subsets), bernoulliEventProb and bernoulliExpectation. The spectral norm is spectralNorm. The definitions of matrix_completion_{basic,svd,bernoulli,tangent} are reused, not restated.
  • Square case. Theorem 1.2 is printed "under the same hypotheses as in Theorem 1.1", for n1×n2n_1\times n_2n1​×n2​ matrices. The paper proves only n1=n2=nn_1=n_2=nn1​=n2​=n (Section I-H), and the goal and milestones 3–7 are square. Theorem 3.2 is quoted from Candès–Recht and is stated rectangular, as printed.
  • Rank. "The same hypotheses" is read as the matrix hypotheses (fixed MMM, strong incoherence, uniform sampling), not as r=O(1)r = O(1)r=O(1): (I.12) carries rrr, the paper calls the result general and nonasymptotic, and Section VI never uses bounded rank. The goal holds for every rrr.
  • Constants. Every "numerical constant" (CCC, CRC_RCR​, c0c_0c0​) and every O(⋅)O(\cdot)O(⋅) is an existential absolute constant quantified before all other variables. The goal's CCC absorbs the standing assumptions n≥C′n \ge C'n≥C′ and m≥2nrm\ge 2nrm≥2nr. Where a milestone needs (I.22), 2nr≤m2nr\le m2nr≤m is an explicit hypothesis, and m≤n2m\le n^2m≤n2 is explicit wherever a probability or p≤1p\le 1p≤1 appears.
  • Correction of Theorem 3.6. The printed bound (III.27) omits the factor nnn and the O(1)j(k+1)O(1)^{j(k+1)}O(1)j(k+1) constant of the paper's own final display (p. 2070), and as printed it is false: for k=0k=0k=0, j=1j=1j=1 and a flat rank-one matrix, the left side exceeds the right by the factor n(1−p)n(1-p)n(1−p). The formal statement is the bound the paper derives, n (C(j(k+1))6nrμ/m)j(k+1)n\,(C(j(k+1))^6nr_\mu/m)^{j(k+1)}n(C(j(k+1))6nrμ​/m)j(k+1), under nrμ≤mnr_\mu\le mnrμ​≤m, which that derivation uses and which (I.12) implies. The milestone text is kept verbatim.
  • Deterministic lemmas. Lemmas 3.3, 8.1 and 8.2 hold for every fixed Ω\OmegaΩ. The event (III.18) is a hypothesis, not a probability.
  • Certificate. YYY of (III.10) exists only when PΩ\mathcal{P}_\OmegaPΩ​ is injective on TTT, so Corollary 3.7's event includes injectivity. YYY is characterized as the minimum-Frobenius-norm solution of PΩ(Y)=Y\mathcal{P}_\Omega(Y)=YPΩ​(Y)=Y, PT(Y)=E\mathcal{P}_T(Y)=EPT​(Y)=E (p. 2061).
  • Ruling out trivialization. The hypothesis m≤n2m\le n^2m≤n2 is there only because successProb is 000 for m>n2m>n^2m>n2; it does not exclude any case the paper covers. The failure probability stays n−3n^{-3}n−3 and is not traded for a constant. The constant CCC may not depend on nnn, rrr, μ\muμ or MMM, so it cannot be chosen to make (I.12) unsatisfiable. For fixed CCC, (I.12) is satisfiable with m≤n2m \le n^2m≤n2 for every large nnn and every r≤n/(Cμ2log⁡6n)r \le n/(C\mu^2\log^6 n)r≤n/(Cμ2log6n).
  • Not covered. Proposition 6.1 (the summand bound on generalized spiders) is the heart of Theorem 3.6. It needs the admissible-quadruplet combinatorics of Sections IV–VI as definitions, and is left to solvers as a lemma of their own. Contributions formalizing Sections IV–VI (the moment expansion (IV.10), admissible pairs, the cancellation identities (VI.1)–(VI.4)) are welcome and reusable for mission I of this series.

Selected references

  • E. J. Candès and T. Tao, The Power of Convex Relaxation: Near-Optimal Matrix Completion, IEEE Trans. Inf. Theory 56(5):2053–2080, 2010. https://doi.org/10.1109/TIT.2010.2044061
  • E. J. Candès and B. Recht, Exact Matrix Completion via Convex Optimization, Found. Comput. Math. 9:717–772, 2009. https://arxiv.org/abs/0805.4471
  • R. H. Keshavan, A. Montanari and S. Oh, Matrix Completion from a Few Entries, IEEE Trans. Inf. Theory 56(6):2980–2998, 2010. https://arxiv.org/abs/0901.3150
  • D. Gross, Recovering Low-Rank Matrices from Few Coefficients in Any Basis, IEEE Trans. Inf. Theory 57(3):1548–1566, 2011. https://arxiv.org/abs/0910.1879
  • B. Recht, A Simpler Approach to Matrix Completion, J. Mach. Learn. Res. 12:3413–3430, 2011. https://arxiv.org/abs/0910.0651
17 thms2 active usersReviewed
Machine LearningProbability·Captain: mikedeng1

Adversarially Robust Generalization Requires More Data 2: A Robust-Error Lower Bound for Linear Classifiers in the Bernoulli ModelResearch Paper

Motivation

Classifiers trained by standard methods reach high accuracy on image benchmarks and yet change their prediction under perturbations of each pixel that are invisible to a human. Training against such perturbations (adversarial training) improves robustness, but on CIFAR10 and SVHN the robust test accuracy stays far below the robust training accuracy: robust models overfit. Schmidt, Santurkar, Tsipras, Talwar and Mądry (arXiv:1804.11285, 2018) asked whether this gap is a failure of current methods or an information-theoretic fact about the number of samples needed. They introduced two simple data models in which a single sample suffices for standard accuracy and proved that robust accuracy needs many more samples.

This mission formalizes their lower bound for the second model, the Bernoulli model on the hypercube, which was designed to resemble MNIST (whose images are close to binary). In this model the lower bound holds for linear classifiers, and the paper shows separately that a non-linear classifier (thresholding followed by a linear rule) escapes it. The result therefore isolates a concrete way in which the model class, not only the amount of data, governs robust generalization.

Setting

Let d≥0d\ge0d≥0 and τ>0\tau>0τ>0. Points are x∈{±1}d⊂Rdx\in\{\pm1\}^d\subset\mathbb R^dx∈{±1}d⊂Rd, labels y∈{±1}y\in\{\pm1\}y∈{±1}. For a parameter θ⋆∈{±1}d\theta^\star\in\{\pm1\}^dθ⋆∈{±1}d, the (θ⋆,τ)(\theta^\star,\tau)(θ⋆,τ)-Bernoulli model draws yyy uniformly from {±1}\{\pm1\}{±1} and then, independently for every coordinate iii, sets xi=yθi⋆x_i=y\theta^\star_ixi​=yθi⋆​ with probability 12+τ\tfrac12+\tau21​+τ and xi=−yθi⋆x_i=-y\theta^\star_ixi​=−yθi⋆​ with probability 12−τ\tfrac12-\tau21​−τ (Definition 7). The two classes are noisy copies of the opposite vertices ±θ⋆\pm\theta^\star±θ⋆.

The adversary may move a test point anywhere in the ℓ∞\ell_\inftyℓ∞​ ball

B∞ε(x)={x′∈Rd:∥x′−x∥∞≤ε},\mathcal B^\varepsilon_\infty(x)=\{x'\in\mathbb R^d:\|x'-x\|_\infty\le\varepsilon\},B∞ε​(x)={x′∈Rd:∥x′−x∥∞​≤ε},

leaving the hypercube. The ℓ∞ε\ell_\infty^\varepsilonℓ∞ε​-robust classification error of a classifier f:Rd→{±1}f:\mathbb R^d\to\{\pm1\}f:Rd→{±1} is (Definition 3)

β(f)=Pr⁡(x,y)[∃x′∈B∞ε(x): f(x′)≠y].\beta(f)=\Pr_{(x,y)}\big[\exists x'\in\mathcal B^\varepsilon_\infty(x):\ f(x')\ne y\big].β(f)=(x,y)Pr​[∃x′∈B∞ε​(x): f(x′)=y].

A linear classifier is fw(x)=sgn⁡⟨w,x⟩f_w(x)=\operatorname{sgn}\langle w,x\ranglefw​(x)=sgn⟨w,x⟩ for w∈Rdw\in\mathbb R^dw∈Rd. A linear-classifier learning algorithm gng_ngn​ is any function from nnn labelled samples to a weight vector w∈Rdw\in\mathbb R^dw∈Rd.

The lower bound is Bayesian: θ⋆\theta^\starθ⋆ is drawn uniformly from {±1}d\{\pm1\}^d{±1}d, the learner receives nnn independent samples SSS from the (θ⋆,τ)(\theta^\star,\tau)(θ⋆,τ)-model, outputs w=gn(S)w=g_n(S)w=gn​(S), and is charged the robust error of fwf_wfw​ on a fresh sample, averaged over θ⋆\theta^\starθ⋆ and SSS. The posterior mean E[θi⋆∣S]=Pr⁡[θi⋆=+1∣S]−Pr⁡[θi⋆=−1∣S]\mathbb E[\theta^\star_i\mid S]=\Pr[\theta^\star_i=+1\mid S]-\Pr[\theta^\star_i=-1\mid S]E[θi⋆​∣S]=Pr[θi⋆​=+1∣S]−Pr[θi⋆​=−1∣S] measures how much the learner can know about coordinate iii.

Formalization targets

Goal: Theorem 31 (p. 35)

For 0<τ≤140<\tau\le\tfrac140<τ≤41​, 0≤ε<3τ0\le\varepsilon<3\tau0≤ε<3τ, 0<γ<120<\gamma<\tfrac120<γ<21​ and every linear learner gng_ngn​: if

n≤ε2γ25000 τ4log⁡(4d/γ),n\le\frac{\varepsilon^2\gamma^2}{5000\,\tau^4\log(4d/\gamma)},n≤5000τ4log(4d/γ)ε2γ2​,

then

Eθ⋆,S[β(fgn(S))]≥12−γ.\mathbb E_{\theta^\star,S}\big[\beta(f_{g_n(S)})\big]\ge\tfrac12-\gamma .Eθ⋆,S​[β(fgn​(S)​)]≥21​−γ.

Milestones

  1. Eqs. (4)–(5), p. 33: in one dimension the posterior odds of θ\thetaθ equal ∏k(1/2+τ1/2−τ)ykxk\prod_k\big(\tfrac{1/2+\tau}{1/2-\tau}\big)^{y_kx_k}∏k​(1/2−τ1/2+τ​)yk​xk​.
  2. Lemma 29, p. 33: for τ≤14\tau\le\tfrac14τ≤41​ and n≤1/τ2n\le1/\tau^2n≤1/τ2, with probability 1−δ1-\delta1−δ,
∣log⁡Pr⁡[θ=+1∣S]Pr⁡[θ=−1∣S]∣≤15τ2nlog⁡(2/δ).\Big|\log\tfrac{\Pr[\theta=+1\mid S]}{\Pr[\theta=-1\mid S]}\Big|\le15\tau\sqrt{2n\log(2/\delta)} .​logPr[θ=−1∣S]Pr[θ=+1∣S]​​≤15τ2nlog(2/δ)​.
  1. Proof of Theorem 31, p. 36: with probability 1−γ/21-\gamma/21−γ/2, ∣E[θi⋆∣S]∣≤15τ2nlog⁡(4d/γ)|\mathbb E[\theta^\star_i\mid S]|\le15\tau\sqrt{2n\log(4d/\gamma)}∣E[θi⋆​∣S]∣≤15τ2nlog(4d/γ)​ for all iii.
  2. §4, p. 10: sup⁡∥Δ∥∞≤ε⟨yw,Δ⟩=ε∥w∥1\sup_{\|\Delta\|_\infty\le\varepsilon}\langle yw,\Delta\rangle=\varepsilon\|w\|_1sup∥Δ∥∞​≤ε​⟨yw,Δ⟩=ε∥w∥1​, so www robustly classifies (x,y)(x,y)(x,y) iff ⟨yw,x⟩>ε∥w∥1\langle yw,x\rangle>\varepsilon\|w\|_1⟨yw,x⟩>ε∥w∥1​.
  3. Proof of Theorem 31, p. 37: when θ⋆\theta^\starθ⋆ has independent coordinates with means bounded by bbb in absolute value, a fresh sample satisfies ⟨w,yx⟩≤2τbγ∥w∥1\langle w,yx\rangle\le\frac{2\tau b}{\gamma}\|w\|_1⟨w,yx⟩≤γ2τb​∥w∥1​ with probability at least (1−γ)/2(1-\gamma)/2(1−γ)/2.

The goal keeps the paper's explicit constants (500050005000, 3τ3\tau3τ, log⁡(4d/γ)\log(4d/\gamma)log(4d/γ)) because Theorem 31 is itself the explicit form of the paper's asymptotic Theorem 9.

Significance

With τ≍d−1/4\tau\asymp d^{-1/4}τ≍d−1/4 a single sample already yields a linear classifier with small standard error (Theorem 8 of the paper), while Theorem 31 shows that for ε\varepsilonε of order τ\tauτ every linear learner needs on the order of d/log⁡d\sqrt d/\log dd​/logd samples to get expected robust error below 12−γ\tfrac12-\gamma21​−γ against an ℓ∞\ell_\inftyℓ∞​ adversary (the paper's Theorem 9 states this as n≤c2ε2γ2d/log⁡(d/γ)n\le c_2\varepsilon^2\gamma^2 d/\log(d/\gamma)n≤c2​ε2γ2d/log(d/γ) for τ=c1d−1/4\tau=c_1d^{-1/4}τ=c1​d−1/4). The companion upper bound (Theorem 10) shows that thresholding the input first makes one sample enough for any ε<1\varepsilon<1ε<1. Together these give a rigorous example in which robust generalization is polynomially harder than standard generalization for a model class, and in which a change of model class removes the gap.

The theorem and its proof are published and not in doubt. The platform holds no statement of this lower bound, of Lemma 29, or of the ℓ∞/ℓ1 robustness criterion for linear classifiers (searched 2026-09-26). The mission produces a checked statement of the result with every hypothesis explicit, including the tie convention and the domain of ε\varepsilonε that the printed statement leaves implicit, and a finite, measure-free encoding of a Bayesian learning lower bound that other hypercube models can reuse.

Difficulty

The obvious attempt bounds the robust error of the best classifier the learner could output, but the learner is arbitrary: it may output any www, including ones that use the samples in unusual ways. The argument must therefore hold for every function of the samples, which is why θ⋆\theta^\starθ⋆ is random and why the error is averaged over it; for a fixed θ⋆\theta^\starθ⋆ the learner gn≡θ⋆g_n\equiv\theta^\stargn​≡θ⋆ is robust and the statement is false. The technical difficulty is to pass from "the posterior of every coordinate is nearly uniform" (a statement about ddd separate one-dimensional problems) to a bound on the margin ⟨w,yx⟩\langle w,yx\rangle⟨w,yx⟩ relative to ∥w∥1\|w\|_1∥w∥1​ that holds for every www at once, uniformly in how www spreads its weight across coordinates. Concentration of ⟨w,yx⟩\langle w,yx\rangle⟨w,yx⟩ is not available for a general www (a single heavy coordinate defeats it), so only a weak, constant-probability tail bound survives, which is why the final error is 12−γ\tfrac12-\gamma21​−γ rather than close to 111.

Formalization scope

Everything is finite. Hypercube points are sign vectors Fin d → Bool, labels are Bool with true ↦ +1+1+1, and every probability is an explicit finite sum of weights; no measure theory is involved. Rd\mathbb R^dRd is EuclideanSpace ℝ (Fin d). Committed conventions:

  • The coordinates of xxx are sampled independently (the reading of "sampling each coordinate" that the paper's proofs use).
  • ∥⋅∥∞≤ε\|\cdot\|_\infty\le\varepsilon∥⋅∥∞​≤ε and ∥w∥1\|w\|_1∥w∥1​ are written coordinatewise; the adversary's ball is the ℓ∞\ell_\inftyℓ∞​ ball, not the Euclidean one.
  • fw(x)=+1f_w(x)=+1fw​(x)=+1 when ⟨w,x⟩=0\langle w,x\rangle=0⟨w,x⟩=0 (the paper's sgn⁡(0)\operatorname{sgn}(0)sgn(0) is not in {±1}\{\pm1\}{±1}).
  • The robust error is Definition 3's event ∃x′∈B∞ε(x), f(x′)≠y\exists x'\in\mathcal B^\varepsilon_\infty(x),\ f(x')\ne y∃x′∈B∞ε​(x), f(x′)=y, not the margin criterion; the equivalence is milestone 4.
  • Added hypotheses: ε≥0\varepsilon\ge0ε≥0 in the goal (for ε<0\varepsilon<0ε<0 the ball is empty and the printed statement fails at n=0n=0n=0), and δ>0\delta>0δ>0 in Lemma 29 (at δ=0\delta=0δ=0 Lean's log⁡(2/0)=0\log(2/0)=0log(2/0)=0 makes it false). Posteriors are defined by Bayes' rule as ratios of joint weights.

The learner is any function of the samples to Rd\mathbb R^dRd; restricting to a specific learner, fixing θ⋆\theta^\starθ⋆, letting the learner output an arbitrary classifier (for which the theorem is false), or bounding only the standard error (ε=0\varepsilon=0ε=0) would each trivialize or falsify the target and are excluded. Useful infrastructure: Hoeffding's inequality for sums of independent ±1\pm1±1 variables, Markov's inequality over finite sums, and the ℓ∞/ℓ1 duality on EuclideanSpace. Related platform work: the other three missions of this series (the Gaussian lower bound, the Gaussian robust upper bound, and the Bernoulli thresholding upper bound). Contributions of general lemmas on finite product measures over the hypercube are welcome.

Selected references

  • L. Schmidt, S. Santurkar, D. Tsipras, K. Talwar, A. Mądry, Adversarially Robust Generalization Requires More Data, arXiv:1804.11285v2, 2018; NeurIPS 2018. https://arxiv.org/abs/1804.11285
  • I. Goodfellow, J. Shlens, C. Szegedy, Explaining and Harnessing Adversarial Examples, ICLR 2015. https://arxiv.org/abs/1412.6572
  • A. Mądry, A. Makelov, L. Schmidt, D. Tsipras, A. Vladu, Towards Deep Learning Models Resistant to Adversarial Attacks, ICLR 2018. https://arxiv.org/abs/1706.06083
  • S. Boucheron, G. Lugosi, P. Massart, Concentration Inequalities: A Nonasymptotic Theory of Independence, Oxford University Press, 2013. https://doi.org/10.1093/acprof:oso/9780199535255.001.0001
7 thms2 active usersReviewed
Machine LearningProbability·Captain: mikedeng1

Adversarially Robust Generalization Requires More Data 1: A Robust-Error Lower Bound in the Gaussian ModelResearch Paper

Motivation

Classifiers trained to high standard accuracy on image benchmarks can be made to fail by perturbations of the input that are small in the ℓ∞\ell_\inftyℓ∞​ norm (Szegedy et al., 2014; Goodfellow et al., 2015). Adversarial training reaches high robust accuracy on the training set, but on CIFAR10 the robust accuracy on held-out data is much lower than on the training set (Madry et al., 2018). That is, robust generalization fails.

Schmidt, Santurkar, Tsipras, Talwar and Mądry (arXiv:1804.11285) ask whether this is a statistical phenomenon: does learning a robust classifier need more samples than learning an accurate one, even in the simplest distributional model? This mission formalizes their answer for a mixture of two Gaussians. In that model, with ∥θ⋆∥2=d\|\theta^\star\|_2 = \sqrt d∥θ⋆∥2​=d​ and σ≤c d1/4\sigma \le c\, d^{1/4}σ≤cd1/4, a single sample suffices for standard generalization (their Theorem 4). Robust generalization, by contrast, needs a number of samples that grows polynomially with the dimension, for every learning algorithm.

Setting

Write Rd\mathbb R^dRd for the feature space and {±1}\{\pm 1\}{±1} for the labels.

  • The ℓ∞\ell_\inftyℓ∞​ perturbation set of radius ε\varepsilonε around xxx is B∞ε(x)={x′∈Rd:∥x′−x∥∞≤ε}\mathcal B_\infty^\varepsilon(x) = \{x' \in \mathbb R^d : \|x' - x\|_\infty \le \varepsilon\}B∞ε​(x)={x′∈Rd:∥x′−x∥∞​≤ε}.
  • For θ∈Rd\theta \in \mathbb R^dθ∈Rd and σ>0\sigma > 0σ>0, the (θ,σ)(\theta, \sigma)(θ,σ)-Gaussian model Pθ,σP_{\theta,\sigma}Pθ,σ​ is the law of (x,y)(x, y)(x,y) obtained by drawing yyy uniformly from {±1}\{\pm1\}{±1} and then x∼N(yθ,σ2I)x \sim \mathcal N(y\theta, \sigma^2 I)x∼N(yθ,σ2I) (Definition 1). Here σ\sigmaσ is a standard deviation.
  • The ℓ∞ε\ell_\infty^\varepsilonℓ∞ε​-robust classification error of a classifier f:Rd→{±1}f : \mathbb R^d \to \{\pm1\}f:Rd→{±1} under a distribution PPP is P(x,y)∼P[∃ x′∈B∞ε(x):f(x′)≠y]\mathbb P_{(x,y) \sim P}[\exists\, x' \in \mathcal B_\infty^\varepsilon(x) : f(x') \ne y]P(x,y)∼P​[∃x′∈B∞ε​(x):f(x′)=y] (Definitions 2–3). With ε=0\varepsilon = 0ε=0 it is the ordinary classification error.
  • A learning algorithm gng_ngn​ maps nnn labelled samples S∈(Rd×{±1})nS \in (\mathbb R^d \times \{\pm 1\})^nS∈(Rd×{±1})n to a classifier fn=gn(S)f_n = g_n(S)fn​=gn​(S).
  • The expected robust error Ξ\XiΞ of gng_ngn​ is the robust error of gn(S)g_n(S)gn​(S) under Pθ,σP_{\theta,\sigma}Pθ,σ​, averaged over S∼Pθ,σ⊗nS \sim P_{\theta,\sigma}^{\otimes n}S∼Pθ,σ⊗n​ and then over a prior θ∼N(0,I)\theta \sim \mathcal N(0, I)θ∼N(0,I). The learner sees SSS but not θ\thetaθ.

Formalization targets

Goal: Corollary 23 (p. 30)

For every learning algorithm gng_ngn​, every σ>0\sigma > 0σ>0 and every ε≥0\varepsilon \ge 0ε≥0,

n≤ε2σ28log⁡d⟹Ξ ≥ (1−1d)12.n \le \frac{\varepsilon^2\sigma^2}{8\log d} \quad\Longrightarrow\quad \Xi \ \ge\ \Big(1 - \frac1d\Big)\frac12 .n≤8logdε2σ2​⟹Ξ ≥ (1−d1​)21​.

Theorem 11 (p. 28)

For every learning algorithm gng_ngn​, every σ>0\sigma > 0σ>0 and every ε≥0\varepsilon \ge 0ε≥0,

Ξ ≥ 12 Pv∼N(0,I)[nσ2+n ∥v∥∞≤ε].\Xi \ \ge\ \frac12\, \mathbb P_{v \sim \mathcal N(0, I)}\Big[\sqrt{\tfrac{n}{\sigma^2+n}}\,\|v\|_\infty \le \varepsilon\Big].Ξ ≥ 21​Pv∼N(0,I)​[σ2+nn​​∥v∥∞​≤ε].

Intermediate statements (milestones)

  1. Eq. (2). Given nnn samples zi∼N(θ,σ2I)z_i \sim \mathcal N(\theta, \sigma^2 I)zi​∼N(θ,σ2I), the posterior of θ∼N(0,I)\theta \sim \mathcal N(0, I)θ∼N(0,I) is N(μ′,Σ′)\mathcal N(\mu', \Sigma')N(μ′,Σ′) with μ′=(σ2+n)−1∑izi\mu' = (\sigma^2+n)^{-1}\sum_i z_iμ′=(σ2+n)−1∑i​zi​ and Σ′=σ2σ2+nI\Sigma' = \frac{\sigma^2}{\sigma^2+n} IΣ′=σ2+nσ2​I. The expectations over θ\thetaθ and over the samples may therefore be exchanged.
  2. Eq. (3). Averaging Pθ,σP_{\theta,\sigma}Pθ,σ​ over θ∼N(m,s2I)\theta \sim \mathcal N(m, s^2 I)θ∼N(m,s2I) gives Pm,s2+σ2P_{m, \sqrt{s^2+\sigma^2}}Pm,s2+σ2​​.
  3. The bound on Ψ\PsiΨ. If ∥m∥∞≤ε\|m\|_\infty \le \varepsilon∥m∥∞​≤ε, every classifier has ℓ∞ε\ell_\infty^\varepsilonℓ∞ε​-robust error at least 12\frac1221​ under Pm,sP_{m,s}Pm,s​.
  4. The law of zˉ\bar zzˉ. The sample mean zˉ\bar zzˉ of the ziz_izi​ is marginally N(0,(1+σ2/n)I)\mathcal N(0, (1+\sigma^2/n) I)N(0,(1+σ2/n)I).
  5. Maximum of ddd Gaussians. Pv∼N(0,Id)[∥v∥∞≤22log⁡d]≥1−1/d\mathbb P_{v\sim\mathcal N(0,I_d)}[\|v\|_\infty \le 2\sqrt{2\log d}] \ge 1 - 1/dPv∼N(0,Id​)​[∥v∥∞​≤22logd​]≥1−1/d.

Corollary 23 is the goal because it is the statement the paper advertises: its main-text Theorem 6 is Corollary 23 with σ=c1d1/4\sigma = c_1 d^{1/4}σ=c1​d1/4.

Significance

In the same model with ∥θ⋆∥2=d\|\theta^\star\|_2 = \sqrt d∥θ⋆∥2​=d​ and σ\sigmaσ of order d1/4d^{1/4}d1/4, a single sample suffices to reach standard error below 1% (Theorem 4 of the paper), and on the order of ε2d\varepsilon^2\sqrt dε2d​ samples suffice for robust error below 1% when ε\varepsilonε is below a small constant (Theorem 5; Corollary 22, formalized in mission 3 of this series). Corollary 23 shows that, up to the logarithmic factor, no learner can do better. Robust generalization then needs ε2d/log⁡d\varepsilon^2 \sqrt d / \log dε2d​/logd times as many samples as standard generalization. The gap is information-theoretic: it concerns every algorithm, not a particular training procedure or model class. The authors present this as a candidate explanation for the robust-generalization gap observed on CIFAR10. The ½ is tight: a constant classifier attains it.

The paper's proof is complete and short. As far as a search of the platform shows, none of its statements has been formalized. A machine-checked version requires multivariate Gaussian conjugacy, Gaussian convolution identities, the outer-measure robust event, and a union bound for the maximum of Gaussians. The Gaussian conjugacy and convolution facts are standard and appear throughout Bayesian statistics. As of this Mathlib version they are not available for stdGaussian on EuclideanSpace.

Difficulty

The obvious attempt fixes θ\thetaθ and bounds the robust error for each θ\thetaθ. That fails: a learner may ignore the data and output the Bayes-optimal robust classifier for one fixed θ\thetaθ, so for each θ\thetaθ some learner does well. The lower bound holds only on average over the prior on θ\thetaθ. The classifier fnf_nfn​ depends on the samples, and the samples depend on θ\thetaθ, so the classifier and the test distribution are correlated through θ\thetaθ. A second obstacle is measure-theoretic. The robust error of a classifier is the probability of an ℓ∞\ell_\inftyℓ∞​-thickening of an arbitrary set {f≠y}\{f \ne y\}{f=y}. Such a set need not be Borel, and it must be bounded below with no structure on the classifier beyond what the learner provides. The same statement with the ℓ2\ell_2ℓ2​ ball is a different theorem.

Formalization scope

  • Rd\mathbb R^dRd is EuclideanSpace ℝ (Fin d) with its Borel σ-algebra; N(0,I)\mathcal N(0, I)N(0,I) is Mathlib's stdGaussian; N(m,s2I)\mathcal N(m, s^2 I)N(m,s2I) is its image under v↦m+svv \mapsto m + s vv↦m+sv, so every Gaussian parameter in the development is a standard deviation.
  • The ℓ∞\ell_\inftyℓ∞​ ball is written coordinatewise (∣xi′−xi∣≤ε|x'_i - x_i| \le \varepsilon∣xi′​−xi​∣≤ε for all iii), because the ambient norm is ℓ2\ell_2ℓ2​. ∥v∥∞≤r\|v\|_\infty \le r∥v∥∞​≤r is written the same way.
  • Labels are Bool, with true for +1+1+1. A classifier is ℝ^d → Bool, and a learning algorithm is (Fin n → ℝ^d × Bool) → ℝ^d → Bool.
  • A model is a measure on Rd×\mathbb R^d \timesRd× Bool; nnn samples form the product measure Measure.pi.
  • The robust event need not be Borel. Its probability is the outer measure, which is its probability under the completion. The expectations are lower Lebesgue integrals.
  • Theorem 11 and Corollary 23 assume the learner is jointly measurable in (samples, input). This is the only condition on it. log⁡\loglog is the natural logarithm. With Lean's conventions log⁡0=log⁡1=0\log 0 = \log 1 = 0log0=log1=0 and x/0=0x/0 = 0x/0=0, the goal's hypothesis forces n=0n = 0n=0 for d≤1d \le 1d≤1, where the statement is still true.
  • The paper writes the posterior mean as nσ2+nzˉ\frac{n}{\sigma^2+n}\bar zσ2+nn​zˉ, with "zˉ=∑izi\bar z = \sum_i z_izˉ=∑i​zi​" on p. 28. The formalization uses (σ2+n)−1∑izi(\sigma^2+n)^{-1}\sum_i z_i(σ2+n)−1∑i​zi​, which is the posterior mean. It agrees with the paper when zˉ\bar zzˉ is read as the sample mean, as it is on p. 30.

A formalization that fixes θ\thetaθ instead of averaging over the prior, that restricts the learner (to linear classifiers, or to classifiers that do not depend on the data), that uses the ℓ2\ell_2ℓ2​ ball, or that assumes the posterior formula as a hypothesis states a different theorem. These are ruled out by the definitions file.

Needed infrastructure: the multivariate Gaussian conjugacy and convolution identities for stdGaussian pushforwards, translation invariance of outer measure under Gaussian shifts, and a sub-Gaussian tail bound for one coordinate. The Gaussian identities are reusable well beyond this mission. Contributions of general Gaussian lemmas, stated for stdGaussian on any finite-dimensional inner product space, are welcome. Related platform work: missions 2–4 of this series formalize the Bernoulli-model lower bound and the two upper bounds of the same paper.

Selected references

  • L. Schmidt, S. Santurkar, D. Tsipras, K. Talwar, A. Mądry, Adversarially Robust Generalization Requires More Data, NeurIPS 2018; arXiv:1804.11285v2, 2018. https://arxiv.org/abs/1804.11285
  • A. Madry, A. Makelov, L. Schmidt, D. Tsipras, A. Vladu, Towards Deep Learning Models Resistant to Adversarial Attacks, ICLR 2018. https://arxiv.org/abs/1706.06083
  • C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, R. Fergus, Intriguing properties of neural networks, ICLR 2014. https://arxiv.org/abs/1312.6199
  • I. Goodfellow, J. Shlens, C. Szegedy, Explaining and Harnessing Adversarial Examples, ICLR 2015. https://arxiv.org/abs/1412.6572
  • S. Boucheron, G. Lugosi, P. Massart, Concentration Inequalities: A Nonasymptotic Theory of Independence, Oxford University Press, 2013 (Theorem 5.8). https://doi.org/10.1093/acprof:oso/9780199535255.001.0001
8 thms2 active usersReviewed
PreviousPage 1 of 2Next

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me