Prove2Me
Navigate
DiscoverFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

Probability

550 missions · 265 completed

Missions

Open285Completed265All550
Machine LearningStatistics·Captain: mikedeng1

Certified Adversarial Robustness via Randomized Smoothing 2: The Certified ℓ2 Radius Cannot Be EnlargedResearch Paper

Motivation

Neural-network classifiers can be made to change their output by perturbations of the input that are imperceptible to a person. A certified defense is a classifier together with a proof that its prediction at a point xxx does not change for any perturbation δ\deltaδ in a stated set, typically an ℓ2\ell_2ℓ2​ ball ∥δ∥2<R\|\delta\|_2<R∥δ∥2​<R. Randomized smoothing turns an arbitrary base classifier into one with such a certificate by classifying Gaussian-noised copies of the input and returning the most likely class. Cohen, Rosenfeld and Kolter (arXiv:1902.02918v2, ICML 2019) gave the certified radius R=σ2(Φ−1(pA‾)−Φ−1(pB‾))R=\frac{\sigma}{2}\big(\Phi^{-1}(\underline{p_A})-\Phi^{-1}(\overline{p_B})\big)R=2σ​(Φ−1(pA​​)−Φ−1(pB​​)) (their Theorem 1) and showed, in their Theorem 2, that this radius cannot be enlarged when only the two class-probability bounds are known about the base classifier. This mission formalizes Theorem 2. Theorem 1 is the subject of the companion mission of this series.

Earlier certificates for the same smoothed classifier, by Lecuyer et al. (2019) via differential privacy and Li et al. (2018) via Rényi divergence, gave smaller radii. Theorem 2 shows that no further analysis that uses only the class-probability bounds can improve on Theorem 1.

Setting

Inputs live in Rd\mathbb R^dRd with the Euclidean norm ∥⋅∥2\|\cdot\|_2∥⋅∥2​; classes form a set Y\mathcal YY. A base classifier is a map f:Rd→Yf:\mathbb R^d\to\mathcal Yf:Rd→Y with Borel decision regions. For a noise level σ>0\sigma>0σ>0, write N(x,σ2I)\mathcal N(x,\sigma^2I)N(x,σ2I) for the isotropic Gaussian law of x+εx+\varepsilonx+ε with ε∼N(0,σ2I)\varepsilon\sim\mathcal N(0,\sigma^2I)ε∼N(0,σ2I). The class probability of ccc at xxx is P(f(x+ε)=c)\mathbb P(f(x+\varepsilon)=c)P(f(x+ε)=c), and the smoothed classifier is

g(x)=arg⁡max⁡c∈Y P(f(x+ε)=c).g(x)=\arg\max_{c\in\mathcal Y}\ \mathbb P(f(x+\varepsilon)=c).g(x)=argc∈Ymax​ P(f(x+ε)=c).

Let Φ\PhiΦ be the standard Gaussian CDF and Φ−1\Phi^{-1}Φ−1 its inverse on (0,1)(0,1)(0,1). A classifier fff is consistent with the observed class probabilities (6) for a top class cAc_AcA​ and numbers pA‾≥pB‾\underline{p_A}\ge\overline{p_B}pA​​≥pB​​ if

P(f(x+ε)=cA) ≥ pA‾ ≥ pB‾ ≥ max⁡c≠cAP(f(x+ε)=c).\mathbb P(f(x+\varepsilon)=c_A)\ \ge\ \underline{p_A}\ \ge\ \overline{p_B}\ \ge\ \max_{c\ne c_A}\mathbb P(f(x+\varepsilon)=c).P(f(x+ε)=cA​) ≥ pA​​ ≥ pB​​ ≥ c=cA​max​P(f(x+ε)=c).

The certified radius is R=σ2(Φ−1(pA‾)−Φ−1(pB‾))R=\frac{\sigma}{2}\big(\Phi^{-1}(\underline{p_A})-\Phi^{-1}(\overline{p_B})\big)R=2σ​(Φ−1(pA​​)−Φ−1(pB​​)). In Lean these are gaussNoise x σ, classProb f σ x c, IsConsistent f σ x cA pA pB and radius σ pA pB in the namespace Cohen2019.Tight, with Phi and PhiInvReal from the series' shared module Cohen2019.Robust; the half-spaces A={z:δT(z−x)≤σ∥δ∥Φ−1(pA‾)}A=\{z:\delta^T(z-x)\le\sigma\|\delta\|\Phi^{-1}(\underline{p_A})\}A={z:δT(z−x)≤σ∥δ∥Φ−1(pA​​)} and B={z:δT(z−x)≥σ∥δ∥Φ−1(1−pB‾)}B=\{z:\delta^T(z-x)\ge\sigma\|\delta\|\Phi^{-1}(1-\overline{p_B})\}B={z:δT(z−x)≥σ∥δ∥Φ−1(1−pB​​)} of the paper's Appendix A are setA and setB.

Quotations write the paper's underlined lower bound as p̲A and its overlined upper bound as p̄B. The PDF has no printed page numbers; every page cited is the PDF page of arXiv:1902.02918v2.

Formalization targets

Goal: Theorem 2 (corrected)

Assume 0<pB‾≤pA‾<10<\overline{p_B}\le\underline{p_A}<10<pB​​≤pA​​<1, pA‾+pB‾≤1\underline{p_A}+\overline{p_B}\le1pA​​+pB​​≤1, and that some finite set sss of classes other than cAc_AcA​ satisfies 1≤pA‾+∣s∣ pB‾1\le\underline{p_A}+|s|\,\overline{p_B}1≤pA​​+∣s∣pB​​. Then for every δ\deltaδ with ∥δ∥2>R\|\delta\|_2>R∥δ∥2​>R there is a base classifier f∗f^*f∗ consistent with (6) and a class c≠cAc\ne c_Ac=cA​ with

P(f∗(x+δ+ε)=cA) < P(f∗(x+δ+ε)=c),\mathbb P(f^*(x+\delta+\varepsilon)=c_A)\ <\ \mathbb P(f^*(x+\delta+\varepsilon)=c),P(f∗(x+δ+ε)=cA​) < P(f∗(x+δ+ε)=c),

so that g(x+δ)≠cAg(x+\delta)\ne c_Ag(x+δ)=cA​ under any tie-breaking. The classifier may depend on δ\deltaδ.

The class-capacity hypothesis is a correction. As printed, with only pA‾+pB‾≤1\underline{p_A}+\overline{p_B}\le1pA​​+pB​​≤1, the theorem fails for two classes: with Y={cA,cB}\mathcal Y=\{c_A,c_B\}Y={cA​,cB​}, pA‾=0.6\underline{p_A}=0.6pA​​=0.6, pB‾=0.1\overline{p_B}=0.1pB​​=0.1 and σ=∥δ∥2=1\sigma=\|\delta\|_2=1σ=∥δ∥2​=1, one has R≈0.767<1R\approx0.767<1R≈0.767<1, yet every consistent fff gives cAc_AcA​ probability at least 0.90.90.9, and Theorem 1 then certifies radius Φ−1(0.9)≈1.28\Phi^{-1}(0.9)\approx1.28Φ−1(0.9)≈1.28.

Milestones

The milestones are the steps the paper itself states, in its order: the Claims P(X∈A)=pA‾\mathbb P(X\in A)=\underline{p_A}P(X∈A)=pA​​ and P(X∈B)=pB‾\mathbb P(X\in B)=\overline{p_B}P(X∈B)=pB​​ for X∼N(x,σ2I)X\sim\mathcal N(x,\sigma^2I)X∼N(x,σ2I); the disjointness of AAA and BBB (corrected to "null" when pA‾+pB‾=1\underline{p_A}+\overline{p_B}=1pA​​+pB​​=1); equations (13) and (14) for Y∼N(x+δ,σ2I)Y\sim\mathcal N(x+\delta,\sigma^2I)Y∼N(x+δ,σ2I),

P(Y∈A)=Φ(Φ−1(pA‾)−∥δ∥σ),P(Y∈B)=Φ(Φ−1(pB‾)+∥δ∥σ);\mathbb P(Y\in A)=\Phi\Big(\Phi^{-1}(\underline{p_A})-\tfrac{\|\delta\|}{\sigma}\Big),\qquad \mathbb P(Y\in B)=\Phi\Big(\Phi^{-1}(\overline{p_B})+\tfrac{\|\delta\|}{\sigma}\Big);P(Y∈A)=Φ(Φ−1(pA​​)−σ∥δ∥​),P(Y∈B)=Φ(Φ−1(pB​​)+σ∥δ∥​);

the equivalence P(Y∈A)<P(Y∈B)  ⟺  ∥δ∥2>R\mathbb P(Y\in A)<\mathbb P(Y\in B)\iff\|\delta\|_2>RP(Y∈A)<P(Y∈B)⟺∥δ∥2​>R; and the existence of the worst-case classifier f∗f^*f∗ satisfying (6) with equalities.

Significance

Theorem 2 makes the guarantee of Theorem 1 exact: when only (6) is known about fff, the set of perturbations under which the Gaussian-smoothed prediction is provably constant is exactly the open ℓ2\ell_2ℓ2​ ball of radius RRR. It settles that improvements to Gaussian-smoothing certificates must use more information about the base classifier than the two bounds, as later work on higher-order and Lipschitz-based certificates does.

The paper's proof is complete in its main lines and has two gaps that this mission records and repairs: the printed statement omits a condition on the number of classes, and the claim A∩B=∅A\cap B=\emptysetA∩B=∅ fails at pA‾+pB‾=1\underline{p_A}+\overline{p_B}=1pA​​+pB​​=1. To our knowledge neither Theorem 1 nor Theorem 2 has a machine-checked proof. Mathlib at the pinned revision has the multivariate standard Gaussian but no normal quantile function and no Gaussian half-space lemma; this mission adds statements for both kinds of fact.

Difficulty

Each step is elementary on paper but rests on facts about Gaussians that Mathlib does not package: the image of the standard Gaussian on Rd\mathbb R^dRd under a linear functional z↦δTzz\mapsto\delta^T zz↦δTz is the one-dimensional Gaussian with variance ∥δ∥2\|\delta\|^2∥δ∥2, and Φ\PhiΦ is a continuous strictly increasing bijection R→(0,1)\mathbb R\to(0,1)R→(0,1) with Φ−1(1−p)=−Φ−1(p)\Phi^{-1}(1-p)=-\Phi^{-1}(p)Φ−1(1−p)=−Φ−1(p). The construction of f∗f^*f∗ has a further step the paper leaves informal: the region between AAA and BBB, of mass 1−pA‾−pB‾1-\underline{p_A}-\overline{p_B}1−pA​​−pB​​, must be shared among "other classes" with none exceeding pB‾\overline{p_B}pB​​, which is where the capacity hypothesis enters. Measurability of the constructed decision regions must be carried along.

Formalization scope

Rd\mathbb R^dRd is EuclideanSpace ℝ (Fin d). N(x,σ2I)\mathcal N(x,\sigma^2I)N(x,σ2I) is the pushforward of Mathlib's stdGaussian under z↦x+σzz\mapsto x+\sigma zz↦x+σz, with σ>0\sigma>0σ>0 a binder. Φ\PhiΦ is cdf (gaussianReal 0 1); Φ−1(p)\Phi^{-1}(p)Φ−1(p) is the generalized inverse inf⁡{t:p≤Φ(t)}\inf\{t:p\le\Phi(t)\}inf{t:p≤Φ(t)}, which is the true inverse on (0,1)(0,1)(0,1) and the junk value 000 at the endpoints, so every statement that evaluates it assumes 0<p<10<p<10<p<1; at pB‾=0\overline{p_B}=0pB​​=0 or pA‾=1\underline{p_A}=1pA​​=1 the paper's radius is infinite and Theorem 2 is vacuous. Class probabilities are real numbers. The base classifier in the conclusion is deterministic with Borel decision regions, which is the stronger existence statement. The conclusion is the strict inequality between class probabilities, not merely the failure of cAc_AcA​ to be a strict unique argmax.

A formalization in which the junk endpoint value of Φ−1\Phi^{-1}Φ−1 makes RRR negative, or in which the classifier's decision regions are non-measurable so that its class probabilities are default values, would make the goal trivial; the hypotheses above exclude both.

Reusable beyond this mission: the Gaussian half-space probabilities and the normal quantile on (0,1)(0,1)(0,1). Contributions of general Mathlib-style lemmas (the law of δTX\delta^T XδTX for X∼N(x,σ2I)X\sim\mathcal N(x,\sigma^2I)X∼N(x,σ2I), properties of Φ−1\Phi^{-1}Φ−1) are welcome.

Selected references

  • J. M. Cohen, E. Rosenfeld, J. Z. Kolter, Certified Adversarial Robustness via Randomized Smoothing, ICML 2019; arXiv:1902.02918v2. https://arxiv.org/abs/1902.02918v2
  • M. Lecuyer, V. Atlidakis, R. Geambasu, D. Hsu, S. Jana, Certified Robustness to Adversarial Examples with Differential Privacy, IEEE S&P 2019. https://arxiv.org/abs/1802.03471
  • B. Li, C. Chen, W. Wang, L. Carin, Certified Adversarial Robustness with Additive Noise, NeurIPS 2019. https://arxiv.org/abs/1809.03113
  • J. Neyman, E. S. Pearson, On the Problem of the Most Efficient Tests of Statistical Hypotheses, Phil. Trans. R. Soc. A 231, 1933. https://doi.org/10.1098/rsta.1933.0009
11 thms2 active usersReviewed
Statistics·Captain: mikedeng1

Weighted Sums of Certain Dependent Random Variables 3: Reversed Weighted Sums of Bounded Martingale Differences Obey a Strong LawResearch Paper

Motivation

A martingale difference sequence is the standard model of a "fair" sequence of observations whose terms may depend on the past: each new term has conditional mean zero given everything observed before it. Laws of large numbers for such sequences underlie the analysis of stochastic approximation, sequential estimation and online learning, where the noise terms are dependent but conditionally centred.

Classical strong laws concern averages in which every observation keeps the same weight as the sample grows. Kazuoki Azuma's 1967 paper Weighted sums of certain dependent random variables (Tôhoku Math. J. 19) studies weighted sums of dependent variables, and is best known for the exponential moment bound that is now called the Azuma inequality (its display (2.4) together with Remark 1). Its Theorem 3 uses that bound to prove a strong law for weighted averages in which the weights are applied in reverse order, so that the oldest observation always receives the newest, largest weight.

Timeline:

  • 1960s: Y. S. Chow (Ann. Math. Statist. 37, 1966) introduces a conditional exponential-moment condition close to Azuma's property [G] in a convergence theorem for independent variables.
  • 1967: Azuma proves the moment bound (2.4) for conditionally sub-Gaussian martingale differences, a law of the iterated logarithm for direct weighted sums (Theorem 2), and the strong law for reversed weighted sums (Theorem 3), the subject of this mission.

Setting

Let (Ω,A,P)(\Omega,\mathfrak A,P)(Ω,A,P) be a probability space and (An)n≥0(\mathfrak A_n)_{n\ge0}(An​)n≥0​ an increasing family of sub-σ\sigmaσ-fields of A\mathfrak AA. A sequence (xn)n≥1(x_n)_{n\ge1}(xn​)n≥1​ of real random variables is a sequence of martingale differences if, for every n≥1n\ge1n≥1, xnx_nxn​ is An\mathfrak A_nAn​-measurable and integrable and E{xn∣An−1}=0E\{x_n\mid\mathfrak A_{n-1}\}=0E{xn​∣An−1​}=0 almost surely. Theorem 3 assumes moreover ∣xn∣≤1|x_n|\le1∣xn​∣≤1 almost surely for every nnn.

Let (an)n≥1(a_n)_{n\ge1}(an​)n≥1​ be positive increasing weights: an>0a_n>0an​>0 and an≤an+1a_n\le a_{n+1}an​≤an+1​. Put

An=a1+a2+⋯+an,Sˉn=anx1+an−1x2+⋯+a1xn=∑j=1nan−j+1xj.A_n = a_1+a_2+\dots+a_n,\qquad \bar S_n = a_nx_1 + a_{n-1}x_2+\dots+a_1x_n=\sum_{j=1}^n a_{n-j+1}x_j .An​=a1​+a2​+⋯+an​,Sˉn​=an​x1​+an−1​x2​+⋯+a1​xn​=j=1∑n​an−j+1​xj​.

The sums Sˉn\bar S_nSˉn​ are the reversed weighted sums. In passing from Sˉn\bar S_nSˉn​ to Sˉn+1\bar S_{n+1}Sˉn+1​ every existing term changes its weight, so (Sˉn)(\bar S_n)(Sˉn​) is in general not a martingale. In the Lean development, AnA_nAn​ is A a n, Sˉn(ω)\bar S_n(\omega)Sˉn​(ω) is Sbar a x n ω, and the martingale-difference property is IsMartingaleDiff μ ℱ x.

Formalization targets

Goal: Theorem 3, (4.9)–(4.10)

If (xn)(x_n)(xn​) is a sequence of martingale differences with ∣xn∣≤1|x_n|\le1∣xn​∣≤1 a.s., (an)(a_n)(an​) is positive and nondecreasing, and

anAn=o(1log⁡log⁡An)(n→∞),(4.9)\frac{a_n}{A_n}=o\Big(\frac{1}{\log\log A_n}\Big)\qquad(n\to\infty),\tag{4.9}An​an​​=o(loglogAn​1​)(n→∞),(4.9)

then

SˉnAn⟶0almost surely.(4.10)\frac{\bar S_n}{A_n}\longrightarrow0\quad\text{almost surely}.\tag{4.10}An​Sˉn​​⟶0almost surely.(4.10)

Milestones

The milestones follow the paper's proof, in attack order.

  1. Remark 1 (p. 358): if ∣xn∣≤Kn|x_n|\le K_n∣xn​∣≤Kn​ a.s., then E{exp⁡(txn)∣An−1}≤cosh⁡(tKn)≤exp⁡(t2Kn2/2)E\{\exp(tx_n)\mid\mathfrak A_{n-1}\}\le\cosh(tK_n)\le\exp(t^2K_n^2/2)E{exp(txn​)∣An−1​}≤cosh(tKn​)≤exp(t2Kn2​/2) a.s.
  2. The tail step of (4.16) (p. 366): for ∣xn∣≤1|x_n|\le1∣xn​∣≤1, real c1,…,cNc_1,\dots,c_Nc1​,…,cN​ with ∑cj2>0\sum c_j^2>0∑cj2​>0 and λ≥0\lambda\ge0λ≥0,
P{∑j=1Ncjxj>λ}≤exp⁡(−λ22∑j=1Ncj2).P\Big\{\sum_{j=1}^Nc_jx_j>\lambda\Big\}\le\exp\Big(-\frac{\lambda^2}{2\sum_{j=1}^Nc_j^2}\Big).P{j=1∑N​cj​xj​>λ}≤exp(−2∑j=1N​cj2​λ2​).
  1. The blocks (4.11)–(4.14) (pp. 364–365): for every ε>0\varepsilon>0ε>0 there are indices n1<n2<⋯n_1<n_2<\cdotsn1​<n2​<⋯ with An1>2(3+ε)/(6+ε)A_{n_1}>2(3+\varepsilon)/(6+\varepsilon)An1​​>2(3+ε)/(6+ε), an/An<ε/(6+ε)a_n/A_n<\varepsilon/(6+\varepsilon)an​/An​<ε/(6+ε) and anlog⁡log⁡An/An<ε2/64a_n\log\log A_n/A_n<\varepsilon^2/64an​loglogAn​/An​<ε2/64 for n>n1n>n_1n>n1​, and Ank−1<Ank≤(1+ε/3)Ank−1<Ank+1A_{n_{k-1}}<A_{n_k}\le(1+\varepsilon/3)A_{n_{k-1}}<A_{n_k+1}Ank−1​​<Ank​​≤(1+ε/3)Ank−1​​<Ank​+1​.
  2. The maximal inequality (4.15) (p. 365): if AN1≤(1+ε/3)AN0A_{N_1}\le(1+\varepsilon/3)A_{N_0}AN1​​≤(1+ε/3)AN0​​ with 1≤N0<N11\le N_0<N_11≤N0​<N1​, then
2P{SˉN1>(ε/2)AN0}≥P{max⁡N0<n≤N1Sˉn>εAN0}.2P\{\bar S_{N_1}>(\varepsilon/2)A_{N_0}\}\ge P\Big\{\max_{N_0<n\le N_1}\bar S_n>\varepsilon A_{N_0}\Big\}.2P{SˉN1​​>(ε/2)AN0​​}≥P{N0​<n≤N1​max​Sˉn​>εAN0​​}.
  1. Block growth (p. 366): (4.11), (4.12) and (4.14) give Ank>(2(3+ε)/(6+ε))k−1A_{n_k}>(2(3+\varepsilon)/(6+\varepsilon))^{k-1}Ank​​>(2(3+ε)/(6+ε))k−1.

Significance

The result. Theorem 3 shows that reversed weighting does not destroy the strong law, under a growth condition on the weights that is strictly weaker than the condition an2/∑j≤naj2→0a_n^2/\sum_{j\le n}a_j^2\to0an2​/∑j≤n​aj2​→0 of the paper's Theorem 2: the paper reproduces an example of T. Tsuchikura satisfying (4.9) but not that condition. The condition allows rapidly growing weights, provided no single weight carries more than an o(1/log⁡log⁡An)o(1/\log\log A_n)o(1/loglogAn​) share of the total. Milestone 2 is the one-sided Azuma inequality in the paper's conditional-expectation form, a tool used throughout probability, combinatorics and learning theory. Milestone 4 is a maximal inequality for a process that is not a martingale, where Doob's inequality cannot be used.

Formalizing it. The theorem is proved in the paper; this mission produces a machine-checked proof. Mathlib contains sub-Gaussian moment-generating-function bounds for martingale differences in a kernel formulation (HasCondSubgaussianMGF, which assumes a standard Borel space), and the Prove2Me platform has two-sided Azuma–Hoeffding inequalities under the same assumption. Neither states Remark 1 or the one-sided tail bound in the paper's conditional-expectation form on an arbitrary probability space, and no strong law for reversed weighted sums is formalized.

Difficulty

The obvious route to a strong law for a martingale, Doob's maximal inequality applied along a geometric subsequence, fails at the first step: (Sˉn)(\bar S_n)(Sˉn​) is not a martingale, because each new step reweights all earlier terms. The maximum of Sˉn\bar S_nSˉn​ over a block of indices therefore needs a separate maximal inequality, and it is there that the monotonicity of the weights is indispensable. A second difficulty is quantitative: the exponential tail bound must be summable over blocks whose growth is controlled only through (4.9), which is weaker than the variance-type condition of Theorem 2, so the block sizes and the constants ε/(6+ε)\varepsilon/(6+\varepsilon)ε/(6+ε), ε2/64\varepsilon^2/64ε2/64 and 1+ε/31+\varepsilon/31+ε/3 have to be chosen to fit together.

Formalization scope

Conventions committed to in Lean:

  • Indices start at 111: sums run over Finset.Icc 1 n, and a0a_0a0​, x0x_0x0​ are never used. Every hypothesis on aaa and xxx is quantified over n≥1n\ge1n≥1.
  • The filtration is a Mathlib Filtration ℕ. Its first σ\sigmaσ-field plays the role of A0\mathfrak A_0A0​ and is arbitrary rather than trivial; the paper's A0={∅,Ω}\mathfrak A_0=\{\emptyset,\Omega\}A0​={∅,Ω} is a special case, so the formal statements are at least as general.
  • "Positive increasing" is read as an>0a_n>0an​>0 and an≤an+1a_n\le a_{n+1}an​≤an+1​ (nondecreasing), the weaker hypothesis.
  • (4.9) is stated literally as a little-ooo relation (IsLittleO along atTop). An→∞A_n\to\inftyAn​→∞ is not a hypothesis, since it follows from positivity and monotonicity.
  • The conclusion is convergence of the real sequence Sˉn(ω)/An\bar S_n(\omega)/A_nSˉn​(ω)/An​ to 000 for almost every ω\omegaω, which contains both the upper and the lower tail; a statement giving only lim sup⁡≤0\limsup\le0limsup≤0 is not the goal. No real-valued limsup is used anywhere.
  • Probabilities are real-valued (μ.real); conditional expectations are Mathlib's μ[f | ℱ n].

A formalization of the goal that replaces Sˉn\bar S_nSˉn​ by the direct sums a1x1+⋯+anxna_1x_1+\dots+a_nx_na1​x1​+⋯+an​xn​, drops the monotonicity of the weights, or strengthens (4.9) to an/An=o(1/log⁡An)a_n/A_n=o(1/\log A_n)an​/An​=o(1/logAn​) or to an2/∑j≤naj2→0a_n^2/\sum_{j\le n}a_j^2\to0an2​/∑j≤n​aj2​→0 states a different theorem and does not count.

A complete development needs the Azuma moment bound in conditional-expectation form, conditional Chebyshev arguments on events, the Borel–Cantelli lemma (in Mathlib) and elementary real analysis of the blocks. The one-sided Azuma inequality and the maximal inequality (4.15) are reusable beyond this mission. Proofs of any milestone are welcome, as are alternative proofs of the goal.

Selected references

  • K. Azuma, Weighted sums of certain dependent random variables, Tôhoku Mathematical Journal 19 (1967), 357–367. https://doi.org/10.2748/tmj/1178243286
  • Y. S. Chow, Some convergence theorems for independent random variables, Annals of Mathematical Statistics 37 (1966), 1482–1493.
  • J. L. Doob, Stochastic Processes, Wiley, New York, 1953.
8 thms2 active usersReviewed
Bandit AlgorithmsMachine LearningOperations Research·Captain: mikedeng1

Analysis of Thompson Sampling for the Multi-armed Bandit Problem 2: Logarithmic Regret for N ArmsResearch Paper

Motivation

Thompson Sampling is the oldest heuristic for the multi-armed bandit problem: proposed by Thompson in 1933, it plays each arm with the posterior probability that the arm is the best one. It is simple to implement, performs well empirically (Chapelle and Li, NIPS 2011), and has been used in production systems such as click-through-rate prediction for search advertising. For a long time, however, no finite-time regret guarantee was known for it: the analyses available before 2012 gave only o(T)o(T)o(T) regret.

Agrawal and Goyal (arXiv:1111.1797, COLT 2012) gave the first logarithmic bounds on the expected regret of Thompson Sampling. This mission formalizes their bound for the general case of NNN arms (their Theorem 2). A companion mission of the same series formalizes their two-armed bound (Theorem 1), whose proof is independent.

Timeline. Lai and Robbins (1985) proved that every consistent algorithm has regret at least of order ∑iΔiD(μi∥μ1)ln⁡T\sum_i \frac{\Delta_i}{D(\mu_i\|\mu_1)}\ln T∑i​D(μi​∥μ1​)Δi​​lnT. Auer, Cesa-Bianchi and Fischer (2002) showed that UCB1 achieves O(∑iln⁡T/Δi)O(\sum_i \ln T/\Delta_i)O(∑i​lnT/Δi​) in finite time. Agrawal and Goyal (2012) proved O((∑a1/Δa2)2ln⁡T)O((\sum_a 1/\Delta_a^2)^2\ln T)O((∑a​1/Δa2​)2lnT) for Thompson Sampling with NNN arms; Kaufmann, Korda and Munos (2012) and Agrawal and Goyal (2013) later proved the asymptotically optimal constant for Bernoulli rewards.

Setting

A stochastic NNN-armed bandit has arms 1,…,N1,\dots,N1,…,N. Arm iii, when played, yields a random reward drawn from a fixed distribution νi\nu_iνi​ supported in [0,1][0,1][0,1], with mean μi\mu_iμi​; rewards of an arm are i.i.d. and independent of the other arms. Arm 111 is assumed to be the unique optimal arm, μ1>μi\mu_1>\mu_iμ1​>μi​ for i≠1i\ne1i=1, and Δi=μ1−μi>0\Delta_i=\mu_1-\mu_i>0Δi​=μ1​−μi​>0 is the gap of arm iii.

Thompson Sampling for general stochastic bandits (Algorithm 2 of the paper) keeps, for each arm iii, a count SiS_iSi​ of successes and FiF_iFi​ of failures, both starting at 000. In each round ttt it draws θi(t)∼Beta(Si+1,Fi+1)\theta_i(t)\sim\mathrm{Beta}(S_i+1,F_i+1)θi​(t)∼Beta(Si​+1,Fi​+1) independently for every arm, plays i(t)=arg⁡max⁡iθi(t)i(t)=\arg\max_i\theta_i(t)i(t)=argmaxi​θi​(t), observes a reward r~t∼νi(t)\tilde r_t\sim\nu_{i(t)}r~t​∼νi(t)​, performs a Bernoulli trial with success probability r~t\tilde r_tr~t​, and increments Si(t)S_{i(t)}Si(t)​ on success and Fi(t)F_{i(t)}Fi(t)​ on failure.

The expected regret in time TTT is

E[R(T)]=E[∑t=1T(μ∗−μi(t))],μ∗=max⁡iμi,\mathbb E[\mathcal R(T)]=\mathbb E\Big[\sum_{t=1}^T(\mu^*-\mu_{i(t)})\Big],\qquad \mu^*=\max_i\mu_i,E[R(T)]=E[t=1∑T​(μ∗−μi(t)​)],μ∗=imax​μi​,

the expectation being over the rewards, the Bernoulli trials and the posterior samples.

The proof works with the reward stacks Zi,mZ_{i,m}Zi,m​: the outcome of the mmm-th Bernoulli trial of arm iii, all independent. Then s(j)=∑m≤jZ1,ms(j)=\sum_{m\le j}Z_{1,m}s(j)=∑m≤j​Z1,m​, the number of successes in the first jjj plays of arm 111, is a Binomial(j,μ1)\mathrm{Binomial}(j,\mu_1)Binomial(j,μ1​) random variable. The other objects of the proof are the threshold Li=24ln⁡T/Δi2L_i=24\ln T/\Delta_i^2Li​=24lnT/Δi2​, the saturated set C(t)C(t)C(t) of suboptimal arms with at least LiL_iLi​ plays before round ttt, the intervals IjI_jIj​ between the jjj-th and (j+1)(j+1)(j+1)-th plays of arm 111, and the counts γj\gamma_jγj​ and Vjℓ,aV_j^{\ell,a}Vjℓ,a​ defined in §4.

Formalization targets

Goal: Theorem 2

There is an absolute constant C>0C>0C>0 such that for every N≥2N\ge2N≥2, every instance as above and every horizon T≥2T\ge2T≥2,

E[R(T)]≤C(∑a=2N1Δa2)2ln⁡T.\mathbb E[\mathcal R(T)]\le C\Big(\sum_{a=2}^N\frac{1}{\Delta_a^2}\Big)^2\ln T .E[R(T)]≤C(a=2∑N​Δa2​1​)2lnT.

CCC does not depend on NNN, on the reward distributions or on TTT.

Milestones

  1. Lemma 4: with E(t)E(t)E(t) the event that every saturated arm's sample lies within Δi/2\Delta_i/2Δi​/2 of its mean, Pr⁡(E(t))≥1−4(N−1)/T2\Pr(E(t))\ge1-4(N-1)/T^2Pr(E(t))≥1−4(N−1)/T2, also conditionally on s(j)=ss(j)=ss(j)=s.
  2. Lemma 5 (Eq. (7)): the expected regret from saturated arms inside IjI_jIj​ is at most E[E[γj+1∣s(j)]∑aΔaE[min⁡{X(j,s(j),μa+Δa/2),T}∣s(j)]]\mathbb E\big[\mathbb E[\gamma_j+1\mid s(j)]\sum_a\Delta_a\mathbb E[\min\{X(j,s(j),\mu_a+\Delta_a/2),T\}\mid s(j)]\big]E[E[γj​+1∣s(j)]∑a​Δa​E[min{X(j,s(j),μa​+Δa​/2),T}∣s(j)]].
  3. Lemma 1: E[X(j,s,y)]=1/Fj+1,yB(s)−1\mathbb E[X(j,s,y)]=1/F^B_{j+1,y}(s)-1E[X(j,s,y)]=1/Fj+1,yB​(s)−1, where X(j,s,y)X(j,s,y)X(j,s,y) counts the trials before an independent Beta(s+1,j−s+1)\mathrm{Beta}(s+1,j-s+1)Beta(s+1,j−s+1) sample exceeds yyy.
  4. Lemma 3: a three-case bound on E[E[min⁡{X(j,s(j),y),T}∣s(j)]]\mathbb E[\mathbb E[\min\{X(j,s(j),y),T\}\mid s(j)]]E[E[min{X(j,s(j),y),T}∣s(j)]] in terms of the Bernoulli KL divergence DDD between yyy and μ1\mu_1μ1​.

Significance

The result. Theorem 2 shows that Thompson Sampling, a randomized Bayesian heuristic, achieves regret logarithmic in the horizon for any number of arms with bounded rewards, matching the order in TTT of the Lai–Robbins lower bound. Its dependence on the gaps, (∑aΔa−2)2(\sum_a\Delta_a^{-2})^2(∑a​Δa−2​)2, is worse than UCB1's; the paper's own Remark 1 and later work improve it. The proof introduced the device of bounding the waiting time between plays of the optimal arm through geometric variables with Beta-cdf parameters (Lemmas 1 and 3), which reappears in later analyses of Thompson Sampling.

Formalizing it. The theorem is proved on paper; it has not been machine-checked. Bandit theory in Lean (bandit environments, regret, UCB-type analyses) is still young, and no Beta–Bernoulli Thompson Sampling result is formalized. The mission produces a Lean model of Algorithm 2 for general [0,1][0,1][0,1] rewards with the paper's stack coupling, the §4 bookkeeping of saturated arms and intervals, and the paper's lemmas as separate targets.

Difficulty

Two difficulties are specific to the NNN-armed analysis. First, the arm that competes with arm 111 changes over time: the set of saturated arms grows, and which saturated arm is "best" depends on the history, so the waiting time between plays of arm 111 cannot be compared with a single geometric variable as in the two-armed case. Second, the number γj\gamma_jγj​ of rounds at which arm 111's sample is large but arm 111 is not played is not independent of the counts Vjℓ,aV_j^{\ell,a}Vjℓ,a​: both depend on the same posterior samples, and Lemma 5 needs a careful conditioning on the history to separate them. The obvious union bound over arms, treating each suboptimal arm as in the two-armed proof, fails because it ignores the interruptions by unsaturated arms, whose number is the source of the squared sum in the bound.

Formalization scope

  • Probability space. Algorithm 2 is realized on a product of three independent i.i.d. tables: Beta draws indexed by (arm, round, successes, failures), rewards indexed by (arm, round) and uniform variables indexed by (arm, round); the Bernoulli trial of a round succeeds when the played arm's uniform variable is below its reward. The law of the run is that of Algorithm 2, which runs for every round t=1,2,…t=1,2,\dotst=1,2,…. s(j)s(j)s(j) is the number of successful trials among the first jjj plays of arm 111 in this infinite run (possibly after the horizon TTT), so it is a Binomial(j,μ1)\mathrm{Binomial}(j,\mu_1)Binomial(j,μ1​) random variable for every jjj, as the paper's independent Z1,mZ_{1,m}Z1,m​ make it. Ties in the arg max (probability 000) go to the smallest index.
  • Indexing. Arms are Fin N, and Lean arm 0 is the paper's arm 111. Rounds are 0,…,T−10,\dots,T-10,…,T−1; Lean round ttt is the paper's round t+1t+1t+1. Sums over a=2,…,Na=2,\dots,Na=2,…,N are sums over a≠0a\ne0a=0.
  • Expectations are lower Lebesgue integrals in [0,∞][0,\infty][0,∞], which has no junk value for non-integrable functions. Conditional expectations given s(j)s(j)s(j) are written as finite sums over the values of s(j)s(j)s(j).
  • The O(⋅)O(\cdot)O(⋅). The paper writes O(⋅)O(\cdot)O(⋅) in the sense of its footnote 1 (f(n)≤c g(n)f(n)\le c\,g(n)f(n)≤cg(n) for n≥n0n\ge n_0n≥n0​). The goal states it with one universal constant CCC, quantified before NNN, the instance and TTT, for every T≥2T\ge2T≥2. The explicit constants printed in App. D are not formalized: expanding the paper's Eq. (21) gives terms 288(N−1)(ln⁡T)∑aΔa−2288(N-1)(\ln T)\sum_a\Delta_a^{-2}288(N−1)(lnT)∑a​Δa−2​ and 48(N−1)248(N-1)^248(N−1)2 where the paper prints 288(ln⁡T)∑iΔi−2288(\ln T)\sum_i\Delta_i^{-2}288(lnT)∑i​Δi−2​, and Eq. (22) drops a factor ln⁡T\ln TlnT in its 192/Δa2192/\Delta_a^2192/Δa2​ term. The O(⋅)O(\cdot)O(⋅) claim does not depend on these slips; a statement pinned to the printed numerals might be false.
  • Ruled out. A constant depending on NNN, on the means or on TTT; a fixed number of arms; Bernoulli rewards only; or any algorithm other than Algorithm 2 would each make the goal a different and weaker theorem. The statement quantifies over all N≥2N\ge2N≥2 and all reward distributions on [0,1][0,1][0,1].
  • Not included. Eq. (8), the bound ∑jE[γj∣s(j)]≤∑uLu+4(N−1)\sum_{j}\mathbb E[\gamma_j\mid s(j)]\le\sum_uL_u+4(N-1)∑j​E[γj​∣s(j)]≤∑u​Lu​+4(N−1) "for all instantiations", is not a milestone: each term is conditioned on a different s(j)s(j)s(j), and the pointwise reading does not follow from the argument given. Remark 1 (an alternate bound) and App. A (several optimal arms) are not part of this mission.
  • Contributions welcome: Beta–Binomial identities, geometric waiting times, Hoeffding bounds for binomial cdfs, and the stopping-time arguments behind Lemma 5. Lemma 1 and Lemma 3 are shared with the two-armed mission of this series.

Selected references

  • S. Agrawal and N. Goyal, Analysis of Thompson Sampling for the Multi-armed Bandit Problem, COLT 2012; arXiv:1111.1797v3, 2012. https://arxiv.org/abs/1111.1797
  • W. R. Thompson, On the likelihood that one unknown probability exceeds another in view of the evidence of two samples, Biometrika 25, 1933. https://doi.org/10.1093/biomet/25.3-4.285
  • T. L. Lai and H. Robbins, Asymptotically efficient adaptive allocation rules, Advances in Applied Mathematics 6, 1985. https://doi.org/10.1016/0196-8858(85)90002-8
  • P. Auer, N. Cesa-Bianchi and P. Fischer, Finite-time analysis of the multiarmed bandit problem, Machine Learning 47, 2002. https://doi.org/10.1023/A:1013689704352
  • O. Chapelle and L. Li, An empirical evaluation of Thompson Sampling, NIPS 2011. https://papers.nips.cc/paper/4321-an-empirical-evaluation-of-thompson-sampling
  • E. Kaufmann, N. Korda and R. Munos, Thompson Sampling: an asymptotically optimal finite-time analysis, ALT 2012. https://arxiv.org/abs/1205.4217
  • S. Agrawal and N. Goyal, Further optimal regret bounds for Thompson Sampling, AISTATS 2013. https://arxiv.org/abs/1209.3353
11 thms2 active usersReviewed
Statistics·Captain: mikedeng1

Weighted Sums of Certain Dependent Random Variables 1: Weighted Sums of Bounded Multiplicative Systems Grow at Most Like √(2Bₙ² log n)Research Paper

Motivation

Weighted sums of dependent random variables occur when the weights change with the observation horizon. Even if each random variable is bounded and centered, allowing the weights in row nnn to be chosen anew makes an almost-sure statement about all large nnn different from a bound on a single finite sum. Kazuo Azuma's 1967 paper treats this situation under a finite-product moment condition called class [M]. Its first main theorem bounds every row of an arbitrary real triangular array at the scale given by that row's Euclidean norm and log⁡n\log nlogn.

The condition is useful because it allows dependence. The paper notes that bounded martingale differences provide examples, but Theorem 1 is stated directly for class [M], without introducing a filtration in the result. The conclusion therefore records the property of the random variables actually used by this part of the paper, rather than restricting the mission to one familiar source of examples. Azuma, §1 and Theorem 1.

Setting

Fix a probability space (Ω,A,P)(\Omega,\mathcal A,P)(Ω,A,P). Let x1,x2,…x_1,x_2,\ldotsx1​,x2​,… be real measurable random variables. They form a bounded multiplicative system with unit bounds if ∣xk∣≤1|x_k|\le1∣xk​∣≤1 almost surely for every k≥1k\ge1k≥1, and

E ⁣[∏k∈Sxk]=0for every nonempty finite S⊆{1,2,…}.E\!\left[\prod_{k\in S}x_k\right]=0 \quad\text{for every nonempty finite }S\subseteq\{1,2,\ldots\}.E[k∈S∏​xk​]=0for every nonempty finite S⊆{1,2,…}.

The indices in SSS are distinct. Taking a singleton shows E[xk]=0E[x_k]=0E[xk​]=0; taking sets of two, three, or more indices imposes the full condition used in the paper. Pairwise zero correlations by themselves do not state class [M]. All bounds and moment conditions are for the positive indices, so x0x_0x0​ is outside the mathematical sequence. Azuma, p. 357, property [M].

For each n≥1n\ge1n≥1, choose real coefficients an1,…,anna_{n1},\ldots,a_{nn}an1​,…,ann​. There is no relation required between different rows. Define the weighted sum TnT_nTn​ and its weight norm BnB_nBn​ by

Tn=∑k=1nankxk,Bn=(∑k=1nank2)1/2.T_n=\sum_{k=1}^{n}a_{nk}x_k, \qquad B_n=\left(\sum_{k=1}^{n}a_{nk}^{2}\right)^{1/2}.Tn​=k=1∑n​ank​xk​,Bn​=(k=1∑n​ank2​)1/2.

Both definitions use the source's 1-based indices. A row of zero weights has Bn=0B_n=0Bn​=0 and Tn=0T_n=0Tn​=0 almost surely; such rows remain within the theorem. Azuma, p. 359, §3.

Formalization targets

The goal is Theorem 1, display (3.1): for every bounded multiplicative system and every real triangular array,

lim sup⁡n→∞∣Tn∣2Bn2log⁡n≤1P-almost surely.\limsup_{n\to\infty} \frac{|T_n|}{\sqrt{2B_n^2\log n}}\le1 \qquad P\text{-almost surely}.n→∞limsup​2Bn2​logn​∣Tn​∣​≤1P-almost surely.

The normalizing constant is exactly 222, and the upper bound is exactly 111. The mission does not assume that BnB_nBn​ grows, converges, or stays positive. In Lean, the target is stated as the equivalent operational bound: for each δ>0\delta>0δ>0, almost every outcome eventually satisfies ∣Tn∣≤(1+δ)2Bn2log⁡n|T_n|\le(1+\delta)\sqrt{2B_n^2\log n}∣Tn​∣≤(1+δ)2Bn2​logn​. This includes zero-weight rows without assigning a meaning to a real quotient 0/00/00/0.

The milestone list follows results displayed in the paper: the corrected convexity inequality (2.2), Lemma 1's exponential moment estimate (2.1), the exponential estimate in the proof of Theorem 1 with its factor 222, and the almost-sure finite exponential series on the next page. Lemma 1 gives, for arbitrary real b1,…,bnb_1,\ldots,b_nb1​,…,bn​ and t∈Rt\in\mathbb Rt∈R,

Eexp⁡ ⁣(t∑k=1nbkxk)≤exp⁡ ⁣(t22∑k=1nbk2).E\exp\!\left(t\sum_{k=1}^{n}b_kx_k\right) \le \exp\!\left(\frac{t^2}{2}\sum_{k=1}^{n}b_k^2\right).Eexp(tk=1∑n​bk​xk​)≤exp(2t2​k=1∑n​bk2​).

The paper prints (2.2) with a missing factor bnkb_{nk}bnk​ in its linear term. The mission records the printed text as provenance and states the corrected inequality in Lean; the printed version fails already when bnk=2b_{nk}=2bnk​=2 and xk=t=1x_k=t=1xk​=t=1. Azuma, pp. 357–360.

Significance

Theorem 1 turns an exponential moment bound for each finite weighted sum into a single almost-sure assertion along an entire triangular array. It gives a scale that adapts to the actual coefficients in each row: two arrays with different row norms receive different bounds, while no regularity across rows is required. The result is also the starting point for the weighted strong-law corollaries that follow it in the paper. Azuma, Theorem 1 and Corollary 1.

The mathematical theorem has been proved since 1967. This mission's remaining work is a machine-checked Lean proof of the exact theorem and its listed intermediate statements. A complete development would add reusable formal statements for bounded multiplicative systems and for their finite exponential moments. Those objects could support later work on dependent sums without importing a filtration or a stronger independence assumption. The proposal statements compile as open goals; compilation alone does not supply proofs.

Difficulty

The usual first step for independent bounded variables is to factor the exponential moment into one-variable expectations. Class [M] does not assume independence, so that factorization is unavailable. The condition controls every product with distinct indices, while allowing other dependence. The almost-sure conclusion must also hold when the coefficients change arbitrarily with nnn: bounds that depend on one fixed row do not by themselves settle what happens for all sufficiently large rows. Finally, rows with Bn=0B_n=0Bn​=0 require a statement that preserves the theorem rather than excluding them by an added positivity hypothesis.

Formalization scope

The Lean development represents the probability law by a measure μ\muμ with IsProbabilityMeasure μ, and a random sequence by x:N→Ω→Rx:\mathbb N\to\Omega\to\mathbb Rx:N→Ω→R. It uses measurable variables and states the unit bound almost surely at every positive index. The class [M] predicate quantifies over every nonempty finite set of positive indices; no conditional expectations, filtration, symmetry, or independence hypotheses enter Theorem 1. Finite products and finite weighted sums use ordinary real multiplication and Finset.Icc 1 n. The triangular weights have type N→N→R\mathbb N\to\mathbb N\to\mathbb RN→N→R, and only entries with 1≤k≤n1\le k\le n1≤k≤n contribute.

The norm BnB_nBn​ is the nonnegative real square root of the sum of squared weights. Real.log is zero at n=0n=0n=0 and n=1n=1n=1 in Lean, but the target is eventually quantified, so its asymptotic content concerns large nnn. The exponential estimate is stated for n≥1n\ge1n≥1. If Bn=0B_n=0Bn​=0, Lean's total division returns zero in the exponent's quotient; all weights in that row are zero, making this extension valid. In the almost-sure series, the term at index zero is set to zero. The source's limsup is represented by eventual inequalities for every positive excess, avoiding a real-valued limsup default on unbounded sequences.

The definition of class [M] includes every finite product, including singletons; replacing it by pairwise orthogonality would change the theorem. Measurability and the almost-sure unit bounds ensure that the finite products and the exponential functions in Lemma 1 are integrable, so their Lean integrals represent expectations. Contributions toward proofs of the corrected convexity bound, the moment estimate, the exponential series, and the final almost-sure step are all within scope. The finite-product predicate and the exponential estimate are reusable beyond this mission.

Selected references

  • Kazuo Azuma, Weighted sums of certain dependent random variables, Tôhoku Mathematical Journal 19 (1967), 357–367. DOI: 10.2748/tmj/1178243286.
7 thms2 active usersReviewed
Convex OptimizationMachine LearningStatistics·Captain: mikedeng1

Stability and Generalization 4: Relative-Entropy Regularization of Mixtures Has Uniform Stability M²/(λm)Research Paper

Motivation

A learning algorithm generalizes when its error on fresh data is close to its error on the training sample. Bousquet and Elisseeff (JMLR 2, 2002) showed that a single property of the algorithm, uniform stability, controls this gap with exponential concentration: if removing any one example from a training set of size mmm changes the loss of the output at every point by at most β\betaβ, the generalization error exceeds the empirical error by roughly 2β+(4mβ+M)ln⁡(1/δ)/(2m)2\beta + (4m\beta + M)\sqrt{\ln(1/\delta)/(2m)}2β+(4mβ+M)ln(1/δ)/(2m)​ with probability 1−δ1-\delta1−δ (their Theorem 12). The bound is useful only when β=O(1/m)\beta = O(1/m)β=O(1/m), and the second half of the paper identifies algorithms with that rate: Tikhonov regularization in a reproducing kernel Hilbert space (Theorem 22), and relative-entropy regularization of mixtures (Theorem 24), the subject of this mission.

Mixtures arise whenever a learner outputs a distribution over a parametric base class instead of a single hypothesis: Bayesian posterior averaging, Gibbs and randomized classifiers, exponential weights. Regularizing by the relative entropy to a prior is the maximum-a-posteriori reading of these procedures, and Theorem 24 is one of the earliest results showing that such posteriors are uniformly stable with rate 1/(λm)1/(\lambda m)1/(λm). The same mechanism (entropic regularization, stability through Pinsker's inequality) reappears in PAC-Bayesian analysis and in the stability of exponential-weights methods.

Setting

Let Θ\ThetaΘ be a measurable space with a reference measure ν\nuν, and write dθd\thetadθ for integration against ν\nuν. A base class H={hθ:θ∈Θ}\mathcal H = \{h_\theta : \theta \in \Theta\}H={hθ​:θ∈Θ} is indexed by Θ\ThetaΘ, and r(hθ,z)∈[0,M]r(h_\theta, z) \in [0, M]r(hθ​,z)∈[0,M] is the loss of the base hypothesis hθh_\thetahθ​ at an example z∈Zz \in Zz∈Z.

The algorithm outputs a density ggg with respect to ν\nuν: a measurable, nonnegative, integrable g:Θ→Rg : \Theta \to \mathbb Rg:Θ→R with ∫Θg dθ=1\int_\Theta g\,d\theta = 1∫Θ​gdθ=1. FFF denotes the set of all densities. A density is scored by the averaged loss

ℓ(g,z)=∫Θr(hθ,z) g(θ) dθ(28),\ell(g, z) = \int_\Theta r(h_\theta, z)\, g(\theta)\, d\theta \qquad (28),ℓ(g,z)=∫Θ​r(hθ​,z)g(θ)dθ(28),

the expected loss of a randomized predictor that draws hθh_\thetahθ​ from ggg. The relative entropy of ggg to g′g'g′ is

K(g,g′)=∫Θg(θ)ln⁡g(θ)g′(θ) dθ∈[0,∞],K(g, g') = \int_\Theta g(\theta) \ln \frac{g(\theta)}{g'(\theta)}\, d\theta \in [0, \infty],K(g,g′)=∫Θ​g(θ)lng′(θ)g(θ)​dθ∈[0,∞],

with K(g,g′)=+∞K(g, g') = +\inftyK(g,g′)=+∞ when g νg\,\nugν is not absolutely continuous with respect to g′ νg'\,\nug′ν or the integrand is not integrable.

Fix a prior f0∈Ff_0 \in Ff0​∈F, a parameter λ>0\lambda > 0λ>0, and a training set S=(z1,…,zm)S = (z_1, \dots, z_m)S=(z1​,…,zm​). The algorithm returns a minimizer over FFF of

Rr(g)=1m∑j=1mℓ(g,zj)+λK(g,f0)(29).R_r(g) = \frac1m \sum_{j=1}^m \ell(g, z_j) + \lambda K(g, f_0) \qquad (29).Rr​(g)=m1​j=1∑m​ℓ(g,zj​)+λK(g,f0​)(29).

For an index iii, the truncated objective is Rr∖i(g)=1m∑j≠iℓ(g,zj)+λK(g,f0)R_r^{\setminus i}(g) = \frac1m \sum_{j \ne i} \ell(g, z_j) + \lambda K(g, f_0)Rr∖i​(g)=m1​∑j=i​ℓ(g,zj​)+λK(g,f0​), and f∖if^{\setminus i}f∖i denotes one of its minimizers over FFF.

Formalization targets

Goal: Theorem 24

For every minimizer fff of (29), every minimizer f∖if^{\setminus i}f∖i of the truncated objective, and every example zzz,

∣ℓ(f,z)−ℓ(f∖i,z)∣≤M2λm.|\ell(f, z) - \ell(f^{\setminus i}, z)| \le \frac{M^2}{\lambda m}.∣ℓ(f,z)−ℓ(f∖i,z)∣≤λmM2​.

Milestones

  1. MMM-admissibility of (28) (§5.2.3, p. 518): ∣ℓ(g,z)−ℓ(g′,z)∣≤M∫Θ∣g−g′∣ dθ|\ell(g,z) - \ell(g',z)| \le M \int_\Theta |g - g'|\,d\theta∣ℓ(g,z)−ℓ(g′,z)∣≤M∫Θ​∣g−g′∣dθ.
  2. Pinsker's inequality, L1L^1L1 form (proof of Theorem 24): 12(∫Θ∣g−g′∣ dθ)2≤K(g,g′)\tfrac12 \bigl(\int_\Theta |g - g'|\,d\theta\bigr)^2 \le K(g, g')21​(∫Θ​∣g−g′∣dθ)2≤K(g,g′) for densities g,g′g, g'g,g′.
  3. Lemma 21 (p. 513): for a differentiable convex regularizer NNN on a vector space and a σ\sigmaσ-admissible loss,
dN(f,f∖i)+dN(f∖i,f)≤1λm(ℓ(f∖i,zi)−ℓ(f,zi)−dℓ(⋅,zi)(f∖i,f))≤σλm∣Δf(xi)∣.d_N(f, f^{\setminus i}) + d_N(f^{\setminus i}, f) \le \frac{1}{\lambda m}\Bigl(\ell(f^{\setminus i}, z_i) - \ell(f, z_i) - d_{\ell(\cdot, z_i)}(f^{\setminus i}, f)\Bigr) \le \frac{\sigma}{\lambda m}|\Delta f(x_i)|.dN​(f,f∖i)+dN​(f∖i,f)≤λm1​(ℓ(f∖i,zi​)−ℓ(f,zi​)−dℓ(⋅,zi​)​(f∖i,f))≤λmσ​∣Δf(xi​)∣.
  1. Bregman divergence of the relative entropy (proof of Theorem 24): dK(⋅,f0)(g,g′)=K(g,g′)d_{K(\cdot, f_0)}(g, g') = K(g, g')dK(⋅,f0​)​(g,g′)=K(g,g′).
  2. L1L^1L1 displacement bound (proof of Theorem 24):
∫Θ∣f−f∖i∣ dθ≤Mλm.\int_\Theta |f - f^{\setminus i}|\,d\theta \le \frac{M}{\lambda m}.∫Θ​∣f−f∖i∣dθ≤λmM​.

Significance

Theorem 24 places entropy-regularized posteriors among the algorithms to which the paper's exponential generalization bound applies: combined with Theorem 12 it gives, for the averaged loss, a deviation of order M2/(λm)+(M2/λ+M)ln⁡(1/δ)/mM^2/(\lambda m) + (M^2/\lambda + M)\sqrt{\ln(1/\delta)/m}M2/(λm)+(M2/λ+M)ln(1/δ)/m​. The proof also yields the L1L^1L1 bound ∫∣f−f∖i∣≤M/(λm)\int |f - f^{\setminus i}| \le M/(\lambda m)∫∣f−f∖i∣≤M/(λm), which by itself gives classification stability M/(λm)M/(\lambda m)M/(λm) for base hypotheses with values in {−1,1}\{-1, 1\}{−1,1} (remark after Theorem 24, p. 518).

The result is proved in the paper; no machine-checked proof is known to exist. A formalization produces reusable pieces that Mathlib does not have: Pinsker's inequality for densities in L1L^1L1 form (Mathlib has the Kullback–Leibler divergence InformationTheory.klDiv, but not Pinsker), the Bregman identity for the relative entropy, and a stability statement for minimizers over a space of probability densities.

Difficulty

The paper derives Theorem 24 from Lemma 21, which is stated for a regularizer that is defined and differentiable on a vector space. The relative entropy K(⋅,f0)K(\cdot, f_0)K(⋅,f0​) is defined only on the convex set of densities and is not differentiable at densities that vanish on a set of positive measure, so the general lemma does not literally apply, and the identity dK(⋅,f0)=Kd_{K(\cdot,f_0)} = KdK(⋅,f0​)​=K needs integrability conditions that the page does not state. A complete proof of the goal must either justify that application on the set of densities, or work directly with the minimizers, which requires identifying them and handling the +∞+\infty+∞ values of KKK. Pinsker's inequality itself requires a separate argument at the level of general measures.

Formalization scope

  • Densities are IsDensity ν g: measurable, nonnegative, integrable, total mass one, with respect to a σ-finite reference measure ν. The integral dθd\thetadθ is always against ν, never Lebesgue measure.
  • The base loss is r : Θ → Z → ℝ, measurable in θ, with 0 ≤ r ≤ M; the paper's costs are nonnegative (p. 502).
  • KKK is InformationTheory.klDiv of the measures g · ν and g' · ν, in ℝ≥0∞. The objectives (29) and its truncation take values in ℝ≥0∞. A formalization that converts KKK to a real number with toReal would send K=+∞K = +\inftyK=+∞ to 000 and make the worst densities minimizers; that reading is excluded.
  • The minimizers are given as hypotheses: f minimizes (29) and f' minimizes the truncated objective over all densities, for the given S : Fin m → Z and i : Fin m.
  • Corrected reading of the algorithm on S∖iS^{\setminus i}S∖i. The goal is stated in the pairwise form of the paper's proof: f∖if^{\setminus i}f∖i minimizes the truncated objective with factor 1/m1/m1/m, the analogue of (20), not (29) run on the m−1m-1m−1 points of S∖iS^{\setminus i}S∖i with factor 1/(m−1)1/(m-1)1/(m−1).
  • Corrected display. The objective displayed before Theorem 24 has ℓ(g,z)\ell(g, z)ℓ(g,z) inside the sum; (29) has ℓ(g,zi)\ell(g, z_i)ℓ(g,zi​), which is used.
  • Lemma 21 is stated as printed, in its differentiable case, on a real normed space whose elements act as functions on XXX through a linear map; the goal does not instantiate it. The Bregman identity is stated with the explicit gradient ln⁡(g′/f0)+1\ln(g'/f_0) + 1ln(g′/f0​)+1, for f0,g′>0f_0, g' > 0f0​,g′>0, finite K(g,f0)K(g, f_0)K(g,f0​), K(g′,f0)K(g', f_0)K(g′,f0​), and integrable gln⁡(g′/f0)g \ln(g'/f_0)gln(g′/f0​).

Contributions welcome: proofs of Pinsker's inequality for klDiv (reusable far beyond this mission), of the Bregman identity, of Lemma 21, and of the goal by any route.

Selected references

  • O. Bousquet and A. Elisseeff, Stability and Generalization, Journal of Machine Learning Research 2 (2002) 499–526. https://www.jmlr.org/papers/v2/bousquet02a.html
  • T. M. Cover and J. A. Thomas, Elements of Information Theory, Wiley, 1991 (Pinsker's inequality). https://doi.org/10.1002/0471200611
  • R. T. Rockafellar, Convex Analysis, Princeton University Press, 1970 (Bregman divergences, Appendix C of the paper). https://doi.org/10.1515/9781400873173
11 thms2 active usersReviewed
Machine LearningStatistics·Captain: mikedeng1

Stability and Generalization 2: Exponential Generalization Bounds for Uniformly Stable AlgorithmsResearch Paper

Motivation

A learning algorithm is judged by its generalization error: the expected loss of the hypothesis it outputs on a fresh example. That quantity depends on an unknown distribution, so it is estimated from the training data, by the empirical error (the average loss on the training set) or the leave-one-out error (the average loss on each training point of the hypothesis trained without it). Classical learning theory controls the gap between these estimates and the true error uniformly over a hypothesis class, through its VC dimension or covering numbers. Such bounds say nothing useful about algorithms that search very large or infinite-dimensional spaces, such as support vector machines and regularization networks in a reproducing kernel Hilbert space.

Bousquet and Elisseeff (JMLR 2 (2002) 499–526) replaced the capacity of the class by a property of the algorithm, its stability: how much its output changes when one training example is removed. Their exponential bound for uniformly stable algorithms is the starting point of the stability approach to generalization, which was later used for stochastic gradient descent (Hardt, Recht and Singer, 2016) and differential privacy, and sharpened by Feldman and Vondrák (2019) and Bousquet, Klochkov and Zhivotovskiy (2020).

Timeline. Rogers and Wagner (1978) and Devroye and Wagner (1979) bounded the leave-one-out error of local rules such as k-nearest neighbours through their stability. McDiarmid (1989) proved the bounded-differences inequality. Lugosi and Pawlak (1994) combined it with smoothed error estimates. Kearns and Ron (1999) named hypothesis and error stability and related them to the VC dimension. Bousquet and Elisseeff (2002) introduced uniform stability and proved the exponential bounds this mission formalizes.

Setting

Let Z=X×YZ = X \times YZ=X×Y be a measurable space of labelled examples with an unknown probability distribution DDD. A training set S=(z1,…,zm)S = (z_1, \dots, z_m)S=(z1​,…,zm​) is drawn from DmD^mDm. A learning algorithm AAA maps a training set to a hypothesis AS:X→Y′A_S : X \to Y'AS​:X→Y′. It is deterministic and symmetric: it depends on the training set only as a multiset, so it is a function Multiset (X × Y) → (X → Y'), defined for training sets of every size. For a cost ccc, the loss of a hypothesis fff at z=(x,y)z = (x, y)z=(x,y) is ℓ(f,z)=c(f(x),y)\ell(f, z) = c(f(x), y)ℓ(f,z)=c(f(x),y).

Given SSS, write S∖iS^{\setminus i}S∖i for SSS with its iii-th example removed, and SiS^iSi for SSS with ziz_izi​ replaced by an independent draw zi′∼Dz_i' \sim Dzi′​∼D. The three errors are

R=Ez∼D[ℓ(AS,z)],Remp=1m∑i=1mℓ(AS,zi),Rloo=1m∑i=1mℓ(AS∖i,zi).R = \mathbb E_{z \sim D}[\ell(A_S, z)], \qquad R_{\mathrm{emp}} = \frac1m \sum_{i=1}^m \ell(A_S, z_i), \qquad R_{\mathrm{loo}} = \frac1m \sum_{i=1}^m \ell(A_{S^{\setminus i}}, z_i).R=Ez∼D​[ℓ(AS​,z)],Remp​=m1​i=1∑m​ℓ(AS​,zi​),Rloo​=m1​i=1∑m​ℓ(AS∖i​,zi​).

An algorithm has uniform stability β\betaβ at sample size mmm (Definition 6) if for every S∈ZmS \in Z^mS∈Zm, every iii and every z∈Zz \in Zz∈Z,

∣ℓ(AS,z)−ℓ(AS∖i,z)∣≤β.|\ell(A_S, z) - \ell(A_{S^{\setminus i}}, z)| \le \beta .∣ℓ(AS​,z)−ℓ(AS∖i​,z)∣≤β.

As a function of the sample size this constant is written βm\beta_mβm​.

Formalization targets

Goal: Theorem 12

If AAA has uniform stability β\betaβ and 0≤ℓ(AS,z)≤M0 \le \ell(A_S, z) \le M0≤ℓ(AS​,z)≤M for all zzz and all training sets SSS, then for every m≥1m \ge 1m≥1 and δ∈(0,1)\delta \in (0,1)δ∈(0,1), each of the following holds, separately, with probability at least 1−δ1 - \delta1−δ over S∼DmS \sim D^mS∼Dm:

R≤Remp+2β+(4mβ+M)ln⁡(1/δ)2m,(11)R \le R_{\mathrm{emp}} + 2\beta + (4m\beta + M)\sqrt{\frac{\ln(1/\delta)}{2m}}, \qquad (11)R≤Remp​+2β+(4mβ+M)2mln(1/δ)​​,(11) R≤Rloo+β+(4mβ+M)ln⁡(1/δ)2m.(12)R \le R_{\mathrm{loo}} + \beta + (4m\beta + M)\sqrt{\frac{\ln(1/\delta)}{2m}}. \qquad (12)R≤Rloo​+β+(4mβ+M)2mln(1/δ)​​.(12)

Milestones

  1. McDiarmid's inequality (Theorem 2): for measurable F:Zm→RF : Z^m \to \mathbb RF:Zm→R with ∣F(S)−F(Si)∣≤ci|F(S) - F(S^i)| \le c_i∣F(S)−F(Si)∣≤ci​, PS[F−ESF≥ϵ]≤e−2ϵ2/∑ici2P_S[F - \mathbb E_S F \ge \epsilon] \le e^{-2\epsilon^2/\sum_i c_i^2}PS​[F−ES​F≥ϵ]≤e−2ϵ2/∑i​ci2​.
  2. Uniform stability β\betaβ implies ∣ℓ(AS,z)−ℓ(ASi,z)∣≤2β|\ell(A_S, z) - \ell(A_{S^i}, z)| \le 2\beta∣ℓ(AS​,z)−ℓ(ASi​,z)∣≤2β (p. 504).
  3. Lemma 7: the bias identities for ES[R−Remp]\mathbb E_S[R - R_{\mathrm{emp}}]ES​[R−Remp​], ES[R(A,S∖i)−Rloo]\mathbb E_S[R(A,S^{\setminus i}) - R_{\mathrm{loo}}]ES​[R(A,S∖i)−Rloo​] and ES[R−Rloo]\mathbb E_S[R - R_{\mathrm{loo}}]ES​[R−Rloo​].
  4. R−RempR - R_{\mathrm{emp}}R−Remp​ and R−RlooR - R_{\mathrm{loo}}R−Rloo​ have bounded differences ci=4β+M/mc_i = 4\beta + M/mci​=4β+M/m.
  5. ES[R−Remp]≤2β\mathbb E_S[R - R_{\mathrm{emp}}] \le 2\betaES​[R−Remp​]≤2β and ES[R−Rloo]≤β\mathbb E_S[R - R_{\mathrm{loo}}] \le \betaES​[R−Rloo​]≤β.
  6. The tail bounds PS[R−Remp>ϵ+2β]≤exp⁡(−2mϵ2/(4mβ+M)2)P_S[R - R_{\mathrm{emp}} > \epsilon + 2\beta] \le \exp(-2m\epsilon^2/(4m\beta+M)^2)PS​[R−Remp​>ϵ+2β]≤exp(−2mϵ2/(4mβ+M)2) and the leave-one-out analogue.

Significance

When β=O(1/m)\beta = O(1/m)β=O(1/m) both bounds are O(1/m)O(1/\sqrt m)O(1/m​), with constants that do not depend on any capacity of the hypothesis space. Later sections of the paper show that Tikhonov regularization in a reproducing kernel Hilbert space has β=O(1/(λm))\beta = O(1/(\lambda m))β=O(1/(λm)), so the theorem gives generalization bounds for support vector regression, kernel ridge regression and, through a smoothed loss, soft-margin classification. The theorem is also the template for later stability bounds: the decomposition into a bias term controlled by stability and a deviation term controlled by a concentration inequality recurs throughout the literature.

The result has been proved since 2002, and replace-one variants appear in textbooks (Mohri, Rostamizadeh and Talwalkar, Foundations of Machine Learning, Theorem 14.2; Shalev-Shwartz and Ben-David, Chapter 13). On Prove2Me the replace-one textbook version is not formalized, and Mathlib at the platform's environment has no McDiarmid inequality. This mission asks for a machine-checked proof of the paper's remove-one version with its exact constants, and a reusable McDiarmid inequality with per-coordinate constants.

Difficulty

The deterministic steps (the bias identity and the bounded-differences estimates) are short on paper. The central difficulty is McDiarmid's inequality itself: it needs a martingale argument along the coordinates of a product measure, or an equivalent tensorization of conditional sub-Gaussian bounds, with the Doob martingale E[F∣z1,…,zk]\mathbb E[F \mid z_1, \dots, z_k]E[F∣z1​,…,zk​] expressed through partial integration over Measure.pi. Hoeffding's inequality for sums, which Mathlib has, does not apply directly: R−RempR - R_{\mathrm{emp}}R−Remp​ is not a sum of independent terms. A second, bookkeeping difficulty is Lemma 7: the identities rest on exchanging ziz_izi​ with zi′z_i'zi′​ and on the symmetry of AAA, which in Lean means measure-preserving coordinate permutations of Dm⊗DD^m \otimes DDm⊗D and multiset equalities such as Si ∖i=S∖iS^{i\,\setminus i} = S^{\setminus i}Si∖i=S∖i.

Formalization scope

Conventions committed to by the Lean statements:

  • An algorithm is a function of a multiset; this is how symmetry is encoded. Samples are Fin m → X × Y, and SiS^iSi is Function.update.
  • The law of SSS is Measure.pi (fun _ => D) with D a probability measure; zi′z_i'zi′​ and zzz are independent draws, integrated against the product (Measure.pi fun _ => D).prod D.
  • The paper's standing assumption that all functions are measurable is one hypothesis: (S,z)↦ℓ(AS,z)(S, z) \mapsto \ell(A_S, z)(S,z)↦ℓ(AS​,z) is measurable for every sample size. With the bound 0≤ℓ(AT,z)≤M0 \le \ell(A_T, z) \le M0≤ℓ(AT​,z)≤M for training sets TTT of every size, every expectation is a genuine integral, so no bound can hold because a non-integrable expectation defaults to 000.
  • Uniform stability quantifies over every sample, every index and every point, not almost every one.
  • "With probability at least 1−δ1 - \delta1−δ" means the DmD^mDm-measure of the failure set is at most δ\deltaδ. The two bounds (11) and (12) are separate statements, joined by a conjunction; they are not claimed for one joint event.
  • The paper assumes βm\beta_mβm​ is non-increasing in mmm and bounds βm−1\beta_{m-1}βm−1​ by βm\beta_mβm​ (p. 504). The leave-one-out bound (12) and its milestones carry the explicit hypothesis of uniform stability β\betaβ at size m−1m-1m−1; the empirical bound (11) does not.
  • McDiarmid's inequality sums ci2c_i^2ci2​ over i=1,…,mi = 1, \dots, mi=1,…,m; the paper's printed upper index nnn is a slip.
  • When a displayed tail bound has a zero denominator, its formal statement uses the limiting bound 000. In McDiarmid's inequality this is the constant-function case; in the stability tails the loss is identically zero.

The stability notion is the paper's remove-one notion. A formalization with replace-one stability would prove a different theorem with different constants, and the published FoundationsML_Stability_UniformlyStable (replace-one) is therefore not used.

The development needs the published loss, empirical-error and generalization-error definitions from Foundations of Machine Learning, a McDiarmid inequality on product measures (reusable for any bounded-differences argument), and the coordinate-exchange lemmas for Measure.pi behind Lemma 7. Contributions of a general McDiarmid inequality, of exchangeability lemmas for product measures, and of proofs of any milestone are welcome.

Selected references

  • O. Bousquet and A. Elisseeff, Stability and Generalization, Journal of Machine Learning Research 2 (2002) 499–526. https://www.jmlr.org/papers/v2/bousquet02a.html
  • C. McDiarmid, On the method of bounded differences, Surveys in Combinatorics, LMS Lecture Note Series 141 (1989) 148–188. https://doi.org/10.1017/CBO9781107359949.008
  • L. Devroye and T. Wagner, Distribution-free performance bounds for potential function rules, IEEE Trans. Inform. Theory 25 (1979) 601–604. https://doi.org/10.1109/TIT.1979.1056087
  • M. Kearns and D. Ron, Algorithmic stability and sanity-check bounds for leave-one-out cross-validation, Neural Computation 11 (1999) 1427–1453. https://doi.org/10.1162/089976699300016304
  • M. Hardt, B. Recht and Y. Singer, Train faster, generalize better: stability of stochastic gradient descent, ICML 2016. https://arxiv.org/abs/1509.01240
  • V. Feldman and J. Vondrák, High probability generalization bounds for uniformly stable algorithms with nearly optimal rate, COLT 2019. https://arxiv.org/abs/1902.10710
  • O. Bousquet, Y. Klochkov and N. Zhivotovskiy, Sharper bounds for uniformly stable algorithms, COLT 2020. https://arxiv.org/abs/1910.07833
  • M. Mohri, A. Rostamizadeh and A. Talwalkar, Foundations of Machine Learning, 2nd ed., MIT Press, 2018, Chapter 14.
13 thms2 active usersReviewed
Machine LearningReinforcement Learning·Captain: mikedeng1

Minimax Regret Bounds for Reinforcement Learning II: High-Probability Regret Bound for UCBVI with a Bernstein–Freedman BonusResearch Paper

Why finite-horizon reinforcement learning needs a variance-sensitive bound

An agent can learn to act in an unknown environment by repeatedly running a finite episode, observing the states reached after its actions, and updating its model of the environment. The agent must trade off rewards in the current episode against information that may improve later decisions. A regret bound measures the cumulative value lost relative to an optimal policy that knows the true transition probabilities. Its dependence on the number of states, actions, episode steps, and interactions says how much exploration that uncertainty can force.

Azar, Osband, and Munos study this question for a finite-horizon Markov decision process with known, bounded rewards and an unknown, stationary transition kernel. Their UCBVI algorithm estimates action values from observed transitions and adds an exploration bonus. Their second version, UCBVI-BF, uses the empirical variance of the next-state value in that bonus. Their Theorem 2 gives an explicit high-probability regret bound whose leading dependence on the horizon is smaller than the bound they give for the simpler UCBVI-CH bonus. The paper states that, in a sufficiently long-run regime, its leading order matches the cited lower-bound scale up to logarithmic factors. This mission targets the explicit theorem, including its lower-order terms, rather than only that asymptotic comparison.

The MDP, interaction, and algorithm

Let S\mathcal SS and A\mathcal AA be nonempty finite state and action sets with cardinalities SSS and AAA. A stationary transition kernel P(y∣x,a)P(y\mid x,a)P(y∣x,a) is a probability distribution on next states yyy for every current state xxx and action aaa. The reward R(x,a)R(x,a)R(x,a) is deterministic, known to the learner, and lies in [0,1][0,1][0,1]. These are the conditions of Assumption 1 and §2. Episodes have H≥1H\ge1H≥1 steps; KKK episodes comprise T=KHT=KHT=KH interactions.

A policy π\piπ chooses an action for each state and step. Its value Vhπ(x)V_h^\pi(x)Vhπ​(x) is the expected reward from step hhh through the final step when the state at hhh is xxx; the terminal value is VH+1π=0V_{H+1}^\pi=0VH+1π​=0. The optimal value Vh∗(x)=sup⁡πVhπ(x)V_h^*(x)=\sup_\pi V_h^\pi(x)Vh∗​(x)=supπ​Vhπ​(x) ranges over all deterministic policies of this form. At the start of episode kkk, the environment may choose the initial state using the completed episodes. The learner then fixes a policy πk\pi_kπk​, observes transitions during the episode, and updates counts for the next episode. Its regret is

Regret⁡(K)=∑k=1K(V1∗(xk,1)−V1πk(xk,1)).\operatorname{Regret}(K)=\sum_{k=1}^{K}\bigl(V_1^*(x_{k,1})-V_1^{\pi_k}(x_{k,1})\bigr).Regret(K)=k=1∑K​(V1∗​(xk,1​)−V1πk​​(xk,1​)).

For each state-action pair, Nk(x,a,y)N_k(x,a,y)Nk​(x,a,y) counts transitions to yyy in episodes before kkk, and Nk(x,a)=∑yNk(x,a,y)N_k(x,a)=\sum_yN_k(x,a,y)Nk​(x,a)=∑y​Nk​(x,a,y). When the latter is positive, P^k(y∣x,a)=Nk(x,a,y)/Nk(x,a)\widehat P_k(y\mid x,a)=N_k(x,a,y)/N_k(x,a)Pk​(y∣x,a)=Nk​(x,a,y)/Nk​(x,a). The count Nk,h′(y)N'_{k,h}(y)Nk,h′​(y) records previous episodes whose state at step hhh was yyy. Algorithms 2 and 4 compute optimistic Qk,hQ_{k,h}Qk,h​ backward from zero terminal value, take a minimum with the previous episode's QQQ estimate and with HHH, and choose a maximizing action at every state. Previously unseen pairs receive Qk,h=HQ_{k,h}=HQk,h​=H. The Bernstein–Freedman bonus uses the empirical variance of Vk,h+1V_{k,h+1}Vk,h+1​ under P^k\widehat P_kPk​ and an additional term based on Nk,h+1′N'_{k,h+1}Nk,h+1′​; the algorithm uses Lalg=ln⁡(5SAT/δ)L_{\rm alg}=\ln(5SAT/\delta)Lalg​=ln(5SAT/δ).

Formalization targets

The goal is Theorem 2 on p. 5. For any MDP and interaction described above and every δ>0\delta>0δ>0, write L=ln⁡(5HSAT/δ)L=\ln(5HSAT/\delta)L=ln(5HSAT/δ). The target is the exact bad-event form of the printed high-probability bound:

Pr⁡ ⁣{Regret⁡(K)>30HLSAK+2500H2S2AL2+4H3/2KL}≤δ.\Pr\!\left\{\operatorname{Regret}(K)>30HL\sqrt{SAK}+2500H^2S^2AL^2+4H^{3/2}\sqrt{KL}\right\}\le\delta.Pr{Regret(K)>30HLSAK​+2500H2S2AL2+4H3/2KL​}≤δ.

The milestone list contains three empirical-transition deviations from the proof of Lemma 1: Eq. (9) for a value-weighted transition error, the displayed count bound before Eq. (11), and Eq. (12) for the full transition row's ℓ1\ell_1ℓ1​ error. It also contains Lemma 2's variance comparison and Eq. (26), which relates cumulative conditional next-value variance to the variance of an episode return. These are source-indexed targets, with their printed constants retained.

What the result and its formalization supply

The theorem gives a quantitative guarantee for a particular executable decision rule: its regret grows sublinearly in KKK in the leading term, with explicit dependence on SSS, AAA, and HHH. The result lets one compare the horizon dependence of a variance-sensitive bonus with a value-agnostic bonus under the same finite-horizon model. It also fixes which logarithm belongs in the algorithm and which appears in the reported bound; replacing either changes the claim.

A formal proof would connect a fully specified adaptive interaction to its finite probability law, empirical counts, backward value iteration, and the stated high-probability conclusion. The local prior-art search found reusable transition-kernel vocabulary and general concentration tools, but no published formal statement of this exact UCBVI-BF algorithm or theorem. The mission's finite path and variance definitions can also support other episodic reinforcement-learning bounds that use conditional variance.

Where the difficulty lies

The bonus is computed using a value function that itself depends on earlier observations and the same episode's backward recursion. A concentration inequality for a fixed transition row and a fixed test function therefore does not directly control every value estimate encountered by the algorithm. The number of samples in a row is also random and changes with the learner's past actions. The regret compares a policy's value at an environment-chosen initial state with a supremum over all policies, while the learner's greedy action must be defined at states it never visits. These dependencies are the central obstacle to turning local concentration statements into the episode-level bound.

Formalization scope and conventions

The Lean model uses finite sums rather than measure theory. A published predicate supplies the stationary, real-valued transition kernel; a local MDP adds the known deterministic reward. State and action types are finite and nonempty. Policies are deterministic and depend on the step. The supremum defining V∗V^*V∗ ranges over their finite function type. A theorem quantifies over every maximizing tie-breaking rule and every initial-state rule that reads only completed episodes. The probability of an event is constructed as a sum over finite outcome sequences, each weighted by the product of true transition probabilities. Counts use all past transitions and no current or future outcomes. These choices rule out a trivialization that assumes the desired law or optimizes over an unbounded class of arbitrary functions.

Lean indexes the HHH steps from zero, while the paper indexes them from one. The last observed next state is kept because Algorithm 4 counts states at the terminal index H+1H+1H+1. At Nk,h+1′(y)=0N'_{k,h+1}(y)=0Nk,h+1′​(y)=0, Algorithm 4's quotient is interpreted as infinite and the capped term is H2H^2H2; Lean's ordinary division by zero would incorrectly produce zero. The algorithm uses Lalg=ln⁡(5SAT/δ)L_{\rm alg}=\ln(5SAT/\delta)Lalg​=ln(5SAT/δ), while Theorem 2's bound uses L=ln⁡(5HSAT/δ)L=\ln(5HSAT/\delta)L=ln(5HSAT/δ). For Eq. (26), the appendix ends its sums at H−1H-1H−1 under a shifted terminal convention; the local statement includes all HHH reward steps and the terminal value VH+1=0V_{H+1}=0VH+1​=0 used by Algorithm 2. The milestone text remains the printed text. The count milestone is the display before Eq. (11), since Eq. (11) drops a factor of 222 under the square root present in that display.

Theorem 2 retains its printed 2500H2S2AL22500H^2S^2AL^22500H2S2AL2 term. The appendix's displayed Lemma 13 calculation does not reproduce that second-order constant when propagated to Lemma 14; this is a source proof gap, not a hypothesis of the theorem. Work on the probability normalization, random-count concentration, adaptive value estimates, variance identity, and a valid route to the printed explicit constants is welcome. A proof with altered constants or an asymptotic-only conclusion would be a different target.

Selected references

  • M. G. Azar, I. Osband, and R. Munos, Minimax Regret Bounds for Reinforcement Learning, arXiv:1703.05449v2, 2017. Preprint.
13 thms2 active usersReviewed
Dynamic ProgrammingMarkov ChainOperations Research·Captain: mikedeng1

An Inventory Model with Limited Production Capacity and Uncertain Demands I. The Average-Cost Criterion: With Finite Storage a Modified Base-Stock Policy Is Strongly Average-Cost OptimalResearch Paper

Motivation

A manufacturer that makes one product to stock faces random demand, can produce at most bbb units per period, and can store at most UUU units. The classical result without the production limit is that a base-stock policy is optimal: raise inventory to a fixed level yˉ\bar yyˉ​ each period. With a production limit, the natural modification is to produce up to yˉ\bar yyˉ​ when that is possible and to produce at full capacity otherwise. Federgruen and Zipkin (1986) proved that this modified base-stock (critical-number) policy is optimal under the long-run average-cost criterion, for discrete demand with a general convex cost. Production-capacity models of this type are standard in operations management texts, and the result underlies the computational and comparative-static work that followed, starting with Part II of the same paper, which treats discounted costs.

Timeline.

  • 1950s–60s: optimality of base-stock (critical-number) policies for uncapacitated periodic-review models; see Heyman and Sobel's Stochastic Models in Operations Research, Vol. II (1984).
  • 1986: Federgruen and Zipkin, Part I (average cost, MOR 11(2):193–207) and Part II (discounted cost, MOR 11(2):208–215) establish the capacitated case. Part I handles the unbounded state space with a general average-cost theory for countable-state Markov decision processes by Federgruen, Schweitzer and Tijms (1983).

Setting

Time is divided into periods t=0,1,…t = 0, 1, \dotst=0,1,…. The demands D0,D1,…D_0, D_1, \dotsD0​,D1​,… are independent copies of a random variable DDD with values in {0,1,2,… }\{0, 1, 2, \dots\}{0,1,2,…} and probability mass function p(j)p(j)p(j); write μ=E(D)\mu = E(D)μ=E(D) and P(j)=Pr⁡{D≤j}P(j) = \Pr\{D \le j\}P(j)=Pr{D≤j}. At the start of period ttt the inventory is an integer xtx_txt​ (negative values are backorders). The decision maker raises it to

yt∈Y(xt)={y∈Z:xt≤y≤xt+b, y≤U},y_t \in Y(x_t) = \{y \in \mathbb Z : x_t \le y \le x_t + b,\ y \le U\},yt​∈Y(xt​)={y∈Z:xt​≤y≤xt​+b, y≤U},

pays the expected one-period cost G(yt)G(y_t)G(yt​), and demand is subtracted: xt+1=yt−Dtx_{t+1} = y_t - D_txt+1​=yt​−Dt​. The order cost per unit is set to zero, as in the paper; this loses no generality because every policy with finite average cost has the same average order cost.

The standing assumptions are: G≥0G \ge 0G≥0 is convex and G(y)→∞G(y) \to \inftyG(y)→∞ as ∣y∣→∞|y| \to \infty∣y∣→∞ (Assumption 1); the characteristic function of DDD is analytic at the origin (Assumption 2), and 0<μ0 < \mu0<μ; G(y)≤A+B∣y∣ρG(y) \le A + B|y|^\rhoG(y)≤A+B∣y∣ρ for some positive integer ρ\rhoρ (Assumption 3); b>μb > \mub>μ and P(b)<1P(b) < 1P(b)<1 (Assumption 4). The smallest global minimizer of GGG is yˉ∞\bar y^\inftyyˉ​∞, and U≥yˉ∞U \ge \bar y^\inftyU≥yˉ​∞.

A Markov policy is a sequence π=(π0,π1,… )\pi = (\pi_0, \pi_1, \dots)π=(π0​,π1​,…) of maps with πt(x)∈Y(x)\pi_t(x) \in Y(x)πt​(x)∈Y(x). The critical-number policy with critical number yˉ\bar yyˉ​ is δ[yˉ](x)=max⁡(x,min⁡(yˉ,x+b))\delta[\bar y](x) = \max(x, \min(\bar y, x + b))δ[yˉ​](x)=max(x,min(yˉ​,x+b)). A stationary policy δ\deltaδ is strongly optimal with average cost ggg if, from every initial state x≤Ux \le Ux≤U, its average cost t−1E{∑i<tG(yi)}t^{-1}E\{\sum_{i<t} G(y_i)\}t−1E{∑i<t​G(yi​)} converges to ggg, while every Markov policy has lim-inf average cost at least ggg from every initial state.

The analysis uses the operators Rv(y)=G(y)+E v(y−D)Rv(y) = G(y) + E\,v(y - D)Rv(y)=G(y)+Ev(y−D) and Sv(x)=min⁡y∈Y(x)Rv(y)Sv(x) = \min_{y \in Y(x)} Rv(y)Sv(x)=miny∈Y(x)​Rv(y), and the optimality equation

g+v(x)=Sv(x),x≤U.(6)g + v(x) = Sv(x),\qquad x \le U. \tag{6}g+v(x)=Sv(x),x≤U.(6)

For an interval ι=[l,u]\iota = [l, u]ι=[l,u], Hιv(x)H_\iota v(x)Hι​v(x) is the largest expected sum of v(yt)v(y_t)v(yt​), over policies forced to produce at capacity below lll and to produce nothing above uuu, until the inventory first returns to ι\iotaι.

Formalization targets

Goal: Theorem 1 (p. 202)

There exist g∗g^*g∗, v∗v^*v∗ and y∗≥yˉ∞y^* \ge \bar y^\inftyy∗≥yˉ​∞ such that (g∗,v∗)(g^*, v^*)(g∗,v∗) solves (6), v∗v^*v∗ is convex with global minimizer y∗y^*y∗, and

δ∗=δ[y∗] is strongly optimal with average cost g∗.\delta^* = \delta[y^*] \text{ is strongly optimal with average cost } g^*.δ∗=δ[y∗] is strongly optimal with average cost g∗.

The y∗y^*y∗ in the optimality claim is the minimizer constructed in part (a).

Milestones

  • Lemma 2(a)–(c) (pp. 196–197): a normal-tail inequality and two series estimates.
  • Lemma 3 (p. 198): if v(x)=O(∣x∣q)v(x) = O(|x|^q)v(x)=O(∣x∣q) then Hιv(x)=O(∣x∣q+3)H_\iota v(x) = O(|x|^{q+3})Hι​v(x)=O(∣x∣q+3).
  • Corollary 1 (p. 200): Hι1=O(∣x∣3)H_\iota 1 = O(|x|^3)Hι​1=O(∣x∣3) and HιG=O(∣x∣ρ+3)H_\iota G = O(|x|^{\rho+3})Hι​G=O(∣x∣ρ+3), both finite.
  • Corollary 2 (p. 201): (t+1)−1P[δ0t]⋯P[δtt](Hι1+HιG)(x)→0(t+1)^{-1}P[\delta_{0t}]\cdots P[\delta_{tt}](H_\iota 1 + H_\iota G)(x) \to 0(t+1)−1P[δ0t​]⋯P[δtt​](Hι​1+Hι​G)(x)→0.
  • Lemma 4 (p. 201): reachability of every state in [L,U−D−][L, U - D_-][L,U−D−​] under some policy that produces at capacity below LLL.
  • Lemma 5 (p. 202): SSS and QQQ preserve the class VVV of convex functions of growth O(∣x∣ρ+3)O(|x|^{\rho+3})O(∣x∣ρ+3) that are nonincreasing below yˉ∞\bar y^\inftyyˉ​∞.

Significance

The result. Theorem 1 reduces an infinite-state average-cost control problem to a one-parameter search over critical numbers. The paper then evaluates the average cost of δ[yˉ]\delta[\bar y]δ[yˉ​] by a renewal formula, proves it convex in yˉ\bar yyˉ​ (Theorem 2), and in §5 extends optimality to unlimited storage. The strong form of optimality matters: it compares with every Markov policy from every starting state, and it compares lim-infs, not only lim-sups.

Formalizing it. The theorem has a published proof, but no machine-checked one, and its proof relies on external results that are themselves unformalized: the countable-state average-cost theory of Federgruen, Schweitzer and Tijms, a fixed-point theorem on a compact convex subset of a product space, and a large-deviation estimate quoted from Feller. A formal development produces reusable infrastructure: expected first-passage sums for integer-valued random walks with a reflecting control, polynomial moment bounds for them, and the convexity-preservation argument for capacitated value iteration.

Difficulty

The state space is unbounded below, so the finite-state theory of average-cost Markov decision processes does not apply, and the one-period cost is unbounded. The obvious approach, letting the discount factor tend to one in the discounted problem, needs uniform bounds on relative value functions. Those bounds come from the expected cost until the inventory returns to a fixed interval, and with capacity limits that expectation must be controlled with growth O(∣x∣ρ+3)O(|x|^{\rho+3})O(∣x∣ρ+3) uniformly over a class of policies. This is the content of Lemma 3, whose proof combines a large-deviation estimate for the demand sums with a renewal-type recursion. A second obstacle is strong optimality: comparing with policies whose lim-inf average cost is smaller requires that the relative value function grows sublinearly along every admissible trajectory (Corollary 2).

Formalization scope

All objects are in the namespace FedergruenZipkin.AvgCost, defined in one file. States x,yx, yx,y and the capacity UUU are integers; demands are natural numbers with a real probability mass function p; bbb is a positive natural number. Convexity on Z\mathbb ZZ is the second-difference inequality. Expectations of a real function are series ∑jp(j) v(y−j)\sum_j p(j)\,v(y-j)∑j​p(j)v(y−j); expected policy costs and hitting sums are [0,∞][0,\infty][0,∞]-valued and need no integrability side condition. Feasibility and all properties of value functions are required only on states x≤Ux \le Ux≤U, which are the only states visited. Assumption 2 is stated literally, as real-analyticity of θ↦∑jp(j)eiθj\theta \mapsto \sum_j p(j)e^{i\theta j}θ↦∑j​p(j)eiθj at 000. The order cost is zero, as in the paper. yˉ∞\bar y^\inftyyˉ​∞ is a parameter characterised as the least minimizer of GGG, not an infimum.

"Strongly optimal" has no displayed definition in the paper; it is read from eq. (7) in the proof of Theorem 1(b): convergence of the average cost of δ∗\delta^*δ∗ to g∗g^*g∗ from every state, together with a lim-inf lower bound for every Markov (memoryless, possibly nonstationary) policy from every state. The class is neither widened to history-dependent policies nor narrowed to stationary ones. The goal additionally records that E v∗(y−D)E\,v^*(y-D)Ev∗(y−D) converges, that v∗v^*v∗ has growth O(∣x∣ρ+3)O(|x|^{\rho+3})O(∣x∣ρ+3), and that g∗≥0g^* \ge 0g∗≥0; all three follow from the paper's proof.

A trivializing reading is ruled out: the existence of ggg, vvv and y∗y^*y∗ is one existential, so y∗y^*y∗ cannot be decoupled from the solution of (6), and strong optimality includes the convergence of δ∗\delta^*δ∗'s own average cost to g∗g^*g∗, so g=0g = 0g=0 does not satisfy it vacuously.

Not posed: Lemma 1 (quoted from Feller, and replaceable by a Chernoff bound); the renewal formulas (10)–(11) and Theorem 2; and §5 (unlimited storage). Useful contributions include a formal theory of expected hitting sums for skip-free-upward random walks, and a proof of Lemma 3 by any route.

Selected references

  • A. Federgruen and P. Zipkin, An Inventory Model with Limited Production Capacity and Uncertain Demands I. The Average-Cost Criterion, Mathematics of Operations Research 11(2):193–207, 1986. https://doi.org/10.1287/moor.11.2.193
  • A. Federgruen and P. Zipkin, An Inventory Model with Limited Production Capacity and Uncertain Demands II. The Discounted-Cost Criterion, Mathematics of Operations Research 11(2):208–215, 1986. https://doi.org/10.1287/moor.11.2.208
  • A. Federgruen, P. J. Schweitzer and H. C. Tijms, Denumerable Undiscounted Semi-Markov Decision Processes with Unbounded Rewards, Mathematics of Operations Research 8(2):298–314, 1983. https://doi.org/10.1287/moor.8.2.298
  • D. P. Heyman and M. J. Sobel, Stochastic Models in Operations Research, Vol. II, McGraw-Hill, 1984.
  • W. Feller, An Introduction to Probability Theory and Its Applications, Vol. II, 2nd ed., Wiley, 1971.
10 thms2 active usersReviewed
Linear OptimizationOperations ResearchOptimization+1·Captain: mikedeng1

Smoothed Analysis of Algorithms: Why the Simplex Algorithm Usually Takes Polynomial Time 2: The Two-Phase Shadow-Vertex Simplex Method Has Polynomial Smoothed ComplexityResearch Paper

Motivation

The simplex method solves linear programs by moving between vertices of a feasible polyhedron. Its worst-case number of moves can grow exponentially, yet it often performs well on ordinary inputs. Worst-case examples alone therefore give an incomplete account of the method’s behavior. Spielman and Teng introduced smoothed analysis to measure expected performance after small random perturbations of an arbitrary input. Their result for a two-phase shadow-vertex simplex method gives a polynomial bound in the input dimensions and inverse perturbation scale. The pinned preprint is the source for every theorem number and constant in this mission.

The paper separates a geometric result about the expected size of a polytope’s shadow (Theorem 4.0.1) from the algorithmic result here (Theorem 5.0.1). That separation matters: a plane chosen before perturbation and a plane chosen by a running algorithm have different distributions. This mission addresses the latter. It complements the standard-form simplex theorems already formalized in the Introduction to Linear Optimization series and the worst-case Klee–Minty result in the Smale’s Ninth Problem mission; those results concern different algorithms or input models and are context rather than imported statements.

Setting

A linear program is specified by vectors a1,…,an∈Rda_1,\ldots,a_n\in\mathbb R^da1​,…,an​∈Rd, right-hand sides y1,…,yn∈Ry_1,\ldots,y_n\in\mathbb Ry1​,…,yn​∈R, and an objective vector z∈Rdz\in\mathbb R^dz∈Rd:

max⁡x⟨z,x⟩subject to⟨ai,x⟩≤yi(1≤i≤n).\max_x\langle z,x\rangle\quad\text{subject to}\quad \langle a_i,x\rangle\le y_i\qquad(1\le i\le n).xmax​⟨z,x⟩subject to⟨ai​,x⟩≤yi​(1≤i≤n).

The paper’s two-phase shadow-vertex method first draws a collection I\mathcal II of ddd-element subsets of [n][n][n] and chooses one whose constraint matrix AIA_IAI​ has the largest smallest singular value. It sets a power-of-two scale MMM from the input norm and a power-of-two scale κ\kappaκ from that singular value. These determine positive relaxed right-hand sides yi′y'_iyi′​: MMM for i∈Ii\in Ii∈I and dM2/(4κ)\sqrt d M^2/(4\kappa)d​M2/(4κ) otherwise. A coefficient vector α\alphaα is chosen uniformly from A1/d2={α:∑i∈Iαi=1, αi≥1/d2}A_{1/d^2}=\{\alpha:\sum_{i\in I}\alpha_i=1,\ \alpha_i\ge1/d^2\}A1/d2​={α:∑i∈I​αi​=1, αi​≥1/d2}. The first phase solves the relaxed program LP′ from the objective AIαA_I\alphaAI​α.

The second phase uses a lifted program LP⁺ in Rd+1\mathbb R^{d+1}Rd+1. For each original constraint it forms ai+=((yi′−yi)/2,ai)a_i^+=((y'_i-y_i)/2,a_i)ai+​=((yi′​−yi​)/2,ai​) and yi+=(yi′+yi)/2y_i^+=(y'_i+y_i)/2yi+​=(yi′​+yi​)/2, together with two artificial constraints at first coordinates 111 and −1-1−1. LP⁺ connects LP′ to the original program and makes infeasibility detectable. Its shadow is taken in the plane of (0,z)(0,z)(0,z) and z+=(1,0,…,0)z^+=(1,0,\ldots,0)z+=(1,0,…,0).

For positive right-hand sides, an optimal polar simplex is a ddd-subset of constraints whose scaled vectors ai/yia_i/y_iai​/yi​ form a facet of ConvHull⁡(0,a1/y1,…,an/yn)\operatorname{ConvHull}(0,a_1/y_1,\ldots,a_n/y_n)ConvHull(0,a1​/y1​,…,an​/yn​) and whose unscaled cone contains an objective qqq. The shadow for objectives t,zt,zt,z is the union of these simplices over all qqq in Span⁡(t,z)\operatorname{Span}(t,z)Span(t,z). Its size bounds the number of polar pivots. In Section 5 the paper writes Sz′S'_zSz′​ for the first-phase shadow size and Sz+S_z^+Sz+​ for the second-phase shadow size without the two artificial pivots.

The input is perturbed by independent Gaussians: each coordinate of aia_iai​ and each yiy_iyi​ has its prescribed center and common standard deviation σR\sigma RσR, where R=max⁡i∥(yˉi,aˉi)∥2R=\max_i\|(\bar y_i,\bar a_i)\|_2R=maxi​∥(yˉ​i​,aˉi​)∥2​. The algorithm has separate random choices of I\mathcal II and α\alphaα.

Formalization targets

The immediate targets bound the two phases: Lemma 5.2.1 gives an explicit expectation bound for Sz′S'_zSz′​ and Lemma 5.3.1 gives one for Sz+S_z^+Sz+​. Lemma 5.1.1 and its corollaries control the chance that the chosen basis has a very small singular value. Corollary 4.3.3 extends the geometric shadow bound to positive, unequal right-hand sides and general Gaussian covariance. These are the mission’s milestone targets.

The goal is the shape of Theorem 5.0.1. With C(A,y,z)=EI,α(Sz′+Sz++2)C(A,y,z)=\mathbb E_{\mathcal I,\alpha}(S'_z+S_z^++2)C(A,y,z)=EI,α​(Sz′​+Sz+​+2), there are a single polynomial P\mathcal PP and a positive constant σ0\sigma_0σ0​ such that, for all n>d≥3n>d\ge3n>d≥3 and all centers and objectives,

EA,yC(A,y,z)≤min⁡{P(d,n,1min⁡(σ,σ0)),(nd)+(nd+1)+2}.\mathbb E_{A,y}C(A,y,z)\le \min\left\{\mathcal P\left(d,n,\frac1{\min(\sigma,\sigma_0)}\right), \binom nd+\binom n{d+1}+2\right\}.EA,y​C(A,y,z)≤min{P(d,n,min(σ,σ0​)1​),(dn​)+(d+1n​)+2}.

The polynomial is uniform over the dimensions and inputs; its coefficients are not prescribed. The bound on CCC implies the corresponding result for the actual pivot count through the paper’s step-to-shadow comparison. The goal is stated with a positive center scale RRR, the case in which the paper’s Gaussian rescaling applies.

Significance

The theorem places the number of pivots of a complete simplex method under one explicit perturbation model, including the work needed to find a starting feasible basis and handle an arbitrary right-hand side. The trivial binomial bound is retained because it controls rare events in the proof and is part of the stated result. The polynomial bound says that even when the unperturbed LP is adversarial, Gaussian noise of a controlled scale makes the expected shadow-size cost polynomial.

The paper proves the mathematical result. This mission asks for machine-checked proofs of its statement and the listed milestones; the draft Lean declarations are targets with sorry, not completed proofs. The reusable formal infrastructure is the finite polar simplex and shadow construction, product Gaussian input law, smallest-singular-value events for sampled minors, and the uniform truncated-simplex coefficient law. The two shadow-size lemmas also require explicit handling of measurable finite-valued counts and their expectations.

Difficulty

The basic shadow estimate fixes its projection plane before perturbing the constraints. In LP′, the initial objective AIαA_I\alphaAI​α uses a basis selected after the perturbation, so the relevant plane depends on the random LP. The fixed-plane theorem cannot be substituted directly. For LP⁺, the normalized lifted vectors ai+/yi+a_i^+/y_i^+ai+​/yi+​ are nonlinear functions of Gaussian data; they are generally not Gaussian vectors. Thus the same shadow estimate does not apply directly to their law either. A further issue is that a poor sampled basis can make y′y'y′ very large. These are distinct obstacles, reflected in the milestone groups from Sections 5.1, 5.2, and 5.3.

Formalization scope

Vectors are EuclideanSpace ℝ (Fin d), constraints are Fin n → EuclideanSpace ℝ (Fin d), and index families are finite sets of Fin n. The paper’s [n][n][n] starts at one; Fin n starts at zero. The Gaussian constructor receives variance σ2\sigma^2σ2, not standard deviation σ\sigmaσ. The 3ndln⁡n3nd\ln n3ndlnn draws are rounded upward and are independent uniform draws with replacement. Equal singular values are resolved by the first sampled set. The uniform law on AδA_\deltaAδ​ is represented by normalized independent exponential weights followed by the affine shift that imposes αi≥δ\alpha_i\ge\deltaαi​≥δ.

The Lean definition of CCC is exactly the Section 5 shadow-size upper bound E(Sz′+Sz++2)\mathbb E(S'_z+S_z^++2)E(Sz′​+Sz+​+2), computed from the sampled LP data. It is not an arbitrary cost variable. The actual algorithmic step bound needs the paper’s polar algorithm and Lemma 3.3.5. The goal explicitly asks for inner and outer integrability so Lean’s default value for a nonintegrable Bochner integral cannot make the result vacuous. The source’s all-zero center scale is excluded because it gives zero perturbation and defeats the rescaling used in Theorem 5.0.1.

For LP⁺ the vectors live in Rd+1\mathbb R^{d+1}Rd+1, so the two LP⁺ milestone bounds use D(n,d+1,⋅)\mathcal D(n,d+1,\cdot)D(n,d+1,⋅). The preprint prints ddd in those calls even though the preceding extension theorem would be applied in dimension d+1d+1d+1. Lemma 5.2.1 is written as an inequality: its printed equality is stronger than the bound established on page 71. These corrections are visible in the theorem titles and notes. Contributions that prove the exact statements, establish the measurability and Gaussian law facts, or formalize the step-to-shadow comparison are welcome.

Selected references

  • Daniel A. Spielman and Shang-Hua Teng, Smoothed Analysis of Algorithms: Why the Simplex Algorithm Usually Takes Polynomial Time, arXiv:cs/0111050v7, 2003, preprint. The PDF used here is the 96-page version with printed and PDF page numbers aligned.
22 thms2 active usersReviewed
Operations ResearchStochastic Systems·Captain: mikedeng1

Quantifying the Bullwhip Effect in a Simple Supply Chain: The Impact of Forecasting, Lead Times, and Information 1: Centralizing Demand Information Does Not Eliminate the Bullwhip EffectResearch Paper

Motivation

The bullwhip effect is the observation that the variability of orders increases as one moves up a supply chain, from the retailer towards the manufacturer and its suppliers. It was documented in industry and in classroom experiments such as the Beer Game (Sterman 1989), and analysed by Lee, Padmanabhan and Whang (1997), who named demand forecasting, lead times, batch ordering, rationing and price variations as its main causes. A remedy often proposed is to centralize demand information: give every stage of the chain the customer demand data, so that no stage forecasts from the distorted orders of its downstream neighbour.

Chen, Drezner, Ryan and Simchi-Levi (2000) quantified the effect for a retailer that forecasts with a moving average and orders with an order-up-to policy. They gave an explicit lower bound on the ratio of the order variance to the demand variance in terms of the lead time, the forecasting window and the demand autocorrelation. They then showed that in a multistage chain with fully centralized demand information this ratio still grows with the total lead time upstream of each stage. This mission formalizes that result, Theorem 3.1 of the paper, together with the single-stage analysis it rests on.

Setting

Time is indexed by the integers. The customer demands DtD_tDt​ seen by the retailer follow the AR(1) model

Dt=μ+ρDt−1+ϵt,(1)D_t = \mu + \rho D_{t-1} + \epsilon_t, \tag{1}Dt​=μ+ρDt−1​+ϵt​,(1)

where μ≥0\mu \ge 0μ≥0, ∣ρ∣<1|\rho| < 1∣ρ∣<1, and the errors ϵt\epsilon_tϵt​ are independent and identically distributed from a symmetric distribution with mean 000 and variance σ2\sigma^2σ2. The demand is in steady state, so that E(Dt)=μ/(1−ρ)E(D_t) = \mu/(1-\rho)E(Dt​)=μ/(1−ρ) and Var(D)=Var(Dt)=σ2/(1−ρ2)\mathrm{Var}(D) = \mathrm{Var}(D_t) = \sigma^2/(1-\rho^2)Var(D)=Var(Dt​)=σ2/(1−ρ2) for every ttt.

The retailer does not know the demand process. With a window of p≥1p \ge 1p≥1 past observations it forms the moving-average estimates

D^tL=L ∑i=1pDt−ip,et=Dt−D^t1,σ^etL=CL,ρ∑i=1pet−i2p,\hat D^L_t = L\,\frac{\sum_{i=1}^p D_{t-i}}{p}, \qquad e_t = D_t - \hat D^1_t, \qquad \hat\sigma^L_{et} = C_{L,\rho}\sqrt{\frac{\sum_{i=1}^p e_{t-i}^2}{p}},D^tL​=Lp∑i=1p​Dt−i​​,et​=Dt​−D^t1​,σ^etL​=CL,ρ​p∑i=1p​et−i2​​​,

where LLL is the lead-time parameter (L=1L = 1L=1 means an order placed at the end of period ttt arrives at the start of period t+1t+1t+1) and CL,ρC_{L,\rho}CL,ρ​ is a constant the paper leaves unspecified. The order-up-to point is yt=D^tL+z σ^etLy_t = \hat D^L_t + z\,\hat\sigma^L_{et}yt​=D^tL​+zσ^etL​ for a safety factor zzz, and the order placed in period ttt is qt=yt−yt−1+Dt−1q_t = y_t - y_{t-1} + D_{t-1}qt​=yt​−yt−1​+Dt−1​. It may be negative: excess inventory is returned without cost.

In the multistage chain with centralized information, stages k=1,2,…k = 1, 2, \dotsk=1,2,… (stage 111 is the retailer) all observe DtD_tDt​ and use the same estimate D^t=∑i=1pDt−i/p\hat D_t = \sum_{i=1}^p D_{t-i}/pD^t​=∑i=1p​Dt−i​/p. Stage kkk has lead time LkL_kLk​ and safety factor zkz_kzk​ and uses the order-up-to point ytk=LkD^t+zkσ^etLky^k_t = L_k\hat D_t + z_k\hat\sigma^{L_k}_{et}ytk​=Lk​D^t​+zk​σ^etLk​​. Following the paper's sequence of events, stage 111 orders qt1=yt1−yt−11+Dt−1q^1_t = y^1_t - y^1_{t-1} + D_{t-1}qt1​=yt1​−yt−11​+Dt−1​, and stage k≥2k \ge 2k≥2, receiving qtk−1q^{k-1}_tqtk−1​, orders qtk=ytk−yt−1k+qtk−1q^k_t = y^k_t - y^k_{t-1} + q^{k-1}_tqtk​=ytk​−yt−1k​+qtk−1​.

Formalization targets

Goal: Theorem 3.1 (p. 441)

For every stage k≥1k \ge 1k≥1 and every period ttt,

Var(qtk)Var(D)≥1+(2∑i=1kLip+2(∑i=1kLi)2p2)(1−ρp),\frac{\mathrm{Var}(q^k_t)}{\mathrm{Var}(D)} \ge 1 + \left(\frac{2\sum_{i=1}^k L_i}{p} + \frac{2\left(\sum_{i=1}^k L_i\right)^2}{p^2}\right)(1-\rho^p),Var(D)Var(qtk​)​≥1+​p2∑i=1k​Li​​+p22(∑i=1k​Li​)2​​(1−ρp),

with equality when z1=⋯=zk=0z_1 = \dots = z_k = 0z1​=⋯=zk​=0. The bound holds for every choice of the constants CLk,ρC_{L_k,\rho}CLk​,ρ​ and of the safety factors.

Milestones (p. 438)

  1. The AR(1) moments Var(Dt)=σ2/(1−ρ2)\mathrm{Var}(D_t) = \sigma^2/(1-\rho^2)Var(Dt​)=σ2/(1−ρ2) and Cov(Dt−1,Dt−p−1)=ρpσ2/(1−ρ2)\mathrm{Cov}(D_{t-1}, D_{t-p-1}) = \rho^p\sigma^2/(1-\rho^2)Cov(Dt−1​,Dt−p−1​)=ρpσ2/(1−ρ2).
  2. Eq. (4): qt=(1+L/p)Dt−1−(L/p)Dt−p−1+z(σ^etL−σ^e,t−1L)q_t = (1 + L/p)D_{t-1} - (L/p)D_{t-p-1} + z(\hat\sigma^L_{et} - \hat\sigma^L_{e,t-1})qt​=(1+L/p)Dt−1​−(L/p)Dt−p−1​+z(σ^etL​−σ^e,t−1L​) for every outcome.
  3. Lemma 2.1: Cov(Dt−i,σ^etL)=0\mathrm{Cov}(D_{t-i}, \hat\sigma^L_{et}) = 0Cov(Dt−i​,σ^etL​)=0 for i=1,…,pi = 1, \dots, pi=1,…,p.
  4. The variance identity after Eq. (4):
Var(qt)=[1+(2Lp+2L2p2)(1−ρp)]Var(D)+2z(1+2Lp)Cov(Dt−1,σ^etL)+z2 Var(σ^etL−σ^e,t−1L).\mathrm{Var}(q_t) = \left[1 + \left(\tfrac{2L}{p} + \tfrac{2L^2}{p^2}\right)(1-\rho^p)\right]\mathrm{Var}(D) + 2z\left(1+\tfrac{2L}{p}\right)\mathrm{Cov}(D_{t-1}, \hat\sigma^L_{et}) + z^2\,\mathrm{Var}(\hat\sigma^L_{et} - \hat\sigma^L_{e,t-1}).Var(qt​)=[1+(p2L​+p22L2​)(1−ρp)]Var(D)+2z(1+p2L​)Cov(Dt−1​,σ^etL​)+z2Var(σ^etL​−σ^e,t−1L​).
  1. Theorem 2.2, the single-stage case:
Var(q)Var(D)≥1+(2Lp+2L2p2)(1−ρp),(5)\frac{\mathrm{Var}(q)}{\mathrm{Var}(D)} \ge 1 + \left(\frac{2L}{p} + \frac{2L^2}{p^2}\right)(1-\rho^p), \tag{5}Var(D)Var(q)​≥1+(p2L​+p22L2​)(1−ρp),(5)

with equality when z=0z = 0z=0.

Significance

Theorem 2.2 shows that forecasting with a positive lead time is enough to make orders more variable than demand, even for independent demands (ρ=0\rho = 0ρ=0). It also says how the effect depends on each parameter: the bound decreases in the window ppp and increases in the lead time LLL. Theorem 3.1 is the paper's answer to the centralization remedy. When every stage sees the true customer demand and uses the same forecast and the same policy, the variability of orders at stage kkk is still bounded below by the single-stage expression with the cumulative lead time ∑i≤kLi\sum_{i\le k}L_i∑i≤k​Li​. Centralization reduces the bullwhip effect but does not remove it. The decentralized comparison (Theorem 3.2, where the bound becomes multiplicative across stages) is a separate mission in this series.

On the formalization side, the Gaussian special case of the single-stage results is on the platform. Snyder and Shen's Fundamentals of Supply Chain Theory states Theorem 2.2, Lemma 2.1, Eq. (4) and the AR(1) moments for normally distributed errors, as the items SupplyChainTheory.bullwhip_signal_processing, bullwhip_lemma_13_1, bullwhip_order_identity and ar1_moments. This mission states them under the paper's weaker hypothesis of a symmetric error distribution. The multistage Theorem 3.1 has no machine-checked counterpart. The paper proves only Theorem 2.2 in print. For the proofs of Lemma 2.1 and Theorem 3.1 it refers to Ryan (1997) and to a working paper, so a formalization supplies arguments the published article does not contain.

Difficulty

Most of the algebra is routine. The difficulty is Lemma 2.1 and the covariances like it. The estimate σ^etL\hat\sigma^L_{et}σ^etL​ is a square root of a quadratic form in past demands, so its covariance with a demand cannot be computed from second moments. Under Gaussian errors one can appeal to properties of Gaussian vectors. With only a symmetric error law, every distributional fact has to come from the symmetry of the errors and from the representation of the steady-state demand as an infinite series in past errors.

The printed derivation also moves faster than a proof. Expanding Var(qt)\mathrm{Var}(q_t)Var(qt​) from Eq. (4) produces the cross terms Cov(Dt−1,σ^e,t−1L)\mathrm{Cov}(D_{t-1}, \hat\sigma^L_{e,t-1})Cov(Dt−1​,σ^e,t−1L​) and Cov(Dt−p−1,σ^etL)\mathrm{Cov}(D_{t-p-1}, \hat\sigma^L_{et})Cov(Dt−p−1​,σ^etL​), which lie outside the lags 1,…,p1, \dots, p1,…,p of Lemma 2.1. The display after Eq. (4) does not account for them. A complete proof of milestone 4 must show that these terms vanish too. For the chain, the stage orders are defined by a recursion across stages, and the variance of qtkq^k_tqtk​ involves the estimates σ^etLi\hat\sigma^{L_i}_{et}σ^etLi​​ of all stages i≤ki \le ki≤k.

Formalization scope

Random variables are real functions on a probability space (Ω,P)(\Omega, P)(Ω,P), and time is Z\mathbb ZZ, so that Dt−p−1D_{t-p-1}Dt−p−1​ exists for every ttt. Variance and covariance are Mathlib's ProbabilityTheory.variance and ProbabilityTheory.covariance. The demand structure ChenBullwhip.Centralized.AR1Demand records (1) for every outcome and the paper's error hypotheses: independence, identical distribution, symmetry, mean 000 and variance σ2\sigma^2σ2. It adds four disclosed conditions:

  1. σ>0\sigma > 0σ>0, since the results divide by Var(D)\mathrm{Var}(D)Var(D);
  2. square integrability of errors and demands, since Mathlib's variance of a non-square-integrable function is 000;
  3. a steady-state condition: every DtD_tDt​ is square integrable with the law of D0D_0D0​, which is the stationary solution the paper's moment formulas presuppose;
  4. p≥1p \ge 1p≥1 in every result.

The published Gaussian structure SupplyChainTheory.AR1Demand satisfies these conditions, so this mission generalizes the Snyder–Shen items rather than referencing them. The constants CL,ρC_{L,\rho}CL,ρ​ are free real parameters, and in the chain CLk,ρC_{L_k,\rho}CLk​,ρ​ is C(Lk)C(L_k)C(Lk​) for an arbitrary function CCC. Lead times are natural numbers, L=0L = 0L=0 included. Sums ∑i=1p\sum_{i=1}^p∑i=1p​ and ∑i=1k\sum_{i=1}^k∑i=1k​ run over {1,…,p}\{1,\dots,p\}{1,…,p} and {1,…,k}\{1,\dots,k\}{1,…,k}, and stages are numbered from 111. The order recursion qtk=ytk−yt−1k+qtk−1q^k_t = y^k_t - y^k_{t-1} + q^{k-1}_tqtk​=ytk​−yt−1k​+qtk−1​ is read from the paper's sequence of events, because the paper prints no formula for qtkq^k_tqtk​.

The orders are computed from the demands through the definitions above. They are never arbitrary random variables with assumed moments. "Tight" is formalized as equality, and orders are never truncated at zero. Without the steady-state condition, a process started from an arbitrary D0D_0D0​ satisfies (1) but has time-dependent moments, and the results fail; with σ=0\sigma = 0σ=0 the ratio form would be false. Both cases are excluded by the structure, not by vacuous hypotheses. The structure is satisfiable: i.i.d. standard Gaussian demands on Z→R\mathbb Z \to \mathbb RZ→R form an instance.

A complete development needs: the L2L^2L2 series representation of a stationary AR(1) process; distributional symmetry facts for i.i.d. sequences with a symmetric law; and covariance bookkeeping for finite linear combinations. The first two are reusable for any linear time-series model with symmetric innovations. Contributions to any milestone, and to general lemmas about stationary AR(1) processes, are welcome.

Selected references

  • F. Chen, Z. Drezner, J. K. Ryan, D. Simchi-Levi, Quantifying the Bullwhip Effect in a Simple Supply Chain: The Impact of Forecasting, Lead Times, and Information, Management Science 46(3):436–443, 2000. https://doi.org/10.1287/mnsc.46.3.436.12069
  • H. L. Lee, V. Padmanabhan, S. Whang, Information Distortion in a Supply Chain: The Bullwhip Effect, Management Science 43(4):546–558, 1997. https://doi.org/10.1287/mnsc.43.4.546
  • J. D. Sterman, Modeling Managerial Behavior: Misperceptions of Feedback in a Dynamic Decision Making Experiment, Management Science 35(3):321–339, 1989. https://doi.org/10.1287/mnsc.35.3.321
  • L. V. Snyder, Z.-J. M. Shen, Fundamentals of Supply Chain Theory, 2nd ed., Wiley, 2019, Chapter 13. https://doi.org/10.1002/9781119584445
  • J. K. Ryan, Analysis of Inventory Models with Limited Demand Information, Ph.D. dissertation, Northwestern University, 1997 (cited by the paper for the proofs of Lemma 2.1 and Theorem 3.1).
9 thms2 active usersReviewed
Algorithmic Game TheoryMechanism DesignOperations Research·Captain: mikedeng1

Multi-parameter Mechanism Design and Sequential Posted Pricing 3: Order-Oblivious Posted Prices 2-Approximate the Optimal Revenue under a Uniform Matroid ConstraintResearch Paper

Motivation

Myerson's optimal auction (Myerson 1981) maximizes a seller's expected revenue when buyers have independent private values, but it is a sealed-bid mechanism: every buyer reports a value, and the allocation and payments are computed from all reports at once. Real sellers more often post prices: buyers arrive, each sees a take-it-or-leave-it price, and buys or leaves. Chawla, Hartline, Malec and Sivan (arXiv:0907.2435) ask how much revenue such simple mechanisms lose. Their strongest notion is the order-oblivious posted-price mechanism (OPM): the prices are fixed in advance, and the guarantee must hold whatever order the buyers arrive in, even an adversarial one.

The tool behind the guarantee for sellers of kkk identical units is a prophet inequality. In the single-choice version, a gambler inspects independent random rewards one at a time and must accept or reject each on the spot; Krengel and Sucheston, and Samuel-Cahn (Ann. Probab. 1984), showed that a single fixed threshold earns at least half of what a prophet who sees all rewards earns. The paper extends Samuel-Cahn's threshold rule to kkk choices (Appendix D.2) and turns it into a revenue guarantee (Theorem 10).

Setting

There are nnn agents [n][n][n]. Agent iii's value viv_ivi​ for being served is drawn independently from a distribution FiF_iFi​ with density fif_ifi​; the virtual valuation is ϕi(v)=v−(1−Fi(v))/fi(v)\phi_i(v) = v - (1 - F_i(v))/f_i(v)ϕi​(v)=v−(1−Fi​(v))/fi​(v) (Definition 1), and FiF_iFi​ is regular if ϕi\phi_iϕi​ is non-decreasing (Definition 2). The seller may serve any set of agents in a downward-closed set system J\mathcal JJ; this mission uses the kkk-uniform matroid, where a set is feasible exactly when it has at most kkk members.

A mechanism MMM maps reported values v\mathbf vv to an allocation M(v)∈JM(\mathbf v) \in \mathcal JM(v)∈J and payments πi(v)\pi_i(\mathbf v)πi​(v). It is truthful if reporting the true value is a dominant strategy and no agent ever gets negative utility. Its expected revenue is RM=Ev[∑iπi(v)]\mathcal R^M = \mathbb E_{\mathbf v}[\sum_i \pi_i(\mathbf v)]RM=Ev​[∑i​πi​(v)], and RM\mathcal R^{\mathcal M}RM denotes the revenue of Myerson's mechanism, the largest over truthful mechanisms (Theorem 19).

Given prices p\mathbf pp and values v\mathbf vv, agent iii desires service if vi≥piv_i \ge p_ivi​≥pi​. Let Sv\mathcal S_{\mathbf v}Sv​ be the class of maximal feasible sets of desiring agents. When agents arrive in an arbitrary order and each buys if it desires service and can still be feasibly served, the set of buyers lies in Sv\mathcal S_{\mathbf v}Sv​. The paper's pessimistic revenue estimate is

Rpobl=Ev∼F min⁡S∈Sv∑i∈Spi.\mathcal R^{\mathrm{obl}}_{\mathbf p} = \mathbb E_{\mathbf v \sim \mathbf F}\ \min_{S \in \mathcal S_{\mathbf v}} \sum_{i \in S} p_i .Rpobl​=Ev∼F​ S∈Sv​min​i∈S∑​pi​.

For the prophet inequality, X1,…,XnX_1, \dots, X_nX1​,…,Xn​ are independent nonnegative random variables with order statistics X(1)≥⋯≥X(n)X_{(1)} \ge \dots \ge X_{(n)}X(1)​≥⋯≥X(n)​, and (x)+=max⁡(0,x)(x)^+ = \max(0, x)(x)+=max(0,x). The threshold rule with threshold ccc picks indices t1(c),…,tk(c)t_1(c), \dots, t_k(c)t1​(c),…,tk​(c), where ti(c)t_i(c)ti​(c) is the lesser of n−k+in-k+in−k+i and the iii-th smallest index jjj with Xj≥cX_j \ge cXj​≥c (or n−k+in - k + in−k+i if there is none). The numbers a∗a^*a∗ and b∗b^*b∗ are the unique solutions of

a=∑i=1kE(X(i)−a/k)+,b=∑i=1nE(Xi−b/k)+.a = \sum_{i=1}^k \mathbb E\big(X_{(i)} - a/k\big)^+, \qquad b = \sum_{i=1}^n \mathbb E\big(X_i - b/k\big)^+ .a=i=1∑k​E(X(i)​−a/k)+,b=i=1∑n​E(Xi​−b/k)+.

Formalization targets

Goal: Theorem 10 (p. 9)

∃ p  ∀M truthful:RM≤2 Rpobl\exists\, \mathbf p\ \ \forall M \text{ truthful}:\qquad \mathcal R^M \le 2\, \mathcal R^{\mathrm{obl}}_{\mathbf p}∃p  ∀M truthful:RM≤2Rpobl​

for every instance with regular distributions and a kkk-uniform matroid constraint. The prices are chosen once, before the mechanism it is compared with; this is the paper's "Rpobl\mathcal R^{\mathrm{obl}}_{\mathbf p}Rpobl​ 2-approximates RM\mathcal R^{\mathcal M}RM".

Milestones

  1. Proposition 1 (p. 5): under regularity, the expected revenue of a truthful mechanism equals its expected virtual surplus E[∑i∈M(v)ϕi(vi)]\mathbb E[\sum_{i \in M(\mathbf v)} \phi_i(v_i)]E[∑i∈M(v)​ϕi​(vi​)] (with the lowest type receiving zero utility).
  2. a∗a^*a∗ and b∗b^*b∗ exist and are unique (App. D.2, p. 18).
  3. The claim a∗≤b∗a^* \le b^*a∗≤b∗ (App. D.2, p. 18).
  4. Theorem 24 (p. 18), the kkk-choice prophet inequality: for a∗≤kc≤b∗a^* \le k c \le b^*a∗≤kc≤b∗,
∑i=1kE[X(i)]≤2∑i=1kE[Xti(c)].\sum_{i=1}^k \mathbb E\big[X_{(i)}\big] \le 2 \sum_{i=1}^k \mathbb E\big[X_{t_i(c)}\big].i=1∑k​E[X(i)​]≤2i=1∑k​E[Xti​(c)​].

Significance

The theorem says that a seller of kkk identical units can fix one price per buyer, ignore the arrival order entirely, and still collect half of the optimal revenue. The factor 2 is tight: Appendix D.2 gives a single-item example with two buyers where no order-oblivious pricing does better. Corollary 11 extends the result to partition matroids, and Theorem 24 is reused for the graphical-matroid result (Theorem 12, App. D.3). Theorem 24 is a statement in optimal stopping independent of mechanism design, and kkk-choice prophet inequalities are now a standard tool for online allocation.

The results are proved in the paper (preprint arXiv:0907.2435v2; a conference version appeared at STOC 2010). To our knowledge none of them, nor any prophet inequality, has a machine-checked proof; Mathlib has independence of random variables but no order statistics, stopping-rule prophet inequalities, or Myerson's revenue characterization in this multi-agent dominant-strategy form. A related single-unit, Bayesian-incentive-compatible form of Proposition 1 exists on the platform (MechanismDesign.Auctions.revenue_eq_virtual_surplus), in a different model.

Difficulty

The threshold rule picks the first values above ccc, not the largest, and its picks are dependent random indices; the expectation E[Xti(c)]\mathbb E[X_{t_i(c)}]E[Xti​(c)​] does not factor. The obvious comparison of the gambler with the prophet term by term fails, because the gambler can exhaust its kkk picks on early, small values. The bound has to balance two events: either at least kkk values reach ccc, or a value is picked whenever it exceeds ccc; independence enters exactly in the second. The rule also has forced picks at the end of the sequence, which must be handled as stated.

On the mechanism side, Rpobl\mathcal R^{\mathrm{obl}}_{\mathbf p}Rpobl​ is a minimum over an adversarially chosen family, not the revenue of one run, so it cannot be read off from a single sequential mechanism. Proposition 1 needs the full revenue-equivalence argument: monotone allocations, the payment identity, and an integration by parts against the density.

Formalization scope

  • Distributions: each FiF_iFi​ has a bounded support [v‾i,v‾i][\underline v_i, \overline v_i][v​i​,vi​] with 0≤v‾i0 \le \underline v_i0≤v​i​, and a measurable density positive on it (a pinned convention; the paper says only "with density fif_ifi​"). Regularity is required on the support. The prior is the product of the marginals.
  • Mechanisms: deterministic, dominant-strategy incentive compatible and ex-post individually rational on the type space, with measurable allocation events and measurable, integrable payments. Payments of unserved agents are not forced to be zero.
  • RM\mathcal R^{\mathcal M}RM is not constructed. The goal is stated against every truthful mechanism, which by Theorem 19 is equivalent. Quantifier order matters: "for every mechanism there are prices" is a weaker statement and is not the goal.
  • Proposition 1 carries the normalization that an agent of the lowest type gets zero utility, which the paper presupposes on p. 12.
  • Rpobl\mathcal R^{\mathrm{obl}}_{\mathbf p}Rpobl​ is a genuine minimum over the finite, nonempty family Sv\mathcal S_{\mathbf v}Sv​; maximality is essential, since without it the empty set makes the estimate 000 and the goal false. Prices are arbitrary reals.
  • a∗a^*a∗ and b∗b^*b∗ are characterised by their equations as hypotheses, not defined by an infimum. Order statistics count multiplicity. Lean indices are 0-based. The threshold rule includes the page's forced picks ti(c)=n−k+it_i(c) = n - k + iti​(c)=n−k+i; it is not replaced by a pure threshold rule. Theorem 24 and the claims about a∗,b∗a^*, b^*a∗,b∗ assume 1≤k≤n1 \le k \le n1≤k≤n; the goal assumes nothing about kkk.
  • Out of scope: non-regular distributions (ironing), Corollary 11, and the p. 19 identity rewriting Rpobl\mathcal R^{\mathrm{obl}}_{\mathbf p}Rpobl​ as a sum of virtual values (a proof step, not a milestone).

Useful reusable infrastructure: order statistics and their measurability, Samuel-Cahn-type threshold rules, and Myerson's payment identity for dominant-strategy mechanisms. Proofs of any milestone, and supporting lemmas on these objects, are welcome.

Selected references

  • S. Chawla, J. D. Hartline, D. Malec, B. Sivan, Multi-parameter Mechanism Design and Sequential Posted Pricing, arXiv:0907.2435v2, 2010 (STOC 2010). https://arxiv.org/abs/0907.2435
  • R. B. Myerson, Optimal Auction Design, Mathematics of Operations Research 6(1), 1981. https://doi.org/10.1287/moor.6.1.58
  • E. Samuel-Cahn, Comparison of Threshold Stop Rules and Maximum for Independent Nonnegative Random Variables, Annals of Probability 12(4), 1984. https://doi.org/10.1214/aop/1176993150
10 thms2 active usersReviewed
Algorithmic Game TheoryCombinatoricsMechanism Design+1·Captain: mikedeng1

Multi-parameter Mechanism Design and Sequential Posted Pricing 2: Sequential Posted Prices e/(e−1)-Approximate the Optimal Revenue under a Partition Matroid ConstraintResearch Paper

Motivation

A seller who knows the distributions of buyers' values can maximise expected revenue with Myerson's optimal mechanism (Myerson 1981): collect bids, compute virtual values, serve a feasible set of maximum virtual surplus, and charge threshold payments. Real sellers seldom run such auctions. They post prices: a buyer is offered a take-it-or-leave-it price and either accepts or walks away. Posted prices need no bidding, involve no competition between buyers, and are trivially truthful. The question is how much revenue they give up.

Chawla, Hartline, Malec and Sivan (arXiv:0907.2435, STOC 2010) answer this for a range of feasibility constraints with a single construction, the sequential posted-price mechanism (SPM) S\mathcal SS. For general matroids it loses at most a factor 222 (Theorem 5). For uniform and partition matroids, that is, multi-unit sales and unions of multi-unit sales, it loses at most a factor e/(e−1)≈1.58e/(e-1)\approx1.58e/(e−1)≈1.58 (Theorem 6), and the paper shows this factor is tight for its mechanism. This mission targets Theorem 6.

Timeline. Myerson (1981) characterised the optimal single-parameter mechanism. Blumrosen and Holenstein (2008) showed that the best single-unit SPM can be a factor π/2\sqrt{\pi/2}π/2​ below Myerson's revenue even with i.i.d. buyers. Chawla, Hartline and Kleinberg (EC 2007) used posted prices to approximate multi-parameter unit-demand pricing. Chawla, Hartline, Malec and Sivan (2010) gave the matroid, partition-matroid and matroid-intersection bounds. Yan (SODA 2011) explained the e/(e−1)e/(e-1)e/(e−1) factor through the correlation gap of submodular functions and sharpened it for kkk units to 1−kke−k/k!1-k^ke^{-k}/k!1−kke−k/k!.

Setting

There are nnn agents. Agent iii has a private value viv_ivi​ for being served, drawn independently from a distribution FiF_iFi​ with density fif_ifi​. The virtual value is φi(v)=v−1−Fi(v)fi(v)\varphi_i(v)=v-\frac{1-F_i(v)}{f_i(v)}φi​(v)=v−fi​(v)1−Fi​(v)​, and FiF_iFi​ is regular if φi\varphi_iφi​ is non-decreasing. The seller may serve any set in a downward-closed family J⊆2[n]\mathcal J\subseteq2^{[n]}J⊆2[n].

A partition matroid assigns each agent iii to a part part(i)\mathrm{part}(i)part(i) and each part bbb a capacity cap(b)∈N\mathrm{cap}(b)\in\mathbb Ncap(b)∈N. A set is feasible iff it contains at most cap(b)\mathrm{cap}(b)cap(b) agents of every part bbb. With one part of capacity kkk this is the kkk-uniform matroid: at most kkk agents are served.

A truthful mechanism MMM maps a value vector v\mathbf vv to a feasible set M(v)M(\mathbf v)M(v) and payments πi(v)\pi_i(\mathbf v)πi​(v). It is dominant-strategy incentive compatible and individually rational. Its expected revenue is RM=E[∑iπi(v)]\mathcal R^M=\mathbb E[\sum_i\pi_i(\mathbf v)]RM=E[∑i​πi​(v)], and qiM=Pr⁡[i∈M(v)]q^M_i=\Pr[i\in M(\mathbf v)]qiM​=Pr[i∈M(v)] is its service probability for agent iii.

The sequential posted-price mechanism S\mathcal SS built from MMM sets the price pi=Fi−1(1−qiM)p_i=F_i^{-1}(1-q^M_i)pi​=Fi−1​(1−qiM​) for agent iii, so that agent iii accepts with probability exactly qiMq^M_iqiM​. It approaches the agents one at a time in decreasing order of price (σ\sigmaσ is the ordering). It offers agent iii the price pip_ipi​ if adding iii to the agents already served keeps the set feasible. The agent accepts iff pi≤vip_i\le v_ipi​≤vi​. Its expected revenue is Rpσ\mathcal R^\sigma_{\mathbf p}Rpσ​.

Formalization targets

Goal: Theorem 6, partition matroids

For every partition matroid, every truthful MMM, and S\mathcal SS built from MMM as above,

RM≤ee−1 Rpσ.\mathcal R^M\le\frac{e}{e-1}\,\mathcal R^\sigma_{\mathbf p}.RM≤e−1e​Rpσ​.

Taking MMM to be Myerson's mechanism gives the paper's statement.

Milestones

  1. Lemma 2 (regular case): RM≤∑ipiMqiM\mathcal R^M\le\sum_ip^M_iq^M_iRM≤∑i​piM​qiM​ with piM=Fi−1(1−qiM)p^M_i=F_i^{-1}(1-q^M_i)piM​=Fi−1​(1−qiM​).
  2. Rank bound (§4): ∑i∈SqiM≤rank⁡(S)\sum_{i\in S}q^M_i\le\operatorname{rank}(S)∑i∈S​qiM​≤rank(S) for every set SSS; for a part bbb this reads ∑part(i)=bqiM≤cap(b)\sum_{\mathrm{part}(i)=b}q^M_i\le\mathrm{cap}(b)∑part(i)=b​qiM​≤cap(b).
  3. Single-unit revenue formula (App. C.2): RS=∑kckpkqk\mathcal R^{\mathcal S}=\sum_kc_kp_kq_kRS=∑k​ck​pk​qk​ with ck=∏j<k(1−qj)c_k=\prod_{j<k}(1-q_j)ck​=∏j<k​(1−qj​), positions in offer order.
  4. Lemma 20: with ppp defined by ∑kpkqk=p∑kqk\sum_kp_kq_k=p\sum_kq_k∑k​pk​qk​=p∑k​qk​ (equation (2)) and prices decreasing, p∑kckqk≤∑kckpkqkp\sum_kc_kq_k\le\sum_kc_kp_kq_kp∑k​ck​qk​≤∑k​ck​pk​qk​.
  5. Display (3): if ∑kqk=s≤1\sum_kq_k=s\le1∑k​qk​=s≤1, then p∑kckqk=p(1−∏k(1−qk))≥p(1−(1−s/n)n)≥(1−1/e)psp\sum_kc_kq_k=p(1-\prod_k(1-q_k))\ge p(1-(1-s/n)^n)\ge(1-1/e)psp∑k​ck​qk​=p(1−∏k​(1−qk​))≥p(1−(1−s/n)n)≥(1−1/e)ps.
  6. Theorem 21: the goal for the 111-uniform matroid.
  7. Theorem 22: the goal for the kkk-uniform matroid, every kkk.

Significance

The theorem shows that a mechanism with no bidding loses at most about 37%37\%37% of the optimal revenue when the constraint is a union of multi-unit supplies. That covers selling several kinds of goods, each in limited stock, to single-minded buyers. The prices are computed once from the distributions. The order is fixed before any value is seen. No agent's payment depends on another agent's report. The factor is tight for this mechanism (App. C.2), and the same template (prices from service probabilities, decreasing order) gives factor 222 for all matroids and m+1m+1m+1 for intersections of mmm matroids.

On the formal side, the paper's results are proved but none is machine-checked as far as we know. The platform has Myerson-type results for a single unit with Bayesian incentive compatibility and a common support (Börgers), and for i.i.d. buyers with a fixed number of units (Talluri and van Ryzin). Neither covers independent, non-identical buyers under a set-system constraint with dominant-strategy truthfulness. A complete development would include the ex-ante revenue bound of Lemma 2 for regular distributions, which is reusable for any posted-price or prophet-inequality argument, and the 1−1/e1-1/e1−1/e correlation-gap inequality.

Difficulty

The obvious argument compares S\mathcal SS with Myerson's mechanism one agent at a time. That fails, because S\mathcal SS may stop offering to an agent once the units of its part are gone, and the agents blocked this way can be the ones Myerson's mechanism serves. The loss has to be bounded in aggregate, using only the ex-ante constraint ∑part(i)=bqi≤cap(b)\sum_{\mathrm{part}(i)=b}q_i\le\mathrm{cap}(b)∑part(i)=b​qi​≤cap(b). The single-unit case reduces to an inequality about products ∏(1−qj)\prod(1-q_j)∏(1−qj​). For kkk units, the printed proof (pp. 15–16) is an induction that compares the run with a hypothetical single-unit instance with probabilities qi/kq_i/kqi​/k. Its second case is informal, so a formal proof needs its own argument for the kkk-unit bound. Passing from uniform to partition matroids needs the observation that with a global order the run inside each part depends only on that part's agents. Lemma 2 needs the revenue-curve concavity that regularity gives, stated through densities rather than derivatives.

Formalization scope

  • Values. Each FiF_iFi​ has a density that is measurable and strictly positive on a bounded interval [v‾i,vˉi][\underline v_i,\bar v_i][v​i​,vˉi​] with 0≤v‾i0\le\underline v_i0≤v​i​, integrates to 111 there, and has no mass outside. The paper says only "with density fif_ifi​". This pin rules out point masses, so the randomised-price variant of S\mathcal SS never arises. The prior is the product of these laws.
  • Regularity is φi\varphi_iφi​ non-decreasing on the support, assumed for every agent in the goal, as in the paper's §4 analyses. The non-regular extension (second paragraph of Lemma 2, Appendix E) is out of scope.
  • Truthful means deterministic, dominant-strategy incentive compatible over the support, ex-post individually rational, feasible, with measurable allocations and integrable payments.
  • Prices are arguments tied to MMM by pi∈[v‾i,vˉi]p_i\in[\underline v_i,\bar v_i]pi​∈[v​i​,vˉi​] and Fi(pi)=1−qiMF_i(p_i)=1-q^M_iFi​(pi​)=1−qiM​. No inverse distribution function is defined.
  • Order. The order σ\sigmaσ is a permutation with σ(0)\sigma(0)σ(0) first, decreasing in price, and ties are arbitrary. It is global across parts. Parts of capacity 000 are allowed.
  • Constant. The constant is exactly e/(e−1)e/(e-1)e/(e−1).
  • No free prices. The theorem is not stated with free or existentially chosen prices. Prices are pinned to MMM's service probabilities, and S\mathcal SS uses the same constraint as MMM. A statement in which the prices could be chosen after the fact, or in which MMM were not required to be individually rational, would be a different or false theorem.
  • Every truthful MMM. The comparison is with every truthful MMM, not with a constructed Myerson mechanism. This form is at least as strong as the paper's, and it is what the paper's proof shows.

Welcome contributions: proofs of the algebraic milestones (Lemma 20, display (3)), the revenue formula, Lemma 2 (reusable payment-identity infrastructure for dominant-strategy mechanisms), and a correlation-gap argument for kkk units.

Selected references

  • S. Chawla, J. D. Hartline, D. L. Malec, B. Sivan, Multi-parameter Mechanism Design and Sequential Posted Pricing, STOC 2010; arXiv:0907.2435v2, 2010. https://arxiv.org/abs/0907.2435
  • R. B. Myerson, Optimal Auction Design, Mathematics of Operations Research 6(1), 1981. https://doi.org/10.1287/moor.6.1.58
  • L. Blumrosen, T. Holenstein, Posted Prices vs. Negotiations: An Asymptotic Analysis, ACM EC 2008.
  • Q. Yan, Mechanism Design via Correlation Gap, ACM-SIAM SODA 2011.
12 thms2 active usersReviewed
Algorithmic Game TheoryCombinatoricsMechanism Design+1·Captain: mikedeng1

Multi-parameter Mechanism Design and Sequential Posted Pricing 1: Sequential Posted Prices 2-Approximate the Optimal Revenue under a Matroid ConstraintResearch Paper

Why posted prices

A seller who must decide whom to serve among several buyers with private values can, in principle, run Myerson's revenue-optimal mechanism: collect bids, compute virtual values, serve the feasible set of largest virtual surplus, and charge threshold payments (Myerson 1981). In practice sellers rarely do this. Retail, ticketing and online platforms mostly use posted prices: each buyer is offered a take-it-or-leave-it price and accepts if and only if the price does not exceed the buyer's value. Posted prices are simple to explain, are trivially truthful, and do not require buyers to reveal their values.

Chawla, Hartline, Malec and Sivan (arXiv:0907.2435, STOC 2010) asked how much revenue is lost by this simplification, and showed that for a wide range of feasibility constraints a sequential posted-price mechanism recovers a constant fraction of the optimal revenue. The matroid case, a factor of 2, is the first and most widely cited of their results. It is a revenue analogue of the prophet inequality and was one of the starting points of the literature on "simple versus optimal" mechanisms.

Setting

There are nnn single-parameter agents, indexed by [n][n][n], and one seller. Agent iii has a private value viv_ivi​ for being served, drawn independently from a distribution FiF_iFi​ with density fif_ifi​. The virtual valuation of agent iii is

ϕi(vi)=vi−1−Fi(vi)fi(vi),\phi_i(v_i) = v_i - \frac{1 - F_i(v_i)}{f_i(v_i)},ϕi​(vi​)=vi​−fi​(vi​)1−Fi​(vi​)​,

and FiF_iFi​ is regular if ϕi\phi_iϕi​ is non-decreasing.

The seller faces a feasibility constraint: a downward-closed family J\mathcal JJ of subsets of [n][n][n], the sets of agents that can be served together. The rank of a set SSS is rank⁡(S)=max⁡S′⊆S, S′∈J∣S′∣\operatorname{rank}(S) = \max_{S' \subseteq S,\, S' \in \mathcal J} |S'|rank(S)=maxS′⊆S,S′∈J​∣S′∣. The constraint is a matroid if it satisfies the augmentation axiom: whenever A,B∈JA, B \in \mathcal JA,B∈J and ∣A∣>∣B∣|A| > |B|∣A∣>∣B∣, some e∈A∖Be \in A \setminus Be∈A∖B has B∪{e}∈JB \cup \{e\} \in \mathcal JB∪{e}∈J. Examples are kkk identical units (kkk-uniform matroids) and disjoint markets with separate capacities (partition matroids).

A mechanism MMM maps reported values v\mathbf vv to a feasible set M(v)∈JM(\mathbf v) \in \mathcal JM(v)∈J of served agents and a payment πi(v)\pi_i(\mathbf v)πi​(v) for each agent. It is truthful if reporting the true value is a dominant strategy and no agent ends with negative utility. Its expected revenue is RM=E[∑iπi(v)]\mathcal R^M = \mathbb E[\sum_i \pi_i(\mathbf v)]RM=E[∑i​πi​(v)], and qiM=Pr⁡[i∈M(v)]q^M_i = \Pr[i \in M(\mathbf v)]qiM​=Pr[i∈M(v)] is the probability that it serves agent iii.

A sequential posted-price mechanism (SPM) with ordering σ\sigmaσ and prices p\mathbf pp approaches the agents in the order σ\sigmaσ. When agent iii's turn comes, if adding iii to the set AAA of agents served so far keeps AAA feasible, iii is offered price pip_ipi​ and is served (and pays pip_ipi​) if pi≤vip_i \le v_ipi​≤vi​; otherwise iii is blocked. Its expected revenue is Rpσ\mathcal R^\sigma_{\mathbf p}Rpσ​.

The mechanism S\mathcal SS of the paper sets pi=Fi−1(1−qiM)p_i = F_i^{-1}(1 - q^M_i)pi​=Fi−1​(1−qiM​), so that agent iii accepts an offer with probability exactly qiMq^M_iqiM​, and approaches the agents in decreasing order of price.

Formalization targets

Goal: Theorem 5

For regular, independent values and a matroid constraint, for every truthful mechanism MMM and the SPM S\mathcal SS built from its service probabilities,

RM≤2 Rpσ.\mathcal R^M \le 2\, \mathcal R^\sigma_{\mathbf p}.RM≤2Rpσ​.

Taking MMM to be Myerson's optimal mechanism gives the paper's statement that S\mathcal SS 2-approximates the optimal revenue.

Milestones

  1. Proposition 1 (p. 5): the expected revenue of a truthful mechanism equals its expected virtual surplus E[∑i∈M(v)ϕi(vi)]\mathbb E[\sum_{i \in M(\mathbf v)} \phi_i(v_i)]E[∑i∈M(v)​ϕi​(vi​)].
  2. Lemma 2 (p. 5): RM≤∑ipiMqiM\mathcal R^M \le \sum_i p^M_i q^M_iRM≤∑i​piM​qiM​ with piM=Fi−1(1−qiM)p^M_i = F_i^{-1}(1 - q^M_i)piM​=Fi−1​(1−qiM​).
  3. Revenue of an SPM (§2.2, p. 4): Rpσ=∑iciqipi\mathcal R^\sigma_{\mathbf p} = \sum_i c_i q_i p_iRpσ​=∑i​ci​qi​pi​, where cic_ici​ is the probability that agent iii is offered service and qi=1−Fi(pi)q_i = 1 - F_i(p_i)qi​=1−Fi​(pi​).
  4. Rank bound (§4, p. 6): ∑i∈SqiM≤rank⁡(S)\sum_{i \in S} q^M_i \le \operatorname{rank}(S)∑i∈S​qiM​≤rank(S) for every set SSS.
  5. Lost revenue (proof of Theorem 5, p. 7): in any run under a matroid, with prices in decreasing order and weights qqq satisfying the rank bound, ∑i blockedpiqi≤∑i servedpi\sum_{i \text{ blocked}} p_i q_i \le \sum_{i \text{ served}} p_i∑i blocked​pi​qi​≤∑i served​pi​.
  6. Half of the benchmark (p. 7): under the same conditions, ∑ipiqi≤2Rpσ\sum_i p_i q_i \le 2 \mathcal R^\sigma_{\mathbf p}∑i​pi​qi​≤2Rpσ​.

Significance

The theorem shows that under a matroid constraint the optimal mechanism's advantage over a single round of posted prices is at most a factor of 2, uniformly over all regular distributions. Prices, rather than an auction, then suffice up to a constant, which justifies posted pricing in settings where an auction is impractical. The same argument, with the matroid replaced by an intersection of mmm matroids, gives the paper's Theorems 7 and 8, and the bound underlies the analysis of VCG with reserve prices (Theorem 32). Lemma 2's benchmark ∑ipiMqiM\sum_i p^M_i q^M_i∑i​piM​qiM​ became a standard tool for "ex ante relaxation" arguments.

The results are proved in the paper; none of them has a machine-checked proof. Formalizing them requires Myerson's revenue characterization in a multi-agent, dominant-strategy setting with a general feasibility constraint, which Lean's libraries do not have, and a probabilistic analysis of a sequential process over a product measure. Related formalizations exist for narrower models: the single-unit, Bayesian incentive compatible revenue identity MechanismDesign.Auctions.revenue_eq_virtual_surplus (Börgers' textbook, common support) and the i.i.d. multi-unit RevenueManagement.revenue_equivalence. Neither covers per-agent supports, set-system constraints or dominant-strategy truthfulness.

Difficulty

The obvious argument compares the SPM with the hypothetical mechanism that ignores the feasibility constraint, whose revenue is exactly ∑ipiqi\sum_i p_i q_i∑i​pi​qi​. The SPM loses the revenue of agents who would have accepted but are blocked. The difficulty is that blocking is correlated with the values of earlier agents, and the lost revenue must be bounded by revenue actually collected. A naive per-agent charge fails in a general matroid, because one served agent can block many others; the bound has to use the matroid's rank structure together with the decreasing price order.

On the mechanism side, Lemma 2 needs the full Myerson theory: monotonicity of truthful allocations, the payment identity, and an optimization over interim allocation rules with a fixed service probability, where regularity is used.

Formalization scope

All objects live in the namespace CHMSPricing.SpmMatroid. Agents are Fin n. The following conventions are fixed.

  • Distributions. Each FiF_iFi​ is given by a measurable density, strictly positive on a bounded interval [v‾i,v‾i][\underline v_i, \overline v_i][v​i​,vi​] with 0≤v‾i<v‾i0 \le \underline v_i < \overline v_i0≤v​i​<vi​, integrating to 111 there, with no mass outside. There are no point masses, so the randomized variant of S\mathcal SS in §4 does not arise. The prior is the product measure.
  • Regularity. Monotone non-decreasing virtual values on the support (Definition 2). All goals assume regular distributions, as the body's analyses do; the non-regular case (the second paragraph of Lemma 2, Appendix E) uses randomized prices and is out of scope.
  • Truthfulness. Deterministic mechanisms, dominant-strategy incentive compatible with deviations within the support, ex-post individually rational, feasible on the type space, with measurable allocation events and measurable integrable payments. Payments of unserved agents are not forced to zero.
  • Benchmark. The goal is stated for every truthful MMM, with S\mathcal SS built from MMM's own service probabilities; this is stronger than comparing with Myerson's mechanism alone and avoids constructing it.
  • Prices. pi=Fi−1(1−qi)p_i = F_i^{-1}(1 - q_i)pi​=Fi−1​(1−qi​) is passed as an argument with the hypotheses pi∈[v‾i,v‾i]p_i \in [\underline v_i, \overline v_i]pi​∈[v​i​,vi​] and Fi(pi)=1−qiF_i(p_i) = 1 - q_iFi​(pi​)=1−qi​, rather than through a generalized inverse.
  • SPM. Positions are 000-based; the price belongs to the agent; acceptance is pi≤vip_i \le v_ipi​≤vi​; ties in the decreasing price order are arbitrary, and the goal holds for every such order.
  • Proposition 1 additionally assumes the normalization that an agent with value v‾i\underline v_iv​i​ has zero utility, which is how the paper's payments are pinned down.

The goal cannot be trivialized by a free choice of prices: the prices are tied to the mechanism's service probabilities, and the SPM uses the same matroid as the mechanism. Individual rationality is essential, since without it a "truthful" mechanism can extract unbounded revenue.

A complete development needs Myerson's lemma for dominant-strategy single-parameter mechanisms, a quantile/revenue-curve argument under regularity, matroid span and rank facts for the paper's finite set systems, and independence arguments for a sequential process on a product measure. The rank bound and the deterministic lost-revenue inequality are independent of the probabilistic parts and are good first contributions; the Myerson-side lemmas are reusable for the other missions of this series.

Selected references

  • S. Chawla, J. D. Hartline, D. Malec, B. Sivan, Multi-parameter Mechanism Design and Sequential Posted Pricing, arXiv:0907.2435v2, 2010; STOC 2010. https://arxiv.org/abs/0907.2435
  • R. B. Myerson, Optimal Auction Design, Mathematics of Operations Research 6(1):58–73, 1981. https://doi.org/10.1287/moor.6.1.58
  • J. Bulow, J. Roberts, The Simple Economics of Optimal Auctions, Journal of Political Economy 97(5):1060–1090, 1989. https://doi.org/10.1086/261643
  • R. Kleinberg, S. M. Weinberg, Matroid Prophet Inequalities, STOC 2012. https://arxiv.org/abs/1201.4764
11 thms2 active usersReviewed
Operations ResearchOptimizationStochastic Systems·Captain: mikedeng1

Dimensioning Large Call Centers I: The Rationalized Staffing Function Is Asymptotically OptimalResearch Paper

Motivation

A call center with NNN agents facing Poisson arrivals at rate λ\lambdaλ and exponential service at rate μ\muμ is the M/M/N (Erlang-C) queue. Choosing NNN trades the cost of agents against the cost of customers waiting, and in practice it is done with the square-root safety-staffing rule N≈R+yRN \approx R + y\sqrt RN≈R+yR​, where R=λ/μR = \lambda/\muR=λ/μ is the offered load. Borst, Mandelbaum and Reiman (CWI Report PNA-R0015, 2000; published in Operations Research 52(1), 2004, doi:10.1287/opre.1030.0081) turned that rule of thumb into an optimization result: for a general convex staffing cost and a general waiting-cost function, they identify the safety factor yyy that makes the rule asymptotically optimal as the arrival rate grows.

Timeline of the asymptotic regime the paper builds on:

  • 1917. Erlang's delay formula π(N,ν)\pi(N,\nu)π(N,ν) for the M/M/N queue.
  • 1981. Halfin and Whitt (Oper. Res. 29(3)) show that with N=R+βRN = R + \beta\sqrt RN=R+βR​ servers the probability of waiting converges to a limit P(β)∈(0,1)P(\beta) \in (0,1)P(β)∈(0,1), the quality-and-efficiency-driven regime.
  • 2000/2004. Borst, Mandelbaum and Reiman classify cost structures into a rationalized, an efficiency-driven and a quality-driven regime, and prove asymptotic optimality of an explicit staffing rule in each.

This mission is the first of a series of four on that paper and covers the rationalized regime (Section 5), where staffing and waiting costs are of the same order.

Setting

The service rate μ>0\mu > 0μ>0 is fixed and the arrival rate λ\lambdaλ grows. A staffing cost FFF, defined on (0,∞)(0,\infty)(0,∞), is convex and strictly increasing; it does not depend on λ\lambdaλ. For each λ>0\lambda > 0λ>0 a waiting-cost function DλD_\lambdaDλ​ satisfies Dλ(0)=0D_\lambda(0)=0Dλ​(0)=0, is strictly increasing on [0,∞)[0,\infty)[0,∞), and makes

G(N,λ)=(Nμ−λ)∫0∞Dλ(t) e−(Nμ−λ)t dtG(N,\lambda) = (N\mu-\lambda)\int_0^\infty D_\lambda(t)\,e^{-(N\mu-\lambda)t}\,dtG(N,λ)=(Nμ−λ)∫0∞​Dλ​(t)e−(Nμ−λ)tdt

finite for every N>λ/μN > \lambda/\muN>λ/μ. With the Erlang-C formula

π(N,ν)=νNN!{(1−νN)∑n=0N−1νnn!+νNN!}−1,\pi(N,\nu) = \frac{\nu^N}{N!}\Big\{\big(1-\tfrac{\nu}{N}\big)\sum_{n=0}^{N-1}\frac{\nu^n}{n!}+\frac{\nu^N}{N!}\Big\}^{-1},π(N,ν)=N!νN​{(1−Nν​)n=0∑N−1​n!νn​+N!νN​}−1,

the expected total cost of staffing N>λ/μN > \lambda/\muN>λ/μ agents is C(N,λ)=F(N)+λ π(N,λ/μ) G(N,λ)C(N,\lambda) = F(N) + \lambda\,\pi(N,\lambda/\mu)\,G(N,\lambda)C(N,λ)=F(N)+λπ(N,λ/μ)G(N,λ), and Nλ∗N^*_\lambdaNλ∗​ is any integer N>λ/μN > \lambda/\muN>λ/μ minimizing it (7).

In normalized units Nλ(x)=λ/μ+xλ/μN_\lambda(x) = \lambda/\mu + x\sqrt{\lambda/\mu}Nλ​(x)=λ/μ+xλ/μ​ the paper defines Fλ(x)=F(Nλ(x))−F(λ/μ)F_\lambda(x) = F(N_\lambda(x)) - F(\lambda/\mu)Fλ​(x)=F(Nλ​(x))−F(λ/μ), Gλ(x)=λG(Nλ(x),λ)G_\lambda(x) = \lambda G(N_\lambda(x),\lambda)Gλ​(x)=λG(Nλ​(x),λ), the continuous delay probability πλ(x)=H(Nλ(x),λ/μ)\pi_\lambda(x) = H(N_\lambda(x),\lambda/\mu)πλ​(x)=H(Nλ​(x),λ/μ) with

H(M,α)={α∫0∞e−αt t (1+t)M−1 dt}−1,H(M,\alpha) = \Big\{\alpha\int_0^\infty e^{-\alpha t}\,t\,(1+t)^{M-1}\,dt\Big\}^{-1},H(M,α)={α∫0∞​e−αtt(1+t)M−1dt}−1,

and Cλ(x)=Fλ(x)+πλ(x)Gλ(x)C_\lambda(x) = F_\lambda(x) + \pi_\lambda(x)G_\lambda(x)Cλ​(x)=Fλ​(x)+πλ​(x)Gλ​(x), minimized at xλ∗x^*_\lambdaxλ∗​ (8). A surrogate C[z;F^,π^,G^]=F^(z)+π^(z)G^(z)C[z;\hat F,\hat\pi,\hat G] = \hat F(z)+\hat\pi(z)\hat G(z)C[z;F^,π^,G^]=F^(z)+π^(z)G^(z) approximates it. Rounding is measured by

Sλ(x)=min⁡{C(⌊Nλ(x)⌋,λ), C(⌈Nλ(x)⌉,λ)}.(10)S_\lambda(x) = \min\{C(\lfloor N_\lambda(x)\rfloor,\lambda),\,C(\lceil N_\lambda(x)\rceil,\lambda)\}. \tag{10}Sλ​(x)=min{C(⌊Nλ​(x)⌋,λ),C(⌈Nλ​(x)⌉,λ)}.(10)

The Halfin–Whitt delay function is P(x)=(1+x/h(−x))−1P(x) = \big(1 + x/h(-x)\big)^{-1}P(x)=(1+x/h(−x))−1, with h=ϕ/(1−Φ)h = \phi/(1-\Phi)h=ϕ/(1−Φ) the standard normal hazard rate (11). Asymptotic equality aλ≈∞bλa_\lambda \stackrel{\infty}{\approx} b_\lambdaaλ​≈∞bλ​ means aλ/bλ→1a_\lambda/b_\lambda \to 1aλ​/bλ​→1 as λ→∞\lambda\to\inftyλ→∞.

Formalization targets

Goal: Theorem 5.1

Assume the rationalized condition (18): for some κ>0\kappa > 0κ>0, Fλ(κ)/Gλ(κ)→γ∈(0,∞)F_\lambda(\kappa)/G_\lambda(\kappa) \to \gamma \in (0,\infty)Fλ​(κ)/Gλ​(κ)→γ∈(0,∞). Let yλ∗y^*_\lambdayλ∗​ minimize Fλ(y)+P(y)Gλ(y)F_\lambda(y) + P(y)G_\lambda(y)Fλ​(y)+P(y)Gλ​(y) over y>0y>0y>0 (19). Then

lim⁡λ→∞Sλ(yλ∗)−F(λ/μ)C(Nλ∗,λ)−F(λ/μ)=1.\lim_{\lambda\to\infty}\frac{S_\lambda(y^*_\lambda) - F(\lambda/\mu)}{C(N^*_\lambda,\lambda) - F(\lambda/\mu)} = 1.λ→∞lim​C(Nλ∗​,λ)−F(λ/μ)Sλ​(yλ∗​)−F(λ/μ)​=1.

The goal fixes no constant and no rate: it asserts only that the excess cost of the explicit rule is asymptotically the optimal excess cost.

Milestones

  • Lemma C.1: GλG_\lambdaGλ​ is strictly convex and strictly decreasing on (0,∞)(0,\infty)(0,∞).
  • Section 3, p. 12: H(N,ν)=π(N,ν)H(N,\nu) = \pi(N,\nu)H(N,ν)=π(N,ν) at integers N>ν>0N > \nu > 0N>ν>0.
  • Lemma 3.1, Lemma 3.2, Corollary 3.3: the approximation principle. If the surrogate approximates CλC_\lambdaCλ​ at both xλ∗x^*_\lambdaxλ∗​ and its own minimizer zλ∗z^*_\lambdazλ∗​, then rounding Nλ(zλ∗)N_\lambda(z^*_\lambda)Nλ​(zλ∗​) is asymptotically optimal.
  • Eqs. (13)–(14): FλF_\lambdaFλ​ preserves lim sup⁡\limsuplimsup-separation of ratios.
  • Lemma 4.1 (Halfin & Whitt): for bounded xλx_\lambdaxλ​, πλ(xλ)/P(xλ)→1\pi_\lambda(x_\lambda)/P(x_\lambda) \to 1πλ​(xλ​)/P(xλ​)→1.

Significance

The theorem justifies the square-root staffing rule from first principles for a broad cost class. In Example 5.3 of the paper (linear staffing cost ccc per agent, linear waiting cost aaa per unit time) it gives N∗≈R+y∗(a/c)RN^* \approx R + y^*(a/c)\sqrt RN∗≈R+y∗(a/c)R​, with y∗(r)y^*(r)y∗(r) the minimizer of y+rP(y)/yy + rP(y)/yy+rP(y)/y, a one-dimensional rule computable once for all loads. Corollary 3.3 is reused verbatim by the efficiency-driven and quality-driven theorems of the paper (missions II and III of this series), and Lemma 4.1 is the analytic input of all three.

The result has been proved since 2000; no machine-checked proof of it, or of the Halfin–Whitt limit for the continuous extension πλ\pi_\lambdaπλ​, is known to exist. The mission produces a formal proof of the regime theorem together with reusable formal statements of the Erlang-C function, its integral representation, and the Halfin–Whitt limit.

Difficulty

The reduction from discrete to continuous staffing (Lemmas 3.1–3.2) is elementary once unimodality of CλC_\lambdaCλ​ is available, but unimodality rests on convexity of πλ\pi_\lambdaπλ​, which the paper cites rather than proves, and on Lemma C.1, which needs differentiation under an improper integral. The central difficulty is Lemma 4.1: the paper derives it from Halfin and Whitt's limit theorem, which is stated for integer server counts, while πλ\pi_\lambdaπλ​ is evaluated at non-integer Nλ(xλ)N_\lambda(x_\lambda)Nλ​(xλ​); a proof needs a uniform Laplace-type asymptotic for the integral defining HHH. A further obstacle is bounding xλ∗x^*_\lambdaxλ∗​: the obvious route through continuity of the optimizer fails because nothing converges, and the paper instead argues by contradiction via (14).

Formalization scope

All objects live in DimCallCenters.Rationalized. The arrival rate is a real lam, and every limit is Filter.atTop on R\mathbb RR with μ\muμ fixed. The queue itself is not modelled; the paper's theorems are statements about the closed-form cost C(N,λ)C(N,\lambda)C(N,λ), and so are these. Committed conventions:

  1. The standing assumptions are a structure WaitModel (μ>0\mu>0μ>0; Dλ(0)=0D_\lambda(0)=0Dλ​(0)=0; DλD_\lambdaDλ​ strictly increasing on [0,∞)[0,\infty)[0,∞); t↦Dλ(t)e−θtt\mapsto D_\lambda(t)e^{-\theta t}t↦Dλ​(t)e−θt integrable on (0,∞)(0,\infty)(0,∞) for every θ>0\theta>0θ>0, which is the paper's finiteness of GGG). FFF is convex and strictly increasing on (0,∞)(0,\infty)(0,∞).
  2. Staffing levels in C(N,λ)C(N,\lambda)C(N,λ) are natural numbers; GGG and HHH take real NNN.
  3. Argmins (Nλ∗N^*_\lambdaNλ∗​, xλ∗x^*_\lambdaxλ∗​, zλ∗z^*_\lambdazλ∗​, yλ∗y^*_\lambdayλ∗​) are hypotheses that a given function is a minimizer, for every λ>0\lambda>0λ>0; ties are allowed and the theorems hold for every choice.
  4. In SλS_\lambdaSλ​ the floor term is omitted when ⌊Nλ(x)⌋≤λ/μ\lfloor N_\lambda(x)\rfloor \le \lambda/\mu⌊Nλ​(x)⌋≤λ/μ, where CCC is undefined.
  5. lim sup⁡\limsuplimsup and lim inf⁡\liminfliminf relations are written with ∃ᶠ/∀ᶠ, not Filter.limsup on R\mathbb RR.
  6. Added hypothesis. The goal assumes G(N,λ)→∞G(N,\lambda)\to\inftyG(N,λ)→∞ as N↓λ/μN\downarrow\lambda/\muN↓λ/μ. The paper asserts this limit on p. 12, but it does not follow from its assumptions (it fails for bounded DλD_\lambdaDλ​); it is equivalent to DλD_\lambdaDλ​ being unbounded and is what makes the continuous optimum exist.

The hypotheses are met by linear staffing and waiting costs (F(N)=cNF(N)=cNF(N)=cN, Dλ(t)=atD_\lambda(t)=atDλ​(t)=at), for which (18) holds with γ=cκ2/a\gamma = c\kappa^2/aγ=cκ2/a, so the goal is not vacuous. It is not trivialized by junk values either: the ratio's denominator is positive at every λ>0\lambda>0λ>0, and SλS_\lambdaSλ​ never evaluates CCC at an unstable level.

Needed infrastructure: Laplace asymptotics for ∫0∞e−αtt(1+t)M−1dt\int_0^\infty e^{-\alpha t}t(1+t)^{M-1}dt∫0∞​e−αtt(1+t)M−1dt, differentiation under the integral sign for GGG, and convexity of πλ\pi_\lambdaπλ​. All of these are reusable for missions II–IV. Proofs of the milestones in any order are welcome, as are proofs of the convexity facts the paper cites from its references [9], [10].

Selected references

  • S. Borst, A. Mandelbaum, M. I. Reiman, Dimensioning Large Call Centers, CWI Report PNA-R0015, 2000; Operations Research 52(1):17–34, 2004. https://doi.org/10.1287/opre.1030.0081
  • S. Halfin, W. Whitt, Heavy-Traffic Limits for Queues with Many Exponential Servers, Operations Research 29(3):567–588, 1981. https://doi.org/10.1287/opre.29.3.567
  • A. K. Erlang, Solution of some problems in the theory of probabilities of significance in automatic telephone exchanges, Elektroteknikeren 13, 1917.
21 thms2 active usersReviewed
Machine LearningStatistics·Captain: mikedeng1

Certified Adversarial Robustness via Randomized Smoothing 1: The Gaussian-Smoothed Classifier Is Constant on the ℓ2 Ball of Radius (σ/2)(Φ⁻¹(p_A) − Φ⁻¹(p_B))Research Paper

Motivation

Classifiers trained on images, speech and text can be made to change their prediction by perturbations of the input that are tiny in norm (Szegedy et al., 2014). Empirical defences against such adversarial examples have repeatedly been broken by stronger attacks (Athalye, Carlini, Wagner, 2018), which motivates certified defences: classifiers that come with a proof that their prediction at a given input cannot change inside a stated ball around it.

Randomized smoothing turns any classifier, however large or opaque, into one with such a certificate in the ℓ2\ell_2ℓ2​ norm. It was introduced with weaker radii by Lecuyer et al. (2019) and Li et al. (2018). Cohen, Rosenfeld and Kolter (arXiv:1902.02918v2, ICML 2019) proved the radius that is now standard, and showed it cannot be enlarged. Their guarantee underlies most later work on certified ℓ2\ell_2ℓ2​ robustness, including Salman et al. (2019).

This mission formalizes the robustness guarantee, Theorem 1 of that paper. Page numbers below are PDF pages of the arXiv v2 preprint, which has no printed page numbers.

Setting

Inputs are points of Rd\mathbb R^dRd with the Euclidean norm ∥⋅∥\|\cdot\|∥⋅∥ and inner product δ⊤z\delta^\top zδ⊤z. Classes form a set Y\mathcal YY. A base classifier is a deterministic or random function f:Rd→Yf : \mathbb R^d \to \mathcal Yf:Rd→Y. A random fff is described by the probabilities P(f(z)=c)\mathbb P(f(z) = c)P(f(z)=c), which for each zzz form a probability distribution on Y\mathcal YY.

Fix a noise level σ>0\sigma > 0σ>0 and let ε∼N(0,σ2I)\varepsilon \sim \mathcal N(0, \sigma^2 I)ε∼N(0,σ2I) be isotropic Gaussian noise. The class probabilities at xxx are P(f(x+ε)=c)\mathbb P(f(x + \varepsilon) = c)P(f(x+ε)=c), and the smoothed classifier is

g(x)=arg⁡max⁡c∈YP(f(x+ε)=c).(1)g(x) = \arg\max_{c \in \mathcal Y} \mathbb P\big(f(x + \varepsilon) = c\big). \qquad (1)g(x)=argc∈Ymax​P(f(x+ε)=c).(1)

The paper leaves g(x)g(x)g(x) undefined when the maximizer is not unique. "g(x)=cg(x) = cg(x)=c" therefore means that every class other than ccc has strictly smaller probability.

Write Φ\PhiΦ for the standard Gaussian cumulative distribution function and Φ−1\Phi^{-1}Φ−1 for its inverse. Φ−1\Phi^{-1}Φ−1 is a real number on (0,1)(0,1)(0,1), and Φ−1(0)=−∞\Phi^{-1}(0) = -\inftyΦ−1(0)=−∞, Φ−1(1)=+∞\Phi^{-1}(1) = +\inftyΦ−1(1)=+∞.

Formalization targets

Goal: Theorem 1 (p. 4; restated p. 13)

Suppose that at a specific xxx there are a class cAc_AcA​ and numbers pA‾,pB‾∈[0,1]\underline{p_A}, \overline{p_B} \in [0,1]pA​​,pB​​∈[0,1] with

P(f(x+ε)=cA) ≥ pA‾ ≥ pB‾ ≥ max⁡c≠cAP(f(x+ε)=c).(6)\mathbb P\big(f(x + \varepsilon) = c_A\big) \ \ge\ \underline{p_A} \ \ge\ \overline{p_B} \ \ge\ \max_{c \ne c_A} \mathbb P\big(f(x + \varepsilon) = c\big). \qquad (6)P(f(x+ε)=cA​) ≥ pA​​ ≥ pB​​ ≥ c=cA​max​P(f(x+ε)=c).(6)

Then g(x+δ)=cAg(x + \delta) = c_Ag(x+δ)=cA​ for every δ\deltaδ with ∥δ∥2<R\|\delta\|_2 < R∥δ∥2​<R, where

R=σ2(Φ−1(pA‾)−Φ−1(pB‾)).(7)R = \frac{\sigma}{2}\Big(\Phi^{-1}(\underline{p_A}) - \Phi^{-1}(\overline{p_B})\Big). \qquad (7)R=2σ​(Φ−1(pA​​)−Φ−1(pB​​)).(7)

The statement covers every base classifier and every set of classes. The radius is infinite when pA‾=1>pB‾\underline{p_A} = 1 > \overline{p_B}pA​​=1>pB​​ or pA‾>0=pB‾\underline{p_A} > 0 = \overline{p_B}pA​​>0=pB​​.

Milestones

The milestones are the paper's own steps, in attack order:

  • Lemma 3 (p. 12): the Neyman–Pearson lemma in both directions, for densities μX\mu_XμX​, μY\mu_YμY​ on Rd\mathbb R^dRd and the likelihood-ratio sets {μY≤tμX}\{\mu_Y \le t\mu_X\}{μY​≤tμX​} and {μY≥tμX}\{\mu_Y \ge t\mu_X\}{μY​≥tμX​}.
  • The likelihood ratio and (5) (proof of Lemma 4, p. 13): for X∼N(x,σ2I)X \sim \mathcal N(x,\sigma^2 I)X∼N(x,σ2I) and Y∼N(x+δ,σ2I)Y \sim \mathcal N(x+\delta,\sigma^2 I)Y∼N(x+δ,σ2I), μY/μX=exp⁡(aδ⊤z+b)\mu_Y/\mu_X = \exp(a\delta^\top z + b)μY​/μX​=exp(aδ⊤z+b), so half-spaces orthogonal to δ\deltaδ are likelihood-ratio sets.
  • Lemma 4 (pp. 12–13): Neyman–Pearson for these two Gaussians and the half-spaces {δ⊤z≤β}\{\delta^\top z \le \beta\}{δ⊤z≤β}, {δ⊤z≥β}\{\delta^\top z \ge \beta\}{δ⊤z≥β}.
  • The four Claims of Appendix A.0.1 (pp. 15–16, with (13) and (14) of p. 14): the probabilities of the half-spaces A={z:δ⊤(z−x)≤σ∥δ∥Φ−1(pA‾)}A = \{z : \delta^\top(z-x) \le \sigma\|\delta\|\Phi^{-1}(\underline{p_A})\}A={z:δ⊤(z−x)≤σ∥δ∥Φ−1(pA​​)} and B={z:δ⊤(z−x)≥σ∥δ∥Φ−1(1−pB‾)}B = \{z : \delta^\top(z-x) \ge \sigma\|\delta\|\Phi^{-1}(1-\overline{p_B})\}B={z:δ⊤(z−x)≥σ∥δ∥Φ−1(1−pB​​)} under XXX and YYY:
P(X∈A)=pA‾,P(X∈B)=pB‾,P(Y∈A)=Φ(Φ−1(pA‾)−∥δ∥σ),P(Y∈B)=Φ(Φ−1(pB‾)+∥δ∥σ).\mathbb P(X \in A) = \underline{p_A},\quad \mathbb P(X \in B) = \overline{p_B},\quad \mathbb P(Y \in A) = \Phi\Big(\Phi^{-1}(\underline{p_A}) - \tfrac{\|\delta\|}{\sigma}\Big),\quad \mathbb P(Y \in B) = \Phi\Big(\Phi^{-1}(\overline{p_B}) + \tfrac{\|\delta\|}{\sigma}\Big).P(X∈A)=pA​​,P(X∈B)=pB​​,P(Y∈A)=Φ(Φ−1(pA​​)−σ∥δ∥​),P(Y∈B)=Φ(Φ−1(pB​​)+σ∥δ∥​).
  • (15) (p. 14): P(Y∈A)>P(Y∈B)\mathbb P(Y \in A) > \mathbb P(Y \in B)P(Y∈A)>P(Y∈B) if and only if ∥δ∥<R\|\delta\| < R∥δ∥<R.

Significance

Theorem 1 turns three numbers at one input into a guarantee over a whole ball: a lower bound on the top-class probability, an upper bound on the other classes, and the noise level. These bounds can be estimated by sampling and certified with confidence intervals (the paper's CERTIFY procedure). That is what lets randomized smoothing certify ImageNet-scale networks, where exact verification methods do not scale. The companion result (Theorem 2, a separate mission of this series) shows that no larger ℓ2\ell_2ℓ2​ ball can be certified from the same information.

The theorem has a short pen-and-paper proof; no machine-checked proof of it is recorded on Prove2Me. A formal development adds:

  • a checked Neyman–Pearson lemma for randomized tests with densities on Rd\mathbb R^dRd, which Mathlib does not have;
  • the Gaussian likelihood-ratio and half-space computations;
  • an explicit treatment of the endpoint cases pA‾=1\underline{p_A} = 1pA​​=1 and pB‾=0\overline{p_B} = 0pB​​=0, where the radius is infinite.

Difficulty

Every step is classical, so the difficulty lies in the missing infrastructure, not in the idea.

  • Densities. Mathlib's multivariate Gaussian is defined as a pushforward of a product measure, not by a density. Identifying it with the density (2πσ2)−d/2e−∥z−x∥2/(2σ2)(2\pi\sigma^2)^{-d/2}e^{-\|z-x\|^2/(2\sigma^2)}(2πσ2)−d/2e−∥z−x∥2/(2σ2), which Lemma 4 needs, is not available.
  • Projections. The Claims need the law of δ⊤X\delta^\top Xδ⊤X for X∼N(x,σ2I)X \sim \mathcal N(x,\sigma^2 I)X∼N(x,σ2I) in closed form, namely a one-dimensional Gaussian with mean δ⊤x\delta^\top xδ⊤x and variance σ2∥δ∥2\sigma^2\|\delta\|^2σ2∥δ∥2.
  • The inverse CDF. Mathlib has no normal quantile, so Φ−1\Phi^{-1}Φ−1 is defined here as an infimum. The identities Φ(Φ−1(p))=p\Phi(\Phi^{-1}(p)) = pΦ(Φ−1(p))=p and Φ−1(1−p)=−Φ−1(p)\Phi^{-1}(1-p) = -\Phi^{-1}(p)Φ−1(1−p)=−Φ−1(p) must be derived.
  • Degenerate cases. The obvious argument through the worst-case half-spaces breaks down when δ=0\delta = 0δ=0 or when pA‾\underline{p_A}pA​​ or pB‾\overline{p_B}pB​​ is 000 or 111. There the half-spaces are empty or everything and Φ−1\Phi^{-1}Φ−1 is infinite, so these cases need a separate argument.

Formalization scope

  • Space and noise. Rd\mathbb R^dRd is EuclideanSpace ℝ (Fin d). N(x,σ2I)\mathcal N(x,\sigma^2 I)N(x,σ2I) is the pushforward of Mathlib's stdGaussian under z↦x+σzz \mapsto x + \sigma zz↦x+σz, with σ>0\sigma > 0σ>0 a hypothesis.
  • Classifiers. A random classifier is a map f : ℝᵈ → PMF 𝒴 with measurable class probabilities; deterministic classifiers are point masses. Y\mathcal YY is an arbitrary type, with no finiteness assumed. The class probability is the published Gaussian smoothing RandomGradFree.Shared.smoothing, applied to z↦P(f(z)=c)z \mapsto \mathbb P(f(z) = c)z↦P(f(z)=c).
  • The prediction. "g(x)=cg(x) = cg(x)=c" is the strict unique-maximizer predicate, and ggg itself is not defined. Defining ggg by an arbitrary choice at ties would make the theorem depend on the tie-break.
  • The radius. RRR is computed in the extended reals with Φ−1(0)=−∞\Phi^{-1}(0) = -\inftyΦ−1(0)=−∞ and Φ−1(1)=+∞\Phi^{-1}(1) = +\inftyΦ−1(1)=+∞. A real-valued Φ−1\Phi^{-1}Φ−1 with junk value 000 at the endpoints would assign a finite, wrong radius there, so it is used only in milestones that assume 0<p<10 < p < 10<p<1. In the two corners pA‾=pB‾∈{0,1}\underline{p_A} = \overline{p_B} \in \{0,1\}pA​​=pB​​∈{0,1}, where (7) reads ∞−∞\infty - \infty∞−∞, the radius is −∞-\infty−∞ and the goal is vacuous, as in the paper.
  • Neyman–Pearson. Random tests are [0,1][0,1][0,1]-valued measurable functions, so the lemmas apply with h(z)=P(f(z)=c)h(z) = \mathbb P(f(z) = c)h(z)=P(f(z)=c) for random fff. The likelihood-ratio sets are written multiplied out (μY≤tμX\mu_Y \le t\mu_XμY​≤tμX​), which avoids division by zero where μX\mu_XμX​ vanishes.
  • Added hypotheses. The milestones about AAA and BBB assume δ≠0\delta \ne 0δ=0 and 0<p<10 < p < 10<p<1, which the paper's computation uses implicitly.

Contributions are welcome at every level: proofs of the milestones, and general lemmas such as the Gaussian density, the law of linear functionals of a Gaussian vector and properties of the normal quantile. These lemmas are reusable well beyond this mission.

Selected references

  • J. M. Cohen, E. Rosenfeld, J. Z. Kolter, Certified Adversarial Robustness via Randomized Smoothing, ICML 2019; arXiv:1902.02918v2. https://arxiv.org/abs/1902.02918
  • J. Neyman, E. S. Pearson, On the Problem of the Most Efficient Tests of Statistical Hypotheses, Phil. Trans. R. Soc. A 231, 1933. https://doi.org/10.1098/rsta.1933.0009
  • M. Lecuyer, V. Atlidakis, R. Geambasu, D. Hsu, S. Jana, Certified Robustness to Adversarial Examples with Differential Privacy, IEEE S&P 2019. https://arxiv.org/abs/1802.03471
  • B. Li, C. Chen, W. Wang, L. Carin, Certified Adversarial Robustness with Additive Noise, NeurIPS 2019. https://arxiv.org/abs/1809.03113
  • H. Salman et al., Provably Robust Deep Learning via Adversarially Trained Smoothed Classifiers, NeurIPS 2019. https://arxiv.org/abs/1906.04584
  • A. Athalye, N. Carlini, D. Wagner, Obfuscated Gradients Give a False Sense of Security, ICML 2018. https://arxiv.org/abs/1802.00420
  • C. Szegedy et al., Intriguing Properties of Neural Networks, ICLR 2014. https://arxiv.org/abs/1312.6199
15 thms2 active usersReviewed
Convex OptimizationOptimization·Captain: mikedeng1

The Entropic Barrier: A Simple and Optimal Universal Self-Concordant Barrier: The Entropic Barrier of a Convex Body in ℝⁿ Is a (1 + εₙ)n-Self-Concordant Barrier with εₙ ≤ 100√(log n / n)Research Paper

Motivation

Interior-point methods minimize a linear function x↦⟨c,x⟩x\mapsto\langle c,x\ranglex↦⟨c,x⟩ over a convex set K⊂Rn\mathcal K\subset\mathbb R^nK⊂Rn by following the minimizers of ⟨c,x⟩+1tg(x)\langle c,x\rangle+\frac1t g(x)⟨c,x⟩+t1​g(x) as t→∞t\to\inftyt→∞, where ggg is a self-concordant barrier for K\mathcal KK. Each Newton step of such a method shrinks 1/t1/t1/t by a factor 1−1/ν1-1/\sqrt\nu1−1/ν​, where ν\nuν is the self-concordance parameter of ggg, so ν\nuν controls the iteration count of every interior-point method built on ggg (Nesterov and Nemirovski 1994; Nesterov 2004).

Timeline:

  • 1994. Nesterov and Nemirovski construct the universal barrier for any convex body and show it is a ν\nuν-self-concordant barrier with ν≤Cn\nu\le Cnν≤Cn for a universal constant CCC. They also show that ν≥n\nu\ge nν≥n is necessary for some bodies (the simplex, the cube).
  • 2014–2015. Hildebrand (Math. Oper. Res. 2014) and Fox (Ann. Mat. Pura Appl. 2015) show that the canonical barrier of a convex cone has parameter equal to the dimension, which gives parameter n+1n+1n+1 for convex bodies.
  • 2015. Bubeck and Eldan (arXiv:1412.1587, COLT 2015) show that the Fenchel dual of the log-Laplace transform of the uniform measure on K\mathcal KK, which they call the entropic barrier, is a (1+o(1))n(1+o(1))n(1+o(1))n-self-concordant barrier, with an explicit o(1)o(1)o(1) term.

Beyond optimization, the entropic barrier is the mirror map that pairs naturally with the exponential-family sampling scheme in bandit linear optimization, which the paper discusses in its §3.1.

Setting

Let K⊂Rn\mathcal K\subset\mathbb R^nK⊂Rn be a convex body: compact, convex, with non-empty interior int⁡(K)\operatorname{int}(\mathcal K)int(K). The log-Laplace transform of K\mathcal KK is

f(θ)=log⁡(∫x∈Kexp⁡(⟨θ,x⟩) dx),θ∈Rn,f(\theta)=\log\left(\int_{x\in\mathcal K}\exp(\langle\theta,x\rangle)\,dx\right),\qquad\theta\in\mathbb R^n,f(θ)=log(∫x∈K​exp(⟨θ,x⟩)dx),θ∈Rn,

and the entropic barrier is its Fenchel dual

f∗(x)=sup⁡θ∈Rn ⟨θ,x⟩−f(θ),x∈int⁡(K).f^*(x)=\sup_{\theta\in\mathbb R^n}\ \langle\theta,x\rangle-f(\theta),\qquad x\in\operatorname{int}(\mathcal K).f∗(x)=θ∈Rnsup​ ⟨θ,x⟩−f(θ),x∈int(K).

For a function g:int⁡(K)→Rg:\operatorname{int}(\mathcal K)\to\mathbb Rg:int(K)→R write ∇g(x)[h]\nabla g(x)[h]∇g(x)[h], ∇2g(x)[h,h]\nabla^2g(x)[h,h]∇2g(x)[h,h], ∇3g(x)[h,h,h]\nabla^3g(x)[h,h,h]∇3g(x)[h,h,h] for its directional derivatives. Following Definition 1 of the paper:

  1. ggg is a barrier for K\mathcal KK if g(x)→+∞g(x)\to+\inftyg(x)→+∞ as x→∂Kx\to\partial\mathcal Kx→∂K;
  2. a C3C^3C3 convex ggg is self-concordant if ∇3g(x)[h,h,h]≤2(∇2g(x)[h,h])3/2\nabla^3g(x)[h,h,h]\le2(\nabla^2g(x)[h,h])^{3/2}∇3g(x)[h,h,h]≤2(∇2g(x)[h,h])3/2 for all x∈int⁡(K)x\in\operatorname{int}(\mathcal K)x∈int(K), h∈Rnh\in\mathbb R^nh∈Rn;
  3. it is ν\nuν-self-concordant if moreover ∇g(x)[h]≤ν⋅∇2g(x)[h,h]\nabla g(x)[h]\le\sqrt{\nu\cdot\nabla^2g(x)[h,h]}∇g(x)[h]≤ν⋅∇2g(x)[h,h]​ for all such x,hx,hx,h.

The proof works with the canonical exponential family pθp_\thetapθ​, the probability measure with density exp⁡(⟨θ,x⟩−f(θ))1{x∈K}\exp(\langle\theta,x\rangle-f(\theta))\mathbb 1\{x\in\mathcal K\}exp(⟨θ,x⟩−f(θ))1{x∈K}, its mean x(θ)x(\theta)x(θ), covariance Σ(θ)\Sigma(\theta)Σ(θ) and third central moment T(θ)T(\theta)T(θ); with Y=⟨θ/∥θ∥,X⟩Y=\langle\theta/\|\theta\|,X\rangleY=⟨θ/∥θ∥,X⟩ for X∼pθX\sim p_\thetaX∼pθ​ and its density ρ\rhoρ; and with the section marginal λ(y)=Voln−1(K∩{yθ/∥θ∥+θ⊥})/Vol(K)\lambda(y)=\mathrm{Vol}_{n-1}(\mathcal K\cap\{y\theta/\|\theta\|+\theta^\perp\})/\mathrm{Vol}(\mathcal K)λ(y)=Voln−1​(K∩{yθ/∥θ∥+θ⊥})/Vol(K).

Formalization targets

Goal: Theorem 1

For every n≥80n\ge80n≥80 and every convex body K⊂Rn\mathcal K\subset\mathbb R^nK⊂Rn, f∗f^*f∗ is a ν\nuν-self-concordant barrier for K\mathcal KK with

ν=(1+εn) n,εn=100log⁡nn.\nu=(1+\varepsilon_n)\,n,\qquad\varepsilon_n=100\sqrt{\frac{\log n}{n}}.ν=(1+εn​)n,εn​=100nlogn​​.

Milestones, in attack order

  1. Lemma 1 (p. 5): strict convexity of fff, f∗f^*f∗; ∇f∗:int⁡(K)→Rn\nabla f^*:\operatorname{int}(\mathcal K)\to\mathbb R^n∇f∗:int(K)→Rn is a bijection; ∇2f=Σ\nabla^2f=\Sigma∇2f=Σ, ∇3f=T\nabla^3f=T∇3f=T (eqs. (4)–(5)); ∇2f∗(x)=Σ(θ(x))−1\nabla^2f^*(x)=\Sigma(\theta(x))^{-1}∇2f∗(x)=Σ(θ(x))−1 (eq. (6)).
  2. f∗f^*f∗ is a barrier (§4, p. 6).
  3. Lemma 2 (p. 7): EX3≤2(EX2)3/2\mathbb EX^3\le2(\mathbb EX^2)^{3/2}EX3≤2(EX2)3/2 for a real centered log-concave XXX; its consequence Epθ⟨X−x(θ),h⟩3≤2(Epθ⟨X−x(θ),h⟩2)3/2\mathbb E_{p_\theta}\langle X-x(\theta),h\rangle^3\le2(\mathbb E_{p_\theta}\langle X-x(\theta),h\rangle^2)^{3/2}Epθ​​⟨X−x(θ),h⟩3≤2(Epθ​​⟨X−x(θ),h⟩2)3/2; f∗f^*f∗ is self-concordant (§4, pp. 6–7).
  4. Reduction of (3) (p. 7): f∗f^*f∗ satisfies (3) with parameter ν\nuν iff ⟨Σ(θ)θ,θ⟩≤ν\langle\Sigma(\theta)\theta,\theta\rangle\le\nu⟨Σ(θ)θ,θ⟩≤ν for all θ\thetaθ.
  5. λ\lambdaλ is nnn-concave on its support (p. 9) and Lemma 5 (p. 9): φ\varphiφ is nnn-concave iff (log⁡φ)′′≤−1n((log⁡φ)′)2(\log\varphi)''\le-\frac1n((\log\varphi)')^2(logφ)′′≤−n1​((logφ)′)2.
  6. Lemma 3 (p. 8): ρ(y+y0)=ρ(y0)ζ(y)e−y2/(2σ2)\rho(y+y_0)=\rho(y_0)\zeta(y)e^{-y^2/(2\sigma^2)}ρ(y+y0​)=ρ(y0​)ζ(y)e−y2/(2σ2) on [−M,M][-M,M][−M,M], with ζ∈[0,1]\zeta\in[0,1]ζ∈[0,1] unimodal, M=7nlog⁡n/∥θ∥M=\sqrt{7n\log n}/\|\theta\|M=7nlogn​/∥θ∥, σ2=n∥θ∥211−7log⁡(n)/n\sigma^2=\frac{n}{\|\theta\|^2}\frac{1}{1-\sqrt{7\log(n)/n}}σ2=∥θ∥2n​1−7log(n)/n​1​; and its consequence (9): E(∣Y−y0∣2∣∣Y−y0∣≤M)≤σ2\mathbb E(|Y-y_0|^2\mid|Y-y_0|\le M)\le\sigma^2E(∣Y−y0​∣2∣∣Y−y0​∣≤M)≤σ2.
  7. Lemma 4 (p. 8): (1−2c(ε)εlog⁡2(1/ε))Var(X)≤∫x1x2(x−x0)2λ(x)dx≤E(∣X−x0∣2∣X∈[x1,x2])(1-2c(\varepsilon)\varepsilon\log^2(1/\varepsilon))\mathrm{Var}(X)\le\int_{x_1}^{x_2}(x-x_0)^2\lambda(x)dx\le\mathbb E(|X-x_0|^2\mid X\in[x_1,x_2])(1−2c(ε)εlog2(1/ε))Var(X)≤∫x1​x2​​(x−x0​)2λ(x)dx≤E(∣X−x0​∣2∣X∈[x1​,x2​]) for log-concave XXX.
  8. (7) (p. 7): Var(Y)≤n∥θ∥2(1+εn)\mathrm{Var}(Y)\le\frac{n}{\|\theta\|^2}(1+\varepsilon_n)Var(Y)≤∥θ∥2n​(1+εn​).

Significance

The result. Theorem 1 gives, for every convex body, an explicit barrier whose parameter is nnn up to a second-order term, against the CnCnCn of the universal barrier, and it is optimal up to that term because ν≥n\nu\ge nν≥n is necessary for some bodies. The barrier is defined by a single formula, its derivatives are moments of an explicit probability measure, and its parameter bound reduces to a variance bound for one-dimensional log-concave marginals. Lemmas 2 and 4 are self-contained facts about log-concave laws on R\mathbb RR (a sharp third-moment bound and a variance-localization bound) that are usable outside this paper.

Formalizing it. The theorem is proved in the paper; nothing here is formalized elsewhere. The platform has a definition of self-concordance (reused here) and results for given self-concordant functions, but no universal or entropic barrier, no exponential family over a convex body, and no moment bounds for log-concave laws. A complete development produces machine-checked versions of the duality facts of Lemma 1, of the two log-concave lemmas, and of the Brunn–Minkowski consequence for section volumes. Two steps of the paper are sketched rather than proved in full: the end of the proof of Lemma 2 ("We omit further details of this proof", p. 12) and, in Lemma 4, a normalization step that cites a lemma stated for isotropic densities. A formal proof either fills or replaces them.

Difficulty

Self-concordance of f∗f^*f∗ reduces to self-concordance of fff by a general duality fact, and that reduces to Lemma 2; the difficulty there is the sharp constant 222, since generic moment comparisons for log-concave laws give a worse constant. The parameter bound is the hard part. The obvious bound ⟨Σ(θ)θ,θ⟩≤Cn\langle\Sigma(\theta)\theta,\theta\rangle\le Cn⟨Σ(θ)θ,θ⟩≤Cn follows from standard concentration for log-concave measures, but any argument that loses a constant factor proves only the 1994 result. The 1+o(1)1+o(1)1+o(1) requires the one-dimensional marginal of the tilted measure to be compared with a Gaussian of variance n/∥θ∥2n/\|\theta\|^2n/∥θ∥2 to within a factor 1+O(log⁡n/n)1+O(\sqrt{\log n/n})1+O(logn/n​), using the fact that λ\lambdaλ is nnn-concave and not merely log-concave. The paper does this pointwise near the mode (Lemma 3) and controls the tails separately (Lemma 4). The pointwise argument assumes ρ\rhoρ smooth, which holds for smooth bodies, and an approximation argument passes to general convex bodies.

Formalization scope

  • Rn\mathbb R^nRn is EuclideanSpace ℝ (Fin n), so nnn is the dimension, not a separate parameter. A convex body is compact, convex, with non-empty interior; a lower-dimensional set is excluded, which rules out a formalization in which the barrier and self-concordance clauses hold vacuously.

  • f∗f^*f∗ is a real supremum. On int⁡(K)\operatorname{int}(\mathcal K)int(K) it is the true supremum; elsewhere Lean returns a junk value that no statement reads. The barrier property is a limit within int⁡(K)\operatorname{int}(\mathcal K)int(K) at every frontier point.

  • Self-concordance (2) is the published ConvexOptimization.IsSelfConcordantOn on interior K, stated by line restrictions with an absolute value. It is equivalent to (2), because h↦−hh\mapsto-hh↦−h flips the sign of the third derivative.

  • The goal states the parameter as the explicit number ν=(1+100log⁡(n)/n) n\nu=(1+100\sqrt{\log(n)/n})\,nν=(1+100log(n)/n​)n. The page says εn≤100log⁡(n)/n\varepsilon_n\le100\sqrt{\log(n)/n}εn​≤100log(n)/n​, and (3) is monotone in ν\nuν, so this is the same claim. An existential ν\nuν is not used.

  • Corrections and implicit hypotheses:

    • Lemma 4 is stated for 0<ε<10<\varepsilon<10<ε<1. The page says ε>0\varepsilon>0ε>0, but the statement is false for ε≥1\varepsilon\ge1ε≥1 and the paper applies it only with ε<1\varepsilon<1ε<1.
    • Lemma 5 assumes φ>0\varphi>0φ>0, which is implicit in ζ=log⁡φ\zeta=\log\varphiζ=logφ.
    • The reduction of (3) assumes ν≥0\nu\ge0ν≥0.
    • Lemma 3 and (9) carry the smoothness of ρ\rhoρ (the paper's own without-loss-of-generality step on p. 7) as a hypothesis, and the theorem's range n≥80n\ge80n≥80.
  • Section volumes use Mathlib's unnormalized (n−1)(n-1)(n−1)-dimensional Hausdorff measure. The normalization constant cancels in ρ\rhoρ and does not affect nnn-concavity. λ\lambdaλ and ρ\rhoρ are fixed pointwise functions, because Lemma 3 evaluates ρ\rhoρ at a maximizer.

  • Log-concavity on R\mathbb RR is the published ConvexOptimization.LogConcaveOn on the whole line.

  • Needed infrastructure that is reusable beyond this mission:

    • differentiation under the integral sign for exponential families on compact sets;
    • Fenchel duality for smooth strictly convex functions;
    • Brunn's concavity theorem for sections of convex bodies;
    • moment and tail bounds for log-concave densities on R\mathbb RR.

    Contributions to any of these, or proofs of single milestones, are welcome.

Selected references

  • S. Bubeck, R. Eldan, The entropic barrier: a simple and optimal universal self-concordant barrier, COLT 2015; arXiv:1412.1587v3. https://arxiv.org/abs/1412.1587
  • Y. Nesterov, A. Nemirovski, Interior-Point Polynomial Algorithms in Convex Programming, SIAM, 1994. https://doi.org/10.1137/1.9781611970791
  • Y. Nesterov, Introductory Lectures on Convex Optimization: A Basic Course, Kluwer, 2004. https://doi.org/10.1007/978-1-4419-8853-9
  • R. Hildebrand, Canonical barriers on convex cones, Mathematics of Operations Research 39:841–850, 2014.
  • D. Fox, A Schwarz lemma for Kähler affine metrics and the canonical potential of a proper convex cone, Annali di Matematica Pura ed Applicata 194:1–42, 2015.
  • B. Klartag, On convex perturbations with a bounded isotropic constant, Geometric and Functional Analysis 16(6):1274–1290, 2006.
  • C. Borell, Convex set functions in d-space, Periodica Mathematica Hungarica 6(2):111–136, 1975.
21 thms2 active usersReviewed
Operations ResearchTheoretical Computer Science·Captain: mikedeng1

On the Power of Randomization in On-Line Algorithms 1: α-Competitiveness Against Adaptive On-Line and β Against Oblivious Adversaries Give a Deterministic α∘β-Competitive AlgorithmResearch Paper

Why randomization matters in online algorithms

An online algorithm must answer each request before it sees the next one. Its performance is compared with an optimum that may choose all its answers after seeing the complete request string. Randomization can improve an online algorithm's guarantee when the request string is fixed in advance. The comparison changes when an adversary chooses later requests after seeing the algorithm's earlier answers. Ben-David, Borodin, Karp, Tardos and Wigderson studied these choices of adversary in a common request-answer model and proved a general relation between their competitive guarantees (Ben-David et al., 1994, manuscript §§2–3).

The paper distinguishes three adversaries. An oblivious adversary fixes the request string before the algorithm's random choices affect any answer. An adaptive off-line adversary chooses the next request from previous answers but serves the resulting request string optimally after the play. An adaptive on-line adversary also chooses its own answer as each request arrives. The ability to react to answers makes the latter two adversaries materially different from the oblivious one for randomized algorithms (Ben-David et al., 1994, manuscript pp. 7–9).

Request-answer games and competitive cost

A request-answer game has a request set RRR, a finite nonempty answer set AAA, and a real cost fn(r,a)f_n(r,a)fn​(r,a) for a request string r∈Rnr\in R^nr∈Rn and an answer string a∈Ana\in A^na∈An. The off-line optimum for rrr is c(r)=min⁡a∈Anfn(r,a)c(r)=\min_{a\in A^n}f_n(r,a)c(r)=mina∈An​fn​(r,a). A deterministic online algorithm DDD returns its iiith answer from the first iii requests alone; it has no access to the rest of rrr or to the eventual stopping time. Its cost on rrr is cD(r)=fn(r,D(r))c_D(r)=f_n(r,D(r))cD​(r)=fn​(r,D(r)).

A randomized online algorithm is a distribution over deterministic online algorithms. With coins ω\omegaω, write GωG_\omegaGω​ for the resulting deterministic algorithm. For a fixed request string rrr, GGG is β\betaβ-competitive against oblivious adversaries when Eω[cGω(r)]≤β(c(r))\mathbb E_\omega[c_{G_\omega}(r)]\leq\beta(c(r))Eω​[cGω​​(r)]≤β(c(r)). The paper calls a transformation “linear” when it has the affine form x↦ux+vx\mapsto ux+vx↦ux+v (Ben-David et al., 1994, manuscript p. 7).

An adaptive off-line adversary QQQ has a rule from prior answer strings to either the next request or a stop signal, together with a common finite upper bound on play length. Let r(Gω,Q)r(G_\omega,Q)r(Gω​,Q) denote its request string and cQ(Gω)=c(r(Gω,Q))c_Q(G_\omega)=c(r(G_\omega,Q))cQ​(Gω​)=c(r(Gω​,Q)). Its competitiveness condition places the transformation inside the expectation: Eω[cGω(Q)]≤Eω[α(cQ(Gω))]\mathbb E_\omega[c_{G_\omega}(Q)]\leq\mathbb E_\omega[\alpha(c_Q(G_\omega))]Eω​[cGω​​(Q)]≤Eω​[α(cQ​(Gω​))]. An adaptive on-line adversary SSS has the same request rule and an additional answer rule; its own cost is cS(Gω)c_S(G_\omega)cS​(Gω​), and the corresponding condition uses Eω[α(cS(Gω))]\mathbb E_\omega[\alpha(c_S(G_\omega))]Eω​[α(cS​(Gω​))] on the right (Ben-David et al., 1994, manuscript pp. 8–9).

Formalization targets

Randomization against adaptive off-line adversaries

The first target is Theorem 2.1: if some randomized algorithm is α\alphaα-competitive against every adaptive off-line adversary, a deterministic algorithm has that same guarantee on every request string:

∃G  ∀Q,E[cG(Q)]≤E[α(cQ(G))]⟹∃D  ∀r,cD(r)≤α(c(r)).\exists G\;\forall Q,\quad \mathbb E[c_G(Q)]\leq\mathbb E[\alpha(c_Q(G))]\quad\Longrightarrow\quad\exists D\;\forall r,\quad c_D(r)\leq\alpha(c(r)).∃G∀Q,E[cG​(Q)]≤E[α(cQ​(G))]⟹∃D∀r,cD​(r)≤α(c(r)).

Composition of two guarantees

Theorem 2.2 takes an α\alphaα guarantee for GGG against adaptive on-line adversaries and a β\betaβ guarantee for another randomized algorithm against oblivious adversaries. It concludes that GGG has the composed guarantee against adaptive off-line adversaries:

E[cG(Q)]≤E[(α∘β)(cQ(G))]for every Q.\mathbb E[c_G(Q)]\leq\mathbb E[(\alpha\circ\beta)(c_Q(G))]\qquad\text{for every }Q.E[cG​(Q)]≤E[(α∘β)(cQ​(G))]for every Q.

The mission goal is Corollary 2.1, the deterministic consequence of these two results:

∃D  ∀r,cD(r)≤(α∘β)(c(r)).\exists D\;\forall r,\qquad c_D(r)\leq(\alpha\circ\beta)(c(r)).∃D∀r,cD​(r)≤(α∘β)(c(r)).

The milestones follow the paper's two theorems and the stated claims in their proofs, including the finite-horizon winning-position formulation and the adversary that simulates a fixed online algorithm (Ben-David et al., 1994, manuscript pp. 9–13).

What the result supplies

The corollary turns the existence of two randomized guarantees under different information rules into the existence of a deterministic online strategy with an explicit composed cost transformation. It is an existence result: it does not say that the deterministic strategy can be computed efficiently from the randomized algorithms. The paper itself notes that such a construction is unavailable in full generality and then examines settings where constructive versions are possible (Ben-David et al., 1994, manuscript p. 13).

The mathematical results were proved in the 1994 paper; this mission asks for machine-checked Lean proofs of the abstract model, the intermediate claims, and Corollary 2.1. The local draft currently contains compiled statements with proof placeholders, so it does not yet provide checked proofs. A completed development would make the adversary distinctions and the exact placement of expectations available for reuse in later online-algorithm formalizations.

Why the proof is difficult

The apparent shortcut is to treat an adaptive request sequence as fixed and apply a guarantee against oblivious adversaries directly. That loses the dependence of later requests on the algorithm's earlier answers. For Theorem 2.1, a winning request strategy must have one finite horizon that works for every answer path; separate finite horizons for each branch do not suffice when the answer set is infinite. For Theorem 2.2, the simulated adversary must make its own answers before the algorithm answers the current request, while still matching a fixed online benchmark along every resulting play. The expectation inequalities must remain valid when the request string itself depends on the algorithm's coins (Ben-David et al., 1994, manuscript pp. 9–11).

Formalization scope

Lean represents requests and answers as oldest-first lists. List index zero is request one in the paper. The general game is a separate definition; the algorithm, adversary, and competitiveness definitions build on it. An off-line adversary's rule returns Option R, where none is the stop signal, and has a uniform finite depth bound. A randomized algorithm consists of a coin probability space and a deterministic prefix algorithm for each coin; its answer events are measurable. Finiteness of AAA and bounded play depth make the cost of each fixed adversarial play take finitely many values, so its real expectation is an ordinary integrable expectation.

The formal game uses real-valued costs, a deliberate restriction of the paper's R∪{∞}\mathbb R\cup\{\infty\}R∪{∞} costs. The answer set is finite and nonempty, while the request set may be infinite. The transformations α\alphaα and β\betaβ are affine. Theorem 2.2 and the goal assume α\alphaα is monotone: the paper applies α\alphaα to an inequality in its proof, and its competitive-ratio examples have positive slope. Theorem 2.1 does not need this added assumption. The two randomized algorithms may have different coin spaces, each an arbitrary Lean type at the declaration's universe level. The off-line and on-line adaptive comparisons retain α\alphaα inside the expectation.

The target ranges over every equal-length request and answer play generated by these rules, including an adversary that stops without a request. It does not allow the deterministic algorithm to see future requests or choose a different policy for each adversary. Reusable contributions include the game interface, bounded adaptive plays, measurable randomized algorithms, and finite-horizon winning positions. The statements of all three principal results, their intervening claims, and proofs of those statements are within scope.

Selected references

  • S. Ben-David, A. Borodin, R. Karp, G. Tardos and A. Wigderson, On the Power of Randomization in On-Line Algorithms, Algorithmica 11, 1994. DOI: 10.1007/BF01294260. The local source is the authors' 20-page manuscript; citations above use its page numbers.
11 thms2 active usersReviewed
Operations ResearchTheoretical Computer Science·Captain: mikedeng1

On the Power of Randomization in On-Line Algorithms 2: The Bound α∘β Against Adaptive Off-Line Adversaries Is TightResearch Paper

Motivation

An on-line algorithm must answer each request as it arrives, without knowing the requests to come; paging, caching, the kkk-server problem and metrical task systems are standard examples. Its quality is measured by competitive analysis: its cost is compared with the cost of an optimal off-line solution that knows the whole request sequence. For randomized on-line algorithms the comparison depends on how much the adversary producing the requests is allowed to see. Ben-David, Borodin, Karp, Tardos and Wigderson (Algorithmica 11, 1994; conference version STOC 1990) introduced the three standard adversaries — oblivious, adaptive on-line and adaptive off-line — and related the competitive ratios achievable against each.

Their Theorem 2.2 (manuscript p. 10) shows that if a randomized algorithm is α\alphaα-competitive against adaptive on-line adversaries and some randomized algorithm is β\betaβ-competitive against oblivious adversaries, then the first algorithm is αβ\alpha\betaαβ-competitive against adaptive off-line adversaries. This mission formalizes the paper's claim (manuscript p. 11) that this product bound cannot be improved in general, together with the explicit construction on pp. 12–13 that proves it.

Setting

A request-answer game consists of a request set RRR, a finite answer set AAA, and cost functions fn:Rn×An→Rf_n : R^n \times A^n \to \mathbb Rfn​:Rn×An→R. For a request sequence r‾∈Rn\underline r \in R^nr​∈Rn, the off-line optimum is c(r‾)=min⁡a‾∈Anfn(r‾,a‾)c(\underline r) = \min_{\underline a \in A^n} f_n(\underline r, \underline a)c(r​)=mina​∈An​fn​(r​,a​). A deterministic on-line algorithm GGG answers the iii-th request with ai=gi(r1,…,ri)a_i = g_i(r_1, \dots, r_i)ai​=gi​(r1​,…,ri​); a randomized one is a probability distribution over deterministic algorithms GxG_xGx​, xxx being the coin tosses.

An adaptive off-line adversary QQQ chooses each request ri+1=qi(a1,…,ai)r_{i+1} = q_i(a_1, \dots, a_i)ri+1​=qi​(a1​,…,ai​) from the answers given so far, stops after at most dQd_QdQ​ requests, and pays the off-line optimum cQ(G)=c(r‾)c_Q(G) = c(\underline r)cQ​(G)=c(r​) of the requests it made; the algorithm pays cG(Q)=fn(r‾,a‾)c_G(Q) = f_n(\underline r, \underline a)cG​(Q)=fn​(r​,a​). An adaptive on-line adversary SSS must in addition answer each request itself, before the algorithm does, with bi+1=pi(a1,…,ai)b_{i+1} = p_i(a_1, \dots, a_i)bi+1​=pi​(a1​,…,ai​), and pays cS(G)=fn(r‾,b‾)c_S(G) = f_n(\underline r, \underline b)cS​(G)=fn​(r​,b​). An oblivious adversary fixes r‾\underline rr​ in advance and pays c(r‾)c(\underline r)c(r​). A randomized GGG is α\alphaα-competitive against oblivious adversaries if Ex[cGx(r‾)]≤α c(r‾)\mathbb E_x[c_{G_x}(\underline r)] \le \alpha\, c(\underline r)Ex​[cGx​​(r​)]≤αc(r​) for all r‾\underline rr​, and against adaptive on-line adversaries if Ex[cGx(S)]≤Ex[α cS(Gx)]\mathbb E_x[c_{G_x}(S)] \le \mathbb E_x[\alpha\, c_S(G_x)]Ex​[cGx​​(S)]≤Ex​[αcS​(Gx​)] for all SSS.

The construction uses the mates game: R=AR = AR=A is a set of 2t2t2t elements split into ttt pairs of mates, and for n≥2n \ge 2n≥2 the cost depends only on the first answer a1a_1a1​ and the second request r2r_2r2​: it is 111 if a1=r2a_1 = r_2a1​=r2​, MMM if a1a_1a1​ is the mate of r2r_2r2​, and mmm otherwise. The algorithm GGG draws a1a_1a1​ uniformly at random. The parameters solve

β=(2t−2)m+M+12t,α=1+(2t−1)M2+(2t−2)m.\beta = \frac{(2t-2)m + M + 1}{2t}, \qquad \alpha = \frac{1 + (2t-1)M}{2 + (2t-2)m}.β=2t(2t−2)m+M+1​,α=2+(2t−2)m1+(2t−1)M​.

Formalization targets

Goal: tightness of Theorem 2.2

For 1<β≤α1 < \beta \le \alpha1<β≤α (or α=β=1\alpha = \beta = 1α=β=1) and every C<αβC < \alpha\betaC<αβ, there are a game and a randomized algorithm GGG such that

G is α-competitive against adaptive on-line adversaries,G is β-competitive against oblivious adversaries,G \text{ is } \alpha\text{-competitive against adaptive on-line adversaries}, \qquad G \text{ is } \beta\text{-competitive against oblivious adversaries},G is α-competitive against adaptive on-line adversaries,G is β-competitive against oblivious adversaries,

and for every randomized algorithm KKK some adaptive off-line adversary QQQ achieves

E[cQ(K)]>0,E[cK(Q)]≥C⋅E[cQ(K)].\mathbb E[c_Q(K)] > 0, \qquad \mathbb E[c_K(Q)] \ge C\cdot \mathbb E[c_Q(K)].E[cQ​(K)]>0,E[cK​(Q)]≥C⋅E[cQ​(K)].

Milestones (pp. 12–13)

  1. The closed forms m(t)m(t)m(t), M(t)M(t)M(t) are the unique solution of the two equations.
  2. m(t)→βm(t) \to \betam(t)→β and M(t)→αβM(t) \to \alpha\betaM(t)→αβ as t→∞t \to \inftyt→∞.
  3. For all large ttt: M(t)≥max⁡(m(t)2,C)M(t) \ge \max(m(t)^2, C)M(t)≥max(m(t)2,C), 1≤m(t)≤M(t)1 \le m(t) \le M(t)1≤m(t)≤M(t), α(m(t)−1)≤M(t)−m(t)\alpha(m(t)-1) \le M(t) - m(t)α(m(t)−1)≤M(t)−m(t).
  4. GGG is β\betaβ-competitive against oblivious adversaries in the mates game.
  5. GGG is α\alphaα-competitive against adaptive on-line adversaries in the mates game.
  6. An adaptive off-line adversary makes every algorithm pay MMM while paying 111.

Significance

The result. Together with Theorem 2.2, the claim pins down exactly how much the adaptive off-line adversary can gain over the other two: the product αβ\alpha\betaαβ is an upper bound for every game and is approached by a single game for every admissible pair (α,β)(\alpha, \beta)(α,β). It shows that no general argument relating the three adversary models can give a bound better than the product, so any improvement for a specific problem (paging, kkk-server) must use the structure of that problem. The paging example cited on p. 11 (RANDOM against the three adversaries) gives one instance of tightness; the mates game gives tightness for every admissible pair.

Formalizing it. The result is proved in the paper, in about one page, with two steps left to the reader ("by inspection of the equations", "a simple case analysis"). No machine-checked proof of this or of any statement about adaptive adversaries is known to us. The formalization makes the model of §2 precise (sequences, stopping, the order in which adversary and algorithm commit, expectations over coins), checks the asymptotics of the parameters, and verifies the case analysis, which on inspection needs an inequality the page does not state. Two printed formulas on p. 12 contain typos; the formal statements carry the correct values.

Difficulty

The construction is explicit, but each competitiveness claim quantifies over all adversaries, which may adapt their requests to the algorithm's random answers, stop at any time, and (for the on-line adversary) commit to their own answers in advance. The algebra of α\alphaα-competitiveness is tight: the adversary's best expected advantage is exactly zero, so every case of its best reply must be checked with no slack. The page's condition M≥m2M \ge m^2M≥m2 does not suffice for this: when a1a_1a1​ is neither the adversary's first answer nor its mate, the reply "mate of a1a_1a1​" beats the reply "the adversary's own answer" only when α(m−1)≤M−m\alpha(m-1) \le M - mα(m−1)≤M−m, which holds for the solved parameters but is not implied by M≥m2M \ge m^2M≥m2. At β=1<α\beta = 1 < \alphaβ=1<α the solved parameter mmm is below 111 for every ttt, and the oblivious bound fails.

Formalization scope

All declarations live in OnlineRandomization.Tightness. The conventions:

  • Costs are real-valued; the paper allows +∞+\infty+∞, so the game class is a special case.
  • Answer sets are nonempty finite types; request sets are arbitrary types.
  • Sequences are Lean lists, oldest first; cost r a is fnf_nfn​ on lists of equal length nnn.
  • Adversaries return none for "stop" and carry a depth bound dQd_QdQ​; an on-line adversary's answer bi+1b_{i+1}bi+1​ depends only on a1,…,aia_1, \dots, a_ia1​,…,ai​.
  • Randomized algorithms are a probability space of coins with a deterministic algorithm per coin and measurable answers; expectations are Bochner integrals, with α\alphaα applied inside the expectation. In the goal, coin spaces range over Type.
  • Competitiveness uses the ratio functions x↦αxx \mapsto \alpha xx↦αx and x↦βxx \mapsto \beta xx↦βx, with no additive constant.
  • The mates game is on Fin t × Bool, with mate (i,b)↦(i,¬b)(i, b) \mapsto (i, \lnot b)(i,b)↦(i,¬b). The paper leaves the costs of plays with fewer than two requests undefined; the formalization sets f0=0f_0 = 0f0​=0 and f1≡1f_1 \equiv 1f1​≡1 (with f1≡0f_1 \equiv 0f1​≡0 the algorithm would not be α\alphaα-competitive).
  • Range. The goal assumes 1<β≤α1 < \beta \le \alpha1<β≤α or α=β=1\alpha = \beta = 1α=β=1; the page's case β=1<α\beta = 1 < \alphaβ=1<α is not covered by its construction and is left out. In fact the claim is false there for 1<C<α1 < C < \alpha1<C<α: an algorithm that is 111-competitive against oblivious adversaries answers optimally, almost surely, on every request sequence (its cost is never below the optimum and its expected cost does not exceed it), and an adaptive off-line adversary reaches only finitely many request sequences, so against K=GK = GK=G every adversary has E[cG(Q)]=E[cQ(G)]\mathbb E[c_G(Q)] = \mathbb E[c_Q(G)]E[cG​(Q)]=E[cQ​(G)], a ratio of 1<C1 < C1<C.

The positivity requirement E[cQ(K)]>0\mathbb E[c_Q(K)] > 0E[cQ​(K)]>0 in the goal is essential: without it the adversary that asks nothing satisfies E[cK(Q)]≥C⋅0\mathbb E[c_K(Q)] \ge C \cdot 0E[cK​(Q)]≥C⋅0 for every KKK, and the third clause would hold vacuously.

A complete development needs: finite expectations over a uniform coin, the evaluation of the play of an adversary against a constant algorithm, and limit and eventual-inequality arguments for rational functions of ttt. The model of §2 is shared with the other missions of this series and is reusable for any request-answer formulation of an on-line problem. Proofs of individual milestones are welcome.

Selected references

  • S. Ben-David, A. Borodin, R. Karp, G. Tardos, A. Wigderson, On the power of randomization in on-line algorithms, Algorithmica 11 (1994) 2–14. https://doi.org/10.1007/BF01294260 (cited from the authors' manuscript, manuscript pp. 7–13).
  • A. Borodin, R. El-Yaniv, Online Computation and Competitive Analysis, Cambridge University Press, 1998. ISBN 0-521-56392-5.
  • P. Raghavan, M. Snir, Memory versus randomization in on-line algorithms, IBM Journal of Research and Development 38 (1994) 683–707. https://doi.org/10.1147/rd.386.0683
9 thms2 active usersReviewed
Machine LearningOptimization·Captain: mikedeng1

Gradient Convergence in Gradient Methods with Errors II: With Zero-Mean Stochastic Errors, Almost Surely Either f(x_t) → −∞ or f(x_t) Converges and ∇f(x_t) → 0Research Paper

Motivation

Stochastic gradient methods minimize a function fff when only noisy estimates of its gradient are available: each step moves along a descent direction corrupted by random noise. They are the standard training algorithm for neural networks and the basic tool of stochastic approximation, and the question every user faces is what can be guaranteed when fff is nonconvex, possibly unbounded below, and the noise is allowed to grow with the gradient.

D. P. Bertsekas and J. N. Tsitsiklis, Gradient Convergence in Gradient Methods with Errors, SIAM J. Optim. 10(3):627–642, 2000 (DOI), answer this under minimal assumptions. Noise with variance growing in ∥∇f(xt)∥\|\nabla f(x_t)\|∥∇f(xt​)∥ had been handled for related methods (Poljak and Tsypkin 1973), but typically together with a lower bound on fff, under which f(xt)f(x_t)f(xt​) is approximately a supermartingale and the supermartingale convergence theorem applies (see the monographs of Kushner and Clark 1978; Benveniste, Métivier and Priouret 1990; Kushner and Yin 1996). Section 4 of the paper (p. 635) removes the lower bound: it proves that, with probability 1, either f(xt)→−∞f(x_t)\to-\inftyf(xt​)→−∞ or f(xt)f(x_t)f(xt​) converges and ∇f(xt)→0\nabla f(x_t)\to 0∇f(xt​)→0, without assuming bounded iterates. Section 5 shows that the randomized incremental gradient method for a finite-sum objective is a special case. This mission formalizes Section 4 and the Section 5 application. A companion mission covers the deterministic counterpart (Proposition 1 of the same paper).

Setting

Let f:Rn→Rf:\mathbb R^n\to\mathbb Rf:Rn→R be continuously differentiable with a Lipschitz gradient: there is L≥0L\ge 0L≥0 with

∥∇f(x)−∇f(xˉ)∥≤L∥x−xˉ∥∀x,xˉ.(2.1)\|\nabla f(x)-\nabla f(\bar x)\|\le L\|x-\bar x\|\qquad\forall x,\bar x. \tag{2.1}∥∇f(x)−∇f(xˉ)∥≤L∥x−xˉ∥∀x,xˉ.(2.1)

Let (Ω,F,P)(\Omega,\mathcal F,P)(Ω,F,P) be a probability space and F0⊆F1⊆⋯\mathcal F_0\subseteq\mathcal F_1\subseteq\cdotsF0​⊆F1​⊆⋯ an increasing sequence of σ\sigmaσ-fields (a filtration; Ft\mathcal F_tFt​ is the history of the algorithm just before the noise wtw_twt​ is drawn). The stochastic gradient method generates random vectors by

xt+1=xt+γt(st+wt),x_{t+1}=x_t+\gamma_t(s_t+w_t),xt+1​=xt​+γt​(st​+wt​),

where γt>0\gamma_t>0γt​>0 is a deterministic stepsize, sts_tst​ is a descent direction and wtw_twt​ is a noise term. The assumptions of Proposition 3 are:

  • (a) xtx_txt​ and sts_tst​ are Ft\mathcal F_tFt​-measurable;
  • (b) there are c1,c2>0c_1,c_2>0c1​,c2​>0 with c1∥∇f(xt)∥2≤−∇f(xt)′stc_1\|\nabla f(x_t)\|^2\le-\nabla f(x_t)'s_tc1​∥∇f(xt​)∥2≤−∇f(xt​)′st​ and ∥st∥≤c2(1+∥∇f(xt)∥)\|s_t\|\le c_2(1+\|\nabla f(x_t)\|)∥st​∥≤c2​(1+∥∇f(xt​)∥) for all ttt; (4.1)
  • (c) for all ttt, with probability 1, E[wt∣Ft]=0E[w_t\mid\mathcal F_t]=0E[wt​∣Ft​]=0 (4.2) and E[∥wt∥2∣Ft]≤A(1+∥∇f(xt)∥2)E[\|w_t\|^2\mid\mathcal F_t]\le A(1+\|\nabla f(x_t)\|^2)E[∥wt​∥2∣Ft​]≤A(1+∥∇f(xt​)∥2) (4.3), with A>0A>0A>0 deterministic;
  • (d) ∑tγt=∞\sum_t\gamma_t=\infty∑t​γt​=∞ and ∑tγt2<∞\sum_t\gamma_t^2<\infty∑t​γt2​<∞.

The noise variance in (c) may grow quadratically with ∥∇f(xt)∥\|\nabla f(x_t)\|∥∇f(xt​)∥ and is therefore unbounded in general. A point xˉ\bar xxˉ is stationary if ∇f(xˉ)=0\nabla f(\bar x)=0∇f(xˉ)=0.

Formalization targets

Goal: Proposition 3 (p. 635)

Under (2.1) and (a)–(d), with probability 1,

f(xt)→−∞or(f(xt)→ℓ∈R  and  ∇f(xt)→0),f(x_t)\to-\infty\quad\text{or}\quad\Bigl(f(x_t)\to\ell\in\mathbb R\ \text{ and }\ \nabla f(x_t)\to 0\Bigr),f(xt​)→−∞or(f(xt​)→ℓ∈R  and  ∇f(xt​)→0),

and every limit point of (xt)(x_t)(xt​) is a stationary point of fff. The dichotomy is per sample path: different paths may take different branches.

Milestones

  1. (4.4), p. 636. The pathwise one-step inequality: if γ 2Lc22≤c1/2\gamma\,2Lc_2^2\le c_1/2γ2Lc22​≤c1​/2 and sss satisfies (4.1) at xxx, then for every www,
f(x+γ(s+w))≤f(x)−γc12∥∇f(x)∥2+γ∇f(x)′w+γ22Lc22+γ2L∥w∥2.f(x+\gamma(s+w))\le f(x)-\gamma\tfrac{c_1}{2}\|\nabla f(x)\|^2+\gamma\nabla f(x)'w+\gamma^2 2Lc_2^2+\gamma^2L\|w\|^2.f(x+γ(s+w))≤f(x)−γ2c1​​∥∇f(x)∥2+γ∇f(x)′w+γ22Lc22​+γ2L∥w∥2.
  1. Lemma 2, p. 637. If rtr_trt​ is Ft+1\mathcal F_{t+1}Ft+1​-measurable with E[rt∣Ft]=0E[r_t\mid\mathcal F_t]=0E[rt​∣Ft​]=0, E[∥rt∥2∣Ft]≤BE[\|r_t\|^2\mid\mathcal F_t]\le BE[∥rt​∥2∣Ft​]≤B and ∑γt2<∞\sum\gamma_t^2<\infty∑γt2​<∞, then ∑t≤Tγtrt\sum_{t\le T}\gamma_tr_t∑t≤T​γt​rt​ and ∑t≤Tγt2∥rt∥2\sum_{t\le T}\gamma_t^2\|r_t\|^2∑t≤T​γt2​∥rt​∥2 converge almost surely.
  2. Lemma 6, p. 640. For every δ>0\delta>0δ>0, almost surely f(xt)f(x_t)f(xt​) converges to a finite value or to −∞-\infty−∞, and if the limit is not −∞-\infty−∞ then lim sup⁡t∥∇f(xt)∥≤δ\limsup_t\|\nabla f(x_t)\|\le\deltalimsupt​∥∇f(xt​)∥≤δ.

Further result: §5, pp. 641–642

For f=1m∑ifif=\frac1m\sum_i f_if=m1​∑i​fi​ with Lipschitz gradients ∇fi\nabla f_i∇fi​ satisfying ∥∇fi(x)∥≤C+D∥∇f(x)∥\|\nabla f_i(x)\|\le C+D\|\nabla f(x)\|∥∇fi​(x)∥≤C+D∥∇f(x)∥ (5.2), the randomized incremental gradient method xt+1=xt−γt∇fk(t)(xt)x_{t+1}=x_t-\gamma_t\nabla f_{k(t)}(x_t)xt+1​=xt​−γt​∇fk(t)​(xt​), with independent uniform indices k(t)k(t)k(t), satisfies the conclusion of Proposition 3.

Significance

Proposition 3 is a convergence guarantee for stochastic gradient descent on smooth nonconvex objectives that needs neither a lower bound on fff, nor bounded iterates, nor bounded noise variance. It contains, as special cases, stochastic gradient descent with unbiased gradient estimates whose variance grows with the gradient, the randomized incremental (single-sample) gradient method for finite sums of Section 5, and scaled or approximate gradient directions through condition (4.1). Its conclusion is the strongest one available at this generality: if f(xt)f(x_t)f(xt​) stays bounded below along a path, then the gradient vanishes along that path and every limit point is stationary.

The result is proved in the paper; to our knowledge it has not been machine-checked. Its formalization requires a working theory of generalized conditional expectations of non-integrable noise, square-integrable martingales in Rn\mathbb R^nRn and pathwise arguments over random interval partitions, which is reusable for other stochastic approximation results (Robbins–Monro type schemes, TD-learning, stochastic subgradient methods).

Difficulty

The natural first idea is to view f(xt)f(x_t)f(xt​) as a supermartingale up to summable errors and apply the supermartingale convergence theorem (Robbins–Siegmund). This fails here: the theorem needs f(xt)f(x_t)f(xt​) bounded below, and fff is not assumed bounded below; in addition the noise term γt2L∥wt∥2\gamma_t^2L\|w_t\|^2γt2​L∥wt​∥2 in (4.4) has conditional mean of order γt2∥∇f(xt)∥2\gamma_t^2\|\nabla f(x_t)\|^2γt2​∥∇f(xt​)∥2, which is not summable when the gradient is unbounded. Any argument must therefore extract a decrease of fff that dominates noise of the same order as the gradient itself, without a lower bound to anchor a supermartingale, and must do so along every sample path while the hypotheses are only conditional-expectation statements about non-integrable noise.

Formalization scope

  • Rn\mathbb R^nRn is EuclideanSpace ℝ (Fin n), ∇f\nabla f∇f is Mathlib's gradient f, fff is ContDiff ℝ 1, and the Lipschitz constant is L : ℝ≥0 with LipschitzWith L (gradient f) (the standing assumption (2.1), equivalent to the page's form).
  • The probability space is a Measure Ω with IsProbabilityMeasure, the σ\sigmaσ-fields are a Mathlib Filtration ℕ, and measurability in (a) is StronglyMeasurable[ℱ t]. The recursion and (4.1) hold on every sample path. Stepsizes are deterministic.
  • Conditional expectations of the noise. The page assumes no integrability of wtw_twt​, of ∥wt∥2\|w_t\|^2∥wt​∥2 or of f(xt)f(x_t)f(xt​), and none is added. Mathlib's condExp is 000 for non-integrable functions, so stating (4.2)–(4.3) with it would make them hold vacuously for any non-integrable noise; that encoding is ruled out. Instead (4.3) says that for every Ft\mathcal F_tFt​-measurable set SSS, ∫S∥wt∥2 dP≤∫SA(1+∥∇f(xt)∥2) dP\int_S\|w_t\|^2\,dP\le\int_S A(1+\|\nabla f(x_t)\|^2)\,dP∫S​∥wt​∥2dP≤∫S​A(1+∥∇f(xt​)∥2)dP (in [0,∞][0,\infty][0,∞]), and (4.2) says that ∫Swt dP=0\int_S w_t\,dP=0∫S​wt​dP=0 for every Ft\mathcal F_tFt​-measurable SSS on which wtw_twt​ is integrable. These are exactly the generalized conditional-expectation statements of the page.
  • In Lemma 2 the bound BBB is a constant, so Mathlib's condExp is used there, with integrability of ∥rt∥2\|r_t\|^2∥rt​∥2 stated explicitly; the page's hypothesis implies it. Lemma 2 is stated over any finite-dimensional real inner-product space, since the paper applies it to real and to vector-valued sequences.
  • ∑γt=∞\sum\gamma_t=\infty∑γt​=∞ is divergence of the partial sums; ∑γt2<∞\sum\gamma_t^2<\infty∑γt2​<∞ is Summable. Convergent random series (Lemma 2) are convergence of partial sums, not Summable, which would mean unconditional convergence. "lim sup⁡∥∇f(xt)∥≤δ\limsup\|\nabla f(x_t)\|\le\deltalimsup∥∇f(xt​)∥≤δ" is "for every δ′>δ\delta'>\deltaδ′>δ, eventually ∥∇f(xt)∥≤δ′\|\nabla f(x_t)\|\le\delta'∥∇f(xt​)∥≤δ′". Limit points are MapClusterPt.
  • In the §5 result the page's references to "section 4" and "(4.1)" are read as section 3 and condition (3.1), the indices k(t)k(t)k(t) run from t=0t=0t=0, x0x_0x0​ is deterministic, and the stepsizes are nonnegative as on the page.
  • Not stated: Lemma 3 (it needs the random interval construction of p. 636 as a definition), Lemmas 4–5 (steps that depend on the proof's own choice of ϵ\epsilonϵ), and the Remarks of §4.

Contributions welcome: a proof of Lemma 2 from Mathlib's martingale convergence theorems (Submartingale.exists_ae_tendsto_of_bdd), a proof of (4.4) from the descent lemma, a general bridge between the set-integral encoding of conditional expectations and Mathlib's condExp on localizing sets, and the interval construction behind Lemmas 3–6.

Selected references

  • D. P. Bertsekas and J. N. Tsitsiklis, Gradient Convergence in Gradient Methods with Errors, SIAM J. Optim. 10(3):627–642, 2000. https://doi.org/10.1137/S1052623497331063
  • B. T. Poljak and Y. Z. Tsypkin, Pseudogradient adaptation and training algorithms, Automat. Remote Control 12 (1973), 83–94.
  • H. J. Kushner and D. S. Clark, Stochastic Approximation Methods for Constrained and Unconstrained Systems, Springer, 1978.
  • H. J. Kushner and G. Yin, Stochastic Approximation Methods, Springer, 1996 (as cited in the paper).
  • A. Benveniste, M. Métivier and P. Priouret, Adaptive Algorithms and Stochastic Approximations, Springer, 1990.
  • D. P. Bertsekas and J. N. Tsitsiklis, Neuro-Dynamic Programming, Athena Scientific, 1996.
4 thms2 active usersReviewed
Machine LearningStatistics·Captain: mikedeng1

The Sample Complexity of Pattern Classification with Neural Networks: The Size of the Weights is More Important than the Size of the Network I: Fat-Shattering Margin Bound with d = fat_H(γ/16)Research Paper

Motivation

Classical generalization bounds for classifiers, built on the VC dimension, grow with the number of adjustable parameters. For neural networks this is at odds with practice: networks with many more weights than training examples often generalize well. Bartlett's 1998 paper (IEEE Trans. Inform. Theory 44(2), 525–536) explains part of this by measuring a real-valued classifier's confidence. If a hypothesis classifies most training examples correctly with a margin γ\gammaγ, its misclassification probability is controlled by a scale-sensitive dimension of the class at scale proportional to γ\gammaγ, not by its VC dimension. Later in the paper this yields bounds for networks with small weights that do not depend on the number of weights.

This mission formalizes the first of the paper's two main technical results, the margin bound of Theorem 2 (p. 527), together with the steps of its proof on pp. 527–528.

The fat-shattering dimension was introduced by Kearns and Schapire (JCSS 1994). Alon, Ben-David, Cesa-Bianchi and Haussler (J. ACM 1997) proved the scale-sensitive Sauer-type covering bound used here (Theorem 5 of the paper). Shawe-Taylor, Bartlett, Williamson and Anthony (IEEE Trans. Inform. Theory 1998) proved the zero-training-error version (Theorem 1 of the paper). Theorem 2 extends it to hypotheses that make margin errors on the training data.

Setting

Let XXX be a set and PPP a probability distribution on X×{−1,1}X\times\{-1,1\}X×{−1,1}. The threshold function is sgn⁡(α)=−1\operatorname{sgn}(\alpha)=-1sgn(α)=−1 for α<0\alpha<0α<0 and sgn⁡(α)=1\operatorname{sgn}(\alpha)=1sgn(α)=1 for α≥0\alpha\ge0α≥0. For a real-valued hypothesis hhh on XXX, the misclassification probability is er⁡P(h)=P{sgn⁡(h(x))≠y}\operatorname{er}_P(h)=P\{\operatorname{sgn}(h(x))\ne y\}erP​(h)=P{sgn(h(x))=y}. For a sample z=((x1,y1),…,(xm,ym))z=((x_1,y_1),\dots,(x_m,y_m))z=((x1​,y1​),…,(xm​,ym​)) drawn independently from PPP and γ>0\gamma>0γ>0, the margin error estimate is

er⁡^zγ(h)=1m ∣{i:yih(xi)<γ}∣.\widehat{\operatorname{er}}{}^{\gamma}_z(h)=\tfrac1m\,|\{i : y_ih(x_i)<\gamma\}|.erzγ​(h)=m1​∣{i:yi​h(xi​)<γ}∣.

Let HHH be a class of real functions on XXX. Points x1,…,xmx_1,\dots,x_mx1​,…,xm​ are γ\gammaγ-shattered by HHH if some r∈Rmr\in\mathbb R^mr∈Rm has the following property: for every sign vector b∈{−1,1}mb\in\{-1,1\}^mb∈{−1,1}m, some h∈Hh\in Hh∈H satisfies (h(xi)−ri)bi≥γ(h(x_i)-r_i)b_i\ge\gamma(h(xi​)−ri​)bi​≥γ for all iii. The fat-shattering dimension fat⁡H(γ)\operatorname{fat}_H(\gamma)fatH​(γ) is the largest such mmm, possibly ∞\infty∞.

The proof uses the following objects:

  • the squashing function πγ(α)=max⁡(−γ,min⁡(γ,α))\pi_\gamma(\alpha)=\max(-\gamma,\min(\gamma,\alpha))πγ​(α)=max(−γ,min(γ,α)) and the class πγ(H)={πγ∘h:h∈H}\pi_\gamma(H)=\{\pi_\gamma\circ h:h\in H\}πγ​(H)={πγ​∘h:h∈H};
  • the sample ℓ∞\ell_\inftyℓ∞​ pseudometric dℓ∞(x)(f,g)=max⁡i∣f(xi)−g(xi)∣d_{\ell_\infty(x)}(f,g)=\max_i|f(x_i)-g(x_i)|dℓ∞​(x)​(f,g)=maxi​∣f(xi​)−g(xi​)∣;
  • the covering number N∞(F,ϵ,m)\mathcal N_\infty(F,\epsilon,m)N∞​(F,ϵ,m), the largest over x∈Xmx\in X^mx∈Xm of the size of the smallest ϵ\epsilonϵ-cover (Definition 3), and the corresponding packing number M∞(F,α,m)\mathcal M_\infty(F,\alpha,m)M∞​(F,α,m);
  • the quantization Qα(x)=⌈(x−α/2)/α⌉αQ_\alpha(x)=\lceil (x-\alpha/2)/\alpha\rceil\alphaQα​(x)=⌈(x−α/2)/α⌉α.

Formalization targets

Goal: Theorem 2

Assume 0<δ<1/20<\delta<1/20<δ<1/2, 0<γ<10<\gamma<10<γ<1, m≥1m\ge1m≥1, and d=fat⁡H(γ/16)d=\operatorname{fat}_H(\gamma/16)d=fatH​(γ/16) finite with d≤34md\le 34md≤34m. With probability at least 1−δ1-\delta1−δ over zzz, every h∈Hh\in Hh∈H satisfies

er⁡P(h)<er⁡^zγ(h)+2m(dln⁡34emdlog⁡2(578m)+ln⁡4δ).\operatorname{er}_P(h)<\widehat{\operatorname{er}}{}^{\gamma}_z(h)+\sqrt{\frac2m\Bigl(d\ln\frac{34em}{d}\log_2(578m)+\ln\frac4\delta\Bigr)} .erP​(h)<erzγ​(h)+m2​(dlnd34em​log2​(578m)+lnδ4​)​.

Milestones, in the order the proof uses them

  1. Lemma 4. er⁡P(h)<er⁡^zγ(h)+(2/m)ln⁡(2N∞(πγ(H),γ/2,2m)/δ)\operatorname{er}_P(h)<\widehat{\operatorname{er}}{}^{\gamma}_z(h)+\sqrt{(2/m)\ln(2\mathcal N_\infty(\pi_\gamma(H),\gamma/2,2m)/\delta)}erP​(h)<erzγ​(h)+(2/m)ln(2N∞​(πγ​(H),γ/2,2m)/δ)​ uniformly over HHH, with probability at least 1−δ1-\delta1−δ.
  2. Theorem 5 (Alon et al.). If F:{1,…,n}→{1,…,b}F:\{1,\dots,n\}\to\{1,\dots,b\}F:{1,…,n}→{1,…,b} and fat⁡F(1)≤d\operatorname{fat}_F(1)\le dfatF​(1)≤d, then log⁡2N∞(F,2,n)<1+log⁡2(nb2)log⁡2∑i≤d(ni)bi\log_2\mathcal N_\infty(F,2,n)<1+\log_2(nb^2)\log_2\sum_{i\le d}\binom ni b^ilog2​N∞​(F,2,n)<1+log2​(nb2)log2​∑i≤d​(in​)bi, provided nnn is large enough.
  3. Writing F=Qγ/8(πγ(H))F=Q_{\gamma/8}(\pi_\gamma(H))F=Qγ/8​(πγ​(H)): fat⁡F(γ/8)≤fat⁡πγ(H)(γ/16)\operatorname{fat}_F(\gamma/8)\le\operatorname{fat}_{\pi_\gamma(H)}(\gamma/16)fatF​(γ/8)≤fatπγ​(H)​(γ/16).
  4. M∞(πγ(H),γ/2,2m)≤M∞(F,γ/2,2m)\mathcal M_\infty(\pi_\gamma(H),\gamma/2,2m)\le\mathcal M_\infty(F,\gamma/2,2m)M∞​(πγ​(H),γ/2,2m)≤M∞​(F,γ/2,2m).
  5. N∞(πγ(H),γ/2,2m)≤N∞(F,γ/4,2m)\mathcal N_\infty(\pi_\gamma(H),\gamma/2,2m)\le\mathcal N_\infty(F,\gamma/4,2m)N∞​(πγ​(H),γ/2,2m)≤N∞​(F,γ/4,2m).
  6. log⁡2N∞(πγ(H),γ/2,2m)<1+dlog⁡2(34em/d)log⁡2(578m)\log_2\mathcal N_\infty(\pi_\gamma(H),\gamma/2,2m)<1+d\log_2(34em/d)\log_2(578m)log2​N∞​(πγ​(H),γ/2,2m)<1+dlog2​(34em/d)log2​(578m) when 1≤d≤2m1\le d\le 2m1≤d≤2m and m≥dlog⁡2(34em/d)+1m\ge d\log_2(34em/d)+1m≥dlog2​(34em/d)+1.
  7. fat⁡πγ(H)(γ/16)≤fat⁡H(γ/16)\operatorname{fat}_{\pi_\gamma(H)}(\gamma/16)\le\operatorname{fat}_H(\gamma/16)fatπγ​(H)​(γ/16)≤fatH​(γ/16).

A further item, Proposition 8 (p. 529), is the probabilistic device the paper uses to make such bounds uniform over γ\gammaγ.

Significance

Theorem 2 is the bound behind the paper's main message. Corollary 9 makes it uniform over γ\gammaγ, and Theorem 28 combines it with fat-shattering estimates for networks with bounded weights. Together they show that a network classifying the training data with a large margin generalizes at a rate governed by the size of its weights, not by its number of weights. The same template, a margin error plus a capacity term at scale γ\gammaγ, underlies later margin analyses of support vector machines and boosting.

All results in this mission are proved in the literature; none is open. None is machine-checked on this platform: the platform has Rademacher-complexity margin bounds, but no statement about fat-shattering dimension or ℓ∞\ell_\inftyℓ∞​ sample covering numbers of real-valued classes. A complete formalization would provide a reusable library of these objects, with their basic inequalities between squashing, quantization, packing and covering. It would also give a checked version of the explicit constants 34em/d34em/d34em/d and 578m578m578m, which differ from those in later textbook treatments.

Difficulty

The bound is uniform over a possibly uncountable class HHH, so a union bound over hypotheses does not apply. The obvious replacement is a union bound over a cover of HHH. Two steps make it hard:

  • Lemma 4. It needs a ghost-sample symmetrization and a random-swap argument, carried out with an ℓ∞\ell_\inftyℓ∞​ cover of the squashed class on the double sample, so the cover depends on the data.
  • Theorem 5. Bounding that covering number by the fat-shattering dimension is a combinatorial counting argument about strongly shattered pairs. It is the scale-sensitive analogue of the Sauer–Shelah lemma, and here the bookkeeping of constants is exact.

The quantization steps look routine but carry the factor-of-two losses that produce the constants γ/16\gamma/16γ/16, 171717 and 578578578.

Formalization scope

The model is in the namespace BartlettNN.Margin.

  • Labels and samples. Labels are Bool, read as ±1\pm1±1 through pm (true is +1+1+1). sgn⁡(0)=1\operatorname{sgn}(0)=1sgn(0)=1. Samples are functions Fin m → X × Bool, indexed from 000, with law Measure.pi (fun _ => P). The margin estimate uses the strict inequality yih(xi)<γy_ih(x_i)<\gammayi​h(xi​)<γ, and shattering uses ≥γ\ge\gamma≥γ.
  • Fat-shattering dimension. fat⁡\operatorname{fat}fat is valued in ℕ∞. A ℕ-valued supremum would be 000 on an unbounded set, so the goal assumes fat H (γ/16) = d with d : ℕ.
  • Covering and packing numbers. Covers are finite and external (centres are arbitrary functions), the cover inequality is strict, and covering numbers are ⊤ when no finite cover exists. N∞\mathcal N_\inftyN∞​ and M∞\mathcal M_\inftyM∞​ are suprema over all samples, with repetitions allowed. "α\alphaα-separated", which the paper leaves undefined, is read as distance ≥α\ge\alpha≥α.
  • Logarithms. ln⁡\lnln is Real.log, log⁡2\log_2log2​ is Real.logb 2, and eee is Real.exp 1.
  • High probability. "With probability at least 1−δ1-\delta1−δ, every hhh" bounds the measure of the event that some h∈Hh\in Hh∈H violates the inequality. It is not a per-hypothesis statement.

Measurability. The paper states "we ignore issues of measurability, and assume that all sets considered are measurable" (p. 526). This is made explicit, not removed, through three hypotheses:

  • every h∈Hh\in Hh∈H is measurable;
  • the bad events {z:∃h∈H, er⁡P(h)≥er⁡^zγ(h)+ϵ}\{z:\exists h\in H,\ \operatorname{er}_P(h)\ge\widehat{\operatorname{er}}{}^{\gamma}_z(h)+\epsilon\}{z:∃h∈H, erP​(h)≥erzγ​(h)+ϵ} are measurable;
  • the double-sample events of display (1) are measurable.

Replacing these by countability of HHH would weaken the theorem.

Corrections of the printed text.

  • Theorem 2. The goal adds d≤34md\le 34md≤34m. Beyond 34m34m34m the term dln⁡(34em/d)d\ln(34em/d)dln(34em/d) decreases, vanishes at d=34emd=34emd=34em and then turns negative, and the printed statement fails for rich classes. Within this range nothing is lost: the proof covers d≤2md\le2md≤2m, and for 2m<d≤34m2m<d\le34m2m<d≤34m the bound exceeds 111.
  • Milestone 6. It carries the hypothesis d≤2md\le 2md≤2m, the range of the binomial estimate behind 34em/d34em/d34em/d.
  • Milestone 3. Its printed justification ∣Qγ/8(a)−Qγ/8(b)∣<∣a−b∣+γ/16|Q_{\gamma/8}(a)-Q_{\gamma/8}(b)|<|a-b|+\gamma/16∣Qγ/8​(a)−Qγ/8​(b)∣<∣a−b∣+γ/16 is false; the correct term is γ/8\gamma/8γ/8. The milestone's conclusion is true as printed, and only the conclusion is formalized.

Trivializing formalizations, ruled out. The following would each make the statements empty or different, and none is used:

  • a ℕ-valued fat dimension or covering number;
  • Real.sign in place of sgn⁡\operatorname{sgn}sgn;
  • a per-hypothesis probability bound;
  • an unrestricted ddd, which makes ⋅\sqrt{\cdot}⋅​ of a negative number equal to 000;
  • a covering number that is 000 on classes without finite covers.

Infrastructure that a complete development needs, and contributions that are welcome:

  • product measures and Hoeffding's inequality, which Mathlib has;
  • a symmetrization (ghost-sample) lemma for margin events;
  • the combinatorics of Theorem 5;
  • the elementary inequalities between packing and covering numbers.

The covering/packing and fat-shattering lemmas apply beyond this mission. Proofs of individual milestones, or of Theorem 5 in the generality of Alon et al., are useful contributions in their own right.

Selected references

  • P. L. Bartlett, The Sample Complexity of Pattern Classification with Neural Networks: The Size of the Weights is More Important than the Size of the Network, IEEE Trans. Inform. Theory 44(2), 525–536, 1998. https://doi.org/10.1109/18.661502
  • N. Alon, S. Ben-David, N. Cesa-Bianchi, D. Haussler, Scale-sensitive dimensions, uniform convergence, and learnability, J. ACM 44(4), 615–631, 1997. https://doi.org/10.1145/263867.263927
  • J. Shawe-Taylor, P. L. Bartlett, R. C. Williamson, M. Anthony, Structural risk minimization over data-dependent hierarchies, IEEE Trans. Inform. Theory 44(5), 1926–1940, 1998. https://doi.org/10.1109/18.705570
  • M. J. Kearns, R. E. Schapire, Efficient distribution-free learning of probabilistic concepts, J. Comput. Syst. Sci. 48(3), 464–497, 1994. https://doi.org/10.1016/S0022-0000(05)80062-5
  • V. N. Vapnik, A. Ya. Chervonenkis, On the uniform convergence of relative frequencies of events to their probabilities, Theory Probab. Appl. 16(2), 264–280, 1971. https://doi.org/10.1137/1116025
12 thms2 active usersReviewed
Operations ResearchTheoretical Computer Science·Captain: mikedeng1

Secretary Problems: Weights and Discounts 3: An O(log n)-Competitive Algorithm for the Discounted Secretary ProblemResearch Paper

Motivation

In the secretary problem, nnn candidates with arbitrary values arrive one at a time in a uniformly random order, and an online decision maker must accept or reject each candidate on arrival, irrevocably, keeping at most one. The rule that observes the first n/en/en/e candidates and then accepts the first one better than everything seen so far selects the best candidate with probability about 1/e1/e1/e (Dynkin, 1963). The problem is a basic model of online selection and, read economically, of posted-price mechanisms for agents who arrive in random order: a rule that compares each agent only against a threshold set by earlier agents is truthful.

Babaioff, Dinitz, Gupta, Immorlica and Talwar (SODA 2009) study a variant in which time costs value. Selecting the candidate who arrives at time ttt earns that candidate's value multiplied by a discount d(t)d(t)d(t), for an arbitrary non-negative discount function ddd known in advance. Earlier work treated only specific discount shapes, such as geometric discounting d(t)=βtd(t)=\beta^td(t)=βt (Rasmussen and Pliska, 1976). For a general ddd the classical rule can fail badly: if all the discount mass sits in the first few time steps, a rule that waits through a sample of size n/en/en/e earns nothing. The paper shows that the best competitive ratio for arbitrary discounts lies between Ω(log⁡n/log⁡log⁡n)\Omega(\log n/\log\log n)Ω(logn/loglogn) (its Theorem 4.3) and O(log⁡n)O(\log n)O(logn) (its Theorem 4.4). This mission formalizes the upper bound.

Setting

There are n≥1n\ge1n≥1 elements, indexed {0,…,n−1}\{0,\dots,n-1\}{0,…,n−1}, with values v(e)≥0v(e)\ge 0v(e)≥0, and nnn times with discounts d(t)≥0d(t)\ge0d(t)≥0. A uniformly random permutation π\piπ fixes the order of arrivals: element π(t)\pi(t)π(t) arrives at time ttt. An algorithm knows ddd but not vvv; it sees each value on arrival and may select the current element, irrevocably, earning d(t) v(π(t))d(t)\,v(\pi(t))d(t)v(π(t)). Expectations over π\piπ are exact averages over the n!n!n! orders.

The offline optimum on order π\piπ is OPT(π)=max⁡td(t) v(π(t))\mathsf{OPT}(\pi)=\max_t d(t)\,v(\pi(t))OPT(π)=maxt​d(t)v(π(t)); it is a random variable, and the benchmark is its expectation Eπ[OPT]\mathbb E_\pi[\mathsf{OPT}]Eπ​[OPT] (p. 4 of the paper).

Let dmax⁡=max⁡td(t)d_{\max}=\max_t d(t)dmax​=maxt​d(t) and vmax⁡=max⁡ev(e)v_{\max}=\max_e v(e)vmax​=maxe​v(e). For c≥1c\ge1c≥1 the ccc-th discount class is the set of times

Pc={ i:2−cdmax⁡<d(i)≤2−(c−1)dmax⁡ }.P_c=\{\,i : 2^{-c}d_{\max}<d(i)\le 2^{-(c-1)}d_{\max}\,\}.Pc​={i:2−cdmax​<d(i)≤2−(c−1)dmax​}.

The quantity OPTc\mathsf{OPT}_cOPTc​ is the part of Eπ[OPT]\mathbb E_\pi[\mathsf{OPT}]Eπ​[OPT] earned when the optimal time (the smallest time attaining the maximum) lies in PcP_cPc​.

The classical secretary rule on mmm arrivals observes the first ⌊m/e⌋\lfloor m/e\rfloor⌊m/e⌋ and then selects the first arrival that ranks above every earlier arrival. Ranks use a fixed tie-break order: larger value first, and smaller element index among equal values.

The algorithm A\mathcal AA sets M=3⌈log⁡2n⌉+2M=3\lceil\log_2 n\rceil+2M=3⌈log2​n⌉+2, draws c∈{1,…,M}c\in\{1,\dots,M\}c∈{1,…,M} uniformly, and runs the classical rule on the subsequence of arrivals at the times of PcP_cPc​, ignoring all other arrivals.

Formalization targets

Goal: Theorem 4.4 with its explicit constant

Eπ[OPT]  ≤  4e (3⌈log⁡2n⌉+2)  E[A](n≥1, d≥0, v≥0).\mathbb E_\pi[\mathsf{OPT}]\;\le\;4e\,\bigl(3\lceil\log_2 n\rceil+2\bigr)\;\mathbb E[\mathcal A]\qquad(n\ge1,\ d\ge0,\ v\ge0).Eπ​[OPT]≤4e(3⌈log2​n⌉+2)E[A](n≥1, d≥0, v≥0).

The paper states E[OPT]/E[A]≤O(log⁡n)\mathbb E[\mathsf{OPT}]/\mathbb E[\mathcal A]\le O(\log n)E[OPT]/E[A]≤O(logn); the constant 4e4e4e is the one its proof yields.

Milestones

  1. The classical secretary rule selects the top-ranked of m≥1m\ge1m≥1 elements with probability at least 1/e1/e1/e (§2, p. 4).
  2. OPT1≥vmax⁡dmax⁡/n\mathsf{OPT}_1\ge v_{\max}d_{\max}/nOPT1​≥vmax​dmax​/n (proof of Theorem 4.4, p. 7).
  3. OPTc≤2−c 2n2dmax⁡vmax⁡\mathsf{OPT}_c\le 2^{-c}\,2n^2d_{\max}v_{\max}OPTc​≤2−c2n2dmax​vmax​ for every c≥1c\ge1c≥1 (p. 7).
  4. ∑c=13⌈log⁡2n⌉+1OPTc≥12Eπ[OPT]\sum_{c=1}^{3\lceil\log_2 n\rceil+1}\mathsf{OPT}_c\ge\tfrac12\mathbb E_\pi[\mathsf{OPT}]∑c=13⌈log2​n⌉+1​OPTc​≥21​Eπ​[OPT] (p. 7).
  5. E[Ac]≥OPTc/2e\mathbb E[\mathcal A_c]\ge\mathsf{OPT}_c/2eE[Ac​]≥OPTc​/2e for every c≥1c\ge1c≥1, where Ac\mathcal A_cAc​ is the classical rule on PcP_cPc​ (p. 7).

Significance

The theorem shows that a general discount function costs only a logarithmic factor against the offline benchmark, and that one algorithm achieves this without any knowledge of the values. Together with the lower bound of Theorem 4.3 it pins the competitive ratio of the discounted secretary problem between log⁡n/log⁡log⁡n\log n/\log\log nlogn/loglogn and log⁡n\log nlogn. The same scale-splitting idea, stated in the paper as Theorem 4.5 without full proof, extends the bound to the weighted discounted problem.

The result is proved in the paper; to our knowledge it has not been formalized. A complete development would also produce a machine-checked proof of the classical secretary guarantee for the rule with sample size exactly ⌊m/e⌋\lfloor m/e\rfloor⌊m/e⌋ at every finite mmm, with an explicit tie-break, which is reusable by every secretary-type mission. Milestone 1 is that statement. Sharper constants or a smaller class range are welcome as additional statements but do not replace the goal, which is about this algorithm with this MMM.

Difficulty

The obvious argument, running the classical rule on all nnn arrivals, fails because the discounts can be concentrated at times the rule spends sampling. Splitting by discount scale fixes this but creates two problems. First, there are unboundedly many scales, and one has to show that the offline optimum's mass outside the top O(log⁡n)O(\log n)O(logn) of them is negligible against E[OPT]\mathbb E[\mathsf{OPT}]E[OPT], a random quantity rather than a fixed maximum. Second, the classical rule on a class sees only a random subset of the elements, in random order, and the guarantee must be transferred to this subsequence, conditioning on which elements land in PcP_cPc​. Neither step is deep, but both require careful bookkeeping of permutations, and the classical 1/e1/e1/e bound at finite mmm with a floor in the sample size is itself a nontrivial estimate.

Formalization scope

Elements and times are Fin n, an order is π : Equiv.Perm (Fin n) read as time ↦\mapsto↦ element, and the paper's time t=1,…,nt=1,\dots,nt=1,…,n is index t−1t-1t−1. Values and discounts are Fin n → ℝ with non-negativity hypotheses. Every expectation is the finite average 1n!∑π\frac1{n!}\sum_\pin!1​∑π​; the algorithm's random class is the explicit average 1M∑c=1M\frac1M\sum_{c=1}^MM1​∑c=1M​. Maxima are suprema over the finite index set. The logarithm is base 2, ⌈log⁡2n⌉\lceil\log_2 n\rceil⌈log2​n⌉ is Nat.clog 2 n, and the sample size is Nat.floor (m / Real.exp 1). Ties are broken by the order on Lex (ℝ × (Fin n)ᵒᵈ) (larger value, then smaller index); distinct values are not assumed. The optimal time is the smallest maximizing time, so that the OPTc\mathsf{OPT}_cOPTc​ add up to E[OPT]\mathbb E[\mathsf{OPT}]E[OPT]. Competitiveness is stated multiplicatively, never as a quotient, so E[A]=0\mathbb E[\mathcal A]=0E[A]=0 is not a loophole.

The goal is a statement about the specific algorithm A\mathcal AA, not "there exists an algorithm": an existential over unrestricted algorithms is witnessed by a clairvoyant rule that reads the values in advance. A\mathcal AA sees the values only through comparisons among arrivals that have already occurred, and E[OPT]\mathbb E[\mathsf{OPT}]E[OPT] is the expected offline maximum over the same random order, not dmax⁡vmax⁡d_{\max}v_{\max}dmax​vmax​.

Needed infrastructure: averages over permutations and the fact that the elements landing at a fixed set of times form a uniformly random subset in uniformly random order; the finite-mmm analysis of the classical rule; and elementary estimates on geometric sums. Contributions of general lemmas about uniform permutations are welcome and reusable.

Selected references

  • M. Babaioff, M. Dinitz, A. Gupta, N. Immorlica, K. Talwar, Secretary Problems: Weights and Discounts, Proc. 20th ACM-SIAM Symposium on Discrete Algorithms (SODA), 2009. https://doi.org/10.1137/1.9781611973068.139
  • E. B. Dynkin, The optimum choice of the instant for stopping a Markov process, Soviet Math. Doklady 4, 1963.
  • T. S. Ferguson, Who solved the secretary problem?, Statistical Science 4(3), 1989. https://doi.org/10.1214/ss/1177012493
  • L. T. Rasmussen, S. R. Pliska, Choosing the maximum from a sequence with a discount function, Applied Mathematics and Optimization 2, 1976. https://doi.org/10.1007/BF01458209
  • M. Babaioff, N. Immorlica, R. Kleinberg, Matroids, secretary problems, and online mechanisms, SODA 2007. https://dl.acm.org/doi/10.5555/1283383.1283429
9 thms2 active usersReviewed
Mathematical Physics·Captain: mikedeng1

The Intermediate Disorder Regime for Directed Polymers in Dimension 1+1: The Rescaled Partition Function Converges in Law to a Wiener ChaosResearch Paper

Motivation

A directed polymer in a random environment is a simple random walk whose paths are reweighted by a random field of energies. It is a basic model of disordered statistical mechanics, and in dimension 1+11+11+1 it belongs to the Kardar–Parisi–Zhang (KPZ) universality class: at fixed temperature the free energy fluctuates on the scale n1/3n^{1/3}n1/3 and is expected to follow Tracy–Widom laws. Comets and Yoshida (Ann. Probab. 34 (2006)) showed that in dimension 1+11+11+1 every positive inverse temperature β\betaβ lies in the strong disorder regime, where the normalised partition function tends to zero. At β=0\beta=0β=0 the polymer is simply the random walk.

Alberts, Khanin and Quastel (arXiv:1202.4398, Ann. Probab. 42 (2014)) identified the regime in between. If the inverse temperature is scaled as βn−1/4\beta n^{-1/4}βn−1/4, the partition function neither concentrates nor vanishes: it converges in law to a universal random variable, a Wiener chaos in a space–time white noise. That variable is the solution at time 111, integrated in space, of the stochastic heat equation with multiplicative noise, whose logarithm is the Hopf–Cole solution of the KPZ equation. The theorem therefore connects discrete polymers to the continuum KPZ equation under weak, nnn-dependent disorder, with no assumption on the environment beyond exponential moments.

Setting

The environment is a family ω=(ω(i,x))i≥1, x∈Z\omega=(\omega(i,x))_{i\ge1,\,x\in\mathbb Z}ω=(ω(i,x))i≥1,x∈Z​ of i.i.d. real random variables on a probability space (Ω,Q)(\Omega,Q)(Ω,Q), with mean 000 and variance 111. Write λ(β)=log⁡Q eβω\lambda(\beta)=\log Q\,e^{\beta\omega}λ(β)=logQeβω. The polymer is the symmetric simple random walk SSS on Z\mathbb ZZ started at 000, independent of ω\omegaω, under the uniform measure P\mathbf PP on the 2n2^n2n step sequences. The energy of an nnn-step path is Hnω(S)=∑i=1nω(i,Si)H_n^\omega(S)=\sum_{i=1}^n\omega(i,S_i)Hnω​(S)=∑i=1n​ω(i,Si​), and the point-to-line partition function is

Znω(β)=P[eβHnω(S)].Z_n^\omega(\beta)=\mathbf P\big[e^{\beta H_n^\omega(S)}\big].Znω​(β)=P[eβHnω​(S)].

The modified partition function replaces eβωe^{\beta\omega}eβω by 1+βω1+\beta\omega1+βω: znω(β)=P[∏i=1n(1+β ω(i,Si))]\mathfrak z_n^\omega(\beta)=\mathbf P\big[\prod_{i=1}^n(1+\beta\,\omega(i,S_i))\big]znω​(β)=P[∏i=1n​(1+βω(i,Si​))].

On the continuum side, a white noise on [0,1]×R[0,1]\times\mathbb R[0,1]×R is a centred Gaussian family {W(A)}\{W(A)\}{W(A)}, indexed by the Borel sets of finite Lebesgue measure, with E[W(A)W(B)]=∣A∩B∣E[W(A)W(B)]=|A\cap B|E[W(A)W(B)]=∣A∩B∣. Its multiple stochastic integrals are continuous linear maps Ik:L2([0,1]k×Rk)→L2I_k:L^2([0,1]^k\times\mathbb R^k)\to L^2Ik​:L2([0,1]k×Rk)→L2 with Ik(1A1×⋯×Ak)=∏jW(Aj)I_k(\mathbf 1_{A_1\times\dots\times A_k})=\prod_jW(A_j)Ik​(1A1​×⋯×Ak​​)=∏j​W(Aj​) for pairwise disjoint AjA_jAj​. Let ϱ(t,x)=e−x2/2t/2πt\varrho(t,x)=e^{-x^2/2t}/\sqrt{2\pi t}ϱ(t,x)=e−x2/2t/2πt​ be the heat kernel and Δk={0=t0<t1<⋯<tk≤1}\Delta_k=\{0=t_0<t_1<\dots<t_k\le1\}Δk​={0=t0​<t1​<⋯<tk​≤1}. Put ϱk(t,x)=∏j=1kϱ(tj−tj−1,xj−xj−1)\varrho_k(\mathbf t,\mathbf x)=\prod_{j=1}^k\varrho(t_j-t_{j-1},x_j-x_{j-1})ϱk​(t,x)=∏j=1k​ϱ(tj​−tj−1​,xj​−xj−1​) on Δk×Rk\Delta_k\times\mathbb R^kΔk​×Rk, with x0=0x_0=0x0​=0, and zero elsewhere. The Wiener chaos (7) is

Zβ=∑k≥0βkIk(ϱk)=1+∑k≥1βk∫Δk∫Rk∏i=1kW(ti,xi) ϱ(ti−ti−1,xi−xi−1) dxi dti.\mathcal Z_\beta=\sum_{k\ge0}\beta^kI_k(\varrho_k)=1+\sum_{k\ge1}\beta^k\int_{\Delta_k}\int_{\mathbb R^k}\prod_{i=1}^kW(t_i,x_i)\,\varrho(t_i-t_{i-1},x_i-x_{i-1})\,dx_i\,dt_i .Zβ​=k≥0∑​βkIk​(ϱk​)=1+k≥1∑​βk∫Δk​​∫Rk​i=1∏k​W(ti​,xi​)ϱ(ti​−ti−1​,xi​−xi−1​)dxi​dti​.

The discrete counterpart of IkI_kIk​ is the weighted U-statistic Skn(g)\mathcal S_k^n(g)Skn​(g) of (29): a sum of products ∏jω(ij,xj)\prod_j\omega(i_j,x_j)∏j​ω(ij​,xj​) over distinct times, weighted by averages of ggg over space–time rectangles of size 1n×2n\frac1n\times\frac2{\sqrt n}n1​×n​2​.

Formalization targets

Goal: Theorem 2.1 (second bullet), proved as Proposition 5.4

If, in addition, λ(b)<∞\lambda(b)<\inftyλ(b)<∞ for all 0<b<β00<b<\beta_00<b<β0​ for some β0>0\beta_0>0β0​>0, then for every β>0\beta>0β>0

e−nλ(βn−1/4) Znω(βn−1/4)→(d)Z2β(n→∞).e^{-n\lambda(\beta n^{-1/4})}\,Z_n^\omega(\beta n^{-1/4})\xrightarrow{(d)}\mathcal Z_{\sqrt2\beta}\qquad(n\to\infty).e−nλ(βn−1/4)Znω​(βn−1/4)(d)​Z2​β​(n→∞).

Modified partition function: Proposition 5.3 (Theorem 2.1, first bullet)

Under mean zero and variance one alone, znω(βn−1/4)→(d)Z2β\mathfrak z_n^\omega(\beta n^{-1/4})\xrightarrow{(d)}\mathcal Z_{\sqrt2\beta}znω​(βn−1/4)(d)​Z2​β​.

Milestones

In the order the proof uses them:

  1. the norm ∥ϱk∥L22=1/(2kΓ(k/2+1))\|\varrho_k\|^2_{L^2}=1/(2^k\Gamma(k/2+1))∥ϱk​∥L22​=1/(2kΓ(k/2+1)) and the convergence of the chaos series (§3.4, Lemma 3.1);
  2. the L2L^2L2 structure of Skn\mathcal S_k^nSkn​ (Lemma 4.1);
  3. an approximation lemma for convergence in law (Lemma 4.2);
  4. convergence of n−3k/4Skn(g)n^{-3k/4}\mathcal S_k^n(g)n−3k/4Skn​(g) to Ik(g)I_k(g)Ik​(g), jointly over finitely many orders (Theorem 4.3);
  5. convergence of whole discrete chaos expansions (Lemma 4.4);
  6. the exact expansion znω(β)=∑k≤n2k/2βkSkn(pkn)\mathfrak z_n^\omega(\beta)=\sum_{k\le n}2^{k/2}\beta^k\mathcal S_k^n(p_k^n)znω​(β)=∑k≤n​2k/2βkSkn​(pkn​) (Lemma 5.2);
  7. a uniform L2L^2L2 bound on the discretised random walk kernels nk/2pknn^{k/2}p_k^nnk/2pkn​ (Lemma A.1);
  8. Proposition 5.3.

Significance

The theorem gives a universal scaling limit for the partition function in a regime where the polymer still moves diffusively but feels the disorder. As a corollary, log⁡Znω(βn−1/4)−nλ(βn−1/4)\log Z_n^\omega(\beta n^{-1/4})-n\lambda(\beta n^{-1/4})logZnω​(βn−1/4)−nλ(βn−1/4) converges in law, so the fluctuation exponent of the free energy is 000 in this regime. The companion results (point-to-point partition functions, Theorem 2.2) show that the rescaled polymer path measure converges to a continuum directed random polymer. Letting β\betaβ grow then interpolates between Gaussian behaviour and the Tracy–Widom GUE fluctuations of the KPZ class. The U-statistic machinery of Section 4 (discrete chaos expansions in an i.i.d. space–time field converging to multiple Wiener–Itô integrals) applies to other discrete models with polynomial chaos expansions.

The theorem is proved in the paper. To our knowledge it has no machine-checked proof, and Mathlib has no space–time white noise, no multiple Wiener integrals, no Wiener chaos, no Lindeberg–Feller theorem for triangular arrays and no local limit theorem for the simple random walk. A complete formalization would build each of these and check the paper's argument, including two statements that are false as printed (see Formalization scope).

Difficulty

The obvious route is to expand Znω(βn−1/4)Z_n^\omega(\beta n^{-1/4})Znω​(βn−1/4) in powers of β\betaβ and pass to the limit term by term. Two steps of that route do not go through as stated. First, the kkk-th term is a degenerate U-statistic of order kkk in the environment. Its convergence to IkI_kIk​ is not a classical central limit theorem: it needs joint control of all orders and a density argument in L2([0,1]k×Rk)L^2([0,1]^k\times\mathbb R^k)L2([0,1]k×Rk), and the space–time discretisation must respect the parity of the walk, which is why the rectangles have spatial length 2/n2/\sqrt n2/n​. Second, exchanging the limit with the infinite sum over kkk needs bounds on the discrete kernels nk/2pknn^{k/2}p_k^nnk/2pkn​ that are uniform in nnn and summable in kkk. Pointwise convergence from the local limit theorem is not enough. Passing from zn\mathfrak z_nzn​ to ZnZ_nZn​ also requires a central limit theorem for a triangular array whose law depends on nnn.

Formalization scope

The environment is ω : ℕ × ℤ → Ω → ℝ, independent and identically distributed, square integrable, with mean 000 and variance 111. Times i≥1i\ge1i≥1 are read. ZnZ_nZn​ and zn\mathfrak z_nzn​ are explicit averages over the 2n2^n2n step sequences. The white noise lives on [0,1]×R[0,1]\times\mathbb R[0,1]×R, since only t≤1t\le1t≤1 enters (7). IkI_kIk​ is a family of continuous linear maps on L2L^2L2 characterised by its values on indicators of products of disjoint sets. Every limit theorem quantifies over all white noises and all such families, and a separate item asserts that one exists. Zβ\mathcal Z_\betaZβ​ and Skn\mathcal S_k^nSkn​ are sums in L2L^2L2; real powers n−1/4n^{-1/4}n−1/4, n−3k/4n^{-3k/4}n−3k/4 are Real.rpow; λ\lambdaλ is Mathlib's cgf; convergence in law is TendstoInDistribution.

Normalisation of IkI_kIk​. The paper's normalisation of multiple integrals is not consistent (§3.2 and the remark on p. 19 differ by k!k!k!). The formalization follows (7) and the variance computation on p. 6: for ggg supported on Δk×Rk\Delta_k\times\mathbb R^kΔk​×Rk, E[Ik(g)2]=∥g∥2E[I_k(g)^2]=\|g\|^2E[Ik​(g)2]=∥g∥2.

Corrected statements. Two milestones are stated in the corrected form that the paper's proof establishes:

  • Lemma 4.1. The bound Q[Skn(g)2]≤n3k/2∥g∥2Q[\mathcal S_k^n(g)^2]\le n^{3k/2}\|g\|^2Q[Skn​(g)2]≤n3k/2∥g∥2 is false for general ggg when k≥2k\ge2k≥2, because permuted index vectors give the same product of environment variables. It is stated for ggg vanishing outside Δk×Rk\Delta_k\times\mathbb R^kΔk​×Rk, the only case Section 5 uses. "Mean zero" is stated for k≥1k\ge1k≥1, since S0n(g0)=g0\mathcal S_0^n(g_0)=g_0S0n​(g0​)=g0​.
  • Lemma 4.4. It is stated for Fock vectors whose components vanish outside the simplices, with square-summable norms, so that ∑kIk(gk)\sum_kI_k(g_k)∑k​Ik​(gk​) converges.

Theorem 4.5 and Lemma 4.6 (perturbed environments) are not part of the milestone list. Theorem 4.5 is false as printed: mean zero and a variance tending to 111 do not give a central limit theorem for a triangular array, and the Lindeberg condition named in its proof must be added. Lemma A.1 is stated for the point-to-line kernels only.

Trivializations ruled out. The limit is the chaos (7) built from a genuine white noise and its multiple integrals, not any random variable with a prescribed property. III is pinned down by WWW. Every infinite sum comes with a milestone proving its summability, and the partition function averages over all 2n2^n2n walk paths.

Infrastructure that would be reusable beyond this mission: space–time white noise and multiple Wiener–Itô integrals with their isometry, a Lindeberg–Feller central limit theorem, the local limit theorem for the simple random walk, and the Cramér–Wold device. Contributions of any of these, or of proofs of the milestones in any order, are welcome.

Selected references

  • T. Alberts, K. Khanin, J. Quastel, The intermediate disorder regime for directed polymers in dimension 1+1, Ann. Probab. 42 (2014), 1212–1256. arXiv:1202.4398
  • T. Alberts, K. Khanin, J. Quastel, The continuum directed random polymer, J. Stat. Phys. 154 (2014), 305–326. MR3162542
  • F. Comets, N. Yoshida, Directed polymers in random environment are diffusive at weak disorder, Ann. Probab. 34 (2006), 1746–1770. MR2271480
  • S. Janson, Gaussian Hilbert Spaces, Cambridge Tracts in Mathematics 129, Cambridge University Press, 1997. MR1474726
  • P. Billingsley, Convergence of Probability Measures, Wiley, 1968. MR0233396
12 thms2 active usersReviewed
Bandit AlgorithmsMachine LearningOperations Research·Captain: mikedeng1

Analysis of Thompson Sampling for the Multi-armed Bandit Problem 1: Logarithmic Regret for Two ArmsResearch Paper

Motivation

Thompson Sampling (TS) is the oldest heuristic for the stochastic multi-armed bandit problem: it was proposed by Thompson in 1933 (Biometrika 25) and is used in practice for online advertising and recommendation, where it often performs as well as or better than upper-confidence-bound methods (Chapelle and Li, NIPS 2011; Scott 2010). Until 2012 its theoretical guarantees for the frequentist regret were weak: earlier analyses gave only o(T)o(T)o(T) regret in time TTT (Granmo 2010; May, Korda, Lee and Leslie 2011).

Agrawal and Goyal, Analysis of Thompson Sampling for the Multi-armed Bandit Problem (arXiv:1111.1797v3, COLT 2012), gave the first logarithmic finite-time bound on the expected regret of TS. This mission formalizes their two-armed result, Theorem 1. A companion mission covers the NNN-armed bound, Theorem 2.

Timeline. Lai and Robbins (1985) proved that every consistent algorithm has regret at least [∑iΔi/D(μi∥μ∗)+o(1)]ln⁡T\big[\sum_i \Delta_i/D(\mu_i\|\mu^*)+o(1)\big]\ln T[∑i​Δi​/D(μi​∥μ∗)+o(1)]lnT. Auer, Cesa-Bianchi and Fischer (2002) gave UCB1 with an O(∑iln⁡T/Δi)O(\sum_i \ln T/\Delta_i)O(∑i​lnT/Δi​) finite-time bound. Agrawal and Goyal (2012) proved O(ln⁡T/Δ+1/Δ3)O(\ln T/\Delta+1/\Delta^3)O(lnT/Δ+1/Δ3) for two-armed TS. Kaufmann, Korda and Munos (ALT 2012) and Agrawal and Goyal (AISTATS 2013) later proved asymptotically optimal bounds for Bernoulli TS.

Setting

There are two arms. Arm i∈{1,2}i\in\{1,2\}i∈{1,2} has a fixed, unknown reward distribution DiD_iDi​ supported in [0,1][0,1][0,1], with mean μi\mu_iμi​. Plays of an arm give i.i.d. rewards, independent of the other arm. Arm 1 is the unique optimal arm, μ1>μ2\mu_1>\mu_2μ1​>μ2​, and Δ=μ1−μ2\Delta=\mu_1-\mu_2Δ=μ1​−μ2​ is the gap.

Thompson Sampling for general stochastic bandits (Algorithm 2 of the paper) keeps, for each arm iii, a success count SiS_iSi​ and a failure count FiF_iFi​, both starting at 000. In each round t=1,2,…t=1,2,\dotst=1,2,… it

  1. samples, independently for each arm, θi(t)∼Beta(Si+1,Fi+1)\theta_i(t)\sim\mathrm{Beta}(S_i+1,F_i+1)θi​(t)∼Beta(Si​+1,Fi​+1);
  2. plays i(t)=arg⁡max⁡iθi(t)i(t)=\arg\max_i\theta_i(t)i(t)=argmaxi​θi​(t) and observes a reward r~t∼Di(t)\tilde r_t\sim D_{i(t)}r~t​∼Di(t)​;
  3. performs a Bernoulli trial with success probability r~t\tilde r_tr~t​, with outcome rt∈{0,1}r_t\in\{0,1\}rt​∈{0,1};
  4. increments Si(t)S_{i(t)}Si(t)​ if rt=1r_t=1rt​=1 and Fi(t)F_{i(t)}Fi(t)​ otherwise.

ki(t)k_i(t)ki​(t) is the number of plays of arm iii before round ttt. The expected regret in time TTT is

E[R(T)]=E[∑t=1T(μ1−μi(t))],\mathbb E[\mathcal R(T)]=\mathbb E\Big[\sum_{t=1}^T(\mu_1-\mu_{i(t)})\Big],E[R(T)]=E[t=1∑T​(μ1​−μi(t)​)],

the expectation being over the rewards and the algorithm's randomness.

The analysis uses the Beta cdf Fα,βbetaF^{beta}_{\alpha,\beta}Fα,βbeta​, the binomial cdf Fn,pBF^B_{n,p}Fn,pB​, and the random variable X(j,s,y)X(j,s,y)X(j,s,y): the number of independent Beta(s+1,j−s+1)\mathrm{Beta}(s+1,j-s+1)Beta(s+1,j−s+1) draws made before one exceeds yyy.

Formalization targets

Goal: Theorem 1 (p. 3)

There is an absolute constant C>0C>0C>0 such that for every two-armed instance with rewards in [0,1][0,1][0,1] and μ1>μ2\mu_1>\mu_2μ1​>μ2​, and every T≥2T\ge 2T≥2,

E[R(T)]≤C(ln⁡TΔ+1Δ3).\mathbb E[\mathcal R(T)]\le C\Big(\frac{\ln T}{\Delta}+\frac1{\Delta^3}\Big).E[R(T)]≤C(ΔlnT​+Δ31​).

The constant is not fixed numerically: the paper states the theorem in O(⋅)O(\cdot)O(⋅) form (footnote 1), and the explicit display it reports on p. 8, 40ln⁡T/Δ+48/Δ3+18Δ40\ln T/\Delta+48/\Delta^3+18\Delta40lnT/Δ+48/Δ3+18Δ, is not the formal claim.

Milestones

  • Fact 1 (p. 12): Fα,βbeta(y)=1−Fα+β−1,yB(α−1)F^{beta}_{\alpha,\beta}(y)=1-F^B_{\alpha+\beta-1,y}(\alpha-1)Fα,βbeta​(y)=1−Fα+β−1,yB​(α−1) for positive integers α,β\alpha,\betaα,β.
  • Lemma 1 (p. 6): E[X(j,s,y)]=1/Fj+1,yB(s)−1\mathbb E[X(j,s,y)]=1/F^B_{j+1,y}(s)-1E[X(j,s,y)]=1/Fj+1,yB​(s)−1.
  • Lemma 6 (p. 13): Hoeffding-type bounds (10)–(11) on binomial cdfs.
  • Fact 2 (p. 13): every median of Binomial(n,p)\mathrm{Binomial}(n,p)Binomial(n,p) is ⌊np⌋\lfloor np\rfloor⌊np⌋ or ⌈np⌉\lceil np\rceil⌈np⌉.
  • Lemma 2 (p. 7): Pr⁡(E2(t))≥1−2/T2\Pr(E_2(t))\ge 1-2/T^2Pr(E2​(t))≥1−2/T2, where E2(t)={θ2(t)≤μ2+Δ/2 or k2(t)<24ln⁡T/Δ2}E_2(t)=\{\theta_2(t)\le\mu_2+\Delta/2\ \text{or}\ k_2(t)<24\ln T/\Delta^2\}E2​(t)={θ2​(t)≤μ2​+Δ/2 or k2​(t)<24lnT/Δ2}.
  • Lemma 3 (p. 7): a three-case bound on E[E[min⁡{X(j,s(j),y),T}∣s(j)]]\mathbb E\big[\mathbb E[\min\{X(j,s(j),y),T\}\mid s(j)]\big]E[E[min{X(j,s(j),y),T}∣s(j)]] for s(j)∼Binomial(j,μ1)s(j)\sim\mathrm{Binomial}(j,\mu_1)s(j)∼Binomial(j,μ1​).
  • Eq. (1) (p. 7): E[k2(T)]≤C(ln⁡T/Δ2+1/Δ4)\mathbb E[k_2(T)]\le C(\ln T/\Delta^2+1/\Delta^4)E[k2​(T)]≤C(lnT/Δ2+1/Δ4).

Significance

The result. Theorem 1 shows that TS, a randomized Bayesian heuristic with no explicit confidence bonus, has regret logarithmic in TTT on every two-armed instance, matching the order in TTT of the Lai–Robbins lower bound. The proof introduced a way to control the optimal arm's waiting time between plays through the Beta–Binomial duality, and later analyses of TS reuse that device.

Formalizing it. The result is proved on paper and has no machine-checked proof that we know of. The platform's existing TS results concern Gaussian TS (Lattimore and Szepesvári, Ch. 36) and Bayesian regret, which are different algorithms or regret notions. A formalization adds a reusable Lean model of Algorithm 2 on [0,1][0,1][0,1]-valued rewards, Beta–Binomial facts (Fact 1, Lemma 1), a binomial-median theorem, and binomial Hoeffding bounds. It also produces a proof with a constant that has been checked, since the printed constants contain an arithmetic slip.

Difficulty

The standard UCB argument does not transfer to TS. For UCB, the optimal arm's index exceeds its mean with high probability however often the arm has been played, because the exploration bonus is deterministic; the analysis then only has to count plays of the suboptimal arm until its own index concentrates, after Θ(ln⁡T/Δ2)\Theta(\ln T/\Delta^2)Θ(lnT/Δ2) plays. Under TS the optimal arm's sample θ1(t)\theta_1(t)θ1​(t) is random and, if the arm has been played rarely or its early rewards were poor, it falls below μ2\mu_2μ2​ with constant probability. The optimal arm may then wait a long, random time between plays, and the length of that wait depends on the arm's posterior, which in turn depends on how long it has waited. Counting plays of the suboptimal arm with a union bound over rounds, under the assumption that the optimal arm is already concentrated, therefore does not work; controlling these waiting times is the central difficulty and is where the 1/Δ31/\Delta^31/Δ3 dependence enters.

Formalization scope

  • Model. The instance is the platform's StochasticBandit 2 (a probability measure on R\mathbb RR per arm, mean banditArmMean), with the hypothesis that each reward law gives mass 111 to [0,1][0,1][0,1]. Lean arm 0 is the paper's arm 1 and Lean arm 1 the paper's arm 2. Lean rounds are indexed from 000.
  • Algorithm. Algorithm 2 is realized on one probability space with three independent i.i.d. tables: Beta draws W(i,t,a,b)∼Beta(a+1,b+1)W(i,t,a,b)\sim\mathrm{Beta}(a+1,b+1)W(i,t,a,b)∼Beta(a+1,b+1), rewards X(i,t)∼DiX(i,t)\sim D_iX(i,t)∼Di​, and uniforms V(i,t)V(i,t)V(i,t). Round ttt uses θi(t)=W(i,t,Si(t),Fi(t))\theta_i(t)=W(i,t,S_i(t),F_i(t))θi​(t)=W(i,t,Si​(t),Fi​(t)), r~t=X(i(t),t)\tilde r_t=X(i(t),t)r~t​=X(i(t),t) and rt=1{V(i(t),t)<r~t}r_t=\mathbf 1\{V(i(t),t)<\tilde r_t\}rt​=1{V(i(t),t)<r~t​}. Ties in the arg max go to the smaller index (a null event).
  • Values. Regret and expectations are lower Lebesgue integrals in [0,∞][0,\infty][0,∞]. X(j,s,y)X(j,s,y)X(j,s,y) is N∪{∞}\mathbb N\cup\{\infty\}N∪{∞}-valued, so Lemma 1 at y=1y=1y=1 reads ∞=∞\infty=\infty∞=∞, as in the paper.
  • O(·). The paper's O(⋅)O(\cdot)O(⋅) (footnote 1: f≤cgf\le cgf≤cg for n≥n0n\ge n_0n≥n0​) is stated with one universal constant C>0C>0C>0, quantified before the instance, the means and the horizon, for all T≥2T\ge 2T≥2. Eq. (1) is stated the same way, without its printed numerals.
  • Not trivial. The goal is about Algorithm 2 itself, with fresh Beta samples, fresh rewards and the Bernoulli coin. A statement about "any policy satisfying Lemma 2's event bound", or one whose constant depends on Δ\DeltaΔ, the reward laws or TTT, would not be Theorem 1.
  • Edge cases. μ1<1\mu_1<1μ1​<1 is assumed only in Lemma 3, where the paper's RRR and DDD require it. It is not a hypothesis of the goal.
  • Infrastructure. A complete proof needs: inverse-transform or order-statistics facts for Beta laws (Fact 1); geometric expectations; Hoeffding's inequality for sums of Bernoulli variables (Mathlib has Hoeffding/Azuma); the binomial median theorem (Jogdeo–Samuels; Kaas–Buhrman); and the coupling from the reward tables to the per-arm i.i.d. output stacks the paper reasons with. Fact 1, Lemma 6 and Fact 2 are reusable beyond this mission. Contributions to any milestone are welcome, and so is a direct proof of the regret bound with an explicit constant.

Selected references

  • S. Agrawal and N. Goyal, Analysis of Thompson Sampling for the Multi-armed Bandit Problem, COLT 2012; arXiv:1111.1797v3. https://arxiv.org/abs/1111.1797
  • W. R. Thompson, On the likelihood that one unknown probability exceeds another in view of the evidence of two samples, Biometrika 25 (1933) 285–294. https://doi.org/10.2307/2332286
  • T. L. Lai and H. Robbins, Asymptotically efficient adaptive allocation rules, Advances in Applied Mathematics 6 (1985) 4–22. https://doi.org/10.1016/0196-8858(85)90002-8
  • P. Auer, N. Cesa-Bianchi and P. Fischer, Finite-time analysis of the multiarmed bandit problem, Machine Learning 47 (2002) 235–256. https://doi.org/10.1023/A:1013689704352
  • K. Jogdeo and S. M. Samuels, Monotone convergence of binomial probabilities and a generalization of Ramanujan's equation, Annals of Mathematical Statistics 39 (1968) 1191–1195. https://doi.org/10.1214/aoms/1177698243
  • R. Kaas and J. M. Buhrman, Mean, median and mode in binomial distributions, Statistica Neerlandica 34 (1980) 13–18. https://doi.org/10.1111/j.1467-9574.1980.tb00681.x
  • O. Chapelle and L. Li, An empirical evaluation of Thompson Sampling, NIPS 2011. https://papers.nips.cc/paper/4321-an-empirical-evaluation-of-thompson-sampling
11 thms2 active usersReviewed
Markov ChainOperations ResearchStochastic Systems·Captain: mikedeng1

Analysis and Algorithms for Service Parts Supply Chains VI: The Shortfall Distribution of Capacity-Limited SystemsTextbook

Motivation

Service parts supply chains are often limited by a capacitated resource, such as a production line or a repair shop, instead of by lead times alone. Once capacity binds, the classical tools for setting stock levels (Palm's theorem and the Poisson distribution of units in resupply) no longer apply, and the quantity that determines how much stock is needed is the shortfall: the amount by which the end-of-period inventory falls below its target because capacity was insufficient. Chapter 8 of Muckstadt, Analysis and Algorithms for Service Parts Supply Chains (Springer 2005, DOI 10.1007/b138879) builds its tactical planning models for capacity-limited systems on the distribution of this random variable, and on a continuous-time repair queue in which item counts are geometric.

The shortfall recursion is the Lindley recursion of queueing theory (Lindley 1952), so its stationary law is the law of the maximum of a random walk with negative drift. The exponential tail of that maximum goes back to Cramér's work on ruin probabilities; for capacitated production–inventory systems it was stated by Glasserman (1997), whose theorem the book quotes as Theorem 11. Glasserman and Tayur (1995) used the shortfall to optimize base-stock levels in multi-echelon capacitated systems, and Roundy and Muckstadt (2000) studied the mass-exponential approximation that the theorem motivates.

Setting

A single item is produced in periods n=1,2,…n = 1, 2, \dotsn=1,2,… of an infinite horizon; at most ccc units can be produced per period. The demand of period nnn is DnD_nDn​; the demands are nonnegative, independent and identically distributed, with generic demand DDD and E[D]<cE[D] < cE[D]<c (the standing assumption of Section 8.1.1).

Under the modified (s−1,s)(s-1, s)(s−1,s) policy with target level sss, the facility observes DnD_nDn​ and produces min⁡{c,s−In−1+Dn}\min\{c, s - I_{n-1} + D_n\}min{c,s−In−1​+Dn​} units, where InI_nIn​ is the end-of-period net inventory and I0=sI_0 = sI0​=s. The shortfall Vn=s−InV_n = s - I_nVn​=s−In​ satisfies V0=0V_0 = 0V0​=0 and

Vn=[Vn−1+Dn−c]+.(8.1)V_n = \left[V_{n-1} + D_n - c\right]^+ . \tag{8.1}Vn​=[Vn−1​+Dn​−c]+.(8.1)

With the random walk Sn=∑k=1n(Dk−c)S_n = \sum_{k=1}^{n} (D_k - c)Sn​=∑k=1n​(Dk​−c) (S0=0S_0 = 0S0​=0), the stationary shortfall is

V=sup⁡n≥0Sn.V = \sup_{n \ge 0} S_n .V=n≥0sup​Sn​.

A law on R\mathbb RR is lattice if it is concentrated on a progression a+dZa + d\mathbb Za+dZ with d>0d > 0d>0.

In the discrete case (ccc and DDD integer valued) (Vn)(V_n)(Vn​) is a Markov chain on {0,1,2,… }\{0, 1, 2, \dots\}{0,1,2,…} with transition probabilities pijp_{ij}pij​ (p. 185). In the repair model of Section 8.3.1, reparable units of item iii arrive at rate λi\lambda_iλi​, λ=∑iλi\lambda = \sum_i \lambda_iλ=∑i​λi​, a single exponential server repairs at rate μ>λ\mu > \lambdaμ>λ, NNN is the number of units in repair and NiN_iNi​ the number of item-iii units, and ηi=λi/(μ−λ+λi)\eta_i = \lambda_i/(\mu - \lambda + \lambda_i)ηi​=λi​/(μ−λ+λi​).

Formalization targets

Goal: Theorem 11, corrected (p. 191)

Assume E[eαD]<∞E[e^{\alpha D}] < \inftyE[eαD]<∞ for all α<δ\alpha < \deltaα<δ, with δ>0\delta > 0δ>0; P[D>c]>0P[D > c] > 0P[D>c]>0; the law of DDD is non-lattice; and E[e−α(c−D)]=1E[e^{-\alpha(c-D)}] = 1E[e−α(c−D)]=1 has a root in (0,δ)(0, \delta)(0,δ). Then there are β>0\beta > 0β>0 and α>0\alpha > 0α>0 with

P{V>v}βe−αv→1(v→∞),α the unique positive root of E[e−α(c−D)]=1.\frac{P\{V > v\}}{\beta e^{-\alpha v}} \to 1 \quad (v \to \infty), \qquad \alpha \text{ the unique positive root of } E\left[e^{-\alpha(c - D)}\right] = 1 .βe−αvP{V>v}​→1(v→∞),α the unique positive root of E[e−α(c−D)]=1.

The constant β\betaβ is left unspecified, as in the book.

Milestones, in attack order

  1. Eq. (8.1): under the modified policy, s−In=Vns - I_n = V_ns−In​=Vn​ for every nnn, independently of sss.
  2. Section 8.1.1: V<∞V < \inftyV<∞ almost surely, P{Vn>v}→P{V>v}P\{V_n > v\} \to P\{V > v\}P{Vn​>v}→P{V>v} for every vvv, and the law of VVV is stationary for (8.1).
  3. Eq. (8.2): for v>0v > 0v>0, P{Vn>v}=P{Dn>v+c}+ED[1(d≤v+c) P{Vn−1>v+c−d}]P\{V_n > v\} = P\{D_n > v + c\} + E_D[1(d \le v + c)\, P\{V_{n-1} > v + c - d\}]P{Vn​>v}=P{Dn​>v+c}+ED​[1(d≤v+c)P{Vn−1​>v+c−d}].
  4. Theorem 11, second sentence: E[e−α(c−D)]=1E[e^{-\alpha(c-D)}] = 1E[e−α(c−D)]=1 has at most one positive root.
  5. Section 8.1.2: with integer demand, (Vn)(V_n)(Vn​) is a Markov chain with transition probabilities pijp_{ij}pij​.
  6. Section 8.1.2: πi=lim⁡nP{Vn=i}\pi_i = \lim_n P\{V_n = i\}πi​=limn​P{Vn​=i} exists and solves πP=π\pi\mathcal P = \piπP=π, ∑iπi=1\sum_i \pi_i = 1∑i​πi​=1, πi≥0\pi_i \ge 0πi​≥0.
  7. Section 8.3.1: if NNN is geometric with parameter λ/μ\lambda/\muλ/μ and NiN_iNi​ given N=jN = jN=j is binomial(j,λi/λ)(j, \lambda_i/\lambda)(j,λi​/λ), then P[Ni=j]=(1−ηi)ηijP[N_i = j] = (1 - \eta_i)\eta_i^jP[Ni​=j]=(1−ηi​)ηij​.
  8. Section 8.3.1: ∑j>spi(j)=ηis+1\sum_{j > s} p_i(j) = \eta_i^{s+1}∑j>s​pi​(j)=ηis+1​, and the smallest cost-minimising stock level is the smallest sss with ηis+1≤hi/(hi+b)\eta_i^{s+1} \le h_i/(h_i + b)ηis+1​≤hi​/(hi​+b).

Significance

The exponential tail is the justification the book gives for approximating the shortfall by a mass-exponential law (an atom at zero plus an exponential tail), from which target stock levels and fill rates are computed in closed form. The decay rate α\alphaα depends only on the demand law and the capacity, so the theorem also says how the stock needed for a given service level grows as utilization approaches one. The discrete-chain milestones justify the exact computation of the shortfall distribution behind the book's Table 8.1 and Figures 8.3–8.8. The geometric law of NiN_iNi​ reduces the multi-item repair problem to independent newsvendor problems with an explicit solution.

The asymptotics of the random-walk maximum are proved in the literature (Cramér–Lundberg theory, Feller Vol. II, XII.5; Asmussen, Applied Probability and Queues, XIII.5); no machine-checked proof is known to exist. Mathlib has neither the Lindley recursion, nor ladder-height decompositions, nor the key renewal theorem for non-lattice laws. The printed Theorem 11 is not correct as stated (see Formalization scope), so the mission also records a corrected statement.

Difficulty

The central step of the goal is the passage from the random walk to an exact asymptotic. An exponential change of measure (Esscher tilt) with the root α\alphaα turns P{V>v}P\{V > v\}P{V>v} into an expectation under a law with positive drift, but it only yields the upper bound P{V>v}≤e−αvP\{V > v\} \le e^{-\alpha v}P{V>v}≤e−αv (Lundberg's inequality); it does not show that eαvP{V>v}e^{\alpha v}P\{V > v\}eαvP{V>v} converges, nor that the limit is positive. Convergence needs a renewal theorem for the overshoot of the tilted walk, which fails for lattice laws. That is why the non-lattice hypothesis cannot be dropped. For the milestones, the existence of the stationary law needs the reversal argument that identifies the law of VnV_nVn​ with that of max⁡k≤nSk\max_{k \le n} S_kmaxk≤n​Sk​, plus the strong law of large numbers to show V<∞V < \inftyV<∞ from E[D]<cE[D] < cE[D]<c.

Formalization scope

  • Model. Demands are real, nonnegative, measurable, i.i.d. (iIndepFun plus IdentDistrib with D1D_1D1​), integrable, with E[D]<cE[D] < cE[D]<c; these are fields of ShortfallModel. Periods are numbered from 111 as in the book (demand 0 is an unused i.i.d. copy). The discrete case is a separate structure with N\mathbb NN-valued demand and capacity.
  • Stationary shortfall. The book's "stationary distribution ... Let VVV represent this random variable" is pinned to V=sup⁡n≥0SnV = \sup_{n \ge 0} S_nV=supn≥0​Sn​, taken in [0,∞][0, \infty][0,∞] and converted to a real number; milestone 2 proves that it is the limit law of VnV_nVn​ from V0=0V_0 = 0V0​=0 and a stationary law of (8.1). The discrete πi\pi_iπi​ is pinned to lim⁡nP{Vn=i}\lim_n P\{V_n = i\}limn​P{Vn​=i}.
  • Corrections to Theorem 11. The printed theorem is false. For integer demand P{V>v}P\{V > v\}P{V>v} is a step function, and no βe−αv\beta e^{-\alpha v}βe−αv is asymptotic to it. If E[eαD]E[e^{\alpha D}]E[eαD] is finite only for α<δ\alpha < \deltaα<δ, the equation E[e−α(c−D)]=1E[e^{-\alpha(c-D)}] = 1E[e−α(c−D)]=1 may have no root in (0,δ)(0,\delta)(0,δ). The goal therefore adds two labelled hypotheses: a non-lattice demand law, and a root in (0,δ)(0, \delta)(0,δ). The mass-exponential demand of Section 8.1.3 (an atom at 000 plus a density) is non-lattice. The approximation β≈e−2(.583)(c−E(D))/σ\beta \approx e^{-2(.583)(c-E(D))/\sigma}β≈e−2(.583)(c−E(D))/σ is not stated.
  • Repair model. The M/M/1 queue is not built. The geometric law of NNN (asserted on p. 202) and the binomial split of NNN (quoted from Chapter 3) enter milestone 7 as hypotheses, exactly as the page's proof uses them. The stability condition λ<μ\lambda < \muλ<μ, not written on the page, is a hypothesis. "The optimal sis_isi​" is read as the smallest minimiser of the cost.
  • Ruled out. Stating Theorem 11 with α\alphaα or β\betaβ allowed to depend on vvv, with β=0\beta = 0β=0 (the ratio would be a division by zero, which Lean evaluates to 000), or for a VVV postulated to have an exponential tail proves nothing. Here β,α\beta, \alphaβ,α are quantified before vvv, both are asserted positive, and VVV is constructed from the demands.
  • Not formalized. The mass-exponential approximations (8.3)–(8.4), the Roundy–Muckstadt refinement, the fill-rate formula η(s)\eta(s)η(s) (a definition, whose steady-state identity needs uniform integrability the book does not discuss), the random-capacity chain on p. 186, and the monotonicity of sis_isi​ in μ\muμ.
  • Reusable infrastructure. Welcome: the Lindley recursion and its reversal identity, the Loynes existence theorem, Lundberg's inequality, and a non-lattice renewal theorem. All of these are needed well beyond this mission, in queueing (GI/G/1 waiting times) and ruin theory.

Selected references

  • J. A. Muckstadt, Analysis and Algorithms for Service Parts Supply Chains, Springer, 2005, Chapter 8. https://doi.org/10.1007/b138879
  • P. Glasserman, Bounds and asymptotics for planning critical safety stocks, Operations Research 45(2), 244–257, 1997. https://doi.org/10.1287/opre.45.2.244
  • P. Glasserman and S. Tayur, Sensitivity analysis for base-stock levels in multiechelon production-inventory systems, Management Science 41(2), 263–281, 1995 (the book's reference [97]). https://doi.org/10.1287/mnsc.41.2.263
  • R. O. Roundy and J. A. Muckstadt, Heuristic computation of periodic-review base stock inventory policies, Management Science 46(1), 104–109, 2000. https://doi.org/10.1287/mnsc.46.1.104.15131
  • D. V. Lindley, The theory of queues with a single server, Mathematical Proceedings of the Cambridge Philosophical Society 48(2), 277–289, 1952. https://doi.org/10.1017/S0305004100027638
  • W. Feller, An Introduction to Probability Theory and Its Applications, Vol. II, 2nd ed., Wiley, 1971, Chapter XII.
  • S. Asmussen, Applied Probability and Queues, 2nd ed., Springer, 2003, Chapter XIII. https://doi.org/10.1007/b97236
12 thms2 active usersReviewed
Operations ResearchStochastic Systems·Captain: mikedeng1

Analysis and Algorithms for Service Parts Supply Chains III: Palm's Theorem for (s–1, s) PoliciesTextbook

Motivation

Service parts (spare engines, avionics modules, repairable components) are usually managed one unit at a time: whenever a customer order removes a unit from stock, a replacement is ordered at once, from a repair shop or an outside supplier. This is the (s–1, s) policy, under which the inventory position (on hand plus on order minus backorders) stays constant at the stock level sss. Every performance measure of such a system (fill rate, expected backorders, availability) is a function of one random variable: the number of units in resupply, i.e. ordered but not yet returned. Chapter 3 of Muckstadt, Analysis and Algorithms for Service Parts Supply Chains (Springer 2005) computes its distribution, and the rest of the book (the METRIC-type multi-echelon models of Chapters 4 and 5, the stock-level optimization of Section 3.4) is built on that computation.

Timeline. C. Palm (1938) showed, in the setting of telephone traffic, that in an infinite-server system with Poisson arrivals the number of busy servers has, in steady state, a Poisson law whose mean is the arrival rate times the mean service time, whatever the service-time distribution. Feeney and Sherbrooke (1966) carried the result to (s–1, s) inventory systems with compound Poisson demand, and treated the lost-sales case. Sherbrooke's METRIC model (1968) made the Poisson law of units in resupply the basis of multi-echelon spare-parts planning.

Setting

A single item is stocked at one location. Customer orders arrive at epochs T0<T1<⋯T_0 < T_1 < \cdotsT0​<T1​<⋯ of a Poisson process with rate λ>0\lambda > 0λ>0, started empty at time 000: the interarrival times AkA_kAk​ are independent and exponential with rate λ\lambdaλ, and Tk=A0+⋯+AkT_k = A_0 + \cdots + A_kTk​=A0​+⋯+Ak​. The kkk-th order triggers a resupply order with resupply time Lk≥0L_k \ge 0Lk​≥0. The resupply times are independent and identically distributed, independent of the arrival process, with a density ggg, distribution function G(u)=P[L≤u]G(u) = P[L \le u]G(u)=P[L≤u] and finite mean

τˉ=E[L]=∫0∞[1−G(u)] du.\bar\tau = E[L] = \int_0^\infty [1 - G(u)]\,du .τˉ=E[L]=∫0∞​[1−G(u)]du.

With backorders allowed, the number of units in resupply at time ttt is

X(t)=#{k:Tk≤t<Tk+Lk},X(t) = \#\{k : T_k \le t < T_k + L_k\},X(t)=#{k:Tk​≤t<Tk​+Lk​},

and N(t)=#{k:Tk≤t}N(t) = \#\{k : T_k \le t\}N(t)=#{k:Tk​≤t} counts the orders placed in [0,t][0,t][0,t]. On-hand stock and backorders at time ttt are (s−X(t))+(s - X(t))^+(s−X(t))+ and (X(t)−s)+(X(t) - s)^+(X(t)−s)+.

In the compound Poisson version the kkk-th order asks for Xk≥1X_k \ge 1Xk​≥1 units, the sizes are i.i.d. with uj=P[Xk=j]u_j = P[X_k = j]uj​=P[Xk​=j], independent of arrivals and resupply times, and all units of one order share its resupply time LkL_kLk​. The units in resupply are Y(t)=∑k:Tk≤t<Tk+LkXkY(t) = \sum_{k : T_k \le t < T_k + L_k} X_kY(t)=∑k:Tk​≤t<Tk​+Lk​​Xk​. Writing un(y)u^{(y)}_nun(y)​ for the probability that yyy orders ask for nnn units in total, the compound Poisson probabilities with parameter μ\muμ are

p(n∣μ)=∑y=0nμye−μy! un(y).p(n \mid \mu) = \sum_{y=0}^{n} \frac{\mu^y e^{-\mu}}{y!}\,u^{(y)}_n .p(n∣μ)=y=0∑n​y!μye−μ​un(y)​.

In the lost-sales version an order that finds no stock on hand is lost, so at most sss units are ever in resupply.

Formalization targets

Goal: Palm's theorem (Theorem 6, p. 39)

For every x≥0x \ge 0x≥0,

lim⁡t→∞P[X(t)=x]=e−λτˉ(λτˉ)xx!.\lim_{t\to\infty} P[X(t) = x] = e^{-\lambda\bar\tau}\frac{(\lambda\bar\tau)^x}{x!}.t→∞lim​P[X(t)=x]=e−λτˉx!(λτˉ)x​.

The resupply-time law enters only through its mean. This is the statement the book's proof establishes and every later chapter uses.

The proof's milestones (pp. 38–41)

  1. Eq. (3.5): P[N(t)=n]=e−λt(λt)n/n!P[N(t) = n] = e^{-\lambda t}(\lambda t)^n/n!P[N(t)=n]=e−λt(λt)n/n!.
  2. Eq. (3.3): given N(t)=nN(t) = nN(t)=n, the epochs (T0,…,Tn−1)(T_0, \dots, T_{n-1})(T0​,…,Tn−1​) have density n!/tnn!/t^nn!/tn on 0<t1<⋯<tn<t0 < t_1 < \cdots < t_n < t0<t1​<⋯<tn​<t.
  3. Eq. (3.7): given N(t)=nN(t) = nN(t)=n, X(t)X(t)X(t) is binomial with parameters nnn and p=1t∫0t[1−G(u)] dup = \frac1t\int_0^t[1-G(u)]\,dup=t1​∫0t​[1−G(u)]du.
  4. Eq. (3.8): for every t>0t > 0t>0, X(t)X(t)X(t) is Poisson with mean λ∫0t[1−G(u)] du\lambda\int_0^t[1-G(u)]\,duλ∫0t​[1−G(u)]du.
  5. Eq. (3.10): ∫0t[1−G(u)] du→τˉ\int_0^t[1-G(u)]\,du \to \bar\tau∫0t​[1−G(u)]du→τˉ.

Extensions in Section 3.1

  1. Theorem 7 (pp. 43–44): with compound Poisson demand, lim⁡t→∞P[Y(t)=n]=p(n∣λτˉ)\lim_{t\to\infty}P[Y(t) = n] = p(n \mid \lambda\bar\tau)limt→∞​P[Y(t)=n]=p(n∣λτˉ).
  2. Theorem 8 (p. 44): in the lost-sales system with exponential resupply times of rate β\betaβ, the probability vectors solving the balance equations are exactly the truncated Poisson law πx∝(λ/β)x/x!\pi_x \propto (\lambda/\beta)^x/x!πx​∝(λ/β)x/x!, 0≤x≤s0 \le x \le s0≤x≤s.
  3. Theorem 9 (pp. 46–47): for a due-date delay T≥0T \ge 0T≥0, the units in resupply that have been there for at least TTT satisfy lim⁡t→∞P[YT(t)=n]=p(n∣λτˉα)\lim_{t\to\infty}P[Y_T(t) = n] = p(n \mid \lambda\bar\tau\alpha)limt→∞​P[YT​(t)=n]=p(n∣λτˉα) with α=1τˉ∫T∞[1−G(t)] dt\alpha = \frac1{\bar\tau}\int_T^\infty[1-G(t)]\,dtα=τˉ1​∫T∞​[1−G(t)]dt.

Significance

The result. Palm's theorem turns an infinite-dimensional object (the whole resupply-time distribution) into one number, τˉ\bar\tauτˉ. This insensitivity is what makes spare-parts planning computable: the expected backorders at stock level sss are ∑x>s(x−s) p(x∣λτˉ)\sum_{x > s}(x - s)\,p(x \mid \lambda\bar\tau)∑x>s​(x−s)p(x∣λτˉ), the fill rate is P[X≤s−1]P[X \le s - 1]P[X≤s−1], and both can be optimized over sss with only the demand rate and mean repair time as data. Theorem 7 extends this to batch demand, Theorem 9 to systems allowed a response time, and Theorem 8 gives the exact law when shortages are lost instead of backordered.

Formalizing it. All of these results are classical and proved. None of them is formalized on the platform, and Mathlib has Poisson and exponential distributions but no Poisson process, no thinning theorem and no infinite-server queue. The mission produces a Poisson arrival stream built from i.i.d. exponential gaps, the conditional-uniformity property of its epochs, independent thinning, and the M/G/∞ transient law, all reusable in queueing and inventory missions.

Difficulty

The algebra of the proof (summing the binomial against the Poisson law of N(t)N(t)N(t)) is short. The work is in the probabilistic step the book treats in a sentence: that, given N(t)=nN(t) = nN(t)=n, the nnn orders behave like independent uniform epochs, each of which independently is still in resupply at time ttt with the same probability ppp. This needs the joint law of the partial sums of exponential variables (Eq. (3.3)), and then a symmetrization argument, since the epochs are ordered while the resupply times are attached to order indices. The naive route of computing P[X(t)=x]P[X(t) = x]P[X(t)=x] by conditioning on individual epochs does not go through without that exchangeability step. The limit t→∞t \to \inftyt→∞ is then elementary; stating a stationary version directly is not a substitute, since the book's "steady state" is exactly this limit.

Formalization scope

The model is a structure on a probability space (Ω,P)(\Omega, P)(Ω,P): exponential interarrival times with rate λ>0\lambda > 0λ>0, nonnegative resupply times with a density and an integrable first coordinate, and mutual independence of the whole family. Orders are indexed from 000, so the book's X1,…,XnX_1, \dots, X_nX1​,…,Xn​ are T0,…,Tn−1T_0, \dots, T_{n-1}T0​,…,Tn−1​. Counts are cardinalities of sets of order indices, with value 000 on the probability-zero event where infinitely many orders fall in a bounded interval.

Commitments and pinnings:

  • "Steady state probability" (Theorems 6, 7, 9) is lim⁡t→∞P[⋅(t)=x]\lim_{t\to\infty}P[\cdot(t) = x]limt→∞​P[⋅(t)=x] for the system empty at time 000, which is what the proofs compute via (3.8)–(3.11).
  • Independence of resupply times from arrivals is not written in Theorem 6 but is used on p. 40; it is part of the model.
  • The stock level sss does not enter the backorder model; it matters only in Theorem 8.
  • Theorem 8 is stated algebraically: a vector on {0,…,s}\{0, \dots, s\}{0,…,s} solves the balance equations (3.26), (3.25) for 0<j<s0 < j < s0<j<s and (3.32), and sums to one, if and only if it is the truncated Poisson law. The book obtains these equations by letting t→∞t \to \inftyt→∞ in the forward equations under the unproved assumption Pj′(t)→0P_j'(t) \to 0Pj′​(t)→0. The book writes (3.25) "for 0≤j≤s0 \le j \le s0≤j≤s", which at j=sj = sj=s contradicts its own (3.32); the boundary equation (3.32) is used. The sentence on p. 46 extending Theorem 8 to arbitrary resupply densities is asserted without proof and is not stated.
  • Theorem 7 identifies the limit law by its probabilities (3.22)–(3.23); its mean λτˉuˉ\lambda\bar\tau\bar uλτˉuˉ is a property of that law. Theorem 9 is stated for compound demand as the book states it, although the book's proof covers only the Poisson case.

A model in which X(t)X(t)X(t) is postulated through its law, or in which resupply times may depend on the arrival epochs, makes the goal empty or false; here X(t)X(t)X(t) is computed from the primitive arrival and resupply times, whose joint law is fully specified.

Welcome contributions: a general Poisson-process library (construction from exponential gaps, Poisson marginals, order-statistics property), independent thinning, and proofs of the milestones in the listed order.

Selected references

  • J. A. Muckstadt, Analysis and Algorithms for Service Parts Supply Chains, Springer Series in Operations Research and Financial Engineering, 2005, Chapter 3. https://doi.org/10.1007/b138879
  • C. Palm, "Analysis of the Erlang traffic formula for busy-signal arrangements", Ericsson Technics 5, 1938, 39–58.
  • G. J. Feeney and C. C. Sherbrooke, "The (s–1, s) inventory policy under compound Poisson demand", Management Science 12(5), 1966, 391–411. https://doi.org/10.1287/mnsc.12.5.391
  • C. C. Sherbrooke, "METRIC: A multi-echelon technique for recoverable item control", Operations Research 16(1), 1968, 122–141. https://doi.org/10.1287/opre.16.1.122
  • S. M. Ross, Stochastic Processes, 2nd ed., Wiley, 1996, Section 2.3 (conditional distribution of arrival times) and Section 2.4 (the M/G/∞ queue).
12 thms2 active usersReviewed
PreviousPage 3 of 12Next

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me