Prove2Me
Navigate
DiscoverFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

Machine Learning

273 missions · 183 completed

The science of systems that learn from data and experience. Its scope runs from the statistical and mathematical foundations of learning, including generalization, expressivity, and computational limits, through the design of learning algorithms, deep learning, reinforcement learning, and probabilistic methods, to the empirical study of large models and the trustworthiness, interpretability, and societal impact of learned systems.

Missions

Open90Completed183All273
Linear algebraMarkov ChainReinforcement Learning·Captain: mikedeng1

Reinforcement Learning: An Introduction VIII: The TD Fixed Point of Linear Semi-gradient TD(0) and Its Error BoundTextbook

Motivation

Reinforcement learning methods estimate the value function vπv_\pivπ​ of a policy π\piπ: the expected discounted sum of future rewards from each state. When the state space is large, vπv_\pivπ​ cannot be stored as a table and is approximated by a parametrized function. The most studied case is linear function approximation, where each state sss carries a feature vector x(s)∈Rd\mathbf x(s) \in \mathbb R^dx(s)∈Rd and the estimate is v^(s,w)=w⊤x(s)\hat v(s, \mathbf w) = \mathbf w^\top \mathbf x(s)v^(s,w)=w⊤x(s). Temporal-difference learning with this approximation, linear semi-gradient TD(0), is one of the basic algorithms of the field, and Chapter 9 of Sutton and Barto's Reinforcement Learning: An Introduction (2nd ed., MIT Press, 2018) presents its analysis: where the algorithm can converge, why that point exists, and how good it is.

The history is short. Sutton (1988, doi:10.1007/BF00115009) introduced TD learning and showed positive definiteness of the matrix governing its expected update. Dayan (1992, doi:10.1007/BF00992701) extended convergence to TD(λ). Tsitsiklis and Van Roy (1997, doi:10.1109/9.580874) proved convergence with probability one for linear TD(λ) under on-policy sampling and bounded the error of the limit. Bradtke and Barto (1996) introduced least-squares TD (LSTD), which computes the same limit directly.

Setting

A finite Markov decision process has finite sets of states S\mathcal SS, actions A\mathcal AA and rewards R⊂R\mathcal R \subset \mathbb RR⊂R, and dynamics p(s′,r∣s,a)p(s', r \mid s, a)p(s′,r∣s,a), the probability of next state s′s's′ and reward rrr after action aaa in state sss. A policy π(a∣s)\pi(a \mid s)π(a∣s) is a probability distribution over actions for each state. It induces a Markov chain on states with transition matrix P\mathbf PP, P(s,s′)=p(s′∣s)=∑aπ(a∣s)∑rp(s′,r∣s,a)\mathbf P(s, s') = p(s' \mid s) = \sum_a \pi(a \mid s) \sum_r p(s', r \mid s, a)P(s,s′)=p(s′∣s)=∑a​π(a∣s)∑r​p(s′,r∣s,a), and expected one-step reward rπ(s)r_\pi(s)rπ​(s). For a discount rate 0≤γ<10 \le \gamma < 10≤γ<1, the true value is vπ(s)=∑k≥0γk(Pkrπ)(s)v_\pi(s) = \sum_{k \ge 0} \gamma^k (\mathbf P^k r_\pi)(s)vπ​(s)=∑k≥0​γk(Pkrπ​)(s), the expected discounted return.

A state distribution μ\muμ is stationary if μ⊤P=μ⊤\mu^\top \mathbf P = \mu^\topμ⊤P=μ⊤; write D=diag(μ)\mathbf D = \mathrm{diag}(\mu)D=diag(μ). The feature matrix X\mathbf XX is the ∣S∣×d|\mathcal S| \times d∣S∣×d matrix with rows x(s)\mathbf x(s)x(s). The mean square value error of a weight vector is

VE‾(w)=∑sμ(s) [vπ(s)−w⊤x(s)]2.\overline{\mathrm{VE}}(\mathbf w) = \sum_{s} \mu(s)\,[v_\pi(s) - \mathbf w^\top \mathbf x(s)]^2 .VE(w)=s∑​μ(s)[vπ​(s)−w⊤x(s)]2.

Linear semi-gradient TD(0) updates wt+1=wt+α(Rt+1+γwt⊤xt+1−wt⊤xt)xt\mathbf w_{t+1} = \mathbf w_t + \alpha(R_{t+1} + \gamma \mathbf w_t^\top \mathbf x_{t+1} - \mathbf w_t^\top \mathbf x_t)\mathbf x_twt+1​=wt​+α(Rt+1​+γwt⊤​xt+1​−wt⊤​xt​)xt​. In steady state its expected update involves

b=E[Rt+1xt],A=E[xt(xt−γxt+1)⊤],\mathbf b = \mathbb E[R_{t+1}\mathbf x_t], \qquad \mathbf A = \mathbb E[\mathbf x_t(\mathbf x_t - \gamma \mathbf x_{t+1})^\top],b=E[Rt+1​xt​],A=E[xt​(xt​−γxt+1​)⊤],

and the TD fixed point is wTD=A−1b\mathbf w_{\mathrm{TD}} = \mathbf A^{-1}\mathbf bwTD​=A−1b. A real square matrix MMM, not necessarily symmetric, is positive definite if y⊤My>0y^\top M y > 0y⊤My>0 for every y≠0y \ne 0y=0. The key matrix is D(I−γP)\mathbf D(\mathbf I - \gamma\mathbf P)D(I−γP).

Formalization targets

Goal: the TD fixed point exists and its error bound (9.12), (9.14)

Under the hypotheses above, with every μ(s)>0\mu(s) > 0μ(s)>0 and linearly independent feature columns, A\mathbf AA is invertible, b=AwTD\mathbf b = \mathbf A \mathbf w_{\mathrm{TD}}b=AwTD​, and

VE‾(wTD)≤11−γmin⁡wVE‾(w).\overline{\mathrm{VE}}(\mathbf w_{\mathrm{TD}}) \le \frac{1}{1-\gamma}\min_{\mathbf w} \overline{\mathrm{VE}}(\mathbf w).VE(wTD​)≤1−γ1​wmin​VE(w).

Milestones

  1. The expected update (9.13): E[wt+1∣wt]=(I−αA)wt+αb\mathbb E[\mathbf w_{t+1} \mid \mathbf w_t] = (\mathbf I - \alpha \mathbf A)\mathbf w_t + \alpha \mathbf bE[wt+1​∣wt​]=(I−αA)wt​+αb.
  2. The matrix form A=X⊤D(I−γP)X\mathbf A = \mathbf X^\top \mathbf D(\mathbf I - \gamma \mathbf P)\mathbf XA=X⊤D(I−γP)X.
  3. The criterion of Sutton (1988): positive diagonal, nonpositive off-diagonal entries, positive row sums and nonnegative column sums give positive definiteness.
  4. The column sums of the key matrix, 1⊤D(I−γP)=(1−γ)μ⊤\mathbf 1^\top \mathbf D(\mathbf I - \gamma \mathbf P) = (1-\gamma)\mu^\top1⊤D(I−γP)=(1−γ)μ⊤.
  5. The key matrix and A\mathbf AA are positive definite.
  6. A positive definite A\mathbf AA is invertible and A−1b\mathbf A^{-1}\mathbf bA−1b is the unique solution of b=Aw\mathbf b = \mathbf A \mathbf wb=Aw (9.12).
  7. The Sherman–Morrison update (9.22) of the LSTD inverse A^t−1\hat{\mathbf A}_t^{-1}A^t−1​.

Significance

Positive definiteness of A\mathbf AA is the reason on-policy linear TD(0) is stable: it makes the expected iteration contract toward the fixed point for small step sizes, and it guarantees that the fixed point exists and is unique. The error bound (9.14) quantifies the price of bootstrapping: the limit of TD can be worse than the best linear approximation, but by at most the factor 1/(1−γ)1/(1-\gamma)1/(1−γ). The same objects A\mathbf AA, b\mathbf bb and the key matrix reappear in LSTD, in the analysis of off-policy divergence (Chapter 11 of the book, where D\mathbf DD is no longer the stationary distribution of P\mathbf PP and positive definiteness fails), and in gradient-TD methods.

All results here are known. The book gives the positive definiteness argument in a box and cites (9.14) without proof. None of them has a machine-checked proof on the platform; the general Woodbury identity (FamousTheorems.woodbury_identity) is available, and (9.22) is its rank-one case written for the LSTD recursion. The mission produces a formal account of the finite-state theory of linear TD(0), with every hypothesis the book leaves implicit stated.

Difficulty

The key matrix D(I−γP)\mathbf D(\mathbf I - \gamma \mathbf P)D(I−γP) is not symmetric, so the usual tools for symmetric positive definite matrices do not apply directly, and A\mathbf AA is positive definite only because of the specific interplay between D\mathbf DD and P\mathbf PP: if μ\muμ is replaced by a non-stationary distribution the claim is false (this is the off-policy counterexample of Chapter 11). The error bound (9.14) is not a consequence of positive definiteness alone. The TD fixed point is not the minimizer of VE‾\overline{\mathrm{VE}}VE, and VE‾(wTD)\overline{\mathrm{VE}}(\mathbf w_{\mathrm{TD}})VE(wTD​) has to be compared with the error of the μ\muμ-weighted projection of vπv_\pivπ​, which requires controlling P\mathbf PP in the μ\muμ-weighted norm. The book gives no argument for this step.

Formalization scope

The Lean development lives in the namespace SuttonBartoRL.LinearTD. The MDP has four-argument dynamics p(s′,r∣s,a)p(s', r \mid s, a)p(s′,r∣s,a) with a finite reward set and one action set for all states; policies are stochastic. vπv_\pivπ​ is defined from expected discounted returns as the series ∑kγkPkrπ\sum_k \gamma^k \mathbf P^k r_\pi∑k​γkPkrπ​, never from a Bellman equation or from wTD\mathbf w_{\mathrm{TD}}wTD​. A\mathbf AA and b\mathbf bb are defined as the book's steady-state expectations (9.11), as finite sums over μ\muμ, π\piπ and ppp; the matrix form is a milestone, not a definition. Features are a matrix Matrix S (Fin d) ℝ with rows x(s)\mathbf x(s)x(s); linear independence of its columns is LinearIndependent ℝ Xᵀ. Positive definiteness is a custom predicate ∀y≠0, 0<y⊤My\forall y \ne 0,\ 0 < y^\top M y∀y=0, 0<y⊤My, not Mathlib's Matrix.PosDef, which requires symmetry. The minimum in (9.14) is expressed by quantifying over every w\mathbf ww. The matrix inverse is Mathlib's, which is zero on singular matrices; the goal therefore asserts invertibility of A\mathbf AA explicitly.

Hypotheses the book leaves implicit and the statements make explicit: 0≤γ<10 \le \gamma < 10≤γ<1 (the continuing case); μ\muμ a stationary distribution of the chain induced by π\piπ with μ(s)>0\mu(s) > 0μ(s)>0 for every sss (otherwise the key matrix is only positive semidefinite); linearly independent feature columns (the book's "degenerate cases", p. 205). The box calls the off-diagonal entries of the key matrix "negative"; they are zero wherever p(s′∣s)=0p(s' \mid s) = 0p(s′∣s)=0, so the criterion is stated with nonpositive entries. The book's sentence that εI\varepsilon\mathbf IεI "ensures that A^t\hat{\mathbf A}_tA^t​ is always invertible" (p. 229) is false in general, because the summands xk(xk−γxk+1)⊤\mathbf x_k(\mathbf x_k - \gamma \mathbf x_{k+1})^\topxk​(xk​−γxk+1​)⊤ are not positive semidefinite: with d=1d = 1d=1, ε=1/10\varepsilon = 1/10ε=1/10, γ=1/2\gamma = 1/2γ=1/2, x0=1\mathbf x_0 = 1x0​=1, x1=11/5\mathbf x_1 = 11/5x1​=11/5 one gets A^1=0\hat{\mathbf A}_1 = 0A^1​=0. It is not stated; (9.22) carries invertibility of A^t−1\hat{\mathbf A}_{t-1}A^t−1​ and a nonzero denominator as hypotheses.

A statement in which vπv_\pivπ​ is defined as the solution of the projected equation, or in which A\mathbf AA is assumed invertible or positive definite, would make the goal trivial or empty; neither is done. Convergence of the stochastic algorithm with probability one is not stated, since the book says it needs conditions and a step-size schedule it does not give. The bound for the episodic case and for other bootstrapping methods (p. 208) is stated only by reference in the book and is not a target.

Useful infrastructure: the μ\muμ-weighted inner product and orthogonal projection onto the column space of X\mathbf XX, the non-expansiveness of a stochastic matrix in the norm of its stationary distribution, and the positive definiteness criterion for non-symmetric matrices. All of these are reusable in the off-policy and average-reward chapters of the book. Contributions of these lemmas, and of alternative proofs of the milestones, are welcome.

Selected references

  • Richard S. Sutton and Andrew G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018, ISBN 9780262039246, §§9.2, 9.4, 9.8.
  • Richard S. Sutton, Learning to predict by the methods of temporal differences, Machine Learning 3, 1988. doi:10.1007/BF00115009
  • John N. Tsitsiklis and Benjamin Van Roy, An analysis of temporal-difference learning with function approximation, IEEE Transactions on Automatic Control 42(5), 1997. doi:10.1109/9.580874
  • Steven J. Bradtke and Andrew G. Barto, Linear least-squares algorithms for temporal difference learning, Machine Learning 22, 1996. doi:10.1007/BF00114723
  • Richard S. Varga, Matrix Iterative Analysis, Prentice-Hall, 1962.
11 thms2 active usersReviewed
Convex OptimizationOptimization·Captain: mikedeng1

SAGA: A Fast Incremental Gradient Method With Support for Non-Strongly Convex Composite Objectives II: The 4n/k Rate of the Averaged Iterate without Strong ConvexityResearch Paper

Motivation

Many problems in statistics and machine learning minimize an average of nnn losses, one per data point, plus a regularizer: least squares, logistic regression, and their ℓ1\ell_1ℓ1​- or ℓ2\ell_2ℓ2​-penalized versions. When nnn is large, a full gradient costs nnn component gradients, while stochastic gradient descent, which uses one component per step, needs decreasing step sizes and converges slowly. Incremental gradient methods with variance reduction (SAG, SVRG, SDCA, Finito, MISO) use one component gradient per step but converge at the rate of a full-gradient method.

SAGA (Defazio, Bach and Lacoste-Julien, NIPS 2014, arXiv:1407.0202) is a method of this family. It handles a non-smooth regularizer through its proximal operator, and it comes with a guarantee when the losses are convex but not strongly convex. This mission covers that second guarantee, Theorem 2 of the paper. A companion mission covers the linear rate under strong convexity (Theorem 1, Corollary 1).

Timeline.

  • 2012: SAG (Le Roux, Schmidt and Bach) gives a linear rate for smooth, strongly convex finite sums. Its analysis does not cover a proximal term.
  • 2013: SVRG (Johnson and Zhang) gives a linear rate for the strongly convex case, using periodic full-gradient passes.
  • 2013: SDCA (Shalev-Shwartz and Zhang) works on the dual and needs strong convexity.
  • 2014: Prox-SVRG (Xiao and Zhang, arXiv:1403.4699) extends SVRG to composite objectives. Its key inequality is reused by SAGA's Theorem 2.
  • 2014: SAGA proves both a linear rate under strong convexity and an O(n/k)O(n/k)O(n/k) rate for the averaged iterate under convexity alone, for composite objectives.

Setting

Let d≥0d\ge 0d≥0 and n≥1n\ge 1n≥1. The components f1,…,fn:Rd→Rf_1,\dots,f_n:\mathbb R^d\to\mathbb Rf1​,…,fn​:Rd→R are convex and differentiable, and each gradient fi′f_i'fi′​ is LLL-Lipschitz (L>0L>0L>0). Write

f(x)=1n∑i=1nfi(x),f′(x)=1n∑i=1nfi′(x).f(x)=\frac1n\sum_{i=1}^n f_i(x),\qquad f'(x)=\frac1n\sum_{i=1}^n f_i'(x).f(x)=n1​i=1∑n​fi​(x),f′(x)=n1​i=1∑n​fi′​(x).

The regularizer h:Rd→Rh:\mathbb R^d\to\mathbb Rh:Rd→R is convex but possibly non-differentiable. The objective is the composite function F=f+hF=f+hF=f+h, and x∗x^*x∗ is any minimizer of FFF. Minimizers need not be unique, and f′(x∗)f'(x^*)f′(x∗) need not vanish.

The proximal operator with parameter γ>0\gamma>0γ>0 is

proxγh(y)=arg⁡min⁡x∈Rd{h(x)+12γ∥x−y∥2}.\mathrm{prox}_\gamma^h(y)=\arg\min_{x\in\mathbb R^d}\Big\{h(x)+\frac1{2\gamma}\|x-y\|^2\Big\}.proxγh​(y)=argx∈Rdmin​{h(x)+2γ1​∥x−y∥2}.

SAGA keeps an iterate xkx^kxk and a table of points ϕ1k,…,ϕnk\phi_1^k,\dots,\phi_n^kϕ1k​,…,ϕnk​, initialized as ϕi0=x0\phi_i^0=x^0ϕi0​=x0. At step k+1k+1k+1 it draws an index jjj uniformly from {1,…,n}\{1,\dots,n\}{1,…,n}, independently of the past, and sets

wk+1=xk−γ[fj′(xk)−fj′(ϕjk)+1n∑i=1nfi′(ϕik)],xk+1=proxγh(wk+1).w^{k+1}=x^k-\gamma\Big[f_j'(x^k)-f_j'(\phi_j^k)+\frac1n\sum_{i=1}^n f_i'(\phi_i^k)\Big],\qquad x^{k+1}=\mathrm{prox}_\gamma^h(w^{k+1}).wk+1=xk−γ[fj′​(xk)−fj′​(ϕjk​)+n1​i=1∑n​fi′​(ϕik​)],xk+1=proxγh​(wk+1).

It then sets ϕjk+1=xk\phi_j^{k+1}=x^kϕjk+1​=xk and leaves the other table entries unchanged. The averaged iterate is xˉk=1k∑t=1kxt\bar x^k=\frac1k\sum_{t=1}^k x^txˉk=k1​∑t=1k​xt, which excludes x0x^0x0.

Formalization targets

Goal: Theorem 2 (p. 11)

With step size γ=1/(3L)\gamma=1/(3L)γ=1/(3L), for every k≥1k\ge1k≥1,

E[F(xˉk)]−F(x∗)≤4nk[2Ln∥x0−x∗∥2+f(x0)−⟨f′(x∗),x0−x∗⟩−f(x∗)].\mathbb E\big[F(\bar x^k)\big]-F(x^*)\le\frac{4n}{k}\Big[\frac{2L}{n}\|x^0-x^*\|^2+f(x^0)-\langle f'(x^*),x^0-x^*\rangle-f(x^*)\Big].E[F(xˉk)]−F(x∗)≤k4n​[n2L​∥x0−x∗∥2+f(x0)−⟨f′(x∗),x0−x∗⟩−f(x∗)].

The expectation is over the indices j1,…,jkj^1,\dots,j^kj1,…,jk. The constants are those printed in the paper.

Milestones (in attack order)

  1. Lemma 1 (p. 6) is an inner-product bound for averages of μ\muμ-strongly convex functions with LLL-Lipschitz gradients. It is stated for μ≥0\mu\ge0μ≥0, and Theorem 2 uses the case μ=0\mu=0μ=0.
  2. Lemma 2 (p. 7): 1n∑i∥fi′(ϕi)−fi′(x∗)∥2≤2L[1n∑ifi(ϕi)−f(x∗)−1n∑i⟨fi′(x∗),ϕi−x∗⟩]\frac1n\sum_i\|f_i'(\phi_i)-f_i'(x^*)\|^2\le 2L\big[\frac1n\sum_i f_i(\phi_i)-f(x^*)-\frac1n\sum_i\langle f_i'(x^*),\phi_i-x^*\rangle\big]n1​∑i​∥fi′​(ϕi​)−fi′​(x∗)∥2≤2L[n1​∑i​fi​(ϕi​)−f(x∗)−n1​∑i​⟨fi′​(x∗),ϕi​−x∗⟩].
  3. The bound on Δ\DeltaΔ (p. 12). Write Δ=−1γ(wk+1−xk)−f′(xk)\Delta=-\frac1\gamma(w^{k+1}-x^k)-f'(x^k)Δ=−γ1​(wk+1−xk)−f′(xk) for the gradient error. For every β>0\beta>0β>0, E∥Δ∥2≤(1+β−1)E∥fj′(ϕjk)−fj′(x∗)∥2+(1+β)E∥fj′(xk)−fj′(x∗)∥2\mathbb E\|\Delta\|^2\le(1+\beta^{-1})\mathbb E\|f_j'(\phi_j^k)-f_j'(x^*)\|^2+(1+\beta)\mathbb E\|f_j'(x^k)-f_j'(x^*)\|^2E∥Δ∥2≤(1+β−1)E∥fj′​(ϕjk​)−fj′​(x∗)∥2+(1+β)E∥fj′​(xk)−fj′​(x∗)∥2.
  4. The prox-SVRG inequality (p. 12): αE∥xk+1−x∗∥2≤α∥xk−x∗∥2−2αγE[F(xk+1)−F(x∗)]+2αγ2E∥Δ∥2\alpha\mathbb E\|x^{k+1}-x^*\|^2\le\alpha\|x^k-x^*\|^2-2\alpha\gamma\mathbb E[F(x^{k+1})-F(x^*)]+2\alpha\gamma^2\mathbb E\|\Delta\|^2αE∥xk+1−x∗∥2≤α∥xk−x∗∥2−2αγE[F(xk+1)−F(x∗)]+2αγ2E∥Δ∥2.
  5. The one-step Lyapunov decrease (p. 12): E[Tk+1]−Tk≤−14nE[F(xk+1)−F(x∗)]\mathbb E[T^{k+1}]-T^k\le-\frac1{4n}\mathbb E[F(x^{k+1})-F(x^*)]E[Tk+1]−Tk≤−4n1​E[F(xk+1)−F(x∗)]. Here T(x,ϕ)=1n∑ifi(ϕi)−f(x∗)−1n∑i⟨fi′(x∗),ϕi−x∗⟩+(c+α)∥x−x∗∥2T(x,\phi)=\frac1n\sum_i f_i(\phi_i)-f(x^*)-\frac1n\sum_i\langle f_i'(x^*),\phi_i-x^*\rangle+(c+\alpha)\|x-x^*\|^2T(x,ϕ)=n1​∑i​fi​(ϕi​)−f(x∗)−n1​∑i​⟨fi′​(x∗),ϕi​−x∗⟩+(c+α)∥x−x∗∥2, with c=3L2nc=\frac{3L}{2n}c=2n3L​ and α=3L8n\alpha=\frac{3L}{8n}α=8n3L​.

In milestones 3–5, E\mathbb EE is the expectation over the single index jjj of the next step, given the current state.

Significance

The result. Theorem 2 shows that one method, with a step size that depends only on LLL, covers composite problems that are not strongly convex. Examples are ℓ1\ell_1ℓ1​-regularized least squares and logistic regression without a ridge term. On these problems the method converges in expected objective value at rate O(n/k)O(n/k)O(n/k). SAG has no proximal analysis, and SDCA requires strong convexity. With the same step size 1/(3L)1/(3L)1/(3L), the paper also states adaptivity to strong convexity, so no strong convexity constant has to be known in advance. The bound is in terms of T0T^0T0, a quantity computable from the starting point.

Formalizing it. The result is proved on paper, but the proof is not self-contained. Its central inequality (milestone 4) is quoted from the prox-SVRG analysis of Xiao and Zhang, with only the remark that their argument uses E[Δ]=0\mathbb E[\Delta]=0E[Δ]=0. A machine-checked proof must therefore reconstruct that argument for SAGA's estimator. To our knowledge, no machine-checked proof of SAGA, SVRG or prox-SVRG exists in Lean or Mathlib. The mission also produces reusable statements about convex functions with Lipschitz gradients (Lemmas 1 and 2) and an explicit finite model of a randomized incremental method.

Difficulty

The naive approach applies the non-expansiveness of the proximal operator to ∥xk+1−x∗∥2\|x^{k+1}-x^*\|^2∥xk+1−x∗∥2, as in the strongly convex proof. That bounds distances, but it produces no term in F(xk+1)−F(x∗)F(x^{k+1})-F(x^*)F(xk+1)−F(x∗). Without strong convexity, the distance terms cannot be traded for function values, so the argument yields no rate.

The function-value term comes from the prox-SVRG inequality (milestone 4), which the paper does not prove. Its difficulty is that xk+1x^{k+1}xk+1 depends on the same random index as Δ\DeltaΔ, so the cross term between them does not vanish in expectation even though E[Δ]=0\mathbb E[\Delta]=0E[Δ]=0. A second difficulty is bookkeeping: wk+1w^{k+1}wk+1 uses the old table, the table entry jjj receives xkx^kxk and not xk+1x^{k+1}xk+1, and the constants must make three coefficients vanish exactly. A final step converts the bound on 1k∑tE[F(xt)]\frac1k\sum_t\mathbb E[F(x^t)]k1​∑t​E[F(xt)] into a bound on E[F(xˉk)]\mathbb E[F(\bar x^k)]E[F(xˉk)], which requires Jensen's inequality for the convex FFF.

Formalization scope

  • Space and indices. Points live in EuclideanSpace ℝ (Fin d). Components are indexed by Fin n (0-based), with n≥1n\ge1n≥1.
  • Gradients and smoothness. The gradients are given maps f' with HasGradientAt (f i) (f' i x) x at every point. Smoothness is the Lipschitz bound ∥fi′(x)−fi′(y)∥≤L∥x−y∥\|f_i'(x)-f_i'(y)\|\le L\|x-y\|∥fi′​(x)−fi′​(y)∥≤L∥x−y∥.
  • Convexity. Convexity is ConvexOn ℝ Set.univ. Lemma 1 uses StrongConvexOn Set.univ μ, whose modulus μ2∥x−y∥2\frac\mu2\|x-y\|^22μ​∥x−y∥2 is the paper's.
  • The regularizer. hhh is real-valued and convex. Extended-valued regularizers such as indicator functions are outside the statement.
  • The proximal map. The proximal operator enters as any map PPP such that P(y)P(y)P(y) minimizes h(z)+12γ∥z−y∥2h(z)+\frac1{2\gamma}\|z-y\|^2h(z)+2γ1​∥z−y∥2 for every yyy. For convex hhh this determines P=proxγhP=\mathrm{prox}_\gamma^hP=proxγh​.
  • State and expectation. The state is the pair (xk,ϕk)(x^k,\phi^k)(xk,ϕk). The expectation over kkk steps is the uniform average over the nkn^knk index sequences, which is exactly the law of kkk independent uniform indices.

Two trivializations are excluded. The averaged-iterate bound carries k≥1k\ge1k≥1, since at k=0k=0k=0 the factor 4n/k4n/k4n/k collapses to 000. The left side is FFF evaluated at the averaged point, not the average of F(xt)F(x^t)F(xt), which is a weaker intermediate step.

A complete development needs the descent lemma and co-coercivity for convex functions with Lipschitz gradients, the characterization and non-expansiveness of the proximal operator, and finite-sum manipulations over index sequences. The lemmas on smooth convex functions and on proximal operators are reusable beyond this mission. Contributions are welcome at every level: proofs of the milestones, a reusable proximal-operator library, and the telescoping argument for the goal.

Selected references

  • A. Defazio, F. Bach, S. Lacoste-Julien, SAGA: A Fast Incremental Gradient Method With Support for Non-Strongly Convex Composite Objectives, NIPS 2014. arXiv:1407.0202
  • L. Xiao, T. Zhang, A Proximal Stochastic Gradient Method with Progressive Variance Reduction, SIAM J. Optim. 24(4), 2014. arXiv:1403.4699
  • N. Le Roux, M. Schmidt, F. Bach, A Stochastic Gradient Method with an Exponential Convergence Rate for Finite Training Sets, NIPS 2012. arXiv:1202.6258
  • R. Johnson, T. Zhang, Accelerating Stochastic Gradient Descent using Predictive Variance Reduction, NIPS 2013. NeurIPS proceedings
  • S. Shalev-Shwartz, T. Zhang, Stochastic Dual Coordinate Ascent Methods for Regularized Loss Minimization, JMLR 14, 2013. arXiv:1209.1873
  • Y. Nesterov, Introductory Lectures on Convex Optimization, Kluwer, 2004. doi:10.1007/978-1-4419-8853-9
11 thms2 active usersReviewed
ProbabilityReinforcement LearningStatistics·Captain: mikedeng1

Reinforcement Learning: An Introduction V: Off-policy Prediction by Importance SamplingTextbook

Motivation

Reinforcement learning methods must explore in order to find good behaviour, yet the quantity they usually want to evaluate is the value of a different, often deterministic, policy. Off-policy prediction separates the two roles: episodes are generated by a behaviour policy bbb, and the goal is the value function vπv_\pivπ​ of a target policy π\piπ. Almost every off-policy method in Sutton and Barto's Reinforcement Learning: An Introduction (2nd ed., MIT Press, 2018), and in the literature that follows it, rests on importance sampling: a return observed under bbb is reweighted by the relative probability of its trajectory under π\piπ and bbb. Section 5.5 of the book introduces the idea for Monte Carlo prediction, §5.6 gives the incremental form of the weighted estimator, and §§5.8–5.9 refine the weights using the internal structure of the return: discounting-aware importance sampling, after Sutton, Mahmood, Precup and van Hasselt (2014), and per-decision importance sampling, introduced by Precup, Sutton and Singh (2000). The book's remarks on the variance of the two estimators (p. 105) cite Precup, Sutton and Dasgupta (2001). Later chapters (7, 11, 12) reuse the same ratios for nnn-step, gradient-TD and eligibility-trace methods.

This mission is the fifth in a series formalizing the book's central mathematical claims. It covers §§5.5–5.9 (pp. 103–115).

Setting

A finite Markov decision process has finite sets of states S\mathcal SS (terminal states included), actions A\mathcal AA and rewards R⊂R\mathcal R \subset \mathbb RR⊂R, and dynamics p(s′,r∣s,a)p(s', r \mid s, a)p(s′,r∣s,a): for each (s,a)(s, a)(s,a) a probability distribution over next state and reward. The state-transition probability is p(s′∣s,a)=∑rp(s′,r∣s,a)p(s' \mid s, a) = \sum_r p(s', r \mid s, a)p(s′∣s,a)=∑r​p(s′,r∣s,a). A policy μ\muμ gives a distribution μ(⋅∣s)\mu(\cdot \mid s)μ(⋅∣s) over actions in every state.

An episode from a start state sss is a sequence S0=s,A0,R1,S1,…,AT−1,RT,STS_0 = s, A_0, R_1, S_1, \dots, A_{T-1}, R_T, S_TS0​=s,A0​,R1​,S1​,…,AT−1​,RT​,ST​ in which S0,…,ST−1S_0, \dots, S_{T-1}S0​,…,ST−1​ are nonterminal and STS_TST​ is terminal. Under μ\muμ it has probability ∏k=0T−1μ(Ak∣Sk) p(Sk+1,Rk+1∣Sk,Ak)\prod_{k=0}^{T-1} \mu(A_k \mid S_k)\, p(S_{k+1}, R_{k+1} \mid S_k, A_k)∏k=0T−1​μ(Ak​∣Sk​)p(Sk+1​,Rk+1​∣Sk​,Ak​). The return is G0=∑k=0T−1γkRk+1G_0 = \sum_{k=0}^{T-1} \gamma^k R_{k+1}G0​=∑k=0T−1​γkRk+1​ with discount rate γ∈[0,1]\gamma \in [0, 1]γ∈[0,1], and the value vπ(s)v_\pi(s)vπ​(s) is the expected return of an episode generated by π\piπ from sss.

The behaviour policy covers the target policy if π(a∣s)>0\pi(a \mid s) > 0π(a∣s)>0 implies b(a∣s)>0b(a \mid s) > 0b(a∣s)>0. The importance-sampling ratio of decisions 0,…,j0, \dots, j0,…,j is

ρ0:j=∏k=0jπ(Ak∣Sk)b(Ak∣Sk),\rho_{0:j} = \prod_{k=0}^{j} \frac{\pi(A_k \mid S_k)}{b(A_k \mid S_k)},ρ0:j​=k=0∏j​b(Ak​∣Sk​)π(Ak​∣Sk​)​,

and the per-decision return weights each reward only by the ratio of the decisions that precede it:

G~0=ρ0:0R1+γρ0:1R2+⋯+γT−1ρ0:T−1RT.\tilde G_0 = \rho_{0:0} R_1 + \gamma \rho_{0:1} R_2 + \dots + \gamma^{T-1} \rho_{0:T-1} R_T .G~0​=ρ0:0​R1​+γρ0:1​R2​+⋯+γT−1ρ0:T−1​RT​.

The book writes these objects at a general time ttt and conditions on St=sS_t = sSt​=s; by the Markov property this is the same as starting the episode at sss, which is what the formal statements do.

Formalization targets

Goal: unbiasedness of ordinary and per-decision importance sampling

For episodes generated by bbb from sss,

Eb[ρ0:T−1G0∣S0=s]=vπ(s)=Eb[G~0∣S0=s].\mathbb E_b\bigl[\rho_{0:T-1} G_0 \mid S_0 = s\bigr] = v_\pi(s) = \mathbb E_b\bigl[\tilde G_0 \mid S_0 = s\bigr].Eb​[ρ0:T−1​G0​∣S0​=s]=vπ​(s)=Eb​[G~0​∣S0​=s].

The first equality is Eq. (5.4) (p. 104); the second is the statement E[ρt:T−1Gt]=E[G~t]\mathbb E[\rho_{t:T-1}G_t] = \mathbb E[\tilde G_t]E[ρt:T−1​Gt​]=E[G~t​] of §5.9 (p. 114).

Milestones

  1. (5.3): the trajectory probability is a product, and the ratio of trajectory probabilities under π\piπ and bbb is ρ0:T−1\rho_{0:T-1}ρ0:T−1​, independent of the dynamics.
  2. (5.4) alone.
  3. (5.13): ∑ab(a∣x) π(a∣x)/b(a∣x)=∑aπ(a∣x)=1\sum_a b(a \mid x)\, \pi(a \mid x)/b(a \mid x) = \sum_a \pi(a \mid x) = 1∑a​b(a∣x)π(a∣x)/b(a∣x)=∑a​π(a∣x)=1 under coverage.
  4. (5.14) and its kkk-th form: Eb[ρ0:T−1Rk]=Eb[ρ0:k−1Rk]\mathbb E_b[\rho_{0:T-1} R_k] = \mathbb E_b[\rho_{0:k-1} R_k]Eb​[ρ0:T−1​Rk​]=Eb​[ρ0:k−1​Rk​] for every k≥1k \ge 1k≥1 (Exercise 5.13).
  5. Example 5.5: in a one-state MDP with a loop, vπ(s)=1v_\pi(s) = 1vπ​(s)=1 and Eb[ρ0:T−1G0]=1\mathbb E_b[\rho_{0:T-1}G_0] = 1Eb​[ρ0:T−1​G0​]=1, yet Eb[(ρ0:T−1G0)2]=∞\mathbb E_b[(\rho_{0:T-1}G_0)^2] = \inftyEb​[(ρ0:T−1​G0​)2]=∞.
  6. (5.7)–(5.8): the incremental rule Vn+1=Vn+(Wn/Cn)(Gn−Vn)V_{n+1} = V_n + (W_n/C_n)(G_n - V_n)Vn+1​=Vn​+(Wn​/Cn​)(Gn​−Vn​) computes the weighted average ∑k<nWkGk/∑k<nWk\sum_{k<n} W_k G_k / \sum_{k<n} W_k∑k<n​Wk​Gk​/∑k<n​Wk​ (Exercise 5.10).
  7. §5.8: Gt=(1−γ)∑h=t+1T−1γh−t−1Gˉt:h+γT−t−1Gˉt:TG_t = (1-\gamma)\sum_{h=t+1}^{T-1}\gamma^{h-t-1}\bar G_{t:h} + \gamma^{T-t-1}\bar G_{t:T}Gt​=(1−γ)∑h=t+1T−1​γh−t−1Gˉt:h​+γT−t−1Gˉt:T​ with flat partial returns Gˉt:h=Rt+1+⋯+Rh\bar G_{t:h} = R_{t+1} + \dots + R_hGˉt:h​=Rt+1​+⋯+Rh​.

Significance

Eq. (5.4) is the reason the first-visit ordinary importance-sampling estimator (5.5) is unbiased, and it is the template for every importance-sampling correction in the rest of the book. The per-decision identity shows that an estimator with fewer ratio factors per reward, (5.15), has the same expectation, which is the starting point for per-decision and control-variate methods for multi-step off-policy learning (Precup, Sutton and Singh 2000). Example 5.5 shows that unbiasedness says nothing about variance: the ordinary estimator can have infinite variance on a two-action problem, which motivates weighted importance sampling and the incremental weighted update of §5.6.

The results of these sections are classical and proved informally in the book, partly as exercises (5.10, 5.13) left without solution. No machine-checked version exists on the platform: a search for importance sampling, off-policy and per-decision returned no statements. The mission produces a formal trajectory model of an episodic MDP under two policies, which later missions on nnn-step off-policy returns and off-policy traces can reuse.

Difficulty

Eq. (5.4) itself is a termwise identity: for every episode, Pr⁡b(episode) ρ0:T−1=Pr⁡π(episode)\Pr_b(\text{episode})\,\rho_{0:T-1} = \Pr_\pi(\text{episode})Prb​(episode)ρ0:T−1​=Prπ​(episode) under coverage. The per-decision identity is not termwise. The later factors of ρ0:T−1\rho_{0:T-1}ρ0:T−1​ multiply a reward that was received before the corresponding decisions, and removing them requires summing over all continuations of an episode prefix, of every remaining length, and using that each factor has conditional expectation one (5.13) and that the continuation terminates with probability one. The obvious attempt, cancelling the factors episode by episode, fails: on a single episode ρ0:T−1R1\rho_{0:T-1}R_1ρ0:T−1​R1​ and ρ0:0R1\rho_{0:0}R_1ρ0:0​R1​ differ.

In Example 5.5 the episodes have no length bound, so the expected square is an infinite series over episode lengths whose divergence must be shown directly.

Formalization scope

  • States, actions and rewards are finite types; the terminal states are a finite subset of the state type. Policies are stochastic, one action set is used in every state, and vπv_\pivπ​ is defined as an expected return, never as the solution of a Bellman equation.
  • Expectations are series over episode lengths of finite sums over episodes. Lean assigns 000 to a divergent series, so the goal and milestones 2 and 4 assume that under bbb every episode from sss terminates within a fixed number HHH of steps with probability one. The book leaves termination implicit; this bounded-horizon hypothesis is a restriction relative to the book's episodic setting and is stated as such. Example 5.5, whose episodes are unbounded, is stated without it, with the expected square in [0,∞][0, \infty][0,∞].
  • The discount rate is kept general in [0,1][0, 1][0,1].
  • The importance-sampling ratio uses real division; a factor with b(Ak∣Sk)=0b(A_k \mid S_k) = 0b(Ak​∣Sk​)=0 evaluates to 000 in Lean, but such episodes have probability 000 under bbb.
  • The flat-partial-return decomposition is an algebraic identity and is stated for every real γ\gammaγ, which is more general than the book's "for any γ∈[0,1)\gamma \in [0,1)γ∈[0,1)".
  • The incremental weighted update is stated with nonnegative weights and W1>0W_1 > 0W1​>0. The book's C0=0C_0 = 0C0​=0 makes (5.8) divide by zero at n=1n = 1n=1 when W1=0W_1 = 0W1​=0; the hypothesis excludes that case.
  • A trivializing formalization is ruled out: vπv_\pivπ​ is the expected return of π\piπ's own episodes, the ratio is computed from the episode, and the per-decision identity, which carries the chapter's content beyond (5.4), is part of the goal.
  • Not stated: the bias and variance comparisons of ordinary and weighted importance sampling (p. 105) and the discounting-aware estimators (5.9)–(5.10) as estimators; these are statistical claims about estimators over a random number of visits that the book does not make precise.

Contributions welcome: proofs of the milestones, a general measure-theoretic version of the trajectory model without the bounded-horizon hypothesis, and variants for action values qπq_\piqπ​ (Exercise 5.6).

Selected references

  • R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018, ISBN 9780262039246, §§5.5–5.9, pp. 103–115. http://incompleteideas.net/book/the-book-2nd.html
  • D. Precup, R. S. Sutton and S. Singh, Eligibility Traces for Off-Policy Policy Evaluation, Proceedings of the 17th International Conference on Machine Learning (ICML), 2000, pp. 759–766 (cited in the book's bibliography).
  • D. Precup, R. S. Sutton and S. Dasgupta, Off-Policy Temporal-Difference Learning with Function Approximation, Proceedings of the 18th International Conference on Machine Learning (ICML), 2001, pp. 417–424 (cited in the book, p. 105).
14 thms2 active usersReviewed
Convex OptimizationOptimization·Captain: mikedeng1

SAGA: A Fast Incremental Gradient Method With Support for Non-Strongly Convex Composite Objectives I: Linear Convergence under Strong ConvexityResearch Paper

Motivation

Many problems in machine learning and statistics are finite sums: an empirical risk f(x)=1n∑i=1nfi(x)f(x)=\frac1n\sum_{i=1}^n f_i(x)f(x)=n1​∑i=1n​fi​(x) over nnn data points, often plus a regulariser hhh such as an ℓ1\ell_1ℓ1​ penalty. When nnn is large, a full gradient of fff costs nnn component gradients, while stochastic gradient descent uses one component per step but converges only sublinearly because its gradient estimate has non-vanishing variance. Incremental gradient methods with variance reduction keep the per-step cost of one component gradient and still converge linearly on strongly convex problems.

SAGA, introduced by Defazio, Bach and Lacoste-Julien at NIPS 2014 (arXiv:1407.0202), is one of the standard methods of this family, alongside SAG, SVRG, SDCA and Finito/MISO. It keeps a table of past component gradients and handles a non-smooth regulariser through its proximal operator.

Timeline. Le Roux, Schmidt and Bach (2012) gave SAG the first linear rate for strongly convex finite sums at the cost of one gradient per step. Shalev-Shwartz and Zhang (2013) proved linear rates for SDCA, a dual method. Johnson and Zhang (2013) introduced SVRG, with periodic full-gradient passes; Xiao and Zhang (2014) extended it to composite objectives (prox-SVRG). SAGA (2014) combines an unbiased SVRG-style estimator with a SAG-style table, and proves a linear rate in the composite strongly convex case with a simple Lyapunov argument.

Setting

Let Rd\mathbb R^dRd carry the Euclidean inner product. There are n≥1n\ge1n≥1 differentiable components f1,…,fn:Rd→Rf_1,\dots,f_n:\mathbb R^d\to\mathbb Rf1​,…,fn​:Rd→R with gradients fi′f_i'fi′​. Each fif_ifi​ is μ\muμ-strongly convex (μ>0\mu>0μ>0): fi(ax+by)≤afi(x)+bfi(y)−abμ2∥x−y∥2f_i(ax+by)\le af_i(x)+bf_i(y)-ab\frac\mu2\|x-y\|^2fi​(ax+by)≤afi​(x)+bfi​(y)−ab2μ​∥x−y∥2 for a,b≥0a,b\ge0a,b≥0, a+b=1a+b=1a+b=1. Each gradient is LLL-Lipschitz: ∥fi′(x)−fi′(y)∥≤L∥x−y∥\|f_i'(x)-f_i'(y)\|\le L\|x-y\|∥fi′​(x)−fi′​(y)∥≤L∥x−y∥. Write f=1n∑ifif=\frac1n\sum_i f_if=n1​∑i​fi​ and f′=1n∑ifi′f'=\frac1n\sum_i f_i'f′=n1​∑i​fi′​. The regulariser h:Rd→Rh:\mathbb R^d\to\mathbb Rh:Rd→R is convex, and the goal is to minimise the composite objective F=f+hF=f+hF=f+h; x∗x^*x∗ denotes its minimiser, which is unique.

The proximal operator with step γ>0\gamma>0γ>0 is

prox⁡γh(y)=argmin⁡x{h(x)+12γ∥x−y∥2}.\operatorname{prox}^h_\gamma(y)=\operatorname*{argmin}_{x}\Big\{h(x)+\tfrac1{2\gamma}\|x-y\|^2\Big\}.proxγh​(y)=xargmin​{h(x)+2γ1​∥x−y∥2}.

SAGA keeps an iterate xkx^kxk and points ϕ1k,…,ϕnk\phi_1^k,\dots,\phi_n^kϕ1k​,…,ϕnk​ at which the stored gradients fi′(ϕik)f_i'(\phi_i^k)fi′​(ϕik​) were taken. It starts from x0x^0x0 with ϕi0=x0\phi_i^0=x^0ϕi0​=x0. At iteration k+1k+1k+1 it draws jjj uniformly from {1,…,n}\{1,\dots,n\}{1,…,n}, independently of the past, and sets

wk+1=xk−γ[fj′(xk)−fj′(ϕjk)+1n∑i=1nfi′(ϕik)],xk+1=prox⁡γh(wk+1),w^{k+1}=x^k-\gamma\Big[f_j'(x^k)-f_j'(\phi_j^k)+\frac1n\sum_{i=1}^n f_i'(\phi_i^k)\Big],\qquad x^{k+1}=\operatorname{prox}^h_\gamma(w^{k+1}),wk+1=xk−γ[fj′​(xk)−fj′​(ϕjk​)+n1​i=1∑n​fi′​(ϕik​)],xk+1=proxγh​(wk+1),

then ϕjk+1=xk\phi_j^{k+1}=x^kϕjk+1​=xk, with every other entry unchanged.

The analysis uses the Lyapunov function

T(x,{ϕi})=1n∑ifi(ϕi)−f(x∗)−1n∑i⟨fi′(x∗),ϕi−x∗⟩+c∥x−x∗∥2.T(x,\{\phi_i\})=\frac1n\sum_i f_i(\phi_i)-f(x^*)-\frac1n\sum_i\langle f_i'(x^*),\phi_i-x^*\rangle+c\|x-x^*\|^2 .T(x,{ϕi​})=n1​i∑​fi​(ϕi​)−f(x∗)−n1​i∑​⟨fi′​(x∗),ϕi​−x∗⟩+c∥x−x∗∥2.

Formalization targets

Goal: Corollary 1 (p. 8)

With γ=12(μn+L)\gamma=\frac1{2(\mu n+L)}γ=2(μn+L)1​, for every k≥0k\ge0k≥0,

E∥xk−x∗∥2≤(1−μ2(μn+L))k[∥x0−x∗∥2+nμn+L(f(x0)−⟨f′(x∗),x0−x∗⟩−f(x∗))],\mathbb E\|x^k-x^*\|^2\le\Big(1-\frac{\mu}{2(\mu n+L)}\Big)^k\Big[\|x^0-x^*\|^2+\frac{n}{\mu n+L}\big(f(x^0)-\langle f'(x^*),x^0-x^*\rangle-f(x^*)\big)\Big],E∥xk−x∗∥2≤(1−2(μn+L)μ​)k[∥x0−x∗∥2+μn+Ln​(f(x0)−⟨f′(x∗),x0−x∗⟩−f(x∗))],

where the expectation is over the indices drawn in the first kkk iterations. The constants are the paper's.

Theorem 1 (p. 7)

With γ\gammaγ as above, c=12γ(1−γμ)nc=\frac1{2\gamma(1-\gamma\mu)n}c=2γ(1−γμ)n1​ and κ=1γμ\kappa=\frac1{\gamma\mu}κ=γμ1​, for every state (xk,{ϕik})(x^k,\{\phi^k_i\})(xk,{ϕik​}),

E[Tk+1]≤(1−1κ)Tk,\mathbb E\big[T^{k+1}\big]\le\Big(1-\frac1\kappa\Big)T^k ,E[Tk+1]≤(1−κ1​)Tk,

with the expectation over the next index only.

Supporting lemmas

Lemma 4 (p. 10), a lower bound combining strong convexity and smoothness; Lemma 1 (pp. 6–7), its average over the components; Lemma 2 (p. 7), which bounds the stale-gradient variance by the table part of TTT; and Lemma 3 (p. 7), a second-moment bound for the SAGA step.

Significance

The result. Corollary 1 gives an ε\varepsilonε-accurate iterate in expectation after O((n+L/μ)log⁡(1/ε))O\big((n+L/\mu)\log(1/\varepsilon)\big)O((n+L/μ)log(1/ε)) component-gradient evaluations. This is the complexity of full-gradient descent with the condition number decoupled from nnn, and it holds in the composite setting, so it covers the lasso and elastic-net problems that SAG's analysis does not reach. The paper notes that the rate improves on the published rates of SAG and SVRG and is within a factor 2 of SDCA's. Theorem 1 is the template of later Lyapunov analyses of variance-reduced methods.

Formalizing it. The result has been proved since 2014, and no machine-checked proof is known to this mission. The work left is to formalize the known proof: the convexity inequalities (Lemmas 4, 1, 2), the variance computation (Lemma 3), the one-step contraction (Theorem 1), and the passage from conditional to total expectation along the random index sequence (Corollary 1). The paper's Lemma 3 has a sign misprint, which the formalization corrects; see the scope section.

Difficulty

The obvious argument for SGD-type methods bounds E∥xk+1−x∗∥2\mathbb E\|x^{k+1}-x^*\|^2E∥xk+1−x∗∥2 in terms of ∥xk−x∗∥2\|x^k-x^*\|^2∥xk−x∗∥2 alone. That fails here: the variance of the SAGA estimator depends on the stale table points ϕik\phi_i^kϕik​, which can be far from x∗x^*x∗ even when xkx^kxk is close. One needs a potential that also measures the table. Balancing the terms of TTT then requires the four round-bracket coefficients in the paper's display (10) to be non-positive for the specific γ\gammaγ, ccc and an auxiliary β=(2μn+L)/L\beta=(2\mu n+L)/Lβ=(2μn+L)/L. Checking these coefficients is routine but long algebra in μ\muμ, LLL, nnn. The composite case adds one step: since f′(x∗)≠0f'(x^*)\neq0f′(x∗)=0 in general, the argument goes through the fixed-point identity x∗=prox⁡γh(x∗−γf′(x∗))x^*=\operatorname{prox}^h_\gamma(x^*-\gamma f'(x^*))x∗=proxγh​(x∗−γf′(x∗)) and the non-expansiveness of the proximal operator, neither of which is a numbered result of the paper.

Formalization scope

  • Space and data. The space is EuclideanSpace ℝ (Fin d). Components are indexed by Fin n (0-based), with 0 < n. The gradients fi′f_i'fi′​ are given maps with HasGradientAt (f i) (f' i x) x. Strong convexity is Mathlib's StrongConvexOn Set.univ μ, whose modulus μ2∥x−y∥2\frac\mu2\|x-y\|^22μ​∥x−y∥2 is the paper's.

  • Regulariser and minimiser. hhh is real-valued and convex; extended-valued regularisers are out of scope, as on the page. A minimiser x∗x^*x∗ of f+hf+hf+h is a hypothesis.

  • Proximal operator. It is any map PPP such that P(y)P(y)P(y) minimises h(x)+12γ∥x−y∥2h(x)+\frac1{2\gamma}\|x-y\|^2h(x)+2γ1​∥x−y∥2 for every yyy (IsProxPoint). The minimiser is unique, so PPP is prox⁡γh\operatorname{prox}^h_\gammaproxγh​.

  • State and expectation. The state is the pair (x,ϕ)(x,\phi)(x,ϕ). The run after kkk steps is a deterministic function of the index sequence in Fin k → Fin n. The expectation in Corollary 1 is the average over all nkn^knk sequences, which is exactly the law of kkk independent uniform indices; no measure theory is involved. Theorem 1's conditional expectation is the average over the next index.

  • Constants and corrections. Constants are as printed and fixed, not "for some constant" and not "for all small enough steps". Lemma 4 carries the hypothesis μ<L\mu<Lμ<L, which its fractions 1/(L−μ)1/(L-\mu)1/(L−μ) require. Lemma 3 is stated with +γf′(x∗)+\gamma f'(x^*)+γf′(x∗), as in its proof and its use in Theorem 1; the printed −γf′(x∗)-\gamma f'(x^*)−γf′(x∗) is false whenever f′(x∗)≠0f'(x^*)\neq0f′(x∗)=0.

  • Trivializing formalizations, ruled out. Taking the proximal step as merely non-expansive, fixing an index sequence instead of averaging over all of them, measuring x∗x^*x∗ against fff instead of f+hf+hf+h, or restricting Theorem 1 to reachable states changes the theorem and is excluded.

  • Infrastructure. A complete development needs:

    • the co-coercivity inequality for convex functions with Lipschitz gradient;
    • existence, uniqueness and non-expansiveness of the proximal map of a finite convex function;
    • the optimality condition x∗=prox⁡γh(x∗−γf′(x∗))x^*=\operatorname{prox}^h_\gamma(x^*-\gamma f'(x^*))x∗=proxγh​(x∗−γf′(x∗));
    • finite-sum variance identities.

    These pieces are reusable well beyond SAGA, by SVRG, SAG and proximal-gradient analyses. Contributions of any of them, or of proofs of the individual milestones, are welcome.

Selected references

  • A. Defazio, F. Bach, S. Lacoste-Julien, SAGA: A Fast Incremental Gradient Method With Support for Non-Strongly Convex Composite Objectives, NIPS 2014. arXiv:1407.0202
  • N. Le Roux, M. Schmidt, F. Bach, A Stochastic Gradient Method with an Exponential Convergence Rate for Finite Training Sets, NIPS 2012. arXiv:1202.6258
  • R. Johnson, T. Zhang, Accelerating Stochastic Gradient Descent using Predictive Variance Reduction, NIPS 2013. NeurIPS proceedings
  • L. Xiao, T. Zhang, A Proximal Stochastic Gradient Method with Progressive Variance Reduction, SIAM J. Optim. 24(4), 2014. arXiv:1403.4699
  • S. Shalev-Shwartz, T. Zhang, Stochastic Dual Coordinate Ascent Methods for Regularized Loss Minimization, JMLR 14, 2013. arXiv:1209.1873
  • Y. Nesterov, Introductory Lectures on Convex Optimization, Kluwer, 2004. doi:10.1007/978-1-4419-8853-9
11 thms2 active usersReviewed
Bandit AlgorithmsOperations ResearchStatistics·Captain: mikedeng1

Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems III: Contextual Bandits and the Banditron Mistake BoundTextbook

Motivation

In many sequential decision problems the learner sees side information before acting. A news site chooses an article for a visitor whose history and location it knows; an ad server chooses an advertisement for a query. Only the reward of the chosen action is observed. These are contextual bandit problems, and Chapter 4 of Bubeck and Cesa-Bianchi's monograph Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems (arXiv:1204.5721v2) surveys several of their formal versions. In a contextual problem the learner is compared with the best policy, a map from contexts to arms, rather than with the best single arm.

This mission covers three of the chapter's models. The first marks each round with a context from a finite set. In the second, NNN experts give advice, as in prediction with expert advice. The third is the bandit multiclass problem: a linear classifier predicts one of KKK labels and then learns only whether its prediction was right. The goal is the mistake bound of the Banditron (Kakade, Shalev-Shwartz and Tewari, ICML 2008). The bound shows that one bit of feedback per round suffices to compete with every linear classifier, at regret O(n2/3)O(n^{2/3})O(n2/3).

Setting

There are K≥2K \ge 2K≥2 arms (or labels) {1,…,K}\{1,\dots,K\}{1,…,K} and rounds t=1,…,nt = 1, \dots, nt=1,…,n.

Adversarial losses. At round ttt an adversary assigns losses ℓi,t∈[0,1]\ell_{i,t} \in [0,1]ℓi,t​∈[0,1] to the arms and may adapt to the forecaster's past plays I1,…,It−1I_1, \dots, I_{t-1}I1​,…,It−1​. The forecaster draws ItI_tIt​ at random from a distribution ptp_tpt​ that depends on what it has observed, and it observes only ℓIt,t\ell_{I_t,t}ℓIt​,t​. Expectations E\mathbb EE are over the forecaster's draws.

Side information. Each round carries a context sts_tst​ from a finite set S\mathcal SS, and the sequence s1,s2,…s_1, s_2, \dotss1​,s2​,… is fixed in advance. The pseudo-regret against context-to-arm maps is

R‾nS=max⁡g:S→{1,…,K}E[∑t=1nℓIt,t−∑t=1nℓg(st),t].\overline R^{\mathcal S}_n = \max_{g:\mathcal S\to\{1,\dots,K\}} \mathbb E\Big[\sum_{t=1}^n \ell_{I_t,t} - \sum_{t=1}^n \ell_{g(s_t),t}\Big].RnS​=g:S→{1,…,K}max​E[t=1∑n​ℓIt​,t​−t=1∑n​ℓg(st​),t​].

The S-Exp3 forecaster runs one instance of Exp3 (Section 3.1 of the book) on each context.

Expert advice. At each round each of NNN experts jjj proposes a distribution ξtj\xi^j_tξtj​ over arms, which may depend on the forecaster's past plays. The contextual pseudo-regret is

R‾nctx=max⁡k=1,…,NE[∑t=1nℓIt,t−∑t=1nEi∼ξtkℓi,t].\overline R^{\mathrm{ctx}}_n = \max_{k=1,\dots,N}\mathbb E\Big[\sum_{t=1}^n \ell_{I_t,t} - \sum_{t=1}^n \mathbb E_{i\sim\xi^k_t}\ell_{i,t}\Big].Rnctx​=k=1,…,Nmax​E[t=1∑n​ℓIt​,t​−t=1∑n​Ei∼ξtk​​ℓi,t​].

Exp4 (Fig. 4.1) runs exponential weights over the experts with importance-weighted loss estimates.

Bandit multiclass. The examples (xt,yt)∈Rd×{1,…,K}(x_t, y_t) \in \mathbb R^d \times \{1,\dots,K\}(xt​,yt​)∈Rd×{1,…,K} are fixed in advance, with ∥xt∥=1\|x_t\| = 1∥xt​∥=1 (Euclidean). A K×dK\times dK×d matrix UUU classifies xxx by arg⁡max⁡i(Ux)i\arg\max_i (Ux)_iargmaxi​(Ux)i​. Its multiclass hinge loss on round ttt is ℓt(U)=[1−(Uxt)yt+max⁡i≠yt(Uxt)i]+\ell_t(U) = [1 - (Ux_t)_{y_t} + \max_{i\neq y_t}(Ux_t)_i]_+ℓt​(U)=[1−(Uxt​)yt​​+maxi=yt​​(Uxt​)i​]+​. Write Ln(U)=∑t≤nℓt(U)L_n(U) = \sum_{t\le n}\ell_t(U)Ln​(U)=∑t≤n​ℓt​(U) for the cumulative hinge loss, Lˉn(U)=Ln(U)/n\bar L_n(U) = L_n(U)/nLˉn​(U)=Ln​(U)/n for its average, and ∥U∥\|U\|∥U∥ for the Frobenius norm. The multiclass Perceptron predicts y^t=arg⁡max⁡i(Wtxt)i\hat y_t = \arg\max_i (W_tx_t)_iy^​t​=argmaxi​(Wt​xt​)i​ and, after seeing yty_tyt​, adds xtx_txt​ to row yty_tyt​ and subtracts it from row y^t\hat y_ty^​t​. The Banditron (p. 58) predicts YtY_tYt​ from pi,t=(1−γ)1y^t=i+γ/Kp_{i,t} = (1-\gamma)\mathbb 1_{\hat y_t = i} + \gamma/Kpi,t​=(1−γ)1y^​t​=i​+γ/K. It observes only 1Yt=yt\mathbb 1_{Y_t = y_t}1Yt​=yt​​ and updates Wt+1=Wt+X~tW_{t+1} = W_t + \widetilde X_tWt+1​=Wt​+Xt​, where (X~t)i,j=xt,j(1Yt=yt1Yt=i/pi,t−1y^t=i)(\widetilde X_t)_{i,j} = x_{t,j}\big(\mathbb 1_{Y_t=y_t}\mathbb 1_{Y_t=i}/p_{i,t} - \mathbb 1_{\hat y_t=i}\big)(Xt​)i,j​=xt,j​(1Yt​=yt​​1Yt​=i​/pi,t​−1y^​t​=i​). Its number of mistakes is Mn=∑t≤n1Yt≠ytM_n = \sum_{t\le n}\mathbb 1_{Y_t\neq y_t}Mn​=∑t≤n​1Yt​=yt​​.

Formalization targets

Goal: Theorem 4.7 (Banditron)

For n≥8Kn \ge 8Kn≥8K, γ=(K/n)1/3\gamma = (K/n)^{1/3}γ=(K/n)1/3, every example sequence as above and every K×dK\times dK×d matrix UUU,

E Mn≤Ln(U)+(1+∥U∥2Lˉn(U))K1/3n2/3+2∥U∥2K2/3n1/3+2 ∥U∥K1/6n1/3.\mathbb E\,M_n \le L_n(U) + \Big(1 + \|U\|\sqrt{2\bar L_n(U)}\Big)K^{1/3}n^{2/3} + 2\|U\|^2K^{2/3}n^{1/3} + \sqrt2\,\|U\|K^{1/6}n^{1/3}.EMn​≤Ln​(U)+(1+∥U∥2Lˉn​(U)​)K1/3n2/3+2∥U∥2K2/3n1/3+2​∥U∥K1/6n1/3.

Milestones

  1. Multiclass Perceptron bound (Section 4.4, p. 57). For every n≥1n \ge 1n≥1 and UUU, ∑t≤n1y^t≠yt≤Ln(U)+2∥U∥2+∥U∥2nLˉn(U)\sum_{t\le n}\mathbb 1_{\hat y_t\ne y_t} \le L_n(U) + 2\|U\|^2 + \|U\|\sqrt{2n\bar L_n(U)}∑t≤n​1y^​t​=yt​​≤Ln​(U)+2∥U∥2+∥U∥2nLˉn​(U)​.
  2. Theorem 4.1 (p. 44). S-Exp3 satisfies R‾nS≤2n∣S∣Kln⁡K\overline R^{\mathcal S}_n \le \sqrt{2n|\mathcal S|K\ln K}RnS​≤2n∣S∣KlnK​.
  3. Theorem 4.2 (p. 46), with corrected constants. Exp4 without mixing satisfies R‾nctx≤2nKln⁡N\overline R^{\mathrm{ctx}}_n \le \sqrt{2nK\ln N}Rnctx​≤2nKlnN​ for ηt=2ln⁡N/(nK)\eta_t = \sqrt{2\ln N/(nK)}ηt​=2lnN/(nK)​, and R‾nctx≤2nKln⁡N\overline R^{\mathrm{ctx}}_n \le 2\sqrt{nK\ln N}Rnctx​≤2nKlnN​ for ηt=ln⁡N/(tK)\eta_t = \sqrt{\ln N/(tK)}ηt​=lnN/(tK)​.
  4. Theorem 4.3 (p. 50), with corrected learning rate. Let the plays be drawn from distributions qtq_tqt​ with qi,t≥ε>0q_{i,t}\ge\varepsilon > 0qi,t​≥ε>0, and let Exp3 run on the estimates ℓi,t1It=i/qi,t\ell_{i,t}\mathbb 1_{I_t=i}/q_{i,t}ℓi,t​1It​=i​/qi,t​ with η=2εln⁡K/n\eta = \sqrt{2\varepsilon\ln K/n}η=2εlnK/n​. Then max⁡kE[∑tEi∼ptℓi,t−∑tℓk,t]≤(2n/ε)ln⁡K\max_k \mathbb E\big[\sum_t \mathbb E_{i\sim p_t}\ell_{i,t} - \sum_t\ell_{k,t}\big] \le \sqrt{(2n/\varepsilon)\ln K}maxk​E[∑t​Ei∼pt​​ℓi,t​−∑t​ℓk,t​]≤(2n/ε)lnK​.

Significance

Theorem 4.7 shows that, on any sequence of examples, the bandit version of online multiclass classification costs at most O(K1/3n2/3)O(K^{1/3}n^{2/3})O(K1/3n2/3) mistakes beyond the hinge loss of the best linear classifier. The full-information Perceptron, by comparison, pays O(n)O(\sqrt n)O(n​). The bound has no stochastic assumption and has explicit constants. Theorems 4.1–4.3 are the basic regret guarantees for side information and expert advice. Theorem 4.3 in particular lets learning algorithms serve as experts inside Exp4, which is the construction behind Theorem 4.5.

The mission produces machine-checked statements, and eventually proofs, of these results with fully explicit constants and an explicit model of adaptive adversaries and adaptive advice. To the curators' knowledge none of the Banditron, the multiclass Perceptron bound, S-Exp3 or Theorem 4.3 is formalized anywhere. The platform's Bandit Algorithms series has a proved Exp4 bound, but only for advice and rewards fixed in advance. The book proves all four milestones and the goal; two printed statements (4.2 and 4.3) contain misprints that this mission corrects.

Difficulty

The Banditron bound concerns a randomized process whose weight matrix depends on all earlier random predictions. The Perceptron argument tracks ⟨U,Wn+1⟩\langle U, W_{n+1}\rangle⟨U,Wn+1​⟩ and ∥Wn+1∥2\|W_{n+1}\|^2∥Wn+1​∥2. It carries over only in conditional expectation, and the second moment of the importance-weighted update is of order K/γK/\gammaK/γ on rounds where y^t≠yt\hat y_t \neq y_ty^​t​=yt​ and of order γ\gammaγ otherwise. Combining these into one inequality for ∑tP(y^t≠yt)\sum_t\mathbb P(\hat y_t\neq y_t)∑t​P(y^​t​=yt​) and then for EMn\mathbb E M_nEMn​ requires solving a quadratic inequality in the presence of expectations, and the constants must come out as printed. For the Exp3/Exp4 results, the obstacle is that losses and advice adapt to past plays. The standard potential argument has to be run conditionally on the history, and a version that fixes the losses in advance proves a weaker theorem.

Formalization scope

  • Rounds and laws. Rounds are numbered from 000 in Lean (Lean round ttt is the book's round t+1t+1t+1). Every forecaster is a sampling rule from past plays to weights on Fin K. The law of the first nnn plays is the product ∏tpt(ωt∣ω<t)\prod_t p_t(\omega_t\mid\omega_{<t})∏t​pt​(ωt​∣ω<t​) over sequences ω:Fin n→Fin K\omega : \mathrm{Fin}\,n\to\mathrm{Fin}\,Kω:Finn→FinK, and expectations are finite sums against it. The adversary and the experts are deterministic functions of past plays; an independent randomized adversary is a mixture of these. The examples of the Banditron are fixed.
  • Argmax. y^t\hat y_ty^​t​ uses any argmax selector; all tie-breaking rules are covered.
  • Norms. ∥xt∥=1\|x_t\| = 1∥xt​∥=1 is the Euclidean condition ∑jxt,j2=1\sum_j x_{t,j}^2 = 1∑j​xt,j2​=1; ∥U∥\|U\|∥U∥ is the Frobenius norm written out explicitly.
  • Infima and maxima. Each "inf⁡U\inf_UinfU​" and "max⁡k\max_kmaxk​" of the book is stated as "for every UUU" or "for every kkk", which is equivalent.
  • Explicit constants. Every bound is the one printed or, for the corrected items, the one the proof yields. No O(⋅)O(\cdot)O(⋅) appears.
  • Corrected misprints. Theorem 4.7 prints the examples in Rd×{−1,+1}\mathbb R^d\times\{-1,+1\}Rd×{−1,+1}; labels are in {1,…,K}\{1,\dots,K\}{1,…,K}. Theorem 4.2 prints 2nNln⁡K\sqrt{2nN\ln K}2nNlnK​ and 2nNln⁡K2\sqrt{nN\ln K}2nNlnK​; the proof gives 2nKln⁡N\sqrt{2nK\ln N}2nKlnN​ and 2nKln⁡N2\sqrt{nK\ln N}2nKlnN​. Theorem 4.3 prints η=2ln⁡K/(nK)\eta = \sqrt{2\ln K/(nK)}η=2lnK/(nK)​; (4.7) follows from the proof with η=2εln⁡K/n\eta = \sqrt{2\varepsilon\ln K/n}η=2εlnK/n​.
  • Parameter range. At n=8Kn = 8Kn=8K the Banditron's γ\gammaγ equals 1/21/21/2, outside the box's open interval (0,1/2)(0,1/2)(0,1/2). The proof uses only γ≤1/2\gamma\le 1/2γ≤1/2, so n=8Kn = 8Kn=8K is included.
  • Ruling out trivial forms. Theorem 4.1 is stated for the explicit S-Exp3 forecaster, not as an existence claim, so no forecaster tuned to the losses can witness it. The losses and the advice are allowed to adapt, so a proof for oblivious sequences does not suffice.
  • Left out. Theorem 4.4 (Exp4 with mixing) is proved in the book only by reference. The argument that reference suggests yields 32γn+Kln⁡N/γ\tfrac32\gamma n + K\ln N/\gamma23​γn+KlnN/γ, not the printed γn/2+Kln⁡N/γ\gamma n/2 + K\ln N/\gammaγn/2+KlnN/γ. Theorem 4.5 is stated with O(⋅)O(\cdot)O(⋅), Theorem 4.6 "for some constant ccc", and Eq. (4.8) is left to the reader.

Useful reusable infrastructure: the path-law expectation for history-dependent sampling, the exponential-weights potential argument under adaptive losses, and Perceptron-type inner-product arguments for matrices. Proofs of any milestone and of the goal are welcome.

Selected references

  • S. Bubeck, N. Cesa-Bianchi, Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems, Foundations and Trends in Machine Learning 5(1), 2012. arXiv:1204.5721v2. https://arxiv.org/abs/1204.5721 ; https://doi.org/10.1561/2200000024
  • S. M. Kakade, S. Shalev-Shwartz, A. Tewari, Efficient Bandit Algorithms for Online Multiclass Prediction, ICML 2008. https://doi.org/10.1145/1390156.1390212
  • P. Auer, N. Cesa-Bianchi, Y. Freund, R. E. Schapire, The Nonstochastic Multiarmed Bandit Problem, SIAM Journal on Computing 32(1), 2002. https://doi.org/10.1137/S0097539701398375
  • O.-A. Maillard, R. Munos, Adaptive Bandits: Towards the Best History-Dependent Strategy, AISTATS 2011. https://proceedings.mlr.press/v15/maillard11a.html
11 thms2 active usersReviewed
Bandit AlgorithmsConvex OptimizationOperations Research+1·Captain: mikedeng1

Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems V: Bandit Convex Optimization with One-Point FeedbackTextbook

Motivation

In bandit convex optimization a forecaster repeatedly picks a point xtx_txt​ of a convex set K⊆Rd\mathcal K\subseteq\mathbb R^dK⊆Rd, and an adversary picks a convex loss ℓt\ell_tℓt​. The forecaster pays ℓt(xt)\ell_t(x_t)ℓt​(xt​) and observes only that number: it never sees the function, its gradient, or its value elsewhere. This is the model of online optimization with only function-value access, as in tuning a system online from measured costs, dynamic pricing with an unknown convex demand-cost curve, or routing with path costs observed only on the route taken. The question is how fast the forecaster can approach the best fixed point in hindsight.

Chapter 6 of Bubeck and Cesa-Bianchi's monograph (arXiv:1204.5721v2, Foundations and Trends in Machine Learning 5(1), 2012) treats the problem through spherical gradient estimates fed to projected gradient descent. The one-point method is due to Flaxman, Kalai and McMahan (SODA 2005, arXiv:cs/0408007), who obtained an O(n3/4)\mathcal O(n^{3/4})O(n3/4) regret bound. Agarwal, Dekel and Xiao (COLT 2010) showed that two function evaluations per round allow O(n)\mathcal O(\sqrt n)O(n​). Whether one-point feedback admits n\sqrt nn​ regret was open when the monograph was written (p. 94); Bubeck, Eldan and Lee (STOC 2017, arXiv:1607.03084) later obtained n\sqrt nn​ regret up to logarithmic and polynomial-in-ddd factors for convex losses, with a different and much more involved algorithm.

Setting

Let B={x∈Rd:∥x∥≤1}\mathbb B=\{x\in\mathbb R^d:\|x\|\le1\}B={x∈Rd:∥x∥≤1} be the closed Euclidean unit ball and S={x:∥x∥=1}\mathbb S=\{x:\|x\|=1\}S={x:∥x∥=1} the unit sphere, with unnormalized spherical measure σ\sigmaσ, so that σ(S)=d Vol(B)\sigma(\mathbb S)=d\,\mathrm{Vol}(\mathbb B)σ(S)=dVol(B). Fix δ>0\delta>0δ>0. For a loss ℓ\ellℓ, the smoothed loss is ℓ~(x)=E ℓ(x+δB)\widetilde\ell(x)=\mathbb E\,\ell(x+\delta B)ℓ(x)=Eℓ(x+δB) with BBB uniform on B\mathbb BB.

The set K\mathcal KK is closed and convex with rB⊆K⊆RBr\mathbb B\subseteq\mathcal K\subseteq R\mathbb BrB⊆K⊆RB. The losses ℓ1,ℓ2,⋯:Rd→R\ell_1,\ell_2,\dots:\mathbb R^d\to\mathbb Rℓ1​,ℓ2​,⋯:Rd→R are GGG-Lipschitz, differentiable and convex, and are fixed before the game (an oblivious adversary).

OSGD (Online Stochastic Gradient Descent) on a set K′\mathcal K'K′ with learning rate η\etaη starts at x1=0x_1=0x1​=0 and sets xt+1=argmin⁡y∈K′∥y−(xt−ηg~t(xt))∥x_{t+1}=\operatorname{argmin}_{y\in\mathcal K'}\|y-(x_t-\eta\widetilde g_t(x_t))\|xt+1​=argminy∈K′​∥y−(xt​−ηg​t​(xt​))∥, where g~t\widetilde g_tg​t​ is a gradient estimate. With S1,S2,…S_1,S_2,\dotsS1​,S2​,… independent and uniform on S\mathbb SS:

  • the two-point estimate (6.1) is g~t(xt)=d2δ(ℓt(Xt+)−ℓt(Xt−))St\widetilde g_t(x_t)=\frac d{2\delta}\big(\ell_t(X_t^+)-\ell_t(X_t^-)\big)S_tg​t​(xt​)=2δd​(ℓt​(Xt+​)−ℓt​(Xt−​))St​ with Xt±=xt±δStX_t^\pm=x_t\pm\delta S_tXt±​=xt​±δSt​; the played point is Xt+X_t^+Xt+​ or Xt−X_t^-Xt−​ by a fair coin;
  • the one-point estimate (6.3) is g~t(xt)=dδ ℓt(X~t)St\widetilde g_t(x_t)=\frac d\delta\,\ell_t(\widetilde X_t)S_tg​t​(xt​)=δd​ℓt​(Xt​)St​ with played point X~t=xt+δSt\widetilde X_t=x_t+\delta S_tXt​=xt​+δSt​.

OSGD runs on the shrunken set K′=(1−δ/r)K\mathcal K'=(1-\delta/r)\mathcal KK′=(1−δ/r)K, so that the perturbed points stay in K\mathcal KK. The pseudo-regret is

R‾n=E∑t=1nℓt(X~t)−min⁡x∈K∑t=1nℓt(x).\overline R_n=\mathbb E\sum_{t=1}^n\ell_t(\widetilde X_t)-\min_{x\in\mathcal K}\sum_{t=1}^n\ell_t(x).Rn​=Et=1∑n​ℓt​(Xt​)−x∈Kmin​t=1∑n​ℓt​(x).

Formalization targets

Goal: Theorem 6.2, tuned

If in addition ∣ℓt∣≤L|\ell_t|\le L∣ℓt​∣≤L on K\mathcal KK, and δ=(2n)−1/4RdL/((3+R/r)G)\delta=(2n)^{-1/4}\sqrt{RdL/((3+R/r)G)}δ=(2n)−1/4RdL/((3+R/r)G)​, η=(2n)−3/4R3/(dL(3+R/r)G)\eta=(2n)^{-3/4}\sqrt{R^3/(dL(3+R/r)G)}η=(2n)−3/4R3/(dL(3+R/r)G)​, then one-point OSGD satisfies

R‾n≤4n3/4RdL (3+R/r) G.\overline R_n\le 4n^{3/4}\sqrt{RdL\,(3+R/r)\,G}.Rn​≤4n3/4RdL(3+R/r)G​.

Milestones

  1. Lemma 6.1: ∇∫Bℓ(x+δb) db=1δ∫Sℓ(x+δs)s dσ(s)\nabla\int_{\mathbb B}\ell(x+\delta b)\,db=\frac1\delta\int_{\mathbb S}\ell(x+\delta s)s\,d\sigma(s)∇∫B​ℓ(x+δb)db=δ1​∫S​ℓ(x+δs)sdσ(s).
  2. Lemma 6.2: dδE[ℓ(x+δS)S]=∇E ℓ(x+δB)\frac d\delta\mathbb E[\ell(x+\delta S)S]=\nabla\mathbb E\,\ell(x+\delta B)δd​E[ℓ(x+δS)S]=∇Eℓ(x+δB).
  3. Eq. (6.2): ∣ℓ(x)−ℓ~(x)∣≤δG|\ell(x)-\widetilde\ell(x)|\le\delta G∣ℓ(x)−ℓ(x)∣≤δG.
  4. Lemma 6.3: the queried points' regret against xxx is at most the smoothed regret of the iterates against (1−ξ)x(1-\xi)x(1−ξ)x, plus 3δGn+ξGRn3\delta Gn+\xi GRn3δGn+ξGRn.
  5. Theorem 6.1: two-point OSGD has R‾n≤R2/η+η(Gd)2n+δ(3+R/r)Gn\overline R_n\le R^2/\eta+\eta(Gd)^2n+\delta(3+R/r)GnRn​≤R2/η+η(Gd)2n+δ(3+R/r)Gn, and R‾n≤2RGdn+δ(3+R/r)Gn\overline R_n\le 2RGd\sqrt n+\delta(3+R/r)GnRn​≤2RGdn​+δ(3+R/r)Gn for η=R/(Gdn)\eta=R/(Gd\sqrt n)η=R/(Gdn​).
  6. Theorem 6.2, first display: one-point OSGD has R‾n≤R2/η+(dL)2δ2ηn+δ(3+R/r)Gn\overline R_n\le R^2/\eta+\frac{(dL)^2}{\delta^2}\eta n+\delta(3+R/r)GnRn​≤R2/η+δ2(dL)2​ηn+δ(3+R/r)Gn for every 0<δ≤r0<\delta\le r0<δ≤r and η>0\eta>0η>0.

Significance

The n3/4n^{3/4}n3/4 bound shows that a single function value per round suffices for sublinear regret against any oblivious sequence of Lipschitz convex losses, with a forecaster whose only operations are a random perturbation and a Euclidean projection. The smoothing identity of Lemmas 6.1–6.2 is the basic tool of zeroth-order (derivative-free) optimization, used well beyond bandits, and Theorem 6.1 is the n\sqrt nn​ benchmark for two-point methods.

All results are proved in the source. To the best of current knowledge none is formalized: the related items of the Introduction to Online Convex Optimization series on Prove2Me (Hazan's Lemma 6.7 and Theorem 6.9) were formalized with missing hypotheses and are recorded as disproved. This mission produces machine-checked statements with every hypothesis explicit, and the formal infrastructure (sphere measure calculus, a projected stochastic gradient analysis) for later zeroth-order results.

Difficulty

Two steps resist a direct formal treatment. First, Lemma 6.1 is a divergence-theorem identity on the ball; Mathlib has the sphere measure and polar coordinates, but its divergence theorem covers boxes rather than balls, so differentiating the ball average in xxx requires either such a theorem or a direct argument about translates of the ball. Second, the regret analysis takes expectations of quantities that depend on the whole past: the iterate xtx_txt​ is a function of S1,…,St−1S_1,\dots,S_{t-1}S1​,…,St−1​, and unbiasedness E[g~t∣xt]=∇ℓ~t(xt)\mathbb E[\widetilde g_t\mid x_t]=\nabla\widetilde\ell_t(x_t)E[g​t​∣xt​]=∇ℓt​(xt​) holds only conditionally, via independence of StS_tSt​ from the past. A pathwise gradient-descent inequality must be combined with this conditional expectation round by round, with measurability of the projected iterates established along the way. The naive approach of treating the estimate as the true gradient of ℓt\ell_tℓt​ fails: it is a gradient of ℓ~t\widetilde\ell_tℓt​, and the gap is handled only by Eq. (6.2) and Lemma 6.3.

Formalization scope

Points are in EuclideanSpace ℝ (Fin d) with d≥1d\ge1d≥1; rounds are t=1,2,…t=1,2,\dotst=1,2,…, sums run over Finset.Icc 1 n. σ\sigmaσ is Mathlib's Measure.toSphere of Lebesgue measure; the uniform laws are normalized restrictions. Randomness lives on an arbitrary probability space; the directions StS_tSt​ are measurable, mutually independent (iIndepFun) and uniform on S\mathbb SS, and in Theorem 6.1 the pairs (St,Ct)(S_t,C_t)(St​,Ct​) are independent with CtC_tCt​ a fair sign independent of StS_tSt​. A run of OSGD is a predicate (start at 000, each iterate a Euclidean projection onto (1−δ/r)K(1-\delta/r)\mathcal K(1−δ/r)K), which determines the run uniquely, so the forecaster uses only observed values and its own randomness. The losses are Lipschitz, differentiable and convex on all of Rd\mathbb R^dRd; the bound ∣ℓt∣≤L|\ell_t|\le L∣ℓt​∣≤L is on K\mathcal KK, because a convex function bounded on Rd\mathbb R^dRd is constant. The minimum over K\mathcal KK is an infimum over the subtype K\mathcal KK, attained in every theorem.

Conventions and corrections, each stated in the item's Formalization Note:

  • Lemma 6.1 carries the factor 1/δ1/\delta1/δ that the printed statement omits and the proof contains (corrected misprint).
  • Theorem 6.1's second display prints η=R/(GDn)\eta=R/(GD\sqrt n)η=R/(GDn​) and a limit "for δ→0\delta\to0δ→0"; the item states R‾n≤2RGdn+δ(3+R/r)Gn\overline R_n\le 2RGd\sqrt n+\delta(3+R/r)GnRn​≤2RGdn​+δ(3+R/r)Gn for η=R/(Gdn)\eta=R/(Gd\sqrt n)η=R/(Gdn​) and every admissible δ\deltaδ, which implies the limit (corrected misprint).
  • Theorems 6.1 and 6.2 add 0<δ≤r0<\delta\le r0<δ≤r, which the proofs need for Xt±,X~t∈KX_t^\pm,\widetilde X_t\in\mathcal KXt±​,Xt​∈K; for the tuned δ\deltaδ of the goal it is a condition on nnn.
  • The goal adds G,L>0G,L>0G,L>0 and n≥1n\ge1n≥1, which its formulas for δ,η\delta,\etaδ,η need; the constant 444 is the book's rounding of 2⋅23/42\cdot2^{3/4}2⋅23/4 and is kept, as is the form R2/ηR^2/\etaR2/η.

The statements cannot be satisfied trivially: the run is pinned by its recursion, the losses are fixed before the randomness, the expectations are of bounded measurable functions (no zero-valued Bochner integrals), and the minimum is over the nonempty compact K\mathcal KK. Section 6.3 (Lemma 6.4, Theorem 6.3) is not included, because its algorithm box and proof use different stage lengths and its unimodality condition is stated on a smaller set than the proof uses.

Needed infrastructure: calculus of ball averages and sphere integrals, symmetry of the uniform sphere law, nonexpansiveness of projections onto closed convex sets, and conditional-expectation bookkeeping for adapted iterates. Each is reusable for zeroth-order optimization; contributions of any of them as separate lemmas are welcome.

Selected references

  • S. Bubeck, N. Cesa-Bianchi, Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems, Foundations and Trends in Machine Learning 5(1), 2012. arXiv:1204.5721v2, doi:10.1561/2200000024
  • A. Flaxman, A. Kalai, H. B. McMahan, Online convex optimization in the bandit setting: gradient descent without a gradient, SODA 2005. arXiv:cs/0408007
  • A. Agarwal, O. Dekel, L. Xiao, Optimal algorithms for online convex optimization with multi-point bandit feedback, COLT 2010. link
  • S. Bubeck, R. Eldan, Y. T. Lee, Kernel-based methods for bandit convex optimization, STOC 2017. arXiv:1607.03084
10 thms2 active usersReviewed
ProbabilityStatistics·Captain: mikedeng1

Certified Adversarial Robustness via Randomized Smoothing 1: The Gaussian-Smoothed Classifier Is Constant on the ℓ2 Ball of Radius (σ/2)(Φ⁻¹(p_A) − Φ⁻¹(p_B))Research Paper

Motivation

Classifiers trained on images, speech and text can be made to change their prediction by perturbations of the input that are tiny in norm (Szegedy et al., 2014). Empirical defences against such adversarial examples have repeatedly been broken by stronger attacks (Athalye, Carlini, Wagner, 2018), which motivates certified defences: classifiers that come with a proof that their prediction at a given input cannot change inside a stated ball around it.

Randomized smoothing turns any classifier, however large or opaque, into one with such a certificate in the ℓ2\ell_2ℓ2​ norm. It was introduced with weaker radii by Lecuyer et al. (2019) and Li et al. (2018). Cohen, Rosenfeld and Kolter (arXiv:1902.02918v2, ICML 2019) proved the radius that is now standard, and showed it cannot be enlarged. Their guarantee underlies most later work on certified ℓ2\ell_2ℓ2​ robustness, including Salman et al. (2019).

This mission formalizes the robustness guarantee, Theorem 1 of that paper. Page numbers below are PDF pages of the arXiv v2 preprint, which has no printed page numbers.

Setting

Inputs are points of Rd\mathbb R^dRd with the Euclidean norm ∥⋅∥\|\cdot\|∥⋅∥ and inner product δ⊤z\delta^\top zδ⊤z. Classes form a set Y\mathcal YY. A base classifier is a deterministic or random function f:Rd→Yf : \mathbb R^d \to \mathcal Yf:Rd→Y. A random fff is described by the probabilities P(f(z)=c)\mathbb P(f(z) = c)P(f(z)=c), which for each zzz form a probability distribution on Y\mathcal YY.

Fix a noise level σ>0\sigma > 0σ>0 and let ε∼N(0,σ2I)\varepsilon \sim \mathcal N(0, \sigma^2 I)ε∼N(0,σ2I) be isotropic Gaussian noise. The class probabilities at xxx are P(f(x+ε)=c)\mathbb P(f(x + \varepsilon) = c)P(f(x+ε)=c), and the smoothed classifier is

g(x)=arg⁡max⁡c∈YP(f(x+ε)=c).(1)g(x) = \arg\max_{c \in \mathcal Y} \mathbb P\big(f(x + \varepsilon) = c\big). \qquad (1)g(x)=argc∈Ymax​P(f(x+ε)=c).(1)

The paper leaves g(x)g(x)g(x) undefined when the maximizer is not unique. "g(x)=cg(x) = cg(x)=c" therefore means that every class other than ccc has strictly smaller probability.

Write Φ\PhiΦ for the standard Gaussian cumulative distribution function and Φ−1\Phi^{-1}Φ−1 for its inverse. Φ−1\Phi^{-1}Φ−1 is a real number on (0,1)(0,1)(0,1), and Φ−1(0)=−∞\Phi^{-1}(0) = -\inftyΦ−1(0)=−∞, Φ−1(1)=+∞\Phi^{-1}(1) = +\inftyΦ−1(1)=+∞.

Formalization targets

Goal: Theorem 1 (p. 4; restated p. 13)

Suppose that at a specific xxx there are a class cAc_AcA​ and numbers pA‾,pB‾∈[0,1]\underline{p_A}, \overline{p_B} \in [0,1]pA​​,pB​​∈[0,1] with

P(f(x+ε)=cA) ≥ pA‾ ≥ pB‾ ≥ max⁡c≠cAP(f(x+ε)=c).(6)\mathbb P\big(f(x + \varepsilon) = c_A\big) \ \ge\ \underline{p_A} \ \ge\ \overline{p_B} \ \ge\ \max_{c \ne c_A} \mathbb P\big(f(x + \varepsilon) = c\big). \qquad (6)P(f(x+ε)=cA​) ≥ pA​​ ≥ pB​​ ≥ c=cA​max​P(f(x+ε)=c).(6)

Then g(x+δ)=cAg(x + \delta) = c_Ag(x+δ)=cA​ for every δ\deltaδ with ∥δ∥2<R\|\delta\|_2 < R∥δ∥2​<R, where

R=σ2(Φ−1(pA‾)−Φ−1(pB‾)).(7)R = \frac{\sigma}{2}\Big(\Phi^{-1}(\underline{p_A}) - \Phi^{-1}(\overline{p_B})\Big). \qquad (7)R=2σ​(Φ−1(pA​​)−Φ−1(pB​​)).(7)

The statement covers every base classifier and every set of classes. The radius is infinite when pA‾=1>pB‾\underline{p_A} = 1 > \overline{p_B}pA​​=1>pB​​ or pA‾>0=pB‾\underline{p_A} > 0 = \overline{p_B}pA​​>0=pB​​.

Milestones

The milestones are the paper's own steps, in attack order:

  • Lemma 3 (p. 12): the Neyman–Pearson lemma in both directions, for densities μX\mu_XμX​, μY\mu_YμY​ on Rd\mathbb R^dRd and the likelihood-ratio sets {μY≤tμX}\{\mu_Y \le t\mu_X\}{μY​≤tμX​} and {μY≥tμX}\{\mu_Y \ge t\mu_X\}{μY​≥tμX​}.
  • The likelihood ratio and (5) (proof of Lemma 4, p. 13): for X∼N(x,σ2I)X \sim \mathcal N(x,\sigma^2 I)X∼N(x,σ2I) and Y∼N(x+δ,σ2I)Y \sim \mathcal N(x+\delta,\sigma^2 I)Y∼N(x+δ,σ2I), μY/μX=exp⁡(aδ⊤z+b)\mu_Y/\mu_X = \exp(a\delta^\top z + b)μY​/μX​=exp(aδ⊤z+b), so half-spaces orthogonal to δ\deltaδ are likelihood-ratio sets.
  • Lemma 4 (pp. 12–13): Neyman–Pearson for these two Gaussians and the half-spaces {δ⊤z≤β}\{\delta^\top z \le \beta\}{δ⊤z≤β}, {δ⊤z≥β}\{\delta^\top z \ge \beta\}{δ⊤z≥β}.
  • The four Claims of Appendix A.0.1 (pp. 15–16, with (13) and (14) of p. 14): the probabilities of the half-spaces A={z:δ⊤(z−x)≤σ∥δ∥Φ−1(pA‾)}A = \{z : \delta^\top(z-x) \le \sigma\|\delta\|\Phi^{-1}(\underline{p_A})\}A={z:δ⊤(z−x)≤σ∥δ∥Φ−1(pA​​)} and B={z:δ⊤(z−x)≥σ∥δ∥Φ−1(1−pB‾)}B = \{z : \delta^\top(z-x) \ge \sigma\|\delta\|\Phi^{-1}(1-\overline{p_B})\}B={z:δ⊤(z−x)≥σ∥δ∥Φ−1(1−pB​​)} under XXX and YYY:
P(X∈A)=pA‾,P(X∈B)=pB‾,P(Y∈A)=Φ(Φ−1(pA‾)−∥δ∥σ),P(Y∈B)=Φ(Φ−1(pB‾)+∥δ∥σ).\mathbb P(X \in A) = \underline{p_A},\quad \mathbb P(X \in B) = \overline{p_B},\quad \mathbb P(Y \in A) = \Phi\Big(\Phi^{-1}(\underline{p_A}) - \tfrac{\|\delta\|}{\sigma}\Big),\quad \mathbb P(Y \in B) = \Phi\Big(\Phi^{-1}(\overline{p_B}) + \tfrac{\|\delta\|}{\sigma}\Big).P(X∈A)=pA​​,P(X∈B)=pB​​,P(Y∈A)=Φ(Φ−1(pA​​)−σ∥δ∥​),P(Y∈B)=Φ(Φ−1(pB​​)+σ∥δ∥​).
  • (15) (p. 14): P(Y∈A)>P(Y∈B)\mathbb P(Y \in A) > \mathbb P(Y \in B)P(Y∈A)>P(Y∈B) if and only if ∥δ∥<R\|\delta\| < R∥δ∥<R.

Significance

Theorem 1 turns three numbers at one input into a guarantee over a whole ball: a lower bound on the top-class probability, an upper bound on the other classes, and the noise level. These bounds can be estimated by sampling and certified with confidence intervals (the paper's CERTIFY procedure). That is what lets randomized smoothing certify ImageNet-scale networks, where exact verification methods do not scale. The companion result (Theorem 2, a separate mission of this series) shows that no larger ℓ2\ell_2ℓ2​ ball can be certified from the same information.

The theorem has a short pen-and-paper proof; no machine-checked proof of it is recorded on Prove2Me. A formal development adds:

  • a checked Neyman–Pearson lemma for randomized tests with densities on Rd\mathbb R^dRd, which Mathlib does not have;
  • the Gaussian likelihood-ratio and half-space computations;
  • an explicit treatment of the endpoint cases pA‾=1\underline{p_A} = 1pA​​=1 and pB‾=0\overline{p_B} = 0pB​​=0, where the radius is infinite.

Difficulty

Every step is classical, so the difficulty lies in the missing infrastructure, not in the idea.

  • Densities. Mathlib's multivariate Gaussian is defined as a pushforward of a product measure, not by a density. Identifying it with the density (2πσ2)−d/2e−∥z−x∥2/(2σ2)(2\pi\sigma^2)^{-d/2}e^{-\|z-x\|^2/(2\sigma^2)}(2πσ2)−d/2e−∥z−x∥2/(2σ2), which Lemma 4 needs, is not available.
  • Projections. The Claims need the law of δ⊤X\delta^\top Xδ⊤X for X∼N(x,σ2I)X \sim \mathcal N(x,\sigma^2 I)X∼N(x,σ2I) in closed form, namely a one-dimensional Gaussian with mean δ⊤x\delta^\top xδ⊤x and variance σ2∥δ∥2\sigma^2\|\delta\|^2σ2∥δ∥2.
  • The inverse CDF. Mathlib has no normal quantile, so Φ−1\Phi^{-1}Φ−1 is defined here as an infimum. The identities Φ(Φ−1(p))=p\Phi(\Phi^{-1}(p)) = pΦ(Φ−1(p))=p and Φ−1(1−p)=−Φ−1(p)\Phi^{-1}(1-p) = -\Phi^{-1}(p)Φ−1(1−p)=−Φ−1(p) must be derived.
  • Degenerate cases. The obvious argument through the worst-case half-spaces breaks down when δ=0\delta = 0δ=0 or when pA‾\underline{p_A}pA​​ or pB‾\overline{p_B}pB​​ is 000 or 111. There the half-spaces are empty or everything and Φ−1\Phi^{-1}Φ−1 is infinite, so these cases need a separate argument.

Formalization scope

  • Space and noise. Rd\mathbb R^dRd is EuclideanSpace ℝ (Fin d). N(x,σ2I)\mathcal N(x,\sigma^2 I)N(x,σ2I) is the pushforward of Mathlib's stdGaussian under z↦x+σzz \mapsto x + \sigma zz↦x+σz, with σ>0\sigma > 0σ>0 a hypothesis.
  • Classifiers. A random classifier is a map f : ℝᵈ → PMF 𝒴 with measurable class probabilities; deterministic classifiers are point masses. Y\mathcal YY is an arbitrary type, with no finiteness assumed. The class probability is the published Gaussian smoothing RandomGradFree.Shared.smoothing, applied to z↦P(f(z)=c)z \mapsto \mathbb P(f(z) = c)z↦P(f(z)=c).
  • The prediction. "g(x)=cg(x) = cg(x)=c" is the strict unique-maximizer predicate, and ggg itself is not defined. Defining ggg by an arbitrary choice at ties would make the theorem depend on the tie-break.
  • The radius. RRR is computed in the extended reals with Φ−1(0)=−∞\Phi^{-1}(0) = -\inftyΦ−1(0)=−∞ and Φ−1(1)=+∞\Phi^{-1}(1) = +\inftyΦ−1(1)=+∞. A real-valued Φ−1\Phi^{-1}Φ−1 with junk value 000 at the endpoints would assign a finite, wrong radius there, so it is used only in milestones that assume 0<p<10 < p < 10<p<1. In the two corners pA‾=pB‾∈{0,1}\underline{p_A} = \overline{p_B} \in \{0,1\}pA​​=pB​​∈{0,1}, where (7) reads ∞−∞\infty - \infty∞−∞, the radius is −∞-\infty−∞ and the goal is vacuous, as in the paper.
  • Neyman–Pearson. Random tests are [0,1][0,1][0,1]-valued measurable functions, so the lemmas apply with h(z)=P(f(z)=c)h(z) = \mathbb P(f(z) = c)h(z)=P(f(z)=c) for random fff. The likelihood-ratio sets are written multiplied out (μY≤tμX\mu_Y \le t\mu_XμY​≤tμX​), which avoids division by zero where μX\mu_XμX​ vanishes.
  • Added hypotheses. The milestones about AAA and BBB assume δ≠0\delta \ne 0δ=0 and 0<p<10 < p < 10<p<1, which the paper's computation uses implicitly.

Contributions are welcome at every level: proofs of the milestones, and general lemmas such as the Gaussian density, the law of linear functionals of a Gaussian vector and properties of the normal quantile. These lemmas are reusable well beyond this mission.

Selected references

  • J. M. Cohen, E. Rosenfeld, J. Z. Kolter, Certified Adversarial Robustness via Randomized Smoothing, ICML 2019; arXiv:1902.02918v2. https://arxiv.org/abs/1902.02918
  • J. Neyman, E. S. Pearson, On the Problem of the Most Efficient Tests of Statistical Hypotheses, Phil. Trans. R. Soc. A 231, 1933. https://doi.org/10.1098/rsta.1933.0009
  • M. Lecuyer, V. Atlidakis, R. Geambasu, D. Hsu, S. Jana, Certified Robustness to Adversarial Examples with Differential Privacy, IEEE S&P 2019. https://arxiv.org/abs/1802.03471
  • B. Li, C. Chen, W. Wang, L. Carin, Certified Adversarial Robustness with Additive Noise, NeurIPS 2019. https://arxiv.org/abs/1809.03113
  • H. Salman et al., Provably Robust Deep Learning via Adversarially Trained Smoothed Classifiers, NeurIPS 2019. https://arxiv.org/abs/1906.04584
  • A. Athalye, N. Carlini, D. Wagner, Obfuscated Gradients Give a False Sense of Security, ICML 2018. https://arxiv.org/abs/1802.00420
  • C. Szegedy et al., Intriguing Properties of Neural Networks, ICLR 2014. https://arxiv.org/abs/1312.6199
15 thms2 active usersReviewed
🏆Completed
Convex OptimizationStatistics·Captain: mikedeng1

Stability and Generalization 3: Tikhonov Regularization in a Reproducing Kernel Hilbert Space Has Uniform Stability σ²κ²/(2λm)Research Paper

Motivation

Learning from a finite sample is useful only if changing the sample has a controlled effect on the learned predictor. Uniform stability asks for a bound on the change in loss at every test point when one training example is removed. Bousquet and Elisseeff use this property to obtain generalization bounds for learning algorithms, and identify regularization as a source of stability in methods built from reproducing kernels. The present mission isolates their result for a squared norm penalty in a reproducing kernel Hilbert space (RKHS). It concerns the sensitivity of the optimizer itself, before any probability bound on a random training sample is applied. The result is Theorem 22 of Bousquet and Elisseeff (2002).

Setting

Let XXX be an input space, YYY a label space, and HHH a real reproducing kernel Hilbert space of real-valued predictors on XXX. A kernel K:X×X→RK:X\times X\to\mathbb RK:X×X→R and a feature representative Φ(x)∈H\Phi(x)\in HΦ(x)∈H express the reproducing identity f(x)=⟨f,Φ(x)⟩Hf(x)=\langle f,\Phi(x)\rangle_Hf(x)=⟨f,Φ(x)⟩H​ and K(x,x′)=⟨Φ(x),Φ(x′)⟩HK(x,x')=\langle\Phi(x),\Phi(x')\rangle_HK(x,x′)=⟨Φ(x),Φ(x′)⟩H​. Thus K(x,x)=∥Φ(x)∥H2K(x,x)=\|\Phi(x)\|_H^2K(x,x)=∥Φ(x)∥H2​. The source assumes that all diagonal kernel values satisfy K(x,x)≤κ2K(x,x)\le\kappa^2K(x,x)≤κ2.

A labeled example is z=(x,y)∈X×Yz=(x,y)\in X\times Yz=(x,y)∈X×Y. Its loss under fff is ℓ(f,z)=c(f(x),y)\ell(f,z)=c(f(x),y)ℓ(f,z)=c(f(x),y), where ccc is a real-valued cost. Let DHD_HDH​ be the set of predictions that some element of HHH can produce at some input. The loss is σ\sigmaσ-admissible when c(⋅,y)c(\cdot,y)c(⋅,y) is convex for every label yyy and ∣c(a,y)−c(b,y)∣≤σ∣a−b∣|c(a,y)-c(b,y)|\le\sigma|a-b|∣c(a,y)−c(b,y)∣≤σ∣a−b∣ for all a,b∈DHa,b\in D_Ha,b∈DH​ and y∈Yy\in Yy∈Y. Here σ\sigmaσ is a nonnegative Lipschitz constant. The definition compares any two attainable predictions, even when they arise at different inputs.

Fix a sample S=(z1,…,zm)S=(z_1,\ldots,z_m)S=(z1​,…,zm​), a deleted index iii, and a regularization weight λ>0\lambda>0λ>0. The paper's full and truncated objectives, with squared RKHS norm regularization, are

Rr(g)=1m∑j=1mℓ(g,zj)+λ∥g∥H2,Rr∖i(g)=1m∑j≠iℓ(g,zj)+λ∥g∥H2.R_r(g)=\frac1m\sum_{j=1}^{m}\ell(g,z_j)+\lambda\|g\|_H^2, \qquad R_r^{\setminus i}(g)=\frac1m\sum_{j\ne i}\ell(g,z_j)+\lambda\|g\|_H^2.Rr​(g)=m1​j=1∑m​ℓ(g,zj​)+λ∥g∥H2​,Rr∖i​(g)=m1​j=i∑​ℓ(g,zj​)+λ∥g∥H2​.

The factor in both objectives is 1/m1/m1/m. Let fff and f∖if^{\setminus i}f∖i be minimizers of these respective objectives over all of HHH. The statements allow any minimizer satisfying the relevant global optimality condition; they do not choose one by an arbitrary fallback rule.

Formalization targets

The central target is the explicit deletion stability estimate of Theorem 22. For every test point z∈X×Yz\in X\times Yz∈X×Y,

∣ℓ(f,z)−ℓ(f∖i,z)∣≤σ2κ22λm.|\ell(f,z)-\ell(f^{\setminus i},z)| \le \frac{\sigma^2\kappa^2}{2\lambda m}.∣ℓ(f,z)−ℓ(f∖i,z)∣≤2λmσ2κ2​.

Its four milestones follow the paper's route through Lemma 20, the point-evaluation inequality (25), and the two quantitative inequalities displayed in the proof of Theorem 22. In particular, the intermediate RKHS distance bound is ∥f∖i−f∥H≤κσ/(2λm)\|f^{\setminus i}-f\|_H\le\kappa\sigma/(2\lambda m)∥f∖i−f∥H​≤κσ/(2λm) when κ≥0\kappa\ge0κ≥0. The main theorem uses κ2\kappa^2κ2, so it does not need a choice of sign for κ\kappaκ. These numerical constants are part of the target, rather than placeholders for unspecified bounds.

Significance

The theorem supplies a deterministic, uniform sensitivity estimate for kernel methods trained by squared norm regularization. The bound applies simultaneously to every test example and decreases as either the sample size or the regularization weight increases. It is one of the ingredients that lets the paper apply its earlier stability-to-generalization results to concrete learning procedures. The loss need not be bounded for this theorem; bounding it is a separate question addressed later in the paper.

The result is proved in the source article. This mission asks for a machine-checked version of its exact pairwise claim and the reusable infrastructure around it: the paper's admissibility condition, the two objectives, the general regularizer inequality, and the RKHS evaluation bound. The existing Prove2Me library already has definitions for loss, empirical error, and the RKHS reproducing identity, so the new definitions concentrate on what is specific to these pages. The related replace-one estimate in Mohri, Rostamizadeh and Talwalkar's Foundations of Machine Learning uses a different perturbation and constant; it is not interchangeable with this result.

Difficulty

The two minimizers solve different objectives, and the deleted example appears in only one of them. A comparison of their objective values alone does not directly give a bound on their distance in the Hilbert norm. The source also distinguishes an abstract convex class of functions in Lemma 20 from the full RKHS used in Theorem 22. A proof must keep those domains straight while preserving the precise normalization of the truncated objective. Another delicate point is that the kernel bound controls evaluations through the reproducing identity; a bound on K(x,x)K(x,x)K(x,x) is not by itself a bound on loss unless the admissibility condition is also used.

Formalization scope

The Lean model uses an abstract complete real inner product space HHH, an evaluation map ev⁡:H→(X→R)\operatorname{ev}:H\to(X\to\mathbb R)ev:H→(X→R), a feature map Φ:X→H\Phi:X\to HΦ:X→H, and the published IsRKHSOf predicate tying these to KKK. The completeness instance matches the source's Hilbert-space assumption. Samples have type Fin m → X × Y, so an index i : Fin m already forces m≥1m\ge1m≥1. Minimization ranges over the entire HHH for Theorem 22 and over the declared convex class for Lemma 20. The objective definitions use the published Loss and EmpiricalError objects. No probability measure or measurability assumption is needed for these deterministic assertions.

There is a printed mismatch that affects what “deletion” means. Theorem 22 names an algorithm defined by equation (26), which, run afresh on m−1m-1m−1 points, would normalize its data term by 1/(m−1)1/(m-1)1/(m−1). Lemma 20 and the proof of Theorem 22 instead compare the full objective with equation (20), whose data term uses 1/m1/m1/m. The formalized goal states that comparison, with its explicit constant, and records the discrepancy for audit. This excludes the tempting shortcut of treating the two normalizations as identical. The regularizer is the genuine squared norm and the second minimizer is required to minimize the genuine truncated objective; neither a restricted hypothesis ball nor an artificially assumed distance bound enters the goal. Contributions that establish minimizer existence or connect the pairwise bound to an algorithmic selection would extend this core without changing its statement.

Selected references

  • Olivier Bousquet and André Elisseeff, Stability and Generalization, Journal of Machine Learning Research 2 (2002), 499–526. Article and PDF.
  • Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar, Foundations of Machine Learning, second edition, MIT Press, 2018, Chapter 14. Book information.
9 thms2 active usersReviewed
Linear algebraReinforcement Learning·Captain: mikedeng1

Reinforcement Learning: An Introduction XI: Dutch Traces and the Equivalence of Forward and Backward Views in Monte Carlo LearningTextbook

Motivation

Eligibility traces are one of the basic mechanisms of reinforcement learning. A trace is a short-term memory vector ztz_tzt​ with one component per weight. It records which components contributed to recent value estimates, so that an error observed now can be credited to the right components without storing the past. Chapter 12 of Sutton and Barto's Reinforcement Learning: An Introduction (2nd ed., 2018) organizes the topic around two ways of describing an algorithm. A forward view updates each state toward a target built from rewards that arrive later. A backward view makes an update at every step from the current error and the trace.

The chapter proves one exact equivalence between the two views itself, in §12.6: "This is the only equivalence of forward- and backward-views that we explicitly demonstrate in this book" (p. 301). The setting is linear Monte Carlo prediction, and the backward view uses a dutch trace. The same trace appears in true online TD(λ\lambdaλ), whose exact equivalence to the online λ\lambdaλ-return algorithm (van Seijen and Sutton, 2014; van Seijen et al., 2016) the book cites without proof. Of that proof, §12.6 "gives some of the flavor ... but is much simpler" (p. 301).

The older equivalence of §§12.1–12.2 goes back to Sutton (1988). If the weights are held fixed during an episode, the summed updates of TD(λ\lambdaλ) with accumulating traces equal the summed updates of the off-line λ\lambdaλ-return algorithm. The book leaves it as Exercises 12.3–12.4.

Setting

Fix a dimension ddd, a step size α\alphaα and an episode of length T≥1T \ge 1T≥1 with feature vectors x0,…,xT−1∈Rdx_0, \dots, x_{T-1} \in \mathbb R^dx0​,…,xT−1​∈Rd. The episode ends with a single return G∈RG \in \mathbb RG∈R ("a single reward received at the end of the episode ... and ... no discounting", p. 301). The forward view is the linear gradient Monte Carlo, or LMS, rule (12.13): from an initial w0w_0w0​,

wt+1=wt+α [G−wt⊤xt] xt,0≤t<T.w_{t+1} = w_t + \alpha\,[G - w_t^\top x_t]\,x_t, \qquad 0 \le t < T .wt+1​=wt​+α[G−wt⊤​xt​]xt​,0≤t<T.

Let Ft=I−αxtxt⊤F_t = I - \alpha x_t x_t^\topFt​=I−αxt​xt⊤​ be the fading matrix. The backward view keeps two vectors that are updated at each step in O(d)O(d)O(d) time without knowledge of GGG. The dutch trace is z0=x0z_0 = x_0z0​=x0​, zt=zt−1+(1−αzt−1⊤xt) xtz_t = z_{t-1} + (1 - \alpha z_{t-1}^\top x_t)\,x_tzt​=zt−1​+(1−αzt−1⊤​xt​)xt​. The auxiliary vector is at=at−1−αxtxt⊤at−1a_t = a_{t-1} - \alpha x_t x_t^\top a_{t-1}at​=at−1​−αxt​xt⊤​at−1​, with a0=F0w0a_0 = F_0 w_0a0​=F0​w0​.

For the second part, an episode S0,R1,S1,…,RT,STS_0, R_1, S_1, \dots, R_T, S_TS0​,R1​,S1​,…,RT​,ST​ carries states and rewards, and v^(s,w)\hat v(s, w)v^(s,w) is a differentiable value function with v^(terminal,⋅)=0\hat v(\text{terminal}, \cdot) = 0v^(terminal,⋅)=0. For one fixed weight vector www, define the following, with γ∈[0,1]\gamma \in [0,1]γ∈[0,1] and λ∈[0,1)\lambda \in [0,1)λ∈[0,1):

  • the return GtG_tGt​;
  • the nnn-step return Gt:t+nG_{t:t+n}Gt:t+n​ (12.1), with Gt:t+n=GtG_{t:t+n} = G_tGt:t+n​=Gt​ once t+n≥Tt + n \ge Tt+n≥T;
  • the λ\lambdaλ-return Gtλ=(1−λ)∑n≥1λn−1Gt:t+nG^\lambda_t = (1-\lambda)\sum_{n \ge 1}\lambda^{n-1} G_{t:t+n}Gtλ​=(1−λ)∑n≥1​λn−1Gt:t+n​ (12.2);
  • the TD error δt=Rt+1+γv^(St+1,w)−v^(St,w)\delta_t = R_{t+1} + \gamma\hat v(S_{t+1}, w) - \hat v(S_t, w)δt​=Rt+1​+γv^(St+1​,w)−v^(St​,w) (12.6);
  • the accumulating trace z−1=0z_{-1} = 0z−1​=0, zt=γλzt−1+∇v^(St,w)z_t = \gamma\lambda z_{t-1} + \nabla\hat v(S_t, w)zt​=γλzt−1​+∇v^(St​,w) (12.5).

Formalization targets

Goal: the dutch-trace equivalence (12.14), corrected

wT=aT−1+αG zT−1.w_T = a_{T-1} + \alpha G\, z_{T-1}.wT​=aT−1​+αGzT−1​.

The left side is the forward view after TTT LMS updates. On the right, aT−1a_{T-1}aT−1​ and zT−1z_{T-1}zT−1​ are produced by the incremental recursions above. The goal is about the two algorithms, not only about the closed-form product identity.

Milestones on the goal's path (§12.6, p. 302)

  1. wt+1=Ftwt+αGxtw_{t+1} = F_t w_t + \alpha G x_twt+1​=Ft​wt​+αGxt​.
  2. wT=FT−1⋯F0w0+αG∑k=0T−1FT−1⋯Fk+1xkw_T = F_{T-1}\cdots F_0 w_0 + \alpha G \sum_{k=0}^{T-1} F_{T-1}\cdots F_{k+1} x_kwT​=FT−1​⋯F0​w0​+αG∑k=0T−1​FT−1​⋯Fk+1​xk​, the first line of (12.14).
  3. zt=∑k=0tFt⋯Fk+1xkz_t = \sum_{k=0}^{t} F_t \cdots F_{k+1} x_kzt​=∑k=0t​Ft​⋯Fk+1​xk​ for the dutch-trace recursion.
  4. at=Ft⋯F0w0a_t = F_t \cdots F_0 w_0at​=Ft​⋯F0​w0​ for the auxiliary-vector recursion (corrected initialization).

Milestones on the λ\lambdaλ-return (§§12.1–12.2)

  1. (12.3): Gtλ=(1−λ)∑n=1T−t−1λn−1Gt:t+n+λT−t−1GtG^\lambda_t = (1-\lambda)\sum_{n=1}^{T-t-1}\lambda^{n-1}G_{t:t+n} + \lambda^{T-t-1}G_tGtλ​=(1−λ)∑n=1T−t−1​λn−1Gt:t+n​+λT−t−1Gt​ for t<Tt < Tt<T.
  2. Exercise 12.1: Gtλ=Rt+1+γ[(1−λ)v^(St+1,w)+λGt+1λ]G^\lambda_t = R_{t+1} + \gamma[(1-\lambda)\hat v(S_{t+1}, w) + \lambda G^\lambda_{t+1}]Gtλ​=Rt+1​+γ[(1−λ)v^(St+1​,w)+λGt+1λ​].
  3. Exercise 12.3: Gtλ−v^(St,w)=∑k=tT−1(γλ)k−tδkG^\lambda_t - \hat v(S_t, w) = \sum_{k=t}^{T-1}(\gamma\lambda)^{k-t}\delta_kGtλ​−v^(St​,w)=∑k=tT−1​(γλ)k−tδk​.
  4. Exercise 12.4: ∑t<Tαδtzt=∑t<Tα[Gtλ−v^(St,w)]∇v^(St,w)\sum_{t<T}\alpha\delta_t z_t = \sum_{t<T}\alpha[G^\lambda_t - \hat v(S_t, w)]\nabla\hat v(S_t, w)∑t<T​αδt​zt​=∑t<T​α[Gtλ​−v^(St​,w)]∇v^(St​,w).

Significance

The goal says that an O(d)O(d)O(d)-per-step algorithm reproduces the Monte Carlo/LMS result exactly. That algorithm never stores the feature vectors or the TTT intermediate weight vectors. The book draws the conclusion that eligibility traces "are not specific to TD learning at all" (p. 303). The dutch trace in the case γλ=1\gamma\lambda = 1γλ=1 is the same object that true online TD(λ\lambdaλ) (12.11) uses for general γλ\gamma\lambdaγλ. The fading-matrix products and their incremental forms are therefore the vocabulary of any later formalization of true online TD(λ\lambdaλ) and of the online λ\lambdaλ-return algorithm.

Exercises 12.3–12.4 are the fixed-weight equivalence of TD(λ\lambdaλ) and the off-line λ\lambdaλ-return algorithm. They are the standard justification for calling TD(λ\lambdaλ) an approximation of the λ\lambdaλ-return algorithm. Exercise 12.1 and (12.3) are the identities the rest of the chapter uses to manipulate λ\lambdaλ-returns.

On status: all of these results are known and elementary on paper, and the book prints the derivation of (12.14). To our knowledge none of them has a machine-checked proof, and the platform has no statement about λ\lambdaλ-returns, eligibility traces or TD(λ\lambdaλ). What this mission adds is formal statements with every convention fixed, including one correction to the printed text. It also adds reusable definitions of nnn-step returns, λ\lambdaλ-returns and traces.

Difficulty

The algebra is elementary. The difficulty lies in the conventions, and a careless reading of the page produces a false statement. The book initializes a0=w0a_0 = w_0a0​=w0​, and taken literally that makes the goal false. The λ\lambdaλ-return is an infinite series, whose tail collapses only because every nnn-step return that reaches past termination equals the full return. That convention has to be built into the definition of Gt:t+nG_{t:t+n}Gt:t+n​, together with the terminal value v^(terminal,⋅)=0\hat v(\text{terminal}, \cdot) = 0v^(terminal,⋅)=0. Exercises 12.3 and 12.4 are true only when the weights stay fixed. With the algorithms' changing weights wtw_twt​ the nnn-step returns (12.1) use wt+n−1w_{t+n-1}wt+n−1​, and neither identity holds. The boundary indices (t=T−1t = T-1t=T−1, the empty product at k=T−1k = T-1k=T−1, GTλ=0G^\lambda_T = 0GTλ​=0) must come out right.

Formalization scope

Namespace SuttonBartoRL.Traces. Vectors of §12.6 are Fin d → ℝ, matrices Matrix (Fin d) (Fin d) ℝ, and xx⊤x x^\topxx⊤ is Matrix.vecMulVec x x. The ordered product fadeProd α x j t is Ft⋯FjF_t \cdots F_jFt​⋯Fj​, the identity when t<jt < jt<j. Feature sequences are indexed by N\mathbb NN; only x0,…,xT−1x_0, \dots, x_{T-1}x0​,…,xT−1​ enter. T≥1T \ge 1T≥1 is a hypothesis wherever T−1T - 1T−1 appears. The identities of §12.6 are stated for every real α\alphaα and GGG, a harmless strengthening of the book's positive step size.

For §§12.1–12.2, weights are EuclideanSpace ℝ (Fin d), ∇\nabla∇ is Mathlib's gradient, and each v^(s,⋅)\hat v(s,\cdot)v^(s,⋅) is assumed differentiable in Exercise 12.4. An episode is a length TTT, states and rewards. The value at time t≥Tt \ge Tt≥T is 000. (12.2) is a tsum over n≥0n \ge 0n≥0 of λnGt:t+n+1\lambda^n G_{t:t+n+1}λnGt:t+n+1​, and λ∈[0,1)\lambda \in [0,1)λ∈[0,1), the range the book gives with (12.2). Milestone 5 concludes summability, so the junk value of a divergent tsum cannot make it trivial. Exercises 12.3–12.4 take one weight binder w, used in every return, TD error and gradient. That is the book's fixed-www assumption, stated in the binders.

Correction. The printed initialization a0=w0a_0 = w_0a0​=w0​ (p. 302) contradicts the printed definition at≐Ft⋯F0w0a_t \doteq F_t\cdots F_0 w_0at​≐Ft​⋯F0​w0​ and (12.14). The counterexample is d=1d = 1d=1, T=1T = 1T=1, x0=1x_0 = 1x0​=1, α=1/2\alpha = 1/2α=1/2, w0=1w_0 = 1w0​=1, G=0G = 0G=0: the forward view gives w1=1/2w_1 = 1/2w1​=1/2, while a0+αGz0=1a_0 + \alpha G z_0 = 1a0​+αGz0​=1. The mission states the corrected result with a0=F0w0a_0 = F_0 w_0a0​=F0​w0​, equivalently the same recursion started from a−1=w0a_{-1} = w_0a−1​=w0​. The printed text is kept verbatim in the milestone.

A trivializing formalization is ruled out: the goal is not the closed-form identity with aT−1a_{T-1}aT−1​ and zT−1z_{T-1}zT−1​ defined as the products and sums. Those vectors are defined by their GGG-free incremental recursions, and the closed forms are separate milestones.

Out of scope: the equivalence of true online TD(λ\lambdaλ) and the online λ\lambdaλ-return algorithm (cited, p. 300), the truncated-return identity (12.10), the error bound (12.8), and all convergence claims. Proofs of the milestones, and reuse of the definitions in later missions on true online TD(λ\lambdaλ), are welcome.

Selected references

  • R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018, ISBN 9780262039246, Chapter 12, pp. 287–320. http://incompleteideas.net/book/the-book-2nd.html
  • R. S. Sutton, "Learning to predict by the methods of temporal differences", Machine Learning 3 (1988), 9–44. https://doi.org/10.1007/BF00115009
  • H. van Seijen and R. S. Sutton, "True online TD(λ)", Proceedings of ICML 2014, PMLR 32, 692–700. https://proceedings.mlr.press/v32/seijen14.html
  • H. van Seijen, A. R. Mahmood, P. M. Pilarski, M. C. Machado and R. S. Sutton, "True online temporal-difference learning", Journal of Machine Learning Research 17 (2016), 1–40. https://jmlr.org/papers/v17/15-599.html
11 thms2 active usersReviewed
OptimizationProbability·Captain: mikedeng1

Gradient Convergence in Gradient Methods with Errors II: With Zero-Mean Stochastic Errors, Almost Surely Either f(x_t) → −∞ or f(x_t) Converges and ∇f(x_t) → 0Research Paper

Motivation

Stochastic gradient methods minimize a function fff when only noisy estimates of its gradient are available: each step moves along a descent direction corrupted by random noise. They are the standard training algorithm for neural networks and the basic tool of stochastic approximation, and the question every user faces is what can be guaranteed when fff is nonconvex, possibly unbounded below, and the noise is allowed to grow with the gradient.

D. P. Bertsekas and J. N. Tsitsiklis, Gradient Convergence in Gradient Methods with Errors, SIAM J. Optim. 10(3):627–642, 2000 (DOI), answer this under minimal assumptions. Noise with variance growing in ∥∇f(xt)∥\|\nabla f(x_t)\|∥∇f(xt​)∥ had been handled for related methods (Poljak and Tsypkin 1973), but typically together with a lower bound on fff, under which f(xt)f(x_t)f(xt​) is approximately a supermartingale and the supermartingale convergence theorem applies (see the monographs of Kushner and Clark 1978; Benveniste, Métivier and Priouret 1990; Kushner and Yin 1996). Section 4 of the paper (p. 635) removes the lower bound: it proves that, with probability 1, either f(xt)→−∞f(x_t)\to-\inftyf(xt​)→−∞ or f(xt)f(x_t)f(xt​) converges and ∇f(xt)→0\nabla f(x_t)\to 0∇f(xt​)→0, without assuming bounded iterates. Section 5 shows that the randomized incremental gradient method for a finite-sum objective is a special case. This mission formalizes Section 4 and the Section 5 application. A companion mission covers the deterministic counterpart (Proposition 1 of the same paper).

Setting

Let f:Rn→Rf:\mathbb R^n\to\mathbb Rf:Rn→R be continuously differentiable with a Lipschitz gradient: there is L≥0L\ge 0L≥0 with

∥∇f(x)−∇f(xˉ)∥≤L∥x−xˉ∥∀x,xˉ.(2.1)\|\nabla f(x)-\nabla f(\bar x)\|\le L\|x-\bar x\|\qquad\forall x,\bar x. \tag{2.1}∥∇f(x)−∇f(xˉ)∥≤L∥x−xˉ∥∀x,xˉ.(2.1)

Let (Ω,F,P)(\Omega,\mathcal F,P)(Ω,F,P) be a probability space and F0⊆F1⊆⋯\mathcal F_0\subseteq\mathcal F_1\subseteq\cdotsF0​⊆F1​⊆⋯ an increasing sequence of σ\sigmaσ-fields (a filtration; Ft\mathcal F_tFt​ is the history of the algorithm just before the noise wtw_twt​ is drawn). The stochastic gradient method generates random vectors by

xt+1=xt+γt(st+wt),x_{t+1}=x_t+\gamma_t(s_t+w_t),xt+1​=xt​+γt​(st​+wt​),

where γt>0\gamma_t>0γt​>0 is a deterministic stepsize, sts_tst​ is a descent direction and wtw_twt​ is a noise term. The assumptions of Proposition 3 are:

  • (a) xtx_txt​ and sts_tst​ are Ft\mathcal F_tFt​-measurable;
  • (b) there are c1,c2>0c_1,c_2>0c1​,c2​>0 with c1∥∇f(xt)∥2≤−∇f(xt)′stc_1\|\nabla f(x_t)\|^2\le-\nabla f(x_t)'s_tc1​∥∇f(xt​)∥2≤−∇f(xt​)′st​ and ∥st∥≤c2(1+∥∇f(xt)∥)\|s_t\|\le c_2(1+\|\nabla f(x_t)\|)∥st​∥≤c2​(1+∥∇f(xt​)∥) for all ttt; (4.1)
  • (c) for all ttt, with probability 1, E[wt∣Ft]=0E[w_t\mid\mathcal F_t]=0E[wt​∣Ft​]=0 (4.2) and E[∥wt∥2∣Ft]≤A(1+∥∇f(xt)∥2)E[\|w_t\|^2\mid\mathcal F_t]\le A(1+\|\nabla f(x_t)\|^2)E[∥wt​∥2∣Ft​]≤A(1+∥∇f(xt​)∥2) (4.3), with A>0A>0A>0 deterministic;
  • (d) ∑tγt=∞\sum_t\gamma_t=\infty∑t​γt​=∞ and ∑tγt2<∞\sum_t\gamma_t^2<\infty∑t​γt2​<∞.

The noise variance in (c) may grow quadratically with ∥∇f(xt)∥\|\nabla f(x_t)\|∥∇f(xt​)∥ and is therefore unbounded in general. A point xˉ\bar xxˉ is stationary if ∇f(xˉ)=0\nabla f(\bar x)=0∇f(xˉ)=0.

Formalization targets

Goal: Proposition 3 (p. 635)

Under (2.1) and (a)–(d), with probability 1,

f(xt)→−∞or(f(xt)→ℓ∈R  and  ∇f(xt)→0),f(x_t)\to-\infty\quad\text{or}\quad\Bigl(f(x_t)\to\ell\in\mathbb R\ \text{ and }\ \nabla f(x_t)\to 0\Bigr),f(xt​)→−∞or(f(xt​)→ℓ∈R  and  ∇f(xt​)→0),

and every limit point of (xt)(x_t)(xt​) is a stationary point of fff. The dichotomy is per sample path: different paths may take different branches.

Milestones

  1. (4.4), p. 636. The pathwise one-step inequality: if γ 2Lc22≤c1/2\gamma\,2Lc_2^2\le c_1/2γ2Lc22​≤c1​/2 and sss satisfies (4.1) at xxx, then for every www,
f(x+γ(s+w))≤f(x)−γc12∥∇f(x)∥2+γ∇f(x)′w+γ22Lc22+γ2L∥w∥2.f(x+\gamma(s+w))\le f(x)-\gamma\tfrac{c_1}{2}\|\nabla f(x)\|^2+\gamma\nabla f(x)'w+\gamma^2 2Lc_2^2+\gamma^2L\|w\|^2.f(x+γ(s+w))≤f(x)−γ2c1​​∥∇f(x)∥2+γ∇f(x)′w+γ22Lc22​+γ2L∥w∥2.
  1. Lemma 2, p. 637. If rtr_trt​ is Ft+1\mathcal F_{t+1}Ft+1​-measurable with E[rt∣Ft]=0E[r_t\mid\mathcal F_t]=0E[rt​∣Ft​]=0, E[∥rt∥2∣Ft]≤BE[\|r_t\|^2\mid\mathcal F_t]\le BE[∥rt​∥2∣Ft​]≤B and ∑γt2<∞\sum\gamma_t^2<\infty∑γt2​<∞, then ∑t≤Tγtrt\sum_{t\le T}\gamma_tr_t∑t≤T​γt​rt​ and ∑t≤Tγt2∥rt∥2\sum_{t\le T}\gamma_t^2\|r_t\|^2∑t≤T​γt2​∥rt​∥2 converge almost surely.
  2. Lemma 6, p. 640. For every δ>0\delta>0δ>0, almost surely f(xt)f(x_t)f(xt​) converges to a finite value or to −∞-\infty−∞, and if the limit is not −∞-\infty−∞ then lim sup⁡t∥∇f(xt)∥≤δ\limsup_t\|\nabla f(x_t)\|\le\deltalimsupt​∥∇f(xt​)∥≤δ.

Further result: §5, pp. 641–642

For f=1m∑ifif=\frac1m\sum_i f_if=m1​∑i​fi​ with Lipschitz gradients ∇fi\nabla f_i∇fi​ satisfying ∥∇fi(x)∥≤C+D∥∇f(x)∥\|\nabla f_i(x)\|\le C+D\|\nabla f(x)\|∥∇fi​(x)∥≤C+D∥∇f(x)∥ (5.2), the randomized incremental gradient method xt+1=xt−γt∇fk(t)(xt)x_{t+1}=x_t-\gamma_t\nabla f_{k(t)}(x_t)xt+1​=xt​−γt​∇fk(t)​(xt​), with independent uniform indices k(t)k(t)k(t), satisfies the conclusion of Proposition 3.

Significance

Proposition 3 is a convergence guarantee for stochastic gradient descent on smooth nonconvex objectives that needs neither a lower bound on fff, nor bounded iterates, nor bounded noise variance. It contains, as special cases, stochastic gradient descent with unbiased gradient estimates whose variance grows with the gradient, the randomized incremental (single-sample) gradient method for finite sums of Section 5, and scaled or approximate gradient directions through condition (4.1). Its conclusion is the strongest one available at this generality: if f(xt)f(x_t)f(xt​) stays bounded below along a path, then the gradient vanishes along that path and every limit point is stationary.

The result is proved in the paper; to our knowledge it has not been machine-checked. Its formalization requires a working theory of generalized conditional expectations of non-integrable noise, square-integrable martingales in Rn\mathbb R^nRn and pathwise arguments over random interval partitions, which is reusable for other stochastic approximation results (Robbins–Monro type schemes, TD-learning, stochastic subgradient methods).

Difficulty

The natural first idea is to view f(xt)f(x_t)f(xt​) as a supermartingale up to summable errors and apply the supermartingale convergence theorem (Robbins–Siegmund). This fails here: the theorem needs f(xt)f(x_t)f(xt​) bounded below, and fff is not assumed bounded below; in addition the noise term γt2L∥wt∥2\gamma_t^2L\|w_t\|^2γt2​L∥wt​∥2 in (4.4) has conditional mean of order γt2∥∇f(xt)∥2\gamma_t^2\|\nabla f(x_t)\|^2γt2​∥∇f(xt​)∥2, which is not summable when the gradient is unbounded. Any argument must therefore extract a decrease of fff that dominates noise of the same order as the gradient itself, without a lower bound to anchor a supermartingale, and must do so along every sample path while the hypotheses are only conditional-expectation statements about non-integrable noise.

Formalization scope

  • Rn\mathbb R^nRn is EuclideanSpace ℝ (Fin n), ∇f\nabla f∇f is Mathlib's gradient f, fff is ContDiff ℝ 1, and the Lipschitz constant is L : ℝ≥0 with LipschitzWith L (gradient f) (the standing assumption (2.1), equivalent to the page's form).
  • The probability space is a Measure Ω with IsProbabilityMeasure, the σ\sigmaσ-fields are a Mathlib Filtration ℕ, and measurability in (a) is StronglyMeasurable[ℱ t]. The recursion and (4.1) hold on every sample path. Stepsizes are deterministic.
  • Conditional expectations of the noise. The page assumes no integrability of wtw_twt​, of ∥wt∥2\|w_t\|^2∥wt​∥2 or of f(xt)f(x_t)f(xt​), and none is added. Mathlib's condExp is 000 for non-integrable functions, so stating (4.2)–(4.3) with it would make them hold vacuously for any non-integrable noise; that encoding is ruled out. Instead (4.3) says that for every Ft\mathcal F_tFt​-measurable set SSS, ∫S∥wt∥2 dP≤∫SA(1+∥∇f(xt)∥2) dP\int_S\|w_t\|^2\,dP\le\int_S A(1+\|\nabla f(x_t)\|^2)\,dP∫S​∥wt​∥2dP≤∫S​A(1+∥∇f(xt​)∥2)dP (in [0,∞][0,\infty][0,∞]), and (4.2) says that ∫Swt dP=0\int_S w_t\,dP=0∫S​wt​dP=0 for every Ft\mathcal F_tFt​-measurable SSS on which wtw_twt​ is integrable. These are exactly the generalized conditional-expectation statements of the page.
  • In Lemma 2 the bound BBB is a constant, so Mathlib's condExp is used there, with integrability of ∥rt∥2\|r_t\|^2∥rt​∥2 stated explicitly; the page's hypothesis implies it. Lemma 2 is stated over any finite-dimensional real inner-product space, since the paper applies it to real and to vector-valued sequences.
  • ∑γt=∞\sum\gamma_t=\infty∑γt​=∞ is divergence of the partial sums; ∑γt2<∞\sum\gamma_t^2<\infty∑γt2​<∞ is Summable. Convergent random series (Lemma 2) are convergence of partial sums, not Summable, which would mean unconditional convergence. "lim sup⁡∥∇f(xt)∥≤δ\limsup\|\nabla f(x_t)\|\le\deltalimsup∥∇f(xt​)∥≤δ" is "for every δ′>δ\delta'>\deltaδ′>δ, eventually ∥∇f(xt)∥≤δ′\|\nabla f(x_t)\|\le\delta'∥∇f(xt​)∥≤δ′". Limit points are MapClusterPt.
  • In the §5 result the page's references to "section 4" and "(4.1)" are read as section 3 and condition (3.1), the indices k(t)k(t)k(t) run from t=0t=0t=0, x0x_0x0​ is deterministic, and the stepsizes are nonnegative as on the page.
  • Not stated: Lemma 3 (it needs the random interval construction of p. 636 as a definition), Lemmas 4–5 (steps that depend on the proof's own choice of ϵ\epsilonϵ), and the Remarks of §4.

Contributions welcome: a proof of Lemma 2 from Mathlib's martingale convergence theorems (Submartingale.exists_ae_tendsto_of_bdd), a proof of (4.4) from the descent lemma, a general bridge between the set-integral encoding of conditional expectations and Mathlib's condExp on localizing sets, and the interval construction behind Lemmas 3–6.

Selected references

  • D. P. Bertsekas and J. N. Tsitsiklis, Gradient Convergence in Gradient Methods with Errors, SIAM J. Optim. 10(3):627–642, 2000. https://doi.org/10.1137/S1052623497331063
  • B. T. Poljak and Y. Z. Tsypkin, Pseudogradient adaptation and training algorithms, Automat. Remote Control 12 (1973), 83–94.
  • H. J. Kushner and D. S. Clark, Stochastic Approximation Methods for Constrained and Unconstrained Systems, Springer, 1978.
  • H. J. Kushner and G. Yin, Stochastic Approximation Methods, Springer, 1996 (as cited in the paper).
  • A. Benveniste, M. Métivier and P. Priouret, Adaptive Algorithms and Stochastic Approximations, Springer, 1990.
  • D. P. Bertsekas and J. N. Tsitsiklis, Neuro-Dynamic Programming, Athena Scientific, 1996.
4 thms2 active usersReviewed
ProbabilityStatistics·Captain: mikedeng1

The Sample Complexity of Pattern Classification with Neural Networks: The Size of the Weights is More Important than the Size of the Network I: Fat-Shattering Margin Bound with d = fat_H(γ/16)Research Paper

Motivation

Classical generalization bounds for classifiers, built on the VC dimension, grow with the number of adjustable parameters. For neural networks this is at odds with practice: networks with many more weights than training examples often generalize well. Bartlett's 1998 paper (IEEE Trans. Inform. Theory 44(2), 525–536) explains part of this by measuring a real-valued classifier's confidence. If a hypothesis classifies most training examples correctly with a margin γ\gammaγ, its misclassification probability is controlled by a scale-sensitive dimension of the class at scale proportional to γ\gammaγ, not by its VC dimension. Later in the paper this yields bounds for networks with small weights that do not depend on the number of weights.

This mission formalizes the first of the paper's two main technical results, the margin bound of Theorem 2 (p. 527), together with the steps of its proof on pp. 527–528.

The fat-shattering dimension was introduced by Kearns and Schapire (JCSS 1994). Alon, Ben-David, Cesa-Bianchi and Haussler (J. ACM 1997) proved the scale-sensitive Sauer-type covering bound used here (Theorem 5 of the paper). Shawe-Taylor, Bartlett, Williamson and Anthony (IEEE Trans. Inform. Theory 1998) proved the zero-training-error version (Theorem 1 of the paper). Theorem 2 extends it to hypotheses that make margin errors on the training data.

Setting

Let XXX be a set and PPP a probability distribution on X×{−1,1}X\times\{-1,1\}X×{−1,1}. The threshold function is sgn⁡(α)=−1\operatorname{sgn}(\alpha)=-1sgn(α)=−1 for α<0\alpha<0α<0 and sgn⁡(α)=1\operatorname{sgn}(\alpha)=1sgn(α)=1 for α≥0\alpha\ge0α≥0. For a real-valued hypothesis hhh on XXX, the misclassification probability is er⁡P(h)=P{sgn⁡(h(x))≠y}\operatorname{er}_P(h)=P\{\operatorname{sgn}(h(x))\ne y\}erP​(h)=P{sgn(h(x))=y}. For a sample z=((x1,y1),…,(xm,ym))z=((x_1,y_1),\dots,(x_m,y_m))z=((x1​,y1​),…,(xm​,ym​)) drawn independently from PPP and γ>0\gamma>0γ>0, the margin error estimate is

er⁡^zγ(h)=1m ∣{i:yih(xi)<γ}∣.\widehat{\operatorname{er}}{}^{\gamma}_z(h)=\tfrac1m\,|\{i : y_ih(x_i)<\gamma\}|.erzγ​(h)=m1​∣{i:yi​h(xi​)<γ}∣.

Let HHH be a class of real functions on XXX. Points x1,…,xmx_1,\dots,x_mx1​,…,xm​ are γ\gammaγ-shattered by HHH if some r∈Rmr\in\mathbb R^mr∈Rm has the following property: for every sign vector b∈{−1,1}mb\in\{-1,1\}^mb∈{−1,1}m, some h∈Hh\in Hh∈H satisfies (h(xi)−ri)bi≥γ(h(x_i)-r_i)b_i\ge\gamma(h(xi​)−ri​)bi​≥γ for all iii. The fat-shattering dimension fat⁡H(γ)\operatorname{fat}_H(\gamma)fatH​(γ) is the largest such mmm, possibly ∞\infty∞.

The proof uses the following objects:

  • the squashing function πγ(α)=max⁡(−γ,min⁡(γ,α))\pi_\gamma(\alpha)=\max(-\gamma,\min(\gamma,\alpha))πγ​(α)=max(−γ,min(γ,α)) and the class πγ(H)={πγ∘h:h∈H}\pi_\gamma(H)=\{\pi_\gamma\circ h:h\in H\}πγ​(H)={πγ​∘h:h∈H};
  • the sample ℓ∞\ell_\inftyℓ∞​ pseudometric dℓ∞(x)(f,g)=max⁡i∣f(xi)−g(xi)∣d_{\ell_\infty(x)}(f,g)=\max_i|f(x_i)-g(x_i)|dℓ∞​(x)​(f,g)=maxi​∣f(xi​)−g(xi​)∣;
  • the covering number N∞(F,ϵ,m)\mathcal N_\infty(F,\epsilon,m)N∞​(F,ϵ,m), the largest over x∈Xmx\in X^mx∈Xm of the size of the smallest ϵ\epsilonϵ-cover (Definition 3), and the corresponding packing number M∞(F,α,m)\mathcal M_\infty(F,\alpha,m)M∞​(F,α,m);
  • the quantization Qα(x)=⌈(x−α/2)/α⌉αQ_\alpha(x)=\lceil (x-\alpha/2)/\alpha\rceil\alphaQα​(x)=⌈(x−α/2)/α⌉α.

Formalization targets

Goal: Theorem 2

Assume 0<δ<1/20<\delta<1/20<δ<1/2, 0<γ<10<\gamma<10<γ<1, m≥1m\ge1m≥1, and d=fat⁡H(γ/16)d=\operatorname{fat}_H(\gamma/16)d=fatH​(γ/16) finite with d≤34md\le 34md≤34m. With probability at least 1−δ1-\delta1−δ over zzz, every h∈Hh\in Hh∈H satisfies

er⁡P(h)<er⁡^zγ(h)+2m(dln⁡34emdlog⁡2(578m)+ln⁡4δ).\operatorname{er}_P(h)<\widehat{\operatorname{er}}{}^{\gamma}_z(h)+\sqrt{\frac2m\Bigl(d\ln\frac{34em}{d}\log_2(578m)+\ln\frac4\delta\Bigr)} .erP​(h)<erzγ​(h)+m2​(dlnd34em​log2​(578m)+lnδ4​)​.

Milestones, in the order the proof uses them

  1. Lemma 4. er⁡P(h)<er⁡^zγ(h)+(2/m)ln⁡(2N∞(πγ(H),γ/2,2m)/δ)\operatorname{er}_P(h)<\widehat{\operatorname{er}}{}^{\gamma}_z(h)+\sqrt{(2/m)\ln(2\mathcal N_\infty(\pi_\gamma(H),\gamma/2,2m)/\delta)}erP​(h)<erzγ​(h)+(2/m)ln(2N∞​(πγ​(H),γ/2,2m)/δ)​ uniformly over HHH, with probability at least 1−δ1-\delta1−δ.
  2. Theorem 5 (Alon et al.). If F:{1,…,n}→{1,…,b}F:\{1,\dots,n\}\to\{1,\dots,b\}F:{1,…,n}→{1,…,b} and fat⁡F(1)≤d\operatorname{fat}_F(1)\le dfatF​(1)≤d, then log⁡2N∞(F,2,n)<1+log⁡2(nb2)log⁡2∑i≤d(ni)bi\log_2\mathcal N_\infty(F,2,n)<1+\log_2(nb^2)\log_2\sum_{i\le d}\binom ni b^ilog2​N∞​(F,2,n)<1+log2​(nb2)log2​∑i≤d​(in​)bi, provided nnn is large enough.
  3. Writing F=Qγ/8(πγ(H))F=Q_{\gamma/8}(\pi_\gamma(H))F=Qγ/8​(πγ​(H)): fat⁡F(γ/8)≤fat⁡πγ(H)(γ/16)\operatorname{fat}_F(\gamma/8)\le\operatorname{fat}_{\pi_\gamma(H)}(\gamma/16)fatF​(γ/8)≤fatπγ​(H)​(γ/16).
  4. M∞(πγ(H),γ/2,2m)≤M∞(F,γ/2,2m)\mathcal M_\infty(\pi_\gamma(H),\gamma/2,2m)\le\mathcal M_\infty(F,\gamma/2,2m)M∞​(πγ​(H),γ/2,2m)≤M∞​(F,γ/2,2m).
  5. N∞(πγ(H),γ/2,2m)≤N∞(F,γ/4,2m)\mathcal N_\infty(\pi_\gamma(H),\gamma/2,2m)\le\mathcal N_\infty(F,\gamma/4,2m)N∞​(πγ​(H),γ/2,2m)≤N∞​(F,γ/4,2m).
  6. log⁡2N∞(πγ(H),γ/2,2m)<1+dlog⁡2(34em/d)log⁡2(578m)\log_2\mathcal N_\infty(\pi_\gamma(H),\gamma/2,2m)<1+d\log_2(34em/d)\log_2(578m)log2​N∞​(πγ​(H),γ/2,2m)<1+dlog2​(34em/d)log2​(578m) when 1≤d≤2m1\le d\le 2m1≤d≤2m and m≥dlog⁡2(34em/d)+1m\ge d\log_2(34em/d)+1m≥dlog2​(34em/d)+1.
  7. fat⁡πγ(H)(γ/16)≤fat⁡H(γ/16)\operatorname{fat}_{\pi_\gamma(H)}(\gamma/16)\le\operatorname{fat}_H(\gamma/16)fatπγ​(H)​(γ/16)≤fatH​(γ/16).

A further item, Proposition 8 (p. 529), is the probabilistic device the paper uses to make such bounds uniform over γ\gammaγ.

Significance

Theorem 2 is the bound behind the paper's main message. Corollary 9 makes it uniform over γ\gammaγ, and Theorem 28 combines it with fat-shattering estimates for networks with bounded weights. Together they show that a network classifying the training data with a large margin generalizes at a rate governed by the size of its weights, not by its number of weights. The same template, a margin error plus a capacity term at scale γ\gammaγ, underlies later margin analyses of support vector machines and boosting.

All results in this mission are proved in the literature; none is open. None is machine-checked on this platform: the platform has Rademacher-complexity margin bounds, but no statement about fat-shattering dimension or ℓ∞\ell_\inftyℓ∞​ sample covering numbers of real-valued classes. A complete formalization would provide a reusable library of these objects, with their basic inequalities between squashing, quantization, packing and covering. It would also give a checked version of the explicit constants 34em/d34em/d34em/d and 578m578m578m, which differ from those in later textbook treatments.

Difficulty

The bound is uniform over a possibly uncountable class HHH, so a union bound over hypotheses does not apply. The obvious replacement is a union bound over a cover of HHH. Two steps make it hard:

  • Lemma 4. It needs a ghost-sample symmetrization and a random-swap argument, carried out with an ℓ∞\ell_\inftyℓ∞​ cover of the squashed class on the double sample, so the cover depends on the data.
  • Theorem 5. Bounding that covering number by the fat-shattering dimension is a combinatorial counting argument about strongly shattered pairs. It is the scale-sensitive analogue of the Sauer–Shelah lemma, and here the bookkeeping of constants is exact.

The quantization steps look routine but carry the factor-of-two losses that produce the constants γ/16\gamma/16γ/16, 171717 and 578578578.

Formalization scope

The model is in the namespace BartlettNN.Margin.

  • Labels and samples. Labels are Bool, read as ±1\pm1±1 through pm (true is +1+1+1). sgn⁡(0)=1\operatorname{sgn}(0)=1sgn(0)=1. Samples are functions Fin m → X × Bool, indexed from 000, with law Measure.pi (fun _ => P). The margin estimate uses the strict inequality yih(xi)<γy_ih(x_i)<\gammayi​h(xi​)<γ, and shattering uses ≥γ\ge\gamma≥γ.
  • Fat-shattering dimension. fat⁡\operatorname{fat}fat is valued in ℕ∞. A ℕ-valued supremum would be 000 on an unbounded set, so the goal assumes fat H (γ/16) = d with d : ℕ.
  • Covering and packing numbers. Covers are finite and external (centres are arbitrary functions), the cover inequality is strict, and covering numbers are ⊤ when no finite cover exists. N∞\mathcal N_\inftyN∞​ and M∞\mathcal M_\inftyM∞​ are suprema over all samples, with repetitions allowed. "α\alphaα-separated", which the paper leaves undefined, is read as distance ≥α\ge\alpha≥α.
  • Logarithms. ln⁡\lnln is Real.log, log⁡2\log_2log2​ is Real.logb 2, and eee is Real.exp 1.
  • High probability. "With probability at least 1−δ1-\delta1−δ, every hhh" bounds the measure of the event that some h∈Hh\in Hh∈H violates the inequality. It is not a per-hypothesis statement.

Measurability. The paper states "we ignore issues of measurability, and assume that all sets considered are measurable" (p. 526). This is made explicit, not removed, through three hypotheses:

  • every h∈Hh\in Hh∈H is measurable;
  • the bad events {z:∃h∈H, er⁡P(h)≥er⁡^zγ(h)+ϵ}\{z:\exists h\in H,\ \operatorname{er}_P(h)\ge\widehat{\operatorname{er}}{}^{\gamma}_z(h)+\epsilon\}{z:∃h∈H, erP​(h)≥erzγ​(h)+ϵ} are measurable;
  • the double-sample events of display (1) are measurable.

Replacing these by countability of HHH would weaken the theorem.

Corrections of the printed text.

  • Theorem 2. The goal adds d≤34md\le 34md≤34m. Beyond 34m34m34m the term dln⁡(34em/d)d\ln(34em/d)dln(34em/d) decreases, vanishes at d=34emd=34emd=34em and then turns negative, and the printed statement fails for rich classes. Within this range nothing is lost: the proof covers d≤2md\le2md≤2m, and for 2m<d≤34m2m<d\le34m2m<d≤34m the bound exceeds 111.
  • Milestone 6. It carries the hypothesis d≤2md\le 2md≤2m, the range of the binomial estimate behind 34em/d34em/d34em/d.
  • Milestone 3. Its printed justification ∣Qγ/8(a)−Qγ/8(b)∣<∣a−b∣+γ/16|Q_{\gamma/8}(a)-Q_{\gamma/8}(b)|<|a-b|+\gamma/16∣Qγ/8​(a)−Qγ/8​(b)∣<∣a−b∣+γ/16 is false; the correct term is γ/8\gamma/8γ/8. The milestone's conclusion is true as printed, and only the conclusion is formalized.

Trivializing formalizations, ruled out. The following would each make the statements empty or different, and none is used:

  • a ℕ-valued fat dimension or covering number;
  • Real.sign in place of sgn⁡\operatorname{sgn}sgn;
  • a per-hypothesis probability bound;
  • an unrestricted ddd, which makes ⋅\sqrt{\cdot}⋅​ of a negative number equal to 000;
  • a covering number that is 000 on classes without finite covers.

Infrastructure that a complete development needs, and contributions that are welcome:

  • product measures and Hoeffding's inequality, which Mathlib has;
  • a symmetrization (ghost-sample) lemma for margin events;
  • the combinatorics of Theorem 5;
  • the elementary inequalities between packing and covering numbers.

The covering/packing and fat-shattering lemmas apply beyond this mission. Proofs of individual milestones, or of Theorem 5 in the generality of Alon et al., are useful contributions in their own right.

Selected references

  • P. L. Bartlett, The Sample Complexity of Pattern Classification with Neural Networks: The Size of the Weights is More Important than the Size of the Network, IEEE Trans. Inform. Theory 44(2), 525–536, 1998. https://doi.org/10.1109/18.661502
  • N. Alon, S. Ben-David, N. Cesa-Bianchi, D. Haussler, Scale-sensitive dimensions, uniform convergence, and learnability, J. ACM 44(4), 615–631, 1997. https://doi.org/10.1145/263867.263927
  • J. Shawe-Taylor, P. L. Bartlett, R. C. Williamson, M. Anthony, Structural risk minimization over data-dependent hierarchies, IEEE Trans. Inform. Theory 44(5), 1926–1940, 1998. https://doi.org/10.1109/18.705570
  • M. J. Kearns, R. E. Schapire, Efficient distribution-free learning of probabilistic concepts, J. Comput. Syst. Sci. 48(3), 464–497, 1994. https://doi.org/10.1016/S0022-0000(05)80062-5
  • V. N. Vapnik, A. Ya. Chervonenkis, On the uniform convergence of relative frequencies of events to their probabilities, Theory Probab. Appl. 16(2), 264–280, 1971. https://doi.org/10.1137/1116025
12 thms2 active usersReviewed
Markov ChainReinforcement Learning·Captain: mikedeng1

Reinforcement Learning: An Introduction VII: The Error Reduction Property of n-step ReturnsTextbook

Motivation

Temporal-difference (TD) learning estimates the value of a policy by moving a current estimate toward a target built from observed rewards and from the estimate itself. Chapter 7 of Sutton and Barto's Reinforcement Learning: An Introduction (2nd ed., MIT Press, 2018) interpolates between the two extreme targets of the preceding chapters: the one-step TD target, which uses one reward and then bootstraps, and the Monte Carlo target, which uses every reward until the end of the episode. The intermediate target, the nnn-step return, uses nnn rewards and then bootstraps from the current estimate. The family underlies nnn-step TD, nnn-step Sarsa, the off-policy per-decision methods and the tree-backup algorithm, and it is the introduction to eligibility traces (Chapter 12).

The book justifies the whole family with one inequality, the error reduction property (7.3), p. 144: the expected nnn-step return is closer to the true value than the estimate it bootstraps from, by a factor γn\gamma^nγn in the worst state. It is the reason given for calling nnn-step TD methods "sound". The same chapter states, mostly as exercises without solutions, a series of exact identities that rewrite each kind of nnn-step return as a sum of one-step TD errors.

Setting

A finite Markov decision process has finite state and action sets S\mathcal SS, A\mathcal AA, a finite reward set R⊂R\mathcal R \subset \mathbb RR⊂R, and dynamics p(s′,r∣s,a)p(s', r \mid s, a)p(s′,r∣s,a), the probability of next state s′s's′ and reward rrr after action aaa in state sss (Eqs. (3.2)–(3.3)). A policy π(a∣s)\pi(a \mid s)π(a∣s) is a probability distribution over actions for each state. Following π\piπ from St=sS_t = sSt​=s produces a random trajectory At,Rt+1,St+1,At+1,Rt+2,…A_t, R_{t+1}, S_{t+1}, A_{t+1}, R_{t+2}, \dotsAt​,Rt+1​,St+1​,At+1​,Rt+2​,… For a discount factor 0≤γ<10 \le \gamma < 10≤γ<1 the state-value function is the expected discounted return (3.12),

vπ(s)=Eπ[∑k=0∞γkRt+k+1 ∣ St=s].v_\pi(s) = \mathbb E_\pi\Big[\sum_{k=0}^{\infty} \gamma^k R_{t+k+1} \,\Big|\, S_t = s\Big].vπ​(s)=Eπ​[k=0∑∞​γkRt+k+1​​St​=s].

Given any function V:S→RV : \mathcal S \to \mathbb RV:S→R (an estimate of vπv_\pivπ​), the nnn-step return (7.1) is

Gt:t+n=Rt+1+γRt+2+⋯+γn−1Rt+n+γnV(St+n).G_{t:t+n} = R_{t+1} + \gamma R_{t+2} + \cdots + \gamma^{n-1} R_{t+n} + \gamma^n V(S_{t+n}).Gt:t+n​=Rt+1​+γRt+2​+⋯+γn−1Rt+n​+γnV(St+n​).

In an episode that terminates at time TTT it is replaced by the complete return GtG_tGt​ when t+n≥Tt + n \ge Tt+n≥T. The TD error (6.5) is δk=Rk+1+γV(Sk+1)−V(Sk)\delta_k = R_{k+1} + \gamma V(S_{k+1}) - V(S_k)δk​=Rk+1​+γV(Sk+1​)−V(Sk​). Off-policy variants use a behavior policy bbb that generates the data, the per-decision ratio ρt=π(At∣St)/b(At∣St)\rho_t = \pi(A_t \mid S_t)/b(A_t \mid S_t)ρt​=π(At​∣St​)/b(At​∣St​), and the return with control variate (7.13), Gt:h=ρt(Rt+1+γGt+1:h)+(1−ρt)V(St)G_{t:h} = \rho_t(R_{t+1} + \gamma G_{t+1:h}) + (1-\rho_t) V(S_t)Gt:h​=ρt​(Rt+1​+γGt+1:h​)+(1−ρt​)V(St​), Gh:h=V(Sh)G_{h:h} = V(S_h)Gh:h​=V(Sh​). The tree-backup return (7.15)–(7.16) uses action values QQQ and the expected approximate value Vˉ(s)=∑aπ(a∣s)Q(s,a)\bar V(s) = \sum_a \pi(a \mid s) Q(s, a)Vˉ(s)=∑a​π(a∣s)Q(s,a) (7.8).

Formalization targets

Goal: the error reduction property (7.3)

For a finite MDP, a policy π\piπ, 0≤γ<10 \le \gamma < 10≤γ<1, any V:S→RV : \mathcal S \to \mathbb RV:S→R and every n≥1n \ge 1n≥1,

max⁡s∣Eπ[Gt:t+n∣St=s]−vπ(s)∣≤γnmax⁡s∣V(s)−vπ(s)∣.\max_s \big|\mathbb E_\pi[G_{t:t+n} \mid S_t = s] - v_\pi(s)\big| \le \gamma^n \max_s \big|V(s) - v_\pi(s)\big|.smax​​Eπ​[Gt:t+n​∣St​=s]−vπ​(s)​≤γnsmax​​V(s)−vπ​(s)​.

Milestones, in the book's order

  1. Exercise 7.1, p. 143: with VVV fixed and V(ST)=0V(S_T) = 0V(ST​)=0, Gt:t+n−V(St)=∑k=tmin⁡(t+n,T)−1γk−tδkG_{t:t+n} - V(S_t) = \sum_{k=t}^{\min(t+n,T)-1} \gamma^{k-t}\delta_kGt:t+n​−V(St​)=∑k=tmin(t+n,T)−1​γk−tδk​.
  2. Exercise 7.4, Eq. (7.6), p. 148: the nnn-step Sarsa return equals Qt−1(St,At)+∑k=tmin⁡(t+n,T)−1γk−t[Rk+1+γQk(Sk+1,Ak+1)−Qk−1(Sk,Ak)]Q_{t-1}(S_t, A_t) + \sum_{k=t}^{\min(t+n,T)-1} \gamma^{k-t}[R_{k+1} + \gamma Q_k(S_{k+1}, A_{k+1}) - Q_{k-1}(S_k, A_k)]Qt−1​(St​,At​)+∑k=tmin(t+n,T)−1​γk−t[Rk+1​+γQk​(Sk+1​,Ak+1​)−Qk−1​(Sk​,Ak​)], with estimates changing from step to step.
  3. Eq. (7.12), p. 150: Gt:h=Rt+1+γGt+1:hG_{t:h} = R_{t+1} + \gamma G_{t+1:h}Gt:h​=Rt+1​+γGt+1:h​ for t<h<Tt < h < Tt<h<T, Gh:h=V(Sh)G_{h:h} = V(S_h)Gh:h​=V(Sh​).
  4. Exercise 7.6, p. 151, for (7.13): under coverage, Eb\mathbb E_bEb​ of the control-variate return equals Eb\mathbb E_bEb​ of the same return without the control variate, and both equal Eπ[Gt:t+n∣St=s]\mathbb E_\pi[G_{t:t+n} \mid S_t = s]Eπ​[Gt:t+n​∣St​=s].
  5. Exercise 7.8, p. 151: Gt:h−V(St)=∑k=th−1γk−t(∏i=tkρi)δkG_{t:h} - V(S_t) = \sum_{k=t}^{h-1} \gamma^{k-t} \big(\prod_{i=t}^{k}\rho_i\big) \delta_kGt:h​−V(St​)=∑k=th−1​γk−t(∏i=tk​ρi​)δk​ for the return (7.13).
  6. Exercise 7.11, p. 153: the tree-backup return equals Q(St,At)+∑k=tmin⁡(t+n−1,T−1)δk∏i=t+1kγπ(Ai∣Si)Q(S_t, A_t) + \sum_{k=t}^{\min(t+n-1,T-1)} \delta_k \prod_{i=t+1}^{k} \gamma\pi(A_i \mid S_i)Q(St​,At​)+∑k=tmin(t+n−1,T−1)​δk​∏i=t+1k​γπ(Ai​∣Si​) with the expectation-based TD error δk=Rk+1+γVˉ(Sk+1)−Q(Sk,Ak)\delta_k = R_{k+1} + \gamma\bar V(S_{k+1}) - Q(S_k, A_k)δk​=Rk+1​+γVˉ(Sk+1​)−Q(Sk​,Ak​).

Significance

The result. The error reduction property makes the expected nnn-step target a γn\gamma^nγn-contraction toward vπv_\pivπ​ in the sup norm, uniformly over the estimate it starts from. It is the one-line reason the book offers for the soundness of every nnn-step TD method, and the same contraction is what the λ\lambdaλ-return of Chapter 12 averages over nnn. The TD-error identities are the algebra behind implementations that accumulate TD errors instead of storing returns, and behind the forward/backward-view equivalences of Chapter 12. Exercise 7.6 is the unbiasedness of the control-variate return, which is what allows (7.13) to replace plain importance weighting without changing the expected update.

Formalizing it. All results are elementary and well known, but the book gives no proofs: (7.3) is asserted, and the identities are exercises without published solutions. None is formalized on Prove2Me. The mission produces machine-checked versions with every hypothesis explicit (discounting, the fixed estimate, terminal values, coverage), and a trajectory-level expectation for finite MDPs that other chapters of the series can reuse.

Difficulty

The obvious proof of (7.3) is a matrix computation: Eπ[Gt:t+n∣St=s]−vπ(s)=γn(Pπn(V−vπ))(s)\mathbb E_\pi[G_{t:t+n} \mid S_t = s] - v_\pi(s) = \gamma^n (P_\pi^n (V - v_\pi))(s)Eπ​[Gt:t+n​∣St​=s]−vπ​(s)=γn(Pπn​(V−vπ​))(s), and a stochastic matrix does not increase the sup norm. The difficulty lies in the step before it. The left side is an expectation over trajectories, and vπv_\pivπ​ is an infinite discounted series; neither is a matrix power by definition. Connecting them requires a Chapman–Kolmogorov identity for the finite-trajectory distribution induced by π\piπ and p(s′,r∣s,a)p(s', r \mid s, a)p(s′,r∣s,a), the splitting of vπv_\pivπ​ at time nnn, and summability of the discounted series. Defining the expected nnn-step return as the matrix expression would reduce the goal to the last line and remove its content; that shortcut is ruled out below.

The TD-error identities are telescoping sums, but each has its own boundary: termination inside the nnn steps, the convention that terminal states have value zero, the index Q−1Q_{-1}Q−1​ at t=0t = 0t=0 in (7.6), the special case GT−1:t+n=RTG_{T-1:t+n} = R_TGT−1:t+n​=RT​ of the tree backup, and ratios with vanishing denominators in (7.13).

Formalization scope

  • Model. The finite MDP, policies and vπv_\pivπ​ follow the series conventions: dynamics p(s′,r∣s,a)p(s', r \mid s, a)p(s′,r∣s,a) over a finite reward set, one action set for all states, vπ(s)=∑kγk(Pπkrπ)(s)v_\pi(s) = \sum_k \gamma^k (P_\pi^k r_\pi)(s)vπ​(s)=∑k​γk(Pπk​rπ​)(s) computed from expected rewards and never defined as a Bellman solution.
  • Expectations are over trajectories. Eπ[ ⋅∣St=s]\mathbb E_\pi[\,\cdot \mid S_t = s]Eπ​[⋅∣St​=s] is the finite sum over nnn-step segments (At+k,St+k+1,Rt+k+1)k<n(A_{t+k}, S_{t+k+1}, R_{t+k+1})_{k<n}(At+k​,St+k+1​,Rt+k+1​)k<n​ weighted by ∏kπ(At+k∣St+k) p(St+k+1,Rt+k+1∣St+k,At+k)\prod_k \pi(A_{t+k}\mid S_{t+k})\,p(S_{t+k+1}, R_{t+k+1}\mid S_{t+k}, A_{t+k})∏k​π(At+k​∣St+k​)p(St+k+1​,Rt+k+1​∣St+k​,At+k​). The expected nnn-step return is not defined as ∑k<nγkPπkrπ+γnPπnV\sum_{k<n}\gamma^k P_\pi^k r_\pi + \gamma^n P_\pi^n V∑k<n​γkPπk​rπ​+γnPπn​V, which would make the goal a two-line matrix inequality.
  • The estimate is fixed. In the algorithm, Vt+n−1V_{t+n-1}Vt+n−1​ is the current random estimate. Every statement takes a fixed function VVV (or QQQ), which is the book's own reading ("if the value estimates don't change"). The only exception is Exercise 7.4, whose estimates QkQ_kQk​ are indexed by time k∈Zk \in \mathbb Zk∈Z exactly as in (7.6).
  • Discounting. The goal assumes 0≤γ<10 \le \gamma < 10≤γ<1 and takes the maximum over all states. Episodic tasks enter through absorbing zero-reward terminal states. The undiscounted episodic case γ=1\gamma = 1γ=1 is not stated.
  • Episodes. Sample-path identities use sequences Sk,Ak,RkS_k, A_k, R_kSk​,Ak​,Rk​ and a termination time TTT. The book's convention that terminal states have value 000 is a hypothesis (V(ST)=0V(S_T) = 0V(ST​)=0, Q(ST,⋅)=0Q(S_T, \cdot) = 0Q(ST​,⋅)=0).
  • Exercise 7.6 is stated for the state-value return (7.13) of p. 150, although the exercise follows the action-value return (7.14). Its conclusion includes, besides the literal "does not change the expected value", equality with the on-policy expected return, the property the book states on p. 150. Coverage (π(a∣s)>0⇒b(a∣s)>0\pi(a\mid s) > 0 \Rightarrow b(a\mid s) > 0π(a∣s)>0⇒b(a∣s)>0) is assumed.
  • Not stated. The convergence of nnn-step TD methods "under appropriate technical conditions" (p. 144), and the programming exercises.

Needed infrastructure: finite sums over function types, Chapman–Kolmogorov for the segment distribution, summability of ∑kγkPπkrπ\sum_k \gamma^k P_\pi^k r_\pi∑k​γkPπk​rπ​. The trajectory layer is reusable for the importance-sampling and eligibility-trace chapters. Alternative proofs of the goal, and proofs of the undiscounted episodic version as a separate theorem, are welcome.

Selected references

  • R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018, ISBN 9780262039246, Chapter 7, pp. 141–158. http://incompleteideas.net/book/the-book-2nd.html
  • C. J. C. H. Watkins, Learning from Delayed Rewards, PhD thesis, University of Cambridge, 1989 (the nnn-step return and its error reduction property, as credited on p. 158 of the book). https://www.cs.rhul.ac.uk/~chrisw/new_thesis.pdf
  • D. Precup, R. S. Sutton and S. Singh, Eligibility traces for off-policy policy evaluation, Proceedings of the 17th International Conference on Machine Learning (ICML), 2000, pp. 759–766 (the tree-backup algorithm, as credited on p. 158 of the book; no DOI).
10 thms2 active usersReviewed
Algorithmic Game TheoryMechanism Design·Captain: mikedeng1

How Much Data Is Sufficient to Learn High-Performing Algorithms? Generalization Guarantees for Data-Driven Algorithm Design 2: Neutral Affine Maximizers Have Pseudo-Dimension at Least ⌊n/2⌋Research Paper

Motivation

In data-driven algorithm design, an algorithm or mechanism comes with a vector of tunable parameters ρ\rhoρ, and the parameters are chosen by optimizing average performance over sample instances drawn from an unknown distribution. How many samples suffice for the empirical optimum to be near-optimal in expectation is governed by the pseudo-dimension of the class U={uρ}\mathcal U=\{u_\rho\}U={uρ​} of performance functions, where uρ(x)u_\rho(x)uρ​(x) is the performance of the parameter ρ\rhoρ on instance xxx. Balcan, DeBlasio, Dick, Kingsford, Sandholm and Vitercik (arXiv:1908.02894v4, 2021) give a general upper bound on Pdim(U)\mathrm{Pdim}(\mathcal U)Pdim(U) for classes whose dual functions are piecewise structured (their Theorem 3.3, the subject of mission 1 of this series), and apply it across integer programming, computational biology, clustering and mechanism design.

A general upper bound raises the question of whether it can be improved. Section 5 of the paper answers it for one application: neutral affine maximizers (NAMs), a family of voting mechanisms studied by Roberts (1979), Mishra and Sen (2012) and Nath and Sandholm (2019). For NAMs with nnn agents and mmm alternatives the general theorem gives Pdim(U)=O(nln⁡m)\mathrm{Pdim}(\mathcal U)=O(n\ln m)Pdim(U)=O(nlnm), and Theorem 5.2 shows a lower bound linear in nnn. So the sample complexity of tuning a NAM by sampling cannot be reduced below order nnn by a sharper analysis, and the general bound is tight up to logarithmic factors.

Setting

There are nnn agents and mmm alternatives. Agent iii has a value vi(j)∈Rv_i(j)\in\mathbb Rvi​(j)∈R for each alternative j∈[m]j\in[m]j∈[m]; a valuation profile is v=(v1,…,vn)∈Rnmv=(v_1,\dots,v_n)\in\mathbb R^{nm}v=(v1​,…,vn​)∈Rnm.

A NAM is specified by a weight vector ρ=(ρ[1],…,ρ[n])∈R≥0n\rho=(\rho[1],\dots,\rho[n])\in\mathbb R^n_{\ge0}ρ=(ρ[1],…,ρ[n])∈R≥0n​ in which at least one agent has weight zero; an agent with ρ[i]=0\rho[i]=0ρ[i]=0 is a sink agent. The NAM's outcome on a profile vvv is an alternative maximizing the weighted value,

ψρ(v)∈argmax⁡j∈[m] ∑i=1nρ[i] vi(j).\psi_\rho(v)\in\operatorname*{argmax}_{j\in[m]}\ \sum_{i=1}^n\rho[i]\,v_i(j).ψρ​(v)∈j∈[m]argmax​ i=1∑n​ρ[i]vi​(j).

Its utility is the social welfare of that outcome,

uρ(v)=∑i=1nvi(ψρ(v)),u_\rho(v)=\sum_{i=1}^n v_i\bigl(\psi_\rho(v)\bigr),uρ​(v)=i=1∑n​vi​(ψρ​(v)),

and the class studied is U={uρ∣ρ∈R≥0n, {i∣ρ[i]=0}≠∅}\mathcal U=\{u_\rho \mid \rho\in\mathbb R^n_{\ge0},\ \{i\mid\rho[i]=0\}\ne\emptyset\}U={uρ​∣ρ∈R≥0n​, {i∣ρ[i]=0}=∅}. NAMs also charge VCG-style payments that are redistributed to the sink agents; these do not enter uρu_\rhouρ​.

A class H\mathcal HH of real-valued functions on a set YYY shatters points y1,…,yNy_1,\dots,y_Ny1​,…,yN​ if there are thresholds z1,…,zN∈Rz_1,\dots,z_N\in\mathbb Rz1​,…,zN​∈R such that every pattern b∈{0,1}Nb\in\{0,1\}^Nb∈{0,1}N is realized by some h∈Hh\in\mathcal Hh∈H, in the sense that h(yℓ)>zℓh(y_\ell)>z_\ellh(yℓ​)>zℓ​ exactly when bℓ=1b_\ell=1bℓ​=1. The pseudo-dimension Pdim(H)\mathrm{Pdim}(\mathcal H)Pdim(H) is the largest NNN for which some NNN points are shattered. In Lean: a profile is v : Fin n → Fin m → ℝ, the outcome rule is ψ, the utility is welfare ψ ρ, the class is namClass ψ, and shattering is the published FoundationsML.Regression.Shatters.

Formalization targets

Goal: Theorem 5.2, corrected

For every n≥1n\ge1n≥1, every m≥2m\ge2m≥2 and every tie-breaking rule,

Pdim(U) ≥ ⌊n2⌋.\mathrm{Pdim}(\mathcal U)\ \ge\ \Bigl\lfloor\frac n2\Bigr\rfloor .Pdim(U) ≥ ⌊2n​⌋.

The printed statement reads Pdim(U)≥n/2\mathrm{Pdim}(\mathcal U)\ge n/2Pdim(U)≥n/2; see Formalization scope for why the floor is needed.

Milestones: the two claims of the proof (p. 24)

The proof fixes N=⌊n/2⌋N=\lfloor n/2\rfloorN=⌊n/2⌋ explicit profiles v(1),…,v(N)v^{(1)},\dots,v^{(N)}v(1),…,v(N) and, for each bit vector b∈{0,1}Nb\in\{0,1\}^Nb∈{0,1}N, an explicit weight vector ρ∈{0,1}n\rho\in\{0,1\}^nρ∈{0,1}n (both are definitions of this mission). The two milestones are the claims the proof makes about them: for every ℓ∈[N]\ell\in[N]ℓ∈[N] and ε∈(0,12)\varepsilon\in(0,\tfrac12)ε∈(0,21​),

bℓ=0 ⟹ uρ(v(ℓ))=ε,bℓ=1 ⟹ uρ(v(ℓ))=1.b_\ell=0\ \Longrightarrow\ u_\rho\bigl(v^{(\ell)}\bigr)=\varepsilon,\qquad b_\ell=1\ \Longrightarrow\ u_\rho\bigl(v^{(\ell)}\bigr)=1 .bℓ​=0 ⟹ uρ​(v(ℓ))=ε,bℓ​=1 ⟹ uρ​(v(ℓ))=1.

Significance

The result. Theorem 5.2 is the paper's evidence that its main upper bound cannot be improved in general by more than logarithmic factors: for NAMs the upper bound is O(nln⁡m)O(n\ln m)O(nlnm) and the lower bound is of order nnn. Through the standard link between pseudo-dimension and uniform convergence, it also means that any learner choosing NAM weights from samples needs a number of samples growing with the number of agents, whatever tie-breaking rule the mechanism uses.

Formalizing it. The theorem is proved in the paper; no machine-checked proof of it, or of any pseudo-dimension lower bound for a mechanism class, was found on Prove2Me. This mission produces a Lean model of NAM outcomes and welfare that does not fix a tie-breaking rule, an explicit lower-bound construction stated as definitions, and a pseudo-dimension lower bound in the vocabulary of the published FoundationsML pseudo-dimension items. It complements mission 1, which formalizes the matching upper-bound machinery.

Difficulty

The mathematical content is a single explicit construction, and the work lies in stating and verifying it at full generality rather than in a deep argument. Three points make the naive transcription wrong. First, the printed bound n/2n/2n/2 is false at n=1n=1n=1 and is not what the proof establishes for odd nnn; the even-nnn reduction must be replaced by a statement in ⌊n/2⌋\lfloor n/2\rfloor⌊n/2⌋. Second, the outcome ψρ(v)\psi_\rho(v)ψρ​(v) is defined by an argmax with unspecified tie-breaking, so the claims must hold for every maximizer; this requires showing the relevant maximizers are unique on the constructed profiles, not reading off a convenient one. Third, the construction embeds ⌊n/2⌋\lfloor n/2\rfloor⌊n/2⌋ indices into the nnn agents twice (as ℓ\ellℓ and ⌊n/2⌋+ℓ\lfloor n/2\rfloor+\ell⌊n/2⌋+ℓ), and the printed index condition for the second alternative is a typo; an off-by-one in this embedding silently breaks both claims. Finally, every constructed ρ\rhoρ must be admissible: it needs a sink agent, which holds only because N≥1N\ge1N≥1 or because nnn is odd and the last agent keeps weight 000.

Formalization scope

Representation. Agents are Fin n, alternatives Fin m, profiles Fin n → Fin m → ℝ, all 0-based: the paper's agent iii is i - 1, and its first and second alternatives are 0 and 1. Bits are Bool with true for 111. An outcome rule is any ψ : (Fin n → ℝ) → (Fin n → Fin m → ℝ) → Fin m with IsArgmaxSelector ψ, which requires ψ ρ v to maximize ∑ i, ρ i * v i j for every ρ and v; the theorem and both claims are stated for every such ψ. "Pdim(U)≥N\mathrm{Pdim}(\mathcal U)\ge NPdim(U)≥N" is the existence of an NNN-tuple of profiles shattered by namClass ψ in the sense of FoundationsML.Regression.Shatters, whose strict threshold t i < g (z i) shatters the same tuples as the paper's sign convention.

Corrections and readings of the printed text.

  1. The bound is ⌊n/2⌋\lfloor n/2\rfloor⌊n/2⌋ (natural-number division n / 2) instead of n/2n/2n/2. The proof assumes nnn even; at n=1n=1n=1 the only admissible ρ\rhoρ is 000, so U\mathcal UU is a single function and Pdim(U)=0<12\mathrm{Pdim}(\mathcal U)=0<\tfrac12Pdim(U)=0<21​.
  2. The hypotheses n≥1n\ge1n≥1 and m≥2m\ge2m≥2 are explicit. For n=0n=0n=0 no weight vector has a zero coordinate and U=∅\mathcal U=\emptysetU=∅; for m=1m=1m=1 the class is a single function. The proof takes m=2m=2m=2; the statement is for every m≥2m\ge2m≥2, with value 000 on every further alternative in the constructed profiles.
  3. The set-builder "{ρ[i]∣i=0}≠∅\{\rho[i]\mid i=0\}\ne\emptyset{ρ[i]∣i=0}=∅" in Theorem 5.2 is read as {i∣ρ[i]=0}≠∅\{i\mid\rho[i]=0\}\ne\emptyset{i∣ρ[i]=0}=∅, as written in Lemma 5.1.
  4. The construction's condition "ℓ=n/2+i\ell=n/2+iℓ=n/2+i" for vi(ℓ)(2)=εv_i^{(\ell)}(2)=\varepsilonvi(ℓ)​(2)=ε is read as i=n/2+ℓi=n/2+\elli=n/2+ℓ, as the paper's own n=6n=6n=6 example shows. The milestone texts are verbatim and keep the printed "vn/2+ℓ(ℓ)(1)v^{(\ell)}_{n/2+\ell}(1)vn/2+ℓ(ℓ)​(1)", which should read "(2)(2)(2)".

Ruled-out trivializations. Fixing one tie-breaking rule would state a special case and is excluded by quantifying over all argmax selectors. Dropping the sink-agent condition, or allowing n=0n=0n=0 or m=1m=1m=1, would change the class or make the statement vacuous or false; the conventions above exclude all three.

Infrastructure. Only Mathlib finite sums over Fin and the published Shatters definition are needed. The NAM model (IsNAMParam, IsArgmaxSelector, welfare, namClass) is reusable for any further statement about learning NAM parameters, including the matching O(nln⁡m)O(n\ln m)O(nlnm) upper bound once mission 1's general theorem is available. Proofs of the two claims, of the admissibility of the constructed weight vectors, and of the goal are all welcome.

Selected references

  • M.-F. Balcan, D. DeBlasio, T. Dick, C. Kingsford, T. Sandholm, E. Vitercik, How Much Data Is Sufficient to Learn High-Performing Algorithms? Generalization Guarantees for Data-Driven Algorithm Design, arXiv:1908.02894v4, 2021. https://arxiv.org/abs/1908.02894v4
  • K. Roberts, The characterization of implementable social choice rules, in J.-J. Laffont (ed.), Aggregation and Revelation of Preferences, North-Holland, 1979 (reference [86] of the paper).
  • D. Mishra, A. Sen, Roberts' theorem with neutrality: a social welfare ordering approach, Games and Economic Behavior 75(1):283–298, 2012 (reference [74] of the paper).
  • S. Nath, T. Sandholm, Efficiency and budget balance in general quasi-linear domains, Games and Economic Behavior 113:673–693, 2019 (reference [78] of the paper).
  • D. Pollard, Convergence of Stochastic Processes, Springer, 1984 (pseudo-dimension; reference [83] of the paper).
6 thms2 active usersReviewed
Dynamic ProgrammingMarkov ChainReinforcement Learning·Captain: mikedeng1

Reinforcement Learning: An Introduction II: The Bellman Optimality Equation and the Existence of an Optimal PolicyTextbook

Motivation

Sutton and Barto's Reinforcement Learning: An Introduction (2nd ed., MIT Press, 2018) is the standard introductory text of the field. Its Chapter 3 sets up the model that the rest of Part I works in: the finite Markov decision process (MDP), the value functions of a policy, and the Bellman equations that relate the value of a state to the values of its successors. Section 3.6 then states the fact every planning and control method of the book relies on: in a finite MDP there is an optimal policy, its value is the unique solution of a system of nonlinear equations, and a policy that acts greedily with respect to that solution is optimal. Dynamic programming (Chapter 4), Monte Carlo control (Chapter 5), Sarsa and Q-learning (Chapter 6) are all methods for solving the Bellman optimality equation; their correctness statements presuppose that it has exactly one solution and that it identifies optimal behaviour.

The book is deliberately informal ("we chose not to produce a rigorous formal treatment", p. xiii): §3.6 asserts these facts without proof. The results themselves are classical, going back to Bellman (1957), Howard (1960) and Blackwell (1965); a textbook proof for discounted finite MDPs is in Puterman, Markov Decision Processes (Wiley, 1994), Chapter 6.

Setting

A finite MDP has a finite set of states S\mathcal SS, a finite nonempty set of actions A\mathcal AA, a finite set of rewards R⊂R\mathcal R \subset \mathbb RR⊂R, and dynamics

p(s′,r∣s,a)=Pr⁡{St=s′,Rt=r∣St−1=s,At−1=a},∑s′∈S∑r∈Rp(s′,r∣s,a)=1,p(s', r \mid s, a) = \Pr\{S_t = s', R_t = r \mid S_{t-1} = s, A_{t-1} = a\}, \qquad \sum_{s' \in \mathcal S}\sum_{r \in \mathcal R} p(s', r \mid s, a) = 1,p(s′,r∣s,a)=Pr{St​=s′,Rt​=r∣St−1​=s,At−1​=a},s′∈S∑​r∈R∑​p(s′,r∣s,a)=1,

Eqs. (3.2)–(3.3). From ppp one derives p(s′∣s,a)=∑rp(s′,r∣s,a)p(s' \mid s, a) = \sum_r p(s', r \mid s, a)p(s′∣s,a)=∑r​p(s′,r∣s,a) and r(s,a)=∑rr∑s′p(s′,r∣s,a)r(s, a) = \sum_r r \sum_{s'} p(s', r \mid s, a)r(s,a)=∑r​r∑s′​p(s′,r∣s,a), Eqs. (3.4)–(3.5).

A policy π\piπ gives a probability π(a∣s)\pi(a \mid s)π(a∣s) of each action in each state. Fix a discount rate 0≤γ<10 \le \gamma < 10≤γ<1. The return of a reward sequence is Gt=∑k≥0γkRt+k+1G_t = \sum_{k \ge 0} \gamma^k R_{t+k+1}Gt​=∑k≥0​γkRt+k+1​ (3.8). The state-value function and action-value function of π\piπ are the expected returns

vπ(s)=Eπ[Gt∣St=s],qπ(s,a)=Eπ[Gt∣St=s,At=a](3.12)–(3.13).v_\pi(s) = \mathbb E_\pi[G_t \mid S_t = s], \qquad q_\pi(s, a) = \mathbb E_\pi[G_t \mid S_t = s, A_t = a] \qquad (3.12)\text{–}(3.13).vπ​(s)=Eπ​[Gt​∣St​=s],qπ​(s,a)=Eπ​[Gt​∣St​=s,At​=a](3.12)–(3.13).

In the Lean development these are stateValue M γ π s and actionValue M γ π s a, computed as ∑kγk(Pπkrπ)(s)\sum_k \gamma^k (P_\pi^k r_\pi)(s)∑k​γk(Pπk​rπ​)(s) from the transition matrix Pπ(s,s′)=∑aπ(a∣s) p(s′∣s,a)P_\pi(s, s') = \sum_a \pi(a \mid s)\,p(s' \mid s, a)Pπ​(s,s′)=∑a​π(a∣s)p(s′∣s,a) and the expected reward rπ(s)=∑aπ(a∣s) r(s,a)r_\pi(s) = \sum_a \pi(a \mid s)\,r(s, a)rπ​(s)=∑a​π(a∣s)r(s,a) of the Markov chain the policy induces. A policy π\piπ is optimal (IsOptimalPolicy) if vπ(s)≥vπ′(s)v_\pi(s) \ge v_{\pi'}(s)vπ​(s)≥vπ′​(s) for every policy π′\pi'π′ and every state sss. The optimal value functions are

v∗(s)=max⁡πvπ(s)(3.15),q∗(s,a)=max⁡πqπ(s,a)(3.16),v_*(s) = \max_\pi v_\pi(s) \quad (3.15), \qquad q_*(s, a) = \max_\pi q_\pi(s, a) \quad (3.16),v∗​(s)=πmax​vπ​(s)(3.15),q∗​(s,a)=πmax​qπ​(s,a)(3.16),

optimalValue and optimalActionValue, with the maximum over all stochastic policies.

Formalization targets

Goal: the Bellman optimality equation and optimal policies (§3.6, pp. 62–64)

For every finite MDP and 0≤γ<10 \le \gamma < 10≤γ<1:

  1. the maximum in (3.15) is attained at every state;
  2. an optimal policy exists;
  3. v∗v_*v∗​ satisfies the Bellman optimality equation
v∗(s)=max⁡a∑s′,rp(s′,r∣s,a)[r+γv∗(s′)]for all s;(3.19)v_*(s) = \max_{a} \sum_{s', r} p(s', r \mid s, a)\big[r + \gamma v_*(s')\big] \quad \text{for all } s; \qquad (3.19)v∗​(s)=amax​s′,r∑​p(s′,r∣s,a)[r+γv∗​(s′)]for all s;(3.19)
  1. v∗v_*v∗​ is the only function on S\mathcal SS satisfying (3.19);
  2. every policy that assigns positive probability only to actions attaining the maximum in (3.19) is optimal.

Milestones

  • (3.9) Gt=Rt+1+γGt+1G_t = R_{t+1} + \gamma G_{t+1}Gt​=Rt+1​+γGt+1​ for bounded rewards, with the series convergent.
  • (3.14) the Bellman equation vπ(s)=∑aπ(a∣s)∑s′,rp(s′,r∣s,a)[r+γvπ(s′)]v_\pi(s) = \sum_a \pi(a \mid s) \sum_{s', r} p(s', r \mid s, a)[r + \gamma v_\pi(s')]vπ​(s)=∑a​π(a∣s)∑s′,r​p(s′,r∣s,a)[r+γvπ​(s′)], and (p. 60) its uniqueness: vπv_\pivπ​ is its only solution.
  • Exercise 3.15 adding a constant ccc to all rewards adds vc=c/(1−γ)v_c = c/(1-\gamma)vc​=c/(1−γ) to every value.
  • Exercises 3.18 and 3.19 vπ(s)=∑aπ(a∣s) qπ(s,a)v_\pi(s) = \sum_a \pi(a \mid s)\,q_\pi(s, a)vπ​(s)=∑a​π(a∣s)qπ​(s,a) and qπ(s,a)=∑s′,rp(s′,r∣s,a)[r+γvπ(s′)]q_\pi(s, a) = \sum_{s', r} p(s', r \mid s, a)[r + \gamma v_\pi(s')]qπ​(s,a)=∑s′,r​p(s′,r∣s,a)[r+γvπ​(s′)].
  • (3.16)–(3.17) the maximum defining q∗q_*q∗​ is attained and q∗(s,a)=∑s′,rp(s′,r∣s,a)[r+γv∗(s′)]q_*(s, a) = \sum_{s', r} p(s', r \mid s, a)[r + \gamma v_*(s')]q∗​(s,a)=∑s′,r​p(s′,r∣s,a)[r+γv∗​(s′)].
  • (3.20) the Bellman optimality equation for action values, q∗(s,a)=∑s′,rp(s′,r∣s,a)[r+γmax⁡a′q∗(s′,a′)]q_*(s, a) = \sum_{s', r} p(s', r \mid s, a)[r + \gamma \max_{a'} q_*(s', a')]q∗​(s,a)=∑s′,r​p(s′,r∣s,a)[r+γmaxa′​q∗​(s′,a′)].

Significance

The goal is what turns "find a good policy" into "solve a system of equations". Parts 3 and 4 identify v∗v_*v∗​ with the unique solution of (3.19), so any procedure that finds a solution of (3.19) has found v∗v_*v∗​; part 5 converts v∗v_*v∗​ into an optimal policy by a one-step search. Parts 1 and 2 say that the book's definition (3.15) makes sense: a single policy is simultaneously best at every state, so "the optimal value function" is well defined and shared by all optimal policies. Chapter 4's policy iteration and value iteration, and the fixed points of Q-learning, are statements about this equation. Exercises 3.18 and 3.19 are used, by number, in the proof of the policy gradient theorem (p. 325).

The mathematics is classical and proved in many texts; what is missing is a machine-checked version in the book's own model. Platform relatives exist in different models: FoundationsML.ReinforcementLearning.bellman_equations_unique_solution (uniqueness for a fixed policy, with an expected-reward kernel instead of p(s′,r∣s,a)p(s', r \mid s, a)p(s′,r∣s,a)), BertsekasDP.discounted_main_theorem (cost minimization over deterministic stationary policies), BanditAlgorithm.mdp_discounted_bellman_solution (existence of a solution with a greedy deterministic policy, rewards in [0,1][0,1][0,1]) and FoundationsRL.RLBasics.bellman_optimality (finite horizon). None of them states the book's result: the four-argument dynamics, stochastic policies, the maximum over all of them, uniqueness of the solution of (3.19), and optimality of every policy supported on greedy actions. This mission produces that statement and, with it, a vocabulary of finite-MDP definitions that the later missions of the series reuse.

Difficulty

The book's derivation of (3.19) (p. 63) starts from v∗(s)=max⁡aqπ∗(s,a)v_*(s) = \max_a q_{\pi_*}(s, a)v∗​(s)=maxa​qπ∗​​(s,a) with a policy π∗\pi_*π∗​ that is optimal at every state at once. The existence of such a policy is the substance of the goal, and it does not follow from the definition: (3.15) takes a separate maximum at each state, and a priori the maximizing policy could depend on the state. Uniqueness for (3.19) is likewise not a consequence of linear algebra, as it is for (3.14): the equation is nonlinear because of the maximum. The fixed point must be related to the value of an actual policy, and every policy's value must be bounded above by it.

Formalization scope

  • Model. S and A are finite types with A nonempty (without an action, max⁡a\max_amaxa​ is undefined). One action set serves every state, as the book's footnote 3 (p. 48) allows. Rewards form a finite set M.R : Finset ℝ and the dynamics are the four-argument M.p s a s' r with the normalization (3.3).
  • Discounting. All statements assume 0≤γ<10 \le \gamma < 10≤γ<1 (the continuing discounted case of §3.3). The episodic case with γ=1\gamma = 1γ=1 is not covered: the book's uniqueness claims then need every episode to terminate under every policy, which the chapter never states, and without it (3.19) can have many solutions (a state that loops to itself with reward 000 satisfies v(s)=v(s)v(s) = v(s)v(s)=v(s) for any value).
  • Value functions from returns. vπv_\pivπ​ and qπq_\piqπ​ are expected discounted returns, computed from the Markov chain the policy induces. They are not defined as solutions of the Bellman equations, and v∗v_*v∗​ is not defined as a solution of (3.19): either would make the goal true by definition. The Bellman equations are theorems.
  • Maxima. v∗v_*v∗​ and q∗q_*q∗​ are real suprema over the type of stochastic policies (Lean gives a supremum that does not exist the value 000); the goal and milestone (3.17) assert that these suprema are attained, so they are the book's maxima.
  • Conditional expectations. (3.17), (3.18) and (3.20) are stated in their finite-sum form over (s′,r)(s', r)(s′,r).
  • Exercises. Exercises 3.15, 3.18 and 3.19 have no printed solutions; the statements give the formalization's answers (vc=c/(1−γ)v_c = c/(1-\gamma)vc​=c/(1−γ) and the two displayed identities).
  • Reusable infrastructure. The definitions (MDP, Policy, trans, expReward, policyTrans, policyReward, stateValue, actionValue, optimalValue, optimalActionValue, IsOptimalPolicy) follow the conventions shared by the whole series and are meant to be merged with the finite-MDP layers of the later chapters. Lemmas on summability of the value series, the Bellman operator as a γ\gammaγ-contraction in the sup norm, and the Markov-chain identities for PπkP_\pi^kPπk​ are welcome as separate contributions.

Selected references

  • R. S. Sutton, A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018, ISBN 9780262039246, Chapter 3. http://incompleteideas.net/book/the-book-2nd.html
  • R. Bellman, Dynamic Programming, Princeton University Press, 1957. https://doi.org/10.2307/j.ctv1nxcw0f
  • R. A. Howard, Dynamic Programming and Markov Processes, MIT Press, 1960.
  • D. Blackwell, "Discounted dynamic programming", Annals of Mathematical Statistics 36(1), 1965, 226–235. https://doi.org/10.1214/aoms/1177700285
  • M. L. Puterman, Markov Decision Processes: Discrete Stochastic Dynamic Programming, Wiley, 1994. https://doi.org/10.1002/9780470316887
11 thms2 active usersReviewed
Bandit AlgorithmsOperations ResearchProbability·Captain: mikedeng1

Analysis of Thompson Sampling for the Multi-armed Bandit Problem 1: Logarithmic Regret for Two ArmsResearch Paper

Motivation

Thompson Sampling (TS) is the oldest heuristic for the stochastic multi-armed bandit problem: it was proposed by Thompson in 1933 (Biometrika 25) and is used in practice for online advertising and recommendation, where it often performs as well as or better than upper-confidence-bound methods (Chapelle and Li, NIPS 2011; Scott 2010). Until 2012 its theoretical guarantees for the frequentist regret were weak: earlier analyses gave only o(T)o(T)o(T) regret in time TTT (Granmo 2010; May, Korda, Lee and Leslie 2011).

Agrawal and Goyal, Analysis of Thompson Sampling for the Multi-armed Bandit Problem (arXiv:1111.1797v3, COLT 2012), gave the first logarithmic finite-time bound on the expected regret of TS. This mission formalizes their two-armed result, Theorem 1. A companion mission covers the NNN-armed bound, Theorem 2.

Timeline. Lai and Robbins (1985) proved that every consistent algorithm has regret at least [∑iΔi/D(μi∥μ∗)+o(1)]ln⁡T\big[\sum_i \Delta_i/D(\mu_i\|\mu^*)+o(1)\big]\ln T[∑i​Δi​/D(μi​∥μ∗)+o(1)]lnT. Auer, Cesa-Bianchi and Fischer (2002) gave UCB1 with an O(∑iln⁡T/Δi)O(\sum_i \ln T/\Delta_i)O(∑i​lnT/Δi​) finite-time bound. Agrawal and Goyal (2012) proved O(ln⁡T/Δ+1/Δ3)O(\ln T/\Delta+1/\Delta^3)O(lnT/Δ+1/Δ3) for two-armed TS. Kaufmann, Korda and Munos (ALT 2012) and Agrawal and Goyal (AISTATS 2013) later proved asymptotically optimal bounds for Bernoulli TS.

Setting

There are two arms. Arm i∈{1,2}i\in\{1,2\}i∈{1,2} has a fixed, unknown reward distribution DiD_iDi​ supported in [0,1][0,1][0,1], with mean μi\mu_iμi​. Plays of an arm give i.i.d. rewards, independent of the other arm. Arm 1 is the unique optimal arm, μ1>μ2\mu_1>\mu_2μ1​>μ2​, and Δ=μ1−μ2\Delta=\mu_1-\mu_2Δ=μ1​−μ2​ is the gap.

Thompson Sampling for general stochastic bandits (Algorithm 2 of the paper) keeps, for each arm iii, a success count SiS_iSi​ and a failure count FiF_iFi​, both starting at 000. In each round t=1,2,…t=1,2,\dotst=1,2,… it

  1. samples, independently for each arm, θi(t)∼Beta(Si+1,Fi+1)\theta_i(t)\sim\mathrm{Beta}(S_i+1,F_i+1)θi​(t)∼Beta(Si​+1,Fi​+1);
  2. plays i(t)=arg⁡max⁡iθi(t)i(t)=\arg\max_i\theta_i(t)i(t)=argmaxi​θi​(t) and observes a reward r~t∼Di(t)\tilde r_t\sim D_{i(t)}r~t​∼Di(t)​;
  3. performs a Bernoulli trial with success probability r~t\tilde r_tr~t​, with outcome rt∈{0,1}r_t\in\{0,1\}rt​∈{0,1};
  4. increments Si(t)S_{i(t)}Si(t)​ if rt=1r_t=1rt​=1 and Fi(t)F_{i(t)}Fi(t)​ otherwise.

ki(t)k_i(t)ki​(t) is the number of plays of arm iii before round ttt. The expected regret in time TTT is

E[R(T)]=E[∑t=1T(μ1−μi(t))],\mathbb E[\mathcal R(T)]=\mathbb E\Big[\sum_{t=1}^T(\mu_1-\mu_{i(t)})\Big],E[R(T)]=E[t=1∑T​(μ1​−μi(t)​)],

the expectation being over the rewards and the algorithm's randomness.

The analysis uses the Beta cdf Fα,βbetaF^{beta}_{\alpha,\beta}Fα,βbeta​, the binomial cdf Fn,pBF^B_{n,p}Fn,pB​, and the random variable X(j,s,y)X(j,s,y)X(j,s,y): the number of independent Beta(s+1,j−s+1)\mathrm{Beta}(s+1,j-s+1)Beta(s+1,j−s+1) draws made before one exceeds yyy.

Formalization targets

Goal: Theorem 1 (p. 3)

There is an absolute constant C>0C>0C>0 such that for every two-armed instance with rewards in [0,1][0,1][0,1] and μ1>μ2\mu_1>\mu_2μ1​>μ2​, and every T≥2T\ge 2T≥2,

E[R(T)]≤C(ln⁡TΔ+1Δ3).\mathbb E[\mathcal R(T)]\le C\Big(\frac{\ln T}{\Delta}+\frac1{\Delta^3}\Big).E[R(T)]≤C(ΔlnT​+Δ31​).

The constant is not fixed numerically: the paper states the theorem in O(⋅)O(\cdot)O(⋅) form (footnote 1), and the explicit display it reports on p. 8, 40ln⁡T/Δ+48/Δ3+18Δ40\ln T/\Delta+48/\Delta^3+18\Delta40lnT/Δ+48/Δ3+18Δ, is not the formal claim.

Milestones

  • Fact 1 (p. 12): Fα,βbeta(y)=1−Fα+β−1,yB(α−1)F^{beta}_{\alpha,\beta}(y)=1-F^B_{\alpha+\beta-1,y}(\alpha-1)Fα,βbeta​(y)=1−Fα+β−1,yB​(α−1) for positive integers α,β\alpha,\betaα,β.
  • Lemma 1 (p. 6): E[X(j,s,y)]=1/Fj+1,yB(s)−1\mathbb E[X(j,s,y)]=1/F^B_{j+1,y}(s)-1E[X(j,s,y)]=1/Fj+1,yB​(s)−1.
  • Lemma 6 (p. 13): Hoeffding-type bounds (10)–(11) on binomial cdfs.
  • Fact 2 (p. 13): every median of Binomial(n,p)\mathrm{Binomial}(n,p)Binomial(n,p) is ⌊np⌋\lfloor np\rfloor⌊np⌋ or ⌈np⌉\lceil np\rceil⌈np⌉.
  • Lemma 2 (p. 7): Pr⁡(E2(t))≥1−2/T2\Pr(E_2(t))\ge 1-2/T^2Pr(E2​(t))≥1−2/T2, where E2(t)={θ2(t)≤μ2+Δ/2 or k2(t)<24ln⁡T/Δ2}E_2(t)=\{\theta_2(t)\le\mu_2+\Delta/2\ \text{or}\ k_2(t)<24\ln T/\Delta^2\}E2​(t)={θ2​(t)≤μ2​+Δ/2 or k2​(t)<24lnT/Δ2}.
  • Lemma 3 (p. 7): a three-case bound on E[E[min⁡{X(j,s(j),y),T}∣s(j)]]\mathbb E\big[\mathbb E[\min\{X(j,s(j),y),T\}\mid s(j)]\big]E[E[min{X(j,s(j),y),T}∣s(j)]] for s(j)∼Binomial(j,μ1)s(j)\sim\mathrm{Binomial}(j,\mu_1)s(j)∼Binomial(j,μ1​).
  • Eq. (1) (p. 7): E[k2(T)]≤C(ln⁡T/Δ2+1/Δ4)\mathbb E[k_2(T)]\le C(\ln T/\Delta^2+1/\Delta^4)E[k2​(T)]≤C(lnT/Δ2+1/Δ4).

Significance

The result. Theorem 1 shows that TS, a randomized Bayesian heuristic with no explicit confidence bonus, has regret logarithmic in TTT on every two-armed instance, matching the order in TTT of the Lai–Robbins lower bound. The proof introduced a way to control the optimal arm's waiting time between plays through the Beta–Binomial duality, and later analyses of TS reuse that device.

Formalizing it. The result is proved on paper and has no machine-checked proof that we know of. The platform's existing TS results concern Gaussian TS (Lattimore and Szepesvári, Ch. 36) and Bayesian regret, which are different algorithms or regret notions. A formalization adds a reusable Lean model of Algorithm 2 on [0,1][0,1][0,1]-valued rewards, Beta–Binomial facts (Fact 1, Lemma 1), a binomial-median theorem, and binomial Hoeffding bounds. It also produces a proof with a constant that has been checked, since the printed constants contain an arithmetic slip.

Difficulty

The standard UCB argument does not transfer to TS. For UCB, the optimal arm's index exceeds its mean with high probability however often the arm has been played, because the exploration bonus is deterministic; the analysis then only has to count plays of the suboptimal arm until its own index concentrates, after Θ(ln⁡T/Δ2)\Theta(\ln T/\Delta^2)Θ(lnT/Δ2) plays. Under TS the optimal arm's sample θ1(t)\theta_1(t)θ1​(t) is random and, if the arm has been played rarely or its early rewards were poor, it falls below μ2\mu_2μ2​ with constant probability. The optimal arm may then wait a long, random time between plays, and the length of that wait depends on the arm's posterior, which in turn depends on how long it has waited. Counting plays of the suboptimal arm with a union bound over rounds, under the assumption that the optimal arm is already concentrated, therefore does not work; controlling these waiting times is the central difficulty and is where the 1/Δ31/\Delta^31/Δ3 dependence enters.

Formalization scope

  • Model. The instance is the platform's StochasticBandit 2 (a probability measure on R\mathbb RR per arm, mean banditArmMean), with the hypothesis that each reward law gives mass 111 to [0,1][0,1][0,1]. Lean arm 0 is the paper's arm 1 and Lean arm 1 the paper's arm 2. Lean rounds are indexed from 000.
  • Algorithm. Algorithm 2 is realized on one probability space with three independent i.i.d. tables: Beta draws W(i,t,a,b)∼Beta(a+1,b+1)W(i,t,a,b)\sim\mathrm{Beta}(a+1,b+1)W(i,t,a,b)∼Beta(a+1,b+1), rewards X(i,t)∼DiX(i,t)\sim D_iX(i,t)∼Di​, and uniforms V(i,t)V(i,t)V(i,t). Round ttt uses θi(t)=W(i,t,Si(t),Fi(t))\theta_i(t)=W(i,t,S_i(t),F_i(t))θi​(t)=W(i,t,Si​(t),Fi​(t)), r~t=X(i(t),t)\tilde r_t=X(i(t),t)r~t​=X(i(t),t) and rt=1{V(i(t),t)<r~t}r_t=\mathbf 1\{V(i(t),t)<\tilde r_t\}rt​=1{V(i(t),t)<r~t​}. Ties in the arg max go to the smaller index (a null event).
  • Values. Regret and expectations are lower Lebesgue integrals in [0,∞][0,\infty][0,∞]. X(j,s,y)X(j,s,y)X(j,s,y) is N∪{∞}\mathbb N\cup\{\infty\}N∪{∞}-valued, so Lemma 1 at y=1y=1y=1 reads ∞=∞\infty=\infty∞=∞, as in the paper.
  • O(·). The paper's O(⋅)O(\cdot)O(⋅) (footnote 1: f≤cgf\le cgf≤cg for n≥n0n\ge n_0n≥n0​) is stated with one universal constant C>0C>0C>0, quantified before the instance, the means and the horizon, for all T≥2T\ge 2T≥2. Eq. (1) is stated the same way, without its printed numerals.
  • Not trivial. The goal is about Algorithm 2 itself, with fresh Beta samples, fresh rewards and the Bernoulli coin. A statement about "any policy satisfying Lemma 2's event bound", or one whose constant depends on Δ\DeltaΔ, the reward laws or TTT, would not be Theorem 1.
  • Edge cases. μ1<1\mu_1<1μ1​<1 is assumed only in Lemma 3, where the paper's RRR and DDD require it. It is not a hypothesis of the goal.
  • Infrastructure. A complete proof needs: inverse-transform or order-statistics facts for Beta laws (Fact 1); geometric expectations; Hoeffding's inequality for sums of Bernoulli variables (Mathlib has Hoeffding/Azuma); the binomial median theorem (Jogdeo–Samuels; Kaas–Buhrman); and the coupling from the reward tables to the per-arm i.i.d. output stacks the paper reasons with. Fact 1, Lemma 6 and Fact 2 are reusable beyond this mission. Contributions to any milestone are welcome, and so is a direct proof of the regret bound with an explicit constant.

Selected references

  • S. Agrawal and N. Goyal, Analysis of Thompson Sampling for the Multi-armed Bandit Problem, COLT 2012; arXiv:1111.1797v3. https://arxiv.org/abs/1111.1797
  • W. R. Thompson, On the likelihood that one unknown probability exceeds another in view of the evidence of two samples, Biometrika 25 (1933) 285–294. https://doi.org/10.2307/2332286
  • T. L. Lai and H. Robbins, Asymptotically efficient adaptive allocation rules, Advances in Applied Mathematics 6 (1985) 4–22. https://doi.org/10.1016/0196-8858(85)90002-8
  • P. Auer, N. Cesa-Bianchi and P. Fischer, Finite-time analysis of the multiarmed bandit problem, Machine Learning 47 (2002) 235–256. https://doi.org/10.1023/A:1013689704352
  • K. Jogdeo and S. M. Samuels, Monotone convergence of binomial probabilities and a generalization of Ramanujan's equation, Annals of Mathematical Statistics 39 (1968) 1191–1195. https://doi.org/10.1214/aoms/1177698243
  • R. Kaas and J. M. Buhrman, Mean, median and mode in binomial distributions, Statistica Neerlandica 34 (1980) 13–18. https://doi.org/10.1111/j.1467-9574.1980.tb00681.x
  • O. Chapelle and L. Li, An empirical evaluation of Thompson Sampling, NIPS 2011. https://papers.nips.cc/paper/4321-an-empirical-evaluation-of-thompson-sampling
11 thms2 active usersReviewed
Quantum InformationTheoretical Computer Science·Captain: mikedeng1

Shadow Tomography of Quantum States 1: Polylogarithmically Many Copies Suffice to Estimate Every Acceptance Probability to Within εResearch Paper

Motivation

Learning an unknown quantum state is expensive. Full quantum state tomography of a DDD-dimensional mixed state ρ\rhoρ to accuracy ε\varepsilonε in trace distance needs on the order of D2/ε2D^2/\varepsilon^2D2/ε2 copies of ρ\rhoρ (O'Donnell–Wright 2016; Haah et al. 2017), and this is optimal. For a system of nnn qubits, D=2nD = 2^nD=2n, so full tomography is out of reach beyond a few dozen qubits.

Often one does not need the whole density matrix, only the behaviour of ρ\rhoρ on a fixed list of tests: acceptance probabilities of verification circuits, expectation values of observables, or the answers a piece of quantum advice gives to a set of questions. Aaronson (arXiv:1711.01053, STOC 2018) named this task shadow tomography and asked whether the number of copies can be polylogarithmic in both the dimension and the number of tests. Measuring each test on separate copies costs O~(M/ε2)\tilde O(M/\varepsilon^2)O~(M/ε2) copies, which is linear in MMM.

Timeline.

  • 2016: the question was posed at a mini-course without a name (Aaronson, The Complexity of Quantum States and Transformations, §8.3.1).
  • 2016: Harrow, Lin and Montanaro gave a correct "quantum OR" test, repairing an earlier flawed claim (arXiv:1607.03236, Corollary 11).
  • 2017–2018: Aaronson proved the first polylogarithmic bound, the theorem of this mission.
  • Later work improved the exponents, notably Bădescu–O'Donnell 2021, and introduced the related "classical shadows" of Huang–Kueng–Preskill 2020.

Setting

A mixed state of dimension DDD is a D×DD\times DD×D Hermitian positive semidefinite matrix ρ\rhoρ with Tr ρ=1\mathrm{Tr}\,\rho = 1Trρ=1. A two-outcome measurement is a D×DD\times DD×D Hermitian matrix EEE with all eigenvalues in [0,1][0,1][0,1]. Equivalently, 0⪯E⪯10 \preceq E \preceq \mathbb 10⪯E⪯1. It accepts ρ\rhoρ with probability Tr(Eρ)\mathrm{Tr}(E\rho)Tr(Eρ).

The state ρ⊗k\rho^{\otimes k}ρ⊗k consists of kkk independent copies of ρ\rhoρ. A measurement of ρ⊗k\rho^{\otimes k}ρ⊗k with classical output is a POVM: a finite family of positive semidefinite matrices PωP_\omegaPω​ on the kkk-register space with ∑ωPω=1\sum_\omega P_\omega = \mathbb 1∑ω​Pω​=1. Outcome ω\omegaω occurs with probability Tr(Pωρ⊗k)\mathrm{Tr}(P_\omega\rho^{\otimes k})Tr(Pω​ρ⊗k). An adaptive procedure that measures the copies one after another is described by one such POVM.

Problem 1 (shadow tomography). Given an unknown ρ\rhoρ and known two-outcome measurements E1,…,EME_1,\dots,E_ME1​,…,EM​, output numbers b1,…,bM∈[0,1]b_1,\dots,b_M\in[0,1]b1​,…,bM​∈[0,1] with ∣bi−Tr(Eiρ)∣≤ε|b_i-\mathrm{Tr}(E_i\rho)|\le\varepsilon∣bi​−Tr(Ei​ρ)∣≤ε for all iii, with success probability at least 1−δ1-\delta1−δ. The output must come from a measurement of ρ⊗k\rho^{\otimes k}ρ⊗k, with k=k(D,M,ε,δ)k=k(D,M,\varepsilon,\delta)k=k(D,M,ε,δ) as small as possible. The measurement may depend on the EiE_iEi​, but not on ρ\rhoρ.

Formalization targets

Goal: Theorem 2, in the explicit form proved in §5

There is a universal constant CCC such that, for M≥2M\ge2M≥2 and 0<ε,δ≤1/20<\varepsilon,\delta\le 1/20<ε,δ≤1/2, Problem 1 is solvable with

k≤C log⁡Dε(log⁡log⁡D+log⁡1εε2)2log⁡4M(log⁡log⁡M+log⁡log⁡D+log⁡1ε+log⁡1δ)=O~(log⁡1/δε5log⁡4Mlog⁡D)k \le C\,\frac{\log D}{\varepsilon}\Big(\frac{\log\log D+\log\frac1\varepsilon}{\varepsilon^{2}}\Big)^{2}\log^4 M\Big(\log\log M+\log\log D+\log\frac1\varepsilon+\log\frac1\delta\Big) = \tilde O\Big(\frac{\log 1/\delta}{\varepsilon^5}\log^4 M\log D\Big)k≤CεlogD​(ε2loglogD+logε1​​)2log4M(loglogM+loglogD+logε1​+logδ1​)=O~(ε5log1/δ​log4MlogD)

copies. This is the last display of the proof (p. 19). The goal fixes no constant, so any improvement of CCC remains consistent with it.

Milestones

  • Theorem 13 (Harrow–Lin–Montanaro). A one-copy test that accepts with probability at least (1−ϵ)2/7(1-\epsilon)^2/7(1−ϵ)2/7 if some Tr(Eiρ)≥1−ϵ\mathrm{Tr}(E_i\rho)\ge1-\epsilonTr(Ei​ρ)≥1−ϵ, and at most 4ΔM4\Delta M4ΔM if ∑iTr(Eiρ)≤ΔM\sum_i\mathrm{Tr}(E_i\rho)\le\Delta M∑i​Tr(Ei​ρ)≤ΔM.
  • Lemma 14 (Quantum OR Bound). Deciding whether max⁡iTr(Eiρ)≥c\max_i\mathrm{Tr}(E_i\rho)\ge cmaxi​Tr(Ei​ρ)≥c or ≤c−ε\le c-\varepsilon≤c−ε with O(log⁡(1/δ)log⁡M/ε2)O(\log(1/\delta)\log M/\varepsilon^2)O(log(1/δ)logM/ε2) copies, independent of DDD.
  • Lemma 15 (Gentle Search). Finding jjj with Tr(Ejρ)≥c−ε\mathrm{Tr}(E_j\rho)\ge c-\varepsilonTr(Ej​ρ)≥c−ε with O(log⁡4Mε2(log⁡log⁡M+log⁡1δ))O(\frac{\log^4M}{\varepsilon^2}(\log\log M+\log\frac1\delta))O(ε2log4M​(loglogM+logδ1​)) copies.
  • Amplification claims (p. 16). The threshold tests Ei,t,±∗E^*_{i,t,\pm}Ei,t,±∗​ on ρ⊗q\rho^{\otimes q}ρ⊗q accept with probability at least 5/65/65/6 when the hypothesis is off by ε\varepsilonε, and at most 1/31/31/3 when it is within ε/2\varepsilon/2ε/2.
  • Markov claim (p. 17). The postselection test FtF_tFt​ on an arbitrary, possibly entangled, qqq-register state accepts with probability at most aq(a+ε/4)q\frac{a q}{(a+\varepsilon/4)q}(a+ε/4)qaq​.
  • Lemma 12 (Quantum Union Bound, probability part). Measurements each accepting with probability at least 1−ε1-\varepsilon1−ε all accept in succession with probability at least 1−2Mε1-2M\sqrt\varepsilon1−2Mε​.
  • Chernoff claim (p. 18). 1−Tr(Ftρ⊗q)≤ε4/log⁡2D1-\mathrm{Tr}(F_t\rho^{\otimes q})\le\varepsilon^4/\log^2D1−Tr(Ft​ρ⊗q)≤ε4/log2D.
  • Proposition 20. Promise-gap thresholds for all iii at once can be decided with O(log⁡(M/δ)/ε2)O(\log(M/\delta)/\varepsilon^2)O(log(M/δ)/ε2) copies.

Significance

The result. Theorem 2 shows that a state of exponential dimension can be learned "for all practical purposes" on exponentially many tests from polynomially many copies. Applications in the paper include a bound on quantum advice and one-way communication, and implications for quantum money and copy-protection. It also shows that the information needed to predict many measurement outcomes is far smaller than the description of ρ\rhoρ.

Formalizing it. The theorem is proved in the paper, and later work improves its exponents. As far as is known it has not been machine-checked. A complete development formalizes the gentle-measurement toolkit (Lemma 12, Lemma 14, Lemma 15), the amplification of two-outcome measurements on tensor powers, and the postselection argument. These are standard tools of quantum learning theory and quantum complexity with no formal counterpart yet. Lemma 14 and Lemma 15 are reusable beyond this mission.

Difficulty

The naive approach measures the EiE_iEi​ directly on shared copies. A measurement that is likely to reject disturbs the state, so later measurements see a damaged state, and separate copies per measurement cost MMM copies.

The proof needs three ingredients:

  • a gentle search that finds a measurement on which the current hypothesis is wrong while damaging the copies only slightly;
  • a potential argument showing that postselection cannot happen too often;
  • a uniform control of the damage.

The potential argument has to hold for the state after postselection, which is correlated or entangled across registers. Independence-based concentration fails there, which is why the Markov claim, not a Chernoff bound, governs that step. Theorem 13 itself rests on a delicate ancilla-based procedure of Harrow, Lin and Montanaro, and the mission cites it as a milestone without its proof.

Formalization scope

  • Representation.
    • Operators are complex matrices over a finite index type, and states use the published WildeQIT.IsDensityOperator (positive semidefinite, trace one).
    • A two-outcome measurement is IsEffect E: both EEE and 1−E\mathbb 1-E1−E are positive semidefinite.
    • ρ⊗k\rho^{\otimes k}ρ⊗k is a matrix indexed by kkk-tuples Fin k → n.
    • A measurement with output is a POVM structure with a finite outcome type. Probabilities are real parts of traces.
  • Quantifier order of the goal. ∃C\exists C∃C, then for all D,M,ε,δD,M,\varepsilon,\deltaD,M,ε,δ there is kkk; then for all EiE_iEi​ there are a POVM and outputs bbb; then for all ρ\rhoρ. Choosing the measurement after ρ\rhoρ would make the goal trivial (output the true values with k=0k=0k=0), and this order rules that out.
  • Disclosed hypotheses.
    • Theorem 2 assumes M≥2M\ge2M≥2, ε≤1/2\varepsilon\le1/2ε≤1/2 and δ≤1/2\delta\le1/2δ≤1/2. These keep the logarithmic factors positive; at M=1M=1M=1 the bound would force k=0k=0k=0.
    • Lemma 14 assumes M≥2M\ge2M≥2, and Lemmas 14 and 15 bound δ\deltaδ.
    • Theorem 13 assumes ϵ≤1/2\epsilon\le1/2ϵ≤1/2, as in Harrow–Lin–Montanaro's Corollary 11.
    • The Chernoff claim assumes D≥2D\ge2D≥2.
  • Conventions.
    • All logarithms are natural, including inside log⁡log⁡\log\logloglog.
    • Amplified tests use real thresholds.
    • "Applied in succession" in Lemma 12 uses Lüders instruments (E\sqrt{E}E​ Kraus operators), in the order E1,E2,…E_1,E_2,\dotsE1​,E2​,….
    • The hypothesis ρt\rho_tρt​ enters the amplification claims only as the number a=Tr(Eρt)a=\mathrm{Tr}(E\rho_t)a=Tr(Eρt​).
  • Printed steps not drafted.
    • The printed ε−4\varepsilon^{-4}ε−4 form of Theorem 2 relies on an external online-learning algorithm that is only sketched.
    • The halting rule of §5 is unspecified, because Lemma 15 always returns an index.
    • The asymptotic claims pt≥0.9/Dqp_t\ge0.9/D^qpt​≥0.9/Dq for t=o(log⁡2D/ε4)t=o(\log^2D/\varepsilon^4)t=o(log2D/ε4) and t=O(qlog⁡D/ε)t=O(q\log D/\varepsilon)t=O(qlogD/ε) use a circular o(⋅)o(\cdot)o(⋅).
    • The trace-distance part of Lemma 12 has an unquantified O(⋅)O(\cdot)O(⋅).
    • Lemma 12's printed bound 1−2Mε1-2M\sqrt\varepsilon1−2Mε​ is weaker than its use on p. 18. It is stated as printed. The proof of the goal must retune constants or use Wilde's stronger 1−2Mε1-2\sqrt{M\varepsilon}1−2Mε​-type bound.
  • Contributions welcome. Proofs of any milestone; a formal Hoeffding bound for binomial counts of product effects; the gentle measurement lemma for Lüders instruments; Naimark dilation for effects.

Selected references

  • S. Aaronson, Shadow Tomography of Quantum States, STOC 2018; arXiv:1711.01053v2, 2018. https://arxiv.org/abs/1711.01053
  • A. W. Harrow, C. Y.-Y. Lin, A. Montanaro, Sequential measurements, disturbance and property testing, SODA 2017. https://arxiv.org/abs/1607.03236
  • M. M. Wilde, Sequential decoding of a general classical-quantum channel, Proc. R. Soc. A, 2013. https://arxiv.org/abs/1303.0808
  • R. O'Donnell, J. Wright, Efficient quantum tomography, STOC 2016. https://arxiv.org/abs/1508.01907
  • J. Haah, A. W. Harrow, Z. Ji, X. Wu, N. Yu, Sample-optimal tomography of quantum states, IEEE Trans. Inf. Theory, 2017. https://arxiv.org/abs/1508.01797
  • C. Bădescu, R. O'Donnell, Improved quantum data analysis, STOC 2021. https://arxiv.org/abs/2011.10908
  • H.-Y. Huang, R. Kueng, J. Preskill, Predicting many properties of a quantum system from very few measurements, Nature Physics, 2020. https://arxiv.org/abs/2002.08953
16 thms2 active usersReviewed
Bandit AlgorithmsOptimizationReinforcement Learning·Captain: mikedeng1

Reinforcement Learning: An Introduction I: The Gradient Bandit Algorithm Is Stochastic Gradient AscentTextbook

Motivation

The multi-armed bandit is the simplest setting in which a learner must trade off exploiting what it knows against exploring what it does not: one situation, kkk actions, and a reward drawn from an unknown distribution each time an action is taken. Chapter 2 of Sutton and Barto's Reinforcement Learning: An Introduction (2nd ed., MIT Press, 2018) uses it to introduce, in the smallest possible setting, ideas that run through the rest of the book: incremental estimation with a step size, the bias introduced by the initial estimate, soft-max policies over learned preferences, and learning by following the gradient of expected reward.

The chapter ends with the gradient bandit algorithm (§2.8), which learns a numerical preference for each action instead of a value estimate. A shaded box on pp. 38–40 shows that its expected update is exactly a gradient-ascent step on the expected reward, so the algorithm is an instance of stochastic gradient ascent. The same argument, a score-function (likelihood-ratio) identity with a baseline, reappears in Chapter 13 as the REINFORCE algorithm and the policy gradient theorem. The bandit case is where the book first carries it out in full.

Setting

Actions are 1,…,k1, \dots, k1,…,k. Each action xxx has a reward distribution νx\nu_xνx​ on R\mathbb RR with finite mean q∗(x)q_*(x)q∗​(x), the true action value. At each step the learner holds a vector of action preferences H=(H(1),…,H(k))∈RkH = (H(1), \dots, H(k)) \in \mathbb R^kH=(H(1),…,H(k))∈Rk and selects action AAA with the soft-max probability

π(a)=eH(a)∑b=1keH(b)(2.11).\pi(a) = \frac{e^{H(a)}}{\sum_{b=1}^k e^{H(b)}} \qquad (2.11).π(a)=∑b=1k​eH(b)eH(a)​(2.11).

Given A=xA = xA=x, a reward R∼νxR \sim \nu_xR∼νx​ is received. The expected reward is E[R]=∑xπ(x) q∗(x)\mathbb E[R] = \sum_x \pi(x)\, q_*(x)E[R]=∑x​π(x)q∗​(x), a smooth function of HHH. With a step size α>0\alpha > 0α>0 and a baseline B∈RB \in \mathbb RB∈R, the gradient bandit update (2.12) is

H′(A)=H(A)+α(R−B)(1−π(A)),H′(a)=H(a)−α(R−B) π(a)  (a≠A).H'(A) = H(A) + \alpha (R - B)(1 - \pi(A)), \qquad H'(a) = H(a) - \alpha (R - B)\,\pi(a) \ \ (a \ne A).H′(A)=H(A)+α(R−B)(1−π(A)),H′(a)=H(a)−α(R−B)π(a)  (a=A).

The chapter's estimation sections use a single action's rewards R1,R2,…R_1, R_2, \dotsR1​,R2​,…. The sample average after n−1n-1n−1 selections is Qn=(R1+⋯+Rn−1)/(n−1)Q_n = (R_1 + \cdots + R_{n-1})/(n-1)Qn​=(R1​+⋯+Rn−1​)/(n−1), with an arbitrary initial value Q1Q_1Q1​. A constant step size α∈(0,1]\alpha \in (0,1]α∈(0,1] updates Qn+1=Qn+α[Rn−Qn]Q_{n+1} = Q_n + \alpha [R_n - Q_n]Qn+1​=Qn​+α[Rn​−Qn​] (2.5). The trace of one oˉ0=0\bar o_0 = 0oˉ0​=0, oˉn=oˉn−1+α(1−oˉn−1)\bar o_n = \bar o_{n-1} + \alpha (1 - \bar o_{n-1})oˉn​=oˉn−1​+α(1−oˉn−1​) defines the step size βn=α/oˉn\beta_n = \alpha / \bar o_nβn​=α/oˉn​ (2.8)–(2.9).

Formalization targets

Goal: the expected update is the gradient step

For every action aaa, with A∼πA \sim \piA∼π and R∣A=x∼νxR \mid A = x \sim \nu_xR∣A=x∼νx​,

E[H′(a)]=H(a)+α ∂ E[R]∂H(a),\mathbb E\bigl[H'(a)\bigr] = H(a) + \alpha\, \frac{\partial\, \mathbb E[R]}{\partial H(a)} ,E[H′(a)]=H(a)+α∂H(a)∂E[R]​,

that is, the update (2.12) equals the exact gradient-ascent step (2.13) in expected value, for every baseline BBB that does not depend on the selected action.

Milestones

  1. (2.3): Qn+1=Qn+1n[Rn−Qn]Q_{n+1} = Q_n + \tfrac1n [R_n - Q_n]Qn+1​=Qn​+n1​[Rn​−Qn​] for n≥1n \ge 1n≥1, including Q2=R1Q_2 = R_1Q2​=R1​ for arbitrary Q1Q_1Q1​.
  2. (2.6): Qn+1=(1−α)nQ1+∑i=1nα(1−α)n−iRiQ_{n+1} = (1-\alpha)^n Q_1 + \sum_{i=1}^n \alpha(1-\alpha)^{n-i} R_iQn+1​=(1−α)nQ1​+∑i=1n​α(1−α)n−iRi​, with weights summing to one.
  3. Exercise 2.7: with βn=α/oˉn\beta_n = \alpha/\bar o_nβn​=α/oˉn​, Qn+1=∑i=1nα(1−α)n−ioˉnRiQ_{n+1} = \sum_{i=1}^n \frac{\alpha(1-\alpha)^{n-i}}{\bar o_n} R_iQn+1​=∑i=1n​oˉn​α(1−α)n−i​Ri​ for n≥1n \ge 1n≥1, weights summing to one, and no dependence on Q1Q_1Q1​.
  4. Shift invariance (p. 37): adding a constant ccc to every preference leaves π\piπ unchanged.
  5. Exercise 2.9: for k=2k = 2k=2, π(1)=σ(H(1)−H(2))\pi(1) = \sigma(H(1) - H(2))π(1)=σ(H(1)−H(2)) with σ(x)=1/(1+e−x)\sigma(x) = 1/(1+e^{-x})σ(x)=1/(1+e−x).
  6. Soft-max derivative (p. 40): ∂π(x)/∂H(a)=π(x)(1a=x−π(a))\partial \pi(x)/\partial H(a) = \pi(x)(\mathbb 1_{a=x} - \pi(a))∂π(x)/∂H(a)=π(x)(1a=x​−π(a)).
  7. Zero-sum gradient (p. 39): ∑x∂π(x)/∂H(a)=0\sum_x \partial \pi(x)/\partial H(a) = 0∑x​∂π(x)/∂H(a)=0.
  8. Performance gradient as an expectation (p. 39): ∂E[R]/∂H(a)=E[(R−B)(1a=A−π(a))]\partial \mathbb E[R]/\partial H(a) = \mathbb E[(R - B)(\mathbb 1_{a=A} - \pi(a))]∂E[R]/∂H(a)=E[(R−B)(1a=A​−π(a))].

Significance

The result. The identity makes a model-free algorithm, which uses only the sampled action and reward, an unbiased estimator of the gradient of a quantity that depends on the unknown q∗q_*q∗​. It therefore places the gradient bandit algorithm within stochastic approximation, where convergence theory for stochastic gradient methods applies. It also explains the role of the baseline: any baseline independent of the action leaves the expected update unchanged, so the choice of baseline can only affect the variance of the update, as Figure 2.5 shows empirically. The estimation milestones make precise two claims the chapter uses repeatedly: sample averages can be maintained incrementally, and constant step sizes produce an exponentially recency-weighted average biased by Q1Q_1Q1​. Exercise 2.7 removes that bias.

Formalizing it. All of these results are elementary and proved (or left as routine exercises) in the book. None of them is formalized on Prove2Me or, as far as is known, in Mathlib. What this mission adds is a machine-checked version of the book's argument with the reward model and baseline condition stated precisely, and a reusable soft-max layer (definition, partial derivatives, shift invariance) for later missions of this series, in particular the policy gradient theorem of Chapter 13.

Difficulty

The mathematics is beginning calculus, as the book says. The formal difficulty lies elsewhere. The goal is an identity between an expectation over a two-stage random experiment (an action from π\piπ, then a reward from νA\nu_AνA​) and a partial derivative in one coordinate of a vector-valued parameter. A proof has to justify exchanging the finite sum with the derivative and splitting the reward integral, and it has to use integrability of each νx\nu_xνx​. It also needs the fact that the baseline term vanishes because ∑x∂π(x)/∂H(a)=0\sum_x \partial\pi(x)/\partial H(a) = 0∑x​∂π(x)/∂H(a)=0. A scalar-parameter version of the log-sum-exp derivative does not suffice: the book differentiates in one coordinate H(a)H(a)H(a) while all other preferences are held fixed. For Exercise 2.7 the obvious unrolling of (2.6) does not apply directly, because the step size βn\beta_nβn​ varies with nnn and the book states neither the weights nor the range of α\alphaα.

Formalization scope

  • Actions are Fin k. Every statement quantifies over some action, so k≥1k \ge 1k≥1 whenever it has content. Preferences are vectors Fin k → ℝ. The partial derivative in coordinate aaa is the derivative of h↦f(update H a h)h \mapsto f(\text{update } H\ a\ h)h↦f(update H a h) at H(a)H(a)H(a). The soft-max derivative milestone is stated with HasDerivAt, so it also asserts differentiability.
  • Rewards: each νx\nu_xνx​ is a probability measure on R\mathbb RR with Integrable identity and mean q∗(x)q_*(x)q∗​(x). The expectation of a function of (A,R)(A, R)(A,R) is ∑xπ(x)∫⋅ dνx\sum_x \pi(x) \int \cdot \, d\nu_x∑x​π(x)∫⋅dνx​. The book's normal-distribution testbed is only an example.
  • The baseline is a fixed real BBB, the book's "any scalar that does not depend on" the action (pp. 39–40). The book's Bt=RˉtB_t = \bar R_tBt​=Rˉt​, the average of past rewards, is covered once one conditions on the past. Footnote 1 on p. 37 states that the chapter's experiments used a Rˉt\bar R_tRˉt​ that also included RtR_tRt​. That baseline depends on AtA_tAt​, and the identity does not cover it.
  • Rewards of one action are a sequence indexed from 111. Q1Q_1Q1​ is arbitrary, and 00=10^0 = 100=1 as in the book (p. 33), so α=1\alpha = 1α=1 is included in (2.6).
  • Exercise 2.7 speaks of "a conventional constant step size α>0\alpha > 0α>0". The formalization takes α∈(0,1]\alpha \in (0,1]α∈(0,1], the range of the constant step size in (2.5). For α=2\alpha = 2α=2 the trace oˉn\bar o_noˉn​ vanishes at every even nnn and βn\beta_nβn​ is undefined. "Without initial bias" is read as "for n≥1n \ge 1n≥1, Qn+1Q_{n+1}Qn+1​ is the displayed weighted average of R1,…,RnR_1, \dots, R_nR1​,…,Rn​ with weights summing to one", which in particular does not involve Q1Q_1Q1​.
  • Exercise 2.9 is read as the two equalities π(1)=σ(H(1)−H(2))\pi(1) = \sigma(H(1)-H(2))π(1)=σ(H(1)−H(2)) and π(2)=σ(H(2)−H(1))\pi(2) = \sigma(H(2)-H(1))π(2)=σ(H(2)−H(1)).
  • A trivializing formalization is ruled out: the goal is about the expected value of the algorithm's update (2.12) under the joint law of action and reward, not the soft-max derivative alone and not a version in which the reward is replaced by its mean or the expectation is taken over AAA only.
  • Not formalized: the UCB rule (2.10) and the 10-armed testbed, which carry no provable claim in the chapter, and the stochastic-approximation conditions (2.7), which the book cites without proof.
  • Welcome contributions: a general soft-max library (derivatives, Jacobian, log-sum-exp) over a finite type, reusable for Chapter 13, and proofs of the milestones in the listed order.

Selected references

  • R. S. Sutton, A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018, ISBN 9780262039246, Chapter 2, pp. 25–46. http://incompleteideas.net/book/the-book-2nd.html
  • R. J. Williams, Simple statistical gradient-following algorithms for connectionist reinforcement learning, Machine Learning 8 (1992) 229–256. https://doi.org/10.1007/BF00992696
  • H. Robbins, S. Monro, A stochastic approximation method, Annals of Mathematical Statistics 22 (1951) 400–407. https://doi.org/10.1214/aoms/1177729586
11 thms2 active usersReviewed
🏆Completed
Dynamic ProgrammingOperations ResearchReinforcement Learning·Captain: mikedeng1

Approximately Optimal Approximate Reinforcement Learning II: Near-Optimality of a Policy with Small Policy AdvantageResearch Paper

Motivation

Approximate policy iteration and policy-gradient methods stop when they can no longer find a direction of improvement. Kakade and Langford (ICML 2002) asked what such a stopping point guarantees. Their algorithm, conservative policy iteration, halts at a policy π\piπ for which no policy can improve much on π\piπ as measured under a restart distribution μ\muμ; the quantity that is small is the optimal policy advantage OPT(Aπ,μ)\mathrm{OPT}(\mathbb A_{\pi,\mu})OPT(Aπ,μ​). Theorem 6.2 of the paper translates this local condition into a global statement: the performance of π\piπ is close to optimal, with a loss controlled by how well μ\muμ covers the states an optimal policy visits.

The bound is the origin of the distribution mismatch coefficient ∥dπ∗,μ~/μ∥∞\|d_{\pi^*,\tilde\mu}/\mu\|_\infty∥dπ∗,μ~​​/μ∥∞​, which reappears in the analysis of approximate dynamic programming (concentrability coefficients, Munos 2003), of conservative and trust-region methods, and of the convergence of policy gradient methods (Agarwal, Kakade, Lee, Mahajan 2021), where it governs the rate. The performance difference lemma (Lemma 6.1) used in its proof has become a standard tool of reinforcement learning theory.

Setting

A finite Markov decision process has a finite nonempty state set SSS, a finite nonempty action set AAA, transition probabilities P(s′;s,a)P(s';s,a)P(s′;s,a) (for each s,as,as,a a probability distribution over s′s's′), a reward function R:S×A→[0,R]\mathcal R:S\times A\to[0,R]R:S×A→[0,R] with R>0R>0R>0, and a discount factor 0≤γ<10\le\gamma<10≤γ<1. A stochastic policy π(a;s)\pi(a;s)π(a;s) is, for each state sss, a probability distribution over actions. A state distribution is a probability vector μ\muμ on SSS.

The normalized value function is Vπ(s)=(1−γ)E[∑t≥0γtR(st,at)∣π,s]V_\pi(s)=(1-\gamma)E[\sum_{t\ge0}\gamma^t\mathcal R(s_t,a_t)\mid\pi,s]Vπ​(s)=(1−γ)E[∑t≥0​γtR(st​,at​)∣π,s], where s0=ss_0=ss0​=s, at∼π(⋅;st)a_t\sim\pi(\cdot;s_t)at​∼π(⋅;st​) and st+1∼P(⋅;st,at)s_{t+1}\sim P(\cdot;s_t,a_t)st+1​∼P(⋅;st​,at​). The state–action value is Qπ(s,a)=(1−γ)R(s,a)+γ∑s′P(s′;s,a)Vπ(s′)Q_\pi(s,a)=(1-\gamma)\mathcal R(s,a)+\gamma\sum_{s'}P(s';s,a)V_\pi(s')Qπ​(s,a)=(1−γ)R(s,a)+γ∑s′​P(s′;s,a)Vπ​(s′) and the advantage is Aπ(s,a)=Qπ(s,a)−Vπ(s)A_\pi(s,a)=Q_\pi(s,a)-V_\pi(s)Aπ​(s,a)=Qπ​(s,a)−Vπ​(s). The discounted future state distribution from μ\muμ is

dπ,μ(s)=(1−γ)∑t≥0γtPr⁡(st=s;π,μ),s0∼μ,d_{\pi,\mu}(s)=(1-\gamma)\sum_{t\ge0}\gamma^t\Pr(s_t=s;\pi,\mu),\qquad s_0\sim\mu,dπ,μ​(s)=(1−γ)t≥0∑​γtPr(st​=s;π,μ),s0​∼μ,

and the performance of π\piπ from μ\muμ is ημ(π)=∑sμ(s)Vπ(s)\eta_\mu(\pi)=\sum_s\mu(s)V_\pi(s)ημ​(π)=∑s​μ(s)Vπ​(s).

The policy advantage of π′\pi'π′ with respect to π\piπ and μ\muμ is Aπ,μ(π′)=∑sdπ,μ(s)∑aπ′(a;s)Aπ(s,a)\mathbb A_{\pi,\mu}(\pi')=\sum_sd_{\pi,\mu}(s)\sum_a\pi'(a;s)A_\pi(s,a)Aπ,μ​(π′)=∑s​dπ,μ​(s)∑a​π′(a;s)Aπ​(s,a): the expected advantage of π′\pi'π′ over π\piπ on the states π\piπ itself visits. Its maximum over all stochastic policies is OPT(Aπ,μ)=max⁡π′Aπ,μ(π′)\mathrm{OPT}(\mathbb A_{\pi,\mu})=\max_{\pi'}\mathbb A_{\pi,\mu}(\pi')OPT(Aπ,μ​)=maxπ′​Aπ,μ​(π′) (Definition 4.3). An optimal policy π∗\pi^*π∗ satisfies Vπ(s)≤Vπ∗(s)V_\pi(s)\le V_{\pi^*}(s)Vπ​(s)≤Vπ∗​(s) for every policy π\piπ and every state sss. For nonnegative f,gf,gf,g on SSS, ∥f/g∥∞=max⁡sf(s)/g(s)\|f/g\|_\infty=\max_sf(s)/g(s)∥f/g∥∞​=maxs​f(s)/g(s) (p. 5).

Formalization targets

Goal: Theorem 6.2 (p. 6)

If OPT(Aπ,μ)<ε\mathrm{OPT}(\mathbb A_{\pi,\mu})<\varepsilonOPT(Aπ,μ​)<ε and π∗\pi^*π∗ is optimal, then for every state distribution μ~\tilde\muμ~​

ημ~(π∗)−ημ~(π)≤ε1−γ∥dπ∗,μ~dπ,μ∥∞≤ε(1−γ)2∥dπ∗,μ~μ∥∞.\eta_{\tilde\mu}(\pi^*)-\eta_{\tilde\mu}(\pi)\le\frac{\varepsilon}{1-\gamma}\left\|\frac{d_{\pi^*,\tilde\mu}}{d_{\pi,\mu}}\right\|_\infty\le\frac{\varepsilon}{(1-\gamma)^2}\left\|\frac{d_{\pi^*,\tilde\mu}}{\mu}\right\|_\infty.ημ~​​(π∗)−ημ~​​(π)≤1−γε​​dπ,μ​dπ∗,μ~​​​​∞​≤(1−γ)2ε​​μdπ∗,μ~​​​​∞​.

The goal states both inequalities and the outer bound. The evaluation distribution μ~\tilde\muμ~​ is arbitrary and unrelated to the restart distribution μ\muμ; taking μ~=D\tilde\mu=Dμ~​=D, the start distribution, gives Corollary 4.5 (p. 5).

Milestone: Lemma 6.1 (p. 6)

For any policies π~\tilde\piπ~, π\piπ and any starting distribution μ\muμ,

ημ(π~)−ημ(π)=11−γE(a,s)∼π~dπ~,μ[Aπ(s,a)].\eta_\mu(\tilde\pi)-\eta_\mu(\pi)=\frac{1}{1-\gamma}E_{(a,s)\sim\tilde\pi d_{\tilde\pi,\mu}}\big[A_\pi(s,a)\big].ημ​(π~)−ημ​(π)=1−γ1​E(a,s)∼π~dπ~,μ​​[Aπ​(s,a)].

The states are weighted by the future state distribution of the new policy π~\tilde\piπ~, the advantage is that of the old policy π\piπ.

Significance

Theorem 6.2 is the quality guarantee for conservative policy iteration: combined with the paper's Theorem 4.4 (the algorithm stops with OPT(Aπ,μ)<2ε\mathrm{OPT}(\mathbb A_{\pi,\mu})<2\varepsilonOPT(Aπ,μ​)<2ε after polynomially many calls), it bounds the suboptimality of the returned policy for any target distribution, independently of the size of the state space except through the mismatch coefficient. It also explains the role of the restart distribution: a more uniform μ\muμ makes ∥dπ∗,μ~/μ∥∞\|d_{\pi^*,\tilde\mu}/\mu\|_\infty∥dπ∗,μ~​​/μ∥∞​ small. Lemma 6.1 is used throughout later theory, from trust-region policy optimization to the global convergence of policy gradient methods.

Both results are proved in the paper, with short arguments. The contribution of this mission is a machine-checked version of the infinite-horizon discounted statement in the paper's normalization, with the ∥⋅∥∞\|\cdot\|_\infty∥⋅∥∞​ ratios handled exactly, including states where a denominator vanishes. Neither the discounted performance difference lemma for stochastic policies nor the distribution mismatch bound is known to be formalized in Mathlib; a finite-horizon performance difference identity has been formalized separately and is a different statement.

Difficulty

The mathematics is short; the difficulty is in the infinite-horizon bookkeeping. The value function and dπ,μd_{\pi,\mu}dπ,μ​ are infinite series, and Lemma 6.1 relates the series of two different policies: its natural one-line argument uses the Bellman equation for VπV_\piVπ​, which is not the definition here, together with interchanges of infinite sums over time with finite sums over states and actions, each of which needs summability. Theorem 6.2 then needs two facts that are not stated as results in the paper: that OPT(Aπ,μ)\mathrm{OPT}(\mathbb A_{\pi,\mu})OPT(Aπ,μ​) equals ∑sdπ,μ(s)max⁡aAπ(s,a)\sum_sd_{\pi,\mu}(s)\max_aA_\pi(s,a)∑s​dπ,μ​(s)maxa​Aπ​(s,a) (the supremum over policies is attained by a greedy policy, and max⁡aAπ(s,a)≥0\max_aA_\pi(s,a)\ge0maxa​Aπ​(s,a)≥0), and that dπ,μ(s)≥(1−γ)μ(s)d_{\pi,\mu}(s)\ge(1-\gamma)\mu(s)dπ,μ​(s)≥(1−γ)μ(s). Reading the ℓ∞\ell_\inftyℓ∞​ ratio with real division would give a false statement when a denominator is zero; the statement avoids this.

Formalization scope

States and actions are finite nonempty types; policies and kernels are real-valued functions π s a (the paper's π(a;s)\pi(a;s)π(a;s)) and P s a s' (the paper's P(s′;s,a)P(s';s,a)P(s′;s,a)), with their distribution properties as explicit hypotheses. The published definitions IsTransitionKernel, IsPolicy, InducedTransition, OccupationDist, InducedReward and PolicyValue from the Foundations of Machine Learning series are reused; VπV_\piVπ​ is (1−γ)(1-\gamma)(1−γ) times PolicyValue, the defining series. OPT\mathrm{OPT}OPT is the supremum of the policy advantages over stochastic policies, which is the paper's maximum. Optimality of π∗\pi^*π∗ is relative to stationary stochastic policies, the paper's policy class; the existence of an optimal policy (the paper's "well known result", p. 2) is not part of this mission.

Every hypothesis is explicit: rewards in [0,R][0,R][0,R] with R>0R>0R>0, 0≤γ<10\le\gamma<10≤γ<1, PPP a kernel, π\piπ and π∗\pi^*π∗ stochastic policies, μ\muμ and μ~\tilde\muμ~​ state distributions. Each ∥f/g∥∞\|f/g\|_\infty∥f/g∥∞​ bound is stated multiplicatively: "X≤K∥f/g∥∞X\le K\|f/g\|_\inftyX≤K∥f/g∥∞​" is "X≤KCX\le KCX≤KC for every CCC with f(s)≤Cg(s)f(s)\le Cg(s)f(s)≤Cg(s) for all sss". When some g(s)=0<f(s)g(s)=0<f(s)g(s)=0<f(s) no such CCC exists and the bound is empty, which matches ∥f/g∥∞=+∞\|f/g\|_\infty=+\infty∥f/g∥∞​=+∞; no full-support assumption is made on μ\muμ or μ~\tilde\muμ~​. The hypothesis OPT(Aπ,μ)<ε\mathrm{OPT}(\mathbb A_{\pi,\mu})<\varepsilonOPT(Aπ,μ​)<ε is on the supremum itself, not on the closed form ∑sdπ,μ(s)max⁡aAπ(s,a)\sum_sd_{\pi,\mu}(s)\max_aA_\pi(s,a)∑s​dπ,μ​(s)maxa​Aπ​(s,a), which is a step of the proof; a formalization that assumed the closed form, or that divided by dπ,μd_{\pi,\mu}dπ,μ​ in real arithmetic, would not be this theorem. The proof of the theorem uses only that π∗\pi^*π∗ is a policy; optimality is kept as a hypothesis because the paper states it.

The proof on p. 7 twice writes dπ,μ(s)≤(1−γ)μ(s)d_{\pi,\mu}(s)\le(1-\gamma)\mu(s)dπ,μ​(s)≤(1−γ)μ(s); the inequality it uses, and the one stated on p. 5, is dπ,μ(s)≥(1−γ)μ(s)d_{\pi,\mu}(s)\ge(1-\gamma)\mu(s)dπ,μ​(s)≥(1−γ)μ(s). This slip is in the proof, not in the statement. Pages are PDF pages; the paper has no printed page numbers.

Useful reusable infrastructure: summability and Bellman equations for the normalized discounted value, dπ,μd_{\pi,\mu}dπ,μ​ as a probability distribution with dπ,μ≥(1−γ)μd_{\pi,\mu}\ge(1-\gamma)\mudπ,μ​≥(1−γ)μ, and attainment of OPT\mathrm{OPT}OPT by a greedy policy. Contributions of any of these as separate lemmas are welcome.

Selected references

  • S. Kakade, J. Langford, Approximately Optimal Approximate Reinforcement Learning, Proceedings of the 19th International Conference on Machine Learning (ICML), 2002. https://dl.acm.org/doi/10.5555/645531.656005
  • R. Munos, Error Bounds for Approximate Policy Iteration, ICML 2003. https://dl.acm.org/doi/10.5555/3041838.3041903
  • A. Agarwal, S. Kakade, J. Lee, G. Mahajan, On the Theory of Policy Gradient Methods: Optimality, Approximation, and Distribution Shift, Journal of Machine Learning Research 22(98), 2021. https://jmlr.org/papers/v22/19-736.html
  • J. Schulman, S. Levine, P. Abbeel, M. Jordan, P. Moritz, Trust Region Policy Optimization, ICML 2015. https://arxiv.org/abs/1502.05477
10 thms2 active usersReviewed
🏆Completed
Captain: Minghui

Certified Federated Unlearning for Linearized ModelsResearch Paper

Removing a client's contribution

Federated learning combines information from several clients without pooling their raw training records. A client may later request removal of its contribution. Retraining on the retained records supplies a natural comparison model, but repeating the training process can be costly. Jin, Chen, Zhang, and Li introduce a linearized learning pipeline and a server-side removal procedure in Forgettable Federated Linear Learning with Certified Data Unlearning, arXiv:2306.02216v3. Their linearization makes the training objective quadratic, so the distinction between an exact Newton correction and an approximate correction can be studied explicitly.

This mission formalizes a corrected finite-run error bound motivated by that analysis. It is not a transcription or validation of the printed Theorem 2. The source audit found that the supplementary argument drops a finite-training term when passing to a limit, uses an invalid general inverse-perturbation inequality, and does not justify its three-term squared-norm constant. The draft preserves the removal problem while stating its error factors explicitly. The source anchors are Section III-C, Theorem 2, PDF pp. 5–6, and supplementary Section C5, PDF p. 16. The preprint first appeared in 2023; this mission fixes the revised May 2026 version so later source changes cannot silently alter its meaning.

Affine features and retained data

A parameter is a vector w∈Rdw\in\mathbb R^dw∈Rd. Record iii has a fixed linear feature map Ai:Rd→RkA_i:\mathbb R^d\to\mathbb R^kAi​:Rd→Rk, an offset aia_iai​, and a target yiy_iyi​. Its prediction is Aiw+aiA_iw+a_iAi​w+ai​. This represents the fixed linearization in the paper's equation (3); arbitrary real targets are permitted, and one-hot classification targets are a special case. Neither approximation accuracy for a nonlinear neural network nor an infinite-width limit is asserted.

Let DDD be the full finite dataset and SSS a nonempty subset of retained indices. Client removal is represented by retaining precisely the indices whose owner differs from the removed client. More general record removals are also allowed. For a fixed regularization parameter μ>0\mu>0μ>0, define

LS(w)=12∣S∣∑i∈S∥Aiw+ai−yi∥2+μ2∥w∥2.L_S(w)=\frac1{2|S|}\sum_{i\in S}\|A_iw+a_i-y_i\|^2+\frac\mu2\|w\|^2.LS​(w)=2∣S∣1​i∈S∑​∥Ai​w+ai​−yi​∥2+2μ​∥w∥2.

Write GS=∣S∣−1∑i∈SAi∗AiG_S=|S|^{-1}\sum_{i\in S}A_i^*A_iGS​=∣S∣−1∑i∈S​Ai∗​Ai​, HS=GS+μIH_S=G_S+\mu IHS​=GS​+μI, and bS=∣S∣−1∑i∈SAi∗(yi−ai)b_S=|S|^{-1}\sum_{i\in S}A_i^*(y_i-a_i)bS​=∣S∣−1∑i∈S​Ai∗​(yi​−ai​). Define uS=HS−1bSu_S=H_S^{-1}b_SuS​=HS−1​bS​ and let uDu_DuD​ use the full dataset. These reference parameters are computed from the data. The accepted child proofs establish the Hessian positivity and invertibility needed for the error bound; the broader unique-minimizer theorem is a separate supporting statement. The construction comes from Section III-A, PDF pp. 3–4, equations (3)–(5).

A separate nonempty server dataset PPP has Gram operator GPG_PGP​ and regularized Hessian HP=GP+μIH_P=G_P+\mu IHP​=GP​+μI. All operator norms below are Euclidean operator norms. The datasets and feature maps are fixed throughout the probability calculation.

Formalization targets

Let WWW be the trained parameter, RRR the parameter returned by retraining on SSS, and VVV an approximate removal correction. The removed parameter is W−VW-VW−V. Their joint probability model has finite outcome space Ω\OmegaΩ, with masses pω≥0p_\omega\ge0pω​≥0 summing to one. They may be dependent. This covers the outputs of finite randomized runs on finite data with fixed initialization; no independence assumption is used.

For each trained parameter www, define the server removal objective and its exact minimizer by

Fw(v)=12⟨v,HPv⟩−⟨HSw−bS,v⟩,vP(w)=HP−1(HSw−bS).F_w(v)=\tfrac12\langle v,H_Pv\rangle-\langle H_Sw-b_S,v\rangle, \qquad v_P(w)=H_P^{-1}(H_Sw-b_S).Fw​(v)=21​⟨v,HP​v⟩−⟨HS​w−bS​,v⟩,vP​(w)=HP−1​(HS​w−bS​).

This is the quadratic surrogate in Section III-B, PDF p. 5, equation (6). Define

Q=E[FW(V)−FW(vP(W))],κ=∥HP−1∥ ∥GP−GS∥,Q=\mathbb E[F_W(V)-F_W(v_P(W))],\quad \kappa=\|H_P^{-1}\|\,\|G_P-G_S\|,Q=E[FW​(V)−FW​(vP​(W))],κ=∥HP−1​∥∥GP​−GS​∥, Etrain=E∥W−uD∥2,Eretrain=E∥R−uS∥2.E_{\rm train}=\mathbb E\|W-u_D\|^2,\qquad E_{\rm retrain}=\mathbb E\|R-u_S\|^2.Etrain​=E∥W−uD​∥2,Eretrain​=E∥R−uS​∥2.

The corrected goal is

E∥W−V−R∥2≤6μQ+6κ2(Etrain+∥uD−uS∥2)+3Eretrain.\boxed{\mathbb E\|W-V-R\|^2\le \frac6\mu Q+6\kappa^2\bigl(E_{\rm train}+\|u_D-u_S\|^2\bigr) +3E_{\rm retrain}.}E∥W−V−R∥2≤μ6​Q+6κ2(Etrain​+∥uD​−uS​∥2)+3Eretrain​.​

The two completed milestones used by the accepted proof are the surrogate gap bound and the corrected signed removal-error identity:

μ2∥v−vP(w)∥2≤Fw(v)−Fw(vP(w)),\frac\mu2\|v-v_P(w)\|^2\le F_w(v)-F_w(v_P(w)),2μ​∥v−vP​(w)∥2≤Fw​(v)−Fw​(vP​(w)), w−v−r=HP−1(GP−GS)(w−uS)+(vP(w)−v)+(uS−r).w-v-r=H_P^{-1}(G_P-G_S)(w-u_S)+(v_P(w)-v)+(u_S-r).w−v−r=HP−1​(GP​−GS​)(w−uS​)+(vP​(w)−v)+(uS​−r).

The broader ridge-structure, exact-Newton-removal and inverse-perturbation statements remain available as separate open theorems. Their milestone entries were removed because the accepted proof does not depend on their full statements.

Formalization note: the completed root is a corrected, paper-derived error bound. Its formal bridge uses the two source-backed child theorems above, anchored to Section III-B (Section 3), PDF p. 5, equation (6), and Section III-C (Section 3), PDF p. 5 and PDF p. 6, Theorem 2; supplementary C5, PDF p. 16, unnumbered displays. The coefficients in the boxed goal are conservative; no optimality claim is made.

What the result supplies

The result connects the removal solver's objective gap, the difference between the server and retained Hessians, and the actual optimization errors to an observable parameter discrepancy. Exact Hessian matching sets κ=0\kappa=0κ=0. Exact removal optimization sets Q=0Q=0Q=0, but finite retraining error still remains. This distinguishes exact optimization of the retained objective from reproducing an unfinished retraining run.

The original paper motivates the comparison; the displayed corrected bound is a new formulation derived from its quadratic setting. The root Lean theorem and its two dependency milestones are now Proved. Their accepted proofs match the original formal statements exactly; the three separate supporting statements remain open. The requested OpenProblem classification describes the formalization task and does not assert that the elementary corrected inequality is an unresolved research conjecture.

The mathematical difficulty

An approximate server Hessian cannot be substituted for the retained Hessian without a sensitivity term. A bound on the difference of the Gram operators alone does not bound its action on every parameter vector. Likewise, a small training error relative to the full-data optimum does not imply that the full and retained optima coincide. The displacement ∥uD−uS∥\|u_D-u_S\|∥uD​−uS​∥ therefore remains visible. Formalization must respect the normalization of each empirical objective, the sign of the correction, and the operator norm used in the perturbation estimate.

Formalization scope

The model uses finite-dimensional real Euclidean spaces, continuous linear maps and adjoints, finite index sets, a total ring inverse, and finite weighted expectations. Positive regularization must justify every use of the inverse; it is not an invertibility assumption hidden inside the dataset. Nonempty retained and server data exclude division by an empty sample count. Zero-dimensional feature or parameter spaces are permitted and harmless. A finite law on an empty outcome type has no inhabitant because its masses cannot sum to one.

The root theorem quantifies over arbitrary output maps W,V,RW,V,RW,V,R. It is an error-propagation theorem in terms of their actual errors and surrogate gap, not a convergence theorem for a particular implementation. Obtaining algorithm-specific bounds on those quantities is separate future work. In particular, the draft does not import the source's unsupported all-smaller-learning-rates FedAvg contraction claim. It also makes no differential-privacy, distributional indistinguishability, nonlinear-network, or empirical accuracy assertion.

Selected references

  • Ruinan Jin, Minghui Chen, Qiong Zhang, Xiaoxiao Li, Forgettable Federated Linear Learning with Certified Data Unlearning, IEEE Transactions on Neural Networks and Learning Systems, early access (2026). arXiv:2306.02216v3, DOI. Main anchors: Section II-B, PDF p. 3, equation (1); Sections III-A–III-C, PDF pp. 3–6, equations (3)–(6), Theorem 2; supplementary Section C5, PDF p. 16, unnumbered displays.
7 thms2 active usersReviewed
🏆Completed
OptimizationStatistics·Captain: mikedeng1

Robustness and Generalization IV: Robustness of the Lasso on a Compact Sample SpaceResearch Paper

Motivation

The Lasso (Tibshirani 1996, doi:10.1111/j.2517-6161.1996.tb02080.x) is ℓ1\ell_1ℓ1​-penalized least squares regression, one of the standard estimators of statistics and machine learning because it selects sparse coefficient vectors. Explaining why a learned Lasso predictor generalizes is less routine than it looks. The two classical routes are uniform convergence over the hypothesis class and algorithmic stability (Bousquet and Elisseeff 2002, JMLR 2:499–526). The stability route is closed for the Lasso: Xu, Caramanis and Mannor (IEEE Trans. Inf. Theory 56(7), 2010, doi:10.1109/TIT.2010.2048503) showed that its uniform stability bound does not decrease with the sample size, a fact reproduced as Theorem 7 of Xu and Mannor (2012).

Xu and Mannor, Robustness and Generalization (Mach Learn 86 (2012) 391–423, doi:10.1007/s10994-011-5268-1), propose a third route, algorithmic robustness: if the sample space can be split into KKK cells such that a test point in the same cell as a training point has nearly the same loss, then the algorithm generalizes (their Theorem 1). Their Example 6 shows that the Lasso is robust in this sense, with a number of cells given by a covering number and a robustness level depending on the training responses. This mission formalizes Example 6 together with the general criterion it rests on (Theorem 6) and the Lipschitz estimate for the Lasso loss (Lemma 3).

Setting

A sample is a point z=(z(y),z(x))z = (z^{(y)}, z^{(x)})z=(z(y),z(x)) with a response z(y)∈Rz^{(y)} \in \mathbb Rz(y)∈R and a feature vector z(x)∈Rmz^{(x)} \in \mathbb R^mz(x)∈Rm, so the samples live in Rm+1\mathbb R^{m+1}Rm+1. The sample space Z⊆Rm+1\mathcal Z \subseteq \mathbb R^{m+1}Z⊆Rm+1 is a compact set, and Rm+1\mathbb R^{m+1}Rm+1 carries the norm ∥z∥∞=max⁡(∣z(y)∣,max⁡j∣zj(x)∣)\|z\|_\infty = \max(|z^{(y)}|, \max_j |z^{(x)}_j|)∥z∥∞​=max(∣z(y)∣,maxj​∣zj(x)​∣). A training set is s=(s1,…,sn)∈Zn\mathbf s = (s_1, \dots, s_n) \in \mathcal Z^ns=(s1​,…,sn​)∈Zn.

A learning algorithm maps each training set s\mathbf ss to a hypothesis As\mathcal A_{\mathbf s}As​; with a loss l(h,z)l(h, z)l(h,z), it is (K,ϵ(⋅))(K, \epsilon(\cdot))(K,ϵ(⋅))-robust (Definition 2, p. 396) if Z\mathcal ZZ can be partitioned into KKK disjoint sets C1,…,CKC_1, \dots, C_KC1​,…,CK​, fixed independently of the data, such that for every s∈Zn\mathbf s \in \mathcal Z^ns∈Zn, every training point s∈ss \in \mathbf ss∈s, every z∈Zz \in \mathcal Zz∈Z and every iii,

s,z∈Ci  ⟹  ∣l(As,s)−l(As,z)∣≤ϵ(s).s, z \in C_i \implies |l(\mathcal A_{\mathbf s}, s) - l(\mathcal A_{\mathbf s}, z)| \le \epsilon(\mathbf s).s,z∈Ci​⟹∣l(As​,s)−l(As​,z)∣≤ϵ(s).

For a metric ρ\rhoρ on Z\mathcal ZZ and ϵ>0\epsilon > 0ϵ>0, a set T^⊆Z\hat T \subseteq \mathcal ZT^⊆Z is an ϵ\epsilonϵ-cover of Z\mathcal ZZ if every point of Z\mathcal ZZ is within distance ≤ϵ\le \epsilon≤ϵ of a point of T^\hat TT^; the covering number N(ϵ,Z,ρ)\mathcal N(\epsilon, \mathcal Z, \rho)N(ϵ,Z,ρ) is the least cardinality of such a cover (Definition 1, p. 394).

For a coefficient vector w∈Rmw \in \mathbb R^mw∈Rm let ∥w∥1=∑j∣wj∣\|w\|_1 = \sum_j |w_j|∥w∥1​=∑j​∣wj​∣. Given c>0c > 0c>0, the Lasso is

min⁡w 1n∑i=1n(si(y)−w⊤si(x))2+c∥w∥1,(5)\min_{w} \ \frac1n \sum_{i=1}^n \big(s_i^{(y)} - w^\top s_i^{(x)}\big)^2 + c\|w\|_1, \tag{5}wmin​ n1​i=1∑n​(si(y)​−w⊤si(x)​)2+c∥w∥1​,(5)

a Lasso algorithm returns a minimizer As=w\mathcal A_{\mathbf s} = wAs​=w of (5) for each s\mathbf ss, and the loss is the absolute prediction error l(w,z)=∣z(y)−w⊤z(x)∣l(w, z) = |z^{(y)} - w^\top z^{(x)}|l(w,z)=∣z(y)−w⊤z(x)∣. Finally Y(s)=1n∑i=1n[si(y)]2Y(\mathbf s) = \frac1n \sum_{i=1}^n [s_i^{(y)}]^2Y(s)=n1​∑i=1n​[si(y)​]2.

Formalization targets

Goal: Example 6 (p. 404)

For every compact Z⊆Rm+1\mathcal Z \subseteq \mathbb R^{m+1}Z⊆Rm+1, every c>0c > 0c>0, every Lasso algorithm A\mathcal AA and every γ>0\gamma > 0γ>0,

A is (N(γ/2,Z,∥⋅∥∞), (Y(s)/c+1)γ)-robust.\mathcal A \text{ is } \Big(\mathcal N(\gamma/2, \mathcal Z, \|\cdot\|_\infty),\ \big(Y(\mathbf s)/c + 1\big)\gamma\Big)\text{-robust}.A is (N(γ/2,Z,∥⋅∥∞​), (Y(s)/c+1)γ)-robust.

The statement holds for every selection of a minimizer, since (5) need not have a unique solution.

Milestones

  1. Optimality bound (proof of Lemma 3, p. 419): every Lasso solution satisfies ∥w∗∥1≤1nc∑i=1n[si(y)]2\|w^*\|_1 \le \frac{1}{nc} \sum_{i=1}^n [s_i^{(y)}]^2∥w∗∥1​≤nc1​∑i=1n​[si(y)​]2.
  2. Lemma 3 (p. 419): for all za,zb∈Rm+1z_a, z_b \in \mathbb R^{m+1}za​,zb​∈Rm+1,
∣l(w∗(s),za)−l(w∗(s),zb)∣≤[1nc∑i=1n[si(y)]2+1]∥za−zb∥∞.|l(w^*(\mathbf s), z_a) - l(w^*(\mathbf s), z_b)| \le \Big[\frac{1}{nc} \sum_{i=1}^n [s_i^{(y)}]^2 + 1\Big] \|z_a - z_b\|_\infty.∣l(w∗(s),za​)−l(w∗(s),zb​)∣≤[nc1​i=1∑n​[si(y)​]2+1]∥za​−zb​∥∞​.
  1. Theorem 6 (p. 402): for a metric ρ\rhoρ on Z\mathcal ZZ and γ>0\gamma > 0γ>0, if ∣l(As,z1)−l(As,z2)∣≤ϵ(s)|l(\mathcal A_{\mathbf s}, z_1) - l(\mathcal A_{\mathbf s}, z_2)| \le \epsilon(\mathbf s)∣l(As​,z1​)−l(As​,z2​)∣≤ϵ(s) whenever z1∈sz_1 \in \mathbf sz1​∈s and ρ(z1,z2)≤γ\rho(z_1, z_2) \le \gammaρ(z1​,z2​)≤γ, and N(γ/2,Z,ρ)<∞\mathcal N(\gamma/2, \mathcal Z, \rho) < \inftyN(γ/2,Z,ρ)<∞, then A\mathcal AA is (N(γ/2,Z,ρ),ϵ(⋅))(\mathcal N(\gamma/2, \mathcal Z, \rho), \epsilon(\cdot))(N(γ/2,Z,ρ),ϵ(⋅))-robust.

Significance

Combined with Theorem 1 of the same paper, Example 6 yields a generalization bound for the Lasso of the form ϵ(s)+M(2Kln⁡2+2ln⁡(1/δ))/n\epsilon(\mathbf s) + M\sqrt{(2K\ln 2 + 2\ln(1/\delta))/n}ϵ(s)+M(2Kln2+2ln(1/δ))/n​ with KKK a covering number of the sample space, a bound that uses no stability of the algorithm and no uniqueness of the minimizer. Theorem 6 is the reusable part: it converts any data-dependent local Lipschitz or continuity estimate of the loss into robustness, and the paper derives its examples for the SVM, the Lasso, neural networks and PCA from it. The authors note (p. 404) that the resulting bound is weaker than VC-dimension bounds for linear predictors, since it depends exponentially on the dimension; the value of the example is the method, not the rate.

The results are proved in the paper, with short arguments. No machine-checked version of Theorem 6, Lemma 3 or Example 6 is known to exist. The formal work is to connect Mathlib's covering numbers to partitions of a set, to handle the ℓ1\ell_1ℓ1​/ℓ∞\ell_\inftyℓ∞​ pairing on R×Rm\mathbb R \times \mathbb R^mR×Rm, and to state robustness so that later missions of this series (the generalization bound of Theorem 1, mission I) can consume it.

Difficulty

The constant in the robustness level depends on the training set through Y(s)Y(\mathbf s)Y(s), while the partition in Definition 2 must be chosen before the training set is seen. A formalization that lets the cells depend on s\mathbf ss proves a much weaker, nearly empty statement, so the data dependence has to be carried entirely by ϵ(s)\epsilon(\mathbf s)ϵ(s) and the cells must depend only on Z\mathcal ZZ and γ\gammaγ. A cover by balls is not a partition, and the radius of the cover (γ/2\gamma/2γ/2) and the closeness threshold in Theorem 6 (γ\gammaγ) differ by the factor that the diameter of a cell requires. The Lipschitz estimate must bound a Lasso solution without any information beyond optimality, and the pairing between ∥w∥1\|w\|_1∥w∥1​ and ∥⋅∥∞\|\cdot\|_\infty∥⋅∥∞​ is the one that makes the constant come out as printed; a Euclidean norm on either side gives a different constant.

Formalization scope

  • Rm+1\mathbb R^{m+1}Rm+1 is ℝ × (Fin m → ℝ), a point being (z^{(y)}, z^{(x)}). Lean's norm on this product is the maximum of the absolute values of all coordinates, which is exactly ∥⋅∥∞\|\cdot\|_\infty∥⋅∥∞​. ∥w∥1\|w\|_1∥w∥1​ is written out as ∑j∣wj∣\sum_j |w_j|∑j​∣wj​∣, since the default norm on Fin m → ℝ is the sup norm; w⊤xw^\top xw⊤x is dotProduct w x.
  • The sample space is a set Z with IsCompact Z. Robustness (IsRobustOn) asks for cells C : Fin K → Set α that lie in Z, cover Z and are pairwise disjoint (empty cells allowed), chosen before the universally quantified training set; training sets are maps Fin n → α with all points in Z. No measurability is involved anywhere in this mission.
  • The covering number is Mathlib's Metric.coveringNumber at radius Real.toNNReal (γ / 2): closed balls, centres in Z (the metric space of Definition 1 is Z\mathcal ZZ itself), value in ℕ∞, converted with toNat. Theorem 6 assumes its finiteness, as the paper does; without that hypothesis toNat would return 000 and the statement would be false for nonempty Z. Example 6 does not assume it: it follows from compactness.
  • A Lasso algorithm is any function A with ∀ s, IsLassoSolution c s (A s); it is not defined by a choice of minimizer. The regularization parameter satisfies c>0c > 0c>0, which the paper leaves implicit. The factor 1/n1/n1/n is a real division; for n=0n = 0n=0 it is 000 in Lean, the objective reduces to c∥w∥1c\|w\|_1c∥w∥1​, and all statements remain true.
  • The robustness level is (Y(s)/c+1)γ(Y(\mathbf s)/c + 1)\gamma(Y(s)/c+1)γ in Example 6 and 1nc∑i[si(y)]2+1\frac{1}{nc}\sum_i [s_i^{(y)}]^2 + 1nc1​∑i​[si(y)​]2+1 in Lemma 3, each in its printed form.

Useful infrastructure beyond this mission: a lemma turning a finite cover of a set into a partition of it with cells of diameter at most twice the radius, and finiteness of Mathlib's internal covering number for compact sets. Contributions of either as separate theorems are welcome.

Selected references

  • H. Xu and S. Mannor, Robustness and Generalization, Machine Learning 86 (2012) 391–423. doi:10.1007/s10994-011-5268-1
  • R. Tibshirani, Regression Shrinkage and Selection via the Lasso, Journal of the Royal Statistical Society, Series B 58(1) (1996) 267–288. doi:10.1111/j.2517-6161.1996.tb02080.x
  • H. Xu, C. Caramanis and S. Mannor, Robust Regression and Lasso, IEEE Transactions on Information Theory 56(7) (2010) 3561–3574. doi:10.1109/TIT.2010.2048503
  • O. Bousquet and A. Elisseeff, Stability and Generalization, Journal of Machine Learning Research 2 (2002) 499–526. jmlr.org/papers/v2/bousquet02a
7 thms2 active usersReviewed
🏆Completed
ProbabilityStatistics·Captain: mikedeng1

Robustness and Generalization III: Quantile-Value and Truncated-Mean Generalization Bounds for Pseudo-Robust AlgorithmsResearch Paper

Motivation

Classical generalization bounds control the gap between the expected loss of a learned hypothesis and its average loss on the training sample. The average is sensitive to outliers: when a non-negligible fraction of the sample is corrupted, the mean loss stops describing the quality of a solution, and quantile-type summaries such as the median become the natural measurement. Quantile losses have long been used for this reason in statistics and econometrics (Koenker and Bassett 1978; Huber 1981). The standard tools for proving generalization bounds — symmetrization, Rademacher and VC arguments — are built around the expected loss and do not extend to quantiles in any direct way.

Xu and Mannor (Mach Learn 86 (2012) 391–423) introduced algorithmic robustness: an algorithm is robust if the sample space can be partitioned into finitely many cells such that a test point falling in the same cell as a training point incurs a similar loss. Because the argument works cell by cell and needs no symmetrization, it transfers to loss functionals other than the mean. Sect. 4.1 of the paper uses this to bound the quantile value and the truncated mean of the testing error, and Sect. 5 relaxes robustness to pseudo robustness, which only asks the cell condition for a subset of the training samples. This mission formalizes the resulting Theorem 5 (p. 402), whose proof is Appendix C (pp. 415–418).

Setting

Let Z\mathcal ZZ be a measurable sample space, H\mathcal HH a set of hypotheses and l:H×Z→[0,M]l : \mathcal H \times \mathcal Z \to [0, M]l:H×Z→[0,M] a loss, with each l(h,⋅)l(h, \cdot)l(h,⋅) measurable. A training set s=(s1,…,sn)\mathbf s = (s_1, \dots, s_n)s=(s1​,…,sn​) consists of nnn i.i.d. draws from a probability measure μ\muμ on Z\mathcal ZZ; its empirical distribution is μemp=1n∑iδsi\mu_{\mathrm{emp}} = \frac1n \sum_i \delta_{s_i}μemp​=n1​∑i​δsi​​. A learning algorithm is a map A:Zn→H\mathcal A : \mathcal Z^n \to \mathcal HA:Zn→H, and As\mathcal A_{\mathbf s}As​ is the hypothesis learned from s\mathbf ss.

For a real random variable XXX and a level β\betaβ, the β\betaβ-quantile value is

Qβ(X)=inf⁡{c∈R:Pr⁡(X≤c)≥β},\mathbb Q^\beta(X) = \inf\{ c \in \mathbb R : \Pr(X \le c) \ge \beta \},Qβ(X)=inf{c∈R:Pr(X≤c)≥β},

and, writing Q=Qβ(X)Q = \mathbb Q^\beta(X)Q=Qβ(X), the β\betaβ-truncated mean is

Tβ(X)=E[X⋅1(X<Q)]+(β−Pr⁡[X<Q]) Q,\mathbb T^\beta(X) = \mathbb E[X \cdot \mathbf 1(X < Q)] + \big(\beta - \Pr[X < Q]\big)\, Q,Tβ(X)=E[X⋅1(X<Q)]+(β−Pr[X<Q])Q,

where the second term vanishes when Pr⁡[X=Q]=0\Pr[X = Q] = 0Pr[X=Q]=0. It is the contribution to EX\mathbb E XEX of the leftmost β\betaβ fraction of the distribution. For a hypothesis hhh and a measure ν\nuν on Z\mathcal ZZ put Q(h,β,ν)=Qβ(l(h,z))\mathcal Q(h, \beta, \nu) = \mathbb Q^\beta(l(h, z))Q(h,β,ν)=Qβ(l(h,z)) and T(h,β,ν)=Tβ(l(h,z))\mathcal T(h, \beta, \nu) = \mathbb T^\beta(l(h, z))T(h,β,ν)=Tβ(l(h,z)) with z∼νz \sim \nuz∼ν.

The algorithm is (K,ϵ(⋅),n^(⋅))(K, \epsilon(\cdot), \hat n(\cdot))(K,ϵ(⋅),n^(⋅)) pseudo robust, with ϵ:Zn→R\epsilon : \mathcal Z^n \to \mathbb Rϵ:Zn→R and n^:Zn→{1,…,n}\hat n : \mathcal Z^n \to \{1, \dots, n\}n^:Zn→{1,…,n}, if Z\mathcal ZZ can be partitioned into KKK disjoint sets C1,…,CKC_1, \dots, C_KC1​,…,CK​, fixed in advance, such that every training set s\mathbf ss has a subset s^\hat{\mathbf s}s^ of n^(s)\hat n(\mathbf s)n^(s) samples with: whenever s∈s^s \in \hat{\mathbf s}s∈s^ and z∈Zz \in \mathcal Zz∈Z lie in a common cell, ∣l(As,s)−l(As,z)∣≤ϵ(s)|l(\mathcal A_{\mathbf s}, s) - l(\mathcal A_{\mathbf s}, z)| \le \epsilon(\mathbf s)∣l(As​,s)−l(As​,z)∣≤ϵ(s). With n^≡n\hat n \equiv nn^≡n this is (K,ϵ(⋅))(K, \epsilon(\cdot))(K,ϵ(⋅))-robustness.

Formalization targets

Goal: Theorem 5 (p. 402)

Let λ0=(2Kln⁡2+2ln⁡(1/δ))/n\lambda_0 = \sqrt{(2K \ln 2 + 2 \ln(1/\delta))/n}λ0​=(2Kln2+2ln(1/δ))/n​ and r(s)=(n−n^(s))/nr(\mathbf s) = (n - \hat n(\mathbf s))/nr(s)=(n−n^(s))/n. If A\mathcal AA is (K,ϵ(⋅),n^(⋅))(K, \epsilon(\cdot), \hat n(\cdot))(K,ϵ(⋅),n^(⋅)) pseudo robust, β∈(0,1)\beta \in (0,1)β∈(0,1) and δ>0\delta > 0δ>0, then with probability at least 1−δ1 - \delta1−δ: whenever 0≤β−λ0−r(s)0 \le \beta - \lambda_0 - r(\mathbf s)0≤β−λ0​−r(s) and β+λ0+r(s)≤1\beta + \lambda_0 + r(\mathbf s) \le 1β+λ0​+r(s)≤1,

Q(As,β−λ0−r(s),μemp)−ϵ(s)≤Q(As,β,μ)≤Q(As,β+λ0+r(s),μemp)+ϵ(s),\mathcal Q(\mathcal A_{\mathbf s}, \beta - \lambda_0 - r(\mathbf s), \mu_{\mathrm{emp}}) - \epsilon(\mathbf s) \le \mathcal Q(\mathcal A_{\mathbf s}, \beta, \mu) \le \mathcal Q(\mathcal A_{\mathbf s}, \beta + \lambda_0 + r(\mathbf s), \mu_{\mathrm{emp}}) + \epsilon(\mathbf s),Q(As​,β−λ0​−r(s),μemp​)−ϵ(s)≤Q(As​,β,μ)≤Q(As​,β+λ0​+r(s),μemp​)+ϵ(s), T(As,β−λ0−r(s),μemp)−ϵ(s)≤T(As,β,μ)≤T(As,β+λ0+r(s),μemp)+ϵ(s).\mathcal T(\mathcal A_{\mathbf s}, \beta - \lambda_0 - r(\mathbf s), \mu_{\mathrm{emp}}) - \epsilon(\mathbf s) \le \mathcal T(\mathcal A_{\mathbf s}, \beta, \mu) \le \mathcal T(\mathcal A_{\mathbf s}, \beta + \lambda_0 + r(\mathbf s), \mu_{\mathrm{emp}}) + \epsilon(\mathbf s).T(As​,β−λ0​−r(s),μemp​)−ϵ(s)≤T(As​,β,μ)≤T(As​,β+λ0​+r(s),μemp​)+ϵ(s).

The constants are the paper's, and KKK, ϵ\epsilonϵ, n^\hat nn^, MMM, μ\muμ, δ\deltaδ and the algorithm are arbitrary.

Milestones (Appendix C)

  1. Property 1 (p. 415): for a nonnegative XXX and levels 0≤β2≤β1≤10 \le \beta_2 \le \beta_1 \le 10≤β2​≤β1​≤1 (with β1=1\beta_1 = 1β1​=1 only for XXX bounded above), Qβ1(X)≥Qβ2(X)\mathbb Q^{\beta_1}(X) \ge \mathbb Q^{\beta_2}(X)Qβ1​(X)≥Qβ2​(X) and Tβ1(X)≥Tβ2(X)\mathbb T^{\beta_1}(X) \ge \mathbb T^{\beta_2}(X)Tβ1​(X)≥Tβ2​(X).
  2. Property 2 (p. 415): if Pr⁡(Y≥a)≥Pr⁡(X≥a)\Pr(Y \ge a) \ge \Pr(X \ge a)Pr(Y≥a)≥Pr(X≥a) for all aaa, then Qβ(Y)≥Qβ(X)\mathbb Q^\beta(Y) \ge \mathbb Q^\beta(X)Qβ(Y)≥Qβ(X) and Tβ(Y)≥Tβ(X)\mathbb T^\beta(Y) \ge \mathbb T^\beta(X)Tβ(Y)≥Tβ(X) for β∈[0,1]\beta \in [0,1]β∈[0,1].
  3. The event E\mathcal EE (pp. 415–416): with NiN_iNi​ the indices of samples in CiC_iCi​, ∑i∣∣Ni∣/n−μ(Ci)∣≤λ0\sum_i \big| |N_i|/n - \mu(C_i) \big| \le \lambda_0∑i​​∣Ni​∣/n−μ(Ci​)​≤λ0​ with probability at least 1−δ1 - \delta1−δ.

Significance

The result. Theorem 5 shows that any pseudo-robust algorithm has a testing-error quantile and truncated mean that are bracketed by the empirical ones at levels shifted by λ0+(n−n^(s))/n\lambda_0 + (n - \hat n(\mathbf s))/nλ0​+(n−n^(s))/n, up to the robustness tolerance ϵ(s)\epsilon(\mathbf s)ϵ(s). The quantile of the testing error can therefore be estimated from training data for every algorithm to which the robustness framework applies — among them majority voting, SVMs, Lasso and principal component analysis (Sect. 6 of the paper) — without a separate complexity analysis of the loss class. The pseudo-robust form covers algorithms that are robust only away from a small set of training samples, which is the typical situation in the presence of outliers. The robust case n^≡n\hat n \equiv nn^≡n is the paper's Theorem 2 (p. 400).

Formalizing it. The paper states Theorem 5 and proves it in Appendix C; no machine-checked proof exists. The appendix contains misprints (see Formalization scope) and the argument uses minimizers of the loss over each cell, which need not exist; a formal proof settles which steps are sound as written. The definitions of quantile value and truncated mean of a law on R\mathbb RR developed here are reusable beyond this mission.

Difficulty

The concentration step is the same as for the expected loss: on the event E\mathcal EE the empirical cell frequencies are close to the cell probabilities. The difficulty is converting this into a statement about quantiles. Quantile values are not linear in the distribution and are discontinuous in the level, so the triangle-inequality argument that bounds the mean-loss gap does not apply. Mass that moves between cells shifts every level of the quantile function, and the up to n−n^(s)n - \hat n(\mathbf s)n−n^(s) samples outside s^\hat{\mathbf s}s^ carry no guarantee at all, so an arbitrary fraction r(s)r(\mathbf s)r(s) of the empirical law is uncontrolled. For the truncated mean this must be done for the whole lower tail up to level β\betaβ, not just at one point, and the atoms of the loss distribution (the second branch of the definition) have to be accounted for exactly.

Formalization scope

  • The Lean namespace is XuMannorRobust.Quantile. Z\mathcal ZZ is a type with a measurable space structure, H\mathcal HH an arbitrary type, a training set a function Fin n → Z, and the i.i.d. law the product measure Measure.pi (fun _ => μ).
  • "With probability at least 1−δ1 - \delta1−δ" is encoded as: the outer measure of the set of training sets on which the claim fails is at most δ\deltaδ. No measurability of s↦As\mathbf s \mapsto \mathcal A_{\mathbf s}s↦As​ is needed.
  • Added measurability. The paper ignores measurability; the formalization requires each l(h,⋅)l(h,\cdot)l(h,⋅) and each cell CiC_iCi​ to be measurable.
  • Corrected Definition 3. The paper prints the second branch of the truncated mean as (β−Pr⁡[X<Q])/Pr⁡[X=Q]⋅Q\big(\beta - \Pr[X < Q]\big)/\Pr[X = Q] \cdot Q(β−Pr[X<Q])/Pr[X=Q]⋅Q. That contradicts its own worked example on p. 399, where the 0.630.630.63-truncated mean of a uniform law on c1<⋯<c10c_1 < \dots < c_{10}c1​<⋯<c10​ is 0.1(∑i≤6ci+0.3c7)0.1(\sum_{i \le 6} c_i + 0.3 c_7)0.1(∑i≤6​ci​+0.3c7​), and its verbal description. The formalization drops the division, as the example requires; with the printed formula Tβ\mathbb T^\betaTβ would not even be monotone in β\betaβ.
  • Qβ\mathbb Q^\betaQβ and Tβ\mathbb T^\betaTβ are defined on the law of the random variable, a measure on R\mathbb RR. Lean returns 000 for the infimum of an empty set or of a set unbounded below, so Q0=0\mathbb Q^0 = 0Q0=0 (the paper's value is −∞-\infty−∞). This never helps: Q(As,β,μ)≥0\mathcal Q(\mathcal A_{\mathbf s}, \beta, \mu) \ge 0Q(As​,β,μ)≥0 and ϵ(s)≥0\epsilon(\mathbf s) \ge 0ϵ(s)≥0, so the goal's inequalities remain meaningful at level 000. The goal keeps every level in [0,1][0,1][0,1] through the paper's side condition, which depends on n^(s)\hat n(\mathbf s)n^(s) and is therefore placed inside the probability event as a premise. The codomain {1,…,n}\{1, \dots, n\}{1,…,n} of n^\hat nn^ is part of the definition: with n^(s)=0\hat n(\mathbf s) = 0n^(s)=0 nothing would constrain ϵ(s)\epsilon(\mathbf s)ϵ(s).
  • The partition is fixed before the training set; the good subset s^\hat{\mathbf s}s^ may depend on s\mathbf ss and is a set of indices. Choosing the partition after s\mathbf ss would make pseudo robustness trivial and is ruled out.
  • Properties 1 and 2 are stated for nonnegative laws and levels in [0,1][0,1][0,1]. The level 111 is admitted only for a variable bounded above (for property 2, the dominating one). For an unbounded variable, Q1\mathbb Q^1Q1 is +∞+\infty+∞ in the paper, where the inequality is trivial, and a junk 000 in Lean. Property 3 of Appendix C (p. 415) is misprinted (with the constraint ∑αi≤β\sum \alpha_i \le \beta∑αi​≤β the minimum is 000) and is not formalized.
  • Needed infrastructure: the Bretagnolle–Huber–Carol inequality for multinomial frequencies (van der Vaart and Wellner 1996, Prop. A.6.6) (or a direct concentration argument), and elementary order properties of lower quantile values and truncated means of laws on R\mathbb RR. Contributions of these as separate lemmas are welcome.

Selected references

  • H. Xu and S. Mannor, Robustness and Generalization, Machine Learning 86(3):391–423, 2012. https://doi.org/10.1007/s10994-011-5268-1
  • R. Koenker and G. Bassett, Regression Quantiles, Econometrica 46(1):33–50, 1978. https://doi.org/10.2307/1913643
  • P. J. Huber, Robust Statistics, Wiley, 1981. https://doi.org/10.1002/0471725250
  • A. W. van der Vaart and J. A. Wellner, Weak Convergence and Empirical Processes, Springer, 1996 (Proposition A.6.6). https://doi.org/10.1007/978-1-4757-2545-2
7 thms2 active usersReviewed
🏆Completed
ProbabilityStatistics·Captain: mikedeng1

Robustness and Generalization II: A Learning Method Generalizes w.r.t. a Training Sequence If and Only If It Is Weakly Robust w.r.t. ItResearch Paper

Motivation

Most generalization guarantees in statistical learning theory bound the gap between training error and expected error through a complexity measure of the hypothesis class: VC dimension, Rademacher complexity, covering numbers. Such bounds are sufficient conditions, and they say little about why a particular algorithm, run on a particular data stream, does or does not generalize. Xu and Mannor (Mach Learn 86 (2012) 391–423) proposed algorithmic robustness as an alternative: an algorithm is robust if a test sample "close to" a training sample incurs a loss close to that training sample's loss. Their first results show that robustness implies generalization (Theorem 1 of the paper, the subject of the first mission in this series).

Section 8 of the paper asks the converse question: is some form of robustness also necessary? The answer is Theorem 8. For a learning method trained on a fixed, growing sequence of samples, generalization is equivalent to a weaker property, weak robustness. The authors present this as evidence that robustness is "an essential property of successful learning", and contrast it with the characterization of learnability by stability (Remark 5 of the paper, citing Shalev-Shwartz et al. 2009; journal version JMLR 11 (2010)): learnability is uniform over all distributions, whereas the generalization studied here is for one distribution and one training sequence.

Setting

Let Z\mathcal ZZ be a measurable space of samples, drawn from an unknown probability measure μ\muμ. Let H\mathcal HH be a set of hypotheses and l:H×Z→Rl : \mathcal H \times \mathcal Z \to \mathbb Rl:H×Z→R a loss with 0≤l(h,z)≤M0 \le l(h, z) \le M0≤l(h,z)≤M for all h,zh, zh,z (the paper's standing assumption, Sect. 1.1).

  • The expected loss of hhh is L(h)=Ez∼μ l(h,z)\mathcal L(h) = \mathbb E_{z \sim \mu}\, l(h, z)L(h)=Ez∼μ​l(h,z) (expectedLoss).
  • The average loss of hhh on an nnn-sample set t(n)=(t1,…,tn)\mathbf t(n) = (t_1, \dots, t_n)t(n)=(t1​,…,tn​) is L(h,t(n))=1n∑i=1nl(h,ti)L(h, \mathbf t(n)) = \frac1n \sum_{i=1}^n l(h, t_i)L(h,t(n))=n1​∑i=1n​l(h,ti​) (avgLoss).
  • A learning method A={An}n∈N\mathcal A = \{\mathcal A^n\}_{n \in \mathbb N}A={An}n∈N​ is a sequence of maps An:Zn→H\mathcal A^n : \mathcal Z^n \to \mathcal HAn:Zn→H; As(n)\mathcal A_{\mathbf s(n)}As(n)​ is the hypothesis learned from s(n)\mathbf s(n)s(n).
  • A training sequence s∗=(s1∗,s2∗,… )\mathbf s^* = (s^*_1, s^*_2, \dots)s∗=(s1∗​,s2∗​,…) is fixed and deterministic, and s∗(n)\mathbf s^*(n)s∗(n) denotes its first nnn elements (firstN).
  • A test sample t(n)\mathbf t(n)t(n) consists of nnn i.i.d. draws from μ\muμ; Pr⁡\PrPr always refers to t(n)∼μn\mathbf t(n) \sim \mu^nt(n)∼μn.

The method generalizes w.r.t. s∗\mathbf s^*s∗ (Definition 8) if

lim⁡n→∞∣L(As∗(n))−L(As∗(n),s∗(n))∣=0.\lim_{n\to\infty} \big| \mathcal L(\mathcal A_{\mathbf s^*(n)}) - L(\mathcal A_{\mathbf s^*(n)}, \mathbf s^*(n)) \big| = 0.n→∞lim​​L(As∗(n)​)−L(As∗(n)​,s∗(n))​=0.

It is weakly robust w.r.t. s∗\mathbf s^*s∗ (Definition 9) if there are sets Dn⊆Zn\mathcal D_n \subseteq \mathcal Z^nDn​⊆Zn with Pr⁡(t(n)∈Dn)→1\Pr(\mathbf t(n) \in \mathcal D_n) \to 1Pr(t(n)∈Dn​)→1 and

lim⁡n→∞{max⁡s^(n)∈Dn∣L(As∗(n),s^(n))−L(As∗(n),s∗(n))∣}=0.(6)\lim_{n\to\infty} \Big\{ \max_{\hat{\mathbf s}(n) \in \mathcal D_n} \big| L(\mathcal A_{\mathbf s^*(n)}, \hat{\mathbf s}(n)) - L(\mathcal A_{\mathbf s^*(n)}, \mathbf s^*(n)) \big| \Big\} = 0. \qquad (6)n→∞lim​{s^(n)∈Dn​max​​L(As∗(n)​,s^(n))−L(As∗(n)​,s∗(n))​}=0.(6)

A set Dn\mathcal D_nDn​ can be read as a family of perturbed copies of the training set that carries almost all of the probability of the test sample.

Formalization targets

Goal: Theorem 8 (p. 409)

A generalizes w.r.t. s∗  ⟺  A is weakly robust w.r.t. s∗,\mathcal A \text{ generalizes w.r.t. } \mathbf s^* \iff \mathcal A \text{ is weakly robust w.r.t. } \mathbf s^*,A generalizes w.r.t. s∗⟺A is weakly robust w.r.t. s∗,

for every probability measure μ\muμ, every loss measurable in zzz with values in [0,M][0, M][0,M], every learning method A\mathcal AA and every training sequence s∗\mathbf s^*s∗.

Milestones

  1. First equality of the proof (p. 410). For n≥1n \ge 1n≥1 and every hhh, Et(n)L(h,t(n))=L(h)\mathbb E_{\mathbf t(n)} L(h, \mathbf t(n)) = \mathcal L(h)Et(n)​L(h,t(n))=L(h).
  2. Sufficiency display (p. 410). If Pr⁡(t(n)∉D)≤δ\Pr(\mathbf t(n) \notin \mathcal D) \le \deltaPr(t(n)∈/D)≤δ and ∣L(h,s^)−L(h,s)∣≤ϵ|L(h, \hat{\mathbf s}) - L(h, \mathbf s)| \le \epsilon∣L(h,s^)−L(h,s)∣≤ϵ on D\mathcal DD, then
∣L(h)−L(h,s)∣≤δM+ϵ.\big|\mathcal L(h) - L(h, \mathbf s)\big| \le \delta M + \epsilon.​L(h)−L(h,s)​≤δM+ϵ.
  1. Lemma 2 (p. 410). If A\mathcal AA is not weakly robust w.r.t. s∗\mathbf s^*s∗, there are ϵ∗,δ∗>0\epsilon^*, \delta^* > 0ϵ∗,δ∗>0 with
Pr⁡(∣L(As∗(n),t(n))−L(As∗(n),s∗(n))∣≥ϵ∗)≥δ∗for infinitely many n.(8)\Pr\big(|L(\mathcal A_{\mathbf s^*(n)}, \mathbf t(n)) - L(\mathcal A_{\mathbf s^*(n)}, \mathbf s^*(n))| \ge \epsilon^*\big) \ge \delta^* \quad\text{for infinitely many } n. \qquad (8)Pr(∣L(As∗(n)​,t(n))−L(As∗(n)​,s∗(n))∣≥ϵ∗)≥δ∗for infinitely many n.(8)
  1. Eq. (9) (p. 411). L(As∗(n),t(n))−L(As∗(n))→0L(\mathcal A_{\mathbf s^*(n)}, \mathbf t(n)) - \mathcal L(\mathcal A_{\mathbf s^*(n)}) \to 0L(As∗(n)​,t(n))−L(As∗(n)​)→0 in probability.

Milestones 1–2 give the sufficiency direction; milestones 3–4 give necessity.

Significance

Theorem 8 is a characterization, not a bound. The sufficiency half says a quantitative robustness property yields generalization. The necessity half says every method that generalizes along a sequence is weakly robust along it, so no generalization argument can avoid something of this shape. The paper remarks that (K,ϵ)(K, \epsilon)(K,ϵ)-robustness for every ϵ\epsilonϵ implies weak robustness, which places Theorem 1's condition inside this characterization. Corollary 6, the almost-sure version (generalization with probability 1 iff almost-sure weak robustness), follows from Theorem 8 applied sequence by sequence.

The result is proved in the paper; no machine-checked proof of it is known. This mission contributes a formal statement of Definitions 8 and 9 in Lean, the two directions of the proof as reusable finite-nnn and asymptotic lemmas, and a place to formalize the bounded-loss law of large numbers for a hypothesis that changes with nnn (Eq. (9)), which Mathlib states for a fixed random variable.

Difficulty

The sufficiency direction is a direct estimate once the expectation of the average test loss is identified with the expected loss; the formal work is in handling the product measure μn\mu^nμn and a set Dn\mathcal D_nDn​ that need not be measurable.

The necessity direction is where care is needed. Eq. (9) is not the weak law of large numbers for a fixed function: the hypothesis As∗(n)\mathcal A_{\mathbf s^*(n)}As∗(n)​ changes with nnn, so the concentration must be uniform in the hypothesis, which holds only because the loss is uniformly bounded. Lemma 2 negates a statement with an existential over sequences of sets and a limit; the naive reading "for each ϵ,δ\epsilon, \deltaϵ,δ some Dn\mathcal D_nDn​ works eventually" does not by itself produce a single sequence Dn\mathcal D_nDn​ satisfying (6) with one limit.

Formalization scope

  • Z\mathcal ZZ is a type with a MeasurableSpace, μ\muμ a Measure with IsProbabilityMeasure, H\mathcal HH an arbitrary type. The learning method is A : (n : ℕ) → (Fin n → Z) → H, the training sequence sStar : ℕ → Z, and t(n)∼\mathbf t(n) \simt(n)∼ Measure.pi (fun _ : Fin n => μ). Indices start at 000.
  • The loss bound 0≤l≤M0 \le l \le M0≤l≤M is a hypothesis of every theorem. Measurability of l(h,⋅)l(h, \cdot)l(h,⋅) is added; the paper explicitly ignores measurability. Expectations are Bochner integrals, well defined here because the loss is bounded and measurable.
  • Probabilities and their limits live in [0,∞][0, \infty][0,∞] (ℝ≥0∞). The sets Dn\mathcal D_nDn​ need not be measurable; their probability is the outer measure. "For infinitely many nnn" is ∃ᶠ n in atTop.
  • Eq. (6) is encoded without a supremum: weak robustness asks for sets DnD_nDn​ and reals ηn→0\eta_n \to 0ηn​→0 with ∣L(As∗(n),s^)−L(As∗(n),s∗(n))∣≤ηn|L(\mathcal A_{\mathbf s^*(n)}, \hat{\mathbf s}) - L(\mathcal A_{\mathbf s^*(n)}, \mathbf s^*(n))| \le \eta_n∣L(As∗(n)​,s^)−L(As∗(n)​,s∗(n))∣≤ηn​ for all nnn and all s^∈Dn\hat{\mathbf s} \in D_ns^∈Dn​. This avoids Lean's junk value sup⁡∅=0\sup \emptyset = 0sup∅=0; since Pr⁡(t(n)∈Dn)→1\Pr(\mathbf t(n) \in D_n) \to 1Pr(t(n)∈Dn​)→1 forces DnD_nDn​ to be nonempty for all large nnn, the bound form is equivalent to the paper's reading.
  • Only part 1 of Definitions 8 and 9 is formalized. Corollary 6 is out of scope.
  • The goal is not trivial in either direction: a constant method on a one-point space satisfies both sides, and a constant method whose hypothesis has training average 111 and expected loss 1/21/21/2 along a fixed sequence fails both, so neither side is vacuous or always true.
  • Needed infrastructure: integrals over Measure.pi of coordinate functions, a Chebyshev or Hoeffding bound for averages of bounded i.i.d. variables uniform over a family of functions, and a diagonal-sequence construction. The uniform concentration lemma is reusable beyond this mission. Proofs of the milestones, and alternative routes to Eq. (9), are welcome.

Selected references

  • Huan Xu, Shie Mannor, Robustness and Generalization, Machine Learning 86 (2012) 391–423. https://doi.org/10.1007/s10994-011-5268-1
  • Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, Karthik Sridharan, Learnability, Stability and Uniform Convergence, Journal of Machine Learning Research 11 (2010) 2635–2670. https://www.jmlr.org/papers/v11/shalev-shwartz10a.html
  • Wassily Hoeffding, Probability Inequalities for Sums of Bounded Random Variables, Journal of the American Statistical Association 58 (1963) 13–30. https://doi.org/10.1080/01621459.1963.10500830
8 thms2 active usersReviewed
🏆Completed
Operations ResearchProbabilityStatistics·Captain: mikedeng1

Robustness and Generalization I: A Generalization Bound for Robust AlgorithmsResearch Paper

Why algorithmic robustness

A learning algorithm maps a training set to a hypothesis. It generalizes when the loss it incurs on the training set is close to its expected loss on fresh data. The classical way to certify this bounds the complexity of the whole hypothesis class the algorithm may output, through its VC dimension, covering numbers or Rademacher complexity. A second approach, algorithmic stability (Bousquet and Elisseeff 2002), looks instead at how the output changes when one training point is replaced.

Huan Xu and Shie Mannor proposed a third notion, algorithmic robustness. An algorithm is robust if the sample space can be cut into finitely many cells such that a test point falling in the same cell as a training point incurs nearly the same loss as that training point. The notion came out of their earlier analyses of support vector machines and the Lasso as robust optimization problems (Xu, Caramanis and Mannor 2009). The conference version appeared at COLT 2010, and the journal version, which this mission follows, is Xu and Mannor, Machine Learning 86 (2012) 391–423.

Robustness is a property of the algorithm and not of its hypothesis class, so it applies to algorithms whose class has infinite VC dimension. The paper's main result for i.i.d. data is Theorem 1 (p. 396). This mission formalizes Theorem 1 together with the steps of its proof.

Setting

Throughout, Z\mathcal ZZ is a measurable space of samples and H\mathcal HH is an arbitrary set of hypotheses. A loss l:H×Z→Rl : \mathcal H \times \mathcal Z \to \mathbb Rl:H×Z→R satisfies 0≤l(h,z)≤M0 \le l(h,z) \le M0≤l(h,z)≤M for a constant MMM. A training set is s=(s1,…,sn)∈Zn\mathbf s = (s_1, \dots, s_n) \in \mathcal Z^ns=(s1​,…,sn​)∈Zn, and a learning algorithm is a map A:Zn→H\mathcal A : \mathcal Z^n \to \mathcal HA:Zn→H, written s↦As\mathbf s \mapsto \mathcal A_{\mathbf s}s↦As​.

For a probability measure μ\muμ on Z\mathcal ZZ, the expected error and the training error of the learned hypothesis are

L(As)=Ez∼μ l(As,z),lemp(As)=1n∑i=1nl(As,si).\mathcal L(\mathcal A_{\mathbf s}) = \mathbb E_{z\sim\mu}\, l(\mathcal A_{\mathbf s}, z), \qquad l_{\mathrm{emp}}(\mathcal A_{\mathbf s}) = \frac1n \sum_{i=1}^n l(\mathcal A_{\mathbf s}, s_i).L(As​)=Ez∼μ​l(As​,z),lemp​(As​)=n1​i=1∑n​l(As​,si​).

Definition 2 (p. 396). For K∈NK \in \mathbb NK∈N and ϵ(⋅):Zn→R\epsilon(\cdot) : \mathcal Z^n \to \mathbb Rϵ(⋅):Zn→R, the algorithm A\mathcal AA is (K,ϵ(⋅))(K, \epsilon(\cdot))(K,ϵ(⋅))-robust if Z\mathcal ZZ can be partitioned into KKK disjoint sets C1,…,CKC_1, \dots, C_KC1​,…,CK​ such that for every s∈Zn\mathbf s \in \mathcal Z^ns∈Zn,

∀s∈s, ∀z∈Z, ∀i:s,z∈Ci  ⟹  ∣l(As,s)−l(As,z)∣≤ϵ(s).\forall s \in \mathbf s,\ \forall z \in \mathcal Z,\ \forall i:\quad s, z \in C_i \implies |l(\mathcal A_{\mathbf s}, s) - l(\mathcal A_{\mathbf s}, z)| \le \epsilon(\mathbf s).∀s∈s, ∀z∈Z, ∀i:s,z∈Ci​⟹∣l(As​,s)−l(As​,z)∣≤ϵ(s).

The partition is chosen once, before the training set. Only the tolerance ϵ(s)\epsilon(\mathbf s)ϵ(s) may depend on s\mathbf ss.

For a partition C1,…,CKC_1,\dots,C_KC1​,…,CK​, the cell count ∣Ni∣|N_i|∣Ni​∣ is the number of training points in CiC_iCi​. The Lean development uses expectedLoss, empiricalLoss, cellCount and IsRobust in the namespace XuMannorRobust.Standard.

Formalization targets

Goal: Theorem 1 (p. 396)

Let A\mathcal AA be (K,ϵ(⋅))(K,\epsilon(\cdot))(K,ϵ(⋅))-robust and let s\mathbf ss consist of n≥1n \ge 1n≥1 i.i.d. draws from μ\muμ. Then for every δ>0\delta > 0δ>0, with probability at least 1−δ1-\delta1−δ,

∣L(As)−lemp(As)∣≤ϵ(s)+M2Kln⁡2+2ln⁡(1/δ)n.|\mathcal L(\mathcal A_{\mathbf s}) - l_{\mathrm{emp}}(\mathcal A_{\mathbf s})| \le \epsilon(\mathbf s) + M\sqrt{\frac{2K\ln 2 + 2\ln(1/\delta)}{n}}.∣L(As​)−lemp​(As​)∣≤ϵ(s)+Mn2Kln2+2ln(1/δ)​​.

The constants are the paper's and are kept as printed. KKK, ϵ(⋅)\epsilon(\cdot)ϵ(⋅), MMM, nnn, δ\deltaδ, μ\muμ and the algorithm are all universally quantified.

Milestones (proof of Theorem 1, pp. 396–397)

  1. Bretagnolle–Huber–Carol inequality for the multinomial vector of cell counts. For every λ≥0\lambda \ge 0λ≥0,
Pr⁡{∑i=1K∣∣Ni∣n−μ(Ci)∣≥λ}≤2Kexp⁡(−nλ22).\Pr\Big\{\sum_{i=1}^K \Big|\frac{|N_i|}{n} - \mu(C_i)\Big| \ge \lambda\Big\} \le 2^K \exp\Big(\frac{-n\lambda^2}{2}\Big).Pr{i=1∑K​​n∣Ni​∣​−μ(Ci​)​≥λ}≤2Kexp(2−nλ2​).
  1. Eq. (3). With probability at least 1−δ1-\delta1−δ,
∑i=1K∣∣Ni∣n−μ(Ci)∣≤2Kln⁡2+2ln⁡(1/δ)n.\sum_{i=1}^K \Big|\frac{|N_i|}{n} - \mu(C_i)\Big| \le \sqrt{\frac{2K\ln 2 + 2\ln(1/\delta)}{n}}.i=1∑K​​n∣Ni​∣​−μ(Ci​)​≤n2Kln2+2ln(1/δ)​​.
  1. Eq. (4). For a partition witnessing robustness and for every training set s\mathbf ss, deterministically,
∣L(As)−lemp(As)∣≤ϵ(s)+M∑i=1K∣∣Ni∣n−μ(Ci)∣.|\mathcal L(\mathcal A_{\mathbf s}) - l_{\mathrm{emp}}(\mathcal A_{\mathbf s})| \le \epsilon(\mathbf s) + M\sum_{i=1}^K \Big|\frac{|N_i|}{n} - \mu(C_i)\Big|.∣L(As​)−lemp​(As​)∣≤ϵ(s)+Mi=1∑K​​n∣Ni​∣​−μ(Ci​)​.

Significance

Theorem 1 is the base result of the robustness framework. The later results of the same paper are extensions of it:

  • Corollary 1: an adaptive number of cells;
  • Corollaries 2 and 3: covering-number instances;
  • Theorem 4: a pseudo-robust version;
  • the Markovian case.

Its complexity term depends only on the number of cells KKK, not on any capacity measure of H\mathcal HH. This is why it gives bounds for algorithms such as support vector machines, Lasso, feed-forward networks and principal component analysis (Sect. 6 of the paper). For those, KKK is a covering number of the sample space. Section 8 of the paper shows that a weak form of robustness is also necessary for generalization.

Theorem 1 is a published result with a short proof. What a formalization adds:

  • a machine-checked statement of the robustness notion, pinning down which quantifier comes first;
  • a formal proof of the multinomial concentration step, which the paper takes from van der Vaart and Wellner rather than proving;
  • a reusable interface for the covering-number examples.

A search of the platform (2026-09-26) found no formal statement of Theorem 1, Definition 2, or the Bretagnolle–Huber–Carol inequality for multinomial vectors. Hoeffding's inequality is already available there in proved form.

Difficulty

The deterministic step, Eq. (4), splits the expected loss over the cells. It then compares the loss within each cell with the loss at the training points in that cell. This needs integration over a partition and some care with cells of μ\muμ-measure zero, where the conditional expectation in the paper's chain is undefined.

The main obstacle is the probabilistic step. The quantity ∑i∣∣Ni∣/n−μ(Ci)∣\sum_i ||N_i|/n - \mu(C_i)|∑i​∣∣Ni​∣/n−μ(Ci​)∣ is an ℓ1\ell_1ℓ1​ deviation of a multinomial vector. A coordinate-wise Hoeffding bound followed by a union bound over the KKK coordinates gives a bound whose deviation level grows linearly in KKK. That is not 2Ke−nλ2/22^K e^{-n\lambda^2/2}2Ke−nλ2/2, and it does not give the constant 2Kln⁡2\sqrt{2K\ln 2}2Kln2​ of Theorem 1. The difficulty is to obtain the exact exponential rate 2Ke−nλ2/22^K e^{-n\lambda^2/2}2Ke−nλ2/2 for the ℓ1\ell_1ℓ1​ deviation as a whole, with no loss in the constant.

Formalization scope

Samples are a type Z with a MeasurableSpace, training sets are Fin n → Z, the algorithm is a function (Fin n → Z) → H, and the loss is H → Z → ℝ. The partition is a family C : Fin K → Set Z that is pairwise disjoint, measurable, and covers Z. Empty cells are allowed, as in the paper. The i.i.d. sample law is Measure.pi (fun _ => μ) with μ a probability measure, and μ(Ci)\mu(C_i)μ(Ci​) enters as a real number.

"With probability at least 1−δ1-\delta1−δ" is encoded as an upper bound δ\deltaδ on the outer measure, under μn\mu^nμn, of the set of training sets where the inequality fails. This needs no measurability of s↦As\mathbf s \mapsto \mathcal A_{\mathbf s}s↦As​.

The paper ignores measurability. The formalization restores it: every l(h,⋅)l(h,\cdot)l(h,⋅) is measurable and every cell is a measurable set. Together with 0≤l≤M0 \le l \le M0≤l≤M this makes the expected error a genuine expectation.

The theorems assume n≥1n \ge 1n≥1. The Bretagnolle–Huber–Carol step assumes λ≥0\lambda \ge 0λ≥0, because the printed inequality is false for λ<0\lambda < 0λ<0. No upper bound on δ\deltaδ is imposed: for δ>2K\delta > 2^Kδ>2K the radicand is negative, the square root evaluates to 000, and the statements remain true.

Two trivializing readings of Definition 2 are ruled out:

  • The partition may not depend on the training set. In IsRobust the existential over the partition precedes the universal over training sets. If the order were swapped, every algorithm with a {0,1}\{0,1\}{0,1}-valued loss would be (2,0)(2,0)(2,0)-robust, since it could take the two level sets of its own learned loss as cells. Theorem 1 would then fail for a memorizing classifier.
  • The tolerance may not depend on the test point, and the condition is required for every z∈Zz \in \mathcal Zz∈Z, not only for zzz equal to a training point.

Beyond the paper's text, a complete development needs the integral over a finite measurable partition, a Hoeffding bound for indicator averages, and a union bound over the subsets of Fin K. The multinomial concentration inequality is reusable beyond this mission, in histogram estimators, discretization arguments and the covering-number examples of the paper. Contributions are welcome at every level: proofs of the milestones, and alternative proofs of the Bretagnolle–Huber–Carol step (for instance via the method of types).

Selected references

  • H. Xu and S. Mannor, Robustness and Generalization, Machine Learning 86 (2012) 391–423. https://doi.org/10.1007/s10994-011-5268-1
  • A. W. van der Vaart and J. A. Wellner, Weak Convergence and Empirical Processes, Springer, 1996 (Proposition A.6.6). https://doi.org/10.1007/978-1-4757-2545-2
  • O. Bousquet and A. Elisseeff, Stability and Generalization, Journal of Machine Learning Research 2 (2002) 499–526. https://www.jmlr.org/papers/v2/bousquet02a.html
  • H. Xu, C. Caramanis and S. Mannor, Robustness and Regularization of Support Vector Machines, Journal of Machine Learning Research 10 (2009) 1485–1510. https://www.jmlr.org/papers/v10/xu09b.html
  • W. Hoeffding, Probability Inequalities for Sums of Bounded Random Variables, Journal of the American Statistical Association 58 (1963) 13–30. https://doi.org/10.1080/01621459.1963.10500830
7 thms2 active usersReviewed
🏆Completed
Dynamic ProgrammingOperations ResearchProbability+1·Captain: mikedeng1

Approximately Optimal Approximate Reinforcement Learning I: Conservative Policy Iteration Improves Monotonically and Returns a Near-Greedy PolicyResearch Paper

Motivation

Approximate policy iteration and policy gradient methods are the two classical families of reinforcement learning algorithms that work with approximate, sampled information instead of an exact model. Kakade and Langford (ICML 2002) observed that neither family answers three basic questions: is there a performance measure that is guaranteed to improve at every step, how hard is it to verify that an update improves it, and what performance is reached after a reasonable number of updates. Greedy approximate policy iteration can make the policy worse when the value estimates are slightly wrong at a few states, and policy gradient methods can stall on plateaus where estimating the gradient needs an enormous number of samples.

Their answer is conservative policy iteration: instead of jumping to a greedy policy, move only a controlled fraction of the way toward it, with a step size chosen from an estimate of how much the greedy policy helps. The paper proves that this update improves a restart-distribution performance measure monotonically, terminates after a number of iterations that depends only on the reward range and the target accuracy, and stops at a policy that the greedy oracle can no longer improve by much. The idea is the direct ancestor of trust-region and proximal policy optimization methods (TRPO, Schulman et al. 2015; PPO, Schulman et al. 2017), whose improvement bounds are refinements of the paper's Theorem 4.1.

Setting

A finite Markov decision process has a finite nonempty set of states SSS, a finite nonempty set of actions AAA, transition probabilities P(s′;s,a)P(s';s,a)P(s′;s,a) (for each state sss and action aaa, a probability distribution over next states s′s's′), a reward function R:S×A→[0,R]\mathcal R : S\times A\to[0,R]R:S×A→[0,R] with R>0R>0R>0, and a discount factor 0≤γ<10\le\gamma<10≤γ<1. A stochastic policy π(a;s)\pi(a;s)π(a;s) gives, for each state sss, a probability distribution over actions.

The normalized value of π\piπ from sss is Vπ(s)=(1−γ)E[∑t≥0γtR(st,at)∣π,s]V_\pi(s) = (1-\gamma)E[\sum_{t\ge0}\gamma^t\mathcal R(s_t,a_t)\mid\pi,s]Vπ​(s)=(1−γ)E[∑t≥0​γtR(st​,at​)∣π,s], where s0=ss_0=ss0​=s, at∼π(⋅ ;st)a_t\sim\pi(\cdot\,;s_t)at​∼π(⋅;st​), st+1∼P(⋅ ;st,at)s_{t+1}\sim P(\cdot\,;s_t,a_t)st+1​∼P(⋅;st​,at​); it lies in [0,R][0,R][0,R]. The state-action value is Qπ(s,a)=(1−γ)R(s,a)+γEs′∼P(s′;s,a)[Vπ(s′)]Q_\pi(s,a) = (1-\gamma)\mathcal R(s,a)+\gamma E_{s'\sim P(s';s,a)}[V_\pi(s')]Qπ​(s,a)=(1−γ)R(s,a)+γEs′∼P(s′;s,a)​[Vπ​(s′)] and the advantage is Aπ(s,a)=Qπ(s,a)−Vπ(s)∈[−R,R]A_\pi(s,a) = Q_\pi(s,a)-V_\pi(s)\in[-R,R]Aπ​(s,a)=Qπ​(s,a)−Vπ​(s)∈[−R,R].

For a state distribution μ\muμ (a restart distribution), the discounted future state distribution is dπ,μ(s)=(1−γ)∑t≥0γtPr⁡(st=s;π,μ)d_{\pi,\mu}(s) = (1-\gamma)\sum_{t\ge0}\gamma^t\Pr(s_t=s;\pi,\mu)dπ,μ​(s)=(1−γ)∑t≥0​γtPr(st​=s;π,μ) (eq. (2.1)), and the performance measure is ημ(π)=Es∼μ[Vπ(s)]\eta_\mu(\pi) = E_{s\sim\mu}[V_\pi(s)]ημ​(π)=Es∼μ​[Vπ​(s)].

The policy advantage of a policy π′\pi'π′ with respect to π\piπ and μ\muμ is

Aπ,μ(π′)=Es∼dπ,μ[Ea∼π′(a;s)[Aπ(s,a)]],\mathbb A_{\pi,\mu}(\pi') = E_{s\sim d_{\pi,\mu}}\big[E_{a\sim\pi'(a;s)}[A_\pi(s,a)]\big],Aπ,μ​(π′)=Es∼dπ,μ​​[Ea∼π′(a;s)​[Aπ​(s,a)]],

and OPT(Aπ,μ)=max⁡π′Aπ,μ(π′)\mathrm{OPT}(\mathbb A_{\pi,\mu}) = \max_{\pi'}\mathbb A_{\pi,\mu}(\pi')OPT(Aπ,μ​)=maxπ′​Aπ,μ​(π′). The conservative update (4.1) is πnew=(1−α)π+απ′\pi_{new} = (1-\alpha)\pi+\alpha\pi'πnew​=(1−α)π+απ′ with α∈[0,1]\alpha\in[0,1]α∈[0,1]. An ε\varepsilonε-greedy policy chooser GεG_\varepsilonGε​ (Definition 4.3) returns, for every policy π\piπ, a policy π′\pi'π′ with Aπ,μ(π′)≥OPT(Aπ,μ)−ε\mathbb A_{\pi,\mu}(\pi')\ge\mathrm{OPT}(\mathbb A_{\pi,\mu})-\varepsilonAπ,μ​(π′)≥OPT(Aπ,μ​)−ε.

Conservative policy iteration (§5) starts from any policy and repeats: call Gε(π,μ)G_\varepsilon(\pi,\mu)Gε​(π,μ) to get π′\pi'π′; form an ε3\frac\varepsilon33ε​-accurate estimate A^\hat{\mathbb A}A^ of Aπ,μ(π′)\mathbb A_{\pi,\mu}(\pi')Aπ,μ​(π′) from μ\muμ-restarts; if A^<2ε3\hat{\mathbb A}<\frac{2\varepsilon}3A^<32ε​, stop and return π\piπ; otherwise apply (4.1) with α=(1−γ)(A^−ε/3)4R\alpha = \frac{(1-\gamma)(\hat{\mathbb A}-\varepsilon/3)}{4R}α=4R(1−γ)(A^−ε/3)​ and repeat.

Formalization targets

Goal: Theorem 4.4 (p. 5)

With probability at least 1−δ1-\delta1−δ, conservative policy iteration (i) strictly improves ημ\eta_\muημ​ with every policy update, (ii) stops after at most 72R2/ε272R^2/\varepsilon^272R2/ε2 policy updates, and (iii) returns a policy π\piπ with

OPT(Aπ,μ)<2ε.\mathrm{OPT}(\mathbb A_{\pi,\mu}) < 2\varepsilon.OPT(Aπ,μ​)<2ε.

The estimation step is represented by its guarantee: each reached loop's estimate fails to be ε3\frac\varepsilon33ε​-accurate with probability at most δ/(N+1)\delta/(N+1)δ/(N+1), N=⌊72R2/ε2⌋N=\lfloor72R^2/\varepsilon^2\rfloorN=⌊72R2/ε2⌋.

Milestones

Lemma 6.1 (p. 6), the performance difference identity:

ημ(π~)−ημ(π)=11−γE(a,s)∼π~dπ~,μ[Aπ(s,a)].\eta_\mu(\tilde\pi)-\eta_\mu(\pi) = \frac1{1-\gamma}E_{(a,s)\sim\tilde\pi d_{\tilde\pi,\mu}}[A_\pi(s,a)].ημ​(π~)−ημ​(π)=1−γ1​E(a,s)∼π~dπ~,μ​​[Aπ​(s,a)].

Theorem 4.1 (p. 4), with ε=max⁡s∣Ea∼π′(a;s)[Aπ(s,a)]∣\varepsilon=\max_s|E_{a\sim\pi'(a;s)}[A_\pi(s,a)]|ε=maxs​∣Ea∼π′(a;s)​[Aπ​(s,a)]∣ and all α∈[0,1]\alpha\in[0,1]α∈[0,1]:

ημ(πnew)−ημ(π)≥α1−γ(A−2αγε1−γ(1−α)).\eta_\mu(\pi_{new})-\eta_\mu(\pi)\ge\frac{\alpha}{1-\gamma}\Big(\mathbb A-\frac{2\alpha\gamma\varepsilon}{1-\gamma(1-\alpha)}\Big).ημ​(πnew​)−ημ​(π)≥1−γα​(A−1−γ(1−α)2αγε​).

Corollary 4.2 (p. 5): if A≥0\mathbb A\ge0A≥0, the step size α=(1−γ)A4R\alpha=\frac{(1-\gamma)\mathbb A}{4R}α=4R(1−γ)A​ gives

ημ(πnew)−ημ(π)≥A28R.\eta_\mu(\pi_{new})-\eta_\mu(\pi)\ge\frac{\mathbb A^2}{8R}.ημ​(πnew​)−ημ​(π)≥8RA2​.

Significance

Theorem 4.4 is the first guarantee of its kind for approximate reinforcement learning: the number of iterations is bounded by 72R2/ε272R^2/\varepsilon^272R2/ε2, independent of the number of states and of the restart distribution, and every iteration provably helps. Lemma 6.1 is the standard performance difference lemma, used throughout the analysis of policy optimization, including natural policy gradient and trust-region methods; Theorem 4.1 is the prototype of the "surrogate objective minus a penalty" bound that TRPO refines.

These results are proved in the paper. As far as a search of the platform shows, none is formalized: the platform's finite-horizon performance difference lemma (Foster and Rakhlin's Lemma 13) is a different statement, for episodic problems with non-stationary policies. This mission produces machine-checked versions of the discounted performance difference identity, the conservative improvement bound with its exact constants, and the high-probability termination and quality guarantee of the algorithm, all on top of an explicit infinite-horizon model rather than an assumed Bellman equation.

Difficulty

The obvious argument for the improvement bound expands ημ(πnew)\eta_\mu(\pi_{new})ημ​(πnew​) to first order in α\alphaα; that only gives α1−γA+O(α2)\frac{\alpha}{1-\gamma}\mathbb A+O(\alpha^2)1−γα​A+O(α2) with an unspecified constant, which cannot fix a step size. The exact bound needs control of how far the state distribution of the mixed policy drifts from that of the old policy, uniformly in time, and the performance difference identity is only useful once the states are weighted by the new policy's distribution. On the formal side, VπV_\piVπ​ and dπ,μd_{\pi,\mu}dπ,μ​ are infinite discounted series, so summability, exchanges of sums and the identities ∑sdπ,μ(s)=1\sum_s d_{\pi,\mu}(s)=1∑s​dπ,μ​(s)=1 and ∑aπ(a;s)Aπ(s,a)=0\sum_a\pi(a;s)A_\pi(s,a)=0∑a​π(a;s)Aπ​(s,a)=0 must all be established from the definitions. For Theorem 4.4, the algorithm is a random process whose policies depend on all earlier estimates; the argument has to be made pathwise on the event that every reached loop is accurate, together with a union bound over the loops that can be reached.

Formalization scope

Policies are functions π : S → A → ℝ with π s a the paper's π(a;s)\pi(a;s)π(a;s), and P s a s' is P(s′;s,a)P(s';s,a)P(s′;s,a); both are constrained by the published predicates IsPolicy and IsTransitionKernel. VπV_\piVπ​ is (1−γ)(1-\gamma)(1−γ) times the published series PolicyValue, so values are normalized as in the paper. OPT\mathrm{OPT}OPT is a real supremum over all stochastic policies; the set is nonempty and bounded, and the maximum is attained. Every theorem carries the standing assumptions of §2: finite nonempty SSS and AAA, a transition kernel, rewards in [0,R][0,R][0,R] with R>0R>0R>0, 0≤γ<10\le\gamma<10≤γ<1, and a state distribution μ\muμ. In Corollary 4.2, RRR is any upper bound on the rewards rather than necessarily the attained maximum.

In Theorem 4.4 the run is formalized pathwise, driven by arbitrary real random estimates on a probability space; the conclusion bounds the probability of the failure event by δ\deltaδ. Two deviations from the printed statement are disclosed. First, (ii) is stated for policy updates: the proof bounds updates, and the algorithm calls GεG_\varepsilonGε​ once more than it updates, so "at most 72R2/ε272R^2/\varepsilon^272R2/ε2 calls" is off by one. Second, the per-loop failure budget is δ/(N+1)\delta/(N+1)δ/(N+1), which covers the N+1N+1N+1 loops that may be reached. The Hoeffding estimate (5.1) is not formalized: as printed it concerns the ε6\frac\varepsilon66ε​-biased target, and its role is taken by the accuracy hypothesis. The step size is clipped at 111, which never binds when the estimate is accurate. No trivializing reading is available: the accuracy hypothesis is satisfied by a perfect estimator and a 000-greedy chooser exists, so the theorem is not vacuous, and strict improvement at every update is required, not merely nonnegative change.

Pages are PDF pages; the paper has no printed page numbers.

A complete development needs summability and algebra of discounted occupation measures, the performance difference identity, and a union bound over the loops of a random process; the first two are reusable for any discounted policy-optimization result. Proofs of the milestones in any order are welcome.

Selected references

  • S. Kakade and J. Langford, Approximately Optimal Approximate Reinforcement Learning, Proceedings of the 19th International Conference on Machine Learning (ICML), 2002. https://dl.acm.org/doi/10.5555/645531.656005
  • J. Schulman, S. Levine, P. Moritz, M. Jordan, P. Abbeel, Trust Region Policy Optimization, ICML 2015. https://arxiv.org/abs/1502.05477
  • J. Schulman, F. Wolski, P. Dhariwal, A. Radford, O. Klimov, Proximal Policy Optimization Algorithms, 2017. https://arxiv.org/abs/1707.06347
  • D. J. Foster and A. Rakhlin, Foundations of Reinforcement Learning and Interactive Decision Making, 2023. https://arxiv.org/abs/2312.16730
13 thms2 active usersReviewed
🏆Completed
Bandit AlgorithmsOperations ResearchStatistics·Captain: mikedeng1

Optimal Best Arm Identification with Fixed Confidence II: Characterization of the Optimal Proportions of Arm DrawsResearch Paper

Motivation

In best arm identification with fixed confidence, a learner samples KKK unknown distributions ("arms") sequentially and must name the arm with the largest mean, with error probability at most a prescribed δ\deltaδ, using as few samples as possible. Garivier and Kaufmann (arXiv:1602.04589, COLT 2016) showed that every δ\deltaδ-PAC strategy needs, in expectation, at least T∗(μ) kl(δ,1−δ)T^*(\boldsymbol\mu)\,\mathrm{kl}(\delta,1-\delta)T∗(μ)kl(δ,1−δ) samples, where the characteristic time T∗(μ)T^*(\boldsymbol\mu)T∗(μ) is the value of a max–min optimization problem over the proportions of draws allocated to the arms. The maximizer of that problem, w∗(μ)w^*(\boldsymbol\mu)w∗(μ), is the allocation any asymptotically optimal strategy must follow; the Track-and-Stop algorithm of the same paper computes w∗(μ^)w^*(\hat{\boldsymbol\mu})w∗(μ^​) at plug-in estimates and tracks it.

A strategy can only track w∗w^*w∗ if w∗w^*w∗ can be computed. This mission formalizes the part of the paper (Section 2.2 and Appendix A) that turns the abstract max–min problem into an explicit recipe: a closed form for the inner infimum, and a characterization of w∗w^*w∗ through the root of one increasing scalar function. The problem had been solved in closed form before only for special cases, such as Poisson rewards with all suboptimal arms equal (Vaidhyan and Sundaresan, 2015); the paper's result covers every one-parameter exponential family.

Setting

A canonical one-parameter exponential family is a family of laws νθ\nu_\thetaνθ​, θ∈Θ\theta\in\Thetaθ∈Θ, on R\mathbb RR with density exp⁡(θx−b(θ))\exp(\theta x-b(\theta))exp(θx−b(θ)) with respect to a reference measure ξ\xiξ. The law νθ\nu_\thetaνθ​ has mean b˙(θ)\dot b(\theta)b˙(θ); the set of attainable means is the mean space b˙(Θ)\dot b(\Theta)b˙(Θ). For means μ=b˙(θ)\mu=\dot b(\theta)μ=b˙(θ) and μ′=b˙(θ′)\mu'=\dot b(\theta')μ′=b˙(θ′) the Kullback–Leibler divergence is

d(μ,μ′)=KL(νθ,νθ′)=b(θ′)−b(θ)−b˙(θ)(θ′−θ).d(\mu,\mu')=\mathrm{KL}(\nu_\theta,\nu_{\theta'})=b(\theta')-b(\theta)-\dot b(\theta)(\theta'-\theta).d(μ,μ′)=KL(νθ​,νθ′​)=b(θ′)−b(θ)−b˙(θ)(θ′−θ).

Bernoulli laws and Gaussian laws with known variance are the standard examples.

A bandit model is identified with its vector of means μ=(μ1,…,μK)∈b˙(Θ)K\boldsymbol\mu=(\mu_1,\dots,\mu_K)\in\dot b(\Theta)^Kμ=(μ1​,…,μK​)∈b˙(Θ)K. S\mathcal SS is the set of models with a unique optimal arm a∗(μ)a^*(\boldsymbol\mu)a∗(μ), and Alt(μ)={λ∈S:a∗(λ)≠a∗(μ)}\mathrm{Alt}(\boldsymbol\mu)=\{\boldsymbol\lambda\in\mathcal S:a^*(\boldsymbol\lambda)\ne a^*(\boldsymbol\mu)\}Alt(μ)={λ∈S:a∗(λ)=a∗(μ)} is the set of alternatives. ΣK\Sigma_KΣK​ is the probability simplex. The transportation cost of proportions w∈ΣKw\in\Sigma_Kw∈ΣK​ and the objects of the paper are

cμ(w)=inf⁡λ∈Alt(μ)∑a=1Kwa d(μa,λa),T∗(μ)−1=sup⁡w∈ΣKcμ(w),w∗(μ)=argmax⁡w∈ΣKcμ(w).c_{\boldsymbol\mu}(w)=\inf_{\boldsymbol\lambda\in\mathrm{Alt}(\boldsymbol\mu)}\sum_{a=1}^Kw_a\,d(\mu_a,\lambda_a),\qquad T^*(\boldsymbol\mu)^{-1}=\sup_{w\in\Sigma_K}c_{\boldsymbol\mu}(w),\qquad w^*(\boldsymbol\mu)=\operatorname*{argmax}_{w\in\Sigma_K}c_{\boldsymbol\mu}(w).cμ​(w)=λ∈Alt(μ)inf​a=1∑K​wa​d(μa​,λa​),T∗(μ)−1=w∈ΣK​sup​cμ​(w),w∗(μ)=w∈ΣK​argmax​cμ​(w).

The arms are sorted so that μ1>μ2≥⋯≥μK\mu_1>\mu_2\ge\dots\ge\mu_Kμ1​>μ2​≥⋯≥μK​. The parameterized Jensen–Shannon divergence is, for α∈[0,1]\alpha\in[0,1]α∈[0,1],

Iα(μ1,μ2)=α d(μ1,αμ1+(1−α)μ2)+(1−α) d(μ2,αμ1+(1−α)μ2).I_\alpha(\mu_1,\mu_2)=\alpha\,d\big(\mu_1,\alpha\mu_1+(1-\alpha)\mu_2\big)+(1-\alpha)\,d\big(\mu_2,\alpha\mu_1+(1-\alpha)\mu_2\big).Iα​(μ1​,μ2​)=αd(μ1​,αμ1​+(1−α)μ2​)+(1−α)d(μ2​,αμ1​+(1−α)μ2​).

For a∈{2,…,K}a\in\{2,\dots,K\}a∈{2,…,K} let ga(x)=(1+x)I1/(1+x)(μ1,μa)g_a(x)=(1+x)I_{1/(1+x)}(\mu_1,\mu_a)ga​(x)=(1+x)I1/(1+x)​(μ1​,μa​) for x≥0x\ge0x≥0, let xa=ga−1x_a=g_a^{-1}xa​=ga−1​, and let x1≡1x_1\equiv1x1​≡1. Finally

Fμ(y)=∑a=2Kd(μ1,ma(y))d(μa,ma(y)),ma(y)=μ1+xa(y)μa1+xa(y).F_{\boldsymbol\mu}(y)=\sum_{a=2}^K\frac{d\big(\mu_1,m_a(y)\big)}{d\big(\mu_a,m_a(y)\big)},\qquad m_a(y)=\frac{\mu_1+x_a(y)\mu_a}{1+x_a(y)}.Fμ​(y)=a=2∑K​d(μa​,ma​(y))d(μ1​,ma​(y))​,ma​(y)=1+xa​(y)μ1​+xa​(y)μa​​.

Formalization targets

Goal: Theorem 5 (p. 5)

With D=d(μ1,μ2)D=d(\mu_1,\mu_2)D=d(μ1​,μ2​): FμF_{\boldsymbol\mu}Fμ​ is continuous and strictly increasing on [0,D[[0,D[[0,D[, Fμ(0)=0F_{\boldsymbol\mu}(0)=0Fμ​(0)=0, Fμ(y)→∞F_{\boldsymbol\mu}(y)\to\inftyFμ​(y)→∞ as y→Dy\to Dy→D, the equation Fμ(y)=1F_{\boldsymbol\mu}(y)=1Fμ​(y)=1 has a unique solution y∗∈[0,D[y^*\in[0,D[y∗∈[0,D[, and

w∈w∗(μ)  ⟺  wa=xa(y∗)∑i=1Kxi(y∗)for every arm a.w\in w^*(\boldsymbol\mu)\iff w_a=\frac{x_a(y^*)}{\sum_{i=1}^Kx_i(y^*)}\quad\text{for every arm }a.w∈w∗(μ)⟺wa​=∑i=1K​xi​(y∗)xa​(y∗)​for every arm a.

The equivalence says at once that the argmax exists, that it is a single point, and that it is given by eq. (5).

Milestones

  1. Lemma 3 (p. 5): for every w∈ΣKw\in\Sigma_Kw∈ΣK​,
cμ(w)=min⁡a≠1(w1+wa) Iw1w1+wa(μ1,μa).c_{\boldsymbol\mu}(w)=\min_{a\ne1}(w_1+w_a)\,I_{\frac{w_1}{w_1+w_a}}(\mu_1,\mu_a).cμ​(w)=a=1min​(w1​+wa​)Iw1​+wa​w1​​​(μ1​,μa​).
  1. Claim after eq. (4) (p. 5): gag_aga​ is a strictly increasing one-to-one mapping from [0,+∞[[0,+\infty[[0,+∞[ onto [0,d(μ1,μa)[[0,d(\mu_1,\mu_a)[[0,d(μ1​,μa​)[.
  2. Lemma 4 (p. 5): for every maximizer w∗w^*w∗ and all a,b∈{2,…,K}a,b\in\{2,\dots,K\}a,b∈{2,…,K},
(w1∗+wa∗)Iw1∗w1∗+wa∗(μ1,μa)=(w1∗+wb∗)Iw1∗w1∗+wb∗(μ1,μb).(w^*_1+w^*_a)I_{\frac{w^*_1}{w^*_1+w^*_a}}(\mu_1,\mu_a)=(w^*_1+w^*_b)I_{\frac{w^*_1}{w^*_1+w^*_b}}(\mu_1,\mu_b).(w1∗​+wa∗​)Iw1∗​+wa∗​w1∗​​​(μ1​,μa​)=(w1∗​+wb∗​)Iw1∗​+wb∗​w1∗​​​(μ1​,μb​).

Significance

The result. Theorem 5 reduces a (K−1)(K-1)(K−1)-dimensional non-smooth max–min problem to finding the root of one continuous increasing function on a bounded interval, each evaluation of which requires K−1K-1K−1 scalar inversions. It gives existence and uniqueness of w∗(μ)w^*(\boldsymbol\mu)w∗(μ), which the paper's lower bound only presupposes, and it is the computational core of Track-and-Stop: without an explicit, well-posed w∗w^*w∗ the tracking strategy is not defined. Lemma 3 alone gives the closed form of the inner infimum used again in the analysis of the stopping rule.

Formalizing it. The results are proved in the paper; nothing here is open. To the best of our knowledge none of them has a machine-checked proof: the platform's existing best-arm-identification rows concern Gaussian arms and state the characteristic time at the level of measures, without this characterization. The mission produces a checked reduction for general one-parameter exponential families, including the edge cases the text passes over (ties among suboptimal arms, zero weights, the behaviour of xax_axa​ near the end of its domain).

Difficulty

The infimum in Lemma 3 ranges over Alt(μ)\mathrm{Alt}(\boldsymbol\mu)Alt(μ), a set of models with a unique best arm, so it is an open condition: the minimizing configuration, in which λ1\lambda_1λ1​ and λa\lambda_aλa​ coincide, lies outside Alt(μ)\mathrm{Alt}(\boldsymbol\mu)Alt(μ) and is only approached. Other arms may also compete for the best position. A statement in which the infimum is taken over the closed relaxation {λa≥λ1}\{\lambda_a\ge\lambda_1\}{λa​≥λ1​} is a lemma of the proof, not Lemma 3.

For Theorem 5, the equalization in Lemma 4 needs an argument that holds for every maximizer, not only for one found by a first-order condition, because the objective is a minimum of functions and is not differentiable. The monotonicity of FμF_{\boldsymbol\mu}Fμ​ needs the monotonicity of each xax_axa​ and of each ratio in the moving point mam_ama​, and the limit at DDD rests on the second-best arm(s) only, which is where the ordering μ1>μ2≥…\mu_1>\mu_2\ge\dotsμ1​>μ2​≥… enters. Finally the analytic facts about ddd (continuity, positivity off the diagonal, monotonicity in each argument) must be derived from the exponential family itself.

Formalization scope

Lean represents a model by μ : Fin K → ℝ with 2 ≤ K and every μ a in the mean space deriv F.b '' F.Θ. The paper's arm 111 is index 0, arm 222 is index 1. The exponential family is a structure ExpFamily whose parameter set is a nonempty open interval and whose b is twice continuously differentiable with b¨>0\ddot b>0b¨>0 on Θ\ThetaΘ. These two conditions are added to the paper's "convex, twice differentiable" and are disclosed: strict convexity is what makes the mean parameterization unique, and openness is what lets alternatives approach the boundary of Alt(μ)\mathrm{Alt}(\boldsymbol\mu)Alt(μ). The reference measure and normalization are part of the structure but unused here.

The transportation cost is an infimum in EReal, so it is the true infimum of the set of values rather than a default 0. w∗(μ)w^*(\boldsymbol\mu)w∗(μ) is never defined by choice: "www is optimal" is a predicate, and Theorem 5 characterizes the set of such www. The functions xax_axa​ are the inverse of gag_aga​ on [0,+∞[[0,+\infty[[0,+∞[ and are evaluated only on [0,d(μ1,μ2)[[0,d(\mu_1,\mu_2)[[0,d(μ1​,μ2​)[. "Increasing" in Theorem 5 is stated as strictly increasing, as proved in Appendix A.2. Lemmas 3 and 4 and the claim on gag_aga​ assume only that arm 111 is the unique best arm, which is weaker than the paper's standing ordering.

A formalization that replaces Alt(μ)\mathrm{Alt}(\boldsymbol\mu)Alt(μ) by {λa≥λ1}\{\lambda_a\ge\lambda_1\}{λa​≥λ1​}, assumes the maximizer exists and is unique, or asserts only existence of some y∗y^*y∗ without the formula for w∗w^*w∗, does not state these results and is ruled out.

A complete development needs: basic calculus of exponential families (the Bregman form of ddd, its continuity and strict positivity off the diagonal, its monotonicity in the second argument), the inverse function of a continuous strictly monotone map on an interval, and compactness of the simplex. The divergence facts are reusable in every bandit mission built on exponential families. Proofs of the milestones, of auxiliary facts about ddd and IαI_\alphaIα​, and a sorry-free instance of ExpFamily (Bernoulli or unit-variance Gaussian) are welcome.

Selected references

  • A. Garivier, E. Kaufmann, Optimal Best Arm Identification with Fixed Confidence, COLT 2016 (JMLR W&CP 49), arXiv:1602.04589v2. https://arxiv.org/abs/1602.04589
  • E. Kaufmann, O. Cappé, A. Garivier, On the Complexity of Best-Arm Identification in Multi-Armed Bandit Models, JMLR 17, 2016. https://arxiv.org/abs/1407.4443
  • N. K. Vaidhyan, R. Sundaresan, Learning to detect an oddball target, arXiv:1508.05572, 2015. https://arxiv.org/abs/1508.05572
  • O. Cappé, A. Garivier, O.-A. Maillard, R. Munos, G. Stoltz, Kullback–Leibler upper confidence bounds for optimal sequential allocation, Annals of Statistics 41(3), 2013. https://arxiv.org/abs/1210.1136
6 thms2 active usersReviewed
PreviousPage 6 of 11Next

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me