Prove2Me
Navigate
DiscoverFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

Machine Learning

273 missions · 182 completed

The science of systems that learn from data and experience. Its scope runs from the statistical and mathematical foundations of learning, including generalization, expressivity, and computational limits, through the design of learning algorithms, deep learning, reinforcement learning, and probabilistic methods, to the empirical study of large models and the trustworthiness, interpretability, and societal impact of learned systems.

Missions

Open91Completed182All273
🏆Completed
Convex OptimizationOptimization·Captain: mikedeng1

Introduction to Online Convex Optimization VII: The Online Conditional Gradient AlgorithmTextbook

Motivation

Every algorithm through Chapter VI updates its iterate by a Euclidean projection onto the decision set KKK. For many decision sets that arise in practice — bounded-nuclear-norm matrices (matrix completion / recommendation systems), the flow polytope (network routing), the Birkhoff–von Neumann polytope (ranking/permutations), matroid polytopes — a projection requires an expensive operation (an SVD, a quadratic program) while a linear minimization over the same set is comparatively cheap (an eigenvector computation via the power method, a shortest-path or minimum-weight-matching computation, a greedy matroid algorithm). Chapter 7 develops an OCO algorithm that replaces every projection with a call to a linear-minimization oracle, at the cost of a worse regret rate.

Setting

The conditional gradient (CG) / Frank–Wolfe method (Algorithm 25) minimizes a β\betaβ-smooth function fff over a convex set KKK (diameter DDD) without ever projecting: at each round it calls the oracle vt=arg⁡min⁡x∈K⟨x,∇f(xt)⟩v_t = \arg\min_{x\in K}\langle x, \nabla f(x_t)\ranglevt​=argminx∈K​⟨x,∇f(xt​)⟩ and steps xt+1=xt+ηt(vt−xt)x_{t+1} = x_t + \eta_t(v_t - x_t)xt+1​=xt​+ηt​(vt​−xt​), staying inside KKK automatically since it is a convex combination of two points of KKK. Theorem 7.1 gives its convergence rate; §7.3.1's matrix completion example and §7.4's routing/ranking/matroid examples motivate why the oracle call is often much cheaper than a projection.

The online conditional gradient (OCG) algorithm (Algorithm 27) lifts this to the OCO setting. Applying CG naively to each ftf_tft​ separately fails (the method only sees gradient direction, and a single round's direction is not enough information); instead, the algorithm builds the aggregate regularized function Ft(x)=η∑τ=1t−1⟨∇τ,x⟩+∥x−x1∥2F_t(x) = \eta\sum_{\tau=1}^{t-1}\langle\nabla_\tau, x\rangle + \|x-x_1\|^2Ft​(x)=η∑τ=1t−1​⟨∇τ​,x⟩+∥x−x1​∥2 from all past gradients, calls the linear oracle on ∇Ft(xt)\nabla F_t(x_t)∇Ft​(xt​), and takes a (1−σt)/σt(1-\sigma_t)/\sigma_t(1−σt​)/σt​-weighted step toward the oracle's answer.

Formalization targets

Theorem 7.1 (offline CG convergence, milestone)

ht≤2βD2t,t≥1,ht:=f(xt)−f(x⋆).h_t \le \frac{2\beta D^2}{t}, \quad t \ge 1, \qquad h_t := f(x_t) - f(x^\star).ht​≤t2βD2​,t≥1,ht​:=f(xt​)−f(x⋆).

Lemma 7.4 (per-round iterate bound, milestone)

ht≤2D2σt,t≥1,ht:=Ft(xt)−Ft(xt⋆),  xt⋆:=arg⁡min⁡x∈KFt(x).h_t \le 2D^2\sigma_t, \quad t \ge 1, \qquad h_t := F_t(x_t) - F_t(x^\star_t),\ \ x^\star_t := \arg\min_{x\in K} F_t(x).ht​≤2D2σt​,t≥1,ht​:=Ft​(xt​)−Ft​(xt⋆​),  xt⋆​:=argx∈Kmin​Ft​(x).

Theorem 7.3 — the mission's goal

Online conditional gradient (Algorithm 27) with η=D/(2GT3/4)\eta = D/(2GT^{3/4})η=D/(2GT3/4), σt=min⁡{1,2/t}\sigma_t = \min\{1, 2/\sqrt t\}σt​=min{1,2/t​} attains

RegretT=∑t=1Tft(xt)−min⁡x⋆∈K∑t=1Tft(x⋆)≤8DGT3/4.\mathrm{Regret}_T = \sum_{t=1}^T f_t(x_t) - \min_{x^\star\in K}\sum_{t=1}^T f_t(x^\star) \le 8DGT^{3/4}.RegretT​=t=1∑T​ft​(xt​)−x⋆∈Kmin​t=1∑T​ft​(x⋆)≤8DGT3/4.

Significance

This is the chapter's central trade: Algorithm 27's O(T3/4)O(T^{3/4})O(T3/4) regret is worse than Chapter III's full-information O(T)O(\sqrt T)O(T​) rate and Chapter V's RFTL rate, but its per-round cost is a single linear-minimization oracle call, not a projection — exactly the trade that makes it the practical choice for the recommendation-system, routing, and ranking applications the chapter develops in detail. Theorem 7.1's offline rate is independently significant as the field's standard Frank–Wolfe convergence guarantee, reused as the analytical engine (via Eq. (7.2)) for both Lemma 7.4's online bound and, historically, for a large family of projection-free methods outside OCO entirely. No prior art was found on the platform for Frank–Wolfe, conditional gradient, or projection-free methods (q=Frank-Wolfe returned 0 hits during planning); this mission drafts the standard textbook account fresh.

Difficulty

Theorem 7.1's proof is a one-step smoothness-plus-convexity inequality (Eq. (7.2)) combined with an induction lemma (Lemma 7.2, not separately drafted — it is a purely algebraic recursion h_{t+1} ≤ h_t(1-η_t) + η_t²c ⟹ h_t ≤ 4c/t, reused verbatim by Lemma 7.4's own induction and not independently central to the chapter's content). Lemma 7.4's proof is the chapter's most delicate step: it applies Theorem 7.1's offline analysis technique to the online aggregate function FtF_tFt​ — not to any single ftf_tft​, and not even to a fixed function across rounds, since FtF_tFt​ itself changes every round as more gradients accumulate — then combines it with a second inequality (comparing Ft(xt⋆)F_t(x^\star_t)Ft​(xt⋆​) to Ft+1(xt+1⋆)F_{t+1}(x^\star_{t+1})Ft+1​(xt+1⋆​) via strong convexity and Cauchy–Schwarz) and a careful algebraic balancing of the η\etaη, GGG, σt\sigma_tσt​ parameters (Eq. (7.6)) to close the induction. Theorem 7.3's own proof is a second reduction: it relates the algorithm's regret against the true cost sequence ftf_tft​ to Lemma 7.4's bound on FtF_tFt​, via an intermediate comparison to xt⋆x^\star_txt⋆​ (playing the role of Chapter V's RFTL iterates applied to a shifted cost sequence f~t\tilde f_tf~​t​).

Formalization scope

IsLinearMinimizer makes the "projection-free" linear-oracle call (Eq. (7.4)) an explicit, first-class object, reused by both Algorithm 25 and Algorithm 27's definitions, rather than silently replaced by a projection anywhere. SmoothOn is redeclared under this chapter's own sub-namespace (not imported from Chapter II, which is not yet a published series definition); see MODERATION_NOTES.md. AggregateFunction/AggregateGradient give FtF_tFt​ and its closed-form gradient explicitly, matching Algorithm 27 line 4's formula exactly (the book computes ∇Ft\nabla F_t∇Ft​ directly rather than leaving it abstract, so this mission does too). This chunk indexes rounds from 1 throughout (not the 0-indexed Finset.range shift used elsewhere in the series), since Algorithm 27's own line 4 sums τ=1\tau=1τ=1 to t−1t-1t−1 and every theorem in this chapter states a per-round or Finset.Icc 1 T-summed bound directly in the book's own round numbers — a deliberate, chunk-local convention choice, not an inconsistency with earlier chapters' definitions (this chunk does not import them). Lemma 7.4 keeps Theorem 7.3's specific parameters and a GGG-Lipschitz hypothesis as explicit premises, since the book's own proof of the lemma uses them, rather than presenting it as a fully parameter-free general fact.

Not formalized: Lemma 7.2 (a routine algebraic recursion, not independently central, and reused identically inside Lemma 7.4's own proof rather than cited as a numbered result on its own); Algorithm 26 and §7.3.1's matrix-completion specialization, §7.4's routing/ranking/matroid examples, and Corollary-level results (illustrative applications, not further formalizable theorems); §7.1's linear-algebra review (singular values, nuclear norm — background, not a formalization target for this mission).

Selected references

  • E. Hazan, Introduction to Online Convex Optimization, 2nd ed., arXiv:1909.05207v3, Chapter 7.
  • M. Frank, P. Wolfe, "An algorithm for quadratic programming," Naval Research Logistics Quarterly 3(1-2), 1956, 95-110.
  • E. Hazan, S. Kale, "Projection-free online learning," ICML 2012 (the chapter's Algorithm 27).
7 thms3 active usersReviewed
🏆Completed
Reinforcement Learning·Captain: mikedeng1

Foundations of Reinforcement Learning VI: Function Approximation and Bellman RankTextbook

Motivation

Every RL guarantee proved earlier in this series — UCB-VI's regret bound, the contextual bandit oracle reductions — scales with the size of the state space SSS, because the algorithms maintain a separate statistic per state. Real environments (images, sensor readouts, natural language) have combinatorially or infinitely many states, so a tabular guarantee is vacuous there: the only hope is to generalize across states via a class of value functions, the way supervised learning generalizes across inputs via a hypothesis class. The chapter develops two algorithms along this line: LSVI-UCB, the linear-function-approximation analogue of UCB-VI whose regret is independent of ∣S∣|S|∣S∣ (a low-rank MDP result originating with Jin, Yang, Wang, and Jordan, Provably Efficient Reinforcement Learning with Linear Function Approximation, COLT 2020, arXiv:1907.05388), and BiLinUCB, which attains an analogous sample-complexity guarantee under the strictly more general structural condition of low Bellman rank (Jiang, Krishnamurthy, Agarwal, Langford, and Schapire, Contextual Decision Processes with Low Bellman Rank are PAC-Learnable, ICML 2017, arXiv:1610.09512; the Q-type variant formalized here follows Du, Kakade, Wang, and Yang, Is a Good Representation Sufficient for Sample Efficient Reinforcement Learning?, ICLR 2020, arXiv:1910.03016). This mission formalizes the second, strictly more general track: BiLinUCB and its Bellman-rank guarantee; LSVI-UCB is left out of scope (see Formalization scope).

Setting

Both algorithms act on the finite-horizon episodic MDP M=(S,A,{Ph}h=1H,{Rh}h=1H,d1)M = (S,A,\{P_h\}_{h=1}^H, \{R_h\}_{h=1}^H,d_1)M=(S,A,{Ph​}h=1H​,{Rh​}h=1H​,d1​) of the earlier chapters of this series (RLBasics.Core), with value fM(π)=Es1∼d1[V1M,π(s1)]f^M(\pi) = \mathbb E_{s_1\sim d_1}[V_1^{M,\pi}(s_1)]fM(π)=Es1​∼d1​​[V1M,π​(s1​)]. For a state-action value function Q=(Qh)h=1HQ = (Q_h)_{h=1}^HQ=(Qh​)h=1H​ (with the convention QH+1≡0Q_{H+1}\equiv 0QH+1​≡0), the Bellman residual of QQQ under policy π\piπ at layer hhh is

Eh(π,Q):=EM,π[Qh(sh,ah)−Rh(sh,ah)−max⁡a′Qh+1(sh+1,a′)],E_h(\pi,Q) := \mathbb E^{M,\pi}\Big[Q_h(s_h,a_h) - R_h(s_h,a_h) - \max_{a'} Q_{h+1}(s_{h+1},a')\Big],Eh​(π,Q):=EM,π[Qh​(sh​,ah​)−Rh​(sh​,ah​)−a′max​Qh+1​(sh+1​,a′)],

which vanishes identically when Q=QM,⋆Q = Q^{M,\star}Q=QM,⋆, the optimal value function (Bellman optimality). Given a class Q\mathcal QQ of candidate value functions, MMM has Bellman rank ddd relative to Q\mathcal QQ (Definition 8) if ddd is the least integer such that, at every layer hhh, there exist embeddings Xh(π),Wh(Q)∈RdX_h(\pi), W_h(Q) \in \mathbb R^dXh​(π),Wh​(Q)∈Rd with Eh(π,Q)=⟨Xh(π),Wh(Q)⟩E_h(\pi,Q) = \langle X_h(\pi), W_h(Q)\rangleEh​(π,Q)=⟨Xh​(π),Wh​(Q)⟩ for every policy π\piπ and Q∈QQ\in\mathcal QQ∈Q — equivalently, the least rank of the Π×Q\Pi\times\mathcal QΠ×Q matrix of Bellman residuals, at any layer. A linear MDP (the setting of LSVI-UCB, §7.2) is the special case where the transition kernel and reward themselves factor through a known feature map ϕ:S×A→Rd\phi : S\times A \to \mathbb R^dϕ:S×A→Rd; every linear MDP has Bellman rank at most ddd relative to the linear value-function class, but Bellman rank captures far more (kernel/neural function classes, and low-rank MDPs whose feature map is unknown).

The chapter presents two algorithms. LSVI-UCB (§7.2.1) runs TTT episodes of ridge regression per layer against the linear feature map, forming a confidence ellipsoid of radius ρ∝d\rho \propto \sqrt dρ∝d​ around each layer's estimated parameter and acting greedily with respect to an upper-confidence bonus built from that ellipsoid — the same optimism-under-uncertainty template as UCB-VI, now regularized rather than tabular; it motivates Bellman rank but is not itself formalized by this mission (see Formalization scope). BiLinUCB (§7.3.1), the algorithm this mission formalizes, instead proceeds in KKK iterations of nnn episodes: each iteration plays the greedy policy for the current optimistic-on-average value function Qk=arg⁡max⁡Q∈QkEs1∼d1[Q1(s1,πQ(s1))]Q_k = \arg\max_{Q\in\mathcal Q_k}\mathbb E_{s_1\sim d_1}[Q_1(s_1, \pi_Q(s_1))]Qk​=argmaxQ∈Qk​​Es1​∼d1​​[Q1​(s1​,πQ​(s1​))], collects nnn fresh episodes, and shrinks the confidence set Qk+1\mathcal Q_{k+1}Qk+1​ by discarding value functions whose empirical Bellman residual along the played policy is large; after KKK iterations it outputs the policy with the best empirical return observed at any iteration.

Formalization targets

Goal — Proposition 47 (BiLinUCB, sample complexity under Bellman rank)

∃ c1,c2,c3>0, ∀ ε,δ>0,  n≳H3dlog⁡(∣Q∣/δ)ε2,  K≳Hdlog⁡(1+n/d),  β∝Klog⁡∣Q∣+log⁡(HK/δ)n ⟹\exists\, c_1,c_2,c_3>0,\ \forall\,\varepsilon,\delta>0,\ \ n\gtrsim \frac{H^3d\log(|\mathcal Q|/\delta)}{\varepsilon^2},\ \ K\gtrsim Hd\log(1+n/d),\ \ \beta\propto\frac{K\log|\mathcal Q|+\log(HK/\delta)}n\ \Longrightarrow∃c1​,c2​,c3​>0, ∀ε,δ>0,  n≳ε2H3dlog(∣Q∣/δ)​,  K≳Hdlog(1+n/d),  β∝nKlog∣Q∣+log(HK/δ)​ ⟹ Pr⁡[fM⋆(πM⋆)−fM⋆(π^)≤ε]≥1−δ,\Pr\big[f^{M^\star}(\pi^{M^\star}) - f^{M^\star}(\hat\pi) \le \varepsilon\big] \ge 1-\delta,Pr[fM⋆(πM⋆)−fM⋆(π^)≤ε]≥1−δ,

for M⋆M^\starM⋆ of Bellman rank ddd relative to Q∋QM⋆,⋆\mathcal Q\ni Q^{M^\star,\star}Q∋QM⋆,⋆, where π^\hat\piπ^ is BiLinUCB's output policy after KKK iterations of nnn episodes. This is the weakest stable form: it fixes the shape of the sample complexity (polynomial in H,d,log⁡∣Q∣,1/εH,d,\log|\mathcal Q|,1/\varepsilonH,d,log∣Q∣,1/ε, logarithmic in 1/δ1/\delta1/δ, independent of ∣S∣|S|∣S∣) and leaves the leading constants — which the book itself introduces only as "a sufficiently large numerical constant" — outside the formal claim. Unlike every other goal in this series, this is a PAC (sample-complexity) guarantee on the algorithm's final output policy, not a bound on cumulative regret accrued while learning: BiLinUCB commits to π^\hat\piπ^ only after the training phase ends, and its suboptimality is measured post-training.

Reaching it rests on two structural facts about BiLinUCB's confidence sets, each formalized as a milestone in attack order:

  • Lemma 29 (confidence-set validity): with the stated threshold β\betaβ, with probability at least 1−δ1-\delta1−δ, simultaneously at every iteration kkk, every retained value function has true (population) Bellman residual along the played policies bounded by β\betaβ up to a constant, and the realizable QM⋆,⋆Q^{M^\star,\star}QM⋆,⋆ is itself always retained.
  • Lemma 30 (optimism and elliptic-norm bound): conditioned on Lemma 29's event, every retained value function's embedding Wh(Q)W_h(Q)Wh​(Q) has bounded norm with respect to the Gram matrix of the played policies' embeddings, and the optimistic value function QkQ_kQk​ BiLinUCB selects at each iteration has initial-state value at least fM⋆(πM⋆)f^{M^\star}(\pi^{M^\star})fM⋆(πM⋆).

Significance

Proposition 47 shows that a single structural parameter — Bellman rank — is sufficient for sample-efficient RL with function approximation, with a sample complexity that depends only on the horizon, the rank, and the value-function class's log-cardinality, never on ∣S∣|S|∣S∣. This subsumes the linear-MDP guarantee of Proposition 46 (LSVI-UCB, §7.2.1; not itself a target of this mission, see Formalization scope) as a special case — every linear MDP has Bellman rank ≤d\le d≤d — while covering strictly more models (kernelized and neural value-function classes with a low-dimensional Bellman-residual factorization that need not come from a known linear feature map). The result is proved in the source and this mission formalizes its statement and the two structural lemmas its proof rests on, as stated; no new mathematics is contributed. Formalizing it commits, for the first time on this platform, to machine-checkable statements of the Bellman rank abstraction, the elliptic-norm confidence-set machinery it drives, and a PAC- (rather than regret-) style learning guarantee, none of which appear in the platform's existing bandit or tabular-RL missions.

Difficulty

The obvious first attempt is to formalize Bellman rank as an unconstrained integer parameter ddd attached to the MDP, sidestepping the actual rank condition on the Bellman-residual matrix; this is a trivializing formalization; Bellman rank must be the least dimension admitting the stated bilinear factorization; see Formalization scope. A second obstacle is that BiLinUCB's optimism is only "on average" with respect to the initial state distribution (initValue), unlike LSVI-UCB's pointwise optimism over every state and action — conflating the two confidence-set constructions collapses the chapter's main conceptual contrast. Finally, the two technical lemmas (29 and 30) separate a purely probabilistic statement (validity of the empirical confidence set, via Hoeffding and a union bound) from a purely deterministic consequence (the elliptic-norm bound and optimism, which hold on any sample path where the probabilistic event occurred); keeping this separation is what lets Proposition 47's proof combine them cleanly, and collapsing it into one monolithic high-probability statement would misrepresent the book's proof structure.

Formalization scope

The MDP, policy, and trajectory/history machinery (EpisodicMDP, Policy, IsPolicy, Trajectory, Learner, probEvent) are reused unchanged from the series' published RLBasics.Core/RLBasics.UCBVI definitions. The value-function class Q\mathcal QQ is realized as an abstract finite, nonempty type Qc together with an evaluation map qeval : Qc → ℕ → S → A → ℝ, kept fully abstract rather than specialized to any concrete function class — specializing it to, e.g., linear functions would collapse Proposition 47 back into a restatement of Proposition 46, which is exactly the trivializing formalization this mission avoids. Bellman rank (IsBellmanRank) is defined as the least natural number admitting the bilinear factorization (an IsLeast over the coercion to HasBellmanRankLE), never as a free parameter. Every realized per-step reward is taken equal to its conditional mean Rh(s,a)R_h(s,a)Rh​(s,a) throughout (exact, by the tower property, and not a restriction to deterministic rewards). The elliptic norm of Lemma 30 is expressed via the direct sum-of-squared-inner-products identity ∥v∥Σ2=∑x∈xs⟨x,v⟩2\|v\|^2_\Sigma = \sum_{x\in xs}\langle x,v\rangle^2∥v∥Σ2​=∑x∈xs​⟨x,v⟩2 rather than introducing Matrix/matrix-inverse machinery, since no milestone here needs an explicit matrix inverse. Every place the book writes ≲/∝/"a sufficiently large numerical constant" is formalized as an existentially quantified universal constant, fixed ahead of every MDP, value-function class, and (ε,δ)(\varepsilon,\delta)(ε,δ) instance — never depending on the instance itself. This mission scopes entirely to the Bellman-rank track (Definition 8, BiLinUCB, Lemmas 29-30, Proposition 47); Proposition 46 (LSVI-UCB) and its supporting Lemmas 27-28 are the chapter's motivating linear special case (Section 7.2) but are left out of this mission's scope for time and are not claimed as proved by it — LSVI-UCB's confidence-ellipsoid construction is materially different from BiLinUCB's empirical-Bellman-residual confidence set (see Difficulty) and would need its own milestone chain. A complete development needs no infrastructure beyond what is already published in RLBasics.Core/RLBasics.UCBVI; contributions completing the sorrys in the two technical lemmas and the goal are welcome.

Selected references

  • Foster, D. J. and Rakhlin, A. Foundations of Reinforcement Learning and Interactive Decision Making. arXiv:2312.16730v1, 2023. arXiv:2312.16730
  • Jin, C., Yang, Z., Wang, Z., and Jordan, M. I. Provably Efficient Reinforcement Learning with Linear Function Approximation. COLT 2020. arXiv:1907.05388
  • Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E. Contextual Decision Processes with Low Bellman Rank are PAC-Learnable. ICML 2017. arXiv:1610.09512
  • Du, S. S., Kakade, S. M., Wang, R., and Yang, L. F. Is a Good Representation Sufficient for Sample Efficient Reinforcement Learning? ICLR 2020. arXiv:1910.03016
7 thms3 active usersReviewed
🏆Completed
Convex OptimizationProbabilityRandom Matrix Theory+1·Captain: mikedeng1

High-Dimensional Probability III: Grothendieck's InequalityTextbook

Motivation

Many hard combinatorial optimization problems — finding the maximum cut of a graph, deciding the ground state of an Ising spin system, bounding the correlation of a physical system — can be written as maximizing a bilinear form over sign vectors xi∈{−1,1}x_i \in \{-1, 1\}xi​∈{−1,1}. Exhaustive search over 2n2^n2n sign patterns is intractable, so practitioners relax the problem: replace each sign xix_ixi​ by a unit vector XiX_iXi​ in a higher-dimensional space and optimize the resulting inner products instead. This relaxation, a semidefinite program, is convex and solvable in polynomial time. The question is how much is lost in the relaxation — whether its optimal value can be far from the true, combinatorial optimum.

Grothendieck's inequality, proved by Alexander Grothendieck in 1953 in the context of Banach space theory (Résumé de la théorie métrique des produits tensoriels topologiques, Bol. Soc. Mat. São Paulo 8 (1953), 1–79), answers this for a broad class of such relaxations: replacing signs by unit vectors in an arbitrary Hilbert space changes the optimal value by at most an absolute, dimension-free constant factor. The inequality has since become a standard tool across combinatorial optimization, Banach space geometry, and (via the Goemans-Williamson algorithm for maximum cut, Section 3.6 of the source) approximation algorithms; see U. Haagerup, The Grothendieck inequality for bilinear forms on C∗C^*C∗-algebras, Adv. Math. 56 (1985) for the tightest known constant, and Alon–Naor, Approximating the cut-norm via Grothendieck's inequality, SIAM J. Comput. 35 (2006), for the algorithmic connection this mission's Theorem 3.5.6 sets up.

Setting

Fix positive integers m,nm, nm,n. Consider a real m×nm \times nm×n matrix A=(aij)A = (a_{ij})A=(aij​). Say AAA is normalized if for every choice of numbers x1,…,xm,y1,…,yn∈{−1,1}x_1, \dots, x_m, y_1, \dots, y_n \in \{-1, 1\}x1​,…,xm​,y1​,…,yn​∈{−1,1},

∣∑i=1m∑j=1naij xiyj∣  ≤  1.\Bigl| \sum_{i=1}^m \sum_{j=1}^n a_{ij}\, x_i y_j \Bigr| \;\le\; 1.​i=1∑m​j=1∑n​aij​xi​yj​​≤1.

This says AAA, viewed as a bilinear form on {−1,1}m×{−1,1}n\{-1,1\}^m \times \{-1,1\}^n{−1,1}m×{−1,1}n, has sup-norm at most 111. Now let HHH be any real Hilbert space — a real vector space equipped with an inner product ⟨⋅,⋅⟩\langle \cdot, \cdot \rangle⟨⋅,⋅⟩ complete in the induced norm — and consider vectors u1,…,um∈Hu_1, \dots, u_m \in Hu1​,…,um​∈H and v1,…,vn∈Hv_1, \dots, v_n \in Hv1​,…,vn​∈H, each of unit norm ∥ui∥=∥vj∥=1\|u_i\| = \|v_j\| = 1∥ui​∥=∥vj​∥=1. Replacing the scalar product xiyjx_i y_jxi​yj​ by the inner product ⟨ui,vj⟩\langle u_i, v_j \rangle⟨ui​,vj​⟩ in the same bilinear form gives ∑i,jaij⟨ui,vj⟩\sum_{i,j} a_{ij} \langle u_i, v_j \rangle∑i,j​aij​⟨ui​,vj​⟩, a real number depending on the choice of HHH and of the unit vectors. The question is how large this can be, uniformly over every such choice.

Formalization targets

Grothendieck's inequality (Theorem 3.5.1)

A normalized  ⟹  ∣∑i,jaij ⟨ui,vj⟩∣  ≤  KA \text{ normalized} \;\Longrightarrow\; \Bigl| \sum_{i,j} a_{ij}\, \langle u_i, v_j\rangle \Bigr| \;\le\; KA normalized⟹​i,j∑​aij​⟨ui​,vj​⟩​≤K

for every real Hilbert space HHH and unit vectors ui,vj∈Hu_i, v_j \in Hui​,vj​∈H, where KKK is a constant depending on neither AAA, its dimensions, nor HHH. This mission's goal formalizes the book's own first-pass bound K≤288K \le 288K≤288 (Section 3.5), proved by a Gaussian truncation argument; it does not fix a numeral for KKK, only that some absolute constant works, matching the shape of the true statement rather than a specific numeral that a sharper argument (the book's own Section 3.7 gives K≤1.783K \le 1.783K≤1.783) would immediately obsolete. See Formalization scope below for why this is the goal, not the sharper bound.

Significance

The result itself. Grothendieck's inequality is the single fact that makes semidefinite relaxation a provably good algorithmic strategy rather than a heuristic: whatever the true, hard-to-compute combinatorial optimum of a {−1,1}\{-1,1\}{−1,1}-valued bilinear optimization is, the tractable Hilbert-space relaxation cannot overshoot it by more than the constant KKK. Milestone Theorem 3.5.6 makes this concrete for positive-semidefinite matrices, showing the semidefinite relaxation SDP(A)(A)(A) of the integer program INT(A)(A)(A) satisfies INT(A)≤(A) \le(A)≤ SDP(A)≤2K⋅(A) \le 2K \cdot(A)≤2K⋅ INT(A)(A)(A) — the guarantee underlying the Goemans-Williamson 0.878-approximation algorithm for maximum cut (Theorem 3.6.5 of the source, out of scope for this mission; see Formalization scope).

Formalizing it. The inequality and its two chapter milestones are proved but not previously formalized on this platform (checked by concept search for "Grothendieck", "semidefinite", "positive-semidefinite", and "max-cut" — no hits beyond the unrelated Grothendieck-Teichmüller group). What remains after this mission is the sharper K≤1.783K \le 1.783K≤1.783 argument of Section 3.7 (the "kernel trick"), a separate, heavier development building on positive-definite kernels, and full proofs of every milestone below (currently open sorry goals).

Difficulty

The statement of Grothendieck's inequality contains no randomness, yet every known elementary proof is probabilistic; this is itself a striking feature of the result. The obvious approach — bound ∑i,jaij⟨ui,vj⟩\sum_{i,j} a_{ij}\langle u_i,v_j\rangle∑i,j​aij​⟨ui​,vj​⟩ directly by exploiting the normalization hypothesis on AAA — fails because the normalization hypothesis only controls AAA against sign vectors, and there is no way to project an arbitrary unit vector in a Hilbert space onto {−1,1}\{-1,1\}{−1,1} without losing information. The book's proof instead represents each unit vector ui,vju_i, v_jui​,vj​ via a scalar Gaussian random variable ⟨g,ui⟩\langle g, u_i\rangle⟨g,ui​⟩ for a single Gaussian vector ggg, recovering the inner products in expectation (Exercise 3.3.5); but these Gaussian variables are unbounded, so the normalization hypothesis (which bounds AAA against bounded ±1\pm 1±1 inputs) cannot be applied to them directly. The core technical step is a truncation argument: splitting each Gaussian variable into a bounded part and a small-L2L^2L2-norm unbounded remainder, applying the hypothesis to the bounded parts, and bounding the remainder terms by treating them as elements of the Hilbert space L2L^2L2 and invoking the very inequality being proved (Theorem 3.5.1 itself, applied with H=L2H = L^2H=L2) as a self-referential bootstrap — this is why the proof fixes KKK as the smallest valid constant before starting, rather than building it up from scratch.

Formalization scope

The goal and both milestones work with the real matrix and real inner product space directly; H is required to be a complete real inner product space (NormedAddCommGroup, InnerProductSpace ℝ, CompleteSpace), matching the book's "any Hilbert space." No dimension bound on HHH is imposed — the inequality's content is exactly that KKK does not grow with dim⁡H\dim HdimH.

This mission does not formalize the sharper K≤1.783K \le 1.783K≤1.783 bound of Section 3.7, nor Theorem 3.6.5 (the 0.878-approximation guarantee for maximum cut via randomized rounding): the latter's statement quantifies over "the result of a randomized rounding of the solution of the semidefinite program," which would drag a specific algorithm into the audited statement rather than keeping it a self-contained mathematical claim (the statement/proof-separation trap this series' triage rubric flags). Grothendieck's identity (Lemma 3.6.6), the key fact behind that rounding step, is included on its own as a milestone, stated with an explicit, named random sign variable rather than an opaque "rounding procedure."

A trivializing formalization would state the goal with KKK allowed to depend on AAA, mmm, nnn, or HHH — every such bound is easy (e.g. K=∑ij∣aij∣K = \sum_{ij} |a_{ij}|K=∑ij​∣aij​∣) and carries none of the theorem's content; the Lean statement rules this out by quantifying KKK before every other object. INT(A)\mathrm{INT}(A)INT(A) and SDP(A)\mathrm{SDP}(A)SDP(A) (Theorem 3.5.6) are defined from scratch in this chunk's namespace, using Matrix.PosSemidef from Mathlib for the positive-semidefiniteness hypothesis (which bundles the real-symmetric condition); Mathlib has no ready-made SDP-value construction to reuse. The sub-gaussian (Orlicz ψ2\psi_2ψ2​) norm used by Theorem 3.1.1 is reused, unchanged, from the 01-concentration mission in this series (HighDimProb.Concentration.SubgaussianNorm) rather than redefined.

Selected references

  • A. Grothendieck, Résumé de la théorie métrique des produits tensoriels topologiques, Bol. Soc. Mat. São Paulo 8 (1953), 1–79.
  • U. Haagerup, The Grothendieck inequality for bilinear forms on C∗C^*C∗-algebras, Adv. Math. 56 (1985), 93–116.
  • N. Alon, A. Naor, Approximating the cut-norm via Grothendieck's inequality, SIAM J. Comput. 35 (2006), 787–803.
  • M. X. Goemans, D. P. Williamson, Improved approximation algorithms for maximum cut and satisfiability problems using semidefinite programming, J. ACM 42 (1995), 1115–1145.
  • R. Vershynin, High-Dimensional Probability: An Introduction with Applications in Data Science, Cambridge University Press, 2018, Chapter 3, DOI 10.1017/9781108231596.
8 thms3 active usersReviewed
🏆Completed
ProbabilityRandom Matrix TheoryStatistics·Captain: mikedeng1

High-Dimensional Probability V: The Johnson-Lindenstrauss LemmaTextbook

Motivation

Any dataset of NNN points can be described exactly by embedding it in Rn\mathbb R^nRn for nnn large enough — but a large nnn is expensive: nearest-neighbor search, clustering, and streaming algorithms all scale with the ambient dimension, not with NNN. The question that opens this mission is whether the dimension can be cut down while leaving the data's geometry — the pairwise distances between points — essentially untouched.

Johnson and Lindenstrauss answered this in 1984, while studying extensions of Lipschitz maps into Hilbert space (W. Johnson, J. Lindenstrauss, Extensions of Lipschitz mappings into a Hilbert space, Contemp. Math. 26 (1984), 189–206): NNN points in any Euclidean space, of any dimension nnn, can be mapped by a single linear map into a space of dimension only O(ε−2log⁡N)O(\varepsilon^{-2}\log N)O(ε−2logN), distorting every pairwise distance by at most a factor of 1±ε1\pm\varepsilon1±ε. The map does not depend on the data beyond its cardinality — a single random object works simultaneously for the whole point set with high probability. This is now one of the standard tools of randomized dimension reduction, cited across nearest-neighbor search, streaming linear algebra, compressed sensing, and machine learning pipelines that need to shrink feature dimension before a downstream algorithm runs.

Setting

Fix a probability space (Ω,F,Prob)(\Omega,\mathcal F,\mathrm{Prob})(Ω,F,Prob). A random orthogonal projection of rank mmm in Rn\mathbb R^nRn is a map P:Ω→(Rn→Rn)P:\Omega\to(\mathbb R^n\to\mathbb R^n)P:Ω→(Rn→Rn), continuous and linear for each ω\omegaω, such that almost surely PωP_\omegaPω​ is idempotent (Pω∘Pω=PωP_\omega\circ P_\omega = P_\omegaPω​∘Pω​=Pω​), self-adjoint, and has range of dimension mmm — i.e. PωP_\omegaPω​ is the orthogonal projection onto some mmm-dimensional subspace Eω⊂RnE_\omega\subset\mathbb R^nEω​⊂Rn. It is uniformly distributed in the Grassmannian Gn,mG_{n,m}Gn,m​ (written E∼Unif(Gn,m)E\sim\mathrm{Unif}(G_{n,m})E∼Unif(Gn,m​)) when its law is rotation invariant: for every orthogonal transformation UUU of Rn\mathbb R^nRn, the conjugated map ω↦U∘Pω∘U−1\omega\mapsto U\circ P_\omega\circ U^{-1}ω↦U∘Pω​∘U−1 has the same law as PPP. Conjugating a projection by UUU is exactly the projection onto the image of its range under UUU, so this says the law of the random subspace E=range(P)E=\mathrm{range}(P)E=range(P) is invariant under the full orthogonal group — the operational definition Vershynin himself uses for a "uniformly distributed" random subspace, since no coordinate-free formula for such a subspace's law is given directly.

A companion notion drives the proof: a random vector XXX is uniform on the Euclidean sphere of radius rrr, X∼Unif(r Sn−1)X\sim\mathrm{Unif}(r\,S^{n-1})X∼Unif(rSn−1), when it lies on that sphere almost surely and its law is likewise rotation invariant. And a real random variable YYY is sub-gaussian with sub-gaussian (ψ2\psi_2ψ2​) norm ∥Y∥ψ2:=inf⁡{t>0:Eexp⁡(Y2/t2)≤2}\|Y\|_{\psi_2} := \inf\{t>0:\mathbb E\exp(Y^2/t^2)\le 2\}∥Y∥ψ2​​:=inf{t>0:Eexp(Y2/t2)≤2}, the standard non-asymptotic measure of how light-tailed YYY's distribution is (a bounded or Gaussian random variable has finite ψ2\psi_2ψ2​ norm; the tail probability P{∣Y∣≥s}\mathbb P\{|Y|\ge s\}P{∣Y∣≥s} then decays at least as fast as 2exp⁡(−cs2/∥Y∥ψ22)2\exp(-cs^2/\|Y\|_{\psi_2}^2)2exp(−cs2/∥Y∥ψ2​2​)).

Formalization targets

Goal (Theorem 5.3.1, Johnson-Lindenstrauss Lemma)

∃ C,c>0:m≥Cε2log⁡∣X∣  ⟹  Prob{∀x,y∈X: (1−ε)∥x−y∥2≤∥nm Pω(x−y)∥2≤(1+ε)∥x−y∥2}  ≥  1−2exp⁡(−cε2m)\exists\,C,c>0:\quad m\ge\frac{C}{\varepsilon^2}\log|X| \;\Longrightarrow\; \mathrm{Prob}\Bigl\{\forall x,y\in X:\ (1-\varepsilon)\|x-y\|_2\le \bigl\|\sqrt{\tfrac nm}\,P_\omega(x-y)\bigr\|_2\le(1+\varepsilon)\|x-y\|_2\Bigr\} \;\ge\;1-2\exp(-c\varepsilon^2 m)∃C,c>0:m≥ε2C​log∣X∣⟹Prob{∀x,y∈X: (1−ε)∥x−y∥2​≤​mn​​Pω​(x−y)​2​≤(1+ε)∥x−y∥2​}≥1−2exp(−cε2m)

for every finite X⊂RnX\subset\mathbb R^nX⊂Rn, every ε>0\varepsilon>0ε>0, and every random orthogonal projection PPP of rank mmm uniformly distributed in Gn,mG_{n,m}Gn,m​. The universal quantifier over pairs x,y∈Xx,y\in Xx,y∈X sits inside the single probability event — this is the union-bound content that makes the statement a genuine simultaneous guarantee for the whole point set, not a restatement of the single-vector lemma below for one fixed pair. Both constants are the book's own unnamed absolute constants, never depending on nnn, mmm, N=∣X∣N=|X|N=∣X∣, or ε\varepsilonε; this is the weakest stable form of the claim (no numeral is hard-coded for CCC or ccc), matching the book's own statement exactly.

Significance

The lemma gives a universal, data-oblivious dimension-reduction guarantee: the target dimension m=O(ε−2log⁡N)m=O(\varepsilon^{-2}\log N)m=O(ε−2logN) depends only on the number of points and the desired distortion, never on the ambient dimension nnn or on the geometry of the specific point set. This is what makes it usable as a black-box preprocessing step ahead of an algorithm whose cost scales with nnn — the projection is drawn once, without looking at the data, and works with high probability for every pairwise distance simultaneously. The bound is also known to be essentially optimal in NNN: Alon (Problems and results in extremal combinatorics, Discrete Math. 273 (2003)) showed a lower bound of Ω(ε−2log⁡N/log⁡(1/ε))\Omega(\varepsilon^{-2}\log N/\log(1/\varepsilon))Ω(ε−2logN/log(1/ε)) on the target dimension, so the log⁡N\log NlogN dependence cannot be removed.

The theorem itself has been proved for decades and admits several proof strategies (this book's route through Lipschitz concentration on the sphere; the original volume/measure-concentration argument; later "sparse" or structured variants of the projection for faster computation). This mission formalizes the classical dense-Gaussian-projection proof route as Vershynin presents it, building the sphere-concentration engine (Theorem 5.1.4) and the single-vector projection lemma (Lemma 5.3.2) that the union-bound argument for the goal rests on. No machine-checked formal proof of this chain is known to exist on the platform prior to this mission (see Formalization scope below); what is contributed is the statement infrastructure — the goal and its two direct supporting lemmas, stated with explicit, unpinned absolute constants — for solvers to close.

Difficulty

The natural first idea — bound the distortion of a single fixed vector under a random projection, then take a union bound over the (N2)\binom N2(2N​) pairwise differences — is exactly the strategy Lemma 5.3.2 and the goal use, but it does not by itself explain why the single-vector concentration bound (Lemma 5.3.2(b)) holds with the stated sub-gaussian-type tail. That bound is not elementary: it reduces to a uniform concentration statement for an arbitrary Lipschitz function of a uniformly random point on a high-dimensional sphere (Theorem 5.1.4), since ∥Pz∥2\|Pz\|_2∥Pz∥2​, viewed as a function of a rotated copy of zzz, is a 111-Lipschitz function on the sphere. Proving that every Lipschitz function concentrates — not just linear ones, for which sub-gaussianity was already established in Chapter 3 — needs a genuinely different tool: comparing the sub-level sets of an arbitrary Lipschitz function to spherical caps via an isoperimetric inequality on the sphere. This geometric input is what makes the concentration phenomenon behind Johnson-Lindenstrauss a dimension-free fact rather than a special property of coordinate projections.

Formalization scope

XXX is a Finset of points in EuclideanSpace ℝ (Fin n), matching "a set of NNN points"; NNN is read off as X.card. The random subspace E∈Gn,mE\in G_{n,m}E∈Gn,m​ is represented throughout by the orthogonal projection PPP onto it (IsUniformProjection), following the book's own statements, which are phrased in terms of PPP rather than EEE; the scaled map Q=n/m PQ=\sqrt{n/m}\,PQ=n/m​P of the goal is written Real.sqrt (n/m) • P ω applied to x - y, using linearity of PωP_\omegaPω​ to realize Qx−Qy=Q(x−y)Qx-Qy = Q(x-y)Qx−Qy=Q(x−y). Both "uniform on the sphere" and "uniform in the Grassmannian" are defined operationally by rotation invariance of the underlying law, since Mathlib has no ready-made normalized surface measure on a general-radius Euclidean sphere or Haar-measure construction on the Grassmannian/orthogonal group to build a canonical uniform object from; rotation invariance uniquely determines the corresponding measure among those supported on the relevant set, so the operational and constructive definitions coincide extensionally. Every "absolute constant" in the book (CCC in Theorem 5.3.1's sample-complexity hypothesis, ccc in every failure-probability bound, and the sub-gaussian constant CCC of Theorem 5.1.4) is existentially quantified ahead of the dimension, sample size, and every other object, and pinned to no numeral — a formalization that hard-coded a specific numeral for any of these would be invalidated by the next sharper constant in the literature and would not match what the book actually proves.

A trivializing formalization is one that states the conclusion for a single fixed pair x,yx,yx,y rather than universally over all pairs inside one event; that would collapse the union-bound content that makes this a dimension-reduction statement for a whole point set (with NNN points), rather than a restatement of the single-vector Lemma 5.3.2(b). This mission's goal statement is built to rule that out explicitly (see Formalization targets above).

Reusable infrastructure: subgaussianNorm (the Orlicz ψ2\psi_2ψ2​ norm, restated per Vershynin Definition 2.5.6) and the rotation-invariance idiom for "uniformly distributed" random geometric objects are of independent interest to any later chapter needing sub-gaussian random vectors or random subspaces/projections (e.g. Chapters 4, 6, 7, 9, 11 of this same book series). Solvers' contributions are welcome on: the isoperimetric inequality on the sphere and its use to prove Theorem 5.1.4 (the mission's hardest open leaf); the coordinate-projection computation underlying Lemma 5.3.2(a); and the concentration-plus-union-bound argument closing the goal from the three supporting lemmas.

Selected references

  • W. Johnson, J. Lindenstrauss, Extensions of Lipschitz mappings into a Hilbert space, Contemporary Mathematics 26 (1984), 189–206.
  • N. Alon, Problems and results in extremal combinatorics, I, Discrete Mathematics 273 (2003), 31–53. https://doi.org/10.1016/S0012-365X(03)00227-9
  • R. Vershynin, High-Dimensional Probability: An Introduction with Applications in Data Science, Cambridge University Press, 2018, Chapter 5. https://doi.org/10.1017/9781108231596
7 thms3 active usersReviewed
🏆Completed
Algorithmic Game TheoryConvex OptimizationLinear Optimization+1·Captain: mikedeng1

Introduction to Online Convex Optimization VIII: Solving Zero-Sum Games and Linear Programs via Regret MinimizationTextbook

Motivation

Two-player zero-sum games and linear programming are, on their surface, unrelated pieces of 20th-century mathematics: von Neumann's minimax theorem for games (1928) was proved with tools from topology, and linear programming duality (Dantzig, 1940s) with convexity and geometry. Yet the two are formally equivalent — Dantzig recounts von Neumann conjecturing the equivalence outright, on first hearing a description of linear programming, because he had "just recently completed a book with Oscar Morgenstern on the theory of games" [Albers, Alexanderson, and Reid, More Mathematical People, 1990]. Freund and Schapire (1999) later showed that both concepts reduce, in one uniform way, to online regret minimization: a decades-old topological existence proof and a decades-old LP-duality argument both become corollaries of a single fact about no-regret learning. This mission formalizes the algorithmic content of that reduction — Hazan's Lemma 8.4, which is not merely an existence statement but a concrete, efficient algorithm with an explicit convergence rate.

Setting

A two-player zero-sum game in normal form is a real matrix A∈Rn×mA \in \mathbb{R}^{n \times m}A∈Rn×m (Hazan restricts entries to [−1,1][-1,1][−1,1] for interpretability as losses/rewards, a convention this mission's theorems drop as inessential — the argument is invariant to scaling and shifting). The row player picks a mixed strategy xxx in the probability simplex Δn={x∈Rn:xi≥0,∑ixi=1}\Delta_n = \{x \in \mathbb{R}^n : x_i \ge 0, \sum_i x_i = 1\}Δn​={x∈Rn:xi​≥0,∑i​xi​=1}; the column player picks y∈Δmy \in \Delta_my∈Δm​. The row player's expected loss, and simultaneously the column player's expected reward, is the bilinear form xTAyx^{\mathsf T} A yxTAy.

The row player's guaranteed loss is λR=min⁡x∈Δnmax⁡y∈ΔmxTAy\lambda_R = \min_{x \in \Delta_n} \max_{y \in \Delta_m} x^{\mathsf T} A yλR​=minx∈Δn​​maxy∈Δm​​xTAy: the smallest loss she can secure no matter what the column player does. Symmetrically, the column player's guaranteed reward is λC=max⁡y∈Δmmin⁡x∈ΔnxTAy\lambda_C = \max_{y \in \Delta_m} \min_{x \in \Delta_n} x^{\mathsf T} A yλC​=maxy∈Δm​​minx∈Δn​​xTAy. Always λR≥λC\lambda_R \ge \lambda_CλR​≥λC​ ("weak duality" — an elementary max-min/min-max inequality, Direction 1 of Section 8.3). Von Neumann's minimax theorem (Theorem 8.3) is the nontrivial converse: λR=λC\lambda_R = \lambda_CλR​=λC​, a common value λ⋆\lambda^\starλ⋆ called the value of the game, whose optimal strategies form a Nash equilibrium — already on the platform as AGT.zero_sum_minimax.

Algorithm 28 ("Simple LP", p. 147) computes an approximate equilibrium constructively. The row player runs a multiplicative-weights / Exponentiated Gradient update against the sequence of best-response losses the column player generates in a repeated TTT-round play of the game: starting from the uniform strategy x1=(1/n,…,1/n)x_1 = (1/n, \dots, 1/n)x1​=(1/n,…,1/n), at each round ttt the column player best-responds with yt∈arg⁡max⁡y∈ΔmxtTAyy_t \in \arg\max_{y \in \Delta_m} x_t^{\mathsf T} A yyt​∈argmaxy∈Δm​​xtT​Ay, and the row player updates xt+1(i)∝xt(i) e−η(Ayt)ix_{t+1}(i) \propto x_t(i)\, e^{-\eta (A y_t)_i}xt+1​(i)∝xt​(i)e−η(Ayt​)i​. The algorithm returns the time-averaged strategy xˉ=1T∑t=1Txt\bar{x} = \frac{1}{T}\sum_{t=1}^T x_txˉ=T1​∑t=1T​xt​.

Formalization targets

Goal — Lemma 8.4

max⁡y′∈ΔmxˉTAy′  ≤  λR(A)+2log⁡nT\max_{y' \in \Delta_m} \bar{x}^{\mathsf T} A y' \;\le\; \lambda_R(A) + \frac{\sqrt{2 \log n}}{\sqrt{T}}y′∈Δm​max​xˉTAy′≤λR​(A)+T​2logn​​

for the vector xˉ\bar{x}xˉ returned by Algorithm 28 after TTT rounds with learning rate η=2log⁡n/T\eta = \sqrt{2 \log n / T}η=2logn/T​. The book calls xˉ\bar{x}xˉ a "2log⁡n/T\sqrt{2 \log n}/\sqrt{T}2logn​/T​-approximate solution" to the zero-sum game — and, via Section 8.2.1's equivalence, to the linear program the game encodes — in exactly this sense. The goal is stated against λR\lambda_RλR​, the quantity the algorithm's own analysis produces; Theorem 8.3 identifies it with λC\lambda_CλC​ and with the book's λ⋆\lambda^\starλ⋆, so nothing about the bound is lost by this choice of rendering.

Supporting milestone — Eq. (8.1)

∑t=0T−1xtTAyt  ≤  min⁡x′∈Δn∑t=0T−1(x′)TAyt  +  2Tlog⁡n\sum_{t=0}^{T-1} x_t^{\mathsf T} A y_t \;\le\; \min_{x' \in \Delta_n} \sum_{t=0}^{T-1} (x')^{\mathsf T} A y_t \;+\; \sqrt{2T \log n}t=0∑T−1​xtT​Ayt​≤x′∈Δn​min​t=0∑T−1​(x′)TAyt​+2Tlogn​

the external-regret bound the row player's multiplicative-weights update achieves against the adaptively-chosen linear loss sequence ft(⋅)=(⋅)TAytf_t(\cdot) = (\cdot)^{\mathsf T} A y_tft​(⋅)=(⋅)TAyt​ — the single analytical fact the goal's proof needs.

Significance

The result itself. Lemma 8.4 gives a genuinely efficient algorithm: O(log⁡n/ε2)O(\log n / \varepsilon^2)O(logn/ε2) rounds of a trivial multiplicative update to reach an ε\varepsilonε-approximate value and equilibrium of an n×mn \times mn×m zero-sum game, and — through the equivalence with LP duality — an approximation algorithm for a broad class of linear programs, predating and prefiguring the multiplicative-weights-based approximation schemes surveyed by Arora, Hazan, and Kale (2012). It is also the constructive engine behind Theorem 8.3: unlike the classical topological proof of the minimax theorem, this one produces the equilibrium, not just its existence.

Formalizing it. The equilibrium-existence half of this story, Theorem 8.3, is already a published, proved-format Prove2Me theorem (AGT.zero_sum_minimax, from the Algorithmic Game Theory series) and is reused here as a reference item rather than redrafted. What that theorem does not capture — and what makes this mission non-trivial rather than a restatement — is the quantitative, algorithmic content: that one specific, simple, Hedge-type update, run for a specific number of rounds, provably gets within a specific, explicit distance of the value, using only the existence of some sublinear-regret online algorithm as a black box.

Difficulty

The tempting shortcut is to formalize only "no-regret learning dynamics converge to an equilibrium" as a qualitative statement, discharging it by citing AGT.zero_sum_minimax (equilibria exist) plus a generic regret bound. That collapses Lemma 8.4 into a restatement of Theorem 8.3 and drops exactly what is new here: the explicit rate 2log⁡n/T\sqrt{2\log n}/\sqrt{T}2logn​/T​, tied to one concrete update rule (Algorithm 28) rather than an arbitrary sublinear-regret black box. The real content is in chaining three quantitative facts — Eq. (8.1)'s specific regret bound for the multiplicative-weights update, the column player's best-response equality (Eq. (8.2)), and the definitional unfolding of λR\lambda_RλR​ — with none of the slack that a purely qualitative "an algorithm with sublinear regret exists" argument would tolerate.

Formalization scope

Matrices are Matrix (Fin n) (Fin m) ℝ with n, m ≥ 1 (empty strategy sets are excluded throughout, matching this mission's reference item AGT.zero_sum_minimax); mixed strategies use Mathlib's stdSimplex ℝ (Fin n). lambdaR/lambdaC are rendered with iInf/iSup over simplex membership, the same convention Introduction to Online Convex Optimization III fixed for RegretT earlier in this series. Algorithm 28's run is packaged as a Prop-valued structure (IsSimpleLPRun) rather than a computable function, in the style of this series' other algorithm-run definitions (IsHedgeRun, IsOnlineGradientDescent): initial uniform strategy, a best-response condition on the column player at every round, and the multiplicative-weights recursion on the row player, with the learning rate η left free and fixed to √(2 log n / T) only at the point the theorems need the book's specific constant.

The trivializing risk here is stating only that some sublinear-regret algorithm secures the bound (already implied, vacuously, by AGT.zero_sum_minimax plus any regret bound); this mission rules that out by fixing the exact update rule of Algorithm 28 in IsSimpleLPRun and proving the bound for that rule specifically, with the book's exact constant √(2 log n)/√T, not an unspecified O(·).

Chapter 5's Corollary 5.7 (the general RFTL/Exponentiated-Gradient regret bound) belongs to a different mission of this series and is not imported; eg_regret_bound restates, locally and self-containedly, exactly the instance of it this chapter's proof needs. A later mission for Chapter 5, once published, could supersede this local restatement by specializing its general bound — a natural contribution for a solver with that mission's Lean available.

Selected references

  • J. von Neumann, "Zur Theorie der Gesellschaftsspiele", Mathematische Annalen, 1928.
  • Y. Freund and R. E. Schapire, "Adaptive Game Playing Using Multiplicative Weights", Games and Economic Behavior, 1999. https://doi.org/10.1006/game.1999.0738
  • E. Hazan, Introduction to Online Convex Optimization, 2nd ed., 2022. arXiv:1909.05207v3
  • N. Nisan, T. Roughgarden, E. Tardos, and V. V. Vazirani (eds.), Algorithmic Game Theory, Cambridge University Press, 2007. https://doi.org/10.1017/CBO9780511800481
  • S. Arora, E. Hazan, and S. Kale, "The Multiplicative Weights Update Method: a Meta-Algorithm and Applications", Theory of Computing, 2012. https://doi.org/10.4086/toc.2012.v008a006
5 thms3 active usersReviewed
🏆Completed
Dynamic ProgrammingReinforcement Learning·Captain: mikedeng1

Foundations of Reinforcement Learning IV: Reinforcement Learning Basics and the UCB-VI AlgorithmTextbook

Motivation

Reinforcement learning (RL) formalizes sequential decision-making under uncertainty: an agent repeatedly observes a state, takes an action, receives a reward, and transitions to a new state, with the goal of maximizing cumulative reward over an unknown environment. It underlies applications from game-playing agents to robotics and adaptive medical treatment. What separates RL from the bandit and contextual bandit problems of earlier chapters in this series is state: the environment carries information forward across time steps, so a good action now can pay off many steps later, and a poor exploration strategy can take exponentially long to discover it. This mission formalizes the foundational results of finite-horizon episodic RL — the Markov Decision Process (MDP) model, the Bellman-optimality principle that makes dynamic programming possible, and the analytical toolkit (the performance-difference and Bellman-residual-decomposition lemmas) used throughout the field's regret analyses — culminating in a polynomial regret guarantee for UCB-VI, the canonical optimism-based algorithm for tabular RL introduced by Azar, Osband, and Munos (Minimax Regret Bounds for Reinforcement Learning, ICML 2017, arXiv:1703.05449).

Setting

A finite-horizon episodic Markov Decision Process M=(S,A,{PhM}h=1H,{RhM}h=1H,d1)M = (S, A, \{P^M_h\}_{h=1}^H, \{R^M_h\}_{h=1}^H, d_1)M=(S,A,{PhM​}h=1H​,{RhM​}h=1H​,d1​) consists of a finite state space SSS, a finite action space AAA, a horizon HHH, per-layer transition kernels PhM:S×A→Δ(S)P^M_h : S\times A \to \Delta(S)PhM​:S×A→Δ(S) and reward distributions RhM:S×A→Δ(R)R^M_h : S\times A \to \Delta(\mathbb R)RhM​:S×A→Δ(R), and an initial state distribution d1∈Δ(S)d_1 \in \Delta(S)d1​∈Δ(S). An episode unrolls for h=1,…,Hh=1,\dots,Hh=1,…,H: the learner selects an action ah∼πh(sh)a_h \sim \pi_h(s_h)ah​∼πh​(sh​) under a randomized non-stationary policy π=(π1,…,πH)∈Πrns\pi = (\pi_1,\dots,\pi_H) \in \Pi^{\mathrm{rns}}π=(π1​,…,πH​)∈Πrns (each πh:S→Δ(A)\pi_h : S \to \Delta(A)πh​:S→Δ(A)), receives reward rh∼RhM(sh,ah)r_h \sim R^M_h(s_h,a_h)rh​∼RhM​(sh​,ah​), and transitions to sh+1∼PhM(sh,ah)s_{h+1}\sim P^M_h(s_h,a_h)sh+1​∼PhM​(sh​,ah​). The value of π\piπ is fM(π):=EM,π[∑hrh]f^M(\pi) := \mathbb E^{M,\pi}[\sum_h r_h]fM(π):=EM,π[∑h​rh​], and the state-action and state value functions QhM,π(s,a)Q^{M,\pi}_h(s,a)QhM,π​(s,a), VhM,π(s)V^{M,\pi}_h(s)VhM,π​(s) are the analogous reward-to-go quantities from layer hhh onward. In the online RL problem, M⋆M^\starM⋆ is unknown, the learner interacts with it for TTT episodes, and the goal is to minimize the regret Reg=∑t=1T(fM⋆(πM⋆)−fM⋆(πt))\mathrm{Reg} = \sum_{t=1}^T \big(f^{M^\star} (\pi^{M^\star}) - f^{M^\star}(\pi^t)\big)Reg=∑t=1T​(fM⋆(πM⋆)−fM⋆(πt)) against the best policy πM⋆∈arg⁡max⁡π∈ΠrnsfM⋆(π)\pi^{M^\star} \in \arg\max_{\pi\in\Pi^{\mathrm{rns}}} f^{M^\star}(\pi)πM⋆∈argmaxπ∈Πrns​fM⋆(π).

Formalization targets

Goal — Theorem 1 (UCB-VI regret)

∃ C>0, ∀ δ∈(0,1],Pr⁡[Reg≤C⋅H⋅S⋅A T⋅log⁡(SAHT/δ)]≥1−δ,\exists\, C>0,\ \forall\, \delta\in(0,1],\quad \Pr\Big[\mathrm{Reg} \le C\cdot H\cdot S\cdot\sqrt{A\,T}\cdot\sqrt{\log(SAHT/\delta)}\Big] \ge 1-\delta,∃C>0, ∀δ∈(0,1],Pr[Reg≤C⋅H⋅S⋅AT​⋅log(SAHT/δ)​]≥1−δ,

for the UCB-VI algorithm run with the explicit bonus bh,δt(s,a)=2log⁡(2SAHT/δ)/nht(s,a)b^t_{h,\delta}(s,a) = 2\sqrt{\log(2SAHT/\delta)/n^t_h(s,a)}bh,δt​(s,a)=2log(2SAHT/δ)/nht​(s,a)​, under Assumption 6 (deterministic, known, [0,1][0,1][0,1]-bounded rewards). This is the weakest stable form of the guarantee: it fixes only the shape of the bound (polynomial in S,A,H,TS,A,H,TS,A,H,T, logarithmic in 1/δ1/\delta1/δ), leaving the exact leading constant — which the book itself does not pin down on this page — outside the formal claim.

Reaching it rests on three structural facts, each formalized as a milestone in attack order:

  • Proposition 25 (Bellman optimality): existence of a single deterministic policy simultaneously optimal at every state, computable by backward induction — the reason dynamic programming solves planning at all.
  • Lemma 13 (Performance difference) and Lemma 14 (Bellman residual decomposition): two "credit assignment" identities decomposing a value gap (between two policies, or one policy under two models) into a sum of per-layer, on-roll-in terms.
  • Lemma 15 (Error decomposition for optimistic policies): the fact that a greedy policy driven by any optimistic value estimate suffers sub-optimality controlled additively — not exponentially — by that estimate's own Bellman residuals, evaluated on-policy.

Significance

The result itself. Theorem 1 is the chapter's headline result: the first guarantee, in this development, that a learning algorithm — one that does not know the environment's transitions in advance — can achieve regret growing only polynomially in the size of the state space, the action space, and the horizon, and only as T\sqrt TT​ in the number of episodes. The chapter's own "combination lock" example (Figure 8) shows this is not automatic: naive exploration strategies (such as ε\varepsilonε-greedy, which suffices for ordinary bandits) incur regret exponential in the horizon on some MDPs with as few as H+2H+2H+2 states. UCB-VI's guarantee is the sample-complexity foundation on which essentially every subsequent result on tabular, linear, and general function-approximation RL in the book is built.

Formalizing it. No part of this chapter's mathematical content already has a faithful counterpart on the platform (see Difficulty, below, and Formalization scope). This mission contributes: (i) a from-scratch Lean formalization of the finite-horizon episodic MDP model and its value functions, faithful to the book's 000/111-indexing and reward-distribution conventions; (ii) faithful statements (drafted with sorry, not yet proved) of Proposition 25 and Lemmas 13–15; and (iii) a faithful statement of Theorem 1 itself, including a from-scratch construction of the finite TTT-episode adaptive interaction process needed to make sense of a high-probability regret guarantee. Proving these — Proposition 25 by backward induction, Lemmas 13–15 by telescoping, and Theorem 1 by combining Lemma 15's optimism bound with a concentration argument bounding the estimation-error and martingale terms of Eqs. (5.30)–(5.33) (omitted here; see Difficulty) — is open work for solvers.

Difficulty

The obvious approach to Theorem 1 — bound the regret episode-by-episode using only the fact that QtQ^tQt is close to QM⋆,⋆Q^{M^\star,\star}QM⋆,⋆ in some fixed sense — fails because QtQ^tQt's error is itself random (it depends on the transitions observed so far) and compounds across HHH layers of dynamic programming. Lemma 15 defuses the compounding: it shows the sub-optimality gap is additive in the per-layer Bellman residuals rather than multiplicative, provided QtQ^tQt is optimistic. Making QtQ^tQt optimistic with high probability, in turn, requires a concentration argument for the empirical transition estimates P^ht\widehat P^t_hPht​ (an application of Freedman's or the Azuma–Hoeffding inequality, not included among this mission's milestones) and a union bound over all (s,a,h,t)(s,a,h,t)(s,a,h,t) — accounting for the SAHTSAHTSAHT inside the bonus's logarithm. The final regret sum further requires bounding ∑t∑h1/nht(sht,aht)\sum_t \sum_h 1/\sqrt{n^t_h(s^t_h,a^t_h)}∑t​∑h​1/nht​(sht​,aht​)​ by a pigeonhole/potential-function argument over the visitation counts, which is where the SATS\sqrt{AT}SAT​ scaling — rather than a naive SATSA\sqrt{T}SAT​ — originates. None of this concentration or counting machinery is included in the current milestones; a solver attempting Theorem 1 needs it as prerequisite lemmas.

Formalization scope

MDP and value functions. States and actions are finite types (Fintype); layers are represented 000-indexed throughout the Lean development (the book's layer hhh is h - 1), with the terminal convention V _ _ H _ = 0 matching VH+1≡0V_{H+1}\equiv 0VH+1​≡0. Since every value/expectation formula in this chapter uses the reward distribution Rh(s,a)∈Δ(R)R_h(s,a)\in\Delta(\mathbb R)Rh​(s,a)∈Δ(R) only through its mean, EpisodicMDP.R records that mean directly — equivalent, by linearity of expectation, to carrying the full distribution, and changing no theorem's content. Transition kernels and policies are represented as plain real-valued functions (S → A → ℝ-style) rather than as Mathlib's PMF, with IsPolicy/the EpisodicMDP structure's own proof fields asserting the probability-distribution properties (nonnegativity, summing to 111) where the book requires membership in Πrns\Pi^{\mathrm{rns}}Πrns or a well-formed kernel; this keeps every expectation a finite Finset.sum, needing no measure theory. The optimal value functions Qstar/Vstar are defined as literal suprema over the entire (uncountable, since ∣A∣≥2|A|\ge2∣A∣≥2) policy type — not via a recursive shortcut — which is what rules out the trivializing formalization of Proposition 25: defining Vstar by the very recursion the proposition asserts would make the proposition a tautology, whereas here it is a genuine claim about a supremum over an enormous space of competitor policies.

UCB-VI and Theorem 1. Because SSS, AAA, HHH, and TTT are all finite, the TTT-episode adaptive interaction (in which round ttt's policy is a function of the realized history of the first t−1t-1t−1 episodes) is modeled as a finite probability space: the outcome type Fin T → Trajectory S A H is a Fintype, "probability" is a finite sum over it, and "with probability ≥1−δ\ge 1-\delta≥1−δ" is a plain inequality between two real numbers — no MeasureTheory is used anywhere in this mission. The universal constant CCC in Theorem 1 is existentially quantified (∃ C > 0, …) rather than given as a literal numeral, since its value is not pinned down by the book on this page and this mission does not carry out the (non-milestoned) concentration argument that would derive it; this is the convention adopted throughout for the book's own "≲\lesssim≲" notation. The empirical-transition estimator inside QhtQ^t_hQht​ is defined to be 000 when nht(s,a)=0n^t_h(s,a)=0nht​(s,a)=0 (no data yet collected for that pair) — a boundary case Eq. (5.25) does not address, resolved here by convention rather than proof.

Prior art. The platform's existing BanditAlgorithm.UCRL2Algorithm mission formalizes UCRL2 for the average-reward, infinite-horizon MDP setting with a diameter parameter, and its Bellman-optimality statement is the average-cost optimality equation — genuinely different from this chapter's finite-horizon episodic Bellman recursion, even though both go by the name "Bellman optimality." No reference item was reused; every definition and theorem in this mission is drafted from scratch. Both missions bound regret via optimism over a confidence set of models, but for different objectives (average reward vs. finite-horizon cumulative reward) and different MDP classes.

Reusable infrastructure and open contributions. EpisodicMDP, Policy, V/Q/Qstar/Vstar, and stateDist/stateExp are reusable by any future mission on finite-horizon episodic RL in this book's later chapters. Contributions welcome: proofs of the four milestone lemmas (by backward induction and telescoping, respectively); the concentration lemmas underlying Theorem 1's optimism guarantee (Eqs. (5.30)–(5.33) of the source, not milestoned here); and the final regret proof combining them.

Selected references

  • Foster, D. J. and Rakhlin, A., Foundations of Reinforcement Learning and Interactive Decision Making, 2023. arXiv:2312.16730v1
  • Azar, M. G., Osband, I., and Munos, R., Minimax Regret Bounds for Reinforcement Learning, ICML 2017. arXiv:1703.05449
  • Jaksch, T., Ortner, R., and Auer, P., Near-optimal Regret Bounds for Reinforcement Learning, JMLR 11 (2010).
7 thms3 active usersReviewed
🏆Completed
Bandit AlgorithmsOperations Research·Captain: Shuze Chen

Bandit Algorithms I: Concentration of MeasureTextbook

How quickly does the empirical mean of independent random variables concentrate around the true mean? This question is the analytic engine of the entire theory of stochastic bandits: every optimistic algorithm (Explore-Then-Commit, UCB and its relatives) is calibrated by a tail bound on the sample mean. This mission formalizes the subgaussian framework of Chapter 5 of Lattimore–Szepesvári's Bandit Algorithms: a random variable XXX is σ\sigmaσ-subgaussian when E[eλX]≤eλ2σ2/2\mathbb{E}[e^{\lambda X}] \le e^{\lambda^2\sigma^2/2}E[eλX]≤eλ2σ2/2 for all λ\lambdaλ, and the Cramér–Chernoff method converts this moment-generating-function control into the exponential tail P(X≥ε)≤e−ε2/(2σ2)\mathbb{P}(X \ge \varepsilon) \le e^{-\varepsilon^2/(2\sigma^2)}P(X≥ε)≤e−ε2/(2σ2). The goal theorem is the Hoeffding-type bound: the sample mean of nnn independent σ\sigmaσ-subgaussian deviations exceeds the true mean by ε\varepsilonε with probability at most exp⁡(−nε2/(2σ2))\exp(-n\varepsilon^2/(2\sigma^2))exp(−nε2/(2σ2)), together with its confidence form P(μ^+2σ2log⁡(1/δ)/n≤μ)≤δ\mathbb{P}\big(\hat\mu + \sqrt{2\sigma^2\log(1/\delta)/n} \le \mu\big) \le \deltaP(μ^​+2σ2log(1/δ)/n​≤μ)≤δ — the exact bound every UCB index is built from. These few lines of analysis are cited by every regret bound in the series.

2 thms3 active usersReviewed
ProbabilityStatistics·Captain: mikedeng1

Rademacher and Gaussian Complexities: Risk Bounds and Structural Results 5: Kernel Expansions with α′Kα ≤ B² Have Rademacher and Gaussian Complexity at Most 2B√(E k(X,X)/n)Research Paper

Motivation

Kernel methods, such as support vector machines, predict with functions of the form x↦∑iαik(x,xi)x \mapsto \sum_i \alpha_i k(x, x_i)x↦∑i​αi​k(x,xi​): finite expansions of a fixed similarity function kkk centred at data points. Their statistical behaviour is governed by the size of the class of such expansions that the method searches. Bartlett and Mendelson, in Rademacher and Gaussian Complexities: Risk Bounds and Structural Results (JMLR 3, 2002), develop risk bounds in terms of the Rademacher and Gaussian complexities of a class, and in §4.3 (pp. 476–478) compute these complexities for the class of kernel expansions whose coefficient vector has quadratic form α′Kα≤B2\alpha' K \alpha \le B^2α′Kα≤B2. The resulting bound depends on the kernel only through its diagonal k(x,x)k(x,x)k(x,x), which is what makes margin bounds for support vector machines dimension-free. This mission formalizes that computation. The source is the published JMLR article (pages cited by the journal's printed numbers).

Setting

Let X\mathcal XX be a compact topological space. A kernel is a continuous function k:X×X→Rk : \mathcal X \times \mathcal X \to \mathbb Rk:X×X→R such that for every mmm and all x1,…,xm∈Xx_1, \dots, x_m \in \mathcal Xx1​,…,xm​∈X the Gram matrix Kij=k(xi,xj)K_{ij} = k(x_i, x_j)Kij​=k(xi​,xj​) is symmetric and positive semidefinite. For B≥0B \ge 0B≥0 the class of kernel expansions is

F={x↦∑i=1mαik(x,xi):m∈N, xi∈X, αi∈R, ∑i,jαiαjk(xi,xj)≤B2},F = \Big\{x \mapsto \sum_{i=1}^m \alpha_i k(x, x_i) : m \in \mathbb N,\ x_i \in \mathcal X,\ \alpha_i \in \mathbb R,\ \sum_{i,j}\alpha_i\alpha_j k(x_i, x_j) \le B^2\Big\},F={x↦i=1∑m​αi​k(x,xi​):m∈N, xi​∈X, αi​∈R, i,j∑​αi​αj​k(xi​,xj​)≤B2},

with centres anywhere in X\mathcal XX (Lean: kernelClass k B).

For a class FFF of real functions on X\mathcal XX and a sample x1,…,xnx_1, \dots, x_nx1​,…,xn​, let σ1,…,σn\sigma_1, \dots, \sigma_nσ1​,…,σn​ be independent uniform signs and g1,…,gng_1, \dots, g_ng1​,…,gn​ independent standard Gaussians. The empirical Rademacher complexity and empirical Gaussian complexity (Definition 2, p. 464) are

R^n(F)=Eσsup⁡f∈F∣2n∑i=1nσif(xi)∣,G^n(F)=Egsup⁡f∈F∣2n∑i=1ngif(xi)∣.\hat R_n(F) = \mathbb E_\sigma \sup_{f\in F}\Big|\frac2n\sum_{i=1}^n \sigma_i f(x_i)\Big|, \qquad \hat G_n(F) = \mathbb E_g \sup_{f\in F}\Big|\frac2n\sum_{i=1}^n g_i f(x_i)\Big|.R^n​(F)=Eσ​f∈Fsup​​n2​i=1∑n​σi​f(xi​)​,G^n​(F)=Eg​f∈Fsup​​n2​i=1∑n​gi​f(xi​)​.

For a probability measure μ\muμ on X\mathcal XX and an i.i.d. sample X1,…,Xn∼μX_1, \dots, X_n \sim \muX1​,…,Xn​∼μ, the Rademacher complexity is Rn(F)=ER^n(F)R_n(F) = \mathbb E \hat R_n(F)Rn​(F)=ER^n​(F) and the Gaussian complexity is Gn(F)=EG^n(F)G_n(F) = \mathbb E \hat G_n(F)Gn​(F)=EG^n​(F) (Lean: empiricalRademacher, empiricalGaussian, rademacherComplexity, gaussianComplexity).

A feature map of kkk is a map Φ:X→H\Phi : \mathcal X \to \mathcal HΦ:X→H into a real Hilbert space with k(x1,x2)=⟨Φ(x1),Φ(x2)⟩k(x_1, x_2) = \langle \Phi(x_1), \Phi(x_2) \ranglek(x1​,x2​)=⟨Φ(x1​),Φ(x2​)⟩.

Formalization targets

Goal: the expected complexity bound (§4.3, p. 478, display after the proof of Lemma 22)

With X∼μX \sim \muX∼μ,

Rn(F)≤2BE k(X,X)n,Gn(F)≤2BE k(X,X)n.R_n(F) \le 2B\sqrt{\frac{\mathbb E\, k(X,X)}{n}}, \qquad G_n(F) \le 2B\sqrt{\frac{\mathbb E\, k(X,X)}{n}}.Rn​(F)≤2BnEk(X,X)​​,Gn​(F)≤2BnEk(X,X)​​.

Milestone 1: feature-map inclusion (p. 477)

For any feature map Φ\PhiΦ of kkk, ∥∑iαiΦ(xi)∥2=∑i,jαiαjk(xi,xj)\|\sum_i \alpha_i \Phi(x_i)\|^2 = \sum_{i,j}\alpha_i\alpha_j k(x_i,x_j)∥∑i​αi​Φ(xi​)∥2=∑i,j​αi​αj​k(xi​,xj​), and hence F⊆{x↦⟨w,Φ(x)⟩:∥w∥≤B}F \subseteq \{x \mapsto \langle w, \Phi(x)\rangle : \|w\| \le B\}F⊆{x↦⟨w,Φ(x)⟩:∥w∥≤B}.

Milestone 2: Lemma 22 (p. 477)

For every sample X1,…,XnX_1, \dots, X_nX1​,…,Xn​,

G^n(F)≤2Bn∑i=1nk(Xi,Xi),R^n(F)≤2Bn∑i=1nk(Xi,Xi).\hat G_n(F) \le \frac{2B}{n}\sqrt{\sum_{i=1}^n k(X_i, X_i)}, \qquad \hat R_n(F) \le \frac{2B}{n}\sqrt{\sum_{i=1}^n k(X_i, X_i)}.G^n​(F)≤n2B​i=1∑n​k(Xi​,Xi​)​,R^n​(F)≤n2B​i=1∑n​k(Xi​,Xi​)​.

Significance

The goal shows that the kernel class has complexity of order n−1/2n^{-1/2}n−1/2, with a constant given by BBB and the quantity E k(X,X)\mathbb E\, k(X,X)Ek(X,X), which is the trace of the integral operator Tkf=∫k(⋅,y)f(y) dμ(y)T_k f = \int k(\cdot, y) f(y)\, d\mu(y)Tk​f=∫k(⋅,y)f(y)dμ(y) on L2(μ)L_2(\mu)L2​(μ). No dimension of the feature space enters. Fed into the paper's margin-cost risk bound (Theorem 21, p. 476), the sample-wise Lemma 22 gives a data-dependent misclassification bound for support vector machines in terms of the trace of the Gram matrix of the training sample. Bounds of this form are the standard complexity estimate for kernel classes in learning theory textbooks.

The result is proved in the paper; nothing in this mission is open mathematics. What the mission adds is a machine-checked version with the paper's normalization (factor 2/n2/n2/n, absolute value inside the supremum), covering both the Rademacher and the Gaussian complexity, for expansions with centres anywhere in X\mathcal XX. Related statements on the platform (Mohri et al.'s Theorem 5.10 and Proposition 9.3) use the 1/n1/n1/n normalization without absolute value, assume a uniform bound sup⁡xk(x,x)≤r2\sup_x k(x,x) \le r^2supx​k(x,x)≤r2, and treat only the Rademacher case, so they do not imply the targets here.

Difficulty

The class FFF is defined through the kernel, not through a feature map, and its expansions have arbitrarily many centres anywhere in X\mathcal XX. The supremum over FFF is therefore a supremum over an infinite-dimensional family, and must be controlled without assuming a separate numerical upper bound on k(x,x)k(x,x)k(x,x) or that the feature space is finite dimensional. For the Gaussian complexity the supremum sits inside an expectation over a continuous random vector, and the passage from the sample-wise bound to the expected bound must move an expectation inside a square root in the right direction. A bound with sup⁡xk(x,x)\sup_x k(x,x)supx​k(x,x) in place of E k(X,X)\mathbb E\, k(X,X)Ek(X,X) is weaker and is not the target.

Formalization scope

Conventions committed to in Lean:

  • Complexities take values in [0,∞][0, \infty][0,∞] (ℝ≥0∞); expectations over the Gaussian vector and over the sample are lower Lebesgue integrals against product measures, and the Rademacher expectation is the exact average over the 2n2^n2n sign vectors σ:Fin n→Z×\sigma : \mathrm{Fin}\,n \to \mathbb Z^\timesσ:Finn→Z×. An unbounded class has complexity +∞+\infty+∞, so a junk value of 000 for a real supremum or a non-integrable expectation cannot make the bounds trivial.
  • A kernel (IsKernel k) carries compactness of X\mathcal XX, joint continuity, and positive semidefiniteness (with symmetry) of every Gram matrix, as in the paper's definition. The goal puts the Borel σ\sigmaσ-algebra on X\mathcal XX and assumes μ\muμ is a probability measure; E k(X,X)\mathbb E\, k(X,X)Ek(X,X) is the Bochner integral of the continuous function x↦k(x,x)x \mapsto k(x,x)x↦k(x,x).
  • B≥0B \ge 0B≥0 is assumed in every statement (the paper fixes B>0B > 0B>0). For B<0B < 0B<0 the right-hand sides are negative while the left-hand sides are not, and the inclusion of milestone 1 fails.
  • No n≥1n \ge 1n≥1 hypothesis: at n=0n = 0n=0 both sides of every bound are 000 in Lean.
  • The feature map in milestone 1 is a hypothesis (any real Hilbert space and any Φ\PhiΦ with k=⟨Φ(⋅),Φ(⋅)⟩k = \langle \Phi(\cdot), \Phi(\cdot)\ranglek=⟨Φ(⋅),Φ(⋅)⟩); its existence is the RKHS theorem, referenced as the supporting platform item FoundationsML.Kernels.RKHS_exists.
  • No measurability or integrability hypotheses on R^n(F)\hat R_n(F)R^n​(F) or G^n(F)\hat G_n(F)G^n​(F) are assumed, and none is needed.

A trivializing formalization is ruled out: the goal is stated for the kernel class FFF itself, not for the larger ball of linear functionals of a feature map, and not for the subclass with centres at the sample points.

Infrastructure a complete development needs: Gaussian integration over Rn\mathbb R^nRn (second moments of a standard Gaussian vector), Jensen's inequality for the square root under a lower Lebesgue integral, and Cauchy–Schwarz for positive semidefinite bilinear forms (or the RKHS feature map). The Definition 2 complexities are shared with the paper's other missions and reusable. Contributions of any of these milestones, and of proofs of the goal that avoid the feature map, are welcome.

Selected references

  • P. L. Bartlett and S. Mendelson, Rademacher and Gaussian Complexities: Risk Bounds and Structural Results, Journal of Machine Learning Research 3 (2002), 463–482. https://www.jmlr.org/papers/v3/bartlett02a.html
  • N. Cristianini and J. Shawe-Taylor, An Introduction to Support Vector Machines, Cambridge University Press, 2000. https://doi.org/10.1017/CBO9780511801389
  • N. Aronszajn, Theory of Reproducing Kernels, Transactions of the American Mathematical Society 68 (1950), 337–404. https://doi.org/10.1090/S0002-9947-1950-0051437-7
  • M. Mohri, A. Rostamizadeh and A. Talwalkar, Foundations of Machine Learning, 2nd ed., MIT Press, 2018. https://mitpress.mit.edu/9780262039406/
8 thms2 active usersReviewed
ProbabilityStatistics·Captain: mikedeng1

Local Rademacher Complexities II: Local Rademacher Averages of the Classification Loss Class Are Bounded by Weighted Empirical Risk Minimization (Theorem 6.3)Research Paper

Motivation

Local Rademacher averages measure the complexity of a learning problem only near the functions that matter, such as those with small empirical error, rather than over the whole function class. Bartlett, Bousquet and Mendelson (Local Rademacher complexities, Ann. Statist. 33 (2005)) show that error bounds for empirical risk minimization are governed by the fixed point of a sub-root upper bound on such local averages, and that these bounds give fast rates (1/n1/n1/n rather than 1/n1/\sqrt n1/n​) under variance conditions. A bound is only useful in practice if it can be computed from the data. For classification with the discrete loss, the paper's Corollary 6.2 states the bound in terms of a localized empirical Rademacher average ψ^n(r)\hat\psi_n(r)ψ^​n​(r). That average is a supremum over a constrained subclass, and it is not obvious how to evaluate it.

Theorem 6.3 of the paper answers this. An upper bound on ψ^n(r)\hat\psi_n(r)ψ^​n​(r) can be computed by any algorithm that minimizes a weighted empirical classification error. A similar reduction was known for the global Rademacher average of a classification class: Bartlett, Boucheron and Lugosi (Model selection and error estimation, Machine Learning 48 (2002)) observed that the empirical Rademacher average equals one half minus an expected empirical risk minimum with random labels, and Lemma 6.4 of the paper is adapted from their argument. Theorem 6.3 shows that localization and the use of star-hulls keep this reduction intact.

Setting

Fix inputs X1,…,XnX_1,\dots,X_nX1​,…,Xn​ in a set X\mathcal XX (n≥1n\ge1n≥1) and labels Y1,…,Yn∈{−1,1}Y_1,\dots,Y_n\in\{-1,1\}Y1​,…,Yn​∈{−1,1}. Everything below is deterministic given this sample. A classifier is a function f:X→{−1,1}f:\mathcal X\to\{-1,1\}f:X→{−1,1}, and F\mathcal FF is a class of classifiers. The discrete loss is ℓ(y,y′)=1[y≠y′]\ell(y,y')=\mathbf 1[y\ne y']ℓ(y,y′)=1[y=y′]. For a vector z∈Rnz\in\mathbb R^nz∈Rn write

Pnℓ(f(X),z)=1n∑i=1nℓ(f(Xi),zi),Pnℓf=Pnℓ(f(X),Y),P_n\ell(f(X),z)=\frac1n\sum_{i=1}^n\ell(f(X_i),z_i),\qquad P_n\ell_f=P_n\ell(f(X),Y),Pn​ℓ(f(X),z)=n1​i=1∑n​ℓ(f(Xi​),zi​),Pn​ℓf​=Pn​ℓ(f(X),Y),

so PnℓfP_n\ell_fPn​ℓf​ is the empirical risk of fff.

A sign vector σ∈{−1,1}n\sigma\in\{-1,1\}^nσ∈{−1,1}n plays the role of Rademacher signs, and Eσ\mathbb E_\sigmaEσ​ is the average over all 2n2^n2n sign vectors. The empirical Rademacher average of the loss functions of the classifiers with empirical risk at most bbb is

EσRn{ℓf:f∈F, Pnℓf≤b}=1n Eσsup⁡f∈F, Pnℓf≤b ∑i=1nσi ℓ(f(Xi),Yi).\mathbb E_\sigma R_n\{\ell_f : f\in\mathcal F,\ P_n\ell_f\le b\}=\frac1n\,\mathbb E_\sigma\sup_{f\in\mathcal F,\ P_n\ell_f\le b}\ \sum_{i=1}^n\sigma_i\,\ell(f(X_i),Y_i).Eσ​Rn​{ℓf​:f∈F, Pn​ℓf​≤b}=n1​Eσ​f∈F, Pn​ℓf​≤bsup​ i=1∑n​σi​ℓ(f(Xi​),Yi​).

For c≥0c\ge0c≥0, x>0x>0x>0 and 0<r≤1/20<r\le1/20<r≤1/2, the empirical local Rademacher complexity of the classification loss class is

ψ^n(r)=csup⁡α∈[2r,1]α EσRn{ℓf:f∈F, Pnℓf≤2r/α2}+26xn.\hat\psi_n(r)=c\sup_{\alpha\in[\sqrt{2r},1]}\alpha\,\mathbb E_\sigma R_n\{\ell_f : f\in\mathcal F,\ P_n\ell_f\le 2r/\alpha^2\}+\frac{26x}{n}.ψ^​n​(r)=cα∈[2r​,1]sup​αEσ​Rn​{ℓf​:f∈F, Pn​ℓf​≤2r/α2}+n26x​.

In Corollary 6.2, c=20c=20c=20. The parameter α\alphaα comes from the star-hull of the loss class: rescaling a loss function by α\alphaα turns the constraint Pn(αℓf)2≤2rP_n(\alpha\ell_f)^2\le 2rPn​(αℓf​)2≤2r into Pnℓf≤2r/α2P_n\ell_f\le 2r/\alpha^2Pn​ℓf​≤2r/α2.

For a sign vector σ\sigmaσ and a multiplier μ≥0\mu\ge0μ≥0, the weighted empirical risk minimum is

J(μ)=min⁡f∈F1n∑i=1n∣σi+μYi∣ ℓ(f(Xi),sign⁡(σi+μYi)).J(\mu)=\min_{f\in\mathcal F}\frac1n\sum_{i=1}^n|\sigma_i+\mu Y_i|\,\ell\big(f(X_i),\operatorname{sign}(\sigma_i+\mu Y_i)\big).J(μ)=f∈Fmin​n1​i=1∑n​∣σi​+μYi​∣ℓ(f(Xi​),sign(σi​+μYi​)).

It is the smallest weighted training error when the labels are corrupted to sign⁡(σi+μYi)\operatorname{sign}(\sigma_i+\mu Y_i)sign(σi​+μYi​) and example iii has weight ∣σi+μYi∣|\sigma_i+\mu Y_i|∣σi​+μYi​∣.

Formalization targets

Goal: Theorem 6.3

If some f∈Ff\in\mathcal Ff∈F has Pnℓf≤2rP_n\ell_f\le 2rPn​ℓf​≤2r, then

ψ^n(r)≤csup⁡α∈[2r,1]α Eσmin⁡μ≥0((2rα2−12)μ+12n∑i=1n∣σi+μYi∣−J(μ))+26xn.\hat\psi_n(r)\le c\sup_{\alpha\in[\sqrt{2r},1]}\alpha\,\mathbb E_\sigma\min_{\mu\ge0}\Big(\Big(\frac{2r}{\alpha^2}-\frac12\Big)\mu+\frac1{2n}\sum_{i=1}^n|\sigma_i+\mu Y_i|-J(\mu)\Big)+\frac{26x}{n}.ψ^​n​(r)≤cα∈[2r​,1]sup​αEσ​μ≥0min​((α22r​−21​)μ+2n1​i=1∑n​∣σi​+μYi​∣−J(μ))+n26x​.

The multiplier ccc is kept general, and the term 26x/n26x/n26x/n appears on both sides as printed.

Milestones

  1. Lemma 6.4. For every b∈[0,1]b\in[0,1]b∈[0,1] with a feasible classifier,
EσRn{ℓf:f∈F, Pnℓf≤b}=12−Eσmin⁡{Pnℓ(f(X),σ):f∈F, Pnℓ(f(X),Y)≤b}.\mathbb E_\sigma R_n\{\ell_f : f\in\mathcal F,\ P_n\ell_f\le b\}=\frac12-\mathbb E_\sigma\min\{P_n\ell(f(X),\sigma) : f\in\mathcal F,\ P_n\ell(f(X),Y)\le b\}.Eσ​Rn​{ℓf​:f∈F, Pn​ℓf​≤b}=21​−Eσ​min{Pn​ℓ(f(X),σ):f∈F, Pn​ℓ(f(X),Y)≤b}.
  1. Weak duality (proof of Theorem 6.3). With L(f,μ)=Pnℓ(f(X),σ)+μ(Pnℓ(f(X),Y)−2r/α2)L(f,\mu)=P_n\ell(f(X),\sigma)+\mu(P_n\ell(f(X),Y)-2r/\alpha^2)L(f,μ)=Pn​ℓ(f(X),σ)+μ(Pn​ℓ(f(X),Y)−2r/α2) and g(μ)=min⁡f∈FL(f,μ)g(\mu)=\min_{f\in\mathcal F}L(f,\mu)g(μ)=minf∈F​L(f,μ), for every μ≥0\mu\ge0μ≥0,
min⁡{Pnℓ(f(X),σ):f∈F, Pnℓ(f(X),Y)≤2r/α2}≥g(μ).\min\{P_n\ell(f(X),\sigma) : f\in\mathcal F,\ P_n\ell(f(X),Y)\le 2r/\alpha^2\}\ge g(\mu).min{Pn​ℓ(f(X),σ):f∈F, Pn​ℓ(f(X),Y)≤2r/α2}≥g(μ).
  1. The identity for g(μ)g(\mu)g(μ) (proof of Theorem 6.3, corrected).
g(μ)=J(μ)−12n∑i=1n∣σi+μYi∣+1+μ2−μ2rα2.g(\mu)=J(\mu)-\frac1{2n}\sum_{i=1}^n|\sigma_i+\mu Y_i|+\frac{1+\mu}2-\mu\frac{2r}{\alpha^2}.g(μ)=J(μ)−2n1​i=1∑n​∣σi​+μYi​∣+21+μ​−μα22r​.

Significance

The theorem turns a quantity defined by a supremum over a data-dependent subclass into one computable by a standard learning primitive. For each sign vector and each multiplier μ\muμ, J(μ)J(\mu)J(μ) is the value of a weighted classification problem, which any weighted empirical risk minimizer solves. The expectation over signs can be estimated by repeated sampling. The paper notes that JJJ is Lipschitz in μ\muμ, so a finite grid of μ\muμ values suffices, and that a sub-root upper bound on ψ^n\hat\psi_nψ^​n​ can then be read off. Combined with Corollary 6.2, this yields error bounds for empirical risk minimization in classification that are computable from the training data and that localize: they depend only on the classifiers with small empirical error.

The result is proved in the paper. It has not been formalized; as far as a search of the Prove2Me catalog shows, neither the classification loss class nor J(μ)J(\mu)J(μ) exists as a formal object. This mission produces machine-checked statements of the theorem and of its three proof steps. These cover the exact identity between Rademacher averages of the discrete loss class and random-label empirical risk minimization, and a Lagrangian duality bound for constrained empirical risk minimization.

Difficulty

The obvious route is to apply Lemma 6.4 and then exchange the constrained minimum for a Lagrangian. Each step has a point where a careless argument fails.

  • Lemma 6.4 needs a change of variables on sign vectors (σi↦−Yiσi\sigma_i\mapsto-Y_i\sigma_iσi​↦−Yi​σi​) that preserves the uniform average. It also needs the identity ℓ(y,y′)=∣y−y′∣/2\ell(y,y')=|y-y'|/2ℓ(y,y′)=∣y−y′∣/2 on {±1}\{\pm1\}{±1}, which fails off {±1}\{\pm1\}{±1}.
  • The Lagrangian step gives only weak duality. The bound is an inequality, and attempts to prove equality in Theorem 6.3 fail in general.
  • The identity for g(μ)g(\mu)g(μ) rests on ℓ(y,y^)=(1−yy^)/2\ell(y,\hat y)=(1-y\hat y)/2ℓ(y,y^​)=(1−yy^​)/2. This holds only for ±1\pm1±1 arguments, while sign⁡(σi+μYi)\operatorname{sign}(\sigma_i+\mu Y_i)sign(σi​+μYi​) is 000 when μ=1\mu=1μ=1 and σi=−Yi\sigma_i=-Y_iσi​=−Yi​. Those terms carry weight zero, and the bookkeeping has to show this.
  • Passing the per-α\alphaα, per-σ\sigmaσ inequalities through the outer supremum and the average requires every supremum and minimum to be over a nonempty, bounded set. This is where the feasibility hypothesis is used.

Formalization scope

  • Representation. Inputs are xs : Fin n → X for an arbitrary type X. Labels and signs are real vectors Fin n → ℝ. Classifiers are functions X → ℝ with values in {±1}\{\pm1\}{±1}, a class is a Set (X → ℝ), and the discrete loss is defined on all real pairs. Sign vectors are indexed by Fin n → Bool through the published UnderstandingML.signVec. Every Eσ\mathbb E_\sigmaEσ​, on both sides of every statement, is the finite average over these 2n2^n2n vectors; no probability measure is used. The empirical Rademacher average is the published UnderstandingML.rademacher applied to the set of loss vectors.
  • Suprema and minima. Every supremum and minimum is Lean's real ⨆/⨅ over a subtype. The hypotheses make each index set nonempty and each family bounded, so these are true suprema and minima. The convention Real.sign 0 = 0 is used where the paper's sign is undefined; it affects only weight-zero terms.
  • Added hypotheses. Each of these is implicit on the page:
    • n≥1n\ge1n≥1;
    • c≥0c\ge0c≥0 (for c<0c<0c<0 the inequality reverses);
    • 0<r≤1/20<r\le1/20<r≤1/2 (otherwise the range of α\alphaα is empty);
    • x>0x>0x>0 (Corollary 6.2's "fix x>0x>0x>0");
    • a classifier with Pnℓf≤2rP_n\ell_f\le 2rPn​ℓf​≤2r in the goal, and a feasible classifier in Lemma 6.4 and in the weak duality step (the page's minima presuppose one);
    • a nonempty F\mathcal FF in the g(μ)g(\mu)g(μ) identity.
  • Corrections of the print. The last display of the proof on p. 30 ends each line with −2r/α2-2r/\alpha^2−2r/α2. From the page's own definition g(μ)=min⁡fL(f,μ)g(\mu)=\min_f L(f,\mu)g(μ)=minf​L(f,μ), the constant is −μ 2r/α2-\mu\,2r/\alpha^2−μ2r/α2, which is the form Theorem 6.3's term (2r/α2−1/2)μ(2r/\alpha^2-1/2)\mu(2r/α2−1/2)μ requires. The milestone states the corrected identity.
  • No trivialization. Without the feasibility hypothesis, Lean would evaluate the empty-class Rademacher average and the unbounded μ\muμ-minimum to the junk value 000, and the goal would compare junk values. The feasibility hypothesis rules this out. The goal is the inequality between the two expressions for ψ^n\hat\psi_nψ^​n​ as printed; it is not restated through g(μ)g(\mu)g(μ), L(f,μ)L(f,\mu)L(f,μ) or Lemma 6.4.
  • Contributions welcome. A reusable lemma that the uniform average over {±1}n\{\pm1\}^n{±1}n is invariant under coordinatewise sign flips would serve beyond this mission, as would general facts about real infima over finite-valued families. Proofs of the three milestones, and of the goal from them, are the main targets.

Selected references

  • P. L. Bartlett, O. Bousquet, S. Mendelson, Local Rademacher complexities, Annals of Statistics 33(4), 1497–1537, 2005. arXiv:math/0508275v1 (cited version): https://arxiv.org/abs/math/0508275, DOI https://doi.org/10.1214/009053605000000282 — §6.2, Corollary 6.2 and Theorem 6.3 (pp. 28–29), Lemma 6.4 (p. 29), proof of Theorem 6.3 (p. 30).
  • P. L. Bartlett, S. Boucheron, G. Lugosi, Model selection and error estimation, Machine Learning 48, 85–113, 2002. https://doi.org/10.1023/A:1013999503812
  • P. L. Bartlett, S. Mendelson, Rademacher and Gaussian complexities: risk bounds and structural results, Journal of Machine Learning Research 3, 463–482, 2002. https://www.jmlr.org/papers/v3/bartlett02a.html
  • S. Shalev-Shwartz, S. Ben-David, Understanding Machine Learning: From Theory to Algorithms, Cambridge University Press, 2014, Chapter 26 (the Rademacher complexity reused here). https://doi.org/10.1017/CBO9781107298019
6 thms2 active usersReviewed
ProbabilityStatistics·Captain: mikedeng1

Rademacher and Gaussian Complexities: Risk Bounds and Structural Results 4: A Fixed Boolean Combination of k Classes Has Gaussian Complexity at Most 2 Σ_j G_n(F_j)Research Paper

Motivation

Data-dependent risk bounds in statistical learning replace combinatorial quantities such as the VC dimension with averages of how well a function class can fit random noise on the observed sample. Bartlett and Mendelson's article (JMLR 3, 2002) established these averages, the Rademacher and Gaussian complexities, as a general tool: a risk bound (their Theorem 8) holds with a complexity penalty, and the complexity of a complicated class can be bounded through structural results that relate it to the complexities of simpler classes.

This mission formalizes two of those structural results, both stated for Gaussian complexities. The first (Theorem 14) controls a Lipschitz function of several real-valued classes at once, the vector-valued analogue of the classical contraction principle. The second (Theorem 16) controls an arbitrary fixed boolean combination of classes of classifiers, such as intersections, unions or majority votes of a fixed number of base classifiers, by the sum of the complexities of the components. Such combinations arise whenever a classifier is assembled from simpler ones, for instance in decision lists, small decision trees over a base class, or voting schemes.

Setting

Let X\mathcal XX be a set, μ\muμ a probability measure on it, and n≥1n \ge 1n≥1 a sample size. For a class FFF of functions X→R\mathcal X \to \mathbb RX→R and a sample x=(x1,…,xn)x = (x_1, \dots, x_n)x=(x1​,…,xn​), the empirical Gaussian complexity is

G^n(F)(x)=E[sup⁡f∈F∣2n∑i=1ngif(xi)∣],\hat G_n(F)(x) = \mathbb E\left[\sup_{f\in F}\left|\frac2n\sum_{i=1}^n g_i f(x_i)\right|\right],G^n​(F)(x)=E[f∈Fsup​​n2​i=1∑n​gi​f(xi​)​],

with g1,…,gng_1, \dots, g_ng1​,…,gn​ independent standard Gaussian N(0,1)N(0,1)N(0,1) variables, and the Gaussian complexity is Gn(F)=E G^n(F)(X1,…,Xn)G_n(F) = \mathbb E\, \hat G_n(F)(X_1, \dots, X_n)Gn​(F)=EG^n​(F)(X1​,…,Xn​) for X1,…,XnX_1, \dots, X_nX1​,…,Xn​ i.i.d. with law μ\muμ (Definition 2, p. 464). In Lean these are empiricalGaussian n F x and gaussianComplexity μ n F.

Three constructions of classes appear.

  • Direct sum. With A=Rm\mathcal A = \mathbb R^mA=Rm carrying the Euclidean distance, a class FFF of maps X→A\mathcal X \to \mathcal AX→A is a subset of the direct sum of real classes F1,…,FmF_1, \dots, F_mF1​,…,Fm​ when each f∈Ff \in Ff∈F is x↦(f1(x),…,fm(x))x \mapsto (f_1(x), \dots, f_m(x))x↦(f1​(x),…,fm​(x)) with fi∈Fif_i \in F_ifi​∈Fi​ (SubsetDirectSum F Fi).
  • Composition. For ϕ:Y×A→R\phi : \mathcal Y \times \mathcal A \to \mathbb Rϕ:Y×A→R, ϕ∘f\phi \circ fϕ∘f is (x,y)↦ϕ(y,f(x))(x, y) \mapsto \phi(y, f(x))(x,y)↦ϕ(y,f(x)) and ϕ∘F\phi\circ Fϕ∘F collects these (compClass φ F).
  • Boolean combination. For g:{±1}k→{±1}g : \{\pm1\}^k \to \{\pm1\}g:{±1}k→{±1} and classes F1,…,FkF_1, \dots, F_kF1​,…,Fk​ of {±1}\{\pm1\}{±1}-valued functions, g(F1,…,Fk)={x↦g(f1(x),…,fk(x)):fj∈Fj}g(F_1, \dots, F_k) = \{x \mapsto g(f_1(x), \dots, f_k(x)) : f_j \in F_j\}g(F1​,…,Fk​)={x↦g(f1​(x),…,fk​(x)):fj​∈Fj​} (boolComb g F).

A centred Gaussian process indexed by a finite set III is a family (Xi)i∈I(X_i)_{i \in I}(Xi​)i∈I​ of real random variables whose finite-dimensional laws are jointly Gaussian with mean zero; ∥Xi−Xj∥2=(E(Xi−Xj)2)1/2\|X_i - X_j\|_2 = (\mathbb E(X_i - X_j)^2)^{1/2}∥Xi​−Xj​∥2​=(E(Xi​−Xj​)2)1/2.

Formalization targets

Goal: Theorem 16 (p. 472)

For a fixed boolean function g:{±1}k→{±1}g : \{\pm1\}^k \to \{\pm1\}g:{±1}k→{±1} with k≥1k \ge 1k≥1 and classes F1,…,FkF_1, \dots, F_kF1​,…,Fk​ of {±1}\{\pm1\}{±1}-valued functions,

Gn(g(F1,…,Fk))≤2∑j=1kGn(Fj).G_n\bigl(g(F_1, \dots, F_k)\bigr) \le 2 \sum_{j=1}^k G_n(F_j).Gn​(g(F1​,…,Fk​))≤2j=1∑k​Gn​(Fj​).

Milestones

  1. Lemma 13 (p. 471), the comparison of Gaussian processes as printed: if ∥Xi−Xj∥2≤∥Yi−Yj∥2\|X_i - X_j\|_2 \le \|Y_i - Y_j\|_2∥Xi​−Xj​∥2​≤∥Yi​−Yj​∥2​ for all i,ji, ji,j, then Esup⁡iXi≤2 Esup⁡iYi\mathbb E\sup_i X_i \le 2\,\mathbb E\sup_i Y_iEsupi​Xi​≤2Esupi​Yi​.
  2. Theorem 14 (p. 471): if each ϕ(y,⋅)\phi(y, \cdot)ϕ(y,⋅) is LLL-Lipschitz for the Euclidean distance, passes through the origin, and ϕ\phiϕ is uniformly bounded, then for every sample (xk,yk)k≤n(x_k, y_k)_{k \le n}(xk​,yk​)k≤n​,
G^n(ϕ∘F)≤2L∑i=1mG^n(Fi).\hat G_n(\phi \circ F) \le 2L \sum_{i=1}^m \hat G_n(F_i).G^n​(ϕ∘F)≤2Li=1∑m​G^n​(Fi​).
  1. The extension of ggg (proof of Theorem 16, p. 472): g(x)=(1−∥x−a∥)g(a)g(x) = (1 - \|x - a\|)g(a)g(x)=(1−∥x−a∥)g(a) when ∥x−a∥<1\|x - a\| < 1∥x−a∥<1 for a cube vertex aaa, and 000 otherwise, is well defined, extends ggg, maps into [−1,1][-1,1][−1,1], vanishes at 000 and is 111-Lipschitz.

Significance

Theorem 16 turns any bound on the Gaussian complexity of base classes into a bound for a fixed boolean combination of them, at the cost of a factor 222 on the sum of their complexities, whatever ggg and kkk are. Together with the comparison between Gaussian and Rademacher complexities (Lemma 4 of the paper) and the risk bound of Theorem 8, it yields generalization bounds for classifiers built as combinations of base classifiers. Theorem 14 is the general tool: it handles any Lipschitz loss of a vector-valued predictor, such as multiclass margins, through the complexities of its coordinate classes.

All three results are proved in the paper (Lemma 13 is classical and cited from Pisier). None of them is formalized on Prove2Me; the finite-dimensional Sudakov–Fernique inequality, with constant 111, is (HighDimProb.RandomProcesses.sudakov_fernique_finite_dim). This mission produces machine-checked versions of the vector contraction for Gaussian averages and of the boolean-combination bound, and records the corrections the printed statements need.

Difficulty

The obvious approach to Theorem 14 compares two Gaussian processes indexed by the class, but Definition 2 takes the supremum of an absolute value scaled by 2/n2/n2/n, while Gaussian comparison inequalities bound the expected supremum of the process itself. The printed proof equates the two; done carefully, the comparison with the printed constant 222 of Lemma 13 gives only 4L4L4L. Reaching the printed 2L2L2L requires a comparison with constant 111 and an argument that handles the absolute value. A second difficulty is that the classes may be infinite and unbounded, so expected suprema must be handled as extended-valued quantities, and the reduction to finite classes ("without loss of generality") must be justified. For Theorem 16 the extension of ggg must be checked to be Lipschitz across the boundaries of the tents in the Euclidean, not the sup, norm.

Formalization scope

The source is the published JMLR article (vol. 3, 2002, pp. 463–482), not the COLT 2001 version, whose numbering differs.

  • Complexities in [0,∞][0, \infty][0,∞]. G^n\hat G_nG^n​ and GnG_nGn​ are lower Lebesgue integrals of an ENNReal supremum against N(0,1)⊗nN(0,1)^{\otimes n}N(0,1)⊗n and μ⊗n\mu^{\otimes n}μ⊗n. An unbounded class has complexity +∞+\infty+∞; a real-valued supremum or Bochner integral would silently return 000 there and make the upper bounds false, so that encoding is ruled out. No finiteness or boundedness of any class is assumed.
  • A=Rm\mathcal A = \mathbb R^mA=Rm is EuclideanSpace ℝ (Fin m), so "Lipschitz" refers to the Euclidean distance as printed. Using Fin m → ℝ (the sup distance) would change the theorem.
  • {±1}\{\pm1\}{±1} is encoded as Z×\mathbb Z^\timesZ× coerced to R\mathbb RR.
  • Lemma 13: centred processes added. As printed the lemma is false: Xi≡5X_i \equiv 5Xi​≡5, Yi≡0Y_i \equiv 0Yi​≡0 satisfy the hypothesis. Both processes are assumed mean zero, as in Slepian's lemma. The two processes may live on different probability spaces; the index set is any finite nonempty type. The constant 222 is kept as printed.
  • Theorem 14: printed proof loose, statement kept. The constant 2L2L2L is kept as printed; the statement is true via the constant-111 comparison. The uniform-boundedness hypothesis on ϕ\phiϕ is kept as printed although the extended-valued formulation does not need it.
  • Theorem 16 and the extension: k≥1k \ge 1k≥1 added. For k=0k = 0k=0 the boolean function is a constant ±1\pm1±1, the right side is 000, and the left side is E2n∣∑igi∣>0\mathbb E\frac2n|\sum_i g_i| > 0En2​∣∑i​gi​∣>0; also the extension would have g(0)=±1g(0) = \pm1g(0)=±1.
  • Theorem 16: measurability guard. Each G^n(Fj)\hat G_n(F_j)G^n​(Fj​) is assumed almost-everywhere measurable as a function of the sample, so that the expectation of ∑jG^n(Fj)\sum_j \hat G_n(F_j)∑j​G^n​(Fj​) is the sum of the expectations. The paper does not discuss measurability.

A complete development needs Gaussian comparison for finite index sets (available on the platform), the reduction from infinite to finite classes for extended-valued suprema, and elementary Euclidean geometry of the cube. The first two are reusable for every Gaussian-average argument in learning theory. Proofs of any milestone, and of the constant-111 variant of Lemma 13 transported between probability spaces, are welcome.

Selected references

  • P. L. Bartlett, S. Mendelson, Rademacher and Gaussian Complexities: Risk Bounds and Structural Results, Journal of Machine Learning Research 3 (2002), 463–482. https://www.jmlr.org/papers/v3/bartlett02a.html
  • G. Pisier, The Volume of Convex Bodies and Banach Space Geometry, Cambridge University Press, 1989. https://doi.org/10.1017/CBO9780511662454
  • M. Ledoux, M. Talagrand, Probability in Banach Spaces: Isoperimetry and Processes, Springer, 1991. https://doi.org/10.1007/978-3-642-20212-4
8 thms2 active usersReviewed
AnalysisFunctional Analysis·Captain: mikedeng1

Theory of Reproducing Kernels IV: The Kernels of a Decreasing Sequence of Reproducing Kernel Classes Converge to the Kernel of the Limit ClassResearch Paper

Motivation

A reproducing kernel Hilbert space is a Hilbert space of functions on a set in which every point evaluation is continuous; the function K(x,y)K(x,y)K(x,y) that represents evaluation at yyy is its reproducing kernel. N. Aronszajn's Theory of Reproducing Kernels (Trans. Amer. Math. Soc. 68 (1950), 337–404, DOI 10.1090/S0002-9947-1950-0051437-7) gave the general theory of these spaces, which today underlies kernel methods in statistics and machine learning, Gaussian-process regression, and the Bergman and Szegő kernels of complex analysis.

Part I of the paper studies how kernels behave under the basic operations on classes of functions: sums, inclusions, products, restrictions, and limits. §9 treats limits. Its case A concerns a decreasing sequence of classes with increasing norms, defined on an increasing sequence of sets. The application in the paper's Part II is the computation of kernels of a domain by approximation from simpler domains: when a domain is exhausted by an increasing sequence of subdomains, the kernels of the subdomains converge to the kernel of the whole domain. This mission formalizes §9, Theorem I and the steps of its proof.

Setting

Let XXX be an arbitrary set and E1⊂E2⊂⋯E_1\subset E_2\subset\cdotsE1​⊂E2​⊂⋯ subsets with union E=E1+E2+⋯=XE = E_1+E_2+\cdots = XE=E1​+E2​+⋯=X. For each nnn let FnF_nFn​ be a complex Hilbert space of functions on EnE_nEn​, with norm ∥⋅∥n\|\cdot\|_n∥⋅∥n​, in which point evaluations are continuous; Kn(x,y)K_n(x,y)Kn​(x,y), for x,y∈Enx,y\in E_nx,y∈En​, is its reproducing kernel, characterized by Kn(⋅,y)∈FnK_n(\cdot,y)\in F_nKn​(⋅,y)∈Fn​ and

f(y)=(f,Kn(⋅,y))n(f∈Fn, y∈En),f(y) = (f, K_n(\cdot,y))_n \qquad (f\in F_n,\ y\in E_n),f(y)=(f,Kn​(⋅,y))n​(f∈Fn​, y∈En​),

with the scalar product (f,g)n(f,g)_n(f,g)n​ linear in fff. For fn∈Fnf_n\in F_nfn​∈Fn​ and m≤nm\le nm≤n, fnmf_{nm}fnm​ denotes the restriction of fnf_nfn​ to EmE_mEm​. The standing assumptions of §9 A (p. 362) are:

  1. E1⊂E2⊂⋯E_1\subset E_2\subset\cdotsE1​⊂E2​⊂⋯ and E=⋃nEnE = \bigcup_n E_nE=⋃n​En​;
  2. the classes decrease: fnm∈Fmf_{nm}\in F_mfnm​∈Fm​ for every fn∈Fnf_n\in F_nfn​∈Fn​ and m≤nm\le nm≤n;
  3. the norms increase: ∥fnm∥m≤∥fn∥n\|f_{nm}\|_m\le\|f_n\|_n∥fnm​∥m​≤∥fn​∥n​ for every fn∈Fnf_n\in F_nfn​∈Fn​ and m≤nm\le nm≤n;

together with the existence of every kernel KnK_nKn​. For two kernels on a set YYY, K1≪KK_1\ll KK1​≪K means that K−K1K-K_1K−K1​ is a positive matrix: ∑i,j(K−K1)(yi,yj) ξˉiξj≥0\sum_{i,j}(K-K_1)(y_i,y_j)\,\bar\xi_i\xi_j\ge 0∑i,j​(K−K1​)(yi​,yj​)ξˉ​i​ξj​≥0 for all finite families yi∈Yy_i\in Yyi​∈Y, ξi∈C\xi_i\in\mathbb Cξi​∈C. KnmK_{nm}Knm​ is the restriction of KnK_nKn​ to Em×EmE_m\times E_mEm​×Em​.

The limit class F0F_0F0​ is the set of functions f0f_0f0​ on EEE such that (1°) every restriction f0nf_{0n}f0n​ belongs to FnF_nFn​ and (2°) lim⁡n∥f0n∥n<∞\lim_n\|f_{0n}\|_n<\inftylimn​∥f0n​∥n​<∞.

Formalization targets

Goal: §9, Theorem I (pp. 362–363)

Under the standing assumptions there is K0:E×E→CK_0 : E\times E\to\mathbb CK0​:E×E→C such that, whenever x,y∈ENx,y\in E_Nx,y∈EN​,

lim⁡n→∞Kn(x,y)=K0(x,y),\lim_{n\to\infty}K_n(x,y)=K_0(x,y),n→∞lim​Kn​(x,y)=K0​(x,y),

and K0K_0K0​ is the reproducing kernel of F0F_0F0​ with the norm

∥f0∥0=lim⁡n→∞∥f0n∥n.\|f_0\|_0=\lim_{n\to\infty}\|f_{0n}\|_n .∥f0​∥0​=n→∞lim​∥f0n​∥n​.

Milestones (in the order the proof uses them)

  1. §9, Eq. (4): Knm≪KmK_{nm}\ll K_mKnm​≪Km​ for m<nm<nm<n.
  2. §9, proof of Theorem I, p. 363: for y∈Eky\in E_ky∈Ek​, {Km(y,y)}m≥k\{K_m(y,y)\}_{m\ge k}{Km​(y,y)}m≥k​ is a decreasing sequence of non-negative numbers.
  3. §9, Eq. (5): for y∈Eky\in E_ky∈Ek​, k≤m≤nk\le m\le nk≤m≤n, ∥Kmk(⋅,y)−Knk(⋅,y)∥k2≤Km(y,y)−Kn(y,y)\|K_{mk}(\cdot,y)-K_{nk}(\cdot,y)\|_k^2\le K_m(y,y)-K_n(y,y)∥Kmk​(⋅,y)−Knk​(⋅,y)∥k2​≤Km​(y,y)−Kn​(y,y).
  4. §9, Eq. (6): with K0K_0K0​ the pointwise limit, K0k(⋅,y)∈FkK_{0k}(\cdot,y)\in F_kK0k​(⋅,y)∈Fk​ and ∥Kmk(⋅,y)−K0k(⋅,y)∥k2≤Km(y,y)−K0(y,y)\|K_{mk}(\cdot,y)-K_{0k}(\cdot,y)\|_k^2\le K_m(y,y)-K_0(y,y)∥Kmk​(⋅,y)−K0k​(⋅,y)∥k2​≤Km​(y,y)−K0​(y,y).
  5. §9, Remark after Theorem I: under 1°, ∥f0n∥n\|f_{0n}\|_n∥f0n​∥n​ is non-decreasing, so its limit exists, possibly infinite.
  6. §9, Eq. (7): if F0F_0F0​ carries the limit norm, then (f0,g0)0=lim⁡n(f0n,g0n)n(f_0,g_0)_0=\lim_n(f_{0n},g_{0n})_n(f0​,g0​)0​=limn​(f0n​,g0n​)n​.

Significance

The result. Theorem I turns a monotone family of function spaces into a single space and identifies its kernel as the pointwise limit of the kernels. It reduces the computation of a kernel on a large set to kernels on an exhausting sequence of subsets, the method Aronszajn uses in Part II for Bergman-type kernels of plane domains. With En=EE_n=EEn​=E for all nnn (explicitly allowed on p. 362) it gives the limit of a decreasing sequence of kernels K1≫K2≫⋯K_1\gg K_2\gg\cdotsK1​≫K2​≫⋯ on one set as the kernel of the intersection class with the limit norm. The milestones (4)–(6) are quantitative: (5) bounds the distance between restricted kernel sections by the decrease of the diagonal values, which yields strong convergence of Km(⋅,y)K_m(\cdot,y)Km​(⋅,y) in every FkF_kFk​.

Formalizing it. The theorem is classical and proved in the paper; to our knowledge no machine-checked proof exists. Mathlib has the RKHS class, the operator-valued kernel, the positive semidefiniteness of kernels and the Moore–Aronszajn construction RKHS.OfKernel, but nothing about restrictions of an RKHS to a subset, the order ≪\ll≪ between kernels, or limits of sequences of reproducing kernel spaces. This mission produces those statements on Mathlib's RKHS vocabulary over C\mathbb CC, with kernels on varying domains.

Difficulty

The kernels KnK_nKn​ live on different sets En×EnE_n\times E_nEn​×En​, so convergence is not convergence of a sequence of functions on one set: a pair x,yx,yx,y enters the sequence only from the first ENE_NEN​ containing both. The identification of the limit class needs three separate facts: that F0F_0F0​ with the limit norm is a Hilbert space (the limit of norms must be shown to come from a scalar product, and completeness requires passing to the limit in two indices), that K0(⋅,y)∈F0K_0(\cdot,y)\in F_0K0​(⋅,y)∈F0​, and that K0K_0K0​ reproduces. The natural first idea, to embed all FnF_nFn​ in one space and take an intersection, fails: the FnF_nFn​ are spaces of functions on different sets, and their norms differ, so there is no common ambient Hilbert space; the comparison goes only through restriction and the inequalities (3). Eq. (4) itself uses §7, Theorem II (a contractively included Hilbert subclass has a dominated kernel) and the restriction theorem of §5, neither of which is in Mathlib.

Formalization scope

  • Scalars and spaces. Complex scalars throughout (Aronszajn works with complex Hilbert spaces from §1 on). Each FnF_nFn​ is a type H n with [InnerProductSpace ℂ (H n)] [CompleteSpace (H n)] [RKHS ℂ (H n) (E n) ℂ], a space of functions on the subtype E n; the set EEE is a type X with no topology, measure or nonemptiness assumption.
  • Kernel. The scalar kernel kernelFn H x y is Mathlib's RKHS.kernel H x y 1. Mathlib's inner product is conjugate-linear in the first slot, so Aronszajn's (f,g)(f,g)(f,g) is ⟪g, f⟫_ℂ.
  • Standing assumptions. (1)–(3) are the structure IsDecreasingRKSequence; every statement takes it as a hypothesis. Restriction is pointwise agreement on EmE_mEm​. Indexing starts at 000.
  • Order. K1≪KK_1\ll KK1​≪K is KernelLE K₁ K := (Matrix.of K - Matrix.of K₁).PosSemidef, with Mathlib's positive semidefiniteness over an arbitrary index type (finitely supported vectors).
  • Comparisons of kernel values (Km(y,y)≥0K_m(y,y)\ge 0Km​(y,y)≥0, the right-hand sides of (5), (6)) are in Mathlib's ComplexOrder, which also asserts that these values are real.
  • Convergence of kernels is stated only where the terms are defined: for x,y∈ENx,y\in E_Nx,y∈EN​, the sequence j↦KN+j(x,y)j\mapsto K_{N+j}(x,y)j↦KN+j​(x,y) converges to K0(x,y)K_0(x,y)K0​(x,y). Kernels are never extended by 000 outside EnE_nEn​.
  • Condition 2° is convergence of ∥f0n∥n\|f_{0n}\|_n∥f0n​∥n​ to a real number, not a supremum, and the norm of F0F_0F0​ is stated as a limit (Tendsto).
  • The goal asserts (a) the convergence, (b) the existence of an RKHS on XXX with kernel K0K_0K0​, and (c) that every RKHS on XXX with kernel K0K_0K0​ has exactly the functions of F0F_0F0​ as its elements and the limit norm. A formalization that defines F0F_0F0​ as RKHS.OfKernel K₀ and then asserts that its kernel is K0K_0K0​ would be a tautology (RKHS.kernel_ofKernel); the goal instead characterizes the space by its functions and norm, as the paper does.
  • Eq. (7) is stated for an inner product space of functions whose norm is assumed to be the limit norm; the paper's derivation that the limit norm is a quadratic form is the content of the goal.
  • Non-vacuity. The constant sequence En=EE_n=EEn​=E, Fn=FF_n=FFn​=F satisfies the standing assumptions (checked in Lean), and the one-point example Fn=CF_n=\mathbb CFn​=C with norms cn∣f∣c_n|f|cn​∣f∣, cnc_ncn​ increasing, satisfies them with Kn=cn−2K_n=c_n^{-2}Kn​=cn−2​.

Needed infrastructure, reusable beyond this mission: restriction of an RKHS to a subset (§5), the dominated-kernel theorem for contractive inclusions (§7, Theorem II), and the passage from a convergent sequence of norms to a convergent sequence of scalar products. Proofs of any milestone, and of these general facts as separate lemmas, are welcome.

Selected references

  • N. Aronszajn, Theory of Reproducing Kernels, Trans. Amer. Math. Soc. 68 (1950), no. 3, 337–404. https://doi.org/10.1090/S0002-9947-1950-0051437-7
  • E. H. Moore, General Analysis, Part I, Memoirs of the American Philosophical Society 1 (1935). (Positive matrices.)
  • Mathlib, Mathlib/Analysis/InnerProductSpace/Reproducing.lean (the RKHS class, RKHS.kernel, RKHS.OfKernel). https://github.com/leanprover-community/mathlib4
11 thms2 active usersReviewed
AnalysisFunctional Analysis·Captain: mikedeng1

Theory of Reproducing Kernels V: A Hermitian Kernel Represents a Bounded Symmetric Operator with Bounds m and M iff mK ≪ Λ ≪ MKResearch Paper

Motivation

Reproducing kernel Hilbert spaces are the function spaces of kernel methods in statistics and machine learning (Gaussian-process regression, support vector machines, kernel mean embeddings), of the Bergman and Szegő spaces of complex analysis, and of the theory of positive-definite functions. In all of these, bounded operators on the space (covariance operators, integral operators, projections onto subspaces, multiplication operators) are handled through functions of two points rather than through abstract operators. N. Aronszajn's Theory of Reproducing Kernels (Trans. Amer. Math. Soc. 68 (1950), 337–404) gives, in its §11, the dictionary between bounded operators on a space with a reproducing kernel and their kernels, and characterizes the kernels of bounded symmetric operators with prescribed bounds. Aronszajn credits the ideas of the section to E. H. Moore.

Setting

Let EEE be an arbitrary set and let FFF be a class of complex-valued functions on EEE that forms a complex Hilbert space with scalar product (f,g)(f, g)(f,g), linear in fff and conjugate-linear in ggg. A reproducing kernel of FFF is a function K:E×E→CK : E \times E \to \mathbb{C}K:E×E→C such that, for every y∈Ey \in Ey∈E, the function K(⋅,y)K(\cdot, y)K(⋅,y) belongs to FFF and

f(y)=(f,K(⋅,y))for every f∈F.f(y) = (f, K(\cdot, y)) \qquad \text{for every } f \in F.f(y)=(f,K(⋅,y))for every f∈F.

Such a kernel exists exactly when every point evaluation f↦f(y)f \mapsto f(y)f↦f(y) is continuous.

For a bounded linear operator LLL on FFF, with adjoint L∗L^*L∗ defined by (Lf,g)=(f,L∗g)(Lf, g) = (f, L^* g)(Lf,g)=(f,L∗g), the kernel of LLL is

Λ(x,y)=Lx∗K(x,y),\Lambda(x, y) = L^*_x K(x, y),Λ(x,y)=Lx∗​K(x,y),

the value at xxx of the element L∗(K(⋅,y))L^*(K(\cdot, y))L∗(K(⋅,y)) of FFF. By the reproducing property, Lf(y)=(f,Λ(⋅,y))Lf(y) = (f, \Lambda(\cdot, y))Lf(y)=(f,Λ(⋅,y)) for every f∈Ff \in Ff∈F and y∈Ey \in Ey∈E, so LLL is determined by Λ\LambdaΛ.

A function P:E×E→CP : E \times E \to \mathbb{C}P:E×E→C is a positive matrix if ∑i,jξi‾ P(yi,yj) ξj≥0\sum_{i,j} \overline{\xi_i}\, P(y_i, y_j)\, \xi_j \ge 0∑i,j​ξi​​P(yi​,yj​)ξj​≥0 for every finite family of points yi∈Ey_i \in Eyi​∈E and complex numbers ξi\xi_iξi​. For two arbitrary functions Λ1,Λ2\Lambda_1, \Lambda_2Λ1​,Λ2​ on E×EE \times EE×E, one writes Λ1≪Λ2\Lambda_1 \ll \Lambda_2Λ1​≪Λ2​ if Λ2−Λ1\Lambda_2 - \Lambda_1Λ2​−Λ1​ is a positive matrix. A bounded operator LLL is symmetric if L=L∗L = L^*L=L∗, and positive if (Lf,f)≥0(Lf, f) \ge 0(Lf,f)≥0 for every fff. A symmetric LLL has lower bound ≥m\ge m≥m and upper bound ≤M\le M≤M if

m (f,f)≤(Lf,f)≤M (f,f)for every f∈F.m\,(f, f) \le (Lf, f) \le M\,(f, f) \qquad \text{for every } f \in F.m(f,f)≤(Lf,f)≤M(f,f)for every f∈F.

A kernel Λ\LambdaΛ is hermitian symmetric if Λ(x,y)=Λ(y,x)‾\Lambda(x, y) = \overline{\Lambda(y, x)}Λ(x,y)=Λ(y,x)​.

Formalization targets

Goal: §11, Theorem I (p. 373)

For an arbitrary hermitian symmetric function Λ:E×E→C\Lambda : E \times E \to \mathbb{C}Λ:E×E→C and real numbers m,Mm, Mm,M:

∃ L bounded, symmetric, with Λ=Lx∗K(x,y) and m(f,f)≤(Lf,f)≤M(f,f)  ∀f⟺mK≪Λ≪MK.\exists\, L \text{ bounded, symmetric, with } \Lambda = L^*_x K(x,y) \text{ and } m(f,f) \le (Lf,f) \le M(f,f)\ \ \forall f \quad\Longleftrightarrow\quad mK \ll \Lambda \ll MK .∃L bounded, symmetric, with Λ=Lx∗​K(x,y) and m(f,f)≤(Lf,f)≤M(f,f)  ∀f⟺mK≪Λ≪MK.

The function Λ\LambdaΛ is not assumed to have Λ(⋅,y)∈F\Lambda(\cdot, y) \in FΛ(⋅,y)∈F; that membership is part of what the condition yields.

Milestones

  1. §11, (3): the kernel of the adjoint, Λ∗(y,z)=Λ(z,y)‾\Lambda^*(y, z) = \overline{\Lambda(z, y)}Λ∗(y,z)=Λ(z,y)​.
  2. §11, (6): LLL is symmetric if and only if Λ\LambdaΛ is hermitian symmetric.
  3. §11, (7): LLL is positive if and only if Λ\LambdaΛ is a positive matrix.
  4. §11, (4): the kernel of a composition, Λ(y,z)=(Λ1(x,z),Λ2(y,x)‾)x\Lambda(y, z) = (\Lambda_1(x, z), \overline{\Lambda_2(y, x)})_xΛ(y,z)=(Λ1​(x,z),Λ2​(y,x)​)x​ for L=L1L2L = L_1 L_2L=L1​L2​.
  5. §11, Theorem II: if Lnu→LuL_n u \to L uLn​u→Lu weakly for every uuu, then Λn→Λ\Lambda_n \to \LambdaΛn​→Λ pointwise; if ∥Ln−L∥→0\|L_n - L\| \to 0∥Ln​−L∥→0, then Λn→Λ\Lambda_n \to \LambdaΛn​→Λ uniformly on every set of couples (x,y)(x, y)(x,y) on which K(x,x)K(x, x)K(x,x) and K(y,y)K(y, y)K(y,y) are uniformly bounded.
  6. §11, Theorem III, first sentence: for complete orthonormal systems {gm′}\{g'_m\}{gm′​}, {gn′′}\{g''_n\}{gn′′​} and αmn=(gn′′,Lgm′)\alpha_{mn} = (g''_n, L g'_m)αmn​=(gn′′​,Lgm′​),
Λ(x,y)=lim⁡p,q→∞∑m=1p∑n=1qαmn gm′(x) gn′′(y)‾.\Lambda(x, y) = \lim_{p, q \to \infty} \sum_{m=1}^{p} \sum_{n=1}^{q} \alpha_{mn}\, g'_m(x)\, \overline{g''_n(y)} .Λ(x,y)=p,q→∞lim​m=1∑p​n=1∑q​αmn​gm′​(x)gn′′​(y)​.

Significance

Theorem I identifies, by finite quadratic-form inequalities alone, which functions of two points are kernels of bounded symmetric operators and with which spectral bounds. It reduces statements about operators (boundedness, positivity, operator inequalities mI≤L≤MImI \le L \le MImI≤L≤MI) to statements about finitely many evaluations of kernels, the form in which they are checked in practice, for instance when a covariance or integral operator is shown to be bounded and positive from its kernel. The milestones make the correspondence L↦ΛL \mapsto \LambdaL↦Λ a usable calculus: adjoints become conjugate transposes, composition becomes a scalar product in the middle variable, and limits of operators become limits of kernels.

All of these results are proved in the paper. None is formalized: Mathlib has reproducing kernel Hilbert spaces (RKHS), adjoints, positive operators and positive semidefinite matrices over arbitrary index types, but no kernel of an operator and none of the statements above. The mission produces machine-checked proofs of the §11 dictionary and of Theorem I.

Difficulty

The necessity half of Theorem I and milestones (3), (6), (4) follow from the reproducing property. The sufficiency half is where the work is: Λ\LambdaΛ is an arbitrary function, and the hypothesis mK≪Λ≪MKmK \ll \Lambda \ll MKmK≪Λ≪MK is only about finite families of points. One has to produce an operator on all of FFF. The obvious attempt, defining LLL on the dense span of the functions K(⋅,y)K(\cdot, y)K(⋅,y) by the kernel and extending by continuity, needs the bound ∣( Lf,g)∣≤C∥f∥∥g∥|(\,L f, g)| \le C\|f\|\|g\|∣(Lf,g)∣≤C∥f∥∥g∥ on that span, which does not follow directly from the two one-sided inequalities on the diagonal forms. The positivity of Λ−mK\Lambda - mKΛ−mK and MK−ΛMK - \LambdaMK−Λ does not by itself give the membership Λ(⋅,y)∈F\Lambda(\cdot, y) \in FΛ(⋅,y)∈F, which the definition of the kernel of an operator requires. In milestone (7), positivity of an operator is a statement about all of FFF, while positivity of the kernel only sees finite combinations of kernel functions; the passage between them uses density of these combinations.

Formalization scope

  • The space is Mathlib's RKHS ℂ H X ℂ: a complex Hilbert space H whose elements are functions X → ℂ on an arbitrary type X (no topology, no measure, not assumed nonempty), with continuous evaluations. The scalar kernel is the series' shared definition AronszajnRK.Sum.kernelFn H x y := RKHS.kernel H x y 1; the function K(⋅,y)K(\cdot, y)K(⋅,y) is the element RKHS.kerFun H y 1.
  • The kernel of L : H →L[ℂ] H is opKernel L x y := (adjoint L) (kerFun H y 1) x. Mathlib's ⟪u, v⟫_ℂ is conjugate-linear in u, so the paper's (f,g)(f, g)(f,g) is ⟪g, f⟫_ℂ, and every formula with a scalar product or a bar has been rewritten in that order. On a one-point EEE with F=CF = \mathbb{C}F=C, K=1K = 1K=1 and L=cIL = cIL=cI, the kernel is cˉ\bar ccˉ.
  • Positive matrices and ≪\ll≪ are Matrix.PosSemidef of Matrix.of Λ over the index type X (finitely supported test vectors, ComplexOrder on ℂ); no finiteness of X is assumed.
  • Symmetric is IsSelfAdjoint L. "Positive" in (7) is ∀ f, 0 ≤ ⟪f, L f⟫_ℂ in ComplexOrder (real and nonnegative), without assuming self-adjointness, as in the paper. The bounds in Theorem I are bounds of the quadratic form, m‖f‖² ≤ Re⟪f, L f⟫ ≤ M‖f‖², not of the operator norm. The paper does not assume m≤Mm \le Mm≤M and neither does the statement: for m>Mm > Mm>M both sides hold exactly when F={0}F = \{0\}F={0} and Λ=0\Lambda = 0Λ=0.
  • In (4) the statement asserts that the functions x↦Λ1(x,z)x \mapsto \Lambda_1(x, z)x↦Λ1​(x,z) and x↦Λ2(y,x)‾x \mapsto \overline{\Lambda_2(y, x)}x↦Λ2​(y,x)​ are elements of H, and that the kernel of L₁ ∘L L₂ is their scalar product.
  • Weak convergence in Theorem II is ⟪v, Lₙ u⟫ → ⟪v, L u⟫ for all u, v; uniform convergence is ‖Lₙ − L‖ → 0 in operator norm.
  • In Theorem III the orthonormal systems are HilbertBasis with arbitrary index types, and the double limit is taken along growing finite sets of indices in both variables. For systems indexed by N\mathbb{N}N this contains the paper's lim⁡p,q\lim_{p,q}limp,q​ over {1..p}×{1..q}\{1..p\}\times\{1..q\}{1..p}×{1..q}; the general form also covers finite-dimensional spaces. The second sentence of Theorem III (kernels in F⊗F‾F \otimes \overline{F}F⊗F correspond to operators of finite norm) needs the direct product F⊗F‾F \otimes \overline FF⊗F and is not stated.
  • A trivializing formalization is ruled out: Theorem I quantifies over every hermitian function Λ\LambdaΛ, not over functions already known to be kernels of operators, and the right-hand side is the finite-matrix condition, not a statement about an operator built from Λ\LambdaΛ.
  • Not stated: the decomposition (8)–(9) into hermitian parts and the remark on general bounded operators (p. 374), and formula (5). Available substrate: RKHS, RKHS.kerFun_inner, RKHS.kerFun_dense, RKHS.posSemidef_kernel, ContinuousLinearMap.adjoint, ContinuousLinearMap.IsPositive and isPositive_iff_complex, HilbertBasis, Matrix.PosSemidef. Contributions welcome: a general lemma that a function Λ\LambdaΛ with 0≪Λ≪K0 \ll \Lambda \ll K0≪Λ≪K is the kernel of an operator 0≤L≤I0 \le L \le I0≤L≤I, and the inclusion theorem K1≪K⇒F1⊂FK_1 \ll K \Rightarrow F_1 \subset FK1​≪K⇒F1​⊂F, both reusable beyond this mission.

Selected references

  • N. Aronszajn, Theory of Reproducing Kernels, Trans. Amer. Math. Soc. 68 (1950), no. 3, 337–404. https://doi.org/10.1090/S0002-9947-1950-0051437-7
  • E. H. Moore, General Analysis, Part II, Mem. Amer. Philos. Soc. 1 (1939).
  • V. I. Paulsen and M. Raghupathi, An Introduction to the Theory of Reproducing Kernel Hilbert Spaces, Cambridge University Press, 2016. https://doi.org/10.1017/CBO9781316219232
10 thms2 active usersReviewed
Convex OptimizationOptimizationProbability+1·Captain: mikedeng1

Variance-based Regularization with Convex Objectives IV: Fast Rates for Approximate Robust Minimizers under a Growth ConditionResearch Paper

Motivation

In stochastic optimization and statistical learning one chooses a parameter θ\thetaθ from a set Θ⊆Rd\Theta\subseteq\mathbb R^dΘ⊆Rd to make the risk R(θ)=EP[ℓ(θ;X)]R(\theta)=\mathbb E_P[\ell(\theta;X)]R(θ)=EP​[ℓ(θ;X)] small, having seen only a sample X1,…,XnX_1,\dots,X_nX1​,…,Xn​ from PPP. Generalization bounds suggest trading empirical risk against its standard deviation, but the variance-penalized objective is non-convex even for convex losses. Duchi and Namkoong (arXiv:1610.02581v3) replace it by the robustly regularized risk, the worst-case expected loss over a χ2\chi^2χ2-divergence ball around the empirical distribution. This objective is convex whenever ℓ\ellℓ is, and it agrees with the variance-penalized objective up to a small error.

When the risk has curvature near its minimizers, empirical risk minimization attains rates faster than 1/n1/\sqrt n1/n​ (Bartlett, Bousquet and Mendelson 2005; Shapiro, Dentcheva and Ruszczyński 2009). Section 4.1 of the paper asks whether minimizers of the robust risk, which carry an extra variance-dependent penalty of order ρ/n\sqrt{\rho/n}ρ/n​, keep these fast rates. Its Theorem 5 answers yes, and does so for approximate minimizers, which is what iterative solvers return.

Setting

A loss ℓ:Rd×X→R\ell:\mathbb R^d\times\mathcal X\to\mathbb Rℓ:Rd×X→R is fixed, with ℓ(⋅;x)\ell(\cdot;x)ℓ(⋅;x) convex and LLL-Lipschitz on a convex set Θ\ThetaΘ for every xxx, and ℓ(θ;⋅)\ell(\theta;\cdot)ℓ(θ;⋅) integrable. The risk is R(θ)=EP[ℓ(θ;X)]R(\theta)=\mathbb E_P[\ell(\theta;X)]R(θ)=EP​[ℓ(θ;X)].

For a radius ρ≥0\rho\ge0ρ≥0, the χ2\chi^2χ2 ball around the empirical distribution P^n\widehat P_nPn​ is the set of weight vectors

Pn={p∈R+n:12∥np−1∥22≤ρ, ⟨1,p⟩=1},\mathcal P_n=\Big\{p\in\mathbb R^n_+:\tfrac12\|np-\mathbf 1\|_2^2\le\rho,\ \langle\mathbf 1,p\rangle=1\Big\},Pn​={p∈R+n​:21​∥np−1∥22​≤ρ, ⟨1,p⟩=1},

and the robust risk is Rn(θ,Pn)=sup⁡p∈Pn∑ipi ℓ(θ;Xi)R_n(\theta,\mathcal P_n)=\sup_{p\in\mathcal P_n}\sum_i p_i\,\ell(\theta;X_i)Rn​(θ,Pn​)=supp∈Pn​​∑i​pi​ℓ(θ;Xi​).

For ϵ≥0\epsilon\ge0ϵ≥0 the ϵ\epsilonϵ-suboptimal sets of the risk and of the robust risk are

S⋆ϵ={θ∈Θ:R(θ)≤inf⁡ΘR+ϵ},S^⋆ϵ={θ∈Θ:Rn(θ,Pn)≤inf⁡ΘRn(⋅,Pn)+ϵ},S_\star^\epsilon=\{\theta\in\Theta:R(\theta)\le\inf_\Theta R+\epsilon\},\qquad\widehat S_\star^\epsilon=\{\theta\in\Theta:R_n(\theta,\mathcal P_n)\le\inf_\Theta R_n(\cdot,\mathcal P_n)+\epsilon\},S⋆ϵ​={θ∈Θ:R(θ)≤Θinf​R+ϵ},S⋆ϵ​={θ∈Θ:Rn​(θ,Pn​)≤Θinf​Rn​(⋅,Pn​)+ϵ},

with S⋆=S⋆0S_\star=S_\star^0S⋆​=S⋆0​ the solution set and πS⋆\pi_{S_\star}πS⋆​​ the Euclidean projection onto it. The risk satisfies a growth condition of order γ>1\gamma>1γ>1 if, for some λ>0\lambda>0λ>0 and r>0r>0r>0,

R(θ)−inf⁡ΘR ≥ λ dist(θ,S⋆)γwhenever dist(θ,S⋆)≤r.(26)R(\theta)-\inf_\Theta R\ \ge\ \lambda\,\mathrm{dist}(\theta,S_\star)^\gamma\quad\text{whenever }\mathrm{dist}(\theta,S_\star)\le r.\tag{26}R(θ)−Θinf​R ≥ λdist(θ,S⋆​)γwhenever dist(θ,S⋆​)≤r.(26)

The complexity of the problem enters through the localized class {x↦ℓ(θ;x)−ℓ(πS⋆(θ);x):θ∈A}\{x\mapsto\ell(\theta;x)-\ell(\pi_{S_\star}(\theta);x):\theta\in A\}{x↦ℓ(θ;x)−ℓ(πS⋆​​(θ);x):θ∈A} and its empirical Rademacher complexity Rn(A)=Eε[sup⁡θ∈A1n∑iεi(ℓ(θ;Xi)−ℓ(πS⋆(θ);Xi))]\mathfrak R_n(A)=\mathbb E_\varepsilon\big[\sup_{\theta\in A}\frac1n\sum_i\varepsilon_i(\ell(\theta;X_i)-\ell(\pi_{S_\star}(\theta);X_i))\big]Rn​(A)=Eε​[supθ∈A​n1​∑i​εi​(ℓ(θ;Xi​)−ℓ(πS⋆​​(θ);Xi​))], with independent uniform signs εi∈{±1}\varepsilon_i\in\{\pm1\}εi​∈{±1}.

Formalization targets

Goal: Theorem 5 (p. 19)

For t>0t>0t>0, ρ≥0\rho\ge0ρ≥0, and 0<ϵ≤12λrγ0<\epsilon\le\frac12\lambda r^\gamma0<ϵ≤21​λrγ satisfying

ϵ≥(28γLγλ)1γ−1(ρn)γ2(γ−1)andϵ2≥2 E[Rn(S⋆2ϵ)]+L(2ϵλ)1γ2tn,(27)\epsilon\ge\Big(2\frac{8^\gamma L^\gamma}{\lambda}\Big)^{\frac1{\gamma-1}}\Big(\frac\rho n\Big)^{\frac\gamma{2(\gamma-1)}}\quad\text{and}\quad\frac\epsilon2\ge2\,\mathbb E[\mathfrak R_n(S_\star^{2\epsilon})]+L\Big(\frac{2\epsilon}\lambda\Big)^{\frac1\gamma}\sqrt{\frac{2t}n},\tag{27}ϵ≥(2λ8γLγ​)γ−11​(nρ​)2(γ−1)γ​and2ϵ​≥2E[Rn​(S⋆2ϵ​)]+L(λ2ϵ​)γ1​n2t​​,(27) P(S^⋆ϵ⊂S⋆2ϵ) ≥ 1−e−t.\mathbb P\big(\widehat S_\star^\epsilon\subset S_\star^{2\epsilon}\big)\ \ge\ 1-e^{-t}.P(S⋆ϵ​⊂S⋆2ϵ​) ≥ 1−e−t.

Milestones, in attack order

  1. Localization (p. 44). Under (26), S⋆2ϵS_\star^{2\epsilon}S⋆2ϵ​ lies in {θ∈Θ:dist(θ,S⋆)≤(2ϵ/λ)1/γ}\{\theta\in\Theta:\mathrm{dist}(\theta,S_\star)\le(2\epsilon/\lambda)^{1/\gamma}\}{θ∈Θ:dist(θ,S⋆​)≤(2ϵ/λ)1/γ}.
  2. Theorem 1, upper half of (10) (p. 7). sup⁡p∈Pn⟨p,z⟩−zˉ≤2ρsn2/n\sup_{p\in\mathcal P_n}\langle p,z\rangle-\bar z\le\sqrt{2\rho s_n^2/n}supp∈Pn​​⟨p,z⟩−zˉ≤2ρsn2​/n​ for every z∈Rnz\in\mathbb R^nz∈Rn.
  3. Claim E.1 (p. 44). If S^⋆ϵ⊄S⋆2ϵ\widehat S_\star^\epsilon\not\subset S_\star^{2\epsilon}S⋆ϵ​⊂S⋆2ϵ​, the localized deviation Δn\Delta_nΔn​ plus a variance term reaches ϵ\epsilonϵ somewhere on S⋆2ϵS_\star^{2\epsilon}S⋆2ϵ​.
  4. Display (43) (p. 45). P(S^⋆ϵ⊄S⋆2ϵ)≤P(sup⁡S⋆2ϵΔn≥ϵ/2)\mathbb P(\widehat S_\star^\epsilon\not\subset S_\star^{2\epsilon})\le\mathbb P(\sup_{S_\star^{2\epsilon}}\Delta_n\ge\epsilon/2)P(S⋆ϵ​⊂S⋆2ϵ​)≤P(supS⋆2ϵ​​Δn​≥ϵ/2).
  5. Concentration (p. 45). P(sup⁡S⋆2ϵΔn≥2E[Rn(S⋆2ϵ)]+u)≤exp⁡(−nu22L2(λ2ϵ)2/γ)\mathbb P(\sup_{S_\star^{2\epsilon}}\Delta_n\ge2\mathbb E[\mathfrak R_n(S_\star^{2\epsilon})]+u)\le\exp(-\frac{nu^2}{2L^2}(\frac\lambda{2\epsilon})^{2/\gamma})P(supS⋆2ϵ​​Δn​≥2E[Rn​(S⋆2ϵ​)]+u)≤exp(−2L2nu2​(2ϵλ​)2/γ).

Significance

The theorem says that the variance penalty implicit in the robust objective does not cost the fast rates available under curvature. The ρ\rhoρ-dependent condition in (27) is of order (ρ/n)γ/(2(γ−1))(\rho/n)^{\gamma/(2(\gamma-1))}(ρ/n)γ/(2(γ−1)), which for quadratic growth (γ=2\gamma=2γ=2) is ρ/n\rho/nρ/n, the same order as the localized complexity term in typical parametric problems. Corollary 4.1 of the paper derives explicit rates of order dnlog⁡nd+tn+ρn\frac dn\log\frac nd+\frac tn+\frac\rho nnd​logdn​+nt​+nρ​ from it for a unique minimizer. The result applies to ϵ\epsilonϵ-approximate minimizers, so it covers the output of the stochastic-gradient methods used to solve the robust problem.

The result is proved in the paper (Appendix E). None of it is formalized: no statement about growth conditions, localized deviations of a robust objective, or fast rates for robust minimizers is on Prove2Me. A formal proof would check the printed constants, settle the boundary case ϵ=0\epsilon=0ϵ=0 (see below), and produce a localization lemma and a reduction from approximate robust minimizers to empirical processes that apply to other estimators.

Difficulty

The obvious argument fails at two points. First, a uniform deviation bound over all of Θ\ThetaΘ gives only the 1/n1/\sqrt n1/n​ rate: the speed-up comes from localizing to S⋆2ϵS_\star^{2\epsilon}S⋆2ϵ​, which requires transferring the growth condition, assumed only within distance rrr of S⋆S_\starS⋆​, to every 2ϵ2\epsilon2ϵ-suboptimal point by convexity. Second, the robust risk is not an empirical average, so standard comparisons between empirical and population minimizers do not apply. Claim E.1 handles this by moving along the segment from a bad approximate minimizer to its projection, which needs the projection to be preserved along that segment (a normal-cone property of πS⋆\pi_{S_\star}πS⋆​​) and the risk to be continuous there. The robust–empirical gap is then controlled by the variance expansion of Theorem 1. The concentration step needs a bounded-differences inequality for a supremum over an uncountable class, together with symmetrization; neither is in Mathlib in this form.

Formalization scope

Parameters live in EuclideanSpace ℝ (Fin d), so norms, distances and projections are Euclidean. The sample is the coordinate process of the product measure P⊗nP^{\otimes n}P⊗n on Fin n → X, n≥1n\ge1n≥1, and probabilities are measures of sample sets (the outer measure for a set that is not measurable). The χ2\chi^2χ2 ball is the weight-vector form (8). The suboptimal sets are written without infima (R(θ)≤R(θ′)+ϵR(\theta)\le R(\theta')+\epsilonR(θ)≤R(θ′)+ϵ for all θ′∈Θ\theta'\in\Thetaθ′∈Θ). Each supremum "sup⁡≥c\sup\ge csup≥c" is written as "for every δ>0\delta>0δ>0 some θ\thetaθ reaches c−δc-\deltac−δ", so no statement relies on the default value of a real supremum. The Rademacher complexity is the published UnderstandingML.rademacher, and its expectation over the sample is assumed integrable, so that it is the true expectation and not the default value 000 of a Bochner integral. Lipschitz continuity is required on Θ\ThetaΘ, as printed.

Corrections and presuppositions:

  • ϵ>0\epsilon>0ϵ>0. The paper prints 0≤ϵ0\le\epsilon0≤ϵ. At ϵ=0\epsilon=0ϵ=0, ρ=0\rho=0ρ=0, both conditions of (27) hold, yet for ℓ(θ;x)=12(θ−x)2\ell(\theta;x)=\frac12(\theta-x)^2ℓ(θ;x)=21​(θ−x)2 on Θ=[−1,1]\Theta=[-1,1]Θ=[−1,1] with XXX uniform on [−12,12][-\frac12,\frac12][−21​,21​] the robust minimizer is the sample mean, which is almost surely not in S⋆={0}S_\star=\{0\}S⋆​={0}. The proof divides by ϵ\epsilonϵ (p. 45). The goal is stated for ϵ>0\epsilon>0ϵ>0.
  • S⋆S_\starS⋆​ nonempty and closed are assumed. The projection πS⋆\pi_{S_\star}πS⋆​​ presupposes them, and Appendix E calls S⋆S_\starS⋆​ closed.
  • Only the upper half of Theorem 1's (10) is stated; it needs no boundedness of the values.

The constant (2⋅8γLγ/λ)1/(γ−1)\big(2\cdot8^\gamma L^\gamma/\lambda\big)^{1/(\gamma-1)}(2⋅8γLγ/λ)1/(γ−1) is the printed one; the proof uses a smaller one, which the printed condition implies. The hypotheses ϵ>0\epsilon>0ϵ>0, γ>1\gamma>1γ>1 and λ>0\lambda>0λ>0 make every power well defined. A formalization that assumed (26) vacuously, took ϵ=0\epsilon=0ϵ=0, or let the Rademacher term be a non-integrable Bochner integral would trivialize the goal; the statements rule these out.

Infrastructure: Euclidean projection onto closed convex sets and its normal-cone characterization (partly in Mathlib), convexity of integral functionals, McDiarmid's bounded-differences inequality, and symmetrization for suprema of empirical processes. The concentration tools and the localization lemma can be reused beyond this mission. Contributions toward McDiarmid's inequality and symmetrization are especially welcome.

Selected references

  • J. C. Duchi and H. Namkoong, Variance-based regularization with convex objectives, arXiv:1610.02581v3, 2017. https://arxiv.org/abs/1610.02581
  • P. L. Bartlett, O. Bousquet and S. Mendelson, Local Rademacher complexities, Annals of Statistics 33(4), 2005. https://doi.org/10.1214/009053605000000282
  • S. Boucheron, G. Lugosi and P. Massart, Concentration Inequalities: A Nonasymptotic Theory of Independence, Oxford University Press, 2013. https://doi.org/10.1093/acprof:oso/9780199535255.001.0001
  • A. Shapiro, D. Dentcheva and A. Ruszczyński, Lectures on Stochastic Programming: Modeling and Theory, SIAM, 2009. https://doi.org/10.1137/1.9780898718751
  • A. Maurer and M. Pontil, Empirical Bernstein bounds and sample variance penalization, COLT, 2009. https://arxiv.org/abs/0907.3740
12 thms2 active usersReviewed
ProbabilityStatistics·Captain: mikedeng1

Certified Adversarial Robustness via Randomized Smoothing 2: The Certified ℓ2 Radius Cannot Be EnlargedResearch Paper

Motivation

Neural-network classifiers can be made to change their output by perturbations of the input that are imperceptible to a person. A certified defense is a classifier together with a proof that its prediction at a point xxx does not change for any perturbation δ\deltaδ in a stated set, typically an ℓ2\ell_2ℓ2​ ball ∥δ∥2<R\|\delta\|_2<R∥δ∥2​<R. Randomized smoothing turns an arbitrary base classifier into one with such a certificate by classifying Gaussian-noised copies of the input and returning the most likely class. Cohen, Rosenfeld and Kolter (arXiv:1902.02918v2, ICML 2019) gave the certified radius R=σ2(Φ−1(pA‾)−Φ−1(pB‾))R=\frac{\sigma}{2}\big(\Phi^{-1}(\underline{p_A})-\Phi^{-1}(\overline{p_B})\big)R=2σ​(Φ−1(pA​​)−Φ−1(pB​​)) (their Theorem 1) and showed, in their Theorem 2, that this radius cannot be enlarged when only the two class-probability bounds are known about the base classifier. This mission formalizes Theorem 2. Theorem 1 is the subject of the companion mission of this series.

Earlier certificates for the same smoothed classifier, by Lecuyer et al. (2019) via differential privacy and Li et al. (2018) via Rényi divergence, gave smaller radii. Theorem 2 shows that no further analysis that uses only the class-probability bounds can improve on Theorem 1.

Setting

Inputs live in Rd\mathbb R^dRd with the Euclidean norm ∥⋅∥2\|\cdot\|_2∥⋅∥2​; classes form a set Y\mathcal YY. A base classifier is a map f:Rd→Yf:\mathbb R^d\to\mathcal Yf:Rd→Y with Borel decision regions. For a noise level σ>0\sigma>0σ>0, write N(x,σ2I)\mathcal N(x,\sigma^2I)N(x,σ2I) for the isotropic Gaussian law of x+εx+\varepsilonx+ε with ε∼N(0,σ2I)\varepsilon\sim\mathcal N(0,\sigma^2I)ε∼N(0,σ2I). The class probability of ccc at xxx is P(f(x+ε)=c)\mathbb P(f(x+\varepsilon)=c)P(f(x+ε)=c), and the smoothed classifier is

g(x)=arg⁡max⁡c∈Y P(f(x+ε)=c).g(x)=\arg\max_{c\in\mathcal Y}\ \mathbb P(f(x+\varepsilon)=c).g(x)=argc∈Ymax​ P(f(x+ε)=c).

Let Φ\PhiΦ be the standard Gaussian CDF and Φ−1\Phi^{-1}Φ−1 its inverse on (0,1)(0,1)(0,1). A classifier fff is consistent with the observed class probabilities (6) for a top class cAc_AcA​ and numbers pA‾≥pB‾\underline{p_A}\ge\overline{p_B}pA​​≥pB​​ if

P(f(x+ε)=cA) ≥ pA‾ ≥ pB‾ ≥ max⁡c≠cAP(f(x+ε)=c).\mathbb P(f(x+\varepsilon)=c_A)\ \ge\ \underline{p_A}\ \ge\ \overline{p_B}\ \ge\ \max_{c\ne c_A}\mathbb P(f(x+\varepsilon)=c).P(f(x+ε)=cA​) ≥ pA​​ ≥ pB​​ ≥ c=cA​max​P(f(x+ε)=c).

The certified radius is R=σ2(Φ−1(pA‾)−Φ−1(pB‾))R=\frac{\sigma}{2}\big(\Phi^{-1}(\underline{p_A})-\Phi^{-1}(\overline{p_B})\big)R=2σ​(Φ−1(pA​​)−Φ−1(pB​​)). In Lean these are gaussNoise x σ, classProb f σ x c, IsConsistent f σ x cA pA pB and radius σ pA pB in the namespace Cohen2019.Tight, with Phi and PhiInvReal from the series' shared module Cohen2019.Robust; the half-spaces A={z:δT(z−x)≤σ∥δ∥Φ−1(pA‾)}A=\{z:\delta^T(z-x)\le\sigma\|\delta\|\Phi^{-1}(\underline{p_A})\}A={z:δT(z−x)≤σ∥δ∥Φ−1(pA​​)} and B={z:δT(z−x)≥σ∥δ∥Φ−1(1−pB‾)}B=\{z:\delta^T(z-x)\ge\sigma\|\delta\|\Phi^{-1}(1-\overline{p_B})\}B={z:δT(z−x)≥σ∥δ∥Φ−1(1−pB​​)} of the paper's Appendix A are setA and setB.

Quotations write the paper's underlined lower bound as p̲A and its overlined upper bound as p̄B. The PDF has no printed page numbers; every page cited is the PDF page of arXiv:1902.02918v2.

Formalization targets

Goal: Theorem 2 (corrected)

Assume 0<pB‾≤pA‾<10<\overline{p_B}\le\underline{p_A}<10<pB​​≤pA​​<1, pA‾+pB‾≤1\underline{p_A}+\overline{p_B}\le1pA​​+pB​​≤1, and that some finite set sss of classes other than cAc_AcA​ satisfies 1≤pA‾+∣s∣ pB‾1\le\underline{p_A}+|s|\,\overline{p_B}1≤pA​​+∣s∣pB​​. Then for every δ\deltaδ with ∥δ∥2>R\|\delta\|_2>R∥δ∥2​>R there is a base classifier f∗f^*f∗ consistent with (6) and a class c≠cAc\ne c_Ac=cA​ with

P(f∗(x+δ+ε)=cA) < P(f∗(x+δ+ε)=c),\mathbb P(f^*(x+\delta+\varepsilon)=c_A)\ <\ \mathbb P(f^*(x+\delta+\varepsilon)=c),P(f∗(x+δ+ε)=cA​) < P(f∗(x+δ+ε)=c),

so that g(x+δ)≠cAg(x+\delta)\ne c_Ag(x+δ)=cA​ under any tie-breaking. The classifier may depend on δ\deltaδ.

The class-capacity hypothesis is a correction. As printed, with only pA‾+pB‾≤1\underline{p_A}+\overline{p_B}\le1pA​​+pB​​≤1, the theorem fails for two classes: with Y={cA,cB}\mathcal Y=\{c_A,c_B\}Y={cA​,cB​}, pA‾=0.6\underline{p_A}=0.6pA​​=0.6, pB‾=0.1\overline{p_B}=0.1pB​​=0.1 and σ=∥δ∥2=1\sigma=\|\delta\|_2=1σ=∥δ∥2​=1, one has R≈0.767<1R\approx0.767<1R≈0.767<1, yet every consistent fff gives cAc_AcA​ probability at least 0.90.90.9, and Theorem 1 then certifies radius Φ−1(0.9)≈1.28\Phi^{-1}(0.9)\approx1.28Φ−1(0.9)≈1.28.

Milestones

The milestones are the steps the paper itself states, in its order: the Claims P(X∈A)=pA‾\mathbb P(X\in A)=\underline{p_A}P(X∈A)=pA​​ and P(X∈B)=pB‾\mathbb P(X\in B)=\overline{p_B}P(X∈B)=pB​​ for X∼N(x,σ2I)X\sim\mathcal N(x,\sigma^2I)X∼N(x,σ2I); the disjointness of AAA and BBB (corrected to "null" when pA‾+pB‾=1\underline{p_A}+\overline{p_B}=1pA​​+pB​​=1); equations (13) and (14) for Y∼N(x+δ,σ2I)Y\sim\mathcal N(x+\delta,\sigma^2I)Y∼N(x+δ,σ2I),

P(Y∈A)=Φ(Φ−1(pA‾)−∥δ∥σ),P(Y∈B)=Φ(Φ−1(pB‾)+∥δ∥σ);\mathbb P(Y\in A)=\Phi\Big(\Phi^{-1}(\underline{p_A})-\tfrac{\|\delta\|}{\sigma}\Big),\qquad \mathbb P(Y\in B)=\Phi\Big(\Phi^{-1}(\overline{p_B})+\tfrac{\|\delta\|}{\sigma}\Big);P(Y∈A)=Φ(Φ−1(pA​​)−σ∥δ∥​),P(Y∈B)=Φ(Φ−1(pB​​)+σ∥δ∥​);

the equivalence P(Y∈A)<P(Y∈B)  ⟺  ∥δ∥2>R\mathbb P(Y\in A)<\mathbb P(Y\in B)\iff\|\delta\|_2>RP(Y∈A)<P(Y∈B)⟺∥δ∥2​>R; and the existence of the worst-case classifier f∗f^*f∗ satisfying (6) with equalities.

Significance

Theorem 2 makes the guarantee of Theorem 1 exact: when only (6) is known about fff, the set of perturbations under which the Gaussian-smoothed prediction is provably constant is exactly the open ℓ2\ell_2ℓ2​ ball of radius RRR. It settles that improvements to Gaussian-smoothing certificates must use more information about the base classifier than the two bounds, as later work on higher-order and Lipschitz-based certificates does.

The paper's proof is complete in its main lines and has two gaps that this mission records and repairs: the printed statement omits a condition on the number of classes, and the claim A∩B=∅A\cap B=\emptysetA∩B=∅ fails at pA‾+pB‾=1\underline{p_A}+\overline{p_B}=1pA​​+pB​​=1. To our knowledge neither Theorem 1 nor Theorem 2 has a machine-checked proof. Mathlib at the pinned revision has the multivariate standard Gaussian but no normal quantile function and no Gaussian half-space lemma; this mission adds statements for both kinds of fact.

Difficulty

Each step is elementary on paper but rests on facts about Gaussians that Mathlib does not package: the image of the standard Gaussian on Rd\mathbb R^dRd under a linear functional z↦δTzz\mapsto\delta^T zz↦δTz is the one-dimensional Gaussian with variance ∥δ∥2\|\delta\|^2∥δ∥2, and Φ\PhiΦ is a continuous strictly increasing bijection R→(0,1)\mathbb R\to(0,1)R→(0,1) with Φ−1(1−p)=−Φ−1(p)\Phi^{-1}(1-p)=-\Phi^{-1}(p)Φ−1(1−p)=−Φ−1(p). The construction of f∗f^*f∗ has a further step the paper leaves informal: the region between AAA and BBB, of mass 1−pA‾−pB‾1-\underline{p_A}-\overline{p_B}1−pA​​−pB​​, must be shared among "other classes" with none exceeding pB‾\overline{p_B}pB​​, which is where the capacity hypothesis enters. Measurability of the constructed decision regions must be carried along.

Formalization scope

Rd\mathbb R^dRd is EuclideanSpace ℝ (Fin d). N(x,σ2I)\mathcal N(x,\sigma^2I)N(x,σ2I) is the pushforward of Mathlib's stdGaussian under z↦x+σzz\mapsto x+\sigma zz↦x+σz, with σ>0\sigma>0σ>0 a binder. Φ\PhiΦ is cdf (gaussianReal 0 1); Φ−1(p)\Phi^{-1}(p)Φ−1(p) is the generalized inverse inf⁡{t:p≤Φ(t)}\inf\{t:p\le\Phi(t)\}inf{t:p≤Φ(t)}, which is the true inverse on (0,1)(0,1)(0,1) and the junk value 000 at the endpoints, so every statement that evaluates it assumes 0<p<10<p<10<p<1; at pB‾=0\overline{p_B}=0pB​​=0 or pA‾=1\underline{p_A}=1pA​​=1 the paper's radius is infinite and Theorem 2 is vacuous. Class probabilities are real numbers. The base classifier in the conclusion is deterministic with Borel decision regions, which is the stronger existence statement. The conclusion is the strict inequality between class probabilities, not merely the failure of cAc_AcA​ to be a strict unique argmax.

A formalization in which the junk endpoint value of Φ−1\Phi^{-1}Φ−1 makes RRR negative, or in which the classifier's decision regions are non-measurable so that its class probabilities are default values, would make the goal trivial; the hypotheses above exclude both.

Reusable beyond this mission: the Gaussian half-space probabilities and the normal quantile on (0,1)(0,1)(0,1). Contributions of general Mathlib-style lemmas (the law of δTX\delta^T XδTX for X∼N(x,σ2I)X\sim\mathcal N(x,\sigma^2I)X∼N(x,σ2I), properties of Φ−1\Phi^{-1}Φ−1) are welcome.

Selected references

  • J. M. Cohen, E. Rosenfeld, J. Z. Kolter, Certified Adversarial Robustness via Randomized Smoothing, ICML 2019; arXiv:1902.02918v2. https://arxiv.org/abs/1902.02918v2
  • M. Lecuyer, V. Atlidakis, R. Geambasu, D. Hsu, S. Jana, Certified Robustness to Adversarial Examples with Differential Privacy, IEEE S&P 2019. https://arxiv.org/abs/1802.03471
  • B. Li, C. Chen, W. Wang, L. Carin, Certified Adversarial Robustness with Additive Noise, NeurIPS 2019. https://arxiv.org/abs/1809.03113
  • J. Neyman, E. S. Pearson, On the Problem of the Most Efficient Tests of Statistical Hypotheses, Phil. Trans. R. Soc. A 231, 1933. https://doi.org/10.1098/rsta.1933.0009
11 thms2 active usersReviewed
Bandit AlgorithmsOperations ResearchProbability·Captain: mikedeng1

Analysis of Thompson Sampling for the Multi-armed Bandit Problem 2: Logarithmic Regret for N ArmsResearch Paper

Motivation

Thompson Sampling is the oldest heuristic for the multi-armed bandit problem: proposed by Thompson in 1933, it plays each arm with the posterior probability that the arm is the best one. It is simple to implement, performs well empirically (Chapelle and Li, NIPS 2011), and has been used in production systems such as click-through-rate prediction for search advertising. For a long time, however, no finite-time regret guarantee was known for it: the analyses available before 2012 gave only o(T)o(T)o(T) regret.

Agrawal and Goyal (arXiv:1111.1797, COLT 2012) gave the first logarithmic bounds on the expected regret of Thompson Sampling. This mission formalizes their bound for the general case of NNN arms (their Theorem 2). A companion mission of the same series formalizes their two-armed bound (Theorem 1), whose proof is independent.

Timeline. Lai and Robbins (1985) proved that every consistent algorithm has regret at least of order ∑iΔiD(μi∥μ1)ln⁡T\sum_i \frac{\Delta_i}{D(\mu_i\|\mu_1)}\ln T∑i​D(μi​∥μ1​)Δi​​lnT. Auer, Cesa-Bianchi and Fischer (2002) showed that UCB1 achieves O(∑iln⁡T/Δi)O(\sum_i \ln T/\Delta_i)O(∑i​lnT/Δi​) in finite time. Agrawal and Goyal (2012) proved O((∑a1/Δa2)2ln⁡T)O((\sum_a 1/\Delta_a^2)^2\ln T)O((∑a​1/Δa2​)2lnT) for Thompson Sampling with NNN arms; Kaufmann, Korda and Munos (2012) and Agrawal and Goyal (2013) later proved the asymptotically optimal constant for Bernoulli rewards.

Setting

A stochastic NNN-armed bandit has arms 1,…,N1,\dots,N1,…,N. Arm iii, when played, yields a random reward drawn from a fixed distribution νi\nu_iνi​ supported in [0,1][0,1][0,1], with mean μi\mu_iμi​; rewards of an arm are i.i.d. and independent of the other arms. Arm 111 is assumed to be the unique optimal arm, μ1>μi\mu_1>\mu_iμ1​>μi​ for i≠1i\ne1i=1, and Δi=μ1−μi>0\Delta_i=\mu_1-\mu_i>0Δi​=μ1​−μi​>0 is the gap of arm iii.

Thompson Sampling for general stochastic bandits (Algorithm 2 of the paper) keeps, for each arm iii, a count SiS_iSi​ of successes and FiF_iFi​ of failures, both starting at 000. In each round ttt it draws θi(t)∼Beta(Si+1,Fi+1)\theta_i(t)\sim\mathrm{Beta}(S_i+1,F_i+1)θi​(t)∼Beta(Si​+1,Fi​+1) independently for every arm, plays i(t)=arg⁡max⁡iθi(t)i(t)=\arg\max_i\theta_i(t)i(t)=argmaxi​θi​(t), observes a reward r~t∼νi(t)\tilde r_t\sim\nu_{i(t)}r~t​∼νi(t)​, performs a Bernoulli trial with success probability r~t\tilde r_tr~t​, and increments Si(t)S_{i(t)}Si(t)​ on success and Fi(t)F_{i(t)}Fi(t)​ on failure.

The expected regret in time TTT is

E[R(T)]=E[∑t=1T(μ∗−μi(t))],μ∗=max⁡iμi,\mathbb E[\mathcal R(T)]=\mathbb E\Big[\sum_{t=1}^T(\mu^*-\mu_{i(t)})\Big],\qquad \mu^*=\max_i\mu_i,E[R(T)]=E[t=1∑T​(μ∗−μi(t)​)],μ∗=imax​μi​,

the expectation being over the rewards, the Bernoulli trials and the posterior samples.

The proof works with the reward stacks Zi,mZ_{i,m}Zi,m​: the outcome of the mmm-th Bernoulli trial of arm iii, all independent. Then s(j)=∑m≤jZ1,ms(j)=\sum_{m\le j}Z_{1,m}s(j)=∑m≤j​Z1,m​, the number of successes in the first jjj plays of arm 111, is a Binomial(j,μ1)\mathrm{Binomial}(j,\mu_1)Binomial(j,μ1​) random variable. The other objects of the proof are the threshold Li=24ln⁡T/Δi2L_i=24\ln T/\Delta_i^2Li​=24lnT/Δi2​, the saturated set C(t)C(t)C(t) of suboptimal arms with at least LiL_iLi​ plays before round ttt, the intervals IjI_jIj​ between the jjj-th and (j+1)(j+1)(j+1)-th plays of arm 111, and the counts γj\gamma_jγj​ and Vjℓ,aV_j^{\ell,a}Vjℓ,a​ defined in §4.

Formalization targets

Goal: Theorem 2

There is an absolute constant C>0C>0C>0 such that for every N≥2N\ge2N≥2, every instance as above and every horizon T≥2T\ge2T≥2,

E[R(T)]≤C(∑a=2N1Δa2)2ln⁡T.\mathbb E[\mathcal R(T)]\le C\Big(\sum_{a=2}^N\frac{1}{\Delta_a^2}\Big)^2\ln T .E[R(T)]≤C(a=2∑N​Δa2​1​)2lnT.

CCC does not depend on NNN, on the reward distributions or on TTT.

Milestones

  1. Lemma 4: with E(t)E(t)E(t) the event that every saturated arm's sample lies within Δi/2\Delta_i/2Δi​/2 of its mean, Pr⁡(E(t))≥1−4(N−1)/T2\Pr(E(t))\ge1-4(N-1)/T^2Pr(E(t))≥1−4(N−1)/T2, also conditionally on s(j)=ss(j)=ss(j)=s.
  2. Lemma 5 (Eq. (7)): the expected regret from saturated arms inside IjI_jIj​ is at most E[E[γj+1∣s(j)]∑aΔaE[min⁡{X(j,s(j),μa+Δa/2),T}∣s(j)]]\mathbb E\big[\mathbb E[\gamma_j+1\mid s(j)]\sum_a\Delta_a\mathbb E[\min\{X(j,s(j),\mu_a+\Delta_a/2),T\}\mid s(j)]\big]E[E[γj​+1∣s(j)]∑a​Δa​E[min{X(j,s(j),μa​+Δa​/2),T}∣s(j)]].
  3. Lemma 1: E[X(j,s,y)]=1/Fj+1,yB(s)−1\mathbb E[X(j,s,y)]=1/F^B_{j+1,y}(s)-1E[X(j,s,y)]=1/Fj+1,yB​(s)−1, where X(j,s,y)X(j,s,y)X(j,s,y) counts the trials before an independent Beta(s+1,j−s+1)\mathrm{Beta}(s+1,j-s+1)Beta(s+1,j−s+1) sample exceeds yyy.
  4. Lemma 3: a three-case bound on E[E[min⁡{X(j,s(j),y),T}∣s(j)]]\mathbb E[\mathbb E[\min\{X(j,s(j),y),T\}\mid s(j)]]E[E[min{X(j,s(j),y),T}∣s(j)]] in terms of the Bernoulli KL divergence DDD between yyy and μ1\mu_1μ1​.

Significance

The result. Theorem 2 shows that Thompson Sampling, a randomized Bayesian heuristic, achieves regret logarithmic in the horizon for any number of arms with bounded rewards, matching the order in TTT of the Lai–Robbins lower bound. Its dependence on the gaps, (∑aΔa−2)2(\sum_a\Delta_a^{-2})^2(∑a​Δa−2​)2, is worse than UCB1's; the paper's own Remark 1 and later work improve it. The proof introduced the device of bounding the waiting time between plays of the optimal arm through geometric variables with Beta-cdf parameters (Lemmas 1 and 3), which reappears in later analyses of Thompson Sampling.

Formalizing it. The theorem is proved on paper; it has not been machine-checked. Bandit theory in Lean (bandit environments, regret, UCB-type analyses) is still young, and no Beta–Bernoulli Thompson Sampling result is formalized. The mission produces a Lean model of Algorithm 2 for general [0,1][0,1][0,1] rewards with the paper's stack coupling, the §4 bookkeeping of saturated arms and intervals, and the paper's lemmas as separate targets.

Difficulty

Two difficulties are specific to the NNN-armed analysis. First, the arm that competes with arm 111 changes over time: the set of saturated arms grows, and which saturated arm is "best" depends on the history, so the waiting time between plays of arm 111 cannot be compared with a single geometric variable as in the two-armed case. Second, the number γj\gamma_jγj​ of rounds at which arm 111's sample is large but arm 111 is not played is not independent of the counts Vjℓ,aV_j^{\ell,a}Vjℓ,a​: both depend on the same posterior samples, and Lemma 5 needs a careful conditioning on the history to separate them. The obvious union bound over arms, treating each suboptimal arm as in the two-armed proof, fails because it ignores the interruptions by unsaturated arms, whose number is the source of the squared sum in the bound.

Formalization scope

  • Probability space. Algorithm 2 is realized on a product of three independent i.i.d. tables: Beta draws indexed by (arm, round, successes, failures), rewards indexed by (arm, round) and uniform variables indexed by (arm, round); the Bernoulli trial of a round succeeds when the played arm's uniform variable is below its reward. The law of the run is that of Algorithm 2, which runs for every round t=1,2,…t=1,2,\dotst=1,2,…. s(j)s(j)s(j) is the number of successful trials among the first jjj plays of arm 111 in this infinite run (possibly after the horizon TTT), so it is a Binomial(j,μ1)\mathrm{Binomial}(j,\mu_1)Binomial(j,μ1​) random variable for every jjj, as the paper's independent Z1,mZ_{1,m}Z1,m​ make it. Ties in the arg max (probability 000) go to the smallest index.
  • Indexing. Arms are Fin N, and Lean arm 0 is the paper's arm 111. Rounds are 0,…,T−10,\dots,T-10,…,T−1; Lean round ttt is the paper's round t+1t+1t+1. Sums over a=2,…,Na=2,\dots,Na=2,…,N are sums over a≠0a\ne0a=0.
  • Expectations are lower Lebesgue integrals in [0,∞][0,\infty][0,∞], which has no junk value for non-integrable functions. Conditional expectations given s(j)s(j)s(j) are written as finite sums over the values of s(j)s(j)s(j).
  • The O(⋅)O(\cdot)O(⋅). The paper writes O(⋅)O(\cdot)O(⋅) in the sense of its footnote 1 (f(n)≤c g(n)f(n)\le c\,g(n)f(n)≤cg(n) for n≥n0n\ge n_0n≥n0​). The goal states it with one universal constant CCC, quantified before NNN, the instance and TTT, for every T≥2T\ge2T≥2. The explicit constants printed in App. D are not formalized: expanding the paper's Eq. (21) gives terms 288(N−1)(ln⁡T)∑aΔa−2288(N-1)(\ln T)\sum_a\Delta_a^{-2}288(N−1)(lnT)∑a​Δa−2​ and 48(N−1)248(N-1)^248(N−1)2 where the paper prints 288(ln⁡T)∑iΔi−2288(\ln T)\sum_i\Delta_i^{-2}288(lnT)∑i​Δi−2​, and Eq. (22) drops a factor ln⁡T\ln TlnT in its 192/Δa2192/\Delta_a^2192/Δa2​ term. The O(⋅)O(\cdot)O(⋅) claim does not depend on these slips; a statement pinned to the printed numerals might be false.
  • Ruled out. A constant depending on NNN, on the means or on TTT; a fixed number of arms; Bernoulli rewards only; or any algorithm other than Algorithm 2 would each make the goal a different and weaker theorem. The statement quantifies over all N≥2N\ge2N≥2 and all reward distributions on [0,1][0,1][0,1].
  • Not included. Eq. (8), the bound ∑jE[γj∣s(j)]≤∑uLu+4(N−1)\sum_{j}\mathbb E[\gamma_j\mid s(j)]\le\sum_uL_u+4(N-1)∑j​E[γj​∣s(j)]≤∑u​Lu​+4(N−1) "for all instantiations", is not a milestone: each term is conditioned on a different s(j)s(j)s(j), and the pointwise reading does not follow from the argument given. Remark 1 (an alternate bound) and App. A (several optimal arms) are not part of this mission.
  • Contributions welcome: Beta–Binomial identities, geometric waiting times, Hoeffding bounds for binomial cdfs, and the stopping-time arguments behind Lemma 5. Lemma 1 and Lemma 3 are shared with the two-armed mission of this series.

Selected references

  • S. Agrawal and N. Goyal, Analysis of Thompson Sampling for the Multi-armed Bandit Problem, COLT 2012; arXiv:1111.1797v3, 2012. https://arxiv.org/abs/1111.1797
  • W. R. Thompson, On the likelihood that one unknown probability exceeds another in view of the evidence of two samples, Biometrika 25, 1933. https://doi.org/10.1093/biomet/25.3-4.285
  • T. L. Lai and H. Robbins, Asymptotically efficient adaptive allocation rules, Advances in Applied Mathematics 6, 1985. https://doi.org/10.1016/0196-8858(85)90002-8
  • P. Auer, N. Cesa-Bianchi and P. Fischer, Finite-time analysis of the multiarmed bandit problem, Machine Learning 47, 2002. https://doi.org/10.1023/A:1013689704352
  • O. Chapelle and L. Li, An empirical evaluation of Thompson Sampling, NIPS 2011. https://papers.nips.cc/paper/4321-an-empirical-evaluation-of-thompson-sampling
  • E. Kaufmann, N. Korda and R. Munos, Thompson Sampling: an asymptotically optimal finite-time analysis, ALT 2012. https://arxiv.org/abs/1205.4217
  • S. Agrawal and N. Goyal, Further optimal regret bounds for Thompson Sampling, AISTATS 2013. https://arxiv.org/abs/1209.3353
11 thms2 active usersReviewed
🏆Completed
AnalysisFunctional Analysis·Captain: mikedeng1

Theory of Reproducing Kernels II: One Reproducing Kernel Class Contains Another iff a Multiple of Its Kernel Dominates the Other KernelResearch Paper

Motivation

A reproducing kernel Hilbert space (RKHS) is a Hilbert space of functions in which evaluation at each point is a continuous functional. Each such space is determined by its kernel, and in practice spaces are usually given by their kernels: Gaussian or Sobolev kernels in kernel methods and Gaussian-process regression, the Bergman and Szegő kernels in complex analysis. A recurring question is therefore how a relation between two function spaces reads on their kernels. When is one space contained in another? When do two kernels define the same space?

N. Aronszajn's Theory of Reproducing Kernels (Trans. Amer. Math. Soc. 68 (1950), 337–404) answers both questions. §7 (pp. 354–356) compares a class with a contractively included subclass. §13 (C) (pp. 382–383) uses Banach's closed graph theorem to remove the contractivity assumption. The answer is an order relation between kernels, checkable on finite point sets. It is the standard tool for comparing RKHSs: Part II of the paper (§2, p. 387) applies §7, Theorem II to compare Bergman kernels of nested plane domains.

Setting

Let XXX be an arbitrary set (no topology, no measure). A function K:X×X→CK : X\times X\to\mathbb CK:X×X→C is a positive matrix if

∑i,j=1nK(yi,yj) ξˉi ξj ≥0\sum_{i,j=1}^n K(y_i,y_j)\,\bar\xi_i\,\xi_j\ \ge 0i,j=1∑n​K(yi​,yj​)ξˉ​i​ξj​ ≥0

for every finite family y1,…,yn∈Xy_1,\dots,y_n\in Xy1​,…,yn​∈X and every ξ1,…,ξn∈C\xi_1,\dots,\xi_n\in\mathbb Cξ1​,…,ξn​∈C. For two kernels, K1≪KK_1\ll KK1​≪K means that K−K1K-K_1K−K1​ is a positive matrix.

A class with a reproducing kernel is a complex Hilbert space FFF of functions X→CX\to\mathbb CX→C for which there is a function KKK with K(⋅,y)∈FK(\cdot,y)\in FK(⋅,y)∈F and f(y)=(f,K(⋅,y))f(y) = (f,K(\cdot,y))f(y)=(f,K(⋅,y)) for all f∈Ff\in Ff∈F and y∈Xy\in Xy∈X. The scalar product (f,g)(f,g)(f,g) is linear in fff. Every such KKK is a positive matrix, and by Moore's theorem (§2 (4)) every positive matrix is the kernel of exactly one such class. A linear class of functions is an (R.K.)-class if some norm makes it a Hilbert space with a reproducing kernel. Such a class carries many admissible norms, and correspondingly many kernels.

In the Lean development, FFF is a Mathlib RKHS ℂ H X ℂ instance on a complete complex inner product space H. The class of functions is Set.range (⇑ : H → X → ℂ). The scalar kernel is kernelFn H x y = RKHS.kernel H x y 1, and K1≪KK_1\ll KK1​≪K is AronszajnRK.Limits.KernelLE K₁ K (the series' shared definition).

Formalization targets

Goal: Corollary IV₂ of §13

For classes F,F1F, F_1F,F1​ with reproducing kernels K,K1K, K_1K,K1​:

F1⊂F  ⟺  ∃ M>0: K1≪MK.F_1\subset F \iff \exists\, M>0:\ K_1\ll MK .F1​⊂F⟺∃M>0: K1​≪MK.

The constant MMM is left unspecified: it measures the norm of the inclusion operator and has no canonical value.

Milestones, in attack order

  1. ≪ is a partial ordering of the positive matrices (§7, p. 354).
  2. §7, Theorem I: if K1≪KK_1\ll KK1​≪K, then F1⊂FF_1\subset FF1​⊂F and ∥f1∥1≥∥f1∥\|f_1\|_1\ge\|f_1\|∥f1​∥1​≥∥f1​∥.
  3. §7, Theorem II: if a linear class F1⊂FF_1\subset FF1​⊂F is a Hilbert space with ∥f1∥1≥∥f1∥\|f_1\|_1\ge\|f_1\|∥f1​∥1​≥∥f1​∥, then F1F_1F1​ has a reproducing kernel K1≪KK_1\ll KK1​≪K.
  4. §13 (C), p. 382: the norm c∥⋅∥c\|\cdot\|c∥⋅∥ corresponds to the same class, with kernel c−2Kc^{-2}Kc−2K.
  5. §13 (C), Lemma: the identity correspondence from F1⋅F2⊂F1F_1\cdot F_2\subset F_1F1​⋅F2​⊂F1​ to F2F_2F2​ is a closed linear transformation.
  6. §13 (C), Theorem IV: if F1⊂FF_1\subset FF1​⊂F are (R.K.)-classes, then ∥f∥≤M∥f∥1\|f\|\le M\|f\|_1∥f∥≤M∥f∥1​ on F1F_1F1​ for some M>0M>0M>0, whatever admissible norms are chosen.
  7. §13 (C), Corollary IV₁: any two admissible norms on one (R.K.)-class are equivalent.
  8. §13 (C), Corollary IV₃: F1=FF_1 = FF1​=F iff mK≪K1≪MKmK\ll K_1\ll MKmK≪K1​≪MK for some m,M>0m,M>0m,M>0.
  9. §13 (D), Theorem VI: intersections and sums of (R.K.)-classes are (R.K.)-classes.

Significance

The result. Corollary IV₂ reduces inclusion of infinite-dimensional function spaces to positivity of finite matrices. Corollary IV₃ says when two kernels give the same space, as sets of functions with equivalent norms. Kernel-method theory uses these facts to compare hypothesis spaces of different kernels. In complex analysis they compare Bergman spaces of nested domains, since the restricted kernel of a larger domain is dominated by the kernel of a smaller one (Part II, §2, p. 387, of the paper). Theorem VI makes (R.K.)-classes a lattice under intersection and sum.

Formalizing it. All statements are classical and have been proved since 1950. To our knowledge none is machine-checked. Mathlib has RKHSs, their kernels, positive semidefiniteness of kernels and the construction of a space from a kernel. It has no comparison theorem between two RKHSs, and no characterization of an RKHS by its kernel up to inclusion or equality of function sets. This mission adds these. As a by-product it produces the closed-graph argument for RKHS inclusions in a form that other missions can reuse.

Difficulty

The sufficiency direction (K1≪MK⇒F1⊂FK_1\ll MK \Rightarrow F_1\subset FK1​≪MK⇒F1​⊂F) does not follow from the definitions alone. Domination is a statement about finite quadratic forms; membership of a function of F1F_1F1​ in FFF is a statement about an infinite-dimensional space, and nothing pointwise connects the two. The paper's argument rests on the sum theorem of §6, the goal of a separate mission in this series.

The necessity direction cannot start from an assumed norm inequality, because the two norms are a priori unrelated. Producing the constant MMM needs the closed graph theorem together with the fact that norm convergence in an RKHS implies pointwise convergence. A direct estimate of ∑K1(yi,yj)ξˉiξj\sum K_1(y_i,y_j)\bar\xi_i\xi_j∑K1​(yi​,yj​)ξˉ​i​ξj​ against ∑K(yi,yj)ξˉiξj\sum K(y_i,y_j)\bar\xi_i\xi_j∑K(yi​,yj​)ξˉ​i​ξj​ from the inclusion alone has no route to a uniform constant.

Formalization scope

  • Scalars and spaces. All scalars are complex and XXX is an arbitrary type. A class with a reproducing kernel is an RKHS ℂ H X ℂ instance on a complete complex inner product space H. The function of f : H is ⇑f.
  • Scalar product. Mathlib's ⟪f, g⟫_ℂ is conjugate-linear in f, so Aronszajn's (f,g)(f,g)(f,g) is ⟪g, f⟫_ℂ, and the reproducing property is ⟪k y, f⟫_ℂ = f y.
  • Inclusion and norm comparison. "F1⊂FF_1\subset FF1​⊂F" is inclusion of the sets of functions. A norm comparison ∥f∥≤M∥f∥1\|f\|\le M\|f\|_1∥f∥≤M∥f∥1​ is stated between the elements of the two spaces that are the same function.
  • Positivity. Positive matrices are Matrix.PosSemidef over the index type X. Its quadratic form over finitely supported vectors is exactly the paper's. Over C\mathbb CC it includes Hermitian symmetry, which nonnegativity implies. MKMKMK is fun x y => (M : ℂ) * K x y with real M>0M>0M>0.
  • The goal and Corollary IV₃. "FFF, F1F_1F1​ the corresponding classes" of positive matrices KKK, K1K_1K1​ is read as: arbitrary RKHSs whose kernels are KKK, K1K_1K1​. By Moore's uniqueness theorem this is the same statement.
  • Theorem II. F1F_1F1​ is not assumed to carry an RKHS instance, since having a kernel is the conclusion. It is a complex Hilbert space with an injective linear map into the functions of FFF.
  • Theorem IV and Corollary IV₁ quantify over all admissible norms, i.e. all RKHSs with the given function sets.
  • Closedness of a transformation is the paper's sequential definition (§13 (C), p. 382), stated on a submodule that need not be closed.
  • (R.K.)-class means: equals the set of functions of some RKHS in the universe of XXX.
  • Rescaling milestone. It asserts that a rescaled RKHS exists, and that every RKHS with the same functions and the rescaled norm has the rescaled kernel.

Trivializations ruled out. No statement defines a class as RKHS.OfKernel K and then asserts that its kernel is K. Classes are always compared through their sets of functions and norms. "F1⊂FF_1\subset FF1​⊂F" is never weakened to the existence of some injection H1→HH_1\to HH1​→H.

Substrate. Mathlib's RKHS, RKHS.kernel, RKHS.kerFun, RKHS.kerFun_inner, RKHS.posSemidef_kernel, RKHS.OfKernel, Matrix.PosSemidef (with .add, .smul), Banach's LinearMap.continuous_of_isClosed_graph, and the continuous functional calculus (CFC.sqrt, for the square root (I−L)1/2(I-L)^{1/2}(I−L)1/2 in the proof of Theorem II). The sum theorem of §6, used in the proof of Theorem I, is the goal of mission I of this series. Contributions of that theorem and of reusable lemmas on RKHS inclusions are welcome.

Selected references

  • N. Aronszajn, Theory of Reproducing Kernels, Trans. Amer. Math. Soc. 68 (1950), 337–404. https://doi.org/10.1090/S0002-9947-1950-0051437-7
  • S. Banach, Théorie des opérations linéaires, Monografje Matematyczne 1, Warsaw, 1932 (closed graph theorem, as cited on p. 382).
  • V. I. Paulsen and M. Raghupathi, An Introduction to the Theory of Reproducing Kernel Hilbert Spaces, Cambridge Stud. Adv. Math. 152, Cambridge University Press, 2016 (modern treatment of the inclusion criterion). https://doi.org/10.1017/CBO9781316219232
14 thms2 active usersReviewed
Convex OptimizationProbabilityStatistics·Captain: mikedeng1

Stability and Generalization 4: Relative-Entropy Regularization of Mixtures Has Uniform Stability M²/(λm)Research Paper

Motivation

A learning algorithm generalizes when its error on fresh data is close to its error on the training sample. Bousquet and Elisseeff (JMLR 2, 2002) showed that a single property of the algorithm, uniform stability, controls this gap with exponential concentration: if removing any one example from a training set of size mmm changes the loss of the output at every point by at most β\betaβ, the generalization error exceeds the empirical error by roughly 2β+(4mβ+M)ln⁡(1/δ)/(2m)2\beta + (4m\beta + M)\sqrt{\ln(1/\delta)/(2m)}2β+(4mβ+M)ln(1/δ)/(2m)​ with probability 1−δ1-\delta1−δ (their Theorem 12). The bound is useful only when β=O(1/m)\beta = O(1/m)β=O(1/m), and the second half of the paper identifies algorithms with that rate: Tikhonov regularization in a reproducing kernel Hilbert space (Theorem 22), and relative-entropy regularization of mixtures (Theorem 24), the subject of this mission.

Mixtures arise whenever a learner outputs a distribution over a parametric base class instead of a single hypothesis: Bayesian posterior averaging, Gibbs and randomized classifiers, exponential weights. Regularizing by the relative entropy to a prior is the maximum-a-posteriori reading of these procedures, and Theorem 24 is one of the earliest results showing that such posteriors are uniformly stable with rate 1/(λm)1/(\lambda m)1/(λm). The same mechanism (entropic regularization, stability through Pinsker's inequality) reappears in PAC-Bayesian analysis and in the stability of exponential-weights methods.

Setting

Let Θ\ThetaΘ be a measurable space with a reference measure ν\nuν, and write dθd\thetadθ for integration against ν\nuν. A base class H={hθ:θ∈Θ}\mathcal H = \{h_\theta : \theta \in \Theta\}H={hθ​:θ∈Θ} is indexed by Θ\ThetaΘ, and r(hθ,z)∈[0,M]r(h_\theta, z) \in [0, M]r(hθ​,z)∈[0,M] is the loss of the base hypothesis hθh_\thetahθ​ at an example z∈Zz \in Zz∈Z.

The algorithm outputs a density ggg with respect to ν\nuν: a measurable, nonnegative, integrable g:Θ→Rg : \Theta \to \mathbb Rg:Θ→R with ∫Θg dθ=1\int_\Theta g\,d\theta = 1∫Θ​gdθ=1. FFF denotes the set of all densities. A density is scored by the averaged loss

ℓ(g,z)=∫Θr(hθ,z) g(θ) dθ(28),\ell(g, z) = \int_\Theta r(h_\theta, z)\, g(\theta)\, d\theta \qquad (28),ℓ(g,z)=∫Θ​r(hθ​,z)g(θ)dθ(28),

the expected loss of a randomized predictor that draws hθh_\thetahθ​ from ggg. The relative entropy of ggg to g′g'g′ is

K(g,g′)=∫Θg(θ)ln⁡g(θ)g′(θ) dθ∈[0,∞],K(g, g') = \int_\Theta g(\theta) \ln \frac{g(\theta)}{g'(\theta)}\, d\theta \in [0, \infty],K(g,g′)=∫Θ​g(θ)lng′(θ)g(θ)​dθ∈[0,∞],

with K(g,g′)=+∞K(g, g') = +\inftyK(g,g′)=+∞ when g νg\,\nugν is not absolutely continuous with respect to g′ νg'\,\nug′ν or the integrand is not integrable.

Fix a prior f0∈Ff_0 \in Ff0​∈F, a parameter λ>0\lambda > 0λ>0, and a training set S=(z1,…,zm)S = (z_1, \dots, z_m)S=(z1​,…,zm​). The algorithm returns a minimizer over FFF of

Rr(g)=1m∑j=1mℓ(g,zj)+λK(g,f0)(29).R_r(g) = \frac1m \sum_{j=1}^m \ell(g, z_j) + \lambda K(g, f_0) \qquad (29).Rr​(g)=m1​j=1∑m​ℓ(g,zj​)+λK(g,f0​)(29).

For an index iii, the truncated objective is Rr∖i(g)=1m∑j≠iℓ(g,zj)+λK(g,f0)R_r^{\setminus i}(g) = \frac1m \sum_{j \ne i} \ell(g, z_j) + \lambda K(g, f_0)Rr∖i​(g)=m1​∑j=i​ℓ(g,zj​)+λK(g,f0​), and f∖if^{\setminus i}f∖i denotes one of its minimizers over FFF.

Formalization targets

Goal: Theorem 24

For every minimizer fff of (29), every minimizer f∖if^{\setminus i}f∖i of the truncated objective, and every example zzz,

∣ℓ(f,z)−ℓ(f∖i,z)∣≤M2λm.|\ell(f, z) - \ell(f^{\setminus i}, z)| \le \frac{M^2}{\lambda m}.∣ℓ(f,z)−ℓ(f∖i,z)∣≤λmM2​.

Milestones

  1. MMM-admissibility of (28) (§5.2.3, p. 518): ∣ℓ(g,z)−ℓ(g′,z)∣≤M∫Θ∣g−g′∣ dθ|\ell(g,z) - \ell(g',z)| \le M \int_\Theta |g - g'|\,d\theta∣ℓ(g,z)−ℓ(g′,z)∣≤M∫Θ​∣g−g′∣dθ.
  2. Pinsker's inequality, L1L^1L1 form (proof of Theorem 24): 12(∫Θ∣g−g′∣ dθ)2≤K(g,g′)\tfrac12 \bigl(\int_\Theta |g - g'|\,d\theta\bigr)^2 \le K(g, g')21​(∫Θ​∣g−g′∣dθ)2≤K(g,g′) for densities g,g′g, g'g,g′.
  3. Lemma 21 (p. 513): for a differentiable convex regularizer NNN on a vector space and a σ\sigmaσ-admissible loss,
dN(f,f∖i)+dN(f∖i,f)≤1λm(ℓ(f∖i,zi)−ℓ(f,zi)−dℓ(⋅,zi)(f∖i,f))≤σλm∣Δf(xi)∣.d_N(f, f^{\setminus i}) + d_N(f^{\setminus i}, f) \le \frac{1}{\lambda m}\Bigl(\ell(f^{\setminus i}, z_i) - \ell(f, z_i) - d_{\ell(\cdot, z_i)}(f^{\setminus i}, f)\Bigr) \le \frac{\sigma}{\lambda m}|\Delta f(x_i)|.dN​(f,f∖i)+dN​(f∖i,f)≤λm1​(ℓ(f∖i,zi​)−ℓ(f,zi​)−dℓ(⋅,zi​)​(f∖i,f))≤λmσ​∣Δf(xi​)∣.
  1. Bregman divergence of the relative entropy (proof of Theorem 24): dK(⋅,f0)(g,g′)=K(g,g′)d_{K(\cdot, f_0)}(g, g') = K(g, g')dK(⋅,f0​)​(g,g′)=K(g,g′).
  2. L1L^1L1 displacement bound (proof of Theorem 24):
∫Θ∣f−f∖i∣ dθ≤Mλm.\int_\Theta |f - f^{\setminus i}|\,d\theta \le \frac{M}{\lambda m}.∫Θ​∣f−f∖i∣dθ≤λmM​.

Significance

Theorem 24 places entropy-regularized posteriors among the algorithms to which the paper's exponential generalization bound applies: combined with Theorem 12 it gives, for the averaged loss, a deviation of order M2/(λm)+(M2/λ+M)ln⁡(1/δ)/mM^2/(\lambda m) + (M^2/\lambda + M)\sqrt{\ln(1/\delta)/m}M2/(λm)+(M2/λ+M)ln(1/δ)/m​. The proof also yields the L1L^1L1 bound ∫∣f−f∖i∣≤M/(λm)\int |f - f^{\setminus i}| \le M/(\lambda m)∫∣f−f∖i∣≤M/(λm), which by itself gives classification stability M/(λm)M/(\lambda m)M/(λm) for base hypotheses with values in {−1,1}\{-1, 1\}{−1,1} (remark after Theorem 24, p. 518).

The result is proved in the paper; no machine-checked proof is known to exist. A formalization produces reusable pieces that Mathlib does not have: Pinsker's inequality for densities in L1L^1L1 form (Mathlib has the Kullback–Leibler divergence InformationTheory.klDiv, but not Pinsker), the Bregman identity for the relative entropy, and a stability statement for minimizers over a space of probability densities.

Difficulty

The paper derives Theorem 24 from Lemma 21, which is stated for a regularizer that is defined and differentiable on a vector space. The relative entropy K(⋅,f0)K(\cdot, f_0)K(⋅,f0​) is defined only on the convex set of densities and is not differentiable at densities that vanish on a set of positive measure, so the general lemma does not literally apply, and the identity dK(⋅,f0)=Kd_{K(\cdot,f_0)} = KdK(⋅,f0​)​=K needs integrability conditions that the page does not state. A complete proof of the goal must either justify that application on the set of densities, or work directly with the minimizers, which requires identifying them and handling the +∞+\infty+∞ values of KKK. Pinsker's inequality itself requires a separate argument at the level of general measures.

Formalization scope

  • Densities are IsDensity ν g: measurable, nonnegative, integrable, total mass one, with respect to a σ-finite reference measure ν. The integral dθd\thetadθ is always against ν, never Lebesgue measure.
  • The base loss is r : Θ → Z → ℝ, measurable in θ, with 0 ≤ r ≤ M; the paper's costs are nonnegative (p. 502).
  • KKK is InformationTheory.klDiv of the measures g · ν and g' · ν, in ℝ≥0∞. The objectives (29) and its truncation take values in ℝ≥0∞. A formalization that converts KKK to a real number with toReal would send K=+∞K = +\inftyK=+∞ to 000 and make the worst densities minimizers; that reading is excluded.
  • The minimizers are given as hypotheses: f minimizes (29) and f' minimizes the truncated objective over all densities, for the given S : Fin m → Z and i : Fin m.
  • Corrected reading of the algorithm on S∖iS^{\setminus i}S∖i. The goal is stated in the pairwise form of the paper's proof: f∖if^{\setminus i}f∖i minimizes the truncated objective with factor 1/m1/m1/m, the analogue of (20), not (29) run on the m−1m-1m−1 points of S∖iS^{\setminus i}S∖i with factor 1/(m−1)1/(m-1)1/(m−1).
  • Corrected display. The objective displayed before Theorem 24 has ℓ(g,z)\ell(g, z)ℓ(g,z) inside the sum; (29) has ℓ(g,zi)\ell(g, z_i)ℓ(g,zi​), which is used.
  • Lemma 21 is stated as printed, in its differentiable case, on a real normed space whose elements act as functions on XXX through a linear map; the goal does not instantiate it. The Bregman identity is stated with the explicit gradient ln⁡(g′/f0)+1\ln(g'/f_0) + 1ln(g′/f0​)+1, for f0,g′>0f_0, g' > 0f0​,g′>0, finite K(g,f0)K(g, f_0)K(g,f0​), K(g′,f0)K(g', f_0)K(g′,f0​), and integrable gln⁡(g′/f0)g \ln(g'/f_0)gln(g′/f0​).

Contributions welcome: proofs of Pinsker's inequality for klDiv (reusable far beyond this mission), of the Bregman identity, of Lemma 21, and of the goal by any route.

Selected references

  • O. Bousquet and A. Elisseeff, Stability and Generalization, Journal of Machine Learning Research 2 (2002) 499–526. https://www.jmlr.org/papers/v2/bousquet02a.html
  • T. M. Cover and J. A. Thomas, Elements of Information Theory, Wiley, 1991 (Pinsker's inequality). https://doi.org/10.1002/0471200611
  • R. T. Rockafellar, Convex Analysis, Princeton University Press, 1970 (Bregman divergences, Appendix C of the paper). https://doi.org/10.1515/9781400873173
11 thms2 active usersReviewed
ProbabilityStatistics·Captain: mikedeng1

Stability and Generalization 2: Exponential Generalization Bounds for Uniformly Stable AlgorithmsResearch Paper

Motivation

A learning algorithm is judged by its generalization error: the expected loss of the hypothesis it outputs on a fresh example. That quantity depends on an unknown distribution, so it is estimated from the training data, by the empirical error (the average loss on the training set) or the leave-one-out error (the average loss on each training point of the hypothesis trained without it). Classical learning theory controls the gap between these estimates and the true error uniformly over a hypothesis class, through its VC dimension or covering numbers. Such bounds say nothing useful about algorithms that search very large or infinite-dimensional spaces, such as support vector machines and regularization networks in a reproducing kernel Hilbert space.

Bousquet and Elisseeff (JMLR 2 (2002) 499–526) replaced the capacity of the class by a property of the algorithm, its stability: how much its output changes when one training example is removed. Their exponential bound for uniformly stable algorithms is the starting point of the stability approach to generalization, which was later used for stochastic gradient descent (Hardt, Recht and Singer, 2016) and differential privacy, and sharpened by Feldman and Vondrák (2019) and Bousquet, Klochkov and Zhivotovskiy (2020).

Timeline. Rogers and Wagner (1978) and Devroye and Wagner (1979) bounded the leave-one-out error of local rules such as k-nearest neighbours through their stability. McDiarmid (1989) proved the bounded-differences inequality. Lugosi and Pawlak (1994) combined it with smoothed error estimates. Kearns and Ron (1999) named hypothesis and error stability and related them to the VC dimension. Bousquet and Elisseeff (2002) introduced uniform stability and proved the exponential bounds this mission formalizes.

Setting

Let Z=X×YZ = X \times YZ=X×Y be a measurable space of labelled examples with an unknown probability distribution DDD. A training set S=(z1,…,zm)S = (z_1, \dots, z_m)S=(z1​,…,zm​) is drawn from DmD^mDm. A learning algorithm AAA maps a training set to a hypothesis AS:X→Y′A_S : X \to Y'AS​:X→Y′. It is deterministic and symmetric: it depends on the training set only as a multiset, so it is a function Multiset (X × Y) → (X → Y'), defined for training sets of every size. For a cost ccc, the loss of a hypothesis fff at z=(x,y)z = (x, y)z=(x,y) is ℓ(f,z)=c(f(x),y)\ell(f, z) = c(f(x), y)ℓ(f,z)=c(f(x),y).

Given SSS, write S∖iS^{\setminus i}S∖i for SSS with its iii-th example removed, and SiS^iSi for SSS with ziz_izi​ replaced by an independent draw zi′∼Dz_i' \sim Dzi′​∼D. The three errors are

R=Ez∼D[ℓ(AS,z)],Remp=1m∑i=1mℓ(AS,zi),Rloo=1m∑i=1mℓ(AS∖i,zi).R = \mathbb E_{z \sim D}[\ell(A_S, z)], \qquad R_{\mathrm{emp}} = \frac1m \sum_{i=1}^m \ell(A_S, z_i), \qquad R_{\mathrm{loo}} = \frac1m \sum_{i=1}^m \ell(A_{S^{\setminus i}}, z_i).R=Ez∼D​[ℓ(AS​,z)],Remp​=m1​i=1∑m​ℓ(AS​,zi​),Rloo​=m1​i=1∑m​ℓ(AS∖i​,zi​).

An algorithm has uniform stability β\betaβ at sample size mmm (Definition 6) if for every S∈ZmS \in Z^mS∈Zm, every iii and every z∈Zz \in Zz∈Z,

∣ℓ(AS,z)−ℓ(AS∖i,z)∣≤β.|\ell(A_S, z) - \ell(A_{S^{\setminus i}}, z)| \le \beta .∣ℓ(AS​,z)−ℓ(AS∖i​,z)∣≤β.

As a function of the sample size this constant is written βm\beta_mβm​.

Formalization targets

Goal: Theorem 12

If AAA has uniform stability β\betaβ and 0≤ℓ(AS,z)≤M0 \le \ell(A_S, z) \le M0≤ℓ(AS​,z)≤M for all zzz and all training sets SSS, then for every m≥1m \ge 1m≥1 and δ∈(0,1)\delta \in (0,1)δ∈(0,1), each of the following holds, separately, with probability at least 1−δ1 - \delta1−δ over S∼DmS \sim D^mS∼Dm:

R≤Remp+2β+(4mβ+M)ln⁡(1/δ)2m,(11)R \le R_{\mathrm{emp}} + 2\beta + (4m\beta + M)\sqrt{\frac{\ln(1/\delta)}{2m}}, \qquad (11)R≤Remp​+2β+(4mβ+M)2mln(1/δ)​​,(11) R≤Rloo+β+(4mβ+M)ln⁡(1/δ)2m.(12)R \le R_{\mathrm{loo}} + \beta + (4m\beta + M)\sqrt{\frac{\ln(1/\delta)}{2m}}. \qquad (12)R≤Rloo​+β+(4mβ+M)2mln(1/δ)​​.(12)

Milestones

  1. McDiarmid's inequality (Theorem 2): for measurable F:Zm→RF : Z^m \to \mathbb RF:Zm→R with ∣F(S)−F(Si)∣≤ci|F(S) - F(S^i)| \le c_i∣F(S)−F(Si)∣≤ci​, PS[F−ESF≥ϵ]≤e−2ϵ2/∑ici2P_S[F - \mathbb E_S F \ge \epsilon] \le e^{-2\epsilon^2/\sum_i c_i^2}PS​[F−ES​F≥ϵ]≤e−2ϵ2/∑i​ci2​.
  2. Uniform stability β\betaβ implies ∣ℓ(AS,z)−ℓ(ASi,z)∣≤2β|\ell(A_S, z) - \ell(A_{S^i}, z)| \le 2\beta∣ℓ(AS​,z)−ℓ(ASi​,z)∣≤2β (p. 504).
  3. Lemma 7: the bias identities for ES[R−Remp]\mathbb E_S[R - R_{\mathrm{emp}}]ES​[R−Remp​], ES[R(A,S∖i)−Rloo]\mathbb E_S[R(A,S^{\setminus i}) - R_{\mathrm{loo}}]ES​[R(A,S∖i)−Rloo​] and ES[R−Rloo]\mathbb E_S[R - R_{\mathrm{loo}}]ES​[R−Rloo​].
  4. R−RempR - R_{\mathrm{emp}}R−Remp​ and R−RlooR - R_{\mathrm{loo}}R−Rloo​ have bounded differences ci=4β+M/mc_i = 4\beta + M/mci​=4β+M/m.
  5. ES[R−Remp]≤2β\mathbb E_S[R - R_{\mathrm{emp}}] \le 2\betaES​[R−Remp​]≤2β and ES[R−Rloo]≤β\mathbb E_S[R - R_{\mathrm{loo}}] \le \betaES​[R−Rloo​]≤β.
  6. The tail bounds PS[R−Remp>ϵ+2β]≤exp⁡(−2mϵ2/(4mβ+M)2)P_S[R - R_{\mathrm{emp}} > \epsilon + 2\beta] \le \exp(-2m\epsilon^2/(4m\beta+M)^2)PS​[R−Remp​>ϵ+2β]≤exp(−2mϵ2/(4mβ+M)2) and the leave-one-out analogue.

Significance

When β=O(1/m)\beta = O(1/m)β=O(1/m) both bounds are O(1/m)O(1/\sqrt m)O(1/m​), with constants that do not depend on any capacity of the hypothesis space. Later sections of the paper show that Tikhonov regularization in a reproducing kernel Hilbert space has β=O(1/(λm))\beta = O(1/(\lambda m))β=O(1/(λm)), so the theorem gives generalization bounds for support vector regression, kernel ridge regression and, through a smoothed loss, soft-margin classification. The theorem is also the template for later stability bounds: the decomposition into a bias term controlled by stability and a deviation term controlled by a concentration inequality recurs throughout the literature.

The result has been proved since 2002, and replace-one variants appear in textbooks (Mohri, Rostamizadeh and Talwalkar, Foundations of Machine Learning, Theorem 14.2; Shalev-Shwartz and Ben-David, Chapter 13). On Prove2Me the replace-one textbook version is not formalized, and Mathlib at the platform's environment has no McDiarmid inequality. This mission asks for a machine-checked proof of the paper's remove-one version with its exact constants, and a reusable McDiarmid inequality with per-coordinate constants.

Difficulty

The deterministic steps (the bias identity and the bounded-differences estimates) are short on paper. The central difficulty is McDiarmid's inequality itself: it needs a martingale argument along the coordinates of a product measure, or an equivalent tensorization of conditional sub-Gaussian bounds, with the Doob martingale E[F∣z1,…,zk]\mathbb E[F \mid z_1, \dots, z_k]E[F∣z1​,…,zk​] expressed through partial integration over Measure.pi. Hoeffding's inequality for sums, which Mathlib has, does not apply directly: R−RempR - R_{\mathrm{emp}}R−Remp​ is not a sum of independent terms. A second, bookkeeping difficulty is Lemma 7: the identities rest on exchanging ziz_izi​ with zi′z_i'zi′​ and on the symmetry of AAA, which in Lean means measure-preserving coordinate permutations of Dm⊗DD^m \otimes DDm⊗D and multiset equalities such as Si ∖i=S∖iS^{i\,\setminus i} = S^{\setminus i}Si∖i=S∖i.

Formalization scope

Conventions committed to by the Lean statements:

  • An algorithm is a function of a multiset; this is how symmetry is encoded. Samples are Fin m → X × Y, and SiS^iSi is Function.update.
  • The law of SSS is Measure.pi (fun _ => D) with D a probability measure; zi′z_i'zi′​ and zzz are independent draws, integrated against the product (Measure.pi fun _ => D).prod D.
  • The paper's standing assumption that all functions are measurable is one hypothesis: (S,z)↦ℓ(AS,z)(S, z) \mapsto \ell(A_S, z)(S,z)↦ℓ(AS​,z) is measurable for every sample size. With the bound 0≤ℓ(AT,z)≤M0 \le \ell(A_T, z) \le M0≤ℓ(AT​,z)≤M for training sets TTT of every size, every expectation is a genuine integral, so no bound can hold because a non-integrable expectation defaults to 000.
  • Uniform stability quantifies over every sample, every index and every point, not almost every one.
  • "With probability at least 1−δ1 - \delta1−δ" means the DmD^mDm-measure of the failure set is at most δ\deltaδ. The two bounds (11) and (12) are separate statements, joined by a conjunction; they are not claimed for one joint event.
  • The paper assumes βm\beta_mβm​ is non-increasing in mmm and bounds βm−1\beta_{m-1}βm−1​ by βm\beta_mβm​ (p. 504). The leave-one-out bound (12) and its milestones carry the explicit hypothesis of uniform stability β\betaβ at size m−1m-1m−1; the empirical bound (11) does not.
  • McDiarmid's inequality sums ci2c_i^2ci2​ over i=1,…,mi = 1, \dots, mi=1,…,m; the paper's printed upper index nnn is a slip.
  • When a displayed tail bound has a zero denominator, its formal statement uses the limiting bound 000. In McDiarmid's inequality this is the constant-function case; in the stability tails the loss is identically zero.

The stability notion is the paper's remove-one notion. A formalization with replace-one stability would prove a different theorem with different constants, and the published FoundationsML_Stability_UniformlyStable (replace-one) is therefore not used.

The development needs the published loss, empirical-error and generalization-error definitions from Foundations of Machine Learning, a McDiarmid inequality on product measures (reusable for any bounded-differences argument), and the coordinate-exchange lemmas for Measure.pi behind Lemma 7. Contributions of a general McDiarmid inequality, of exchangeability lemmas for product measures, and of proofs of any milestone are welcome.

Selected references

  • O. Bousquet and A. Elisseeff, Stability and Generalization, Journal of Machine Learning Research 2 (2002) 499–526. https://www.jmlr.org/papers/v2/bousquet02a.html
  • C. McDiarmid, On the method of bounded differences, Surveys in Combinatorics, LMS Lecture Note Series 141 (1989) 148–188. https://doi.org/10.1017/CBO9781107359949.008
  • L. Devroye and T. Wagner, Distribution-free performance bounds for potential function rules, IEEE Trans. Inform. Theory 25 (1979) 601–604. https://doi.org/10.1109/TIT.1979.1056087
  • M. Kearns and D. Ron, Algorithmic stability and sanity-check bounds for leave-one-out cross-validation, Neural Computation 11 (1999) 1427–1453. https://doi.org/10.1162/089976699300016304
  • M. Hardt, B. Recht and Y. Singer, Train faster, generalize better: stability of stochastic gradient descent, ICML 2016. https://arxiv.org/abs/1509.01240
  • V. Feldman and J. Vondrák, High probability generalization bounds for uniformly stable algorithms with nearly optimal rate, COLT 2019. https://arxiv.org/abs/1902.10710
  • O. Bousquet, Y. Klochkov and N. Zhivotovskiy, Sharper bounds for uniformly stable algorithms, COLT 2020. https://arxiv.org/abs/1910.07833
  • M. Mohri, A. Rostamizadeh and A. Talwalkar, Foundations of Machine Learning, 2nd ed., MIT Press, 2018, Chapter 14.
13 thms2 active usersReviewed
ProbabilityReinforcement Learning·Captain: mikedeng1

Minimax Regret Bounds for Reinforcement Learning II: High-Probability Regret Bound for UCBVI with a Bernstein–Freedman BonusResearch Paper

Why finite-horizon reinforcement learning needs a variance-sensitive bound

An agent can learn to act in an unknown environment by repeatedly running a finite episode, observing the states reached after its actions, and updating its model of the environment. The agent must trade off rewards in the current episode against information that may improve later decisions. A regret bound measures the cumulative value lost relative to an optimal policy that knows the true transition probabilities. Its dependence on the number of states, actions, episode steps, and interactions says how much exploration that uncertainty can force.

Azar, Osband, and Munos study this question for a finite-horizon Markov decision process with known, bounded rewards and an unknown, stationary transition kernel. Their UCBVI algorithm estimates action values from observed transitions and adds an exploration bonus. Their second version, UCBVI-BF, uses the empirical variance of the next-state value in that bonus. Their Theorem 2 gives an explicit high-probability regret bound whose leading dependence on the horizon is smaller than the bound they give for the simpler UCBVI-CH bonus. The paper states that, in a sufficiently long-run regime, its leading order matches the cited lower-bound scale up to logarithmic factors. This mission targets the explicit theorem, including its lower-order terms, rather than only that asymptotic comparison.

The MDP, interaction, and algorithm

Let S\mathcal SS and A\mathcal AA be nonempty finite state and action sets with cardinalities SSS and AAA. A stationary transition kernel P(y∣x,a)P(y\mid x,a)P(y∣x,a) is a probability distribution on next states yyy for every current state xxx and action aaa. The reward R(x,a)R(x,a)R(x,a) is deterministic, known to the learner, and lies in [0,1][0,1][0,1]. These are the conditions of Assumption 1 and §2. Episodes have H≥1H\ge1H≥1 steps; KKK episodes comprise T=KHT=KHT=KH interactions.

A policy π\piπ chooses an action for each state and step. Its value Vhπ(x)V_h^\pi(x)Vhπ​(x) is the expected reward from step hhh through the final step when the state at hhh is xxx; the terminal value is VH+1π=0V_{H+1}^\pi=0VH+1π​=0. The optimal value Vh∗(x)=sup⁡πVhπ(x)V_h^*(x)=\sup_\pi V_h^\pi(x)Vh∗​(x)=supπ​Vhπ​(x) ranges over all deterministic policies of this form. At the start of episode kkk, the environment may choose the initial state using the completed episodes. The learner then fixes a policy πk\pi_kπk​, observes transitions during the episode, and updates counts for the next episode. Its regret is

Regret⁡(K)=∑k=1K(V1∗(xk,1)−V1πk(xk,1)).\operatorname{Regret}(K)=\sum_{k=1}^{K}\bigl(V_1^*(x_{k,1})-V_1^{\pi_k}(x_{k,1})\bigr).Regret(K)=k=1∑K​(V1∗​(xk,1​)−V1πk​​(xk,1​)).

For each state-action pair, Nk(x,a,y)N_k(x,a,y)Nk​(x,a,y) counts transitions to yyy in episodes before kkk, and Nk(x,a)=∑yNk(x,a,y)N_k(x,a)=\sum_yN_k(x,a,y)Nk​(x,a)=∑y​Nk​(x,a,y). When the latter is positive, P^k(y∣x,a)=Nk(x,a,y)/Nk(x,a)\widehat P_k(y\mid x,a)=N_k(x,a,y)/N_k(x,a)Pk​(y∣x,a)=Nk​(x,a,y)/Nk​(x,a). The count Nk,h′(y)N'_{k,h}(y)Nk,h′​(y) records previous episodes whose state at step hhh was yyy. Algorithms 2 and 4 compute optimistic Qk,hQ_{k,h}Qk,h​ backward from zero terminal value, take a minimum with the previous episode's QQQ estimate and with HHH, and choose a maximizing action at every state. Previously unseen pairs receive Qk,h=HQ_{k,h}=HQk,h​=H. The Bernstein–Freedman bonus uses the empirical variance of Vk,h+1V_{k,h+1}Vk,h+1​ under P^k\widehat P_kPk​ and an additional term based on Nk,h+1′N'_{k,h+1}Nk,h+1′​; the algorithm uses Lalg=ln⁡(5SAT/δ)L_{\rm alg}=\ln(5SAT/\delta)Lalg​=ln(5SAT/δ).

Formalization targets

The goal is Theorem 2 on p. 5. For any MDP and interaction described above and every δ>0\delta>0δ>0, write L=ln⁡(5HSAT/δ)L=\ln(5HSAT/\delta)L=ln(5HSAT/δ). The target is the exact bad-event form of the printed high-probability bound:

Pr⁡ ⁣{Regret⁡(K)>30HLSAK+2500H2S2AL2+4H3/2KL}≤δ.\Pr\!\left\{\operatorname{Regret}(K)>30HL\sqrt{SAK}+2500H^2S^2AL^2+4H^{3/2}\sqrt{KL}\right\}\le\delta.Pr{Regret(K)>30HLSAK​+2500H2S2AL2+4H3/2KL​}≤δ.

The milestone list contains three empirical-transition deviations from the proof of Lemma 1: Eq. (9) for a value-weighted transition error, the displayed count bound before Eq. (11), and Eq. (12) for the full transition row's ℓ1\ell_1ℓ1​ error. It also contains Lemma 2's variance comparison and Eq. (26), which relates cumulative conditional next-value variance to the variance of an episode return. These are source-indexed targets, with their printed constants retained.

What the result and its formalization supply

The theorem gives a quantitative guarantee for a particular executable decision rule: its regret grows sublinearly in KKK in the leading term, with explicit dependence on SSS, AAA, and HHH. The result lets one compare the horizon dependence of a variance-sensitive bonus with a value-agnostic bonus under the same finite-horizon model. It also fixes which logarithm belongs in the algorithm and which appears in the reported bound; replacing either changes the claim.

A formal proof would connect a fully specified adaptive interaction to its finite probability law, empirical counts, backward value iteration, and the stated high-probability conclusion. The local prior-art search found reusable transition-kernel vocabulary and general concentration tools, but no published formal statement of this exact UCBVI-BF algorithm or theorem. The mission's finite path and variance definitions can also support other episodic reinforcement-learning bounds that use conditional variance.

Where the difficulty lies

The bonus is computed using a value function that itself depends on earlier observations and the same episode's backward recursion. A concentration inequality for a fixed transition row and a fixed test function therefore does not directly control every value estimate encountered by the algorithm. The number of samples in a row is also random and changes with the learner's past actions. The regret compares a policy's value at an environment-chosen initial state with a supremum over all policies, while the learner's greedy action must be defined at states it never visits. These dependencies are the central obstacle to turning local concentration statements into the episode-level bound.

Formalization scope and conventions

The Lean model uses finite sums rather than measure theory. A published predicate supplies the stationary, real-valued transition kernel; a local MDP adds the known deterministic reward. State and action types are finite and nonempty. Policies are deterministic and depend on the step. The supremum defining V∗V^*V∗ ranges over their finite function type. A theorem quantifies over every maximizing tie-breaking rule and every initial-state rule that reads only completed episodes. The probability of an event is constructed as a sum over finite outcome sequences, each weighted by the product of true transition probabilities. Counts use all past transitions and no current or future outcomes. These choices rule out a trivialization that assumes the desired law or optimizes over an unbounded class of arbitrary functions.

Lean indexes the HHH steps from zero, while the paper indexes them from one. The last observed next state is kept because Algorithm 4 counts states at the terminal index H+1H+1H+1. At Nk,h+1′(y)=0N'_{k,h+1}(y)=0Nk,h+1′​(y)=0, Algorithm 4's quotient is interpreted as infinite and the capped term is H2H^2H2; Lean's ordinary division by zero would incorrectly produce zero. The algorithm uses Lalg=ln⁡(5SAT/δ)L_{\rm alg}=\ln(5SAT/\delta)Lalg​=ln(5SAT/δ), while Theorem 2's bound uses L=ln⁡(5HSAT/δ)L=\ln(5HSAT/\delta)L=ln(5HSAT/δ). For Eq. (26), the appendix ends its sums at H−1H-1H−1 under a shifted terminal convention; the local statement includes all HHH reward steps and the terminal value VH+1=0V_{H+1}=0VH+1​=0 used by Algorithm 2. The milestone text remains the printed text. The count milestone is the display before Eq. (11), since Eq. (11) drops a factor of 222 under the square root present in that display.

Theorem 2 retains its printed 2500H2S2AL22500H^2S^2AL^22500H2S2AL2 term. The appendix's displayed Lemma 13 calculation does not reproduce that second-order constant when propagated to Lemma 14; this is a source proof gap, not a hypothesis of the theorem. Work on the probability normalization, random-count concentration, adaptive value estimates, variance identity, and a valid route to the printed explicit constants is welcome. A proof with altered constants or an asymptotic-only conclusion would be a different target.

Selected references

  • M. G. Azar, I. Osband, and R. Munos, Minimax Regret Bounds for Reinforcement Learning, arXiv:1703.05449v2, 2017. Preprint.
13 thms2 active usersReviewed
Linear algebraOptimizationReinforcement Learning·Captain: mikedeng1

Reinforcement Learning: An Introduction X: Off-policy Divergence and the Gradient of the Projected Bellman ErrorTextbook

Motivation

Off-policy learning estimates the value function of a target policy π\piπ from data generated by a different behavior policy bbb. It is how an agent learns about a greedy policy while exploring, and how many policies can be evaluated from one stream of experience. Combined with linear function approximation and bootstrapping (updating an estimate toward other estimates, as temporal-difference methods do), off-policy learning can be unstable: the weights can diverge even when every quantity involved is well defined. Sutton and Barto call this combination the deadly triad (Chapter 11 of Reinforcement Learning: An Introduction, 2nd ed., 2018).

Chapter 11 does two things. It exhibits the instability with small, fully computable examples, and it identifies an objective that can be minimized stably from off-policy data: the mean square projected Bellman error (PBE), whose gradient the Gradient-TD methods (GTD2, TDC) follow in expectation. This mission formalizes the chapter's exact, finite-dimensional claims.

A short history, following the book's bibliographical remarks (pp. 285–286): Baird (1995) gave the seven-state counterexample for off-policy semi-gradient TD; Tsitsiklis and Van Roy (1996) gave the earliest w-to-2w example and the counterexample of Example 11.1, showing that even a least-squares fit at every step can diverge; Gradient-TD methods, which follow the gradient of the PBE, were introduced by Sutton, Szepesvári and Maei and by Sutton et al. (2009); Sutton, Mahmood and White (2016) introduced Emphatic-TD. The learnability discussion of §11.6 is the book's own.

Setting

A finite Markov decision process has states S\mathcal SS, actions A\mathcal AA, a finite reward set R\mathcal RR and dynamics p(s′,r∣s,a)p(s',r\mid s,a)p(s′,r∣s,a). A policy π(a∣s)\pi(a\mid s)π(a∣s) induces the transition matrix Pπ(s,s′)=∑aπ(a∣s) p(s′∣s,a)P_\pi(s,s') = \sum_a \pi(a\mid s)\,p(s'\mid s,a)Pπ​(s,s′)=∑a​π(a∣s)p(s′∣s,a) and expected reward rπ(s)r_\pi(s)rπ​(s). For 0≤γ<10\le\gamma<10≤γ<1 the true value function is the expected discounted return vπ(s)=∑k≥0γk(Pπkrπ)(s)v_\pi(s) = \sum_{k\ge0}\gamma^k (P_\pi^k r_\pi)(s)vπ​(s)=∑k≥0​γk(Pπk​rπ​)(s).

A state weighting μ\muμ is a probability distribution on S\mathcal SS, with D=diag⁡(μ)\mathbf D = \operatorname{diag}(\mu)D=diag(μ) and norm ∥v∥μ2=∑sμ(s)v(s)2\|v\|_\mu^2 = \sum_s \mu(s)v(s)^2∥v∥μ2​=∑s​μ(s)v(s)2. The ∣S∣×d|\mathcal S|\times d∣S∣×d matrix X\mathbf XX has the feature vectors x(s)⊤\mathbf x(s)^\topx(s)⊤ as rows, and a weight vector w∈Rd\mathbf w\in\mathbb R^dw∈Rd gives the linear value function vw=Xwv_{\mathbf w} = \mathbf X\mathbf wvw​=Xw. The Bellman operator is

(Bπv)(s)=∑aπ(a∣s)∑s′,rp(s′,r∣s,a) [r+γv(s′)].(B_\pi v)(s) = \sum_a \pi(a\mid s)\sum_{s',r} p(s',r\mid s,a)\,[r+\gamma v(s')].(Bπ​v)(s)=a∑​π(a∣s)s′,r∑​p(s′,r∣s,a)[r+γv(s′)].

The Bellman error vector is δˉw=Bπvw−vw\bar\delta_{\mathbf w} = B_\pi v_{\mathbf w} - v_{\mathbf w}δˉw​=Bπ​vw​−vw​, the projection is Π=X(X⊤DX)−1X⊤D\Pi = \mathbf X(\mathbf X^\top\mathbf D\mathbf X)^{-1}\mathbf X^\top\mathbf DΠ=X(X⊤DX)−1X⊤D, and the two objectives are

BE(w)=∥δˉw∥μ2,PBE(w)=∥Πδˉw∥μ2.\mathrm{BE}(\mathbf w) = \|\bar\delta_{\mathbf w}\|_\mu^2,\qquad \mathrm{PBE}(\mathbf w) = \|\Pi\bar\delta_{\mathbf w}\|_\mu^2 .BE(w)=∥δˉw​∥μ2​,PBE(w)=∥Πδˉw​∥μ2​.

Off-policy samples are reweighted by the importance-sampling ratio ρt=π(At∣St)/b(At∣St)\rho_t = \pi(A_t\mid S_t)/b(A_t\mid S_t)ρt​=π(At​∣St​)/b(At​∣St​).

Formalization targets

Goal: the gradient of the PBE

If X⊤DX\mathbf X^\top\mathbf D\mathbf XX⊤DX is invertible, then for every w\mathbf ww

PBE(w)=(X⊤Dδˉw)⊤(X⊤DX)−1(X⊤Dδˉw),\mathrm{PBE}(\mathbf w) = (\mathbf X^\top\mathbf D\bar\delta_{\mathbf w})^\top(\mathbf X^\top\mathbf D\mathbf X)^{-1}(\mathbf X^\top\mathbf D\bar\delta_{\mathbf w}),PBE(w)=(X⊤Dδˉw​)⊤(X⊤DX)−1(X⊤Dδˉw​), ∇PBE(w)=2 (γPπX−X)⊤DX (X⊤DX)−1 X⊤Dδˉw.\nabla\mathrm{PBE}(\mathbf w) = 2\,(\gamma P_\pi\mathbf X-\mathbf X)^\top\mathbf D\mathbf X\,(\mathbf X^\top\mathbf D\mathbf X)^{-1}\,\mathbf X^\top\mathbf D\bar\delta_{\mathbf w}.∇PBE(w)=2(γPπ​X−X)⊤DX(X⊤DX)−1X⊤Dδˉw​.

With μ\muμ the state distribution under bbb this is (11.27), ∇PBE(w)=2 E[ρt(γxt+1−xt)xt⊤] E[xtxt⊤]−1 E[ρtδtxt]\nabla\mathrm{PBE}(\mathbf w) = 2\,\mathbb E[\rho_t(\gamma\mathbf x_{t+1}-\mathbf x_t)\mathbf x_t^\top]\,\mathbb E[\mathbf x_t\mathbf x_t^\top]^{-1}\,\mathbb E[\rho_t\delta_t\mathbf x_t]∇PBE(w)=2E[ρt​(γxt+1​−xt​)xt⊤​]E[xt​xt⊤​]−1E[ρt​δt​xt​].

Milestones, in attack order

  1. The w-to-2w example (p. 260): repeated off-policy semi-gradient TD(0) on one transition multiplies www by 1+α(2γ−1)1+\alpha(2\gamma-1)1+α(2γ−1), so wt→±∞w_t\to\pm\inftywt​→±∞ for every α>0\alpha>0α>0 once γ>12\gamma>\tfrac12γ>21​.
  2. Example 11.1, (11.10) (p. 263): the least-squares iteration wk+1=6−4ε5γwkw_{k+1} = \tfrac{6-4\varepsilon}{5}\gamma w_kwk+1​=56−4ε​γwk​ diverges when γ>5/(6−4ε)\gamma>5/(6-4\varepsilon)γ>5/(6−4ε) and w0≠0w_0\ne0w0​=0.
  3. (11.12)–(11.13): Πv\Pi vΠv is the unique μ\muμ-closest representable function, and Π⊤DΠ=DX(X⊤DX)−1X⊤D\Pi^\top\mathbf D\Pi = \mathbf D\mathbf X(\mathbf X^\top\mathbf D\mathbf X)^{-1}\mathbf X^\top\mathbf DΠ⊤DΠ=DX(X⊤DX)−1X⊤D.
  4. (11.21): vπv_\pivπ​ is the unique fixed point of BπB_\piBπ​ (γ<1\gamma<1γ<1).
  5. (11.24), Exercise 11.4: RE(w)=VE(w)+E[(Gt−vπ(St))2]\mathrm{RE}(\mathbf w) = \mathrm{VE}(\mathbf w) + \mathbb E[(G_t - v_\pi(S_t))^2]RE(w)=VE(w)+E[(Gt​−vπ​(St​))2] in the on-policy case.
  6. Example 11.4 (p. 276): two Markov reward processes that generate the same observable data distribution (every finite prefix of the stream of feature vectors and rewards has the same probability, each process started from its stationary distribution) have BE(0)=0\mathrm{BE}(\mathbf 0) = 0BE(0)=0 and BE(0)=23\mathrm{BE}(\mathbf 0) = \tfrac23BE(0)=32​.
  7. (11.25)–(11.26): the PBE in matrix terms.
  8. The three factors (p. 278): X⊤Dδˉw=E[ρtδtxt]\mathbf X^\top\mathbf D\bar\delta_{\mathbf w} = \mathbb E[\rho_t\delta_t\mathbf x_t]X⊤Dδˉw​=E[ρt​δt​xt​], (γPπX−X)⊤DX=E[ρt(γxt+1−xt)xt⊤](\gamma P_\pi\mathbf X-\mathbf X)^\top\mathbf D\mathbf X = \mathbb E[\rho_t(\gamma\mathbf x_{t+1}-\mathbf x_t)\mathbf x_t^\top](γPπ​X−X)⊤DX=E[ρt​(γxt+1​−xt​)xt⊤​], X⊤DX=E[xtxt⊤]\mathbf X^\top\mathbf D\mathbf X = \mathbb E[\mathbf x_t\mathbf x_t^\top]X⊤DX=E[xt​xt⊤​], under coverage.

Significance

The results. The divergence examples show that no step-size choice rescues semi-gradient TD off-policy, and that exact least-squares fitting does not either. The learnability results separate objectives that can be estimated from observed features and rewards (RE, PBE) from one that cannot (BE), and (11.24) shows that the unobservable VE has the same minimizer as the observable RE. The gradient formula (11.27) is the expected update of the Gradient-TD family; its factorization into three expectations is what makes an O(d)O(d)O(d) stochastic-gradient method possible.

Formalizing them. All of these are proved or computed in the book, in informal notation that leaves hypotheses implicit (invertibility, coverage, the distribution μ\muμ, the domain of γ\gammaγ). A machine-checked version fixes them, and gives a reusable layer of linear value-function geometry (Bellman operator, μ\muμ-norm, projection, BE, PBE, behavior-policy expectations) for later work on Gradient-TD convergence and on the TD fixed point. No formalization of this chapter is known to exist.

Difficulty

The individual steps are finite linear algebra, but they are easy to get wrong. The gradient of a quadratic form h⊤C−1h\mathbf h^\top\mathbf C^{-1}\mathbf hh⊤C−1h in an affine h(w)\mathbf h(\mathbf w)h(w) produces C−1+C−⊤\mathbf C^{-1}+\mathbf C^{-\top}C−1+C−⊤, and collapsing it to 2C−12\mathbf C^{-1}2C−1 uses the symmetry of X⊤DX\mathbf X^\top\mathbf D\mathbf XX⊤DX. The Jacobian of w↦X⊤Dδˉw\mathbf w\mapsto\mathbf X^\top\mathbf D\bar\delta_{\mathbf w}w↦X⊤Dδˉw​ has to be identified through the Bellman operator's affine form rπ+γPπvr_\pi+\gamma P_\pi vrπ​+γPπ​v, which is a theorem about the four-argument dynamics, not a definition. Passing from matrix to expectation form needs coverage, since otherwise b(a∣s)ρ=π(a∣s)b(a\mid s)\rho = \pi(a\mid s)b(a∣s)ρ=π(a∣s) fails. The divergence examples need the conclusion ∣wk∣→∞|w_k|\to\infty∣wk​∣→∞, not merely "the multiplier exceeds one".

Formalization scope

States are a finite type; values are functions S→R\mathcal S\to\mathbb RS→R; weights are Rd\mathbb R^dRd as Fin d → ℝ; matrices are Mathlib Matrix. The dynamics are the book's p(s′,r∣s,a)p(s',r\mid s,a)p(s′,r∣s,a) with a finite reward set, and one action type serves all states. μ\muμ is a probability distribution (nonnegative, summing to one). vπv_\pivπ​ is defined from expected returns, not as a Bellman solution. The gradient is a Fréchet derivative whose linear map is u↦g⊤u\mathbf u\mapsto\mathbf g^\top\mathbf uu↦g⊤u.

Explicit hypotheses the book leaves implicit:

  • X⊤DX\mathbf X^\top\mathbf D\mathbf XX⊤DX invertible. The book substitutes a pseudoinverse otherwise; that case is not formalized, and without the hypothesis Lean's matrix inverse is zero and the PBE is trivially 000.
  • Coverage (π(a∣s)>0⇒b(a∣s)>0\pi(a\mid s)>0\Rightarrow b(a\mid s)>0π(a∣s)>0⇒b(a∣s)>0) for the expectation identities.
  • 0≤γ<10\le\gamma<10≤γ<1 for (11.21); the episodic γ=1\gamma=1γ=1 case is not stated.
  • In (11.24) the return enters through conditional distributions νs\nu_sνs​ with finite second moment and mean vπ(s)v_\pi(s)vπ​(s).

Readings and corrections:

  • (11.10) is introduced as minimizing "the VE", but its displayed objective is an unweighted sum over the two states. The displayed sum is formalized.
  • The sentence on p. 269, "there always exists an approximate value function with zero PBE", needs X⊤D(I−γPπ)X\mathbf X^\top\mathbf D(I-\gamma P_\pi)\mathbf XX⊤D(I−γPπ​)X invertible, which can fail off-policy. Counterexample: two states with features x=1,2x = 1, 2x=1,2, both moving to the second state under π\piπ, rewards rπ=(1,0)r_\pi = (1, 0)rπ​=(1,0), γ=34\gamma = \tfrac34γ=43​ and μ=(23,13)\mu = (\tfrac23, \tfrac13)μ=(32​,31​). Then X⊤Dδˉw=23\mathbf X^\top\mathbf D\bar\delta_{\mathbf w} = \tfrac23X⊤Dδˉw​=32​ for every w\mathbf ww, so PBE(w)=(23)2/2>0\mathrm{PBE}(\mathbf w) = (\tfrac23)^2/2 > 0PBE(w)=(32​)2/2>0 everywhere. The sentence is not formalized.
  • The Example 11.4 MRPs are transcribed from the figure on p. 276. Their equal data distributions are stated through all finite prefixes of the observed stream (feature vectors and rewards) from the stationary distributions; the BE minimizer of the second MRP as γ→1\gamma\to1γ→1 is not formalized.
  • Baird's counterexample (divergence shown by simulation), Gradient-TD and Emphatic-TD convergence (asserted with references) are out of scope.

A trivializing formalization is ruled out: vπv_\pivπ​ is not defined as a fixed point of BπB_\piBπ​, the PBE is not stated without the invertibility hypothesis, and the gradient is a derivative of the defined PBE, not a restated formula.

Welcome contributions: proofs of the milestones, and a general lemma on the gradient of h⊤C−1h\mathbf h^\top\mathbf C^{-1}\mathbf hh⊤C−1h for affine h\mathbf hh and symmetric invertible C\mathbf CC, which is reusable well beyond this mission.

Selected references

  • R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018, ISBN 9780262039246, Chapter 11. http://incompleteideas.net/book/the-book-2nd.html
  • L. Baird, Residual algorithms: Reinforcement learning with function approximation, ICML 1995. https://doi.org/10.1016/B978-1-55860-377-6.50013-X
  • J. N. Tsitsiklis and B. Van Roy, Feature-based methods for large scale dynamic programming, Machine Learning 22, 1996. https://doi.org/10.1007/BF00114724
  • R. S. Sutton, H. R. Maei, D. Precup, S. Bhatnagar, D. Silver, Cs. Szepesvári, E. Wiewiora, Fast gradient-descent methods for temporal-difference learning with linear function approximation, ICML 2009. https://doi.org/10.1145/1553374.1553501
  • R. S. Sutton, A. R. Mahmood, M. White, An emphatic approach to the problem of off-policy temporal-difference learning, JMLR 17, 2016. https://jmlr.org/papers/v17/14-488.html
14 thms2 active usersReviewed
Numerical AnalysisOptimization·Captain: mikedeng1

Gradient Convergence in Gradient Methods with Errors I: With Deterministic Errors Proportional to the Stepsize, Either f(x_t) → −∞ or f(x_t) Converges and ∇f(x_t) → 0Research Paper

Motivation

Gradient methods are the workhorse of large-scale nonlinear optimization and of the training of statistical models. In practice the direction actually used is rarely the exact negative gradient: it may be scaled, computed incrementally one data component at a time, or perturbed by approximation error. The classical convergence theory of such methods often assumes that the iterates stay bounded, that the objective is bounded below, or that the errors vanish at a prescribed rate, and these assumptions must then be checked separately for each method.

Bertsekas and Tsitsiklis (2000) proved convergence results for gradient methods with errors that need none of these assumptions. Their deterministic result (Proposition 1) allows a general descent direction together with an error whose size is proportional to the stepsize, and concludes that either the objective values diverge to −∞-\infty−∞ or they converge and the gradients tend to zero. It applies, among others, to the incremental gradient method for a sum of functions (Proposition 2 of the same paper), which underlies backpropagation-style training. The stochastic counterpart (Proposition 3, zero-mean errors) is the subject of a companion mission.

Setting

Throughout, f:Rn→Rf:\mathbb R^n\to\mathbb Rf:Rn→R is a continuously differentiable function whose gradient is Lipschitz continuous: for some constant LLL,

∥∇f(x)−∇f(xˉ)∥≤L∥x−xˉ∥∀x,xˉ∈Rn.(2.1)\|\nabla f(x)-\nabla f(\bar x)\|\le L\|x-\bar x\|\qquad\forall x,\bar x\in\mathbb R^n. \tag{2.1}∥∇f(x)−∇f(xˉ)∥≤L∥x−xˉ∥∀x,xˉ∈Rn.(2.1)

Here ∥⋅∥\|\cdot\|∥⋅∥ is the Euclidean norm and x′yx'yx′y the standard inner product. The gradient method with errors generates a sequence of iterates

xt+1=xt+γt(st+wt),t=0,1,…,x_{t+1}=x_t+\gamma_t(s_t+w_t),\qquad t=0,1,\dots,xt+1​=xt​+γt​(st​+wt​),t=0,1,…,

where γt>0\gamma_t>0γt​>0 is the stepsize, sts_tst​ is a descent direction and wtw_twt​ is an error vector. Nothing is assumed about how sts_tst​ and wtw_twt​ are produced, beyond the two conditions below, which hold for some positive scalars c1,c2,p,qc_1,c_2,p,qc1​,c2​,p,q and every ttt:

c1∥∇f(xt)∥2≤−∇f(xt)′st,∥st∥≤c2(1+∥∇f(xt)∥),(2.2)c_1\|\nabla f(x_t)\|^2\le-\nabla f(x_t)'s_t,\qquad\|s_t\|\le c_2\bigl(1+\|\nabla f(x_t)\|\bigr), \tag{2.2}c1​∥∇f(xt​)∥2≤−∇f(xt​)′st​,∥st​∥≤c2​(1+∥∇f(xt​)∥),(2.2) ∥wt∥≤γt(q+p∥∇f(xt)∥).(2.3)\|w_t\|\le\gamma_t\bigl(q+p\|\nabla f(x_t)\|\bigr). \tag{2.3}∥wt​∥≤γt​(q+p∥∇f(xt​)∥).(2.3)

The stepsizes are diminishing in the standard sense:

∑t=0∞γt=∞,∑t=0∞γt2<∞.\sum_{t=0}^\infty\gamma_t=\infty,\qquad\sum_{t=0}^\infty\gamma_t^2<\infty.t=0∑∞​γt​=∞,t=0∑∞​γt2​<∞.

A stationary point of fff is a point xˉ\bar xxˉ with ∇f(xˉ)=0\nabla f(\bar x)=0∇f(xˉ)=0; a limit point of (xt)(x_t)(xt​) is the limit of some subsequence.

Formalization targets

Goal: Proposition 1 (p. 630)

Under (2.1), (2.2), (2.3) and the stepsize conditions, either

f(xt)→−∞,f(x_t)\to-\infty,f(xt​)→−∞,

or else f(xt)f(x_t)f(xt​) converges to a finite value and

lim⁡t→∞∇f(xt)=0.\lim_{t\to\infty}\nabla f(x_t)=0.t→∞lim​∇f(xt​)=0.

Furthermore, every limit point of (xt)(x_t)(xt​) is a stationary point of fff.

Milestones

The milestones follow the paper's own argument, in order.

  1. Lemma 1 (p. 629). For real sequences with Wt≥0W_t\ge0Wt​≥0, Yt+1≤Yt−Wt+ZtY_{t+1}\le Y_t-W_t+Z_tYt+1​≤Yt​−Wt​+Zt​ and ∑t=0TZt\sum_{t=0}^T Z_t∑t=0T​Zt​ convergent, either Yt→−∞Y_t\to-\inftyYt​→−∞, or YtY_tYt​ converges and ∑tWt<∞\sum_t W_t<\infty∑t​Wt​<∞.
  2. (2.4) (p. 630). Under (2.1), f(x+z)≤f(x)+z′∇f(x)+L2∥z∥2f(x+z)\le f(x)+z'\nabla f(x)+\tfrac L2\|z\|^2f(x+z)≤f(x)+z′∇f(x)+2L​∥z∥2 for all x,zx,zx,z.
  3. (2.5) (p. 631). For some β1,β2>0\beta_1,\beta_2>0β1​,β2​>0 and all sufficiently large ttt, f(xt+1)≤f(xt)−γtβ1∥∇f(xt)∥2+γt2β2f(x_{t+1})\le f(x_t)-\gamma_t\beta_1\|\nabla f(x_t)\|^2+\gamma_t^2\beta_2f(xt+1​)≤f(xt​)−γt​β1​∥∇f(xt​)∥2+γt2​β2​.
  4. (2.6) (p. 631). Either f(xt)→−∞f(x_t)\to-\inftyf(xt​)→−∞, or f(xt)f(x_t)f(xt​) converges and ∑tγt∥∇f(xt)∥2<∞\sum_t\gamma_t\|\nabla f(x_t)\|^2<\infty∑t​γt​∥∇f(xt​)∥2<∞.
  5. After (2.6) (p. 631). If f(xt)↛−∞f(x_t)\not\to-\inftyf(xt​)→−∞, then lim inf⁡t→∞∥∇f(xt)∥=0\liminf_{t\to\infty}\|\nabla f(x_t)\|=0liminft→∞​∥∇f(xt​)∥=0.

Significance

The result. Proposition 1 separates two concerns that are usually entangled: what the method guarantees, and what must be known about fff. It concludes stationarity of all limit points and convergence of the gradients to zero without assuming that the iterates are bounded or that fff is bounded below; when fff is bounded below the first alternative is excluded and ∇f(xt)→0\nabla f(x_t)\to0∇f(xt​)→0 follows outright. Because sts_tst​ need not be the negative gradient and wtw_twt​ need not vanish faster than the stepsize, the result covers scaled gradient methods, incremental gradient methods for sums of functions, and gradient methods with deterministic approximation error. The descent inequality (2.4) and the deterministic supermartingale-type Lemma 1 are standard tools that recur throughout optimization theory.

Formalizing it. The proposition is proved in the paper; it has not been machine-checked. A formal proof would give a reusable, verified convergence theorem for a broad class of first-order methods on Rn\mathbb R^nRn, together with a formal descent lemma for functions with Lipschitz gradient, which Mathlib does not currently state in this form, and a deterministic Robbins–Siegmund-type lemma for sequences.

Difficulty

The summability estimate (2.6) gives only lim inf⁡∥∇f(xt)∥=0\liminf\|\nabla f(x_t)\|=0liminf∥∇f(xt​)∥=0. Passing to lim⁡∇f(xt)=0\lim\nabla f(x_t)=0lim∇f(xt​)=0 is the main step: the obvious argument (a summable series ∑tγt∥∇f(xt)∥2\sum_t\gamma_t\|\nabla f(x_t)\|^2∑t​γt​∥∇f(xt​)∥2 with ∑tγt=∞\sum_t\gamma_t=\infty∑t​γt​=∞ forces the gradient norms to zero) is false in general, because a nonnegative sequence with these two properties may still have infinitely many large terms. What is missing is a bound on how far ∥∇f(xt)∥\|\nabla f(x_t)\|∥∇f(xt​)∥ can travel while the stepsizes are small, and only (2.1) and (2.2)–(2.3) together supply it. A second difficulty is that there is no boundedness of the iterates: every estimate must hold globally, and the error wtw_twt​ is controlled only relative to ∥∇f(xt)∥\|\nabla f(x_t)\|∥∇f(xt​)∥, which may be unbounded along the sequence.

Formalization scope

  • Rn\mathbb R^nRn is EuclideanSpace ℝ (Fin n), x′yx'yx′y is the real inner product ⟪x, y⟫_ℝ, and ∇f\nabla f∇f is Mathlib's gradient f; the hypothesis ContDiff ℝ 1 f makes it the true gradient. The paper's statement is for Rn\mathbb R^nRn and the formalization does not generalize to Hilbert spaces.
  • The standing assumption (2.1) of §2 is part of every statement about fff, as LipschitzWith L (gradient f) with L : ℝ≥0; this is equivalent to (2.1) for some real constant.
  • The sequences xt,st,wtx_t,s_t,w_txt​,st​,wt​ and γt\gamma_tγt​ are arbitrary data indexed from t=0t=0t=0, constrained only by the recursion and by (2.2), (2.3), γt>0\gamma_t>0γt​>0 (all four constants c1,c2,p,qc_1,c_2,p,qc1​,c2​,p,q are positive, as printed).
  • ∑tγt=∞\sum_t\gamma_t=\infty∑t​γt​=∞ is divergence of the partial sums to +∞+\infty+∞; ∑tγt2<∞\sum_t\gamma_t^2<\infty∑t​γt2​<∞ is summability of nonnegative terms. In Lemma 1 the convergence of ∑tZt\sum_t Z_t∑t​Zt​ is convergence of the partial sums, not absolute convergence, since ZtZ_tZt​ may change sign.
  • lim⁡inf⁡\lim\infliminf is written out as "for every ϵ>0\epsilon>0ϵ>0, infinitely often below ϵ\epsilonϵ", avoiding junk values of a lim inf of an unbounded sequence. Limit points are cluster points of the sequence.
  • "Every limit point is stationary" is a separate conjunct, outside the dichotomy, exactly as on the page.
  • The constants β1,β2\beta_1,\beta_2β1​,β2​ in (2.5) are existential. The milestones (2.5) and (2.6) retain the stepsize hypotheses of Proposition 1, where the paper derives them.
  • A trivializing formalization would drop the "−∞-\infty−∞" alternative or require fff bounded below; neither is done. Instances satisfying all hypotheses with f(xt)→−∞f(x_t)\to-\inftyf(xt​)→−∞ and ∇f↛0\nabla f\not\to0∇f→0 exist (a linear fff), so the first alternative is genuinely needed.
  • Out of scope: Proposition 2 (the incremental gradient method of §3), which is a corollary of the goal, and the stochastic results of §4–§5.

Contributions welcome: proofs of the descent lemma and Lemma 1 (both reusable well beyond this mission), and of the steps (2.5)–(2.6) and the excursion argument.

Selected references

  • D. P. Bertsekas and J. N. Tsitsiklis, Gradient Convergence in Gradient Methods with Errors, SIAM Journal on Optimization 10(3):627–642, 2000. https://doi.org/10.1137/S1052623497331063
  • D. P. Bertsekas, Nonlinear Programming, 2nd ed., Athena Scientific, 1999.
  • H. Robbins and D. Siegmund, A convergence theorem for non negative almost supermartingales and some applications, in Optimizing Methods in Statistics, Academic Press, 1971, 233–257. https://doi.org/10.1016/B978-0-12-604550-5.50015-8
6 thms2 active usersReviewed
Linear algebraMarkov ChainReinforcement Learning·Captain: mikedeng1

Reinforcement Learning: An Introduction VIII: The TD Fixed Point of Linear Semi-gradient TD(0) and Its Error BoundTextbook

Motivation

Reinforcement learning methods estimate the value function vπv_\pivπ​ of a policy π\piπ: the expected discounted sum of future rewards from each state. When the state space is large, vπv_\pivπ​ cannot be stored as a table and is approximated by a parametrized function. The most studied case is linear function approximation, where each state sss carries a feature vector x(s)∈Rd\mathbf x(s) \in \mathbb R^dx(s)∈Rd and the estimate is v^(s,w)=w⊤x(s)\hat v(s, \mathbf w) = \mathbf w^\top \mathbf x(s)v^(s,w)=w⊤x(s). Temporal-difference learning with this approximation, linear semi-gradient TD(0), is one of the basic algorithms of the field, and Chapter 9 of Sutton and Barto's Reinforcement Learning: An Introduction (2nd ed., MIT Press, 2018) presents its analysis: where the algorithm can converge, why that point exists, and how good it is.

The history is short. Sutton (1988, doi:10.1007/BF00115009) introduced TD learning and showed positive definiteness of the matrix governing its expected update. Dayan (1992, doi:10.1007/BF00992701) extended convergence to TD(λ). Tsitsiklis and Van Roy (1997, doi:10.1109/9.580874) proved convergence with probability one for linear TD(λ) under on-policy sampling and bounded the error of the limit. Bradtke and Barto (1996) introduced least-squares TD (LSTD), which computes the same limit directly.

Setting

A finite Markov decision process has finite sets of states S\mathcal SS, actions A\mathcal AA and rewards R⊂R\mathcal R \subset \mathbb RR⊂R, and dynamics p(s′,r∣s,a)p(s', r \mid s, a)p(s′,r∣s,a), the probability of next state s′s's′ and reward rrr after action aaa in state sss. A policy π(a∣s)\pi(a \mid s)π(a∣s) is a probability distribution over actions for each state. It induces a Markov chain on states with transition matrix P\mathbf PP, P(s,s′)=p(s′∣s)=∑aπ(a∣s)∑rp(s′,r∣s,a)\mathbf P(s, s') = p(s' \mid s) = \sum_a \pi(a \mid s) \sum_r p(s', r \mid s, a)P(s,s′)=p(s′∣s)=∑a​π(a∣s)∑r​p(s′,r∣s,a), and expected one-step reward rπ(s)r_\pi(s)rπ​(s). For a discount rate 0≤γ<10 \le \gamma < 10≤γ<1, the true value is vπ(s)=∑k≥0γk(Pkrπ)(s)v_\pi(s) = \sum_{k \ge 0} \gamma^k (\mathbf P^k r_\pi)(s)vπ​(s)=∑k≥0​γk(Pkrπ​)(s), the expected discounted return.

A state distribution μ\muμ is stationary if μ⊤P=μ⊤\mu^\top \mathbf P = \mu^\topμ⊤P=μ⊤; write D=diag(μ)\mathbf D = \mathrm{diag}(\mu)D=diag(μ). The feature matrix X\mathbf XX is the ∣S∣×d|\mathcal S| \times d∣S∣×d matrix with rows x(s)\mathbf x(s)x(s). The mean square value error of a weight vector is

VE‾(w)=∑sμ(s) [vπ(s)−w⊤x(s)]2.\overline{\mathrm{VE}}(\mathbf w) = \sum_{s} \mu(s)\,[v_\pi(s) - \mathbf w^\top \mathbf x(s)]^2 .VE(w)=s∑​μ(s)[vπ​(s)−w⊤x(s)]2.

Linear semi-gradient TD(0) updates wt+1=wt+α(Rt+1+γwt⊤xt+1−wt⊤xt)xt\mathbf w_{t+1} = \mathbf w_t + \alpha(R_{t+1} + \gamma \mathbf w_t^\top \mathbf x_{t+1} - \mathbf w_t^\top \mathbf x_t)\mathbf x_twt+1​=wt​+α(Rt+1​+γwt⊤​xt+1​−wt⊤​xt​)xt​. In steady state its expected update involves

b=E[Rt+1xt],A=E[xt(xt−γxt+1)⊤],\mathbf b = \mathbb E[R_{t+1}\mathbf x_t], \qquad \mathbf A = \mathbb E[\mathbf x_t(\mathbf x_t - \gamma \mathbf x_{t+1})^\top],b=E[Rt+1​xt​],A=E[xt​(xt​−γxt+1​)⊤],

and the TD fixed point is wTD=A−1b\mathbf w_{\mathrm{TD}} = \mathbf A^{-1}\mathbf bwTD​=A−1b. A real square matrix MMM, not necessarily symmetric, is positive definite if y⊤My>0y^\top M y > 0y⊤My>0 for every y≠0y \ne 0y=0. The key matrix is D(I−γP)\mathbf D(\mathbf I - \gamma\mathbf P)D(I−γP).

Formalization targets

Goal: the TD fixed point exists and its error bound (9.12), (9.14)

Under the hypotheses above, with every μ(s)>0\mu(s) > 0μ(s)>0 and linearly independent feature columns, A\mathbf AA is invertible, b=AwTD\mathbf b = \mathbf A \mathbf w_{\mathrm{TD}}b=AwTD​, and

VE‾(wTD)≤11−γmin⁡wVE‾(w).\overline{\mathrm{VE}}(\mathbf w_{\mathrm{TD}}) \le \frac{1}{1-\gamma}\min_{\mathbf w} \overline{\mathrm{VE}}(\mathbf w).VE(wTD​)≤1−γ1​wmin​VE(w).

Milestones

  1. The expected update (9.13): E[wt+1∣wt]=(I−αA)wt+αb\mathbb E[\mathbf w_{t+1} \mid \mathbf w_t] = (\mathbf I - \alpha \mathbf A)\mathbf w_t + \alpha \mathbf bE[wt+1​∣wt​]=(I−αA)wt​+αb.
  2. The matrix form A=X⊤D(I−γP)X\mathbf A = \mathbf X^\top \mathbf D(\mathbf I - \gamma \mathbf P)\mathbf XA=X⊤D(I−γP)X.
  3. The criterion of Sutton (1988): positive diagonal, nonpositive off-diagonal entries, positive row sums and nonnegative column sums give positive definiteness.
  4. The column sums of the key matrix, 1⊤D(I−γP)=(1−γ)μ⊤\mathbf 1^\top \mathbf D(\mathbf I - \gamma \mathbf P) = (1-\gamma)\mu^\top1⊤D(I−γP)=(1−γ)μ⊤.
  5. The key matrix and A\mathbf AA are positive definite.
  6. A positive definite A\mathbf AA is invertible and A−1b\mathbf A^{-1}\mathbf bA−1b is the unique solution of b=Aw\mathbf b = \mathbf A \mathbf wb=Aw (9.12).
  7. The Sherman–Morrison update (9.22) of the LSTD inverse A^t−1\hat{\mathbf A}_t^{-1}A^t−1​.

Significance

Positive definiteness of A\mathbf AA is the reason on-policy linear TD(0) is stable: it makes the expected iteration contract toward the fixed point for small step sizes, and it guarantees that the fixed point exists and is unique. The error bound (9.14) quantifies the price of bootstrapping: the limit of TD can be worse than the best linear approximation, but by at most the factor 1/(1−γ)1/(1-\gamma)1/(1−γ). The same objects A\mathbf AA, b\mathbf bb and the key matrix reappear in LSTD, in the analysis of off-policy divergence (Chapter 11 of the book, where D\mathbf DD is no longer the stationary distribution of P\mathbf PP and positive definiteness fails), and in gradient-TD methods.

All results here are known. The book gives the positive definiteness argument in a box and cites (9.14) without proof. None of them has a machine-checked proof on the platform; the general Woodbury identity (FamousTheorems.woodbury_identity) is available, and (9.22) is its rank-one case written for the LSTD recursion. The mission produces a formal account of the finite-state theory of linear TD(0), with every hypothesis the book leaves implicit stated.

Difficulty

The key matrix D(I−γP)\mathbf D(\mathbf I - \gamma \mathbf P)D(I−γP) is not symmetric, so the usual tools for symmetric positive definite matrices do not apply directly, and A\mathbf AA is positive definite only because of the specific interplay between D\mathbf DD and P\mathbf PP: if μ\muμ is replaced by a non-stationary distribution the claim is false (this is the off-policy counterexample of Chapter 11). The error bound (9.14) is not a consequence of positive definiteness alone. The TD fixed point is not the minimizer of VE‾\overline{\mathrm{VE}}VE, and VE‾(wTD)\overline{\mathrm{VE}}(\mathbf w_{\mathrm{TD}})VE(wTD​) has to be compared with the error of the μ\muμ-weighted projection of vπv_\pivπ​, which requires controlling P\mathbf PP in the μ\muμ-weighted norm. The book gives no argument for this step.

Formalization scope

The Lean development lives in the namespace SuttonBartoRL.LinearTD. The MDP has four-argument dynamics p(s′,r∣s,a)p(s', r \mid s, a)p(s′,r∣s,a) with a finite reward set and one action set for all states; policies are stochastic. vπv_\pivπ​ is defined from expected discounted returns as the series ∑kγkPkrπ\sum_k \gamma^k \mathbf P^k r_\pi∑k​γkPkrπ​, never from a Bellman equation or from wTD\mathbf w_{\mathrm{TD}}wTD​. A\mathbf AA and b\mathbf bb are defined as the book's steady-state expectations (9.11), as finite sums over μ\muμ, π\piπ and ppp; the matrix form is a milestone, not a definition. Features are a matrix Matrix S (Fin d) ℝ with rows x(s)\mathbf x(s)x(s); linear independence of its columns is LinearIndependent ℝ Xᵀ. Positive definiteness is a custom predicate ∀y≠0, 0<y⊤My\forall y \ne 0,\ 0 < y^\top M y∀y=0, 0<y⊤My, not Mathlib's Matrix.PosDef, which requires symmetry. The minimum in (9.14) is expressed by quantifying over every w\mathbf ww. The matrix inverse is Mathlib's, which is zero on singular matrices; the goal therefore asserts invertibility of A\mathbf AA explicitly.

Hypotheses the book leaves implicit and the statements make explicit: 0≤γ<10 \le \gamma < 10≤γ<1 (the continuing case); μ\muμ a stationary distribution of the chain induced by π\piπ with μ(s)>0\mu(s) > 0μ(s)>0 for every sss (otherwise the key matrix is only positive semidefinite); linearly independent feature columns (the book's "degenerate cases", p. 205). The box calls the off-diagonal entries of the key matrix "negative"; they are zero wherever p(s′∣s)=0p(s' \mid s) = 0p(s′∣s)=0, so the criterion is stated with nonpositive entries. The book's sentence that εI\varepsilon\mathbf IεI "ensures that A^t\hat{\mathbf A}_tA^t​ is always invertible" (p. 229) is false in general, because the summands xk(xk−γxk+1)⊤\mathbf x_k(\mathbf x_k - \gamma \mathbf x_{k+1})^\topxk​(xk​−γxk+1​)⊤ are not positive semidefinite: with d=1d = 1d=1, ε=1/10\varepsilon = 1/10ε=1/10, γ=1/2\gamma = 1/2γ=1/2, x0=1\mathbf x_0 = 1x0​=1, x1=11/5\mathbf x_1 = 11/5x1​=11/5 one gets A^1=0\hat{\mathbf A}_1 = 0A^1​=0. It is not stated; (9.22) carries invertibility of A^t−1\hat{\mathbf A}_{t-1}A^t−1​ and a nonzero denominator as hypotheses.

A statement in which vπv_\pivπ​ is defined as the solution of the projected equation, or in which A\mathbf AA is assumed invertible or positive definite, would make the goal trivial or empty; neither is done. Convergence of the stochastic algorithm with probability one is not stated, since the book says it needs conditions and a step-size schedule it does not give. The bound for the episodic case and for other bootstrapping methods (p. 208) is stated only by reference in the book and is not a target.

Useful infrastructure: the μ\muμ-weighted inner product and orthogonal projection onto the column space of X\mathbf XX, the non-expansiveness of a stochastic matrix in the norm of its stationary distribution, and the positive definiteness criterion for non-symmetric matrices. All of these are reusable in the off-policy and average-reward chapters of the book. Contributions of these lemmas, and of alternative proofs of the milestones, are welcome.

Selected references

  • Richard S. Sutton and Andrew G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018, ISBN 9780262039246, §§9.2, 9.4, 9.8.
  • Richard S. Sutton, Learning to predict by the methods of temporal differences, Machine Learning 3, 1988. doi:10.1007/BF00115009
  • John N. Tsitsiklis and Benjamin Van Roy, An analysis of temporal-difference learning with function approximation, IEEE Transactions on Automatic Control 42(5), 1997. doi:10.1109/9.580874
  • Steven J. Bradtke and Andrew G. Barto, Linear least-squares algorithms for temporal difference learning, Machine Learning 22, 1996. doi:10.1007/BF00114723
  • Richard S. Varga, Matrix Iterative Analysis, Prentice-Hall, 1962.
11 thms2 active usersReviewed
Convex OptimizationOptimization·Captain: mikedeng1

SAGA: A Fast Incremental Gradient Method With Support for Non-Strongly Convex Composite Objectives II: The 4n/k Rate of the Averaged Iterate without Strong ConvexityResearch Paper

Motivation

Many problems in statistics and machine learning minimize an average of nnn losses, one per data point, plus a regularizer: least squares, logistic regression, and their ℓ1\ell_1ℓ1​- or ℓ2\ell_2ℓ2​-penalized versions. When nnn is large, a full gradient costs nnn component gradients, while stochastic gradient descent, which uses one component per step, needs decreasing step sizes and converges slowly. Incremental gradient methods with variance reduction (SAG, SVRG, SDCA, Finito, MISO) use one component gradient per step but converge at the rate of a full-gradient method.

SAGA (Defazio, Bach and Lacoste-Julien, NIPS 2014, arXiv:1407.0202) is a method of this family. It handles a non-smooth regularizer through its proximal operator, and it comes with a guarantee when the losses are convex but not strongly convex. This mission covers that second guarantee, Theorem 2 of the paper. A companion mission covers the linear rate under strong convexity (Theorem 1, Corollary 1).

Timeline.

  • 2012: SAG (Le Roux, Schmidt and Bach) gives a linear rate for smooth, strongly convex finite sums. Its analysis does not cover a proximal term.
  • 2013: SVRG (Johnson and Zhang) gives a linear rate for the strongly convex case, using periodic full-gradient passes.
  • 2013: SDCA (Shalev-Shwartz and Zhang) works on the dual and needs strong convexity.
  • 2014: Prox-SVRG (Xiao and Zhang, arXiv:1403.4699) extends SVRG to composite objectives. Its key inequality is reused by SAGA's Theorem 2.
  • 2014: SAGA proves both a linear rate under strong convexity and an O(n/k)O(n/k)O(n/k) rate for the averaged iterate under convexity alone, for composite objectives.

Setting

Let d≥0d\ge 0d≥0 and n≥1n\ge 1n≥1. The components f1,…,fn:Rd→Rf_1,\dots,f_n:\mathbb R^d\to\mathbb Rf1​,…,fn​:Rd→R are convex and differentiable, and each gradient fi′f_i'fi′​ is LLL-Lipschitz (L>0L>0L>0). Write

f(x)=1n∑i=1nfi(x),f′(x)=1n∑i=1nfi′(x).f(x)=\frac1n\sum_{i=1}^n f_i(x),\qquad f'(x)=\frac1n\sum_{i=1}^n f_i'(x).f(x)=n1​i=1∑n​fi​(x),f′(x)=n1​i=1∑n​fi′​(x).

The regularizer h:Rd→Rh:\mathbb R^d\to\mathbb Rh:Rd→R is convex but possibly non-differentiable. The objective is the composite function F=f+hF=f+hF=f+h, and x∗x^*x∗ is any minimizer of FFF. Minimizers need not be unique, and f′(x∗)f'(x^*)f′(x∗) need not vanish.

The proximal operator with parameter γ>0\gamma>0γ>0 is

proxγh(y)=arg⁡min⁡x∈Rd{h(x)+12γ∥x−y∥2}.\mathrm{prox}_\gamma^h(y)=\arg\min_{x\in\mathbb R^d}\Big\{h(x)+\frac1{2\gamma}\|x-y\|^2\Big\}.proxγh​(y)=argx∈Rdmin​{h(x)+2γ1​∥x−y∥2}.

SAGA keeps an iterate xkx^kxk and a table of points ϕ1k,…,ϕnk\phi_1^k,\dots,\phi_n^kϕ1k​,…,ϕnk​, initialized as ϕi0=x0\phi_i^0=x^0ϕi0​=x0. At step k+1k+1k+1 it draws an index jjj uniformly from {1,…,n}\{1,\dots,n\}{1,…,n}, independently of the past, and sets

wk+1=xk−γ[fj′(xk)−fj′(ϕjk)+1n∑i=1nfi′(ϕik)],xk+1=proxγh(wk+1).w^{k+1}=x^k-\gamma\Big[f_j'(x^k)-f_j'(\phi_j^k)+\frac1n\sum_{i=1}^n f_i'(\phi_i^k)\Big],\qquad x^{k+1}=\mathrm{prox}_\gamma^h(w^{k+1}).wk+1=xk−γ[fj′​(xk)−fj′​(ϕjk​)+n1​i=1∑n​fi′​(ϕik​)],xk+1=proxγh​(wk+1).

It then sets ϕjk+1=xk\phi_j^{k+1}=x^kϕjk+1​=xk and leaves the other table entries unchanged. The averaged iterate is xˉk=1k∑t=1kxt\bar x^k=\frac1k\sum_{t=1}^k x^txˉk=k1​∑t=1k​xt, which excludes x0x^0x0.

Formalization targets

Goal: Theorem 2 (p. 11)

With step size γ=1/(3L)\gamma=1/(3L)γ=1/(3L), for every k≥1k\ge1k≥1,

E[F(xˉk)]−F(x∗)≤4nk[2Ln∥x0−x∗∥2+f(x0)−⟨f′(x∗),x0−x∗⟩−f(x∗)].\mathbb E\big[F(\bar x^k)\big]-F(x^*)\le\frac{4n}{k}\Big[\frac{2L}{n}\|x^0-x^*\|^2+f(x^0)-\langle f'(x^*),x^0-x^*\rangle-f(x^*)\Big].E[F(xˉk)]−F(x∗)≤k4n​[n2L​∥x0−x∗∥2+f(x0)−⟨f′(x∗),x0−x∗⟩−f(x∗)].

The expectation is over the indices j1,…,jkj^1,\dots,j^kj1,…,jk. The constants are those printed in the paper.

Milestones (in attack order)

  1. Lemma 1 (p. 6) is an inner-product bound for averages of μ\muμ-strongly convex functions with LLL-Lipschitz gradients. It is stated for μ≥0\mu\ge0μ≥0, and Theorem 2 uses the case μ=0\mu=0μ=0.
  2. Lemma 2 (p. 7): 1n∑i∥fi′(ϕi)−fi′(x∗)∥2≤2L[1n∑ifi(ϕi)−f(x∗)−1n∑i⟨fi′(x∗),ϕi−x∗⟩]\frac1n\sum_i\|f_i'(\phi_i)-f_i'(x^*)\|^2\le 2L\big[\frac1n\sum_i f_i(\phi_i)-f(x^*)-\frac1n\sum_i\langle f_i'(x^*),\phi_i-x^*\rangle\big]n1​∑i​∥fi′​(ϕi​)−fi′​(x∗)∥2≤2L[n1​∑i​fi​(ϕi​)−f(x∗)−n1​∑i​⟨fi′​(x∗),ϕi​−x∗⟩].
  3. The bound on Δ\DeltaΔ (p. 12). Write Δ=−1γ(wk+1−xk)−f′(xk)\Delta=-\frac1\gamma(w^{k+1}-x^k)-f'(x^k)Δ=−γ1​(wk+1−xk)−f′(xk) for the gradient error. For every β>0\beta>0β>0, E∥Δ∥2≤(1+β−1)E∥fj′(ϕjk)−fj′(x∗)∥2+(1+β)E∥fj′(xk)−fj′(x∗)∥2\mathbb E\|\Delta\|^2\le(1+\beta^{-1})\mathbb E\|f_j'(\phi_j^k)-f_j'(x^*)\|^2+(1+\beta)\mathbb E\|f_j'(x^k)-f_j'(x^*)\|^2E∥Δ∥2≤(1+β−1)E∥fj′​(ϕjk​)−fj′​(x∗)∥2+(1+β)E∥fj′​(xk)−fj′​(x∗)∥2.
  4. The prox-SVRG inequality (p. 12): αE∥xk+1−x∗∥2≤α∥xk−x∗∥2−2αγE[F(xk+1)−F(x∗)]+2αγ2E∥Δ∥2\alpha\mathbb E\|x^{k+1}-x^*\|^2\le\alpha\|x^k-x^*\|^2-2\alpha\gamma\mathbb E[F(x^{k+1})-F(x^*)]+2\alpha\gamma^2\mathbb E\|\Delta\|^2αE∥xk+1−x∗∥2≤α∥xk−x∗∥2−2αγE[F(xk+1)−F(x∗)]+2αγ2E∥Δ∥2.
  5. The one-step Lyapunov decrease (p. 12): E[Tk+1]−Tk≤−14nE[F(xk+1)−F(x∗)]\mathbb E[T^{k+1}]-T^k\le-\frac1{4n}\mathbb E[F(x^{k+1})-F(x^*)]E[Tk+1]−Tk≤−4n1​E[F(xk+1)−F(x∗)]. Here T(x,ϕ)=1n∑ifi(ϕi)−f(x∗)−1n∑i⟨fi′(x∗),ϕi−x∗⟩+(c+α)∥x−x∗∥2T(x,\phi)=\frac1n\sum_i f_i(\phi_i)-f(x^*)-\frac1n\sum_i\langle f_i'(x^*),\phi_i-x^*\rangle+(c+\alpha)\|x-x^*\|^2T(x,ϕ)=n1​∑i​fi​(ϕi​)−f(x∗)−n1​∑i​⟨fi′​(x∗),ϕi​−x∗⟩+(c+α)∥x−x∗∥2, with c=3L2nc=\frac{3L}{2n}c=2n3L​ and α=3L8n\alpha=\frac{3L}{8n}α=8n3L​.

In milestones 3–5, E\mathbb EE is the expectation over the single index jjj of the next step, given the current state.

Significance

The result. Theorem 2 shows that one method, with a step size that depends only on LLL, covers composite problems that are not strongly convex. Examples are ℓ1\ell_1ℓ1​-regularized least squares and logistic regression without a ridge term. On these problems the method converges in expected objective value at rate O(n/k)O(n/k)O(n/k). SAG has no proximal analysis, and SDCA requires strong convexity. With the same step size 1/(3L)1/(3L)1/(3L), the paper also states adaptivity to strong convexity, so no strong convexity constant has to be known in advance. The bound is in terms of T0T^0T0, a quantity computable from the starting point.

Formalizing it. The result is proved on paper, but the proof is not self-contained. Its central inequality (milestone 4) is quoted from the prox-SVRG analysis of Xiao and Zhang, with only the remark that their argument uses E[Δ]=0\mathbb E[\Delta]=0E[Δ]=0. A machine-checked proof must therefore reconstruct that argument for SAGA's estimator. To our knowledge, no machine-checked proof of SAGA, SVRG or prox-SVRG exists in Lean or Mathlib. The mission also produces reusable statements about convex functions with Lipschitz gradients (Lemmas 1 and 2) and an explicit finite model of a randomized incremental method.

Difficulty

The naive approach applies the non-expansiveness of the proximal operator to ∥xk+1−x∗∥2\|x^{k+1}-x^*\|^2∥xk+1−x∗∥2, as in the strongly convex proof. That bounds distances, but it produces no term in F(xk+1)−F(x∗)F(x^{k+1})-F(x^*)F(xk+1)−F(x∗). Without strong convexity, the distance terms cannot be traded for function values, so the argument yields no rate.

The function-value term comes from the prox-SVRG inequality (milestone 4), which the paper does not prove. Its difficulty is that xk+1x^{k+1}xk+1 depends on the same random index as Δ\DeltaΔ, so the cross term between them does not vanish in expectation even though E[Δ]=0\mathbb E[\Delta]=0E[Δ]=0. A second difficulty is bookkeeping: wk+1w^{k+1}wk+1 uses the old table, the table entry jjj receives xkx^kxk and not xk+1x^{k+1}xk+1, and the constants must make three coefficients vanish exactly. A final step converts the bound on 1k∑tE[F(xt)]\frac1k\sum_t\mathbb E[F(x^t)]k1​∑t​E[F(xt)] into a bound on E[F(xˉk)]\mathbb E[F(\bar x^k)]E[F(xˉk)], which requires Jensen's inequality for the convex FFF.

Formalization scope

  • Space and indices. Points live in EuclideanSpace ℝ (Fin d). Components are indexed by Fin n (0-based), with n≥1n\ge1n≥1.
  • Gradients and smoothness. The gradients are given maps f' with HasGradientAt (f i) (f' i x) x at every point. Smoothness is the Lipschitz bound ∥fi′(x)−fi′(y)∥≤L∥x−y∥\|f_i'(x)-f_i'(y)\|\le L\|x-y\|∥fi′​(x)−fi′​(y)∥≤L∥x−y∥.
  • Convexity. Convexity is ConvexOn ℝ Set.univ. Lemma 1 uses StrongConvexOn Set.univ μ, whose modulus μ2∥x−y∥2\frac\mu2\|x-y\|^22μ​∥x−y∥2 is the paper's.
  • The regularizer. hhh is real-valued and convex. Extended-valued regularizers such as indicator functions are outside the statement.
  • The proximal map. The proximal operator enters as any map PPP such that P(y)P(y)P(y) minimizes h(z)+12γ∥z−y∥2h(z)+\frac1{2\gamma}\|z-y\|^2h(z)+2γ1​∥z−y∥2 for every yyy. For convex hhh this determines P=proxγhP=\mathrm{prox}_\gamma^hP=proxγh​.
  • State and expectation. The state is the pair (xk,ϕk)(x^k,\phi^k)(xk,ϕk). The expectation over kkk steps is the uniform average over the nkn^knk index sequences, which is exactly the law of kkk independent uniform indices.

Two trivializations are excluded. The averaged-iterate bound carries k≥1k\ge1k≥1, since at k=0k=0k=0 the factor 4n/k4n/k4n/k collapses to 000. The left side is FFF evaluated at the averaged point, not the average of F(xt)F(x^t)F(xt), which is a weaker intermediate step.

A complete development needs the descent lemma and co-coercivity for convex functions with Lipschitz gradients, the characterization and non-expansiveness of the proximal operator, and finite-sum manipulations over index sequences. The lemmas on smooth convex functions and on proximal operators are reusable beyond this mission. Contributions are welcome at every level: proofs of the milestones, a reusable proximal-operator library, and the telescoping argument for the goal.

Selected references

  • A. Defazio, F. Bach, S. Lacoste-Julien, SAGA: A Fast Incremental Gradient Method With Support for Non-Strongly Convex Composite Objectives, NIPS 2014. arXiv:1407.0202
  • L. Xiao, T. Zhang, A Proximal Stochastic Gradient Method with Progressive Variance Reduction, SIAM J. Optim. 24(4), 2014. arXiv:1403.4699
  • N. Le Roux, M. Schmidt, F. Bach, A Stochastic Gradient Method with an Exponential Convergence Rate for Finite Training Sets, NIPS 2012. arXiv:1202.6258
  • R. Johnson, T. Zhang, Accelerating Stochastic Gradient Descent using Predictive Variance Reduction, NIPS 2013. NeurIPS proceedings
  • S. Shalev-Shwartz, T. Zhang, Stochastic Dual Coordinate Ascent Methods for Regularized Loss Minimization, JMLR 14, 2013. arXiv:1209.1873
  • Y. Nesterov, Introductory Lectures on Convex Optimization, Kluwer, 2004. doi:10.1007/978-1-4419-8853-9
11 thms2 active usersReviewed
ProbabilityReinforcement LearningStatistics·Captain: mikedeng1

Reinforcement Learning: An Introduction V: Off-policy Prediction by Importance SamplingTextbook

Motivation

Reinforcement learning methods must explore in order to find good behaviour, yet the quantity they usually want to evaluate is the value of a different, often deterministic, policy. Off-policy prediction separates the two roles: episodes are generated by a behaviour policy bbb, and the goal is the value function vπv_\pivπ​ of a target policy π\piπ. Almost every off-policy method in Sutton and Barto's Reinforcement Learning: An Introduction (2nd ed., MIT Press, 2018), and in the literature that follows it, rests on importance sampling: a return observed under bbb is reweighted by the relative probability of its trajectory under π\piπ and bbb. Section 5.5 of the book introduces the idea for Monte Carlo prediction, §5.6 gives the incremental form of the weighted estimator, and §§5.8–5.9 refine the weights using the internal structure of the return: discounting-aware importance sampling, after Sutton, Mahmood, Precup and van Hasselt (2014), and per-decision importance sampling, introduced by Precup, Sutton and Singh (2000). The book's remarks on the variance of the two estimators (p. 105) cite Precup, Sutton and Dasgupta (2001). Later chapters (7, 11, 12) reuse the same ratios for nnn-step, gradient-TD and eligibility-trace methods.

This mission is the fifth in a series formalizing the book's central mathematical claims. It covers §§5.5–5.9 (pp. 103–115).

Setting

A finite Markov decision process has finite sets of states S\mathcal SS (terminal states included), actions A\mathcal AA and rewards R⊂R\mathcal R \subset \mathbb RR⊂R, and dynamics p(s′,r∣s,a)p(s', r \mid s, a)p(s′,r∣s,a): for each (s,a)(s, a)(s,a) a probability distribution over next state and reward. The state-transition probability is p(s′∣s,a)=∑rp(s′,r∣s,a)p(s' \mid s, a) = \sum_r p(s', r \mid s, a)p(s′∣s,a)=∑r​p(s′,r∣s,a). A policy μ\muμ gives a distribution μ(⋅∣s)\mu(\cdot \mid s)μ(⋅∣s) over actions in every state.

An episode from a start state sss is a sequence S0=s,A0,R1,S1,…,AT−1,RT,STS_0 = s, A_0, R_1, S_1, \dots, A_{T-1}, R_T, S_TS0​=s,A0​,R1​,S1​,…,AT−1​,RT​,ST​ in which S0,…,ST−1S_0, \dots, S_{T-1}S0​,…,ST−1​ are nonterminal and STS_TST​ is terminal. Under μ\muμ it has probability ∏k=0T−1μ(Ak∣Sk) p(Sk+1,Rk+1∣Sk,Ak)\prod_{k=0}^{T-1} \mu(A_k \mid S_k)\, p(S_{k+1}, R_{k+1} \mid S_k, A_k)∏k=0T−1​μ(Ak​∣Sk​)p(Sk+1​,Rk+1​∣Sk​,Ak​). The return is G0=∑k=0T−1γkRk+1G_0 = \sum_{k=0}^{T-1} \gamma^k R_{k+1}G0​=∑k=0T−1​γkRk+1​ with discount rate γ∈[0,1]\gamma \in [0, 1]γ∈[0,1], and the value vπ(s)v_\pi(s)vπ​(s) is the expected return of an episode generated by π\piπ from sss.

The behaviour policy covers the target policy if π(a∣s)>0\pi(a \mid s) > 0π(a∣s)>0 implies b(a∣s)>0b(a \mid s) > 0b(a∣s)>0. The importance-sampling ratio of decisions 0,…,j0, \dots, j0,…,j is

ρ0:j=∏k=0jπ(Ak∣Sk)b(Ak∣Sk),\rho_{0:j} = \prod_{k=0}^{j} \frac{\pi(A_k \mid S_k)}{b(A_k \mid S_k)},ρ0:j​=k=0∏j​b(Ak​∣Sk​)π(Ak​∣Sk​)​,

and the per-decision return weights each reward only by the ratio of the decisions that precede it:

G~0=ρ0:0R1+γρ0:1R2+⋯+γT−1ρ0:T−1RT.\tilde G_0 = \rho_{0:0} R_1 + \gamma \rho_{0:1} R_2 + \dots + \gamma^{T-1} \rho_{0:T-1} R_T .G~0​=ρ0:0​R1​+γρ0:1​R2​+⋯+γT−1ρ0:T−1​RT​.

The book writes these objects at a general time ttt and conditions on St=sS_t = sSt​=s; by the Markov property this is the same as starting the episode at sss, which is what the formal statements do.

Formalization targets

Goal: unbiasedness of ordinary and per-decision importance sampling

For episodes generated by bbb from sss,

Eb[ρ0:T−1G0∣S0=s]=vπ(s)=Eb[G~0∣S0=s].\mathbb E_b\bigl[\rho_{0:T-1} G_0 \mid S_0 = s\bigr] = v_\pi(s) = \mathbb E_b\bigl[\tilde G_0 \mid S_0 = s\bigr].Eb​[ρ0:T−1​G0​∣S0​=s]=vπ​(s)=Eb​[G~0​∣S0​=s].

The first equality is Eq. (5.4) (p. 104); the second is the statement E[ρt:T−1Gt]=E[G~t]\mathbb E[\rho_{t:T-1}G_t] = \mathbb E[\tilde G_t]E[ρt:T−1​Gt​]=E[G~t​] of §5.9 (p. 114).

Milestones

  1. (5.3): the trajectory probability is a product, and the ratio of trajectory probabilities under π\piπ and bbb is ρ0:T−1\rho_{0:T-1}ρ0:T−1​, independent of the dynamics.
  2. (5.4) alone.
  3. (5.13): ∑ab(a∣x) π(a∣x)/b(a∣x)=∑aπ(a∣x)=1\sum_a b(a \mid x)\, \pi(a \mid x)/b(a \mid x) = \sum_a \pi(a \mid x) = 1∑a​b(a∣x)π(a∣x)/b(a∣x)=∑a​π(a∣x)=1 under coverage.
  4. (5.14) and its kkk-th form: Eb[ρ0:T−1Rk]=Eb[ρ0:k−1Rk]\mathbb E_b[\rho_{0:T-1} R_k] = \mathbb E_b[\rho_{0:k-1} R_k]Eb​[ρ0:T−1​Rk​]=Eb​[ρ0:k−1​Rk​] for every k≥1k \ge 1k≥1 (Exercise 5.13).
  5. Example 5.5: in a one-state MDP with a loop, vπ(s)=1v_\pi(s) = 1vπ​(s)=1 and Eb[ρ0:T−1G0]=1\mathbb E_b[\rho_{0:T-1}G_0] = 1Eb​[ρ0:T−1​G0​]=1, yet Eb[(ρ0:T−1G0)2]=∞\mathbb E_b[(\rho_{0:T-1}G_0)^2] = \inftyEb​[(ρ0:T−1​G0​)2]=∞.
  6. (5.7)–(5.8): the incremental rule Vn+1=Vn+(Wn/Cn)(Gn−Vn)V_{n+1} = V_n + (W_n/C_n)(G_n - V_n)Vn+1​=Vn​+(Wn​/Cn​)(Gn​−Vn​) computes the weighted average ∑k<nWkGk/∑k<nWk\sum_{k<n} W_k G_k / \sum_{k<n} W_k∑k<n​Wk​Gk​/∑k<n​Wk​ (Exercise 5.10).
  7. §5.8: Gt=(1−γ)∑h=t+1T−1γh−t−1Gˉt:h+γT−t−1Gˉt:TG_t = (1-\gamma)\sum_{h=t+1}^{T-1}\gamma^{h-t-1}\bar G_{t:h} + \gamma^{T-t-1}\bar G_{t:T}Gt​=(1−γ)∑h=t+1T−1​γh−t−1Gˉt:h​+γT−t−1Gˉt:T​ with flat partial returns Gˉt:h=Rt+1+⋯+Rh\bar G_{t:h} = R_{t+1} + \dots + R_hGˉt:h​=Rt+1​+⋯+Rh​.

Significance

Eq. (5.4) is the reason the first-visit ordinary importance-sampling estimator (5.5) is unbiased, and it is the template for every importance-sampling correction in the rest of the book. The per-decision identity shows that an estimator with fewer ratio factors per reward, (5.15), has the same expectation, which is the starting point for per-decision and control-variate methods for multi-step off-policy learning (Precup, Sutton and Singh 2000). Example 5.5 shows that unbiasedness says nothing about variance: the ordinary estimator can have infinite variance on a two-action problem, which motivates weighted importance sampling and the incremental weighted update of §5.6.

The results of these sections are classical and proved informally in the book, partly as exercises (5.10, 5.13) left without solution. No machine-checked version exists on the platform: a search for importance sampling, off-policy and per-decision returned no statements. The mission produces a formal trajectory model of an episodic MDP under two policies, which later missions on nnn-step off-policy returns and off-policy traces can reuse.

Difficulty

Eq. (5.4) itself is a termwise identity: for every episode, Pr⁡b(episode) ρ0:T−1=Pr⁡π(episode)\Pr_b(\text{episode})\,\rho_{0:T-1} = \Pr_\pi(\text{episode})Prb​(episode)ρ0:T−1​=Prπ​(episode) under coverage. The per-decision identity is not termwise. The later factors of ρ0:T−1\rho_{0:T-1}ρ0:T−1​ multiply a reward that was received before the corresponding decisions, and removing them requires summing over all continuations of an episode prefix, of every remaining length, and using that each factor has conditional expectation one (5.13) and that the continuation terminates with probability one. The obvious attempt, cancelling the factors episode by episode, fails: on a single episode ρ0:T−1R1\rho_{0:T-1}R_1ρ0:T−1​R1​ and ρ0:0R1\rho_{0:0}R_1ρ0:0​R1​ differ.

In Example 5.5 the episodes have no length bound, so the expected square is an infinite series over episode lengths whose divergence must be shown directly.

Formalization scope

  • States, actions and rewards are finite types; the terminal states are a finite subset of the state type. Policies are stochastic, one action set is used in every state, and vπv_\pivπ​ is defined as an expected return, never as the solution of a Bellman equation.
  • Expectations are series over episode lengths of finite sums over episodes. Lean assigns 000 to a divergent series, so the goal and milestones 2 and 4 assume that under bbb every episode from sss terminates within a fixed number HHH of steps with probability one. The book leaves termination implicit; this bounded-horizon hypothesis is a restriction relative to the book's episodic setting and is stated as such. Example 5.5, whose episodes are unbounded, is stated without it, with the expected square in [0,∞][0, \infty][0,∞].
  • The discount rate is kept general in [0,1][0, 1][0,1].
  • The importance-sampling ratio uses real division; a factor with b(Ak∣Sk)=0b(A_k \mid S_k) = 0b(Ak​∣Sk​)=0 evaluates to 000 in Lean, but such episodes have probability 000 under bbb.
  • The flat-partial-return decomposition is an algebraic identity and is stated for every real γ\gammaγ, which is more general than the book's "for any γ∈[0,1)\gamma \in [0,1)γ∈[0,1)".
  • The incremental weighted update is stated with nonnegative weights and W1>0W_1 > 0W1​>0. The book's C0=0C_0 = 0C0​=0 makes (5.8) divide by zero at n=1n = 1n=1 when W1=0W_1 = 0W1​=0; the hypothesis excludes that case.
  • A trivializing formalization is ruled out: vπv_\pivπ​ is the expected return of π\piπ's own episodes, the ratio is computed from the episode, and the per-decision identity, which carries the chapter's content beyond (5.4), is part of the goal.
  • Not stated: the bias and variance comparisons of ordinary and weighted importance sampling (p. 105) and the discounting-aware estimators (5.9)–(5.10) as estimators; these are statistical claims about estimators over a random number of visits that the book does not make precise.

Contributions welcome: proofs of the milestones, a general measure-theoretic version of the trajectory model without the bounded-horizon hypothesis, and variants for action values qπq_\piqπ​ (Exercise 5.6).

Selected references

  • R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018, ISBN 9780262039246, §§5.5–5.9, pp. 103–115. http://incompleteideas.net/book/the-book-2nd.html
  • D. Precup, R. S. Sutton and S. Singh, Eligibility Traces for Off-Policy Policy Evaluation, Proceedings of the 17th International Conference on Machine Learning (ICML), 2000, pp. 759–766 (cited in the book's bibliography).
  • D. Precup, R. S. Sutton and S. Dasgupta, Off-Policy Temporal-Difference Learning with Function Approximation, Proceedings of the 18th International Conference on Machine Learning (ICML), 2001, pp. 417–424 (cited in the book, p. 105).
14 thms2 active usersReviewed
Convex OptimizationOptimization·Captain: mikedeng1

SAGA: A Fast Incremental Gradient Method With Support for Non-Strongly Convex Composite Objectives I: Linear Convergence under Strong ConvexityResearch Paper

Motivation

Many problems in machine learning and statistics are finite sums: an empirical risk f(x)=1n∑i=1nfi(x)f(x)=\frac1n\sum_{i=1}^n f_i(x)f(x)=n1​∑i=1n​fi​(x) over nnn data points, often plus a regulariser hhh such as an ℓ1\ell_1ℓ1​ penalty. When nnn is large, a full gradient of fff costs nnn component gradients, while stochastic gradient descent uses one component per step but converges only sublinearly because its gradient estimate has non-vanishing variance. Incremental gradient methods with variance reduction keep the per-step cost of one component gradient and still converge linearly on strongly convex problems.

SAGA, introduced by Defazio, Bach and Lacoste-Julien at NIPS 2014 (arXiv:1407.0202), is one of the standard methods of this family, alongside SAG, SVRG, SDCA and Finito/MISO. It keeps a table of past component gradients and handles a non-smooth regulariser through its proximal operator.

Timeline. Le Roux, Schmidt and Bach (2012) gave SAG the first linear rate for strongly convex finite sums at the cost of one gradient per step. Shalev-Shwartz and Zhang (2013) proved linear rates for SDCA, a dual method. Johnson and Zhang (2013) introduced SVRG, with periodic full-gradient passes; Xiao and Zhang (2014) extended it to composite objectives (prox-SVRG). SAGA (2014) combines an unbiased SVRG-style estimator with a SAG-style table, and proves a linear rate in the composite strongly convex case with a simple Lyapunov argument.

Setting

Let Rd\mathbb R^dRd carry the Euclidean inner product. There are n≥1n\ge1n≥1 differentiable components f1,…,fn:Rd→Rf_1,\dots,f_n:\mathbb R^d\to\mathbb Rf1​,…,fn​:Rd→R with gradients fi′f_i'fi′​. Each fif_ifi​ is μ\muμ-strongly convex (μ>0\mu>0μ>0): fi(ax+by)≤afi(x)+bfi(y)−abμ2∥x−y∥2f_i(ax+by)\le af_i(x)+bf_i(y)-ab\frac\mu2\|x-y\|^2fi​(ax+by)≤afi​(x)+bfi​(y)−ab2μ​∥x−y∥2 for a,b≥0a,b\ge0a,b≥0, a+b=1a+b=1a+b=1. Each gradient is LLL-Lipschitz: ∥fi′(x)−fi′(y)∥≤L∥x−y∥\|f_i'(x)-f_i'(y)\|\le L\|x-y\|∥fi′​(x)−fi′​(y)∥≤L∥x−y∥. Write f=1n∑ifif=\frac1n\sum_i f_if=n1​∑i​fi​ and f′=1n∑ifi′f'=\frac1n\sum_i f_i'f′=n1​∑i​fi′​. The regulariser h:Rd→Rh:\mathbb R^d\to\mathbb Rh:Rd→R is convex, and the goal is to minimise the composite objective F=f+hF=f+hF=f+h; x∗x^*x∗ denotes its minimiser, which is unique.

The proximal operator with step γ>0\gamma>0γ>0 is

prox⁡γh(y)=argmin⁡x{h(x)+12γ∥x−y∥2}.\operatorname{prox}^h_\gamma(y)=\operatorname*{argmin}_{x}\Big\{h(x)+\tfrac1{2\gamma}\|x-y\|^2\Big\}.proxγh​(y)=xargmin​{h(x)+2γ1​∥x−y∥2}.

SAGA keeps an iterate xkx^kxk and points ϕ1k,…,ϕnk\phi_1^k,\dots,\phi_n^kϕ1k​,…,ϕnk​ at which the stored gradients fi′(ϕik)f_i'(\phi_i^k)fi′​(ϕik​) were taken. It starts from x0x^0x0 with ϕi0=x0\phi_i^0=x^0ϕi0​=x0. At iteration k+1k+1k+1 it draws jjj uniformly from {1,…,n}\{1,\dots,n\}{1,…,n}, independently of the past, and sets

wk+1=xk−γ[fj′(xk)−fj′(ϕjk)+1n∑i=1nfi′(ϕik)],xk+1=prox⁡γh(wk+1),w^{k+1}=x^k-\gamma\Big[f_j'(x^k)-f_j'(\phi_j^k)+\frac1n\sum_{i=1}^n f_i'(\phi_i^k)\Big],\qquad x^{k+1}=\operatorname{prox}^h_\gamma(w^{k+1}),wk+1=xk−γ[fj′​(xk)−fj′​(ϕjk​)+n1​i=1∑n​fi′​(ϕik​)],xk+1=proxγh​(wk+1),

then ϕjk+1=xk\phi_j^{k+1}=x^kϕjk+1​=xk, with every other entry unchanged.

The analysis uses the Lyapunov function

T(x,{ϕi})=1n∑ifi(ϕi)−f(x∗)−1n∑i⟨fi′(x∗),ϕi−x∗⟩+c∥x−x∗∥2.T(x,\{\phi_i\})=\frac1n\sum_i f_i(\phi_i)-f(x^*)-\frac1n\sum_i\langle f_i'(x^*),\phi_i-x^*\rangle+c\|x-x^*\|^2 .T(x,{ϕi​})=n1​i∑​fi​(ϕi​)−f(x∗)−n1​i∑​⟨fi′​(x∗),ϕi​−x∗⟩+c∥x−x∗∥2.

Formalization targets

Goal: Corollary 1 (p. 8)

With γ=12(μn+L)\gamma=\frac1{2(\mu n+L)}γ=2(μn+L)1​, for every k≥0k\ge0k≥0,

E∥xk−x∗∥2≤(1−μ2(μn+L))k[∥x0−x∗∥2+nμn+L(f(x0)−⟨f′(x∗),x0−x∗⟩−f(x∗))],\mathbb E\|x^k-x^*\|^2\le\Big(1-\frac{\mu}{2(\mu n+L)}\Big)^k\Big[\|x^0-x^*\|^2+\frac{n}{\mu n+L}\big(f(x^0)-\langle f'(x^*),x^0-x^*\rangle-f(x^*)\big)\Big],E∥xk−x∗∥2≤(1−2(μn+L)μ​)k[∥x0−x∗∥2+μn+Ln​(f(x0)−⟨f′(x∗),x0−x∗⟩−f(x∗))],

where the expectation is over the indices drawn in the first kkk iterations. The constants are the paper's.

Theorem 1 (p. 7)

With γ\gammaγ as above, c=12γ(1−γμ)nc=\frac1{2\gamma(1-\gamma\mu)n}c=2γ(1−γμ)n1​ and κ=1γμ\kappa=\frac1{\gamma\mu}κ=γμ1​, for every state (xk,{ϕik})(x^k,\{\phi^k_i\})(xk,{ϕik​}),

E[Tk+1]≤(1−1κ)Tk,\mathbb E\big[T^{k+1}\big]\le\Big(1-\frac1\kappa\Big)T^k ,E[Tk+1]≤(1−κ1​)Tk,

with the expectation over the next index only.

Supporting lemmas

Lemma 4 (p. 10), a lower bound combining strong convexity and smoothness; Lemma 1 (pp. 6–7), its average over the components; Lemma 2 (p. 7), which bounds the stale-gradient variance by the table part of TTT; and Lemma 3 (p. 7), a second-moment bound for the SAGA step.

Significance

The result. Corollary 1 gives an ε\varepsilonε-accurate iterate in expectation after O((n+L/μ)log⁡(1/ε))O\big((n+L/\mu)\log(1/\varepsilon)\big)O((n+L/μ)log(1/ε)) component-gradient evaluations. This is the complexity of full-gradient descent with the condition number decoupled from nnn, and it holds in the composite setting, so it covers the lasso and elastic-net problems that SAG's analysis does not reach. The paper notes that the rate improves on the published rates of SAG and SVRG and is within a factor 2 of SDCA's. Theorem 1 is the template of later Lyapunov analyses of variance-reduced methods.

Formalizing it. The result has been proved since 2014, and no machine-checked proof is known to this mission. The work left is to formalize the known proof: the convexity inequalities (Lemmas 4, 1, 2), the variance computation (Lemma 3), the one-step contraction (Theorem 1), and the passage from conditional to total expectation along the random index sequence (Corollary 1). The paper's Lemma 3 has a sign misprint, which the formalization corrects; see the scope section.

Difficulty

The obvious argument for SGD-type methods bounds E∥xk+1−x∗∥2\mathbb E\|x^{k+1}-x^*\|^2E∥xk+1−x∗∥2 in terms of ∥xk−x∗∥2\|x^k-x^*\|^2∥xk−x∗∥2 alone. That fails here: the variance of the SAGA estimator depends on the stale table points ϕik\phi_i^kϕik​, which can be far from x∗x^*x∗ even when xkx^kxk is close. One needs a potential that also measures the table. Balancing the terms of TTT then requires the four round-bracket coefficients in the paper's display (10) to be non-positive for the specific γ\gammaγ, ccc and an auxiliary β=(2μn+L)/L\beta=(2\mu n+L)/Lβ=(2μn+L)/L. Checking these coefficients is routine but long algebra in μ\muμ, LLL, nnn. The composite case adds one step: since f′(x∗)≠0f'(x^*)\neq0f′(x∗)=0 in general, the argument goes through the fixed-point identity x∗=prox⁡γh(x∗−γf′(x∗))x^*=\operatorname{prox}^h_\gamma(x^*-\gamma f'(x^*))x∗=proxγh​(x∗−γf′(x∗)) and the non-expansiveness of the proximal operator, neither of which is a numbered result of the paper.

Formalization scope

  • Space and data. The space is EuclideanSpace ℝ (Fin d). Components are indexed by Fin n (0-based), with 0 < n. The gradients fi′f_i'fi′​ are given maps with HasGradientAt (f i) (f' i x) x. Strong convexity is Mathlib's StrongConvexOn Set.univ μ, whose modulus μ2∥x−y∥2\frac\mu2\|x-y\|^22μ​∥x−y∥2 is the paper's.

  • Regulariser and minimiser. hhh is real-valued and convex; extended-valued regularisers are out of scope, as on the page. A minimiser x∗x^*x∗ of f+hf+hf+h is a hypothesis.

  • Proximal operator. It is any map PPP such that P(y)P(y)P(y) minimises h(x)+12γ∥x−y∥2h(x)+\frac1{2\gamma}\|x-y\|^2h(x)+2γ1​∥x−y∥2 for every yyy (IsProxPoint). The minimiser is unique, so PPP is prox⁡γh\operatorname{prox}^h_\gammaproxγh​.

  • State and expectation. The state is the pair (x,ϕ)(x,\phi)(x,ϕ). The run after kkk steps is a deterministic function of the index sequence in Fin k → Fin n. The expectation in Corollary 1 is the average over all nkn^knk sequences, which is exactly the law of kkk independent uniform indices; no measure theory is involved. Theorem 1's conditional expectation is the average over the next index.

  • Constants and corrections. Constants are as printed and fixed, not "for some constant" and not "for all small enough steps". Lemma 4 carries the hypothesis μ<L\mu<Lμ<L, which its fractions 1/(L−μ)1/(L-\mu)1/(L−μ) require. Lemma 3 is stated with +γf′(x∗)+\gamma f'(x^*)+γf′(x∗), as in its proof and its use in Theorem 1; the printed −γf′(x∗)-\gamma f'(x^*)−γf′(x∗) is false whenever f′(x∗)≠0f'(x^*)\neq0f′(x∗)=0.

  • Trivializing formalizations, ruled out. Taking the proximal step as merely non-expansive, fixing an index sequence instead of averaging over all of them, measuring x∗x^*x∗ against fff instead of f+hf+hf+h, or restricting Theorem 1 to reachable states changes the theorem and is excluded.

  • Infrastructure. A complete development needs:

    • the co-coercivity inequality for convex functions with Lipschitz gradient;
    • existence, uniqueness and non-expansiveness of the proximal map of a finite convex function;
    • the optimality condition x∗=prox⁡γh(x∗−γf′(x∗))x^*=\operatorname{prox}^h_\gamma(x^*-\gamma f'(x^*))x∗=proxγh​(x∗−γf′(x∗));
    • finite-sum variance identities.

    These pieces are reusable well beyond SAGA, by SVRG, SAG and proximal-gradient analyses. Contributions of any of them, or of proofs of the individual milestones, are welcome.

Selected references

  • A. Defazio, F. Bach, S. Lacoste-Julien, SAGA: A Fast Incremental Gradient Method With Support for Non-Strongly Convex Composite Objectives, NIPS 2014. arXiv:1407.0202
  • N. Le Roux, M. Schmidt, F. Bach, A Stochastic Gradient Method with an Exponential Convergence Rate for Finite Training Sets, NIPS 2012. arXiv:1202.6258
  • R. Johnson, T. Zhang, Accelerating Stochastic Gradient Descent using Predictive Variance Reduction, NIPS 2013. NeurIPS proceedings
  • L. Xiao, T. Zhang, A Proximal Stochastic Gradient Method with Progressive Variance Reduction, SIAM J. Optim. 24(4), 2014. arXiv:1403.4699
  • S. Shalev-Shwartz, T. Zhang, Stochastic Dual Coordinate Ascent Methods for Regularized Loss Minimization, JMLR 14, 2013. arXiv:1209.1873
  • Y. Nesterov, Introductory Lectures on Convex Optimization, Kluwer, 2004. doi:10.1007/978-1-4419-8853-9
11 thms2 active usersReviewed
PreviousPage 5 of 11Next

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me