Prove2Me
Navigate
DiscoverFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

Machine Learning

291 missions · 186 completed

The science of systems that learn from data and experience. Its scope runs from the statistical and mathematical foundations of learning, including generalization, expressivity, and computational limits, through the design of learning algorithms, deep learning, reinforcement learning, and probabilistic methods, to the empirical study of large models and the trustworthiness, interpretability, and societal impact of learned systems.

Missions

Open105Completed186All291
🏆Completed
ProbabilityRandom Matrix TheoryStatistics·Captain: mikedeng1

High-Dimensional Probability I: Approximate Carathéodory's TheoremTextbook

Motivation

Many arguments in high-dimensional geometry, statistics and computer science need to approximate a point of a convex set by an average of a handful of extreme points, rather than represent it exactly. The classical Carathéodory theorem (1907) answers the exact question: every point of the convex hull of a set T⊆RnT \subseteq \mathbb{R}^nT⊆Rn is a convex combination of at most n+1n+1n+1 points of TTT. That bound is tight and grows with the dimension nnn, which makes it useless whenever nnn is large — exactly the regime of interest in high-dimensional probability.

B. Maurey's empirical method — an unpublished 1980–81 result reported by G. Pisier, "Remarques sur un résultat non publié de B. Maurey," Séminaire d'Analyse Fonctionnelle 1980–1981 — and later applied by B. Carl to bound covering numbers of operators between Banach spaces (Inequalities of Bernstein-Jackson-type and the degree of compactness of operators in Banach spaces, Ann. Inst. Fourier 35(3), 1985, 79–118), replaces the exact question with an approximate one and removes the dimension dependence entirely: to approximate xxx to accuracy ε\varepsilonε, the number of points needed depends only on ε\varepsilonε, never on nnn. Vershynin's High-Dimensional Probability opens with this result as its "Appetizer," using it to illustrate the book's central theme — that randomness is a tool for constructing deterministic combinatorial objects — before any probabilistic machinery has been introduced.

Setting

A convex combination of finitely many points z1,…,zm∈Rnz_1, \dots, z_m \in \mathbb{R}^nz1​,…,zm​∈Rn is a sum ∑i=1mλizi\sum_{i=1}^m \lambda_i z_i∑i=1m​λi​zi​ with λi≥0\lambda_i \ge 0λi​≥0 and ∑iλi=1\sum_i \lambda_i = 1∑i​λi​=1. The convex hull conv⁡(T)\operatorname{conv}(T)conv(T) of a set T⊆RnT \subseteq \mathbb{R}^nT⊆Rn is the set of all convex combinations of all finite collections of points of TTT. The diameter of TTT is diam⁡(T)=sup⁡{∥s−t∥2:s,t∈T}\operatorname{diam}(T) = \sup\{\|s-t\|_2 : s, t \in T\}diam(T)=sup{∥s−t∥2​:s,t∈T}, the Euclidean norm throughout.

The classical Carathéodory theorem states that every x∈conv⁡(T)x \in \operatorname{conv}(T)x∈conv(T) is a convex combination of at most n+1n+1n+1 points of TTT — with n+1n+1n+1 generally unavoidable, attained by a simplex. The question this mission answers is different: given that we are willing to approximate xxx rather than represent it exactly, and willing to use only combinations with equal coefficients 1/k1/k1/k (an average of kkk points, with repetition allowed), how large must kkk be as a function of the desired accuracy?

Formalization targets

Goal — Theorem 0.0.2, Approximate Carathéodory's theorem

diam(T)≤1, x∈conv⁡(T), k∈Z>0 ⟹ ∃ x1,…,xk∈T:∥x−1k∑j=1kxj∥2≤1k.\text{diam}(T) \le 1,\ x \in \operatorname{conv}(T),\ k \in \mathbb{Z}_{>0} \ \Longrightarrow\ \exists\, x_1,\dots,x_k \in T:\quad \left\| x - \frac{1}{k}\sum_{j=1}^{k} x_j \right\|_2 \le \frac{1}{\sqrt{k}}.diam(T)≤1, x∈conv(T), k∈Z>0​ ⟹ ∃x1​,…,xk​∈T:​x−k1​j=1∑k​xj​​2​≤k​1​.

The quantifiers are exactly this order: for every bounded TTT, every point of its convex hull, and every target kkk, such an averaging set exists. This is the weakest stable statement carrying the theorem's content — the number of points kkk does not depend on the dimension nnn, and the coefficients are forced to be uniform — and it is the form the corollary below invokes directly.

Milestone — Corollary 0.0.4, Covering polytopes by balls

P=conv⁡(T), ∣T∣=N, diam⁡(P)≤1, ε>0 ⟹ ∃ C, ∣C∣≤N⌈1/ε2⌉:P⊆⋃c∈CB‾(c,ε).P = \operatorname{conv}(T),\ |T| = N,\ \operatorname{diam}(P) \le 1,\ \varepsilon > 0 \ \Longrightarrow\ \exists\, C,\ |C| \le N^{\lceil 1/\varepsilon^2 \rceil}:\quad P \subseteq \bigcup_{c \in C} \overline{B}(c, \varepsilon).P=conv(T), ∣T∣=N, diam(P)≤1, ε>0 ⟹ ∃C, ∣C∣≤N⌈1/ε2⌉:P⊆c∈C⋃​B(c,ε).

This is a direct application of the goal to computational geometry's covering problem: how many balls of radius ε\varepsilonε are needed to cover a polytope, and where should they be centered.

Significance

The result itself. The approximate Carathéodory theorem is the prototype of a dimension-free approximation result: whenever a set is bounded, a fixed number of points (depending only on the target accuracy, not the ambient dimension) suffices to approximate any point of its convex hull. This is what makes possible dimension-independent covering-number bounds such as Corollary 0.0.4, which in turn are the starting point for the book's later treatment of entropy, packing and generic chaining (Chapters 4, 7–8). The technique generalizes far beyond Rn\mathbb{R}^nRn: it underlies covering-number bounds for operators between Banach spaces (Carl's original application) and is a recurring device in learning theory for bounding the size of an ε\varepsilonε-net of a hypothesis class.

Formalizing it. Both results are elementary and already fully proved in the literature; no open mathematical content remains. What this mission contributes is a machine-checked, faithful Lean statement of Maurey's construction and its corollary, phrased over Mathlib's existing convex-hull and metric-diameter machinery, so that later missions in this series (concentration inequalities, Johnson–Lindenstrauss, chaining) can build on a verified base case of "probability constructs a deterministic covering," and so that the empirical method itself becomes a reusable, linked component on the platform. The proof of the goal (via the probabilistic argument sketched by the book: interpret a convex combination as a probability distribution, average kkk i.i.d. samples, and bound the variance) is left open for solvers.

Difficulty

The identity that makes the proof work — averaging kkk independent copies of a random vector concentrates around its mean at rate 1/k1/\sqrt{k}1/k​ in mean-square — is a two-line computation once the convex combination is reinterpreted probabilistically. The step that is easy to miss is this reinterpretation itself: nothing in the statement mentions probability, so the "obvious" attack of manipulating the convex-combination weights directly, or trying to construct x1,…,xkx_1,\dots,x_kx1​,…,xk​ by some explicit combinatorial recipe, does not see a path to a bound independent of nnn. The probabilistic argument produces the points non-constructively, via an averaging/existence argument (the expected squared distance is small, so some realization achieves it) rather than an explicit formula — a solver has to introduce a probability space and a random vector that does not appear anywhere in the formal statement to be proved.

Formalization scope

Both results are stated over EuclideanSpace ℝ (Fin n) for an explicit dimension n : ℕ, so ‖·‖ is the Euclidean norm and Mathlib's Metric.diam is used directly for diam⁡(T)=sup⁡{∥s−t∥2}\operatorname{diam}(T) = \sup\{\|s-t\|_2\}diam(T)=sup{∥s−t∥2​}. The convex hull is Mathlib's convexHull ℝ T; by Mathlib's convexHull_eq, this already coincides with the book's own definition of a convex combination of finitely many points of TTT, so no bespoke convex-combination definition is introduced — this mission needs no supporting definitions of its own. In the corollary, "a polytope PPP with NNN vertices" is formalized, following the book's own proof, as P=conv⁡(T)P = \operatorname{conv}(T)P=conv(T) for a finite vertex set TTT with #T=N\#T = N#T=N, rather than via a separate Polytope structure (which Mathlib does not provide and the book's argument does not need). The covering bound N⌈1/ε2⌉N^{\lceil 1/\varepsilon^2\rceil}N⌈1/ε2⌉ is an exponent, not a product with NNN — matching the book's own proof, which counts the NkN^kNk ordered kkk-tuples of vertices with repetition, k:=⌈1/ε2⌉k := \lceil 1/\varepsilon^2\rceilk:=⌈1/ε2⌉; the typeset "N⌈1/ε2⌉N\lceil 1/\varepsilon^2\rceilN⌈1/ε2⌉" in the corollary statement is the same juxtaposition-as-exponent notation the proof uses for "NkN^kNk" one line earlier.

A trivializing formalization is ruled out explicitly: the goal must hold for every integer k>0k > 0k>0 and every x∈conv⁡(T)x \in \operatorname{conv}(T)x∈conv(T), not merely some convenient choice — e.g. k=1k = 1k=1 together with x∈Tx \in Tx∈T trivially satisfies the inequality but proves nothing about the theorem's actual content, that a fixed, dimension-independent kkk works uniformly over all points of the hull. The formal statement quantifies TTT, then xxx, then kkk, and only then asserts existence of the x1,…,xkx_1,\dots,x_kx1​,…,xk​, exactly in that order.

Classical Carathéodory (Theorem 0.0.1, stated for context in the source but not used by either formalized result's proof) is not drafted here: Mathlib already proves the corresponding statement via affine independence (Caratheodory.eq_pos_convex_span_of_mem_convexHull, Analysis/Convex/Caratheodory.lean), from which the book's "n+1n+1n+1 points" bound follows via AffineIndependent.card_le_finrank_succ. It is not added as a kind: reference milestone because no corresponding theorem is yet published on the Prove2Me platform to point at (checked 2026-09-17: GET /theorems?q=Caratheodory returns only unrelated tropical-convexity results), and re-drafting existing Mathlib content as a new platform theorem would duplicate rather than reuse it.

Selected references

  • R. Vershynin, High-Dimensional Probability: An Introduction with Applications in Data Science, Cambridge University Press, 2018, DOI 10.1017/9781108231596, Appetizer (pp. 1–5).
  • G. Pisier, "Remarques sur un résultat non publié de B. Maurey," Séminaire d'Analyse Fonctionnelle (Maurey–Schwartz), 1980–1981, exposé no. 5. numdam.org/item/SAF_1980-1981____A5_0
  • B. Carl, "Inequalities of Bernstein-Jackson-type and the degree of compactness of operators in Banach spaces," Annales de l'Institut Fourier, 35(3), 1985, 79–118. numdam.org/item/AIF_1985__35_3_79_0
2 thms2 active usersReviewed
🏆Completed
Convex OptimizationOptimization·Captain: mikedeng1

Introduction to Online Convex Optimization III: Online Gradient DescentTextbook

Motivation

Online convex optimization (OCO) models a repeated decision process: at each round a learner picks a point in a convex set, an adversary (or the world) reveals a convex cost function, the learner pays that cost at its own point, and the process repeats. No statistical assumption on the sequence of costs is made. This model, introduced by Zinkevich [Zinkevich, Online Convex Programming and Generalized Infinitesimal Gradient Ascent, ICML 2003], underlies most of modern online learning: portfolio selection, online routing, and — through its special case of stochastic optimization — the training of essentially every large machine-learning model in current use, since stochastic gradient descent (the subject of §3.4 of this chapter) is exactly an application of the regret bounds proved here.

The algorithm this chapter introduces, online gradient descent (OGD), is the field's default answer: take a gradient step against the most recently observed cost, project back onto the feasible set. It predates OCO itself as a heuristic, but Zinkevich's contribution — and this chapter's — is the regret analysis: a guarantee that holds against every sequence of costs, adversarially chosen, with an explicit, small constant. Precursors for less general settings appear in Kivinen and Warmuth [1997]; logarithmic-regret algorithms for OCO, the subject of §3.3 here, are due to Hazan, Agarwal and Kale [2007].

The online convex optimization protocol

Fix a convex set KKK in a real inner product space, playing the role of the decision (or "action") space, and a sequence of cost functions f1,f2,⋯:K→Rf_1, f_2, \dots : K \to \mathbb{R}f1​,f2​,⋯:K→R, each convex. At round ttt, the learner (not knowing ftf_tft​) plays a point xt∈Kx_t \in Kxt​∈K, then observes ftf_tft​ and pays ft(xt)f_t(x_t)ft​(xt​). Regret after TTT rounds compares the learner's cumulative cost to that of the single best fixed decision made with hindsight of the whole sequence:

RegretT=∑t=1Tft(xt)−min⁡x⋆∈K∑t=1Tft(x⋆).\mathrm{Regret}_T = \sum_{t=1}^{T} f_t(x_t) - \min_{x^\star \in K} \sum_{t=1}^{T} f_t(x^\star).RegretT​=t=1∑T​ft​(xt​)−x⋆∈Kmin​t=1∑T​ft​(x⋆).

A learner with regret o(T)o(T)o(T) is, on average, eventually as good as the best fixed point in KKK, even though it never knew the cost sequence in advance.

Two chapter-wide parameters bound how hard an instance can be: DDD, the diameter of KKK (dist⁡(x,y)≤D\operatorname{dist}(x,y) \le Ddist(x,y)≤D for all x,y∈Kx, y \in Kx,y∈K), and GGG, a common bound on the gradient norm of every ftf_tft​ over KKK (∥∇ft(x)∥≤G\|\nabla f_t(x)\| \le G∥∇ft​(x)∥≤G for all x∈Kx \in Kx∈K), which implies but is strictly stronger than GGG-Lipschitzness on KKK. Online gradient descent (Algorithm 8) plays x1∈Kx_1 \in Kx1​∈K arbitrarily, then at every round sets yt+1=xt−ηt∇ft(xt)y_{t+1} = x_t - \eta_t \nabla f_t(x_t)yt+1​=xt​−ηt​∇ft​(xt​) and projects, xt+1=ΠK(yt+1)x_{t+1} = \Pi_K(y_{t+1})xt+1​=ΠK​(yt+1​), for a sequence of step sizes ηt\eta_tηt​ chosen in advance.

Formalization targets

Goal: Theorem 3.1 (online gradient descent regret)

RegretT≤32GDTfor all T≥1,\mathrm{Regret}_T \le \frac{3}{2} G D \sqrt{T} \quad \text{for all } T \ge 1,RegretT​≤23​GDT​for all T≥1,

using step sizes ηt=D/(Gt)\eta_t = D / (G\sqrt{t})ηt​=D/(Gt​). This is the chapter's — and arguably the book's — central result: the simplest algorithm for the fully general OCO protocol already attains O(T)O(\sqrt{T})O(T​) regret, with an explicit small constant, against convex Lipschitz costs with no further structure.

Milestone: Theorem 3.2 (matching lower bound)

Any algorithm for OCO incurs Ω(DGT)\Omega(DG\sqrt{T})Ω(DGT​) regret in the worst case: no algorithm, however clever, can improve asymptotically on Theorem 3.1's rate. This is the weaker, worst-case-existence half of the theorem (see Formalization scope).

Milestone: Theorem 3.3 (logarithmic regret under strong convexity)

If every ftf_tft​ is additionally α\alphaα-strongly convex, the same algorithm — with only the step sizes changed to ηt=1/(αt)\eta_t = 1/(\alpha t)ηt​=1/(αt) — achieves

RegretT≤G22α(1+log⁡T).\mathrm{Regret}_T \le \frac{G^2}{2\alpha}(1 + \log T).RegretT​≤2αG2​(1+logT).

Strong convexity is a strictly stronger hypothesis than convexity, so this target does not subsume the goal; it sits alongside it as the chapter's second, sharper regime.

Significance

Theorem 3.1 is the reference point against which every later algorithm and every later chapter's improvement (Online Newton Step, RFTL, adaptive-regret methods) is measured: any new algorithm for OCO is judged first by whether it matches this O(T)O(\sqrt{T})O(T​) rate, then by what extra structure lets it do better. Theorem 3.2 closes the question for the general convex-Lipschitz class: O(T)O(\sqrt{T})O(T​) is not an artifact of a loose analysis, it is information-theoretically necessary. Theorem 3.3 identifies the one extra hypothesis (strong convexity) that buys an exponential improvement in the horizon dependence, from T\sqrt{T}T​ to log⁡T\log TlogT, without any other change to the algorithm — the same phenomenon that in the book's Chapter 2 separated well-conditioned from general convex offline optimization, now transplanted to the online, adversarial setting.

None of these three statements has a machine-checked proof on the platform prior to this mission (see Formalization scope below for what was checked). Formalizing them establishes the regret protocol and the OGD algorithm as reusable definitions for the rest of this thirteen-chapter series, several chapters of which (Online Newton Step, RFTL, bandit convex optimization) build directly on Algorithm 8 or its regret guarantee.

Difficulty

The regret bound's proof (Theorem 3.1) is short but not naive: bounding ∇t⊤(xt−x⋆)\nabla_t^\top(x_t - x^\star)∇t⊤​(xt​−x⋆) by convexity alone gives no telescoping structure, so the argument instead bounds it using the projection step — the Pythagorean inequality ∥ΠK(z)−x⋆∥≤∥z−x⋆∥\|\Pi_K(z) - x^\star\| \le \|z - x^\star\|∥ΠK​(z)−x⋆∥≤∥z−x⋆∥ for x⋆∈Kx^\star \in Kx⋆∈K — applied to the specific point z=xt−ηt∇tz = x_t - \eta_t \nabla_tz=xt​−ηt​∇t​. This turns the per-round convexity bound into a telescoping sum in ∥xt−x⋆∥2\|x_t - x^\star\|^2∥xt​−x⋆∥2, and only the resulting sum, evaluated with the specific step-size schedule ηt=D/(Gt)\eta_t = D/(G\sqrt{t})ηt​=D/(Gt​), produces the T\sqrt{T}T​ rate; a constant or linearly growing step size does not. The same projection argument is reused for Theorem 3.3, where the strong-convexity inequality is engineered to make the ∥x⋆−xt∥2\|x^\star - x_t\|^2∥x⋆−xt​∥2 terms cancel exactly against the projection telescoping, leaving a harmonic sum. Theorem 3.2's difficulty is of a different kind: it is a lower bound over every algorithm, proved by exhibiting a randomized hard instance (the hypercube with 2n2^n2n sign-vector linear costs) on which no algorithm can do better than random guessing in expectation.

Formalization scope

KKK is formalized as a subset of an arbitrary real, complete inner product space (not fixed to Rn\mathbb{R}^nRn), since the chapter's argument uses only Hilbert-space structure. Rounds are 0-indexed (Finset.range T) rather than the book's 1-indexed rounds, so a step size stated here at round ttt is the book's step size at round t+1t+1t+1. D and G are carried as shared section hypotheses (the chapter-wide diameter and gradient-norm bounds), not re-derived or re-stated per theorem; G is formalized exactly as the book defines it (p. 20: a bound on ∥∇ft(x)∥\|\nabla f_t(x)\|∥∇ft​(x)∥ over KKK, via Mathlib's HasGradientAt), not as the weaker two-point Lipschitz condition it implies — an earlier draft used the weaker Lipschitz hypothesis and was corrected during moderation, since it made the drafted theorems strictly stronger than the book's own (true by an added argument the book does not give, but not faithful to the stated proof). The projection step is formalized relationally (IsMetricProjection, an arbitrary closest point) rather than as a canonical function, since a general convex set need not come with one built into Mathlib.

For Theorem 3.2, this mission formalizes the theorem's main sentence — the worst-case existence claim — quantifying over "any algorithm" as a non-anticipating map from the full cost sequence to the play sequence, with the hard cost sequence existentially quantified after the algorithm and the horizon: for every algorithm and every horizon there is a cost sequence forcing Ω(DGT)\Omega(DG\sqrt{T})Ω(DGT​) regret against it. (An earlier draft quantified the cost sequence first — one fixed sequence defeating every algorithm — which is false: a constant algorithm playing a minimizer of that one sequence has zero regret against it; this was corrected during moderation.) It does not formalize the theorem's parenthetical strengthening, that the same Ω(DGT)\Omega(DG\sqrt{T})Ω(DGT​) bound holds even when costs are drawn from a fixed stationary distribution; that claim is about expected regret of a deterministic algorithm against a random cost sequence, and would need a probability-space formalization of the OCO protocol that this mission's definitions do not build. A formalization limited to the deterministic worst case does not trivialize the theorem: it is exactly the inequality "O(T)O(\sqrt{T})O(T​) cannot be improved," stated without the randomization machinery of its proof.

Theorem 3.4 (the stochastic gradient descent corollary, via a noisy gradient oracle with bounded second moment) is not included: it needs an expectation over a random oracle applied at a random, round-dependent point, which is a substantially heavier probabilistic object than the deterministic protocol built here, and is left for a future mission or an extension of this one.

Selected references

  • Zinkevich, M. Online Convex Programming and Generalized Infinitesimal Gradient Ascent. ICML 2003. https://www.aaai.org/Papers/ICML/2003/ICML03-120.pdf
  • Kivinen, J. and Warmuth, M. K. Exponentiated Gradient Versus Gradient Descent for Linear Predictors. Information and Computation, 1997. https://doi.org/10.1006/inco.1996.2612
  • Hazan, E., Agarwal, A. and Kale, S. Logarithmic Regret Algorithms for Online Convex Optimization. Machine Learning 69, 2007. https://doi.org/10.1007/s10994-007-5016-8
  • Hazan, E. Introduction to Online Convex Optimization, 2nd ed. arXiv:1909.05207v3, Chapter 3. https://arxiv.org/abs/1909.05207
5 thms2 active usersReviewed
🏆Completed
Bandit AlgorithmsReinforcement LearningStatistics·Captain: mikedeng1

Foundations of Reinforcement Learning II: Contextual Bandits and Inverse Gap WeightingTextbook

Motivation

Decision-making problems rarely present the same fixed choice twice. A doctor prescribing a treatment sees each patient's medical history and symptoms before deciding; a website choosing which article to show sees the visitor's profile first. The multi-armed bandit model — where the learner repeatedly picks from a fixed set of arms with no side information — cannot express this: it is blind to the covariates that any real decision-maker actually observes. The contextual bandit model closes this gap by letting the learner see a context before acting, and asks for a decision rule that generalizes across contexts rather than memorizing a policy per context. Foster and Rakhlin's Foundations of Reinforcement Learning and Interactive Decision Making (arXiv:2312.16730v1, Section 3, pp. 38–53) develops this model and its algorithms as the bridge between supervised learning and sequential decision making, en route to general reinforcement learning. Contextual bandits with a learned reward-function class underlie production systems for content recommendation, online advertising, and adaptive clinical trial design (Li et al., A Contextual-Bandit Approach to Personalized News Article Recommendation, 2010, https://arxiv.org/abs/1003.0146; Agarwal et al., Making Contextual Decisions with Low Technical Debt, 2016, https://arxiv.org/abs/1606.03966).

The algorithmic history in this chapter runs through two distinct principles. The optimism principle (LinUCB, Section 3.2) generalizes the UCB algorithm to contexts under a linear reward model, but the chapter's own Example 3.1 (Section 3.3) shows optimism fails outside such structured classes, incurring regret linear in the size of the context space or the class. Foster and Rakhlin then present two "black-box" alternatives that use any function class FFF through an abstract regression subroutine: the naive ε\varepsilonε-Greedy method (Section 3.4), and the Inverse Gap Weighting (IGW) strategy underlying the SquareCB algorithm (Bietti, Agarwal & Langford, A Contextual Bandit Bake-off, 2018, https://arxiv.org/abs/1802.04064; Foster & Rakhlin, Beyond UCB: Optimal and Efficient Contextual Bandits with Regression Oracles, 2020, https://arxiv.org/abs/2002.04926). SquareCB attains a regret rate that both generalizes across contexts (no dependence on the size of the context space) and matches the optimal T\sqrt{T}T​ rate — improving on ε\varepsilonε-Greedy's T2/3T^{2/3}T2/3 rate — while remaining agnostic to the internal structure of FFF.

Setting

Over TTT rounds, a decision-maker faces the contextual bandit protocol: at each round ttt, it observes a context xt∈Xx_t \in Xxt​∈X, selects a decision πt\pi_tπt​ from a finite action set Π={1,…,A}\Pi = \{1,\dots,A\}Π={1,…,A}, and observes a reward rt∈Rr_t \in \mathbb{R}rt​∈R. Rewards are generated independently as rt∼M⋆(⋅∣xt,πt)r_t \sim M^\star(\cdot \mid x_t, \pi_t)rt​∼M⋆(⋅∣xt​,πt​) for a fixed, unknown conditional model M⋆M^\starM⋆; write f⋆(x,π):=E[r∣x,π]f^\star(x,\pi) := \mathbb{E}[r \mid x, \pi]f⋆(x,π):=E[r∣x,π] for the mean reward function and π⋆(x):=arg⁡max⁡πf⋆(x,π)\pi^\star(x) := \arg\max_\pi f^\star(x,\pi)π⋆(x):=argmaxπ​f⋆(x,π) for the optimal, context-dependent policy. The context sequence x1,…,xTx_1,\dots,x_Tx1​,…,xT​ is arbitrary — fixed in advance or adversarially chosen — while rewards remain stochastic. Performance is measured by regret against π⋆\pi^\starπ⋆:

Reg:=∑t=1Tf⋆(xt,π⋆(xt))−∑t=1TEπt∼pt[f⋆(xt,πt)],\mathrm{Reg} := \sum_{t=1}^T f^\star(x_t,\pi^\star(x_t)) - \sum_{t=1}^T \mathbb{E}_{\pi_t\sim p_t}[f^\star(x_t,\pi_t)],Reg:=t=1∑T​f⋆(xt​,π⋆(xt​))−t=1∑T​Eπt​∼pt​​[f⋆(xt​,πt​)],

where ptp_tpt​ is the learner's (possibly randomized) action distribution at round ttt.

To generalize across contexts, the learner is given a class F⊆{f:X×Π→R}F \subseteq \{f : X\times\Pi \to \mathbb{R}\}F⊆{f:X×Π→R} with f⋆∈Ff^\star \in Ff⋆∈F, and aims for regret scaling with the statistical complexity log⁡∣F∣\log|F|log∣F∣ rather than with ∣X∣|X|∣X∣. Both algorithms in this mission access FFF only through an online regression oracle (Definition 3, p. 47): given the history (x1,π1,r1),…,(xt−1,πt−1,rt−1)(x_1,\pi_1,r_1),\dots,(x_{t-1},\pi_{t-1},r_{t-1})(x1​,π1​,r1​),…,(xt−1​,πt−1​,rt−1​), it returns an estimate f^t:X×Π→R\hat f_t : X\times\Pi\to\mathbb{R}f^​t​:X×Π→R satisfying, with probability at least 1−δ1-\delta1−δ, ∑t=1TEπt∼pt[(f^t(xt,πt)−f⋆(xt,πt))2]≤EstSq(F,T,δ)\sum_{t=1}^T \mathbb{E}_{\pi_t\sim p_t}[(\hat f_t(x_t,\pi_t)-f^\star(x_t,\pi_t))^2] \le \mathrm{EstSq}(F,T,\delta)∑t=1T​Eπt​∼pt​​[(f^​t​(xt​,πt​)−f⋆(xt​,πt​))2]≤EstSq(F,T,δ) — for instance, exponential weights on a finite class FFF achieves EstSq(F,T,δ)=log⁡(∣F∣/δ)\mathrm{EstSq}(F,T,\delta) = \log(|F|/\delta)EstSq(F,T,δ)=log(∣F∣/δ). SquareCB (p. 50–51) then samples its action from the Inverse Gap Weighting distribution (Definition 4, p. 50): given a vector of estimated values f^∈RA\hat f \in \mathbb{R}^Af^​∈RA with greedy action πˉ=arg⁡max⁡πf^(π)\bar\pi = \arg\max_\pi \hat f(\pi)πˉ=argmaxπ​f^​(π), and an exploration parameter γ≥0\gamma \ge 0γ≥0, p=IGWγ(f^)p = \mathrm{IGW}_\gamma(\hat f)p=IGWγ​(f^​) is p(π)=1/(λ+2γ(f^(πˉ)−f^(π)))p(\pi) = 1/(\lambda + 2\gamma(\hat f(\bar\pi)-\hat f(\pi)))p(π)=1/(λ+2γ(f^​(πˉ)−f^​(π))) for the unique λ∈[1,A]\lambda \in [1,A]λ∈[1,A] making ppp a probability distribution.

Formalization targets

Milestone — Proposition 9 (IGW estimation-to-regret inequality)

Eπ∼p[f⋆(π⋆)−f⋆(π)]≤Aγ+γ⋅Eπ∼p[(f^(π)−f⋆(π))2],p=IGWγ(f^).\mathbb{E}_{\pi\sim p}[f^\star(\pi^\star)-f^\star(\pi)] \le \frac{A}{\gamma} + \gamma\cdot\mathbb{E}_{\pi\sim p}[(\hat f(\pi)-f^\star(\pi))^2], \qquad p = \mathrm{IGW}_\gamma(\hat f).Eπ∼p​[f⋆(π⋆)−f⋆(π)]≤γA​+γ⋅Eπ∼p​[(f^​(π)−f⋆(π))2],p=IGWγ​(f^​).

This holds for any f^,f⋆∈RA\hat f, f^\star \in \mathbb{R}^Af^​,f⋆∈RA and any γ>0\gamma>0γ>0, with no reference to FFF or to how f^\hat ff^​ was produced — it is the purely algebraic core the goal theorem invokes at every round.

Goal — Proposition 10 (SquareCB regret bound)

Reg≤2A T EstSq(F,T,δ)\mathrm{Reg} \le 2\sqrt{A\,T\,\mathrm{EstSq}(F,T,\delta)}Reg≤2ATEstSq(F,T,δ)​

with probability at least 1−δ1-\delta1−δ, for SquareCB run with γ=TA/EstSq(F,T,δ)\gamma = \sqrt{TA/\mathrm{EstSq}(F,T,\delta)}γ=TA/EstSq(F,T,δ)​, for any context sequence x1,…,xTx_1,\dots,x_Tx1​,…,xT​. This is the weakest stable target level in the chapter's oracle-based development: it is stated for an arbitrary class FFF and oracle, so it survives any future improvement to the oracle's own EstSq\mathrm{EstSq}EstSq bound, unlike a version hard-coded to a specific class or oracle.

Significance

Proposition 10 shows that Inverse Gap Weighting converts any estimation-error guarantee into a regret guarantee with the same statistical rate, with no algorithm-side dependence on the structure of FFF or the size of XXX: the same SquareCB template, driven by a plug-in regression oracle, is minimax optimal whenever the oracle itself is. When FFF is finite, this yields Reg≲ATlog⁡(∣F∣/δ)\mathrm{Reg} \lesssim \sqrt{AT\log(|F|/\delta)}Reg≲ATlog(∣F∣/δ)​, matching the optimal rate for stochastic multi-armed bandits (Section 2) while generalizing across contexts — a guarantee that optimism (Proposition 7) provably cannot deliver outside linear classes (Example 3.1), and that the simpler ε\varepsilonε-Greedy baseline (Proposition 8) only delivers at a slower T2/3T^{2/3}T2/3 rate. Foster and Rakhlin describe Proposition 9 itself as being "at the core of the development for the rest of the course": the same IGW mechanism reappears, generalized, in the book's treatment of general decision-making and the Decision-Estimation Coefficient.

Both propositions are proved results, not open questions; this mission's contribution is a machine-checked formalization of their exact statements and hypotheses — the precise OracleGuarantee hypothesis Proposition 10 requires, the exact constant (222, not a bare ≲\lesssim≲) its proof yields at the stated optimal γ\gammaγ, and the universally-quantified form of the IGW inequality (Proposition 9) that makes it reusable independently of any particular oracle or class.

Difficulty

The obvious first idea for exploiting an estimator f^t\hat f_tf^​t​ is a UCB-style optimism approach: build a confidence set around f^t\hat f_tf^​t​ and act greedily on its upper envelope, as in LinUCB (Proposition 7). Example 3.1 shows this fails in general: a class FFF can force the confidence set to remain wide on a fresh action at every new context, driving regret linear in min⁡{∣F∣,∣X∣}\min\{|F|,|X|\}min{∣F∣,∣X∣} — the confidence width in the regret bound does not shrink merely because the oracle's cumulative estimation error is small, since that error is not localized to the specific action the confidence-set approach tries next. Uniform exploration (ε\varepsilonε-Greedy) sidesteps this but wastes exploration budget on actions already known to be far from optimal, which is what caps its rate at T2/3T^{2/3}T2/3 (Proposition 8). Inverse Gap Weighting instead ties the sampling probability itself to the estimated gap from the greedy action, so cheap-to-rule-out actions are down-weighted continuously rather than either fully explored (ε-Greedy) or trusted outright (optimism); the technical content of Proposition 9 is showing this specific reciprocal-gap form gives a bound with no hidden dependence on FFF or XXX, for every pair (f^,f⋆)(\hat f, f^\star)(f^​,f⋆) simultaneously — a guarantee optimism cannot match because its confidence sets are class-dependent by construction.

Formalization scope

Contexts form an arbitrary type X; actions are Fin A for A : ℕ. A finite probability distribution over Fin A is represented directly as p : Fin A → ℝ with ∀ π, 0 ≤ p π and ∑ π, p π = 1, and Eπ∼p[g]\mathbb{E}_{\pi\sim p}[g]Eπ∼p​[g] as the finite sum ∑ π, p π * g π, rather than via Mathlib's PMF (which is ℝ≥0∞-valued) — an equivalent and lighter-weight representation of a distribution on a finite type. The normalizing constant λ\lambdaλ of Definition 4 and the optimal actions π⋆\pi^\starπ⋆, πˉ\bar\piπˉ are each specified by their defining property (existence of λ∈[1,A]\lambda \in [1,A]λ∈[1,A] realizing the IGW formula; ∀π,f(π)≤f(argmax)\forall\pi, f(\pi)\le f(\text{argmax})∀π,f(π)≤f(argmax)) rather than constructed explicitly via an intermediate-value or Finset.argmax argument, avoiding committing to one choice function for a value the book itself leaves implicit. The class FFF enters neither proposition's statement directly: it appears in the source only through the abstract bound EstSq(F,T,δ)\mathrm{EstSq}(F,T,\delta)EstSq(F,T,δ), which is carried as an explicit real-valued parameter and hypothesis (OracleGuarantee) rather than as a literal subset of a function space, since no property of FFF beyond producing this bound is ever used. The probability-(1−δ)(1-\delta)(1−δ) qualifier attached to the online regression oracle's guarantee is likewise the explicit hypothesis OracleGuarantee ... EstSq on a fixed realized run, rather than a statement quantified over an underlying probability space of histories — every subsequent step in both propositions' proofs is deterministic given that this event holds, so this does not weaken either conclusion. A trivializing formalization would fix A=1A=1A=1 (a single ever-optimal action, making both Reg and the IGW inequality vacuous) or take EstSq as an unconstrained free variable with no positivity hypothesis (making γ\gammaγ in Proposition 10 undefined); this mission's statements require 0 < EstSq and leave AAA, TTT, XXX, FFF-via-EstSq fully general.

This mission omits Proposition 7 (LinUCB): its proof rests on an entirely disjoint apparatus (finite linear parameter sets, least-squares confidence sets, the elliptic potential lemma) that neither Proposition 9 nor 10 requires, and Example 3.1 (the failure of optimism) is a worked example rather than a numbered, formalizable claim. It also omits Proposition 8 (ε\varepsilonε-Greedy): the source leaves the optimal ε\varepsilonε unspecified ("choosing ε\varepsilonε appropriately"), and deriving its own optimal value and matching constant independently — rather than reusing the book's own explicit constant, as Rule 7 of this formalization effort requires — was judged too likely to introduce an unfaithful, invented constant within this mission's time budget; both are natural extensions for a follow-up mission or contribution. Reusable infrastructure: the Fin A-indexed finite-distribution convention and the OracleGuarantee/optimal-action-by-property pattern extend directly to any later chapter built on the same online-regression-oracle abstraction.

Selected references

  • Foster, D. J. and Rakhlin, A. Foundations of Reinforcement Learning and Interactive Decision Making. 2023. https://arxiv.org/abs/2312.16730
  • Foster, D. J. and Rakhlin, A. Beyond UCB: Optimal and Efficient Contextual Bandits with Regression Oracles. ICML 2020. https://arxiv.org/abs/2002.04926
  • Bietti, A., Agarwal, A., and Langford, J. A Contextual Bandit Bake-off. JMLR 2021 (arXiv 2018). https://arxiv.org/abs/1802.04064
  • Li, L., Chu, W., Langford, J., and Schapire, R. E. A Contextual-Bandit Approach to Personalized News Article Recommendation. WWW 2010. https://arxiv.org/abs/1003.0146
  • Agarwal, A. et al. Making Contextual Decisions with Low Technical Debt. 2016. https://arxiv.org/abs/1606.03966
5 thms2 active usersReviewed
🏆Completed
Bandit AlgorithmsReinforcement Learning·Captain: mikedeng1

Foundations of Reinforcement Learning I: Multi-Armed Bandits and the UCB AlgorithmTextbook

Motivation

The multi-armed bandit is the simplest model of sequential decision-making under partial feedback: a learner repeatedly picks one of finitely many options and observes a reward only for the option chosen, never for the alternatives. It formalizes problems ranging from clinical trial design (which treatment to offer a patient) to online advertising (which ad to show) and A/B testing more generally. The framework dates to Robbins' 1952 paper on sequential design, and the algorithm this mission's goal theorem concerns — the Upper Confidence Bound (UCB) algorithm of Lai and Robbins [1985] and Auer, Cesa-Bianchi and Fischer [2002] — is the canonical answer to how to explore efficiently: instead of exploring uniformly at random, act optimistically with respect to the current uncertainty about each option's value. This mission draws its formalization from Chapter 2 of Foster and Rakhlin's 2023 lecture notes, Foundations of Reinforcement Learning and Interactive Decision Making, which develops the bandit problem as the first rung of a ladder of increasingly general interactive decision-making settings (contextual bandits, structured bandits, reinforcement learning) that the book's later chapters build.

Setting

Fix a finite decision (action) space Π={1,…,A}\Pi = \{1,\dots,A\}Π={1,…,A}. In the multi-armed bandit protocol, for each round t=1,…,Tt = 1,\dots,Tt=1,…,T the learner selects a decision πt∈Π\pi_t \in \Piπt​∈Π, possibly at random according to a distribution ptp_tpt​ depending on the history Ht−1=((π1,r1),…,(πt−1,rt−1))H_{t-1} = ((\pi_1,r_1),\dots,(\pi_{t-1},r_{t-1}))Ht−1​=((π1​,r1​),…,(πt−1​,rt−1​)) observed so far, and then observes a reward rt∈Rr_t \in \mathbb{R}rt​∈R drawn independently from a fixed conditional distribution M⋆(⋅∣πt)M^\star(\cdot \mid \pi_t)M⋆(⋅∣πt​) (the stochastic rewards assumption). Writing f⋆(π):=E[r∣π]f^\star(\pi) := \mathbb{E}[r \mid \pi]f⋆(π):=E[r∣π] for the mean reward function and π⋆:=arg⁡max⁡πf⋆(π)\pi^\star := \arg\max_\pi f^\star(\pi)π⋆:=argmaxπ​f⋆(π) for an optimal decision, the learner's performance is measured by the regret

Reg:=∑t=1Tf⋆(π⋆)−∑t=1TEπt∼pt[f⋆(πt)].\mathrm{Reg} := \sum_{t=1}^T f^\star(\pi^\star) - \sum_{t=1}^T \mathbb{E}_{\pi_t \sim p_t}[f^\star(\pi_t)].Reg:=t=1∑T​f⋆(π⋆)−t=1∑T​Eπt​∼pt​​[f⋆(πt​)].

Because the learner observes a reward only for the action played (bandit feedback), a purely greedy strategy that always plays the current empirical maximizer can commit to a suboptimal action forever, incurring linear regret; some form of deliberate exploration is necessary. The chapter's central construction is the confidence interval: a pair of functions f‾t,fˉt:Π→R\underline{f}_t, \bar f_t : \Pi \to \mathbb{R}f​t​,fˉ​t​:Π→R such that, with probability at least 1−δ1-\delta1−δ, f⋆(π)∈[f‾t(π),fˉt(π)]f^\star(\pi) \in [\underline{f}_t(\pi), \bar f_t(\pi)]f⋆(π)∈[f​t​(π),fˉ​t​(π)] for every round ttt and decision π\piπ simultaneously. The UCB algorithm plays the optimistic action πt=arg⁡max⁡πfˉt(π)\pi_t = \arg\max_\pi \bar f_t(\pi)πt​=argmaxπ​fˉ​t​(π) at every round, using the confidence interval built from Hoeffding's inequality around the empirical mean f^t(π)\hat f_t(\pi)f^​t​(π).

Formalization targets

Goal — Proposition 5 (UCB regret)

Reg  ≲  ATlog⁡(AT/δ)\mathrm{Reg} \;\lesssim\; \sqrt{AT\log(AT/\delta)}Reg≲ATlog(AT/δ)​

holding with probability at least 1−δ1-\delta1−δ, for the UCB algorithm using the confidence radius 2log⁡(2T2A/δ)/nt(π)\sqrt{2\log(2T^2A/\delta)/n_t(\pi)}2log(2T2A/δ)/nt​(π)​ of Eq. (2.19). This is the weakest stable statement the chapter proves: it is optimal up to the log factor, and strengthening it (e.g. to the sharper instance-dependent bound of Remark 10) is explicitly left to later work by the book itself.

Milestones

  • Proposition 4 (ε-Greedy regret): Reg≲A1/3T2/3log⁡1/3(AT/δ)\mathrm{Reg} \lesssim A^{1/3}T^{2/3}\log^{1/3}(AT/\delta)Reg≲A1/3T2/3log1/3(AT/δ) — the book's preceding, weaker result, establishing that naive forced exploration already gives sublinear regret, and motivating why an adaptive strategy (UCB) does better.
  • Lemma 7 (Optimism): the per-round regret of the optimistic action is bounded by the confidence width at that action.
  • Lemma 8 (Confidence width potential lemma): ∑t=1T(1/nt(πt)∧1)≲AT\sum_{t=1}^T (1/\sqrt{n_t(\pi_t)} \wedge 1) \lesssim \sqrt{AT}∑t=1T​(1/nt​(πt​)​∧1)≲AT​, a pigeonhole bound on how often any one action's confidence interval can still be wide.

Significance

UCB is the prototype of the "optimism in the face of uncertainty" principle that recurs, in increasingly abstract form, throughout the rest of the book: the same two-step argument (Lemma 7 + Lemma 8) reappears for linear bandits, structured bandits via the Decision-Estimation Coefficient, and UCB-VI for tabular reinforcement learning. Formalizing Chapter 2 in full therefore front-loads the proof pattern every later chapter in this series specializes. The result itself is also of standalone interest: the AT\sqrt{AT}AT​ minimax rate is the benchmark every subsequent bandit algorithm in the literature is compared against, and the A1/3T2/3A^{1/3}T^{2/3}A1/3T2/3-vs-AT\sqrt{AT}AT​ contrast between ε-Greedy and UCB is the standard illustration, in any course on the subject, of why adaptive exploration matters.

No formalization of this exact statement — realizability with respect to a function class f⋆∈F=RΠf^\star \in \mathcal{F} = \mathbb{R}^\Pif⋆∈F=RΠ and a generic confidence interval, rather than a per-arm sub-Gaussian empirical mean — currently exists on the platform (see Formalization scope below); the mission both proves this specific regret bound and seeds the generic optimism/potential lemma pair (Lemma 7, Lemma 8) that the book's later, more structured settings specialize.

Difficulty

The natural first attempt — bound the regret of the empirical-mean-greedy algorithm directly — fails outright: on a two-armed instance where one arm is deterministic and the other only slightly better in expectation, the greedy algorithm can commit to the worse arm forever with constant probability, giving linear, not sublinear, regret (§2.1). The obvious fix, ε-Greedy, forces exploration uniformly across all actions regardless of how much is already known about each, so the exploration cost scales with εT\varepsilon TεT even for actions whose value is already well determined — this is exactly what caps ε-Greedy at the T2/3T^{2/3}T2/3 rate. UCB's optimism principle resolves this by exploring an action only in proportion to how uncertain it still is; the technical core, isolated in Lemma 7 and Lemma 8, is disentangling "the algorithm made a mistake" from "the algorithm is still uncertain," which are conflated in the naive per-round regret decomposition used for ε-Greedy.

Formalization scope

Both the goal and the milestones fix a finite decision space Fin A, a mean reward function fStar : Fin A → ℝ with fStar π ∈ [0,1], and an optimal decision piStar. Regret is defined generically (Eq. (2.3)) via per-round decision weights p : ℕ → Fin A → ℝ, so it applies uniformly to a randomized algorithm (ε-Greedy) and a deterministic one (UCB, via the point mass at the played action). The book's "with probability at least 1−δ1-\delta1−δ" qualifier on both Proposition 4 and Proposition 5 is formalized as the deterministic consequence of the underlying concentration event (Eq. (2.9) and Eq. (2.18) respectively) holding — exactly the move the book's own proofs make ("Let us condition on the event in (2.18) ... "). The concentration events themselves rest on Hoeffding's inequality for adaptive stopping times (Lemma 33) and Bernstein's inequality (Lemma 5), both stated in the book's technical appendix outside this chapter, and are not drafted here; a solver may either take them as a hypothesis (as this mission's statements do) or import/prove them separately. A trivializing formalization is ruled out explicitly: taking δ outside (0,1)(0,1)(0,1), or dropping the fStar π ∈ [0,1] hypothesis, would make the stated constants vacuous or false, so both are retained as explicit hypotheses in every theorem. In every ≲ statement (Prop. 4, Lemma 8, Prop. 5) the witnessed constant C is quantified before the instance parameters (A, T, δ, and the realized sequences): ∃ C, 0 < C ∧ ∀ A T δ ..., Reg ≤ C * (rate), not the other order. This is deliberate, not stylistic: quantifying C after the instance lets it depend on A, T, δ, making the bound satisfiable by an arbitrarily large C chosen per instance and hence content-free, which is not what the book's ≲ means (a single constant working uniformly over all instances). Proposition 5's UCB decision rule is stated in the book's own two clauses, not collapsed into a single "maximize the upper confidence bound" rule: the confidence radius of Eq. (2.19) is +∞+\infty+∞ at nt(π)=0n_t(\pi)=0nt​(π)=0 (an action never yet sampled), so the book's UCB always plays an unsampled action before ever comparing indices, and only compares finite upper confidence bounds once every action has been sampled at least once; the confidence event of Eq. (2.18) is correspondingly assumed only at sampled actions, since the book's own bound is vacuous otherwise. An earlier draft instead capped the radius at 111 when nt(π)=0n_t(\pi)=0nt​(π)=0, which is a true statement about a different algorithm (a sampled action can have index above the capped unsampled index), and was corrected to the book's own rule after moderation. Reuse from the platform's existing bandit library (BanditAlgorithm, Lattimore & Szepesvári) is deliberately avoided: that library's UCB (bandit_ucb_regret_bound, bandit_ucb_minimax_regret_bound) is stated for per-arm 1-sub-Gaussian rewards with δ=1/n2\delta = 1/n^2δ=1/n2 fixed by the horizon, whereas this chapter's UCB is stated for a free failure probability δ\deltaδ and a generic confidence-interval abstraction (the multi-armed case being F=RΠ\mathcal{F} = \mathbb{R}^\PiF=RΠ of the book's general realizability framework) — the two are related but not the same statement. Contributions extending the mission with the generic confidence-interval form of Lemma 7/8 applied to other chapters in this series (contextual and structured bandits) are welcome.

Selected references

  • T. Lai and H. Robbins, Asymptotically Efficient Adaptive Allocation Rules, Advances in Applied Mathematics, 1985.
  • P. Auer, N. Cesa-Bianchi, and P. Fischer, Finite-time Analysis of the Multiarmed Bandit Problem, Machine Learning, 2002.
  • D. Foster and A. Rakhlin, Foundations of Reinforcement Learning and Interactive Decision Making, arXiv:2312.16730, 2023. https://arxiv.org/abs/2312.16730
  • T. Lattimore and C. Szepesvári, Bandit Algorithms, Cambridge University Press, 2020.
7 thms2 active usersReviewed
🏆Completed
Probability·Captain: naimengye

Speculative Actions: Cost-Latency Analysis for Agentic SpeculationResearch Paper

Motivation

An LLM agent acting in an environment spends most of its wall-clock time waiting. Each step — a model call, a tool or MCP request, a browser action, sometimes a human reply — must complete before the next can be issued, and the round trips dominate end-to-end latency: a chess game between two reasoning agents runs for hours, and an operating-system tuning task for tens of minutes. When a training or prompt-optimization loop repeats such a run thousands of times, the waiting is the cost.

Speculative actions (Ye, Ahuja, Liargkovas, Lu, Kaffes, Peng, ICLR 2026) transplants a classical systems idea — speculative execution in microprocessors, and speculative decoding for LLM inference — to the agent's environment loop. A cheap, fast speculator guesses the action a slow, authoritative actor is about to produce, the guess is used to launch the next environment call early, and the work is committed only when the actor's real action confirms the guess. The interface stays sequential and lossless; the internals run in parallel.

What makes this a formalization target rather than an engineering report is the paper's §5 cost–latency analysis. Speculating more branches buys hit probability but costs tokens, and the paper derives closed-form expressions for both sides of that trade — a self-contained piece of applied probability sitting underneath a systems paper. This mission asks for those expressions, machine-checked.

Setting

Fix a horizon TTT and index steps t=0,1,…,T−1t = 0, 1, \dots, T-1t=0,1,…,T−1. At each step a policy maps the state to an API call; the actor executes it with latency Exp(β)\mathrm{Exp}(\beta)Exp(β), while the speculator proposes candidate actions with latency Exp(α)\mathrm{Exp}(\alpha)Exp(α), where β<α\beta < \alphaβ<α (the speculator is faster in expectation). A speculative branch hits when the action it guesses implies the same next call the actor's true action would have implied; branches hit independently across steps with probability ppp.

Two knobs define the two regimes analyzed. Breadth kkk: at each step, launch kkk independent one-step speculations in parallel, each immediately followed by a real call. At least one of the kkk succeeds with probability

p(k)  =  1−(1−p)k.p(k) \;=\; 1 - (1-p)^k .p(k)=1−(1−p)k.

Depth: follow a single branch, extending it whenever a speculative or real call returns and pruning subtrees the actor contradicts.

The quantity driving both results is SnS_nSn​, the expected number of hits by round nnn. A hit consumes the following step's speculation window — after a correct guess the next call is already cached, so no new speculation is launched there — which yields the two-term recursion

S0=0,S1=p,Sn=p (1+Sn−2)+(1−p) Sn−1.S_0 = 0, \qquad S_1 = p, \qquad S_n = p\,(1 + S_{n-2}) + (1-p)\,S_{n-1}.S0​=0,S1​=p,Sn​=p(1+Sn−2​)+(1−p)Sn−1​.

Write Tseq,MseqT_{\mathrm{seq}}, M_{\mathrm{seq}}Tseq​,Mseq​ for the latency and token cost of strictly sequential execution, and Tspec,MspecT_{\mathrm{spec}}, M_{\mathrm{spec}}Tspec​,Mspec​ for their speculative counterparts. In the depth regime latencies are taken deterministic: aaa for a real call, b<ab < ab<a for a speculative one.

Target

The goal theorem is the finite-horizon latency ratio for breadth-focused speculation (Proposition 1), with p(k)p(k)p(k) abbreviated pkp_kpk​:

E[Tspec]E[Tseq]=1−1T αα+β[(T−1)pk1+pk+pk2(1+pk)2−pk2(1+pk)2(−pk)T−1].\frac{\mathbb{E}[T_{\mathrm{spec}}]}{\mathbb{E}[T_{\mathrm{seq}}]} = 1 - \frac{1}{T}\,\frac{\alpha}{\alpha+\beta} \left[\frac{(T-1)p_k}{1+p_k} + \frac{p_k^2}{(1+p_k)^2} - \frac{p_k^2}{(1+p_k)^2}(-p_k)^{T-1}\right].E[Tseq​]E[Tspec​]​=1−T1​α+βα​[1+pk​(T−1)pk​​+(1+pk​)2pk2​​−(1+pk​)2pk2​​(−pk​)T−1].

The supporting targets, ordered as the analysis builds them:

  1. the closed form Sn=p1+pn+p2(1+p)2(1−(−p)n)S_n = \frac{p}{1+p}n + \frac{p^2}{(1+p)^2}\bigl(1 - (-p)^n\bigr)Sn​=1+pp​n+(1+p)2p2​(1−(−p)n) solving the recursion;
  2. the per-hit saving E[(B−A)+]=αβ(α+β)\mathbb{E}[(B-A)^+] = \frac{\alpha}{\beta(\alpha+\beta)}E[(B−A)+]=β(α+β)α​ for independent A∼Exp(α)A \sim \mathrm{Exp}(\alpha)A∼Exp(α), B∼Exp(β)B \sim \mathrm{Exp}(\beta)B∼Exp(β);
  3. the T→∞T \to \inftyT→∞ limit 1−pk1+pk⋅αα+β1 - \frac{p_k}{1+p_k}\cdot\frac{\alpha}{\alpha+\beta}1−1+pk​pk​​⋅α+βα​, and the resulting 50% ceiling: the latency reduction is strictly below 12\tfrac1221​ for every pk≤1p_k \le 1pk​≤1;
  4. the cost counterpart (Theorem 4), finite-horizon and in the limit, with k~\tilde kk~ the number of distinct actions across the kkk branches;
  5. the depth-focused time and cost identities (Theorem 6), whose latency coefficient is ppp rather than p1+p\frac{p}{1+p}1+pp​ — raising the speedup ceiling from 12\tfrac1221​ to 111;
  6. the structure of confidence-aware selective speculation (Theorem 3 and Corollary 5): with sorted per-branch confidences, the marginal hit-probability gain is non-increasing, so the optimal breadth is the greedy threshold rule "add a branch while Δ⋆δq(m)≥c\Delta^\star \delta q(m) \ge cΔ⋆δq(m)≥c".

Significance

The analysis is what turns speculation from a trick into a tunable system. Proposition 1 and Theorem 4 are governed by the same quantity pkp_kpk​, so a practitioner who can estimate hit probability can choose kkk offline against a latency/cost budget rather than by trial. The 50% ceiling is a genuine negative result — it says breadth alone cannot do better, and motivates the depth regime, where the ceiling becomes 1. Theorem 3 explains why confidence-based branch selection is cheap in practice: the whole dynamic program collapses to one scalar continuation value, so a runtime system sorts confidences and adds branches greedily in O(k)O(k)O(k) per step.

The paper's proofs are pen-and-paper and, as far as we are aware, none of these results has a machine-checked proof. Three parts reward formalization specifically. The recursion's closed form is derived by a characteristic-equation argument with a particular solution that collides with the homogeneous part — routine but error-prone. The per-hit saving is an honest two-dimensional integral over independent exponentials. And Theorem 6's cost expression is stated in the paper with a floor function and then immediately replaced by an approximation, so formalizing it forces a decision about which claim is actually being asserted (see Formalization scope).

Difficulty

The obvious first move on the recursion — guess a constant particular solution — fails, because r=1r = 1r=1 is a root of the characteristic polynomial r2−(1−p)r−pr^2 - (1-p)r - pr2−(1−p)r−p and a constant trial collides with the homogeneous family; the particular solution is linear in nnn, and the p2(1+p)2\frac{p^2}{(1+p)^2}(1+p)2p2​ coefficient comes out of matching both initial conditions, not one.

The interesting hypothesis is the one the recursion's shape encodes and the prose states only in passing: a hit at round ttt removes the speculation window at round t+1t+1t+1. Drop it and the recursion becomes one-term and the answer changes.

For the per-hit saving, the difficulty is analytic rather than algebraic: the inner antiderivative of (b−a)αe−αa(b-a)\alpha e^{-\alpha a}(b−a)αe−αa must be handled, and the outer integral runs over an unbounded interval, so integrability has to be established rather than assumed.

The asymptotic statements need the oscillating term (−pk)T−1(-p_k)^{T-1}(−pk​)T−1 controlled uniformly — it is bounded, not vanishing termwise in an obvious way — before the 1T\tfrac1TT1​ prefactor can be taken to zero.

Formalization scope

Everything is over R\mathbb{R}R. The model lives in one definition bundle, Def_SpecActions_model, in namespace SpecActions; the mission's Lean names match the prose symbols (SnS_nSn​ is hits, p(k)p(k)p(k) is phit, k~\tilde kk~ is kt).

The model is formalized at the level the paper's own proofs use: E[T]\mathbb{E}[T]E[T] and E[M]\mathbb{E}[M]E[M] are defined by the expressions Appendix A derives for them (specTime, specCost, and their depth analogues), and the theorems assert the algebraic and asymptotic identities relating those quantities. Deriving those expressions from a measure-theoretic model of the execution trace is deliberately not in scope — with one exception: milestone 2 states the per-hit saving as a genuine iterated integral against the exponential densities, so the one probabilistic step the paper actually computes is formalized as an integral rather than assumed.

Conventions a solver should know before starting:

  • Statements are quantified over α,β>0\alpha, \beta > 0α,β>0 and 0≤pk≤10 \le p_k \le 10≤pk​≤1; the standing assumption β<α\beta < \alphaβ<α is not imposed, since none of the identities need it.
  • Finite-horizon statements carry 1≤T1 \le T1≤T, and T−1T-1T−1 is natural-number subtraction — the T=0T = 0T=0 case is excluded rather than silently truncated.
  • hits takes pkp_kpk​ (the per-step hit probability p(k)p(k)p(k)), not the per-branch ppp; phit relates the two, and Thm_SpecActions_phit_bounds supplies the 0≤p(k)≤10 \le p(k) \le 10≤p(k)≤1 range facts the other statements assume.
  • Theorem 6's cost is stated as the exact identity, not the paper's approximation. The paper gives an exact expression involving ⌊a/b⌋\lfloor a/b \rfloor⌊a/b⌋ and then an ≈\approx≈ form with a2b−12\frac{a}{2b} - \frac122ba​−21​; these coincide only when a/ba/ba/b is an integer. The milestone asserts the exact floor version, which is what the proof establishes.
  • The 50% ceiling is stated as the strict bound pk1+pk⋅αα+β<12\frac{p_k}{1+p_k}\cdot\frac{\alpha}{\alpha+\beta} < \frac121+pk​pk​​⋅α+βα​<21​, which holds for all admissible parameters; the paper's "upper bound of 50%, occurring when p=1p=1p=1 and α=∞\alpha = \inftyα=∞" describes an unattained supremum.
  • Theorem 3's dynamic program is formalized as the two facts that carry its content — diminishing marginal returns, and optimality of the greedy threshold breadth — rather than as a Bellman recursion over a mode process, which would require a full MDP development.

Reusable beyond this mission: the two-term linear recursion solved in milestone 1, and the E[(B−A)+]\mathbb{E}[(B-A)^+]E[(B−A)+] computation for independent exponentials, which is a standard fact absent from Mathlib. Contributions extending the model toward an actual measure on execution traces — deriving specTime rather than defining it — are welcome as follow-on work.

Selected references

  • Naimeng Ye, Arnav Ahuja, Georgios Liargkovas, Yunan Lu, Kostis Kaffes, Tianyi Peng. Speculative Actions: A Lossless Framework for Faster Agentic Systems. ICLR 2026. arXiv:2510.04371 — Proposition 1 (p. 4), Appendix A (pp. 13–14), Theorem 3 (p. 10), Theorem 4 (p. 19), Corollary 5 (p. 22), Theorem 6 (p. 23).
  • Yaniv Leviathan, Matan Kalman, Yossi Matias. Fast Inference from Transformers via Speculative Decoding. ICML 2023. arXiv:2211.17192 — the speculate-verify pattern at token level.
  • Wenyue Hua, Mengting Wan, Shashank Vadrevu, Ryan Nadel, Yongfeng Zhang, Chi Wang. Interactive Speculative Planning. 2024. arXiv:2410.00079 — depth-oriented speculation on a single planning branch.
  • Yilin Guan et al. Dynamic Speculative Agent Planning. 2025. arXiv:2509.01920 — online RL for choosing speculation depth under a cost-latency trade-off.
  • Robert M. Tomasulo. An Efficient Algorithm for Exploiting Multiple Arithmetic Units. IBM Journal of Research and Development, 1967. DOI:10.1147/rd.111.0025 — speculative execution in hardware.
13 thms2 active usersReviewed
🏆Completed
Bandit AlgorithmsOperations Research·Captain: Shuze Chen

Bandit Algorithms XII: Follow-the-Regularised-Leader and Mirror DescentTextbook

Beneath Exp3, Exp4 and their relatives lies one algorithm: minimize past losses plus a convex regularizer. Chapters 26–28 of Lattimore–Szepesvári develop this unifying view. For a Legendre potential FFF with Bregman divergence DFD_FDF​, both mirror descent and follow-the-regularised-leader satisfy the master bound Rn(a)≤F(a)−F(a1)η+1η∑tDF(at,a~t+1)R_n(a) \le \frac{F(a) - F(a_1)}{\eta} + \frac{1}{\eta}\sum_t D_F(a_t, \tilde a_{t+1})Rn​(a)≤ηF(a)−F(a1​)​+η1​∑t​DF​(at​,a~t+1​); the negentropy potential on the simplex recovers Exp3 exactly. The goal theorem is the payoff for adversarial linear bandits: FTRL on the unit ball with the self-concordant-flavoured potential F(a)=−log⁡(1−∥a∥)−∥a∥F(a) = -\log(1-\|a\|) - \|a\|F(a)=−log(1−∥a∥)−∥a∥ achieves Rn≤23ndlog⁡nR_n \le 2\sqrt{3nd\log n}Rn​≤23ndlogn​ — improving the d\sqrt{d}d​ factor over the Exp3-style approach of Chapter 27 and matching the Ω(dn)\Omega(d\sqrt{n})Ω(dn​) lower bound of Mission XI up to logarithms.

5 thms2 active users
🏆Completed
Bandit AlgorithmsOperations Research·Captain: Shuze Chen

Bandit Algorithms VIII: Contextual Bandits and Exp4Textbook

Real decisions come with context: a news site chooses an article for a particular user. Competing with the single best arm is then meaningless; the right benchmark is the best mapping from contexts to arms, or more generally the best of MMM expert policies. Chapter 18 of Lattimore–Szepesvári formalizes this via Exp4 — exponential weighting over experts, fed by the importance-weighted estimator of Mission V. The goal theorem: with learning rate η=2log⁡(M)/(nk)\eta = \sqrt{2\log(M)/(nk)}η=2log(M)/(nk)​, Exp4 satisfies Rn≤2nklog⁡MR_n \le \sqrt{2nk\log M}Rn​≤2nklogM​ against the best of MMM experts. Since MMM enters only logarithmically, the learner can compete with exponentially large policy classes — the conceptual gateway from bandits to reinforcement learning with function approximation.

9 thms2 active usersReviewed
ProbabilityStatistics·Captain: mikedeng1

Rademacher and Gaussian Complexities: Risk Bounds and Structural Results 6: The Expected Maximum Discrepancy Lies Between R_n(F)/2 − 2√(2/n) and R_n(F) + 4√(2/n)Research Paper

Motivation

Data-dependent risk bounds in statistical learning theory control the gap between the expected loss of a learned function and its empirical loss by a complexity penalty that is computed from the training data. The first such penalties were the maximum discrepancy of a function class (Bartlett, Boucheron and Lugosi, Model selection and error estimation, Machine Learning 48, 2002) and its Rademacher complexity (Koltchinskii, Rademacher penalties and structural risk minimization, IEEE Trans. Inf. Theory 47, 2001; Koltchinskii and Panchenko 2000). The maximum discrepancy compares the behaviour of the class on two fixed halves of the sample; the Rademacher complexity compares it on two random halves. Bartlett and Mendelson (JMLR 3, 2002), Lemma 3, show that these two quantities are equivalent up to a factor 2 and an additive O(1/n)O(1/\sqrt n)O(1/n​). This mission formalizes that lemma from the published JMLR article (pp. 463–482); the proof is its Appendix A.

Setting

Let μ\muμ be a probability measure on a measurable space X\mathcal XX and let X1,…,XnX_1,\dots,X_nX1​,…,Xn​ be independent samples from μ\muμ. Let FFF be a class of measurable functions f:X→[−1,1]f:\mathcal X\to[-1,1]f:X→[−1,1]. Let σ1,…,σn\sigma_1,\dots,\sigma_nσ1​,…,σn​ be independent uniform {±1}\{\pm1\}{±1}-valued random variables, independent of the sample.

The Rademacher complexity of FFF is

Rn(F)=Esup⁡f∈F∣2n∑i=1nσif(Xi)∣.R_n(F) = \mathbf E\sup_{f\in F}\left|\frac2n\sum_{i=1}^n\sigma_i f(X_i)\right|.Rn​(F)=Ef∈Fsup​​n2​i=1∑n​σi​f(Xi​)​.

For even nnn, the maximum discrepancy of FFF is the random variable

D^n(F)=sup⁡f∈F(2n∑i=1n/2f(Xi)−2n∑i=n/2+1nf(Xi)),\hat D_n(F) = \sup_{f\in F}\left(\frac2n\sum_{i=1}^{n/2}f(X_i) - \frac2n\sum_{i=n/2+1}^n f(X_i)\right),D^n​(F)=f∈Fsup​​n2​i=1∑n/2​f(Xi​)−n2​i=n/2+1∑n​f(Xi​)​,

with no absolute value, and the expected maximum discrepancy is Dn(F)=ED^n(F)D_n(F)=\mathbf E\hat D_n(F)Dn​(F)=ED^n​(F). The class is closed under negation if f∈Ff\in Ff∈F implies −f∈F-f\in F−f∈F, and −F={−f:f∈F}-F=\{-f:f\in F\}−F={−f:f∈F}.

The proof works with the conditional supremum function

s(N)=2n E[sup⁡f∈F∑i=1nσif(Xi)  |  ∑i=1nσi=N],s(N) = \frac2n\,\mathbf E\left[\sup_{f\in F}\sum_{i=1}^n\sigma_i f(X_i)\;\middle|\;\sum_{i=1}^n\sigma_i=N\right],s(N)=n2​E[f∈Fsup​i=1∑n​σi​f(Xi​)​i=1∑n​σi​=N],

defined for the values NNN that ∑iσi\sum_i\sigma_i∑i​σi​ can take.

Formalization targets

Goal: Lemma 3, first and second displays

For every even n≥2n\ge2n≥2,

Rn(F)2−22n≤Dn(F)≤Rn(F)+42n,\frac{R_n(F)}{2} - 2\sqrt{\frac2n} \le D_n(F) \le R_n(F) + 4\sqrt{\frac2n},2Rn​(F)​−2n2​​≤Dn​(F)≤Rn​(F)+4n2​​,

and if FFF is closed under negation,

Rn(F)−42n≤Dn(F).R_n(F) - 4\sqrt{\frac2n} \le D_n(F).Rn​(F)−4n2​​≤Dn​(F).

Milestones (Appendix A, pp. 479–480)

  1. Rn(F)≥E s(∑iσi)R_n(F)\ge\mathbf E\,s(\sum_i\sigma_i)Rn​(F)≥Es(∑i​σi​), with equality when FFF is closed under negation.
  2. Dn(F)=s(0)D_n(F) = s(0)Dn​(F)=s(0).
  3. ∣s(N1)−s(N2)∣≤4∣N2−N1∣/n|s(N_1)-s(N_2)|\le 4|N_2-N_1|/n∣s(N1​)−s(N2​)∣≤4∣N2​−N1​∣/n.
  4. ∣Es(N)−s(EN)∣≤E∣s(N)−s(EN)∣≤42/n|\mathbf E s(N)-s(\mathbf EN)|\le\mathbf E|s(N)-s(\mathbf EN)|\le4\sqrt{2/n}∣Es(N)−s(EN)∣≤E∣s(N)−s(EN)∣≤42/n​ for N=∑iσiN=\sum_i\sigma_iN=∑i​σi​.
  5. Rn(F)=Rn(F∪−F)≤Dn(F∪−F)+42/nR_n(F)=R_n(F\cup-F)\le D_n(F\cup-F)+4\sqrt{2/n}Rn​(F)=Rn​(F∪−F)≤Dn​(F∪−F)+42/n​.
  6. Dn(F∪−F)≤2Dn(F)+Dn({f0,−f0})D_n(F\cup-F)\le 2D_n(F)+D_n(\{f_0,-f_0\})Dn​(F∪−F)≤2Dn​(F)+Dn​({f0​,−f0​}) for any f0∈Ff_0\in Ff0​∈F (a corrected form of the printed step, see below).

Significance

Lemma 3 makes the maximum discrepancy and the Rademacher complexity interchangeable in risk bounds: a bound in terms of one gives a bound in terms of the other with an explicit additive loss. The maximum discrepancy can be computed by a single empirical risk minimization on a relabelled sample, while the Rademacher complexity has the structural properties (monotonicity, convex-hull invariance, contraction) that make it easy to bound for concrete classes; the lemma transfers the second kind of estimate to the first quantity.

The lemma is proved in the paper; no machine-checked proof of it, or of the comparison between fixed and random half-sample splits, is known to exist. Formalizing it requires the exchangeability argument for i.i.d. samples, the conditioning of a uniform sign vector on its sum, and a moment bound for the Rademacher sum ∑iσi\sum_i\sigma_i∑i​σi​, all with explicit constants.

Difficulty

The heart of the proof is that, conditioned on the number of positive signs, a uniform sign vector splits the i.i.d. sample into two random subsets of fixed sizes, and every split of the same sizes has the same law as the fixed split. Making this precise requires a permutation-invariance argument for product measures applied to a supremum over an arbitrary class, where measurability is not automatic. The step from classes closed under negation to general classes is where the printed argument is loose: since D^n\hat D_nD^n​ has no absolute value, D^n(F∪−F)\hat D_n(F\cup-F)D^n​(F∪−F) is a maximum of two suprema that may be negative, and the naive bound Dn(F∪−F)≤2Dn(F)D_n(F\cup-F)\le2D_n(F)Dn​(F∪−F)≤2Dn​(F) fails.

Formalization scope

The Lean development lives in the namespace RadGauss.Discrepancy. Sign vectors are Fin n → Bool (true ↦ 1, false ↦ -1) and expectations over signs are finite averages over all 2n2^n2n sign vectors; s(N)s(N)s(N) is the average over the sign vectors with sum NNN. RnR_nRn​ takes values in [0,∞][0,\infty][0,∞] (a lower Lebesgue integral of an [0,∞][0,\infty][0,∞]-valued supremum), while D^n\hat D_nD^n​, DnD_nDn​ and sss are real, because the maximum discrepancy is signed. Inequalities of the form a−c≤Da-c\le Da−c≤D are written a≤D+ca\le D+ca≤D+c with real terms embedded by ENNReal.ofReal; this is equivalent to the printed form since Dn(F)≥0D_n(F)\ge0Dn​(F)≥0 for nonempty FFF.

Hypotheses added to the page, all disclosed in each item:

  • the sample size is even, n=2mn=2mn=2m with m≥1m\ge1m≥1, since D^n\hat D_nD^n​ needs half sums;
  • FFF is nonempty (the supremum over the empty class is −∞-\infty−∞ in the paper and 000 in Lean);
  • every f∈Ff\in Ff∈F is measurable, and for every sign vector σ\sigmaσ the map x↦sup⁡f∈F∑iσif(xi)x\mapsto\sup_{f\in F}\sum_i\sigma_if(x_i)x↦supf∈F​∑i​σi​f(xi​) is measurable. This is the measurability guard: without it the Bochner integrals defining DnD_nDn​ and sss would silently be 000.

Corrections of printed statements:

  • The Lipschitz bound on sss is printed for 0≤n2<n1≤n0\le n_2<n_1\le n0≤n2​<n1​≤n but used for negative values of ∑iσi\sum_i\sigma_i∑i​σi​; it is stated for every pair of attainable values.
  • The printed step Dn(F∪−F)≤2Dn(F)D_n(F\cup-F)\le2D_n(F)Dn​(F∪−F)≤2Dn​(F) is false for the signed D^n\hat D_nD^n​ of p. 464 (F={f}F=\{f\}F={f}, f(X)f(X)f(X) uniform on {±1}\{\pm1\}{±1}, n=2n=2n=2 gives 1≤01\le01≤0); the milestone states it with the additional term Dn({f0,−f0})≤2/nD_n(\{f_0,-f_0\})\le2/\sqrt nDn​({f0​,−f0​})≤2/n​. The goal itself remains true.
  • The third display of Lemma 3, P{∣D^n(F)−Dn(F)∣≥ϵ}≤2exp⁡(−ϵ2n/2)P\{|\hat D_n(F)-D_n(F)|\ge\epsilon\}\le2\exp(-\epsilon^2n/2)P{∣D^n​(F)−Dn​(F)∣≥ϵ}≤2exp(−ϵ2n/2), is false as printed (F={f}F=\{f\}F={f} as above, n=2n=2n=2, ϵ=2\epsilon=2ϵ=2: the probability is 1/2>2e−41/2>2e^{-4}1/2>2e−4) and is not part of the mission.

A formalization in which DnD_nDn​ or sss is a junk value (non-integrable or non-measurable suprema, an empty class, an odd sample size with truncated n/2n/2n/2) would make the goal trivial or meaningless; the hypotheses above rule that out, and the class F={0}F=\{0\}F={0} satisfies all of them.

Welcome contributions include general lemmas on the invariance of Esup⁡f∈FΦf(Xπ(1),…,Xπ(n))\mathbf E\sup_{f\in F}\Phi_f(X_{\pi(1)},\dots,X_{\pi(n)})Esupf∈F​Φf​(Xπ(1)​,…,Xπ(n)​) under permutations π\piπ of an i.i.d. sample, conditioning of uniform sign vectors on their sum, and the bound E∣∑iσi∣≤n\mathbf E|\sum_i\sigma_i|\le\sqrt nE∣∑i​σi​∣≤n​. These are reusable well beyond this mission.

Selected references

  • P. L. Bartlett, S. Mendelson, Rademacher and Gaussian Complexities: Risk Bounds and Structural Results, Journal of Machine Learning Research 3 (2002) 463–482. https://www.jmlr.org/papers/v3/bartlett02a.html
  • P. L. Bartlett, S. Boucheron, G. Lugosi, Model selection and error estimation, Machine Learning 48 (2002) 85–113. https://doi.org/10.1023/A:1013999503812
  • V. Koltchinskii, Rademacher penalties and structural risk minimization, IEEE Transactions on Information Theory 47 (2001) 1902–1914. https://doi.org/10.1109/18.930926
  • L. Devroye, L. Györfi, G. Lugosi, A Probabilistic Theory of Pattern Recognition, Springer, 1996. https://doi.org/10.1007/978-1-4612-0711-5
10 thms1 active userReviewed
CombinatoricsGroup Theory·Captain: mikedeng1

A Characterization of Multiclass Learnability 2: A Concept Class with Natarajan Dimension 1 and Infinite DS DimensionResearch Paper

Motivation

In binary classification the VC dimension decides PAC learnability: a class of {0,1}\{0,1\}{0,1}-valued functions is learnable from finitely many examples exactly when its VC dimension is finite. Multiclass classification, where a predictor outputs one of many labels, arises whenever the label set is large: language models choosing a next token, image recognition over open vocabularies, structured prediction. For finitely many labels the Natarajan dimension plays the role of the VC dimension (Natarajan 1989; Ben-David, Cesa-Bianchi, Haussler and Long 1995). Whether it still characterizes learnability when the label set is infinite stayed open for three decades.

Brukhim, Carmon, Dinur, Moran and Yehudayoff (arXiv:2203.01550, FOCS 2022) settled both directions. Their Theorem A shows that the DS dimension of Daniely and Shalev-Shwartz (COLT 2014, PMLR 35) characterizes multiclass PAC learnability for every label set. Their Theorem 2, the goal of this mission, shows that the Natarajan dimension does not: there is a class whose Natarajan dimension is 111 and whose DS dimension is infinite.

Timeline:

  • 1989: Natarajan introduces his dimension and proves it gives sample-complexity bounds when the label set is finite.
  • 1995: Ben-David, Cesa-Bianchi, Haussler and Long show that, for finite label sets, every "reasonable" extension of the VC dimension characterizes learnability.
  • 2003: Januszkiewicz and Świątkowski construct, for every dimension, finite simplicial complexes without empty squares from coset complexes of finite groups (Comment. Math. Helv. 78(3), 555–583); the multiclass paper uses this construction for its separation.
  • 2014: Daniely and Shalev-Shwartz introduce the DS dimension, prove that finite DS dimension is necessary for learnability, and ask whether it is sufficient.
  • 2022: Brukhim et al. prove that finite DS dimension is sufficient and that the Natarajan dimension fails to characterize learnability for infinite label sets.

Setting

A concept class is a set H⊆YX\mathcal H \subseteq \mathcal Y^{\mathcal X}H⊆YX of functions from a domain X\mathcal XX to a label set Y\mathcal YY, with no finiteness assumption on either. For a sequence S=(x1,…,xn)∈XnS = (x_1, \dots, x_n) \in \mathcal X^nS=(x1​,…,xn​)∈Xn, the projection H∣S⊆Yn\mathcal H|_S \subseteq \mathcal Y^nH∣S​⊆Yn is the set of words (h(x1),…,h(xn))(h(x_1), \dots, h(x_n))(h(x1​),…,h(xn​)), h∈Hh \in \mathcal Hh∈H.

  • SSS is N-shattered if there are f,g:[n]→Yf, g : [n] \to \mathcal Yf,g:[n]→Y with f(i)≠g(i)f(i) \ne g(i)f(i)=g(i) for every iii and H∣S⊇{f(1),g(1)}×⋯×{f(n),g(n)}\mathcal H|_S \supseteq \{f(1), g(1)\} \times \dots \times \{f(n), g(n)\}H∣S​⊇{f(1),g(1)}×⋯×{f(n),g(n)}: the projection contains a copy of the Boolean cube. The Natarajan dimension dN(H)d_N(\mathcal H)dN​(H) is the largest nnn for which some S∈XnS \in \mathcal X^nS∈Xn is N-shattered, or ∞\infty∞.
  • A pseudo-cube of dimension ddd is a non-empty, finite B⊆YdB \subseteq \mathcal Y^dB⊆Yd in which every word hhh has, for every coordinate iii, an iii-neighbour: a word g∈Bg \in Bg∈B with g(i)≠h(i)g(i) \ne h(i)g(i)=h(i) and g(j)=h(j)g(j) = h(j)g(j)=h(j) for j≠ij \ne ij=i. SSS is DS-shattered if H∣S\mathcal H|_SH∣S​ contains an nnn-dimensional pseudo-cube, and the DS dimension dDS(H)d_{DS}(\mathcal H)dDS​(H) is the largest such nnn, or ∞\infty∞.

Every Boolean cube is a pseudo-cube, so dN≤dDSd_N \le d_{DS}dN​≤dDS​. The hexagon {12,32,34,54,56,16}⊆{1,…,6}2\{12, 32, 34, 54, 56, 16\} \subseteq \{1,\dots,6\}^2{12,32,34,54,56,16}⊆{1,…,6}2 is a 2-dimensional pseudo-cube that contains no Boolean square.

The milestones pass through simplicial complexes: downward-closed families of finite sets. A complex is good if it is finite, pure, has a proper coloring rrr of its vertices with dim⁡(C)+1\dim(C)+1dim(C)+1 colors, and satisfies replacement (every vertex of every face can be exchanged for a new vertex). A good complex CCC with coloring rrr defines the class B(C,r)B(C, r)B(C,r) of its top faces, each written as the word listing its vertices by color. A square is a 4-cycle of distinct vertices in the 1-skeleton; it is empty if neither diagonal is an edge. The coset complex CF(H1,…,Hd)C_F(H_1, \dots, H_d)CF​(H1​,…,Hd​) of subgroups of a group FFF has the cosets gHigH_igHi​ as vertices and the sets of cosets with a common point as faces.

Formalization targets

Goal: Theorem 2 (p. 4)

∃ X,Y, H⊆YX:dN(H)=1anddDS(H)=∞.\exists\, \mathcal X, \mathcal Y,\ \mathcal H \subseteq \mathcal Y^{\mathcal X}:\qquad d_N(\mathcal H) = 1 \quad\text{and}\quad d_{DS}(\mathcal H) = \infty.∃X,Y, H⊆YX:dN​(H)=1anddDS​(H)=∞.

Milestones

  1. Theorem 45 (p. 30; Januszkiewicz–Świątkowski): for every d>1d > 1d>1 a finite group FFF and subgroups H1,…,HdH_1, \dots, H_dH1​,…,Hd​ with (⋂j≠iHj)∖Hi≠∅(\bigcap_{j\ne i} H_j) \setminus H_i \ne \emptyset(⋂j=i​Hj​)∖Hi​=∅ for all iii, whose coset complex has no empty squares.
  2. Proposition 46 (p. 31): such a coset complex has dimension d−1d - 1d−1, is good and has no empty squares.
  3. Proposition 42 (p. 28): a ddd-dimensional good complex with a proper coloring rrr yields the (d+1)(d+1)(d+1)-dimensional pseudo-cube B(C,r)B(C, r)B(C,r); conversely every pseudo-cube yields a good complex C(B)C(B)C(B).
  4. Proposition 43 (p. 29): dN(B(C,r))≥2d_N(B(C,r)) \ge 2dN​(B(C,r))≥2 iff CCC has a square v0v1v2v3v_0 v_1 v_2 v_3v0​v1​v2​v3​ with r(v0)=r(v2)r(v_0) = r(v_2)r(v0​)=r(v2​) and r(v1)=r(v3)r(v_1) = r(v_3)r(v1​)=r(v3​).
  5. Corollary 44 (p. 29): a good complex without empty squares gives dN(B(C,r))≤1d_N(B(C, r)) \le 1dN​(B(C,r))≤1 for every proper coloring.
  6. Proof of Theorem 2 (p. 32): for every d≥1d \ge 1d≥1, a ddd-dimensional pseudo-cube with Natarajan dimension exactly 111.

Significance

Theorem 2 shows that the classical generalization of the VC dimension to many labels is the wrong invariant once the label set is infinite: a class can contain no Boolean square at all and still be unlearnable, because it contains pseudo-cubes of every dimension. Combined with the necessity of finite DS dimension, it gives a class that is not PAC learnable although its Natarajan dimension is 111, and it identifies pseudo-cubes, not Boolean cubes, as the relevant combinatorial obstruction. It also links learning theory to a problem studied in geometric group theory, finite "flag-no-square" complexes.

The paper's proof is complete modulo Theorem 45, which it imports from Januszkiewicz–Świątkowski 2003. None of these results is formalized. A formalization would give machine-checked versions of the dictionary between concept classes and properly colored complexes (Propositions 42–44), of the coset-complex translation (Proposition 46), and of the final disjoint-union argument; Theorem 45 itself, which rests on Coxeter-group and topological arguments, is a separate and substantial formalization target.

Difficulty

Infinite complexes that are pure, properly colored, satisfy replacement and have no empty squares are easy to build: grow a tree of faces indefinitely. The definition of a pseudo-cube demands finiteness, and the difficulty is entirely there: one must "fold" such an infinite object into a finite one without creating an empty square. The obvious finite candidate, the group (Z/2)d(\mathbb Z/2)^d(Z/2)d with its coordinate subgroups, produces the Boolean cube, whose complex is full of empty squares. Theorem 45 is the input that resolves this, and it is far beyond the rest of the argument.

Formalization scope

All declarations live in the namespace MulticlassDS.NatGap.

  • Concept classes are Set (X → Y) with arbitrary types; [n][n][n] is Fin n (0-based), and shattering is defined for sequences Fin n → X, as in the paper.
  • Both dimensions are ℕ∞-valued suprema, so "infinite DS dimension" is dsDim H = ⊤. An ℕ-valued supremum would silently return 000 on an unbounded family and would trivialize the goal.
  • The goal requires the Natarajan dimension to be exactly 111; an upper bound alone holds for any class with at most one element.
  • Pseudo-cubes are required to be finite (Definition 5). Without finiteness, the tree classes of Example 8 would already have infinite "DS dimension".
  • Complexes are Set (Finset V). The dimension is the predicate HasDim C d, not a natural-number subtraction, and colors are Fin (d + 1).
  • Replacement is stated with a new vertex u∉fu \notin fu∈/f. The page writes "u≠vu \ne vu=v", but read literally that allows u∈fu \in fu∈f, which makes the condition hold by downward closure and makes Proposition 42 false; the proofs of Propositions 42 and 46 use a new vertex.
  • Coset-complex vertices are left cosets as subsets of the group, not pairs (index, coset).
  • Proposition 46 states dimension d−1d - 1d−1 under d>1d > 1d>1, where the subtraction is exact; the converse of Proposition 42 is indexed by d+1d + 1d+1 and ddd to avoid it.
  • The proof-of-Theorem-2 milestone says "for every ddd"; it is posed for d≥1d \ge 1d≥1, because at d=0d = 0d=0 the only pseudo-cube has Natarajan dimension 000.

Welcome contributions: proofs of Propositions 42–44 and 46 and of the goal from the milestones, which need only finite combinatorics and elementary group theory; and, separately, a formalization of the Januszkiewicz–Świątkowski construction behind Theorem 45. The definitions of pseudo-cubes, the DS dimension and good complexes are reusable by the companion mission on sample compression and by any later work on multiclass learnability.

Selected references

  • N. Brukhim, D. Carmon, I. Dinur, S. Moran, A. Yehudayoff, A Characterization of Multiclass Learnability, arXiv:2203.01550v1, 2022 (FOCS 2022). https://arxiv.org/abs/2203.01550
  • T. Januszkiewicz, J. Świątkowski, Hyperbolic Coxeter groups of large dimension, Comment. Math. Helv. 78(3) (2003), 555–583 (reference [Januszkiewicz and Świątkowski 2003] of arXiv:2203.01550v1, p. 33).
  • A. Daniely, S. Shalev-Shwartz, Optimal learners for multiclass problems, COLT 2014, PMLR 35, 287–316. https://proceedings.mlr.press/v35/
  • B. K. Natarajan, On learning sets and functions, Machine Learning 4 (1989), 67–97. https://doi.org/10.1007/BF00114804
  • S. Ben-David, N. Cesa-Bianchi, D. Haussler, P. M. Long, Characterizations of learnability for classes of {0,…,n}-valued functions, J. Comput. Syst. Sci. 50(1) (1995), 74–86. https://doi.org/10.1006/jcss.1995.1008
10 thms1 active userReviewed
CombinatoricsTheoretical Computer Science·Captain: mikedeng1

A Characterization of Multiclass Learnability 1: Classes of Finite DS Dimension Have n → r Sample Compression Schemes with r Polylogarithmic in nResearch Paper

Motivation

In multiclass classification a learner sees examples (x,y)(x, y)(x,y) with xxx in a domain X\mathcal XX and a label yyy in a set Y\mathcal YY, and must predict labels of new points. When Y\mathcal YY is finite, the Natarajan dimension characterizes PAC learnability, extending the role of the VC dimension in binary classification (Natarajan 1989; Ben-David, Cesa-Bianchi, Haussler, Long 1995). Label sets in practice are often unbounded: structured prediction, ranking, and language modelling all predict from very large or infinite label spaces. For infinite Y\mathcal YY the Natarajan dimension fails to characterize learnability, and the question of which combinatorial parameter does was left open by Daniely and Shalev-Shwartz.

Timeline:

  • 1989–1995. Natarajan, then Ben-David et al. and Haussler–Long: for finite Y\mathcal YY, learnability is equivalent to finite Natarajan dimension, with sample complexity depending on log⁡∣Y∣\log|\mathcal Y|log∣Y∣.
  • 2011–2015. Daniely, Sabato, Ben-David and Shalev-Shwartz show that ERM can fail for multiclass problems with many labels. Daniely and Shalev-Shwartz (COLT 2014) introduce the DS dimension, prove that finite DS dimension is necessary for learnability, and ask whether it is sufficient.
  • 2022. Brukhim, Carmon, Dinur, Moran, Yehudayoff prove sufficiency, so the DS dimension characterizes multiclass PAC learnability, and show that the Natarajan dimension does not.

Setting

A concept class is a set H⊆YX\mathcal H\subseteq\mathcal Y^{\mathcal X}H⊆YX of functions. For a sequence S=(x1,…,xn)S=(x_1,\dots,x_n)S=(x1​,…,xn​) the projection H∣S⊆Yn\mathcal H|_S\subseteq\mathcal Y^nH∣S​⊆Yn is the set of label words (h(x1),…,h(xn))(h(x_1),\dots,h(x_n))(h(x1​),…,h(xn​)), h∈Hh\in\mathcal Hh∈H. A finite non-empty set B⊆YdB\subseteq\mathcal Y^dB⊆Yd is a pseudo-cube if every h∈Bh\in Bh∈B has, in every coordinate iii, a neighbour g∈Bg\in Bg∈B that differs from hhh exactly in coordinate iii. The sequence SSS is DS-shattered if H∣S\mathcal H|_SH∣S​ contains an nnn-dimensional pseudo-cube, and the DS dimension dDS(H)d_{DS}(\mathcal H)dDS​(H) is the maximum length of a DS-shattered sequence. The Natarajan dimension dN(H)≤dDS(H)d_N(\mathcal H)\le d_{DS}(\mathcal H)dN​(H)≤dDS​(H) is the same with Boolean cubes ∏i{f(i),g(i)}\prod_i\{f(i),g(i)\}∏i​{f(i),g(i)}, f(i)≠g(i)f(i)\ne g(i)f(i)=g(i), in place of pseudo-cubes.

A sample S∈(X×Y)nS\in(\mathcal X\times\mathcal Y)^nS∈(X×Y)n is H\mathcal HH-realizable if some h∈Hh\in\mathcal Hh∈H is consistent with it. An n→rn\to rn→r sample compression scheme for H\mathcal HH (Littlestone and Warmuth 1986) is a single reconstruction function ρ:(X×Y)r→YX\rho:(\mathcal X\times\mathcal Y)^r\to\mathcal Y^{\mathcal X}ρ:(X×Y)r→YX such that every realizable sample of size nnn contains rrr of its examples S′S'S′ with ρ(S′)\rho(S')ρ(S′) consistent with the whole sample. Logarithms are base 222 throughout.

Formalization targets

Goal: Theorem 36 (p. 22)

For H\mathcal HH with dDS(H)=dDS<∞d_{DS}(\mathcal H)=d_{DS}<\inftydDS​(H)=dDS​<∞ and dN(H)=dNd_N(\mathcal H)=d_NdN​(H)=dN​, and all integers n,t>0n,t>0n,t>0, there is an n→rn\to rn→r sample compression scheme, r≤nr\le nr≤n, with

r≤(dDS+t+1t+1(dDS+t)+103dNlog⁡((dDS+t+1t+1)log⁡(2n)))log⁡(2n).r\le\left(\frac{d_{DS}+t+1}{t+1}(d_{DS}+t)+10^3d_N\log\left(\binom{d_{DS}+t+1}{t+1}\log(2n)\right)\right)\log(2n).r≤(t+1dDS​+t+1​(dDS​+t)+103dN​log((t+1dDS​+t+1​)log(2n)))log(2n).

Milestones

The scheme combines two components, each with its own chain of results:

  • List learning from the DS dimension. Lemma 13 (orientations of out-degree ≤d\le d≤d on Yd+1\mathcal Y^{d+1}Yd+1), Claim 16 (the one-inclusion algorithm is right on some leave-one-out example), Fact 14 (leave-one-out symmetrization), Proposition 32 (a list PAC learner with list size (d+tt)\binom{d+t}{t}(td+t​) and success probability t+1d+t+1\frac{t+1}{d+t+1}d+t+1t+1​), Lemma 39 (an n→r1n\to r_1n→r1​ list compression scheme with r1≤dDS+t+1t+1(dDS+t)log⁡(2n)r_1\le\frac{d_{DS}+t+1}{t+1}(d_{DS}+t)\log(2n)r1​≤t+1dDS​+t+1​(dDS​+t)log(2n) and menu size ≤(dDS+t+1t+1)log⁡(2n)\le\binom{d_{DS}+t+1}{t+1}\log(2n)≤(t+1dDS​+t+1​)log(2n)).
  • Learning from a menu via shifting. Claim 22, Corollary 23, Claim 26, Proposition 27 (avd⁡≤4dE\operatorname{avd}\le4d_Eavd≤4dE​), Corollary 28, Lemma 29 (dE≤5dNlog⁡pd_E\le5d_N\log pdE​≤5dN​logp), Lemma 17 (orientations of out-degree ≤20dNlog⁡p\le20d_N\log p≤20dN​logp on [p]n[p]^n[p]n), Proposition 34 (error ≤20dNlog⁡(p)/n\le20d_N\log(p)/n≤20dN​log(p)/n given a ppp-menu), Lemma 40 (an n→r2n\to r_2n→r2​ compression scheme given a ppp-menu with r2≤103dNlog⁡(p)log⁡(2n)r_2\le10^3d_N\log(p)\log(2n)r2​≤103dN​log(p)log(2n)).

Significance

Theorem 36 is the algorithmic heart of the characterization: by the standard "compression implies generalization" argument it gives PAC learnability of every class of finite DS dimension, with sample complexity O~(dDS3/2/ϵ)\tilde O(d_{DS}^{3/2}/\epsilon)O~(dDS3/2​/ϵ) in the realizable case (t=⌈dDS1/2⌉t=\lceil d_{DS}^{1/2}\rceilt=⌈dDS1/2​⌉), and with the agnostic case following by known reductions. It also exhibits sample compression schemes of size polylogarithmic in nnn for multiclass classes with infinitely many labels, in contrast to the constant-size schemes known for finite VC classes.

The result is proved in the paper; none of it is formalized. The formalization would produce a machine-checked theory of one-inclusion graphs and their orientations, multiclass shifting, the exponential dimension, list learning, and sample compression schemes for arbitrary label sets. These objects recur throughout learning theory (one-inclusion graphs in optimal PAC learning, shifting in VC theory), so the infrastructure is reusable beyond this mission.

Difficulty

The natural first idea, running empirical risk minimization or bounding the Natarajan dimension, fails: classes with Natarajan dimension 111 and infinitely many labels can be unlearnable, and ERM can fail even for learnable classes. The DS dimension gives only a weak guarantee: by Claim 16, among d+1d+1d+1 leave-one-out runs, one is correct. Turning this into a learner requires a list learner whose menus are still of unbounded total size, and then learning with a menu of size ppp, where the obstacle is controlling one-inclusion graph orientations over [p]n[p]^n[p]n by the Natarajan dimension. Multiclass shifting does not preserve the average degree (Example 20), so the binary argument breaks down, and a new potential (avd⁡′\operatorname{avd}'avd′) and a new dimension (dEd_EdE​) are needed. Lemma 13 for infinite classes needs a compactness argument.

Formalization scope

Lean conventions:

  • A class is H : Set (X → Y) with arbitrary types X, Y; sequences and samples are functions on Fin n ([n][n][n] is 0-based).
  • The DS, Natarajan and exponential dimensions are suprema in ℕ∞, so unbounded families give ⊤; hypotheses are written dsDim H = dDS with dDS : ℕ. A pseudo-cube is required to be finite.
  • Logarithms are Real.logb 2. Menu sizes use Set.encard.
  • A compression scheme is a reconstruction function fixed before the sample (∃ ρ, ∀ S, ∃ S'); a subsample may repeat and reorder examples. Theorem 36 states r≤nr\le nr≤n explicitly.
  • Orientations of the one-inclusion graph of V⊆YmV\subseteq\mathcal Y^mV⊆Ym are maps sending a direction iii and a vertex vvv to the head of the edge of direction iii through vvv; the out-degree of vvv counts directions whose head is not vvv.
  • Classes over [p][p][p] use labels Fin p; the shifting condition 1≤g(i)≤∣ef∣1\le g(i)\le|e_f|1≤g(i)≤∣ef​∣ becomes g(i)<∣ef∣g(i)<|e_f|g(i)<∣ef​∣.
  • The one-inclusion algorithm (Algorithms 1 and 3) is parametrized by a permutation-equivariant choice of minimal orientations, the reading under which the paper's leave-one-out proofs are valid; statements about the algorithm hold for every such choice. Its default output on non-realizable input requires a non-empty label set, assumed in Claim 16 and Propositions 32 and 34. Lemma 40 assumes a non-empty label set because it is false for X≠∅=Y\mathcal X\ne\emptyset=\mathcal YX=∅=Y.
  • Distributions are discrete (PMF), and i.i.d. probabilities are sums over Zm\mathcal Z^mZm. The measure-theoretic generality of the paper is not attempted.

A trivial formalization is ruled out by these choices. Placing the reconstruction function after the sample would let it output the consistent hypothesis. A dimension in ℕ defined by sSup would be 000 for infinite dimension. Pseudo-cubes without finiteness would change the dimension (Example 8).

Contributions welcome: proofs of any milestone, in particular the shifting results of §3 (self-contained combinatorics on finite classes), Fact 14 (pure discrete probability), and Lemma 13; general-purpose lemmas about one-inclusion graphs, orientations and sample compression schemes are reusable by other learning-theory missions.

Selected references

  • N. Brukhim, D. Carmon, I. Dinur, S. Moran, A. Yehudayoff, A Characterization of Multiclass Learnability, FOCS 2022; arXiv:2203.01550v1 (2022). https://arxiv.org/abs/2203.01550
  • A. Daniely, S. Shalev-Shwartz, Optimal Learners for Multiclass Problems, COLT 2014. https://arxiv.org/abs/1405.2690
  • N. Littlestone, M. Warmuth, Relating Data Compression and Learnability, unpublished technical report, University of California, Santa Cruz, 1986 (no stable link).
  • D. Haussler, N. Littlestone, M. Warmuth, Predicting {0,1}-Functions on Randomly Drawn Points, Information and Computation 115(2), 1994. https://doi.org/10.1006/inco.1994.1097
  • D. Haussler, P. M. Long, A Generalization of Sauer's Lemma, Journal of Combinatorial Theory, Series A 71(2), 1995. https://doi.org/10.1016/0097-3165(95)90006-3
  • S. Ben-David, N. Cesa-Bianchi, D. Haussler, P. M. Long, Characterizations of Learnability for Classes of {0,…,n}-Valued Functions, JCSS 50(1), 1995. https://doi.org/10.1006/jcss.1995.1008
  • B. K. Natarajan, On Learning Sets and Functions, Machine Learning 4, 1989. https://doi.org/10.1007/BF00114804
20 thms1 active userReviewed
OptimizationProbabilityStatistics·Captain: mikedeng1

Non-Strongly-Convex Smooth Stochastic Approximation with Convergence Rate O(1/n): Averaged Constant-Step-Size LMS Has Expected Excess Risk at Most (1/2n)[σ√d/(1−√(γR²)) + R‖θ₀−θ*‖/√(γR²)]²Research Paper

Motivation

Least-squares regression fitted by stochastic gradient descent — the least-mean-square (LMS) algorithm — is the basic large-scale learning procedure: each observation is touched once, at a cost linear in the dimension. Classical analyses of stochastic approximation give the rate O(1/n)O(1/\sqrt n)O(1/n​) for non-strongly-convex objectives, and O(1/(μn))O(1/(\mu n))O(1/(μn)) when the objective is μ\muμ-strongly convex. For least squares, μ\muμ is the smallest eigenvalue of the input covariance, which in high-dimensional problems is close to zero, so the strongly convex rate is often worse than the non-strongly-convex one.

F. Bach and E. Moulines (arXiv:1306.2119, NeurIPS 2013) showed that for the square loss this dichotomy disappears: averaged LMS with a constant step size reaches the rate O(1/n)O(1/n)O(1/n) with no strong-convexity assumption, and with a constant that does not involve the smallest eigenvalue. Averaging of stochastic approximation iterates goes back to Polyak and Juditsky (SIAM J. Control Optim. 1992), whose guarantees are asymptotic and use decreasing step sizes. The proof technique for the expansion of the noise process is adapted from Aguech, Moulines and Priouret (SIAM J. Control Optim. 2000). This mission formalizes the non-asymptotic bound in expectation (Theorem 1 of the paper) and the chain of lemmas of its Appendix A.

Setting

Let H=Rd\mathcal H=\mathbb R^dH=Rd with d≥1d\ge1d≥1, inner product ⟨⋅,⋅⟩\langle\cdot,\cdot\rangle⟨⋅,⋅⟩ and norm ∥⋅∥\|\cdot\|∥⋅∥. For a∈Ha\in\mathcal Ha∈H, a⊗aa\otimes aa⊗a is the operator b↦⟨a,b⟩ab\mapsto\langle a,b\rangle ab↦⟨a,b⟩a. For self-adjoint operators, A≼BA\preccurlyeq BA≼B means that B−AB-AB−A is positive semi-definite.

The data are independent and identically distributed pairs (xn,zn)∈H×H(x_n,z_n)\in\mathcal H\times\mathcal H(xn​,zn​)∈H×H, n≥1n\ge1n≥1, with finite second moments. The covariance operator is H=E[xn⊗xn]H=\mathbb E[x_n\otimes x_n]H=E[xn​⊗xn​], assumed invertible (its eigenvalues may be arbitrarily small). The least-squares objective

f(θ)=12 E[⟨θ,xn⟩2−2⟨θ,zn⟩]f(\theta)=\tfrac12\,\mathbb E\big[\langle\theta,x_n\rangle^2-2\langle\theta,z_n\rangle\big]f(θ)=21​E[⟨θ,xn​⟩2−2⟨θ,zn​⟩]

attains its global minimum at θ∗\theta^*θ∗, and ξn=zn−⟨θ∗,xn⟩xn\xi_n=z_n-\langle\theta^*,x_n\rangle x_nξn​=zn​−⟨θ∗,xn​⟩xn​ is the residual. The model need not be well specified: E[ξn∣xn]\mathbb E[\xi_n\mid x_n]E[ξn​∣xn​] need not vanish. Two constants R,σ>0R,\sigma>0R,σ>0 satisfy

E[ξn⊗ξn]≼σ2H,E[∥xn∥2xn⊗xn]≼R2H.\mathbb E[\xi_n\otimes\xi_n]\preccurlyeq\sigma^2H,\qquad \mathbb E\big[\|x_n\|^2x_n\otimes x_n\big]\preccurlyeq R^2H .E[ξn​⊗ξn​]≼σ2H,E[∥xn​∥2xn​⊗xn​]≼R2H.

These are assumptions (A1)–(A6) of §2.1. The LMS recursion with constant step size γ\gammaγ, started at θ0∈H\theta_0\in\mathcal Hθ0​∈H, is

θn=θn−1−γ(⟨θn−1,xn⟩xn−zn)=(I−γxn⊗xn)θn−1+γzn,\theta_n=\theta_{n-1}-\gamma\big(\langle\theta_{n-1},x_n\rangle x_n-z_n\big)=(I-\gamma x_n\otimes x_n)\theta_{n-1}+\gamma z_n ,θn​=θn−1​−γ(⟨θn−1​,xn​⟩xn​−zn​)=(I−γxn​⊗xn​)θn−1​+γzn​,

and its average is θˉn−1=n−1∑k=0n−1θk\bar\theta_{n-1}=n^{-1}\sum_{k=0}^{n-1}\theta_kθˉn−1​=n−1∑k=0n−1​θk​.

Formalization targets

Goal: Theorem 1, Eq. (2)

For every step size 0<γ<1/R20<\gamma<1/R^20<γ<1/R2 and every n≥1n\ge1n≥1,

E[f(θˉn−1)−f(θ∗)]≤12n[σd1−γR2+R∥θ0−θ∗∥1γR2]2.\mathbb E\big[f(\bar\theta_{n-1})-f(\theta^*)\big]\le\frac{1}{2n}\left[\frac{\sigma\sqrt d}{1-\sqrt{\gamma R^2}}+R\|\theta_0-\theta^*\|\frac{1}{\sqrt{\gamma R^2}}\right]^2 .E[f(θˉn−1​)−f(θ∗)]≤2n1​[1−γR2​σd​​+R∥θ0​−θ∗∥γR2​1​]2.

The constants are the paper's. A companion item states the case γ=1/(4R2)\gamma=1/(4R^2)γ=1/(4R2), where the bound reads 2n[σd+R∥θ0−θ∗∥]2\frac2n\big[\sigma\sqrt d+R\|\theta_0-\theta^*\|\big]^2n2​[σd​+R∥θ0​−θ∗∥]2.

Milestones (Appendix A)

  1. The excess risk is a quadratic form: f(θ)−f(θ∗)=12⟨θ−θ∗,H(θ−θ∗)⟩f(\theta)-f(\theta^*)=\tfrac12\langle\theta-\theta^*,H(\theta-\theta^*)\ranglef(θ)−f(θ∗)=21​⟨θ−θ∗,H(θ−θ∗)⟩.
  2. Consequences of (A6): E∥xn∥2≤R2\mathbb E\|x_n\|^2\le R^2E∥xn​∥2≤R2, tr⁡H≤R2\operatorname{tr}H\le R^2trH≤R2, H≼R2IH\preccurlyeq R^2IH≼R2I, and γH≼I\gamma H\preccurlyeq IγH≼I for γ≤1/R2\gamma\le1/R^2γ≤1/R2.
  3. Lemma 1: for a recursion αn=(I−γxn⊗xn)αn−1+γξn\alpha_n=(I-\gamma x_n\otimes x_n)\alpha_{n-1}+\gamma\xi_nαn​=(I−γxn​⊗xn​)αn−1​+γξn​ with martingale-difference noise and γR2≤1\gamma R^2\le1γR2≤1,
(1−γR2) E⟨αˉn−1,Hαˉn−1⟩+12nγE∥αn∥2≤12nγ∥α0∥2+γn∑k=1nE∥ξk∥2.(1-\gamma R^2)\,\mathbb E\langle\bar\alpha_{n-1},H\bar\alpha_{n-1}\rangle+\tfrac{1}{2n\gamma}\mathbb E\|\alpha_n\|^2\le\tfrac{1}{2n\gamma}\|\alpha_0\|^2+\tfrac{\gamma}{n}\textstyle\sum_{k=1}^{n}\mathbb E\|\xi_k\|^2 .(1−γR2)E⟨αˉn−1​,Hαˉn−1​⟩+2nγ1​E∥αn​∥2≤2nγ1​∥α0​∥2+nγ​∑k=1n​E∥ξk​∥2.
  1. Lemma 3: (1−(1−u)n)2≤nu(1-(1-u)^n)^2\le nu(1−(1−u)n)2≤nu for u∈[0,1]u\in[0,1]u∈[0,1] and n>0n>0n>0.
  2. Lemma 2: for αn=(I−γH)αn−1+γξn\alpha_n=(I-\gamma H)\alpha_{n-1}+\gamma\xi_nαn​=(I−γH)αn−1​+γξn​ with E[ξn⊗ξn]≼C\mathbb E[\xi_n\otimes\xi_n]\preccurlyeq CE[ξn​⊗ξn​]≼C, the second-moment bound (13) and
E⟨αˉn−1,Hαˉn−1⟩≤1nγ∥α0∥2+1ntr⁡(CH−1).\mathbb E\langle\bar\alpha_{n-1},H\bar\alpha_{n-1}\rangle\le\tfrac{1}{n\gamma}\|\alpha_0\|^2+\tfrac1n\operatorname{tr}(CH^{-1}).E⟨αˉn−1​,Hαˉn−1​⟩≤nγ1​∥α0​∥2+n1​tr(CH−1).
  1. The pathwise decomposition θn−θ∗=M1n(θ0−θ∗)+γ∑k=1nMk+1nξk\theta_n-\theta^*=M^n_1(\theta_0-\theta^*)+\gamma\sum_{k=1}^nM^n_{k+1}\xi_kθn​−θ∗=M1n​(θ0​−θ∗)+γ∑k=1n​Mk+1n​ξk​ (A.2).
  2. The initial-condition bound E⟨ηˉn−1,Hηˉn−1⟩≤∥η0∥2/(nγ)\mathbb E\langle\bar\eta_{n-1},H\bar\eta_{n-1}\rangle\le\|\eta_0\|^2/(n\gamma)E⟨ηˉ​n−1​,Hηˉ​n−1​⟩≤∥η0​∥2/(nγ) for the noise-free process (A.3).
  3. The expansion of the noise process (A.4): the remainder recursion (16), the covariance bound (17) E[ηn−1r⊗ηn−1r]≼γr+1R2rσ2I\mathbb E[\eta^r_{n-1}\otimes\eta^r_{n-1}]\preccurlyeq\gamma^{r+1}R^{2r}\sigma^2IE[ηn−1r​⊗ηn−1r​]≼γr+1R2rσ2I, the order-rrr bound 1nγrR2rdσ2\frac1n\gamma^rR^{2r}d\sigma^2n1​γrR2rdσ2, the remainder bound γr+2σ2R2r+41−γR2\frac{\gamma^{r+2}\sigma^2R^{2r+4}}{1-\gamma R^2}1−γR2γr+2σ2R2r+4​, and the noise bound
(E⟨ηˉn−1,Hηˉn−1⟩)1/2≤σdn⋅11−γR2(η0=0, γR2<1).\big(\mathbb E\langle\bar\eta_{n-1},H\bar\eta_{n-1}\rangle\big)^{1/2}\le\frac{\sigma\sqrt d}{\sqrt n}\cdot\frac{1}{1-\sqrt{\gamma R^2}}\quad(\eta_0=0,\ \gamma R^2<1).(E⟨ηˉ​n−1​,Hηˉ​n−1​⟩)1/2≤n​σd​​⋅1−γR2​1​(η0​=0, γR2<1).

Significance

The result. Theorem 1 gives a finite-sample, dimension-explicit bound with two terms: a variance term σ2d/n\sigma^2d/nσ2d/n, which matches the minimax rate for least-squares regression, and a bias term R2∥θ0−θ∗∥2/(γn)R^2\|\theta_0-\theta^*\|^2/(\gamma n)R2∥θ0​−θ∗∥2/(γn). Neither involves the smallest eigenvalue of HHH, so the guarantee survives ill-conditioning, which is the regime of high-dimensional learning. The bound is the basis for the paper's later results: the high-probability bound (Theorem 2) and the constant-step algorithm for logistic regression (Theorem 3), whose analysis invokes Theorem 1 for the quadratic approximations.

Formalizing it. The result is proved in the paper; nothing here is open. To our knowledge no part of it has been machine-checked. A formalization yields a reusable layer for linear stochastic approximation in finite dimension: martingale-difference noise in Rd\mathbb R^dRd, second-moment bounds for linear recursions driven by random operators, Loewner-order arguments, and averaging. The lemmas are stated for an abstract filtration and an abstract operator HHH, so they apply beyond this model. The mission also records the corrections the appendix needs (an "===" that should be "≼\preccurlyeq≼" in (13), an index in (16), and the exponent of ∥η0∥\|\eta_0\|∥η0​∥ in A.5).

Difficulty

Two steps resist the naive approach. First, the obvious one-step analysis — expand ∥θn−θ∗∥2\|\theta_n-\theta^*\|^2∥θn​−θ∗∥2 and take expectations — yields the bias part and Lemma 1, but on the noise it gives only γ∑kE∥ξk∥2/n\gamma\sum_k\mathbb E\|\xi_k\|^2/nγ∑k​E∥ξk​∥2/n, which does not decrease with nnn. The σ2d/n\sigma^2d/nσ2d/n rate requires averaging to cancel the noise, and this cancellation is visible only for the recursion with xn⊗xnx_n\otimes x_nxn​⊗xn​ replaced by its mean HHH. The random recursion is therefore expanded in powers of γ\gammaγ around the mean recursion, and each term ηr\eta^rηr needs its own covariance bound, by induction on rrr, using the independence of xnx_nxn​ from ηn−1r\eta^{r}_{n-1}ηn−1r​. Second, the induction relies on Loewner-order bookkeeping: sums of (I−γH)2kH(I-\gamma H)^{2k}H(I−γH)2kH must be bounded uniformly in nnn without dividing by small eigenvalues.

Formalization scope

The space H\mathcal HH is EuclideanSpace ℝ (Fin d). Operators are continuous linear maps, and H−1H^{-1}H−1 is an explicit two-sided inverse. Observations are indexed from 111. Averages are pˉn−1=n−1∑k=0n−1pk\bar p_{n-1}=n^{-1}\sum_{k=0}^{n-1}p_kpˉ​n−1​=n−1∑k=0n−1​pk​, with n≥1n\ge1n≥1 in every statement that uses them. The covariance operator is defined by its bilinear form, ⟨v,Hw⟩=E[⟨x1,v⟩⟨x1,w⟩]\langle v,Hw\rangle=\mathbb E[\langle x_1,v\rangle\langle x_1,w\rangle]⟨v,Hw⟩=E[⟨x1​,v⟩⟨x1​,w⟩]. Every Loewner inequality whose sides are expectations is an inequality of quadratic forms (for example E⟨ξ1,v⟩2≤σ2⟨v,Hv⟩\mathbb E\langle\xi_1,v\rangle^2\le\sigma^2\langle v,Hv\rangleE⟨ξ1​,v⟩2≤σ2⟨v,Hv⟩ for all vvv), which is the same order for self-adjoint operators.

Lean's Bochner integral is 000 on non-integrable functions, so every moment assumption carries the integrability of its integrand, and every bounded expectation in a conclusion is paired with an integrability conjunct. Without these, a heavy-tailed xnx_nxn​ would satisfy (A6) vacuously and a conclusion could hold through the value 000; neither formalization is acceptable. Independence is of the pairs (xn,zn)(x_n,z_n)(xn​,zn​), not of xnx_nxn​ and znz_nzn​ separately. (A4) is attainment of the minimum, not a gradient condition.

The following hypotheses are added to the page and disclosed in each item:

  • γ>0\gamma>0γ>0 (a step size, and γR2\sqrt{\gamma R^2}γR2​ is a denominator);
  • n≥1n\ge1n≥1;
  • the positivity and self-adjointness of HHH in Lemma 2;
  • γR2<1\gamma R^2<1γR2<1 instead of ≤1\le1≤1 in the remainder bound, which divides by 1−γR21-\gamma R^21−γR2.

A complete development needs:

  • conditional expectations of Rd\mathbb R^dRd-valued martingale differences, and the orthogonality of their sums;
  • independence of a fresh observation from the past iterates;
  • spectral calculus for (I−γH)k(I-\gamma H)^k(I−γH)k;
  • Minkowski's inequality in L2L^2L2.

All of these are reusable for other stochastic-approximation missions. Proofs of any milestone are welcome, as are alternative arguments for the noise bound that avoid the expansion.

Selected references

  • F. Bach and E. Moulines, Non-strongly-convex smooth stochastic approximation with convergence rate O(1/n), Advances in Neural Information Processing Systems 26, 2013. https://arxiv.org/abs/1306.2119
  • B. T. Polyak and A. B. Juditsky, Acceleration of stochastic approximation by averaging, SIAM Journal on Control and Optimization 30(4), 1992. https://doi.org/10.1137/0330046
  • R. Aguech, E. Moulines and P. Priouret, On a perturbation approach for the analysis of stochastic tracking algorithms, SIAM Journal on Control and Optimization 39(3), 2000. https://doi.org/10.1137/S0363012997331639
  • F. Bach and E. Moulines, Non-asymptotic analysis of stochastic approximation algorithms for machine learning, Advances in Neural Information Processing Systems 24, 2011. https://hal.science/hal-00608041
15 thms1 active userReviewed
OptimizationReinforcement Learning·Captain: mikedeng1

On the Theory of Policy Gradient Methods: Optimality, Approximation, and Distribution Shift 1: Projected Gradient Ascent on the Simplex Is ε-Optimal After 64γ|S||A|D∞²/((1−γ)⁶ε²) IterationsResearch Paper

Motivation

Policy gradient methods optimize a parameterized policy of a Markov decision process by gradient ascent on its expected discounted return. They are among the most widely used methods in reinforcement learning, from REINFORCE (Williams 1992) and the policy gradient theorem to natural policy gradient and trust-region methods. The objective is not concave in the policy, even when the policy is a raw table of action probabilities, so standard optimization theory guarantees at best convergence to a stationary point, and it was long unclear whether or how fast these methods find an optimal policy.

Agarwal, Kakade, Lee and Mahajan (JMLR 2021) give a systematic answer for the tabular and function-approximation settings. This mission formalizes their warm-up result, Theorem 4.1: projected gradient ascent over the simplex of stochastic policies reaches an ϵ\epsilonϵ-optimal policy after a number of iterations polynomial in the sizes of the MDP, the effective horizon 1/(1−γ)1/(1-\gamma)1/(1−γ), 1/ϵ1/\epsilon1/ϵ, and a distribution mismatch coefficient. The gradient domination idea it rests on goes back to the analysis of conservative policy iteration by Kakade and Langford (2002) and to Scherrer and Geist (2014).

Setting

A finite discounted MDP consists of finite sets S\mathcal SS of states and A\mathcal AA of actions, a transition kernel P(s′∣s,a)P(s'\mid s,a)P(s′∣s,a), rewards r(s,a)∈[0,1]r(s,a)\in[0,1]r(s,a)∈[0,1] and a discount factor γ∈[0,1)\gamma\in[0,1)γ∈[0,1). A policy π\piπ assigns to each state a probability distribution π(⋅∣s)\pi(\cdot\mid s)π(⋅∣s) over actions. Its value from a start state s0s_0s0​ is

Vπ(s0)=E[∑t=0∞γtr(st,at) ∣ s0],at∼π(⋅∣st), st+1∼P(⋅∣st,at),V^\pi(s_0)=\mathbb E\Big[\sum_{t=0}^\infty\gamma^t r(s_t,a_t)\,\Big|\,s_0\Big],\qquad a_t\sim\pi(\cdot\mid s_t),\ s_{t+1}\sim P(\cdot\mid s_t,a_t),Vπ(s0​)=E[t=0∑∞​γtr(st​,at​)​s0​],at​∼π(⋅∣st​), st+1​∼P(⋅∣st​,at​),

and for a start distribution ρ\rhoρ, Vπ(ρ)=∑sρ(s)Vπ(s)V^\pi(\rho)=\sum_s\rho(s)V^\pi(s)Vπ(ρ)=∑s​ρ(s)Vπ(s). The action value is Qπ(s,a)=r(s,a)+γ∑s′P(s′∣s,a)Vπ(s′)Q^\pi(s,a)=r(s,a)+\gamma\sum_{s'}P(s'\mid s,a)V^\pi(s')Qπ(s,a)=r(s,a)+γ∑s′​P(s′∣s,a)Vπ(s′) and the advantage is Aπ(s,a)=Qπ(s,a)−Vπ(s)A^\pi(s,a)=Q^\pi(s,a)-V^\pi(s)Aπ(s,a)=Qπ(s,a)−Vπ(s). The discounted state visitation distribution is dρπ(s)=(1−γ)∑t≥0γtPr⁡π(st=s∣s0∼ρ)d^\pi_\rho(s)=(1-\gamma)\sum_{t\ge0}\gamma^t\Pr^\pi(s_t=s\mid s_0\sim\rho)dρπ​(s)=(1−γ)∑t≥0​γtPrπ(st​=s∣s0​∼ρ). An optimal policy π⋆\pi^\starπ⋆ maximizes Vπ(s)V^\pi(s)Vπ(s) at every state simultaneously; V⋆=Vπ⋆V^\star=V^{\pi^\star}V⋆=Vπ⋆.

In the direct parameterization the parameter is the table itself, πs,a=π(a∣s)\pi_{s,a}=\pi(a\mid s)πs,a​=π(a∣s), a point of the product simplex Δ(A)∣S∣⊆RS×A\Delta(\mathcal A)^{|\mathcal S|}\subseteq\mathbb R^{\mathcal S\times\mathcal A}Δ(A)∣S∣⊆RS×A. The algorithm optimizes Vπ(μ)V^\pi(\mu)Vπ(μ) for a chosen start distribution μ\muμ by projected gradient ascent

π(t+1)=PΔ(A)∣S∣(π(t)+η∇πV(t)(μ)),\pi^{(t+1)}=P_{\Delta(\mathcal A)^{|\mathcal S|}}\big(\pi^{(t)}+\eta\nabla_\pi V^{(t)}(\mu)\big),π(t+1)=PΔ(A)∣S∣​(π(t)+η∇π​V(t)(μ)),

where PΔ(A)∣S∣P_{\Delta(\mathcal A)^{|\mathcal S|}}PΔ(A)∣S∣​ is the Euclidean projection and V(t)=Vπ(t)V^{(t)}=V^{\pi^{(t)}}V(t)=Vπ(t). Performance is measured under a possibly different distribution ρ\rhoρ, and the distribution mismatch coefficient ∥dρπ⋆/μ∥∞\|d^{\pi^\star}_\rho/\mu\|_\infty∥dρπ⋆​/μ∥∞​ (componentwise ratio) measures how well μ\muμ covers the states an optimal policy visits from ρ\rhoρ.

Formalization targets

Goal: Theorem 4.1

With step size η=(1−γ)3/(2γ∣A∣)\eta=(1-\gamma)^3/(2\gamma|\mathcal A|)η=(1−γ)3/(2γ∣A∣), from any initial policy, for every ρ∈Δ(S)\rho\in\Delta(\mathcal S)ρ∈Δ(S) and ϵ>0\epsilon>0ϵ>0,

min⁡t≤T{V⋆(ρ)−V(t)(ρ)}≤ϵwheneverT>64γ∣S∣∣A∣(1−γ)6ϵ2∥dρπ⋆μ∥∞2.\min_{t\le T}\big\{V^\star(\rho)-V^{(t)}(\rho)\big\}\le\epsilon\qquad\text{whenever}\qquad T>\frac{64\gamma|\mathcal S||\mathcal A|}{(1-\gamma)^6\epsilon^2}\Big\|\frac{d^{\pi^\star}_\rho}{\mu}\Big\|_\infty^2 .t≤Tmin​{V⋆(ρ)−V(t)(ρ)}≤ϵwheneverT>(1−γ)6ϵ264γ∣S∣∣A∣​​μdρπ⋆​​​∞2​.

Milestones, in the order of the proof

  1. Lemma 3.2 (performance difference): Vπ(s0)−Vπ′(s0)=11−γEs∼ds0πEa∼π(⋅∣s)[Aπ′(s,a)]V^\pi(s_0)-V^{\pi'}(s_0)=\frac1{1-\gamma}\mathbb E_{s\sim d^\pi_{s_0}}\mathbb E_{a\sim\pi(\cdot\mid s)}[A^{\pi'}(s,a)]Vπ(s0​)−Vπ′(s0​)=1−γ1​Es∼ds0​π​​Ea∼π(⋅∣s)​[Aπ′(s,a)].
  2. (7), the gradient of the direct parameterization: ∂Vπ(μ)/∂π(a∣s)=11−γdμπ(s)Qπ(s,a)\partial V^\pi(\mu)/\partial\pi(a\mid s)=\frac1{1-\gamma}d^\pi_\mu(s)Q^\pi(s,a)∂Vπ(μ)/∂π(a∣s)=1−γ1​dμπ​(s)Qπ(s,a).
  3. Lemma 4.1 (gradient domination): V⋆(ρ)−Vπ(ρ)≤11−γ∥dρπ⋆/μ∥∞max⁡πˉ(πˉ−π)⊤∇πVπ(μ)V^\star(\rho)-V^\pi(\rho)\le\frac1{1-\gamma}\|d^{\pi^\star}_\rho/\mu\|_\infty\max_{\bar\pi}(\bar\pi-\pi)^\top\nabla_\pi V^\pi(\mu)V⋆(ρ)−Vπ(ρ)≤1−γ1​∥dρπ⋆​/μ∥∞​maxπˉ​(πˉ−π)⊤∇π​Vπ(μ), together with the sharper form with dμπd^\pi_\mudμπ​ in place of (1−γ)μ(1-\gamma)\mu(1−γ)μ.
  4. Lemma D.3 (smoothness): ∥∇πVπ(s0)−∇πVπ′(s0)∥2≤2γ∣A∣(1−γ)3∥π−π′∥2\|\nabla_\pi V^\pi(s_0)-\nabla_\pi V^{\pi'}(s_0)\|_2\le\frac{2\gamma|\mathcal A|}{(1-\gamma)^3}\|\pi-\pi'\|_2∥∇π​Vπ(s0​)−∇π​Vπ′(s0​)∥2​≤(1−γ)32γ∣A∣​∥π−π′∥2​.
  5. Theorem E.1(3) (Beck 2017, Theorem 10.15): projected gradient descent with step 1/β1/\beta1/β on a β\betaβ-smooth function over a closed convex set has min⁡t<T∥Gη(xt)∥≤2β(f(x0)−f(x∗))/T\min_{t<T}\|G^\eta(x_t)\|\le\sqrt{2\beta(f(x_0)-f(x^*))}/\sqrt Tmint<T​∥Gη(xt​)∥≤2β(f(x0​)−f(x∗))​/T​, with GηG^\etaGη the gradient mapping.
  6. Proposition B.1: a gradient mapping of norm at most ϵ\epsilonϵ at π\piπ makes the next iterate π+\pi^+π+ ϵ(ηβ+1)\epsilon(\eta\beta+1)ϵ(ηβ+1)-stationary over feasible unit directions.

Significance

The theorem shows that, for the simplest constrained parameterization, a first-order method finds a globally optimal policy at a polynomial rate in spite of non-concavity. The guarantee holds for every performance distribution ρ\rhoρ at once, and it isolates the role of exploration in a single quantity, the mismatch coefficient; Section 4.3 of the paper shows that without a well-covering μ\muμ gradient methods can need exponentially many steps. Lemma 4.1 and the smoothness bound are reused across the rest of the paper, and the performance difference lemma underlies essentially all of its analyses.

All results here are proved in the paper, with Theorem E.1 and Theorem E.2 cited from Beck (2017) and Ghadimi–Lan (2016). None of them has a machine-checked proof on the platform. A formal development provides a verified link between the policy gradient expression of the direct parameterization and a standard nonconvex projected-gradient rate, and a reusable formal library of discounted visitation distributions, the performance difference identity, and projected gradient methods on Euclidean spaces.

Difficulty

The obvious argument, "projected gradient ascent converges to a stationary point, and stationary points are optimal", fails on both counts as stated. Stationary points of Vπ(μ)V^\pi(\mu)Vπ(μ) need not be optimal when μ\muμ does not cover the relevant states; the quantitative replacement is gradient domination, which only controls suboptimality through the mismatch coefficient. The convergence rate itself requires smoothness of the value as a function of the policy table, which is a bound on second derivatives of a matrix inverse (I−γPπ)−1(I-\gamma P_\pi)^{-1}(I−γPπ​)−1 with the dependence (1−γ)−3(1-\gamma)^{-3}(1−γ)−3 and the factor ∣A∣|\mathcal A|∣A∣ made explicit. Finally, the near-stationarity delivered by the gradient-mapping rate is at the next iterate, not the current one, which is why the conclusion is over t∈{0,…,T}t\in\{0,\dots,T\}t∈{0,…,T}. On the Lean side, the value is an infinite series in the policy entries, so its differentiability and the exact gradient formula have to be established for a function defined on the whole parameter space.

Formalization scope

Policies are parameter vectors in EuclideanSpace ℝ (S × A), so norms are ℓ2\ell_2ℓ2​ and Mathlib's gradient is ∇π\nabla_\pi∇π​; the objective π↦Vπ(μ)\pi\mapsto V^\pi(\mu)π↦Vπ(μ) is defined on the whole space and is only ever evaluated, with its gradient, at policies. The MDP layer (transition kernels, policies, VπV^\piVπ, QπQ^\piQπ, occupation distributions, optimal policies) is the published FoundationsML.ReinforcementLearning library; VπV^\piVπ is the unnormalized discounted sum. The projection is any map satisfying the nearest-point property. The optimal policy is a hypothesis IsOptimalPolicy (optimal from every state), not a supremum over all functions.

Conventions committed to:

  • The mismatch coefficient is any constant DDD with dρπ⋆(s)≤Dμ(s)d^{\pi^\star}_\rho(s)\le D\mu(s)dρπ⋆​(s)≤Dμ(s) for all sss; this avoids Lean's x/0=0x/0=0x/0=0 and is equivalent to the page's statement when the coefficient is finite.
  • γ>0\gamma>0γ>0 and ϵ>0\epsilon>0ϵ>0 are explicit hypotheses of the goal (the step size divides by γ\gammaγ, the threshold by ϵ\epsilonϵ).
  • The goal concludes ∃ t≤T\exists\,t\le T∃t≤T. The printed min⁡t<T\min_{t<T}mint<T​ fails at T=1T=1T=1 (one state, two actions with rewards 111 and 000, γ=0.001\gamma=0.001γ=0.001, ϵ=1/2\epsilon=1/2ϵ=1/2, initial policy on the bad action); the proof on p. 50 establishes the range 0≤t≤T0\le t\le T0≤t≤T.
  • Proposition B.1 bounds the directions feasible at π+\pi^+π+, as its proof does; Theorem E.1 assumes smoothness on CCC only and uses the radicand 2β(f(x0)−f(x∗))2\beta(f(x_0)-f(x^*))2β(f(x0​)−f(x∗)) of Beck and of p. 49.

A formalization that defines the update through formula (7), that assumes gradient domination or smoothness as hypotheses of the goal, or that takes the gradient off the simplex where the value series may diverge, would trivialize the goal; the goal mentions none of these, and (7) is a milestone theorem.

A complete development needs: summability and differentiability of the value series near the simplex; the performance difference lemma; the Euclidean projection inequality on a closed convex set; the descent lemma for functions smooth on a convex set; and the gradient-mapping argument. The projection and gradient-mapping results are independent of reinforcement learning and reusable. Contributions to any milestone are welcome, as are proofs of Theorem E.2 (Ghadimi–Lan) as a stepping stone to Proposition B.1.

Selected references

  • A. Agarwal, S. M. Kakade, J. D. Lee, G. Mahajan, On the Theory of Policy Gradient Methods: Optimality, Approximation, and Distribution Shift, JMLR 22(98), 2021; arXiv:1908.00261v5. https://arxiv.org/abs/1908.00261
  • A. Beck, First-Order Methods in Optimization, MOS-SIAM Series on Optimization, SIAM, 2017. https://doi.org/10.1137/1.9781611974997
  • S. Ghadimi, G. Lan, Accelerated gradient methods for nonconvex nonlinear and stochastic programming, Mathematical Programming 156, 2016. https://doi.org/10.1007/s10107-015-0871-8
  • S. Kakade, J. Langford, Approximately optimal approximate reinforcement learning, ICML 2002. https://dl.acm.org/doi/10.5555/645531.656005
  • B. Scherrer, M. Geist, Local Policy Search in a Convex Space and Conservative Policy Iteration as Boosted Policy Search, ECML PKDD 2014, pp. 35–50 (arXiv version: https://arxiv.org/abs/1306.1520)
  • R. J. Williams, Simple statistical gradient-following algorithms for connectionist reinforcement learning, Machine Learning 8, 1992. https://doi.org/10.1007/BF00992696
18 thms1 active userReviewed
Optimal TransportOptimizationStatistics·Captain: mikedeng1

Robust Wasserstein Profile Inference and Applications to Machine Learning 1: Square-Root LASSO Is Wasserstein DRO — the Worst-Case Squared Loss over D_c(P, P_n) ≤ δ Equals (√MSE_n(β) + √δ‖β‖_p)²Research Paper

Motivation

Regularized least squares is the standard tool of high-dimensional linear regression. The square-root LASSO of Belloni, Chernozhukov and Wang (Biometrika, 2011) minimizes MSEn(β)+λ∥β∥1\sqrt{\mathrm{MSE}_n(\beta)} + \lambda\|\beta\|_1MSEn​(β)​+λ∥β∥1​. Unlike the LASSO, its optimal regularization parameter does not depend on the unknown noise level. Regularization is usually justified through sparsity or bias–variance arguments. Blanchet, Kang and Murthy (arXiv:1610.05627, J. Appl. Probab. 56(3), 2019) give a different justification. The square-root LASSO, and every ℓp\ell_pℓp​-penalized square-root least-squares estimator, is exactly a distributionally robust estimator. It minimizes the worst-case expected square loss over all data distributions within a given optimal-transport distance of the empirical distribution.

The rest of the paper builds on this representation: the radius of the transport ball is the regularization parameter, which the paper's Robust Wasserstein Profile function selects by a statistical criterion (mission 3 of this series). The duality theorem underneath, Proposition 1, is due to Blanchet and Murthy (Math. Oper. Res., 2019). Closely related representations for logistic regression appear in Shafieezadeh-Abadeh, Mohajerin Esfahani and Kuhn (NeurIPS 2015), where they are approximate. The cost function introduced in this paper makes them exact.

Setting

The training data are n≥1n \ge 1n≥1 pairs (X1,Y1),…,(Xn,Yn)(X_1, Y_1), \dots, (X_n, Y_n)(X1​,Y1​),…,(Xn​,Yn​) with predictors Xi∈RdX_i \in \mathbb R^dXi​∈Rd and responses Yi∈RY_i \in \mathbb RYi​∈R. No distributional assumption is made; the data are fixed vectors. The empirical distribution is Pn=1n∑i=1nδ(Xi,Yi)P_n = \frac1n \sum_{i=1}^n \delta_{(X_i, Y_i)}Pn​=n1​∑i=1n​δ(Xi​,Yi​)​. For β∈Rd\beta \in \mathbb R^dβ∈Rd the square loss is l(x,y;β)=(y−βTx)2l(x, y; \beta) = (y - \beta^T x)^2l(x,y;β)=(y−βTx)2 and the mean square error is MSEn(β)=1n∑i=1n(Yi−βTXi)2\mathrm{MSE}_n(\beta) = \frac1n\sum_{i=1}^n (Y_i - \beta^T X_i)^2MSEn​(β)=n1​∑i=1n​(Yi​−βTXi​)2.

A cost function ccc assigns to two points z,wz, wz,w of Rd×R\mathbb R^d \times \mathbb RRd×R a value c(z,w)∈[0,∞]c(z, w) \in [0, \infty]c(z,w)∈[0,∞], the cost of moving a unit of mass from zzz to www. The optimal transport cost between probability measures PPP and QQQ is

Dc(P,Q)=inf⁡{Eπ[c(U,W)]:π a probability measure on pairs (U,W), πU=P, πW=Q}.(7)D_c(P, Q) = \inf\Big\{ \mathbb E_\pi[c(U, W)] : \pi \text{ a probability measure on pairs } (U, W),\ \pi_U = P,\ \pi_W = Q \Big\}. \qquad (7)Dc​(P,Q)=inf{Eπ​[c(U,W)]:π a probability measure on pairs (U,W), πU​=P, πW​=Q}.(7)

The worst-case expected loss at radius δ≥0\delta \ge 0δ≥0 is sup⁡P:Dc(P,Pn)≤δEP[l(X,Y;β)]\sup_{P : D_c(P, P_n) \le \delta} \mathbb E_P[l(X, Y; \beta)]supP:Dc​(P,Pn​)≤δ​EP​[l(X,Y;β)], and the distributionally robust regression problem (8) minimizes it over β\betaβ.

Two costs are used. With q∈(1,∞]q \in (1, \infty]q∈(1,∞]:

  • the squared ℓq\ell_qℓq​ cost on Rd+1\mathbb R^{d+1}Rd+1, c((x,y),(u,v))=∥(x,y)−(u,v)∥q2c((x, y), (u, v)) = \|(x, y) - (u, v)\|_q^2c((x,y),(u,v))=∥(x,y)−(u,v)∥q2​ (Proposition 2);
  • the cost Nq2N_q^2Nq2​, where (14) Nq((x,y),(u,v))=∥x−u∥qN_q((x, y), (u, v)) = \|x - u\|_qNq​((x,y),(u,v))=∥x−u∥q​ if y=vy = vy=v and +∞+\infty+∞ otherwise. Under this cost the responses cannot be moved, and only the predictors are perturbed (Theorem 1).

The exponent ppp is the dual of qqq, 1/p+1/q=11/p + 1/q = 11/p+1/q=1, and βˉ=(−β,1)\bar\beta = (-\beta, 1)βˉ​=(−β,1).

Formalization targets

Goal: Theorem 1 (p. 11)

For the cost c=Nq2c = N_q^2c=Nq2​, every δ≥0\delta \ge 0δ≥0 and every β∈Rd\beta \in \mathbb R^dβ∈Rd,

sup⁡P: Dc(P,Pn)≤δEP[(Y−βTX)2]=(MSEn(β)+δ ∥β∥p)2,\sup_{P :\, D_c(P, P_n) \le \delta} \mathbb E_P\big[(Y - \beta^T X)^2\big] = \Big(\sqrt{\mathrm{MSE}_n(\beta)} + \sqrt\delta\,\|\beta\|_p\Big)^2 ,P:Dc​(P,Pn​)≤δsup​EP​[(Y−βTX)2]=(MSEn​(β)​+δ​∥β∥p​)2,

and consequently

inf⁡β∈Rdsup⁡P: Dc(P,Pn)≤δEP[(Y−βTX)2]=inf⁡β∈Rd(MSEn(β)+δ ∥β∥p)2.\inf_{\beta \in \mathbb R^d} \sup_{P :\, D_c(P, P_n) \le \delta} \mathbb E_P\big[(Y - \beta^T X)^2\big] = \inf_{\beta \in \mathbb R^d} \Big(\sqrt{\mathrm{MSE}_n(\beta)} + \sqrt\delta\,\|\beta\|_p\Big)^2 .β∈Rdinf​P:Dc​(P,Pn​)≤δsup​EP​[(Y−βTX)2]=β∈Rdinf​(MSEn​(β)​+δ​∥β∥p​)2.

The second identity is the printed theorem; the first is what its proof establishes for each β\betaβ. The goal states both.

Milestones

  1. Proposition 1 (p. 10): strong duality. For a lower semicontinuous cost vanishing on the diagonal, an upper semicontinuous loss and δ>0\delta > 0δ>0, the worst-case expected loss equals min⁡γ≥0{γδ+1n∑iφγ(Xi,Yi)}\min_{\gamma \ge 0} \{\gamma\delta + \frac1n \sum_i \varphi_\gamma(X_i, Y_i)\}minγ≥0​{γδ+n1​∑i​φγ​(Xi​,Yi​)}, with φγ(z)=sup⁡u{l(u)−γc(u,z)}\varphi_\gamma(z) = \sup_u \{l(u) - \gamma c(u, z)\}φγ​(z)=supu​{l(u)−γc(u,z)} (11).
  2. (28) (pp. 28–29): the closed form of φγ\varphi_\gammaφγ​ for the square loss and the squared ℓq\ell_qℓq​ cost.
  3. (29) and the display after it (p. 29): inf⁡γ>b2{γδ+γγ−b2M}=(M+bδ)2\inf_{\gamma > b^2} \{\gamma\delta + \frac{\gamma}{\gamma - b^2} M\} = (\sqrt M + b\sqrt\delta)^2infγ>b2​{γδ+γ−b2γ​M}=(M​+bδ​)2 for M,b,δ≥0M, b, \delta \ge 0M,b,δ≥0.
  4. Proposition 2 (p. 10): the analogue of the goal for the squared ℓq\ell_qℓq​ cost, with ∥βˉ∥p\|\bar\beta\|_p∥βˉ​∥p​ in place of ∥β∥p\|\beta\|_p∥β∥p​ (13).
  5. Outline of the proof of Theorem 1, last display (p. 29): the closed form of φγ\varphi_\gammaφγ​ for the cost Nq2N_q^2Nq2​.

Significance

The result. Theorem 1 identifies ℓp\ell_pℓp​-penalized square-root least squares with a min–max problem over data distributions. For q=∞q = \inftyq=∞, p=1p = 1p=1 the minimizers are those of the square-root LASSO with λ=δ\lambda = \sqrt\deltaλ=δ​. The regularization parameter therefore acquires a meaning: it is the square root of the transport budget an adversary may spend perturbing the predictors. This is the basis of the paper's choice of δ\deltaδ by the Robust Wasserstein Profile function (§4), and of the interpretation of regularized estimators as robust to covariate perturbations. Proposition 2 shows that letting the adversary also move the responses changes the penalty to ∥(−β,1)∥p\|(-\beta, 1)\|_p∥(−β,1)∥p​, which is why the label-preserving cost NqN_qNq​ is needed for an exact match.

Formalizing it. All results are proved on paper; none is formalized. A complete development gives a machine-checked strong-duality theorem for optimal-transport balls with possibly infinite costs (Proposition 1), two explicit worst-case computations, and corrected boundary cases of the closed forms (28) and the outline display, which print +∞+\infty+∞ for all γ≤∥βˉ∥p2\gamma \le \|\bar\beta\|_p^2γ≤∥βˉ​∥p2​ although the value can be finite at equality. The corrections do not affect the theorems.

Difficulty

The obvious argument fails in two places. The first is the duality step: the supremum ranges over all Borel probability measures on Rd+1\mathbb R^{d+1}Rd+1 within transport cost δ\deltaδ, an infinite-dimensional set that is not compact in any convenient topology, with a loss that is unbounded above. Exchanging the supremum with the Lagrange multiplier of the budget constraint is Proposition 1, a theorem in its own right (Blanchet–Murthy), and its attainment claim needs δ>0\delta > 0δ>0.

The second is the cost NqN_qNq​, which is +∞+\infty+∞ off {y=v}\{y = v\}{y=v}, so the standard Wasserstein duality theorems, which assume a finite metric cost, do not apply. The degenerate cases β=0\beta = 0β=0, MSEn(β)=0\mathrm{MSE}_n(\beta) = 0MSEn​(β)=0, δ=0\delta = 0δ=0, where the objective in γ\gammaγ does not blow up at both ends, must be covered separately.

Formalization scope

  • Spaces. A data point is a pair in (Fin d → ℝ) × ℝ with the product σ-algebra and topology. Proposition 2's cost uses the stacked vector in Fin (d+1) → ℝ (response last, built with Fin.snoc), and βˉ\bar\betaβˉ​ is the stacked vector of (−β,1)(-\beta, 1)(−β,1).
  • Norms. ∥⋅∥q\|\cdot\|_q∥⋅∥q​ and ∥⋅∥p\|\cdot\|_p∥⋅∥p​ are the norms of PiLp, with exponents in ℝ≥0∞, so q=∞q = \inftyq=∞ (the square-root LASSO case) is included. The exponents are linked by p.HolderConjugate q, and q∈(1,∞]q \in (1, \infty]q∈(1,∞] throughout. Theorem 1 does not print a range for qqq; the range is taken from Proposition 2, which the paper calls essentially the same result.
  • Transport cost and worst case. Costs are ℝ≥0∞-valued, and DcD_cDc​ is an infimum over probability couplings with both marginals fixed. Expectations of the nonnegative losses are lower Lebesgue integrals, and the worst case is a supremum in ℝ≥0∞ over all probability measures in the ball. No integrability side condition removes measures from the ball. Identities with a real right-hand side are stated after embedding it with ENNReal.ofReal.
  • The empirical distribution is the published definition WassersteinDRO.Regularization.empiricalDistribution, applied to i↦(Xi,Yi)i \mapsto (X_i, Y_i)i↦(Xi​,Yi​), with n>0n > 0n>0.
  • φγ\varphi_\gammaφγ​. A point at infinite cost contributes −∞-\infty−∞ for every γ≥0\gamma \ge 0γ≥0, including γ=0\gamma = 0γ=0, as in the paper's treatment of NqN_qNq​. With the convention 0⋅∞=00 \cdot \infty = 00⋅∞=0 instead, Proposition 1's minimum would not be attained for the cost Nq2N_q^2Nq2​ at β=0\beta = 0β=0.
  • Proposition 1 is stated for a nonnegative loss and δ>0\delta > 0δ>0; both are restrictions of the page, recorded in the item.
  • Corrections. (28) and the outline display are stated with their corrected boundary cases. The one-dimensional lemma behind (29) is stated as a greatest lower bound over γ>b2\gamma > b^2γ>b2, including b=0b = 0b=0, M=0M = 0M=0, δ=0\delta = 0δ=0.

A formalization in which the transport infimum did not fix both marginals, allowed sub-probability couplings, or used a Bochner integral would make the worst case trivially +∞+\infty+∞ or 000. The conventions above rule this out: at δ=0\delta = 0δ=0 the ball is {Pn}\{P_n\}{Pn​} and both sides of the goal equal MSEn(β)\mathrm{MSE}_n(\beta)MSEn​(β).

The work needs Kantorovich-type duality for lower semicontinuous costs on Rm\mathbb R^mRm (absent from Mathlib), Hölder's inequality with its equality case for PiLp, and elementary one-variable optimization. The duality theorem and the transport-cost definition are reusable beyond this mission: mission 2 of this series (classification) uses Proposition 1 with the cost NqN_qNq​, ρ=1\rho = 1ρ=1. Contributions that prove Proposition 1, or its weak-duality half, are particularly welcome.

Selected references

  • J. Blanchet, Y. Kang, K. Murthy, Robust Wasserstein Profile Inference and Applications to Machine Learning, J. Appl. Probab. 56(3), 2019; arXiv:1610.05627v4. https://arxiv.org/abs/1610.05627
  • J. Blanchet, K. Murthy, Quantifying distributional model risk via optimal transport, Math. Oper. Res. 44(2), 2019. https://doi.org/10.1287/moor.2018.0936
  • A. Belloni, V. Chernozhukov, L. Wang, Square-root lasso: pivotal recovery of sparse signals via conic programming, Biometrika 98(4), 2011. https://doi.org/10.1093/biomet/asr043
  • S. Shafieezadeh-Abadeh, P. Mohajerin Esfahani, D. Kuhn, Distributionally robust logistic regression, NeurIPS 2015. https://arxiv.org/abs/1509.09259
  • C. Villani, Optimal Transport: Old and New, Springer, 2009. https://doi.org/10.1007/978-3-540-71050-9
14 thms1 active userReviewed
Operations ResearchOptimization·Captain: mikedeng1

Oracle-Based Robust Optimization via Online Learning 2: Follow the Perturbed Leader with an ε-Approximate Linear Oracle Has Expected Regret at Most 2√(DRAT) + 2εTResearch Paper

Motivation

Many decision problems are solved repeatedly against data that arrive over time: routing traffic, allocating budgets, choosing portfolios or combinatorial structures. Online linear optimization models this. At each round t=1,…,Tt = 1, \ldots, Tt=1,…,T a learner picks a decision xtx_txt​ from a fixed domain K⊆Rn\mathcal K\subseteq\mathbb R^nK⊆Rn, then a reward vector ftf_tft​ is revealed and the learner earns ft⋅xtf_t\cdot x_tft​⋅xt​. Performance is measured by regret, the gap to the best fixed decision in hindsight. When K\mathcal KK is combinatorial (paths, spanning trees, assignments), the only computationally reasonable access to K\mathcal KK is a procedure that optimizes a linear function over it, and in practice such procedures are often only approximate.

Follow the Perturbed Leader (FPL), introduced by Hannan (1957) and analysed for linear optimization by Kalai and Vempala (JCSS 2005), uses exactly one call to an exact linear optimizer per round and achieves regret O(T)O(\sqrt T)O(T​) over arbitrary, not necessarily convex, domains. Ben-Tal, Hazan, Koren and Mannor (arXiv:1402.6361, Operations Research 2015) needed a version of FPL that works with an additively approximate linear optimizer, as a building block for oracle-based robust optimization with linearly parametrized uncertainty sets. Their §3.3 analyses this variant and proves Theorem 6, the goal of this mission.

Setting

Fix a dimension nnn, a domain K⊆Rn\mathcal K\subseteq\mathbb R^nK⊆Rn (arbitrary: not necessarily convex, closed or bounded) and ϵ>0\epsilon > 0ϵ>0. An ϵ\epsilonϵ-approximate linear optimization procedure over K\mathcal KK is a map Mϵ:Rn→RnM_\epsilon:\mathbb R^n\to\mathbb R^nMϵ​:Rn→Rn such that, for every g∈Rng\in\mathbb R^ng∈Rn,

Mϵ(g)∈Kandg⋅Mϵ(g)  ≥  g⋅x−ϵfor all x∈K.M_\epsilon(g)\in\mathcal K \qquad\text{and}\qquad g\cdot M_\epsilon(g)\;\ge\; g\cdot x-\epsilon\quad\text{for all }x\in\mathcal K .Mϵ​(g)∈Kandg⋅Mϵ​(g)≥g⋅x−ϵfor all x∈K.

Reward vectors f1,…,fT∈Rnf_1,\ldots,f_T\in\mathbb R^nf1​,…,fT​∈Rn are fixed in advance (an oblivious adversary). Write f1:t=∑τ=1tfτf_{1:t}=\sum_{\tau=1}^t f_\tauf1:t​=∑τ=1t​fτ​, with f1:0=0f_{1:0}=0f1:0​=0, and ∥v∥1=∑i∣vi∣\|v\|_1=\sum_i|v_i|∥v∥1​=∑i​∣vi​∣.

Follow the Approximate Perturbed Leader with parameter η>0\eta>0η>0 plays at round ttt

xt=Mϵ(f1:t−1+pt),pt uniform on the cube [0,1/η]n.x_t = M_\epsilon\big(f_{1:t-1}+p_t\big),\qquad p_t \text{ uniform on the cube } [0,1/\eta]^n .xt​=Mϵ​(f1:t−1​+pt​),pt​ uniform on the cube [0,1/η]n.

Three scale parameters enter the bound: DDD bounds the ℓ1\ell_1ℓ1​ diameter of K\mathcal KK, ∥x−y∥1≤D\|x-y\|_1\le D∥x−y∥1​≤D for x,y∈Kx,y\in\mathcal Kx,y∈K; AAA bounds ∥ft∥1\|f_t\|_1∥ft​∥1​; and RRR bounds how much each reward varies over the domain, ∣ft⋅x−ft⋅y∣≤R|f_t\cdot x-f_t\cdot y|\le R∣ft​⋅x−ft​⋅y∣≤R for x,y∈Kx,y\in\mathcal Kx,y∈K.

Formalization targets

Goal: Theorem 6 (p. 11)

With η=D/(RAT)\eta=\sqrt{D/(RAT)}η=D/(RAT)​, for every x∗∈Kx^*\in\mathcal Kx∗∈K,

∑t=1Tft⋅x∗−E[∑t=1Tft⋅xt]  ≤  2DRAT+2ϵT.\sum_{t=1}^T f_t\cdot x^* - \mathbf E\Big[\sum_{t=1}^T f_t\cdot x_t\Big]\;\le\;2\sqrt{DRAT}+2\epsilon T .t=1∑T​ft​⋅x∗−E[t=1∑T​ft​⋅xt​]≤2DRAT​+2ϵT.

The bound for every η\etaη (proof of Theorem 6, p. 13)

For every η>0\eta>0η>0 and x∈Kx\in\mathcal Kx∈K,

E[∑t=1Tft⋅xt]  ≥  f1:T⋅x−Dη−ηRAT−2ϵT.\mathbf E\Big[\sum_{t=1}^T f_t\cdot x_t\Big]\;\ge\; f_{1:T}\cdot x-\frac D\eta-\eta RAT-2\epsilon T .E[t=1∑T​ft​⋅xt​]≥f1:T​⋅x−ηD​−ηRAT−2ϵT.

Supporting lemmas (pp. 12–13)

  • Lemma 7 (approximate be-the-leader): ∑t=1TMϵ(f1:t)⋅ft≥Mϵ(f1:T)⋅f1:T−ϵT\sum_{t=1}^T M_\epsilon(f_{1:t})\cdot f_t\ge M_\epsilon(f_{1:T})\cdot f_{1:T}-\epsilon T∑t=1T​Mϵ​(f1:t​)⋅ft​≥Mϵ​(f1:T​)⋅f1:T​−ϵT.
  • Lemma 8 (be the approximate perturbed leader): for T≥2T\ge2T≥2, p∈[0,1/η]np\in[0,1/\eta]^np∈[0,1/η]n and x∈Kx\in\mathcal Kx∈K, ∑t=1TMϵ(f1:t+p)⋅ft≥f1:T⋅x−D/η−2ϵT\sum_{t=1}^T M_\epsilon(f_{1:t}+p)\cdot f_t\ge f_{1:T}\cdot x-D/\eta-2\epsilon T∑t=1T​Mϵ​(f1:t​+p)⋅ft​≥f1:T​⋅x−D/η−2ϵT.
  • Lemma 9 (stability): for ppp uniform on [0,1/η]n[0,1/\eta]^n[0,1/η]n, E[Mϵ(f1:t−1+p)⋅ft]−E[Mϵ(f1:t+p)⋅ft]≥−ηRA\mathbf E[M_\epsilon(f_{1:t-1}+p)\cdot f_t]-\mathbf E[M_\epsilon(f_{1:t}+p)\cdot f_t]\ge-\eta RAE[Mϵ​(f1:t−1​+p)⋅ft​]−E[Mϵ​(f1:t​+p)⋅ft​]≥−ηRA.

Significance

Theorem 6 shows that perturbed-leader online linear optimization is robust to additive error in its optimization subroutine: an ϵ\epsilonϵ-approximate oracle costs only 2ϵT2\epsilon T2ϵT extra regret, so the average regret is 2DRA/T+2ϵ2\sqrt{DRA/T}+2\epsilon2DRA/T​+2ϵ. This allows the algorithm to be run over domains where exact linear optimization is intractable but a good additive approximation is available, and the paper invokes it as the online-learning primitive of its oracle-based scheme for linearly parametrized uncertainty in §3.2 (that application is not part of this mission). Unlike online gradient methods, it requires no convexity of K\mathcal KK and no projection.

On the formal side, no regret bound for Follow the Perturbed Leader, exact or approximate, is currently formalized on the platform, and the Kalai–Vempala stability argument (comparing a uniform distribution on a cube with its translate) is a reusable piece of measure theory. The mission's statements are proved on paper; the work here is to formalize those proofs, with one correction to a hypothesis, explained under Formalization scope.

Difficulty

Lemmas 7 and 8 are deterministic and combinatorial. The substance is Lemma 9. It compares the expectations of one bounded function of Mϵ(⋅)M_\epsilon(\cdot)Mϵ​(⋅) under the uniform law on a cube and under its translate by ftf_tft​. The natural first attempt, a pointwise comparison of Mϵ(f1:t−1+p)M_\epsilon(f_{1:t-1}+p)Mϵ​(f1:t−1​+p) and Mϵ(f1:t+p)M_\epsilon(f_{1:t}+p)Mϵ​(f1:t​+p), fails: an approximate (even an exact) maximizer can jump arbitrarily under an arbitrarily small change of its input, and MϵM_\epsilonMϵ​ is not assumed continuous or even consistent between nearby inputs. Any valid argument must therefore control the two distributions as a whole rather than the decisions point by point, which in the formal development involves Lebesgue measure on Rn\mathbb R^nRn, conditioning on a box and translation invariance.

A second subtlety is that the stability bound depends on how RRR is read, which is the reason for the correction below.

Formalization scope

Vectors are Fin n → ℝ with dotProduct. All ℓ1\ell_1ℓ1​ quantities are written as ∑i∣vi∣\sum_i|v_i|∑i​∣vi​∣, never with the default norm (the sup norm). Rewards are a function f : ℕ → Fin n → ℝ read at t=1,…,Tt=1,\ldots,Tt=1,…,T, and f1:tf_{1:t}f1:t​ is prefixSum f t. The perturbation law is Lebesgue measure conditioned on the cube [0,1/η]n[0,1/\eta]^n[0,1/η]n (ProbabilityTheory.cond volume), a probability measure for η>0\eta>0η>0. Maxima over K\mathcal KK are expressed as "for every x∈Kx\in\mathcal Kx∈K", so neither attainment nor boundedness of K\mathcal KK is presupposed.

Conventions and deviations, each also stated in the affected item:

  1. RRR is an oscillation bound. The paper takes R≥max⁡t,x∣ft⋅x∣R\ge\max_{t,x}|f_t\cdot x|R≥maxt,x​∣ft​⋅x∣. With that reading Lemma 9 is false (for K={−1,1}\mathcal K=\{-1,1\}K={−1,1}, the exact maximizer, f1:t−1=−Af_{1:t-1}=-Af1:t−1​=−A, ft=A=Rf_t=A=Rft​=A=R, ηA≤1\eta A\le1ηA≤1, the left side is −2ηRA-2\eta RA−2ηRA), and the printed constant in Theorem 6 does not follow. The proof's step "they can differ by at most RRR" is correct when R≥∣ft⋅x−ft⋅y∣R\ge|f_t\cdot x-f_t\cdot y|R≥∣ft​⋅x−ft​⋅y∣ for x,y∈Kx,y\in\mathcal Kx,y∈K; Lemma 9, the display and Theorem 6 are stated with that hypothesis. The printed hypothesis implies it with 2R2R2R; for non-negative rewards the two coincide.
  2. Expected reward. E[∑tft⋅xt]\mathbf E[\sum_t f_t\cdot x_t]E[∑t​ft​⋅xt​] is written as ∑t∫ft⋅Mϵ(f1:t−1+p) dμη(p)\sum_t\int f_t\cdot M_\epsilon(f_{1:t-1}+p)\,d\mu_\eta(p)∑t​∫ft​⋅Mϵ​(f1:t−1​+p)dμη​(p), which by linearity of expectation is the same for independent or shared perturbations (the paper makes the same observation).
  3. Printed typos. In (13) the summand ftf_tft​ is fτf_\taufτ​ and round ttt uses f1:t−1f_{1:t-1}f1:t−1​; in Lemma 8 and the display, max⁡xf1:t⋅x\max_{x}f_{1:t}\cdot xmaxx​f1:t​⋅x means f1:Tf_{1:T}f1:T​.
  4. Added hypotheses. MϵM_\epsilonMϵ​ is measurable (otherwise every expectation would be a Bochner integral of a non-measurable function and equal 000); an approximate maximizer can always be chosen measurable. D,R,A>0D,R,A>0D,R,A>0 and T≥1T\ge1T≥1 make η=D/(RAT)\eta=\sqrt{D/(RAT)}η=D/(RAT)​ a positive real. The display is stated for T≥1T\ge1T≥1 (Lemma 8 needs T≥2T\ge2T≥2 as printed; the case T=1T=1T=1 also holds).
  5. No O(⋅)O(\cdot)O(⋅) appears: all constants are the paper's explicit ones.

A trivializing formalization is ruled out: MϵM_\epsilonMϵ​ must return points of K\mathcal KK (otherwise DDD would not bound ∥Mϵ(⋅)−Mϵ(⋅)∥1\|M_\epsilon(\cdot)-M_\epsilon(\cdot)\|_1∥Mϵ​(⋅)−Mϵ​(⋅)∥1​), it must be measurable, and the perturbation law is the normalized uniform distribution, not Lebesgue measure restricted to the cube (which is not a probability measure for η≠1\eta\neq1η=1).

Needed infrastructure: the overlap estimate for a cube and its translate, vol([0,1/η]n∩(v+[0,1/η]n))≥(1−η∥v∥1) η−n\mathrm{vol}([0,1/\eta]^n\cap(v+[0,1/\eta]^n))\ge(1-\eta\|v\|_1)\,\eta^{-n}vol([0,1/η]n∩(v+[0,1/η]n))≥(1−η∥v∥1​)η−n, and the integrability of bounded measurable functions of MϵM_\epsilonMϵ​. Both are reusable for any perturbation-based online-learning analysis; contributions of these as standalone lemmas are welcome.

Selected references

  • A. Ben-Tal, E. Hazan, T. Koren, S. Mannor, Oracle-Based Robust Optimization via Online Learning, Operations Research 63(3), 2015; preprint arXiv:1402.6361v1, 2014. https://arxiv.org/abs/1402.6361
  • A. Kalai, S. Vempala, Efficient algorithms for online decision problems, Journal of Computer and System Sciences 71(3), 291–307, 2005. https://doi.org/10.1016/j.jcss.2004.10.016
  • J. Hannan, Approximation to Bayes risk in repeated play, Contributions to the Theory of Games III, Annals of Mathematics Studies 39, 97–139, 1957.
7 thms1 active userReviewed
Convex OptimizationOperations ResearchOptimization·Captain: mikedeng1

Oracle-Based Robust Optimization via Online Learning 1: The Dual-Subgradient Meta-Algorithm Returns a 2ε-Approximate Robust Solution or Certifies Infeasibility within ⌈G²D²/ε²⌉ Oracle CallsResearch Paper

Motivation

Robust optimization protects a decision against every realization of uncertain data in a prescribed uncertainty set. The standard approach replaces the uncertain constraints by a deterministic robust counterpart and solves that counterpart directly (Ben-Tal, El Ghaoui, Nemirovski, Robust Optimization, 2009). The counterpart is often a harder problem than the original: a robust linear program with ellipsoidal uncertainty becomes a second-order cone program, and a robust quadratic program can become a semidefinite program. A practitioner who has an efficient, specialised solver for the nominal problem may therefore have no efficient solver for its robust version.

Ben-Tal, Hazan, Koren and Mannor (arXiv:1402.6361, Operations Research 2015) ask whether the robust problem can be solved by repeatedly calling a solver of the nominal problem, with the number of calls independent of the dimension. Their first answer, the dual-subgradient meta-algorithm of §3.1, does so whenever the constraints are concave in the noise and the uncertainty set is convex. It is a primal–dual scheme: an online-learning algorithm picks the noise, and the nominal solver answers. This mission formalizes that result, Theorem 3.

Setting

Let D⊆Rn\mathcal D\subseteq\mathbb R^nD⊆Rn be a convex domain, U⊆Rd\mathcal U\subseteq\mathbb R^dU⊆Rd a convex uncertainty set, and f1,…,fm:Rn×Rd→Rf_1,\dots,f_m:\mathbb R^n\times\mathbb R^d\to\mathbb Rf1​,…,fm​:Rn×Rd→R constraint functions. The robust feasibility problem (3) is

∃ x∈D:fi(x,ui)≤0∀ui∈U, i=1,…,m.\exists\,x\in\mathcal D:\qquad f_i(x,u_i)\le 0\quad\forall u_i\in\mathcal U,\ i=1,\dots,m .∃x∈D:fi​(x,ui​)≤0∀ui​∈U, i=1,…,m.

(An objective is handled by binary search on its value, so feasibility is the core question.) A point x∈Dx\in\mathcal Dx∈D is an ϵ\epsilonϵ-approximate solution if fi(x,u)≤ϵf_i(x,u)\le\epsilonfi​(x,u)≤ϵ for all u∈Uu\in\mathcal Uu∈U and all iii.

An ϵ\epsilonϵ-approximate oracle Oϵ\mathcal O_\epsilonOϵ​ (Figure 1) takes a noise vector u=(u1,…,um)∈Umu=(u_1,\dots,u_m)\in\mathcal U^mu=(u1​,…,um​)∈Um and either returns some x∈Dx\in\mathcal Dx∈D with fi(x,ui)≤ϵf_i(x,u_i)\le\epsilonfi​(x,ui​)≤ϵ for all iii, or answers "infeasible", which it may do only if no x∈Dx\in\mathcal Dx∈D has fi(x,ui)≤0f_i(x,u_i)\le 0fi​(x,ui​)≤0 for all iii.

The standing assumptions of §3.1 are: each fi(⋅,u)f_i(\cdot,u)fi​(⋅,u) is convex on D\mathcal DD; each fi(x,⋅)f_i(x,\cdot)fi​(x,⋅) is concave on U\mathcal UU for x∈Dx\in\mathcal Dx∈D; D≥∥u−v∥2D\ge\|u-v\|_2D≥∥u−v∥2​ for all u,v∈Uu,v\in\mathcal Uu,v∈U; and ∥∇ufi(x,u)∥2≤G\|\nabla_u f_i(x,u)\|_2\le G∥∇u​fi​(x,u)∥2​≤G for x∈Dx\in\mathcal Dx∈D, u∈Uu\in\mathcal Uu∈U. Write PPP for the Euclidean projection onto U\mathcal UU.

Algorithm 1 sets T=⌈G2D2/ϵ2⌉T=\lceil G^2D^2/\epsilon^2\rceilT=⌈G2D2/ϵ2⌉ and η=D/(GT)\eta=D/(G\sqrt T)η=D/(GT​), starts from u10,…,um0∈Uu^0_1,\dots,u^0_m\in\mathcal Uu10​,…,um0​∈U, and for t=1,…,Tt=1,\dots,Tt=1,…,T updates

uit=P(uit−1+η ∇ufi(xt−1,uit−1)),xt=Oϵ(u1t,…,umt),u^t_i=P\bigl(u^{t-1}_i+\eta\,\nabla_u f_i(x^{t-1},u^{t-1}_i)\bigr),\qquad x^t=\mathcal O_\epsilon(u^t_1,\dots,u^t_m),uit​=P(uit−1​+η∇u​fi​(xt−1,uit−1​)),xt=Oϵ​(u1t​,…,umt​),

stopping with "infeasible" as soon as the oracle says so, and otherwise returning xˉ=1T∑t=1Txt\bar x=\frac1T\sum_{t=1}^T x^txˉ=T1​∑t=1T​xt. In Lean these are alg1T, alg1Eta, alg1U, alg1X, alg1Output and alg1Calls in the namespace OracleRO.DualSubgrad.

Formalization targets

Goal: Theorem 3 (p. 7)

For every ϵ\epsilonϵ-approximate oracle,

output="infeasible" ⟹ ¬ ∃x∈D ∀i ∀u∈U: fi(x,u)≤0,\text{output}=\text{"infeasible"}\ \Longrightarrow\ \neg\,\exists x\in\mathcal D\ \forall i\ \forall u\in\mathcal U:\ f_i(x,u)\le 0,output="infeasible" ⟹ ¬∃x∈D ∀i ∀u∈U: fi​(x,u)≤0, output=xˉ ⟹ xˉ∈D  and  fi(xˉ,u)≤2ϵ  ∀i, ∀u∈U,\text{output}=\bar x\ \Longrightarrow\ \bar x\in\mathcal D\ \text{ and }\ f_i(\bar x,u)\le 2\epsilon\ \ \forall i,\ \forall u\in\mathcal U,output=xˉ ⟹ xˉ∈D  and  fi​(xˉ,u)≤2ϵ  ∀i, ∀u∈U,

and the number of oracle calls is at most ⌈G2D2/ϵ2⌉\lceil G^2D^2/\epsilon^2\rceil⌈G2D2/ϵ2⌉.

Milestones

  1. Lemma 1 (p. 5, Zinkevich 2003): projected online gradient ascent with step η=D/(GT)\eta=D/(G\sqrt T)η=D/(GT​) on concave rewards has regret ∑tft(x∗)−∑tft(xt)≤GDT\sum_t f_t(x^*)-\sum_t f_t(x_t)\le GD\sqrt T∑t​ft​(x∗)−∑t​ft​(xt​)≤GDT​ for every x∗x^*x∗ in the decision set.
  2. (6) (p. 7): if a point is returned, 1T∑t=1Tfi(xt,uit)≤ϵ\frac1T\sum_{t=1}^T f_i(x^t,u^t_i)\le\epsilonT1​∑t=1T​fi​(xt,uit​)≤ϵ for every iii.
  3. (7) (p. 8): for every iii and u∈Uu\in\mathcal Uu∈U, 1T∑tfi(xt,u)−1T∑tfi(xt,uit)≤GD/T≤ϵ\frac1T\sum_t f_i(x^t,u)-\frac1T\sum_t f_i(x^t,u^t_i)\le GD/\sqrt T\le\epsilonT1​∑t​fi​(xt,u)−T1​∑t​fi​(xt,uit​)≤GD/T​≤ϵ.
  4. Final inequality of the proof (p. 8): fi(xˉ,u)≤1T∑tfi(xt,u)f_i(\bar x,u)\le\frac1T\sum_t f_i(x^t,u)fi​(xˉ,u)≤T1​∑t​fi​(xt,u) for u∈Uu\in\mathcal Uu∈U.

Significance

The result. Theorem 3 turns any approximate solver of the nominal problem into an approximate solver of its robust counterpart, at a cost of ⌈G2D2/ϵ2⌉\lceil G^2D^2/\epsilon^2\rceil⌈G2D2/ϵ2⌉ solver calls, a number that depends on the geometry of U\mathcal UU and the sensitivity of the constraints to the noise but not on nnn, ddd or mmm. It is the prototype of the paper's oracle-based reductions: the same primal–dual template, with a different online learner, gives the dual-perturbation algorithm of §3.2–3.3 for non-convex uncertainty sets, and the applications of §4 (robust linear programs, quadratic programs, semidefinite programs) instantiate it.

Formalizing it. The theorem is proved in the paper; none of it is machine-checked. A formal development adds a checked statement of the reduction with an explicit call count in place of the paper's O(⋅)O(\cdot)O(⋅), and a reusable regret bound for projected online gradient ascent on concave rewards (Lemma 1), which the paper quotes from Zinkevich without proof and which many other online-learning results rest on.

Difficulty

The obvious argument for the dual side fails at one point: in round ttt the primal point xtx^txt is computed from utu^tut, so the reward fi(xt,⋅)f_i(x^t,\cdot)fi​(xt,⋅) that the dual player faces depends on its own current move. A regret bound that assumed rewards fixed in advance, or drawn independently of the learner's play, would not apply. Lemma 1 must be used in its adversarial form, valid for every sequence of reward functions, including adaptively chosen ones. A second point is that the projection step requires the variational characterization of a nearest point in a convex set, which a mere "map into U\mathcal UU" does not provide.

Formalization scope

Points are elements of EuclideanSpace ℝ (Fin k), so every norm is the ℓ2\ell_2ℓ2​ norm. The projection is a predicate IsProjOnto U P (each P(y)P(y)P(y) is a nearest point of U\mathcal UU to yyy), not a construction; the oracle is a function (Fin m → E d) → Option (E n) with none for "infeasible", constrained by the predicate IsApproxOracle on inputs in Um\mathcal U^mUm. The goal is quantified over every oracle meeting that specification. The gradient ∇ufi(x,u)\nabla_u f_i(x,u)∇u​fi​(x,u) is a given map gradU with HasGradientAt at points of U\mathcal UU; no differentiability in xxx is assumed. Rounds are indexed by natural numbers with index 000 for the initialization; the starting primal point x0∈Dx^0\in\mathcal Dx0∈D, used by the first update and left undefined by the algorithm, is an input. Hypotheses D>0D>0D>0 and G>0G>0G>0 are added so that η\etaη and T≥1T\ge1T≥1 are meaningful. Maxima over U\mathcal UU are stated as "for every u∈Uu\in\mathcal Uu∈U".

Explicit instantiations and corrections:

  • The paper's "O(G2D2/ϵ2)O(G^2D^2/\epsilon^2)O(G2D2/ϵ2) calls" is stated as at most ⌈G2D2/ϵ2⌉\lceil G^2D^2/\epsilon^2\rceil⌈G2D2/ϵ2⌉ calls (one call per round, TTT rounds).
  • Lemma 1's "G≥max⁡t∥ft(xt)∥G\ge\max_t\|f_t(x_t)\|G≥maxt​∥ft​(xt​)∥" is read as the gradient bound ∥∇ft(xt)∥≤G\|\nabla f_t(x_t)\|\le G∥∇ft​(xt​)∥≤G, as the same sentence describes it.
  • The proof's "Combining (10) and (12)" refers to (6) and (7).

Trivializing formalizations are ruled out: an oracle specification under which "infeasible" is never returned, or an output that is not the average of the oracle's answers, would not be Theorem 3. The "infeasible" conclusion is about the robust problem, not the nominal one.

A complete development needs the variational inequality for nearest points in a convex set, the gradient (supergradient) inequality for a concave function differentiable at a point of a convex set, Zinkevich's telescoping argument, and Jensen's inequality for finite averages. The first two and Lemma 1 are reusable beyond this mission. Proofs of the milestones, in any order, are welcome.

Selected references

  • A. Ben-Tal, E. Hazan, T. Koren, S. Mannor, Oracle-Based Robust Optimization via Online Learning, arXiv:1402.6361v1, 2014; Operations Research 63(3), 2015. https://arxiv.org/abs/1402.6361v1
  • M. Zinkevich, Online Convex Programming and Generalized Infinitesimal Gradient Ascent, ICML 2003. https://dl.acm.org/doi/10.5555/3041838.3041955
  • A. Ben-Tal, L. El Ghaoui, A. Nemirovski, Robust Optimization, Princeton University Press, 2009. https://doi.org/10.1515/9781400831050
  • E. Hazan, Introduction to Online Convex Optimization, Foundations and Trends in Optimization, 2016. https://arxiv.org/abs/1909.05207
8 thms1 active userReviewed
Operations ResearchOptimizationReinforcement Learning·Captain: mikedeng1

Twice Regularized MDPs and the Equivalence Between Robustness and Regularization 2: The Greedy Policy of the R2 Optimal Value Is the Unique Optimal R2 PolicyResearch Paper

Motivation

A robust Markov decision process (robust MDP) evaluates a policy against the worst transition kernel and reward in an uncertainty set around a nominal model (P0,r0)(P_0, r_0)(P0​,r0​). It is the standard model for planning when the dynamics are estimated from data (Iyengar 2005; Nilim and El Ghaoui 2005; Wiesemann, Kuhn and Rustem 2013). Its Bellman update contains an inner optimization over the uncertainty set at every state, which makes robust planning costly when the sets are not (s,a)(s,a)(s,a)-rectangular.

Derman, Geist and Mannor (arXiv:2110.06267, NeurIPS 2021) show that, for sss-rectangular ball uncertainty sets, this inner optimization can be replaced by an explicit penalty that depends both on the policy and on the value function. The resulting twice regularized (R²) MDPs have Bellman operators with no inner optimization over models. The first mission of this series formalizes the robust–regularized equivalence (Theorem 4.1 of the paper). This mission formalizes Section 5: the R² Bellman operators are monotone and contracting under a bound on the transition radius, and the greedy policy of the R² optimal value is optimal.

Setting

Let S\mathcal SS and A\mathcal AA be finite nonempty sets of states and actions, γ∈(0,1)\gamma\in(0,1)γ∈(0,1) a discount factor, P0(s′∣s,a)P_0(s'\mid s,a)P0​(s′∣s,a) a transition kernel and r0(s,a)r_0(s,a)r0​(s,a) a reward. A policy π∈ΔAS\pi\in\Delta_{\mathcal A}^{\mathcal S}π∈ΔAS​ assigns to each state a probability distribution πs\pi_sπs​ on A\mathcal AA. For v∈RSv\in\mathbb R^{\mathcal S}v∈RS write qs(a)=r0(s,a)+γ∑s′P0(s′∣s,a)v(s′)q_s(a)=r_0(s,a)+\gamma\sum_{s'}P_0(s'\mid s,a)v(s')qs​(a)=r0​(s,a)+γ∑s′​P0​(s′∣s,a)v(s′) and

[T(P0,r0)πv](s)=∑aπs(a) qs(a).[T^\pi_{(P_0,r_0)}v](s)=\sum_a\pi_s(a)\,q_s(a).[T(P0​,r0​)π​v](s)=a∑​πs​(a)qs​(a).

All norms ∥⋅∥\|\cdot\|∥⋅∥ below are ℓ2\ell_2ℓ2​-norms, ∥a∥=(∑za(z)2)1/2\|a\|=\big(\sum_z a(z)^2\big)^{1/2}∥a∥=(∑z​a(z)2)1/2; ∥⋅∥∞\|\cdot\|_\infty∥⋅∥∞​ is the sup norm.

Fix nonnegative radii αsr,αsP\alpha^r_s,\alpha^P_sαsr​,αsP​ for each state. The R² regularizer is Ωv,R2(πs)=∥πs∥ (αsr+αsPγ∥v∥)\Omega_{v,\mathrm R^2}(\pi_s)=\|\pi_s\|\,(\alpha^r_s+\alpha^P_s\gamma\|v\|)Ωv,R2​(πs​)=∥πs​∥(αsr​+αsP​γ∥v∥), and the R² Bellman operators are

[Tπ,R2v](s)=[T(P0,r0)πv](s)−Ωv,R2(πs),[T∗,R2v](s)=max⁡π∈ΔAS[Tπ,R2v](s).[T^{\pi,\mathrm R^2}v](s)=[T^\pi_{(P_0,r_0)}v](s)-\Omega_{v,\mathrm R^2}(\pi_s),\qquad [T^{*,\mathrm R^2}v](s)=\max_{\pi\in\Delta^{\mathcal S}_{\mathcal A}}[T^{\pi,\mathrm R^2}v](s).[Tπ,R2v](s)=[T(P0​,r0​)π​v](s)−Ωv,R2​(πs​),[T∗,R2v](s)=π∈ΔAS​max​[Tπ,R2v](s).

A policy π\piπ is greedy for vvv when Tπ,R2v=T∗,R2vT^{\pi,\mathrm R^2}v=T^{*,\mathrm R^2}vTπ,R2v=T∗,R2v.

Assumption 5.1 (bounded radius). For each sss there is ϵs>0\epsilon_s>0ϵs​>0 with

αsP≤min⁡(1−γ−ϵsγ∣S∣ ; min⁡u∈R+A,∥u∥=1, w∈R+S,∥w∥=1 ∑a,s′u(a)P0(s′∣s,a)w(s′)),\alpha^P_s\le\min\Big(\frac{1-\gamma-\epsilon_s}{\gamma\sqrt{|\mathcal S|}}\ ;\ \min_{u\in\mathbb R^{\mathcal A}_+,\|u\|=1,\ w\in\mathbb R^{\mathcal S}_+,\|w\|=1}\ \sum_{a,s'}u(a)P_0(s'\mid s,a)w(s')\Big),αsP​≤min(γ∣S∣​1−γ−ϵs​​ ; u∈R+A​,∥u∥=1, w∈R+S​,∥w∥=1min​ a,s′∑​u(a)P0​(s′∣s,a)w(s′)),

and ϵ∗=min⁡sϵs\epsilon_*=\min_s\epsilon_sϵ∗​=mins​ϵs​. The R² value function vπ,R2v^{\pi,\mathrm R^2}vπ,R2 of a policy and the R² optimal value v∗,R2v^{*,\mathrm R^2}v∗,R2 are the fixed points of Tπ,R2T^{\pi,\mathrm R^2}Tπ,R2 and T∗,R2T^{*,\mathrm R^2}T∗,R2.

Formalization targets

Goal: Theorem 5.1 (p. 8)

Under Assumption 5.1, T∗,R2T^{*,\mathrm R^2}T∗,R2 and every Tπ,R2T^{\pi,\mathrm R^2}Tπ,R2 have unique fixed points; a greedy policy π∗,R2\pi^{*,\mathrm R^2}π∗,R2 for v∗,R2v^{*,\mathrm R^2}v∗,R2 exists, and every such policy satisfies

vπ∗,R2,R2=v∗,R2 ≥ vπ,R2for all π∈ΔAS;v^{\pi^{*,\mathrm R^2},\mathrm R^2}=v^{*,\mathrm R^2}\ \ge\ v^{\pi,\mathrm R^2}\qquad\text{for all }\pi\in\Delta^{\mathcal S}_{\mathcal A};vπ∗,R2,R2=v∗,R2 ≥ vπ,R2for all π∈ΔAS​;

every optimal policy is greedy; and when αsr>0\alpha^r_s>0αsr​>0 for all sss the greedy policy is unique, hence the unique optimal R² policy.

Milestones

  1. Proposition 2.1 (p. 3): for Ω\OmegaΩ strongly convex on the simplex, Ω∗(y)=max⁡a∈Δ⟨a,y⟩−Ω(a)\Omega^*(y)=\max_{a\in\Delta}\langle a,y\rangle-\Omega(a)Ω∗(y)=maxa∈Δ​⟨a,y⟩−Ω(a) is differentiable with Lipschitz gradient equal to the unique maximizer, satisfies Ω∗(y+c1)=Ω∗(y)+c\Omega^*(y+c\mathbb 1)=\Omega^*(y)+cΩ∗(y+c1)=Ω∗(y)+c, and is non-decreasing.
  2. Proposition 5.1 (i) (p. 8): v1≤v2v_1\le v_2v1​≤v2​ implies Tπ,R2v1≤Tπ,R2v2T^{\pi,\mathrm R^2}v_1\le T^{\pi,\mathrm R^2}v_2Tπ,R2v1​≤Tπ,R2v2​ and T∗,R2v1≤T∗,R2v2T^{*,\mathrm R^2}v_1\le T^{*,\mathrm R^2}v_2T∗,R2v1​≤T∗,R2v2​.
  3. Proposition 5.1 (iii) (p. 8):
∥Tπ,R2v1−Tπ,R2v2∥∞≤(1−ϵ∗)∥v1−v2∥∞,∥T∗,R2v1−T∗,R2v2∥∞≤(1−ϵ∗)∥v1−v2∥∞.\|T^{\pi,\mathrm R^2}v_1-T^{\pi,\mathrm R^2}v_2\|_\infty\le(1-\epsilon_*)\|v_1-v_2\|_\infty,\qquad \|T^{*,\mathrm R^2}v_1-T^{*,\mathrm R^2}v_2\|_\infty\le(1-\epsilon_*)\|v_1-v_2\|_\infty.∥Tπ,R2v1​−Tπ,R2v2​∥∞​≤(1−ϵ∗​)∥v1​−v2​∥∞​,∥T∗,R2v1​−T∗,R2v2​∥∞​≤(1−ϵ∗​)∥v1​−v2​∥∞​.

Significance

Theorem 5.1 is the R² counterpart of the fundamental theorem of discounted dynamic programming: optimal R² values are achieved by stationary policies obtained by a single greedy step. Together with the contraction of Proposition 5.1 (iii) it justifies the R² modified policy iteration algorithm of the paper, whose greedy step is a projection onto the simplex rather than a robust max–min problem. Combined with the first mission of the series, which identifies the robust value of an sss-rectangular ball-constrained MDP with the optimum of an R²-regularized program, it gives a route to robust planning at the cost of regularized planning.

The results are proved in the paper (App. C), partly by reference to Geist, Scherrer and Pietquin (2019) for the optimality operator. No machine-checked proof of any of them exists; this mission produces the first. Prop. 2.1 is a general fact of convex analysis (Danskin-type smoothness of a conjugate on the simplex) that is reusable for any regularized MDP or entropy-regularized game.

Difficulty

The R² evaluation operator is not affine: the value regularizer −αsPγ∥πs∥ ∥v∥-\alpha^P_s\gamma\|\pi_s\|\,\|v\|−αsP​γ∥πs​∥∥v∥ is concave in vvv and decreases as ∥v∥\|v\|∥v∥ grows. Monotonicity therefore does not follow from the positivity of P0P_0P0​ as in the standard case; it requires the second bound of Assumption 5.1, which compares the ℓ2\ell_2ℓ2​ variation of ∥v∥\|v\|∥v∥ with the minimal nonnegative bilinear form of P0(⋅∣s,⋅)P_0(\cdot\mid s,\cdot)P0​(⋅∣s,⋅). Likewise the contraction modulus is not γ\gammaγ but 1−ϵ∗1-\epsilon_*1−ϵ∗​, because the regularizer is ∣S∣\sqrt{|\mathcal S|}∣S∣​-Lipschitz between the ℓ2\ell_2ℓ2​ and sup norms. The optimality step of the classical proof uses linearity of TπT^\piTπ when comparing values of policies; here only monotonicity and contraction are available. Uniqueness of the greedy policy rests on strict concavity on the simplex, which holds only when the regularization weight is positive.

Formalization scope

States and actions are finite nonempty types; transitions are arrays P₀ : S → A → S → ℝ with the published predicate IsTransitionKernel; value functions are S → ℝ with the pointwise order. The ℓ2\ell_2ℓ2​-norm is an explicit l2norm (Mathlib's norm on S → ℝ is the sup norm, used only for ∥⋅∥∞\|\cdot\|_\infty∥⋅∥∞​). T∗,R2v(s)T^{*,\mathrm R^2}v(s)T∗,R2v(s) is the real supremum over the simplex ΔA\Delta_{\mathcal A}ΔA​ (attained), and the inner minimum of Assumption 5.1 is the real infimum over nonnegative ℓ2\ell_2ℓ2​-unit vectors; the witnesses ϵs\epsilon_sϵs​ are explicit. Greedy policies are a predicate, never a function, and the R² value functions are not defined by choice: the goal asserts their existence and uniqueness and speaks about the fixed points.

Disclosed deviations from the page. Assumption 5.1 is a hypothesis of Theorem 5.1 (its proof assumes it). The uniqueness clause of Theorem 5.1 additionally assumes αsr>0\alpha^r_s>0αsr​>0 for all sss: with one state, two actions, zero reward and zero radii every policy is greedy and optimal. Proposition 2.1 assumes Ω\OmegaΩ continuous on the simplex, without which the maximum need not be attained, and strong convexity is Mathlib's StrongConvexOn for some modulus (norm-independent in finite dimension). Proposition 5.1 (ii) is false as printed and is not drafted: with one state, one action, P0=1P_0=1P0​=1, r0=0r_0=0r0​=0, γ=1/2\gamma=1/2γ=1/2, αr=0\alpha^r=0αr=0, αP=1/2\alpha^P=1/2αP=1/2, ϵ=1/4\epsilon=1/4ϵ=1/4, one has Tv=v/2−∣v∣/4Tv=v/2-|v|/4Tv=v/2−∣v∣/4, and v1=−1v_1=-1v1​=−1, c=1c=1c=1 give T(v1+c)=0>−1/4=Tv1+γcT(v_1+c)=0>-1/4=Tv_1+\gamma cT(v1​+c)=0>−1/4=Tv1​+γc. Remark 5.1, Algorithm 1 and the ℓp\ell_pℓp​ variant of App. C.1 are out of scope. The inner minimum of Assumption 5.1 is 000 whenever some P0(s′∣s,a)=0P_0(s'\mid s,a)=0P0​(s′∣s,a)=0, forcing αsP=0\alpha^P_s=0αsP​=0; this is the assumption as printed.

A formalization in which ∥⋅∥\|\cdot\|∥⋅∥ is the sup norm, the inner minimum ranges over all unit vectors (making the assumption unsatisfiable), or the value functions are postulated rather than shown to exist would be trivial or wrong; the drafted statements avoid all three. Contributions welcome: Prop. 2.1 as a general convex-analysis lemma, Banach fixed-point plumbing for S → ℝ with the sup norm, and the strict concavity of p↦⟨p,q⟩−c∥p∥p\mapsto\langle p,q\rangle-c\|p\|p↦⟨p,q⟩−c∥p∥ on the simplex.

Selected references

  • E. Derman, M. Geist, S. Mannor, Twice regularized MDPs and the equivalence between robustness and regularization, NeurIPS 2021. arXiv:2110.06267v1
  • M. Geist, B. Scherrer, O. Pietquin, A theory of regularized Markov decision processes, ICML 2019. arXiv:1901.11275
  • A. Nilim, L. El Ghaoui, Robust control of Markov decision processes with uncertain transition matrices, Operations Research 53(5), 2005. doi:10.1287/opre.1050.0216
  • G. N. Iyengar, Robust dynamic programming, Mathematics of Operations Research 30(2), 2005. doi:10.1287/moor.1040.0129
  • W. Wiesemann, D. Kuhn, B. Rustem, Robust Markov decision processes, Mathematics of Operations Research 38(1), 2013. doi:10.1287/moor.1120.0566
  • A. Mensch, M. Blondel, Differentiable dynamic programming for structured prediction and attention, ICML 2018. arXiv:1802.03676
9 thms1 active userReviewed
Bandit AlgorithmsOperations ResearchProbability·Captain: mikedeng1

The Best of Both Worlds: Stochastic and Adversarial Bandits: SAO Has Pseudo-Regret O(K log K log²β/Δ) on Stochastic Rewards and Regret Õ(√(nK)) Against Adaptive AdversariesResearch Paper

Motivation

In a multi-armed bandit problem a learner chooses one of KKK actions in each of nnn rounds and observes only the reward of the chosen action. Two models of the rewards have separate theories. In the stochastic model, each arm pays independent draws from a fixed distribution; algorithms such as UCB1 (Auer, Cesa-Bianchi & Fischer 2002) have regret of order ∑ilog⁡(n)/Δi\sum_i \log(n)/\Delta_i∑i​log(n)/Δi​, logarithmic in nnn. In the adversarial model, an adversary chooses the rewards; Exp3 and its variants (Auer, Cesa-Bianchi, Freund & Schapire 2002) have regret of order nK\sqrt{nK}nK​, which is optimal there. An algorithm tuned for one model fails in the other: stochastic algorithms can suffer linear regret against an adversary, and adversarial algorithms pay n\sqrt nn​ even when the rewards are i.i.d.

Bubeck and Slivkins (arXiv:1202.4473, COLT 2012) asked whether one algorithm can be near-optimal in both models without knowing which one it faces. They answered yes with the algorithm SAO. This result started the "best of both worlds" line of work on bandits. Later contributions include EXP3++ (Seldin & Slivkins 2014) and Tsallis-INF (Zimmert & Seldin 2021).

Setting

There are K≥2K\ge2K≥2 arms and n≥Kn\ge Kn≥K rounds. On round ttt the algorithm draws an arm ItI_tIt​ from a probability vector pt=(p1,t,…,pK,t)p_t=(p_{1,t},\dots,p_{K,t})pt​=(p1,t​,…,pK,t​) computed from the history it has observed. At the same time a reward vector gt∈[0,1]Kg_t\in[0,1]^Kgt​∈[0,1]K is fixed, and the algorithm observes only gIt,tg_{I_t,t}gIt​,t​.

  • Adversarial model. The vector gtg_tgt​ is chosen by an adaptive adversary: a function of the arms I1,…,It−1I_1,\dots,I_{t-1}I1​,…,It−1​ played earlier, but not of ItI_tIt​. The regret is Rn=max⁡i∑t=1ngi,t−∑t=1ngIt,tR_n=\max_i\sum_{t=1}^n g_{i,t}-\sum_{t=1}^n g_{I_t,t}Rn​=maxi​∑t=1n​gi,t​−∑t=1n​gIt​,t​.
  • Stochastic model. There are distributions ν1,…,νK\nu_1,\dots,\nu_Kν1​,…,νK​ on [0,1][0,1][0,1] with means μi\mu_iμi​, and all gi,t∼νig_{i,t}\sim\nu_igi,t​∼νi​ are independent. The pseudo-regret is R‾n=∑t=1n(max⁡iμi−μIt)\overline R_n=\sum_{t=1}^n(\max_i\mu_i-\mu_{I_t})Rn​=∑t=1n​(maxi​μi​−μIt​​). The gap of arm iii is Δi=max⁡jμj−μi\Delta_i=\max_j\mu_j-\mu_iΔi​=maxj​μj​−μi​, and the minimal gap is Δ=min⁡i:Δi>0Δi\Delta=\min_{i:\Delta_i>0}\Delta_iΔ=mini:Δi​>0​Δi​.

The analysis uses importance-weighted estimates H~i,t=1t∑s≤tgi,s1{Is=i}/pi,s\widetilde H_{i,t}=\frac1t\sum_{s\le t}g_{i,s}\mathbb 1_{\{I_s=i\}}/p_{i,s}Hi,t​=t1​∑s≤t​gi,s​1{Is​=i}​/pi,s​, the sample means H^i,t\widehat H_{i,t}Hi,t​, the averages Hi,t=1t∑s≤tgi,sH_{i,t}=\frac1t\sum_{s\le t}g_{i,s}Hi,t​=t1​∑s≤t​gi,s​, and the play counts Ti(t)T_i(t)Ti​(t).

SAO (Algorithm 1 of the paper) takes a parameter β>1\beta>1β>1. It keeps a set of active arms, initially all arms, and samples them uniformly at first. On each round it applies a test, (12), that deactivates an arm whose estimate H~i,t\widetilde H_{i,t}Hi,t​ falls far below the best active one. The probability of a deactivated arm then decays as qiτi/tq_i\tau_i/tqi​τi​/t, where τi\tau_iτi​ is the deactivation time and qiq_iqi​ the arm's probability at that moment. Three further tests, (13)–(15), check that the observations stay consistent with stochastic rewards. If any of them fails on round τ0\tau_0τ0​, SAO switches permanently to the adversarial algorithm Exp3.P (Bubeck & Cesa-Bianchi 2012, Fig. 3.1) for the remaining rounds.

Formalization targets

Goal: Theorem 4.1, high-probability form

For every δ∈(0,1)\delta\in(0,1)δ∈(0,1) let β=10Kn3δ−1\beta=10Kn^3\delta^{-1}β=10Kn3δ−1. With probability at least 1−δ1-\delta1−δ, SAO with parameter β\betaβ satisfies, in the stochastic model (whenever some arm has Δi>0\Delta_i>0Δi​>0),

R‾n≤260K(1+log⁡K)log⁡2(β)Δ,\overline R_n\le\frac{260K(1+\log K)\log^2(\beta)}{\Delta},Rn​≤Δ260K(1+logK)log2(β)​,

and, against every adaptive adversary with rewards in [0,1][0,1][0,1],

Rn≤60(1+log⁡K)(1+log⁡n)nKlog⁡(β)+5K2log⁡2(β)+200K2log⁡2(β).R_n\le60(1+\log K)(1+\log n)\sqrt{nK\log(\beta)+5K^2\log^2(\beta)}+200K^2\log^2(\beta).Rn​≤60(1+logK)(1+logn)nKlog(β)+5K2log2(β)​+200K2log2(β).

Milestones

The milestones follow the paper's proof in order:

  • Freedman's inequality (Theorem 4.3) in the paper's two-sided form, and its variance-adaptive form, Lemma 4.4.
  • The concentration lemmas for SAO's estimates (Lemmas 4.5, 4.6, 4.7) and the Exp3.P phase (Lemma 4.8).
  • The two good events of §4.1, (21)–(25).
  • The deterministic consequences on those events: Exp3.P is never started in the stochastic model; suboptimal arms are deactivated by time 260Klog⁡(β)/Δi2260K\log(\beta)/\Delta_i^2260Klog(β)/Δi2​; ∑iqi≤1+log⁡K\sum_iq_i\le1+\log K∑i​qi​≤1+logK, (27); and the adversarial regret bound of §4.3.
  • The two halves of Theorem 4.1.

Significance

The theorem shows that the stochastic and adversarial regret rates are not in conflict. A single algorithm, with no information about the model, gets O(Klog⁡Klog⁡2(n/δ)/Δ)O(K\log K\log^2(n/\delta)/\Delta)O(KlogKlog2(n/δ)/Δ) pseudo-regret on stochastic rewards and O~(nK)\tilde O(\sqrt{nK})O~(nK​) regret against adaptive adversaries. Each rate is within polylogarithmic factors of optimal for its model. Later algorithms improved the logarithmic factors and removed the explicit switching, but they are compared against this result.

The theorem is proved in the paper. It is not known to have a machine-checked proof. Formalizing it requires a precise model of an adaptive adversary interacting with a randomized algorithm, martingale concentration with random variance (Lemma 4.4), and an exact statement of SAO including its boundary cases. The pieces are reusable: the interaction model, the estimators, Exp3.P and its high-probability guarantee all apply to other adversarial bandit results.

Difficulty

Neither standard analysis carries over. In the stochastic model, SAO's sampling probabilities are random and depend on the past, and a deactivated arm's probability keeps changing. Hoeffding-type bounds for a fixed sampling scheme therefore do not apply to H~i,t\widetilde H_{i,t}Hi,t​. The variance of the importance-weighted estimate grows like ∑s1/pi,s\sum_s1/p_{i,s}∑s​1/pi,s​, which is controlled only through the algorithm's own schedule (16). This is why Lemma 4.5 has the two-part radius with max⁡(t−τi,0)/(qiτit)\max(t-\tau_i,0)/(q_i\tau_it)max(t−τi​,0)/(qi​τi​t). In the adversarial model, the deterministic argument has to show that whenever the consistency tests pass, the regret accumulated before the switch is already small, for an adversary that adapts to the arms played. A union bound over all quantities, all arms and all times (§4.1) is needed before any deterministic reasoning, so every constant in the event matters.

Formalization scope

All declarations live in the namespace BestBothWorlds.SAO.

  • Arms and paths. Arms are Fin K and rounds are 1,…,n1,\dots,n1,…,n. An arm path is Fin n → Fin K.
  • Algorithms and adversaries. An algorithm is a deterministic map from the observed history to a probability vector. A deterministic adaptive adversary is a map from the list of earlier arms to a reward vector; randomized adversaries are mixtures of these.
  • Probabilities. For a fixed adversary, the probability of an event is ∑I∈E∏tpIt,t\sum_{I\in E}\prod_tp_{I_t,t}∑I∈E​∏t​pIt​,t​. In the stochastic model this is integrated against the product law of the reward table.
  • Logarithms and constants. Real.log is the natural logarithm. All constants of Theorem 4.1 are explicit, with β=10Kn3δ−1\beta=10Kn^3\delta^{-1}β=10Kn3δ−1.
  • SAO. It is defined exactly as Algorithm 1. Arms are tested in order within a round, and the active set changes during the loop. Test (13) is false when Ti(t)=0T_i(t)=0Ti​(t)=0, and test (14) is false when τi=1\tau_i=1τi​=1.
  • Exp3.P. After the switch, Exp3.P runs from scratch for n−τ0n-\tau_0n−τ0​ rounds. Its parameters are those of Bubeck–Cesa-Bianchi Theorem 3.2 with confidence K/βK/\betaK/β, and γ\gammaγ and βP\beta_{\mathrm P}βP​ are clipped at 111.
  • §4 notation. τ0\tau_0τ0​, τi←min⁡(τi,τ0)\tau_i\leftarrow\min(\tau_i,\tau_0)τi​←min(τi​,τ0​) and qi=pi,min⁡(τi,τ0)q_i=p_{i,\min(\tau_i,\tau_0)}qi​=pi,min(τi​,τ0​)​ are computed from the run, never assumed.

A trivializing formalization is ruled out. The goal's hypotheses concern only the instance (KKK, nnn, δ\deltaδ, the distributions or the adversary). The algorithm's quantities (τ0\tau_0τ0​, τi\tau_iτi​, qiq_iqi​, the sampling probabilities) are computed by the definition of SAO and are never free variables or hypotheses. The adversarial half covers adaptive adversaries, not only oblivious reward tables.

The expectation form of Theorem 4.1 (O(⋅)O(\cdot)O(⋅) bounds with β=n4\beta=n^4β=n4), Theorem 1.1 and the two-armed warm-up of §3 are out of scope. Proofs of any milestone are welcome, as are alternative proofs of the concentration lemmas from Mathlib's martingale library.

Selected references

  • S. Bubeck and A. Slivkins, The best of both worlds: stochastic and adversarial bandits, COLT 2012; arXiv:1202.4473v1. https://arxiv.org/abs/1202.4473
  • S. Bubeck and N. Cesa-Bianchi, Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems, Foundations and Trends in Machine Learning 5(1), 2012. https://doi.org/10.1561/2200000024
  • D. A. Freedman, On tail probabilities for martingales, Annals of Probability 3(1), 1975. https://doi.org/10.1214/aop/1176996452
  • P. Auer, N. Cesa-Bianchi and P. Fischer, Finite-time analysis of the multiarmed bandit problem, Machine Learning 47, 2002. https://doi.org/10.1023/A:1013689704352
  • P. Auer, N. Cesa-Bianchi, Y. Freund and R. E. Schapire, The nonstochastic multiarmed bandit problem, SIAM Journal on Computing 32(1), 2002. https://doi.org/10.1137/S0097539701398375
  • Y. Seldin and A. Slivkins, One practical algorithm for both stochastic and adversarial bandits, ICML 2014. https://proceedings.mlr.press/v32/seldinb14.html
  • J. Zimmert and Y. Seldin, Tsallis-INF: an optimal algorithm for stochastic and adversarial bandits, JMLR 22, 2021. https://jmlr.org/papers/v22/19-753.html
20 thms1 active userReviewed
ProbabilityStatistics·Captain: mikedeng1

The Sample Complexity of Pattern Classification with Neural Networks: The Size of the Weights is More Important than the Size of the Network III: Sigmoid Networks with Small Weights GeneralizeResearch Paper

Motivation

A classifier built from a neural network produces a real score and predicts a binary label from its sign. A network can have many hidden units, so a guarantee based only on the number of parameters can be uninformative even when its output weights are small. Bartlett's 1998 paper asks whether a classifier's margin on training examples and the total magnitude of its weights can control its probability of error without fixing the number of units. Its Theorem 28 gives such a statement for two-layer networks whose activation is bounded and nondecreasing. The paper also discusses why this parameter-magnitude view supports weight decay and early stopping as learning heuristics, while leaving their algorithmic behavior outside the theorem's scope (Bartlett 1998, pp. 526, 534–535).

The theorem combines two results in the same paper. Theorem 2 turns the fat-shattering dimension of a real-valued function class into a margin generalization bound. Corollary 24 controls that dimension for finite combinations of affine-input units when the sum of the absolute combination weights is bounded. Lemmas 19, 22, and 23 supply covering estimates along that path. These are the milestones of this mission, with the source statements preserved in the milestone record (Bartlett 1998, pp. 527, 532–534).

Setting

An input is a vector x∈Rnx\in\mathbb R^nx∈Rn, represented in Lean as Fin n → ℝ. A label is y∈{−1,+1}y\in\{-1,+1\}y∈{−1,+1}; Lean's Bool is converted by pm, where true means +1+1+1. A probability distribution PPP lives on labeled inputs. From an independent sample z=((xi,yi))i=1mz=((x_i,y_i))_{i=1}^mz=((xi​,yi​))i=1m​, the empirical margin error at scale γ>0\gamma>0γ>0 is the fraction of indices with yih(xi)<γy_i h(x_i)<\gammayi​h(xi​)<γ. The population error is the probability that sgn⁡(h(x))≠y\operatorname{sgn}(h(x))\ne ysgn(h(x))=y, where sgn⁡(0)=+1\operatorname{sgn}(0)=+1sgn(0)=+1. The inequality in the empirical error is strict, as in the paper's definition (Bartlett 1998, p. 526).

Fix a bounded nondecreasing activation σ:R→[−1,1]\sigma:\mathbb R\to[-1,1]σ:R→[−1,1]. The first-layer class FFF contains every x↦σ(w⋅x+w0)x\mapsto\sigma(w\cdot x+w_0)x↦σ(w⋅x+w0​), with an arbitrary weight vector and bias. The network class HHH contains all finite sums ∑i=1Nαifi\sum_{i=1}^N\alpha_i f_i∑i=1N​αi​fi​ with fi∈Ff_i\in Ffi​∈F and ∑i∣αi∣≤A\sum_i|\alpha_i|\le A∑i​∣αi​∣≤A. Thus AAA bounds the output layer's total weight magnitude, while NNN can vary without an imposed width limit. The bias w0w_0w0​ is part of every unit. For a class GGG, fat⁡G(η)\operatorname{fat}_G(\eta)fatG​(η) records the largest length of an input sequence whose every sign pattern can be realized with separation at least η\etaη around one vector of thresholds (Bartlett 1998, pp. 526, 533–534).

Formalization targets

Two-layer generalization

For 0<γ≤10<\gamma\le10<γ≤1, 0<δ<1/20<\delta<1/20<δ<1/2, A≥1A\ge1A≥1, and an independent sample of length m≥1m\ge1m≥1, the goal is one universal c>0c>0c>0 such that, with probability at least 1−δ1-\delta1−δ, every h∈Hh\in Hh∈H satisfies

er⁡P(h)<er⁡^zγ(h)+cm(A2nγ2log⁡ ⁣(32Aγ)(log⁡m)2+log⁡ ⁣(1δ)).\operatorname{er}_P(h)<\widehat{\operatorname{er}}_z^\gamma(h)+ \sqrt{\frac{c}{m}\left( \frac{A^2n}{\gamma^2}\log\!\left(\frac{32A}{\gamma}\right)(\log m)^2+ \log\!\left(\frac1\delta\right)\right)}.erP​(h)<erzγ​(h)+mc​(γ2A2n​log(γ32A​)(logm)2+log(δ1​))​.

The paper prints log⁡(A/γ)\log(A/\gamma)log(A/γ) in this display. That term vanishes at A=γ=1A=\gamma=1A=γ=1, although the class can then contain halfspace classifiers with nonzero sample complexity. The proof obtains a positive factor at that corner through Corollary 24 at scale γ/16\gamma/16γ/16, giving log⁡(32A/γ)\log(32A/\gamma)log(32A/γ). The goal states this correction and records the printed statement separately in the moderation notes. The constant precedes all network, distribution, margin, confidence, and sample parameters in Lean; it cannot be selected after observing the instance (Bartlett 1998, pp. 533–534).

Capacity and margin milestones

Corollary 24 bounds fat⁡H(η)\operatorname{fat}_H(\eta)fatH​(η) by a constant multiple of M2A2nη−2log⁡(MA/η)M^2A^2n\eta^{-2}\log(MA/\eta)M2A2nη−2log(MA/η) when the activation has range [−M/2,M/2][-M/2,M/2][−M/2,M/2]. Theorem 2 then converts a finite fat dimension at scale γ/16\gamma/16γ/16 into a simultaneous bound on population error for all members of HHH. The three covering lemmas track how shattering, pseudodimension, and an ℓ1\ell_1ℓ1​ weight budget affect covers in sample ℓ1\ell_1ℓ1​, ℓ∞\ell_\inftyℓ∞​, and ℓ2\ell_2ℓ2​ distances. Each bound retains the scale and explicit constants printed by the paper, subject to the stated corrections to undefined or false boundary cases (Bartlett 1998, pp. 527, 532–533).

Significance

The goal gives a width-independent generalization guarantee for a chosen network when its empirical margin error and total output weight are small. It applies to the entire class HHH at once, so choosing a network after inspecting the sample does not turn the bound into a claim about only one fixed predictor. It does not assert that a learning algorithm finds such a network or that the displayed constants are optimal. Bartlett notes that later work had improved a logarithmic factor, and that empirical agreement with neural-network performance remained an open experimental question at the time (Bartlett 1998, pp. 534–535).

The paper proves the mathematical result. This mission asks for a machine-checked proof of its corrected formal statement and the stated supporting results; the draft theorem files currently contain proof obligations. A completed development would also make the fat dimension and strict external sample-cover definitions available for other margin analyses. Those objects differ from the platform's fixed-architecture neural networks and closed-ball covering numbers, so they are defined here with the conventions of this paper.

Difficulty

Counting hidden units gives no finite width-independent capacity bound, because HHH permits arbitrarily many terms. Bounding each unit separately also does not control the full combination class: different small contributions can produce distinct values on a sample. The challenging step is relating covers of the base class to covers of all finite combinations under the total absolute-weight constraint, and then relating those covers back to fat-shattering. Even once a finite capacity estimate is available, the probability statement must hold simultaneously for every h∈Hh\in Hh∈H, including a network selected after sampling (Bartlett 1998, pp. 532–534).

Formalization scope

Lean uses N∪{∞}\mathbb N\cup\{\infty\}N∪{∞} for fat dimensions and covering numbers, so an unbounded class cannot acquire a spurious dimension zero. Covers are external finite sets of real functions and use the strict distance <ε<\varepsilon<ε of Definition 3. Sample ℓ1\ell_1ℓ1​ and ℓ2\ell_2ℓ2​ distances are normalized by mmm. Pseudodimension is the supremum of positive-scale fat dimensions, matching the paper's right limit. Theorems assume m≥1m\ge1m≥1, and sample indices are zero-based. The network class is generated from its weights rather than supplied as an arbitrary set satisfying the desired bound.

The paper says it ignores measurability issues and assumes all sets considered are measurable (Bartlett 1998, p. 526). The goal makes the event of a violating network measurable. Its individual network functions are measurable from monotonicity of σ\sigmaσ and finite sums; the restated Theorem 2 has explicit hypotheses for measurable class members, the violating event, and the double-sample event in its proof. Theorem 2 additionally restricts d=fat⁡H(γ/16)d=\operatorname{fat}_H(\gamma/16)d=fatH​(γ/16) to d≤34md\le34md≤34m, where its printed logarithmic bound remains valid. Corollary 24 uses n≥1n\ge1n≥1 because a zero-dimensional input still permits a biased constant unit. Lemma 23 uses d≥1d\ge1d≥1 and 0<γ<emM/d0<\gamma<emM/d0<γ<emM/d in place of the printed γ≥0\gamma\ge0γ≥0: at γ=0\gamma=0γ=0 or d=0d=0d=0, or for γ≥emM/d\gamma\ge emM/dγ≥emM/d, the printed strict inequality fails, while on the rest of the printed range it is kept. These are recorded as corrections rather than attributed to the printed wording.

The deeper-network part of Theorem 28 is outside this proposal. Its printed chain through Corollary 27 has an unresolved range issue when the input box bound BBB is smaller than the activation range, and the displayed log⁡n\log nlogn factor also vanishes at n=1n=1n=1. This mission's goal is Part 1 and uses none of those claims. Contributions that establish the corrected covering lemmas, the capacity corollary, or the simultaneous margin bound fit the present proof frontier (Bartlett 1998, pp. 533–534).

Selected references

  • P. L. Bartlett, The Sample Complexity of Pattern Classification with Neural Networks: The Size of the Weights is More Important than the Size of the Network, IEEE Transactions on Information Theory 44(2), 525–536, 1998. DOI.
15 thms1 active userReviewed
OptimizationProbabilityStatistics·Captain: mikedeng1

Variance-based Regularization with Convex Objectives II: A Covering-Number Certificate and Oracle Inequality for the Robust MinimizerResearch Paper

Motivation

Empirical risk minimization (ERM) chooses, from a class F\mathcal FF of loss functions, the one with the smallest average loss on a sample X1,…,XnX_1,\dots,X_nX1​,…,Xn​. Its standard guarantees bound the excess population risk by a term of order 1/n1/\sqrt n1/n​, whatever the variance of the losses. When good functions in F\mathcal FF have small variance, a better trade-off is available in principle: minimize the empirical risk plus a standard-deviation penalty 2ρ VarP^n(f)/n\sqrt{2\rho\,\mathrm{Var}_{\widehat P_n}(f)/n}2ρVarPn​​(f)/n​. Maurer and Pontil (COLT 2009) showed that this sample variance penalization enjoys faster rates, but the penalized objective is non-convex even for convex losses, so it cannot be minimized efficiently in general.

Duchi and Namkoong (arXiv:1610.02581v3, 2017; NIPS 2017) replace the variance penalty by a distributionally robust objective: the worst-case average loss over all reweightings of the sample within a χ2\chi^2χ2-divergence ball of radius ρ/n\rho/nρ/n. This objective is convex whenever the loss is convex, and it equals the empirical risk plus the standard-deviation penalty up to an error of order 1/n1/n1/n. Theorem 3 of the paper turns this into a guarantee for the minimizer of the robust objective, using covering numbers of the class. This mission formalizes Theorem 3 and the lemmas its proof rests on.

Setting

Let X\mathcal XX be a measurable space, PPP a probability measure on it, and X1,…,XnX_1,\dots,X_nX1​,…,Xn​ (n≥1n\ge1n≥1) an i.i.d. sample from PPP with empirical distribution P^n\widehat P_nPn​. Let F\mathcal FF be a nonempty class of measurable functions f:X→[M0,M1]f:\mathcal X\to[M_0,M_1]f:X→[M0​,M1​], and set M=M1−M0M = M_1-M_0M=M1​−M0​. Write E[f]=∫f dP\mathbb E[f]=\int f\,dPE[f]=∫fdP, Var(f)\mathrm{Var}(f)Var(f) for the variance of f(X)f(X)f(X), and

EP^n[f]=1n∑i=1nf(Xi),VarP^n(f)=1n∑i=1nf(Xi)2−(EP^n[f])2.\mathbb E_{\widehat P_n}[f] = \frac1n\sum_{i=1}^n f(X_i),\qquad \mathrm{Var}_{\widehat P_n}(f) = \frac1n\sum_{i=1}^n f(X_i)^2 - \big(\mathbb E_{\widehat P_n}[f]\big)^2 .EPn​​[f]=n1​i=1∑n​f(Xi​),VarPn​​(f)=n1​i=1∑n​f(Xi​)2−(EPn​​[f])2.

For ρ≥0\rho\ge0ρ≥0, the χ2\chi^2χ2 ball Pn\mathcal P_nPn​ is the set of weight vectors p∈Rnp\in\mathbb R^np∈Rn with pi≥0p_i\ge0pi​≥0, ∑ipi=1\sum_i p_i = 1∑i​pi​=1 and 12∑i(npi−1)2≤ρ\frac12\sum_i (np_i-1)^2\le\rho21​∑i​(npi​−1)2≤ρ: the distributions PPP on the sample with Dϕ(P∥P^n)≤ρ/nD_\phi(P\|\widehat P_n)\le\rho/nDϕ​(P∥Pn​)≤ρ/n for ϕ(t)=12(t−1)2\phi(t)=\frac12(t-1)^2ϕ(t)=21​(t−1)2. The robust risk of fff is

Rn(f)=sup⁡P: Dϕ(P∥P^n)≤ρ/nEP[f(X)]=max⁡p∈Pn∑i=1npif(Xi),R_n(f) = \sup_{P:\,D_\phi(P\|\widehat P_n)\le \rho/n}\mathbb E_P[f(X)] = \max_{p\in\mathcal P_n}\sum_{i=1}^n p_i f(X_i),Rn​(f)=P:Dϕ​(P∥Pn​)≤ρ/nsup​EP​[f(X)]=p∈Pn​max​i=1∑n​pi​f(Xi​),

and a robust minimizer is any f^∈argmin⁡f∈FRn(f)\widehat f\in\operatorname{argmin}_{f\in\mathcal F} R_n(f)f​∈argminf∈F​Rn​(f).

Complexity is measured by empirical ℓ∞\ell_\inftyℓ∞​ covering numbers. For V⊂RmV\subset\mathbb R^mV⊂Rm, N(V,ϵ,∥⋅∥∞)N(V,\epsilon,\|\cdot\|_\infty)N(V,ϵ,∥⋅∥∞​) is the least number of points v1,…,vN∈Vv_1,\dots,v_N\in Vv1​,…,vN​∈V such that every v∈Vv\in Vv∈V lies within sup-distance ϵ\epsilonϵ of some viv_ivi​. For x∈Xmx\in\mathcal X^mx∈Xm let F(x)={(f(x1),…,f(xm)):f∈F}\mathcal F(x)=\{(f(x_1),\dots,f(x_m)) : f\in\mathcal F\}F(x)={(f(x1​),…,f(xm​)):f∈F}, and

N∞(F,ϵ,m)=sup⁡x∈XmN(F(x),ϵ,∥⋅∥∞)∈N∪{∞}.N_\infty(\mathcal F,\epsilon,m) = \sup_{x\in\mathcal X^m} N\big(\mathcal F(x),\epsilon,\|\cdot\|_\infty\big)\in\mathbb N\cup\{\infty\}.N∞​(F,ϵ,m)=x∈Xmsup​N(F(x),ϵ,∥⋅∥∞​)∈N∪{∞}.

Formalization targets

Goal: the oracle inequality (16)

Let n≥8M2/tn\ge 8M^2/tn≥8M2/t, t≥log⁡12t\ge\log 12t≥log12, ϵ>0\epsilon>0ϵ>0 and ρ≥9t\rho\ge 9tρ≥9t. With probability at least 1−2(3N∞(F,ϵ,2n)+1)e−t1-2(3N_\infty(\mathcal F,\epsilon,2n)+1)e^{-t}1−2(3N∞​(F,ϵ,2n)+1)e−t, every robust minimizer f^\widehat ff​ satisfies

E[f^(X)]≤inf⁡f∈F{E[f]+22ρnVar(f)}+19Mρ3n+(2+42tn)ϵ.\mathbb E[\widehat f(X)] \le \inf_{f\in\mathcal F}\left\{\mathbb E[f] + 2\sqrt{\frac{2\rho}{n}\mathrm{Var}(f)}\right\} + \frac{19M\rho}{3n} + \left(2+4\sqrt{\frac{2t}{n}}\right)\epsilon .E[f​(X)]≤f∈Finf​{E[f]+2n2ρ​Var(f)​}+3n19Mρ​+(2+4n2t​​)ϵ.

The certificate (15)

Under the same hypotheses and with the same probability, simultaneously for all f∈Ff\in\mathcal Ff∈F,

E[f(X)]≤Rn(f)+113Mρn+(2+42tn)ϵ.\mathbb E[f(X)] \le R_n(f) + \frac{11}{3}\frac{M\rho}{n} + \left(2+4\sqrt{\frac{2t}{n}}\right)\epsilon .E[f(X)]≤Rn​(f)+311​nMρ​+(2+4n2t​​)ϵ.

Supporting results (milestones)

  1. Theorem 1, inequality (10): for every vector z∈[M0,M1]nz\in[M_0,M_1]^nz∈[M0​,M1​]n, the robust mean minus the sample mean lies between (2ρsn2/n−2Mρ/n)+\big(\sqrt{2\rho s_n^2/n}-2M\rho/n\big)_+(2ρsn2​/n​−2Mρ/n)+​ and 2ρsn2/n\sqrt{2\rho s_n^2/n}2ρsn2​/n​.
  2. Lemma C.1: a uniform empirical Bernstein bound over F\mathcal FF with probability 1−6N∞(F,ϵ,2n)e−t1-6N_\infty(\mathcal F,\epsilon,2n)e^{-t}1−6N∞​(F,ϵ,2n)e−t.
  3. Lemma A.1, first bound: P(sn≥Esn2+t)≤exp⁡(−nt2/(2M2))\mathbb P(s_n\ge\sqrt{\mathbb E s_n^2}+t)\le\exp(-nt^2/(2M^2))P(sn​≥Esn2​​+t)≤exp(−nt2/(2M2)).
  4. Bernstein's inequality for one fixed fff, as displayed in the proof (p. 38).
  5. The certificate (15).

Significance

Inequality (15) says the robust risk is a uniform upper confidence bound on the population risk, with an O(1/n)O(1/n)O(1/n) slack instead of the O(1/n)O(1/\sqrt n)O(1/n​) slack of the empirical risk. Inequality (16) says the robust minimizer competes with the best variance-penalized population risk in the class. When some f∈Ff\in\mathcal Ff∈F has small risk and small variance, the excess risk of f^\widehat ff​ is of order 1/n1/n1/n up to the covering term, a rate ERM does not achieve in general (§3.3 of the paper gives an example). For a parametric class with N∞(F,ϵ,2n)N_\infty(\mathcal F,\epsilon,2n)N∞​(F,ϵ,2n) polynomial in 1/ϵ1/\epsilon1/ϵ, choosing ϵ=M/n\epsilon=M/nϵ=M/n gives Corollaries 3.1 and 3.2 of the paper.

The results are proved in the paper; none of them has a machine-checked proof that we know of. The mission's output is a formal proof of Theorem 3 and its ingredients: a deterministic analysis of the χ2\chi^2χ2-constrained linear program (Theorem 1 (10)), a covering-number empirical Bernstein inequality (Lemma C.1, from Maurer and Pontil), concentration of the sample standard deviation (Lemma A.1), and the scalar Bernstein inequality in the form used. Each of these is reusable outside distributionally robust optimization.

Difficulty

The deterministic part, (10), is a short analysis of a quadratically constrained linear program. The main obstacle is Lemma C.1. A union bound over a cover of F\mathcal FF fails directly: the cover depends on the sample, and a population-level cover of F\mathcal FF need not be finite. The standard route goes through a ghost sample of size nnn (hence covering at 2n2n2n points), a symmetrization that must preserve the sample variance rather than only the mean, and a concentration bound for the sample variance itself. Lemma A.1 needs concentration of sns_nsn​, a non-linear and non-smooth function of the sample, at the sub-Gaussian rate M/nM/\sqrt nM/n​. Finally, the oracle inequality (16) holds for an infimum over the whole class, while the concentration step for the comparison function is only proved for one fixed fff at a time.

Formalization scope

The sample is the coordinate process of the product measure P⊗nP^{\otimes n}P⊗n on Xn\mathcal X^nXn (Measure.pi). Each probability statement bounds the probability of the bad event, the set of samples where the inequality fails for some fff (or some minimizer). This set need not be measurable, and its measure is then the outer measure, as is standard in empirical-process theory. Probability bounds are computed in [0,∞][0,\infty][0,∞], and the covering number is valued in N∪{∞}\mathbb N\cup\{\infty\}N∪{∞}, so an infinite covering number makes the bound trivial rather than collapsing to zero. Covering numbers are internal (centres in F(x)\mathcal F(x)F(x)) and use closed sup-norm balls, as on p. 9; this is Mathlib's Metric.coveringNumber. The empirical variance is normalized by 1/n1/n1/n. The χ2\chi^2χ2 ball is encoded as weight vectors on the sample points; with tied sample values this gives the same supremum as the paper's distributions on the sample. Statement (16) is formalized for every minimizer of the robust risk, and the event is empty if no minimizer exists. The infimum ranges over the nonempty class F\mathcal FF, on which every term is at least M0M_0M0​. Population moments are those of bounded measurable functions, hence finite.

Deviations from the printed text:

  • Lemma A.1 is stated only for its first (upper-tail) bound. The paper derives the second bound from Lemma A.4, which is false as printed; the second bound is not stated. M>0M>0M>0 is assumed because M2M^2M2 is a denominator.
  • Lemma C.1 is the paper's restatement of Maurer and Pontil's Theorem 6, with a general radius ϵ\epsilonϵ. It is formalized as printed, with the implicit assumption ϵ>0\epsilon>0ϵ>0 made explicit.
  • n≥1n\ge1n≥1 is assumed throughout. The hypothesis n≥8M2/tn\ge 8M^2/tn≥8M2/t is kept as printed.

A trivializing formalization is ruled out: the bound is not taken over all functions, a probability bound is not formed from the real part of an infinite covering number, and the minimizer is not a hypothesis that can fail to exist for the given sample.

Needed infrastructure: product-measure concentration (Bernstein, and a bounded-difference or convex-Lipschitz inequality for sns_nsn​), symmetrization with a ghost sample, and finite union bounds over a cover. Contributions are welcome on any milestone, in particular a general covering-number empirical Bernstein inequality, which is reusable on its own.

Selected references

  • J. C. Duchi and H. Namkoong, Variance-based regularization with convex objectives, arXiv:1610.02581v3, 2017 (NIPS 2017; JMLR 20, 2019). https://arxiv.org/abs/1610.02581
  • A. Maurer and M. Pontil, Empirical Bernstein bounds and sample variance penalization, COLT 2009. https://arxiv.org/abs/0907.3740
  • S. Boucheron, G. Lugosi and P. Massart, Concentration Inequalities: A Nonasymptotic Theory of Independence, Oxford University Press, 2013. https://doi.org/10.1093/acprof:oso/9780199535255.001.0001
  • A. W. van der Vaart and J. A. Wellner, Weak Convergence and Empirical Processes, Springer, 1996. https://doi.org/10.1007/978-1-4757-2545-2
11 thms1 active userReviewed
CombinatoricsFunctional Analysis·Captain: mikedeng1

The Sample Complexity of Pattern Classification with Neural Networks: The Size of the Weights is More Important than the Size of the Network II: Fat-Shattering Bound for Bounded-Weight NetworksResearch Paper

Motivation

In the mid-1990s, neural networks trained by gradient descent were observed to generalize well even when the number of weights far exceeded the number of training examples. The classical theory could not explain this: VC-dimension bounds for networks grow with the number of parameters, so for large networks they are vacuous. Bartlett's paper (IEEE Trans. Inform. Theory 44 (1998)) showed that, for classification with a margin, what controls generalization is the size of the weights, not the size of the network. Its two main technical results are a margin bound in terms of the fat-shattering dimension (Theorem 2, the subject of mission I of this series) and a bound on the fat-shattering dimension of networks with bounded weights (Theorem 17, the subject of this mission). The same idea of weight-norm capacity control underlies much of the later theory of margins, boosting and kernel methods.

Timeline. Kearns and Schapire (1994) introduced the fat-shattering dimension. Alon, Ben-David, Cesa-Bianchi and Haussler (1997) bounded ℓ∞ covering numbers by it. Bartlett, Kulkarni and Posner (1997) gave the matching lower bound on ℓ1 covering numbers used here as Lemma 19. Maurey's approximation lemma (reported by Pisier, 1981) was used by Jones (1992) and Barron (1993) for approximation by networks, and by Lee, Bartlett and Williamson (1996) for covering numbers of convex hulls. Bartlett (1998) combined these into Theorem 17.

Setting

Let XXX be a set and HHH a class of functions X→RX\to\mathbb RX→R. For γ>0\gamma>0γ>0, a sequence x=(x1,…,xm)∈Xmx=(x_1,\dots,x_m)\in X^mx=(x1​,…,xm​)∈Xm is γ\gammaγ-shattered by HHH if there is r∈Rmr\in\mathbb R^mr∈Rm such that for every b∈{−1,1}mb\in\{-1,1\}^mb∈{−1,1}m some h∈Hh\in Hh∈H satisfies (h(xi)−ri)bi≥γ(h(x_i)-r_i)b_i\ge\gamma(h(xi​)−ri​)bi​≥γ for all iii. The fat-shattering dimension is

fat⁡H(γ)=max⁡{m: H γ-shatters some x∈Xm}∈N∪{∞}.\operatorname{fat}_H(\gamma)=\max\{m:\ H\ \gamma\text{-shatters some }x\in X^m\}\in\mathbb N\cup\{\infty\}.fatH​(γ)=max{m: H γ-shatters some x∈Xm}∈N∪{∞}.

A cover of a class FFF at scale ε\varepsilonε for a pseudometric ρ\rhoρ on functions is a set TTT of functions such that every f∈Ff\in Ff∈F has some t∈Tt\in Tt∈T with ρ(t,f)<ε\rho(t,f)<\varepsilonρ(t,f)<ε; N(F,ε,ρ)\mathcal N(F,\varepsilon,\rho)N(F,ε,ρ) is the least size of a cover. For a sample x∈Xmx\in X^mx∈Xm the pseudometrics dℓ∞(x)d_{\ell_\infty(x)}dℓ∞​(x)​, dℓ1(x)d_{\ell_1(x)}dℓ1​(x)​, dℓ2(x)d_{\ell_2(x)}dℓ2​(x)​ are the maximum, the mean, and the root mean square of ∣f(xi)−g(xi)∣|f(x_i)-g(x_i)|∣f(xi​)−g(xi​)∣ over iii, and the uniform covering numbers are Np(F,ε,m)=max⁡x∈XmN(F,ε,dℓp(x))\mathcal N_p(F,\varepsilon,m)=\max_{x\in X^m}\mathcal N(F,\varepsilon,d_{\ell_p(x)})Np​(F,ε,m)=maxx∈Xm​N(F,ε,dℓp​(x)​).

The hidden units form a nonempty class FFF of functions X→[−M/2,M/2]X\to[-M/2,M/2]X→[−M/2,M/2]. For A>0A>0A>0 the two-layer network class with ℓ1-bounded output weights is

H={∑i=1Nwifi: N∈N, fi∈F, ∑i=1N∣wi∣≤A}.H=\Big\{\sum_{i=1}^Nw_if_i:\ N\in\mathbb N,\ f_i\in F,\ \sum_{i=1}^N|w_i|\le A\Big\}.H={i=1∑N​wi​fi​: N∈N, fi​∈F, i=1∑N​∣wi​∣≤A}.

In Lean these are BartlettNN.Margin.fat, BartlettNN.Margin.coverNum and BartlettNN.Margin.Ninf (shared with mission I), and BartlettNN.FatNet.N1, BartlettNN.FatNet.N2 and BartlettNN.FatNet.combos F A.

Formalization targets

Goal: Theorem 17

There is a universal constant ccc such that for every XXX, FFF, MMM, A>0A>0A>0 and γ>0\gamma>0γ>0 with d=fat⁡F(γ/(32A))≥1d=\operatorname{fat}_F(\gamma/(32A))\ge1d=fatF​(γ/(32A))≥1,

fat⁡H(γ)≤cM2A2dγ2ln⁡2(MAdγ).\operatorname{fat}_H(\gamma)\le\frac{cM^2A^2d}{\gamma^2}\ln^2\Big(\frac{MAd}{\gamma}\Big).fatH​(γ)≤γ2cM2A2d​ln2(γMAd​).

The constant is left unspecified, as in the paper, so that the goal survives any improvement of the numerical constants.

Milestones (in the order the proof uses them)

  1. Lemma 19 (cited from Bartlett–Kulkarni–Posner): for [0,1][0,1][0,1]-valued FFF with fat⁡F(4γ)≥d\operatorname{fat}_F(4\gamma)\ge dfatF​(4γ)≥d, log⁡2N1(F,γ,d)≥d/32\log_2\mathcal N_1(F,\gamma,d)\ge d/32log2​N1​(F,γ,d)≥d/32.
  2. Lemma 20, (5): for d=fat⁡F(γ/4)d=\operatorname{fat}_F(\gamma/4)d=fatF​(γ/4) and m≥2+2dlog⁡2(32M/γ)m\ge2+2d\log_2(32M/\gamma)m≥2+2dlog2​(32M/γ), log⁡2N2(F,γ,m)<1+dlog⁡2(4emM/(dγ))log⁡2(9mM2/γ2)\log_2\mathcal N_2(F,\gamma,m)<1+d\log_2(4emM/(d\gamma))\log_2(9mM^2/\gamma^2)log2​N2​(F,γ,m)<1+dlog2​(4emM/(dγ))log2​(9mM2/γ2).
  3. Lemma 21 (Maurey): in a Hilbert space, a point of the closed convex hull of a set of norm at most bbb is within c/k\sqrt{c/k}c/k​ of an average of kkk points of the set, for every c>b2−∥h∥2c>b^2-\|h\|^2c>b2−∥h∥2.
  4. Lemma 22: log⁡2N2(H,γ,m)≤(2M2A2/γ2)log⁡2(2N2(F,γ/(2A),m)+1)\log_2\mathcal N_2(H,\gamma,m)\le(2M^2A^2/\gamma^2)\log_2(2\mathcal N_2(F,\gamma/(2A),m)+1)log2​N2​(H,γ,m)≤(2M2A2/γ2)log2​(2N2​(F,γ/(2A),m)+1).
  5. Inequality (6): if m=fat⁡H(4γ)≥2+2dlog⁡2(64MA/γ)m=\operatorname{fat}_H(4\gamma)\ge2+2d\log_2(64MA/\gamma)m=fatH​(4γ)≥2+2dlog2​(64MA/γ) with d=fat⁡F(γ/(8A))d=\operatorname{fat}_F(\gamma/(8A))d=fatF​(γ/(8A)), then m≤(64M2A2/γ2)(3+dlog⁡2(8emMA/γ)log⁡2(36mM2A2/γ2))m\le(64M^2A^2/\gamma^2)(3+d\log_2(8emMA/\gamma)\log_2(36mM^2A^2/\gamma^2))m≤(64M2A2/γ2)(3+dlog2​(8emMA/γ)log2​(36mM2A2/γ2)).

Significance

Theorem 17 bounds the capacity of a network class without reference to the number of hidden units NNN. With Theorem 2 (mission I) it gives misclassification bounds for networks with small weights that hold for networks of any size, and by iteration it yields the bounds for deep sigmoid networks of Theorem 28 (mission III). Its method — upper-bound ℓ2 covering numbers through Maurey's lemma and compare with a lower bound in terms of fat-shattering — is a template for bounding the fat-shattering dimension of convex hulls in general.

The results are proved in the paper (Lemmas 19 and 21 by citation). None of them is formalized: the platform has no fat-shattering dimension, no uniform sample covering numbers of function classes, and only a finite-dimensional, diameter-based form of Maurey's lemma (HighDimProb.Appetizer.approx_caratheodory), which is not Lemma 21. This mission produces machine-checked statements of all five ingredients and of the theorem.

Difficulty

The obvious route would bound fat⁡H\operatorname{fat}_HfatH​ through a VC-type count of the network's parameters, which fails because NNN is unbounded. The paper's route needs a lower bound on covering numbers by the fat-shattering dimension (Lemma 19, a combinatorial packing argument not proved in the paper), an upper bound by the fat-shattering dimension at a finer scale (Lemma 20, which goes through the Alon et al. scale-sensitive Sauer lemma and a quantization argument), and a probabilistic approximation argument in the empirical L2L_2L2​ space (Lemmas 21 and 22). The final step solves a transcendental inequality (6) for mmm, with care at the boundary where the logarithm is small.

Formalization scope

  • Functions are maps X → ℝ; fat is valued in ℕ∞, so an unbounded shattering is ∞\infty∞, not a junk 000. Labels ±1\pm1±1 are Bool read through pm (true ↦ 1); sequences are indexed by Fin m (0-based).
  • Covers are external (finite sets of arbitrary functions X→RX\to\mathbb RX→R) with the strict inequality of Definition 3; the covering number is ∞\infty∞ when no finite cover exists. The ℓ1 and ℓ2 distances carry the factor 1/m1/m1/m. Mathlib's Metric.coveringNumber (closed balls, metric types) is not used.
  • Bounds of the form "log⁡2N≤B\log_2\mathcal N\le Blog2​N≤B" are stated for every finite value of N\mathcal NN, and where the paper's bound implies finiteness (Lemmas 20, 22, Theorem 17) finiteness is part of the conclusion.
  • The constant ccc of Theorem 17 is quantified before XXX, FFF, MMM, AAA, γ\gammaγ and ddd; a constant chosen after them would make the statement trivially true.
  • Corrections of the printed statement. Theorem 17 is stated for γ>0\gamma>0γ>0 (printed γ≥0\gamma\ge0γ≥0, under which the bound is meaningless) and A>0A>0A>0 (printed A≥0A\ge0A≥0, under which γ/(32A)\gamma/(32A)γ/(32A) is undefined). Implicit positivity (M>0M>0M>0, γ>0\gamma>0γ>0, A>0A>0A>0) in Lemmas 20, 22 and (6) is stated as hypotheses. Inequality (6) is copied as printed, with log⁡2(8emMA/γ)\log_2(8emMA/\gamma)log2​(8emMA/γ).
  • In Theorem 17 the logarithm is natural (the base is absorbed by ccc); (5) and (6) use log⁡2\log_2log2​.
  • Lemma 21 is stated in a complete real inner product space, with "convex closure" read as the closure of the convex hull.

Contributions welcome: proofs of the five milestones and of the goal, and reusable infrastructure on fat-shattering and covering numbers of function classes.

Selected references

  • P. L. Bartlett, The Sample Complexity of Pattern Classification with Neural Networks: The Size of the Weights is More Important than the Size of the Network, IEEE Trans. Inform. Theory 44(2), 1998, 525–536. https://doi.org/10.1109/18.661502
  • N. Alon, S. Ben-David, N. Cesa-Bianchi, D. Haussler, Scale-sensitive dimensions, uniform convergence, and learnability, J. ACM 44(4), 1997, 615–631. https://doi.org/10.1145/263867.263927
  • P. L. Bartlett, S. R. Kulkarni, S. E. Posner, Covering numbers for real-valued function classes, IEEE Trans. Inform. Theory 43(5), 1997, 1721–1724. https://doi.org/10.1109/18.623181
  • M. J. Kearns, R. E. Schapire, Efficient distribution-free learning of probabilistic concepts, J. Comput. Syst. Sci. 48(3), 1994, 464–497. https://doi.org/10.1016/S0022-0000(05)80062-5
  • W. S. Lee, P. L. Bartlett, R. C. Williamson, Efficient agnostic learning of neural networks with bounded fan-in, IEEE Trans. Inform. Theory 42(6), 1996, 2118–2132. https://doi.org/10.1109/18.556601
  • A. R. Barron, Universal approximation bounds for superpositions of a sigmoidal function, IEEE Trans. Inform. Theory 39(3), 1993, 930–945. https://doi.org/10.1109/18.256500
11 thms1 active userReviewed
Algorithmic Game TheoryConvex OptimizationOptimization·Captain: mikedeng1

Blackwell Approachability and No-Regret Learning are Equivalent 2: A No-Regret Algorithm and a Valid Halfspace Oracle Approach a Compact Convex Set at Rate 2·Regret_T/TResearch Paper

Motivation

Blackwell approachability is the vector-payoff analogue of von Neumann's minimax theorem. In a repeated game where each round's outcome is a vector u(xt,yt)∈Rdu(x_t, y_t) \in \mathbb R^du(xt​,yt​)∈Rd, a player wants the running average of these vectors to converge to a target set SSS, whatever the opponent does. Blackwell (1956) showed when this is possible, and approachability has since become a standard tool for calibrated forecasting, regret minimization with respect to general benchmarks, and learning in games.

Online linear optimization (OLO) is the problem of choosing points θt\theta_tθt​ in a fixed decision set K\mathcal KK against a sequence of linear losses ⟨ft,⋅⟩\langle f_t, \cdot\rangle⟨ft​,⋅⟩, with performance measured by regret against the best fixed point in hindsight. Algorithms with regret o(T)o(T)o(T) — "no-regret" algorithms such as online gradient descent (Zinkevich, 2003) — are among the most studied objects of machine learning.

Abernethy, Bartlett and Hazan (COLT 2011) showed that the two problems are algorithmically equivalent: each can be converted into the other with explicit control of the rates. This mission covers the direction from OLO to approachability.

Timeline:

  • 1956: Blackwell proves the approachability theorem for convex sets, via a geometric projection strategy.
  • 2003: Zinkevich introduces online gradient descent, a no-regret algorithm for any bounded convex decision set.
  • 2009: Even-Dar, Kleinberg, Mannor and Mansour state approachability in the response-satisfiability form (as cited on p. 32 of the 2011 paper).
  • 2011: Abernethy, Bartlett and Hazan give the two reductions, with explicit rates, and apply them to efficient calibration.

Setting

A Blackwell instance (X,Y,u,S)(\mathcal X, \mathcal Y, u, S)(X,Y,u,S) consists of compact convex sets X⊆Rn\mathcal X \subseteq \mathbb R^nX⊆Rn, Y⊆Rm\mathcal Y \subseteq \mathbb R^mY⊆Rm, a payoff u:X×Y→Rdu : \mathcal X \times \mathcal Y \to \mathbb R^du:X×Y→Rd that is affine in each argument (biaffine), and a closed convex target set S⊆RdS \subseteq \mathbb R^dS⊆Rd. Write dist(z,U)=inf⁡w∈U∥z−w∥\mathtt{dist}(z, U) = \inf_{w \in U}\|z - w\|dist(z,U)=infw∈U​∥z−w∥ for the Euclidean distance to a set, and B2(r)B_2(r)B2​(r) for the closed Euclidean ball of radius rrr.

A halfspace oracle takes a halfspace H={z:⟨a,z⟩≤c}H = \{z : \langle a, z\rangle \le c\}H={z:⟨a,z⟩≤c} and returns a point O(H)∈X\mathcal O(H) \in \mathcal XO(H)∈X; it is valid if for every halfspace H⊇SH \supseteq SH⊇S, u(O(H),y)∈Hu(\mathcal O(H), y) \in Hu(O(H),y)∈H for all y∈Yy \in \mathcal Yy∈Y.

A set X⊆RdX \subseteq \mathbb R^dX⊆Rd is a cone if αz∈X\alpha z \in Xαz∈X for all z∈Xz \in Xz∈X, α≥0\alpha \ge 0α≥0. For K⊆RdK \subseteq \mathbb R^dK⊆Rd, cone(K)={αx:α≥0,x∈K}\mathtt{cone}(K) = \{\alpha x : \alpha \ge 0, x \in K\}cone(K)={αx:α≥0,x∈K}, and the polar cone of CCC is C0={θ:⟨θ,x⟩≤0 ∀x∈C}C^0 = \{\theta : \langle \theta, x\rangle \le 0 \ \forall x \in C\}C0={θ:⟨θ,x⟩≤0 ∀x∈C}.

An OLO algorithm L\mathcal LL maps past loss vectors (f1,…,ft−1)(f_1, \dots, f_{t-1})(f1​,…,ft−1​) to a point θt∈K\theta_t \in \mathcal Kθt​∈K, and its regret is

RegretT=∑t=1T⟨ft,θt⟩−min⁡θ∈K∑t=1T⟨ft,θ⟩.\mathrm{Regret}_T = \sum_{t=1}^T \langle f_t, \theta_t\rangle - \min_{\theta \in \mathcal K} \sum_{t=1}^T \langle f_t, \theta\rangle .RegretT​=t=1∑T​⟨ft​,θt​⟩−θ∈Kmin​t=1∑T​⟨ft​,θ⟩.

Algorithm 2 runs L\mathcal LL on K=S0∩B2(1)\mathcal K = S^0 \cap B_2(1)K=S0∩B2​(1) when SSS is a cone: at round ttt it sets θt=L(f1,…,ft−1)\theta_t = \mathcal L(f_1, \dots, f_{t-1})θt​=L(f1​,…,ft−1​), plays xt=O({z:⟨θt,z⟩≤0})x_t = \mathcal O(\{z : \langle \theta_t, z\rangle \le 0\})xt​=O({z:⟨θt​,z⟩≤0}), observes yt∈Yy_t \in \mathcal Yyt​∈Y, and feeds ft=−u(xt,yt)f_t = -u(x_t, y_t)ft​=−u(xt​,yt​) back to L\mathcal LL.

When SSS is compact but not a cone, it is lifted: with κ=max⁡s∈S∥s∥\kappa = \max_{s\in S}\|s\|κ=maxs∈S​∥s∥ and κ⊕z∈Rd+1\kappa \oplus z \in \mathbb R^{d+1}κ⊕z∈Rd+1 the concatenation, put u′(x,y)=κ⊕u(x,y)u'(x, y) = \kappa \oplus u(x, y)u′(x,y)=κ⊕u(x,y) and S′=cone({κ}×S)S' = \mathtt{cone}(\{\kappa\} \times S)S′=cone({κ}×S), and run Algorithm 2 on (X,Y,u′,S′)(\mathcal X, \mathcal Y, u', S')(X,Y,u′,S′).

Formalization targets

Goal: Corollary 18 (p. 39)

For a Blackwell instance with SSS nonempty and compact, any valid halfspace oracle for the lifted instance, any OLO algorithm with values in K′=(S′)0∩B2(1)\mathcal K' = (S')^0 \cap B_2(1)K′=(S′)0∩B2​(1), any T≥1T \ge 1T≥1 and any y1,…,yT∈Yy_1, \dots, y_T \in \mathcal Yy1​,…,yT​∈Y, the run of Algorithm 2 on the lifted instance satisfies

dist(1T∑t=1Tu(xt,yt),S)≤2 dist(1T∑t=1Tu′(xt,yt),S′)≤2T RegretT.\mathtt{dist}\Big(\frac1T\sum_{t=1}^T u(x_t,y_t), S\Big) \le 2\,\mathtt{dist}\Big(\frac1T\sum_{t=1}^T u'(x_t,y_t), S'\Big) \le \frac2T\,\mathrm{Regret}_T .dist(T1​t=1∑T​u(xt​,yt​),S)≤2dist(T1​t=1∑T​u′(xt​,yt​),S′)≤T2​RegretT​.

The bound holds for every TTT and every adversary, with no rate assumed for L\mathcal LL; a no-regret L\mathcal LL then gives approachability.

Milestones

  1. Lemma 13 (p. 35): for a nonempty convex cone CCC, dist(x,C)=max⁡θ∈C0∩B2(1)⟨θ,x⟩\mathtt{dist}(x, C) = \max_{\theta \in C^0 \cap B_2(1)} \langle \theta, x\rangledist(x,C)=maxθ∈C0∩B2​(1)​⟨θ,x⟩.
  2. Theorem 17 (p. 38): if SSS is a cone, Algorithm 2 achieves dist(1T∑tu(xt,yt),S)≤Regret(LK;f1:T)/T\mathtt{dist}\big(\frac1T\sum_t u(x_t,y_t), S\big) \le \mathrm{Regret}(\mathcal L_{\mathcal K}; f_{1:T})/Tdist(T1​∑t​u(xt​,yt​),S)≤Regret(LK​;f1:T​)/T.
  3. Lemma 14 (p. 35): for nonempty compact convex K\mathcal KK, κ=max⁡K∥⋅∥\kappa = \max_{\mathcal K}\|\cdot\|κ=maxK​∥⋅∥ and x∉Kx \notin \mathcal Kx∈/K, dist(κ⊕x,cone({κ}×K))≤dist(x,K)≤2 dist(κ⊕x,cone({κ}×K))\mathtt{dist}(\kappa\oplus x, \mathtt{cone}(\{\kappa\}\times\mathcal K)) \le \mathtt{dist}(x, \mathcal K) \le 2\,\mathtt{dist}(\kappa\oplus x, \mathtt{cone}(\{\kappa\}\times\mathcal K))dist(κ⊕x,cone({κ}×K))≤dist(x,K)≤2dist(κ⊕x,cone({κ}×K)).

Significance

The result. Corollary 18 turns any no-regret algorithm into an approachability strategy for a compact convex target, provided a valid halfspace oracle is available, with rate 2 RegretT/T2\,\mathrm{Regret}_T/T2RegretT​/T. Combined with online gradient descent it gives an O(1/T)O(1/\sqrt T)O(1/T​) approachability rate, and through the choice of OLO algorithm it lets approachability inherit the computational efficiency of online learning. The paper uses this route to build an efficient calibrated forecaster (Section 5). Together with the converse reduction (Theorem 16), it shows that the two problems are equivalent.

Formalizing it. The results are proved in the paper; none of them has been machine-checked. Formalizing them requires the conic duality formula for distances (Lemma 13), a quantitative lifting lemma (Lemma 14) and the bookkeeping of an interactive protocol. The proof of Lemma 14 on the page is a sketch: it refers to an undefined point and uses a triangle-similarity argument, so a complete proof is new work.

Difficulty

The reduction's core is Lemma 13: the distance to a cone is a maximum of a linear function over the polar cone's unit ball. Lemma 13 needs projection onto a cone in Euclidean space; for a non-closed cone the projection may not exist, and the argument must go through the closure. The lifting Lemma 14 is a geometric statement whose page proof relies on a picture and an undefined point, so the factor 2 has no complete written argument. Finally, connecting the average lifted payoff to the lift of the average payoff, and the halfspace guarantee ⟨θt,ft⟩≥0\langle\theta_t, f_t\rangle \ge 0⟨θt​,ft​⟩≥0 to the regret, requires keeping the round indexing and the oracle's validity domain exactly aligned.

Formalization scope

All spaces are EuclideanSpace ℝ (Fin d). The concatenation κ⊕z\kappa\oplus zκ⊕z lives in EuclideanSpace ℝ (Fin (d+1)) with coordinate 0 equal to κ\kappaκ, so ∥κ⊕z∥2=κ2+∥z∥2\|\kappa\oplus z\|^2 = \kappa^2 + \|z\|^2∥κ⊕z∥2=κ2+∥z∥2; a product type with the sup norm would change every distance and is ruled out. Distances are Metric.infDist. The polar cone uses the paper's sign (≤0\le 0≤0), the negative of Mathlib's innerDual. A halfspace is the pair (a,c)(a, c)(a,c); a valid oracle must answer every halfspace containing SSS, including a=0a = 0a=0, not only the halfspaces the algorithm happens to query. The OLO algorithm is a map from histories Fin t → ℝᴰ with values in S0∩B2(1)S^0 \cap B_2(1)S0∩B2​(1) at every history. Rounds are t=1,…,Tt = 1, \dots, Tt=1,…,T, and the run of Algorithm 2 is given as hypotheses on sequences θ,x,f\theta, x, fθ,x,f, which exist and are unique by recursion. The minimum in the regret and κ\kappaκ are written as sInf/sSup of images over nonempty compact sets, where they are attained.

Hypotheses added relative to the page: S≠∅S \neq \emptysetS=∅ and T≥1T \ge 1T≥1 in the goal; C≠∅C \ne \emptysetC=∅ in Lemma 13 (the empty set is a cone under Definition 11 and the identity fails for it); K≠∅\mathcal K \ne \emptysetK=∅ in Lemma 14. Corrected misprints, each disclosed in the item's note: "RegretT(A)\mathrm{Regret}_T(\mathcal A)RegretT​(A)" in Corollary 18 and (9) denotes the regret of the OLO algorithm L\mathcal LL on the lifted losses; Lemma 14's "K⊆H\mathcal K \subseteq \mathcal HK⊆H" has a stray H\mathcal HH; κ\kappaκ is the maximal norm of the set, not its diameter. The oracle in the goal is a valid oracle for the lifted instance, which is what applying Algorithm 2 to (X,Y,u′,S′)(\mathcal X, \mathcal Y, u', S')(X,Y,u′,S′) requires.

A formalization in which the oracle is valid only at the run's own queries, the OLO algorithm is unconstrained, the regret's minimum ranges over all of Rd+1\mathbb R^{d+1}Rd+1, or the middle term of the goal is dropped, is a different statement and is ruled out.

The development needs: the dual formula for the distance to a convex cone, nearest-point projection onto closed convex sets (in Mathlib), compactness of polar-cone slices, and finite sums of biaffine payoffs. The cone layer (Lemma 13) is reusable for the converse direction of the paper and for conic duality generally. Proofs of any milestone, and of the bridge from an oracle for the original instance to one for the lifted instance, are welcome.

Selected references

  • J. Abernethy, P. L. Bartlett, E. Hazan, Blackwell Approachability and No-Regret Learning are Equivalent, JMLR W&CP 19 (COLT 2011), pp. 27–46. https://proceedings.mlr.press/v19/abernethy11b.html
  • D. Blackwell, An analog of the minimax theorem for vector payoffs, Pacific Journal of Mathematics 6(1), 1956, pp. 1–8. https://doi.org/10.2140/pjm.1956.6.1
  • M. Zinkevich, Online convex programming and generalized infinitesimal gradient ascent, ICML 2003. https://www.aaai.org/Papers/ICML/2003/ICML03-120.pdf
6 thms1 active userReviewed
Dynamic ProgrammingReinforcement Learning·Captain: mikedeng1

Reinforcement Learning: An Introduction IV: Policy Iteration for ε-Soft PoliciesTextbook

Motivation

Policy iteration alternates two steps: evaluate the current policy, then replace it by a policy that is greedy with respect to the evaluated action values. The policy improvement theorem guarantees that each greedy step does not make the policy worse, and that the process stops only at an optimal policy. When the action values are estimated from experience rather than computed from a model, as in Monte Carlo control, a greedy policy is a problem: it never tries the actions it does not currently prefer, so their values are never re-estimated. Sutton and Barto, Reinforcement Learning: An Introduction (2nd ed., 2018), §5.4, resolve this without the unrealistic assumption of exploring starts by moving the policy only toward a greedy one, to an ε-greedy policy that keeps every action's probability at least ε/|A|.

The question this mission formalizes is whether policy iteration still works under that restriction. The book's answer (pp. 101–102) is yes, in a precise sense: an ε-greedy step never makes an ε-soft policy worse, and it fails to make it strictly better only when the policy is already the best among all ε-soft policies. This is the dynamic-programming fact that justifies on-policy first-visit Monte Carlo control for ε-soft policies, and more broadly every ε-greedy on-policy control scheme that is analysed with exact action values.

Setting

A finite Markov decision process has a finite state set S, a finite nonempty action set A used in every state, a finite reward set R ⊂ ℝ and dynamics p(s′, r | s, a) ≥ 0 with ∑s′,rp(s′,r∣s,a)=1\sum_{s', r} p(s', r \mid s, a) = 1∑s′,r​p(s′,r∣s,a)=1. A policy π gives, for each state s, a probability distribution π(· | s) on A. With a discount rate 0≤γ<10 \le \gamma < 10≤γ<1, the state value vπ(s)v_\pi(s)vπ​(s) is the expected discounted return Eπ[∑k≥0γkRt+k+1∣St=s]\mathbb E_\pi[\sum_{k \ge 0} \gamma^k R_{t+k+1} \mid S_t = s]Eπ​[∑k≥0​γkRt+k+1​∣St​=s], and the action value is

qπ(s,a)=∑s′,rp(s′,r∣s,a) [r+γvπ(s′)].q_\pi(s, a) = \sum_{s', r} p(s', r \mid s, a)\,[r + \gamma v_\pi(s')].qπ​(s,a)=s′,r∑​p(s′,r∣s,a)[r+γvπ​(s′)].

For ε > 0, a policy is ε-soft if π(a∣s)≥ε/∣A∣\pi(a \mid s) \ge \varepsilon/|A|π(a∣s)≥ε/∣A∣ for all s and a. A policy π′ is ε-greedy with respect to qπq_\piqπ​ if at each state some maximizer A∗(s)A^*(s)A∗(s) of qπ(s,⋅)q_\pi(s, \cdot)qπ​(s,⋅) receives probability 1−ε+ε/∣A∣1 - \varepsilon + \varepsilon/|A|1−ε+ε/∣A∣ and every other action receives ε/∣A∣\varepsilon/|A|ε/∣A∣; ties among maximizers are broken arbitrarily. A policy π is optimal among the ε-soft policies if it is ε-soft and vπ′′(s)≤vπ(s)v_{\pi''}(s) \le v_\pi(s)vπ′′​(s)≤vπ​(s) for every ε-soft π″ and every state s.

The book's analysis uses a new environment with the same states, actions and rewards, in which with probability 1 − ε the chosen action is executed and with probability ε a uniformly random action replaces it:

p~(s′,r∣s,a)=(1−ε) p(s′,r∣s,a)+∑a′ε∣A∣ p(s′,r∣s,a′).\tilde p(s', r \mid s, a) = (1 - \varepsilon)\, p(s', r \mid s, a) + \sum_{a'} \frac{\varepsilon}{|A|}\, p(s', r \mid s, a').p~​(s′,r∣s,a)=(1−ε)p(s′,r∣s,a)+a′∑​∣A∣ε​p(s′,r∣s,a′).

Its optimal value function is written v~∗\tilde v_*v~∗​.

Formalization targets

Goal: ε-greedy improvement with the equality case

For 0≤γ<10 \le \gamma < 10≤γ<1, 0<ε≤10 < \varepsilon \le 10<ε≤1, an ε-soft policy π and any ε-greedy policy π′ with respect to qπq_\piqπ​,

vπ′(s)≥vπ(s)for all s,andvπ′=vπ  ⟹  π,π′ are optimal among the ε-soft policies.v_{\pi'}(s) \ge v_\pi(s) \quad \text{for all } s, \qquad \text{and} \qquad v_{\pi'} = v_\pi \;\Longrightarrow\; \pi, \pi' \text{ are optimal among the ε-soft policies}.vπ′​(s)≥vπ​(s)for all s,andvπ′​=vπ​⟹π,π′ are optimal among the ε-soft policies.

The goal is stated in terms of the original MDP and ε-soft policies only; the new environment appears only in the milestones.

Milestones

  1. Policy improvement theorem for stochastic policies (4.7)–(4.8), p. 78: ∑aπ′(a∣s)qπ(s,a)≥vπ(s)\sum_a \pi'(a \mid s) q_\pi(s, a) \ge v_\pi(s)∑a​π′(a∣s)qπ​(s,a)≥vπ​(s) for all s implies vπ′≥vπv_{\pi'} \ge v_\pivπ′​≥vπ​, strictly at every state where the hypothesis is strict.
  2. Eq. (5.2), pp. 101–102: ∑aπ′(a∣s)qπ(s,a)=ε∣A∣∑aqπ(s,a)+(1−ε)max⁡aqπ(s,a)≥vπ(s)\sum_a \pi'(a \mid s) q_\pi(s, a) = \frac{\varepsilon}{|A|}\sum_a q_\pi(s, a) + (1-\varepsilon)\max_a q_\pi(s, a) \ge v_\pi(s)∑a​π′(a∣s)qπ​(s,a)=∣A∣ε​∑a​qπ​(s,a)+(1−ε)maxa​qπ​(s,a)≥vπ​(s).
  3. Characterization, p. 102: an ε-soft π is optimal among ε-soft policies if and only if vπ=v~∗v_\pi = \tilde v_*vπ​=v~∗​.
  4. Uniqueness, p. 102: v~∗\tilde v_*v~∗​ is the unique solution of the Bellman optimality equation with the altered transition probabilities p~\tilde pp~​, and that equation splits as (1−ε)max⁡a(⋅)+ε∣A∣∑a(⋅)(1-\varepsilon)\max_a(\cdot) + \frac{\varepsilon}{|A|}\sum_a(\cdot)(1−ε)maxa​(⋅)+∣A∣ε​∑a​(⋅).
  5. Fixed-point equation, p. 102: if vπ′=vπv_{\pi'} = v_\pivπ′​=vπ​, then vπ(s)=(1−ε)max⁡aqπ(s,a)+ε∣A∣∑aqπ(s,a)v_\pi(s) = (1-\varepsilon)\max_a q_\pi(s, a) + \frac{\varepsilon}{|A|}\sum_a q_\pi(s, a)vπ​(s)=(1−ε)maxa​qπ​(s,a)+∣A∣ε​∑a​qπ​(s,a).

Significance

The result is what makes ε-greedy on-policy control a form of generalized policy iteration: monotone improvement at every step, and a characterization of where the process can stop. It also locates precisely what is lost by exploring, namely that the fixed point is optimal among ε-soft policies, not among all policies. The value v~∗\tilde v_*v~∗​ of the new environment is the benchmark against which ε-greedy methods converge when action values are exact.

The book presents the argument informally and states the stochastic policy improvement theorem without proof ("we will not go through the details", p. 79). A formalization supplies the missing proof of the stochastic case, the identification of the best ε-soft policy value with the optimal value of a modified MDP, and the uniqueness of that value. As far as a search of the platform shows, no statement about ε-soft or ε-greedy policies has been formalized there; existing finite-MDP results (Bellman optimality in the Foundations of Machine Learning and Bertsekas series) use different reward models and do not cover the modified environment.

Difficulty

The improvement half follows from (5.2) and the policy improvement theorem, but both need work in the return-based model: the theorem requires comparing infinite discounted sums under two different Markov chains, and (5.2) uses the identity vπ(s)=∑aπ(a∣s)qπ(s,a)v_\pi(s) = \sum_a \pi(a \mid s) q_\pi(s, a)vπ​(s)=∑a​π(a∣s)qπ​(s,a), which is a theorem about returns, not a definition. The equality half is where the obvious argument fails. The deterministic-policy argument of Chapter 4 shows that an unimproved greedy policy satisfies the ordinary Bellman optimality equation; here the unimproved policy satisfies a different equation, and nothing in the original MDP identifies its solution with the best ε-soft value. That identification needs two further facts: every policy of the new environment corresponds to an ε-soft policy of the original one with the same values, and conversely (at ε = 1 only the uniform policy is ε-soft); and the altered optimality equation has exactly one solution.

Formalization scope

Everything is stated in the namespace SuttonBartoRL.EpsSoft on a finite MDP with four-argument dynamics, one finite nonempty action set for all states (so ∣A(s)∣=∣A∣|A(s)| = |A|∣A(s)∣=∣A∣, the book's footnote 3, p. 48), and a finite reward set. The conventions are:

  • vπv_\pivπ​ is defined from expected discounted returns as ∑kγk(Pπkrπ)(s)\sum_k \gamma^k (P_\pi^k r_\pi)(s)∑k​γk(Pπk​rπ​)(s) with 0≤γ<10 \le \gamma < 10≤γ<1; Bellman equations are theorems, never definitions. qπq_\piqπ​ is the one-step lookahead (4.6) on this vπv_\pivπ​.
  • v~∗\tilde v_*v~∗​ is the supremum of the new environment's policy values over all stochastic policies, a bounded family for γ<1\gamma < 1γ<1.
  • ε ranges over (0, 1]: the book requires ε > 0, and for ε > 1 no ε-soft policy exists. The equality in (5.2) is stated without the book's intermediate division by 1 − ε, so the case ε = 1 is included.
  • "Any ε-greedy policy" is encoded by quantifying over every choice of maximizer at every state.
  • "Optimal among ε-soft policies" means ε-soft and pointwise at least as good as every ε-soft policy.

Defining vπv_\pivπ​ as the solution of the Bellman expectation equation, or v~∗\tilde v_*v~∗​ as the solution of the altered optimality equation, would make milestones 3–5 and the goal's equality half hold by definition; the formalization does neither. The goal is not the statement "vπ′≥vπv_{\pi'} \ge v_\pivπ′​≥vπ​" alone: the equality case is part of the book's claim and part of the goal.

The finite-MDP definitions duplicate those of other missions in this series and are expected to be merged later. Useful contributions include the Neumann-series identity vπ=(I−γPπ)−1rπv_\pi = (I - \gamma P_\pi)^{-1} r_\pivπ​=(I−γPπ​)−1rπ​, the Bellman expectation equation, the contraction property of Bellman operators, and the correspondence between policies of the new environment and ε-soft policies of the original one; these are reusable for the other finite-MDP missions of the series.

Selected references

  • R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018, ISBN 9780262039246, §4.2 (pp. 76–79) and §5.4 (pp. 100–103). http://incompleteideas.net/book/the-book-2nd.html
  • R. A. Howard, Dynamic Programming and Markov Processes, MIT Press, 1960 (policy iteration).
  • M. L. Puterman, Markov Decision Processes: Discrete Stochastic Dynamic Programming, Wiley, 1994, doi:10.1002/9780470316887.
11 thms1 active userReviewed
OptimizationProbabilityStatistics·Captain: mikedeng1

Variance-based Regularization with Convex Objectives I: The χ²-Robust Risk Equals Empirical Risk plus a Standard-Deviation PenaltyResearch Paper

Motivation

Many statistical procedures minimize an average observed loss. This treats two candidates with the same average as equally attractive even when one has much more variable losses across the sample. Adding a multiple of the empirical standard deviation can distinguish them, but the resulting objective need not be convex even when each individual loss is convex. Duchi and Namkoong study a distributionally robust alternative: they maximize expected loss over a small neighborhood of the empirical distribution, then minimize that worst-case value. Their paper identifies when this convex robust value agrees exactly with the mean-plus-standard-deviation expression and how far apart the two can be otherwise. The finite-sample statement is Theorem 1 of the pinned preprint.

The relation matters to someone choosing a loss function for stochastic optimization. The variance expression has a direct statistical interpretation, while the robust expression preserves convexity in a decision parameter when the loss is convex. Theorem 1 makes the relationship quantitative for a single bounded random variable, before the paper turns to uniform guarantees over whole classes of losses. This mission isolates that first step and its finite optimization model.

Setting

Take observed real values z1,…,znz_1,\ldots,z_nz1​,…,zn​, with n≥1n\ge1n≥1. Their empirical mean and empirical variance are

zˉ=1n∑i=1nzi,sn2=1n∑i=1nzi2−zˉ2.\bar z=\frac1n\sum_{i=1}^n z_i,\qquad s_n^2=\frac1n\sum_{i=1}^n z_i^2-\bar z^2.zˉ=n1​i=1∑n​zi​,sn2​=n1​i=1∑n​zi2​−zˉ2.

The variance uses 1/n1/n1/n, not the unbiased-estimator factor 1/(n−1)1/(n-1)1/(n−1). A weight vector p=(p1,…,pn)p=(p_1,\ldots,p_n)p=(p1​,…,pn​) is feasible when its entries are nonnegative, sum to one, and satisfy

12∑i=1n(npi−1)2≤ρ,ρ≥0.\frac12\sum_{i=1}^n(np_i-1)^2\le\rho,\qquad \rho\ge0.21​i=1∑n​(npi​−1)2≤ρ,ρ≥0.

This is the paper's χ² neighborhood Pn(ρ)\mathcal P_n(\rho)Pn​(ρ) of the uniform empirical weights. Its robust sample expectation is

Rn(z,ρ)=sup⁡p∈Pn(ρ)∑i=1npizi.R_n(z,\rho)=\sup_{p\in\mathcal P_n(\rho)}\sum_{i=1}^n p_i z_i.Rn​(z,ρ)=p∈Pn​(ρ)sup​i=1∑n​pi​zi​.

For a random variable ZZZ with law PPP supported on [M0,M1][M_0,M_1][M0​,M1​], write M=M1−M0M=M_1-M_0M=M1​−M0​ and σ2=Var⁡P(Z)\sigma^2=\operatorname{Var}_P(Z)σ2=VarP​(Z). An independent sample Z1,…,ZnZ_1,\ldots,Z_nZ1​,…,Zn​ supplies the vector zzz. The paper describes Pn\mathcal P_nPn​ through a ϕ\phiϕ-divergence from the empirical distribution, with ϕ(t)=12(t−1)2\phi(t)=\tfrac12(t-1)^2ϕ(t)=21​(t−1)2; its finite maximization problem (8) is the weight-vector form used here. The preprint, pp. 2 and 5–7 fixes these conventions.

Formalization targets

Deterministic bound

For every sample in [M0,M1][M_0,M_1][M0​,M1​], the robust value lies between the empirical mean plus a corrected variance penalty and the full penalty:

(2ρsn2n−2Mρn)+≤Rn(z,ρ)−zˉ≤2ρsn2n.\left(\sqrt{\frac{2\rho s_n^2}{n}}-\frac{2M\rho}{n}\right)_+\le R_n(z,\rho)-\bar z\le\sqrt{\frac{2\rho s_n^2}{n}}.(n2ρsn2​​​−n2Mρ​)+​≤Rn​(z,ρ)−zˉ≤n2ρsn2​​​.

This is inequality (10). The correction is explicit, so this target records more than an asymptotic approximation.

Exact expansion

When σ2>0\sigma^2>0σ2>0 and the sample size obeys

n≥max⁡{5,M2σ2max⁡{8σ,44,44ρ}},n\ge\max\left\{5,\frac{M^2}{\sigma^2}\max\{8\sigma,44,44\rho\}\right\},n≥max{5,σ2M2​max{8σ,44,44ρ}},

the goal is the high-probability equality

Pr⁡{Rn(Z1:n,ρ)≠Zˉ+2ρsn2n}≤exp⁡(−nσ211M2).\Pr\left\{R_n(Z_{1:n},\rho)\ne\bar Z+\sqrt{\frac{2\rho s_n^2}{n}}\right\}\le\exp\left(-\frac{n\sigma^2}{11M^2}\right).Pr{Rn​(Z1:n​,ρ)=Zˉ+n2ρsn2​​​}≤exp(−11M2nσ2​).

This is Theorem 1's equality (11) with the missing ρ\rhoρ-dependent sample-size requirement supplied from the proof. The exact expansion is the mission goal; display (30), inequality (10), and Lemma A.2 form the milestone list, and the exact value under condition (9) is a further statement of the mission.

Significance

The deterministic result states how large the discrepancy between a convex robust risk and a variance penalty can be for any bounded sample. The equality says that, with the stated confidence, no discrepancy remains once the population variance and sample size make the penalty compatible with nonnegative probability weights. These are the numerical facts later sections need when they move from one loss variable to families of losses and minimizers. The claims and constants come from Theorem 1 and Section 2.1.

The paper develops arguments for these results, although its printed (11) needs the correction described below; the statements in this mission have no machine-checked proofs yet. The formalization work includes the finite χ² feasible set, its real supremum, exact handling of tied observations, empirical moments with the paper's normalization, and a product-law event for the probability estimate. The Samson concentration milestone is reusable for other bounded independent-coordinate models. Solvers can also contribute a different route to the corrected exact expansion; the goal concerns the statement, not one chosen argument.

Difficulty

Without the nonnegativity requirement on ppp, optimizing a linear function over the centered Euclidean ball gives the mean plus a standard-deviation term. The candidate weights can become negative when a sample coordinate is far below the mean, so that calculation alone cannot certify the robust value. Condition (9) records precisely when the candidate is feasible. The probability target then needs a quantitative guarantee that the sample variance is large enough often enough, with the stated exponential constant. A pointwise inequality for a fixed sample does not by itself yield that probability estimate. These are separate obligations in Section 2.1 and Appendix A.

Formalization scope

The sample is a function Fin n → ℝ; feasible weights have the same type. chiSqBall, robustSup, empMean, and empVar mirror equations (8) and the definitions on p. 6. Every theorem assumes n>0n>0n>0 and ρ≥0\rho\ge0ρ≥0, so the weight ball is nonempty and its real supremum is bounded. The high-probability theorem uses a probability measure PPP on the reals, supported on [M0,M1][M_0,M_1][M0​,M1​], and the independent product measure on Fin n → ℝ. Its conclusion bounds the measure of the event on which equality fails. The positive population variance hypothesis makes division by σ2\sigma^2σ2 and M2M^2M2 meaningful. The deterministic bounds include every sample in the interval and use x+=max⁡{x,0}x_+=\max\{x,0\}x+​=max{x,0}.

The paper prints the threshold without 44ρ44\rho44ρ in (11), but its Appendix A invokes the corresponding inequality, and the printed claim fails for sufficiently large ρ\rhoρ. The goal includes that term. The paper's route through Lemmas A.1 and A.4 contains misprinted lower-tail and moment claims, so those are not milestones. Lemma A.3's displayed (31b) is also omitted because its correction term has the wrong scaling; the corrected goal stands as a target to establish independently. These discrepancies are detailed in the local moderation notes and the pinned source, pp. 7 and 32–35.

No hypothesis may force the bad event to be empty, and the robust value must optimize over all feasible weights, not a selected optimizer. The supporting definitions are intended for reuse in later missions on uniform variance expansions. Contributions to the finite optimization facts, the concentration statement, and the probability goal are welcome.

Selected references

  • J. C. Duchi and H. Namkoong, Variance-based regularization with convex objectives, arXiv preprint arXiv:1610.02581v3, 2017. Pinned preprint.
8 thms1 active userReviewed
Bandit AlgorithmsConvex OptimizationOperations Research+1·Captain: mikedeng1

Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems IV: Online Stochastic Mirror Descent for Combinatorial Semi-BanditsTextbook

Motivation

Many sequential decision problems ask a learner to choose, round after round, a combination of items: a set of mmm ads out of ddd, a path in a network, a matching. After each choice the learner sees the loss of the items it used, not of those it did not. This is online combinatorial optimization with semi-bandit feedback. It contains the classical adversarial multi-armed bandit (choose one of ddd arms) and is a standard model in online advertising, routing and ranking.

Chapter 5 of Bubeck and Cesa-Bianchi's monograph arXiv:1204.5721v2 treats this problem with one algorithm, Online Stochastic Mirror Descent (OSMD). Every regret bound in the chapter comes from a single mirror-descent inequality, specialized through the choice of a convex "regularizer". The chapter's capstone, Theorem 5.7, shows that a polynomial regularizer gives pseudo-regret O(mdn)O(\sqrt{mdn})O(mdn​) with no logarithmic factor. For m=1m=1m=1 this is the minimax-optimal rate of the adversarial bandit, first attained by the INF strategy of Audibert and Bubeck (2009). The semi-bandit version is due to Audibert, Bubeck and Lugosi (2014).

Setting

Vectors live in Rd\mathbb R^dRd. The arm set is a nonempty C⊆{0,1}d\mathcal C\subseteq\{0,1\}^dC⊆{0,1}d with ∥v∥1=m\|v\|_1=m∥v∥1​=m for every v∈Cv\in\mathcal Cv∈C, and K=Conv(C)\mathcal K=\mathrm{Conv}(\mathcal C)K=Conv(C). An oblivious adversary fixes loss vectors ℓ1,…,ℓn∈[0,1]d\ell_1,\dots,\ell_n\in[0,1]^dℓ1​,…,ℓn​∈[0,1]d. In round ttt the learner plays a random arm vt∈Cv_t\in\mathcal Cvt​∈C, pays ℓt⊤vt\ell_t^\top v_tℓt⊤​vt​, and observes (ℓt(1)vt(1),…,ℓt(d)vt(d))(\ell_t(1)v_t(1),\dots,\ell_t(d)v_t(d))(ℓt​(1)vt​(1),…,ℓt​(d)vt​(d)). The pseudo-regret is

Rˉn=E∑t=1nℓt⊤vt−min⁡x∈K∑t=1nℓt⊤x.\bar R_n=\mathbb E\sum_{t=1}^n\ell_t^\top v_t-\min_{x\in\mathcal K}\sum_{t=1}^n\ell_t^\top x .Rˉn​=Et=1∑n​ℓt⊤​vt​−x∈Kmin​t=1∑n​ℓt⊤​x.

A Legendre function on Dˉ\bar DDˉ, for a nonempty open convex DDD, is a continuous F:Dˉ→RF:\bar D\to\mathbb RF:Dˉ→R that is strictly convex and C1C^1C1 on DDD and whose gradient norm tends to +∞+\infty+∞ at Dˉ∖D\bar D\setminus DDˉ∖D. Its Bregman divergence is DF(x,y)=F(x)−F(y)−(x−y)⊤∇F(y)D_F(x,y)=F(x)-F(y)-(x-y)^\top\nabla F(y)DF​(x,y)=F(x)−F(y)−(x−y)⊤∇F(y), and its Legendre–Fenchel transform is F∗(u)=sup⁡x∈Dˉ(x⊤u−F(x))F^*(u)=\sup_{x\in\bar D}(x^\top u-F(x))F∗(u)=supx∈Dˉ​(x⊤u−F(x)).

Online Mirror Descent with learning rate η>0\eta>0η>0 and vectors gtg_tgt​ starts at x1∈arg⁡min⁡KFx_1\in\arg\min_{\mathcal K}Fx1​∈argminK​F. It then sets ∇F(wt+1)=∇F(xt)−ηgt\nabla F(w_{t+1})=\nabla F(x_t)-\eta g_t∇F(wt+1​)=∇F(xt​)−ηgt​ and xt+1=arg⁡min⁡y∈KDF(y,wt+1)x_{t+1}=\arg\min_{y\in\mathcal K}D_F(y,w_{t+1})xt+1​=argminy∈K​DF​(y,wt+1​). OSMD uses a random estimate gt=ℓ~tg_t=\tilde\ell_tgt​=ℓ~t​ of the loss. In the semi-bandit case it plays vtv_tvt​ with E[vt∣xt]=xt\mathbb E[v_t\mid x_t]=x_tE[vt​∣xt​]=xt​ and uses

ℓ~t(i)=ℓt(i) vt(i)xt(i).(5.5)\tilde\ell_t(i)=\frac{\ell_t(i)\,v_t(i)}{x_t(i)}. \tag{5.5}ℓ~t​(i)=xt​(i)ℓt​(i)vt​(i)​.(5.5)

A 000-potential is a convex, C1C^1C1, increasing ψ:(−∞,a)→(0,∞)\psi:(-\infty,a)\to(0,\infty)ψ:(−∞,a)→(0,∞) with ψ(−∞)=0\psi(-\infty)=0ψ(−∞)=0, ψ(a−)=+∞\psi(a^-)=+\inftyψ(a−)=+∞ and ∫01∣ψ−1∣<∞\int_0^1|\psi^{-1}|<\infty∫01​∣ψ−1∣<∞. It defines the Legendre function Fψ(x)=∑i∫0xiψ−1(s) dsF_\psi(x)=\sum_i\int_0^{x_i}\psi^{-1}(s)\,dsFψ​(x)=∑i​∫0xi​​ψ−1(s)ds on [0,∞)d[0,\infty)^d[0,∞)d. With ψ=exp⁡\psi=\expψ=exp this is the negative entropy.

Formalization targets

Goal: Theorem 5.7 (p. 80)

For every 000-potential ψ\psiψ and non-negative unbiased estimates,

Rˉn≤sup⁡KFψ−Fψ(x1)η+η2∑t=1n∑i=1dE[ℓ~t(i)2(ψ−1)′(xt(i))].\bar R_n\le\frac{\sup_{\mathcal K}F_\psi-F_\psi(x_1)}{\eta}+\frac\eta2\sum_{t=1}^n\sum_{i=1}^d\mathbb E\left[\frac{\tilde\ell_t(i)^2}{(\psi^{-1})'(x_t(i))}\right].Rˉn​≤ηsupK​Fψ​−Fψ​(x1​)​+2η​t=1∑n​i=1∑d​E[(ψ−1)′(xt​(i))ℓ~t​(i)2​].

For ψ(x)=(−x)−q\psi(x)=(-x)^{-q}ψ(x)=(−x)−q with q>1q>1q>1, the estimate (5.5) and η=2q−1 m1−2/q/(n d1−2/q)\eta=\sqrt{\tfrac{2}{q-1}\,m^{1-2/q}/(n\,d^{1-2/q})}η=q−12​m1−2/q/(nd1−2/q)​,

Rˉn≤q2q−1 mdn,and  Rˉn≤22mdn  at q=2.\bar R_n\le q\sqrt{\tfrac{2}{q-1}\,mdn},\qquad\text{and }\ \bar R_n\le2\sqrt{2mdn}\ \text{ at }q=2.Rˉn​≤qq−12​mdn​,and  Rˉn​≤22mdn​  at q=2.

Milestones

  1. Lemma 5.1: F∗∗=FF^{**}=FF∗∗=F, ∇F∗=(∇F)−1\nabla F^*=(\nabla F)^{-1}∇F∗=(∇F)−1 on D∗D^*D∗, and DF(x,y)=DF∗(∇F(y),∇F(x))D_F(x,y)=D_{F^*}(\nabla F(y),\nabla F(x))DF​(x,y)=DF∗​(∇F(y),∇F(x)).
  2. Lemma 5.2: existence, uniqueness and the Pythagorean inequality of Bregman projections.
  3. Theorem 5.3: ∑tℓt(xt)−∑tℓt(x)≤F(x)−F(x1)η+1η∑tDF∗(∇F(xt)−η∇ℓt(xt),∇F(xt))\sum_t\ell_t(x_t)-\sum_t\ell_t(x)\le\frac{F(x)-F(x_1)}\eta+\frac1\eta\sum_tD_{F^*}(\nabla F(x_t)-\eta\nabla\ell_t(x_t),\nabla F(x_t))∑t​ℓt​(xt​)−∑t​ℓt​(x)≤ηF(x)−F(x1​)​+η1​∑t​DF∗​(∇F(xt​)−η∇ℓt​(xt​),∇F(xt​)).
  4. Theorem 5.5, linear losses, and its corrected general form.
  5. Lemma 5.3: FψF_\psiFψ​ is Legendre and DFψ∗(u,v)≤12∑iψ′(vi)(ui−vi)2D_{F_\psi^*}(u,v)\le\frac12\sum_i\psi'(v_i)(u_i-v_i)^2DFψ∗​​(u,v)≤21​∑i​ψ′(vi​)(ui​−vi​)2 for u≤vu\le vu≤v.
  6. Theorem 5.6: with the negative entropy, Rˉn≤2mdnln⁡(d/m)\bar R_n\le\sqrt{2mdn\ln(d/m)}Rˉn​≤2mdnln(d/m)​.

Significance

Theorem 5.7 is the sharpest semi-bandit bound in the monograph. It shows that removing the ln⁡(d/m)\sqrt{\ln(d/m)}ln(d/m)​ factor of the exponential-weights analysis (Theorem 5.6) is a matter of the regularizer, not of a new algorithm. The same OSMD template gives the Euclidean-ball bound of Theorem 5.8 and is reused for bandit convex optimization in Chapter 6. Lemma 5.1, Lemma 5.2 and Theorem 5.3 are the standard mirror-descent toolkit, used throughout online learning and optimization.

All results of the chapter are proved in the book. Lemmas 5.1 and 5.2 are cited from Cesa-Bianchi and Lugosi (2006). None of them is formalized on Prove2Me. The published mirror-descent bound of Bandit Algorithms XII treats linear losses with a comparator inside DDD and Euclidean-space vectors; it is not Theorem 5.3. The mission adds a machine-checked version of the whole chain, from Legendre duality to the explicit constant q2mdn/(q−1)q\sqrt{2mdn/(q-1)}q2mdn/(q−1)​, with two of the printed statements corrected (below).

Difficulty

The pathwise mirror-descent inequality is a telescoping argument, but several of its steps rest on convex analysis that Mathlib does not package. One is the existence and interior location of Bregman projections onto a set that touches the boundary of DDD. Another is the differentiability of F∗F^*F∗ on the open dual space and the identity ∇F∗=(∇F)−1\nabla F^*=(\nabla F)^{-1}∇F∗=(∇F)−1. A third is the closed form of Fψ∗F_\psi^*Fψ∗​ for a potential defined through an improper integral of ψ−1\psi^{-1}ψ−1.

The probabilistic step is not a martingale argument. Only conditioning on the current iterate xtx_txt​ is available. The estimate (5.5) divides by xt(i)x_t(i)xt​(i), so its integrability and unbiasedness have to be derived from the fact that the iterates stay in the open orthant. Finally, the explicit constant requires a Hölder step, ∑ix1(i)1−1/q≤m(q−1)/qd1/q\sum_ix_1(i)^{1-1/q}\le m^{(q-1)/q}d^{1/q}∑i​x1​(i)1−1/q≤m(q−1)/qd1/q, and the matching bound ∑ixt(i)1/q≤m1/qd1−1/q\sum_ix_t(i)^{1/q}\le m^{1/q}d^{1-1/q}∑i​xt​(i)1/q≤m1/qd1−1/q.

Formalization scope

Vectors are Fin d → ℝ. The arm set is a Set of 0/10/10/1 vectors with coordinate sum mmm, and K\mathcal KK is convexHull ℝ C. Rounds are t=1,…,nt=1,\dots,nt=1,…,n, sums run over Finset.Icc 1 n, and index 000 is unused. A randomized run is a family of measurable processes xt,vt,ℓ~t,wtx_t, v_t, \tilde\ell_t, w_txt​,vt​,ℓ~t​,wt​ on a probability space, with the deterministic OMD recursion holding on every sample path. E[⋅∣xt]\mathbb E[\cdot\mid x_t]E[⋅∣xt​] is the coordinatewise conditional expectation given σ(xt)\sigma(x_t)σ(xt​), which is exactly what the book's proofs use. Losses are oblivious, so Rˉn≤B\bar R_n\le BRˉn​≤B is stated as "for every x∈Kx\in\mathcal Kx∈K, E∑tℓt⊤vt−∑tℓt⊤x≤B\mathbb E\sum_t\ell_t^\top v_t-\sum_t\ell_t^\top x\le BE∑t​ℓt⊤​vt​−∑t​ℓt⊤​x≤B". F∗F^*F∗ is valued in EReal, and DF∗D_{F^*}DF∗​ is evaluated only on the open dual space, where F∗F^*F∗ is finite. Wherever an expectation of a possibly non-integrable quantity appears on a right-hand side, its integrability is assumed: the book's bound is then +∞+\infty+∞ and trivial, while Lean's integral would be 000.

Corrections and instantiations, each labelled in the item's Formalization Note:

  • Theorem 5.7, corrected misprint. The book prints η=2q−1m1−2/qd1−2/q\eta=\sqrt{\frac2{q-1}\frac{m^{1-2/q}}{d^{1-2/q}}}η=q−12​d1−2/qm1−2/q​​. The proof (p. 81) gives the stated bound only for η=2q−1m1−2/qn d1−2/q\eta=\sqrt{\frac2{q-1}\frac{m^{1-2/q}}{n\,d^{1-2/q}}}η=q−12​nd1−2/qm1−2/q​​, which is stated. At q=2q=2q=2 this is η=2/n\eta=\sqrt{2/n}η=2/n​.
  • Theorem 5.5, corrected misprint. In the first bound the book prints E[∥xt−x~t∥ ∥g~t∥∗]\mathbb E[\|x_t-\tilde x_t\|\,\|\tilde g_t\|_*]E[∥xt​−x~t​∥∥g~​t​∥∗​]. That statement fails for ℓt(x)=x2\ell_t(x)=x^2ℓt​(x)=x2 on [−1,1][-1,1][−1,1] with F=x2/2F=x^2/2F=x2/2 and x~t=±1\tilde x_t=\pm1x~t​=±1. The version stated uses ∥∇ℓt(x~t)∥∗\|\nabla\ell_t(\tilde x_t)\|_*∥∇ℓt​(x~t​)∥∗​, as the proof's first inequality does. The linear-loss bound is stated as printed.
  • Lemma 5.2. "For all z∈K∩Dz\in K\cap Dz∈K∩D" is read as "for the projection zzz", which lies in K∩DK\cap DK∩D.
  • Hypotheses made explicit: q>1q>1q>1; non-negativity of the estimates in Theorem 5.6 (used in its proof); unbiasedness E[ℓ~t∣xt]=ℓt\mathbb E[\tilde\ell_t\mid x_t]=\ell_tE[ℓ~t​∣xt​]=ℓt​ in the general parts of Theorems 5.6 and 5.7; K∩(0,∞)d≠∅\mathcal K\cap(0,\infty)^d\ne\emptysetK∩(0,∞)d=∅ (OMD's requirement K∩D≠∅K\cap D\ne\emptysetK∩D=∅); a subgradient selection as an explicit input.
  • Theorem 5.6's particular bound uses the book's η=2mndln⁡dm\eta=\sqrt{\frac{2m}{nd}\ln\frac dm}η=nd2m​lnmd​​ as printed. There are no O(·) constants in the chapter's statements.

A trivializing formalization would let η\etaη, xtx_txt​ or the estimate be junk values: an OSMD step at η=0\eta=0η=0, a Lean division x/0=0x/0=0x/0=0, or a regret written as a real infimum over an unbounded set. Here every run is the book's algorithm on the open orthant, and each bound is stated against every comparator in K\mathcal KK.

Reusable beyond this mission: the Legendre/Bregman layer, the OMD run predicate and the ω\omegaω-potential layer. Proofs of Lemmas 5.1 and 5.2 in this generality would be welcome additions to the library.

Selected references

  • S. Bubeck, N. Cesa-Bianchi, Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems, Foundations and Trends in Machine Learning 5(1), 2012; arXiv:1204.5721v2. https://arxiv.org/abs/1204.5721
  • N. Cesa-Bianchi, G. Lugosi, Prediction, Learning, and Games, Cambridge University Press, 2006. https://doi.org/10.1017/CBO9780511546921
  • J.-Y. Audibert, S. Bubeck, Regret bounds and minimax policies under partial monitoring, Journal of Machine Learning Research 11, 2010. https://www.jmlr.org/papers/v11/audibert10a.html
  • J.-Y. Audibert, S. Bubeck, G. Lugosi, Regret in online combinatorial optimization, Mathematics of Operations Research 39(1), 2014. https://doi.org/10.1287/moor.2013.0598
12 thms1 active userReviewed
Bandit AlgorithmsOperations Research·Captain: mikedeng1

Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems II: High-Probability and Expected Regret of Exp3.PTextbook

Motivation

In the adversarial (non-stochastic) multi-armed bandit problem a forecaster repeatedly chooses one of KKK actions while an opponent sets the rewards, and only the reward of the chosen action is revealed. The model was proposed as a way of playing an unknown repeated game: Baños (1968) studied the repeated game in which the player observes only its own payoff, which is exactly the bandit problem against an opponent who reacts to the player's past moves. It is the basic model of online decision making under partial feedback without statistical assumptions, and it underlies regret minimization in games, adversarial routing and online advertising. Chapter 3 of Bubeck and Cesa-Bianchi's monograph (arXiv:1204.5721v2) collects its fundamental results: the Exp3 forecaster of Auer, Cesa-Bianchi, Freund and Schapire (SIAM J. Comput. 2002), its high-probability variant Exp3.P, and the nK\sqrt{nK}nK​ minimax lower bound.

Setting

There are K≥2K \ge 2K≥2 arms and rounds t=1,2,…,nt = 1, 2, \dots, nt=1,2,…,n. At each round an adversary assigns a gain gi,t∈[0,1]g_{i,t} \in [0,1]gi,t​∈[0,1] to every arm iii; the forecaster picks an arm ItI_tIt​, possibly at random, and observes only gIt,tg_{I_t,t}gIt​,t​. The adversary may be non-oblivious (adaptive): gi,t=gi,t(I1,…,It−1)g_{i,t} = g_{i,t}(I_1,\dots,I_{t-1})gi,t​=gi,t​(I1​,…,It−1​) may depend on the forecaster's past actions. A forecaster rule maps the past actions to a probability vector ptp_tpt​ on the arms, and a run is a sequence of random arms with It∼ptI_t \sim p_tIt​∼pt​ given the past. The regret is the random variable

Rn=max⁡i=1,…,K∑t=1ngi,t−∑t=1ngIt,t,R_n = \max_{i=1,\dots,K}\sum_{t=1}^n g_{i,t} - \sum_{t=1}^n g_{I_t,t},Rn​=i=1,…,Kmax​t=1∑n​gi,t​−t=1∑n​gIt​,t​,

and, in the loss version ℓi,t∈[0,1]\ell_{i,t} \in [0,1]ℓi,t​∈[0,1], the pseudo-regret is R‾n=E∑tℓIt,t−min⁡iE∑tℓi,t\overline R_n = \mathbb E\sum_t \ell_{I_t,t} - \min_i \mathbb E\sum_t \ell_{i,t}Rn​=E∑t​ℓIt​,t​−mini​E∑t​ℓi,t​. Since the maximum sits inside the expectation, R‾n≤ERn\overline R_n \le \mathbb E R_nRn​≤ERn​ in the gain version, and against an adaptive adversary the two can differ.

Exp3 draws ItI_tIt​ from exponential weights pi,t+1∝exp⁡(−ηtL~i,t)p_{i,t+1} \propto \exp(-\eta_t \tilde L_{i,t})pi,t+1​∝exp(−ηt​L~i,t​) of importance-weighted cumulative loss estimates L~i,t=∑s≤tℓi,s1Is=i/pi,s\tilde L_{i,t} = \sum_{s \le t} \ell_{i,s}\mathbb 1_{I_s = i}/p_{i,s}L~i,t​=∑s≤t​ℓi,s​1Is​=i​/pi,s​. Exp3.P uses biased gain estimates g~i,t=(gi,t1It=i+β)/pi,t\tilde g_{i,t} = (g_{i,t}\mathbb 1_{I_t=i} + \beta)/p_{i,t}g~​i,t​=(gi,t​1It​=i​+β)/pi,t​ and mixes in the uniform distribution:

pi,t+1=(1−γ)exp⁡(ηG~i,t)∑kexp⁡(ηG~k,t)+γK,G~i,t=∑s=1tg~i,s.p_{i,t+1} = (1-\gamma)\frac{\exp(\eta\tilde G_{i,t})}{\sum_k \exp(\eta \tilde G_{k,t})} + \frac{\gamma}{K}, \qquad \tilde G_{i,t} = \sum_{s=1}^t \tilde g_{i,s}.pi,t+1​=(1−γ)∑k​exp(ηG~k,t​)exp(ηG~i,t​)​+Kγ​,G~i,t​=s=1∑t​g~​i,s​.

Formalization targets

Goal: Theorem 3.3 (expected regret of Exp3.P)

With β=ln⁡K/(nK)\beta = \sqrt{\ln K/(nK)}β=lnK/(nK)​, η=0.95ln⁡K/(nK)\eta = 0.95\sqrt{\ln K/(nK)}η=0.95lnK/(nK)​, γ=1.05Kln⁡K/n\gamma = 1.05\sqrt{K\ln K/n}γ=1.05KlnK/n​, against every adaptive adversary,

ERn≤5.15nKln⁡K+nKln⁡K.\mathbb E R_n \le 5.15\sqrt{nK\ln K} + \sqrt{\frac{nK}{\ln K}}.ERn​≤5.15nKlnK​+lnKnK​​.

Milestones

  • Lemma 3.1: for β∈(0,1]\beta \in (0,1]β∈(0,1] and a fixed arm iii, with probability at least 1−δ1-\delta1−δ, ∑tgi,t≤∑tg~i,t+ln⁡(δ−1)/β\sum_t g_{i,t} \le \sum_t \tilde g_{i,t} + \ln(\delta^{-1})/\beta∑t​gi,t​≤∑t​g~​i,t​+ln(δ−1)/β.
  • Eq. (3.12): if γ≤1/2\gamma \le 1/2γ≤1/2 and (1+β)Kη≤γ(1+\beta)K\eta \le \gamma(1+β)Kη≤γ, then with probability at least 1−δ1-\delta1−δ,
Rn≤βnK+γn+(1+β)ηKn+ln⁡(Kδ−1)β+ln⁡Kη.R_n \le \beta nK + \gamma n + (1+\beta)\eta Kn + \frac{\ln(K\delta^{-1})}{\beta} + \frac{\ln K}{\eta}.Rn​≤βnK+γn+(1+β)ηKn+βln(Kδ−1)​+ηlnK​.
  • Theorem 3.2: with β=ln⁡(Kδ−1)/(nK)\beta = \sqrt{\ln(K\delta^{-1})/(nK)}β=ln(Kδ−1)/(nK)​, Rn≤5.15nKln⁡(Kδ−1)R_n \le 5.15\sqrt{nK\ln(K\delta^{-1})}Rn​≤5.15nKln(Kδ−1)​ (3.10); with β=ln⁡K/(nK)\beta = \sqrt{\ln K/(nK)}β=lnK/(nK)​, Rn≤nK/ln⁡K ln⁡(δ−1)+5.15nKln⁡KR_n \le \sqrt{nK/\ln K}\,\ln(\delta^{-1}) + 5.15\sqrt{nK\ln K}Rn​≤nK/lnK​ln(δ−1)+5.15nKlnK​ (3.11), each with probability at least 1−δ1-\delta1−δ.
  • Theorem 3.1: Exp3 with η=2ln⁡K/(nK)\eta = \sqrt{2\ln K/(nK)}η=2lnK/(nK)​ has R‾n≤2nKln⁡K\overline R_n \le \sqrt{2nK\ln K}Rn​≤2nKlnK​ (3.2); with ηt=ln⁡K/(tK)\eta_t = \sqrt{\ln K/(tK)}ηt​=lnK/(tK)​, R‾n≤2nKln⁡K\overline R_n \le 2\sqrt{nK\ln K}Rn​≤2nKlnK​ (3.3).
  • Lemma 3.2 and Theorem 3.4: for n≥K≥2n \ge K \ge 2n≥K≥2 and every forecaster there is a Bernoulli instance with max⁡iE∑tYi,t−E∑tYIt,t≥nK/20\max_i \mathbb E\sum_t Y_{i,t} - \mathbb E\sum_t Y_{I_t,t} \ge \sqrt{nK}/20maxi​E∑t​Yi,t​−E∑t​YIt​,t​≥nK​/20.

Significance

The goal bounds the expected regret, not the pseudo-regret, against an opponent that adapts to the forecaster's randomized past choices. A pseudo-regret bound says nothing about ERn\mathbb E R_nERn​ in that setting, and the book obtains the expected-regret bound by first proving a high-probability bound valid at every confidence level, (3.11), and integrating its tail. Together with Theorem 3.4 the chapter shows that nK\sqrt{nK}nK​ is the minimax rate of adversarial bandits up to a ln⁡K\sqrt{\ln K}lnK​ factor. Lemma 3.1, the concentration of biased importance-weighted estimates, holds for any forecaster rule and is the step that turns exponential weights into a high-probability guarantee.

All results are proved in the book. On the formal side, the platform has the pseudo-regret bound of Exp3 against an oblivious adversary (a fixed reward table, Bandit Algorithms V) and an Exp3-IX high-probability bound; it has no Exp3.P, no regret bound against adaptive adversaries and no Bernoulli nK/20\sqrt{nK}/20nK​/20 lower bound. This mission adds an explicit model of adaptive adversaries and randomized forecaster runs, and the chapter's statements with the book's exact constants.

Difficulty

Against an adaptive adversary the gains are random and depend on the forecaster's own past draws, so the argument used for a fixed reward table (take expectations of an inequality that holds for every fixed sequence) does not control ERn\mathbb E R_nERn​: the maximum over arms does not commute with the expectation. Unbiased estimates do not help either, because the variance of ℓi,t/pi,t\ell_{i,t}/p_{i,t}ℓi,t​/pi,t​ is of order 1/pi,t1/p_{i,t}1/pi,t​, which can be arbitrarily large; even with uniform mixing at rate n−1/2n^{-1/2}n−1/2 the cumulative variance is of order n3/2n^{3/2}n3/2. The bias β\betaβ and the mixing γ\gammaγ have to be tuned jointly so that the estimate concentrates while the exponential-weights analysis survives, and the constants 0.950.950.95, 1.051.051.05 and 5.155.155.15 come out of that tuning. The lower bound needs an information-theoretic comparison of a forecaster's behaviour on K+1K+1K+1 Bernoulli instances, against forecasters that may be randomized.

Formalization scope

Arms are Fin K with K≥2K \ge 2K≥2; rounds are numbered 1,…,n1,\dots,n1,…,n; logarithms are natural. Action sequences are functions N→\mathbb N \toN→ Fin K whose entry 000 is ignored. An adversary is a structure holding values in [0,1][0,1][0,1] that may depend on the past actions only (gains for Exp3.P, losses for Exp3); a randomized adversary with independent external randomness reduces to this case by conditioning. A run of a forecaster rule ppp on a probability space is pinned down by the cylinder identity P(I1=h1,…,It=ht)=P(I1=h1,…,It−1=ht−1) pt(h)(ht)\mathbb P(I_1 = h_1,\dots,I_t = h_t) = \mathbb P(I_1=h_1,\dots,I_{t-1}=h_{t-1})\,p_t(h)(h_t)P(I1​=h1​,…,It​=ht​)=P(I1​=h1​,…,It−1​=ht−1​)pt​(h)(ht​), which determines the law of (I1,…,In)(I_1,\dots,I_n)(I1​,…,In​). "With probability at least 1−δ1-\delta1−δ" is P(event)≥1−δ\mathbb P(\text{event}) \ge 1-\deltaP(event)≥1−δ for δ∈(0,1)\delta \in (0,1)δ∈(0,1), and ERn\mathbb E R_nERn​ is the Bochner integral of the bounded, measurable regret. The lower bounds use a stochastic model in which the forecaster sees past actions and the rewards of the played arms, and rewards are i.i.d. product Bernoulli.

Constants and conventions:

  • Every constant is the book's exact one: 0.950.950.95, 1.051.051.05, 5.155.155.15, 1/201/201/20. No O(⋅)O(\cdot)O(⋅) is involved.
  • Exp3.P with 1.05Kln⁡K/n>11.05\sqrt{K\ln K/n} > 11.05KlnK/n​>1 is outside the box's range γ∈[0,1]\gamma \in [0,1]γ∈[0,1]; its vector can then have negative entries, and if it does on a history of positive probability no run exists. This happens only when n<1.11 Kln⁡Kn < 1.11\,K\ln Kn<1.11KlnK, where the printed bounds already follow from Rn≤nR_n \le nRn​≤n, so the statements are true there whether or not a run exists.
  • Corrected misprints: the Exp3 box's ℓ~i,s\tilde\ell_{i,s}ℓ~i,s​ is ℓ~i,t\tilde\ell_{i,t}ℓ~i,t​; the sign in (3.16) is the box's exp⁡(+ηG~)\exp(+\eta\tilde G)exp(+ηG~); the proof of (3.10) says the bound is trivial "if n≥5.15⋯n \ge 5.15\sqrt{\cdots}n≥5.15⋯​", which should be n≤n \len≤. The statements carry no lower bound on nnn.
  • Added standing hypotheses: K≥2K \ge 2K≥2 everywhere, n≥Kn \ge Kn≥K in Theorem 3.4 (from the protocol box, p. 6; Theorem 3.4 is false without it), β>0\beta > 0β>0 and pi,t>0p_{i,t} > 0pi,t​>0 in Lemma 3.1.
  • Theorem 3.4 is stated as "for every forecaster there is a Bernoulli instance with regret at least nK/20\sqrt{nK}/20nK​/20", which implies the book's inf⁡sup⁡\inf\supinfsup (3.18).

A trivializing formalization is ruled out: the forecasters are fixed rules of the observed history drawn with fresh randomness, the adversary is not restricted to a fixed sequence, and the lower bounds quantify over all forecasters and exhibit the instance.

Welcome contributions: a reusable construction of runs (existence of a probability space carrying a run for every rule), the supermartingale form of Lemma 3.1, the exponential-weights potential argument, a tail-integration lemma EW≤∫01δ−1P(W>ln⁡δ−1) dδ\mathbb E W \le \int_0^1 \delta^{-1}\mathbb P(W > \ln\delta^{-1})\,d\deltaEW≤∫01​δ−1P(W>lnδ−1)dδ, and a KL/Pinsker comparison for bandit runs.

Selected references

  • S. Bubeck, N. Cesa-Bianchi, Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems, Foundations and Trends in Machine Learning 5(1), 2012. arXiv:1204.5721v2, doi:10.1561/2200000024
  • P. Auer, N. Cesa-Bianchi, Y. Freund, R. E. Schapire, The nonstochastic multiarmed bandit problem, SIAM Journal on Computing 32(1), 2002. doi:10.1137/S0097539701398375
  • J.-Y. Audibert, S. Bubeck, Regret bounds and minimax policies under partial monitoring, Journal of Machine Learning Research 11, 2010. jmlr.org/papers/v11/audibert10a
  • N. Cesa-Bianchi, G. Lugosi, Prediction, Learning, and Games, Cambridge University Press, 2006. doi:10.1017/CBO9780511546921
13 thms1 active userReviewed
PreviousPage 10 of 12Next

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me