Prove2Me
Navigate
DiscoverFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

Reinforcement Learning

31 missions · 9 completed

Missions

Open22Completed9All31
🏆Completed
Machine LearningStatistics·Captain: mikedeng1

Foundations of Reinforcement Learning III: Structured Bandits and the Decision-Estimation CoefficientTextbook

Motivation

Every algorithm in the first three chapters of Foster and Rakhlin's Foundations of Reinforcement Learning and Interactive Decision Making — ε-Greedy and UCB for the multi-armed bandit, Inverse Gap Weighting and SquareCB for contextual bandits — is a special case of the same two-step recipe: estimate a model of the world with an online regression oracle, then convert the estimate into a decision that trades exploration against exploitation. Chapter 4 asks whether this recipe can be made generic: given any structured decision-making problem, specified only by a function class FFF and a decision space Π\PiΠ, is there a single quantity that governs the best achievable regret, the way A/γ\sqrt{A/\gamma}A/γ​ governs the multi-armed bandit and d/γ\sqrt{d/\gamma}d/γ​ governs the linear bandit? The chapter's answer is the Decision-Estimation Coefficient (DEC), introduced by Foster, Kakade, Qian, and Rakhlin [40] as a complexity measure that both upper- and lower-bounds achievable regret for a general decision-making protocol, unifying results that were previously proved from scratch, case by case, for each structured setting. This mission formalizes the chapter's central upper bound (Proposition 13) together with the machinery that makes it computable in two concrete cases — the multi-armed bandit (Proposition 14) and the linear bandit (Propositions 16–17).

Setting

Fix a finite decision space Π\PiΠ and a class F⊆RΠF \subseteq \mathbb{R}^\PiF⊆RΠ of candidate mean-reward functions, with a ground-truth f⋆∈Ff^\star \in Ff⋆∈F (realizability). Over TTT rounds, at each round ttt the learner observes an estimate f^t\hat f_tf^​t​ produced by an online regression oracle, plays a decision distribution pt∈Δ(Π)p_t \in \Delta(\Pi)pt​∈Δ(Π) (possibly depending on f^t\hat f_tf^​t​ and the history), and the regret is

Reg:=∑t=1Tf⋆(π⋆)−∑t=1TEπ∼pt[f⋆(π)],\mathrm{Reg} := \sum_{t=1}^T f^\star(\pi^\star) - \sum_{t=1}^T \mathbb{E}_{\pi \sim p_t}[f^\star(\pi)],Reg:=t=1∑T​f⋆(π⋆)−t=1∑T​Eπ∼pt​​[f⋆(π)],

where π⋆=arg⁡max⁡πf⋆(π)\pi^\star = \arg\max_\pi f^\star(\pi)π⋆=argmaxπ​f⋆(π). The oracle's cumulative estimation error is assumed bounded: ∑t=1TEπ∼pt[(f^t(π)−f⋆(π))2]≤EstSq(F,T,δ)\sum_{t=1}^T \mathbb{E}_{\pi \sim p_t}[(\hat f_t(\pi) - f^\star(\pi))^2] \le \mathrm{EstSq}(F,T,\delta)∑t=1T​Eπ∼pt​​[(f^​t​(π)−f⋆(π))2]≤EstSq(F,T,δ) with probability at least 1−δ1-\delta1−δ (Definition 7). Writing πf:=arg⁡max⁡πf(π)\pi_f := \arg\max_\pi f(\pi)πf​:=argmaxπ​f(π), the DEC game value at a reference model f^\hat ff^​ and scale γ>0\gamma > 0γ>0 is the min-max quantity

decγ(F,f^):=min⁡p∈Δ(Π)max⁡f∈F  Eπ∼p[f(πf)−f(π)−γ(f(π)−f^(π))2],\mathrm{dec}_\gamma(F, \hat f) := \min_{p \in \Delta(\Pi)} \max_{f \in F} \; \mathbb{E}_{\pi \sim p}\bigl[f(\pi_f) - f(\pi) - \gamma(f(\pi) - \hat f(\pi))^2\bigr],decγ​(F,f^​):=p∈Δ(Π)min​f∈Fmax​Eπ∼p​[f(πf​)−f(π)−γ(f(π)−f^​(π))2],

and the DEC of FFF itself is decγ(F):=sup⁡f^∈co(F)decγ(F,f^)\mathrm{dec}_\gamma(F) := \sup_{\hat f \in \mathrm{co}(F)} \mathrm{dec}_\gamma(F, \hat f)decγ​(F):=supf^​∈co(F)​decγ​(F,f^​). The Estimation-to-Decisions (E2D) algorithm plays, at each round, a ptp_tpt​ certifying (i.e. attaining or beating) the value of this min-max game at f^t\hat f_tf^​t​.

Formalization targets

Goal — Proposition 13 (E2D regret bound)

Reg≤decγ(F)⋅T+γ⋅EstSq(F,T,δ)\mathrm{Reg} \le \mathrm{dec}_\gamma(F) \cdot T + \gamma \cdot \mathrm{EstSq}(F, T, \delta)Reg≤decγ​(F)⋅T+γ⋅EstSq(F,T,δ)

with probability at least 1−δ1-\delta1−δ, for any exploration parameter γ>0\gamma > 0γ>0. This is the weakest stable statement the chapter proves about E2D: it holds for an arbitrary function class and an arbitrary regression oracle, with no structural assumption on FFF beyond realizability, and the chapter's later sections instantiate it rather than strengthen it.

Milestones

  • Lemma 9 (Decoupling), general form: for any distribution ν\nuν over a finite model class and any fˉ\bar ffˉ​, Ef∼ν[f(πf)−fˉ(πf)]≤A⋅Ef∼νEπ∼p[(f(π)−fˉ(π))2]\mathbb{E}_{f\sim\nu}[f(\pi_f) - \bar f(\pi_f)] \le \sqrt{A \cdot \mathbb{E}_{f\sim\nu}\mathbb{E}_{\pi\sim p}[(f(\pi)-\bar f(\pi))^2]}Ef∼ν​[f(πf​)−fˉ​(πf​)]≤A⋅Ef∼ν​Eπ∼p​[(f(π)−fˉ​(π))2]​ — the estimation-to-decisions bridge the whole chapter's approach rests on, decoupling the model index from the played decision.
  • Proposition 14 (IGW minimizes the DEC): for the multi-armed bandit (Π=[A]\Pi=[A]Π=[A], F=RAF=\mathbb{R}^AF=RA), Inverse Gap Weighting is the exact minimizer of the DEC game, giving decγ(F)=(A−1)/(4γ)\mathrm{dec}_\gamma(F) = (A-1)/(4\gamma)decγ​(F)=(A−1)/(4γ) — the first concrete computation of an abstract quantity, recovering Chapter 3's rate from Proposition 13 alone.
  • Proposition 16 (G-optimal design): existence, for any compact full-dimensional-span set Z⊆RdZ \subseteq \mathbb{R}^dZ⊆Rd, of a distribution ppp with sup⁡z∈Z⟨Σp−1z,z⟩≤d\sup_{z\in Z}\langle \Sigma_p^{-1}z,z\rangle \le dsupz∈Z​⟨Σp−1​z,z⟩≤d — the classical convex-analysis primitive Proposition 17 needs.
  • Proposition 17 (DEC for linear bandits): combining the G-optimal design with inverse gap weighting gives decγ(F)≲d/γ\mathrm{dec}_\gamma(F) \lesssim d/\gammadecγ​(F)≲d/γ for the linear bandit function class, leading via Proposition 13 to a dT\sqrt{dT}dT​ regret bound.

Significance

The Decision-Estimation Coefficient is, in the book's own words, "the main result" of this line of work: Foster, Kakade, Qian, and Rakhlin [40] show it is not merely an upper bound but (in a suitable localized form, developed further in Chapter 6) a tight characterization of the minimax regret for structured bandits and, more generally, for the interactive decision-making protocol the rest of the book studies. Proposition 13 is the mechanism that makes this useful in practice: it reduces regret analysis for a new structured problem to a single, purely convex-analytic computation of decγ(F)\mathrm{dec}_\gamma(F)decγ​(F), in place of a bespoke exploration argument. Propositions 14–17 are the demonstration that this reduction is not vacuous — they recompute, via the DEC alone, the two rates (multi-armed and linear bandit) that earlier chapters of the book derived by direct, setting-specific arguments, and the match is exact. Formalizing this chapter therefore captures the book's unifying abstraction itself, not just one more instance of it. No formalization of the Decision-Estimation Coefficient, in any form, currently exists on the platform (see Formalization scope).

Difficulty

The obvious formalization mistake is to state Proposition 13's conclusion with decγ(F)\mathrm{dec}_\gamma(F)decγ​(F) left as an unconstrained free real-number parameter satisfying only the inequality the theorem asserts — a formalization under which the "theorem" would be a triviality about an arbitrary real number, since nothing about the actual min-max game would ever be checked. The chapter's content is precisely the opposite: that this specific minimax quantity can be computed (Proposition 14) or bounded via a concrete strategy (Proposition 17), and — as Chapter 6 shows for a lower bound outside this chunk's scope — that no smaller quantity would do. A second difficulty is proof-theoretic rather than notational: the book's own proof of Proposition 13 bounds regret by an unconstrained supremum over all reference functions f^:Π→R\hat f : \Pi \to \mathbb{R}f^​:Π→R, and only identifies this with the official, co(F)\mathrm{co}(F)co(F)-restricted decγ(F)\mathrm{dec}_\gamma(F)decγ​(F) of Eq. (4.16) via Proposition 24 — a fact stated on p. 80, outside this chapter's numbered range, whose own proof the book defers to an exercise. A formalization that quietly imports Proposition 24 to close this gap would rest the goal theorem on an unverified fact; this mission instead states the hypothesis the book's own text uses to motivate restricting to co(F)\mathrm{co}(F)co(F) in the first place (online estimation algorithms produce f^t∈co(F)\hat f_t \in \mathrm{co}(F)f^​t​∈co(F)), so the goal is faithful to what is actually established within the chapter's own pages.

Formalization scope

Every item fixes a finite decision space (Fin A, Fin n, or a generic Fintype S) and states the DEC as the literal sInf-of-sSup transcription of the min-max game (Eqs. (4.15)–(4.16)), never as an opaque bound — this is the trivializing formalization the chunk's own reading of the chapter rules out (see Difficulty). piStar : (S → ℝ) → S is a hypothesized global maximizer selector throughout, constrained to be a genuine argmax only on the function class in scope (F or Set.univ), matching how the book treats πf\pi_fπf​ as a fixed but arbitrary tie-breaking choice. The goal theorem (Proposition 13) adds the explicit hypothesis hfhat : ∀ t, fhat t ∈ convexHull ℝ F, replacing an appeal to the out-of-range Proposition 24 (see Difficulty); this is the one place this mission's statement is not a line-by-line transcription of the book's own displayed proof steps, and it is recorded here and in MODERATION_NOTES.md. Proposition 14's and Proposition 17's ≲\lesssim≲ are replaced by the explicit constants the book's own proofs establish ((A−1)/(4γ)(A-1)/(4\gamma)(A−1)/(4γ) exactly, and (4d+1)/(2γ)(4d+1)/(2\gamma)(4d+1)/(2γ) respectively — the latter obtained by summing the three terms the proof of Proposition 17 isolates). Proposition 14's Lean statement splits the book's single equality decγ(F,f^)=(A−1)/(4γ)\mathrm{dec}_\gamma(F,\hat f) = (A-1)/(4\gamma)decγ​(F,f^​)=(A−1)/(4γ) into an upper bound on the literal decGf, a lower bound restricted to full-support distributions, and IGW's own exact game value, because the book's min over the whole simplex is not provable as a literal Lean equality: a distribution with a zero-weight arm makes the inner supremum genuinely unbounded, and Lean's total Real.sSup returns a junk value smaller than (A−1)/(4γ)(A-1)/(4\gamma)(A−1)/(4γ) there (caught in moderation, MODERATION_NOTES.md); the three-conjunct statement recovers exactly the book's real content without asserting that false literal equality. Lemma 9 is restated inside FoundationsRL.Structured rather than imported from the Chapter 2 mission, since draft items across chunks cannot import one another; its source citation still points to its original location (p. 32). Proposition 16 is not drafted: the platform's existing BanditAlgorithm.kiefer_wolfowitz_equivalence (Lattimore & Szepesvári, Theorem 21.1) states the identical existence claim — compact set with full-dimensional span, a design with G-value at most ddd — as one clause of a larger equivalence, and is reused as a reference item rather than redrafted. Proposition 22 (primal/dual DEC equivalence, §4.4) is deliberately excluded: the book states it "under mild regularity conditions" it does not pin down in the statement itself, which is exactly the kind of unquantified hypothesis this series' faithfulness standard excludes from a goal or milestone. Contributions extending this mission with Chapter 6's lower bound (matching decγ(F)\mathrm{dec}_\gamma(F)decγ​(F) from below, establishing tightness) or with a formalization of Proposition 24 itself (removing this mission's hfhat hypothesis) are welcome.

Selected references

  • D. Foster, S. Kakade, J. Qian, and A. Rakhlin, The Statistical Complexity of Interactive Decision Making, arXiv:2112.13487, 2021. https://arxiv.org/abs/2112.13487
  • D. Foster and A. Rakhlin, Foundations of Reinforcement Learning and Interactive Decision Making, arXiv:2312.16730, 2023. https://arxiv.org/abs/2312.16730
  • T. Lattimore and C. Szepesvári, Bandit Algorithms, Cambridge University Press, 2020.
  • J. Kiefer and J. Wolfowitz, The Equivalence of Two Extremum Problems, Canadian Journal of Mathematics, 1960.
9 thms4 active usersReviewed
🏆Completed
Experimental DesignOperations ResearchProbability+1·Captain: Shuze Chen

Treatment Locality in A/B TestingResearch Paper

Modern A/B tests must infer lifetime treatment effects — e.g. customer lifetime value under a new feature — from short-horizon experiment data. Chen, Simchi-Levi and Wang (arXiv:2407.19618) model the experiment as a Markov decision process and exploit a structural fact of many practical interventions: the treatment is local, modifying the system at a single crucial state only. This mission formalizes the core asymptotic theory of the paper: for any differentiable estimator built from the experiment's transition and reward statistics, information sharing — pooling across test arms the samples collected away from the treated state — keeps the estimator asymptotically normal with the same asymptotic bias and never increases its asymptotic variance (Theorem 9), and is asymptotically efficient among unbiased estimators (Theorem 5). The route runs through a Markov chain central limit theorem with the asymptotic variance identified as the autocovariance series, and the linearization/delta method for functionals of chain statistics.

42 thms4 active users
🏆Completed
Machine Learning·Captain: mikedeng1

Foundations of Reinforcement Learning VI: Function Approximation and Bellman RankTextbook

Motivation

Every RL guarantee proved earlier in this series — UCB-VI's regret bound, the contextual bandit oracle reductions — scales with the size of the state space SSS, because the algorithms maintain a separate statistic per state. Real environments (images, sensor readouts, natural language) have combinatorially or infinitely many states, so a tabular guarantee is vacuous there: the only hope is to generalize across states via a class of value functions, the way supervised learning generalizes across inputs via a hypothesis class. The chapter develops two algorithms along this line: LSVI-UCB, the linear-function-approximation analogue of UCB-VI whose regret is independent of ∣S∣|S|∣S∣ (a low-rank MDP result originating with Jin, Yang, Wang, and Jordan, Provably Efficient Reinforcement Learning with Linear Function Approximation, COLT 2020, arXiv:1907.05388), and BiLinUCB, which attains an analogous sample-complexity guarantee under the strictly more general structural condition of low Bellman rank (Jiang, Krishnamurthy, Agarwal, Langford, and Schapire, Contextual Decision Processes with Low Bellman Rank are PAC-Learnable, ICML 2017, arXiv:1610.09512; the Q-type variant formalized here follows Du, Kakade, Wang, and Yang, Is a Good Representation Sufficient for Sample Efficient Reinforcement Learning?, ICLR 2020, arXiv:1910.03016). This mission formalizes the second, strictly more general track: BiLinUCB and its Bellman-rank guarantee; LSVI-UCB is left out of scope (see Formalization scope).

Setting

Both algorithms act on the finite-horizon episodic MDP M=(S,A,{Ph}h=1H,{Rh}h=1H,d1)M = (S,A,\{P_h\}_{h=1}^H, \{R_h\}_{h=1}^H,d_1)M=(S,A,{Ph​}h=1H​,{Rh​}h=1H​,d1​) of the earlier chapters of this series (RLBasics.Core), with value fM(π)=Es1∼d1[V1M,π(s1)]f^M(\pi) = \mathbb E_{s_1\sim d_1}[V_1^{M,\pi}(s_1)]fM(π)=Es1​∼d1​​[V1M,π​(s1​)]. For a state-action value function Q=(Qh)h=1HQ = (Q_h)_{h=1}^HQ=(Qh​)h=1H​ (with the convention QH+1≡0Q_{H+1}\equiv 0QH+1​≡0), the Bellman residual of QQQ under policy π\piπ at layer hhh is

Eh(π,Q):=EM,π[Qh(sh,ah)−Rh(sh,ah)−max⁡a′Qh+1(sh+1,a′)],E_h(\pi,Q) := \mathbb E^{M,\pi}\Big[Q_h(s_h,a_h) - R_h(s_h,a_h) - \max_{a'} Q_{h+1}(s_{h+1},a')\Big],Eh​(π,Q):=EM,π[Qh​(sh​,ah​)−Rh​(sh​,ah​)−a′max​Qh+1​(sh+1​,a′)],

which vanishes identically when Q=QM,⋆Q = Q^{M,\star}Q=QM,⋆, the optimal value function (Bellman optimality). Given a class Q\mathcal QQ of candidate value functions, MMM has Bellman rank ddd relative to Q\mathcal QQ (Definition 8) if ddd is the least integer such that, at every layer hhh, there exist embeddings Xh(π),Wh(Q)∈RdX_h(\pi), W_h(Q) \in \mathbb R^dXh​(π),Wh​(Q)∈Rd with Eh(π,Q)=⟨Xh(π),Wh(Q)⟩E_h(\pi,Q) = \langle X_h(\pi), W_h(Q)\rangleEh​(π,Q)=⟨Xh​(π),Wh​(Q)⟩ for every policy π\piπ and Q∈QQ\in\mathcal QQ∈Q — equivalently, the least rank of the Π×Q\Pi\times\mathcal QΠ×Q matrix of Bellman residuals, at any layer. A linear MDP (the setting of LSVI-UCB, §7.2) is the special case where the transition kernel and reward themselves factor through a known feature map ϕ:S×A→Rd\phi : S\times A \to \mathbb R^dϕ:S×A→Rd; every linear MDP has Bellman rank at most ddd relative to the linear value-function class, but Bellman rank captures far more (kernel/neural function classes, and low-rank MDPs whose feature map is unknown).

The chapter presents two algorithms. LSVI-UCB (§7.2.1) runs TTT episodes of ridge regression per layer against the linear feature map, forming a confidence ellipsoid of radius ρ∝d\rho \propto \sqrt dρ∝d​ around each layer's estimated parameter and acting greedily with respect to an upper-confidence bonus built from that ellipsoid — the same optimism-under-uncertainty template as UCB-VI, now regularized rather than tabular; it motivates Bellman rank but is not itself formalized by this mission (see Formalization scope). BiLinUCB (§7.3.1), the algorithm this mission formalizes, instead proceeds in KKK iterations of nnn episodes: each iteration plays the greedy policy for the current optimistic-on-average value function Qk=arg⁡max⁡Q∈QkEs1∼d1[Q1(s1,πQ(s1))]Q_k = \arg\max_{Q\in\mathcal Q_k}\mathbb E_{s_1\sim d_1}[Q_1(s_1, \pi_Q(s_1))]Qk​=argmaxQ∈Qk​​Es1​∼d1​​[Q1​(s1​,πQ​(s1​))], collects nnn fresh episodes, and shrinks the confidence set Qk+1\mathcal Q_{k+1}Qk+1​ by discarding value functions whose empirical Bellman residual along the played policy is large; after KKK iterations it outputs the policy with the best empirical return observed at any iteration.

Formalization targets

Goal — Proposition 47 (BiLinUCB, sample complexity under Bellman rank)

∃ c1,c2,c3>0, ∀ ε,δ>0,  n≳H3dlog⁡(∣Q∣/δ)ε2,  K≳Hdlog⁡(1+n/d),  β∝Klog⁡∣Q∣+log⁡(HK/δ)n ⟹\exists\, c_1,c_2,c_3>0,\ \forall\,\varepsilon,\delta>0,\ \ n\gtrsim \frac{H^3d\log(|\mathcal Q|/\delta)}{\varepsilon^2},\ \ K\gtrsim Hd\log(1+n/d),\ \ \beta\propto\frac{K\log|\mathcal Q|+\log(HK/\delta)}n\ \Longrightarrow∃c1​,c2​,c3​>0, ∀ε,δ>0,  n≳ε2H3dlog(∣Q∣/δ)​,  K≳Hdlog(1+n/d),  β∝nKlog∣Q∣+log(HK/δ)​ ⟹ Pr⁡[fM⋆(πM⋆)−fM⋆(π^)≤ε]≥1−δ,\Pr\big[f^{M^\star}(\pi^{M^\star}) - f^{M^\star}(\hat\pi) \le \varepsilon\big] \ge 1-\delta,Pr[fM⋆(πM⋆)−fM⋆(π^)≤ε]≥1−δ,

for M⋆M^\starM⋆ of Bellman rank ddd relative to Q∋QM⋆,⋆\mathcal Q\ni Q^{M^\star,\star}Q∋QM⋆,⋆, where π^\hat\piπ^ is BiLinUCB's output policy after KKK iterations of nnn episodes. This is the weakest stable form: it fixes the shape of the sample complexity (polynomial in H,d,log⁡∣Q∣,1/εH,d,\log|\mathcal Q|,1/\varepsilonH,d,log∣Q∣,1/ε, logarithmic in 1/δ1/\delta1/δ, independent of ∣S∣|S|∣S∣) and leaves the leading constants — which the book itself introduces only as "a sufficiently large numerical constant" — outside the formal claim. Unlike every other goal in this series, this is a PAC (sample-complexity) guarantee on the algorithm's final output policy, not a bound on cumulative regret accrued while learning: BiLinUCB commits to π^\hat\piπ^ only after the training phase ends, and its suboptimality is measured post-training.

Reaching it rests on two structural facts about BiLinUCB's confidence sets, each formalized as a milestone in attack order:

  • Lemma 29 (confidence-set validity): with the stated threshold β\betaβ, with probability at least 1−δ1-\delta1−δ, simultaneously at every iteration kkk, every retained value function has true (population) Bellman residual along the played policies bounded by β\betaβ up to a constant, and the realizable QM⋆,⋆Q^{M^\star,\star}QM⋆,⋆ is itself always retained.
  • Lemma 30 (optimism and elliptic-norm bound): conditioned on Lemma 29's event, every retained value function's embedding Wh(Q)W_h(Q)Wh​(Q) has bounded norm with respect to the Gram matrix of the played policies' embeddings, and the optimistic value function QkQ_kQk​ BiLinUCB selects at each iteration has initial-state value at least fM⋆(πM⋆)f^{M^\star}(\pi^{M^\star})fM⋆(πM⋆).

Significance

Proposition 47 shows that a single structural parameter — Bellman rank — is sufficient for sample-efficient RL with function approximation, with a sample complexity that depends only on the horizon, the rank, and the value-function class's log-cardinality, never on ∣S∣|S|∣S∣. This subsumes the linear-MDP guarantee of Proposition 46 (LSVI-UCB, §7.2.1; not itself a target of this mission, see Formalization scope) as a special case — every linear MDP has Bellman rank ≤d\le d≤d — while covering strictly more models (kernelized and neural value-function classes with a low-dimensional Bellman-residual factorization that need not come from a known linear feature map). The result is proved in the source and this mission formalizes its statement and the two structural lemmas its proof rests on, as stated; no new mathematics is contributed. Formalizing it commits, for the first time on this platform, to machine-checkable statements of the Bellman rank abstraction, the elliptic-norm confidence-set machinery it drives, and a PAC- (rather than regret-) style learning guarantee, none of which appear in the platform's existing bandit or tabular-RL missions.

Difficulty

The obvious first attempt is to formalize Bellman rank as an unconstrained integer parameter ddd attached to the MDP, sidestepping the actual rank condition on the Bellman-residual matrix; this is a trivializing formalization; Bellman rank must be the least dimension admitting the stated bilinear factorization; see Formalization scope. A second obstacle is that BiLinUCB's optimism is only "on average" with respect to the initial state distribution (initValue), unlike LSVI-UCB's pointwise optimism over every state and action — conflating the two confidence-set constructions collapses the chapter's main conceptual contrast. Finally, the two technical lemmas (29 and 30) separate a purely probabilistic statement (validity of the empirical confidence set, via Hoeffding and a union bound) from a purely deterministic consequence (the elliptic-norm bound and optimism, which hold on any sample path where the probabilistic event occurred); keeping this separation is what lets Proposition 47's proof combine them cleanly, and collapsing it into one monolithic high-probability statement would misrepresent the book's proof structure.

Formalization scope

The MDP, policy, and trajectory/history machinery (EpisodicMDP, Policy, IsPolicy, Trajectory, Learner, probEvent) are reused unchanged from the series' published RLBasics.Core/RLBasics.UCBVI definitions. The value-function class Q\mathcal QQ is realized as an abstract finite, nonempty type Qc together with an evaluation map qeval : Qc → ℕ → S → A → ℝ, kept fully abstract rather than specialized to any concrete function class — specializing it to, e.g., linear functions would collapse Proposition 47 back into a restatement of Proposition 46, which is exactly the trivializing formalization this mission avoids. Bellman rank (IsBellmanRank) is defined as the least natural number admitting the bilinear factorization (an IsLeast over the coercion to HasBellmanRankLE), never as a free parameter. Every realized per-step reward is taken equal to its conditional mean Rh(s,a)R_h(s,a)Rh​(s,a) throughout (exact, by the tower property, and not a restriction to deterministic rewards). The elliptic norm of Lemma 30 is expressed via the direct sum-of-squared-inner-products identity ∥v∥Σ2=∑x∈xs⟨x,v⟩2\|v\|^2_\Sigma = \sum_{x\in xs}\langle x,v\rangle^2∥v∥Σ2​=∑x∈xs​⟨x,v⟩2 rather than introducing Matrix/matrix-inverse machinery, since no milestone here needs an explicit matrix inverse. Every place the book writes ≲/∝/"a sufficiently large numerical constant" is formalized as an existentially quantified universal constant, fixed ahead of every MDP, value-function class, and (ε,δ)(\varepsilon,\delta)(ε,δ) instance — never depending on the instance itself. This mission scopes entirely to the Bellman-rank track (Definition 8, BiLinUCB, Lemmas 29-30, Proposition 47); Proposition 46 (LSVI-UCB) and its supporting Lemmas 27-28 are the chapter's motivating linear special case (Section 7.2) but are left out of this mission's scope for time and are not claimed as proved by it — LSVI-UCB's confidence-ellipsoid construction is materially different from BiLinUCB's empirical-Bellman-residual confidence set (see Difficulty) and would need its own milestone chain. A complete development needs no infrastructure beyond what is already published in RLBasics.Core/RLBasics.UCBVI; contributions completing the sorrys in the two technical lemmas and the goal are welcome.

Selected references

  • Foster, D. J. and Rakhlin, A. Foundations of Reinforcement Learning and Interactive Decision Making. arXiv:2312.16730v1, 2023. arXiv:2312.16730
  • Jin, C., Yang, Z., Wang, Z., and Jordan, M. I. Provably Efficient Reinforcement Learning with Linear Function Approximation. COLT 2020. arXiv:1907.05388
  • Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E. Contextual Decision Processes with Low Bellman Rank are PAC-Learnable. ICML 2017. arXiv:1610.09512
  • Du, S. S., Kakade, S. M., Wang, R., and Yang, L. F. Is a Good Representation Sufficient for Sample Efficient Reinforcement Learning? ICLR 2020. arXiv:1910.03016
7 thms3 active usersReviewed
🏆Completed
Dynamic ProgrammingMachine Learning·Captain: mikedeng1

Foundations of Reinforcement Learning IV: Reinforcement Learning Basics and the UCB-VI AlgorithmTextbook

Motivation

Reinforcement learning (RL) formalizes sequential decision-making under uncertainty: an agent repeatedly observes a state, takes an action, receives a reward, and transitions to a new state, with the goal of maximizing cumulative reward over an unknown environment. It underlies applications from game-playing agents to robotics and adaptive medical treatment. What separates RL from the bandit and contextual bandit problems of earlier chapters in this series is state: the environment carries information forward across time steps, so a good action now can pay off many steps later, and a poor exploration strategy can take exponentially long to discover it. This mission formalizes the foundational results of finite-horizon episodic RL — the Markov Decision Process (MDP) model, the Bellman-optimality principle that makes dynamic programming possible, and the analytical toolkit (the performance-difference and Bellman-residual-decomposition lemmas) used throughout the field's regret analyses — culminating in a polynomial regret guarantee for UCB-VI, the canonical optimism-based algorithm for tabular RL introduced by Azar, Osband, and Munos (Minimax Regret Bounds for Reinforcement Learning, ICML 2017, arXiv:1703.05449).

Setting

A finite-horizon episodic Markov Decision Process M=(S,A,{PhM}h=1H,{RhM}h=1H,d1)M = (S, A, \{P^M_h\}_{h=1}^H, \{R^M_h\}_{h=1}^H, d_1)M=(S,A,{PhM​}h=1H​,{RhM​}h=1H​,d1​) consists of a finite state space SSS, a finite action space AAA, a horizon HHH, per-layer transition kernels PhM:S×A→Δ(S)P^M_h : S\times A \to \Delta(S)PhM​:S×A→Δ(S) and reward distributions RhM:S×A→Δ(R)R^M_h : S\times A \to \Delta(\mathbb R)RhM​:S×A→Δ(R), and an initial state distribution d1∈Δ(S)d_1 \in \Delta(S)d1​∈Δ(S). An episode unrolls for h=1,…,Hh=1,\dots,Hh=1,…,H: the learner selects an action ah∼πh(sh)a_h \sim \pi_h(s_h)ah​∼πh​(sh​) under a randomized non-stationary policy π=(π1,…,πH)∈Πrns\pi = (\pi_1,\dots,\pi_H) \in \Pi^{\mathrm{rns}}π=(π1​,…,πH​)∈Πrns (each πh:S→Δ(A)\pi_h : S \to \Delta(A)πh​:S→Δ(A)), receives reward rh∼RhM(sh,ah)r_h \sim R^M_h(s_h,a_h)rh​∼RhM​(sh​,ah​), and transitions to sh+1∼PhM(sh,ah)s_{h+1}\sim P^M_h(s_h,a_h)sh+1​∼PhM​(sh​,ah​). The value of π\piπ is fM(π):=EM,π[∑hrh]f^M(\pi) := \mathbb E^{M,\pi}[\sum_h r_h]fM(π):=EM,π[∑h​rh​], and the state-action and state value functions QhM,π(s,a)Q^{M,\pi}_h(s,a)QhM,π​(s,a), VhM,π(s)V^{M,\pi}_h(s)VhM,π​(s) are the analogous reward-to-go quantities from layer hhh onward. In the online RL problem, M⋆M^\starM⋆ is unknown, the learner interacts with it for TTT episodes, and the goal is to minimize the regret Reg=∑t=1T(fM⋆(πM⋆)−fM⋆(πt))\mathrm{Reg} = \sum_{t=1}^T \big(f^{M^\star} (\pi^{M^\star}) - f^{M^\star}(\pi^t)\big)Reg=∑t=1T​(fM⋆(πM⋆)−fM⋆(πt)) against the best policy πM⋆∈arg⁡max⁡π∈ΠrnsfM⋆(π)\pi^{M^\star} \in \arg\max_{\pi\in\Pi^{\mathrm{rns}}} f^{M^\star}(\pi)πM⋆∈argmaxπ∈Πrns​fM⋆(π).

Formalization targets

Goal — Theorem 1 (UCB-VI regret)

∃ C>0, ∀ δ∈(0,1],Pr⁡[Reg≤C⋅H⋅S⋅A T⋅log⁡(SAHT/δ)]≥1−δ,\exists\, C>0,\ \forall\, \delta\in(0,1],\quad \Pr\Big[\mathrm{Reg} \le C\cdot H\cdot S\cdot\sqrt{A\,T}\cdot\sqrt{\log(SAHT/\delta)}\Big] \ge 1-\delta,∃C>0, ∀δ∈(0,1],Pr[Reg≤C⋅H⋅S⋅AT​⋅log(SAHT/δ)​]≥1−δ,

for the UCB-VI algorithm run with the explicit bonus bh,δt(s,a)=2log⁡(2SAHT/δ)/nht(s,a)b^t_{h,\delta}(s,a) = 2\sqrt{\log(2SAHT/\delta)/n^t_h(s,a)}bh,δt​(s,a)=2log(2SAHT/δ)/nht​(s,a)​, under Assumption 6 (deterministic, known, [0,1][0,1][0,1]-bounded rewards). This is the weakest stable form of the guarantee: it fixes only the shape of the bound (polynomial in S,A,H,TS,A,H,TS,A,H,T, logarithmic in 1/δ1/\delta1/δ), leaving the exact leading constant — which the book itself does not pin down on this page — outside the formal claim.

Reaching it rests on three structural facts, each formalized as a milestone in attack order:

  • Proposition 25 (Bellman optimality): existence of a single deterministic policy simultaneously optimal at every state, computable by backward induction — the reason dynamic programming solves planning at all.
  • Lemma 13 (Performance difference) and Lemma 14 (Bellman residual decomposition): two "credit assignment" identities decomposing a value gap (between two policies, or one policy under two models) into a sum of per-layer, on-roll-in terms.
  • Lemma 15 (Error decomposition for optimistic policies): the fact that a greedy policy driven by any optimistic value estimate suffers sub-optimality controlled additively — not exponentially — by that estimate's own Bellman residuals, evaluated on-policy.

Significance

The result itself. Theorem 1 is the chapter's headline result: the first guarantee, in this development, that a learning algorithm — one that does not know the environment's transitions in advance — can achieve regret growing only polynomially in the size of the state space, the action space, and the horizon, and only as T\sqrt TT​ in the number of episodes. The chapter's own "combination lock" example (Figure 8) shows this is not automatic: naive exploration strategies (such as ε\varepsilonε-greedy, which suffices for ordinary bandits) incur regret exponential in the horizon on some MDPs with as few as H+2H+2H+2 states. UCB-VI's guarantee is the sample-complexity foundation on which essentially every subsequent result on tabular, linear, and general function-approximation RL in the book is built.

Formalizing it. No part of this chapter's mathematical content already has a faithful counterpart on the platform (see Difficulty, below, and Formalization scope). This mission contributes: (i) a from-scratch Lean formalization of the finite-horizon episodic MDP model and its value functions, faithful to the book's 000/111-indexing and reward-distribution conventions; (ii) faithful statements (drafted with sorry, not yet proved) of Proposition 25 and Lemmas 13–15; and (iii) a faithful statement of Theorem 1 itself, including a from-scratch construction of the finite TTT-episode adaptive interaction process needed to make sense of a high-probability regret guarantee. Proving these — Proposition 25 by backward induction, Lemmas 13–15 by telescoping, and Theorem 1 by combining Lemma 15's optimism bound with a concentration argument bounding the estimation-error and martingale terms of Eqs. (5.30)–(5.33) (omitted here; see Difficulty) — is open work for solvers.

Difficulty

The obvious approach to Theorem 1 — bound the regret episode-by-episode using only the fact that QtQ^tQt is close to QM⋆,⋆Q^{M^\star,\star}QM⋆,⋆ in some fixed sense — fails because QtQ^tQt's error is itself random (it depends on the transitions observed so far) and compounds across HHH layers of dynamic programming. Lemma 15 defuses the compounding: it shows the sub-optimality gap is additive in the per-layer Bellman residuals rather than multiplicative, provided QtQ^tQt is optimistic. Making QtQ^tQt optimistic with high probability, in turn, requires a concentration argument for the empirical transition estimates P^ht\widehat P^t_hPht​ (an application of Freedman's or the Azuma–Hoeffding inequality, not included among this mission's milestones) and a union bound over all (s,a,h,t)(s,a,h,t)(s,a,h,t) — accounting for the SAHTSAHTSAHT inside the bonus's logarithm. The final regret sum further requires bounding ∑t∑h1/nht(sht,aht)\sum_t \sum_h 1/\sqrt{n^t_h(s^t_h,a^t_h)}∑t​∑h​1/nht​(sht​,aht​)​ by a pigeonhole/potential-function argument over the visitation counts, which is where the SATS\sqrt{AT}SAT​ scaling — rather than a naive SATSA\sqrt{T}SAT​ — originates. None of this concentration or counting machinery is included in the current milestones; a solver attempting Theorem 1 needs it as prerequisite lemmas.

Formalization scope

MDP and value functions. States and actions are finite types (Fintype); layers are represented 000-indexed throughout the Lean development (the book's layer hhh is h - 1), with the terminal convention V _ _ H _ = 0 matching VH+1≡0V_{H+1}\equiv 0VH+1​≡0. Since every value/expectation formula in this chapter uses the reward distribution Rh(s,a)∈Δ(R)R_h(s,a)\in\Delta(\mathbb R)Rh​(s,a)∈Δ(R) only through its mean, EpisodicMDP.R records that mean directly — equivalent, by linearity of expectation, to carrying the full distribution, and changing no theorem's content. Transition kernels and policies are represented as plain real-valued functions (S → A → ℝ-style) rather than as Mathlib's PMF, with IsPolicy/the EpisodicMDP structure's own proof fields asserting the probability-distribution properties (nonnegativity, summing to 111) where the book requires membership in Πrns\Pi^{\mathrm{rns}}Πrns or a well-formed kernel; this keeps every expectation a finite Finset.sum, needing no measure theory. The optimal value functions Qstar/Vstar are defined as literal suprema over the entire (uncountable, since ∣A∣≥2|A|\ge2∣A∣≥2) policy type — not via a recursive shortcut — which is what rules out the trivializing formalization of Proposition 25: defining Vstar by the very recursion the proposition asserts would make the proposition a tautology, whereas here it is a genuine claim about a supremum over an enormous space of competitor policies.

UCB-VI and Theorem 1. Because SSS, AAA, HHH, and TTT are all finite, the TTT-episode adaptive interaction (in which round ttt's policy is a function of the realized history of the first t−1t-1t−1 episodes) is modeled as a finite probability space: the outcome type Fin T → Trajectory S A H is a Fintype, "probability" is a finite sum over it, and "with probability ≥1−δ\ge 1-\delta≥1−δ" is a plain inequality between two real numbers — no MeasureTheory is used anywhere in this mission. The universal constant CCC in Theorem 1 is existentially quantified (∃ C > 0, …) rather than given as a literal numeral, since its value is not pinned down by the book on this page and this mission does not carry out the (non-milestoned) concentration argument that would derive it; this is the convention adopted throughout for the book's own "≲\lesssim≲" notation. The empirical-transition estimator inside QhtQ^t_hQht​ is defined to be 000 when nht(s,a)=0n^t_h(s,a)=0nht​(s,a)=0 (no data yet collected for that pair) — a boundary case Eq. (5.25) does not address, resolved here by convention rather than proof.

Prior art. The platform's existing BanditAlgorithm.UCRL2Algorithm mission formalizes UCRL2 for the average-reward, infinite-horizon MDP setting with a diameter parameter, and its Bellman-optimality statement is the average-cost optimality equation — genuinely different from this chapter's finite-horizon episodic Bellman recursion, even though both go by the name "Bellman optimality." No reference item was reused; every definition and theorem in this mission is drafted from scratch. Both missions bound regret via optimism over a confidence set of models, but for different objectives (average reward vs. finite-horizon cumulative reward) and different MDP classes.

Reusable infrastructure and open contributions. EpisodicMDP, Policy, V/Q/Qstar/Vstar, and stateDist/stateExp are reusable by any future mission on finite-horizon episodic RL in this book's later chapters. Contributions welcome: proofs of the four milestone lemmas (by backward induction and telescoping, respectively); the concentration lemmas underlying Theorem 1's optimism guarantee (Eqs. (5.30)–(5.33) of the source, not milestoned here); and the final regret proof combining them.

Selected references

  • Foster, D. J. and Rakhlin, A., Foundations of Reinforcement Learning and Interactive Decision Making, 2023. arXiv:2312.16730v1
  • Azar, M. G., Osband, I., and Munos, R., Minimax Regret Bounds for Reinforcement Learning, ICML 2017. arXiv:1703.05449
  • Jaksch, T., Ortner, R., and Auer, P., Near-optimal Regret Bounds for Reinforcement Learning, JMLR 11 (2010).
7 thms3 active usersReviewed
Machine LearningProbability·Captain: mikedeng1

Minimax Regret Bounds for Reinforcement Learning II: High-Probability Regret Bound for UCBVI with a Bernstein–Freedman BonusResearch Paper

Why finite-horizon reinforcement learning needs a variance-sensitive bound

An agent can learn to act in an unknown environment by repeatedly running a finite episode, observing the states reached after its actions, and updating its model of the environment. The agent must trade off rewards in the current episode against information that may improve later decisions. A regret bound measures the cumulative value lost relative to an optimal policy that knows the true transition probabilities. Its dependence on the number of states, actions, episode steps, and interactions says how much exploration that uncertainty can force.

Azar, Osband, and Munos study this question for a finite-horizon Markov decision process with known, bounded rewards and an unknown, stationary transition kernel. Their UCBVI algorithm estimates action values from observed transitions and adds an exploration bonus. Their second version, UCBVI-BF, uses the empirical variance of the next-state value in that bonus. Their Theorem 2 gives an explicit high-probability regret bound whose leading dependence on the horizon is smaller than the bound they give for the simpler UCBVI-CH bonus. The paper states that, in a sufficiently long-run regime, its leading order matches the cited lower-bound scale up to logarithmic factors. This mission targets the explicit theorem, including its lower-order terms, rather than only that asymptotic comparison.

The MDP, interaction, and algorithm

Let S\mathcal SS and A\mathcal AA be nonempty finite state and action sets with cardinalities SSS and AAA. A stationary transition kernel P(y∣x,a)P(y\mid x,a)P(y∣x,a) is a probability distribution on next states yyy for every current state xxx and action aaa. The reward R(x,a)R(x,a)R(x,a) is deterministic, known to the learner, and lies in [0,1][0,1][0,1]. These are the conditions of Assumption 1 and §2. Episodes have H≥1H\ge1H≥1 steps; KKK episodes comprise T=KHT=KHT=KH interactions.

A policy π\piπ chooses an action for each state and step. Its value Vhπ(x)V_h^\pi(x)Vhπ​(x) is the expected reward from step hhh through the final step when the state at hhh is xxx; the terminal value is VH+1π=0V_{H+1}^\pi=0VH+1π​=0. The optimal value Vh∗(x)=sup⁡πVhπ(x)V_h^*(x)=\sup_\pi V_h^\pi(x)Vh∗​(x)=supπ​Vhπ​(x) ranges over all deterministic policies of this form. At the start of episode kkk, the environment may choose the initial state using the completed episodes. The learner then fixes a policy πk\pi_kπk​, observes transitions during the episode, and updates counts for the next episode. Its regret is

Regret⁡(K)=∑k=1K(V1∗(xk,1)−V1πk(xk,1)).\operatorname{Regret}(K)=\sum_{k=1}^{K}\bigl(V_1^*(x_{k,1})-V_1^{\pi_k}(x_{k,1})\bigr).Regret(K)=k=1∑K​(V1∗​(xk,1​)−V1πk​​(xk,1​)).

For each state-action pair, Nk(x,a,y)N_k(x,a,y)Nk​(x,a,y) counts transitions to yyy in episodes before kkk, and Nk(x,a)=∑yNk(x,a,y)N_k(x,a)=\sum_yN_k(x,a,y)Nk​(x,a)=∑y​Nk​(x,a,y). When the latter is positive, P^k(y∣x,a)=Nk(x,a,y)/Nk(x,a)\widehat P_k(y\mid x,a)=N_k(x,a,y)/N_k(x,a)Pk​(y∣x,a)=Nk​(x,a,y)/Nk​(x,a). The count Nk,h′(y)N'_{k,h}(y)Nk,h′​(y) records previous episodes whose state at step hhh was yyy. Algorithms 2 and 4 compute optimistic Qk,hQ_{k,h}Qk,h​ backward from zero terminal value, take a minimum with the previous episode's QQQ estimate and with HHH, and choose a maximizing action at every state. Previously unseen pairs receive Qk,h=HQ_{k,h}=HQk,h​=H. The Bernstein–Freedman bonus uses the empirical variance of Vk,h+1V_{k,h+1}Vk,h+1​ under P^k\widehat P_kPk​ and an additional term based on Nk,h+1′N'_{k,h+1}Nk,h+1′​; the algorithm uses Lalg=ln⁡(5SAT/δ)L_{\rm alg}=\ln(5SAT/\delta)Lalg​=ln(5SAT/δ).

Formalization targets

The goal is Theorem 2 on p. 5. For any MDP and interaction described above and every δ>0\delta>0δ>0, write L=ln⁡(5HSAT/δ)L=\ln(5HSAT/\delta)L=ln(5HSAT/δ). The target is the exact bad-event form of the printed high-probability bound:

Pr⁡ ⁣{Regret⁡(K)>30HLSAK+2500H2S2AL2+4H3/2KL}≤δ.\Pr\!\left\{\operatorname{Regret}(K)>30HL\sqrt{SAK}+2500H^2S^2AL^2+4H^{3/2}\sqrt{KL}\right\}\le\delta.Pr{Regret(K)>30HLSAK​+2500H2S2AL2+4H3/2KL​}≤δ.

The milestone list contains three empirical-transition deviations from the proof of Lemma 1: Eq. (9) for a value-weighted transition error, the displayed count bound before Eq. (11), and Eq. (12) for the full transition row's ℓ1\ell_1ℓ1​ error. It also contains Lemma 2's variance comparison and Eq. (26), which relates cumulative conditional next-value variance to the variance of an episode return. These are source-indexed targets, with their printed constants retained.

What the result and its formalization supply

The theorem gives a quantitative guarantee for a particular executable decision rule: its regret grows sublinearly in KKK in the leading term, with explicit dependence on SSS, AAA, and HHH. The result lets one compare the horizon dependence of a variance-sensitive bonus with a value-agnostic bonus under the same finite-horizon model. It also fixes which logarithm belongs in the algorithm and which appears in the reported bound; replacing either changes the claim.

A formal proof would connect a fully specified adaptive interaction to its finite probability law, empirical counts, backward value iteration, and the stated high-probability conclusion. The local prior-art search found reusable transition-kernel vocabulary and general concentration tools, but no published formal statement of this exact UCBVI-BF algorithm or theorem. The mission's finite path and variance definitions can also support other episodic reinforcement-learning bounds that use conditional variance.

Where the difficulty lies

The bonus is computed using a value function that itself depends on earlier observations and the same episode's backward recursion. A concentration inequality for a fixed transition row and a fixed test function therefore does not directly control every value estimate encountered by the algorithm. The number of samples in a row is also random and changes with the learner's past actions. The regret compares a policy's value at an environment-chosen initial state with a supremum over all policies, while the learner's greedy action must be defined at states it never visits. These dependencies are the central obstacle to turning local concentration statements into the episode-level bound.

Formalization scope and conventions

The Lean model uses finite sums rather than measure theory. A published predicate supplies the stationary, real-valued transition kernel; a local MDP adds the known deterministic reward. State and action types are finite and nonempty. Policies are deterministic and depend on the step. The supremum defining V∗V^*V∗ ranges over their finite function type. A theorem quantifies over every maximizing tie-breaking rule and every initial-state rule that reads only completed episodes. The probability of an event is constructed as a sum over finite outcome sequences, each weighted by the product of true transition probabilities. Counts use all past transitions and no current or future outcomes. These choices rule out a trivialization that assumes the desired law or optimizes over an unbounded class of arbitrary functions.

Lean indexes the HHH steps from zero, while the paper indexes them from one. The last observed next state is kept because Algorithm 4 counts states at the terminal index H+1H+1H+1. At Nk,h+1′(y)=0N'_{k,h+1}(y)=0Nk,h+1′​(y)=0, Algorithm 4's quotient is interpreted as infinite and the capped term is H2H^2H2; Lean's ordinary division by zero would incorrectly produce zero. The algorithm uses Lalg=ln⁡(5SAT/δ)L_{\rm alg}=\ln(5SAT/\delta)Lalg​=ln(5SAT/δ), while Theorem 2's bound uses L=ln⁡(5HSAT/δ)L=\ln(5HSAT/\delta)L=ln(5HSAT/δ). For Eq. (26), the appendix ends its sums at H−1H-1H−1 under a shifted terminal convention; the local statement includes all HHH reward steps and the terminal value VH+1=0V_{H+1}=0VH+1​=0 used by Algorithm 2. The milestone text remains the printed text. The count milestone is the display before Eq. (11), since Eq. (11) drops a factor of 222 under the square root present in that display.

Theorem 2 retains its printed 2500H2S2AL22500H^2S^2AL^22500H2S2AL2 term. The appendix's displayed Lemma 13 calculation does not reproduce that second-order constant when propagated to Lemma 14; this is a source proof gap, not a hypothesis of the theorem. Work on the probability normalization, random-count concentration, adaptive value estimates, variance identity, and a valid route to the printed explicit constants is welcome. A proof with altered constants or an asymptotic-only conclusion would be a different target.

Selected references

  • M. G. Azar, I. Osband, and R. Munos, Minimax Regret Bounds for Reinforcement Learning, arXiv:1703.05449v2, 2017. Preprint.
13 thms2 active usersReviewed
Linear algebraMachine LearningOptimization·Captain: mikedeng1

Reinforcement Learning: An Introduction X: Off-policy Divergence and the Gradient of the Projected Bellman ErrorTextbook

Motivation

Off-policy learning estimates the value function of a target policy π\piπ from data generated by a different behavior policy bbb. It is how an agent learns about a greedy policy while exploring, and how many policies can be evaluated from one stream of experience. Combined with linear function approximation and bootstrapping (updating an estimate toward other estimates, as temporal-difference methods do), off-policy learning can be unstable: the weights can diverge even when every quantity involved is well defined. Sutton and Barto call this combination the deadly triad (Chapter 11 of Reinforcement Learning: An Introduction, 2nd ed., 2018).

Chapter 11 does two things. It exhibits the instability with small, fully computable examples, and it identifies an objective that can be minimized stably from off-policy data: the mean square projected Bellman error (PBE), whose gradient the Gradient-TD methods (GTD2, TDC) follow in expectation. This mission formalizes the chapter's exact, finite-dimensional claims.

A short history, following the book's bibliographical remarks (pp. 285–286): Baird (1995) gave the seven-state counterexample for off-policy semi-gradient TD; Tsitsiklis and Van Roy (1996) gave the earliest w-to-2w example and the counterexample of Example 11.1, showing that even a least-squares fit at every step can diverge; Gradient-TD methods, which follow the gradient of the PBE, were introduced by Sutton, Szepesvári and Maei and by Sutton et al. (2009); Sutton, Mahmood and White (2016) introduced Emphatic-TD. The learnability discussion of §11.6 is the book's own.

Setting

A finite Markov decision process has states S\mathcal SS, actions A\mathcal AA, a finite reward set R\mathcal RR and dynamics p(s′,r∣s,a)p(s',r\mid s,a)p(s′,r∣s,a). A policy π(a∣s)\pi(a\mid s)π(a∣s) induces the transition matrix Pπ(s,s′)=∑aπ(a∣s) p(s′∣s,a)P_\pi(s,s') = \sum_a \pi(a\mid s)\,p(s'\mid s,a)Pπ​(s,s′)=∑a​π(a∣s)p(s′∣s,a) and expected reward rπ(s)r_\pi(s)rπ​(s). For 0≤γ<10\le\gamma<10≤γ<1 the true value function is the expected discounted return vπ(s)=∑k≥0γk(Pπkrπ)(s)v_\pi(s) = \sum_{k\ge0}\gamma^k (P_\pi^k r_\pi)(s)vπ​(s)=∑k≥0​γk(Pπk​rπ​)(s).

A state weighting μ\muμ is a probability distribution on S\mathcal SS, with D=diag⁡(μ)\mathbf D = \operatorname{diag}(\mu)D=diag(μ) and norm ∥v∥μ2=∑sμ(s)v(s)2\|v\|_\mu^2 = \sum_s \mu(s)v(s)^2∥v∥μ2​=∑s​μ(s)v(s)2. The ∣S∣×d|\mathcal S|\times d∣S∣×d matrix X\mathbf XX has the feature vectors x(s)⊤\mathbf x(s)^\topx(s)⊤ as rows, and a weight vector w∈Rd\mathbf w\in\mathbb R^dw∈Rd gives the linear value function vw=Xwv_{\mathbf w} = \mathbf X\mathbf wvw​=Xw. The Bellman operator is

(Bπv)(s)=∑aπ(a∣s)∑s′,rp(s′,r∣s,a) [r+γv(s′)].(B_\pi v)(s) = \sum_a \pi(a\mid s)\sum_{s',r} p(s',r\mid s,a)\,[r+\gamma v(s')].(Bπ​v)(s)=a∑​π(a∣s)s′,r∑​p(s′,r∣s,a)[r+γv(s′)].

The Bellman error vector is δˉw=Bπvw−vw\bar\delta_{\mathbf w} = B_\pi v_{\mathbf w} - v_{\mathbf w}δˉw​=Bπ​vw​−vw​, the projection is Π=X(X⊤DX)−1X⊤D\Pi = \mathbf X(\mathbf X^\top\mathbf D\mathbf X)^{-1}\mathbf X^\top\mathbf DΠ=X(X⊤DX)−1X⊤D, and the two objectives are

BE(w)=∥δˉw∥μ2,PBE(w)=∥Πδˉw∥μ2.\mathrm{BE}(\mathbf w) = \|\bar\delta_{\mathbf w}\|_\mu^2,\qquad \mathrm{PBE}(\mathbf w) = \|\Pi\bar\delta_{\mathbf w}\|_\mu^2 .BE(w)=∥δˉw​∥μ2​,PBE(w)=∥Πδˉw​∥μ2​.

Off-policy samples are reweighted by the importance-sampling ratio ρt=π(At∣St)/b(At∣St)\rho_t = \pi(A_t\mid S_t)/b(A_t\mid S_t)ρt​=π(At​∣St​)/b(At​∣St​).

Formalization targets

Goal: the gradient of the PBE

If X⊤DX\mathbf X^\top\mathbf D\mathbf XX⊤DX is invertible, then for every w\mathbf ww

PBE(w)=(X⊤Dδˉw)⊤(X⊤DX)−1(X⊤Dδˉw),\mathrm{PBE}(\mathbf w) = (\mathbf X^\top\mathbf D\bar\delta_{\mathbf w})^\top(\mathbf X^\top\mathbf D\mathbf X)^{-1}(\mathbf X^\top\mathbf D\bar\delta_{\mathbf w}),PBE(w)=(X⊤Dδˉw​)⊤(X⊤DX)−1(X⊤Dδˉw​), ∇PBE(w)=2 (γPπX−X)⊤DX (X⊤DX)−1 X⊤Dδˉw.\nabla\mathrm{PBE}(\mathbf w) = 2\,(\gamma P_\pi\mathbf X-\mathbf X)^\top\mathbf D\mathbf X\,(\mathbf X^\top\mathbf D\mathbf X)^{-1}\,\mathbf X^\top\mathbf D\bar\delta_{\mathbf w}.∇PBE(w)=2(γPπ​X−X)⊤DX(X⊤DX)−1X⊤Dδˉw​.

With μ\muμ the state distribution under bbb this is (11.27), ∇PBE(w)=2 E[ρt(γxt+1−xt)xt⊤] E[xtxt⊤]−1 E[ρtδtxt]\nabla\mathrm{PBE}(\mathbf w) = 2\,\mathbb E[\rho_t(\gamma\mathbf x_{t+1}-\mathbf x_t)\mathbf x_t^\top]\,\mathbb E[\mathbf x_t\mathbf x_t^\top]^{-1}\,\mathbb E[\rho_t\delta_t\mathbf x_t]∇PBE(w)=2E[ρt​(γxt+1​−xt​)xt⊤​]E[xt​xt⊤​]−1E[ρt​δt​xt​].

Milestones, in attack order

  1. The w-to-2w example (p. 260): repeated off-policy semi-gradient TD(0) on one transition multiplies www by 1+α(2γ−1)1+\alpha(2\gamma-1)1+α(2γ−1), so wt→±∞w_t\to\pm\inftywt​→±∞ for every α>0\alpha>0α>0 once γ>12\gamma>\tfrac12γ>21​.
  2. Example 11.1, (11.10) (p. 263): the least-squares iteration wk+1=6−4ε5γwkw_{k+1} = \tfrac{6-4\varepsilon}{5}\gamma w_kwk+1​=56−4ε​γwk​ diverges when γ>5/(6−4ε)\gamma>5/(6-4\varepsilon)γ>5/(6−4ε) and w0≠0w_0\ne0w0​=0.
  3. (11.12)–(11.13): Πv\Pi vΠv is the unique μ\muμ-closest representable function, and Π⊤DΠ=DX(X⊤DX)−1X⊤D\Pi^\top\mathbf D\Pi = \mathbf D\mathbf X(\mathbf X^\top\mathbf D\mathbf X)^{-1}\mathbf X^\top\mathbf DΠ⊤DΠ=DX(X⊤DX)−1X⊤D.
  4. (11.21): vπv_\pivπ​ is the unique fixed point of BπB_\piBπ​ (γ<1\gamma<1γ<1).
  5. (11.24), Exercise 11.4: RE(w)=VE(w)+E[(Gt−vπ(St))2]\mathrm{RE}(\mathbf w) = \mathrm{VE}(\mathbf w) + \mathbb E[(G_t - v_\pi(S_t))^2]RE(w)=VE(w)+E[(Gt​−vπ​(St​))2] in the on-policy case.
  6. Example 11.4 (p. 276): two Markov reward processes that generate the same observable data distribution (every finite prefix of the stream of feature vectors and rewards has the same probability, each process started from its stationary distribution) have BE(0)=0\mathrm{BE}(\mathbf 0) = 0BE(0)=0 and BE(0)=23\mathrm{BE}(\mathbf 0) = \tfrac23BE(0)=32​.
  7. (11.25)–(11.26): the PBE in matrix terms.
  8. The three factors (p. 278): X⊤Dδˉw=E[ρtδtxt]\mathbf X^\top\mathbf D\bar\delta_{\mathbf w} = \mathbb E[\rho_t\delta_t\mathbf x_t]X⊤Dδˉw​=E[ρt​δt​xt​], (γPπX−X)⊤DX=E[ρt(γxt+1−xt)xt⊤](\gamma P_\pi\mathbf X-\mathbf X)^\top\mathbf D\mathbf X = \mathbb E[\rho_t(\gamma\mathbf x_{t+1}-\mathbf x_t)\mathbf x_t^\top](γPπ​X−X)⊤DX=E[ρt​(γxt+1​−xt​)xt⊤​], X⊤DX=E[xtxt⊤]\mathbf X^\top\mathbf D\mathbf X = \mathbb E[\mathbf x_t\mathbf x_t^\top]X⊤DX=E[xt​xt⊤​], under coverage.

Significance

The results. The divergence examples show that no step-size choice rescues semi-gradient TD off-policy, and that exact least-squares fitting does not either. The learnability results separate objectives that can be estimated from observed features and rewards (RE, PBE) from one that cannot (BE), and (11.24) shows that the unobservable VE has the same minimizer as the observable RE. The gradient formula (11.27) is the expected update of the Gradient-TD family; its factorization into three expectations is what makes an O(d)O(d)O(d) stochastic-gradient method possible.

Formalizing them. All of these are proved or computed in the book, in informal notation that leaves hypotheses implicit (invertibility, coverage, the distribution μ\muμ, the domain of γ\gammaγ). A machine-checked version fixes them, and gives a reusable layer of linear value-function geometry (Bellman operator, μ\muμ-norm, projection, BE, PBE, behavior-policy expectations) for later work on Gradient-TD convergence and on the TD fixed point. No formalization of this chapter is known to exist.

Difficulty

The individual steps are finite linear algebra, but they are easy to get wrong. The gradient of a quadratic form h⊤C−1h\mathbf h^\top\mathbf C^{-1}\mathbf hh⊤C−1h in an affine h(w)\mathbf h(\mathbf w)h(w) produces C−1+C−⊤\mathbf C^{-1}+\mathbf C^{-\top}C−1+C−⊤, and collapsing it to 2C−12\mathbf C^{-1}2C−1 uses the symmetry of X⊤DX\mathbf X^\top\mathbf D\mathbf XX⊤DX. The Jacobian of w↦X⊤Dδˉw\mathbf w\mapsto\mathbf X^\top\mathbf D\bar\delta_{\mathbf w}w↦X⊤Dδˉw​ has to be identified through the Bellman operator's affine form rπ+γPπvr_\pi+\gamma P_\pi vrπ​+γPπ​v, which is a theorem about the four-argument dynamics, not a definition. Passing from matrix to expectation form needs coverage, since otherwise b(a∣s)ρ=π(a∣s)b(a\mid s)\rho = \pi(a\mid s)b(a∣s)ρ=π(a∣s) fails. The divergence examples need the conclusion ∣wk∣→∞|w_k|\to\infty∣wk​∣→∞, not merely "the multiplier exceeds one".

Formalization scope

States are a finite type; values are functions S→R\mathcal S\to\mathbb RS→R; weights are Rd\mathbb R^dRd as Fin d → ℝ; matrices are Mathlib Matrix. The dynamics are the book's p(s′,r∣s,a)p(s',r\mid s,a)p(s′,r∣s,a) with a finite reward set, and one action type serves all states. μ\muμ is a probability distribution (nonnegative, summing to one). vπv_\pivπ​ is defined from expected returns, not as a Bellman solution. The gradient is a Fréchet derivative whose linear map is u↦g⊤u\mathbf u\mapsto\mathbf g^\top\mathbf uu↦g⊤u.

Explicit hypotheses the book leaves implicit:

  • X⊤DX\mathbf X^\top\mathbf D\mathbf XX⊤DX invertible. The book substitutes a pseudoinverse otherwise; that case is not formalized, and without the hypothesis Lean's matrix inverse is zero and the PBE is trivially 000.
  • Coverage (π(a∣s)>0⇒b(a∣s)>0\pi(a\mid s)>0\Rightarrow b(a\mid s)>0π(a∣s)>0⇒b(a∣s)>0) for the expectation identities.
  • 0≤γ<10\le\gamma<10≤γ<1 for (11.21); the episodic γ=1\gamma=1γ=1 case is not stated.
  • In (11.24) the return enters through conditional distributions νs\nu_sνs​ with finite second moment and mean vπ(s)v_\pi(s)vπ​(s).

Readings and corrections:

  • (11.10) is introduced as minimizing "the VE", but its displayed objective is an unweighted sum over the two states. The displayed sum is formalized.
  • The sentence on p. 269, "there always exists an approximate value function with zero PBE", needs X⊤D(I−γPπ)X\mathbf X^\top\mathbf D(I-\gamma P_\pi)\mathbf XX⊤D(I−γPπ​)X invertible, which can fail off-policy. Counterexample: two states with features x=1,2x = 1, 2x=1,2, both moving to the second state under π\piπ, rewards rπ=(1,0)r_\pi = (1, 0)rπ​=(1,0), γ=34\gamma = \tfrac34γ=43​ and μ=(23,13)\mu = (\tfrac23, \tfrac13)μ=(32​,31​). Then X⊤Dδˉw=23\mathbf X^\top\mathbf D\bar\delta_{\mathbf w} = \tfrac23X⊤Dδˉw​=32​ for every w\mathbf ww, so PBE(w)=(23)2/2>0\mathrm{PBE}(\mathbf w) = (\tfrac23)^2/2 > 0PBE(w)=(32​)2/2>0 everywhere. The sentence is not formalized.
  • The Example 11.4 MRPs are transcribed from the figure on p. 276. Their equal data distributions are stated through all finite prefixes of the observed stream (feature vectors and rewards) from the stationary distributions; the BE minimizer of the second MRP as γ→1\gamma\to1γ→1 is not formalized.
  • Baird's counterexample (divergence shown by simulation), Gradient-TD and Emphatic-TD convergence (asserted with references) are out of scope.

A trivializing formalization is ruled out: vπv_\pivπ​ is not defined as a fixed point of BπB_\piBπ​, the PBE is not stated without the invertibility hypothesis, and the gradient is a derivative of the defined PBE, not a restated formula.

Welcome contributions: proofs of the milestones, and a general lemma on the gradient of h⊤C−1h\mathbf h^\top\mathbf C^{-1}\mathbf hh⊤C−1h for affine h\mathbf hh and symmetric invertible C\mathbf CC, which is reusable well beyond this mission.

Selected references

  • R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018, ISBN 9780262039246, Chapter 11. http://incompleteideas.net/book/the-book-2nd.html
  • L. Baird, Residual algorithms: Reinforcement learning with function approximation, ICML 1995. https://doi.org/10.1016/B978-1-55860-377-6.50013-X
  • J. N. Tsitsiklis and B. Van Roy, Feature-based methods for large scale dynamic programming, Machine Learning 22, 1996. https://doi.org/10.1007/BF00114724
  • R. S. Sutton, H. R. Maei, D. Precup, S. Bhatnagar, D. Silver, Cs. Szepesvári, E. Wiewiora, Fast gradient-descent methods for temporal-difference learning with linear function approximation, ICML 2009. https://doi.org/10.1145/1553374.1553501
  • R. S. Sutton, A. R. Mahmood, M. White, An emphatic approach to the problem of off-policy temporal-difference learning, JMLR 17, 2016. https://jmlr.org/papers/v17/14-488.html
14 thms2 active usersReviewed
Linear algebraMachine LearningMarkov Chain·Captain: mikedeng1

Reinforcement Learning: An Introduction VIII: The TD Fixed Point of Linear Semi-gradient TD(0) and Its Error BoundTextbook

Motivation

Reinforcement learning methods estimate the value function vπv_\pivπ​ of a policy π\piπ: the expected discounted sum of future rewards from each state. When the state space is large, vπv_\pivπ​ cannot be stored as a table and is approximated by a parametrized function. The most studied case is linear function approximation, where each state sss carries a feature vector x(s)∈Rd\mathbf x(s) \in \mathbb R^dx(s)∈Rd and the estimate is v^(s,w)=w⊤x(s)\hat v(s, \mathbf w) = \mathbf w^\top \mathbf x(s)v^(s,w)=w⊤x(s). Temporal-difference learning with this approximation, linear semi-gradient TD(0), is one of the basic algorithms of the field, and Chapter 9 of Sutton and Barto's Reinforcement Learning: An Introduction (2nd ed., MIT Press, 2018) presents its analysis: where the algorithm can converge, why that point exists, and how good it is.

The history is short. Sutton (1988, doi:10.1007/BF00115009) introduced TD learning and showed positive definiteness of the matrix governing its expected update. Dayan (1992, doi:10.1007/BF00992701) extended convergence to TD(λ). Tsitsiklis and Van Roy (1997, doi:10.1109/9.580874) proved convergence with probability one for linear TD(λ) under on-policy sampling and bounded the error of the limit. Bradtke and Barto (1996) introduced least-squares TD (LSTD), which computes the same limit directly.

Setting

A finite Markov decision process has finite sets of states S\mathcal SS, actions A\mathcal AA and rewards R⊂R\mathcal R \subset \mathbb RR⊂R, and dynamics p(s′,r∣s,a)p(s', r \mid s, a)p(s′,r∣s,a), the probability of next state s′s's′ and reward rrr after action aaa in state sss. A policy π(a∣s)\pi(a \mid s)π(a∣s) is a probability distribution over actions for each state. It induces a Markov chain on states with transition matrix P\mathbf PP, P(s,s′)=p(s′∣s)=∑aπ(a∣s)∑rp(s′,r∣s,a)\mathbf P(s, s') = p(s' \mid s) = \sum_a \pi(a \mid s) \sum_r p(s', r \mid s, a)P(s,s′)=p(s′∣s)=∑a​π(a∣s)∑r​p(s′,r∣s,a), and expected one-step reward rπ(s)r_\pi(s)rπ​(s). For a discount rate 0≤γ<10 \le \gamma < 10≤γ<1, the true value is vπ(s)=∑k≥0γk(Pkrπ)(s)v_\pi(s) = \sum_{k \ge 0} \gamma^k (\mathbf P^k r_\pi)(s)vπ​(s)=∑k≥0​γk(Pkrπ​)(s), the expected discounted return.

A state distribution μ\muμ is stationary if μ⊤P=μ⊤\mu^\top \mathbf P = \mu^\topμ⊤P=μ⊤; write D=diag(μ)\mathbf D = \mathrm{diag}(\mu)D=diag(μ). The feature matrix X\mathbf XX is the ∣S∣×d|\mathcal S| \times d∣S∣×d matrix with rows x(s)\mathbf x(s)x(s). The mean square value error of a weight vector is

VE‾(w)=∑sμ(s) [vπ(s)−w⊤x(s)]2.\overline{\mathrm{VE}}(\mathbf w) = \sum_{s} \mu(s)\,[v_\pi(s) - \mathbf w^\top \mathbf x(s)]^2 .VE(w)=s∑​μ(s)[vπ​(s)−w⊤x(s)]2.

Linear semi-gradient TD(0) updates wt+1=wt+α(Rt+1+γwt⊤xt+1−wt⊤xt)xt\mathbf w_{t+1} = \mathbf w_t + \alpha(R_{t+1} + \gamma \mathbf w_t^\top \mathbf x_{t+1} - \mathbf w_t^\top \mathbf x_t)\mathbf x_twt+1​=wt​+α(Rt+1​+γwt⊤​xt+1​−wt⊤​xt​)xt​. In steady state its expected update involves

b=E[Rt+1xt],A=E[xt(xt−γxt+1)⊤],\mathbf b = \mathbb E[R_{t+1}\mathbf x_t], \qquad \mathbf A = \mathbb E[\mathbf x_t(\mathbf x_t - \gamma \mathbf x_{t+1})^\top],b=E[Rt+1​xt​],A=E[xt​(xt​−γxt+1​)⊤],

and the TD fixed point is wTD=A−1b\mathbf w_{\mathrm{TD}} = \mathbf A^{-1}\mathbf bwTD​=A−1b. A real square matrix MMM, not necessarily symmetric, is positive definite if y⊤My>0y^\top M y > 0y⊤My>0 for every y≠0y \ne 0y=0. The key matrix is D(I−γP)\mathbf D(\mathbf I - \gamma\mathbf P)D(I−γP).

Formalization targets

Goal: the TD fixed point exists and its error bound (9.12), (9.14)

Under the hypotheses above, with every μ(s)>0\mu(s) > 0μ(s)>0 and linearly independent feature columns, A\mathbf AA is invertible, b=AwTD\mathbf b = \mathbf A \mathbf w_{\mathrm{TD}}b=AwTD​, and

VE‾(wTD)≤11−γmin⁡wVE‾(w).\overline{\mathrm{VE}}(\mathbf w_{\mathrm{TD}}) \le \frac{1}{1-\gamma}\min_{\mathbf w} \overline{\mathrm{VE}}(\mathbf w).VE(wTD​)≤1−γ1​wmin​VE(w).

Milestones

  1. The expected update (9.13): E[wt+1∣wt]=(I−αA)wt+αb\mathbb E[\mathbf w_{t+1} \mid \mathbf w_t] = (\mathbf I - \alpha \mathbf A)\mathbf w_t + \alpha \mathbf bE[wt+1​∣wt​]=(I−αA)wt​+αb.
  2. The matrix form A=X⊤D(I−γP)X\mathbf A = \mathbf X^\top \mathbf D(\mathbf I - \gamma \mathbf P)\mathbf XA=X⊤D(I−γP)X.
  3. The criterion of Sutton (1988): positive diagonal, nonpositive off-diagonal entries, positive row sums and nonnegative column sums give positive definiteness.
  4. The column sums of the key matrix, 1⊤D(I−γP)=(1−γ)μ⊤\mathbf 1^\top \mathbf D(\mathbf I - \gamma \mathbf P) = (1-\gamma)\mu^\top1⊤D(I−γP)=(1−γ)μ⊤.
  5. The key matrix and A\mathbf AA are positive definite.
  6. A positive definite A\mathbf AA is invertible and A−1b\mathbf A^{-1}\mathbf bA−1b is the unique solution of b=Aw\mathbf b = \mathbf A \mathbf wb=Aw (9.12).
  7. The Sherman–Morrison update (9.22) of the LSTD inverse A^t−1\hat{\mathbf A}_t^{-1}A^t−1​.

Significance

Positive definiteness of A\mathbf AA is the reason on-policy linear TD(0) is stable: it makes the expected iteration contract toward the fixed point for small step sizes, and it guarantees that the fixed point exists and is unique. The error bound (9.14) quantifies the price of bootstrapping: the limit of TD can be worse than the best linear approximation, but by at most the factor 1/(1−γ)1/(1-\gamma)1/(1−γ). The same objects A\mathbf AA, b\mathbf bb and the key matrix reappear in LSTD, in the analysis of off-policy divergence (Chapter 11 of the book, where D\mathbf DD is no longer the stationary distribution of P\mathbf PP and positive definiteness fails), and in gradient-TD methods.

All results here are known. The book gives the positive definiteness argument in a box and cites (9.14) without proof. None of them has a machine-checked proof on the platform; the general Woodbury identity (FamousTheorems.woodbury_identity) is available, and (9.22) is its rank-one case written for the LSTD recursion. The mission produces a formal account of the finite-state theory of linear TD(0), with every hypothesis the book leaves implicit stated.

Difficulty

The key matrix D(I−γP)\mathbf D(\mathbf I - \gamma \mathbf P)D(I−γP) is not symmetric, so the usual tools for symmetric positive definite matrices do not apply directly, and A\mathbf AA is positive definite only because of the specific interplay between D\mathbf DD and P\mathbf PP: if μ\muμ is replaced by a non-stationary distribution the claim is false (this is the off-policy counterexample of Chapter 11). The error bound (9.14) is not a consequence of positive definiteness alone. The TD fixed point is not the minimizer of VE‾\overline{\mathrm{VE}}VE, and VE‾(wTD)\overline{\mathrm{VE}}(\mathbf w_{\mathrm{TD}})VE(wTD​) has to be compared with the error of the μ\muμ-weighted projection of vπv_\pivπ​, which requires controlling P\mathbf PP in the μ\muμ-weighted norm. The book gives no argument for this step.

Formalization scope

The Lean development lives in the namespace SuttonBartoRL.LinearTD. The MDP has four-argument dynamics p(s′,r∣s,a)p(s', r \mid s, a)p(s′,r∣s,a) with a finite reward set and one action set for all states; policies are stochastic. vπv_\pivπ​ is defined from expected discounted returns as the series ∑kγkPkrπ\sum_k \gamma^k \mathbf P^k r_\pi∑k​γkPkrπ​, never from a Bellman equation or from wTD\mathbf w_{\mathrm{TD}}wTD​. A\mathbf AA and b\mathbf bb are defined as the book's steady-state expectations (9.11), as finite sums over μ\muμ, π\piπ and ppp; the matrix form is a milestone, not a definition. Features are a matrix Matrix S (Fin d) ℝ with rows x(s)\mathbf x(s)x(s); linear independence of its columns is LinearIndependent ℝ Xᵀ. Positive definiteness is a custom predicate ∀y≠0, 0<y⊤My\forall y \ne 0,\ 0 < y^\top M y∀y=0, 0<y⊤My, not Mathlib's Matrix.PosDef, which requires symmetry. The minimum in (9.14) is expressed by quantifying over every w\mathbf ww. The matrix inverse is Mathlib's, which is zero on singular matrices; the goal therefore asserts invertibility of A\mathbf AA explicitly.

Hypotheses the book leaves implicit and the statements make explicit: 0≤γ<10 \le \gamma < 10≤γ<1 (the continuing case); μ\muμ a stationary distribution of the chain induced by π\piπ with μ(s)>0\mu(s) > 0μ(s)>0 for every sss (otherwise the key matrix is only positive semidefinite); linearly independent feature columns (the book's "degenerate cases", p. 205). The box calls the off-diagonal entries of the key matrix "negative"; they are zero wherever p(s′∣s)=0p(s' \mid s) = 0p(s′∣s)=0, so the criterion is stated with nonpositive entries. The book's sentence that εI\varepsilon\mathbf IεI "ensures that A^t\hat{\mathbf A}_tA^t​ is always invertible" (p. 229) is false in general, because the summands xk(xk−γxk+1)⊤\mathbf x_k(\mathbf x_k - \gamma \mathbf x_{k+1})^\topxk​(xk​−γxk+1​)⊤ are not positive semidefinite: with d=1d = 1d=1, ε=1/10\varepsilon = 1/10ε=1/10, γ=1/2\gamma = 1/2γ=1/2, x0=1\mathbf x_0 = 1x0​=1, x1=11/5\mathbf x_1 = 11/5x1​=11/5 one gets A^1=0\hat{\mathbf A}_1 = 0A^1​=0. It is not stated; (9.22) carries invertibility of A^t−1\hat{\mathbf A}_{t-1}A^t−1​ and a nonzero denominator as hypotheses.

A statement in which vπv_\pivπ​ is defined as the solution of the projected equation, or in which A\mathbf AA is assumed invertible or positive definite, would make the goal trivial or empty; neither is done. Convergence of the stochastic algorithm with probability one is not stated, since the book says it needs conditions and a step-size schedule it does not give. The bound for the episodic case and for other bootstrapping methods (p. 208) is stated only by reference in the book and is not a target.

Useful infrastructure: the μ\muμ-weighted inner product and orthogonal projection onto the column space of X\mathbf XX, the non-expansiveness of a stochastic matrix in the norm of its stationary distribution, and the positive definiteness criterion for non-symmetric matrices. All of these are reusable in the off-policy and average-reward chapters of the book. Contributions of these lemmas, and of alternative proofs of the milestones, are welcome.

Selected references

  • Richard S. Sutton and Andrew G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018, ISBN 9780262039246, §§9.2, 9.4, 9.8.
  • Richard S. Sutton, Learning to predict by the methods of temporal differences, Machine Learning 3, 1988. doi:10.1007/BF00115009
  • John N. Tsitsiklis and Benjamin Van Roy, An analysis of temporal-difference learning with function approximation, IEEE Transactions on Automatic Control 42(5), 1997. doi:10.1109/9.580874
  • Steven J. Bradtke and Andrew G. Barto, Linear least-squares algorithms for temporal difference learning, Machine Learning 22, 1996. doi:10.1007/BF00114723
  • Richard S. Varga, Matrix Iterative Analysis, Prentice-Hall, 1962.
11 thms2 active usersReviewed
Machine LearningProbabilityStatistics·Captain: mikedeng1

Reinforcement Learning: An Introduction V: Off-policy Prediction by Importance SamplingTextbook

Motivation

Reinforcement learning methods must explore in order to find good behaviour, yet the quantity they usually want to evaluate is the value of a different, often deterministic, policy. Off-policy prediction separates the two roles: episodes are generated by a behaviour policy bbb, and the goal is the value function vπv_\pivπ​ of a target policy π\piπ. Almost every off-policy method in Sutton and Barto's Reinforcement Learning: An Introduction (2nd ed., MIT Press, 2018), and in the literature that follows it, rests on importance sampling: a return observed under bbb is reweighted by the relative probability of its trajectory under π\piπ and bbb. Section 5.5 of the book introduces the idea for Monte Carlo prediction, §5.6 gives the incremental form of the weighted estimator, and §§5.8–5.9 refine the weights using the internal structure of the return: discounting-aware importance sampling, after Sutton, Mahmood, Precup and van Hasselt (2014), and per-decision importance sampling, introduced by Precup, Sutton and Singh (2000). The book's remarks on the variance of the two estimators (p. 105) cite Precup, Sutton and Dasgupta (2001). Later chapters (7, 11, 12) reuse the same ratios for nnn-step, gradient-TD and eligibility-trace methods.

This mission is the fifth in a series formalizing the book's central mathematical claims. It covers §§5.5–5.9 (pp. 103–115).

Setting

A finite Markov decision process has finite sets of states S\mathcal SS (terminal states included), actions A\mathcal AA and rewards R⊂R\mathcal R \subset \mathbb RR⊂R, and dynamics p(s′,r∣s,a)p(s', r \mid s, a)p(s′,r∣s,a): for each (s,a)(s, a)(s,a) a probability distribution over next state and reward. The state-transition probability is p(s′∣s,a)=∑rp(s′,r∣s,a)p(s' \mid s, a) = \sum_r p(s', r \mid s, a)p(s′∣s,a)=∑r​p(s′,r∣s,a). A policy μ\muμ gives a distribution μ(⋅∣s)\mu(\cdot \mid s)μ(⋅∣s) over actions in every state.

An episode from a start state sss is a sequence S0=s,A0,R1,S1,…,AT−1,RT,STS_0 = s, A_0, R_1, S_1, \dots, A_{T-1}, R_T, S_TS0​=s,A0​,R1​,S1​,…,AT−1​,RT​,ST​ in which S0,…,ST−1S_0, \dots, S_{T-1}S0​,…,ST−1​ are nonterminal and STS_TST​ is terminal. Under μ\muμ it has probability ∏k=0T−1μ(Ak∣Sk) p(Sk+1,Rk+1∣Sk,Ak)\prod_{k=0}^{T-1} \mu(A_k \mid S_k)\, p(S_{k+1}, R_{k+1} \mid S_k, A_k)∏k=0T−1​μ(Ak​∣Sk​)p(Sk+1​,Rk+1​∣Sk​,Ak​). The return is G0=∑k=0T−1γkRk+1G_0 = \sum_{k=0}^{T-1} \gamma^k R_{k+1}G0​=∑k=0T−1​γkRk+1​ with discount rate γ∈[0,1]\gamma \in [0, 1]γ∈[0,1], and the value vπ(s)v_\pi(s)vπ​(s) is the expected return of an episode generated by π\piπ from sss.

The behaviour policy covers the target policy if π(a∣s)>0\pi(a \mid s) > 0π(a∣s)>0 implies b(a∣s)>0b(a \mid s) > 0b(a∣s)>0. The importance-sampling ratio of decisions 0,…,j0, \dots, j0,…,j is

ρ0:j=∏k=0jπ(Ak∣Sk)b(Ak∣Sk),\rho_{0:j} = \prod_{k=0}^{j} \frac{\pi(A_k \mid S_k)}{b(A_k \mid S_k)},ρ0:j​=k=0∏j​b(Ak​∣Sk​)π(Ak​∣Sk​)​,

and the per-decision return weights each reward only by the ratio of the decisions that precede it:

G~0=ρ0:0R1+γρ0:1R2+⋯+γT−1ρ0:T−1RT.\tilde G_0 = \rho_{0:0} R_1 + \gamma \rho_{0:1} R_2 + \dots + \gamma^{T-1} \rho_{0:T-1} R_T .G~0​=ρ0:0​R1​+γρ0:1​R2​+⋯+γT−1ρ0:T−1​RT​.

The book writes these objects at a general time ttt and conditions on St=sS_t = sSt​=s; by the Markov property this is the same as starting the episode at sss, which is what the formal statements do.

Formalization targets

Goal: unbiasedness of ordinary and per-decision importance sampling

For episodes generated by bbb from sss,

Eb[ρ0:T−1G0∣S0=s]=vπ(s)=Eb[G~0∣S0=s].\mathbb E_b\bigl[\rho_{0:T-1} G_0 \mid S_0 = s\bigr] = v_\pi(s) = \mathbb E_b\bigl[\tilde G_0 \mid S_0 = s\bigr].Eb​[ρ0:T−1​G0​∣S0​=s]=vπ​(s)=Eb​[G~0​∣S0​=s].

The first equality is Eq. (5.4) (p. 104); the second is the statement E[ρt:T−1Gt]=E[G~t]\mathbb E[\rho_{t:T-1}G_t] = \mathbb E[\tilde G_t]E[ρt:T−1​Gt​]=E[G~t​] of §5.9 (p. 114).

Milestones

  1. (5.3): the trajectory probability is a product, and the ratio of trajectory probabilities under π\piπ and bbb is ρ0:T−1\rho_{0:T-1}ρ0:T−1​, independent of the dynamics.
  2. (5.4) alone.
  3. (5.13): ∑ab(a∣x) π(a∣x)/b(a∣x)=∑aπ(a∣x)=1\sum_a b(a \mid x)\, \pi(a \mid x)/b(a \mid x) = \sum_a \pi(a \mid x) = 1∑a​b(a∣x)π(a∣x)/b(a∣x)=∑a​π(a∣x)=1 under coverage.
  4. (5.14) and its kkk-th form: Eb[ρ0:T−1Rk]=Eb[ρ0:k−1Rk]\mathbb E_b[\rho_{0:T-1} R_k] = \mathbb E_b[\rho_{0:k-1} R_k]Eb​[ρ0:T−1​Rk​]=Eb​[ρ0:k−1​Rk​] for every k≥1k \ge 1k≥1 (Exercise 5.13).
  5. Example 5.5: in a one-state MDP with a loop, vπ(s)=1v_\pi(s) = 1vπ​(s)=1 and Eb[ρ0:T−1G0]=1\mathbb E_b[\rho_{0:T-1}G_0] = 1Eb​[ρ0:T−1​G0​]=1, yet Eb[(ρ0:T−1G0)2]=∞\mathbb E_b[(\rho_{0:T-1}G_0)^2] = \inftyEb​[(ρ0:T−1​G0​)2]=∞.
  6. (5.7)–(5.8): the incremental rule Vn+1=Vn+(Wn/Cn)(Gn−Vn)V_{n+1} = V_n + (W_n/C_n)(G_n - V_n)Vn+1​=Vn​+(Wn​/Cn​)(Gn​−Vn​) computes the weighted average ∑k<nWkGk/∑k<nWk\sum_{k<n} W_k G_k / \sum_{k<n} W_k∑k<n​Wk​Gk​/∑k<n​Wk​ (Exercise 5.10).
  7. §5.8: Gt=(1−γ)∑h=t+1T−1γh−t−1Gˉt:h+γT−t−1Gˉt:TG_t = (1-\gamma)\sum_{h=t+1}^{T-1}\gamma^{h-t-1}\bar G_{t:h} + \gamma^{T-t-1}\bar G_{t:T}Gt​=(1−γ)∑h=t+1T−1​γh−t−1Gˉt:h​+γT−t−1Gˉt:T​ with flat partial returns Gˉt:h=Rt+1+⋯+Rh\bar G_{t:h} = R_{t+1} + \dots + R_hGˉt:h​=Rt+1​+⋯+Rh​.

Significance

Eq. (5.4) is the reason the first-visit ordinary importance-sampling estimator (5.5) is unbiased, and it is the template for every importance-sampling correction in the rest of the book. The per-decision identity shows that an estimator with fewer ratio factors per reward, (5.15), has the same expectation, which is the starting point for per-decision and control-variate methods for multi-step off-policy learning (Precup, Sutton and Singh 2000). Example 5.5 shows that unbiasedness says nothing about variance: the ordinary estimator can have infinite variance on a two-action problem, which motivates weighted importance sampling and the incremental weighted update of §5.6.

The results of these sections are classical and proved informally in the book, partly as exercises (5.10, 5.13) left without solution. No machine-checked version exists on the platform: a search for importance sampling, off-policy and per-decision returned no statements. The mission produces a formal trajectory model of an episodic MDP under two policies, which later missions on nnn-step off-policy returns and off-policy traces can reuse.

Difficulty

Eq. (5.4) itself is a termwise identity: for every episode, Pr⁡b(episode) ρ0:T−1=Pr⁡π(episode)\Pr_b(\text{episode})\,\rho_{0:T-1} = \Pr_\pi(\text{episode})Prb​(episode)ρ0:T−1​=Prπ​(episode) under coverage. The per-decision identity is not termwise. The later factors of ρ0:T−1\rho_{0:T-1}ρ0:T−1​ multiply a reward that was received before the corresponding decisions, and removing them requires summing over all continuations of an episode prefix, of every remaining length, and using that each factor has conditional expectation one (5.13) and that the continuation terminates with probability one. The obvious attempt, cancelling the factors episode by episode, fails: on a single episode ρ0:T−1R1\rho_{0:T-1}R_1ρ0:T−1​R1​ and ρ0:0R1\rho_{0:0}R_1ρ0:0​R1​ differ.

In Example 5.5 the episodes have no length bound, so the expected square is an infinite series over episode lengths whose divergence must be shown directly.

Formalization scope

  • States, actions and rewards are finite types; the terminal states are a finite subset of the state type. Policies are stochastic, one action set is used in every state, and vπv_\pivπ​ is defined as an expected return, never as the solution of a Bellman equation.
  • Expectations are series over episode lengths of finite sums over episodes. Lean assigns 000 to a divergent series, so the goal and milestones 2 and 4 assume that under bbb every episode from sss terminates within a fixed number HHH of steps with probability one. The book leaves termination implicit; this bounded-horizon hypothesis is a restriction relative to the book's episodic setting and is stated as such. Example 5.5, whose episodes are unbounded, is stated without it, with the expected square in [0,∞][0, \infty][0,∞].
  • The discount rate is kept general in [0,1][0, 1][0,1].
  • The importance-sampling ratio uses real division; a factor with b(Ak∣Sk)=0b(A_k \mid S_k) = 0b(Ak​∣Sk​)=0 evaluates to 000 in Lean, but such episodes have probability 000 under bbb.
  • The flat-partial-return decomposition is an algebraic identity and is stated for every real γ\gammaγ, which is more general than the book's "for any γ∈[0,1)\gamma \in [0,1)γ∈[0,1)".
  • The incremental weighted update is stated with nonnegative weights and W1>0W_1 > 0W1​>0. The book's C0=0C_0 = 0C0​=0 makes (5.8) divide by zero at n=1n = 1n=1 when W1=0W_1 = 0W1​=0; the hypothesis excludes that case.
  • A trivializing formalization is ruled out: vπv_\pivπ​ is the expected return of π\piπ's own episodes, the ratio is computed from the episode, and the per-decision identity, which carries the chapter's content beyond (5.4), is part of the goal.
  • Not stated: the bias and variance comparisons of ordinary and weighted importance sampling (p. 105) and the discounting-aware estimators (5.9)–(5.10) as estimators; these are statistical claims about estimators over a random number of visits that the book does not make precise.

Contributions welcome: proofs of the milestones, a general measure-theoretic version of the trajectory model without the bounded-horizon hypothesis, and variants for action values qπq_\piqπ​ (Exercise 5.6).

Selected references

  • R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018, ISBN 9780262039246, §§5.5–5.9, pp. 103–115. http://incompleteideas.net/book/the-book-2nd.html
  • D. Precup, R. S. Sutton and S. Singh, Eligibility Traces for Off-Policy Policy Evaluation, Proceedings of the 17th International Conference on Machine Learning (ICML), 2000, pp. 759–766 (cited in the book's bibliography).
  • D. Precup, R. S. Sutton and S. Dasgupta, Off-Policy Temporal-Difference Learning with Function Approximation, Proceedings of the 18th International Conference on Machine Learning (ICML), 2001, pp. 417–424 (cited in the book, p. 105).
14 thms2 active usersReviewed
Convex OptimizationOperations ResearchOptimization·Captain: mikedeng1

Twice Regularized MDPs and the Equivalence Between Robustness and Regularization 1: The Robust Value Function Is the Optimum of a Policy- and Value-Regularized Convex ProgramResearch Paper

Motivation

A Markov decision process (MDP) is solved for one model of its dynamics and rewards, but in practice that model is estimated from data, and a policy that is optimal for the estimate can perform poorly on the true system (Mannor et al., 2007). Robust MDPs address this by evaluating a policy against the worst model in an uncertainty set U\mathcal UU (Iyengar, 2005; Nilim and El Ghaoui, 2005; Wiesemann, Kuhn and Rustem, 2013). Robust planning, however, solves an inner optimization over U\mathcal UU at every Bellman update, which is expensive and does not scale to learning settings.

A separate line of work regularizes the policy (entropy, KL, Tsallis penalties) and observes empirically that regularized policies are robust to perturbations (Geist, Scherrer and Pietquin, 2019). Derman, Geist and Mannor (arXiv:2110.06267, NeurIPS 2021) make this precise: for uncertainty sets centred at a nominal model, the robust value function is the solution of a regularized problem posed on the nominal model alone, with a regularizer that is the support function of the uncertainty set. This mission formalizes that equivalence: Proposition 3.1, Theorem 3.1 and Theorem 4.1 of the paper.

Setting

Let S\mathcal SS and A\mathcal AA be finite sets of states and actions, A\mathcal AA nonempty, and X:=S×A\mathcal X := \mathcal S\times\mathcal AX:=S×A. Fix a discount factor γ∈(0,1)\gamma\in(0,1)γ∈(0,1) and a strictly positive initial distribution μ0∈ΔS\mu_0\in\Delta_{\mathcal S}μ0​∈ΔS​. A transition kernel PPP assigns to every pair (s,a)(s,a)(s,a) a probability distribution P(⋅∣s,a)P(\cdot\mid s,a)P(⋅∣s,a) on S\mathcal SS; a reward is r∈RXr\in\mathbb R^{\mathcal X}r∈RX. A policy π∈ΔAS\pi\in\Delta_{\mathcal A}^{\mathcal S}π∈ΔAS​ assigns to every state an action distribution πs\pi_sπs​.

For v∈RSv\in\mathbb R^{\mathcal S}v∈RS write rπ(s)=∑aπs(a)r(s,a)r^\pi(s) = \sum_a\pi_s(a)r(s,a)rπ(s)=∑a​πs​(a)r(s,a), Pπ(s′∣s)=∑aπs(a)P(s′∣s,a)P^\pi(s'\mid s) = \sum_a\pi_s(a)P(s'\mid s,a)Pπ(s′∣s)=∑a​πs​(a)P(s′∣s,a), and define the evaluation Bellman operator

T(P,r)πv:=rπ+γPπv.T^\pi_{(P,r)}v := r^\pi + \gamma P^\pi v .T(P,r)π​v:=rπ+γPπv.

The inner product on RS\mathbb R^{\mathcal S}RS is ⟨v,μ⟩=∑sv(s)μ(s)\langle v,\mu\rangle = \sum_s v(s)\mu(s)⟨v,μ⟩=∑s​v(s)μ(s), and the support function of a set C⊆RιC\subseteq\mathbb R^{\iota}C⊆Rι is σC(y)=max⁡a∈C⟨a,y⟩\sigma_C(y) = \max_{a\in C}\langle a,y\rangleσC​(y)=maxa∈C​⟨a,y⟩.

Given a set U\mathcal UU of models (P,r)(P,r)(P,r), the robust Bellman operator is

[Tπ,Uv](s):=min⁡(P,r)∈UT(P,r)πv(s),[T^{\pi,\mathcal U}v](s) := \min_{(P,r)\in\mathcal U}T^\pi_{(P,r)}v(s),[Tπ,Uv](s):=(P,r)∈Umin​T(P,r)π​v(s),

and the robust value function vπ,Uv^{\pi,\mathcal U}vπ,U is its fixed point. Around a nominal model (P0,r0)(P_0,r_0)(P0​,r0​), an s-rectangular uncertainty set U=(P0+P)×(r0+R)\mathcal U = (P_0+\mathcal P)\times(r_0+\mathcal R)U=(P0​+P)×(r0​+R) is given by sets Ps⊆RX\mathcal P_s\subseteq\mathbb R^{\mathcal X}Ps​⊆RX and Rs⊆RA\mathcal R_s\subseteq\mathbb R^{\mathcal A}Rs​⊆RA, one per state: its models are P(s′∣s,a)=P0(s′∣s,a)+Ps(s′,a)P(s'\mid s,a) = P_0(s'\mid s,a)+P_s(s',a)P(s′∣s,a)=P0​(s′∣s,a)+Ps​(s′,a) and r(s,a)=r0(s,a)+rs(a)r(s,a) = r_0(s,a)+r_s(a)r(s,a)=r0​(s,a)+rs​(a), with Ps∈PsP_s\in\mathcal P_sPs​∈Ps​ and rs∈Rsr_s\in\mathcal R_srs​∈Rs​ chosen independently for each sss. Finally [v⋅πs](s′,a):=v(s′)πs(a)[v\cdot\pi_s](s',a) := v(s')\pi_s(a)[v⋅πs​](s′,a):=v(s′)πs​(a).

Formalization targets

Goal: Theorem 4.1 (general robust MDP)

For U=(P0+P)×(r0+R)\mathcal U = (P_0+\mathcal P)\times(r_0+\mathcal R)U=(P0​+P)×(r0​+R) and every policy π\piπ, Tπ,UT^{\pi,\mathcal U}Tπ,U has a unique fixed point vπ,Uv^{\pi,\mathcal U}vπ,U, and it is the optimal solution of

max⁡v∈RS⟨v,μ0⟩s.t.v(s)≤T(P0,r0)πv(s)−σRs(−πs)−σPs(−γv⋅πs)∀s∈S.(2)\max_{v\in\mathbb R^{\mathcal S}}\langle v,\mu_0\rangle\quad\text{s.t.}\quad v(s)\le T^\pi_{(P_0,r_0)}v(s)-\sigma_{\mathcal R_s}(-\pi_s)-\sigma_{\mathcal P_s}(-\gamma v\cdot\pi_s)\quad\forall s\in\mathcal S. \tag{2}v∈RSmax​⟨v,μ0​⟩s.t.v(s)≤T(P0​,r0​)π​v(s)−σRs​​(−πs​)−σPs​​(−γv⋅πs​)∀s∈S.(2)

Milestones

  1. Proposition 3.1. For any uncertainty set U=P×R\mathcal U = \mathcal P\times\mathcal RU=P×R with P\mathcal PP a nonempty compact set of kernels and R\mathcal RR a nonempty compact set of rewards, vπ,Uv^{\pi,\mathcal U}vπ,U is the optimal solution of the robust program \max_{v}\langle v,\mu_0\rangle\quad\text{s.t.}\quad v\le T^\pi_{(P,r)}v\ \ \forall(P,r)\in\mathcal U. \tag{$P_{\mathcal U}$}
  2. Theorem 3.1. For U={P0}×(r0+R)\mathcal U=\{P_0\}\times(r_0+\mathcal R)U={P0​}×(r0​+R), vπ,Uv^{\pi,\mathcal U}vπ,U is the optimal solution of max⁡v⟨v,μ0⟩\max_v\langle v,\mu_0\ranglemaxv​⟨v,μ0​⟩ s.t. v(s)≤T(P0,r0)πv(s)−σRs(−πs)v(s)\le T^\pi_{(P_0,r_0)}v(s)-\sigma_{\mathcal R_s}(-\pi_s)v(s)≤T(P0​,r0​)π​v(s)−σRs​​(−πs​) for all sss.
  3. Robust counterpart (proof of Theorem 4.1, App. B.1). For every vvv and sss,
max⁡(P,r)∈U{v(s)−rπ(s)−γPπv(s)}=σPs(−γv⋅πs)+σRs(−πs)+v(s)−T(P0,r0)πv(s).\max_{(P,r)\in\mathcal U}\{v(s)-r^\pi(s)-\gamma P^\pi v(s)\} = \sigma_{\mathcal P_s}(-\gamma v\cdot\pi_s)+\sigma_{\mathcal R_s}(-\pi_s)+v(s)-T^\pi_{(P_0,r_0)}v(s).(P,r)∈Umax​{v(s)−rπ(s)−γPπv(s)}=σPs​​(−γv⋅πs​)+σRs​​(−πs​)+v(s)−T(P0​,r0​)π​v(s).

Theorem 3.1 is the special case Ps={0}\mathcal P_s=\{0\}Ps​={0} of the goal; it is listed separately because it is the paper's statement that policy regularization is equivalent to reward uncertainty.

Significance

The goal says that a robust MDP with s-rectangular uncertainty in both reward and transitions is a regularized MDP on the nominal model, with two regularizers: a policy regularizer σRs(−πs)\sigma_{\mathcal R_s}(-\pi_s)σRs​​(−πs​) coming from reward uncertainty, and a regularizer σPs(−γv⋅πs)\sigma_{\mathcal P_s}(-\gamma v\cdot\pi_s)σPs​​(−γv⋅πs​) coming from transition uncertainty that depends on both the policy and the value. For ball-shaped sets these support functions are explicit (αsr∥πs∥\alpha^r_s\|\pi_s\|αsr​∥πs​∥ and αsPγ∥v∥∥πs∥\alpha^P_s\gamma\|v\|\|\pi_s\|αsP​γ∥v∥∥πs​∥, Corollary 4.1 of the paper), which leads to the twice regularized (R²) Bellman operators of Section 5 and to robust planning at the cost of non-robust planning. Theorem 3.1 also explains why standard policy regularizers (negative entropy, KL, Tsallis) yield robustness: each is the support function of a reward uncertainty set.

The results are proved in the paper (appendices A.1, A.2, B.1); none has a machine-checked proof. The mission produces formal statements and proofs of the equivalence, the robust Bellman operator's fixed-point theory for stochastic policies and general compact uncertainty sets, and a closed-form robust counterpart that later R² results can import. The paper's printed proof of Proposition 3.1 treats Tπ,UT^{\pi,\mathcal U}Tπ,U as linear in one step; a formal proof settles the statement independently of that step.

Difficulty

The obvious argument reads Proposition 3.1 as linear-programming duality, as for a single MDP. That fails: Tπ,UT^{\pi,\mathcal U}Tπ,U is a minimum of affine maps, hence concave and not affine, and the feasible set of (PU)(P_{\mathcal U})(PU​) is an intersection of infinitely many half-space systems; the argument has to go through monotonicity and contraction of Tπ,UT^{\pi,\mathcal U}Tπ,U, which in turn requires every model in U\mathcal UU to be a genuine transition kernel. For the goal, the paper invokes Fenchel–Rockafellar duality to evaluate the inner maximum; the work in Lean is to separate the maximum over the product set U\mathcal UU into per-state maxima, which needs the s-rectangular structure and attainment of every maximum (compactness), and to track the index order of the perturbation Ps(s′,a)P_s(s',a)Ps​(s′,a) against the kernel P(s′∣s,a)P(s'\mid s,a)P(s′∣s,a).

Formalization scope

  • States and actions are finite types, A nonempty; values are S → ℝ ordered pointwise; a transition array is P : S → A → S → ℝ with P s a s' =P(s′∣s,a)=P(s'\mid s,a)=P(s′∣s,a), and the kernel property is the published IsTransitionKernel; Pπ(s′∣s)P^\pi(s'\mid s)Pπ(s′∣s) is the published InducedTransition. A policy has π s ∈ stdSimplex ℝ A for every s.
  • Perturbations PsP_sPs​ are functions S × A → ℝ indexed (s′,a)(s',a)(s′,a), as in the paper's RX\mathbb R^{\mathcal X}RX; rewards perturbations are A → ℝ.
  • Minima and maxima (in Tπ,UT^{\pi,\mathcal U}Tπ,U and in σ\sigmaσ) are real sInf/sSup. Every theorem assumes the sets nonempty and compact, so these are attained; nothing is quantified over an unbounded set.
  • The robust value function is encoded as the fixed point of Tπ,UT^{\pi,\mathcal U}Tπ,U, and each theorem asserts its existence and uniqueness. The paper's definition vπ,U(s)=min⁡(P,r)∈Uv(P,r)π(s)v^{\pi,\mathcal U}(s)=\min_{(P,r)\in\mathcal U}v^\pi_{(P,r)}(s)vπ,U(s)=min(P,r)∈U​v(P,r)π​(s) (p. 4) coincides with it for rectangular sets by a cited result; the proofs use only the fixed-point property. For the non-rectangular sets of Proposition 3.1 the pointwise minimum can be strictly larger than the fixed point and is then not the optimum of (PU)(P_{\mathcal U})(PU​), so the fixed point is the object the proposition is true for.
  • "The optimal solution" means: feasible, objective-maximal, and the unique maximizer (uniqueness uses μ0>0\mu_0>0μ0​>0).
  • Disclosed hypotheses: U=P×R\mathcal U=\mathcal P\times\mathcal RU=P×R with P\mathcal PP, R\mathcal RR nonempty and compact and every transition in P\mathcal PP a kernel (Prop. 3.1); Ps\mathcal P_sPs​, Rs\mathcal R_sRs​ nonempty and compact and every perturbed row P0(⋅∣s,a)+Ps(⋅,a)P_0(\cdot\mid s,a)+P_s(\cdot,a)P0​(⋅∣s,a)+Ps​(⋅,a) in ΔS\Delta_{\mathcal S}ΔS​ (Thm 4.1); reward sets rectangular in Thm 3.1, as its proof uses. These are the robust-MDP standing assumptions of p. 4 (P⊆ΔSX\mathcal P\subseteq\Delta^{\mathcal X}_{\mathcal S}P⊆ΔSX​) and what makes "min" and "max" well defined.
  • Not drafted: Corollary 4.1, whose ℓ²-ball Ps\mathcal P_sPs​ contains perturbations that leave the simplex, so P0+PP_0+\mathcal PP0​+P is not a set of kernels; Corollary 3.1 and Proposition 3.2 (consequences after the goal; Prop. 3.2 depends on an unspecified policy parametrization).
  • A formalization that asserts only that the feasible sets of (PU)(P_{\mathcal U})(PU​) and (2) coincide, or that drops the kernel condition or the existence of the fixed point, does not count: the goal names the robust value function and its optimality.
  • "Convex" in the statement of Theorem 4.1 is descriptive and is not part of the formal goal.

Contributions welcome: the monotone-contraction fixed-point lemma for Tπ,UT^{\pi,\mathcal U}Tπ,U and the per-state separation of maxima over rectangular sets are reusable for any robust MDP mission.

Selected references

  • E. Derman, M. Geist, S. Mannor, Twice regularized MDPs and the equivalence between robustness and regularization, NeurIPS 2021. arXiv:2110.06267v1
  • G. N. Iyengar, Robust dynamic programming, Mathematics of Operations Research 30(2), 2005. doi:10.1287/moor.1040.0129
  • A. Nilim, L. El Ghaoui, Robust control of Markov decision processes with uncertain transition matrices, Operations Research 53(5), 2005. doi:10.1287/opre.1050.0216
  • W. Wiesemann, D. Kuhn, B. Rustem, Robust Markov decision processes, Mathematics of Operations Research 38(1), 2013. doi:10.1287/moor.1120.0566
  • M. Geist, B. Scherrer, O. Pietquin, A theory of regularized Markov decision processes, ICML 2019. PMLR 97
  • S. Mannor, D. Simester, P. Sun, J. N. Tsitsiklis, Bias and variance approximation in value function estimates, Management Science 53(2), 2007. doi:10.1287/mnsc.1060.0614
8 thms2 active usersReviewed
Linear algebraMachine Learning·Captain: mikedeng1

Reinforcement Learning: An Introduction XI: Dutch Traces and the Equivalence of Forward and Backward Views in Monte Carlo LearningTextbook

Motivation

Eligibility traces are one of the basic mechanisms of reinforcement learning. A trace is a short-term memory vector ztz_tzt​ with one component per weight. It records which components contributed to recent value estimates, so that an error observed now can be credited to the right components without storing the past. Chapter 12 of Sutton and Barto's Reinforcement Learning: An Introduction (2nd ed., 2018) organizes the topic around two ways of describing an algorithm. A forward view updates each state toward a target built from rewards that arrive later. A backward view makes an update at every step from the current error and the trace.

The chapter proves one exact equivalence between the two views itself, in §12.6: "This is the only equivalence of forward- and backward-views that we explicitly demonstrate in this book" (p. 301). The setting is linear Monte Carlo prediction, and the backward view uses a dutch trace. The same trace appears in true online TD(λ\lambdaλ), whose exact equivalence to the online λ\lambdaλ-return algorithm (van Seijen and Sutton, 2014; van Seijen et al., 2016) the book cites without proof. Of that proof, §12.6 "gives some of the flavor ... but is much simpler" (p. 301).

The older equivalence of §§12.1–12.2 goes back to Sutton (1988). If the weights are held fixed during an episode, the summed updates of TD(λ\lambdaλ) with accumulating traces equal the summed updates of the off-line λ\lambdaλ-return algorithm. The book leaves it as Exercises 12.3–12.4.

Setting

Fix a dimension ddd, a step size α\alphaα and an episode of length T≥1T \ge 1T≥1 with feature vectors x0,…,xT−1∈Rdx_0, \dots, x_{T-1} \in \mathbb R^dx0​,…,xT−1​∈Rd. The episode ends with a single return G∈RG \in \mathbb RG∈R ("a single reward received at the end of the episode ... and ... no discounting", p. 301). The forward view is the linear gradient Monte Carlo, or LMS, rule (12.13): from an initial w0w_0w0​,

wt+1=wt+α [G−wt⊤xt] xt,0≤t<T.w_{t+1} = w_t + \alpha\,[G - w_t^\top x_t]\,x_t, \qquad 0 \le t < T .wt+1​=wt​+α[G−wt⊤​xt​]xt​,0≤t<T.

Let Ft=I−αxtxt⊤F_t = I - \alpha x_t x_t^\topFt​=I−αxt​xt⊤​ be the fading matrix. The backward view keeps two vectors that are updated at each step in O(d)O(d)O(d) time without knowledge of GGG. The dutch trace is z0=x0z_0 = x_0z0​=x0​, zt=zt−1+(1−αzt−1⊤xt) xtz_t = z_{t-1} + (1 - \alpha z_{t-1}^\top x_t)\,x_tzt​=zt−1​+(1−αzt−1⊤​xt​)xt​. The auxiliary vector is at=at−1−αxtxt⊤at−1a_t = a_{t-1} - \alpha x_t x_t^\top a_{t-1}at​=at−1​−αxt​xt⊤​at−1​, with a0=F0w0a_0 = F_0 w_0a0​=F0​w0​.

For the second part, an episode S0,R1,S1,…,RT,STS_0, R_1, S_1, \dots, R_T, S_TS0​,R1​,S1​,…,RT​,ST​ carries states and rewards, and v^(s,w)\hat v(s, w)v^(s,w) is a differentiable value function with v^(terminal,⋅)=0\hat v(\text{terminal}, \cdot) = 0v^(terminal,⋅)=0. For one fixed weight vector www, define the following, with γ∈[0,1]\gamma \in [0,1]γ∈[0,1] and λ∈[0,1)\lambda \in [0,1)λ∈[0,1):

  • the return GtG_tGt​;
  • the nnn-step return Gt:t+nG_{t:t+n}Gt:t+n​ (12.1), with Gt:t+n=GtG_{t:t+n} = G_tGt:t+n​=Gt​ once t+n≥Tt + n \ge Tt+n≥T;
  • the λ\lambdaλ-return Gtλ=(1−λ)∑n≥1λn−1Gt:t+nG^\lambda_t = (1-\lambda)\sum_{n \ge 1}\lambda^{n-1} G_{t:t+n}Gtλ​=(1−λ)∑n≥1​λn−1Gt:t+n​ (12.2);
  • the TD error δt=Rt+1+γv^(St+1,w)−v^(St,w)\delta_t = R_{t+1} + \gamma\hat v(S_{t+1}, w) - \hat v(S_t, w)δt​=Rt+1​+γv^(St+1​,w)−v^(St​,w) (12.6);
  • the accumulating trace z−1=0z_{-1} = 0z−1​=0, zt=γλzt−1+∇v^(St,w)z_t = \gamma\lambda z_{t-1} + \nabla\hat v(S_t, w)zt​=γλzt−1​+∇v^(St​,w) (12.5).

Formalization targets

Goal: the dutch-trace equivalence (12.14), corrected

wT=aT−1+αG zT−1.w_T = a_{T-1} + \alpha G\, z_{T-1}.wT​=aT−1​+αGzT−1​.

The left side is the forward view after TTT LMS updates. On the right, aT−1a_{T-1}aT−1​ and zT−1z_{T-1}zT−1​ are produced by the incremental recursions above. The goal is about the two algorithms, not only about the closed-form product identity.

Milestones on the goal's path (§12.6, p. 302)

  1. wt+1=Ftwt+αGxtw_{t+1} = F_t w_t + \alpha G x_twt+1​=Ft​wt​+αGxt​.
  2. wT=FT−1⋯F0w0+αG∑k=0T−1FT−1⋯Fk+1xkw_T = F_{T-1}\cdots F_0 w_0 + \alpha G \sum_{k=0}^{T-1} F_{T-1}\cdots F_{k+1} x_kwT​=FT−1​⋯F0​w0​+αG∑k=0T−1​FT−1​⋯Fk+1​xk​, the first line of (12.14).
  3. zt=∑k=0tFt⋯Fk+1xkz_t = \sum_{k=0}^{t} F_t \cdots F_{k+1} x_kzt​=∑k=0t​Ft​⋯Fk+1​xk​ for the dutch-trace recursion.
  4. at=Ft⋯F0w0a_t = F_t \cdots F_0 w_0at​=Ft​⋯F0​w0​ for the auxiliary-vector recursion (corrected initialization).

Milestones on the λ\lambdaλ-return (§§12.1–12.2)

  1. (12.3): Gtλ=(1−λ)∑n=1T−t−1λn−1Gt:t+n+λT−t−1GtG^\lambda_t = (1-\lambda)\sum_{n=1}^{T-t-1}\lambda^{n-1}G_{t:t+n} + \lambda^{T-t-1}G_tGtλ​=(1−λ)∑n=1T−t−1​λn−1Gt:t+n​+λT−t−1Gt​ for t<Tt < Tt<T.
  2. Exercise 12.1: Gtλ=Rt+1+γ[(1−λ)v^(St+1,w)+λGt+1λ]G^\lambda_t = R_{t+1} + \gamma[(1-\lambda)\hat v(S_{t+1}, w) + \lambda G^\lambda_{t+1}]Gtλ​=Rt+1​+γ[(1−λ)v^(St+1​,w)+λGt+1λ​].
  3. Exercise 12.3: Gtλ−v^(St,w)=∑k=tT−1(γλ)k−tδkG^\lambda_t - \hat v(S_t, w) = \sum_{k=t}^{T-1}(\gamma\lambda)^{k-t}\delta_kGtλ​−v^(St​,w)=∑k=tT−1​(γλ)k−tδk​.
  4. Exercise 12.4: ∑t<Tαδtzt=∑t<Tα[Gtλ−v^(St,w)]∇v^(St,w)\sum_{t<T}\alpha\delta_t z_t = \sum_{t<T}\alpha[G^\lambda_t - \hat v(S_t, w)]\nabla\hat v(S_t, w)∑t<T​αδt​zt​=∑t<T​α[Gtλ​−v^(St​,w)]∇v^(St​,w).

Significance

The goal says that an O(d)O(d)O(d)-per-step algorithm reproduces the Monte Carlo/LMS result exactly. That algorithm never stores the feature vectors or the TTT intermediate weight vectors. The book draws the conclusion that eligibility traces "are not specific to TD learning at all" (p. 303). The dutch trace in the case γλ=1\gamma\lambda = 1γλ=1 is the same object that true online TD(λ\lambdaλ) (12.11) uses for general γλ\gamma\lambdaγλ. The fading-matrix products and their incremental forms are therefore the vocabulary of any later formalization of true online TD(λ\lambdaλ) and of the online λ\lambdaλ-return algorithm.

Exercises 12.3–12.4 are the fixed-weight equivalence of TD(λ\lambdaλ) and the off-line λ\lambdaλ-return algorithm. They are the standard justification for calling TD(λ\lambdaλ) an approximation of the λ\lambdaλ-return algorithm. Exercise 12.1 and (12.3) are the identities the rest of the chapter uses to manipulate λ\lambdaλ-returns.

On status: all of these results are known and elementary on paper, and the book prints the derivation of (12.14). To our knowledge none of them has a machine-checked proof, and the platform has no statement about λ\lambdaλ-returns, eligibility traces or TD(λ\lambdaλ). What this mission adds is formal statements with every convention fixed, including one correction to the printed text. It also adds reusable definitions of nnn-step returns, λ\lambdaλ-returns and traces.

Difficulty

The algebra is elementary. The difficulty lies in the conventions, and a careless reading of the page produces a false statement. The book initializes a0=w0a_0 = w_0a0​=w0​, and taken literally that makes the goal false. The λ\lambdaλ-return is an infinite series, whose tail collapses only because every nnn-step return that reaches past termination equals the full return. That convention has to be built into the definition of Gt:t+nG_{t:t+n}Gt:t+n​, together with the terminal value v^(terminal,⋅)=0\hat v(\text{terminal}, \cdot) = 0v^(terminal,⋅)=0. Exercises 12.3 and 12.4 are true only when the weights stay fixed. With the algorithms' changing weights wtw_twt​ the nnn-step returns (12.1) use wt+n−1w_{t+n-1}wt+n−1​, and neither identity holds. The boundary indices (t=T−1t = T-1t=T−1, the empty product at k=T−1k = T-1k=T−1, GTλ=0G^\lambda_T = 0GTλ​=0) must come out right.

Formalization scope

Namespace SuttonBartoRL.Traces. Vectors of §12.6 are Fin d → ℝ, matrices Matrix (Fin d) (Fin d) ℝ, and xx⊤x x^\topxx⊤ is Matrix.vecMulVec x x. The ordered product fadeProd α x j t is Ft⋯FjF_t \cdots F_jFt​⋯Fj​, the identity when t<jt < jt<j. Feature sequences are indexed by N\mathbb NN; only x0,…,xT−1x_0, \dots, x_{T-1}x0​,…,xT−1​ enter. T≥1T \ge 1T≥1 is a hypothesis wherever T−1T - 1T−1 appears. The identities of §12.6 are stated for every real α\alphaα and GGG, a harmless strengthening of the book's positive step size.

For §§12.1–12.2, weights are EuclideanSpace ℝ (Fin d), ∇\nabla∇ is Mathlib's gradient, and each v^(s,⋅)\hat v(s,\cdot)v^(s,⋅) is assumed differentiable in Exercise 12.4. An episode is a length TTT, states and rewards. The value at time t≥Tt \ge Tt≥T is 000. (12.2) is a tsum over n≥0n \ge 0n≥0 of λnGt:t+n+1\lambda^n G_{t:t+n+1}λnGt:t+n+1​, and λ∈[0,1)\lambda \in [0,1)λ∈[0,1), the range the book gives with (12.2). Milestone 5 concludes summability, so the junk value of a divergent tsum cannot make it trivial. Exercises 12.3–12.4 take one weight binder w, used in every return, TD error and gradient. That is the book's fixed-www assumption, stated in the binders.

Correction. The printed initialization a0=w0a_0 = w_0a0​=w0​ (p. 302) contradicts the printed definition at≐Ft⋯F0w0a_t \doteq F_t\cdots F_0 w_0at​≐Ft​⋯F0​w0​ and (12.14). The counterexample is d=1d = 1d=1, T=1T = 1T=1, x0=1x_0 = 1x0​=1, α=1/2\alpha = 1/2α=1/2, w0=1w_0 = 1w0​=1, G=0G = 0G=0: the forward view gives w1=1/2w_1 = 1/2w1​=1/2, while a0+αGz0=1a_0 + \alpha G z_0 = 1a0​+αGz0​=1. The mission states the corrected result with a0=F0w0a_0 = F_0 w_0a0​=F0​w0​, equivalently the same recursion started from a−1=w0a_{-1} = w_0a−1​=w0​. The printed text is kept verbatim in the milestone.

A trivializing formalization is ruled out: the goal is not the closed-form identity with aT−1a_{T-1}aT−1​ and zT−1z_{T-1}zT−1​ defined as the products and sums. Those vectors are defined by their GGG-free incremental recursions, and the closed forms are separate milestones.

Out of scope: the equivalence of true online TD(λ\lambdaλ) and the online λ\lambdaλ-return algorithm (cited, p. 300), the truncated-return identity (12.10), the error bound (12.8), and all convergence claims. Proofs of the milestones, and reuse of the definitions in later missions on true online TD(λ\lambdaλ), are welcome.

Selected references

  • R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018, ISBN 9780262039246, Chapter 12, pp. 287–320. http://incompleteideas.net/book/the-book-2nd.html
  • R. S. Sutton, "Learning to predict by the methods of temporal differences", Machine Learning 3 (1988), 9–44. https://doi.org/10.1007/BF00115009
  • H. van Seijen and R. S. Sutton, "True online TD(λ)", Proceedings of ICML 2014, PMLR 32, 692–700. https://proceedings.mlr.press/v32/seijen14.html
  • H. van Seijen, A. R. Mahmood, P. M. Pilarski, M. C. Machado and R. S. Sutton, "True online temporal-difference learning", Journal of Machine Learning Research 17 (2016), 1–40. https://jmlr.org/papers/v17/15-599.html
11 thms2 active usersReviewed
Machine LearningMarkov Chain·Captain: mikedeng1

Reinforcement Learning: An Introduction VII: The Error Reduction Property of n-step ReturnsTextbook

Motivation

Temporal-difference (TD) learning estimates the value of a policy by moving a current estimate toward a target built from observed rewards and from the estimate itself. Chapter 7 of Sutton and Barto's Reinforcement Learning: An Introduction (2nd ed., MIT Press, 2018) interpolates between the two extreme targets of the preceding chapters: the one-step TD target, which uses one reward and then bootstraps, and the Monte Carlo target, which uses every reward until the end of the episode. The intermediate target, the nnn-step return, uses nnn rewards and then bootstraps from the current estimate. The family underlies nnn-step TD, nnn-step Sarsa, the off-policy per-decision methods and the tree-backup algorithm, and it is the introduction to eligibility traces (Chapter 12).

The book justifies the whole family with one inequality, the error reduction property (7.3), p. 144: the expected nnn-step return is closer to the true value than the estimate it bootstraps from, by a factor γn\gamma^nγn in the worst state. It is the reason given for calling nnn-step TD methods "sound". The same chapter states, mostly as exercises without solutions, a series of exact identities that rewrite each kind of nnn-step return as a sum of one-step TD errors.

Setting

A finite Markov decision process has finite state and action sets S\mathcal SS, A\mathcal AA, a finite reward set R⊂R\mathcal R \subset \mathbb RR⊂R, and dynamics p(s′,r∣s,a)p(s', r \mid s, a)p(s′,r∣s,a), the probability of next state s′s's′ and reward rrr after action aaa in state sss (Eqs. (3.2)–(3.3)). A policy π(a∣s)\pi(a \mid s)π(a∣s) is a probability distribution over actions for each state. Following π\piπ from St=sS_t = sSt​=s produces a random trajectory At,Rt+1,St+1,At+1,Rt+2,…A_t, R_{t+1}, S_{t+1}, A_{t+1}, R_{t+2}, \dotsAt​,Rt+1​,St+1​,At+1​,Rt+2​,… For a discount factor 0≤γ<10 \le \gamma < 10≤γ<1 the state-value function is the expected discounted return (3.12),

vπ(s)=Eπ[∑k=0∞γkRt+k+1 ∣ St=s].v_\pi(s) = \mathbb E_\pi\Big[\sum_{k=0}^{\infty} \gamma^k R_{t+k+1} \,\Big|\, S_t = s\Big].vπ​(s)=Eπ​[k=0∑∞​γkRt+k+1​​St​=s].

Given any function V:S→RV : \mathcal S \to \mathbb RV:S→R (an estimate of vπv_\pivπ​), the nnn-step return (7.1) is

Gt:t+n=Rt+1+γRt+2+⋯+γn−1Rt+n+γnV(St+n).G_{t:t+n} = R_{t+1} + \gamma R_{t+2} + \cdots + \gamma^{n-1} R_{t+n} + \gamma^n V(S_{t+n}).Gt:t+n​=Rt+1​+γRt+2​+⋯+γn−1Rt+n​+γnV(St+n​).

In an episode that terminates at time TTT it is replaced by the complete return GtG_tGt​ when t+n≥Tt + n \ge Tt+n≥T. The TD error (6.5) is δk=Rk+1+γV(Sk+1)−V(Sk)\delta_k = R_{k+1} + \gamma V(S_{k+1}) - V(S_k)δk​=Rk+1​+γV(Sk+1​)−V(Sk​). Off-policy variants use a behavior policy bbb that generates the data, the per-decision ratio ρt=π(At∣St)/b(At∣St)\rho_t = \pi(A_t \mid S_t)/b(A_t \mid S_t)ρt​=π(At​∣St​)/b(At​∣St​), and the return with control variate (7.13), Gt:h=ρt(Rt+1+γGt+1:h)+(1−ρt)V(St)G_{t:h} = \rho_t(R_{t+1} + \gamma G_{t+1:h}) + (1-\rho_t) V(S_t)Gt:h​=ρt​(Rt+1​+γGt+1:h​)+(1−ρt​)V(St​), Gh:h=V(Sh)G_{h:h} = V(S_h)Gh:h​=V(Sh​). The tree-backup return (7.15)–(7.16) uses action values QQQ and the expected approximate value Vˉ(s)=∑aπ(a∣s)Q(s,a)\bar V(s) = \sum_a \pi(a \mid s) Q(s, a)Vˉ(s)=∑a​π(a∣s)Q(s,a) (7.8).

Formalization targets

Goal: the error reduction property (7.3)

For a finite MDP, a policy π\piπ, 0≤γ<10 \le \gamma < 10≤γ<1, any V:S→RV : \mathcal S \to \mathbb RV:S→R and every n≥1n \ge 1n≥1,

max⁡s∣Eπ[Gt:t+n∣St=s]−vπ(s)∣≤γnmax⁡s∣V(s)−vπ(s)∣.\max_s \big|\mathbb E_\pi[G_{t:t+n} \mid S_t = s] - v_\pi(s)\big| \le \gamma^n \max_s \big|V(s) - v_\pi(s)\big|.smax​​Eπ​[Gt:t+n​∣St​=s]−vπ​(s)​≤γnsmax​​V(s)−vπ​(s)​.

Milestones, in the book's order

  1. Exercise 7.1, p. 143: with VVV fixed and V(ST)=0V(S_T) = 0V(ST​)=0, Gt:t+n−V(St)=∑k=tmin⁡(t+n,T)−1γk−tδkG_{t:t+n} - V(S_t) = \sum_{k=t}^{\min(t+n,T)-1} \gamma^{k-t}\delta_kGt:t+n​−V(St​)=∑k=tmin(t+n,T)−1​γk−tδk​.
  2. Exercise 7.4, Eq. (7.6), p. 148: the nnn-step Sarsa return equals Qt−1(St,At)+∑k=tmin⁡(t+n,T)−1γk−t[Rk+1+γQk(Sk+1,Ak+1)−Qk−1(Sk,Ak)]Q_{t-1}(S_t, A_t) + \sum_{k=t}^{\min(t+n,T)-1} \gamma^{k-t}[R_{k+1} + \gamma Q_k(S_{k+1}, A_{k+1}) - Q_{k-1}(S_k, A_k)]Qt−1​(St​,At​)+∑k=tmin(t+n,T)−1​γk−t[Rk+1​+γQk​(Sk+1​,Ak+1​)−Qk−1​(Sk​,Ak​)], with estimates changing from step to step.
  3. Eq. (7.12), p. 150: Gt:h=Rt+1+γGt+1:hG_{t:h} = R_{t+1} + \gamma G_{t+1:h}Gt:h​=Rt+1​+γGt+1:h​ for t<h<Tt < h < Tt<h<T, Gh:h=V(Sh)G_{h:h} = V(S_h)Gh:h​=V(Sh​).
  4. Exercise 7.6, p. 151, for (7.13): under coverage, Eb\mathbb E_bEb​ of the control-variate return equals Eb\mathbb E_bEb​ of the same return without the control variate, and both equal Eπ[Gt:t+n∣St=s]\mathbb E_\pi[G_{t:t+n} \mid S_t = s]Eπ​[Gt:t+n​∣St​=s].
  5. Exercise 7.8, p. 151: Gt:h−V(St)=∑k=th−1γk−t(∏i=tkρi)δkG_{t:h} - V(S_t) = \sum_{k=t}^{h-1} \gamma^{k-t} \big(\prod_{i=t}^{k}\rho_i\big) \delta_kGt:h​−V(St​)=∑k=th−1​γk−t(∏i=tk​ρi​)δk​ for the return (7.13).
  6. Exercise 7.11, p. 153: the tree-backup return equals Q(St,At)+∑k=tmin⁡(t+n−1,T−1)δk∏i=t+1kγπ(Ai∣Si)Q(S_t, A_t) + \sum_{k=t}^{\min(t+n-1,T-1)} \delta_k \prod_{i=t+1}^{k} \gamma\pi(A_i \mid S_i)Q(St​,At​)+∑k=tmin(t+n−1,T−1)​δk​∏i=t+1k​γπ(Ai​∣Si​) with the expectation-based TD error δk=Rk+1+γVˉ(Sk+1)−Q(Sk,Ak)\delta_k = R_{k+1} + \gamma\bar V(S_{k+1}) - Q(S_k, A_k)δk​=Rk+1​+γVˉ(Sk+1​)−Q(Sk​,Ak​).

Significance

The result. The error reduction property makes the expected nnn-step target a γn\gamma^nγn-contraction toward vπv_\pivπ​ in the sup norm, uniformly over the estimate it starts from. It is the one-line reason the book offers for the soundness of every nnn-step TD method, and the same contraction is what the λ\lambdaλ-return of Chapter 12 averages over nnn. The TD-error identities are the algebra behind implementations that accumulate TD errors instead of storing returns, and behind the forward/backward-view equivalences of Chapter 12. Exercise 7.6 is the unbiasedness of the control-variate return, which is what allows (7.13) to replace plain importance weighting without changing the expected update.

Formalizing it. All results are elementary and well known, but the book gives no proofs: (7.3) is asserted, and the identities are exercises without published solutions. None is formalized on Prove2Me. The mission produces machine-checked versions with every hypothesis explicit (discounting, the fixed estimate, terminal values, coverage), and a trajectory-level expectation for finite MDPs that other chapters of the series can reuse.

Difficulty

The obvious proof of (7.3) is a matrix computation: Eπ[Gt:t+n∣St=s]−vπ(s)=γn(Pπn(V−vπ))(s)\mathbb E_\pi[G_{t:t+n} \mid S_t = s] - v_\pi(s) = \gamma^n (P_\pi^n (V - v_\pi))(s)Eπ​[Gt:t+n​∣St​=s]−vπ​(s)=γn(Pπn​(V−vπ​))(s), and a stochastic matrix does not increase the sup norm. The difficulty lies in the step before it. The left side is an expectation over trajectories, and vπv_\pivπ​ is an infinite discounted series; neither is a matrix power by definition. Connecting them requires a Chapman–Kolmogorov identity for the finite-trajectory distribution induced by π\piπ and p(s′,r∣s,a)p(s', r \mid s, a)p(s′,r∣s,a), the splitting of vπv_\pivπ​ at time nnn, and summability of the discounted series. Defining the expected nnn-step return as the matrix expression would reduce the goal to the last line and remove its content; that shortcut is ruled out below.

The TD-error identities are telescoping sums, but each has its own boundary: termination inside the nnn steps, the convention that terminal states have value zero, the index Q−1Q_{-1}Q−1​ at t=0t = 0t=0 in (7.6), the special case GT−1:t+n=RTG_{T-1:t+n} = R_TGT−1:t+n​=RT​ of the tree backup, and ratios with vanishing denominators in (7.13).

Formalization scope

  • Model. The finite MDP, policies and vπv_\pivπ​ follow the series conventions: dynamics p(s′,r∣s,a)p(s', r \mid s, a)p(s′,r∣s,a) over a finite reward set, one action set for all states, vπ(s)=∑kγk(Pπkrπ)(s)v_\pi(s) = \sum_k \gamma^k (P_\pi^k r_\pi)(s)vπ​(s)=∑k​γk(Pπk​rπ​)(s) computed from expected rewards and never defined as a Bellman solution.
  • Expectations are over trajectories. Eπ[ ⋅∣St=s]\mathbb E_\pi[\,\cdot \mid S_t = s]Eπ​[⋅∣St​=s] is the finite sum over nnn-step segments (At+k,St+k+1,Rt+k+1)k<n(A_{t+k}, S_{t+k+1}, R_{t+k+1})_{k<n}(At+k​,St+k+1​,Rt+k+1​)k<n​ weighted by ∏kπ(At+k∣St+k) p(St+k+1,Rt+k+1∣St+k,At+k)\prod_k \pi(A_{t+k}\mid S_{t+k})\,p(S_{t+k+1}, R_{t+k+1}\mid S_{t+k}, A_{t+k})∏k​π(At+k​∣St+k​)p(St+k+1​,Rt+k+1​∣St+k​,At+k​). The expected nnn-step return is not defined as ∑k<nγkPπkrπ+γnPπnV\sum_{k<n}\gamma^k P_\pi^k r_\pi + \gamma^n P_\pi^n V∑k<n​γkPπk​rπ​+γnPπn​V, which would make the goal a two-line matrix inequality.
  • The estimate is fixed. In the algorithm, Vt+n−1V_{t+n-1}Vt+n−1​ is the current random estimate. Every statement takes a fixed function VVV (or QQQ), which is the book's own reading ("if the value estimates don't change"). The only exception is Exercise 7.4, whose estimates QkQ_kQk​ are indexed by time k∈Zk \in \mathbb Zk∈Z exactly as in (7.6).
  • Discounting. The goal assumes 0≤γ<10 \le \gamma < 10≤γ<1 and takes the maximum over all states. Episodic tasks enter through absorbing zero-reward terminal states. The undiscounted episodic case γ=1\gamma = 1γ=1 is not stated.
  • Episodes. Sample-path identities use sequences Sk,Ak,RkS_k, A_k, R_kSk​,Ak​,Rk​ and a termination time TTT. The book's convention that terminal states have value 000 is a hypothesis (V(ST)=0V(S_T) = 0V(ST​)=0, Q(ST,⋅)=0Q(S_T, \cdot) = 0Q(ST​,⋅)=0).
  • Exercise 7.6 is stated for the state-value return (7.13) of p. 150, although the exercise follows the action-value return (7.14). Its conclusion includes, besides the literal "does not change the expected value", equality with the on-policy expected return, the property the book states on p. 150. Coverage (π(a∣s)>0⇒b(a∣s)>0\pi(a\mid s) > 0 \Rightarrow b(a\mid s) > 0π(a∣s)>0⇒b(a∣s)>0) is assumed.
  • Not stated. The convergence of nnn-step TD methods "under appropriate technical conditions" (p. 144), and the programming exercises.

Needed infrastructure: finite sums over function types, Chapman–Kolmogorov for the segment distribution, summability of ∑kγkPπkrπ\sum_k \gamma^k P_\pi^k r_\pi∑k​γkPπk​rπ​. The trajectory layer is reusable for the importance-sampling and eligibility-trace chapters. Alternative proofs of the goal, and proofs of the undiscounted episodic version as a separate theorem, are welcome.

Selected references

  • R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018, ISBN 9780262039246, Chapter 7, pp. 141–158. http://incompleteideas.net/book/the-book-2nd.html
  • C. J. C. H. Watkins, Learning from Delayed Rewards, PhD thesis, University of Cambridge, 1989 (the nnn-step return and its error reduction property, as credited on p. 158 of the book). https://www.cs.rhul.ac.uk/~chrisw/new_thesis.pdf
  • D. Precup, R. S. Sutton and S. Singh, Eligibility traces for off-policy policy evaluation, Proceedings of the 17th International Conference on Machine Learning (ICML), 2000, pp. 759–766 (the tree-backup algorithm, as credited on p. 158 of the book; no DOI).
10 thms2 active usersReviewed
Dynamic ProgrammingMachine LearningMarkov Chain·Captain: mikedeng1

Reinforcement Learning: An Introduction II: The Bellman Optimality Equation and the Existence of an Optimal PolicyTextbook

Motivation

Sutton and Barto's Reinforcement Learning: An Introduction (2nd ed., MIT Press, 2018) is the standard introductory text of the field. Its Chapter 3 sets up the model that the rest of Part I works in: the finite Markov decision process (MDP), the value functions of a policy, and the Bellman equations that relate the value of a state to the values of its successors. Section 3.6 then states the fact every planning and control method of the book relies on: in a finite MDP there is an optimal policy, its value is the unique solution of a system of nonlinear equations, and a policy that acts greedily with respect to that solution is optimal. Dynamic programming (Chapter 4), Monte Carlo control (Chapter 5), Sarsa and Q-learning (Chapter 6) are all methods for solving the Bellman optimality equation; their correctness statements presuppose that it has exactly one solution and that it identifies optimal behaviour.

The book is deliberately informal ("we chose not to produce a rigorous formal treatment", p. xiii): §3.6 asserts these facts without proof. The results themselves are classical, going back to Bellman (1957), Howard (1960) and Blackwell (1965); a textbook proof for discounted finite MDPs is in Puterman, Markov Decision Processes (Wiley, 1994), Chapter 6.

Setting

A finite MDP has a finite set of states S\mathcal SS, a finite nonempty set of actions A\mathcal AA, a finite set of rewards R⊂R\mathcal R \subset \mathbb RR⊂R, and dynamics

p(s′,r∣s,a)=Pr⁡{St=s′,Rt=r∣St−1=s,At−1=a},∑s′∈S∑r∈Rp(s′,r∣s,a)=1,p(s', r \mid s, a) = \Pr\{S_t = s', R_t = r \mid S_{t-1} = s, A_{t-1} = a\}, \qquad \sum_{s' \in \mathcal S}\sum_{r \in \mathcal R} p(s', r \mid s, a) = 1,p(s′,r∣s,a)=Pr{St​=s′,Rt​=r∣St−1​=s,At−1​=a},s′∈S∑​r∈R∑​p(s′,r∣s,a)=1,

Eqs. (3.2)–(3.3). From ppp one derives p(s′∣s,a)=∑rp(s′,r∣s,a)p(s' \mid s, a) = \sum_r p(s', r \mid s, a)p(s′∣s,a)=∑r​p(s′,r∣s,a) and r(s,a)=∑rr∑s′p(s′,r∣s,a)r(s, a) = \sum_r r \sum_{s'} p(s', r \mid s, a)r(s,a)=∑r​r∑s′​p(s′,r∣s,a), Eqs. (3.4)–(3.5).

A policy π\piπ gives a probability π(a∣s)\pi(a \mid s)π(a∣s) of each action in each state. Fix a discount rate 0≤γ<10 \le \gamma < 10≤γ<1. The return of a reward sequence is Gt=∑k≥0γkRt+k+1G_t = \sum_{k \ge 0} \gamma^k R_{t+k+1}Gt​=∑k≥0​γkRt+k+1​ (3.8). The state-value function and action-value function of π\piπ are the expected returns

vπ(s)=Eπ[Gt∣St=s],qπ(s,a)=Eπ[Gt∣St=s,At=a](3.12)–(3.13).v_\pi(s) = \mathbb E_\pi[G_t \mid S_t = s], \qquad q_\pi(s, a) = \mathbb E_\pi[G_t \mid S_t = s, A_t = a] \qquad (3.12)\text{–}(3.13).vπ​(s)=Eπ​[Gt​∣St​=s],qπ​(s,a)=Eπ​[Gt​∣St​=s,At​=a](3.12)–(3.13).

In the Lean development these are stateValue M γ π s and actionValue M γ π s a, computed as ∑kγk(Pπkrπ)(s)\sum_k \gamma^k (P_\pi^k r_\pi)(s)∑k​γk(Pπk​rπ​)(s) from the transition matrix Pπ(s,s′)=∑aπ(a∣s) p(s′∣s,a)P_\pi(s, s') = \sum_a \pi(a \mid s)\,p(s' \mid s, a)Pπ​(s,s′)=∑a​π(a∣s)p(s′∣s,a) and the expected reward rπ(s)=∑aπ(a∣s) r(s,a)r_\pi(s) = \sum_a \pi(a \mid s)\,r(s, a)rπ​(s)=∑a​π(a∣s)r(s,a) of the Markov chain the policy induces. A policy π\piπ is optimal (IsOptimalPolicy) if vπ(s)≥vπ′(s)v_\pi(s) \ge v_{\pi'}(s)vπ​(s)≥vπ′​(s) for every policy π′\pi'π′ and every state sss. The optimal value functions are

v∗(s)=max⁡πvπ(s)(3.15),q∗(s,a)=max⁡πqπ(s,a)(3.16),v_*(s) = \max_\pi v_\pi(s) \quad (3.15), \qquad q_*(s, a) = \max_\pi q_\pi(s, a) \quad (3.16),v∗​(s)=πmax​vπ​(s)(3.15),q∗​(s,a)=πmax​qπ​(s,a)(3.16),

optimalValue and optimalActionValue, with the maximum over all stochastic policies.

Formalization targets

Goal: the Bellman optimality equation and optimal policies (§3.6, pp. 62–64)

For every finite MDP and 0≤γ<10 \le \gamma < 10≤γ<1:

  1. the maximum in (3.15) is attained at every state;
  2. an optimal policy exists;
  3. v∗v_*v∗​ satisfies the Bellman optimality equation
v∗(s)=max⁡a∑s′,rp(s′,r∣s,a)[r+γv∗(s′)]for all s;(3.19)v_*(s) = \max_{a} \sum_{s', r} p(s', r \mid s, a)\big[r + \gamma v_*(s')\big] \quad \text{for all } s; \qquad (3.19)v∗​(s)=amax​s′,r∑​p(s′,r∣s,a)[r+γv∗​(s′)]for all s;(3.19)
  1. v∗v_*v∗​ is the only function on S\mathcal SS satisfying (3.19);
  2. every policy that assigns positive probability only to actions attaining the maximum in (3.19) is optimal.

Milestones

  • (3.9) Gt=Rt+1+γGt+1G_t = R_{t+1} + \gamma G_{t+1}Gt​=Rt+1​+γGt+1​ for bounded rewards, with the series convergent.
  • (3.14) the Bellman equation vπ(s)=∑aπ(a∣s)∑s′,rp(s′,r∣s,a)[r+γvπ(s′)]v_\pi(s) = \sum_a \pi(a \mid s) \sum_{s', r} p(s', r \mid s, a)[r + \gamma v_\pi(s')]vπ​(s)=∑a​π(a∣s)∑s′,r​p(s′,r∣s,a)[r+γvπ​(s′)], and (p. 60) its uniqueness: vπv_\pivπ​ is its only solution.
  • Exercise 3.15 adding a constant ccc to all rewards adds vc=c/(1−γ)v_c = c/(1-\gamma)vc​=c/(1−γ) to every value.
  • Exercises 3.18 and 3.19 vπ(s)=∑aπ(a∣s) qπ(s,a)v_\pi(s) = \sum_a \pi(a \mid s)\,q_\pi(s, a)vπ​(s)=∑a​π(a∣s)qπ​(s,a) and qπ(s,a)=∑s′,rp(s′,r∣s,a)[r+γvπ(s′)]q_\pi(s, a) = \sum_{s', r} p(s', r \mid s, a)[r + \gamma v_\pi(s')]qπ​(s,a)=∑s′,r​p(s′,r∣s,a)[r+γvπ​(s′)].
  • (3.16)–(3.17) the maximum defining q∗q_*q∗​ is attained and q∗(s,a)=∑s′,rp(s′,r∣s,a)[r+γv∗(s′)]q_*(s, a) = \sum_{s', r} p(s', r \mid s, a)[r + \gamma v_*(s')]q∗​(s,a)=∑s′,r​p(s′,r∣s,a)[r+γv∗​(s′)].
  • (3.20) the Bellman optimality equation for action values, q∗(s,a)=∑s′,rp(s′,r∣s,a)[r+γmax⁡a′q∗(s′,a′)]q_*(s, a) = \sum_{s', r} p(s', r \mid s, a)[r + \gamma \max_{a'} q_*(s', a')]q∗​(s,a)=∑s′,r​p(s′,r∣s,a)[r+γmaxa′​q∗​(s′,a′)].

Significance

The goal is what turns "find a good policy" into "solve a system of equations". Parts 3 and 4 identify v∗v_*v∗​ with the unique solution of (3.19), so any procedure that finds a solution of (3.19) has found v∗v_*v∗​; part 5 converts v∗v_*v∗​ into an optimal policy by a one-step search. Parts 1 and 2 say that the book's definition (3.15) makes sense: a single policy is simultaneously best at every state, so "the optimal value function" is well defined and shared by all optimal policies. Chapter 4's policy iteration and value iteration, and the fixed points of Q-learning, are statements about this equation. Exercises 3.18 and 3.19 are used, by number, in the proof of the policy gradient theorem (p. 325).

The mathematics is classical and proved in many texts; what is missing is a machine-checked version in the book's own model. Platform relatives exist in different models: FoundationsML.ReinforcementLearning.bellman_equations_unique_solution (uniqueness for a fixed policy, with an expected-reward kernel instead of p(s′,r∣s,a)p(s', r \mid s, a)p(s′,r∣s,a)), BertsekasDP.discounted_main_theorem (cost minimization over deterministic stationary policies), BanditAlgorithm.mdp_discounted_bellman_solution (existence of a solution with a greedy deterministic policy, rewards in [0,1][0,1][0,1]) and FoundationsRL.RLBasics.bellman_optimality (finite horizon). None of them states the book's result: the four-argument dynamics, stochastic policies, the maximum over all of them, uniqueness of the solution of (3.19), and optimality of every policy supported on greedy actions. This mission produces that statement and, with it, a vocabulary of finite-MDP definitions that the later missions of the series reuse.

Difficulty

The book's derivation of (3.19) (p. 63) starts from v∗(s)=max⁡aqπ∗(s,a)v_*(s) = \max_a q_{\pi_*}(s, a)v∗​(s)=maxa​qπ∗​​(s,a) with a policy π∗\pi_*π∗​ that is optimal at every state at once. The existence of such a policy is the substance of the goal, and it does not follow from the definition: (3.15) takes a separate maximum at each state, and a priori the maximizing policy could depend on the state. Uniqueness for (3.19) is likewise not a consequence of linear algebra, as it is for (3.14): the equation is nonlinear because of the maximum. The fixed point must be related to the value of an actual policy, and every policy's value must be bounded above by it.

Formalization scope

  • Model. S and A are finite types with A nonempty (without an action, max⁡a\max_amaxa​ is undefined). One action set serves every state, as the book's footnote 3 (p. 48) allows. Rewards form a finite set M.R : Finset ℝ and the dynamics are the four-argument M.p s a s' r with the normalization (3.3).
  • Discounting. All statements assume 0≤γ<10 \le \gamma < 10≤γ<1 (the continuing discounted case of §3.3). The episodic case with γ=1\gamma = 1γ=1 is not covered: the book's uniqueness claims then need every episode to terminate under every policy, which the chapter never states, and without it (3.19) can have many solutions (a state that loops to itself with reward 000 satisfies v(s)=v(s)v(s) = v(s)v(s)=v(s) for any value).
  • Value functions from returns. vπv_\pivπ​ and qπq_\piqπ​ are expected discounted returns, computed from the Markov chain the policy induces. They are not defined as solutions of the Bellman equations, and v∗v_*v∗​ is not defined as a solution of (3.19): either would make the goal true by definition. The Bellman equations are theorems.
  • Maxima. v∗v_*v∗​ and q∗q_*q∗​ are real suprema over the type of stochastic policies (Lean gives a supremum that does not exist the value 000); the goal and milestone (3.17) assert that these suprema are attained, so they are the book's maxima.
  • Conditional expectations. (3.17), (3.18) and (3.20) are stated in their finite-sum form over (s′,r)(s', r)(s′,r).
  • Exercises. Exercises 3.15, 3.18 and 3.19 have no printed solutions; the statements give the formalization's answers (vc=c/(1−γ)v_c = c/(1-\gamma)vc​=c/(1−γ) and the two displayed identities).
  • Reusable infrastructure. The definitions (MDP, Policy, trans, expReward, policyTrans, policyReward, stateValue, actionValue, optimalValue, optimalActionValue, IsOptimalPolicy) follow the conventions shared by the whole series and are meant to be merged with the finite-MDP layers of the later chapters. Lemmas on summability of the value series, the Bellman operator as a γ\gammaγ-contraction in the sup norm, and the Markov-chain identities for PπkP_\pi^kPπk​ are welcome as separate contributions.

Selected references

  • R. S. Sutton, A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018, ISBN 9780262039246, Chapter 3. http://incompleteideas.net/book/the-book-2nd.html
  • R. Bellman, Dynamic Programming, Princeton University Press, 1957. https://doi.org/10.2307/j.ctv1nxcw0f
  • R. A. Howard, Dynamic Programming and Markov Processes, MIT Press, 1960.
  • D. Blackwell, "Discounted dynamic programming", Annals of Mathematical Statistics 36(1), 1965, 226–235. https://doi.org/10.1214/aoms/1177700285
  • M. L. Puterman, Markov Decision Processes: Discrete Stochastic Dynamic Programming, Wiley, 1994. https://doi.org/10.1002/9780470316887
11 thms2 active usersReviewed
Bandit AlgorithmsMachine LearningOptimization·Captain: mikedeng1

Reinforcement Learning: An Introduction I: The Gradient Bandit Algorithm Is Stochastic Gradient AscentTextbook

Motivation

The multi-armed bandit is the simplest setting in which a learner must trade off exploiting what it knows against exploring what it does not: one situation, kkk actions, and a reward drawn from an unknown distribution each time an action is taken. Chapter 2 of Sutton and Barto's Reinforcement Learning: An Introduction (2nd ed., MIT Press, 2018) uses it to introduce, in the smallest possible setting, ideas that run through the rest of the book: incremental estimation with a step size, the bias introduced by the initial estimate, soft-max policies over learned preferences, and learning by following the gradient of expected reward.

The chapter ends with the gradient bandit algorithm (§2.8), which learns a numerical preference for each action instead of a value estimate. A shaded box on pp. 38–40 shows that its expected update is exactly a gradient-ascent step on the expected reward, so the algorithm is an instance of stochastic gradient ascent. The same argument, a score-function (likelihood-ratio) identity with a baseline, reappears in Chapter 13 as the REINFORCE algorithm and the policy gradient theorem. The bandit case is where the book first carries it out in full.

Setting

Actions are 1,…,k1, \dots, k1,…,k. Each action xxx has a reward distribution νx\nu_xνx​ on R\mathbb RR with finite mean q∗(x)q_*(x)q∗​(x), the true action value. At each step the learner holds a vector of action preferences H=(H(1),…,H(k))∈RkH = (H(1), \dots, H(k)) \in \mathbb R^kH=(H(1),…,H(k))∈Rk and selects action AAA with the soft-max probability

π(a)=eH(a)∑b=1keH(b)(2.11).\pi(a) = \frac{e^{H(a)}}{\sum_{b=1}^k e^{H(b)}} \qquad (2.11).π(a)=∑b=1k​eH(b)eH(a)​(2.11).

Given A=xA = xA=x, a reward R∼νxR \sim \nu_xR∼νx​ is received. The expected reward is E[R]=∑xπ(x) q∗(x)\mathbb E[R] = \sum_x \pi(x)\, q_*(x)E[R]=∑x​π(x)q∗​(x), a smooth function of HHH. With a step size α>0\alpha > 0α>0 and a baseline B∈RB \in \mathbb RB∈R, the gradient bandit update (2.12) is

H′(A)=H(A)+α(R−B)(1−π(A)),H′(a)=H(a)−α(R−B) π(a)  (a≠A).H'(A) = H(A) + \alpha (R - B)(1 - \pi(A)), \qquad H'(a) = H(a) - \alpha (R - B)\,\pi(a) \ \ (a \ne A).H′(A)=H(A)+α(R−B)(1−π(A)),H′(a)=H(a)−α(R−B)π(a)  (a=A).

The chapter's estimation sections use a single action's rewards R1,R2,…R_1, R_2, \dotsR1​,R2​,…. The sample average after n−1n-1n−1 selections is Qn=(R1+⋯+Rn−1)/(n−1)Q_n = (R_1 + \cdots + R_{n-1})/(n-1)Qn​=(R1​+⋯+Rn−1​)/(n−1), with an arbitrary initial value Q1Q_1Q1​. A constant step size α∈(0,1]\alpha \in (0,1]α∈(0,1] updates Qn+1=Qn+α[Rn−Qn]Q_{n+1} = Q_n + \alpha [R_n - Q_n]Qn+1​=Qn​+α[Rn​−Qn​] (2.5). The trace of one oˉ0=0\bar o_0 = 0oˉ0​=0, oˉn=oˉn−1+α(1−oˉn−1)\bar o_n = \bar o_{n-1} + \alpha (1 - \bar o_{n-1})oˉn​=oˉn−1​+α(1−oˉn−1​) defines the step size βn=α/oˉn\beta_n = \alpha / \bar o_nβn​=α/oˉn​ (2.8)–(2.9).

Formalization targets

Goal: the expected update is the gradient step

For every action aaa, with A∼πA \sim \piA∼π and R∣A=x∼νxR \mid A = x \sim \nu_xR∣A=x∼νx​,

E[H′(a)]=H(a)+α ∂ E[R]∂H(a),\mathbb E\bigl[H'(a)\bigr] = H(a) + \alpha\, \frac{\partial\, \mathbb E[R]}{\partial H(a)} ,E[H′(a)]=H(a)+α∂H(a)∂E[R]​,

that is, the update (2.12) equals the exact gradient-ascent step (2.13) in expected value, for every baseline BBB that does not depend on the selected action.

Milestones

  1. (2.3): Qn+1=Qn+1n[Rn−Qn]Q_{n+1} = Q_n + \tfrac1n [R_n - Q_n]Qn+1​=Qn​+n1​[Rn​−Qn​] for n≥1n \ge 1n≥1, including Q2=R1Q_2 = R_1Q2​=R1​ for arbitrary Q1Q_1Q1​.
  2. (2.6): Qn+1=(1−α)nQ1+∑i=1nα(1−α)n−iRiQ_{n+1} = (1-\alpha)^n Q_1 + \sum_{i=1}^n \alpha(1-\alpha)^{n-i} R_iQn+1​=(1−α)nQ1​+∑i=1n​α(1−α)n−iRi​, with weights summing to one.
  3. Exercise 2.7: with βn=α/oˉn\beta_n = \alpha/\bar o_nβn​=α/oˉn​, Qn+1=∑i=1nα(1−α)n−ioˉnRiQ_{n+1} = \sum_{i=1}^n \frac{\alpha(1-\alpha)^{n-i}}{\bar o_n} R_iQn+1​=∑i=1n​oˉn​α(1−α)n−i​Ri​ for n≥1n \ge 1n≥1, weights summing to one, and no dependence on Q1Q_1Q1​.
  4. Shift invariance (p. 37): adding a constant ccc to every preference leaves π\piπ unchanged.
  5. Exercise 2.9: for k=2k = 2k=2, π(1)=σ(H(1)−H(2))\pi(1) = \sigma(H(1) - H(2))π(1)=σ(H(1)−H(2)) with σ(x)=1/(1+e−x)\sigma(x) = 1/(1+e^{-x})σ(x)=1/(1+e−x).
  6. Soft-max derivative (p. 40): ∂π(x)/∂H(a)=π(x)(1a=x−π(a))\partial \pi(x)/\partial H(a) = \pi(x)(\mathbb 1_{a=x} - \pi(a))∂π(x)/∂H(a)=π(x)(1a=x​−π(a)).
  7. Zero-sum gradient (p. 39): ∑x∂π(x)/∂H(a)=0\sum_x \partial \pi(x)/\partial H(a) = 0∑x​∂π(x)/∂H(a)=0.
  8. Performance gradient as an expectation (p. 39): ∂E[R]/∂H(a)=E[(R−B)(1a=A−π(a))]\partial \mathbb E[R]/\partial H(a) = \mathbb E[(R - B)(\mathbb 1_{a=A} - \pi(a))]∂E[R]/∂H(a)=E[(R−B)(1a=A​−π(a))].

Significance

The result. The identity makes a model-free algorithm, which uses only the sampled action and reward, an unbiased estimator of the gradient of a quantity that depends on the unknown q∗q_*q∗​. It therefore places the gradient bandit algorithm within stochastic approximation, where convergence theory for stochastic gradient methods applies. It also explains the role of the baseline: any baseline independent of the action leaves the expected update unchanged, so the choice of baseline can only affect the variance of the update, as Figure 2.5 shows empirically. The estimation milestones make precise two claims the chapter uses repeatedly: sample averages can be maintained incrementally, and constant step sizes produce an exponentially recency-weighted average biased by Q1Q_1Q1​. Exercise 2.7 removes that bias.

Formalizing it. All of these results are elementary and proved (or left as routine exercises) in the book. None of them is formalized on Prove2Me or, as far as is known, in Mathlib. What this mission adds is a machine-checked version of the book's argument with the reward model and baseline condition stated precisely, and a reusable soft-max layer (definition, partial derivatives, shift invariance) for later missions of this series, in particular the policy gradient theorem of Chapter 13.

Difficulty

The mathematics is beginning calculus, as the book says. The formal difficulty lies elsewhere. The goal is an identity between an expectation over a two-stage random experiment (an action from π\piπ, then a reward from νA\nu_AνA​) and a partial derivative in one coordinate of a vector-valued parameter. A proof has to justify exchanging the finite sum with the derivative and splitting the reward integral, and it has to use integrability of each νx\nu_xνx​. It also needs the fact that the baseline term vanishes because ∑x∂π(x)/∂H(a)=0\sum_x \partial\pi(x)/\partial H(a) = 0∑x​∂π(x)/∂H(a)=0. A scalar-parameter version of the log-sum-exp derivative does not suffice: the book differentiates in one coordinate H(a)H(a)H(a) while all other preferences are held fixed. For Exercise 2.7 the obvious unrolling of (2.6) does not apply directly, because the step size βn\beta_nβn​ varies with nnn and the book states neither the weights nor the range of α\alphaα.

Formalization scope

  • Actions are Fin k. Every statement quantifies over some action, so k≥1k \ge 1k≥1 whenever it has content. Preferences are vectors Fin k → ℝ. The partial derivative in coordinate aaa is the derivative of h↦f(update H a h)h \mapsto f(\text{update } H\ a\ h)h↦f(update H a h) at H(a)H(a)H(a). The soft-max derivative milestone is stated with HasDerivAt, so it also asserts differentiability.
  • Rewards: each νx\nu_xνx​ is a probability measure on R\mathbb RR with Integrable identity and mean q∗(x)q_*(x)q∗​(x). The expectation of a function of (A,R)(A, R)(A,R) is ∑xπ(x)∫⋅ dνx\sum_x \pi(x) \int \cdot \, d\nu_x∑x​π(x)∫⋅dνx​. The book's normal-distribution testbed is only an example.
  • The baseline is a fixed real BBB, the book's "any scalar that does not depend on" the action (pp. 39–40). The book's Bt=RˉtB_t = \bar R_tBt​=Rˉt​, the average of past rewards, is covered once one conditions on the past. Footnote 1 on p. 37 states that the chapter's experiments used a Rˉt\bar R_tRˉt​ that also included RtR_tRt​. That baseline depends on AtA_tAt​, and the identity does not cover it.
  • Rewards of one action are a sequence indexed from 111. Q1Q_1Q1​ is arbitrary, and 00=10^0 = 100=1 as in the book (p. 33), so α=1\alpha = 1α=1 is included in (2.6).
  • Exercise 2.7 speaks of "a conventional constant step size α>0\alpha > 0α>0". The formalization takes α∈(0,1]\alpha \in (0,1]α∈(0,1], the range of the constant step size in (2.5). For α=2\alpha = 2α=2 the trace oˉn\bar o_noˉn​ vanishes at every even nnn and βn\beta_nβn​ is undefined. "Without initial bias" is read as "for n≥1n \ge 1n≥1, Qn+1Q_{n+1}Qn+1​ is the displayed weighted average of R1,…,RnR_1, \dots, R_nR1​,…,Rn​ with weights summing to one", which in particular does not involve Q1Q_1Q1​.
  • Exercise 2.9 is read as the two equalities π(1)=σ(H(1)−H(2))\pi(1) = \sigma(H(1)-H(2))π(1)=σ(H(1)−H(2)) and π(2)=σ(H(2)−H(1))\pi(2) = \sigma(H(2)-H(1))π(2)=σ(H(2)−H(1)).
  • A trivializing formalization is ruled out: the goal is about the expected value of the algorithm's update (2.12) under the joint law of action and reward, not the soft-max derivative alone and not a version in which the reward is replaced by its mean or the expectation is taken over AAA only.
  • Not formalized: the UCB rule (2.10) and the 10-armed testbed, which carry no provable claim in the chapter, and the stochastic-approximation conditions (2.7), which the book cites without proof.
  • Welcome contributions: a general soft-max library (derivatives, Jacobian, log-sum-exp) over a finite type, reusable for Chapter 13, and proofs of the milestones in the listed order.

Selected references

  • R. S. Sutton, A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018, ISBN 9780262039246, Chapter 2, pp. 25–46. http://incompleteideas.net/book/the-book-2nd.html
  • R. J. Williams, Simple statistical gradient-following algorithms for connectionist reinforcement learning, Machine Learning 8 (1992) 229–256. https://doi.org/10.1007/BF00992696
  • H. Robbins, S. Monro, A stochastic approximation method, Annals of Mathematical Statistics 22 (1951) 400–407. https://doi.org/10.1214/aoms/1177729586
11 thms2 active usersReviewed
🏆Completed
Dynamic ProgrammingMachine LearningOperations Research·Captain: mikedeng1

Approximately Optimal Approximate Reinforcement Learning II: Near-Optimality of a Policy with Small Policy AdvantageResearch Paper

Motivation

Approximate policy iteration and policy-gradient methods stop when they can no longer find a direction of improvement. Kakade and Langford (ICML 2002) asked what such a stopping point guarantees. Their algorithm, conservative policy iteration, halts at a policy π\piπ for which no policy can improve much on π\piπ as measured under a restart distribution μ\muμ; the quantity that is small is the optimal policy advantage OPT(Aπ,μ)\mathrm{OPT}(\mathbb A_{\pi,\mu})OPT(Aπ,μ​). Theorem 6.2 of the paper translates this local condition into a global statement: the performance of π\piπ is close to optimal, with a loss controlled by how well μ\muμ covers the states an optimal policy visits.

The bound is the origin of the distribution mismatch coefficient ∥dπ∗,μ~/μ∥∞\|d_{\pi^*,\tilde\mu}/\mu\|_\infty∥dπ∗,μ~​​/μ∥∞​, which reappears in the analysis of approximate dynamic programming (concentrability coefficients, Munos 2003), of conservative and trust-region methods, and of the convergence of policy gradient methods (Agarwal, Kakade, Lee, Mahajan 2021), where it governs the rate. The performance difference lemma (Lemma 6.1) used in its proof has become a standard tool of reinforcement learning theory.

Setting

A finite Markov decision process has a finite nonempty state set SSS, a finite nonempty action set AAA, transition probabilities P(s′;s,a)P(s';s,a)P(s′;s,a) (for each s,as,as,a a probability distribution over s′s's′), a reward function R:S×A→[0,R]\mathcal R:S\times A\to[0,R]R:S×A→[0,R] with R>0R>0R>0, and a discount factor 0≤γ<10\le\gamma<10≤γ<1. A stochastic policy π(a;s)\pi(a;s)π(a;s) is, for each state sss, a probability distribution over actions. A state distribution is a probability vector μ\muμ on SSS.

The normalized value function is Vπ(s)=(1−γ)E[∑t≥0γtR(st,at)∣π,s]V_\pi(s)=(1-\gamma)E[\sum_{t\ge0}\gamma^t\mathcal R(s_t,a_t)\mid\pi,s]Vπ​(s)=(1−γ)E[∑t≥0​γtR(st​,at​)∣π,s], where s0=ss_0=ss0​=s, at∼π(⋅;st)a_t\sim\pi(\cdot;s_t)at​∼π(⋅;st​) and st+1∼P(⋅;st,at)s_{t+1}\sim P(\cdot;s_t,a_t)st+1​∼P(⋅;st​,at​). The state–action value is Qπ(s,a)=(1−γ)R(s,a)+γ∑s′P(s′;s,a)Vπ(s′)Q_\pi(s,a)=(1-\gamma)\mathcal R(s,a)+\gamma\sum_{s'}P(s';s,a)V_\pi(s')Qπ​(s,a)=(1−γ)R(s,a)+γ∑s′​P(s′;s,a)Vπ​(s′) and the advantage is Aπ(s,a)=Qπ(s,a)−Vπ(s)A_\pi(s,a)=Q_\pi(s,a)-V_\pi(s)Aπ​(s,a)=Qπ​(s,a)−Vπ​(s). The discounted future state distribution from μ\muμ is

dπ,μ(s)=(1−γ)∑t≥0γtPr⁡(st=s;π,μ),s0∼μ,d_{\pi,\mu}(s)=(1-\gamma)\sum_{t\ge0}\gamma^t\Pr(s_t=s;\pi,\mu),\qquad s_0\sim\mu,dπ,μ​(s)=(1−γ)t≥0∑​γtPr(st​=s;π,μ),s0​∼μ,

and the performance of π\piπ from μ\muμ is ημ(π)=∑sμ(s)Vπ(s)\eta_\mu(\pi)=\sum_s\mu(s)V_\pi(s)ημ​(π)=∑s​μ(s)Vπ​(s).

The policy advantage of π′\pi'π′ with respect to π\piπ and μ\muμ is Aπ,μ(π′)=∑sdπ,μ(s)∑aπ′(a;s)Aπ(s,a)\mathbb A_{\pi,\mu}(\pi')=\sum_sd_{\pi,\mu}(s)\sum_a\pi'(a;s)A_\pi(s,a)Aπ,μ​(π′)=∑s​dπ,μ​(s)∑a​π′(a;s)Aπ​(s,a): the expected advantage of π′\pi'π′ over π\piπ on the states π\piπ itself visits. Its maximum over all stochastic policies is OPT(Aπ,μ)=max⁡π′Aπ,μ(π′)\mathrm{OPT}(\mathbb A_{\pi,\mu})=\max_{\pi'}\mathbb A_{\pi,\mu}(\pi')OPT(Aπ,μ​)=maxπ′​Aπ,μ​(π′) (Definition 4.3). An optimal policy π∗\pi^*π∗ satisfies Vπ(s)≤Vπ∗(s)V_\pi(s)\le V_{\pi^*}(s)Vπ​(s)≤Vπ∗​(s) for every policy π\piπ and every state sss. For nonnegative f,gf,gf,g on SSS, ∥f/g∥∞=max⁡sf(s)/g(s)\|f/g\|_\infty=\max_sf(s)/g(s)∥f/g∥∞​=maxs​f(s)/g(s) (p. 5).

Formalization targets

Goal: Theorem 6.2 (p. 6)

If OPT(Aπ,μ)<ε\mathrm{OPT}(\mathbb A_{\pi,\mu})<\varepsilonOPT(Aπ,μ​)<ε and π∗\pi^*π∗ is optimal, then for every state distribution μ~\tilde\muμ~​

ημ~(π∗)−ημ~(π)≤ε1−γ∥dπ∗,μ~dπ,μ∥∞≤ε(1−γ)2∥dπ∗,μ~μ∥∞.\eta_{\tilde\mu}(\pi^*)-\eta_{\tilde\mu}(\pi)\le\frac{\varepsilon}{1-\gamma}\left\|\frac{d_{\pi^*,\tilde\mu}}{d_{\pi,\mu}}\right\|_\infty\le\frac{\varepsilon}{(1-\gamma)^2}\left\|\frac{d_{\pi^*,\tilde\mu}}{\mu}\right\|_\infty.ημ~​​(π∗)−ημ~​​(π)≤1−γε​​dπ,μ​dπ∗,μ~​​​​∞​≤(1−γ)2ε​​μdπ∗,μ~​​​​∞​.

The goal states both inequalities and the outer bound. The evaluation distribution μ~\tilde\muμ~​ is arbitrary and unrelated to the restart distribution μ\muμ; taking μ~=D\tilde\mu=Dμ~​=D, the start distribution, gives Corollary 4.5 (p. 5).

Milestone: Lemma 6.1 (p. 6)

For any policies π~\tilde\piπ~, π\piπ and any starting distribution μ\muμ,

ημ(π~)−ημ(π)=11−γE(a,s)∼π~dπ~,μ[Aπ(s,a)].\eta_\mu(\tilde\pi)-\eta_\mu(\pi)=\frac{1}{1-\gamma}E_{(a,s)\sim\tilde\pi d_{\tilde\pi,\mu}}\big[A_\pi(s,a)\big].ημ​(π~)−ημ​(π)=1−γ1​E(a,s)∼π~dπ~,μ​​[Aπ​(s,a)].

The states are weighted by the future state distribution of the new policy π~\tilde\piπ~, the advantage is that of the old policy π\piπ.

Significance

Theorem 6.2 is the quality guarantee for conservative policy iteration: combined with the paper's Theorem 4.4 (the algorithm stops with OPT(Aπ,μ)<2ε\mathrm{OPT}(\mathbb A_{\pi,\mu})<2\varepsilonOPT(Aπ,μ​)<2ε after polynomially many calls), it bounds the suboptimality of the returned policy for any target distribution, independently of the size of the state space except through the mismatch coefficient. It also explains the role of the restart distribution: a more uniform μ\muμ makes ∥dπ∗,μ~/μ∥∞\|d_{\pi^*,\tilde\mu}/\mu\|_\infty∥dπ∗,μ~​​/μ∥∞​ small. Lemma 6.1 is used throughout later theory, from trust-region policy optimization to the global convergence of policy gradient methods.

Both results are proved in the paper, with short arguments. The contribution of this mission is a machine-checked version of the infinite-horizon discounted statement in the paper's normalization, with the ∥⋅∥∞\|\cdot\|_\infty∥⋅∥∞​ ratios handled exactly, including states where a denominator vanishes. Neither the discounted performance difference lemma for stochastic policies nor the distribution mismatch bound is known to be formalized in Mathlib; a finite-horizon performance difference identity has been formalized separately and is a different statement.

Difficulty

The mathematics is short; the difficulty is in the infinite-horizon bookkeeping. The value function and dπ,μd_{\pi,\mu}dπ,μ​ are infinite series, and Lemma 6.1 relates the series of two different policies: its natural one-line argument uses the Bellman equation for VπV_\piVπ​, which is not the definition here, together with interchanges of infinite sums over time with finite sums over states and actions, each of which needs summability. Theorem 6.2 then needs two facts that are not stated as results in the paper: that OPT(Aπ,μ)\mathrm{OPT}(\mathbb A_{\pi,\mu})OPT(Aπ,μ​) equals ∑sdπ,μ(s)max⁡aAπ(s,a)\sum_sd_{\pi,\mu}(s)\max_aA_\pi(s,a)∑s​dπ,μ​(s)maxa​Aπ​(s,a) (the supremum over policies is attained by a greedy policy, and max⁡aAπ(s,a)≥0\max_aA_\pi(s,a)\ge0maxa​Aπ​(s,a)≥0), and that dπ,μ(s)≥(1−γ)μ(s)d_{\pi,\mu}(s)\ge(1-\gamma)\mu(s)dπ,μ​(s)≥(1−γ)μ(s). Reading the ℓ∞\ell_\inftyℓ∞​ ratio with real division would give a false statement when a denominator is zero; the statement avoids this.

Formalization scope

States and actions are finite nonempty types; policies and kernels are real-valued functions π s a (the paper's π(a;s)\pi(a;s)π(a;s)) and P s a s' (the paper's P(s′;s,a)P(s';s,a)P(s′;s,a)), with their distribution properties as explicit hypotheses. The published definitions IsTransitionKernel, IsPolicy, InducedTransition, OccupationDist, InducedReward and PolicyValue from the Foundations of Machine Learning series are reused; VπV_\piVπ​ is (1−γ)(1-\gamma)(1−γ) times PolicyValue, the defining series. OPT\mathrm{OPT}OPT is the supremum of the policy advantages over stochastic policies, which is the paper's maximum. Optimality of π∗\pi^*π∗ is relative to stationary stochastic policies, the paper's policy class; the existence of an optimal policy (the paper's "well known result", p. 2) is not part of this mission.

Every hypothesis is explicit: rewards in [0,R][0,R][0,R] with R>0R>0R>0, 0≤γ<10\le\gamma<10≤γ<1, PPP a kernel, π\piπ and π∗\pi^*π∗ stochastic policies, μ\muμ and μ~\tilde\muμ~​ state distributions. Each ∥f/g∥∞\|f/g\|_\infty∥f/g∥∞​ bound is stated multiplicatively: "X≤K∥f/g∥∞X\le K\|f/g\|_\inftyX≤K∥f/g∥∞​" is "X≤KCX\le KCX≤KC for every CCC with f(s)≤Cg(s)f(s)\le Cg(s)f(s)≤Cg(s) for all sss". When some g(s)=0<f(s)g(s)=0<f(s)g(s)=0<f(s) no such CCC exists and the bound is empty, which matches ∥f/g∥∞=+∞\|f/g\|_\infty=+\infty∥f/g∥∞​=+∞; no full-support assumption is made on μ\muμ or μ~\tilde\muμ~​. The hypothesis OPT(Aπ,μ)<ε\mathrm{OPT}(\mathbb A_{\pi,\mu})<\varepsilonOPT(Aπ,μ​)<ε is on the supremum itself, not on the closed form ∑sdπ,μ(s)max⁡aAπ(s,a)\sum_sd_{\pi,\mu}(s)\max_aA_\pi(s,a)∑s​dπ,μ​(s)maxa​Aπ​(s,a), which is a step of the proof; a formalization that assumed the closed form, or that divided by dπ,μd_{\pi,\mu}dπ,μ​ in real arithmetic, would not be this theorem. The proof of the theorem uses only that π∗\pi^*π∗ is a policy; optimality is kept as a hypothesis because the paper states it.

The proof on p. 7 twice writes dπ,μ(s)≤(1−γ)μ(s)d_{\pi,\mu}(s)\le(1-\gamma)\mu(s)dπ,μ​(s)≤(1−γ)μ(s); the inequality it uses, and the one stated on p. 5, is dπ,μ(s)≥(1−γ)μ(s)d_{\pi,\mu}(s)\ge(1-\gamma)\mu(s)dπ,μ​(s)≥(1−γ)μ(s). This slip is in the proof, not in the statement. Pages are PDF pages; the paper has no printed page numbers.

Useful reusable infrastructure: summability and Bellman equations for the normalized discounted value, dπ,μd_{\pi,\mu}dπ,μ​ as a probability distribution with dπ,μ≥(1−γ)μd_{\pi,\mu}\ge(1-\gamma)\mudπ,μ​≥(1−γ)μ, and attainment of OPT\mathrm{OPT}OPT by a greedy policy. Contributions of any of these as separate lemmas are welcome.

Selected references

  • S. Kakade, J. Langford, Approximately Optimal Approximate Reinforcement Learning, Proceedings of the 19th International Conference on Machine Learning (ICML), 2002. https://dl.acm.org/doi/10.5555/645531.656005
  • R. Munos, Error Bounds for Approximate Policy Iteration, ICML 2003. https://dl.acm.org/doi/10.5555/3041838.3041903
  • A. Agarwal, S. Kakade, J. Lee, G. Mahajan, On the Theory of Policy Gradient Methods: Optimality, Approximation, and Distribution Shift, Journal of Machine Learning Research 22(98), 2021. https://jmlr.org/papers/v22/19-736.html
  • J. Schulman, S. Levine, P. Abbeel, M. Jordan, P. Moritz, Trust Region Policy Optimization, ICML 2015. https://arxiv.org/abs/1502.05477
10 thms2 active usersReviewed
🏆Completed
Dynamic ProgrammingMachine LearningOperations Research+1·Captain: mikedeng1

Approximately Optimal Approximate Reinforcement Learning I: Conservative Policy Iteration Improves Monotonically and Returns a Near-Greedy PolicyResearch Paper

Motivation

Approximate policy iteration and policy gradient methods are the two classical families of reinforcement learning algorithms that work with approximate, sampled information instead of an exact model. Kakade and Langford (ICML 2002) observed that neither family answers three basic questions: is there a performance measure that is guaranteed to improve at every step, how hard is it to verify that an update improves it, and what performance is reached after a reasonable number of updates. Greedy approximate policy iteration can make the policy worse when the value estimates are slightly wrong at a few states, and policy gradient methods can stall on plateaus where estimating the gradient needs an enormous number of samples.

Their answer is conservative policy iteration: instead of jumping to a greedy policy, move only a controlled fraction of the way toward it, with a step size chosen from an estimate of how much the greedy policy helps. The paper proves that this update improves a restart-distribution performance measure monotonically, terminates after a number of iterations that depends only on the reward range and the target accuracy, and stops at a policy that the greedy oracle can no longer improve by much. The idea is the direct ancestor of trust-region and proximal policy optimization methods (TRPO, Schulman et al. 2015; PPO, Schulman et al. 2017), whose improvement bounds are refinements of the paper's Theorem 4.1.

Setting

A finite Markov decision process has a finite nonempty set of states SSS, a finite nonempty set of actions AAA, transition probabilities P(s′;s,a)P(s';s,a)P(s′;s,a) (for each state sss and action aaa, a probability distribution over next states s′s's′), a reward function R:S×A→[0,R]\mathcal R : S\times A\to[0,R]R:S×A→[0,R] with R>0R>0R>0, and a discount factor 0≤γ<10\le\gamma<10≤γ<1. A stochastic policy π(a;s)\pi(a;s)π(a;s) gives, for each state sss, a probability distribution over actions.

The normalized value of π\piπ from sss is Vπ(s)=(1−γ)E[∑t≥0γtR(st,at)∣π,s]V_\pi(s) = (1-\gamma)E[\sum_{t\ge0}\gamma^t\mathcal R(s_t,a_t)\mid\pi,s]Vπ​(s)=(1−γ)E[∑t≥0​γtR(st​,at​)∣π,s], where s0=ss_0=ss0​=s, at∼π(⋅ ;st)a_t\sim\pi(\cdot\,;s_t)at​∼π(⋅;st​), st+1∼P(⋅ ;st,at)s_{t+1}\sim P(\cdot\,;s_t,a_t)st+1​∼P(⋅;st​,at​); it lies in [0,R][0,R][0,R]. The state-action value is Qπ(s,a)=(1−γ)R(s,a)+γEs′∼P(s′;s,a)[Vπ(s′)]Q_\pi(s,a) = (1-\gamma)\mathcal R(s,a)+\gamma E_{s'\sim P(s';s,a)}[V_\pi(s')]Qπ​(s,a)=(1−γ)R(s,a)+γEs′∼P(s′;s,a)​[Vπ​(s′)] and the advantage is Aπ(s,a)=Qπ(s,a)−Vπ(s)∈[−R,R]A_\pi(s,a) = Q_\pi(s,a)-V_\pi(s)\in[-R,R]Aπ​(s,a)=Qπ​(s,a)−Vπ​(s)∈[−R,R].

For a state distribution μ\muμ (a restart distribution), the discounted future state distribution is dπ,μ(s)=(1−γ)∑t≥0γtPr⁡(st=s;π,μ)d_{\pi,\mu}(s) = (1-\gamma)\sum_{t\ge0}\gamma^t\Pr(s_t=s;\pi,\mu)dπ,μ​(s)=(1−γ)∑t≥0​γtPr(st​=s;π,μ) (eq. (2.1)), and the performance measure is ημ(π)=Es∼μ[Vπ(s)]\eta_\mu(\pi) = E_{s\sim\mu}[V_\pi(s)]ημ​(π)=Es∼μ​[Vπ​(s)].

The policy advantage of a policy π′\pi'π′ with respect to π\piπ and μ\muμ is

Aπ,μ(π′)=Es∼dπ,μ[Ea∼π′(a;s)[Aπ(s,a)]],\mathbb A_{\pi,\mu}(\pi') = E_{s\sim d_{\pi,\mu}}\big[E_{a\sim\pi'(a;s)}[A_\pi(s,a)]\big],Aπ,μ​(π′)=Es∼dπ,μ​​[Ea∼π′(a;s)​[Aπ​(s,a)]],

and OPT(Aπ,μ)=max⁡π′Aπ,μ(π′)\mathrm{OPT}(\mathbb A_{\pi,\mu}) = \max_{\pi'}\mathbb A_{\pi,\mu}(\pi')OPT(Aπ,μ​)=maxπ′​Aπ,μ​(π′). The conservative update (4.1) is πnew=(1−α)π+απ′\pi_{new} = (1-\alpha)\pi+\alpha\pi'πnew​=(1−α)π+απ′ with α∈[0,1]\alpha\in[0,1]α∈[0,1]. An ε\varepsilonε-greedy policy chooser GεG_\varepsilonGε​ (Definition 4.3) returns, for every policy π\piπ, a policy π′\pi'π′ with Aπ,μ(π′)≥OPT(Aπ,μ)−ε\mathbb A_{\pi,\mu}(\pi')\ge\mathrm{OPT}(\mathbb A_{\pi,\mu})-\varepsilonAπ,μ​(π′)≥OPT(Aπ,μ​)−ε.

Conservative policy iteration (§5) starts from any policy and repeats: call Gε(π,μ)G_\varepsilon(\pi,\mu)Gε​(π,μ) to get π′\pi'π′; form an ε3\frac\varepsilon33ε​-accurate estimate A^\hat{\mathbb A}A^ of Aπ,μ(π′)\mathbb A_{\pi,\mu}(\pi')Aπ,μ​(π′) from μ\muμ-restarts; if A^<2ε3\hat{\mathbb A}<\frac{2\varepsilon}3A^<32ε​, stop and return π\piπ; otherwise apply (4.1) with α=(1−γ)(A^−ε/3)4R\alpha = \frac{(1-\gamma)(\hat{\mathbb A}-\varepsilon/3)}{4R}α=4R(1−γ)(A^−ε/3)​ and repeat.

Formalization targets

Goal: Theorem 4.4 (p. 5)

With probability at least 1−δ1-\delta1−δ, conservative policy iteration (i) strictly improves ημ\eta_\muημ​ with every policy update, (ii) stops after at most 72R2/ε272R^2/\varepsilon^272R2/ε2 policy updates, and (iii) returns a policy π\piπ with

OPT(Aπ,μ)<2ε.\mathrm{OPT}(\mathbb A_{\pi,\mu}) < 2\varepsilon.OPT(Aπ,μ​)<2ε.

The estimation step is represented by its guarantee: each reached loop's estimate fails to be ε3\frac\varepsilon33ε​-accurate with probability at most δ/(N+1)\delta/(N+1)δ/(N+1), N=⌊72R2/ε2⌋N=\lfloor72R^2/\varepsilon^2\rfloorN=⌊72R2/ε2⌋.

Milestones

Lemma 6.1 (p. 6), the performance difference identity:

ημ(π~)−ημ(π)=11−γE(a,s)∼π~dπ~,μ[Aπ(s,a)].\eta_\mu(\tilde\pi)-\eta_\mu(\pi) = \frac1{1-\gamma}E_{(a,s)\sim\tilde\pi d_{\tilde\pi,\mu}}[A_\pi(s,a)].ημ​(π~)−ημ​(π)=1−γ1​E(a,s)∼π~dπ~,μ​​[Aπ​(s,a)].

Theorem 4.1 (p. 4), with ε=max⁡s∣Ea∼π′(a;s)[Aπ(s,a)]∣\varepsilon=\max_s|E_{a\sim\pi'(a;s)}[A_\pi(s,a)]|ε=maxs​∣Ea∼π′(a;s)​[Aπ​(s,a)]∣ and all α∈[0,1]\alpha\in[0,1]α∈[0,1]:

ημ(πnew)−ημ(π)≥α1−γ(A−2αγε1−γ(1−α)).\eta_\mu(\pi_{new})-\eta_\mu(\pi)\ge\frac{\alpha}{1-\gamma}\Big(\mathbb A-\frac{2\alpha\gamma\varepsilon}{1-\gamma(1-\alpha)}\Big).ημ​(πnew​)−ημ​(π)≥1−γα​(A−1−γ(1−α)2αγε​).

Corollary 4.2 (p. 5): if A≥0\mathbb A\ge0A≥0, the step size α=(1−γ)A4R\alpha=\frac{(1-\gamma)\mathbb A}{4R}α=4R(1−γ)A​ gives

ημ(πnew)−ημ(π)≥A28R.\eta_\mu(\pi_{new})-\eta_\mu(\pi)\ge\frac{\mathbb A^2}{8R}.ημ​(πnew​)−ημ​(π)≥8RA2​.

Significance

Theorem 4.4 is the first guarantee of its kind for approximate reinforcement learning: the number of iterations is bounded by 72R2/ε272R^2/\varepsilon^272R2/ε2, independent of the number of states and of the restart distribution, and every iteration provably helps. Lemma 6.1 is the standard performance difference lemma, used throughout the analysis of policy optimization, including natural policy gradient and trust-region methods; Theorem 4.1 is the prototype of the "surrogate objective minus a penalty" bound that TRPO refines.

These results are proved in the paper. As far as a search of the platform shows, none is formalized: the platform's finite-horizon performance difference lemma (Foster and Rakhlin's Lemma 13) is a different statement, for episodic problems with non-stationary policies. This mission produces machine-checked versions of the discounted performance difference identity, the conservative improvement bound with its exact constants, and the high-probability termination and quality guarantee of the algorithm, all on top of an explicit infinite-horizon model rather than an assumed Bellman equation.

Difficulty

The obvious argument for the improvement bound expands ημ(πnew)\eta_\mu(\pi_{new})ημ​(πnew​) to first order in α\alphaα; that only gives α1−γA+O(α2)\frac{\alpha}{1-\gamma}\mathbb A+O(\alpha^2)1−γα​A+O(α2) with an unspecified constant, which cannot fix a step size. The exact bound needs control of how far the state distribution of the mixed policy drifts from that of the old policy, uniformly in time, and the performance difference identity is only useful once the states are weighted by the new policy's distribution. On the formal side, VπV_\piVπ​ and dπ,μd_{\pi,\mu}dπ,μ​ are infinite discounted series, so summability, exchanges of sums and the identities ∑sdπ,μ(s)=1\sum_s d_{\pi,\mu}(s)=1∑s​dπ,μ​(s)=1 and ∑aπ(a;s)Aπ(s,a)=0\sum_a\pi(a;s)A_\pi(s,a)=0∑a​π(a;s)Aπ​(s,a)=0 must all be established from the definitions. For Theorem 4.4, the algorithm is a random process whose policies depend on all earlier estimates; the argument has to be made pathwise on the event that every reached loop is accurate, together with a union bound over the loops that can be reached.

Formalization scope

Policies are functions π : S → A → ℝ with π s a the paper's π(a;s)\pi(a;s)π(a;s), and P s a s' is P(s′;s,a)P(s';s,a)P(s′;s,a); both are constrained by the published predicates IsPolicy and IsTransitionKernel. VπV_\piVπ​ is (1−γ)(1-\gamma)(1−γ) times the published series PolicyValue, so values are normalized as in the paper. OPT\mathrm{OPT}OPT is a real supremum over all stochastic policies; the set is nonempty and bounded, and the maximum is attained. Every theorem carries the standing assumptions of §2: finite nonempty SSS and AAA, a transition kernel, rewards in [0,R][0,R][0,R] with R>0R>0R>0, 0≤γ<10\le\gamma<10≤γ<1, and a state distribution μ\muμ. In Corollary 4.2, RRR is any upper bound on the rewards rather than necessarily the attained maximum.

In Theorem 4.4 the run is formalized pathwise, driven by arbitrary real random estimates on a probability space; the conclusion bounds the probability of the failure event by δ\deltaδ. Two deviations from the printed statement are disclosed. First, (ii) is stated for policy updates: the proof bounds updates, and the algorithm calls GεG_\varepsilonGε​ once more than it updates, so "at most 72R2/ε272R^2/\varepsilon^272R2/ε2 calls" is off by one. Second, the per-loop failure budget is δ/(N+1)\delta/(N+1)δ/(N+1), which covers the N+1N+1N+1 loops that may be reached. The Hoeffding estimate (5.1) is not formalized: as printed it concerns the ε6\frac\varepsilon66ε​-biased target, and its role is taken by the accuracy hypothesis. The step size is clipped at 111, which never binds when the estimate is accurate. No trivializing reading is available: the accuracy hypothesis is satisfied by a perfect estimator and a 000-greedy chooser exists, so the theorem is not vacuous, and strict improvement at every update is required, not merely nonnegative change.

Pages are PDF pages; the paper has no printed page numbers.

A complete development needs summability and algebra of discounted occupation measures, the performance difference identity, and a union bound over the loops of a random process; the first two are reusable for any discounted policy-optimization result. Proofs of the milestones in any order are welcome.

Selected references

  • S. Kakade and J. Langford, Approximately Optimal Approximate Reinforcement Learning, Proceedings of the 19th International Conference on Machine Learning (ICML), 2002. https://dl.acm.org/doi/10.5555/645531.656005
  • J. Schulman, S. Levine, P. Moritz, M. Jordan, P. Abbeel, Trust Region Policy Optimization, ICML 2015. https://arxiv.org/abs/1502.05477
  • J. Schulman, F. Wolski, P. Dhariwal, A. Radford, O. Klimov, Proximal Policy Optimization Algorithms, 2017. https://arxiv.org/abs/1707.06347
  • D. J. Foster and A. Rakhlin, Foundations of Reinforcement Learning and Interactive Decision Making, 2023. https://arxiv.org/abs/2312.16730
13 thms2 active usersReviewed
Dynamical SystemsProbabilityStochastic Systems·Captain: mikedeng1

The O.D.E. Method for Convergence of Stochastic Approximation and Reinforcement Learning I: Stability and Almost-Sure Convergence under Tapering StepsizesResearch Paper

Motivation

Stochastic approximation is the family of recursive algorithms that locate a zero of a function observed only through noisy evaluations. It goes back to Robbins and Monro (1951) and today underlies stochastic gradient descent, temporal-difference learning, Q-learning and actor–critic methods in reinforcement learning, and models of learning by boundedly rational agents.

The standard analysis is the O.D.E. method (Ljung 1977; see Kushner and Yin 1997): the interpolated iterates are compared with the solutions of an ordinary differential equation, and convergence of the algorithm follows from the stability of that ODE. The method has one well-known gap. It assumes, rather than proves, that the iterates remain bounded with probability one. In applications this stability hypothesis is often the hardest part: for asynchronous Q-learning and adaptive critic algorithms, almost sure boundedness had been proved only for discounted cost or after adding a projection step (Borkar and Meyn, p. 460).

Borkar and Meyn (SIAM J. Control Optim. 38 (2000)) close this gap with a scaling argument borrowed from the fluid-model approach to the stability of queueing networks (Dai 1995; Dai and Meyn 1995). They show that boundedness itself follows from the asymptotic stability of the origin for a second, "fluid-limit" ODE obtained by rescaling the drift. This mission formalizes that stability theorem for tapering step sizes, and the convergence theorem that follows from it.

Setting

Fix d≥0d\ge 0d≥0 and work in Rd\mathbb R^dRd with the Euclidean norm. Let h:Rd→Rdh:\mathbb R^d\to\mathbb R^dh:Rd→Rd and let {a(n)}n≥0\{a(n)\}_{n\ge0}{a(n)}n≥0​ be a deterministic sequence of positive step sizes. On a probability space (Ω,F,P)(\Omega,\mathcal F,\mathsf P)(Ω,F,P), random vectors X(n)X(n)X(n) and M(n)M(n)M(n) satisfy the stochastic approximation recursion

X(n+1)=X(n)+a(n)[h(X(n))+M(n+1)],n≥0.(1.1)X(n+1) = X(n) + a(n)\big[h(X(n)) + M(n+1)\big], \qquad n\ge0. \tag{1.1}X(n+1)=X(n)+a(n)[h(X(n))+M(n+1)],n≥0.(1.1)

Its mean ODE is x˙=h(x)\dot x = h(x)x˙=h(x) (1.2). For r>0r>0r>0 the scaled field is hr(x)=h(rx)/rh_r(x)=h(rx)/rhr​(x)=h(rx)/r, with the scaled ODE x˙=hr(x)\dot x = h_r(x)x˙=hr​(x) (1.4).

  • (A1) hhh is Lipschitz; hr(x)→h∞(x)h_r(x)\to h_\infty(x)hr​(x)→h∞​(x) as r→∞r\to\inftyr→∞ for every xxx; and the origin is an asymptotically stable equilibrium of the fluid-limit ODE x˙=h∞(x)\dot x = h_\infty(x)x˙=h∞​(x) (1.5).
  • (A2) With Fn\mathcal F_nFn​ the history of the iterates up to time nnn, {M(n)}\{M(n)\}{M(n)} is a martingale difference sequence, E[M(n+1)∣Fn]=0\mathsf E[M(n+1)\mid\mathcal F_n]=0E[M(n+1)∣Fn​]=0, and for some constant C0<∞C_0<\inftyC0​<∞, E[∥M(n+1)∥2∣Fn]≤C0(1+∥X(n)∥2)\mathsf E[\|M(n+1)\|^2\mid\mathcal F_n]\le C_0(1+\|X(n)\|^2)E[∥M(n+1)∥2∣Fn​]≤C0​(1+∥X(n)∥2).
  • (TS) Tapering step sizes: 0<a(n)≤10<a(n)\le10<a(n)≤1, ∑na(n)=∞\sum_n a(n)=\infty∑n​a(n)=∞, ∑na(n)2<∞\sum_n a(n)^2<\infty∑n​a(n)2<∞.

A point x∗x^*x∗ is globally asymptotically stable for x˙=h(x)\dot x = h(x)x˙=h(x) if it is a Lyapunov-stable equilibrium and every solution converges to it.

Formalization targets

Goal: Theorem 2.2 (almost sure convergence)

Under (A1), (A2) and (TS), if x˙=h(x)\dot x=h(x)x˙=h(x) has a unique globally asymptotically stable equilibrium x∗x^*x∗, then for every initial condition X(0)∈RdX(0)\in\mathbb R^dX(0)∈Rd,

X(n)⟶x∗almost surely.X(n)\longrightarrow x^* \qquad \text{almost surely.}X(n)⟶x∗almost surely.

The goal contains no constants and no rates, only the qualitative conclusion.

Milestone: Theorem 2.1 (i) (almost sure boundedness)

Under (A1), (A2) and (TS), for every initial condition,

sup⁡n∥X(n)∥<∞almost surely.\sup_n \|X(n)\| < \infty \qquad \text{almost surely.}nsup​∥X(n)∥<∞almost surely.

Milestones: the lemmas of Section 4.1

  • Lemma 4.1: the fluid-limit ODE is globally exponentially asymptotically stable.
  • Lemma 4.2: the piecewise ODE solutions ϕ^\hat\phiϕ^​, ϕ∞\phi^\inftyϕ∞ used for comparison are bounded by a constant independent of the initial condition.
  • Lemma 4.3 (i), (ii): two discrete Bellman–Gronwall inequalities.
  • Lemma 4.4: for large scale rrr, every solution of x˙=hr(x)\dot x = h_r(x)x˙=hr​(x) from the unit ball is ϵ\epsilonϵ-small on a window [T,T+1][T,T+1][T,T+1].
  • Lemma 4.5: the rescaled iterates have uniformly bounded second moments, and the rescaled noise sum ξ\xiξ is an L2L^2L2-bounded martingale.
  • Lemma 4.6: almost surely the rescaled interpolated iterates ϕ\phiϕ track ϕ^\hat\phiϕ^​ and stay bounded.

Significance

The result. Theorem 2.1 (i) turns the stability hypothesis of the O.D.E. method into a checkable condition on a deterministic ODE. Theorem 2.2 then gives convergence to x∗x^*x∗ with no a priori boundedness assumption. The paper applies this to reinforcement learning, obtaining the first convergence proof for asynchronous Q-learning and adaptive critic algorithms for average-cost Markov decision processes (the asynchronous extension, Theorem 2.5, is sketched in the paper and is not part of this mission). The same fluid-limit criterion is now a textbook tool; see Borkar, Stochastic Approximation: A Dynamical Systems Viewpoint (2008), Chapter 3.

Formalizing it. The theorems are proved in the paper, and the proofs are short but rely on several standard facts stated informally: uniform convergence of hrh_rhr​ to h∞h_\inftyh∞​ on compact sets, continuous dependence of ODE solutions on initial data and on the vector field, and the martingale convergence theorem. No machine-checked version of the O.D.E. method or of this stability criterion is known to exist. A formal development would give a verified link between discrete-time stochastic recursions, martingale convergence in Mathlib, and the stability theory of Lipschitz ODEs.

Difficulty

The obvious approach is to compare the iterates with solutions of x˙=h(x)\dot x = h(x)x˙=h(x) over windows of fixed ODE time and to control the accumulated noise by martingale convergence. This fails without boundedness: the noise bound in (A2) grows with ∥X(n)∥\|X(n)\|∥X(n)∥, so the deviation from the ODE can only be controlled relative to the current size of the iterate, and nothing prevents the iterates from escaping to infinity.

A second difficulty is that the hypothesis (A1) concerns only the fluid limit h∞h_\inftyh∞​, which describes the drift at infinite scale. It says nothing directly about hhh at any finite state, and nothing about the noise. Any argument therefore has to transfer information from the limit r→∞r\to\inftyr→∞ to the recursion at random, path-dependent scales, uniformly over those scales, while the noise is controlled only relative to the current size of the iterate. In Lean this involves ODE comparison and Gronwall estimates on a random partition of the time axis, conditional second-moment estimates for a rescaled recursion, and a vector-valued L2L^2L2 martingale convergence argument, none of which is available off the shelf for this setting.

Formalization scope

The state space is EuclideanSpace ℝ (Fin d). An ODE solution is a forward solution on [0,∞)[0,\infty)[0,∞): the derivative is taken within [0,∞)[0,\infty)[0,∞) at each t≥0t\ge0t≥0, which makes solutions continuous there. Stability notions are the standard ones (Lyapunov stability; asymptotic, global asymptotic and global exponential stability, the last in the form ∥x(t)−x∗∥≤be−δt∥x(0)−x∗∥\|x(t)-x^*\|\le b e^{-\delta t}\|x(0)-x^*\|∥x(t)−x∗∥≤be−δt∥x(0)−x∗∥). All vector fields in the mission are Lipschitz, so forward solutions exist and are unique, and quantifying over "every solution" is meaningful.

The filtration in (A2) is the natural filtration of the iterates. Because a(n)>0a(n)>0a(n)>0, it carries the same information as the paper's σ(X(i),M(i),i≤n)\sigma(X(i),M(i),i\le n)σ(X(i),M(i),i≤n). (A2) includes integrability of M(n+1)M(n+1)M(n+1) and ∥M(n+1)∥2\|M(n+1)\|^2∥M(n+1)∥2, so that the conditional expectations are meaningful. The theorems quantify over every probability space and every noise process satisfying (A2); the goal and Theorem 2.1 (i) take a deterministic initial condition, as the paper does. Stating the goal for a particular noise model (no noise, or i.i.d. noise) would be a different and much weaker theorem, and is ruled out. "sup⁡n∥X(n)∥<∞\sup_n\|X(n)\|<\inftysupn​∥X(n)∥<∞" is boundedness above of the set of norms, not a real supremum, which Lean sets to 000 on unbounded sets. Second-moment suprema in Lemma 4.5 are taken in [0,∞][0,\infty][0,∞].

The proof objects of Section 4.1 (time grid t(n)t(n)t(n), blocks m(j)m(j)m(j) and T(j)T(j)T(j), scales r(j)r(j)r(j), the interpolation ϕ\phiϕ, the rescaled iterates and noise sum) are separate definitions built from the step sizes and the sample path, as on the page. The piecewise ODE solutions ϕ^\hat\phiϕ^​ and ϕ∞\phi^\inftyϕ∞ are characterized by a predicate, and the lemmas hold for every function satisfying it.

Useful infrastructure, reusable beyond this mission: Lipschitz ODE comparison and continuous-dependence estimates in Mathlib's ODE library, uniform convergence of hrh_rhr​ on compact sets, the discrete Gronwall lemmas, and L2L^2L2-bounded vector-valued martingale convergence. Contributions are welcome on any milestone. The two Gronwall lemmas and Lemma 4.1 are self-contained entry points.

Selected references

  • V. S. Borkar and S. P. Meyn, The O.D.E. Method for Convergence of Stochastic Approximation and Reinforcement Learning, SIAM J. Control Optim. 38(2):447–469, 2000. https://doi.org/10.1137/S0363012997331639
  • H. Robbins and S. Monro, A Stochastic Approximation Method, Ann. Math. Statist. 22(3):400–407, 1951. https://doi.org/10.1214/aoms/1177729586
  • L. Ljung, Analysis of Recursive Stochastic Algorithms, IEEE Trans. Automat. Control 22(4):551–575, 1977. https://doi.org/10.1109/TAC.1977.1101561
  • H. J. Kushner and G. G. Yin, Stochastic Approximation Algorithms and Applications, Springer, 1997. https://doi.org/10.1007/978-1-4899-2696-8
  • J. G. Dai, On Positive Harris Recurrence of Multiclass Queueing Networks: A Unified Approach via Fluid Limit Models, Ann. Appl. Probab. 5(1):49–77, 1995. https://doi.org/10.1214/aoap/1177004828
  • J. G. Dai and S. P. Meyn, Stability and Convergence of Moments for Multiclass Queueing Networks via Fluid Limit Models, IEEE Trans. Automat. Control 40(11):1889–1904, 1995. https://doi.org/10.1109/9.471210
  • V. S. Borkar, Stochastic Approximation: A Dynamical Systems Viewpoint, Cambridge University Press / Hindustan Book Agency, 2008. https://doi.org/10.1007/978-93-86279-38-5
17 thms2 active usersReviewed
🏆Completed
Machine LearningStatistics·Captain: mikedeng1

Foundations of Reinforcement Learning V: General Decision Making and the Decision-Estimation Coefficient Lower BoundTextbook

Motivation

Online decision-making problems — multi-armed bandits, contextual bandits, structured bandits, and episodic reinforcement learning — look superficially different but share a common shape: a learner repeatedly acts, observes feedback, and is scored by regret against the best action in hindsight. Foster, Kakade, Qian and Rakhlin's Foundations of Reinforcement Learning and Interactive Decision Making (Foster & Rakhlin, arXiv:2312.16730v1) develops a unifying account of this shape and asks a sharper question than "does this specific algorithm work?": for a given class of possible environments, what is the best regret any algorithm can achieve? The Decision-Estimation Coefficient (DEC), introduced by Foster, Kakade, Qian and Rakhlin (2021, "The Statistical Complexity of Interactive Decision Making") and refined by Foster, Golowich, Qian, Rakhlin and Sekhari (2023), was proposed as the answer: a single real-valued complexity measure of a model class that simultaneously (i) drives a generic optimal-up-to-constants algorithm (Estimation-to-Decisions, E2D), and (ii) lower-bounds the regret of every algorithm. Item (ii) is what turns the DEC from "a complexity measure that happens to work for the algorithms we know" into a genuine characterization of statistical difficulty, in the same sense that minimax rates characterize the difficulty of estimation problems in classical statistics. This mission formalizes that lower bound.

Setting

Chapter 6 of the book (pp. 93–128) introduces Decision Making with Structured Observations (DMSO), a protocol general enough to subsume the contextual-bandit, structured-bandit and episodic tabular-RL protocols of earlier chapters. Over TTT rounds, the learner selects a decision πt\pi_tπt​ from a decision space Π\PiΠ; nature draws a reward-observation pair (rt,ot)(r_t, o_t)(rt​,ot​) from a fixed, unknown model M⋆(⋅∣πt)M^\star(\cdot \mid \pi_t)M⋆(⋅∣πt​), where a model MMM maps each decision to a distribution over a reward space RRR and an observation space OOO. The learner has access to a model class M\mathcal{M}M containing M⋆M^\starM⋆ (realizability). For M∈MM \in \mathcal{M}M∈M, write fM(π):=EM,π[r]f^M(\pi) := \mathbb{E}_{M,\pi}[r]fM(π):=EM,π​[r] for the mean reward function and πM:=arg⁡max⁡πfM(π)\pi_M := \arg\max_\pi f^M(\pi)πM​:=argmaxπ​fM(π) for the optimal decision; regret is Reg:=∑t=1TfM⋆(πM⋆)−Eπt∼pt[fM⋆(πt)]\mathrm{Reg} := \sum_{t=1}^T f^{M^\star}(\pi_{M^\star}) - \mathbb{E}_{\pi_t \sim p_t}[f^{M^\star}(\pi_t)]Reg:=∑t=1T​fM⋆(πM⋆​)−Eπt​∼pt​​[fM⋆(πt​)], exactly as in the bandit chapters, now for the general model class.

Because observations, not just mean rewards, now carry information, the DEC needs a way to measure distance between the full conditional distributions M(π)M(\pi)M(π) and M^(π)\hat M(\pi)M^(π), not just between scalars fM(π)f^M(\pi)fM(π) and fM^(π)f^{\hat M}(\pi)fM^(π). The chapter uses the squared Hellinger distance DH2D_H^2DH2​, one of a family of Csiszár fff-divergences that also includes total variation (DTVD_{TV}DTV​) and Kullback-Leibler (DKLD_{KL}DKL​) divergence. For a reference model M^\hat MM^ and scale γ>0\gamma > 0γ>0, the general Decision-Estimation Coefficient is the min-max game value

decγ(M,M^):=inf⁡p∈Δ(Π)sup⁡M∈MEπ∼p[fM(πM)−fM(π)−γ⋅DH2(M(π),M^(π))],\mathrm{dec}_\gamma(\mathcal{M}, \hat M) := \inf_{p \in \Delta(\Pi)} \sup_{M \in \mathcal{M}} \mathbb{E}_{\pi \sim p}\bigl[f^M(\pi_M) - f^M(\pi) - \gamma \cdot D_H^2(M(\pi), \hat M(\pi))\bigr],decγ​(M,M^):=p∈Δ(Π)inf​M∈Msup​Eπ∼p​[fM(πM​)−fM(π)−γ⋅DH2​(M(π),M^(π))],

and decγ(M):=sup⁡M^∈co(M)decγ(M,M^)\mathrm{dec}_\gamma(\mathcal{M}) := \sup_{\hat M \in \mathrm{co}(\mathcal{M})} \mathrm{dec}_\gamma(\mathcal{M}, \hat M)decγ​(M):=supM^∈co(M)​decγ​(M,M^). This mission's Lean development (FoundationsRL.GeneralDM) formalizes discrete versions of DTVD_{TV}DTV​, DH2D_H^2DH2​, DKLD_{KL}DKL​ for a finite outcome type, the DMSO regret, and this DEC.

Formalization targets

The goal is Proposition 28 (DEC Lower Bound), p. 105:

∃ c>0 (sufficiently small):∀ T with decεTc(M)≥10 εT,  εT:=c/T,  ∀ algorithm  p,  ∃ M∈M:regret(M,p)≥120 decεTc(M)⋅T.\exists\, c > 0 \text{ (sufficiently small)} : \forall\, T \text{ with } \mathrm{dec}^c_{\varepsilon_T}(\mathcal{M}) \ge 10\,\varepsilon_T,\; \varepsilon_T := c/\sqrt{T},\; \forall\, \text{algorithm}\; p,\; \exists\, M \in \mathcal{M} : \mathrm{regret}(M, p) \ge \tfrac{1}{20}\, \mathrm{dec}^c_{\varepsilon_T}(\mathcal{M}) \cdot T.∃c>0 (sufficiently small):∀T with decεT​c​(M)≥10εT​,εT​:=c/T​,∀algorithmp,∃M∈M:regret(M,p)≥201​decεT​c​(M)⋅T.

Here decεc\mathrm{dec}^c_\varepsilondecεc​ is the constrained DEC (§6.5.1), a variant of the offset DEC above that hard-constrains the information gain rather than subtracting it — a technical refinement needed to make the lower-bound direction go through — and the "localization condition" decεTc(M)≥10εT\mathrm{dec}^c_{\varepsilon_T}(\mathcal{M}) \ge 10\varepsilon_TdecεT​c​(M)≥10εT​ is a genuine hypothesis of the proposition, not a footnote. Unlike almost every other target in this series of missions, the statement quantifies over every algorithm rather than naming one: it is a genuine impossibility result. Two supporting divergence facts are included as milestones because the DEC's information-theoretic argument rests on them: Lemma 19 (DTV2≤DH2≤DKLD_{TV}^2 \le D_H^2 \le D_{KL}DTV2​≤DH2​≤DKL​) and Lemma 20 (a bounded-likelihood-ratio refinement bounding DKLD_{KL}DKL​ in terms of DH2D_H^2DH2​). The chapter's own matching upper bound, Proposition 26 (the E2D regret bound for the general DMSO protocol, the direct analogue of Chapter 4's Proposition 13), is included as a milestone to give the reader the matching pair the chapter presents together. Finally, Corollary 1 restates the lower bound in terms of the localized offset DEC (combining Proposition 28 with Proposition 27), included as a milestone showing the lower bound's reach beyond the constrained DEC alone.

Significance

Proposition 28 is what makes the DEC a genuine characterization of the statistical complexity of interactive decision making, rather than merely a sufficient condition for a particular algorithm family to succeed. Combined with the (uncited, technically deeper) matching upper bound for the constrained DEC — Proposition 29, stated but not proved in the book — it shows that for any finite model class, the constrained DEC is necessary and sufficient for low regret up to a log⁡∣M∣\sqrt{\log|\mathcal{M}|}log∣M∣​ factor in the localization radius: no complexity measure that is substantially different from the DEC can characterize the same problems. This is the general decision-making analogue of how minimax rates pin down statistical estimation, now for interactive protocols with adaptive feedback.

Formalizing the lower bound is new work: no result of this shape exists on the Prove2Me platform (searches for "decision-estimation", "general divergence", "constrained DEC" and "Hellinger" — the last of which surfaces two related-but-distinct affinity/Le Cam bounds from a different mission on bandit lower bounds — return no faithful prior art; see MODERATION_NOTES.md). The formal statement is the boxed proposition; the book gives a self-contained but simplified proof (two named simplifying assumptions, §6.5.3) and cites Foster, Golowich, Qian, Rakhlin & Sekhari (2023) for the unrestricted argument. This mission's Lean items are draft statements (:= by sorry), not proofs; formalizing the proof itself — a two-point adaptive testing argument using the chain rule for KL divergence and a change-of-measure step — is the open contribution this mission proposes.

Difficulty

The obvious first attempt is to try to prove the lower bound by exhibiting one fixed pair of hard models M,M^M, \hat MM,M^, as in classical two-point minimax lower bounds (Le Cam's method, Fano's inequality). This fails here because the decision-making protocol is interactive and adaptive: the algorithm's queries depend on what it has observed, so a model pair chosen obliviously (before seeing the algorithm) cannot in general be made indistinguishable to every algorithm — an adaptive algorithm can be constructed that distinguishes any two fixed models quickly by querying where they differ. The book's proof instead selects the "hard" alternative model MMM as a function of the algorithm's own strategy (via the constrained DEC's arg max, Eq. (6.36)), so that the pair is hard specifically for the algorithm under consideration, then uses the chain rule for KL divergence plus the change-of-measure identity between the algorithm's induced distributions under MMM and M^\hat MM^ to conclude that the algorithm's realized decisions must look similar under both models — hence it cannot get low regret on both simultaneously. Every step of this argument depends on the exact game structure of the constrained DEC, not just its numerical value; a formalization that leaves decεc\mathrm{dec}^c_\varepsilondecεc​ as an unconstrained real parameter (rather than the actual inf⁡\infinf-sup⁡\supsup game with its information-gain constraint) would make the lower bound's conclusion vacuous, since the hypothesis decεTc(M)≥10εT\mathrm{dec}^c_{\varepsilon_T}(\mathcal{M}) \ge 10\varepsilon_TdecεT​c​(M)≥10εT​ would no longer track any actual property of M\mathcal{M}M.

Formalization scope

The decision space Π\PiΠ and the outcome (reward, observation) alphabet YYY are both taken as finite types (Fintype); a model m:Π→Y→Rm : \Pi \to Y \to \mathbb{R}m:Π→Y→R is a conditional probability vector, and a reward-extraction map rew:Y→R\mathrm{rew} : Y \to \mathbb{R}rew:Y→R recovers the mean reward fm(π)=∑ym(π)(y)⋅rew(y)f^m(\pi) = \sum_y m(\pi)(y)\cdot\mathrm{rew}(y)fm(π)=∑y​m(π)(y)⋅rew(y). hellingerSq, totalVariationDiscrete, klDivDiscrete specialize the book's general dominating-measure divergence formula (Eq. (6.5)) to the counting measure on this finite type; klDivDiscrete returns an ENNReal so its +∞+\infty+∞ case (when PPP is not absolutely continuous w.r.t. QQQ) is represented honestly. The DEC, the constrained DEC and the localized subclass are literal sInf-of-sSup/sSup-of-sSup transcriptions of the book's min-max games — the same convention this series uses for the Chapter-4 DEC — not opaque free real numbers, which rules out the trivializing formalization named above.

Three deviations from this series' usual convention of pinning every constant to the value the book's own proof derives are deliberate and disclosed. First, the numerical constant ccc in εT:=c/T\varepsilon_T := c/\sqrt{T}εT​:=c/T​ is explicitly called "not important" by the authors themselves (footnote a, p. 105); it is existentially quantified (∃ c > 0) rather than pinned to a numeral. Second — added at moderation, round 2, 2026-09-19, after the constant was found to be pinned incorrectly — the lower bound's own multiplicative constant is also existentially quantified (∃ c' > 0) rather than pinned to 1/20. The book's printed proof (§6.5.3, pp. 107–110) derives 1/20 (p. 110, not p. 109 as an earlier draft of this mission stated) only under two named simplifying assumptions the theorem's hypotheses do not carry (p. 107, "Simplifications": a class-wide bounded-curvature hypothesis, Eq. (6.34); and a bound on the unaugmented sup⁡M^∈Mdeccε(M,M^)\sup_{\hat M\in\mathcal M}\mathrm{decc}_\varepsilon(M,\hat M)supM^∈M​deccε​(M,M^) rather than the officially-defined, augmented deccε(M)=sup⁡M^∈co(M)deccε(M∪{M^},M^)\mathrm{decc}_\varepsilon(M) = \sup_{\hat M\in\mathrm{co}(\mathcal M)} \mathrm{decc}_\varepsilon(M\cup\{\hat M\},\hat M)deccε​(M)=supM^∈co(M)​deccε​(M∪{M^},M^) this mission's decC implements). Since augmenting either supremum's domain can only raise its value, the printed proof's bound on the narrower, unaugmented quantity does not license a pinned 1/20 against the fully general decC this theorem states; the book itself attributes the proof of the general statement to an external reference (Foster, Golowich, Qian, Rakhlin & Sekhari 2023) not in this document. The existential c' matches the book's own unpinned ≳\gtrsim≳ for Proposition 28 as printed on pp. 105–106. Third, "any algorithm" and E[Reg(T)]\mathbb{E}[\mathrm{Reg}(T)]E[Reg(T)] are formalized, as throughout this series, without a full stochastic-process/history model: regret is a deterministic quantity evaluated at a fixed realized decision-distribution sequence p:Fin T→Π→Rp : \mathrm{Fin}\,T \to \Pi \to \mathbb{R}p:FinT→Π→R, rather than an expectation over an adaptive, history-dependent algorithm's own randomness. Formalizing the fully adaptive, measure-theoretic version of "any algorithm" — with an explicit filtration and expectation over the induced process law PMP_MPM​ — is future work a solver could add; the current statement is faithful to the book's deterministic-per-realization content but not to its full generality over randomized, history-dependent strategies. The DMSO protocol (Def_FoundationsRL_GeneralDM_Protocol) and the DEC (Def_FoundationsRL_GeneralDM_DEC) are restated locally rather than imported from Chapter 4's mission (FoundationsRL.Structured), since draft items cannot import another chunk's drafts; contributions extending either mission to reuse the other's substrate once both are published are welcome.

Selected references

  • Foster, D. J., Kakade, S. M., Qian, J., & Rakhlin, A. (2023). Foundations of Reinforcement Learning and Interactive Decision Making. arXiv:2312.16730.
  • Foster, D. J., Kakade, S. M., Qian, J., & Rakhlin, A. (2021). The Statistical Complexity of Interactive Decision Making. arXiv:2112.13487.
  • Foster, D. J., Golowich, N., Qian, J., Rakhlin, A., & Sekhari, A. (2023). A Unified Model and Dimension for Interactive Estimation. arXiv:2306.06184.
  • Polyanskiy, Y., & Wu, Y. Information Theory: From Coding to Learning. Cambridge University Press (draft edition cited by the book as [68]).
10 thms2 active usersReviewed
🏆Completed
Bandit AlgorithmsMachine LearningStatistics·Captain: mikedeng1

Foundations of Reinforcement Learning II: Contextual Bandits and Inverse Gap WeightingTextbook

Motivation

Decision-making problems rarely present the same fixed choice twice. A doctor prescribing a treatment sees each patient's medical history and symptoms before deciding; a website choosing which article to show sees the visitor's profile first. The multi-armed bandit model — where the learner repeatedly picks from a fixed set of arms with no side information — cannot express this: it is blind to the covariates that any real decision-maker actually observes. The contextual bandit model closes this gap by letting the learner see a context before acting, and asks for a decision rule that generalizes across contexts rather than memorizing a policy per context. Foster and Rakhlin's Foundations of Reinforcement Learning and Interactive Decision Making (arXiv:2312.16730v1, Section 3, pp. 38–53) develops this model and its algorithms as the bridge between supervised learning and sequential decision making, en route to general reinforcement learning. Contextual bandits with a learned reward-function class underlie production systems for content recommendation, online advertising, and adaptive clinical trial design (Li et al., A Contextual-Bandit Approach to Personalized News Article Recommendation, 2010, https://arxiv.org/abs/1003.0146; Agarwal et al., Making Contextual Decisions with Low Technical Debt, 2016, https://arxiv.org/abs/1606.03966).

The algorithmic history in this chapter runs through two distinct principles. The optimism principle (LinUCB, Section 3.2) generalizes the UCB algorithm to contexts under a linear reward model, but the chapter's own Example 3.1 (Section 3.3) shows optimism fails outside such structured classes, incurring regret linear in the size of the context space or the class. Foster and Rakhlin then present two "black-box" alternatives that use any function class FFF through an abstract regression subroutine: the naive ε\varepsilonε-Greedy method (Section 3.4), and the Inverse Gap Weighting (IGW) strategy underlying the SquareCB algorithm (Bietti, Agarwal & Langford, A Contextual Bandit Bake-off, 2018, https://arxiv.org/abs/1802.04064; Foster & Rakhlin, Beyond UCB: Optimal and Efficient Contextual Bandits with Regression Oracles, 2020, https://arxiv.org/abs/2002.04926). SquareCB attains a regret rate that both generalizes across contexts (no dependence on the size of the context space) and matches the optimal T\sqrt{T}T​ rate — improving on ε\varepsilonε-Greedy's T2/3T^{2/3}T2/3 rate — while remaining agnostic to the internal structure of FFF.

Setting

Over TTT rounds, a decision-maker faces the contextual bandit protocol: at each round ttt, it observes a context xt∈Xx_t \in Xxt​∈X, selects a decision πt\pi_tπt​ from a finite action set Π={1,…,A}\Pi = \{1,\dots,A\}Π={1,…,A}, and observes a reward rt∈Rr_t \in \mathbb{R}rt​∈R. Rewards are generated independently as rt∼M⋆(⋅∣xt,πt)r_t \sim M^\star(\cdot \mid x_t, \pi_t)rt​∼M⋆(⋅∣xt​,πt​) for a fixed, unknown conditional model M⋆M^\starM⋆; write f⋆(x,π):=E[r∣x,π]f^\star(x,\pi) := \mathbb{E}[r \mid x, \pi]f⋆(x,π):=E[r∣x,π] for the mean reward function and π⋆(x):=arg⁡max⁡πf⋆(x,π)\pi^\star(x) := \arg\max_\pi f^\star(x,\pi)π⋆(x):=argmaxπ​f⋆(x,π) for the optimal, context-dependent policy. The context sequence x1,…,xTx_1,\dots,x_Tx1​,…,xT​ is arbitrary — fixed in advance or adversarially chosen — while rewards remain stochastic. Performance is measured by regret against π⋆\pi^\starπ⋆:

Reg:=∑t=1Tf⋆(xt,π⋆(xt))−∑t=1TEπt∼pt[f⋆(xt,πt)],\mathrm{Reg} := \sum_{t=1}^T f^\star(x_t,\pi^\star(x_t)) - \sum_{t=1}^T \mathbb{E}_{\pi_t\sim p_t}[f^\star(x_t,\pi_t)],Reg:=t=1∑T​f⋆(xt​,π⋆(xt​))−t=1∑T​Eπt​∼pt​​[f⋆(xt​,πt​)],

where ptp_tpt​ is the learner's (possibly randomized) action distribution at round ttt.

To generalize across contexts, the learner is given a class F⊆{f:X×Π→R}F \subseteq \{f : X\times\Pi \to \mathbb{R}\}F⊆{f:X×Π→R} with f⋆∈Ff^\star \in Ff⋆∈F, and aims for regret scaling with the statistical complexity log⁡∣F∣\log|F|log∣F∣ rather than with ∣X∣|X|∣X∣. Both algorithms in this mission access FFF only through an online regression oracle (Definition 3, p. 47): given the history (x1,π1,r1),…,(xt−1,πt−1,rt−1)(x_1,\pi_1,r_1),\dots,(x_{t-1},\pi_{t-1},r_{t-1})(x1​,π1​,r1​),…,(xt−1​,πt−1​,rt−1​), it returns an estimate f^t:X×Π→R\hat f_t : X\times\Pi\to\mathbb{R}f^​t​:X×Π→R satisfying, with probability at least 1−δ1-\delta1−δ, ∑t=1TEπt∼pt[(f^t(xt,πt)−f⋆(xt,πt))2]≤EstSq(F,T,δ)\sum_{t=1}^T \mathbb{E}_{\pi_t\sim p_t}[(\hat f_t(x_t,\pi_t)-f^\star(x_t,\pi_t))^2] \le \mathrm{EstSq}(F,T,\delta)∑t=1T​Eπt​∼pt​​[(f^​t​(xt​,πt​)−f⋆(xt​,πt​))2]≤EstSq(F,T,δ) — for instance, exponential weights on a finite class FFF achieves EstSq(F,T,δ)=log⁡(∣F∣/δ)\mathrm{EstSq}(F,T,\delta) = \log(|F|/\delta)EstSq(F,T,δ)=log(∣F∣/δ). SquareCB (p. 50–51) then samples its action from the Inverse Gap Weighting distribution (Definition 4, p. 50): given a vector of estimated values f^∈RA\hat f \in \mathbb{R}^Af^​∈RA with greedy action πˉ=arg⁡max⁡πf^(π)\bar\pi = \arg\max_\pi \hat f(\pi)πˉ=argmaxπ​f^​(π), and an exploration parameter γ≥0\gamma \ge 0γ≥0, p=IGWγ(f^)p = \mathrm{IGW}_\gamma(\hat f)p=IGWγ​(f^​) is p(π)=1/(λ+2γ(f^(πˉ)−f^(π)))p(\pi) = 1/(\lambda + 2\gamma(\hat f(\bar\pi)-\hat f(\pi)))p(π)=1/(λ+2γ(f^​(πˉ)−f^​(π))) for the unique λ∈[1,A]\lambda \in [1,A]λ∈[1,A] making ppp a probability distribution.

Formalization targets

Milestone — Proposition 9 (IGW estimation-to-regret inequality)

Eπ∼p[f⋆(π⋆)−f⋆(π)]≤Aγ+γ⋅Eπ∼p[(f^(π)−f⋆(π))2],p=IGWγ(f^).\mathbb{E}_{\pi\sim p}[f^\star(\pi^\star)-f^\star(\pi)] \le \frac{A}{\gamma} + \gamma\cdot\mathbb{E}_{\pi\sim p}[(\hat f(\pi)-f^\star(\pi))^2], \qquad p = \mathrm{IGW}_\gamma(\hat f).Eπ∼p​[f⋆(π⋆)−f⋆(π)]≤γA​+γ⋅Eπ∼p​[(f^​(π)−f⋆(π))2],p=IGWγ​(f^​).

This holds for any f^,f⋆∈RA\hat f, f^\star \in \mathbb{R}^Af^​,f⋆∈RA and any γ>0\gamma>0γ>0, with no reference to FFF or to how f^\hat ff^​ was produced — it is the purely algebraic core the goal theorem invokes at every round.

Goal — Proposition 10 (SquareCB regret bound)

Reg≤2A T EstSq(F,T,δ)\mathrm{Reg} \le 2\sqrt{A\,T\,\mathrm{EstSq}(F,T,\delta)}Reg≤2ATEstSq(F,T,δ)​

with probability at least 1−δ1-\delta1−δ, for SquareCB run with γ=TA/EstSq(F,T,δ)\gamma = \sqrt{TA/\mathrm{EstSq}(F,T,\delta)}γ=TA/EstSq(F,T,δ)​, for any context sequence x1,…,xTx_1,\dots,x_Tx1​,…,xT​. This is the weakest stable target level in the chapter's oracle-based development: it is stated for an arbitrary class FFF and oracle, so it survives any future improvement to the oracle's own EstSq\mathrm{EstSq}EstSq bound, unlike a version hard-coded to a specific class or oracle.

Significance

Proposition 10 shows that Inverse Gap Weighting converts any estimation-error guarantee into a regret guarantee with the same statistical rate, with no algorithm-side dependence on the structure of FFF or the size of XXX: the same SquareCB template, driven by a plug-in regression oracle, is minimax optimal whenever the oracle itself is. When FFF is finite, this yields Reg≲ATlog⁡(∣F∣/δ)\mathrm{Reg} \lesssim \sqrt{AT\log(|F|/\delta)}Reg≲ATlog(∣F∣/δ)​, matching the optimal rate for stochastic multi-armed bandits (Section 2) while generalizing across contexts — a guarantee that optimism (Proposition 7) provably cannot deliver outside linear classes (Example 3.1), and that the simpler ε\varepsilonε-Greedy baseline (Proposition 8) only delivers at a slower T2/3T^{2/3}T2/3 rate. Foster and Rakhlin describe Proposition 9 itself as being "at the core of the development for the rest of the course": the same IGW mechanism reappears, generalized, in the book's treatment of general decision-making and the Decision-Estimation Coefficient.

Both propositions are proved results, not open questions; this mission's contribution is a machine-checked formalization of their exact statements and hypotheses — the precise OracleGuarantee hypothesis Proposition 10 requires, the exact constant (222, not a bare ≲\lesssim≲) its proof yields at the stated optimal γ\gammaγ, and the universally-quantified form of the IGW inequality (Proposition 9) that makes it reusable independently of any particular oracle or class.

Difficulty

The obvious first idea for exploiting an estimator f^t\hat f_tf^​t​ is a UCB-style optimism approach: build a confidence set around f^t\hat f_tf^​t​ and act greedily on its upper envelope, as in LinUCB (Proposition 7). Example 3.1 shows this fails in general: a class FFF can force the confidence set to remain wide on a fresh action at every new context, driving regret linear in min⁡{∣F∣,∣X∣}\min\{|F|,|X|\}min{∣F∣,∣X∣} — the confidence width in the regret bound does not shrink merely because the oracle's cumulative estimation error is small, since that error is not localized to the specific action the confidence-set approach tries next. Uniform exploration (ε\varepsilonε-Greedy) sidesteps this but wastes exploration budget on actions already known to be far from optimal, which is what caps its rate at T2/3T^{2/3}T2/3 (Proposition 8). Inverse Gap Weighting instead ties the sampling probability itself to the estimated gap from the greedy action, so cheap-to-rule-out actions are down-weighted continuously rather than either fully explored (ε-Greedy) or trusted outright (optimism); the technical content of Proposition 9 is showing this specific reciprocal-gap form gives a bound with no hidden dependence on FFF or XXX, for every pair (f^,f⋆)(\hat f, f^\star)(f^​,f⋆) simultaneously — a guarantee optimism cannot match because its confidence sets are class-dependent by construction.

Formalization scope

Contexts form an arbitrary type X; actions are Fin A for A : ℕ. A finite probability distribution over Fin A is represented directly as p : Fin A → ℝ with ∀ π, 0 ≤ p π and ∑ π, p π = 1, and Eπ∼p[g]\mathbb{E}_{\pi\sim p}[g]Eπ∼p​[g] as the finite sum ∑ π, p π * g π, rather than via Mathlib's PMF (which is ℝ≥0∞-valued) — an equivalent and lighter-weight representation of a distribution on a finite type. The normalizing constant λ\lambdaλ of Definition 4 and the optimal actions π⋆\pi^\starπ⋆, πˉ\bar\piπˉ are each specified by their defining property (existence of λ∈[1,A]\lambda \in [1,A]λ∈[1,A] realizing the IGW formula; ∀π,f(π)≤f(argmax)\forall\pi, f(\pi)\le f(\text{argmax})∀π,f(π)≤f(argmax)) rather than constructed explicitly via an intermediate-value or Finset.argmax argument, avoiding committing to one choice function for a value the book itself leaves implicit. The class FFF enters neither proposition's statement directly: it appears in the source only through the abstract bound EstSq(F,T,δ)\mathrm{EstSq}(F,T,\delta)EstSq(F,T,δ), which is carried as an explicit real-valued parameter and hypothesis (OracleGuarantee) rather than as a literal subset of a function space, since no property of FFF beyond producing this bound is ever used. The probability-(1−δ)(1-\delta)(1−δ) qualifier attached to the online regression oracle's guarantee is likewise the explicit hypothesis OracleGuarantee ... EstSq on a fixed realized run, rather than a statement quantified over an underlying probability space of histories — every subsequent step in both propositions' proofs is deterministic given that this event holds, so this does not weaken either conclusion. A trivializing formalization would fix A=1A=1A=1 (a single ever-optimal action, making both Reg and the IGW inequality vacuous) or take EstSq as an unconstrained free variable with no positivity hypothesis (making γ\gammaγ in Proposition 10 undefined); this mission's statements require 0 < EstSq and leave AAA, TTT, XXX, FFF-via-EstSq fully general.

This mission omits Proposition 7 (LinUCB): its proof rests on an entirely disjoint apparatus (finite linear parameter sets, least-squares confidence sets, the elliptic potential lemma) that neither Proposition 9 nor 10 requires, and Example 3.1 (the failure of optimism) is a worked example rather than a numbered, formalizable claim. It also omits Proposition 8 (ε\varepsilonε-Greedy): the source leaves the optimal ε\varepsilonε unspecified ("choosing ε\varepsilonε appropriately"), and deriving its own optimal value and matching constant independently — rather than reusing the book's own explicit constant, as Rule 7 of this formalization effort requires — was judged too likely to introduce an unfaithful, invented constant within this mission's time budget; both are natural extensions for a follow-up mission or contribution. Reusable infrastructure: the Fin A-indexed finite-distribution convention and the OracleGuarantee/optimal-action-by-property pattern extend directly to any later chapter built on the same online-regression-oracle abstraction.

Selected references

  • Foster, D. J. and Rakhlin, A. Foundations of Reinforcement Learning and Interactive Decision Making. 2023. https://arxiv.org/abs/2312.16730
  • Foster, D. J. and Rakhlin, A. Beyond UCB: Optimal and Efficient Contextual Bandits with Regression Oracles. ICML 2020. https://arxiv.org/abs/2002.04926
  • Bietti, A., Agarwal, A., and Langford, J. A Contextual Bandit Bake-off. JMLR 2021 (arXiv 2018). https://arxiv.org/abs/1802.04064
  • Li, L., Chu, W., Langford, J., and Schapire, R. E. A Contextual-Bandit Approach to Personalized News Article Recommendation. WWW 2010. https://arxiv.org/abs/1003.0146
  • Agarwal, A. et al. Making Contextual Decisions with Low Technical Debt. 2016. https://arxiv.org/abs/1606.03966
5 thms2 active usersReviewed
🏆Completed
Bandit AlgorithmsMachine Learning·Captain: mikedeng1

Foundations of Reinforcement Learning I: Multi-Armed Bandits and the UCB AlgorithmTextbook

Motivation

The multi-armed bandit is the simplest model of sequential decision-making under partial feedback: a learner repeatedly picks one of finitely many options and observes a reward only for the option chosen, never for the alternatives. It formalizes problems ranging from clinical trial design (which treatment to offer a patient) to online advertising (which ad to show) and A/B testing more generally. The framework dates to Robbins' 1952 paper on sequential design, and the algorithm this mission's goal theorem concerns — the Upper Confidence Bound (UCB) algorithm of Lai and Robbins [1985] and Auer, Cesa-Bianchi and Fischer [2002] — is the canonical answer to how to explore efficiently: instead of exploring uniformly at random, act optimistically with respect to the current uncertainty about each option's value. This mission draws its formalization from Chapter 2 of Foster and Rakhlin's 2023 lecture notes, Foundations of Reinforcement Learning and Interactive Decision Making, which develops the bandit problem as the first rung of a ladder of increasingly general interactive decision-making settings (contextual bandits, structured bandits, reinforcement learning) that the book's later chapters build.

Setting

Fix a finite decision (action) space Π={1,…,A}\Pi = \{1,\dots,A\}Π={1,…,A}. In the multi-armed bandit protocol, for each round t=1,…,Tt = 1,\dots,Tt=1,…,T the learner selects a decision πt∈Π\pi_t \in \Piπt​∈Π, possibly at random according to a distribution ptp_tpt​ depending on the history Ht−1=((π1,r1),…,(πt−1,rt−1))H_{t-1} = ((\pi_1,r_1),\dots,(\pi_{t-1},r_{t-1}))Ht−1​=((π1​,r1​),…,(πt−1​,rt−1​)) observed so far, and then observes a reward rt∈Rr_t \in \mathbb{R}rt​∈R drawn independently from a fixed conditional distribution M⋆(⋅∣πt)M^\star(\cdot \mid \pi_t)M⋆(⋅∣πt​) (the stochastic rewards assumption). Writing f⋆(π):=E[r∣π]f^\star(\pi) := \mathbb{E}[r \mid \pi]f⋆(π):=E[r∣π] for the mean reward function and π⋆:=arg⁡max⁡πf⋆(π)\pi^\star := \arg\max_\pi f^\star(\pi)π⋆:=argmaxπ​f⋆(π) for an optimal decision, the learner's performance is measured by the regret

Reg:=∑t=1Tf⋆(π⋆)−∑t=1TEπt∼pt[f⋆(πt)].\mathrm{Reg} := \sum_{t=1}^T f^\star(\pi^\star) - \sum_{t=1}^T \mathbb{E}_{\pi_t \sim p_t}[f^\star(\pi_t)].Reg:=t=1∑T​f⋆(π⋆)−t=1∑T​Eπt​∼pt​​[f⋆(πt​)].

Because the learner observes a reward only for the action played (bandit feedback), a purely greedy strategy that always plays the current empirical maximizer can commit to a suboptimal action forever, incurring linear regret; some form of deliberate exploration is necessary. The chapter's central construction is the confidence interval: a pair of functions f‾t,fˉt:Π→R\underline{f}_t, \bar f_t : \Pi \to \mathbb{R}f​t​,fˉ​t​:Π→R such that, with probability at least 1−δ1-\delta1−δ, f⋆(π)∈[f‾t(π),fˉt(π)]f^\star(\pi) \in [\underline{f}_t(\pi), \bar f_t(\pi)]f⋆(π)∈[f​t​(π),fˉ​t​(π)] for every round ttt and decision π\piπ simultaneously. The UCB algorithm plays the optimistic action πt=arg⁡max⁡πfˉt(π)\pi_t = \arg\max_\pi \bar f_t(\pi)πt​=argmaxπ​fˉ​t​(π) at every round, using the confidence interval built from Hoeffding's inequality around the empirical mean f^t(π)\hat f_t(\pi)f^​t​(π).

Formalization targets

Goal — Proposition 5 (UCB regret)

Reg  ≲  ATlog⁡(AT/δ)\mathrm{Reg} \;\lesssim\; \sqrt{AT\log(AT/\delta)}Reg≲ATlog(AT/δ)​

holding with probability at least 1−δ1-\delta1−δ, for the UCB algorithm using the confidence radius 2log⁡(2T2A/δ)/nt(π)\sqrt{2\log(2T^2A/\delta)/n_t(\pi)}2log(2T2A/δ)/nt​(π)​ of Eq. (2.19). This is the weakest stable statement the chapter proves: it is optimal up to the log factor, and strengthening it (e.g. to the sharper instance-dependent bound of Remark 10) is explicitly left to later work by the book itself.

Milestones

  • Proposition 4 (ε-Greedy regret): Reg≲A1/3T2/3log⁡1/3(AT/δ)\mathrm{Reg} \lesssim A^{1/3}T^{2/3}\log^{1/3}(AT/\delta)Reg≲A1/3T2/3log1/3(AT/δ) — the book's preceding, weaker result, establishing that naive forced exploration already gives sublinear regret, and motivating why an adaptive strategy (UCB) does better.
  • Lemma 7 (Optimism): the per-round regret of the optimistic action is bounded by the confidence width at that action.
  • Lemma 8 (Confidence width potential lemma): ∑t=1T(1/nt(πt)∧1)≲AT\sum_{t=1}^T (1/\sqrt{n_t(\pi_t)} \wedge 1) \lesssim \sqrt{AT}∑t=1T​(1/nt​(πt​)​∧1)≲AT​, a pigeonhole bound on how often any one action's confidence interval can still be wide.

Significance

UCB is the prototype of the "optimism in the face of uncertainty" principle that recurs, in increasingly abstract form, throughout the rest of the book: the same two-step argument (Lemma 7 + Lemma 8) reappears for linear bandits, structured bandits via the Decision-Estimation Coefficient, and UCB-VI for tabular reinforcement learning. Formalizing Chapter 2 in full therefore front-loads the proof pattern every later chapter in this series specializes. The result itself is also of standalone interest: the AT\sqrt{AT}AT​ minimax rate is the benchmark every subsequent bandit algorithm in the literature is compared against, and the A1/3T2/3A^{1/3}T^{2/3}A1/3T2/3-vs-AT\sqrt{AT}AT​ contrast between ε-Greedy and UCB is the standard illustration, in any course on the subject, of why adaptive exploration matters.

No formalization of this exact statement — realizability with respect to a function class f⋆∈F=RΠf^\star \in \mathcal{F} = \mathbb{R}^\Pif⋆∈F=RΠ and a generic confidence interval, rather than a per-arm sub-Gaussian empirical mean — currently exists on the platform (see Formalization scope below); the mission both proves this specific regret bound and seeds the generic optimism/potential lemma pair (Lemma 7, Lemma 8) that the book's later, more structured settings specialize.

Difficulty

The natural first attempt — bound the regret of the empirical-mean-greedy algorithm directly — fails outright: on a two-armed instance where one arm is deterministic and the other only slightly better in expectation, the greedy algorithm can commit to the worse arm forever with constant probability, giving linear, not sublinear, regret (§2.1). The obvious fix, ε-Greedy, forces exploration uniformly across all actions regardless of how much is already known about each, so the exploration cost scales with εT\varepsilon TεT even for actions whose value is already well determined — this is exactly what caps ε-Greedy at the T2/3T^{2/3}T2/3 rate. UCB's optimism principle resolves this by exploring an action only in proportion to how uncertain it still is; the technical core, isolated in Lemma 7 and Lemma 8, is disentangling "the algorithm made a mistake" from "the algorithm is still uncertain," which are conflated in the naive per-round regret decomposition used for ε-Greedy.

Formalization scope

Both the goal and the milestones fix a finite decision space Fin A, a mean reward function fStar : Fin A → ℝ with fStar π ∈ [0,1], and an optimal decision piStar. Regret is defined generically (Eq. (2.3)) via per-round decision weights p : ℕ → Fin A → ℝ, so it applies uniformly to a randomized algorithm (ε-Greedy) and a deterministic one (UCB, via the point mass at the played action). The book's "with probability at least 1−δ1-\delta1−δ" qualifier on both Proposition 4 and Proposition 5 is formalized as the deterministic consequence of the underlying concentration event (Eq. (2.9) and Eq. (2.18) respectively) holding — exactly the move the book's own proofs make ("Let us condition on the event in (2.18) ... "). The concentration events themselves rest on Hoeffding's inequality for adaptive stopping times (Lemma 33) and Bernstein's inequality (Lemma 5), both stated in the book's technical appendix outside this chapter, and are not drafted here; a solver may either take them as a hypothesis (as this mission's statements do) or import/prove them separately. A trivializing formalization is ruled out explicitly: taking δ outside (0,1)(0,1)(0,1), or dropping the fStar π ∈ [0,1] hypothesis, would make the stated constants vacuous or false, so both are retained as explicit hypotheses in every theorem. In every ≲ statement (Prop. 4, Lemma 8, Prop. 5) the witnessed constant C is quantified before the instance parameters (A, T, δ, and the realized sequences): ∃ C, 0 < C ∧ ∀ A T δ ..., Reg ≤ C * (rate), not the other order. This is deliberate, not stylistic: quantifying C after the instance lets it depend on A, T, δ, making the bound satisfiable by an arbitrarily large C chosen per instance and hence content-free, which is not what the book's ≲ means (a single constant working uniformly over all instances). Proposition 5's UCB decision rule is stated in the book's own two clauses, not collapsed into a single "maximize the upper confidence bound" rule: the confidence radius of Eq. (2.19) is +∞+\infty+∞ at nt(π)=0n_t(\pi)=0nt​(π)=0 (an action never yet sampled), so the book's UCB always plays an unsampled action before ever comparing indices, and only compares finite upper confidence bounds once every action has been sampled at least once; the confidence event of Eq. (2.18) is correspondingly assumed only at sampled actions, since the book's own bound is vacuous otherwise. An earlier draft instead capped the radius at 111 when nt(π)=0n_t(\pi)=0nt​(π)=0, which is a true statement about a different algorithm (a sampled action can have index above the capped unsampled index), and was corrected to the book's own rule after moderation. Reuse from the platform's existing bandit library (BanditAlgorithm, Lattimore & Szepesvári) is deliberately avoided: that library's UCB (bandit_ucb_regret_bound, bandit_ucb_minimax_regret_bound) is stated for per-arm 1-sub-Gaussian rewards with δ=1/n2\delta = 1/n^2δ=1/n2 fixed by the horizon, whereas this chapter's UCB is stated for a free failure probability δ\deltaδ and a generic confidence-interval abstraction (the multi-armed case being F=RΠ\mathcal{F} = \mathbb{R}^\PiF=RΠ of the book's general realizability framework) — the two are related but not the same statement. Contributions extending the mission with the generic confidence-interval form of Lemma 7/8 applied to other chapters in this series (contextual and structured bandits) are welcome.

Selected references

  • T. Lai and H. Robbins, Asymptotically Efficient Adaptive Allocation Rules, Advances in Applied Mathematics, 1985.
  • P. Auer, N. Cesa-Bianchi, and P. Fischer, Finite-time Analysis of the Multiarmed Bandit Problem, Machine Learning, 2002.
  • D. Foster and A. Rakhlin, Foundations of Reinforcement Learning and Interactive Decision Making, arXiv:2312.16730, 2023. https://arxiv.org/abs/2312.16730
  • T. Lattimore and C. Szepesvári, Bandit Algorithms, Cambridge University Press, 2020.
7 thms2 active usersReviewed
Machine LearningOptimization·Captain: mikedeng1

On the Theory of Policy Gradient Methods: Optimality, Approximation, and Distribution Shift 1: Projected Gradient Ascent on the Simplex Is ε-Optimal After 64γ|S||A|D∞²/((1−γ)⁶ε²) IterationsResearch Paper

Motivation

Policy gradient methods optimize a parameterized policy of a Markov decision process by gradient ascent on its expected discounted return. They are among the most widely used methods in reinforcement learning, from REINFORCE (Williams 1992) and the policy gradient theorem to natural policy gradient and trust-region methods. The objective is not concave in the policy, even when the policy is a raw table of action probabilities, so standard optimization theory guarantees at best convergence to a stationary point, and it was long unclear whether or how fast these methods find an optimal policy.

Agarwal, Kakade, Lee and Mahajan (JMLR 2021) give a systematic answer for the tabular and function-approximation settings. This mission formalizes their warm-up result, Theorem 4.1: projected gradient ascent over the simplex of stochastic policies reaches an ϵ\epsilonϵ-optimal policy after a number of iterations polynomial in the sizes of the MDP, the effective horizon 1/(1−γ)1/(1-\gamma)1/(1−γ), 1/ϵ1/\epsilon1/ϵ, and a distribution mismatch coefficient. The gradient domination idea it rests on goes back to the analysis of conservative policy iteration by Kakade and Langford (2002) and to Scherrer and Geist (2014).

Setting

A finite discounted MDP consists of finite sets S\mathcal SS of states and A\mathcal AA of actions, a transition kernel P(s′∣s,a)P(s'\mid s,a)P(s′∣s,a), rewards r(s,a)∈[0,1]r(s,a)\in[0,1]r(s,a)∈[0,1] and a discount factor γ∈[0,1)\gamma\in[0,1)γ∈[0,1). A policy π\piπ assigns to each state a probability distribution π(⋅∣s)\pi(\cdot\mid s)π(⋅∣s) over actions. Its value from a start state s0s_0s0​ is

Vπ(s0)=E[∑t=0∞γtr(st,at) ∣ s0],at∼π(⋅∣st), st+1∼P(⋅∣st,at),V^\pi(s_0)=\mathbb E\Big[\sum_{t=0}^\infty\gamma^t r(s_t,a_t)\,\Big|\,s_0\Big],\qquad a_t\sim\pi(\cdot\mid s_t),\ s_{t+1}\sim P(\cdot\mid s_t,a_t),Vπ(s0​)=E[t=0∑∞​γtr(st​,at​)​s0​],at​∼π(⋅∣st​), st+1​∼P(⋅∣st​,at​),

and for a start distribution ρ\rhoρ, Vπ(ρ)=∑sρ(s)Vπ(s)V^\pi(\rho)=\sum_s\rho(s)V^\pi(s)Vπ(ρ)=∑s​ρ(s)Vπ(s). The action value is Qπ(s,a)=r(s,a)+γ∑s′P(s′∣s,a)Vπ(s′)Q^\pi(s,a)=r(s,a)+\gamma\sum_{s'}P(s'\mid s,a)V^\pi(s')Qπ(s,a)=r(s,a)+γ∑s′​P(s′∣s,a)Vπ(s′) and the advantage is Aπ(s,a)=Qπ(s,a)−Vπ(s)A^\pi(s,a)=Q^\pi(s,a)-V^\pi(s)Aπ(s,a)=Qπ(s,a)−Vπ(s). The discounted state visitation distribution is dρπ(s)=(1−γ)∑t≥0γtPr⁡π(st=s∣s0∼ρ)d^\pi_\rho(s)=(1-\gamma)\sum_{t\ge0}\gamma^t\Pr^\pi(s_t=s\mid s_0\sim\rho)dρπ​(s)=(1−γ)∑t≥0​γtPrπ(st​=s∣s0​∼ρ). An optimal policy π⋆\pi^\starπ⋆ maximizes Vπ(s)V^\pi(s)Vπ(s) at every state simultaneously; V⋆=Vπ⋆V^\star=V^{\pi^\star}V⋆=Vπ⋆.

In the direct parameterization the parameter is the table itself, πs,a=π(a∣s)\pi_{s,a}=\pi(a\mid s)πs,a​=π(a∣s), a point of the product simplex Δ(A)∣S∣⊆RS×A\Delta(\mathcal A)^{|\mathcal S|}\subseteq\mathbb R^{\mathcal S\times\mathcal A}Δ(A)∣S∣⊆RS×A. The algorithm optimizes Vπ(μ)V^\pi(\mu)Vπ(μ) for a chosen start distribution μ\muμ by projected gradient ascent

π(t+1)=PΔ(A)∣S∣(π(t)+η∇πV(t)(μ)),\pi^{(t+1)}=P_{\Delta(\mathcal A)^{|\mathcal S|}}\big(\pi^{(t)}+\eta\nabla_\pi V^{(t)}(\mu)\big),π(t+1)=PΔ(A)∣S∣​(π(t)+η∇π​V(t)(μ)),

where PΔ(A)∣S∣P_{\Delta(\mathcal A)^{|\mathcal S|}}PΔ(A)∣S∣​ is the Euclidean projection and V(t)=Vπ(t)V^{(t)}=V^{\pi^{(t)}}V(t)=Vπ(t). Performance is measured under a possibly different distribution ρ\rhoρ, and the distribution mismatch coefficient ∥dρπ⋆/μ∥∞\|d^{\pi^\star}_\rho/\mu\|_\infty∥dρπ⋆​/μ∥∞​ (componentwise ratio) measures how well μ\muμ covers the states an optimal policy visits from ρ\rhoρ.

Formalization targets

Goal: Theorem 4.1

With step size η=(1−γ)3/(2γ∣A∣)\eta=(1-\gamma)^3/(2\gamma|\mathcal A|)η=(1−γ)3/(2γ∣A∣), from any initial policy, for every ρ∈Δ(S)\rho\in\Delta(\mathcal S)ρ∈Δ(S) and ϵ>0\epsilon>0ϵ>0,

min⁡t≤T{V⋆(ρ)−V(t)(ρ)}≤ϵwheneverT>64γ∣S∣∣A∣(1−γ)6ϵ2∥dρπ⋆μ∥∞2.\min_{t\le T}\big\{V^\star(\rho)-V^{(t)}(\rho)\big\}\le\epsilon\qquad\text{whenever}\qquad T>\frac{64\gamma|\mathcal S||\mathcal A|}{(1-\gamma)^6\epsilon^2}\Big\|\frac{d^{\pi^\star}_\rho}{\mu}\Big\|_\infty^2 .t≤Tmin​{V⋆(ρ)−V(t)(ρ)}≤ϵwheneverT>(1−γ)6ϵ264γ∣S∣∣A∣​​μdρπ⋆​​​∞2​.

Milestones, in the order of the proof

  1. Lemma 3.2 (performance difference): Vπ(s0)−Vπ′(s0)=11−γEs∼ds0πEa∼π(⋅∣s)[Aπ′(s,a)]V^\pi(s_0)-V^{\pi'}(s_0)=\frac1{1-\gamma}\mathbb E_{s\sim d^\pi_{s_0}}\mathbb E_{a\sim\pi(\cdot\mid s)}[A^{\pi'}(s,a)]Vπ(s0​)−Vπ′(s0​)=1−γ1​Es∼ds0​π​​Ea∼π(⋅∣s)​[Aπ′(s,a)].
  2. (7), the gradient of the direct parameterization: ∂Vπ(μ)/∂π(a∣s)=11−γdμπ(s)Qπ(s,a)\partial V^\pi(\mu)/\partial\pi(a\mid s)=\frac1{1-\gamma}d^\pi_\mu(s)Q^\pi(s,a)∂Vπ(μ)/∂π(a∣s)=1−γ1​dμπ​(s)Qπ(s,a).
  3. Lemma 4.1 (gradient domination): V⋆(ρ)−Vπ(ρ)≤11−γ∥dρπ⋆/μ∥∞max⁡πˉ(πˉ−π)⊤∇πVπ(μ)V^\star(\rho)-V^\pi(\rho)\le\frac1{1-\gamma}\|d^{\pi^\star}_\rho/\mu\|_\infty\max_{\bar\pi}(\bar\pi-\pi)^\top\nabla_\pi V^\pi(\mu)V⋆(ρ)−Vπ(ρ)≤1−γ1​∥dρπ⋆​/μ∥∞​maxπˉ​(πˉ−π)⊤∇π​Vπ(μ), together with the sharper form with dμπd^\pi_\mudμπ​ in place of (1−γ)μ(1-\gamma)\mu(1−γ)μ.
  4. Lemma D.3 (smoothness): ∥∇πVπ(s0)−∇πVπ′(s0)∥2≤2γ∣A∣(1−γ)3∥π−π′∥2\|\nabla_\pi V^\pi(s_0)-\nabla_\pi V^{\pi'}(s_0)\|_2\le\frac{2\gamma|\mathcal A|}{(1-\gamma)^3}\|\pi-\pi'\|_2∥∇π​Vπ(s0​)−∇π​Vπ′(s0​)∥2​≤(1−γ)32γ∣A∣​∥π−π′∥2​.
  5. Theorem E.1(3) (Beck 2017, Theorem 10.15): projected gradient descent with step 1/β1/\beta1/β on a β\betaβ-smooth function over a closed convex set has min⁡t<T∥Gη(xt)∥≤2β(f(x0)−f(x∗))/T\min_{t<T}\|G^\eta(x_t)\|\le\sqrt{2\beta(f(x_0)-f(x^*))}/\sqrt Tmint<T​∥Gη(xt​)∥≤2β(f(x0​)−f(x∗))​/T​, with GηG^\etaGη the gradient mapping.
  6. Proposition B.1: a gradient mapping of norm at most ϵ\epsilonϵ at π\piπ makes the next iterate π+\pi^+π+ ϵ(ηβ+1)\epsilon(\eta\beta+1)ϵ(ηβ+1)-stationary over feasible unit directions.

Significance

The theorem shows that, for the simplest constrained parameterization, a first-order method finds a globally optimal policy at a polynomial rate in spite of non-concavity. The guarantee holds for every performance distribution ρ\rhoρ at once, and it isolates the role of exploration in a single quantity, the mismatch coefficient; Section 4.3 of the paper shows that without a well-covering μ\muμ gradient methods can need exponentially many steps. Lemma 4.1 and the smoothness bound are reused across the rest of the paper, and the performance difference lemma underlies essentially all of its analyses.

All results here are proved in the paper, with Theorem E.1 and Theorem E.2 cited from Beck (2017) and Ghadimi–Lan (2016). None of them has a machine-checked proof on the platform. A formal development provides a verified link between the policy gradient expression of the direct parameterization and a standard nonconvex projected-gradient rate, and a reusable formal library of discounted visitation distributions, the performance difference identity, and projected gradient methods on Euclidean spaces.

Difficulty

The obvious argument, "projected gradient ascent converges to a stationary point, and stationary points are optimal", fails on both counts as stated. Stationary points of Vπ(μ)V^\pi(\mu)Vπ(μ) need not be optimal when μ\muμ does not cover the relevant states; the quantitative replacement is gradient domination, which only controls suboptimality through the mismatch coefficient. The convergence rate itself requires smoothness of the value as a function of the policy table, which is a bound on second derivatives of a matrix inverse (I−γPπ)−1(I-\gamma P_\pi)^{-1}(I−γPπ​)−1 with the dependence (1−γ)−3(1-\gamma)^{-3}(1−γ)−3 and the factor ∣A∣|\mathcal A|∣A∣ made explicit. Finally, the near-stationarity delivered by the gradient-mapping rate is at the next iterate, not the current one, which is why the conclusion is over t∈{0,…,T}t\in\{0,\dots,T\}t∈{0,…,T}. On the Lean side, the value is an infinite series in the policy entries, so its differentiability and the exact gradient formula have to be established for a function defined on the whole parameter space.

Formalization scope

Policies are parameter vectors in EuclideanSpace ℝ (S × A), so norms are ℓ2\ell_2ℓ2​ and Mathlib's gradient is ∇π\nabla_\pi∇π​; the objective π↦Vπ(μ)\pi\mapsto V^\pi(\mu)π↦Vπ(μ) is defined on the whole space and is only ever evaluated, with its gradient, at policies. The MDP layer (transition kernels, policies, VπV^\piVπ, QπQ^\piQπ, occupation distributions, optimal policies) is the published FoundationsML.ReinforcementLearning library; VπV^\piVπ is the unnormalized discounted sum. The projection is any map satisfying the nearest-point property. The optimal policy is a hypothesis IsOptimalPolicy (optimal from every state), not a supremum over all functions.

Conventions committed to:

  • The mismatch coefficient is any constant DDD with dρπ⋆(s)≤Dμ(s)d^{\pi^\star}_\rho(s)\le D\mu(s)dρπ⋆​(s)≤Dμ(s) for all sss; this avoids Lean's x/0=0x/0=0x/0=0 and is equivalent to the page's statement when the coefficient is finite.
  • γ>0\gamma>0γ>0 and ϵ>0\epsilon>0ϵ>0 are explicit hypotheses of the goal (the step size divides by γ\gammaγ, the threshold by ϵ\epsilonϵ).
  • The goal concludes ∃ t≤T\exists\,t\le T∃t≤T. The printed min⁡t<T\min_{t<T}mint<T​ fails at T=1T=1T=1 (one state, two actions with rewards 111 and 000, γ=0.001\gamma=0.001γ=0.001, ϵ=1/2\epsilon=1/2ϵ=1/2, initial policy on the bad action); the proof on p. 50 establishes the range 0≤t≤T0\le t\le T0≤t≤T.
  • Proposition B.1 bounds the directions feasible at π+\pi^+π+, as its proof does; Theorem E.1 assumes smoothness on CCC only and uses the radicand 2β(f(x0)−f(x∗))2\beta(f(x_0)-f(x^*))2β(f(x0​)−f(x∗)) of Beck and of p. 49.

A formalization that defines the update through formula (7), that assumes gradient domination or smoothness as hypotheses of the goal, or that takes the gradient off the simplex where the value series may diverge, would trivialize the goal; the goal mentions none of these, and (7) is a milestone theorem.

A complete development needs: summability and differentiability of the value series near the simplex; the performance difference lemma; the Euclidean projection inequality on a closed convex set; the descent lemma for functions smooth on a convex set; and the gradient-mapping argument. The projection and gradient-mapping results are independent of reinforcement learning and reusable. Contributions to any milestone are welcome, as are proofs of Theorem E.2 (Ghadimi–Lan) as a stepping stone to Proposition B.1.

Selected references

  • A. Agarwal, S. M. Kakade, J. D. Lee, G. Mahajan, On the Theory of Policy Gradient Methods: Optimality, Approximation, and Distribution Shift, JMLR 22(98), 2021; arXiv:1908.00261v5. https://arxiv.org/abs/1908.00261
  • A. Beck, First-Order Methods in Optimization, MOS-SIAM Series on Optimization, SIAM, 2017. https://doi.org/10.1137/1.9781611974997
  • S. Ghadimi, G. Lan, Accelerated gradient methods for nonconvex nonlinear and stochastic programming, Mathematical Programming 156, 2016. https://doi.org/10.1007/s10107-015-0871-8
  • S. Kakade, J. Langford, Approximately optimal approximate reinforcement learning, ICML 2002. https://dl.acm.org/doi/10.5555/645531.656005
  • B. Scherrer, M. Geist, Local Policy Search in a Convex Space and Conservative Policy Iteration as Boosted Policy Search, ECML PKDD 2014, pp. 35–50 (arXiv version: https://arxiv.org/abs/1306.1520)
  • R. J. Williams, Simple statistical gradient-following algorithms for connectionist reinforcement learning, Machine Learning 8, 1992. https://doi.org/10.1007/BF00992696
18 thms1 active userReviewed
Operations ResearchProbabilityStochastic Systems·Captain: mikedeng1

Learning in Structured MDPs with Convex Cost Functions: Improved Regret Bounds for Inventory Management: Base-Stock Values from Any Two Starting States Differ by at Most 36 max(h,p)LxResearch Paper

Motivation

The lost-sales inventory problem with lead times is a basic model of operations management. A retailer reviews one product's stock each period and places an order that arrives LLL periods later. Demand that cannot be met from stock on hand is lost, and the retailer pays a holding cost hhh per unit left on the shelf and a penalty ppp per unit of lost demand. The optimal policy depends on the whole pipeline of outstanding orders, so the state space grows with LLL, and the problem is computationally hard for long lead times. Simple base-stock (order-up-to) policies are therefore the standard heuristic, and Huh, Janakiraman, Muckstadt and Rusmevichientong (Management Science 2009) showed they are asymptotically optimal as the lost-sales penalty grows.

Agrawal and Jia (arXiv:1905.04337) study the learning version, in which the demand distribution is unknown and only sales, not demands, are observed. They give an algorithm whose regret against the best base-stock policy is O~(LT)\tilde O(L\sqrt T)O~(LT​), improving the earlier bound of Zhang, Chao and Shi, which grows exponentially in LLL. The improvement rests on one structural fact: started from two different states, the base-stock system accumulates expected costs that differ by an amount linear in LLL and independent of the horizon. That fact, Lemma 2.5 of the paper, is the goal of this mission.

Setting

Fix a lead time L≥0L\ge 0L≥0 and a base-stock level xxx. A state is a vector s=(s(0),s(1),…,s(L))\mathbf s=(s(0),s(1),\dots,s(L))s=(s(0),s(1),…,s(L)) of real numbers. Its entry s(0)s(0)s(0) is the on-hand inventory after the current period's arrival, and s(1),…,s(L)s(1),\dots,s(L)s(1),…,s(L) are the outstanding orders, s(L)s(L)s(L) the most recent. Under a base-stock policy with level xxx the states lie in

Sx={s:s(i)≥0 for all i, ∑i=0Ls(i)=x}.\mathcal S^x=\Big\{\mathbf s : s(i)\ge 0\ \text{for all } i,\ \sum_{i=0}^{L}s(i)=x\Big\}.Sx={s:s(i)≥0 for all i, i=0∑L​s(i)=x}.

In each period ttt a demand dt≥0d_t\ge 0dt​≥0 is drawn, independently across periods, from a distribution FFF on [0,∞)[0,\infty)[0,∞). The sales are yt=min⁡{st(0),dt}y_t=\min\{s_t(0),d_t\}yt​=min{st​(0),dt​} and the on-hand inventory is It=st(0)I_t=s_t(0)It​=st​(0). The policy reorders exactly what was sold, so for L≥1L\ge 1L≥1 the next state is

st+1=(st(0)−yt+st(1), st(2), …, st(L), yt),\mathbf s_{t+1}=\big(s_t(0)-y_t+s_t(1),\ s_t(2),\ \dots,\ s_t(L),\ y_t\big),st+1​=(st​(0)−yt​+st​(1), st​(2), …, st​(L), yt​),

and for L=0L=0L=0 the state (x)(x)(x) never changes. The pseudo-cost of period ttt is Ctx=h(st(0)−yt)−p ytC^x_t=h(s_t(0)-y_t)-p\,y_tCtx​=h(st​(0)−yt​)−pyt​, and the value over horizon TTT from the start state s\mathbf ss is

VTx(s)=E[∑t=1TCtx ∣ s1=s].V^x_T(\mathbf s)=\mathbb E\Big[\sum_{t=1}^{T}C^x_t\ \Big|\ \mathbf s_1=\mathbf s\Big].VTx​(s)=E[t=1∑T​Ctx​ ​ s1​=s].

Along a demand path, nTx(s)=∑t=1Tytn^x_T(\mathbf s)=\sum_{t=1}^T y_tnTx​(s)=∑t=1T​yt​ is the total sales and mTx(s)=∑t=1TItm^x_T(\mathbf s)=\sum_{t=1}^T I_tmTx​(s)=∑t=1T​It​ the total on-hand inventory.

States are compared by the order of Definition B.1: s′⪰s\mathbf s'\succeq\mathbf ss′⪰s if s′−s=δ\mathbf s'-\mathbf s=\deltas′−s=δ with ∑iδi=0\sum_i\delta_i=0∑i​δi​=0 and some 0≤k≤L−10\le k\le L-10≤k≤L−1 such that δi≥0\delta_i\ge 0δi​≥0 for i≤ki\le ki≤k and δi≤0\delta_i\le 0δi​≤0 for i>ki>ki>k. Thus s′\mathbf s's′ holds the same total, shifted toward the shelf. The state s^=(x,0,…,0)\hat{\mathbf s}=(x,0,\dots,0)s^=(x,0,…,0) dominates every state of Sx\mathcal S^xSx.

Formalization targets

Goal: Lemma 2.5 (p. 8)

For every xxx, every horizon TTT, all costs h,p≥0h,p\ge 0h,p≥0, every demand law FFF and all s,s′∈Sx\mathbf s,\mathbf s'\in\mathcal S^xs,s′∈Sx,

VTx(s)−VTx(s′)≤36max⁡(h,p) L x.V^x_T(\mathbf s)-V^x_T(\mathbf s')\le 36\max(h,p)\,L\,x .VTx​(s)−VTx​(s′)≤36max(h,p)Lx.

The constant is the paper's printed one. The proof's last display gives 18(h+p)Lx18(h+p)Lx18(h+p)Lx, a stronger bound, which is deliberately not the goal.

Milestones (Appendix B and the proof of Lemma 2.5)

All of the following hold for L≥1L\ge 1L≥1, along any single demand path that drives both chains:

  1. Lemma B.2 (p. 20). If s1′⪰s1\mathbf s'_1\succeq\mathbf s_1s1′​⪰s1​ then for t≤L+1t\le L+1t≤L+1 the cumulative sales satisfy Yt′−Yt≤max⁡0≤k≤t−1(δ0+⋯+δk)Y'_t-Y_t\le\max_{0\le k\le t-1}(\delta_0+\dots+\delta_k)Yt′​−Yt​≤max0≤k≤t−1​(δ0​+⋯+δk​).
  2. Lemma B.3 (p. 20). If moreover It′≥ItI'_t\ge I_tIt′​≥It​ for t=1,…,L+1t=1,\dots,L+1t=1,…,L+1, then nT(sL+1′)=nT(sL+1)n_T(\mathbf s'_{L+1})=n_T(\mathbf s_{L+1})nT​(sL+1′​)=nT​(sL+1​) for every TTT.
  3. Lemma B.5 (p. 21). At the successive first crossing times σi,τi\sigma_i,\tau_iσi​,τi​ of Definition B.4, the state order alternates: sσi′⪰sσi\mathbf s'_{\sigma_i}\succeq\mathbf s_{\sigma_i}sσi​′​⪰sσi​​ and sτi′⪯sτi\mathbf s'_{\tau_i}\preceq\mathbf s_{\tau_i}sτi​′​⪯sτi​​ whenever these times exist.
  4. Lemma B.6 (p. 21). If s′⪰s\mathbf s'\succeq\mathbf ss′⪰s in Sx\mathcal S^xSx then ∣nTx(s′)−nTx(s)∣≤3x|n^x_T(\mathbf s')-n^x_T(\mathbf s)|\le 3x∣nTx​(s′)−nTx​(s)∣≤3x.
  5. Lemma B.7 (p. 23). If s′⪰s\mathbf s'\succeq\mathbf ss′⪰s in Sx\mathcal S^xSx then ∣mTx(s)−mTx(s′)∣≤6Lx|m^x_T(\mathbf s)-m^x_T(\mathbf s')|\le 6Lx∣mTx​(s)−mTx​(s′)∣≤6Lx.
  6. Proof of Lemma 2.5 (p. 9). s^⪰s\hat{\mathbf s}\succeq\mathbf ss^⪰s for every s∈Sx\mathbf s\in\mathcal S^xs∈Sx.
  7. Proof of Lemma 2.5 (p. 9). ∣VTx(s)−VTx(s^)∣≤9(h+p)Lx|V^x_T(\mathbf s)-V^x_T(\hat{\mathbf s})|\le 9(h+p)Lx∣VTx​(s)−VTx​(s^)∣≤9(h+p)Lx.

Significance

Lemma 2.5 bounds the dependence of the base-stock chain's finite-horizon cost on its starting state, uniformly in the horizon. In the paper it yields three consequences: the long-run average cost (the loss) of a base-stock policy does not depend on the initial state (Lemma 2.6), the bias of the chain is bounded by 36max⁡(h,p)Lx36\max(h,p)Lx36max(h,p)Lx (Lemma 2.8), and finite-horizon average costs concentrate around the loss (Lemma 2.10). These feed the regret bound of Theorem 1.3. The lemma is also a statement about the base-stock lost-sales system alone, without any learning, so it is of independent interest for coupling arguments on lost-sales chains.

The paper's proof is complete on paper, but nothing in it has a machine-checked proof. Neither the lost-sales base-stock chain with lead times started from an arbitrary pipeline state nor any of the coupling lemmas of Appendix B is formalized elsewhere. This mission produces a checked proof of the goal and of the pathwise comparison lemmas. Theorem 1.3 is not posed: its supporting lemmas rely on limits whose existence the paper settles only by an informal discretization (Remark 4).

Difficulty

The obvious argument couples the two chains on a common demand path and waits until they coalesce. Coalescence is guaranteed only after LLL consecutive periods of zero demand, an event of probability exponentially small in LLL, so this argument gives a bound exponential in LLL. That is the bound of earlier work.

The linear bound needs a finer pathwise accounting. The two coupled chains do not stay ordered: the one that starts with more inventory on the shelf sells more at first, then runs short and sells less. The order ⪰\succeq⪰ between the two states alternates along a sequence of times, and the sales gained in one phase must be shown to be lost again in the next, so that the cumulative difference stays bounded by a constant multiple of xxx for every horizon. Turning this alternation into a bound requires tracking how the pipeline vectors evolve between alternation times, including the boundary cases in which the chains coalesce or the horizon ends inside a phase.

Formalization scope

All declarations live in the namespace LostSalesLearning.ValueGap. A state is a function Fin (L + 1) → ℝ, a demand path is a function ℕ → ℝ≥0, and time is 0-based: traj s d 0 is the paper's s1\mathbf s_1s1​, traj s d t is st+1\mathbf s_{t+1}st+1​, and ∑t=1T\sum_{t=1}^T∑t=1T​ is a sum over Finset.range T. The demand law FFF is a probability measure on ℝ≥0, and the demand path has the product law Measure.infinitePi (fun _ => F). The value is the expectation of the summed pseudo-costs, which equals Definition 2.4 by the tower property and is the form used in the paper's proof.

Committed conventions:

  • The costs satisfy h≥0h\ge 0h≥0 and p≥0p\ge 0p≥0, the reading of "per unit holding cost and per unit lost sales penalty".
  • No assumption is placed on FFF. The paper's assumptions F(0)>0F(0)>0F(0)>0 and bounded demand belong to other results.
  • The goal holds for every L≥0L\ge 0L≥0; the Appendix B milestones carry L≥1L\ge 1L≥1, as Appendix B does.
  • The order ⪰\succeq⪰ is Definition B.1 verbatim, with the equal-sum clause and the split index k≤L−1k\le L-1k≤L−1.
  • The pathwise milestones quantify over every demand path and drive both chains with the same path.
  • The first crossing times of Definition B.4 are represented by alternationTimes; an absent next crossing is none.

Two trivializing formalizations are ruled out. A comparison of the two values on different or fixed demand paths would be a different statement: the goal compares two expectations under the same law, and each pathwise milestone uses one common path. A Bochner integral of a non-integrable function would be 000. The integrand here is measurable and bounded by T(h+p)xT(h+p)xT(h+p)x on Sx\mathcal S^xSx, so the values are genuine expectations.

A complete development needs the elementary dynamics of the chain, including invariance of Sx\mathcal S^xSx and the shift of trajectories, which is reusable for other lost-sales models. It also needs the alternation times of Definition B.4, and measurability of the trajectory in the demand path. Proofs of individual milestones, alternative proofs of the goal, and sharper constants as separate statements are all welcome.

Selected references

  • S. Agrawal and R. Jia, Learning in Structured MDPs with Convex Cost Functions: Improved Regret Bounds for Inventory Management, arXiv:1905.04337v1, 2019. https://arxiv.org/abs/1905.04337
  • W. T. Huh, G. Janakiraman, J. A. Muckstadt and P. Rusmevichientong, Asymptotic Optimality of Order-Up-To Policies in Lost Sales Inventory Systems, Management Science 55(3), 2009. https://doi.org/10.1287/mnsc.1080.0945
  • H. Zhang, X. Chao and C. Shi, Closing the Gap: A Learning Algorithm for Lost-Sales Inventory Systems with Lead Times, Management Science 66(5), 2020. https://doi.org/10.1287/mnsc.2019.3288
  • M. L. Puterman, Markov Decision Processes: Discrete Stochastic Dynamic Programming, Wiley, 1994. https://doi.org/10.1002/9780470316887
9 thms1 active userReviewed
Machine LearningOperations ResearchOptimization·Captain: mikedeng1

Twice Regularized MDPs and the Equivalence Between Robustness and Regularization 2: The Greedy Policy of the R2 Optimal Value Is the Unique Optimal R2 PolicyResearch Paper

Motivation

A robust Markov decision process (robust MDP) evaluates a policy against the worst transition kernel and reward in an uncertainty set around a nominal model (P0,r0)(P_0, r_0)(P0​,r0​). It is the standard model for planning when the dynamics are estimated from data (Iyengar 2005; Nilim and El Ghaoui 2005; Wiesemann, Kuhn and Rustem 2013). Its Bellman update contains an inner optimization over the uncertainty set at every state, which makes robust planning costly when the sets are not (s,a)(s,a)(s,a)-rectangular.

Derman, Geist and Mannor (arXiv:2110.06267, NeurIPS 2021) show that, for sss-rectangular ball uncertainty sets, this inner optimization can be replaced by an explicit penalty that depends both on the policy and on the value function. The resulting twice regularized (R²) MDPs have Bellman operators with no inner optimization over models. The first mission of this series formalizes the robust–regularized equivalence (Theorem 4.1 of the paper). This mission formalizes Section 5: the R² Bellman operators are monotone and contracting under a bound on the transition radius, and the greedy policy of the R² optimal value is optimal.

Setting

Let S\mathcal SS and A\mathcal AA be finite nonempty sets of states and actions, γ∈(0,1)\gamma\in(0,1)γ∈(0,1) a discount factor, P0(s′∣s,a)P_0(s'\mid s,a)P0​(s′∣s,a) a transition kernel and r0(s,a)r_0(s,a)r0​(s,a) a reward. A policy π∈ΔAS\pi\in\Delta_{\mathcal A}^{\mathcal S}π∈ΔAS​ assigns to each state a probability distribution πs\pi_sπs​ on A\mathcal AA. For v∈RSv\in\mathbb R^{\mathcal S}v∈RS write qs(a)=r0(s,a)+γ∑s′P0(s′∣s,a)v(s′)q_s(a)=r_0(s,a)+\gamma\sum_{s'}P_0(s'\mid s,a)v(s')qs​(a)=r0​(s,a)+γ∑s′​P0​(s′∣s,a)v(s′) and

[T(P0,r0)πv](s)=∑aπs(a) qs(a).[T^\pi_{(P_0,r_0)}v](s)=\sum_a\pi_s(a)\,q_s(a).[T(P0​,r0​)π​v](s)=a∑​πs​(a)qs​(a).

All norms ∥⋅∥\|\cdot\|∥⋅∥ below are ℓ2\ell_2ℓ2​-norms, ∥a∥=(∑za(z)2)1/2\|a\|=\big(\sum_z a(z)^2\big)^{1/2}∥a∥=(∑z​a(z)2)1/2; ∥⋅∥∞\|\cdot\|_\infty∥⋅∥∞​ is the sup norm.

Fix nonnegative radii αsr,αsP\alpha^r_s,\alpha^P_sαsr​,αsP​ for each state. The R² regularizer is Ωv,R2(πs)=∥πs∥ (αsr+αsPγ∥v∥)\Omega_{v,\mathrm R^2}(\pi_s)=\|\pi_s\|\,(\alpha^r_s+\alpha^P_s\gamma\|v\|)Ωv,R2​(πs​)=∥πs​∥(αsr​+αsP​γ∥v∥), and the R² Bellman operators are

[Tπ,R2v](s)=[T(P0,r0)πv](s)−Ωv,R2(πs),[T∗,R2v](s)=max⁡π∈ΔAS[Tπ,R2v](s).[T^{\pi,\mathrm R^2}v](s)=[T^\pi_{(P_0,r_0)}v](s)-\Omega_{v,\mathrm R^2}(\pi_s),\qquad [T^{*,\mathrm R^2}v](s)=\max_{\pi\in\Delta^{\mathcal S}_{\mathcal A}}[T^{\pi,\mathrm R^2}v](s).[Tπ,R2v](s)=[T(P0​,r0​)π​v](s)−Ωv,R2​(πs​),[T∗,R2v](s)=π∈ΔAS​max​[Tπ,R2v](s).

A policy π\piπ is greedy for vvv when Tπ,R2v=T∗,R2vT^{\pi,\mathrm R^2}v=T^{*,\mathrm R^2}vTπ,R2v=T∗,R2v.

Assumption 5.1 (bounded radius). For each sss there is ϵs>0\epsilon_s>0ϵs​>0 with

αsP≤min⁡(1−γ−ϵsγ∣S∣ ; min⁡u∈R+A,∥u∥=1, w∈R+S,∥w∥=1 ∑a,s′u(a)P0(s′∣s,a)w(s′)),\alpha^P_s\le\min\Big(\frac{1-\gamma-\epsilon_s}{\gamma\sqrt{|\mathcal S|}}\ ;\ \min_{u\in\mathbb R^{\mathcal A}_+,\|u\|=1,\ w\in\mathbb R^{\mathcal S}_+,\|w\|=1}\ \sum_{a,s'}u(a)P_0(s'\mid s,a)w(s')\Big),αsP​≤min(γ∣S∣​1−γ−ϵs​​ ; u∈R+A​,∥u∥=1, w∈R+S​,∥w∥=1min​ a,s′∑​u(a)P0​(s′∣s,a)w(s′)),

and ϵ∗=min⁡sϵs\epsilon_*=\min_s\epsilon_sϵ∗​=mins​ϵs​. The R² value function vπ,R2v^{\pi,\mathrm R^2}vπ,R2 of a policy and the R² optimal value v∗,R2v^{*,\mathrm R^2}v∗,R2 are the fixed points of Tπ,R2T^{\pi,\mathrm R^2}Tπ,R2 and T∗,R2T^{*,\mathrm R^2}T∗,R2.

Formalization targets

Goal: Theorem 5.1 (p. 8)

Under Assumption 5.1, T∗,R2T^{*,\mathrm R^2}T∗,R2 and every Tπ,R2T^{\pi,\mathrm R^2}Tπ,R2 have unique fixed points; a greedy policy π∗,R2\pi^{*,\mathrm R^2}π∗,R2 for v∗,R2v^{*,\mathrm R^2}v∗,R2 exists, and every such policy satisfies

vπ∗,R2,R2=v∗,R2 ≥ vπ,R2for all π∈ΔAS;v^{\pi^{*,\mathrm R^2},\mathrm R^2}=v^{*,\mathrm R^2}\ \ge\ v^{\pi,\mathrm R^2}\qquad\text{for all }\pi\in\Delta^{\mathcal S}_{\mathcal A};vπ∗,R2,R2=v∗,R2 ≥ vπ,R2for all π∈ΔAS​;

every optimal policy is greedy; and when αsr>0\alpha^r_s>0αsr​>0 for all sss the greedy policy is unique, hence the unique optimal R² policy.

Milestones

  1. Proposition 2.1 (p. 3): for Ω\OmegaΩ strongly convex on the simplex, Ω∗(y)=max⁡a∈Δ⟨a,y⟩−Ω(a)\Omega^*(y)=\max_{a\in\Delta}\langle a,y\rangle-\Omega(a)Ω∗(y)=maxa∈Δ​⟨a,y⟩−Ω(a) is differentiable with Lipschitz gradient equal to the unique maximizer, satisfies Ω∗(y+c1)=Ω∗(y)+c\Omega^*(y+c\mathbb 1)=\Omega^*(y)+cΩ∗(y+c1)=Ω∗(y)+c, and is non-decreasing.
  2. Proposition 5.1 (i) (p. 8): v1≤v2v_1\le v_2v1​≤v2​ implies Tπ,R2v1≤Tπ,R2v2T^{\pi,\mathrm R^2}v_1\le T^{\pi,\mathrm R^2}v_2Tπ,R2v1​≤Tπ,R2v2​ and T∗,R2v1≤T∗,R2v2T^{*,\mathrm R^2}v_1\le T^{*,\mathrm R^2}v_2T∗,R2v1​≤T∗,R2v2​.
  3. Proposition 5.1 (iii) (p. 8):
∥Tπ,R2v1−Tπ,R2v2∥∞≤(1−ϵ∗)∥v1−v2∥∞,∥T∗,R2v1−T∗,R2v2∥∞≤(1−ϵ∗)∥v1−v2∥∞.\|T^{\pi,\mathrm R^2}v_1-T^{\pi,\mathrm R^2}v_2\|_\infty\le(1-\epsilon_*)\|v_1-v_2\|_\infty,\qquad \|T^{*,\mathrm R^2}v_1-T^{*,\mathrm R^2}v_2\|_\infty\le(1-\epsilon_*)\|v_1-v_2\|_\infty.∥Tπ,R2v1​−Tπ,R2v2​∥∞​≤(1−ϵ∗​)∥v1​−v2​∥∞​,∥T∗,R2v1​−T∗,R2v2​∥∞​≤(1−ϵ∗​)∥v1​−v2​∥∞​.

Significance

Theorem 5.1 is the R² counterpart of the fundamental theorem of discounted dynamic programming: optimal R² values are achieved by stationary policies obtained by a single greedy step. Together with the contraction of Proposition 5.1 (iii) it justifies the R² modified policy iteration algorithm of the paper, whose greedy step is a projection onto the simplex rather than a robust max–min problem. Combined with the first mission of the series, which identifies the robust value of an sss-rectangular ball-constrained MDP with the optimum of an R²-regularized program, it gives a route to robust planning at the cost of regularized planning.

The results are proved in the paper (App. C), partly by reference to Geist, Scherrer and Pietquin (2019) for the optimality operator. No machine-checked proof of any of them exists; this mission produces the first. Prop. 2.1 is a general fact of convex analysis (Danskin-type smoothness of a conjugate on the simplex) that is reusable for any regularized MDP or entropy-regularized game.

Difficulty

The R² evaluation operator is not affine: the value regularizer −αsPγ∥πs∥ ∥v∥-\alpha^P_s\gamma\|\pi_s\|\,\|v\|−αsP​γ∥πs​∥∥v∥ is concave in vvv and decreases as ∥v∥\|v\|∥v∥ grows. Monotonicity therefore does not follow from the positivity of P0P_0P0​ as in the standard case; it requires the second bound of Assumption 5.1, which compares the ℓ2\ell_2ℓ2​ variation of ∥v∥\|v\|∥v∥ with the minimal nonnegative bilinear form of P0(⋅∣s,⋅)P_0(\cdot\mid s,\cdot)P0​(⋅∣s,⋅). Likewise the contraction modulus is not γ\gammaγ but 1−ϵ∗1-\epsilon_*1−ϵ∗​, because the regularizer is ∣S∣\sqrt{|\mathcal S|}∣S∣​-Lipschitz between the ℓ2\ell_2ℓ2​ and sup norms. The optimality step of the classical proof uses linearity of TπT^\piTπ when comparing values of policies; here only monotonicity and contraction are available. Uniqueness of the greedy policy rests on strict concavity on the simplex, which holds only when the regularization weight is positive.

Formalization scope

States and actions are finite nonempty types; transitions are arrays P₀ : S → A → S → ℝ with the published predicate IsTransitionKernel; value functions are S → ℝ with the pointwise order. The ℓ2\ell_2ℓ2​-norm is an explicit l2norm (Mathlib's norm on S → ℝ is the sup norm, used only for ∥⋅∥∞\|\cdot\|_\infty∥⋅∥∞​). T∗,R2v(s)T^{*,\mathrm R^2}v(s)T∗,R2v(s) is the real supremum over the simplex ΔA\Delta_{\mathcal A}ΔA​ (attained), and the inner minimum of Assumption 5.1 is the real infimum over nonnegative ℓ2\ell_2ℓ2​-unit vectors; the witnesses ϵs\epsilon_sϵs​ are explicit. Greedy policies are a predicate, never a function, and the R² value functions are not defined by choice: the goal asserts their existence and uniqueness and speaks about the fixed points.

Disclosed deviations from the page. Assumption 5.1 is a hypothesis of Theorem 5.1 (its proof assumes it). The uniqueness clause of Theorem 5.1 additionally assumes αsr>0\alpha^r_s>0αsr​>0 for all sss: with one state, two actions, zero reward and zero radii every policy is greedy and optimal. Proposition 2.1 assumes Ω\OmegaΩ continuous on the simplex, without which the maximum need not be attained, and strong convexity is Mathlib's StrongConvexOn for some modulus (norm-independent in finite dimension). Proposition 5.1 (ii) is false as printed and is not drafted: with one state, one action, P0=1P_0=1P0​=1, r0=0r_0=0r0​=0, γ=1/2\gamma=1/2γ=1/2, αr=0\alpha^r=0αr=0, αP=1/2\alpha^P=1/2αP=1/2, ϵ=1/4\epsilon=1/4ϵ=1/4, one has Tv=v/2−∣v∣/4Tv=v/2-|v|/4Tv=v/2−∣v∣/4, and v1=−1v_1=-1v1​=−1, c=1c=1c=1 give T(v1+c)=0>−1/4=Tv1+γcT(v_1+c)=0>-1/4=Tv_1+\gamma cT(v1​+c)=0>−1/4=Tv1​+γc. Remark 5.1, Algorithm 1 and the ℓp\ell_pℓp​ variant of App. C.1 are out of scope. The inner minimum of Assumption 5.1 is 000 whenever some P0(s′∣s,a)=0P_0(s'\mid s,a)=0P0​(s′∣s,a)=0, forcing αsP=0\alpha^P_s=0αsP​=0; this is the assumption as printed.

A formalization in which ∥⋅∥\|\cdot\|∥⋅∥ is the sup norm, the inner minimum ranges over all unit vectors (making the assumption unsatisfiable), or the value functions are postulated rather than shown to exist would be trivial or wrong; the drafted statements avoid all three. Contributions welcome: Prop. 2.1 as a general convex-analysis lemma, Banach fixed-point plumbing for S → ℝ with the sup norm, and the strict concavity of p↦⟨p,q⟩−c∥p∥p\mapsto\langle p,q\rangle-c\|p\|p↦⟨p,q⟩−c∥p∥ on the simplex.

Selected references

  • E. Derman, M. Geist, S. Mannor, Twice regularized MDPs and the equivalence between robustness and regularization, NeurIPS 2021. arXiv:2110.06267v1
  • M. Geist, B. Scherrer, O. Pietquin, A theory of regularized Markov decision processes, ICML 2019. arXiv:1901.11275
  • A. Nilim, L. El Ghaoui, Robust control of Markov decision processes with uncertain transition matrices, Operations Research 53(5), 2005. doi:10.1287/opre.1050.0216
  • G. N. Iyengar, Robust dynamic programming, Mathematics of Operations Research 30(2), 2005. doi:10.1287/moor.1040.0129
  • W. Wiesemann, D. Kuhn, B. Rustem, Robust Markov decision processes, Mathematics of Operations Research 38(1), 2013. doi:10.1287/moor.1120.0566
  • A. Mensch, M. Blondel, Differentiable dynamic programming for structured prediction and attention, ICML 2018. arXiv:1802.03676
9 thms1 active userReviewed
Markov ChainProbability·Captain: mikedeng1

Linear Least-Squares Algorithms for Temporal Difference Learning II: Probability-One Convergence of LS TD on Ergodic Markov ChainsResearch Paper

Motivation

Temporal-difference (TD) learning estimates the value function of a Markov chain — the expected discounted sum of future rewards from each state — from a single stream of observed transitions, without knowing the transition probabilities. With a linear function approximator the value of state xxx is represented as ϕx′θ\phi_x'\thetaϕx′​θ for a feature vector ϕx\phi_xϕx​ and a parameter θ\thetaθ. Classical TD(λ\lambdaλ) updates θ\thetaθ by stochastic approximation, and its behaviour depends on a step-size schedule that must be tuned.

Bradtke and Barto (Machine Learning 22, 1996) replaced the stochastic-approximation update by a least-squares solve: LS TD (Eq. (11)) recomputes θt\theta_tθt​ at every step as the instrumental-variable least-squares solution of the empirical consistency condition. The method, later generalized as LSTD(λ\lambdaλ) by Boyan (Machine Learning 49, 2002), is the basis of least-squares policy iteration and of the "LSTD" methods in standard reinforcement-learning texts (Sutton and Barto, Reinforcement Learning, 2nd ed., 2018, §9.8). Its appeal is that it has no step size; the question this mission formalizes is whether it nonetheless converges, with probability one, to the true parameter.

Timeline. Sutton (1988) introduced TD(λ\lambdaλ). Watkins and Dayan (1992) and Tsitsiklis (1994) proved probability-one convergence of tabular TD(0) and Q-learning. Bradtke and Barto (1996) proved probability-one convergence of LS TD on absorbing chains (Theorem 1) and on ergodic chains (Theorem 2). Tsitsiklis and Van Roy (IEEE TAC 42, 1997) proved convergence of linear TD(λ\lambdaλ) with general features on ergodic chains.

Setting

A finite Markov chain on a finite nonempty set XXX is a matrix PPP with P(x,y)≥0P(x,y)\ge0P(x,y)≥0 and ∑yP(x,y)=1\sum_yP(x,y)=1∑y​P(x,y)=1. A transition x→yx\to yx→y earns reward R(x,y)R(x,y)R(x,y); the expected reward out of xxx is rˉx=∑yP(x,y)R(x,y)\bar r_x=\sum_yP(x,y)R(x,y)rˉx​=∑y​P(x,y)R(x,y). For a discount factor γ\gammaγ the value function is

V(x)=E{∑k=0∞γkrk ∣ x0=x}=∑k=0∞γk(Pkrˉ)(x).V(x)=E\Big\{\sum_{k=0}^\infty\gamma^kr_k\ \Big|\ x_0=x\Big\}=\sum_{k=0}^\infty\gamma^k(P^k\bar r)(x).V(x)=E{k=0∑∞​γkrk​ ​ x0​=x}=k=0∑∞​γk(Pkrˉ)(x).

The chain is ergodic (Kemeny and Snell) if every state can be reached from every state: for all x,yx,yx,y there is nnn with Pn(x,y)>0P^n(x,y)>0Pn(x,y)>0. An invariant distribution is a probability vector π\piπ with πP=π\pi P=\piπP=π; write Π=diag⁡(π)\Pi=\operatorname{diag}(\pi)Π=diag(π).

Each state has a feature vector ϕx∈Rm\phi_x\in\mathbb R^mϕx​∈Rm; Φ\PhiΦ is the matrix with rows ϕx\phi_xϕx​. The true parameter θ∗\theta^*θ∗ is a vector with V(x)=ϕx′θ∗V(x)=\phi_x'\theta^*V(x)=ϕx′​θ∗ for all xxx.

The algorithm (Figure 3) starts at an arbitrary state x0x_0x0​, lets the chain move x0→x1→⋯x_0\to x_1\to\cdotsx0​→x1​→⋯, and after ttt transitions computes

θt=[1t∑kϕxk(ϕxk−γϕxk+1)′]−1[1t∑kϕxkR(xk,xk+1)],(11)\theta_t=\Big[\frac1t\sum_{k}\phi_{x_k}(\phi_{x_k}-\gamma\phi_{x_{k+1}})'\Big]^{-1}\Big[\frac1t\sum_k\phi_{x_k}R(x_k,x_{k+1})\Big],\tag{11}θt​=[t1​k∑​ϕxk​​(ϕxk​​−γϕxk+1​​)′]−1[t1​k∑​ϕxk​​R(xk​,xk+1​)],(11)

the sums running over the ttt transitions observed so far.

Formalization targets

Goal: Theorem 2 (p. 44)

If PPP is ergodic, (1) {ϕx}\{\phi_x\}{ϕx​} is linearly independent, (2) each ϕx\phi_xϕx​ has dimension ∣X∣|X|∣X∣, and (3) 0<γ<10<\gamma<10<γ<1, then θ∗\theta^*θ∗ is finite and, from any initial law,

θt⟶θ∗with probability 1.\theta_t\longrightarrow\theta^*\qquad\text{with probability }1 .θt​⟶θ∗with probability 1.

The goal leaves the chain, the rewards, the features and the initial law arbitrary.

Milestones, in the order of the proof

  1. Visit frequencies (Proof of Theorem 2, p. 45): an ergodic chain visits every state infinitely often and #{k<t:xk=x}/t→πx\#\{k<t:x_k=x\}/t\to\pi_x#{k<t:xk​=x}/t→πx​ almost surely.
  2. Invertibility (Proof of Theorem 2, p. 45): πx>0\pi_x>0πx​>0 for all xxx, and Φ′Π(I−γP)Φ\Phi'\Pi(I-\gamma P)\PhiΦ′Π(I−γP)Φ is invertible.
  3. The pathwise limit (Proof of Lemma 5, pp. 54–55): along any path whose transition frequencies converge to πxP(x,y)\pi_xP(x,y)πx​P(x,y), θt→[Φ′Π(I−γP)Φ]−1[Φ′Πrˉ]\theta_t\to[\Phi'\Pi(I-\gamma P)\Phi]^{-1}[\Phi'\Pi\bar r]θt​→[Φ′Π(I−γP)Φ]−1[Φ′Πrˉ].
  4. Lemma 5 (p. 43): for any chain, if almost surely every state is visited infinitely often and in proportion π\piπ, and Φ′Π(I−γP)Φ\Phi'\Pi(I-\gamma P)\PhiΦ′Π(I−γP)Φ is invertible, then θt→[Φ′Π(I−γP)Φ]−1[Φ′Πrˉ]\theta_t\to[\Phi'\Pi(I-\gamma P)\Phi]^{-1}[\Phi'\Pi\bar r]θt​→[Φ′Π(I−γP)Φ]−1[Φ′Πrˉ] almost surely.
  5. Eq. (12) (p. 44): the value series converges and rˉ=(I−γP)Φθ∗\bar r=(I-\gamma P)\Phi\theta^*rˉ=(I−γP)Φθ∗.

Significance

The result. Theorem 2 shows that an LS TD learner running on one long trajectory recovers the exact value function whenever the features can represent every function on the states, with no step-size schedule. It is the step-size-free counterpart of the tabular TD(0) convergence theorems and the starting point for the later analysis of LSTD with fewer features than states, where the limit is the TD fixed point [Φ′Π(I−γP)Φ]−1Φ′Πrˉ[\Phi'\Pi(I-\gamma P)\Phi]^{-1}\Phi'\Pi\bar r[Φ′Π(I−γP)Φ]−1Φ′Πrˉ rather than θ∗\theta^*θ∗. Lemma 5 is the general identification of that fixed point as the almost-sure limit of LSTD.

Formalizing it. The result is proved in the paper; it has not been machine-checked. A formalization adds three things the paper delegates: the strong law of large numbers for occupation times of a finite irreducible Markov chain, which the paper cites to Kemeny and Snell and which is not in Mathlib; the per-state transition frequencies used in the first sentence of the proof of Lemma 5; and the linear algebra of the limit. The first is reusable well beyond reinforcement learning.

Difficulty

The algebra is short once the empirical averages in (11) are known to converge. The difficulty is probabilistic: the averages are over a dependent sequence, so the ordinary strong law of large numbers does not apply. Two facts are needed: that the fraction of time in each state converges to πx\pi_xπx​ almost surely for any starting law, including periodic chains, where PnP^nPn itself does not converge; and that, among the visits to xxx, the fraction followed by a move to yyy converges to P(x,y)P(x,y)P(x,y), which needs the strong Markov property at successive visit times. Neither follows from convergence of the chain's distribution, and neither holds for a chain started at a fixed state without an argument that every state is reached.

Formalization scope

  • The model is a finite state type X with Fintype, DecidableEq, Nonempty, a row-stochastic matrix P : Matrix X X ℝ (structure Chain), rewards R : X → X → ℝ, features φ : X → Fin m → ℝ. Condition (2) is m = Fintype.card X; condition (1) is LinearIndependent ℝ φ.
  • "Ergodic" is read as irreducible, periodic chains allowed (Kemeny–Snell's aperiodic case is "regular"). "Arbitrary initial state" is read as every initial law ν\nuν, which contains every point mass.
  • The path is any process ZZZ on any probability space whose finite-dimensional distributions are ν(x0)P(x0,x1)⋯P(xn−1,xn)\nu(x_0)P(x_0,x_1)\cdots P(x_{n-1},x_n)ν(x0​)P(x0​,x1​)⋯P(xn−1​,xn​), with measurable events {Zt=x}\{Z_t=x\}{Zt​=x}.
  • VVV is the discounted series, never (I−γP)−1rˉ(I-\gamma P)^{-1}\bar r(I−γP)−1rˉ; Lean's tsum is 000 on a divergent series, so "θ* is finite" is stated as convergence of the series together with existence of θ∗\theta^*θ∗ with V=Φθ∗V=\Phi\theta^*V=Φθ∗. θ∗\theta^*θ∗ is existential, never defined as Lemma 5's limit.
  • (11) uses the transitions k=0,…,t−1k=0,\dots,t-1k=0,…,t−1 (the paper prints k=1,…,tk=1,\dots,tk=1,…,t with ϕt+1\phi_{t+1}ϕt+1​; an index shift), keeps the factors 1/t1/t1/t, and uses Lean's matrix inverse, which is 000 on a singular matrix: the paper notes θt\theta_tθt​ is undefined for small ttt, and finitely many junk values do not affect convergence. No εI\varepsilon IεI regularization, no pseudo-inverse.
  • θLSTD=lim⁡tθt\theta_{\rm LSTD}=\lim_t\theta_tθLSTD​=limt​θt​ is formalized as convergence of θt\theta_tθt​ (existence of the limit is part of the claim).
  • The convergence is almost sure. A formalization that assumes the visit frequencies converge in the goal, starts the chain from π\piπ, or weakens the conclusion to convergence in probability or along a subsequence is a different theorem.

Needed infrastructure: the strong law for occupation times of a finite irreducible chain under an arbitrary initial law (milestone 1), the strong Markov property at visit times, positivity of the invariant distribution of an irreducible chain, and the invertibility of I−γPI-\gamma PI−γP for ∣γ∣<1|\gamma|<1∣γ∣<1. Contributions to any of these, as standalone lemmas, are welcome.

Selected references

  • S. J. Bradtke and A. G. Barto, Linear Least-Squares Algorithms for Temporal Difference Learning, Machine Learning 22, 33–57, 1996. https://doi.org/10.1023/A:1018056104778
  • J. G. Kemeny and J. L. Snell, Finite Markov Chains, Springer, 1976.
  • J. A. Boyan, Technical Update: Least-Squares Temporal Difference Learning, Machine Learning 49, 233–246, 2002. https://doi.org/10.1023/A:1017936530646
  • J. N. Tsitsiklis and B. Van Roy, An Analysis of Temporal-Difference Learning with Function Approximation, IEEE Transactions on Automatic Control 42(5), 674–690, 1997. https://doi.org/10.1109/9.580874
  • R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018. http://incompleteideas.net/book/the-book-2nd.html
8 thms1 active userReviewed
Markov ChainProbability·Captain: mikedeng1

Linear Least-Squares Algorithms for Temporal Difference Learning I: Probability-One Convergence of Trial-Based LS TD on Absorbing Markov ChainsResearch Paper

Motivation

Temporal-difference learning estimates the value of a policy from observed state transitions and rewards. In a finite Markov decision process, fixing a policy produces a Markov chain, so policy evaluation becomes the task of estimating the expected return from each state. Bradtke and Barto's 1996 paper introduced a least-squares temporal-difference method, LS TD, that uses each observed transition in a linear system instead of selecting a learning-rate schedule. Their Theorem 1 states probability-one convergence for trials that end at absorbing states under explicit conditions on state access, rewards, and features. This mission formalizes that result and the statements the authors use to reach it. Bradtke and Barto, 1996.

The result matters for episodic policy evaluation: a learner may collect many short trajectories, each begun from a prescribed start distribution, and update the same estimate as data accumulate. The theorem identifies conditions under which the limit is the true value parameter even when the discount factor is one. That endpoint is useful for undiscounted tasks ending in an absorbing goal state; it also makes the convergence claim more delicate than the standard discounted case. The paper proves the result mathematically. The Lean statements in this mission are targets for machine-checked proofs, not claims of proofs already present in Mathlib. Bradtke and Barto, Theorem 1, pp. 43–44.

Setting

Let XXX be a finite, nonempty set of states. After a policy is fixed, P(x,y)P(x,y)P(x,y) is the probability of a transition from xxx to yyy, so each row of PPP is nonnegative and sums to one. A transition earns a deterministic real reward R(x,y)R(x,y)R(x,y). A state is absorbing when P(x,x)=1P(x,x)=1P(x,x)=1; let T\mathcal TT be the absorbing states and N=X∖T\mathcal N=X\setminus\mathcal TN=X∖T the others. The chain is absorbing when some absorbing state can be reached with positive probability from every state. A start distribution SSS gives the state at the beginning of each trial. No state is inaccessible when every state can be reached from the positive support of SSS.

For a discount γ\gammaγ, the expected immediate reward is rˉ(x)=∑yP(x,y)R(x,y)\bar r(x)=\sum_yP(x,y)R(x,y)rˉ(x)=∑y​P(x,y)R(x,y). The true value function is defined by the expected return

V(x)=∑k=0∞γk(Pkrˉ)(x).V(x)=\sum_{k=0}^{\infty}\gamma^k(P^k\bar r)(x).V(x)=k=0∑∞​γk(Pkrˉ)(x).

A feature vector ϕx∈Rm\phi_x\in\mathbb R^mϕx​∈Rm represents state xxx. The matrix Φ\PhiΦ has row xxx equal to ϕx⊤\phi_x^\topϕx⊤​. The target parameter θ∗\theta^*θ∗ is a vector for which V(x)=ϕx⊤θ∗V(x)=\phi_x^\top\theta^*V(x)=ϕx⊤​θ∗ at every state; it is something the theorem must establish, not an input chosen by a formula. Equation (11) forms an LS TD estimate θn\theta_nθn​ from the observed feature differences and rewards. Bradtke and Barto, §2, Table 1, Eq. (11).

Figure 2 collects trials. Each starts from SSS, follows PPP while the current state is non-absorbing, and ends upon entry into T\mathcal TT. The next trial starts with a fresh draw from SSS. The estimator includes transitions taken within trials; a draw that starts the next trial is not an observed transition for Eq. (11). Bradtke and Barto, Figure 2, p. 42.

Formalization targets

Theorem 1: convergence of trial-based LS TD

If every state is accessible from SSS, rewards between absorbing states vanish, the feature vectors on N\mathcal NN are linearly independent, features on T\mathcal TT are zero, m=∣N∣m=|\mathcal N|m=∣N∣, and 0≤γ≤10\le\gamma\le10≤γ≤1, then the expected-return series converges and there is a parameter θ∗\theta^*θ∗ satisfying

V(x)=ϕx⊤θ∗(x∈X),θn⟶θ∗with probability one.V(x)=\phi_x^\top\theta^*\quad(x\in X),\qquad \theta_n\longrightarrow\theta^*\quad\text{with probability one}.V(x)=ϕx⊤​θ∗(x∈X),θn​⟶θ∗with probability one.

The theorem keeps the paper's endpoint γ=1\gamma=1γ=1. The return series' convergence is explicit because a real infinite sum in Lean has a default value when it diverges. Bradtke and Barto, Theorem 1, p. 43.

Supporting targets

The milestone list follows the statements used in the paper: almost-sure visits and departure proportions for the trials; invertibility of the non-absorbing block of I−γPI-\gamma PI−γP; invertibility of Φ⊤Π(I−γP)Φ\Phi^\top\Pi(I-\gamma P)\PhiΦ⊤Π(I−γP)Φ for positive non-absorbing weights; Lemma 5's probability-one limit [Φ⊤Π(I−γP)Φ]−1Φ⊤Πrˉ[\Phi^\top\Pi(I-\gamma P)\Phi]^{-1}\Phi^\top\Pi\bar r[Φ⊤Π(I−γP)Φ]−1Φ⊤Πrˉ; and Eq. (12), rˉ=(I−γP)Φθ∗\bar r=(I-\gamma P)\Phi\theta^*rˉ=(I−γP)Φθ∗, together with finiteness of the true parameter. Here Π=diag⁡(π)\Pi=\operatorname{diag}(\pi)Π=diag(π). Bradtke and Barto, Lemma 5, p. 43; Proof of Theorem 1, p. 44.

Significance

Theorem 1 identifies the target of the asymptotic LS TD estimate: the value function defined from rewards, rather than merely a vector satisfying a sampled linear system. It covers an undiscounted absorbing chain, where a general fixed-point equation for values would fail to determine the values of absorbing states. The zero-reward and zero-feature conditions determine that boundary correctly. The result also explains the dimension condition: one independent feature vector for each non-absorbing state permits exact representation of the return. Bradtke and Barto, pp. 43–44.

A complete formal development would connect finite-state stochastic-process laws, visit frequencies, matrix limits, and the return-defined value function in one checked statement. The reusable parts include a finite row-stochastic chain model, a path-law description of restarts, a filtered least-squares estimator, and results about transient blocks of stochastic matrices. The paper's mathematical proof exists; this mission asks for formal proofs of its Lean targets. It also leaves room for alternative proofs and sharper, separately stated variants without weakening Theorem 1.

Difficulty

Ordinary matrix convergence cannot be applied until the observed transition frequencies are known to converge and the limiting matrix is invertible. A trial has random length, and the process resets after absorption, so a sequence indexed by all restart-process steps does not have the same raw state proportions as a count indexed by trials. The proof must account for both clocks while retaining the in-trial data of Eq. (11). At γ=1\gamma=1γ=1, a direct geometric-series argument for the value function is unavailable; its finiteness depends on absorption and the reward convention. The matrix I−γPI-\gamma PI−γP itself is singular at the undiscounted endpoint because of absorbing states, while its non-absorbing block is the relevant invertible matrix. Bradtke and Barto, Proof of Theorem 1, p. 44.

Formalization scope

The Lean state type is finite and nonempty. The paper evaluates one fixed policy, so PPP is a real row-stochastic matrix and RRR is a deterministic real reward on transitions; there is no action type in the formal statement. Absorbing states are exactly those with P(x,x)=1P(x,x)=1P(x,x)=1, and “absorbing chain” means that an absorbing state is reachable from every state. The paper does not define “inaccessible”; the formalization reads it as unreachable from the positive support of SSS. The state space carries the discrete measurable structure. Theorem 1's restart process and Lemma 5's ordinary Markov chain are each constrained by their finite-dimensional cylinder probabilities, not by assumed transition frequencies.

The feature space is Rm\mathbb R^mRm, and mmm equals the cardinality of the subtype N\mathcal NN. LS TD uses only departures from N\mathcal NN. Index nnn counts restart-process steps, so the estimate repeats at a restart draw; the paper counts in-trial transitions. These indices have the same asymptotic estimate when transitions continue. The 1/t1/t1/t factors in Eq. (11) cancel, and early singular inverses take Lean's total-inverse default. The value function is the return series, and the goal explicitly asserts its summability. The true parameter is existential, never defined by the formula whose convergence the theorem is meant to prove.

The paper defines πx\pi_xπx​ for absorbing chains as expected departures from xxx per trial. The Theorem 1 visit-frequency milestone normalizes by restart-process steps, which rescales all weights by one positive common factor; Lemma 5's matrix expression is invariant under that rescaling. Lemma 5 itself counts every ordinary-chain transition and carries the paper's “any Markov chain” scope. The milestone on invertibility allows arbitrary weights at absorbing states because their feature rows are zero. These conventions are recorded with each Lean item. Contributions toward the path-law frequency theorem, transient-matrix invertibility, return-series summability, and the matrix limit are all within scope. A vacuous path law or a value function defined from the desired linear equation would not establish the stated goal.

Selected references

  • S. J. Bradtke and A. G. Barto, Linear Least-Squares Algorithms for Temporal Difference Learning, Machine Learning 22, 33–57 (1996). DOI: 10.1023/A:1018056104778.
8 thms1 active userReviewed
Dynamic ProgrammingMachine LearningMarkov Chain·Captain: mikedeng1

Reinforcement Learning: An Introduction IX: Average Reward and the Futility of Discounting in Continuing ProblemsTextbook

Motivation

Reinforcement learning formulates control as maximizing reward accumulated over time. For continuing tasks, where interaction never terminates, the standard textbook objective is the discounted return with a discount rate γ<1\gamma < 1γ<1. Chapter 10 of Sutton and Barto, Reinforcement Learning: An Introduction (2nd ed., MIT Press, 2018), argues that once values are approximated by a function of features rather than stored per state, this objective is the wrong one, and it proposes the average-reward setting, long standard in dynamic programming (Puterman, Markov Decision Processes, 1994), as its replacement.

The central piece of evidence is a short calculation printed in the box The Futility of Discounting in Continuing Problems (p. 254). If one tries to rescue discounting by averaging discounted values over the states the policy actually visits, the resulting objective is a constant multiple of the average reward, so the discount rate has no effect on which policy is preferred. This mission formalizes that calculation together with the definitions of §10.3 that it rests on, and the two exercises of §10.3 that illustrate the differential value (10.13).

Setting

A finite MDP has finite state set S\mathcal SS, finite action set A\mathcal AA, finite reward set R⊂R\mathcal R \subset \mathbb RR⊂R and dynamics p(s′,r∣s,a)p(s', r \mid s, a)p(s′,r∣s,a), a probability distribution over next state and reward for each state–action pair (3.2)–(3.3). Write p(s′∣s,a)=∑rp(s′,r∣s,a)p(s' \mid s, a) = \sum_r p(s', r\mid s,a)p(s′∣s,a)=∑r​p(s′,r∣s,a). A policy π(a∣s)\pi(a \mid s)π(a∣s) is a distribution over actions for each state. It turns the MDP into a Markov chain with transition matrix Pπ(s,s′)=∑aπ(a∣s) p(s′∣s,a)P_\pi(s, s') = \sum_a \pi(a\mid s)\, p(s'\mid s,a)Pπ​(s,s′)=∑a​π(a∣s)p(s′∣s,a) and expected one-step reward rπ(s)=∑aπ(a∣s)∑s′,rp(s′,r∣s,a) rr_\pi(s) = \sum_a \pi(a\mid s)\sum_{s',r} p(s',r\mid s,a)\, rrπ​(s)=∑a​π(a∣s)∑s′,r​p(s′,r∣s,a)r, so that Pr⁡{St=s∣S0=s0}=Pπt(s0,s)\Pr\{S_t = s\mid S_0=s_0\} = P_\pi^t(s_0,s)Pr{St​=s∣S0​=s0​}=Pπt​(s0​,s) and E[Rt+1∣S0=s0]=(Pπtrπ)(s0)\mathbb E[R_{t+1}\mid S_0 = s_0] = (P_\pi^t r_\pi)(s_0)E[Rt+1​∣S0​=s0​]=(Pπt​rπ​)(s0​).

A stationary distribution of π\piπ is a probability vector μ\muμ on S\mathcal SS with

∑sμ(s)∑aπ(a∣s) p(s′∣s,a)=μ(s′)for all s′(10.8).\sum_s \mu(s)\sum_a \pi(a\mid s)\,p(s'\mid s,a) = \mu(s') \quad\text{for all } s' \qquad (10.8).s∑​μ(s)a∑​π(a∣s)p(s′∣s,a)=μ(s′)for all s′(10.8).

The average reward of π\piπ is

r(π)=lim⁡h→∞1h∑t=1hE[Rt∣S0,A0:t−1∼π](10.6),r(π)=∑sμπ(s)∑aπ(a∣s)∑s′,rp(s′,r∣s,a) r(10.7),r(\pi) = \lim_{h\to\infty}\frac1h\sum_{t=1}^h \mathbb E[R_t\mid S_0, A_{0:t-1}\sim\pi] \quad (10.6), \qquad r(\pi) = \sum_s \mu_\pi(s)\sum_a\pi(a\mid s)\sum_{s',r}p(s',r\mid s,a)\,r \quad (10.7),r(π)=h→∞lim​h1​t=1∑h​E[Rt​∣S0​,A0:t−1​∼π](10.6),r(π)=s∑​μπ​(s)a∑​π(a∣s)s′,r∑​p(s′,r∣s,a)r(10.7),

the second form holding when the steady-state distribution μπ(s)=lim⁡t→∞Pr⁡{St=s}\mu_\pi(s) = \lim_{t\to\infty}\Pr\{S_t = s\}μπ​(s)=limt→∞​Pr{St​=s} exists and does not depend on S0S_0S0​ (an ergodic MDP).

The discounted value function is vπγ(s)=Eπ[∑k≥0γkRt+k+1∣St=s]v^\gamma_\pi(s) = \mathbb E_\pi\big[\sum_{k\ge0}\gamma^k R_{t+k+1}\mid S_t = s\big]vπγ​(s)=Eπ​[∑k≥0​γkRt+k+1​∣St​=s], and the objective of the box is

J(π)=∑sμπ(s) vπγ(s).J(\pi) = \sum_s \mu_\pi(s)\, v^\gamma_\pi(s).J(π)=s∑​μπ​(s)vπγ​(s).

Finally, the differential value of a state (10.13) is vπ(s)=lim⁡γ→1lim⁡h→∞∑t=0hγt(Eπ[Rt+1∣S0=s]−r(π))v_\pi(s) = \lim_{\gamma\to1}\lim_{h\to\infty}\sum_{t=0}^h\gamma^t\big(\mathbb E_\pi[R_{t+1}\mid S_0 = s] - r(\pi)\big)vπ​(s)=limγ→1​limh→∞​∑t=0h​γt(Eπ​[Rt+1​∣S0​=s]−r(π)).

Formalization targets

Goal: the futility of discounting

For 0≤γ<10 \le \gamma < 10≤γ<1, every policy π\piπ and every stationary distribution μπ\mu_\piμπ​ of π\piπ,

J(π)=∑sμπ(s) vπγ(s)=11−γ r(π),J(\pi) = \sum_s \mu_\pi(s)\, v^\gamma_\pi(s) = \frac{1}{1-\gamma}\, r(\pi),J(π)=s∑​μπ​(s)vπγ​(s)=1−γ1​r(π),

and therefore, for a fixed γ\gammaγ and policies π,π′\pi, \pi'π,π′ with their own stationary distributions,

J(π)≤J(π′)  ⟺  r(π)≤r(π′).J(\pi) \le J(\pi') \iff r(\pi) \le r(\pi').J(π)≤J(π′)⟺r(π)≤r(π′).

Milestones

  1. (10.6)–(10.7): if Pr⁡{St=s∣S0=s0}→μ(s)\Pr\{S_t = s\mid S_0 = s_0\}\to\mu(s)Pr{St​=s∣S0​=s0​}→μ(s) for all s0,ss_0, ss0​,s, then from every start both lim⁡tE[Rt]\lim_t \mathbb E[R_t]limt​E[Rt​] and the Cesàro limit (10.6) exist and equal the μ\muμ-sum.
  2. (10.8): such a limit μ\muμ is a stationary distribution.
  3. (3.14): vπγv^\gamma_\pivπγ​ satisfies the Bellman equation, the step "(Bellman Eq.)" of the box.
  4. Offset invariance (p. 250): the differential Bellman equations for vπ,qπ,v∗,q∗v_\pi, q_\pi, v_*, q_*vπ​,qπ​,v∗​,q∗​ and the differential TD errors (10.10)–(10.11) are unchanged when all values are shifted by a constant.
  5. Exercise 10.6: for expected rewards 1,0,1,0,…1,0,1,0,\dots1,0,1,0,… from A\mathsf AA and 0,1,0,1,…0,1,0,1,\dots0,1,0,1,… from B\mathsf BB, the average reward is 12\tfrac1221​, the limit (10.7) and a steady-state distribution do not exist, and vπ(A)=14v_\pi(\mathsf A) = \tfrac14vπ​(A)=41​, vπ(B)=−14v_\pi(\mathsf B) = -\tfrac14vπ​(B)=−41​.
  6. Exercise 10.7: in the three-state ring with reward +1+1+1 on arrival in A\mathsf AA, the average reward is 13\tfrac1331​ and the differential values are v(A)=−13v(\mathsf A) = -\tfrac13v(A)=−31​, v(B)=0v(\mathsf B) = 0v(B)=0, v(C)=13v(\mathsf C) = \tfrac13v(C)=31​.

The book prints no answers to Exercises 10.6 and 10.7; the values above were computed for this mission.

Significance

The result itself. The identity shows that the discount rate cannot enter the ranking of policies once performance is measured over the on-policy state distribution: any objective of the form "discounted value averaged over where the policy goes" is the average reward up to a positive factor. The book draws from it the conclusion (p. 253) that γ\gammaγ changes from a problem parameter to a solution-method parameter, and that discounting algorithms with function approximation, which do not optimize this averaged objective, are not guaranteed to optimize average reward either. Milestones 1–2 connect the closed form of r(π)r(\pi)r(π) to its definition as a long-run rate; milestone 4 explains why differential methods determine values only up to an offset; the exercises show that (10.13) gives finite differential values in periodic chains where the differential return (10.9) has no limit.

Formalizing it. The results are classical and elementary, but the book's derivation is informal: it takes the Bellman equation for granted, sums a geometric series of equalities, and leaves implicit which distribution μπ\mu_\piμπ​ is meant and when (10.6) and (10.7) agree. This mission fixes each of those points: vπγv^\gamma_\pivπγ​ is defined from returns, the identity is proved for every stationary distribution, and the ergodicity hypothesis is made the exact limit condition the text states. No machine-checked version of these statements is known to exist; the platform's Markov-chain and average-reward libraries (Puterman's unichain optimality equation, Doeblin convergence) state different results.

Difficulty

The box reads as a chain of equalities, and each individual step is short. The one step that is not algebra is "(Bellman Eq.)": with vπγv^\gamma_\pivπγ​ defined as an expected discounted return, the Bellman equation requires exchanging an infinite sum with a finite expectation and shifting the index of a convergent series, which needs absolute convergence for γ<1\gamma < 1γ<1 and bounded rewards. The final line, which unrolls J=r+γJJ = r + \gamma JJ=r+γJ into a geometric series, is only valid because JJJ is finite. For milestone 1, the book's hypothesis is the existence of an S0S_0S0​-independent limiting distribution, not irreducibility or aperiodicity; replacing it by either would change the statement. The exercises require the iterated limit in (10.13) to be computed explicitly: the inner limit is a periodic series summed in closed form, and the outer limit is a removable singularity at γ=1\gamma = 1γ=1.

Formalization scope

The Lean development lives in the namespace SuttonBartoRL.AverageReward. States, actions and rewards are finite; one action set serves every state; the dynamics are the four-argument p(s′,r∣s,a)p(s', r\mid s,a)p(s′,r∣s,a). Probabilities Pr⁡{St=s∣S0=s0}\Pr\{S_t = s\mid S_0 = s_0\}Pr{St​=s∣S0​=s0​} and E[Rt+1∣S0=s0]\mathbb E[R_{t+1}\mid S_0=s_0]E[Rt+1​∣S0​=s0​] are computed from the matrix powers PπtP_\pi^tPπt​. The discounted value vπγv^\gamma_\pivπγ​ is a real series ∑kγk(Pπkrπ)(s)\sum_k \gamma^k (P_\pi^k r_\pi)(s)∑k​γk(Pπk​rπ​)(s); it is not defined as the solution of the Bellman equation, since that definition would reduce the goal to three lines of algebra. The average reward and JJJ take the state distribution μ\muμ as an explicit argument; the goal holds for every stationary distribution of π\piπ and assumes no ergodicity, which is all the box uses (under ergodicity μπ\mu_\piμπ​ is the unique one). γ=0\gamma = 0γ=0 is allowed. The differential value (10.13) is stated as the existence of both limits with the given value, with γ→1\gamma \to 1γ→1 from below, so no default value of a nonexistent limit can make a statement true. The optimality equations use a nonempty action set. The sum ∑t=1hE[Rt]\sum_{t=1}^h \mathbb E[R_t]∑t=1h​E[Rt​] of (10.6) is written with the index shifted to t=0,…,h−1t = 0,\dots,h-1t=0,…,h−1.

The Lean does not formalize a trajectory probability space: expectations and probabilities are the matrix expressions above, which is what they equal for a Markov chain. The Exercise 10.6 hypothesis constrains expected rewards of one policy, which is all the exercise uses.

Reusable parts: the finite-MDP layer duplicates the one drafted for the other missions of this series and is expected to be merged with it; the Cesàro and stationary-distribution lemmas of milestones 1–2 are general facts about finite Markov chains. Contributions welcome: proofs of any milestone, and a proof of the goal from milestone 3.

Selected references

  • R. S. Sutton, A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018, ISBN 9780262039246, §§10.3–10.4, pp. 249–254. http://incompleteideas.net/book/the-book-2nd.html
  • M. L. Puterman, Markov Decision Processes: Discrete Stochastic Dynamic Programming, Wiley, 1994. https://doi.org/10.1002/9780470316887
  • S. Mahadevan, "Average reward reinforcement learning: foundations, algorithms, and empirical results", Machine Learning 22, 159–195, 1996. https://doi.org/10.1007/BF00114727
11 thms1 active userReviewed
PreviousPage 1 of 2Next

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me