Prove2Me
Navigate
DiscoverFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

Markov Chain

77 missions · 29 completed

Missions

Open48Completed29All77
🏆Completed
ProbabilityStochastic Systems·Captain: Shuze Chen

Markov Chains and Mixing Times II: The Convergence TheoremTextbook

Motivation

The first mission of this series established that an irreducible finite Markov chain has a unique stationary distribution π\piπ. The present mission, covering Chapters 3–4 of Levin–Peres–Wilmer, Markov Chains and Mixing Times (AMS, 2009), answers the two questions that make that fact useful. First, the inverse problem of sampling: given a target distribution π\piπ — uniform over proper colorings, a Gibbs measure, a posterior — how does one build a chain whose stationary distribution is π\piπ? The Metropolis and Glauber constructions of Chapter 3 are the universal answers, and they are the engine of Markov chain Monte Carlo across statistical physics, Bayesian statistics, and approximate counting. Second, the convergence question: in what sense, and how fast, does an irreducible aperiodic chain approach π\piπ? Chapter 4 introduces the total variation distance, proves the Convergence Theorem — geometric convergence to stationarity — and defines the mixing time, the parameter the entire remainder of the book estimates.

Setting

All chains live on a finite state space VVV and are presented by row-stochastic matrices, with the definitions of Mission I. The total variation distance between distributions μ\muμ and ν\nuν is

∥μ−ν∥TV=max⁡A⊆V ∣μ(A)−ν(A)∣,\|\mu-\nu\|_{\mathrm{TV}} = \max_{A\subseteq V}\,|\mu(A)-\nu(A)|,∥μ−ν∥TV​=A⊆Vmax​∣μ(A)−ν(A)∣,

the maximal discrepancy over events. A coupling of μ\muμ and ν\nuν is a distribution on V×VV\times VV×V whose marginals are μ\muμ and ν\nuν. For a chain PPP with stationary π\piπ one sets

d(t)=max⁡x∥Pt(x,⋅)−π∥TV,dˉ(t)=max⁡x,y∥Pt(x,⋅)−Pt(y,⋅)∥TV,d(t)=\max_x \|P^t(x,\cdot)-\pi\|_{\mathrm{TV}},\qquad \bar d(t)=\max_{x,y}\|P^t(x,\cdot)-P^t(y,\cdot)\|_{\mathrm{TV}},d(t)=xmax​∥Pt(x,⋅)−π∥TV​,dˉ(t)=x,ymax​∥Pt(x,⋅)−Pt(y,⋅)∥TV​,

and the mixing time is tmix(ε)=min⁡{t:d(t)≤ε}t_{\mathrm{mix}}(\varepsilon)=\min\{t : d(t)\le\varepsilon\}tmix​(ε)=min{t:d(t)≤ε}, with tmix=tmix(1/4)t_{\mathrm{mix}}=t_{\mathrm{mix}}(1/4)tmix​=tmix​(1/4).

The Metropolis chain for a target π\piπ and a symmetric proposal chain Ψ\PsiΨ accepts a proposed move x→yx\to yx→y with probability 1∧π(y)/π(x)1\wedge \pi(y)/\pi(x)1∧π(y)/π(x); a general (not necessarily symmetric) base chain is handled by the ratio (π(y)Ψ(y,x))/(π(x)Ψ(x,y))∧1\bigl(\pi(y)\Psi(y,x)\bigr)/\bigl(\pi(x)\Psi(x,y)\bigr)\wedge 1(π(y)Ψ(y,x))/(π(x)Ψ(x,y))∧1. The Glauber dynamics for a distribution π\piπ on configurations VsitesV^{\text{sites}}Vsites picks a uniform site and re-samples its value from π\piπ conditioned on the rest.

Formalization targets

Goal

P irreducible and aperiodic  ⟹  ∃ α∈(0,1), C>0:d(t)≤Cαt.\text{$P$ irreducible and aperiodic}\;\Longrightarrow\;\exists\,\alpha\in(0,1),\ C>0:\quad d(t)\le C\alpha^{t}.P irreducible and aperiodic⟹∃α∈(0,1), C>0:d(t)≤Cαt.

This is Theorem 4.9, the Convergence Theorem. It asserts only the geometric shape of convergence, leaving all quantitative rates to later missions, which is why it is the goal.

Milestones

The milestones are the chapter's working parts: stationarity and reversibility of the Metropolis chain for symmetric and general base chains (§3.2, Exercise 3.1), stationarity and reversibility of the Glauber dynamics (§3.3, Exercise 3.2); the three characterizations of total variation distance — the half-ℓ1\ell^1ℓ1 formula (Proposition 4.2 with Remark 4.3), the supremum over [−1,1][-1,1][−1,1]-bounded test functions (Proposition 4.5), and the coupling characterization with an optimal coupling attaining it (Proposition 4.7 with Remark 4.8); the comparison d≤dˉ≤2dd\le\bar d\le 2dd≤dˉ≤2d (Lemma 4.11) and submultiplicativity dˉ(s+t)≤dˉ(s)dˉ(t)\bar d(s+t)\le\bar d(s)\bar d(t)dˉ(s+t)≤dˉ(s)dˉ(t) (Lemma 4.12); the standard mixing-time consequences d(ℓ tmix(ε))≤(2ε)ℓd(\ell\, t_{\mathrm{mix}}(\varepsilon))\le(2\varepsilon)^\elld(ℓtmix​(ε))≤(2ε)ℓ and tmix(ε)≤⌈log⁡2ε−1⌉ tmixt_{\mathrm{mix}}(\varepsilon)\le\lceil\log_2\varepsilon^{-1}\rceil\, t_{\mathrm{mix}}tmix​(ε)≤⌈log2​ε−1⌉tmix​ (§4.5); and the equality of distance to stationarity for a group walk and its inverse walk (Lemma 4.13 and Corollary 4.14).

Significance

The results. The Convergence Theorem is the qualitative foundation on which quantitative mixing theory stands: it guarantees that tmix(ε)t_{\mathrm{mix}}(\varepsilon)tmix​(ε) is finite, so every bound in Missions III–XIII is a bound on a well-defined quantity. The TV characterizations are used constantly — the coupling characterization is the engine of Mission III, the half-ℓ1\ell^1ℓ1 formula of every explicit computation. The Metropolis and Glauber stationarity results justify the chains analyzed in Missions III (colorings, hardcore), VIII (path coupling) and IX (Ising). Submultiplicativity of dˉ\bar ddˉ is what makes tmixt_{\mathrm{mix}}tmix​ a meaningful single number.

Formalizing them. None of this exists in Mathlib: there is no total variation distance for finitely supported distributions, no coupling theory, no mixing time, no MCMC correctness statement. The definition layer published here (TV distance, ddd, dˉ\bar ddˉ, tmixt_{\mathrm{mix}}tmix​, couplings, Metropolis, Glauber) is imported by every subsequent mission of the series.

Difficulty

The tempting proof of Theorem 4.9 via spectral decomposition fails twice: it needs reversibility, which the theorem does not assume, and spectral machinery that arrives only in Mission VII. The book's proof is the Doeblin decomposition: by Proposition 1.7 some power satisfies Pr(x,y)≥δπ(y)P^r(x,y)\ge\delta\pi(y)Pr(x,y)≥δπ(y), so Pr=(1−θ)Π+θQP^r=(1-\theta)\Pi+\theta QPr=(1−θ)Π+θQ with Π\PiΠ the rank-one matrix of rows π\piπ, and induction gives Prk=(1−θk)Π+θkQkP^{rk}=(1-\theta^k)\Pi+\theta^kQ^kPrk=(1−θk)Π+θkQk. The formal work is matrix algebra with careful bookkeeping of the remainder chain QQQ, plus the monotonicity of ddd needed to interpolate between multiples of rrr. For Proposition 4.7 the delicate half is constructing the optimal coupling: mass μ∧ν\mu\wedge\nuμ∧ν on the diagonal and the normalized product of the positive parts off it, with the degenerate case μ=ν\mu=\nuμ=ν handled separately. The Glauber stationarity statement must be phrased with care because configurations outside the support of π\piπ have junk rows; the formalization asserts stochasticity only at supported configurations, and detailed balance globally.

Formalization scope

Total variation distance is defined as the supremum over events, ⨆A ∣μ(A)−ν(A)∣\bigsqcup_{A}\,|\mu(A)-\nu(A)|⨆A​∣μ(A)−ν(A)∣ over Finset V, exactly as in (4.1); the half-ℓ1\ell^1ℓ1 formula is a milestone, not the definition. The mixing time is sInf of the set {t:d(t)≤ε}\{t : d(t)\le\varepsilon\}{t:d(t)≤ε} in N\mathbb NN (junk value 000 if empty — impossible under the goal theorem). Couplings are distributions on the product with prescribed marginals; no probability-space machinery is used. The mixing-time inequalities are stated with the integer-rounding slack made explicit (e.g. ⌈log⁡2ε−1⌉\lceil\log_2\varepsilon^{-1}\rceil⌈log2​ε−1⌉ via Nat.ceil of a real logarithm) so that no statement is true only "up to rounding". The Metropolis definitions use total real division, so the hypotheses require π>0\pi>0π>0 pointwise; this matches the book, which divides by π(x)\pi(x)π(x) throughout.

Welcome contributions beyond the milestones: simp lemmas for tvDist, monotonicity of ddd and dˉ\bar ddˉ in ttt, and triangle-inequality infrastructure — all reused by Missions III–XIII.

Selected references

  • D. A. Levin, Y. Peres, E. L. Wilmer, Markov Chains and Mixing Times, American Mathematical Society, 2009. https://documents.epfl.ch/groups/i/ip/ipg/www/2013-2014/Random_Walks/markovmixing.pdf
  • N. Metropolis, A. Rosenbluth, M. Rosenbluth, A. Teller, E. Teller, Equation of state calculations by fast computing machines, J. Chem. Phys. 21 (1953). https://doi.org/10.1063/1.1699114
  • W. Doeblin, Exposé de la théorie des chaînes simples constantes de Markov à un nombre fini d'états, Rev. Math. Union Interbalkan. 2 (1938).
14 thms2 active usersReviewed
Dynamic ProgrammingOperations Research·Captain: mikedeng1

Discrete-Time Controlled Markov Processes with Average Cost Criterion: A Survey 2: Uniformly Bounded Mean Return Times Make the Differential Discounted Values Uniformly BoundedResearch Paper

Motivation

Controlled Markov processes with the average cost criterion model systems that run indefinitely, such as queues, inventories, maintenance and communication networks, where only the long-run cost per unit time matters. The standard route to an optimal stationary policy goes through the average cost optimality equation (ACOE). The ACOE is usually obtained by the vanishing discount method: solve the discounted problem for each discount factor β<1\beta<1β<1 and let β→1\beta\to1β→1. The method works only when the differences of discounted values stay bounded as β→1\beta\to1β→1. Conditions that guarantee this are therefore central in the survey of Arapostathis, Borkar, Fernández-Gaucherand, Ghosh and Marcus (SIAM J. Control Optim. 31 (1993), §5).

This mission formalizes one such condition, due to Ross: if the mean return time to a fixed state is bounded uniformly over all stationary policies and initial states, the differential discounted value functions are bounded uniformly in the discount factor and the state.

Timeline. Derman (Management Sci. 9 (1962); survey reference [38]) and Derman–Veinott (Ann. Math. Statist. 38 (1967); survey reference [43]) introduced recurrence conditions of this kind for countable-state processes. Ross (Ann. Math. Statist. 39 (1968), survey reference [147]; Introduction to Stochastic Dynamic Programming, 1983, survey reference [150]) showed, for bounded costs, that under a Derman–Veinott type recurrence condition hβh_\betahβ​ is bounded uniformly in β\betaβ, and obtained a bounded solution of the ACOE by letting β↑1\beta\uparrow1β↑1 (survey, pp. 291 and 301). Later work replaced the condition with weaker ones (survey Assumptions 5.1–5.3) and with Sennott's conditions (survey Theorem 5.9).

Setting

The state space is S={0,1,2,… }S=\{0,1,2,\dots\}S={0,1,2,…}. In each state iii, an action aaa is chosen from a nonempty compact set U(i)U(i)U(i) of a metric space AAA. The one-stage cost c(i,a)c(i,a)c(i,a) is nonnegative and the next state is drawn from the transition law P(⋅∣i,a)P(\cdot\mid i,a)P(⋅∣i,a). For fixed i,ji,ji,j, the maps a↦c(i,a)a\mapsto c(i,a)a↦c(i,a) and a↦P(j∣i,a)a\mapsto P(j\mid i,a)a↦P(j∣i,a) are continuous on U(i)U(i)U(i). A policy π∈Π\pi\in\Piπ∈Π chooses the action at time ttt at random, given the whole history, and must choose from U(Xt)U(X_t)U(Xt​). A stationary deterministic policy f∈ΠSDf\in\Pi_{SD}f∈ΠSD​ is a map f:S→Af:S\to Af:S→A with f(i)∈U(i)f(i)\in U(i)f(i)∈U(i). PiπP^\pi_iPiπ​ and EiπE^\pi_iEiπ​ denote the law and the expectation of the controlled process (Xt,At)(X_t,A_t)(Xt​,At​) started at iii.

For a discount factor β∈(0,1)\beta\in(0,1)β∈(0,1), the discounted cost and the optimal discounted cost are

Jβ(i,π)=Eiπ[∑t=0∞βtc(Xt,At)],Jβ∗(i)=inf⁡π∈ΠJβ(i,π).J_\beta(i,\pi)=E^\pi_i\Big[\sum_{t=0}^\infty\beta^t c(X_t,A_t)\Big],\qquad J^*_\beta(i)=\inf_{\pi\in\Pi}J_\beta(i,\pi).Jβ​(i,π)=Eiπ​[t=0∑∞​βtc(Xt​,At​)],Jβ∗​(i)=π∈Πinf​Jβ​(i,π).

A policy f∈ΠSDf\in\Pi_{SD}f∈ΠSD​ is β\betaβ-discount optimal if Jβ(i,f)=Jβ∗(i)J_\beta(i,f)=J^*_\beta(i)Jβ​(i,f)=Jβ∗​(i) for all iii. The differential discounted value function is

hβ(i)=Jβ∗(i)−Jβ∗(0),h_\beta(i)=J^*_\beta(i)-J^*_\beta(0),hβ​(i)=Jβ∗​(i)−Jβ∗​(0),

measured relative to the fixed state 000. The return time to 000 is

τ=min⁡{t≥1: Xt=0},\tau=\min\{t\ge1:\ X_t=0\},τ=min{t≥1: Xt​=0},

with τ=∞\tau=\inftyτ=∞ if the process never returns. Throughout, as in §5.1 of the survey, the cost is bounded: c(i,a)≤Mc(i,a)\le Mc(i,a)≤M on admissible pairs.

Formalization targets

Goal: Theorem 5.3

If there is a constant K>0K>0K>0 with

Eif[τ]<Kfor all f∈ΠSD, i∈S,(5.7)E^f_i[\tau]<K\qquad\text{for all } f\in\Pi_{SD},\ i\in S, \tag{5.7}Eif​[τ]<Kfor all f∈ΠSD​, i∈S,(5.7)

then there is a constant BBB such that

∣hβ(i)∣≤Bfor all β∈(0,1), i∈S.|h_\beta(i)|\le B\qquad\text{for all }\beta\in(0,1),\ i\in S.∣hβ​(i)∣≤Bfor all β∈(0,1), i∈S.

This is the theorem as printed: it asserts only uniform boundedness and fixes no constant.

Milestones

  1. Theorem 2.1 (iii): for every β∈(0,1)\beta\in(0,1)β∈(0,1), a β\betaβ-discount optimal fβ∈ΠSDf_\beta\in\Pi_{SD}fβ​∈ΠSD​ exists.
  2. (5.8): for such an fβf_\betafβ​, Jβ∗(i)≤M Eifβ[τ]+Jβ∗(0) Eifβ[βτ]J^*_\beta(i)\le M\,E^{f_\beta}_i[\tau]+J^*_\beta(0)\,E^{f_\beta}_i[\beta^\tau]Jβ∗​(i)≤MEifβ​​[τ]+Jβ∗​(0)Eifβ​​[βτ].
  3. (5.9): Jβ∗(i)−βJβ∗(0)≤MKJ^*_\beta(i)-\beta J^*_\beta(0)\le MKJβ∗​(i)−βJβ∗​(0)≤MK.
  4. Jensen step: Jβ∗(i)≥Jβ∗(0) Eifβ[βτ]≥Jβ∗(0) βKJ^*_\beta(i)\ge J^*_\beta(0)\,E^{f_\beta}_i[\beta^\tau]\ge J^*_\beta(0)\,\beta^KJβ∗​(i)≥Jβ∗​(0)Eifβ​​[βτ]≥Jβ∗​(0)βK.
  5. (5.10): Jβ∗(0)−Jβ∗(i)≤(1−βK)Jβ∗(0)≤(1−βK)M1−β≤MKJ^*_\beta(0)-J^*_\beta(i)\le(1-\beta^K)J^*_\beta(0)\le(1-\beta^K)\frac{M}{1-\beta}\le MKJβ∗​(0)−Jβ∗​(i)≤(1−βK)Jβ∗​(0)≤(1−βK)1−βM​≤MK.
  6. Explicit bound: ∣hβ(i)∣≤MK|h_\beta(i)|\le MK∣hβ​(i)∣≤MK. This is stronger than the goal and is the constant the survey's argument yields.

Significance

The result. Theorem 5.3 verifies the hypothesis of the vanishing discount theorem (Theorem 5.2 of the survey) from a condition on the uncontrolled dynamics of stationary policies. Theorem 5.2 then gives a bounded solution (ρ,h)(\rho,h)(ρ,h) of the ACOE, an average optimal stationary policy, and the limit lim⁡β→1(1−β)Jβ∗(i)=ρ\lim_{\beta\to1}(1-\beta)J^*_\beta(i)=\rholimβ→1​(1−β)Jβ∗​(i)=ρ. Mean return times can often be estimated directly, for instance through Foster–Lyapunov drift arguments on queues, which makes (5.7) checkable in applications. The explicit bound MKMKMK also controls the span of the relative value function.

Formalizing it. The result is proved in the literature; to our knowledge it has not been machine-checked. A formal proof needs discounted dynamic programming on a countable state space with compact action sets, the existence of optimal stationary policies (Theorem 2.1 (iii), which the survey cites without proof), and the strong Markov property of the controlled chain at a return time. All of these are reusable well beyond this mission.

Difficulty

The estimates (5.9) and (5.10) are elementary once (5.8) and the existence of fβf_\betafβ​ are available. The weight lies elsewhere.

  • Optimal stationary policies. The infimum defining Jβ∗J^*_\betaJβ∗​ ranges over all history-dependent randomized policies. Bringing it down to a single stationary deterministic policy requires the discounted optimality equation, a measurable selection of minimizers on compact action sets, and a verification argument against arbitrary policies.
  • Restarting at τ\tauτ. (5.8) splits the discounted cost at the random time τ\tauτ. The tail must be identified with βτ\beta^\tauβτ times the discounted cost from state 000. This is the strong Markov property for the process built by the Ionescu-Tulcea theorem, applied at a stopping time that may be infinite.

Formalization scope

  • The state space is ℕ; the action space is a metric space with its Borel σ\sigmaσ-algebra. The model CMP carries compact nonempty U(i)U(i)U(i), a nonnegative measurable cost, and continuity of c(i,⋅)c(i,\cdot)c(i,⋅) and P(j∣i,⋅)P(j\mid i,\cdot)P(j∣i,⋅) on U(i)U(i)U(i), the standing assumptions of §5.
  • Policies are history-dependent, randomized and admissible. The path measure is Mathlib's Kernel.trajMeasure. Jβ∗J^*_\betaJβ∗​ is an infimum over all such policies, not over Markov or stationary policies only.
  • Costs are lower Lebesgue integrals in [0,∞][0,\infty][0,∞]. hβh_\betahβ​ is the difference of the real parts of Jβ∗(i)J^*_\beta(i)Jβ∗​(i) and Jβ∗(0)J^*_\beta(0)Jβ∗​(0). This is exact here because bounded cost gives Jβ∗≤M/(1−β)<∞J^*_\beta\le M/(1-\beta)<\inftyJβ∗​≤M/(1−β)<∞.
  • Explicit choices:
    • The bounded-cost hypothesis c≤Mc\le Mc≤M on admissible pairs is a binder of every §5.1 statement. It is the section's standing assumption, and without it the theorem is false.
    • τ\tauτ counts from t≥1t\ge1t≥1 and takes values in N∪{∞}\mathbb N\cup\{\infty\}N∪{∞}, so (5.7) applies from i=0i=0i=0 and forces τ<∞\tau<\inftyτ<∞ almost surely; βτ=0\beta^\tau=0βτ=0 on {τ=∞}\{\tau=\infty\}{τ=∞}.
    • The typo βn\beta^nβn in (5.8) is read as βt\beta^tβt.
    • K≥1K\ge1K≥1 in (5.10) is not assumed; it follows from (5.7).
  • Theorem 2.1 is stated in the survey for Borel models under Assumptions 2.1–2.3. Here it is posed in the countable model, where those assumptions follow from the §5 continuity and compactness assumptions.
  • The goal's bound BBB is quantified before β\betaβ and iii. A per-β\betaβ or per-state bound would be trivial, since every hβ(i)h_\beta(i)hβ​(i) is a finite number.
  • Welcome contributions include discounted dynamic programming on countable state spaces, the strong Markov property for trajMeasure, and return-time estimates.

Selected references

  • A. Arapostathis, V. S. Borkar, E. Fernández-Gaucherand, M. K. Ghosh, S. I. Marcus, Discrete-time controlled Markov processes with average cost criterion: a survey, SIAM J. Control Optim. 31(2) (1993) 282–344. https://doi.org/10.1137/0331018
  • C. Derman, On sequential decisions and Markov chains, Management Sci. 9 (1962) 16–24. https://doi.org/10.1287/mnsc.9.1.16
  • C. Derman, A. F. Veinott Jr., A solution to a countable system of equations arising in Markovian decision processes, Ann. Math. Statist. 38 (1967) 582–584 (cited as [43] in the survey, https://doi.org/10.1137/0331018).
  • S. M. Ross, Non-discounted denumerable Markovian decision models, Ann. Math. Statist. 39 (1968) 412–423 (cited as [147] in the survey, https://doi.org/10.1137/0331018).
  • S. M. Ross, Introduction to Stochastic Dynamic Programming, Academic Press, 1983 (cited as [150] in the survey, https://doi.org/10.1137/0331018).
10 thms1 active userReviewed
Dynamic ProgrammingOperations Research·Captain: mikedeng1

Discrete-Time Controlled Markov Processes with Average Cost Criterion: A Survey 3: A Uniform Lower Bound on the Probability of Moving to State 0 Reduces Average Cost to Discounted CostResearch Paper

Why average cost is a control problem

A controller acting over an indefinite horizon must decide whether a lower cost today is worth a higher cost later. Average cost measures the expected expenditure per stage as the horizon grows. It is appropriate when operation has no natural terminal date, but its limiting definition makes it difficult to compute an optimal policy directly. Discounted cost assigns less weight to distant stages and has a more direct optimality equation. Arapostathis, Borkar, Fernández-Gaucherand, Ghosh and Marcus survey these criteria for controlled Markov processes and state a condition under which solving one discounted problem yields a solution to an average-cost problem (Arapostathis et al., 1993, §5.1).

The condition is a common lower bound on the one-step probability of moving to a distinguished state. Ross's reduction, reported as Theorem 5.6 of the survey, uses that bound to define a new transition law and a specific discount factor. The survey's theorem states the reduction informally; its proof specifies the transformed law, the optimality equation and the resulting average-cost policy. Those claims are the targets of this mission (Arapostathis et al., 1993, pp. 303–304).

Controlled process and criteria

The state space is S={0,1,2,…}S=\{0,1,2,\ldots\}S={0,1,2,…}. In state iii, the controller may choose an action aaa from a nonempty compact set U(i)U(i)U(i) in a metric action space. Choosing aaa incurs the one-stage cost c(i,a)≥0c(i,a)\ge0c(i,a)≥0 and moves the state to jjj with probability P(j∣i,a)P(j\mid i,a)P(j∣i,a). The cost is measurable, and for each fixed i,ji,ji,j its value and P(j∣i,a)P(j\mid i,a)P(j∣i,a) vary continuously with aaa on U(i)U(i)U(i). Section 5 imposes these state and continuity conventions; §5.1 additionally assumes the costs are bounded (Arapostathis et al., 1993, pp. 284–288, 299, 301).

An admissible policy π\piπ chooses a probability law for the next action from the entire observed history, and assigns probability one to admissible actions. A stationary deterministic policy is a map fff with f(i)∈U(i)f(i)\in U(i)f(i)∈U(i); it always takes action f(i)f(i)f(i) in state iii. These are distinct classes. Let JN(i,π)J_N(i,\pi)JN​(i,π) be expected cost over the first NNN stages from iii, and let Jβ(i,π)J_\beta(i,\pi)Jβ​(i,π) be expected cost when stage ttt is weighted by βt\beta^tβt, where 0<β<10<\beta<10<β<1. The average-cost criterion is J(i,π)=lim sup⁡N→∞JN(i,π)/NJ(i,\pi)=\limsup_{N\to\infty}J_N(i,\pi)/NJ(i,π)=limsupN→∞​JN​(i,π)/N. The optimal values J∗(i)J^*(i)J∗(i) and Jβ∗(i)J_\beta^*(i)Jβ∗​(i) are infima over all admissible policies, including randomized and history-dependent ones (Arapostathis et al., 1993, pp. 285–287).

The average cost optimality equation, or ACOE, asks for a scalar ρ\rhoρ and a real function hhh such that, for each state iii,

ρ+h(i)=min⁡a∈U(i){c(i,a)+∑j∈SP(j∣i,a)h(j)}.\rho+h(i)=\min_{a\in U(i)}\left\{c(i,a)+\sum_{j\in S}P(j\mid i,a)h(j)\right\}.ρ+h(i)=a∈U(i)min​⎩⎨⎧​c(i,a)+j∈S∑​P(j∣i,a)h(j)⎭⎬⎫​.

The minimum is attained. The survey's verification theorem identifies ρ\rhoρ with the optimal average cost when the terminal contribution of h(Xt)h(X_t)h(Xt​) vanishes after division by ttt (Arapostathis et al., 1993, p. 299, (5.1), Theorem 5.1).

Formalization targets

The reduction

Assume that P(0∣i,a)≥αP(0\mid i,a)\ge\alphaP(0∣i,a)≥α on every admissible state-action pair for a single constant 0<α<10<\alpha<10<α<1. The transformed process M~\widetilde MM has the same admissible actions and costs and has transition probabilities

P~(j∣i,a)=P(j∣i,a)−α1{j=0}1−α.\widetilde P(j\mid i,a)=\frac{P(j\mid i,a)-\alpha\mathbf1_{\{j=0\}}}{1-\alpha}.P(j∣i,a)=1−αP(j∣i,a)−α1{j=0}​​.

Write J~1−α∗\widetilde J^*_{1-\alpha}J1−α∗​ for its discounted value at discount factor 1−α1-\alpha1−α. The goal states that this value is finite, a stationary deterministic discounted-optimal policy exists, and every such policy is average-cost optimal for the original process. It also identifies a constant optimal average cost for every initial state:

J∗(i)=αJ~1−α∗(0),i∈S.J^*(i)=\alpha\widetilde J^*_{1-\alpha}(0),\qquad i\in S.J∗(i)=αJ1−α∗​(0),i∈S.

The goal does not assume the average optimality that it asserts. It requires the transformed law to satisfy the displayed formula at every admissible state-action pair (Arapostathis et al., 1993, Theorem 5.6 and proof, pp. 303–304).

Supporting targets

The milestones establish that the transformed law gives a controlled Markov process, that the discounted problem has an optimal stationary deterministic policy, and that the transformed discounted value satisfies the original ACOE with ρ=αJ~1−α∗(0)\rho=\alpha\widetilde J^*_{1-\alpha}(0)ρ=αJ1−α∗​(0). The final milestone is the ACOE verification theorem needed to identify the average cost. The source gives the first two discounted claims through Theorem 2.1 and writes out the transformed equation in the proof of Theorem 5.6 (Arapostathis et al., 1993, pp. 289, 299, 304).

What the result supplies

The theorem replaces an average-cost optimization problem by one discounted problem with a prescribed discount factor and transition law. It yields a stationary deterministic policy that is optimal even when compared with every history-dependent randomized policy, and it shows that the optimal average cost is independent of the initial state. Without the common return probability, neither this particular law nor this fixed discount factor follows from the survey's argument (Arapostathis et al., 1993, Theorem 5.6).

The survey reports this as a known result of Ross rather than an open conjecture. The formalization work is to give the path measures, value functions, transformed process and verification result machine-checkable meanings. The Lean declarations here are open theorem statements awaiting proofs; compiling a statement with sorry does not establish the mathematical theorem. The policy and cost definitions can also support the other countable-state missions drawn from §5.

Difficulty

The transformed probabilities have to form a measurable stochastic kernel, preserve the action continuity assumptions and yield a controlled process with the original feasible actions and costs. The discounted value must be finite and uniformly bounded before its real form can enter the ACOE. A simple comparison of numerical optimal values is insufficient: a policy chosen in the transformed model must be shown optimal for the original model against the full policy class. The verification theorem also requires control of the terminal expectation of h(Xt)h(X_t)h(Xt​) for arbitrary admissible policies, rather than only the stationary policies named in its printed condition (Arapostathis et al., 1993, pp. 299–300, 304).

Formalization scope

Lean uses N\mathbb NN for the countable state space and Mathlib probability kernels for the transition and randomized decision rules. Strategic path measures are constructed from those kernels. Nonnegative expected costs and their infima live in [0,∞][0,\infty][0,∞], so an unbounded integral cannot silently become a finite real value. The transformed discounted value is converted to a real only in conclusions that also assert its finiteness. The ACOE uses real sums with explicit summability and an attained minimum. The model requires a Borel metric action space, nonempty compact action sets, measurable costs, and coordinatewise action continuity of the transition probabilities.

The source's §5.1 bounded-cost assumption is explicit in the goal and its discounted milestones. The displayed transformation needs α<1\alpha<1α<1; the paper's theorem sentence gives only α>0\alpha>0α>0, so the case α=1\alpha=1α=1 is excluded from this version. The paper prints (5.2) for stationary deterministic policies, but the proof uses it for all admissible policies to compare with J∗J^*J∗; the verification milestone takes the stronger, proof-supported hypothesis. Expectations in that condition are explicitly integrable. The converse part of Theorem 5.1 is not used and is outside this mission.

The transformed process is represented by another CMP constrained to have exactly the source's action sets, admissible costs and transition formula. A milestone poses its existence. This representation keeps the definition layer free of an unproved stochastic-kernel construction. The discount-optimal policy conclusion ranges over every stationary deterministic policy attaining the transformed discounted value, while the values themselves take infima over all admissible policies. Definitions of admissible path laws and the verification theorem are reusable contributions; proofs of the transformed kernel, stationary discounted existence and ACOE identity are welcome.

Selected references

  • A. Arapostathis, V. S. Borkar, E. Fernández-Gaucherand, M. K. Ghosh and S. I. Marcus, Discrete-time controlled Markov processes with average cost criterion: a survey, SIAM Journal on Control and Optimization 31(2), 282–344, 1993. DOI: 10.1137/0331018.
7 thms1 active userReviewed
Control TheoryDynamic ProgrammingOperations Research·Captain: mikedeng1

Optimal Control of Markov Processes with Incomplete State Information 1: Reduction to Complete State Information on the Conditional State Distributions, with the Same Optimal LawResearch Paper

Motivation

A controller often cannot see the state of the system it steers. It sees only measurements that are noisy functions of that state. In operations research this happens in machine maintenance and inspection, in queues observed only partially, and in inventory systems with inexact stock records. In control engineering it is the usual case. The question is what the controller should base its decisions on. The full record of past measurements is the obvious choice, but that record grows with time, so a law that uses it is a function on a space whose dimension grows with the horizon.

K. J. Åström's 1965 paper Optimal Control of Markov Processes with Incomplete State Information answered this question for finite Markov chains. The answer is that the conditional distribution of the hidden state given the measurements is a sufficient statistic. The problem with incomplete information is equivalent to a problem with complete information whose state is that distribution. This is the model now called a partially observable Markov decision process (POMDP), and the conditional distribution is now called the belief state.

Timeline. For linear systems with quadratic cost and Gaussian noise, the separation theorem of Joseph and Tou (1961) and Gunckel and Franklin (1963) says that the optimal control is a fixed function of the conditional mean of the state. Åström (1965) proved the reduction for finite-state Markov chains with arbitrary costs, with the conditional distribution as the new state. Smallwood and Sondik (1973) showed that for finite horizons the value function is piecewise linear and concave in the belief, which made exact computation possible. Bertsekas and Shreve (1978) and Bäuerle and Rieder (2011) gave the reduction for general Borel models.

Setting

The hidden state xtx_txt​, t=1,…,Nt = 1, \dots, Nt=1,…,N, takes values in a finite set SSS. The controls u=(u1,…,ur)u = (u_1, \dots, u_r)u=(u1​,…,ur​) range over a compact nonempty set U⊂RrU \subset \mathbb R^rU⊂Rr. The state moves by the transition probabilities pij(u,t)=P{xt=j∣xt−1=i}p_{ij}(u, t) = P\{x_t = j \mid x_{t-1} = i\}pij​(u,t)=P{xt​=j∣xt−1​=i}, which are continuous in uuu. The state is observed through outputs yty_tyt​ in a finite set YYY, with qij=P{yt=j∣xt=i}q_{ij} = P\{y_t = j \mid x_t = i\}qij​=P{yt​=j∣xt​=i}, conditionally independent given the states. The law of x1x_1x1​ is p1p^1p1. An instantaneous cost g(u,i,t)g(u, i, t)g(u,i,t), continuous in uuu, is paid at each time.

A control law chooses u(t)=c(η1,…,ηt,t)∈Uu(t) = c(\eta_1, \dots, \eta_t, t) \in Uu(t)=c(η1​,…,ηt​,t)∈U from the outputs observed so far, η(t)=(η1,…,ηt)\eta(t) = (\eta_1, \dots, \eta_t)η(t)=(η1​,…,ηt​). With u(t)u(t)u(t) moving xtx_txt​ to xt+1x_{t+1}xt+1​, a law determines the joint law of (x1,…,xN,y1,…,yN)(x_1, \dots, x_N, y_1, \dots, y_N)(x1​,…,xN​,y1​,…,yN​) and the expected cost

EL=E∑t=1Ng(u(t),xt,t).(2.6)EL = E \sum_{t=1}^N g(u(t), x_t, t). \tag{2.6}EL=Et=1∑N​g(u(t),xt​,t).(2.6)

Problem P.1 is to find an admissible law minimizing (2.6).

The conditional state distribution is wi(t)=P{xt=i∣η(t)}w_i(t) = P\{x_t = i \mid \eta(t)\}wi​(t)=P{xt​=i∣η(t)}. It is updated by Bayes' rule: with zj(u,w)i=∑sqij psi(u,t+1) wsz^j(u, w)_i = \sum_s q_{ij}\, p_{si}(u, t+1)\, w_szj(u,w)i​=∑s​qij​psi​(u,t+1)ws​ and ∥z∥=∑i∣zi∣\|z\| = \sum_i |z_i|∥z∥=∑i​∣zi​∣, the output ηt+1=j\eta_{t+1} = jηt+1​=j gives w(t+1)=zj(u(t),w(t))/∥zj(u(t),w(t))∥w(t+1) = z^j(u(t), w(t)) / \|z^j(u(t), w(t))\|w(t+1)=zj(u(t),w(t))/∥zj(u(t),w(t))∥, and ∥zj∥\|z^j\|∥zj∥ is the probability of that output. The cost-to-go Vk(w)V_k(w)Vk​(w) is the minimal expected cost of the steps k,…,Nk, \dots, Nk,…,N when xkx_kxk​ has distribution www, with VN+1=0V_{N+1} = 0VN+1​=0. Problem P.2 controls the process w(t)w(t)w(t) directly: a law chooses u(t)u(t)u(t) from w(1),…,w(t)w(1), \dots, w(t)w(1),…,w(t) to minimize E∑t=1N∑ig(u(t),i,t) wi(t)E\sum_{t=1}^N \sum_i g(u(t), i, t)\, w_i(t)E∑t=1N​∑i​g(u(t),i,t)wi​(t).

Formalization targets

Goal: Theorem 3

P.1 has a solution if and only if P.2 has one. For every solution (V,c0)(V, c^0)(V,c0) of the functional equation

Vk(w)=min⁡u∈U{∑ig(u,i,k) wi+∑jVk+1(zj(u,w)∥zj(u,w)∥)∥zj(u,w)∥},VN+1=0,(3.28)V_k(w) = \min_{u \in U} \Big\{ \sum_i g(u, i, k)\, w_i + \sum_j V_{k+1}\Big(\frac{z^j(u, w)}{\|z^j(u, w)\|}\Big) \|z^j(u, w)\| \Big\}, \qquad V_{N+1} = 0, \tag{3.28}Vk​(w)=u∈Umin​{i∑​g(u,i,k)wi​+j∑​Vk+1​(∥zj(u,w)∥zj(u,w)​)∥zj(u,w)∥},VN+1​=0,(3.28)

with c0(w,k)c^0(w, k)c0(w,k) attaining the minimum, the law

u(t)=c0(w(t),t)u(t) = c^0(w(t), t)u(t)=c0(w(t),t)

is optimal for P.1 and for P.2, among all admissible laws of each, and both minimal values equal Eη1V1(w(1))E_{\eta_1} V_1(w(1))Eη1​​V1​(w(1)).

Milestones

  1. (3.20)–(3.25): the conditional distributions obey the Bayes recursion, and ∥zj∥=P[yt+1=j∣η(t)]\|z^j\| = P[y_{t+1} = j \mid \eta(t)]∥zj∥=P[yt+1​=j∣η(t)].
  2. Theorem 1: the cost-to-go satisfies (3.28) with the minimum attained, and an optimal Markov law attains it.
  3. Theorem 2: a solution of (3.28) gives an optimal law for P.1 with value (3.29).
  4. Lemma 1: under u(t)=c(w(t),t)u(t) = c(w(t), t)u(t)=c(w(t),t), {w(t)}\{w(t)\}{w(t)} is a Markov process with transition probabilities P(y,Γ,u)=∑k∈K∥zk(u,y)∥P(y, \Gamma, u) = \sum_{k \in K} \|z^k(u, y)\|P(y,Γ,u)=∑k∈K​∥zk(u,y)∥.
  5. Proof of Theorem 3: the integral against this kernel is the sum in (3.28).

Significance

The result. Theorem 3 replaces a minimization over functions of ever longer measurement records with a recursion over a fixed space, the probability simplex over the states. Every exact and approximate POMDP algorithm starts from it: value iteration on beliefs, the piecewise-linear representation of Smallwood and Sondik, point-based methods. It also splits the controller in two. A filter computes w(t)w(t)w(t) in real time, and the function c0c^0c0 can be computed off-line. This is the decomposition the paper draws on p. 189, and it extends the linear-quadratic separation theorem to arbitrary finite chains.

Formalizing it. The theorem is proved. The platform has the reduction in Bäuerle and Rieder's discounted Borel model with an observable state component and rewards in extended reals. It does not have Åström's model: finite chains, time-dependent transition matrices, an unobservable state, costs, and laws of the raw output history. This mission formalizes Åström's statements as he gives them. The cost (2.6) is defined from the joint law of states and outputs, and the comparison classes are all laws of the outputs (P.1) and all laws of the distribution history (P.2). The finite setting makes every expectation a finite sum, so a complete development needs no measure theory.

Difficulty

The obvious argument is backward induction on the conditional distributions. The difficulty is that w(t)w(t)w(t) depends on the controls already used, so it is not given in advance: the state of the reduced problem is produced by the law being optimized. It has to be shown that the expected cost of an arbitrary law of the outputs, computed from the joint law, splits as the reduced recursion says. In particular, laws that use more of the record than w(t)w(t)w(t) must gain nothing. Restricting the comparison class to laws of the form c(w(t),t)c(w(t), t)c(w(t),t) assumes this conclusion.

A second difficulty is attainment. "Min" in (3.28) and "has a solution" presuppose that minima over UUU are attained, which needs continuity of Vk+1V_{k+1}Vk+1​ on the simplex. The weights ∥zj(u,w)∥\|z^j(u, w)\|∥zj(u,w)∥ can vanish, and then the update zj/∥zj∥z^j/\|z^j\|zj/∥zj∥ is undefined.

Formalization scope

States and outputs are finite types, St and Obs, with the chain given by the structure Model. Controls are Fin r → ℝ, and UUU is compact and nonempty. The law p1p^1p1 of x1x_1x1​ is the datum in place of the paper's law of x0x_0x0​, since no control u(0)u(0)u(0) exists. The transition from xtx_txt​ to xt+1x_{t+1}xt+1​ uses u(t)u(t)u(t) and the matrix p(u(t),t+1)p(u(t), t+1)p(u(t),t+1). Times 1,…,N1, \dots, N1,…,N are indexed by Fin N as 0,…,N−10, \dots, N-10,…,N−1.

The norm ∥⋅∥\|\cdot\|∥⋅∥ is the ℓ1\ell^1ℓ1 norm l1, not Mathlib's sup norm. Conditional distributions are ratios of path sums, condState, and are claimed only on output histories of positive probability. When ∥zj∥=0\|z^j\| = 0∥zj∥=0 the update is the zero vector and is always multiplied by 000.

The cost-to-go costToGo is an infimum over admissible tail laws. Its index set is nonempty and the costs are bounded below, so the real infimum is a true infimum. It is never defined through (3.28), since that would make Theorem 1 circular. The P.2 functional sums branch by branch over the outputs, with weights ∥zj∥\|z^j\|∥zj∥. "Given by Theorem 1" is read as "c0(w,k)∈Uc^0(w, k) \in Uc0(w,k)∈U attains the minimum in (3.28)" (IsSolution328).

The goal is not the bare equivalence of solvability. In this compact, continuous, finite setting both problems always have solutions, so that sentence alone is trivially true. The goal also requires the law c0(w(t),t)c^0(w(t), t)c0(w(t),t) to be optimal in both problems, against every admissible law, with equal minimal values.

Reusable beyond this mission are the finite POMDP model, the joint path law, the Bayes filter and the belief-MDP kernel. Welcome contributions include proofs of the milestones, the continuity of VkV_kVk​ on the simplex, and existence of solutions of (3.28).

Selected references

  • K. J. Åström, Optimal control of Markov processes with incomplete state information, Journal of Mathematical Analysis and Applications 10(1):174–205, 1965. https://doi.org/10.1016/0022-247X(65)90154-X
  • R. D. Smallwood and E. J. Sondik, The optimal control of partially observable Markov processes over a finite horizon, Operations Research 21(5):1071–1088, 1973. https://doi.org/10.1287/opre.21.5.1071
  • D. P. Bertsekas and S. E. Shreve, Stochastic Optimal Control: The Discrete-Time Case, Academic Press, 1978, Chapter 10. https://web.mit.edu/dimitrib/www/soc.html
  • N. Bäuerle and U. Rieder, Markov Decision Processes with Applications to Finance, Springer, 2011, Chapter 5. https://doi.org/10.1007/978-3-642-18324-9
  • P. D. Joseph and J. T. Tou, On linear control theory, Transactions of the AIEE, Part II 80(4):193–196, 1961. https://doi.org/10.1109/TAI.1961.6371743
10 thms1 active userReviewed
Dynamic ProgrammingOperations Research·Captain: mikedeng1

Discrete-Time Controlled Markov Processes with Average Cost Criterion: A Survey 1: Uniformly Bounded Differential Discounted Values Give a Bounded Solution of the Average Cost Optimality EquationResearch Paper

Motivation

Many control problems in queueing, inventory and communication systems run indefinitely, and the quantity of interest is the long-run cost per unit time rather than a discounted total. The average cost criterion is harder to analyse than the discounted one: the discounted dynamic programming operator is a contraction, while the average cost problem has no contraction, and on an infinite state space its behaviour depends on the recurrence structure of the controlled chain. The survey of Arapostathis, Borkar, Fernández-Gaucherand, Ghosh and Marcus (SIAM J. Control Optim. 31 (1993)) organises the theory around the average cost optimality equation (ACOE) and the conditions under which it has a solution.

Timeline (as recorded in the survey's §3 and §5). Derman studied the ACOE and characterized optimal stationary policies by its solutions (Derman, On sequential decisions and Markov chains, Management Sci. 1962; Denumerable state Markovian decision processes — average cost criterion, Ann. Math. Statist. 1966). Taylor introduced a vanishing discount argument for a replacement problem (Ann. Math. Statist. 1965). Ross extended it to general countable models, showing that uniformly bounded differential discounted value functions yield a bounded solution of the ACOE (Ann. Math. Statist. 1968, two papers; Introduction to Stochastic Dynamic Programming, 1983). Sennott replaced the uniform bound by one-sided bounds and obtained the average cost optimality inequality (Oper. Res. 1989). The survey presents Ross's result as Theorem 5.2, following the 1983 book; this mission formalizes it.

Setting

A controlled Markov process on the countable state space S={0,1,2,… }S=\{0,1,2,\dots\}S={0,1,2,…} consists of a metric space A\mathbf AA of actions; for each state iii a nonempty compact set U(i)⊆AU(i)\subseteq\mathbf AU(i)⊆A of admissible actions; a cost c(i,a)≥0c(i,a)\ge0c(i,a)≥0; and transition probabilities P(j∣i,a)P(j\mid i,a)P(j∣i,a). For fixed i,ji,ji,j, the maps a↦c(i,a)a\mapsto c(i,a)a↦c(i,a) and a↦P(j∣i,a)a\mapsto P(j\mid i,a)a↦P(j∣i,a) are continuous on U(i)U(i)U(i).

An admissible policy π\piπ chooses, at each time ttt, a probability distribution on U(Xt)U(X_t)U(Xt​) that may depend on the whole past (X0,A0,…,Xt)(X_0,A_0,\dots,X_t)(X0​,A0​,…,Xt​). The class of all of them is Π\PiΠ, and ΠSD\Pi_{SD}ΠSD​ is the class of stationary deterministic policies, maps fff with f(i)∈U(i)f(i)\in U(i)f(i)∈U(i). Each initial state iii and policy π\piπ define a law PiπP^\pi_iPiπ​ of the trajectory, with expectation EiπE^\pi_iEiπ​. For a discount factor 0<β<10<\beta<10<β<1,

Jβ(i,π)=Eiπ∑t=0∞βtc(Xt,At),J(i,π)=lim sup⁡N→∞1NEiπ∑t=0N−1c(Xt,At),J_\beta(i,\pi)=E^\pi_i\sum_{t=0}^\infty\beta^tc(X_t,A_t),\qquad J(i,\pi)=\limsup_{N\to\infty}\frac1N E^\pi_i\sum_{t=0}^{N-1}c(X_t,A_t),Jβ​(i,π)=Eiπ​t=0∑∞​βtc(Xt​,At​),J(i,π)=N→∞limsup​N1​Eiπ​t=0∑N−1​c(Xt​,At​),

and Jβ∗(i)=inf⁡π∈ΠJβ(i,π)J^*_\beta(i)=\inf_{\pi\in\Pi}J_\beta(i,\pi)Jβ∗​(i)=infπ∈Π​Jβ​(i,π), J∗(i)=inf⁡π∈ΠJ(i,π)J^*(i)=\inf_{\pi\in\Pi}J(i,\pi)J∗(i)=infπ∈Π​J(i,π). The differential discounted value function is hβ(i)=Jβ∗(i)−Jβ∗(0)h_\beta(i)=J^*_\beta(i)-J^*_\beta(0)hβ​(i)=Jβ∗​(i)−Jβ∗​(0). A pair (ρ,h)(\rho,h)(ρ,h), ρ∈R\rho\in\mathbb Rρ∈R, h:S→Rh:S\to\mathbb Rh:S→R, solves the ACOE if

ρ+h(i)=min⁡a∈U(i){c(i,a)+∑j∈SP(j∣i,a)h(j)},i∈S.(5.1)\rho+h(i)=\min_{a\in U(i)}\Big\{c(i,a)+\sum_{j\in S}P(j\mid i,a)h(j)\Big\},\qquad i\in S.\tag{5.1}ρ+h(i)=a∈U(i)min​{c(i,a)+j∈S∑​P(j∣i,a)h(j)},i∈S.(5.1)

In the Lean development these objects are CMP, Policy, StationaryPolicy, pathMeasure, discCost, avgCost, discValue (Jβ∗J^*_\betaJβ∗​), optAvg (J∗J^*J∗), hRel (hβh_\betahβ​) and ACOE.

Formalization targets

Goal: Theorem 5.2 (p. 301)

Assume Jβ∗(i)<∞J^*_\beta(i)<\inftyJβ∗​(i)<∞ for all β∈(0,1)\beta\in(0,1)β∈(0,1) and i∈Si\in Si∈S, and that there is K>0K>0K>0 with ∣hβ(i)∣≤K|h_\beta(i)|\le K∣hβ​(i)∣≤K for all such β\betaβ and iii. Then there are ρ∈R\rho\in\mathbb Rρ∈R, a bounded h:S→Rh:S\to\mathbb Rh:S→R and a sequence βn∈(0,1)\beta_n\in(0,1)βn​∈(0,1), βn→1\beta_n\to1βn​→1, with

(ρ,h) solves (5.1),h(i)=lim⁡n→∞hβn(i),lim⁡β↑1(1−β)Jβ∗(i)=ρ(i∈S).(\rho,h)\text{ solves (5.1)},\qquad h(i)=\lim_{n\to\infty}h_{\beta_n}(i),\qquad \lim_{\beta\uparrow1}(1-\beta)J^*_\beta(i)=\rho\qquad(i\in S).(ρ,h) solves (5.1),h(i)=n→∞lim​hβn​​(i),β↑1lim​(1−β)Jβ∗​(i)=ρ(i∈S).

The goal does not assert that ρ\rhoρ is the optimal average cost; that follows from Theorem 5.1 and Remark 5.1(a), which are milestones.

Milestones

  1. Lemma 2.1 (p. 289): the dynamic programming map T(v)(i)=inf⁡a∈U(i){c(i,a)+∑jP(j∣i,a)v(j)}T(v)(i)=\inf_{a\in U(i)}\{c(i,a)+\sum_jP(j\mid i,a)v(j)\}T(v)(i)=infa∈U(i)​{c(i,a)+∑j​P(j∣i,a)v(j)} satisfies T(v+k)=T(v)+kT(v+k)=T(v)+kT(v+k)=T(v)+k and is monotone.
  2. Theorem 2.1 (i), (iii) (p. 289), in the countable model: Jβ∗=TβJβ∗J^*_\beta=T_\beta J^*_\betaJβ∗​=Tβ​Jβ∗​ and a β\betaβ-discount optimal f∈ΠSDf\in\Pi_{SD}f∈ΠSD​ exists.
  3. Equation (5.6) (p. 301): (1−β)Jβ∗(0)+hβ(i)=min⁡a∈U(i){c(i,a)+β∑jP(j∣i,a)hβ(j)}(1-\beta)J^*_\beta(0)+h_\beta(i)=\min_{a\in U(i)}\{c(i,a)+\beta\sum_jP(j\mid i,a)h_\beta(j)\}(1−β)Jβ∗​(0)+hβ​(i)=mina∈U(i)​{c(i,a)+β∑j​P(j∣i,a)hβ​(j)}.
  4. Theorem 5.1 (p. 299): a solution of (5.1) with lim⁡t1tEiπh(Xt)=0\lim_t\frac1tE^\pi_ih(X_t)=0limt​t1​Eiπ​h(Xt​)=0 gives ρ=J(i,f)=J∗(i)\rho=J(i,f)=J^*(i)ρ=J(i,f)=J∗(i) for a minimizing selector fff; minimizing selectors are average optimal; conversely, an average optimal fff with an irreducible positive recurrent chain is a minimizing selector.
  5. Remark 5.1(a) (p. 300): a bounded solution of (5.1) satisfies the growth condition of Theorem 5.1.

Significance

Theorem 5.2 is the template of the vanishing discount method. Under its hypothesis the average cost problem has a bounded solution of the ACOE, so (through Theorem 5.1) the optimal average cost is a constant ρ\rhoρ independent of the initial state, it is attained by a stationary deterministic policy, and it is the Abelian limit of the scaled discounted values. Recurrence conditions on the controlled chain, such as uniformly bounded mean return times to a fixed state (Theorem 5.3 of the survey), are verified by checking the hypothesis of Theorem 5.2; the later results of §5 refine its conclusion under weaker hypotheses.

The theorem itself is classical. What this mission adds is a machine-checked statement and, eventually, proof, on a model with history-dependent randomized policies, compact action sets and unbounded costs, together with the supporting verification theorem (Theorem 5.1) and the discounted optimality equation. To our knowledge none of these results has been formalized in Lean; Mathlib has the Ionescu-Tulcea construction of the path measure but no controlled Markov processes.

Difficulty

The obvious argument fixes a sequence βn↑1\beta_n\uparrow1βn​↑1, extracts a pointwise convergent subsequence of the bounded functions hβnh_{\beta_n}hβn​​ and of the bounded numbers (1−βn)Jβn∗(0)(1-\beta_n)J^*_{\beta_n}(0)(1−βn​)Jβn​∗​(0), and passes to the limit in (5.6). Two steps resist this. First, the limit of a minimum over U(i)U(i)U(i) is not in general the minimum of the limits: the convergence of a↦∑jP(j∣i,a)hβn(j)a\mapsto\sum_jP(j\mid i,a)h_{\beta_n}(j)a↦∑j​P(j∣i,a)hβn​​(j) must be shown to be uniform on the compact set U(i)U(i)U(i), which requires more than pointwise continuity of each P(j∣i,⋅)P(j\mid i,\cdot)P(j∣i,⋅). Second, the subsequential limit ρ\rhoρ could depend on the subsequence, so part (iii), a limit along all β↑1\beta\uparrow1β↑1, needs an independent identification of ρ\rhoρ, here as the optimal average cost through Theorem 5.1, which in turn needs the comparison with arbitrary history-dependent policies. The discounted optimality equation behind (5.6) also has to be established for unbounded costs, where Jβ∗J^*_\betaJβ∗​ is not the unique fixed point of TβT_\betaTβ​.

Formalization scope

The state space is ℕ; state 0 is the reference state of hβh_\betahβ​. Policies are history-dependent randomized stochastic kernels with the admissibility constraint πt(U(xt)∣ht)=1\pi_t(U(x_t)\mid h_t)=1πt​(U(xt​)∣ht​)=1, and J∗J^*J∗, Jβ∗J^*_\betaJβ∗​ are infima over all of them. Costs are lower Lebesgue integrals with values in [0,∞][0,\infty][0,∞], built from Mathlib's Kernel.trajMeasure. The following choices make implicit hypotheses explicit:

  • Finiteness of Jβ∗J^*_\betaJβ∗​. The paper's bound ∣hβ∣≤K|h_\beta|\le K∣hβ​∣≤K presupposes finite values; the goal assumes Jβ∗(i)<∞J^*_\beta(i)<\inftyJβ∗​(i)<∞, and (5.6) assumes it for its β\betaβ.
  • Convergent series in the ACOE. A solution of (5.1) requires every series ∑jP(j∣i,a)h(j)\sum_jP(j\mid i,a)h(j)∑j​P(j∣i,a)h(j), a∈U(i)a\in U(i)a∈U(i), to converge, and the minimum to be attained.
  • (5.2) over all policies. The paper prints the growth condition of Theorem 5.1 for π∈ΠSD\pi\in\Pi_{SD}π∈ΠSD​, but its conclusion ρ=J∗(i)\rho=J^*(i)ρ=J∗(i) concerns all policies, and the proof uses the condition for arbitrary π\piπ. It is stated for every π∈Π\pi\in\Piπ∈Π, with integrability of h(Xt)h(X_t)h(Xt​) explicit.
  • Theorem 2.1 is cited without proof in the survey for Borel models; in the countable model its Assumptions 2.2–2.3 follow from the continuity assumptions of §5. Only parts (i) and (iii) are stated.
  • Lemma 2.1 is stated for functions bounded below (on discrete ℕ these are the lower semicontinuous functions bounded below), with convergent series.
  • Irreducible, positive recurrent (converse of Theorem 5.1): every state is reached with positive probability from every state, and every state has finite expected return time.

A formalization in which J∗J^*J∗ is an infimum over stationary policies only, in which the ACOE is an inequality or holds for one fixed action, or in which hβh_\betahβ​ is computed from +∞+\infty+∞ values through a junk conversion, would trivialize the goal; all three are excluded by the definitions above.

A complete development needs the Ionescu-Tulcea path measure for history-dependent policies, the discounted optimality equation for nonnegative unbounded costs, Scheffé-type uniform convergence on compact action sets, and the martingale identity behind Theorem 5.1. The model file is reusable by the other missions of this series and by any countable-state average cost result; proofs of the milestones are welcome independently of the goal.

Selected references

  • A. Arapostathis, V. S. Borkar, E. Fernández-Gaucherand, M. K. Ghosh, S. I. Marcus, Discrete-time controlled Markov processes with average cost criterion: a survey, SIAM J. Control Optim. 31(2) (1993) 282–344. https://doi.org/10.1137/0331018
  • C. Derman, On sequential decisions and Markov chains, Management Sci. 9 (1962) 16–24 (reference [38] of the survey).
  • C. Derman, Denumerable state Markovian decision processes — average cost criterion, Ann. Math. Statist. 37 (1966) 1545–1553 (reference [39]).
  • H. M. Taylor, Markovian sequential replacement processes, Ann. Math. Statist. 36 (1965) 1677–1694 (reference [177]).
  • S. M. Ross, Non-discounted denumerable Markovian decision models, Ann. Math. Statist. 39 (1968) 412–423, and Arbitrary state Markovian decision processes, Ann. Math. Statist. 39 (1968) 2118–2122 (references [147], [148]).
  • S. M. Ross, Introduction to Stochastic Dynamic Programming, Academic Press, New York, 1983 (reference [150]).
  • L. I. Sennott, Average cost optimal stationary policies in infinite state Markov decision processes with unbounded costs, Oper. Res. 37 (1989) 626–633 (reference [156]).
9 thms1 active userReviewed
Operations ResearchProbabilityStochastic Systems·Captain: mikedeng1

On the Stochastic Matrices Associated with Certain Queuing Processes 2: The GI/M/1 Imbedded Chain Is Ergodic iff ρ < 1 and Recurrent iff ρ ≤ 1Research Paper

Motivation

A single-server queue in which customers arrive according to a renewal process and are served in exponentially distributed times is the system GI/M/1. Observed just before successive arrivals, its queue length is a Markov chain on {0,1,2,… }\{0, 1, 2, \dots\}{0,1,2,…}, the imbedded chain introduced by D. G. Kendall (Kendall 1953, Ann. Math. Statist. 24, pp. 338–354). Whether this chain settles into a statistical equilibrium, keeps returning to the empty state without one, or drifts off to infinity is the first question asked about the queue, and every later quantity (stationary queue lengths, waiting-time distributions) presupposes the answer.

F. G. Foster's 1953 paper (Foster 1953) answers it for GI/M/1 and for M/G/1 by a different route from Kendall's direct analysis: it first proves general criteria, stated in terms of solutions of linear equations and inequalities in the transition matrix, for a countable Markov chain to be ergodic, recurrent or transient, and then checks them on the two queueing matrices. The criteria are of independent use; one of them (Theorem 2 of the paper) is now known as Foster's criterion, the starting point of the drift (Lyapunov-function) method for stability of Markov chains and queueing networks.

Timeline. Kendall (1951, J. Roy. Statist. Soc. B 13) studied queue-length processes directly, including a recurrence argument for M/G/1 that Foster's §3 reproduces; Kendall (1953) introduced the imbedded-chain method and, for GI/M/1, proved by it that ρ<1\rho < 1ρ<1 is sufficient for ergodicity (Foster 1953, p. 359); Foster (1953) proved the full classification, ergodic iff ρ<1\rho < 1ρ<1 and recurrent iff ρ≤1\rho \le 1ρ≤1, by the general criteria. This mission treats the GI/M/1 half; a companion mission treats M/G/1.

Setting

A transition matrix on the states {0,1,2,… }\{0, 1, 2, \dots\}{0,1,2,…} is an array [pij][p_{ij}][pij​] of nonnegative reals whose rows sum to 111. For a state jjj, fjjf_{jj}fjj​ is the probability that the chain started at jjj returns to jjj at some later step. The chain is recurrent if fjj=1f_{jj} = 1fjj​=1 for every jjj, transient if fjj<1f_{jj} < 1fjj​<1 for every jjj, and ergodic (recurrent-nonnull, positive recurrent) if moreover every mean recurrence time ∑nnfjj(n)\sum_n n f^{(n)}_{jj}∑n​nfjj(n)​ is finite. Foster's general theorems concern an irreducible chain (every state reachable from every state), assumed aperiodic for simplicity.

The GI/M/1 chain is described by a sequence a=(an)n≥0a = (a_n)_{n \ge 0}a=(an​)n≥0​ of positive numbers with ∑nan=1\sum_n a_n = 1∑n​an​=1: ana_nan​ is the probability that exactly nnn services are completed between two arrivals. With the tails αi=∑j≥i+1aj\alpha_i = \sum_{j \ge i+1} a_jαi​=∑j≥i+1​aj​,

[pij]=[α0a000⋯α1a1a00⋯α2a2a1a0⋯⋮⋮⋮⋮],[p_{ij}] = \begin{bmatrix} \alpha_0 & a_0 & 0 & 0 & \cdots \\ \alpha_1 & a_1 & a_0 & 0 & \cdots \\ \alpha_2 & a_2 & a_1 & a_0 & \cdots \\ \vdots & \vdots & \vdots & \vdots & \end{bmatrix},[pij​]=​α0​α1​α2​⋮​a0​a1​a2​⋮​0a0​a1​⋮​00a0​⋮​⋯⋯⋯​​,

that is pi0=αip_{i0} = \alpha_ipi0​=αi​, pij=ai+1−jp_{ij} = a_{i+1-j}pij​=ai+1−j​ for 1≤j≤i+11 \le j \le i+11≤j≤i+1, and pij=0p_{ij} = 0pij​=0 for j>i+1j > i+1j>i+1. In Lean this matrix is gim1Matrix a. The traffic parameter ρ\rhoρ is defined through its inverse,

ρ−1=∑n=1∞n an∈(0,∞],\rho^{-1} = \sum_{n=1}^{\infty} n\, a_n \in (0, \infty],ρ−1=n=1∑∞​nan​∈(0,∞],

the mean number of service completions per interarrival interval (rhoInv a, and rho a =ρ= \rho=ρ).

Formalization targets

Goal: the classification of GI/M/1 (§4, p. 359)

the chain is ergodic  ⟺  ρ<1,the chain is recurrent  ⟺  ρ≤1.\text{the chain is ergodic} \iff \rho < 1, \qquad \text{the chain is recurrent} \iff \rho \le 1 .the chain is ergodic⟺ρ<1,the chain is recurrent⟺ρ≤1.

Together: ergodic for ρ<1\rho < 1ρ<1, recurrent-null for ρ=1\rho = 1ρ=1, transient for ρ>1\rho > 1ρ>1. The statement carries no constants and leaves the sequence aaa free apart from positivity and normalization.

Milestones

  1. Theorem 7 (p. 358): for a probability distribution {pn}\{p_n\}{pn​} with p0>0p_0 > 0p0​>0, the equation ∑n≥0znpn=z\sum_{n \ge 0} z^n p_n = z∑n≥0​znpn​=z has a root in (0,1)(0, 1)(0,1) iff ∑n≥1npn>1\sum_{n\ge1} n p_n > 1∑n≥1​npn​>1.
  2. Theorem 1, sufficiency (p. 355): a nonnull solution of ∑ixipij=xj\sum_i x_i p_{ij} = x_j∑i​xi​pij​=xj​ with ∑i∣xi∣<∞\sum_i |x_i| < \infty∑i​∣xi​∣<∞ makes the system ergodic.
  3. Theorem 1, necessity (p. 355): in an ergodic system every nonnegative solution of ∑ixipij≤xj\sum_i x_i p_{ij} \le x_j∑i​xi​pij​≤xj​ has ∑ixi<∞\sum_i x_i < \infty∑i​xi​<∞.
  4. Theorem 4 (pp. 356–357): the system is transient iff ∑jpijyj=yi\sum_j p_{ij} y_j = y_i∑j​pij​yj​=yi​ (i≠0i \ne 0i=0) has a bounded nonconstant solution.

Milestones 2–4 are stated for a general irreducible aperiodic chain.

Significance

The classification tells exactly when the GI/M/1 queue is stable: the stationary distribution of the imbedded chain, which is geometric, exists precisely in the ergodic case ρ<1\rho < 1ρ<1, and for ρ>1\rho > 1ρ>1 the queue grows without bound. Theorems 1 and 4 are general tools, reusable for any countable chain: Theorem 1 characterizes ergodicity by summable invariant vectors, Theorem 4 characterizes transience by bounded harmonic functions off one state. Theorem 7 is the extinction criterion of branching processes and recurs throughout applied probability.

All of these results are proved in the literature (Foster 1953; Feller's textbook for Theorem 7 and a version of Theorem 4). As far as a search of the platform shows, none of them has a machine-checked proof; the platform holds related special cases for the G/M/1 queue with a specific interarrival law (QueueingFundamentals.GM1.unique_root_unit_interval, open), but not the general lemma or the classification. A formalization would provide the general criteria as reusable library results and the first verified stability classification of a non-Markovian queue's imbedded chain.

Difficulty

The matrix is explicit, but none of the three properties is a finite computation: ergodicity and recurrence are statements about return times over all horizons, so each direction must go through an existence or nonexistence statement about infinite systems of equations. For the converse directions the obvious argument fails: exhibiting a candidate solution such as xi≡1x_i \equiv 1xi​≡1 shows nothing until it is known that ergodicity forces every such solution to be summable, and showing that no bounded nonconstant solution of (7) exists when ρ<1\rho < 1ρ<1 requires control of all solutions, not of one. The general criteria themselves rest on limit theorems for pij(n)p_{ij}^{(n)}pij(n)​ and on interchanging infinite sums, and the infinite-mean case ∑nan=∞\sum n a_n = \infty∑nan​=∞ has to be carried along everywhere.

Formalization scope

  • The Markov-chain vocabulary is the published definition QueueingFundamentals_Foundations_MarkovChain: TransitionMatrix (entries p, nonnegativity, rows summing to 111 via HasSum), returnProb, meanRecurrenceTime, Irreducible, Aperiodic, PositiveRecurrent. "Ergodic" is PositiveRecurrent. IsRecurrent and IsTransient are defined state by state from returnProb; their complementarity for irreducible chains is a theorem, not a definition.
  • States are indexed from 000, as in the paper. The goal quantifies over every TransitionMatrix whose entries equal gim1Matrix a; such a matrix exists for every admissible aaa (rows sum to 111), so the statement is not vacuous.
  • ρ−1\rho^{-1}ρ−1 and ρ\rhoρ live in [0,∞][0, \infty][0,∞] (ℝ≥0∞), with ∞−1=0\infty^{-1} = 0∞−1=0: an infinite mean gives ρ=0\rho = 0ρ=0, and that chain is ergodic.
  • The goal does not assume irreducibility or aperiodicity: they follow from an>0a_n > 0an​>0. Milestones 2–4 carry them, as the paper's standing assumptions (§1).
  • Every infinite series appearing in a hypothesis is required to converge (HasSum or Summable), so that a divergent series cannot satisfy an equation or inequality vacuously. In Theorem 1's sufficiency half the xix_ixi​ may be of either sign. In Theorem 7 the distribution is renamed qqq to avoid a clash with pijp_{ij}pij​.
  • Ruled out as trivializing: defining ρ\rhoρ by a real inverse of a real series, defining "ergodic" as the existence of a summable invariant vector (which is Theorem 1's condition), or stating the goal over a matrix that need not exist.
  • Not included: the paper's explicit description of the solutions of (7) for ρ≥1\rho \ge 1ρ≥1 via the generating function (1−z){A(z)−z}−1(1 - z)\{A(z) - z\}^{-1}(1−z){A(z)−z}−1, and the M/G/1 half (Theorems 2, 3, 5), which is the companion mission. Contributions welcome: proofs of the general criteria (reusable for any countable chain), of Theorem 7, and lemmas on the GI/M/1 matrix such as irreducibility and aperiodicity.

Selected references

  • F. G. Foster, On the stochastic matrices associated with certain queuing processes, Ann. Math. Statist. 24 (1953), 355–360. https://doi.org/10.1214/aoms/1177728976
  • D. G. Kendall, Stochastic processes occurring in the theory of queues and their analysis by the method of the imbedded Markov chain, Ann. Math. Statist. 24 (1953), 338–354 (the paper immediately preceding Foster's in the same issue).
  • D. G. Kendall, Some problems in the theory of queues, J. Roy. Statist. Soc. B 13 (1951), 151–185.
  • W. Feller, An Introduction to Probability Theory and Its Applications, Vol. 1, Wiley, 1950.
8 thms1 active userReviewed
CombinatoricsOperations ResearchProbability+1·Captain: mikedeng1

Reversibility and Stochastic Networks VI: The Ewens Sampling Distribution Is Consistent Under Sampling Without ReplacementTextbook

Motivation

The neutral theory of molecular evolution holds that much of the genetic variation observed at the molecular level is caused by selectively neutral mutations rather than by selection. To test it against data one needs the distribution of allele frequencies that a neutral model predicts, and in practice that distribution has to be compared with a sample from the population, never with the whole population. Ewens (Ewens 1972) derived the equilibrium distribution of allele counts under the infinite alleles model, now called the Ewens sampling formula; it underlies classical tests of neutrality and appears throughout combinatorics and probability as the law of the cycle type of an Ewens-distributed random permutation and of the Chinese restaurant process.

Chapter 7 of F. P. Kelly, Reversibility and Stochastic Networks (Wiley, 1979) obtains the infinite alleles model as a limit of the reversible migration processes of Chapters 2 and 6, and uses reversibility to answer questions about allele ages and fixation. The mission formalizes the finite, combinatorial results of that chapter.

Timeline. Kimura and Crow (1964) introduced the infinite alleles model. Ewens (1972) found its equilibrium sampling distribution (7.6). Kingman (1978, J. London Math. Soc.) characterized the consistency of random partitions under sampling, the property Theorem 7.1 asserts for the Ewens family. Kelly (1979, Chapter 7) derived (7.6) as a limit of reversible migration processes, and the consistency and the allele-age results from the reversibility of a labelled population process.

Setting

A population consists of M≥2M\ge2M≥2 individuals, each carrying an allelic type. Its description is M=(M1,…,MM)\mathbf M=(M_1,\dots,M_M)M=(M1​,…,MM​), where MiM_iMi​ is the number of allelic types carried by exactly iii individuals, so that

∑i=1MiMi=M.(7.3)\sum_{i=1}^{M} iM_i=M. \qquad (7.3)i=1∑M​iMi​=M.(7.3)

For a real parameter ν>0\nu>0ν>0, the Ewens distribution on descriptions is

πM(M)=(ν+M−1M)−1∏i=1M(νi)Mi1Mi!,(7.6)\pi_M(\mathbf M)=\binom{\nu+M-1}{M}^{-1}\prod_{i=1}^{M}\Big(\frac{\nu}{i}\Big)^{M_i}\frac{1}{M_i!}, \qquad (7.6)πM​(M)=(Mν+M−1​)−1i=1∏M​(iν​)Mi​Mi​!1​,(7.6)

where (xk)=x(x−1)⋯(x−k+1)/k!\binom{x}{k}=x(x-1)\cdots(x-k+1)/k!(kx​)=x(x−1)⋯(x−k+1)/k! is the binomial coefficient for real xxx. In the infinite alleles model, individuals die at rate μ\muμ, each death is followed by the birth of an offspring of a uniformly chosen survivor, and the offspring is a mutant of an entirely new type with probability uuu; then (7.6) is the equilibrium distribution with ν=(M−1)u/(1−u)\nu=(M-1)u/(1-u)ν=(M−1)u/(1−u) (7.5).

A random sample of size 1≤m≤M1\le m\le M1≤m≤M without replacement is a uniformly random mmm-element subset of the MMM labelled individuals, each of the (Mm)\binom Mm(mM​) subsets being equally likely; the sample has a description in the same sense.

The number jjj of individuals carrying one given allele performs a random walk on {0,…,M}\{0,\dots,M\}{0,…,M} with intensities

q(j,j−1)=μjM(M−jM−1+j−1M−1u),q(j,j+1)=μM−jMjM−1(1−u).(7.8)q(j,j-1)=\mu\frac jM\Big(\frac{M-j}{M-1}+\frac{j-1}{M-1}u\Big),\qquad q(j,j+1)=\mu\frac{M-j}{M}\frac{j}{M-1}(1-u). \qquad (7.8)q(j,j−1)=μMj​(M−1M−j​+M−1j−1​u),q(j,j+1)=μMM−j​M−1j​(1−u).(7.8)

An allele is quasi-fixed when it is the only allele present (j=Mj=Mj=M).

Formalization targets

Goal: consistency under sampling (Theorem 7.1)

If M≥2M\ge2M≥2 and the population description is distributed as πM\pi_MπM​, then a random sample of size 1≤m≤M1\le m\le M1≤m≤M drawn without replacement has description m\mathbf mm with probability πm(m)\pi_m(\mathbf m)πm​(m), the same ν\nuν being used for both sizes:

∑MπM(M) P(sample has description m∣population has description M)=πm(m).\sum_{\mathbf M}\pi_M(\mathbf M)\,P\big(\text{sample has description }\mathbf m\mid\text{population has description }\mathbf M\big)=\pi_m(\mathbf m).M∑​πM​(M)P(sample has description m∣population has description M)=πm​(m).

Milestones

  1. (7.6) is a distribution: πM(M)>0\pi_M(\mathbf M)>0πM​(M)>0 and ∑MπM(M)=1\sum_{\mathbf M}\pi_M(\mathbf M)=1∑M​πM​(M)=1 (Exercise 7.1.3).
  2. Theorem 7.1 for m=M−1m=M-1m=M−1, the case the book's proof establishes first.
  3. Corollary 7.5, the identity of its proof: the probability that a uniformly chosen individual's allele is carried by exactly iii individuals is
∑MiMiMπM(M)=νM(ν+M−1i)−1(Mi).(7.9)\sum_{\mathbf M}\frac{iM_i}{M}\pi_M(\mathbf M)=\frac{\nu}{M}\binom{\nu+M-1}{i}^{-1}\binom Mi. \qquad (7.9)M∑​MiMi​​πM​(M)=Mν​(iν+M−1​)−1(iM​).(7.9)
  1. Theorem 7.9: the probability QQQ that the walk (7.8) started at 111 reaches MMM before 000 satisfies
Q−1=∑i=0M−1(M−1i)−1(ν+M−1i).Q^{-1}=\sum_{i=0}^{M-1}\binom{M-1}{i}^{-1}\binom{\nu+M-1}{i}.Q−1=i=0∑M−1​(iM−1​)−1(iν+M−1​).

Significance

The results. Consistency under sampling is what makes the Ewens formula usable as a statistical model: the predicted distribution for an observed sample does not depend on the unknown population size, only on ν\nuν. Kelly deduces from it the sufficiency of the number of alleles in a sample for ν\nuν and the heterozygosity ν/(ν+1)\nu/(\nu+1)ν/(ν+1) (Exercises 7.1.5, 7.1.8). The formula (7.9) gives the equilibrium frequency of the oldest allele, and Theorem 7.9 gives the quasi-fixation probability from which the mean time between quasi-fixations follows (Corollary 7.10).

Formalizing them. All four results are classical and proved; none has a machine-checked proof on the platform or in Mathlib as of this writing. The mission produces a reusable formal Ewens distribution over integer partitions, a definition of sampling without replacement by counting labelled subsets, and an absorption probability for an explicit birth–death walk. Proofs independent of Kelly's process argument are welcome.

Difficulty

The book's proof of Theorem 7.1 is a process argument: in a population whose size fluctuates between M−1M-1M−1 and MMM, a drop in size acts as a random deletion, and the truncated equilibrium (7.7) restricted to each size gives πM−1\pi_{M-1}πM−1​ and πM\pi_MπM​. Turning that into a statement about finite sets requires the equilibrium of a truncated reversible process, which is not available here, so a formal proof must either build that process or find a direct combinatorial route. A direct route has to relate, for each description of the sample, the number of mmm-subsets of a labelled population with a given description to products of binomial coefficients, and sum the result against (7.6); the bookkeeping over partitions is where the work lies. Theorem 7.9 needs a solution of the first-step equations of a non-symmetric walk and the identification of that solution with a hitting probability defined as a limit.

Formalization scope

  • Descriptions of nnn individuals are integer partitions Nat.Partition n, with MiM_iMi​ the multiplicity of the part iii; the product in (7.6) runs over i=1,…,ni=1,\dots,ni=1,…,n. The real binomial coefficient is the published definition AppliedComb.GenFun.binomReal.
  • The population is Fin M with allelic types Fin M → ℕ; the description of a labelled set is computed from the labelling. The sampling probability is (Mm)−1\binom Mm^{-1}(mM​)−1 times the number of mmm-subsets whose restricted labelling has the given description. It is not defined by a formula on descriptions, and a definition that removed individuals one at a time in proportion to class sizes (the book's proof route) is ruled out as a definition because it presupposes the reduction the proof must supply.
  • The goal and Corollary 7.5 quantify over an arbitrary choice of labelling for each population description. They assume M≥2M\ge2M≥2, as required by the chapter's rule that a parent is chosen among the other M−1M-1M−1 individuals; the goal also assumes 1≤m≤M1\le m\le M1≤m≤M. Because πM>0\pi_M>0πM​>0, this forces the conditional sampling law to depend on the population only through its description. Types are natural numbers, so every description is realized and the hypothesis is never vacuous.
  • The quasi-fixation probability is defined through the jump chain of (7.8): the limit of the probabilities of reaching MMM within nnn jumps without reaching 000. The theorem assumes M≥2M\ge2M≥2, μ>0\mu>0μ>0, 0<u<10<u<10<u<1 and ν=(M−1)u/(1−u)\nu=(M-1)u/(1-u)ν=(M−1)u/(1−u).
  • Corollary 7.5 is formalized as the identity of its proof. The identification of the oldest allele's frequency with that of a randomly chosen individual uses allele ages and the reversibility of the labelled process (Theorem 7.2) and is not formalized. Theorem 7.2 itself, whose state space orders the allele labels within each class, and the allele-age results (Corollaries 7.3, 7.4, 7.7, 7.8, Theorem 7.6, Corollary 7.10, Theorem 7.11) are not part of the mission.

Contributions of general partition and sampling lemmas (counting subsets with a given description, the generating function identity (1−x)−ν=∏jeνxj/j(1-x)^{-\nu}=\prod_j e^{\nu x^j/j}(1−x)−ν=∏j​eνxj/j) are reusable beyond this mission.

Selected references

  • F. P. Kelly, Reversibility and Stochastic Networks, Wiley, 1979, Chapter 7. https://www.statslab.cam.ac.uk/~frank/BOOKS/kelly_book.html
  • W. J. Ewens, The sampling theory of selectively neutral alleles, Theoretical Population Biology 3 (1972), 87–112. https://doi.org/10.1016/0040-5809(72)90035-4
  • J. F. C. Kingman, The representation of partition structures, Journal of the London Mathematical Society (2) 18 (1978), 374–380. https://doi.org/10.1112/jlms/s2-18.2.374
  • M. Kimura and J. F. Crow, The number of alleles that can be maintained in a finite population, Genetics 49 (1964), 725–738. https://doi.org/10.1093/genetics/49.4.725
9 thms1 active userReviewed
Operations ResearchProbabilityStochastic Systems·Captain: mikedeng1

On the Stochastic Matrices Associated with Certain Queuing Processes 1: The M/G/1 Imbedded Chain Is Ergodic iff ρ < 1 and Recurrent iff ρ ≤ 1Research Paper

Motivation

Many queues observed at well-chosen instants are Markov chains on the nonnegative integers. For the single-server queue with Poisson arrivals and general service times (M/G/1), D. G. Kendall showed in 1951 that the number of customers left behind at successive departure epochs is such a chain, the imbedded Markov chain (Kendall 1951; Kendall 1953). Whether the queue settles into a steady state, keeps returning to empty without settling, or grows without bound is then a question about this chain: is it ergodic, null recurrent, or transient?

F. G. Foster's 1953 paper (doi:10.1214/aoms/1177728976) answers this question by first proving general criteria for an irreducible chain on {0,1,2,… }\{0,1,2,\dots\}{0,1,2,…}, stated as solvability conditions for linear inequalities in the transition matrix, and then applying them to the M/G/1 and GI/M/1 chains. Theorem 2 of the paper is the drift condition now known as Foster's criterion, the starting point of the Lyapunov-function method for the stability of queues and stochastic networks (Meyn and Tweedie 2009). This mission is the M/G/1 half of the paper.

Timeline:

  • 1951–1953, Kendall. Introduces the imbedded chains of M/G/1 and GI/M/1 and obtains most of their classification by direct methods.
  • 1953, Foster. Derives the classification from general criteria: Theorem 2 (ergodicity), Theorems 4–6 (transience and recurrence).
  • 1950s onward. The criteria become the standard tools (Feller's text; later the drift conditions of Meyn and Tweedie).

Setting

A Markov chain on the states {0,1,2,… }\{0,1,2,\dots\}{0,1,2,…} is given by a transition matrix P=[pij]P=[p_{ij}]P=[pij​]: pij≥0p_{ij}\ge0pij​≥0 and ∑jpij=1\sum_j p_{ij}=1∑j​pij​=1 for every row iii. Write fij(n)f_{ij}^{(n)}fij(n)​ for the probability that the chain started in iii first reaches jjj (for i=ji=ji=j, first returns to jjj) at step n≥1n\ge1n≥1. The chain is irreducible if every state can be reached from every other, and aperiodic if for every state the return times have greatest common divisor 111. A state jjj is recurrent if fjj=∑nfjj(n)=1f_{jj}=\sum_n f_{jj}^{(n)}=1fjj​=∑n​fjj(n)​=1 and transient if fjj<1f_{jj}<1fjj​<1; a recurrent state is ergodic (positive recurrent, "recurrent-nonnull") if in addition its mean recurrence time ∑nnfjj(n)\sum_n n f_{jj}^{(n)}∑n​nfjj(n)​ is finite. The mean first-passage time from iii to jjj is μij=∑n≥1nfij(n)∈[0,∞]\mu_{ij}=\sum_{n\ge1} n f_{ij}^{(n)}\in[0,\infty]μij​=∑n≥1​nfij(n)​∈[0,∞].

The M/G/1 matrix is built from a sequence k0,k1,…k_0,k_1,\dotsk0​,k1​,… of positive numbers summing to one (knk_nkn​ is the probability of nnn arrivals during one service):

[pij]=[k0k1k2⋯k0k1k2⋯0k0k1⋯00k0⋯⋮⋮⋮],[p_{ij}] = \begin{bmatrix} k_0 & k_1 & k_2 & \cdots \\ k_0 & k_1 & k_2 & \cdots \\ 0 & k_0 & k_1 & \cdots \\ 0 & 0 & k_0 & \cdots \\ \vdots & \vdots & \vdots & \end{bmatrix},[pij​]=​k0​k0​00⋮​k1​k1​k0​0⋮​k2​k2​k1​k0​⋮​⋯⋯⋯⋯​​,

that is, p0j=kjp_{0j}=k_jp0j​=kj​ and, for i≥1i\ge1i≥1, pij=kj−i+1p_{ij}=k_{j-i+1}pij​=kj−i+1​ when j≥i−1j\ge i-1j≥i−1 and 000 otherwise. The traffic intensity is

ρ=∑n=1∞n kn∈[0,∞],\rho=\sum_{n=1}^{\infty}n\,k_n\in[0,\infty],ρ=n=1∑∞​nkn​∈[0,∞],

the mean number of arrivals per service.

Formalization targets

Goal: the M/G/1 classification (§3, p. 358)

the chain is ergodic  ⟺  ρ<1,the chain is recurrent  ⟺  ρ≤1.\text{the chain is ergodic}\iff\rho<1,\qquad\text{the chain is recurrent}\iff\rho\le1 .the chain is ergodic⟺ρ<1,the chain is recurrent⟺ρ≤1.

The goal leaves kkk arbitrary apart from positivity and normalization; in particular ρ=∞\rho=\inftyρ=∞ is allowed and falls in the transient case.

Milestones (the paper's general theorems and the step of §3 they feed)

  1. Theorem 2 (drift criterion): a nonnegative solution of ∑jpijyj≤yi−1\sum_j p_{ij}y_j\le y_i-1∑j​pij​yj​≤yi​−1 (i≠0i\ne0i=0) with ∑jp0jyj<∞\sum_j p_{0j}y_j<\infty∑j​p0j​yj​<∞ makes the system ergodic. Already posed on the platform and referenced here.
  2. Theorem 3: in an ergodic system the mean first-passage times dj=μj0d_j=\mu_{j0}dj​=μj0​ are finite and satisfy ∑j≥1pijdj=di−1\sum_{j\ge1}p_{ij}d_j=d_i-1∑j≥1​pij​dj​=di​−1 (i≠0i\ne0i=0), ∑j≥1p0jdj<∞\sum_{j\ge1}p_{0j}d_j<\infty∑j≥1​p0j​dj​<∞.
  3. §3 display: for the ergodic M/G/1 chain, μi,i−1=μ10\mu_{i,i-1}=\mu_{10}μi,i−1​=μ10​ and μi0=iμ10\mu_{i0}=i\mu_{10}μi0​=iμ10​ (i≠0i\ne0i=0).
  4. Theorem 5: a solution of ∑jpijyj≤yi\sum_j p_{ij}y_j\le y_i∑j​pij​yj​≤yi​ (i≠0i\ne0i=0) with yi→∞y_i\to\inftyyi​→∞ makes the system recurrent.
  5. Theorem 7: for a probability distribution {pn}\{p_n\}{pn​} with p0>0p_0>0p0​>0, ∑nznpn=z\sum_n z^np_n=z∑n​znpn​=z has a root in (0,1)(0,1)(0,1) iff ∑n≥1npn>1\sum_{n\ge1}np_n>1∑n≥1​npn​>1.
  6. Theorem 4: the system is transient iff ∑jpijyj=yi\sum_j p_{ij}y_j=y_i∑j​pij​yj​=yi​ (i≠0i\ne0i=0) has a bounded nonconstant solution.

Significance

The result. The classification is the stability theorem for the M/G/1 queue: for ρ<1\rho<1ρ<1 the departure-epoch queue length has a stationary distribution, which is what the Pollaczek–Khinchine formula describes; for ρ=1\rho=1ρ=1 the queue empties infinitely often but has no steady state; for ρ>1\rho>1ρ>1 it grows without bound. The general criteria behind it (Theorems 2, 4, 5) apply to any chain on the nonnegative integers and are reused in the companion GI/M/1 mission and throughout queueing and Markov-chain stability theory.

Formalizing it. All results here are proved on paper (Kendall and Foster, 1951–1953, with Theorems 3 and 7 classical lemmas from Feller). None of them is known to have a machine-checked proof against a Lean development of countable-state Markov chains. The mission produces such proofs on the published discrete-chain vocabulary (transition matrices, first-passage probabilities, return probabilities, positive recurrence), together with the general Foster criteria as reusable theorems. Theorem 2 is already posed as an open platform theorem and is reused here.

Difficulty

The queue-specific part of the argument is short once the general criteria are available; the weight of the mission is in those criteria. They relate qualitative properties of an infinite chain (ergodicity, recurrence, transience) to solvability of infinite systems of linear inequalities, and this needs limit behaviour of the nnn-step probabilities pij(n)p_{ij}^{(n)}pij(n)​ and of hitting probabilities of state 000, none of which follows from finite-state arguments. Two further points resist the naive approach. The converse directions (ergodic ⇒ρ<1\Rightarrow\rho<1⇒ρ<1, recurrent ⇒ρ≤1\Rightarrow\rho\le1⇒ρ≤1) need exact identities for mean first-passage times, not just bounds, and these must be handled in [0,∞][0,\infty][0,∞] because the means may be infinite. And the boundary case ρ=1\rho=1ρ=1 (null recurrence) separates the two equivalences: an argument that only compares the mean drift ρ−1\rho-1ρ−1 with 000, such as a law of large numbers for the increments, cannot tell recurrence from transience there.

Formalization scope

  • The chain is the published QueueingFundamentals.Foundations.TransitionMatrix (entries P.p i j, rows summing to 111 as a HasSum), with its firstPassage, returnProb, Irreducible, Aperiodic and PositiveRecurrent. "Ergodic" is P.PositiveRecurrent; aperiodicity is the paper's standing assumption and is not folded into it a second time.
  • States are indexed from 000, as in the paper; "i≠0i\ne0i=0" is i ≠ 0.
  • The M/G/1 matrix is a function mg1Matrix k : ℕ → ℕ → ℝ; the goal and the §3 display quantify over every TransitionMatrix P with P.p = mg1Matrix k. Such a P exists for every admissible k (checked in a sorry-free local file for ki=2−(i+1)k_i=2^{-(i+1)}ki​=2−(i+1)).
  • ∑nkn=1\sum_n k_n=1∑n​kn​=1 is added as the meaning of "stochastic matrix"; §3 writes only ki>0k_i>0ki​>0.
  • ρ\rhoρ and all mean first-passage times are extended nonnegative reals ([0,∞][0,\infty][0,∞]), so divergent means are ∞\infty∞, never 000. Theorem 7's mean is also taken in [0,∞][0,\infty][0,∞].
  • Recurrent means fjj=1f_{jj}=1fjj​=1 for every state jjj; transient means fjj<1f_{jj}<1fjj​<1 for every state. For irreducible chains these are complementary, which is a theorem, not a definition.
  • The general Theorems 3, 4 and 5 assume irreducibility and aperiodicity, the paper's standing assumption of §1. The goal does not assume them: they follow from ki>0k_i>0ki​>0.
  • Every series in a hypothesis carries its convergence (Summable or HasSum); Theorem 3's equation (6) is written as di=1+∑j≥1pijdjd_i=1+\sum_{j\ge1}p_{ij}d_jdi​=1+∑j≥1​pij​dj​ in [0,∞][0,\infty][0,∞] together with finiteness of the djd_jdj​, j≠0j\ne0j=0.
  • Theorem 7's distribution is renamed qqq in Lean to avoid a clash with pijp_{ij}pij​. Theorem 1 of the paper (§2) and Theorem 6 are not targets of this mission.

Ruled out: ρ\rhoρ as a real tsum (which is 000 for a divergent series and would call a heavy-tailed chain ergodic); defining "ergodic" or "recurrent" through the existence of Lyapunov or drift functions (which would make the criteria tautological); a goal over a matrix PPP that need not exist.

Contributions welcome: proofs of the general criteria (Theorems 2–5) on the published chain vocabulary, the limit theorem pij(n)→πjp_{ij}^{(n)}\to\pi_jpij(n)​→πj​ for irreducible aperiodic chains, first-step analysis for hitting times, and Theorem 7 as a lemma on probability generating functions; all of these are reusable beyond this mission.

Selected references

  • F. G. Foster, On the stochastic matrices associated with certain queuing processes, The Annals of Mathematical Statistics 24(3), 355–360, 1953. https://doi.org/10.1214/aoms/1177728976
  • D. G. Kendall, Some problems in the theory of queues, Journal of the Royal Statistical Society B 13(2), 151–185, 1951. https://doi.org/10.1111/j.2517-6161.1951.tb00093.x
  • D. G. Kendall, Stochastic processes occurring in the theory of queues and their analysis by the method of the imbedded Markov chain, The Annals of Mathematical Statistics 24(3), 338–354, 1953. https://doi.org/10.1214/aoms/1177728975
  • W. Feller, An Introduction to Probability Theory and Its Applications, Vol. I, Wiley, 1950.
  • S. Meyn and R. L. Tweedie, Markov Chains and Stochastic Stability, 2nd ed., Cambridge University Press, 2009. https://doi.org/10.1017/CBO9780511626630
9 thms1 active userReviewed
ProbabilityReinforcement Learning·Captain: mikedeng1

Linear Least-Squares Algorithms for Temporal Difference Learning II: Probability-One Convergence of LS TD on Ergodic Markov ChainsResearch Paper

Motivation

Temporal-difference (TD) learning estimates the value function of a Markov chain — the expected discounted sum of future rewards from each state — from a single stream of observed transitions, without knowing the transition probabilities. With a linear function approximator the value of state xxx is represented as ϕx′θ\phi_x'\thetaϕx′​θ for a feature vector ϕx\phi_xϕx​ and a parameter θ\thetaθ. Classical TD(λ\lambdaλ) updates θ\thetaθ by stochastic approximation, and its behaviour depends on a step-size schedule that must be tuned.

Bradtke and Barto (Machine Learning 22, 1996) replaced the stochastic-approximation update by a least-squares solve: LS TD (Eq. (11)) recomputes θt\theta_tθt​ at every step as the instrumental-variable least-squares solution of the empirical consistency condition. The method, later generalized as LSTD(λ\lambdaλ) by Boyan (Machine Learning 49, 2002), is the basis of least-squares policy iteration and of the "LSTD" methods in standard reinforcement-learning texts (Sutton and Barto, Reinforcement Learning, 2nd ed., 2018, §9.8). Its appeal is that it has no step size; the question this mission formalizes is whether it nonetheless converges, with probability one, to the true parameter.

Timeline. Sutton (1988) introduced TD(λ\lambdaλ). Watkins and Dayan (1992) and Tsitsiklis (1994) proved probability-one convergence of tabular TD(0) and Q-learning. Bradtke and Barto (1996) proved probability-one convergence of LS TD on absorbing chains (Theorem 1) and on ergodic chains (Theorem 2). Tsitsiklis and Van Roy (IEEE TAC 42, 1997) proved convergence of linear TD(λ\lambdaλ) with general features on ergodic chains.

Setting

A finite Markov chain on a finite nonempty set XXX is a matrix PPP with P(x,y)≥0P(x,y)\ge0P(x,y)≥0 and ∑yP(x,y)=1\sum_yP(x,y)=1∑y​P(x,y)=1. A transition x→yx\to yx→y earns reward R(x,y)R(x,y)R(x,y); the expected reward out of xxx is rˉx=∑yP(x,y)R(x,y)\bar r_x=\sum_yP(x,y)R(x,y)rˉx​=∑y​P(x,y)R(x,y). For a discount factor γ\gammaγ the value function is

V(x)=E{∑k=0∞γkrk ∣ x0=x}=∑k=0∞γk(Pkrˉ)(x).V(x)=E\Big\{\sum_{k=0}^\infty\gamma^kr_k\ \Big|\ x_0=x\Big\}=\sum_{k=0}^\infty\gamma^k(P^k\bar r)(x).V(x)=E{k=0∑∞​γkrk​ ​ x0​=x}=k=0∑∞​γk(Pkrˉ)(x).

The chain is ergodic (Kemeny and Snell) if every state can be reached from every state: for all x,yx,yx,y there is nnn with Pn(x,y)>0P^n(x,y)>0Pn(x,y)>0. An invariant distribution is a probability vector π\piπ with πP=π\pi P=\piπP=π; write Π=diag⁡(π)\Pi=\operatorname{diag}(\pi)Π=diag(π).

Each state has a feature vector ϕx∈Rm\phi_x\in\mathbb R^mϕx​∈Rm; Φ\PhiΦ is the matrix with rows ϕx\phi_xϕx​. The true parameter θ∗\theta^*θ∗ is a vector with V(x)=ϕx′θ∗V(x)=\phi_x'\theta^*V(x)=ϕx′​θ∗ for all xxx.

The algorithm (Figure 3) starts at an arbitrary state x0x_0x0​, lets the chain move x0→x1→⋯x_0\to x_1\to\cdotsx0​→x1​→⋯, and after ttt transitions computes

θt=[1t∑kϕxk(ϕxk−γϕxk+1)′]−1[1t∑kϕxkR(xk,xk+1)],(11)\theta_t=\Big[\frac1t\sum_{k}\phi_{x_k}(\phi_{x_k}-\gamma\phi_{x_{k+1}})'\Big]^{-1}\Big[\frac1t\sum_k\phi_{x_k}R(x_k,x_{k+1})\Big],\tag{11}θt​=[t1​k∑​ϕxk​​(ϕxk​​−γϕxk+1​​)′]−1[t1​k∑​ϕxk​​R(xk​,xk+1​)],(11)

the sums running over the ttt transitions observed so far.

Formalization targets

Goal: Theorem 2 (p. 44)

If PPP is ergodic, (1) {ϕx}\{\phi_x\}{ϕx​} is linearly independent, (2) each ϕx\phi_xϕx​ has dimension ∣X∣|X|∣X∣, and (3) 0<γ<10<\gamma<10<γ<1, then θ∗\theta^*θ∗ is finite and, from any initial law,

θt⟶θ∗with probability 1.\theta_t\longrightarrow\theta^*\qquad\text{with probability }1 .θt​⟶θ∗with probability 1.

The goal leaves the chain, the rewards, the features and the initial law arbitrary.

Milestones, in the order of the proof

  1. Visit frequencies (Proof of Theorem 2, p. 45): an ergodic chain visits every state infinitely often and #{k<t:xk=x}/t→πx\#\{k<t:x_k=x\}/t\to\pi_x#{k<t:xk​=x}/t→πx​ almost surely.
  2. Invertibility (Proof of Theorem 2, p. 45): πx>0\pi_x>0πx​>0 for all xxx, and Φ′Π(I−γP)Φ\Phi'\Pi(I-\gamma P)\PhiΦ′Π(I−γP)Φ is invertible.
  3. The pathwise limit (Proof of Lemma 5, pp. 54–55): along any path whose transition frequencies converge to πxP(x,y)\pi_xP(x,y)πx​P(x,y), θt→[Φ′Π(I−γP)Φ]−1[Φ′Πrˉ]\theta_t\to[\Phi'\Pi(I-\gamma P)\Phi]^{-1}[\Phi'\Pi\bar r]θt​→[Φ′Π(I−γP)Φ]−1[Φ′Πrˉ].
  4. Lemma 5 (p. 43): for any chain, if almost surely every state is visited infinitely often and in proportion π\piπ, and Φ′Π(I−γP)Φ\Phi'\Pi(I-\gamma P)\PhiΦ′Π(I−γP)Φ is invertible, then θt→[Φ′Π(I−γP)Φ]−1[Φ′Πrˉ]\theta_t\to[\Phi'\Pi(I-\gamma P)\Phi]^{-1}[\Phi'\Pi\bar r]θt​→[Φ′Π(I−γP)Φ]−1[Φ′Πrˉ] almost surely.
  5. Eq. (12) (p. 44): the value series converges and rˉ=(I−γP)Φθ∗\bar r=(I-\gamma P)\Phi\theta^*rˉ=(I−γP)Φθ∗.

Significance

The result. Theorem 2 shows that an LS TD learner running on one long trajectory recovers the exact value function whenever the features can represent every function on the states, with no step-size schedule. It is the step-size-free counterpart of the tabular TD(0) convergence theorems and the starting point for the later analysis of LSTD with fewer features than states, where the limit is the TD fixed point [Φ′Π(I−γP)Φ]−1Φ′Πrˉ[\Phi'\Pi(I-\gamma P)\Phi]^{-1}\Phi'\Pi\bar r[Φ′Π(I−γP)Φ]−1Φ′Πrˉ rather than θ∗\theta^*θ∗. Lemma 5 is the general identification of that fixed point as the almost-sure limit of LSTD.

Formalizing it. The result is proved in the paper; it has not been machine-checked. A formalization adds three things the paper delegates: the strong law of large numbers for occupation times of a finite irreducible Markov chain, which the paper cites to Kemeny and Snell and which is not in Mathlib; the per-state transition frequencies used in the first sentence of the proof of Lemma 5; and the linear algebra of the limit. The first is reusable well beyond reinforcement learning.

Difficulty

The algebra is short once the empirical averages in (11) are known to converge. The difficulty is probabilistic: the averages are over a dependent sequence, so the ordinary strong law of large numbers does not apply. Two facts are needed: that the fraction of time in each state converges to πx\pi_xπx​ almost surely for any starting law, including periodic chains, where PnP^nPn itself does not converge; and that, among the visits to xxx, the fraction followed by a move to yyy converges to P(x,y)P(x,y)P(x,y), which needs the strong Markov property at successive visit times. Neither follows from convergence of the chain's distribution, and neither holds for a chain started at a fixed state without an argument that every state is reached.

Formalization scope

  • The model is a finite state type X with Fintype, DecidableEq, Nonempty, a row-stochastic matrix P : Matrix X X ℝ (structure Chain), rewards R : X → X → ℝ, features φ : X → Fin m → ℝ. Condition (2) is m = Fintype.card X; condition (1) is LinearIndependent ℝ φ.
  • "Ergodic" is read as irreducible, periodic chains allowed (Kemeny–Snell's aperiodic case is "regular"). "Arbitrary initial state" is read as every initial law ν\nuν, which contains every point mass.
  • The path is any process ZZZ on any probability space whose finite-dimensional distributions are ν(x0)P(x0,x1)⋯P(xn−1,xn)\nu(x_0)P(x_0,x_1)\cdots P(x_{n-1},x_n)ν(x0​)P(x0​,x1​)⋯P(xn−1​,xn​), with measurable events {Zt=x}\{Z_t=x\}{Zt​=x}.
  • VVV is the discounted series, never (I−γP)−1rˉ(I-\gamma P)^{-1}\bar r(I−γP)−1rˉ; Lean's tsum is 000 on a divergent series, so "θ* is finite" is stated as convergence of the series together with existence of θ∗\theta^*θ∗ with V=Φθ∗V=\Phi\theta^*V=Φθ∗. θ∗\theta^*θ∗ is existential, never defined as Lemma 5's limit.
  • (11) uses the transitions k=0,…,t−1k=0,\dots,t-1k=0,…,t−1 (the paper prints k=1,…,tk=1,\dots,tk=1,…,t with ϕt+1\phi_{t+1}ϕt+1​; an index shift), keeps the factors 1/t1/t1/t, and uses Lean's matrix inverse, which is 000 on a singular matrix: the paper notes θt\theta_tθt​ is undefined for small ttt, and finitely many junk values do not affect convergence. No εI\varepsilon IεI regularization, no pseudo-inverse.
  • θLSTD=lim⁡tθt\theta_{\rm LSTD}=\lim_t\theta_tθLSTD​=limt​θt​ is formalized as convergence of θt\theta_tθt​ (existence of the limit is part of the claim).
  • The convergence is almost sure. A formalization that assumes the visit frequencies converge in the goal, starts the chain from π\piπ, or weakens the conclusion to convergence in probability or along a subsequence is a different theorem.

Needed infrastructure: the strong law for occupation times of a finite irreducible chain under an arbitrary initial law (milestone 1), the strong Markov property at visit times, positivity of the invariant distribution of an irreducible chain, and the invertibility of I−γPI-\gamma PI−γP for ∣γ∣<1|\gamma|<1∣γ∣<1. Contributions to any of these, as standalone lemmas, are welcome.

Selected references

  • S. J. Bradtke and A. G. Barto, Linear Least-Squares Algorithms for Temporal Difference Learning, Machine Learning 22, 33–57, 1996. https://doi.org/10.1023/A:1018056104778
  • J. G. Kemeny and J. L. Snell, Finite Markov Chains, Springer, 1976.
  • J. A. Boyan, Technical Update: Least-Squares Temporal Difference Learning, Machine Learning 49, 233–246, 2002. https://doi.org/10.1023/A:1017936530646
  • J. N. Tsitsiklis and B. Van Roy, An Analysis of Temporal-Difference Learning with Function Approximation, IEEE Transactions on Automatic Control 42(5), 674–690, 1997. https://doi.org/10.1109/9.580874
  • R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018. http://incompleteideas.net/book/the-book-2nd.html
8 thms1 active userReviewed
ProbabilityReinforcement Learning·Captain: mikedeng1

Linear Least-Squares Algorithms for Temporal Difference Learning I: Probability-One Convergence of Trial-Based LS TD on Absorbing Markov ChainsResearch Paper

Motivation

Temporal-difference learning estimates the value of a policy from observed state transitions and rewards. In a finite Markov decision process, fixing a policy produces a Markov chain, so policy evaluation becomes the task of estimating the expected return from each state. Bradtke and Barto's 1996 paper introduced a least-squares temporal-difference method, LS TD, that uses each observed transition in a linear system instead of selecting a learning-rate schedule. Their Theorem 1 states probability-one convergence for trials that end at absorbing states under explicit conditions on state access, rewards, and features. This mission formalizes that result and the statements the authors use to reach it. Bradtke and Barto, 1996.

The result matters for episodic policy evaluation: a learner may collect many short trajectories, each begun from a prescribed start distribution, and update the same estimate as data accumulate. The theorem identifies conditions under which the limit is the true value parameter even when the discount factor is one. That endpoint is useful for undiscounted tasks ending in an absorbing goal state; it also makes the convergence claim more delicate than the standard discounted case. The paper proves the result mathematically. The Lean statements in this mission are targets for machine-checked proofs, not claims of proofs already present in Mathlib. Bradtke and Barto, Theorem 1, pp. 43–44.

Setting

Let XXX be a finite, nonempty set of states. After a policy is fixed, P(x,y)P(x,y)P(x,y) is the probability of a transition from xxx to yyy, so each row of PPP is nonnegative and sums to one. A transition earns a deterministic real reward R(x,y)R(x,y)R(x,y). A state is absorbing when P(x,x)=1P(x,x)=1P(x,x)=1; let T\mathcal TT be the absorbing states and N=X∖T\mathcal N=X\setminus\mathcal TN=X∖T the others. The chain is absorbing when some absorbing state can be reached with positive probability from every state. A start distribution SSS gives the state at the beginning of each trial. No state is inaccessible when every state can be reached from the positive support of SSS.

For a discount γ\gammaγ, the expected immediate reward is rˉ(x)=∑yP(x,y)R(x,y)\bar r(x)=\sum_yP(x,y)R(x,y)rˉ(x)=∑y​P(x,y)R(x,y). The true value function is defined by the expected return

V(x)=∑k=0∞γk(Pkrˉ)(x).V(x)=\sum_{k=0}^{\infty}\gamma^k(P^k\bar r)(x).V(x)=k=0∑∞​γk(Pkrˉ)(x).

A feature vector ϕx∈Rm\phi_x\in\mathbb R^mϕx​∈Rm represents state xxx. The matrix Φ\PhiΦ has row xxx equal to ϕx⊤\phi_x^\topϕx⊤​. The target parameter θ∗\theta^*θ∗ is a vector for which V(x)=ϕx⊤θ∗V(x)=\phi_x^\top\theta^*V(x)=ϕx⊤​θ∗ at every state; it is something the theorem must establish, not an input chosen by a formula. Equation (11) forms an LS TD estimate θn\theta_nθn​ from the observed feature differences and rewards. Bradtke and Barto, §2, Table 1, Eq. (11).

Figure 2 collects trials. Each starts from SSS, follows PPP while the current state is non-absorbing, and ends upon entry into T\mathcal TT. The next trial starts with a fresh draw from SSS. The estimator includes transitions taken within trials; a draw that starts the next trial is not an observed transition for Eq. (11). Bradtke and Barto, Figure 2, p. 42.

Formalization targets

Theorem 1: convergence of trial-based LS TD

If every state is accessible from SSS, rewards between absorbing states vanish, the feature vectors on N\mathcal NN are linearly independent, features on T\mathcal TT are zero, m=∣N∣m=|\mathcal N|m=∣N∣, and 0≤γ≤10\le\gamma\le10≤γ≤1, then the expected-return series converges and there is a parameter θ∗\theta^*θ∗ satisfying

V(x)=ϕx⊤θ∗(x∈X),θn⟶θ∗with probability one.V(x)=\phi_x^\top\theta^*\quad(x\in X),\qquad \theta_n\longrightarrow\theta^*\quad\text{with probability one}.V(x)=ϕx⊤​θ∗(x∈X),θn​⟶θ∗with probability one.

The theorem keeps the paper's endpoint γ=1\gamma=1γ=1. The return series' convergence is explicit because a real infinite sum in Lean has a default value when it diverges. Bradtke and Barto, Theorem 1, p. 43.

Supporting targets

The milestone list follows the statements used in the paper: almost-sure visits and departure proportions for the trials; invertibility of the non-absorbing block of I−γPI-\gamma PI−γP; invertibility of Φ⊤Π(I−γP)Φ\Phi^\top\Pi(I-\gamma P)\PhiΦ⊤Π(I−γP)Φ for positive non-absorbing weights; Lemma 5's probability-one limit [Φ⊤Π(I−γP)Φ]−1Φ⊤Πrˉ[\Phi^\top\Pi(I-\gamma P)\Phi]^{-1}\Phi^\top\Pi\bar r[Φ⊤Π(I−γP)Φ]−1Φ⊤Πrˉ; and Eq. (12), rˉ=(I−γP)Φθ∗\bar r=(I-\gamma P)\Phi\theta^*rˉ=(I−γP)Φθ∗, together with finiteness of the true parameter. Here Π=diag⁡(π)\Pi=\operatorname{diag}(\pi)Π=diag(π). Bradtke and Barto, Lemma 5, p. 43; Proof of Theorem 1, p. 44.

Significance

Theorem 1 identifies the target of the asymptotic LS TD estimate: the value function defined from rewards, rather than merely a vector satisfying a sampled linear system. It covers an undiscounted absorbing chain, where a general fixed-point equation for values would fail to determine the values of absorbing states. The zero-reward and zero-feature conditions determine that boundary correctly. The result also explains the dimension condition: one independent feature vector for each non-absorbing state permits exact representation of the return. Bradtke and Barto, pp. 43–44.

A complete formal development would connect finite-state stochastic-process laws, visit frequencies, matrix limits, and the return-defined value function in one checked statement. The reusable parts include a finite row-stochastic chain model, a path-law description of restarts, a filtered least-squares estimator, and results about transient blocks of stochastic matrices. The paper's mathematical proof exists; this mission asks for formal proofs of its Lean targets. It also leaves room for alternative proofs and sharper, separately stated variants without weakening Theorem 1.

Difficulty

Ordinary matrix convergence cannot be applied until the observed transition frequencies are known to converge and the limiting matrix is invertible. A trial has random length, and the process resets after absorption, so a sequence indexed by all restart-process steps does not have the same raw state proportions as a count indexed by trials. The proof must account for both clocks while retaining the in-trial data of Eq. (11). At γ=1\gamma=1γ=1, a direct geometric-series argument for the value function is unavailable; its finiteness depends on absorption and the reward convention. The matrix I−γPI-\gamma PI−γP itself is singular at the undiscounted endpoint because of absorbing states, while its non-absorbing block is the relevant invertible matrix. Bradtke and Barto, Proof of Theorem 1, p. 44.

Formalization scope

The Lean state type is finite and nonempty. The paper evaluates one fixed policy, so PPP is a real row-stochastic matrix and RRR is a deterministic real reward on transitions; there is no action type in the formal statement. Absorbing states are exactly those with P(x,x)=1P(x,x)=1P(x,x)=1, and “absorbing chain” means that an absorbing state is reachable from every state. The paper does not define “inaccessible”; the formalization reads it as unreachable from the positive support of SSS. The state space carries the discrete measurable structure. Theorem 1's restart process and Lemma 5's ordinary Markov chain are each constrained by their finite-dimensional cylinder probabilities, not by assumed transition frequencies.

The feature space is Rm\mathbb R^mRm, and mmm equals the cardinality of the subtype N\mathcal NN. LS TD uses only departures from N\mathcal NN. Index nnn counts restart-process steps, so the estimate repeats at a restart draw; the paper counts in-trial transitions. These indices have the same asymptotic estimate when transitions continue. The 1/t1/t1/t factors in Eq. (11) cancel, and early singular inverses take Lean's total-inverse default. The value function is the return series, and the goal explicitly asserts its summability. The true parameter is existential, never defined by the formula whose convergence the theorem is meant to prove.

The paper defines πx\pi_xπx​ for absorbing chains as expected departures from xxx per trial. The Theorem 1 visit-frequency milestone normalizes by restart-process steps, which rescales all weights by one positive common factor; Lemma 5's matrix expression is invariant under that rescaling. Lemma 5 itself counts every ordinary-chain transition and carries the paper's “any Markov chain” scope. The milestone on invertibility allows arbitrary weights at absorbing states because their feature rows are zero. These conventions are recorded with each Lean item. Contributions toward the path-law frequency theorem, transient-matrix invertibility, return-series summability, and the matrix limit are all within scope. A vacuous path law or a value function defined from the desired linear equation would not establish the stated goal.

Selected references

  • S. J. Bradtke and A. G. Barto, Linear Least-Squares Algorithms for Temporal Difference Learning, Machine Learning 22, 33–57 (1996). DOI: 10.1023/A:1018056104778.
8 thms1 active userReviewed
Operations ResearchProbabilityStochastic Systems·Captain: mikedeng1

Optimization of Multiclass Queueing Networks: Polyhedral and Nonlinear Characterizations of Achievable Performance I: Quadratic Potential Functions Bound Mean Response Times in Open NetworksResearch Paper

Motivation

Scheduling in a multiclass queueing network asks which waiting job a server should work on next when jobs of several types share stations and revisit them along fixed routes. Such networks model semiconductor wafer fabs, job shops and communication switches. Optimal policies are rarely computable: the state space is countably infinite, and even deciding properties of optimal policies is hard (Papadimitriou and Tsitsiklis 1999). A practical substitute is the achievable region approach: describe, by constraints that every policy must satisfy, a set containing all performance vectors any policy can achieve, then optimize a linear cost over that set to get a lower bound on the optimal cost.

Bertsimas, Paschalidis and Tsitsiklis (MIT Sloan working paper 1992; Ann. Appl. Probab. 1994) gave a general method for producing such constraints for open networks, by computing the steady-state drift of quadratic potential functions. This mission formalizes their first-order bounds (Section 4).

Timeline:

  • 1980–1988: Coffman and Mitrani, then Federgruen and Groenevelt — the achievable performance vectors of a single-station multiclass queue form a polytope described by conservation laws.
  • Early 1990s: Kumar (reference [Kuma] of the paper), using a potential-function argument he attributes to Meyn, derives a single lower bound on the mean number in system for re-entrant lines with deterministic routing (described on p. 16 of the paper).
  • 1992–1994: Bertsimas, Paschalidis and Tsitsiklis — parametric families of linear bounds for general open networks with Markovian routing (Theorem 4.1), and the nonparametric polyhedron (Theorems 4.2–4.4), shown to be at least as tight.

Setting

A network has NNN single-server stations and RRR job classes. Class rrr is served at station σ(r)\sigma(r)σ(r), and CiC_iCi​ is the set of classes served at station iii. Class-rrr jobs arrive from outside as a Poisson stream of rate λ0r\lambda_{0r}λ0r​, service times are exponential with rate μr\mu_rμr​, and after service a class-rrr job becomes a class-sss job with probability prsp_{rs}prs​ or leaves with probability pr0=1−∑sprsp_{r0}=1-\sum_s p_{rs}pr0​=1−∑s​prs​. The traffic equations

λr=λ0r+∑r′λr′pr′r(15)\lambda_r=\lambda_{0r}+\sum_{r'}\lambda_{r'}p_{r'r}\qquad(15)λr​=λ0r​+r′∑​λr′​pr′r​(15)

have a unique solution λ\lambdaλ (the network is open), and ∑r∈Ciλr/μr<1\sum_{r\in C_i}\lambda_r/\mu_r<1∑r∈Ci​​λr​/μr​<1 at every station.

The state n⃗=(n1,…,nR)\vec n=(n_1,\dots,n_R)n=(n1​,…,nR​) counts the jobs of each class. A Markovian policy decides from the current state which classes are in service, at most one per station and only classes with jobs present; idling is allowed. Write BrB_rBr​ for the event that station σ(r)\sigma(r)σ(r) serves class rrr, and B0iB_{0i}B0i​ for the event that station iii is idle. Under such a policy n⃗(t)\vec n(t)n(t) is a continuous-time Markov chain. Assumption A requires that it has a unique invariant distribution π\piπ and that Eπ[nr2]<∞E_\pi[n_r^2]<\inftyEπ​[nr2​]<∞ for all rrr. Let nˉr=Eπ[nr]\bar n_r=E_\pi[n_r]nˉr​=Eπ​[nr​], which equals λrxr\lambda_rx_rλr​xr​ with xrx_rxr​ the mean response time of class rrr (Little's law), and define

Irr′=Eπ[1{Br}nr′],Nir′=Eπ[1{B0i}nr′].I_{rr'}=E_\pi[1\{B_r\}n_{r'}],\qquad N_{ir'}=E_\pi[1\{B_{0i}\}n_{r'}].Irr′​=Eπ​[1{Br​}nr′​],Nir′​=Eπ​[1{B0i​}nr′​].

For a set SSS of classes, f-parameters are reals f(r)≥0f(r)\ge 0f(r)≥0 for r∈Sr\in Sr∈S such that μr[∑r′∈Sprr′(f(r)−f(r′))+∑r′∉Sprr′f(r)]\mu_r\big[\sum_{r'\in S}p_{rr'}(f(r)-f(r'))+\sum_{r'\notin S}p_{rr'}f(r)\big]μr​[∑r′∈S​prr′​(f(r)−f(r′))+∑r′∈/S​prr′​f(r)] is nonnegative and the same for all r∈Ci∩Sr\in C_i\cap Sr∈Ci​∩S; that common value is fif_ifi​, and fi=0f_i=0fi​=0 when Ci∩S=∅C_i\cap S=\emptysetCi​∩S=∅ (restriction (17)). The sums over r′∉Sr'\notin Sr′∈/S include the exit r′=0r'=0r′=0.

Formalization targets

Goal: Theorem 4.1

For every policy satisfying Assumption A, every SSS and every f-parameters satisfying (17),

∑r∈Sλrf(r)xr ≥ N′(S)D′(S),\sum_{r\in S}\lambda_rf(r)x_r\ \ge\ \frac{N'(S)}{D'(S)},r∈S∑​λr​f(r)xr​ ≥ D′(S)N′(S)​,

where

N′(S)=∑r∈Sλ0rf2(r)+∑r∉Sλr∑r′∈Sprr′f2(r′)+∑r∈Sλr[∑r′∈Sprr′(f(r)−f(r′))2+∑r′∉Sprr′f2(r)],N'(S)=\sum_{r\in S}\lambda_{0r}f^2(r)+\sum_{r\notin S}\lambda_r\sum_{r'\in S}p_{rr'}f^2(r')+\sum_{r\in S}\lambda_r\Big[\sum_{r'\in S}p_{rr'}(f(r)-f(r'))^2+\sum_{r'\notin S}p_{rr'}f^2(r)\Big],N′(S)=r∈S∑​λ0r​f2(r)+r∈/S∑​λr​r′∈S∑​prr′​f2(r′)+r∈S∑​λr​[r′∈S∑​prr′​(f(r)−f(r′))2+r′∈/S∑​prr′​f2(r)], D′(S)=2[∑i=1Nfi−∑r∈Sλ0rf(r)].D'(S)=2\Big[\sum_{i=1}^Nf_i-\sum_{r\in S}\lambda_{0r}f(r)\Big].D′(S)=2[i=1∑N​fi​−r∈S∑​λ0r​f(r)].

The formal goal is the product form N′(S)≤D′(S)∑r∈Sf(r)nˉrN'(S)\le D'(S)\sum_{r\in S}f(r)\bar n_rN′(S)≤D′(S)∑r∈S​f(r)nˉr​.

Milestones

  1. The utilization identity Eπ[1{Br}]=λr/μrE_\pi[1\{B_r\}]=\lambda_r/\mu_rEπ​[1{Br​}]=λr​/μr​ (pp. 16 and 19).
  2. Theorem 4.2: the linear equalities (24), (25) between nˉr\bar n_rnˉr​ and Irr′I_{rr'}Irr′​.
  3. Theorem 4.3: ∑r∈CiIrr′+Nir′=nˉr′\sum_{r\in C_i}I_{rr'}+N_{ir'}=\bar n_{r'}∑r∈Ci​​Irr′​+Nir′​=nˉr′​ (28).
  4. Theorem 4.4: any nonnegative (x,I,N)(x,I,N)(x,I,N) satisfying (24), (25), (28), with nˉr=λrxr\bar n_r=\lambda_rx_rnˉr​=λr​xr​ in those equalities, satisfies every inequality of Theorem 4.1. This statement is deterministic.

Significance

Theorem 4.1 gives, for each choice of SSS and fff, a linear inequality on mean response times valid for all admissible policies. Minimizing a linear holding cost ∑rcrxr\sum_r c_rx_r∑r​cr​xr​ subject to these inequalities is a linear program whose value bounds the optimal scheduling cost from below; the paper reports numerical values of such bounds in its Section 9. Theorems 4.2–4.4 show that a polynomial-size polyhedron in the variables (nˉ,I,N)(\bar n,I,N)(nˉ,I,N) implies all of these inequalities at once, so the parametric search over fff is unnecessary.

The results are proved in the paper. As far as is known, none of them has a machine-checked proof. Formalizing them requires a Lean treatment of invariant distributions of controlled countable-state Markov chains with unbounded test functions, which is currently absent from Mathlib, and then the algebra of the drift identities. The definitions here (network data, Markovian sequencing policies, the generator, Assumption A) are the substrate that the paper's later results on routing, closed networks and higher-order bounds would reuse.

Difficulty

Every statement except Theorem 4.4 rests on taking expectations of the generator applied to unbounded functions (nrn_rnr​, nrnr′n_rn_{r'}nr​nr′​) under the invariant distribution. The invariance condition is stated only for indicators of single states; extending ∑nπ(n)(Gg)(n)=0\sum_n\pi(n)(\mathcal Gg)(n)=0∑n​π(n)(Gg)(n)=0 to quadratic ggg needs an interchange of summations justified by the second-moment condition of Assumption A. The utilization identity additionally needs uniqueness of the traffic solution to identify μrEπ[1{Br}]\mu_rE_\pi[1\{B_r\}]μr​Eπ​[1{Br​}] with λr\lambda_rλr​. Theorem 4.1 then needs the sign bookkeeping that turns an identity into an inequality: the terms dropped are nonnegative only because f≥0f\ge0f≥0 on SSS, fi≥0f_i\ge0fi​≥0 and at most one class per station is in service.

Formalization scope

Classes are Fin R, stations Fin N, states Fin R → ℕ, all rates and probabilities real. A policy is a Bool-valued function of the state with the two admissibility constraints; work conservation is not assumed. Invariance is global balance of the generator on the countable state space; expectations are tsums. The uniformized chain and the epochs τk\tau_kτk​ of the paper are not built: the paper notes that its expectations at τk\tau_kτk​ are expectations under the invariant distribution of n⃗(t)\vec n(t)n(t).

Conventions fixed in Lean:

  • λrxr\lambda_rx_rλr​xr​ appears only as the mean number in system nˉr\bar n_rnˉr​ (Little's law, used by the paper on pp. 11 and 20); response times are not formalized.
  • Sums over r′∉Sr'\notin Sr′∈/S include the exit r′=0r'=0r′=0 (p. 15).
  • f-parameters are nonnegative on SSS (p. 9).
  • The network is open: (15) has a unique solution, and λ\lambdaλ is an input constrained by (15), never defined from the policy.
  • (18) is stated multiplied by D′(S)D'(S)D′(S), which avoids Lean's x/0=0x/0=0x/0=0 and is (18) whenever D′(S)>0D'(S)>0D′(S)>0.

A quotient-form statement of (18) would be trivially true when D′(S)=0D'(S)=0D′(S)=0, and defining λr\lambda_rλr​ as μrEπ[1{Br}]\mu_rE_\pi[1\{B_r\}]μr​Eπ​[1{Br​}] would make the utilization identity hold by definition; both are excluded.

Welcome contributions: a general lemma extending global balance to test functions of polynomial growth under moment conditions; proofs of the drift identities; the deterministic Theorem 4.4.

Selected references

  • D. Bertsimas, I. Ch. Paschalidis, J. N. Tsitsiklis, Optimization of Multiclass Queueing Networks: Polyhedral and Nonlinear Characterizations of Achievable Performance, MIT Sloan WP #3509-92-MSA, 1992; Ann. Appl. Probab. 4(1), 1994. https://doi.org/10.1214/aoap/1177005200
  • C. H. Papadimitriou, J. N. Tsitsiklis, The complexity of optimal queuing network control, Math. Oper. Res. 24(2), 1999. https://doi.org/10.1287/moor.24.2.293
8 thms1 active userReviewed
Dynamic ProgrammingOperations Research·Captain: mikedeng1

Discrete Dynamic Programming 2: A Stationary Policy Is Nearly Optimal as the Discount Factor Tends to 1 Exactly When It Maximizes x(g) and, Among Those, y(g)Research Paper

Motivation

A finite Markov decision problem with discounting is solved by Howard's policy improvement routine: start from a stationary policy, switch to actions that do better against its value, repeat. When the discount factor β\betaβ tends to 111 the total discounted income typically diverges, and the natural targets become the long-run average income and, among policies with the best average, the policy that does best in the transient phase. Howard treated this undiscounted case directly (Howard, 1960). David Blackwell's 1962 paper (Blackwell, 1962) treats β=1\beta = 1β=1 as a limit of β<1\beta < 1β<1: it expands the discounted return of a stationary policy in powers of 1−β1-\beta1−β and reads off which policies remain good as β→1\beta \to 1β→1. The two leading coefficients of that expansion, the gain x(f)x(f)x(f) and the bias y(f)y(f)y(f), became the standard objects of average-reward and sensitive-discount optimality (Veinott, 1969; Puterman, 1994, Ch. 8–10).

Timeline. Howard (1960) gives policy iteration for discounted and average-income problems. Blackwell (1962) proves that some stationary policy is optimal for all β\betaβ near 111 (his Theorem 5, the subject of a companion mission) and, in Theorem 4, characterizes the nearly optimal stationary policies through xxx and yyy. Miller and Veinott (1969) and Veinott (1969) extend the expansion to all orders (nnn-discount optimality).

Setting

There are finitely many states s∈Ss \in Ss∈S and a finite nonempty set AAA of actions. Action aaa in state sss pays an income i(s,a)∈Ri(s,a) \in \mathbb Ri(s,a)∈R and moves the system to s′s's′ with probability q(s′∣s,a)q(s' \mid s,a)q(s′∣s,a). FFF is the finite set of decision rules f:S→Af : S \to Af:S→A. A policy is a sequence π={f1,f2,… }\pi = \{f_1, f_2, \dots\}π={f1​,f2​,…} of decision rules; f(∞)f^{(\infty)}f(∞) uses fff every day, and (g,π)(g, \pi)(g,π) uses ggg first and then π\piπ. For f∈Ff \in Ff∈F, r(f)r(f)r(f) is the vector (i(s,f(s)))s(i(s,f(s)))_s(i(s,f(s)))s​ and Q(f)Q(f)Q(f) the Markov matrix (q(s′∣s,f(s)))s,s′(q(s' \mid s,f(s)))_{s,s'}(q(s′∣s,f(s)))s,s′​. The discounted return of π\piπ is the vector

Vβ(π)=∑n=0∞βnQ(f1)⋯Q(fn) r(fn+1),0≤β<1,V_\beta(\pi) = \sum_{n=0}^\infty \beta^n Q(f_1)\cdots Q(f_n)\, r(f_{n+1}), \qquad 0 \le \beta < 1,Vβ​(π)=n=0∑∞​βnQ(f1​)⋯Q(fn​)r(fn+1​),0≤β<1,

and Vβ(f)V_\beta(f)Vβ​(f) abbreviates Vβ(f(∞))V_\beta(f^{(\infty)})Vβ​(f(∞)). Vectors are compared coordinatewise; w1>w2w_1 > w_2w1​>w2​ means w1≥w2w_1 \ge w_2w1​≥w2​ and w1≠w2w_1 \neq w_2w1​=w2​. A policy is β-optimal if its return dominates that of every policy, and U(β)U(\beta)U(β) is the return of a β-optimal policy. It is optimal if it is β-optimal for all β\betaβ sufficiently near 111, and nearly optimal if U(β)−Vβ(π)→0U(\beta) - V_\beta(\pi) \to 0U(β)−Vβ​(π)→0 as β→1\beta \to 1β→1.

For any Markov matrix QQQ, the limit matrix Q∗Q^*Q∗ is the limit of (I+Q+⋯+QN)/(N+1)(I + Q + \cdots + Q^N)/(N+1)(I+Q+⋯+QN)/(N+1), and the deviation matrix is H=(I−Q+Q∗)−1−Q∗H = (I - Q + Q^*)^{-1} - Q^*H=(I−Q+Q∗)−1−Q∗. For a rule fff, Q∗(f)Q^*(f)Q∗(f) and H(f)H(f)H(f) are those of Q(f)Q(f)Q(f), and

x(f)=Q∗(f) r(f),y(f)=H(f) r(f).x(f) = Q^*(f)\, r(f), \qquad y(f) = H(f)\, r(f).x(f)=Q∗(f)r(f),y(f)=H(f)r(f).

With p(s,a)w=∑s′q(s′∣s,a)ws′p(s,a)w = \sum_{s'} q(s' \mid s,a) w_{s'}p(s,a)w=∑s′​q(s′∣s,a)ws′​, the set G(s,f)G(s,f)G(s,f) consists of the actions aaa with p(s,a)x(f)>xs(f)p(s,a)x(f) > x_s(f)p(s,a)x(f)>xs​(f), or with p(s,a)x(f)=xs(f)p(s,a)x(f) = x_s(f)p(s,a)x(f)=xs​(f) and i(s,a)+p(s,a)y(f)>xs(f)+ys(f)i(s,a) + p(s,a)y(f) > x_s(f) + y_s(f)i(s,a)+p(s,a)y(f)>xs​(f)+ys​(f); E(s,f)E(s,f)E(s,f) consists of those with equality in both.

Formalization targets

Goal: Theorem 4(e)

For any f0f_0f0​ with G(s,f0)=∅G(s,f_0) = \varnothingG(s,f0​)=∅ for all sss:

x(f0)≥x(g)  ∀g∈F;∃f∗∈F∗:={g:x(g)=x(f0)} with y(f∗)≥y(g) ∀g∈F∗;x(f_0) \ge x(g)\ \ \forall g \in F;\qquad \exists f^* \in F^* := \{g : x(g) = x(f_0)\}\ \text{with}\ y(f^*) \ge y(g)\ \forall g \in F^*;x(f0​)≥x(g)  ∀g∈F;∃f∗∈F∗:={g:x(g)=x(f0​)} with y(f∗)≥y(g) ∀g∈F∗; g(∞) is nearly optimal  ⟺  x(g)=x(f∗) and y(g)=y(f∗).g^{(\infty)} \text{ is nearly optimal} \iff x(g) = x(f^*) \text{ and } y(g) = y(f^*).g(∞) is nearly optimal⟺x(g)=x(f∗) and y(g)=y(f∗).

Milestones and intermediate results

Milestones: Lemma 1(b) (rank⁡(I−Q)+rank⁡Q∗=S\operatorname{rank}(I-Q) + \operatorname{rank} Q^* = Srank(I−Q)+rankQ∗=S), Theorem 4(b) (improvement for β near 1), 4(c) (a sufficient condition for optimality), Lemma 2, and 4(d) (a sufficient condition for near optimality).

The mission also states, as intermediate results:

  • Lemma 1(a), (c), (d): for every Markov matrix, convergence of the Cesàro means to a Markov Q∗Q^*Q∗ with QQ∗=Q∗Q=Q∗Q∗=Q∗QQ^* = Q^*Q = Q^*Q^* = Q^*QQ∗=Q∗Q=Q∗Q∗=Q∗; unique solvability of Qx=xQx = xQx=x, Q∗x=Q∗cQ^*x = Q^*cQ∗x=Q∗c; nonsingularity of I−Q+Q∗I - Q + Q^*I−Q+Q∗, ∑nβn(Qn−Q∗)→H\sum_n \beta^n (Q^n - Q^*) \to H∑n​βn(Qn−Q∗)→H and the identities for HHH.
  • Theorem 4(a): Vβ(f)=x(f)/(1−β)+y(f)+o(1)V_\beta(f) = x(f)/(1-\beta) + y(f) + o(1)Vβ​(f)=x(f)/(1−β)+y(f)+o(1), with x(f),y(f)x(f), y(f)x(f),y(f) the unique solutions of their linear systems; display (2), the same expansion for (g,f(∞))(g, f^{(\infty)})(g,f(∞)).
  • Theorem 3 and its Corollary for fixed β<1\beta < 1β<1, and the first assertion of 4(e).

Significance

Theorem 4(e) says that near optimality for β near 1 is exactly lexicographic maximization: first of the average income xxx, then of the bias yyy. It justifies the two-level optimality equations used throughout average-reward dynamic programming and shows that, once the β = 1 improvement routine stops, the remaining problem is a bias maximization over the gain-optimal rules. Theorem 4(a) is the first two terms of the Laurent expansion of discounted values, the starting point of sensitive-discount optimality.

The results are classical and proved in the paper (Lemma 1 with a reference to Kemeny and Snell); no machine-checked proof of them is known on the platform. A complete development produces a multichain theory of Cesàro limit and deviation matrices of arbitrary finite Markov matrices, which Mathlib does not have, and the expansion of discounted returns near β = 1.

Difficulty

Lemma 1 must be proved for every Markov matrix, including reducible and periodic ones, where QnQ^nQn does not converge and the stationary distribution is not unique; arguments through the Perron–Frobenius eigenvector of an irreducible chain do not apply. In Theorem 4(e) the hard part is the existence of a single f∗f^*f∗ whose bias dominates every gain-optimal rule in every coordinate at once; a rule maximizing each coordinate separately is not enough. The final characterization compares a stationary policy with all policies, including time-dependent ones, through U(β)U(\beta)U(β).

Formalization scope

States and actions are finite nonempty types; incomes are real of any sign; a policy is a sequence ℕ → (St → Act) with π 0 the paper's f1f_1f1​. VβV_\betaVβ​ is a real tsum. Q∗Q^*Q∗ is limUnder of the Cesàro means, and its existence is Lemma 1(a), not an assumption; H(β)H(\beta)H(β) is a matrix tsum, whose summability for 0≤β<10 \le \beta < 10≤β<1 is part of Lemma 1(d); HHH uses Mathlib's total inverse, whose nonsingularity is also part of Lemma 1(d). x(f)x(f)x(f) and y(f)y(f)y(f) are defined by the closed forms Q∗(f)r(f)Q^*(f)r(f)Q∗(f)r(f) and H(f)r(f)H(f)r(f)H(f)r(f) from the paper's proof, and Theorem 4(a) asserts that they are the unique solutions of the paper's defining systems. Limits "as β → 1" are along β→1−\beta \to 1^-β→1−. "Nearly optimal" is encoded without UUU: for every ε>0\varepsilon > 0ε>0, for all β in some interval (β0,1)(\beta_0, 1)(β0​,1), every policy's return is at most Vβ(π)+εV_\beta(\pi) + \varepsilonVβ​(π)+ε in every coordinate; this is equivalent to U(β)−Vβ(π)→0U(\beta) - V_\beta(\pi) \to 0U(β)−Vβ​(π)→0 because a β-optimal policy exists. "Optimal" (§4) and "β-optimal" (§3) are distinct definitions, and Theorem 3's β-dependent improvement set is distinct from the §4 set G(s,f)G(s,f)G(s,f).

A formalization in which optimality or near optimality is tested only against stationary policies, or in which Q∗Q^*Q∗ is assumed to exist or the chain to be irreducible, proves a different and easier theorem and does not meet the targets.

Contributions are welcome at every level: the Cesàro and Abel limit theory of finite Markov matrices (reusable well beyond this paper), the policy improvement theorem for fixed β, and the comparison arguments of Theorem 4. Theorem 3 and the Corollary are also drafted in the companion mission on Theorem 5 in another namespace.

Selected references

  • D. Blackwell, Discrete Dynamic Programming, Ann. Math. Statist. 33(2):719–726, 1962. https://doi.org/10.1214/aoms/1177704593
  • R. A. Howard, Dynamic Programming and Markov Processes, MIT Press, 1960.
  • J. G. Kemeny and J. L. Snell, Finite Markov Chains, Van Nostrand, 1960.
  • B. L. Miller and A. F. Veinott, Discrete Dynamic Programming with a Small Interest Rate, Ann. Math. Statist. 40(2):366–370, 1969.
  • A. F. Veinott, Discrete Dynamic Programming with Sensitive Discount Optimality Criteria, Ann. Math. Statist. 40(5):1635–1660, 1969. https://doi.org/10.1214/aoms/1177697379
  • M. L. Puterman, Markov Decision Processes: Discrete Stochastic Dynamic Programming, Wiley, 1994. https://doi.org/10.1002/9780470316887
9 thms1 active userReviewed
Dynamic ProgrammingMachine LearningReinforcement Learning·Captain: mikedeng1

Reinforcement Learning: An Introduction IX: Average Reward and the Futility of Discounting in Continuing ProblemsTextbook

Motivation

Reinforcement learning formulates control as maximizing reward accumulated over time. For continuing tasks, where interaction never terminates, the standard textbook objective is the discounted return with a discount rate γ<1\gamma < 1γ<1. Chapter 10 of Sutton and Barto, Reinforcement Learning: An Introduction (2nd ed., MIT Press, 2018), argues that once values are approximated by a function of features rather than stored per state, this objective is the wrong one, and it proposes the average-reward setting, long standard in dynamic programming (Puterman, Markov Decision Processes, 1994), as its replacement.

The central piece of evidence is a short calculation printed in the box The Futility of Discounting in Continuing Problems (p. 254). If one tries to rescue discounting by averaging discounted values over the states the policy actually visits, the resulting objective is a constant multiple of the average reward, so the discount rate has no effect on which policy is preferred. This mission formalizes that calculation together with the definitions of §10.3 that it rests on, and the two exercises of §10.3 that illustrate the differential value (10.13).

Setting

A finite MDP has finite state set S\mathcal SS, finite action set A\mathcal AA, finite reward set R⊂R\mathcal R \subset \mathbb RR⊂R and dynamics p(s′,r∣s,a)p(s', r \mid s, a)p(s′,r∣s,a), a probability distribution over next state and reward for each state–action pair (3.2)–(3.3). Write p(s′∣s,a)=∑rp(s′,r∣s,a)p(s' \mid s, a) = \sum_r p(s', r\mid s,a)p(s′∣s,a)=∑r​p(s′,r∣s,a). A policy π(a∣s)\pi(a \mid s)π(a∣s) is a distribution over actions for each state. It turns the MDP into a Markov chain with transition matrix Pπ(s,s′)=∑aπ(a∣s) p(s′∣s,a)P_\pi(s, s') = \sum_a \pi(a\mid s)\, p(s'\mid s,a)Pπ​(s,s′)=∑a​π(a∣s)p(s′∣s,a) and expected one-step reward rπ(s)=∑aπ(a∣s)∑s′,rp(s′,r∣s,a) rr_\pi(s) = \sum_a \pi(a\mid s)\sum_{s',r} p(s',r\mid s,a)\, rrπ​(s)=∑a​π(a∣s)∑s′,r​p(s′,r∣s,a)r, so that Pr⁡{St=s∣S0=s0}=Pπt(s0,s)\Pr\{S_t = s\mid S_0=s_0\} = P_\pi^t(s_0,s)Pr{St​=s∣S0​=s0​}=Pπt​(s0​,s) and E[Rt+1∣S0=s0]=(Pπtrπ)(s0)\mathbb E[R_{t+1}\mid S_0 = s_0] = (P_\pi^t r_\pi)(s_0)E[Rt+1​∣S0​=s0​]=(Pπt​rπ​)(s0​).

A stationary distribution of π\piπ is a probability vector μ\muμ on S\mathcal SS with

∑sμ(s)∑aπ(a∣s) p(s′∣s,a)=μ(s′)for all s′(10.8).\sum_s \mu(s)\sum_a \pi(a\mid s)\,p(s'\mid s,a) = \mu(s') \quad\text{for all } s' \qquad (10.8).s∑​μ(s)a∑​π(a∣s)p(s′∣s,a)=μ(s′)for all s′(10.8).

The average reward of π\piπ is

r(π)=lim⁡h→∞1h∑t=1hE[Rt∣S0,A0:t−1∼π](10.6),r(π)=∑sμπ(s)∑aπ(a∣s)∑s′,rp(s′,r∣s,a) r(10.7),r(\pi) = \lim_{h\to\infty}\frac1h\sum_{t=1}^h \mathbb E[R_t\mid S_0, A_{0:t-1}\sim\pi] \quad (10.6), \qquad r(\pi) = \sum_s \mu_\pi(s)\sum_a\pi(a\mid s)\sum_{s',r}p(s',r\mid s,a)\,r \quad (10.7),r(π)=h→∞lim​h1​t=1∑h​E[Rt​∣S0​,A0:t−1​∼π](10.6),r(π)=s∑​μπ​(s)a∑​π(a∣s)s′,r∑​p(s′,r∣s,a)r(10.7),

the second form holding when the steady-state distribution μπ(s)=lim⁡t→∞Pr⁡{St=s}\mu_\pi(s) = \lim_{t\to\infty}\Pr\{S_t = s\}μπ​(s)=limt→∞​Pr{St​=s} exists and does not depend on S0S_0S0​ (an ergodic MDP).

The discounted value function is vπγ(s)=Eπ[∑k≥0γkRt+k+1∣St=s]v^\gamma_\pi(s) = \mathbb E_\pi\big[\sum_{k\ge0}\gamma^k R_{t+k+1}\mid S_t = s\big]vπγ​(s)=Eπ​[∑k≥0​γkRt+k+1​∣St​=s], and the objective of the box is

J(π)=∑sμπ(s) vπγ(s).J(\pi) = \sum_s \mu_\pi(s)\, v^\gamma_\pi(s).J(π)=s∑​μπ​(s)vπγ​(s).

Finally, the differential value of a state (10.13) is vπ(s)=lim⁡γ→1lim⁡h→∞∑t=0hγt(Eπ[Rt+1∣S0=s]−r(π))v_\pi(s) = \lim_{\gamma\to1}\lim_{h\to\infty}\sum_{t=0}^h\gamma^t\big(\mathbb E_\pi[R_{t+1}\mid S_0 = s] - r(\pi)\big)vπ​(s)=limγ→1​limh→∞​∑t=0h​γt(Eπ​[Rt+1​∣S0​=s]−r(π)).

Formalization targets

Goal: the futility of discounting

For 0≤γ<10 \le \gamma < 10≤γ<1, every policy π\piπ and every stationary distribution μπ\mu_\piμπ​ of π\piπ,

J(π)=∑sμπ(s) vπγ(s)=11−γ r(π),J(\pi) = \sum_s \mu_\pi(s)\, v^\gamma_\pi(s) = \frac{1}{1-\gamma}\, r(\pi),J(π)=s∑​μπ​(s)vπγ​(s)=1−γ1​r(π),

and therefore, for a fixed γ\gammaγ and policies π,π′\pi, \pi'π,π′ with their own stationary distributions,

J(π)≤J(π′)  ⟺  r(π)≤r(π′).J(\pi) \le J(\pi') \iff r(\pi) \le r(\pi').J(π)≤J(π′)⟺r(π)≤r(π′).

Milestones

  1. (10.6)–(10.7): if Pr⁡{St=s∣S0=s0}→μ(s)\Pr\{S_t = s\mid S_0 = s_0\}\to\mu(s)Pr{St​=s∣S0​=s0​}→μ(s) for all s0,ss_0, ss0​,s, then from every start both lim⁡tE[Rt]\lim_t \mathbb E[R_t]limt​E[Rt​] and the Cesàro limit (10.6) exist and equal the μ\muμ-sum.
  2. (10.8): such a limit μ\muμ is a stationary distribution.
  3. (3.14): vπγv^\gamma_\pivπγ​ satisfies the Bellman equation, the step "(Bellman Eq.)" of the box.
  4. Offset invariance (p. 250): the differential Bellman equations for vπ,qπ,v∗,q∗v_\pi, q_\pi, v_*, q_*vπ​,qπ​,v∗​,q∗​ and the differential TD errors (10.10)–(10.11) are unchanged when all values are shifted by a constant.
  5. Exercise 10.6: for expected rewards 1,0,1,0,…1,0,1,0,\dots1,0,1,0,… from A\mathsf AA and 0,1,0,1,…0,1,0,1,\dots0,1,0,1,… from B\mathsf BB, the average reward is 12\tfrac1221​, the limit (10.7) and a steady-state distribution do not exist, and vπ(A)=14v_\pi(\mathsf A) = \tfrac14vπ​(A)=41​, vπ(B)=−14v_\pi(\mathsf B) = -\tfrac14vπ​(B)=−41​.
  6. Exercise 10.7: in the three-state ring with reward +1+1+1 on arrival in A\mathsf AA, the average reward is 13\tfrac1331​ and the differential values are v(A)=−13v(\mathsf A) = -\tfrac13v(A)=−31​, v(B)=0v(\mathsf B) = 0v(B)=0, v(C)=13v(\mathsf C) = \tfrac13v(C)=31​.

The book prints no answers to Exercises 10.6 and 10.7; the values above were computed for this mission.

Significance

The result itself. The identity shows that the discount rate cannot enter the ranking of policies once performance is measured over the on-policy state distribution: any objective of the form "discounted value averaged over where the policy goes" is the average reward up to a positive factor. The book draws from it the conclusion (p. 253) that γ\gammaγ changes from a problem parameter to a solution-method parameter, and that discounting algorithms with function approximation, which do not optimize this averaged objective, are not guaranteed to optimize average reward either. Milestones 1–2 connect the closed form of r(π)r(\pi)r(π) to its definition as a long-run rate; milestone 4 explains why differential methods determine values only up to an offset; the exercises show that (10.13) gives finite differential values in periodic chains where the differential return (10.9) has no limit.

Formalizing it. The results are classical and elementary, but the book's derivation is informal: it takes the Bellman equation for granted, sums a geometric series of equalities, and leaves implicit which distribution μπ\mu_\piμπ​ is meant and when (10.6) and (10.7) agree. This mission fixes each of those points: vπγv^\gamma_\pivπγ​ is defined from returns, the identity is proved for every stationary distribution, and the ergodicity hypothesis is made the exact limit condition the text states. No machine-checked version of these statements is known to exist; the platform's Markov-chain and average-reward libraries (Puterman's unichain optimality equation, Doeblin convergence) state different results.

Difficulty

The box reads as a chain of equalities, and each individual step is short. The one step that is not algebra is "(Bellman Eq.)": with vπγv^\gamma_\pivπγ​ defined as an expected discounted return, the Bellman equation requires exchanging an infinite sum with a finite expectation and shifting the index of a convergent series, which needs absolute convergence for γ<1\gamma < 1γ<1 and bounded rewards. The final line, which unrolls J=r+γJJ = r + \gamma JJ=r+γJ into a geometric series, is only valid because JJJ is finite. For milestone 1, the book's hypothesis is the existence of an S0S_0S0​-independent limiting distribution, not irreducibility or aperiodicity; replacing it by either would change the statement. The exercises require the iterated limit in (10.13) to be computed explicitly: the inner limit is a periodic series summed in closed form, and the outer limit is a removable singularity at γ=1\gamma = 1γ=1.

Formalization scope

The Lean development lives in the namespace SuttonBartoRL.AverageReward. States, actions and rewards are finite; one action set serves every state; the dynamics are the four-argument p(s′,r∣s,a)p(s', r\mid s,a)p(s′,r∣s,a). Probabilities Pr⁡{St=s∣S0=s0}\Pr\{S_t = s\mid S_0 = s_0\}Pr{St​=s∣S0​=s0​} and E[Rt+1∣S0=s0]\mathbb E[R_{t+1}\mid S_0=s_0]E[Rt+1​∣S0​=s0​] are computed from the matrix powers PπtP_\pi^tPπt​. The discounted value vπγv^\gamma_\pivπγ​ is a real series ∑kγk(Pπkrπ)(s)\sum_k \gamma^k (P_\pi^k r_\pi)(s)∑k​γk(Pπk​rπ​)(s); it is not defined as the solution of the Bellman equation, since that definition would reduce the goal to three lines of algebra. The average reward and JJJ take the state distribution μ\muμ as an explicit argument; the goal holds for every stationary distribution of π\piπ and assumes no ergodicity, which is all the box uses (under ergodicity μπ\mu_\piμπ​ is the unique one). γ=0\gamma = 0γ=0 is allowed. The differential value (10.13) is stated as the existence of both limits with the given value, with γ→1\gamma \to 1γ→1 from below, so no default value of a nonexistent limit can make a statement true. The optimality equations use a nonempty action set. The sum ∑t=1hE[Rt]\sum_{t=1}^h \mathbb E[R_t]∑t=1h​E[Rt​] of (10.6) is written with the index shifted to t=0,…,h−1t = 0,\dots,h-1t=0,…,h−1.

The Lean does not formalize a trajectory probability space: expectations and probabilities are the matrix expressions above, which is what they equal for a Markov chain. The Exercise 10.6 hypothesis constrains expected rewards of one policy, which is all the exercise uses.

Reusable parts: the finite-MDP layer duplicates the one drafted for the other missions of this series and is expected to be merged with it; the Cesàro and stationary-distribution lemmas of milestones 1–2 are general facts about finite Markov chains. Contributions welcome: proofs of any milestone, and a proof of the goal from milestone 3.

Selected references

  • R. S. Sutton, A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018, ISBN 9780262039246, §§10.3–10.4, pp. 249–254. http://incompleteideas.net/book/the-book-2nd.html
  • M. L. Puterman, Markov Decision Processes: Discrete Stochastic Dynamic Programming, Wiley, 1994. https://doi.org/10.1002/9780470316887
  • S. Mahadevan, "Average reward reinforcement learning: foundations, algorithms, and empirical results", Machine Learning 22, 159–195, 1996. https://doi.org/10.1007/BF00114727
11 thms1 active userReviewed
Machine LearningOptimizationReinforcement Learning·Captain: mikedeng1

Reinforcement Learning: An Introduction XII: The Policy Gradient TheoremTextbook

Motivation

Policy gradient methods learn a parameterized policy π(a∣s,θ)\pi(a \mid s, \theta)π(a∣s,θ) directly, by stochastic gradient ascent on a scalar performance measure J(θ)J(\theta)J(θ), instead of deriving the policy from learned action values. They are how reinforcement learning handles continuous action spaces, stochastic optimal policies and prior knowledge built into the policy's form, and they underlie REINFORCE (Williams, 1992) and the actor–critic family. Every such method needs an estimate of ∇J(θ)\nabla J(\theta)∇J(θ). The difficulty is that JJJ depends on θ\thetaθ in two ways: through the action choices in each state, and through the distribution of states those choices produce. The second effect depends on the unknown environment dynamics.

The policy gradient theorem (Sutton, McAllester, Singh and Mansour, 2000; Marbach and Tsitsiklis, 2001) gives ∇J(θ)\nabla J(\theta)∇J(θ) as an expectation over the on-policy state distribution that involves no derivative of that distribution. Chapter 13 of Sutton and Barto's Reinforcement Learning: An Introduction (2nd ed., 2018) states it as Eq. (13.5), proves it in a box for the episodic case (p. 325) and in a second box for the continuing case (pp. 334–335), and builds REINFORCE, REINFORCE with baseline and actor–critic methods on it. This mission formalizes that chapter's theorem and the identities around it, in the book's own model.

Setting

A finite episodic MDP has a finite set S\mathcal SS of nonterminal states, a terminal state, a finite action set A\mathcal AA, a finite reward set R⊂R\mathcal R \subset \mathbb RR⊂R and dynamics p(s′,r∣s,a)p(s', r \mid s, a)p(s′,r∣s,a): for each nonterminal sss and action aaa, a probability distribution over next state s′∈S+=S∪{terminal}s' \in \mathcal S^+ = \mathcal S \cup \{\text{terminal}\}s′∈S+=S∪{terminal} and reward rrr. The terminal state is absorbing and pays nothing. Write p(s′∣s,a)=∑rp(s′,r∣s,a)p(s' \mid s, a) = \sum_r p(s', r \mid s, a)p(s′∣s,a)=∑r​p(s′,r∣s,a) and r(s,a)r(s, a)r(s,a) for the expected reward.

A differentiable policy parameterization assigns to every θ∈Rd′\theta \in \mathbb R^{d'}θ∈Rd′ and state sss a distribution π(⋅∣s,θ)\pi(\cdot \mid s, \theta)π(⋅∣s,θ) over actions, with θ↦π(a∣s,θ)\theta \mapsto \pi(a \mid s, \theta)θ↦π(a∣s,θ) differentiable. Under πθ\pi_\thetaπθ​ the nonterminal states form a substochastic chain with matrix Pθ(s,s′)=∑aπ(a∣s,θ)p(s′∣s,a)P_\theta(s, s') = \sum_a \pi(a \mid s, \theta) p(s' \mid s, a)Pθ​(s,s′)=∑a​π(a∣s,θ)p(s′∣s,a); Pr⁡(s→x,k,π)=Pθk(s,x)\Pr(s \to x, k, \pi) = P_\theta^k(s, x)Pr(s→x,k,π)=Pθk​(s,x). Episodes terminate when ∑kPθk(s,s′)<∞\sum_{k} P_\theta^k(s, s') < \infty∑k​Pθk​(s,s′)<∞ for all s,s′s, s's,s′.

There is no discounting (γ=1\gamma = 1γ=1, p. 324). The state value vπ(s)=∑k≥0(Pθkrθ)(s)v_{\pi}(s) = \sum_{k \ge 0} (P_\theta^k r_\theta)(s)vπ​(s)=∑k≥0​(Pθk​rθ​)(s) is the expected total reward from sss, with rθ(s)=∑aπ(a∣s,θ)r(s,a)r_\theta(s) = \sum_a \pi(a\mid s,\theta) r(s,a)rθ​(s)=∑a​π(a∣s,θ)r(s,a); the action value qπ(s,a)q_\pi(s,a)qπ​(s,a) is the expected total reward after taking aaa in sss. The episode starts in a fixed state s0s_0s0​, and the performance is J(θ)=vπθ(s0)J(\theta) = v_{\pi_\theta}(s_0)J(θ)=vπθ​​(s0​) (13.4). The expected number of visits to sss in an episode is η(s)=∑k≥0Pr⁡(s0→s,k,π)\eta(s) = \sum_{k \ge 0} \Pr(s_0 \to s, k, \pi)η(s)=∑k≥0​Pr(s0​→s,k,π), and the on-policy distribution is μ(s)=η(s)/∑s′η(s′)\mu(s) = \eta(s) / \sum_{s'} \eta(s')μ(s)=η(s)/∑s′​η(s′) (9.3).

In the continuing case there is no terminal state, J(θ)=r(π)J(\theta) = r(\pi)J(θ)=r(π) is the average reward per step (13.15), μ\muμ is the steady-state distribution, and vπv_\pivπ​, qπq_\piqπ​ are differential values, defined from the return ∑k(Rt+k+1−r(π))\sum_k (R_{t+k+1} - r(\pi))∑k​(Rt+k+1​−r(π)) (13.17).

Formalization targets

Goal: the policy gradient theorem, episodic case (13.5)

If episodes terminate under πθ0\pi_{\theta_0}πθ0​​, then JJJ is differentiable at θ0\theta_0θ0​ and

∇J(θ0)=∑sη(s)∑aqπ(s,a) ∇π(a∣s,θ0)=(∑s′η(s′))∑sμ(s)∑aqπ(s,a) ∇π(a∣s,θ0),\nabla J(\theta_0) = \sum_s \eta(s) \sum_a q_\pi(s,a)\, \nabla \pi(a \mid s, \theta_0) = \Big(\sum_{s'} \eta(s')\Big) \sum_s \mu(s) \sum_a q_\pi(s,a)\, \nabla \pi(a \mid s, \theta_0),∇J(θ0​)=s∑​η(s)a∑​qπ​(s,a)∇π(a∣s,θ0​)=(s′∑​η(s′))s∑​μ(s)a∑​qπ​(s,a)∇π(a∣s,θ0​),

with ∑s′η(s′)≥1\sum_{s'} \eta(s') \ge 1∑s′​η(s′)≥1. The book writes ∇J(θ)∝∑sμ(s)∑aqπ(s,a)∇π(a∣s,θ)\nabla J(\theta) \propto \sum_s \mu(s) \sum_a q_\pi(s,a) \nabla \pi(a \mid s,\theta)∇J(θ)∝∑s​μ(s)∑a​qπ​(s,a)∇π(a∣s,θ) and names the constant, the average length of an episode, in words (p. 326). The goal states it.

Milestones

  1. Exercises 3.18–3.19 with γ=1\gamma = 1γ=1: vπ(s)=∑aπ(a∣s)qπ(s,a)v_\pi(s) = \sum_a \pi(a\mid s) q_\pi(s,a)vπ​(s)=∑a​π(a∣s)qπ​(s,a) and qπ(s,a)=∑s′,rp(s′,r∣s,a)(r+vπ(s′))q_\pi(s,a) = \sum_{s',r} p(s',r\mid s,a)(r + v_\pi(s'))qπ​(s,a)=∑s′,r​p(s′,r∣s,a)(r+vπ​(s′)).
  2. The recursion ∇vπ(s)=∑a[∇π(a∣s)qπ(s,a)+π(a∣s)∑s′p(s′∣s,a)∇vπ(s′)]\nabla v_\pi(s) = \sum_a [\nabla\pi(a\mid s) q_\pi(s,a) + \pi(a\mid s) \sum_{s'} p(s'\mid s,a) \nabla v_\pi(s')]∇vπ​(s)=∑a​[∇π(a∣s)qπ​(s,a)+π(a∣s)∑s′​p(s′∣s,a)∇vπ​(s′)], including the differentiability of vπv_\pivπ​.
  3. The unrolled gradient ∇vπ(s)=∑x∑k=0∞Pr⁡(s→x,k,π)∑a∇π(a∣x)qπ(x,a)\nabla v_\pi(s) = \sum_{x} \sum_{k=0}^\infty \Pr(s \to x, k, \pi) \sum_a \nabla\pi(a\mid x) q_\pi(x,a)∇vπ​(s)=∑x​∑k=0∞​Pr(s→x,k,π)∑a​∇π(a∣x)qπ​(x,a) for every sss.
  4. The theorem with a baseline (13.10): ∑ab(s)∇π(a∣s,θ)=0\sum_a b(s) \nabla \pi(a\mid s,\theta) = 0∑a​b(s)∇π(a∣s,θ)=0, hence qπq_\piqπ​ may be replaced by qπ−bq_\pi - bqπ​−b.
  5. The log form behind REINFORCE: where π(⋅∣s,θ)>0\pi(\cdot\mid s,\theta) > 0π(⋅∣s,θ)>0, ∑aqπ(s,a)∇π(a∣s,θ)=∑aπ(a∣s,θ)qπ(s,a)∇ln⁡π(a∣s,θ)\sum_a q_\pi(s,a) \nabla\pi(a\mid s,\theta) = \sum_a \pi(a\mid s,\theta) q_\pi(s,a) \nabla \ln \pi(a\mid s,\theta)∑a​qπ​(s,a)∇π(a∣s,θ)=∑a​π(a∣s,θ)qπ​(s,a)∇lnπ(a∣s,θ), and hence ∇J(θ)=(∑s′η(s′))∑sμ(s)∑aπ(a∣s,θ)qπ(s,a)∇ln⁡π(a∣s,θ)\nabla J(\theta) = (\sum_{s'}\eta(s')) \sum_s \mu(s) \sum_a \pi(a\mid s,\theta) q_\pi(s,a) \nabla \ln \pi(a\mid s,\theta)∇J(θ)=(∑s′​η(s′))∑s​μ(s)∑a​π(a∣s,θ)qπ​(s,a)∇lnπ(a∣s,θ), the exact form of ∇J∝Eπ[qπ(St,At)∇π(At∣St,θ)/π(At∣St,θ)]\nabla J \propto \mathbb E_\pi[q_\pi(S_t,A_t) \nabla\pi(A_t\mid S_t,\theta)/\pi(A_t\mid S_t,\theta)]∇J∝Eπ​[qπ​(St​,At​)∇π(At​∣St​,θ)/π(At​∣St​,θ)].
  6. Exercise 13.3, (13.9): for the linear soft-max, ∇ln⁡π(a∣s,θ)=x(s,a)−∑bπ(b∣s,θ)x(s,b)\nabla \ln \pi(a\mid s,\theta) = x(s,a) - \sum_b \pi(b\mid s,\theta) x(s,b)∇lnπ(a∣s,θ)=x(s,a)−∑b​π(b∣s,θ)x(s,b).
  7. Exercise 13.4: the eligibility vectors of the Gaussian policy (13.19)–(13.20).
  8. The continuing case: under ergodicity, ∇r(πθ)=∑sμ(s)∑a∇π(a∣s,θ)qπ(s,a)\nabla r(\pi_\theta) = \sum_s \mu(s) \sum_a \nabla\pi(a\mid s,\theta) q_\pi(s,a)∇r(πθ​)=∑s​μ(s)∑a​∇π(a∣s,θ)qπ​(s,a) with differential qπq_\piqπ​.

Significance

The theorem turns ∇J\nabla J∇J into a quantity that can be sampled by following the policy: weighting states by μ\muμ is what visiting them under π\piπ does, and the log form makes the action sum an expectation over At∼πA_t \sim \piAt​∼π. REINFORCE (13.8), REINFORCE with baseline (13.11) and one-step and eligibility-trace actor–critic methods all rest on it, and so does their claim that the expected update is in the direction of the performance gradient (p. 329). The baseline identity is why a learned state value can reduce variance without introducing bias.

The results are proved, in the book and in the literature. What this mission adds is a machine-checked version in the book's model: random episode lengths with γ=1\gamma = 1γ=1, vector parameters θ∈Rd′\theta \in \mathbb R^{d'}θ∈Rd′, four-argument dynamics, and values defined from expected returns. The platform already has a proved finite-horizon policy gradient theorem (policy_gradient_finite_horizon, with a baseline and log-form companion) for a fixed horizon TTT, a scalar parameter θ∈R\theta \in \mathbb Rθ∈R and an expected-reward kernel; it does not cover the book's statement. The mission also makes explicit two points the text leaves informal: that episodes terminate, and what exact constant hides behind "∝\propto∝".

Difficulty

The book's proof is a formal manipulation: differentiate the Bellman equation, substitute it into itself, and "unroll" infinitely often. Two steps are not justified on the page. First, it presupposes that ∇vπ(s)\nabla v_\pi(s)∇vπ​(s) exists; with γ=1\gamma = 1γ=1 the value is an infinite series whose convergence depends on θ\thetaθ through termination, so differentiability of vπv_\pivπ​ at θ0\theta_0θ0​ has to be established, and termination is assumed only at θ0\theta_0θ0​. Second, "repeated unrolling" is a limit: after nnn unrollings a remainder ∑xPθn(s,x)∇vπ(x)\sum_x P_\theta^{n}(s,x) \nabla v_\pi(x)∑x​Pθn​(s,x)∇vπ​(x) is left over, and it vanishes only because Pθn→0P_\theta^n \to 0Pθn​→0. Differentiating the series for vπv_\pivπ​ term by term is not an alternative shortcut without a uniform bound on the derivatives of PθkP_\theta^kPθk​.

In the continuing case the corresponding obstacle is the differentiability of the steady-state distribution and of the differential values, which the book's proof uses without comment; here they are part of what is to be proved, from ergodicity at θ0\theta_0θ0​ alone.

Formalization scope

  • Model. S+\mathcal S^+S+ is Option S, with none the single terminal state (several terminal states can be merged, all having value 0). One action type for all states. θ\thetaθ lives in EuclideanSpace ℝ (Fin d), and ∇\nabla∇ is Mathlib's gradient; conclusions are HasGradientAt, so differentiability is asserted, not assumed.
  • Values from returns. vπv_\pivπ​, qπq_\piqπ​, η\etaη are series in powers of PθP_\thetaPθ​; Bellman equations are theorems (milestone 1). The continuing-case average reward and steady-state distribution are the limits of (13.15), and the differential values are the series of (13.17).
  • Implicit hypotheses made explicit. Termination under πθ0\pi_{\theta_0}πθ0​​ is a hypothesis of every episodic result that involves values; the continuing case assumes the book's ergodicity (the limit of Pr⁡{St=s′}\Pr\{S_t = s'\}Pr{St​=s′} exists and does not depend on S0S_0S0​) at θ0\theta_0θ0​. The positivity of π(a∣s,θ0)\pi(a \mid s,\theta_0)π(a∣s,θ0​) is assumed where a logarithm is differentiated.
  • "∝". The episodic goal states the exact equality with the constant ∑s′η(s′)\sum_{s'} \eta(s')∑s′​η(s′) and proves it is at least 1. A formalization of the form "∃c, ∇J=c⋅…\exists c,\ \nabla J = c \cdot \ldots∃c, ∇J=c⋅…" is ruled out: it holds with c=0c = 0c=0 and loses the book's constant.
  • Fixed start state. s0s_0s0​ is a fixed state, as in the book (p. 324); no start distribution.
  • Not included. Convergence of REINFORCE or actor–critic under stochastic-approximation conditions (p. 329) rests on unstated conditions and is not an item. The baseline is a deterministic function of the state, not the random variable the book also allows.

Reusable infrastructure: the episodic value layer (substochastic chains, expected visits, termination) is needed by any undiscounted episodic RL result; the soft-max and Gaussian eligibility computations are needed by every policy-gradient algorithm. Proofs of any milestone, and general lemmas on the differentiability of values and stationary distributions of parameterized finite Markov chains, are welcome.

Selected references

  • R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018, ISBN 9780262039246, Chapter 13. http://incompleteideas.net/book/the-book-2nd.html
  • R. S. Sutton, D. McAllester, S. Singh and Y. Mansour, Policy Gradient Methods for Reinforcement Learning with Function Approximation, NeurIPS 12, 2000. https://proceedings.neurips.cc/paper/1999/hash/464d828b85b0bed98e80ade0a5c43b0f-Abstract.html
  • P. Marbach and J. N. Tsitsiklis, Simulation-Based Optimization of Markov Reward Processes, IEEE Transactions on Automatic Control 46(2), 2001. https://doi.org/10.1109/9.905687
  • R. J. Williams, Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning, Machine Learning 8, 1992. https://doi.org/10.1007/BF00992696
14 thms1 active userReviewed
Machine LearningReinforcement LearningStatistics·Captain: mikedeng1

Reinforcement Learning: An Introduction VI: Batch TD(0) Converges to the Certainty-Equivalence EstimateTextbook

Why batch TD(0) and batch Monte Carlo disagree

Temporal-difference (TD) learning estimates the value of each state of a Markov reward process from observed experience, updating an estimate toward a target built from the next reward and the current estimate of the next state. Monte Carlo (MC) methods instead update toward the full observed return. Both are standard prediction methods in reinforcement learning, and their relationship is a recurring question of the field (Sutton 1988).

When only a finite amount of experience is available, a common practice is to present the same data repeatedly until the estimates stop changing. Chapter 6 of Sutton and Barto's Reinforcement Learning: An Introduction (2nd ed., MIT Press, 2018) uses this setting to explain why TD(0) is often faster: under such batch updating, both methods converge deterministically, but to different answers. Batch MC finds the least-squares fit to the observed returns; batch TD(0) finds the value function of the maximum-likelihood Markov model of the data, the certainty-equivalence estimate. The comparison appears in §6.3, Optimality of TD(0) (pp. 126–128), and is illustrated by Example 6.4, You are the Predictor. The book states these conclusions without proof. This mission formalizes them.

Setting

Let S\mathcal SS be a finite set of nonterminal states and S+=S∪{terminal}\mathcal S^+ = \mathcal S \cup \{\text{terminal}\}S+=S∪{terminal}. An episode is a finite sequence S0,R1,S1,…,ST−1,RT,STS_0, R_1, S_1, \dots, S_{T-1}, R_T, S_TS0​,R1​,S1​,…,ST−1​,RT​,ST​ with S0,…,ST−1∈SS_0, \dots, S_{T-1} \in \mathcal SS0​,…,ST−1​∈S, real rewards R1,…,RTR_1, \dots, R_TR1​,…,RT​, and STS_TST​ terminal. A batch is a finite list of episodes. A visit of sss is an (episode, time t<Tt < Tt<T) pair with St=sS_t = sSt​=s, and n(s)n(s)n(s) counts all visits (every-visit counting).

A value array V:S→RV : \mathcal S \to \mathbb RV:S→R is extended by V(terminal)=0V(\text{terminal}) = 0V(terminal)=0. For a discount rate γ∈[0,1]\gamma \in [0,1]γ∈[0,1], the return is Gt=∑k=t+1Tγk−t−1RkG_t = \sum_{k=t+1}^{T}\gamma^{k-t-1}R_kGt​=∑k=t+1T​γk−t−1Rk​ and the TD error is δt=Rt+1+γV(St+1)−V(St)\delta_t = R_{t+1} + \gamma V(S_{t+1}) - V(S_t)δt​=Rt+1​+γV(St+1​)−V(St​).

Batch TD(0) with step size α\alphaα computes the TD(0) increment for every visit in the batch and changes VVV once, by their sum:

Vm+1(s)=Vm(s)+α∑visits t of s[Rt+1+γVm(St+1)−Vm(St)].V_{m+1}(s) = V_m(s) + \alpha\sum_{\text{visits } t \text{ of } s}\big[R_{t+1} + \gamma V_m(S_{t+1}) - V_m(S_t)\big].Vm+1​(s)=Vm​(s)+αvisits t of s∑​[Rt+1​+γVm​(St+1​)−Vm​(St​)].

Batch constant-α\alphaα MC is the same iteration with the increment Gt−Vm(St)G_t - V_m(S_t)Gt​−Vm​(St​).

The maximum-likelihood model of the batch has transition probabilities p^(j∣i)=N(i,j)/n(i)\hat p(j \mid i) = N(i,j)/n(i)p^​(j∣i)=N(i,j)/n(i), where N(i,j)N(i,j)N(i,j) counts the observed transitions from iii to j∈S+j \in \mathcal S^+j∈S+, and expected rewards r^(i,j)\hat r(i,j)r^(i,j) equal to the average reward observed on those transitions. With P^=(p^(s′∣s))s,s′∈S\hat P = (\hat p(s'\mid s))_{s,s'\in\mathcal S}P^=(p^​(s′∣s))s,s′∈S​ and r^(s)=∑jp^(j∣s)r^(s,j)\hat r(s) = \sum_j \hat p(j\mid s)\hat r(s,j)r^(s)=∑j​p^​(j∣s)r^(s,j), the certainty-equivalence estimate is the value function of this Markov reward process,

v^(s)=∑k≥0γk(P^kr^)(s).\hat v(s) = \sum_{k\ge0}\gamma^k\big(\hat P^k\hat r\big)(s).v^(s)=k≥0∑​γk(P^kr^)(s).

Formalization targets

Goal: batch TD(0) converges to the certainty-equivalence estimate

For every finite batch and every γ∈[0,1]\gamma \in [0,1]γ∈[0,1], the series defining v^\hat vv^ converges, and there is αˉ>0\bar\alpha > 0αˉ>0 such that for all α∈(0,αˉ)\alpha \in (0,\bar\alpha)α∈(0,αˉ) and all initial arrays V0V_0V0​,

lim⁡m→∞Vm(s)=v^(s)for every visited s,Vm(s)=V0(s) otherwise.\lim_{m\to\infty} V_m(s) = \hat v(s)\quad\text{for every visited } s, \qquad V_m(s) = V_0(s)\ \text{otherwise}.m→∞lim​Vm​(s)=v^(s)for every visited s,Vm​(s)=V0​(s) otherwise.

The limit depends neither on α\alphaα nor on V0V_0V0​ at visited states.

Milestones

  1. (6.6): with VVV held fixed, Gt−V(St)=∑k=tT−1γk−tδkG_t - V(S_t) = \sum_{k=t}^{T-1}\gamma^{k-t}\delta_kGt​−V(St​)=∑k=tT−1​γk−tδk​.
  2. Exercise 6.8: the same identity for action values, δt=Rt+1+γQ(St+1,At+1)−Q(St,At)\delta_t = R_{t+1} + \gamma Q(S_{t+1},A_{t+1}) - Q(S_t,A_t)δt​=Rt+1​+γQ(St+1​,At+1​)−Q(St​,At​).
  3. Least squares: the sample averages Gˉ(s)\bar G(s)Gˉ(s) of the returns after the visits to sss minimize ∑visits(Gt−V(St))2\sum_{\text{visits}}(G_t - V(S_t))^2∑visits​(Gt​−V(St​))2 over all arrays VVV.
  4. Batch MC: for small α\alphaα, batch constant-α\alphaα MC converges to Gˉ(s)\bar G(s)Gˉ(s) at every visited sss.
  5. Fixed points: the batch TD(0) increments vanish everywhere if and only if V=v^V = \hat vV=v^ on visited states.
  6. Example 6.4: on the eight episodes A,0,B,0A,0,B,0A,0,B,0; B,1B,1B,1 (six times); B,0B,0B,0 with γ=1\gamma = 1γ=1, the certainty-equivalence estimate is v^(A)=v^(B)=3/4\hat v(A) = \hat v(B) = 3/4v^(A)=v^(B)=3/4 and batch TD(0) converges to it, while batch MC converges to V(A)=0V(A) = 0V(A)=0, V(B)=3/4V(B) = 3/4V(B)=3/4.

Significance

The result explains the empirical observation of Figure 6.2 in the book: batch TD(0) has lower error than batch MC on Markov data, because it computes the certainty-equivalence estimate, while batch MC fits the training returns. It also gives a precise meaning to the claim that TD methods approximate the certainty-equivalence solution with memory linear in the number of states, where computing it directly needs a model of quadratic size and cubic time (p. 128). Identity (6.6) is the starting point of the nnn-step and eligibility-trace methods of later chapters.

The comparison under repeated presentation of a finite training set goes back to Sutton 1988, but the textbook states the conclusions without proof, and no machine-checked version is known to exist. The formalization pins down every hypothesis the text leaves implicit: the step-size threshold, the treatment of unvisited states, every-visit counting, and the undiscounted case.

Difficulty

The fixed-point equation of batch TD(0) is D(r^+γP^V−V)=0D(\hat r + \gamma\hat PV - V) = 0D(r^+γP^V−V)=0 on visited states, with DDD the diagonal of visit counts, and convergence of the iteration V↦V+αD(r^+γP^V−V)V \mapsto V + \alpha D(\hat r + \gamma\hat P V - V)V↦V+αD(r^+γP^V−V) requires every eigenvalue of D(I−γP^)D(I - \gamma\hat P)D(I−γP^) to have positive real part. For γ<1\gamma < 1γ<1 this follows from P^\hat PP^ being substochastic. For γ=1\gamma = 1γ=1, the case of Example 6.4, P^\hat PP^ is only substochastic and the naive contraction argument fails: invertibility of I−P^I - \hat PI−P^ must be derived from the structure of the data, since every episode ends in the terminal state. The matrix D(I−γP^)D(I - \gamma\hat P)D(I−γP^) is not symmetric, so symmetric positive-definiteness arguments do not apply. The same issue makes convergence of the series defining v^\hat vv^ nontrivial at γ=1\gamma = 1γ=1.

Formalization scope

An episode is a Lean List (X × ℝ) of transitions (St,Rt+1)(S_t, R_{t+1})(St​,Rt+1​), with the terminal state represented by none : Option X; a batch is a list of episodes; states form a Fintype. Values at the terminal state are 000 by definition. The certainty-equivalence estimate is defined from returns as the series ∑kγkP^kr^\sum_k\gamma^k\hat P^k\hat r∑k​γkP^kr^, not as the solution of a Bellman equation, and its convergence is part of the goal, not assumed. "Sufficiently small α\alphaα" is an existential threshold αˉ>0\bar\alpha > 0αˉ>0 quantified before α\alphaα and V0V_0V0​; a statement for one fixed α\alphaα, or for some α\alphaα, would be weaker than the book's and is ruled out. Unvisited states receive no increment and keep their initial value; the goal records this rather than claiming convergence to v^\hat vv^ there. The standing assumption γ∈[0,1]\gamma \in [0,1]γ∈[0,1] includes γ=1\gamma = 1γ=1. Every visit is counted in both the TD increments and the model; mixing first-visit and every-visit counts would make the goal false.

The mission needs only finite sums, matrix powers and limits of real sequences; Mathlib's Matrix and Filter.Tendsto suffice. A lemma that a nonnegative matrix whose rows reach an absorbing mass has spectral radius below one would be reusable beyond this mission, as would a convergence criterion for V↦V+α(b−MV)V \mapsto V + \alpha(b - MV)V↦V+α(b−MV) when MMM is a nonsingular M-matrix. Contributions of either kind, and of the elementary milestones 1–3, are welcome.

Selected references

  • Richard S. Sutton and Andrew G. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press, 2018, ISBN 9780262039246, §6.1 and §6.3, pp. 119–129. http://incompleteideas.net/book/the-book-2nd.html
  • Richard S. Sutton, Learning to predict by the methods of temporal differences, Machine Learning 3, 9–44, 1988. https://doi.org/10.1007/BF00115009
9 thms1 active userReviewed
Operations ResearchProbabilityStochastic Systems·Captain: mikedeng1

Fundamentals of Queueing Theory VII: The Geometric Arrival-Point Law of the G/M/1 QueueTextbook

Motivation

Most queueing models with a closed-form answer assume Poisson arrivals. In practice the times between arrivals are often far from exponential: scheduled appointments, batch releases from an upstream process, or arrivals timed by a machine cycle. The G/M/1 queue keeps the service side exponential and makes no assumption about the arrival stream beyond independent, identically distributed interarrival times. It is the standard counterpart of the M/G/1 queue, and its solution is the one used in teaching and in practice whenever the input is not Poisson (Gross, Shortle, Thompson & Harris, Fundamentals of Queueing Theory, 4th ed., Wiley 2008, §5.3.1, DOI 10.1002/9781118625651).

The answer has an unusually clean form. The number of customers that an arriving customer finds in the system is geometric, exactly as in the M/M/1 queue, with the traffic intensity ρ\rhoρ replaced by a number r0r_0r0​ that depends on the whole interarrival distribution through a single scalar equation. This mission is the seventh of a series formalizing the book chapter by chapter; it covers the G/M/1 half of §5.3 (printed pp.259–263).

Setting

Customers arrive at a single server. The interarrival times are independent with common law AAA, a probability distribution on [0,∞)[0,\infty)[0,∞) with CDF A(t)A(t)A(t) and finite mean E[T]=1/λE[T] = 1/\lambdaE[T]=1/λ, λ>0\lambda > 0λ>0. Service times are independent exponential random variables with rate μ>0\mu > 0μ>0, and the discipline is first come, first served.

Let XnX_nXn​ be the number of customers in the system just before the nnnth arrival. Between two arrivals the server completes a Poisson number of services (truncated by the number present), so {Xn}\{X_n\}{Xn​} is a Markov chain on {0,1,2,… }\{0,1,2,\dots\}{0,1,2,…}. Its transition probabilities are built from

bk=∫0∞e−μt(μt)kk! dA(t)(k≥0),b_k = \int_0^\infty \frac{e^{-\mu t}(\mu t)^k}{k!}\,dA(t) \qquad (k \ge 0),bk​=∫0∞​k!e−μt(μt)k​dA(t)(k≥0),

the probability of exactly kkk completions during one interarrival time (Eq. (5.50)): pi0=1−∑k=0ibkp_{i0} = 1 - \sum_{k=0}^{i} b_kpi0​=1−∑k=0i​bk​, pij=bi+1−jp_{ij} = b_{i+1-j}pij​=bi+1−j​ for 1≤j≤i+11 \le j \le i+11≤j≤i+1, and pij=0p_{ij} = 0pij​=0 otherwise (Eq. (5.51)). A stationary arrival-point distribution is a probability vector q={qn}q = \{q_n\}q={qn​} with qP=qqP = qqP=q and qe=1qe = 1qe=1 (Eq. (5.52)); qnq_nqn​ is the long-run probability that an arrival finds nnn customers present.

The characteristic equation of the chain is

z=β(z),β(z)=∑n≥0bnzn,z = \beta(z), \qquad \beta(z) = \sum_{n \ge 0} b_n z^n ,z=β(z),β(z)=n≥0∑​bn​zn,

where β\betaβ is the probability generating function of {bn}\{b_n\}{bn​} (Eq. (5.55)). Equivalently z=A∗[μ(1−z)]z = A^*[\mu(1-z)]z=A∗[μ(1−z)] (Eq. (5.56)), where A∗(s)=∫0∞e−sx dA(x)A^*(s) = \int_0^\infty e^{-sx}\,dA(x)A∗(s)=∫0∞​e−sxdA(x) is the Laplace–Stieltjes transform of the interarrival law. The traffic intensity is ρ=λ/μ\rho = \lambda/\muρ=λ/μ.

Formalization targets

Goal: Eq. (5.60), the geometric arrival-point law

If ρ=λ/μ<1\rho = \lambda/\mu < 1ρ=λ/μ<1, there is a number r0r_0r0​ with 0<r0<10 < r_0 < 10<r0​<1 and r0=β(r0)r_0 = \beta(r_0)r0​=β(r0​), it is the only complex root of z=β(z)z = \beta(z)z=β(z) in the open unit disk, and

qn=(1−r0) r0 n(n≥0)q_n = (1 - r_0)\, r_0^{\,n} \qquad (n \ge 0)qn​=(1−r0​)r0n​(n≥0)

is a stationary arrival-point distribution and the only one. The root is part of the conclusion, not an assumption.

Milestones

  1. Eqs. (5.51)–(5.53): for a probability vector qqq, qP=qqP = qqP=q is equivalent to qi=∑k≥0qi+k−1bkq_i = \sum_{k\ge0} q_{i+k-1}b_kqi​=∑k≥0​qi+k−1​bk​ (i≥1i \ge 1i≥1) and q0=∑j≥0qj(1−∑k=0jbk)q_0 = \sum_{j\ge0} q_j\bigl(1 - \sum_{k=0}^{j} b_k\bigr)q0​=∑j≥0​qj​(1−∑k=0j​bk​).
  2. p.261: 0<b0<10 < b_0 < 10<b0​<1, bn>0b_n > 0bn​>0 for all nnn, β(1)=1\beta(1) = 1β(1)=1, and β′(1)=∑nnbn=μ/λ\beta'(1) = \sum_n n b_n = \mu/\lambdaβ′(1)=∑n​nbn​=μ/λ.
  3. Eq. (5.56): β(z)=A∗[μ(1−z)]\beta(z) = A^*[\mu(1-z)]β(z)=A∗[μ(1−z)] for ∣z∣≤1|z| \le 1∣z∣≤1.
  4. Eq. (5.58), Figure 5.2: z=β(z)z = \beta(z)z=β(z) has at most one root in (0,1)(0,1)(0,1), and one exists if and only if λ/μ<1\lambda/\mu < 1λ/μ<1.
  5. p.262: when λ/μ<1\lambda/\mu < 1λ/μ<1, z=β(z)z = \beta(z)z=β(z) has exactly one root with ∣z∣<1|z| < 1∣z∣<1.
  6. Eq. (5.59): successive substitution z(k+1)=β(z(k))z^{(k+1)} = \beta(z^{(k)})z(k+1)=β(z(k)) from any 0<z(0)<10 < z^{(0)} < 10<z(0)<1 converges to r0r_0r0​.
  7. Eq. (5.61): L(A)=r0/(1−r0)L^{(A)} = r_0/(1-r_0)L(A)=r0​/(1−r0​) and Lq(A)=r02/(1−r0)L_q^{(A)} = r_0^2/(1-r_0)Lq(A)​=r02​/(1−r0​).
  8. Eq. (5.62): Wq(t)=1−r0e−μ(1−r0)tW_q(t) = 1 - r_0 e^{-\mu(1-r_0)t}Wq​(t)=1−r0​e−μ(1−r0​)t and W(t)=1−e−μ(1−r0)tW(t) = 1 - e^{-\mu(1-r_0)t}W(t)=1−e−μ(1−r0​)t for t≥0t \ge 0t≥0.
  9. Eq. (5.63): Wq=r0/(μ(1−r0))W_q = r_0/(\mu(1-r_0))Wq​=r0​/(μ(1−r0​)) and W=1/(μ(1−r0))W = 1/(\mu(1-r_0))W=1/(μ(1−r0​)).

Significance

The result. Equation (5.60) reduces the analysis of a queue with arbitrary renewal input to one scalar root. Every arrival-point performance measure of the M/M/1 queue then carries over with ρ\rhoρ replaced by r0r_0r0​: the mean number found by an arrival, the mean queue found by an arrival, and the full distributions of line delay and system time seen by arrivals (Eqs. (5.61)–(5.63)). The same root drives the multiserver G/M/c analysis later in §5.3 and the relation between arrival-point and time-average probabilities in §6.3. The result also illustrates a point the book stresses: qnq_nqn​ is the distribution seen by arrivals, and it equals the time-average distribution pnp_npn​ only when the input is Poisson.

Formalizing it. The mathematics is classical (the embedded-chain method goes back to Kendall, 1953) and fully proved in the textbook literature; nothing here is open. To our knowledge none of it has a machine-checked proof: the platform had no G/M/1, embedded-chain, or Rouché-type statement when this mission was drafted. The work is to formalize the known argument, which touches analytic facts about power series with nonnegative coefficients, a mixture-of-Poisson computation, a counting of roots in the unit disk, and the uniqueness of the stationary law of an irreducible countable chain.

Difficulty

Locating a real root in (0,1)(0,1)(0,1) is a one-variable question. The hard step is excluding every other complex root inside the unit disk: a real-variable argument says nothing about complex roots, and the book's route relies on Rouché's theorem, which Mathlib does not have. A second point is uniqueness of the stationary vector: showing that the geometric vector solves qP=qqP = qqP=q does not show that no other probability vector does, and the goal asserts both. Computing ∑nnbn=μ/λ\sum_n n b_n = \mu/\lambda∑n​nbn​=μ/λ requires interchanging a sum with the integral against AAA, which is where the finite mean of the interarrival law enters.

Formalization scope

The interarrival law is a measure A : Measure ℝ with IsProbabilityMeasure A, A (Set.Iio 0) = 0, integrable identity, and ∫ x ∂A = 1/λ (the structure IsInterarrivalLaw). Every theorem also assumes λ>0\lambda > 0λ>0 and μ>0\mu > 0μ>0. The integrals defining bkb_kbk​ and A∗A^*A∗ are over [0,∞)[0,\infty)[0,∞), closed at 000. The generating function β\betaβ takes complex arguments; real roots are written with the real-to-complex coercion. A stationary vector is a function q : ℕ → ℝ with qn≥0q_n \ge 0qn​≥0, HasSum q 1, and HasSum (fun i => q i * p i j) (q j) for every jjj.

The explicit closed forms the statements carry are: the transition matrix (5.51); the equations (5.53); β(z)=A∗[μ(1−z)]\beta(z) = A^*[\mu(1-z)]β(z)=A∗[μ(1−z)] (5.56); β′(1)=μ/λ\beta'(1) = \mu/\lambdaβ′(1)=μ/λ; qn=(1−r0)r0nq_n = (1-r_0)r_0^nqn​=(1−r0​)r0n​ (5.60); r0/(1−r0)r_0/(1-r_0)r0​/(1−r0​) and r02/(1−r0)r_0^2/(1-r_0)r02​/(1−r0​) (5.61); 1−r0e−μ(1−r0)t1 - r_0e^{-\mu(1-r_0)t}1−r0​e−μ(1−r0​)t and 1−e−μ(1−r0)t1 - e^{-\mu(1-r_0)t}1−e−μ(1−r0​)t (5.62); r0/(μ(1−r0))r_0/(\mu(1-r_0))r0​/(μ(1−r0​)) and 1/(μ(1−r0))1/(\mu(1-r_0))1/(μ(1−r0​)) (5.63). The waiting-time CDFs are defined as in §2.2.5 of the book: Wq(t)=q0+∑n≥1qnPr⁡{n completions in≤t}W_q(t) = q_0 + \sum_{n\ge1} q_n \Pr\{n \text{ completions in} \le t\}Wq​(t)=q0​+∑n≥1​qn​Pr{n completions in≤t} with the Erlang type-nnn CDF, and W(t)W(t)W(t) likewise with n+1n+1n+1 completions. The means in (5.63) are ∫0∞[1−Wq(t)] dt\int_0^\infty [1 - W_q(t)]\,dt∫0∞​[1−Wq​(t)]dt and ∫0∞[1−W(t)] dt\int_0^\infty [1 - W(t)]\,dt∫0∞​[1−W(t)]dt.

A trivializing formalization would take "r0∈(0,1)r_0 \in (0,1)r0​∈(0,1) solves z=β(z)z = \beta(z)z=β(z)" as a hypothesis of the goal, which turns (5.60) into a geometric-series check; here existence, location and uniqueness of the root, and uniqueness of the stationary vector, are all conclusions.

Out of scope for this mission: the M/G/c and M/G/∞ results of §5.2 and the multiserver G/M/c analysis of §5.3.2. Reusable pieces include a Rouché-type or fixed-point counting lemma for power series with nonnegative coefficients summing to one, and the uniqueness of stationary laws for irreducible chains on N\mathbb NN. Contributions of either kind are welcome.

Selected references

  • D. Gross, J. F. Shortle, J. M. Thompson, C. M. Harris, Fundamentals of Queueing Theory, 4th ed., Wiley, 2008, §5.3.1, pp.259–263. https://doi.org/10.1002/9781118625651
  • D. G. Kendall, "Stochastic processes occurring in the theory of queues and their analysis by the method of the imbedded Markov chain", Annals of Mathematical Statistics 24(3), 1953, 338–354. https://doi.org/10.1214/aoms/1177728975
12 thms1 active userReviewed
Operations ResearchProbabilityStochastic Systems·Captain: mikedeng1

Fundamentals of Queueing Theory VI: The Pollaczek–Khintchine Transform for the M/G/1 QueueTextbook

Motivation

The M/G/1 queue is the single-server queue with Poisson arrivals and an arbitrary service-time distribution. It is the first queueing model beyond the birth–death family in which exact formulas survive. It is also the model a practitioner reaches for when service times are measured and visibly not exponential: repair times, transmission times of variable-length packets, machining times. Its central result is the Pollaczek–Khintchine formula, first obtained by Pollaczek (1930) and Khintchine (1932). It expresses the stationary queue in terms of the service distribution, and it shows that the mean wait grows linearly in the squared coefficient of variation of service. That makes variability, and not only load, a measurable driver of congestion.

The textbook treatment followed here is Gross, Shortle, Thompson and Harris, Fundamentals of Queueing Theory, 4th ed. (Wiley 2008), §5.1. It derives the result through Kendall's (1953) imbedded Markov chain of system sizes at departure epochs. It then obtains the transforms of the waiting times and the busy-period functional equation of Takács (1962).

Setting

Customers arrive in a Poisson stream of rate λ>0\lambda > 0λ>0. Service times SSS are independent with distribution BBB, a probability distribution on [0,∞)[0,\infty)[0,∞) with mean E[S]\mathrm E[S]E[S], and the discipline is first-come first-served. The traffic intensity is ρ=λ E[S]\rho = \lambda\,\mathrm E[S]ρ=λE[S].

Let XnX_nXn​ be the number of customers the nnnth departing customer leaves behind. The number of arrivals during one service time equals iii with probability

ki=∫0∞e−λt(λt)ii! dB(t),k_i = \int_0^\infty \frac{e^{-\lambda t}(\lambda t)^i}{i!}\,dB(t),ki​=∫0∞​i!e−λt(λt)i​dB(t),

and (Xn)(X_n)(Xn​) is a Markov chain on {0,1,2,… }\{0,1,2,\dots\}{0,1,2,…} whose transition matrix PPP has first row (k0,k1,k2,… )(k_0,k_1,k_2,\dots)(k0​,k1​,k2​,…) and, for i≥1i \ge 1i≥1, entries pij=kj−i+1p_{ij} = k_{j-i+1}pij​=kj−i+1​ for j≥i−1j \ge i-1j≥i−1 and 000 otherwise. A stationary distribution is a probability vector π\piπ with πP=π\pi P = \piπP=π. Its generating function is Π(z)=∑iπizi\Pi(z) = \sum_i \pi_i z^iΠ(z)=∑i​πi​zi, and that of the arrivals per service is K(z)=∑ikiziK(z) = \sum_i k_i z^iK(z)=∑i​ki​zi, for complex ∣z∣≤1|z| \le 1∣z∣≤1. The Laplace–Stieltjes transform of a distribution FFF on [0,∞)[0,\infty)[0,∞) is F∗(s)=∫0∞e−st dF(t)F^*(s) = \int_0^\infty e^{-st}\,dF(t)F∗(s)=∫0∞​e−stdF(t). In the Lean development these are arrivalProb, transitionMatrix, IsStationaryDist, pgf, utilization and lst in the namespace QueueingFundamentals.MG1.

Formalization targets

Goal: the Pollaczek–Khintchine transform formula (5.15)–(5.16)

If E[S]<∞\mathrm E[S] < \inftyE[S]<∞ and ρ<1\rho < 1ρ<1, the chain has a stationary distribution, and every stationary distribution satisfies π0=1−ρ\pi_0 = 1-\rhoπ0​=1−ρ and

Π(z)=(1−ρ)(1−z)K(z)K(z)−z,∣z∣≤1, z≠1,\Pi(z) = \frac{(1-\rho)(1-z)K(z)}{K(z)-z}, \qquad |z| \le 1,\ z \ne 1,Π(z)=K(z)−z(1−ρ)(1−z)K(z)​,∣z∣≤1, z=1,

with K(z)≠zK(z) \ne zK(z)=z at each such zzz. It leaves the service distribution completely general.

Milestones

  1. The stationary equations (5.12): πi=π0ki+∑j=1i+1πjki−j+1\pi_i = \pi_0 k_i + \sum_{j=1}^{i+1}\pi_j k_{i-j+1}πi​=π0​ki​+∑j=1i+1​πj​ki−j+1​.
  2. The transform (5.14), Π(z)=π0(1−z)K(z)/(K(z)−z)\Pi(z) = \pi_0(1-z)K(z)/(K(z)-z)Π(z)=π0​(1−z)K(z)/(K(z)−z), with π0\pi_0π0​ free and no condition on ρ\rhoρ.
  3. Ergodicity (§5.1.4): a unique stationary distribution exists if and only if ρ<1\rho < 1ρ<1.
  4. The departure-point mean (5.7): L(D)=ρ+(ρ2+λ2σB2)/(2(1−ρ))L^{(D)} = \rho + (\rho^2+\lambda^2\sigma_B^2)/(2(1-\rho))L(D)=ρ+(ρ2+λ2σB2​)/(2(1−ρ)).
  5. K(z)=B∗[λ(1−z)]K(z) = B^*[\lambda(1-z)]K(z)=B∗[λ(1−z)] (5.32).
  6. The system-wait transform (5.29), (5.33): Π(z)=W∗[λ(1−z)]\Pi(z) = W^*[\lambda(1-z)]Π(z)=W∗[λ(1−z)] and W∗(s)=(1−ρ)sB∗(s)/(s−λ[1−B∗(s)])W^*(s) = (1-\rho)sB^*(s)/(s-\lambda[1-B^*(s)])W∗(s)=(1−ρ)sB∗(s)/(s−λ[1−B∗(s)]).
  7. The line-wait transform (5.34): Wq∗(s)=(1−ρ)s/(s−λ[1−B∗(s)])W_q^*(s) = (1-\rho)s/(s-\lambda[1-B^*(s)])Wq∗​(s)=(1−ρ)s/(s−λ[1−B∗(s)]).
  8. The busy-period equation (5.37): G∗(s)=B∗[s+λ−λG∗(s)]G^*(s) = B^*[s+\lambda-\lambda G^*(s)]G∗(s)=B∗[s+λ−λG∗(s)].
  9. The mean busy period: E[X]=1/(μ−λ)\mathrm E[X] = 1/(\mu-\lambda)E[X]=1/(μ−λ) with μ=1/E[S]\mu = 1/\mathrm E[S]μ=1/E[S].

Significance

The transform formula determines the whole stationary departure-point distribution from the service distribution. Its derivatives at z=1z = 1z=1 give every moment of the system size, including the mean-value formula (5.7). Combined with the transform identity (5.32), it gives the waiting-time transforms (5.33)–(5.34). Those in turn give the classical geometric-series representation of the line-wait distribution through the residual service time. The busy-period equation is the starting point for busy-period moments and for the M/G/1 analysis of priority and vacation models later in the book.

All results here are classical and proved in the literature. As far as a search of the platform shows (2026-09-28), none is machine-checked: there is no M/G/1 queue, imbedded departure-point chain, Laplace–Stieltjes transform of a service distribution, or busy-period equation on Prove2Me. Mathlib has Poisson distributions and measure convolution but no generating-function theory for countable Markov chains, no Laplace–Stieltjes transform, and no identity theorem in the form these statements need. The mission produces a checked statement of the Pollaczek–Khintchine formulas that later queueing developments (vacations, priorities, M/G/1-type chains) can build on.

Difficulty

Turning the stationary equations into (5.14) is formal power-series algebra. The difficulties lie elsewhere. First, the formula must hold for complex zzz on the closed disk, which needs the non-vanishing of K(z)−zK(z)-zK(z)−z away from z=1z = 1z=1. That fact fails for ρ>1\rho > 1ρ>1, where KKK has a fixed point inside the disk. Second, (5.15) evaluates π0\pi_0π0​ from Π(1)=1\Pi(1) = 1Π(1)=1 by a limit at the point where the formula is 0/00/00/0, and this uses K′(1)=ρK'(1) = \rhoK′(1)=ρ, an interchange of sum and integral. Third, the existence half of the goal requires positive recurrence of a chain with unbounded jumps. The book obtains it from Foster's criterion, which is not in Mathlib. Fourth, the waiting-time and busy-period transforms are stated for all real s>0s > 0s>0, while the generating-function route reaches only s=λ(1−z)∈(0,2λ]s = \lambda(1-z) \in (0, 2\lambda]s=λ(1−z)∈(0,2λ]. Extending the identity requires either analyticity arguments or a direct derivation. A formal proof of (5.14) alone does not touch any of these.

Formalization scope

The service distribution is a Measure ℝ with IsProbabilityMeasure B and B (Set.Iio 0) = 0; no density is assumed. The arrival rate is lam : ℝ with 0 < lam. Stationarity is IsStationaryDist P π: nonnegative entries, HasSum π 1, and HasSum (fun i => π i * P i j) (π j) for every j. That is global balance on ℕ, as the book writes it. Generating functions take a complex argument with ‖z‖ ≤ 1; transforms take a complex argument, and the waiting-time and busy-period statements use real s. The mean and variance of B are Bochner integrals, and every statement that uses them assumes Integrable. The mean busy period assumes 0 < E[S], so that μ=1/E[S]\mu = 1/\mathrm E[S]μ=1/E[S] is the book's service rate.

Closed forms carried by the statements: π0=1−ρ\pi_0 = 1-\rhoπ0​=1−ρ (5.15); (1−ρ)(1−z)K(z)/(K(z)−z)(1-\rho)(1-z)K(z)/(K(z)-z)(1−ρ)(1−z)K(z)/(K(z)−z) (5.16); π0(1−z)K(z)/(K(z)−z)\pi_0(1-z)K(z)/(K(z)-z)π0​(1−z)K(z)/(K(z)−z) (5.14); ρ+(ρ2+λ2σB2)/(2(1−ρ))\rho + (\rho^2+\lambda^2\sigma_B^2)/(2(1-\rho))ρ+(ρ2+λ2σB2​)/(2(1−ρ)) (5.7); B∗[λ(1−z)]B^*[\lambda(1-z)]B∗[λ(1−z)] (5.32); (1−ρ)sB∗(s)/(s−λ[1−B∗(s)])(1-\rho)sB^*(s)/(s-\lambda[1-B^*(s)])(1−ρ)sB∗(s)/(s−λ[1−B∗(s)]) (5.33); (1−ρ)s/(s−λ[1−B∗(s)])(1-\rho)s/(s-\lambda[1-B^*(s)])(1−ρ)s/(s−λ[1−B∗(s)]) (5.34); B∗[s+λ−λG∗(s)]B^*[s+\lambda-\lambda G^*(s)]B∗[s+λ−λG∗(s)] (5.37); 1/(μ−λ)1/(\mu-\lambda)1/(μ−λ) for the mean busy period.

The waiting-time distribution WWW enters through the book's FCFS relation πn=1n!∫(λt)ne−λt dW(t)\pi_n = \frac1{n!}\int(\lambda t)^n e^{-\lambda t}\,dW(t)πn​=n!1​∫(λt)ne−λtdW(t), and WqW_qWq​ through W=Wq∗BW = W_q * BW=Wq​∗B; both are hypotheses, as in the book. The busy-period distribution GGG enters through the equation (5.36) in CDF form, with nnn-fold convolutions built from Mathlib's Measure.conv.

The goal is not the algebraic consequence of (5.12) for an arbitrary sequence: π\piπ must be a probability vector, π0\pi_0π0​ is determined as 1−ρ1-\rho1−ρ, and the existence of a stationary distribution is part of the conclusion, so the statement cannot hold vacuously. The departure-point/time-average equality (§5.1.3, via PASTA) is not formalized.

Useful infrastructure, reusable beyond this mission: generating functions of stationary distributions on ℕ, Poisson mixtures, Laplace–Stieltjes transforms of measures on [0,∞)[0,\infty)[0,∞), and a Foster-type drift criterion for countable chains. Contributions proving any milestone, or those tools, are welcome.

Selected references

  • D. Gross, J. F. Shortle, J. M. Thompson, C. M. Harris, Fundamentals of Queueing Theory, 4th ed., Wiley, 2008, §5.1. https://doi.org/10.1002/9781118625651
  • D. G. Kendall, Stochastic processes occurring in the theory of queues and their analysis by the method of the imbedded Markov chain, Annals of Mathematical Statistics 24 (1953) 338–354. https://doi.org/10.1214/aoms/1177728975
  • F. G. Foster, On the stochastic matrices associated with certain queuing processes, Annals of Mathematical Statistics 24 (1953) 355–360. https://doi.org/10.1214/aoms/1177728976
  • L. Takács, Introduction to the Theory of Queues, Oxford University Press, 1962.
  • F. Pollaczek, Über eine Aufgabe der Wahrscheinlichkeitstheorie, Mathematische Zeitschrift 32 (1930) 64–100. https://doi.org/10.1007/BF01194620
12 thms1 active userReviewed
Operations ResearchProbabilityStochastic Systems·Captain: mikedeng1

Fundamentals of Queueing Theory V: Closed Jackson Networks and the Mean-Value RecursionTextbook

Motivation

Networks of queues model systems in which a job visits several service stations in turn: jobs in a computer system alternating between CPU and disks, machines cycling between operation and repair, parts routed through a job shop. In a closed network no job enters or leaves; a fixed population of NNN customers circulates among kkk nodes. Closed networks are the standard model of multiprogrammed computer systems and of machine-repair and finite-source systems, and they are the setting of chapter 4 of Gross, Shortle, Thompson and Harris, Fundamentals of Queueing Theory (4th ed., Wiley 2008, doi:10.1002/9781118625651).

The chapter's results form a short line of computational ideas:

  • Jackson (1957, 1963) showed that open networks of exponential servers with Markovian routing have a product-form steady state; Gordon and Newell (1967) gave the closed-network version, (4.15)–(4.18) of the book.
  • Buzen (1973) gave a convolution recursion for the normalizing constant G(N)G(N)G(N) and for marginal distributions, (4.19)–(4.22).
  • Reiser and Lavenberg (1980) introduced mean-value analysis (MVA), which computes mean queue lengths, waiting times and throughputs population by population without ever forming G(N)G(N)G(N), (4.23)–(4.25); the book presents it following Bruell and Balbo (1980).
  • The book closes the section with a recursion for the full marginal distributions, (4.26), which it proves from the product form (pp.207–209).

This mission formalizes that line, ending at (4.26).

Setting

A closed Jackson network has nodes i=1,…,ki = 1, \dots, ki=1,…,k, each with a single server whose service times are exponential with rate μi>0\mu_i > 0μi​>0. A customer finishing service at node iii moves to node jjj with probability rijr_{ij}rij​; the routing matrix R=(rij)R = (r_{ij})R=(rij​) has nonnegative entries and rows summing to one, and it is irreducible: every node can be reached from every other. The state is nˉ=(n1,…,nk)\bar n = (n_1, \dots, n_k)nˉ=(n1​,…,nk​), the number of customers at each node, with n1+⋯+nk=Nn_1 + \cdots + n_k = Nn1​+⋯+nk​=N; this state space is finite.

The steady-state distribution pnˉp_{\bar n}pnˉ​ is the probability vector on the state space that solves the flow-balance equations (4.14),

∑j=1k∑i=1i≠jkμirij pnˉ;i+j−=∑i=1kμi(1−rii) pnˉ,\sum_{j=1}^{k}\sum_{\substack{i=1\\ i\ne j}}^{k} \mu_i r_{ij}\, p_{\bar n;i^+j^-} = \sum_{i=1}^{k}\mu_i(1-r_{ii})\,p_{\bar n},j=1∑k​i=1i=j​∑k​μi​rij​pnˉ;i+j−​=i=1∑k​μi​(1−rii​)pnˉ​,

where nˉ;i+j−\bar n;i^+j^-nˉ;i+j− has one more customer at iii and one fewer at jjj, and terms with a negative subscript or with μi\mu_iμi​ at an empty node vanish. The traffic equations (4.16) are μiρi=∑jμjrjiρj\mu_i\rho_i = \sum_j \mu_j r_{ji}\rho_jμi​ρi​=∑j​μj​rji​ρj​; they determine ρ=(ρ1,…,ρk)\rho = (\rho_1, \dots, \rho_k)ρ=(ρ1​,…,ρk​) up to a positive factor. The normalizing constant is

G(N)=∑n1+⋯+nk=Nρ1n1⋯ρknk,G(N) = \sum_{n_1+\cdots+n_k=N}\rho_1^{n_1}\cdots\rho_k^{n_k},G(N)=n1​+⋯+nk​=N∑​ρ1n1​​⋯ρknk​​,

and more generally, with fi(n)=ρi n/ai(n)f_i(n) = \rho_i^{\,n}/a_i(n)fi​(n)=ρin​/ai​(n) for cic_ici​-server nodes ((4.13)), G(N)=∑∏ifi(ni)G(N) = \sum \prod_i f_i(n_i)G(N)=∑∏i​fi​(ni​) and Buzen's function gm(n)=∑n1+⋯+nm=n∏i≤mfi(ni)g_m(n) = \sum_{n_1+\cdots+n_m=n}\prod_{i\le m} f_i(n_i)gm​(n)=∑n1​+⋯+nm​=n​∏i≤m​fi​(ni​).

For each population NNN write pi(n,N)=Pr⁡{Ni=n}p_i(n, N) = \Pr\{N_i = n\}pi​(n,N)=Pr{Ni​=n} for the marginal distribution at node iii, Pˉi(n;N)=Pr⁡{Ni≥n}\bar P_i(n; N) = \Pr\{N_i \ge n\}Pˉi​(n;N)=Pr{Ni​≥n}, Li(N)L_i(N)Li​(N) for the mean number at node iii, and

λi(N)=Pr⁡{server busy at node i}⋅μi\lambda_i(N) = \Pr\{\text{server busy at node } i\}\cdot\mu_iλi​(N)=Pr{server busy at node i}⋅μi​

for the throughput of node iii.

Formalization targets

Goal: the marginal recursion (4.26)

For every node iii,

pi(0,0)=1,pi(n,N)=λi(N)μi pi(n−1,N−1)(n,N≥1).p_i(0,0) = 1, \qquad p_i(n, N) = \frac{\lambda_i(N)}{\mu_i}\,p_i(n-1, N-1) \quad (n, N \ge 1).pi​(0,0)=1,pi​(n,N)=μi​λi​(N)​pi​(n−1,N−1)(n,N≥1).

It involves only the steady-state distributions and quantities computed from them; it holds for every irreducible routing matrix and every choice of rates.

Milestones

  1. Product form (4.14)–(4.16). For any positive solution ρ\rhoρ of (4.16), a probability distribution solves (4.14) if and only if pnˉ=G(N)−1ρ1n1⋯ρknkp_{\bar n} = G(N)^{-1}\rho_1^{n_1}\cdots\rho_k^{n_k}pnˉ​=G(N)−1ρ1n1​​⋯ρknk​​.
  2. Buzen's algorithm (4.19)–(4.21). G(N)=gk(N)G(N) = g_k(N)G(N)=gk​(N), gm(n)=∑i=0nfm(i) gm−1(n−i)g_m(n) = \sum_{i=0}^{n} f_m(i)\,g_{m-1}(n-i)gm​(n)=∑i=0n​fm​(i)gm−1​(n−i), g1=f1g_1 = f_1g1​=f1​, gm(0)=1g_m(0) = 1gm​(0)=1.
  3. Marginal at the last node (4.22). pk(n)=fk(n) gk−1(N−n)/G(N)p_k(n) = f_k(n)\,g_{k-1}(N-n)/G(N)pk​(n)=fk​(n)gk−1​(N−n)/G(N) for 0≤n≤N0 \le n \le N0≤n≤N.
  4. Complementary marginal (p.208). Pˉi(ni;N)=ρi niG(N−ni)/G(N)\bar P_i(n_i; N) = \rho_i^{\,n_i}G(N-n_i)/G(N)Pˉi​(ni​;N)=ρini​​G(N−ni​)/G(N).
  5. Mean-value analysis (4.23)–(4.25). Li(0)=0L_i(0) = 0Li​(0)=0; Li(N)=λi(N)Wi(N)L_i(N) = \lambda_i(N)W_i(N)Li​(N)=λi​(N)Wi​(N) with Wi(N)=(1+Li(N−1))/μiW_i(N) = (1 + L_i(N-1))/\mu_iWi​(N)=(1+Li​(N−1))/μi​; and for vvv solving vi=∑jvjrjiv_i = \sum_j v_j r_{ji}vi​=∑j​vj​rji​ with vl=1v_l = 1vl​=1, λl(N)=N/∑iviWi(N)\lambda_l(N) = N/\sum_i v_iW_i(N)λl​(N)=N/∑i​vi​Wi​(N) and λi(N)=λl(N)vi\lambda_i(N) = \lambda_l(N)v_iλi​(N)=λl​(N)vi​.

Significance

The product form reduces a (N+k−1N)\binom{N+k-1}{N}(NN+k−1​)-state Markov chain to the constants G(0),…,G(N)G(0), \dots, G(N)G(0),…,G(N), and Buzen's recursion computes them in O(kN2)O(kN^2)O(kN2) operations. Mean-value analysis goes further and avoids G(N)G(N)G(N), whose magnitude can overflow or underflow for large populations; it is the method used in capacity planning of computer systems. The recursion (4.26) extends MVA from means to full marginal distributions, so a single pass over NNN yields every nodal distribution.

All of these results are classical and proved in the literature; the book proves (4.26) itself. What the mission adds is a machine-checked development of them from the global balance equations: the product form with its uniqueness, the convolution identities, the marginal formulas, and the correctness of the MVA iteration as stated by the book, all over one shared definition layer. A search of the platform on 2026-09-28 found no formal statement of Buzen's algorithm or of MVA. The platform has Kelly's closed migration process theorem (KellyStochasticNetworks.closed_migration_equilibrium), which shows that the unnormalized product form satisfies the equilibrium equations under Kelly's conventions; the normalization, uniqueness and everything downstream of the product form are new here.

Difficulty

The combinatorial identities (Buzen's recursion, the tail marginal) are reindexings of finite sums over compositions of NNN; in Lean the work is in bijections between the state spaces {n1+⋯+nk=N}\{n_1+\cdots+n_k = N\}{n1​+⋯+nk​=N} for different kkk and NNN. The substantive step is uniqueness in the product-form theorem: the global balance equations have a one-dimensional solution space only because the chain on the NNN-customer states is irreducible on the population level, which is a property of the network chain and not of the routing matrix alone. The goal and MVA also need a positive solution of the traffic equations, which is not among the hypotheses and has to come from irreducibility of RRR. The book's own intuitive derivation of MVA via the arrival theorem is not the route the statements require; they are stated in terms of the steady-state distributions alone.

Formalization scope

Nodes are Fin k (book node iii is index i−1i-1i−1); states are n : Fin k → ℕ with ∑ i, n i = N, collected in a Finset, and all sums are finite. A distribution is a real function on Nk\mathbb N^kNk that is nonnegative, vanishes off the NNN-customer states and sums to one there. The balance equations are (4.14) verbatim with the book's boundary convention (p.188), not detailed balance. All results except Buzen's algorithm and (4.22) are for single-server nodes, as in the book; (4.13)'s multiserver factor ai(n)a_i(n)ai​(n) enters only (4.19)–(4.22).

Closed forms instantiated in the statements: the product form G(N)−1∏iρiniG(N)^{-1}\prod_i\rho_i^{n_i}G(N)−1∏i​ρini​​ ((4.15)); G(N)G(N)G(N) as the explicit sum (4.18)/(4.19); ai(n)a_i(n)ai​(n) from (4.13); gmg_mgm​ from (4.20); pk(n)=fk(n)gk−1(N−n)/G(N)p_k(n) = f_k(n)g_{k-1}(N-n)/G(N)pk​(n)=fk​(n)gk−1​(N−n)/G(N) ((4.22)); Pˉi(n;N)=ρinG(N−n)/G(N)\bar P_i(n;N) = \rho_i^nG(N-n)/G(N)Pˉi​(n;N)=ρin​G(N−n)/G(N) (p.208); Wi(N)=(1+Li(N−1))/μiW_i(N) = (1+L_i(N-1))/\mu_iWi​(N)=(1+Li​(N−1))/μi​ ((4.23)); λl(N)=N/∑iviWi(N)\lambda_l(N) = N/\sum_i v_iW_i(N)λl​(N)=N/∑i​vi​Wi​(N) (MVA step (iii)(b)).

Two trivializing formalizations are ruled out: λi(N)\lambda_i(N)λi​(N) in (4.26) and (4.24) is the throughput computed from the steady-state distribution, not a free constant (which would make (4.26) a definition); and gmg_mgm​ is defined by the sum (4.20), so the recursion (4.21) is a theorem rather than rfl. The product-form statement is an equivalence, so it asserts both that the product form is a steady state and that it is the only one.

Needed infrastructure: bijections between compositions of NNN into kkk and k−1k-1k−1 parts, uniqueness of stationary distributions of irreducible finite continuous-time chains (stated directly via the balance equations), and existence of positive solutions of v=vRv = vRv=vR for irreducible stochastic RRR. The last two are reusable beyond this mission. Contributions welcome: proofs of the milestones in any order, and helper lemmas on these three points.

Not formalized: open Jackson networks (4.11) and Burke's theorem (4.5)–(4.6), multiclass networks (§4.2.1), the multiserver recursion (4.27) and cyclic queues (§4.4).

Selected references

  • D. Gross, J. F. Shortle, J. M. Thompson, C. M. Harris, Fundamentals of Queueing Theory, 4th ed., Wiley, 2008, §4.3, pp.195–209. https://doi.org/10.1002/9781118625651
  • J. R. Jackson, "Jobshop-like queueing systems", Management Science 10(1), 1963. https://doi.org/10.1287/mnsc.10.1.131
  • W. J. Gordon, G. F. Newell, "Closed queuing systems with exponential servers", Operations Research 15(2), 1967. https://doi.org/10.1287/opre.15.2.254
  • J. P. Buzen, "Computational algorithms for closed queueing networks with exponential servers", Communications of the ACM 16(9), 1973. https://doi.org/10.1145/362342.362345
  • M. Reiser, S. S. Lavenberg, "Mean-value analysis of closed multichain queuing networks", Journal of the ACM 27(2), 1980. https://doi.org/10.1145/322186.322195
  • S. C. Bruell, G. Balbo, Computational Algorithms for Closed Queueing Networks, North-Holland, 1980.
7 thms1 active userReviewed
Operations ResearchProbabilityStochastic Systems·Captain: mikedeng1

Fundamentals of Queueing Theory IV: The Stationary Distribution of the M/M/1 Retrial QueueTextbook

Motivation

In many service systems a customer who finds every server busy does not join a queue. A caller who hears a busy signal hangs up and redials later; a request rejected by a saturated server is resent after a timeout; an aircraft that cannot land circles and tries again. These retrial queues are the subject of a substantial literature in telephone traffic engineering, computer networks and call-centre design, surveyed in the monograph of Falin and Templeton (1997) and the bibliography of Artalejo (1999). Their analysis is harder than that of ordinary queues: the blocked customers form an orbit whose size is part of the state, so even the simplest model is a two-dimensional Markov chain, and explicit stationary distributions are rare.

This mission is the fourth of a series formalizing Gross, Shortle, Thompson and Harris, Fundamentals of Queueing Theory (4th ed., Wiley 2008). Its goal is the explicit stationary distribution of the single-server retrial queue, Eq. (3.57) of §3.5.1, one of the few retrial models solvable in closed form. Chapter 3 of the book treats Markovian queues that are not birth–death processes: bulk arrivals, bulk service, Erlang phases, priority disciplines and retrials. The milestones also collect three capstone formulas from the chapter's other sections: the bulk-input queue, the partial-batch bulk-service queue, and Cobham's formula for nonpreemptive priorities (Cobham, 1954).

Setting

In the M/M/1M/M/1M/M/1 retrial queue customers arrive according to a Poisson process with rate λ\lambdaλ and are served one at a time by a single server, with exponential service times of mean 1/μ1/\mu1/μ. An arrival that finds the server busy enters the orbit and stays there for an exponential time with mean 1/γ1/\gamma1/γ, after which it tries again; each customer in orbit retries independently. No customer leaves because of impatience. With Ns(t)∈{0,1}N_s(t) \in \{0,1\}Ns​(t)∈{0,1} the number in service and No(t)N_o(t)No​(t) the number in orbit, the pair is a continuous-time Markov chain on states {i,n}\{i, n\}{i,n}, i∈{0,1}i \in \{0,1\}i∈{0,1}, n∈{0,1,2,… }n \in \{0,1,2,\dots\}n∈{0,1,2,…}. Writing pi,np_{i,n}pi,n​ for the steady-state probability of {i,n}\{i,n\}{i,n}, the rate-balance equations are

(λ+nγ)p0,n=μp1,n,n≥0,(3.47)(λ+μ)p1,n=λp0,n+(n+1)γp0,n+1+λp1,n−1,n≥1,(3.48)(λ+μ)p1,0=λp0,0+γp0,1.(3.49)\begin{aligned} (\lambda + n\gamma)p_{0,n} &= \mu p_{1,n}, && n \ge 0, && (3.47)\\ (\lambda+\mu)p_{1,n} &= \lambda p_{0,n} + (n+1)\gamma p_{0,n+1} + \lambda p_{1,n-1}, && n \ge 1, && (3.48)\\ (\lambda+\mu)p_{1,0} &= \lambda p_{0,0} + \gamma p_{0,1}. && && (3.49) \end{aligned}(λ+nγ)p0,n​(λ+μ)p1,n​(λ+μ)p1,0​​=μp1,n​,=λp0,n​+(n+1)γp0,n+1​+λp1,n−1​,=λp0,0​+γp0,1​.​​n≥0,n≥1,​​(3.47)(3.48)(3.49)​

Following the book's convention (§1.9, and the footnote on p.118), a steady-state solution is a nonnegative solution of these equations whose total mass ∑n(p0,n+p1,n)\sum_n (p_{0,n} + p_{1,n})∑n​(p0,n​+p1,n​) equals 111. The traffic intensity is ρ=λ/μ\rho = \lambda/\muρ=λ/μ, and the partial generating functions are P0(z)=∑nznp0,nP_0(z) = \sum_n z^n p_{0,n}P0​(z)=∑n​znp0,n​ and P1(z)=∑nznp1,nP_1(z) = \sum_n z^n p_{1,n}P1​(z)=∑n​znp1,n​.

The other models of the mission use the same convention. In the bulk-input queue M[X]/M/1M^{[X]}/M/1M[X]/M/1, batches arrive at rate λ\lambdaλ with batch-size probabilities cn=Pr⁡{X=n}c_n = \Pr\{X = n\}cn​=Pr{X=n}, n≥1n \ge 1n≥1, and batch-size generating function C(z)=∑ncnznC(z) = \sum_n c_n z^nC(z)=∑n​cn​zn. In the partial-batch bulk-service queue M/M[K]/1M/M^{[K]}/1M/M[K]/1, single arrivals come at rate λ\lambdaλ and the server serves up to KKK customers together in an exponential time of mean 1/μ1/\mu1/μ. In the nonpreemptive priority queue there are rrr classes with rates λk\lambda_kλk​ and μk\mu_kμk​, loads ρk=λk/μk\rho_k = \lambda_k/\mu_kρk​=λk​/μk​ and cumulative loads σk=ρ1+⋯+ρk\sigma_k = \rho_1 + \cdots + \rho_kσk​=ρ1​+⋯+ρk​.

Formalization targets

Goal: the stationary distribution (3.57)

For λ,μ,γ>0\lambda, \mu, \gamma > 0λ,μ,γ>0 and ρ<1\rho < 1ρ<1, the numbers

p0,n=(1−ρ)(λ/γ)+1ρnn! γn∏i=0n−1(λ+iγ),p1,n=(1−ρ)(λ/γ)+1ρn+1n! γn∏i=1n(λ+iγ)p_{0,n} = (1-\rho)^{(\lambda/\gamma)+1}\frac{\rho^n}{n!\,\gamma^n}\prod_{i=0}^{n-1}(\lambda+i\gamma), \qquad p_{1,n} = (1-\rho)^{(\lambda/\gamma)+1}\frac{\rho^{n+1}}{n!\,\gamma^n}\prod_{i=1}^{n}(\lambda+i\gamma)p0,n​=(1−ρ)(λ/γ)+1n!γnρn​i=0∏n−1​(λ+iγ),p1,n​=(1−ρ)(λ/γ)+1n!γnρn+1​i=1∏n​(λ+iγ)

form a steady-state solution of (3.47)–(3.49), and every steady-state solution equals them.

Milestones on the retrial queue

The generating functions satisfy (3.50)–(3.52) on (−1,1)(-1,1)(−1,1), including the separable equation

P0′(z)=λργ(1−ρz)P0(z),P_0'(z) = \frac{\lambda\rho}{\gamma(1-\rho z)}P_0(z),P0′​(z)=γ(1−ρz)λρ​P0​(z),

their closed form is (3.55),

P0(z)=(1−ρz)(1−ρ1−ρz)(λ/γ)+1,P1(z)=ρ(1−ρ1−ρz)(λ/γ)+1,P_0(z) = (1-\rho z)\left(\frac{1-\rho}{1-\rho z}\right)^{(\lambda/\gamma)+1}, \qquad P_1(z) = \rho\left(\frac{1-\rho}{1-\rho z}\right)^{(\lambda/\gamma)+1},P0​(z)=(1−ρz)(1−ρz1−ρ​)(λ/γ)+1,P1​(z)=ρ(1−ρz1−ρ​)(λ/γ)+1,

and the mean orbit size is (3.58), Lo=ρ21−ρ⋅μ+γγL_o = \frac{\rho^2}{1-\rho}\cdot\frac{\mu+\gamma}{\gamma}Lo​=1−ρρ2​⋅γμ+γ​.

Milestones from the rest of Chapter 3

The bulk-input generating function (3.3), p0=1−ρp_0 = 1 - \rhop0​=1−ρ with ρ=λE[X]/μ\rho = \lambda\mathrm E[X]/\muρ=λE[X]/μ, and the mean (3.4); the unique root r0∈(0,1)r_0 \in (0,1)r0​∈(0,1) of μrK+1−(λ+μ)r+λ=0\mu r^{K+1} - (\lambda+\mu)r + \lambda = 0μrK+1−(λ+μ)r+λ=0 and the geometric law pn=(1−r0)r0np_n = (1-r_0)r_0^npn​=(1−r0​)r0n​ (3.9); and Cobham's formula (3.41)/(3.43), the unique solution of the linear system (3.40).

Significance

The closed form (3.57) makes every performance measure of the M/M/1M/M/1M/M/1 retrial queue explicit. The server is busy a fraction ρ\rhoρ of the time, exactly as without retrials. The mean orbit size (3.58) is the M/M/1M/M/1M/M/1 mean queue length multiplied by (μ+γ)/γ(\mu+\gamma)/\gamma(μ+γ)/γ, and the mean time in orbit (3.59) follows from Little's law. These formulas quantify the cost of retrials against an ordinary queue and are the reference case against which approximations for multi-server retrial systems are checked.

The results are classical and proved in the book, partly through exercises (Problems 3.39–3.41). None of them is formalized in any proof assistant, as far as the platform's catalogue shows: there is no retrial, bulk or priority queue on Prove2Me. The mission produces machine-checked statements and, once solved, proofs of the chapter's main closed forms. It also produces a small reusable layer: generating functions of probability sequences on the closed unit disc, and the "probability solution of the balance equations" pattern for chains with countable state spaces.

Difficulty

The derivation in the book is formal. It differentiates power series term by term, divides by 1−z1 - z1−z, integrates ln⁡P0\ln P_0lnP0​, and fixes the constant by setting z=1z = 1z=1, without justifying any of these steps. A formal proof has to show that the series converge and are differentiable on (−1,1)(-1,1)(−1,1), that the differential equation determines P0P_0P0​ up to a constant, and that the values at z=1z = 1z=1 are the limits of the values inside the disc (Abel's theorem). The uniqueness half of the goal is the hardest part. The book never proves it; it follows from the ODE argument only once every step is shown to hold for an arbitrary probability solution. Verifying that (3.57) solves (3.47)–(3.49) is only the easy half. The same pattern recurs in the bulk-input queue, where z=1z = 1z=1 is a removable singularity of (3.3). In the bulk-service queue the root r0r_0r0​ is only characterized as the unique root in (0,1)(0,1)(0,1), so existence and uniqueness of the root are part of the claim.

Formalization scope

A steady-state solution is a pair p0 p1 : ℕ → ℝ (resp. one sequence p : ℕ → ℝ) that is pointwise nonnegative, has total mass 111 as a HasSum, and solves the book's balance equations exactly as printed, global balance and not detailed balance. Every "the steady-state solution is X" is stated with both halves: X is a steady-state solution, and every steady-state solution equals X. Stating only that (3.57) solves (3.47)–(3.49), without normalization or uniqueness, would be a trivializing formalization. So would taking r0r_0r0​ as a given root in (3.9), or taking the Wq(i)W_q^{(i)}Wq(i)​ in (3.41) as numbers assumed to satisfy it. None of these is used. The closed forms instantiated are (3.52), (3.55), (3.57), (3.58), (3.3), (3.4), (3.9), (3.41) and (3.43), each written out in full, with the real power (1−ρ)(λ/γ)+1(1-\rho)^{(\lambda/\gamma)+1}(1−ρ)(λ/γ)+1 as Real.rpow.

The conventions are as follows. The retrial generating functions take real arguments, on (−1,1)(-1,1)(−1,1) for the differential equations and on [−1,1][-1,1][−1,1] for the closed form. The bulk-input generating function takes complex arguments with ∣z∣≤1|z| \le 1∣z∣≤1, z≠1z \ne 1z=1, because (3.3) is 0/00/00/0 at z=1z = 1z=1. The condition ρ<1\rho < 1ρ<1 is a hypothesis of every retrial statement. For the bulk-service queue the book's unnamed condition is stated as λ<Kμ\lambda < K\muλ<Kμ. For bulk input, E[X]<∞\mathrm E[X] < \inftyE[X]<∞ is assumed throughout, and the mean (3.4) is asserted under the further condition E[X2]<∞\mathrm E[X^2] < \inftyE[X2]<∞, which it requires. For Cobham's formula only the algebraic content is formalized; the mean-value argument that yields (3.40) and (3.42) is not.

Needed infrastructure: power series of summable nonnegative sequences on the closed unit disc (convergence, term-by-term differentiation, Abel continuity), the binomial series (1−x)−a=∑na(a+1)⋯(a+n−1)n!xn(1 - x)^{-a} = \sum_n \frac{a(a+1)\cdots(a+n-1)}{n!}x^n(1−x)−a=∑n​n!a(a+1)⋯(a+n−1)​xn for real aaa, and uniqueness of invariant probability vectors for irreducible chains. All of this is reusable beyond the mission. Proofs of any milestone, of the easy half of the goal, or of the needed series facts are welcome contributions.

Selected references

  • D. Gross, J. F. Shortle, J. M. Thompson, C. M. Harris, Fundamentals of Queueing Theory, 4th ed., Wiley, 2008, §§3.1, 3.2.0.1, 3.4.2, 3.5.1. https://doi.org/10.1002/9781118625651
  • G. I. Falin, J. G. C. Templeton, Retrial Queues, Chapman & Hall, 1997. https://doi.org/10.1007/978-1-4899-2977-8
  • J. R. Artalejo, Accessible bibliography on retrial queues, Mathematical and Computer Modelling 30 (1999) 1–6. https://doi.org/10.1016/S0895-7177(99)00128-4
  • A. Cobham, Priority assignment in waiting line problems, Journal of the Operations Research Society of America 2 (1954) 70–76. https://doi.org/10.1287/opre.2.1.70
10 thms1 active userReviewed
AnalysisOperations ResearchProbability+1·Captain: mikedeng1

Fundamentals of Queueing Theory III: The Transient M/M/1 Queue via Modified Bessel FunctionsTextbook

Motivation

Steady-state formulas describe a queue that has been running forever. Many practical questions are about a queue that has not: a call centre just after opening, a server just after a reset, a system under a burst of load. For these, the relevant quantity is the transient distribution pn(t)=Pr⁡{N(t)=n}p_n(t) = \Pr\{N(t) = n\}pn​(t)=Pr{N(t)=n} of the number N(t)N(t)N(t) in the system at a finite time ttt. It is also what determines how fast the steady state is approached, and it is needed for the busy period: the length of time a server stays busy once a customer arrives at an idle server.

For the single-server Markovian queue M/M/1 the transient distribution has an explicit closed form in modified Bessel functions. Its history is short and well documented. Ledermann and Reuter (1954) obtained it by spectral analysis of the birth–death process. Bailey (1954) found it by generating functions and Laplace transforms, and Champernowne (1956) by combinatorial methods. Bailey's route is the standard textbook derivation, and it is the one Gross, Shortle, Thompson and Harris outline in §2.11 of Fundamentals of Queueing Theory (4th ed., 2008). Abate and Whitt (1989) showed that computing with the resulting series is numerically delicate, which is one reason for having the formula pinned down exactly.

This mission formalizes §§2.11–2.12 of that book: the transient laws of M/M/1/1, M/M/1 and M/M/∞, and the M/M/1 busy period.

Setting

Customers arrive in a Poisson stream of rate λ>0\lambda > 0λ>0. Each service takes an exponential time of rate μ>0\mu > 0μ>0, and ρ=λ/μ\rho = \lambda/\muρ=λ/μ. The number in the system is a continuous-time Markov chain on {0,1,2,… }\{0, 1, 2, \dots\}{0,1,2,…}, and its state probabilities pn(t)p_n(t)pn​(t) satisfy the forward (differential–difference) equations. For M/M/1 started with N(0)=iN(0) = iN(0)=i they are, for t≥0t \ge 0t≥0,

pn′(t)=−(λ+μ)pn(t)+λpn−1(t)+μpn+1(t) (n>0),p0′(t)=−λp0(t)+μp1(t),(2.72)p_n'(t) = -(\lambda+\mu)p_n(t) + \lambda p_{n-1}(t) + \mu p_{n+1}(t)\ (n > 0), \qquad p_0'(t) = -\lambda p_0(t) + \mu p_1(t), \tag{2.72}pn′​(t)=−(λ+μ)pn​(t)+λpn−1​(t)+μpn+1​(t) (n>0),p0′​(t)=−λp0​(t)+μp1​(t),(2.72)

with pn(0)=1p_n(0) = 1pn​(0)=1 if n=in = in=i and 000 otherwise. The other systems are variants:

  • M/M/1/1, no waiting room: two states and equations (2.70).
  • M/M/∞, ample service: the death rate in state nnn is nμn\munμ, giving (2.76).
  • The busy-period system: (2.72) with 000 made absorbing (λ0=0\lambda_0 = 0λ0​=0) and N(0)=1N(0) = 1N(0)=1. Its p0(t)p_0(t)p0​(t) is the distribution function of the busy period TbpT_{bp}Tbp​.

A family (pn)(p_n)(pn​) solves a system on [0,∞)[0,\infty)[0,∞) when each pnp_npn​ has, at every t≥0t \ge 0t≥0, the prescribed derivative (a right derivative at t=0t = 0t=0). It is a probability solution when pn(t)≥0p_n(t) \ge 0pn​(t)≥0 and ∑npn(t)=1\sum_n p_n(t) = 1∑n​pn​(t)=1 for every t≥0t \ge 0t≥0. The modified Bessel function of the first kind is

In(y)=∑k=0∞(y/2)n+2kk! (n+k)!,I−n=In,I_n(y) = \sum_{k=0}^{\infty} \frac{(y/2)^{n+2k}}{k!\,(n+k)!}, \qquad I_{-n} = I_n,In​(y)=k=0∑∞​k!(n+k)!(y/2)n+2k​,I−n​=In​,

and the Laplace transform of fff is fˉ(s)=∫0∞e−stf(t) dt\bar f(s) = \int_0^\infty e^{-st} f(t)\,dtfˉ​(s)=∫0∞​e−stf(t)dt for Re⁡s>0\operatorname{Re} s > 0Res>0.

Formalization targets

Goal: the transient M/M/1 law, (2.75)

With y=2tλμy = 2t\sqrt{\lambda\mu}y=2tλμ​,

pn(t)=e−(λ+μ)t[ρ(n−i)/2In−i(y)+ρ(n−i−1)/2In+i+1(y)+(1−ρ)ρn∑j=n+i+2∞ρ−j/2Ij(y)].p_n(t) = e^{-(\lambda+\mu)t}\Big[\rho^{(n-i)/2} I_{n-i}(y) + \rho^{(n-i-1)/2} I_{n+i+1}(y) + (1-\rho)\rho^n \sum_{j=n+i+2}^{\infty} \rho^{-j/2} I_j(y)\Big].pn​(t)=e−(λ+μ)t[ρ(n−i)/2In−i​(y)+ρ(n−i−1)/2In+i+1​(y)+(1−ρ)ρnj=n+i+2∑∞​ρ−j/2Ij​(y)].

The goal asserts five things for every λ,μ>0\lambda, \mu > 0λ,μ>0 and every iii, with no restriction on ρ\rhoρ:

  1. the series converges;
  2. these functions solve (2.72);
  3. they meet the initial condition;
  4. they form a probability distribution at every ttt;
  5. they are the only probability solution.

Milestones

  1. (2.71): the M/M/1/1 solution p1(t)=λλ+μ(1−e−(λ+μ)t)+p1(0)e−(λ+μ)tp_1(t) = \frac{\lambda}{\lambda+\mu}(1-e^{-(\lambda+\mu)t}) + p_1(0)e^{-(\lambda+\mu)t}p1​(t)=λ+μλ​(1−e−(λ+μ)t)+p1​(0)e−(λ+μ)t, and the matching formula for p0p_0p0​.
  2. (2.74) and Rouché's theorem: for Re⁡s>0\operatorname{Re} s > 0Res>0, the quadratic (λ+μ+s)z−μ−λz2(\lambda+\mu+s)z - \mu - \lambda z^2(λ+μ+s)z−μ−λz2 has exactly one zero in ∣z∣<1|z| < 1∣z∣<1, namely z1=(λ+μ+s−(λ+μ+s)2−4λμ)/(2λ)z_1 = (\lambda+\mu+s-\sqrt{(\lambda+\mu+s)^2-4\lambda\mu})/(2\lambda)z1​=(λ+μ+s−(λ+μ+s)2−4λμ​)/(2λ).
  3. The transform of p0p_0p0​: pˉ0(s)=z1i+1/(μ(1−z1))\bar p_0(s) = z_1^{i+1}/(\mu(1-z_1))pˉ​0​(s)=z1i+1​/(μ(1−z1​)).
  4. The limit of (2.75): pn(t)→(1−ρ)ρnp_n(t) \to (1-\rho)\rho^npn​(t)→(1−ρ)ρn if ρ<1\rho < 1ρ<1, and pn(t)→0p_n(t) \to 0pn​(t)→0 if ρ≥1\rho \ge 1ρ≥1.
  5. (2.77), M/M/∞: started empty, pn(t)=a(t)ne−a(t)/n!p_n(t) = a(t)^n e^{-a(t)}/n!pn​(t)=a(t)ne−a(t)/n! with a(t)=(1−e−μt)λ/μa(t) = (1-e^{-\mu t})\lambda/\mua(t)=(1−e−μt)λ/μ. The statement says that this family solves (2.76), is the unique probability solution, and has generating function exp⁡((z−1)a(t))\exp((z-1)a(t))exp((z−1)a(t)).
  6. The busy-period transform: pˉ0(s)=2μ/(s[λ+μ+s+(λ+μ+s)2−4λμ])\bar p_0(s) = 2\mu/(s[\lambda+\mu+s+\sqrt{(\lambda+\mu+s)^2-4\lambda\mu}])pˉ​0​(s)=2μ/(s[λ+μ+s+(λ+μ+s)2−4λμ​]).
  7. The busy-period density: p0′(t)=μ/λ e−(λ+μ)tI1(2λμ t)/tp_0'(t) = \sqrt{\mu/\lambda}\,e^{-(\lambda+\mu)t} I_1(2\sqrt{\lambda\mu}\,t)/tp0′​(t)=μ/λ​e−(λ+μ)tI1​(2λμ​t)/t.
  8. (2.79): for λ<μ\lambda < \muλ<μ, E[Tbp]=1/(μ−λ)E[T_{bp}] = 1/(\mu-\lambda)E[Tbp​]=1/(μ−λ) and E[Tbc]=1/λ+1/(μ−λ)E[T_{bc}] = 1/\lambda + 1/(\mu-\lambda)E[Tbc​]=1/λ+1/(μ−λ).

Significance

The formula (2.75) is the exact finite-time law of the most basic queue. It gives the rate at which M/M/1 approaches equilibrium, and it gives the distribution of the queue under overload (ρ≥1\rho \ge 1ρ≥1), where no steady state exists. It is the reference against which numerical transient methods, such as the uniformization of Chapter 8 of the same book, are checked. The busy-period density and its mean (2.79) enter server-utilisation and vacation models, and the Laplace-transform method used here recurs in the M/G/1 analysis of Chapter 5.

All of these results are classical and proved in the literature. None of them is machine-checked, as far as the platform's catalogue and Mathlib show. The chain from a countable system of linear ODEs, through generating functions and a root-location argument, to a Bessel series is a standard pattern in applied probability, and a formal version of it is what this mission adds. The formal statements also make explicit what the book leaves implicit: the sense in which the equations hold at t=0t = 0t=0, and the class in which the solution is unique.

Difficulty

The forward equations (2.72) form an infinite linear system. The obvious approach is to treat it like a finite system of ODEs, whose solution is a matrix exponential, and read off (2.75). That fails for two reasons. The generator is an infinite matrix, so its exponential needs a functional-analytic setting. And uniqueness is not automatic for infinite systems: it needs a class, such as probability solutions, and an argument that works in that class.

The Bessel form is a second, independent difficulty. The transform pˉ0(s)\bar p_0(s)pˉ​0​(s) is fixed by a root-location argument in the complex plane. Inverting the transform, or verifying (2.75) directly, requires manipulating the three-term Bessel recurrence and exchanging infinite sums. The tail sum ∑jρ−j/2Ij\sum_{j} \rho^{-j/2} I_j∑j​ρ−j/2Ij​ has to be controlled uniformly enough to be differentiated term by term. For ρ≥1\rho \ge 1ρ≥1 the factor (1−ρ)(1-\rho)(1−ρ) is non-positive, so the nonnegativity of pn(t)p_n(t)pn​(t) is not visible from the formula.

Formalization scope

Conventions committed to:

  • Parameters. Rates are real with λ,μ>0\lambda, \mu > 0λ,μ>0, and ρ=λ/μ\rho = \lambda/\muρ=λ/μ. States are ℕ (Fin 2 for M/M/1/1).
  • Solutions. "Solves on [0,∞)[0,\infty)[0,∞)" is HasDerivWithinAt on Set.Ici 0 at every t≥0t \ge 0t≥0. Uniqueness is asserted among solutions that are probability distributions at every time.
  • Special functions. Half-integer powers of ρ\rhoρ are real powers, and I−m=ImI_{-m} = I_mI−m​=Im​ is part of the definition. Laplace transforms are complex Bochner integrals over (0,∞)(0,\infty)(0,∞), and each statement also asserts the integrability it needs. Square roots with positive real part are hypotheses r2=(λ+μ+s)2−4λμr^2 = (\lambda+\mu+s)^2 - 4\lambda\mur2=(λ+μ+s)2−4λμ, Re⁡r>0\operatorname{Re} r > 0Rer>0.

The closed forms stated exactly as in the book are:

  • (2.71);
  • z1z_1z1​ and z2z_2z2​ of (2.74);
  • pˉ0(s)=z1i+1/(μ(1−z1))\bar p_0(s) = z_1^{i+1}/(\mu(1-z_1))pˉ​0​(s)=z1i+1​/(μ(1−z1​));
  • (2.75), with the Bessel series of p.101;
  • the M/M/∞ law and (2.77);
  • the busy-period transform and density of p.102;
  • (2.79).

The book derives (2.79) by a steady-state ratio argument valid for M/G/1. Here it is stated for M/M/1, as the mean of the explicit density.

A statement of (2.75) that only asserts the right-hand side is well defined, or checks only n=0n = 0n=0, is ruled out: the goal requires the ODE system, the initial condition, the probability property and uniqueness. For the same reason, the M/M/∞ law is tied to the system (2.76) and does not reduce to a Taylor expansion.

Needed infrastructure that Mathlib lacks:

  • modified Bessel functions of integer order;
  • Laplace transforms;
  • a Rouché-type zero count or a direct root-location lemma;
  • uniqueness for countable linear ODE systems with bounded or linearly growing rates.

The Bessel and Laplace definitions, and the uniqueness lemma for birth–death forward equations, are reusable beyond this mission. Contributions of those as separate lemmas are welcome.

Selected references

  • D. Gross, J. F. Shortle, J. M. Thompson, C. M. Harris, Fundamentals of Queueing Theory, 4th ed., Wiley, 2008, §§2.11–2.12, pp.97–103. https://doi.org/10.1002/9781118625651
  • N. T. J. Bailey, "A continuous time treatment of a simple queue using generating functions", J. Royal Statistical Society B 16 (1954) 288–291. https://doi.org/10.1111/j.2517-6161.1954.tb00172.x
  • W. Ledermann, G. E. H. Reuter, "Spectral theory for the differential equations of simple birth and death processes", Phil. Trans. Royal Society A 246 (1954) 321–369. https://doi.org/10.1098/rsta.1954.0001
  • D. G. Champernowne, "An elementary method of solution of the queueing problem with a single server and constant parameters", J. Royal Statistical Society B 18 (1956) 125–128. https://doi.org/10.1111/j.2517-6161.1956.tb00217.x
  • J. Abate, W. Whitt, "Calculating time-dependent performance measures for the M/M/1 queue", IEEE Trans. Communications 37 (1989) 1102–1104. https://doi.org/10.1109/26.41165
15 thms1 active userReviewed
Operations ResearchProbabilityStochastic Systems·Captain: mikedeng1

Fundamentals of Queueing Theory I: Foster's Criterion for Positive RecurrenceTextbook

Motivation

Almost every model in queueing theory is analysed through a Markov chain. The number of customers in an M/M/c queue is a continuous-time birth–death chain; the number left behind by departing customers of an M/G/1 queue is a discrete-parameter chain on {0,1,2,… }\{0,1,2,\dots\}{0,1,2,…} (the imbedded Markov chain); networks of queues are chains on vectors of queue lengths. Before any steady-state formula (Erlang's formulas, the Pollaczek–Khintchine formula, product forms) can be used, one has to know that the chain has a steady state at all: that it is positive recurrent, so that a stationary distribution exists and equals the limiting distribution.

Chapter 1 of Gross, Shortle, Thompson and Harris, Fundamentals of Queueing Theory (4th ed., Wiley 2008, DOI 10.1002/9781118625651), collects the two ingredients the rest of the book stands on: the Poisson process with its exponential interarrival times (§§1.7–1.8), and the classification theory of discrete-parameter Markov chains (§1.9), ending with Foster's criterion (Theorem 1.2), a sufficient condition for positive recurrence in terms of a drift inequality. The criterion goes back to F. G. Foster, On the stochastic matrices associated with certain queuing processes, Ann. Math. Statist. 24 (1953) (DOI 10.1214/aoms/1177728976), and is the ancestor of the Foster–Lyapunov method used for stability of queueing networks and stochastic systems.

This mission is the first of a series formalizing the book chapter by chapter.

Setting

A homogeneous discrete-parameter Markov chain on {0,1,2,… }\{0,1,2,\dots\}{0,1,2,…} is given by a transition matrix P={pij}P=\{p_{ij}\}P={pij​} with pij≥0p_{ij}\ge0pij​≥0 and ∑jpij=1\sum_j p_{ij}=1∑j​pij​=1 for every iii. The mmm-step transition probabilities pij(m)p_{ij}^{(m)}pij(m)​ are the entries of PmP^mPm.

The first-passage probability fij(n)f_{ij}^{(n)}fij(n)​ is the probability that the chain started in iii enters jjj for the first time at step n≥1n\ge1n≥1; for i=ji=ji=j it is the probability of first return at step nnn. The return probability is fjj=∑n≥1fjj(n)f_{jj}=\sum_{n\ge1}f_{jj}^{(n)}fjj​=∑n≥1​fjj(n)​ and the mean recurrence time is mjj=∑n≥1nfjj(n)∈[0,∞]m_{jj}=\sum_{n\ge1}n f_{jj}^{(n)}\in[0,\infty]mjj​=∑n≥1​nfjj(n)​∈[0,∞]. A state is positive recurrent if fjj=1f_{jj}=1fjj​=1 and mjj<∞m_{jj}<\inftymjj​<∞; the chain is positive recurrent if every state is.

The chain is irreducible if for every pair of states (i,j)(i,j)(i,j) some pij(n)p_{ij}^{(n)}pij(n)​ is positive, and aperiodic if for every state kkk the greatest common divisor of {n≥1:pkk(n)>0}\{n\ge1:p_{kk}^{(n)}>0\}{n≥1:pkk(n)​>0} is 111. A stationary distribution is a probability vector π\piπ with π=πP\pi=\pi Pπ=πP, i.e. πj=∑iπipij\pi_j=\sum_i\pi_i p_{ij}πj​=∑i​πi​pij​ for every jjj.

For the Poisson part, T0,T1,…T_0,T_1,\dotsT0​,T1​,… are independent interarrival times, each exponentially distributed with rate λ>0\lambda>0λ>0; the arrival epochs are Sn=T0+⋯+Tn−1S_n=T_0+\dots+T_{n-1}Sn​=T0​+⋯+Tn−1​, and N(t)=#{n≥1:Sn≤t}N(t)=\#\{n\ge1:S_n\le t\}N(t)=#{n≥1:Sn​≤t} counts the arrivals in [0,t][0,t][0,t].

Formalization targets

Goal: Theorem 1.2 (Foster's criterion)

An irreducible, aperiodic chain is positive recurrent if there exist xj≥0x_j\ge0xj​≥0 with

∑j=0∞pijxj≤xi−1(i≠0),∑j=0∞p0jxj<∞.\sum_{j=0}^\infty p_{ij}x_j\le x_i-1\quad(i\ne0),\qquad\sum_{j=0}^\infty p_{0j}x_j<\infty .j=0∑∞​pij​xj​≤xi​−1(i=0),j=0∑∞​p0j​xj​<∞.

Milestones: the Markov chain theorems

  • Theorem 1.1(a). In an irreducible, positive recurrent chain, πj=1/mjj\pi_j=1/m_{jj}πj​=1/mjj​ is a stationary distribution, and it is the only one.
  • Theorem 1.1(c). If moreover the chain is aperiodic and all moments of π\piπ are finite, then lim⁡m→∞pij(m)=πj\lim_{m\to\infty}p_{ij}^{(m)}=\pi_jlimm→∞​pij(m)​=πj​ for all i,ji,ji,j.

Milestones: the Poisson process and the exponential distribution

  • Eqs. (1.11)–(1.14). The unique solution of p0′=−λp0p_0'=-\lambda p_0p0′​=−λp0​, pn′=−λpn+λpn−1p_n'=-\lambda p_n+\lambda p_{n-1}pn′​=−λpn​+λpn−1​ with p0(0)=1p_0(0)=1p0​(0)=1, pn(0)=0p_n(0)=0pn​(0)=0 is pn(t)=(λt)ne−λt/n!p_n(t)=(\lambda t)^n e^{-\lambda t}/n!pn​(t)=(λt)ne−λt/n!.
  • Eq. (1.15). With exponential interarrival times,
Pr⁡{N(t)≤n}=∫t∞λ(λx)nn!e−λxdx=∑i=0n(λt)ie−λti!.\Pr\{N(t)\le n\}=\int_t^\infty\frac{\lambda(\lambda x)^n}{n!}e^{-\lambda x}dx=\sum_{i=0}^n\frac{(\lambda t)^ie^{-\lambda t}}{i!}.Pr{N(t)≤n}=∫t∞​n!λ(λx)n​e−λxdx=i=0∑n​i!(λt)ie−λt​.
  • Eq. (1.16). Given N(L)=kN(L)=kN(L)=k, the arrival epochs have density k!/Lkk!/L^kk!/Lk on {0<t1<⋯<tk<L}\{0<t_1<\dots<t_k<L\}{0<t1​<⋯<tk​<L}.
  • Eq. (1.17) and its converse (p.21). The exponential law satisfies Pr⁡{T≤t1∣T≥t0}=Pr⁡{0≤T≤t1−t0}\Pr\{T\le t_1\mid T\ge t_0\}=\Pr\{0\le T\le t_1-t_0\}Pr{T≤t1​∣T≥t0​}=Pr{0≤T≤t1​−t0​}, and it is the only continuous distribution on [0,∞)[0,\infty)[0,∞) that does.
  • Nonhomogeneous Poisson law (p.22). With a continuous rate λ(t)\lambda(t)λ(t) the forward equations have the unique solution pn(t)=e−m(t)m(t)n/n!p_n(t)=e^{-m(t)}m(t)^n/n!pn​(t)=e−m(t)m(t)n/n!, m(t)=∫0tλ(s) dsm(t)=\int_0^t\lambda(s)\,dsm(t)=∫0t​λ(s)ds.

Significance

Foster's criterion reduces positive recurrence, a statement about return times, to exhibiting one test function xxx with negative drift outside a single state. In the book it is the tool that establishes the existence of steady state for imbedded chains of the M/G/1 and G/M/1 queues (Chapter 5); its generalizations are the standard stability proofs for queueing networks. Theorem 1.1 then supplies what positive recurrence buys: the stationary distribution exists, is unique, equals 1/mjj1/m_{jj}1/mjj​, and is the limit of the transition probabilities. The Poisson results justify the "Markovian" arrivals and services of Chapters 2–4.

All of these results are classical and proved in the literature; the book states Theorems 1.1 and 1.2 without proof. The Prove2Me platform already holds machine-checked versions of related Markov chain theorems in other missions (Levin–Peres–Wilmer's and Durrett's countable-chain convergence theorems), stated with different definitions and hypotheses. What this mission adds is a formal development in the book's own terms — first-passage probabilities fjj(n)f_{jj}^{(n)}fjj(n)​, mean recurrence times mjjm_{jj}mjj​, gcd periodicity — on which the later missions of the series (imbedded chains, birth–death processes) can build, together with a formal proof of Foster's criterion, which is not on the platform.

Difficulty

For Foster's criterion the natural first step, taking expectations of the drift inequality along the chain, only shows that the expected value of xxx decreases while the chain stays away from 000. Turning that into a bound on the expected return time to 000 requires an optional-stopping or telescoping argument over a random time, with the value xxx possibly unbounded, and a separate argument that positive recurrence of state 000 propagates to all states of an irreducible chain. The book's hypotheses include aperiodicity, which the argument does not use.

For Theorem 1.1, identifying the stationary distribution with 1/mjj1/m_{jj}1/mjj​ requires relating the matrix powers PnP^nPn to the first-passage probabilities (a renewal decomposition), and uniqueness over countably many states needs care with infinite sums. For the Poisson results, the difficulty is measure-theoretic: the distribution of the sum of n+1n+1n+1 exponential variables, and conditioning on the event {N(L)=k}\{N(L)=k\}{N(L)=k} for the order-statistics property.

Formalization scope

States are natural numbers; the transition matrix is a real function p:N×N→Rp:\mathbb N\times\mathbb N\to\mathbb Rp:N×N→R with nonnegative entries and rows summing to one (as a convergent series). The return probability and the mean recurrence time are valued in [0,∞][0,\infty][0,∞], so null recurrence (mjj=∞m_{jj}=\inftymjj​=∞) is representable. Irreducibility is the per-pair notion. Stationary equations are stated componentwise with convergent series.

In Foster's criterion the series ∑jpijxj\sum_j p_{ij}x_j∑j​pij​xj​ are required to converge for every iii, which is the book's condition ∑jp0jxj<∞\sum_j p_{0j}x_j<\infty∑j​p0j​xj​<∞ together with the finiteness implicit in the inequalities for i≠0i\ne0i=0; xxx is real-valued and nonnegative. Dropping the convergence requirement would let a divergent row series (whose Lean sum is 000) satisfy the inequality vacuously; allowing xj=∞x_j=\inftyxj​=∞ would make the hypothesis trivially satisfiable. Neither is permitted.

The closed forms stated explicitly are: πj=1/mjj\pi_j=1/m_{jj}πj​=1/mjj​ (Theorem 1.1(a), with both existence and uniqueness), the Poisson probabilities (λt)ne−λt/n!(\lambda t)^ne^{-\lambda t}/n!(λt)ne−λt/n! (1.14), the Erlang tail integral and the Poisson CDF (1.15), the density k!/Lkk!/L^kk!/Lk (1.16), and e−m(t)m(t)n/n!e^{-m(t)}m(t)^n/n!e−m(t)m(t)n/n! for the nonhomogeneous law. Equations (1.14) and the nonhomogeneous law are stated as "solves the equations with the initial conditions if and only if equals the closed form", so both existence and uniqueness are asserted.

The Poisson results use random variables on a probability space, with Mathlib's expMeasure for the exponential law and cond for conditional probability. The derivation of the forward equations from the o(Δt)o(\Delta t)o(Δt) axioms of §1.7 is not formalized; the Poisson law is reached from the equations and, separately, from exponential interarrival times.

Not formalized: Theorem 1.1(b) and the word "ergodic" in 1.1(c), which rest on the book's informal notion of ergodicity; Theorem 1.3, whose phrase "for Theorem 1.1 to be valid" for a continuous-time chain is not pinned down.

The Markov chain definitions are reusable by every later mission that studies an imbedded chain. Contributions welcome: proofs of the milestones, and supporting lemmas (Chapman–Kolmogorov, renewal decomposition of pjj(n)p_{jj}^{(n)}pjj(n)​, class properties of recurrence).

Selected references

  • D. Gross, J. F. Shortle, J. M. Thompson, C. M. Harris, Fundamentals of Queueing Theory, 4th ed., Wiley, 2008. https://doi.org/10.1002/9781118625651
  • F. G. Foster, On the stochastic matrices associated with certain queuing processes, Annals of Mathematical Statistics 24 (1953), 355–360. https://doi.org/10.1214/aoms/1177728976
11 thms1 active userReviewed
Control TheoryDynamic ProgrammingLinear algebra+2·Captain: mikedeng1

Bellman's Dynamic Programming IX: Markovian Decision Processes and the Maximal Perron RootTextbook

Motivation

Chapter XI of Richard Bellman's Dynamic Programming (Princeton University Press, 1957; DOI 10.2307/j.ctv1nxcw0f) studies decision processes whose state is a vector of nonnegative quantities, for example the probabilities that a system is in each of NNN states, or the stocks of NNN commodities, and whose transitions are linear maps chosen stage by stage by a controller. Maximizing a linear functional of the state at every stage leads to the nonlinear difference equation

xi(n+1)=max⁡q∑j=1Naij(q) xj(n),xi(0)=ci,x_i(n+1) = \max_q \sum_{j=1}^N a_{ij}(q)\, x_j(n), \qquad x_i(0) = c_i,xi​(n+1)=qmax​j=1∑N​aij​(q)xj​(n),xi​(0)=ci​,

and, in the limit of small time steps, to differential equations of the form dx/dt=max⁡q[A(q,t)x+b(q,t)]dx/dt = \max_q [A(q,t)x + b(q,t)]dx/dt=maxq​[A(q,t)x+b(q,t)] and, when two opposing controllers act, dx/dt=max⁡pmin⁡q[… ]dx/dt = \max_p \min_q[\dots]dx/dt=maxp​minq​[…].

These equations are the multiplicative counterpart of the additive Bellman equation. Their growth rate is the natural object for controlled population models, controlled Markov chains observed through their unnormalized state vectors, and economic growth models with a choice of technology. Bellman announced the discrete results in "A Markovian decision process" (J. Math. Mech. 6, 1957) the same year as the book, and R. A. Howard's Dynamic Programming and Markov Processes (MIT Press, 1960) developed policy iteration for the related average-reward problem. The central discrete result of the chapter, Theorem 2, is an early instance of what is now called nonlinear Perron–Frobenius theory (Lemmens and Nussbaum, 2012).

Setting

Fix N≥1N \ge 1N≥1. Row iii of the matrix carries its own control qiq_iqi​, ranging over a set SiS_iSi​; the joint control is q=(q1,…,qN)q = (q_1, \dots, q_N)q=(q1​,…,qN​) in S=S1×⋯×SNS = S_1 \times \dots \times S_NS=S1​×⋯×SN​, and A(q)=(aij(qi))A(q) = (a_{ij}(q_i))A(q)=(aij​(qi​)). Bellman insists on this row-wise structure (§ 3): "the set of q's for each row is distinct from the corresponding set for any other row ... so that there is no interaction between the various maximizations". The maximum of a vector over qqq is then taken row by row.

The Perron root φ(q)\varphi(q)φ(q) is the characteristic root of A(q)A(q)A(q) of largest absolute value, the spectral radius of A(q)A(q)A(q) as a complex matrix. The conditions (10.3) of the chapter are:

  1. for every yyy and every row the maximum of ∑jaij(qi)yj\sum_j a_{ij}(q_i) y_j∑j​aij​(qi​)yj​ over SiS_iSi​ is attained;
  2. 0<aij(q)≤m<∞0 < a_{ij}(q) \le m < \infty0<aij​(q)≤m<∞ on SSS;
  3. φ\varphiφ attains its maximum on SSS.

For the continuous processes, ∥x∥=∑i∣xi∣\|x\| = \sum_i |x_i|∥x∥=∑i​∣xi​∣ and ∥A∥=∑i,j∣aij∣\|A\| = \sum_{i,j}|a_{ij}|∥A∥=∑i,j​∣aij​∣, and a solution of dx/dt=F(t,x)dx/dt = F(t,x)dx/dt=F(t,x), x(0)=cx(0)=cx(0)=c, on [0,T][0,T][0,T] is a continuous xxx with x(t)=c+∫0tF(s,x(s)) dsx(t) = c + \int_0^t F(s, x(s))\,dsx(t)=c+∫0t​F(s,x(s))ds, which is the book's "satisfying the equation almost everywhere". The successive approximations are x0=cx_0 = cx0​=c, xn+1(t)=c+∫0tF(s,xn(s)) dsx_{n+1}(t) = c + \int_0^t F(s, x_n(s))\,dsxn+1​(t)=c+∫0t​F(s,xn​(s))ds.

Formalization targets

Goal: Chapter XI, Theorem 2

Under (10.3) there is exactly one λ>0\lambda > 0λ>0 for which

λyi=max⁡q∑j=1Naij(q) yj,i=1,…,N,\lambda y_i = \max_q \sum_{j=1}^N a_{ij}(q)\, y_j, \qquad i = 1,\dots,N,λyi​=qmax​j=1∑N​aij​(q)yj​,i=1,…,N,

has a solution with all yi>0y_i > 0yi​>0. That solution is unique up to a positive factor, and

λ=max⁡q∈Sφ(q).\lambda = \max_{q \in S} \varphi(q).λ=q∈Smax​φ(q).

Milestones

  1. § 4, Lemma. For row-wise maximized operators T1(x)=max⁡q[b1(q,t)+∫0tA(q,s)x ds]T_1(x) = \max_q[b_1(q,t) + \int_0^t A(q,s)x\,ds]T1​(x)=maxq​[b1​(q,t)+∫0t​A(q,s)xds] and T2(y)T_2(y)T2​(y) likewise, ∥T1(x)−T2(y)∥≤max⁡q[∥b1−b2∥+∫0t∥A(q,s)∥ ∥x−y∥ ds]\|T_1(x) - T_2(y)\| \le \max_q[\|b_1 - b_2\| + \int_0^t \|A(q,s)\|\,\|x-y\|\,ds]∥T1​(x)−T2​(y)∥≤maxq​[∥b1​−b2​∥+∫0t​∥A(q,s)∥∥x−y∥ds].
  2. Theorem 1. If ∥A(q,t)∥,∥b(q,t)∥≤f(t)\|A(q,t)\|, \|b(q,t)\| \le f(t)∥A(q,t)∥,∥b(q,t)∥≤f(t) with fff locally integrable and the maximum is attained, then dx/dt=max⁡q[A(q,t)x+b(q,t)]dx/dt = \max_q[A(q,t)x + b(q,t)]dx/dt=maxq​[A(q,t)x+b(q,t)], x(0)=cx(0) = cx(0)=c, has a unique solution, the uniform limit of the successive approximations.
  3. Theorem 3 (corrected). If moreover φ\varphiφ has a unique maximizer on SSS and c≥0c \ge 0c≥0, c≠0c \ne 0c=0, then the recurrence satisfies xi(n)∼a yi λnx_i(n) \sim a\,y_i\,\lambda^nxi​(n)∼ayi​λn with a=a(c)>0a = a(c) > 0a=a(c)>0.
  4. Theorem 4. The same well-posedness for dx/dt=max⁡pmin⁡q[A(p,q,t)x+b(p,q,t)]=min⁡qmax⁡p[… ]dx/dt = \max_p\min_q[A(p,q,t)x + b(p,q,t)] = \min_q\max_p[\dots]dx/dt=maxp​minq​[A(p,q,t)x+b(p,q,t)]=minq​maxp​[…] on [0,T][0,T][0,T].
  5. Theorem 5. If (Bp,q)≥d>0(Bp,q) \ge d > 0(Bp,q)≥d>0 on probability vectors, the solution of du/dt=max⁡pmin⁡q[(Ap,q)−(Bp,q)u]du/dt = \max_p\min_q[(Ap,q) - (Bp,q)u]du/dt=maxp​minq​[(Ap,q)−(Bp,q)u] satisfies
lim⁡t→∞u(t)=max⁡pmin⁡q(Ap,q)(Bp,q)=min⁡qmax⁡p(Ap,q)(Bp,q).\lim_{t\to\infty} u(t) = \max_p \min_q \frac{(Ap,q)}{(Bp,q)} = \min_q \max_p \frac{(Ap,q)}{(Bp,q)} .t→∞lim​u(t)=pmax​qmin​(Bp,q)(Ap,q)​=qmin​pmax​(Bp,q)(Ap,q)​.

Significance

Theorem 2 identifies the optimal long-run growth rate of a controlled multiplicative process with the largest Perron root among the admissible matrices, and shows that the optimal process has a single positive stationary direction. Theorem 3 turns this into the asymptotics of the value iteration x(n+1)=max⁡qA(q)x(n)x(n+1) = \max_q A(q)x(n)x(n+1)=maxq​A(q)x(n): after normalization by λn\lambda^nλn the iterates converge to a multiple of the eigenvector. Theorems 1 and 4 are the existence and uniqueness results that justify defining continuous-time controlled processes and differential games by these equations. Theorem 5 recovers the min-max theorem for ratios of bilinear forms (Chapter X) as the long-run limit of a scalar differential game.

The results are classical, and none of them is formalized. Mathlib has the spectral radius and irreducible matrices but no Perron–Frobenius theorem and no Brouwer fixed point theorem; the platform has a statement of the Perron theorem for a single positive matrix (ClassicalGaps.perron_positive_matrix). A formal proof of the goal therefore also produces a reusable monotone, positively homogeneous eigenvector theorem on the positive orthant.

Difficulty

The map y↦max⁡qA(q)yy \mapsto \max_q A(q)yy↦maxq​A(q)y is not linear, so the linear-algebra proof of the Perron theorem through the characteristic polynomial does not apply. Existence of a positive eigenvector needs a fixed point argument for a nonlinear map of the simplex (Bellman uses Brouwer's theorem). The identification λ=max⁡qφ(q)\lambda = \max_q \varphi(q)λ=maxq​φ(q) must connect the nonlinear eigenvalue with the spectra of the individual matrices, which requires the Perron theory of each A(q)A(q)A(q), including the fact that the Perron root dominates every complex eigenvalue in modulus. For Theorem 3, the iterates may switch controls infinitely often when SSS is infinite, so an argument that the optimal control is eventually constant does not settle convergence. For Theorems 1 and 4, the right-hand side is only measurable in ttt and Lipschitz in xxx with an integrable constant, so the classical Picard–Lindelöf theorem with a continuous right-hand side does not apply directly.

Formalization scope

Everything lives in the namespace BellmanDP.Markovian. Vectors are Fin N → ℝ and matrices are Matrix (Fin N) (Fin N) ℝ. Row iii's control type is Q i with admissible set S i, and the joint admissible set is Set.pi Set.univ S. The Perron root is (spectralRadius ℂ (A.map (algebraMap ℝ ℂ))).toReal, the largest modulus of a complex eigenvalue; it is not defined as a positive eigenvalue with a positive eigenvector, which would make the Perron–Frobenius content of the goal definitional. The maximized eigen-equation is stated with IsGreatest, so the maxima are attained. The goal and Theorem 3 assume N≥1N \ge 1N≥1; for N=0N = 0N=0 every λ\lambdaλ would qualify.

Conventions and repairs:

  • Theorem 3 prints "a unique q for which the maximum value of q is assumed". A control has no maximum value; the proof uses "q∗q^*q∗ ... the value of qqq for which λ=φ(q∗)\lambda = \varphi(q^*)λ=φ(q∗)", so the hypothesis is uniqueness of the maximizer of φ\varphiφ. For c=0c = 0c=0 the iterates vanish and xi(n)∼ayiλnx_i(n) \sim a y_i\lambda^nxi​(n)∼ayi​λn fails, so c≠0c \ne 0c=0 is assumed (the proof takes c>0c > 0c>0 "without loss of generality"). The asymptotic is stated as xi(n)/λn→ayix_i(n)/\lambda^n \to a y_ixi​(n)/λn→ayi​ with a>0a > 0a>0.
  • Theorems 1 and 4: the book's controls are functions of ttt with the maximum outside the integral; since the maximization is pointwise (§ 4), the statements use pointwise sets and the integral of the pointwise maximum. Measurability of t↦F(t,x)t \mapsto F(t,x)t↦F(t,x) is not stated in the book and is assumed. In Theorem 4 the max-min is taken row by row, and (2a) is encoded as the existence of a saddle point in each row.
  • § 4 Lemma: "≤max⁡q[… ]\le \max_q[\dots]≤maxq​[…]" is stated as "≤[… ]\le [\dots]≤[…] at some admissible joint qqq".
  • Theorem 5: the right-hand side is the max-min form; the equality of the two ratio values is part of the conclusion.

Degenerate readings are ruled out: the maxima are attained or taken over nonempty compact sets, never Lean's junk sSup of an unbounded set, and the Perron root is spectral rather than defined through the conclusion. Contributions welcome: a proof of the single-matrix Perron theorem in the form needed here, a Brouwer or Kakutani fixed point theorem for the simplex, and a Carathéodory existence theorem for dx/dt=F(t,x)dx/dt = F(t,x)dx/dt=F(t,x) with an integrable Lipschitz constant, each reusable well beyond this mission.

Selected references

  • R. Bellman, Dynamic Programming, Princeton University Press, 1957; Princeton Landmarks in Mathematics ed., 2010, Chapter XI. https://doi.org/10.2307/j.ctv1nxcw0f
  • R. Bellman, "A Markovian decision process", Journal of Mathematics and Mechanics 6 (1957), 679–684.
  • R. A. Howard, Dynamic Programming and Markov Processes, MIT Press, 1960.
  • O. Perron, "Zur Theorie der Matrices", Mathematische Annalen 64 (1907), 248–263. https://doi.org/10.1007/BF01449896
  • B. Lemmens and R. Nussbaum, Nonlinear Perron–Frobenius Theory, Cambridge University Press, 2012. https://doi.org/10.1017/CBO9781139026079
9 thms1 active userReviewed
Dynamic ProgrammingOperations ResearchProbability+1·Captain: mikedeng1

Stochastic Dynamic Programming and the Control of Queueing Systems VII: The (BOR) Assumptions and Positive Recurrence of Optimal PoliciesTextbook

Motivation

Queueing control problems (admission control, routing, service rate selection) are naturally modelled as Markov decision chains with a countably infinite state space and unbounded costs, for instance a holding cost that grows with the queue length. For such models the long-run average cost criterion is often the relevant one, and the central question is whether an optimal stationary policy exists and can be computed from an average cost optimality equation (ACOE). Chapter 7 of Linn I. Sennott, Stochastic Dynamic Programming and the Control of Queueing Systems (Wiley, 1999, doi:10.1002/9780470317037) develops a verifiable set of conditions, the (SEN) assumptions, under which an average cost optimality inequality (ACOI) holds and yields an optimal stationary policy. The inequality may be strict (Example 7.3.1), and an optimal policy may induce a Markov chain without positive recurrent states.

Sections 7.4 and 7.5 answer two practical questions: when is the ACOI in fact an equation, and how can (SEN) be checked in a concrete model? The answer culminates in the (BOR) assumptions, which require only one well-behaved stationary policy and the finiteness of a set of low-cost states.

According to the book's bibliographic notes (p. 163): the (BOR) assumptions modify a line of development due to Borkar (SIAM J. Control Optim. 22, 1984, and 27, 1989; monograph 1991) and are weaker than his original conditions; the proof that (BOR) implies (SEN) is from Cavazos-Cadena and Sennott (Oper. Res. Letters 11, 1992), and the version of (BOR) used here is from Sennott (Prob. Eng. Inform. Sci. 7, 1993). Proposition 7.5.5 and the (CAV*) assumptions go back to Cavazos-Cadena (Kybernetika 25, 1989); Proposition 7.5.3 and Corollary 7.5.4 to Sennott (Oper. Res. 37, 1989).

Setting

A Markov decision chain consists of a countable state space SSS, finite nonempty action sets AiA_iAi​, nonnegative finite costs C(i,a)C(i,a)C(i,a) and transition probabilities Pij(a)P_{ij}(a)Pij​(a). A policy θ\thetaθ may use the whole history and randomize. For α∈(0,1)\alpha\in(0,1)α∈(0,1) the discount value function is Vα(i)=inf⁡θVθ,α(i)V_\alpha(i)=\inf_\theta V_{\theta,\alpha}(i)Vα​(i)=infθ​Vθ,α​(i), the infimum of ∑tαtEθ[C(Xt,At)∣X0=i]\sum_t\alpha^tE_\theta[C(X_t,A_t)\mid X_0=i]∑t​αtEθ​[C(Xt​,At​)∣X0​=i]; the average cost of θ\thetaθ is Jθ(i)=lim sup⁡n1nEθ[∑t<nC(Xt,At)∣X0=i]J_\theta(i)=\limsup_n\frac1nE_\theta[\sum_{t<n}C(X_t,A_t)\mid X_0=i]Jθ​(i)=limsupn​n1​Eθ​[∑t<n​C(Xt​,At​)∣X0​=i] and the minimum average cost is J(i)=inf⁡θJθ(i)J(i)=\inf_\theta J_\theta(i)J(i)=infθ​Jθ​(i). All of these lie in [0,∞][0,\infty][0,∞].

For a distinguished state zzz the relative value is hα(i)=Vα(i)−Vα(z)h_\alpha(i)=V_\alpha(i)-V_\alpha(z)hα​(i)=Vα​(i)−Vα​(z). The (SEN) assumptions are: (SEN1) (1−α)Vα(z)(1-\alpha)V_\alpha(z)(1−α)Vα​(z) is bounded on (0,1)(0,1)(0,1); (SEN2) hα≤Mh_\alpha\le Mhα​≤M for a finite function M≥0M\ge0M≥0; (SEN3) hα≥−Lh_\alpha\ge-Lhα​≥−L for a finite constant L≥0L\ge0L≥0. Under (SEN), J=lim⁡α→1−(1−α)Vα(i)J=\lim_{\alpha\to1^-}(1-\alpha)V_\alpha(i)J=limα→1−​(1−α)Vα​(i) is a finite constant, and a limit function hhh is a pointwise limit of hβnh_{\beta_n}hβn​​ along some βn→1−\beta_n\to1^-βn​→1−. The ACOI and ACOE read

J+h(i) ≥ (resp. =) min⁡a∈Ai{C(i,a)+∑jPij(a)h(j)},i∈S.J+h(i)\ \ge\ (\text{resp. }=)\ \min_{a\in A_i}\Big\{C(i,a)+\sum_jP_{ij}(a)h(j)\Big\},\qquad i\in S.J+h(i) ≥ (resp. =) a∈Ai​min​{C(i,a)+j∑​Pij​(a)h(j)},i∈S.

For a nonempty set GGG the first passage time is T=min⁡{n≥1:Xn∈G}T=\min\{n\ge1:X_n\in G\}T=min{n≥1:Xn​∈G}. The class ℜ(i,G)\Re(i,G)ℜ(i,G) consists of the policies that, from iii, enter GGG with probability one in finite expected time miG(θ)m_{iG}(\theta)miG​(θ); ℜ∗(i,G)\Re^*(i,G)ℜ∗(i,G) adds a finite expected first passage cost ciG(θ)=Eθ[∑t<TC(Xt,At)]c_{iG}(\theta)=E_\theta[\sum_{t<T}C(X_t,A_t)]ciG​(θ)=Eθ​[∑t<T​C(Xt​,At​)]. A (randomized) stationary policy ddd is zzz standard if the Markov chain it induces has miz<∞m_{iz}<\inftymiz​<∞ and ciz<∞c_{iz}<\inftyciz​<∞ for every iii; it then has a single positive recurrent class Rd∋zR_d\ni zRd​∋z and a finite constant average cost JdJ_dJd​.

Formalization targets

Goal: Theorem 7.5.6

Assume (BOR): (BOR1) a zzz standard policy ddd exists; (BOR2) for some ε>0\varepsilon>0ε>0 the set D={i:C(i,a)≤Jd+ε for some a}D=\{i: C(i,a)\le J_d+\varepsilon\text{ for some }a\}D={i:C(i,a)≤Jd​+ε for some a} is finite; (BOR3) every i∈D−Rdi\in D-R_di∈D−Rd​ can be reached from zzz by some θi∈ℜ∗(z,i)\theta_i\in\Re^*(z,i)θi​∈ℜ∗(z,i). Then (SEN) holds and every limit function satisfies the ACOE; every average cost optimal stationary policy eee has a positive recurrent state in

D(e)={i:C(i,e)≤J+ε},D(e)=\{i: C(i,e)\le J+\varepsilon\},D(e)={i:C(i,e)≤J+ε},

at most ∣D(e)∣|D(e)|∣D(e)∣ positive recurrent classes and no null recurrent class; and a policy realizing the minimum in the ACOE satisfies e∈ℜ∗(i,D(e)∩R(e))e\in\Re^*(i,D(e)\cap R(e))e∈ℜ∗(i,D(e)∩R(e)) for every iii.

Milestones

  • Lemma 7.4.1: hα(i)≤ciz(θi)h_\alpha(i)\le c_{iz}(\theta_i)hα​(i)≤ciz​(θi​) for θi∈ℜ∗(i,z)\theta_i\in\Re^*(i,z)θi​∈ℜ∗(i,z), hence (SEN2).
  • Lemma 7.4.2: h(i)≤ciG(θ)−JmiG(θ)+Eθ[h(XT)]h(i)\le c_{iG}(\theta)-Jm_{iG}(\theta)+E_\theta[h(X_T)]h(i)≤ciG​(θ)−JmiG​(θ)+Eθ​[h(XT​)] for θ∈ℜ(i,G)\theta\in\Re(i,G)θ∈ℜ(i,G) under an integrability condition.
  • Theorem 7.4.3: four sufficient conditions for equality in the ACOI at a state.
  • Lemma 7.5.2: Jd=(1−α)∑i∈Rπi(d)Vd,α(i)J_d=(1-\alpha)\sum_{i\in R}\pi_i(d)V_{d,\alpha}(i)Jd​=(1−α)∑i∈R​πi​(d)Vd,α​(i) for a zzz standard ddd.
  • Proposition 7.5.3: a zzz standard policy gives (SEN1–2).
  • Corollary 7.5.4: on S={0,1,… }S=\{0,1,\dots\}S={0,1,…}, increasing VαV_\alphaVα​ plus a 000 standard policy gives (SEN), with nonnegative increasing limit functions.
  • Proposition 7.5.5: an optimal stationary policy has a positive recurrent state of cost at most J+εJ+\varepsilonJ+ε, reachable from iii, when (7.33) holds.
  • Corollaries 7.5.9 and 7.5.10: the (CAV) and (CAV*) conditions imply (BOR).

Significance

Theorem 7.5.6 reduces the verification of the ACOE for a queueing model to three checks that do not involve the discount value function: exhibit one stationary policy with finite mean return times and costs to a fixed state (typically a stable "serve at maximal rate" policy), check that low costs occur on a finite set (automatic when the holding cost grows without bound, Corollaries 7.5.9–7.5.10), and check reachability of finitely many states. Its conclusions go beyond existence: optimal stationary policies induce chains with positive recurrent classes located in a known finite set, and ACOE-realizing policies reach them in finite expected time and cost. This is what makes value iteration and approximating-sequence methods in later chapters of the book applicable to these models.

The results are proved in the book. The present mission produces machine-checked statements of the first passage calculus for general (history-dependent, randomized) policies, of (SEN) and limit functions, and of the chain of implications from (CAV*) to the ACOE. No machine-checked version of these statements is known.

Difficulty

The obvious approach to the ACOE is to pass to the limit α→1−\alpha\to1^-α→1− in the discount optimality equation. Exchanging this limit with ∑jPij(a)hα(j)\sum_jP_{ij}(a)h_\alpha(j)∑j​Pij​(a)hα​(j) requires a dominating function, and (SEN2) only gives a pointwise bound MMM whose expectation may be infinite; Fatou's lemma then yields only the inequality. Obtaining equality requires tracking first passages to sets and showing that the discrepancy Φ\PhiΦ vanishes along them, which in turn needs finiteness of ciGc_{iG}ciG​ that is not assumed but has to be derived. On the recurrence side, the average cost criterion is a limit superior of Cesàro averages over a countable state space, and mass can escape to infinity; the finiteness of the set DDD is what prevents an optimal policy from spending its time in transient or null recurrent states, and turning that into positive recurrence requires the renewal-type identities of Appendix C.

Formalization scope

States form a countable type SSS; action sets are nonempty Finsets; costs are in ℝ≥0; transition probabilities are ℝ≥0∞-valued with row sums one on admissible actions. Policies are general: a history is a state sequence and an action sequence, and all probabilities and expectations (hitting probabilities, miGm_{iG}miG​, ciGc_{iG}ciG​, Pθ(XT=j)P_\theta(X_T=j)Pθ​(XT​=j), Qij(n)Q^{(n)}_{ij}Qij(n)​) are computed from the history probabilities of the process. VαV_\alphaVα​, JθJ_\thetaJθ​, miGm_{iG}miG​ and ciGc_{iG}ciG​ take values in [0,∞][0,\infty][0,∞]; miG=∞m_{iG}=\inftymiG​=∞ when GGG is missed with positive probability; the first passage time satisfies T≥1T\ge1T≥1. hαh_\alphahα​ and ∑jPij(a)h(j)\sum_jP_{ij}(a)h(j)∑j​Pij​(a)h(j) are in the extended reals, with the book's convention that a function bounded below has an expectation in (−∞,+∞](-\infty,+\infty](−∞,+∞]. Limit functions are real valued. Positive recurrence, communicating classes and steady state probabilities πj=(mjj)−1\pi_j=(m_{jj})^{-1}πj​=(mjj​)−1 are the notions for the chain induced by a (randomized) stationary policy. JdJ_dJd​ is the average cost of ddd from zzz.

A formalization in which the ACOE is asserted for some convenient function instead of every limit function, or in which ∣D(e)∣|D(e)|∣D(e)∣ is a natural-number cardinality that vanishes on infinite sets, would trivialize part of the goal; the statements quantify over all limit functions and use Set.encard.

A complete development needs: history-dependent policies and their path laws on countable spaces; first passage decompositions (strong Markov property at TTT); Abelian limits of ∑tαtP(T=t)\sum_t\alpha^tP(T=t)∑t​αtP(T=t); Fatou and dominated convergence for series; and the renewal reward theorem for positive recurrent classes (Appendix C of the book). The first passage and Markov chain layer is reusable beyond this mission. Proofs of individual milestones, and sharper statements of the Appendix C facts they use, are welcome.

Selected references

  • L. I. Sennott, Stochastic Dynamic Programming and the Control of Queueing Systems, Wiley Series in Probability and Statistics, John Wiley & Sons, 1999. doi:10.1002/9780470317037
  • V. S. Borkar, "On minimum cost per unit time control of Markov chains", SIAM J. Control Optim. 22 (1984), 965–978.
  • V. S. Borkar, "Control of Markov chains with long-run average cost criterion: the dynamic programming equations", SIAM J. Control Optim. 27 (1989), 642–657.
  • V. S. Borkar, Topics in Controlled Markov Chains, Pitman Research Notes in Mathematics 240, Longman, 1991.
  • R. Cavazos-Cadena, "Weak conditions for the existence of optimal stationary policies in average Markov decision chains with unbounded costs", Kybernetika 25 (1989), 145–156.
  • R. Cavazos-Cadena and L. I. Sennott, "Comparing recent assumptions for the existence of average optimal stationary policies", Oper. Res. Letters 11 (1992), 33–37.
  • L. I. Sennott, "The average cost optimality equation and critical number policies", Prob. Eng. Inform. Sci. 7 (1993).
  • L. I. Sennott, "Average cost optimal stationary policies in infinite state Markov decision processes with unbounded costs", Operations Research 37 (1989), 626–633. doi:10.1287/opre.37.4.626
  • K. L. Chung, Markov Chains with Stationary Transition Probabilities, 2nd ed., Springer, 1967.
15 thms1 active userReviewed
Dynamic ProgrammingOperations ResearchProbability·Captain: mikedeng1

Stochastic Dynamic Programming and the Control of Queueing Systems V: The Average Cost Optimality Equation and Value Iteration for Finite State SpacesTextbook

Motivation

Average cost Markov decision chains model systems that run indefinitely and are judged by their long-run cost per step: admission and routing control in queues, inventory replenishment, machine maintenance. For a finite state space the classical tool is the average cost optimality equation (ACOE)

J+h(i)=min⁡a∈Ai{C(i,a)+∑jPij(a) h(j)},J + h(i) = \min_{a \in A_i}\Big\{C(i,a) + \sum_j P_{ij}(a)\,h(j)\Big\},J+h(i)=a∈Ai​min​{C(i,a)+j∑​Pij​(a)h(j)},

whose solution gives both the minimum average cost JJJ and an optimal stationary policy. To be useful the equation has to be solved numerically, and the method used in practice is value iteration: compute the minimum nnn-horizon costs vnv_nvn​ and extract JJJ and hhh from their growth. This mission formalizes Sections 6.4–6.6 of L. I. Sennott, Stochastic Dynamic Programming and the Control of Queueing Systems (Wiley, 1999, doi:10.1002/9780470317037): when the minimum average cost is constant, the ACOE holds, any solution of it is optimal, and value iteration converges, provided the optimal policies are aperiodic. When they are not, a transformation of the model makes them so.

Related classical work includes Blackwell's discrete dynamic programming (1962) and Schweitzer–Federgruen's analysis of undiscounted value iteration (1977); Sennott's treatment derives the ACOE from the discounted value function VαV_\alphaVα​ as α→1−\alpha \to 1^-α→1−, which is the route that extends to countable state spaces in later chapters of the book.

Setting

A Markov decision chain (MDC) Δ\DeltaΔ has a finite state space SSS; in each state iii a finite nonempty action set AiA_iAi​; nonnegative costs C(i,a)C(i,a)C(i,a); and transition probabilities Pij(a)P_{ij}(a)Pij​(a). A policy θ\thetaθ may use the whole history and randomize; a stationary policy eee always chooses e(i)∈Aie(i) \in A_ie(i)∈Ai​ in state iii and induces a Markov chain with transitions Pij(e)=Pij(e(i))P_{ij}(e) = P_{ij}(e(i))Pij​(e)=Pij​(e(i)).

For a policy θ\thetaθ and initial state iii: vθ,n(i)v_{\theta,n}(i)vθ,n​(i) is the expected cost of the first nnn steps, Vθ,α(i)V_{\theta,\alpha}(i)Vθ,α​(i) the expected α\alphaα-discounted cost, and Jθ(i)=lim sup⁡nvθ,n(i)/nJ_\theta(i) = \limsup_n v_{\theta,n}(i)/nJθ​(i)=limsupn​vθ,n​(i)/n the average cost. The value functions are the infima over all policies: vnv_nvn​, VαV_\alphaVα​ and the minimum average cost J(i)J(i)J(i). A policy is average cost optimal if Jθ≡JJ_\theta \equiv JJθ​≡J.

Section 6.2 of the book provides a stationary policy fff that is α\alphaα discount optimal for all α\alphaα close to 111 (a Blackwell optimal policy), and Section 6.3 builds from it a relative value function w∗w^*w∗. For a distinguished state zzz put

hα(i)=Vα(i)−Vα(z),h(i)=lim⁡α→1−hα(i),dn(i)=h(i)+nJ−vn(i).h_\alpha(i) = V_\alpha(i) - V_\alpha(z), \qquad h(i) = \lim_{\alpha\to1^-} h_\alpha(i), \qquad d_n(i) = h(i) + nJ - v_n(i).hα​(i)=Vα​(i)−Vα​(z),h(i)=α→1−lim​hα​(i),dn​(i)=h(i)+nJ−vn​(i).

For a distinguished state xxx the finite horizon relative value function is rn(i)=vn(i)−vn(x)r_n(i) = v_n(i) - v_n(x)rn​(i)=vn​(i)−vn​(x).

A positive recurrent class RRR of a Markov chain is aperiodic if Pij(n)→πjP^{(n)}_{ij} \to \pi_jPij(n)​→πj​ for i,j∈Ri, j \in Ri,j∈R, where π\piπ is the steady state distribution. Assumption OPA ("optimal policies are aperiodic") requires every positive recurrent class of every average cost optimal stationary policy to be aperiodic. The aperiodicity transformation Δ∗\Delta^*Δ∗ with 0<τ<10<\tau<10<τ<1 keeps states and actions, scales costs by τ\tauτ, and sets Pij∗(a)=τPij(a)P^*_{ij}(a) = \tau P_{ij}(a)Pij∗​(a)=τPij​(a) for j≠ij \ne ij=i, Pii∗(a)=τPii(a)+(1−τ)P^*_{ii}(a) = \tau P_{ii}(a) + (1-\tau)Pii∗​(a)=τPii​(a)+(1−τ).

Formalization targets

Goal: convergence of value iteration (Proposition 6.6.3)

If J(i)≡JJ(i) \equiv JJ(i)≡J and Assumption OPA holds, then for any distinguished state xxx

lim⁡n→∞[vn(x)−vn−1(x)]=J,lim⁡n→∞rn(i)=:r(i) exists,\lim_{n\to\infty}[v_n(x) - v_{n-1}(x)] = J, \qquad \lim_{n\to\infty} r_n(i) =: r(i) \text{ exists},n→∞lim​[vn​(x)−vn−1​(x)]=J,n→∞lim​rn​(i)=:r(i) exists,

(J,r)(J, r)(J,r) solves the ACOE, and every limit point of the finite horizon optimal stationary policies is average cost optimal.

Milestones

  1. Proposition 6.4.1: unichain structure, bounded ∣Vα(i)−Vα(z)∣|V_\alpha(i) - V_\alpha(z)|∣Vα​(i)−Vα​(z)∣, or pairwise reachability imply J(i)≡JJ(i) \equiv JJ(i)≡J, with the implication diagram (6.26).
  2. Theorem 6.4.2: under J(i)≡JJ(i) \equiv JJ(i)≡J, hhh exists, solves the ACOE (6.31), yields optimal policies, ∣dn∣≤L|d_n| \le L∣dn​∣≤L and vn/n→Jv_n/n \to Jvn​/n→J.
  3. Proposition 6.5.1: any finite solution (F,r)(F, r)(F,r) of the ACOE (or of the inequality (6.36)) gives J≡FJ \equiv FJ≡F and optimal policies, and differs from hhh by constants on recurrent classes.
  4. Lemma 6.6.2: on an aperiodic positive recurrent class of an optimal policy, dnd_ndn​ converges to a constant.
  5. Lemma 6.6.5 and Proposition 6.6.6: Δ∗\Delta^*Δ∗ has the same recurrent classes and steady states, all of them aperiodic, costs scaled by τ\tauτ; value iteration on Δ∗\Delta^*Δ∗ produces a solution (J∗/τ,r∗)(J^*/\tau, r^*)(J∗/τ,r∗) of the ACOE of Δ\DeltaΔ.

Significance

The ACOE with constant JJJ is the standard certificate of optimality for finite average cost models, and Proposition 6.5.1 is what allows any numerical solution of it to be trusted. Proposition 6.6.3 is the correctness theorem of the value iteration algorithm (VIA 6.6.4 of the book), and Proposition 6.6.6 removes its one extra hypothesis at the price of a model transformation. Chapter 8 of the book runs this algorithm on a sequence of finite truncations to compute optimal policies for countable-state queueing models, so these results are the base of the book's computational method.

All results in this mission are proved in the book; none has a machine-checked proof. Existing formalizations on the platform treat average reward models under a unichain hypothesis with a single action set type; this mission assumes only a constant minimum average cost (multichain models allowed) and uses the general policy class throughout.

Difficulty

The ACOE itself is not the obstacle; convergence of vn(x)−vn−1(x)v_n(x) - v_{n-1}(x)vn​(x)−vn−1​(x) is. Theorem 6.4.2 bounds dnd_ndn​ but does not make it converge, and Example 6.6.1 of the book (a two-state periodic chain) shows that without aperiodicity vn(x)−vn−1(x)v_n(x) - v_{n-1}(x)vn​(x)−vn−1​(x) oscillates. The naive argument, passing to the limit in the finite horizon optimality equation, assumes the limits exist, which is exactly what is in question. Chain structure is the obstruction: a multichain optimal policy has several recurrent classes, and the Cesàro-type convergence that suffices for the ACOE itself is weaker than the pointwise convergence value iteration needs. The policy statement is also delicate, since the finite horizon minimizers fnf_nfn​ need not converge.

Formalization scope

  • The state type S is finite ([Fintype S]); actions are a type Act with per-state nonempty Finset action sets. Costs are in ℝ≥0, transition probabilities in ℝ≥0∞, and all value functions are defined in [0,∞] as infima over all history-dependent randomized policies, then converted to ℝ (they are finite for finite SSS).
  • JJJ constant is stated as J(i)=JJ(i) = JJ(i)=J for all iii, with J∈R≥0J \in \mathbb R_{\ge 0}J∈R≥0​. The relative value hhh is defined as the limit α→1−\alpha \to 1^-α→1− of hαh_\alphahα​, not taken as an arbitrary solution of the ACOE; Theorem 6.4.2(i) asserts the limit exists. The Blackwell optimal policy fff enters as a hypothesis: any stationary policy discount optimal on an interval (α0,1)(\alpha_0,1)(α0​,1).
  • min_a is Finset.inf' over AiA_iAi​. Limit points of policy sequences follow Definition B.1 (a subsequence agreeing eventually in every state). Finite horizon optimal policies fnf_nfn​ are any minimizers of vn(i)=min⁡a{C(i,a)+∑jPij(a)vn−1(j)}v_n(i) = \min_a\{C(i,a) + \sum_j P_{ij}(a) v_{n-1}(j)\}vn​(i)=mina​{C(i,a)+∑j​Pij​(a)vn−1​(j)}.
  • Aperiodicity of a class is the book's definition (Pij(n)→πjP^{(n)}_{ij} \to \pi_jPij(n)​→πj​ on the class), with πj=1/mjj\pi_j = 1/m_{jj}πj​=1/mjj​. Assumption OPA quantifies over average cost optimal stationary policies only, not over all stationary policies.
  • A trivializing formalization is ruled out: hhh, rnr_nrn​, dnd_ndn​ and vnv_nvn​ are computed from the model, not free functions constrained by the ACOE, and the ACOE conclusions are equalities of real numbers with the minimum over the actual action sets.
  • The model, criteria and Markov chain definitions restate those of mission IV of this series in their own namespace; they are reusable for any finite average cost result. Contributions of general Markov chain facts (convergence of P(n)P^{(n)}P(n) on aperiodic classes, Cesàro limits 1n∑tP(t)\frac1n\sum_t P^{(t)}n1​∑t​P(t)) are welcome.

Selected references

  • L. I. Sennott, Stochastic Dynamic Programming and the Control of Queueing Systems, Wiley Series in Probability and Statistics, Wiley, 1999. https://doi.org/10.1002/9780470317037
  • D. Blackwell, Discrete dynamic programming, Annals of Mathematical Statistics 33 (1962), 719–726. https://doi.org/10.1214/aoms/1177704593
  • P. J. Schweitzer and A. Federgruen, The asymptotic behavior of undiscounted value iteration in Markov decision problems, Mathematics of Operations Research 2 (1977), 360–381. https://doi.org/10.1287/moor.2.4.360
  • M. L. Puterman, Markov Decision Processes: Discrete Stochastic Dynamic Programming, Wiley, 1994. https://doi.org/10.1002/9780470316887
13 thms1 active userReviewed
Operations ResearchProbabilityStochastic Systems·Captain: mikedeng1

Stochastic Dynamic Programming and the Control of Queueing Systems XIV: Conforming Approximating Sequences for Markov ChainsTextbook

Motivation

Countable-state Markov chains are the standard model of queues with unbounded buffers, but any numerical computation of their long-run behaviour works on a finite state space. The usual remedy is truncation: restrict the chain to a finite set SNS_NSN​ and redistribute the probability of leaving SNS_NSN​ back into it. Whether the steady state probabilities and average costs of the truncated chains converge to those of the original chain depends on how that probability is redistributed. Gibson and Seneta studied this question for the stationary distributions of chains without costs (Gibson and Seneta, J. Appl. Prob., 1987). Sennott extended it to chains with costs and expected first passage costs (Sennott, Adv. Appl. Prob. 29, 1997; ZOR Math. Meth. Oper. Res. 45, 1997), and used it as the basis of the approximating sequence method for average-cost Markov decision chains (Sennott, 1999, Chapter 8). This mission covers Appendix C, Sections C.4–C.5 of the 1999 book, the Markov-chain results that the book's average-cost approximation theorems use.

Setting

A Markov chain with costs Γ\GammaΓ on a denumerable state space SSS has transition probabilities PijP_{ij}Pij​ with ∑jPij=1\sum_jP_{ij}=1∑j​Pij​=1 and a finite nonnegative cost C(i)C(i)C(i) at each state. For a set G⊆SG\subseteq SG⊆S and a start iii, TiG≥1T_{iG}\ge 1TiG​≥1 is the first passage time to GGG. The taboo probability GPik(t){}_GP^{(t)}_{ik}G​Pik(t)​ is the probability of moving from iii to kkk in ttt steps with no intermediate state in GGG. The expected visits Guik{}_Gu_{ik}G​uik​ count the visits to kkk at times 0≤t<TiG0\le t<T_{iG}0≤t<TiG​. The mean first passage time is miG=E[TiG]m_{iG}=E[T_{iG}]miG​=E[TiG​], infinite when GGG is missed with positive probability. The first passage cost is ciG=E[∑t<TiGC(Xt)]c_{iG}=E\big[\sum_{t<T_{iG}}C(X_t)\big]ciG​=E[∑t<TiG​​C(Xt​)]. A state iii is positive recurrent when mii<∞m_{ii}<\inftymii​<∞, and the steady state probability is πi=mii−1\pi_i=m_{ii}^{-1}πi​=mii−1​. On a positive recurrent class RRR the average cost is JR=∑j∈RπjC(j)J_R=\sum_{j\in R}\pi_jC(j)JR​=∑j∈R​πj​C(j). The chain is zzz standard when miz<∞m_{iz}<\inftymiz​<∞ and ciz<∞c_{iz}<\inftyciz​<∞ for every iii. Such a chain has one positive recurrent class R∋zR\ni zR∋z with JR<∞J_R<\inftyJR​<∞, and every other state is transient.

An approximating sequence (AS) (ΓN)N≥N0(\Gamma_N)_{N\ge N_0}(ΓN​)N≥N0​​ consists of increasing nonempty finite sets SNS_NSN​ with ⋃NSN=S\bigcup_NS_N=S⋃N​SN​=S and, for each NNN, a chain ΓN\Gamma_NΓN​ on SNS_NSN​ with the same costs and transition probabilities Pij(N)→PijP_{ij}(N)\to P_{ij}Pij​(N)→Pij​. The quantities of ΓN\Gamma_NΓN​ are written miG(N)m_{iG}(N)miG​(N), ciG(N)c_{iG}(N)ciG​(N), πi(N)\pi_i(N)πi​(N) and J(i)(N)J(i)(N)J(i)(N). An AS is conforming (for a zzz standard Γ\GammaΓ) if, for large NNN, ΓN\Gamma_NΓN​ is unichain with zzz in its positive recurrent class, and miz(N)→mizm_{iz}(N)\to m_{iz}miz​(N)→miz​ and ciz(N)→cizc_{iz}(N)\to c_{iz}ciz​(N)→ciz​ for all iii. It is conforming on RRR if πi(N)→πi\pi_i(N)\to\pi_iπi​(N)→πi​ and J(i)(N)→JRJ(i)(N)\to J_RJ(i)(N)→JR​ on RRR.

An augmentation type approximating sequence (ATAS) keeps the original probabilities inside SNS_NSN​ and redistributes the probability of each excluded target r∉SNr\notin S_Nr∈/SN​ according to an augmentation distribution q⋅(i,r,N)q_\cdot(i,r,N)q⋅​(i,r,N) on SNS_NSN​:

Pij(N)=Pij+∑r∈S−SNPir qj(i,r,N),j∈SN.P_{ij}(N)=P_{ij}+\sum_{r\in S-S_N}P_{ir}\,q_j(i,r,N),\qquad j\in S_N.Pij​(N)=Pij​+r∈S−SN​∑​Pir​qj​(i,r,N),j∈SN​.

It sends excess probability to GGG if every q⋅(i,r,N)q_\cdot(i,r,N)q⋅​(i,r,N) is concentrated on GGG.

Formalization targets

Goal: Proposition C.5.2

For a zzz standard chain Γ\GammaΓ and a finite nonempty G⊆SG\subseteq SG⊆S,

every ATAS that sends excess probability to G is conforming,\text{every ATAS that sends excess probability to } G \text{ is conforming},every ATAS that sends excess probability to G is conforming,

and if G⊆RG\subseteq RG⊆R it is also conforming on RRR. No rate of convergence and no constants are involved, and GGG need not contain zzz.

Milestones

  1. Proposition C.4.2: for fixed ttt, lim⁡NGPik(t)(N)=GPik(t)\lim_N{}_GP^{(t)}_{ik}(N)={}_GP^{(t)}_{ik}limN​G​Pik(t)​(N)=G​Pik(t)​; also lim inf⁡NGuik(N)≥Guik\liminf_N{}_Gu_{ik}(N)\ge{}_Gu_{ik}liminfN​G​uik​(N)≥G​uik​ and lim inf⁡NmiG(N)≥miG\liminf_Nm_{iG}(N)\ge m_{iG}liminfN​miG​(N)≥miG​.
  2. Proposition C.4.3: πi(N)→0\pi_i(N)\to0πi​(N)→0 off the positive recurrent states, and along subsequences πi(Ns)→bπi\pi_i(N_s)\to b\pi_iπi​(Ns​)→bπi​ on a class, with 0≤b≤10\le b\le10≤b≤1.
  3. Proposition C.4.5: lim inf⁡NciG(N)≥ciG\liminf_Nc_{iG}(N)\ge c_{iG}liminfN​ciG​(N)≥ciG​.
  4. Proposition C.4.6: on a positive recurrent class, convergence of π\piπ, of mzzm_{zz}mzz​ and of all miGm_{iG}miG​ are equivalent. Given these, convergence of J(i)J(i)J(i), of czzc_{zz}czz​ and of all ciGc_{iG}ciG​ are equivalent.
  5. Proposition C.4.9: conformity implies πi(N)→πi\pi_i(N)\to\pi_iπi​(N)→πi​ for all iii, and that the constant average costs J(N)J(N)J(N) of ΓN\Gamma_NΓN​ converge to JRJ_RJR​.

Further results

  1. Proposition C.5.3: an ATAS is conforming when, for N≥N∗N\ge N^*N≥N∗, the augmentation distributions satisfy ∑j≠zqj(i,r,N)mjz≤mrz\sum_{j\ne z}q_j(i,r,N)m_{jz}\le m_{rz}∑j=z​qj​(i,r,N)mjz​≤mrz​ and ∑j≠zqj(i,r,N)cjz≤crz\sum_{j\ne z}q_j(i,r,N)c_{jz}\le c_{rz}∑j=z​qj​(i,r,N)cjz​≤crz​.
  2. Corollary C.5.4: for a 000 standard chain on {0,1,2,… }\{0,1,2,\dots\}{0,1,2,…} with an upper Hessenberg transition matrix, truncated to SN={0,…,N}S_N=\{0,\dots,N\}SN​={0,…,N} with the excess sent to NNN, the ATAS is conforming.

Significance

The result. Conformity is the hypothesis under which the book's approximating sequence method works for average-cost queueing control (Chapter 8). The method computes optimal policies for finite truncations and passes to the limit. That argument needs the first passage times and costs to a distinguished state to converge along the chains induced by fixed policies. Propositions C.5.2 and C.5.3 turn this analytic requirement into conditions on the truncation scheme that can be checked in practice: send the overflow to a fixed finite set, or to states from which reaching zzz is no more expensive. Examples C.4.4 and C.4.7 of the book show that an arbitrary approximating sequence can fail. The limit of the steady state probabilities can be a strict multiple bπb\pibπ with b<1b<1b<1. First passage costs can converge to the wrong value even when the steady state probabilities converge.

Formalizing it. The results are proved in the book, some in abbreviated form ("the proof for the costs is similar and is omitted"). The Prove2Me library had no statement on truncation or augmentation of countable Markov chains when this mission was drafted (September 2026). A formalization supplies the omitted cost arguments, makes the passage between lim inf⁡\liminfliminf bounds and limits in [0,∞][0,\infty][0,∞] explicit, and produces a reusable library of first passage quantities for countable chains.

Difficulty

The lower bounds of Propositions C.4.2 and C.4.5 are the routine part. The difficulty is the matching upper bound: in ΓN\Gamma_NΓN​, a first passage that leaves SNS_NSN​ is restarted elsewhere, which can lengthen it without bound. Taking limits termwise in the first passage equation miz(N)=1+∑j≠zPij(N)mjz(N)m_{iz}(N)=1+\sum_{j\ne z}P_{ij}(N)m_{jz}(N)miz​(N)=1+∑j=z​Pij​(N)mjz​(N) fails, because no dominating function is available and mass can escape to infinity. Example C.4.4 exhibits exactly this. Unichain structure is also not automatic: ΓN\Gamma_NΓN​ may have several recurrent classes, or a recurrent class not containing zzz, and ruling this out is part of the conclusion rather than an assumption.

Formalization scope

The Lean development works in SennottDP.ChainASM. A chain is a structure MC S with P : S → S → ℝ≥0∞, ∑' j, P i j = 1 and C : S → ℝ≥0. Theorems assume [Countable S] [Infinite S], matching the book's denumerable state space. Taboo probabilities, expected visits, miGm_{iG}miG​, ciGc_{iG}ciG​, πj=(mjj)−1\pi_j=(m_{jj})^{-1}πj​=(mjj​)−1 and JR=∑j∈RπjC(j)J_R=\sum_{j\in R}\pi_jC(j)JR​=∑j∈R​πj​C(j) are defined as sums in [0,∞][0,\infty][0,∞]. miG=∑t≥0P(TiG>t)m_{iG}=\sum_{t\ge0}P(T_{iG}>t)miG​=∑t≥0​P(TiG​>t) is infinite whenever GGG is missed with positive probability. The average cost J(i)J(i)J(i) is the lim sup⁡\limsuplimsup of the Cesàro cost averages.

An AS is a structure carrying N0N_0N0​, the finite sets SNS_NSN​ (as Finset S) and Pij(N)P_{ij}(N)Pij​(N). ΓN\Gamma_NΓN​ is built as an MC on the subtype of SNS_NSN​, and a set GGG is read in ΓN\Gamma_NΓN​ as G∩SNG\cap S_NG∩SN​. Quantities of ΓN\Gamma_NΓN​ are lifted to functions of NNN and of states of SSS with the value 000 where they are undefined (N<N0N<N_0N<N0​ or a state outside SNS_NSN​). For fixed states this affects finitely many NNN, and all statements are limits, lim inf⁡\liminfliminfs or eventual equalities. All convergence is in [0,∞][0,\infty][0,∞]. The conformity predicate includes the standing assumption that Γ\GammaΓ is zzz standard. The positive recurrent class RRR of a zzz standard chain is the communicating class of zzz.

A trivializing formalization is excluded: the AS of Example C.4.4, whose positive recurrent class {N}\{N\}{N} excludes z=0z=0z=0, is not conforming under these definitions. The ATAS predicate requires the augmentation distributions to be probability distributions and to reproduce Pij(N)P_{ij}(N)Pij​(N) exactly by (C.27).

A complete development needs first passage decompositions for countable chains, the renewal-reward identity JR=czz/mzzJ_R=c_{zz}/m_{zz}JR​=czz​/mzz​, and dominated and Fatou-type limit theorems for sums (the book's Appendix A). The first passage library and the lifted-quantity conventions can be reused by the average-cost approximation chapters. Contributions of intermediate lemmas are welcome, especially the finite-state unichain facts of Section C.3 and the identities of Propositions C.1.4 and C.2.2.

Proposition C.5.5 (lower Hessenberg chains, from Gibson and Seneta) is stated in the book without proof and without naming the distinguished state, and is not included.

Selected references

  • L. I. Sennott, Stochastic Dynamic Programming and the Control of Queueing Systems, Wiley, 1999, Appendix C, Sections C.4–C.5. https://doi.org/10.1002/9780470317037
  • L. I. Sennott, "The computation of average optimal policies in denumerable state Markov decision chains", Advances in Applied Probability 29 (1997) 114–137 (cited in the book as Sennott 1997a).
  • L. I. Sennott, "On computing average cost optimal policies with application to routing to parallel queues", ZOR Mathematical Methods of Operations Research 45 (1997) 45–62 (cited in the book as Sennott 1997b).
  • D. Gibson and E. Seneta, "Augmented truncations of infinite stochastic matrices", Journal of Applied Probability (1987).
10 thms1 active userReviewed
PreviousPage 3 of 4Next

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me