Prove2Me
Navigate
DiscoverFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

Probability

550 missions · 275 completed

Missions

Open275Completed275All550
Operations ResearchOptimizationStochastic Systems·Captain: mikedeng1

Dimensioning Large Call Centers IV: Asymptotically Optimal Staffing under a Waiting-Cost ConstraintResearch Paper

Motivation

A call center has to decide how many agents to staff. In practice the decision is often posed as a service-level constraint rather than a cost trade-off: use the fewest agents for which the expected waiting cost, or the fraction of customers who wait, stays below a target. Borst, Mandelbaum and Reiman (CWI Report PNA-R0015, 2000; journal version in Operations Research 52(1), 2004, doi:10.1287/opre.1030.0081) treat this constraint problem in Section 8 of their paper, alongside the cost-minimization problem of Sections 5–7, and show that a simple square-root staffing rule solves it asymptotically as the arrival rate grows.

The rule matters because it is what practitioners use. Under the classical Erlang-C model, the exact optimum requires evaluating the Erlang-C formula over many staffing levels. The asymptotic rule replaces this with a single equation in the Halfin–Whitt function PPP: when the target is a delay probability ε\varepsilonε (Example 8.5 of the paper), it reduces to staffing λ/μ+P−1(ε)λ/μ\lambda/\mu + P^{-1}(\varepsilon)\sqrt{\lambda/\mu}λ/μ+P−1(ε)λ/μ​ servers.

Timeline. Erlang's formula for the M/M/N delay probability dates from 1917. Halfin and Whitt (Operations Research 29, 1981) identified the limit P(x)P(x)P(x) of the delay probability under square-root staffing N=λ/μ+xλ/μN = \lambda/\mu + x\sqrt{\lambda/\mu}N=λ/μ+xλ/μ​ with integer NNN. Jagers and Van Doorn (Operations Research Letters 5, 1986; SIAM Review 33, 1991) studied the continued Erlang loss and delay functions at non-integer numbers of servers, including their convexity, which is what lets the staffing problem be relaxed to a continuous one. Borst, Mandelbaum and Reiman (2000/2004) used these to prove asymptotic optimality of square-root rules for both the cost and the constraint formulations.

Setting

Customers arrive at rate λ\lambdaλ to NNN identical servers, each with service rate μ>0\mu > 0μ>0; μ\muμ is fixed while λ→∞\lambda \to \inftyλ→∞. Stability requires N>λ/μN > \lambda/\muN>λ/μ. A customer who waits ttt time units costs Dλ(t)D_\lambda(t)Dλ​(t), where Dλ(0)=0D_\lambda(0) = 0Dλ​(0)=0, DλD_\lambdaDλ​ is strictly increasing on [0,∞)[0,\infty)[0,∞) and ∫0∞Dλ(t)e−θt dt<∞\int_0^\infty D_\lambda(t)e^{-\theta t}\,dt < \infty∫0∞​Dλ​(t)e−θtdt<∞ for all θ>0\theta > 0θ>0.

The Erlang-C probability of waiting is

π(N,ν)=νNN!{(1−ν/N)∑n=0N−1νnn!+νNN!}−1,\pi(N,\nu) = \frac{\nu^N}{N!}\Big\{(1-\nu/N)\sum_{n=0}^{N-1}\frac{\nu^n}{n!} + \frac{\nu^N}{N!}\Big\}^{-1},π(N,ν)=N!νN​{(1−ν/N)n=0∑N−1​n!νn​+N!νN​}−1,

and the conditional waiting cost is G(N,λ)=(Nμ−λ)∫0∞Dλ(t)e−(Nμ−λ)t dtG(N,\lambda) = (N\mu-\lambda)\int_0^\infty D_\lambda(t)e^{-(N\mu-\lambda)t}\,dtG(N,λ)=(Nμ−λ)∫0∞​Dλ​(t)e−(Nμ−λ)tdt. The waiting cost per unit time with NNN servers is

K(N,λ)=λ π(N,λ/μ) G(N,λ).K(N,\lambda) = \lambda\,\pi(N,\lambda/\mu)\,G(N,\lambda).K(N,λ)=λπ(N,λ/μ)G(N,λ).

Given a target Mλ>0M_\lambda > 0Mλ​>0, the optimal staffing level is the least integer N>λ/μN > \lambda/\muN>λ/μ with K(N,λ)≤MλK(N,\lambda) \le M_\lambdaK(N,λ)≤Mλ​; call it Nλ∗N^*_\lambdaNλ∗​.

In the continuous parametrization Nλ(x)=λ/μ+xλ/μN_\lambda(x) = \lambda/\mu + x\sqrt{\lambda/\mu}Nλ​(x)=λ/μ+xλ/μ​, define Gλ(x)=λG(Nλ(x),λ)G_\lambda(x) = \lambda G(N_\lambda(x),\lambda)Gλ​(x)=λG(Nλ​(x),λ), the continuous Erlang-C function πλ(x)=H(Nλ(x),λ/μ)\pi_\lambda(x) = H(N_\lambda(x),\lambda/\mu)πλ​(x)=H(Nλ​(x),λ/μ) with H(M,α)={α∫0∞e−αtt(1+t)M−1dt}−1H(M,\alpha) = \{\alpha\int_0^\infty e^{-\alpha t}t(1+t)^{M-1}dt\}^{-1}H(M,α)={α∫0∞​e−αtt(1+t)M−1dt}−1, and Kλ(x)=πλ(x)Gλ(x)K_\lambda(x) = \pi_\lambda(x)G_\lambda(x)Kλ​(x)=πλ​(x)Gλ​(x). The Halfin–Whitt function is P(x)=1/(1+x/h(−x))P(x) = 1/(1 + x/h(-x))P(x)=1/(1+x/h(−x)) with h=ϕ/(1−Φ)h = \phi/(1-\Phi)h=ϕ/(1−Φ) the standard normal hazard rate. A staffing function xλ>0x_\lambda > 0xλ​>0 is judged by the rounding gap

Tλ(x)=min⁡{∣K(⌊Nλ(x)⌋,λ)−Mλ∣, ∣K(⌈Nλ(x)⌉,λ)−Mλ∣, ∣K(⌈Nλ(x)⌉,λ)−K(Nλ∗,λ)∣}.T_\lambda(x) = \min\big\{|K(\lfloor N_\lambda(x)\rfloor,\lambda) - M_\lambda|,\ |K(\lceil N_\lambda(x)\rceil,\lambda) - M_\lambda|,\ |K(\lceil N_\lambda(x)\rceil,\lambda) - K(N^*_\lambda,\lambda)|\big\}.Tλ​(x)=min{∣K(⌊Nλ​(x)⌋,λ)−Mλ​∣, ∣K(⌈Nλ​(x)⌉,λ)−Mλ​∣, ∣K(⌈Nλ​(x)⌉,λ)−K(Nλ∗​,λ)∣}.

It is asymptotically optimal when Tλ(xλ)/Mλ→0T_\lambda(x_\lambda)/M_\lambda \to 0Tλ​(xλ​)/Mλ​→0 as λ→∞\lambda\to\inftyλ→∞.

Formalization targets

Goal: Theorem 8.2 (rationalized regime)

Suppose that for some κ>0\kappa > 0κ>0 and γ∈(0,∞)\gamma \in (0,\infty)γ∈(0,∞), Gλ(κ)/Mλ→γG_\lambda(\kappa)/M_\lambda \to \gammaGλ​(κ)/Mλ​→γ, i.e. the waiting cost is comparable to the target. Let yλ∗>0y^*_\lambda > 0yλ∗​>0 solve P(y)Gλ(y)=MλP(y)G_\lambda(y) = M_\lambdaP(y)Gλ​(y)=Mλ​. Then

lim⁡λ→∞Tλ(yλ∗)Mλ=0.\lim_{\lambda\to\infty}\frac{T_\lambda(y^*_\lambda)}{M_\lambda} = 0.λ→∞lim​Mλ​Tλ​(yλ∗​)​=0.

Supporting milestones

  • Lemma C.1: GλG_\lambdaGλ​ is strictly convex and decreasing on (0,∞)(0,\infty)(0,∞).
  • Section 3: πλ(x)=π(Nλ(x),λ/μ)\pi_\lambda(x) = \pi(N_\lambda(x),\lambda/\mu)πλ​(x)=π(Nλ​(x),λ/μ) when Nλ(x)N_\lambda(x)Nλ​(x) is an integer.
  • Lemma 8.1: if zλ∗>0z^*_\lambda > 0zλ∗​>0 solves π^λ(z)G^λ(z)=Mλ\hat\pi_\lambda(z)\hat G_\lambda(z) = M_\lambdaπ^λ​(z)G^λ​(z)=Mλ​ and Kλ(zλ∗)/(π^λG^λ)(zλ∗)→1K_\lambda(z^*_\lambda)/(\hat\pi_\lambda\hat G_\lambda)(z^*_\lambda) \to 1Kλ​(zλ∗​)/(π^λ​G^λ​)(zλ∗​)→1, then Tλ(zλ∗)/Mλ→0T_\lambda(z^*_\lambda)/M_\lambda \to 0Tλ​(zλ∗​)/Mλ​→0.
  • Lemma B.1: PPP is strictly convex and decreasing on (0,∞)(0,\infty)(0,∞).
  • Eq. (17): lim sup⁡aλ/b=∞\limsup a_\lambda/b = \inftylimsupaλ​/b=∞ implies lim inf⁡P(aλ)/P(b)=0\liminf P(a_\lambda)/P(b) = 0liminfP(aλ​)/P(b)=0 and lim inf⁡πλ(aλ)/πλ(b)=0\liminf \pi_\lambda(a_\lambda)/\pi_\lambda(b) = 0liminfπλ​(aλ​)/πλ​(b)=0.
  • Lemma 4.1 (Halfin–Whitt): for bounded xλ>0x_\lambda > 0xλ​>0, πλ(xλ)/P(xλ)→1\pi_\lambda(x_\lambda)/P(x_\lambda) \to 1πλ​(xλ​)/P(xλ​)→1; with xλ→xx_\lambda \to xxλ​→x, πλ(xλ)/P(x)→1\pi_\lambda(x_\lambda)/P(x)\to 1πλ​(xλ​)/P(x)→1.

Further target: Theorem 8.6 (efficiency-driven regime)

If Gλ(κ)/Mλ→0G_\lambda(\kappa)/M_\lambda \to 0Gλ​(κ)/Mλ​→0 for every κ>0\kappa > 0κ>0 and yλ∗>0y^*_\lambda > 0yλ∗​>0 solves Gλ(y)=MλG_\lambda(y) = M_\lambdaGλ​(y)=Mλ​, then Tλ(yλ∗)/Mλ→0T_\lambda(y^*_\lambda)/M_\lambda \to 0Tλ​(yλ∗​)/Mλ​→0.

Significance

The theorem certifies the staffing rule used in workforce-management practice: the excess staffing is determined by one scalar equation involving the Gaussian function PPP and the scaled waiting cost, and rounding the resulting staffing level misses the constraint by a vanishing fraction of the target. Lemma 8.1 is a reusable framework: any approximation π^λG^λ\hat\pi_\lambda\hat G_\lambdaπ^λ​G^λ​ that is asymptotically exact at the proposed staffing level yields an asymptotically optimal rule, and the paper instantiates it in three regimes (Theorems 8.2, 8.6, 8.9).

The results are proved on paper. To the best of current knowledge none of them, nor the Halfin–Whitt limit for the continuous Erlang-C extension, has a machine-checked proof. A formalization would produce the first verified heavy-traffic limit of the Erlang-C delay probability, a verified continuous Erlang-C extension with its integer identity, and the convexity facts about PPP and GλG_\lambdaGλ​ that many staffing papers cite without proof.

Difficulty

The obvious argument is to quote Halfin and Whitt: the delay probability converges to P(x)P(x)P(x) under square-root staffing, so PPP can replace the Erlang-C formula. That limit, as published in 1981, is about integer server counts along sequences with a convergent excess-staffing parameter. The paper needs it for the continuous function HHH at non-integer server counts and for staffing functions that are merely bounded, and it also needs the identity H(N,ν)=π(N,ν)H(N,\nu) = \pi(N,\nu)H(N,ν)=π(N,ν) at integers and the monotonicity of πλ\pi_\lambdaπλ​ in xxx, both cited from Jagers and Van Doorn rather than proved. None of these is in Mathlib. A second obstacle is that the staffing function yλ∗y^*_\lambdayλ∗​ is defined only implicitly by an equation involving GλG_\lambdaGλ​, which depends on the arbitrary cost functions DλD_\lambdaDλ​; nothing a priori prevents it from escaping to infinity, outside the range where the Halfin–Whitt approximation applies. Finally, TλT_\lambdaTλ​ compares integer-level costs given by the Erlang-C formula with a continuous approximation, so both representations of the delay probability are in play at once.

Formalization scope

The queue itself is not formalized: there is no Markov chain and no waiting-time distribution. Every statement is about the closed-form waiting cost K(N,λ)K(N,\lambda)K(N,λ) with π\piπ given by the Erlang-C formula, exactly as the paper's analysis is. Conventions, all in the namespace DimCallCenters.Constraint:

  • lam : ℝ is the arrival rate (λ is a Lean keyword); limits are Filter.atTop in lam, with μ fixed. Objects indexed by λ (MλM_\lambdaMλ​, Nλ∗N^*_\lambdaNλ∗​, yλ∗y^*_\lambdayλ∗​) are functions of lam constrained only for lam > 0.
  • WaitModel packages μ > 0 and DλD_\lambdaDλ​ with Dλ(0)=0D_\lambda(0) = 0Dλ​(0)=0, strict monotonicity on [0,∞)[0,\infty)[0,∞), and integrability of Dλ(t)e−θtD_\lambda(t)e^{-\theta t}Dλ​(t)e−θt on (0,∞)(0,\infty)(0,∞) for θ > 0 (the paper's finiteness of GGG; integrability is required because Lean's integral of a non-integrable function is 0).
  • Nλ∗N^*_\lambdaNλ∗​ is a function Nstar : ℝ → ℕ given with its two defining properties (feasible; below every feasible integer level above λ/μ). yλ∗y^*_\lambdayλ∗​ and zλ∗z^*_\lambdazλ∗​ are any positive solutions of their equations; existence and uniqueness are not hypotheses.
  • In TλT_\lambdaTλ​ the round-down term is dropped when ⌊Nλ(x)⌋≤λ/μ\lfloor N_\lambda(x)\rfloor \le \lambda/\mu⌊Nλ​(x)⌋≤λ/μ (an unstable level where KKK is undefined). This can only enlarge TλT_\lambdaTλ​.
  • Asymptotic relations are limits of ratios. lim sup⁡=∞\limsup = \inftylimsup=∞ and lim inf⁡=0\liminf = 0liminf=0 are stated with ∃ᶠ ("frequently"), lim sup⁡<∞\limsup < \inftylimsup<∞ as eventual boundedness.
  • PPP is defined through explicit ϕ\phiϕ, Φ\PhiΦ, hhh; the formula also gives P(0)=1P(0) = 1P(0)=1, used in Lemma 4.1(2) at x=0x = 0x=0.
  • No hypothesis lim⁡N↓λ/μG(N,λ)=∞\lim_{N\downarrow\lambda/\mu}G(N,\lambda) = \inftylimN↓λ/μ​G(N,λ)=∞ is added: it is not needed for the statements here.

A trivializing formalization is ruled out: TλT_\lambdaTλ​ keeps all of the paper's terms and is never replaced by a smaller quantity, and the hypotheses are jointly satisfiable — Dλ(t)=aλ/μ tD_\lambda(t) = a\sqrt{\lambda/\mu}\,tDλ​(t)=aλ/μ​t with Mλ=MλM_\lambda = M\lambdaMλ​=Mλ satisfies (33) for every κ\kappaκ with γ=a/(μκM)\gamma = a/(\mu\kappa M)γ=a/(μκM).

Infrastructure needed: the continuous Erlang-C function and its integer identity; the Halfin–Whitt limit (a Gaussian approximation of Poisson/gamma tails); calculus facts about the normal hazard rate. These are reusable beyond this mission, notably by the sibling missions on the cost-minimization problem. Example 8.5 (delay-probability target with Dλ=1t>0D_\lambda = 1_{t>0}Dλ​=1t>0​) motivates the rule but violates the strict monotonicity of DλD_\lambdaDλ​, so it is not an instance of the theorem as stated. Contributions on any milestone, and on Theorem 8.9 (quality-driven regime, which needs Lemma 4.2), are welcome.

Selected references

  • S. Borst, A. Mandelbaum, M. I. Reiman, Dimensioning Large Call Centers, CWI Report PNA-R0015, 2000; Operations Research 52(1):17–34, 2004. https://doi.org/10.1287/opre.1030.0081
  • S. Halfin, W. Whitt, Heavy-Traffic Limits for Queues with Many Exponential Servers, Operations Research 29(3):567–588, 1981. https://doi.org/10.1287/opre.29.3.567
  • A. A. Jagers, E. A. Van Doorn, On the Continued Erlang Loss Function, Operations Research Letters 5:43–46, 1986.
  • A. A. Jagers, E. A. Van Doorn, Convexity of Functions which are Generalizations of the Erlang Loss Function and the Erlang Delay Function, SIAM Review 33:281–282, 1991.
18 thms1 active userReviewed
Machine LearningStatistics·Captain: mikedeng1

The Sample Complexity of Pattern Classification with Neural Networks: The Size of the Weights is More Important than the Size of the Network III: Sigmoid Networks with Small Weights GeneralizeResearch Paper

Motivation

A classifier built from a neural network produces a real score and predicts a binary label from its sign. A network can have many hidden units, so a guarantee based only on the number of parameters can be uninformative even when its output weights are small. Bartlett's 1998 paper asks whether a classifier's margin on training examples and the total magnitude of its weights can control its probability of error without fixing the number of units. Its Theorem 28 gives such a statement for two-layer networks whose activation is bounded and nondecreasing. The paper also discusses why this parameter-magnitude view supports weight decay and early stopping as learning heuristics, while leaving their algorithmic behavior outside the theorem's scope (Bartlett 1998, pp. 526, 534–535).

The theorem combines two results in the same paper. Theorem 2 turns the fat-shattering dimension of a real-valued function class into a margin generalization bound. Corollary 24 controls that dimension for finite combinations of affine-input units when the sum of the absolute combination weights is bounded. Lemmas 19, 22, and 23 supply covering estimates along that path. These are the milestones of this mission, with the source statements preserved in the milestone record (Bartlett 1998, pp. 527, 532–534).

Setting

An input is a vector x∈Rnx\in\mathbb R^nx∈Rn, represented in Lean as Fin n → ℝ. A label is y∈{−1,+1}y\in\{-1,+1\}y∈{−1,+1}; Lean's Bool is converted by pm, where true means +1+1+1. A probability distribution PPP lives on labeled inputs. From an independent sample z=((xi,yi))i=1mz=((x_i,y_i))_{i=1}^mz=((xi​,yi​))i=1m​, the empirical margin error at scale γ>0\gamma>0γ>0 is the fraction of indices with yih(xi)<γy_i h(x_i)<\gammayi​h(xi​)<γ. The population error is the probability that sgn⁡(h(x))≠y\operatorname{sgn}(h(x))\ne ysgn(h(x))=y, where sgn⁡(0)=+1\operatorname{sgn}(0)=+1sgn(0)=+1. The inequality in the empirical error is strict, as in the paper's definition (Bartlett 1998, p. 526).

Fix a bounded nondecreasing activation σ:R→[−1,1]\sigma:\mathbb R\to[-1,1]σ:R→[−1,1]. The first-layer class FFF contains every x↦σ(w⋅x+w0)x\mapsto\sigma(w\cdot x+w_0)x↦σ(w⋅x+w0​), with an arbitrary weight vector and bias. The network class HHH contains all finite sums ∑i=1Nαifi\sum_{i=1}^N\alpha_i f_i∑i=1N​αi​fi​ with fi∈Ff_i\in Ffi​∈F and ∑i∣αi∣≤A\sum_i|\alpha_i|\le A∑i​∣αi​∣≤A. Thus AAA bounds the output layer's total weight magnitude, while NNN can vary without an imposed width limit. The bias w0w_0w0​ is part of every unit. For a class GGG, fat⁡G(η)\operatorname{fat}_G(\eta)fatG​(η) records the largest length of an input sequence whose every sign pattern can be realized with separation at least η\etaη around one vector of thresholds (Bartlett 1998, pp. 526, 533–534).

Formalization targets

Two-layer generalization

For 0<γ≤10<\gamma\le10<γ≤1, 0<δ<1/20<\delta<1/20<δ<1/2, A≥1A\ge1A≥1, and an independent sample of length m≥1m\ge1m≥1, the goal is one universal c>0c>0c>0 such that, with probability at least 1−δ1-\delta1−δ, every h∈Hh\in Hh∈H satisfies

er⁡P(h)<er⁡^zγ(h)+cm(A2nγ2log⁡ ⁣(32Aγ)(log⁡m)2+log⁡ ⁣(1δ)).\operatorname{er}_P(h)<\widehat{\operatorname{er}}_z^\gamma(h)+ \sqrt{\frac{c}{m}\left( \frac{A^2n}{\gamma^2}\log\!\left(\frac{32A}{\gamma}\right)(\log m)^2+ \log\!\left(\frac1\delta\right)\right)}.erP​(h)<erzγ​(h)+mc​(γ2A2n​log(γ32A​)(logm)2+log(δ1​))​.

The paper prints log⁡(A/γ)\log(A/\gamma)log(A/γ) in this display. That term vanishes at A=γ=1A=\gamma=1A=γ=1, although the class can then contain halfspace classifiers with nonzero sample complexity. The proof obtains a positive factor at that corner through Corollary 24 at scale γ/16\gamma/16γ/16, giving log⁡(32A/γ)\log(32A/\gamma)log(32A/γ). The goal states this correction and records the printed statement separately in the moderation notes. The constant precedes all network, distribution, margin, confidence, and sample parameters in Lean; it cannot be selected after observing the instance (Bartlett 1998, pp. 533–534).

Capacity and margin milestones

Corollary 24 bounds fat⁡H(η)\operatorname{fat}_H(\eta)fatH​(η) by a constant multiple of M2A2nη−2log⁡(MA/η)M^2A^2n\eta^{-2}\log(MA/\eta)M2A2nη−2log(MA/η) when the activation has range [−M/2,M/2][-M/2,M/2][−M/2,M/2]. Theorem 2 then converts a finite fat dimension at scale γ/16\gamma/16γ/16 into a simultaneous bound on population error for all members of HHH. The three covering lemmas track how shattering, pseudodimension, and an ℓ1\ell_1ℓ1​ weight budget affect covers in sample ℓ1\ell_1ℓ1​, ℓ∞\ell_\inftyℓ∞​, and ℓ2\ell_2ℓ2​ distances. Each bound retains the scale and explicit constants printed by the paper, subject to the stated corrections to undefined or false boundary cases (Bartlett 1998, pp. 527, 532–533).

Significance

The goal gives a width-independent generalization guarantee for a chosen network when its empirical margin error and total output weight are small. It applies to the entire class HHH at once, so choosing a network after inspecting the sample does not turn the bound into a claim about only one fixed predictor. It does not assert that a learning algorithm finds such a network or that the displayed constants are optimal. Bartlett notes that later work had improved a logarithmic factor, and that empirical agreement with neural-network performance remained an open experimental question at the time (Bartlett 1998, pp. 534–535).

The paper proves the mathematical result. This mission asks for a machine-checked proof of its corrected formal statement and the stated supporting results; the draft theorem files currently contain proof obligations. A completed development would also make the fat dimension and strict external sample-cover definitions available for other margin analyses. Those objects differ from the platform's fixed-architecture neural networks and closed-ball covering numbers, so they are defined here with the conventions of this paper.

Difficulty

Counting hidden units gives no finite width-independent capacity bound, because HHH permits arbitrarily many terms. Bounding each unit separately also does not control the full combination class: different small contributions can produce distinct values on a sample. The challenging step is relating covers of the base class to covers of all finite combinations under the total absolute-weight constraint, and then relating those covers back to fat-shattering. Even once a finite capacity estimate is available, the probability statement must hold simultaneously for every h∈Hh\in Hh∈H, including a network selected after sampling (Bartlett 1998, pp. 532–534).

Formalization scope

Lean uses N∪{∞}\mathbb N\cup\{\infty\}N∪{∞} for fat dimensions and covering numbers, so an unbounded class cannot acquire a spurious dimension zero. Covers are external finite sets of real functions and use the strict distance <ε<\varepsilon<ε of Definition 3. Sample ℓ1\ell_1ℓ1​ and ℓ2\ell_2ℓ2​ distances are normalized by mmm. Pseudodimension is the supremum of positive-scale fat dimensions, matching the paper's right limit. Theorems assume m≥1m\ge1m≥1, and sample indices are zero-based. The network class is generated from its weights rather than supplied as an arbitrary set satisfying the desired bound.

The paper says it ignores measurability issues and assumes all sets considered are measurable (Bartlett 1998, p. 526). The goal makes the event of a violating network measurable. Its individual network functions are measurable from monotonicity of σ\sigmaσ and finite sums; the restated Theorem 2 has explicit hypotheses for measurable class members, the violating event, and the double-sample event in its proof. Theorem 2 additionally restricts d=fat⁡H(γ/16)d=\operatorname{fat}_H(\gamma/16)d=fatH​(γ/16) to d≤34md\le34md≤34m, where its printed logarithmic bound remains valid. Corollary 24 uses n≥1n\ge1n≥1 because a zero-dimensional input still permits a biased constant unit. Lemma 23 uses d≥1d\ge1d≥1 and 0<γ<emM/d0<\gamma<emM/d0<γ<emM/d in place of the printed γ≥0\gamma\ge0γ≥0: at γ=0\gamma=0γ=0 or d=0d=0d=0, or for γ≥emM/d\gamma\ge emM/dγ≥emM/d, the printed strict inequality fails, while on the rest of the printed range it is kept. These are recorded as corrections rather than attributed to the printed wording.

The deeper-network part of Theorem 28 is outside this proposal. Its printed chain through Corollary 27 has an unresolved range issue when the input box bound BBB is smaller than the activation range, and the displayed log⁡n\log nlogn factor also vanishes at n=1n=1n=1. This mission's goal is Part 1 and uses none of those claims. Contributions that establish the corrected covering lemmas, the capacity corollary, or the simultaneous margin bound fit the present proof frontier (Bartlett 1998, pp. 533–534).

Selected references

  • P. L. Bartlett, The Sample Complexity of Pattern Classification with Neural Networks: The Size of the Weights is More Important than the Size of the Network, IEEE Transactions on Information Theory 44(2), 525–536, 1998. DOI.
15 thms1 active userReviewed
AnalysisDynamic ProgrammingOperations Research+2·Captain: mikedeng1

Stochastic Optimal Control: The Discrete-Time Case VI: Lower Semianalytic Functions — Analytically Measurable ε-Optimal Selectors (Jankov–von Neumann)Textbook

Motivation

Dynamic programming over uncountable state and control spaces needs two things at every stage: the optimal cost-to-go, obtained by minimizing over the control, must be a function that can be integrated against the next stage's transition probabilities, and a policy that nearly attains the minimum must be measurable, so that it defines a stochastic process. With Borel-measurable costs and Borel-measurable policies both requirements fail. Minimizing a Borel function of (x,y)(x,y)(x,y) over yyy produces a function whose level sets are projections of Borel sets, and such projections need not be Borel (Suslin, 1917). The repair, developed by Blackwell, Freedman and Orkin (1974), Shreve and Bertsekas, and set out in Chapter 7 of Bertsekas and Shreve's Stochastic Optimal Control: The Discrete-Time Case (1978), is to enlarge the class of costs to the lower semianalytic functions and the class of policies to the analytically or universally measurable ones. Sections 7.6–7.7 of the book establish that this class is closed under partial minimization and admits measurable ε-optimal selectors. Chapters 8–10 of the book, and much of the later literature on Borel-space Markov decision processes (Hernández-Lerma and Lasserre; Feinberg and coauthors), build on these results.

Timeline:

  • 1917: Suslin shows that projections of Borel sets need not be Borel and introduces analytic sets; Lusin proves that analytic sets are universally measurable.
  • 1941–1949: Jankov and von Neumann independently prove that an analytic subset of a product admits a selector measurable with respect to the σ-algebra generated by analytic sets.
  • 1974: Blackwell, Freedman and Orkin use analytic sets to construct ε-optimal policies in Borel dynamic programming.
  • 1978: Bertsekas and Shreve give the treatment used here (§7.6–7.7), including the selection theorem for lower semianalytic functions, Proposition 7.50.

Setting

A Borel space is a topological space homeomorphic to a Borel subset of a complete separable metric space (Definition 7.7); its Borel σ-algebra is BX\mathscr B_XBX​. The Baire space is N=NN\mathscr N=\mathbb N^{\mathbb N}N=NN with the product topology. A set A⊆XA\subseteq XA⊆X is analytic if it is empty or the image of N\mathscr NN under a continuous map; by Proposition 7.41 this is the book's Definition 7.16 (the Suslin operation applied to closed sets). Every Borel set is analytic, and the converse fails when XXX is uncountable.

Three σ-algebras on XXX are in play. The analytic σ-algebra AX\mathscr A_XAX​ is generated by the analytic sets (Definition 7.19). The universal σ-algebra is UX=⋂pBX(p)\mathscr U_X=\bigcap_{p}\mathscr B_X(p)UX​=⋂p​BX​(p), the intersection over all probability measures ppp on (X,BX)(X,\mathscr B_X)(X,BX​) of the ppp-completions of BX\mathscr B_XBX​ (Definition 7.18). For a function fff from D⊆XD\subseteq XD⊆X into a Borel space YYY, fff is analytically measurable if D∈AXD\in\mathscr A_XD∈AX​ and f−1(B)∈AXf^{-1}(B)\in\mathscr A_Xf−1(B)∈AX​ for every B∈BYB\in\mathscr B_YB∈BY​, and universally measurable if the same holds with UX\mathscr U_XUX​ (Definition 7.20).

Let R∗=[−∞,∞]R^*=[-\infty,\infty]R∗=[−∞,∞]. A function f:D→R∗f:D\to R^*f:D→R∗ is lower semianalytic if DDD is analytic and {x∈D∣f(x)<c}\{x\in D\mid f(x)<c\}{x∈D∣f(x)<c} is analytic for every real ccc (Definition 7.21). For D⊆X×YD\subseteq X\times YD⊆X×Y write Dx={y∣(x,y)∈D}D_x=\{y\mid (x,y)\in D\}Dx​={y∣(x,y)∈D}, projX(D)={x∣Dx≠∅}\mathrm{proj}_X(D)=\{x\mid D_x\neq\emptyset\}projX​(D)={x∣Dx​=∅}, and define the partial infimum

f∗(x)=inf⁡y∈Dxf(x,y),x∈projX(D).f^*(x)=\inf_{y\in D_x}f(x,y),\qquad x\in\mathrm{proj}_X(D).f∗(x)=y∈Dx​inf​f(x,y),x∈projX​(D).

A selector is a function φ:projX(D)→Y\varphi:\mathrm{proj}_X(D)\to Yφ:projX​(D)→Y whose graph Gr(φ)\mathrm{Gr}(\varphi)Gr(φ) lies in DDD.

Formalization targets

Goal: Proposition 7.50

Let X,YX,YX,Y be Borel spaces, D⊆X×YD\subseteq X\times YD⊆X×Y analytic, and f:D→R∗f:D\to R^*f:D→R∗ lower semianalytic.

(a) For every ε>0\varepsilon>0ε>0 there is an analytically measurable selector φ\varphiφ with

f[x,φ(x)]≤{f∗(x)+εif f∗(x)>−∞,−1/εif f∗(x)=−∞.f[x,\varphi(x)]\le\begin{cases}f^*(x)+\varepsilon&\text{if }f^*(x)>-\infty,\\-1/\varepsilon&\text{if }f^*(x)=-\infty.\end{cases}f[x,φ(x)]≤{f∗(x)+ε−1/ε​if f∗(x)>−∞,if f∗(x)=−∞.​

(b) The set III of points where the infimum is attained is universally measurable, and for every ε>0\varepsilon>0ε>0 there is a universally measurable selector φ\varphiφ with f[x,φ(x)]=f∗(x)f[x,\varphi(x)]=f^*(x)f[x,φ(x)]=f∗(x) on III and the bounds of (a) off III.

The goal fixes no constant beyond the book's ε\varepsilonε and −1/ε-1/\varepsilon−1/ε.

Milestones

In attack order: Proposition 7.40 (Borel images and preimages of analytic sets are analytic), Corollary 7.42.1 (AX⊆UX\mathscr A_X\subseteq\mathscr U_XAX​⊆UX​), Corollary 7.44.2 (composites of analytically measurable maps are universally measurable), and Proposition 7.49, the Jankov–von Neumann theorem:

A⊆X×Y analytic ⟹ ∃ φ:projX(A)→Y analytically measurable, Gr(φ)⊆A.A\subseteq X\times Y\text{ analytic}\ \Longrightarrow\ \exists\,\varphi:\mathrm{proj}_X(A)\to Y\ \text{analytically measurable},\ \mathrm{Gr}(\varphi)\subseteq A.A⊆X×Y analytic ⟹ ∃φ:projX​(A)→Y analytically measurable, Gr(φ)⊆A.

Further items of the mission, on the same definitions: Proposition 7.39 (projections of analytic sets are analytic, and every analytic set is a projection of a Borel set), Lemma 7.30(1) (strict and non-strict, real and extended level sets give the same class) and Proposition 7.47 (lower semianalytic functions are exactly partial infima of Borel functions).

Significance

Proposition 7.50 is the selection theorem behind the existence of ε-optimal policies in Borel-space dynamic programming. In the finite-horizon model of Chapter 8 the optimal cost-to-go at each stage is lower semianalytic, by Propositions 7.47 and 7.48. Proposition 7.50 then turns the one-stage minimization into a measurable policy, analytically measurable when only ε-optimality is required and universally measurable when the minimum is attained. Chapters 8–9 of the book (the finite-horizon recursion JK∗=TK(J0)J^*_K=T^K(J_0)JK∗​=TK(J0​) and the optimality equation under (P), (N), (D)) use it at every step. Downstream catalog papers on average-cost and stochastic shortest-path problems over Borel spaces cite these results.

All results here are proved in the book and in the descriptive set theory literature (Kechris, Classical Descriptive Set Theory, §18 and §29). None is formalized on Prove2Me. Mathlib has analytic sets in Polish-type settings, the Lusin separation theorem and Suslin's theorem, but it has no universal σ-algebra, no analytic σ-algebra, no lower semianalytic functions and no Jankov–von Neumann uniformization. The definitions in this mission are reusable by the later missions of the series (Chapters 8–10), which restate them locally until these are published.

Difficulty

The obvious route to a selector is to choose, for each xxx, a minimizing or near-minimizing yyy. The axiom of choice provides such a function, but nothing makes it measurable, and the conclusion of the theorem is exactly that measurability. The Borel route fails too: the set {x∣f∗(x)<c}\{x\mid f^*(x)<c\}{x∣f∗(x)<c} is a projection of a Borel set, which is analytic but in general not Borel, so no Borel-measurable selector exists in general. The Jankov–von Neumann theorem needs a lexicographically least branch of a continuous parametrization of AAA by N\mathscr NN, and an argument that the resulting map is measurable with respect to AX\mathscr A_XAX​, which is generated by sets that are not closed under complementation. Part (b) adds a further obstacle: the composite of two analytically measurable maps need not be analytically measurable, so the exact selector is only universally measurable. Proving that requires Lusin's theorem that analytic sets are measurable for every completed probability measure.

Formalization scope

  • A Borel space is a type with a topology satisfying the class IsBorelSpace (Definition 7.7, the ambient complete separable metric space taken in the same universe), together with Mathlib's [MeasurableSpace X] [BorelSpace X], so measurable sets are exactly the Borel sets. On X×YX\times YX×Y the product σ-algebra is used; it coincides with BX×Y\mathscr B_{X\times Y}BX×Y​ for separable metrizable spaces (Proposition 7.13).
  • Analytic sets are Mathlib's MeasureTheory.AnalyticSet (empty or a continuous image of ℕ → ℕ).
  • R∗R^*R∗ is EReal. The book uses ∞−∞=∞\infty-\infty=\infty∞−∞=∞, and Mathlib's EReal uses ⊥+⊤=⊥\bot+\top=\bot⊥+⊤=⊥. No statement of this mission adds infinities of opposite sign; f∗(x)+εf^*(x)+\varepsilonf∗(x)+ε adds a real number.
  • Functions on DDD and on projX(D)\mathrm{proj}_X(D)projX​(D) are functions on subtypes. The graph condition Gr(φ)⊆D\mathrm{Gr}(\varphi)\subseteq DGr(φ)⊆D is part of every selector statement.
  • Universally measurable means NullMeasurableSet E p for every probability measure p.
  • "Analytically measurable" refers to the σ-algebra generated by analytic sets. Replacing it by the power set, dropping the graph condition, or dropping the −1/ε-1/\varepsilon−1/ε case would make the selection theorems a consequence of the axiom of choice. The statements rule all three out.

Not included: Lusin's theorem in Suslin-scheme form (Proposition 7.42, which needs the Suslin operation as a definition), Proposition 7.43 on P(X)P(X)P(X), the integration results of Propositions 7.46 and 7.48, and Lemma 7.30(2)–(4). None is used in the proof of the goal. Contributions welcome: the bridge between IsBorelSpace and Mathlib's StandardBorelSpace, the universal σ-algebra API, and the Jankov–von Neumann theorem itself.

Selected references

  • D. P. Bertsekas and S. E. Shreve, Stochastic Optimal Control: The Discrete-Time Case, Academic Press 1978; Athena Scientific 1996, §7.6–7.7. https://web.mit.edu/dimitrib/www/soc.html
  • D. Blackwell, D. Freedman and M. Orkin, The optimal reward operator in dynamic programming, Annals of Probability 2 (1974) 926–941. https://doi.org/10.1214/aop/1176996558
  • A. S. Kechris, Classical Descriptive Set Theory, Graduate Texts in Mathematics 156, Springer 1995, §18 (Jankov–von Neumann uniformization), §29 (measurability of analytic sets). https://doi.org/10.1007/978-1-4612-4190-4
  • S. E. Shreve and D. P. Bertsekas, Universally measurable policies in dynamic programming, Mathematics of Operations Research 4 (1979) 15–30. https://doi.org/10.1287/moor.4.1.15
8 thms1 active userReviewed
AnalysisDynamic ProgrammingOperations Research+2·Captain: mikedeng1

Stochastic Optimal Control: The Discrete-Time Case V: Semicontinuous Functions — a Borel-Measurable Minimizing Selector for Lower Semicontinuous CostsTextbook

Motivation

Every step of the dynamic programming algorithm on a general state space does three things: it takes a conditional expectation of the cost-to-go under a transition kernel, it minimizes the resulting function of state and control over the control, and, if a policy is to be produced, it picks a control for each state that attains or nearly attains that minimum. On a finite or countable state space all three are harmless. On an uncountable state space each can destroy the measurability needed to take the next expectation: the infimum over an uncountable family of measurable functions need not be measurable, and a minimizer chosen state by state need not be a measurable function of the state, so it does not define a policy at all.

Section 7.5 of Bertsekas and Shreve, Stochastic Optimal Control: The Discrete-Time Case (1978; Athena Scientific reprint 1996), settles the three operations for semicontinuous costs and continuous kernels. The results are the topological half of the book's measurability theory; the descriptive set theory half (lower semianalytic functions and analytically measurable selectors, §7.6–7.7) is a separate mission in this series. The semicontinuous results are what Propositions 8.6–8.7 and Corollaries 9.17.2–9.17.3 of the book use to obtain Borel-measurable optimal policies for finite-horizon and infinite-horizon models with lower semicontinuous costs and compact control sets.

Timeline. The exact selection theorem for lower semicontinuous functions (Proposition 7.33 below) is credited by the book's notes to Dubins and Savage, How to Gamble If You Must (1965). The Hausdorff metric on closed sets goes back to Hausdorff's Set Theory. Measurable selection in the closed-valued setting was later systematized by Kuratowski and Ryll-Nardzewski (1965), whose theorem gives a different route to results of this kind.

Setting

Throughout, R∗=[−∞,+∞]R^*=[-\infty,+\infty]R∗=[−∞,+∞] is the extended real line. A function f:X→R∗f:X\to R^*f:X→R∗ on a metrizable space XXX is lower semicontinuous if every sublevel set {x∣f(x)≤c}\{x\mid f(x)\le c\}{x∣f(x)≤c}, c∈Rc\in\mathbb Rc∈R, is closed, and upper semicontinuous if every superlevel set {x∣f(x)≥c}\{x\mid f(x)\ge c\}{x∣f(x)≥c} is closed (Definition 7.13). C(X)C(X)C(X) is the space of bounded continuous real-valued functions on XXX.

For a separable metrizable space YYY, P(Y)P(Y)P(Y) is the set of Borel probability measures on YYY with the weak topology (convergence of integrals of functions in C(Y)C(Y)C(Y)). A stochastic kernel q(dy∣x)q(dy\mid x)q(dy∣x) on YYY given XXX is a map x↦q(dy∣x)x\mapsto q(dy\mid x)x↦q(dy∣x) from XXX to P(Y)P(Y)P(Y), and it is continuous if this map is continuous (Definition 7.12). The integral of a Borel-measurable f:Y→R∗f:Y\to R^*f:Y→R∗ is ∫f dp=∫f+dp−∫f−dp\int f\,dp=\int f^+dp-\int f^-dp∫fdp=∫f+dp−∫f−dp with the convention −∞+∞=+∞−∞=+∞-\infty+\infty=+\infty-\infty=+\infty−∞+∞=+∞−∞=+∞ (Eq. (43) of Chapter 7).

For a compact metric space YYY, 2Y2^Y2Y is the collection of closed subsets of YYY with the topology of the Hausdorff metric (Appendix C). For D⊆X×YD\subseteq X\times YD⊆X×Y, the section at xxx is Dx={y∣(x,y)∈D}D_x=\{y\mid (x,y)\in D\}Dx​={y∣(x,y)∈D}, the projection is projX(D)={x∣Dx≠∅}\mathrm{proj}_X(D)=\{x\mid D_x\neq\emptyset\}projX​(D)={x∣Dx​=∅}, and a function φ:projX(D)→Y\varphi:\mathrm{proj}_X(D)\to Yφ:projX​(D)→Y has its graph in DDD if (x,φ(x))∈D(x,\varphi(x))\in D(x,φ(x))∈D for every x∈projX(D)x\in\mathrm{proj}_X(D)x∈projX​(D). "Borel-measurable" refers to the Borel σ-algebras of the topologies in question; on projX(D)\mathrm{proj}_X(D)projX​(D) this is the Borel σ-algebra of the subspace topology.

Formalization targets

Goal: Proposition 7.33

Let XXX be metrizable, YYY compact metrizable, D⊆X×YD\subseteq X\times YD⊆X×Y closed, and f:D→R∗f:D\to R^*f:D→R∗ lower semicontinuous. Put

f∗(x)=min⁡y∈Dxf(x,y),x∈projX(D).f^*(x)=\min_{y\in D_x}f(x,y),\qquad x\in\mathrm{proj}_X(D).f∗(x)=y∈Dx​min​f(x,y),x∈projX​(D).

Then projX(D)\mathrm{proj}_X(D)projX​(D) is closed, f∗f^*f∗ is lower semicontinuous, and there is a Borel-measurable φ:projX(D)→Y\varphi:\mathrm{proj}_X(D)\to Yφ:projX​(D)→Y with graph in DDD and

f(x,φ(x))=f∗(x)∀x∈projX(D).f\bigl(x,\varphi(x)\bigr)=f^*(x)\qquad\forall x\in\mathrm{proj}_X(D).f(x,φ(x))=f∗(x)∀x∈projX​(D).

Milestones

  • Proposition 7.32: for f∗(x)=inf⁡y∈Yf(x,y)f^*(x)=\inf_{y\in Y}f(x,y)f∗(x)=infy∈Y​f(x,y), lower semicontinuity of fff and compactness of YYY give lower semicontinuity of f∗f^*f∗ and attainment; upper semicontinuity of fff gives upper semicontinuity of f∗f^*f∗.
  • Lemma 7.18: there is a Borel-measurable σ:2Y−{∅}→Y\sigma:2^Y-\{\emptyset\}\to Yσ:2Y−{∅}→Y with σ(A)∈A\sigma(A)\in Aσ(A)∈A.
  • Lemma 7.20: for lower semicontinuous fff on a nonempty compact YYY, the argmin map x↦{y∣f(x,y)≤f∗(x)}x\mapsto\{y\mid f(x,y)\le f^*(x)\}x↦{y∣f(x,y)≤f∗(x)} is Borel-measurable into 2Y2^Y2Y.
  • Lemma 7.14: fff is lower semicontinuous and bounded below iff fn↑ff_n\uparrow ffn​↑f for some fn∈C(X)f_n\in C(X)fn​∈C(X) (and dually).
  • Proposition 7.30: x↦∫f(x,y) q(dy∣x)x\mapsto\int f(x,y)\,q(dy\mid x)x↦∫f(x,y)q(dy∣x) is continuous for f∈C(X×Y)f\in C(X\times Y)f∈C(X×Y) and continuous qqq.
  • Proposition 7.31: the same map is lower (upper) semicontinuous and bounded below (above) when fff is.
  • Lemma 7.21: an open G⊆X×YG\subseteq X\times YG⊆X×Y, YYY separable, has open projection and a Borel-measurable selector with graph in GGG.
  • Proposition 7.34: for open DDD and upper semicontinuous fff, projX(D)\mathrm{proj}_X(D)projX​(D) is open, f∗=inf⁡Dxff^*=\inf_{D_x}ff∗=infDx​​f is upper semicontinuous, and for each ε>0\varepsilon>0ε>0 there is a Borel-measurable φε\varphi_\varepsilonφε​ with graph in DDD and
f(x,φε(x))≤{f∗(x)+εif f∗(x)>−∞,−1/εif f∗(x)=−∞.f\bigl(x,\varphi_\varepsilon(x)\bigr)\le\begin{cases}f^*(x)+\varepsilon&\text{if }f^*(x)>-\infty,\\-1/\varepsilon&\text{if }f^*(x)=-\infty.\end{cases}f(x,φε​(x))≤{f∗(x)+ε−1/ε​if f∗(x)>−∞,if f∗(x)=−∞.​

Significance

The results. Propositions 7.31–7.33 are the closure properties that make the dynamic programming recursion stay inside the class of lower semicontinuous functions bounded below: the expectation step preserves the class (7.31), the minimization step preserves it (7.32, 7.33), and the minimization admits a Borel-measurable exact minimizer (7.33). This is why, in semicontinuous models, the optimal cost functions are lower semicontinuous and optimal policies can be taken Borel-measurable and nonrandomized. Proposition 7.34 gives the weaker, ε\varepsilonε-optimal counterpart for upper semicontinuous costs, where the infimum need not be attained.

Formalizing them. All of these results are proved in the book; none is open. As far as is known, none has a machine-checked proof: Mathlib has semicontinuity, the Hausdorff extended metric on closed and on nonempty compact sets, and the weak topology on probability measures, but no theorem combining them into a measurable selection result of this kind. A formal development would supply measurable selectors for semicontinuous minimization in Lean and the Borel-measurability of set-valued maps into the hyperspace of closed sets, both reusable well beyond dynamic programming.

Difficulty

The obvious attempt at the goal is to pick, for each xxx, some minimizer yyy of f(x,⋅)f(x,\cdot)f(x,⋅) over the compact section DxD_xDx​. The minimizer exists by compactness and lower semicontinuity, but the choice is made pointwise and gives no control on measurability: a minimizer chosen by the axiom of choice need not be Borel-measurable. The argmin sets F∗(x)F^*(x)F∗(x) vary with xxx only semicontinuously: they can jump from a single point to a large set, so a continuous selection generally does not exist, and continuity arguments cannot replace measurability. Lemma 7.18 isolates the hardest part: a choice of a point of each nonempty closed set that is measurable as a function of the set itself.

A second difficulty is bookkeeping at infinity. Values ±∞\pm\infty±∞ are allowed throughout, so sublevel sets, minima, integrals and ε\varepsilonε-bounds must all be handled in R∗R^*R∗; the integral in Proposition 7.31 uses the convention ∞−∞=+∞\infty-\infty=+\infty∞−∞=+∞, which is not Mathlib's.

Formalization scope

  • Extended reals. Values are in EReal. The only place where values of opposite infinite sign are combined is the integral, which is the published definition DupacovaWets.Consistency.expect (reused, not restated): ∫f+−∫f−\int f^+-\int f^-∫f+−∫f− with an explicit case returning +∞+\infty+∞ when ∫f+=∞\int f^+=\infty∫f+=∞, exactly the book's convention (42). The ε\varepsilonε-bound of Proposition 7.34 adds a real ε\varepsilonε to a value different from −∞-\infty−∞, which is safe in EReal.
  • Semicontinuity is Mathlib's LowerSemicontinuous/UpperSemicontinuous, equivalent to Definition 7.13 for EReal-valued functions. Lemma 7.13 of the book (the sequential characterization) is Mathlib's lowerSemicontinuous_iff_le_liminf together with first countability of metrizable spaces, and is not restated here.
  • Functions on DDD. Functions "on DDD" are functions on X×YX\times YX×Y with LowerSemicontinuousOn f D (resp. UpperSemicontinuousOn); values off DDD play no role. projX(D)\mathrm{proj}_X(D)projX​(D) is Prod.fst '' D, selectors are functions on that subtype, and its σ-algebra is the Borel σ-algebra of the subspace topology.
  • Hyperspace. 2Y2^Y2Y is Closeds Y, and 2Y−{∅}2^Y-\{\emptyset\}2Y−{∅} for compact YYY is NonemptyCompacts Y, each with the Hausdorff extended metric and the Borel σ-algebra of its topology. This topology agrees with the book's (the exponential topology of Appendix C, independent of the metric).
  • Boundedness. "Bounded below/above" is by a real constant. BddBelow in EReal would be vacuous and is not used.
  • Edge cases. Proposition 7.32(a)'s attainment clause is stated for nonempty YYY, since for Y=∅Y=\emptysetY=∅ the infimum is +∞+\infty+∞ and nothing attains it.
  • Argmin minimum. Lemma 7.20 assumes nonempty YYY because its defining formula uses a minimum; for empty YYY there is no minimizer.
  • Ruling out trivial readings. The graph condition (x,φ(x))∈D(x,\varphi(x))\in D(x,φ(x))∈D is part of every selection statement; without it the goal would follow from the unconstrained case. The selector must be Borel-measurable on projX(D)\mathrm{proj}_X(D)projX​(D) and must attain the minimum exactly, not up to ε\varepsilonε.

A complete development needs the Borel structure of the hyperspace (measurability of maps into Closeds Y from upper semicontinuity in the sense of Kuratowski, Proposition C.4 of the book), the construction of a measurable choice function on NonemptyCompacts Y, and approximation of semicontinuous functions by monotone sequences in C(X)C(X)C(X). Each of these is reusable on its own; proofs of individual milestones by any route are welcome.

Selected references

  • D. P. Bertsekas and S. E. Shreve, Stochastic Optimal Control: The Discrete-Time Case, Academic Press, 1978; Athena Scientific reprint, 1996, Section 7.5 and Appendix C. https://web.mit.edu/dimitrib/www/soc.html
  • L. E. Dubins and L. J. Savage, How to Gamble If You Must: Inequalities for Stochastic Processes, McGraw-Hill, 1965.
  • K. Kuratowski and C. Ryll-Nardzewski, "A general theorem on selectors," Bull. Acad. Polon. Sci. 13 (1965), 397–403.
  • F. Hausdorff, Set Theory, Chelsea, New York, 1957.
10 thms1 active userReviewed
Machine LearningOptimizationStatistics·Captain: mikedeng1

Variance-based Regularization with Convex Objectives II: A Covering-Number Certificate and Oracle Inequality for the Robust MinimizerResearch Paper

Motivation

Empirical risk minimization (ERM) chooses, from a class F\mathcal FF of loss functions, the one with the smallest average loss on a sample X1,…,XnX_1,\dots,X_nX1​,…,Xn​. Its standard guarantees bound the excess population risk by a term of order 1/n1/\sqrt n1/n​, whatever the variance of the losses. When good functions in F\mathcal FF have small variance, a better trade-off is available in principle: minimize the empirical risk plus a standard-deviation penalty 2ρ VarP^n(f)/n\sqrt{2\rho\,\mathrm{Var}_{\widehat P_n}(f)/n}2ρVarPn​​(f)/n​. Maurer and Pontil (COLT 2009) showed that this sample variance penalization enjoys faster rates, but the penalized objective is non-convex even for convex losses, so it cannot be minimized efficiently in general.

Duchi and Namkoong (arXiv:1610.02581v3, 2017; NIPS 2017) replace the variance penalty by a distributionally robust objective: the worst-case average loss over all reweightings of the sample within a χ2\chi^2χ2-divergence ball of radius ρ/n\rho/nρ/n. This objective is convex whenever the loss is convex, and it equals the empirical risk plus the standard-deviation penalty up to an error of order 1/n1/n1/n. Theorem 3 of the paper turns this into a guarantee for the minimizer of the robust objective, using covering numbers of the class. This mission formalizes Theorem 3 and the lemmas its proof rests on.

Setting

Let X\mathcal XX be a measurable space, PPP a probability measure on it, and X1,…,XnX_1,\dots,X_nX1​,…,Xn​ (n≥1n\ge1n≥1) an i.i.d. sample from PPP with empirical distribution P^n\widehat P_nPn​. Let F\mathcal FF be a nonempty class of measurable functions f:X→[M0,M1]f:\mathcal X\to[M_0,M_1]f:X→[M0​,M1​], and set M=M1−M0M = M_1-M_0M=M1​−M0​. Write E[f]=∫f dP\mathbb E[f]=\int f\,dPE[f]=∫fdP, Var(f)\mathrm{Var}(f)Var(f) for the variance of f(X)f(X)f(X), and

EP^n[f]=1n∑i=1nf(Xi),VarP^n(f)=1n∑i=1nf(Xi)2−(EP^n[f])2.\mathbb E_{\widehat P_n}[f] = \frac1n\sum_{i=1}^n f(X_i),\qquad \mathrm{Var}_{\widehat P_n}(f) = \frac1n\sum_{i=1}^n f(X_i)^2 - \big(\mathbb E_{\widehat P_n}[f]\big)^2 .EPn​​[f]=n1​i=1∑n​f(Xi​),VarPn​​(f)=n1​i=1∑n​f(Xi​)2−(EPn​​[f])2.

For ρ≥0\rho\ge0ρ≥0, the χ2\chi^2χ2 ball Pn\mathcal P_nPn​ is the set of weight vectors p∈Rnp\in\mathbb R^np∈Rn with pi≥0p_i\ge0pi​≥0, ∑ipi=1\sum_i p_i = 1∑i​pi​=1 and 12∑i(npi−1)2≤ρ\frac12\sum_i (np_i-1)^2\le\rho21​∑i​(npi​−1)2≤ρ: the distributions PPP on the sample with Dϕ(P∥P^n)≤ρ/nD_\phi(P\|\widehat P_n)\le\rho/nDϕ​(P∥Pn​)≤ρ/n for ϕ(t)=12(t−1)2\phi(t)=\frac12(t-1)^2ϕ(t)=21​(t−1)2. The robust risk of fff is

Rn(f)=sup⁡P: Dϕ(P∥P^n)≤ρ/nEP[f(X)]=max⁡p∈Pn∑i=1npif(Xi),R_n(f) = \sup_{P:\,D_\phi(P\|\widehat P_n)\le \rho/n}\mathbb E_P[f(X)] = \max_{p\in\mathcal P_n}\sum_{i=1}^n p_i f(X_i),Rn​(f)=P:Dϕ​(P∥Pn​)≤ρ/nsup​EP​[f(X)]=p∈Pn​max​i=1∑n​pi​f(Xi​),

and a robust minimizer is any f^∈argmin⁡f∈FRn(f)\widehat f\in\operatorname{argmin}_{f\in\mathcal F} R_n(f)f​∈argminf∈F​Rn​(f).

Complexity is measured by empirical ℓ∞\ell_\inftyℓ∞​ covering numbers. For V⊂RmV\subset\mathbb R^mV⊂Rm, N(V,ϵ,∥⋅∥∞)N(V,\epsilon,\|\cdot\|_\infty)N(V,ϵ,∥⋅∥∞​) is the least number of points v1,…,vN∈Vv_1,\dots,v_N\in Vv1​,…,vN​∈V such that every v∈Vv\in Vv∈V lies within sup-distance ϵ\epsilonϵ of some viv_ivi​. For x∈Xmx\in\mathcal X^mx∈Xm let F(x)={(f(x1),…,f(xm)):f∈F}\mathcal F(x)=\{(f(x_1),\dots,f(x_m)) : f\in\mathcal F\}F(x)={(f(x1​),…,f(xm​)):f∈F}, and

N∞(F,ϵ,m)=sup⁡x∈XmN(F(x),ϵ,∥⋅∥∞)∈N∪{∞}.N_\infty(\mathcal F,\epsilon,m) = \sup_{x\in\mathcal X^m} N\big(\mathcal F(x),\epsilon,\|\cdot\|_\infty\big)\in\mathbb N\cup\{\infty\}.N∞​(F,ϵ,m)=x∈Xmsup​N(F(x),ϵ,∥⋅∥∞​)∈N∪{∞}.

Formalization targets

Goal: the oracle inequality (16)

Let n≥8M2/tn\ge 8M^2/tn≥8M2/t, t≥log⁡12t\ge\log 12t≥log12, ϵ>0\epsilon>0ϵ>0 and ρ≥9t\rho\ge 9tρ≥9t. With probability at least 1−2(3N∞(F,ϵ,2n)+1)e−t1-2(3N_\infty(\mathcal F,\epsilon,2n)+1)e^{-t}1−2(3N∞​(F,ϵ,2n)+1)e−t, every robust minimizer f^\widehat ff​ satisfies

E[f^(X)]≤inf⁡f∈F{E[f]+22ρnVar(f)}+19Mρ3n+(2+42tn)ϵ.\mathbb E[\widehat f(X)] \le \inf_{f\in\mathcal F}\left\{\mathbb E[f] + 2\sqrt{\frac{2\rho}{n}\mathrm{Var}(f)}\right\} + \frac{19M\rho}{3n} + \left(2+4\sqrt{\frac{2t}{n}}\right)\epsilon .E[f​(X)]≤f∈Finf​{E[f]+2n2ρ​Var(f)​}+3n19Mρ​+(2+4n2t​​)ϵ.

The certificate (15)

Under the same hypotheses and with the same probability, simultaneously for all f∈Ff\in\mathcal Ff∈F,

E[f(X)]≤Rn(f)+113Mρn+(2+42tn)ϵ.\mathbb E[f(X)] \le R_n(f) + \frac{11}{3}\frac{M\rho}{n} + \left(2+4\sqrt{\frac{2t}{n}}\right)\epsilon .E[f(X)]≤Rn​(f)+311​nMρ​+(2+4n2t​​)ϵ.

Supporting results (milestones)

  1. Theorem 1, inequality (10): for every vector z∈[M0,M1]nz\in[M_0,M_1]^nz∈[M0​,M1​]n, the robust mean minus the sample mean lies between (2ρsn2/n−2Mρ/n)+\big(\sqrt{2\rho s_n^2/n}-2M\rho/n\big)_+(2ρsn2​/n​−2Mρ/n)+​ and 2ρsn2/n\sqrt{2\rho s_n^2/n}2ρsn2​/n​.
  2. Lemma C.1: a uniform empirical Bernstein bound over F\mathcal FF with probability 1−6N∞(F,ϵ,2n)e−t1-6N_\infty(\mathcal F,\epsilon,2n)e^{-t}1−6N∞​(F,ϵ,2n)e−t.
  3. Lemma A.1, first bound: P(sn≥Esn2+t)≤exp⁡(−nt2/(2M2))\mathbb P(s_n\ge\sqrt{\mathbb E s_n^2}+t)\le\exp(-nt^2/(2M^2))P(sn​≥Esn2​​+t)≤exp(−nt2/(2M2)).
  4. Bernstein's inequality for one fixed fff, as displayed in the proof (p. 38).
  5. The certificate (15).

Significance

Inequality (15) says the robust risk is a uniform upper confidence bound on the population risk, with an O(1/n)O(1/n)O(1/n) slack instead of the O(1/n)O(1/\sqrt n)O(1/n​) slack of the empirical risk. Inequality (16) says the robust minimizer competes with the best variance-penalized population risk in the class. When some f∈Ff\in\mathcal Ff∈F has small risk and small variance, the excess risk of f^\widehat ff​ is of order 1/n1/n1/n up to the covering term, a rate ERM does not achieve in general (§3.3 of the paper gives an example). For a parametric class with N∞(F,ϵ,2n)N_\infty(\mathcal F,\epsilon,2n)N∞​(F,ϵ,2n) polynomial in 1/ϵ1/\epsilon1/ϵ, choosing ϵ=M/n\epsilon=M/nϵ=M/n gives Corollaries 3.1 and 3.2 of the paper.

The results are proved in the paper; none of them has a machine-checked proof that we know of. The mission's output is a formal proof of Theorem 3 and its ingredients: a deterministic analysis of the χ2\chi^2χ2-constrained linear program (Theorem 1 (10)), a covering-number empirical Bernstein inequality (Lemma C.1, from Maurer and Pontil), concentration of the sample standard deviation (Lemma A.1), and the scalar Bernstein inequality in the form used. Each of these is reusable outside distributionally robust optimization.

Difficulty

The deterministic part, (10), is a short analysis of a quadratically constrained linear program. The main obstacle is Lemma C.1. A union bound over a cover of F\mathcal FF fails directly: the cover depends on the sample, and a population-level cover of F\mathcal FF need not be finite. The standard route goes through a ghost sample of size nnn (hence covering at 2n2n2n points), a symmetrization that must preserve the sample variance rather than only the mean, and a concentration bound for the sample variance itself. Lemma A.1 needs concentration of sns_nsn​, a non-linear and non-smooth function of the sample, at the sub-Gaussian rate M/nM/\sqrt nM/n​. Finally, the oracle inequality (16) holds for an infimum over the whole class, while the concentration step for the comparison function is only proved for one fixed fff at a time.

Formalization scope

The sample is the coordinate process of the product measure P⊗nP^{\otimes n}P⊗n on Xn\mathcal X^nXn (Measure.pi). Each probability statement bounds the probability of the bad event, the set of samples where the inequality fails for some fff (or some minimizer). This set need not be measurable, and its measure is then the outer measure, as is standard in empirical-process theory. Probability bounds are computed in [0,∞][0,\infty][0,∞], and the covering number is valued in N∪{∞}\mathbb N\cup\{\infty\}N∪{∞}, so an infinite covering number makes the bound trivial rather than collapsing to zero. Covering numbers are internal (centres in F(x)\mathcal F(x)F(x)) and use closed sup-norm balls, as on p. 9; this is Mathlib's Metric.coveringNumber. The empirical variance is normalized by 1/n1/n1/n. The χ2\chi^2χ2 ball is encoded as weight vectors on the sample points; with tied sample values this gives the same supremum as the paper's distributions on the sample. Statement (16) is formalized for every minimizer of the robust risk, and the event is empty if no minimizer exists. The infimum ranges over the nonempty class F\mathcal FF, on which every term is at least M0M_0M0​. Population moments are those of bounded measurable functions, hence finite.

Deviations from the printed text:

  • Lemma A.1 is stated only for its first (upper-tail) bound. The paper derives the second bound from Lemma A.4, which is false as printed; the second bound is not stated. M>0M>0M>0 is assumed because M2M^2M2 is a denominator.
  • Lemma C.1 is the paper's restatement of Maurer and Pontil's Theorem 6, with a general radius ϵ\epsilonϵ. It is formalized as printed, with the implicit assumption ϵ>0\epsilon>0ϵ>0 made explicit.
  • n≥1n\ge1n≥1 is assumed throughout. The hypothesis n≥8M2/tn\ge 8M^2/tn≥8M2/t is kept as printed.

A trivializing formalization is ruled out: the bound is not taken over all functions, a probability bound is not formed from the real part of an infinite covering number, and the minimizer is not a hypothesis that can fail to exist for the given sample.

Needed infrastructure: product-measure concentration (Bernstein, and a bounded-difference or convex-Lipschitz inequality for sns_nsn​), symmetrization with a ghost sample, and finite union bounds over a cover. Contributions are welcome on any milestone, in particular a general covering-number empirical Bernstein inequality, which is reusable on its own.

Selected references

  • J. C. Duchi and H. Namkoong, Variance-based regularization with convex objectives, arXiv:1610.02581v3, 2017 (NIPS 2017; JMLR 20, 2019). https://arxiv.org/abs/1610.02581
  • A. Maurer and M. Pontil, Empirical Bernstein bounds and sample variance penalization, COLT 2009. https://arxiv.org/abs/0907.3740
  • S. Boucheron, G. Lugosi and P. Massart, Concentration Inequalities: A Nonasymptotic Theory of Independence, Oxford University Press, 2013. https://doi.org/10.1093/acprof:oso/9780199535255.001.0001
  • A. W. van der Vaart and J. A. Wellner, Weak Convergence and Empirical Processes, Springer, 1996. https://doi.org/10.1007/978-1-4757-2545-2
11 thms1 active userReviewed
Markov ChainOperations ResearchStochastic Systems·Captain: mikedeng1

Optimization of Multiclass Queueing Networks: Polyhedral and Nonlinear Characterizations of Achievable Performance I: Quadratic Potential Functions Bound Mean Response Times in Open NetworksResearch Paper

Motivation

Scheduling in a multiclass queueing network asks which waiting job a server should work on next when jobs of several types share stations and revisit them along fixed routes. Such networks model semiconductor wafer fabs, job shops and communication switches. Optimal policies are rarely computable: the state space is countably infinite, and even deciding properties of optimal policies is hard (Papadimitriou and Tsitsiklis 1999). A practical substitute is the achievable region approach: describe, by constraints that every policy must satisfy, a set containing all performance vectors any policy can achieve, then optimize a linear cost over that set to get a lower bound on the optimal cost.

Bertsimas, Paschalidis and Tsitsiklis (MIT Sloan working paper 1992; Ann. Appl. Probab. 1994) gave a general method for producing such constraints for open networks, by computing the steady-state drift of quadratic potential functions. This mission formalizes their first-order bounds (Section 4).

Timeline:

  • 1980–1988: Coffman and Mitrani, then Federgruen and Groenevelt — the achievable performance vectors of a single-station multiclass queue form a polytope described by conservation laws.
  • Early 1990s: Kumar (reference [Kuma] of the paper), using a potential-function argument he attributes to Meyn, derives a single lower bound on the mean number in system for re-entrant lines with deterministic routing (described on p. 16 of the paper).
  • 1992–1994: Bertsimas, Paschalidis and Tsitsiklis — parametric families of linear bounds for general open networks with Markovian routing (Theorem 4.1), and the nonparametric polyhedron (Theorems 4.2–4.4), shown to be at least as tight.

Setting

A network has NNN single-server stations and RRR job classes. Class rrr is served at station σ(r)\sigma(r)σ(r), and CiC_iCi​ is the set of classes served at station iii. Class-rrr jobs arrive from outside as a Poisson stream of rate λ0r\lambda_{0r}λ0r​, service times are exponential with rate μr\mu_rμr​, and after service a class-rrr job becomes a class-sss job with probability prsp_{rs}prs​ or leaves with probability pr0=1−∑sprsp_{r0}=1-\sum_s p_{rs}pr0​=1−∑s​prs​. The traffic equations

λr=λ0r+∑r′λr′pr′r(15)\lambda_r=\lambda_{0r}+\sum_{r'}\lambda_{r'}p_{r'r}\qquad(15)λr​=λ0r​+r′∑​λr′​pr′r​(15)

have a unique solution λ\lambdaλ (the network is open), and ∑r∈Ciλr/μr<1\sum_{r\in C_i}\lambda_r/\mu_r<1∑r∈Ci​​λr​/μr​<1 at every station.

The state n⃗=(n1,…,nR)\vec n=(n_1,\dots,n_R)n=(n1​,…,nR​) counts the jobs of each class. A Markovian policy decides from the current state which classes are in service, at most one per station and only classes with jobs present; idling is allowed. Write BrB_rBr​ for the event that station σ(r)\sigma(r)σ(r) serves class rrr, and B0iB_{0i}B0i​ for the event that station iii is idle. Under such a policy n⃗(t)\vec n(t)n(t) is a continuous-time Markov chain. Assumption A requires that it has a unique invariant distribution π\piπ and that Eπ[nr2]<∞E_\pi[n_r^2]<\inftyEπ​[nr2​]<∞ for all rrr. Let nˉr=Eπ[nr]\bar n_r=E_\pi[n_r]nˉr​=Eπ​[nr​], which equals λrxr\lambda_rx_rλr​xr​ with xrx_rxr​ the mean response time of class rrr (Little's law), and define

Irr′=Eπ[1{Br}nr′],Nir′=Eπ[1{B0i}nr′].I_{rr'}=E_\pi[1\{B_r\}n_{r'}],\qquad N_{ir'}=E_\pi[1\{B_{0i}\}n_{r'}].Irr′​=Eπ​[1{Br​}nr′​],Nir′​=Eπ​[1{B0i​}nr′​].

For a set SSS of classes, f-parameters are reals f(r)≥0f(r)\ge 0f(r)≥0 for r∈Sr\in Sr∈S such that μr[∑r′∈Sprr′(f(r)−f(r′))+∑r′∉Sprr′f(r)]\mu_r\big[\sum_{r'\in S}p_{rr'}(f(r)-f(r'))+\sum_{r'\notin S}p_{rr'}f(r)\big]μr​[∑r′∈S​prr′​(f(r)−f(r′))+∑r′∈/S​prr′​f(r)] is nonnegative and the same for all r∈Ci∩Sr\in C_i\cap Sr∈Ci​∩S; that common value is fif_ifi​, and fi=0f_i=0fi​=0 when Ci∩S=∅C_i\cap S=\emptysetCi​∩S=∅ (restriction (17)). The sums over r′∉Sr'\notin Sr′∈/S include the exit r′=0r'=0r′=0.

Formalization targets

Goal: Theorem 4.1

For every policy satisfying Assumption A, every SSS and every f-parameters satisfying (17),

∑r∈Sλrf(r)xr ≥ N′(S)D′(S),\sum_{r\in S}\lambda_rf(r)x_r\ \ge\ \frac{N'(S)}{D'(S)},r∈S∑​λr​f(r)xr​ ≥ D′(S)N′(S)​,

where

N′(S)=∑r∈Sλ0rf2(r)+∑r∉Sλr∑r′∈Sprr′f2(r′)+∑r∈Sλr[∑r′∈Sprr′(f(r)−f(r′))2+∑r′∉Sprr′f2(r)],N'(S)=\sum_{r\in S}\lambda_{0r}f^2(r)+\sum_{r\notin S}\lambda_r\sum_{r'\in S}p_{rr'}f^2(r')+\sum_{r\in S}\lambda_r\Big[\sum_{r'\in S}p_{rr'}(f(r)-f(r'))^2+\sum_{r'\notin S}p_{rr'}f^2(r)\Big],N′(S)=r∈S∑​λ0r​f2(r)+r∈/S∑​λr​r′∈S∑​prr′​f2(r′)+r∈S∑​λr​[r′∈S∑​prr′​(f(r)−f(r′))2+r′∈/S∑​prr′​f2(r)], D′(S)=2[∑i=1Nfi−∑r∈Sλ0rf(r)].D'(S)=2\Big[\sum_{i=1}^Nf_i-\sum_{r\in S}\lambda_{0r}f(r)\Big].D′(S)=2[i=1∑N​fi​−r∈S∑​λ0r​f(r)].

The formal goal is the product form N′(S)≤D′(S)∑r∈Sf(r)nˉrN'(S)\le D'(S)\sum_{r\in S}f(r)\bar n_rN′(S)≤D′(S)∑r∈S​f(r)nˉr​.

Milestones

  1. The utilization identity Eπ[1{Br}]=λr/μrE_\pi[1\{B_r\}]=\lambda_r/\mu_rEπ​[1{Br​}]=λr​/μr​ (pp. 16 and 19).
  2. Theorem 4.2: the linear equalities (24), (25) between nˉr\bar n_rnˉr​ and Irr′I_{rr'}Irr′​.
  3. Theorem 4.3: ∑r∈CiIrr′+Nir′=nˉr′\sum_{r\in C_i}I_{rr'}+N_{ir'}=\bar n_{r'}∑r∈Ci​​Irr′​+Nir′​=nˉr′​ (28).
  4. Theorem 4.4: any nonnegative (x,I,N)(x,I,N)(x,I,N) satisfying (24), (25), (28), with nˉr=λrxr\bar n_r=\lambda_rx_rnˉr​=λr​xr​ in those equalities, satisfies every inequality of Theorem 4.1. This statement is deterministic.

Significance

Theorem 4.1 gives, for each choice of SSS and fff, a linear inequality on mean response times valid for all admissible policies. Minimizing a linear holding cost ∑rcrxr\sum_r c_rx_r∑r​cr​xr​ subject to these inequalities is a linear program whose value bounds the optimal scheduling cost from below; the paper reports numerical values of such bounds in its Section 9. Theorems 4.2–4.4 show that a polynomial-size polyhedron in the variables (nˉ,I,N)(\bar n,I,N)(nˉ,I,N) implies all of these inequalities at once, so the parametric search over fff is unnecessary.

The results are proved in the paper. As far as is known, none of them has a machine-checked proof. Formalizing them requires a Lean treatment of invariant distributions of controlled countable-state Markov chains with unbounded test functions, which is currently absent from Mathlib, and then the algebra of the drift identities. The definitions here (network data, Markovian sequencing policies, the generator, Assumption A) are the substrate that the paper's later results on routing, closed networks and higher-order bounds would reuse.

Difficulty

Every statement except Theorem 4.4 rests on taking expectations of the generator applied to unbounded functions (nrn_rnr​, nrnr′n_rn_{r'}nr​nr′​) under the invariant distribution. The invariance condition is stated only for indicators of single states; extending ∑nπ(n)(Gg)(n)=0\sum_n\pi(n)(\mathcal Gg)(n)=0∑n​π(n)(Gg)(n)=0 to quadratic ggg needs an interchange of summations justified by the second-moment condition of Assumption A. The utilization identity additionally needs uniqueness of the traffic solution to identify μrEπ[1{Br}]\mu_rE_\pi[1\{B_r\}]μr​Eπ​[1{Br​}] with λr\lambda_rλr​. Theorem 4.1 then needs the sign bookkeeping that turns an identity into an inequality: the terms dropped are nonnegative only because f≥0f\ge0f≥0 on SSS, fi≥0f_i\ge0fi​≥0 and at most one class per station is in service.

Formalization scope

Classes are Fin R, stations Fin N, states Fin R → ℕ, all rates and probabilities real. A policy is a Bool-valued function of the state with the two admissibility constraints; work conservation is not assumed. Invariance is global balance of the generator on the countable state space; expectations are tsums. The uniformized chain and the epochs τk\tau_kτk​ of the paper are not built: the paper notes that its expectations at τk\tau_kτk​ are expectations under the invariant distribution of n⃗(t)\vec n(t)n(t).

Conventions fixed in Lean:

  • λrxr\lambda_rx_rλr​xr​ appears only as the mean number in system nˉr\bar n_rnˉr​ (Little's law, used by the paper on pp. 11 and 20); response times are not formalized.
  • Sums over r′∉Sr'\notin Sr′∈/S include the exit r′=0r'=0r′=0 (p. 15).
  • f-parameters are nonnegative on SSS (p. 9).
  • The network is open: (15) has a unique solution, and λ\lambdaλ is an input constrained by (15), never defined from the policy.
  • (18) is stated multiplied by D′(S)D'(S)D′(S), which avoids Lean's x/0=0x/0=0x/0=0 and is (18) whenever D′(S)>0D'(S)>0D′(S)>0.

A quotient-form statement of (18) would be trivially true when D′(S)=0D'(S)=0D′(S)=0, and defining λr\lambda_rλr​ as μrEπ[1{Br}]\mu_rE_\pi[1\{B_r\}]μr​Eπ​[1{Br​}] would make the utilization identity hold by definition; both are excluded.

Welcome contributions: a general lemma extending global balance to test functions of polynomial growth under moment conditions; proofs of the drift identities; the deterministic Theorem 4.4.

Selected references

  • D. Bertsimas, I. Ch. Paschalidis, J. N. Tsitsiklis, Optimization of Multiclass Queueing Networks: Polyhedral and Nonlinear Characterizations of Achievable Performance, MIT Sloan WP #3509-92-MSA, 1992; Ann. Appl. Probab. 4(1), 1994. https://doi.org/10.1214/aoap/1177005200
  • C. H. Papadimitriou, J. N. Tsitsiklis, The complexity of optimal queuing network control, Math. Oper. Res. 24(2), 1999. https://doi.org/10.1287/moor.24.2.293
8 thms1 active userReviewed
CombinatoricsGraph TheoryOperations Research+1·Captain: mikedeng1

Secretary Problems: Weights and Discounts 5: A 3e-Competitive Algorithm for the Graphic Matroid Secretary ProblemResearch Paper

Motivation

In the secretary problem, nnn items with nonnegative values arrive one at a time in a uniformly random order, and an online algorithm must decide on each arrival, irrevocably, whether to keep it. The classical version keeps one item; the rule that observes a 1/e1/e1/e fraction of the arrivals and then takes the first item better than everything seen picks the best item with probability at least 1/e1/e1/e (Ferguson 1989).

Babaioff, Immorlica and Kleinberg (SODA 2007; journal version J. ACM 2018) introduced the matroid secretary problem: the kept set must be independent in a known matroid. It models online auctions in which the feasible sets of winners have matroid structure, for example hiring along the edges of a network without closing a cycle. They gave a 161616-competitive algorithm when the matroid is graphic, i.e. the items are the edges of a graph and a set is feasible when it contains no cycle.

Timeline for graphic matroids:

  • 2007, Babaioff–Immorlica–Kleinberg: 161616-competitive.
  • 2009, Babaioff–Dinitz–Gupta–Immorlica–Talwar (SODA 2009, Theorem 1.5): 3e≈8.153e\approx 8.153e≈8.15-competitive, through a random reduction to partition matroids. This mission formalizes that result.
  • 2009, Korula–Pál (ICALP 2009): 2e2e2e-competitive, by a different reduction.

Setting

Let G=(V,E)G=(V,E)G=(V,E) be a finite simple graph. Each edge eee has a value v(e)≥0v(e)\ge 0v(e)≥0. A set S⊆ES\subseteq ES⊆E is independent in the graphic matroid of GGG if the graph (V,S)(V,S)(V,S) has no cycle. The offline optimum is

OPT(G,v)=max⁡{∑e∈Sv(e):S⊆E acyclic}.\mathrm{OPT}(G,v)=\max\Big\{\sum_{e\in S}v(e): S\subseteq E\ \text{acyclic}\Big\}.OPT(G,v)=max{e∈S∑​v(e):S⊆E acyclic}.

The edges arrive in a uniformly random order. An algorithm sees each edge and its value on arrival and decides at once whether to select it. The selected set must be acyclic. The algorithm is α\alphaα-competitive if OPT(G,v)≤α⋅E[value of the selected set]\mathrm{OPT}(G,v)\le\alpha\cdot\mathbb E[\text{value of the selected set}]OPT(G,v)≤α⋅E[value of the selected set] for every GGG and every v≥0v\ge 0v≥0.

A partition matroid on a subset U′⊆EU'\subseteq EU′⊆E is given by a family PPP of nonempty, pairwise disjoint parts with union U′U'U′: a set is independent when it lies in U′U'U′ and meets each part at most once. Its max-weight base has value val(P,v)=∑p∈Pmax⁡e∈pv(e)\mathrm{val}(P,v)=\sum_{p\in P}\max_{e\in p}v(e)val(P,v)=∑p∈P​maxe∈p​v(e).

Definition 5.1. A random partition μ\muμ (a probability distribution on such families, chosen from GGG alone) is an α\alphaα-partition scheme if every partition in its support has only acyclic independent sets, and for every v≥0v\ge 0v≥0,

OPT(G,v)≤α⋅EP∼μ[val(P,v)].\mathrm{OPT}(G,v)\le \alpha\cdot\mathbb E_{P\sim\mu}[\mathrm{val}(P,v)].OPT(G,v)≤α⋅EP∼μ​[val(P,v)].

The random partition of Lemma 5.3. Pick an edge {u,w}\{u,w\}{u,w} uniformly at random. With probability 12\tfrac1221​ colour uuu red and www blue, otherwise the reverse. Colour every other vertex red or blue independently with probability 12\tfrac1221​. Each red vertex xxx gets a part: the red-blue edges at xxx. Then repeat on the edges with both endpoints blue, with fresh randomness.

The algorithm. Draw the partition, let the edges arrive, and on each part run the classical secretary rule on that part's arrivals. Output all selected edges.

Formalization targets

Goal: Theorem 1.5

For every finite simple graph GGG and every v≥0v\ge 0v≥0:

  1. every possible output of the algorithm is an acyclic set of edges of GGG;
OPT(G,v)≤3e⋅E[ALG].\mathrm{OPT}(G,v)\le 3e\cdot\mathbb E[\mathrm{ALG}].OPT(G,v)≤3e⋅E[ALG].

Part 1 is needed for the statement to have content: an algorithm that selects every edge would otherwise satisfy part 2.

Milestones

  • Section 2, p. 4. On m≥1m\ge1m≥1 arrivals, the classical rule selects the maximum with probability at least 1/e1/e1/e.
  • Theorem 5.4, first clause. For a fixed partition PPP, the per-part rule outputs a set independent in the partition matroid, and val(P,v)≤e⋅Eπ[ALG]\mathrm{val}(P,v)\le e\cdot\mathbb E_\pi[\mathrm{ALG}]val(P,v)≤e⋅Eπ​[ALG].
  • Lemma 5.3, independence. Every partition the random construction can produce is a partition matroid on a subset of EEE, and each of its independent sets is a forest.
  • Lemma 5.3. The construction is a 333-partition scheme.
  • Section 5, p. 10. Any α\alphaα-partition scheme for a graphic matroid, combined with the per-part rule, gives a feasible, eαe\alphaeα-competitive algorithm.

Significance

The theorem shows that the graphic matroid secretary problem admits a constant-competitive algorithm with a small explicit constant. It does so through a reduction: a random partition matroid that is feasible for the original matroid and loses only a constant factor in expectation. The reduction separates the combinatorics (Lemma 5.3) from the online part (Theorem 5.4). The same framework gives algorithms for uniform and transversal matroids and for the weighted and discounted variants on any matroid with an α\alphaα-partition property.

The result is proved in the paper; it has not been formalized. The mission contributes a machine-checked version of the reduction, a formal treatment of a recursively defined random partition, and the classical secretary bound in a reusable finite form. The constant 3e3e3e is not the best known for graphic matroids (Korula–Pál improve it to 2e2e2e), so the formal goal is this algorithm's guarantee, not the best possible ratio.

Difficulty

The online half is routine once the classical bound is available: the relative order of the edges in each part is uniform, and the parts are disjoint. The difficulty is Lemma 5.3. The natural idea of using a fixed optimal forest to build the partition is ruled out because the partition must be chosen before the values are seen. The expectation bound must therefore hold for every valuation at once, for a law that depends on the graph only. The construction is recursive and random: its expected value is not a closed-form sum, and any bound has to be carried through the random sequence of blue-blue subgraphs. Feasibility needs an invariant across rounds: the parts created later live inside the blue-blue edges of every earlier round.

Formalization scope

  • Graph. A SimpleGraph on a Fintype vertex type with decidable adjacency. The edges are G.edgeFinset, and acyclicity of SSS is (SimpleGraph.fromEdgeSet S).IsAcyclic. Multigraphs are not covered.
  • Values. Values are a real function v : Sym2 V → ℝ with ∀ e, 0 ≤ v e; only the values on edges matter.
  • OPT is a Finset.sup' over acyclic subsets of the edge set. A partition is a finite family of nonempty, pairwise disjoint parts inside the edge set. Its max-weight base value is the sum of the part maxima.
  • Random partition. A PMF defined by well-founded recursion on the number of edges. Empty parts are dropped, and edges with two red endpoints are discarded.
  • Random order. The edges are numbered by a fixed enumeration. An arrival order is a permutation of the numbers, and expectation over the order is the average over all ∣E∣!|E|!∣E∣! permutations.
  • Classical rule. It samples ⌊m/e⌋\lfloor m/e\rfloor⌊m/e⌋ arrivals of a part with mmm edges. Ties are broken by preferring the smaller edge number among equal values.
  • Constants. Competitiveness is multiplicative (OPT≤3e⋅E[ALG]\mathrm{OPT}\le 3e\cdot\mathbb E[\mathrm{ALG}]OPT≤3e⋅E[ALG]), so a zero expectation is not a loophole.
  • Ruling out trivial formalizations. In Definition 5.1 the random partition is fixed before the valuation, and the independence requirement holds for every partition in its support. A partition allowed to depend on vvv would make every matroid 111-partitionable.

A complete development needs the classical secretary bound in finite form, the uniformity of induced sub-orders of a uniform permutation, expectations of PMF.bind along a well-founded recursion, and facts about forests in SimpleGraph. The first two, and a general graphic-matroid layer, are reusable beyond this mission. Proofs of any milestone, alternative proofs of Lemma 5.3, and extensions to the uniform and transversal cases of Theorem 5.2 are welcome.

Selected references

  • M. Babaioff, M. Dinitz, A. Gupta, N. Immorlica, K. Talwar, Secretary Problems: Weights and Discounts, Proc. 20th ACM-SIAM Symposium on Discrete Algorithms (SODA), 2009. https://doi.org/10.1137/1.9781611973068.135
  • M. Babaioff, N. Immorlica, R. Kleinberg, Matroids, secretary problems, and online mechanisms, SODA 2007, pp. 434–443. https://dl.acm.org/doi/10.5555/1283383.1283429
  • M. Babaioff, N. Immorlica, D. Kempe, R. Kleinberg, Matroid Secretary Problems, Journal of the ACM 65(6), 2018. https://doi.org/10.1145/3212512
  • N. Korula, M. Pál, Algorithms for Secretary Problems on Graphs and Hypergraphs, ICALP 2009, LNCS 5556. https://doi.org/10.1007/978-3-642-02930-1_42
  • T. S. Ferguson, Who solved the secretary problem?, Statistical Science 4(3), 1989. https://doi.org/10.1214/ss/1177012493
10 thms1 active userReviewed
Algorithmic Game TheoryOperations ResearchTheoretical Computer Science·Captain: mikedeng1

Secretary Problems: Weights and Discounts 1: An (8+3e)-Competitive Algorithm for the Weighted Secretary ProblemResearch Paper

Motivation

The classical secretary problem asks how to select one valuable candidate when candidates arrive in random order and a decision must be made when each candidate appears. Many allocation settings have several goods of unequal quality instead of a single position. An employer may have roles of different desirability, or a seller may have placements with different visibility. In the weighted secretary problem, an agent's value is multiplied by the weight of the good assigned to that agent. The algorithm must decide irrevocably as agents arrive, while the benchmark sees every value before assigning goods. Babaioff, Dinitz, Gupta, Immorlica and Talwar study this model with arbitrary fixed agent values and a uniformly random arrival order, and give a constant competitive ratio independent of the number of agents and goods (authors' version, §§2–3).

The paper also studies time discounts and matroid constraints. This mission concerns its weighted-goods result, Theorem 3.4. The result combines an online allocation rule for several comparably valuable agents with the familiar one-choice secretary rule for an unusually valuable agent. These are distinct ways in which the sorted offline assignment can earn value; both are present even when the weights are fixed in advance. The weighted model matters because matching a valuable agent to an unsuitable good can lose value despite accepting the right agent.

Setting

There are nnn agents e∈Ue\in Ue∈U, each with a nonnegative value v(e)v(e)v(e), and KKK goods indexed in decreasing order of nonnegative weight:

w(1)≥w(2)≥⋯≥w(K)≥0.w(1)\ge w(2)\ge\cdots\ge w(K)\ge0.w(1)≥w(2)≥⋯≥w(K)≥0.

An assignment sss gives each good to at most one agent, and each agent receives at most one good. A good may remain unassigned, represented by ⊥\bot⊥ with v(⊥)=0v(\bot)=0v(⊥)=0. Its value is ∑k=1Kv(s(k))w(k)\sum_{k=1}^K v(s(k))w(k)∑k=1K​v(s(k))w(k). Agent values are arbitrary, not drawn independently from a distribution. The uncertainty is the arrival order π\piπ, chosen uniformly from all permutations; an agent's value becomes visible on arrival, and an allocation decision cannot be revised.

The offline optimum, OPT\mathrm{OPT}OPT, assigns the heaviest good to the highest-valued agent, the next good to the next agent, and so on. If K>nK>nK>n, the extra goods remain unassigned. A consistent tie break makes the ordering unique without changing the numerical value. This sorted assignment is defined directly; the mission does not replace it with an unconstrained variable said to be optimal.

The reservation algorithm draws a sample size τ∼Binom(n,1/2)\tau\sim\mathrm{Binom}(n,1/2)τ∼Binom(n,1/2), observes the first τ\tauτ agents without allocation, and retains the best min⁡(K,τ)\min(K,\tau)min(K,τ) sampled agents. Positive values are grouped into value classes [2i−1,2i)[2^{i-1},2^i)[2i−1,2i) for integer iii. A sampled agent in class iii reserves one good in that class's contiguous block, with higher classes receiving heavier blocks. A later agent receives the heaviest unassigned good reserved for its class when one is available. The classical secretary rule instead observes the first ⌊n/e⌋\lfloor n/e\rfloor⌊n/e⌋ agents, then selects the first later arrival better than every predecessor; its winner receives good 111.

Formalization targets

The mission's goal is the exact guarantee of Theorem 3.4 for Algorithm AAA, which runs the reservation algorithm with probability 8/(3e+8)8/(3e+8)8/(3e+8) and the classical rule with probability 3e/(3e+8)3e/(3e+8)3e/(3e+8):

OPT≤(8+3e) E[A].\mathrm{OPT}\le(8+3e)\,\mathbb E[A].OPT≤(8+3e)E[A].

Here the expectation covers the uniform arrival permutation, the independent binomial sample size used by the reservation branch, and the mixing coin. The multiplicative inequality expresses competitiveness even when an expected payoff is zero. It uses the explicit constant in the paper's proof rather than an instance-dependent or unspecified constant.

Four source results form the milestones. The classical secretary rule selects the maximum with probability at least 1/e1/e1/e. Lemma 3.2 compares the starting indices bib_ibi​ and oio_ioi​ of class-iii blocks in the reservation and optimum assignments. Lemma 3.1 says that if the optimum assigns at least two agents from class iii, the reservation rule assigns at least ui/4u_i/4ui​/4 agents from that class in expectation. Lemma 3.3 converts this to expected value at least OPTi/8\mathrm{OPT}_i/8OPTi​/8. The target retains the paper's class condition and both numerical fractions (authors' version, pp. 4–5).

Significance

Theorem 3.4 supplies a constant factor guarantee for irrevocable allocation when goods have different weights and agents arrive in random order. The factor does not grow with nnn or KKK. It separates the effects of uncertain arrivals from the offline matching of high values to high weights, and it supplies a benchmark for later variants with more complicated feasibility constraints. The paper extends the reservation idea to additional combinatorial settings, including partition-matroid variants in Appendix C (authors' version, Appendix C).

The theorem is proved in the source paper, while the Lean statements in this mission are proof obligations. Formalizing them requires checking that the random-order model, sample distribution, tie convention and assignments jointly express the same algorithm. A complete development will also establish reusable finite-average facts for random permutations and binomial samples, and structural facts about sorted assignments and reserved blocks. Those pieces can support other secretary problems in the series; the mission's specific promise remains the weighted algorithm's exact bound.

Difficulty

A count of how many agents a class receives does not by itself control the weighted value of those goods. Goods have unequal weights, and the value of assigning the next good changes with its position in a block. A class whose offline optimum receives several agents can also lose all its sampled members from the allocation phase. Thus a direct comparison of expected class counts with expected class values is insufficient. The paper's separate count, block-position and value statements identify the claims a solver must establish; the final theorem must also account for classes represented only once in the offline assignment (authors' version, p. 5).

Formalization scope

Agents and goods are Fin n and Fin K; their indices start at zero in Lean, so paper time ttt corresponds to Lean index t−1t-1t−1. An arrival permutation maps time to agent. Values and weights are real and explicitly nonnegative, and weights are antitone in the good index. The finite sums defining expectations are normalized by n!n!n! for permutations and by (nτ)/2n\binom n\tau/2^n(τn​)/2n for sample sizes. No measurability or integration convention is needed. For the goal, K≥1K\ge1K≥1 makes the heaviest good available; K>nK>nK>n is allowed.

Equal values are ordered by smaller original agent index throughout the sorted optimum, the sample's top agents and the classical rule. The classical rule observes exactly ⌊n/e⌋\lfloor n/e\rfloor⌊n/e⌋ arrivals, and zero-valued agents reserve no value-class goods. Positive values below one use negative integer class indices. The paper says only that class iii holds the values “between” 2i−12^{i-1}2i−1 and 2i2^i2i (p. 4, and again in Appendix C, p. 12); the mission fixes the half-open interval [2i−1,2i)[2^{i-1},2^i)[2i−1,2i), so that the classes partition the positive reals (authors' version, pp. 4, 12). A reservation assignment is built from each post-sample agent's rank within its class, so a good is offered to at most one such agent. The theorem is about this concrete algorithm and the concrete sorted offline assignment; an arbitrary favorable policy or an optimum supplied as a hypothesis would not express the source result.

The development needs a finite assignment interface, a tie-aware rank order, value classes, the two online rules, and normalized finite expectations. The assignment and finite-average definitions are reusable. Contributions that prove the structural validity of the reservation assignment, the classical success guarantee, Lemmas 3.1–3.3, or the final combination all advance the stated target.

Selected references

  • Moshe Babaioff, Michael Dinitz, Anupam Gupta, Nicole Immorlica and Kunal Talwar, Secretary Problems: Weights and Discounts, Proceedings of SODA 2009; authors' full version, proceedings DOI.
7 thms1 active userReviewed
Dynamical SystemsOperations ResearchStochastic Systems·Captain: mikedeng1

Dynamics of Stochastic Approximation Algorithms 6: An Attractor Whose Basin Meets the Attainable Set Contains the Limit Set with Positive ProbabilityResearch Paper

Motivation

A stochastic approximation algorithm is a recursion xn+1=xn+γn+1(F(xn)+Un+1)x_{n+1}=x_n+\gamma_{n+1}\big(F(x_n)+U_{n+1}\big)xn+1​=xn​+γn+1​(F(xn​)+Un+1​) with decreasing step sizes γn\gamma_nγn​ and a noise term Un+1U_{n+1}Un+1​. Recursions of this form appear in stochastic gradient methods, adaptive control, learning in games (fictitious play, reinforcement learning) and urn models. The ODE method compares the iterates with the solutions of x˙=F(x)\dot x=F(x)x˙=F(x). In Benaïm's lecture notes (Benaïm 1999), the comparison is phrased through the continuous-time interpolated process XXX. Under standard noise conditions, XXX is almost surely an asymptotic pseudotrajectory of the flow of FFF, and its limit set is almost surely internally chain transitive.

That theorem constrains where the process may end up. It does not say which of several candidate sets the process actually reaches. When the ODE has several attractors, for example several stable equilibria of a learning dynamic or several stable compositions of an urn, an application needs to know that each attractor is reached with positive probability. Section 7 of the notes answers this question. The answer is a criterion of attainability: if the process can, with positive probability and at arbitrarily late times, enter the basin of an attractor, then it converges to that attractor with positive probability.

Timeline.

  • Kushner and Clark (1978) proved convergence statements for processes that visit a compact subset of the domain of attraction of an asymptotically stable equilibrium infinitely often.
  • Arthur, Ermoliev and Kaniovski (1983) and Pemantle (1990) studied urn processes whose limit points depend on the trajectory.
  • Benaïm (1997) and Duflo (1997, Random Iterative Models) developed the attainability argument for general stochastic approximation processes.
  • Benaïm (1999) states it for arbitrary attractors of a semiflow on a locally compact metric space, under a single conditional shadowing condition (24).

Setting

Let (M,d)(M,d)(M,d) be a metric space and let Φ=(Φt)t≥0\Phi=(\Phi_t)_{t\ge0}Φ=(Φt​)t≥0​ be a semiflow on MMM: a continuous map (t,x)↦Φt(x)(t,x)\mapsto\Phi_t(x)(t,x)↦Φt​(x) with Φ0=Id\Phi_0=\mathrm{Id}Φ0​=Id and Φt+s=Φt∘Φs\Phi_{t+s}=\Phi_t\circ\Phi_sΦt+s​=Φt​∘Φs​.

  • A set AAA is invariant if Φt(A)=A\Phi_t(A)=AΦt​(A)=A for all t≥0t\ge0t≥0, and positively invariant if Φt(A)⊂A\Phi_t(A)\subset AΦt​(A)⊂A.
  • An attractor is a nonempty compact invariant set AAA with a neighbourhood WWW on which dist⁡(Φtx,A)→0\operatorname{dist}(\Phi_tx,A)\to0dist(Φt​x,A)→0 uniformly. Its basin B(A)B(A)B(A) is the set of points xxx with dist⁡(Φtx,A)→0\operatorname{dist}(\Phi_tx,A)\to0dist(Φt​x,A)→0.
  • A continuous curve X:R+→MX:\mathbb R_+\to MX:R+​→M is an asymptotic pseudotrajectory if sup⁡0≤h≤Td(X(t+h),Φh(X(t)))→0\sup_{0\le h\le T}d(X(t+h),\Phi_h(X(t)))\to0sup0≤h≤T​d(X(t+h),Φh​(X(t)))→0 as t→∞t\to\inftyt→∞, for every T>0T>0T>0.
  • The limit set of XXX is L(X)=⋂t≥0X([t,∞))‾L(X)=\bigcap_{t\ge0}\overline{X([t,\infty))}L(X)=⋂t≥0​X([t,∞))​.
  • For T>0T>0T>0, dX(T)=sup⁡k∈Nd(ΦT(X(kT)),X(kT+T))d_X(T)=\sup_{k\in\mathbb N}d(\Phi_T(X(kT)),X(kT+T))dX​(T)=supk∈N​d(ΦT​(X(kT)),X(kT+T)).

Now let X=(X(t))t≥0X=(X(t))_{t\ge0}X=(X(t))t≥0​ be a process on a probability space (Ω,F,P)(\Omega,\mathcal F,P)(Ω,F,P) with continuous paths in MMM, adapted to a filtration (Ft)t≥0(\mathcal F_t)_{t\ge0}(Ft​)t≥0​. The standing assumption of Section 7 is that for all δ>0\delta>0δ>0, T>0T>0T>0 and t≥0t\ge0t≥0,

P(sup⁡s≥t sup⁡0≤h≤Td(X(s+h),Φh(X(s)))≥δ ∣ Ft)≤w(t,δ,T)(24)P\Big(\sup_{s\ge t}\ \sup_{0\le h\le T}d\big(X(s+h),\Phi_h(X(s))\big)\ge\delta\ \Big|\ \mathcal F_t\Big)\le w(t,\delta,T)\tag{24}P(s≥tsup​ 0≤h≤Tsup​d(X(s+h),Φh​(X(s)))≥δ ​ Ft​)≤w(t,δ,T)(24)

for a function w≥0w\ge0w≥0 with w(t,δ,T)↓0w(t,\delta,T)\downarrow0w(t,δ,T)↓0 as t→∞t\to\inftyt→∞.

A point ppp is attainable if P(∃s≥t:X(s)∈U)>0P(\exists s\ge t: X(s)\in U)>0P(∃s≥t:X(s)∈U)>0 for every t>0t>0t>0 and every open neighbourhood UUU of ppp. Att(X)\mathrm{Att}(X)Att(X) is the set of attainable points.

Formalization targets

Goal: Theorem 7.3, first statement

If MMM is locally compact, AAA is an attractor of Φ\PhiΦ, and Att(X)∩B(A)≠∅\mathrm{Att}(X)\cap B(A)\neq\emptysetAtt(X)∩B(A)=∅, then

P(L(X)⊂A)>0.P\big(L(X)\subset A\big)>0 .P(L(X)⊂A)>0.

This statement contains no constants and no rates, so it does not depend on how (24) is quantified for a particular algorithm.

Theorem 7.3, second statement

If UUU is open and relatively compact with U‾⊂B(A)\overline U\subset B(A)U⊂B(A), there are T,δ>0T,\delta>0T,δ>0, depending only on UUU (and on Φ\PhiΦ, AAA), such that for every process satisfying the standing assumption and every t≥0t\ge0t≥0

P(L(X)⊂A)≥(1−w(t,δ,T)) P(∃s≥t: X(s)∈U).P\big(L(X)\subset A\big)\ge\big(1-w(t,\delta,T)\big)\,P\big(\exists s\ge t:\ X(s)\in U\big).P(L(X)⊂A)≥(1−w(t,δ,T))P(∃s≥t: X(s)∈U).

Milestones

  • Lemma 6.8. For a nonempty compact K⊂B(A)K\subset B(A)K⊂B(A) there are T,δ>0T,\delta>0T,δ>0 such that every asymptotic pseudotrajectory with X(0)∈KX(0)\in KX(0)∈K and dX(T)<δd_X(T)<\deltadX​(T)<δ has L(X)⊂AL(X)\subset AL(X)⊂A.
  • Lemma 7.1, in three parts:
    • Att(X)\mathrm{Att}(X)Att(X) is closed;
    • it is positively invariant;
    • it contains L(X)L(X)L(X) almost surely.

Significance

The result. Theorem 7.3 turns a question about the long-run limit of a random process into a question about where the process can go. Attainability is usually checked by a controllability argument: the noise can push the iterates in every direction. For urn processes with an urn function mapping the simplex into its interior, every point is attainable (Example 7.2 of the notes). Then every attractor of the mean ODE is reached with positive probability. Combined with nonconvergence results for unstable sets (Section 9 of the notes), this characterizes the possible limits of many learning and urn processes. Theorem 7.3 is the positive half of that picture.

Formalizing it. The theorem has a published proof (p. 32 of the notes) and no machine-checked version. A formal proof needs the following:

  • a precise reading of the conditional shadowing condition (24) as a conditional expectation of an indicator;
  • the stopping-time decomposition of the event {∃s≥t:X(s)∈U}\{\exists s\ge t: X(s)\in U\}{∃s≥t:X(s)∈U} over dyadic times;
  • the deterministic Lemma 6.8, which rests on the limit set theorem for precompact asymptotic pseudotrajectories (Theorem 5.7 of the notes, the subject of mission 1 of this series).

Difficulty

The obvious argument says: once XXX enters a compact part of the basin, the flow carries it into AAA. That fails because XXX is not a trajectory of the flow. Each window of length TTT adds an error, and errors over infinitely many windows can push the process out of the basin.

Two things are needed instead:

  • A uniform version of the deterministic statement, with TTT and δ\deltaδ fixed in advance from the compact set alone. This is Lemma 6.8, which needs local compactness of MMM and the structure of limit sets of asymptotic pseudotrajectories.
  • A probabilistic step that applies (24) at the random time when XXX first enters UUU. That time is not a stopping time on a continuum, and conditioning at it needs care.

A naive union bound over all times is useless: it does not use the conditional form of (24).

Formalization scope

  • The semiflow is Mathlib's Flow ℝ≥0 M on a metric space; local compactness is LocallyCompactSpace M.
  • The process is X : ℝ≥0 → Ω → M with continuous paths. The paper's alternative of càdlàg paths is not covered.
  • Adaptedness is Borel measurability of X(t)X(t)X(t) with respect to Ft\mathcal F_tFt​, for a Mathlib Filtration ℝ≥0. PPP is a probability measure.
  • The suprema in (24) and in dX(T)d_X(T)dX​(T) are computed in [0,∞][0,\infty][0,∞] with the extended distance, so that "sup ≥δ\ge\delta≥δ" and "sup <δ<\delta<δ" are exact even when the supremum is infinite or not attained.
  • The conditional probability in (24) is the conditional expectation of the indicator of the event. The event is required to be measurable, so the condition cannot hold vacuously through a junk conditional expectation.
  • www is required to be both nonincreasing in ttt and convergent to 000.
  • The events {L(X)⊂A}\{L(X)\subset A\}{L(X)⊂A} and {∃s≥t:X(s)∈U}\{\exists s\ge t: X(s)\in U\}{∃s≥t:X(s)∈U} are measured with PPP as an outer measure, so no measurability hypothesis is added for them.
  • In the second statement of Theorem 7.3, TTT and δ\deltaδ are chosen before the probability space, the process, www and ttt.
  • Invariance in the definition of an attractor is the equality Φt(A)=A\Phi_t(A)=AΦt​(A)=A, not inclusion.
  • The almost-sure clause of Lemma 7.1 is stated for separable MMM. Without separability it cannot be proved in ordinary set theory.

These choices rule out the trivializing formalizations: a conditional-probability hypothesis that holds vacuously, invariance read as inclusion, constants T,δT,\deltaT,δ that depend on the process or on ttt, and a probability bound www without monotonicity.

A complete development needs:

  • limit sets of asymptotic pseudotrajectories and the fact that an internally chain transitive set meeting the basin of an attractor lies in the attractor (shared with missions 1 and 5 of this series);
  • measurability of path functionals of continuous processes;
  • conditioning on events of the form {τ=tn(k)}\{\tau=t_n(k)\}{τ=tn​(k)} at dyadic times.

The first and second are reusable well beyond this mission. Contributions toward either, and alternative proofs of Lemma 6.8, are welcome.

Selected references

  • M. Benaïm, Dynamics of stochastic approximation algorithms, Séminaire de Probabilités XXXIII, Lecture Notes in Mathematics 1709, Springer, 1999, pp. 1–68. https://doi.org/10.1007/BFb0096509
  • M. Benaïm, M. W. Hirsch, Asymptotic pseudotrajectories and chain recurrent flows, with applications, Journal of Dynamics and Differential Equations 8 (1996), 141–176. https://doi.org/10.1007/BF02218617
  • M. Benaïm, Vertex-reinforced random walks and a conjecture of Pemantle, Annals of Probability 25 (1997), 361–392. https://doi.org/10.1214/aop/1024404292
  • M. Duflo, Random Iterative Models, Applications of Mathematics 34, Springer, 1997. https://doi.org/10.1007/978-3-662-12880-0
  • H. J. Kushner, D. S. Clark, Stochastic Approximation Methods for Constrained and Unconstrained Systems, Springer, 1978. https://doi.org/10.1007/978-1-4684-9352-8
  • C. Conley, Isolated Invariant Sets and the Morse Index, CBMS Regional Conference Series in Mathematics 38, AMS, 1978. https://doi.org/10.1090/cbms/038
11 thms1 active userReviewed
Dynamical SystemsOperations ResearchStochastic Systems·Captain: mikedeng1

Dynamics of Stochastic Approximation Algorithms 3: Martingale Noise with Bounded q-th Moments and Summable γ_n^(1+q/2) Satisfies Assumption A1 Almost SurelyResearch Paper

Motivation

A stochastic approximation algorithm is a recursion

xn+1−xn=γn+1(F(xn)+Un+1)x_{n+1}-x_n=\gamma_{n+1}\big(F(x_n)+U_{n+1}\big)xn+1​−xn​=γn+1​(F(xn​)+Un+1​)

in Rd\mathbb R^dRd, where FFF is a vector field, γn\gamma_nγn​ are small step sizes and Un+1U_{n+1}Un+1​ is noise. Such recursions go back to Robbins and Monro's root-finding scheme (Robbins–Monro 1951) and underlie stochastic gradient descent, temporal-difference learning, adaptive control and learning in games. The ODE method studies them by comparing the iterates with the trajectories of x˙=F(x)\dot x=F(x)x˙=F(x).

Benaïm's lecture notes (Benaïm 1999) organize the ODE method in two steps. A deterministic step, Proposition 4.1, shows that whenever the noise satisfies a condition called A1 (and the iterates are bounded, or FFF is Lipschitz and bounded on a neighbourhood of them), the interpolated process is an asymptotic pseudotrajectory of the flow of FFF. A probabilistic step then verifies A1 for concrete noise models. This mission formalizes the first such verification, Proposition 4.2: martingale difference noise with bounded qqq-th moments and step sizes with ∑nγn1+q/2<∞\sum_n\gamma_n^{1+q/2}<\infty∑n​γn1+q/2​<∞. The result is described as a particular case of a general theorem of Métivier and Priouret (1987); the same estimates reappear later in the notes.

Setting

Let {γn}n≥1\{\gamma_n\}_{n\ge1}{γn​}n≥1​ be a deterministic sequence with γn≥0\gamma_n\ge0γn​≥0, ∑nγn=∞\sum_n\gamma_n=\infty∑n​γn​=∞ and γn→0\gamma_n\to0γn​→0 (a step sequence). Put τ0=0\tau_0=0τ0​=0, τn=∑i=1nγi\tau_n=\sum_{i=1}^n\gamma_iτn​=∑i=1n​γi​, and let

m(t)=sup⁡{k≥0: t≥τk}m(t)=\sup\{k\ge0:\ t\ge\tau_k\}m(t)=sup{k≥0: t≥τk​}

be the index of the step that contains time t≥0t\ge0t≥0. For a sequence {Un}n≥1\{U_n\}_{n\ge1}{Un​}n≥1​ define the piecewise constant processes Uˉ(t)=Um(t)+1\bar U(t)=U_{m(t)+1}Uˉ(t)=Um(t)+1​ and γˉ(t)=γm(t)+1\bar\gamma(t)=\gamma_{m(t)+1}γˉ​(t)=γm(t)+1​, so that step n+1n+1n+1 occupies the time interval [τn,τn+1)[\tau_n,\tau_{n+1})[τn​,τn+1​) of length γn+1\gamma_{n+1}γn+1​.

Assumption A1 asks that for every T>0T>0T>0

lim⁡n→∞sup⁡{∥∑i=nk−1γi+1Ui+1∥: k=n+1,…,m(τn+T)}=0,\lim_{n\to\infty}\sup\Big\{\Big\|\sum_{i=n}^{k-1}\gamma_{i+1}U_{i+1}\Big\|:\ k=n+1,\dots,m(\tau_n+T)\Big\}=0,n→∞lim​sup{​i=n∑k−1​γi+1​Ui+1​​: k=n+1,…,m(τn​+T)}=0,

or, in the form the notes call equivalent, lim⁡t→∞Δ(t,T)=0\lim_{t\to\infty}\Delta(t,T)=0limt→∞​Δ(t,T)=0 for every T>0T>0T>0, where

Δ(t,T)=sup⁡0≤h≤T∥∫tt+hUˉ(s) ds∥.\Delta(t,T)=\sup_{0\le h\le T}\Big\|\int_t^{t+h}\bar U(s)\,ds\Big\|.Δ(t,T)=0≤h≤Tsup​​∫tt+h​Uˉ(s)ds​.

Let (Ω,F,P)(\Omega,\mathcal F,P)(Ω,F,P) be a probability space with a nondecreasing sequence {Fn}\{\mathcal F_n\}{Fn​} of sub-σ\sigmaσ-algebras, and F:Rd→RdF:\mathbb R^d\to\mathbb R^dF:Rd→Rd continuous. A sequence {xn}\{x_n\}{xn​} given by the recursion above is a Robbins–Monro algorithm if γ\gammaγ is deterministic, UnU_nUn​ is Fn\mathcal F_nFn​-measurable, and E(Un+1∣Fn)=0E(U_{n+1}\mid\mathcal F_n)=0E(Un+1​∣Fn​)=0.

Formalization targets

Goal: Proposition 4.2

For a Robbins–Monro algorithm and some real q≥2q\ge2q≥2, if

sup⁡nE(∥Un+1∥q)<∞and∑nγn1+q/2<∞,\sup_nE\big(\|U_{n+1}\|^q\big)<\infty\qquad\text{and}\qquad\sum_n\gamma_n^{1+q/2}<\infty,nsup​E(∥Un+1​∥q)<∞andn∑​γn1+q/2​<∞,

then with probability one the realised noise sequence satisfies A1, in both of its forms, simultaneously for all T>0T>0T>0.

Milestones

  1. Eq. (13), an instance of Burkholder's inequality with a universal constant CqC_qCq​:
E{sup⁡n<k≤m(τn+T)∥∑i=nk−1γi+1Ui+1∥q}≤Cq E{[∑i=nm(τn+T)−1γi+12∥Ui+1∥2]q/2}.E\Big\{\sup_{n<k\le m(\tau_n+T)}\Big\|\sum_{i=n}^{k-1}\gamma_{i+1}U_{i+1}\Big\|^q\Big\}\le C_q\,E\Big\{\Big[\sum_{i=n}^{m(\tau_n+T)-1}\gamma_{i+1}^2\|U_{i+1}\|^2\Big]^{q/2}\Big\}.E{n<k≤m(τn​+T)sup​​i=n∑k−1​γi+1​Ui+1​​q}≤Cq​E{[i=n∑m(τn​+T)−1​γi+12​∥Ui+1​∥2]q/2}.
  1. Inequality (14), for finite families with αi≥0\alpha_i\ge0αi​≥0, u>1u>1u>1, 0<δ<10<\delta<10<δ<1:
(∑i∣αiβi∣)u≤(∑iαiδu/(u−1))u−1∑iαi(1−δ)u∣βi∣u.\Big(\sum_i|\alpha_i\beta_i|\Big)^u\le\Big(\sum_i\alpha_i^{\delta u/(u-1)}\Big)^{u-1}\sum_i\alpha_i^{(1-\delta)u}|\beta_i|^u.(i∑​∣αi​βi​∣)u≤(i∑​αiδu/(u−1)​)u−1i∑​αi(1−δ)u​∣βi​∣u.
  1. Eq. (16): for every T>0T>0T>0 there is C(q,T)C(q,T)C(q,T) with E(Δ(t,T)q)≤C(q,T)∫tt+Tγˉq/2(s) dsE(\Delta(t,T)^q)\le C(q,T)\int_t^{t+T}\bar\gamma^{q/2}(s)\,dsE(Δ(t,T)q)≤C(q,T)∫tt+T​γˉ​q/2(s)ds for all t≥0t\ge0t≥0.
  2. Eq. (17): ∑k≥0E(Δ(kT,T)q)<∞\sum_{k\ge0}E(\Delta(kT,T)^q)<\infty∑k≥0​E(Δ(kT,T)q)<∞ for every T>0T>0T>0.
  3. Block comparison: Δ(t,T)≤2Δ(kT,T)+Δ((k+1)T,T)\Delta(t,T)\le2\Delta(kT,T)+\Delta((k+1)T,T)Δ(t,T)≤2Δ(kT,T)+Δ((k+1)T,T) for kT≤t<(k+1)TkT\le t<(k+1)TkT≤t<(k+1)T.

Significance

Proposition 4.2 is the standard sufficient condition under which the ODE method applies to stochastic gradient-type recursions with martingale noise. With q=2q=2q=2 it covers step sizes with ∑γn2<∞\sum\gamma_n^2<\infty∑γn2​<∞ (for example γn=1/n\gamma_n=1/nγn​=1/n) and noise with bounded variance; larger qqq trades stronger moment assumptions for slower decay of the steps, down to ∑γn1+q/2<∞\sum\gamma_n^{1+q/2}<\infty∑γn1+q/2​<∞. Combined with Proposition 4.1 it shows that the interpolated process of a Robbins–Monro algorithm with bounded iterates is almost surely an asymptotic pseudotrajectory of the flow of FFF; the limit set theorems of the notes then locate the limit points of the algorithm.

The result is proved in the notes and in the cited literature; it has not, to our knowledge, been machine-checked. A formal proof would supply reusable pieces that Mathlib currently lacks, most notably a Burkholder (or Burkholder–Davis–Gundy) inequality for discrete-time vector martingales in LqL^qLq, and the continuous-time bookkeeping of the step processes Uˉ\bar UUˉ, γˉ\bar\gammaγˉ​ and the noise deviation Δ\DeltaΔ, shared by the other missions of this series.

Difficulty

The obvious argument controls each window by Doob's L2L^2L2 maximal inequality and sums over windows. That works for q=2q=2q=2 only. For q>2q>2q>2 the second moment of the window sums is not summable under ∑γn1+q/2<∞\sum\gamma_n^{1+q/2}<\infty∑γn1+q/2​<∞, and one needs an LqL^qLq maximal inequality whose right-hand side is the q/2q/2q/2-th moment of the square function. That inequality, Burkholder's, is not in Mathlib. Converting the square function into the moment bound requires a Hölder-type inequality with tuned exponents, and passing from the discrete sums to Δ(t,T)\Delta(t,T)Δ(t,T) requires handling partial steps at both ends of [t,t+h][t,t+h][t,t+h]. A second subtlety is that A1 quantifies over all T>0T>0T>0: the almost-sure statement must hold on a single event of full probability for every TTT, not on an event that depends on TTT.

Formalization scope

The space is Rd\mathbb R^dRd as EuclideanSpace ℝ (Fin d); time is real; qqq is a real number with q≥2q\ge2q≥2, and all powers are real powers of nonnegative quantities. The sequences γ\gammaγ and UUU are indexed by N\mathbb NN, and their values at 000 are unused, as the paper indexes them from 111. The filtration is a Mathlib Filtration ℕ, Un+1U_{n+1}Un+1​ is required to be Fn+1\mathcal F_{n+1}Fn+1​-strongly measurable and integrable, and the martingale difference condition is E(Un+1∣Fn)=0E(U_{n+1}\mid\mathcal F_n)=0E(Un+1​∣Fn​)=0 almost surely. Expectations of nonnegative quantities, the suprema in A1 and Δ\DeltaΔ, and the moment bound are taken in [0,∞][0,\infty][0,∞], so no default value of a non-integrable expectation or of an empty supremum can make a statement hold vacuously; the supremum over an empty range of kkk is 000.

The following readings are excluded and are not acceptable formalizations: a moment hypothesis that holds vacuously, a conditional expectation hypothesis on non-integrable noise, the conclusion "for each TTT, A1 holds almost surely" in place of "almost surely, A1 holds for all TTT", and qqq fixed to 222 or restricted to integers.

All hypotheses are satisfiable: U=0U=0U=0, x=0x=0x=0, F=0F=0F=0 and γn=1/n\gamma_n=1/nγn​=1/n with q=2q=2q=2 satisfy every one of them.

Contributions welcome: a general Burkholder inequality for discrete-time martingales in finite-dimensional spaces (reusable well beyond this mission), lemmas on the step processes and Δ\DeltaΔ (measurability, local integrability, additivity), and the proofs of the milestones.

Selected references

  • M. Benaïm, Dynamics of Stochastic Approximation Algorithms, Séminaire de Probabilités XXXIII, Lecture Notes in Mathematics 1709, Springer, 1999, pp. 1–68. https://doi.org/10.1007/BFb0096509
  • M. Benaïm and M. W. Hirsch, Asymptotic pseudotrajectories and chain recurrent flows, with applications, Journal of Dynamics and Differential Equations 8 (1996), 141–176. https://doi.org/10.1007/BF02218617
  • M. Métivier and P. Priouret, Théorèmes de convergence presque sûre pour une classe d'algorithmes stochastiques à pas décroissant, Probability Theory and Related Fields 74 (1987), 403–428.
  • D. L. Burkholder, Distribution function inequalities for martingales, Annals of Probability 1 (1973), 19–42. https://doi.org/10.1214/aop/1176997023
  • D. W. Stroock, Probability Theory: An Analytic View, Cambridge University Press, 1993.
  • H. Robbins and S. Monro, A stochastic approximation method, Annals of Mathematical Statistics 22 (1951), 400–407. https://doi.org/10.1214/aoms/1177729586
  • H. J. Kushner and G. G. Yin, Stochastic Approximation Algorithms and Applications, Springer, 1997.
10 thms1 active userReviewed
Operations ResearchOptimizationStatistics·Captain: mikedeng1

Asymptotic Behavior of Statistical Estimators and of Optimal Solutions of Stochastic Optimization Problems: Optimal Solutions Under Estimated Distributions Are Strongly ConsistentResearch Paper

Motivation

Many estimation procedures in statistics, and most stochastic optimization models in operations research, have the same shape: a decision or parameter x∈Rnx\in\mathbb R^nx∈Rn is chosen to minimize an expected loss Ef(x)=∫f(x,ξ) P(dξ)Ef(x)=\int f(x,\xi)\,P(d\xi)Ef(x)=∫f(x,ξ)P(dξ) under a distribution PPP that is not known. In practice PPP is replaced by an estimate PνP^\nuPν built from the information available at stage ν\nuν (an empirical measure, a smoothed or parametric fit, a Bayesian posterior), and the minimizer of the estimated problem is used in place of the true one. The basic question is whether this is justified: do the estimated solutions converge to a true solution, and the estimated optimal values to the true optimal value, as information accumulates?

For maximum likelihood this is Wald's consistency theorem (Wald 1949); Huber extended it to M-estimators under non-standard conditions (Huber 1967). Both settings are unconstrained, or constrained to an open set, and assume finite-valued criteria. Constrained least squares, L1L^1L1 and Huber regression with inequality constraints, variance-component models with Heywood cases, and two-stage stochastic programs with recourse all lead instead to criteria that take the value +∞+\infty+∞ off a closed feasible set and are only lower semicontinuous in xxx.

J. Dupačová and R. Wets (IIASA WP-86-41, 1986; journal version Ann. Statist. 16 (1988)) proved consistency in this generality by combining epi-convergence of functions with the theory of measurable multifunctions and normal integrands. This mission formalizes their §3.

Setting

Ξ\XiΞ is a Polish space with its Borel σ\sigmaσ-field and PPP is a probability measure on it. The integrand is f:Rn×Ξ→(−∞,∞]f:\mathbb R^n\times\Xi\to(-\infty,\infty]f:Rn×Ξ→(−∞,∞], and the true problem is to minimize

Ef(x)=∫Ξf(x,ξ) P(dξ),Ef(x)=\int_\Xi f(x,\xi)\,P(d\xi),Ef(x)=∫Ξ​f(x,ξ)P(dξ),

with the convention that Ef(x)=+∞Ef(x)=+\inftyEf(x)=+∞ whenever ξ↦f(x,ξ)\xi\mapsto f(x,\xi)ξ↦f(x,ξ) is not bounded above by a summable function. The effective domain of a function h:Rn→[−∞,∞]h:\mathbb R^n\to[-\infty,\infty]h:Rn→[−∞,∞] is dom⁡h={x:h(x)<∞}\operatorname{dom}h=\{x: h(x)<\infty\}domh={x:h(x)<∞}, and argmin⁡h={x:h(x)=inf⁡h}\operatorname{argmin}h=\{x: h(x)=\inf h\}argminh={x:h(x)=infh}.

Information arrives on a probability space (Z,F,μ)(Z,\mathcal F,\mu)(Z,F,μ) with an increasing sequence of σ\sigmaσ-fields F1⊆F2⊆⋯⊆F\mathcal F^1\subseteq\mathcal F^2\subseteq\dots\subseteq\mathcal FF1⊆F2⊆⋯⊆F. Each sample ζ∈Z\zeta\in Zζ∈Z yields probability measures Pν(⋅,ζ)P^\nu(\cdot,\zeta)Pν(⋅,ζ) on Ξ\XiΞ, and ζ↦Pν(A,ζ)\zeta\mapsto P^\nu(A,\zeta)ζ↦Pν(A,ζ) is Fν\mathcal F^\nuFν-measurable for every Borel AAA: the estimate at stage ν\nuν uses only stage-ν\nuν information. The estimated problem minimizes

Eνf(x,ζ)=∫Ξf(x,ξ) Pν(dξ,ζ).E^\nu f(x,\zeta)=\int_\Xi f(x,\xi)\,P^\nu(d\xi,\zeta).Eνf(x,ζ)=∫Ξ​f(x,ξ)Pν(dξ,ζ).

A sequence gνg^\nugν epi-converges to ggg if, at every xxx, lim inf⁡gν(xν)≥g(x)\liminf g^\nu(x^\nu)\ge g(x)liminfgν(xν)≥g(x) along every sequence xν→xx^\nu\to xxν→x, and lim sup⁡gν(xν)≤g(x)\limsup g^\nu(x^\nu)\le g(x)limsupgν(xν)≤g(x) along some sequence xν→xx^\nu\to xxν→x.

The standing hypotheses are Assumption 3.4: dom⁡f=S×Ξ\operatorname{dom}f=S\times\Xidomf=S×Ξ with SSS closed and nonempty; f(x,⋅)f(x,\cdot)f(x,⋅) is continuous for x∈Sx\in Sx∈S; f(⋅,ξ)f(\cdot,\xi)f(⋅,ξ) is lower semicontinuous; and fff is locally lower Lipschitz on SSS with a bounded continuous modulus β(ξ)\beta(\xi)β(ξ). Assumption 3.5 asks that, for μ\muμ-almost every ζ\zetaζ, Pν(⋅,ζ)P^\nu(\cdot,\zeta)Pν(⋅,ζ) converge in distribution to PPP, that ∣f(x,⋅)∣|f(x,\cdot)|∣f(x,⋅)∣ be uniformly tight along P=P0,P1,…P=P^0,P^1,\dotsP=P0,P1,… for each x∈Sx\in Sx∈S, and that ∫inf⁡xf(x,ξ) Pν(dξ,ζ)>−∞\int\inf_x f(x,\xi)\,P^\nu(d\xi,\zeta)>-\infty∫infx​f(x,ξ)Pν(dξ,ζ)>−∞ for all ν\nuν.

Formalization targets

Goal: Theorem 3.9, "In particular" (pp. 21–22)

Let D⊆RnD\subseteq\mathbb R^nD⊆Rn be compact, suppose (argmin⁡Eνf)∩D≠∅(\operatorname{argmin}E^\nu f)\cap D\neq\emptyset(argminEνf)∩D=∅ μ\muμ-a.s. for every ν\nuν, and suppose {x∗}=argmin⁡Ef∩D\{x^*\}=\operatorname{argmin}Ef\cap D{x∗}=argminEf∩D. Then there are Fν\mathcal F^\nuFν-measurable selections xνx^\nuxν of argmin⁡Eνf\operatorname{argmin}E^\nu fargminEνf with

xν(ζ)→x∗andinf⁡Eνf(⋅,ζ)→inf⁡Effor μ-almost every ζ.x^\nu(\zeta)\to x^*\quad\text{and}\quad \inf E^\nu f(\cdot,\zeta)\to\inf Ef\qquad\text{for }\mu\text{-almost every }\zeta .xν(ζ)→x∗andinfEνf(⋅,ζ)→infEffor μ-almost every ζ.

The goal does not assume that EfEfEf has a unique global minimizer, and it does not assume convexity.

Milestones

In attack order:

  • Proposition 3.3: epi-convergence gives lim sup⁡(inf⁡gν)≤inf⁡g\limsup(\inf g^\nu)\le\inf glimsup(infgν)≤infg, limits of minimizers are minimizers, and the minimum is attained in the closure of a bounded DDD.
  • Lemma 3.6: almost surely, EfEfEf and every EνfE^\nu fEνf are proper and l.s.c., with domain SSS.
  • Theorem 3.7: almost surely, EνfE^\nu fEνf epi-converges and converges pointwise to EfEfEf.
  • Theorem 3.8: almost surely, the epigraphs of EνfE^\nu fEνf are closed, and they depend Fν\mathcal F^\nuFν-measurably on ζ\zetaζ.
  • Theorem 3.9:
    • (3.14) lim sup⁡(inf⁡Eνf)≤inf⁡Ef\limsup(\inf E^\nu f)\le\inf Eflimsup(infEνf)≤infEf a.s.;
    • (i) cluster points of estimated minimizers minimize EfEfEf;
    • (ii) ζ↦argmin⁡Eνf(⋅,ζ)\zeta\mapsto\operatorname{argmin}E^\nu f(\cdot,\zeta)ζ↦argminEνf(⋅,ζ) is closed-valued and Fν\mathcal F^\nuFν-measurable.
  • Proposition 3.1: the measurable selection theorem.

Significance

The result separates two things: the statistical input, which is only convergence in distribution of PνP^\nuPν plus a tightness condition, and the variational output, which is convergence of optimal values and solutions. It therefore applies to any estimator PνP^\nuPν that converges weakly almost surely: empirical measures, kernel estimates, parametric fits. It also covers constrained and nonsmooth problems: the feasible set enters through f=+∞f=+\inftyf=+∞ off SSS, and only lower semicontinuity in xxx is required. Asymptotic distribution results for constrained estimators, such as the second part of the same paper and the subsequent literature on sample average approximation, start from this consistency.

The theorem is proved on paper. To the best of current knowledge none of it is machine-checked. Mathlib has weak convergence of probability measures, lower semicontinuity and extended-real integrals. It does not have epi-convergence, Effros-measurable multifunctions, normal integrands or the Kuratowski–Ryll-Nardzewski selection theorem. A formal proof produces these as reusable components. It also has to supply the details that the paper's proof of Theorem 3.8 leaves as a sketch.

Difficulty

Pointwise convergence Eνf(x)→Ef(x)E^\nu f(x)\to Ef(x)Eνf(x)→Ef(x) is not enough to move minimizers to the limit, and uniform convergence fails because fff is +∞+\infty+∞ off SSS and need not be bounded. Epi-convergence is the right notion. Proving it needs a liminf inequality along moving points xν→xx^\nu\to xxν→x under moving measures PνP^\nuPν. That combines Fatou's lemma, the lower Lipschitz bound and the tightness condition, and the integrands are extended-real-valued, so care is needed.

The second difficulty is measurability. The exceptional null set lies in F\mathcal FF but not in Fν\mathcal F^\nuFν, so "Fν\mathcal F^\nuFν-measurable" has to be understood on a full-measure set in the trace σ\sigmaσ-field. The paper's argument for Theorem 3.8 appeals to continuity of P↦epi⁡EPfP\mapsto\operatorname{epi}E_PfP↦epiEP​f in the epi-topology, and it remarks itself that Theorem 3.7 gives this only along sequences satisfying Assumption 3.5. A solver will have to rebuild this step, for example through the normal-integrand structure of EνfE^\nu fEνf.

Formalization scope

  • Rn\mathbb R^nRn is EuclideanSpace ℝ (Fin n) with its Euclidean norm.
  • Ξ\XiΞ is a Polish space with its Borel σ\sigmaσ-algebra. This is exactly a closed subset of a Polish space with the relative Borel field.
  • fff is EReal-valued. Every expectation is the mission's expect: +∞+\infty+∞ when ∫f+=∞\int f^+=\infty∫f+=∞, and ∫f+−∫f−\int f^+-\int f^-∫f+−∫f− otherwise, both computed as Lebesgue integrals of [0,∞][0,\infty][0,∞]-valued functions. A Bochner integral, which would assign 000 to non-integrable functions, is never used for EfEfEf or EνfE^\nu fEνf.
  • The sample index is shifted: Lean's Pν k and 𝔽 k are the paper's Pk+1P^{k+1}Pk+1 and Fk+1\mathcal F^{k+1}Fk+1, and P=P0P=P^0P=P0 is a separate argument.
  • Infima, lim inf⁡\liminfliminf and lim sup⁡\limsuplimsup are taken in [−∞,∞][-\infty,\infty][−∞,∞].
  • Measurability on the full-measure set Z0Z_0Z0​ uses the trace σ\sigmaσ-field.
  • The selections in the goal are total, Fν\mathcal F^\nuFν-measurable maps Z→RnZ\to\mathbb R^nZ→Rn that select almost surely. This is equivalent to the paper's maps Z0→RnZ_0\to\mathbb R^nZ0​→Rn.
  • "Random l.s.c. function" in Theorem 3.8 is encoded by the equivalent conditions (3.4i)–(3.4ii): nonempty, closed and measurable epigraphs.
  • Lower Lipschitz (3.10) is written additively.
  • The hypothesis that Ξ\XiΞ is the support of PPP is omitted. It is unused in §3, and omitting it strengthens every statement.
  • Nothing beyond the page is assumed: no convexity, no compact SSS, no bounded fff, no unique minimizer, no i.i.d. sampling, no empirical PνP^\nuPν, no completeness of μ\muμ or Fν\mathcal F^\nuFν.

The hypotheses are not vacuous. A sorry-free check verifies all of them, including those of the goal, for f(x,ξ)=∥x∥2f(x,\xi)=\|x\|^2f(x,ξ)=∥x∥2 with Dirac measures. Defining the expectation through a Bochner integral, or dropping S≠∅S\neq\emptysetS=∅ (which makes every argmin⁡\operatorname{argmin}argmin all of Rn\mathbb R^nRn), would trivialize or change the statements; the definitions above rule both out.

Contributions are welcome on any milestone. Proposition 3.1 (Kuratowski–Ryll-Nardzewski for Rm\mathbb R^mRm-valued multifunctions) and Proposition 3.3 (deterministic epi-convergence facts) are independent of the probabilistic setting and reusable beyond this mission.

Selected references

  • J. Dupačová, R. Wets, Asymptotic Behavior of Statistical Estimators and Optimal Solutions for Stochastic Optimization Problems, IIASA Working Paper WP-86-41, 1986. https://pure.iiasa.ac.at/id/eprint/2818/ — journal version: Ann. Statist. 16(4), 1517–1549, 1988. https://doi.org/10.1214/aos/1176351052
  • A. Wald, Note on the consistency of the maximum likelihood estimate, Ann. Math. Statist. 20, 595–601, 1949. https://doi.org/10.1214/aoms/1177729938
  • P. J. Huber, The behavior of maximum likelihood estimates under nonstandard conditions, Proc. Fifth Berkeley Symp. Math. Statist. Probab. 1, 221–233, 1967. https://projecteuclid.org/euclid.bsmsp/1200512988
  • R. T. Rockafellar, R. J.-B. Wets, Variational Analysis, Springer, 1998 (Ch. 7 epi-convergence; Ch. 14 measurable multifunctions and normal integrands). https://doi.org/10.1007/978-3-642-02431-3
13 thms1 active userReviewed
Discrete GeometryLinear OptimizationOperations Research+1·Captain: mikedeng1

Smoothed Analysis of Algorithms: Why the Simplex Algorithm Usually Takes Polynomial Time 1: The Expected Shadow of a Gaussian-Perturbed Polytope Has Polynomially Many VerticesResearch Paper

Why the shadow of a perturbed polytope matters

The simplex method solves linear programs very fast in practice, yet for most pivot rules there are inputs on which it takes exponentially many steps (Klee and Minty, 1972, for Dantzig's rule; Goldfarb, 1983, for the shadow-vertex rule). Average-case analyses (Borgwardt, 1980s; Smale, 1983) explained good behaviour on random inputs, but random inputs look nothing like real ones. Spielman and Teng introduced smoothed analysis to close this gap: the input is chosen by an adversary and then perturbed by a small Gaussian, and the running time is measured in expectation over the perturbation. They proved that the shadow-vertex simplex method has smoothed complexity polynomial in the number of constraints nnn, the dimension ddd and 1/σ1/\sigma1/σ (Spielman–Teng, J. ACM 2004; this mission follows the preprint arXiv:cs/0111050v7). The work received the Gödel Prize (2008) and the Fulkerson Prize (2009).

Timeline. Borgwardt (1977–1987) bounded the expected number of shadow-vertex pivots for rotationally symmetric random data. Spielman and Teng (2001, STOC; journal 2004) proved the first smoothed bound, with a shadow bound of order nd3/σ6nd^3/\sigma^6nd3/σ6 — the theorem of this mission. Deshpande and Spielman (FOCS 2005) improved the shadow bound, Vershynin (2009) reduced the dependence on nnn to polylogarithmic, and Dadush and Huiberts (STOC 2018) obtained O(d2log⁡n σ−2)O(d^2\sqrt{\log n}\,\sigma^{-2})O(d2logn​σ−2) for small σ\sigmaσ.

Setting

Fix d≥3d\ge3d≥3 and n>dn>dn>d. The data are vectors a1,…,an∈Rda_1,\dots,a_n\in\mathbb R^da1​,…,an​∈Rd, the constraint vectors of the linear program max⁡⟨z∣x⟩\max\langle z|x\ranglemax⟨z∣x⟩ subject to ⟨ai∣x⟩≤1\langle a_i|x\rangle\le1⟨ai​∣x⟩≤1 for all iii. Each aia_iai​ is a Gaussian of standard deviation σ\sigmaσ centered at a point aˉi\bar a_iaˉi​ with ∥aˉi∥≤1\|\bar a_i\|\le1∥aˉi​∥≤1: it has density

μi(a)=(12π σ)de−∥a−aˉi∥2/2σ2,\mu_i(a)=\Big(\tfrac{1}{\sqrt{2\pi}\,\sigma}\Big)^d e^{-\|a-\bar a_i\|^2/2\sigma^2},μi​(a)=(2π​σ1​)de−∥a−aˉi​∥2/2σ2,

and the aia_iai​ are independent (joint density ∏iμi(ai)\prod_i\mu_i(a_i)∏i​μi​(ai​)).

For a direction q∈Rdq\in\mathbb R^dq∈Rd, optSimpq(a1,…,an)\mathrm{optSimp}_q(a_1,\dots,a_n)optSimpq​(a1​,…,an​) is the set of index sets I⊆{1,…,n}I\subseteq\{1,\dots,n\}I⊆{1,…,n} with ∣I∣=d|I|=d∣I∣=d such that (ai)i∈I(a_i)_{i\in I}(ai​)i∈I​ is linearly independent, the simplex △(AI)=ConvHull(ai:i∈I)\triangle(A_I)=\mathrm{ConvHull}(a_i:i\in I)△(AI​)=ConvHull(ai​:i∈I) is a facet of ConvHull(0,a1,…,an)\mathrm{ConvHull}(0,a_1,\dots,a_n)ConvHull(0,a1​,…,an​), and qqq lies in the cone {∑i∈Iαiai:αi≥0}\{\sum_{i\in I}\alpha_ia_i:\alpha_i\ge0\}{∑i∈I​αi​ai​:αi​≥0}. In polar terms, III is the set of tight constraints at the vertex of the feasible polyhedron that maximizes ⟨q∣x⟩\langle q|x\rangle⟨q∣x⟩.

For linearly independent t,zt,zt,z, the shadow Shadowt,z(a1,…,an)\mathrm{Shadow}_{t,z}(a_1,\dots,a_n)Shadowt,z​(a1​,…,an​) is the set of index sets III that belong to optSimpq\mathrm{optSimp}_qoptSimpq​ for some nonzero q∈Span(t,z)q\in\mathrm{Span}(t,z)q∈Span(t,z). Its size is the number of vertices of the projection of the feasible polyhedron onto the plane Span(t,z)\mathrm{Span}(t,z)Span(t,z); the shadow-vertex method walks along this polygon, one pivot per vertex. Finally

D(n,d,σ)=58,888,678 nd3min⁡(σ, 1/(3dln⁡n))6.\mathcal D(n,d,\sigma)=\frac{58{,}888{,}678\,nd^3}{\min\big(\sigma,\,1/(3\sqrt{d\ln n})\big)^6}.D(n,d,σ)=min(σ,1/(3dlnn​))658,888,678nd3​.

Formalization targets

Goal: Theorem 4.0.1 (Shadow Size)

Ea1,…,an[ ∣Shadowt,z(a1,…,an)∣ ]≤D(n,d,σ)\mathbb E_{a_1,\dots,a_n}\big[\,|\mathrm{Shadow}_{t,z}(a_1,\dots,a_n)|\,\big]\le\mathcal D(n,d,\sigma)Ea1​,…,an​​[∣Shadowt,z​(a1​,…,an​)∣]≤D(n,d,σ)

for every d≥3d\ge3d≥3, n>dn>dn>d, every pair of linearly independent t,zt,zt,z, every σ>0\sigma>0σ>0 and all centers of norm at most 111.

Milestones

The milestones follow the paper's proof, leaves first.

  • Probability tools: the chi-square bound (Corollary 2.4.6), the combination lemma (Lemma 2.3.5), almost polynomial densities (Lemma 2.3.7), and comparing Gaussian tails (Lemma 2.4.11).
  • Reduction: the measure of the event P={∥ai∥≤2 ∀i}P=\{\|a_i\|\le2\ \forall i\}P={∥ai​∥≤2 ∀i} (Proposition 4.0.5), and the discretization of the shadow into mmm equally spaced directions (Lemma 4.0.6).
  • Angle bound: the probability, conditioned on PPP, that the ray through a fixed unit vector qqq passes within angle ε\varepsilonε of the boundary of its optimal facet is O(nd3ε/σ6)O(nd^3\varepsilon/\sigma^6)O(nd3ε/σ6) (Lemma 4.0.7, from Lemma 4.0.11).
  • Distance and incidence: in Blaschke coordinates ai=Rωbi+sqa_i=R_\omega b_i+sqai​=Rω​bi​+sq, a deterministic split (Lemma 4.0.12), a distance bound (Lemmas 4.1.1–4.1.3) and an angle-of-incidence bound (Lemmas 4.2.1–4.2.3).

Significance

The result. Theorem 4.0.1 is the geometric heart of the smoothed analysis of the simplex method. Section 4.3 of the paper extends it to arbitrary centers, covariances and right-hand sides, and Section 5 combines these extensions with a two-phase method to show that the simplex method has polynomial smoothed complexity. The same shadow bound underlies later analyses of the simplex method, of perturbed polytopes' diameters, and of condition numbers of random linear programs.

Formalizing it. The theorem has been proved, and improved constants are known, but none of this is machine-checked. A formal proof would verify a long and delicate argument: a change of variables of integral geometry (Blaschke's formula), several conditional-density estimates, and explicit constants in the millions. The mission also produces reusable statements about Gaussian vectors and convex hulls of random points.

Difficulty

The obvious approach is to count, for each candidate facet III, the probability that III appears in the shadow; there are (nd)\binom nd(dn​) candidates, so a union bound is exponential in ddd. The paper avoids this by discretizing the angle of qqq (Lemma 4.0.6) and bounding, for each fixed direction, the probability that the optimal facet changes within a small angular step. That needs a lower bound on the angle between qqq and the boundary of its optimal facet, conditioned on the facet being optimal. The conditioning changes the distribution of a1,…,ada_1,\dots,a_da1​,…,ad​, so the bound cannot come from the Gaussian density alone. The proof changes variables to the facet's normal ω\omegaω, offset sss and in-plane coordinates bib_ibi​ (Corollary 2.5.3), whose Jacobian contributes the factors ⟨ω∣q⟩\langle\omega|q\rangle⟨ω∣q⟩ and Vol(△(b))\mathrm{Vol}(\triangle(b))Vol(△(b)). It then shows that both the distance of the origin to a face of the in-plane simplex and the angle of incidence ⟨ω∣q⟩\langle\omega|q\rangle⟨ω∣q⟩ are unlikely to be small. Measure-theoretic bookkeeping is as hard as the geometry: densities known only up to normalization, conditioning on events of positive measure, and the measure-zero degeneracies the paper sets aside.

Formalization scope

Points live in EuclideanSpace ℝ (Fin d). Constraint vectors are indexed by Fin n (0-based), so the paper's {1,…,d}\{1,\dots,d\}{1,…,d} is {i:i<d}\{i:i<d\}{i:i<d}. The Gaussian of standard deviation σ\sigmaσ centered at ccc is Lebesgue measure with the density above, and the joint law is the product measure. Lemma 4.0.6 also uses Mathlib's multivariateGaussian with a positive definite covariance. Expectations of shadow sizes are lower Lebesgue integrals of [0,∞][0,\infty][0,∞]-valued counts, and their measurability is part of each conclusion. "Density proportional to ν\nuν" and conditional probabilities are stated cross-multiplied, ∫Eν≤bound⋅∫ν\int_{E}\nu\le\text{bound}\cdot\int\nu∫E​ν≤bound⋅∫ν, so no 0/00/00/0 appears.

The shadow is the set of index sets III, and the direction q=0q=0q=0 is excluded. Including it would add every facet of ConvHull(0,a1,…,an)\mathrm{ConvHull}(0,a_1,\dots,a_n)ConvHull(0,a1​,…,an​) to the shadow, since 000 lies in every cone, and make the goal false. ang(q,∅)=∞\mathrm{ang}(q,\emptyset)=\inftyang(q,∅)=∞ is represented exactly in [0,∞][0,\infty][0,∞], never by a real infimum. Where the paper omits a hypothesis it uses, it is added and recorded in the item: the standing assumptions d≥3d\ge3d≥3, n>dn>dn>d and σ≤1/(3dln⁡n)\sigma\le1/(3\sqrt{d\ln n})σ≤1/(3dlnn​) (Lemma 4.2.3 is false without a bound on σ\sigmaσ), unit length of the reference vector qqq, s≥0s\ge0s≥0, and ε>0\varepsilon>0ε>0 for strict inequalities. Lemma 2.3.7 is stated with ≤\le≤ rather than the page's <<<, which fails in an edge case.

Infrastructure a complete development needs: Gaussian tail and chi-square estimates; faces and facets of convex hulls; the Blaschke change of variables and the latitude–longitude change of variables on the sphere (not in Mathlib); surface measure on Sd−1S^{d-1}Sd−1 (Mathlib's Measure.toSphere); and the disintegration of the joint law used in the combination lemma. The Gaussian estimates, the combination lemma and the Blaschke formula are useful beyond this mission. Proofs of any milestone, and of supporting lemmas such as the change-of-variables formulas, are welcome.

Selected references

  • D. A. Spielman, S.-H. Teng, Smoothed Analysis of Algorithms: Why the Simplex Algorithm Usually Takes Polynomial Time, arXiv:cs/0111050v7, 2003. https://arxiv.org/abs/cs/0111050v7
  • D. A. Spielman, S.-H. Teng, Smoothed analysis of algorithms: Why the simplex algorithm usually takes polynomial time, J. ACM 51(3):385–463, 2004. https://doi.org/10.1145/990308.990310
  • K. H. Borgwardt, The Simplex Method: A Probabilistic Analysis, Springer, 1987.
  • V. Klee, G. J. Minty, How good is the simplex algorithm?, in Inequalities III, Academic Press, 1972, 159–175.
  • A. Deshpande, D. A. Spielman, Improved smoothed analysis of the shadow vertex simplex method, FOCS 2005, 387–396.
  • R. Vershynin, Beyond Hirsch conjecture: walks on random polytopes and smoothed complexity of the simplex method, SIAM J. Comput. 39(2):646–678, 2009. https://doi.org/10.1137/070683386
  • D. Dadush, S. Huiberts, A friendly smoothed analysis of the simplex method, STOC 2018; arXiv:1711.05667. https://arxiv.org/abs/1711.05667
29 thms1 active userReviewed
Operations ResearchStochastic Systems·Captain: mikedeng1

Open Queueing Networks in Heavy Traffic: Reflected Brownian Motion Limit for the Queue Length ProcessResearch Paper

Motivation

Open networks of single-server queues with general interarrival and service distributions are the standard model of job shops, communication networks and service systems. Outside the product-form (Jackson) case their queue-length distributions are not known in closed form. When every station is close to saturation, a heavy-traffic limit replaces the network by a diffusion process. Martin I. Reiman's paper Open Queueing Networks in Heavy Traffic (Mathematics of Operations Research 9(3), 1984) proves such a limit for the vector of queue lengths of a general open network. The limit is a reflected Brownian motion on the nonnegative orthant. That process has since become the default diffusion approximation for open networks, and it is the starting point of later work on its stationary distribution and on control of networks in heavy traffic.

Timeline:

  • Iglehart and Whitt (1970a,b) proved heavy-traffic limits for a single multiple-server station and for acyclic networks, in which no customer visits a station twice.
  • Harrison (1973, 1978) treated tandem queues; the 1978 paper introduced reflected Brownian motion on the nonnegative orthant as the diffusion limit.
  • Harrison and Reiman (1981a, Ann. Probab. 9:302–308) constructed reflected Brownian motion on the orthant through a continuous reflection mapping. That paper is the source of Lemma 1 here.

(These attributions follow Reiman's own account, pp. 441–442 of the 1984 paper.)

  • Reiman (1984) proved the limit for general open networks with Markovian routing (Theorem 1). The paper also proves a limit for sojourn times along fixed routes (Theorem 2).

Setting

There are KKK single-server stations and a nonempty set J⊆{1,…,K}\mathcal J\subseteq\{1,\dots,K\}J⊆{1,…,K} of stations that receive customers from outside. The primitives are mutually independent sequences of IID random variables: interarrival times uki>0u_k^i>0uki​>0 (k∈Jk\in\mathcal Jk∈J), service times vki>0v_k^i>0vki​>0, and routing indicators ϕki∈{0,1,…,K}\phi_k^i\in\{0,1,\dots,K\}ϕki​∈{0,1,…,K}. When the iiith customer served at station kkk finishes, it moves to station ϕki\phi_k^iϕki​, or leaves if ϕki=0\phi_k^i=0ϕki​=0. The parameters are the service rates μk=(Evk1)−1\mu_k=(E v_k^1)^{-1}μk​=(Evk1​)−1, the service-time variances sk=var⁡vk1s_k=\operatorname{var} v_k^1sk​=varvk1​, the arrival rates λk=(Euk1)−1\lambda_k=(E u_k^1)^{-1}λk​=(Euk1​)−1 (with λk=0\lambda_k=0λk​=0 for k∉Jk\notin\mathcal Jk∈/J), and the interarrival variances ak=var⁡uk1a_k=\operatorname{var} u_k^1ak​=varuk1​. The routing matrix P=(pkj)P=(p_{kj})P=(pkj​), pkj=P{ϕk1=j}p_{kj}=P\{\phi_k^1=j\}pkj​=P{ϕk1​=j}, has spectral radius strictly less than one, so every customer eventually leaves.

Let Ak(t)A_k(t)Ak​(t) be the number of exogenous arrivals to station kkk by time ttt, and Sk(t)S_k(t)Sk​(t) the number of service completions at kkk in ttt units of busy time. Let S^k(t)=∑i≤Sk(t)eϕki−Sk(t)ek\hat S_k(t)=\sum_{i\le S_k(t)}e_{\phi_k^i}-S_k(t)e_kS^k​(t)=∑i≤Sk​(t)​eϕki​​−Sk​(t)ek​, with e0=0e_0=0e0​=0. The queue length Q(t)∈Z+KQ(t)\in\mathbb Z_+^KQ(t)∈Z+K​ and the busy time B(t)B(t)B(t) are the unique solution of

Q(t)=A(t)+∑k=1KS^k(Bk(t)),Bk(t)=∫0t1{Qk(s)>0} ds,B(0)=0.Q(t)=A(t)+\sum_{k=1}^K\hat S_k(B_k(t)),\qquad B_k(t)=\int_0^t1_{\{Q_k(s)>0\}}\,ds,\qquad B(0)=0 .Q(t)=A(t)+k=1∑K​S^k​(Bk​(t)),Bk​(t)=∫0t​1{Qk​(s)>0}​ds,B(0)=0.

A sequence of such networks, indexed by nnn, shares KKK, J\mathcal JJ and PPP. Its parameters μ(n),s(n),λ(n),a(n)\mu(n),s(n),\lambda(n),a(n)μ(n),s(n),λ(n),a(n) converge to finite limits μ,s,λ,a\mu,s,\lambda,aμ,s,λ,a. With ν(n)=λ(n)+μ(n)P\nu(n)=\lambda(n)+\mu(n)Pν(n)=λ(n)+μ(n)P, the heavy-traffic condition is

ck(n)=n (νk(n)−μk(n))→ck.c_k(n)=\sqrt n\,(\nu_k(n)-\mu_k(n))\to c_k .ck​(n)=n​(νk​(n)−μk​(n))→ck​.

Moments of order 2+ϵ2+\epsilon2+ϵ of the interarrival and service times are bounded uniformly in nnn. The scaled queue length is Zn(t)=n−1/2Qn(nt)Z^n(t)=n^{-1/2}Q^n(nt)Zn(t)=n−1/2Qn(nt), 0≤t≤10\le t\le10≤t≤1.

Formalization targets

Goal: Theorem 1

Let ξ\xiξ be a Brownian motion with drift ccc and covariance matrix A\mathcal AA, where

Aii=λi3ai+μi3si(1−2pii)+∑jμjpji(1−pji+pjiμj2sj),\mathcal A_{ii}=\lambda_i^3a_i+\mu_i^3s_i(1-2p_{ii})+\sum_j\mu_jp_{ji}(1-p_{ji}+p_{ji}\mu_j^2s_j),Aii​=λi3​ai​+μi3​si​(1−2pii​)+j∑​μj​pji​(1−pji​+pji​μj2​sj​), Aij=−[μi3sipij+μj3sjpji+∑kμkpkipkj(1−μk2sk)](i≠j).\mathcal A_{ij}=-\Big[\mu_i^3s_ip_{ij}+\mu_j^3s_jp_{ji}+\sum_k\mu_kp_{ki}p_{kj}(1-\mu_k^2s_k)\Big]\quad(i\ne j).Aij​=−[μi3​si​pij​+μj3​sj​pji​+k∑​μk​pki​pkj​(1−μk2​sk​)](i=j).

Let Z=ϕ(ξ)Z=\phi(\xi)Z=ϕ(ξ) be its reflection with reflection matrix I−PI-PI−P. Then

Zn⇒Zin D[0,1] (Skorohod topology).Z^n\Rightarrow Z\quad\text{in } D[0,1]\text{ (Skorohod topology)}.Zn⇒Zin D[0,1] (Skorohod topology).

The goal fixes no constants beyond the parameters' limits. It is stated for every network sequence satisfying (20)–(26).

Milestones

The milestones follow the paper's proof, in order:

  • the existence and uniqueness claim for (1)–(3);
  • the representation Q=X~+Y(I−P)Q=\tilde X+Y(I-P)Q=X~+Y(I−P) (Eq. (13));
  • the least-element map fff (Proposition 1);
  • the reflection mapping ϕ\phiϕ (Lemma 1) and f=ϕf=\phif=ϕ on continuous paths (Proposition 2);
  • the netput limit ζn⇒ζ\zeta^n\Rightarrow\zetaζn⇒ζ (Proposition 3);
  • stochastic boundedness of ZnZ^nZn (Lemma 6);
  • vanishing scaled idleness n−1Ikn(n)→0n^{-1}I^n_k(n)\to0n−1Ikn​(n)→0 (Proposition 4);
  • the centred limit ζ~n⇒ζ\tilde\zeta^n\Rightarrow\zetaζ~​n⇒ζ (Proposition 5).

Significance

Theorem 1 justifies the diffusion approximation of a heavily loaded open network. Writing Qn(t)≈n Z(t/n)Q^n(t)\approx\sqrt n\,Z(t/n)Qn(t)≈n​Z(t/n) reduces questions about the network to questions about one reflected Brownian motion, whose data are explicit functions of the first two moments of the primitives and of the routing matrix. The same limit, with Lemma 2, gives the paper's Theorem 2 on sojourn times. It is the model case for the multiclass heavy-traffic theory that followed.

The result has been proved since 1984. No machine-checked version exists. The mission's contributions would be:

  • a formal statement of the network, of its Harrison representation, and of weak convergence in DDD;
  • a formal proof of the reflection-mapping facts (Proposition 1, Lemma 1, Proposition 2), which are deterministic and reusable;
  • eventually, a formal proof of the full limit theorem.

Difficulty

The obvious route applies a functional central limit theorem to QnQ^nQn directly. That fails because QnQ^nQn is not a sum of independent terms: each station serves only while its queue is nonempty, so the service process is evaluated at the random busy time Bk(t)B_k(t)Bk​(t), which depends on the whole network. The proof therefore has to separate the netput process, which obeys a central limit theorem, from the regulator YYY. It then has to show that the random time change Bkn(nt)/nB^n_k(nt)/nBkn​(nt)/n converges to the identity, i.e. that idleness vanishes on the diffusion scale. Weak convergence must also be transported through a reflection map that is defined on all of DDD but is known to be continuous only at continuous paths.

Formalization scope

The Lean development uses the following conventions:

  • Stations are Fin K, vectors are row vectors Fin K → ℝ, and a row vector times a matrix is Matrix.vecMul.
  • A routing indicator lives in Fin (K+1), with 0 meaning "leaves" and j.succ meaning station jjj.
  • The primitives are mutually independent (iIndep of their σ-algebras), IID within each sequence, everywhere positive and square integrable.
  • "Spectral radius <1<1<1" is stated as Pm→0P^m\to0Pm→0.
  • (Qn,Bn)(Q^n,B^n)(Qn,Bn) is any pair solving (1)–(3) almost surely, with measurable paths so that (2) is a Lebesgue integral.
  • The networks are indexed by ℕ; (25)–(26) are imposed for n≥1n\ge1n≥1, (22) and (26) over k∈Jk\in\mathcal Jk∈J, and J\mathcal JJ is the same for all nnn.
  • Brownian motion with drift ccc and covariance A\mathcal AA lives on [0,∞)[0,\infty)[0,∞). It is defined by continuity, ξ(0)=0\xi(0)=0ξ(0)=0, independent increments, and the Gaussian characteristic function of increments.
  • ZZZ is the reflection of ξ\xiξ in the sense of (14)–(17).
  • Weak convergence in DDD is stated in Skorohod-representation form: a coupling with almost-sure J1_11​ convergence on [0,1][0,1][0,1]. This form accommodates a separate probability space for each nnn.

Added hypotheses, each implicit on the page:

  1. The existence item assumes Uk(l),Vk(l)→∞U_k(l),V_k(l)\to\inftyUk​(l),Vk​(l)→∞ at the sample point; without it the maxima defining Ak(t)A_k(t)Ak​(t) and Sk(t)S_k(t)Sk​(t) need not exist.
  2. Solutions of (1)–(3) have measurable paths.

No positivity hypothesis on the limits μk\mu_kμk​ is added: (25) and (26) bound the means of the service and interarrival times, so the limits are positive.

The statement is not to be weakened. Ruled out are:

  • convergence of finite-dimensional distributions only;
  • a single network without the index nnn;
  • uniform convergence used in place of the Skorohod topology without the coupling;
  • a Brownian motion that is not required to have independent Gaussian increments.

Each of these is a different theorem.

Useful contributions, all reusable beyond this mission:

  • the deterministic reflection-map results;
  • Donsker-type theorems for renewal counting processes in DDD;
  • the random time-change lemma (Billingsley);
  • the continuous mapping theorem in coupling form.

Selected references

  • M. I. Reiman, Open Queueing Networks in Heavy Traffic, Mathematics of Operations Research 9(3):441–458, 1984. https://doi.org/10.1287/moor.9.3.441
  • J. M. Harrison and M. I. Reiman, Reflected Brownian Motion on an Orthant, Annals of Probability 9:302–308, 1981 (cited in Reiman 1984 as [6]).
  • J. M. Harrison, The Diffusion Approximation for Tandem Queues in Heavy Traffic, Advances in Applied Probability 10:886–905, 1978 (Reiman 1984, [5]).
  • J. M. Harrison, The Heavy Traffic Approximation for Single Server Queues in Series, Journal of Applied Probability 10:613–629, 1973 (Reiman 1984, [4]).
  • D. L. Iglehart and W. Whitt, Multiple Channel Queues in Heavy Traffic, I and II: Sequences, Networks, and Batches, Advances in Applied Probability 2:150–177 and 355–364, 1970 (Reiman 1984, [8], [9]).
  • P. Billingsley, Convergence of Probability Measures, Wiley, New York, 1968 (Reiman 1984, [1]).
13 thms1 active userReviewed
Algorithmic Game TheoryMechanism DesignOperations Research·Captain: mikedeng1

Multi-parameter Mechanism Design and Sequential Posted Pricing 4: A 6.75-Approximate Truthful Posted-Price Menu for Unit-Demand Buyers of Multiple ItemsResearch Paper

Motivation

A hotel sells rooms of several types, in limited numbers, to guests who each want one room. The revenue-optimal way to sell is known only in special cases: for buyers with several private values, optimal mechanisms can be randomized, involve lotteries, and lack a closed form (Manelli–Vincent 2007; Chawla, Hartline, Kleinberg 2007). In practice sellers post prices. The question is how much revenue posting prices gives up.

Chawla, Hartline, Malec and Sivan (arXiv:0907.2435v2, STOC 2010) answer it for a broad class of single- and multi-parameter problems. For unit-demand buyers of multiple copies of multiple items they show that a menu of posted prices, offered to the buyers in whatever order they arrive, earns at least 1/6.751/6.751/6.75 of the revenue of any deterministic truthful mechanism (Theorem 14). This mission formalizes that result together with the two steps it is built from: a reduction from the multi-parameter problem to a single-parameter one with "copies" of each buyer (Lemma 3, Theorem 4), and an order-oblivious pricing for the intersection of two partition matroids (Theorem 13).

Setting

Single-parameter problem (BSMD). Finitely many agents iii have independent private values vi∼Fiv_i \sim F_ivi​∼Fi​, each with a density on a bounded interval. A seller may serve any set in a downward-closed set system J\mathcal JJ. A deterministic mechanism MMM maps reported values vvv to a served set M(v)∈JM(v) \in \mathcal JM(v)∈J and payments πi(v)\pi_i(v)πi​(v); it is truthful if reporting the true value is a dominant strategy and no agent ends with negative utility. Its expected revenue is RM=Ev[∑iπi(v)]\mathcal R^M = \mathbb E_v[\sum_i \pi_i(v)]RM=Ev​[∑i​πi​(v)]. For prices ppp, agent iii desires service if pi≤vip_i \le v_ipi​≤vi​, and Sv\mathcal S_vSv​ is the class of maximal feasible sets of desiring agents. The order-oblivious revenue is

Rpobl=Ev[min⁡S∈Sv∑i∈Spi],\mathcal R^{\mathrm{obl}}_{\mathbf p} = \mathbb E_{v}\Big[\min_{S \in \mathcal S_v} \sum_{i \in S} p_i\Big],Rpobl​=Ev​[S∈Sv​min​i∈S∑​pi​],

a lower bound on the revenue of posting the prices ppp to the agents in an adversarial order.

Multi-parameter unit-demand problem (BMUMD). There are mmm buyers and a finite set JJJ of services, partitioned into the groups JiJ_iJi​ of services targeted at buyer iii. Buyer iii has value vjv_jvj​ for each j∈Jij \in J_ij∈Ji​, all values independent with vj∼Fjv_j \sim F_jvj​∼Fj​, and the set system J⊆2J\mathcal J \subseteq 2^JJ⊆2J is unit-demand: ∣S∩Ji∣≤1|S \cap J_i| \le 1∣S∩Ji​∣≤1 for feasible SSS. A mechanism A\mathcal AA is truthful if no buyer gains by misreporting its whole vector (vj)j∈Ji(v_j)_{j \in J_i}(vj​)j∈Ji​​, and individually rational if a buyer receiving jjj pays at most vjv_jvj​ and a buyer receiving nothing pays 000.

Copies. The instance Icopies\mathcal I^{\mathrm{copies}}Icopies replaces each buyer iii by ∣Ji∣|J_i|∣Ji​∣ single-parameter agents, one per service j∈Jij \in J_ij∈Ji​ with value vjv_jvj​, under the same J\mathcal JJ.

Price menus. Given prices (pj)(p_j)(pj​) and an arrival order σ\sigmaσ, the price-menu mechanism approaches the buyers in order; buyer iii is offered the services of JiJ_iJi​ that can still be feasibly allocated, at prices pjp_jpj​, and buys a utility-maximizing one if some has pj≤vjp_j \le v_jpj​≤vj​.

Multiple copies of items. With items KKK and cap(k)\mathrm{cap}(k)cap(k) copies of item kkk, services are pairs (i,k)(i,k)(i,k) and a set of services is feasible if it gives each buyer at most one item and uses at most cap(k)\mathrm{cap}(k)cap(k) copies of kkk: the intersection of two partition matroids.

Formalization targets

Goal: Theorem 14

For regular distributions there are prices ppp such that, for every arrival order σ\sigmaσ, the price-menu mechanism Pσ\mathcal P_\sigmaPσ​ is truthful and

RA≤274 RPσ\mathcal R^{\mathcal A} \le \tfrac{27}{4}\,\mathcal R^{\mathcal P_\sigma}RA≤427​RPσ​

for every individually rational, truthful deterministic mechanism A\mathcal AA.

Milestones

  • Truthful BMUMD mechanisms are weakly monotone (p. 13), and the allocation of Acopies\mathcal A^{\mathrm{copies}}Acopies is monotone in each vjv_jvj​ (p. 13).
  • Lemma 3: RA≤RA′\mathcal R^{\mathcal A} \le \mathcal R^{\mathcal A'}RA≤RA′ for some truthful A′\mathcal A'A′ on Icopies\mathcal I^{\mathrm{copies}}Icopies.
  • The price-menu mechanism allocates a maximal feasible set of services (p. 14).
  • Theorem 4: if RM′≤α Rpobl\mathcal R^{M'} \le \alpha\,\mathcal R^{\mathrm{obl}}_{\mathbf p}RM′≤αRpobl​ for every truthful M′M'M′ on Icopies\mathcal I^{\mathrm{copies}}Icopies, then RA≤α RPσ\mathcal R^{\mathcal A} \le \alpha\,\mathcal R^{\mathcal P_\sigma}RA≤αRPσ​ for every σ\sigmaσ and every truthful IR A\mathcal AA.
  • Lemma 2 (regular part): RM≤∑ipiMqiM\mathcal R^M \le \sum_i p^M_i q^M_iRM≤∑i​piM​qiM​, with qiMq^M_iqiM​ the probability that MMM serves iii and Fi(piM)=1−qiMF_i(p^M_i) = 1 - q^M_iFi​(piM​)=1−qiM​.
  • Theorem 19 (existence form): a revenue-optimal truthful mechanism exists.
  • The claim ci≥4/9c_i \ge 4/9ci​≥4/9 of App. D.4: under ∑i′∈Pqi′≤cap(P)/3\sum_{i' \in P} q_{i'} \le \mathrm{cap}(P)/3∑i′∈P​qi′​≤cap(P)/3 in every part, with probability at least 4/94/94/9 neither part of iii is full without iii.
  • Theorem 13: for two partition matroids there are prices with RM≤274 Rpobl\mathcal R^M \le \tfrac{27}{4}\,\mathcal R^{\mathrm{obl}}_{\mathbf p}RM≤427​Rpobl​ for every truthful MMM.

Significance

The result shows that for unit-demand buyers, a seller loses at most a constant factor by replacing the optimal, possibly opaque, truthful mechanism with a menu of prices that does not depend on the order in which buyers arrive. The reduction of Theorem 4 is generic: any order-oblivious pricing for the single-parameter instance with copies, under any unit-demand constraint, transfers to the multi-parameter instance with the same factor. Theorem 13 supplies one such pricing for the intersection of two partition matroids, which is exactly the shape of the multi-unit, multi-item constraint.

All results here are proved in the paper and none is formalized elsewhere; the platform has Myerson's single-unit optimal auction and weak monotonicity in an abstract quasilinear model (Börgers), but no posted-price approximation, no copies reduction, and no order-oblivious revenue. The formal development adds a machine-checked account of the reduction (in particular that the price-menu mechanism is truthful and allocates a maximal feasible set for every order), a precise version of the probabilistic claim behind the constant 6.756.756.75, and reusable definitions of order-oblivious revenue and of multi-parameter truthfulness with the paper's individual rationality.

Difficulty

Lemma 3 needs more than the observation that the copies instance has more competition: one must build a truthful single-parameter mechanism with at least the same revenue. The allocation is copied, but the payments must be threshold payments of the copies mechanism, and showing they dominate the original payments uses both weak monotonicity and the paper's individual rationality, through the taxation principle.

Theorem 13 compares order-oblivious revenue with Myerson's revenue through the bound of Lemma 2, at prices built from Myerson's service probabilities scaled by 1/31/31/3. The step that is easy to get wrong is the probability that an agent is considered: the events "part P1P_1P1​ is not full" and "part P2P_2P2​ is not full" depend on overlapping agents, so the product bound (2/3)(2/3)(2/3)(2/3)(2/3)(2/3) does not follow from Markov's inequality alone; it holds because both events are decreasing in the set of desiring agents (Harris' inequality). The comparison must also be uniform: one set of prices must serve against every truthful mechanism, which requires an optimal mechanism to exist.

Formalization scope

  • Distributions (P1): each FjF_jFj​ has a measurable density, strictly positive on a bounded interval [v‾j,v‾j]⊆[0,∞)[\underline v_j, \overline v_j] \subseteq [0, \infty)[v​j​,vj​]⊆[0,∞), with no mass outside. Values are independent (product prior).
  • Regularity (P2): the virtual value ϕ(v)=v−(1−F(v))/f(v)\phi(v) = v - (1 - F(v))/f(v)ϕ(v)=v−(1−F(v))/f(v) is non-decreasing on the support. It is assumed in Lemma 2, Theorem 19, Theorem 13 and the goal. Theorem 14 does not state it, but its proof goes through Theorem 13, which the paper proves for regular distributions; the non-regular extension (App. E, randomized prices) is out of scope, as is the second paragraph of Lemma 2.
  • Mechanisms (P3): deterministic; dominant-strategy truthful with misreports in the support (a buyer misreports all coordinates of JiJ_iJi​ at once); single-parameter IR is ex-post nonnegative utility; multi-parameter IR is the paper's (πi≤vj\pi_i \le v_jπi​≤vj​ if served jjj, πi=0\pi_i = 0πi​=0 if unserved); allocation events and payments measurable, payments integrable.
  • Benchmarks (P4): Myerson's mechanism is not constructed. "Approximates RM\mathcal R^{\mathcal M}RM" is stated against every truthful mechanism, and Lemma 3 and Theorem 19 in existence form.
  • Price menus: ties between utility-maximizing services are broken by a fixed enumeration of JJJ; a service of utility 000 is bought. Theorem 4 assumes α≥0\alpha \ge 0α≥0.
  • Dropped: the last sentence of Theorem 14 (polynomial-time computability of the prices) has no cost model here.
  • Constant: 6.756.756.75 is written 27/427/427/4 everywhere.
  • Not trivializable: the prices in Theorem 13 and the goal are chosen before the mechanism, and the benchmark includes every truthful mechanism, so a degenerate price vector cannot meet the bound; Rpobl\mathcal R^{\mathrm{obl}}_{\mathbf p}Rpobl​ is a genuine minimum over a nonempty finite class.

Needed infrastructure, reusable beyond this mission: Myerson's characterization of truthful single-parameter mechanisms and the revenue–virtual-surplus identity for densities on intervals, Harris' inequality for product measures, and the taxation principle for deterministic multi-parameter mechanisms. Contributions on any of these are welcome.

Selected references

  • S. Chawla, J. D. Hartline, D. Malec, B. Sivan, Multi-parameter Mechanism Design and Sequential Posted Pricing, STOC 2010; arXiv:0907.2435v2, 2010. https://arxiv.org/abs/0907.2435
  • R. Myerson, Optimal Auction Design, Mathematics of Operations Research 6(1), 1981. https://doi.org/10.1287/moor.6.1.58
  • S. Chawla, J. D. Hartline, R. Kleinberg, Algorithmic Pricing via Virtual Valuations, EC 2007. https://arxiv.org/abs/0711.3203
  • A. M. Manelli, D. R. Vincent, Multidimensional mechanism design: Revenue maximization and the multiple-good monopoly, Journal of Economic Theory 137(1), 2007. https://doi.org/10.1016/j.jet.2006.12.007
  • T. E. Harris, A lower bound for the critical probability in a certain percolation process, Proc. Cambridge Philos. Soc. 56, 1960. https://doi.org/10.1017/S0305004100034241
15 thms1 active userReviewed
Machine LearningOptimizationStatistics·Captain: mikedeng1

Variance-based Regularization with Convex Objectives I: The χ²-Robust Risk Equals Empirical Risk plus a Standard-Deviation PenaltyResearch Paper

Motivation

Many statistical procedures minimize an average observed loss. This treats two candidates with the same average as equally attractive even when one has much more variable losses across the sample. Adding a multiple of the empirical standard deviation can distinguish them, but the resulting objective need not be convex even when each individual loss is convex. Duchi and Namkoong study a distributionally robust alternative: they maximize expected loss over a small neighborhood of the empirical distribution, then minimize that worst-case value. Their paper identifies when this convex robust value agrees exactly with the mean-plus-standard-deviation expression and how far apart the two can be otherwise. The finite-sample statement is Theorem 1 of the pinned preprint.

The relation matters to someone choosing a loss function for stochastic optimization. The variance expression has a direct statistical interpretation, while the robust expression preserves convexity in a decision parameter when the loss is convex. Theorem 1 makes the relationship quantitative for a single bounded random variable, before the paper turns to uniform guarantees over whole classes of losses. This mission isolates that first step and its finite optimization model.

Setting

Take observed real values z1,…,znz_1,\ldots,z_nz1​,…,zn​, with n≥1n\ge1n≥1. Their empirical mean and empirical variance are

zˉ=1n∑i=1nzi,sn2=1n∑i=1nzi2−zˉ2.\bar z=\frac1n\sum_{i=1}^n z_i,\qquad s_n^2=\frac1n\sum_{i=1}^n z_i^2-\bar z^2.zˉ=n1​i=1∑n​zi​,sn2​=n1​i=1∑n​zi2​−zˉ2.

The variance uses 1/n1/n1/n, not the unbiased-estimator factor 1/(n−1)1/(n-1)1/(n−1). A weight vector p=(p1,…,pn)p=(p_1,\ldots,p_n)p=(p1​,…,pn​) is feasible when its entries are nonnegative, sum to one, and satisfy

12∑i=1n(npi−1)2≤ρ,ρ≥0.\frac12\sum_{i=1}^n(np_i-1)^2\le\rho,\qquad \rho\ge0.21​i=1∑n​(npi​−1)2≤ρ,ρ≥0.

This is the paper's χ² neighborhood Pn(ρ)\mathcal P_n(\rho)Pn​(ρ) of the uniform empirical weights. Its robust sample expectation is

Rn(z,ρ)=sup⁡p∈Pn(ρ)∑i=1npizi.R_n(z,\rho)=\sup_{p\in\mathcal P_n(\rho)}\sum_{i=1}^n p_i z_i.Rn​(z,ρ)=p∈Pn​(ρ)sup​i=1∑n​pi​zi​.

For a random variable ZZZ with law PPP supported on [M0,M1][M_0,M_1][M0​,M1​], write M=M1−M0M=M_1-M_0M=M1​−M0​ and σ2=Var⁡P(Z)\sigma^2=\operatorname{Var}_P(Z)σ2=VarP​(Z). An independent sample Z1,…,ZnZ_1,\ldots,Z_nZ1​,…,Zn​ supplies the vector zzz. The paper describes Pn\mathcal P_nPn​ through a ϕ\phiϕ-divergence from the empirical distribution, with ϕ(t)=12(t−1)2\phi(t)=\tfrac12(t-1)^2ϕ(t)=21​(t−1)2; its finite maximization problem (8) is the weight-vector form used here. The preprint, pp. 2 and 5–7 fixes these conventions.

Formalization targets

Deterministic bound

For every sample in [M0,M1][M_0,M_1][M0​,M1​], the robust value lies between the empirical mean plus a corrected variance penalty and the full penalty:

(2ρsn2n−2Mρn)+≤Rn(z,ρ)−zˉ≤2ρsn2n.\left(\sqrt{\frac{2\rho s_n^2}{n}}-\frac{2M\rho}{n}\right)_+\le R_n(z,\rho)-\bar z\le\sqrt{\frac{2\rho s_n^2}{n}}.(n2ρsn2​​​−n2Mρ​)+​≤Rn​(z,ρ)−zˉ≤n2ρsn2​​​.

This is inequality (10). The correction is explicit, so this target records more than an asymptotic approximation.

Exact expansion

When σ2>0\sigma^2>0σ2>0 and the sample size obeys

n≥max⁡{5,M2σ2max⁡{8σ,44,44ρ}},n\ge\max\left\{5,\frac{M^2}{\sigma^2}\max\{8\sigma,44,44\rho\}\right\},n≥max{5,σ2M2​max{8σ,44,44ρ}},

the goal is the high-probability equality

Pr⁡{Rn(Z1:n,ρ)≠Zˉ+2ρsn2n}≤exp⁡(−nσ211M2).\Pr\left\{R_n(Z_{1:n},\rho)\ne\bar Z+\sqrt{\frac{2\rho s_n^2}{n}}\right\}\le\exp\left(-\frac{n\sigma^2}{11M^2}\right).Pr{Rn​(Z1:n​,ρ)=Zˉ+n2ρsn2​​​}≤exp(−11M2nσ2​).

This is Theorem 1's equality (11) with the missing ρ\rhoρ-dependent sample-size requirement supplied from the proof. The exact expansion is the mission goal; display (30), inequality (10), and Lemma A.2 form the milestone list, and the exact value under condition (9) is a further statement of the mission.

Significance

The deterministic result states how large the discrepancy between a convex robust risk and a variance penalty can be for any bounded sample. The equality says that, with the stated confidence, no discrepancy remains once the population variance and sample size make the penalty compatible with nonnegative probability weights. These are the numerical facts later sections need when they move from one loss variable to families of losses and minimizers. The claims and constants come from Theorem 1 and Section 2.1.

The paper develops arguments for these results, although its printed (11) needs the correction described below; the statements in this mission have no machine-checked proofs yet. The formalization work includes the finite χ² feasible set, its real supremum, exact handling of tied observations, empirical moments with the paper's normalization, and a product-law event for the probability estimate. The Samson concentration milestone is reusable for other bounded independent-coordinate models. Solvers can also contribute a different route to the corrected exact expansion; the goal concerns the statement, not one chosen argument.

Difficulty

Without the nonnegativity requirement on ppp, optimizing a linear function over the centered Euclidean ball gives the mean plus a standard-deviation term. The candidate weights can become negative when a sample coordinate is far below the mean, so that calculation alone cannot certify the robust value. Condition (9) records precisely when the candidate is feasible. The probability target then needs a quantitative guarantee that the sample variance is large enough often enough, with the stated exponential constant. A pointwise inequality for a fixed sample does not by itself yield that probability estimate. These are separate obligations in Section 2.1 and Appendix A.

Formalization scope

The sample is a function Fin n → ℝ; feasible weights have the same type. chiSqBall, robustSup, empMean, and empVar mirror equations (8) and the definitions on p. 6. Every theorem assumes n>0n>0n>0 and ρ≥0\rho\ge0ρ≥0, so the weight ball is nonempty and its real supremum is bounded. The high-probability theorem uses a probability measure PPP on the reals, supported on [M0,M1][M_0,M_1][M0​,M1​], and the independent product measure on Fin n → ℝ. Its conclusion bounds the measure of the event on which equality fails. The positive population variance hypothesis makes division by σ2\sigma^2σ2 and M2M^2M2 meaningful. The deterministic bounds include every sample in the interval and use x+=max⁡{x,0}x_+=\max\{x,0\}x+​=max{x,0}.

The paper prints the threshold without 44ρ44\rho44ρ in (11), but its Appendix A invokes the corresponding inequality, and the printed claim fails for sufficiently large ρ\rhoρ. The goal includes that term. The paper's route through Lemmas A.1 and A.4 contains misprinted lower-tail and moment claims, so those are not milestones. Lemma A.3's displayed (31b) is also omitted because its correction term has the wrong scaling; the corrected goal stands as a target to establish independently. These discrepancies are detailed in the local moderation notes and the pinned source, pp. 7 and 32–35.

No hypothesis may force the bad event to be empty, and the robust value must optimize over all feasible weights, not a selected optimizer. The supporting definitions are intended for reuse in later missions on uniform variance expansions. Contributions to the finite optimization facts, the concentration statement, and the probability goal are welcome.

Selected references

  • J. C. Duchi and H. Namkoong, Variance-based regularization with convex objectives, arXiv preprint arXiv:1610.02581v3, 2017. Pinned preprint.
8 thms1 active userReviewed
Dynamic ProgrammingOperations ResearchOptimization·Captain: mikedeng1

Optimal Policies for a Multi-Echelon Inventory Problem: The Two-Echelon Optimal Cost Splits into the Isolated Installation-1 Cost Plus a Function of Echelon StockResearch Paper

Motivation

Most physical supply chains hold stock at several levels: a factory warehouse feeds a regional depot, which feeds a retail outlet. Each level orders from the one above it, and a shortage upstream delays replenishment downstream. Optimizing such a multi-echelon system by dynamic programming looks hopeless, because the state is a vector of stock levels and stock in transit at every installation, and the value function of a two-installation system with a two-period shipping lag already depends on three continuous variables.

Andrew J. Clark and Herbert Scarf (Management Science 6(4):475–490, 1960) showed that for a serial system this curse of dimensionality disappears. Working with echelon stock (the stock at a level plus everything below it or in transit to a lower level), the optimal system cost separates into the cost of the lowest installation, optimized as if it stood alone, plus a function of echelon stock only. The result is the foundation of multi-echelon inventory theory: the echelon base-stock policies used in practice, the stationary analyses of Federgruen and Zipkin (1984) and Chen and Zheng (1994), and textbook treatments (Zipkin, Foundations of Inventory Management, 2000; Snyder and Shen, Fundamentals of Supply Chain Theory) all descend from it.

Timeline. Arrow, Harris and Marschak (1951) and Arrow, Karlin and Scarf (1958) set up periodic-review inventory models with discounted costs. Karlin and Scarf (1958) treated a single installation with a delivery lag, reducing it to a problem without lag (the paper's facts 1–3). Clark and Scarf (1960) proved the decomposition for serial systems with linear shipping costs and a setup cost permitted only at the top. Federgruen and Zipkin (1984) extended it to infinite horizons and Chen and Zheng (1994) gave a lower-bound proof that reaches more general structures.

Setting

Two installations are in series. Customer demand occurs only at installation 1; its demand in each period is non-negative with density φ\varphiφ on (0,∞)(0,\infty)(0,∞), independent across periods, and excess demand is backlogged. Installation 2 ships to installation 1 with a two-period lead time at unit cost c1≥0c_1\ge0c1​≥0. The system orders z≥0z\ge0z≥0 units from outside at cost c(z)=K+czc(z)=K+czc(z)=K+cz for z>0z>0z>0 and c(0)=0c(0)=0c(0)=0 (eq. (5)); these arrive at installation 2 one period later. Costs nnn periods ahead are discounted by αn\alpha^nαn, α≥0\alpha\ge0α≥0.

The state at the start of a period is (x1,w1,x2)(x_1,w_1,x_2)(x1​,w1​,x2​): x1x_1x1​ is the stock on hand at installation 1, w1w_1w1​ the stock that reaches installation 1 next period, and x2x_2x2​ the echelon-2 stock (on hand at both installations plus in transit), so x1+w1≤x2x_1+w_1\le x_2x1​+w1​≤x2​. Installation 1 pays the expected holding and shortage cost (1),

L(x)={hx+p∫x∞(t−x)φ(t) dt,x>0,p∫0∞(t−x)φ(t) dt,x≤0,L(x)=\begin{cases}hx+p\int_x^\infty(t-x)\varphi(t)\,dt,&x>0,\\ p\int_0^\infty(t-x)\varphi(t)\,dt,&x\le0,\end{cases}L(x)={hx+p∫x∞​(t−x)φ(t)dt,p∫0∞​(t−x)φ(t)dt,​x>0,x≤0,​

and echelon 2 pays a natural one-period cost L~(x2)\tilde L(x_2)L~(x2​) (Assumption 3).

With nnn periods remaining, the optimal system cost Cn(x1,w1,x2)C_n(x_1,w_1,x_2)Cn​(x1​,w1​,x2​) satisfies, with C0≡0C_0\equiv0C0​≡0,

Cn(x1,w1,x2)=min⁡x1+w1≤y≤x20≤z{c(z)+c1(y−x1−w1)+L~(x2)+L(x1)+α∫0∞Cn−1(x1+w1−t, y−x1−w1, x2+z−t)φ(t) dt}(14)C_n(x_1,w_1,x_2)=\min_{\substack{x_1+w_1\le y\le x_2\\0\le z}}\Big\{c(z)+c_1(y-x_1-w_1)+\tilde L(x_2)+L(x_1)+\alpha\int_0^\infty C_{n-1}(x_1+w_1-t,\,y-x_1-w_1,\,x_2+z-t)\varphi(t)\,dt\Big\}\qquad(14)Cn​(x1​,w1​,x2​)=x1​+w1​≤y≤x2​0≤z​min​{c(z)+c1​(y−x1​−w1​)+L~(x2​)+L(x1​)+α∫0∞​Cn−1​(x1​+w1​−t,y−x1​−w1​,x2​+z−t)φ(t)dt}(14)

where yyy is installation 1's target (stock on hand plus in transit after shipping). Installation 1 in isolation, buying at unit cost c1c_1c1​ with a two-period lag, has optimal cost C^n(x1,w1)\hat C_n(x_1,w_1)C^n​(x1​,w1​), C^0≡0\hat C_0\equiv0C^0​≡0:

C^n(x1,w1)=min⁡y≥x1+w1{c1(y−x1−w1)+L(x1)+α∫0∞C^n−1(x1+w1−t, y−x1−w1)φ(t) dt}.(15)\hat C_n(x_1,w_1)=\min_{y\ge x_1+w_1}\Big\{c_1(y-x_1-w_1)+L(x_1)+\alpha\int_0^\infty\hat C_{n-1}(x_1+w_1-t,\,y-x_1-w_1)\varphi(t)\,dt\Big\}.\qquad(15)C^n​(x1​,w1​)=y≥x1​+w1​min​{c1​(y−x1​−w1​)+L(x1​)+α∫0∞​C^n−1​(x1​+w1​−t,y−x1​−w1​)φ(t)dt}.(15)

In Lean these are ClarkScarf.Serial.Model.sysCost and isoCost; the expressions in braces are sysObj and isoObj, indexed by nnn for the problem with n+1n+1n+1 periods remaining.

Formalization targets

Goal: Theorem 1 (p. 482)

There are functions gng_ngn​ with g1=L~g_1=\tilde Lg1​=L~ such that, for all n≥1n\ge1n≥1 and x1+w1≤x2x_1+w_1\le x_2x1​+w1​≤x2​,

Cn(x1,w1,x2)=C^n(x1,w1)+gn(x2),(16)C_n(x_1,w_1,x_2)=\hat C_n(x_1,w_1)+g_n(x_2),\qquad(16)Cn​(x1​,w1​,x2​)=C^n​(x1​,w1​)+gn​(x2​),(16)

and installation 1 acts optimally by aiming at an isolated-optimal target y^\hat yy^​ and taking min⁡(x2,y^)\min(x_2,\hat y)min(x2​,y^​), as much as installation 2 can supply. The goal fixes no form for gng_ngn​ and needs no critical numbers.

Milestones

  1. Convexity of y↦α∫ ⁣ ⁣∫L(y−t1−t2)φ(t1)φ(t2)y\mapsto\alpha\int\!\!\int L(y-t_1-t_2)\varphi(t_1)\varphi(t_2)y↦α∫∫L(y−t1​−t2​)φ(t1​)φ(t2​) (§2 item 2, p. 478).
  2. The isolated decomposition C^n(x1,w1)=L(x1)+α∫0∞L(x1+w1−t)φ(t) dt+fn(x1+w1)\hat C_n(x_1,w_1)=L(x_1)+\alpha\int_0^\infty L(x_1+w_1-t)\varphi(t)\,dt+f_n(x_1+w_1)C^n​(x1​,w1​)=L(x1​)+α∫0∞​L(x1​+w1​−t)φ(t)dt+fn​(x1​+w1​) for n≥2n\ge2n≥2, with fnf_nfn​ of (7) (p. 480).
  3. Convexity of every fnf_nfn​ (§2 item 3, p. 478).
  4. Eqs. (18)–(19) (p. 483): the system cost when echelon-2 stock is above or below the isolated critical number xˉn\bar x_nxˉn​.
  5. Eqs. (21)–(25) (pp. 483–484): the shortfall cost Λn\Lambda_nΛn​ depends on x2x_2x2​ alone,
Λn(x2)=c1(x2−xˉn)+α2∫0∞ ⁣ ⁣∫0∞[L(x2−t−y)−L(xˉn−t−y)]φ(t)φ(y) dy dt+α∫0∞[fn−1(x2−t)−fn−1(xˉn−t)]φ(t) dt.\Lambda_n(x_2)=c_1(x_2-\bar x_n)+\alpha^2\int_0^\infty\!\!\int_0^\infty[L(x_2-t-y)-L(\bar x_n-t-y)]\varphi(t)\varphi(y)\,dy\,dt+\alpha\int_0^\infty[f_{n-1}(x_2-t)-f_{n-1}(\bar x_n-t)]\varphi(t)\,dt.Λn​(x2​)=c1​(x2​−xˉn​)+α2∫0∞​∫0∞​[L(x2​−t−y)−L(xˉn​−t−y)]φ(t)φ(y)dydt+α∫0∞​[fn−1​(x2​−t)−fn−1​(xˉn​−t)]φ(t)dt.
  1. Theorem 2 (p. 484), the explicit form: given critical numbers, gng_ngn​ is computed by (26), gn(x2)=min⁡z≥0{c(z)+L~(x2)+Λn(x2)+α∫gn−1(x2+z−t)φ(t) dt}g_n(x_2)=\min_{z\ge0}\{c(z)+\tilde L(x_2)+\Lambda_n(x_2)+\alpha\int g_{n-1}(x_2+z-t)\varphi(t)\,dt\}gn​(x2​)=minz≥0​{c(z)+L~(x2​)+Λn​(x2​)+α∫gn−1​(x2​+z−t)φ(t)dt}.

Significance

The result. Theorem 1 replaces one three-dimensional dynamic program by two one-dimensional ones. Installation 1 solves its own problem (15), whose solution is a critical-number policy, and echelon 2 solves a single-installation problem in x2x_2x2​ with one-period cost L~+Λn\tilde L+\Lambda_nL~+Λn​. When L~\tilde LL~ is convex the augmented cost is convex (the paper remarks this for Expression (10)), so the echelon-2 policy is of (S,s)(S,s)(S,s) type by Scarf's theorem, and the whole system runs on echelon base-stock rules. Every later serial-system result, finite or infinite horizon, uses this decomposition or its proof idea, and the "induced penalty" Λn\Lambda_nΛn​ is the prototype of the penalty functions used in the multi-echelon literature.

Formalizing it. The theorem is classical and proved, but no machine-checked version exists. The published platform items on Clark–Scarf are a stationary single-period decomposition with normal demand and a disproved infinite-horizon base-stock recursion, neither of which is this finite-horizon dynamic program. A formal development produces the value functions (14)–(15) with real infima and set integrals, the measurability and integrability of value functions defined by infima, the convexity propagation through the recursion (7), and the decomposition itself, which are reusable for any finite-horizon inventory recursion with lead times.

Difficulty

The obvious induction on nnn substitutes (16) into (14) and separates the minimizations over yyy and zzz. The separation is immediate; the hard step is that the constrained minimum over x1+w1≤y≤x2x_1+w_1\le y\le x_2x1​+w1​≤y≤x2​ differs from the unconstrained one by an amount that a priori depends on (x1,w1)(x_1,w_1)(x1​,w1​). Showing that it depends on x2x_2x2​ alone is the content of Theorem 1; nothing in the separation step itself rules out a dependence on (x1,w1)(x_1,w_1)(x1​,w1​). On the measure-theoretic side, every value function is defined by an infimum over an uncountable set and then integrated against φ\varphiφ. Its measurability and integrability are not automatic, and they must be established before any identity between integrals can be manipulated.

Formalization scope

Everything lives in ClarkScarf.Serial, one definition file Def_ClarkScarf_Serial_Model and seven theorem files. Conventions committed to:

  • The model is a structure Model whose fields carry the data and the standing hypotheses: h,p,α,c1,K,c≥0h,p,\alpha,c_1,K,c\ge0h,p,α,c1​,K,c≥0; φ≥0\varphi\ge0φ≥0 with ∫0∞φ=1\int_0^\infty\varphi=1∫0∞​φ=1; and two additions the page leaves implicit, disclosed in each statement: a finite demand mean (otherwise (1) is infinite for x≤0x\le0x≤0) and L~\tilde LL~ non-negative, continuous and of at most linear growth (Assumption 3 leaves L~\tilde LL~ unspecified; these make every expectation in (14) finite and measurable). No discount bound α<1\alpha<1α<1, no convexity of L~\tilde LL~, no K=0K=0K=0 and no sign condition on w1w_1w1​ is assumed.
  • Expectations are set integrals ∫(0,∞)F(t)φ(t) dt\int_{(0,\infty)}F(t)\varphi(t)\,dt∫(0,∞)​F(t)φ(t)dt; "Min" is a real infimum over a nonempty feasible set of a non-negative objective.
  • Every statement about CnC_nCn​ is restricted to the state domain x1+w1≤x2x_1+w_1\le x_2x1​+w1​≤x2​; outside it the feasible set of (14) is empty.
  • The horizon index counts periods remaining, C0≡C^0≡0C_0\equiv\hat C_0\equiv0C0​≡C^0​≡0, and fn≡0f_n\equiv0fn​≡0 for n≤2n\le2n≤2.

A formalization in which the feasible set of (14) is empty, in which the expectations are junk zeros of non-integrable integrands, or in which gng_ngn​ may depend on (x1,w1)(x_1,w_1)(x1​,w1​) would make (16) trivial; the domain restriction, the integrability conditions and the order ∃g ∀x1,w1,x2\exists g\,\forall x_1,w_1,x_2∃g∀x1​,w1​,x2​ rule these out. A sorry-free check (not part of the mission) verifies C1=L(x1)+L~(x2)C_1=L(x_1)+\tilde L(x_2)C1​=L(x1​)+L~(x2​) and C^1=L(x1)\hat C_1=L(x_1)C^1​=L(x1​) and exhibits a model with exponential demand satisfying all hypotheses.

Needed infrastructure: Fubini-type rearrangement of iterated set integrals against a density, integrability of functions of linear growth against a finite-mean density, convexity preserved under infimal projection u↦inf⁡y≥uu\mapsto\inf_{y\ge u}u↦infy≥u​ and under convolution with a density, and measurability of infimum-defined functions. Contributions of these general lemmas, of the base cases n=1,2n=1,2n=1,2, and of any milestone are welcome.

Selected references

  • A. J. Clark and H. Scarf, Optimal Policies for a Multi-Echelon Inventory Problem, Management Science 6(4):475–490, 1960. https://doi.org/10.1287/mnsc.6.4.475
  • S. Karlin and H. Scarf, Inventory Models of the Arrow-Harris-Marschak Type with Time Lag, in Arrow, Karlin, Scarf (eds.), Studies in the Mathematical Theory of Inventory and Production, Stanford University Press, 1958.
  • H. Scarf, The Optimality of (S, s) Policies in the Dynamic Inventory Problem, in Mathematical Methods in the Social Sciences, Stanford University Press, 1960.
  • A. Federgruen and P. Zipkin, Computational Issues in an Infinite-Horizon, Multiechelon Inventory Model, Operations Research 32(4):818–836, 1984. https://doi.org/10.1287/opre.32.4.818
  • F. Chen and Y.-S. Zheng, Lower Bounds for Multi-Echelon Stochastic Inventory Systems, Management Science 40(11):1426–1443, 1994. https://doi.org/10.1287/mnsc.40.11.1426
8 thms1 active userReviewed
Operations ResearchStochastic Systems·Captain: mikedeng1

Fundamentals of Queueing Theory IX: Kingman's Upper Bound on the G/G/1 Queue WaitTextbook

Motivation

The single-server queue with general independent interarrival and service times, the G/G/1 queue, is the basic model of a congested resource: a machine, a link, a checkout. For Markovian arrivals or services the mean wait has a closed form (the Pollaczek–Khintchine formula for M/G/1, the geometric law for G/M/1). For general distributions it has none, and the mean wait depends on the whole distributions of the interarrival and service times, not only on their moments. Capacity planning still needs numbers. Bounds that use only the first two moments are therefore the practical tool. They say how bad congestion can be for any queue with a given arrival rate, service rate and variabilities, and they become exact as the traffic intensity approaches one.

This mission formalizes Chapter 7, §7.1 of Gross, Shortle, Thompson and Harris, Fundamentals of Queueing Theory (4th ed., Wiley 2008, DOI 10.1002/9781118625651), together with the heavy-traffic Theorem 7.1 of §7.2.3.

Timeline. Lindley (Proc. Cambridge Philos. Soc., 1952) derived the recursion for successive waiting times and characterized the stationary law. Kingman (Proc. Cambridge Philos. Soc., 1961, 1962) proved the heavy-traffic exponential limit. In "Some inequalities for the queue GI/G/1" (Biometrika, 1962) he proved the two-moment upper bound. Marshall (1968) derived further moment relations and bounds (Operations Research, 1968). Marchal (Operations Research, 1978) gave the lower bound (7.14).

Setting

A G/G/1 queue is specified by two probability laws on [0,∞)[0,\infty)[0,∞): the law AAA of an interarrival time TTT and the law BBB of a service time SSS. Both have finite second moments, and

E[T]=1λ,E[S]=1μ,σA2=Var[T],σB2=Var[S],ρ=λμ.E[T] = \frac1\lambda,\quad E[S] = \frac1\mu,\quad \sigma_A^2 = \mathrm{Var}[T],\quad \sigma_B^2 = \mathrm{Var}[S],\quad \rho = \frac{\lambda}{\mu}.E[T]=λ1​,E[S]=μ1​,σA2​=Var[T],σB2​=Var[S],ρ=μλ​.

The pairs (S(n),T(n))(S^{(n)}, T^{(n)})(S(n),T(n)) are independent and identically distributed, and S(n)S^{(n)}S(n) is independent of T(n)T^{(n)}T(n). Customers are served first come, first served. The line delay Wq(n)W_q^{(n)}Wq(n)​ of the nnnth customer obeys Lindley's recursion

Wq(n+1)=max⁡(0, Wq(n)+U(n)),U(n)=S(n)−T(n),(7.1)W_q^{(n+1)} = \max\bigl(0,\ W_q^{(n)} + U^{(n)}\bigr), \qquad U^{(n)} = S^{(n)} - T^{(n)}, \tag{7.1}Wq(n+1)​=max(0, Wq(n)​+U(n)),U(n)=S(n)−T(n),(7.1)

and Wq(n)W_q^{(n)}Wq(n)​ is independent of (S(n),T(n))(S^{(n)}, T^{(n)})(S(n),T(n)). The idle gap X(n)=−min⁡(0,Wq(n)+U(n))X^{(n)} = -\min(0, W_q^{(n)} + U^{(n)})X(n)=−min(0,Wq(n)​+U(n)) is the time between the nnnth departure and the next start of service.

The queue is stationary when the law ν\nuν of Wq(n)W_q^{(n)}Wq(n)​ does not depend on nnn, that is, when one step of (7.1) maps ν\nuν to itself. The mean stationary line delay is Wq=E[Wq(n)]=∫w dν(w)W_q = E[W_q^{(n)}] = \int w\,d\nu(w)Wq​=E[Wq(n)​]=∫wdν(w). In the Lean development these objects are IsGG1Input A B lam mu, lindley, idleX, IsStationaryWaitLaw A B ν and meanWait ν, in the namespace QueueingFundamentals.Bounds.

Formalization targets

Goal: Kingman's upper bound (7.13)

For every stationary G/G/1 queue with ρ<1\rho < 1ρ<1, WqW_qWq​ is finite and

Wq≤λ(σA2+σB2)2(1−ρ).W_q \le \frac{\lambda(\sigma_A^2 + \sigma_B^2)}{2(1-\rho)}.Wq​≤2(1−ρ)λ(σA2​+σB2​)​.

Milestones

  • the idle-gap identity E[X]=−E[U]=1/λ−1/μE[X] = -E[U] = 1/\lambda - 1/\muE[X]=−E[U]=1/λ−1/μ (7.4), and the mean-wait formula (7.7)
Wq=E[X2]−E[U2]2E[U];W_q = \frac{E[X^2] - E[U^2]}{2E[U]};Wq​=2E[U]E[X2]−E[U2]​;
  • the variance of the interdeparture time D=S(n+1)+X(n)D = S^{(n+1)} + X^{(n)}D=S(n+1)+X(n) (7.12): Var[D]=2σB2+σA2−2Wq(1/λ−1/μ)\mathrm{Var}[D] = 2\sigma_B^2 + \sigma_A^2 - 2W_q(1/\lambda - 1/\mu)Var[D]=2σB2​+σA2​−2Wq​(1/λ−1/μ);
  • Marchal's lower bound (7.14), Wq≥(λ2σB2+ρ(ρ−2))/(2λ(1−ρ))W_q \ge (\lambda^2\sigma_B^2 + \rho(\rho-2))/(2\lambda(1-\rho))Wq​≥(λ2σB2​+ρ(ρ−2))/(2λ(1−ρ));
  • the distributional lower bound Wq≥r0W_q \ge r_0Wq​≥r0​, with r0r_0r0​ the unique nonnegative root of f(z)=z−∫−z∞[1−U(t)] dtf(z) = z - \int_{-z}^\infty [1 - U(t)]\,dtf(z)=z−∫−z∞​[1−U(t)]dt and U(t)U(t)U(t) the CDF of S−TS - TS−T ((7.15), (7.16));
  • the two-sided estimate (7.17), max⁡(0,r0,λ2σB2+ρ(ρ−2)2λ(1−ρ))≤Wq≤λ(σA2+σB2)2(1−ρ)\max\bigl(0, r_0, \tfrac{\lambda^2\sigma_B^2 + \rho(\rho-2)}{2\lambda(1-\rho)}\bigr) \le W_q \le \tfrac{\lambda(\sigma_A^2+\sigma_B^2)}{2(1-\rho)}max(0,r0​,2λ(1−ρ)λ2σB2​+ρ(ρ−2)​)≤Wq​≤2(1−ρ)λ(σA2​+σB2​)​;
  • Theorem 7.1 (heavy traffic): for a sequence of G/G/1 queues with ρj→1\rho_j \to 1ρj​→1, αj=−E[Sj−Tj]\alpha_j = -E[S_j - T_j]αj​=−E[Sj​−Tj​] and βj2=Var[Sj−Tj]\beta_j^2 = \mathrm{Var}[S_j - T_j]βj2​=Var[Sj​−Tj​], under convergence of the input laws, Var[S−T]>0\mathrm{Var}[S - T] > 0Var[S−T]>0 and uniformly bounded (2+δ)(2+\delta)(2+δ)-moments,
2αjβj2 Wq,j→dExp(1).\frac{2\alpha_j}{\beta_j^2}\,W_{q,j} \xrightarrow{d} \mathrm{Exp}(1).βj2​2αj​​Wq,j​d​Exp(1).

The goal is (7.13) rather than the stronger (7.17) because it depends only on the first two moments of the input.

Significance

Kingman's bound is the most widely used performance estimate for single-server queues. It needs no distributional form, only two means and two variances. It yields the "Kingman formula" approximation used across manufacturing and service operations, and it is asymptotically exact as ρ→1\rho \to 1ρ→1 (Theorem 7.1). The departure variance (7.12) drives the decomposition approximations for networks of §7.3. The heavy-traffic theorem is the entry point to diffusion approximations of queues.

All results here are proved in the literature (Theorem 7.1 is stated in the book without proof). This mission produces the first machine-checked versions. As far as a search of the platform shows, none of these statements, and no stationary Lindley recursion, has been formalized. The substrate it needs is reusable for any mission on G/G/1, G/G/c or random walks: stationary laws of a recursion on distributions, moment identities for max⁡(0,⋅)\max(0,\cdot)max(0,⋅), and convergence in distribution.

Difficulty

The book's derivation squares (7.3) and takes expectations, using E[(Wq(n+1))2]=E[(Wq(n))2]E[(W_q^{(n+1)})^2] = E[(W_q^{(n)})^2]E[(Wq(n+1)​)2]=E[(Wq(n)​)2]. That step is valid only if the stationary wait has a finite second moment. It is not assumed here and fails in general: with finite second moments of SSS and TTT the stationary wait has a finite mean, but its second moment is finite only if E[S3]<∞E[S^3] < \inftyE[S3]<∞. So the moment identity (7.7) cannot be obtained by cancelling second moments. A truncation or limiting argument is needed, and even the finiteness of WqW_qWq​ has to be proved rather than assumed. The lower bound Wq≥r0W_q \ge r_0Wq​≥r0​ further needs a Jensen argument for the conditional mean of one Lindley step. Theorem 7.1 needs a uniform-integrability argument across a sequence of queues.

Formalization scope

Conventions committed to in Lean:

  • laws, not random variables: AAA, BBB and the stationary law ν\nuν are Measure ℝ; independence of Wq(n),S(n),T(n)W_q^{(n)}, S^{(n)}, T^{(n)}Wq(n)​,S(n),T(n) (and S(n+1)S^{(n+1)}S(n+1) for DDD) is the product measure;
  • the input laws are probability measures on [0,∞)[0,\infty)[0,∞) with finite second moments (MemLp id 2), E[T]=1/λE[T] = 1/\lambdaE[T]=1/λ, E[S]=1/μE[S] = 1/\muE[S]=1/μ, λ,μ>0\lambda, \mu > 0λ,μ>0, ρ=λ/μ<1\rho = \lambda/\mu < 1ρ=λ/μ<1;
  • stationarity is invariance of the whole law ν\nuν under one step of (7.1), not equality of means;
  • WqW_qWq​, the variances (Mathlib variance) and f1f_1f1​ are Lebesgue integrals. Every theorem therefore asserts, as part of its conclusion, that ν\nuν has a finite mean, and none assumes a finite second moment of ν\nuν;
  • U(t)U(t)U(t) is Mathlib's cdf of the law of S−TS - TS−T;
  • convergence in distribution is convergence of ∫g\int g∫g for all bounded continuous ggg, and Exp(1)\mathrm{Exp}(1)Exp(1) is expMeasure 1.

Closed forms carried by the statements: (7.4), (7.7), (7.12), (7.13), (7.14) and (7.17) exactly as printed, and the scaling 2αj/βj22\alpha_j/\beta_j^22αj​/βj2​ of Theorem 7.1.

Stating (7.13) with WqW_qWq​, E[X2]E[X^2]E[X2] or the idle probability as free real numbers constrained by (7.7) would reduce it to algebra. Here WqW_qWq​ is always the mean of a stationary law of the queue.

Not formalized: (7.5) and (7.8), which need the idle-period law III and the arrival-point probability q0q_0q0​ as separate objects, and the multiserver bounds of §7.1.3. Proofs of any milestone are welcome, as are reusable lemmas on stationary laws of Lindley's recursion (existence, uniqueness, and finiteness of the mean under E[S2]<∞E[S^2] < \inftyE[S2]<∞).

Selected references

  • D. Gross, J. F. Shortle, J. M. Thompson, C. M. Harris, Fundamentals of Queueing Theory, 4th ed., Wiley, 2008. https://doi.org/10.1002/9781118625651
  • D. V. Lindley, The theory of queues with a single server, Math. Proc. Cambridge Philos. Soc. 48 (1952)
  • J. F. C. Kingman, The single server queue in heavy traffic, Math. Proc. Cambridge Philos. Soc. 57 (1961)
  • J. F. C. Kingman, Some inequalities for the queue GI/G/1, Biometrika 49 (1962)
  • K. T. Marshall, Some inequalities in queuing, Operations Research 16 (1968)
  • W. G. Marchal, Some simpler bounds on the mean queuing time, Operations Research 26 (1978)
10 thms1 active userReviewed
Operations ResearchTheoretical Computer Science·Captain: mikedeng1

On the Power of Randomization in On-Line Algorithms 3: An Augmented Potential Function Yields an Explicit Deterministic α∘β-Competitive AlgorithmResearch Paper

Motivation

Competitive analysis measures an online algorithm, which must answer each request before seeing the next, against the best off-line answer to the whole request sequence. Randomized online algorithms are often much better than deterministic ones against an oblivious adversary, who fixes the requests in advance; the paging problem is the standard example. Against an adaptive adversary, who sees the algorithm's answers before choosing the next request, the advantage can disappear.

Ben-David, Borodin, Karp, Tardos and Wigderson (Algorithmica 11, 1994; preliminary version STOC 1990) made this precise in an abstract framework of request-answer games. Their Corollary 2.1 says: if a game has a randomized algorithm that is α\alphaα-competitive against adaptive on-line adversaries and one that is β\betaβ-competitive against oblivious adversaries, then it has a deterministic α∘β\alpha\circ\betaα∘β-competitive algorithm. That proof is non-constructive: it goes through a game-theoretic determinacy argument. Section 3 of the paper gives a constructive version. Most competitive analyses of randomized algorithms against adaptive adversaries are carried out with a potential function, in the style of Manasse, McGeoch and Sleator (J. Algorithms 11, 1990). The paper shows that such a potential function, together with any oblivious-competitive algorithm HHH, determines an explicit deterministic algorithm MMM, answer by answer. This mission formalizes that construction and its guarantee.

Setting

A request-answer game has a request set RRR, a finite answer set AAA and cost functions fn:Rn×An→Rf_n : R^n \times A^n \to \mathbb Rfn​:Rn×An→R. The off-line optimum of r∈Rnr \in R^nr∈Rn is c(r)=min⁡a∈Anfn(r,a)c(r) = \min_{a \in A^n} f_n(r, a)c(r)=mina∈An​fn​(r,a). A deterministic online algorithm MMM is a sequence of maps mi:Ri→Am_i : R^i \to Ami​:Ri→A; on r=(r1,…,rn)r = (r_1,\dots,r_n)r=(r1​,…,rn​) it answers M(r)=(m1(r1),m2(r1,r2),…,mn(r))M(r) = (m_1(r_1), m_2(r_1,r_2), \dots, m_n(r))M(r)=(m1​(r1​),m2​(r1​,r2​),…,mn​(r)), at cost cM(r)=fn(r,M(r))c_M(r) = f_n(r, M(r))cM​(r)=fn​(r,M(r)). It is α\alphaα-competitive if cM(r)≤α(c(r))c_M(r) \le \alpha(c(r))cM​(r)≤α(c(r)) for all rrr. Throughout, α\alphaα and β\betaβ are affine maps R→R\mathbb R \to \mathbb RR→R (the paper's "linear functions").

A randomized online algorithm HHH is a probability distribution over deterministic algorithms HyH_yHy​; it is β\betaβ-competitive against any oblivious adversary if Ey[fn(r,Hy(r))]≤β(c(r))\mathbb E_y[f_n(r, H_y(r))] \le \beta(c(r))Ey​[fn​(r,Hy​(r))]≤β(c(r)) for all rrr. The algorithm GGG analysed by the potential function is described by its next-answer laws gn+1(rrn+1,a)g_{n+1}(r r_{n+1}, a)gn+1​(rrn+1​,a) on AAA, given the requests so far, the new request and its own past answers. An adaptive on-line adversary SSS chooses each request from the algorithm's past answers and answers it itself, before the algorithm does, for at most dQd_QdQ​ rounds; a configuration after nnn rounds is (r,a,b)∈Rn×An×An(r, a, b) \in R^n \times A^n \times A^n(r,a,b)∈Rn×An×An: requests, algorithm's answers, adversary's answers.

An augmented potential function for α\alphaα and GGG (Definition 3.1) is a family Φn:Rn×An×An→R\Phi_n : R^n \times A^n \times A^n \to \mathbb RΦn​:Rn×An×An→R with (1) Φ0=0\Phi_0 = 0Φ0​=0; (2) Φn(r,a,b)≤α(fn(r,b))−fn(r,a)\Phi_n(r,a,b) \le \alpha(f_n(r,b)) - f_n(r,a)Φn​(r,a,b)≤α(fn​(r,b))−fn​(r,a) for every configuration; (3) Ean+1∼gn+1(rrn+1,a)[Φn+1(rrn+1,aan+1,bbn+1)]≥Φn(r,a,b)\mathbb E_{a_{n+1} \sim g_{n+1}(r r_{n+1}, a)}[\Phi_{n+1}(r r_{n+1}, a a_{n+1}, b b_{n+1})] \ge \Phi_n(r,a,b)Ean+1​∼gn+1​(rrn+1​,a)​[Φn+1​(rrn+1​,aan+1​,bbn+1​)]≥Φn​(r,a,b) for every configuration, every rn+1∈Rr_{n+1} \in Rrn+1​∈R and every bn+1∈Ab_{n+1} \in Abn+1​∈A.

Formalization targets

Goal: Theorem 3.1 (p. 15)

Let Φ\PhiΦ be an augmented potential function for α\alphaα and GGG, and HHH a β\betaβ-competitive algorithm against oblivious adversaries. Say that MMM obeys the potential rule if for every r∈Rnr \in R^nr∈Rn and r′=rtr' = rtr′=rt,

Ey[Φn+1(r′,M(r) mn+1(r′),Hy(r′))] ≥ Ey[Φn(r,M(r),Hy(r))].\mathbb E_y\big[\Phi_{n+1}(r', M(r)\,m_{n+1}(r'), H_y(r'))\big] \ \ge\ \mathbb E_y\big[\Phi_n(r, M(r), H_y(r))\big].Ey​[Φn+1​(r′,M(r)mn+1​(r′),Hy​(r′))] ≥ Ey​[Φn​(r,M(r),Hy​(r))].

Then such an MMM exists, and every such MMM satisfies

cM(r)≤α(β(c(r)))for all r.c_M(r) \le \alpha\big(\beta(c(r))\big) \quad \text{for all } r .cM​(r)≤α(β(c(r)))for all r.

Both parts are part of the goal: the rule can be followed, and following it guarantees α∘β\alpha\circ\betaα∘β-competitiveness.

Milestones

  1. In every play of GGG against an adaptive on-line adversary, the expected final potential is nonnegative (proof of Lemma 3.1).
  2. Lemma 3.1, "if" direction: an augmented potential function for α\alphaα and GGG makes GGG α\alphaα-competitive against any adaptive on-line adversary, E[cG(S)]≤E[α(cS(G))]\mathbb E[c_G(S)] \le \mathbb E[\alpha(c_S(G))]E[cG​(S)]≤E[α(cS​(G))].
  3. For every rrr, ttt and every a∈Ana \in A^na∈An, some a′∈Aa' \in Aa′∈A satisfies Ey[Φn+1(rt,aa′,Hy(rt))]≥Ey[Φn(r,a,Hy(r))]\mathbb E_y[\Phi_{n+1}(rt, aa', H_y(rt))] \ge \mathbb E_y[\Phi_n(r, a, H_y(r))]Ey​[Φn+1​(rt,aa′,Hy​(rt))]≥Ey​[Φn​(r,a,Hy​(r))].
  4. If MMM obeys the rule, Ey[Φn(r,M(r),Hy(r))]≥0\mathbb E_y[\Phi_n(r, M(r), H_y(r))] \ge 0Ey​[Φn​(r,M(r),Hy​(r))]≥0 for every rrr.
  5. If MMM obeys the rule, fn(r,M(r))≤Ey[α(fn(r,Hy(r)))]f_n(r, M(r)) \le \mathbb E_y[\alpha(f_n(r, H_y(r)))]fn​(r,M(r))≤Ey​[α(fn​(r,Hy​(r)))] for every rrr: MMM is α\alphaα-competitive against the randomized adaptive adversary that serves its requests with HHH.

Significance

The theorem turns two separate analyses into one deterministic algorithm with an explicit description. The potential function certifies GGG against the strongest on-line adversary; the oblivious algorithm HHH need not be related to GGG, and the paper remarks that HHH may be GGG itself. The next answer of MMM is computable whenever the expected potential under HHH is (Corollary 3.1, stated informally in the paper), and the paper notes that for the potential functions used in the KKK-server literature this expectation is computable in time polynomial in the number of nodes and KKK. Read in this light, a potential-function proof for a randomized algorithm doubles as a deterministic algorithm.

The result is proved in the paper. As far as is known, neither this theorem nor the abstract framework of request-answer games with adaptive adversaries has a machine-checked formalization. The mission produces that framework and a checked derandomization principle that applies to every request-answer game with real costs, not to one problem.

Difficulty

The obvious argument for the existence of mn+1(r′)m_{n+1}(r')mn+1​(r′) averages property (3) of Φ\PhiΦ; the work is in seeing which configuration to apply it to. The rule compares MMM's configuration against HyH_yHy​'s answers, not against an adversary playing GGG, and the paper argues through an auxiliary on-line adversary that asks r′r'r′ and serves it with HyH_yHy​. Making this rigorous requires interchanging the expectation over HHH's coins with the finite expectation over GGG's next answer, and checking that the needed expectations are finite.

The second difficulty is the two kinds of randomness. GGG enters only through its next-answer laws, while HHH must be a single distribution over deterministic algorithms: the rule evaluates Hy(r)H_y(r)Hy​(r) and Hy(r′)H_y(r')Hy​(r′) with the same coins yyy. Replacing HHH by a behavioural description breaks the coupling between consecutive rounds.

Formalization scope

Requests and answers are Lean lists, oldest first, and fn(r,a)f_n(r,a)fn​(r,a) is F.cost r a on lists of common length; values on lists of different lengths are never used. Costs are real: the paper allows fn=+∞f_n = +\inftyfn​=+∞, so every statement here is about the real-valued games. The answer type is finite and nonempty, so the minimum c(r)c(r)c(r) exists. Affine maps are written α(x)=cx+d\alpha(x) = c x + dα(x)=cx+d. The goal additionally assumes α\alphaα nondecreasing: the last step of the paper's proof applies α\alphaα to an inequality, which needs it, and the paper's examples are positive ratios. In Lemma 3.1 and milestone 5 linearity of α\alphaα is kept as the paper's standing convention, although with α\alphaα inside the expectation the argument does not use it.

GGG is a map from (requests, own answers) to a probability mass function on AAA (behavioural form); its play against an adaptive on-line adversary is a probability mass function on final configurations, with finite support, and its expectations are finite sums. HHH is a probability measure on a coin space with a deterministic algorithm per coin, each answer measurable in the coins; expectations over HHH are Bochner integrals of functions with finitely many values. α\alphaα stays inside expectations, as in the paper's definition of competitiveness against adaptive adversaries. Adversaries stop by returning none and have a uniform depth bound.

The goal cannot be satisfied vacuously: it states the existence of an algorithm obeying the rule alongside the guarantee for every such algorithm, and Definition 3.1 is required at every configuration, not only at reachable ones.

Not formalized: the "only if" direction of Lemma 3.1, which the paper only sketches, and Corollary 3.1, whose notion of computability the paper leaves unspecified. The definitions of request-answer games, online algorithms, adversaries and competitiveness are reusable for the other missions of this paper and for any problem-specific competitive analysis. Contributions welcome: proofs of the milestones, and a lemma relating the mixed and behavioural forms of a randomized algorithm.

Selected references

  • S. Ben-David, A. Borodin, R. Karp, G. Tardos, A. Wigderson, On the power of randomization in on-line algorithms, Algorithmica 11 (1994), 2–14. https://doi.org/10.1007/BF01294260
  • M. Manasse, L. McGeoch, D. Sleator, Competitive algorithms for server problems, Journal of Algorithms 11 (1990), 208–230. https://doi.org/10.1016/0196-6774(90)90003-W
  • D. Sleator, R. Tarjan, Amortized efficiency of list update and paging rules, Communications of the ACM 28 (1985), 202–208. https://doi.org/10.1145/2786.2793
  • A. Borodin, R. El-Yaniv, Online Computation and Competitive Analysis, Cambridge University Press, 1998. ISBN 0-521-56392-5
9 thms1 active userReviewed
Machine LearningStatistics·Captain: mikedeng1

Stability and Generalization 1: Polynomial Generalization Bounds from Hypothesis Stability for the Empirical and Leave-One-Out ErrorsResearch Paper

Why stability bounds

A learning algorithm is judged by its generalization error, its expected loss on a fresh example, which cannot be computed because the data distribution is unknown. Practitioners estimate it either by the empirical error on the training set or by the leave-one-out error, which retrains the algorithm once per example. Classical learning theory justifies these estimates through uniform convergence over the whole hypothesis space (VC dimension, covering numbers). That route says nothing useful about algorithms such as nearest-neighbour rules or regularized kernel methods, whose effective hypothesis space is huge or unknown.

An alternative is to bound the deviation through a property of the algorithm itself: how much its output changes when one training example is removed. This idea goes back to Rogers and Wagner (1978) and Devroye and Wagner (1979) for local rules, and Kearns and Ron (1999) gave it a name. Bousquet and Elisseeff (JMLR 2002) systematized it with several stability notions and corresponding bounds; their paper is the standard reference for algorithmic stability in learning theory. This mission formalizes its first family of results, the polynomial bounds of §4.1.

Setting

Let Z=X×YZ = X \times YZ=X×Y and let DDD be a probability distribution on ZZZ. A training set S={z1,…,zm}S = \{z_1, \dots, z_m\}S={z1​,…,zm​} consists of mmm examples drawn i.i.d. from DDD. A learning algorithm AAA maps a training set SSS to a hypothesis AS:X→Y′A_S : X \to Y'AS​:X→Y′; it is deterministic and symmetric, meaning it does not depend on the order of the examples. A cost ccc with 0≤c(y′,y)≤M0 \le c(y', y) \le M0≤c(y′,y)≤M defines the loss ℓ(f,z)=c(f(x),y)\ell(f, z) = c(f(x), y)ℓ(f,z)=c(f(x),y) of a hypothesis fff at z=(x,y)z = (x, y)z=(x,y).

For each index iii, S∖iS^{\setminus i}S∖i is SSS with ziz_izi​ removed, and SiS^iSi is SSS with ziz_izi​ replaced by an independent fresh draw zi′∼Dz'_i \sim Dzi′​∼D. The three error quantities are

R(A,S)=Ez[ℓ(AS,z)],Remp(A,S)=1m∑i=1mℓ(AS,zi),Rloo(A,S)=1m∑i=1mℓ(AS∖i,zi).R(A,S) = \mathbb E_z[\ell(A_S, z)], \qquad R_{\mathrm{emp}}(A,S) = \frac1m \sum_{i=1}^m \ell(A_S, z_i), \qquad R_{\mathrm{loo}}(A,S) = \frac1m \sum_{i=1}^m \ell(A_{S^{\setminus i}}, z_i).R(A,S)=Ez​[ℓ(AS​,z)],Remp​(A,S)=m1​i=1∑m​ℓ(AS​,zi​),Rloo​(A,S)=m1​i=1∑m​ℓ(AS∖i​,zi​).

Two stability notions (Definitions 3 and 4) control them. AAA has hypothesis stability β1\beta_1β1​ if ES,z[∣ℓ(AS,z)−ℓ(AS∖i,z)∣]≤β1\mathbb E_{S,z}[|\ell(A_S,z) - \ell(A_{S^{\setminus i}},z)|] \le \beta_1ES,z​[∣ℓ(AS​,z)−ℓ(AS∖i​,z)∣]≤β1​ for every iii, and pointwise hypothesis stability β2\beta_2β2​ if ES[∣ℓ(AS,zi)−ℓ(AS∖i,zi)∣]≤β2\mathbb E_{S}[|\ell(A_S,z_i) - \ell(A_{S^{\setminus i}},z_i)|] \le \beta_2ES​[∣ℓ(AS​,zi​)−ℓ(AS∖i​,zi​)∣]≤β2​ for every iii.

Formalization targets

Goal: Theorem 11

For m≥1m \ge 1m≥1, under hypothesis stability β1\beta_1β1​ and pointwise hypothesis stability β2\beta_2β2​, for every δ>0\delta > 0δ>0, each of the following holds with probability at least 1−δ1 - \delta1−δ over S∼DmS \sim D^mS∼Dm:

R(A,S)≤Remp(A,S)+M2+6Mm(β1+β2)2mδ,R(A,S)≤Rloo(A,S)+M2+6Mmβ12mδ.R(A,S) \le R_{\mathrm{emp}}(A,S) + \sqrt{\frac{M^2 + 6Mm(\beta_1+\beta_2)}{2m\delta}}, \qquad R(A,S) \le R_{\mathrm{loo}}(A,S) + \sqrt{\frac{M^2 + 6Mm\beta_1}{2m\delta}} .R(A,S)≤Remp​(A,S)+2mδM2+6Mm(β1​+β2​)​​,R(A,S)≤Rloo​(A,S)+2mδM2+6Mmβ1​​​.

Milestones

  1. Lemma 25 (p. 520), a generalized Rogers–Wagner identity: upper bounds on ES[(R−Remp)2]\mathbb E_S[(R - R_{\mathrm{emp}})^2]ES​[(R−Remp​)2] and ES[(R−Rloo)2]\mathbb E_S[(R - R_{\mathrm{loo}})^2]ES​[(R−Rloo​)2] by correlations of the loss.
  2. Lemma 9, (8) and (9) (p. 505): ES[(R−Remp)2]≤M22m+3M ES,zi′[∣ℓ(AS,zi)−ℓ(ASi,zi)∣]\mathbb E_S[(R - R_{\mathrm{emp}})^2] \le \frac{M^2}{2m} + 3M\,\mathbb E_{S,z'_i}[|\ell(A_S,z_i) - \ell(A_{S^i},z_i)|]ES​[(R−Remp​)2]≤2mM2​+3MES,zi′​​[∣ℓ(AS​,zi​)−ℓ(ASi​,zi​)∣] and ES[(R−Rloo)2]≤M22m+3M ES,z[∣ℓ(AS,z)−ℓ(AS∖i,z)∣]\mathbb E_S[(R - R_{\mathrm{loo}})^2] \le \frac{M^2}{2m} + 3M\,\mathbb E_{S,z}[|\ell(A_S,z) - \ell(A_{S^{\setminus i}},z)|]ES​[(R−Rloo​)2]≤2mM2​+3MES,z​[∣ℓ(AS​,z)−ℓ(AS∖i​,z)∣].
  3. The replace-one term (proof of Theorem 11): ES,zi′[∣ℓ(AS,zi)−ℓ(ASi,zi)∣]≤β1+β2\mathbb E_{S,z'_i}[|\ell(A_S,z_i) - \ell(A_{S^i},z_i)|] \le \beta_1 + \beta_2ES,zi′​​[∣ℓ(AS​,zi​)−ℓ(ASi​,zi​)∣]≤β1​+β2​.
  4. The second-moment bounds (proof of Theorem 11): ES[(R−Remp)2]≤M22m+3M(β1+β2)\mathbb E_S[(R - R_{\mathrm{emp}})^2] \le \frac{M^2}{2m} + 3M(\beta_1+\beta_2)ES​[(R−Remp​)2]≤2mM2​+3M(β1​+β2​) and ES[(R−Rloo)2]≤M22m+3Mβ1\mathbb E_S[(R - R_{\mathrm{loo}})^2] \le \frac{M^2}{2m} + 3M\beta_1ES​[(R−Rloo​)2]≤2mM2​+3Mβ1​.

Significance

Theorem 11 is the weakest-assumption bound in the paper: it requires only average-case stability, not the uniform (worst-case) stability behind the exponential bounds of §4.2. It shows that both the resubstitution and the deleted estimate are within O(1/mδ)O(1/\sqrt{m\delta})O(1/mδ​) of the risk whenever the stability parameters decay like 1/m1/m1/m, with no reference to the size of the hypothesis class. It also extends Devroye and Wagner's leave-one-out analysis for classification to bounded regression losses and to the empirical estimator. Later work on average stability and on generalization of stochastic gradient methods (for example Hardt, Recht and Singer, 2016) starts from these notions.

The result is proved in the paper; as far as is known it has no machine-checked proof. A formal development has two concrete payoffs. First, it fixes the constants: in checking the argument, two printed slips were found (the empirical constant in Theorem 11 and the third term of Lemma 25's first inequality), and the formal statements record the versions that the paper's proof actually establishes. Second, the Lemma 25 and Lemma 9 machinery — exchangeability of i.i.d. samples under renaming, and second-moment control through stability — is reusable for any later stability result.

Difficulty

The obvious route is the Efron–Stein (Steele) variance inequality, Theorem 1 of the paper. It bounds the variance of R−RempR - R_{\mathrm{emp}}R−Remp​, not its second moment, and leaves the bias to be handled separately; the paper notes that it gives worse constants. The direct route of Appendix A instead expands ES[(R−Remp)2]\mathbb E_S[(R - R_{\mathrm{emp}})^2]ES​[(R−Remp​)2] and rewrites each correlation term by renaming i.i.d. variables: training points, fresh test points and replacement points are exchanged with one another, and the algorithm is retrained on sets T∪{z,z′}T \cup \{z, z'\}T∪{z,z′} with T=S∖{i,j}T = S^{\setminus \{i,j\}}T=S∖{i,j}. Every renaming is a measure-preserving map on a product of m+2m + 2m+2 copies of DDD, and each must be justified by the symmetry of AAA. Doing this rigorously, rather than as "a matter of renaming", is the core of the work. The leave-one-out case is only sketched in the paper ("it is easy to see"), so its formal proof has to be reconstructed.

Formalization scope

  • An algorithm is a function Multiset (X × Y) → (X → Y'). Symmetry in the training set is built into the type, and the same algorithm acts on sets of every size, as SSS and S∖iS^{\setminus i}S∖i require. A sample is S : Fin m → X × Y with law DmD^mDm (Measure.pi); fresh points zzz, z′z'z′, zi′z'_izi′​ are further independent coordinates, via product measures Dm⊗DD^m \otimes DDm⊗D and (Dm⊗D)⊗D(D^m \otimes D) \otimes D(Dm⊗D)⊗D.
  • The loss, empirical error and generalization error are the published FoundationsML.Stability definitions (Loss, EmpiricalError, GeneralizationError).
  • The cost satisfies 0≤c≤M0 \le c \le M0≤c≤M everywhere. The paper's assumption that "all functions are measurable" becomes one hypothesis: for every nnn, (S,z)↦ℓ(AS,z)(S, z) \mapsto \ell(A_S, z)(S,z)↦ℓ(AS​,z) is measurable on (X×Y)n×(X×Y)(X \times Y)^n \times (X \times Y)(X×Y)n×(X×Y). Both stability definitions also require their integrands to be integrable. Together these rule out the trivializing reading in which a non-integrable expectation equals Lean's default value 000 and the stability hypotheses hold vacuously.
  • "With probability 1−δ1 - \delta1−δ" is stated as a bound on the failure event: Dm{S:R>Remp+⋯ }≤δD^m\{S : R > R_{\mathrm{emp}} + \cdots\} \le \deltaDm{S:R>Remp​+⋯}≤δ for every δ>0\delta > 0δ>0, separately for each estimator.
  • m≥2m \ge 2m≥2 is assumed in Lemmas 9 and 25 and in the two second-moment steps of the proof, because the lemmas refer to two distinct indices. Theorem 11 itself is stated for every m≥1m \ge 1m≥1, as printed.
  • Corrected statements. (i) Theorem 11's empirical bound is stated with 6Mm(β1+β2)6Mm(\beta_1+\beta_2)6Mm(β1​+β2​), not the printed 12Mmβ212Mm\beta_212Mmβ2​: the proof bounds a hypothesis-stability term by β2\beta_2β2​ when it is bounded by β1\beta_1β1​. The two coincide when β1=β2\beta_1 = \beta_2β1​=β2​. Accordingly the replace-one milestone is stated as ≤β1+β2\le \beta_1 + \beta_2≤β1​+β2​ (printed 2β22\beta_22β2​), and the empirical second-moment bound as M22m+3M(β1+β2)\frac{M^2}{2m} + 3M(\beta_1+\beta_2)2mM2​+3M(β1​+β2​) (printed 6Mβ26M\beta_26Mβ2​). (ii) Lemma 25's empirical inequality has ES[ℓ(AS,zi)ℓ(AS,zj)]\mathbb E_S[\ell(A_S,z_i)\ell(A_S,z_j)]ES​[ℓ(AS​,zi​)ℓ(AS​,zj​)] as its third term, as its proof gives, not the printed leave-one-out term. (iii) The leave-one-out second-moment bound follows from (9), not from (10) as printed.

Contributions are welcome at every level: proofs of the milestones, a general exchangeability lemma for symmetric algorithms on product measures, and Markov/Chebyshev glue for the final step.

Selected references

  • O. Bousquet and A. Elisseeff, Stability and Generalization, Journal of Machine Learning Research 2 (2002), 499–526. https://www.jmlr.org/papers/v2/bousquet02a.html
  • W. H. Rogers and T. J. Wagner, A finite sample distribution-free performance bound for local discrimination rules, Annals of Statistics 6(3) (1978), 506–514. https://doi.org/10.1214/aos/1176344196
  • L. Devroye and T. J. Wagner, Distribution-free performance bounds for potential function rules, IEEE Transactions on Information Theory 25(5) (1979), 601–604. https://doi.org/10.1109/TIT.1979.1056087
  • M. Kearns and D. Ron, Algorithmic stability and sanity-check bounds for leave-one-out cross-validation, Neural Computation 11(6) (1999), 1427–1453. https://doi.org/10.1162/089976699300016304
  • M. Hardt, B. Recht and Y. Singer, Train faster, generalize better: stability of stochastic gradient descent, ICML 2016. https://arxiv.org/abs/1509.01240
13 thms1 active userReviewed
Control TheoryOperations ResearchStochastic Systems·Captain: mikedeng1

Dynamic Scheduling of a System with Two Parallel Servers in Heavy Traffic with Resource Pooling: The Threshold Policy Is Asymptotically OptimalResearch Paper

Motivation

Many service systems route several classes of work to servers with overlapping skills: call centers with cross-trained agents, manufacturing cells with flexible machines, computing clusters with heterogeneous processors. Choosing which server works on which class at each moment is a dynamic scheduling problem. Exact optimal policies are out of reach except in toy cases, so heavy-traffic theory replaces the queueing system by a Brownian control problem, solves that limit problem, and then asks for a policy in the original system whose performance converges to the Brownian optimum. This programme was proposed by Harrison (Harrison 1988), and the parallel server system studied here is the example Harrison used (Harrison, Ann. Appl. Probab. 1998) to show that the greedy static priority rule can be very inefficient.

Bell and Williams (2001) gave the first proof of asymptotic optimality of a continuous-review policy for this system, with renewal arrivals and general service times. Harrison (1998) had treated Poisson arrivals and deterministic service times with a discrete-review policy and a pathwise criterion. Harrison and López (Queueing Systems, 1999) identified the complete resource pooling condition for general parallel server systems. The threshold policy and the proof method of Bell and Williams were later extended to multiserver systems (Bell and Williams, Electron. J. Probab., 2005).

Setting

There are two job classes and two servers. Server 1 serves class 1 (activity 1); server 2 serves class 1 (activity 2) and class 2 (activity 3). A sequence of such systems is indexed by r→∞r\to\inftyr→∞. On a probability space, i.i.d. sequences uˇk(i)\check u_k(i)uˇk​(i) (k=1,2k=1,2k=1,2) and vˇj(i)\check v_j(i)vˇj​(i) (j=1,2,3j=1,2,3j=1,2,3), i≥1i\ge1i≥1, are fixed: strictly positive, mutually independent, with mean one and finite variances αk2,βj2\alpha_k^2,\beta_j^2αk2​,βj2​. In system rrr the interarrival times are ukr(i)=uˇk(i)/λkru_k^r(i)=\check u_k(i)/\lambda_k^rukr​(i)=uˇk​(i)/λkr​ and the service times are vjr(i)=vˇj(i)/μjrv_j^r(i)=\check v_j(i)/\mu_j^rvjr​(i)=vˇj​(i)/μjr​. The renewal processes Akr(t)A_k^r(t)Akr​(t) and Sjr(t)S_j^r(t)Sjr​(t) count arrivals and potential service completions.

A scheduling control policy is an allocation T=(T1,T2,T3)T=(T_1,T_2,T_3)T=(T1​,T2​,T3​), where Tj(t)T_j(t)Tj​(t) is the time devoted to activity jjj in [0,t][0,t][0,t]. Each Tj(t)T_j(t)Tj​(t) is a random variable, each TjT_jTj​ is continuous and nondecreasing from 000, and so are the idle times I1=t−T1I_1=t-T_1I1​=t−T1​ and I2=t−T2−T3I_2=t-T_2-T_3I2​=t−T2​−T3​. The queue lengths

Q1(t)=A1(t)−S1(T1(t))−S2(T2(t)),Q2(t)=A2(t)−S3(T3(t))Q_1(t)=A_1(t)-S_1(T_1(t))-S_2(T_2(t)),\qquad Q_2(t)=A_2(t)-S_3(T_3(t))Q1​(t)=A1​(t)−S1​(T1​(t))−S2​(T2​(t)),Q2​(t)=A2​(t)−S3​(T3​(t))

must be nonnegative. Policies may anticipate the future. The rates satisfy Assumption 3.1: λ1>μ1\lambda_1>\mu_1λ1​>μ1​, 1−(λ1−μ1)/μ2=λ2/μ31-(\lambda_1-\mu_1)/\mu_2=\lambda_2/\mu_31−(λ1​−μ1​)/μ2​=λ2​/μ3​, and the rates converge at rate 1/r1/r1/r to limits with second-order parameters θ1,θ2\theta_1,\theta_2θ1​,θ2​. Assumption 3.2 is h1μ2≥h2μ3h_1\mu_2\ge h_2\mu_3h1​μ2​≥h2​μ3​, and Assumption 3.3 gives finite exponential moments near 000. With Q^r(t)=r−1Qr(r2t)\hat Q^r(t)=r^{-1}Q^r(r^2t)Q^​r(t)=r−1Qr(r2t) the cost is

J^r(Tr)=E(∫0∞e−γt h⋅Q^r(t) dt).\hat J^r(T^r)=\mathbf E\Big(\int_0^\infty e^{-\gamma t}\,h\cdot\hat Q^r(t)\,dt\Big).J^r(Tr)=E(∫0∞​e−γth⋅Q^​r(t)dt).

The threshold policy with Lr=[clog⁡r]L^r=[c\log r]Lr=[clogr] works as follows. Server 1 works whenever it has a class 1 job available. Server 2 serves class 1 with preemptive-resume priority when more than LrL^rLr class 1 jobs are present, and otherwise serves class 2. The Brownian benchmark is built from a two-dimensional Brownian motion X~\tilde XX~ with drift θ\thetaθ and diagonal covariance, from y=(1,μ2/μ3)y=(1,\mu_2/\mu_3)y=(1,μ2​/μ3​), and from the reflected process W~∗=y⋅X~+V~∗\tilde W^*=y\cdot\tilde X+\tilde V^*W~∗=y⋅X~+V~∗ with V~∗(t)=−inf⁡s≤ty⋅X~(s)\tilde V^*(t)=-\inf_{s\le t}y\cdot\tilde X(s)V~∗(t)=−infs≤t​y⋅X~(s). Its cost is J∗=E∫0∞e−γth2 W~∗(t)/y2 dtJ^*=\mathbf E\int_0^\infty e^{-\gamma t}h_2\,\tilde W^*(t)/y_2\,dtJ∗=E∫0∞​e−γth2​W~∗(t)/y2​dt.

Formalization targets

Goal: Theorem 5.3

For ccc larger than a constant c0c_0c0​ that depends only on the model data, and for every sequence {Tr}\{T^r\}{Tr} of scheduling control policies,

lim inf⁡r→∞J^r(Tr) ≥ J∗ = lim⁡r→∞J^r(Tr,∗),J∗<∞.\liminf_{r\to\infty}\hat J^r(T^r)\ \ge\ J^*\ =\ \lim_{r\to\infty}\hat J^r(T^{r,*}),\qquad J^*<\infty .r→∞liminf​J^r(Tr) ≥ J∗ = r→∞lim​J^r(Tr,∗),J∗<∞.

Milestones

  • Proposition B.1: the one-dimensional Skorokhod problem, its explicit solution and its minimality.
  • Appendix A, (181) and (184): Cramér-type deviation bounds for delayed renewal processes.
  • Theorem 7.2: after first reaching LrL^rLr, the class 1 queue stays within Lr−1L^r-1Lr−1 of the threshold, with probability tending to one.
  • Theorem 7.1: (Q^1r,I^1r)⇒(0,0)(\hat Q_1^r,\hat I_1^r)\Rightarrow(0,0)(Q^​1r​,I^1r​)⇒(0,0) under the threshold policy.
  • Lemma 8.1: the fluid-scaled threshold allocations converge to Tˉ∗(t)=(t,λ1−μ1μ2t,λ2μ3t)\bar T^*(t)=(t,\frac{\lambda_1-\mu_1}{\mu_2}t,\frac{\lambda_2}{\mu_3}t)Tˉ∗(t)=(t,μ2​λ1​−μ1​​t,μ3​λ2​​t).
  • Theorem 5.2 (state-space collapse): (Q^1r,Q^2r,I^1r,I^2r)⇒(0,Q~2∗,0,I~2∗)(\hat Q_1^r,\hat Q_2^r,\hat I_1^r,\hat I_2^r)\Rightarrow(0,\tilde Q_2^*,0,\tilde I_2^*)(Q^​1r​,Q^​2r​,I^1r​,I^2r​)⇒(0,Q~​2∗​,0,I~2∗​).
  • Lemma 9.3: along a subsequence achieving a finite lim inf⁡\liminfliminf cost, the fluid-scaled processes converge to (0,λt,μt,Tˉ∗,0)(0,\lambda t,\mu t,\bar T^*,0)(0,λt,μt,Tˉ∗,0).

A further draft theorem states that Definition 5.1 determines an admissible allocation, unique pathwise, whenever Lr≥1L^r\ge1Lr≥1.

Significance

The theorem proves that a simple state-dependent rule, which sends server 2 to class 1 only when the class 1 queue exceeds a logarithmic safety stock, is asymptotically optimal among all policies, including those that anticipate the future. The limiting cost is the explicit optimum of the Brownian control problem. The proof gives a template for heavy-traffic asymptotic optimality under complete resource pooling: a lower bound valid for every policy, and state-space collapse under the proposed policy. The residual process analysis of Section 7 shows how a threshold of order log⁡r\log rlogr makes starvation of server 1 negligible on the diffusion time scale.

The paper's results are proved but not machine-checked; no formal proof exists in any proof assistant. The mission asks for formal statements of the paper's main theorem and its supporting lemmas, followed by formal proofs. Parts of the development are independent of the paper: the one-dimensional Skorokhod map, renewal large deviation bounds, and convergence encodings on path space.

Difficulty

The lower bound must hold for arbitrary, possibly anticipating, policies, so no Markov structure is available. The argument has to pass through fluid limits of an arbitrary cost-minimizing subsequence and a pathwise minimality property, and Fatou's lemma for the limit needs uniform control. For the upper bound, the obvious approach, a static priority rule, is known to fail: it starves server 1 and produces a large class 1 queue. With a threshold policy, the hard step is to show that the class 1 queue, once at the threshold, rarely moves Lr−1L^r-1Lr−1 away from it over a time interval of length r2tr^2tr2t. That requires large deviation estimates for renewal processes started at random, multiparameter stopping times. Showing that J^r(Tr,∗)\hat J^r(T^{r,*})J^r(Tr,∗) converges to J∗J^*J∗, rather than only that the processes converge in distribution, also requires uniform integrability of the scaled queue lengths.

Formalization scope

Classes and activities are indexed by Fin 2 and Fin 3. The i.i.d. sequences keep the paper's index base i≥1i\ge1i≥1, and the systems are indexed by n∈Nn\in\mathbb Nn∈N with r=rn∈[1,∞)r=r_n\in[1,\infty)r=rn​∈[1,∞), rn→∞r_n\to\inftyrn​→∞. Time is real, and every condition is imposed for t≥0t\ge0t≥0. Admissibility is exactly (11)–(14). Measurability in (11) is with respect to the completion of P\mathbf PP, since the paper's space is complete. Finiteness of the renewal processes everywhere on Ω\OmegaΩ, which the paper obtains by discarding a null set, is a hypothesis. Queue lengths are real, costs are lower Lebesgue integrals in [0,∞][0,\infty][0,∞], counting processes take values in N∪{∞}\mathbb N\cup\{\infty\}N∪{∞}, and Λ\LambdaΛ, Λ∗\Lambda^*Λ∗ take values in the extended reals.

The constant c0c_0c0​ is existential and is chosen after the model data and before ccc, the policies and the Brownian motions. The threshold relations are required only for the systems with Lr≥1L^r\ge1Lr≥1, which are all but finitely many. Each convergence to a deterministic limit (Theorem 7.1, Lemmas 8.1 and 9.3) is stated as u.o.c. convergence in probability, the paper's own equivalence (p. 633). Theorem 5.2 is stated in coupling form: there are copies of the processes on one probability space, with Skorokhod paths and the same laws, that converge almost surely uniformly on compacts. This is equivalent to weak convergence in D4\mathbf D^4D4 to a limit with continuous paths. J∗J^*J∗ is defined by (44) from an arbitrary pair of independent standard Brownian motions (Mathlib's IsBrownianReal), not by a closed form.

Two formalizations would make the goal trivial, and both are excluded. Leaving out the requirement that Tr,∗T^{r,*}Tr,∗ actually follow the policy would make the goal false or empty. Narrowing the class of competing policies, for example to non-anticipating ones, would weaken the theorem. A draft theorem also states that the threshold allocation exists and is unique pathwise, so the hypothesis on Tr,∗T^{r,*}Tr,∗ can be satisfied.

The development needs renewal theory (functional central limit theorems, Cramér bounds), multiparameter stopping times, tightness in D\mathbf DD, the Skorokhod representation theorem, the reflection map, and properties of reflected Brownian motion. Contributions are welcome at every level: proofs of milestones, reusable lemmas on renewal processes and the Skorokhod map, and further lemmas of the paper (Lemmas 7.5, 7.6 and 9.2 are not yet stated).

Selected references

  • S. L. Bell and R. J. Williams, Dynamic scheduling of a system with two parallel servers in heavy traffic with resource pooling: asymptotic optimality of a threshold policy, Ann. Appl. Probab. 11 (2001) 608–649. https://doi.org/10.1214/aoap/1015345343
  • J. M. Harrison, Heavy traffic analysis of a system with parallel servers: asymptotic optimality of discrete-review policies, Ann. Appl. Probab. 8 (1998) 822–848.
  • J. M. Harrison and M. J. López, Heavy traffic resource pooling in parallel-server systems, Queueing Systems 33 (1999) 339–368.
  • J. M. Harrison, Brownian models of queueing networks with heterogeneous customer populations, in Stochastic Differential Systems, Stochastic Control Theory and Their Applications, Springer (1988) 147–186.
  • S. L. Bell and R. J. Williams, Dynamic scheduling of a parallel server system in heavy traffic with complete resource pooling: asymptotic optimality of a threshold policy, Electron. J. Probab. 10 (2005) 1044–1115.
  • J. M. Harrison, Brownian Motion and Stochastic Flow Systems, Wiley (1985).
14 thms1 active userReviewed
Dynamical SystemsOperations ResearchStochastic Systems·Captain: mikedeng1

Dynamics of Stochastic Approximation Algorithms 7: Weak Limit Points of the Occupation Measures of a Weak Asymptotic Pseudotrajectory Are InvariantResearch Paper

Motivation

Stochastic approximation algorithms are recursions xn+1−xn=γn+1(F(xn)+Un+1)x_{n+1}-x_n=\gamma_{n+1}(F(x_n)+U_{n+1})xn+1​−xn​=γn+1​(F(xn​)+Un+1​) driven by small steps γn\gamma_nγn​ and noise Un+1U_{n+1}Un+1​; they include the Robbins–Monro scheme, stochastic gradient methods and learning dynamics in games. The ODE method studies their long-run behaviour by comparing a time-interpolation of the iterates with the trajectories of a deterministic dynamical system. In Benaïm's lecture notes (Benaïm 1999) this comparison is formalized by the notion of an asymptotic pseudotrajectory, introduced in Benaïm and Hirsch (1996): a path that, over every window of fixed length, shadows the deterministic orbit started at its current position with an error that vanishes as time goes to infinity.

The pathwise results of the earlier sections of the notes concern algorithms whose step sizes decrease fast enough, typically γn=o(1/log⁡n)\gamma_n=o(1/\log n)γn​=o(1/logn) or γn=O(n−α)\gamma_n=O(n^{-\alpha})γn​=O(n−α). When the step sizes go to zero more slowly, the limit sets of the process can no longer be characterized precisely: with steps of order 1/log⁡n1/\log n1/logn the process may fail to converge even when the chain recurrent set of the ODE consists of isolated equilibria. Section 10, which is mainly based on work of Benaïm and Schreiber, describes instead the statistical behaviour of such processes in terms of the deterministic dynamics. It introduces a weaker, conditional notion, the weak asymptotic pseudotrajectory, and proves in Theorem 10.1 that the empirical distribution of the time spent by the process in different regions of the state space accumulates only on invariant measures of the deterministic dynamics. This is an ergodic-theoretic counterpart of the limit-set theorems of Section 5.

Setting

A semiflow on a metric space (M,d)(M,d)(M,d) is a continuous map Φ:R+×M→M\Phi:\mathbb R_+\times M\to MΦ:R+​×M→M, (t,x)↦Φt(x)(t,x)\mapsto\Phi_t(x)(t,x)↦Φt​(x), with Φ0=Id\Phi_0=\mathrm{Id}Φ0​=Id and Φt+s=Φt∘Φs\Phi_{t+s}=\Phi_t\circ\Phi_sΦt+s​=Φt​∘Φs​ for t,s≥0t,s\ge0t,s≥0. Throughout, MMM is a separable metric space with its Borel σ\sigmaσ-algebra.

Let (Ω,F,P)(\Omega,\mathcal F,P)(Ω,F,P) be a probability space and {Ft}t≥0\{\mathcal F_t\}_{t\ge0}{Ft​}t≥0​ a nondecreasing family of sub-σ\sigmaσ-algebras. A process X:R+×Ω→MX:\mathbb R_+\times\Omega\to MX:R+​×Ω→M is a weak asymptotic pseudotrajectory of Φ\PhiΦ if

  1. it is progressively measurable: for every T>0T>0T>0 the restriction of XXX to [0,T]×Ω[0,T]\times\Omega[0,T]×Ω is measurable for the product of the Borel σ\sigmaσ-field of [0,T][0,T][0,T] and FT\mathcal F_TFT​;
  2. for each α>0\alpha>0α>0 and T>0T>0T>0, almost surely
lim⁡t→∞P{sup⁡0≤h≤Td(X(t+h),Φh(X(t)))≥α ∣ Ft}=0.\lim_{t\to\infty}P\Big\{\sup_{0\le h\le T}d\big(X(t+h),\Phi_h(X(t))\big)\ge\alpha\ \Big|\ \mathcal F_t\Big\}=0 .t→∞lim​P{0≤h≤Tsup​d(X(t+h),Φh​(X(t)))≥α ​ Ft​}=0.

Let P(M)\mathcal P(M)P(M) be the space of Borel probability measures on MMM with the topology of weak convergence. A measure μ∈P(M)\mu\in\mathcal P(M)μ∈P(M) is Φ\PhiΦ-invariant if (Φt)∗μ=μ(\Phi_t)_*\mu=\mu(Φt​)∗​μ=μ for every t≥0t\ge0t≥0; the set of invariant measures is M(Φ)\mathcal M(\Phi)M(Φ). The occupation measure of the process at time t>0t>0t>0 is the random probability measure

μt(ω)=1t∫0tδX(s,ω) ds,\mu_t(\omega)=\frac1t\int_0^t\delta_{X(s,\omega)}\,ds ,μt​(ω)=t1​∫0t​δX(s,ω)​ds,

and M(X,ω)⊂P(M)\mathcal M(X,\omega)\subset\mathcal P(M)M(X,ω)⊂P(M) is the set of its weak limit points as t→∞t\to\inftyt→∞.

Formalization targets

Goal: Theorem 10.1

If XXX is a weak asymptotic pseudotrajectory of Φ\PhiΦ, there is a set Ω~⊂Ω\tilde\Omega\subset\OmegaΩ~⊂Ω with P(Ω~)=1P(\tilde\Omega)=1P(Ω~)=1 such that for all ω∈Ω~\omega\in\tilde\Omegaω∈Ω~

M(X,ω)⊂M(Φ).\mathcal M(X,\omega)\subset\mathcal M(\Phi).M(X,ω)⊂M(Φ).

No tightness is assumed, so M(X,ω)\mathcal M(X,\omega)M(X,ω) may be empty; the statement asserts the inclusion, not nonemptiness.

Milestones

Fix a uniformly continuous f:M→[0,1]f:M\to[0,1]f:M→[0,1] and T>0T>0T>0, and set Un(f,T)=∫(n−1)TnTf(X(s)) dsU_n(f,T)=\int_{(n-1)T}^{nT}f(X(s))\,dsUn​(f,T)=∫(n−1)TnT​f(X(s))ds for n≥1n\ge1n≥1. The milestones are the numbered displays of the proof on pp. 62–63:

  • Eq. (47): 1n∑i=1n[Ui(f,T)−E(Ui(f,T)∣F(i−1)T)]→0\frac1n\sum_{i=1}^n[U_i(f,T)-E(U_i(f,T)\mid\mathcal F_{(i-1)T})]\to0n1​∑i=1n​[Ui​(f,T)−E(Ui​(f,T)∣F(i−1)T​)]→0 almost surely (stated for every continuous fff with values in [0,1][0,1][0,1], since the proof also applies it to f∘ΦTf\circ\Phi_Tf∘ΦT​);
  • Eq. (50): the same with Ui+1(f,T)U_{i+1}(f,T)Ui+1​(f,T) conditioned on F(i−1)T\mathcal F_{(i-1)T}F(i−1)T​;
  • Eq. (51): E(Ui+1(f,T)−Ui(f∘ΦT,T)∣F(i−1)T)→0E(U_{i+1}(f,T)-U_i(f\circ\Phi_T,T)\mid\mathcal F_{(i-1)T})\to0E(Ui+1​(f,T)−Ui​(f∘ΦT​,T)∣F(i−1)T​)→0 almost surely;
  • Eq. (52): 1n∑i=1nUi+1(f,T)−1n∑i=1nUi(f∘ΦT,T)→0\frac1n\sum_{i=1}^nU_{i+1}(f,T)-\frac1n\sum_{i=1}^nU_i(f\circ\Phi_T,T)\to0n1​∑i=1n​Ui+1​(f,T)−n1​∑i=1n​Ui​(f∘ΦT​,T)→0 almost surely;
  • Eq. (53): for a single measurable path whose occupation measures converge weakly to μ\muμ along tj→∞t_j\to\inftytj​→∞, the averages 1njT∑i=0nj−1∫iT(i+1)Tf(xs) ds\frac1{n_jT}\sum_{i=0}^{n_j-1}\int_{iT}^{(i+1)T}f(x_s)\,dsnj​T1​∑i=0nj​−1​∫iT(i+1)T​f(xs​)ds with nj=⌊tj/T⌋n_j=\lfloor t_j/T\rfloornj​=⌊tj​/T⌋ converge to ∫f dμ\int f\,d\mu∫fdμ for every bounded continuous fff.

Significance

The result. Theorem 10.1 locates the long-run statistics of a stochastic process that only shadows a deterministic semiflow in conditional probability. When the occupation measures are tight, for example when the path has compact closure, M(X,ω)\mathcal M(X,\omega)M(X,ω) is nonempty, and the theorem restricts where the process spends its time to the supports of invariant measures. Right after the theorem the notes define the minimal center of attraction of the process from the supports of the measures in M(X,ω)\mathcal M(X,\omega)M(X,ω); the conclusion applies to processes, such as slowly decreasing step-size algorithms, for which the pathwise limit-set theorem of Section 5 is not available.

Formalizing it. The theorem has a complete published proof. No machine-checked version of it, of weak asymptotic pseudotrajectories, or of occupation-measure limit theorems for continuous-time processes is known to exist. The mission produces a formal definition of progressively measurable weak asymptotic pseudotrajectories, occupation measures of measurable paths and their weak limit points, and a proof that combines a martingale law of large numbers in discrete time with weak convergence in P(M)\mathcal P(M)P(M).

Difficulty

The obvious route is to apply the pathwise argument for asymptotic pseudotrajectories along each path. It fails, because condition 2 controls only conditional probabilities: the deviation events may occur infinitely often along almost every path while their conditional probabilities tend to zero. The proof therefore has to work with averages and conditional expectations instead of with individual paths: a strong law of large numbers for bounded martingale differences transfers conditional statements to time averages, and this must be done for one test function and one horizon at a time. Passing from countably many test functions to invariance requires a countable family of uniformly continuous functions that determines weak convergence on the separable space MMM, and the a.s. sets must be intersected over that family and over rational horizons. Measurability is a second difficulty: paths are not assumed continuous, so the integrals, suprema and conditional expectations involved must be shown to be well defined from progressive measurability alone.

Formalization scope

Time is R≥0\mathbb R_{\ge0}R≥0​; the semiflow is Mathlib's Flow ℝ≥0 M; the filtration is a Filtration ℝ≥0; P(M)\mathcal P(M)P(M) is ProbabilityMeasure M with its topology of weak convergence. MMM is a separable metric space with its Borel σ\sigmaσ-algebra; it is not assumed compact, complete or Polish. Progressive measurability is stated literally for every T>0T>0T>0. The conditional probability in condition 2 is the conditional expectation of the indicator of the deviation event, which is required to be measurable (the paper's P{⋅∣Ft}P\{\cdot\mid\mathcal F_t\}P{⋅∣Ft​} presupposes an event); the supremum over h∈[0,T]h\in[0,T]h∈[0,T] is taken in [0,∞][0,\infty][0,∞]. Invariance for the semiflow is (Φt)∗μ=μ(\Phi_t)_*\mu=\mu(Φt​)∗​μ=μ for all t≥0t\ge0t≥0, the form the proof establishes; for a flow it agrees with the definition μ(A)=μ(Φt(A))\mu(A)=\mu(\Phi_t(A))μ(A)=μ(Φt​(A)) of Section 8.3. Weak limit points are cluster points of t↦μt(ω)t\mapsto\mu_t(\omega)t↦μt​(ω) as t→∞t\to\inftyt→∞; the occupation measure is a genuine probability measure for every measurable path and t>0t>0t>0.

The following formalizations would trivialize the statement and are excluded by the definitions: an "occupation measure" equal to the zero measure for a non-measurable path; a conditional probability of a non-measurable event, which Lean evaluates to 000 and which would make condition 2 vacuous; invariance defined through images Φt(A)\Phi_t(A)Φt​(A), which need not be Borel for a semiflow; and a compactness or Polish assumption on MMM, which the theorem does not make.

A complete development needs: Fubini-type measurability for progressively measurable processes, square-integrable martingale convergence and Kronecker's lemma (both largely in Mathlib), conditional expectations of time integrals, a convergence-determining countable family of uniformly continuous functions on a separable metric space, and the identification of weak convergence with convergence of integrals of bounded continuous functions. The martingale law of large numbers (Eqs. (47), (50)) and Eq. (53) are reusable outside this mission. Proofs of any milestone, and alternative arguments for the goal, are welcome.

Selected references

  • M. Benaïm, Dynamics of Stochastic Approximation Algorithms, Séminaire de Probabilités XXXIII, Lecture Notes in Mathematics 1709, Springer, 1999, pp. 1–68. Section 10, Theorem 10.1, pp. 60–63. https://doi.org/10.1007/BFb0096509
  • M. Benaïm and M. W. Hirsch, Asymptotic pseudotrajectories and chain recurrent flows, with applications, Journal of Dynamics and Differential Equations 8 (1996), 141–176. https://doi.org/10.1007/BF02218617
10 thms1 active userReviewed
Machine LearningReinforcement Learning·Captain: mikedeng1

Minimax Regret Bounds for Reinforcement Learning I: High-Probability Regret Bound for UCBVI with a Chernoff–Hoeffding BonusResearch Paper

Motivation

An agent learning to control an unknown environment must balance rewards it can collect now against information that improves later decisions. In a finite Markov decision process (MDP), every action changes the distribution of the next state, so a mistaken transition estimate can affect decisions many steps later. Regret measures this loss against a policy that already knows the transition probabilities. The paper of Azar, Osband and Munos gives high-probability regret bounds for two variants of upper confidence bound value iteration (UCBVI) in finite-horizon reinforcement learning. This mission targets its Chernoff–Hoeffding variant, UCBVI-CH, whose bonus depends only on the horizon and the visit count. Theorem 1 improves the paper's cited earlier dependence on the number of states from SSS to S\sqrt SS​ in the leading term for sufficiently many interactions. Azar, Osband and Munos, 2017, pp. 2, 4–5.

The paper was released in 2017 alongside work on the attainable dependence of episodic regret on the horizon HHH, state count SSS, action count AAA, and total interaction time TTT. Its second algorithm, UCBVI-BF, uses a variance-dependent bonus and is the subject of the next mission in this series. UCBVI-CH has a simpler bonus and its own explicit bound, making it a distinct mathematical target. Azar, Osband and Munos, 2017, pp. 1–5.

Setting

The state set S\mathcal SS and action set A\mathcal AA are finite and nonempty, with cardinalities SSS and AAA. A stationary transition kernel P(y∣x,a)P(y\mid x,a)P(y∣x,a) gives the probability of moving to state yyy after action aaa in state xxx; each row is nonnegative and sums to one. The known, deterministic reward R(x,a)R(x,a)R(x,a) lies in [0,1][0,1][0,1]. An episode lasts H≥1H\ge1H≥1 steps. The environment chooses its starting state xk,1x_{k,1}xk,1​ before episode kkk and may base that choice on earlier episodes. It cannot see the current episode's future random draws. Azar, Osband and Munos, 2017, §2 and Assumption 1, pp. 2–3.

A policy π\piπ selects an action from the current state and the step number. Its value Vhπ(x)V_h^\pi(x)Vhπ​(x) is the expected sum of rewards from step hhh through step HHH when starting in state xxx. The terminal value is VH+1π=0V_{H+1}^\pi=0VH+1π​=0, and Vh∗(x)V_h^*(x)Vh∗​(x) is the maximum of Vhπ(x)V_h^\pi(x)Vhπ​(x) over all such policies. Since the state, action and step sets are finite, this maximum is over a finite nonempty policy class. The paper's sentence describing H−hH-hH−h rewards uses a shifted terminal convention; this series follows the HHH reward steps of Algorithms 1–2. Azar, Osband and Munos, 2017, pp. 3–4.

At the start of episode kkk, UCBVI-CH forms visit counts Nk(x,a,y)N_k(x,a,y)Nk​(x,a,y) and Nk(x,a)N_k(x,a)Nk​(x,a) from earlier completed transitions. On a visited pair it uses the empirical row P^k(y∣x,a)=Nk(x,a,y)/Nk(x,a)\widehat P_k(y\mid x,a)=N_k(x,a,y)/N_k(x,a)Pk​(y∣x,a)=Nk​(x,a,y)/Nk​(x,a). Algorithm 2 computes values backward from zero at the terminal step. For a visited pair, Qk,h(x,a)Q_{k,h}(x,a)Qk,h​(x,a) is the minimum of the preceding episode's Qk−1,h(x,a)Q_{k-1,h}(x,a)Qk−1,h​(x,a), HHH, and the empirical Bellman value plus Algorithm 3's bonus. For an unvisited pair, Qk,h(x,a)=HQ_{k,h}(x,a)=HQk,h​(x,a)=H. A maximizing action is chosen at every state, including states outside the realized path. Azar, Osband and Munos, 2017, Algorithms 1–3, pp. 3–4.

Formalization targets

Theorem 1: UCBVI-CH regret

For KKK episodes and T=KHT=KHT=KH, regret sums the gap V1∗(xk,1)−V1πk(xk,1)V_1^*(x_{k,1})-V_1^{\pi_k}(x_{k,1})V1∗​(xk,1​)−V1πk​​(xk,1​). The goal is the paper's printed bound, with its constants:

Pr⁡ ⁣{Regret⁡(K)>20H3/2LSAK+250H2S2AL2}≤δ,L=ln⁡(5HSAT/δ),δ>0.\Pr\!\left\{\operatorname{Regret}(K)>20H^{3/2}L\sqrt{SAK}+250H^2S^2AL^2\right\}\le\delta, \qquad L=\ln(5HSAT/\delta),\quad \delta>0.Pr{Regret(K)>20H3/2LSAK​+250H2S2AL2}≤δ,L=ln(5HSAT/δ),δ>0.

Algorithm 3 itself uses Lalg=ln⁡(5SAT/δ)L_{\rm alg}=\ln(5SAT/\delta)Lalg​=ln(5SAT/δ) in its bonus 7HLalg/Nk(x,a)7HL_{\rm alg}/\sqrt{N_k(x,a)}7HLalg​/Nk​(x,a)​. Both logarithms remain as printed. The probability is over the MDP's next-state draws, for every admissible starting-state rule and every way of breaking ties between maximizing actions. Azar, Osband and Munos, 2017, Algorithm 3, p. 4; Theorem 1, p. 5.

Supporting results

Four milestones retain the source's indexed attack path: the Bernstein bound (9) for the empirical value error, the count-deviation display before (11), Lemma 18 on optimism, and the weighted recursion displayed in the proof of Lemma 3. The last milestone preserves the signed weights that appear before the paper's final simplification. Azar, Osband and Munos, 2017, pp. 17, 20–21, 28.

Significance

Theorem 1 gives a finite-sample failure probability with explicit dependence on H,S,A,KH,S,A,KH,S,A,K and δ\deltaδ. It covers a learner whose initial state can change between episodes, a feature that matters in episodic learning where the experimenter does not fix a single starting distribution. For the regime stated after Theorem 1, the leading rate is O~(HSAT)\widetilde O(H\sqrt{SAT})O(HSAT​). This is a result claimed by the paper; the present Lean declarations are open proof targets, not machine-checked proofs of that claim. Azar, Osband and Munos, 2017, p. 5.

Formalizing the result creates reusable finite objects for adaptive interaction: a constructed probability law on complete paths, empirical transition counts pooled across steps, a policy value defined by its expected reward, and confidence events with their domains stated explicitly. The concentration and optimism milestones can then be investigated independently of the final regret bound. The later UCBVI-BF mission uses the same paper's model with a different bonus. Azar, Osband and Munos, 2017, pp. 3–5, 14–17.

Difficulty

The visit count Nk(x,a)N_k(x,a)Nk​(x,a) is random and depends on earlier observations and decisions. A concentration inequality for a predetermined number of samples therefore does not immediately give a statement that holds at every episode start. The algorithm also reuses the previous episode's QQQ estimate through a minimum. Any optimism claim must account for this dependence across episodes as well as the backward dependence across steps. In the regret analysis, the terms called martingale differences can have either sign, so replacing a positive weight by a larger common bound can reverse an inequality. These are concrete obstacles to the printed chain of estimates. Azar, Osband and Munos, 2017, pp. 4, 17, 20–21, 28.

Formalization scope

States, actions, steps, episodes and complete outcome arrays are finite. Probabilities are finite sums of products of transition rows. The transition-row predicate is a published general definition; this mission defines the paper-specific reward-bounded MDP, policies, UCBVI-CH recursion, and path law on top of it. The starting-state rule can inspect only earlier episodes. Greedy tie-breaking is universally quantified. V∗V^*V∗ is a maximum over policies, and the bonus is read only at positive counts. A model that assigns an arbitrary probability law, fixes one starting state, or omits Algorithm 2's minimum does not represent this target. Azar, Osband and Munos, 2017, pp. 2–4.

Lean uses steps 0,…,H−10,\dots,H-10,…,H−1 and terminal index HHH in place of the paper's algorithmic 1,…,H+11,\dots,H+11,…,H+1. The appendix sometimes puts the terminal value at HHH. The weighted recursion therefore runs through the final reward step, rather than ending one step early. Its typical-state threshold is 4H2L4H^2L4H2L, as required by (34)–(36), whereas Appendix B.1 prints 2H2L2H^2L2H2L. The proof's correction term c4c_4c4​ dominates its other terms under A≥2A\ge2A≥2, which is made explicit in that milestone. The printed (11) loses a factor of two from the count display before it; only the preceding display is a milestone. Lemma 18 is stated under the empirical-model part of the confidence event and δ≤1\delta\le1δ≤1, the domain on which its bonus comparison holds. The weighted milestone retains its coefficients because the bracketed martingale terms can be negative. Azar, Osband and Munos, 2017, pp. 14–17, 20–21, 28.

The goal retains Theorem 1's constant 202020. Appendix C.1 cites Lemmas 15 and 18, but the sketch of Lemma 15 does not track that constant explicitly. Formalizing the printed bound may therefore expose a gap in its proof; the mission records the claim without weakening its constants. Contributions establishing or repairing the explicit bound, as well as the four stated milestones and reusable finite concentration results, are within scope. Azar, Osband and Munos, 2017, pp. 5, 27, 29.

Selected references

  • M. G. Azar, I. Osband and R. Munos, Minimax Regret Bounds for Reinforcement Learning, arXiv:1703.05449v2, 2017. Pinned preprint.
9 thms1 active userReviewed
Operations ResearchTheoretical Computer Science·Captain: mikedeng1

Secretary Problems: Weights and Discounts 2: An Ω(log n / log log n) Lower Bound on the Competitive Ratio of the Discounted Secretary ProblemResearch Paper

Motivation

In the classical secretary problem a decision maker sees nnn candidates in uniformly random order, learns each candidate's value on arrival, and must accept or reject it on the spot; the goal is to pick a valuable one. A simple sample-then-select rule picks the best candidate with probability at least 1/e1/e1/e, so the problem is constant-competitive. The secretary problem is also a model of online mechanism design: a rule that accepts the first agent above a threshold computed from earlier agents is a truthful posted-price mechanism (as the paper notes in §1).

Babaioff, Dinitz, Gupta, Immorlica and Talwar (SODA 2009; authors' version) study the discounted secretary problem, where accepting at time ttt is worth d(t) v(e)d(t)\,v(e)d(t)v(e) for a known discount function ddd. Discounts model settings where a sale is worth more at some times than at others. The case d(t)=βtd(t)=\beta^td(t)=βt had been studied before (Rasmussen and Pliska 1976); the paper asks what happens for arbitrary ddd. Its answer has two sides: an O(log⁡n)O(\log n)O(logn)-competitive algorithm, and the result of this mission, a lower bound showing that no online algorithm is better than Ω(log⁡n/log⁡log⁡n)\Omega(\log n/\log\log n)Ω(logn/loglogn)-competitive. So, unlike the classical problem, the discounted problem with a general discount is not constant-competitive.

Setting

There are nnn elements e∈{0,…,n−1}e\in\{0,\dots,n-1\}e∈{0,…,n−1} with values v(e)≥0v(e)\ge 0v(e)≥0, and a discount function ddd on the times. The elements arrive in a uniformly random order π\piπ: element π(t)\pi(t)π(t) arrives at time ttt. A randomized online stopping rule AAA specifies, for each time ttt and each sequence of values seen so far h=(v(π(0)),…,v(π(t)))h=(v(\pi(0)),\dots,v(\pi(t)))h=(v(π(0)),…,v(π(t))), a probability pt(h)∈[0,1]p_t(h)\in[0,1]pt​(h)∈[0,1] of stopping at ttt if it has not stopped yet. Stopping at ttt selects π(t)\pi(t)π(t) and earns d(t) v(π(t))d(t)\,v(\pi(t))d(t)v(π(t)); the rule selects at most one element and may select none. The rule knows nnn and ddd, but it sees only values, only as they arrive, and it is not told which instance it is facing.

The expected value of AAA is

E[A]=Eπ[∑td(t) v(π(t)) pt(ht)∏s<t(1−ps(hs))],\mathbb E[A]=\mathbb E_\pi\Bigl[\sum_t d(t)\,v(\pi(t))\,p_t(h_t)\prod_{s<t}\bigl(1-p_s(h_s)\bigr)\Bigr],E[A]=Eπ​[t∑​d(t)v(π(t))pt​(ht​)s<t∏​(1−ps​(hs​))],

and the benchmark is the expected offline optimum

E[OPT]=Eπ[max⁡td(t) v(π(t))],\mathbb E[\mathrm{OPT}]=\mathbb E_\pi\Bigl[\max_t d(t)\,v(\pi(t))\Bigr],E[OPT]=Eπ​[tmax​d(t)v(π(t))],

which is itself a random variable averaged over the order. AAA is α\alphaα-competitive on an instance when E[OPT]≤α E[A]\mathbb E[\mathrm{OPT}]\le\alpha\,\mathbb E[A]E[OPT]≤αE[A].

The hard family (§4.1.1 of the paper): fix an integer c≥1c\ge1c≥1 and put L=cL=cL=c, n=L4cn=L^{4c}n=L4c, nt=L2tn_t=L^{2t}nt​=L2t for t≤2ct\le 2ct≤2c, and K=n2K=n^2K=n2. The step discount is d(j)=L−1d(j)=L^{-1}d(j)=L−1 on the times 1≤j≤n11\le j\le n_11≤j≤n1​ and d(j)=L−td(j)=L^{-t}d(j)=L−t on nt−1<j≤ntn_{t-1}<j\le n_tnt−1​<j≤nt​. The instance I1\mathcal I_1I1​ has n/n1n/n_1n/n1​ elements of value KKK and the rest 000; It+1\mathcal I_{t+1}It+1​ is obtained from It\mathcal I_tIt​ by raising n/nt+1n/n_{t+1}n/nt+1​ of its values KtK^tKt to Kt+1K^{t+1}Kt+1, so It\mathcal I_tIt​ has n/ntn/n_tn/nt​ elements of value KtK^tKt.

Formalization targets

Goal: Theorem 4.3 in the form its proof establishes

For every integer c≥1c\ge1c≥1 and every randomized online stopping rule AAA for horizon n=c4cn=c^{4c}n=c4c and the step discount,

∃ t∈{1,…,2c}:c⋅E[A(It)] < 10⋅E[OPT(It)].\exists\,t\in\{1,\dots,2c\}:\qquad c\cdot\mathbb E[A(\mathcal I_t)]\ <\ 10\cdot\mathbb E[\mathrm{OPT}(\mathcal I_t)].∃t∈{1,…,2c}:c⋅E[A(It​)] < 10⋅E[OPT(It​)].

That is, no online rule is c/10c/10c/10-competitive on all of I1,…,I2c\mathcal I_1,\dots,\mathcal I_{2c}I1​,…,I2c​.

Milestones

  1. Lemma 4.1: E[OPT(It)]≥(1−1/e)KtL−t\mathbb E[\mathrm{OPT}(\mathcal I_t)]\ge(1-1/e)K^tL^{-t}E[OPT(It​)]≥(1−1/e)KtL−t for 1≤t≤2c1\le t\le 2c1≤t≤2c.
  2. Coupling step of Lemma 4.2's proof: for every rule and 1≤t<2c1\le t<2c1≤t<2c, the probability of stopping among the first ntn_tnt​ arrivals drops by at most 1/L21/L^21/L2 from It\mathcal I_tIt​ to It+1\mathcal I_{t+1}It+1​.
  3. Lemma 4.2: a rule that is c/10c/10c/10-competitive on I1,…,I2c\mathcal I_1,\dots,\mathcal I_{2c}I1​,…,I2c​ stops among the first ntn_tnt​ arrivals of It\mathcal I_tIt​ with probability at least t/ct/ct/c.
  4. Theorem 4.3, asymptotic form: for c≥2c\ge2c≥2 and n=c4cn=c^{4c}n=c4c, every rule has some It\mathcal I_tIt​ with
140⋅log⁡nlog⁡log⁡n⋅E[A(It)]<E[OPT(It)].\frac1{40}\cdot\frac{\log n}{\log\log n}\cdot\mathbb E[A(\mathcal I_t)]<\mathbb E[\mathrm{OPT}(\mathcal I_t)].401​⋅loglognlogn​⋅E[A(It​)]<E[OPT(It​)].

Significance

The result separates the discounted secretary problem from its classical and weighted relatives, which admit constant-competitive algorithms (the paper's Theorem 3.4 and the eee-competitive classical rule). Together with the paper's O(log⁡n)O(\log n)O(logn) upper bound (Theorem 4.4) it pins the competitive ratio for general discounts between log⁡n/log⁡log⁡n\log n/\log\log nlogn/loglogn and log⁡n\log nlogn up to constants, and it motivates the paper's known-OPT\mathrm{OPT}OPT model (§4.2), where an estimate of E[OPT]\mathbb E[\mathrm{OPT}]E[OPT] restores a constant ratio. The construction is a template for lower bounds against randomized online algorithms in random-order models: geometrically nested instances that a rule cannot tell apart early, played against a discount that punishes waiting.

The theorem is proved in the paper, in about a page. To our knowledge no part of it has a machine-checked proof. This mission produces the formal model of randomized online stopping rules in the random-order discounted setting, a reusable object for the paper's other discounted results (the O(log⁡n)O(\log n)O(logn) upper bound, and the 2\sqrt22​ lower bound with known values of Theorem 4.6), and a checked version of the lower bound with explicit constants.

Difficulty

The obvious attempt is to fix one instance and show that every rule loses on it. That fails: for any single instance there is a rule tuned to it (a rule that waits exactly as long as that instance warrants). The lower bound has to play the 2c2c2c instances against each other. A rule that does well on It\mathcal I_tIt​ must commit early, within the first ntn_tnt​ steps, yet the rule cannot distinguish It\mathcal I_tIt​ from It+1\mathcal I_{t+1}It+1​ during those steps except with probability L−2L^{-2}L−2. Making "cannot distinguish" precise is the central step: it needs a coupling of the two runs over the same random order and the same internal randomness, which works only because the rule's decision at time ttt depends on the values observed so far and nothing else. The accounting then has to show that the rule's early earnings on It+1\mathcal I_{t+1}It+1​ and its late earnings are both small compared with E[OPT(It+1)]\mathbb E[\mathrm{OPT}(\mathcal I_{t+1})]E[OPT(It+1​)], which uses L≥2L\ge 2L≥2 and that K=n2K=n^2K=n2 dwarfs L2cL^{2c}L2c.

Formalization scope

  • Elements and times are Fin n, 0-based: index jjj is the paper's time j+1j+1j+1, so the paper's block (nt−1,nt](n_{t-1},n_t](nt−1​,nt​] is the index range [nt−1,nt)[n_{t-1},n_t)[nt−1​,nt​). The random order is π : Equiv.Perm (Fin n) read as time ↦\mapsto↦ element, and every expectation over it is the finite average 1n!∑π\frac1{n!}\sum_\pin!1​∑π​. Values and discounts are real.
  • Algorithms are the structure StoppingRule n: stopping probabilities pt(h)∈[0,1]p_t(h)\in[0,1]pt​(h)∈[0,1] indexed by time and the arrival-ordered value sequence, with the non-anticipation condition that pt(h)p_t(h)pt​(h) depends only on h0,…,hth_0,\dots,h_th0​,…,ht​. The theorem quantifies over all such rules, so it covers deterministic and randomized online algorithms that observe values only. A rule may depend on nnn and ddd but not on the instance index.
  • OPT is Eπ[max⁡td(t)v(π(t))]\mathbb E_\pi[\max_t d(t)v(\pi(t))]Eπ​[maxt​d(t)v(π(t))] (a supremum over the finite type Fin n), and competitiveness is multiplicative, E[OPT]≤α E[A]\mathbb E[\mathrm{OPT}]\le\alpha\,\mathbb E[A]E[OPT]≤αE[A], never a quotient.
  • Constants. The goal uses the paper's constant 101010 (from "if AAA is c/10c/10c/10-competitive"); the asymptotic form uses 1/401/401/40, from log⁡n/log⁡log⁡n≤4c\log n/\log\log n\le 4clogn/loglogn≤4c for c≥2c\ge2c≥2, with the natural logarithm. K=n2K=n^2K=n2, the value the paper suggests.
  • The construction (nnn, ntn_tnt​, ddd, KKK, It\mathcal I_tIt​) is fixed by explicit formulas in the definition file. A solver cannot choose the discount or the instances, and the goal is not stated for a restricted class of algorithms; a formalization that let the rule see the instance index or future values, or quantified only over threshold rules, would be a different and trivial or weaker theorem. For c<10c<10c<10 the goal is immediate, since E[A]≤E[OPT]\mathbb E[A]\le\mathbb E[\mathrm{OPT}]E[A]≤E[OPT] and E[OPT(It)]>0\mathbb E[\mathrm{OPT}(\mathcal I_t)]>0E[OPT(It​)]>0; the content lies in c≥10c\ge10c≥10. The bound is stated only for the horizons n=c4cn=c^{4c}n=c4c the paper constructs.
  • Needed infrastructure: counting arguments over permutations of Fin n (the probability that a set of mmm elements misses the first kkk positions), the coupling of two value sequences that agree on a prefix, and elementary estimates on geometric sums. The rule model and the permutation-counting lemmas are reusable for the paper's other discounted results. Contributions of these supporting lemmas, as well as proofs of the milestones, are welcome.

Selected references

  • M. Babaioff, M. Dinitz, A. Gupta, N. Immorlica, K. Talwar, Secretary Problems: Weights and Discounts, Proceedings of the 20th ACM-SIAM Symposium on Discrete Algorithms (SODA), 2009. https://doi.org/10.1137/1.9781611973068.135 (authors' full version, the one cited here: https://www.cs.jhu.edu/~mdinitz/papers/secretary.pdf)
  • E. B. Dynkin, Optimal choice of the stopping moment of a Markov process, Doklady Akademii Nauk SSSR, 1963.
  • W. T. Rasmussen, S. R. Pliska, Choosing the maximum from a sequence with a discount function, Applied Mathematics and Optimization 2(3), 1976.
  • T. S. Ferguson, Who solved the secretary problem?, Statistical Science 4(3), 1989. https://doi.org/10.1214/ss/1177012493
7 thms1 active userReviewed
Statistics·Captain: mikedeng1

Weighted Sums of Certain Dependent Random Variables 2: An Iterated-Logarithm Upper Bound for Weighted Conditionally Sub-Gaussian Martingale DifferencesResearch Paper

Motivation

The law of the iterated logarithm (LIL) gives the exact almost-sure size of the fluctuations of a sum of random variables. For independent fair ±1\pm1±1 increments x1,x2,…x_1,x_2,\dotsx1​,x2​,… with partial sums SnS_nSn​, Khinchin (1924) showed that lim sup⁡n∣Sn∣/2nlog⁡log⁡n=1\limsup_n |S_n|/\sqrt{2n\log\log n}=1limsupn​∣Sn​∣/2nloglogn​=1 almost surely; Kolmogorov (1929) extended this to bounded independent increments, and Hartman and Wintner (1941) to independent, identically distributed increments with variance 111. Weighted sums a1x1+⋯+anxna_1x_1+\dots+a_nx_na1​x1​+⋯+an​xn​ appear in summability theory, in stochastic approximation and in the analysis of orthogonal series, and there the natural normalisation replaces nnn by the sum of squared weights.

Independence is often not available. In martingale settings (sequential estimation, online learning, adaptive algorithms) the increments are only conditionally centred given the past. In 1967 Kazuoki Azuma (Azuma 1967) introduced a conditional sub-Gaussian condition on martingale differences, called property [G], and proved for it an iterated-logarithm upper bound for weighted sums. The moment-generating-function bound that drives his proof, display (2.4), is the inequality now known as the Azuma–Hoeffding inequality. This mission formalizes Theorem 2 of that paper and the lemmas its proof rests on.

Timeline.

  • 1924: Khinchin proves the LIL for fair coin tossing.
  • 1929: Kolmogorov proves it for bounded independent increments under a growth condition.
  • 1941: Hartman and Wintner prove it for i.i.d. increments with finite variance.
  • 1963: Hoeffding proves the exponential tail bound for sums of bounded independent variables.
  • 1965: Gaposhkin proves a LIL for weighted (Cesàro and Abel) means of independent variables.
  • 1967: Azuma proves the conditional mgf bound (2.4), its maximal version (Lemma 2) and the upper-half LIL for weighted sums of class [G] martingale differences (Theorem 2).

Setting

Let (Ω,A,P)(\Omega,\mathfrak A,P)(Ω,A,P) be a probability space and (An)n≥0(\mathfrak A_n)_{n\ge0}(An​)n≥0​ an increasing family of sub-σ\sigmaσ-fields of A\mathfrak AA (a filtration). A sequence of real random variables (xn)n≥1(x_n)_{n\ge1}(xn​)n≥1​ is a martingale-difference sequence if each xnx_nxn​ is An\mathfrak A_nAn​-measurable and integrable and E{xn∣An−1}=0E\{x_n\mid\mathfrak A_{n-1}\}=0E{xn​∣An−1​}=0 almost surely.

The sequence satisfies [G] with τ(xn)≤1\tau(x_n)\le1τ(xn​)≤1 if, in addition, for every n≥1n\ge1n≥1 and every real ttt,

E{exp⁡(txn)∣An−1}≤exp⁡(t2/2)a.s.E\{\exp(tx_n)\mid\mathfrak A_{n-1}\}\le\exp(t^2/2)\quad\text{a.s.}E{exp(txn​)∣An−1​}≤exp(t2/2)a.s.

Every martingale-difference sequence with ∣xn∣≤1|x_n|\le1∣xn​∣≤1 almost surely has this property, but the class also contains unbounded increments, for instance conditionally standard Gaussian ones.

Fix real weights (an)n≥1(a_n)_{n\ge1}(an​)n≥1​ of arbitrary sign and write

Dn2=∑j=1naj2,Sn=a1x1+⋯+anxn.D_n^2=\sum_{j=1}^n a_j^2,\qquad S_n=a_1x_1+\dots+a_nx_n .Dn2​=j=1∑n​aj2​,Sn​=a1​x1​+⋯+an​xn​.

For the lemmas the weights are called (bk)(b_k)(bk​), and the maximal partial sum is Sn∗(ω)=max⁡1≤m≤n∣∑k=1mbkxk(ω)∣S_n^*(\omega)=\max_{1\le m\le n}\big|\sum_{k=1}^m b_kx_k(\omega)\big|Sn∗​(ω)=max1≤m≤n​​∑k=1m​bk​xk​(ω)​.

In Lean these objects are IsMartingaleDiff, IsCondSubgaussianOne, weightedSum (SnS_nSn​), sqWeightSum (Dn2D_n^2Dn2​) and maxAbsWeightedSum (Sn∗S_n^*Sn∗​), all in the namespace AzumaWeightedSums.IteratedLog.

Formalization targets

Goal: Theorem 2, display (4.2)

If (xn)(x_n)(xn​) satisfies [G] with τ(xn)≤1\tau(x_n)\le1τ(xn​)≤1 and the weights satisfy

an2/Dn2→0,Dn2→∞,a_n^2/D_n^2\to0,\qquad D_n^2\to\infty,an2​/Dn2​→0,Dn2​→∞,

then

lim sup⁡n→∞∣Sn∣2Dn2log⁡log⁡Dn2≤1a.s.\limsup_{n\to\infty}\frac{|S_n|}{\sqrt{2D_n^2\log\log D_n^2}}\le1\quad\text{a.s.}n→∞limsup​2Dn2​loglogDn2​​∣Sn​∣​≤1a.s.

The constant 111 is sharp, as Gaussian increments show, so the goal is stated with the paper's constant and in no weaker form.

Milestones

  1. Display (2.4). For every nnn, every real (bk)(b_k)(bk​) and every real ttt,
E{exp⁡(t∑k=1nbkxk)}≤exp⁡(t22∑k=1nbk2).E\Big\{\exp\Big(t\sum_{k=1}^n b_kx_k\Big)\Big\}\le\exp\Big(\frac{t^2}{2}\sum_{k=1}^n b_k^2\Big).E{exp(tk=1∑n​bk​xk​)}≤exp(2t2​k=1∑n​bk2​).
  1. Doob's LαL^\alphaLα maximal inequality, cited on p. 359: for a nonnegative submartingale (fm)(f_m)(fm​) and α>1\alpha>1α>1, E{(max⁡m≤nfm)α}≤(α/(α−1))αE{fnα}E\{(\max_{m\le n}f_m)^\alpha\}\le(\alpha/(\alpha-1))^\alpha E\{f_n^\alpha\}E{(maxm≤n​fm​)α}≤(α/(α−1))αE{fnα​}.
  2. Lemma 2, display (2.3). E{exp⁡(tSn∗)}≤8exp⁡(t22∑k=1nbk2)E\{\exp(tS_n^*)\}\le8\exp\big(\frac{t^2}{2}\sum_{k=1}^n b_k^2\big)E{exp(tSn∗​)}≤8exp(2t2​∑k=1n​bk2​) for every real ttt.
  3. Maximal tail bound. If Vn=∑k=1nbk2>0V_n=\sum_{k=1}^n b_k^2>0Vn​=∑k=1n​bk2​>0 and λ≥0\lambda\ge0λ≥0, then P{Sn∗>λ}≤8exp⁡(−λ2/(2Vn))P\{S_n^*>\lambda\}\le8\exp(-\lambda^2/(2V_n))P{Sn∗​>λ}≤8exp(−λ2/(2Vn​)).

Significance

The result. Theorem 2 controls weighted sums of dependent increments almost surely, uniformly in nnn, at the iterated-logarithm scale. It needs no independence and no boundedness: a conditional sub-Gaussian bound is enough. Bounded martingale differences are a special case, so the theorem covers martingale noise in stochastic approximation and the error terms of adaptive estimators. Milestones 1, 3 and 4 are reusable concentration inequalities: the conditional-expectation form of the Azuma–Hoeffding bound, and its maximal version with an explicit constant.

Formalizing it. The theorem is proved in the literature; nothing here is open. Mathlib has the sub-Gaussian mgf bound for sums in kernel form (HasSubgaussianMGF.sum_of_hasCondSubgaussianMGF, under a standard Borel assumption that this mission does not make), and Doob's weak-type maximal inequality (Submartingale.maximal_ineq). It has no LpL^pLp form of Doob's inequality and no law of the iterated logarithm of any kind. Prove2Me has an Azuma–Hoeffding tail bound without the maximum, and nothing at the iterated-logarithm scale. A complete development would therefore add Doob's LαL^\alphaLα inequality and the first machine-checked iterated-logarithm upper bound, for a dependent class.

Difficulty

The tail bound (milestone 4) is not enough on its own: applied at each fixed nnn and summed over nnn, it gives a divergent series at the 2Dn2log⁡log⁡Dn2\sqrt{2D_n^2\log\log D_n^2}2Dn2​loglogDn2​​ scale. The proof has to pass to a subsequence of times and control the maximum over each block, and the blocks have to be fine enough that no constant is lost. The paper's printed proof chooses blocks along which Dn2D_n^2Dn2​ roughly doubles. Its third displayed estimate uses the increment Dnk+12−Dnk2D_{n_{k+1}}^2-D_{n_k}^2Dnk+1​2​−Dnk​2​ where the maximal inequality actually delivers the full Dnk+12D_{n_{k+1}}^2Dnk+1​2​; with that correction, blocks of ratio 222 prove (4.2) only with 2\sqrt22​ in place of 111. A faithful formal proof must recover the constant 111, so the block ratio has to be tuned to ε\varepsilonε. Here the hypothesis an2/Dn2→0a_n^2/D_n^2\to0an2​/Dn2​→0 is essential, because it means a single term cannot carry Dn2D_n^2Dn2​ past the next block boundary.

Lemma 2 needs Doob's inequality in LαL^\alphaLα form for every even integer α=2j\alpha=2jα=2j, uniformly enough to sum an exponential series. The weak-type inequality available in Mathlib does not give this directly.

Formalization scope

  • Index base and filtration. Sequences are ℕ → Ω → ℝ with sums over Finset.Icc 1 n; x0x_0x0​ and a0a_0a0​ are ignored. The paper fixes A0={∅,Ω}\mathfrak A_0=\{\emptyset,\Omega\}A0​={∅,Ω}; here A0\mathfrak A_0A0​ is arbitrary, which makes every statement at least as strong.
  • [G]. For each n≥1n\ge1n≥1 and each real ttt, the conditional bound holds almost surely, in the paper's quantifier order, and exp⁡(txn)\exp(tx_n)exp(txn​) is assumed integrable. Without integrability, Lean's conditional expectation is 000 and the hypothesis would be empty. τ\tauτ is not defined as an infimum: "τ(xn)≤1\tau(x_n)\le1τ(xn​)≤1" is stated as admissibility of the constant 111.
  • Expectations. Expectations of exponentials and powers are lower Lebesgue integrals in [0,∞][0,\infty][0,∞], so the inequalities also assert finiteness. A Bochner integral, which is 000 for non-integrable functions, would make them trivially true.
  • The lim sup is not Lean's real-valued limsup, which is 000 on unbounded sequences. The goal states the equivalent: for every ε>0\varepsilon>0ε>0, almost surely, ∣Sn∣≤(1+ε)2Dn2log⁡log⁡Dn2|S_n|\le(1+\varepsilon)\sqrt{2D_n^2\log\log D_n^2}∣Sn​∣≤(1+ε)2Dn2​loglogDn2​​ for all sufficiently large nnn. Because Dn2→∞D_n^2\to\inftyDn2​→∞, log⁡log⁡Dn2>0\log\log D_n^2>0loglogDn2​>0 for those nnn, so Real.log is never evaluated at a junk argument that matters.
  • Ruled-out trivialisations. Dropping an2/Dn2→0a_n^2/D_n^2\to0an2​/Dn2​→0, replacing [G] by boundedness, weakening the constant 111, or stating the bound with Lean's real limsup would each change the theorem. None of these is used.
  • Infrastructure. A complete proof needs conditional-expectation pull-out lemmas for the induction in (2.4), Doob's LpL^pLp inequality (reusable well beyond this mission), Chernoff's bound and the first Borel–Cantelli lemma (both in Mathlib), and a block construction. Proofs of any milestone are welcome, as is a proof of Doob's LαL^\alphaLα inequality in Mathlib's own form.

Selected references

  • K. Azuma, Weighted sums of certain dependent random variables, Tôhoku Mathematical Journal 19 (1967) 357–367. https://doi.org/10.2748/tmj/1178243286
  • J. L. Doob, Stochastic Processes, Wiley, New York, 1953 (the LαL^\alphaLα maximal inequality, p. 317; reference [2] of Azuma 1967).
  • W. Hoeffding, Probability inequalities for sums of bounded random variables, Journal of the American Statistical Association 58 (1963) 13–30. https://doi.org/10.1080/01621459.1963.10500830
  • P. Hartman and A. Wintner, On the law of the iterated logarithm, American Journal of Mathematics 63 (1941) 169–176. https://doi.org/10.2307/2371287
  • V. F. Gaposhkin, The law of the iterated logarithm for Cesàro's and Abel's methods of summation, Theory of Probability and its Applications 10 (1965) 411–420 (reference [3] of Azuma 1967).
7 thms1 active userReviewed
Control TheoryMathematical PhysicsOptimization+2·Captain: mikedeng1

Optimization of Mean-field Spin Glasses III: The Lagrangian Value of the Stochastic Control Problem Equals the Parisi FunctionalResearch Paper

Motivation

The ground-state energy of a mixed ppp-spin spin glass, OPTN=max⁡σ∈{±1}NHN(σ)/N\mathsf{OPT}_N=\max_{\sigma\in\{\pm1\}^N}H_N(\sigma)/NOPTN​=maxσ∈{±1}N​HN​(σ)/N, converges almost surely to the infimum of the Parisi functional over non-decreasing order parameters (Auffinger–Chen 2017). El Alaoui, Montanari and Sellke (arXiv:2001.00904v1) study algorithms that find near-optimal configurations. They introduce incremental approximate message passing (IAMP) and show that, among a broad class of such algorithms, the best achievable energy is the infimum of the Parisi functional over a larger space of order parameters (their Theorem 4).

The upper bound in Theorem 4 is reduced, in Section 4 of the paper, to a stochastic optimal control problem. The energy reached by a message-passing algorithm becomes the objective of a control problem driven by a Brownian motion, with a terminal constraint and a variance constraint. The variance constraint is removed by a Lagrange multiplier 12ξ′′γ\tfrac12\xi''\gamma21​ξ′′γ. Proposition 4.1 states that the resulting Lagrangian value is exactly the Parisi functional P(γ)\mathsf P(\gamma)P(γ). This mission formalizes that duality and the verification argument behind it (Section 7).

Setting

Mixture. Real coefficients (ck)k≥2(c_k)_{k\ge2}(ck​)k≥2​ define the mixture ξ(t)=∑k≥2ck2tk\xi(t)=\sum_{k\ge2}c_k^2t^kξ(t)=∑k≥2​ck2​tk, with the standing assumption ξ(1+ε)<∞\xi(1+\varepsilon)<\inftyξ(1+ε)<∞ for some ε>0\varepsilon>0ε>0. Its derivatives ξ′\xi'ξ′, ξ′′\xi''ξ′′ are nonnegative and nondecreasing on [0,1][0,1][0,1].

Order parameters. SF+\mathsf{SF}_+SF+​ is the set of nonnegative step functions

γ=∑i=1mγi I[ti−1,ti),0=t0<t1<⋯<tm=1, γi≥0.\gamma=\sum_{i=1}^m\gamma_i\,\mathbb I_{[t_{i-1},t_i)},\qquad 0=t_0<t_1<\dots<t_m=1,\ \gamma_i\ge0 .γ=i=1∑m​γi​I[ti−1​,ti​)​,0=t0​<t1​<⋯<tm​=1, γi​≥0.

Put ν(t)=∫t1ξ′′(s)γ(s) ds\nu(t)=\int_t^1\xi''(s)\gamma(s)\,dsν(t)=∫t1​ξ′′(s)γ(s)ds.

Parisi PDE and functional. Φγ:[0,1]×R→R\Phi_\gamma:[0,1]\times\mathbb R\to\mathbb RΦγ​:[0,1]×R→R solves

∂tΦγ+12ξ′′(t)(∂x2Φγ+γ(t)(∂xΦγ)2)=0,Φγ(1,x)=∣x∣.\partial_t\Phi_\gamma+\tfrac12\xi''(t)\big(\partial_x^2\Phi_\gamma+\gamma(t)(\partial_x\Phi_\gamma)^2\big)=0,\qquad\Phi_\gamma(1,x)=|x| .∂t​Φγ​+21​ξ′′(t)(∂x2​Φγ​+γ(t)(∂x​Φγ​)2)=0,Φγ​(1,x)=∣x∣.

For γ∈SF+\gamma\in\mathsf{SF}_+γ∈SF+​ it is given explicitly by the Cole–Hopf recursion: with r(t)=ξ′(1)−ξ′(t)r(t)=\xi'(1)-\xi'(t)r(t)=ξ′(1)−ξ′(t) and G∼N(0,1)G\sim\mathsf N(0,1)G∼N(0,1), for t∈[ti−1,ti)t\in[t_{i-1},t_i)t∈[ti−1​,ti​),

Φγ(t,x)=1γilog⁡Eexp⁡{γiΦγ(ti,x+r(t)−r(ti) G)}.\Phi_\gamma(t,x)=\frac1{\gamma_i}\log\mathbb E\exp\big\{\gamma_i\Phi_\gamma(t_i,x+\sqrt{r(t)-r(t_i)}\,G)\big\}.Φγ​(t,x)=γi​1​logEexp{γi​Φγ​(ti​,x+r(t)−r(ti​)​G)}.

The Parisi functional is P(γ)=Φγ(0,0)−12∫01t ξ′′(t)γ(t) dt\mathsf P(\gamma)=\Phi_\gamma(0,0)-\tfrac12\int_0^1t\,\xi''(t)\gamma(t)\,dtP(γ)=Φγ​(0,0)−21​∫01​tξ′′(t)γ(t)dt.

Control problem. Let BBB be a standard Brownian motion. A control u∈D[t,1]u\in D[t,1]u∈D[t,1] is a process on [t,1][t,1][t,1], progressively measurable for the filtration of (Br)r∈[t,1](B_r)_{r\in[t,1]}(Br​)r∈[t,1]​, with E∫t1ξ′′(s)us2 ds<∞\mathbb E\int_t^1\xi''(s)u_s^2\,ds<\inftyE∫t1​ξ′′(s)us2​ds<∞. The value is

Jγ(t,z)=sup⁡u∈D[t,1]E[∫t1ξ′′(s)us ds+12∫t1ν(s)(ξ′′(s)us2−1)ds]s.t.z+∫t1ξ′′(s) us dBs∈(−1,1) a.s.\mathcal J_\gamma(t,z)=\sup_{u\in D[t,1]}\mathbb E\Big[\int_t^1\xi''(s)u_s\,ds+\frac12\int_t^1\nu(s)\big(\xi''(s)u_s^2-1\big)ds\Big]\quad\text{s.t.}\quad z+\int_t^1\sqrt{\xi''(s)}\,u_s\,dB_s\in(-1,1)\ \text{a.s.}Jγ​(t,z)=u∈D[t,1]sup​E[∫t1​ξ′′(s)us​ds+21​∫t1​ν(s)(ξ′′(s)us2​−1)ds]s.t.z+∫t1​ξ′′(s)​us​dBs​∈(−1,1) a.s.

Candidate value function. With Φγ∗(t,z)=inf⁡x{Φγ(t,x)−xz}\Phi^*_\gamma(t,z)=\inf_x\{\Phi_\gamma(t,x)-xz\}Φγ∗​(t,z)=infx​{Φγ​(t,x)−xz},

V(t,z)=Φγ∗(t,z)−12ν(t)z2−12∫t1ν(s) ds.V(t,z)=\Phi^*_\gamma(t,z)-\tfrac12\nu(t)z^2-\tfrac12\int_t^1\nu(s)\,ds .V(t,z)=Φγ∗​(t,z)−21​ν(t)z2−21​∫t1​ν(s)ds.

Formalization targets

Goal: Proposition 4.1

Jγ(0,0)=P(γ)for every γ∈SF+.\mathcal J_\gamma(0,0)=\mathsf P(\gamma)\qquad\text{for every }\gamma\in\mathsf{SF}_+ .Jγ​(0,0)=P(γ)for every γ∈SF+​.

Milestones

  1. Lemma 7.2 (a)–(e): Φγ(t,⋅)\Phi_\gamma(t,\cdot)Φγ​(t,⋅) is smooth for t<1t<1t<1, with derivatives jointly continuous on [0,1)×R[0,1)\times\mathbb R[0,1)×R and C1C^1C1 in time where γ\gammaγ is constant. The range of ∂xΦγ(t,⋅)\partial_x\Phi_\gamma(t,\cdot)∂x​Φγ​(t,⋅) is (−1,1)(-1,1)(−1,1), it is strictly increasing, and 0<∂x2Φγ(t′,x)≤C(t,γ)0<\partial_x^2\Phi_\gamma(t',x)\le C(t,\gamma)0<∂x2​Φγ​(t′,x)≤C(t,γ) for t′≤tt'\le tt′≤t.
  2. Envelope identities (proof of Lemma 7.3): ∂zΦγ∗(t,z)=−xt∗(z)\partial_z\Phi^*_\gamma(t,z)=-x^*_t(z)∂z​Φγ∗​(t,z)=−xt∗​(z) and ∂z2Φγ∗(t,z)=−1/∂x2Φγ(t,xt∗(z))\partial_z^2\Phi^*_\gamma(t,z)=-1/\partial_x^2\Phi_\gamma(t,x^*_t(z))∂z2​Φγ∗​(t,z)=−1/∂x2​Φγ​(t,xt∗​(z)), where xt∗(z)x^*_t(z)xt∗​(z) is the unique root of ∂xΦγ(t,x)=z\partial_x\Phi_\gamma(t,x)=z∂x​Φγ​(t,x)=z.
  3. Lemma 7.3: VVV solves the HJB equation
∂tV+ξ′′(t)sup⁡λ∈R{λ+λ22(ν(t)+∂z2V)}−12ν(t)=0,V(1,z)=0.\partial_tV+\xi''(t)\sup_{\lambda\in\mathbb R}\Big\{\lambda+\frac{\lambda^2}{2}\big(\nu(t)+\partial_z^2V\big)\Big\}-\frac12\nu(t)=0,\qquad V(1,z)=0 .∂t​V+ξ′′(t)λ∈Rsup​{λ+2λ2​(ν(t)+∂z2​V)}−21​ν(t)=0,V(1,z)=0.
  1. Evaluation at the origin: V(0,0)=P(γ)V(0,0)=\mathsf P(\gamma)V(0,0)=P(γ).
  2. Proposition 7.1: Jγ(t,z)=V(t,z)\mathcal J_\gamma(t,z)=V(t,z)Jγ​(t,z)=V(t,z) for all (t,z)∈[0,1]×(−1,1)(t,z)\in[0,1]\times(-1,1)(t,z)∈[0,1]×(−1,1).

Proposition 7.1 at (0,0)(0,0)(0,0) together with milestone 4 gives the goal.

Significance

The result. By integration by parts (Eq. (4.4) of the paper), Jγ(0,0)\mathcal J_\gamma(0,0)Jγ​(0,0) bounds the value of the constrained control problem (4.2). That problem in turn bounds the asymptotic energy of every message-passing algorithm in the class of Theorem 4. Proposition 4.1 turns the bound into inf⁡γ∈SF+P(γ)\inf_{\gamma\in\mathsf{SF}_+}\mathsf P(\gamma)infγ∈SF+​​P(γ), which is the analytic core of the optimality statement for IAMP. It is also an instance of a broader principle: the Parisi functional has a stochastic-control representation (Jagannath–Tobasco 2016).

Formalizing it. The result is proved in the paper; no machine-checked version exists. A complete formalization needs a verification theorem for a control problem with a state constraint (M1∈(−1,1)M_1\in(-1,1)M1​∈(−1,1)), Itô's formula for a C1,2C^{1,2}C1,2 function that is only piecewise C1C^1C1 in time, and quantitative regularity of the Cole–Hopf solution. Each of these is reusable well beyond spin glasses.

Difficulty

The value function Jγ\mathcal J_\gammaJγ​ is not known to be smooth, and the dynamic-programming equation (4.6) is only heuristic. The proof therefore guesses a solution and verifies it. Two steps carry the difficulty.

First, the guess VVV is a Legendre transform. Its regularity, and the sign ν+∂z2V<0\nu+\partial_z^2V<0ν+∂z2​V<0 that makes the HJB supremum finite, rest on strict convexity and bounded curvature of Φγ(t,⋅)\Phi_\gamma(t,\cdot)Φγ​(t,⋅) (Lemma 7.2). These must be proved by induction through the Cole–Hopf recursion, including the steps with γi=0\gamma_i=0γi​=0.

Second, the verification argument applies Itô's formula to V(s,Msu)V(s,M^u_s)V(s,Msu​), where MuM^uMu is a martingale confined to (−1,1)(-1,1)(−1,1) and VVV is only C1C^1C1 in time between the jumps of γ\gammaγ. The boundary θ→1\theta\to1θ→1 needs a dominated-convergence argument, and attaining the supremum needs an explicit optimal feedback control built from an SDE.

Formalization scope

  • Mixture. ξ\xiξ is a coefficient sequence c:N→Rc:\mathbb N\to\mathbb Rc:N→R with c0=c1=0c_0=c_1=0c0​=c1​=0 imposed. ξ′\xi'ξ′ and ξ′′\xi''ξ′′ are explicit termwise series.
  • Step functions. SF+\mathsf{SF}_+SF+​ is represented by its data (breakpoints and values). γ\gammaγ is extended by 000 outside [0,1)[0,1)[0,1); its value at t=1t=1t=1 never matters.
  • Cole–Hopf. Φγ\Phi_\gammaΦγ​ is defined by the recursion. When γi=0\gamma_i=0γi​=0, the recursion uses its limit, the heat semigroup, instead of dividing by zero. Expectations over GGG are integrals against gaussianReal 0 1.
  • Derivatives. Space derivatives are deriv/iteratedDeriv. Time derivatives are right derivatives, because γ\gammaγ jumps.
  • Legendre transform. Φγ∗\Phi^*_\gammaΦγ∗​ is a real infimum, used only for ∣z∣<1|z|<1∣z∣<1, where it is bounded below.
  • Brownian motion and filtration. BBB is a Mathlib IsBrownianReal process on R≥0\mathbb R_{\ge0}R≥0​, with each BrB_rBr​ measurable. The filtration is Fst=σ(Br:t≤r≤s)\mathcal F^t_s=\sigma(B_r:t\le r\le s)Fst​=σ(Br​:t≤r≤s).
  • Stochastic integral. It is the L2L^2L2 Itô integral of the published definition Peng1990.SMP.IsItoIntegral (horizon 111), whose integrability class is exactly E∫ξ′′u2<∞\mathbb E\int\xi''u^2<\inftyE∫ξ′′u2<∞.
  • Supremum. Jγ(t,z)=v\mathcal J_\gamma(t,z)=vJγ​(t,z)=v is stated as "vvv is the least upper bound of the objective values of admissible controls" (IsLUB), never as a real sSup. A default value of an empty or unbounded supremum therefore cannot make a statement trivially true.
  • Disclosed hypothesis. Lemma 7.2, the envelope identities and Lemma 7.3 assume that ξ\xiξ is not identically zero (some ck≠0c_k\neq0ck​=0). For ξ≡0\xi\equiv0ξ≡0 one has Φγ(t,x)=∣x∣\Phi_\gamma(t,x)=|x|Φγ​(t,x)=∣x∣ for all ttt, and these statements fail. Proposition 7.1, the evaluation at the origin and the goal need no such hypothesis.
  • Lemma 7.3. The statement includes the inequality ν+∂z2V<0\nu+\partial_z^2V<0ν+∂z2​V<0, which the page proves. This rules out reading the HJB supremum as a default value.

The paper's algorithmic results (Theorems 2–4, Corollary 2.2) are out of scope. They need the AMP and state-evolution machinery of Section 5 and Appendix A, and an informal model of computation. The optional bound (4.4) is not stated.

Welcome contributions include the regularity of Cole–Hopf solutions (Gaussian convolution, log-moment-generating functions), a general verification theorem for one-dimensional controlled martingales with a terminal state constraint, and Itô's formula for C1,2C^{1,2}C1,2 functions.

Selected references

  • A. El Alaoui, A. Montanari, M. Sellke, Optimization of Mean-field Spin Glasses, arXiv:2001.00904v1, 2020. https://arxiv.org/abs/2001.00904
  • A. Auffinger, W.-K. Chen, Parisi formula for the ground state energy in the mixed p-spin model, Ann. Probab. 45(6b), 2017. https://arxiv.org/abs/1606.05335
  • A. Jagannath, I. Tobasco, A dynamic programming approach to the Parisi functional, Proc. AMS 144, 2016. https://arxiv.org/abs/1502.04398
  • N. Touzi, Optimal Stochastic Control, Stochastic Target Problems, and Backward SDE, Fields Institute Monographs 29, Springer, 2012 (cited as [Tou12] in the paper; the verification argument of Section 7 follows its Theorem 4.1).
14 thms1 active userReviewed
PreviousPage 8 of 11Next

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me