Prove2Me
Navigate
DiscoverCollectionsFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

Operations Research

1,660 missions · 832 completed

The discipline of applying mathematical analysis to complex decision problems in operations: allocating scarce resources, scheduling, routing, inventory, and the design of service and production systems. Drawing on mathematical programming, stochastic modeling, queueing, simulation, and game-theoretic reasoning, it seeks policies that perform provably well in systems shaped by constraints, congestion, and uncertainty.

Missions

Open828Completed832All1660
OptimizationProbability·Captain: mikedeng1

Supply Chain Coordination Under Channel Rebates with Sales Effort Effects IV: A Larger Return Credit Strictly Increases the Retailer's Sales EffortResearch Paper

Returns and the retailer's incentive for sales effort

A manufacturer that sells through a retailer can accept returns: it pays a return credit bbb for every unit the retailer has not sold at the end of the season. Returns policies are common in publishing, software and computer hardware, and since Pasternack (Marketing Science, 1985) they have been studied as a way to make the retailer order more. Retailers also raise demand through sales effort: merchandising, shelf space, point-of-sale advertising. A recurring view in the marketing and operations literature is that returns weaken this incentive. Padmanabhan and Png (1995, p. 70) write that "by reducing the risk of losses due to excess inventory, a returns policy lessens some of the retailer's incentive to invest in such efforts", and Kandel (1996, p. 348) makes the same argument for consignment.

T. A. Taylor, Supply Chain Coordination Under Channel Rebates with Sales Effort Effects (Management Science 48(8), 2002), studies a newsvendor retailer who chooses an order quantity and a sales effort level before demand is observed. Its Proposition 4 shows that in this model the conventional view is reversed: a larger return credit makes the retailer exert strictly more effort. This mission formalizes that proposition and the solution of the retailer's problem it rests on.

Setting

Prices satisfy 0<c<w<p0<c<w<p0<c<w<p and s<cs<cs<c (Assumption A1): ppp is the retail price, www the wholesale price, ccc the manufacturing cost and sss the salvage value, which may be negative. A demand factor ξ≥0\xi\ge 0ξ≥0 has a density φ\varphiφ with φ(ξ)>0\varphi(\xi)>0φ(ξ)>0 for every ξ≥0\xi\ge0ξ≥0 (Assumption A4), distribution function Φ(Q)=∫0Qφ\Phi(Q)=\int_0^Q\varphiΦ(Q)=∫0Q​φ and partial mean Γ(Q)=∫0Qξ dΦ(ξ)\Gamma(Q)=\int_0^Q\xi\,d\Phi(\xi)Γ(Q)=∫0Q​ξdΦ(ξ).

The retailer chooses an effort level e≥0e\ge0e≥0, and demand is eξe\xieξ. Effort costs V(e)V(e)V(e), where VVV is strictly convex and strictly increasing with V(0)=0V(0)=0V(0)=0 (Assumption A5; the paper declares every convexity and monotonicity statement strict). Write Λ(γ)=γV′(γ)−V(γ)\Lambda(\gamma)=\gamma V'(\gamma)-V(\gamma)Λ(γ)=γV′(γ)−V(γ).

Under a returns-only contract (w,b)(w,b)(w,b) with b∈[s,w)b\in[s,w)b∈[s,w), the retailer orders Q≥0Q\ge0Q≥0, exerts effort e≥0e\ge 0e≥0, and earns the expected profit

R‾b(Q,e)=−wQ+p Emin⁡(Q,eξ)+b E(Q−eξ)+−V(e).\underline R_b(Q,e)=-wQ+p\,E\min(Q,e\xi)+b\,E(Q-e\xi)^+-V(e).R​b​(Q,e)=−wQ+pEmin(Q,eξ)+bE(Q−eξ)+−V(e).

This is the integrated channel's profit Π(Q,e)=−cQ+pEmin⁡(Q,eξ)+sE(Q−eξ)+−V(e)\Pi(Q,e)=-cQ+pE\min(Q,e\xi)+sE(Q-e\xi)^+-V(e)Π(Q,e)=−cQ+pEmin(Q,eξ)+sE(Q−eξ)+−V(e) with www in place of ccc and bbb in place of sss. An optimal pair is a maximizer of R‾b\underline R_bR​b​ over Q≥0Q\ge0Q≥0, e≥0e\ge0e≥0. The paper's e‾\underline ee​ is the effort of such a pair, and Q‾0\underline Q_0Q​0​ is the critical fractile Φ(Q‾0)=(p−w)/(p−b)\Phi(\underline Q_0)=(p-w)/(p-b)Φ(Q​0​)=(p−w)/(p−b).

Formalization targets

Goal: Proposition 4 (p. 1004)

If s≤b2<b1<ws\le b_2<b_1<ws≤b2​<b1​<w, (Q1,e1)(Q_1,e_1)(Q1​,e1​) is an optimal pair under the credit b1b_1b1​, (Q2,e2)(Q_2,e_2)(Q2​,e2​) is an optimal pair under b2b_2b2​, and e2>0e_2>0e2​>0, then

e2<e1.e_2<e_1 .e2​<e1​.

The paper writes this as ∂e‾/∂b>0\partial\underline e/\partial b>0∂e​/∂b>0. Its proof shows strict monotonicity, and that is the statement here.

Milestone 1: the order for a given effort (§4.2, p. 1000)

For a fixed e>0e>0e>0, the unique maximizer of Q↦R‾b(Q,e)Q\mapsto\underline R_b(Q,e)Q↦R​b​(Q,e) over Q≥0Q\ge0Q≥0 is Q2=e Q‾0Q_2=e\,\underline Q_0Q2​=eQ​0​, and

max⁡Q≥0R‾b(Q,e)=e (p−b) Γ(Q‾0)−V(e).\max_{Q\ge0}\underline R_b(Q,e)=e\,(p-b)\,\Gamma(\underline Q_0)-V(e).Q≥0max​R​b​(Q,e)=e(p−b)Γ(Q​0​)−V(e).

Milestone 2: the optimal effort (§4.2, p. 1000)

An optimal pair (Q,e‾)(Q,\underline e)(Q,e​) with e‾>0\underline e>0e​>0 satisfies

V′(e‾)=(p−b) Γ(Q‾0),Q=e‾ Q‾0,R‾b(Q,e‾)=Λ(e‾),V'(\underline e)=(p-b)\,\Gamma(\underline Q_0),\qquad Q=\underline e\,\underline Q_0,\qquad \underline R_b(Q,\underline e)=\Lambda(\underline e),V′(e​)=(p−b)Γ(Q​0​),Q=e​Q​0​,R​b​(Q,e​)=Λ(e​),

and conversely every e>0e>0e>0 solving this first-order condition gives the optimal pair (eQ‾0,e)(e\underline Q_0,e)(eQ​0​,e).

Significance

The result. Proposition 4 separates two effects of a return credit. For a fixed order quantity, a larger credit lowers the marginal value of effort: it pays the retailer for unsold units, and effort only matters through sold units. This is the fixed-quantity comparison on which the conventional view rests; Cachon's survey states it for buy-backs as Eq. (19) (on Prove2Me as CachonCoord.EffortNewsvendor.sec_6_4_1_effort_coordination, clause 1). Once the order quantity is chosen together with the effort, the larger credit raises the order, and the effort follows. For the multiplicative demand model the net effect is unambiguous for every demand density and every convex effort cost. A manufacturer that wants more retailer effort can therefore use the return credit as a lever, and a model that ignores the quantity response gets the sign wrong.

Formalizing it. The result is proved in the paper, in a four-line appendix argument that relies on the §4.2 solution of the retailer's problem; neither is machine-checked anywhere. The mission produces a checked newsvendor-with-effort solution (the scaling Q2=eQ‾0Q_2=e\underline Q_0Q2​=eQ​0​ and the profit Λ(e‾)\Lambda(\underline e)Λ(e​)), which the other missions of this series and any multiplicative-effort supply-chain model can reuse, and a checked statement of the comparative statics in the return credit.

Difficulty

The obvious argument differentiates the retailer's first-order condition in bbb with the order held fixed, and it gives the wrong sign; the order quantity must be re-optimized. The correct comparison runs through the reduced problem in the effort alone, and it needs the §4.2 solution in full: the scaling of the optimal order with the effort, which turns Emin⁡(Q,eξ)E\min(Q,e\xi)Emin(Q,eξ) and E(Q−eξ)+E(Q-e\xi)^+E(Q−eξ)+ into newsvendor expectations of ξ\xiξ at Q/eQ/eQ/e; the value of the optimal order as a function of eee; and the first-order characterization of the optimal effort. The comparison across credits then involves two objects: the function Λ\LambdaΛ, and the retailer's reduced profits under the two credits at a common effort level. The statement compares maximizers of two different optimization problems, not roots of an equation, so a formal proof must connect optimality of the pair to the first-order condition, which uses the strict convexity of VVV and the positivity of the density.

Formalization scope

Everything lives in the namespace ChannelRebate.ReturnsEffort. The definitions file holds:

  • Demand: a measurable density φ\varphiφ, zero on (−∞,0)(-\infty,0)(−∞,0), positive on [0,∞)[0,\infty)[0,∞), integrating to 111, with finite mean;
  • the law of ξ\xiξ;
  • Φ\PhiΦ and Γ\GammaΓ as interval integrals;
  • EffortCost: A5 with strict convexity and strict monotonicity on [0,∞)[0,\infty)[0,∞) and V(0)=0V(0)=0V(0)=0;
  • Λ\LambdaΛ;
  • the retailer's profit returnsProfit;
  • optimal orders and optimal pairs as maximizers over Q≥0Q\ge0Q≥0 (and e≥0e\ge0e≥0).

The expectations are the published CachonCoord.Newsvendor.expSales and expLeftover applied to the law of eξe\xieξ (the image of the law of ξ\xiξ under x↦exx\mapsto exx↦ex).

The formalization makes the following commitments, each recorded on its item:

  • Φ−1\Phi^{-1}Φ−1 is never an inverse function: Q‾0\underline Q_0Q​0​ is a positive solution of Φ(Q‾0)=(p−w)/(p−b)\Phi(\underline Q_0)=(p-w)/(p-b)Φ(Q​0​)=(p−w)/(p−b).
  • The paper writes (∂/∂e)V(\partial/\partial e)V(∂/∂e)V without stating that VVV is differentiable. Differentiability on (0,∞)(0,\infty)(0,∞), with derivative V′V'V′, is a hypothesis.
  • The paper restricts attention to strictly positive effort (p. 1000). The goal therefore assumes e2>0e_2>0e2​>0; without it both optimal efforts could be 000 when V′(0+)V'(0^+)V′(0+) is large, and the strict inequality would fail.
  • Existence of optimal pairs is a hypothesis, as on p. 999 ("assume the cost of effort function and demand distribution are chosen such that the existence of an optimal solution is assured").
  • The credit satisfies b<wb<wb<w strictly; at b=wb=wb=w the critical fractile is undefined.

The goal is about optimal pairs of the joint problem. Defining e‾\underline ee​ as the root of the first-order condition, or fixing the order quantity and varying bbb, would give a different and, in the second case, false statement; neither is the target.

Contributions are welcome:

  • a proof of the scaling identity Emin⁡(Q,eξ)=e Emin⁡(Q/e,ξ)E\min(Q,e\xi)=e\,E\min(Q/e,\xi)Emin(Q,eξ)=eEmin(Q/e,ξ);
  • a proof of the newsvendor value identity (p−b)Emin⁡(q,ξ)−(w−b)q=(p−b)Γ(q)(p-b)E\min(q,\xi)-(w-b)q=(p-b)\Gamma(q)(p−b)Emin(q,ξ)−(w−b)q=(p−b)Γ(q) at the critical fractile, reusable for any newsvendor with a density;
  • proofs of the two milestones;
  • the goal theorem.

The two facts the paper's proof uses — that Λ\LambdaΛ is strictly increasing on (0,∞)(0,\infty)(0,∞), and that the retailer's reduced profit at a fixed effort increases with the credit — may be posted as supporting theorems.

Selected references

  • T. A. Taylor, Supply Chain Coordination Under Channel Rebates with Sales Effort Effects, Management Science 48(8):992–1007, 2002. https://doi.org/10.1287/mnsc.48.8.992.168
  • B. A. Pasternack, Optimal Pricing and Return Policies for Perishable Commodities, Marketing Science 4(2):166–176, 1985. https://doi.org/10.1287/mksc.4.2.166
  • V. Padmanabhan, I. P. L. Png, Returns Policies: Make Money by Making Good, Sloan Management Review 37(1):65–72, 1995.
  • E. Kandel, The Right to Return, Journal of Law and Economics 39(1):329–356, 1996. https://doi.org/10.1086/467352
  • G. P. Cachon, Supply Chain Coordination with Contracts, in Handbooks in OR & MS 11, Elsevier, 2003, §6.4.1. https://doi.org/10.1016/S0927-0507(03)11006-7
5 thms1 active userReviewed
OptimizationProbability·Captain: mikedeng1

Supply Chain Coordination Under Channel Rebates with Sales Effort Effects II: A Target Rebate and Returns Contract Coordinates Effort and Quantity Under Uniform DemandResearch Paper

Motivation

Manufacturers of computer hardware, software and automobiles routinely pay their retailers channel rebates: a payment per unit the retailer sells to end customers. A target rebate pays only for units sold beyond a target level. These industries also commonly offer returns: a credit for each unsold unit. Both instruments are used to change the retailer's behaviour, and the retailer controls two things that matter to the manufacturer: how much stock she orders, and how much sales effort she exerts to raise demand. Effort cannot be observed or written into a contract, but sales can.

T. A. Taylor's paper (Management Science 48(8), 2002) asks whether a contract built on sales and returns can make an independent retailer choose the effort level and the order quantity that maximize the profit of the whole supply chain, while splitting that profit in any desired proportion. Its Proposition 2 shows that returns alone, linear rebates alone, or target rebates alone cannot do this. Its Theorem 2, the goal of this mission, shows that a target rebate combined with returns can, when demand is uniform and effort cost is quadratic.

Setting

A manufacturer with unit production cost ccc sells to a retailer at wholesale price www; the retailer sells at the fixed retail price ppp, and unsold units have salvage value sss. The standing assumption is 0<c<w<p0<c<w<p0<c<w<p and s<cs<cs<c (sss may be negative). Before demand is seen, the retailer chooses an order quantity Q≥0Q\ge0Q≥0 and an effort level e≥0e\ge0e≥0. Demand is eξe\xieξ, where ξ\xiξ is uniform on [0,1][0,1][0,1], with distribution function Φ\PhiΦ and Γ(x)=∫0xξ dΦ(ξ)\Gamma(x)=\int_0^x\xi\,d\Phi(\xi)Γ(x)=∫0x​ξdΦ(ξ). Effort costs V(e)=ae2/2V(e)=ae^2/2V(e)=ae2/2, a>0a>0a>0.

The integrated channel, which owns both firms, earns

Π(Q,e)=−cQ+pEmin⁡(Q,eξ)+sE(Q−eξ)+−V(e).\Pi(Q,e)=-cQ+pE\min(Q,e\xi)+sE(Q-e\xi)^+-V(e).Π(Q,e)=−cQ+pEmin(Q,eξ)+sE(Q−eξ)+−V(e).

Its optimum uses the critical fractile Qˉ0\bar Q_0Qˉ​0​, defined by Φ(Qˉ0)=(p−c)/(p−s)\Phi(\bar Q_0)=(p-c)/(p-s)Φ(Qˉ​0​)=(p−c)/(p−s), the effort eˉ=(p−s)Γ(Qˉ0)/a\bar e=(p-s)\Gamma(\bar Q_0)/aeˉ=(p−s)Γ(Qˉ​0​)/a and the order Qˉ=eˉQˉ0\bar Q=\bar e\bar Q_0Qˉ​=eˉQˉ​0​. Its optimal profit is Π=Λ(eˉ)\Pi=\Lambda(\bar e)Π=Λ(eˉ), where Λ(γ)=γV′(γ)−V(γ)\Lambda(\gamma)=\gamma V'(\gamma)-V(\gamma)Λ(γ)=γV′(γ)−V(γ).

A target rebate and returns contract (w,u,b,T)(w,u,b,T)(w,u,b,T) pays the retailer u>0u>0u>0 for every unit sold beyond the target TTT, and credits her b∈[s,w)b\in[s,w)b∈[s,w) for every unsold unit. Her profit is

R(Q,e∣T)=−wQ+pEmin⁡(Q,eξ)+uE(min⁡(Q,eξ)−T)++bE(Q−eξ)+−V(e).R(Q,e\mid T)=-wQ+pE\min(Q,e\xi)+uE(\min(Q,e\xi)-T)^+ +bE(Q-e\xi)^+-V(e).R(Q,e∣T)=−wQ+pEmin(Q,eξ)+uE(min(Q,eξ)−T)++bE(Q−eξ)+−V(e).

The manufacturer earns M(Q,e∣T)=(w−c)Q−uE(min⁡(Q,eξ)−T)+−(b−s)E(Q−eξ)+M(Q,e\mid T)=(w-c)Q-uE(\min(Q,e\xi)-T)^+-(b-s)E(Q-e\xi)^+M(Q,e∣T)=(w−c)Q−uE(min(Q,eξ)−T)+−(b−s)E(Q−eξ)+. The other quantities the statements use are:

  • the fractiles Q‾0\underline Q_0Q​0​ and Q‾1\underline Q_1Q​1​, with Φ(Q‾0)=(p−w)/(p−b)\Phi(\underline Q_0)=(p-w)/(p-b)Φ(Q​0​)=(p−w)/(p−b) and Φ(Q‾1)=(p+u−w)/(p+u−b)\Phi(\underline Q_1)=(p+u-w)/(p+u-b)Φ(Q​1​)=(p+u−w)/(p+u−b);
  • the returns-only effort e‾=(p−b)Γ(Q‾0)/a\underline e=(p-b)\Gamma(\underline Q_0)/ae​=(p−b)Γ(Q​0​)/a;
  • a threshold τ∈[Q‾0,Q‾1]\tau\in[\underline Q_0,\underline Q_1]τ∈[Q​0​,Q​1​] that separates the retailer's low and high orders;
  • the retailer's best profit at effort eee, A(e∣T)=max⁡Q≥0R(Q,e∣T)A(e\mid T)=\max_{Q\ge0}R(Q,e\mid T)A(e∣T)=maxQ≥0​R(Q,e∣T).

Contract terms of Theorem 2. Put ζ(T)=4a2(p−s)3T2\zeta(T)=4a^2(p-s)^3T^2ζ(T)=4a2(p−s)3T2 and

u(T)=(w−c)(p−c)5(p−c)5−ζ(T),b(T)=s+(w−c)(p−c)6−(p−s)ζ(T)(p−c)6−(p−c)ζ(T),T3=(p−c)32a(p−s)2.u(T)=(w-c)\frac{(p-c)^5}{(p-c)^5-\zeta(T)},\qquad b(T)=s+(w-c)\frac{(p-c)^6-(p-s)\zeta(T)}{(p-c)^6-(p-c)\zeta(T)},\qquad T_3=\frac{(p-c)^3}{2a(p-s)^2}.u(T)=(w−c)(p−c)5−ζ(T)(p−c)5​,b(T)=s+(w−c)(p−c)6−(p−c)ζ(T)(p−c)6−(p−s)ζ(T)​,T3​=2a(p−s)2(p−c)3​.

T1T_1T1​ and T2T_2T2​ are the fixed points on [0,T3][0,T_3][0,T3​] of T↦e‾τT\mapsto\underline e\tauT↦e​τ and T↦eˉτT\mapsto\bar e\tauT↦eˉτ, with e‾\underline ee​ and τ\tauτ evaluated at u(T)u(T)u(T), b(T)b(T)b(T). L‾(T,w)\underline L(T,w)L​(T,w) and Lˉ(T)\bar L(T)Lˉ(T) are the retailer's profits at effort e‾\underline ee​ and at effort eˉ\bar eeˉ.

Formalization targets

Goal: Theorem 2 (p. 1002)

For every κ∈(0,Π)\kappa\in(0,\Pi)κ∈(0,Π) there is ε0>0\varepsilon_0>0ε0​>0 such that for every ε∈(0,min⁡(κ,ε0))\varepsilon\in(0,\min(\kappa,\varepsilon_0))ε∈(0,min(κ,ε0​)) the following holds. Pairs (w∗,T∗)(w^*,T^*)(w∗,T∗) with

w∗∈(c,p),T∗∈(T1,T2),L‾(T∗,w∗)=κ−ε,Lˉ(T∗)=κw^*\in(c,p),\qquad T^*\in(T_1,T_2),\qquad \underline L(T^*,w^*)=\kappa-\varepsilon,\qquad \bar L(T^*)=\kappaw∗∈(c,p),T∗∈(T1​,T2​),L​(T∗,w∗)=κ−ε,Lˉ(T∗)=κ

exist. Every such pair gives u∗=u(T∗)>0u^*=u(T^*)>0u∗=u(T∗)>0, b∗=b(T∗)∈(s,w∗)b^*=b(T^*)\in(s,w^*)b∗=b(T∗)∈(s,w∗) and T∗>0T^*>0T∗>0. The pair (Qˉ,eˉ)(\bar Q,\bar e)(Qˉ​,eˉ) is the retailer's unique optimum under (w∗,u∗,b∗,T∗)(w^*,u^*,b^*,T^*)(w∗,u∗,b∗,T∗), and

R(Qˉ,eˉ∣T∗)=κ,M(Qˉ,eˉ∣T∗)=Π−κ.R(\bar Q,\bar e\mid T^*)=\kappa,\qquad M(\bar Q,\bar e\mid T^*)=\Pi-\kappa.R(Qˉ​,eˉ∣T∗)=κ,M(Qˉ​,eˉ∣T∗)=Π−κ.

Milestones

  1. §4.1: (Qˉ,eˉ)(\bar Q,\bar e)(Qˉ​,eˉ) is the unique maximizer of Π\PiΠ, with value Λ(eˉ)\Lambda(\bar e)Λ(eˉ).
  2. §4.2: under returns alone, (e‾ Q‾0,e‾)(\underline e\,\underline Q_0,\underline e)(e​Q​0​,e​) is the retailer's unique optimum, with value Λ(e‾)\Lambda(\underline e)Λ(e​).
  3. Lemma 2: at effort eee the retailer orders eQ‾0e\underline Q_0eQ​0​ if e<T/τe<T/\taue<T/τ, eQ‾1e\underline Q_1eQ​1​ if e>T/τe>T/\taue>T/τ, and either if e=T/τe=T/\taue=T/τ.
  4. The two-branch formula for A(e∣T)A(e\mid T)A(e∣T). Its derivative jumps up at T/τT/\tauT/τ, so T/τT/\tauT/τ is never optimal.
  5. Lemma 3: A(⋅∣T)A(\cdot\mid T)A(⋅∣T) is concave on [0,T/τ)[0,T/\tau)[0,T/τ), and either convex then concave, or concave, on (T/τ,∞)(T/\tau,\infty)(T/τ,∞).
  6. Lemma 4: a unique threshold Υ\UpsilonΥ decides whether the optimal effort lies above T/τT/\tauT/τ (at e^\hat ee^) or equals e‾\underline ee​.
  7. Lemma 5: the retailer's optimal (Q,e)(Q,e)(Q,e) in the three cases T<ΥT<\UpsilonT<Υ, T>ΥT>\UpsilonT>Υ, T=ΥT=\UpsilonT=Υ.
  8. Lemma 6: T1T_1T1​ and T2T_2T2​ exist, are unique, and 0<T1<T2<T30<T_1<T_2<T_30<T1​<T2​<T3​.

Significance

The result. Theorem 2 shows that two contractible instruments can align two decisions, one of them not contractible. The rebate pushes effort and quantity up when sales are high, and the return credit raises both when demand is low. Together they reproduce the integrated channel's optimum (Qˉ,eˉ)(\bar Q,\bar e)(Qˉ​,eˉ), and the free parameter κ\kappaκ allocates the profit Π\PiΠ between the firms in any proportion. Proposition 2 of the same paper rules out each instrument alone. This makes Theorem 2 the basis for the paper's recommendation that rebates and returns be used together. The uniform instance is the only one in which the paper proves this; for normal demand it gives only numerical evidence.

Formalizing it. The theorem is proved in the paper; it has not been formalized. The printed proof treats the wholesale price as fixed while it varies: u(T)u(T)u(T) and b(T)b(T)b(T) contain www, so T1T_1T1​, T2T_2T2​ and Lˉ\bar LLˉ move with w∗w^*w∗. A machine-checked proof therefore has to repair the argument, not only transcribe it. A numerical check (p = 10, c = 4, s = 1, a = 1; for example κ = 1.2, ε = 0.012 gives w* ≈ 4.889, T* ≈ 1.095) found the statement itself consistent.

Difficulty

The retailer's problem is not concave. For fixed effort, R(⋅,e∣T)R(\cdot,e\mid T)R(⋅,e∣T) has a kink at Q=TQ=TQ=T, and the optimal order jumps from eQ‾0e\underline Q_0eQ​0​ to eQ‾1e\underline Q_1eQ​1​ at e=T/τe=T/\taue=T/τ. After the order is optimized out, the profit in effort, A(⋅∣T)A(\cdot\mid T)A(⋅∣T), has an upward kink at T/τT/\tauT/τ and can be convex and then concave beyond it. Checking the first-order condition at eˉ\bar eeˉ is therefore not enough. Coordination is a statement about the global maximum of a kinked, non-concave function, and comparing the two local candidates is what the threshold Υ\UpsilonΥ and the profits L‾\underline LL​, Lˉ\bar LLˉ do. The contract is then pinned down by two equations in (w,T)(w,T)(w,T) whose coefficients themselves depend on www through u(T)u(T)u(T) and b(T)b(T)b(T).

Formalization scope

Everything is stated for ξ∼Uniform(0,1)\xi\sim\mathrm{Uniform}(0,1)ξ∼Uniform(0,1), taken as Lebesgue measure on [0,1][0,1][0,1], and for V(e)=ae2/2V(e)=ae^2/2V(e)=ae2/2, as in all of Lemmas 3–6 and Theorem 2. Lemma 2 and the §4.1–4.2 solutions, which the paper states for general demand, are specialised to this instance. The uniform law has bounded support, so the paper's Assumption A4 (positive density on [0,∞)[0,\infty)[0,∞)) is not imposed; the paper allows that relaxation on p. 995.

The expectations use the published expSales and expLeftover (Cachon's newsvendor definitions) of the image law of ξ\xiξ under x↦exx\mapsto exx↦ex. Φ\PhiΦ and Γ\GammaΓ are integrals of the uniform density, not hard-coded closed forms. Every inverse (Qˉ0\bar Q_0Qˉ​0​, Q‾0\underline Q_0Q​0​, Q‾1\underline Q_1Q​1​), the threshold τ\tauτ, the function jjj and its root Υ\UpsilonΥ, the fixed points T1T_1T1​, T2T_2T2​ and the profits L‾\underline LL​, Lˉ\bar LLˉ are stated by their defining equations, never by a choice function. "Optimal" means a maximizer over Q≥0Q\ge0Q≥0, e≥0e\ge0e≥0. Channel coordination means that the retailer's set of maximizers is exactly {(Qˉ,eˉ)}\{(\bar Q,\bar e)\}{(Qˉ​,eˉ)}. The manufacturer's profit MMM, which the paper does not display, is written from the contract's cash flows: units bought back at bbb are salvaged at sss.

A trivializing formalization is ruled out: coordination is global optimality of (Qˉ,eˉ)(\bar Q,\bar e)(Qˉ​,eˉ), not a first-order condition at eˉ\bar eeˉ. ε0\varepsilon_0ε0​ may depend on κ\kappaκ but not on ε\varepsilonε. T1T_1T1​ and T2T_2T2​ are tied to their fixed-point equations and are not free variables.

A complete development needs:

  • the closed forms of Φ\PhiΦ, Γ\GammaΓ and the three expectations under the uniform law;
  • scaling identities in eee;
  • maximization of piecewise concave functions with a kink;
  • one-dimensional intermediate value arguments, with the monotonicity needed to repair the proof of Theorem 2.

The uniform newsvendor identities and the kinked-maximization lemmas are reusable beyond this mission. Contributions to any milestone, and alternative proofs of the existence part of Theorem 2, are welcome.

Selected references

  • T. A. Taylor, Supply Chain Coordination Under Channel Rebates with Sales Effort Effects, Management Science 48(8):992–1007, 2002. https://doi.org/10.1287/mnsc.48.8.992.168
  • G. P. Cachon, Supply Chain Coordination with Contracts, in Handbooks in Operations Research and Management Science, Vol. 11, 2003. https://doi.org/10.1016/S0927-0507(03)11006-7
  • B. A. Pasternack, Optimal Pricing and Return Policies for Perishable Commodities, Marketing Science 4(2):166–176, 1985. https://doi.org/10.1287/mksc.4.2.166
11 thms1 active userReviewed
Control TheoryProbabilityStochastic Systems·Captain: mikedeng1

On Minimizing the Ruin Probability by Investment and Reinsurance I: An Increasing Solution f of the HJB Equation Is Bounded, δ(u) = f(u)/f(∞), and the Feedback Strategy A*, b* Is OptimalResearch Paper

Motivation

An insurance company that collects premiums and pays random claims is ruined when its surplus becomes negative. The probability of ultimate ruin is the classical solvency criterion of risk theory, going back to Lundberg and Cramér. Real insurers have two levers on this probability: they can invest part of the surplus in a risky asset, and they can cede part of every claim to a reinsurer, at a price. The question of how to use both levers dynamically, as a function of the current surplus, so as to make the probability of ruin as small as possible, is a stochastic control problem for a jump-diffusion.

H. Schmidli, On minimizing the ruin probability by investment and reinsurance, Ann. Appl. Probab. 12 (2002), solves this problem for the Cramér–Lundberg model with proportional reinsurance and investment in a Black–Scholes asset. The main results are a verification theorem (Theorem 1, the subject of this mission) and an existence theorem for the associated integral equation (Theorem 2, the companion mission).

Timeline. Hipp and Plum (2000) treated investment in the Cramér–Lundberg model through the Hamilton–Jacobi–Bellman equation, assuming a bounded solution. Schmidli (2001) treated optimal proportional reinsurance alone. The 2002 paper combines investment and reinsurance, and proves the verification theorem without assuming the solution to be bounded.

Setting

Claims arrive at the jumps T1<T2<…T_1<T_2<\dotsT1​<T2​<… of a Poisson process NNN with rate λ>0\lambda>0λ>0; the claim sizes Y1,Y2,…Y_1,Y_2,\dotsY1​,Y2​,… are i.i.d. with continuous distribution function GGG, G(0)=0G(0)=0G(0)=0, independent of NNN. The insurer receives premiums at rate c>0c>0c>0. A risky asset Zt=exp⁡{σWt+(μ−12σ2)t}Z_t=\exp\{\sigma W_t+(\mu-\frac12\sigma^2)t\}Zt​=exp{σWt​+(μ−21​σ2)t}, with μ,σ>0\mu,\sigma>0μ,σ>0 and WWW a standard Brownian motion independent of the claims, is available.

A strategy is a pair of processes (At,bt)(A_t,b_t)(At​,bt​), predictable for the smallest right-continuous filtration generated by the claims and WWW: At∈RA_t\in\mathbb RAt​∈R is the amount invested in the risky asset (locally bounded), and bt∈[0,1]b_t\in[0,1]bt​∈[0,1] is the retention level, meaning the insurer pays btYb_tYbt​Y of a claim YYY occurring at time ttt, at a reinsurance premium rate c(bt)c(b_t)c(bt​). The surplus X=XAbX=X^{Ab}X=XAb starting from u≥0u\ge0u≥0 solves

dXt=(c−c(bt)+μAt) dt+σAt dWt−bt dSt,X0=u,dX_t=\big(c-c(b_t)+\mu A_t\big)\,dt+\sigma A_t\,dW_t-b_t\,dS_t,\qquad X_0=u,dXt​=(c−c(bt​)+μAt​)dt+σAt​dWt​−bt​dSt​,X0​=u,

where St=∑i≤NtYiS_t=\sum_{i\le N_t}Y_iSt​=∑i≤Nt​​Yi​ is the aggregate claims process. The ruin time is τ=inf⁡{t≥0:Xt<0}\tau=\inf\{t\ge0:X_t<0\}τ=inf{t≥0:Xt​<0}, the survival probability is δAb(u)=P[τ=∞]\delta^{Ab}(u)=\mathbb P[\tau=\infty]δAb(u)=P[τ=∞], and the value function is δ(u)=sup⁡A,bδAb(u)\delta(u)=\sup_{A,b}\delta^{Ab}(u)δ(u)=supA,b​δAb(u).

The premium function c(b)c(b)c(b) is decreasing and continuous on [0,1][0,1][0,1], c(1)=0c(1)=0c(1)=0, lim inf⁡b↑1c(b)/(1−b)>0\liminf_{b\uparrow1}c(b)/(1-b)>0liminfb↑1​c(b)/(1−b)>0, and there is b‾>0\underline b>0b​>0 with c(b)>cc(b)>cc(b)>c for b<b‾b<\underline bb<b​ and c(b)≤cc(b)\le cc(b)≤c for b≥b‾b\ge\underline bb≥b​ (full reinsurance is too expensive).

The Hamilton–Jacobi–Bellman equation for δ\deltaδ is, with the convention f(u)=0f(u)=0f(u)=0 for u<0u<0u<0,

sup⁡b∈[0,1]sup⁡A≥0[12σ2A2f′′(u)+(c−c(b)+μA)f′(u)+λ(E[f(u−bY)]−f(u))]=0.(1)\sup_{b\in[0,1]}\sup_{A\ge0}\Big[\tfrac12\sigma^2A^2f''(u)+\big(c-c(b)+\mu A\big)f'(u)+\lambda\big(\mathbb E[f(u-bY)]-f(u)\big)\Big]=0.\tag{1}b∈[0,1]sup​A≥0sup​[21​σ2A2f′′(u)+(c−c(b)+μA)f′(u)+λ(E[f(u−bY)]−f(u))]=0.(1)

For strictly concave fff the maximum over AAA is attained at A∗(u)=−μf′(u)/(σ2f′′(u))A^*(u)=-\mu f'(u)/(\sigma^2f''(u))A∗(u)=−μf′(u)/(σ2f′′(u)) (equation (2)); b∗(u)b^*(u)b∗(u) denotes a measurable maximiser in bbb.

Formalization targets

Goal: Theorem 1 (p. 896)

Let fff be strictly increasing and nonnegative on [0,∞)[0,\infty)[0,∞), continuous there, twice continuously differentiable on (0,∞)(0,\infty)(0,∞), zero on (−∞,0)(-\infty,0)(−∞,0), and a solution of (1) at every u>0u>0u>0. Then fff is bounded, f(∞)=lim⁡x→∞f(x)∈(0,∞)f(\infty)=\lim_{x\to\infty}f(x)\in(0,\infty)f(∞)=limx→∞​f(x)∈(0,∞),

δ(u)=f(u)f(∞)(u≥0),\delta(u)=\frac{f(u)}{f(\infty)}\qquad(u\ge0),δ(u)=f(∞)f(u)​(u≥0),

and every surplus process from u>0u>0u>0 following At=A∗(Xt−)A_t=A^*(X_{t-})At​=A∗(Xt−​), bt=b∗(Xt−)b_t=b^*(X_{t-})bt​=b∗(Xt−​) has survival probability δ(u)\delta(u)δ(u).

Milestones

The proof of Theorem 1 runs through: Lemma 3 (for small surplus the optimal retention is b∗=1b^*=1b∗=1); the identity E[f(Xτ∗∧t∗)]=f(u)\mathbb E[f(X^*_{\tau^*\wedge t})]=f(u)E[f(Xτ∗∧t∗​)]=f(u) along the feedback strategy; the inequality E[f(Xt∧τ)]≤f(u)\mathbb E[f(X_{t\wedge\tau})]\le f(u)E[f(Xt∧τ​)]≤f(u) for every strategy; Lemma 1 (almost surely, either ruin occurs or Xt→∞X_t\to\inftyXt​→∞); the existence of a strategy with P[τ=∞]>0\mathbb P[\tau=\infty]>0P[τ=∞]>0; the comparison δAb(u)≤f(u)/f(∞)\delta^{Ab}(u)\le f(u)/f(\infty)δAb(u)≤f(u)/f(∞) with equality for the feedback strategy; and the absence of ruin by creeping, P[τ∗<∞,Xτ∗∗=0]=0\mathbb P[\tau^*<\infty,X^*_{\tau^*}=0]=0P[τ∗<∞,Xτ∗∗​=0]=0.

Companion statements

The second sentence of Theorem 1 (that f(x)=1+∫0xg(z) dzf(x)=1+\int_0^xg(z)\,dzf(x)=1+∫0x​g(z)dz satisfies the hypotheses whenever ggg is a decreasing solution of the integral equation (5)) and its last sentence (at most one such solution of (1) with f(0)=1f(0)=1f(0)=1) are separate items.

Significance

Theorem 1 reduces a control problem over all predictable strategies to an analytic question: find one increasing smooth solution of (1). Combined with Theorem 2 (existence of a solution of (5)), it shows that the optimal survival probability is f(u)/f(∞)f(u)/f(\infty)f(u)/f(∞) and that the optimal strategy is a Markov feedback rule in the current surplus. It also yields that this solution is bounded, which Hipp and Plum (2000) had to assume, and closes the gap they left open on ruin by creeping (Remark (iii), p. 897). Numerical procedures for the optimal strategy (§5 of the paper) rest on this identification.

The theorem is proved in the paper; nothing in it is formalized. The mission produces a machine-checked verification theorem for a controlled jump-diffusion with an Itô integral against Brownian motion, a Poisson claim stream and a predictable control, together with a formal model of the Cramér–Lundberg process with investment and reinsurance that later results in risk theory can reuse.

Difficulty

The standard verification argument applies Itô's formula to f(Xt)f(X_t)f(Xt​) and uses (1) to show that f(Xt∧τ)f(X_{t\wedge\tau})f(Xt∧τ​) is a supermartingale, and a martingale under the feedback strategy. Here fff is not C2C^2C2 at 000 (f′′(0+)=−∞f''(0+)=-\inftyf′′(0+)=−∞), is discontinuous at 000 after the extension by 000, and is not assumed bounded, so the stochastic integral is only a local martingale and the passage t→∞t\to\inftyt→∞ needs both Lemma 1 and a separate argument that the feedback process does not creep through 000. Lemma 1 itself is only sketched in the paper ("We will just describe the argument"). The optimality part needs a surplus process following the feedback rule; its existence is not proved in the paper.

Formalization scope

The model is formalized on a measurable space with a probability measure: i.i.d. exponential interarrival times (realising the Poisson process), i.i.d. claim sizes with law ν\nuν, and a real Brownian motion (IsBrownianReal), the three families mutually independent. Independence of WWW from the claims is not printed in the paper but is used by (1) and its proof; it is a disclosed standing assumption. The filtration is Ft=⋂s>tσ(Sr,Wr:r≤s)\mathcal F_t=\bigcap_{s>t}\sigma(S_r,W_r:r\le s)Ft​=⋂s>t​σ(Sr​,Wr​:r≤s), not completed. The Itô integral is the published relation EthierKurtz.HasBrownianItoIntegral, and its paths are required to be continuous, so that the pathwise ruin event is well defined. The surplus equation holds for every ttt and every sample point.

Conventions: the premium c(b)c(b)c(b) is real valued (so c(0)<∞c(0)<\inftyc(0)<∞, which the paper allows to fail when E[Y]<∞\mathbb E[Y]<\inftyE[Y]<∞); the supremum in (1) is stated as a least upper bound equal to 000, so an unbounded family never counts as a solution; fff is C2C^2C2 on (0,∞)(0,\infty)(0,∞) and (1) is required for u>0u>0u>0; δ\deltaδ is a supremum over all admissible triples (strategy and surplus process); the controls of the feedback rule at non-positive pre-jump surplus are left free, and optimality is stated for u>0u>0u>0. Expectations of f(Xt∧τ)f(X_{t\wedge\tau})f(Xt∧τ​) are lower Lebesgue integrals of nonnegative functions, so no default value of a non-integrable expectation satisfies an inequality. The optimality clause is stated for every process following the feedback rule; it would be vacuous if no such process existed, and its existence is not part of the target.

Infrastructure needed: Itô's formula for a semimartingale with Brownian and compound-Poisson parts, the compensator of the compound Poisson random measure, localisation, and optional stopping. The model definitions, the HJB predicate and the comparison lemmas are reusable for other ruin-minimisation and dividend problems. Contributions to any milestone, and to general stochastic-calculus lemmas they need, are welcome.

Selected references

  • H. Schmidli, On minimizing the ruin probability by investment and reinsurance, Ann. Appl. Probab. 12(3) (2002), 890–907. https://doi.org/10.1214/aoap/1031863173
  • C. Hipp, M. Plum, Optimal investment for insurers, Insurance Math. Econom. 27 (2000), 215–228. https://doi.org/10.1016/S0167-6687(00)00049-4
  • H. Schmidli, Optimal proportional reinsurance policies in a dynamic setting, Scand. Actuar. J. (2001), 55–68. https://doi.org/10.1080/034612301750077338
  • P. Brémaud, Point Processes and Queues: Martingale Dynamics, Springer (1981). https://doi.org/10.1007/978-1-4684-9477-8
11 thms1 active userReviewed
OptimizationProbability·Captain: mikedeng1

Supply Chain Coordination Under Channel Rebates with Sales Effort Effects I: A Target Rebate Coordinates the Newsvendor Channel and Splits Its Profit in Any ProportionResearch Paper

Motivation

Manufacturers in computer hardware, software and automobiles routinely pay channel rebates to their retailers: a payment per unit the retailer sells to end consumers, as opposed to per unit the retailer buys. A target rebate pays only for units sold beyond a target level, a linear rebate pays for every unit sold. T. A. Taylor, Supply Chain Coordination Under Channel Rebates with Sales Effort Effects, Management Science 48(8) (2002), doi:10.1287/mnsc.48.8.992.168, asks whether such payments can make an independent retailer act in the interest of the whole supply chain.

The question belongs to the supply-chain contracting literature built on the newsvendor model. A wholesale price above cost makes the retailer under-order relative to the integrated channel (double marginalization, Spengler 1950). Returns contracts (Pasternack 1985) and revenue sharing (Pasternack 1999; Cachon and Lariviere 2000, as cited by Taylor) are known remedies. Cachon's survey, Supply Chain Coordination with Contracts (Handbooks in OR & MS 11, 2003), discusses sales-rebate contracts through a first-order condition. Taylor's Theorem 1 is the global statement for the target rebate, with an explicit contract and an arbitrary profit split, and without any lump-sum side payment.

This mission covers §3 of the paper, the model without sales effort. Missions II–IV of the series cover §4, where the retailer's effort also affects demand.

Setting

A manufacturer sells to a retailer who places one order of size Q≥0Q \ge 0Q≥0 before observing demand ξ\xiξ. Demand has a density φ\varphiφ with φ(x)=0\varphi(x) = 0φ(x)=0 for x<0x < 0x<0 and φ(x)>0\varphi(x) > 0φ(x)>0 for every x≥0x \ge 0x≥0 (Assumption A4), and a finite mean. Write

Φ(Q)=∫0Qφ(x) dx,Γ(Q)=∫0Qx φ(x) dx.\Phi(Q) = \int_0^Q \varphi(x)\,dx, \qquad \Gamma(Q) = \int_0^Q x\,\varphi(x)\,dx .Φ(Q)=∫0Q​φ(x)dx,Γ(Q)=∫0Q​xφ(x)dx.

The retail price ppp, the production cost ccc and the salvage value sss are fixed with 0<c<p0 < c < p0<c<p and s<cs < cs<c; sss may be negative. A contract specifies a wholesale price www, a rebate u>0u > 0u>0 and a target T≥0T \ge 0T≥0 (Assumption A1: 0<c<w<p0 < c < w < p0<c<w<p, s<cs < cs<c, u>0u > 0u>0, T≥0T \ge 0T≥0). No other transfer is allowed (Assumption A3).

  • The integrated channel earns π(Q)=−cQ+pEmin⁡(Q,ξ)+sE(Q−ξ)+\pi(Q) = -cQ + pE\min(Q,\xi) + sE(Q-\xi)^+π(Q)=−cQ+pEmin(Q,ξ)+sE(Q−ξ)+. Its optimal order Qˉ0\bar Q_0Qˉ​0​ solves Φ(Qˉ0)=(p−c)/(p−s)\Phi(\bar Q_0) = (p-c)/(p-s)Φ(Qˉ​0​)=(p−c)/(p−s), and its optimal profit is π=(p−s)Γ(Qˉ0)\pi = (p-s)\Gamma(\bar Q_0)π=(p−s)Γ(Qˉ​0​).
  • Under the target rebate (w,u,T)(w,u,T)(w,u,T) the retailer earns
r(Q∣T)=−wQ+pEmin⁡(Q,ξ)+sE(Q−ξ)++uE(min⁡(Q,ξ)−T)+,r(Q\mid T) = -wQ + pE\min(Q,\xi) + sE(Q-\xi)^+ + uE(\min(Q,\xi)-T)^+,r(Q∣T)=−wQ+pEmin(Q,ξ)+sE(Q−ξ)++uE(min(Q,ξ)−T)+,

and the manufacturer earns m(Q∣T)=(w−c)Q−uE(min⁡(Q,ξ)−T)+m(Q\mid T) = (w-c)Q - uE(\min(Q,\xi)-T)^+m(Q∣T)=(w−c)Q−uE(min(Q,ξ)−T)+.

  • Under the wholesale price-only contract (u=0u = 0u=0) the retailer orders Q0Q_0Q0​ with Φ(Q0)=(p−w)/(p−s)\Phi(Q_0) = (p-w)/(p-s)Φ(Q0​)=(p−w)/(p−s) and earns r‾=(p−s)Γ(Q0)\underline r = (p-s)\Gamma(Q_0)r​=(p−s)Γ(Q0​).

The contract coordinates the channel when Qˉ0\bar Q_0Qˉ​0​ is the retailer's unique optimal order. Finally u^(w)=(w−c)(p−s)/(c−s)\hat u(w) = (w-c)(p-s)/(c-s)u^(w)=(w−c)(p−s)/(c−s).

Formalization targets

Goal: Theorem 1 (p. 996)

Fix κ∈(0,π)\kappa \in (0,\pi)κ∈(0,π). For every sufficiently small ε∈(0,κ)\varepsilon \in (0,\kappa)ε∈(0,κ) there is exactly one triple (w∗,u∗,T∗)(w^*,u^*,T^*)(w∗,u∗,T∗) with r‾(w∗)=κ−ε\underline r(w^*) = \kappa-\varepsilonr​(w∗)=κ−ε, u∗=u^(w∗)u^* = \hat u(w^*)u∗=u^(w∗), T∗≥0T^* \ge 0T∗≥0 and

(p+u∗−s)Γ(Qˉ0)−u∗(Γ(T∗)+T∗[1−Φ(T∗)])=κ,(1)(p+u^*-s)\Gamma(\bar Q_0) - u^*\big(\Gamma(T^*) + T^*[1-\Phi(T^*)]\big) = \kappa, \qquad (1)(p+u∗−s)Γ(Qˉ​0​)−u∗(Γ(T∗)+T∗[1−Φ(T∗)])=κ,(1)

and this triple has w∗∈(c,p)w^* \in (c,p)w∗∈(c,p), u∗>0u^*>0u∗>0, T∗>0T^*>0T∗>0, coordinates the channel, and gives the retailer r∗=κr^* = \kappar∗=κ and the manufacturer m∗=π−κm^* = \pi-\kappam∗=π−κ.

Milestones

  1. §3.1, p. 995. The newsvendor solutions: Qˉ0\bar Q_0Qˉ​0​ and Q0Q_0Q0​ are the unique optimal orders, with values (p−s)Γ(Qˉ0)(p-s)\Gamma(\bar Q_0)(p−s)Γ(Qˉ​0​) and (p−s)Γ(Q0)(p-s)\Gamma(Q_0)(p−s)Γ(Q0​), and Q0<Qˉ0Q_0 < \bar Q_0Q0​<Qˉ​0​.
  2. §3.2, p. 995. The piecewise form of r(⋅∣T)r(\cdot\mid T)r(⋅∣T), its derivative p−w−(p−s)Φ(Q)p-w-(p-s)\Phi(Q)p−w−(p−s)Φ(Q) below TTT and p+u−w−(p+u−s)Φ(Q)p+u-w-(p+u-s)\Phi(Q)p+u−w−(p+u−s)Φ(Q) above, strict concavity on [0,T)[0,T)[0,T) and (T,∞)(T,\infty)(T,∞), and an upward kink of the derivative at TTT.
  3. Lemma 1, p. 995. With Φ(Q1)=(p+u−w)/(p+u−s)\Phi(Q_1) = (p+u-w)/(p+u-s)Φ(Q1​)=(p+u−w)/(p+u−s), the equation r(Q0∣τ0)=r(Q1∣τ0)r(Q_0\mid\tau_0) = r(Q_1\mid\tau_0)r(Q0​∣τ0​)=r(Q1​∣τ0​) has exactly one root τ0∈[Q0,Q1]\tau_0 \in [Q_0,Q_1]τ0​∈[Q0​,Q1​], lying in (Q0,Q1)(Q_0,Q_1)(Q0​,Q1​); the retailer's optimal order set is {Q1}\{Q_1\}{Q1​} for T<τ0T<\tau_0T<τ0​, {Q0}\{Q_0\}{Q0​} for T>τ0T>\tau_0T>τ0​, and {Q0,Q1}\{Q_0,Q_1\}{Q0​,Q1​} for T=τ0T=\tau_0T=τ0​.
  4. §3.2, p. 996. The retailer's optimal profit is (p+u−s)Γ(Q1)−u(Γ(T)+T[1−Φ(T)])(p+u-s)\Gamma(Q_1) - u(\Gamma(T)+T[1-\Phi(T)])(p+u−s)Γ(Q1​)−u(Γ(T)+T[1−Φ(T)]) if T<τ0T<\tau_0T<τ0​ and (p−s)Γ(Q0)(p-s)\Gamma(Q_0)(p−s)Γ(Q0​) if T≥τ0T \ge \tau_0T≥τ0​.
  5. Proposition 1, p. 996. If T=0T = 0T=0, channel coordination requires m∗<0m^* < 0m∗<0.

Significance

Theorem 1 shows that a single sales-based instrument, a target rebate, achieves what a wholesale price alone cannot: the retailer orders the channel-optimal quantity, and the channel profit π\piπ is split in any proportion κ:π−κ\kappa : \pi-\kappaκ:π−κ. The parameter κ\kappaκ can be set to the retailer's opportunity cost, so a manufacturer with all the bargaining power extracts the rest of the channel profit without a side payment. Proposition 1 explains why the target is needed: a linear rebate that coordinates the channel leaves the manufacturer with a loss. The quantity-only analysis is also the base case of §4, where the paper shows that a target rebate combined with returns coordinates effort and quantity as well.

The results are proved in the paper (Appendix, p. 1005). No machine-checked proof of any of them is known to exist. The mission produces a Lean statement of the model and of each result, and invites formal proofs of the newsvendor solution, the shape of the kinked profit, the threshold lemma and the coordination theorem.

Difficulty

The retailer's profit r(⋅∣T)r(\cdot\mid T)r(⋅∣T) is not concave: its derivative jumps up at the target. The usual newsvendor argument, a stationary point of a concave function is a global maximum, therefore fails. A first-order condition at Qˉ0\bar Q_0Qˉ​0​ shows only that Qˉ0\bar Q_0Qˉ​0​ is a local candidate on (T,∞)(T,\infty)(T,∞). The retailer may still prefer the smaller candidate Q0Q_0Q0​ on [0,T][0,T][0,T], and coordination fails exactly when she does. Theorem 1(b) requires comparing the two branches globally, through the threshold τ0\tau_0τ0​ of Lemma 1, and then showing that the contract's target lies strictly below τ0\tau_0τ0​.

The contract itself is defined implicitly. w∗w^*w∗ solves an equation involving the critical fractile Q0(w)Q_0(w)Q0​(w), and T∗T^*T∗ solves (1). Existence and uniqueness of both, and the placement of T∗T^*T∗ relative to τ0\tau_0τ0​, are part of the claim.

Formalization scope

All definitions are in ChannelRebate.Quantity.Setting. Demand is a structure Demand carrying the density φ\varphiφ with the properties listed under Setting; its law is Lebesgue measure with density φ\varphiφ. Emin⁡(Q,ξ)E\min(Q,\xi)Emin(Q,ξ) and E(Q−ξ)+E(Q-\xi)^+E(Q−ξ)+ are expSales and expLeftover of the published definition CachonCoord_Newsvendor_Contracts, applied to this law. E(min⁡(Q,ξ)−T)+E(\min(Q,\xi)-T)^+E(min(Q,ξ)−T)+ is the local rebateUnits. Φ\PhiΦ and Γ\GammaΓ are interval integrals from 000.

Conventions:

  • Φ−1\Phi^{-1}Φ−1 is never a function. Each critical fractile (Qˉ0\bar Q_0Qˉ​0​, Q0Q_0Q0​, Q1Q_1Q1​) is a hypothesis giving its defining equation, and τ0\tau_0τ0​ is any solution of f0(τ0)=0f_0(\tau_0) = 0f0​(τ0​)=0 in [Q0,Q1][Q_0,Q_1][Q0​,Q1​].
  • An optimal order maximizes over Q≥0Q \ge 0Q≥0. "Coordination" and the order sets of Lemma 1 are equalities of the set of maximizers, so a tie with Q0Q_0Q0​ does not count as coordination.
  • "For ε\varepsilonε sufficiently small" is "there exists ε0>0\varepsilon_0 > 0ε0​>0 such that for all ε∈(0,min⁡(ε0,κ))\varepsilon \in (0,\min(\varepsilon_0,\kappa))ε∈(0,min(ε0​,κ))".
  • In Theorem 1 the uniqueness ranges over all real triples satisfying the specification, with www unrestricted; w∗∈(c,p)w^* \in (c,p)w∗∈(c,p) is a conclusion.
  • The manufacturer's profit is its own formula, never channel profit minus retailer profit; m∗=π−κm^* = \pi-\kappam∗=π−κ is a conclusion.
  • The standing assumptions A1, A3 and A4 appear as hypotheses or in Demand; no continuity of φ\varphiφ is assumed.

A formalization that reads coordination as the first-order condition at Qˉ0\bar Q_0Qˉ​0​, or that defines T∗T^*T∗ or w∗w^*w∗ by a choice function, would make Theorem 1 vacuous or weaker, and is ruled out by the statements above.

A complete development needs the newsvendor calculus for a density (derivatives of Emin⁡(Q,ξ)E\min(Q,\xi)Emin(Q,ξ) and E(Q−ξ)+E(Q-\xi)^+E(Q−ξ)+ in QQQ, strict monotonicity of Φ\PhiΦ on [0,∞)[0,\infty)[0,∞)), the intermediate value theorem for the two implicit equations, and the global comparison of the two concave branches. The newsvendor calculus is reusable for missions II–IV of this series and for other newsvendor contracts. Proofs of the milestones in any order are welcome.

Selected references

  • T. A. Taylor, Supply Chain Coordination Under Channel Rebates with Sales Effort Effects, Management Science 48(8):992–1007, 2002. doi:10.1287/mnsc.48.8.992.168
  • G. P. Cachon, Supply Chain Coordination with Contracts, in Handbooks in Operations Research and Management Science 11, 2003. doi:10.1016/S0927-0507(03)11006-7
  • B. A. Pasternack, Optimal Pricing and Return Policies for Perishable Commodities, Marketing Science 4(2):166–176, 1985. doi:10.1287/mksc.4.2.166
  • J. J. Spengler, Vertical Integration and Antitrust Policy, Journal of Political Economy 58(4):347–352, 1950. doi:10.1086/256964
8 thms1 active userReviewed
OptimizationProbability·Captain: mikedeng1

Quantitative Stability in Stochastic Programming: The Method of Probability Metrics 3: Mixed-Integer Two-Stage Programs Are Hölder Stable in the Polyhedral DiscrepancyResearch Paper

Motivation

Two-stage stochastic programs are solved with an approximation of the true distribution of the random data: an empirical measure, a scenario tree, or a discretization. Whether the computed decisions mean anything depends on how the optimal value and the solution set react when the underlying probability measure is replaced by a nearby one, and on which distance between measures "nearby" refers to. Rachev and Römisch (preprint, edoc.hu-berlin.de; published in Math. Oper. Res. 27 (2002) 792–818, doi:10.1287/moor.27.4.792.304) organize this question in two steps. First, a minimal information distance dFUd_{\mathcal F_{\mathcal U}}dFU​​, built from the integrands of the problem itself, controls optimal values and solution sets (their Theorems 2.2 and 2.3). Second, for a given class of models, dFUd_{\mathcal F_{\mathcal U}}dFU​​ is bounded by a canonical probability metric that depends only on the analytical properties of the integrands.

For linear two-stage programs the integrands are locally Lipschitz in the random parameter and the canonical metric is a Fortet–Mourier metric. When the second stage has integer variables, the recourse function is discontinuous, and metrics of Fortet–Mourier type no longer control the problem. This mission formalizes the paper's answer for that case (§3.2): mixed-integer two-stage programs are Hölder stable with respect to a polyhedral discrepancy d1,phkd_{1,phk}d1,phk​ built from Lipschitz functions restricted to polyhedra. The structure of mixed-integer value functions that makes this possible goes back to Blair and Jeroslow (1977), Bank et al. (1982) and Schultz (1996), all cited in the paper.

Setting

Let c∈Rmc\in\mathbb R^mc∈Rm, let X⊆RmX\subseteq\mathbb R^mX⊆Rm be closed and Ξ⊆Rs\Xi\subseteq\mathbb R^sΞ⊆Rs a polyhedron, i.e. an intersection of finitely many closed half-spaces. The second stage has costs q∈Rm^q\in\mathbb R^{\hat m}q∈Rm^, qˉ∈Rmˉ\bar q\in\mathbb R^{\bar m}qˉ​∈Rmˉ and (r,m^)(r,\hat m)(r,m^)- and (r,mˉ)(r,\bar m)(r,mˉ)-matrices WWW, Wˉ\bar WWˉ. Its value function is

Φ(t)=min⁡{qy+qˉyˉ: Wy+Wˉyˉ=t, y∈Z+m^, yˉ∈R+mˉ}(t∈Rr),\Phi(t)=\min\{qy+\bar q\bar y:\ Wy+\bar W\bar y=t,\ y\in\mathbb Z^{\hat m}_+,\ \bar y\in\mathbb R^{\bar m}_+\}\qquad(t\in\mathbb R^r),Φ(t)=min{qy+qˉ​yˉ​: Wy+Wˉyˉ​=t, y∈Z+m^​, yˉ​∈R+mˉ​}(t∈Rr),

and T\mathcal TT is the set of right-hand sides ttt for which the constraint set is nonempty. The right-hand side h(ξ)∈Rrh(\xi)\in\mathbb R^rh(ξ)∈Rr and the technology matrix T(ξ)T(\xi)T(ξ) depend affinely on ξ∈Rs\xi\in\mathbb R^sξ∈Rs. For a Borel probability measure μ\muμ on Ξ\XiΞ the program is

min⁡{∫Ξf0(ξ,x) μ(dξ): x∈X},f0(ξ,x)=cx+Φ(h(ξ)−T(ξ)x).(9)\min\Big\{\int_\Xi f_0(\xi,x)\,\mu(d\xi):\ x\in X\Big\},\qquad f_0(\xi,x)=cx+\Phi(h(\xi)-T(\xi)x). \tag{9}min{∫Ξ​f0​(ξ,x)μ(dξ): x∈X},f0​(ξ,x)=cx+Φ(h(ξ)−T(ξ)x).(9)

Three conditions make (9) well defined: (B1) WWW and Wˉ\bar WWˉ have rational entries; (B2) h(ξ)−T(ξ)x∈Th(\xi)-T(\xi)x\in\mathcal Th(ξ)−T(ξ)x∈T for all (ξ,x)∈Ξ×X(\xi,x)\in\Xi\times X(ξ,x)∈Ξ×X (relatively complete recourse); (B3) there is uuu with W′u≤qW'u\le qW′u≤q and Wˉ′u≤qˉ\bar W'u\le\bar qWˉ′u≤qˉ​ (dual feasibility).

v(ν)v(\nu)v(ν) and S(ν)S(\nu)S(ν) denote the optimal value and solution set of (9) under ν\nuν; for an open bounded U⊆Rm\mathcal U\subseteq\mathbb R^mU⊆Rm, vU(ν)v_{\mathcal U}(\nu)vU​(ν) and SU(ν)S_{\mathcal U}(\nu)SU​(ν) are the same objects with XXX replaced by X∩cl⁡UX\cap\operatorname{cl}\mathcal UX∩clU. SU(ν)S_{\mathcal U}(\nu)SU​(ν) is a complete local minimizing (CLM) set if it is nonempty and contained in U\mathcal UU. The minimal information distance is dFU(μ,ν)=sup⁡x∈X∩cl⁡U∣∫Ξf0(ξ,x)(μ−ν)(dξ)∣d_{\mathcal F_{\mathcal U}}(\mu,\nu)=\sup_{x\in X\cap\operatorname{cl}\mathcal U}|\int_\Xi f_0(\xi,x)(\mu-\nu)(d\xi)|dFU​​(μ,ν)=supx∈X∩clU​∣∫Ξ​f0​(ξ,x)(μ−ν)(dξ)∣. The moment classes are Pp,K(Ξ)={ν:∫Ξ∥ξ∥pν(dξ)≤K}\mathcal P_{p,K}(\Xi)=\{\nu:\int_\Xi\|\xi\|^p\nu(d\xi)\le K\}Pp,K​(Ξ)={ν:∫Ξ​∥ξ∥pν(dξ)≤K}, and the polyhedral metric is

d1,phk(μ,ν)=sup⁡{∣∫Pf(ξ)(μ−ν)(dξ)∣: P a polyhedron with at most k faces, f 1-Lipschitz on P, ∣f(ξ)∣≤max⁡{1,∥ξ∥}}.d_{1,phk}(\mu,\nu)=\sup\Big\{\Big|\int_P f(\xi)(\mu-\nu)(d\xi)\Big|:\ P \text{ a polyhedron with at most } k \text{ faces},\ f \text{ 1-Lipschitz on } P,\ |f(\xi)|\le\max\{1,\|\xi\|\}\Big\}.d1,phk​(μ,ν)=sup{​∫P​f(ξ)(μ−ν)(dξ)​: P a polyhedron with at most k faces, f 1-Lipschitz on P, ∣f(ξ)∣≤max{1,∥ξ∥}}.

Finally, ψ(τ)=inf⁡{∫Ξf0(ξ,x)μ(dξ)−v(μ):d(x,S(μ))≥τ, x∈X∩cl⁡U}\psi(\tau)=\inf\{\int_\Xi f_0(\xi,x)\mu(d\xi)-v(\mu): d(x,S(\mu))\ge\tau,\ x\in X\cap\operatorname{cl}\mathcal U\}ψ(τ)=inf{∫Ξ​f0​(ξ,x)μ(dξ)−v(μ):d(x,S(μ))≥τ, x∈X∩clU} is the growth function at μ\muμ, ψ−1(t)=sup⁡{τ≥0:ψ(τ)≤t}\psi^{-1}(t)=\sup\{\tau\ge0:\psi(\tau)\le t\}ψ−1(t)=sup{τ≥0:ψ(τ)≤t}, and Ψ(η)=η+ψ−1(2η)\Psi(\eta)=\eta+\psi^{-1}(2\eta)Ψ(η)=η+ψ−1(2η).

Formalization targets

Goal: Theorem 3.6

Under (B1)–(B3), with μ∈Pp,K(Ξ)\mu\in\mathcal P_{p,K}(\Xi)μ∈Pp,K​(Ξ) for some p>1p>1p>1, K>0K>0K>0, S(μ)≠∅S(\mu)\ne\emptysetS(μ)=∅ and U\mathcal UU an open bounded neighbourhood of S(μ)S(\mu)S(μ), there are L>0L>0L>0, δ>0\delta>0δ>0 and k∈Nk\in\mathbb Nk∈N such that for all ν∈Pp,K(Ξ)\nu\in\mathcal P_{p,K}(\Xi)ν∈Pp,K​(Ξ) with d1,phk(μ,ν)<δd_{1,phk}(\mu,\nu)<\deltad1,phk​(μ,ν)<δ

∣v(μ)−vU(ν)∣≤L d1,phk(μ,ν)11+rp−1,∅≠SU(ν)⊆S(μ)+Ψ(L d1,phk(μ,ν)11+rp−1)B,|v(\mu)-v_{\mathcal U}(\nu)|\le L\,d_{1,phk}(\mu,\nu)^{\frac{1}{1+\frac{r}{p-1}}},\qquad \emptyset\ne S_{\mathcal U}(\nu)\subseteq S(\mu)+\Psi\Big(L\,d_{1,phk}(\mu,\nu)^{\frac{1}{1+\frac{r}{p-1}}}\Big)\mathbb B,∣v(μ)−vU​(ν)∣≤Ld1,phk​(μ,ν)1+p−1r​1​,∅=SU​(ν)⊆S(μ)+Ψ(Ld1,phk​(μ,ν)1+p−1r​1​)B,

and SU(ν)S_{\mathcal U}(\nu)SU​(ν) is a CLM set. The constants and the number of faces are existential; only the exponent is fixed.

Milestones

  1. Lemma 3.5 (i)–(ii): a countable, locally finite Borel partition of T\mathcal TT into translates of pos⁡Wˉ\operatorname{pos}\bar WposWˉ minus finitely many translates, on each piece of which Φ\PhiΦ is Lipschitz with a common constant.
  2. Lemma 3.5, last assertion: Φ\PhiΦ is finite and lower semicontinuous on T\mathcal TT, and ∣Φ(t)−Φ(t~)∣≤α∥t−t~∥+β|\Phi(t)-\Phi(\tilde t)|\le\alpha\|t-\tilde t\|+\beta∣Φ(t)−Φ(t~)∣≤α∥t−t~∥+β.
  3. Growth bound (proof of Theorem 3.6): ∣f0(ξ,x)∣≤∥c∥∥x∥+α(∥h(ξ)∥+∥T(ξ)∥∥x∥)+β|f_0(\xi,x)|\le\|c\|\|x\|+\alpha(\|h(\xi)\|+\|T(\xi)\|\|x\|)+\beta∣f0​(ξ,x)∣≤∥c∥∥x∥+α(∥h(ξ)∥+∥T(ξ)∥∥x∥)+β, hence every measure with finite first moment lies in the domain of dFUd_{\mathcal F_{\mathcal U}}dFU​​.
  4. Theorem 2.2 for d=0d=0d=0: ∣v(μ)−vU(ν)∣≤dFU(μ,ν)|v(\mu)-v_{\mathcal U}(\nu)|\le d_{\mathcal F_{\mathcal U}}(\mu,\nu)∣v(μ)−vU​(ν)∣≤dFU​​(μ,ν) for all admissible ν\nuν, Berge upper semicontinuity of SUS_{\mathcal U}SU​, CLM sets near μ\muμ.
  5. Theorem 2.3 for d=0d=0d=0: ∅≠SU(ν)⊆S(μ)+Ψ(L^dFU(μ,ν))B\emptyset\ne S_{\mathcal U}(\nu)\subseteq S(\mu)+\Psi(\hat L d_{\mathcal F_{\mathcal U}}(\mu,\nu))\mathbb B∅=SU​(ν)⊆S(μ)+Ψ(L^dFU​​(μ,ν))B with Ψ(η)=η+ψ−1(η)\Psi(\eta)=\eta+\psi^{-1}(\eta)Ψ(η)=η+ψ−1(η).
  6. Estimate (12): dFU(μ,ν)≤C d1,phk(μ,ν)1/(1+r/(p−1))d_{\mathcal F_{\mathcal U}}(\mu,\nu)\le C\,d_{1,phk}(\mu,\nu)^{1/(1+r/(p-1))}dFU​​(μ,ν)≤Cd1,phk​(μ,ν)1/(1+r/(p−1)) for small d1,phk(μ,ν)d_{1,phk}(\mu,\nu)d1,phk​(μ,ν).

Companion: Corollary 3.7

For bounded Ξ\XiΞ and every μ∈P(Ξ)\mu\in\mathcal P(\Xi)μ∈P(Ξ) the same conclusions hold with the Lipschitz rate L αphk(μ,ν)L\,\alpha_{phk}(\mu,\nu)Lαphk​(μ,ν), where αphk(μ,ν)=sup⁡{∣μ(P)−ν(P)∣}\alpha_{phk}(\mu,\nu)=\sup\{|\mu(P)-\nu(P)|\}αphk​(μ,ν)=sup{∣μ(P)−ν(P)∣} over polyhedra with at most kkk faces.

Significance

Theorem 3.6 says that, for mixed-integer recourse, closeness of the distributions in d1,phkd_{1,phk}d1,phk​ is enough for closeness of optimal values and of localized solution sets, with an explicit Hölder exponent (p−1)/(p−1+r)(p-1)/(p-1+r)(p−1)/(p−1+r) that degrades with the second-stage dimension rrr and improves with the number of moments. Two consequences follow. Scenario reduction and discretization methods for mixed-integer models can be judged by a polyhedral discrepancy; and, through the entropy estimates of §4 of the paper, empirical approximations converge at rates governed by uniform laws for polyhedra. Corollary 3.7 shows that for bounded supports the rate becomes Lipschitz in the polyhedral discrepancy.

The results are proved in the paper, partly by reference: Lemma 3.5 is quoted from Bank et al., Schultz, and Blair–Jeroslow, and the partition argument behind (12) refers to Schultz (1996). None of it has a machine-checked proof. Formalizing the mission produces the first formal treatment of mixed-integer value functions with rational data, a formal version of the minimal-information-distance stability theorems in the objective-only case, and the polyhedral metric itself.

Difficulty

The obvious route to a stability estimate, a Lipschitz bound on ξ↦f0(ξ,x)\xi\mapsto f_0(\xi,x)ξ↦f0​(ξ,x) followed by a Fortet–Mourier or Wasserstein bound, fails at the first step: Φ\PhiΦ jumps across the boundaries of the pieces Bi\mathcal B_iBi​, so f0(⋅,x)f_0(\cdot,x)f0​(⋅,x) is only piecewise Lipschitz, with pieces that move with xxx. Any metric that controls dFUd_{\mathcal F_{\mathcal U}}dFU​​ must therefore see the indicator functions of these pieces, which is why polyhedra enter. The second obstacle is that there are countably many unbounded pieces: a single polyhedral bound covers only a bounded region of T\mathcal TT, and the tail outside it must be paid for with moments, which is the origin of the exponent. Lemma 3.5 itself, in particular the uniform Lipschitz constant and the local finiteness of the partition, rests on rationality of WWW and Wˉ\bar WWˉ; without (B1) the value function need not be lower semicontinuous.

Formalization scope

Rm\mathbb R^mRm, Rs\mathbb R^sRs, Rr\mathbb R^rRr are Euclidean spaces (the paper never fixes its norm, and every constant in the statements is existential). Measures are Borel probability measures on Rs\mathbb R^sRs with ν(Rs∖Ξ)=0\nu(\mathbb R^s\setminus\Xi)=0ν(Rs∖Ξ)=0. The integral of the extended-real integrand f0f_0f0​ is the published DupacovaWets.Consistency.expect (+∞+\infty+∞ when ∫f0+=∞\int f_0^+=\infty∫f0+​=∞). Φ\PhiΦ, optimal values and ψ\psiψ are extended-real infima, +∞+\infty+∞ on empty sets; the paper's "min" is read as an infimum. Distances and Ψ\PsiΨ take values in [0,∞][0,\infty][0,∞], and every estimate between optimal values states explicitly that both are finite, so that no convention for ∞−∞\infty-\infty∞−∞ can make a bound vacuous. T(ξ)T(\xi)T(ξ) is a linear map Rm→Rr\mathbb R^m\to\mathbb R^rRm→Rr and ∥T(ξ)∥\|T(\xi)\|∥T(ξ)∥ its operator norm. Integer vectors are tuples of natural numbers; WWW, Wˉ\bar WWˉ are real matrices with rational entries.

Disclosed reading: a "polyhedron with at most kkk faces" is an intersection of kkk closed half-spaces. Because kkk is existential in Theorem 3.6 and Corollary 3.7, the theorems are equivalent under this reading and under a facet count. The clause "piecewise polyhedral" of Lemma 3.5 is not defined in the paper and is not formalized. The page states "PFU(Ξ)⊆P1(Ξ)\mathcal P_{\mathcal F_{\mathcal U}}(\Xi)\subseteq\mathcal P_1(\Xi)PFU​​(Ξ)⊆P1​(Ξ)"; the argument proves and uses P1(Ξ)⊆PFU(Ξ)\mathcal P_1(\Xi)\subseteq\mathcal P_{\mathcal F_{\mathcal U}}(\Xi)P1​(Ξ)⊆PFU​​(Ξ), and that direction is the one formalized. Theorems 2.2 and 2.3 are stated in their objective-only (d=0d=0d=0) form for a general normal integrand.

The integrals in d1,phkd_{1,phk}d1,phk​ are Bochner integrals over PPP, which are meaningful because measures in Pp,K(Ξ)\mathcal P_{p,K}(\Xi)Pp,K​(Ξ) with p>1p>1p>1 integrate max⁡{1,∥ξ∥}\max\{1,\|\xi\|\}max{1,∥ξ∥}. A formalization that took the supremum over all of Rs\mathbb R^sRs instead of PPP, quantified ν\nuν over all of P(Ξ)\mathcal P(\Xi)P(Ξ), or let a non-integrable function contribute 000 would not state the paper's theorem, and these readings are excluded.

Needed infrastructure: the structure theory of mixed-integer value functions with rational data (the partition of Lemma 3.5 and the attainment of the minimum in (10)), lower semicontinuity of integral functionals of normal integrands via Fatou's lemma, Berge's upper semicontinuity of parametric argmin maps, and approximation of polyhedral pieces by polyhedra. The value-function results and the general d=0d=0d=0 stability theorems are reusable beyond this mission. Proofs of any milestone, and alternative arguments for Lemma 3.5, are welcome.

Selected references

  • S. T. Rachev, W. Römisch, Quantitative stability in stochastic programming: The method of probability metrics, Math. Oper. Res. 27(4) (2002) 792–818. doi:10.1287/moor.27.4.792.304. Preprint used here: edoc.hu-berlin.de.
  • B. Bank, J. Guddat, D. Klatte, B. Kummer, K. Tammer, Non-Linear Parametric Optimization, Akademie-Verlag, Berlin, 1982. doi:10.1007/978-3-0348-6328-5
  • R. Schultz, Rates of convergence in stochastic programs with complete integer recourse, SIAM J. Optim. 6 (1996) 1138–1152. doi:10.1137/S1052623494271655
  • C. E. Blair, R. G. Jeroslow, The value function of a mixed integer program: I, Discrete Math. 19 (1977) 121–138. doi:10.1016/0012-365X(77)90028-0
  • R. T. Rockafellar, R. J.-B. Wets, Variational Analysis, Springer, 1998. doi:10.1007/978-3-642-02431-3
11 thms1 active userReviewed
Convex OptimizationOptimizationProbability·Captain: mikedeng1

Convex Quadratic and Semidefinite Programming Relaxations in Scheduling II: Randomized Rounding of the Time-Slot Convex Quadratic Relaxation (CQP) Is Within 2 for R | rᵢⱼ | Σ wⱼCⱼResearch Paper

Motivation

Scheduling jobs on unrelated parallel machines to minimize the total weighted completion time is a basic model of machine scheduling: each job may take a different time on each machine, and a job becomes available on a machine only at a machine-dependent release date. In the three-field notation the problem is R∣rij∣∑wjCjR \mid r_{ij} \mid \sum w_jC_jR∣rij​∣∑wj​Cj​. Already the sequencing problem on a single machine with release dates is strongly NP-hard, so the question is how well it can be approximated in polynomial time.

Timeline (as surveyed in §1 of the paper):

  • Phillips, Stein and Wein gave a performance guarantee of O(log⁡2n)O(\log^2 n)O(log2n) for a network-scheduling generalization.
  • Hall, Shmoys and Wein (1996/1997) gave the first constant factor, 16/316/316/3, from an interval-indexed linear programming relaxation (Math. Oper. Res. 22, 1997).
  • Schulz and Skutella (1997–1999) gave a randomized (2+ε)(2+\varepsilon)(2+ε)-approximation from a time-indexed linear programming relaxation in intervals of geometrically increasing length; its size depends on pmax⁡p_{\max}pmax​ and on ε\varepsilonε.
  • Skutella (J. ACM 48, 2001) replaced the linear program by a convex quadratic program in assignment variables of strongly polynomial size and showed, in §3, that randomized rounding of it is a 222-approximation.

This mission formalizes that last result.

Setting

There are nnn jobs JJJ and mmm machines. Job jjj has a weight wj≥0w_j \ge 0wj​≥0; on machine iii it has a processing time pij>0p_{ij} > 0pij​>0 and a release date rij≥0r_{ij} \ge 0rij​≥0. A nonpreemptive schedule runs every job without interruption on one machine, starting no earlier than its release date there, and each machine processes at most one job at a time. With completion times CjC_jCj​, the objective is ∑jwjCj\sum_j w_jC_j∑j​wj​Cj​.

Smith's order on machine iii: j≺ikj \prec_i kj≺i​k if wj/pij>wk/pikw_j/p_{ij} > w_k/p_{ik}wj​/pij​>wk​/pik​, or the ratios are equal and j<kj < kj<k.

Time slots. Sort the release dates on machine iii as ρi1≤⋯≤ρin\rho_{i_1} \le \dots \le \rho_{i_n}ρi1​​≤⋯≤ρin​​ and set ρin+1=∞\rho_{i_{n+1}} = \inftyρin+1​​=∞. The kkk-th time slot iki_kik​ holds the jobs started on machine iii within [ρik,ρik+1)[\rho_{i_k}, \rho_{i_{k+1}})[ρik​​,ρik+1​​). An assignment of jobs to slots is feasible if job jjj goes to a slot iki_kik​ with ρik≥rij\rho_{i_k} \ge r_{ij}ρik​​≥rij​. For assignment variables aikja_{i_kj}aik​j​, the slot starts are

si1=ρi1,sik+1=max⁡{ρik+1, sik+∑jaikjpij},s_{i_1} = \rho_{i_1}, \qquad s_{i_{k+1}} = \max\Big\{\rho_{i_{k+1}},\ s_{i_k} + \sum_j a_{i_kj}p_{ij}\Big\},si1​​=ρi1​​,sik+1​​=max{ρik+1​​, sik​​+j∑​aik​j​pij​},

and for a 0/1 assignment the completion time of j∈ikj \in i_kj∈ik​ is Cj=sik+pij+∑j′≺ij, j′∈ikpij′C_j = s_{i_k} + p_{ij} + \sum_{j' \prec_i j,\ j' \in i_k} p_{ij'}Cj​=sik​​+pij​+∑j′≺i​j, j′∈ik​​pij′​ (17).

(CQP). Minimize ZCQP(a)=∑jwjCˉj(a)Z_{CQP}(a) = \sum_j w_j \bar C_j(a)ZCQP​(a)=∑j​wj​Cˉj​(a) with

Cˉj(a)=∑i,kaikj(ρik+1+aikj2 pij+∑j′≺ijaikj′pij′)\bar C_j(a) = \sum_{i,k} a_{i_kj}\Big(\rho_{i_k} + \frac{1 + a_{i_kj}}{2}\,p_{ij} + \sum_{j' \prec_i j} a_{i_kj'}p_{ij'}\Big)Cˉj​(a)=i,k∑​aik​j​(ρik​​+21+aik​j​​pij​+j′≺i​j∑​aik​j′​pij′​)

subject to ∑i,kaikj=1\sum_{i,k} a_{i_kj} = 1∑i,k​aik​j​=1 (14), aikj=0a_{i_kj} = 0aik​j​=0 if ρik<rij\rho_{i_k} < r_{ij}ρik​​<rij​ (18), a≥0a \ge 0a≥0 (20), and ∑jaikjpij≤ρik+1−ρik\sum_j a_{i_kj}p_{ij} \le \rho_{i_{k+1}} - \rho_{i_k}∑j​aik​j​pij​≤ρik+1​​−ρik​​ (21). In matrix form the objective is bTa+12cTa+12aT(D+diag(c))ab^Ta + \tfrac12 c^Ta + \tfrac12 a^T(D + \mathrm{diag}(c))abTa+21​cTa+21​aT(D+diag(c))a, which is convex.

Randomized rounding. Each job jjj is put into exactly one slot iki_kik​, with probability aikja_{i_kj}aik​j​, the choices being pairwise independent across jobs. The resulting slot assignment is turned into a schedule by (15)–(17).

Formalization targets

Goal: Theorem 3.4 (pp. 19–20), formal core

For every instance with p>0p > 0p>0, w≥0w \ge 0w≥0, r≥0r \ge 0r≥0:

  1. every feasible slot assignment yields a feasible schedule with completion times (17);
  2. for every (CQP)-feasible aaa and every pairwise independent rounding μ\muμ of aaa,
Eμ[∑jwjCj]≤2 ZCQP(a);E_\mu\Big[\sum_j w_jC_j\Big] \le 2\,Z_{CQP}(a);Eμ​[j∑​wj​Cj​]≤2ZCQP​(a);
  1. every feasible schedule SSS admits a (CQP)-feasible aaa with ZCQP(a)≤∑jwjCj(S)Z_{CQP}(a) \le \sum_j w_jC_j(S)ZCQP​(a)≤∑j​wj​Cj​(S).

The bound is stated for every feasible aaa, so it needs no optimum of (CQP), and with (3) it gives the positive half of Corollary 3.6.

Milestones

  • Lemma 3.1 (p. 15): an optimal schedule exists whose slots are sequenced by ≺i\prec_i≺i​ without interruption.
  • p. 16: the schedule built from a feasible slot assignment is feasible, with completion times (17).
  • Lemma 3.2 (p. 16): rebuilding such a schedule from its slot assignment gives a feasible schedule with the same property that completes no job later.
  • p. 17: for 0/1 assignments the relaxed completion times (19) equal (17).
  • Lemma 3.3 (p. 18): the quadratic programming relaxation has an optimal solution with sik=ρiks_{i_k} = \rho_{i_k}sik​​=ρik​​ for all i,ki, ki,k.
  • Proof of Lemma 3.5 (p. 20): Ej↦ik[sik]≤2ρikE_{j \mapsto i_k}[s_{i_k}] \le 2\rho_{i_k}Ej↦ik​​[sik​​]≤2ρik​​.
  • Lemma 3.5 (p. 20): E[Cj]≤2Cˉj(a)E[C_j] \le 2\bar C_j(a)E[Cj​]≤2Cˉj​(a) for every job.

Significance

The result gives a 222-approximation for R∣rij∣∑wjCjR \mid r_{ij} \mid \sum w_jC_jR∣rij​∣∑wj​Cj​ from a program of polynomial size, with a short analysis that holds job by job. The relaxation (CQP) is also the basis of §4, where the same rounding compares nonpreemptive with preemptive schedules, and its relaxation property bounds the optimum of (CQP) against the optimal schedule within a factor 222.

The result is proved in the paper; none of it is formalized. A formal development adds a checked model of release-date schedules and time slots on unrelated machines, the exchange argument behind Smith's rule in the presence of release dates, and a measure-free treatment of pairwise independent rounding. These pieces are reusable for other slot-based relaxations.

Difficulty

The analysis of Lemma 3.5 is short; the work lies in the reduction to slot assignments. Lemma 3.1 is an exchange argument, but swapping two jobs can push a job into a different time slot, so the naive bubble sort within a slot does not obviously terminate in a slot-sequenced schedule. Lemma 3.2 needs an induction across slots that accounts for empty slots and ties ρik=ρik+1\rho_{i_k} = \rho_{i_{k+1}}ρik​​=ρik+1​​. Lemma 3.3 is a shifting argument on fractional solutions whose objective change is a quadratic computation, and the step size printed on p. 18 must be divided by piȷ^p_{i\hat\jmath}pi^​​ to land exactly on sik=ρiks_{i_k} = \rho_{i_k}sik​​=ρik​​. The relaxation conjunct of the goal chains Lemmas 3.1, 3.2, 3.3 and the p. 17 identity.

Formalization scope

Jobs are Fin n, machines Fin m, and slots Fin m × Fin n, all 0-based. ≺i\prec_i≺i​ is cross-multiplied, which equals the page's ratio order because p>0p > 0p>0. ρ\rhoρ is the sorted tuple of release dates; ρin+1=∞\rho_{i_{n+1}} = \inftyρin+1​​=∞ is never represented by a real number. Constraint (21) is imposed only for slots before the last, and a job's slot is the largest kkk with ρik≤Sj\rho_{i_k} \le S_jρik​​≤Sj​. Schedules are a machine and a real start time per job; two jobs on one machine must satisfy Cj≤SkC_j \le S_kCj​≤Sk​ or Ck≤SjC_k \le S_jCk​≤Sj​. Randomized rounding is a probability weight on slot assignments with marginals aikja_{i_kj}aik​j​ and product joint probabilities for distinct jobs; full independence is not required, and the product weight is one instance. The standing assumptions pij>0p_{ij} > 0pij​>0, wj≥0w_j \ge 0wj​≥0, rij≥0r_{ij} \ge 0rij​≥0 of p. 2 are hypotheses of every theorem; Lemmas 3.1 and 3.3, which assert existence, also assume m≥1m \ge 1m≥1.

A statement of the bound alone, E≤2ZCQP(a)E \le 2 Z_{CQP}(a)E≤2ZCQP​(a), would say nothing about schedules: conjuncts 1 and 3 of the goal are what tie it to the scheduling problem, and neither may be dropped.

Not formalized: polynomial-time solvability of (CQP), the running time of the algorithm, the tightness half of Corollary 3.6, and the first sentence of Lemma 3.1 as a statement about every optimal schedule (false when some wj=0w_j = 0wj​=0).

Contributions welcome: proofs of any milestone; a general library of time-slot schedules; and the convexity of (22) (positive semidefiniteness of D+diag(c)D + \mathrm{diag}(c)D+diag(c)), which is not a milestone here.

Selected references

  • M. Skutella, Convex quadratic and semidefinite programming relaxations in scheduling, Journal of the ACM 48(2), 2001. https://doi.org/10.1145/375827.375840
  • L. A. Hall, A. S. Schulz, D. B. Shmoys and J. Wein, Scheduling to minimize average completion time: off-line and on-line approximation algorithms, Mathematics of Operations Research 22(3), 1997. https://doi.org/10.1287/moor.22.3.513
  • A. S. Schulz and M. Skutella, Scheduling unrelated machines by randomized rounding, SIAM Journal on Discrete Mathematics 15(4), 2002. https://doi.org/10.1137/S0895480199357078
  • W. E. Smith, Various optimizers for single-stage production, Naval Research Logistics Quarterly 3, 1956. https://doi.org/10.1002/nav.3800030106
9 thms1 active userReviewed
Graph TheoryOptimization·Captain: mikedeng1

Solving Project Scheduling Problems by Minimum Cut Computations: Finite-Capacity n-Cuts Are Exactly the Feasible Schedules, and a Minimum a-b-Cut Has the Optimal CostResearch Paper

Motivation

Time-indexed formulations are a standard way to model scheduling problems as integer programs: a binary variable xjtx_{jt}xjt​ records whether job jjj starts at time ttt. They were introduced by Pritsker, Watters and Wolfe (1969) and give strong linear relaxations, at the price of many variables. In project scheduling, jobs are linked by time lags: a lag (i,j)(i,j)(i,j) of length dijd_{ij}dij​ requires Sj≥Si+dijS_j \ge S_i + d_{ij}Sj​≥Si​+dij​, and since dijd_{ij}dij​ may be negative, lags express both minimal and maximal distances between start times, i.e. time windows. Many project scheduling objectives (weighted completion times, net present value, earliness–tardiness, and the Lagrangian subproblems of resource-constrained scheduling) reduce to minimizing a sum of start-time dependent costs wjtw_{jt}wjt​ subject to such lags.

Möhring, Schulz, Stork and Uetz showed that this problem, despite being an integer program, is solved by one minimum cut computation in a directed graph built from the time-indexed variables. This makes the Lagrangian relaxation of resource-constrained project scheduling computationally practical, which is the use the paper puts it to.

Timeline.

  • 1969: Pritsker, Watters and Wolfe give a time-indexed 0/1 formulation of multiproject scheduling with the weak form of the precedence constraints.
  • 1970: Rhys and Balinski show that selection / minimum-weight closure problems reduce to minimum cut.
  • 1985: Chang and Edmonds reduce the case of ordinary precedence constraints and unit processing times to minimum-weight closure, hence to minimum cut.
  • 1987: Christofides, Alvarez-Valdés and Tamarit propose the strong (disaggregated) form (3) of the temporal constraints.
  • 2003: Möhring, Schulz, Stork and Uetz give a direct transformation for arbitrary time lags and arbitrary processing times (Management Science 49(3)).

Setting

Jobs form the set J={0,…,n}J = \{0, \dots, n\}J={0,…,n}; job jjj has an integral processing time pj≥0p_j \ge 0pj​≥0. A set L⊆J×JL \subseteq J \times JL⊆J×J of time lags is given, with integral lengths dijd_{ij}dij​. A schedule is an integral vector S=(S0,…,Sn)S = (S_0, \dots, S_n)S=(S0​,…,Sn​) of start times; it is feasible if Sj≥Si+dijS_j \ge S_i + d_{ij}Sj​≥Si​+dij​ for all (i,j)∈L(i,j) \in L(i,j)∈L and it lies within the horizon TTT: 0≤Sj0 \le S_j0≤Sj​ and Sj+pj≤TS_j + p_j \le TSj​+pj​≤T. Starting job jjj at time ttt costs wjt≥0w_{jt} \ge 0wjt​≥0. The earliest and latest feasible start times e(j)e(j)e(j) and ℓ(j)\ell(j)ℓ(j) are the minimum and the maximum of SjS_jSj​ over feasible schedules. Throughout, a feasible schedule is assumed to exist.

The integer program (1)–(5) has variables xjt∈Zx_{jt} \in \mathbb{Z}xjt​∈Z:

min⁡ w(x)=∑j∑t=0Twjtxjts.t.∑t=0Txjt=1,∑s=tTxis+∑s=0t+dij−1xjs≤1,xjt≥0,\min\ w(x) = \sum_{j} \sum_{t=0}^{T} w_{jt} x_{jt} \quad\text{s.t.}\quad \sum_{t=0}^{T} x_{jt} = 1,\qquad \sum_{s=t}^{T} x_{is} + \sum_{s=0}^{t+d_{ij}-1} x_{js} \le 1,\qquad x_{jt} \ge 0,min w(x)=j∑​t=0∑T​wjt​xjt​s.t.t=0∑T​xjt​=1,s=t∑T​xis​+s=0∑t+dij​−1​xjs​≤1,xjt​≥0,

for all j∈Jj \in Jj∈J, (i,j)∈L(i,j) \in L(i,j)∈L, t=0,…,Tt = 0, \dots, Tt=0,…,T, and all variables with ttt outside [e(j),ℓ(j)][e(j), \ell(j)][e(j),ℓ(j)] are zero.

The minimum cut digraph D=(V,A)D = (V, A)D=(V,A) has a source aaa, a sink bbb, and nodes vjtv_{jt}vjt​ for t=e(j),…,ℓ(j)+1t = e(j), \dots, \ell(j)+1t=e(j),…,ℓ(j)+1. Its arcs are the assignment arcs (vjt,vj,t+1)(v_{jt}, v_{j,t+1})(vjt​,vj,t+1​) for t=e(j),…,ℓ(j)t = e(j), \dots, \ell(j)t=e(j),…,ℓ(j), of capacity wjtw_{jt}wjt​; the temporal arcs (vit,vj,t+dij)(v_{it}, v_{j,t+d_{ij}})(vit​,vj,t+dij​​) for (i,j)∈L(i,j) \in L(i,j)∈L and e(i)+1≤t≤ℓ(i)e(i)+1 \le t \le \ell(i)e(i)+1≤t≤ℓ(i), e(j)+1≤t+dij≤ℓ(j)e(j)+1 \le t+d_{ij} \le \ell(j)e(j)+1≤t+dij​≤ℓ(j); and the auxiliary arcs (a,vj,e(j))(a, v_{j,e(j)})(a,vj,e(j)​) and (vj,ℓ(j)+1,b)(v_{j,\ell(j)+1}, b)(vj,ℓ(j)+1​,b). Temporal and auxiliary arcs have infinite capacity. An aaa-bbb-cut (X,Xˉ)(X, \bar X)(X,Xˉ) splits VVV with a∈Xa \in Xa∈X, b∈Xˉb \in \bar Xb∈Xˉ; its capacity c(X,Xˉ)c(X, \bar X)c(X,Xˉ) is the total capacity of the arcs from XXX to Xˉ\bar XXˉ. An nnn-cut is an aaa-bbb-cut containing exactly one assignment arc of every job. The mapping (7) sends a cut to xxx with xjt=1x_{jt} = 1xjt​=1 iff (vjt,vj,t+1)(v_{jt}, v_{j,t+1})(vjt​,vj,t+1​) is in the cut.

Formalization targets

Goal: Theorem 1 (p. 7)

(7) is a bijection {n-cuts with c(X,Xˉ)<∞}→{feasible x of (1)–(5)},c(X,Xˉ)=w(x),\text{(7) is a bijection } \{n\text{-cuts with } c(X,\bar X) < \infty\} \to \{\text{feasible } x \text{ of (1)–(5)}\},\qquad c(X, \bar X) = w(x),(7) is a bijection {n-cuts with c(X,Xˉ)<∞}→{feasible x of (1)–(5)},c(X,Xˉ)=w(x), andc(X,Xˉ)=w(x∗) for every minimum a-b-cut (X,Xˉ) and some optimal x∗.\text{and}\quad c(X, \bar X) = w(x^*) \text{ for every minimum } a\text{-}b\text{-cut } (X, \bar X) \text{ and some optimal } x^*.andc(X,Xˉ)=w(x∗) for every minimum a-b-cut (X,Xˉ) and some optimal x∗.

Milestones (in the paper's order of proof)

  1. Proof of Lemma 1 (p. 7): for every lag (i,j)∈L(i,j) \in L(i,j)∈L, e(i)+dij≤e(j)e(i) + d_{ij} \le e(j)e(i)+dij​≤e(j).
  2. Lemma 1 (p. 6): every minimum aaa-bbb-cut has an nnn-cut of the same capacity.
  3. Lemma 2 (p. 8): every feasible xxx is the image of an nnn-cut (X,Xˉ)(X, \bar X)(X,Xˉ) with w(x)=c(X,Xˉ)w(x) = c(X, \bar X)w(x)=c(X,Xˉ).
  4. p. 8: the mapping (7) is injective on finite-capacity nnn-cuts.
  5. Lemma 3 (p. 8): the image of a finite-capacity nnn-cut is feasible, with w(x)=c(X,Xˉ)w(x) = c(X, \bar X)w(x)=c(X,Xˉ).
  6. Lemma 4 (p. 8): the minimum cut capacity equals the optimal value.

A companion statement (p. 8) covers strictly positive weights: then every minimum aaa-bbb-cut is an nnn-cut, and (7) is a bijection between minimum aaa-bbb-cuts and optimal solutions.

Significance

The result. Theorem 1 shows that the time-indexed integer program with arbitrary (also negative) time lags has an integral linear relaxation in disguise: it is a minimum cut problem, solvable in strongly polynomial time, O(nmT2log⁡(n2T/m))O(nmT^2 \log(n^2T/m))O(nmT2log(n2T/m)) with push-relabel (Corollary 1). This is what allows the paper to compute Lagrangian lower bounds for resource-constrained project scheduling with time windows by repeated minimum cut computations, and to use the resulting dual information for list scheduling heuristics (§§3–4). The construction is a direct reduction, sparser than the closure-based route of Chang and Edmonds.

Formalizing it. The result is proved on paper; no machine-checked proof is known. The mission produces a formal model of the time-indexed project scheduling IP with time lags and of its cut digraph, together with checked versions of the correspondence and of each lemma of its proof. The model of cuts with infinite capacities and tagged parallel arcs is reusable for other cut-based reductions (selection, closure, and image segmentation problems).

Difficulty

The bijection itself is bookkeeping. The substance is in two places. First, Lemma 1: a minimum cut need not be an nnn-cut (a zero-weight assignment arc can be cut twice), and repairing it to the canonical prefix cut must not introduce a temporal arc into the cut. The repair argument shifts a hypothetical violating temporal arc back along the lag, and needs that the lags are consistent with the earliest and latest start times, e(i)+dij≤e(j)e(i) + d_{ij} \le e(j)e(i)+dij​≤e(j) and ℓ(i)+dij≤ℓ(j)\ell(i) + d_{ij} \le \ell(j)ℓ(i)+dij​≤ℓ(j), facts about the feasible schedules rather than about the graph. Second, Lemma 3: temporal arcs exist only for ttt in a restricted window, so a finite-capacity nnn-cut satisfies the temporal constraint (3) only after a boundary case analysis at t=e(i)t = e(i)t=e(i) and t+dij>ℓ(j)t + d_{ij} > \ell(j)t+dij​>ℓ(j). The naive reading "every temporal constraint is an infinite arc" is false at the boundary.

Formalization scope

An instance is a Lean structure with fields p,L,d,T,wp, L, d, T, wp,L,d,T,w; jobs are Fin (n + 1), start times are integers, and xxx is a function J→N→ZJ \to \mathbb{N} \to \mathbb{Z}J→N→Z (integrality by type). The standing assumptions of §2 are hypotheses of every theorem: wjt≥0w_{jt} \ge 0wjt​≥0 (p. 5) and the existence of a feasible schedule (p. 4). The horizon condition 0≤Sj0 \le S_j0≤Sj​, Sj+pj≤TS_j + p_j \le TSj​+pj​≤T is read from "t=0,…,Tt = 0, \dots, Tt=0,…,T", "e(j)≥0e(j) \ge 0e(j)≥0" and "ℓ(j)≤T−pj\ell(j) \le T - p_jℓ(j)≤T−pj​" (p. 5), which the paper never writes as a formula. The paper's artificial jobs 000 and nnn with zero processing time are not assumed: §2.2 never uses them, so the statements are stronger. e(j)e(j)e(j) and ℓ(j)\ell(j)ℓ(j) are defined as the minimum and maximum feasible start times, never taken as parameters. Capacities live in [0,∞][0, \infty][0,∞] with ∞\infty∞ on temporal and auxiliary arcs; arcs are a tagged type, so parallel arcs and self-lags remain distinct. A cut is its source side X⊆VX \subseteq VX⊆V. Running-time claims (the O(nT)O(nT)O(nT) clause of Lemma 1, Corollary 1, the Goldberg–Rao bound) are not formalized.

Trivializing formalizations are ruled out: the integer program is not replaced by schedules, infinite capacities are not big-M constants, the minimum is taken over all aaa-bbb-cuts and not only over nnn-cuts, and e,ℓe, \elle,ℓ are computed from the instance rather than quantified over.

Needed infrastructure: finite sums in [0,∞][0, \infty][0,∞], minimum and maximum of finite sets of integers, and elementary reasoning on integer intervals. Proofs of any milestone, and of the auxiliary fact ℓ(i)+dij≤ℓ(j)\ell(i) + d_{ij} \le \ell(j)ℓ(i)+dij​≤ℓ(j) used in Lemma 3, are welcome.

Selected references

  • R. H. Möhring, A. S. Schulz, F. Stork, M. Uetz, Solving project scheduling problems by minimum cut computations, Management Science 49(3):330–350, 2003. https://doi.org/10.1287/mnsc.49.3.330.12737 (formalized from the authors' manuscript, July 2000, revised April and November 2002).
  • A. A. B. Pritsker, L. J. Watters, P. M. Wolfe, Multiproject scheduling with limited resources: a zero-one programming approach, Management Science 16(1):93–108, 1969. https://doi.org/10.1287/mnsc.16.1.93
  • G. J. Chang, J. Edmonds, The poset scheduling problem, Order 2(2):113–118, 1985. https://doi.org/10.1007/BF00334849
  • N. Christofides, R. Alvarez-Valdés, J. M. Tamarit, Project scheduling with resource constraints: a branch and bound approach, European Journal of Operational Research 29(3):262–273, 1987. https://doi.org/10.1016/0377-2217(87)90240-2
  • J. M. W. Rhys, A selection problem of shared fixed costs and network flows, Management Science 17(3):200–207, 1970. https://doi.org/10.1287/mnsc.17.3.200
  • A. V. Goldberg, R. E. Tarjan, A new approach to the maximum-flow problem, Journal of the ACM 35(4):921–940, 1988. https://doi.org/10.1145/48014.61051
8 thms1 active userReviewed
ProbabilityStochastic Systems·Captain: mikedeng1

Is Network Traffic Approximated by Stable Lévy Motion or Fractional Brownian Motion? 1: Infinite Source Poisson Input Under Slow Growth Converges (fidi) to Totally Skewed α-Stable Lévy MotionResearch Paper

Motivation

Measurements of Ethernet and Internet traffic in the 1990s showed that the amount of data offered to a link is bursty on every time scale and that the lengths of transmissions (file sizes, connection durations) have heavy, regularly varying tails. Two families of approximations for the cumulative input were proposed: fractional Brownian motion, which has dependent Gaussian increments, and α-stable Lévy motion, which has independent, heavy-tailed increments. The two lead to very different predictions for buffer overflow and link dimensioning.

T. Mikosch, S. Resnick, H. Rootzén and A. Stegeman (Ann. Appl. Probab. 12 (2002) 23–68) showed that both answers are correct in different regimes, and that the regime is decided by how fast the connection rate grows relative to the time scale. This mission formalizes their result for the infinite source Poisson model under slow growth, where the limit is a totally skewed α-stable Lévy motion. Companion missions treat the ON/OFF model and the fast-growth regime with its fractional Brownian limit.

Setting

A transmission length has law FonF_{\mathrm{on}}Fon​ on [0,∞)[0,\infty)[0,∞) with tail Fˉon(x)=Fon((x,∞))\bar F_{\mathrm{on}}(x)=F_{\mathrm{on}}((x,\infty))Fˉon​(x)=Fon​((x,∞)). Condition (2.8) asks

Fˉon(x)=x−αL(x),x>0,1<α<2,\bar F_{\mathrm{on}}(x)=x^{-\alpha}L(x),\qquad x>0,\quad 1<\alpha<2,Fˉon​(x)=x−αL(x),x>0,1<α<2,

with LLL slowly varying (L(cx)/L(x)→1L(cx)/L(x)\to1L(cx)/L(x)→1 for every c>0c>0c>0). The mean μon\mu_{\mathrm{on}}μon​ is finite and the variance infinite. The quantile function (2.9) is b(t)=(1/Fˉon)←(t)=inf⁡{x:1/Fˉon(x)≥t}b(t)=(1/\bar F_{\mathrm{on}})^{\leftarrow}(t)=\inf\{x:1/\bar F_{\mathrm{on}}(x)\ge t\}b(t)=(1/Fˉon​)←(t)=inf{x:1/Fˉon​(x)≥t}.

In the TTT-th model, connections start at the points (Γk)k∈Z(\Gamma_k)_{k\in\mathbb Z}(Γk​)k∈Z​ of a homogeneous Poisson process on R\mathbb RR with rate λ=λ(T)\lambda=\lambda(T)λ=λ(T), labelled so that Γ0<0<Γ1\Gamma_0<0<\Gamma_1Γ0​<0<Γ1​. Connection kkk transmits at unit rate for a time XkX_kXk​; the XkX_kXk​ are iid with law FonF_{\mathrm{on}}Fon​ and independent of the Γk\Gamma_kΓk​. The number of active connections and the cumulative input are

N(t)=∑k∈Z1[Γk≤t<Γk+Xk],A(t)=∫0tN(s) ds.N(t)=\sum_{k\in\mathbb Z}\mathbf 1[\Gamma_k\le t<\Gamma_k+X_k],\qquad A(t)=\int_0^tN(s)\,ds .N(t)=k∈Z∑​1[Γk​≤t<Γk​+Xk​],A(t)=∫0t​N(s)ds.

The rate λ(T)\lambda(T)λ(T) is a positive non-decreasing function of TTT. Slow Growth Condition 1 is

lim⁡T→∞b(λT)T=0,\lim_{T\to\infty}\frac{b(\lambda T)}{T}=0,T→∞lim​Tb(λT)​=0,

where b(λT)b(\lambda T)b(λT) is bbb evaluated at λ(T) T\lambda(T)\,Tλ(T)T.

The stable law Sα(σ,β,μ)S_\alpha(\sigma,\beta,\mu)Sα​(σ,β,μ) is the law with characteristic function exp⁡{−σα∣θ∣α(1−iβ sign(θ)tan⁡(πα/2))+iμθ}\exp\{-\sigma^\alpha|\theta|^\alpha(1-i\beta\,\mathrm{sign}(\theta)\tan(\pi\alpha/2))+i\mu\theta\}exp{−σα∣θ∣α(1−iβsign(θ)tan(πα/2))+iμθ} for α≠1\alpha\ne1α=1. α-stable Lévy motion Xα,σ,βX_{\alpha,\sigma,\beta}Xα,σ,β​ has independent stationary increments with X(t)−X(s)∼Sα(σ(t−s)1/α,β,0)X(t)-X(s)\sim S_\alpha(\sigma(t-s)^{1/\alpha},\beta,0)X(t)−X(s)∼Sα​(σ(t−s)1/α,β,0). Write

Cα=1−αΓ(2−α)cos⁡(πα/2)(5.1).C_\alpha=\frac{1-\alpha}{\Gamma(2-\alpha)\cos(\pi\alpha/2)}\quad(5.1).Cα​=Γ(2−α)cos(πα/2)1−α​(5.1).

Formalization targets

Goal: Theorem 1 (corrected scale)

Under (2.8) and Condition 1,

A(T⋅)−Tλμon(⋅)b(λT)→ fidi Xα,Cα−1/α,1(⋅),\frac{A(T\cdot)-T\lambda\mu_{\mathrm{on}}(\cdot)}{b(\lambda T)}\xrightarrow{\ fidi\ }X_{\alpha,C_\alpha^{-1/\alpha},1}(\cdot),b(λT)A(T⋅)−Tλμon​(⋅)​ fidi ​Xα,Cα−1/α​,1​(⋅),

convergence of the finite-dimensional distributions to a totally skewed (β=1\beta=1β=1) α-stable Lévy motion.

On the scale. The paper prints the limit Xα,1,1X_{\alpha,1,1}Xα,1,1​. Its own proof establishes λT P(j1>b(λT)x)→x−α\lambda T\,P(j_1>b(\lambda T)x)\to x^{-\alpha}λTP(j1​>b(λT)x)→x−α (pp. 37–38), which makes αx−α−1dx\alpha x^{-\alpha-1}dxαx−α−1dx the Lévy measure of the limit. The totally skewed stable law with that Lévy measure has σα=Γ(1−α)cos⁡(πα/2)=Cα−1\sigma^\alpha=\Gamma(1-\alpha)\cos(\pi\alpha/2)=C_\alpha^{-1}σα=Γ(1−α)cos(πα/2)=Cα−1​ (Samorodnitsky–Taqqu, Property 1.2.15; also the paper's own criterion on p. 47 with c=1c=1c=1, and the σ\sigmaσ the paper prints in Theorem 2). Since CαC_\alphaCα​ decreases from 2/π2/\pi2/π to 000 on (1,2)(1,2)(1,2) (C1.5≈0.399C_{1.5}\approx0.399C1.5​≈0.399), the printed limit has the wrong scale and the theorem as printed is false. The goal states the corrected scale Cα−1/αC_\alpha^{-1/\alpha}Cα−1/α​; milestones quote the page as printed.

Milestones, in attack order

(2.14), the covariance of NNN; Lemma 1 part 1 (Condition 1   ⟺  λTFˉon(T)→0  ⟺  Cov(NT(0),NT(T))→0\iff\lambda T\bar F_{\mathrm{on}}(T)\to0\iff\mathrm{Cov}(N_T(0),N_T(T))\to0⟺λTFˉon​(T)→0⟺Cov(NT​(0),NT​(T))→0); Lemma 2, slow part (λT2Fˉon(T)/b(λT)→0\lambda T^2\bar F_{\mathrm{on}}(T)/b(\lambda T)\to0λT2Fˉon​(T)/b(λT)→0); (4.15), negligibility of the boundary pieces A2,A3,A4A_2,A_3,A_4A2​,A3​,A4​; the tail limit of j1j_1j1​; (4.17), A12=OP([λT]1/2)=oP(b(λT))A_{12}=O_P([\lambda T]^{1/2})=o_P(b(\lambda T))A12​=OP​([λT]1/2)=oP​(b(λT)); (4.20)–(4.21), A13=o(b(λT))A_{13}=o(b(\lambda T))A13​=o(b(λT)); (4.19), corrected, the stable limit of A11A_{11}A11​; the corrected one-dimensional limit of A(T)A(T)A(T); and EA22=o(b(λT))EA_{22}=o(b(\lambda T))EA22​=o(b(λT)) of §4.5, which reduces the fidi convergence to the one-dimensional one.

Significance

The theorem shows that under slow connection growth the cumulative input is asymptotically a process with independent, infinite-variance increments. Long-range dependence present in each model (the covariance (2.14) decays like h−(α−1)L(h)h^{-(\alpha-1)}L(h)h−(α−1)L(h)) disappears on the time scale TTT, and the heavy tail of the transmission lengths dominates. Together with the fast-growth result (fractional Brownian motion) it explains why both approximations are found in measured traffic, and it identifies the critical quantity b(λT)/Tb(\lambda T)/Tb(λT)/T.

The result is proved in the paper, modulo the scale. No part of it is machine-checked. The mission produces a formal statement of the infinite source Poisson model with its marked point process structure, a formal statement of stable laws and stable Lévy motion through their characteristic functions, and machine-checked versions of the paper's estimates. Correcting the printed scale is part of the output.

Difficulty

The pieces A2,A3,A4,A12,A13,A22A_2,A_3,A_4,A_{12},A_{13},A_{22}A2​,A3​,A4​,A12​,A13​,A22​ are controlled by first moments, regular variation and Lemma 2. The central step is the stable limit of A11A_{11}A11​: a sum of a Poisson number of iid, heavy-tailed summands whose law changes with TTT. A classical central limit theorem does not apply because the variance is infinite, and the summands form a triangular array. The paper cites a point-process limit (Resnick's Exercise 4.4.2.8) for this step; a formal proof needs a convergence theorem for row-wise iid triangular arrays to stable laws, with the constant CαC_\alphaCα​ computed exactly.

Formalization scope

Each TTT has its own probability space carrying the TTT-th model, and every statement quantifies over all such families. The Poisson points are given by their iid exponential spacings −Γ0,Γ1,Γk+1−Γk-\Gamma_0,\Gamma_1,\Gamma_{k+1}-\Gamma_k−Γ0​,Γ1​,Γk+1​−Γk​ (k≠0)(k\ne0)(k=0), indexed by Z\mathbb ZZ. (2.8) is stated as "x↦xαFˉon(x)x\mapsto x^\alpha\bar F_{\mathrm{on}}(x)x↦xαFˉon​(x) is slowly varying", with lengths non-negative. b(t)b(t)b(t) is inf⁡{x>0:tFˉon(x)≤1}\inf\{x>0:t\bar F_{\mathrm{on}}(x)\le1\}inf{x>0:tFˉon​(x)≤1}, which equals the page's generalized inverse for t>0t>0t>0. N(t)N(t)N(t) is a cardinality and the region sums are sums of non-negative terms; their junk values (an infinite set, a non-summable family) occur only on null events. AAA is an interval integral. The growth hypothesis is: λ>0\lambda>0λ>0, non-decreasing, Condition 1. No other hypothesis is added.

The fidi limit is pinned by its characteristic function on Rk\mathbb R^kRk for sorted times 0≤t1≤⋯≤tk0\le t_1\le\dots\le t_k0≤t1​≤⋯≤tk​: the product of the characteristic functions of Sα(σ(tj−tj−1)1/α,1,0)S_\alpha(\sigma(t_j-t_{j-1})^{1/\alpha},1,0)Sα​(σ(tj​−tj−1​)1/α,1,0) evaluated at θj+⋯+θk\theta_j+\dots+\theta_kθj​+⋯+θk​. The existence of the limit law is part of the claim. One-dimensional limits are stated the same way on R\mathbb RR. Convergence in probability to 000 is PT(∣YT∣>η)→0P_T(|Y_T|>\eta)\to0PT​(∣YT​∣>η)→0 for every η>0\eta>0η>0.

Trivializing formalizations are excluded:

  • bbb is never the junk value 000, since its defining set is nonempty and bounded below.
  • The model hypothesis is satisfiable for every rate and length law; a sanity file builds it from product measures, together with a Pareto law satisfying (2.8) and a rate satisfying Condition 1.
  • The limit object is fixed by its characteristic function, not merely asserted to exist.
  • Only fidi convergence is claimed, as on the page. Convergence in the Skorokhod J1J_1J1​ topology fails for this model (Remark after Theorem 1).

A complete development needs regular variation (Karamata's theorem, Potter bounds, generalized inverses), the Poisson random measure of the marked points, and stable limits for triangular arrays. All three are reusable beyond this mission. Proofs of the milestones are welcome separately, as are general Mathlib-style lemmas on regular variation.

Selected references

  • T. Mikosch, S. Resnick, H. Rootzén, A. Stegeman, Is network traffic approximated by stable Lévy motion or fractional Brownian motion?, Ann. Appl. Probab. 12(1) (2002) 23–68. https://doi.org/10.1214/aoap/1015961155
  • G. Samorodnitsky, M. S. Taqqu, Stable Non-Gaussian Random Processes, Chapman & Hall, 1994. https://doi.org/10.1201/9780203738818
  • N. H. Bingham, C. M. Goldie, J. L. Teugels, Regular Variation, Cambridge University Press, 1987. https://doi.org/10.1017/CBO9780511721434
  • S. I. Resnick, Extreme Values, Regular Variation, and Point Processes, Springer, 1987. https://doi.org/10.1007/978-0-387-75953-1
12 thms1 active userReviewed
Dynamic ProgrammingOptimizationReinforcement Learning·Captain: mikedeng1

Bounded-parameter Markov Decision Processes 2: Interval Policy Evaluation Converges to the Interval Value Function of Every PolicyResearch Paper

Motivation

A Markov decision process (MDP) is evaluated by solving a linear fixed-point equation built from its transition probabilities. In practice those probabilities are rarely known exactly: they are estimated from data, elicited from experts, or obtained by aggregating the states of a larger model, and each estimate comes with an error band. Givan, Leach and Dean (Bounded-parameter Markov decision processes, Artificial Intelligence, 2000) replace each transition probability by a closed interval and ask what can still be said about the value of a policy. Their answer for a fixed policy is an interval value: at each state, the smallest and the largest expected discounted reward the policy can obtain over all exact MDPs consistent with the intervals.

The same model reappears under other names. Robust MDPs with rectangular uncertainty sets (Nilim and El Ghaoui, Operations Research, 2005; Iyengar, Mathematics of Operations Research, 2005) take the pessimistic end of the interval, and optimistic model-based reinforcement learning (for example UCRL-type algorithms) takes the optimistic end over a confidence set. Interval transition models are also used to bound the error of state aggregation. In all of these, the basic computational question is how to evaluate a fixed policy against a whole family of MDPs at once.

Setting

Fix a finite set QQQ of states and a finite nonempty set AAA of actions. An exact MDP M=⟨Q,A,F,R⟩M=\langle Q,A,F,R\rangleM=⟨Q,A,F,R⟩ has transition probabilities Fpq(α)F_{pq}(\alpha)Fpq​(α), each row q↦Fpq(α)q\mapsto F_{pq}(\alpha)q↦Fpq​(α) a probability distribution, a reward R(q)∈RR(q)\in\mathbb RR(q)∈R per state, and a discount rate 0≤γ<10\le\gamma<10≤γ<1. A policy is a map π:Q→A\pi:Q\to Aπ:Q→A. Its value function VM,πV_{M,\pi}VM,π​ is the expected discounted cumulative reward of the Markov chain with transition matrix Fpq(π(p))F_{pq}(\pi(p))Fpq​(π(p)):

VM,π(p)=∑t≥0γt (PπtR)(p),Pπ(p,q)=Fpq(π(p)).V_{M,\pi}(p)=\sum_{t\ge0}\gamma^t\,(P_\pi^tR)(p),\qquad P_\pi(p,q)=F_{pq}(\pi(p)).VM,π​(p)=t≥0∑​γt(Pπt​R)(p),Pπ​(p,q)=Fpq​(π(p)).

On the space V‾\overline VV of value functions v:Q→Rv:Q\to\mathbb Rv:Q→R, with the sup norm ∥v∥=max⁡q∣v(q)∣\|v\|=\max_q|v(q)|∥v∥=maxq​∣v(q)∣, the value-iteration operator is VIM,π(v)(p)=R(p)+γ∑qFpq(π(p)) v(q)VI_{M,\pi}(v)(p)=R(p)+\gamma\sum_{q}F_{pq}(\pi(p))\,v(q)VIM,π​(v)(p)=R(p)+γ∑q​Fpq​(π(p))v(q), and VIM,αVI_{M,\alpha}VIM,α​ is the same with the action α\alphaα in place of π(p)\pi(p)π(p). An operator TTT on V‾\overline VV is a contraction mapping if ∥Tv−Tu∥≤λ∥v−u∥\|Tv-Tu\|\le\lambda\|v-u\|∥Tv−Tu∥≤λ∥v−u∥ for all u,vu,vu,v and some 0≤λ<10\le\lambda<10≤λ<1.

A bounded-parameter MDP (BMDP) M↕M_\updownarrowM↕​ gives intervals [F↓pq(α),F↑pq(α)]⊆[0,1][F_{\downarrow pq}(\alpha),F_{\uparrow pq}(\alpha)]\subseteq[0,1][F↓pq​(α),F↑pq​(α)]⊆[0,1] with ∑qF↓pq(α)≤1≤∑qF↑pq(α)\sum_qF_{\downarrow pq}(\alpha)\le1\le\sum_qF_{\uparrow pq}(\alpha)∑q​F↓pq​(α)≤1≤∑q​F↑pq​(α), a reward RRR and a discount rate γ\gammaγ. An exact MDP belongs to M↕M_\updownarrowM↕​ if it has the same RRR and γ\gammaγ and every transition probability lies in its interval. The interval value of a policy π\piπ (Definition 3) is

V↕π(q)=[V↓π(q),V↑π(q)]=[min⁡M∈M↕VM,π(q), max⁡M∈M↕VM,π(q)].V_{\updownarrow\pi}(q)=\Big[V_{\downarrow\pi}(q),V_{\uparrow\pi}(q)\Big]=\Big[\min_{M\in M_\updownarrow}V_{M,\pi}(q),\ \max_{M\in M_\updownarrow}V_{M,\pi}(q)\Big].V↕π​(q)=[V↓π​(q),V↑π​(q)]=[M∈M↕​min​VM,π​(q), M∈M↕​max​VM,π​(q)].

Interval policy evaluation acts on an interval value function V↕=[V↓,V↑]V_\updownarrow=[V_\downarrow,V_\uparrow]V↕​=[V↓​,V↑​] by

IVI↕π(V↕)(p)=[min⁡M∈M↕VIM,π(V↓)(p), max⁡M∈M↕VIM,π(V↑)(p)],IVI_{\updownarrow\pi}(V_\updownarrow)(p)=\Big[\min_{M\in M_\updownarrow}VI_{M,\pi}(V_\downarrow)(p),\ \max_{M\in M_\updownarrow}VI_{M,\pi}(V_\uparrow)(p)\Big],IVI↕π​(V↕​)(p)=[M∈M↕​min​VIM,π​(V↓​)(p), M∈M↕​max​VIM,π​(V↑​)(p)],

with lower and upper bound maps IVI↓π,IVI↑π:V‾→V‾IVI_{\downarrow\pi},IVI_{\uparrow\pi}:\overline V\to\overline VIVI↓π​,IVI↑π​:V→V.

For the optional part, interval value iteration adds a maximization over actions: IVI↕optIVI_{\updownarrow opt}IVI↕opt​ and IVI↕pesIVI_{\updownarrow pes}IVI↕pes​ take, at each state, the largest of the candidate intervals [min⁡MVIM,α(V↓)(p),max⁡MVIM,α(V↑)(p)]\big[\min_MVI_{M,\alpha}(V_\downarrow)(p),\max_MVI_{M,\alpha}(V_\uparrow)(p)\big][minM​VIM,α​(V↓​)(p),maxM​VIM,α​(V↑​)(p)] for the lexicographic orders ≤opt\le_{opt}≤opt​ (upper end first) and ≤pes\le_{pes}≤pes​ (lower end first). For a value function VVV, ρV(p)\rho_V(p)ρV​(p) and σV(p)\sigma_V(p)σV​(p) are the actions maximizing max⁡MVIM,α(V)(p)\max_MVI_{M,\alpha}(V)(p)maxM​VIM,α​(V)(p) and min⁡MVIM,α(V)(p)\min_MVI_{M,\alpha}(V)(p)minM​VIM,α​(V)(p), and IVI↓opt,V(V′)=IVI↓opt([V′,V])IVI_{\downarrow opt,V}(V')=IVI_{\downarrow opt}([V',V])IVI↓opt,V​(V′)=IVI↓opt​([V′,V]), IVI↑pes,V(V′)=IVI↑pes([V,V′])IVI_{\uparrow pes,V}(V')=IVI_{\uparrow pes}([V,V'])IVI↑pes,V​(V′)=IVI↑pes​([V,V′]).

Formalization targets

Goal: convergence of interval policy evaluation (Section 5.1, p. 22)

For every BMDP, every policy π\piπ, and every interval value function V↕0=[V↓0,V↑0]V_{\updownarrow0}=[V_{\downarrow0},V_{\uparrow0}]V↕0​=[V↓0​,V↑0​] with V↓0≤V↑0V_{\downarrow0}\le V_{\uparrow0}V↓0​≤V↑0​,

lim⁡n→∞IVI↕π n(V↕0)=V↕π.\lim_{n\to\infty}IVI_{\updownarrow\pi}^{\,n}(V_{\updownarrow0})=V_{\updownarrow\pi}.n→∞lim​IVI↕πn​(V↕0​)=V↕π​.

The limit is the interval value of Definition 3, an extremum over infinitely many MDPs; nothing about the starting point is assumed beyond its being an interval value function.

Milestones on the goal's path

  • Theorem 10 (p. 21): IVI↓πIVI_{\downarrow\pi}IVI↓π​ and IVI↑πIVI_{\uparrow\pi}IVI↑π​ are contraction mappings on V‾\overline VV.
  • Theorem 11 (p. 22): IVI↓π(V↓π)=V↓πIVI_{\downarrow\pi}(V_{\downarrow\pi})=V_{\downarrow\pi}IVI↓π​(V↓π​)=V↓π​, IVI↑π(V↑π)=V↑πIVI_{\uparrow\pi}(V_{\uparrow\pi})=V_{\uparrow\pi}IVI↑π​(V↑π​)=V↑π​, and hence IVI↕π(V↕π)=V↕πIVI_{\updownarrow\pi}(V_{\updownarrow\pi})=V_{\updownarrow\pi}IVI↕π​(V↕π​)=V↕π​.

Further results on interval value iteration (Section 5.2)

  • Lemma 4 (p. 25): IVI↓opt,V(V′)(p)=max⁡α∈ρV(p)min⁡MVIM,α(V′)(p)IVI_{\downarrow opt,V}(V')(p)=\max_{\alpha\in\rho_V(p)}\min_{M}VI_{M,\alpha}(V')(p)IVI↓opt,V​(V′)(p)=maxα∈ρV​(p)​minM​VIM,α​(V′)(p) and IVI↑pes,V(V′)(p)=max⁡α∈σV(p)max⁡MVIM,α(V′)(p)IVI_{\uparrow pes,V}(V')(p)=\max_{\alpha\in\sigma_V(p)}\max_{M}VI_{M,\alpha}(V')(p)IVI↑pes,V​(V′)(p)=maxα∈σV​(p)​maxM​VIM,α​(V′)(p).
  • Theorem 13(a) (p. 25): IVI↑optIVI_{\uparrow opt}IVI↑opt​ and IVI↓pesIVI_{\downarrow pes}IVI↓pes​ are contraction mappings.
  • Theorem 13(b) (p. 25): for every value function VVV, IVI↓opt,VIVI_{\downarrow opt,V}IVI↓opt,V​ and IVI↑pes,VIVI_{\uparrow pes,V}IVI↑pes,V​ are contraction mappings.

Significance

The result. The interval value is defined as a minimum and a maximum over a continuum of MDPs, and there is no a priori reason it should be computable by a dynamic program. The goal says it is: one operator, whose per-state step is a small linear program over a box intersected with a simplex, iterated from any start, converges to both ends of the interval simultaneously. The pessimistic end is the robust value of the policy under rectangular uncertainty, so the goal is also a policy-evaluation theorem for interval-rectangular robust MDPs; the optimistic end is the value used by optimism-based exploration. Theorem 11 is the bridge between the "game" definition (nature picks the MDP) and the fixed-point characterization; Theorem 13 is the corresponding first step for the optimal interval values of the companion mission.

Formalizing it. The results are proved in the paper (2000), with Theorem 11 relying on the existence of a single member MDP that is π\piπ-minimizing at every state simultaneously (the paper's Theorem 7). No machine-checked version of the BMDP model or of these results is known. Nearby statements formalized elsewhere concern other models: robust MDPs with costs and general row sets, and optimistic Bellman equations over compact confidence sets. This mission produces the BMDP model and its interval operators, the contraction and fixed-point theorems, and the convergence theorem, all stated against the member-MDP definition of V↕πV_{\updownarrow\pi}V↕π​.

Difficulty

The contraction estimates (Theorems 10 and 13) follow the familiar pattern for Bellman operators, with the extra point that the minimum or maximum over the infinite family M↕M_\updownarrowM↕​ must be attained. The real difficulty is Theorem 11. The obvious argument compares V↓π(q)V_{\downarrow\pi}(q)V↓π​(q) with VIM,π(V↓π)(q)VI_{M,\pi}(V_{\downarrow\pi})(q)VIM,π​(V↓π​)(q) state by state; but V↓π(q)V_{\downarrow\pi}(q)V↓π​(q) is a minimum taken separately at each state, possibly by a different MDP at each state, and V↓πV_{\downarrow\pi}V↓π​ need not a priori be the value function of any single member MDP. Without a member MDP that is minimizing at all states at once, the inequality IVI↓π(V↓π)≥V↓πIVI_{\downarrow\pi}(V_{\downarrow\pi})\ge V_{\downarrow\pi}IVI↓π​(V↓π​)≥V↓π​ does not follow. Establishing that simultaneous minimizer is the central step.

Formalization scope

All statements live in the namespace BoundedParamMDP.IntervalEval. The exact MDP, the BMDP, its member MDPs, VM,πV_{M,\pi}VM,π​, VIM,πVI_{M,\pi}VIM,π​, VIM,αVI_{M,\alpha}VIM,α​ and Definition 3's V↓πV_{\downarrow\pi}V↓π​, V↑πV_{\uparrow\pi}V↑π​ are the shared definitions BoundedParamMDP.Optimal.MDP and BoundedParamMDP.Optimal.BMDP used by the companion mission, so its theorems apply to the same objects. Conventions:

  • QQQ and AAA are finite types, AAA is nonempty; QQQ may be empty (every statement then holds for the paper's own reasons). Policies are deterministic and stationary, Q → A.
  • An exact MDP is a structure with F p α q =Fpq(α)=F_{pq}(\alpha)=Fpq​(α), rows nonnegative and summing to 111, a state reward R, and the discount rate 0≤γ<10\le\gamma<10≤γ<1 as a field. Rewards are tight (the paper's footnote 3): members share the BMDP's R and γ.
  • VM,πV_{M,\pi}VM,π​ is the discounted series ∑tγtPπtR\sum_t\gamma^t P_\pi^tR∑t​γtPπt​R, not a solution chosen from the Bellman equation.
  • Minima and maxima over M↕M_\updownarrowM↕​ are the real infimum and supremum over the nonempty subtype Member B; all quantities are bounded, so these are the genuine values (and attained, by the paper's Lemma 1).
  • An interval value function is a pair (V↓, V↑) of functions Q → ℝ. The operators are defined on all pairs; the goal assumes V↓0≤V↑0V_{\downarrow0}\le V_{\uparrow0}V↓0​≤V↑0​ as the paper's interval notion, though the conclusion does not need it.
  • The norm is Mathlib's norm on Q → ℝ (the sup norm (4)); convergence of pairs is in the product topology. A contraction mapping is "there exists λ∈[0,1)\lambda\in[0,1)λ∈[0,1)", as on p. 6; the modulus is not fixed to γ\gammaγ.
  • max⁡α,≤opt\max_{\alpha,\le_{opt}}maxα,≤opt​​ is the supremum over the finite action set in the lexicographic order on (upper end, lower end); ≤pes\le_{pes}≤pes​ uses (lower end, upper end). ρV(p)\rho_V(p)ρV​(p), σV(p)\sigma_V(p)σV​(p) are sets of actions.
  • Theorem 13(a) is stated for IVI↑optIVI_{\uparrow opt}IVI↑opt​ as the upper end of IVI↕optIVI_{\updownarrow opt}IVI↕opt​ with arbitrary lower inputs on both sides: the bound by λ∥U′−U∥\lambda\|U'-U\|λ∥U′−U∥ includes the independence from the lower input that the paper uses (p. 23) to regard IVI↑optIVI_{\uparrow opt}IVI↑opt​ as a map on V‾\overline VV.
  • Lemma 4's second line is printed with min⁡M\min_{M}minM​; the evident reading max⁡M\max_MmaxM​ (eq. (29), proof of Theorem 13(b), p. 47) is stated.

A trivializing formalization is excluded: the limit in the goal is Definition 3's infimum and supremum over member MDPs, not the fixed point of IVI↕πIVI_{\updownarrow\pi}IVI↕π​, so the goal is not Banach's theorem alone; and the member type is nonempty for every BMDP, so the infima and suprema are not the junk value 000.

The proof of Theorem 11 uses the paper's Lemma 1 (attainment in the order-maximizing family), Theorem 3 (VM,πV_{M,\pi}VM,π​ is the unique fixed point of VIM,πVI_{M,\pi}VIM,π​), Theorem 6 (comparison principle) and Theorem 7 (existence of π\piπ-minimizing and π\piπ-maximizing MDPs). These are milestones of the companion mission Bounded-parameter Markov Decision Processes 1 and are not restated here. The Banach fixed-point theorem is Mathlib's ContractingWith. Contributions welcome: proofs of the milestones, a reusable lemma on the attained minimum of a linear function over a box intersected with the simplex, and the order-maximizing construction itself.

Selected references

  • R. Givan, S. Leach, T. Dean, Bounded-parameter Markov decision processes, Artificial Intelligence 122 (2000). https://doi.org/10.1016/S0004-3702(00)00047-3
  • A. Nilim, L. El Ghaoui, Robust control of Markov decision processes with uncertain transition matrices, Operations Research 53(5), 2005. https://doi.org/10.1287/opre.1050.0216
  • G. N. Iyengar, Robust dynamic programming, Mathematics of Operations Research 30(2), 2005. https://doi.org/10.1287/moor.1040.0129
  • M. L. Puterman, Markov Decision Processes: Discrete Stochastic Dynamic Programming, Wiley, 1994. https://doi.org/10.1002/9780470316887
10 thms1 active userReviewed
Dynamic ProgrammingOptimization·Captain: mikedeng1

Algorithms for Scheduling Runway Operations Under Constrained Position Shifting 1: Dynamic Programming on the CPS Network Gives the Minimum MakespanResearch Paper

Motivation

A busy runway is the bottleneck of an airport, and the order in which arriving or departing aircraft use it determines how many operations fit into an hour. Consecutive aircraft must be spaced by minimum separations that depend on both aircraft (a light aircraft behind a heavy one needs a long gap because of wake turbulence), so reordering the first-come-first-served (FCFS) queue can raise throughput considerably. Unrestricted reordering is unacceptable in practice: it makes the controllers' job harder and can push one aircraft to the back of the queue indefinitely. Following Dear (1976), constrained position shifting (CPS) allows each aircraft to move at most kkk positions away from its FCFS position, with kkk between 1 and 3 in practice.

Earlier algorithms for CPS either exploited a small number of aircraft types and ignored time windows and precedence constraints (Psaraftis 1980), or needed exponentially many parallel processors (Trivizas 1998), and Carr (2004) conjectured that runway scheduling under CPS has exponential complexity in general. Balakrishnan and Chandran (Oper. Res. 58(6), 2010) showed that, for fixed kkk, the problem with time windows and precedence constraints is solved by dynamic programming on a network whose size is linear in the number of aircraft. This mission formalizes the core of that result: the network, its precedence pruning, and the dynamic program for the minimum makespan.

Setting

There are n≥1n\ge 1n≥1 aircraft, labelled 0,…,n−10,\dots,n-10,…,n−1 in FCFS order, and the landing positions are also 0,…,n−10,\dots,n-10,…,n−1. An instance consists of

  1. the maximum position shift k∈Nk\in\mathbb Nk∈N;
  2. separations δab≥0\delta_{ab}\ge 0δab​≥0, the minimum time between a leading aircraft aaa and a trailing aircraft bbb, satisfying the triangle inequality δac≤δab+δbc\delta_{ac}\le\delta_{ab}+\delta_{bc}δac​≤δab​+δbc​ (which holds for wake-vortex separations in arrivals-only or departures-only operation);
  3. a time window [e(a),l(a)][e(a),l(a)][e(a),l(a)] for each aircraft;
  4. a finite set of precedence pairs (x,y)(x,y)(x,y) of distinct aircraft, meaning that xxx lands before yyy.

A kkk-CPS sequence is a bijection σ\sigmaσ from positions to aircraft with ∣σ(p)−p∣≤k|\sigma(p)-p|\le k∣σ(p)−p∣≤k for all ppp. A feasible schedule is such a σ\sigmaσ with landing times tpt_ptp​ such that e(σ(p))≤tp≤l(σ(p))e(\sigma(p))\le t_p\le l(\sigma(p))e(σ(p))≤tp​≤l(σ(p)), δσ(p)σ(q)≤tq−tp\delta_{\sigma(p)\sigma(q)}\le t_q-t_pδσ(p)σ(q)​≤tq​−tp​ for all positions p<qp<qp<q, and every precedence pair is respected. The makespan is tn−1t_{n-1}tn−1​, the landing time of the last aircraft.

The CPS network has stages p=1,…,np=1,\dots,np=1,…,n. A node of stage ppp is a list of min⁡{2k+1,p}\min\{2k+1,p\}min{2k+1,p} distinct aircraft occupying the positions that end at p−1p-1p−1, each within kkk of its FCFS position; its last entry is its final aircraft. An arc joins a stage-ppp node iii to a stage-(p+1)(p+1)(p+1) node jjj when the first min⁡{2k,p}\min\{2k,p\}min{2k,p} entries of jjj are the last min⁡{2k,p}\min\{2k,p\}min{2k,p} entries of iii. A source-sink path picks one node per stage, joined by arcs, and its sequence puts the final aircraft of the stage-(q+1)(q+1)(q+1) node at position qqq. The pruned network GGG deletes the nodes that violate a precedence pair (x,y)(x,y)(x,y): those containing yyy before xxx, or yyy at a position less than x−kx-kx−k, or xxx at a position greater than y+ky+ky+k.

On GGG, with e(j),l(j)e(j),l(j)e(j),l(j) the window of the final aircraft of jjj, δi(j)\delta_i(j)δi​(j) the separation between the final aircraft of iii and jjj, and P(j)P(j)P(j) the predecessors of jjj, the dynamic program is T∗(j)=e(j)T^*(j)=e(j)T∗(j)=e(j) at stage 1 and

T∗(j)=max⁡{e(j), min⁡i∈P(j): T∗(i)≤l(i)(T∗(i)+δi(j))}(1)T^*(j)=\max\Big\{e(j),\ \min_{i\in P(j):\,T^*(i)\le l(i)}\big(T^*(i)+\delta_i(j)\big)\Big\}\qquad(1)T∗(j)=max{e(j), i∈P(j):T∗(i)≤l(i)min​(T∗(i)+δi​(j))}(1)

with min⁡∅=+∞\min\emptyset=+\inftymin∅=+∞.

Formalization targets

Goal: §4.1, the minimum makespan rule

∃ feasible schedule  ⟺  ∃ j∈Gn: T∗(j)≤l(j),\exists\ \text{feasible schedule}\iff \exists\, j\in G_n:\ T^*(j)\le l(j),∃ feasible schedule⟺∃j∈Gn​: T∗(j)≤l(j), min⁡{makespan of a feasible schedule}=min⁡{T∗(j):j∈Gn, T∗(j)≤l(j)},\min\{\text{makespan of a feasible schedule}\}=\min\{T^*(j): j\in G_n,\ T^*(j)\le l(j)\},min{makespan of a feasible schedule}=min{T∗(j):j∈Gn​, T∗(j)≤l(j)},

where GnG_nGn​ is stage nnn of GGG; the second identity is stated as "τ\tauτ is the least makespan iff τ\tauτ is the least admissible T∗(j)T^*(j)T∗(j)" for every real τ\tauτ.

Milestones

  1. Theorem 1. σ\sigmaσ is a kkk-CPS sequence iff it is the sequence of a source-sink path of the CPS network.
  2. Lemma 1 (Case I). For a<ba<ba<b, a path whose sequence puts aaa after bbb has a node containing bbb before aaa.
  3. Lemma 2 (Case II). For a<ba<ba<b, in the position-constrained network, a path whose sequence puts aaa before bbb has a node containing aaa before bbb.
  4. §3.1 conclusion. The source-sink paths of GGG are exactly the kkk-CPS sequences respecting all precedence pairs.
  5. Lemma 3. For every node jjj of GGG, T∗(j)T^*(j)T∗(j) is the earliest landing time of the final aircraft of jjj over all partial schedules ending at jjj, and +∞+\infty+∞ exactly when there is none.

Significance

The network turns a search over permutations into a shortest-path-like computation over O(n(2k+1)2k+1)O(n(2k+1)^{2k+1})O(n(2k+1)2k+1) nodes, so for the small values of kkk used in practice the minimum makespan with time windows and precedence constraints is computed in time linear in nnn. The same network underlies the paper's algorithms for total delay (§5.2) and, in a time-expanded form, for arbitrary separable costs without the triangle inequality (§6).

The paper's proofs are short and informal. A formalization pins down exactly what T∗(j)T^*(j)T∗(j) means (the paper describes it only as "arrival time of the final aircraft of node iii in an optimal solution"), makes precise which nodes the precedence pruning removes, and checks the correctness of the stage-nnn rule. The result has not been machine-checked before.

Difficulty

Two features of the model make the correctness of the network nontrivial. First, a node records only a window of at most 2k+12k+12k+1 consecutive positions, so nothing local prevents a path from using the same aircraft twice or omitting one; Theorem 1 asserts that the global sequence is nevertheless a permutation, and Lemmas 1 and 2 assert that precedence, a global property of the sequence, is detected inside single nodes. Second, the recursion (1) adds the separation only between consecutive aircraft, while feasibility constrains every pair of aircraft; the triangle inequality is what reconciles the two, and without it Lemma 3 fails.

The obvious reading of §4.1 also fails: taking the least T∗(j)T^*(j)T∗(j) over all stage-nnn nodes, as the paper prints, can select a node whose final aircraft lands after its latest time. The goal filters stage nnn by T∗(j)≤l(j)T^*(j)\le l(j)T∗(j)≤l(j).

Formalization scope

All data are real numbers; aircraft and positions are Fin n, 0-based; stages are 1-based natural numbers; [NeZero n] encodes n≥1n\ge 1n≥1. Separations are pairwise in the definition of feasibility, never only consecutive. The recursion takes values in WithTop ℝ, with ⊤=+∞\top=+\infty⊤=+∞; optima are stated with IsLeast over sets of makespans, never with a real infimum. Precedence pairs are a Finset of ordered pairs of distinct aircraft, covering both of the paper's cases.

Explicit readings of loose phrases:

  • "corresponding source-sink path" (Theorem 1) is the path whose final aircraft, stage by stage, form the sequence, as defined in the proof;
  • "the values of T∗(⋅)T^*(\cdot)T∗(⋅)" (Lemma 3) are earliest landing times over partial schedules ending at the node: paths in GGG with times satisfying the windows of the earlier aircraft, the earliest time of the last one, and all pairwise separations;
  • "the minimum makespan is the lowest value of T∗(⋅)T^*(\cdot)T∗(⋅) among all nodes in stage nnn" is corrected to the nodes with T∗(j)≤l(j)T^*(j)\le l(j)T∗(j)≤l(j), which is the paper's own feasibility check;
  • the Case II position filter uses the printed rule (<b−k<b-k<b−k, >a+k>a+k>a+k).

Not formalized: running times (Proposition 1 and the O(k)O(k)O(k) preprocessing remark), the predecessor pointers and tie-breaking that reconstruct an optimal sequence, the removal of nodes unreachable from the source or the sink (it does not change the set of source-sink paths), asymmetric shifts (§3.2) and disjoint time windows.

A formalization that defines feasibility with consecutive separations only, defines T∗T^*T∗ by the recursion and states Lemma 3 as an unfolding, or states Theorem 1 as "some CPS sequence exists iff some path exists" would be trivial; the statements here avoid all three.

The definitions (instance, kkk-CPS sequences, the CPS network and its pruning) are shared with the two other missions of this series (total delay, discrete time) and may be consolidated later. Proofs of the milestones, and lemmas about the window structure of the network, are welcome.

Selected references

  • H. Balakrishnan, B. G. Chandran, Algorithms for Scheduling Runway Operations Under Constrained Position Shifting, Operations Research 58(6), 1650–1665, 2010. https://doi.org/10.1287/opre.1100.0869
  • H. N. Psaraftis, A Dynamic Programming Approach for Sequencing Groups of Identical Jobs, Operations Research 28(6), 1347–1359, 1980. https://doi.org/10.1287/opre.28.6.1347
  • R. G. Dear, The Dynamic Scheduling of Aircraft in the Near Terminal Area, MIT Flight Transportation Laboratory Report R76-9, 1976.
  • D. A. Trivizas, Optimal Scheduling with Maximum Position Shift (MPS) Constraints: A Runway Scheduling Application, Journal of Navigation 51(2), 250–266, 1998.
  • F. R. Carr, Robust Decision-Support Tools for Airport Surface Traffic, Ph.D. thesis, Massachusetts Institute of Technology, 2004.
8 thms1 active userReviewed
ProbabilityStochastic Systems·Captain: mikedeng1

Inventory Management of a Fast-Fashion Retail Network 1: Expected Sales Under the Major-Size Display Policy Are Non-Decreasing, Discretely Concave and SupermodularResearch Paper

Motivation

Fast-fashion retailers such as Zara replenish their stores several times a week from a central warehouse, and the warehouse holds a limited stock of each garment that must be split across hundreds of stores. Caro and Gallien built the store-level sales model behind the shipment-allocation system they developed with Zara and tested in a live pilot (Caro and Gallien, Inventory Management of a Fast-Fashion Retail Network, working paper, August 2, 2007; published in Operations Research 58(2), 2010). One display policy shapes the model. A garment (a reference) comes in several sizes, a few of which are designated major sizes. As soon as one major size runs out in a store, the whole reference is removed from the shop floor, so the remaining stock of the other sizes stops selling.

The allocation optimization only works if expected store sales have the right shape. They must grow with inventory, show decreasing marginal returns, and show complementarity across sizes. Proposition 1 of the paper establishes these properties for the model. The paper also cites the analogous result of Lu and Song (2003) for assemble-to-order systems as the template for its proof.

Setting

Fix a finite set of sizes S\mathcal SS, split into the major sizes S+\mathcal S^+S+ and the minor sizes S−=S∖S+\mathcal S^- = \mathcal S \setminus \mathcal S^+S−=S∖S+. Time t≥0t \ge 0t≥0 is measured from the last replenishment, and the next replenishment comes at time T>0T > 0T>0.

Demand is random. On a probability space, each size sss has a counting process Ns(t)N_s(t)Ns​(t), the number of sale opportunities for sss in [0,t][0, t][0,t]. It is a Poisson process with rate λs>0\lambda_s > 0λs​>0:

  • Ns(0)=0N_s(0) = 0Ns​(0)=0;
  • its paths are non-decreasing and right-continuous with jumps of size one;
  • the increment Ns(t)−Ns(u)N_s(t) - N_s(u)Ns​(t)−Ns​(u) is Poisson with mean λs(t−u)\lambda_s (t-u)λs​(t−u);
  • increments over disjoint intervals are independent.

The processes of different sizes are mutually independent. In Lean this family is IsPoissonFamily lam N P.

An inventory vector q=(qs)∈NSq = (q_s) \in \mathbb N^{\mathcal S}q=(qs​)∈NS gives the stock of each size right after replenishment. The virtual stockout time of size sss is

τs(qs)=inf⁡{t≥0:Ns(t)=qs},\tau_s(q_s) = \inf\{t \ge 0 : N_s(t) = q_s\},τs​(qs​)=inf{t≥0:Ns​(t)=qs​},

the time at which size sss would run out if it stayed on display. For a set of sizes A\mathcal AA, write τA=min⁡s∈Aτs(qs)\tau_{\mathcal A} = \min_{s \in \mathcal A} \tau_s(q_s)τA​=mins∈A​τs​(qs​) and a∧b=min⁡(a,b)a \wedge b = \min(a, b)a∧b=min(a,b). Under the major-size policy, the random sales over one period are

G(q)=∑s∈S+Ns(τS+∧T)+∑s∈S−Ns(τS+∪{s}∧T).(1)G(q) = \sum_{s \in \mathcal S^+} N_s(\tau_{\mathcal S^+} \wedge T) + \sum_{s \in \mathcal S^-} N_s(\tau_{\mathcal S^+ \cup \{s\}} \wedge T). \tag{1}G(q)=s∈S+∑​Ns​(τS+​∧T)+s∈S−∑​Ns​(τS+∪{s}​∧T).(1)

The expected sales function is g(q)=E[G(q)]g(q) = \mathbb E[G(q)]g(q)=E[G(q)] (expectedSales). The proof also uses hA(q)=E[τA∧T]h^{\mathcal A}(q) = \mathbb E[\tau_{\mathcal A} \wedge T]hA(q)=E[τA​∧T] (hA) and the marginal difference Δsf(q)=f(q+es)−f(q)\Delta_s f(q) = f(q + e_s) - f(q)Δs​f(q)=f(q+es​)−f(q) (delta), where ese_ses​ is the sss-th unit vector.

Formalization targets

Goal: Proposition 1

For every q∈NSq \in \mathbb N^{\mathcal S}q∈NS and all sizes s≠s′s \ne s's=s′:

g is non-decreasing,Δsg(q+es)≤Δsg(q),Δsg(q)≤Δsg(q+es′),g \text{ is non-decreasing}, \qquad \Delta_s g(q + e_s) \le \Delta_s g(q), \qquad \Delta_s g(q) \le \Delta_s g(q + e_{s'}),g is non-decreasing,Δs​g(q+es​)≤Δs​g(q),Δs​g(q)≤Δs​g(q+es′​),

and ggg is supermodular on the lattice NS\mathbb N^{\mathcal S}NS:

g(q)+g(q′)≤g(q∨q′)+g(q∧q′).g(q) + g(q') \le g(q \vee q') + g(q \wedge q').g(q)+g(q′)≤g(q∨q′)+g(q∧q′).

The statement is uniform in the rates, the horizon and the choice of major sizes, so it covers every instance of the model.

Milestones (Appendix §5.1 and §3.1.3, in the order the proof uses them)

  1. On every sample path, q↦G(q)q \mapsto G(q)q↦G(q) is non-decreasing.
  2. Identity (2): g(q)=λS+hS+(q)+∑s∈S−λshS+∪{s}(q)g(q) = \lambda_{\mathcal S^+} h^{\mathcal S^+}(q) + \sum_{s \in \mathcal S^-} \lambda_s h^{\mathcal S^+ \cup \{s\}}(q)g(q)=λS+​hS+(q)+∑s∈S−​λs​hS+∪{s}(q), where λS+=∑s∈S+λs\lambda_{\mathcal S^+} = \sum_{s \in \mathcal S^+} \lambda_sλS+​=∑s∈S+​λs​.
  3. For s∈As \in \mathcal As∈A, ΔshA(q)\Delta_s h^{\mathcal A}(q)Δs​hA(q) equals an integral over [0,T][0, T][0,T] of Poisson probabilities, and that integral equals 1λsP(τs(qs+1)≤τA∖{s}∧T)\frac{1}{\lambda_s}\mathbb P(\tau_s(q_s+1) \le \tau_{\mathcal A \setminus \{s\}} \wedge T)λs​1​P(τs​(qs​+1)≤τA∖{s}​∧T).
  4. ΔshA\Delta_s h^{\mathcal A}Δs​hA is non-increasing in qsq_sqs​ and non-decreasing in qs′q_{s'}qs′​ for s′≠ss' \ne ss′=s.
  5. On every sample path, q↦τA∧Tq \mapsto \tau_{\mathcal A} \wedge Tq↦τA​∧T is supermodular.
  6. hAh^{\mathcal A}hA is supermodular.

Significance

The result. Proposition 1 justifies treating each store's expected sales as a concave, complementary function of the size profile. The paper's allocation MIP is built on that structure. It approximates ggg by a lower envelope of tangents and embeds the approximation in a mixed integer program that allocates warehouse stock across stores. The proposition explains why sending a unit of a major size can raise the value of the minor sizes already in a store. It also explains why the marginal value of a size falls as more of it is shipped. Without these properties a greedy or envelope-based allocation would have no structural footing.

Formalizing it. The result is proved on paper, with one error in the printed proof. The last line of the Appendix display gives ΔshA(q)\Delta_s h^{\mathcal A}(q)Δs​hA(q) as a probability, but the quantity is a time. The correct value carries a factor 1/λs1/\lambda_s1/λs​ and has qs+1q_s + 1qs​+1 in place of qsq_sqs​. This mission states the corrected identity. As far as is known, no machine-checked proof of any part of the result exists. A complete development would give a verified link from a continuous-time stochastic model to the discrete convexity properties that inventory optimization relies on. It would also produce reusable facts about Poisson processes and their hitting times along the way.

Difficulty

The monotonicity of ggg is pathwise and elementary. The other properties are not pathwise properties of GGG. Under the major-size policy, a realisation of GGG need not have decreasing increments in qsq_sqs​. The concavity and the complementarity appear only after taking expectations and rewriting ggg through the stopping times, using identity (2). That identity is an optional-sampling statement for the compensated Poisson processes at the bounded random times τA∧T\tau_{\mathcal A} \wedge TτA​∧T. These times depend on several independent processes at once, so the relevant filtration is the joint one. Mathlib has optional sampling only in discrete time.

The marginal-difference formula needs three ingredients:

  • the tail-integral representation of E[τA∧T]\mathbb E[\tau_{\mathcal A} \wedge T]E[τA​∧T];
  • the factorisation of P(τA>t)\mathbb P(\tau_{\mathcal A} > t)P(τA​>t) through independence across sizes;
  • the Erlang law of the hitting times τs(k)\tau_s(k)τs​(k).

The first idea, proving concavity of ggg directly from (1) path by path, fails for the reason above.

Formalization scope

All objects live in the namespace FastFashion.Structure:

  • Sizes: a Fintype with decidable equality. The major sizes are a Finset Sp and the minor sizes are its complement. No nonemptiness assumption is placed on the sizes or on S+\mathcal S^+S+; S+=∅\mathcal S^+ = \emptysetS+=∅ is allowed, as in the paper.
  • Inventories: vectors q:S→Nq : \mathcal S \to \mathbb Nq:S→N with the componentwise order and lattice operations. ese_ses​ is Pi.single s 1.
  • Stockout times: τs(qs)\tau_s(q_s)τs​(qs​) takes values in WithTop ℝ. It is +∞+\infty+∞ on paths that never reach qsq_sqs​, so it is never Lean's junk value 000, and τ∅=+∞\tau_\emptyset = +\inftyτ∅​=+∞. The truncation τA∧T\tau_{\mathcal A} \wedge TτA​∧T (stopMin) is a real number in [0,T][0, T][0,T].
  • Expectations and probabilities: expectations are Bochner integrals. They need no integrability hypothesis, because 0≤G(q)≤∑sNs(T)0 \le G(q) \le \sum_s N_s(T)0≤G(q)≤∑s​Ns​(T) and 0≤τA∧T≤T0 \le \tau_{\mathcal A} \wedge T \le T0≤τA​∧T≤T. Probabilities are real numbers.
  • Supermodularity: Topkis's lattice definition, the published platform definition Supermodularity.Monotonicity.SupermodularOn, applied on the whole lattice.
  • Poisson family: the processes are independent as whole processes. The σ-algebras generated by all times are independent, not only the one-dimensional marginals.

Two formalizations would trivialize the mission, and neither is used. ggg is defined as the expectation of GGG from (1), never by formula (2), and hAh^{\mathcal A}hA is never defined by a product-of-probabilities integral. Proposition 1 states both readings of "supermodular": the per-coordinate marginal-difference inequalities and the lattice inequality.

Deviations from the printed text:

  • The last line of the Appendix display is replaced by 1λsP(τs(qs+1)≤τA∖{s}∧T)\frac{1}{\lambda_s}\mathbb P(\tau_s(q_s+1) \le \tau_{\mathcal A\setminus\{s\}} \wedge T)λs​1​P(τs​(qs​+1)≤τA∖{s}​∧T).
  • "non-decreasing in xsx_sxs​" in Proposition 1 is read as qsq_sqs​.

Infrastructure a complete development needs:

  • hitting times of counting processes and their Erlang laws;
  • a continuous-time optional sampling theorem for the compensated Poisson process, or a direct computation from independent increments;
  • the tail formula E[X]=∫0∞P(X>t) dt\mathbb E[X] = \int_0^\infty \mathbb P(X > t)\,dtE[X]=∫0∞​P(X>t)dt for bounded non-negative XXX;
  • lattice facts about minima of single-variable increasing functions.

The Poisson-process and hitting-time lemmas are reusable well beyond this mission and are welcome as separate contributions. The companion mission on the paper's tangent-envelope approximation uses the same model layer.

Selected references

  • F. Caro, J. Gallien, Inventory Management of a Fast-Fashion Retail Network, working paper, August 2, 2007; published in Operations Research 58(2):257–273, 2010. https://doi.org/10.1287/opre.1090.0698
  • D. M. Topkis, Supermodularity and Complementarity, Princeton University Press, 1998. https://doi.org/10.1515/9781400822539
  • Y. Lu, J.-S. Song, Order-Based Cost Optimization in Assemble-to-Order Systems, Operations Research 53(1):151–169, 2005 (working paper 2003). https://doi.org/10.1287/opre.1040.0146
  • I. Karatzas, S. E. Shreve, Brownian Motion and Stochastic Calculus, 2nd ed., Springer, 1991. https://doi.org/10.1007/978-1-4612-0949-2
9 thms1 active userReviewed
Dynamic ProgrammingOptimization·Captain: mikedeng1

Algorithms for Scheduling Runway Operations Under Constrained Position Shifting 2: The Minimum Total Delay Is a Shortest Source-Sink Path in the CPS NetworkResearch Paper

Motivation

Runway capacity is the binding constraint at many busy airports. An aircraft landing or taking off behind another must wait a minimum separation time that depends on the two aircraft's weight classes, because of wake vortices; the US Federal Aviation Administration publishes these spacing requirements. Reordering aircraft changes the sum of the separations incurred and hence the delays, but unrestricted reordering is unacceptable to airlines and controllers: an aircraft that arrived early could be pushed to the end of the queue. Constrained position shifting (CPS), introduced by Psaraftis (1980) and used in the United States and Europe since, allows each aircraft to move at most kkk positions from its first-come-first-served (FCFS) position.

Balakrishnan and Chandran (Operations Research 58(6), 2010) recast CPS scheduling as dynamic programming on a layered network, the CPS network, whose size is polynomial in the number of aircraft nnn for fixed kkk. Their §4 minimizes the makespan (the landing time of the last aircraft). §5.2 minimizes the total delay, equivalently the average delay, which measures passenger and airline cost more directly. This mission formalizes §5.2.

Setting

There are nnn aircraft labelled 1,…,n1,\dots,n1,…,n in FCFS order, a maximum position shift k∈Nk\in\mathbb Nk∈N, separations δab≥0\delta_{ab}\ge 0δab​≥0 (the minimum time between leading aircraft aaa and trailing aircraft bbb), and a finite set of precedence pairs (x,y)(x,y)(x,y) with x≠yx\ne yx=y, meaning that xxx must land before yyy. The separations satisfy the triangle inequality δac≤δab+δbc\delta_{ac}\le\delta_{ab}+\delta_{bc}δac​≤δab​+δbc​, as wake-vortex separations do for arrivals-only or departures-only operations.

A kkk-CPS sequence is a permutation σ\sigmaσ of the aircraft (position ppp receives aircraft σ(p)\sigma(p)σ(p)) with ∣σ(p)−p∣≤k|\sigma(p)-p|\le k∣σ(p)−p∣≤k for every ppp. A feasible schedule is a kkk-CPS sequence that places xxx before yyy for each precedence pair, together with landing times t1,…,tnt_1,\dots,t_nt1​,…,tn​ (by position) such that

tp≥0,tq−tp≥δσ(p)σ(q)(p<q).t_p\ge 0,\qquad t_q-t_p\ge\delta_{\sigma(p)\sigma(q)}\quad(p<q).tp​≥0,tq​−tp​≥δσ(p)σ(q)​(p<q).

There are no time windows. Its total delay is t1+⋯+tnt_1+\dots+t_nt1​+⋯+tn​, measured from time 000, at which all aircraft are available.

The CPS network has stages 1,…,n1,\dots,n1,…,n. A node of stage ppp is a sequence of min⁡{2k+1,p}\min\{2k+1,p\}min{2k+1,p} distinct aircraft that can occupy positions p−min⁡{2k+1,p}+1,…,pp-\min\{2k+1,p\}+1,\dots,pp−min{2k+1,p}+1,…,p (each within kkk of its position). Its last aircraft is its final aircraft. An arc joins a stage-ppp node iii to a stage-(p+1)(p+1)(p+1) node jjj when the sequences overlap: the first min⁡{2k,p}\min\{2k,p\}min{2k,p} aircraft of jjj are the last min⁡{2k,p}\min\{2k,p\}min{2k,p} aircraft of iii. A source-sink path picks one node per stage, and its sequence assigns position ppp to the final aircraft of the stage-ppp node. Precedence is incorporated by deleting every node that places yyy before xxx for some pair (x,y)(x,y)(x,y), or that places an aircraft outside the positions a pair allows it (§3.1). The result is the pruned network GGG. For nodes i,ji,ji,j, δi(j)\delta_i(j)δi​(j) is the separation between their final aircraft.

For a node jjj in stage ppp and times t1,…,tpt_1,\dots,t_pt1​,…,tp​ along a path to jjj, the paper uses the partial objective

θj(p)=t1+⋯+tp−1+(n−p+1) tp.\theta_j(p)=t_1+\dots+t_{p-1}+(n-p+1)\,t_p .θj​(p)=t1​+⋯+tp−1​+(n−p+1)tp​.

Formalization targets

Goal: the shortest-path equivalence (§5.2, p. 1656)

Give the arc (i,j)(i,j)(i,j) from stage p−1p-1p−1 to stage ppp the distance (n−p+1)δi(j)(n-p+1)\delta_i(j)(n−p+1)δi​(j), and source and sink arcs distance 000. Then a feasible schedule exists if and only if GGG has a source-sink path, and

min⁡(σ,t) feasible ∑p=1ntp  =  min⁡v source-sink path of G ∑p=2n(n−p+1) δvp−1(vp),\min_{(\sigma,t)\ \text{feasible}}\ \sum_{p=1}^{n}t_p\;=\;\min_{v\ \text{source-sink path of }G}\ \sum_{p=2}^{n}(n-p+1)\,\delta_{v_{p-1}}(v_p),(σ,t) feasiblemin​ p=1∑n​tp​=v source-sink path of Gmin​ p=2∑n​(n−p+1)δvp−1​​(vp​),

where each minimum exists exactly when the other does.

Milestones

  1. Theorem 1 with Lemmas 1–2 (pp. 1652–1654): the source-sink paths of GGG are exactly the kkk-CPS sequences respecting the precedence pairs.
  2. Proposition 2 (p. 1656): in a schedule of minimum total delay, tp−tp−1=δσ(p−1)σ(p)t_p-t_{p-1}=\delta_{\sigma(p-1)\sigma(p)}tp​−tp−1​=δσ(p−1)σ(p)​ for every p≥2p\ge 2p≥2.
  3. Proposition 3 (p. 1656): all paths from the source to a node lying on some source-sink path contain the same set of aircraft.
  4. The recursion of §5.2 (p. 1656): with θj∗(p)\theta^*_j(p)θj∗​(p) the minimum of θj(p)\theta_j(p)θj​(p) over partial schedules ending at jjj, θj∗(1)=0\theta^*_j(1)=0θj∗​(1)=0 and
θj∗(p)=min⁡i∈P(j)(θi∗(p−1)+(n−p+1) δi(j)),\theta^*_j(p)=\min_{i\in P(j)}\big(\theta^*_i(p-1)+(n-p+1)\,\delta_i(j)\big),θj∗​(p)=i∈P(j)min​(θi∗​(p−1)+(n−p+1)δi​(j)),

where P(j)P(j)P(j) is the set of predecessors of jjj in GGG.

Significance

The result turns a sequencing problem over up to n!n!n! orders into a shortest-path computation on a directed acyclic graph with O(n(2k+1)2k+1)O(n(2k+1)^{2k+1})O(n(2k+1)2k+1) nodes, and it does so while respecting precedence constraints, which earlier CPS algorithms (Psaraftis 1980; Trivizas 1998) could not handle. It is the average-delay counterpart of the makespan algorithm of §4 and shares its network. The same network carries the discrete-time models of §6.

The paper's proof of the equivalence is one sentence ("follows from properties of shortest paths and is omitted here"). Propositions 2 and 3 are proved in a few lines, and Theorem 1 by an argument about windows of 2k+12k+12k+1 consecutive positions. None of these results has a machine-checked proof. A formal development makes explicit the conventions the prose leaves implicit (the time origin, pairwise separations, which nodes Proposition 3 refers to) and checks that the recursion and the shortest-path value really compute the minimum total delay.

Difficulty

The arc distances weight the separation into stage ppp by n−p+1n-p+1n−p+1, the number of aircraft whose landing it delays. The equivalence needs two directions. A path's length must be the total delay of the schedule that lands each aircraft at its predecessor's time plus the separation, and that schedule must satisfy every pairwise separation, not only the consecutive ones; this is where the triangle inequality enters. Conversely, every feasible schedule must have total delay at least the length of its sequence's path, and every feasible sequence must be a path of GGG. The last point is Theorem 1 with precedence pruning: a node sees only 2k+12k+12k+1 consecutive positions, so detecting a violated precedence pair inside one node requires the position filter of §3.1 for pairs that reverse FCFS order. Proposition 3 is false for nodes that lie on no source-sink path, which the printed statement does not exclude.

Formalization scope

Aircraft and positions are Fin n, 0-based, and stages are 0-based natural numbers, so the arc into stage sss has distance (n−s)δi(j)(n-s)\delta_i(j)(n−s)δi​(j). Separations are real, nonnegative and satisfy the triangle inequality; feasibility requires them between every pair of positions. Landing times are real and nonnegative. The paper fixes no time origin, but its path length equals the total delay exactly when delays are measured from 000, and without a lower bound on the times the objective has no minimum. Minima are stated with IsLeast on sets of values, never with a real infimum. Feasibility and the existence of a path are stated separately, because the equality of least elements alone does not force them to exist.

The following loose phrases are given explicit readings. "In an optimal solution" (Proposition 2) means a feasible schedule whose total delay is at most that of every feasible schedule. "The CPS network" in Proposition 3 means nodes lying on a source-sink path, stated for the network without precedence pruning, which also covers the pruned one. "θ∗\theta^*θ∗" means the least element of the set of values of θj(p)\theta_j(p)θj​(p) over partial schedules ending at jjj, and the printed θj∗(s)\theta^*_j(s)θj∗​(s) is read as θj∗(p)\theta^*_j(p)θj∗​(p). The boundary value θj∗(1)=0\theta^*_j(1)=0θj∗​(1)=0 is not printed and is stated explicitly. Precedence pruning uses the printed Case II rule. Running times (the O(n(2k+1)2k+2)O(n(2k+1)^{2k+2})O(n(2k+1)2k+2) bound) are not formalized.

A trivializing formalization would require separations only between consecutive aircraft, define the minimum by the recursion itself, or use a network whose nodes need not be distinct or admissible. Here the minimum is defined directly over feasible schedules, separations are pairwise, and the nodes are checked against the paper's Figure 1 (stage sizes 2, 4, 7, 14, 14, 7 and 13 paths for n=6n=6n=6, k=1k=1k=1).

The development needs finite sums, permutations of Fin n and lists. The network layer (stages, arcs, precedence pruning, Theorem 1) is shared with the makespan and discrete-time missions of this series and reusable for any CPS objective. Proofs of any milestone, alternative proofs of the goal, and generalizations (asymmetric shifts, §3.2) are welcome.

Selected references

  • H. Balakrishnan, B. G. Chandran, Algorithms for Scheduling Runway Operations Under Constrained Position Shifting, Operations Research 58(6), 1650–1665, 2010. https://doi.org/10.1287/opre.1100.0869
  • H. N. Psaraftis, A Dynamic Programming Approach for Sequencing Groups of Identical Jobs, Operations Research 28(6), 1347–1359, 1980. https://doi.org/10.1287/opre.28.6.1347
  • R. G. Dear, Y. S. Sherif, An Algorithm for Computer Assisted Sequencing and Scheduling of Terminal Area Operations, Transportation Research Part A 25(2–3), 129–139, 1991.
  • D. A. Trivizas, Optimal Scheduling with Maximum Position Shift (MPS) Constraints: A Runway Scheduling Application, Journal of Navigation 51(2), 250–266, 1998.
  • R. de Neufville, A. Odoni, Airport Systems: Planning, Design, and Management, McGraw-Hill, 2003.
6 thms1 active userReviewed
OptimizationProbability·Captain: mikedeng1

Reliable Facility Location Design Under the Risk of Disruptions 1: With R = J the Compact Level-Assignment Formulation Equals the Scenario-Based Stochastic ProgramResearch Paper

Motivation

Facility location models decide where to open depots, warehouses or service centres and which customers each one serves. Classical models assume that an open facility always works. In practice facilities are disrupted by weather, strikes, power loss or supplier failure, and a network designed for normal operation can become very expensive when its nearest facilities go down. Reliable facility location models place facilities so that the sum of the fixed cost and the expected transportation cost, taken over random facility failures, is minimal.

Cui, Ouyang and Shen (UCTC-FR-2010-02, 2010; published as Oper. Res. 58(4):998–1011) study the case in which each site fails independently with its own probability. The most direct model of this situation is a scenario-based stochastic program: list every pattern of failures, give it its probability, and let each customer be served optimally in each pattern. That program has 2J2^J2J scenarios for JJJ candidate sites and cannot be written down for realistic JJJ. The paper proposes instead a compact formulation in which each customer receives an ordered list of backup facilities ("levels") and the probability that each level is the one that serves her is carried by a small set of recursive variables. This mission is about the claim that the compact model loses nothing.

The level-assignment technique comes from Snyder and Daskin (Transp. Sci. 39(3):400–416, 2005), who treated the case of one common failure probability for all sites. The uniform-probability model of Snyder and Shen's textbook is on the platform as SupplyChainTheory_disruptions; it is a different model (one qqq, no cap on the number of levels) and is not reused here.

Setting

There are III customers i=0,…,I−1i = 0,\dots,I-1i=0,…,I−1 with demand rates λi≥0\lambda_i \ge 0λi​≥0 and JJJ candidate sites j=0,…,J−1j = 0,\dots,J-1j=0,…,J−1 with fixed costs fjf_jfj​ and failure probabilities 0≤qj<10 \le q_j < 10≤qj​<1. Failures are independent. Shipping one unit from site jjj to customer iii costs dijd_{ij}dij​, and each unit of missed demand of customer iii costs a penalty ϕi\phi_iϕi​. An emergency facility with index JJJ never fails (qJ=0q_J = 0qJ​=0), costs nothing to open (fJ=0f_J = 0fJ​=0), and "serves" at the penalty cost diJ=ϕid_{iJ} = \phi_idiJ​=ϕi​.

The compact model (RUFL). Each customer is assigned to facilities at levels r=0,…,Rr = 0,\dots,Rr=0,…,R, with R≥1R \ge 1R≥1; a level-rrr facility serves her exactly when all her facilities at levels 0,…,r−10,\dots,r-10,…,r−1 have failed. The variables are Xj∈{0,1}X_j\in\{0,1\}Xj​∈{0,1} (site jjj open), Yijr∈{0,1}Y_{ijr}\in\{0,1\}Yijr​∈{0,1} (facility jjj is customer iii's level-rrr facility) and PijrP_{ijr}Pijr​, the probability that jjj serves iii at level rrr. Constraints (1b)–(1d) say that each level holds one facility until the emergency facility appears, which happens exactly once, and that only open sites are used, each at most once. Constraints (1e)–(1f) fix PPP: Pij0=1−qjP_{ij0} = 1-q_jPij0​=1−qj​ and Pijr=(1−qj)∑k<Jqk1−qkPi,k,r−1Yi,k,r−1P_{ijr} = (1-q_j)\sum_{k<J}\frac{q_k}{1-q_k}P_{i,k,r-1}Y_{i,k,r-1}Pijr​=(1−qj​)∑k<J​1−qk​qk​​Pi,k,r−1​Yi,k,r−1​. The objective is

Φ(X,Y,P)=∑j=0J−1fjXj+∑i=0I−1∑j=0J∑r=0RλidijPijrYijr.\Phi(X,Y,P) = \sum_{j=0}^{J-1} f_jX_j + \sum_{i=0}^{I-1}\sum_{j=0}^{J}\sum_{r=0}^{R}\lambda_i d_{ij}P_{ijr}Y_{ijr}.Φ(X,Y,P)=j=0∑J−1​fj​Xj​+i=0∑I−1​j=0∑J​r=0∑R​λi​dij​Pijr​Yijr​.

The scenario model (SSP). A scenario is ω∈Ω={0,1}J\omega\in\Omega=\{0,1\}^Jω∈Ω={0,1}J, with δjω=1\delta_{j\omega}=1δjω​=1 if site jjj operates in ω\omegaω (and δJω=1\delta_{J\omega}=1δJω​=1 always). Its probability is pω=∏j<J(1−qj)δjωqj1−δjωp_\omega = \prod_{j<J}(1-q_j)^{\delta_{j\omega}}q_j^{1-\delta_{j\omega}}pω​=∏j<J​(1−qj​)δjω​qj1−δjω​​. With Yijω∈{0,1}Y_{ij\omega}\in\{0,1\}Yijω​∈{0,1} (customer iii served by jjj in ω\omegaω), (SSP) minimizes

Ψ(X,Y)=∑j=0J−1fjXj+∑i=0I−1∑j=0J∑ω∈ΩλidijpωYijω\Psi(X,Y) = \sum_{j=0}^{J-1} f_jX_j + \sum_{i=0}^{I-1}\sum_{j=0}^{J}\sum_{\omega\in\Omega}\lambda_i d_{ij}p_\omega Y_{ij\omega}Ψ(X,Y)=j=0∑J−1​fj​Xj​+i=0∑I−1​j=0∑J​ω∈Ω∑​λi​dij​pω​Yijω​

subject to ∑j=0JYijω=1\sum_{j=0}^{J}Y_{ij\omega}=1∑j=0J​Yijω​=1 and Yijω≤δjωXjY_{ij\omega}\le\delta_{j\omega}X_jYijω​≤δjω​Xj​ for regular jjj.

Formalization targets

Goal: Proposition 1 (p. 10)

If R=JR = JR=J, both programs have optimal solutions, and for every optimal (X,Y,P)(X,Y,P)(X,Y,P) of (RUFL) and every optimal (X′,Y′)(X',Y')(X′,Y′) of (SSP),

Φ(X,Y,P)=Ψ(X′,Y′).\Phi(X,Y,P) = \Psi(X',Y').Φ(X,Y,P)=Ψ(X′,Y′).

The paper states this as "formulation (1a)–(1g) is equivalent to the stochastic programming formulation that covers all failure scenarios"; its proof shows equality of optimal values, and that is the reading formalized.

Milestones (Appendix A.1, pp. 32–34)

  1. Scenario-probability identity. For distinct regular sites j(0),…,j(r−1)j(0),\dots,j(r-1)j(0),…,j(r−1) and a facility j(r)j(r)j(r) not among them, the scenarios in which j(0),…,j(r−1)j(0),\dots,j(r-1)j(0),…,j(r−1) fail and j(r)j(r)j(r) operates have total probability
∑ω∈Ω(i,r)pω=(1−qj(r))∏ℓ=0r−1qj(ℓ).\sum_{\omega\in\Omega(i,r)}p_\omega = (1-q_{j(r)})\prod_{\ell=0}^{r-1}q_{j(\ell)}.ω∈Ω(i,r)∑​pω​=(1−qj(r)​)ℓ=0∏r−1​qj(ℓ)​.
  1. (RUFL) → (SSP). Every (RUFL)-feasible (X,Y,P)(X,Y,P)(X,Y,P) has an (SSP)-feasible (X,Y′)(X,Y')(X,Y′) with Ψ(X,Y′)=Φ(X,Y,P)\Psi(X,Y') = \Phi(X,Y,P)Ψ(X,Y′)=Φ(X,Y,P), for any RRR.
  2. Normalization. Some optimal solution of (SSP), with the same XXX, serves every customer in every scenario by her closest operating open facility, ties to the lowest index.
  3. (SSP) → (RUFL). If R=JR = JR=J, every normalized (SSP)-feasible (X,Y)(X,Y)(X,Y) has a (RUFL)-feasible (X,Y′,P′)(X,Y',P')(X,Y′,P′) with Φ(X,Y′,P′)=Ψ(X,Y)\Phi(X,Y',P') = \Psi(X,Y)Φ(X,Y′,P′)=Ψ(X,Y).

Significance

The proposition says that, with enough levels, the compact model is exact: its optimal value is the true minimum expected cost over all failure patterns. The compact model has O(IJR)O(IJR)O(IJR) variables and constraints against O(IJ2J)O(IJ2^J)O(IJ2J) for (SSP), and after the standard linearization of the products PijrYijrP_{ijr}Y_{ijr}Pijr​Yijr​ it is an ordinary mixed-integer program to which the paper's Lagrangian relaxation applies. For R<JR<JR<J the paper notes the compact model is in general not equivalent; Proposition 1 is the anchor that says what the parameter RRR trades away.

The result is proved in the paper; as far as we know it has not been formalized. The work here is to formalize that proof: the product-measure computation behind milestone 1, the two explicit solution maps, and the exchange argument behind the normalization. The probability identity for "first success in an ordered list of independent trials" and the scenario-to-level bookkeeping are reusable for other reliability models built on Snyder–Daskin levels.

Difficulty

The cost identities are routine once the right bookkeeping is in place; the obstacle is that bookkeeping. In direction (RUFL) → (SSP) one must show that the recursion (1e)–(1f), which multiplies by qk/(1−qk)q_k/(1-q_k)qk​/(1−qk​) and sums over all regular kkk, collapses to the closed product (1−qj(r))∏ℓ<rqj(ℓ)(1-q_{j(r)})\prod_{\ell<r}q_{j(\ell)}(1−qj(r)​)∏ℓ<r​qj(ℓ)​ along a customer's assigned list, using that (1b)–(1d) put exactly one facility on each level until the emergency facility and that (1c) forbids repeats. Then the sum over the 2J2^J2J scenarios must be regrouped by the first operating facility on that list, which needs milestone 1 on a product space.

The converse direction is where R=JR=JR=J enters, and the naive attempt fails: an arbitrary optimal (SSP) solution may serve a customer by different, equally good facilities in different scenarios in ways that no single ordered list reproduces. The normalization step removes this freedom; only after it can one read off a single list per customer, and R=JR=JR=J guarantees that this list (all open sites no farther than the penalty, then JJJ) fits on the available levels.

Formalization scope

Indices are 0-based; facilities are Fin (J+1) with Fin.last J the emergency facility, and levels are Fin (R+1). The extended data diJ=ϕid_{iJ}=\phi_idiJ​=ϕi​, qJ=0q_J=0qJ​=0, and the conventions δJω=1\delta_{J\omega}=1δJω​=1, XJ=1X_J=1XJ​=1 are definitions, not hypotheses. Scenarios are Fin J → Bool and pωp_\omegapω​ is the explicit finite product; no measure theory is involved, and (17a) is a finite sum. Binary variables are real numbers constrained to {0,1}\{0,1\}{0,1}. R=JR = JR=J is expressed by taking the instance type Instance I J J; with the standing R≥1R\ge 1R≥1 this means J≥1J\ge1J≥1. Demand rates are nonnegative; no sign is assumed on ddd, ϕ\phiϕ or fff.

Two printed constraints are corrected, and the formalization commits to the corrections. The first sum in (1b) is printed over the regular sites j≤J−1j\le J-1j≤J−1; it runs here over all J+1J+1J+1 facilities, as the paper's own reading of (1b) requires. Constraint (17c) is printed as ∑iYijω≤δjωXj\sum_i Y_{ij\omega}\le\delta_{j\omega}X_j∑i​Yijω​≤δjω​Xj​, which gives each site a capacity of one customer per scenario and makes Proposition 1 false; it is stated here per customer, Yijω≤δjωXjY_{ij\omega}\le\delta_{j\omega}X_jYijω​≤δjω​Xj​.

The goal asserts existence of optimal solutions of both programs before comparing their values, so it cannot hold vacuously. Contributions welcome: proofs of the milestones, a general lemma on the probability that the first success among independent Bernoulli trials occurs at a given position, and an existence-of-optimum lemma for finite binary programs.

Selected references

  • X. Cui, Y. Ouyang, Z.-J. M. Shen, Reliable Facility Location Design under the Risk of Disruptions, UCTC-FR-2010-02, University of California Transportation Center, 2010; Operations Research 58(4):998–1011, 2010. https://doi.org/10.1287/opre.1090.0801
  • L. V. Snyder, M. S. Daskin, Reliability models for facility location: the expected failure cost case, Transportation Science 39(3):400–416, 2005. https://doi.org/10.1287/trsc.1040.0107
  • H. D. Sherali, A. Alameddine, A new reformulation-linearization technique for bilinear programming problems, Journal of Global Optimization 2(4):379–410, 1992. https://doi.org/10.1007/BF00122429
8 thms1 active userReviewed
Optimization·Captain: mikedeng1

Reliable Facility Location Design Under the Risk of Disruptions 2: Optimal Solutions Order Each Customer's Assignment Levels by DistanceResearch Paper

Why backup assignments matter

A facility location plan decides which sites to open and which customers each site serves. Real facilities fail: plants are shut by strikes, warehouses by floods, distribution centres by power outages. When a customer's facility is down, she is served by a backup facility, or not served at all at a penalty. Reliable facility location chooses sites and backup assignments together, minimizing fixed costs plus the expected transportation and penalty cost over the random failures.

Snyder and Daskin (2005) introduced the level-assignment formulation in which every customer receives a ranked list of facilities, and they assumed all sites fail with the same probability. Cui, Ouyang and Shen (2010) allow each site its own failure probability qjq_jqj​, which is the realistic case: a site in a flood plain and a site inland do not fail equally often. Their compact mixed-integer program (RUFL) is the subject of this mission.

Timeline.

  • Snyder and Daskin (2005), Transportation Science 39(3): level-assignment formulation of the reliability P-median and UFL problems, equal failure probability qqq; consecutive assignments are ordered by distance in an optimal solution.
  • Snyder and Shen, Fundamentals of Supply Chain Theory, Theorem 9.10 (textbook treatment): the same ordering for the reliable fixed-charge location problem with 0<q<10<q<10<q<1 and positive demands.
  • Cui, Ouyang and Shen (2010), UCTC-FR-2010-02 / Operations Research 58(4): site-dependent failure probabilities qjq_jqj​, a cap RRR on the number of levels, and Proposition 2, the ordering result for this model.

The model (RUFL)

There are customers i=0,…,I−1i = 0,\dots,I-1i=0,…,I−1 with demand rates λi\lambda_iλi​, and candidate sites j=0,…,J−1j = 0,\dots,J-1j=0,…,J−1 with fixed costs fjf_jfj​ and failure probabilities 0≤qj<10 \le q_j < 10≤qj​<1; failures are independent. Shipping one unit from jjj to iii costs dijd_{ij}dij​, and each unserved unit of customer iii costs a penalty ϕi\phi_iϕi​. An emergency facility with index JJJ represents non-service: fJ=0f_J = 0fJ​=0, qJ=0q_J = 0qJ​=0 and diJ=ϕid_{iJ} = \phi_idiJ​=ϕi​.

Each customer is assigned at levels r=0,…,Rr = 0,\dots,Rr=0,…,R, R≥1R\ge 1R≥1. Her level-rrr facility serves her exactly when all her facilities at levels 0,…,r−10,\dots,r-10,…,r−1 have failed. The variables are Xj∈{0,1}X_j\in\{0,1\}Xj​∈{0,1} (site jjj open), Yijr∈{0,1}Y_{ijr}\in\{0,1\}Yijr​∈{0,1} (facility jjj is customer iii's level-rrr facility) and PijrP_{ijr}Pijr​, the probability that jjj serves iii at level rrr. (RUFL) minimizes

Φ(X,Y,P)=∑j=0J−1fjXj+∑i=0I−1∑j=0J∑r=0RλidijPijrYijr\Phi(X,Y,P) = \sum_{j=0}^{J-1} f_jX_j + \sum_{i=0}^{I-1}\sum_{j=0}^{J}\sum_{r=0}^{R}\lambda_i d_{ij}P_{ijr}Y_{ijr}Φ(X,Y,P)=j=0∑J−1​fj​Xj​+i=0∑I−1​j=0∑J​r=0∑R​λi​dij​Pijr​Yijr​

subject to: at every level a customer has exactly one facility unless she has already reached the emergency facility (1b); only open sites are used, each at most once (1c); the emergency facility is assigned exactly once (1d); and the probabilities follow the transition equations

Pij0=1−qj,Pijr=(1−qj)∑k=0J−1qk1−qkPi,k,r−1Yi,k,r−1(1≤r≤R).P_{ij0} = 1-q_j,\qquad P_{ijr} = (1-q_j)\sum_{k=0}^{J-1}\frac{q_k}{1-q_k}P_{i,k,r-1}Y_{i,k,r-1}\quad (1\le r\le R).Pij0​=1−qj​,Pijr​=(1−qj​)k=0∑J−1​1−qk​qk​​Pi,k,r−1​Yi,k,r−1​(1≤r≤R).

So if a customer's levels are j0,j1,…j_0, j_1, \dotsj0​,j1​,…, then Pijrr=(1−qjr) qj0⋯qjr−1P_{i j_r r} = (1-q_{j_r})\,q_{j_0}\cdots q_{j_{r-1}}Pijr​r​=(1−qjr​​)qj0​​⋯qjr−1​​.

Formalization targets

Goal: Proposition 2

Assume λi>0\lambda_i > 0λi​>0 for every customer and qk>0q_k > 0qk​>0 for every regular site. In every optimal solution (X,Y,P)(X,Y,P)(X,Y,P) of (RUFL),

Yijr=1, Yik,r+1=1 ⟹ dij≤dik(0≤r, r+1≤R, 0≤j,k≤J),Y_{ijr} = 1,\ Y_{ik,r+1} = 1 \ \Longrightarrow\ d_{ij}\le d_{ik}\qquad (0\le r,\ r+1\le R,\ 0\le j,k\le J),Yijr​=1, Yik,r+1​=1 ⟹ dij​≤dik​(0≤r, r+1≤R, 0≤j,k≤J),

where diJ=ϕid_{iJ} = \phi_idiJ​=ϕi​.

Milestones (the proof of Appendix A.2)

  1. Swap identity. If (X,Y,P)(X,Y,P)(X,Y,P) is feasible with Yijr=Yik,r+1=1Y_{ijr} = Y_{ik,r+1} = 1Yijr​=Yik,r+1​=1 for regular j,kj,kj,k, exchanging jjj and kkk and recomputing PPP gives a feasible solution whose cost changes by
λi(1−qk)(dik−dij)Pijr.\lambda_i(1-q_k)(d_{ik}-d_{ij})P_{ijr}.λi​(1−qk​)(dik​−dij​)Pijr​.
  1. Emergency case. If instead k=Jk = Jk=J, moving JJJ up to level rrr and dropping jjj gives a feasible solution whose cost changes by λiPijr(ϕi−dij)\lambda_i P_{ijr}(\phi_i - d_{ij})λi​Pijr​(ϕi​−dij​).
  2. Positivity. If every regular qk>0q_k>0qk​>0, an assigned facility has Pijr>0P_{ijr}>0Pijr​>0.

Significance

Proposition 2 says that for a given set of facilities assigned to a customer, the optimal order of the levels depends only on the distances, not on the failure probabilities. A solution method may therefore sort each customer's assigned facilities by distance instead of searching over orders. The paper's Lagrangian subproblem (RSPi_ii​) uses exactly this: "following a similar argument to Proposition 2" (§3.3.1, p. 13), its objective depends only on the set of facilities chosen, which is what makes the set function of Proposition 3 (a separate mission of this series) well defined. The proposition does not say that the RRR nearest open facilities are the right ones: the paper's Example 1 (p. 11) shows a farther but more reliable facility can be optimal.

The result is proved in the paper; no machine-checked proof exists. The equal-probability special case is published on Prove2Me as SupplyChainTheory.rflp_ordered_assignments (Snyder–Shen Theorem 9.10), in a different model (one qqq, no level cap). This mission adds the site-dependent case with a level cap, and it makes explicit two hypotheses the printed statement omits.

Difficulty

The exchange argument is short on paper, but its bookkeeping is exactly where a formal proof has to work. The paper's modified P′P'P′ keeps the old values at unlisted indices, and that array does not satisfy the transition equations, so feasibility of the swapped solution has to be shown with the recomputed probabilities, level by level, including the fact that levels after r+1r+1r+1 are unchanged. The emergency case is dismissed in one sentence; its cost change must be computed. Finally, "<0<0<0" needs λi>0\lambda_i>0λi​>0 and Pijr>0P_{ijr}>0Pijr​>0; the naive reading of the printed proposition, with qj=0q_j = 0qj​=0 allowed, is false, because a level-0 facility that never fails leaves all later levels with probability zero and in arbitrary order.

Formalization scope

Indices are 0-based. Facilities are Fin (J+1) with the emergency facility Fin.last J; levels are Fin (R+1); the extended data diJ=ϕid_{iJ}=\phi_idiJ​=ϕi​, qJ=0q_J=0qJ​=0 are definitions. XXX, YYY, PPP are real arrays, with (1g) requiring XXX, YYY to be 000 or 111. Constraint (1b) is printed with its first sum ending at J−1J-1J−1; it is formalized with the sum over all j≤Jj \le Jj≤J, the reading the paper's own explanation of (1b) gives. "Optimal" means feasible with objective at most that of every feasible solution; the feasible set is nonempty (send every customer to JJJ at level 0), so the goal is not vacuous. Proposition 2 carries the hypotheses λi>0\lambda_i>0λi​>0 and qk>0q_k>0qk​>0; without them it is false, and a statement that dropped them would be unprovable rather than trivial. The level rrr ranges over 0≤r≤R−10\le r\le R-10≤r≤R−1, the range in which level r+1r+1r+1 exists.

The development needs only finite sums and a recursion over levels; no measure theory is required. The recomputed probabilities transProb and the two moves swapY, emergencyUpY are reusable for any exchange argument on this formulation. Proofs of the milestones, of the goal, and of the equivalence of the transition equations with the closed product form are welcome.

Selected references

  • Cui, T., Ouyang, Y., Shen, Z.-J. M., Reliable Facility Location Design under the Risk of Disruptions, UCTC-FR-2010-02, University of California Transportation Center, 2010; published in Operations Research 58(4):998–1011, 2010. https://doi.org/10.1287/opre.1090.0801
  • Snyder, L. V., Daskin, M. S., Reliability Models for Facility Location: The Expected Failure Cost Case, Transportation Science 39(3):400–416, 2005. https://doi.org/10.1287/trsc.1040.0107
  • Snyder, L. V., Shen, Z.-J. M., Fundamentals of Supply Chain Theory, 2nd ed., Wiley, 2019, §9.6. https://doi.org/10.1002/9781119584445
6 thms1 active userReviewed
OptimizationProbability·Captain: mikedeng1

Quantitative Stability in Stochastic Programming: The Method of Probability Metrics 2: Linear Two-Stage Programs Are Lipschitz Stable in the Fortet–Mourier Metric ζ₂Research Paper

Why stability of two-stage programs matters

A linear two-stage stochastic program with fixed recourse chooses a here-and-now decision xxx before a random vector ξ\xiξ is observed, and pays a recourse cost afterwards. It is the basic model of stochastic programming, used in capacity planning, energy and supply-chain models. In practice the distribution μ\muμ of ξ\xiξ is never known exactly: it is estimated from data, replaced by a discrete scenario approximation, or perturbed for robustness. The question is how much the optimal value and the optimal decisions can move when μ\muμ is replaced by a nearby ν\nuν, and in which distance between probability measures "nearby" should be measured.

Rachev and Römisch (Math. Oper. Res. 27(4), 2002, doi:10.1287/moor.27.4.792.304; the mission cites the authors' preprint from edoc.hu-berlin.de) answer this with the method of probability metrics: they derive from the structure of the integrand a canonical distance on measures for each class of models. For linear two-stage programs the canonical distance is the Fortet–Mourier metric of order 2.

Timeline. Robinson and Wets (1987) studied qualitative stability of two-stage programs with respect to weak convergence of measures. Walkup and Wets (1969) had earlier established the piecewise-bilinear structure of the second-stage value function used here. Römisch and Schultz (1991) proved a quantitative stability result for two-stage programs with complete recourse in the Wasserstein metric W2W_2W2​. The present paper (published 2002) replaces complete recourse by relatively complete recourse plus dual feasibility, and W2W_2W2​ by the metric ζ2\zeta_2ζ2​, which it bounds by a multiple of W2W_2W2​ (p. 13).

Setting

Let X⊆RmX\subseteq\mathbb R^mX⊆Rm be a nonempty polyhedron and Ξ⊆Rs\Xi\subseteq\mathbb R^sΞ⊆Rs a polyhedron. Fix c∈Rmc\in\mathbb R^mc∈Rm, an (r,m‾)(r,\overline m)(r,m)-matrix WWW, and data q(ξ)∈Rm‾q(\xi)\in\mathbb R^{\overline m}q(ξ)∈Rm, h(ξ)∈Rrh(\xi)\in\mathbb R^rh(ξ)∈Rr and an (r,m)(r,m)(r,m)-matrix T(ξ)T(\xi)T(ξ) depending affine linearly on ξ\xiξ. The second-stage value is

Φ(u,t)=inf⁡{uy: Wy=t, y≥0},\Phi(u,t)=\inf\{uy:\ Wy=t,\ y\ge0\},Φ(u,t)=inf{uy: Wy=t, y≥0},

with pos⁡W={Wy:y≥0}\operatorname{pos}W=\{Wy:y\ge0\}posW={Wy:y≥0} and D={u:{z:W′z≤u}≠∅}D=\{u:\{z:W'z\le u\}\ne\emptyset\}D={u:{z:W′z≤u}=∅}. The integrand is f0(ξ,x)=cx+Φ(q(ξ),h(ξ)−T(ξ)x)f_0(\xi,x)=cx+\Phi(q(\xi),h(\xi)-T(\xi)x)f0​(ξ,x)=cx+Φ(q(ξ),h(ξ)−T(ξ)x) when h(ξ)−T(ξ)x∈pos⁡Wh(\xi)-T(\xi)x\in\operatorname{pos}Wh(ξ)−T(ξ)x∈posW and q(ξ)∈Dq(\xi)\in Dq(ξ)∈D, and +∞+\infty+∞ otherwise. For a Borel probability measure ν\nuν on Ξ\XiΞ the problem is

min⁡{∫Ξf0(ξ,x) ν(dξ):x∈X},\min\Big\{\int_\Xi f_0(\xi,x)\,\nu(d\xi):x\in X\Big\},min{∫Ξ​f0​(ξ,x)ν(dξ):x∈X},

with optimal value v(ν)v(\nu)v(ν) and solution set S(ν)S(\nu)S(ν). Two assumptions are made: (A1) h(ξ)−T(ξ)x∈pos⁡Wh(\xi)-T(\xi)x\in\operatorname{pos}Wh(ξ)−T(ξ)x∈posW and q(ξ)∈Dq(\xi)\in Dq(ξ)∈D for all (ξ,x)∈Ξ×X(\xi,x)\in\Xi\times X(ξ,x)∈Ξ×X; (A2) μ\muμ has a finite second moment.

The Fortet–Mourier metric ζ2\zeta_2ζ2​ on P2(Ξ)\mathcal P_2(\Xi)P2​(Ξ), the probability measures on Ξ\XiΞ with finite second moment, is

ζ2(μ,ν)=sup⁡f∈F2(Ξ)∣∫Ξf dμ−∫Ξf dν∣,F2(Ξ)={f:∣f(ξ)−f(ξ~)∣≤max⁡{1,∥ξ∥,∥ξ~∥}∥ξ−ξ~∥}.\zeta_2(\mu,\nu)=\sup_{f\in\mathcal F_2(\Xi)}\Big|\int_\Xi f\,d\mu-\int_\Xi f\,d\nu\Big|,\quad \mathcal F_2(\Xi)=\{f:|f(\xi)-f(\tilde\xi)|\le\max\{1,\|\xi\|,\|\tilde\xi\|\}\|\xi-\tilde\xi\|\}.ζ2​(μ,ν)=f∈F2​(Ξ)sup​​∫Ξ​fdμ−∫Ξ​fdν​,F2​(Ξ)={f:∣f(ξ)−f(ξ~​)∣≤max{1,∥ξ∥,∥ξ~​∥}∥ξ−ξ~​∥}.

For an open bounded U⊇S(μ)\mathcal U\supseteq S(\mu)U⊇S(μ), the growth function is ψ(τ)=inf⁡{∫Ξf0(ξ,x) μ(dξ)−v(μ):d(x,S(μ))≥τ, x∈X∩cl⁡U}\psi(\tau)=\inf\{\int_\Xi f_0(\xi,x)\,\mu(d\xi)-v(\mu): d(x,S(\mu))\ge\tau,\ x\in X\cap\operatorname{cl}\mathcal U\}ψ(τ)=inf{∫Ξ​f0​(ξ,x)μ(dξ)−v(μ):d(x,S(μ))≥τ, x∈X∩clU}, and Ψ(η)=η+ψ−1(2η)\Psi(\eta)=\eta+\psi^{-1}(2\eta)Ψ(η)=η+ψ−1(2η) with ψ−1(t)=sup⁡{τ≥0:ψ(τ)≤t}\psi^{-1}(t)=\sup\{\tau\ge0:\psi(\tau)\le t\}ψ−1(t)=sup{τ≥0:ψ(τ)≤t}.

Formalization targets

Goal: Theorem 3.3 (p. 13)

Under (A1), (A2), S(μ)≠∅S(\mu)\ne\emptysetS(μ)=∅ and U\mathcal UU an open bounded neighbourhood of S(μ)S(\mu)S(μ), there are L>0L>0L>0 and δ>0\delta>0δ>0 such that for every ν∈P2(Ξ)\nu\in\mathcal P_2(\Xi)ν∈P2​(Ξ) with ζ2(μ,ν)<δ\zeta_2(\mu,\nu)<\deltaζ2​(μ,ν)<δ:

∣v(μ)−v(ν)∣≤L ζ2(μ,ν),∅≠S(ν)⊆S(μ)+Ψ(L ζ2(μ,ν)) B.|v(\mu)-v(\nu)|\le L\,\zeta_2(\mu,\nu),\qquad \emptyset\ne S(\nu)\subseteq S(\mu)+\Psi(L\,\zeta_2(\mu,\nu))\,\mathbb B.∣v(μ)−v(ν)∣≤Lζ2​(μ,ν),∅=S(ν)⊆S(μ)+Ψ(Lζ2​(μ,ν))B.

The constants are existential, so the goal does not depend on any particular estimate of them.

Milestones

  1. Lemma 3.1 (p. 11): Φ\PhiΦ is finite and continuous on the polyhedral cone D×pos⁡WD\times\operatorname{pos}WD×posW, piecewise of the form Cju⋅tC_ju\cdot tCj​u⋅t on finitely many polyhedral cones with disjoint interiors, convex in ttt and concave in uuu.
  2. Proposition 3.2 (p. 11): f0f_0f0​ is a normal convex integrand, and for ∥x∥≤r\|x\|\le r∥x∥≤r it satisfies ∣f0(ξ,x)−f0(ξ~,x)∣≤Lrmax⁡{1,∥ξ∥,∥ξ~∥}∥ξ−ξ~∥|f_0(\xi,x)-f_0(\tilde\xi,x)|\le Lr\max\{1,\|\xi\|,\|\tilde\xi\|\}\|\xi-\tilde\xi\|∣f0​(ξ,x)−f0​(ξ~​,x)∣≤Lrmax{1,∥ξ∥,∥ξ~​∥}∥ξ−ξ~​∥, a Lipschitz bound in xxx with factor L^max⁡{1,∥ξ∥2}\hat L\max\{1,\|\xi\|^2\}L^max{1,∥ξ∥2}, and ∣f0(ξ,x)∣≤Krmax⁡{1,∥ξ∥2}|f_0(\xi,x)|\le Kr\max\{1,\|\xi\|^2\}∣f0​(ξ,x)∣≤Krmax{1,∥ξ∥2}.
  3. PFU⊇P2(Ξ)\mathcal P_{\mathcal F_{\mathcal U}}\supseteq\mathcal P_2(\Xi)PFU​​⊇P2​(Ξ) (p. 13): every measure with finite second moment satisfies the integrability conditions of the general theory.
  4. Corollary 2.8 (p. 10): the general Lipschitz stability theorem in the metric ζg\zeta_gζg​ for convex models whose integrands lie, up to a factor LLL, in the class Fg\mathcal F_gFg​.

Significance

The theorem says that, under the two natural well-posedness conditions of linear recourse models, optimal values are Lipschitz continuous and solution sets move by at most Ψ(Lζ2)\Psi(L\zeta_2)Ψ(Lζ2​), where ζ2\zeta_2ζ2​ only sees test functions with the local Lipschitz growth of the integrand. Since ζ2≤(1+∫∥ξ∥2dμ+∫∥ξ∥2dν)1/2W2\zeta_2\le(1+\int\|\xi\|^2d\mu+\int\|\xi\|^2d\nu)^{1/2}W_2ζ2​≤(1+∫∥ξ∥2dμ+∫∥ξ∥2dν)1/2W2​, the result contains the earlier W2W_2W2​ stability theorem.

The results are proved in the paper, except Lemma 3.1, which it cites from Walkup and Wets (1969). None of them is formalized in Mathlib or on Prove2Me. A formalization would provide a checked model of the two-stage recourse function, the Fortet–Mourier metrics, and an end-to-end quantitative stability theorem for stochastic programs, none of which exist in Mathlib.

Difficulty

The obvious argument bounds ∣v(μ)−v(ν)∣|v(\mu)-v(\nu)|∣v(μ)−v(ν)∣ by sup⁡x∣∫f0(⋅,x) d(μ−ν)∣\sup_x|\int f_0(\cdot,x)\,d(\mu-\nu)|supx​∣∫f0​(⋅,x)d(μ−ν)∣ and then by Lζ2L\zeta_2Lζ2​. Two steps fail. First, the supremum ranges over all of XXX, which may be unbounded, while f0(⋅,x)/Lf_0(\cdot,x)/Lf0​(⋅,x)/L lies in F2\mathcal F_2F2​ only for xxx in a bounded set; the argument must localize to X∩cl⁡UX\cap\operatorname{cl}\mathcal UX∩clU and then show, using convexity, that for ν\nuν close to μ\muμ the localized and global problems coincide. Second, the Lipschitz estimate in ξ\xiξ for f0f_0f0​ requires the piecewise-bilinear structure of Φ\PhiΦ (Lemma 3.1) and a chaining argument across the polyhedral pieces of Ξ\XiΞ; the integrand is not differentiable.

Formalization scope

  • Rn\mathbb R^nRn is EuclideanSpace ℝ (Fin n). The paper never fixes its norm and all its constants are existential, so the Euclidean norm is a faithful instance. Second-stage vectors u,t,y,zu,t,y,zu,t,y,z are coordinate vectors Fin k → ℝ; index sets are 0-based.
  • A polyhedron is a finite intersection of closed half-spaces. The standing assumptions are: XXX a nonempty polyhedron, Ξ\XiΞ a polyhedron, q,h,Tq,h,Tq,h,T affine maps.
  • A measure in P(Ξ)\mathcal P(\Xi)P(Ξ) is a Borel probability measure on Rs\mathbb R^sRs with ν(Ξc)=0\nu(\Xi^c)=0ν(Ξc)=0; ∫Ξf0(ξ,x) ν(dξ)\int_\Xi f_0(\xi,x)\,\nu(d\xi)∫Ξ​f0​(ξ,x)ν(dξ) is the extended-real integral of the published DupacovaWets.Consistency.expect over ν\nuν restricted to Ξ\XiΞ.
  • Φ\PhiΦ, f0f_0f0​, vvv and ψ\psiψ are extended-real infima (+∞+\infty+∞ over an empty set); "min" in ψ\psiψ is read as an infimum. ζ2\zeta_2ζ2​, ζg\zeta_gζg​, ψ−1\psi^{-1}ψ−1 and Ψ\PsiΨ take values in [0,∞][0,\infty][0,∞], and the distance to a set is the extended distance (+∞+\infty+∞ to ∅\emptyset∅).
  • To rule out trivializations: every value bound states that v(μ)v(\mu)v(μ), v(ν)v(\nu)v(ν) are finite before comparing them, so the extended-real arithmetic ∞−∞\infty-\infty∞−∞ cannot make it vacuous; in ζ2\zeta_2ζ2​ a test function that is not integrable for both measures contributes +∞+\infty+∞, never a junk 000.
  • Disclosed readings: Proposition 3.2 leaves its radius rrr unquantified; it is stated with the constants first and then every r≥1r\ge1r≥1. Lemma 3.1's "pos⁡W×D\operatorname{pos}W\times DposW×D" is read as D×pos⁡WD\times\operatorname{pos}WD×posW, in the order of Φ\PhiΦ's arguments. Corollary 2.8 adds μ,ν∈PFU(Ξ)\mu,\nu\in\mathcal P_{\mathcal F_{\mathcal U}}(\Xi)μ,ν∈PFU​​(Ξ), an assumption of Theorem 2.2 that its proof invokes; Theorem 3.3 needs no such addition, because milestone 3 supplies it.
  • Needed infrastructure: LP duality and the Walkup–Wets decomposition of Φ\PhiΦ; normal integrands and measurability of their infima; integrability under moment conditions; the general stability theorems (Theorems 2.2–2.3 of the paper) for d=0d=0d=0. The definitions of ζp\zeta_pζp​, Fg\mathcal F_gFg​ and normal integrands are reusable well beyond this mission. Contributions toward any milestone, and proofs of the auxiliary facts (for example, that every f∈F2(Ξ)f\in\mathcal F_2(\Xi)f∈F2​(Ξ) is integrable for ν∈P2(Ξ)\nu\in\mathcal P_2(\Xi)ν∈P2​(Ξ)), are welcome.

Selected references

  • S. T. Rachev, W. Römisch, Quantitative stability in stochastic programming: The method of probability metrics, Mathematics of Operations Research 27(4) (2002) 792–818. https://doi.org/10.1287/moor.27.4.792.304 (cited here from the authors' preprint, edoc.hu-berlin.de)
  • D. W. Walkup, R. J-B Wets, Lifting projections of convex polyhedra, Pacific Journal of Mathematics 28 (1969) 465–475 (the paper's [55]).
  • S. M. Robinson, R. J-B Wets, Stability in two-stage stochastic programming, SIAM Journal on Control and Optimization 25 (1987) 1409–1416 (the paper's [37]).
  • W. Römisch, R. Schultz, Stability analysis for stochastic programs, Annals of Operations Research 30 (1991) 241–266 (the paper's [40]).
  • R. T. Rockafellar, R. J-B Wets, Variational Analysis, Springer (the paper's [38]; normal integrands, Chapter 14).
  • S. T. Rachev, Probability Metrics and the Stability of Stochastic Models, Wiley, Chichester 1991 (the paper's [32]).
8 thms1 active userReviewed
AnalysisStochastic Systems·Captain: mikedeng1

The Power of Two Choices in Randomized Load Balancing II: T_d(λ)/log T_1(λ) → 1/log d as λ → 1⁻, an Exponential Improvement over One ChoiceResearch Paper

Motivation

A dispatcher assigns each arriving job to one of nnn servers. Sending it to a uniformly random server is simple but produces long queues near capacity; sending it to the globally shortest queue requires full state information. The power of ddd choices sits between the two: each job samples ddd servers at random and joins the shortest of them. Mitzenmacher's analysis of this supermarket model (IEEE TPDS 12(10), 2001) together with the balls-into-bins result of Azar, Broder, Karlin and Upfal (SIAM J. Comput. 29(1), 1999), established that a second random choice changes the behaviour of the system qualitatively. The idea underlies load balancers, hashing schemes and distributed schedulers.

This mission formalizes the paper's quantitative statement of that improvement in heavy traffic, the regime where the arrival rate per server λ\lambdaλ approaches the service rate 111.

Setting

Customers arrive as a Poisson stream of rate λn\lambda nλn with 0≤λ<10\le\lambda<10≤λ<1, service times are exponential with mean 111, and each customer joins the shortest of ddd queues sampled uniformly with replacement. As n→∞n\to\inftyn→∞ the fraction of queues with at least iii customers follows a deterministic limiting system with fixed point πi=λ(di−1)/(d−1)\pi_i=\lambda^{(d^i-1)/(d-1)}πi​=λ(di−1)/(d−1). Corollary 2 of the paper shows that, for d≥2d\ge2d≥2, the expected time a customer spends in that limiting system converges to

Td(λ)=∑i=1∞λdi−dd−1.T_d(\lambda)=\sum_{i=1}^{\infty}\lambda^{\frac{d^i-d}{d-1}} .Td​(λ)=i=1∑∞​λd−1di−d​.

With a single choice the system is nnn independent M/M/1 queues, and the expected time is

T1(λ)=11−λ.T_1(\lambda)=\frac{1}{1-\lambda}.T1​(λ)=1−λ1​.

This mission takes these two formulas as its definitions; the limiting system itself and the convergence to Td(λ)T_d(\lambda)Td​(λ) are the subject of the companion mission (part I). The auxiliary function of Lemma 3 is

Fd(λ)=∑i=0∞λdilog⁡11−λ.F_d(\lambda)=\frac{\sum_{i=0}^{\infty}\lambda^{d^i}}{\log\frac{1}{1-\lambda}} .Fd​(λ)=log1−λ1​∑i=0∞​λdi​.

Throughout, d≥2d\ge2d≥2 is an integer and logarithms are natural.

Formalization targets

Goal: Theorem 4, limit clause (p. 1099)

lim⁡λ→1−Td(λ)log⁡T1(λ)=1log⁡d(d≥2).\lim_{\lambda\to1^-}\frac{T_d(\lambda)}{\log T_1(\lambda)}=\frac{1}{\log d}\qquad(d\ge2).λ→1−lim​logT1​(λ)Td​(λ)​=logd1​(d≥2).

Milestones, in the order the paper's argument uses them

  1. The rewriting of TdT_dTd​ (proof of Theorem 4): with λ′=λ1/(d−1)\lambda'=\lambda^{1/(d-1)}λ′=λ1/(d−1) and 0<λ<10<\lambda<10<λ<1,
Td(λ)=∑i=1∞(λ′)diλd/(d−1).T_d(\lambda)=\frac{\sum_{i=1}^{\infty}(\lambda')^{d^i}}{\lambda^{d/(d-1)}} .Td​(λ)=λd/(d−1)∑i=1∞​(λ′)di​.
  1. The product identity (proof of Lemma 3): for 0≤λ<10\le\lambda<10≤λ<1,
∏i=0∞(1+λdi+λ2di+⋯+λ(d−1)di)=11−λ.\prod_{i=0}^{\infty}\bigl(1+\lambda^{d^i}+\lambda^{2d^i}+\dots+\lambda^{(d-1)d^i}\bigr)=\frac{1}{1-\lambda}.i=0∏∞​(1+λdi+λ2di+⋯+λ(d−1)di)=1−λ1​.
  1. Lemma 3:
lim⁡λ→1−Fd(λ)=1log⁡d.\lim_{\lambda\to1^-}F_d(\lambda)=\frac{1}{\log d}.λ→1−lim​Fd​(λ)=logd1​.

The goal is a statement about the shape of the growth, not a constant: it pins the leading behaviour Td(λ)∼log⁡d11−λT_d(\lambda)\sim\log_d\frac{1}{1-\lambda}Td​(λ)∼logd​1−λ1​.

Significance

Theorem 4 makes "exponential improvement" precise. With one choice the expected time grows like 11−λ\frac1{1-\lambda}1−λ1​ as the load approaches capacity; with d≥2d\ge2d≥2 choices it grows like log⁡11−λlog⁡d\frac{\log\frac1{1-\lambda}}{\log d}logdlog1−λ1​​. It also quantifies the diminishing return of extra choices: going from d=2d=2d=2 to ddd only divides the leading term by log⁡2d\log_2 dlog2​d, while going from one choice to two changes its order. These constants are the benchmark against which later heavy-traffic analyses of join-the-shortest-of-ddd policies are compared.

The result is proved in the paper (the limit clause; see Difficulty for the other clause). No machine-checked version is known to exist. The formalization yields a self-contained Lean development of the lacunary series ∑ixdi\sum_i x^{d^i}∑i​xdi near x=1x=1x=1 and of the base-ddd product identity, both of which are classical analytic facts with uses outside queueing (lacunary power series, digit expansions).

Difficulty

The obvious approach compares ∑iλdi\sum_i\lambda^{d^i}∑i​λdi with an integral, but the terms are not monotone images of a smooth function of a continuous index in a way that controls the error uniformly as λ→1−\lambda\to1^-λ→1−: the number of terms close to 111 grows like log⁡d11−λ\log_d\frac{1}{1-\lambda}logd​1−λ1​ and each of them contributes nearly 111, so the error terms are of the same order as the quantity being measured. The difficulty is to obtain matching upper and lower bounds with constants tending to 1/log⁡d1/\log d1/logd, uniformly as λ→1−\lambda\to1^-λ→1−. Passing from Lemma 3 to Theorem 4 additionally requires controlling the change of variable λ↦λ1/(d−1)\lambda\mapsto\lambda^{1/(d-1)}λ↦λ1/(d−1) inside the logarithm.

Theorem 4 as printed also contains a first clause, Td(λ)≤cdlog⁡T1(λ)T_d(\lambda)\le c_d\log T_1(\lambda)Td​(λ)≤cd​logT1​(λ) for all λ∈[0,1]\lambda\in[0,1]λ∈[0,1]. It is false as printed near λ=0\lambda=0λ=0: Td(λ)≥1T_d(\lambda)\ge1Td​(λ)≥1 while log⁡T1(λ)→0\log T_1(\lambda)\to0logT1​(λ)→0. The paper does not prove it, and this mission does not pose it.

Formalization scope

  • Representation. λ\lambdaλ is a real number (lam), ddd a natural number with the hypothesis 2≤d2\le d2≤d on every statement. TdT_dTd​, T1T_1T1​ and FdF_dFd​ are real-valued definitions in one definition file.
  • Exponents. TdT_dTd​'s iii-th term uses the natural-number exponent ∑1≤k<idk=di−dd−1\sum_{1\le k<i}d^k=\frac{d^i-d}{d-1}∑1≤k<i​dk=d−1di−d​, so its i=1i=1i=1 term is λ0=1\lambda^0=1λ0=1. The series runs over i∈Ni\in\mathbb Ni∈N with the i=0i=0i=0 term equal to 000. λ′=λ1/(d−1)\lambda'=\lambda^{1/(d-1)}λ′=λ1/(d−1) and λd/(d−1)\lambda^{d/(d-1)}λd/(d−1) are real powers.
  • Series and junk values. All series are real tsums; they converge for 0≤λ<10\le\lambda<10≤λ<1. Outside that range Lean's conventions (a divergent series sums to 000, x/0=0x/0=0x/0=0, log⁡x=0\log x=0logx=0 for x≤0x\le0x≤0) produce meaningless values, so every statement either restricts λ\lambdaλ to [0,1)[0,1)[0,1) or (0,1)(0,1)(0,1) or is a limit along λ<1\lambda<1λ<1.
  • Limits. "λ→1−\lambda\to1^-λ→1−" is the one-sided filter 𝓝[<] 1; the limit 1/log⁡d1/\log d1/logd uses Real.log.
  • Infinite product. The product identity is stated with HasProd (unconditional convergence of finite subproducts), which for factors ≥1\ge1≥1 is the same as convergence of the partial products.
  • Not trivializable. Because every limit is one-sided at 111 and the denominators log⁡T1(λ)\log T_1(\lambda)logT1​(λ) and log⁡11−λ\log\frac1{1-\lambda}log1−λ1​ are positive there, no statement can be satisfied by Lean's junk values at λ≥1\lambda\ge1λ≥1 or at λ=0\lambda=0λ=0; the summability of the rewritten series is asserted explicitly.
  • Infrastructure. Useful general lemmas: asymptotics of ∑ixdi\sum_i x^{d^i}∑i​xdi as x→1−x\to1^-x→1−, products of finite geometric sums, change of variables in one-sided limits. These are reusable beyond this mission. Contributions of any of the milestones, or of alternative proofs of Lemma 3, are welcome.

Selected references

  • M. Mitzenmacher, The Power of Two Choices in Randomized Load Balancing, IEEE Transactions on Parallel and Distributed Systems 12(10), 2001, pp. 1094–1104. https://doi.org/10.1109/71.963420
  • Y. Azar, A. Z. Broder, A. R. Karlin, E. Upfal, Balanced Allocations, SIAM Journal on Computing 29(1), 1999, pp. 180–200. https://doi.org/10.1137/S0097539795288490
5 thms1 active userReviewed
Dynamic ProgrammingOptimizationReinforcement Learning·Captain: mikedeng1

Bounded-parameter Markov Decision Processes 1: The Optimistic and Pessimistic Optimal Interval Value Functions Satisfy Bellman-like EquationsResearch Paper

Motivation

A Markov decision process is specified by numbers: transition probabilities and rewards. In practice these numbers are estimated from data, elicited from experts, or produced by aggregating the states of a larger model, and in each case what is actually known is a range for each parameter rather than its value. Givan, Leach and Dean (Artificial Intelligence, 2000) introduced bounded-parameter MDPs (BMDPs) to reason about a whole family of MDPs at once: every transition probability is only known to lie in a closed interval. Their motivation was state-space aggregation, where grouping states with similar but not identical dynamics yields exactly such interval bounds, and the same object later reappeared as the interval or "box" uncertainty set of robust MDPs (Nilim and El Ghaoui, 2005; Iyengar, 2005) and of optimistic exploration in reinforcement learning (Strehl and Littman, 2008).

Over such a family, the value of a policy is no longer a number but an interval, and "optimal" needs a definition. This mission formalizes the paper's answer: two total orders on intervals, the corresponding optimal policies, and the Bellman-like equations their interval value functions satisfy.

Setting

An exact MDP M=⟨Q,A,F,R⟩M=\langle Q,A,F,R\rangleM=⟨Q,A,F,R⟩ has a finite set QQQ of states, a finite nonempty set AAA of actions, transition probabilities Fpq(α)≥0F_{pq}(\alpha)\ge 0Fpq​(α)≥0 with ∑qFpq(α)=1\sum_{q}F_{pq}(\alpha)=1∑q​Fpq​(α)=1, a reward R(q)∈RR(q)\in\mathbb RR(q)∈R for each state and a discount rate 0≤γ<10\le\gamma<10≤γ<1. A policy is a map π:Q→A\pi:Q\to Aπ:Q→A; Π\PiΠ is the set of all policies. The value function VM,π(p)V_{M,\pi}(p)VM,π​(p) is the expected discounted sum of rewards ∑t≥0γt E[R(Xt)∣X0=p]\sum_{t\ge 0}\gamma^t\,\mathbb E[R(X_t)\mid X_0=p]∑t≥0​γtE[R(Xt​)∣X0​=p] along the Markov chain with transitions Fpq(π(p))F_{pq}(\pi(p))Fpq​(π(p)). The operators

VIM,π(v)(p)=R(p)+γ∑qFpq(π(p)) v(q),VIM,α(v)(p)=R(p)+γ∑qFpq(α) v(q)VI_{M,\pi}(v)(p)=R(p)+\gamma\sum_{q}F_{pq}(\pi(p))\,v(q),\qquad VI_{M,\alpha}(v)(p)=R(p)+\gamma\sum_{q}F_{pq}(\alpha)\,v(q)VIM,π​(v)(p)=R(p)+γq∑​Fpq​(π(p))v(q),VIM,α​(v)(p)=R(p)+γq∑​Fpq​(α)v(q)

act on value functions v:Q→Rv:Q\to\mathbb Rv:Q→R. V1≤domV2V_1\le_{\mathrm{dom}}V_2V1​≤dom​V2​ is the statewise order.

A BMDP M↕M_\updownarrowM↕​ gives for each p,q,αp,q,\alphap,q,α an interval [F↓pq(α),F↑pq(α)]⊆[0,1][F_\downarrow{}_{pq}(\alpha),F_\uparrow{}_{pq}(\alpha)]\subseteq[0,1][F↓​pq​(α),F↑​pq​(α)]⊆[0,1] with ∑qF↓pq(α)≤1≤∑qF↑pq(α)\sum_q F_\downarrow{}_{pq}(\alpha)\le 1\le\sum_q F_\uparrow{}_{pq}(\alpha)∑q​F↓​pq​(α)≤1≤∑q​F↑​pq​(α), a reward RRR and a discount rate γ\gammaγ. Its members M∈M↕M\in M_\updownarrowM∈M↕​ are the exact MDPs with that RRR and γ\gammaγ and with FpqM(α)∈[F↓pq(α),F↑pq(α)]F^M_{pq}(\alpha)\in[F_\downarrow{}_{pq}(\alpha),F_\uparrow{}_{pq}(\alpha)]FpqM​(α)∈[F↓​pq​(α),F↑​pq​(α)]. The interval value of a policy is

V↕π(q)=[min⁡M∈M↕VM,π(q), max⁡M∈M↕VM,π(q)]=[V↓π(q),V↑π(q)].V_\updownarrow{}_\pi(q)=\Big[\min_{M\in M_\updownarrow}V_{M,\pi}(q),\ \max_{M\in M_\updownarrow}V_{M,\pi}(q)\Big]=[V_\downarrow{}_\pi(q),V_\uparrow{}_\pi(q)].V↕​π​(q)=[M∈M↕​min​VM,π​(q), M∈M↕​max​VM,π​(q)]=[V↓​π​(q),V↑​π​(q)].

Intervals are compared by the optimistic order, which compares upper bounds first and breaks ties by lower bounds, and the pessimistic order, which compares lower bounds first:

[l1,u1]≤opt[l2,u2]  ⟺  u1<u2 ∨ (u1=u2∧l1≤l2),[l1,u1]≤pes[l2,u2]  ⟺  l1<l2 ∨ (l1=l2∧u1≤u2).[l_1,u_1]\le_{\mathrm{opt}}[l_2,u_2]\iff u_1<u_2\ \vee\ (u_1=u_2\wedge l_1\le l_2),\qquad [l_1,u_1]\le_{\mathrm{pes}}[l_2,u_2]\iff l_1<l_2\ \vee\ (l_1=l_2\wedge u_1\le u_2).[l1​,u1​]≤opt​[l2​,u2​]⟺u1​<u2​ ∨ (u1​=u2​∧l1​≤l2​),[l1​,u1​]≤pes​[l2​,u2​]⟺l1​<l2​ ∨ (l1​=l2​∧u1​≤u2​).

A policy πopt\pi_{\mathrm{opt}}πopt​ is optimistically optimal if V↕πopt(q)≥optV↕π(q)V_\updownarrow{}_{\pi_{\mathrm{opt}}}(q)\ge_{\mathrm{opt}}V_\updownarrow{}_\pi(q)V↕​πopt​​(q)≥opt​V↕​π​(q) for all π\piπ and qqq; pessimistically optimal is defined with ≥pes\ge_{\mathrm{pes}}≥pes​. The optimal interval value functions are V↕opt=V↕πoptV_\updownarrow{}_{\mathrm{opt}}=V_\updownarrow{}_{\pi_{\mathrm{opt}}}V↕​opt​=V↕​πopt​​ and V↕pes=V↕πpesV_\updownarrow{}_{\mathrm{pes}}=V_\updownarrow{}_{\pi_{\mathrm{pes}}}V↕​pes​=V↕​πpes​​.

Formalization targets

Goal: Theorem 9, equations (25) and (26)

At every state ppp,

V↕opt(p)=max⁡α∈A, ≤opt[min⁡M∈M↕VIM,α(V↓opt)(p), max⁡M∈M↕VIM,α(V↑opt)(p)],V_\updownarrow{}_{\mathrm{opt}}(p)=\max_{\alpha\in A,\ \le_{\mathrm{opt}}}\Big[\min_{M\in M_\updownarrow}VI_{M,\alpha}(V_\downarrow{}_{\mathrm{opt}})(p),\ \max_{M\in M_\updownarrow}VI_{M,\alpha}(V_\uparrow{}_{\mathrm{opt}})(p)\Big],V↕​opt​(p)=α∈A, ≤opt​max​[M∈M↕​min​VIM,α​(V↓​opt​)(p), M∈M↕​max​VIM,α​(V↑​opt​)(p)],

and the same with pes\mathrm{pes}pes in place of opt\mathrm{opt}opt. The maximum over actions is for the total order on intervals: attained by some action and an upper bound for every action.

Milestones

In the order the proof uses them: Theorem 3 (second sentence: VM,πV_{M,\pi}VM,π​ is the unique fixed point of VIM,πVI_{M,\pi}VIM,π​); Theorem 6 (comparison: u≤domVIM,π(u)u\le_{\mathrm{dom}}VI_{M,\pi}(u)u≤dom​VIM,π​(u) implies u≤domVM,πu\le_{\mathrm{dom}}V_{M,\pi}u≤dom​VM,π​, and its three variants); Lemma 1 (every member's values and backups are bracketed by finitely many order-maximizing MDPs); Lemma 2 (statewise composition of member MDPs dominates, or is dominated by, both components); Theorem 7 (π\piπ-maximizing and π\piπ-minimizing MDPs exist among the order-maximizing ones); Corollary 1 (V↓πV_\downarrow{}_\piV↓​π​ and V↑πV_\uparrow{}_\piV↑​π​ are a minimum and a maximum for ≤dom\le_{\mathrm{dom}}≤dom​); Lemma 3 (statewise composition of policies improves the upper, resp. lower, bounds); Theorem 8 (optimistically and pessimistically optimal policies exist).

Significance

Theorem 9 is the interval analogue of the Bellman optimality equation. It is what makes the optimal interval value functions computable: the paper's interval value iteration algorithms IVI↕optIVI_\updownarrow{}_{\mathrm{opt}}IVI↕​opt​ and IVI↕pesIVI_\updownarrow{}_{\mathrm{pes}}IVI↕​pes​ (Section 5) iterate exactly the right-hand sides of (25) and (26). The lower half of (26) is the max-min equation of robust dynamic programming with rectangular box uncertainty, and the upper half of (25) is the max-max equation behind optimistic planning; Theorem 9 states both, together with the tie-breaking second component, in one model. Theorems 7 and 8 are of independent use: a single member MDP is worst (or best) for a policy at all states at once, and the partial orders on interval value functions nevertheless have maxima over policies.

None of these results has a machine-checked proof. The platform has robust Bellman equations in other models (costs, general rectangular sets, a sup over nature) but nothing on BMDPs, interval value functions, or the orders ≤opt\le_{\mathrm{opt}}≤opt​ and ≤pes\le_{\mathrm{pes}}≤pes​. The mission produces the paper's optimality theory end to end, from the comparison principle for exact MDPs to the two Bellman-like equations, on a model shared with the companion mission on interval policy evaluation.

Difficulty

The orders ≤opt\le_{\mathrm{opt}}≤opt​ and ≤pes\le_{\mathrm{pes}}≤pes​ are total on intervals but only partial on interval value functions, and they are lexicographic. The obvious construction of an optimal policy, switching statewise to whichever of two policies has the better interval, fails: the composed policy need not dominate either component in ≤opt\le_{\mathrm{opt}}≤opt​, because lower bounds can get worse at states where the upper bounds do not change (the paper says so explicitly, eq. (20)). Existence therefore needs a two-stage argument. A second obstacle is that V↓πV_\downarrow{}_\piV↓​π​ and V↑πV_\uparrow{}_\piV↑​π​ are defined statewise, as minima over an uncountable family; that one member attains them at every state simultaneously is a theorem (Theorem 7), and the Bellman-like equations rest on it. Finally, the inner minima in (25) act on V↓optV_\downarrow{}_{\mathrm{opt}}V↓​opt​, the lower bound of an optimistically chosen policy, which is not the componentwise best lower bound over policies; the equation couples the two components through the tie-break.

Formalization scope

QQQ and AAA are finite types, AAA is nonempty; QQQ may be empty. An exact MDP is a structure with transition function F p α q =Fpq(α)=F_{pq}(\alpha)=Fpq​(α), stochastic rows, reward R:Q→RR:Q\to\mathbb RR:Q→R and discount 0≤γ<10\le\gamma<10≤γ<1. VM,πV_{M,\pi}VM,π​ is defined as the convergent series ∑tγtPπtR\sum_t\gamma^tP_\pi^tR∑t​γtPπt​R, not as a solution of (2). Rewards of the BMDP are tight, as the paper assumes from footnote 3 on; reward intervals are not formalized. Members of M↕M_\updownarrowM↕​ form a subtype of exact MDPs with the BMDP's RRR and γ\gammaγ. The minima and maxima of Definition 3 and of (25)–(26) are the real infimum and supremum over that subtype, which is nonempty and on which values are bounded, so no junk value enters. Intervals are pairs (lower, upper); the orders (17) are spelled out. Policies are deterministic stationary maps Q→AQ\to AQ→A, the paper's Π\PiΠ. An ordering of QQQ is a bijection Fin |Q| ≃ Q, and the index rrr of Definition 1 is the largest index at which expression (9) does not exceed 1 (the page's phrase admits ties; only the largest index keeps the remaining mass inside its interval). Definition 8 is read as V↕opt=V↕πoptV_\updownarrow{}_{\mathrm{opt}}=V_\updownarrow{}_{\pi_{\mathrm{opt}}}V↕​opt​=V↕​πopt​​ for an optimistically optimal πopt\pi_{\mathrm{opt}}πopt​, so the goal is quantified over all optimistically (pessimistically) optimal policies; Theorem 8 is the milestone that makes this quantification non-vacuous. Lemma 3 is stated as in the Appendix (p. 36): part (d) concerns π4=π1⊕pesπ2\pi_4=\pi_1\oplus_{\mathrm{pes}}\pi_2π4​=π1​⊕pes​π2​, where p. 17 misprints π3\pi_3π3​.

A formalization in which members carry their own discount, rows need not sum to one, the minima range over all functions Q→A→Q→RQ\to A\to Q\to\mathbb RQ→A→Q→R, or V↕optV_\updownarrow{}_{\mathrm{opt}}V↕​opt​ is the componentwise supremum over policies would make several statements false or trivial; each is ruled out above.

Contributions welcome beyond proofs of the milestones: a reusable theory of discounted policy evaluation for finite MDPs (the series representation, the fixed-point characterization, monotonicity of VIM,πVI_{M,\pi}VIM,π​), which other missions on finite MDPs can import.

Selected references

  • R. Givan, S. Leach, T. Dean, Bounded-parameter Markov decision processes, Artificial Intelligence 122 (2000). https://doi.org/10.1016/S0004-3702(00)00047-3
  • A. Nilim, L. El Ghaoui, Robust control of Markov decision processes with uncertain transition matrices, Operations Research 53(5), 2005. https://doi.org/10.1287/opre.1050.0216
  • G. Iyengar, Robust dynamic programming, Mathematics of Operations Research 30(2), 2005. https://doi.org/10.1287/moor.1040.0129
  • A. Strehl, M. Littman, An analysis of model-based interval estimation for Markov decision processes, Journal of Computer and System Sciences 74(8), 2008. https://doi.org/10.1016/j.jcss.2007.08.009
  • M. Puterman, Markov Decision Processes: Discrete Stochastic Dynamic Programming, Wiley, 1994. https://doi.org/10.1002/9780470316887
11 thms1 active userReviewed
Dynamical SystemsStochastic SystemsTheoretical Computer Science·Captain: mikedeng1

The Power of Two Choices in Randomized Load Balancing I: In the Limiting Supermarket System the Expected Time Converges to T_d(λ) = Σ λ^((d^i−d)/(d−1))Research Paper

Motivation

Randomized load balancing asks how much a dispatcher gains from a little information. In the supermarket model, jobs arrive at nnn servers and each job inspects a few servers chosen at random before joining one. With one random choice every server is an independent M/M/1 queue, and the expected time a job spends in the system is 1/(1−λ)1/(1-\lambda)1/(1−λ) at load λ\lambdaλ. With two choices it is exponentially smaller as λ→1−\lambda \to 1^-λ→1−. This "power of two choices" underlies the design of distributed schedulers, hashing schemes and server farms, where querying every server is too expensive but querying two is cheap.

Timeline:

  • 1994/1999. Azar, Broder, Karlin and Upfal proved the static version, balanced allocations: throwing nnn balls into nnn bins with d≥2d \ge 2d≥2 choices per ball gives a maximum load of log⁡log⁡n/log⁡d+O(1)\log\log n/\log d + O(1)loglogn/logd+O(1) w.h.p., against Θ(log⁡n/log⁡log⁡n)\Theta(\log n/\log\log n)Θ(logn/loglogn) for one choice (SIAM J. Comput. 1999).
  • 1996. Vvedenskaya, Dobrushin and Karpelevich derived the dynamic limiting equations for d=2d = 2d=2 and showed convergence of their trajectories to the fixed point, without a rate (Problems of Information Transmission 32(1), 1996).
  • 1996/2001. Mitzenmacher, independently, introduced the limiting system for general ddd, proved doubly exponential tails along every trajectory and exponential convergence to the fixed point in a weighted L1L_1L1​ potential, and derived the limiting expected time Td(λ)T_d(\lambda)Td​(λ) (IEEE TPDS 2001). This mission formalizes Section 2 of that paper.

Setting

Customers arrive as a Poisson stream of rate λn\lambda nλn, 0<λ<10<\lambda<10<λ<1, at nnn FIFO servers. Each customer chooses ddd servers independently and uniformly at random with replacement and joins the one holding the fewest customers; service times are exponential with mean 111. Let si(t)s_i(t)si​(t) be the fraction of servers holding at least iii customers at time ttt.

A state is a sequence s=(s0,s1,s2,… )s = (s_0, s_1, s_2, \dots)s=(s0​,s1​,s2​,…) of reals with s0=1s_0 = 1s0​=1, si≥0s_i \ge 0si​≥0 and s0≥s1≥s2≥⋯s_0 \ge s_1 \ge s_2 \ge \cdotss0​≥s1​≥s2​≥⋯. The empty state has s0=1s_0 = 1s0​=1 and si=0s_i = 0si​=0 for i≥1i \ge 1i≥1. As n→∞n \to \inftyn→∞ the tails follow the limiting system

dsidt=λ (si−1d−sid)−(si−si+1)(i≥1),s0=1.(1)\frac{ds_i}{dt} = \lambda\,(s_{i-1}^d - s_i^d) - (s_i - s_{i+1}) \quad (i \ge 1), \qquad s_0 = 1. \tag{1}dtdsi​​=λ(si−1d​−sid​)−(si​−si+1​)(i≥1),s0​=1.(1)

A trajectory is a map t↦s(t)t \mapsto s(t)t↦s(t), t≥0t \ge 0t≥0, whose values are states and whose coordinates satisfy (1) (from the right at t=0t = 0t=0). Throughout, d≥2d \ge 2d≥2 is an integer.

The fixed point is πi=λ(di−1)/(d−1)\pi_i = \lambda^{(d^i - 1)/(d - 1)}πi​=λ(di−1)/(d−1), so π0=1\pi_0 = 1π0​=1, π1=λ\pi_1 = \lambdaπ1​=λ, π2=λ1+d\pi_2 = \lambda^{1+d}π2​=λ1+d. A customer arriving in state sss becomes the iii-th customer of its queue with probability si−1d−sids_{i-1}^d - s_i^dsi−1d​−sid​ and then waits for iii services, so the expected time it spends in the system is

E(s)=∑i=1∞i (si−1d−sid).E(s) = \sum_{i=1}^{\infty} i\,(s_{i-1}^d - s_i^d).E(s)=i=1∑∞​i(si−1d​−sid​).

The target constant is

Td(λ)=∑i=1∞λdi−dd−1.T_d(\lambda) = \sum_{i=1}^{\infty} \lambda^{\frac{d^i - d}{d-1}}.Td​(λ)=i=1∑∞​λd−1di−d​.

Formalization targets

Goal: Corollary 2

For d≥2d \ge 2d≥2 and 0<λ<10 < \lambda < 10<λ<1, Td(λ)<∞T_d(\lambda) < \inftyTd​(λ)<∞, and:

if sj(0)=0 for some j, then E(s(t))→t→∞Td(λ);if s(0) is empty, then E(s(t))≤Td(λ)  ∀t≥0.\text{if } s_j(0) = 0 \text{ for some } j, \text{ then } E(s(t)) \xrightarrow[t\to\infty]{} T_d(\lambda);\qquad \text{if } s(0) \text{ is empty, then } E(s(t)) \le T_d(\lambda) \ \ \forall t \ge 0.if sj​(0)=0 for some j, then E(s(t))t→∞​Td​(λ);if s(0) is empty, then E(s(t))≤Td​(λ)  ∀t≥0.

The hypothesis sj(0)=0s_j(0) = 0sj​(0)=0 holds for every initial state that comes from a finite system.

Milestones, in the order the proof uses them

  • Lemma 2. π\piπ is the unique fixed point of (1) with ∑i≥1si<∞\sum_{i\ge1} s_i < \infty∑i≥1​si​<∞.
  • Theorem 2. If sj(0)=0s_j(0) = 0sj​(0)=0 for some jjj, there are constants N≥1N \ge 1N≥1, 0<α<10<\alpha<10<α<1, β>1\beta>1β>1, γ>0\gamma>0γ>0, independent of ttt, with si(t)≤γ αβis_i(t) \le \gamma\,\alpha^{\beta^i}si​(t)≤γαβi for all i≥Ni \ge Ni≥N, t≥0t \ge 0t≥0. From the empty state, si(t)≤πis_i(t) \le \pi_isi​(t)≤πi​ for all iii and t≥0t \ge 0t≥0.
  • Theorem 3. There are weights wi≥1w_i \ge 1wi​≥1 such that Φ(t)=∑i≥1wi∣si(t)−πi∣\Phi(t) = \sum_{i\ge1} w_i |s_i(t) - \pi_i|Φ(t)=∑i≥1​wi​∣si​(t)−πi​∣ satisfies Φ(t)≤c0e−δt\Phi(t) \le c_0 e^{-\delta t}Φ(t)≤c0​e−δt (δ>0\delta > 0δ>0) whenever Φ(0)<∞\Phi(0) < \inftyΦ(0)<∞, and in particular whenever sj(0)=0s_j(0) = 0sj​(0)=0 for some jjj.
  • Corollary 1. Under the conditions of Theorem 3 (weights wi≥1w_i \ge 1wi​≥1 fixed before the trajectory, with Φ(0)<∞\Phi(0) < \inftyΦ(0)<∞, or sj(0)=0s_j(0) = 0sj​(0)=0 for some jjj), ∑i≥1∣si(t)−πi∣≤c0e−δt\sum_{i\ge1} |s_i(t) - \pi_i| \le c_0 e^{-\delta t}∑i≥1​∣si​(t)−πi​∣≤c0​e−δt.
  • Proof of Corollary 2, display. ∑i≥1i(si−1d−sid)=∑i≥0sid\sum_{i\ge1} i(s_{i-1}^d - s_i^d) = \sum_{i\ge0} s_i^d∑i≥1​i(si−1d​−sid​)=∑i≥0​sid​ for a state with si→0s_i \to 0si​→0.
  • Lemma 4. The drift of (1) is Lipschitz in ℓ1\ell^1ℓ1 on states, with constant 2+2dλ2 + 2d\lambda2+2dλ.

Significance

The comparison T1(λ)=1/(1−λ)T_1(\lambda) = 1/(1-\lambda)T1​(λ)=1/(1−λ) against Td(λ)T_d(\lambda)Td​(λ) is the quantitative content of the power of two choices: the companion mission shows Td(λ)/log⁡T1(λ)→1/log⁡dT_d(\lambda)/\log T_1(\lambda) \to 1/\log dTd​(λ)/logT1​(λ)→1/logd as λ→1−\lambda \to 1^-λ→1−, an exponential improvement. Corollary 2 is what connects that number to the dynamics: it says the fixed point's expected time is actually reached from realistic initial states, and is an upper bound from an empty start. Theorem 2's comparison argument and Theorem 3's weighted potential are reused across the mean-field analysis of load-balancing variants (threshold policies, heterogeneous servers, work stealing), so the formal statements here are templates for that literature.

All results are proved in the paper (Theorem 3's proof is written out for d=2d = 2d=2 with a remark that it extends). None has a machine-checked proof that we know of. The work is formalizing the known arguments for all d≥2d \ge 2d≥2, including the parts the paper delegates to a citation: quasimonotone comparison for countable ODE systems in Theorem 2, and the upper right Dini derivative handling of ∣ϵi∣|\epsilon_i|∣ϵi​∣ in Theorem 3.

Difficulty

The system is infinite-dimensional, so standard finite-dimensional ODE comparison and Lyapunov theorems do not apply as stated. Theorem 2 rests on a monotonicity property (raising the initial tails raises them for all time) that the paper justifies by a coupling intuition and a citation to a comparison theorem for quasimonotone systems; that theorem must be supplied for ℓ∞\ell^\inftyℓ∞-bounded countable systems. The plain L1L_1L1​ distance is nonincreasing along trajectories but does not decay exponentially by a direct estimate; the weights wiw_iwi​ must be built so that every coordinate's contribution to dΦ/dtd\Phi/dtdΦ/dt is dominated by −δwi∣ϵi∣-\delta w_i |\epsilon_i|−δwi​∣ϵi​∣, while staying geometrically bounded so that Φ(0)<∞\Phi(0) < \inftyΦ(0)<∞ under the doubly exponential tails of Theorem 2. Exchanging t→∞t \to \inftyt→∞ with the infinite sum defining EEE needs uniform tail control, which is again Theorem 2.

Formalization scope

  • Namespace PowerTwoChoices.Limit; λ is lam : ℝ, d:Nd : ℕd:N with 2≤d2 \le d2≤d, 0<λ<10 < \lambda < 10<λ<1 in every theorem.
  • A state is x : ℕ → ℝ with x 0 = 1, nonnegative and antitone. A trajectory is s : ℝ → ℕ → ℝ, a state for every t≥0t \ge 0t≥0, with HasDerivWithinAt on Set.Ici 0 for each coordinate i≥1i \ge 1i≥1: right derivative at t=0t = 0t=0, two-sided for t>0t > 0t>0.
  • Exponents are natural numbers: πi=λ∑k<idk\pi_i = \lambda^{\sum_{k<i} d^k}πi​=λ∑k<i​dk, and TdT_dTd​'s iii-th term is λ∑1≤k<idk\lambda^{\sum_{1\le k<i} d^k}λ∑1≤k<i​dk.
  • Φ\PhiΦ, the L1L_1L1​ distance, EEE and TdT_dTd​ are sums in [0,∞][0,\infty][0,∞], so a divergent series is ∞\infty∞, never a default 000. "Converges exponentially" (Definition 2) means Φ(t)≤c0e−δt\Phi(t) \le c_0 e^{-\delta t}Φ(t)≤c0​e−δt for all t≥0t \ge 0t≥0 with real c0c_0c0​ and δ>0\delta > 0δ>0; δ\deltaδ may depend on the trajectory. The paper's d(t)d(t)d(t) is renamed l1Dist.
  • Time is real and t→∞t \to \inftyt→∞ is Filter.atTop.
  • Every theorem quantifies over trajectories; existence of a trajectory from a given state (Picard iteration, cited by the paper) is not part of the mission. The class is nonempty: the constant trajectory at π\piπ is one.
  • Ruled out: summing EEE or TdT_dTd​ as a real tsum (a divergent series would make the upper bound free and the limit attainable by a junk 000), dropping the summability condition in Lemma 2 ((1,1,… )(1,1,\dots)(1,1,…) is also a fixed point), choosing Theorem 3's weights after the trajectory, and stating anything for d=2d = 2d=2 only.
  • Needed infrastructure: comparison principles for quasimonotone countable ODE systems, Dini-derivative Grönwall estimates, and interchange of limits and ℓ1\ell^1ℓ1 sums. All of these are reusable beyond this mission, and contributions of such lemmas as separate theorems are welcome.

Selected references

  • M. Mitzenmacher, The Power of Two Choices in Randomized Load Balancing, IEEE Transactions on Parallel and Distributed Systems 12(10), 2001, pp. 1094–1104. https://doi.org/10.1109/71.963420
  • Y. Azar, A. Z. Broder, A. R. Karlin, E. Upfal, Balanced Allocations, SIAM Journal on Computing 29(1), 1999, pp. 180–200. https://doi.org/10.1137/S0097539795288490
  • N. D. Vvedenskaya, R. L. Dobrushin, F. I. Karpelevich, Queueing System with Selection of the Shortest of Two Queues: An Asymptotic Approach, Problems of Information Transmission 32(1), 1996, pp. 15–27.
  • T. G. Kurtz, Approximation of Population Processes, SIAM, 1981. https://doi.org/10.1137/1.9781611970333
8 thms1 active userReviewed
ProbabilityStochastic Systems·Captain: mikedeng1

Is Network Traffic Approximated by Stable Lévy Motion or Fractional Brownian Motion? 4: Superposed ON/OFF Input Under Fast Growth Converges in C[0,∞) to Fractional Brownian MotionResearch Paper

Why heavy-tailed input matters

Measurements of Ethernet and Internet traffic in the 1990s showed that cumulative traffic is self-similar and long-range dependent: correlations of the input rate decay so slowly that they are not summable, and fluctuations at large time scales do not average out the way Poisson-type models predict (Leland et al. 1994). A widely accepted explanation is that the lengths of individual transmissions (file sizes, session durations) are heavy tailed, with infinite variance. Queueing and capacity-planning calculations then depend on which stochastic process approximates the cumulative input over long horizons.

Two approximations had been proposed: fractional Brownian motion, a Gaussian self-similar process with dependent increments, and α-stable Lévy motion, a heavy-tailed process with independent increments. Mikosch, Resnick, Rootzén and Stegeman (Ann. Appl. Probab. 12 (2002) 23–68) showed that both arise from the same models and that the answer depends on how fast the number of sources grows relative to the time scale. This mission formalizes their Gaussian answer for the superposition of ON/OFF sources: Theorem 4.

  • 1995–1997: Willinger, Taqqu, Sherman and Wilson derive fractional Brownian motion from superposed ON/OFF sources with heavy-tailed periods, as an iterated limit: first the number of sources M→∞M\to\inftyM→∞, then the time scale T→∞T\to\inftyT→∞, and for finite-dimensional distributions only.
  • 1998: Heath, Resnick and Samorodnitsky (Math. Oper. Res. 23) give the exact decay of the covariance of a single stationary ON/OFF source.
  • 2002: Mikosch, Resnick, Rootzén and Stegeman let MMM and TTT grow together and show that fast growth gives fractional Brownian motion as a functional limit in C[0,∞)\mathbb C[0,\infty)C[0,∞), while slow growth gives stable Lévy motion.

The superposed ON/OFF model

A single ON/OFF source alternates between ON-periods, during which it sends work at rate 111, and silent OFF-periods. The ON-periods X1,X2,…X_1,X_2,\dotsX1​,X2​,… are iid with law FonF_{\mathrm{on}}Fon​ and the OFF-periods Yoff,Y1,Y2,…Y_{\mathrm{off}},Y_1,Y_2,\dotsYoff​,Y1​,Y2​,… iid with law FoffF_{\mathrm{off}}Foff​, both on [0,∞)[0,\infty)[0,∞), all independent, with means μon\mu_{\mathrm{on}}μon​, μoff\mu_{\mathrm{off}}μoff​ and μ=μon+μoff\mu=\mu_{\mathrm{on}}+\mu_{\mathrm{off}}μ=μon​+μoff​. The tails are regularly varying: for x>0x>0x>0,

Fˉon(x)=x−αLon(x),Fˉoff(x)=x−αoffLoff(x),1<α<αoff<2,\bar F_{\mathrm{on}}(x)=x^{-\alpha}L_{\mathrm{on}}(x),\qquad \bar F_{\mathrm{off}}(x)=x^{-\alpha_{\mathrm{off}}}L_{\mathrm{off}}(x),\qquad 1<\alpha<\alpha_{\mathrm{off}}<2,Fˉon​(x)=x−αLon​(x),Fˉoff​(x)=x−αoff​Loff​(x),1<α<αoff​<2,

with Lon,LoffL_{\mathrm{on}},L_{\mathrm{off}}Lon​,Loff​ slowly varying (L(cx)/L(x)→1L(cx)/L(x)\to1L(cx)/L(x)→1 for every c>0c>0c>0). So both periods have finite mean and infinite variance, and the ON-periods have the heavier tail.

To make the source stationary, a delay T0=B(Xon(0)+Yoff)+(1−B)Yoff(0)T_0=B(X^{(0)}_{\mathrm{on}}+Y_{\mathrm{off}})+(1-B)Y^{(0)}_{\mathrm{off}}T0​=B(Xon(0)​+Yoff​)+(1−B)Yoff(0)​ is used, where BBB is Bernoulli with P(B=1)=μon/μP(B=1)=\mu_{\mathrm{on}}/\muP(B=1)=μon​/μ and Xon(0),Yoff(0)X^{(0)}_{\mathrm{on}},Y^{(0)}_{\mathrm{off}}Xon(0)​,Yoff(0)​ have the integrated-tail laws F(0)(x)=μF−1∫0xFˉ(s) dsF^{(0)}(x)=\mu_F^{-1}\int_0^x\bar F(s)\,dsF(0)(x)=μF−1​∫0x​Fˉ(s)ds. With Tn=T0+∑i=1n(Xi+Yi)T_n=T_0+\sum_{i=1}^n(X_i+Y_i)Tn​=T0​+∑i=1n​(Xi​+Yi​) the source's activity is

Wt=B 1[0,Xon(0))(t)+∑n≥01[Tn,Tn+Xn+1)(t),t≥0,W_t=B\,\mathbf 1_{[0,X^{(0)}_{\mathrm{on}})}(t)+\sum_{n\ge0}\mathbf 1_{[T_n,T_n+X_{n+1})}(t),\qquad t\ge0,Wt​=B1[0,Xon(0)​)​(t)+n≥0∑​1[Tn​,Tn​+Xn+1​)​(t),t≥0,

a stationary process with EWt=μon/μEW_t=\mu_{\mathrm{on}}/\muEWt​=μon​/μ.

The TTT-th model superposes M=M(T)M=M(T)M=M(T) independent copies W(1),…,W(M)W^{(1)},\dots,W^{(M)}W(1),…,W(M), where MMM is integer valued, non-decreasing and M(T)→∞M(T)\to\inftyM(T)→∞. The cumulative input is A(t)=∫0t∑m=1MWs(m) dsA(t)=\int_0^t\sum_{m=1}^MW^{(m)}_s\,dsA(t)=∫0t​∑m=1M​Ws(m)​ds. With the quantile b(t)=(1/Fˉon)←(t)b(t)=(1/\bar F_{\mathrm{on}})^{\leftarrow}(t)b(t)=(1/Fˉon​)←(t), the fast growth condition is

Condition 2:lim⁡T→∞b(MT)T=∞,\text{Condition 2:}\qquad \lim_{T\to\infty}\frac{b(MT)}{T}=\infty,Condition 2:T→∞lim​Tb(MT)​=∞,

equivalently MTFˉon(T)→∞MT\bar F_{\mathrm{on}}(T)\to\inftyMTFˉon​(T)→∞. The normalisation and limit constants are

dT=[T3−αLon(T)M]1/2,σ02=2μoff2Γ(2−α)/(α−1)μ3Γ(4−α),H=3−α2∈(12,1).d_T=[T^{3-\alpha}L_{\mathrm{on}}(T)M]^{1/2},\qquad \sigma_0^2=\frac{2\mu_{\mathrm{off}}^2\Gamma(2-\alpha)/(\alpha-1)}{\mu^3\Gamma(4-\alpha)},\qquad H=\frac{3-\alpha}{2}\in(\tfrac12,1).dT​=[T3−αLon​(T)M]1/2,σ02​=μ3Γ(4−α)2μoff2​Γ(2−α)/(α−1)​,H=23−α​∈(21​,1).

A standard fractional Brownian motion BHB_HBH​ is a mean-zero Gaussian process on [0,∞)[0,\infty)[0,∞) with continuous paths and Cov(BH(t),BH(s))=12(t2H+s2H−∣t−s∣2H)\mathrm{Cov}(B_H(t),B_H(s))=\tfrac12(t^{2H}+s^{2H}-|t-s|^{2H})Cov(BH​(t),BH​(s))=21​(t2H+s2H−∣t−s∣2H).

Formalization targets

Goal: Theorem 4

If Condition 2 holds, then

A(T⋅)−TMμ−1μon(⋅)dT →d σ0BH(⋅)weakly in C[0,∞).\frac{A(T\cdot)-TM\mu^{-1}\mu_{\mathrm{on}}(\cdot)}{d_T}\ \xrightarrow{d}\ \sigma_0B_H(\cdot)\qquad\text{weakly in }\mathbb C[0,\infty).dT​A(T⋅)−TMμ−1μon​(⋅)​ d​ σ0​BH​(⋅)weakly in C[0,∞).

Intermediate targets

In the order of the paper's argument: the mean EWt=μon/μEW_t=\mu_{\mathrm{on}}/\muEWt​=μon​/μ (p. 27); the covariance decay (2.5) γW(h)∼μoff2(α−1)μ3h−(α−1)Lon(h)\gamma_W(h)\sim\frac{\mu_{\mathrm{off}}^2}{(\alpha-1)\mu^3}h^{-(\alpha-1)}L_{\mathrm{on}}(h)γW​(h)∼(α−1)μ3μoff2​​h−(α−1)Lon​(h); Condition 2 ⇔T=o(dT)\Leftrightarrow T=o(d_T)⇔T=o(dT​) (p. 61); the variance asymptotic (7.1) Var(GT)∼σ02T3−αLon(T)\mathrm{Var}(G_T)\sim\sigma_0^2T^{3-\alpha}L_{\mathrm{on}}(T)Var(GT​)∼σ02​T3−αLon​(T) for GT=∫0T(Wu−EWu) duG_T=\int_0^T(W_u-EW_u)\,duGT​=∫0T​(Wu​−EWu​)du; Lemma 13, the one-dimensional limit dT−1∑mGTt(m)→N(0,σ02t3−α)d_T^{-1}\sum_mG^{(m)}_{Tt}\to N(0,\sigma_0^2t^{3-\alpha})dT−1​∑m​GTt(m)​→N(0,σ02​t3−α); the covariance limit (7.4); and the second-moment bound E∣dT−1∑mGTu(m)∣2≤c u1+εE|d_T^{-1}\sum_mG^{(m)}_{Tu}|^2\le c\,u^{1+\varepsilon}E∣dT−1​∑m​GTu(m)​∣2≤cu1+ε behind tightness (p. 63). (2.5) and (7.1) are cited by the paper from Heath–Resnick–Samorodnitsky and Willinger–Taqqu–Sherman–Wilson; they are stated here as the paper states them.

Significance

Theorem 4 justifies fractional Brownian motion as the model of aggregate traffic from many heavy-tailed sources in a single, simultaneous limit, replacing the iterated limit of earlier work, and it upgrades finite-dimensional convergence to weak convergence of paths. Functionals of the path, such as the supremum of A(Tt)−ctA(Tt)-ctA(Tt)−ct that governs the buffer content of a fluid queue, therefore inherit Gaussian approximations. Together with Theorem 2 it shows that a single growth condition on MMM decides between Gaussian and stable approximations.

The result is proved in the paper; it is not formalized anywhere. A formalization adds a machine-checked statement of the theorem and reusable developments of regularly varying functions, the stationary alternating renewal process, fractional Brownian motion, and convergence in C[0,∞)\mathbb C[0,\infty)C[0,∞).

Difficulty

The obvious route is a central limit theorem for ∑mGTt(m)\sum_{m}G^{(m)}_{Tt}∑m​GTt(m)​, a sum of MMM iid bounded terms; but the number of terms and their distribution change with TTT, so one needs a triangular-array limit theorem whose variance condition rests on the exact asymptotic (7.1). That asymptotic in turn needs the covariance decay (2.5) of a single stationary source, which requires renewal theory for the alternating process with infinite-variance periods. Convergence of finite-dimensional distributions alone does not give the theorem: tightness in C[0,K]\mathbb C[0,K]C[0,K] needs a moment bound on increments that is uniform in TTT for small time lags, which requires Potter-type bounds for the regularly varying function x↦EGx2x\mapsto EG_x^2x↦EGx2​.

Formalization scope

  • Time and models. T→∞T\to\inftyT→∞ along atTop on R\mathbb RR. Each model TTT has its own probability space carrying M(T)M(T)M(T) independent sources; the theorems hold for every such family. Lean's X n, Y n are the paper's Xn+1X_{n+1}Xn+1​, Yn+1Y_{n+1}Yn+1​ and Lean's sources 0,…,M−10,\dots,M-10,…,M−1 are the paper's 1,…,M1,\dots,M1,…,M. Time in the normalised process runs over R≥0\mathbb R_{\ge0}R≥0​.
  • Standing hypotheses. Fon,FoffF_{\mathrm{on}},F_{\mathrm{off}}Fon​,Foff​ are probability measures on [0,∞)[0,\infty)[0,∞) (non-negativity is implicit in "lengths"); (2.1) is stated as "x↦xαFˉ(x)x\mapsto x^{\alpha}\bar F(x)x↦xαFˉ(x) is slowly varying"; 1<α<αoff<21<\alpha<\alpha_{\mathrm{off}}<21<α<αoff​<2; MMM non-decreasing with M→∞M\to\inftyM→∞ (§3.2); Condition 2 where §7 assumes it. No other hypothesis is added.
  • Junk values ruled out. bbb is an infimum over {x>0:tFˉon(x)≤1}\{x>0:t\bar F_{\mathrm{on}}(x)\le1\}{x>0:tFˉon​(x)≤1}, never a division by zero; the series defining WtW_tWt​ has non-negative terms and fails to converge only on a null event; equalities of laws come with measurability. The limit is pinned completely: Gaussian, mean zero, the covariance above with σH=1\sigma_H=1σH​=1 and H=(3−α)/2H=(3-\alpha)/2H=(3−α)/2, continuous paths, multiplied by σ0\sigma_0σ0​. Fidi convergence in place of the functional limit, a free scale in the limit, or a limit identified only by its marginals would each be a different, weaker theorem.
  • Weak convergence is stated in coupling form: for every sequence Tn→∞T_n\to\inftyTn​→∞ there are one probability space, copies YnY_nYn​ with continuous paths of the laws of the normalised processes at TnT_nTn​, and a standard fractional Brownian motion Y′Y'Y′ such that Yn→σ0Y′Y_n\to\sigma_0Y'Yn​→σ0​Y′ uniformly on every [0,K][0,K][0,K] almost surely. On the Polish space C[0,∞)\mathbb C[0,\infty)C[0,∞) this is equivalent to weak convergence.
  • Infrastructure needed: Karamata's theorem and Potter bounds; the stationary alternating renewal process and its covariance asymptotics; a triangular-array central limit theorem; existence of fractional Brownian motion; a moment criterion for tightness in C[0,K]\mathbb C[0,K]C[0,K]. The regular-variation and path-space material is reusable well beyond this mission; contributions of any of these pieces, and of any milestone, are welcome.

Selected references

  • T. Mikosch, S. Resnick, H. Rootzén and A. Stegeman, Is network traffic approximated by stable Lévy motion or fractional Brownian motion?, Ann. Appl. Probab. 12(1) (2002), 23–68. https://doi.org/10.1214/aoap/1015961155
  • D. Heath, S. Resnick and G. Samorodnitsky, Heavy tails and long range dependence in on/off processes and associated fluid models, Math. Oper. Res. 23 (1998), 145–165. https://doi.org/10.1287/moor.23.1.145
  • W. Willinger, M. S. Taqqu, R. Sherman and D. V. Wilson, Self-similarity through high-variability: statistical analysis of Ethernet LAN traffic at the source level, IEEE/ACM Trans. Networking 5 (1997), 71–86. https://doi.org/10.1109/90.554723
  • W. E. Leland, M. S. Taqqu, W. Willinger and D. V. Wilson, On the self-similar nature of Ethernet traffic (extended version), IEEE/ACM Trans. Networking 2 (1994), 1–15. https://doi.org/10.1109/90.282603
  • N. H. Bingham, C. M. Goldie and J. L. Teugels, Regular Variation, Cambridge University Press, 1987. https://doi.org/10.1017/CBO9780511721434
  • P. Billingsley, Convergence of Probability Measures, Wiley, 1968 (Theorem 12.3, moment criterion for tightness).
9 thms1 active userReviewed
ProbabilityStochastic Systems·Captain: mikedeng1

Designing a Call Center with Impatient Customers II: In the Halfin–Whitt Regime the Scaled Erlang-A Queue Converges to a Diffusion with Piecewise-Linear DriftResearch Paper

Motivation

Telephone call centers are large service systems: hundreds of agents answer callers who wait in a queue and hang up when their patience runs out. Staffing them is a trade-off between the cost of agents and the quality of service, and in large centers both the efficiency (agents busy almost all the time) and the service level (most callers answered almost immediately) can be high at once. Garnett, Mandelbaum and Reiman (M&SOM 4(3), 2002) study the simplest model that captures customer abandonment, the Erlang-A or M/M/N+MM/M/N+MM/M/N+M queue, and derive approximations for it in the regime where the number of agents is large.

Their Theorem 2 is the process-level justification of these approximations: the queue-length process, centred at the number of agents and scaled by its square root, converges to a one-dimensional diffusion. The steady-state performance formulas of the paper (delay probability, abandonment probability, waiting times) are computed from this diffusion and its stationary law.

Timeline.

  • Halfin and Whitt (1981) proved the corresponding limit for the M/M/NM/M/NM/M/N queue without abandonment, in the regime N(1−ρN)→β\sqrt N(1-\rho_N) \to \betaN​(1−ρN​)→β with β>0\beta > 0β>0 (Oper. Res. 29(3)). The limit is a diffusion with drift −μ(β+x)-\mu(\beta + x)−μ(β+x) below 000 and constant drift −μβ-\mu\beta−μβ above.
  • Fleming, Stolyar and Simon (1994) conjectured the limit with abandonment, with a slightly different centering, and proved the weak limit of the stationary distributions.
  • Garnett, Mandelbaum and Reiman (2002) state the process limit for 0<θ<∞0 < \theta < \infty0<θ<∞ as Theorem 2, with the extensions θ=0\theta = 0θ=0 and θ=∞\theta = \inftyθ=∞ as Theorem 2* in Appendix C.

Setting

Fix a service rate μ>0\mu > 0μ>0. For each N≥1N \ge 1N≥1 the NNN-th system has NNN statistically identical agents, Poisson arrivals of rate λN>0\lambda_N > 0λN​>0, exponential service times of rate μ\muμ, an unlimited waiting room served first come, first served, and exponential patience of rate θN>0\theta_N > 0θN​>0 for every caller. A caller whose wait in queue reaches their patience abandons and does not return. The number of callers in the system, QN(t)Q_N(t)QN​(t), is a birth–death process on {0,1,2,… }\{0, 1, 2, \dots\}{0,1,2,…} with birth rate λN\lambda_NλN​ and death rate min⁡(k,N)μ+(k−N)+θN\min(k, N)\mu + (k-N)^+\theta_Nmin(k,N)μ+(k−N)+θN​ in state kkk. Its initial value QN(0)Q_N(0)QN​(0) is an arbitrary random variable.

The traffic intensity is ρN=λN/(Nμ)\rho_N = \lambda_N/(N\mu)ρN​=λN​/(Nμ). The centred and scaled process is

qN(t)=QN(t)−NN,t≥0.q_N(t) = \frac{Q_N(t) - N}{\sqrt N}, \qquad t \ge 0 .qN​(t)=N​QN​(t)−N​,t≥0.

When qN≥0q_N \ge 0qN​≥0 it counts waiting callers and when qN≤0q_N \le 0qN​≤0 it counts idle agents, both in units of N\sqrt NN​.

For β∈R\beta \in \mathbb Rβ∈R and θ>0\theta > 0θ>0, the drift is the piecewise-linear function

f(x)={−μ(β+x),x≤0,−(μβ+θx),x>0,f(x) = \begin{cases} -\mu(\beta + x), & x \le 0,\\ -(\mu\beta + \theta x), & x > 0,\end{cases}f(x)={−μ(β+x),−(μβ+θx),​x≤0,x>0,​

and the limit diffusion qqq solves dq(t)=f(q(t)) dt+2μ db(t)dq(t) = f(q(t))\,dt + \sqrt{2\mu}\,db(t)dq(t)=f(q(t))dt+2μ​db(t), where bbb is a standard Brownian motion independent of q(0)q(0)q(0). The noise is additive, so this means: qqq has continuous paths and q(t)=q(0)+∫0tf(q(s)) ds+2μ b(t)q(t) = q(0) + \int_0^t f(q(s))\,ds + \sqrt{2\mu}\,b(t)q(t)=q(0)+∫0t​f(q(s))ds+2μ​b(t) for all t≥0t \ge 0t≥0. Below 000 it behaves like an Ornstein–Uhlenbeck process with restraining force μ\muμ, above 000 like one with restraining force θ\thetaθ.

Formalization targets

Goal: Theorem 2 (p. 216)

Assume

lim⁡N→∞N(1−ρN)=β∈(−∞,∞),lim⁡N→∞θN=θ∈(0,∞),\lim_{N\to\infty}\sqrt N(1-\rho_N) = \beta \in (-\infty,\infty), \qquad \lim_{N\to\infty}\theta_N = \theta \in (0,\infty),N→∞lim​N​(1−ρN​)=β∈(−∞,∞),N→∞lim​θN​=θ∈(0,∞),

and qN(0)⇒νq_N(0) \Rightarrow \nuqN​(0)⇒ν for a probability law ν\nuν on R\mathbb RR. Then

qN⇒qin D[0,∞),q_N \Rightarrow q \quad\text{in } D[0,\infty),qN​⇒qin D[0,∞),

where qqq is the solution of dq=f(q) dt+2μ dbdq = f(q)\,dt + \sqrt{2\mu}\,dbdq=f(q)dt+2μ​db with q(0)∼νq(0) \sim \nuq(0)∼ν, and this solution is unique in law.

Milestone: the limit equation is well posed (Theorem 2, p. 216)

On every probability space with a standard Brownian motion bbb and an independent initial value X0X_0X0​, the equation has a solution adapted to σ(X0,b(s):s≤t)\sigma(X_0, b(s): s \le t)σ(X0​,b(s):s≤t), and any two solutions with the same bbb and initial value are indistinguishable.

Milestone: infinitesimal moments (Appendix C, p. 224)

The infinitesimal expectation μN(x)\mu_N(x)μN​(x) and variance σN2(x)\sigma_N^2(x)σN2​(x) of qNq_NqN​ at the point xxx, as displayed on p. 224, satisfy

lim⁡N→∞μN(x)=f(x),lim⁡N→∞σN2(x)=2μ(x∈R).\lim_{N\to\infty}\mu_N(x) = f(x), \qquad \lim_{N\to\infty}\sigma_N^2(x) = 2\mu \qquad (x \in \mathbb R).N→∞lim​μN​(x)=f(x),N→∞lim​σN2​(x)=2μ(x∈R).

Significance

The result. Theorem 2 is the basis for the paper's QED (quality- and efficiency-driven) approximations QN≈N+qNQ_N \approx N + q\sqrt NQN​≈N+qN​, which give explicit formulas for the probability of delay, the probability of abandonment and the waiting-time distribution in terms of β\betaβ, μ\muμ and θ\thetaθ. Unlike the M/M/NM/M/NM/M/N case, β\betaβ may be of either sign: abandonment keeps the system stable even when the offered load exceeds the number of agents. The resulting square-root staffing rule N=R+βRN = R + \beta\sqrt RN=R+βR​, with R=λ/μR = \lambda/\muR=λ/μ, is used in call-center practice.

Formalizing it. The result is proved in the paper by Stone's criteria for birth–death processes, with uniqueness of the limit equation delegated to the literature. No many-server process limit exists on Prove2Me or in Mathlib. A formal proof would give the first machine-checked heavy-traffic process limit for a many-server queue, and the infrastructure (Poisson time-change representation of a birth–death process, path-space weak convergence to a continuous diffusion, well-posedness of a Lipschitz additive-noise SDE) would serve every other diffusion limit in queueing theory.

Difficulty

The infinitesimal-moment computation is elementary; it identifies the limit but does not prove convergence. The hard step is the passage from generator convergence to weak convergence of processes on D[0,∞)D[0,\infty)D[0,∞): tightness of {qN}\{q_N\}{qN​} in the Skorokhod space, which needs control of the jumps (of size 1/N1/\sqrt N1/N​) and of the unbounded state space, and identification of every limit point as a solution of the equation, which needs uniqueness of the limit. A pointwise limit of the drifts at each fixed xxx is not enough: the convergence has to be controlled locally uniformly along the paths. A finite-dimensional limit at fixed times is weaker than the theorem.

Formalization scope

  • Queue. QNQ_NQN​ is realized by the time-change representation of Mandelbaum, Massey and Reiman (1998), cited by the paper: QN(t)=QN(0)+A(λNt)−S(μ∫0tmin⁡(QN(s),N) ds)−R(θN∫0t(QN(s)−N)+ ds)Q_N(t) = Q_N(0) + A(\lambda_N t) - S(\mu\int_0^t \min(Q_N(s), N)\,ds) - R(\theta_N \int_0^t (Q_N(s)-N)^+\,ds)QN​(t)=QN​(0)+A(λN​t)−S(μ∫0t​min(QN​(s),N)ds)−R(θN​∫0t​(QN​(s)−N)+ds), with A,S,RA, S, RA,S,R independent unit-rate Poisson processes independent of QN(0)Q_N(0)QN​(0). This has the birth–death law above.
  • Time. Paths are functions on R\mathbb RR; every condition is imposed for t≥0t \ge 0t≥0 only. Brownian time is R≥0\mathbb R_{\ge 0}R≥0​.
  • Indexing. NNN is both the index and the number of agents; every hypothesis on the NNN-th system is imposed for N≥1N \ge 1N≥1. λN\lambda_NλN​, θN\theta_NθN​ are real sequences and μ\muμ is fixed.
  • Weak convergence. All QNQ_NQN​ live on one probability space, which is no loss since only their laws matter. qN(0)⇒νq_N(0) \Rightarrow \nuqN​(0)⇒ν is convergence of E g(qN(0))E\,g(q_N(0))Eg(qN​(0)) for every bounded continuous ggg. qN⇒qq_N \Rightarrow qqN​⇒q is stated in coupling form (Skorokhod representation with almost-sure uniform convergence on compact time intervals), which is equivalent to J1J_1J1​ weak convergence when the limit is continuous.
  • Limit. The SDE is pathwise (Lebesgue integral, no Itô integral). Brownian motion is Mathlib's IsBrownianReal with every path continuous and every coordinate measurable. q(0)q(0)q(0) is independent of bbb, which is implicit in the paper. The goal asserts existence of a solution with qN⇒qq_N \Rightarrow qqN​⇒q and uniqueness of its law; dropping the uniqueness clause, fixing ν\nuν to a point mass, restricting to β>0\beta > 0β>0, μ=1\mu = 1μ=1, θ=μ\theta = \muθ=μ, or the stationary case is a different theorem.
  • Infinitesimal moments. Defined exactly as displayed, with ⌊⋅⌋\lfloor\cdot\rfloor⌊⋅⌋ the integer part. The display is printed for θ=0\theta = 0θ=0; the milestone is its 0<θ<∞0 < \theta < \infty0<θ<∞ instance with the limit fff of Theorem 2, pointwise in xxx.
  • Not covered. The cases θ=0\theta = 0θ=0 and θ=∞\theta = \inftyθ=∞ of Theorem 2* and the interchange of limits (Part 2).

Contributions welcome: a Poisson time-change existence theorem, tightness criteria in D[0,∞)D[0,\infty)D[0,∞), well-posedness of Lipschitz SDEs with additive noise, and the martingale-problem approach to identifying limits.

Selected references

  • O. Garnett, A. Mandelbaum, M. Reiman, Designing a Call Center with Impatient Customers, Manufacturing & Service Operations Management 4(3):208–227, 2002. https://doi.org/10.1287/msom.4.3.208.7753
  • S. Halfin, W. Whitt, Heavy-Traffic Limits for Queues with Many Exponential Servers, Operations Research 29(3):567–588, 1981. https://doi.org/10.1287/opre.29.3.567
  • P. J. Fleming, A. Stolyar, B. Simon, Heavy Traffic Limit for a Mobile Phone System Loss Model, Proc. 2nd Int. Conf. on Telecommunication Systems Modeling and Analysis, 1994.
  • A. Mandelbaum, W. A. Massey, M. I. Reiman, Strong Approximations for Markovian Service Networks, Queueing Systems 30:149–201, 1998. https://doi.org/10.1023/A:1019112920622
  • S. N. Ethier, T. G. Kurtz, Markov Processes: Characterization and Convergence, Wiley, 1986. https://doi.org/10.1002/9780470316658
9 thms1 active userReviewed
ProbabilityStochastic Systems·Captain: mikedeng1

Revenue Management Without Forecasting or Optimization: An Adaptive Algorithm for Determining Airline Seat Protection Levels: Fill-Event Updates Converge to the Optimal Protection LevelsResearch Paper

Motivation

Airlines sell seats on a single flight in several fare classes. Discount classes book first, so the carrier must decide how many seats to protect for later, higher-paying passengers. The classical answer, from Littlewood's rule for two classes to the nested optimality conditions of Brumelle and McGill (1993), computes optimal protection levels from a forecast of the demand distribution of each class. Those forecasts must be built from censored booking data, which is where most of the practical difficulty lies.

van Ryzin and McGill (2000) observed that the optimality conditions themselves can drive an adaptive rule. After each flight one only records whether a fill event occurred (whether demand reached each protection level) and moves the levels by a stochastic-approximation step. No forecast and no optimization is needed. Their Theorem 1 asserts that this rule converges almost surely to the optimal protection levels, with an explicit mean-square rate. It is one of the standard references for model-free revenue management. The convergence proof combines Robbins–Monro theory, the Robbins–Siegmund supermartingale lemma and an induction over fare classes.

Setting

There are k+1k+1k+1 fare classes with fares f1>f2>⋯>fk+1>0f_1 > f_2 > \cdots > f_{k+1} > 0f1​>f2​>⋯>fk+1​>0 and discount ratios ri+1=fi+1/f1r_{i+1} = f_{i+1}/f_1ri+1​=fi+1​/f1​. On a flight the class demands X1,…,Xk+1X_1, \dots, X_{k+1}X1​,…,Xk+1​ are independent, nonnegative and continuously distributed. Successive flights X1,X2,…X^1, X^2, \dotsX1,X2,… are independent with the same law.

A protection vector is θ=(θ1,…,θk)\theta = (\theta_1, \dots, \theta_k)θ=(θ1​,…,θk​). The fill events are

Ai(θ,X)={X1>θ1, X1+X2>θ2, …, X1+⋯+Xi>θi},i=1,…,k.A_i(\theta, X) = \{X_1 > \theta_1,\ X_1 + X_2 > \theta_2,\ \dots,\ X_1 + \cdots + X_i > \theta_i\}, \qquad i = 1, \dots, k.Ai​(θ,X)={X1​>θ1​, X1​+X2​>θ2​, …, X1​+⋯+Xi​>θi​},i=1,…,k.

A vector θ∗\theta^*θ∗ is characterised by the optimality condition (3), P(Ai(θ∗,X))=ri+1P(A_i(\theta^*, X)) = r_{i+1}P(Ai​(θ∗,X))=ri+1​ for i=1,…,ki = 1, \dots, ki=1,…,k. The adjustments are Hi(θ,X)=ri+1−1(Ai(θ,X))H_i(\theta, X) = r_{i+1} - \mathbf 1(A_i(\theta, X))Hi​(θ,X)=ri+1​−1(Ai​(θ,X)), and their means are hi(θ)=ri+1−P(Ai(θ,X))h_i(\theta) = r_{i+1} - P(A_i(\theta, X))hi​(θ)=ri+1​−P(Ai​(θ,X)). From an arbitrary θ1\theta^1θ1 the algorithm updates

θn+1=θn−γnH(θn,Xn),γn=An+B, A>0, B≥0.\theta^{n+1} = \theta^n - \gamma_n H(\theta^n, X^n), \qquad \gamma_n = \frac{A}{n + B},\ A > 0,\ B \ge 0.θn+1=θn−γn​H(θn,Xn),γn​=n+BA​, A>0, B≥0.

The interim protection levels pi(θ)=max⁡{θj:1≤j≤i}p_i(\theta) = \max\{\theta_j : 1 \le j \le i\}pi​(θ)=max{θj​:1≤j≤i} are what the airline actually uses to control bookings. The updates themselves monitor Ai(θn,Xn)A_i(\theta^n, X^n)Ai​(θn,Xn).

Formalization targets

Goal: Theorem 1 (p. 765)

Assume:

  • A1: every XiX_iXi​ has bounded support.
  • A2 (windowed): for each R>0R > 0R>0 there is δR>0\delta_R > 0δR​>0 with (θi−θi∗) hi(θi,θi−1∗,…,θ1∗)≥δR∣θi−θi∗∣2(\theta_i - \theta^*_i)\,h_i(\theta_i, \theta^*_{i-1}, \dots, \theta^*_1) \ge \delta_R|\theta_i - \theta^*_i|^2(θi​−θi∗​)hi​(θi​,θi−1∗​,…,θ1∗​)≥δR​∣θi​−θi∗​∣2 whenever ∣θi−θi∗∣≤R|\theta_i - \theta^*_i| \le R∣θi​−θi∗​∣≤R.
  • A3: the distribution functions of the partial sums X1+⋯+XiX_1 + \cdots + X_iX1​+⋯+Xi​ are Lipschitz.

Then for i=1,…,ki = 1, \dots, ki=1,…,k,

θin→θi∗ a.s.,pi(θ∗)=θi∗,\theta^n_i \to \theta^*_i \ \text{a.s.}, \qquad p_i(\theta^*) = \theta^*_i,θin​→θi∗​ a.s.,pi​(θ∗)=θi∗​,

and there are β>0\beta > 0β>0 and CCC with

E∣θin−θi∗∣2≤C γnβ/2i−1.E|\theta^n_i - \theta^*_i|^2 \le C\,\gamma_n^{\beta/2^{i-1}}.E∣θin​−θi∗​∣2≤Cγnβ/2i−1​.

The exponent is left existential, as in the paper, so any improvement of the rate constant still proves the goal.

Milestones

The milestones follow the paper's own proof, in attack order:

  • the base case i=1i = 1i=1 (p. 765);
  • the Lipschitz bound on hi+1h_{i+1}hi+1​ (p. 766);
  • Lemma 3, boundedness of the iterates (p. 764);
  • the almost-supermartingale inequality (12) (p. 766);
  • Lemma 1 (p. 764);
  • the a.s. summability of γn∣θin−θi∗∣\gamma_n|\theta^n_i - \theta^*_i|γn​∣θin​−θi∗​∣ (p. 766);
  • Lemma 2, Robbins–Siegmund (p. 764);
  • the induction step for (9) (p. 766);
  • Lemma 4, the rate recursion (p. 764).

Significance

Theorem 1 says that the nested optimality conditions of Brumelle and McGill can be reached by observing binary fill events alone. It also says that the limit is automatically ordered, θ1∗≤⋯≤θk∗\theta^*_1 \le \cdots \le \theta^*_kθ1∗​≤⋯≤θk∗​, so the interim levels coincide with the limit. The algorithm is distribution-free and needs neither demand forecasts nor uncensoring, which is why it is used in practice and as a baseline for later data-driven and learning-based revenue management.

The theorem is proved in the paper. As far as is known it has no machine-checked proof. Formalizing it requires a vector Robbins–Monro argument with coupled coordinates, the Robbins–Siegmund lemma, and an explicit polynomial rate. Two of the milestones are general probability facts reusable far beyond this mission: Lemma 1, and Robbins–Siegmund, which no platform item states yet. The formalization also corrects the record on two printed slips, described under Formalization scope.

Difficulty

Coordinate i+1i+1i+1 is not a Robbins–Monro process by itself. Its fill event Ai+1(θ,X)A_{i+1}(\theta, X)Ai+1​(θ,X) depends on all of θ1,…,θi+1\theta_1, \dots, \theta_{i+1}θ1​,…,θi+1​, so the drift of θi+1n\theta^n_{i+1}θi+1n​ is perturbed by the errors of the lower coordinates. Applying a scalar convergence theorem coordinate by coordinate fails at this point. The perturbation has to be bounded via the Lipschitz assumption A3 and shown to be summable along the path, which uses the mean-square rate already established for class iii. The rate and the almost-sure convergence must therefore be carried through the induction together. The exponent halves at each class. A second difficulty is that A2 is only usable on a bounded set, so boundedness of the iterates (Lemma 3) has to come first and must be uniform.

Formalization scope

Classes are indexed from 111 in Lean (N\mathbb NN-indexed vectors whose index 000 and indices above k+1k+1k+1 are never read). A flight's demand vector has the product law ⨂iνi\bigotimes_i \nu_i⨂i​νi​, and the sample path has the countable product of copies of it. Lean's iterate … n is the paper's θn+1\theta^{n+1}θn+1, so the rate pairs it with γn+1\gamma_{n+1}γn+1​. Continuity of a demand law is "every singleton is null". Expectations are Bochner integrals, and integrability is part of each rate conclusion. The fill event reuses the published NestedSeatAlloc.ProbCond.nestEvent.

Pinned and corrected statements:

  • A2 as printed is unsatisfiable. It asks for one δ\deltaδ valid for all θi\theta_iθi​. Since ∣hi∣≤1|h_i| \le 1∣hi​∣≤1, the inequality fails once ∣θi−θi∗∣>1/δ|\theta_i - \theta^*_i| > 1/\delta∣θi​−θi∗​∣>1/δ, so the printed Theorem 1 is vacuously true. The mission uses the windowed A2 above. That is what the proof applies to bounded iterates, and what the paper's sufficient condition (a density bounded below near θ∗\theta^*θ∗) yields. A formalization with the printed A2, or with any other hypothesis no instance satisfies, is ruled out.
  • The base case. The page claims the rate "for all 0<β<10 < \beta < 10<β<1". That is false in general. The milestone states "for some β>0\beta > 0β>0", which is all Theorem 1 uses.
  • Integrability. Lemma 2 assumes each ZnZ_nZn​ integrable, so that the conditional expectation is genuine. Lemma 3's "bounded (a.s.)" is read as a bound uniform in nnn.
  • Standing assumptions. Nonnegative demands and positive fares are added where the argument uses them.

Not posed: §4's modifications (projection onto the capacity, randomized integer levels, booking lead times) and the simulations of §5.

Contributions welcome: proofs of the general lemmas (Lemma 1, Robbins–Siegmund) stated for arbitrary probability spaces, measurability facts for the recursion, and the elementary Lemma 4.

Selected references

  • G. van Ryzin and J. McGill, Revenue Management Without Forecasting or Optimization: An Adaptive Algorithm for Determining Airline Seat Protection Levels, Management Science 46(6), 760–775, 2000. https://doi.org/10.1287/mnsc.46.6.760.11936
  • S. L. Brumelle and J. I. McGill, Airline Seat Allocation with Multiple Nested Fare Classes, Operations Research 41(1), 127–137, 1993. https://doi.org/10.1287/opre.41.1.127
  • H. Robbins and D. Siegmund, A convergence theorem for non negative almost supermartingales and some applications, in Optimizing Methods in Statistics, Academic Press, 233–257, 1971. https://doi.org/10.1016/B978-0-12-604550-5.50015-8
  • A. Benveniste, M. Métivier and P. Priouret, Adaptive Algorithms and Stochastic Approximations, Springer, 1990. https://doi.org/10.1007/978-3-642-75894-2
  • E. Lukacs, Stochastic Convergence, 2nd ed., Academic Press, 1975.
12 thms1 active userReviewed
Experimental DesignProbabilityStatistics+1·Captain: mikedeng1

Stochastic Kriging for Simulation Metamodeling II: The MSE of the Optimal Stochastic Kriging Predictor Exceeds the Kriging MSE by a Positive Definite FormResearch Paper

Motivation

A simulation metamodel is a cheap surrogate for an expensive stochastic simulation: after running the simulation at a few design settings, one predicts the mean response at settings never simulated. Kriging, the interpolation method of geostatistics and of the design and analysis of computer experiments (DACE; Sacks, Welch, Mitchell and Wynn 1989; Santner, Williams and Notz 2003), treats the unknown response surface as a realization of a random field and predicts by the best linear predictor. Kriging was developed for deterministic computer codes, where an observation is the exact response. A stochastic simulation returns the response plus sampling noise, and the noise may be correlated across settings when the experimenter uses common random numbers (CRN).

Ankenman, Nelson and Staum, Stochastic Kriging for Simulation Metamodeling (Proc. 2008 Winter Simulation Conference, pp. 362–370), extend kriging to this setting. The mission formalizes §2 of that paper: the form of the optimal linear predictor, called stochastic kriging, and the formula for its mean squared error, which shows how much the simulation noise costs. The source used throughout is the WSC 2008 proceedings paper, not the later Operations Research 58(2) 2010 article, whose numbering differs.

Setting

Fix a probability space (Ω,F,P)(\Omega,\mathcal F,P)(Ω,F,P) and a set of design settings x\mathbf xx. On replication j=1,2,…j=1,2,\dotsj=1,2,… at setting x\mathbf xx the simulation outputs (display (3), with constant trend)

Yj(x)=β0+M(x)+εj(x).\mathcal Y_j(\mathbf x)=\beta_0+\mathsf M(\mathbf x)+\varepsilon_j(\mathbf x).Yj​(x)=β0​+M(x)+εj​(x).

The constant β0\beta_0β0​ is the overall mean. M\mathsf MM is a mean-zero random field, the extrinsic uncertainty imposed by the modeller. εj(x)\varepsilon_j(\mathbf x)εj​(x) is mean-zero sampling noise, the intrinsic uncertainty of the simulation; its variance may depend on x\mathbf xx, and noises at different settings may be correlated (CRN). All of these have finite second moments, and the field is uncorrelated with the noise.

An experiment design runs ni≥1n_i\ge1ni​≥1 replications at each of kkk design points x1,…,xk\mathbf x_1,\dots,\mathbf x_kx1​,…,xk​. The data are the sample means (4)

Yˉ(xi)=1ni∑j=1niYj(xi),Yˉ=(Yˉ(x1),…,Yˉ(xk))⊤,\bar{\mathcal Y}(\mathbf x_i)=\frac1{n_i}\sum_{j=1}^{n_i}\mathcal Y_j(\mathbf x_i),\qquad\bar{\mathcal Y}=(\bar{\mathcal Y}(\mathbf x_1),\dots,\bar{\mathcal Y}(\mathbf x_k))^\top,Yˉ​(xi​)=ni​1​j=1∑ni​​Yj​(xi​),Yˉ​=(Yˉ​(x1​),…,Yˉ​(xk​))⊤,

and the target at a point x0\mathbf x_0x0​ is the noise-free response Y(x0)=β0+M(x0)\mathsf Y(\mathbf x_0)=\beta_0+\mathsf M(\mathbf x_0)Y(x0​)=β0​+M(x0​).

Write ΣM(x,x′)=Cov[M(x),M(x′)]\Sigma_{\mathsf M}(\mathbf x,\mathbf x')=\mathrm{Cov}[\mathsf M(\mathbf x),\mathsf M(\mathbf x')]ΣM​(x,x′)=Cov[M(x),M(x′)]. Then ΣM\Sigma_{\mathsf M}ΣM​ is the k×kk\times kk×k matrix of these covariances at the design points, ΣM(x0,⋅)\Sigma_{\mathsf M}(\mathbf x_0,\cdot)ΣM​(x0​,⋅) is the vector (Cov[M(x0),M(xi)])i(\mathrm{Cov}[\mathsf M(\mathbf x_0),\mathsf M(\mathbf x_i)])_{i}(Cov[M(x0​),M(xi​)])i​, and Σε\Sigma_\varepsilonΣε​ is the k×kk\times kk×k covariance matrix of the averaged noises ∑j=1niεj(xi)/ni\sum_{j=1}^{n_i}\varepsilon_j(\mathbf x_i)/n_i∑j=1ni​​εj​(xi​)/ni​. A linear predictor (5) is λ0+λ⊤Yˉ\lambda_0+\lambda^\top\bar{\mathcal Y}λ0​+λ⊤Yˉ​ with arbitrary weights (λ0,λ)∈R×Rk(\lambda_0,\lambda)\in\mathbb R\times\mathbb R^k(λ0​,λ)∈R×Rk; its mean squared error is E[(λ0+λ⊤Yˉ−Y(x0))2]\mathrm E[(\lambda_0+\lambda^\top\bar{\mathcal Y}-\mathsf Y(\mathbf x_0))^2]E[(λ0​+λ⊤Yˉ​−Y(x0​))2]. The optimal MSE MSE⋆\mathrm{MSE}^\starMSE⋆ is the infimum of this quantity over all weights.

Formalization targets

Goal: display (7)

There is a rule Ξ\XiΞ, assigning a positive definite matrix Ξ(ΣM,Σε)\Xi(\Sigma_{\mathsf M},\Sigma_\varepsilon)Ξ(ΣM​,Σε​) to every pair of positive definite matrices, such that in every model as above with ΣM≻0\Sigma_{\mathsf M}\succ0ΣM​≻0 and Σε≻0\Sigma_\varepsilon\succ0Σε​≻0, and at every x0\mathbf x_0x0​,

MSE⋆=ΣM(x0,x0)−ΣM(x0,⋅)⊤[ΣM+Σε]−1ΣM(x0,⋅)=[ΣM(x0,x0)−ΣM(x0,⋅)⊤ΣM−1ΣM(x0,⋅)]+ΣM(x0,⋅)⊤ Ξ ΣM(x0,⋅).\begin{aligned}\mathrm{MSE}^\star&=\Sigma_{\mathsf M}(\mathbf x_0,\mathbf x_0)-\Sigma_{\mathsf M}(\mathbf x_0,\cdot)^\top[\Sigma_{\mathsf M}+\Sigma_\varepsilon]^{-1}\Sigma_{\mathsf M}(\mathbf x_0,\cdot)\\&=\Big[\Sigma_{\mathsf M}(\mathbf x_0,\mathbf x_0)-\Sigma_{\mathsf M}(\mathbf x_0,\cdot)^\top\Sigma_{\mathsf M}^{-1}\Sigma_{\mathsf M}(\mathbf x_0,\cdot)\Big]+\Sigma_{\mathsf M}(\mathbf x_0,\cdot)^\top\,\Xi\,\Sigma_{\mathsf M}(\mathbf x_0,\cdot).\end{aligned}MSE⋆​=ΣM​(x0​,x0​)−ΣM​(x0​,⋅)⊤[ΣM​+Σε​]−1ΣM​(x0​,⋅)=[ΣM​(x0​,x0​)−ΣM​(x0​,⋅)⊤ΣM−1​ΣM​(x0​,⋅)]+ΣM​(x0​,⋅)⊤ΞΣM​(x0​,⋅).​

The bracket is the usual kriging MSE. The goal also states that MSE⋆\mathrm{MSE}^\starMSE⋆ strictly exceeds it whenever ΣM(x0,⋅)≠0\Sigma_{\mathsf M}(\mathbf x_0,\cdot)\neq0ΣM​(x0​,⋅)=0. The matrix Ξ\XiΞ is existential and depends only on the two covariance matrices, not on x0\mathbf x_0x0​.

Milestone: display (6)

If ΣM+Σε≻0\Sigma_{\mathsf M}+\Sigma_\varepsilon\succ0ΣM​+Σε​≻0, the stochastic kriging predictor

Y^(x0)=β0+ΣM(x0,⋅)⊤[ΣM+Σε]−1(Yˉ−β01k)\widehat{\mathsf Y}(\mathbf x_0)=\beta_0+\Sigma_{\mathsf M}(\mathbf x_0,\cdot)^\top[\Sigma_{\mathsf M}+\Sigma_\varepsilon]^{-1}(\bar{\mathcal Y}-\beta_0\mathbf 1_k)Y(x0​)=β0​+ΣM​(x0​,⋅)⊤[ΣM​+Σε​]−1(Yˉ​−β0​1k​)

is of the form (5), attains the smallest MSE among all predictors of that form, and is the only one that does.

Further items (not milestones)

  • (9)–(10): the two-point example, with closed forms for the predictor and its MSE in terms of τ2\tau^2τ2, r12r_{12}r12​, r0r_0r0​, V\mathsf VV, ρ\rhoρ and nnn.
  • (14): the equicorrelated example, MSE⋆=τ2(1−kr02/(1+(k−1)r+γ/n))\mathrm{MSE}^\star=\tau^2\big(1-k r_0^2/(1+(k-1)r+\gamma/n)\big)MSE⋆=τ2(1−kr02​/(1+(k−1)r+γ/n)).
  • §3, p. 365: under the Gaussian Assumption 1, Y^(x0)=E[Y(x0)∣Yˉ]\widehat{\mathsf Y}(\mathbf x_0)=\mathrm E[\mathsf Y(\mathbf x_0)\mid\bar{\mathcal Y}]Y(x0​)=E[Y(x0​)∣Yˉ​] almost surely.

Significance

Display (7) quantifies the price of simulation noise. With exact observations the best achievable error is the kriging MSE; with noisy observations it is larger by a positive definite form in the cross-covariance vector. Every point x0\mathbf x_0x0​ that is correlated with the design therefore loses accuracy. The two examples make this concrete. In (10) the MSE increases with the intrinsic correlation ρ\rhoρ, which is why the paper concludes that common random numbers, a standard variance-reduction device for comparing systems, do not help prediction. Display (14) is the baseline against which the paper measures the cost of estimating the noise variance. The conditional-expectation statement shows that under Gaussian assumptions stochastic kriging is optimal among all predictors, not only linear ones.

The results are classical in substance: they are the best-linear-prediction formulas of second-order random-field theory, applied to data whose covariance is ΣM+Σε\Sigma_{\mathsf M}+\Sigma_\varepsilonΣM​+Σε​. The paper states them without proof ("we can show"). No machine-checked version exists. Formalizing them yields a reusable, measure-theoretic best-linear-predictor theorem for random vectors defined from a model rather than postulated, together with the Loewner-order fact behind the positive definiteness of Ξ\XiΞ.

Difficulty

The algebra is short; the difficulty is in deriving the second-moment structure from the model, not assuming it. The covariance of Yˉ\bar{\mathcal Y}Yˉ​ must be computed from the sample-mean definition, using bilinearity of covariance over finite sums of square-integrable variables, and shown to equal ΣM+Σε\Sigma_{\mathsf M}+\Sigma_\varepsilonΣM​+Σε​. The cross-covariance with Y(x0)\mathsf Y(\mathbf x_0)Y(x0​) must be shown to equal ΣM(x0,⋅)\Sigma_{\mathsf M}(\mathbf x_0,\cdot)ΣM​(x0​,⋅). The MSE of an arbitrary affine predictor then has to be written as a quadratic in (λ0,λ)(\lambda_0,\lambda)(λ0​,λ), and its infimum identified.

A first idea is to take "a random vector with mean β01\beta_0\mathbf 1β0​1 and covariance Σ\SigmaΣ" as the hypothesis. That proves a different, more abstract statement and leaves the model's content unproved.

For the second line of (7), the positive definiteness of Ξ\XiΞ does not follow from a scalar argument. It is a matrix inequality, (A+B)−1≺A−1(A+B)^{-1}\prec A^{-1}(A+B)−1≺A−1 for A,B≻0A,B\succ0A,B≻0, and it fails without Σε≻0\Sigma_\varepsilon\succ0Σε​≻0: when Σε=0\Sigma_\varepsilon=0Σε​=0 the two lines coincide.

Formalization scope

  • The model. Settings form an arbitrary type X; design points are x : Fin k → X, so the paper's index iii is i.val + 1. Replication j≥1j\ge1j≥1 is index j - 1 of a sequence ε : ℕ → X → Ω → ℝ.
  • Moments and covariances. Field and noise are in L2L^2L2 and have mean zero. Covariances are Mathlib's covariance. The optimal MSE is an infimum over all of R×Rk\mathbb R\times\mathbb R^kR×Rk, with no unbiasedness constraint and no sign constraint.
  • Hypotheses not printed in §2. Three are added and disclosed:
    1. uncorrelatedness of M\mathsf MM and ε\varepsilonε, the second-order content of Assumption 1's "independent of M\mathsf MM", without which the covariance of Yˉ\bar{\mathcal Y}Yˉ​ is not ΣM+Σε\Sigma_{\mathsf M}+\Sigma_\varepsilonΣM​+Σε​;
    2. square integrability;
    3. ΣM≻0\Sigma_{\mathsf M}\succ0ΣM​≻0 and Σε≻0\Sigma_\varepsilon\succ0Σε​≻0 in the goal. Lean's matrix inverse returns 000 for a singular matrix, and Ξ≻0\Xi\succ0Ξ≻0 is false when Σε\Sigma_\varepsilonΣε​ is singular.
  • Generality kept. No Gaussianity, no independence across replications and no "no CRN" restriction is imposed on (6) and (7).
  • Examples. ∣r12∣<1|r_{12}|<1∣r12​∣<1 in (10) and r≤1r\le1r≤1 in (14) are added, since both are correlations.
  • Non-trivializing. A positive semidefinite Ξ\XiΞ, or a Ξ\XiΞ chosen after x0\mathbf x_0x0​, would make the goal trivial or strictly weaker; the statement fixes Ξ\XiΞ as a function of (ΣM,Σε)(\Sigma_{\mathsf M},\Sigma_\varepsilon)(ΣM​,Σε​) before the model and the point.

Infrastructure needed and reusable beyond this mission:

  • covariance of averages of L2L^2L2 random variables;
  • the MSE of an affine predictor as a quadratic form;
  • minimization of a positive definite quadratic;
  • the Loewner antitonicity of the inverse;
  • for the conditional-expectation item, Gaussian conditioning via uncorrelated-hence-independent residuals.

Contributions to any of these are welcome.

Selected references

  • B. Ankenman, B. L. Nelson, J. Staum, Stochastic Kriging for Simulation Metamodeling, Proceedings of the 2008 Winter Simulation Conference, IEEE, pp. 362–370, 2008. http://www.informs-sim.org/wsc08papers/042.pdf
  • B. Ankenman, B. L. Nelson, J. Staum, Stochastic Kriging for Simulation Metamodeling, Operations Research 58(2), 371–382, 2010 (later journal version, not used here). https://doi.org/10.1287/opre.1090.0754
  • J. Sacks, W. J. Welch, T. J. Mitchell, H. P. Wynn, Design and Analysis of Computer Experiments, Statistical Science 4(4), 409–423, 1989. https://doi.org/10.1214/ss/1177012413
  • T. J. Santner, B. J. Williams, W. I. Notz, The Design and Analysis of Computer Experiments, Springer, 2003. https://doi.org/10.1007/978-1-4757-3799-8
4 thms1 active userReviewed
Experimental DesignProbabilityStatistics+1·Captain: mikedeng1

Stochastic Kriging for Simulation Metamodeling I: The Plug-In Stochastic Kriging Predictor Is UnbiasedResearch Paper

Motivation

Stochastic simulation models of queues, supply chains and manufacturing systems are often too slow to run at every input setting of interest. A metamodel fitted to a moderate number of simulation runs then stands in for the simulation, for instance inside an optimization or a sensitivity analysis. Kriging, developed in geostatistics and in the design and analysis of deterministic computer experiments (DACE), predicts an unknown response surface by treating it as a realization of a Gaussian random field. Deterministic kriging interpolates its data exactly, which is wrong for stochastic simulation, where each run returns the response plus sampling noise.

Ankenman, Nelson and Staum's stochastic kriging separates the two sources of uncertainty: the extrinsic uncertainty of the unknown surface and the intrinsic uncertainty of the simulation output. Its optimal predictor requires the intrinsic noise variances, which are never known in practice and are replaced by sample variances computed from the same replications. This mission formalizes the paper's first key result: this plug-in step introduces no prediction bias.

The source is the Winter Simulation Conference 2008 proceedings version of the paper (pp. 362–370). The later Operations Research 58(2) (2010) article numbers its results differently and is not the reference here.

Setting

Fix a probability space (Ω,F,P)(\Omega,\mathcal F,P)(Ω,F,P) and design variables x∈Rd\mathbf x\in\mathbb R^dx∈Rd. On replication j=1,2,…j=1,2,\dotsj=1,2,… at setting x\mathbf xx the simulation returns

Yj(x)=β0+M(x)+εj(x),\mathcal Y_j(\mathbf x)=\beta_0+\mathsf M(\mathbf x)+\varepsilon_j(\mathbf x),Yj​(x)=β0​+M(x)+εj​(x),

where β0∈R\beta_0\in\mathbb Rβ0​∈R is a constant trend, M\mathsf MM is a mean-zero random field on Rd\mathbb R^dRd (the unknown surface) and εj(x)\varepsilon_j(\mathbf x)εj​(x) is the noise of replication jjj. The quantity to predict at any point x0\mathbf x_0x0​, simulated or not, is the noise-free response Y(x0)=β0+M(x0)\mathsf Y(\mathbf x_0)=\beta_0+\mathsf M(\mathbf x_0)Y(x0​)=β0​+M(x0​).

An experiment design is a list of pairs (xi,ni)(\mathbf x_i,n_i)(xi​,ni​), i=1,…,ki=1,\dots,ki=1,…,k, with distinct design points xi\mathbf x_ixi​ and nin_ini​ replications at xi\mathbf x_ixi​. The data are summarized by the sample means Yˉ(xi)=1ni∑j=1niYj(xi)\bar{\mathcal Y}(\mathbf x_i)=\frac1{n_i}\sum_{j=1}^{n_i}\mathcal Y_j(\mathbf x_i)Yˉ​(xi​)=ni​1​∑j=1ni​​Yj​(xi​), collected in the vector Yˉ∈Rk\bar{\mathcal Y}\in\mathbb R^kYˉ​∈Rk, and the sample variances

S2(xi)=1ni−1∑j=1ni(Yj(xi)−Yˉ(xi))2.\mathcal S^2(\mathbf x_i)=\frac1{n_i-1}\sum_{j=1}^{n_i}\big(\mathcal Y_j(\mathbf x_i)-\bar{\mathcal Y}(\mathbf x_i)\big)^2 .S2(xi​)=ni​−11​j=1∑ni​​(Yj​(xi​)−Yˉ​(xi​))2.

The extrinsic covariances are the k×kk\times kk×k matrix ΣM=(Cov[M(xh),M(xi)])h,i\Sigma_{\mathsf M}=(\mathrm{Cov}[\mathsf M(\mathbf x_h),\mathsf M(\mathbf x_i)])_{h,i}ΣM​=(Cov[M(xh​),M(xi​)])h,i​ and the vector ΣM(x0,⋅)=(Cov[M(x0),M(xi)])i\Sigma_{\mathsf M}(\mathbf x_0,\cdot)=(\mathrm{Cov}[\mathsf M(\mathbf x_0),\mathsf M(\mathbf x_i)])_{i}ΣM​(x0​,⋅)=(Cov[M(x0​),M(xi​)])i​.

Assumption 1. M\mathsf MM is a stationary Gaussian random field: all its finite-dimensional laws are multivariate normal with mean 000, the covariance is τ2R(x−x′)\tau^2R(\mathbf x-\mathbf x')τ2R(x−x′) with τ2>0\tau^2>0τ2>0 and R(0)=1R(\mathbf 0)=1R(0)=1, and the covariance matrix at finitely many distinct points is positive definite. At each design point the noises ε1(xi),ε2(xi),…\varepsilon_1(\mathbf x_i),\varepsilon_2(\mathbf x_i),\dotsε1​(xi​),ε2​(xi​),… are i.i.d. N(0,V(xi))N(0,\mathsf V(\mathbf x_i))N(0,V(xi​)) with V(xi)>0\mathsf V(\mathbf x_i)>0V(xi​)>0; noises at different design points are independent (no common random numbers); and the whole noise family is independent of M\mathsf MM.

The intrinsic variance at a design point is estimated by V^(xi)=S2(xi)\widehat{\mathsf V}(\mathbf x_i)=\mathcal S^2(\mathbf x_i)V(xi​)=S2(xi​), giving Σ^ε=Diag{V^(x1)/n1,…,V^(xk)/nk}\widehat\Sigma_\varepsilon=\mathrm{Diag}\{\widehat{\mathsf V}(\mathbf x_1)/n_1,\dots,\widehat{\mathsf V}(\mathbf x_k)/n_k\}Σε​=Diag{V(x1​)/n1​,…,V(xk​)/nk​} and the plug-in stochastic kriging predictor

Y^^(x0)=β0+ΣM(x0,⋅)⊤[ΣM+Σ^ε]−1(Yˉ−β01k).(13)\widehat{\widehat{\mathsf Y}}(\mathbf x_0)=\beta_0+\Sigma_{\mathsf M}(\mathbf x_0,\cdot)^\top\big[\Sigma_{\mathsf M}+\widehat\Sigma_\varepsilon\big]^{-1}\big(\bar{\mathcal Y}-\beta_0\mathbf 1_k\big).\tag{13}Y(x0​)=β0​+ΣM​(x0​,⋅)⊤[ΣM​+Σε​]−1(Yˉ​−β0​1k​).(13)

Formalization targets

Goal: Theorem 1 (p. 366)

Under Assumption 1, with ni≥2n_i\ge2ni​≥2 replications at every design point, the prediction error is integrable and

E[Y^^(x0)−Y(x0)]=0.\mathrm E\Big[\widehat{\widehat{\mathsf Y}}(\mathbf x_0)-\mathsf Y(\mathbf x_0)\Big]=0 .E[Y(x0​)−Y(x0​)]=0.

Milestones (p. 365)

  1. Under Assumption 1, (Y(x0),Yˉ(x1),…,Yˉ(xk))(\mathsf Y(\mathbf x_0),\bar{\mathcal Y}(\mathbf x_1),\dots,\bar{\mathcal Y}(\mathbf x_k))(Y(x0​),Yˉ​(x1​),…,Yˉ​(xk​)) is multivariate normal.
  2. Under Assumption 1, S2(xi)\mathcal S^2(\mathbf x_i)S2(xi​) has a scaled chi-squared distribution: (ni−1)S2(xi)/V(xi)∼χni−12(n_i-1)\mathcal S^2(\mathbf x_i)/\mathsf V(\mathbf x_i)\sim\chi^2_{n_i-1}(ni​−1)S2(xi​)/V(xi​)∼χni​−12​.
  3. Under Assumption 1, S2(xi)\mathcal S^2(\mathbf x_i)S2(xi​) is strongly consistent for V(xi)\mathsf V(\mathbf x_i)V(xi​): the sample variance of the first mmm replications converges to V(xi)\mathsf V(\mathbf x_i)V(xi​) almost surely as m→∞m\to\inftym→∞.

Significance

Theorem 1 says that the practitioner's shortcut, estimating the noise variances from the replications and plugging them into the optimal predictor, keeps the predictor unbiased; the cost of not knowing the intrinsic variance is paid entirely in mean squared error. This is what makes the paper's subsequent comparison of the plug-in MSE with the known-variance MSE the right measure of the penalty for estimation.

None of these statements has a machine-checked proof, and Mathlib has neither the chi-squared law of the normal sample variance nor the independence of the normal sample mean and sample variance. A complete development produces these classical facts of normal sampling theory for the first time in Lean, together with a reusable model of noisy observations of a Gaussian random field.

Difficulty

The predictor is a nonlinear function of the data: its weight vector [ΣM+Σ^ε]−1ΣM(x0,⋅)[\Sigma_{\mathsf M}+\widehat\Sigma_\varepsilon]^{-1}\Sigma_{\mathsf M}(\mathbf x_0,\cdot)[ΣM​+Σε​]−1ΣM​(x0​,⋅) is random and depends on the same replications as Yˉ\bar{\mathcal Y}Yˉ​. Linearity of expectation therefore does not apply directly: the obvious argument, which treats the weights as constants, fails, and the dependence between the estimated variances, the sample means and the field has to be controlled exactly. The joint independence structure of Assumption 1 — across replications, across design points, and between noise and field — is load-bearing; a weaker pairwise or pointwise independence does not suffice. Integrability of the error must also be established, since the weights depend on random variances.

Formalization scope

Design points are indexed by Fin k (the paper's iii is i.val + 1) and lie in EuclideanSpace ℝ (Fin d); replication j≥1j\ge1j≥1 of the paper is index j - 1 of a sequence, so sums run over Finset.range. The field is M : EuclideanSpace ℝ (Fin d) → Ω → ℝ with Mathlib's IsGaussianProcess; covariances are ProbabilityTheory.covariance; stationarity is Cov[M(y),M(y′)]=τ2R(y−y′)\mathrm{Cov}[\mathsf M(\mathbf y),\mathsf M(\mathbf y')]=\tau^2R(\mathbf y-\mathbf y')Cov[M(y),M(y′)]=τ2R(y−y′) with R(0)=1R(\mathbf 0)=1R(0)=1. The noise law is HasLaw … (gaussianReal 0 (V (x i))); independence is one iIndepFun over the family indexed by Fin k × ℕ plus one IndepFun between that whole family and the whole field. The variance function is ℝ≥0-valued with V(xi)>0\mathsf V(\mathbf x_i)>0V(xi​)>0 assumed; replication counts satisfy ni≥2n_i\ge2ni​≥2 wherever S2\mathcal S^2S2 appears and ni≥1n_i\ge1ni​≥1 where only means appear. The chi-squared law χν2\chi^2_\nuχν2​ is gammaMeasure (ν/2) (1/2).

The covariance quantities ΣM\Sigma_{\mathsf M}ΣM​ and ΣM(x0,⋅)\Sigma_{\mathsf M}(\mathbf x_0,\cdot)ΣM​(x0​,⋅) in (13) are the true ones; only Σε\Sigma_\varepsilonΣε​ is estimated, and Σ^ε\widehat\Sigma_\varepsilonΣε​ is built from the same replications as Yˉ\bar{\mathcal Y}Yˉ​. A formalization in which Σ^ε\widehat\Sigma_\varepsilonΣε​ is deterministic or independent of the data, in which the matrix ΣM+Σ^ε\Sigma_{\mathsf M}+\widehat\Sigma_\varepsilonΣM​+Σε​ may be singular (Lean's Matrix.inv returns 000 there, reducing the claim to E[M(x0)]=0\mathrm E[\mathsf M(\mathbf x_0)]=0E[M(x0​)]=0), or in which the expectation is asserted without integrability, would trivialize the goal and is ruled out: positive definiteness of ΣM\Sigma_{\mathsf M}ΣM​ at distinct design points is part of Assumption 1, and integrability is part of the conclusion. x0\mathbf x_0x0​ may coincide with a design point.

A complete development needs: the law of linear images of independent Gaussian vectors; independence of the sample mean and the residual vector for i.i.d. normal samples (Cochran's theorem); the chi-squared law of the normal sample variance; the strong law of large numbers for the sample variance; and a measurability and boundedness argument for the random weight vector. The normal-sampling results are reusable well beyond this mission and are welcome as separate contributions.

Selected references

  • B. Ankenman, B. L. Nelson, J. Staum, Stochastic Kriging for Simulation Metamodeling, Proceedings of the 2008 Winter Simulation Conference, IEEE, pp. 362–370, 2008. http://www.informs-sim.org/wsc08papers/042.pdf
  • B. Ankenman, B. L. Nelson, J. Staum, Stochastic Kriging for Simulation Metamodeling, Operations Research 58(2):371–382, 2010. https://doi.org/10.1287/opre.1090.0754
  • T. J. Santner, B. J. Williams, W. I. Notz, The Design and Analysis of Computer Experiments, Springer, 2003. https://doi.org/10.1007/978-1-4757-3799-8
6 thms1 active userReviewed
Convex OptimizationNumerical AnalysisOptimization·Captain: mikedeng1

A Coordinate Gradient Descent Method for Nonsmooth Separable Minimization 3: The Local Lipschitzian Error Bound Holds for Polyhedral P with Quadratic or Composite f, and for Strongly Convex fResearch Paper

Motivation

Many problems in statistics, signal processing and machine learning minimize the sum of a smooth function and a convex but nonsmooth one: ℓ1\ell_1ℓ1​-regularized least squares (the Lasso), bound-constrained problems, and group-sparse regression are standard examples. Tseng and Yun (Math. Program. Ser. B 117 (2009) 387–423) proposed the coordinate gradient descent (CGD) method for this class and proved that it converges linearly under a local Lipschitzian error bound, Assumption 2(a) of their paper. An assumption is only as useful as the list of problems known to satisfy it. Section 6 of the paper supplies that list, and this mission formalizes it.

Error bounds of this kind have a history in smooth constrained optimization. Luo and Tseng proved them for quadratic objectives over polyhedral sets (Luo–Tseng 1992, SIAM J. Optim.), for strongly convex functions composed with a linear map (Luo–Tseng 1992, SIAM J. Control Optim.), and for dual functionals of linearly constrained strictly convex programs (Luo–Tseng 1993, Math. Oper. Res.), and used them to derive linear rates for feasible descent methods (Luo–Tseng 1993, Ann. Oper. Res.). Tseng and Yun carry these results over to the nonsmooth problem (1) by reformulating it as a smooth problem over the epigraph of the nonsmooth part.

Setting

Throughout, Rn\mathbb R^nRn carries the Euclidean norm ∥⋅∥\|\cdot\|∥⋅∥. The problem is

min⁡x  Fc(x)=f(x)+cP(x),(1)\min_x\; F_c(x) = f(x) + cP(x), \qquad (1)xmin​Fc​(x)=f(x)+cP(x),(1)

where c>0c > 0c>0, P:Rn→(−∞,∞]P:\mathbb R^n \to (-\infty,\infty]P:Rn→(−∞,∞] is proper, convex and lower semicontinuous with effective domain dom⁡P={x∣P(x)<∞}\operatorname{dom}P = \{x \mid P(x) < \infty\}domP={x∣P(x)<∞}, and fff is continuously differentiable on an open set containing dom⁡P\operatorname{dom}PdomP.

The residual at x∈dom⁡Px \in \operatorname{dom}Px∈domP is

dI(x)=arg⁡min⁡d{∇f(x)⊤d+12∥d∥2+cP(x+d)},d_I(x) = \arg\min_d \Big\{\nabla f(x)^\top d + \tfrac12\|d\|^2 + cP(x+d)\Big\},dI​(x)=argdmin​{∇f(x)⊤d+21​∥d∥2+cP(x+d)},

the unique minimizer of a strongly convex function. A point x∈dom⁡Px \in \operatorname{dom}Px∈domP is stationary if the one-sided directional derivative satisfies Fc′(x;d)≥0F_c'(x;d) \ge 0Fc′​(x;d)≥0 for every direction ddd; Xˉ\bar XXˉ denotes the set of stationary points and dist⁡(x,Xˉ)\operatorname{dist}(x,\bar X)dist(x,Xˉ) the distance to it. A point is stationary exactly when its residual vanishes.

Assumption 2(a) asks that Xˉ≠∅\bar X \neq \emptysetXˉ=∅ and that for every level ζ\zetaζ there be constants τ,ϵ>0\tau, \epsilon > 0τ,ϵ>0 with

dist⁡(x,Xˉ)≤τ∥dI(x)∥whenever Fc(x)≤ζ, ∥dI(x)∥≤ϵ.\operatorname{dist}(x,\bar X) \le \tau\|d_I(x)\| \quad\text{whenever } F_c(x) \le \zeta,\ \|d_I(x)\| \le \epsilon.dist(x,Xˉ)≤τ∥dI​(x)∥whenever Fc​(x)≤ζ, ∥dI​(x)∥≤ϵ.

Writing epi⁡P={(x,ξ)∣P(x)≤ξ}\operatorname{epi}P = \{(x,\xi) \mid P(x) \le \xi\}epiP={(x,ξ)∣P(x)≤ξ}, problem (1) is equivalent to the smooth problem min⁡{f(x)+cξ∣(x,ξ)∈epi⁡P}\min\{f(x) + c\xi \mid (x,\xi) \in \operatorname{epi}P\}min{f(x)+cξ∣(x,ξ)∈epiP} (41). Its projection residual at (x,ξ)(x,\xi)(x,ξ) is the optimal solution (d~,δ~)(\tilde d,\tilde\delta)(d~,δ~) of

min⁡(d,δ){∇f(x)⊤d+12∥d∥2+12δ2+cδ  ∣  (x+d,ξ+δ)∈epi⁡P}.(42)\min_{(d,\delta)}\Big\{\nabla f(x)^\top d + \tfrac12\|d\|^2 + \tfrac12\delta^2 + c\delta \;\Big|\; (x+d,\xi+\delta) \in \operatorname{epi}P\Big\}. \qquad (42)(d,δ)min​{∇f(x)⊤d+21​∥d∥2+21​δ2+cδ​(x+d,ξ+δ)∈epiP}.(42)

PPP is polyhedral if epi⁡P\operatorname{epi}PepiP is the solution set of finitely many linear inequalities in (x,ξ)(x,\xi)(x,ξ); fff is quadratic if f(x)=12x⊤Ax+b⊤x+c0f(x) = \tfrac12 x^\top Ax + b^\top x + c_0f(x)=21​x⊤Ax+b⊤x+c0​ with AAA symmetric, not necessarily positive semidefinite. The four problem classes are:

  • C1: fff quadratic, PPP polyhedral;
  • C2: f(x)=g(Ex)+q⊤xf(x) = g(Ex) + q^\top xf(x)=g(Ex)+q⊤x with ggg strongly convex and differentiable on Rm\mathbb R^mRm, ∇g\nabla g∇g Lipschitz, PPP polyhedral;
  • C3: f(x)=max⁡y∈Y{(Ex)⊤y−g(y)}+q⊤xf(x) = \max_{y\in Y}\{(Ex)^\top y - g(y)\} + q^\top xf(x)=maxy∈Y​{(Ex)⊤y−g(y)}+q⊤x with YYY polyhedral and ggg as in C2, PPP polyhedral;
  • C4: fff strongly convex with ∇f\nabla f∇f Lipschitz on dom⁡P\operatorname{dom}PdomP (condition (22)).

Formalization targets

Goal: Theorem 4 (p. 412)

(Xˉ≠∅ ∧ (C1∨C2∨C3)) ∨ C4  ⟹  Assumption 2(a).\Big(\bar X \neq \emptyset \ \wedge\ (\mathrm{C1} \vee \mathrm{C2} \vee \mathrm{C3})\Big) \ \vee\ \mathrm{C4} \;\Longrightarrow\; \text{Assumption 2(a)}.(Xˉ=∅ ∧ (C1∨C2∨C3)) ∨ C4⟹Assumption 2(a).

Under C4, nonemptiness of Xˉ\bar XXˉ is part of the conclusion. The constants τ,ϵ\tau, \epsilonτ,ϵ are existential, so the goal does not depend on any particular estimate.

Milestone: Lemma 6 (p. 410)

If PPP is Lipschitz on dom⁡P\operatorname{dom}PdomP with constant KKK, there is κ>0\kappa > 0κ>0 depending only on KKK with

∥(d~,δ~)∥≤κ ∥dI(x)∥(x∈dom⁡P, ξ=P(x)).\|(\tilde d,\tilde\delta)\| \le \kappa\,\|d_I(x)\| \qquad (x \in \operatorname{dom}P,\ \xi = P(x)).∥(d~,δ~)∥≤κ∥dI​(x)∥(x∈domP, ξ=P(x)).

Milestone: Lemma 7 (p. 411)

If Xˉ≠∅\bar X \neq \emptysetXˉ=∅ and C1, C2 or C3 holds, then for every ζ\zetaζ there are τ′,ϵ′>0\tau', \epsilon' > 0τ′,ϵ′>0 with

dist⁡(x,Xˉ)≤τ′∥(d~,δ~)∥whenever Fc(x)≤ζ, ∥(d~,δ~)∥≤ϵ′,(44)\operatorname{dist}(x,\bar X) \le \tau'\|(\tilde d,\tilde\delta)\| \quad\text{whenever } F_c(x) \le \zeta,\ \|(\tilde d,\tilde\delta)\| \le \epsilon', \qquad (44)dist(x,Xˉ)≤τ′∥(d~,δ~)∥whenever Fc​(x)≤ζ, ∥(d~,δ~)∥≤ϵ′,(44)

where (d~,δ~)(\tilde d,\tilde\delta)(d~,δ~) solves (42) with ξ=P(x)\xi = P(x)ξ=P(x).

A further item, not a milestone, states the remark after Assumption 2 (p. 404) that Assumption 2(b) (stationary points with different objective values are uniformly separated) holds whenever fff is convex.

Significance

Theorem 4 is what makes the linear convergence theorem of the paper (Theorem 2, the subject of the second mission of this series) applicable. Through C1 it covers every problem with a quadratic fff and a polyhedral PPP, including the Lasso and ℓ1\ell_1ℓ1​-regularized or bound-constrained quadratic programs, without convexity of fff. Through C2 it covers losses of the form g(Ex)g(Ex)g(Ex) with ggg strongly convex, such as least squares 12∥Ex−b∥2\tfrac12\|Ex - b\|^221​∥Ex−b∥2, where fff itself is not strongly convex because EEE may have a nontrivial kernel. These error bounds became a standard tool for linear rates of proximal and coordinate methods without strong convexity.

The paper's result is proved, but Lemma 7 is proved by citation: it applies three error bounds of Luo and Tseng to the reformulation (41). A complete formalization therefore requires those error bounds for smooth problems over polyhedral sets, which are not in Mathlib. To our knowledge none of the results of this mission has a machine-checked proof.

Difficulty

Lemma 6 and the C4 case of Theorem 4 are short inequality arguments. The difficulty is Lemma 7. The natural first idea, proving an error bound from strong convexity, fails for C1–C3: fff may be nonconvex (C1) or have a degenerate Hessian (C2, C3), and the objective f(x)+cξf(x) + c\xif(x)+cξ of (41) is never strongly convex in (x,ξ)(x,\xi)(x,ξ). What replaces strong convexity is the polyhedral structure of epi⁡P\operatorname{epi}PepiP, and the cited Luo–Tseng error bounds that exploit it are substantial results in their own right, each of which has to be established for a projection residual over a general polyhedron in Rn+1\mathbb R^{n+1}Rn+1.

Formalization scope

Rn\mathbb R^nRn is EuclideanSpace ℝ (Fin n), with indices 0,…,n−10,\dots,n-10,…,n−1. PPP is the pair (D,P)(D, P)(D,P), its effective domain and its finite values there, with "proper, convex, lsc" given by the published ProxNewton.Inexact.IsProperClosedConvex. The value +∞+\infty+∞ is never computed: every statement quantifies over x∈Dx \in Dx∈D, and membership in epi⁡P\operatorname{epi}PepiP is x + d ∈ D ∧ P (x + d) ≤ ξ + δ. dI(x)d_I(x)dI​(x) is a chosen minimizer of its subproblem, which at x∈Dx \in Dx∈D is the paper's unique one. The solutions of (42) enter through a predicate, and every bound is asserted for every optimal solution. Stationarity is the liminf form of Fc′(x;d)≥0F_c'(x;d) \ge 0Fc′​(x;d)≥0. dist⁡\operatorname{dist}dist is the infimum distance. Assumption 2(a) quantifies over every real ζ\zetaζ, which agrees with the paper's "ζ≥min⁡Fc\zeta \ge \min F_cζ≥minFc​" because the condition is vacuous below inf⁡Fc\inf F_cinfFc​. In Lemma 6 the constant κ\kappaκ is quantified before the dimension and the data, so it depends on the Lipschitz constant only.

Explicit choices relative to the printed text:

  • In C4, strong convexity and (22) are required on dom⁡P\operatorname{dom}PdomP only, since fff is only assumed smooth near dom⁡P\operatorname{dom}PdomP. This hypothesis is weaker than the paper's, so the formal theorem is at least as strong.
  • In C2 and C3, EEE is a continuous linear map Rn→Rm\mathbb R^n \to \mathbb R^mRn→Rm (equivalently an m×nm \times nm×n matrix). Strong convexity has a positive modulus and ∇g\nabla g∇g is globally Lipschitz.
  • In C3, the maximum must be attained at every xxx, which forces Y≠∅Y \neq \emptysetY=∅, as the paper's formula presumes.
  • Assumption 2(b) for convex fff is stated with convexity on dom⁡P\operatorname{dom}PdomP.

A polyhedrality notion that admits only affine or constant PPP, an Assumption 2(a) whose constants are chosen after xxx, and a Lemma 7 whose residual is dI(x)d_I(x)dI​(x) instead of the solution of (42) would all trivialize or change the statements. The encoding rules out each of them: polyhedral means a finite system of linear inequalities in (x,ξ)(x,\xi)(x,ξ), the constants precede xxx, and Lemma 7 is stated for (42).

Welcome contributions: proofs of Lemma 6 and of the C4 case; Lipschitz continuity of polyhedral functions on their domain (Rockafellar–Wets, Example 9.35); the Luo–Tseng error bounds for affine variational inequalities and for composite strongly convex objectives over polyhedra, which can be reused well beyond this mission.

Selected references

  • P. Tseng, S. Yun, A coordinate gradient descent method for nonsmooth separable minimization, Math. Program. Ser. B 117 (2009) 387–423. https://doi.org/10.1007/s10107-007-0170-0
  • Z.-Q. Luo, P. Tseng, Error bounds and the convergence analysis of matrix splitting algorithms for the affine variational inequality problem, SIAM J. Optim. 2 (1992) 43–54. https://doi.org/10.1137/0802004
  • Z.-Q. Luo, P. Tseng, On the linear convergence of descent methods for convex essentially smooth minimization, SIAM J. Control Optim. 30 (1992) 408–425. https://doi.org/10.1137/0330025
  • Z.-Q. Luo, P. Tseng, On the convergence rate of dual ascent methods for linearly constrained convex minimization, Math. Oper. Res. 18 (1993) 846–867. https://doi.org/10.1287/moor.18.4.846
  • Z.-Q. Luo, P. Tseng, Error bounds and convergence analysis of feasible descent methods: a general approach, Ann. Oper. Res. 46 (1993) 157–178. https://doi.org/10.1007/BF02096261
  • R. T. Rockafellar, R. J.-B. Wets, Variational Analysis, Springer, 1998. https://doi.org/10.1007/978-3-642-02431-3
7 thms1 active userReviewed
Dynamic ProgrammingOptimizationProbability·Captain: mikedeng1

On the Structure of Lost-Sales Inventory Models 3: If Demand Increases in the Convex Order, the Optimal Lost-Sales Cost f̄_t(v; φ) Is Nondecreasing in φResearch Paper

Motivation

In the periodic-review inventory model with lost sales, demand that cannot be met from stock is lost rather than backordered. With a positive order lead time LLL the state of the system is a vector of length LLL: on-hand stock plus the L−1L-1L−1 orders still in the pipeline. This model is standard in retail, where a customer facing an empty shelf buys elsewhere. Its optimal policy is not a base-stock policy and depends on the whole pipeline vector, which makes it much harder to analyse than the backorder model.

A basic question for any stochastic inventory model is how its optimal cost responds to uncertainty in demand. For the backorder model it is classical that more variable demand costs more (Song 1994). Zipkin's paper (Oper. Res. 56(4), 2008, 937–944) recasts the lost-sales model in a transformed state in which the optimal cost functions are L♮^\natural♮-convex. Its §5 uses this structure to prove the lost-sales analogue: if demand becomes more variable in the convex order, the optimal cost does not decrease. The paper presents this as new in the lost-sales setting.

Timeline of the structural results the mission rests on:

  • 1958: Karlin and Scarf analyse the lost-sales model with lead time one and show that the optimal order decreases in the on-hand stock, with sensitivity less than one.
  • 1969: Morton extends these monotonicity and bounded-sensitivity properties to general lead times and derives bounds on the optimal order.
  • 2008: Zipkin reproves these properties through L♮^\natural♮-convexity in the transformed state (Theorem 4) and adds the parametric result of §5 (Theorem 11).

Setting

The order lead time is a positive integer LLL. The cost factors are the unit procurement cost ccc, the unit holding cost h^\hat hh^, the unit lost-sales penalty ppp and the discount factor γ\gammaγ. Demands in different periods are independent and nonnegative, with a common law. The paper treats states and orders as continuous.

Transformed state. A state is a vector v=(v0,…,vL−1)v=(v_0,\dots,v_{L-1})v=(v0​,…,vL−1​), where vlv_lvl​ is the stock on hand plus all orders due to arrive lll or more periods hence, and vL=0v_L=0vL​=0 by convention. So v0v_0v0​ is the inventory position and v0−v1v_0-v_1v0​−v1​ is the stock on hand. The state space is

V={v∈RL: v0≥v1≥⋯≥vL−1≥0}.V=\{v\in\mathbb R^L:\ v_0\ge v_1\ge\cdots\ge v_{L-1}\ge0\}.V={v∈RL: v0​≥v1​≥⋯≥vL−1​≥0}.

The action is ζ=−z≤0\zeta=-z\le0ζ=−z≤0, where z≥0z\ge0z≥0 is the order quantity, and eee is the all-ones vector. After demand ddd, the next state is

v+=([v0−v1−d]++v1, v2, …, vL−1, 0)−ζe.v_+=\big([v_0-v_1-d]^++v_1,\ v_2,\ \dots,\ v_{L-1},\ 0\big)-\zeta e.v+​=([v0​−v1​−d]++v1​, v2​, …, vL−1​, 0)−ζe.

Costs and recursion. The end-of-period holding and penalty cost is q^(u)=h^u++pu−\hat q(u)=\hat hu^++pu^-q^​(u)=h^u++pu−, and its expectation given on-hand stock yyy is q^0(y)=E[q^(y−d)]\hat q^0(y)=E[\hat q(y-d)]q^​0(y)=E[q^​(y−d)]. The optimal cost functions satisfy

gˉt(v,ζ)=−γLcζ+q^0(v0−v1)+γE[fˉt+1(v+)],fˉt(v)=min⁡ζ≤0gˉt(v,ζ),\bar g_t(v,\zeta)=-\gamma^Lc\zeta+\hat q^0(v_0-v_1)+\gamma E[\bar f_{t+1}(v_+)],\qquad \bar f_t(v)=\min_{\zeta\le0}\bar g_t(v,\zeta),gˉ​t​(v,ζ)=−γLcζ+q^​0(v0​−v1​)+γE[fˉ​t+1​(v+​)],fˉ​t​(v)=ζ≤0min​gˉ​t​(v,ζ),

with fˉT+L+1=0\bar f_{T+L+1}=0fˉ​T+L+1​=0. This is the paper's recursion (1)–(2), written in the state vvv.

Program (4). Fix a demand ddd. Choosing the next stock level v+v_+v+​, with −d≤v+−v0≤0-d\le v_+-v_0\le0−d≤v+​−v0​≤0 and v+≥v1v_+\ge v_1v+​≥v1​, gives

κˉt(v,ζ∣d)=min⁡v+{h^(v+−v1)+p(v+−v0+d)+γfˉt+1[(v+,v2,…,vL−1,0)−ζe]}.\bar\kappa_t(v,\zeta\mid d)=\min_{v_+}\Big\{\hat h(v_+-v_1)+p(v_+-v_0+d)+\gamma\bar f_{t+1}\big[(v_+,v_2,\dots,v_{L-1},0)-\zeta e\big]\Big\}.κˉt​(v,ζ∣d)=v+​min​{h^(v+​−v1​)+p(v+​−v0​+d)+γfˉ​t+1​[(v+​,v2​,…,vL−1​,0)−ζe]}.

Here v+−v1v_+-v_1v+​−v1​ is the leftover stock and d−(v0−v+)d-(v_0-v_+)d−(v0​−v+​) the unmet demand.

Parametric demand. The demand d(ϕ)d(\phi)d(ϕ) depends on a real parameter ϕ\phiϕ, with law μϕ\mu_\phiμϕ​, and fˉt(v;ϕ)\bar f_t(v;\phi)fˉ​t​(v;ϕ) is the optimal cost under μϕ\mu_\phiμϕ​. A random variable XXX is smaller than YYY in the convex order if E g(X)≤E g(Y)E\,g(X)\le E\,g(Y)Eg(X)≤Eg(Y) for every convex g:R→Rg:\mathbb R\to\mathbb Rg:R→R for which both expectations exist. This forces equal means and a variance that does not decrease.

Formalization targets

Goal: Theorem 11 (p. 941)

If d(ϕ)d(\phi)d(ϕ) is increasing in ϕ\phiϕ with respect to the convex order, then for all ttt and all v∈Vv\in Vv∈V,

ϕ≤ϕ′ ⟹ fˉt(v;ϕ)≤fˉt(v;ϕ′).\phi\le\phi'\ \Longrightarrow\ \bar f_t(v;\phi)\le\bar f_t(v;\phi').ϕ≤ϕ′ ⟹ fˉ​t​(v;ϕ)≤fˉ​t​(v;ϕ′).

Milestones, in attack order

  1. Reformulation (proof of Theorem 4, p. 939). For v∈Vv\in Vv∈V, ζ≤0\zeta\le0ζ≤0, d≥0d\ge0d≥0, the stock level v+=[v0−v1−d]++v1v_+=[v_0-v_1-d]^++v_1v+​=[v0​−v1​−d]++v1​ is optimal in program (4):
κˉt(v,ζ∣d)=q^(v0−v1−d)+γfˉt+1(v+).\bar\kappa_t(v,\zeta\mid d)=\hat q(v_0-v_1-d)+\gamma\bar f_{t+1}(v_+).κˉt​(v,ζ∣d)=q^​(v0​−v1​−d)+γfˉ​t+1​(v+​).
  1. Convexity of the optimal cost (Theorem 4, p. 939, with the remark on p. 938). fˉt\bar f_tfˉ​t​ is convex on VVV for every ttt.
  2. The key step (proof of Theorem 11, p. 941). If FFF is convex and nonnegative on VVV, then program (4) with continuation FFF is jointly convex in (v,ζ,d)(v,\zeta,d)(v,ζ,d) on {v∈V, ζ≤0, d≥0}\{v\in V,\ \zeta\le0,\ d\ge0\}{v∈V, ζ≤0, d≥0}.
  3. The induction step (proof of Theorem 11, p. 941). If fˉt+1(v;ϕ)\bar f_{t+1}(v;\phi)fˉ​t+1​(v;ϕ) is nondecreasing in ϕ\phiϕ on VVV, then so is gˉt(v,ζ;ϕ)\bar g_t(v,\zeta;\phi)gˉ​t​(v,ζ;ϕ) for every ζ≤0\zeta\le0ζ≤0.

Significance

Theorem 11 is a comparative-statics statement: replacing demand by a mean-preserving spread cannot lower the optimal cost from any starting state, at any horizon. It justifies valuing variance reduction in lost-sales systems, for example forecasting or demand pooling, without solving the high-dimensional dynamic program. The induction pattern of the proof transfers to other parametric questions about the model; §6 of the paper notes that the parametric analysis remains valid in several extensions.

The paper proves Theorem 11 in a few lines on top of Theorem 4. No machine-checked proof of the result or of its structural prerequisites is known: neither the lost-sales recursion in the transformed state nor its convexity is formalized. A complete development checks the measure-theoretic steps the paper leaves implicit: that the optimal costs are finite, the existence of the expectations, and the passage from convexity in ddd on [0,∞)[0,\infty)[0,∞) to the convex order's test functions on R\mathbb RR.

Difficulty

The argument looks immediate: the costs are convex, so the convex order should apply directly. It does not. The function to which the convex order must be applied is the end-of-period cost as a function of demand. In recursion (1) that function is q^(v0−v1−d)+γfˉt+1(v+(d))\hat q(v_0-v_1-d)+\gamma\bar f_{t+1}(v_+(d))q^​(v0​−v1​−d)+γfˉ​t+1​(v+​(d)), where the next state depends on ddd through the kink [v0−v1−d]+[v_0-v_1-d]^+[v0​−v1​−d]+. Convexity of fˉt+1\bar f_{t+1}fˉ​t+1​ alone does not make this composition convex in ddd. The paper avoids the issue by passing to program (4), where the amount of demand filled is a decision. That requires two facts: that selling as much as possible is optimal (milestone 1), and that the program is jointly convex in state, action and demand (milestone 3), which needs convexity of fˉt+1\bar f_{t+1}fˉ​t+1​ on all of VVV (milestone 2). A second difficulty is analytic: the optimal costs are defined by infima and expectations, and one must show that these are finite and integrable before any order comparison applies.

Formalization scope

  • Model. States are Fin L → ℝ with 0-based indices matching the paper. vL=0v_L=0vL​=0 is the helper vext. VVV is the set of antitone, nonnegative vectors.
  • Time. Data are stationary, so the optimal cost is indexed by the number of periods to go kkk, with fˉ=0\bar f=0fˉ​=0 at k=0k=0k=0. The paper's "for all ttt" is "for all kkk".
  • Demand. One-period demand has law μ\muμ, a probability measure on R\mathbb RR with μ((−∞,0))=0\mu((-\infty,0))=0μ((−∞,0))=0 and finite mean. The finite mean is not in the paper; it is added so that q^0\hat q^0q^​0 is finite.
  • Parametric family. A family μϕ\mu_\phiμϕ​, ϕ∈R\phi\in\mathbb Rϕ∈R, with every hypothesis imposed for every ϕ\phiϕ. Every object depending on demand takes the law as an argument, so fˉt(v;ϕ)\bar f_t(v;\phi)fˉ​t​(v;ϕ) is fbar … (μ φ) k v.
  • Costs. "Unit cost" and "discount rate" are read as c,h^,p≥0c,\hat h,p\ge0c,h^,p≥0 and 0<γ≤10<\gamma\le10<γ≤1.
  • Minima. "Min" over orders is an infimum over ζ≤0\zeta\le0ζ≤0, and program (4) is an infimum over the interval [max⁡(v0−d,v1),v0][\max(v_0-d,v_1),v_0][max(v0​−d,v1​),v0​]. Lean's infimum is 000 on sets that are unbounded below or empty. Nonnegative costs and demand keep the infima genuine, and the generic key-step milestone assumes the continuation is nonnegative.
  • Expectations. Bochner integrals.
  • Convex order. The platform definition StochasticOrders.Convex.ConvexOrder (Shaked–Shanthikumar (3.A.1)), applied to the identity on (R,μϕ)(\mathbb R,\mu_\phi)(R,μϕ​). Equal means are a consequence and are not assumed.
  • Convexity milestone. Stated for the model's fˉ\bar ffˉ​ only, not as "L♮^\natural♮-convex implies convex" in general, which fails without regularity.

Ruled out. Value functions that collapse to a junk constant would make Theorem 11 trivially true. This is excluded because the recursion is defined, not assumed, and every cost and demand hypothesis is imposed for every ϕ\phiϕ. Replacing the convex order by "variance nondecreasing", by the increasing convex order, or by test functions convex only on [0,∞)[0,\infty)[0,∞) changes the theorem and is not accepted.

Infrastructure and contributions. A complete development needs:

  • finiteness and integrability of the optimal costs;
  • convexity of partial infima over convex fibers;
  • monotone extension of a function convex on [0,∞)[0,\infty)[0,∞) to all of R\mathbb RR.

The model file is shared in shape with the series' missions 1 (L♮^\natural♮-convexity) and 2 (policy bounds), and its convexity lemmas are reusable for other parametric comparisons of the lost-sales model. Proofs of any milestone, and alternative arguments that bypass program (4), are welcome.

Selected references

  • P. Zipkin, On the structure of lost-sales inventory models, Operations Research 56(4), 937–944, 2008. https://doi.org/10.1287/opre.1070.0482
  • S. Karlin, H. Scarf, Inventory models of the Arrow–Harris–Marschak type with time lag, in Studies in the Mathematical Theory of Inventory and Production, Chapter 10, Stanford University Press, 1958.
  • T. E. Morton, Bounds on the solution of the lagged optimal inventory equation with no demand backlogging and proportional costs, SIAM Review 11, 572–576, 1969 (as cited in Zipkin 2008).
  • J.-S. Song, The effect of leadtime uncertainty in a simple stochastic inventory model, Management Science 40(5), 603–613, 1994. https://doi.org/10.1287/mnsc.40.5.603
  • D. Stoyan, Comparison Methods for Queues and Other Stochastic Models, Wiley, 1983.
  • M. Shaked, J. G. Shanthikumar, Stochastic Orders, Springer, 2007. https://doi.org/10.1007/978-0-387-34675-5
8 thms1 active userReviewed
PreviousPage 57 of 67Next

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me