Prove2Me
Navigate
DiscoverCollectionsFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

Operations Research

1,703 missions · 839 completed

The discipline of applying mathematical analysis to complex decision problems in operations: allocating scarce resources, scheduling, routing, inventory, and the design of service and production systems. Drawing on mathematical programming, stochastic modeling, queueing, simulation, and game-theoretic reasoning, it seeks policies that perform provably well in systems shaped by constraints, congestion, and uncertainty.

Missions

Open864Completed839All1703
Probability·Captain: mikedeng1

A Re-solving Heuristic with Uniformly Bounded Loss for Network Revenue Management 3: On a Degenerate Two-Class Instance, Frequent Re-solving Has Regret at Least Ω(√T)Research Paper

Motivation

Network revenue managers sell access to limited capacity over time. An airline, for example, may have several products that use the same seat inventory, with some products earning more revenue than others. The decision to accept a request must be made when it arrives, before future requests are known. A standard way to guide that decision is to solve a deterministic linear program using expected demand, then use its optimal allocation as a probability of acceptance. As inventory changes, one can solve the program again. Bumpensanti and Wang study how often to do so, and show that frequent re-solving has a real limitation when the linear program is degenerate: on a particular instance its expected loss grows at least as the square root of the selling horizon. Bumpensanti and Wang, arXiv:1802.06192v3, Sections 3.1 and 5.1.

The same paper gives an infrequent re-solving policy with uniformly bounded regret and an upper bound of order T\sqrt TT​ for the frequent re-solving policy. The lower-bound result is the counterpart that identifies a case where that order cannot be removed by the frequent policy itself. It uses a two-class example small enough to expose the issue without other network complications. Bumpensanti and Wang, arXiv:1802.06192v3, Theorem 1 and Propositions 2–3.

Setting

There is one resource, initially with TTT units of capacity, and two customer classes. Class jjj arrives through an independent Poisson process of rate one. Accepting a customer uses one unit of capacity and earns price rjr_jrj​, where 0<r2<r10<r_2<r_10<r2​<r1​. Rejected requests earn nothing, and unused capacity has no terminal value. The higher-price class has priority in the deterministic allocation, but its realized demand is random. The horizon TTT is a positive integer, divided into TTT periods of length one. Bumpensanti and Wang, arXiv:1802.06192v3, Section 2, pp. 7–9, and Appendix C.1, p. 32.

At the start of a period with kkk periods left and capacity ccc, the deterministic linear program (DLP) chooses rates x1,x2x_1,x_2x1​,x2​ that maximize r1x1+r2x2r_1x_1+r_2x_2r1​x1​+r2​x2​ subject to x1+x2≤c/kx_1+x_2\le c/kx1​+x2​≤c/k and 0≤xj≤10\le x_j\le10≤xj​≤1. In this instance its first coordinate is x1=min⁡{c/k,1}x_1=\min\{c/k,1\}x1​=min{c/k,1}. The frequent re-solving policy (FR) accepts each class-jjj request in that period with probability xjx_jxj​, provided capacity remains. It re-solves the DLP at the next period with the new capacity. This is Algorithm 2 specialized to the example. Bumpensanti and Wang, arXiv:1802.06192v3, Algorithm 2, p. 11, and Appendix D, p. 42.

The hindsight optimum knows the total demand from both classes at the end of the horizon and chooses the best feasible allocation using that information. Its expected value is vHO(T,T)v^{\mathrm{HO}}(T,T)vHO(T,T); the first argument is horizon length and the second is initial capacity. Let vFR(T,T)v^{\mathrm{FR}}(T,T)vFR(T,T) be the frequent policy's expected revenue. Since hindsight has more information, their difference is a regret benchmark. The formulation uses the paper's hindsight linear program, which is exact on this unit-consumption instance. Bumpensanti and Wang, arXiv:1802.06192v3, Eq. (3) and Definition 1, p. 9.

Formalization targets

Main result

For every pair of prices 0<r2<r10<r_2<r_10<r2​<r1​, there are M>0M>0M>0 and T0≥1T_0\ge1T0​≥1 such that, for all integer T≥T0T\ge T_0T≥T0​ and every optimal DLP selector used by FR,

vHO(T,T)−vFR(T,T)≥MT.v^{\mathrm{HO}}(T,T)-v^{\mathrm{FR}}(T,T)\ge M\sqrt T.vHO(T,T)−vFR(T,T)≥MT​.

This specializes Proposition 2's existence statement to the explicit family in its Appendix C.1 proof. The constant may depend on the prices, but is chosen before the selector and the horizon. The benchmark is the hindsight value, as in the proposition. Bumpensanti and Wang, arXiv:1802.06192v3, Proposition 2, p. 18, and Appendix C.1, pp. 32–34.

Supporting results

The milestones record the optimal DLP allocation on the example and two clauses of Lemma 7. The latter estimate the probabilities that the high-price arrival count is moderately below its mean in the first third and moderately above its mean in the last third. With T′T'T′ an integer phase length and N∼Poisson⁡(T′)N\sim\operatorname{Poisson}(T')N∼Poisson(T′), the first bound is

P(T′−4T′≤N≤T′−3T′)≥0.0013−0.9496/T′.\mathbb P(T'-4\sqrt{T'}\le N\le T'-3\sqrt{T'}) \ge 0.0013-0.9496/\sqrt{T'}.P(T′−4T′​≤N≤T′−3T′​)≥0.0013−0.9496/T′​.

The third-phase bound uses the standard normal cumulative distribution function Φ\PhiΦ and the unrounded constant Φ(7)−Φ(6)\Phi(7)-\Phi(6)Φ(7)−Φ(6). Bumpensanti and Wang, arXiv:1802.06192v3, Lemma 7 and Eqs. (49)–(53), p. 40.

Significance

This result shows that solving the DLP every period does not, by itself, give a horizon-independent expected loss. The example has only one resource and two customer classes; the gap cannot be attributed to a large network. The paper's separate bounded-regret result uses a different re-solving schedule and acceptance rule, so formalizing this lower bound helps distinguish guarantees for the two policies. Together with the upper bound for FR, it identifies the square-root order as the relevant scale for this policy on the example. Bumpensanti and Wang, arXiv:1802.06192v3, Sections 4–5.

The mathematical proposition is proved in the cited paper; the Lean goal here remains an open theorem statement. Formalizing it requires precise interfaces for a Poisson arrival model, a capacity-limited randomized policy, a hindsight LP benchmark, and asymptotic lower bounds. Those interfaces can be reused in later revenue-management results, while the two-class instance gives a concrete check on the conventions.

Difficulty

The initial DLP solution allocates the average capacity to the high-price class. A simple intuition might therefore predict that repeated re-solving preserves capacity for that class. The remaining-capacity ratio, however, responds to realized arrivals; when the ratio rises above one, the re-solved LP assigns positive acceptance probability to the lower-price class. The policy then spends capacity before later high-price arrivals are known. The lower-bound analysis needs a path event that controls arrivals through the middle phase and a joint event involving the policy's accepted requests. Marginal Poisson counts alone do not determine those admissions. Bumpensanti and Wang, arXiv:1802.06192v3, Appendix C.1, pp. 32–34.

Formalization scope

Lean uses Fin 2 for classes and Fin 1 for the resource. The general model takes positive Poisson rates, nonnegative prices and nonnegative consumption. On the lower-bound instance the rates and consumptions equal one, capacity equals TTT, and prices satisfy 0<r2<r10<r_2<r_10<r2​<r1​. These are the operational standing assumptions of the paper's Section 2 and the positive-price reading used by the example. Optimal DLP ties are represented by quantifying over every selector. The LP value is the published RLPBidPrice.Unbiased.piValue definition, evaluated on nonnegative capacity vectors; the rate and capacity conventions make its real supremum well posed.

The window law superposes the independent class Poisson processes into a Poisson total count, independent class labels, and independent Bernoulli acceptance marks. A fold over ordered arrivals applies Algorithm 2's capacity check after every acceptance. The expected hindsight value sums over independent total demand counts, and the FR value is a backward recursion over the remaining unit periods. The target's MMM is strictly positive and fixed before the horizon, so neither a zero constant nor a horizon-dependent constant can satisfy it trivially.

The two Lemma 7 milestones cover only Q1Q_1Q1​ and Q3Q_3Q3​. Its continuous-time Q2Q_2Q2​ event and the joint admission event require a path probability space and are outside this draft. Equation (52) yields Φ(7)−Φ(6)\Phi(7)-\Phi(6)Φ(7)−Φ(6); the printed 9.8531×10−109.8531\times10^{-10}9.8531×10−10 rounds that value upward. The source proof divides TTT into three exact integer-length phases, while Proposition 2 is stated for all sufficiently large TTT; this gap requires attention in a complete proof. The source here is arXiv:1802.06192v3, whose printed and PDF page numbers coincide.

Selected references

  • Bumpensanti, P. and Wang, H., A Re-solving Heuristic with Uniformly Bounded Loss for Network Revenue Management, arXiv preprint arXiv:1802.06192v3, 2018. PDF.
7 thms1 active userReviewed
Dynamical SystemsStochastic Systems·Captain: mikedeng1

Stability and Instability of Fluid Models for Reentrant Lines 2: In Every Reentrant Line the First-Buffer-First-Served Fluid Model Is StableResearch Paper

Why priority rules in reentrant lines matter

Semiconductor wafer fabrication is the standard example of a reentrant line: every wafer follows the same long route through a set of machines, and visits several machines many times. Each visit is a separate processing step, and a machine with work waiting from several steps must choose which step to serve. Kumar (1993) proposed studying such lines through their buffer priority rules, and Lu and Kumar (1991) showed that a priority rule can make a line unstable even though every machine has spare capacity. The question is therefore not academic: a scheduling rule that looks reasonable can let queues grow without bound.

Dai (1995) reduced the positive Harris recurrence of a multiclass queueing network to the stability of a deterministic fluid model, and Dai and Weiss (1996) used this reduction to prove stability results for several reentrant lines. This mission formalizes their Theorem 4.3: under the First-Buffer-First-Served (FBFS) discipline, the fluid model of every reentrant line is stable whenever each station's nominal workload is below one. Kumar (1993) had proved the analogous statement for discrete deterministic systems.

Timeline:

  • 1991. Lu and Kumar exhibit a two-station reentrant line that is unstable under a static priority rule with every workload below one.
  • 1993. Kumar proves stability of FBFS and LBFS for discrete deterministic reentrant lines.
  • 1995. Dai shows that stability of the fluid model implies positive Harris recurrence of the stochastic network.
  • 1996. Dai and Weiss prove that the FBFS and LBFS fluid models of every reentrant line are stable (Theorems 4.3 and 4.4), and with Dai's theorem obtain stability of the stochastic networks.

Setting

A reentrant line has III stations and KKK classes. All fluid follows one route: it enters as class 111, becomes class k+1k+1k+1 when class kkk is completed, and leaves after class KKK. Class kkk is served at station σ(k)\sigma(k)σ(k) with mean service time mk>0m_k > 0mk​>0, and μk=1/mk\mu_k = 1/m_kμk​=1/mk​. The constituency of station iii is Ci={k:σ(k)=i}C_i = \{k : \sigma(k) = i\}Ci​={k:σ(k)=i}, and its nominal workload is ρi=∑k∈Cimk\rho_i = \sum_{k\in C_i} m_kρi​=∑k∈Ci​​mk​. The exogenous arrival rate is 111. The standing assumption of the paper is

ρi<1(i=1,…,I).(1.7)\rho_i < 1 \qquad (i = 1,\dots,I). \tag{1.7}ρi​<1(i=1,…,I).(1.7)

A fluid model solution is a pair of paths Q(t)=(Qk(t))kQ(t) = (Q_k(t))_kQ(t)=(Qk​(t))k​, the fluid levels, and T(t)=(Tk(t))kT(t) = (T_k(t))_kT(t)=(Tk​(t))k​, the cumulative service time given to each class, such that for t≥0t \ge 0t≥0:

Qk(t)=Qk(0)+μk−1Tk−1(t)−μkTk(t),Qk(t)≥0,Q_k(t) = Q_k(0) + \mu_{k-1}T_{k-1}(t) - \mu_k T_k(t), \qquad Q_k(t) \ge 0,Qk​(t)=Qk​(0)+μk−1​Tk−1​(t)−μk​Tk​(t),Qk​(t)≥0,

with μ0T0(t)=t\mu_0T_0(t) = tμ0​T0​(t)=t; each TkT_kTk​ starts at 000 and is nondecreasing; and the idle time Ui(t)=t−∑k∈CiTk(t)U_i(t) = t - \sum_{k\in C_i}T_k(t)Ui​(t)=t−∑k∈Ci​​Tk​(t) of each station is nondecreasing.

A buffer priority discipline is a permutation π\piπ of the classes; a class kkk has priority over a class lll at the same station when π(k)<π(l)\pi(k) < \pi(l)π(k)<π(l). Write Hk={l∈Cσ(k):π(l)≤π(k)}H_k = \{l \in C_{\sigma(k)} : \pi(l) \le \pi(k)\}Hk​={l∈Cσ(k)​:π(l)≤π(k)}, Tk+=∑l∈HkTlT_k^+ = \sum_{l\in H_k}T_lTk+​=∑l∈Hk​​Tl​, Uk+(t)=t−Tk+(t)U_k^+(t) = t - T_k^+(t)Uk+​(t)=t−Tk+​(t) and Wk+=∑l∈HkmlQlW_k^+ = \sum_{l\in H_k} m_l Q_lWk+​=∑l∈Hk​​ml​Ql​. Under preemptive resume, Uk+U_k^+Uk+​ may increase only at times when Wk+=0W_k^+ = 0Wk+​=0 (condition (4.4)): a station never idles toward the classes of priority at least that of kkk while any of them holds fluid. FBFS is π(k)=k\pi(k) = kπ(k)=k, so earlier steps of the route have priority.

The fluid model is stable (Definition 1.3) if there is a time δ>0\delta > 0δ>0 such that every solution with ∣Q(0)∣=∑kQk(0)=1|Q(0)| = \sum_k Q_k(0) = 1∣Q(0)∣=∑k​Qk​(0)=1 has Q(t)=0Q(t) = 0Q(t)=0 for all t≥δt \ge \deltat≥δ.

Formalization targets

Goal: Theorem 4.3

For every reentrant line with mk>0m_k > 0mk​>0 and (1.7),

the fluid model (1.8)–(1.12), (4.4) with π(k)=k is stable.\text{the fluid model (1.8)–(1.12), (4.4) with } \pi(k) = k \text{ is stable.}the fluid model (1.8)–(1.12), (4.4) with π(k)=k is stable.

The goal asserts existence of an emptying time; it does not fix the value.

Milestones

  1. Lemma 2.2 (i): a nonnegative function has zero derivative wherever it vanishes and is differentiable.
  2. Lemma 2.2 (ii): an absolutely continuous nonnegative ggg whose derivative is at most −ε-\varepsilon−ε wherever g>0g > 0g>0 vanishes from g(0)/εg(0)/\varepsilong(0)/ε on, and is nonincreasing.
  3. Proposition 4.2: at a regular time of a priority fluid model, an empty buffer has equal in- and out-flow rates, and at each station the highest-priority nonempty class satisfies ∑k∈Hk0mkdk=1\sum_{k\in H_{k_0}} m_kd_k = 1∑k∈Hk0​​​mk​dk​=1 while lower-priority classes have zero out-flow.
  4. Inductive step (proof of Theorem 4.3): once buffers 1,…,k−11,\dots,k-11,…,k−1 stay empty from tk−1t_{k-1}tk−1​ on, buffer kkk is empty from tk=tk−1+Qk(tk−1)mk/(1−∑l∈Hkml)t_k = t_{k-1} + Q_k(t_{k-1})m_k/(1 - \sum_{l\in H_k} m_l)tk​=tk−1​+Qk​(tk−1​)mk​/(1−∑l∈Hk​​ml​) on.
  5. Emptying time (proof of Theorem 4.3): every solution with ∣Q(0)∣=1|Q(0)| = 1∣Q(0)∣=1 is empty from
δ=∑k=1Kmk∏l=1k−1(1−∑j∈Hl∖{l}mj)∏l=1k(1−∑j∈Hlmj)\delta = \sum_{k=1}^K m_k\frac{\prod_{l=1}^{k-1}(1-\sum_{j\in H_l\setminus\{l\}}m_j)}{\prod_{l=1}^{k}(1-\sum_{j\in H_l}m_j)}δ=k=1∑K​mk​∏l=1k​(1−∑j∈Hl​​mj​)∏l=1k−1​(1−∑j∈Hl​∖{l}​mj​)​

on, which gives the goal with an explicit constant.

Significance

Combined with Dai's theorem (Theorem 1.1 of the paper), Theorem 4.3 gives positive Harris recurrence of every multiclass reentrant line operated under FBFS, under the distributional assumptions of that theorem, whenever (1.7) holds. Since Lu–Kumar-type examples show that priority rules can destabilize a network with spare capacity, a rule that is provably stable on every reentrant line is a meaningful guarantee. FBFS, together with LBFS, is the basic positive case against which the instability examples in the same paper (Theorem 5.1) are compared.

Formalizing the result adds a machine-checked version of the fluid model with priority constraints, stated for an arbitrary line rather than a fixed network, and a checked explicit emptying time. The result is proved in the paper; no machine-checked proof is known. The rate identities of Proposition 4.2 and the extinction lemma are reused by the LBFS mission of this series.

Difficulty

The fluid model gives the paths only as nondecreasing functions subject to equations; nothing says they are differentiable. Every argument about rates holds only at regular times, and the conclusion must be integrated back through absolute continuity, which itself must be derived from the monotonicity conditions. The priority condition (4.4) is an "increases only when empty" condition; turning it into the rate identity (4.6) at a regular point requires continuity of the fluid levels and a careful local argument. The induction over classes needs uniform control of every earlier buffer for all later times, not only at one instant, so a pointwise reading of "buffer k−1k-1k−1 is empty" is not enough. Finally, the explicit constant δ\deltaδ requires bounding the content of buffer kkk at time tk−1t_{k-1}tk−1​ by the total content, which may have grown since time 000.

Formalization scope

  • Classes and stations are Fin K and Fin I, 0-based: the paper's class kkk is Lean index k−1k-1k−1.
  • Paths are total functions ℝ → Fin K → ℝ; every equation is imposed on t≥0t \ge 0t≥0 only, and derivatives are taken at t>0t > 0t>0.
  • Conditions (1.13) and (4.4) are in interval form (ConstantWhile): the idle process is constant on every interval of [0,∞)[0,\infty)[0,∞) on which the content stays positive. This is equivalent to the paper's Stieltjes integral form for continuous paths.
  • Lipschitz continuity of the paths is not assumed; it follows from the monotonicity conditions.
  • mk>0m_k > 0mk​>0 is an explicit hypothesis; the paper takes it for granted.
  • Priorities are Equiv.Perm (Fin K), FBFS is the identity; ∣Q(0)∣|Q(0)|∣Q(0)∣ is ∑kQk(0)\sum_k Q_k(0)∑k​Qk​(0).
  • In Lemma 2.2 (ii) the printed strict "g˙(t)<−ε\dot g(t) < -\varepsilong˙​(t)<−ε" is relaxed to "≤−ε\le -\varepsilon≤−ε", as every application requires; absolute continuity is dropped from Lemma 2.2 (i), where it is not needed. Both changes strengthen the lemma.
  • In the inductive step, "stay empty for t>tk−1t > t_{k-1}t>tk−1​" is read as t≥tk−1t \ge t_{k-1}t≥tk−1​, and the case Qk(tk−1)=0Q_k(t_{k-1}) = 0Qk​(tk−1​)=0 is included.
  • The stochastic network is not formalized; only the fluid model is.

A trivializing formalization is ruled out: the goal uses the priority fluid model, not the work-conserving one alone (which would claim that every work-conserving fluid model is stable, false by Theorem 5.1), and a sorry-free witness shows that the FBFS solution predicate has solutions with ∣Q(0)∣=1|Q(0)| = 1∣Q(0)∣=1, so stability is not vacuous.

Contributions welcome: the extinction lemma for absolutely continuous functions, the derivation of Lipschitz continuity from (1.10)–(1.12), and the rate identities of Proposition 4.2 are reusable beyond this mission.

Selected references

  • J. G. Dai and G. Weiss, Stability and instability of fluid models for reentrant lines, Mathematics of Operations Research 21(1) (1996) 115–134. https://doi.org/10.1287/moor.21.1.115
  • J. G. Dai, On positive Harris recurrence of multiclass queueing networks: a unified approach via fluid limit models, Annals of Applied Probability 5(1) (1995) 49–77. https://doi.org/10.1214/aoap/1177004828
  • P. R. Kumar, Re-entrant lines, Queueing Systems 13 (1993) 87–110. https://doi.org/10.1007/BF01149327
  • S. H. Lu and P. R. Kumar, Distributed scheduling based on due dates and buffer priorities, IEEE Transactions on Automatic Control 36(12) (1991) 1406–1416. https://doi.org/10.1109/9.106159
8 thms1 active userReviewed
CombinatoricsLinear OptimizationTheoretical Computer Science·Captain: mikedeng1

Online Primal-Dual Algorithms for Covering and Packing 3: A Deterministic O(log d log(n/OPT))-Competitive Algorithm for Online Unweighted Set CoverResearch Paper

Motivation

Online set cover is the basic covering problem in which the requests arrive over time. A ground set of elements and a family of sets are known in advance, but which elements must be covered is revealed one element at a time, and each arriving element has to be covered at once by a set chosen irrevocably. The problem models resource placement under unknown demand (facilities, servers, sensors that must serve clients as they appear) and is the prototype for a family of online covering problems.

Alon, Awerbuch, Azar, Buchbinder and Naor (SIAM J. Comput. 2009) gave the first deterministic algorithm, with competitive ratio O(log⁡mlog⁡n)O(\log m\log n)O(logmlogn) for nnn elements and mmm sets, and showed that no deterministic algorithm does better than Ω(log⁡mlog⁡n/(log⁡log⁡m+log⁡log⁡n))\Omega(\log m\log n/(\log\log m+\log\log n))Ω(logmlogn/(loglogm+loglogn)) on some instances. Buchbinder and Naor (Math. Oper. Res. 2009) recast the fractional part of that algorithm as an instance of a general online primal-dual scheme for covering and packing linear programs, and turned an offline pessimistic estimator of Srinivasan into an online potential function. The result, in their Section 5.1, is a deterministic algorithm whose ratio O(log⁡dlog⁡(n/OPT))O(\log d\log(n/OPT))O(logdlog(n/OPT)) depends on the maximum element frequency ddd instead of the number of sets mmm, and on the ratio n/OPTn/OPTn/OPT instead of nnn.

Timeline:

  • 2003 (conference), 2009 (journal): Alon et al., deterministic O(log⁡mlog⁡n)O(\log m\log n)O(logmlogn) for unweighted online set cover, with a potential ∑j∉Cn2wj\sum_{j\notin C} n^{2w_j}∑j∈/C​n2wj​.
  • 2005 (conference), 2009 (journal): Buchbinder and Naor, the general online fractional covering/packing scheme, and the derandomized rounding of this mission.

Setting

A set-cover instance consists of a finite ground set XXX of nnn elements and a finite family S\mathcal SS of mmm sets. For an element eee, Se\mathcal S_eSe​ is the collection of sets containing eee, and ddd bounds its size: ∣Se∣≤d|\mathcal S_e|\le d∣Se​∣≤d for every eee (the frequency). In the unweighted problem every set costs 111.

Elements arrive in a list σ\sigmaσ. The algorithm maintains:

  • fractional weights w(s)≥0w(s)\ge 0w(s)≥0 for the sets, produced by the paper's Section 3 scheme with {0,1}\{0,1\}{0,1} coefficients: when an element eee arrives that is not yet fractionally covered (∑s∈Sew(s)<1\sum_{s\in\mathcal S_e}w(s)<1∑s∈Se​​w(s)<1), its dual variable y(e)y(e)y(e) is raised to the least value at which
w(s)=max⁡{w(s), 1d(exp⁡(B2c(s)∑k: s∋eky(ek))−1)}(s∋e)w(s)=\max\Big\{w(s),\ \tfrac1d\Big(\exp\Big(\tfrac{B}{2c(s)}\textstyle\sum_{k:\ s\ni e_k}y(e_k)\Big)-1\Big)\Big\}\qquad(s\ni e)w(s)=max{w(s), d1​(exp(2c(s)B​∑k: s∋ek​​y(ek​))−1)}(s∋e)

gives ∑s∈Sew(s)≥1\sum_{s\in\mathcal S_e}w(s)\ge 1∑s∈Se​​w(s)≥1; here B>0B>0B>0 is a parameter;

  • a cover C⊆S\mathcal C\subseteq\mathcal SC⊆S that only grows; CCC is the set of elements covered by C\mathcal CC.

With f(e)=min⁡{1,exp⁡(−α+α∑s∋ew(s))}f(e)=\min\{1,\exp(-\alpha+\alpha\sum_{s\ni e}w(s))\}f(e)=min{1,exp(−α+α∑s∋e​w(s))}, the potential is Φ=Φ1+Φ2\Phi=\Phi_1+\Phi_2Φ=Φ1​+Φ2​ with

Φ1=1−∏e∈X∖C(1−f(e)),Φ2=exp⁡(∑s∈S((ln⁡2)χC(s)−αw(s))−OPT),\Phi_1=1-\prod_{e\in X\setminus C}\big(1-f(e)\big),\qquad \Phi_2=\exp\Big(\sum_{s\in\mathcal S}\big((\ln 2)\chi_{\mathcal C}(s)-\alpha w(s)\big)-OPT\Big),Φ1​=1−e∈X∖C∏​(1−f(e)),Φ2​=exp(s∈S∑​((ln2)χC​(s)−αw(s))−OPT),

where OPTOPTOPT is the optimum number of sets covering the arrived elements, assumed known, r=eln⁡(e/(e−1))r=e\ln(e/(e-1))r=eln(e/(e−1)) and α=max⁡{1,ln⁡(rn/OPT)}\alpha=\max\{1,\ln(rn/OPT)\}α=max{1,ln(rn/OPT)}. The rounding rule: each time the weight of a set sss is augmented, sss is added to C\mathcal CC if this does not increase Φ\PhiΦ.

Formalization targets

Goal: Lemma 5.2

Every arriving element is covered by C\mathcal CC, and at every time

∣C∣ ≤ (2αln⁡(1+d)+1) OPTln⁡2,α=max⁡{1,ln⁡rnOPT}.|\mathcal C|\ \le\ \frac{\big(2\alpha\ln(1+d)+1\big)\,OPT}{\ln 2},\qquad \alpha=\max\Big\{1,\ln\frac{rn}{OPT}\Big\}.∣C∣ ≤ ln2(2αln(1+d)+1)OPT​,α=max{1,lnOPTrn​}.

This is the paper's OPT⋅O(log⁡dlog⁡(n/OPT))OPT\cdot O(\log d\log(n/OPT))OPT⋅O(logdlog(n/OPT)) with the constant its proof gives.

Milestones

  1. Theorem 3.2 (covering half): for every B>0B>0B>0, the {0,1}\{0,1\}{0,1} scheme with frequency bound ℓ\ellℓ yields a fractional cover of cost at most 2ln⁡(1+ℓ)2\ln(1+\ell)2ln(1+ℓ) times that of any fractional cover.
  2. Lemma 5.1 (i): initially Φ≤1\Phi\le 1Φ≤1; Φ>0\Phi>0Φ>0 in every state.
  3. Lemma 5.1 (ii): after the weight of a set is augmented by δ≥0\delta\ge0δ≥0, taking the set or excluding it leaves Φ\PhiΦ no larger than before.

Significance

The bound improves Alon et al.'s O(log⁡mlog⁡n)O(\log m\log n)O(logmlogn) whenever sets are many but each element lies in few of them (d≪md\ll md≪m), and whenever the optimum is large compared with nnn. It also shows that the fractional part and the rounding part of an online covering algorithm can be designed separately: any online fractional solution with a competitive guarantee can be rounded deterministically by an online potential function. The same method gives the routing result of Section 5.2 of the paper.

As far as is known, none of these results is machine-checked. Related formal work on the platform covers Alon et al.'s algorithm and its O(log⁡mlog⁡n)O(\log m\log n)O(logmlogn) bound (a different potential and a different fractional update), and the general covering scheme of Buchbinder and Naor's monograph; neither states Lemma 5.1 or Lemma 5.2, nor Theorem 3.2 for the {0,1}\{0,1\}{0,1} scheme with ℓ\ellℓ in place of nnn. A complete development here would give a verified derandomized rounding argument, reusable for other online covering problems.

Difficulty

The difficulty is not in the final inequality, which follows from Φ2≤1\Phi_2\le1Φ2​≤1 in one line, but in keeping Φ≤1\Phi\le1Φ≤1 throughout. The decision to take a set must be made with no knowledge of future elements, and the potential has to account for elements that may never arrive: Φ1\Phi_1Φ1​ ranges over the whole ground set. Lemma 5.1 (ii) asks that, for every current state, one of the two decisions does not increase a non-linear function of all uncovered elements at once and this must hold for every state, not only for the states a particular run reaches. On the fractional side, Theorem 3.2 is only asserted, "along the same lines" as Theorem 3.1, so its constant has to be re-derived with ℓ\ellℓ in place of nnn and with the scheme's continuous increase made discrete.

A tempting shortcut is to bound ∣C∣|\mathcal C|∣C∣ by the number of rounds or by ∑sw(s)\sum_s w(s)∑s​w(s) directly; neither gives a logarithmic factor in n/OPTn/OPTn/OPT, which comes only from the choice of α\alphaα in Φ\PhiΦ.

Formalization scope

The instance is the published OnlinePrimalDual.OnlineSetCover.SetCoverInstance (elements E, set indices T, incidence elemSets, positive costs c), with elementWeight and coveredBy. Unit costs are the hypothesis ∀ s, inst.c s = 1 in Lemma 5.2; Theorem 3.2 is stated for general positive costs. Logarithms are natural (Real.log), because they invert Real.exp.

Committed conventions:

  • The fractional scheme is the discrete form of the continuous increase: in each round y(e)y(e)y(e) is the least t≥0t\ge0t≥0 (an sInf) at which the new constraint holds. Arrival lists may repeat elements; an element already covered changes nothing.
  • The rounding treats each set's increase within a round as one augmentation; the augmented sets are processed one at a time in the order of a list ord containing every set, and a set is added when Φ(w+δs1s,C∪{s})≤Φ(w,C)\Phi(w+\delta_s\mathbf 1_s,\mathcal C\cup\{s\})\le\Phi(w,\mathcal C)Φ(w+δs​1s​,C∪{s})≤Φ(w,C). Lemma 5.2 holds for every order.
  • The algorithm is a function, so every run exists.
  • OPTOPTOPT is a natural number ≥1\ge1≥1 bounding the size of some cover of the arrived elements; d≥1d\ge1d≥1 bounds the frequency of every element. At the true optimum and the maximum frequency this is the paper's statement.
  • Explicit constants replacing O(⋅)O(\cdot)O(⋅): Theorem 3.2's O(log⁡ℓ)O(\log\ell)O(logℓ) is 2ln⁡(1+ℓ)2\ln(1+\ell)2ln(1+ℓ); Lemma 5.2's OPT⋅O(log⁡dlog⁡(n/OPT))OPT\cdot O(\log d\log(n/OPT))OPT⋅O(logdlog(n/OPT)) is (2αln⁡(1+d)+1) OPT/ln⁡2(2\alpha\ln(1+d)+1)\,OPT/\ln2(2αln(1+d)+1)OPT/ln2 with α=max⁡{1,ln⁡(rn/OPT)}\alpha=\max\{1,\ln(rn/OPT)\}α=max{1,ln(rn/OPT)}.
  • The proof of Lemma 5.1 (i) on p. 13 prints ddd where rrr is meant ("exp⁡(−dne−α)\exp(-dne^{-\alpha})exp(−dne−α)", "α≥ln⁡(dn/OPT)\alpha\ge\ln(dn/OPT)α≥ln(dn/OPT)"); the statement uses rrr, and rrr is printed "eln⁡(e/e−1)e\ln(e/e-1)eln(e/e−1)" for eln⁡(e/(e−1))e\ln(e/(e-1))eln(e/(e−1)).

Not in scope: the packing half of Theorem 3.2, Theorem 3.1 for general coefficients, and the doubling wrapper of p. 12 that removes the assumption that OPTOPTOPT is known. The guarantee is for the algorithm run with the stated α\alphaα; a formalization in which Φ\PhiΦ's product ranges only over arrived elements, or in which the weights or the chosen family are free variables constrained by hypotheses instead of being produced by the algorithm (OPTOPTOPT is an input of the algorithm, as the page assumes it known), would be a different and weaker statement and is ruled out.

Contributions welcome: proofs of the three milestones and of the goal, and general lemmas about the fractional round (attainment of the least ttt, monotonicity of the weights) that other online covering missions can reuse.

Selected references

  • N. Buchbinder, J. Naor, Online Primal-Dual Algorithms for Covering and Packing, Mathematics of Operations Research 34(2), 2009. https://doi.org/10.1287/moor.1080.0363
  • N. Alon, B. Awerbuch, Y. Azar, N. Buchbinder, J. Naor, The Online Set Cover Problem, SIAM Journal on Computing 39(2), 2009. https://doi.org/10.1137/060661946
  • A. Srinivasan, Improved approximation guarantees for packing and covering integer programs, SIAM Journal on Computing 29(2), 1999. https://doi.org/10.1137/S0097539796314240
  • N. Buchbinder, J. Naor, The Design of Competitive Online Algorithms via a Primal-Dual Approach, Foundations and Trends in Theoretical Computer Science 3(2–3), 2009. https://doi.org/10.1561/0400000024
10 thms1 active userReviewed
Linear OptimizationProbability·Captain: mikedeng1

A Re-solving Heuristic with Uniformly Bounded Loss for Network Revenue Management 1: Infrequent Re-solving with Thresholding Has Regret O(1), Uniformly in the Horizon T and the Capacities CResearch Paper

Motivation

Network revenue management decides, in real time, which customer requests to accept when every request consumes a bundle of scarce, perishable resources: seats on several flight legs of an itinerary, room-nights of a multi-night hotel stay, bandwidth on several links. The exact dynamic program is intractable for realistic networks, so practice and theory rely on heuristics built from a deterministic linear program (DLP) that replaces random demand by its mean. The question is how much revenue such heuristics lose.

The classical answer (Gallego and van Ryzin 1994, 1997; Talluri and van Ryzin 1998) is that solving the DLP once and following its solution loses O(T)O(\sqrt T)O(T​) over a horizon of length TTT. Re-solving the DLP as capacity is consumed was long believed to help, and Jasin and Kumar (2012) proved that re-solving in every period has bounded loss if the DLP solution is nondegenerate. Bumpensanti and Wang (arXiv:1802.06192v3) showed that the nondegeneracy condition matters: frequent re-solving can lose Ω(T)\Omega(\sqrt T)Ω(T​) on degenerate instances. They then proposed a policy, Infrequent Re-solving with Thresholding (IRT), whose loss is bounded by a constant independent of the horizon and of the capacities, with no nondegeneracy assumption. This mission formalizes that result.

Setting

There are nnn customer classes j∈[n]j \in [n]j∈[n] and mmm resources l∈[m]l \in [m]l∈[m]. Class-jjj customers arrive as independent Poisson processes of rate λj>0\lambda_j > 0λj​>0 on [0,T][0, T][0,T]. Accepting a class-jjj customer earns rj≥0r_j \ge 0rj​≥0 and consumes alj≥0a_{lj} \ge 0alj​≥0 units of resource lll; A=(alj)A = (a_{lj})A=(alj​) is the bill-of-materials matrix with columns AjA_jAj​, and C∈R≥0mC \in \mathbb R^m_{\ge 0}C∈R≥0m​ is the initial capacity. A customer can be accepted only if Aj≤C′A_j \le C'Aj​≤C′ componentwise, C′C'C′ being the remaining capacity; leftover capacity is worthless at TTT.

The DLP with right-hand side bbb is

max⁡x{∑jrjxj ∣ ∑jAjxj≤b, 0≤xj≤λj},\max_x \Big\{ \sum_j r_j x_j \ \Big|\ \sum_j A_j x_j \le b,\ 0 \le x_j \le \lambda_j \Big\},xmax​{j∑​rj​xj​ ​ j∑​Aj​xj​≤b, 0≤xj​≤λj​},

and vDLP(T,C)v^{\mathrm{DLP}}(T, C)vDLP(T,C) is TTT times its value at b=C/Tb = C/Tb=C/T. The hindsight optimum is vHO(T,C)=E[VHO]v^{\mathrm{HO}}(T, C) = \mathbb E[V^{\mathrm{HO}}]vHO(T,C)=E[VHO], where VHOV^{\mathrm{HO}}VHO is the value of the same LP with capacity CCC and the realized demands Λj(T)∼Poisson(λjT)\Lambda_j(T) \sim \mathrm{Poisson}(\lambda_j T)Λj​(T)∼Poisson(λj​T) as upper bounds. It bounds the expected revenue of every non-anticipating policy, and the regret of a policy π\piπ is vHO−vπv^{\mathrm{HO}} - v^\pivHO−vπ.

A probabilistic allocation on a window accepts each class-jjj arrival with a fixed probability pjp_jpj​ when capacity allows. The IRT policy uses τu=T(5/6)u\tau_u = T^{(5/6)^u}τu​=T(5/6)u and re-solving times tu∗=T−τut^*_u = T - \tau_utu∗​=T−τu​ for u=0,…,Ku = 0, \dots, Ku=0,…,K, where K=⌈log⁡log⁡T/log⁡(6/5)⌉K = \lceil \log\log T / \log(6/5)\rceilK=⌈loglogT/log(6/5)⌉. At tu∗t^*_utu∗​ it solves the DLP with right-hand side C(tu∗)/τuC(t^*_u)/\tau_uC(tu∗​)/τu​ (remaining capacity over remaining time), obtaining xux^uxu. In the epochs u<Ku < Ku<K it accepts class jjj with probability 000 if xju<λjτu−1/4x^u_j < \lambda_j\tau_u^{-1/4}xju​<λj​τu−1/4​, else 111 if xju>λj(1−τu−1/4)x^u_j > \lambda_j(1 - \tau_u^{-1/4})xju​>λj​(1−τu−1/4​), else xju/λjx^u_j/\lambda_jxju​/λj​; in the last epoch [tK∗,T][t^*_K, T][tK∗​,T] it uses xjK/λjx^K_j/\lambda_jxjK​/λj​. The family IRTK′\mathrm{IRT}^{K'}IRTK′ re-solves K′K'K′ times on the same schedule; IRT0\mathrm{IRT}^0IRT0 is static probabilistic allocation (SPA), and HOK′\mathrm{HO}^{K'}HOK′ follows IRTK′\mathrm{IRT}^{K'}IRTK′ until tK′∗t^*_{K'}tK′∗​ and then earns the hindsight optimum of what remains.

Formalization targets

Goal: Theorem 1 (p. 16)

There is a constant M=M(λ,r,A)M = M(\lambda, r, A)M=M(λ,r,A) such that for every horizon T∈{1,2,… }T \in \{1, 2, \dots\}T∈{1,2,…} and every capacity vector C≥0C \ge 0C≥0,

vHO(T,C)−vIRT(T,C)≤M.v^{\mathrm{HO}}(T, C) - v^{\mathrm{IRT}}(T, C) \le M .vHO(T,C)−vIRT(T,C)≤M.

The constant is not fixed; the content is its independence of TTT, of CCC, and of the optimal LP solution chosen at each re-solve.

Milestones

  • vHO≤vDLPv^{\mathrm{HO}} \le v^{\mathrm{DLP}}vHO≤vDLP (Sec. 2.2.2, p. 9).
  • Lemma 4 (p. 37): P(∣X−μ∣≥x)≤2e−x2/(3μ)\mathbb P(|X - \mu| \ge x) \le 2e^{-x^2/(3\mu)}P(∣X−μ∣≥x)≤2e−x2/(3μ) for X∼Poisson(μ)X \sim \mathrm{Poisson}(\mu)X∼Poisson(μ), 0<x≤μ0 < x \le \mu0<x≤μ.
  • Proposition 5 (p. 27): vDLP−vSPA≤MTv^{\mathrm{DLP}} - v^{\mathrm{SPA}} \le M\sqrt TvDLP−vSPA≤MT​, uniformly in CCC.
  • Proposition 1 (p. 16): with one re-solve at T−T5/6T - T^{5/6}T−T5/6,
vHO−vHO1≤MTe−κT1/6,vHO−vIRT1≤MTe−κT1/6+MT5/12.v^{\mathrm{HO}} - v^{\mathrm{HO}^1} \le M T e^{-\kappa T^{1/6}},\qquad v^{\mathrm{HO}} - v^{\mathrm{IRT}^1} \le M T e^{-\kappa T^{1/6}} + M T^{5/12}.vHO−vHO1≤MTe−κT1/6,vHO−vIRT1≤MTe−κT1/6+MT5/12.
  • Eq. (13) (p. 29): for every K′K'K′,
vHO−vIRTK′≤M∑u=0K′−1T(5/6)ue−κT(5/6)u/6+MT(5/6)K′/2.v^{\mathrm{HO}} - v^{\mathrm{IRT}^{K'}} \le M\sum_{u=0}^{K'-1} T^{(5/6)^u} e^{-\kappa T^{(5/6)^u/6}} + M T^{(5/6)^{K'}/2}.vHO−vIRTK′≤Mu=0∑K′−1​T(5/6)ue−κT(5/6)u/6+MT(5/6)K′/2.
  • p. 30: T(5/6)K≤eT^{(5/6)^{K}} \le eT(5/6)K≤e, and the right-hand side of (13) at K′=K(T)K' = K(T)K′=K(T) is bounded uniformly in T≥1T \ge 1T≥1.

Significance

The result separates two design choices in re-solving heuristics: how often to re-solve and how to turn an LP solution into a control. It shows that re-solving only O(log⁡log⁡T)O(\log\log T)O(loglogT) times, combined with rounding nearly-degenerate acceptance probabilities to 000 or 111, is enough for bounded regret, and that the constant is uniform over all capacity-to-horizon ratios, so degenerate DLP solutions, which occur only at particular ratios, cause no loss of order. Since vHO≥v∗v^{\mathrm{HO}} \ge v^*vHO≥v∗, it also shows that the hindsight optimum is within a constant of the optimal policy's value in this model.

The result is proved on paper but not machine-checked. A formalization would produce a reusable Poisson-arrival revenue-management model, a precise account of the regret decomposition over re-solving epochs, and an independent check of a proof with known gaps (see Difficulty). No part of it is formalized on the platform.

Difficulty

The obvious argument compares the policy with the DLP solution and controls the deviation of Poisson demand by its standard deviation; this gives only O(T)O(\sqrt T)O(T​), because the DLP value exceeds the hindsight optimum by order T\sqrt TT​ and the lost sales accumulate. Bounded regret needs a comparison with the hindsight LP, path by path: the acceptances made before the last re-solve must remain extendable to a hindsight-optimal solution with overwhelming probability. This requires LP sensitivity of the hindsight solution to the random right-hand side, uniformly over the degenerate cases, and the printed proof's statement of it (Lemma 5) is false as printed: on a two-class, single-resource instance its interval for zˉ1\bar z_1zˉ1​ is empty. A correct argument needs a proximity bound for optimal LP solutions under a change of the demand bounds, summed over all classes, not only over Jλ={j:xj∗=λj}J_\lambda = \{j : x^*_j = \lambda_j\}Jλ​={j:xj∗​=λj​}. The analysis of the schedule (the sum over epochs in (13)) is elementary but the printed chain of inequalities on p. 30 uses a monotonicity that fails for large arguments.

Formalization scope

All objects are in the definition item ResolvingNRM.IRT.Model. Classes are Fin n, resources Fin m, A : Matrix (Fin m) (Fin n) ℝ. LP values reuse piValue from the published RLPBidPrice.Unbiased.Model; the hindsight LP is over real zzz, as in (3). Expectations are explicit sums of Poisson weights. A window of probabilistic allocation is represented by its Poisson count of arrivals, i.i.d. classes with law λj/∑iλi\lambda_j/\sum_i\lambda_iλj​/∑i​λi​ and independent Bernoulli acceptance coins, the standard representation for controls constant on the window. Policies are backward recursions over the re-solving epochs. Ties in the LP are left open: a policy takes an arbitrary optimal-solution selector, and every theorem quantifies over all selectors after its constants.

Conventions and deviations from the page, each disclosed in the item concerned:

  • Standing assumptions of Sec. 2 (p. 7), left implicit there: λj>0\lambda_j > 0λj​>0, r≥0r \ge 0r≥0, A≥0A \ge 0A≥0, C≥0C \ge 0C≥0.
  • "O(g)O(g)O(g)" is read as: there is MMM, depending only on (λ,r,A)(\lambda, r, A)(λ,r,A), with the bound ≤Mg(T)\le M g(T)≤Mg(T) for all T≥1T \ge 1T≥1 (or T>0T > 0T>0) and all C≥0C \ge 0C≥0.
  • Milestones applied to sub-horizons T(5/6)uT^{(5/6)^u}T(5/6)u take a real TTT; the goal takes T∈NT \in \mathbb NT∈N.
  • Lemma 4 is stated for 0<x≤μ0 < x \le \mu0<x≤μ; as printed (all x>0x > 0x>0) it is false, e.g. μ=1\mu = 1μ=1, x=5x = 5x=5.
  • In Proposition 1 and (13), κ>0\kappa > 0κ>0 is existential and uniform in CCC; the printed κ\kappaκ depends on JλJ_\lambdaJλ​ and is derived through Lemma 5.
  • (13) is stated for every number K′K'K′ of re-solves.
  • Algorithm 3 prints the re-solve right-hand side as C(tk∗)/τkC(t^*_k)/\tau_kC(tk∗​)/τk​; it is read with index uuu.
  • For 1≤T≤e1 \le T \le e1≤T≤e the formula for KKK is undefined or negative; the formalization takes K=0K = 0K=0, so IRT is SPA there.
  • The optimal policy value v∗v^*v∗ and the paper's constant α\alphaα are not defined; Lemmas 5 and 6 are not stated.

A statement that lets MMM depend on TTT or CCC, compares IRT with the DLP instead of the hindsight optimum, uses a fixed number of re-solves, or assumes a nondegenerate or vertex LP solution is trivial or a different theorem, and is ruled out by the quantifier order of the goal. The milestone T(5/6)K≤eT^{(5/6)^K} \le eT(5/6)K≤e has a sorry-free local check.

Welcome contributions: LP proximity results (Cook–Gerards–Schrijver–Tardos type), Poisson concentration, superposition and thinning of Poisson processes, and lemmas on expectations of LP values. The source is arXiv:1802.06192v3; its printed page numbers equal the PDF page numbers.

Selected references

  • P. Bumpensanti, H. Wang, A Re-solving Heuristic with Uniformly Bounded Loss for Network Revenue Management, arXiv:1802.06192v3, 2018; Management Science 66(7), 2020. https://arxiv.org/abs/1802.06192
  • S. Jasin, S. Kumar, A Re-Solving Heuristic with Bounded Revenue Loss for Network Revenue Management with Customer Choice, Mathematics of Operations Research 37(2), 2012. https://doi.org/10.1287/moor.1110.0530
  • G. Gallego, G. van Ryzin, A Multiproduct Dynamic Pricing Problem and Its Applications to Network Yield Management, Operations Research 45(1), 1997. https://doi.org/10.1287/opre.45.1.24
  • K. Talluri, G. van Ryzin, An Analysis of Bid-Price Controls for Network Revenue Management, Management Science 44(11), 1998. https://doi.org/10.1287/mnsc.44.11.1577
  • M. Reiman, Q. Wang, An Asymptotically Optimal Policy for a Quantity-Based Network Revenue Management Problem, Mathematics of Operations Research 33(2), 2008. https://doi.org/10.1287/moor.1070.0288
  • W. Cook, A. M. H. Gerards, A. Schrijver, É. Tardos, Sensitivity Theorems in Integer Linear Programming, Mathematical Programming 34, 1986. https://doi.org/10.1007/BF01582230
10 thms1 active userReviewed
Machine LearningOptimizationProbability+1·Captain: mikedeng1

From Predictive to Prescriptive Analytics 1: The k-Nearest-Neighbor Predictive Prescription Is Asymptotically Optimal and ConsistentResearch Paper

Motivation

Operations research has long studied the stochastic problem min⁡z∈ZE[c(z;Y)]\min_{z\in\mathcal Z}\mathbb E[c(z;Y)]minz∈Z​E[c(z;Y)]: choose a decision zzz before an uncertain quantity YYY (demand, prices, returns) is revealed. In practice the decision maker also observes covariates XXX before deciding (web-search volume before stocking a product, weather before routing a shipment) and should solve the conditional problem instead. Bertsimas and Kallus (arXiv:1402.5481v4; Management Science 66(3), 2020) proposed predictive prescriptions: reweight historical observations (xi,yi)(x^i,y^i)(xi,yi) by how relevant they are to the current xxx, using weights borrowed from a nonparametric regression method, and minimize the reweighted sample cost. Their asymptotic theorems guarantee that this procedure converges to the decision that full knowledge of the conditional distribution would give.

The unconditional special case, sample average approximation (SAA) with i.i.d. draws of YYY and no covariate, has a classical consistency theory (Dupačová–Wets 1988; Shapiro 2003). This mission conditions that theory on XXX, with kkk-nearest-neighbour weights.

Setting

Decisions zzz lie in a set Z⊆Rdz\mathcal Z\subseteq\mathbb R^{d_z}Z⊆Rdz​, covariates XXX in X⊆Rdx\mathcal X\subseteq\mathbb R^{d_x}X⊆Rdx​, uncertainty YYY in Y⊆Rdy\mathcal Y\subseteq\mathbb R^{d_y}Y⊆Rdy​; all three spaces carry the Euclidean norm. A cost c(z;y)∈Rc(z;y)\in\mathbb Rc(z;y)∈R is given. Let μ\muμ be the joint law of (X,Y)(X,Y)(X,Y), μX\mu_XμX​ the law of XXX and μY∣x\mu_{Y|x}μY∣x​ the conditional law of YYY given X=xX=xX=x. The conditional cost and the full-information problem (2) are

C(z∣x)=E[c(z;Y)∣X=x],v∗(x)=min⁡z∈ZC(z∣x),Z∗(x)=arg⁡min⁡z∈ZC(z∣x).C(z\mid x)=\mathbb E\big[c(z;Y)\mid X=x\big],\qquad v^*(x)=\min_{z\in\mathcal Z}C(z\mid x),\qquad \mathcal Z^*(x)=\arg\min_{z\in\mathcal Z}C(z\mid x).C(z∣x)=E[c(z;Y)∣X=x],v∗(x)=z∈Zmin​C(z∣x),Z∗(x)=argz∈Zmin​C(z∣x).

The data are SN={(x1,y1),…,(xN,yN)}S_N=\{(x^1,y^1),\dots,(x^N,y^N)\}SN​={(x1,y1),…,(xN,yN)}, the first NNN terms of an i.i.d. sequence with law μ\muμ. The kNN weights (12) are wN,i(x)=1kI[xiw_{N,i}(x)=\frac1k\mathbb I[x^iwN,i​(x)=k1​I[xi is one of the kkk nearest neighbours of xxx among x1,…,xN]x^1,\dots,x^N]x1,…,xN], with ties among equidistant points broken by lower index first. The predictive prescription (3) is any

z^N(x)∈arg⁡min⁡z∈Z C^N(z∣x),C^N(z∣x)=∑i=1NwN,i(x) c(z;yi).\hat z_N(x)\in\arg\min_{z\in\mathcal Z}\ \widehat C_N(z\mid x),\qquad \widehat C_N(z\mid x)=\sum_{i=1}^N w_{N,i}(x)\,c(z;y^i).z^N​(x)∈argz∈Zmin​ CN​(z∣x),CN​(z∣x)=i=1∑N​wN,i​(x)c(z;yi).

The paper's standing assumptions: Assumption 3 (E∣c(z;Y)∣<∞\mathbb E|c(z;Y)|<\inftyE∣c(z;Y)∣<∞ for every z∈Zz\in\mathcal Zz∈Z, and Z∗(x)≠∅\mathcal Z^*(x)\ne\emptysetZ∗(x)=∅ for a.e. xxx); Assumption 4 (c(z;y)c(z;y)c(z;y) is equicontinuous in zzz, uniformly over y∈Yy\in\mathcal Yy∈Y); Assumption 5 (Z\mathcal ZZ closed and nonempty, and either bounded, or ccc is bounded below for large ∥z∥\|z\|∥z∥ and for each xxx there is a set DxD_xDx​ of positive conditional probability on which c(z;y)→∞c(z;y)\to\inftyc(z;y)→∞ uniformly as ∥z∥→∞\|z\|\to\infty∥z∥→∞).

Definition 1. z^N\hat z_Nz^N​ is asymptotically optimal if, with probability 1, for μX\mu_XμX​-a.e. xxx, C(z^N(x)∣x)→v∗(x)C(\hat z_N(x)\mid x)\to v^*(x)C(z^N​(x)∣x)→v∗(x); it is consistent if, with probability 1, for μX\mu_XμX​-a.e. xxx, inf⁡z∈Z∗(x)∥z^N(x)−z∥→0\inf_{z\in\mathcal Z^*(x)}\|\hat z_N(x)-z\|\to0infz∈Z∗(x)​∥z^N​(x)−z∥→0.

Formalization targets

Goal: Theorem 5 (kNN), p. 19 (= Theorem 15, p. 43)

Under Assumptions 3, 4, 5 and i.i.d. sampling, with k=min⁡{⌈CNδ⌉,N−1}k=\min\{\lceil CN^\delta\rceil,N-1\}k=min{⌈CNδ⌉,N−1}, C>0C>0C>0, 0<δ<10<\delta<10<δ<1: with probability 1, for μX\mu_XμX​-a.e. xxx, Z∗(x)\mathcal Z^*(x)Z∗(x) is nonempty, the argmin is eventually nonempty, and every selection z^N(x)\hat z_N(x)z^N​(x) from it satisfies

lim⁡N→∞C(z^N(x)∣x)=v∗(x)andlim⁡N→∞inf⁡z∈Z∗(x)∥z^N(x)−z∥=0.\lim_{N\to\infty}C(\hat z_N(x)\mid x)=v^*(x)\qquad\text{and}\qquad\lim_{N\to\infty}\inf_{z\in\mathcal Z^*(x)}\|\hat z_N(x)-z\|=0 .N→∞lim​C(z^N​(x)∣x)=v∗(x)andN→∞lim​z∈Z∗(x)inf​∥z^N​(x)−z∥=0.

The goal covers both cases of Assumption 5, including unbounded Z\mathcal ZZ such as the newsvendor's [0,∞)[0,\infty)[0,∞).

Milestones, in the order the proof uses them

  1. Walk step (proof of Theorem 15, p. 47). For any measurable hhh with E∣h(Y)∣<∞\mathbb E|h(Y)|<\inftyE∣h(Y)∣<∞, ∑iwN,i(x)h(yi)→E[h(Y)∣X=x]\sum_i w_{N,i}(x)h(y^i)\to\mathbb E[h(Y)\mid X=x]∑i​wN,i​(x)h(yi)→E[h(Y)∣X=x] a.s., for μX\mu_XμX​-a.e. xxx.
  2. Lemma 7 (p. 47). Fixed-zzz and fixed-DDD almost-sure convergence upgrade to convergence for all z∈Zz\in\mathcal Zz∈Z simultaneously and weak convergence μ^Y∣x,N→μY∣x\hat\mu_{Y|x,N}\to\mu_{Y|x}μ^​Y∣x,N​→μY∣x​, off one null set.
  3. Lemma 5 (p. 45). Along one sample path, pointwise convergence C^N(⋅∣x)→C(⋅∣x)\widehat C_N(\cdot\mid x)\to C(\cdot\mid x)CN​(⋅∣x)→C(⋅∣x) is uniform on compact subsets of Z\mathcal ZZ.
  4. Lemma 6, case 1 (pp. 45–47). For bounded Z\mathcal ZZ, the minima and minimizers of C^N(⋅∣x)\widehat C_N(\cdot\mid x)CN​(⋅∣x) converge to v∗(x)v^*(x)v∗(x) and to Z∗(x)\mathcal Z^*(x)Z∗(x).

Significance

Theorem 5 says that a decision computed from data alone, with no model of the conditional distribution, performs asymptotically as well as the decision a decision maker with full knowledge of μY∣x\mu_{Y|x}μY∣x​ would take, for almost every covariate value and under mild conditions on the cost. It justifies the kNN prescription in the paper's newsvendor and shipment-planning experiments, and its proof skeleton (a pointwise strong law for the weights, then Lemmas 5–7) is reused for kernel, local-linear and recursive-kernel weights (Theorems 6–9), which could follow as missions of the same shape.

The result is proved in the paper; it has not been machine-checked anywhere. Formalizing it requires a strong law for nearest-neighbour regression with integrable responses (Walk 2010, building on Devroye, Györfi, Krzyżak and Lugosi 1994), which Mathlib does not have, and a careful treatment of conditional laws at a point. One of the paper's lemmas (Lemma 6) is false as printed for unbounded Z\mathcal ZZ; the formalization isolates the correct statement.

Difficulty

The deterministic optimization part (Lemmas 5 and 6 for bounded Z\mathcal ZZ) is a compactness argument. The central difficulty is the probabilistic input. A natural first idea, applying the strong law of large numbers to C^N(z∣x)\widehat C_N(z\mid x)CN​(z∣x), fails: the kNN weights depend on all of x1,…,xNx^1,\dots,x^Nx1,…,xN and on xxx, the effective sample size kNk_NkN​ grows sublinearly, and convergence is required for almost every xxx simultaneously, including xxx that are atoms of μX\mu_XμX​ where ties are the rule. A second difficulty is the exchange of quantifiers: the strong law gives a null set depending on zzz and on the set DDD, and the conclusion needs one null set for all z∈Zz\in\mathcal Zz∈Z. Finally, for unbounded Z\mathcal ZZ the minimizers of C^N(⋅∣x)\widehat C_N(\cdot\mid x)CN​(⋅∣x) must be kept bounded, and weak convergence of μ^Y∣x,N\hat\mu_{Y|x,N}μ^​Y∣x,N​ alone does not do so.

Formalization scope

All spaces are EuclideanSpace ℝ (Fin d), so kNN distances and ∥z−z′∥\|z-z'\|∥z−z′∥ are Euclidean. The data are a measurable, mutually independent sequence S i : Ω → ℝ^{d_x} × ℝ^{d_y} with each term of law μ\muμ; samples are indexed from 000. The following readings are explicit in the Lean:

  • μY∣x\mu_{Y|x}μY∣x​ is the fixed version μ.condKernel x of the conditional law, and C(z∣x)C(z\mid x)C(z∣x) is its Bochner integral; Assumption 5's "for every x∈Xx\in\mathcal Xx∈X" refers to this version, and DxD_xDx​ is measurable.
  • Quantifier order is "with probability 1, for μX\mu_XμX​-a.e. xxx" (Definition 1), and the selection z^N\hat z_Nz^N​ is quantified inside both a.e. quantifiers: Z∗(x)\mathcal Z^*(x)Z∗(x) is nonempty, the argmin is nonempty for all large NNN, and every sequence lying in it for all large NNN is covered.
  • v∗(x)v^*(x)v∗(x) is never a real infimum; it is read as C(z⋆∣x)C(z^\star\mid x)C(z⋆∣x) for z⋆∈Z∗(x)z^\star\in\mathcal Z^*(x)z⋆∈Z∗(x), nonempty a.e. by Assumption 3.
  • Ties are broken lower-index-first, so the weights sum to 111 for N≥2N\ge2N≥2; at N≤1N\le1N≤1 the formula gives k=0k=0k=0 and all weights 000.
  • Lemmas 5–7 assume weights that are eventually nonnegative and sum to 111, the reading of the paper's Eμ^Y∣x,N\mathbb E_{\hat\mu_{Y|x,N}}Eμ^​Y∣x,N​​ notation; weak convergence is tested against bounded continuous functions.
  • Lemma 6 is posed for bounded Z\mathcal ZZ only, without its unused weak-convergence hypothesis; the goal is posed for both cases.

A formalization that replaces the conditional law by an arbitrary kernel unrelated to μ\muμ, takes z^N\hat z_Nz^N​ to be a fixed measurable selection chosen outside the almost-sure quantifier, or reads inf⁡z∈Z∗(x)\inf_{z\in\mathcal Z^*(x)}infz∈Z∗(x)​ on an empty set as 000 without Assumption 3 would trivialize or weaken the result and is ruled out by the statements above.

Needed infrastructure: nearest-neighbour ranks and their combinatorics (Stone's lemma: a point is among the kkk nearest neighbours of at most a bounded number of others), a strong law for kNN regression, portmanteau-type characterizations of weak convergence for weighted empirical measures, and stability of minimizers under locally uniform convergence. The kNN strong law and the stability lemmas are reusable well beyond this mission. Related platform items, credited but not reused (they concern unconditional SAA): SolutionQuality.SRP.prop1_i, SolutionQuality.SRP.prop1_ii, SolutionQuality.SRP.fbar_tendstoUniformlyOn (Bayraksan–Morton 2006) and DupacovaWets.Consistency.consistency_of_measurable_estimates (Dupačová–Wets 1988). Proofs of any milestone, and alternative routes to the goal, are welcome.

Selected references

  • D. Bertsimas, N. Kallus, From Predictive to Prescriptive Analytics, arXiv:1402.5481v4, 2018; Management Science 66(3), 2020. https://arxiv.org/abs/1402.5481v4 , https://doi.org/10.1287/mnsc.2018.3253
  • H. Walk, Strong laws of large numbers and nonparametric estimation, in Recent Developments in Applied Probability and Statistics, Physica-Verlag, 2010. https://doi.org/10.1007/978-3-7908-2598-5_8
  • L. Devroye, L. Györfi, A. Krzyżak, G. Lugosi, On the strong universal consistency of nearest neighbor regression function estimates, Annals of Statistics 22(3), 1994. https://doi.org/10.1214/aos/1176325633
  • J. Dupačová, R. Wets, Asymptotic behavior of statistical estimators and of optimal solutions of stochastic optimization problems, Annals of Statistics 16(4), 1988. https://doi.org/10.1214/aos/1176351052
  • A. Shapiro, Monte Carlo sampling methods, in Handbooks in OR & MS 10, 2003. https://doi.org/10.1016/S0927-0507(03)10006-0
6 thms1 active userReviewed
Machine LearningProbabilityStatistics·Captain: mikedeng1

From Predictive to Prescriptive Analytics 2: For Costs in [0, c̄] That Are L-Lipschitz, w.p. ≥ 1 − δ Every Decision Rule in F Costs ≤ Its Empirical Cost + c̄√(log(1/δ)/2N) + L·ℜ_N(F)Research Paper

Motivation

Many operations decisions are made after observing side information: an inventory manager sees weather, search trends and calendar effects before ordering; a supply-chain planner sees market signals before shipping. Bertsimas and Kallus, in From Predictive to Prescriptive Analytics (arXiv:1402.5481v4, Management Science 2020), study how to turn such data into decisions. Besides their local, weight-based prescriptions (the subject of mission 1 of this series), their §8 analyses a second approach: choose a decision rule z(⋅)z(\cdot)z(⋅) from a structured class, such as norm-bounded linear rules, by empirical risk minimization — minimizing the average historical cost of the rule.

The question this mission formalizes is the one any practitioner of that approach must answer: how far can the true expected cost of the chosen rule exceed its in-sample cost? Classical statistical learning answers it with Rademacher complexity (Bartlett and Mendelson, JMLR 2002), but only for real-valued predictors. Operations decisions are vectors (order quantities for several products, shipments to several locations), so the paper extends the theory to multivariate decision rules.

Setting

Covariates xxx take values in a measurable space X\mathcal XX, the uncertain quantity yyy in a measurable space Y\mathcal YY, and decisions zzz in Rd\mathbb R^{d}Rd, unconstrained. Data are an i.i.d. sample SN=((x1,y1),…,(xN,yN))S_N=((x^1,y^1),\dots,(x^N,y^N))SN​=((x1,y1),…,(xN,yN)) from a probability measure μ\muμ on X×Y\mathcal X\times\mathcal YX×Y, and SNx=(x1,…,xN)S_N^x=(x^1,\dots,x^N)SNx​=(x1,…,xN) is its covariate part. A cost c(z;y)c(z;y)c(z;y) is incurred when decision zzz meets outcome yyy. A decision rule is a measurable map z(⋅):X→Rdz(\cdot):\mathcal X\to\mathbb R^dz(⋅):X→Rd; F\mathcal FF is a class of decision rules. Its expected cost is E[c(z(X);Y)]\mathbb E[c(z(X);Y)]E[c(z(X);Y)], and its empirical cost is 1N∑i=1Nc(z(xi);yi)\frac1N\sum_{i=1}^N c(z(x^i);y^i)N1​∑i=1N​c(z(xi);yi), the objective of problem (4).

The empirical multivariate Rademacher complexity of F\mathcal FF (Definition 3) on a sample s1,…,sNs_1,\dots,s_Ns1​,…,sN​ is

R^N(F;SN)=Eσ[2Nsup⁡g∈F∑i=1N∑k=1dσik gk(si)],\widehat{\mathfrak R}_N(\mathcal F;S_N)=\mathbb E_\sigma\Big[\frac2N\sup_{g\in\mathcal F}\sum_{i=1}^N\sum_{k=1}^d\sigma_{ik}\,g_k(s_i)\Big],RN​(F;SN​)=Eσ​[N2​g∈Fsup​i=1∑N​k=1∑d​σik​gk​(si​)],

with σik\sigma_{ik}σik​ independent uniform signs, and the marginal complexity RN(F)\mathfrak R_N(\mathcal F)RN​(F) is its expectation over the sample. For real-valued classes (d=1d=1d=1) these are the usual Rademacher complexities with factor 2/N2/N2/N and no absolute value. Lean notation: empRademacher, margRademacher, expCost, empCost, costClass in PrescAnalytics.ERM.

Formalization targets

Goal: Theorem 13 (out-of-sample guarantees)

If 0≤c(z;y)≤cˉ0\le c(z;y)\le\bar c0≤c(z;y)≤cˉ and ∣c(z;y)−c(z′;y)∣≤L∥z−z′∥∞|c(z;y)-c(z';y)|\le L\|z-z'\|_\infty∣c(z;y)−c(z′;y)∣≤L∥z−z′∥∞​, then for any δ>0\delta>0δ>0 each of the following holds with probability at least 1−δ1-\delta1−δ:

E[c(z(X);Y)]≤1N∑i=1Nc(z(xi);yi)+cˉlog⁡(1/δ)2N+L RN(F)∀z∈F,(26)\mathbb E[c(z(X);Y)]\le\frac1N\sum_{i=1}^N c(z(x^i);y^i)+\bar c\sqrt{\frac{\log(1/\delta)}{2N}}+L\,\mathfrak R_N(\mathcal F)\quad\forall z\in\mathcal F,\tag{26}E[c(z(X);Y)]≤N1​i=1∑N​c(z(xi);yi)+cˉ2Nlog(1/δ)​​+LRN​(F)∀z∈F,(26) E[c(z(X);Y)]≤1N∑i=1Nc(z(xi);yi)+3cˉlog⁡(2/δ)2N+L R^N(F;SNx)∀z∈F.(27)\mathbb E[c(z(X);Y)]\le\frac1N\sum_{i=1}^N c(z(x^i);y^i)+3\bar c\sqrt{\frac{\log(2/\delta)}{2N}}+L\,\widehat{\mathfrak R}_N(\mathcal F;S_N^x)\quad\forall z\in\mathcal F.\tag{27}E[c(z(X);Y)]≤N1​i=1∑N​c(z(xi);yi)+3cˉ2Nlog(2/δ)​​+LRN​(F;SNx​)∀z∈F.(27)

In particular they hold for the empirical risk minimizer. The goal is stated for a general class F\mathcal FF, not only the linear rules (25).

Milestones

  1. Lemma 1 (comparison): for G={(x,y)↦c(f(x);y):f∈F}\mathcal G=\{(x,y)\mapsto c(f(x);y):f\in\mathcal F\}G={(x,y)↦c(f(x);y):f∈F}, R^N(G;SN)≤L R^N(F;SNx)\widehat{\mathfrak R}_N(\mathcal G;S_N)\le L\,\widehat{\mathfrak R}_N(\mathcal F;S_N^x)RN​(G;SN​)≤LRN​(F;SNx​) and RN(G)≤L RN(F)\mathfrak R_N(\mathcal G)\le L\,\mathfrak R_N(\mathcal F)RN​(G)≤LRN​(F).
  2. Theorem 14 (29): for a class G\mathcal GG of functions with values in [0,gˉ][0,\bar g][0,gˉ​], with probability at least 1−δ1-\delta1−δ, Eg≤1N∑ig(ui)+gˉlog⁡(1/δ)/(2N)+RN(G)\mathbb E g\le\frac1N\sum_i g(u^i)+\bar g\sqrt{\log(1/\delta)/(2N)}+\mathfrak R_N(\mathcal G)Eg≤N1​∑i​g(ui)+gˉ​log(1/δ)/(2N)​+RN​(G) for all g∈Gg\in\mathcal Gg∈G.
  3. Theorem 14 (30): the data-dependent version with 3gˉlog⁡(2/δ)/(2N)+R^N(G;SN)3\bar g\sqrt{\log(2/\delta)/(2N)}+\widehat{\mathfrak R}_N(\mathcal G;S_N)3gˉ​log(2/δ)/(2N)​+RN​(G;SN​).

Significance

Theorem 13 says that the quantity empirical risk minimization optimizes is, up to confidence terms that do not depend on the rule, an upper bound on the true expected cost of every rule in the class, uniformly. Bound (27) is computable from data, so it certifies a learned policy without a holdout set. Combined with complexity bounds for norm-restricted linear rules (the paper's Lemmas 2 and 3, not part of this mission), it shows the confidence terms vanish as N→∞N\to\inftyN→∞.

The results are known: (29) and (30) for [0,gˉ][0,\bar g][0,gˉ​]-valued classes are the standard Rademacher bounds (Mohri, Rostamizadeh and Talwalkar, Foundations of Machine Learning, Thm 3.3, after rescaling). What is new in the paper is the ∞\infty∞-norm vector comparison inequality of Lemma 1 with constant 111. None of these statements has a machine-checked proof on the platform. A formal development would supply a reusable symmetrization-plus-McDiarmid pipeline and a vector contraction lemma. Related items posed in this repository are Bartlett–Mendelson's Theorem 8 (RadGauss.RiskBound.theorem_8, with an absolute value in the complexity and constant 8ln⁡(2/δ)/n\sqrt{8\ln(2/\delta)/n}8ln(2/δ)/n​), Maurer's Euclidean vector contraction with constant 2\sqrt22​ (SPOBounds.Margin.maurer_vector_contraction), and the open uniform law HighDimStat.UniformLaws.uniform_law_rademacher_complexity; none is the same statement as a target here.

Difficulty

The uniform bound needs three ingredients: a concentration inequality for the supremum of the deviation over the class (bounded differences), a symmetrization step relating the expected supremum to the Rademacher complexity, and a comparison inequality to pass from the cost class to the decision rules. The scalar contraction principle of Ledoux and Talagrand does not apply directly, since each cost depends on a ddd-dimensional output; applying it coordinatewise loses a factor of ddd, and Euclidean vector contraction loses 2\sqrt22​ and uses the wrong norm. Lemma 1 needs an argument specific to the ∞\infty∞-norm. Measure theory adds its own difficulty: the supremum over an uncountable class must be shown measurable before its expectation or probability means anything.

Formalization scope

Decisions are Fin d → ℝ, whose Lean norm is the ∞\infty∞-norm; X\mathcal XX, Y\mathcal YY are general measurable spaces. Samples are drawn from the product measure Measure.pi (fun _ : Fin N => μ), with N≥1N\ge1N≥1. The following readings are explicit:

  • Each "with probability at least 1−δ1-\delta1−δ, … for all z∈Fz\in\mathcal Fz∈F" is a bound ≤δ\le\delta≤δ on the outer measure of the set of samples where the inequality fails for some zzz; the quantifier over the class is inside the event.
  • Empirical complexities are in EReal (an unbounded class gives +∞+\infty+∞), marginal ones are lower Lebesgue integrals in [0,∞][0,\infty][0,∞]; 0⋅∞=00\cdot\infty=00⋅∞=0.
  • The paper assumes only sup⁡c≤cˉ\sup c\le\bar csupc≤cˉ (resp. ∣g∣≤gˉ|g|\le\bar g∣g∣≤gˉ​). With that alone (26) and (29) are false — a two-point outcome with a constant rule violates them with probability 1/21/21/2 — so the costs are assumed to take values in [0,cˉ][0,\bar c][0,cˉ], keeping every printed constant.
  • The undefined δ′,δ′′\delta',\delta''δ′,δ′′ are δ\deltaδ (the i.i.d. case of the paper's Theorem 21); the Lipschitz denominator printed as ∥zk−zk′∥∞\|z_k-z'_k\|_\infty∥zk​−zk′​∥∞​ is ∥z−z′∥∞\|z-z'\|_\infty∥z−z′∥∞​.
  • Decision rules and the cost are measurable and the class is pointwise separable (a countable subclass approximates every member pointwise). The paper takes probabilities and expectations of suprema over the class without comment; without such a convention the statement fails for pathological classes.

A trivializing formalization is ruled out: the quantifier over the class sits inside the probability, the complexity has no absolute value and is never a junk 000 for unbounded classes, and the hypotheses are satisfied by a clipped newsvendor cost. Solvers will need McDiarmid's inequality, symmetrization with a ghost sample, and the comparison lemma; each is reusable for any Rademacher-based generalization bound. Formalizations of these tools as separate lemmas are welcome.

Selected references

  • D. Bertsimas, N. Kallus, From Predictive to Prescriptive Analytics, arXiv:1402.5481v4, 2018; Management Science 66(3), 2020. https://arxiv.org/abs/1402.5481 , https://doi.org/10.1287/mnsc.2018.3253
  • P. L. Bartlett, S. Mendelson, Rademacher and Gaussian Complexities: Risk Bounds and Structural Results, JMLR 3, 2002. https://www.jmlr.org/papers/v3/bartlett02a.html
  • M. Ledoux, M. Talagrand, Probability in Banach Spaces, Springer, 1991. https://doi.org/10.1007/978-3-642-20212-4
  • M. Mohri, A. Rostamizadeh, A. Talwalkar, Foundations of Machine Learning, 2nd ed., MIT Press, 2018. https://mitpress.mit.edu/9780262039406/
  • A. Maurer, A Vector-Contraction Inequality for Rademacher Complexities, ALT 2016. https://arxiv.org/abs/1605.00251
5 thms1 active userReviewed
Convex OptimizationNumerical AnalysisOptimization·Captain: mikedeng1

Projected Newton Methods for Optimization Problems with Simple Constraints: The Projected Newton Method Converges Superlinearly to the Minimum of a Convex Function over the Nonnegative OrthantResearch Paper

Motivation

Smooth optimization with nonnegative variables appears when the variables are Lagrange multipliers for inequality constraints or when an objective incorporates an augmented Lagrangian or an exact penalty. Bertsekas studies how to retain a Newton-like convergence rate in this setting while using an iteration that projects a scaled gradient step onto the nonnegative orthant. His 1982 paper also discusses large optimal-control examples, where repeatedly solving a quadratic subproblem may be costly. These are the applications motivating the method, rather than assumptions of the theorem. Bertsekas, 1982.

The paper contrasts a partly diagonal scaling matrix with two other approaches: a fully diagonal projected gradient step, which generally has a linear rate, and a constrained Newton step defined by a quadratic program. The main theoretical question is whether the simpler projected iteration can still identify binding coordinates and converge superlinearly near the solution. The mission covers the nonnegative-orthant method in §2 and its stated convergence results. The paper's later extension to general linear constraints and its computational examples concern different objects. Bertsekas, 1982.

Setting

Fix a dimension n≥1n\ge1n≥1 and a continuously differentiable function f:Rn→Rf:\mathbb R^n\to\mathbb Rf:Rn→R. Problem (1) minimizes f(x)f(x)f(x) over the nonnegative orthant R+n={x:xi≥0 for every i}\mathbb R_+^n=\{x:x^i\ge0\text{ for every }i\}R+n​={x:xi≥0 for every i}. The positive part [z]+[z]^+[z]+ replaces each negative coordinate of zzz by zero. A feasible xxx is critical when every partial derivative ∂if(x)\partial_i f(x)∂i​f(x) is nonnegative and ∂if(x)=0\partial_i f(x)=0∂i​f(x)=0 wherever xi>0x^i>0xi>0. This is the paper's componentwise first-order condition.

For a feasible iterate xkx_kxk​, the binding set is B(xk)={i:xki=0}B(x_k)=\{i:x_k^i=0\}B(xk​)={i:xki​=0}. The paper selects a larger working set Ik+I_k^+Ik+​ using a fixed ε>0\varepsilon>0ε>0 and the projected-gradient residual wk=∥xk−[xk−M∇f(xk)]+∥w_k=\|x_k-[x_k-M\nabla f(x_k)]^+\|wk​=∥xk​−[xk​−M∇f(xk​)]+∥, where MMM is fixed, diagonal and positive definite. More precisely, Ik+I_k^+Ik+​ contains the indices with 0≤xki≤min⁡(ε,wk)0\le x_k^i\le\min(\varepsilon,w_k)0≤xki​≤min(ε,wk​) and ∂if(xk)>0\partial_i f(x_k)>0∂i​f(xk​)>0. The matrix DkD_kDk​ is symmetric positive definite and has zero off-diagonal entries in the rows indexed by Ik+I_k^+Ik+​. The projected arc and update are

pk=Dk∇f(xk),xk(a)=[xk−apk]+,xk+1=xk(βmk).p_k=D_k\nabla f(x_k),\qquad x_k(a)=[x_k-a p_k]^+,\qquad x_{k+1}=x_k(\beta^{m_k}).pk​=Dk​∇f(xk​),xk​(a)=[xk​−apk​]+,xk+1​=xk​(βmk​).

Here 0<β<10<\beta<10<β<1, and mkm_kmk​ is the first nonnegative integer passing the paper's two-sum Armijo test (37), with 0<σ<1/20<\sigma<1/20<σ<1/2. One sum uses the scaled gradient outside Ik+I_k^+Ik+​; the other uses the actual projected displacement on Ik+I_k^+Ik+​. These details define the algorithm whose rate is at issue. Bertsekas, 1982, pp. 228–229.

Formalization targets

The first targets establish the fixed-point and descent properties of the projected arc, positivity of the Armijo right-hand side, well-defined step selection, criticality of limit points, and local attraction with finite binding-set identification. Proposition 3 identifies both the working set and the actual binding set:

Ik+=B(xk)=B(x∗)for all sufficiently late k.I_k^+=B(x_k)=B(x^*)\quad\text{for all sufficiently late }k.Ik+​=B(xk​)=B(x∗)for all sufficiently late k.

The goal is Proposition 4. Let fff be convex and C2C^2C2, let x∗x^*x∗ be the unique minimizer on R+n\mathbb R_+^nR+n​ satisfying Assumption (C), and suppose the Hessian quadratic form has uniform positive lower and finite upper bounds on the initial, unrestricted sublevel set. The scaling matrix is Dk=Hk−1D_k=H_k^{-1}Dk​=Hk−1​, where HkH_kHk​ retains Hessian entries except for off-diagonal entries touching Ik+I_k^+Ik+​. Then

xk⟶x∗,∀c>0 ∃K ∀k≥K: ∥xk+1−x∗∥≤c∥xk−x∗∥.x_k\longrightarrow x^*,\qquad \forall c>0\ \exists K\ \forall k\ge K:\ \|x_{k+1}-x^*\|\le c\|x_k-x^*\|.xk​⟶x∗,∀c>0 ∃K ∀k≥K: ∥xk+1​−x∗∥≤c∥xk​−x∗∥.

If the Hessian is Lipschitz in a neighborhood of x∗x^*x∗, the same proposition states an at-least-quadratic error bound: ∥xk+1−x∗∥≤C∥xk−x∗∥2\|x_{k+1}-x^*\|\le C\|x_k-x^*\|^2∥xk+1​−x∗∥≤C∥xk​−x∗∥2 eventually for some C>0C>0C>0. The separate final milestone records the paper's claim that the initial unit trial is eventually accepted. Bertsekas, 1982, pp. 233–236.

Significance

Proposition 4 says that this specific orthant-projected Newton iteration converges to the unique constrained optimum with a superlinear rate under its stated smoothness and curvature conditions. Proposition 3 supplies a distinct finite-identification statement: the working and binding coordinates eventually match those at the optimum. These facts specify the behavior of an algorithm that does not define its direction by solving a quadratic program at every step. The paper proves the preceding propositions and states Proposition 4 with its proof left to the reader, referring to its earlier discussion and standard unconstrained Newton results. Bertsekas, 1982, pp. 234–236.

Formalizing the claims would provide machine-checked statements and, once solved, proofs for the exact projected arc, two-part line search, matrix selection, identification result and rate conclusion. The local mission items are open proof obligations; their compilation checks the definitions and theorem types, not the mathematical claims. The geometry definitions can also support other orthant-constrained algorithms, while the enlarged working set and Armijo rule belong to this paper's method.

Difficulty

A positive definite matrix by itself does not make projection along [x−aD∇f(x)]+[x-aD\nabla f(x)]^+[x−aD∇f(x)]+ a descent move. An off-diagonal coupling can push coordinates against the boundary in a way that defeats the usual unconstrained descent calculation; the paper gives such a situation before Proposition 1. The partly diagonal condition is therefore substantive. A second obstacle is that the exact active set I+(x)I^+(x)I+(x) can jump at a boundary point: iterates approaching that point from the interior need not have the same indexed rows. The enlarged Ik+I_k^+Ik+​ and its dependence on the residual are central to the finite-identification and rate claims. Bertsekas, 1982, pp. 225–229.

Formalization scope

Lean represents Rn\mathbb R^nRn as EuclideanSpace ℝ (Fin n), with the Euclidean norm and zero-based coordinates. The dimension is positive. gradient and the derivative of gradient represent first and second derivatives; the paper's MMM is diag μ with every μi>0\mu^i>0μi>0. The run predicate includes a feasible initial point and the first acceptable integer mkm_kmk​, and the matrix choices are indexed by iteration. Propositions 1–3 use the paper's explicit admissibility condition. Proposition 4 constructs DkD_kDk​ from HkH_kHk​ and does not assume its invertibility, admissibility, convergence or eventual active-set equality.

Assumption (C) includes local C2C^2C2 smoothness, curvature bounds on directions zero at the binding coordinates, and strict complementarity. A local minimum is required to be feasible as well as locally minimal on the orthant. Limit points use subsequential convergence. Superlinearity uses a uniform eventual error inequality, which also covers an iterate that reaches the solution exactly; a ratio with a zero denominator would distort this case. The paper's Proposition 4 display mistakenly binds a direction to the level set while leaving the Hessian's point free. The formalization states the intended reading: every point in the unrestricted initial level set and every direction satisfy the Hessian bounds. The displayed level set has no x≥0x\ge0x≥0 restriction, so neither does the Lean hypothesis.

A definition that accepts any Armijo exponent or a goal that assumes the eventual identification or convergence conclusion would erase the paper's claim. Contributions can build the matrix and projected-arc lemmas, the convergence and identification proofs, and the final rate proof. The local inverse-Hessian observation before Proposition 4 is discussed in the notes but has no separate milestone until its full local setting can be captured without weakening it.

Selected references

  • Dimitri P. Bertsekas, Projected Newton Methods for Optimization Problems with Simple Constraints, SIAM Journal on Control and Optimization 20(2), 221–246, 1982. DOI: 10.1137/0320018.
9 thms1 active userReviewed
OptimizationProbability·Captain: mikedeng1

On the Power of Robust Solutions in Two-Stage Stochastic and Adaptive Optimization Problems 2: On the Non-Symmetric Uniform Simplex, the Robust Optimum Is at Least n + 1 Times the Stochastic OneResearch Paper

Motivation

Many operations problems are decided in two stages: a first-stage decision is fixed before an uncertain quantity is revealed, and a second-stage (recourse) decision is chosen afterwards. Two classical ways to optimize such a problem differ in how they treat the uncertainty. Two-stage stochastic optimization minimizes the expected cost under a probability distribution over scenarios, with a recourse decision adapted to each scenario. Robust optimization fixes a single solution that must be feasible for every scenario in an uncertainty set, and minimizes the worst-case cost. The robust problem is usually far easier to solve, but it is conservative; the question is how much it can lose.

Bertsimas and Goyal (Math. Oper. Res. 35(2), 2010) answer this for uncertain right-hand sides. Their main positive result (Theorem 2.1) bounds the stochasticity gap: if the uncertainty set and the probability measure are both symmetric, the robust optimum is at most twice the stochastic optimum. This mission formalizes the companion negative result, Theorem 2.6: once the uncertainty set is not symmetric, the gap is not bounded by any constant. The example is the simplest non-symmetric set in practice, the corner of the unit simplex, with the uniform distribution on it. It shows that the symmetry hypothesis in the positive result is not an artefact of the proof.

Setting

Fix matrices A∈Rm×n1A\in\mathbb R^{m\times n_1}A∈Rm×n1​, B∈Rm×n2B\in\mathbb R^{m\times n_2}B∈Rm×n2​ and costs c∈R+n1c\in\mathbb R^{n_1}_+c∈R+n1​​, d∈R+n2d\in\mathbb R^{n_2}_+d∈R+n2​​. Let Ω\OmegaΩ be a set of scenarios, b:Ω→R+mb:\Omega\to\mathbb R^m_+b:Ω→R+m​ the right-hand side realized in each scenario, and Ib(Ω)={b(ω)∣ω∈Ω}I_b(\Omega)=\{b(\omega)\mid\omega\in\Omega\}Ib​(Ω)={b(ω)∣ω∈Ω} the uncertainty set. Let μ\muμ be a probability measure on Ω\OmegaΩ.

  • The stochastic problem ΠStoch(b)\Pi_{\mathrm{Stoch}}(b)ΠStoch​(b) (1.1) chooses x≥0x\ge 0x≥0 and a second-stage policy y(ω)≥0y(\omega)\ge 0y(ω)≥0 with Ax+By(ω)≥b(ω)Ax+By(\omega)\ge b(\omega)Ax+By(ω)≥b(ω) for every ω∈Ω\omega\in\Omegaω∈Ω, minimizing cTx+Eμ[dTy(ω)]c^Tx+\mathbb E_\mu[d^Ty(\omega)]cTx+Eμ​[dTy(ω)]. Its optimal value is zStoch(b)z_{\mathrm{Stoch}}(b)zStoch​(b).
  • The robust problem ΠRob(b)\Pi_{\mathrm{Rob}}(b)ΠRob​(b) (1.2) chooses a single pair x≥0x\ge 0x≥0, y≥0y\ge 0y≥0 with Ax+By≥b(ω)Ax+By\ge b(\omega)Ax+By≥b(ω) for every ω\omegaω, minimizing cTx+dTyc^Tx+d^TycTx+dTy. Its optimal value is zRob(b)z_{\mathrm{Rob}}(b)zRob​(b).

In general some coordinates of xxx and yyy are required to be integers; in this mission's instance there are none. A set P⊆RnP\subseteq\mathbb R^nP⊆Rn is symmetric (Definition 1.2) if some u0∈Pu^0\in Pu0∈P satisfies u0+z∈P  ⟺  u0−z∈Pu^0+z\in P\iff u^0-z\in Pu0+z∈P⟺u0−z∈P for every z∈Rnz\in\mathbb R^nz∈Rn.

The instance of Theorem 2.6 has no first stage (n1=0n_1=0n1​=0, so A=0A=0A=0, c=0c=0c=0), n2=m=n≥3n_2=m=n\ge 3n2​=m=n≥3, B=InB=I_nB=In​, and d=en=(0,…,0,1)d=e_n=(0,\dots,0,1)d=en​=(0,…,0,1). The uncertainty set is the corner simplex (2.29)

Ib(Ω)={ b∈Rn  :  ∑j=1nbj≤1, b≥0 },I_b(\Omega)=\Big\{\,b\in\mathbb R^n \;:\; \sum_{j=1}^n b_j\le 1,\ b\ge 0\,\Big\},Ib​(Ω)={b∈Rn:j=1∑n​bj​≤1, b≥0},

and μ\muμ is the uniform probability measure on it: μ(S)=volume⁡({b(ω)∣ω∈S})/volume⁡(Ib(Ω))\mu(S)=\operatorname{volume}(\{b(\omega)\mid\omega\in S\})/\operatorname{volume}(I_b(\Omega))μ(S)=volume({b(ω)∣ω∈S})/volume(Ib​(Ω)).

Formalization targets

Goal: Theorem 2.6

On this instance,

zRob(b) ≥ (n+1)⋅zStoch(b).z_{\mathrm{Rob}}(b)\ \ge\ (n+1)\cdot z_{\mathrm{Stoch}}(b).zRob​(b) ≥ (n+1)⋅zStoch​(b).

The constant n+1n+1n+1 is the paper's, and it is exact: the robust optimum is 111 and the stochastic optimum is at most 1/(n+1)1/(n+1)1/(n+1).

Milestones

  1. Lemma 2.4. The corner simplex is not symmetric for n≥2n\ge 2n≥2.
  2. Robust side (p. 20). Every robust-feasible yyy has yj≥1y_j\ge 1yj​≥1 for all jjj, so zRob(b)≥1z_{\mathrm{Rob}}(b)\ge 1zRob​(b)≥1.
  3. Eq. (2.30). The policy y^(ω)=b(ω)\hat y(\omega)=b(\omega)y^​(ω)=b(ω) is feasible for ΠStoch(b)\Pi_{\mathrm{Stoch}}(b)ΠStoch​(b), so zStoch(b)≤Eμ[bn(ω)]z_{\mathrm{Stoch}}(b)\le\mathbb E_\mu[b_n(\omega)]zStoch​(b)≤Eμ​[bn​(ω)].
  4. Eq. (2.32), denominator. vol⁡(Ib(Ω))=1/n!\operatorname{vol}(I_b(\Omega))=1/n!vol(Ib​(Ω))=1/n!.
  5. Eq. (2.32), numerator. ∫Ib(Ω)xn dx=1/(n+1)!\int_{I_b(\Omega)}x_n\,dx=1/(n+1)!∫Ib​(Ω)​xn​dx=1/(n+1)!.
  6. Eqs. (2.31)–(2.32). Eμ[bn(ω)]=1/(n+1)\mathbb E_\mu[b_n(\omega)]=1/(n+1)Eμ​[bn​(ω)]=1/(n+1).

Significance

The result. Together with Theorem 2.1, Theorem 2.6 locates the boundary of the paper's positive theory. Theorem 2.1 says a static robust solution loses at most a factor of 222 against the fully adaptive stochastic optimum under symmetry; Theorem 2.6 says that without symmetry the factor can be n+1n+1n+1, hence arbitrarily large as the dimension grows. The paper's later results (the bound for "positive" uncertainty sets, Theorem 2.7) are motivated by this example: some structural condition replacing symmetry is necessary. The same instance also separates the adaptive problem from the stochastic one (p. 21), which shows that the gap comes from comparing a worst case with an expectation, not from the lack of adaptivity.

Formalizing it. The theorem is proved in the paper; no machine-checked version exists. The formalization adds two things beyond the paper. First, it pins down the problems as optimization problems over integrable policies with values in the extended reals, so that the inequality cannot hold for a degenerate reason. Second, it requires the volume of the corner simplex and the first moment of a coordinate on it, which the paper calls "standard computation". Neither is in Mathlib at the time of writing: Mathlib has the standard simplex as a convex set but no volume formula for it.

Difficulty

The optimization part is short: one feasible stochastic policy and the vertices of the simplex give both bounds. The work is in the measure theory. The volume 1/n!1/n!1/n! and the moment 1/(n+1)!1/(n+1)!1/(n+1)! are iterated integrals with variable upper limits, and the paper calls them "standard computation"; as statements about Lebesgue measure on Rn\mathbb R^nRn they are dimension-dependent identities that Mathlib does not contain, for the simplex or for the convex hull of n+1n+1n+1 points. The other technical point is the scenario model: the uniform measure lives on Ω\OmegaΩ, while the integrals live on Rn\mathbb R^nRn, so the expectation must be moved through the push-forward of μ\muμ by bbb.

Formalization scope

  • Vectors are Fin k → ℝ with the componentwise order; products are A *ᵥ x and c ⬝ᵥ x. The paper's nnn-th coordinate is index n−1n-1n−1 of Fin n, and ene_nen​ is lastUnit n.
  • The optimal values zRobz_{\mathrm{Rob}}zRob​ and zStochz_{\mathrm{Stoch}}zStoch​ are EReal infima over the feasible set, equal to +∞+\infty+∞ when the problem is infeasible. No optimal solution is assumed to exist.
  • Policies of ΠStoch(b)\Pi_{\mathrm{Stoch}}(b)ΠStoch​(b) are μ\muμ-integrable functions Ω→Rn\Omega\to\mathbb R^nΩ→Rn (the paper takes their expectation), and the constraints hold for every scenario, as printed, not almost surely.
  • The integer coordinates are given as a set of indices; the instance uses the empty set (p2=0p_2=0p2​=0, as in the displays on p. 20). The generic definitions keep the mixed-integer domain so that they agree with the other missions of this series.
  • The uniform measure is encoded by quantifying over every scenario model (Ω,μ,b)(\Omega,\mu,b)(Ω,μ,b) in which μ\muμ is a probability measure, bbb is measurable, the range of bbb is the corner simplex, and the push-forward of μ\muμ by bbb equals Lebesgue measure restricted to the simplex divided by its volume. The model Ω=\Omega=Ω= simplex, b=b=b= identity satisfies these hypotheses; the paper's formula for μ(S)\mu(S)μ(S) is read as a statement about this push-forward.
  • The hypothesis n≥3n\ge 3n≥3 is the paper's and is kept; the argument appears to need only n≥1n\ge 1n≥1.
  • Ruled out: a real-valued infimum for zStochz_{\mathrm{Stoch}}zStoch​, which would be a junk 000 on an infeasible problem and make the goal trivially true; an arbitrary measure, or a Dirac mass at the mean, in place of the uniform measure; and constraints only μ\muμ-almost surely.

Needed infrastructure: the volume of the corner simplex in Rn\mathbb R^nRn and the integral of a coordinate over it (reusable well beyond this mission, for example for Dirichlet distributions and order statistics), and the change of variables from Ω\OmegaΩ to Rn\mathbb R^nRn through the push-forward. Contributions of either simplex integral, in any form that implies the stated milestones, are welcome.

Selected references

  • D. Bertsimas, V. Goyal, On the power of robust solutions in two-stage stochastic and adaptive optimization problems, Mathematics of Operations Research 35(2), 284–305, 2010. https://doi.org/10.1287/moor.1090.0440 (formalized from the authors' manuscript, MIT DSpace / MIT Open Access Articles)
  • A. Ben-Tal, A. Nemirovski, Robust solutions of uncertain linear programs, Operations Research Letters 25(1), 1–13, 1999. https://doi.org/10.1016/S0167-6377(99)00016-4
  • D. Bertsimas, M. Sim, The price of robustness, Operations Research 52(1), 35–53, 2004. https://doi.org/10.1287/opre.1030.0065
  • J. R. Birge, F. Louveaux, Introduction to Stochastic Programming, 2nd ed., Springer, 2011. https://doi.org/10.1007/978-1-4614-0237-4
11 thms1 active userReviewed
Analysis·Captain: mikedeng1

Regret in Decision Making under Uncertainty 2: With u(x, y) = x + f(x − y) and f Decreasingly Concave, the Same Decision Maker Takes a Fair Long-Odds Bet and Buys Fair InsuranceResearch Paper

Motivation

Expected utility theory explains a single attitude toward risk by the curvature of a utility function of final wealth: a concave utility makes a decision maker risk averse, a convex one risk seeking. Several well documented patterns of behaviour do not fit this picture. The same people buy insurance against small-probability losses and buy lottery tickets with small-probability gains (Friedman and Savage, 1948, doi:10.1086/256692); subjects who are risk averse for gains are risk seeking for the mirrored losses, the reflection effect; and subjects who are indifferent between full insurance and bearing a risk reject half-price insurance that pays only half the time, probabilistic insurance (Kahneman and Tversky, 1979, doi:10.2307/1914185).

David E. Bell's 1982 paper Regret in Decision Making under Uncertainty (doi:10.1287/opre.30.5.961) proposes that a decision maker cares not only about what she obtains but also about what she would have obtained had she chosen differently. Its first part (the companion mission of this series) derives the representation u(x,y)=αv(x)+f(v(x)−v(y))u(x,y)=\alpha v(x)+f(v(x)-v(y))u(x,y)=αv(x)+f(v(x)−v(y)); its Section 2 shows that the single additive form with linear vvv explains all of the patterns above at once, provided the regret function fff is decreasingly concave. Regret theory, developed at the same time by Loomes and Sugden (1982, doi:10.2307/2232669), remains one of the standard alternatives to expected utility in decision analysis and behavioural economics.

Setting

A decision maker chooses between two alternatives aaa and bbb. Uncertainty is described by finitely many states iii with probabilities πi\pi_iπi​; in state iii alternative aaa gives final assets aia_iai​ and bbb gives bib_ibi​. Choosing aaa means forgoing bbb, and an outcome is valued by the regret utility

u(x,y)=x+f(x−y),u(x,y)=x+f(x-y),u(x,y)=x+f(x−y),

where xxx is the final assets received, yyy the assets the foregone alternative would have given in the same state, and f:R→Rf:\mathbb R\to\mathbb Rf:R→R the regret function. The expected utility of selecting aaa when bbb is foregone is

EUπ(a ∣ b)=∑iπi u(ai,bi),\mathrm{EU}_\pi(a\,|\,b)=\sum_i \pi_i\,u(a_i,b_i),EUπ​(a∣b)=i∑​πi​u(ai​,bi​),

and aaa is strictly preferred to bbb when EUπ(b ∣ a)<EUπ(a ∣ b)\mathrm{EU}_\pi(b\,|\,a)<\mathrm{EU}_\pi(a\,|\,b)EUπ​(b∣a)<EUπ​(a∣b) (the paper's inequality (1)); aaa and bbb are indifferent when the two expected utilities are equal.

The comparisons are given by three payoff tables. Table II: a horse wins with probability ppp; betting \pgivesgivesgives1-pifitwinsandif it wins andifitwinsand-pifitloses,notbettinggivesif it loses, not betting givesifitloses,notbettinggives0.TableIII:acarisdamagedwithprobability. Table III: a car is damaged with probability .TableIII:acarisdamagedwithprobabilitypatacostofat a cost ofatacostof1;insuringgives; insuring gives ;insuringgives-pinbothstates,notinsuringgivesin both states, not insuring givesinbothstates,notinsuringgives-1ororor0.TableIV:anaccidenthasprobability. Table IV: an accident has probability .TableIV:anaccidenthasprobabilityq;selfinsurancegives; self insurance gives ;selfinsurancegives-1ororor0,fullinsurance, full insurance ,fullinsurance-palways,andprobabilisticinsurance,whichchargeshalfthepremiumandpayswithprobabilityonehalfgivenanaccident,givesalways, and probabilistic insurance, which charges half the premium and pays with probability one half given an accident, givesalways,andprobabilisticinsurance,whichchargeshalfthepremiumandpayswithprobabilityonehalfgivenanaccident,gives-p,, ,-1ororor-p/2$.

The regret function fff is decreasingly concave if it is twice differentiable with strictly increasing second derivative f′′f''f′′. Section 3 also uses g(r)=r+f(r)−f(−r)g(r)=r+f(r)-f(-r)g(r)=r+f(r)−f(−r) and the negative exponential f(r)=1−e−γrf(r)=1-e^{-\gamma r}f(r)=1−e−γr, γ>0\gamma>0γ>0.

Formalization targets

Goal: coexistence of insurance and gambling

If fff is decreasingly concave and 0<p<120<p<\tfrac120<p<21​, then the bet of Table II is strictly preferred to not betting and insuring in Table III is strictly preferred to not insuring:

p [f(1−p)−f(p−1)]>(1−p) [f(p)−f(−p)].p\,[f(1-p)-f(p-1)]>(1-p)\,[f(p)-f(-p)].p[f(1−p)−f(p−1)]>(1−p)[f(p)−f(−p)].

Both comparisons have equal expected values, so the statement is a pure regret effect. The goal fixes no constants and assumes nothing about fff beyond decreasing concavity; it is the weakest hypothesis the page offers.

Milestones

  1. Display (5): the bet is preferred iff the inequality above; display (6): insurance is preferred iff the same inequality, so the two preferences coincide.
  2. The inequality holds for 0<p<120<p<\tfrac120<p<21​ whenever (f(x)−f(−x))/x(f(x)-f(-x))/x(f(x)−f(−x))/x is strictly increasing on x>0x>0x>0; and that ratio is strictly increasing whenever f′′(x)>f′′(−x)f''(x)>f''(-x)f′′(x)>f′′(−x) for x>0x>0x>0.
  3. The reflection effect (Sec. 2(ii)): x2x_2x2​ for sure is indifferent to the lottery (p:x1; 1−p:x3)(p:x_1;\,1-p:x_3)(p:x1​;1−p:x3​) iff −x2-x_2−x2​ for sure is indifferent to (p:−x1; 1−p:−x3)(p:-x_1;\,1-p:-x_3)(p:−x1​;1−p:−x3​).
  4. Probabilistic insurance (Sec. 2(iii)): display (9) characterises indifference between full and self insurance; under (9) and 0≤q<10\le q<10≤q<1,
f(p/2)−f(−p/2)<12 [f(p)−f(−p)](11)f(p/2)-f(-p/2)<\tfrac12\,[f(p)-f(-p)]\qquad(11)f(p/2)−f(−p/2)<21​[f(p)−f(−p)](11)

holds iff full insurance is strictly preferred to probabilistic insurance, iff probabilistic insurance is strictly preferred to self insurance; and (11) follows from the same ratio condition. 5. The negative exponential (Sec. 3): fff concave, ggg convex on r≥0r\ge0r≥0, g(r)/rg(r)/rg(r)/r strictly increasing on r>0r>0r>0, and fff decreasingly concave.

Significance

The goal is the paper's central behavioural claim: one regret function explains why the same decision maker gambles on long odds and buys insurance at fair prices, behaviour that a single concave or convex utility of wealth cannot produce. The milestones show that the reflection effect and the rejection of probabilistic insurance follow from the same model, and two of them reduce to the same inequality on the odd part f(x)−f(−x)f(x)-f(-x)f(x)−f(−x) of the regret function. The negative exponential example shows the hypotheses are met by a concrete, commonly used function, so the explanation is not vacuous.

The results are proved in the paper by short calculations, and none of them has been machine-checked before. The mission produces a formal statement of the regret model of Section 2 with its payoff tables, which pins down several points the printed text leaves loose: the meaning of "decreasingly concave", the domain on which "increasing" and "convex" are meant, and three misprinted intermediate displays.

Difficulty

The comparisons (5), (6), (9) and the reflection identity are algebraic once the payoff tables are expanded; the work is in expanding them correctly, since the foregone argument of uuu changes with the alternative selected and three printed expansions are wrong. The substance lies in the analytic step: passing from a pointwise condition on f′′f''f′′ to strict monotonicity of (f(x)−f(−x))/x(f(x)-f(-x))/x(f(x)−f(−x))/x on (0,∞)(0,\infty)(0,∞), with strict inequalities throughout and through a quotient that is undefined at 000. The naive reading of the page's condition, "f′′(x)>f′′(−x)f''(x)>f''(-x)f′′(x)>f′′(−x) for all xxx", is unsatisfiable and cannot be used as a hypothesis.

Formalization scope

All quantities are real numbers. The utility is regretU f x y = x + f (x - y), i.e. form (4) with v(x)=xv(x)=xv(x)=x exactly; the page assumes vvv only "approximately linear" but computes every display with v(x)=xv(x)=xv(x)=x. A comparison is a finite state space Fin n with a real weight vector; the theorems instantiate it with (p,1−p)(p,1-p)(p,1−p) for Tables II–III and with (q/2,q/2,1−q)(q/2,q/2,1-q)(q/2,q/2,1−q) for Table IV, the accident state split by whether the insurer pays. Strict preference is strict inequality of expected regret utilities; indifference is equality.

Conventions and corrections, each disclosed in the item concerned:

  • "Decreasingly concave" is not defined in the paper; it is read as Differentiable ℝ f, Differentiable ℝ (deriv f) and StrictMono (deriv (deriv f)). The page's presumption that fff is increasing and concave is not assumed.
  • "Increasing" is read strictly, because the conclusions (5), (6), (11) are strict, and on x>0x>0x>0, because (f(x)−f(−x))/x(f(x)-f(-x))/x(f(x)−f(−x))/x is undefined at 000 and even.
  • "f′′(x)>f′′(−x)f''(x)>f''(-x)f′′(x)>f′′(−x) for all xxx" is read for x>0x>0x>0.
  • The convexity of ggg in the exponential example is stated on r≥0r\ge0r≥0: g(r)=r+2sinh⁡(γr)g(r)=r+2\sinh(\gamma r)g(r)=r+2sinh(γr) is concave for r<0r<0r<0, as the paper itself notes on p. 977.
  • Misprints corrected: the last bracket of the second reflection equality on p. 973; f(p/2)f(p/2)f(p/2) for f(−p/2)f(-p/2)f(−p/2) in the expansion before (10) and in the probabilistic-versus-self display on p. 975, which also lacks a factor qqq. The normalisation f(0)=0f(0)=0f(0)=0 is not needed.
  • The equivalences (5), (6), (9) and the reflection identity hold for every real probability and are stated without a range; the goal keeps 0<p<120<p<\tfrac120<p<21​ and the probabilistic-insurance statement keeps 0≤q<10\le q<10≤q<1.

The goal cannot be satisfied vacuously: the negative exponential milestone exhibits a decreasingly concave fff, and the conclusion is a strict preference between two alternatives of equal expected value. Section 2(iv) (preference reversals) is not formalized: its printed rewriting of (12) in terms of hhh drops a linear term and its conclusion is unquantified ("for sufficiently small units of money").

Contributions welcome: proofs of the algebraic reductions, the slope argument for (f(x)−f(−x))/x(f(x)-f(-x))/x(f(x)−f(−x))/x, and the convexity facts for sinh⁡\sinhsinh; the latter two are reusable beyond this mission.

Selected references

  • D. E. Bell, Regret in Decision Making under Uncertainty, Operations Research 30(5), 961–981, 1982. doi:10.1287/opre.30.5.961
  • D. Kahneman and A. Tversky, Prospect Theory: An Analysis of Decision under Risk, Econometrica 47(2), 263–291, 1979. doi:10.2307/1914185
  • G. Loomes and R. Sugden, Regret Theory: An Alternative Theory of Rational Choice under Uncertainty, The Economic Journal 92(368), 805–824, 1982. doi:10.2307/2232669
  • M. Friedman and L. J. Savage, The Utility Analysis of Choices Involving Risk, Journal of Political Economy 56(4), 279–304, 1948. doi:10.1086/256692
12 thms1 active userReviewed
OptimizationProbability·Captain: mikedeng1

On the Power of Robust Solutions in Two-Stage Stochastic and Adaptive Optimization Problems 3: With Uniform Hypercube Cost Uncertainty, the Robust Optimum Is at Least n + 1 Times the Stochastic OneResearch Paper

Motivation

Many planning problems are made in two stages: a first decision xxx is fixed before an uncertain parameter is revealed, and a second decision yyy is taken afterwards. Two models compete for such problems. Two-stage stochastic optimization assumes a probability distribution over the scenarios and minimizes the expected cost, letting the second-stage decision depend on the scenario. Robust optimization asks for one solution that is feasible in every scenario and minimizes the worst-case cost. The robust problem is usually far easier to solve: it is a single deterministic problem, while the stochastic problem optimizes over policies. The question is how much cost one gives up by solving the easy problem instead of the hard one.

Bertsimas and Goyal (Math. Oper. Res. 2010) measure this loss by the stochasticity gap, the ratio of the robust optimum to the stochastic optimum. Their Theorem 2.1 shows that when only the right-hand side of the constraints is uncertain, the uncertainty set is symmetric and the distribution is centred at the point of symmetry, the robust optimum is at most twice the stochastic one. Section 3 asks whether the same holds when the second-stage costs are uncertain as well, and answers no with an explicit instance: Theorem 3.1. This mission formalizes that instance.

Setting

There are no first-stage variables. The second-stage decision is a vector y∈R+ny\in\mathbb R^n_+y∈R+n​ with n≥1n\ge1n≥1 continuous coordinates, subject to the single covering constraint

y1+y2+⋯+yn ≥ 1,y_1+y_2+\dots+y_n\ \ge\ 1 ,y1​+y2​+⋯+yn​ ≥ 1,

the constraint By≥bBy\ge bBy≥b with B=[1,1,…,1]∈R1×nB=[1,1,\dots,1]\in\mathbb R^{1\times n}B=[1,1,…,1]∈R1×n and b=1b=1b=1. A set Ω\OmegaΩ of scenarios carries a cost map d:Ω→Rnd:\Omega\to\mathbb R^nd:Ω→Rn; in scenario ω\omegaω the second-stage cost is d(ω)Tyd(\omega)^{\mathsf T}yd(ω)Ty. The uncertainty set is I(b,d)(Ω)={(b(ω),d(ω)):ω∈Ω}I_{(b,d)}(\Omega)=\{(b(\omega),d(\omega)) : \omega\in\Omega\}I(b,d)​(Ω)={(b(ω),d(ω)):ω∈Ω}, here {1}×[0,1]n\{1\}\times[0,1]^n{1}×[0,1]n: the range of ddd is the whole cube [0,1]n[0,1]^n[0,1]n. A probability measure μ\muμ on Ω\OmegaΩ makes the coordinates d1,…,dnd_1,\dots,d_nd1​,…,dn​ independent, each uniformly distributed on [0,1][0,1][0,1].

The stochastic problem ΠStoch(b,d)\Pi_{\mathrm{Stoch}}(b,d)ΠStoch​(b,d), display (1.4) of the paper, chooses a policy ω↦y(ω)≥0\omega\mapsto y(\omega)\ge0ω↦y(ω)≥0 satisfying the constraint in every scenario and minimizes Eμ[d(ω)Ty(ω)]\mathbb E_\mu[d(\omega)^{\mathsf T}y(\omega)]Eμ​[d(ω)Ty(ω)]; its optimal value is zStoch(b,d)z_{\mathrm{Stoch}}(b,d)zStoch​(b,d). The robust problem ΠRob(b,d)\Pi_{\mathrm{Rob}}(b,d)ΠRob​(b,d), display (1.5), chooses one y≥0y\ge0y≥0 satisfying the constraint and minimizes max⁡ω∈Ωd(ω)Ty\max_{\omega\in\Omega} d(\omega)^{\mathsf T}ymaxω∈Ω​d(ω)Ty; its optimal value is zRob(b,d)z_{\mathrm{Rob}}(b,d)zRob​(b,d). A set PPP is symmetric (Definition 1.2) if there is u0∈Pu^0\in Pu0∈P with u0+z∈P  ⟺  u0−z∈Pu^0+z\in P\iff u^0-z\in Pu0+z∈P⟺u0−z∈P for every zzz.

The Lean development names these zStochBD and zRobBD (the general problems (1.4) and (1.5) with data AAA, BBB, bbb, ccc, ddd and integer coordinate sets), and IsSymmetricAbout (Definition 1.2).

Formalization targets

Goal: Theorem 3.1 (p. 22)

zRob(b,d) ≥ (n+1)⋅zStoch(b,d).z_{\mathrm{Rob}}(b,d)\ \ge\ (n+1)\cdot z_{\mathrm{Stoch}}(b,d).zRob​(b,d) ≥ (n+1)⋅zStoch​(b,d).

The constant n+1n+1n+1 is the paper's. Both sides are pinned down separately by the milestones, so the goal cannot be satisfied by a degenerate value of either side.

Milestones

  1. Symmetry of the instance (proof of Theorem 3.1, pp. 22–23): the uncertainty set is symmetric about (1,(12,…,12))(1,(\tfrac12,\dots,\tfrac12))(1,(21​,…,21​)) and Eμ[d(ω)]=(12,…,12)\mathbb E_\mu[d(\omega)]=(\tfrac12,\dots,\tfrac12)Eμ​[d(ω)]=(21​,…,21​), so the hypotheses of the symmetric theory hold.
  2. Robust side (p. 23): zRob(b,d)≥1z_{\mathrm{Rob}}(b,d)\ge1zRob​(b,d)≥1.
  3. Eq. (3.1) (p. 23): zStoch(b,d)≤Eμ[min⁡(d1(ω),…,dn(ω))]z_{\mathrm{Stoch}}(b,d)\le\mathbb E_\mu[\min(d_1(\omega),\dots,d_n(\omega))]zStoch​(b,d)≤Eμ​[min(d1​(ω),…,dn​(ω))].
  4. Eq. (3.2) (p. 23): Eμ[min⁡(d1(ω),…,dn(ω))]=1n+1\mathbb E_\mu[\min(d_1(\omega),\dots,d_n(\omega))]=\dfrac1{n+1}Eμ​[min(d1​(ω),…,dn​(ω))]=n+11​.

Significance

Theorem 3.1 marks the boundary of the paper's positive results. The bound zRob≤2 zStochz_{\mathrm{Rob}}\le2\,z_{\mathrm{Stoch}}zRob​≤2zStoch​ of Theorem 2.1 needs only symmetry of the uncertainty set and a centred distribution. The instance here has both, has no integer variables and a single constraint, and still has a gap that grows linearly in the dimension. So symmetry alone does not make a robust solution a good approximation of the stochastic optimum when costs are uncertain. The paper's abstract states this as one of its main conclusions, and it is why Sections 4 and 5 compare the robust problem with the adaptive problem, whose worst-case objective does not suffer from it.

The result is proved in the paper; to our knowledge it has no machine-checked proof. What this mission adds is a formal proof of the instance against the general definitions of the two-stage problems, including the probabilistic computation (3.2), the expected minimum of independent uniform random variables, which Mathlib does not contain. That computation is reusable wherever order statistics of uniforms appear, for example in auctions, secretary problems and random-assignment bounds.

Difficulty

The robust half is short. The stochastic half has two parts. Eq. (3.1) needs a measurable policy that selects a cheapest coordinate; the policy printed in the paper puts a unit on every minimizing coordinate, which at ties overpays, so a formal proof must choose a single minimizer measurably and use integrability on the probability space. Eq. (3.2) is the main work: it is a genuine integral over the nnn-dimensional cube against a product measure, for every nnn, and Mathlib has no lemma on the distribution or the expectation of the minimum of independent random variables.

Formalization scope

  • Vectors are Fin k → ℝ with the componentwise order; BBB is the all-ones Matrix (Fin 1) (Fin n) ℝ, AAA and ccc live on Fin 0, b(ω)=1b(\omega)=1b(ω)=1, and the integer coordinate sets are empty (p1=p2=0p_1=p_2=0p1​=p2​=0).
  • Optimal values are infima in EReal, +∞+\infty+∞ when infeasible; the robust worst-case cost is an EReal supremum. No optimal solution is assumed to exist: the paper's "consider an optimal solution" is a proof device, and the statements are about the infima.
  • Stochastic policies must be integrable, together with their cost ω↦d(ω)Ty(ω)\omega\mapsto d(\omega)^{\mathsf T}y(\omega)ω↦d(ω)Ty(ω), and satisfy the constraint in every scenario, as on the page, not only almost surely.
  • The scenario space is an arbitrary probability space (Ω,μ)(\Omega,\mu)(Ω,μ) with a measurable ddd whose range is exactly [0,1]n[0,1]^n[0,1]n and whose law is the product of nnn uniform distributions on [0,1][0,1][0,1]. This is the paper's "each djd_jdj​ is distributed uniformly at random between 0 and 1 and independent of other coefficients" together with its description of I(b,d)(Ω)I_{(b,d)}(\Omega)I(b,d)​(Ω).
  • n≥1n\ge1n≥1 is assumed; the paper leaves it implicit.
  • The paper prints Z+n2\mathbb Z^{n_2}_+Z+n2​​ for the second-stage integer block in (1.4)–(1.5); Z+p2\mathbb Z^{p_2}_+Z+p2​​ is meant, and here p2=0p_2=0p2​=0. Eq. (3.2) is called an inequality on the page; it is formalized as the equality it is.

A real-valued infimum would assign the value 000 to an infeasible or ill-posed problem and make the goal trivially true; the EReal infima, and the separate milestones fixing zRob≥1z_{\mathrm{Rob}}\ge1zRob​≥1 and zStoch≤E[min⁡jdj]=1/(n+1)z_{\mathrm{Stoch}}\le\mathbb E[\min_j d_j]=1/(n+1)zStoch​≤E[minj​dj​]=1/(n+1), rule that out.

Contributions are welcome on each milestone separately. The expected-minimum computation (3.2) is self-contained and of independent use; a general lemma on the law of the minimum of independent random variables would serve it and other missions.

Selected references

  • D. Bertsimas, V. Goyal, On the Power of Robust Solutions in Two-Stage Stochastic and Adaptive Optimization Problems, Mathematics of Operations Research 35(2), 2010. https://doi.org/10.1287/moor.1090.0440 (cited here from the authors' manuscript, MIT DSpace).
  • J. R. Birge, F. Louveaux, Introduction to Stochastic Programming, 2nd ed., Springer, 2011. https://doi.org/10.1007/978-1-4614-0237-4
  • D. Bertsimas, M. Sim, The Price of Robustness, Operations Research 52(1), 2004. https://doi.org/10.1287/opre.1030.0065
8 thms1 active userReviewed
Analysis·Captain: mikedeng1

Regret in Decision Making under Uncertainty 1: Under Assumptions 2 and 3, Utility over Final and Foregone Assets Takes the Regret Form u(x, y) = αv(x) + f(v(x) − v(y))Research Paper

Why regret enters the utility function

Expected utility theory predicts that a decision maker's choice between two lotteries depends only on the distribution of final wealth under each. Observed behaviour contradicts this in systematic ways: the Allais paradox, the coexistence of insurance and gambling, and the reflection effect documented by Kahneman and Tversky (1979). David Bell's 1982 paper (Operations Research 30(5), 961–981) explains these patterns by regret: the dissatisfaction of learning that the alternative not chosen would have done better. In the same year Loomes and Sugden (1982) proposed a closely related theory; together the two papers are the origin of regret theory in decision analysis.

Bell's contribution is axiomatic. Instead of postulating a regret term, he derives the shape of a utility function that incorporates regret from three behavioural assumptions about how preferences respond to shifting all outcomes by an equal increment of value. This mission formalizes that derivation, Section 1 of the paper (pp. 965–970), culminating in Theorem 1.

Setting

An endpoint is described by two attributes: the final assets xxx and the foregone assets yyy, the asset level the decision maker would have had under the alternative not chosen. Utility is a function u(x,y)u(x,y)u(x,y), strictly increasing in xxx and strictly decreasing in yyy.

A simple comparison is a choice between two alternatives whose uncertainties are resolved by one predetermined joint distribution. With nnn states, a probability vector p=(p1,…,pn)p=(p_1,\dots,p_n)p=(p1​,…,pn​) and alternatives A,B∈RnA,B\in\mathbb R^nA,B∈Rn (AiA_iAi​ the final assets of AAA in state iii), the expected utility of selecting AAA while BBB is foregone is

EUp(A ∣ B)=∑i=1npi u(Ai,Bi),\mathrm{EU}_p(A\,|\,B)=\sum_{i=1}^n p_i\,u(A_i,B_i),EUp​(A∣B)=i=1∑n​pi​u(Ai​,Bi​),

and AAA is preferred to BBB when EUp(B ∣ A)<EUp(A ∣ B)\mathrm{EU}_p(B\,|\,A)<\mathrm{EU}_p(A\,|\,B)EUp​(B∣A)<EUp​(A∣B) (inequality (1) of the paper).

A value function v:R→Rv:\mathbb R\to\mathbb Rv:R→R, strictly increasing and onto R\mathbb RR, measures incremental value: moving from aaa to bbb is worth v(b)−v(a)v(b)-v(a)v(b)−v(a). An alternative A′A'A′ is obtained from AAA by a shift of incremental value δ\deltaδ if v(Ai′)=v(Ai)+δv(A'_i)=v(A_i)+\deltav(Ai′​)=v(Ai​)+δ for every state.

  • Assumption 1. Shifting every outcome of both alternatives by the same incremental value leaves the preferred alternative unchanged.
  • Assumption 2. If v(x2)−v(x1)=v(x3)−v(x2)v(x_2)-v(x_1)=v(x_3)-v(x_2)v(x2​)−v(x1​)=v(x3​)−v(x2​), the decision maker is indifferent between x2x_2x2​ for sure and a 50-50 lottery between x1x_1x1​ and x3x_3x3​.
  • Assumption 3. If selecting L1L_1L1​ over L2L_2L2​ is preferred to selecting L3L_3L3​ over L4L_4L4​, that is EUp(L3 ∣ L4)<EUp(L1 ∣ L2)\mathrm{EU}_p(L_3\,|\,L_4)<\mathrm{EU}_p(L_1\,|\,L_2)EUp​(L3​∣L4​)<EUp​(L1​∣L2​), this persists after all outcomes of all four alternatives are shifted by the same incremental value.

Formalization targets

Goal: the regret form (Theorem 1, p. 969, corrected)

Under Assumptions 2 and 3, there are a constant α\alphaα and a function fff with

u(x,y)=α v(x)+f(v(x)−v(y))for all x,y.u(x,y)=\alpha\,v(x)+f\big(v(x)-v(y)\big)\quad\text{for all }x,y.u(x,y)=αv(x)+f(v(x)−v(y))for all x,y.

Utility separates additively into a term in the value of final assets and a function of the regret v(x)−v(y)v(x)-v(y)v(x)−v(y). The coefficient α\alphaα is left free (see Formalization scope).

Milestones, in the order of the paper's argument

  1. Proof of Lemma 1 (p. 967): for v(x)=xv(x)=xv(x)=x, w(x+h,y+h)=k(h) w(x,y)w(x+h,y+h)=k(h)\,w(x,y)w(x+h,y+h)=k(h)w(x,y) with k(h)>0k(h)>0k(h)>0, where w(x,y)=u(x,y)−u(y,x)w(x,y)=u(x,y)-u(y,x)w(x,y)=u(x,y)−u(y,x).
  2. Lemma 1, linear vvv (p. 967): u(x,y)−u(y,x)=[u(0,y−x)−u(y−x,0)] e−cxu(x,y)-u(y,x)=[u(0,y-x)-u(y-x,0)]\,e^{-cx}u(x,y)−u(y,x)=[u(0,y−x)−u(y−x,0)]e−cx.
  3. Lemma 1 (p. 967): Assumption 1 implies u(x,y)−u(y,x)=g(v(x)−v(y)) e−c v(x)u(x,y)-u(y,x)=g(v(x)-v(y))\,e^{-c\,v(x)}u(x,y)−u(y,x)=g(v(x)−v(y))e−cv(x).
  4. Lemma 2, linear vvv (p. 968): u(x,y)−u(y,x)=u(0,y−x)−u(y−x,0)u(x,y)-u(y,x)=u(0,y-x)-u(y-x,0)u(x,y)−u(y,x)=u(0,y−x)−u(y−x,0).
  5. Lemma 2 (p. 968): Assumptions 1 and 2 imply u(x,y)−u(y,x)=g(v(x)−v(y))u(x,y)-u(y,x)=g(v(x)-v(y))u(x,y)−u(y,x)=g(v(x)−v(y)).
  6. Assumption 1 is a special case of Assumption 3 (p. 969).
  7. Proof of Theorem 1 (p. 970): for v(x)=xv(x)=xv(x)=x, u(x+h,y+h)−u(x,y)=j(h)u(x+h,y+h)-u(x,y)=j(h)u(x+h,y+h)−u(x,y)=j(h).
  8. Theorem 1 for v(x)=xv(x)=xv(x)=x (p. 970): u(x,y)=αx+u(0,y−x)u(x,y)=\alpha x+u(0,y-x)u(x,y)=αx+u(0,y−x).

Significance

The additive form is the basis of the rest of Bell's paper: Section 2 uses it to show that a single decision maker with fff decreasingly concave both buys fair insurance and takes long-odds bets, and to account for the Allais paradox and the reflection effect. The same additive structure, a value term plus a function of the regret difference, is the form in which regret theory has been used since, alongside the parallel proposal of Loomes and Sugden (1982). Lemmas 1 and 2 have independent content: simple comparisons identify only the antisymmetric part u(x,y)−u(y,x)u(x,y)-u(y,x)u(x,y)−u(y,x), and the lemmas pin that part down to a function of the regret difference.

The results are published with proofs in prose; none has a machine-checked proof. Formalizing them checks the argument as printed, which turns out to need a correction (the coefficient of v(x)v(x)v(x) in (4)), and fills in the regularity steps the paper leaves implicit: the solutions of the Cauchy equations that appear, and the passage from v(x)=xv(x)=xv(x)=x to a general value function.

Difficulty

The obvious argument reads each assumption as a functional equation and solves it. Two steps resist this. First, the functional equations k(x+h)=k(x)k(h)k(x+h)=k(x)k(h)k(x+h)=k(x)k(h) and j(x+h)=j(x)+j(h)j(x+h)=j(x)+j(h)j(x+h)=j(x)+j(h) have non-measurable solutions; excluding them needs boundedness of kkk and jjj on intervals, which must be extracted from the monotonicity of uuu and is never stated on the page. Second, the proof of Theorem 1 compares two 50-50 lotteries that are indifferent and requires g(a1−b1)g(a_1-b_1)g(a1​−b1​) to take an arbitrary value u(c1,c2)−u(a2,b2)u(c_1,c_2)-u(a_2,b_2)u(c1​,c2​)−u(a2​,b2​); when ggg is bounded no such indifference exists, so the printed step does not cover every case. Passing from v(x)=xv(x)=xv(x)=x to a general vvv requires transporting all three assumptions along v−1v^{-1}v−1.

Formalization scope

  • Assets, probabilities and utilities are real numbers; uuu is u : ℝ → ℝ → ℝ with u x y =u(x,y)=u(x,y)=u(x,y). Every theorem assumes uuu strictly increasing in xxx and strictly decreasing in yyy (the paper's "increasing"/"decreasing", p. 965, read strictly).
  • A simple comparison is a probability vector on Fin n and two alternatives on the same states. Assumption 3 quantifies over four alternatives on one common state space, the setting of the paper's own use of it.
  • The value function is strictly increasing (p. 968) and, as an added reading, maps R\mathbb RR onto R\mathbb RR: Assumptions 1 and 3 shift outcomes by arbitrary equal incremental values, which presupposes that every shifted value is attained. "vvv linear" (Lemmas 1, 2) is read as v(x)=ax+bv(x)=ax+bv(x)=ax+b with a>0a>0a>0.
  • Indifference in Assumption 2 is "neither alternative is preferred", equivalent to the first display on p. 969.
  • Corrected slip. The paper's display (4) and the end of its proof have coefficient 111 on v(x)v(x)v(x) (resp. xxx). This is false as printed: v(x)=xv(x)=xv(x)=x, u(x,y)=x−yu(x,y)=x-yu(x,y)=x−y satisfies every hypothesis, but x−y=x+f(x−y)x-y=x+f(x-y)x−y=x+f(x−y) forces f(0)=0f(0)=0f(0)=0 at x=y=0x=y=0x=y=0 and f(0)=−1f(0)=-1f(0)=−1 at x=y=1x=y=1x=y=1. The proof derives j(h)=αhj(h)=\alpha hj(h)=αh and silently sets α=1\alpha=1α=1. The goal and milestone 8 keep α\alphaα free, with no sign or normalisation imposed; milestone texts stay verbatim.
  • All existential objects (kkk, ccc, ggg, jjj, α\alphaα, fff) are chosen before the variables x,y,hx,y,hx,y,h. The assumptions are statements about preferences between lotteries, never functional equations on uuu; a formalization that states an assumption as "u(x+h,y+h)−u(x,y)u(x+h,y+h)-u(x,y)u(x+h,y+h)−u(x,y) is independent of (x,y)(x,y)(x,y)" would put the conclusion into the hypothesis and is ruled out.
  • Needed infrastructure: monotone or locally bounded solutions of the Cauchy additive and multiplicative equations on R\mathbb RR, and transport of the assumptions along an order isomorphism of R\mathbb RR. Both are reusable; contributions proving either are welcome.

Selected references

  • D. E. Bell, Regret in Decision Making under Uncertainty, Operations Research 30(5), 961–981, 1982. https://doi.org/10.1287/opre.30.5.961
  • G. Loomes and R. Sugden, Regret Theory: An Alternative Theory of Rational Choice under Uncertainty, The Economic Journal 92(368), 805–824, 1982. https://doi.org/10.2307/2232669
  • D. Kahneman and A. Tversky, Prospect Theory: An Analysis of Decision under Risk, Econometrica 47(2), 263–291, 1979. https://doi.org/10.2307/1914185
10 thms1 active userReviewed
CombinatoricsProbabilityTheoretical Computer Science·Captain: mikedeng1

Online Stochastic Matching: Beating 1-1/e 1: When OPT = Ω(n), the Two Suggested Matchings Algorithm Achieves ALG/OPT ≥ (1 − 2/e²)/(4/3 − 2/(3e)) − ε ≈ 0.670 with Probability 1 − e^(−Ω(n))Research Paper

Motivation

Online bipartite matching models a platform that must commit each arriving request to a resource immediately. The motivating application of Feldman, Mehta, Mirrokni and Muthukrishnan is display advertising: an ad server knows from past traffic how many impressions of each type (web page, audience segment) to expect, sells them to advertisers in advance, and must assign each impression to an interested advertiser the moment a user loads the page. The goal is to fill as many contracted impressions as possible.

When arrivals are chosen by an adversary, the best ratio an online algorithm can guarantee is 1−1/e≈0.6321 - 1/e \approx 0.6321−1/e≈0.632, achieved by the RANKING algorithm of Karp, Vazirani and Vazirani (STOC 1990). The ad server, however, is not facing an adversary: it has a forecast. The i.i.d. model captures this: the graph and the distribution of impression types are known in advance, and the impressions are independent draws. The paper asks whether this knowledge allows an online algorithm to beat 1−1/e1 - 1/e1−1/e, and answers yes.

Timeline.

  • 1990: Karp, Vazirani and Vazirani give RANKING, with ratio 1−1/e1 - 1/e1−1/e for adversarial arrivals, and show this is optimal in that model.
  • 2005: Mehta, Saberi, Vazirani and Vazirani obtain 1−1/e1 - 1/e1−1/e for the budgeted AdWords generalization.
  • 2009: Feldman, Mehta, Mirrokni and Muthukrishnan (arXiv:0905.4100, FOCS 2009) show that in the i.i.d. model the two suggested matchings algorithm achieves about 0.6700.6700.670 with high probability when OPT is linear in nnn, the first ratio above 1−1/e1 - 1/e1−1/e for this model, and that no online algorithm reaches 26/2726/2726/27 in expectation.

Setting

An instance is a bipartite graph G=(A,I,E)G = (A, I, E)G=(A,I,E) with a finite set AAA of advertisers, a finite set III of impression types, and edges E⊆A×IE \subseteq A \times IE⊆A×I recording which advertisers want which types. The mission treats the case analysed throughout §4.2 of the paper, in which one impression of each type is expected (ei=1e_i = 1ei​=1). So n=∣I∣n = |I|n=∣I∣ impressions arrive one at a time, with types ω(0),…,ω(n−1)\omega(0), \dots, \omega(n-1)ω(0),…,ω(n−1) drawn independently and uniformly from III. On arrival an impression must be assigned at once and irrevocably to a still unassigned advertiser adjacent to its type, or discarded. ALG(ω)\mathrm{ALG}(\omega)ALG(ω) is the number of impressions an algorithm assigns. OPT(ω)\mathrm{OPT}(\omega)OPT(ω) is the size of a maximum matching of the realization graph, which has one node per arrival ttt, joined to every advertiser aaa with (a,ω(t))∈E(a, \omega(t)) \in E(a,ω(t))∈E.

The two suggested matchings (TSM) algorithm works offline first. Its boosted flow graph GfG_fGf​ has a source arc of capacity 222 into every advertiser, a unit-capacity arc along every edge of EEE, and an arc of capacity 222 from every type to a sink. The algorithm takes the edge set EfE_fEf​ of an integral maximum flow. Every vertex then has at most two edges of EfE_fEf​, so EfE_fEf​ splits into vertex-disjoint paths and cycles. The algorithm colours each component blue and red:

  • on cycles, the colours alternate;
  • on odd paths, the colours alternate, with more blue than red;
  • on even paths between advertisers, the colours alternate;
  • on even paths between types, the first two edges are blue, then the colours alternate, ending in blue.

Online, the first arrival of type iii tries the advertiser along iii's blue edge, the second tries the one along its red edge, and later arrivals are discarded. A tried advertiser that is already taken is not reassigned. The advertisers fall into four classes by their coloured edges: ABRA_{BR}ABR​ (one blue, one red), ABBA_{BB}ABB​ (two blue), ABA_BAB​ (one blue only) and ARA_RAR​ (one red only).

Formalization targets

Goal: Theorem 5, first sentence, ei=1e_i = 1ei​=1

Let

α=1−2/e24/3−2/(3e)≈0.67029.\alpha = \frac{1 - 2/e^2}{4/3 - 2/(3e)} \approx 0.67029 .α=4/3−2/(3e)1−2/e2​≈0.67029.

The goal has three parts. First, every maximum flow edge set admits a colouring that follows the rules. Second, for every ε>0\varepsilon > 0ε>0 and c>0c > 0c>0 there are δ>0\delta > 0δ>0 and NNN such that, for every instance with n≥Nn \ge Nn≥N, every maximum flow edge set and every rule-following colouring,

Pr⁡ω[ OPT≥c n  ⟹  ALG≥(α−ε) OPT ]  ≥  1−e−δn.\Pr_\omega\big[\ \mathrm{OPT} \ge c\,n \implies \mathrm{ALG} \ge (\alpha - \varepsilon)\,\mathrm{OPT}\ \big] \;\ge\; 1 - e^{-\delta n}.ωPr​[ OPT≥cn⟹ALG≥(α−ε)OPT ]≥1−e−δn.

Third, α>1−1/e\alpha > 1 - 1/eα>1−1/e.

Milestones

  1. Facts 1 and 2: concentration for two balls-in-bins statistics.
  2. The note of §4.2.1: each type has no coloured edge, one blue edge, or one blue and one red edge.
  3. Equation (1): ∣Ef∣=2∣ABR∣+2∣ABB∣+∣AB∣+∣AR∣|E_f| = 2|A_{BR}| + 2|A_{BB}| + |A_B| + |A_R|∣Ef​∣=2∣ABR​∣+2∣ABB​∣+∣AB​∣+∣AR​∣.
  4. Equation (2): with high probability, ALG≥(1−1/e2)∣ABB∣+(1−2/e2)∣ABR∣+(1−3/(2e))(∣AB∣+∣AR∣)−4εn\mathrm{ALG} \ge (1 - 1/e^2)|A_{BB}| + (1 - 2/e^2)|A_{BR}| + (1 - 3/(2e))(|A_B| + |A_R|) - 4\varepsilon nALG≥(1−1/e2)∣ABB​∣+(1−2/e2)∣ABR​∣+(1−3/(2e))(∣AB​∣+∣AR​∣)−4εn.
  5. Equation (3): ∣Ef∣=2(∣AT∣+∣IS∣)+∣Eδ∣|E_f| = 2(|A_T| + |I_S|) + |E_\delta|∣Ef​∣=2(∣AT​∣+∣IS​∣)+∣Eδ​∣ for the surgered residual cut (S,T)(S,T)(S,T) of GfG_fGf​.
  6. Equation (4): with high probability, OPT≤∣ABR∣+∣ABB∣+12(∣AB∣+∣AR∣)+(12−1e)∣Eδ∣+εn\mathrm{OPT} \le |A_{BR}| + |A_{BB}| + \tfrac12(|A_B| + |A_R|) + (\tfrac12 - \tfrac1e)|E_\delta| + \varepsilon nOPT≤∣ABR​∣+∣ABB​∣+21​(∣AB​∣+∣AR​∣)+(21​−e1​)∣Eδ​∣+εn.
  7. Lemma 1: ∣Eδ∣≤23∣ABR∣+43∣ABB∣+∣AB∣+13∣AR∣|E_\delta| \le \tfrac23|A_{BR}| + \tfrac43|A_{BB}| + |A_B| + \tfrac13|A_R|∣Eδ​∣≤32​∣ABR​∣+34​∣ABB​∣+∣AB​∣+31​∣AR​∣.

Significance

The theorem separates the i.i.d. model from the adversarial one: knowing the distribution is worth a constant factor above 1−1/e1 - 1/e1−1/e. The suggested matching algorithm of the same paper (Theorem 4) shows that following a single offline matching gets exactly 1−1/e1 - 1/e1−1/e, so the second, red matching is what crosses the barrier. The paper's question started a line of work on the i.i.d. and random-order models, with later improvements to the constant by other authors under further assumptions.

The result has a written proof but, as far as the platform record shows, no machine-checked one. The mission formalizes the paper's own argument: the flow-and-colouring construction, the balls-in-bins concentration facts, the cut-based bound on OPT and the combinatorial Lemma 1. It also fixes two slips in the printed statements (see Formalization scope). The pieces are reusable beyond this paper. The occupancy concentration (Fact 1) and the satisfied-sequences bound (Fact 2) recur in analyses of online algorithms with stochastic input. The degree-capped flow encoding and its path/cycle decomposition are standard tools for 2-matchings.

Difficulty

The upper bound on OPT is the delicate part. A cut of the flow graph bounds the maximum matching of the realization graph only after a second surgery that depends on the random arrivals. Its size must then be compared with the colour classes, which are defined by a different structure (the components of EfE_fEf​). Lemma 1 bridges the two, and it depends on the exact colouring rules: a colouring that only satisfies local degree conditions can put red edges at both ends of an even advertiser path, which breaks the inequality ∣AB∣≥∣AR∣|A_B| \ge |A_R|∣AB​∣≥∣AR​∣ behind (2). On the probabilistic side, the advertisers of ABRA_{BR}ABR​ share impression types with each other, so the success events are dependent, and Fact 2 needs a bounded-differences argument in which one ball affects up to ddd sequences.

Formalization scope

All declarations live in the namespace OnlineStochMatching.TSM. Advertisers and types are finite types A I : Type, and EEE is a Finset (A × I). Probabilities are counting ratios #{ω:Fin n→I∣P ω}/∣I∣n\#\{\omega : \mathrm{Fin}\ n \to I \mid P\,\omega\}/|I|^n#{ω:Fin n→I∣Pω}/∣I∣n, so there are no measurability side conditions. OPT is a maximum over the finite, nonempty set of partial injective assignments. An integral flow of GfG_fGf​ is its set of saturated middle edges, i.e. a subset of EEE with at most two edges per vertex; EfE_fEf​ is such a set of maximum cardinality. A colouring is given by a listing of the components of EfE_fEf​ as vertex sequences. Every theorem quantifies over every maximum EfE_fEf​ and every colouring the rules allow, since the paper fixes neither.

The paper's asymptotic phrases are replaced by explicit quantifiers that come from its own proofs:

  • "with probability 1−e−Ω(n)1 - e^{-\Omega(n)}1−e−Ω(n)" and "with high probability" (Theorem 5, (2), (4)) become: ∃ δ>0, ∃ N\exists\, \delta > 0,\ \exists\, N∃δ>0, ∃N, chosen before the instance, with probability at least 1−e−δn1 - e^{-\delta n}1−e−δn for all n≥Nn \ge Nn≥N;
  • "as long as OPT =Ω(n)= \Omega(n)=Ω(n)" becomes the event OPT≥c n\mathrm{OPT} \ge c\,nOPT≥cn for an arbitrary c>0c > 0c>0 fixed before δ\deltaδ and NNN;
  • the O(1)O(1)O(1) term in the bound on ∣Aδ∗∣|A^*_\delta|∣Aδ∗​∣ (p. 8) is absorbed into εn\varepsilon nεn for n≥Nn \ge Nn≥N.

Corrections to the printed statements:

  • Theorem 5 prints ALG/OPT−ϵ≥α\mathrm{ALG}/\mathrm{OPT} - \epsilon \ge \alphaALG/OPT−ϵ≥α; the proof concludes ALG/OPT+ϵ≥α\mathrm{ALG}/\mathrm{OPT} + \epsilon \ge \alphaALG/OPT+ϵ≥α, so the goal states ALG≥(α−ε)OPT\mathrm{ALG} \ge (\alpha - \varepsilon)\mathrm{OPT}ALG≥(α−ε)OPT;
  • Fact 1 prints the failure probability 2e−ϵn/22e^{-\epsilon n/2}2e−ϵn/2; its proof gives 2e−ϵ2n/22e^{-\epsilon^2 n/2}2e−ϵ2n/2, which is used;
  • Fact 2 states two hypotheses its proof uses: the bins of a sequence are distinct, and c2<nc^2 < nc2<n.

The ratio is multiplied out, so no division by OPT occurs. A colouring condition that is unsatisfiable, or a flow set that is not maximum, would make the goal vacuous or false; part (a) of the goal rules out the first, and every statement requires maximality. Not included: the reduction to general integer eie_iei​ (§4.2.4), the tightness sentence of Theorem 5 (§4.2.5), and footnote 7's variant of the algorithm.

Useful infrastructure: bounded-differences (McDiarmid/Azuma) inequalities for functions of i.i.d. uniform variables, which exist on the platform as separate theorems; path/cycle decomposition of graphs of maximum degree two; and max-flow min-cut for unit-capacity bipartite networks. Contributions are welcome on Facts 1 and 2 independently of the combinatorics, and on Lemma 1 and equations (1) and (3), which are deterministic.

Selected references

  • J. Feldman, A. Mehta, V. Mirrokni, S. Muthukrishnan, Online Stochastic Matching: Beating 1-1/e, FOCS 2009; arXiv:0905.4100v1. https://arxiv.org/abs/0905.4100
  • R. M. Karp, U. V. Vazirani, V. V. Vazirani, An optimal algorithm for on-line bipartite matching, STOC 1990. https://doi.org/10.1145/100216.100262
  • A. Mehta, A. Saberi, U. Vazirani, V. Vazirani, AdWords and generalized online matching, FOCS 2005; J. ACM 54(5), 2007. https://doi.org/10.1145/1284320.1284321
13 thms1 active userReviewed
Optimization·Captain: mikedeng1

Assortment Optimization under Variants of the Nested Logit Model 5: For General Nests, the Nested-by-Preference-and-Revenue LP Optimum Scaled by the Factor (12) Is Feasible for the Full LPResearch Paper

Assortment planning with nested choice

A retailer that groups its products into categories (brands, store sections, flight classes) and decides which products to display in each faces the assortment problem: offering more products attracts more customers but also diverts sales away from the most profitable products. The nested logit model is the standard description of customer choice in this setting. A customer first selects a category (a nest), then a product within it. Choice-based models of this kind are the basis of revenue management under customer choice (Talluri and van Ryzin 2004).

Davis, Gallego and Topaloglu (DGT 2014) study the assortment problem under the nested logit model with two features that earlier work excluded: dissimilarity parameters larger than one, under which products in a nest act as complements rather than substitutes, and a no-purchase option inside each nest, under which a customer may enter a nest and still leave without buying. They show that the problem is NP-hard once either feature is present. For each regime they give a small linear program whose solution yields an assortment with a provable performance guarantee. This mission formalizes the guarantee for the most general instances, where both features occur together (§6.1, Theorem 11).

Timeline. Rusmevichientong, Shmoys and Topaloglu (2010) bound nested-by-revenue assortments under a multinomial logit mixture. [DGT 2014] prove that nested-by-revenue assortments are optimal for dissimilarity parameters at most one without within-nest no-purchase options (Theorem 4). They give factor-(6) guarantees with synergistic products, a factor-two guarantee via knapsack relaxations for partially-captured nests (Theorem 10), and the general factor (12) of Theorem 11. Li, Rusmevichientong and Topaloglu (2015) extend the nested-by-revenue result to ddd-level nested logit models.

The model

There are nests i∈M={1,…,m}i\in M=\{1,\dots,m\}i∈M={1,…,m} and, in each nest, products j∈N={1,…,n}j\in N=\{1,\dots,n\}j∈N={1,…,n}. Product jjj of nest iii has revenue rij≥0r_{ij}\ge0rij​≥0 and preference weight vij>0v_{ij}>0vij​>0, with ri1≥ri2≥⋯≥rinr_{i1}\ge r_{i2}\ge\dots\ge r_{in}ri1​≥ri2​≥⋯≥rin​. Nest iii has a no-purchase weight vi0≥0v_{i0}\ge0vi0​≥0 and a dissimilarity parameter γi>0\gamma_i>0γi​>0, and v0≥0v_0\ge0v0​≥0 is the weight of choosing no nest at all. For an assortment Si⊆NS_i\subseteq NSi​⊆N,

Vi(Si)=vi0+∑j∈Sivij,Ri(Si)=∑j∈SirijvijVi(Si).V_i(S_i)=v_{i0}+\sum_{j\in S_i}v_{ij},\qquad R_i(S_i)=\frac{\sum_{j\in S_i}r_{ij}v_{ij}}{V_i(S_i)} .Vi​(Si​)=vi0​+j∈Si​∑​vij​,Ri​(Si​)=Vi​(Si​)∑j∈Si​​rij​vij​​.

A customer chooses nest iii with probability Vi(Si)γi/(v0+∑lVl(Sl)γl)V_i(S_i)^{\gamma_i}/(v_0+\sum_l V_l(S_l)^{\gamma_l})Vi​(Si​)γi​/(v0​+∑l​Vl​(Sl​)γl​), and the expected revenue is

Π(S1,…,Sm)=∑iVi(Si)γiRi(Si)v0+∑iVi(Si)γi.\Pi(S_1,\dots,S_m)=\frac{\sum_{i}V_i(S_i)^{\gamma_i}R_i(S_i)}{v_0+\sum_{i}V_i(S_i)^{\gamma_i}} .Π(S1​,…,Sm​)=v0​+∑i​Vi​(Si​)γi​∑i​Vi​(Si​)γi​Ri​(Si​)​.

The optimal value Z∗Z^*Z∗ of max⁡Π\max\PimaxΠ equals the optimal value of the linear program

(3)min⁡ xs.t.v0x≥∑iyi,yi≥Vi(Si)γi(Ri(Si)−x)  ∀Si⊆N, i∈M,\text{(3)}\qquad \min\ x\quad\text{s.t.}\quad v_0x\ge\sum_i y_i,\qquad y_i\ge V_i(S_i)^{\gamma_i}\big(R_i(S_i)-x\big)\ \ \forall S_i\subseteq N,\ i\in M,(3)min xs.t.v0​x≥i∑​yi​,yi​≥Vi​(Si​)γi​(Ri​(Si​)−x)  ∀Si​⊆N, i∈M,

which has 2n2^n2n constraints per nest. Problem (4) keeps only the constraints for a chosen candidate collection of assortments in each nest.

A nest is fully captured if vi0=0v_{i0}=0vi0​=0 (i∈Mfi\in M^fi∈Mf) and partially captured if vi0>0v_{i0}>0vi0​>0 (i∈Mpi\in M^pi∈Mp). Nij={1,…,j}N_{ij}=\{1,\dots,j\}Nij​={1,…,j} is the nested-by-revenue assortment. NijkN^k_{ij}Nijk​ is the set of the jjj highest-revenue products among the kkk products of nest iii with the smallest preference weights, with Ni0k=∅N^k_{i0}=\emptysetNi0k​=∅ and Nijn=NijN^n_{ij}=N_{ij}Nijn​=Nij​.

Formalization targets

Goal: Theorem 11

Let (x^,y^)(\hat x,\hat y)(x^,y^​) be an optimal solution of (4) when the candidate collection of every nest is {Nijk:k∈N, j=0,…,k}∪{{j}:j∈N}\{N^k_{ij}:k\in N,\ j=0,\dots,k\}\cup\{\{j\}:j\in N\}{Nijk​:k∈N, j=0,…,k}∪{{j}:j∈N}, and let

β=max⁡i∈Mf, j=2,…,n{Vi(Nij)Vi(Ni,j−1)}∨max⁡i∈Mp, j=1,…,n{Vi(Nij)Vi(Ni,j−1)}∨2(12).\beta=\max_{i\in M^f,\ j=2,\dots,n}\left\{\frac{V_i(N_{ij})}{V_i(N_{i,j-1})}\right\}\vee\max_{i\in M^p,\ j=1,\dots,n}\left\{\frac{V_i(N_{ij})}{V_i(N_{i,j-1})}\right\}\vee2 \qquad (12).β=i∈Mf, j=2,…,nmax​{Vi​(Ni,j−1​)Vi​(Nij​)​}∨i∈Mp, j=1,…,nmax​{Vi​(Ni,j−1​)Vi​(Nij​)​}∨2(12).

Then (βx^,βy^)(\beta\hat x,\beta\hat y)(βx^,βy^​) is feasible for problem (3).

The theorem assumes γˉ=max⁡iγi>1\bar\gamma=\max_i\gamma_i>1γˉ​=maxi​γi​>1, as all of §6 does. Otherwise it places no restriction on the γi\gamma_iγi​ or the vi0v_{i0}vi0​.

Milestones

  1. x^≥0\hat x\ge0x^≥0 (A.4, p. 46).
  2. Every greedy knapsack assortment S^i(ϵi)\hat S_i(\epsilon_i)S^i​(ϵi​) of §5 is one of the NijkN^k_{ij}Nijk​ (pp. 24–25).
  3. The relaxed nest problem over [0,1]n[0,1]^n[0,1]n has an optimal solution of fractional-prefix form (A.4 Case 1, p. 47).
  4. Inequality (30): for a nest with γi>1\gamma_i>1γi​>1 and y^i≥0\hat y_i\ge0y^​i​≥0, βy^i\beta\hat y_iβy^​i​ bounds the relaxed objective at every fractional prefix (p. 47).
  5. Case 1: γi>1\gamma_i>1γi​>1, y^i≥0\hat y_i\ge0y^​i​≥0 gives the constraints of (3) for nest iii (pp. 47–48).
  6. Problem (31) has a nested-by-revenue optimal solution when γi>1\gamma_i>1γi​>1 and its coefficient b=βy^ib=\beta\hat y_ib=βy^​i​ is negative (p. 48).
  7. Case 2: γi>1\gamma_i>1γi​>1, y^i<0\hat y_i<0y^​i​<0 (p. 48).
  8. Case 3: γi≤1\gamma_i\le1γi​≤1, through the factor-two argument of Theorem 10 (pp. 48–49).

Two companions follow the goal. One is the resulting guarantee β Π(S^)≥Z∗≥Π(S^)\beta\,\Pi(\hat S)\ge Z^*\ge\Pi(\hat S)βΠ(S^)≥Z∗≥Π(S^), through Theorem 1. The other is the bound β≤2κ\beta\le2\kappaβ≤2κ when the preference weights within a nest differ by at most a factor κ\kappaκ (p. 26).

Significance

Theorem 11, combined with Theorem 1 of the paper, gives a polynomial-size method for an NP-hard problem. The method solves one linear program with 1+m1+m1+m variables and 1+m(1+n+n2)1+m(1+n+n^2)1+m(1+n+n2) constraints, then reads off an assortment whose expected revenue is within the factor β\betaβ of the optimum. This holds for every nested logit instance, including nests where customers may walk away and nests whose products are complements. When the weights inside each nest are within a factor κ\kappaκ of each other, the guarantee is at most 2κ2\kappa2κ.

The theorem is proved in the paper's appendix. No part of it is machine-checked. Formalizing it checks a case analysis that reuses, by reference, arguments from two other theorems: Theorem 7 (synergistic, fully-captured nests) and Theorem 10 (competitive, partially-captured nests). It makes precise what these arguments need when the two regimes are mixed in one instance. The formalization also fixes the boundary conventions the printed proof leaves implicit: fully-captured nests with k=1k=1k=1, zero-weight denominators, and the sign of y^i\hat y_iy^​i​.

Difficulty

Each nest falls into one of three regimes, and a different argument controls each. With γi≤1\gamma_i\le1γi​≤1 the nest behaves like a knapsack problem. Its guarantee of two needs the knapsack collection of §5 to sit inside {Nijk}\{N^k_{ij}\}{Nijk​}. With γi>1\gamma_i>1γi​>1 and y^i≥0\hat y_i\ge0y^​i​≥0, the constraint must be extended from nested-by-revenue sets to every subset. This goes through a continuous relaxation whose optimum has a fractional coordinate, and it costs the ratio Vi(Nik)/Vi(Ni,k−1)V_i(N_{ik})/V_i(N_{i,k-1})Vi​(Nik​)/Vi​(Ni,k−1​), which is where (12) comes from. With γi>1\gamma_i>1γi​>1 and y^i<0\hat y_i<0y^​i​<0, the scaling argument of Case 1 fails because multiplying by a factor at most one no longer preserves the inequality. The proof switches to the different objective (31), whose convexity in one coordinate forces an integral optimum.

The first idea, bounding every assortment by a nested-by-revenue one, is false here. With γi>1\gamma_i>1γi​>1 or vi0>0v_{i0}>0vi0​>0, nested-by-revenue assortments are not optimal, and the loss is exactly the factor β\betaβ.

Formalization scope

Products are Fin n; NijN_{ij}Nij​ is nbr n j. Powers are Real.rpow, and x/0=0x/0=0x/0=0, so Ri(∅)=0R_i(\emptyset)=0Ri​(∅)=0. Problems (3) and (4) are stated in constraint form: LP4Optimal means feasible and with xxx minimal among feasible points. β\betaβ is the greatest element of the finite set betaSet I, which contains 222 and the ratios of (12). Fully-captured nests skip j=1j=1j=1, as on the page. The collection is constructed: nestedPR breaks weight ties by index and revenue ties by index.

Standing assumptions, all disclosed:

  • vij>0v_{ij}>0vij​>0, rij≥0r_{ij}\ge0rij​≥0 and γi>0\gamma_i>0γi​>0. The page allows zero-weight padding products and γi=0\gamma_i=0γi​=0, but its arguments do not cover them.
  • γˉ>1\bar\gamma>1γˉ​>1 on every statement set in Theorem 11's context.
  • n≥1n\ge1n≥1 for the collection claim and the prefix claim.
  • vi0>0v_{i0}>0vi0​>0 for the statement about (31). That is the only kind of nest where Case 2 arises. For vi0=0v_{i0}=0vi0​=0, Lean's 01−γi=00^{1-\gamma_i}=001−γi​=0 would remove the page's +∞+\infty+∞.
  • v0>0v_0>0v0​>0 for the guarantee, where Theorem 1 fails otherwise.
  • κ≥1\kappa\ge1κ≥1, and vi0v_{i0}vi0​ counted among the weights of a partially-captured nest (vij≤κvi0v_{ij}\le\kappa v_{i0}vij​≤κvi0​, vi0≤κvijv_{i0}\le\kappa v_{ij}vi0​≤κvij​), for the 2κ2\kappa2κ bound.

A trivializing formalization is ruled out. The goal states only feasibility for (3), with β\betaβ the maximum of (12), not any upper bound. The collection is the page's, not an arbitrary family containing it. The goal mentions none of the cases or the relaxations.

A complete development needs continuous knapsack solutions (greedy optimality, fractional prefixes), convexity of t↦t1−γt\mapsto t^{1-\gamma}t↦t1−γ on (0,∞)(0,\infty)(0,∞), and the factor-two argument of Theorem 10. The knapsack and fractional-prefix lemmas are reusable for the companion missions of this series. Proofs of individual cases, and proofs of milestones in greater generality, are welcome.

Selected references

  • J. M. Davis, G. Gallego, H. Topaloglu, Assortment optimization under variants of the nested logit model, Operations Research 62(2), 2014 (revised manuscript of June 18, 2013). https://doi.org/10.1287/opre.2014.1256
  • P. Rusmevichientong, D. B. Shmoys, H. Topaloglu, Assortment optimization with mixtures of logits, technical report, Cornell University, 2010. http://legacy.orie.cornell.edu/~huseyin/publications/publications.html
  • G. Li, P. Rusmevichientong, H. Topaloglu, The d-level nested logit model: assortment and price optimization problems, Operations Research 63(2), 2015.
  • K. Talluri, G. van Ryzin, Revenue management under a general discrete choice model of consumer behavior, Management Science 50(1), 15–33, 2004. https://doi.org/10.1287/mnsc.1030.0147
  • D. P. Williamson, D. B. Shmoys, The Design of Approximation Algorithms, Cambridge University Press, 2011. https://doi.org/10.1017/CBO9780511921735
14 thms1 active userReviewed
Dynamic ProgrammingMarkov Chain·Captain: mikedeng1

Denumerable State Markovian Decision Processes—Average Cost Criterion: Every Limit Point of Policy Improvement Is a Deterministic Stationary Rule That Is Optimal over All RulesResearch Paper

Motivation

Markovian decision processes with the long-run average cost criterion model systems run indefinitely: inventories, queues, maintenance schedules, communication links. For a finite state space the theory was settled by the early 1960s. Howard's policy improvement (policy iteration) procedure (1960) finds an optimal stationary rule in finitely many steps, and Gillette (1957) and Derman (1962) showed that a stationary deterministic rule is optimal over all rules, history-dependent and randomized included (Derman 1962).

Many models of interest, such as queues with unbounded buffers and inventories with unbounded backlog, have a denumerable state space, and there the finite theory breaks down. Derman's paper (Ann. Math. Statist. 37 (1966) 1545–1553) gives two counterexamples in §2, both under bounded costs and finitely many decisions per state. In the first, due to Maitra, no optimal rule exists. In the second, a randomized stationary rule beats every deterministic one. It then gives sufficient conditions under which a stationary deterministic optimal rule exists and policy improvement finds it.

Timeline:

  • 1960: Howard introduces policy iteration for finite average-cost problems.
  • 1962: Derman proves that deterministic stationary rules are optimal for finite state spaces; Blackwell develops the finite discounted and near-discount theory.
  • 1963–1965: Iglehart (inventory) and Taylor (replacement) treat the average cost criterion in special infinite-state models; Derman's §3 proof follows part of Iglehart's argument.
  • 1964–1965: Blackwell, Maitra, Strauch and Derman (J. Math. Anal. Appl. 1965) treat infinite state spaces under the discounted criterion, where with Ki<∞K_i < \inftyKi​<∞ and bounded costs an optimal rule of C′′C''C′′ always exists.
  • 1966: Derman (this paper) gives the bounded-solution verification theorem and the convergence of policy improvement on a denumerable state space.
  • 1967 onward: Derman and Veinott (announced in §5 of this paper), and later Ross, Sennott, and Arapostathis, Borkar, Fernández-Gaucherand, Ghosh and Marcus (1993), give conditions for the existence of solutions of the optimality equation.

Setting

The system is observed at times t=0,1,2,…t = 0, 1, 2, \dotst=0,1,2,… in a state YtY_tYt​ of a denumerable set III. At state iii one of Ki<∞K_i < \inftyKi​<∞ decisions kkk is made (condition (A)). It costs wikw_{ik}wik​, and the next state is jjj with probability qij(k)q_{ij}(k)qij​(k). The costs are bounded (condition (B)) and may have either sign. A rule RRR chooses the decision at each time with probabilities that may depend on the whole history. The class of all rules is CCC, and C′′C''C′′ is the class of stationary deterministic rules ("make decision kik_iki​ at state iii"). The average cost of RRR from Y0=iY_0 = iY0​=i is

QR(i)=lim sup⁡T→∞1T+1∑t=0TERWt,Wt=wYtΔt.Q_R(i) = \limsup_{T\to\infty} \frac{1}{T+1}\sum_{t=0}^{T} E_R W_t, \qquad W_t = w_{Y_t \Delta_t}.QR​(i)=T→∞limsup​T+11​t=0∑T​ER​Wt​,Wt​=wYt​Δt​​.

A rule is optimal over CCC if QR(i)≤QR′(i)Q_R(i) \le Q_{R'}(i)QR​(i)≤QR′​(i) for every R′∈CR' \in CR′∈C and every iii.

The optimality equation (1) asks for a number ggg and a bounded {vj}\{v_j\}{vj​} with

g+vi=min⁡k{wik+∑j∈Iqij(k)vj},i∈I,g + v_i = \min_k \Big\{w_{ik} + \sum_{j\in I} q_{ij}(k) v_j\Big\}, \qquad i \in I,g+vi​=kmin​{wik​+j∈I∑​qij​(k)vj​},i∈I,

and equation (2) is its version for one rule R∈C′′R \in C''R∈C′′, with kik_iki​ in place of the minimum. Condition (C): every R∈C′′R \in C''R∈C′′ induces an irreducible Markov chain all of whose states are positive recurrent. Condition (E): every R∈C′′R \in C''R∈C′′ has a solution {gR,vjR}\{g^R, v^R_j\}{gR,vjR​} of (2), bounded uniformly in jjj and RRR. Condition (F): for every jjj the Cesàro limits πij(R)\pi_{ij}(R)πij​(R) of P{Yt=j∣Y0=i}P\{Y_t = j \mid Y_0 = i\}P{Yt​=j∣Y0​=i} satisfy inf⁡R∈C′′,i∈Iπij(R)>0\inf_{R\in C'', i \in I} \pi_{ij}(R) > 0infR∈C′′,i∈I​πij​(R)>0. One policy improvement iteration replaces RRR by a rule R′R'R′ whose decisions minimize wik+∑jqij(k)vjRw_{ik} + \sum_j q_{ij}(k) v_j^Rwik​+∑j​qij​(k)vjR​ at every state.

Formalization targets

Goal: Theorem 4

Under (A), (B), (C), (E) and (F), let R1,R2,…R_1, R_2, \dotsR1​,R2​,… be any sequence of policy improvement iterations from an arbitrary R1∈C′′R_1 \in C''R1​∈C′′. Then the sequence has a limit point R∗∈C′′R^* \in C''R∗∈C′′, and every limit point satisfies

QR∗(i)≤QR(i)for all R∈C, i∈I,Q_{R^*}(i) \le Q_R(i) \qquad \text{for all } R \in C,\ i \in I,QR∗​(i)≤QR​(i)for all R∈C, i∈I,

with gRn→QR∗(i)g^{R_n} \to Q_{R^*}(i)gRn​→QR∗​(i).

Milestones

  • Display (4): a rule of C′′C''C′′ solving (2) with bounded vvv has QR≡gQ_R \equiv gQR​≡g, the limit existing.
  • Display (6): a bounded solution of (1) gives ∣gn(i)−ng−vi∣≤M|g_n(i) - ng - v_i| \le M∣gn​(i)−ng−vi​∣≤M for the value iteration gng_ngn​ of (5).
  • §3: gn(i)g_n(i)gn​(i) is the least expected cost over the periods 0,…,n0, \dots, n0,…,n among all rules.
  • Theorem 1: a bounded solution of (1) makes every minimizing rule R∗∈C′′R^* \in C''R∗∈C′′ optimal over CCC, with QR∗≡gQ_{R^*} \equiv gQR∗​≡g.
  • Lemma 1 and the remark after it: a strict improvement at a nonempty set of states lowers QQQ at every initial state, by exactly ∑iπiεi\sum_i \pi_i \varepsilon_i∑i​πi​εi​.
  • Theorems 2 and 3: under (C), optimality over C′′C''C′′ implies optimality over CCC; under (C) and (E), an optimal rule of C′′C''C′′ exists.
  • Lemma 2: along policy improvement the gaps εiRn\varepsilon_i^{R_n}εiRn​​ tend to 000 at every state.

Significance

Theorem 1 is the countable-state verification theorem. It is the template for all later average-cost theory on infinite state spaces, which replaces boundedness of vvv by growth or Lyapunov conditions. Theorem 4 extends policy improvement, the standard algorithm for finite average-cost problems, to denumerable state spaces, under conditions that do not reduce the problem to a finite one.

Derman proved these results in 1966. None of them has been machine-checked: the platform has finite-state average-cost policy iteration and finite-state verification theorems, but no countable-state result of this kind. A formal development checks the analytic steps that the paper treats briefly. These are the interchange of limits and infinite sums in (3), (9) and (11), and the identification of the Cesàro limits of a positive recurrent chain with its steady-state probabilities.

Difficulty

On a finite state space, policy improvement terminates because there are finitely many rules and the average cost strictly decreases. On a denumerable state space C′′C''C′′ is uncountable, so neither step works. The sequence need not terminate. With ties in the minimization it need not converge either. A strict decrease of gRng^{R_n}gRn​ does not force the gaps to vanish. That needs (F), a uniform positive lower bound on the long-run occupation of each state. Passing to the limit in (2) along a subsequence requires a limit interchange in an infinite sum, which works only because {vR}\{v^R\}{vR} is uniformly bounded. Finally, comparing with history-dependent randomized rules needs the finite-horizon argument behind (6). A stationary-rule comparison is not enough.

Formalization scope

The model is the published Markov decision chain SennottDP.AvgFinite.MDC: a countable state type, a finite nonempty Finset of decisions at each state (so (A) is built in), and transition probabilities in [0,∞][0,\infty][0,∞] summing to one. History-dependent randomized rules are Policy M, and C′′C''C′′ is StationaryPolicy M. The chain notions (irreducibility, positive recurrence, steady state 1/mjj1/m_{jj}1/mjj​) come from the published SennottDP.MarkovCost.Chain. The conventions are as follows.

  • Costs are a separate signed function w with (B). The nonnegative cost field of MDC is unused, and w≥0w \ge 0w≥0 is not assumed.
  • ERWtE_R W_tER​Wt​ is the genuine expectation against the history law, and QRQ_RQR​ is the real limsup with the paper's normalization (T+1)−1(T+1)^{-1}(T+1)−1.
  • Every use of (1) or (2) assumes vvv bounded, so the series ∑jqij(k)vj\sum_j q_{ij}(k) v_j∑j​qij​(k)vj​ converge absolutely. The bound is part of the paper's hypotheses: without it Theorem 1 is false on infinite III.
  • (E) is an explicit family gR,vRg^R, v^RgR,vR, and the improvement step is taken against it, with arbitrary tie-breaking. (D) follows from (E) and is not a separate hypothesis.
  • (F) is stated with Cesàro limits. That they equal the steady-state probabilities under (C) is a proof obligation.
  • "Converges" in Theorem 4 is read as: a limit point exists, and every limit point is optimal over CCC. Whole-sequence convergence is not claimed, because the paper's proof does not establish it.
  • Sequences are indexed from 000.

Two trivializing formalizations are ruled out: optimality is over all history-dependent randomized rules, never only over C′′C''C′′, and QRQ_RQR​ is computed from the process law, never defined through (2).

A complete development needs Cesàro convergence of ttt-step probabilities to 1/mjj1/m_{jj}1/mjj​ for positive recurrent chains, dominated convergence for bounded functions against stochastic kernels, the finite-horizon dynamic programming principle for history-dependent rules, and a diagonal (Tychonoff) argument in ∏i{1,…,Ki}\prod_i \{1,\dots,K_i\}∏i​{1,…,Ki​}. The first and third are reusable well beyond this paper. Contributions of these lemmas as separate theorems are welcome.

Selected references

  • C. Derman, Denumerable State Markovian Decision Processes—Average Cost Criterion, Ann. Math. Statist. 37(6) (1966) 1545–1553. https://doi.org/10.1214/aoms/1177699146
  • C. Derman, On Sequential Decisions and Markov Chains, Management Sci. 9(1) (1962) 16–24. https://doi.org/10.1287/mnsc.9.1.16
  • R. A. Howard, Dynamic Programming and Markov Processes, Wiley, New York, 1960.
  • D. L. Iglehart, Dynamic programming and stationary analysis of inventory problems, Ch. 1 of Multistage Inventory Models and Techniques (H. Scarf, D. Gilford, M. Shelly, eds.), Stanford Univ. Press, 1963.
  • K. L. Chung, Markov Chains with Stationary Transition Probabilities, Springer, 1960. https://doi.org/10.1007/978-3-642-49686-8
  • A. Arapostathis, V. S. Borkar, E. Fernández-Gaucherand, M. K. Ghosh, S. I. Marcus, Discrete-time controlled Markov processes with average cost criterion: a survey, SIAM J. Control Optim. 31(2) (1993) 282–344. https://doi.org/10.1137/0331018
13 thms1 active userReviewed
Convex OptimizationOptimizationProbability·Captain: mikedeng1

Optimality and Duality Theory for Stochastic Optimization Problems with Nonlinear Dominance Constraints 1: Under Uniform Dominance, Optimal Solutions Have Concave Utility and L∞ MultipliersResearch Paper

Motivation

Stochastic programs often optimize a decision that changes several random outcomes at once. A reference outcome may be acceptable even when no fixed threshold captures its risk: one wants the new outcome to be preferable under every increasing concave assessment of gains. Second order stochastic dominance expresses that comparison. Dentcheva and Ruszczyński study optimization with several such constraints, each imposed on a nonlinear outcome operator, and show how the constraint multipliers can be represented by utility functions rather than scalar penalties (Dentcheva–Ruszczyński, 2004). Their earlier paper, Optimization with stochastic dominance constraints, treats the pure dominance case without the nonlinear decision map; the present result adds decision dependent outcomes, multiple constraints, and split variables. Ogryczak and Ruszczyński's second performance function supplies the stochastic order used here (Ogryczak–Ruszczyński, 2002).

The utility interpretation matters when a modeler wants a certificate explaining why a solution satisfies a risk preference expressed by dominance. The theorem identifies a concave utility for each binding dominance constraint and an essentially bounded multiplier for each comparison between the split outcome and the outcome produced by the decision. The source is a revised April 2003 author manuscript, later published in Mathematical Programming in 2004; the page and equation numbers below follow that manuscript (author manuscript).

Setting

Work on a probability space (Ω,F,P)(\Omega,\mathcal F,P)(Ω,F,P). An integrable random outcome is a measurable real function with finite expected absolute value; L1\mathcal L^1L1 denotes these outcomes, and L∞\mathcal L^\inftyL∞ denotes essentially bounded ones. The decisions lie in a convex set ZZZ inside a separable locally convex Hausdorff real vector space Z\mathcal ZZ. An integrable objective outcome H(z)H(z)H(z) and integrable constraint outcomes Gi(z)G_i(z)Gi​(z) depend continuously in the L1\mathcal L^1L1 norm on zzz. Almost every realized map z↦H(z)(ω)z\mapsto H(z)(\omega)z↦H(z)(ω) and z↦Gi(z)(ω)z\mapsto G_i(z)(\omega)z↦Gi​(z)(ω) is concave and continuous on all of Z\mathcal ZZ. Fixed integrable outcomes YiY_iYi​ serve as references; the iiith comparison is required over a bounded interval [ai,bi][a_i,b_i][ai​,bi​].

For an outcome XXX, its second performance function is the area below its distribution function:

F2(X;η)=∫−∞ηP{X≤ξ} dξ.F_2(X;\eta)=\int_{-\infty}^{\eta}P\{X\le\xi\}\,d\xi.F2​(X;η)=∫−∞η​P{X≤ξ}dξ.

The split program (11)–(14) chooses z∈Zz\in Zz∈Z and X=(X1,…,Xm)∈(L1)mX=(X_1,\ldots,X_m)\in(\mathcal L^1)^mX=(X1​,…,Xm​)∈(L1)m to maximize EH(z)\mathbb E H(z)EH(z), subject to F2(Xi;η)≤F2(Yi;η)F_2(X_i;\eta)\le F_2(Y_i;\eta)F2​(Xi​;η)≤F2​(Yi​;η) for every η∈[ai,bi]\eta\in[a_i,b_i]η∈[ai​,bi​], and Xi≤Gi(z)X_i\le G_i(z)Xi​≤Gi​(z) almost surely. Larger outcomes are preferred, so a dominating XiX_iXi​ has the smaller F2F_2F2​ curve. The split variables expose the dominance and decision coupling as separate constraints (manuscript, pp. 3–4).

The utility cone U1([a,b])\mathcal U_1([a,b])U1​([a,b]) consists of concave nondecreasing functions u:R→Ru:\mathbb R\to\mathbb Ru:R→R that vanish for t≥bt\ge bt≥b and are affine with a nonnegative slope for t≤at\le at≤a. Given uiu_iui​ in these cones and θi∈L∞\theta_i\in\mathcal L^\inftyθi​∈L∞, the Lagrangian is

L(z,X,u,θ)=E ⁣[H(z)+∑i=1m(ui(Xi)−ui(Yi)+θi(Gi(z)−Xi))].L(z,X,u,\theta)=\mathbb E\!\left[H(z)+\sum_{i=1}^m\bigl(u_i(X_i)-u_i(Y_i)+\theta_i(G_i(z)-X_i)\bigr)\right].L(z,X,u,θ)=E[H(z)+i=1∑m​(ui​(Xi​)−ui​(Yi​)+θi​(Gi​(z)−Xi​))].

Uniform dominance means one decision z~∈Z\tilde z\in Zz~∈Z makes every dominance inequality uniformly strict on its interval: for each iii, F2(Yi;η)−F2(Gi(z~);η)F_2(Y_i;\eta)-F_2(G_i(\tilde z);\eta)F2​(Yi​;η)−F2​(Gi​(z~);η) has a positive lower bound over [ai,bi][a_i,b_i][ai​,bi​] (Definition 1, p. 7).

Formalization targets

Utility and bounded multiplier characterization

Theorem 2 is the goal. Under uniform dominance, every optimum (z^,X^)(\hat z,\hat X)(z^,X^) of the split program admits u^i∈U1([ai,bi])\hat u_i\in\mathcal U_1([a_i,b_i])u^i​∈U1​([ai​,bi​]) and nonnegative θ^i∈L∞\hat\theta_i\in\mathcal L^\inftyθ^i​∈L∞ with

L(z^,X^,u^,θ^)=max⁡z∈Z, X∈(L1)mL(z,X,u^,θ^),L(\hat z,\hat X,\hat u,\hat\theta)=\max_{z\in Z,\,X\in(\mathcal L^1)^m}L(z,X,\hat u,\hat\theta),L(z^,X^,u^,θ^)=z∈Z,X∈(L1)mmax​L(z,X,u^,θ^), Eu^i(X^i)=Eu^i(Yi),θ^i(X^i−Gi(z^))=0almost surely.\mathbb E\hat u_i(\hat X_i)=\mathbb E\hat u_i(Y_i),\qquad \hat\theta_i\bigl(\hat X_i-G_i(\hat z)\bigr)=0\quad\text{almost surely}.Eu^i​(X^i​)=Eu^i​(Yi​),θ^i​(X^i​−Gi​(z^))=0almost surely.

Conversely, an attained Lagrangian maximum satisfying the split constraints and these complementarity equations is a primal optimum. The milestone list follows the source's measure multiplier equations (23)–(24), the measure to utility identity (25), Theorem 1's expected concave subgradient characterization, and the converse's weak duality inequality (manuscript, pp. 5, 8–10).

Significance

The result gives a concrete optimality certificate in a program whose constraints compare entire outcome distributions. Each utility multiplier represents the active part of one dominance constraint. Each θi\theta_iθi​ accounts for the almost sure inequality linking a split outcome to the decision. The equalities show exactly where those constraints are complementary, while the Lagrangian maximum compares the proposed solution with all integrable split outcomes. The paper derives a dual problem from the same Lagrangian in its following section (manuscript, p. 11).

The mathematical theorem is proved in the paper. This mission seeks a machine checked version of its definitions, measure identity, subgradient statement, and both directions of Theorem 2. The published second performance definition is reused as a reference; the nonlinear split program and its utility and measure Lagrangians require a development specific to this paper. The 2003 pure dominance mission contains related local drafts, but those items are not published and cannot currently be imported as platform theorems.

Difficulty

The dominance inequality contains a continuum of thresholds for each outcome. A scalar multiplier at one threshold cannot capture the whole constraint, while the dual object for continuous functions on [ai,bi][a_i,b_i][ai​,bi​] is a measure. The split inequality lives in L1\mathcal L^1L1, where the nonnegative cone has empty interior, so an ordinary interior point argument applied to all constraints at once does not match the paper's setting. The source also needs a subgradient of expected concave utility represented by an almost surely selected, essentially bounded random vector; the conclusion is stronger than merely knowing that the expected objective has a deterministic supporting functional (manuscript, pp. 5–9).

Formalization scope

The Lean development keeps the general separable locally convex Hausdorff decision space, the convex set ZZZ, and a finite index type for the mmm dominance constraints. Operators are function representatives with explicit integrability, continuity in L1\mathcal L^1L1, and samplewise concavity and continuity. Almost sure comparisons use the probability measure PPP; the null set for each realization condition precedes the quantifier over decisions. Split outcomes range only over integrable functions, and utility multipliers range over the exact cone U1([ai,bi])\mathcal U_1([a_i,b_i])U1​([ai​,bi​]). The L∞\mathcal L^\inftyL∞ condition includes almost sure strong measurability and essential boundedness. Maxima in Theorems 1 and 2 are attained maxima, expressed by membership and comparison against every competitor, never a real supremum with a default value.

The source prints a strictly positive affine slope in its definition of U1\mathcal U_1U1​, but immediately calls this class a cone and later uses the zero measure. The formalization uses c≥0c\ge0c≥0; with c>0c>0c>0, Theorem 2 is false for a slack dominance constraint. Uniform dominance is expressed as a positive lower bound rather than a real infimum. The measure milestone uses finite nonnegative measures supported on closed intervals, including endpoint atoms. These conditions exclude default zero integrals, an empty interval disguised by an infimum, and a vacuous utility class. Contributions to the measure to utility correspondence, integration identities, and expected concave subgradient infrastructure can be reused beyond this program.

Selected references

  • D. Dentcheva and A. Ruszczyński, Optimality and duality theory for stochastic optimization problems with nonlinear dominance constraints, Mathematical Programming (2004), DOI; revised author manuscript, April 2003.
  • D. Dentcheva and A. Ruszczyński, Optimization with stochastic dominance constraints, manuscript submitted for publication (2002), cited as reference [6] in the 2003 author manuscript.
  • W. Ogryczak and A. Ruszczyński, Dual stochastic dominance and related mean risk models, SIAM Journal on Optimization 13 (2002), DOI.
7 thms1 active userReviewed
Linear OptimizationTheoretical Computer Science·Captain: mikedeng1

Online Primal-Dual Algorithms for Covering and Packing 4: Online Rounding of the Fractional Routing Scheme Respects Capacities and Is O(log P(max)·[exp(1 + 2 ln m/u(min)) − 1])-CompetitiveResearch Paper

Motivation

In online routing of virtual circuits, connection requests between pairs of nodes of a capacitated network arrive one at a time. Each request must be accepted and routed on a single path with bandwidth 111, or rejected, immediately and irrevocably, and no edge may carry more than its capacity. The goal is to maximize the number of accepted requests (the throughput). The model goes back to Awerbuch, Azar and Plotkin (FOCS 1993), whose deterministic algorithm has a logarithmic competitive ratio when edge capacities are at least logarithmic in the size of the network, and it underlies the analysis of admission control in circuit-switched and bandwidth-reserved networks.

Buchbinder and Naor (Math. Oper. Res. 2009) recover an algorithm with the same competitive factor from a general recipe: an online primal–dual scheme first produces a feasible fractional routing online, and an online version of Raghavan's pessimistic estimator (J. Comput. Syst. Sci. 1988) then rounds it, also online. This mission formalizes that construction and its guarantee (Section 5.2 of the paper, with the Section 3 scheme it uses).

Timeline:

  • 1987–1988: Raghavan and Thompson introduce randomized rounding for multicommodity flow; Raghavan derandomizes it with pessimistic estimators.
  • 1993: Awerbuch, Azar and Plotkin give the deterministic throughput-competitive online routing algorithm.
  • 2005–2009: Buchbinder and Naor's primal–dual framework (ESA 2005; MOR 2009) derives an algorithm with the same factor systematically.

Setting

Let EEE be a finite set of mmm edges with capacities u(e)>0u(e) > 0u(e)>0, and u(min⁡)=min⁡eu(e)u(\min) = \min_e u(e)u(min)=mine​u(e). Requests r1,r2,…r_1, r_2, \dotsr1​,r2​,… arrive online; request rir_iri​ comes with a finite list P(ri)\mathcal P(r_i)P(ri​) of admissible paths, each a set of edges, all of size at most P(max⁡)P(\max)P(max).

A fractional routing assigns flows f(ri,P)≥0f(r_i, P) \ge 0f(ri​,P)≥0; it is feasible when ∑P∈P(ri)f(ri,P)≤1\sum_{P \in \mathcal P(r_i)} f(r_i, P) \le 1∑P∈P(ri​)​f(ri​,P)≤1 for every request and the load ∑ri∑P∋ef(ri,P)\sum_{r_i}\sum_{P \ni e} f(r_i, P)∑ri​​∑P∋e​f(ri​,P) of every edge is at most u(e)u(e)u(e). Its value is val(f)=∑ri∑Pf(ri,P)\mathrm{val}(f) = \sum_{r_i}\sum_P f(r_i, P)val(f)=∑ri​​∑P​f(ri​,P); OPT\mathrm{OPT}OPT is the largest value of a feasible routing of the arrived requests, an upper bound on the integral optimum.

The fractional scheme. The covering LP paired with the routing LP (the paper's primal, Fig. 3) has variables x(e)x(e)x(e) (cost u(e)u(e)u(e)) and Z(ri)Z(r_i)Z(ri​) (cost 111) with constraints ∑e∈Px(e)+Z(ri)≥1\sum_{e \in P} x(e) + Z(r_i) \ge 1∑e∈P​x(e)+Z(ri​)≥1. When rir_iri​ arrives, its paths are visited in order; for each path whose constraint fails, f(ri,P)f(r_i,P)f(ri​,P) is raised from 000 to the least value restoring it, while x(e)=max⁡(x(e),1ℓ(eB′Fe/(2u(e))−1))x(e) = \max\big(x(e), \tfrac1\ell(e^{B' F_e/(2u(e))}-1)\big)x(e)=max(x(e),ℓ1​(eB′Fe​/(2u(e))−1)) for e∈Pe \in Pe∈P and Z(ri)=max⁡(Z(ri),1ℓ(eB′f(ri)/2−1))Z(r_i) = \max\big(Z(r_i), \tfrac1\ell(e^{B' f(r_i)/2}-1)\big)Z(ri​)=max(Z(ri​),ℓ1​(eB′f(ri​)/2−1)) follow the flow (FeF_eFe​ the load of eee, f(ri)f(r_i)f(ri​) the flow of rir_iri​). The parameters are ℓ=P(max⁡)+1\ell = P(\max)+1ℓ=P(max)+1 and B′=2ln⁡(1+ℓ)B' = 2\ln(1+\ell)B′=2ln(1+ℓ).

The rounding. With the rounding scale B=exp⁡(1+ln⁡(2m)/u(min⁡))−1B = \exp(1 + \ln(2m)/u(\min)) - 1B=exp(1+ln(2m)/u(min))−1, the integral edge usage χ(e)\chi(e)χ(e) and the number sss of served requests, the potential is Φ=Φ1+Φ2\Phi = \Phi_1 + \Phi_2Φ=Φ1​+Φ2​,

Φ1=12exp⁡(val(f)2B−sln⁡2),Φ2=12m∑eexp⁡((1+ln⁡2mu(e))χ(e)−Fe).\Phi_1 = \tfrac12\exp\Big(\frac{\mathrm{val}(f)}{2B} - s\ln 2\Big), \qquad \Phi_2 = \frac1{2m}\sum_{e}\exp\Big(\Big(1+\frac{\ln 2m}{u(e)}\Big)\chi(e) - F_e\Big).Φ1​=21​exp(2Bval(f)​−sln2),Φ2​=2m1​e∑​exp((1+u(e)ln2m​)χ(e)−Fe​).

After the fractional round of rir_iri​, the algorithm serves rir_iri​ on a path P∈P(ri)P \in \mathcal P(r_i)P∈P(ri​) (adding 111 to χ(e)\chi(e)χ(e) for e∈Pe \in Pe∈P) if this gives potential at most the potential Φstart\Phi^{\mathrm{start}}Φstart before the round; otherwise it rejects rir_iri​.

Formalization targets

Goal: Lemma 5.4

For every request sequence, the algorithm never exceeds a capacity, and for every feasible fractional routing fff,

χ(e)≤u(e)  ∀e,∑riχ(ri) ≥ val(f)4Bln⁡2⋅ln⁡(P(max⁡)+2)−1.\chi(e) \le u(e)\ \ \forall e, \qquad \sum_{r_i}\chi(r_i) \ \ge\ \frac{\mathrm{val}(f)}{4B\ln 2\cdot\ln(P(\max)+2)} - 1 .χ(e)≤u(e)  ∀e,ri​∑​χ(ri​) ≥ 4Bln2⋅ln(P(max)+2)val(f)​−1.

Milestone: Theorem 3.2 (packing half, on routing)

The fractional scheme's flows falgf^{\mathrm{alg}}falg are feasible and val(f)≤2ln⁡(P(max⁡)+2) val(falg)\mathrm{val}(f) \le 2\ln(P(\max)+2)\,\mathrm{val}(f^{\mathrm{alg}})val(f)≤2ln(P(max)+2)val(falg) for every feasible fff.

Milestone: Lemma 5.3

Φ≤1\Phi \le 1Φ≤1 initially, Φ>0\Phi > 0Φ>0 always, and whenever the flows of a request are raised by a non-negative amount of total at most 111, serving the request on some path or rejecting it does not increase Φ\PhiΦ.

Significance

The result shows that a deterministic online algorithm for throughput-competitive routing, previously designed by hand, falls out of two generic components: an online fractional packing scheme and an online pessimistic estimator. A side product is that the fractional phase alone produces, online, a near-optimal routing that respects all capacities exactly, independently of their size. When u(min⁡)≥log⁡nu(\min) \ge \log nu(min)≥logn the rounding loses only a constant factor and the algorithm is O(log⁡P(max⁡))O(\log P(\max))O(logP(max))-competitive, as in Awerbuch–Azar–Plotkin.

The results are proved in the paper; none is machine-checked. A formalization supplies a checked instance of the online primal–dual method together with derandomized online rounding, and fixes the constants the paper leaves inside O(⋅)O(\cdot)O(⋅). Related platform content: the monograph's OnlinePrimalDual.Routing.per_copy_guarantee and routing_competitive concern the Buchbinder–Naor (1,O(log⁡n))(1, O(\log n))(1,O(logn))-competitive algorithm with copies of the graph, a different scheme.

Difficulty

The fractional guarantee is argued in the paper continuously (rates of change of the primal and dual values), while the scheme as formalized is discrete: each flow is the least value restoring a constraint, and every primal variable is a maximum whose branch may switch during the increase. The continuous argument does not transfer verbatim, and feasibility depends on the least value restoring the constraint with equality.

The rounding is a derandomization. The existence of a good path or a good rejection is established in the paper as an expectation over a random trial; a deterministic statement about finitely many alternatives is what the mission asks for, with all m+1m+1m+1 exponential terms of Φ\PhiΦ under control at once. A frequent first attempt compares with the potential after the fractional round; the rule compares with Φstart\Phi^{\mathrm{start}}Φstart, before the flow increase, and the guarantee is stated for that comparison.

Formalization scope

  • Edges are a non-empty Fintype E; capacities are real and positive. A request is a List (Finset E) of paths; the request sequence is a list. Simple paths of a graph are a special case; nothing in the argument uses graph structure. P(max⁡)P(\max)P(max) is a parameter with every path of size at most P(max⁡)P(\max)P(max).
  • A routing is a List (List ℝ) of the shape of the request sequence. OPT\mathrm{OPT}OPT is quantified as "every feasible fractional routing".
  • The continuous increase is its discrete equivalent (an attained sInf). Ties among good paths are broken by list order; only paths that received flow in the current round are candidates for serving, so a request whose flow was not increased is rejected.
  • Explicit constants replacing O(⋅)O(\cdot)O(⋅): Theorem 3.2's O(log⁡ℓ)O(\log \ell)O(logℓ) becomes 2ln⁡(1+ℓ)=2ln⁡(P(max⁡)+2)2\ln(1+\ell) = 2\ln(P(\max)+2)2ln(1+ℓ)=2ln(P(max)+2); Lemma 5.4's O(log⁡P(max⁡)⋅[exp⁡(1+2ln⁡m/u(min⁡))−1])O(\log P(\max)\cdot[\exp(1+2\ln m/u(\min))-1])O(logP(max)⋅[exp(1+2lnm/u(min))−1]) becomes 4Bln⁡2⋅ln⁡(P(max⁡)+2)4B\ln 2\cdot\ln(P(\max)+2)4Bln2⋅ln(P(max)+2) with additive −1-1−1, where B=exp⁡(1+ln⁡(2m)/u(min⁡))−1B = \exp(1+\ln(2m)/u(\min))-1B=exp(1+ln(2m)/u(min))−1 is the scale chosen on p. 15. The printed "2ln⁡m2\ln m2lnm" differs from the proof's "ln⁡2m\ln 2mln2m"; the proof's constant is used (it is at least as strong for m≥2m \ge 2m≥2). All logarithms are natural.
  • The algorithm is a fully specified function: a formalization in which requests are never served, or OPT\mathrm{OPT}OPT is a free variable pinned by hypotheses, would trivialize the goal and is ruled out.
  • Not included: the covering half of Theorem 3.2, and the remark on u(min⁡)≥log⁡nu(\min) \ge \log nu(min)≥logn.

Contributions welcome: proofs of the milestones, a reusable lemma "convex combination ≤\le≤ value ⇒\Rightarrow⇒ some outcome ≤\le≤ value" for derandomization, and the discrete-to-continuous bridge for the Section 3 scheme.

Selected references

  • N. Buchbinder, J. Naor, Online Primal-Dual Algorithms for Covering and Packing, Mathematics of Operations Research, 2009. https://doi.org/10.1287/moor.1080.0363
  • B. Awerbuch, Y. Azar, S. Plotkin, Throughput-Competitive On-Line Routing, Proc. 34th FOCS, pp. 32–40, 1993. https://doi.org/10.1109/SFCS.1993.366884
  • P. Raghavan, Probabilistic construction of deterministic algorithms: approximating packing integer programs, J. Comput. Syst. Sci. 37(2), 1988. https://doi.org/10.1016/0022-0000(88)90003-7
  • P. Raghavan, C. D. Thompson, Randomized rounding: a technique for provably good algorithms and algorithmic proofs, Combinatorica 7(4), 1987. https://doi.org/10.1007/BF02579324
7 thms1 active userReviewed
Linear OptimizationTheoretical Computer Science·Captain: mikedeng1

Online Primal-Dual Algorithms for Covering and Packing 1: The Online Fractional Packing Scheme Is B-Competitive and Violates Each Packing Constraint by at Most 2 log(1 + n·a_i(max)/a_i(min))/BResearch Paper

Motivation

Many resource-allocation problems arrive one request at a time and must be answered immediately: a bandwidth request is admitted or refused when it appears, an advertiser's budget is charged when a query arrives, a job is accepted before later jobs are seen. Their linear-programming relaxations are packing problems: maximize a total profit subject to capacity constraints, where the variables are revealed online and each must be set irrevocably when it is revealed. A standard yardstick for such an online algorithm is its competitive ratio, the worst-case ratio between the offline optimum and the algorithm's value.

Buchbinder and Naor (Math. Oper. Res. 2009) gave a single online primal-dual scheme for the general online fractional packing problem, together with a matching scheme for covering. The scheme raises the newly revealed packing variable while increasing the dual covering variables along an exponential curve, and its analysis is a short primal-dual argument. Earlier online algorithms for throughput-competitive routing and for set cover (Alon et al. 2009) can be read as instances of it, and the same template later became the basis of a monograph on the primal-dual approach to online algorithms (Buchbinder, Naor 2009).

This mission formalizes the paper's headline result, Theorem 3.1, for the scheme exactly as the paper defines it.

Setting

Fix a finite set III of n≥1n\ge1n≥1 packing constraints (equivalently, primal covering variables), with known capacities c(i)>0c(i)>0c(i)>0. Packing variables y(1),…,y(m)y(1),\dots,y(m)y(1),…,y(m) arrive one per round; in round jjj the variable y(j)y(j)y(j) is revealed together with its non-negative column a(i,j)a(i,j)a(i,j), i∈Ii\in Ii∈I. The offline problems form the primal-dual pair of Figure 1 of the paper:

(P) min⁡∑ic(i)x(i)  s.t. ∑ia(i,j)x(i)≥1 ∀j, x≥0;(D) max⁡∑jy(j)  s.t. ∑ja(i,j)y(j)≤c(i) ∀i, y≥0.\text{(P)}\ \min\sum_i c(i)x(i)\ \text{ s.t. } \sum_i a(i,j)x(i)\ge1\ \forall j,\ x\ge0;\qquad \text{(D)}\ \max\sum_j y(j)\ \text{ s.t. } \sum_j a(i,j)y(j)\le c(i)\ \forall i,\ y\ge0.(P) mini∑​c(i)x(i)  s.t. i∑​a(i,j)x(i)≥1 ∀j, x≥0;(D) maxj∑​y(j)  s.t. j∑​a(i,j)y(j)≤c(i) ∀i, y≥0.

The profit of every y(j)y(j)y(j) is normalized to 111. Every column is assumed to have a positive entry; otherwise the packing problem is unbounded. An online algorithm may set y(j)y(j)y(j) only in round jjj and never changes it later.

The scheme with parameter B>0B>0B>0 keeps a primal vector xxx (initially 000) and the dual vector yyy. In round jjj it computes the prefix maximum ai(max⁡)=max⁡k≤ja(i,k)a_i(\max)=\max_{k\le j}a(i,k)ai​(max)=maxk≤j​a(i,k). If the new covering constraint ∑ia(i,j)x(i)≥1\sum_i a(i,j)x(i)\ge1∑i​a(i,j)x(i)≥1 already holds, it sets y(j)=0y(j)=0y(j)=0. Otherwise it sets y(j)y(j)y(j) to the least t≥0t\ge0t≥0 at which the constraint holds after every x(i)x(i)x(i) is replaced by

max⁡{x(i), 1n ai(max⁡)[exp⁡(B2c(i)∑k=1ja(i,k)y(k))−1]},y(j)=t.\max\Big\{x(i),\ \frac{1}{n\,a_i(\max)}\Big[\exp\Big(\frac{B}{2c(i)}\sum_{k=1}^{j}a(i,k)y(k)\Big)-1\Big]\Big\},\qquad y(j)=t.max{x(i), nai​(max)1​[exp(2c(i)B​k=1∑j​a(i,k)y(k))−1]},y(j)=t.

After rrr rounds, X(r)=∑ic(i)x(i)X(r)=\sum_i c(i)x(i)X(r)=∑i​c(i)x(i) is the primal value and Y(r)=∑k≤ry(k)Y(r)=\sum_{k\le r}y(k)Y(r)=∑k≤r​y(k) the dual value. For the analysis, ai(max⁡)a_i(\max)ai​(max) and ai(min⁡)a_i(\min)ai​(min) also denote the largest and the smallest non-zero coefficient of row iii over all mmm columns.

Formalization targets

Goal: Theorem 3.1

For every B>0B>0B>0, the scheme's dual solution is non-negative; after every round rrr, every non-negative y′y'y′ that satisfies the packing constraints restricted to the first rrr columns has ∑k≤ry′(k)≤B Y(r)\sum_{k\le r}y'(k)\le B\,Y(r)∑k≤r​y′(k)≤BY(r) (the scheme is BBB-competitive); and after all mmm rounds, for every iii,

∑k=1ma(i,k) y(k) ≤ c(i)⋅2log⁡(1+n ai(max⁡)/ai(min⁡))B.\sum_{k=1}^{m}a(i,k)\,y(k)\ \le\ c(i)\cdot\frac{2\log\big(1+n\,a_i(\max)/a_i(\min)\big)}{B}.k=1∑m​a(i,k)y(k) ≤ c(i)⋅B2log(1+nai​(max)/ai​(min))​.

The paper states the second part as c(i)⋅O((log⁡n+log⁡(ai(max⁡)/ai(min⁡)))/B)c(i)\cdot O\big((\log n+\log(a_i(\max)/a_i(\min)))/B\big)c(i)⋅O((logn+log(ai​(max)/ai​(min)))/B); the bound above is the one its proof establishes.

Milestones: the three claims of the proof

  1. Claim (i): X(r)≤B⋅Y(r)X(r)\le B\cdot Y(r)X(r)≤B⋅Y(r) after every round rrr.
  2. Claim (ii): after every round, x≥0x\ge0x≥0 and xxx satisfies every covering constraint revealed so far; no x(i)x(i)x(i) ever decreases.
  3. Claim (iii): the violation bound of the goal.

Significance

Theorem 3.1 says that a solution within factor BBB of the optimum can be maintained online at the price of overloading each packing constraint by a factor of order (log⁡n+log⁡(amax⁡/amin⁡))/B(\log n+\log(a_{\max}/a_{\min}))/B(logn+log(amax​/amin​))/B. Scaling the output down by the overload gives a feasible online packing solution with competitive ratio O(log⁡n+log⁡(amax⁡/amin⁡))O(\log n+\log(a_{\max}/a_{\min}))O(logn+log(amax​/amin​)), and Lemma 3.1 of the paper shows that no online algorithm does better up to constant factors. The same trade-off underlies the paper's online rounding results for routing (§5.2) and, through the covering counterpart, for set cover (§5.1).

The result is proved in the paper. As far as the platform's record shows, it is not formalized: the published OnlinePrimalDual.GeneralPacking.theorem14_1 (a restatement of the monograph's version) takes the inequality X≤BYX\le BYX≤BY and primal feasibility as hypotheses on arbitrary vectors x,yx,yx,y and derives competitiveness by weak duality; it does not mention the scheme and has no violation bound. The published OnlinePrimalDual.GeneralPacking.lemma14_2 is the matching lower bound (Lemma 3.1) and is not part of this mission. A formalization here would give the first machine-checked guarantee for the scheme itself and a reusable analysis pattern (a potential bound integrated along a monotone path) for the other schemes of the paper.

Difficulty

The paper's argument is a derivative comparison along a continuous process: while y(j)y(j)y(j) rises, ∂X/∂y(j)≤B\partial X/\partial y(j)\le B∂X/∂y(j)≤B. In the discrete formulation each round jumps directly to the least admissible y(j)y(j)y(j), and each x(i)x(i)x(i) is a maximum of its old value and an exponential, so it is continuous but not differentiable where the maximum switches; the comparison must be turned into an integral inequality over [0,y(j)][0,y(j)][0,y(j)] for such functions. The prefix maximum ai(max⁡)a_i(\max)ai​(max) changes between rounds, and the claim that this never lowers or raises the primal value needs an invariant (x(i)x(i)x(i) is always at least the current increment value). The violation bound rests on a second invariant, x(i)≤1/ai(min⁡)x(i)\le1/a_i(\min)x(i)≤1/ai​(min), which holds because y(j)y(j)y(j) is the least admissible value; any formalization that loses minimality (for example by taking an arbitrary admissible ttt) loses claim (iii).

Formalization scope

  • The instance is the published OnlinePrimalDual.GeneralPacking.GeneralInstance I (Fin m) (costs c>0c>0c>0, coefficients a≥0a\ge0a≥0); n=∣I∣n=|I|n=∣I∣ with [Nonempty I]; columns are Fin m, arrive in index order, and are 0-based in Lean, so "the first rrr columns" is (k : ℕ) < r. m≥1m\ge1m≥1 is [NeZero m].
  • ai(max⁡)a_i(\max)ai​(max), ai(min⁡)a_i(\min)ai​(min) over all columns are the published aMax and aMin. For a row with no non-zero coefficient aMin is 000 and both sides of the violation bound are 000.
  • The scheme is a function, stateAfter inst B r. The continuous loop of the paper is replaced by its discrete implementation, which the paper itself prescribes (p. 4): y(j)y(j)y(j) is the least t≥0t\ge0t≥0 restoring the new covering constraint, written as sInf. Every theorem assumes that every column has a positive entry (the paper's standing assumption, p. 4); this makes the infimum attained. If the prefix maximum is 000, Lean's 1/0=01/0=01/0=0 gives increment 000, which agrees with the paper's bracket being 000.
  • Explicit constant. The paper's c(i)⋅O((log⁡n+log⁡(ai(max⁡)/ai(min⁡)))/B)c(i)\cdot O((\log n+\log(a_i(\max)/a_i(\min)))/B)c(i)⋅O((logn+log(ai​(max)/ai​(min)))/B) is instantiated as c(i)⋅2log⁡(1+n ai(max⁡)/ai(min⁡))/Bc(i)\cdot 2\log(1+n\,a_i(\max)/a_i(\min))/Bc(i)⋅2log(1+nai​(max)/ai​(min))/B, from claim (iii) of the proof. Logarithms are natural (Real.log), since they invert Real.exp.
  • BBB-competitiveness is stated against every non-negative feasible packing solution of every prefix of the input, not only at the end.
  • A formalization in which the inequality X≤BYX\le BYX≤BY, primal feasibility, or the bound x(i)≤1/ai(min⁡)x(i)\le1/a_i(\min)x(i)≤1/ai​(min) is assumed rather than derived from the scheme is ruled out: every statement here is about the vectors the scheme computes from the instance.
  • Needed infrastructure: monotonicity and continuity of the per-round primal path, attainment of the infimum, an integral (or mean-value) form of the derivative comparison for maxima of exponentials, and weak duality for finite LPs. The weak-duality step and the per-round integration lemma are reusable for missions 2–4 of this series. Proofs of the milestones, and of helper lemmas such as the invariant x(i)≤1/ai(min⁡)x(i)\le1/a_i(\min)x(i)≤1/ai​(min), are welcome.

Selected references

  • N. Buchbinder, J. Naor, Online Primal-Dual Algorithms for Covering and Packing, Mathematics of Operations Research, 2009. https://doi.org/10.1287/moor.1080.0363
  • N. Buchbinder, J. Naor, The Design of Competitive Online Algorithms via a Primal-Dual Approach, Foundations and Trends in Theoretical Computer Science 3(2–3), 2009. https://doi.org/10.1561/0400000024
  • N. Alon, B. Awerbuch, Y. Azar, N. Buchbinder, J. Naor, The Online Set Cover Problem, SIAM Journal on Computing 39(2), 2009. https://doi.org/10.1137/060661946
8 thms1 active userReviewed
Dynamic ProgrammingMarkov Chain·Captain: mikedeng1

Discrete-Time Controlled Markov Processes with Average Cost Criterion: A Survey 2: Uniformly Bounded Mean Return Times Make the Differential Discounted Values Uniformly BoundedResearch Paper

Motivation

Controlled Markov processes with the average cost criterion model systems that run indefinitely, such as queues, inventories, maintenance and communication networks, where only the long-run cost per unit time matters. The standard route to an optimal stationary policy goes through the average cost optimality equation (ACOE). The ACOE is usually obtained by the vanishing discount method: solve the discounted problem for each discount factor β<1\beta<1β<1 and let β→1\beta\to1β→1. The method works only when the differences of discounted values stay bounded as β→1\beta\to1β→1. Conditions that guarantee this are therefore central in the survey of Arapostathis, Borkar, Fernández-Gaucherand, Ghosh and Marcus (SIAM J. Control Optim. 31 (1993), §5).

This mission formalizes one such condition, due to Ross: if the mean return time to a fixed state is bounded uniformly over all stationary policies and initial states, the differential discounted value functions are bounded uniformly in the discount factor and the state.

Timeline. Derman (Management Sci. 9 (1962); survey reference [38]) and Derman–Veinott (Ann. Math. Statist. 38 (1967); survey reference [43]) introduced recurrence conditions of this kind for countable-state processes. Ross (Ann. Math. Statist. 39 (1968), survey reference [147]; Introduction to Stochastic Dynamic Programming, 1983, survey reference [150]) showed, for bounded costs, that under a Derman–Veinott type recurrence condition hβh_\betahβ​ is bounded uniformly in β\betaβ, and obtained a bounded solution of the ACOE by letting β↑1\beta\uparrow1β↑1 (survey, pp. 291 and 301). Later work replaced the condition with weaker ones (survey Assumptions 5.1–5.3) and with Sennott's conditions (survey Theorem 5.9).

Setting

The state space is S={0,1,2,… }S=\{0,1,2,\dots\}S={0,1,2,…}. In each state iii, an action aaa is chosen from a nonempty compact set U(i)U(i)U(i) of a metric space AAA. The one-stage cost c(i,a)c(i,a)c(i,a) is nonnegative and the next state is drawn from the transition law P(⋅∣i,a)P(\cdot\mid i,a)P(⋅∣i,a). For fixed i,ji,ji,j, the maps a↦c(i,a)a\mapsto c(i,a)a↦c(i,a) and a↦P(j∣i,a)a\mapsto P(j\mid i,a)a↦P(j∣i,a) are continuous on U(i)U(i)U(i). A policy π∈Π\pi\in\Piπ∈Π chooses the action at time ttt at random, given the whole history, and must choose from U(Xt)U(X_t)U(Xt​). A stationary deterministic policy f∈ΠSDf\in\Pi_{SD}f∈ΠSD​ is a map f:S→Af:S\to Af:S→A with f(i)∈U(i)f(i)\in U(i)f(i)∈U(i). PiπP^\pi_iPiπ​ and EiπE^\pi_iEiπ​ denote the law and the expectation of the controlled process (Xt,At)(X_t,A_t)(Xt​,At​) started at iii.

For a discount factor β∈(0,1)\beta\in(0,1)β∈(0,1), the discounted cost and the optimal discounted cost are

Jβ(i,π)=Eiπ[∑t=0∞βtc(Xt,At)],Jβ∗(i)=inf⁡π∈ΠJβ(i,π).J_\beta(i,\pi)=E^\pi_i\Big[\sum_{t=0}^\infty\beta^t c(X_t,A_t)\Big],\qquad J^*_\beta(i)=\inf_{\pi\in\Pi}J_\beta(i,\pi).Jβ​(i,π)=Eiπ​[t=0∑∞​βtc(Xt​,At​)],Jβ∗​(i)=π∈Πinf​Jβ​(i,π).

A policy f∈ΠSDf\in\Pi_{SD}f∈ΠSD​ is β\betaβ-discount optimal if Jβ(i,f)=Jβ∗(i)J_\beta(i,f)=J^*_\beta(i)Jβ​(i,f)=Jβ∗​(i) for all iii. The differential discounted value function is

hβ(i)=Jβ∗(i)−Jβ∗(0),h_\beta(i)=J^*_\beta(i)-J^*_\beta(0),hβ​(i)=Jβ∗​(i)−Jβ∗​(0),

measured relative to the fixed state 000. The return time to 000 is

τ=min⁡{t≥1: Xt=0},\tau=\min\{t\ge1:\ X_t=0\},τ=min{t≥1: Xt​=0},

with τ=∞\tau=\inftyτ=∞ if the process never returns. Throughout, as in §5.1 of the survey, the cost is bounded: c(i,a)≤Mc(i,a)\le Mc(i,a)≤M on admissible pairs.

Formalization targets

Goal: Theorem 5.3

If there is a constant K>0K>0K>0 with

Eif[τ]<Kfor all f∈ΠSD, i∈S,(5.7)E^f_i[\tau]<K\qquad\text{for all } f\in\Pi_{SD},\ i\in S, \tag{5.7}Eif​[τ]<Kfor all f∈ΠSD​, i∈S,(5.7)

then there is a constant BBB such that

∣hβ(i)∣≤Bfor all β∈(0,1), i∈S.|h_\beta(i)|\le B\qquad\text{for all }\beta\in(0,1),\ i\in S.∣hβ​(i)∣≤Bfor all β∈(0,1), i∈S.

This is the theorem as printed: it asserts only uniform boundedness and fixes no constant.

Milestones

  1. Theorem 2.1 (iii): for every β∈(0,1)\beta\in(0,1)β∈(0,1), a β\betaβ-discount optimal fβ∈ΠSDf_\beta\in\Pi_{SD}fβ​∈ΠSD​ exists.
  2. (5.8): for such an fβf_\betafβ​, Jβ∗(i)≤M Eifβ[τ]+Jβ∗(0) Eifβ[βτ]J^*_\beta(i)\le M\,E^{f_\beta}_i[\tau]+J^*_\beta(0)\,E^{f_\beta}_i[\beta^\tau]Jβ∗​(i)≤MEifβ​​[τ]+Jβ∗​(0)Eifβ​​[βτ].
  3. (5.9): Jβ∗(i)−βJβ∗(0)≤MKJ^*_\beta(i)-\beta J^*_\beta(0)\le MKJβ∗​(i)−βJβ∗​(0)≤MK.
  4. Jensen step: Jβ∗(i)≥Jβ∗(0) Eifβ[βτ]≥Jβ∗(0) βKJ^*_\beta(i)\ge J^*_\beta(0)\,E^{f_\beta}_i[\beta^\tau]\ge J^*_\beta(0)\,\beta^KJβ∗​(i)≥Jβ∗​(0)Eifβ​​[βτ]≥Jβ∗​(0)βK.
  5. (5.10): Jβ∗(0)−Jβ∗(i)≤(1−βK)Jβ∗(0)≤(1−βK)M1−β≤MKJ^*_\beta(0)-J^*_\beta(i)\le(1-\beta^K)J^*_\beta(0)\le(1-\beta^K)\frac{M}{1-\beta}\le MKJβ∗​(0)−Jβ∗​(i)≤(1−βK)Jβ∗​(0)≤(1−βK)1−βM​≤MK.
  6. Explicit bound: ∣hβ(i)∣≤MK|h_\beta(i)|\le MK∣hβ​(i)∣≤MK. This is stronger than the goal and is the constant the survey's argument yields.

Significance

The result. Theorem 5.3 verifies the hypothesis of the vanishing discount theorem (Theorem 5.2 of the survey) from a condition on the uncontrolled dynamics of stationary policies. Theorem 5.2 then gives a bounded solution (ρ,h)(\rho,h)(ρ,h) of the ACOE, an average optimal stationary policy, and the limit lim⁡β→1(1−β)Jβ∗(i)=ρ\lim_{\beta\to1}(1-\beta)J^*_\beta(i)=\rholimβ→1​(1−β)Jβ∗​(i)=ρ. Mean return times can often be estimated directly, for instance through Foster–Lyapunov drift arguments on queues, which makes (5.7) checkable in applications. The explicit bound MKMKMK also controls the span of the relative value function.

Formalizing it. The result is proved in the literature; to our knowledge it has not been machine-checked. A formal proof needs discounted dynamic programming on a countable state space with compact action sets, the existence of optimal stationary policies (Theorem 2.1 (iii), which the survey cites without proof), and the strong Markov property of the controlled chain at a return time. All of these are reusable well beyond this mission.

Difficulty

The estimates (5.9) and (5.10) are elementary once (5.8) and the existence of fβf_\betafβ​ are available. The weight lies elsewhere.

  • Optimal stationary policies. The infimum defining Jβ∗J^*_\betaJβ∗​ ranges over all history-dependent randomized policies. Bringing it down to a single stationary deterministic policy requires the discounted optimality equation, a measurable selection of minimizers on compact action sets, and a verification argument against arbitrary policies.
  • Restarting at τ\tauτ. (5.8) splits the discounted cost at the random time τ\tauτ. The tail must be identified with βτ\beta^\tauβτ times the discounted cost from state 000. This is the strong Markov property for the process built by the Ionescu-Tulcea theorem, applied at a stopping time that may be infinite.

Formalization scope

  • The state space is ℕ; the action space is a metric space with its Borel σ\sigmaσ-algebra. The model CMP carries compact nonempty U(i)U(i)U(i), a nonnegative measurable cost, and continuity of c(i,⋅)c(i,\cdot)c(i,⋅) and P(j∣i,⋅)P(j\mid i,\cdot)P(j∣i,⋅) on U(i)U(i)U(i), the standing assumptions of §5.
  • Policies are history-dependent, randomized and admissible. The path measure is Mathlib's Kernel.trajMeasure. Jβ∗J^*_\betaJβ∗​ is an infimum over all such policies, not over Markov or stationary policies only.
  • Costs are lower Lebesgue integrals in [0,∞][0,\infty][0,∞]. hβh_\betahβ​ is the difference of the real parts of Jβ∗(i)J^*_\beta(i)Jβ∗​(i) and Jβ∗(0)J^*_\beta(0)Jβ∗​(0). This is exact here because bounded cost gives Jβ∗≤M/(1−β)<∞J^*_\beta\le M/(1-\beta)<\inftyJβ∗​≤M/(1−β)<∞.
  • Explicit choices:
    • The bounded-cost hypothesis c≤Mc\le Mc≤M on admissible pairs is a binder of every §5.1 statement. It is the section's standing assumption, and without it the theorem is false.
    • τ\tauτ counts from t≥1t\ge1t≥1 and takes values in N∪{∞}\mathbb N\cup\{\infty\}N∪{∞}, so (5.7) applies from i=0i=0i=0 and forces τ<∞\tau<\inftyτ<∞ almost surely; βτ=0\beta^\tau=0βτ=0 on {τ=∞}\{\tau=\infty\}{τ=∞}.
    • The typo βn\beta^nβn in (5.8) is read as βt\beta^tβt.
    • K≥1K\ge1K≥1 in (5.10) is not assumed; it follows from (5.7).
  • Theorem 2.1 is stated in the survey for Borel models under Assumptions 2.1–2.3. Here it is posed in the countable model, where those assumptions follow from the §5 continuity and compactness assumptions.
  • The goal's bound BBB is quantified before β\betaβ and iii. A per-β\betaβ or per-state bound would be trivial, since every hβ(i)h_\beta(i)hβ​(i) is a finite number.
  • Welcome contributions include discounted dynamic programming on countable state spaces, the strong Markov property for trajMeasure, and return-time estimates.

Selected references

  • A. Arapostathis, V. S. Borkar, E. Fernández-Gaucherand, M. K. Ghosh, S. I. Marcus, Discrete-time controlled Markov processes with average cost criterion: a survey, SIAM J. Control Optim. 31(2) (1993) 282–344. https://doi.org/10.1137/0331018
  • C. Derman, On sequential decisions and Markov chains, Management Sci. 9 (1962) 16–24. https://doi.org/10.1287/mnsc.9.1.16
  • C. Derman, A. F. Veinott Jr., A solution to a countable system of equations arising in Markovian decision processes, Ann. Math. Statist. 38 (1967) 582–584 (cited as [43] in the survey, https://doi.org/10.1137/0331018).
  • S. M. Ross, Non-discounted denumerable Markovian decision models, Ann. Math. Statist. 39 (1968) 412–423 (cited as [147] in the survey, https://doi.org/10.1137/0331018).
  • S. M. Ross, Introduction to Stochastic Dynamic Programming, Academic Press, 1983 (cited as [150] in the survey, https://doi.org/10.1137/0331018).
10 thms1 active userReviewed
CombinatoricsTheoretical Computer Science·Captain: mikedeng1

On the Approximability of Single-Machine Scheduling with Precedence Constraints 6: The Optimal Value of S_G Lies Between n² − an²(ln 1/a + 2) and n² − an²Research Paper

Motivation

The problem 1 ∣ prec ∣ ∑wjCj1\,|\,\mathrm{prec}\,|\,\sum w_jC_j1∣prec∣∑wj​Cj​ asks for a single-machine sequence of jobs, respecting precedence constraints, that minimizes the weighted sum of completion times. It has been known to be strongly NP-hard since Lawler (1978) and Lenstra and Rinnooy Kan (1978), several different 2-approximation algorithms are known, and closing the approximability gap is listed by Schuurman and Woeginger (1999) as one of ten outstanding open problems in scheduling theory. Ambühl, Mastrolilli, Mutsanas and Svensson (Math. Oper. Res. 36(4), 2011) give the first inapproximability result for this problem: under a widely believed complexity assumption it has no polynomial-time approximation scheme (PTAS). The bridge to that result is a quantitative link, Lemma 9.1, between the optimal value of a special bipartite scheduling instance and the maximum edge biclique of a bipartite graph, a problem whose hardness of approximation was established by Ambühl, Mastrolilli and Svensson (FOCS 2007). This mission formalizes that link.

Setting

A schedule of a finite job set is a sequence σ\sigmaσ listing every job once; the machine processes the jobs in that order from time 000 without idle time or pre-emption. Job jjj has a processing time pjp_jpj​ and a weight wjw_jwj​; its completion time CjC_jCj​ is the sum of the processing times of the jobs up to and including jjj, and the value of σ\sigmaσ is val(σ)=∑jwjCj\mathrm{val}(\sigma)=\sum_j w_jC_jval(σ)=∑j​wj​Cj​. Precedence constraints are a relation PPP on jobs: (i,j)∈P(i,j)\in P(i,j)∈P with i≠ji\ne ji=j means job iii must be completed before job jjj starts. A schedule respecting all of them is feasible, and a feasible schedule σ∗\sigma^*σ∗ of least value is optimal.

Let G=(U,V,E)G=(U,V,E)G=(U,V,E) be an nnn-by-nnn bipartite graph: ∣U∣=∣V∣=n|U|=|V|=n∣U∣=∣V∣=n and E⊆U×VE\subseteq U\times VE⊆U×V. An edge biclique is a pair A⊆UA\subseteq UA⊆U, B⊆VB\subseteq VB⊆V with A×B⊆EA\times B\subseteq EA×B⊆E, of value ∣A∣⋅∣B∣|A|\cdot|B|∣A∣⋅∣B∣; the maximum edge biclique problem (Definition 9.1) asks for one of largest value. The bipartite scheduling instance SGS_GSG​ has jobs U∪VU\cup VU∪V and precedence constraints

P=(U×V)∖E,P=(U\times V)\setminus E,P=(U×V)∖E,

so u∈Uu\in Uu∈U must precede v∈Vv\in Vv∈V exactly when (u,v)(u,v)(u,v) is not an edge. Jobs of UUU have p=1p=1p=1, w=0w=0w=0; jobs of VVV have p=0p=0p=0, w=1w=1w=1. Thus val(σ)=∑v∈VCv\mathrm{val}(\sigma)=\sum_{v\in V}C_vval(σ)=∑v∈V​Cv​, where CvC_vCv​ is the number of UUU-jobs scheduled before vvv. For i≥1i\ge1i≥1, σ(i)\sigma(i)σ(i) denotes the number of VVV-jobs scheduled before iii jobs of UUU have been scheduled.

In the Lean development these are weightedCompletion, IsOptimalSchedule, IsEdgeBiclique, maxBicliqueValue, precSG, procSG, weightSG, valSG, IsOptimalSG and vBefore in the namespace SingleMachinePrec.Biclique.

Formalization targets

Goal: Lemma 9.1 (p. 666)

If a maximum edge biclique of GGG has value an2an^2an2 with a∈(0,1]a\in(0,1]a∈(0,1], then SGS_GSG​ has an optimal schedule and every optimal schedule σ∗\sigma^*σ∗ satisfies

n2−an2(ln⁡1a+2)≤val(σ∗)≤n2−an2.n^2-an^2\Bigl(\ln\frac1a+2\Bigr)\le\mathrm{val}(\sigma^*)\le n^2-an^2 .n2−an2(lna1​+2)≤val(σ∗)≤n2−an2.

Milestones (proof of Lemma 9.1, §9, p. 666)

  1. For every edge biclique (A,B)(A,B)(A,B), a schedule in the block order U∖A→B→A→V∖BU\setminus A\to B\to A\to V\setminus BU∖A→B→A→V∖B exists, and every such schedule is feasible with
val(σ)=n2−∣A∣⋅∣B∣.\mathrm{val}(\sigma)=n^2-|A|\cdot|B| .val(σ)=n2−∣A∣⋅∣B∣.
  1. For every schedule, σ(n+1)=n\sigma(n+1)=nσ(n+1)=n and
val(σ)=∑i=1n(σ(i+1)−σ(i))i=n2−∑i=1nσ(i).\mathrm{val}(\sigma)=\sum_{i=1}^n\bigl(\sigma(i+1)-\sigma(i)\bigr)i=n^2-\sum_{i=1}^n\sigma(i).val(σ)=i=1∑n​(σ(i+1)−σ(i))i=n2−i=1∑n​σ(i).
  1. For every feasible schedule and i=1,…,ni=1,\dots,ni=1,…,n,
σ(i)(n−i+1)≤an2,σ(i)≤n.\sigma(i)(n-i+1)\le an^2,\qquad \sigma(i)\le n .σ(i)(n−i+1)≤an2,σ(i)≤n.

Significance

Lemma 9.1 shows that the optimal value of SGS_GSG​ determines the maximum edge biclique of GGG up to a factor of order ln⁡(1/a)\ln(1/a)ln(1/a) in the "area above the work line" n2−val(σ∗)n^2-\mathrm{val}(\sigma^*)n2−val(σ∗). Combined with the hardness of approximating maximum edge biclique (Theorem 9.1, cited from Ambühl, Mastrolilli and Svensson 2007) it yields Theorem 9.2: 1 ∣ prec ∣ ∑wjCj1\,|\,\mathrm{prec}\,|\,\sum w_jC_j1∣prec∣∑wj​Cj​ has no PTAS unless SAT can be decided by a probabilistic algorithm in time 2Nϵ2^{N^\epsilon}2Nϵ for every ϵ>0\epsilon>0ϵ>0. It also makes precise the two-dimensional Gantt chart picture of Eastman, Even and Isaacs (1964) and of Goemans and Williamson (2000), in which every point on the work line of a schedule defines an edge biclique.

The lemma is proved in the paper; it is not formalized anywhere to our knowledge. A formal proof certifies the combinatorial core of the no-PTAS result independently of the complexity-theoretic layer, and its definitions (the bipartite instance SGS_GSG​, edge bicliques, the profile σ(i)\sigma(i)σ(i)) are reusable for the gap inequality behind Theorem 9.2.

Difficulty

The upper bound is a direct computation on one explicit schedule. The lower bound is a statement about every feasible schedule, of which there are exponentially many, and it must hold with the explicit constant 222 and the factor ln⁡(1/a)\ln(1/a)ln(1/a) for every a∈(0,1]a\in(0,1]a∈(0,1]. The printed argument splits the sum at i=(1−a)ni=(1-a)ni=(1−a)n and uses ⌊an⌋\lfloor an\rfloor⌊an⌋, treating ananan as an integer; for general aaa (for example n=3n=3n=3, value 222, an=2/3an=2/3an=2/3) the split point is not an integer, so the printed estimate does not apply verbatim and the constant 222 has to be re-checked for non-integral ananan. On the formal side, the value identity requires relating completion times in a list to counting UUU-jobs before each VVV-job, with ties among zero-length jobs.

Formalization scope

Jobs are the disjoint union U ⊕ V of two finite types with Fintype.card U = Fintype.card V = n; EEE is a relation U → V → Prop. A schedule is a duplicate-free list containing every job; feasibility is the published LawlerPrec.MinMax.IsFeasible and completion times are the published MooreLateJobs.Shared.completionTime (time 000 start, no idle time). Processing times and weights are reals, here in {0,1}\{0,1\}{0,1}. The maximum edge biclique value is the maximum of ∣A∣⋅∣B∣|A|\cdot|B|∣A∣⋅∣B∣ over all edge bicliques, the empty ones included, so the hypothesis a>0a>0a>0 means E≠∅E\ne\emptysetE=∅. The logarithm is natural (Real.log).

Conventions and readings committed to:

  • The goal is stated for every optimal schedule, and the existence of an optimal schedule is a separate conclusion, so the bounds cannot hold vacuously. Proving the bounds for one particular schedule, or for an optimal value defined as an infimum that could be a junk default, would not be this lemma.
  • No integrality hypothesis on ananan is added.
  • Milestones 2 and 3 are stated for every schedule (respectively every feasible schedule), not only for σ∗\sigma^*σ∗; milestone 1 states the value of the block-order schedule as an equality, where the paper writes "≤⋯=\le\cdots=≤⋯=".
  • The paper's P=(U×V)∖EP=(U\times V)\setminus EP=(U×V)∖E is irreflexive; feasibility only constrains distinct jobs, so it agrees with the reflexive partial order of §1.

Not formalized: Theorem 9.1 (cited hardness of maximum edge biclique) and Theorem 9.2 (no PTAS under a complexity assumption); no polynomial-time or complexity-theoretic statement appears in the mission. Contributions welcome: proofs of the three milestones and of the goal; Mathlib's bounds on harmonic numbers (Mathlib/NumberTheory/Harmonic/Bounds.lean) are the relevant library.

Selected references

  • C. Ambühl, M. Mastrolilli, N. Mutsanas, O. Svensson, On the Approximability of Single-Machine Scheduling with Precedence Constraints, Mathematics of Operations Research 36(4):653–669, 2011. https://doi.org/10.1287/moor.1110.0512
  • C. Ambühl, M. Mastrolilli, O. Svensson, Inapproximability results for sparsest cut, optimal linear arrangement, and precedence constraint scheduling, Proc. 48th IEEE FOCS, 329–337, 2007 (reference [4] of the paper).
  • W. L. Eastman, S. Even, I. M. Isaacs, Bounds for the optimal scheduling of n jobs on m processors, Management Science 11(2):268–279, 1964 (reference [11]).
  • M. X. Goemans, D. P. Williamson, Two-dimensional Gantt charts and a scheduling algorithm of Lawler, SIAM J. Discrete Math. 13(3):281–294, 2000 (reference [15]).
  • P. Schuurman, G. J. Woeginger, Polynomial time approximation algorithms for machine scheduling: ten open problems, J. Scheduling 2(5):203–213, 1999 (reference [36]).
8 thms1 active userReviewed
Dynamic ProgrammingMarkov Chain·Captain: mikedeng1

Discrete-Time Controlled Markov Processes with Average Cost Criterion: A Survey 3: A Uniform Lower Bound on the Probability of Moving to State 0 Reduces Average Cost to Discounted CostResearch Paper

Why average cost is a control problem

A controller acting over an indefinite horizon must decide whether a lower cost today is worth a higher cost later. Average cost measures the expected expenditure per stage as the horizon grows. It is appropriate when operation has no natural terminal date, but its limiting definition makes it difficult to compute an optimal policy directly. Discounted cost assigns less weight to distant stages and has a more direct optimality equation. Arapostathis, Borkar, Fernández-Gaucherand, Ghosh and Marcus survey these criteria for controlled Markov processes and state a condition under which solving one discounted problem yields a solution to an average-cost problem (Arapostathis et al., 1993, §5.1).

The condition is a common lower bound on the one-step probability of moving to a distinguished state. Ross's reduction, reported as Theorem 5.6 of the survey, uses that bound to define a new transition law and a specific discount factor. The survey's theorem states the reduction informally; its proof specifies the transformed law, the optimality equation and the resulting average-cost policy. Those claims are the targets of this mission (Arapostathis et al., 1993, pp. 303–304).

Controlled process and criteria

The state space is S={0,1,2,…}S=\{0,1,2,\ldots\}S={0,1,2,…}. In state iii, the controller may choose an action aaa from a nonempty compact set U(i)U(i)U(i) in a metric action space. Choosing aaa incurs the one-stage cost c(i,a)≥0c(i,a)\ge0c(i,a)≥0 and moves the state to jjj with probability P(j∣i,a)P(j\mid i,a)P(j∣i,a). The cost is measurable, and for each fixed i,ji,ji,j its value and P(j∣i,a)P(j\mid i,a)P(j∣i,a) vary continuously with aaa on U(i)U(i)U(i). Section 5 imposes these state and continuity conventions; §5.1 additionally assumes the costs are bounded (Arapostathis et al., 1993, pp. 284–288, 299, 301).

An admissible policy π\piπ chooses a probability law for the next action from the entire observed history, and assigns probability one to admissible actions. A stationary deterministic policy is a map fff with f(i)∈U(i)f(i)\in U(i)f(i)∈U(i); it always takes action f(i)f(i)f(i) in state iii. These are distinct classes. Let JN(i,π)J_N(i,\pi)JN​(i,π) be expected cost over the first NNN stages from iii, and let Jβ(i,π)J_\beta(i,\pi)Jβ​(i,π) be expected cost when stage ttt is weighted by βt\beta^tβt, where 0<β<10<\beta<10<β<1. The average-cost criterion is J(i,π)=lim sup⁡N→∞JN(i,π)/NJ(i,\pi)=\limsup_{N\to\infty}J_N(i,\pi)/NJ(i,π)=limsupN→∞​JN​(i,π)/N. The optimal values J∗(i)J^*(i)J∗(i) and Jβ∗(i)J_\beta^*(i)Jβ∗​(i) are infima over all admissible policies, including randomized and history-dependent ones (Arapostathis et al., 1993, pp. 285–287).

The average cost optimality equation, or ACOE, asks for a scalar ρ\rhoρ and a real function hhh such that, for each state iii,

ρ+h(i)=min⁡a∈U(i){c(i,a)+∑j∈SP(j∣i,a)h(j)}.\rho+h(i)=\min_{a\in U(i)}\left\{c(i,a)+\sum_{j\in S}P(j\mid i,a)h(j)\right\}.ρ+h(i)=a∈U(i)min​⎩⎨⎧​c(i,a)+j∈S∑​P(j∣i,a)h(j)⎭⎬⎫​.

The minimum is attained. The survey's verification theorem identifies ρ\rhoρ with the optimal average cost when the terminal contribution of h(Xt)h(X_t)h(Xt​) vanishes after division by ttt (Arapostathis et al., 1993, p. 299, (5.1), Theorem 5.1).

Formalization targets

The reduction

Assume that P(0∣i,a)≥αP(0\mid i,a)\ge\alphaP(0∣i,a)≥α on every admissible state-action pair for a single constant 0<α<10<\alpha<10<α<1. The transformed process M~\widetilde MM has the same admissible actions and costs and has transition probabilities

P~(j∣i,a)=P(j∣i,a)−α1{j=0}1−α.\widetilde P(j\mid i,a)=\frac{P(j\mid i,a)-\alpha\mathbf1_{\{j=0\}}}{1-\alpha}.P(j∣i,a)=1−αP(j∣i,a)−α1{j=0}​​.

Write J~1−α∗\widetilde J^*_{1-\alpha}J1−α∗​ for its discounted value at discount factor 1−α1-\alpha1−α. The goal states that this value is finite, a stationary deterministic discounted-optimal policy exists, and every such policy is average-cost optimal for the original process. It also identifies a constant optimal average cost for every initial state:

J∗(i)=αJ~1−α∗(0),i∈S.J^*(i)=\alpha\widetilde J^*_{1-\alpha}(0),\qquad i\in S.J∗(i)=αJ1−α∗​(0),i∈S.

The goal does not assume the average optimality that it asserts. It requires the transformed law to satisfy the displayed formula at every admissible state-action pair (Arapostathis et al., 1993, Theorem 5.6 and proof, pp. 303–304).

Supporting targets

The milestones establish that the transformed law gives a controlled Markov process, that the discounted problem has an optimal stationary deterministic policy, and that the transformed discounted value satisfies the original ACOE with ρ=αJ~1−α∗(0)\rho=\alpha\widetilde J^*_{1-\alpha}(0)ρ=αJ1−α∗​(0). The final milestone is the ACOE verification theorem needed to identify the average cost. The source gives the first two discounted claims through Theorem 2.1 and writes out the transformed equation in the proof of Theorem 5.6 (Arapostathis et al., 1993, pp. 289, 299, 304).

What the result supplies

The theorem replaces an average-cost optimization problem by one discounted problem with a prescribed discount factor and transition law. It yields a stationary deterministic policy that is optimal even when compared with every history-dependent randomized policy, and it shows that the optimal average cost is independent of the initial state. Without the common return probability, neither this particular law nor this fixed discount factor follows from the survey's argument (Arapostathis et al., 1993, Theorem 5.6).

The survey reports this as a known result of Ross rather than an open conjecture. The formalization work is to give the path measures, value functions, transformed process and verification result machine-checkable meanings. The Lean declarations here are open theorem statements awaiting proofs; compiling a statement with sorry does not establish the mathematical theorem. The policy and cost definitions can also support the other countable-state missions drawn from §5.

Difficulty

The transformed probabilities have to form a measurable stochastic kernel, preserve the action continuity assumptions and yield a controlled process with the original feasible actions and costs. The discounted value must be finite and uniformly bounded before its real form can enter the ACOE. A simple comparison of numerical optimal values is insufficient: a policy chosen in the transformed model must be shown optimal for the original model against the full policy class. The verification theorem also requires control of the terminal expectation of h(Xt)h(X_t)h(Xt​) for arbitrary admissible policies, rather than only the stationary policies named in its printed condition (Arapostathis et al., 1993, pp. 299–300, 304).

Formalization scope

Lean uses N\mathbb NN for the countable state space and Mathlib probability kernels for the transition and randomized decision rules. Strategic path measures are constructed from those kernels. Nonnegative expected costs and their infima live in [0,∞][0,\infty][0,∞], so an unbounded integral cannot silently become a finite real value. The transformed discounted value is converted to a real only in conclusions that also assert its finiteness. The ACOE uses real sums with explicit summability and an attained minimum. The model requires a Borel metric action space, nonempty compact action sets, measurable costs, and coordinatewise action continuity of the transition probabilities.

The source's §5.1 bounded-cost assumption is explicit in the goal and its discounted milestones. The displayed transformation needs α<1\alpha<1α<1; the paper's theorem sentence gives only α>0\alpha>0α>0, so the case α=1\alpha=1α=1 is excluded from this version. The paper prints (5.2) for stationary deterministic policies, but the proof uses it for all admissible policies to compare with J∗J^*J∗; the verification milestone takes the stronger, proof-supported hypothesis. Expectations in that condition are explicitly integrable. The converse part of Theorem 5.1 is not used and is outside this mission.

The transformed process is represented by another CMP constrained to have exactly the source's action sets, admissible costs and transition formula. A milestone poses its existence. This representation keeps the definition layer free of an unproved stochastic-kernel construction. The discount-optimal policy conclusion ranges over every stationary deterministic policy attaining the transformed discounted value, while the values themselves take infima over all admissible policies. Definitions of admissible path laws and the verification theorem are reusable contributions; proofs of the transformed kernel, stationary discounted existence and ACOE identity are welcome.

Selected references

  • A. Arapostathis, V. S. Borkar, E. Fernández-Gaucherand, M. K. Ghosh and S. I. Marcus, Discrete-time controlled Markov processes with average cost criterion: a survey, SIAM Journal on Control and Optimization 31(2), 282–344, 1993. DOI: 10.1137/0331018.
7 thms1 active userReviewed
CombinatoricsGraph TheoryTheoretical Computer Science·Captain: mikedeng1

On the Approximability of Single-Machine Scheduling with Precedence Constraints 4: Vertex Cover in Connected Graphs of Degree ≤ 3 Reduces to Weighted Vertex Cover for Interval-Order InstancesResearch Paper

Motivation

In the single-machine scheduling problem 1∣prec∣∑wjCj1|\mathrm{prec}|\sum w_jC_j1∣prec∣∑wj​Cj​, a set NNN of nnn jobs, each with a processing time pj≥0p_j\ge 0pj​≥0 and a weight wj≥0w_j\ge 0wj​≥0, is processed on one machine without interruption, subject to precedence constraints given by a partial order PPP on NNN. The aim is to minimize the weighted sum of completion times ∑jwjCj\sum_j w_jC_j∑j​wj​Cj​. The problem is strongly NP-hard for general precedence constraints (Lawler 1978; Lenstra and Rinnooy Kan 1978), and its approximability was a recurring open question in scheduling theory (Schuurman and Woeginger 1999).

A line of work by Chudak and Hochbaum, Correa and Schulz, and Ambühl and Mastrolilli showed that the problem is a special case of minimum weighted vertex cover in a graph GPSG^S_PGPS​ built from the instance. Many problems on partial orders become polynomial when the order is an interval order, so it is natural to ask whether this one does too. Section 7 of Ambühl, Mastrolilli, Mutsanas and Svensson (Math. Oper. Res. 2011) answers no: the problem stays NP-hard on interval orders. The proof is a reduction from vertex cover in connected graphs of maximum degree 3. This mission formalizes the correctness of that reduction.

Setting

A poset P=(N,P)P=(N,P)P=(N,P) is read as a reflexive relation: (x,y)∈P(x,y)\in P(x,y)∈P means x≤yx\le yx≤y. Jobs x,yx,yx,y are incomparable if neither (x,y)(x,y)(x,y) nor (y,x)(y,x)(y,x) is in PPP, and inc⁡(P)\operatorname{inc}(P)inc(P) is the set of ordered incomparable pairs. The vertex cover graph GPSG^S_PGPS​ has one node (i,j)(i,j)(i,j) for each incomparable pair, weighted piwjp_iw_jpi​wj​. Two nodes (i,j)(i,j)(i,j) and (k,ℓ)(k,\ell)(k,ℓ) are adjacent if j=kj=kj=k and i=ℓi=\elli=ℓ, or j=kj=kj=k and (i,ℓ)∈P(i,\ell)\in P(i,ℓ)∈P, or (i,ℓ)∈P(i,\ell)\in P(i,ℓ)∈P and (k,j)∈P(k,j)\in P(k,j)∈P. Write w(CI)w(C_I)w(CI​) for the minimum weight of a vertex cover of GISG^S_IGIS​.

A poset is an interval order if each element xxx can be assigned a closed real interval [ax,bx][a_x,b_x][ax​,bx​] such that x<yx<yx<y if and only if bx<ayb_x<a_ybx​<ay​.

The reduction starts from a graph G=(V,E)G=(V,E)G=(V,E) with vertices v1,…,vNv_1,\dots,v_Nv1​,…,vN​ and a spanning tree T=(V,ET)T=(V,E_T)T=(V,ET​) rooted at v1v_1v1​, numbered so that each parent comes before its children. The paper uses a breadth-first search tree.

  • Stage 1. The graph G′G'G′ is built from TTT. Each viv_ivi​ gets a pendant path vi−u1i−u2iv_i - u^i_1 - u^i_2vi​−u1i​−u2i​. Each non-tree edge {vi,vj}∈E∖ET\{v_i,v_j\}\in E\setminus E_T{vi​,vj​}∈E∖ET​ with i<ji<ji<j gets the path vi−e1ij−e2ij−u2jv_i - e^{ij}_1 - e^{ij}_2 - u^j_2vi​−e1ij​−e2ij​−u2j​. The non-tree edges themselves are not edges of G′G'G′.
  • Stage 2. The scheduling instance SSS has jobs s0s_0s0​, s1,…,sNs_1,\dots,s_Ns1​,…,sN​, m1,…,mNm_1,\dots,m_Nm1​,…,mN​, e1,…,eNe_1,\dots,e_Ne1​,…,eN​, and bijb_{ij}bij​ for each non-tree edge. Their intervals, processing times and weights are given in a table on p. 662. For example, sjs_jsj​ has interval [i,j][i,j][i,j], processing time 1/kj1/k^j1/kj and weight kik^iki, where viv_ivi​ is the parent of vjv_jvj​. The precedence constraints III are the interval order of these intervals. With nnn the number of jobs, the parameter is k=n2+1k=n^2+1k=n2+1.
  • The set DDD. It is {(s0,s1)}∪{(si,sj):vi parent of vj}∪{(si,mi),(mi,ei)}∪{(si,bij),(bij,mj)}\{(s_0,s_1)\}\cup\{(s_i,s_j): v_i \text{ parent of } v_j\}\cup\{(s_i,m_i),(m_i,e_i)\}\cup\{(s_i,b_{ij}),(b_{ij},m_j)\}{(s0​,s1​)}∪{(si​,sj​):vi​ parent of vj​}∪{(si​,mi​),(mi​,ei​)}∪{(si​,bij​),(bij​,mj​)}. The graph GI′G'_IGI′​ is the subgraph of GISG^S_IGIS​ induced by DDD.

Formalization targets

Goal: Theorem 7.1 (p. 661)

For every connected graph GGG of maximum degree at most 333, every parent-first spanning tree TTT and every m∈Nm\in\mathbb Nm∈N, the precedence constraints III of SSS form an interval order, and

G has a vertex cover of size≤m  ⟺  ⌊w(CI)⌋≤m+∣V∣+∣E∖ET∣.G \text{ has a vertex cover of size} \le m \iff \lfloor w(C_I)\rfloor \le m + |V| + |E\setminus E_T|.G has a vertex cover of size≤m⟺⌊w(CI​)⌋≤m+∣V∣+∣E∖ET​∣.

Milestones, in the order the proof uses them

  • Claim 1 (p. 662): τ(G′)=τ(G)+∣V∣+∣E∖ET∣\tau(G') = \tau(G)+|V|+|E\setminus E_T|τ(G′)=τ(G)+∣V∣+∣E∖ET​∣, where τ\tauτ is the vertex cover number.
  • Remark 7.1 (p. 662): for jobs with intervals [a,b][a,b][a,b] and [c,d][c,d][c,d] and a≤da\le da≤d, pi≤1/k⌈b⌉p_i\le 1/k^{\lceil b\rceil}pi​≤1/k⌈b⌉ and wj≤k⌈c⌉w_j\le k^{\lceil c\rceil}wj​≤k⌈c⌉. On incomparable pairs piwj∈{1}∪[0,1/k]p_iw_j\in\{1\}\cup[0,1/k]pi​wj​∈{1}∪[0,1/k]. Moreover, piwj≥kp_iw_j\ge kpi​wj​≥k forces b<cb<cb<c, and piwj=1p_iw_j=1pi​wj​=1 forces ⌈b⌉=⌈c⌉\lceil b\rceil=\lceil c\rceil⌈b⌉=⌈c⌉.
  • Claim 2 (p. 663): an incomparable pair (i,j)(i,j)(i,j) has piwj=1p_iw_j=1pi​wj​=1 if it is in DDD, and piwj≤1/kp_iw_j\le 1/kpi​wj​≤1/k otherwise.
  • Claim 3 (p. 663): GI′≅G′G'_I\cong G'GI′​≅G′.
  • §7, p. 664: with k=n2+1k=n^2+1k=n2+1, ∑(i,j)∈inc⁡(I)∖Dpiwj<1\sum_{(i,j)\in\operatorname{inc}(I)\setminus D}p_iw_j<1∑(i,j)∈inc(I)∖D​pi​wj​<1, and hence w(CI′)=⌊w(CI)⌋w(C'_I)=\lfloor w(C_I)\rfloorw(CI′​)=⌊w(CI​)⌋.

Significance

The result. Interval orders are a standard tractable class: several scheduling and order-theoretic problems that are hard in general become polynomial on them (Papadimitriou and Yannakakis 1979). Theorem 7.1 puts 1∣prec∣∑wjCj1|\mathrm{prec}|\sum w_jC_j1∣prec∣∑wj​Cj​ outside this pattern. Section 6 of the same paper shows that the problem nonetheless has a 3/23/23/2-approximation on interval orders, so hardness and approximability are separated on this class. The paper also remarks that the proof makes weighted vertex cover NP-hard to approximate within some factor r>1r>1r>1 on the graphs GISG^S_IGIS​ arising from interval orders.

Formalizing it. The theorem is proved in the paper. Nothing in this mission is open, and none of it has been machine-checked before. The work splits into the following parts:

  • a gadget argument on unweighted vertex covers (Claim 1, after Alimonti and Kann);
  • an exact case analysis of incomparable pairs in a concrete interval order (Remark 7.1, Claim 2);
  • a graph isomorphism (Claim 3);
  • a rounding argument that links weighted and unweighted optima.

The definitions of GPSG^S_PGPS​ and of minimum-weight vertex covers are shared with the other missions of this series.

Difficulty

The construction is explicit, and each step is elementary. The work is in the bookkeeping. Claim 2 requires classifying every incomparable pair of jobs, including pairs of different kinds such as (bij,sℓ)(b_{ij}, s_\ell)(bij​,sℓ​), by comparing ceilings of interval endpoints. Half-integer endpoints (mim_imi​, bijb_{ij}bij​) are exactly what separates weight-one pairs from comparable ones. Claim 3 requires checking adjacency in GISG^S_IGIS​ for all pairs of nodes of DDD in both directions. The paper writes out two cases in each direction and calls the rest similar.

Claim 1 has a direction that is not simply local. A vertex cover of G′G'G′ that misses both endpoints of a non-tree edge has to be repaired by swapping gadget vertices, and the repair must be repeated without increasing the size.

A natural first idea is to treat the light nodes (weight at most 1/k1/k1/k) as negligible one at a time. This does not suffice: the argument needs their total weight to stay below 111, which is what forces kkk to grow with n2n^2n2.

Formalization scope

  • Graph and tree. GGG is a SimpleGraph (Fin N); vertex vi+1v_{i+1}vi+1​ is i, and the root is index 0. The tree is a TreeLayout: a parent function returning none exactly at the root, with each parent of smaller index and adjacent in GGG. The statements hold for every such layout. This is stronger than the paper's breadth-first tree, and the proof uses only "parent before child".
  • Hypotheses of the goal. Connectivity and the degree bound ((G.neighborSet v).ncard ≤ 3) are kept as in the paper. They matter only for the NP-completeness of the source problem.
  • Jobs. The jobs form an inductive type with one constructor per row of the table. Their order is a PartialOrder instance: x≤yx\le yx≤y iff x=yx=yx=y or bx<ayb_x<a_ybx​<ay​. Processing times and weights are real numbers. Section 1 of the paper asks for nonnegative integers, but the instance uses 1/kj1/k^j1/kj and the formalization follows the instance as printed.
  • Constants. The constants are explicit: k=n2+1k=n^2+1k=n2+1 with nnn the cardinality of the job type, and c=∣V∣+∣E∖ET∣c=|V|+|E\setminus E_T|c=∣V∣+∣E∖ET​∣. Remark 7.1 and Claim 2 are stated for every real k>1k>1k>1.
  • Optimum values. w(CI)w(C_I)w(CI​) is a minimum over the finite family of vertex covers. Unweighted cover numbers are Mathlib's SimpleGraph.vertexCoverNum. The floor is Nat.floor, which agrees with the integer floor because w(CI)≥0w(C_I)\ge 0w(CI​)≥0.
  • Not formalized. The goal's wording ("NP-hard") is not formalized. Neither are the NP-completeness of degree-3 vertex cover (Garey, Johnson and Stockmeyer), the polynomial size of the construction, or Theorem 2.1 (cited), which turns a vertex cover of GISG^S_IGIS​ into a schedule. What is stated is the correctness of the reduction: the instance has interval-order constraints, and its optimum decides the vertex cover question.
  • No trivialization. The instance SSS is built from GGG and TTT exactly as in the table. The goal quantifies over all graphs and layouts, never over an instance SSS assumed to have the properties.
  • Contributions. Contributions are welcome on any milestone. Claims 1 and 3 are independent of the weights, and Claim 2 is independent of the graph theory.

Selected references

  • C. Ambühl, M. Mastrolilli, N. Mutsanas, O. Svensson, On the Approximability of Single-Machine Scheduling with Precedence Constraints, Math. Oper. Res. 36(4):653–669, 2011. https://doi.org/10.1287/moor.1110.0512
  • C. Ambühl, M. Mastrolilli, Single machine precedence constrained scheduling is a vertex cover problem, Algorithmica 53(4):488–503, 2009. https://doi.org/10.1007/s00453-008-9251-1
  • J. R. Correa, A. S. Schulz, Single machine scheduling with precedence constraints, Math. Oper. Res. 30(4):1005–1021, 2005. https://doi.org/10.1287/moor.1050.0158
  • P. Alimonti, V. Kann, Some APX-completeness results for cubic graphs, Theoret. Comput. Sci. 237(1–2):123–134, 2000. https://doi.org/10.1016/S0304-3975(98)00158-3
  • M. R. Garey, D. S. Johnson, L. Stockmeyer, Some simplified NP-complete graph problems, Theoret. Comput. Sci. 1(3):237–267, 1976. https://doi.org/10.1016/0304-3975(76)90059-1
  • C. H. Papadimitriou, M. Yannakakis, Scheduling interval-ordered tasks, SIAM J. Comput. 8(3):405–409, 1979. https://doi.org/10.1137/0208031
10 thms1 active userReviewed
Graph TheoryTheoretical Computer Science·Captain: mikedeng1

On the Approximability of Single-Machine Scheduling with Precedence Constraints 5: An r-Approximate Vertex Cover of the Variable-Cost Graph Yields an (r + ε)-Approximate Vertex Cover of GResearch Paper

Why the variable cost matters

The problem 1∣prec∣∑wjCj1|\mathrm{prec}|\sum w_jC_j1∣prec∣∑wj​Cj​ asks for an order in which to process jobs on one machine, respecting precedence constraints, so as to minimize the weighted sum of completion times. It is strongly NP-hard, and for decades the best approximation ratio known has been 222, achieved by several unrelated algorithms (LP relaxations, Sidney decompositions, primal–dual methods).

Correa and Schulz (2005) and Ambühl and Mastrolilli (2009) showed that the problem is a special case of weighted vertex cover: its objective splits into a fixed cost, the same for every feasible solution, and a variable cost, which equals the weight of a vertex cover in an auxiliary graph GPSG^S_{\mathbf P}GPS​. Approximating vertex cover in GPSG^S_{\mathbf P}GPS​ within a factor α\alphaα therefore approximates the scheduling problem within α\alphaα. Uhan observed that the classical 2-approximations owe their guarantee to the fixed cost and can be arbitrarily bad on the variable cost alone.

Section 8 of Ambühl, Mastrolilli, Mutsanas and Svensson, Math. Oper. Res. 36(4) (2011) (DOI), proves the converse: approximating the variable cost is as hard as approximating vertex cover itself. A better-than-2 algorithm for 1∣prec∣∑wjCj1|\mathrm{prec}|\sum w_jC_j1∣prec∣∑wj​Cj​ must therefore either exploit the fixed cost or improve on the best known approximation for vertex cover, a long-standing open question.

Setting

Scheduling instance. A finite set NNN of jobs, a partial order P=(N,P)\mathbf P = (N,P)P=(N,P) (reflexive; (i,j)∈P(i,j) \in P(i,j)∈P, i≠ji \ne ji=j, means iii precedes jjj), processing times pj≥0p_j \ge 0pj​≥0 and weights wj≥0w_j \ge 0wj​≥0.

Incomparable pairs. Jobs x,yx,yx,y are incomparable, x∥yx \parallel yx∥y, if neither (x,y)(x,y)(x,y) nor (y,x)(y,x)(y,x) lies in PPP. The set inc⁡(P)\operatorname{inc}(\mathbf P)inc(P) consists of the ordered pairs (x,y)(x,y)(x,y) with x∥yx \parallel yx∥y.

The vertex cover graph GPSG^S_{\mathbf P}GPS​. One node per incomparable pair (i,j)(i,j)(i,j), of weight w(i,j)=piwjw_{(i,j)} = p_i w_jw(i,j)​=pi​wj​. Distinct nodes (i,j)(i,j)(i,j) and (k,ℓ)(k,\ell)(k,ℓ) are adjacent when, in one of the two orders, j=kj=kj=k and i=ℓi=\elli=ℓ, or j=kj=kj=k and (i,ℓ)∈P(i,\ell)\in P(i,ℓ)∈P, or (i,ℓ),(k,j)∈P(i,\ell),(k,j) \in P(i,ℓ),(k,j)∈P. For a set CCC of nodes, w(C)=∑u∈Cwuw(C) = \sum_{u\in C} w_uw(C)=∑u∈C​wu​; for a vertex cover CCC this is the variable cost, and τw(GPS)\tau_w(G^S_{\mathbf P})τw​(GPS​) is its minimum over all vertex covers.

The instance S(G,k)S(G,k)S(G,k). Given a graph G=(V,E)G=(V,E)G=(V,E) with V={v1,…,vn}V = \{v_1,\dots,v_n\}V={v1​,…,vn​} and k>0k > 0k>0, the instance has jobs vi′v'_ivi′​ (processing time k−ik^{-i}k−i, weight 000) and vi′′v''_ivi′′​ (processing time 000, weight kik^{i}ki), and precedence constraints vi′<vj′′v'_i < v''_jvi′​<vj′′​ and vj′<vi′′v'_j < v''_ivj′​<vi′′​ for each edge {vi,vj}∈E\{v_i,v_j\} \in E{vi​,vj​}∈E, plus vi′<vj′′v'_i < v''_jvi′​<vj′′​ for all i<ji<ji<j. The nodes (vi′,vi′′)(v'_i, v''_i)(vi′​,vi′′​) of GPSG^S_{\mathbf P}GPS​ have weight 111 and are called heavy; all others are light. For a set CCC of nodes, CG={vi:(vi′,vi′′)∈C}C_G = \{v_i : (v'_i,v''_i)\in C\}CG​={vi​:(vi′​,vi′′​)∈C}. The vertex cover number of GGG is τ(G)\tau(G)τ(G).

Formalization targets

Goal: Theorem 8.1

For every graph GGG on nnn vertices, every r≥1r \ge 1r≥1, ε>0\varepsilon>0ε>0 and every k≥1k \ge 1k≥1 with k>n2r/εk > n^2r/\varepsilonk>n2r/ε: if CCC is a vertex cover of GPSG^S_{\mathbf P}GPS​ for S=S(G,k)S = S(G,k)S=S(G,k) with w(C)≤r τw(GPS)w(C) \le r\,\tau_w(G^S_{\mathbf P})w(C)≤rτw​(GPS​), then CGC_GCG​ is a vertex cover of GGG,

∣CG∣≤r(τ(G)+n2k),|C_G| \le r\Bigl(\tau(G) + \frac{n^2}{k}\Bigr),∣CG​∣≤r(τ(G)+kn2​),

and, when E≠∅E \ne \emptysetE=∅,

∣CG∣≤r(1+n2k)τ(G)<(r+ε) τ(G).|C_G| \le r\Bigl(1+\frac{n^2}{k}\Bigr)\tau(G) < (r+\varepsilon)\,\tau(G).∣CG​∣≤r(1+kn2​)τ(G)<(r+ε)τ(G).

The paper words the theorem as "approximating the variable cost of 1∣prec∣∑wjCj1|\mathrm{prec}|\sum w_jC_j1∣prec∣∑wj​Cj​ is as hard as approximating vertex cover"; the statement above is the mathematical content its proof establishes.

Milestones (§8, p. 664)

  1. In GPSG^S_{\mathbf P}GPS​, every heavy node has weight 111, every light node has weight at most 1/k1/k1/k, and the light nodes have total weight at most n2/kn^2/kn2/k (for k≥1k \ge 1k≥1).
  2. Heavy nodes (vi′,vi′′)(v'_i,v''_i)(vi′​,vi′′​) and (vj′,vj′′)(v'_j,v''_j)(vj′​,vj′′​) are adjacent if and only if {vi,vj}∈E\{v_i,v_j\}\in E{vi​,vj​}∈E; for k>1k>1k>1 the subgraph induced by the weight-1 nodes is isomorphic to GGG via (vi′,vi′′)↦vi(v'_i,v''_i) \mapsto v_i(vi′​,vi′′​)↦vi​.

Significance

The result. Theorem 8.1 is one half of an equivalence: by Theorem 2.1 (Correa–Schulz, Ambühl–Mastrolilli), minimizing the variable cost is a special case of weighted vertex cover; by Theorem 8.1, it is also as hard to approximate. Any hardness of approximation for vertex cover (NP-hardness of factor 1.361.361.36 by Dinur and Safra; factor 2−δ2-\delta2−δ under the unique games conjecture by Khot and Regev) transfers to the variable cost. It also explains why the known 2-approximations must rely on the fixed cost, and it frames the later result of Bansal and Khot that the full objective is hard to approximate within 2−δ2-\delta2−δ under a variant of the unique games conjecture.

Formalizing it. The theorem is proved in the paper; no machine-checked version exists. The mission produces a checked account of the reduction: the vertex cover graph of an arbitrary precedence-constrained instance, the adjacency-poset instance built from a graph, and the quantitative transfer of approximation ratios. The definition of GPSG^S_{\mathbf P}GPS​ is shared with the other missions of this series.

Difficulty

The construction is short; the care is in the bookkeeping. One must check that the precedence relation is a partial order, determine exactly which ordered pairs are incomparable, verify that two heavy nodes are adjacent only through the third clause of the adjacency rule and only when the corresponding vertices are adjacent in GGG, and bound the weights of all remaining nodes, including the many nodes of weight 000. The transfer then compares an approximate cover of GPSG^S_{\mathbf P}GPS​ with an optimal one whose heavy part comes from an optimal cover of GGG; the additive error n2/kn^2/kn2/k must be converted into a multiplicative one, which requires τ(G)≥1\tau(G)\ge 1τ(G)≥1.

A first reading of the page suggests that GPSG^S_{\mathbf P}GPS​ has at most n2n^2n2 nodes; it does not. The pairs (vi′,vj′)(v'_i,v'_j)(vi′​,vj′​) and (vi′′,vj′′)(v''_i,v''_j)(vi′′​,vj′′​) with i≠ji\ne ji=j are incomparable nodes of weight 000, so there can be up to 4n2−2n4n^2-2n4n2−2n nodes. Only nodes of positive weight are few.

Formalization scope

  • Model. Jobs form a finite type; precedence constraints are an explicit reflexive partial-order relation P : N → N → Prop. Processing times and weights are nonnegative reals. The vertex cover graph is a SimpleGraph on the subtype of incomparable ordered pairs, using the symmetric closure of the printed adjacency rule without loops. Vertex covers are Mathlib's SimpleGraph.IsVertexCover; τ(G)\tau(G)τ(G) is Mathlib's vertexCoverNum, finite for a finite graph and converted with toNat; τw(GPS)\tau_w(G^S_{\mathbf P})τw​(GPS​) is a minimum over finite vertex covers.
  • The instance. The graph is a SimpleGraph (Fin n); i : Fin n stands for vi+1v_{i+1}vi+1​, so exponents are i+1i+1i+1 and the order i<ji<ji<j is that of Fin n. Jobs are Fin n ⊕ Fin n (v′v'v′ left, v′′v''v′′ right). The parameter kkk is in R≥0\mathbb R_{\ge 0}R≥0​.
  • Added hypotheses. k≥1k \ge 1k≥1, implicit in the page ("k>n2r/εk > n^2r/\varepsilonk>n2r/ε" does not imply it when ε\varepsilonε is large, and for k<1k<1k<1 the light nodes outweigh the heavy ones). The isomorphism with GGG is stated for k>1k > 1k>1, since at k=1k=1k=1 some light nodes also have weight 111. The multiplicative bound requires E≠∅E \ne \emptysetE=∅; the additive bound holds for every graph.
  • Not formalized. The phrases "approximation algorithm", "polynomial time" and "as hard as"; the passage from vertex covers of GPSG^S_{\mathbf P}GPS​ to schedules (Theorem 2.1, cited from Correa–Schulz and Ambühl–Mastrolilli); the fixed cost. What is stated instead is the explicit map C↦CGC \mapsto C_GC↦CG​ and the ratio it achieves, r(1+n2/k)<r+εr(1+n^2/k) < r+\varepsilonr(1+n2/k)<r+ε. The false count "at most n2n^2n2 vertices" is not stated.
  • Ruled out. The goal is not a statement about an arbitrary graph or an assumed cover of GGG: it concerns the specific instance S(G,k)S(G,k)S(G,k) and every CCC that is an rrr-approximate vertex cover of its graph, and the fact that CGC_GCG​ covers GGG is a conclusion, not a hypothesis.
  • Welcome contributions. Proofs of the two milestones and of the goal; general lemmas on GPSG^S_{\mathbf P}GPS​ (weights of vertex covers, behaviour under induced subgraphs) are reusable across the series.

Selected references

  • C. Ambühl, M. Mastrolilli, N. Mutsanas, O. Svensson, On the Approximability of Single-Machine Scheduling with Precedence Constraints, Mathematics of Operations Research 36(4):653–669, 2011. https://doi.org/10.1287/moor.1110.0512
  • J. R. Correa, A. S. Schulz, Single-Machine Scheduling with Precedence Constraints, Mathematics of Operations Research 30(4):1005–1021, 2005. https://doi.org/10.1287/moor.1050.0158
  • C. Ambühl, M. Mastrolilli, Single Machine Precedence Constrained Scheduling Is a Vertex Cover Problem, Algorithmica 53(4):488–503, 2009. https://doi.org/10.1007/s00453-008-9251-6
  • I. Dinur, S. Safra, On the Hardness of Approximating Minimum Vertex Cover, Annals of Mathematics 162(1):439–485, 2005. https://doi.org/10.4007/annals.2005.162.439
  • S. Khot, O. Regev, Vertex Cover Might Be Hard to Approximate to within 2 − ε, Journal of Computer and System Sciences 74(3):335–349, 2008. https://doi.org/10.1016/j.jcss.2007.06.019
  • N. Bansal, S. Khot, Optimal Long Code Test with One Free Bit, FOCS 2009, 453–462. https://doi.org/10.1109/FOCS.2009.23
5 thms1 active userReviewed
ProbabilityStatistics·Captain: mikedeng1

Assessing Solution Quality in Stochastic Programs: The Single-Replication Confidence Interval on the Optimality Gap Is Asymptotically ValidResearch Paper

Motivation

Most stochastic programs of practical size, such as two-stage recourse models in energy, finance or supply-chain planning, cannot be solved exactly: the expectation in the objective is a high-dimensional integral. The standard remedy is sample average approximation (SAA): replace the expectation by an average over a Monte Carlo sample and solve the resulting deterministic problem. This produces a candidate solution x^\hat xx^ but says nothing about how good it is. A decision maker needs a statistical certificate: an interval that contains the candidate's optimality gap with a prescribed probability.

Mak, Morton and Wood (Oper. Res. Lett. 24, 1999) built such a certificate from ng≥30n_g\ge 30ng​≥30 independent SAA replications, which requires solving at least 30 optimization problems. Bayraksan and Morton (preprint January 2005, published in Math. Program. 108, 2006) showed that a single replication suffices asymptotically, and gave two variants that use two replications. This mission formalizes their validity theorems.

Setting

Let ξ~\tilde\xiξ~​ be a random vector with distribution μ\muμ on a measurable space Ξ\XiΞ, let X⊆RdX\subseteq\mathbb R^dX⊆Rd be a set of decisions, and let f:Rd×Ξ→Rf:\mathbb R^d\times\Xi\to\mathbb Rf:Rd×Ξ→R be a cost. The stochastic program is

z∗=min⁡x∈XEf(x,ξ~).(SP)z^*=\min_{x\in X} Ef(x,\tilde\xi). \qquad\text{(SP)}z∗=x∈Xmin​Ef(x,ξ~​).(SP)

Its optimal set is X∗X^*X∗, and the optimality gap of a candidate x^∈X\hat x\in Xx^∈X is μx^=Ef(x^,ξ~)−z∗≥0\mu_{\hat x}=Ef(\hat x,\tilde\xi)-z^*\ge 0μx^​=Ef(x^,ξ~​)−z∗≥0. The paper assumes throughout:

  • (A1) f(⋅,ξ~)f(\cdot,\tilde\xi)f(⋅,ξ~​) is continuous on XXX, with probability one;
  • (A2) Esup⁡x∈Xf2(x,ξ~)<∞E\sup_{x\in X} f^2(x,\tilde\xi)<\inftyEsupx∈X​f2(x,ξ~​)<∞;
  • (A3) XXX is nonempty and compact.

Let ξ~1,ξ~2,…\tilde\xi^1,\tilde\xi^2,\dotsξ~​1,ξ~​2,… be i.i.d. copies of ξ~\tilde\xiξ~​, and write fˉn(x)=1n∑i=1nf(x,ξ~i)\bar f_n(x)=\frac1n\sum_{i=1}^n f(x,\tilde\xi^i)fˉ​n​(x)=n1​∑i=1n​f(x,ξ~​i). The SAA problem is zn∗=min⁡x∈Xfˉn(x)z_n^*=\min_{x\in X}\bar f_n(x)zn∗​=minx∈X​fˉ​n​(x) (SPn_nn​), with an optimal solution xn∗x_n^*xn∗​. The gap estimator is Gn(x^)=fˉn(x^)−zn∗G_n(\hat x)=\bar f_n(\hat x)-z_n^*Gn​(x^)=fˉ​n​(x^)−zn∗​ (display (2)), and the sample variance of the differences f(x^,ξ~i)−f(x,ξ~i)f(\hat x,\tilde\xi^i)-f(x,\tilde\xi^i)f(x^,ξ~​i)−f(x,ξ~​i) is

sn2(x)=1n−1∑i=1n[(f(x^,ξ~i)−f(x,ξ~i))−(fˉn(x^)−fˉn(x))]2,s_n^2(x)=\frac1{n-1}\sum_{i=1}^n\Big[\big(f(\hat x,\tilde\xi^i)-f(x,\tilde\xi^i)\big)-\big(\bar f_n(\hat x)-\bar f_n(x)\big)\Big]^2,sn2​(x)=n−11​i=1∑n​[(f(x^,ξ~​i)−f(x,ξ~​i))−(fˉ​n​(x^)−fˉ​n​(x))]2,

with population counterpart σx^2(x)=var⁡[f(x^,ξ~)−f(x,ξ~)]\sigma^2_{\hat x}(x)=\operatorname{var}[f(\hat x,\tilde\xi)-f(x,\tilde\xi)]σx^2​(x)=var[f(x^,ξ~​)−f(x,ξ~​)]. Finally zαz_\alphazα​ is defined by P(N(0,1)≤zα)=1−αP(N(0,1)\le z_\alpha)=1-\alphaP(N(0,1)≤zα​)=1−α.

The single replication procedure (SRP) solves (SPn_nn​) once and reports the one-sided interval [0, Gn(x^)+zαsn(xn∗)/n]\big[0,\ G_n(\hat x)+z_\alpha s_n(x_n^*)/\sqrt n\big][0, Gn​(x^)+zα​sn​(xn∗​)/n​] (display (5)). The I2RP takes the variance from a second, independent sample ξ~n+1,…,ξ~2n\tilde\xi^{n+1},\dots,\tilde\xi^{2n}ξ~​n+1,…,ξ~​2n and its own minimizer xn2∗x_n^{2*}xn2∗​. The A2RP runs the SRP on both halves of a sample of size 2n2n2n, averages the gaps and the variances as in (10), and scales by 2n\sqrt{2n}2n​.

Formalization targets

Goal: Theorem 2 (p. 7)

Under (A1)–(A3), for x^∈X\hat x\in Xx^∈X and 0<α<10<\alpha<10<α<1, provided α≤1/2\alpha\le1/2α≤1/2 or σx^2(xmax⁡∗)>0\sigma^2_{\hat x}(x^*_{\max})>0σx^2​(xmax∗​)>0 (see Formalization scope),

lim inf⁡n→∞P(μx^≤Gn(x^)+zαsn(xn∗)n)≥1−α.(6)\liminf_{n\to\infty}P\left(\mu_{\hat x}\le G_n(\hat x)+\frac{z_\alpha s_n(x_n^*)}{\sqrt n}\right)\ge 1-\alpha. \qquad(6)n→∞liminf​P(μx^​≤Gn​(x^)+n​zα​sn​(xn∗​)​)≥1−α.(6)

Consistency (Proposition 1, p. 6)

The milestones follow the paper's own proof:

  1. the uniform strong law sup⁡x∈X∣fˉn(x)−Ef(x,ξ~)∣→0\sup_{x\in X}|\bar f_n(x)-Ef(x,\tilde\xi)|\to 0supx∈X​∣fˉ​n​(x)−Ef(x,ξ~​)∣→0 w.p.1;
  2. (i) zn∗→z∗z_n^*\to z^*zn∗​→z∗ w.p.1;
  3. (ii) every limit point of {xn∗}\{x_n^*\}{xn∗​} lies in X∗X^*X∗ w.p.1;
  4. the uniform convergence sn2→σx^2s_n^2\to\sigma^2_{\hat x}sn2​→σx^2​ on XXX w.p.1;
  5. (iii) σx^2(xmin⁡∗)≤lim inf⁡nsn2(xn∗)≤lim sup⁡nsn2(xn∗)≤σx^2(xmax⁡∗)\sigma^2_{\hat x}(x^*_{\min})\le\liminf_n s_n^2(x_n^*)\le\limsup_n s_n^2(x_n^*)\le\sigma^2_{\hat x}(x^*_{\max})σx^2​(xmin∗​)≤liminfn​sn2​(xn∗​)≤limsupn​sn2​(xn∗​)≤σx^2​(xmax∗​) w.p.1, where xmin⁡∗x^*_{\min}xmin∗​ and xmax⁡∗x^*_{\max}xmax∗​ minimize and maximize σx^2\sigma^2_{\hat x}σx^2​ over X∗X^*X∗;
  6. the ε\varepsilonε-bound of the proof of Theorem 2: if α≤1/2\alpha\le 1/2α≤1/2 and σx^2(xmin⁡∗)>0\sigma^2_{\hat x}(x^*_{\min})>0σx^2​(xmin∗​)>0, then for 0<ε<10<\varepsilon<10<ε<1 the liminf in (6) is at least Φ((1−ε)zα)\Phi((1-\varepsilon)z_\alpha)Φ((1−ε)zα​).

Companions

Theorem 3 (p. 9) and Theorem 4 (p. 10) are the same coverage statement for the I2RP and the A2RP. Three further statements are included: the negative bias Ezn∗≤z∗Ez_n^*\le z^*Ezn∗​≤z∗ of display (1), the pathwise bound Gn(x^)≥fˉn(x^)−fˉn(x)G_n(\hat x)\ge\bar f_n(\hat x)-\bar f_n(x)Gn​(x^)≥fˉ​n​(x^)−fˉ​n​(x) for x∈Xx\in Xx∈X, and the consistency lim inf⁡nsn′2≥σx^2(xmin⁡∗)\liminf_n s_n'^2\ge\sigma^2_{\hat x}(x^*_{\min})liminfn​sn′2​≥σx^2​(xmin∗​) of the pooled variance.

Significance

Theorem 2 makes a single SAA solve enough for an asymptotically valid upper confidence bound on the optimality gap. It cuts the computational cost of the multiple-replication procedure by a factor of about thirty, and it needs no asymptotic normality of Gn(x^)G_n(\hat x)Gn​(x^), which typically fails when (SP) has several optimal solutions. The two-replication variants lessen the small-sample under-coverage of the SRP. The single- and two-replication estimators were later reused in sequential sampling procedures for SAA.

As far as is known, none of these results has been machine-checked. A complete formalization needs a uniform strong law of large numbers over a compact parameter set, which is a reusable result in its own right, together with the SAA consistency theory and a central-limit argument for a statistic that is not itself asymptotically normal.

Difficulty

The obvious route would be to show that Gn(x^)G_n(\hat x)Gn​(x^) is asymptotically normal and apply a standard confidence-interval argument. That fails: zn∗z_n^*zn∗​ is a minimum of sample averages, and when X∗X^*X∗ is not a singleton its limit law is the law of a minimum of correlated Gaussians, not a Gaussian. The paper's argument has to bound the coverage from below without that limit law. It also needs to control the sample variance at a random, non-convergent minimizer xn∗x_n^*xn∗​, which only accumulates on X∗X^*X∗. The uniform strong law (Rubinstein–Shapiro, Lemma A1) on which both consistency statements rest is not in Mathlib.

Formalization scope

The Lean development uses these conventions:

  • Decisions live in EuclideanSpace ℝ (Fin d). The paper's Rn\mathbb R^nRn is renamed Rd\mathbb R^dRd because nnn is the sample size.
  • μ\muμ is a probability measure on Ξ\XiΞ (the law of ξ~\tilde\xiξ~​), and Ef(x,ξ~)Ef(x,\tilde\xi)Ef(x,ξ~​) is the Bochner integral ∫f(x,⋅) dμ\int f(x,\cdot)\,d\mu∫f(x,⋅)dμ.
  • The sample is one infinite i.i.d. sequence ξ : ℕ → Ω → Ξ on a probability space (Ω,P)(\Omega,P)(Ω,P), 0-based: ξ~i\tilde\xi^iξ~​i is ξ (i-1). The second sample of Theorems 3–4 is ξ n, …, ξ (2n-1), exactly as printed, and the A2RP's "random" partition is this fixed one, which has the same joint law.
  • Estimators are functions of a sample path. z∗z^*z∗ and zn∗z_n^*zn∗​ are infima of images of XXX, and X∗X^*X∗ is an argmin set.
  • Probabilities are ℝ≥0∞-valued, so the liminf in (6) is genuine. Proposition 1 (iii) is stated in its equivalent ε\varepsilonε-form, which avoids real liminf/limsup junk values.
  • zαz_\alphazα​ is any real with cdf (gaussianReal 0 1) zα = 1 - α.

Standing assumptions and pins. Every goal-level statement carries (A1)–(A3) and the i.i.d. hypothesis. Three hypotheses are made explicit that the paper leaves implicit:

  1. f(x,⋅)f(x,\cdot)f(x,⋅) is measurable for each xxx ("f(x,ξ~)f(x,\tilde\xi)f(x,ξ~​) is a random variable");
  2. xn∗x_n^*xn∗​ is a measurable map that, almost surely, lies in XXX and minimizes fˉn\bar f_nfˉ​n​ over XXX on the same sample;
  3. (A2) is read as "sup⁡x∈Xf2(x,⋅)\sup_{x\in X}f^2(x,\cdot)supx∈X​f2(x,⋅) has an integrable majorant", which avoids proving that the supremum is measurable.

At n≤1n\le 1n≤1 the factors 1/n1/n1/n, 1/(n−1)1/(n-1)1/(n−1) and 1/n1/\sqrt n1/n​ evaluate to Lean's 000; every coverage statement is a liminf and ignores them.

Several encodings would trivialize the statement, and all are ruled out. The minimizer xn∗x_n^*xn∗​ must minimize the SAA problem of its own sample: a free xn∗x_n^*xn∗​, or one fitted to the other sample, would change the theorem. The second sample must not be replaced by an independent sequence. The quantile must not be pinned through an sInf. Positivity of σx^2(xmin⁡∗)\sigma^2_{\hat x}(x^*_{\min})σx^2​(xmin∗​) is a hypothesis only of the ε\varepsilonε-bound, as on p. 8.

One correction of the paper. Theorems 2 and 4 are stated for every 0<α<10<\alpha<10<α<1, but for α>1/2\alpha>1/2α>1/2 the paper's argument (replace xmin⁡∗x^*_{\min}xmin∗​ by xmax⁡∗x^*_{\max}xmax∗​) needs σx^2(xmax⁡∗)>0\sigma^2_{\hat x}(x^*_{\max})>0σx^2​(xmax∗​)>0, and without it both statements are false: for X=[−1,1]X=[-1,1]X=[−1,1], f(x,ξ)=x2−2xξf(x,\xi)=x^2-2x\xif(x,ξ)=x2−2xξ, ξ~∼N(0,1)\tilde\xi\sim N(0,1)ξ~​∼N(0,1), x^=0\hat x=0x^=0 and α=0.9\alpha=0.9α=0.9, the SRP coverage tends to about 0.0100.0100.010 and the A2RP coverage to e−2zα2≈0.037e^{-2z_\alpha^2}\approx0.037e−2zα2​≈0.037, both below 0.10.10.1. The Lean goal and Theorem 4 therefore carry the hypothesis "α≤1/2\alpha\le1/2α≤1/2, or σx^2(x)>0\sigma^2_{\hat x}(x)>0σx^2​(x)>0 for some x∈X∗x\in X^*x∈X∗". Theorem 3 is stated as printed.

Contributions are welcome at every level. The most reusable one is the uniform strong law of large numbers for Carathéodory integrands on a compact set with an integrable envelope, which also serves other SAA consistency results.

Selected references

  • G. Bayraksan, D. P. Morton, Assessing Solution Quality in Stochastic Programs, preprint (January 26, 2005); published in Math. Program. 108 (2006). https://doi.org/10.1007/s10107-006-0720-x
  • W. K. Mak, D. P. Morton, R. K. Wood, Monte Carlo bounding techniques for determining solution quality in stochastic programs, Oper. Res. Lett. 24 (1999) 47–56. https://doi.org/10.1016/S0167-6377(98)00054-6
  • R. Y. Rubinstein, A. Shapiro, Discrete Event Systems: Sensitivity Analysis and Stochastic Optimization by the Score Function Method, Wiley, 1993 (Lemma A1, p. 67; Theorem A1, p. 69).
  • A. Shapiro, Monte Carlo sampling methods, in: Handbooks in OR & MS 10, Stochastic Programming, Elsevier, 2003, 353–425. https://doi.org/10.1016/S0927-0507(03)10006-0
8 thms1 active userReviewed
Control TheoryProbabilityStochastic Systems·Captain: mikedeng1

Scheduling a Multi Class Queue with Many Exponential Servers: Asymptotic Optimality in Heavy Traffic: The HJB-Based Preemptive Policy Is Asymptotically Optimal Among Work-Conserving PoliciesResearch Paper

Motivation

Large call centers route several types of customers to a common pool of agents. When the pool is large and highly utilized, the relevant asymptotic regime is the quality-and-efficiency-driven (QED) or Halfin–Whitt regime (Halfin & Whitt 1981). The number of servers nnn grows while the offered load stays within O(n)O(\sqrt n)O(n​) of nnn. Waiting is then neither negligible nor overwhelming (Gans, Koole & Mandelbaum 2003).

Which class should a freed agent serve next? Exact optimization of a multi-class many-server queue with abandonment is intractable. The standard route is to solve a limiting diffusion control problem and translate its optimal control back into a policy for the queue. Atar, Mandelbaum and Reiman (Ann. Appl. Probab. 2004) carried this out for kkk customer classes, exponential service and abandonment, general renewal arrivals and general convex-type holding costs. They proved that the translated policy is asymptotically optimal. This mission formalizes that result for the preemptive policy.

Context:

  • Harrison & Zeevi (2004) studied the same multi-class many-server problem.
  • Bell & Williams (2001) proved asymptotic optimality of a threshold policy for a two-server system in conventional heavy traffic.
  • The present paper is the first to cover the QED regime with general costs and abandonment.

Setting

There are k≥1k\ge1k≥1 customer classes and nnn identical servers.

Primitives.

  • Arrivals. Class-iii customers arrive according to a renewal process AinA^n_iAin​ with interarrival times Uˇi(j)/λin\check U_i(j)/\lambda^n_iUˇi​(j)/λin​. Here the Uˇi(j)\check U_i(j)Uˇi​(j) are i.i.d., positive, of mean one and squared coefficient of variation CU,i2C^2_{U,i}CU,i2​.
  • Service. Service times are exponential with rate μin\mu^n_iμin​, represented by Poisson processes SinS^n_iSin​.
  • Abandonment. Waiting customers abandon at rate θin≥0\theta^n_i\ge0θin​≥0, represented by Poisson processes RinR^n_iRin​.

State. Xin(t)X^n_i(t)Xin​(t) is the number of class-iii customers in the system, Ψin(t)\Psi^n_i(t)Ψin​(t) the number in service and Φin=Xin−Ψin\Phi^n_i=X^n_i-\Psi^n_iΦin​=Xin​−Ψin​ the number waiting. The dynamics are

Xin(t)=Xi0,n+Ain(t)−Rin(∫0tΦin)−Sin(∫0tΨin),Ψn,Φn∈Z+k,∑iΨin≤n.X^n_i(t)=X^{0,n}_i+A^n_i(t)-R^n_i\Big(\int_0^t\Phi^n_i\Big)-S^n_i\Big(\int_0^t\Psi^n_i\Big),\qquad \Psi^n,\Phi^n\in\mathbb Z^k_+,\quad \textstyle\sum_i\Psi^n_i\le n .Xin​(t)=Xi0,n​+Ain​(t)−Rin​(∫0t​Φin​)−Sin​(∫0t​Ψin​),Ψn,Φn∈Z+k​,∑i​Ψin​≤n.

Policies.

  • A scheduling control policy (SCP) is the process Ψn\Psi^nΨn.
  • It is admissible if it does not anticipate the future beyond the time of the next arrival: past information is independent of future primitive increments.
  • It is work-conserving if no server idles while customers wait: (1⋅Xn−n)+=1⋅Φn(\mathbb 1\cdot X^n-n)^+=\mathbb 1\cdot\Phi^n(1⋅Xn−n)+=1⋅Φn.

Scaling and cost. In the QED scaling n−1λin→λin^{-1}\lambda^n_i\to\lambda_in−1λin​→λi​ with ∑iλi/μi=1\sum_i\lambda_i/\mu_i=1∑i​λi​/μi​=1. With ρi=λi/μi\rho_i=\lambda_i/\mu_iρi​=λi​/μi​ the centred processes are X^n=n−1/2(Xn−ρn)\hat X^n=n^{-1/2}(X^n-\rho n)X^n=n−1/2(Xn−ρn), Φ^n=n−1/2Φn\hat\Phi^n=n^{-1/2}\Phi^nΦ^n=n−1/2Φn and Ψ^n=n−1/2(Ψn−ρn)\hat\Psi^n=n^{-1/2}(\Psi^n-\rho n)Ψ^n=n−1/2(Ψn−ρn). The cost is

Cn=E∫0∞e−γtL~(Φ^n(t),Ψ^n(t)) dt.C^n=E\int_0^\infty e^{-\gamma t}\tilde L(\hat\Phi^n(t),\hat\Psi^n(t))\,dt .Cn=E∫0∞​e−γtL~(Φ^n(t),Ψ^n(t))dt.

The limiting control problem. It controls

X(t)=x+rW(t)+∫0tb(X(s),u(s)) ds,b(x,u)=ℓ+(μ−θ)(1⋅x)+u−μx,X(t)=x+rW(t)+\int_0^t b(X(s),u(s))\,ds,\qquad b(x,u)=\ell+(\mu-\theta)(\mathbb 1\cdot x)^+u-\mu x,X(t)=x+rW(t)+∫0t​b(X(s),u(s))ds,b(x,u)=ℓ+(μ−θ)(1⋅x)+u−μx,

where the control uuu takes values in the simplex Sk\mathbb S^kSk and WWW is a kkk-dimensional Brownian motion. The data are ri=(λiCU,i2+λi)1/2r_i=(\lambda_iC^2_{U,i}+\lambda_i)^{1/2}ri​=(λi​CU,i2​+λi​)1/2 and ℓi=λ^i−ρiμ^i\ell_i=\hat\lambda_i-\rho_i\hat\mu_iℓi​=λ^i​−ρi​μ^​i​. Its value V(x)V(x)V(x) is the infimum of E∫0∞e−γtL(X,u) dtE\int_0^\infty e^{-\gamma t}L(X,u)\,dtE∫0∞​e−γtL(X,u)dt, with L(x,u)=L~((1⋅x)+u,x−(1⋅x)+u)L(x,u)=\tilde L((\mathbb 1\cdot x)^+u,x-(\mathbb 1\cdot x)^+u)L(x,u)=L~((1⋅x)+u,x−(1⋅x)+u).

HJB equation and the proposed policy. The HJB equation is 12∑iri2∂iif+H(x,Df)−γf=0\tfrac12\sum_ir_i^2\partial_{ii}f+H(x,Df)-\gamma f=021​∑i​ri2​∂ii​f+H(x,Df)−γf=0 with H(x,p)=inf⁡u∈Sk[b(x,u)⋅p+L(x,u)]H(x,p)=\inf_{u\in\mathbb S^k}[b(x,u)\cdot p+L(x,u)]H(x,p)=infu∈Sk​[b(x,u)⋅p+L(x,u)]. Let hhh be a measurable selection of its minimizers. The proposed preemptive policy (P-SCP) sets the queue vector to Θ[(1⋅Xn−n)+h(X^n)]\Theta[(\mathbb 1\cdot X^n-n)^+h(\hat X^n)]Θ[(1⋅Xn−n)+h(X^n)], an integer rounding, and falls back to a static priority rule when that is infeasible.

Formalization targets

Goal: Theorem 2(i)

For a Cpol2C^2_{\mathrm{pol}}Cpol2​ solution fff of the HJB equation, a measurable minimizer selection hhh, and initial states with X^0,n→x\hat X^{0,n}\to xX^0,n→x:

lim⁡n→∞E∫0∞e−γtL~(Φ^tn,∗,Ψ^tn,∗) dt  ≤  lim inf⁡n→∞E∫0∞e−γtL~(Φ^tn,Ψ^tn) dt\lim_{n\to\infty}E\int_0^\infty e^{-\gamma t}\tilde L(\hat\Phi^{n,*}_t,\hat\Psi^{n,*}_t)\,dt\;\le\;\liminf_{n\to\infty}E\int_0^\infty e^{-\gamma t}\tilde L(\hat\Phi^n_t,\hat\Psi^n_t)\,dtn→∞lim​E∫0∞​e−γtL~(Φ^tn,∗​,Ψ^tn,∗​)dt≤n→∞liminf​E∫0∞​e−γtL~(Φ^tn​,Ψ^tn​)dt

This holds for every sequence of work-conserving admissible SCPs, and the left-hand limit exists and is finite. No constants are hard-coded.

Milestones

The milestones follow the proof:

  • on the diffusion side, Proposition 2 (well-posedness), Proposition 4 (stability and moment bounds), Proposition 5(i)–(ii) (growth and continuity of VVV) and Theorem 3 (VVV is the unique Cpol2C^2_{\mathrm{pol}}Cpol2​ HJB solution, and an optimal Markov policy exists);
  • on the queueing side, Proposition 1 (feedback rules give admissible SCPs), Lemmas 2–3 (moment bounds), Lemma 4(i)–(ii) (FCLT for the primitives and the fluid limit (Ψˉn,Φˉn)⇒(ρ,0)(\bar\Psi^n,\bar\Phi^n)\Rightarrow(\rho,0)(Ψˉn,Φˉn)⇒(ρ,0)), and Theorem 4(i)–(ii): lim inf⁡≥V(x)\liminf\ge V(x)liminf≥V(x) always, and lim sup⁡≤V(x)\limsup\le V(x)limsup≤V(x) under condition (49).

Significance

The result. Theorem 2(i) justifies using the diffusion control problem as a design tool for multi-class many-server systems. The policy is explicit given hhh, and it is optimal in the limit against all non-anticipating work-conserving policies, including those that use the full history and the time of the next arrival. The proof also identifies the limit cost with V(x)V(x)V(x).

Formalizing it. The result is proved on paper, with some steps (Proposition 1, the principle of optimality, the time-change and martingale limit theorems) given as sketches or citations. No part of it is machine-checked. A formalization requires:

  • a counting-process model of the queue;
  • a careful definition of non-anticipation;
  • a pathwise controlled SDE;
  • classical solvability of a semilinear elliptic HJB equation on Rk\mathbb R^kRk;
  • a weak-convergence argument in Skorokhod space.

Each of these is reusable well beyond this paper.

Difficulty

The obvious argument would show that X^n\hat X^nX^n converges to the controlled diffusion and pass the costs to the limit. This fails for two reasons:

  • the comparison class contains arbitrary non-Markov, history-dependent policies, so the queue does not converge to a single controlled diffusion;
  • the optimal selector hhh is in general discontinuous (for linear costs it is), so the proposed policy is not a continuous function of the state.

The proof instead compares every policy with the HJB solution through Itô's formula on the prelimit processes. This needs:

  • uniform moment bounds;
  • tightness of the integral processes;
  • the convergence of stochastic integrals of Kurtz and Protter;
  • and, for the proposed policy, the fact that the rounding Θ\ThetaΘ and the priority fallback perturb the minimizer by O(n−1/2)O(n^{-1/2})O(n−1/2).

Existence of a classical HJB solution on all of Rk\mathbb R^kRk, with only Hölder-continuous costs and polynomial growth, rests on a bounded-domain existence theorem for fully nonlinear elliptic equations.

Formalization scope

The Lean development commits to the following conventions.

  • Indexing and norms. Classes are Fin k with k≥1k\ge1k≥1; paper class iii is index i−1i-1i−1, so "class kkk" (highest priority, rounding remainder of Θ\ThetaΘ) is the last index. Vectors are Fin k → ℝ and ∥⋅∥\|\cdot\|∥⋅∥ is the paper's ℓ1\ell^1ℓ1 norm; the paper's ∣⋅∣|\cdot|∣⋅∣ on vectors is read the same way.
  • Probability space and paths. All systems share one complete probability space. Time is real and every condition is for t≥0t\ge0t≥0. The paper's "without loss" path regularity (finite arrival counts, Poisson paths Z+\mathbb Z_+Z+​-valued, nondecreasing and càdlàg) holds for every ω\omegaω.
  • Poisson processes are defined by independent Poisson increments; rate 000 gives the zero process.
  • Policies. A policy is a pair of real processes (Ψn,Xn)(\Psi^n,X^n)(Ψn,Xn) with integer values. Admissibility is Definition 2 verbatim, with the future σ\sigmaσ-field built from the next arrival time τin(t)\tau^n_i(t)τin​(t). Work conservation is (18).
  • Costs and value are lower Lebesgue integrals in [0,∞][0,\infty][0,∞], and lim⁡\limlim/lim inf⁡\liminfliminf are taken there. The integrands are nonnegative under work conservation.
  • Admissible systems range over sample spaces Ω : Type (universe 0). "Complete filtered probability space" means PPP complete with all null sets in F0\mathcal F_0F0​. Brownian motion is Mathlib's IsBrownianReal per coordinate, with independence and the (Ft)(\mathcal F_t)(Ft​)-Brownian property stated explicitly. VVV is the infimum over systems and their controlled processes.
  • Discount rate. γ>0\gamma>0γ>0 is a hypothesis; the paper leaves it implicit.
  • Initial states are integer vectors X0,n∈Z+kX^{0,n}\in\mathbb Z^k_+X0,n∈Z+k​ with n−1/2(X0,n−ρn)→xn^{-1/2}(X^{0,n}-\rho n)\to xn−1/2(X0,n−ρn)→x. The literal "X^0,n∈n−1/2Zk\hat X^{0,n}\in n^{-1/2}\mathbb Z^kX^0,n∈n−1/2Zk" would require ρin∈Z\rho_in\in\mathbb Zρi​n∈Z. Assumption 1(ii) is not imposed: each policy chooses its own initial split.
  • Lemma 3 is stated for all nnn beyond a threshold that depends on the sequence, with constants c,mˉc,\bar mc,mˉ chosen before xxx and the sequence. The printed all-nnn bound with ccc independent of xxx fails when the early terms X^0,n\hat X^{0,n}X^0,n are large.
  • Weak convergence to a continuous limit uses the coupling form CouplingConverges of the published BellWilliams2001.ThresholdPolicy.Paths; convergence to a deterministic limit is UocInProb.

The goal hypothesizes fff and hhh with the pointwise identity b(x,h(x))⋅Df(x)+L(x,h(x))=H(x,Df(x))b(x,h(x))\cdot Df(x)+L(x,h(x))=H(x,Df(x))b(x,h(x))⋅Df(x)+L(x,h(x))=H(x,Df(x)) for all xxx. An arbitrary "optimal Markov control policy" may differ from a minimizer selection on the Lebesgue-null lattice where X^n\hat X^nX^n lives, and that formalization would make the goal false. Restricting the comparators to feedback, Markov or nonpreemptive policies, fixing kkk, dropping abandonment, specializing to Poisson arrivals or linear costs, or imposing a common initial split would each trivialize or weaken the statement and is ruled out.

Not formalized:

  • Lemma 4(iii) (tightness);
  • Lemma 5 (Kurtz–Protter, which needs semimartingale theory absent from Mathlib);
  • Lemma 6 (convergence of Stieltjes integrals at limit points);
  • Proposition 5(iii);
  • the nonpreemptive results, Theorem 2(ii)–(iii).

Contributions are welcome on any milestone, and especially on infrastructure: Poisson and renewal processes, functional central limit theorems in Skorokhod space, classical solvability of elliptic HJB equations, and measurable selection of minimizers.

Selected references

  • R. Atar, A. Mandelbaum, M. I. Reiman, Scheduling a multi class queue with many exponential servers: asymptotic optimality in heavy traffic, Ann. Appl. Probab. 14(3), 2004. https://arxiv.org/abs/math/0407058
  • S. Halfin, W. Whitt, Heavy-traffic limits for queues with many exponential servers, Oper. Res. 29(3), 1981. https://doi.org/10.1287/opre.29.3.567
  • N. Gans, G. Koole, A. Mandelbaum, Telephone call centers: tutorial, review, and research prospects, Manuf. Serv. Oper. Manag. 5(2), 2003. https://doi.org/10.1287/msom.5.2.79.16071
  • J. M. Harrison, A. Zeevi, Dynamic scheduling of a multiclass queue in the Halfin–Whitt heavy traffic regime, Oper. Res. 52(2), 2004. https://doi.org/10.1287/opre.1040.0109
  • S. L. Bell, R. J. Williams, Dynamic scheduling of a system with two parallel servers in heavy traffic with resource pooling: asymptotic optimality of a threshold policy, Ann. Appl. Probab. 11(3), 2001. https://doi.org/10.1214/aoap/1015345343
  • T. G. Kurtz, P. Protter, Weak limit theorems for stochastic integrals and stochastic differential equations, Ann. Probab. 19(3), 1991. https://doi.org/10.1214/aop/1176990334
16 thms1 active userReviewed
Dynamic ProgrammingProbability·Captain: mikedeng1

Computational Issues in an Infinite-Horizon, Multiechelon Inventory Model 1: The Decomposition Policy Is Optimal for Discounted CostsResearch Paper

Motivation

Distribution systems often move stock in two stages. A depot orders from an outside supplier and ships to a retail outlet, where customer demand arrives and unmet demand is backordered. Stock held anywhere costs money, a shortage at the outlet costs more, and each order carries a fixed charge. The basic question is what ordering and shipping rule minimizes total cost.

Clark and Scarf (Management Science 6, 1960) showed that over a finite planning horizon this two-echelon problem decomposes. The outlet solves its own single-location problem, and the depot solves a second single-location problem in which the outlet's shortfall is charged through an induced penalty cost. Federgruen and Zipkin (Operations Research 32(4), 1984) carried the decomposition to the infinite horizon. In the infinite-horizon problems the induced penalty becomes stationary and explicit, which makes the system computable with single-location tools. This mission covers the discounted-cost half of that paper (§§1–2).

Timeline:

  • 1960: Clark and Scarf, finite-horizon decomposition, with a nonstationary penalty P^n\hat P_nP^n​ built from the outlet's optimal cost functions.
  • 1963: Iglehart (Management Science 9) proved, for the single-location discounted problem, that the finite-horizon value functions converge uniformly and that an (s,S)(s,S)(s,S) policy is optimal.
  • 1984: Federgruen and Zipkin combine the two results and prove that a stationary policy built from the decomposition is optimal for the infinite-horizon discounted and average-cost problems.

Setting

Time is discrete. The cost data are a fixed order cost KKK, an order cost rate cdc^dcd, a shipment cost rate crc^rcr, a holding cost rate hdh^dhd on all system stock, an extra holding cost rate hrh^rhr at the outlet, and a backorder penalty rate prp^rpr; all are positive. The discount factor α\alphaα satisfies 0≤α<10 \le \alpha < 10≤α<1, shipments take lll periods and orders take LLL periods. One-period demands are independent copies of a nonnegative continuous random variable uuu with mean μ<∞\mu < \inftyμ<∞, and u(i)u^{(i)}u(i) denotes the sum of iii copies.

The state is (y^,vd,xr)(\hat y, v^d, x^r)(y^​,vd,xr):

  • y^=(y1,…,yL)\hat y = (y^1, \dots, y^L)y^​=(y1,…,yL) lists the outstanding orders, yiy^iyi placed iii periods ago;
  • vdv^dvd is the depot's echelon inventory (its own stock plus xrx^rxr);
  • xrx^rxr is the outlet's stock plus shipments in transit.

An action is an order y≥0y \ge 0y≥0 and a shipment z≥0z \ge 0z≥0 with xr+z≤vd+yLx^r + z \le v^d + y^Lxr+z≤vd+yL. With demand uuu, the next state is ((y,y1,…,yL−1),vd+yL−u,xr+z−u)((y, y^1, \dots, y^{L-1}), v^d + y^L - u, x^r + z - u)((y,y1,…,yL−1),vd+yL−u,xr+z−u). The one-period cost is

cd(y)+hd(vd+yL)+crz+R(xr+z),c^d(y) + h^d(v^d + y^L) + c^r z + R(x^r + z),cd(y)+hd(vd+yL)+crz+R(xr+z),

where cd(y)=K+cdyc^d(y) = K + c^d ycd(y)=K+cdy for y>0y > 0y>0, cd(0)=0c^d(0) = 0cd(0)=0, and

R(x)=αl{−hd(x−lμ)+prE[u(l+1)−x]++(hd+hr)E[x−u(l+1)]+}.R(x) = \alpha^l\{-h^d(x - l\mu) + p^r E[u^{(l+1)} - x]^+ + (h^d + h^r)E[x - u^{(l+1)}]^+\}.R(x)=αl{−hd(x−lμ)+prE[u(l+1)−x]++(hd+hr)E[x−u(l+1)]+}.

Bα(s∣π)B^\alpha(s \mid \pi)Bα(s∣π) is the expected total discounted cost of a policy π\piπ from state sss.

The critical number xr∗x^{r*}xr∗ minimizes (1−α)crx+R(x)(1-\alpha)c^r x + R(x)(1−α)crx+R(x). The stationary induced penalty is P(x)=0P(x) = 0P(x)=0 for x≥xr∗x \ge x^{r*}x≥xr∗ and P(x)=(1−α)cr(x−xr∗)+R(x)−R(xr∗)P(x) = (1-\alpha)c^r(x - x^{r*}) + R(x) - R(x^{r*})P(x)=(1−α)cr(x−xr∗)+R(x)−R(xr∗) otherwise. The depot problem IHαdIH^d_\alphaIHαd​ has state (y^,vd)(\hat y, v^d)(y^​,vd), action y≥0y \ge 0y≥0 and one-period cost cd(y)+hd(vd+yL)+P(vd+yL)c^d(y) + h^d(v^d + y^L) + P(v^d + y^L)cd(y)+hd(vd+yL)+P(vd+yL). The policy πα∗\pi_\alpha^*πα∗​ orders by an optimal stationary policy σd\sigma^dσd of IHαdIH^d_\alphaIHαd​ and ships z=max⁡{0,min⁡{xr∗,vd+yL}−xr}z = \max\{0, \min\{x^{r*}, v^d + y^L\} - x^r\}z=max{0,min{xr∗,vd+yL}−xr}: up to the critical number when the depot has the stock, otherwise as much as it has.

Formalization targets

Goal: Theorem 1 (p. 827)

Assume αlpr≥(1−αl)hd\alpha^l p^r \ge (1-\alpha^l)h^dαlpr≥(1−αl)hd. For every state with y^≥0\hat y \ge 0y^​≥0 and xr≤vdx^r \le v^dxr≤vd, and every admissible policy π\piπ,

Bα(y^,vd,xr∣πα∗)≤Bα(y^,vd,xr∣π).B^\alpha(\hat y, v^d, x^r \mid \pi_\alpha^*) \le B^\alpha(\hat y, v^d, x^r \mid \pi).Bα(y^​,vd,xr∣πα∗​)≤Bα(y^​,vd,xr∣π).

The goal leaves the form of σd\sigma^dσd open: any optimal stationary depot policy will do, and no (s,S)(s,S)(s,S) structure is assumed.

Milestones

The milestones follow the paper's own route. Write g^n\hat g_ng^​n​, gnrg_n^rgnr​, g^nd\hat g_n^dg^​nd​, gndg_n^dgnd​ for the nnn-period optimal costs of the system, of the outlet, of the depot with penalties P^n\hat P_nP^n​, and of the depot with penalty PPP.

  • Eq. (4): g^n=g^nd+gnr\hat g_n = \hat g_n^d + g_n^rg^​n​=g^​nd​+gnr​.
  • Property (e): gnr→gr=Brαg_n^r \to g^r = B^{r\alpha}gnr​→gr=Brα.
  • §2 claim (Iglehart): gnr→grg_n^r \to g^rgnr​→gr uniformly on (−∞,xr∗](-\infty, x^{r*}](−∞,xr∗].
  • Lemma 1: P^n→P\hat P_n \to PP^n​→P uniformly on R\mathbb RR.
  • Lemma 2: g^nd−gnd→0\hat g_n^d - g_n^d \to 0g^​nd​−gnd​→0 uniformly.
  • Lemma 3: g^n→gd+gr\hat g_n \to g^d + g^rg^​n​→gd+gr.
  • Lemma 4: ggg satisfies the optimality equation (8), and πα∗\pi_\alpha^*πα∗​ attains it.

Significance

The theorem shows that, under discounting, the infinite-horizon two-echelon problem is solved by two single-location problems, with a penalty PPP that is written in terms of RRR alone. Computing PPP does not require the outlet's optimal cost functions. The rest of the paper relies on this: its computational sections evaluate PPP in closed form for normal demand, and they treat several outlets by relaxation. A machine-checked version also gives an infinite-horizon decomposition theorem against which future multi-echelon formalizations can be checked.

The result was proved in 1984 and is not open. It has not been formalized. The paper's proof is short only because it cites Iglehart's convergence results and Propositions 9.12 and 9.16 of Bertsekas and Shreve (1978) for its last step, so a formal proof must also supply these.

Difficulty

The obvious argument passes to the limit in the finite-horizon decomposition (4). That fails as stated, because the depot program (3) has nonstationary penalties P^n\hat P_nP^n​, built from the outlet's optimal costs gn−1rg_{n-1}^rgn−1r​, and its value functions are not those of any stationary problem. The comparison of P^n\hat P_nP^n​ with PPP needs uniform control over the whole real line. The first few P^n−P\hat P_n - PP^n​−P are in fact unbounded, since g0r=0g_0^r = 0g0r​=0 has the wrong slope. The uniform control therefore holds only for large nnn, and the error has to be propagated through the depot recursion.

The second obstacle is that the one-period costs are unbounded in both directions: hdvh^d vhdv is negative for negative vvv. Contraction arguments for bounded costs therefore do not apply. Lower boundedness on the feasible set needs the cost relation αlpr≥(1−αl)hd\alpha^l p^r \ge (1-\alpha^l)h^dαlpr≥(1−αl)hd, and passing from the optimality equation to optimality of a policy needs the theory of models with costs bounded below.

Formalization scope

Everything lives in the namespace FZEchelon.Discounted.

  • Model. The data form a structure Model. The pipeline y^\hat yy^​ is a vector indexed by {0,…,L−1}\{0, \dots, L-1\}{0,…,L−1}, whose index kkk is the paper's yk+1y^{k+1}yk+1. For L=0L = 0L=0 the current order arrives at once.
  • Policies and cost. Time runs forward with weight αk\alpha^kαk; the paper counts periods remaining. Policies are measurable, non-anticipative, deterministic and history dependent, and they must be feasible along every demand path. BαB^\alphaBα is an extended real: the expectation of the positive part of the discounted cost sum minus that of the negative part, under the product law of the demands.
  • Finite-horizon programs. These are real infima over the feasible actions.
  • Hypotheses. Statements quantify over the physical states y^≥0\hat y \ge 0y^​≥0, xr≤vdx^r \le v^dxr≤vd. The standing assumptions of §1 are bundled in StandingAssumptions: positive costs, 0≤α≤10 \le \alpha \le 10≤α≤1, demand nonnegative, atomless and of finite mean. The §2 statements add α<1\alpha < 1α<1 and the cost relation, which the paper names in the proof of Theorem 1. The critical numbers xr∗x^{r*}xr∗ and xnr∗x_n^{r*}xnr∗​ enter as minimizers. The depot policy σd\sigma^dσd enters as a measurable, nonnegative stationary policy that is optimal for IHαdIH_\alpha^dIHαd​; that is the paper's definition of πα∗\pi_\alpha^*πα∗​, and its existence is Iglehart's.
  • Ruled out. Comparing πα∗\pi_\alpha^*πα∗​ only against stationary policies, or reading BαB^\alphaBα as a bare series or a truncated sum, would trivialize or change the theorem. The comparison class is all admissible history-dependent policies.
  • Corrections. Where the paper says "bounded" for every nnn (§2 claim, Lemmas 1 and 2), the statements claim boundedness only where it holds: n≥1n \ge 1n≥1, n≥2n \ge 2n≥2, and eventually, respectively. The moderation notes give the counterexample at n=1n = 1n=1. Lemma 2 also carries the standing assumption of p. 821 that never ordering is not optimal. The statement is false without it.
  • Infrastructure. A complete development needs the convexity theory of the single-location newsvendor function RRR, value iteration for discounted models with costs bounded below, and the Markov property for the product measure on demand sequences. The control-system file is reusable for other inventory and queueing missions. Formalizations of Iglehart's theorem and of Bertsekas–Shreve Propositions 9.12 and 9.16 are welcome.

Selected references

  • A. Federgruen, P. Zipkin, Computational Issues in an Infinite-Horizon, Multiechelon Inventory Model, Operations Research 32(4):818–836, 1984. https://doi.org/10.1287/opre.32.4.818
  • A. J. Clark, H. Scarf, Optimal Policies for a Multi-Echelon Inventory Problem, Management Science 6(4):475–490, 1960. https://doi.org/10.1287/mnsc.6.4.475
  • D. L. Iglehart, Optimality of (s, S) Policies in the Infinite Horizon Dynamic Inventory Problem, Management Science 9(2):259–267, 1963. https://doi.org/10.1287/mnsc.9.2.259
  • D. P. Bertsekas, S. E. Shreve, Stochastic Optimal Control: The Discrete-Time Case, Academic Press, 1978. https://web.mit.edu/dimitrib/www/soc.html
11 thms1 active userReviewed
PreviousPage 64 of 69Next

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me