Nonzero-Sum Stochastic Differential Games with Impulse Controls: A Verification Theorem with Applications 1: Regular Solutions of the Quasi-Variational Inequalities Give Nash Equilibrium PayoffsResearch Paper
Motivation
Many economic and engineering systems are steered by agents who act at discrete instants rather than continuously: a central bank intervenes on an exchange rate, two energy producers adjust a shared stock, a firm rebalances inventory. Each action has a fixed cost, so continuous control is not realistic. The mathematical model is impulse control: the state follows a diffusion, and a controller may shift it at chosen stopping times by paying a cost. The single-controller theory is classical (Øksendal and Sulem, Applied Stochastic Control of Jump Diffusions, 2007). When two controllers with different objectives act on the same state, the result is a nonzero-sum stochastic differential game with impulse controls.
Before Aïd, Basei, Callegaro, Campi and Vargiolu, the literature on games with impulse controls was almost entirely zero-sum: Cosso (SIAM J. Control Optim., 2013) characterised the value of zero-sum impulse games through a double-obstacle quasi-variational inequality in the viscosity sense. Their paper (Math. Oper. Res. 45(1), 2020; arXiv:1605.00039) gives the first general formulation of the nonzero-sum case together with a verification theorem: a system of quasi-variational inequalities (QVIs) whose sufficiently regular solutions are the equilibrium payoffs. Its Section 4 then computes Nash equilibria in closed form for a one-dimensional game.
Setting
A -dimensional Brownian motion on a filtered probability space satisfying the usual conditions drives the state equation
with globally Lipschitz and . The game takes place in an open set and ends at the exit time of the state from . Each of two players has a nonempty impulse set and a continuous impulse map : an intervention with impulse moves the state from to .
A strategy of player is a pair with open and continuous. Player intervenes as soon as the state leaves , with impulse at the current state . Player 1 has priority on ties, and several interventions may happen at the same instant. This defines the controlled process , the intervention times and impulses of each player, and the states just before each intervention. The payoff of player is
where , , is a running payoff, the cost of one's own interventions, the gain from the opponent's, and a terminal payoff on . A pair of strategies is -admissible, , when these four terms are integrable, has all moments, and the interventions do not accumulate before . A Nash equilibrium is a pair in from which no player gains by a unilateral admissible deviation.
Given candidate payoff functions on , let be the unique maximiser of over . Define , , the continuation region and the generator . The QVI system is
Formalization targets
Goal: Theorem 3.3 (verification theorem)
Suppose solve the QVI system, with polynomial growth, is a Lipschitz surface near which has locally bounded first and second derivatives, , and the threshold pair is -admissible. Then
Milestones
- Lemma 2.3: the controlled process is the concatenation of diffusion pieces, it jumps only at interventions, and between interventions it stays in .
- Remark 3.6, (3.8b), (3.8d), (3.8f): against , the state stays in , and player 2 intervenes only on with impulse .
- Step 1 of the proof: for every admissible deviation .
- Step 2 of the proof: .
The goal follows from Steps 1 and 2 and their mirror images for player 2.
Significance
The theorem turns the search for Nash equilibria of nonzero-sum impulse games, an infinite-dimensional fixed-point problem over strategy pairs, into a deterministic problem: find functions satisfying a system of coupled QVIs with prescribed regularity. The regularity conditions become smooth-pasting conditions, hence a system of algebraic equations; Section 4 of the paper solves it explicitly for a one-dimensional game with linear payoffs. A further consequence is structural: equilibrium payoffs need only be on the opponent's continuation region, which is what lets non-smooth, piecewise-defined candidates qualify.
The result is proved in the paper; no machine-checked version exists. A complete formalization would be the first verified verification theorem for impulse control, single-player or game, and would expose every convention of the model: priority on ties, simultaneous interventions, the treatment of exit, and the integrability of the payoff. Several of these conventions need correction on the page, as listed below.
Difficulty
The heuristic argument applies Itô's formula to and uses the QVIs term by term. This fails on two counts. First, is only across the free boundary , so Itô's formula does not apply directly. The paper mollifies (following Øksendal's proof of his verification theorem for optimal stopping) and must control the second derivatives near a Lipschitz boundary. Second, the sums over interventions may be infinite and the horizon unbounded, so expectations and limits do not commute. The passage to the limit needs the integrability built into and the polynomial growth of . The stochastic-calculus infrastructure itself (Itô's formula for continuous semimartingales stopped at random times, strong solutions of Lipschitz SDEs restarted at stopping times) is largely missing from Mathlib.
Formalization scope
The Lean model is pathwise. A realization is a sequence of diffusion pieces, each solving the state equation from the random restart time with the restart value. The Itô integral is the published relation EthierKurtz.HasBrownianItoIntegral, and the stochastic basis uses the published You2015.Shared.UsualConditions and IsFBrownian. Times take values in with , and the state space is EuclideanSpace ℝ (Fin d). asks for one admissible realization, while the Nash inequalities and the payoff identity hold on every admissible realization; strong uniqueness makes these readings equivalent.
The following deviate from the page and are disclosed in the items:
- is assumed continuous, so that is a strategy.
- The fourth QVI is imposed on , where exists.
- The exit time is , not the printed .
- The gain term of (2.7) is read with the opponent's interventions.
- The supremum in (2.8) is taken over .
- Interventions accumulating at a finite are excluded from , since the page leaves undefined there.
- Lemma 2.3 is stated in corrected form for simultaneous and boundary interventions.
The payoff is a Bochner expectation only inside , which requires the conditions of (2.7) and the integrability of each payoff. A non-integrable deviation therefore never receives the junk payoff , and the Nash inequality cannot hold vacuously.
Contributions are welcome on the stochastic-calculus layer this needs (Itô's formula, strong existence and uniqueness for Lipschitz SDEs, optional stopping for stochastic integrals) and on the mollification lemma for functions that are with piecewise bounded second derivatives across a Lipschitz surface. These are reusable well beyond this mission.
Selected references
- R. Aïd, M. Basei, G. Callegaro, L. Campi, T. Vargiolu, Nonzero-sum stochastic differential games with impulse controls: a verification theorem with applications, Math. Oper. Res. 45(1), 2020 (accepted manuscript, arXiv:1605.00039v4). https://arxiv.org/abs/1605.00039
- A. Cosso, Stochastic differential games involving impulse controls and double-obstacle quasi-variational inequalities, SIAM J. Control Optim. 51(3), 2102–2131, 2013.
- B. Øksendal, Stochastic Differential Equations, 6th ed., Springer, 2003. https://doi.org/10.1007/978-3-642-14394-6
- B. Øksendal, A. Sulem, Applied Stochastic Control of Jump Diffusions, 2nd ed., Springer, 2007. https://doi.org/10.1007/978-3-540-69826-5