Motivation
A McKean–Vlasov control problem asks a planner to steer a stochastic differential equation whose coefficients depend not only on the current path of the state but also on its (conditional) distribution. Such problems describe the cooperative optimum of a large population of interacting agents: when each agent reacts to the empirical distribution of all others and a social planner chooses everybody's control, the limit as the population grows is a McKean–Vlasov control problem (Lacker 2017). When all agents are also exposed to a common noise (a shared source of randomness, such as a market factor), the relevant distribution is the conditional law of the state given that noise.
Bellman's dynamic programming principle (DPP) is the standard route from a control problem to its Hamilton–Jacobi–Bellman equation and to verification arguments. For McKean–Vlasov dynamics it fails in the naive state variable, because the problem is time-inconsistent; it can be recovered by taking the (conditional) law of the state as the state.
Timeline. Laurière and Pironneau (2014) and Bensoussan, Frehse and Yam (2013–2017) proved a DPP for densities of the state law. Pham and Wei (2017, SICON) proved a DPP over closed-loop controls in a Markovian setting with a common-noise filtration, under regularity assumptions; Pham and Wei (2018, ESAIM:COCV) treated the case without common noise. Bayraktar, Cosso and Pham (2018, Trans. AMS) obtained a randomised DPP, again Markovian and without common noise. Djete, Possamaï and Tan (2022) proved the DPP for the general non-Markovian problem with common noise, under Borel measurability of the coefficients only. This mission formalizes their DPP for the weak formulation.
Setting
Fix a horizon T>0, dimensions n,d,ℓ∈N, a nonempty Polish space (U,ρ) of control values, a point u0∈U and an exponent p≥0. Let Ck:=C([0,T],Rk) with the uniform norm, P(E) the Borel probability measures on E with the weak topology, and xt∧⋅ the path x stopped at time t. The coefficients b,σ,σ0,L are maps [0,T]×Cn×P(Cn×U)×U→Rn,Rn×d,Rn×ℓ,R, and g:Cn×P(Cn)→R; all are Borel measurable, and b,σ,σ0,L depend on the path and the measure only through their restrictions to [0,t] (non-anticipative).
A weak control γ with initial condition (t,ν)∈[0,T]×P(Cn) consists of a probability space carrying two filtrations Fγ⊇Gγ (the second models the common noise), a state Xγ, Brownian motions Wγ (idiosyncratic) and Bγ (common) on [t,T], a predictable U-valued control αγ, and measure-valued processes μsγ=L(Xs∧⋅γ∣Gsγ) and μˉsγ=L((Xs∧⋅γ,αsγ)∣Gsγ), such that Xt∧⋅γ∼ν(t), an integrability condition of order p holds, a compatibility condition between the two filtrations holds, and
dXsγ=b(s,Xγ,μˉsγ,αsγ)ds+σ(s,Xγ,μˉsγ,αsγ)dWsγ+σ0(s,Xγ,μˉsγ,αsγ)dBsγ,s∈[t,T].
The reward and the value function are
J(t,γ)=E[∫tTL(s,Xs∧⋅γ,μˉsγ,αsγ)ds+g(XT∧⋅γ,μTγ)],VW(t,ν)=γ∈ΓW(t,ν)supJ(t,γ),
with ∞−∞:=−∞ and sup∅=−∞. Each weak control also carries Asγ=∫ts∨tπ(αrγ)dr, where π is a Borel isomorphism of U onto a subset of [0,1], and the continuous conditional-law process μ^γ of (Xγ,Aγ,Wγ,Bγ) given Gγ.
Formalization targets
Goal: Theorem 3.1
VW:[0,T]×P(Cn)→[−∞,∞] is upper semi-analytic (every superlevel set {VW>c} is analytic), and for every stopping time τ⋆ of the canonical filtration of Ω⋆:=Cℓ×C([0,T],P(Cn×C×Cd×Cℓ)) with values in [t,T], setting τγ:=τ⋆(Bγ,t,μ^γ),
VW(t,ν)=γ∈ΓW(t,ν)supE[∫tτγL(s,Xs∧⋅γ,μˉsγ,αsγ)ds+VW(τγ,μτγγ)].
Milestones
The canonical reformulation on Ωˉ:=Ω^×C([0,T],P(Ω^)) with weak control rules PˉW(t,ν) (Definition 4.1): Lemma 4.4(i) (weak controls and weak control rules correspond), Corollary 4.6 (VW(t,ν)=supPˉ∈PˉW(t,ν)J(t,Pˉ)), Lemma 4.7 (analytic graphs, VW u.s.a.), Lemma 4.8 (stability under conditioning at a stopping time), Lemma 4.9 (dependence on the initial law only through ν^(t)), (4.10) (moment truncation VWM↗VW, analytic and u.s.a.) and Lemma 4.10 (universally measurable ε-optimal families and concatenation).
Significance
The result. The theorem gives the DPP for McKean–Vlasov control with common noise in a non-Markovian setting, with no regularity on the coefficients beyond Borel measurability and no integrability on the rewards. It is the starting point for the dynamic programming (HJB) equation on the space of probability measures, for the Markovian corollaries of §3.2 of the paper, and, through the equivalence of formulations in the companion work, for the strong formulations treated in the second mission of this series.
Formalizing it. The result is proved on paper; nothing of it is machine-checked. A formal development needs a working theory of weak controls with conditional laws, a controlled martingale problem on a space of measure-valued paths, and analytic-set measurability in continuous time. None of these exists in Mathlib, and the analytic-selection layer on Prove2Me (Bertsekas–Shreve, Chapter 7) is stated but not yet proved.
Difficulty
The obvious argument conditions an optimal-ish control on its information at the stopping time and pastes in near-optimal controls from the conditioned states. Each step fails without additional structure: VW is not known to be measurable, so E[VW(τγ,μτγγ)] is not obviously meaningful; conditioning a weak control on its information at τγ need not yield a weak control, because the conditional-law constraint involves the common-noise filtration and a compatibility condition between filtrations; and choosing a near-optimal control for each conditioned state is a selection problem over an uncountable family of probability spaces. The state of the problem is a law, which is itself defined only up to null sets, so random-time evaluations such as μτγγ need a continuous version.
Formalization scope
The Lean development lives in the namespace MKVDPP.Weak. Conventions:
- Times are in R≥0; paths are elements of C([0,T],Rk) with Rk Euclidean; a path is read at times beyond T as its value at T.
- Expectations and time integrals are E[ξ+]−E[ξ−] in [−∞,∞] with ∞−∞=−∞, computed as lower integrals.
- The supremum defining VW ranges over weak controls on probability spaces in
Type (universe 0).
- "The integrals are implicitly assumed to be well-defined" is pinned to local Itô integrals (
EthierKurtz.HasBrownianItoIntegral) of integrands that are a.s. square integrable on [t,T], taken against the shifted Brownian motions Wγ,t, Bγ,t; no L2(P) condition, which would shrink ΓW.
- The conditional-law requirement "for dP⊗ds-a.e. (s,ω)" is read in Fubini form; μTγ and μτγγ are the X-marginals of the continuous process μ^γ.
- "Local martingale on [t,T]" in Definition 4.1 is pinned to: every localised process Sφ,m, m≥1, is an (Fˉ,Pˉ)-martingale on [t,T]. The page integrates ∣Sˉ∣ and Sˉφ from 0. Here they run from the initial time t, because before t the canonical control may be ∂, where the coefficients are undefined, and a coefficient that is not integrable before t would empty P^W(t,⋅).
- The cemetery value ∂ of the canonical control is read as u0 inside the coefficients (a null set under Definition 4.1).
- In (3.2), (+∞)+(−∞) inside the expectation is −∞.
- Lemma 4.4(i)'s converse carries the hypothesis that the canonical A is a.s. ∫t⋅∨tπ(αˉr)dr; without it the converse is false, because Definition 4.1 leaves A before t unconstrained.
The goal does not mention the canonical space Ωˉ, weak control rules or any selection; a formalization of ΓW that is empty, or a value function that is constant, would trivialize it. A sanity file exhibits a weak control in the degenerate case n=d=ℓ=0 for every t≤T, so the structure is satisfiable.
Infrastructure a complete development needs: conditional laws and r.c.p.d. on Polish spaces, the predictable σ-algebra, stochastic integrals against a filtration, the Stroock–Varadhan martingale problem, Itô's formula, and analytic sets with the Jankov–von Neumann selection theorem. The analytic-selection results and the martingale-problem/SDE correspondence are reusable well beyond this mission. Contributions on any milestone, and on proving the referenced Bertsekas–Shreve selection theorems, are welcome. Out of scope: Theorem 3.4 (it rests on the companion paper), the Markovian results of §3.2, and the HJB discussion of §3.3.
Selected references
- M. F. Djete, D. Possamaï, X. Tan, McKean–Vlasov optimal control: the dynamic programming principle, Ann. Probab. 50(2), 2022; cited as arXiv:1907.08860v2. https://arxiv.org/abs/1907.08860v2
- H. Pham, X. Wei, Dynamic programming for optimal control of stochastic McKean–Vlasov dynamics, SIAM J. Control Optim. 55(2):1069–1101, 2017. https://arxiv.org/abs/1604.04057
- D. Lacker, Limit theory for controlled McKean–Vlasov dynamics, SIAM J. Control Optim. 55(3):1641–1672, 2017. https://arxiv.org/abs/1609.08064
- N. El Karoui, X. Tan, Capacities, measurable selection and dynamic programming part II: application in stochastic control problems, 2013. https://arxiv.org/abs/1310.3364
- E. Bayraktar, A. Cosso, H. Pham, Randomized dynamic programming principle and Feynman–Kac representation for optimal control of McKean–Vlasov dynamics, Trans. Amer. Math. Soc. 370(3):2115–2160, 2018. https://doi.org/10.1090/tran/7118
- M. Laurière, O. Pironneau, Dynamic programming for mean-field type control, C. R. Math. Acad. Sci. Paris 352(9):707–713, 2014. https://doi.org/10.1016/j.crma.2014.07.008
- A. Bensoussan, J. Frehse, S. Yam, Mean Field Games and Mean Field Type Control Theory, SpringerBriefs in Mathematics, Springer, 2013. https://doi.org/10.1007/978-1-4614-8508-7
- D. Bertsekas, S. Shreve, Stochastic Optimal Control: The Discrete-Time Case, Academic Press, 1978. https://web.mit.edu/dimitrib/www/SOC_1978.pdf
- D. Stroock, S. Varadhan, Multidimensional Diffusion Processes, Springer, 1997. https://doi.org/10.1007/3-540-28999-2