Markov Decision Processes XVII: Random-Horizon Consumption-Investment and the De Finetti Dividend ProblemTextbook
Motivation
An insurance company collects premia and pays claims each period; the difference is a random, signed quantity that can push the company's risk reserve up or down. At the start of every period, before that period's premia and claims are realized, the company's owners may pay themselves a dividend out of the current reserve — but once the reserve goes negative the company is ruined and stops operating for good. How should the owners time and size these payments to maximize the total expected discounted dividend paid out before ruin? This is the classical De Finetti dividend problem, one of risk theory's oldest optimization questions, and Chapter 9 §9.2 of Bäuerle and Rieder's Markov Decision Processes with Applications to Finance (Springer, 2011) solves its fully discrete-time version by identifying the exact combinatorial shape of the optimal policy — not just proving one exists. This mission also covers §9.1, a different application of Chapter 7's contracting theory to a consumption-investment problem whose planning horizon is itself random rather than fixed or infinite.
Setting
The dividend model is a stationary Markov Decision Model on the integers: the state is the current risk reserve, the action (for ; only is available once ruined) is the dividend paid, the reward is , and the reserve evolves by i.i.d. increments (premia minus claims) after the dividend is deducted. Because the reward is bounded by an explicit function of the state (Lemma 9.2.2), Chapter 7's general existence theory applies directly, and the value function satisfies a genuine Bellman equation. The chapter's real content begins once existence is established: Theorem 9.2.3 pins down enough analytic structure of and its largest-maximizing policy (monotonicity, a Lipschitz-type inequality, and a self-consistency identity) to drive a purely combinatorial argument that 's shape is a finite alternation of "pay nothing" and "pay down to a fixed level" intervals — a band-policy (Definition 9.2.5). Section 9.1's random-horizon consumption-investment model reuses the same Chapter 7 machinery in a different setting: the usual (consumption, portfolio) decision each period, but where the horizon itself ends after each period with probability , making the effective one-period discount rather than .
Formalization targets
The goal, Theorem 9.2.9, states the section's main claim in one sentence: the stationary policy is optimal and is a band-policy. Short as it is stated, its proof assembles every earlier result of the section. The milestones supply that assembly, in order: Lemma 9.2.2 gives the model's bounding function and the resulting integrability/convergence facts; Theorem 9.2.3 gives the value-function bounds and the self-consistency identity ; Corollary 9.2.4 checks the two sign-definite degenerate cases directly from Theorem 9.2.3; Proposition 9.2.6 proves the top threshold is finite (not merely well-defined) and that is a simple barrier above it; Proposition 9.2.8 proves the increment property below that forces each band's shape; and Theorem 9.2.10 (a postscript refinement, stated after the goal) shows the wave lengths are bounded once the reserve's downward jumps are themselves bounded, collapsing to a single barrier-policy in the extreme case. Theorem 9.1.1, the random-horizon consumption-investment verification theorem, is included as a full item but is not a milestone of this goal, since its content and proof belong to a different, disjoint model — see Difficulty.
Significance
Band-policies and the discrete-time De Finetti dividend problem have no substrate anywhere in Mathlib or on the platform, and the result is a genuinely deep, classical one: a discrete-time analogue of the continuous-time De Finetti barrier-strategy theory, obtained here by pure dynamic-programming argument rather than the stochastic-calculus techniques the continuous-time theory usually relies on. The mission is explicit that the goal's conclusion is the general band-policy structure, not the strictly weaker barrier-policy special case that Theorem 9.2.10 b) proves only under an extra hypothesis (bounded downward jumps) — stating the goal with a barrier-policy conclusion instead would understate what Theorem 9.2.9 actually proves.
Difficulty
The central formalization challenge is Definition 9.2.5's own combinatorial intricacy: a
band-policy is specified by an alternating chain of thresholds with a positive-width gap condition on every wave, and the policy's four
piecewise branches case-split on which wave (if any) the current state falls into. This mission
renders it existentially over the witnessing rather than as one closed-form function, a
faithful but more verbose transcription that avoids conflating the different branch conditions.
A second difficulty is Proposition 9.2.6's own finiteness claim: is a supremum over a
subset of that could, in principle, be unbounded, and Mathlib's convention for
sSup over the naturals returns a finite junk value () even for an unbounded set — using it
directly would silently trivialize "" into a claim that is true regardless of the
proposition's actual mathematical content. This mission instead states the proposition by
exhibiting the finite value of directly, so that " is finite" survives as genuine
content that the theorem's proof must establish. A third difficulty is scope: Theorem 9.1.1's
random-horizon consumption-investment model shares no state space, action space, or definitions
with the dividend model of the goal, despite both appearing in this chunk's assigned page range;
it is formalized as a genuine application of a locally-restated copy of Chapter 7's contracting
theory, but is excluded from the milestone list proper since it plays no role in the goal's own
proof.
Formalization scope
The dividend model's transition law is built from Mathlib's PMF (probability mass function)
type on , which supplies the "probabilities sum to one" fact automatically rather than
as a separate hypothesis. J_\infty, \delta, and every finite-horizon value function throughout
this mission use this whole book series' Filter.limsup-of-truncations convention for
infinite-horizon reward, restated locally (own namespace copy, per this series' file-ownership
boundary) from chunk 07a's identical apparatus rather than imported. The consumption-investment
model of §9.1 is formalized with the number of risky assets as an explicit type parameter and
its admissible-portfolio and domain restrictions as separate, citable fields rather than folded
silently into the reward or transition definitions.
Selected references
- N. Bäuerle and U. Rieder, Markov Decision Processes with Applications to Finance, Universitext, Springer, 2011. DOI: 10.1007/978-3-642-18324-9.
- B. De Finetti, "Su un'impostazione alternativa della teoria collettiva del rischio", Transactions of the XVth International Congress of Actuaries, 1957 (the original continuous-time dividend problem this chapter's discrete-time analogue is modeled on).
- H. Schmidli, Stochastic Control in Insurance, Springer, 2008 (cited by Remark 9.2.1 for the reduction from a continuous dividend-payout action space to the integer setting used throughout this section).
- H. U. Gerber, "Games of economic survival with discrete- and continuous-income processes", Operations Research, 1972 (an early discrete-time treatment of the same class of problems, in the spirit this chapter's own model follows).