Markov Decision Processes XII: Terminal Wealth and Mean-Variance under Partial ObservationTextbook
Motivation
Every portfolio-choice model treated so far in this series assumes the investor knows the exact law governing the market's returns. Real investors do not: the drift of a stock, the regime a market is in, or the probability of an up-move in a simplified binomial model is itself uncertain and must be learned from the very prices being observed. Bäuerle and Rieder's Chapter 6 (Markov Decision Processes with Applications to Finance, Springer, 2011) is the book's synthesis of two threads developed separately earlier: Chapter 5's reduction of a partially observable decision problem to an ordinary one on an enlarged "belief" state space, and Chapter 4's classical solutions of terminal-wealth utility maximization and dynamic mean-variance portfolio choice. Combining them answers a natural question with no simple a priori answer: how does not knowing which market you are in change the qualitatively optimal way to invest, and can the closed-form solutions of the fully-observed theory be recovered, term for term, once the unknown factor is replaced by a belief about it?
Setting
The market has an unobservable factor Y (state space E_Y) driving the vector of relative
risks Z ∈ ℝ^d of d risky assets: given Y_n = y, the next return Z_{n+1} has a density
q_R(y,\cdot), and Y itself evolves via its own transition density q_Y(y,\cdot), jointly —
crucially, this joint law depends only on y, never on wealth or the action taken. An investor
observes only the stock prices (equivalently, the return history), never Y itself. Bayes' rule
turns this into a filtering problem: the investor's belief ρ_n \in \mathbb P(E_Y) about the
current factor is updated one return at a time by an operator Φ(ρ,z) that depends only on the
current belief and the newly observed return — a genuine simplification of Chapter 5's general
Bayes operator, forced by the market's own structure. The pair (x_n,ρ_n) — observable wealth
and current belief — is then an ordinary, fully observed state for an ordinary Markov Decision
Model, and every value function and optimal policy of this chapter lives on that enlarged state
space.
Formalization targets
The goal, Theorem 6.2.3, solves the dynamic mean-variance problem (MV): minimize the
variance of terminal wealth X_N subject to a target expectation \mathbb E[X_N] \ge \mu, under
partial observation. It is reached by a Lagrangian embedding into an auxiliary quadratic-loss
problem QP(b), solved explicitly in Theorem 6.2.2, whose value function factors as
((xS^0_N/S^0_n)-b)^2 d_n(\rho) for a belief-only sequence (d_n) satisfying the backward
recursion (6.7); Lemma 6.2.1 shows this sequence always lies strictly between 0 and 1, which
is exactly what makes the final variance formula and Lagrange multiplier well-posed. The
remaining milestones develop the parallel terminal-wealth theory of §6.1: the general structure
theorem (Theorem 6.1.1), its power- and logarithmic-utility closed forms (Theorems 6.1.2, 6.1.7),
and — for the specific binomial market with an unknown up-probability — a likelihood-ratio
monotonicity result for the filter update (Lemma 6.1.4) and a comparison between the partially
and completely observed optimal investment fractions (Theorem 6.1.5).
Significance
The chapter's organizing insight is that partial observation does not require a new theory: once
the belief is added as a state coordinate, every general result already proved for fully
observed Markov Decision Models — the Bellman equation, the existence of optimal Markov
policies, the Lagrangian embedding technique for mean-variance problems — applies unchanged. What
is genuinely new, and genuinely non-trivial, is checking that the reduced model inherits the
structural hypotheses (monotonicity, boundedness, positive-definiteness of covariance matrices)
those general theorems require, expressed now as conditions on the belief-indexed quantities
Φ(ρ,z), d_n(\rho), \ell_n(\rho), C_n(\rho) rather than on the original, unobserved
factor. Theorem 6.1.5's comparison result is a genuinely new phenomenon with no fully-observed
analogue at all: it quantifies, in the two opposite directions dictated by the sign of the
risk-aversion parameter γ, how residual uncertainty about the market itself changes the
qualitatively optimal amount to invest — the discrete-time analogue of a continuous-time result
in the literature this book cites (Sass and Haussmann 2004).
Difficulty
The recurring difficulty across every result in this mission is that the reduced model's state
space E_X \times \mathbb P(E_Y) includes a space of probability measures as one coordinate,
and every quantity that must be shown well-defined, monotone, or bounded is a functional on that
space, not a function on a concrete Euclidean set. Formalizing the mean-variance recursion (6.7)
in particular is a three-way mutual computation — a scalar d_n(\rho), a vector \ell_n(\rho),
and a matrix C_n(\rho), each an integral against the same belief-dependent predictive law of
the next return, each feeding the next stage's version of all three — where Lemma 6.2.1's
strict-inequality bound is not a bookkeeping detail but exactly the fact that keeps C_n(\rho)
invertible and the whole construction from breaking down. Theorem 6.2.3 itself is the hardest
single step: verifying that the specific constant b^* the Lagrangian method selects makes the
mean constraint bind at exact equality, and that the resulting variance is the true constrained
minimum (not merely a feasible value), is exactly the non-trivial content a superficial
restatement of Theorem 6.2.2 at an unspecified b would silently discard.
Formalization scope
Every value function of this chapter — the terminal-wealth maximization of §6.1, the quadratic
loss QP(b) and the mean-variance problem (MV) of §6.2 — is built from one shared
history-dependent value-function scaffold, parametrized by its terminal payoff (the utility U,
a quadratic loss, or the raw first/second moment), its rate sequence (constant in §6.1,
non-stationary in §6.2), and its feasible-action correspondence, rather than four separately
re-derived constructions. Optimal fractions in the binomial sub-model (Lemma 6.1.4, Theorem
6.1.5) are characterized as any maximizer of the relevant one-step concave problem rather than
through the closed-form solution the book's own proof derives via machinery from a different,
unavailable chunk (Lemma 4.2.9) — the comparison and monotonicity results proved here are facts
about any such maximizer, not about that specific formula. A formalization that assumed the
reduced model's filter update or covariance structure directly, rather than deriving it from the
market's own return and factor densities via the Bayes operator Φ, would trivialize every
result in this mission; none of the items here take that shortcut.
Selected references
- N. Bäuerle and U. Rieder, Markov Decision Processes with Applications to Finance, Universitext, Springer, 2011. DOI: 10.1007/978-3-642-18324-9.
- R. Sass and U. G. Haussmann, "Optimizing trading strategies with respect to drawdown in the hidden Markov model," Statistics & Decisions, 2004.
- N. Bäuerle and U. Rieder, "Portfolio optimization with unobservable Markov-modulated drift process," Journal of Applied Probability, 2007.
- V. Runggaldier, W. Trivellato, and T. Vargiolu, "A Bayesian adaptive control approach to parameter estimation and optimal portfolio selection," in Mathematical Finance, Trends in Mathematics, Birkhäuser, 2002 (the binomial-market source this chapter's §6.1 example specializes).