Markov Decision Processes VIII: Transaction Costs and the Dynamic Mean-Variance ProblemTextbook
Motivation
Two of the oldest simplifying assumptions in portfolio theory are that trading is frictionless
and that risk means variance. Neither survives contact with practice: every real market charges
a transaction cost proportional to the size of a trade, and variance penalizes upside deviations
exactly as much as downside ones, which is not what an investor actually fears. Bäuerle and
Rieder's §4.5 reopens the multiperiod terminal-wealth problem of chunk 04a with proportional
transaction costs added to every trade, and finds that the qualitative shape of the solution
survives — a buy/hold/sell rule with explicit thresholds, still obtained from the Structure
Theorem of chunk 02a. Their §4.6 then leaves expected-utility maximization altogether and
solves the classical Markowitz mean-variance problem in its genuinely dynamic, multiperiod form:
choose a self-financing trading strategy that attains a target expected terminal wealth
while minimizing the variance of that terminal wealth. This is Markowitz's one-period portfolio
selection problem (H. Markowitz, Portfolio Selection, Journal of Finance, 1952) transplanted
into a stage-by-stage trading horizon, and it earns its own solution technique: the objective is
not linear in the underlying probability measure, so no direct Bellman equation applies, and the
chapter instead builds a Lagrangian-embedding argument from scratch. Section §4.7 closes the
chapter by replacing variance with the Average-Value-at-Risk, an axiomatically better-behaved
risk measure (P. Artzner, F. Delbaen, J.-M. Eber, D. Heath, Coherent Measures of Risk,
Mathematical Finance, 1999), and solves the resulting mean-risk problem in the binomial model by
the same Lagrangian route.
Setting
The transaction-cost model (§4.5): state (bond and stock holdings), action (the stock holding chosen after the trade), bond holding after the trade if and if , for a proportional cost rate ; transition ; terminal reward for a utility homogeneous of degree .
The mean-variance model (§4.6): state (wealth), action (amounts invested in risky assets, short-selling allowed), transition . Writing for the terminal wealth reached from under a strategy , the problem is
Because is not linear in the law of , (MV) is solved via the Lagrangian , whose saddle points give (MV)'s value and optimizer, reduced in turn to the tractable auxiliary quadratic problem : minimize , a stochastic linear-quadratic control problem.
The mean-risk model (§4.7): the binomial (Cox–Ross–Rubinstein) market with one bond (interest rate ) and one stock with relative return w.p. or w.p. ; the Average-Value-at-Risk at level , ; the problem subject to , solved via the same Lagrangian route through an auxiliary problem .
Formalization targets
Goal — Theorem 4.6.6 (the mean-variance problem)
where is a recursively-defined sequence in (Lemma 4.6.4) built from the one-period return moments . This closes the loop the chapter opens: it is the exact value and optimal strategy of the dynamic mean-variance problem, obtained by specializing the auxiliary problem 's closed-form solution (Theorem 4.6.5) at the Lagrange multiplier that Lemma 4.6.2's saddle-point argument selects.
Supporting milestones
The Lagrangian route itself: the equivalence of (MV) and its equality-constrained form (Lemma 4.6.1), the saddle-point value identity (Lemma 4.6.2), the reduction of the Lagrange problem to (Lemma 4.6.3), the boundedness of (Lemma 4.6.4), and 's own explicit solution (Theorem 4.6.5) — the four-step argument the goal theorem is the payoff of. Upstream of §4.6: the transaction-cost model's upper bounding function (Proposition 4.5.1), its Structure Assumption via buy/hold/sell decision rules (Proposition 4.5.2), and the resulting explicit three-region optimal policy (Theorem 4.5.4). Downstream: the Two-Fund Theorem (Corollary 4.6.7), and the parallel mean-risk development — the auxiliary problem 's solution (Theorem 4.7.1), the binomial value of (Proposition 4.7.2), and the mean-risk problem's own explicit solution in both orderings of and (Theorems 4.7.3 and 4.7.4).
Significance
Theorem 4.6.6 is the multiperiod extension of the single most-used result in portfolio theory: the mean-variance efficient frontier, here derived stage by stage rather than assumed static, and it recovers the classical Two-Fund Theorem (every investor holds the same risky portfolio, scaled by wealth) as an immediate corollary rather than a separate argument. The transaction-cost results answer a standing objection to frictionless portfolio theory by showing that its qualitative conclusions — a threshold trading rule derived from a value function via the same abstract Structure Theorem — survive costs, with the thresholds now depending on the current value function rather than being fixed. The mean-risk results extend the whole technique to a risk measure that, unlike variance, is coherent in the sense of Artzner et al., showing the Lagrangian-embedding method is not an accident of the quadratic case.
None of this chapter's results have machine-checked proofs on Prove2Me at the time of writing
(the platform's saddle-point sufficiency results, VectorSpaceOpt.lagrangian_saddle_sufficient_pointed
and ConvexOptimization.lagrangian_saddle_iff_strong_duality, are stated over a closed convex
cone in a normed vector space, not over the finite-horizon admissible-policy space that
Lemma 4.6.2 needs, and were checked and ruled out as reusable for this mission). Formalizing this
chapter means building the Lagrangian-embedding argument for a dynamic (rather than static)
optimization problem from scratch: no existing platform infrastructure covers a saddle point of
a Lagrangian defined over a sequence of Markov policies.
Difficulty
The obvious first attempt at (MV) is to apply the Structure Theorem of chunk 02a directly to
the variance objective, exactly as chunk 04a does for expected utility. This fails outright:
is not
additive over time and has no Bellman recursion of the usual form, because the square of an
expectation over the whole horizon cannot be decomposed into a sum of one-period rewards. The
chapter's actual route — Lagrangian relaxation to , then a further reduction to the
quadratic (and hence tractable) — is not a shortcut around this obstacle but the only way
the mean-variance problem admits a Markov Decision Process reformulation at all. A correct
formalization of the goal theorem must go through this exact chain (saddle_point_value,
plambda_implies_qp, qp_solution), not around it.
Formalization scope
The financial market and the four named optimization problems (MV), (MV=), ,
are formalized as explicit structures and Prop-valued predicates in MDPFinance.MeanVariance
(none of them is a numbered definition in the book — each is introduced only in prose — so each
gets its own precise Lean definition rather than being left implicit). Wealth is real-valued,
policies are sequences of measurable Markov maps , and values that can be in the book (the value of , of
, and of (MR) itself) are typed EReal rather than ℝ, matching the book's own use
of infinite values as legitimate outcomes rather than failure states. A formalization that solved
the goal theorem by first proving a Bellman equation for directly
would not be proving Theorem 4.6.6 — no such recursion exists — and the goal statement is phrased
purely in terms of IsOptimalMV, varXN, and meanXN, independent of any intermediate value
function, precisely so that only the actual saddle-point argument can discharge it. The
transaction-cost model's buy/hold/sell threshold functions are
represented by their defining maximizing property rather than a closed form, since the book
itself only pins them down as an argmax. Reusable beyond this mission: the MVMarket/
MeanRiskMarket structures and the Lagrangian-saddle-point machinery are natural substrate for
any later mission that needs a dynamic risk-constrained portfolio problem. Contributions
completing any milestone's sorry are welcome, particularly a sorry-free proof of Lemma 4.6.2
(the saddle-point value identity), since it is the one genuinely general technique this mission
introduces.
Selected references
- H. Markowitz, Portfolio Selection, The Journal of Finance 7(1), 1952, https://doi.org/10.2307/2975974
- P. Artzner, F. Delbaen, J.-M. Eber, D. Heath, Coherent Measures of Risk, Mathematical Finance 9(3), 1999, https://doi.org/10.1111/1467-9965.00068
- N. Bäuerle, U. Rieder, Markov Decision Processes with Applications to Finance, Universitext, Springer, 2011, https://doi.org/10.1007/978-3-642-18324-9, Chapter 4, §§4.5-4.7