Conditional Logit Analysis of Qualitative Choice Behavior 3: The Conditional Logit Likelihood Has a Maximum Exactly When No Direction Makes Every Observed Choice Weakly BestResearch Paper
Motivation
The conditional logit model is the workhorse of discrete choice analysis in transportation, marketing, labour and industrial organization. McFadden's 1974 chapter derived it from a theory of population choice behaviour and showed how to estimate it by maximum likelihood; this line of work was recognized by his 2000 Nobel Prize in Economic Sciences, awarded for theory and methods of discrete choice analysis. Every applied logit estimation rests on a basic question: does the maximum likelihood estimate exist for the sample at hand? In small samples it may not. When one alternative is always chosen whenever it is available, the likelihood keeps increasing as a parameter tends to infinity, and numerical optimizers report diverging coefficients. This failure is known in the binary case as complete or quasi-complete separation. McFadden's Lemma 3 gives the exact condition, for the multinomial conditional logit model with general alternative sets, under which a maximizer exists.
Timeline. Berkson (1951, 1955) popularized binomial logit; multinomial versions were developed by Gurland (1960), Bloch (1967), Rassam (1971), McFadden (1968) and Theil (1969, 1970). McFadden (1974) stated the existence criterion for the conditional logit likelihood (Lemma 3) together with a quadratic-programming test for it (Lemma 4). Albert and Anderson (1984) later classified separation patterns for binary and multinomial logistic regression, and Haberman (1974) treated existence for log-linear models.
Setting
A choice experiment has trials. Trial offers an alternative set of alternatives, indexed , each described by an attribute vector (the values of specified functions of the individual's and the alternative's characteristics). Trial is repeated times, and alternative is chosen times, so .
For a parameter , with the inner product, the selection probabilities are
and the log-likelihood of the sample is
Write for the probability-weighted mean attribute vector of trial .
Axiom 5 (Full Rank). The matrix with rows has rank .
Axiom 6. There is no nonzero with for all and . Equivalently, no nonzero direction makes every observed choice weakly best in its alternative set.
Formalization targets
Goal: Lemma 3
Under Axiom 5,
Milestones
- Equation (19): the gradient .
- Equation (20): the Hessian .
- is concave, and every critical point is a global maximizer.
- A Hessian that is nonsingular everywhere makes strictly concave with at most one maximizer.
- Axiom 5 holds at if and only if the Hessian at is negative definite.
- Necessity: under Axiom 5, a maximizer forces Axiom 6.
- Equation (21): under Axiom 6, has a positive lower bound on the unit sphere.
- The bound for all .
- Sufficiency: Axiom 6 gives a maximizer.
Significance
Lemma 3 tells the practitioner when the conditional logit maximum likelihood estimate exists, before any numerical optimization is attempted. It is a linear-inequality condition on the data alone, so it can be checked by linear or quadratic programming (Lemma 4 of the same paper). The existence of the estimator is also the first step of McFadden's asymptotic theory: Lemma 5 shows that Axiom 6 holds with probability tending to one, and Lemma 6, consistency and asymptotic normality, concerns the estimator whose existence Lemma 3 characterizes. The concavity and Hessian formulas (19)–(20) are the basis of the Newton–Raphson computation of the estimator and of its asymptotic covariance matrix.
The result has been proved since 1974 and is classical. To our knowledge it has no machine-checked proof; Mathlib has no statement about the existence of logit or softmax-regression maximum likelihood estimates. Formalizing it produces a verified existence criterion for the multinomial logit likelihood, verified gradient and Hessian formulas for log-sum-exp likelihoods with repeated observations, and a verified link between full column rank and strict concavity.
Difficulty
The likelihood is concave, and concave functions on need not attain their supremum. Concavity alone therefore gives nothing, and existence must come from a growth condition. The obvious approach, "the likelihood is bounded above by , hence attains its maximum", fails: is bounded but can approach its supremum only at infinity, which is exactly the separation case. Sufficiency needs a quantitative rate at which decreases, uniform over all directions; a direction-by-direction argument does not suffice. Necessity requires strict concavity, which is where Axiom 5 and the requirement that every trial be observed enter. A trial with can supply the rank of Axiom 5 while contributing nothing to , so with such a trial necessity fails. The calculus part, (19)–(20), involves differentiating sums of log-sum-exp terms over dependent index types and identifying the result with a weighted covariance operator.
Formalization scope
- Representation. is
EuclideanSpace ℝ (Fin K), so is the Euclidean norm and is the inner product⟪z, θ⟫. Trials areFin N, alternatives of trial areFin (J n), and the counts are natural numbers. - Data structure. The structure
Data Kbundles , , , and the standing assumptions and for every trial; these make the trial and alternative index sets nonempty. - Axioms 1–4 are built in. The model is the logit form (16) with linear in (Axiom 4), so "Suppose Axioms 1–5 hold" becomes "
Dataplus Axiom 5". - Axiom 5 is read at every . The row space of the matrix does not depend on .
- Hessian. The Hessian is the Fréchet derivative of the gradient vector field (19), as a continuous linear map.
- The maximizer is global over all of . Neither a local maximizer nor "" is acceptable as the goal; that would make it trivial.
- Infrastructure. Gradients and Hessians of log-sum-exp with dependent finite index types; positive definiteness from full column rank; attainment of the maximum of a coercive continuous function on a finite-dimensional space. The calculus lemmas are reusable for any multinomial logit or softmax likelihood. Missions 4 and 5 of this series reuse the same model. Contributions of general log-sum-exp lemmas, independent of this mission's definitions, are welcome.
Selected references
- D. McFadden, Conditional logit analysis of qualitative choice behavior, in P. Zarembka (ed.), Frontiers in Econometrics, Academic Press, New York, 1974, pp. 105–142. https://eml.berkeley.edu/reprints/mcfadden/zarembka.pdf
- A. Albert and J. A. Anderson, On the existence of maximum likelihood estimates in logistic regression models, Biometrika 71(1), 1984, pp. 1–10. https://doi.org/10.1093/biomet/71.1.1
- S. J. Haberman, The Analysis of Frequency Data, University of Chicago Press, 1974.
- J. Berkson, Maximum likelihood and minimum χ² estimates of the logistic function, Journal of the American Statistical Association 50, 1955, pp. 130–162. https://doi.org/10.1080/01621459.1955.10501255