Prove2Me
Navigate
DiscoverFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in
← Formalpedia

Lemma A.7 - OOD risk and the uncentered second moment

Proved
FeatureDistortion.OODRiskIdentity

by Minghui · Sep 26, 2026 · Mathlib c5ea003 (Lean v4.30.0)

linear-algebramachine-learningprobability

Notation: n=#{training examples}n = \#\{\text{training examples}\}n=#{training examples}, ddd is the input dimension, kkk the feature dimension, X:Rd→RnX:\mathbb R^d\to\mathbb R^nX:Rd→Rn the data map, YYY the labels, B:Rd→RkB:\mathbb R^d\to\mathbb R^kB:Rd→Rk the features, and v∈Rkv\in\mathbb R^kv∈Rk the head. Adjoint means Euclidean transpose. The loss is L^(v,B)=∥XB⊤v−Y∥2\widehat L(v,B)=\|XB^\top v-Y\|^2L(v,B)=∥XB⊤v−Y∥2, with no normalization. The probability model, when present, is explicitly specified below; deterministic flow statements involve no random data assumption.

For every pair of natural numbers d,kd,kd,k, every probability measure μ\muμ on Rd\mathbb R^dRd with integrable squared norm and ⟨z,Σμz⟩>0\langle z,\Sigma_\mu z\rangle>0⟨z,Σμ​z⟩>0 for every nonzero z∈Rdz\in\mathbb R^dz∈Rd, every w∈Rdw\in\mathbb R^dw∈Rd, every v∈Rkv\in\mathbb R^kv∈Rk, and every continuous real-linear map B:Rd→RkB:\mathbb R^d\to\mathbb R^kB:Rd→Rk, let Σμ=∫(x⊗x) dμ(x)\Sigma_\mu=\int(x\otimes x)\,d\mu(x)Σμ​=∫(x⊗x)dμ(x), where (x⊗x)(z)=⟨x,z⟩x(x\otimes x)(z)=\langle x,z\rangle x(x⊗x)(z)=⟨x,z⟩x, and let e=B∗v−we=B^*v-we=B∗v−w and L=∫⟨e,x⟩2 dμ(x)L=\int\langle e,x\rangle^2\,d\mu(x)L=∫⟨e,x⟩2dμ(x). The assertion is the conjunction L=⟨e,Σμe⟩L=\langle e,\Sigma_\mu e\rangleL=⟨e,Σμ​e⟩, (L=0⟺B∗v=w)(L=0\Longleftrightarrow B^*v=w)(L=0⟺B∗v=w), and (0<L⟺B∗v≠w)(0<L\Longleftrightarrow B^*v\neq w)(0<L⟺B∗v=w). The moment integral is operator-valued, the adjoint is Euclidean, and no mean-zero assumption is made. If d=0d=0d=0, the positivity hypothesis is vacuous, e=0e=0e=0, and L=0L=0L=0. If k=0k=0k=0, then B∗v=0B^*v=0B∗v=0, so the equivalences test whether w=0w=0w=0.

Formalization note: Source-derived risk identity and positive-definiteness consequences. The reversed inequality in (A.28) is excluded. Source: Kumar, Raghunathan, Jones, Ma, and Liang, Fine-Tuning can Distort Pretrained Features and Underperform Out-of-Distribution, ICLR 2022, https://arxiv.org/pdf/2202.10054v1. Appendix A.2, PDF p. 26, Lemma A.7, equations (A.29)--(A.32). Source-backed parent: Section 3.4, PDF p. 10, Proposition 3.7, equations (3.10)--(3.11); Appendix A.7, PDF pp. 45--47.

Preamble
import Definitions.Def_FeatureDistortion_Model
open MeasureTheory Filter
open scoped Topology
Formal statement
namespace FeatureDistortion
theorem OODRiskIdentity :
  ∀ (d k : ℕ) (D : OODLaw d) (w : Vec d) (v : Vec k) (B : Features d k),
    oodLoss D w v B =
      inner ℝ (weights B v - w) (secondMoment D.measure (weights B v - w)) ∧
    (oodLoss D w v B = 0 ↔ weights B v = w) ∧
    (0 < oodLoss D w v B ↔ weights B v ≠ w) := by sorry
end FeatureDistortion
Source
Kumar, Raghunathan, Jones, Ma, and Liang, Fine-Tuning can Distort Pretrained Features and Underperform Out-of-Distribution, ICLR 2022, https://arxiv.org/pdf/2202.10054v1. Appendix A.2, PDF p. 26, Lemma A.7, equations (A.29)--(A.32). Source-backed parent: Section 3.4, PDF p. 10, Proposition 3.7, equations (3.10)--(3.11); Appendix A.7, PDF pp. 45--47.
Read-back

What the Lean code literally says, in plain math · gpt-6

For every pair of natural numbers d,kd,kd,k, every probability measure μ\muμ on Rd\mathbb R^dRd with integrable squared norm and ⟨z,Σμz⟩>0\langle z,\Sigma_\mu z\rangle>0⟨z,Σμ​z⟩>0 for every nonzero z∈Rdz\in\mathbb R^dz∈Rd, every w∈Rdw\in\mathbb R^dw∈Rd, every v∈Rkv\in\mathbb R^kv∈Rk, and every continuous real-linear map B:Rd→RkB:\mathbb R^d\to\mathbb R^kB:Rd→Rk, let Σμ=∫(x⊗x) dμ(x)\Sigma_\mu=\int(x\otimes x)\,d\mu(x)Σμ​=∫(x⊗x)dμ(x), where (x⊗x)(z)=⟨x,z⟩x(x\otimes x)(z)=\langle x,z\rangle x(x⊗x)(z)=⟨x,z⟩x, and let e=B∗v−we=B^*v-we=B∗v−w and L=∫⟨e,x⟩2 dμ(x)L=\int\langle e,x\rangle^2\,d\mu(x)L=∫⟨e,x⟩2dμ(x). The assertion is the conjunction L=⟨e,Σμe⟩L=\langle e,\Sigma_\mu e\rangleL=⟨e,Σμ​e⟩, (L=0⟺B∗v=w)(L=0\Longleftrightarrow B^*v=w)(L=0⟺B∗v=w), and (0<L⟺B∗v≠w)(0<L\Longleftrightarrow B^*v\neq w)(0<L⟺B∗v=w). The moment integral is operator-valued, the adjoint is Euclidean, and no mean-zero assumption is made. If d=0d=0d=0, the positivity hypothesis is vacuous, e=0e=0e=0, and L=0L=0L=0. If k=0k=0k=0, then B∗v=0B^*v=0B∗v=0, so the equivalences test whether w=0w=0w=0.

Human review
  • Endorsed by Shuze Chen · Sep 27, 2026

    Confirmed by the moderator at approval.

  • Endorsed by Minghui · Sep 27, 2026

    Confirmed by the mission captain (proposal self-audit).

View graph

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me