Prove2Me
Navigate
DiscoverFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in
← Formalpedia

Proposition 3.7 - Perfect-feature LP-FT separation

Proved
FeatureDistortion.PerfectFeatureLPFTSeparation

by Minghui · Sep 26, 2026 · Mathlib c5ea003 (Lean v4.30.0)

linear-algebramachine-learningprobability

Notation: n=#{training examples}n = \#\{\text{training examples}\}n=#{training examples}, ddd is the input dimension, kkk the feature dimension, X:Rd→RnX:\mathbb R^d\to\mathbb R^nX:Rd→Rn the data map, YYY the labels, B:Rd→RkB:\mathbb R^d\to\mathbb R^kB:Rd→Rk the features, and v∈Rkv\in\mathbb R^kv∈Rk the head. Adjoint means Euclidean transpose. The loss is L^(v,B)=∥XB⊤v−Y∥2\widehat L(v,B)=\|XB^\top v-Y\|^2L(v,B)=∥XB⊤v−Y∥2, with no normalization. The probability model, when present, is explicitly specified below; deterministic flow statements involve no random data assumption.

For every triple of natural numbers n,d,kn,d,kn,d,k, every choice of continuous real-linear maps X:Rd→RnX:\mathbb R^d\to\mathbb R^nX:Rd→Rn and B⋆:Rd→RkB_\star:\mathbb R^d\to\mathbb R^kB⋆​:Rd→Rk, real-linear isometric bijection R:Rk→RkR:\mathbb R^k\to\mathbb R^kR:Rk→Rk, and u⋆∈Rku_\star\in\mathbb R^ku⋆​∈Rk, and every probability measure μ\muμ on Rd\mathbb R^dRd with integrable squared norm and ⟨z,Σμz⟩>0\langle z,\Sigma_\mu z\rangle>0⟨z,Σμ​z⟩>0 for every nonzero z∈Rdz\in\mathbb R^dz∈Rd, where Σμ=∫(x⊗x) dμ(x)\Sigma_\mu=\int(x\otimes x)\,d\mu(x)Σμ​=∫(x⊗x)dμ(x) and (x⊗x)(z)=⟨x,z⟩x(x\otimes x)(z)=\langle x,z\rangle x(x⊗x)(z)=⟨x,z⟩x, let B0=RB⋆B_0=RB_\starB0​=RB⋆​, a⋆=Ru⋆a_\star=Ru_\stara⋆​=Ru⋆​, w⋆=B⋆∗u⋆w_\star=B_\star^*u_\starw⋆​=B⋆∗​u⋆​, Y=Xw⋆Y=Xw_\starY=Xw⋆​, S={X∗z:z∈Rn}S=\{X^*z:z\in\mathbb R^n\}S={X∗z:z∈Rn}, and r=dim⁡RSr=\dim_{\mathbb R}Sr=dimR​S. Assume 0<k0<k0<k, k≤rk\leq rk≤r, r+k<dr+k<dr+k<d, B⋆B⋆∗=IRkB_\star B_\star^*=I_{\mathbb R^k}B⋆​B⋆∗​=IRk​, u⋆≠0u_\star\neq0u⋆​=0, and injectivity on Rk\mathbb R^kRk of both v↦ΠS(B0∗v)v\mapsto\Pi_S(B_0^*v)v↦ΠS​(B0∗​v) and v↦ΠS⊥(B0∗v)v\mapsto\Pi_{S^\perp}(B_0^*v)v↦ΠS⊥​(B0∗​v), with Π\PiΠ denoting orthogonal projection. For every real σ>0\sigma>0σ>0, four assertions hold together. Here a fine-tuning pair from v0v_0v0​ means functions a:R→Rka:\mathbb R\to\mathbb R^ka:R→Rk and F:R→L(Rd,Rk)F:\mathbb R\to\mathcal L(\mathbb R^d,\mathbb R^k)F:R→L(Rd,Rk) with a(0)=v0a(0)=v_0a(0)=v0​, F(0)=B0F(0)=B_0F(0)=B0​, and, at every real t≥0t\geq0t≥0, derivatives within [0,∞)[0,\infty)[0,∞) equal to a˙(t)=−2F(t)X∗(XF(t)∗a(t)−Y)\dot a(t)=-2F(t)X^*(XF(t)^*a(t)-Y)a˙(t)=−2F(t)X∗(XF(t)∗a(t)−Y) and F˙(t)=[x↦−2⟨X∗(XF(t)∗a(t)−Y),x⟩a(t)]\dot F(t)=\bigl[x\mapsto-2\langle X^*(XF(t)^*a(t)-Y),x\rangle a(t)\bigr]F˙(t)=[x↦−2⟨X∗(XF(t)∗a(t)−Y),x⟩a(t)], the latter being a derivative in the space of continuous linear maps. A probing curve from v0v_0v0​ means b:R→Rkb:\mathbb R\to\mathbb R^kb:R→Rk with b(0)=v0b(0)=v_0b(0)=v0​ and derivative within [0,∞)[0,\infty)[0,∞) equal to −2B0X∗(XB0∗b(t)−Y)-2B_0X^*(XB_0^*b(t)-Y)−2B0​X∗(XB0∗​b(t)−Y) at every real t≥0t\geq0t≥0. The four assertions are: (1) for every v0∈Rkv_0\in\mathbb R^kv0​∈Rk there exists a fine-tuning pair from v0v_0v0​; (2) for every v0∈Rkv_0\in\mathbb R^kv0​∈Rk there exists a probing curve from v0v_0v0​, and every probing curve from that v0v_0v0​ tends to a⋆a_\stara⋆​ in the Euclidean topology as t→+∞t\to+\inftyt→+∞; (3) for every fine-tuning pair from a⋆a_\stara⋆​ and every real t≥0t\geq0t≥0, ∫⟨F(t)∗a(t)−w⋆,x⟩2 dμ(x)=0\int\langle F(t)^*a(t)-w_\star,x\rangle^2\,d\mu(x)=0∫⟨F(t)∗a(t)−w⋆​,x⟩2dμ(x)=0; and (4) for almost every v0v_0v0​ under the pushforward of standard Gaussian measure on Rk\mathbb R^kRk by z↦σzz\mapsto\sigma zz↦σz, every fine-tuning pair from that v0v_0v0​ satisfies ∫⟨F(t)∗a(t)−w⋆,x⟩2 dμ(x)>0\int\langle F(t)^*a(t)-w_\star,x\rangle^2\,d\mu(x)>0∫⟨F(t)∗a(t)−w⋆​,x⟩2dμ(x)>0 for every real t≥0t\geq0t≥0. All adjoints are Euclidean. The exceptional null set in (4) is chosen before the universal quantifiers over pairs and times; this includes time zero and asserts positivity at each nonnegative real time, without a uniform positive lower bound or a statement about a limiting loss. The admissibility hypotheses exclude zero dimensions and require n≥kn\geq kn≥k and d≥2k+1d\geq2k+1d≥2k+1. No conclusion is required for σ≤0\sigma\leq0σ≤0, and the curves have no conditions at negative times.

Formalization note: Source-derived Proposition 3.7 in the explicit nonzero-signal, identifiable perfect-feature regime, including flow existence and LP convergence. No uniform time-infimum, quantitative Theorem 3.3 constant, or imperfect-feature guarantee is claimed. Source: Kumar, Raghunathan, Jones, Ma, and Liang, Fine-Tuning can Distort Pretrained Features and Underperform Out-of-Distribution, ICLR 2022, https://arxiv.org/pdf/2202.10054v1. Section 3.4, PDF p. 10, Proposition 3.7, equations (3.10)--(3.11); Appendix A.7, PDF pp. 45--47, equations (A.208)--(A.218). Source-backed parent: Section 3.4, PDF p. 10, Proposition 3.7, equations (3.10)--(3.11); Appendix A.7, PDF pp. 45--47.

Preamble
import Definitions.Def_FeatureDistortion_Model
open MeasureTheory Filter
open scoped Topology
Formal statement
namespace FeatureDistortion
theorem PerfectFeatureLPFTSeparation :
  ∀ (n d k : ℕ) (P : Problem n d k) (D : OODLaw d), Admissible P →
    ∀ σ : ℝ, 0 < σ →
      (∀ v₀ : Vec k, ∃ γ : Trajectory d k,
        IsFineTuningFlow P.data (labels P) v₀ (initialFeatures P) γ) ∧
      (∀ v₀ : Vec k,
        (∃ v : ℝ → Vec k, IsLinearProbingFlow P.data (labels P) v₀ (initialFeatures P) v) ∧
        ∀ v : ℝ → Vec k,
          IsLinearProbingFlow P.data (labels P) v₀ (initialFeatures P) v →
          Tendsto v atTop (𝓝 (alignedHead P))) ∧
      (∀ γ : Trajectory d k,
        IsFineTuningFlow P.data (labels P) (alignedHead P) (initialFeatures P) γ →
        ∀ t : ℝ, 0 ≤ t →
          oodLoss D (targetWeights P) (γ.head t) (γ.features t) = 0) ∧
      (∀ᵐ v₀ ∂gaussianHead k σ, ∀ γ : Trajectory d k,
        IsFineTuningFlow P.data (labels P) v₀ (initialFeatures P) γ →
        ∀ t : ℝ, 0 ≤ t →
          0 < oodLoss D (targetWeights P) (γ.head t) (γ.features t)) := by sorry
end FeatureDistortion
Source
Kumar, Raghunathan, Jones, Ma, and Liang, Fine-Tuning can Distort Pretrained Features and Underperform Out-of-Distribution, ICLR 2022, https://arxiv.org/pdf/2202.10054v1. Section 3.4, PDF p. 10, Proposition 3.7, equations (3.10)--(3.11); Appendix A.7, PDF pp. 45--47, equations (A.208)--(A.218). Source-backed parent: Section 3.4, PDF p. 10, Proposition 3.7, equations (3.10)--(3.11); Appendix A.7, PDF pp. 45--47.
Read-back

What the Lean code literally says, in plain math · gpt-6

For every triple of natural numbers n,d,kn,d,kn,d,k, every choice of continuous real-linear maps X:Rd→RnX:\mathbb R^d\to\mathbb R^nX:Rd→Rn and B⋆:Rd→RkB_\star:\mathbb R^d\to\mathbb R^kB⋆​:Rd→Rk, real-linear isometric bijection R:Rk→RkR:\mathbb R^k\to\mathbb R^kR:Rk→Rk, and u⋆∈Rku_\star\in\mathbb R^ku⋆​∈Rk, and every probability measure μ\muμ on Rd\mathbb R^dRd with integrable squared norm and ⟨z,Σμz⟩>0\langle z,\Sigma_\mu z\rangle>0⟨z,Σμ​z⟩>0 for every nonzero z∈Rdz\in\mathbb R^dz∈Rd, where Σμ=∫(x⊗x) dμ(x)\Sigma_\mu=\int(x\otimes x)\,d\mu(x)Σμ​=∫(x⊗x)dμ(x) and (x⊗x)(z)=⟨x,z⟩x(x\otimes x)(z)=\langle x,z\rangle x(x⊗x)(z)=⟨x,z⟩x, let B0=RB⋆B_0=RB_\starB0​=RB⋆​, a⋆=Ru⋆a_\star=Ru_\stara⋆​=Ru⋆​, w⋆=B⋆∗u⋆w_\star=B_\star^*u_\starw⋆​=B⋆∗​u⋆​, Y=Xw⋆Y=Xw_\starY=Xw⋆​, S={X∗z:z∈Rn}S=\{X^*z:z\in\mathbb R^n\}S={X∗z:z∈Rn}, and r=dim⁡RSr=\dim_{\mathbb R}Sr=dimR​S. Assume 0<k0<k0<k, k≤rk\leq rk≤r, r+k<dr+k<dr+k<d, B⋆B⋆∗=IRkB_\star B_\star^*=I_{\mathbb R^k}B⋆​B⋆∗​=IRk​, u⋆≠0u_\star\neq0u⋆​=0, and injectivity on Rk\mathbb R^kRk of both v↦ΠS(B0∗v)v\mapsto\Pi_S(B_0^*v)v↦ΠS​(B0∗​v) and v↦ΠS⊥(B0∗v)v\mapsto\Pi_{S^\perp}(B_0^*v)v↦ΠS⊥​(B0∗​v), with Π\PiΠ denoting orthogonal projection. For every real σ>0\sigma>0σ>0, four assertions hold together. Here a fine-tuning pair from v0v_0v0​ means functions a:R→Rka:\mathbb R\to\mathbb R^ka:R→Rk and F:R→L(Rd,Rk)F:\mathbb R\to\mathcal L(\mathbb R^d,\mathbb R^k)F:R→L(Rd,Rk) with a(0)=v0a(0)=v_0a(0)=v0​, F(0)=B0F(0)=B_0F(0)=B0​, and, at every real t≥0t\geq0t≥0, derivatives within [0,∞)[0,\infty)[0,∞) equal to a˙(t)=−2F(t)X∗(XF(t)∗a(t)−Y)\dot a(t)=-2F(t)X^*(XF(t)^*a(t)-Y)a˙(t)=−2F(t)X∗(XF(t)∗a(t)−Y) and F˙(t)=[x↦−2⟨X∗(XF(t)∗a(t)−Y),x⟩a(t)]\dot F(t)=\bigl[x\mapsto-2\langle X^*(XF(t)^*a(t)-Y),x\rangle a(t)\bigr]F˙(t)=[x↦−2⟨X∗(XF(t)∗a(t)−Y),x⟩a(t)], the latter being a derivative in the space of continuous linear maps. A probing curve from v0v_0v0​ means b:R→Rkb:\mathbb R\to\mathbb R^kb:R→Rk with b(0)=v0b(0)=v_0b(0)=v0​ and derivative within [0,∞)[0,\infty)[0,∞) equal to −2B0X∗(XB0∗b(t)−Y)-2B_0X^*(XB_0^*b(t)-Y)−2B0​X∗(XB0∗​b(t)−Y) at every real t≥0t\geq0t≥0. The four assertions are: (1) for every v0∈Rkv_0\in\mathbb R^kv0​∈Rk there exists a fine-tuning pair from v0v_0v0​; (2) for every v0∈Rkv_0\in\mathbb R^kv0​∈Rk there exists a probing curve from v0v_0v0​, and every probing curve from that v0v_0v0​ tends to a⋆a_\stara⋆​ in the Euclidean topology as t→+∞t\to+\inftyt→+∞; (3) for every fine-tuning pair from a⋆a_\stara⋆​ and every real t≥0t\geq0t≥0, ∫⟨F(t)∗a(t)−w⋆,x⟩2 dμ(x)=0\int\langle F(t)^*a(t)-w_\star,x\rangle^2\,d\mu(x)=0∫⟨F(t)∗a(t)−w⋆​,x⟩2dμ(x)=0; and (4) for almost every v0v_0v0​ under the pushforward of standard Gaussian measure on Rk\mathbb R^kRk by z↦σzz\mapsto\sigma zz↦σz, every fine-tuning pair from that v0v_0v0​ satisfies ∫⟨F(t)∗a(t)−w⋆,x⟩2 dμ(x)>0\int\langle F(t)^*a(t)-w_\star,x\rangle^2\,d\mu(x)>0∫⟨F(t)∗a(t)−w⋆​,x⟩2dμ(x)>0 for every real t≥0t\geq0t≥0. All adjoints are Euclidean. The exceptional null set in (4) is chosen before the universal quantifiers over pairs and times; this includes time zero and asserts positivity at each nonnegative real time, without a uniform positive lower bound or a statement about a limiting loss. The admissibility hypotheses exclude zero dimensions and require n≥kn\geq kn≥k and d≥2k+1d\geq2k+1d≥2k+1. No conclusion is required for σ≤0\sigma\leq0σ≤0, and the curves have no conditions at negative times.

Human review
  • Endorsed by Shuze Chen · Sep 27, 2026

    Confirmed by the moderator at approval.

  • Endorsed by Minghui · Sep 27, 2026

    Confirmed by the mission captain (proposal self-audit).

View graph

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me