Prove2Me
Navigate
DiscoverFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in
← Formalpedia

Proposition A.20 - Perfect-feature LP recovery

Proved
FeatureDistortion.PerfectFeatureLinearProbing

by Minghui · Sep 26, 2026 · Mathlib c5ea003 (Lean v4.30.0)

linear-algebramachine-learningprobability

Notation: n=#{training examples}n = \#\{\text{training examples}\}n=#{training examples}, ddd is the input dimension, kkk the feature dimension, X:Rd→RnX:\mathbb R^d\to\mathbb R^nX:Rd→Rn the data map, YYY the labels, B:Rd→RkB:\mathbb R^d\to\mathbb R^kB:Rd→Rk the features, and v∈Rkv\in\mathbb R^kv∈Rk the head. Adjoint means Euclidean transpose. The loss is L^(v,B)=∥XB⊤v−Y∥2\widehat L(v,B)=\|XB^\top v-Y\|^2L(v,B)=∥XB⊤v−Y∥2, with no normalization. The probability model, when present, is explicitly specified below; deterministic flow statements involve no random data assumption.

For every triple of natural numbers n,d,kn,d,kn,d,k and every choice of continuous real-linear maps X:Rd→RnX:\mathbb R^d\to\mathbb R^nX:Rd→Rn and B⋆:Rd→RkB_\star:\mathbb R^d\to\mathbb R^kB⋆​:Rd→Rk, real-linear isometric bijection R:Rk→RkR:\mathbb R^k\to\mathbb R^kR:Rk→Rk, and u⋆∈Rku_\star\in\mathbb R^ku⋆​∈Rk, put B0=RB⋆B_0=RB_\starB0​=RB⋆​, a⋆=Ru⋆a_\star=Ru_\stara⋆​=Ru⋆​, w⋆=B⋆∗u⋆w_\star=B_\star^*u_\starw⋆​=B⋆∗​u⋆​, Y=Xw⋆Y=Xw_\starY=Xw⋆​, S={X∗z:z∈Rn}S=\{X^*z:z\in\mathbb R^n\}S={X∗z:z∈Rn}, and r=dim⁡RSr=\dim_{\mathbb R}Sr=dimR​S. Assume 0<k0<k0<k, k≤rk\leq rk≤r, r+k<dr+k<dr+k<d, B⋆B⋆∗=IRkB_\star B_\star^*=I_{\mathbb R^k}B⋆​B⋆∗​=IRk​, u⋆≠0u_\star\neq0u⋆​=0, and injectivity on Rk\mathbb R^kRk of both v↦ΠS(B0∗v)v\mapsto\Pi_S(B_0^*v)v↦ΠS​(B0∗​v) and v↦ΠS⊥(B0∗v)v\mapsto\Pi_{S^\perp}(B_0^*v)v↦ΠS⊥​(B0∗​v), where Π\PiΠ denotes orthogonal projection. Then two assertions hold: for every v∈Rkv\in\mathbb R^kv∈Rk, ∥XB0∗v−Y∥2=0\|XB_0^*v-Y\|^2=0∥XB0∗​v−Y∥2=0 if and only if v=a⋆v=a_\starv=a⋆​; and for every v0∈Rkv_0\in\mathbb R^kv0​∈Rk and every function b:R→Rkb:\mathbb R\to\mathbb R^kb:R→Rk with b(0)=v0b(0)=v_0b(0)=v0​ and derivative within [0,∞)[0,\infty)[0,∞) equal to −2B0X∗(XB0∗b(t)−Y)-2B_0X^*(XB_0^*b(t)-Y)−2B0​X∗(XB0∗​b(t)−Y) at every real t≥0t\geq0t≥0, b(t)→a⋆b(t)\to a_\starb(t)→a⋆​ in the Euclidean topology as t→+∞t\to+\inftyt→+∞. All adjoints are Euclidean. The hypotheses exclude zero dimensions and require n≥kn\geq kn≥k and d≥2k+1d\geq2k+1d≥2k+1; if there are no qualifying data the implication is vacuous. Existence of bbb is not asserted here, and its values at negative times are unconstrained.

Formalization note: Source-derived Proposition A.20 and its rotated-feature extension; convergence is a conclusion, with explicit identifiability assumptions. Source: Kumar, Raghunathan, Jones, Ma, and Liang, Fine-Tuning can Distort Pretrained Features and Underperform Out-of-Distribution, ICLR 2022, https://arxiv.org/pdf/2202.10054v1. Appendix A.7, PDF pp. 45--46, Proposition A.20, equations (A.208)--(A.211), and the following rotation paragraph on PDF p. 46. Source-backed parent: Section 3.4, PDF p. 10, Proposition 3.7, equations (3.10)--(3.11); Appendix A.7, PDF pp. 45--47.

Preamble
import Definitions.Def_FeatureDistortion_Model
open MeasureTheory Filter
open scoped Topology
Formal statement
namespace FeatureDistortion
theorem PerfectFeatureLinearProbing :
  ∀ (n d k : ℕ) (P : Problem n d k), Admissible P →
    (∀ v : Vec k,
      trainingLoss P.data (labels P) v (initialFeatures P) = 0 ↔ v = alignedHead P) ∧
    (∀ (v₀ : Vec k) (v : ℝ → Vec k),
      IsLinearProbingFlow P.data (labels P) v₀ (initialFeatures P) v →
      Tendsto v atTop (𝓝 (alignedHead P))) := by sorry
end FeatureDistortion
Source
Kumar, Raghunathan, Jones, Ma, and Liang, Fine-Tuning can Distort Pretrained Features and Underperform Out-of-Distribution, ICLR 2022, https://arxiv.org/pdf/2202.10054v1. Appendix A.7, PDF pp. 45--46, Proposition A.20, equations (A.208)--(A.211), and the following rotation paragraph on PDF p. 46. Source-backed parent: Section 3.4, PDF p. 10, Proposition 3.7, equations (3.10)--(3.11); Appendix A.7, PDF pp. 45--47.
Read-back

What the Lean code literally says, in plain math · gpt-6

For every triple of natural numbers n,d,kn,d,kn,d,k and every choice of continuous real-linear maps X:Rd→RnX:\mathbb R^d\to\mathbb R^nX:Rd→Rn and B⋆:Rd→RkB_\star:\mathbb R^d\to\mathbb R^kB⋆​:Rd→Rk, real-linear isometric bijection R:Rk→RkR:\mathbb R^k\to\mathbb R^kR:Rk→Rk, and u⋆∈Rku_\star\in\mathbb R^ku⋆​∈Rk, put B0=RB⋆B_0=RB_\starB0​=RB⋆​, a⋆=Ru⋆a_\star=Ru_\stara⋆​=Ru⋆​, w⋆=B⋆∗u⋆w_\star=B_\star^*u_\starw⋆​=B⋆∗​u⋆​, Y=Xw⋆Y=Xw_\starY=Xw⋆​, S={X∗z:z∈Rn}S=\{X^*z:z\in\mathbb R^n\}S={X∗z:z∈Rn}, and r=dim⁡RSr=\dim_{\mathbb R}Sr=dimR​S. Assume 0<k0<k0<k, k≤rk\leq rk≤r, r+k<dr+k<dr+k<d, B⋆B⋆∗=IRkB_\star B_\star^*=I_{\mathbb R^k}B⋆​B⋆∗​=IRk​, u⋆≠0u_\star\neq0u⋆​=0, and injectivity on Rk\mathbb R^kRk of both v↦ΠS(B0∗v)v\mapsto\Pi_S(B_0^*v)v↦ΠS​(B0∗​v) and v↦ΠS⊥(B0∗v)v\mapsto\Pi_{S^\perp}(B_0^*v)v↦ΠS⊥​(B0∗​v), where Π\PiΠ denotes orthogonal projection. Then two assertions hold: for every v∈Rkv\in\mathbb R^kv∈Rk, ∥XB0∗v−Y∥2=0\|XB_0^*v-Y\|^2=0∥XB0∗​v−Y∥2=0 if and only if v=a⋆v=a_\starv=a⋆​; and for every v0∈Rkv_0\in\mathbb R^kv0​∈Rk and every function b:R→Rkb:\mathbb R\to\mathbb R^kb:R→Rk with b(0)=v0b(0)=v_0b(0)=v0​ and derivative within [0,∞)[0,\infty)[0,∞) equal to −2B0X∗(XB0∗b(t)−Y)-2B_0X^*(XB_0^*b(t)-Y)−2B0​X∗(XB0∗​b(t)−Y) at every real t≥0t\geq0t≥0, b(t)→a⋆b(t)\to a_\starb(t)→a⋆​ in the Euclidean topology as t→+∞t\to+\inftyt→+∞. All adjoints are Euclidean. The hypotheses exclude zero dimensions and require n≥kn\geq kn≥k and d≥2k+1d\geq2k+1d≥2k+1; if there are no qualifying data the implication is vacuous. Existence of bbb is not asserted here, and its values at negative times are unconstrained.

Human review
  • Endorsed by Shuze Chen · Sep 27, 2026

    Confirmed by the moderator at approval.

  • Endorsed by Minghui · Sep 27, 2026

    Confirmed by the mission captain (proposal self-audit).

View graph

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me