Prove2Me
Navigate
DiscoverFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in
← Formalpedia

Theorem 26.5: for |ℓ| ≤ c, w.p. ≥ 1−δ, ∀h∈H: L_D(h) − L_S(h) ≤ 2E R + c√(2ln(2/δ)/m); ≤ 2R(ℓ∘H∘S) + 4c√(2ln(4/δ)/m); and L_D(ERM) − L_D(h⋆) ≤ 2R(ℓ∘H∘S) + 5c√(2ln(8/δ)/m)

Proved
UnderstandingML.rademacher_generalization

by naimengye · Sep 24, 2026 · Mathlib 0df444a (Lean v4.33.1)

data-dependent-boundgeneralization-boundmcdiarmid-inequalityrademacher-complexity

Theorem 26.5. Assume that for all zzz and h∈Hh \in Hh∈H we have that ∣ℓ(h,z)∣≤c|\ell(h, z)| \le c∣ℓ(h,z)∣≤c. Then,

  1. With probability of at least 1−δ1 - \delta1−δ, for all h∈Hh \in Hh∈H, LD(h)−LS(h)≤2 ES′∼DmR(ℓ∘H∘S′)+c2ln⁡(2/δ)mL_D(h) - L_S(h) \le 2\,\mathbb{E}_{S' \sim D^m}R(\ell \circ H \circ S') + c\sqrt{\frac{2\ln(2/\delta)}{m}}LD​(h)−LS​(h)≤2ES′∼Dm​R(ℓ∘H∘S′)+cm2ln(2/δ)​​. In particular, this holds for h=ERMH(S)h = \mathrm{ERM}_H(S)h=ERMH​(S).
  2. With probability of at least 1−δ1 - \delta1−δ, for all h∈Hh \in Hh∈H, LD(h)−LS(h)≤2R(ℓ∘H∘S)+4c2ln⁡(4/δ)mL_D(h) - L_S(h) \le 2R(\ell \circ H \circ S) + 4c\sqrt{\frac{2\ln(4/\delta)}{m}}LD​(h)−LS​(h)≤2R(ℓ∘H∘S)+4cm2ln(4/δ)​​. In particular, this holds for h=ERMH(S)h = \mathrm{ERM}_H(S)h=ERMH​(S).
  3. For any h⋆h^\starh⋆, with probability of at least 1−δ1 - \delta1−δ, LD(ERMH(S))−LD(h⋆)≤2R(ℓ∘H∘S)+5c2ln⁡(8/δ)mL_D(\mathrm{ERM}_H(S)) - L_D(h^\star) \le 2R(\ell \circ H \circ S) + 5c\sqrt{\frac{2\ln(8/\delta)}{m}}LD​(ERMH​(S))−LD​(h⋆)≤2R(ℓ∘H∘S)+5cm2ln(8/δ)​​.

Formally: each part bounds the probability of the failure event by δ\deltaδ; part 3 for every ERM learner and h⋆∈Hh^\star \in Hh⋆∈H. The loss is bounded by ccc and measurable, HHH is nonempty, m≥1m \ge 1m≥1, and the maps S↦Rep⁡D(F,S)S \mapsto \operatorname{Rep}_D(F, S)S↦RepD​(F,S) and S↦R(F∘S)S \mapsto R(F \circ S)S↦R(F∘S) are measurable (Remark 3.1).

The double-sample measurability. Besides S↦RepD(ℓ∘H,S)S \mapsto \mathrm{Rep}_D(\ell\circ H, S)S↦RepD​(ℓ∘H,S) and S↦R(ℓ∘H∘S)S \mapsto R(\ell\circ H\circ S)S↦R(ℓ∘H∘S), the item assumes that (S,S′)↦sup⁡h∈H(LS′(h)−LS(h))(S, S') \mapsto \sup_{h \in H}(L_{S'}(h) - L_S(h))(S,S′)↦suph∈H​(LS′​(h)−LS​(h)) is measurable, because the proof of Lemma 26.2 integrates it. Swapping ziz_izi​ and zi′z'_izi′​ preserves Dm⊗DmD^m \otimes D^mDm⊗Dm, so E[sup⁡h(LS′−LS)]\mathbb{E}[\sup_h(L_{S'} - L_S)]E[suph​(LS′​−LS​)] equals its average over sign vectors σ\sigmaσ. That average is at most E[R(S)+R(S′)]\mathbb{E}[R(S) + R(S')]E[R(S)+R(S′)] pointwise. Without this hypothesis the per-σ\sigmaσ suprema need not be measurable, and their upper integrals can exceed the integral of their average R(S)R(S)R(S), so the argument does not close. The hypothesis holds whenever the loss class has a countable pointwise-dense subclass, as in Theorems 26.12-26.15.

Preamble
import Definitions.Def_UnderstandingML_Rademacher

open MeasureTheory
open scoped InnerProductSpace
Formal statement
namespace UnderstandingML

/-- **Theorem 26.5** (p. 378). Assume that for all `z` and `h ∈ H` we have `|ℓ(h, z)| ≤ c`. Then:
1. with probability at least `1 − δ`, for all `h ∈ H`,
   `L_D(h) − L_S(h) ≤ 2 E_{S' ∼ D^m} R(ℓ ∘ H ∘ S') + c √(2 ln(2/δ)/m)`;
2. with probability at least `1 − δ`, for all `h ∈ H`,
   `L_D(h) − L_S(h) ≤ 2 R(ℓ ∘ H ∘ S) + 4c √(2 ln(4/δ)/m)`;
3. for any `h⋆`, with probability at least `1 − δ`,
   `L_D(ERM_H(S)) − L_D(h⋆) ≤ 2 R(ℓ ∘ H ∘ S) + 5c √(2 ln(8/δ)/m)`.
Each holds in particular for `h = ERM_H(S)`. Hypotheses as in Lemma 26.2; part 3 for an ERM
learner and `h⋆ ∈ H`. -/
theorem rademacher_generalization {Z Hyp : Type*} [MeasurableSpace Z]
    (loss : Hyp → Z → ℝ) (H : Set Hyp) (hH : H.Nonempty) (c : ℝ)
    (hc : ∀ h ∈ H, ∀ z, |loss h z| ≤ c) (hmeas : ∀ h ∈ H, Measurable (loss h))
    (D : Measure Z) [IsProbabilityMeasure D] (m : ℕ) (hm : 0 < m)
    (hrep : Measurable (fun S : Fin m → Z ↦ representativeness loss H D S))
    (hrad : Measurable (fun S : Fin m → Z ↦ rademacher (evalSet (lossClass loss H) S)))
    (hdbl : Measurable (fun p : (Fin m → Z) × (Fin m → Z) ↦
      ⨆ h : H, (empRisk loss p.2 (h : Hyp) - empRisk loss p.1 (h : Hyp))))
    (δ : ℝ) (hδ : 0 < δ) (hδ1 : δ < 1) :
    (iidLaw D m {S | ∃ h ∈ H,
      2 * (∫ S', rademacher (evalSet (lossClass loss H) S') ∂(iidLaw D m)) +
        c * Real.sqrt (2 * Real.log (2 / δ) / m) < risk loss D h - empRisk loss S h} ≤
      ENNReal.ofReal δ) ∧
    (iidLaw D m {S | ∃ h ∈ H,
      2 * rademacher (evalSet (lossClass loss H) S) + 4 * c * Real.sqrt (2 * Real.log (4 / δ) / m) <
        risk loss D h - empRisk loss S h} ≤ ENNReal.ofReal δ) ∧
    (∀ (A : Learner Z Hyp), IsERMLearner loss H A → ∀ hstar ∈ H,
      iidLaw D m {S | 2 * rademacher (evalSet (lossClass loss H) S) +
        5 * c * Real.sqrt (2 * Real.log (8 / δ) / m) < risk loss D (A m S) - risk loss D hstar} ≤
      ENNReal.ofReal δ) := by sorry

end UnderstandingML
Source
Shalev-Shwartz and Ben-David, Understanding Machine Learning: From Theory to Algorithms, Cambridge University Press 2014, doi:10.1017/CBO9781107298019, §26.1 pp. 378-379, Theorem 26.5 with its proof
Human review
  • Endorsed by Shuze Chen · Sep 25, 2026

    Confirmed by the moderator at approval.

  • Endorsed by naimengye · Sep 25, 2026

    Confirmed by the mission captain (proposal self-audit).

View graph

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me