Prove2Me
Navigate
DiscoverFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in
← Formalpedia

Lemma 3 — exp-concave functions lie above a gradient paraboloid

Proved
LogRegretOCO.FTAL.exp_concave_approx_lower_bound

by mikedeng1 · Sep 26, 2026 · Mathlib 0df444a (Lean v4.33.1)

convex-analysisexp-concavityonline-convex-optimizationp2o-batch-p100ap2o-gran-per-chapterp2o-plan-paperp2o-v1

Let P⊆RnP \subseteq \mathbb{R}^nP⊆Rn be convex with ∥x−y∥≤D\|x - y\| \le D∥x−y∥≤D for all x,y∈Px, y \in Px,y∈P, where D>0D > 0D>0. Let fff be a real function, differentiable at every point of PPP, with ∥∇f(x)∥≤G\|\nabla f(x)\| \le G∥∇f(x)∥≤G for x∈Px \in Px∈P (G>0G > 0G>0), and such that x↦exp⁡(−αf(x))x \mapsto \exp(-\alpha f(x))x↦exp(−αf(x)) is concave on PPP for some α>0\alpha > 0α>0. Then for every β\betaβ with 0<β≤12min⁡{14GD,α}0 < \beta \le \tfrac12 \min\{\tfrac{1}{4GD}, \alpha\}0<β≤21​min{4GD1​,α},

f(x)≥f(y)+∇f(y)⊤(x−y)+β2(x−y)⊤∇f(y)∇f(y)⊤(x−y)for all x,y∈P.f(x) \ge f(y) + \nabla f(y)^\top (x - y) + \frac{\beta}{2} (x - y)^\top \nabla f(y) \nabla f(y)^\top (x - y) \qquad \text{for all } x, y \in P.f(x)≥f(y)+∇f(y)⊤(x−y)+2β​(x−y)⊤∇f(y)∇f(y)⊤(x−y)for all x,y∈P.

The lemma says that an exp-concave function with bounded gradients is bounded below, on the whole decision set, by a rank-one quadratic that touches it at yyy. This is what lets Follow the Approximate Leader replace each cost by its quadratic model without increasing the regret.

Formalization Note The quadratic term is written as β2⟨∇f(y),x−y⟩2\frac{\beta}{2}\langle \nabla f(y), x - y\rangle^22β​⟨∇f(y),x−y⟩2, which equals β2(x−y)⊤∇f(y)∇f(y)⊤(x−y)\frac{\beta}{2}(x-y)^\top \nabla f(y)\nabla f(y)^\top(x-y)2β​(x−y)⊤∇f(y)∇f(y)⊤(x−y). The paper states β≤12min⁡{⋅}\beta \le \frac12\min\{\cdot\}β≤21​min{⋅}; the hypothesis β>0\beta > 0β>0 is added because the proof divides by β\betaβ (at β=0\beta = 0β=0 the claim would be the tangent inequality for a convex function). G>0G > 0G>0 and D>0D > 0D>0 make 1/(4GD)1/(4GD)1/(4GD) meaningful (in Lean 1/0=01/0 = 01/0=0). fff is a function on all of Rn\mathbb{R}^nRn, differentiable at the points of PPP; the diameter is used only as an upper bound.

Preamble
import Mathlib
Formal statement
namespace LogRegretOCO.FTAL
theorem exp_concave_approx_lower_bound {n : ℕ} (P : Set (EuclideanSpace ℝ (Fin n)))
    (f : EuclideanSpace ℝ (Fin n) → ℝ) (D G α β : ℝ)
    (hPconv : Convex ℝ P) (hD : 0 < D) (hG : 0 < G) (hα : 0 < α)
    (hdiam : ∀ x ∈ P, ∀ y ∈ P, ‖x - y‖ ≤ D)
    (hdiff : ∀ x ∈ P, DifferentiableAt ℝ f x)
    (hgrad : ∀ x ∈ P, ‖gradient f x‖ ≤ G)
    (hexp : ConcaveOn ℝ P (fun x => Real.exp (-α * f x)))
    (hβ0 : 0 < β) (hβ : β ≤ 1 / 2 * min (1 / (4 * G * D)) α) :
    ∀ x ∈ P, ∀ y ∈ P,
      f y + inner ℝ (gradient f y) (x - y) + β / 2 * (inner ℝ (gradient f y) (x - y)) ^ 2
        ≤ f x := by sorry
end LogRegretOCO.FTAL
Source
Hazan, Agarwal, Kale, Logarithmic regret algorithms for online convex optimization, Mach Learn 69 (2007), p. 177, Lemma 3
Read-back

What the Lean code literally says, in plain math · claude-opus-5-5

Setting. Let n∈Nn \in \mathbb{N}n∈N and let Rn\mathbb{R}^nRn carry its Euclidean norm and inner product. The theorem concerns a set P⊆RnP \subseteq \mathbb{R}^nP⊆Rn, a function f:Rn→Rf : \mathbb{R}^n \to \mathbb{R}f:Rn→R, and real numbers D,G,α,βD, G, \alpha, \betaD,G,α,β.

Hypotheses.

  1. PPP is convex.
  2. D>0D > 0D>0, G>0G > 0G>0 and α>0\alpha > 0α>0.
  3. ∥x−y∥≤D\|x - y\| \le D∥x−y∥≤D for all x,y∈Px, y \in Px,y∈P.
  4. fff is differentiable at every point of PPP. This is differentiability as a function on all of Rn\mathbb{R}^nRn, at those points.
  5. ∥∇f(x)∥≤G\|\nabla f(x)\| \le G∥∇f(x)∥≤G for every x∈Px \in Px∈P.
  6. The function x↦e−αf(x)x \mapsto e^{-\alpha f(x)}x↦e−αf(x) is concave on PPP.
  7. The parameter β\betaβ satisfies
0<β≤12min⁡ ⁣(14GD, α).0 < \beta \le \tfrac12 \min\!\Big(\frac{1}{4GD},\, \alpha\Big).0<β≤21​min(4GD1​,α).

Conclusion. For all x,y∈Px, y \in Px,y∈P,

f(y)+⟨∇f(y), x−y⟩+β2 ⟨∇f(y), x−y⟩2  ≤  f(x).f(y) + \langle \nabla f(y),\, x - y\rangle + \frac{\beta}{2}\,\langle \nabla f(y),\, x - y\rangle^2 \;\le\; f(x).f(y)+⟨∇f(y),x−y⟩+2β​⟨∇f(y),x−y⟩2≤f(x).

Nothing requires PPP to be closed, bounded beyond the diameter condition, or nonempty.

Degenerate cases.

  • P=∅P = \emptysetP=∅. The conclusion holds vacuously.
  • PPP is a single point. Then x=yx = yx=y, and the conclusion reads f(x)≤f(x)f(x) \le f(x)f(x)≤f(x).
  • n=0n = 0n=0. PPP is empty or a single point, and the gradient is 000. The conclusion is again trivial.
  • Division by zero. It cannot occur, because G>0G > 0G>0 and D>0D > 0D>0.
  • Gradient convention. Differentiability is assumed on PPP, so the gradient values in the conclusion are never the default zero value used at non-differentiable points.
Human review
  • Endorsed by Shuze Chen · Sep 27, 2026

    Confirmed by the moderator at approval.

  • Endorsed by mikedeng1 · Sep 27, 2026

    Confirmed by the mission captain (proposal self-audit).

View graph

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me