Flatland: The Adventures of Gradient Descent with Large Step Sizes
Signal
72
Hype
15
In three linesTheoretical paper on gradient descent convergence with large step sizes. Authors formally define "large" steps requiring only local Lipschitz continuity of gradients, design first-order adaptive methods operating at edge of stability from training start, and show that pursuing global flatness too early slows convergence and harms generalization.Read source
Your take?
Summary generated by Claude — human-verified