Back to feed
arXiv cs.LG·

Flatland: The Adventures of Gradient Descent with Large Step Sizes

Signal
72
Hype
15
In three linesTheoretical paper on gradient descent convergence with large step sizes. Authors formally define "large" steps requiring only local Lipschitz continuity of gradients, design first-order adaptive methods operating at edge of stability from training start, and show that pursuing global flatness too early slows convergence and harms generalization.
Read source
Your take?
Reinforcement learningReasoning

Summary generated by Claude — human-verified