Any-stepsize Gradient Descent for Separable Data under Fenchel–Young Losses
NeurIPSSpotlight2025
TL;DR
The gradient descent (GD) has been one of the most common optimizer in machine learning…
Opening excerpt from the authors’ abstract. source
Read the paper
← All NeurIPS 2025 Spotlight papers · Browse the whole archive