I keep using gradient descent without really understanding the update rule. Can you derive it?
Start from a loss function \(L(\theta)\). The gradient \(\nabla L(\theta)\) points in the direction of steepest increase, so moving against it decreases the loss:
\[\theta_{t+1} = \theta_t - \eta \nabla L(\theta_t)\]The learning rate \(\eta\) controls the step. For a quadratic loss \(L(\theta) = \frac{1}{2}\theta^2\) the gradient is simply \(\theta\), so the update becomes \(\theta_{t+1} = (1 - \eta)\theta_t\). Convergence requires \(0 < \eta < 2\), which is the cleanest way to see why a large learning rate diverges.
Where does momentum fit in?
Momentum keeps a running average of past gradients:
\[v_{t+1} = \beta v_t + (1 - \beta) \nabla L(\theta_t)\] \[\theta_{t+1} = \theta_t - \eta v_{t+1}\]With \(\beta = 0.9\) the effective step is roughly ten times the raw gradient in directions that stay consistent, while noisy directions cancel out.
And Adam on top of that?
Adam adds a second moment, so each coordinate gets its own step size:
\[m_t = \beta_1 m_{t-1} + (1-\beta_1) g_t\] \[v_t = \beta_2 v_{t-1} + (1-\beta_2) g_t^2\] \[\theta_{t+1} = \theta_t - \eta \frac{\hat{m}_t}{\sqrt{\hat{v}_t} + \epsilon}\]The hats are bias corrections, \(\hat{m}_t = m_t / (1 - \beta_1^t)\), which matter only for the first few hundred steps when both averages start at zero.
The gap between the two curves after step four is the usual signature of overfitting: training loss keeps falling while validation turns upward. Early stopping around the minimum of the validation curve is the cheapest fix before touching regularisation.
