Gradient descent, but explain the math

Back to list2026-08-01
You

I keep using gradient descent without really understanding the update rule. Can you derive it?

ChatGPT

Start from a loss function \(L(\theta)\). The gradient \(\nabla L(\theta)\) points in the direction of steepest increase, so moving against it decreases the loss:

\[\theta_{t+1} = \theta_t - \eta \nabla L(\theta_t)\]

The learning rate \(\eta\) controls the step. For a quadratic loss \(L(\theta) = \frac{1}{2}\theta^2\) the gradient is simply \(\theta\), so the update becomes \(\theta_{t+1} = (1 - \eta)\theta_t\). Convergence requires \(0 < \eta < 2\), which is the cleanest way to see why a large learning rate diverges.

You

Where does momentum fit in?

ChatGPT

Momentum keeps a running average of past gradients:

\[v_{t+1} = \beta v_t + (1 - \beta) \nabla L(\theta_t)\] \[\theta_{t+1} = \theta_t - \eta v_{t+1}\]

With \(\beta = 0.9\) the effective step is roughly ten times the raw gradient in directions that stay consistent, while noisy directions cancel out.

You

And Adam on top of that?

ChatGPT

Adam adds a second moment, so each coordinate gets its own step size:

\[m_t = \beta_1 m_{t-1} + (1-\beta_1) g_t\] \[v_t = \beta_2 v_{t-1} + (1-\beta_2) g_t^2\] \[\theta_{t+1} = \theta_t - \eta \frac{\hat{m}_t}{\sqrt{\hat{v}_t} + \epsilon}\]

The hats are bias corrections, \(\hat{m}_t = m_t / (1 - \beta_1^t)\), which matter only for the first few hundred steps when both averages start at zero.

You
training-loss.png
ChatGPT

The gap between the two curves after step four is the usual signature of overfitting: training loss keeps falling while validation turns upward. Early stopping around the minimum of the validation curve is the cheapest fix before touching regularisation.