How Machines Learn
Loading simulation…
flagWhat you'll discover
- arrow_forwardExplain the train-and-test loop of machine learning
- arrow_forwardDescribe a loss function as a measure of wrongness
- arrow_forwardShow how gradient descent nudges weights to reduce error
- arrow_forwardDistinguish underfitting from overfitting
The learning loop
Machine learning, however fancy the model, almost always boils down to the same loop. You show the model an example, it makes a guess, you compare the guess to the right answer, and you measure how wrong it was with a single number called the loss. Then you tweak the model's internal numbers (the weights) a tiny step in the direction that would have made the loss smaller. Repeat thousands or millions of times, and the model slowly gets better.
This loop — predict, measure, adjust — is the heartbeat of training. A modern neural network might run it billions of times across millions of examples, each time shaving a sliver off the error. In the simulation you can watch a real (tiny) network train before your eyes: press Train, and you will see the loss number fall as the model's predictions creep closer to the true pattern.
Loss and gradient descent
The loss is a score for how badly the model is doing — zero means perfect. There are many ways to compute it (mean squared error for numbers, cross-entropy for categories), but the goal is always the same: make it smaller. The clever part is knowing which way to nudge each weight to achieve that.
Imagine you are blindfolded on a hilly landscape, trying to reach the lowest valley. You feel the slope under your feet and take a step downhill. That is gradient descent: the model computes the slope of the loss with respect to each weight, then takes a small step (set by the learning rate) downhill. Too big a step and you leap across the valley; too small and you crawl. Getting this right is much of the art of training AI.
Underfitting and overfitting
A model that has barely learned anything is too simple to capture the pattern — it draws a straight line through curvy data. This is underfitting: high loss on both training and new data. At the other extreme, a model with too much capacity can memorise the training examples, including their noise, and fail completely on anything new. This is overfitting: tiny training loss but terrible real-world performance.
Good machine learning finds the sweet spot: a model flexible enough to capture the real pattern, but not so flexible that it memorises accidents. We guard against overfitting by splitting data into training and test sets, by adding regularisation, and — most honestly — by always checking the model on data it has never seen. In the simulation you can crank up the complexity and watch overfitting appear as the line wiggles violently to chase every point.