Training
The learning phase: the AI makes guesses, gets told how wrong it was, and adjusts itself a little. Repeated millions of times.
In everyday terms
Training is slow and expensive: weeks on thousands of chips for big models. It happens once (or occasionally), before you ever use the model.
For professionals
Iterative optimisation, typically stochastic gradient descent, that updates parameters to reduce a loss function. For LLMs: pre-training on next-token prediction, then post-training (instruction tuning, RLHF).
Think of it like…
Practising darts: throw, see how far off you were, adjust your aim, repeat.
You've already seen it
When people say a model has a "knowledge cut-off date", that's when its training data ended.
Myth vs reality
Myth: ChatGPT learns from every conversation as it happens.
Reality: Models don't update themselves while you chat. Changes come from separate, later training runs.
Quick check
During training, the model…
- Answers user questions
- Repeatedly adjusts its numbers to reduce mistakes
- Searches Google
- Is switched off
Show answer
Repeatedly adjusts its numbers to reduce mistakes: Training = guess, measure error, adjust.