A walk through the four learning settings — supervised, unsupervised, reinforcement, self-supervised — organised around what feedback the world actually returns and when, with MNIST, tic-tac-toe and next-word prediction as the worked cases. Closes with the GLM-5.3 release read as an example of reward-driven training.
Episode page & show notes
The one question behind four names
Supervised, unsupervised, reinforcement, self-supervised are not four families of algorithms. They are four answers to one question: what does the system get told, and when. Everything here follows from that.
Answers already written down
MNIST as the worked case: pixels paired with a digit a person chose. The correct answer sits beside every guess, which makes the feedback immediate — and paid for in human labour. Sixty thousand images means sixty thousand decisions, and the price climbs when the judgment is a chest X-ray two radiologists may read differently, an email that is angry or merely brisk, or a cyclist boxed frame by frame.
The same pile of pixels with nothing attached. What remains is structure: which images resemble each other, and how much of the seven hundred eighty-four dimensions is redundant. The awkward part gets stated plainly — with no answer key there is no clean test. Six customer groups or eight? Nothing in the data holds an opinion, so you judge by downstream usefulness or by whether someone who knows the domain recognises the groups.
A number that arrives at the end
Tic-tac-toe as Sutton and Barto set it up in Reinforcement Learning: An Introduction: pick a free square, receive nothing, and after three to five of your own moves one number arrives. It says how things went, not what you should have done — the credit assignment problem. Values propagate backward along the path actually walked, one step at a time. And because the learner only sees consequences of moves it chose, it must sometimes act against its own current judgment to find out.
Targets hidden in the data
Cover a word and ask for it. A five-hundred-word page becomes roughly five hundred examples, and nobody annotated anything. It uses supervised machinery on invented targets, and it usually is not the last step — pretraining is followed by fine-tuning on far fewer real labels.
What you have, not what sounds advanced
Loan repayment records, uncategorised support transcripts, and a home-screen recommender that can be read either way. The four settings are not a ranking.
GLM-5.3, read through its reward
Z.ai's GLM-5.3 release keeps GLM-5.2's 743-billion-parameter base with no further pretraining and scales the reinforcement learning: verified multi-step environments. Reported jumps on Terminal-Bench 3.0 and DeepSWE are evidence about those benchmarks and nothing else. The open weights were held back around two weeks after cyber capability scaled faster than expected.