The Map Before the Journey
All models are wrong, but some are useful.— George Box
Meera found me at the tea stall outside the office, holding her phone out like evidence. On the screen a chatbot had just written her a perfectly decent poem about monsoon traffic.
“Explain this,” she said. “Not the marketing version. The real version. What is actually happening inside?”
I have been asked that question many times, and I have learned that the honest answer begins with a disappointment and ends with a wonder. The disappointment is that there is no magic in there. There is no little person, no hidden database of poems, no rulebook of grammar written by linguists. The wonder is that what is in there — a very large collection of numbers, tuned by a very simple rule, repeated an enormous number of times — turns out to be enough.
So before we climb, let me show you the whole mountain from a distance.
1.1A tower of ideas
Modern AI looks, from the outside, like a jungle of hundreds of unrelated techniques: transformers, diffusion, embeddings, reinforcement learning, retrieval, agents. From the inside it looks more like a tower, where each floor rests on the one below. Illustration 1.1 is the map for this entire book. We will climb it floor by floor.
Look at the bottom floor. Arithmetic. The same arithmetic you did in school — add, multiply, compare. Every single thing a language model does, every poem, every line of code, is at the lowest level a vast number of multiplications and additions. What makes it interesting is which numbers get multiplied, and how those numbers were chosen.
1.2The one equation to rule them
If I could keep only one line of mathematics from this entire book, it would be this one. Almost everything in modern machine learning is a variation on it.
Read it out loud with me: “theta-star equals the arg-min over theta of the loss of theta, given the data D.” That is still gibberish, so let us take it apart.
So the whole of “training an AI” is: write down a number that measures how wrong you are, then turn the knobs until that number is as small as you can make it. The rest of this book is about three questions hidden inside that sentence. What should the knobs be connected to? (That is architecture.) How do we measure wrongness? (That is the loss.) And how do we turn a hundred billion knobs in a sensible direction? (That is calculus and optimisation.)
It does sound too simple. That is the first surprise, and I will not pretend to fully explain it away. The rule is simple; the consequences of applying it, at scale, to the right kind of data, are not. We will return to this puzzle at the very end of the book.
1.3Two modes of life: training and inference
A model lives two lives. In the first, it is a student. In the second, it is at work.
Training is expensive, slow and happens once (or a few times): data flows forward, a loss is computed, and a correction flows backward to adjust the knobs. Inference is what happens when you type into a chatbot: the knobs are frozen, and your words simply flow forward through the function. No learning happens while you chat. The model you talk to on Tuesday has exactly the same θ∗ as on Monday, unless its makers trained and released a new one.
Meera finished her tea. “Fine,” she said. “Start at the bottom.”
So we will.