The Secret Language of Machines Get the book

Free preview · Chapter 1

The Map Before the Journey

Read the opening chapter exactly as it is printed — or switch to the easy-read edition, made for small screens.

Page 1 of 5
Chapter 1

The Map Before the Journey

All models are wrong, but some are useful.— George Box

Meera found me at the tea stall outside the office, holding her phone out like evidence. On the screen a chatbot had just written her a perfectly decent poem about monsoon traffic.

“Explain this,” she said. “Not the marketing version. The real version. What is actually happening inside?”

I have been asked that question many times, and I have learned that the honest answer begins with a disappointment and ends with a wonder. The disappointment is that there is no magic in there. There is no little person, no hidden database of poems, no rulebook of grammar written by linguists. The wonder is that what is in there — a very large collection of numbers, tuned by a very simple rule, repeated an enormous number of times — turns out to be enough.

So before we climb, let me show you the whole mountain from a distance.

1.1A tower of ideas

Modern AI looks, from the outside, like a jungle of hundreds of unrelated techniques: transformers, diffusion, embeddings, reinforcement learning, retrieval, agents. From the inside it looks more like a tower, where each floor rests on the one below. Illustration 1.1 is the map for this entire book. We will climb it floor by floor.

A tower of ten floors. From the bottom: arithmetic, algebra and functions; vectors, matrices and tensors; probability, statistics and information; derivatives, gradients and the chain rule; optimisation, gradient descent and backprop; neurons, embeddings and attention; transformers and language models; vision, diffusion, audio and music; scaling, post-training and reasoning; and at the top retrieval, memory, tools and agents.
Illustration 1.1. The tower of AI mathematics. Nothing on a high floor is truly new; it is the lower floors, arranged cleverly.

Look at the bottom floor. Arithmetic. The same arithmetic you did in school — add, multiply, compare. Every single thing a language model does, every poem, every line of code, is at the lowest level a vast number of multiplications and additions. What makes it interesting is which numbers get multiplied, and how those numbers were chosen.

1.2The one equation to rule them

If I could keep only one line of mathematics from this entire book, it would be this one. Almost everything in modern machine learning is a variation on it.

θ∗  =  arg minθ  𝓛(θ; 𝒟)

Read it out loud with me: “theta-star equals the arg-min over theta of the loss of theta, given the data D.” That is still gibberish, so let us take it apart.

θthetaEvery adjustable number inside the model. For a large language model, this is hundreds of billions of numbers. Think of it as the position of every knob on an impossibly large mixing desk.
𝒟curly D, “the data”The examples we learn from: text, images, recordings.
𝓛(θ; 𝒟)the lossA single number that measures how badly the model with knobs θ does on the data. Big is bad, small is good. The semicolon says: θ is what we change; 𝒟 is held fixed.
arg minarg-min over theta“Find the setting of θ that makes what follows as small as possible.” Not the smallest loss itself — the knob positions that produce it.
θ∗theta-starThe winning knob settings: the trained model.

So the whole of “training an AI” is: write down a number that measures how wrong you are, then turn the knobs until that number is as small as you can make it. The rest of this book is about three questions hidden inside that sentence. What should the knobs be connected to? (That is architecture.) How do we measure wrongness? (That is the loss.) And how do we turn a hundred billion knobs in a sensible direction? (That is calculus and optimisation.)

Meera asks“That sounds too simple. Surely something so clever needs a clever rule.”

It does sound too simple. That is the first surprise, and I will not pretend to fully explain it away. The rule is simple; the consequences of applying it, at scale, to the right kind of data, are not. We will return to this puzzle at the very end of the book.

1.3Two modes of life: training and inference

A model lives two lives. In the first, it is a student. In the second, it is at work.

Training: an input example goes into the model, which makes a prediction; the prediction and the true answer give a loss, and a gradient flows back to the model telling it which way to turn every knob. Inference: a new input goes into the frozen model, which gives an output.
Illustration 1.2. Training changes θ so that predictions improve. Inference keeps θ fixed and simply runs the function on something new.

Training is expensive, slow and happens once (or a few times): data flows forward, a loss is computed, and a correction flows backward to adjust the knobs. Inference is what happens when you type into a chatbot: the knobs are frozen, and your words simply flow forward through the function. No learning happens while you chat. The model you talk to on Tuesday has exactly the same θ∗ as on Monday, unless its makers trained and released a new one.

Key ideaIf you understand vectors, probability, derivatives, gradients, matrix multiplication and optimisation, you understand the mathematical skeleton of modern AI. Everything else is organs and muscle hung on that skeleton.

Meera finished her tea. “Fine,” she said. “Start at the bottom.”

So we will.

That was Chapter 1 of 29

Keep climbing the tower.

Next comes the bottom floor — functions, vectors, probability, calculus — and then, floor by floor, neurons, attention, transformers, diffusion, voices and agents. Every symbol, pronounced out loud.