Build it up, step by step
The big picture: build once, use many times
Two phases are easy to blur together, so separate them first.
Building the model happens once and takes months. A training program reads trillions of words and slowly adjusts a huge list of numbers. When it finishes, the result is literally a file of numbers, called the weights or parameters: hundreds of gigabytes for a large model.
Using the model happens every time you send a message and takes seconds. Your text is turned into numbers and pushed through a fixed recipe of arithmetic that uses those weights. Out comes a probability for every possible next piece of text. One piece is picked and added to the text, and the whole thing runs again.
Building the model once · months · enormous compute
- Trillions of words of text
- Training adjusts the numbers, step by step
- A file of numbers: the weights
the same weights are loaded onto GPUs ↓
Using the model every message · seconds
- Your text, as tokens
- Arithmetic using the weights
- A probability for every next token
- Pick one, add it, repeat
So a model is not a database of stored sentences, and it is not a program of rules that someone wrote. It is arithmetic whose billions of numbers were tuned by example. The rest of this layer builds that arithmetic up from scratch.
A model is a function with knobs
Start tiny. Say you want to predict an apartment's monthly rent from its size. A simple guess: rent = w × size + b. The numbers w (dollars per square foot) and b (a base amount) are parameters: knobs you can turn. Each setting draws a different line.
Learning means finding the knob settings that make predictions match known examples. We score the mismatch with a single number, the loss; here it's the typical size of the error (the dashed lines). Drag the knobs yourself, then press Auto-fit. The computer repeatedly nudges each knob in whichever direction lowers the loss. That's gradient descent, and it is exactly how large models are trained, just with billions of knobs instead of two.
A neuron: a weighted sum, then a bend
An artificial neuron is the same idea with more inputs. It multiplies each input by its own weight, adds them up along with a bias, then passes the total through a simple bend. The networks on this page use an S-shaped curve that squashes any number into the range 0 to 1, so a neuron acts like a soft switch: off for big negative sums, on for big positive ones.
Why bend at all? Without it, stacking neurons only ever gives another straight line, because a weighted sum of weighted sums is still a weighted sum. The bend is what lets a network represent curves, thresholds and "only if" behavior. Try the weights below: thicker lines are bigger weights, and dashed red lines are negative ones.
A network: many bends add up to any shape
Put many neurons side by side and you have a layer. Feed one layer's outputs into another and you have a network. Each neuron can only draw one soft boundary, but the output combines them, so together they can carve out shapes no single neuron can.
The figure below is a real network with two inputs (a point's position) and up to six hidden neurons, coloring every point on a map. Tune its weights by hand, then press Learn and watch gradient descent do it for you. The XOR pattern needs at least two neurons. In Layer 8 you can watch a network learn to trace a curve you draw.
A large language model is this recipe at enormous scale. Instead of two inputs it takes thousands of numbers per token. Instead of one hidden layer it has around a hundred stacked blocks. Instead of one output it produces a score for every token in its vocabulary. The principle doesn't change: weighted sums, bends, and weights tuned by gradient descent.