How ChatGPT Works
Loading simulation…
flagWhat you'll discover
- arrow_forwardExplain a language model as next-token prediction
- arrow_forwardDescribe how text is broken into tokens
- arrow_forwardShow how attention lets a model focus on relevant context
- arrow_forwardDistinguish training a model from using (prompting) it
Guess the next word
Despite all the magic, a large language model like ChatGPT is doing essentially one thing: predicting the next piece of text. Given "The cat sat on the…", it computes the most likely next word — almost certainly "mat". Then it takes "The cat sat on the mat" as the new input and predicts the next word again, and again, building up a response one piece at a time. That is the entire trick.
What makes this powerful is scale. Trained on billions of pages of text, the model has effectively memorised the statistical patterns of human language — grammar, facts, reasoning styles, even tone. Predict the next word well enough, across enough context, and the result looks astonishingly like genuine understanding. In the simulation you can type the start of a sentence and watch the model's ranked guesses for what comes next.
Tokens, not words
Language models do not actually work with whole words — they work with tokens, small chunks that might be a word, part of a word, or even a single character. "ChatGPT" might be one token; "unbelievable" might be split into "un", "believ" and "able". A typical token is about four characters. Breaking text this way lets one model handle any language, including ones with no spaces between words, and lets it spell out even rare words it has never seen.
The number of tokens also dictates the model's context window — how much recent text it can keep in mind at once. A small model might remember a few thousand tokens; the largest models today handle over a million, enough to read a whole book in one go. Beyond the window, earlier text simply falls out of view.
Attention is all you need
In 2017 a single research paper titled "Attention Is All You Need" changed AI forever. It introduced the Transformer, an architecture built around a mechanism called attention that lets every token look back at every earlier token and decide which ones matter most for its prediction. To predict "it" in "the dog chased the cat because it…", attention lets the model focus on "dog" or "cat" to decide what "it" refers to.
This ability to weigh relevant context, no matter how far back, is what made Transformers so good at language that they swept the field. Every major modern model — GPT, Gemini, Llama — is a Transformer. They are still, at heart, predicting the next token; attention is what makes that prediction context-aware enough to produce fluent, useful, and sometimes astonishingly intelligent-seeming text.