NEBians
search Sign In Register

How ChatGPT Works

Lesson 7 of 8 Simulation schedule15 min

Loading simulation…

tuneAdjust the controls and watch what happens

flagWhat you'll discover

  • arrow_forwardExplain a language model as next-token prediction
  • arrow_forwardDescribe how text is broken into tokens
  • arrow_forwardShow how attention lets a model focus on relevant context
  • arrow_forwardDistinguish training a model from using (prompting) it

Guess the next word

Despite all the magic, a large language model like ChatGPT is doing essentially one thing: predicting the next piece of text. Given "The cat sat on the…", it computes the most likely next word — almost certainly "mat". Then it takes "The cat sat on the mat" as the new input and predicts the next word again, and again, building up a response one piece at a time. That is the entire trick.

What makes this powerful is scale. Trained on billions of pages of text, the model has effectively memorised the statistical patterns of human language — grammar, facts, reasoning styles, even tone. Predict the next word well enough, across enough context, and the result looks astonishingly like genuine understanding. In the simulation you can type the start of a sentence and watch the model's ranked guesses for what comes next.

Tokens, not words

Language models do not actually work with whole words — they work with tokens, small chunks that might be a word, part of a word, or even a single character. "ChatGPT" might be one token; "unbelievable" might be split into "un", "believ" and "able". A typical token is about four characters. Breaking text this way lets one model handle any language, including ones with no spaces between words, and lets it spell out even rare words it has never seen.

The number of tokens also dictates the model's context window — how much recent text it can keep in mind at once. A small model might remember a few thousand tokens; the largest models today handle over a million, enough to read a whole book in one go. Beyond the window, earlier text simply falls out of view.

Attention is all you need

In 2017 a single research paper titled "Attention Is All You Need" changed AI forever. It introduced the Transformer, an architecture built around a mechanism called attention that lets every token look back at every earlier token and decide which ones matter most for its prediction. To predict "it" in "the dog chased the cat because it…", attention lets the model focus on "dog" or "cat" to decide what "it" refers to.

This ability to weigh relevant context, no matter how far back, is what made Transformers so good at language that they swept the field. Every major modern model — GPT, Gemini, Llama — is a Transformer. They are still, at heart, predicting the next token; attention is what makes that prediction context-aware enough to produce fluent, useful, and sometimes astonishingly intelligent-seeming text.

quizCheck your knowledge

1. At its core, a large language model is trained to…
2. A "token" in a language model is…
3. The Transformer architecture is powerful because of a mechanism called…