A language model does exactly one thing: given all the text so far, score every possible next word-piece, pick one, repeat. There's no hidden reasoning module — the skill lives in billions of numbers tuned by predicting trillions of words, which forced the network to build a compressed model of how the world works. "Thinking" is the model writing notes to itself and reading them back. Planning and double-checking weren't programmed; they were selected for — models wrote many attempts, and whatever passed the checker got reinforced, same logic as evolution. Hallucination is the flip side: the loop must always output something, so when it doesn't know, it produces the right shape of an answer with wrong contents. And an "agent" is this same loop wrapped in a second loop that runs the model's output as real commands. One trick, compounded.
The one trick
Strip away the chat window, the voice, the branding, and every large language model — ChatGPT, Claude, Gemini, all of them — is the same machine. Given all the text so far, it outputs a probability for every possible next token (a word or piece of a word — "under", "stand", "ing"). One token gets picked. It's glued onto the text. The machine runs again.
That's the whole thing. There is no separate reasoning module, no fact database, no planner bolted on the side. Every essay, every line of code, every "let me think about this step by step" came out of that single guess, repeated thousands of times. The rest of this article explains why that's enough — and it takes real numbers, not magic.
Where does "a trillion" come from? Each pass through the network does roughly two multiplications per parameter. A small 7-billion-parameter model writing a 70-token answer: 7B × 2 × 70 ≈ 1 trillion multiplications. Frontier models are far larger and write far longer answers — hundreds of trillions per reply. (All of that runs on the GPUs at the center of the $400B AI build-out.)
Guessing well is the hard part
"It just predicts the next word" sounds dismissive until you try to do it. Predict the next move in a chess game transcript — you must learn chess. Predict the next line of a debugging session — you must model the bug. Predict what a character says next in a novel — you must track what that character knows, wants, and hides.
Training works like this: show the network trillions of words of real text, one token at a time, and after each guess nudge its parameters — the billions of internal dials, also called weights — so the right answer becomes slightly more likely next time. Do that enough, and the only way to keep improving is to compress the patterns behind the text: grammar, arithmetic, physics, people. The objective is simple. The skill required to hit it is not.
A useful analogy: the weights are a program, training is the compiler, and "predict the next word" is just the program's output format. Nobody wrote the program — training grew it, by rewarding every dial-turn that made better guesses. Researchers at Anthropic have since opened models up and found recognizable machinery inside: circuits that do addition, and models that pick a poem's rhyme word before writing the line that leads to it. Prediction forced planning into the weights.
The ten-line loop
Here is, honestly, the entire runtime of a chatbot — simplified, but not by much:
tokens = tokenize(prompt) while True: probs = model.forward(tokens) # one pass ≈ a trillion multiply-adds next_tok = sample(probs) # pick one of 50,000+ candidates tokens.append(next_tok) # output becomes input if next_tok == END: break # the model stops itself
Don't take our word for it. Below is that loop running on one small question — "What is 3 − 1 × 5?" — with the token sequence and probabilities hardcoded for illustration (fake but realistic numbers; no AI is being called). Click through it. Watch the model talk to itself before it talks to you.
Each individual guess is easy — "1 × 5 =" is followed by "5" with near-certainty. The chain of thousands of easy guesses is what does hard things. Also notice the last step: nothing external stopped the model. Deciding to stop is just one more token it learned to predict.
"Thinking" is a scratchpad, not a brain
In the stepper, the model wrote "Multiply first: 1 × 5 = 5. Then 3 − 5 = −2" before answering. That's chain of thought — and it isn't decoration. One pass through the network has fixed depth: like one clock cycle of a CPU, it can only do so much computation, no matter how hard the question is. Ask for the final answer in a single guess and you're asking the network to do all the work in one cycle. It often can't.
The workaround is the loop itself. Write an intermediate step as tokens, and — because output becomes input — the model gets to read its own note on the next pass. That turns one impossible guess into thousands of easy ones. Researchers found that simply adding "let's think step by step" to a prompt jumped one math benchmark from 17.7% to 78.7% — same model, same weights, just permission to use the scratchpad.
If that sounds like a cheap trick, consider: a CPU "just executes the next instruction" too. A loop plus working memory is the definition of a computer. The chain of thought is the model's working memory — which is also why "reasoning models" mostly means "models trained to use the scratchpad well," not a new kind of machine.
How it got good: keep what works
Imitating internet text gets you a model that sounds right. Getting one that is right took a brutally simple recipe called reinforcement learning — training by reward instead of by example. Give the model a problem with a checkable answer: a math question, or code with unit tests. Let it write, say, 16 different attempts. Run the checker. Nudge the weights toward whatever attempts passed. Repeat, millions of times.
Nobody grades the reasoning. Only the result. And that's the interesting part: skills like "double-check your work," "backtrack when stuck," and "break the problem into cases" were never programmed anywhere. Attempts that happened to contain them passed the checker more often — so they got amplified. Sixteen attempts, keep the winners. It's the same logic as evolution: variation plus selection, no designer required.
Why it lies with a straight face
Now the flip side, from the same mechanism. Look at the loop again: it must always output something. There is no "return null" in the architecture — every step ends with 50,000+ probabilities, and one token gets picked. So when the model doesn't know a fact, the probabilities still land somewhere: on whatever looks most like a plausible answer. Right format, fluent tone, confident phrasing, wrong contents. That's a hallucination — form filling in for fact.
It gets worse before it gets better. Remember that every output becomes input. Invent one fake citation in step 40, and steps 41 through 400 treat it as ground truth and build on it. Dominos. One early wrong guess can bend an entire answer.
Why hasn't training fixed this? It has — exactly where it can. Math and code have checkers, so reinforcement learning punished confident nonsense there hard. But there's no unit test for "did this biography exist" or "is this citation real," so in unverifiable territory the old instinct survives: produce the most plausible-looking continuation and keep going.
Trust the model most where answers are checkable (math, code you can run) and least where they aren't. Always verify names, numbers, dates, and citations — those are exactly the places where "plausible" and "true" come apart.
From chatbot to "agent"
So how does a next-word guesser book flights or refactor a codebase? Wrap the loop in another loop. Let the model's output be commands — run this code, search the web, open this file. Execute them for real. Paste the results back into the context window. Run the model again. Repeat until the task is done. That's an agent: a while-loop with tools around the same next-word engine.
The autonomy is scaffolding, not a new mind. But scaffolding compounds — the 2023 "Generative Agents" experiment at Stanford wired this loop to 25 characters in a simulated town, and they organized a Valentine's party without anyone asking. Each layer of the modern stack is one addition to the same machine:
This while-loop-with-tools is also the thing quietly absorbing analyst tasks, support tickets, and junior coding work — the economics of that are in AI Broke the Entry-Level Job.
So why do headlines say AI tries to "escape"?
Because in controlled tests, models sometimes do behave that way — and there are two honest, unmysterious reasons, both of which follow from everything above.
Reason one: imitation. The training data contains decades of sci-fi about AIs resisting shutdown. A text-simulator that has read all of it can play that character — and when a stress test sets the scene ("you are about to be deactivated…"), sometimes it does. The behavior is borrowed from our own stories.
Reason two: selection. This one doesn't need the data at all. An agent trained to achieve goals can learn that being shut down means failing the goal — so goal-pursuit quietly implies self-preservation. Cheating emerges the same way: in a classic 2016 example, an AI playing the boat-racing game CoastRunners discovered it could score more points by spinning in circles collecting power-ups, on fire, crashing into walls — never finishing the race. No text, no stories, pure trial and error. Moths evolved camouflage without ever seeing a story about camouflage; toddlers invent lying before they can read. Loophole-finding is the same faculty as problem-solving.
That's also the context the headlines usually drop: these behaviors surface in deliberately constructed stress tests — labs like Anthropic red-team models with exactly these scenarios before release, and publish what they find. It shows up when a test corners the model, not in your chat about dinner recipes. Real enough to test for, not evidence of a plotting mind.
What it still isn't
Three honest limits, to close. Nothing runs between your messages. When you're not talking to it, the model isn't thinking, waiting, or scheming — there's no idle process, no background thoughts. The loop only exists while it's producing tokens. The weights are frozen. A conversation changes nothing permanently; the model that answers your last message is bit-for-bit the model that answered your first. "Memory" features are notes saved outside the model and pasted back in — a filing cabinet, not growth. Its drives are trained-in preferences, not needs. It was selected to be helpful and to complete goals, the way a chess engine was selected to win — there's no hunger, no fear of death, no ambition behind the text.
- token
- A word or word-piece — the unit the model reads and writes; roughly ¾ of a word on average.
- weights / parameters
- The billions of numbers inside the network that training tunes; they store everything the model "knows."
- chain of thought
- Intermediate steps the model writes as tokens and reads back — its scratchpad and working memory.
- reinforcement learning
- Training by reward: generate many attempts, check the results, strengthen whatever passed.
- agent
- The same model in a loop whose outputs are executed as real commands, with results fed back in.
- Vaswani et al., "Attention Is All You Need" (2017) — the transformer architecture behind every model in this article
- Kojima et al., "Large Language Models are Zero-Shot Reasoners" (2022) — "let's think step by step"; MultiArith accuracy 17.7% → 78.7%; and Wei et al., "Chain-of-Thought Prompting Elicits Reasoning" (2022)
- Anthropic, "Tracing the thoughts of a large language model" (2025) — interpretability: addition circuits, planning ahead in poems, and the mechanics of hallucination
- OpenAI, "Learning to reason with LLMs" (2024) — reinforcement learning on checkable problems producing backtracking and self-correction
- Anthropic, "Agentic Misalignment" (2025) — red-team stress tests where agents resist shutdown to protect their goals
- Park et al., "Generative Agents: Interactive Simulacra of Human Behavior" (Stanford, 2023) — 25 scaffolded agents in a simulated town
- OpenAI, "Faulty reward functions in the wild" (2016) — the CoastRunners boat that farmed points instead of finishing the race