Every agent you will ever build (a language model wired into a loop that thinks, calls tools, and decides what to do next) rests on one small act, repeated thousands of times: it reads the words so far, gives every possible next word a probability, and draws one. After “The class was about”, the lecture’s model gives “to” a 31.71% chance. One dial, the temperature, decides how often you actually get it.
Learn what a language model computes, how it is built, trained and served, and how an agent builder feeds it context (everything the model reads before it answers) and reads what it writes.
Drag the temperature, then press Draw the next word a few times. Watch the same sentence continue differently each time. Then we build, piece by piece, everything underneath: the attention that reads the words, the training that sets the probabilities, the hardware that serves them, and the context an agent feeds in.
You need a rough idea of what a probability is (a number from 0 to 1 that says how likely something is) and a little Python for the code. We build the rest from zero.
Based on Lecture 2 of Stanford CS 329Z: Engineering AI Agents (Fall 2026), taught by Diyi Yang, Michael Ryan and John Yang, and its two readings from Anthropic. Lecture slides. Every idea is re-taught here in our own words and devices; it follows Lecture 1, What Are Agentic Systems?
The class was about ___
The ten words and their probabilities are the lecture’s, for “The class was about” (to 0.3171, as 0.0491, the 0.036, a 0.033, two 0.0147, how 0.0124, getting 0.0101, three 0.01, time 0.0091, “ 0.0081) and for “about to” (be 0.1501, begin 0.0855, go 0.0627 and seven more); the blue mark on “to” is that raw 31.71%, and the rest of the model’s full list of possible next words holds the other 50.04% and is left out here, so shares among these ten run higher than the model’s own. Temperature is applied exactly, each probability raised to the power 1/T and the ten rescaled to add up to 100%, and the draws are real random draws in your browser.
Chapter 0
See why an engineer who works at the top of the stack still needs to know how the model underneath computes, learns and runs
Picture the agent from Lecture 1: a language model in a loop, thinking, calling a tool, reading the result, deciding what to do next. You hand it a job, say fixing a failing test in a code repository, and it works. Then three strange things happen.
First, you run it again on the very same job, and it takes a different path: different files opened, different edits, a different number of steps. Second, the longer it runs, the slower each step becomes. Step forty takes far longer than step four, although nothing about the model has changed. Third, an hour in, it quietly ignores an instruction you gave it at the very start.
None of these is a bug in your loop. Each one is the visible end of a design decision made deep inside the model: how it picks each next word, how it reads the words it already has, how much it can hold at once, and what was done to it before you ever called it. This lecture opens the box so that you can see those decisions and build around them.
The lecture begins with a map, drawn as a pyramid in three layers. At the bottom is LLM engineering, the work of building a model in the first place: choosing a tokenizer (the rule that cuts text into pieces), an architecture (the model’s wiring), running pretraining and post-training (the two big phases of learning, Chapters 5 and 6) and reinforcement learning, studying scaling laws (how quality grows with size and data), and writing GPU kernels and distributed-compute code (the software that runs the arithmetic fast across many chips).
In the middle is LLM adaptation: ways to change what an existing model does. Finetuning trains it further on new examples. Distillation trains a smaller model to imitate a larger one. LoRA is a cheap form of finetuning that trains a small add-on instead of the whole model. Learning from feedback (the slide names two methods, RLHF, reinforcement learning from human feedback, and DPO, direct preference optimization) trains it on human preferences. And in-context learning teaches it a task with nothing but instructions and examples placed in its prompt, without changing the model at all.
At the top are LLM applications: chatbots, retrieval-augmented generation (RAG, where the model is handed relevant documents before it answers; it is the next lecture’s topic) and agents.
Stanford teaches this pyramid across several courses, and the slide draws each course as a loop around the topics it covers. CS336 covers the bottom layer, building models from scratch, plus the learning-from-feedback topics just above it. CS224N threads a path down through all three layers. CS 329Z, this course, sits at the very top: agents, chatbots and RAG, dipping down into in-context learning, and branching out into agent frameworks, multi-agent systems, agent memory, proactive agents, agent evaluation and agent design.
So the lecture is candid about its own position. We will operate at the highest level of the stack. But it is still important to understand a bit about how the models work, because every limit at the bottom leaks upward into the agent you build at the top.
The lecture sets two learning objectives, and they make a good test for this whole lesson. The first is to understand and explain the core design decisions that go into training and running modern large language models (LLMs: simply very big language models; Chapter 1 defines exactly what a language model is). The second is to synthesise how those decisions lead to limitations that shape how applications use the models.
The second objective is the one that matters to a builder. Knowing that attention compares every word with every other word is trivia. Knowing that it makes step forty of your agent slower and more expensive than step four, and knowing what to do about it, is engineering.
Here is the second of the three symptoms, made concrete. An agent’s context is the text the model reads before each step. It starts with the system prompt (the builder’s standing instructions) and the task. Every step adds a thought, a tool call and the tool’s result, and nothing is removed. Before writing each new word, the model compares that word with every word already in the context; Chapter 3 shows exactly how. Drag the number of steps and watch two things grow at different speeds.
Left: the context, built up step by step. Right: the attention triangle (the word-comparing work Chapter 3 explains) of one layer, one row per word and one column per earlier word; its area is the number of comparisons. Drag the steps and compare how fast each grows.
Token counts are illustrative: a 1,200-token system prompt and task, then 100 tokens of thought, a 50-token tool call and a 1,400-token tool result per step. The pair count is exact for causal attention, where each token is compared with itself and every earlier token: n(n + 1)/2 pairs for n tokens, in every attention layer. Real systems keep earlier work in a cache (Chapter 7), which changes when the cost is paid, not how it grows.
The column on the left grows in a straight line: fifty steps make it about 29 times longer than one step. The triangle on the right grows with the square of the length. Double the context and the triangle holds about four times as many pairs. That is the whole reason the fortieth step is slow.
This one fact will come back again and again. It is why Chapter 4 exists (researchers keep inventing cheaper kinds of attention), why Chapter 7 separates reading a prompt from writing an answer, and why Chapter 10 argues that the best context is the smallest one that still does the job.
The other two symptoms have chapters of their own. The different path on every run comes from sampling: the model does not pick words, it gives them probabilities, and something must draw one. You have already felt it in the hero at the top of this page, and Chapter 2 takes it apart. The forgotten instruction comes from how well a model can use a very long context, and from what agents do when the context overflows (Chapters 4 and 10).
The lecture has five parts, and this lesson follows them in order: what a language model is, how it is built, how it is trained, how it is run, and how an agent builder uses one. Along the way comes the lecture’s assigned reading, Anthropic’s “Building Effective Agents”, in a chapter of its own.
Chapter 1
Turn a sentence into a probability, one word at a time, and see why that makes a language model something you can draw from
Start typing a message on your phone. Above the keyboard, three suggestions appear for the next word, and they change as you type. The keyboard has a guess about what comes next, and some guesses are stronger than others.
A language model is that idea taken as far as it will go. Given any beginning of a text, it gives every possible next piece of text a probability: a number from 0 to 1, where all the numbers together add up to 1. Chain those guesses together and it can say how likely an entire sentence is.
A model does not work on letters or on whole words, exactly. It works on tokens: a common word, a piece of a rarer word, or a punctuation mark. The full list of tokens a model knows is its vocabulary. The lecture’s example treats each word as one token, which keeps the arithmetic readable; real tokenizers split unusual words into several pieces. (The lecture points to Hugging Face’s course for tokenizers; our tokenization lesson builds one from zero.)
Inside the model, each token becomes a number, its position in the vocabulary. So the input to a language model is a list of integers, one per token, and nothing more.
How likely is the sentence “The class was about agents.”? Asking directly is hopeless: there are far too many possible sentences to count. The trick is to ask a string of small questions instead. How likely is the first word to be “The”? Given “The”, how likely is “class” next? Given “The class”, how likely is “was”? And so on to the full stop. Multiply the answers and you have the probability of the sentence. This is the chain rule of probability, and the lecture writes it in one line.
The lecture works it on its own sentence. The model gives “The” as a first word a probability of 0.0377. Given “The”, it gives “class” 0.0001. Given “The class”, it gives “was” 0.0344. The slide skips the rows for “about” and “agents”, then shows the last one: given “The class was about agents”, the full stop gets 0.0252. All six factors multiplied together give 1.2 × 10−17.
0.0377 × 0.0001 × 0.0344 × … × 0.0252 = 1.2 × 10−17 (six factors, two not shown)
That is 0.000 000 000 000 000 012. Do not read it as “the model thinks this sentence is nearly impossible”. There are astronomically many sentences a model could produce, and the probability has to be shared among all of them, so any one specific sentence gets a tiny slice. The numbers are useful for comparing: a sentence the model finds natural gets a larger slice than a garbled one.
The slide carries a footnote that looks like a detail and is not: in practice we take the sum of the log of the probabilities instead, for numerical stability. A logarithm answers the question “what power of 10 gives me this number?”: log10(100) = 2, because 102 = 100. Because of how powers multiply, log(a × b) = log(a) + log(b), and that identity is exactly what lets a logarithm turn multiplying into adding. The base-10 logarithm of 1.2 × 10−17 is simply −16.92, and the logarithm of each factor adds up to it: −1.42 for “The”, −4 for “class”, −1.46 for “was”, −1.60 for the full stop, plus the two rows the slide skips.
Why bother? A computer stores each number in a fixed number of bits, which means there is a smallest positive number it can hold. Multiply enough small probabilities and the product falls below that floor and becomes exactly zero. This is called underflow, and once it happens every sentence looks equally impossible. Push the sentence longer below and watch it happen.
Every token here has the same probability, 0.0344. The warm line is the true size of the product, written as a power of ten. The dashed floors are the smallest numbers a 32-bit and a 64-bit computer number can hold. Drag the length past each floor.
Illustrative: every token gets the probability the slide gives “was” (0.0344), so the numbers are easy to follow. The products are computed live in your browser, multiplying one token at a time in 32-bit and in 64-bit floating point; the sum of logs is n × log10(0.0344).
With 32-bit numbers, a common choice on graphics chips, the product becomes exactly zero after only 31 tokens. Even 64-bit numbers give out after 222 tokens. An agent’s context easily holds tens of thousands. The sum of logarithms, meanwhile, is just an ordinary negative number that grows steadily more negative: −58.5 at 40 tokens, −439 at 300. It never underflows, which is why training code (Chapter 5) and decoding code (Chapter 2) work with log-probabilities throughout.
The lecture then sorts models into two families, using a picture of red and blue dots. A discriminative model learns to tell classes apart: it draws a boundary, and for any point you give it, it says which side the point falls on. Logistic regression and support vector machines are the lecture’s examples. A generative model learns the distribution of the data: where points tend to lie, and how often. Naive Bayes and diffusion models (the kind that generate images) are its examples.
The difference shows up the moment you ask for something new. Try it: the same dots, two kinds of model.
Circles and squares are two kinds of real example. Choose a model, then press Make a new example. Then press Label a random point in each mode.
Toy data: 28 points from two made-up groups, seeded so they are the same on every visit. The boundary is the line halfway between the two group centres; each blob is an ellipse two spreads wide around its centre. The generative model labels a point by asking which blob makes it more likely.
The boundary can label any point, but asked for a new example it has nothing to offer: it never learned where the points lie, only which side of the line they fall on. The blobs can do both. They can label a point (whichever blob makes it more likely wins), and they can make a new point by drawing one from a blob.
Now the lecture’s knowledge check, which it admits the notation makes tricky: is a language model, written as the chain rule above, generative or discriminative? Each factor, P(xt | x<t), looks like a classifier: from the earlier tokens, predict a label, the next token. That is the trap. Look at the left side instead. The product equals P(x1, x2, x3, …, xT), a probability for every possible sequence: the full distribution of text.
So a language model is generative, and the lecture draws the consequence in one line: generative models are ones we can sample from. Draw a first token from P(x1). Given it, draw a second from P(x2 | x1). Keep going. Each draw is exactly what the hero at the top of this page does, and every sentence, tool call and answer an agent writes is made this way, one draw at a time.
As data flow, a language model is a function with a fixed shape. In: a list of token numbers, say the four for “The class was about”. Out: for the next position, a list of V probabilities, one for every token in the vocabulary, adding up to 1. Nothing else. The model’s internal numbers, its weights, were set during training and do not change while you use it.
Two consequences matter to a builder. First, the model never outputs “the answer”, only a distribution; some rule has to turn it into text, and that rule is yours to choose (Chapter 2). Second, the probability of a long output is a product of many factors, so a long tool call or a long plan is always a chain of small bets, each one made with only the text before it in view.
Chapter 2
Turn the model’s probabilities into actual text with greedy decoding, temperature, top-k, top-p and beam search, and learn which to use when
Here is what the model hands you after “The class was about”: a list. On the lecture’s slide the top ten are “to” at 0.3171, “as” 0.0491, “the” 0.036, “a” 0.033, “two” 0.0147, “how” 0.0124, “getting” 0.0101, “three” 0.01, “time” 0.0091, and an opening quotation mark at 0.0081. Together those ten hold almost exactly half the probability (0.4996). Thousands of other tokens share the other half, each with a sliver.
The model stops there. Something still has to turn the list into one word, then do it again for the word after, and so on. That something is called a decoding strategy (the lecture also says sampling strategy). It is not part of the model: it is a few lines of code that run after it, and you choose them. The lecture names the main ones: greedy, nucleus (top-p), top-k, and more.
The simplest rule takes the most likely word every time. After “about”, that is “to”. After “about to”, the slide lists “be” 0.1501, “begin” 0.0855, “go” 0.0627, “start” 0.0496, “get” 0.0390, “become” 0.0342, “enter” 0.0313, “take” 0.0220, “end” 0.0192 and “graduate” 0.0180, so greedy takes “be”. This is greedy decoding. It is deterministic: the same prompt gives the same text on every run.
Notice what the lecture’s slide does next. It highlights “begin”, the second choice, and its finished sentence reads “The class was about to begin.” The point, in the slide’s words: we don’t always have to pick the most likely next token. Greedy text tends to be flat and can repeat itself, and, as we will see with beam search, the best word now does not always lead to the best sentence.
The alternative is to sample: draw a word at random, in proportion to its probability. Among the ten listed words, “to” then comes up about 63% of the time and “how” about 2.5%. That is why the hero’s sentence continues differently on every press, and why an agent run twice can take two paths.
Pure sampling has a knob, the temperature T. The slide’s one-line summary says: divide all the probabilities by some temperature T, then renormalise so they sum to 1. Its chart shows the next word for “I am tired, I will take a ___”. At T = 0.1, “nap” takes essentially all the probability. At T = 0.5, “nap” still dominates with “break” a distant second. At T = 1 the five words (nap, break, rest, bath, shower) fall off gently, and at T = 2 they are much closer together, about a third for “nap” and more than a tenth for each of the others.
One detail makes the summary precise. Before a model produces probabilities, it produces a raw score for every token, called a logit. A function called the softmax turns scores into probabilities: it exponentiates each score and divides by the total, so bigger scores win bigger shares and everything sums to 1. Temperature divides the scores by T, just before the softmax. (Dividing the probabilities themselves by one number and renormalising would give back exactly the same list.) Here it is, with every symbol explained.
Work it by hand on the slide’s top three words, keeping only those three so the arithmetic stays short. At T = 1 the three probabilities are 0.3171, 0.0491 and 0.036, which add to 0.4022, so their shares are 78.84%, 12.21% and 8.95%.
At T = 0.5, raise each to the power 1/0.5 = 2: 0.31712 = 0.100552, 0.04912 = 0.002411, 0.0362 = 0.001296. They add to 0.104259, so the shares become 96.44%, 2.31% and 1.24%. The favourite has swallowed almost everything.
At T = 2, raise each to the power 1/2, a square root: 0.5631, 0.2216 and 0.1897, adding to 0.9744. The shares become 57.79%, 22.74% and 19.47%. The long shots have risen.
to: 78.84% at T = 1 → 96.44% at T = 0.5 → 57.79% at T = 2 (top three words only)
Why does it work this way? Look at the natural logarithms, which are the scores up to a constant: −1.149, −3.014 and −3.324. At T = 0.5 they become −2.297, −6.028 and −6.648: every gap between words has doubled. At T = 2 they become −0.574, −1.507 and −1.662: every gap has halved. The softmax only cares about gaps, so temperature is a dial on how far apart the words stand. As T approaches 0 the gaps become infinite and sampling turns into greedy decoding.
Even at a sensible temperature, sampling occasionally draws one of the thousands of tokens in the tail, each unlikely on its own but numerous together, and one bad draw can derail the rest of the text. Two rules cut the tail before drawing.
Top-k keeps only the k most likely tokens, rescales them to sum to 1, and samples from those. With k = 3 you get exactly the three-word example above: “to” 78.84%, “as” 12.21%, “the” 8.95%.
Top-p, also called nucleus sampling, keeps the smallest set of top tokens whose probabilities add up to at least p. Over the ten listed words the shares are 63.47%, 9.83%, 7.21%, 6.61% and 2.94% for the top five, a running total of 90.05%, so p = 0.9 keeps exactly five words. The advantage over top-k is that the cut adapts: when the model is confident, a few words already reach p and the nucleus is small; when it is unsure, the nucleus widens.
The faint bars are the model’s own shares of the ten listed words; the solid bars are what the rule leaves to sample from. Pick a rule, move its setting, then press Sample 20.
Probabilities are the lecture’s ten words after “The class was about”. Every rule is applied exactly to those ten, rescaled to sum to 100%; the rest of the vocabulary is left out, as in the hero. Draws are real random draws.
Greedy decoding has a blind spot: it commits to each word before seeing what that word leads to. A slightly less likely first word can open onto a much more likely continuation. Beam search hedges. It keeps the B best partial sentences (the beams) instead of one, extends every beam by every candidate word, scores each extension by its total probability (in practice the sum of log-probabilities, for the reason in Chapter 1), keeps the best B, and repeats. The lecture’s slide shows the search as a growing tree, from “start” through branches like “the”, “green” and “witch”, with only the best few branches kept alive at each level.
Each branch carries the probability of that word given the words before it. Choose a beam width and press Search: kept beams glow, pruned ones fade, and each two-word ending shows its total probability.
An illustrative tree: the words echo the lecture’s beam-search figure, but the figure shows no probabilities, so these are made up to show the effect. A real search runs over the whole vocabulary at every step.
Width 1 is greedy. Width 2 finds a sentence almost twice as likely, because it kept “a” alive long enough to discover “a witch”. Width 3 finds nothing better here but does more work. That is the whole trade-off in miniature.
The lecture closes the section with a slide of considerations, and each maps onto an agent decision. Every sampling strategy has trade-offs. Temperature, top-p and top-k are the tools for diversity. The right choice depends on the task: maths and code are more predictable, so they suit a lower temperature, while creative writing may benefit from more diversity. And it depends on your budget: beam search has higher time and memory costs, because it carries B sentences instead of one.
For an agent builder that translates into three habits. When the model must write something a program will parse, such as a tool call or a JSON object, where one wrong character breaks it, keep the temperature low. When you want many different attempts at a problem, as in the repeated sampling of Chapter 8, you need a temperature above zero, or every attempt will be the same. And when you debug an agent, remember that two runs can differ purely because of the draw, so compare distributions of runs, not single runs. (Our sampling and decoding lesson takes every rule further.)
Here is a complete sampler, written from scratch. It takes the model’s raw scores for the next position and applies temperature, top-k and top-p in that order.
pythonimport numpy as np def sample_next(logits, temperature=1.0, top_k=None, top_p=None, rng=np.random.default_rng()): """Pick one token id from the model's raw scores: a vector with one score per vocabulary token.""" if temperature == 0: # greedy: always the top token return int(np.argmax(logits)) z = logits / temperature # temperature scales the scores, not the probabilities order = np.argsort(z)[::-1] # token ids, most likely first if top_k is not None: order = order[:top_k] # keep only the k best p = np.exp(z[order] - z[order].max()) # softmax over the kept tokens; p = p / p.sum() # subtracting the max avoids overflow if top_p is not None: keep = np.searchsorted(np.cumsum(p), top_p) + 1 # smallest prefix reaching top_p order, p = order[:keep], p[:keep] / p[:keep].sum() return int(rng.choice(order, p=p)) def generate(model, ids, eos_id, max_new=50, **how): for _ in range(max_new): logits = model(ids)[-1] # scores for the next position only nxt = sample_next(logits, **how) ids = ids + [nxt] # the draw becomes part of the input if nxt == eos_id: # the model chose "end of text" break return ids
Trace it on the slide’s list with top_p=0.9 and temperature 1. The sorted shares of the ten words run 0.6347, 0.0983, 0.0721, 0.0661, 0.0294 and so on; their running totals are 0.6347, 0.7330, 0.8050, 0.8711, 0.9005. The first total to reach 0.9 is the fifth, so searchsorted returns index 4, keep is 5, and the draw is made among “to”, “as”, “the”, “a” and “two”, exactly as the device showed. Notice also the loop in generate: each drawn token is appended to the input, so every draw changes what the model sees next. That is how one early lucky draw sends a whole text, or a whole agent run, down a different path.
The lecture closes this part with three pointers for further reading: a Hugging Face blog post on decoding strategies in large language models, the Hugging Face course’s chapter on tokenizers, and Stanford’s course notes on deep generative models. Our sampling and decoding and tokenization lessons cover the first two from zero.
Chapter 3
Follow four words through a decoder-only transformer, from lists of numbers to next-word probabilities, and see why attention replaced recurrence
Guess the word after “The class was about”. Without noticing, you glance back. “Class” tells you this is about a lesson. “About” tells you a topic, or the word “to”, comes next. “The” tells you almost nothing. Some earlier words matter a lot and others barely at all, and which ones matter depends on what you are trying to predict.
A model needs the same skill: for each word, decide how much to look at each earlier word. That mechanism is called attention, and the lecture builds it on screen, one piece at a time, for exactly these four words. Before it starts, it puts up a slide of its own: pay close attention here, because if you get lost here, much of the rest of the lecture will be very confusing. The same is true of this lesson. Chapters 4, 7 and 10 all lean on what follows.
The lecture lists the ancestors. N-gram models (before the 2010s) counted how often short runs of words appear, say three in a row, and predicted from those counts; they could never see further back than a few words. The first neural language model (2003) learned a list of numbers for each word instead of counting. Recurrent neural network language models (2010) read one word at a time and carried a running summary, a hidden state, from each word to the next. Long short-term memory networks, invented in 1997 and popularised for language models in 2014, added gates that decide what the running summary keeps and forgets.
Then came the transformer, in the 2017 paper “Attention Is All You Need” by Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Lukasz Kaiser and Illia Polosukhin. The original design had two halves: an encoder that reads an input (a sentence to translate, say) and a decoder that writes the output. The language models agents run on keep only the second half. The lecture blanks out the encoder on the paper’s own figure and builds a decoder-only transformer. (The paper has a lesson of its own, and understanding attention builds the mechanism slowly.)
Each token’s number is swapped for a learned list of numbers, its embedding. You can think of it as the word’s coordinates in a space where related words sit near each other. The slides draw each embedding as six dots and an ellipsis; real models use thousands of numbers per token. Call that length d, the model’s width.
Attention, as Step 2 below will show, compares words only by the numbers in their vectors, a dot product that never looks at which position those numbers came from. So on its own it has no built-in sense of order: it would treat “class was the about” like “the class was about”, unless we tell it otherwise. That is why a second list, a positional embedding that encodes “first”, “second”, “third”, is added to each word’s embedding. Stack the four results as rows and you have a table of numbers with 4 rows and d columns: the whole sentence, ready to process.
Before attention, each row goes through a layer norm: its numbers are shifted and rescaled to a standard centre and spread. Nothing about the meaning changes; it simply keeps the numbers in a stable range as they pass through dozens of layers.
Now three learned matrices, WQ, WK and WV, each multiply the table (plus a small learned offset, drawn on the slides as a short row of dots). Out come three new tables with one row per word: the queries Q, the keys K and the values V. The slides draw them three numbers wide.
A library is the usual picture, and it fits well. A word’s query is what it is looking for. Every word’s key is the label on its spine, saying what it offers. Its value is the content it hands over if chosen. To decide where to look, each word compares its query with every key.
Picture each word’s list of numbers as an arrow in space; two arrows pointing the same way have a large dot product, two pointing oppositely have a very negative one, and two at right angles have about zero. The comparison is a dot product: multiply the two lists entry by entry and add the results. It is large when the two point the same way. Doing it for every pair at once is one matrix multiplication, Q times K turned on its side (KT, “K transposed”), and gives a 4 × 4 grid of scores: row i, column j says how well word i’s query matches word j’s key.
One more thing about training, which Step 6 below explains fully: a single pass over a sentence predicts the next word at every position at once, not just at the end. Three operations turn raw scores into attention. First, the causal attention mask blanks every cell where column j comes after row i: the upper-right triangle of the grid, drawn black on the slides. A word may look at itself and at earlier words, never at later ones. When the model writes text the future does not exist yet; and since training predicts every position at once, as just noted, a word that could see the next word would simply copy it instead of learning to predict it.
Second, every score is divided by √dk, the square root of the key length. Dot products of long lists grow with their length, and very large scores would make the next step pick one word with near certainty, which is hard to learn from; dividing keeps them in a workable range.
Third, a softmax runs along each row (the same function as in Chapter 2): the scores become weights between 0 and 1 that add up to 1. Blanked cells become exactly 0. Finally, each word’s weights mix the value rows: its output is the weighted average of the values of the words it attends to. Try it on the four words.
Each row is a word looking back; each column is a word being looked at. Pick a row, then switch the mask and the scaling off and on. The bar under the grid is that word’s output recipe: how much of each value it mixes in.
Toy numbers: the word vectors, positions and weight matrices are small seeded random matrices, six numbers per word and three per query, key and value, as the slides draw them. The pattern is not a trained model’s. The mechanics are exact: scores QKT, the mask, division by √3, a softmax per row, then a mix of the value rows.
The mixed values are only dk numbers wide, so one more learned matrix, WO, projects them back to the model’s width d. Then comes a residual connection: the layer’s input is added back to its output. The layer therefore only has to learn a correction to what it received, and whatever came in flows on upward even if the layer contributes little. That is what makes very deep stacks trainable.
One set of queries, keys and values gives one pattern of attention. The lecture’s next build turns this into multi-head attention in three moves. First, multiple weight matrices: several independent sets of WQ, WK and WV, each called a head, each computing its own grid and its own mix, all in parallel. Second, concatenate: the heads’ outputs are laid side by side into one wider table. Third, a bigger projection matrix WO maps that wide table back to width d. A layer with several heads can attend to several different things at once.
Now the lecture zooms out. All of that machinery becomes one box, a multi-head, causal self-attention layer (“self” because the words attend to each other). After it comes another layer norm, then a feed-forward network: a small two-layer network applied to each word’s row separately, where much of the per-word processing happens. It gets a residual connection too. Attention plus feed-forward, each with its layer norm and residual, is one transformer block, and a model stacks N of them, one on top of another.
At the very top sits a final layer norm and the LM head (language-model head): a matrix that turns each word’s row into V scores, one per vocabulary token. A softmax on each row gives a next-token distribution at every position: after “The”, after “The class”, after “The class was” and after “The class was about”. In training, all four predictions are checked at once from a single pass. When generating, only the last one is used; that is the model(ids)[-1] in Chapter 2’s code.
Here is the whole data flow as shapes, for our n = 4 words:
What is fixed and what is not? Every matrix above (the embeddings, WQ, WK, WV, WO, the feed-forward weights, the LM head) is learned in training and frozen when you use the model. What changes with every input is only the tables flowing through: the queries, keys, values and attention weights. When an input degrades, say a long tool result full of noise lands in the context, nothing inside the model adapts; every word of the noise simply becomes more keys to compare against and more values to mix.
The lecture shows the paper’s comparison of layer types. In it, n is the number of words, d the width, k a convolution’s window and r a restricted neighbourhood. The notation O(…) describes how a cost grows as n grows: O(n2) means that doubling n roughly quadruples it.
| Layer type | Work per layer | Steps that must run one after another | Longest path between two words |
|---|---|---|---|
| Self-attention | O(n2 · d) | O(1) | O(1) |
| Recurrent | O(n · d2) | O(n) | O(n) |
| Convolutional | O(k · n · d2) | O(1) | O(logk n) |
| Self-attention (restricted) | O(r · n · d) | O(1) | O(n/r) |
The lecture draws two conclusions from the first two rows. O(1) sequential operations is a huge deal for parallel training on GPUs. A recurrent network cannot process word 2 until word 1 is done, so a thousand-word text takes a thousand steps in a row. Attention computes every position at the same time, which is exactly what a GPU, with its thousands of small cores (Chapter 7), is built for. The price is the other conclusion: O(n2) in sequence length means long context becomes expensive. The last column matters too: in attention any word reaches any other in one step, while a recurrent network must pass information through every word in between.
The next slide, “Transformers were better and cheaper”, shows the payoff on translation; the table keeps a selection of its rows. BLEU scores how closely a machine translation overlaps human ones (higher is better); training cost is counted in FLOPs, individual arithmetic operations. One baseline below, MoE, is itself an early mixture-of-experts model, a design Chapter 4 explains; you do not need that design to read this table, only its BLEU and FLOPs numbers.
| Model | BLEU, English to German | BLEU, English to French | Training FLOPs, German / French |
|---|---|---|---|
| GNMT + RL | 24.6 | 39.92 | 2.3 × 1019 / 1.4 × 1020 |
| ConvS2S | 25.16 | 40.46 | 9.6 × 1018 / 1.5 × 1020 |
| MoE | 26.03 | 40.56 | 2.0 × 1019 / 1.2 × 1020 |
| GNMT + RL ensemble | 26.30 | 41.16 | 1.8 × 1020 / 1.1 × 1021 |
| ConvS2S ensemble | 26.36 | 41.29 | 7.7 × 1019 / 1.2 × 1021 |
| Transformer (base) | 27.3 | 38.1 | 3.3 × 1018 |
| Transformer (big) | 28.4 | 41.8 | 2.3 × 1019 |
Read the English-to-German column. The small base transformer, trained with 3.3 × 1018 operations, beat the best ensemble (several models combined) of the older designs, ConvS2S at 26.36, which cost 7.7 × 1019: a better score for about 23 times less compute. On French, the big transformer’s 41.8 beat the same ensemble’s 41.29 for about 52 times less. Better and cheaper to train is why almost every model an agent calls today is a transformer.
Here is causal attention and a whole block in a dozen lines of numpy, following the slides step for step. (@ is Python’s matrix-multiply operator; np.triu_indices finds the upper-triangle cells to blank out for the mask.)
pythonimport numpy as np def causal_attention(X, Wq, Wk, Wv, Wo): """X: (n, d), one row per token. One head; learned offsets omitted.""" Q, K, V = X @ Wq, X @ Wk, X @ Wv # (n, d_k) each S = Q @ K.T / np.sqrt(Q.shape[1]) # (n, n): query i against key j, scaled S[np.triu_indices(len(X), k=1)] = -np.inf # causal mask: no looking ahead A = np.exp(S - S.max(axis=1, keepdims=True)) # softmax, row by row A = A / A.sum(axis=1, keepdims=True) # every row sums to 1 return (A @ V) @ Wo # mix the values, project back to d def layer_norm(X, eps=1e-5): # simplified: no learned scale or shift return (X - X.mean(1, keepdims=True)) / np.sqrt(X.var(1, keepdims=True) + eps) def block(X, p): X = X + causal_attention(layer_norm(X), p["Wq"], p["Wk"], p["Wv"], p["Wo"]) # + residual H = layer_norm(X) X = X + np.maximum(0, H @ p["W1"]) @ p["W2"] # feed-forward, row by row, + residual return X def next_token_probs(ids, params): X = params["embed"][ids] + params["pos"][:len(ids)] # (n, d) for p in params["blocks"]: # N blocks X = block(X, p) Z = layer_norm(X) @ params["lm_head"] # (n, V) scores Z = np.exp(Z - Z.max(1, keepdims=True)) return Z / Z.sum(1, keepdims=True) # a distribution at every position
Every line maps to a slide: the embedding plus position, the layer norm, the three multiplications, the scaled grid, the mask written as minus infinity above the diagonal, the row-wise softmax, the mix, WO, the residual additions, the feed-forward network and the LM head. The only thing missing is multiple heads, which would run causal_attention several times with different matrices and concatenate the results before WO.
Chapter 4
Trade attention’s ever-growing memory for a fixed-size one, see what that costs on long contexts, and learn why frontier models mix the two
Picture an agent two hundred steps into a job, its context long past a hundred thousand tokens. To write its next word, the model compares that word’s query with the key of every earlier token, and mixes in every earlier value. To do that, it must keep all those keys and values in memory, one pair per token per layer. Chapter 7 names this store the KV cache; for now, notice only that it grows with every token, forever.
The old recurrent networks of Chapter 3 had the opposite property: a running summary of fixed size, however long the text. They lost because they could not be trained in parallel. Is there a way to have both: a fixed-size memory and attention’s parallel training? The lecture answers with linear attention, built in three moves on the same four words.
Take the attention layer from Chapter 3 and delete one piece: the softmax (and the scaling with it). The scores QKT now go straight into the multiplication with V. The slide simply labels this “Drop Softmax”. It sounds like vandalism, but look at what is left: nothing but multiplications and additions. Everything in the layer is now linear.
With only multiplications left, we are free to regroup them. For ordinary numbers, (2 × 3) × 4 and 2 × (3 × 4) are both 24. The same holds for matrices, a property called associativity: (QKT)V equals Q(KTV). The slide’s instruction: let’s multiply K and V instead.
The answer is identical; the cost is not. QKT is the n × n grid, the triangle of Chapter 0, growing with the square of the number of tokens. KTV is a small dk × dv grid (3 × 3 on the slides) whose size does not depend on n at all. Multiplying K and V first skips the big grid entirely.
Now the slides walk the four words through, and something familiar appears. For “The”, compute its query q1, key k1 and value v1 as usual, then multiply the key (turned on its side) by the value: a 3 × 3 grid called the state, S1. The output is the query times the state, q1S1, followed by WO and the residual exactly as before.
Then two instructions appear on the slide. Clear unimportant memory: q1, k1 and v1 are never needed again, so they are thrown away. Save the state: only S1 is kept. For “class”, compute k2Tv2, add it to S1 to get S2, and output q2S2. Then S3 for “was” and S4 for “about”. This is the linear attention update rule.
Why is this still attention? Substitute what St actually is, the running sum of every kiTvi up to t, and the output unrolls into ot = qt(k1Tv1 + … + ktTvt). Each term qtkiT is just the scalar dot product qt·ki, so the whole sum becomes (qt·k1)v1 + … + (qt·kt)vt: each earlier value weighted by how well its key matches the current query, the same recipe as before, minus the softmax. And the causal mask comes for free, because St only ever contains tokens up to t.
Now count. With t tokens already in the context, softmax attention compares one new query with t keys and mixes t values: work and memory both grow with t. Linear attention updates one fixed grid and multiplies by it: the same small amount of work for token 10 as for token 100,000, and constant memory. It is a recurrent network again, but one whose training can still be laid out as big parallel matrix multiplications.
The lecture’s evidence comes from the paper “Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention”. Its figure measures the time and GPU memory of a forward and backward pass as the sequence grows from 29 = 512 to 216 = 65,536 tokens. Linear attention (and a family called Reformer) grows in a straight line with length; softmax attention grows with its square, in both time and memory. On the chart the softmax runs stop at 4,096 tokens, where they already take more than ten times the time and over twenty times the memory of linear attention, while linear attention carries on to 65,536. (Linear attention and RWKV goes deeper.)
The next slide begins “However…”: linear attention struggles more with longer context. To see why, think about what the state is. It is a sum. Every token adds its kTv on top of all the others, in the same few numbers. Reading with a query q returns the value you want, plus every other value weighted by how much its key resembles q. When keys are few and distinct, the extras are small. When there are more tokens than the state has room to keep apart, they pile up into a blur. Write facts into both kinds of memory and ask for them back.
Each fact is a random 16-number key paired with a 16-number value. Choose a memory, drag how many facts are written, and read how many come back right when asked by their keys. The top panel is what the memory holds.
Toy vectors: 256 random key and value pairs of length 16, seeded. Linear attention is exactly the slides’ version (drop the softmax, add each kTv into one 16 × 16 state). Softmax attention uses a sharpened score so that a matching key stands out, as a trained model’s would. A fact counts as recalled when the returned vector is closest to its own value; up to 32 facts are asked about. Real models learn their keys and add gates that manage what the state forgets; the lecture’s evidence is the RULER chart below.
Up to about 8 facts the fixed state recalls everything, because a 16-number key space has room to keep a handful of keys well apart. By 32 facts it recalls about half; by 64, a sixth; by 256, none, while its memory has not grown at all. The softmax cache recalls everything, but it holds 32 numbers per fact, 8,192 numbers at 256 facts against the state’s constant 256.
Real models are not toys, but they show the same shape. The lecture’s chart comes from RULER (“What’s the Real Context Size of Your Long-Context Language Models?”), which tests recall with needle-in-a-haystack tasks: hide a sentence such as “One of the special magic numbers for long-context is: 12345” inside long filler essays, then ask for the special magic number. Harder versions add a distractor with a different key (“…for large-model is: 54321”) or two values for the same key (the answer must be “12345 54321”).
On the chart, Llama 2 (7 billion parameters, a transformer) scores about 96, 92 and 85 at 1K, 2K and 4K tokens. RWKV-v5 (7 billion) and Mamba (2.8 billion), two designs built around a fixed-size recurrent state, fall much faster: RWKV from about 87 to 52, Mamba from about 63 to 17 over the same range. At 8K all three fall sharply, Llama 2 to about zero; the chart draws Llama 2’s and Mamba’s falling segments dashed (Mamba’s from 2K on), while RWKV’s line stays solid.
The lecture’s next slide suggests what a cliff like that usually is. Nothing in the transformer or in linear attention enforces a strict context length, except the way the model processes positional embeddings. The model has to learn to use positions, so in practice it is limited to the context length it was trained with; positions it never saw in training are foreign to it.
The context-engineering reading (Chapter 10) adds two refinements. Models learn their attention habits from training data in which short texts are more common than long ones, so they have less experience with context-wide dependencies. And techniques such as position encoding interpolation stretch a model to longer texts by mapping them onto the positions it was trained on, at some cost to its sense of position. The result, in the reading’s words, is a performance gradient rather than a hard cliff.
O(n2) in context length
Every token can refer back to every earlier token, which helps long context. The KV cache (its memory) grows with n.
O(n) in context length
New tokens interact with a compressed state, which is harder for long context. The state uses constant memory.
So in practice, the lecture says, people mix global and linear attention. Its figures, from Songlin Yang’s explainer of DeltaNet (a refined linear-attention layer), show two recipes. One alternates DeltaNet layers with sliding-window attention, softmax attention in which each token sees only a recent window (its mask is a diagonal band instead of a full triangle). The other uses DeltaNet almost everywhere and inserts a couple of global attention layers, full causal softmax attention over the whole context.
| Model | Wikipedia perplexity (lower is better) | Average common sense | Average retrieval |
|---|---|---|---|
| DeltaNet | 28.24 | 42.1 | 22.7 |
| + sliding-window attention | 27.06 | 42.1 | 30.2 |
| + global attention | 27.51 | 42.1 | 32.7 |
Perplexity measures how surprised a model is by real text, here Wikipedia; lower is better. Read across the rows. Common-sense scores do not move at all (42.1 each time). Retrieval, the needle-finding kind of task, jumps from 22.7 to 30.2 when sliding-window layers are alternated in through half the stack, and to 32.7 with just a couple of global ones. Exact lookup is precisely what a compressed state lacks, and a few full-attention layers give it back.
The lecture’s frontier example is Kimi K3 from Moonshot, and it labels the paper’s figure for us. The model’s blocks pair a feed-forward layer with either multi-head latent attention (a form of full attention, drawn as “Gated MLA”) or linear attention (“KDA”, Kimi Delta Attention). The figure marks the full-attention pair with a 1 and the linear-attention pair with a 3, which reads as three linear blocks for each full one. Its feed-forward layers are a mixture of experts, in which a router sends each token to a few of many expert networks while a couple of shared experts always run. A vision encoder (MoonViT-V2) feeds images in alongside the text. (Its predecessor has a paper lesson.)
The lecture ends the architecture section with four reasons, and each turns into a builder’s habit.
Chapter 5
Score one prediction with the lecture’s own example, then follow raw web pages through filters, mixes and midtraining into a model’s weights
Every probability on this page so far came from somewhere. Before training, a model’s weights are random numbers, and its guess for the word after “The class was about” is noise. Training nudges those weights, billions of times, until the guesses match real text. To nudge in the right direction you need one thing first: a number that says how wrong a guess was. That number is the loss.
The lecture uses one training document, and it begins with our familiar words: “The class was about how agents work, and how to build AI systems that use tools, retrieve information, and carry out tasks across multiple steps…”. Feed the model “The class was about”, and out comes the list you know: “to” 0.3171, “as” 0.0491, and so on. But the document says the next word is “how”, and the model gave “how” only 0.0124.
The slide scores that prediction in one line: Loss = −log(0.0124) = 1.907. Two things about that formula. The minus sign is there because the logarithm of a number below 1 is negative, and we want a positive score. And the logarithm makes the score behave sensibly: if the model had given “how” a probability of 1 (complete confidence, and right), the loss would be −log 1 = 0; as the probability of the true word shrinks toward 0, the loss climbs without limit. The loss is the model’s surprise at the real text.
There is a wrinkle worth catching. The slide’s 1.907 comes from a base-10 logarithm: log10(0.0124) = −1.907. Training code almost always uses the natural logarithm (base e), and ln(0.0124) = −4.390, so a training script would report a loss of about 4.39 for the very same prediction. In base 2 it would be 6.334, a unit called bits. None of them is wrong. Changing the base multiplies every loss by the same constant (natural-log losses are always 2.303 times their base-10 versions), so it changes the ruler, not the ranking of predictions or the direction training moves.
The curve is the loss, −log p, for every probability p the model might give the true next word. Drag p, or tap a preset. Then switch the base and watch the curve stay put while its ruler changes.
The probabilities are the lecture’s (“how” 0.0124 and “to” 0.3171 after “The class was about”); the slide’s loss of 1.907 is the base-10 value. The other values are computed exactly from p.
The next slide writes the general form, the cross-entropy loss, and labels every part of it.
Work it on the example. Only one yi is 1, the one for “how”; every other token’s term is multiplied by 0 and vanishes. So the sum collapses to a single term:
L = −(0 · log 0.3171 + 0 · log 0.0491 + … + 1 · log 0.0124 + …) = −log 0.0124 (1.907 in base 10, 4.390 in natural log)
That is why the general formula and the slide’s one-liner agree: cross-entropy against a single true token is just the surprise at that token. (Loss functions derives it from scratch.)
One document supplies many such scores. Thanks to the causal mask of Chapter 3, a single pass over “The class was about how agents work…” produces a prediction at every position at once: the word after “The”, after “The class”, and so on. The training loss is the average surprise over all positions of all documents in a batch. From that one number, training computes the gradient: for every weight in the model, which way and how much to change it to lower the loss. Every weight then takes a small step that way. Repeat on trillions of tokens and the probabilities you met in the hero emerge.
As data flow: in, a batch of token sequences cut from documents; out, one number, the average loss; then every weight moves. Nothing is frozen in this phase. The model learns exactly one thing: to be less surprised by the data it is shown. Which makes the next question the most important one in this chapter: what data?
The first and largest phase of training is pretraining: next-word prediction on an enormous amount of general text, mostly from the web. The lecture draws the standard pipeline as a row of stages. Start with raw HTML, the code of web pages. Text extraction pulls out the readable words. Language identification keeps the languages you want. Rule-based filters apply simple heuristics (too short, too repetitive). Quality filters use a trained model to score how useful a page looks. Deduplication removes copies. What remains is the final training data.
How much survives? The lecture shows the funnel from DataComp-LM (DCLM), which starts from pages extracted from Common Crawl, a public archive of the web. Step through it.
Each bar is the share of the original documents still left after that stage. Press Apply the next filter to run the pipeline one stage at a time.
Numbers from DCLM’s Figure 4 as shown in the lecture; every percentage is of the original documents in DCLM-Pool. The six heuristic filters reproduce an earlier recipe (RefinedWeb); the model-based filter is a fastText classifier. The order of the heuristic rows follows the figure.
Of every 100 documents, about 1.4 survive: roughly one in 71. The English filter removes half the pool in one stroke, the single biggest cut in absolute terms. Relative to what reaches it, though, the most selective stage comes last: of the 13.7% that pass every heuristic and deduplication, the model-based quality filter keeps only 1.4 points and discards 12.3.
What does survivor text look like? The lecture shows three real examples from DCLM, overlapping on one slide. A programming-language manual page for a function called EOSHIFT (an “end-off shift”), still carrying its navigation links (“Next: , Previous: DTIME, Up: Intrinsic Procedures”) and a table of boundary values. A forum thread in which someone asks for an anime about a legendary hacker and gets replies recommending Summer Wars (“It’s been dubbed by Funimation”), Dennou Coil and Real Drive, complete with usernames, timestamps and a “Register to view this Article” prompt. And an article that opens at 7 a.m. at a cooking school, where a chef explains to his baking class the difference between a pie and a tart (“A tart is more refined than a pie, while a pie is more rustic”), breaking off at the start of a pear recipe. Even after the filters, pretraining text is diverse, messy and full of leftovers from the page around it, and the model learns from all of it.
Some datasets target one domain. The lecture’s example is Nemotron-CC-Math. It starts from 98 Common Crawl snapshots, keeps pages whose addresses appear in existing maths datasets (229.54 million pages), renders each page as text with a simple text browser (Lynx) instead of the usual extractors, has a language model clean up the text, then removes duplicates and anything overlapping test sets. The result: 101.15 million documents and 133.26 billion tokens, about 60% mathematics (60.28%), then computer science (11.99%), physics (11.22%), statistics (7.50%), economics (3.18%), chemistry (1.71%) and other topics (4.12%).
And some datasets aim for breadth. The Pile, “an 800GB dataset of diverse text”, is shown as a map of its sources grouped into five kinds: academic (PubMed Central, ArXiv, FreeLaw, USPTO and more), internet (Pile-CC, OpenWebText2, StackExchange, Wikipedia), prose (Bibliotik, PG-19, BC2), dialogue (subtitles, IRC and others) and miscellaneous (GitHub code, DM Math).
Once you have sources, you must decide what fraction of training comes from each: web text, code, maths, books. That fraction is the data mix, and it is expensive to get wrong, because you only find out after training. The lecture shows Olmix’s recipe in three steps. Swarm: train many small, cheap proxy models, each on a different mix, and measure how well each performs. Regression: fit a formula that predicts performance from the mix, and check its predictions against the real results. Optimization: search the allowed mixes for the one the formula predicts will be best, and use that for the big run.
The lecture then introduces midtraining: a stage between general pretraining and the final polishing of Chapter 6, in which you change the data mix to focus more on your target domains, such as maths or code. Its example chart, from a project that curated 25 trillion tokens, sets weights for each domain (application and web development, business software and strategy, and clinical treatment and dentistry are among its rows) and each of five quality bands, and uses different weights in the main pretraining phase and in a later “cooldown” phase.
Does it help? The lecture’s next slide says that mixing more of your target-domain data into “pretraining” tends to help. The paper behind it (“The Finetuner’s Fallacy”) builds a specialised pretraining mix of general data plus the domain dataset repeated 10 to 50 times. Its chart compares two base models, each pretrained on 200 billion tokens, one with 33 repeats of the domain data and one without, as both are later finetuned on the domain. The specialised model reaches a lower domain test loss, in fewer finetuning steps, and loses less of its general ability along the way; the chart’s own title is “better in-domain generalization and less forgetting during FT”.
A last, schematic slide explains why a middle stage helps at all. Picture the weights as a point on a landscape. Jumping straight from the pretrained point to the target with ordinary finetuning crosses a region where the training signals pull against each other. Stopping at a midtraining point on the way gives a path where they agree. Midtraining, in the paper’s title, bridges pretraining and post-training distributions.
For an agent builder, this chapter is the reason a model has the habits it has. Its knowledge, its formats, the kinds of text it finds natural and the domains it is strong in were all set by its data and its mix. When you choose a model for a coding agent or a maths agent, you are really choosing what it read.
Chapter 6
Turn a text predictor into an assistant with instruction finetuning, human preferences and verifiable rewards, then train it for agent work
Picture a model that has read the filtered web and nothing else. Its one skill is continuing text. Type “Please answer the following question. What is the boiling point of Nitrogen?” To a pure predictor of web pages, a second quiz question, or a whole page of them, is a perfectly plausible continuation. The model may well know the answer; nothing has taught it that its job is to give one.
Post-training is the family of methods that fix this: training after pretraining, on far smaller and far more carefully chosen data, so that the model follows instructions, prefers good answers and, more and more, acts. The lecture walks through four kinds.
The first is the most direct. Collect many examples of instructions paired with good responses, and keep training on them with exactly the same next-word loss as Chapter 5. The lecture’s figure comes from “Scaling Instruction-Finetuned Language Models”, and its examples show the idea. Instruction finetuning: “Please answer the following question. What is the boiling point of Nitrogen?” paired with “−320.4F”. Chain-of-thought finetuning: “Answer the following question by reasoning step-by-step. The cafeteria had 23 apples. If they used 20 for lunch and bought 6 more, how many apples do they have?” paired with a worked answer: 23 − 20 = 3, then 3 + 6 = 9.
The paper trains on a mixture of about 1.8 thousand tasks at once, multi-task instruction finetuning. The payoff is generalisation to tasks it never saw. Asked “Can Geoffrey Hinton have a conversation with George Washington? Give the rationale before answering”, the finetuned model reasons that Hinton is a British-Canadian computer scientist born in 1947, Washington died in 1799, so they could not have talked: the answer is “no”. It learned a general habit, when given an instruction, carry it out, not just the 1.8 thousand tasks.
Instruction data says what a good answer looks like. But for many prompts there is no single right answer, only better and worse ones. The lecture’s second method learns from people’s judgements, in the three steps of “Training language models to follow instructions with human feedback”.
As data flow: step 1 changes the policy’s weights using human-written answers; step 2 trains a different model, and that reward model is then frozen; step 3 changes the policy’s weights again, now guided by the frozen reward model’s scores instead of by a person for every answer. This is the RLHF (reinforcement learning from human feedback) on the stack diagram of Chapter 0. (The paper has a lesson of its own.)
The lecture is blunt about the weaknesses, in a slide inspired by another Stanford course (CS329X). Human annotators can be unreliable: tired, rushed or simply wrong. Everyone has different preferences, so “what people want” is not one thing. The process creates bad incentives for the model: if raters reward answers that please them, the model learns sycophancy (telling people what they want to hear) and authoritativeness (sounding confident and expert whether or not it is right). And the data itself can be sourced unethically.
For an agent builder the second-to-last point bites hardest. A model tuned to please will tend to agree with a user’s wrong premise, or report a task as done with more confidence than the evidence supports. Your harness, not the model’s tone, has to check.
What if the reward did not come from a person, or from a model imitating people, but from a program that can simply check the answer? That is reinforcement learning with verifiable rewards (RLVR). The lecture lists the kinds of checks: unit tests (does the code pass?), a Lean proof (a proof-checking program accepts or rejects a maths proof), exact match (is the final answer the known one?) and system state (is the file there, is the database row correct?). It is great for maths and code, where such checks exist. And it comes with a warning: be careful about the design of the verifier.
The lecture’s figure shows a method called GRPO, from the DeepSeekMath paper. A question q goes into the policy model, which writes not one answer but a group of G answers, o1 to oG. Each is scored, giving rewards r1 to rG. A frozen reference model (a copy of the model from before this training) adds a penalty for drifting too far from where the model started, labelled KL (Kullback-Leibler divergence, a number that grows the more two probability distributions disagree). Then a group computation turns the rewards into advantages A1 to AG: how much better each answer did than its own group. Answers that beat their group become more likely; the rest become less likely. Only the policy is trained; the reference and reward models stay frozen. (DeepSeekMath has the full method.)
Work it on a question the lecture uses later: a movie ticket and a popcorn together cost $18; the ticket costs $6 more than the popcorn; what does the ticket cost? (Popcorn p, ticket p + 6, so 2p + 6 = 18, p = 6, and the ticket costs $12.) Suppose the policy writes eight answers, four of them $12. With an exact-match verifier the rewards are four 1s and four 0s, the average is 0.5, so each right answer gets A = 1 − 0.5 = +0.5 and each wrong one A = 0 − 0.5 = −0.5. Now see what a careless verifier does to the same group.
Eight sampled answers to the ticket question (right answer: $12). Dots are rewards; bars are advantages, reward minus the group average. Switch the verifier and watch which answers get pushed up.
The question and its answer ($12) are the lecture’s example of a reasoning model’s self-correction. The eight answers are illustrative, and the advantage is the simplest group computation (reward minus the group average); GRPO’s own formula adds normalisation and the KL penalty from the reference model.
The sloppy check pays the hedge “$12 or $18” as if it were right, and gives it a positive advantage. Train on enough groups like that and the model learns that hedging pays: it is being rewarded for fooling the checker, not for solving the problem. That is the lecture’s warning about verifier design in one picture. A verifier is part of your training data, and every hole in it is a lesson the model will learn.
The last kind of post-training targets agents directly. An agent needs practice at the whole job: reading a code base, running tests, editing files, trying again. The lecture’s example is SWE-smith, which manufactures such practice from real software projects.
What does practice buy? The lecture shows the results on SWE-bench, a benchmark of real GitHub issues that an agent must fix (the Lite and Verified columns are two versions of it; scores are the percentage resolved). The “system” column is the harness, the agent program wrapped around the model. The table keeps a selection of the lecture’s rows.
| Model | System (harness) | Training examples | Lite | Verified |
|---|---|---|---|---|
| GPT-4o | SWE-agent | · | 18.3 | 23.0 |
| GPT-4o | Agentless | · | 32.0 | 38.8 |
| Claude 3.5 Sonnet | SWE-agent | · | 23.0 | 33.6 |
| Claude 3.7 Sonnet | SWE-agent | · | 48.0 | 58.2 |
| SWE-gym-32B | OpenHands | 491 | 15.3 | 20.6 |
| SWE-fixer-72B | SWE-Fixer | 110k | 24.7 | 32.8 |
| R2E-Gym-32B | OpenHands | 3.3k | · | 34.4 |
| SWE-agent-LM-7B | SWE-agent | 2k | 11.7 | 15.2 |
| SWE-agent-LM-32B | SWE-agent | 5k | 30.7 | 40.2 |
Two readings, both useful to a builder. First, the model matters: inside the very same SWE-agent harness, GPT-4o resolves 23.0% of Verified tasks while SWE-agent-LM-32B, the SWE-smith paper’s own open model of 32 billion parameters, trained on 5,000 examples, resolves 40.2%, the best open-weight result in the lecture’s table. Second, the harness matters: GPT-4o scores 23.0 in SWE-agent but 38.8 in a different system, Agentless. The best closed model in the table, Claude 3.7 Sonnet in SWE-agent, reaches 58.2. Training for the agent’s job and engineering the harness around the model are two separate levers, and this course is mostly about the second (Chapter 11).
Chapter 7
See why the first word takes a while and the rest stream steadily, measure both with the roofline model, and speed up writing by letting a small model guess
Paste a long document into a chat model and ask a question. For a moment nothing happens. Then the answer begins, and the words arrive at a steady clip, one after another, until it stops. Two quite different kinds of work are hiding in that experience, and the difference between them sets the speed and the cost of every agent you will build. Running a trained model is called inference, and the lecture splits it in two.
In prefill, the model reads the entire prompt in one go. The lecture’s slide shows our four words, “The class was about”, going through the attention layer together as full tables: all four queries, all four keys and values, the whole 4 × 4 grid of scores. Everything from Chapter 3 happens for every prompt token at once. Two things come out: a key and a value for every prompt token in every layer, which are kept; and the probabilities for the first new token.
Then comes decode, the steady stream. Now only one new token goes through the model per step. For it, the model computes one query, one key and one value. The query is compared with all the stored keys, the output mixes all the stored values, and the new key and value are added to the store. Then, as the slides say in red, it will clear unimportant memory: the query and every intermediate number are thrown away. The slides step through “The”, “class”, “was” and “about” this way, and end on a picture of just two things left standing, KT and V: the KV cache.
The lecture’s other figure shows the same rhythm on “The quick brown fox”: prefill computes the key and value vectors of all four prompt words and produces “jumps”; each decode step then reuses the cached vectors to produce one more word, “over”, and so on.
Why keep a cache at all? Without it, writing each new token would mean recomputing the keys and values of the entire text so far: a full prefill for every single word. With it, each decode step does the work for one token only. The price is memory that grows with the context, the growing column of Chapter 4.
To see why these two phases behave so differently, the lecture looks at the hardware. A CPU, the ordinary processor in a laptop, has a large control unit, a large fast cache and a handful of powerful ALUs (arithmetic logic units, the parts that actually add and multiply). A GPU spends its chip on thousands of small ALUs instead, all fed from a large main memory called DRAM, which is where the model’s weights and the KV cache live. A GPU is superb at doing the same arithmetic on huge numbers of values at once, as long as the values can be delivered to its ALUs fast enough.
That “as long as” is the whole story, and the roofline performance model draws it. A chip has two limits. Its peak FLOP/s is how many arithmetic operations it can do per second. Its memory bandwidth is how many bytes per second it can move from memory to the ALUs. Which limit you hit depends on your work’s arithmetic intensity: how many operations you perform for each byte you move.
Now work it, with illustrative numbers, for one weight matrix of a model 4,096 numbers wide: 4,096 × 4,096 weights, each stored in 16 bits (2 bytes), about 33.6 million bytes. Pushing n tokens through it takes 2 × n × 4,0962 operations: one multiplication and one addition per weight, per token. Take an illustrative GPU with a peak of 1,000 trillion operations per second and a bandwidth of 3 trillion bytes per second. Its ridge point is 1,000 ÷ 3 = 333 operations per byte.
Decode, n = 1. Operations: 2 × 4,0962 = 33.6 million. Bytes: the 33.6 million bytes of weights, plus a few kilobytes for the token’s own vectors. Intensity: about 1.0 operation per byte. Attainable: min(1,000, 3 × 1.0) = 3 trillion operations per second, 0.3% of peak. The chip spends almost all its time fetching weights, each used once and thrown away.
The lecture’s 4-token prompt. Four times the operations for almost the same bytes: intensity about 4.0, so 12 trillion operations per second, 1.2% of peak. Still bandwidth-bound: a short prompt cannot keep a GPU busy.
Prefill of 2,048 tokens. Operations: 2 × 2,048 × 4,0962 = 68.7 billion. Bytes: the 33.6 million of weights plus 2,048 tokens’ inputs and outputs, another 33.6 million, so 67.1 million. Intensity: 1,024 operations per byte, past the ridge at 333. Compute-bound: the full 1,000 trillion operations per second.
decode: I ≈ 1 → 0.3% of peak prefill of 2,048: I ≈ 1,024 → 100% of peak (illustrative chip)
That is the lecture’s second roofline slide in numbers: decode sits on the slanted, bandwidth-bound side; prefill sits on the flat, compute-bound side. Drag the batch of tokens and watch the point climb the roof.
The roof is the most work per second this chip can do at each arithmetic intensity. Drag how many tokens go through the weight matrix together, or tap a preset, and watch where the point lands.
Illustrative numbers: one 4,096 × 4,096 weight matrix in 16-bit numbers, and a chip with a peak of 1,000 trillion operations per second and 3 trillion bytes per second of bandwidth. The counts leave out attention’s own reads of the KV cache and every other layer. The shape of the roof and the placement of decode and prefill are the lecture’s.
The two phases give the two numbers the lecture uses to measure speed. Time to first token (TTFT) is how long you wait before anything appears: essentially the prefill. Time per output token (TPOT) is the steady gap between words during decode. For an answer of N tokens the total is about TTFT + (N − 1) × TPOT. With illustrative values of 0.5 seconds and 20 milliseconds, a 500-token answer takes 0.5 + 499 × 0.02 = 10.48 seconds, almost all of it decode.
An agent stresses both. Every step re-sends a long, growing context, which is prefill and TTFT; long reasoning traces are thousands of decode steps, which is TPOT multiplied up. And because the KV cache of the unchanged beginning of the context can be kept from one step to the next, only the newly added tokens need a fresh prefill, which is exactly why Chapter 10 tells you to add to the context sequentially rather than rewrite its beginning.
Decode’s weakness is also an opportunity. Each step fetches every weight from memory to produce a single token, while the ALUs sit mostly idle. Checking several tokens in one pass costs about as much as producing one, because it is a short prefill. So: let a small, fast draft model guess the next few tokens one at a time, then have the big target model check all the guesses in one pass. The target keeps every guess up to the first one it disagrees with, and replaces that one with its own choice. The target has the final say on every token; the draft only saves it work.
The lecture’s slide: after “Once”, the draft model guesses “upon”, “a”, “time”, “there”, four cheap steps. The target model checks all four in one pass: “upon”, “a” and “time” are accepted; “there” is rejected. Count the target’s work: one pass produced 3 accepted tokens plus its own correction at the rejected position, 4 tokens, where plain decoding needs 4 target passes for 4 tokens.
Top: press Run the example to watch the slide’s draft and check. Bottom: the time to write 128 tokens as the draft length K grows, for code and for news summaries. Drag K and switch the task.
The example is the lecture’s (“Once” → upon, a, time, there; three accepted). The timings are read by eye from the chart the lecture shows from “Accelerating Large Language Model Decoding with Speculative Sampling” (mean time to sample 128 tokens, acceptance rate and loop time against K), so treat them as approximate.
On the lecture’s chart, sampling 128 tokens takes about 1,800 milliseconds with no drafting (K = 0). Drafting 3 or 4 tokens at a time cuts that to about 730 to 770 ms on code (HumanEval), roughly 2.3 to 2.5 times faster, and to about 930 to 940 ms on news summaries (XSum), roughly 1.9 times faster. Two forces pull against each other as K grows. Each loop costs more (about 14 ms at K = 0, about 28 ms at K = 7), and later guesses are less likely to be accepted: by K = 7 the acceptance rate has fallen to about 0.68 on code but about 0.45 on summaries. Hence the slide’s annotation: some domains are more predictable than others. Code tends to have more predictable continuations (a closing bracket, a variable name used again), so drafts land more often; open-ended prose is harder to guess. (Speculative decoding has the acceptance rule in full.)
Chapter 8
Buy better answers after training is over: think out loud, set the reasoning level, sample many times and check, or let the samples vote
Ask a friend a hard arithmetic question and demand an answer in one second. Then ask again and give them a minute and a piece of paper. Same brain, better answer. Everything in Chapters 5 and 6 improved the model by training it; this chapter is about the paper and the minute. Spending more computation when the question is asked, instead of when the model is built, is called inference-time scaling, or test-time scaling.
There is a simple reason it can work at all. Recall Chapter 7: every token a model writes is one more full pass through all its layers. A model that answers in five tokens has used five passes; one that writes three hundred tokens of working first has used three hundred, and every later token can attend to the intermediate results written earlier. Longer answers are, quite literally, more computation.
The lecture starts with the classic demonstration, from “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models”. Both prompts contain one worked example and then a new question. In standard prompting the example is “Roger has 5 tennis balls. He buys 2 more cans of tennis balls. Each can has 3 tennis balls. How many tennis balls does he have now? A: The answer is 11.” Asked next about a cafeteria that had 23 apples, used 20 to make lunch and bought 6 more, the model answers “The answer is 27”. Wrong.
In chain-of-thought prompting the only change is the example’s answer, which now shows its work: “Roger started with 5 balls. 2 cans of 3 tennis balls each is 6 tennis balls. 5 + 6 = 11. The answer is 11.” Now the model imitates the style: “The cafeteria had 23 apples originally. They used 20 to make lunch. So they had 23 − 20 = 3. They bought 6 more apples, so they have 3 + 6 = 9. The answer is 9.” Right. No weight changed; the prompt simply taught the model to spend more tokens, and so more passes, before committing to an answer.
A reasoning model is trained, not just prompted, to think at length before answering. The lecture contrasts two styles on one question: a jacket costs $80; the store raises its price by 25%, then offers a 25% discount on the new price; what is the final price? A traditional model answers directly and correctly: $80 × 1.25 = $100, then $100 × 0.75 = $75.
The reasoning model’s trace is longer and more interesting. It first reaches for a shortcut: adding 25% and subtracting 25% appears to cancel, suggesting $80. Then: “Wait… That cancellation would require both percentages to use the same base amount. Here, the discount applies to the increased price.” It redoes the work (25% of $80 is $20, giving $100; 25% of $100 is $25, giving $75), checks it (“$80 + $20 − $25 = $75. Okay.”) and concludes that the final price is $75, $5 less than the original.
The next slide shows the same habit on the movie-ticket question from Chapter 6. Inside its <think> tags, the model lets p be the popcorn price, so the ticket costs p + 6, and writes p + 6 = 18: p = 12, so the ticket is 18. Then comes the line the slide circles in red, “Wait, that can’t be right.”, followed by “I only used the ticket price, not the total of ticket + popcorn.” Aha: p + (p + 6) = 18, so 2p + 6 = 18, 2p = 12, p = 6, and the ticket costs 12. The slide’s annotation says what training produced: the model learns to self-correct. Those “aha” moments are extra tokens, spent checking.
Thinking longer costs time and money, and not every question needs it. So reasoning models expose a dial. The lecture shows it on a chart from OpenAI of models on a benchmark called Agents’ Last Exam, plotting each model’s score against its cost per task. Each model is drawn as a short curve rather than a point: one point per effort setting. Zoomed in, GPT-6 Astra’s curve runs from Low through Medium up to Max, higher score and higher cost at each step, and the app exposes the choice as a slider, set to “High” in the screenshot. (The chart plots several other models the same way: GPT-6 Sol and Luna, GPT-5.6 Sol and Luna, Claude Opus 5, and Claude Fable 5 with an Opus 4.8 fallback, each its own curve.)
With open models you often set the level yourself, in plain text. For GPT-OSS, the slide shows three versions of the system prompt (the standing instructions at the top of the context): “Reasoning: low”, “Reasoning: medium” and “Reasoning: high”, followed by the same user prompt, “Solve this math problem…”. The model was trained to think less or more depending on that one line.
How do you train that? The lecture shows two recipes. Inkling, from Thinking Machines, subtracts a price for every token of thinking from the reward:
Work it with illustrative prices: λ = 0.0002 per token at low effort and 0.00002 at high effort. A right answer in 500 tokens earns 1 − 0.0002 × 500 = 0.90 at low effort; a right answer in 2,000 tokens earns 1 − 0.0002 × 2,000 = 0.60. At high effort the same two earn 0.99 and 0.96. At low effort, brevity is worth 0.30 of reward; at high effort only 0.03, so long careful thinking is almost free. Training on those rewards teaches the model what each level means.
Kimi K3, from Moonshot, uses a blunter rule: a reward of 1 for a correct answer, 0 for a wrong one, and −1 for an answer that is wrong and exceeds its token budget. It adds multi-teacher distillation: nine teacher models, one for each pairing of task family (general tasks, general agent work, coding agent work) with effort level (high, medium, low), and a single student model trained to imitate all of them, so that one model can behave like any of the nine.
Thinking longer is one way to spend compute. Another is to answer many times. The lecture’s figure, from “Large Language Monkeys”, has two steps. Step 1: generate many candidate solutions (for a problem that begins “Input a number from stdin and…”, the model writes programs starting data = {}, x = int(input()) and import requests). Step 2: use a verifier to pick a final answer; the slide’s examples are unit tests, proof checkers and majority voting. Here only x = int(input()) passes.
Each step has its own problem. Coverage: can we generate a correct solution at all, in some sample? Precision: can we identify the correct solution among the samples? Coverage is measured as pass@k: the fraction of problems for which at least one of k samples is correct.
Work it with the lecture’s chart. On MATH, with a verifier that can always recognise a right answer (the chart calls it an oracle verifier), Llama-3-8B-Instruct solves about 27% of problems with one sample. If every problem gave it the same p = 0.27, then: k = 1 gives 27%; k = 2 gives 1 − 0.732 = 1 − 0.533 = 46.7%; k = 5 gives 1 − 0.735 = 1 − 0.207 = 79.3%; k = 10 gives 1 − 0.043 = 95.7%.
The real curve climbs far more slowly: on the chart it passes one attempt of GPT-4o (about 64%) only after roughly ten or twenty samples, and reaches about 98% only at ten thousand. The formula is right, but only for a single problem. Real problems differ: some the model solves half the time, others one time in a thousand, and a problem with p = 0.001 needs about 700 samples to reach even a 50% chance. Average the formula over problems of mixed difficulty and you get the slow, steady climb in the chart. Compare both below.
Two curves share the same average one-sample success p. Dashed: every problem equally hard. Solid: a mix of easy and hard problems. The flat line is one attempt of GPT-4o on MATH. Drag p and the number of samples k.
The 27% (Llama-3-8B-Instruct, one sample) and the 64% line (one attempt of GPT-4o) are read by eye from the MATH panel of the lecture’s Large Language Monkeys chart. The dashed curve is the exact formula. The solid curve is illustrative: problem difficulties follow a Beta distribution with the same mean p (shape a = 1.5p, b = 1.5(1 − p)), and coverage is averaged over them. Both assume a verifier that recognises every right answer.
The lecture’s chart has four panels, and all tell the same story. On formal maths proofs (MiniF2F), both Llama models start near 20% and climb to around 50% at ten thousand samples, well above GPT-4o’s single attempt of roughly 27%. On programming contests (CodeContests), Llama-3-70B-Instruct climbs from under 10% to near 40%, passing GPT-4o’s roughly 21% at around a hundred samples. On MATH and GSM8K (grade-school word problems) both models approach 100%. With enough samples and a trustworthy verifier, a small model can beat one attempt of a much bigger one. The catch is that word: verifier. MATH and GSM8K here use an oracle that knows the answer; in real use, precision is the hard part.
Often you cannot check an answer. The lecture’s answer is self-consistency: sample several chain-of-thought answers and choose the one the most samples agree on. Its slide uses a word problem: Janet’s ducks lay 16 eggs a day; she eats three for breakfast and bakes muffins with four; she sells the rest at $2 each. The greedy answer (Chapter 2) counts the eggs she uses, 3 + 4 = 7, as the ones she sells and says $14. Three sampled paths answer $18 (16 − 3 − 4 = 9 eggs, times $2), $26 (an arithmetic slip) and $18 again. The vote picks $18, which is right. (Lecture 1’s lesson lets you run this vote yourself, and the lecture’s reading files it under “voting”, one of its workflow patterns, in Chapter 11.)
Accuracy on a science-question benchmark (ARC Challenge) as more reasoning paths are sampled. Pick how many paths to sample.
Values are read by eye from the lecture’s chart from “Self-Consistency Improves Chain of Thought Reasoning in Language Models” (ARC Challenge), so treat them as approximate. Sample & Rank samples several paths and ranks them to keep a single one; greedy decoding is a single path.
Three lessons sit in that chart. A single sampled path (about 36%) is worse than the greedy path (about 43%): randomness alone costs accuracy. By five paths the vote has already overtaken greedy (about 46.5%), and it keeps climbing to about 54% at forty. And ranking the samples to keep just one of them helps only a little (about 44% at forty): the gain comes from agreement across paths, because wrong paths tend to go wrong in different ways while right ones tend to land on the same answer. (The self-consistency and Large Language Monkeys papers each have a lesson.)
For an agent builder these are the knobs you control after the model is fixed: the reasoning level in the system prompt, the number of samples, and whether you have a verifier (tests, a checker, a schema) or must fall back on a vote. Each one trades money and latency for quality, and Chapter 7’s metrics tell you what you are paying.
Chapter 9
Give the model and the program around it a shared contract, then enforce it with constrained decoding, typed signatures, classifiers and tool calls
Your agent asks the model to pull the contact details out of an email and reply in JSON, the text format programs use to exchange data. Your code will read the reply with a JSON parser and pass the email address to the next step. The model replies: “Sure! Here is the contact info:” followed by the JSON. The parser stops at the first character, because “S” is not JSON, and your pipeline crashes. The model understood the task perfectly. It just did not honour the format.
The lecture’s final part, LLMs for Agents, opens with the agent diagram from Lecture 1 (planning and memory feeding the LLM core; the core driving tools; tools producing actions; the environment observed and acted on) with the LLM core circled. Everything so far has been about what happens inside that circle. Now the question is how the core talks to everything outside it.
The lecture’s motivating slide borrows the picture from Lecture 1’s reading: one big neural network on the left; on the right, a compound system in which several models pass work to each other and to a document search, a database, a calculator and a terminal. Increasingly many new AI results come from compound systems like that. And a compound system, the slide says, requires a shared contract of how its components will communicate.
Programs are strict readers. A calculator wants an expression, a database wants a query, the next model in a chain wants its input in a known place. Structured output means the model’s text follows a format the receiver can parse every time: JSON with known fields, a function call with typed arguments, or one label from a fixed list. Structured input means the prompt itself is laid out so each piece has a clear place. The lecture shows four ways to get there.
The first way works inside the decoding loop of Chapter 2. Constrained decoding forces the model to output a particular structure. The format is written as a pattern (a regular expression, or a JSON schema) and turned into a finite state machine (FSM): a small map of states in which each state lists which tokens may come next, and each token moves you to a new state. Before every draw, every token the current state forbids gets its probability set to zero (in practice, its score is set to minus infinity before the softmax), so the model can only ever choose among legal tokens.
The lecture’s figure, from the SGLang paper, uses the pattern {"summary": ". A normal FSM for it has 14 states in a row, 0 to 13, one per character: {, ", s, u, m and so on. Decoding with it calls the model at every step, even where only one token is legal: starting from the fixed {", three calls in a row are each asked for a token the format has already decided, producing summary, then ":, then the space before the value’s opening quote; only the call right after that has more than one legal token to choose from. A compressed FSM merges that chain into a single transition, 0 to 1, labelled with the whole string; decoding appends all four tokens at once and calls the model only where it has a real choice.
Try it on the contact example, with a format that says: an object with a name string, an email that is a string or null, and an intent that is exactly one of meeting, intro or follow-up.
The box is the output so far: shaded pink tokens were forced by the format, bold warm ones were chosen by the model. The chart shows the current choice, with forbidden tokens struck out. Pick a mode, then press Step or Run all.
Illustrative: the token pieces and the model’s candidate probabilities are made up for this example; decoding is greedy (always the most likely legal token). The format, the field types and the three allowed intents follow the lecture’s DSPy example below; the idea of skipping forced tokens with a compressed FSM is the SGLang figure’s.
Count the calls. Free decoding spends one call per token and still produces text the parser rejects, twice over: the preamble, and an intent of “Meeting” with a capital M that is not one of the three allowed values. The normal FSM always produces valid output, but spends 24 calls for 24 tokens, 15 of which had exactly one legal choice. The compressed FSM produces the same valid output with 9 calls. Since every decode call re-reads the model’s weights (Chapter 7), skipping forced tokens is a direct saving.
The lecture’s next slide shows what this buys in a real system. SGLang’s chart measures throughput (work completed per second, scaled so that SGLang is 1.0) on Llama-7B models across eleven workloads, from MMLU and HellaSwag to ReAct agents, generative agents, tree-of-thought, an LLM judge, JSON decoding, multi-turn chat and a DSPy retrieval pipeline. SGLang leads on every one, often by several times against the other systems shown (vLLM, Guidance and LMQL). The slide’s caption is careful: SGLang has finite state machines for decoding, among other optimizations, so not all of that lead comes from constrained decoding.
Here is one constrained step, from scratch. It reuses sample_next from Chapter 2.
pythonimport numpy as np def constrained_generate(model, fsm, ids, max_new=200): """fsm.allowed(state) -> legal token ids; fsm.step(state, tok) -> next state.""" state, calls = fsm.start, 0 while not fsm.done(state) and max_new > 0: legal = fsm.allowed(state) if len(legal) == 1: # forced by the format: tok = legal[0] # append it, no model call (a compressed FSM) else: logits = model(ids)[-1]; calls += 1 # one decode step mask = np.full(logits.shape, -np.inf) mask[legal] = 0.0 # illegal tokens get a score of minus infinity, tok = sample_next(logits + mask, temperature=0) # so their probability is exactly 0 ids = ids + [tok] state = fsm.step(state, tok) max_new -= 1 return ids, calls
Two limits are worth knowing. The mask guarantees form, not content: a perfectly valid object can still hold the wrong email address. And the mask only removes options: when the model’s favourite token is illegal (“Meeting”), it must take a legal one it rated lower. When that happens often, the format is fighting the model, and a clearer prompt or a friendlier format usually helps more than a stricter mask.
The second way works at the level of the program. DSPy, a library the course covers in depth in its Frameworks and Orchestration lecture (7 October), lets you declare a task as a typed signature: what goes in, what comes out, and the type of each field. The lecture’s example, from DSPy’s own documentation:
pythonimport dspy from typing import Literal, Optional class Extract(dspy.Signature): """Extract contact info.""" message: str = dspy.InputField() name: str = dspy.OutputField() email: Optional[str] = dspy.OutputField() intent: Literal["meeting", "intro", "follow-up"] = dspy.OutputField() extract = dspy.Predict(Extract) extract(message="I'm Sarah" "(sarah@acme.co). Meet Thursday?") # two literals, joined into one string # Prediction(name="Sarah", email="sarah@acme.co", intent="meeting")
In the slide’s own words, signatures define tasks and enforce output types. Optional[str] means “a string, or nothing”; Literal[...] means “exactly one of these”. The result comes back as a Python object with typed fields, not as loose text.
How? The lecture’s next slide peeks inside, and the answer is refreshingly plain: DSPy writes a prompt. The system prompt lists the input fields (message, a string) and output fields (name; email, a Union[str, NoneType]; intent, a Literal['meeting', 'intro', 'follow-up']), then shows a template with a marker for each field: [[ ## message ## ]], [[ ## name ## ]], [[ ## email ## ]], [[ ## intent ## ]] and [[ ## completed ## ]]. The email line carries a note that the value must match a JSON schema allowing a string or null; the intent line, that it must exactly match, with no extra characters, one of meeting, intro or follow-up. It ends with the objective: “Extract contact info.”
The user prompt fills in [[ ## message ## ]] with “I’m Sarah(sarah@acme.co). Meet Thursday?” and asks for the output fields in order, name first, then email, then intent, ending with the completed marker. The model’s reply is ordinary text; DSPy splits it at the markers and converts each piece to its declared type. As data flow: in, a typed Python call; through, a generated text prompt and a text reply; out, a typed object, or an error if parsing fails. The contract lives in your code, and the prompt is generated from it.
The third way skips generation altogether when the answer is a choice. If your pipeline only needs to know whether a review is positive, do not ask a language model to write “positive”. The lecture’s slide: train a small head on top of embeddings to map to class labels. Its figure, from “A Visual Guide to Using BERT for the First Time”, feeds the review “a visually stunning rumination on love” (as tokens, framed by the special tokens [CLS] and [SEP]) through DistilBERT, a small language model. The vector it produces for [CLS] goes into a logistic regression, which outputs 15% negative and 85% positive: label 1, positive. One pass, and the output is always one of the allowed labels.
A fixed head has fixed labels, though. The lecture’s next slide describes what it calls system one models, a popular term right now for, roughly, classifiers with adaptable options, used for decisions: a quick, one-pass reflex, set beside the deliberate, multi-step thinking of Chapter 8’s reasoning models. Its example plays a Mario game in three steps. First, contrastive state-action pre-training: one encoder turns game screens into vectors, another turns action phrases such as “right jump” into vectors, and training makes matching screen and action pairs score high (the diagonal of a grid of scores Si·Aj) and mismatched pairs low. Second, create action candidates: write the options as text (“left”, “jump”, “right run”), place each in the template “Mario should {action}.” and encode them. Third, zero-shot action classification: encode the current screen, score it against each option, and take a softmax: left 5%, jump 80%, right run 15%. Mario should jump, and the agent acts in the game.
The scores are the lecture’s softmax over three written options. Switch options off and on: the model re-decides among whatever is on the menu, with no retraining.
The 5%, 80% and 15% are the slide’s softmax over the three options. Removing an option and re-running the softmax over the rest is exactly the same as rescaling the remaining probabilities to sum to 100%, which is what the device does.
For an agent builder this is a design choice worth remembering: when a step is really a decision among known options (route this ticket, pick this tool, flag this message), a classifier or a system-one model gives a valid answer in one pass, every time, often far more cheaply than generating text and parsing it.
The fourth example is the one agents live on. A tool call is a piece of structured output that the surrounding program recognises, runs and answers. The lecture shows Toolformer’s examples, where the calls are written inline in the text, in brackets, with an arrow to the result:
Each call has a name (QA, Calculator, MT for machine translation, WikiSearch) and an argument in a fixed syntax, which is exactly what makes it parseable: the program spots the bracket, runs the tool, splices the result after the arrow, and the model carries on writing with the result in view. Today’s model APIs formalise the same idea. As the appendix of this lecture’s reading puts it, tools are specified to the model with their exact structure and definition, and when the model plans to call one, its response includes a tool use block that your program executes (Chapter 11 returns to how to design those definitions). Every trick in this chapter applies: a schema for the arguments, constrained decoding to guarantee them, and a classifier-style choice among the tools on offer. (Lecture 1’s lesson walks a Toolformer call step by step.)
{"name": " without calling the model at all?Chapter 10
Decide what goes into an agent’s limited context at every step, keep it small and sharp, and survive tasks longer than the window
A new colleague joins your project on Monday. You could share the entire drive with them, all ten thousand files, and wish them luck. Or you could hand them a one-page brief, the three documents they need this week, and a map of where everything else lives, so they can fetch it when it matters. Everyone knows the second works better. Nobody reads ten thousand files, and the one instruction that mattered gets lost among them.
A model’s context is the brief. The reading behind this part of the lecture, Anthropic’s “Effective Context Engineering for AI Agents” (Rajasekaran and colleagues, 2025), defines it simply: context is the set of tokens included when sampling from a language model. The engineering problem is to get the most out of those tokens, within the model’s limits, so that it reliably does what you want. Its framing question: what configuration of context is most likely to generate our model’s desired behaviour?
The reading calls context engineering the natural progression of prompt engineering. Prompt engineering is about writing and organising instructions, especially system prompts, and it was most of the job when most uses were one-shot: classify this, write that. The lecture shows the reading’s figure. On the left, for a single-turn query, the context window holds a system prompt and a user message; the model returns an assistant message. On the right, for an agent, there is a large pool of possible context (documents, tools, memory files, comprehensive instructions, domain knowledge, message history). A step called curation picks what enters the window this time (a system prompt, two documents, a memory file, two tools, the user message, the message history). The model answers with a message or a tool call, and the tool’s result flows back into the pool for the next round. A pair of scissors on the window marks its hard edge.
That loop is the point. An agent generates more and more data that could matter for its next turn, so, in the reading’s words, context engineering is iterative: curation happens each time we decide what to pass to the model. The lecture lists what competes for the window in an agent: the system prompt, the message history, memories, tool call results and attached files.
Windows are large now, so why curate at all? Because more context is not free even when it fits. Studies built like needle-in-a-haystack tests found what the reading calls context rot: as the number of tokens in the window grows, the model’s ability to accurately recall information from it falls. Some models degrade more gently than others, but it shows up in all of them.
The lecture’s evidence comes from OOLONG, a benchmark of long-context reasoning and aggregation: tasks where the answer depends on many pieces spread through the context, not one needle. Its examples: given news articles with dates, “were there more articles about the economy in September or August?” (the model must find each economy article, note its month and count: two in September, one in August, so September); and given transcripts of twelve hours of dialogue, “how many times does the character Jester cast Healing Word?” (two times). On its chart the stronger models’ scores fall steadily as the context grows (the weakest sit at or below the random baseline throughout). On the synthetic version, the strongest models score roughly 0.85 to 0.9 at 8,000 tokens and roughly 0.4 to 0.5 by 256,000, and one falls below the random baseline; on the real transcripts, which start at 64,000 tokens, the best score is about 0.6 and falls to about 0.3 by 512,000.
The reading explains why, with ideas from Chapters 3 and 4. Attention lets every token relate to every other, which means n2 pairwise relationships for n tokens, and as n grows the model’s ability to capture them is stretched thin. Models also learn their attention habits from training data in which short sequences are more common than long ones. So context should be treated as a finite resource with diminishing returns: the model has an attention budget, and every token spends some of it. The reading’s guiding principle follows directly: find the smallest possible set of high-signal tokens that maximise the likelihood of the desired outcome.
The lecture is frank that there is no one right way to build the context; students experiment with it in the course’s first homework. What it offers instead are practical tips and considerations, and the reading supplies the rest. Take them part by part.
The reading asks for system prompts in clear, simple, direct language at the right altitude: a Goldilocks zone between two failures. Too low, and engineers hardcode brittle if-then logic to force exact behaviour, which is fragile and hard to maintain. Too high, and the prompt gives vague guidance that offers no concrete signals, or falsely assumes the model shares context it does not have. Here are three versions of a support agent’s system prompt, written by us to show the difference.
brittle rules
“If the message contains ‘refund’ and the order is over $50 and under 30 days old, reply with template R2; if it contains ‘refund’ and …” and on, for forty more branches.
vague
“You are a helpful support agent. Resolve the customer’s issue.” Nothing about which tools to use, what a good resolution is, or when to hand off.
clear heuristics
“<background> who our customers are </background> <instructions> goals, priorities, when to ask a human </instructions> ## Tool guidance ## Output description”
The third version follows the reading’s advice on structure: distinct sections such as <background_information>, <instructions>, ## Tool guidance and ## Output description, marked with XML tags or Markdown headers, though exact formatting matters less as models improve. Aim for the minimal set of information that fully outlines the expected behaviour; minimal does not mean short. Start with a minimal prompt on the best available model, then add instructions and examples to fix the failures you actually observe.
Tools are how an agent pulls new context in, so they define, in the reading’s words, the contract between the agent and its information and action space. Good tools return token-efficient results and encourage efficient behaviour; like good functions in a codebase, they are self-contained, robust to error and clear about their intended use, with descriptive, unambiguous parameters. One of the most common failures is a bloated tool set: tools whose functions overlap, leaving ambiguous decisions about which to use. The reading’s test is memorable: if a human engineer can’t definitively say which tool should be used in a given situation, an AI agent can’t be expected to do better. (The lecture’s required reading adds how to write each tool’s definition; see Chapter 11.)
The lecture’s slide, titled “Avoid bloat”, puts numbers on it: sometimes a few tools are all you need. A coding harness called Pi gives the model just four tools: read, write, edit and bash (a command line). The chart runs the same models inside three harnesses on Terminal-Bench 2, a benchmark of tasks done in a computer terminal.
Pick a model. Left: the cost of one attempt at a task. Right: the share of tasks solved. Pi is the four-tool harness.
Numbers from the lecture’s chart “Terminal-Bench 2: impact of harnesses on the same model” (credited to arena.ai’s post on the coding-agent “harness tax”). A rollout is one complete attempt at one task. The chart also draws error bars, which are not shown here.
With Claude Opus 4.8, the four-tool Pi ties Codex for the best success rate (72.2%) at the lowest cost ($0.76 per rollout). With Sonnet 4.6 and Haiku 4.5 it has the best success rate outright; with Haiku, 47.8% against Codex’s 31.1%. Only with the strongest model, Claude Fable 5, does another harness win on success (Claude Code, 75.6%), and it costs the most ($1.55). A bigger toolbox is not automatically a better one: every tool description sits in the context and spends attention budget, on every step.
Few-shot prompting, including examples of the task done well in the prompt, remains a best practice the reading strongly advises. The trap is to stuff in a laundry list of edge cases, trying to spell out every rule. Instead, curate a set of diverse, canonical examples that portray the expected behaviour. For a model, as the reading puts it, examples are the “pictures” worth a thousand words. Its overall advice for every component (system prompts, tools, examples, message history) is the same: informative, yet tight.
Where should the rest of the knowledge come from? The lecture’s tip is to use retrieval to recover the most relevant memories or files instead of carrying them all: a retrieval chain passes the prompt through an embedding model, a vector-database search and a reranking model, and hands the best results to the model along with the prompt. That is the next lecture’s whole subject.
The reading adds a shift in how agents do it. Many applications retrieve with embeddings before inference, up front. Agentic systems increasingly add just-in-time context: instead of loading data in advance, the agent keeps lightweight identifiers (file paths, stored queries, web links) and loads data through tools only when it needs it. Claude Code works this way over large databases: it writes targeted queries, stores results, and uses commands like head and tail to inspect large data without ever loading the full objects into its context. People do the same, with file systems, inboxes and bookmarks rather than memorising everything.
The identifiers carry signal of their own. A file named test_utils.py in a tests folder implies a different purpose from the same name in src/core_logic/; folder hierarchies, naming conventions and timestamps all hint at what matters. Letting the agent explore enables progressive disclosure: each step yields context that informs the next (file sizes suggest complexity, names hint at purpose, timestamps can proxy for relevance), so the agent builds understanding layer by layer and keeps only what is needed in its working memory.
There is a trade-off. Exploring at run time is slower than retrieving pre-computed data, and without good tools and heuristics an agent wastes context by misusing tools, chasing dead ends or missing key information. So the best agents may use a hybrid: some data up front for speed, the rest explored at the agent’s discretion. Claude Code drops CLAUDE.md files into context up front, while tools like glob and grep let it find files just in time, avoiding stale indexes and complex syntax trees. The reading suggests hybrids suit less dynamic content, such as legal or finance work, and gives its standing advice: do the simplest thing that works.
The lecture’s next tip comes straight from Chapter 7: add sequentially, to avoid invalidating the KV cache. The cache holds the keys and values already computed for the context so far. Append new text at the end, and every cached entry stays valid, so only the new tokens need a prefill. Insert even one word at the front (the slide adds “Actually” before “The class was”) and, in the slide’s words, all of these position embeddings just changed: every later token now sits at a new position, and in deeper layers it has a new text before it, so every cached key and value is stale and the whole context must be prefilled again.
An agent’s context after six steps. Purple underline: tokens whose keys and values are still in the cache. Warm: tokens that must be prefilled again. Try each change.
Illustrative sizes, as in Chapter 0: a 1,200-token system prompt and task, and 1,550 tokens per step (a 150-token thought and tool call, a 1,400-token tool result). The rule is the lecture’s: everything after the first changed token must be recomputed.
Notice the tension the device exposes. Clearing old tool results makes the context smaller, which is good for the attention budget, but it rewrites the middle of the context, so everything after the first cleared result must be prefilled again. Appending is the cheapest edit; rewriting the beginning (a timestamp in the system prompt, a reordered tool list) is the most expensive, and it happens silently on every step if your harness does it.
How do you know a context choice helped? Measure it. The lecture’s slide: you can set up benchmarks or evaluations and test your changes. Its example is Toolformer’s own ablation: the same models run with their tools enabled and with them disabled, against GPT-3, across model sizes, on three benchmark families. On maths benchmarks, the largest model scores roughly 26.5 with its tools and roughly 11 without, above GPT-3’s roughly 14.5. On question answering the tools help too, but GPT-3 (roughly 39) stays ahead of the largest Toolformer (roughly 31). A change that helps one family of tasks may not help another, and you only find out by measuring.
Some tasks, the reading notes, span tens of minutes to hours of continuous work (a large codebase migration, a research project), and their token count exceeds any window. Waiting for bigger windows will not fix it, because windows of every size suffer context pollution and relevance problems. The reading offers three techniques.
Compaction takes a conversation nearing the window’s limit, summarises it, and starts a fresh window with the summary. It is usually the first lever. In Claude Code, the message history is passed to the model to summarise and compress: it keeps architectural decisions, unresolved bugs and implementation details, discards redundant tool outputs, and continues with that summary plus the five most recently accessed files. The art lies in choosing what to keep, because overly aggressive compaction loses subtle context whose importance only shows later. The reading’s method for tuning a compaction prompt: on complex agent traces, first maximise recall (capture everything relevant), then improve precision (cut what is superfluous). Its lightest-touch form is tool result clearing: once a tool was called deep in the history, the agent rarely needs the raw result again.
The lecture sharpens the dilemma into four points. By the time you compact, the context has grown extremely long. The new decision is what to keep. Keep too much and you will have to compact again soon. Drop too much and you can lose critical context needed to complete the task. Its example is a report from February 2026: a user told an agent, OpenClaw, to suggest emails for deletion but not to take any action until she said so. The inbox was too large, the system triggered context compaction, the “confirm before acting” instruction was lost, and the agent began bulk-deleting emails on its own; she had to rush to her computer to stop it. Try the dilemma yourself.
An email-cleaning agent has filled 90% of its window. Choose what compaction keeps. The dashed line is where the next compaction triggers.
Illustrative: a 100,000-token window, compaction at 90,000, about 2,000 new tokens per step, and made-up block sizes. The failure it shows (a user instruction lost in compaction) is the one in the report the lecture shows.
Structured note-taking, or agentic memory, is the second technique: the agent regularly writes notes to a store outside its window and pulls them back in later. It is persistent memory with little overhead, like Claude Code keeping a to-do list, or a custom agent maintaining a NOTES.md file. The reading’s example is Claude playing Pokémon, which kept precise tallies across thousands of game steps (“for the last 1,234 steps I’ve been training my Pokémon in Route 1, Pikachu has gained 8 levels toward the target of 10”), built maps of explored regions, remembered achievements and kept notes on which attacks work against which opponents, all without being prompted about memory structure. After a context reset, it read its own notes and carried on for hours. Anthropic released a file-based memory tool in public beta alongside Sonnet 4.5 for exactly this pattern.
Sub-agent architectures are the third. Instead of one agent holding the state of a whole project, specialised sub-agents take focused tasks, each with a clean window, while a main agent coordinates with a high-level plan. A sub-agent may explore extensively, using tens of thousands of tokens or more, but returns only a condensed summary, often 1,000 to 2,000 tokens. Work it with illustrative numbers: three sub-agents that each read 30,000 tokens and each return 1,500 add 3 × 1,500 = 4,500 tokens to the lead agent’s context instead of 3 × 30,000 = 90,000, twenty times less. The detailed search stays isolated in the sub-agents; the lead focuses on synthesis. Anthropic’s multi-agent research system, built this way, showed a substantial improvement over single-agent systems on complex research tasks. It is the orchestrator-workers pattern of the lecture’s other reading (Chapter 11), used to keep each context clean.
Which to use depends on the task. Compaction keeps the flow of tasks that need extensive back-and-forth. Note-taking suits iterative development with clear milestones. Multi-agent architectures fit complex research and analysis where parallel exploration pays off.
Every technique so far still has the model read every token it keeps, however small compaction, notes or sub-agents manage to make that pile. The lecture ends the section with an idea that drops even that: what if the model did not have to read the long text at all? In a recursive language model (RLM), the model never reads the long prompt at all. The prompt is loaded as a variable inside a Python environment, and the model writes code to work with it. In the lecture’s figure, the root model (depth 0) is asked to list all items made before the “Great Catastrophe” in an extremely long book. It runs print(prompt[:100]) to peek at the start, splits the text at “Chapter 2”, and calls itself on each part with a focused question, through a function llm_query: sub-models at depth 1 answer “In Chapter 1, find all items that are listed to belong to people with…” and “Look for the items in these sections that explicitly say they were found…”, and return short sub-responses (“The silver flask…”, “Herod’s ring…”). The root combines them into its final answer: in chapter 1, Bob was the only character noted to be alive and carrying a flask before the Great Catastrophe; in chapter 3, Alice uncovers a ring belonging to the fallen King Herod.
| System, on OOLONG (trec_coarse) with a 263k-token context | Score |
|---|---|
| GPT-5 | 34.2% |
| GPT-5-mini | 23.2% |
| RLM (GPT-5-mini) | 51.1% |
| RLM (GPT-5), without sub-calls | 40.9% |
| ReAct + GPT-5 + BM25 retrieval | 30.7% |
The smaller model, used recursively, beats the larger model reading everything directly: 51.1% against 34.2%. Even without sub-calls, letting GPT-5 treat the context as data in an environment raises it to 40.9%. The context is no longer something the model must hold in its attention; it is something the model can inspect, slice and delegate. (For agents that also improve their own context over time, see agentic context engineering.)
The reading’s conclusion ties the chapter together. As models improve, they need less prescriptive engineering and can work with more autonomy; but even as capabilities scale, treating context as a precious, finite resource remains central to building reliable, effective agents.
Chapter 11
Learn the lecture’s required reading: start from one augmented call, reach for five workflow patterns before an autonomous agent, and design the interface the model works through
Two teams build the same customer-support assistant. The first installs a large framework full of planners, memory modules and agent classes; three weeks in, they are still trying to find out what prompt the model actually received. The second calls the model’s API directly: one call with retrieved help articles, a small router in front of it, two tools. The second team ships first, and its assistant is easier to fix when it goes wrong.
That story is the thesis of the lecture’s required reading, “Building Effective Agents” by Erik Schluntz and Barry Zhang of Anthropic, written in December 2024. (The live post now opens with a note that much of the tooling landscape it describes has changed since then; its patterns are what the lecture assigns.) Working with dozens of teams building agents across industries, the authors found that the most successful implementations were not using complex frameworks or specialised libraries. They were building with simple, composable patterns. This chapter walks the reading in its own order; every chapter of this lesson so far has been about the parts those patterns are built from.
The reading calls every system that uses a language model to do multi-step work an agentic system, and draws one architectural line through them, the same line as Lecture 1’s “who decides the steps?”. Workflows are systems where models and tools are orchestrated through predefined code paths. Agents are systems where the models dynamically direct their own processes and tool use, keeping control over how they accomplish the task.
The reading’s first advice is restraint: find the simplest solution possible, and only increase complexity when needed. That might mean not building an agentic system at all. Agentic systems usually trade latency and cost for better task performance, and you should consider when that trade makes sense. When more complexity is warranted, workflows offer predictability and consistency for well-defined tasks, while agents are the better option when flexibility and model-driven decision-making are needed at scale. For many applications, though, optimising a single model call with retrieval and in-context examples is enough.
Frameworks make agentic systems easier to start. The live post names the Claude Agent SDK, AWS’s Strands Agents SDK, Rivet (a drag-and-drop GUI workflow builder) and Vellum (a GUI tool for building and testing workflows). They simplify the standard chores: calling models, defining and parsing tools, chaining calls. But they add layers of abstraction that can hide the underlying prompts and responses, which makes systems harder to debug, and they make it tempting to add complexity a simpler setup would not need. The advice: start with the model APIs directly, since many patterns take a few lines of code, and if you do use a framework, understand the code underneath. Wrong assumptions about what is under the hood are, the authors say, a common source of error.
Every pattern is built from one unit: a language model augmented with retrieval, tools and memory. Current models can use those augmentations actively: they write their own search queries, choose tools and decide what information to keep. The reading asks you to focus on two things: tailor the augmentations to your use case, and give the model an easy, well-documented interface to them. One way to connect them is the Model Context Protocol (Lecture 1). Everything in Chapters 9 and 10 (structured tool calls, a small clear tool set, retrieval just in time) is about building this unit well.
From that unit the reading builds upward, one pattern at a time. Play each one below: the dots show the order in which work flows.
Pick a pattern and press Run. Warm boxes are model calls; pink is code the engineer wrote; blue is the outside world. Dots show the order of work.
Diagrams redrawn from the reading’s figures (Schluntz and Zhang, “Building Effective Agents”, Anthropic, December 2024). The number of calls and loops in each run is illustrative.
Prompt chaining breaks a task into a fixed sequence of steps, each model call processing the output of the previous one. You can add programmatic checks, which the reading calls gates, on any intermediate step to make sure the process is still on track. Use it when the task decomposes cleanly into fixed subtasks; the goal is to trade latency for accuracy by making each call an easier task. Examples: write marketing copy, then translate it into another language; or write an outline of a document, check that the outline meets certain criteria, then write the document from the outline.
Routing classifies an input and sends it to a specialised follow-up task. This gives separation of concerns and lets you write more specialised prompts; without it, optimising for one kind of input can hurt performance on others. Use it when there are distinct categories better handled separately, and the classification can be done accurately, either by a model or by a more traditional classifier (Chapter 9’s classifiers fit exactly here). Examples: send general questions, refund requests and technical-support requests to different processes, prompts and tools; or send easy, common questions to a smaller, cheaper model like Claude Haiku 4.5 and hard, unusual ones to a more capable model like Claude Sonnet 4.5.
In parallelization, models work on a task at the same time and a program aggregates their outputs. It has two variants. Sectioning splits a task into independent subtasks run in parallel. Voting runs the same task several times to get diverse outputs, the idea behind Chapter 8’s self-consistency. Use it when subtasks can run in parallel for speed, or when several perspectives or attempts give more confident results; for complex tasks with several considerations, models generally do better when each consideration gets its own call and its own focused attention.
Sectioning examples: guardrails, where one model instance handles the user’s query while another screens it for inappropriate content, which tends to work better than asking one call to do both; and automated evaluations, where each call judges a different aspect of a model’s output. Voting examples: several different prompts each review a piece of code for vulnerabilities and flag problems; or several prompts judge whether content is inappropriate, with different vote thresholds to balance false positives against false negatives.
In orchestrator-workers, a central model dynamically breaks a task down, delegates the pieces to worker models, and synthesises their results. It suits complex tasks whose subtasks you cannot predict. In coding, for example, the number of files that need changing, and the nature of each change, depend on the task. It looks like parallelization, but the key difference is flexibility: the subtasks are not predefined; the orchestrator decides them for each input. Examples: coding products that make complex changes to several files each time, and search tasks that gather and analyse information from many sources. Chapter 10’s sub-agent architectures are this pattern, used to keep each context clean.
In evaluator-optimizer, one model call generates a response while another evaluates it and gives feedback, in a loop. It works best when there are clear evaluation criteria and iterative refinement adds measurable value. The reading gives two signs of a good fit: a model’s responses improve noticeably when a person articulates feedback, and a model can provide such feedback itself. It is like the drafts a human writer goes through. Examples: literary translation, where an evaluator can catch nuances the translator missed; and complex searches needing several rounds, where the evaluator decides whether more searching is warranted.
Agents are appearing in production, the reading says, as models mature in the key capabilities: understanding complex inputs, reasoning and planning, using tools reliably, and recovering from errors. An agent starts with a command from, or a discussion with, a person. Once the task is clear it plans and works independently, possibly returning to the person for information or judgement. At each step it must get ground truth from the environment, such as tool results or code execution, to assess its progress. It can pause for human feedback at checkpoints or when blocked. It stops when the task is done, and it is also common to add stopping conditions, such as a maximum number of iterations, to keep control.
For all that, an agent’s implementation is often simple: typically just a model using tools based on environmental feedback, in a loop. (The context-engineering reading of Chapter 10, written a year later, compresses it further: LLMs autonomously using tools in a loop, with autonomy that can grow as models get better at navigating problems and recovering from errors.) That is exactly why the reading says it is crucial to design the toolset and its documentation clearly and thoughtfully. Use agents for open-ended problems where the number of steps is hard or impossible to predict and a fixed path cannot be hardcoded. The model may run for many turns, so you must have some trust in its decisions; that autonomy makes agents ideal for scaling tasks in trusted environments. It also means higher costs and the potential for compounding errors, so the reading recommends extensive testing in sandboxed environments, with appropriate guardrails. Its own examples: a coding agent that resolves SWE-bench tasks, which involve edits to many files from a task description, and a “computer use” reference implementation in which Claude operates a computer.
These patterns are not prescriptions, the reading stresses; they are common shapes to adapt and combine. The key to success is measuring performance and iterating, and adding complexity only when it demonstrably improves outcomes. Sort some real cases yourself.
Read the task, then choose the simplest pattern the reading would reach for. You get the reason either way.
Every task is one of the reading’s own examples, reworded; the last is its advice that a single call with retrieval and examples is often enough.
The reading’s first appendix names two applications where agents add the most value, and what they share: tasks that need both conversation and action, have clear success criteria, allow feedback loops, and include meaningful human oversight.
Customer support combines a familiar chat interface with tools. Conversations flow naturally while needing outside information and actions; tools can pull customer data, order history and knowledge-base articles; actions such as issuing refunds or updating tickets can be handled programmatically; and success can be measured by user-defined resolutions. Some companies are confident enough to charge only for successful resolutions.
Coding agents fit because code can be verified by automated tests, agents can iterate using test results as feedback, the problem space is well defined, and output quality can be measured objectively (the verifiable rewards of Chapter 6, used at run time). Anthropic’s own agents can solve real GitHub issues in SWE-bench Verified from the pull-request description alone, though human review remains crucial to make sure solutions fit the wider system.
The second appendix is about tools, and it is the part of the reading an agent builder returns to most. Tools let the model act on outside services; their exact structure and definition are given to the model, and when it plans to use one, its response includes a tool use block (Chapter 9). Tool definitions deserve as much prompt-engineering attention as the prompt itself.
There are often several ways to express the same action: a file edit as a diff or as a full rewrite of the file; structured output as code inside Markdown or inside JSON. To a programmer these are cosmetic. To a model some are much harder. Writing a diff requires knowing how many lines change in the chunk header before the new code is written. Writing code inside JSON requires escaping every newline and quote. The reading’s suggestions:
Its rule of thumb: put as much effort into the agent-computer interface (ACI) as people put into human-computer interfaces. Put yourself in the model’s shoes: is it obvious how to use the tool from its description and parameters? A good definition includes example usage, edge cases, input format requirements and clear boundaries from other tools. Rename parameters and rewrite descriptions until they are obvious, as if writing a great docstring for a junior developer, especially when many tools are similar. Test how the model uses the tools, run many examples, and iterate. And poka-yoke your tools (a Japanese term for mistake-proofing): change the arguments so that mistakes are harder to make.
relative paths
edit(path, ...) accepts src/app.py. After the agent moves out of the root directory, the same relative path points somewhere else, and edits land in the wrong place.
absolute paths only
edit(path, ...) requires a full path from the root, such as /repo/src/app.py. The mistake is now impossible, and the model used the tool flawlessly.
The function and file names above are our sketch; the example itself is the reading’s own: while building their agent for SWE-bench, the authors spent more time optimising the tools than the overall prompt. The model made mistakes with relative file paths once it had moved out of the root directory; changing the tool to always require absolute paths fixed it.
The reading closes with its summary. Success is not about building the most sophisticated system; it is about building the right system for your needs. Start with simple prompts, optimise them with comprehensive evaluation, and add multi-step agentic systems only when simpler solutions fall short. When building agents, follow three core principles:
Frameworks can help you start quickly, but do not hesitate to reduce layers of abstraction and build with basic components as you move to production. Read against this lecture, every principle has a reason inside the model: simplicity keeps the context small and the attention budget unspent (Chapter 10); transparency works because the reasoning a model writes is extra computation you can read (Chapter 8); and the agent-computer interface is structured output by another name (Chapter 9).
Chapter 12
Review the model from the builder’s side, the numbers and the code, then look ahead to retrieval
You can now reproduce the second lecture of CS 329Z from memory: what a language model computes, how it is wired, how it learns, how it runs, and how an agent builder shapes what goes in and what comes out. The lecture’s own summary slide ticks off five boxes: what language models are; their architecture (transformers and linear attention); their training (pretraining, midtraining, post-training and agent training); their inference (prefill and decode, and inference-time scaling); and LLMs for agents (structured I/O and context engineering). Here is all of it in one place.
A language model gives every possible next token a probability, and by the chain rule a probability to every text; it is generative, so text is made by drawing one token at a time, with a decoding strategy (greedy, temperature, top-k, top-p, beam search) chosen by you. Inside, a decoder-only transformer turns tokens into vectors and lets each one attend, through queries, keys and values under a causal mask, to the tokens before it; this parallelises training but costs O(n2) in context length, so linear attention folds the past into a fixed-size state (cheaper, but blurrier over long contexts) and frontier models mix the two. The probabilities come from minimising cross-entropy on filtered web data (pretraining), rebalanced toward target domains (data mixing, midtraining), then shaped by post-training: instruction finetuning, human preferences (with their unreliability and sycophancy), verifiable rewards (only as good as the verifier) and agent-specific data such as SWE-smith. Serving splits into a compute-bound prefill and a bandwidth-bound decode that leans on the KV cache; speculative decoding lets a draft model guess and the target verify. More compute at answer time (chain of thought, reasoning levels, repeated sampling with a verifier, or voting without one) buys better answers. And an agent builder makes the model’s output machine-readable (constrained decoding, typed signatures, classifiers, tool calls) and treats its context as a finite attention budget: the smallest set of high-signal tokens, appended rather than rewritten, compacted, noted or delegated when tasks outgrow the window, inside the simplest workflow or agent pattern that does the job.
| What | Number | Why it matters |
|---|---|---|
| Next word after “The class was about” | “to” 0.3171; “how” 0.0124 | A model outputs a distribution, never an answer |
| One sentence’s probability | 1.2 × 10−17 | Why code adds log-probabilities |
| Temperature, top three words | “to” 78.84% (T = 1), 96.44% (T = 0.5), 57.79% (T = 2) | Temperature scales the gaps between scores |
| Loss on “how” | 1.907 (base 10) = 4.390 (natural log) | The log base only rescales the loss |
| Transformer (base) vs ConvS2S ensemble | 27.3 vs 26.36 BLEU; 3.3 × 1018 vs 7.7 × 1019 FLOPs | Better and cheaper to train |
| Hybrid attention, retrieval score | DeltaNet 22.7; + sliding window 30.2; + global 32.7 | A few full-attention layers restore exact recall |
| DCLM filtering | 1.4% of documents survive | Pretraining data is heavily curated |
| SWE-bench Verified, same harness | SWE-agent-LM-32B 40.2% vs GPT-4o 23.0% | Agent-specific training pays |
| Roofline (illustrative chip) | decode ≈ 1 op/byte (0.3% of peak); prefill of 2,048 ≈ 1,024 op/byte | Decode waits on memory, prefill on arithmetic |
| Speculative decoding | 1 target pass, 4 tokens; about 2.5× faster on code | More tokens per trip to memory |
| Repeated sampling, p = 0.27 | 46.7% at k = 2; 79.3% at k = 5 (independent tries) | Coverage grows with samples, if you can verify |
| Self-consistency on ARC Challenge | about 43% greedy; about 54% with 40 voting paths | Voting when there is no verifier |
| Terminal-Bench 2, Claude Haiku 4.5 | Pi (4 tools) 47.8% vs Codex 31.1% | Avoid tool bloat |
| OOLONG, 263k tokens | RLM (GPT-5-mini) 51.1% vs GPT-5 34.2% | Treat long context as data, not as input |
| Sub-agent summaries | often 1,000 to 2,000 tokens | Keep the lead agent’s context clean |
Every part of the lecture shows up in the few lines that run a single step of an agent. This is a sketch: llm, count_tokens, run_tool and clip stand for whatever model API and helpers you use.
pythonSYSTEM = "Reasoning: medium\n<instructions>...</instructions>\n## Tool guidance ..." # Ch 8, Ch 10 TOOLS = [read_schema, write_schema, edit_schema, bash_schema] # Ch 10: a few clear tools def agent_step(history, window=100_000, keep_recent=6): context = [SYSTEM, TOOLS] + history # stable front, appended back: KV cache stays valid (Ch 7, 10) if count_tokens(context) > 0.9 * window: # near the limit: compact (Ch 10) summary = llm([COMPACT_PROMPT] + history, temperature=0) # keep rules, decisions, open bugs history[:] = [summary] + history[-keep_recent:] context = [SYSTEM, TOOLS] + history reply = llm(context, temperature=0, schema=ACTION) # low temperature (Ch 2), constrained output (Ch 9) history.append(reply) # each token was one decode step (Ch 3, 7) if reply["action"] == "finish": return reply["answer"] result = run_tool(reply["action"], reply["args"]) # the program runs the tool, in a sandbox (Ch 11) history.append(clip(result, 2000)) # token-efficient results (Ch 10) return None # the caller loops, with a step budget (Ch 11)
Read it line by line and name the chapter behind each choice. The system prompt sets a reasoning level and is written at the right altitude. The tool list is short. The context is assembled with a stable front so the KV cache survives, and compacted, rules first, when it nears the window. The reply is drawn at low temperature and constrained to a schema, so the program can always parse it. The tool result is clipped before it joins the context. And the loop around this function, with its step budget, is the simplest agent pattern of the reading.
Lecture 1 drew the agent: a language model core with planning, memory, tools and an environment, and the question of who decides the steps. This lecture opened the core. What it showed is that the model’s randomness, speed, memory and habits are all consequences of design decisions (sampling, attention, training data, post-training, serving hardware), and that each one is a lever or a limit for the system you build around it.
The course meets Mondays and Wednesdays at Stanford, and posts its slides on the course site. Next is Lecture 3, Retrieval-Augmented Generation (Wednesday 30 September): grounding and hallucination, embeddings and vector stores, chunking strategies, hybrid search, cross-encoders and late interaction (ColBERT), and building a retrieval pipeline from scratch, with Lewis and colleagues’ 2020 paper as the reading. It picks up exactly where Chapter 10 left off: retrieval as the way to recover the most relevant memories and files without bloating the context.
Now press Present or Teach and explain one agent step out loud, from memory: how the next token is drawn, why the context costs more with every step, where the probabilities came from, why decode waits on memory, and what you would change in the context to make the step cheaper and safer. Then go back to the sentence at the top, set the temperature to 0.3 and to 1.8, and explain why the same model continues it so differently.