A vision-language model that stores 32 layers but runs 128, by looping the same layers. Looping lifts its MMStar score from 55.3 to 63.5 with no extra layers. Halfway through, its attention snaps onto the part of the picture that holds the answer.
Press Play and watch where the model looks while the loop runs. Then we build, piece by piece, why the same weights look somewhere new the second time.
You need what a neural network layer is and the idea that a transformer reads a list of tokens. We build the rest from zero.
The scenes are toys after Figure 8 (b, c, h). The heatmaps are illustrative, built so that their concentration follows the Gini curve of Figure 7, traced by eye (the jump from 0.405 to 0.858 between calls 63 and 64 is printed in the paper). The answer bars follow the shape of Figure 11: the answer forms in the last H call. The 55.3 and 63.5 are Table 1.
Chapter 0
Why a model that reads a picture in one pass can run out of steps, and what a loop offers instead
Show a model a picture of a cluttered table and ask: how many cyan squares are there? A person does not answer in one glance. You scan the picture, find the shapes, then look again at the cyan ones and count them.
A vision-language model (VLM) answers questions about pictures. It cuts the picture into small patches and turns each patch into a list of numbers called a token. It puts those image tokens in a row with the tokens of the question. Then a stack of layers processes the row. Each layer updates every token a little, and after the last layer the model writes its answer.
Inside each layer, attention lets a token read from the other tokens. The token gives every other token a weight, the weights add up to 1, and it takes in a weighted mix of them. When the answer token puts large weights on a few image tokens, we say the model is looking at those patches.
In an ordinary model, every layer runs once. A 32-layer model gets exactly 32 updates to find the cyan squares, count them and write the answer. If a question needs more steps, the usual fix is a deeper model. But every extra layer has its own weights, so depth costs memory and training data.
A looped transformer takes another route. It stores a few layers and runs them again and again, each time on its own last output. Depth now costs time instead of weights. The idea is not new: the Universal Transformer did it in 2019. It has come back for language models that need more steps of thought.
Two ways to get the same number of layer calls. Each coloured block is one stored layer. Drag the slider and compare what each model must store.
The weight counts are our arithmetic, not the paper's: about 12 × 1,536² numbers per layer (attention plus a SwiGLU feed-forward), at LoopVL's width of 1,536. Embeddings and the vision encoder are left out. The real LoopVL loops two different 16-layer stacks; the drawing shows the stored layers as one set of 32.
There is a catch, and it is the reason this paper exists. In a loop, the weights never change. If the input to the loop never changed either, every pass would compute the same thing, and the loop would be useless. A loop only helps because the state it works on changes from pass to pass.
For pictures, that raises a new question. After one pass, the model has read the question and the image together. Perhaps it now knows that the cyan squares matter and the big pink square does not. On the next pass, can the same layers look at different patches?
LoopVL (Qian, Gong, Xu and colleagues, 2026) tests exactly that. It is a VLM at the 1B scale whose language part is a loop. The loop has two stacks of 16 layers, run in a fixed order for 128 layer calls. The authors pretrain that loop from scratch and train the whole model on 0.14 trillion tokens. Then they measure where the model looks on each pass.
Two findings carry the paper. First, the loop works. With the same 32 stored layers, looping lifts MMStar from 55.3 to 63.5 (Table 1). MMStar is a benchmark: a fixed set of picture questions, scored as the percentage answered correctly. The paper uses about 16 of them, and every score in this lesson is a percentage of that kind. The 1B model also beats several 2B to 4B models trained on 79 to 260 times more tokens.
Second, the loop looks again. At the boundary between the first and the second cycle, the attention on the image jumps from spread out to concentrated. The authors call this the Visual Aha Moment.
Here is the road, in order. Each chapter explains one part of the instrument in the hero.
Chapter 1
How a function with fixed weights can do new work on every pass
Start with the smallest possible loop. You have one block of layers, a function with weights θ. It takes a state and the input, and it returns a new state. Run it, then feed its output back in, and run it again.
Let's run one with numbers, in one dimension. Take the block f(h, x) = 0.5 h + 0.5 x, the input x = 2, and the start h = 0. Each pass moves the state halfway to 2.
0 → 1 → 1.5 → 1.75 → 1.875 (the same rule each time; the steps are 1, 0.5, 0.25, 0.125)
The rule never changed, but every pass did something different: it moved the state by a different amount. That is the whole trick. A shared block is a rule, and a rule applied to a new state gives a new result. The device shows the same idea in two dimensions.
The faint arrows are one fixed rule: from any state, the arrow shows where one pass moves it. Add passes and watch the state travel. Then change the question and run the same rule again.
A toy rule in two dimensions, not the paper's model: each pass moves the state 35% of the way toward a point set by the question, plus a small turn. A real block works on thousands of numbers per token, but the logic is the same.
Notice two things. First, the steps shrink as the state nears the point the rule leads to. The same weights do a big job on pass 1 and a small one on pass 8. Second, a new question moves the point, so the same rule sends the same start somewhere else.
Why would a model want this? Some problems are naturally step by step: counting, following a path, checking a guess and fixing it. A loop gives the model more steps for them without more weights. On such tasks, earlier work found that a shallow shared block, looped, comes close to a much deeper model (Saunshi and colleagues, 2025).
The cost moves elsewhere. A loop of 16 layers run 8 times does as much arithmetic as 128 separate layers. It just stores fewer of them. So a loop trades memory for time: it is small to hold but not cheaper to run. Chapter 4 counts this cost for LoopVL.
One more property matters later. A loop can run more passes at test time than it saw in training, or fewer. Does that help?
For some loops, running longer keeps improving the answer. For others, extra passes push the state somewhere it never went in training, and the answer gets worse. Researchers call this overthinking. Chapter 5 shows which kind LoopVL is.
Chapter 2
LoopVL loops two stacks at two speeds, and the order is fixed
LoopVL does not loop one block. It loops two, and it runs them at two speeds. The design comes from the Hierarchical Reasoning Model (HRM) and its language version, HRM-Text. Their idea comes from the brain: fast local work happens under a slower, high-level goal that changes less often.
A picture question needs both kinds of work. Some of it is fine detail: a colour, a letter, an edge, which shape sits left of which. Some of it is the big picture: what the question asks, and what the scene is. LoopVL gives each kind its own stack of 16 layers.
One call means one run of a whole 16-layer stack. One cycle is three L calls then one H call. LoopVL runs two cycles. Written out, the order is the paper's Equation 1:
The paper names this schedule H2L3: two H calls (two cycles), three L calls per cycle. Repeating L inside a cycle is the Module-Loop. Repeating the whole cycle is the Model-Loop. Count the layer calls: 2 cycles × 4 calls × 16 layers = 128. The model stores only 2 × 16 = 32 layers.
Each stack carries its own state. The H state starts as the input itself: the image and text tokens, scaled. The L state starts at zero and is kept from one cycle to the next. An L call reads the L state together with the current H state, the goal it works under. The H state stays fixed during the three L calls. Then the H call reads the refined L state and writes a new H state, the goal for the next cycle.
Two lanes, one per state. Step through the eight calls and watch which stack runs, what it reads and what it writes. The counter compares layer calls with stored layers.
The order, the state rules and the re-injection follow Section 2.1. How exactly two states are combined (HRM adds them) is not spelled out for LoopVL; the arrows show what each call reads, not the arithmetic.
Toggle Re-inject the picture and look at the start of each cycle. LoopVL adds back a fixed copy of the projected image tokens into the image positions of the H state. This copy is the visual anchor. A gate looks at the question and sets a strength for each image token, and a learned scale per cycle sets the overall strength. The anchor keeps the original picture within reach, however far the state has moved.
The final H state goes to the language-model head, the layer that turns a vector into a probability for every word in the vocabulary. The answer is read out there, and only there.
Inside each stack, every one of the 16 layers has the same shape. It is a standard transformer layer with two changes worth knowing. First, recall how attention works inside a layer. Each token makes three vectors: a query (what it looks for), a key (what it offers) and a value (what it passes on). Each query is compared with every key, the matches set the attention weights, and the token takes in the values mixed by those weights.
The gate lets each layer turn its attention output down, token by token. This helps in a loop. A layer that runs six times may have nothing to add on some passes, and the gate lets it say so.
Chapter 3
How image patches become tokens, where they sit, and who may read whom
The authors pretrained the loop on text. To read pictures it needs an entrance: a way to turn pixels into tokens that look, to the loop, like the word tokens it already knows. LoopVL builds that entrance in three parts.
The encoder stays frozen through all of training: its weights never change. Only the projector learns to map its features into the loop's space. One patch gives one token. LoopVL does not merge or drop image tokens, so the image keeps a stable grid. That grid is what lets the authors draw attention back onto the picture in Chapter 7.
The image size decides the number of tokens: between 256 and 1,024 per image. Each must fit in the sequence budget, 2,048 tokens in the first training stage and 3,072 in the second. A 1,024-token image plus a 60-token question leaves 964 tokens of room in the first stage.
The calibration step deserves a number. The size of a vector here is its RMS, the square root of the mean of its squared entries. Projected image features can be much larger or smaller than word embeddings. If an image token arrived four times louder than a word, attention would treat it very differently. So LoopVL first normalises each image token to RMS 1. Then it multiplies by the median RMS of the loop's word embeddings.
RMS 3.2 → ÷ 3.2 → RMS 1 → × 0.8 → RMS 0.8 (illustrative numbers: the image token now matches a typical word)
A transformer needs to know where each token is. LoopVL uses RoPE, rotary position embedding: it rotates parts of the query and key vectors by an angle that grows with the position. Two tokens then compare their positions through the difference of their angles.
Image tokens get two positions, a row and a column, so a patch knows its place on the grid: this is 2D RoPE. Text tokens also get two coordinates, but both advance together: (5, 5), (6, 6), (7, 7). With equal coordinates, the 2D rotation acts exactly like the ordinary 1D rotation the loop learned on text. So the text side of the pretrained loop sees nothing new.
The last piece is the attention mask: a table of which token may read which. A language model is usually causal: each token reads only the tokens before it. LoopVL uses a PrefixLM mask instead. The header, the image and the question form the prefix, and inside the prefix every token reads every other, in both directions. The answer tokens read the whole prefix and the answer tokens before them. No prefix token may read the answer.
A tiny sequence: two header tokens, a 2 × 3 image, three question tokens and three answer tokens. Each row is a reader, each column a token it may read. Tap a row, then switch the mask.
The mask rules and the position scheme follow Section 2.1 and Section 6.3. The token counts are shrunk to fit; a real image has 256 to 1,024 tokens.
Why does the mask matter for a loop? Under a causal mask, the first image patch could never read the question, because the question comes later. Under PrefixLM, every image token reads the question on every pass. So an image token can change because of what was asked. That is the precondition for the whole paper: the image state can only evolve toward the question if it can see the question.
Chapter 4
The same data and four model shapes, compared fairly
A loop makes a model deeper without more weights. But is a loop the best way to spend that depth? Perhaps you would do better to store more layers, or wider ones. The fair test holds everything else fixed: the same data, the same training budget, the same layer design. Only the shape changes.
The authors train four models on the same 0.14 trillion tokens. All use the same transformer layer as LoopVL. They differ in how many layers they store, how wide the layers are, and whether they loop. The last column counts training cost in FLOPs, floating-point operations: the total number of multiplications and additions the training run performed.
| Model | Stored layers | Width | Layer calls | Training FLOPs |
|---|---|---|---|---|
| LoopVL (H2L3) | 32 (16 L + 16 H) | 1,536 | 128 | 2.47 × 1021 |
| Transformer-VL 1B | 32 | 1,536 | 32 | 0.92 × 1021 |
| Transformer-VL 4B, deep | 78 | 1,792 | 78 | 2.89 × 1021 |
| Transformer-VL 4B, wide | 32 | 2,816 | 32 | 2.99 × 1021 |
Read the first two rows together. LoopVL and the 1B model store the same 32 layers at the same width. The only difference is that LoopVL runs them four times as often. The other two rows spend about three times the weights, one on depth and one on width.
Choose a benchmark. Each bar is one model, all trained on the same 0.14 trillion tokens. The labels under the bars show what each model stores and what it costs to train.
Scores and FLOPs: Table 1. The average is our mean of all eight benchmarks in Table 1 (also VMCBench, AI2D and VisuLogic). Weights in the layers are our estimate at 12 × width² per layer.
Against its twin, the loop wins every benchmark, by 6 to 23 points. MMStar goes from 55.3 to 63.5, RealWorldQA from 55.3 to 71.0, ChartQA from 51.1 to 74.5. Since the stored layers are identical, the gain comes from running them again.
Against the 4B models, LoopVL wins five of eight benchmarks and loses three (VMCBench, AI2D, ChartQA) by at most 1.1 points. On the eight-benchmark average it leads both. So a model with about a third of the stored weights matches or beats models three times its size, on the same data.
Now the honest part. Looping is not free. LoopVL used 2.7 times the training compute of its 1B twin (2.47 against 0.92, in units of 1021 FLOPs). It is the stored weights that stay small, not the arithmetic. Against the 4B models, LoopVL used 14 to 17% fewer FLOPs, so it is slightly cheaper there too, but in the same range.
The same holds when the model answers. Each answer token runs 128 layer calls, as many as a 128-layer model. The calls must happen in order, one after another, so a loop cannot run its passes in parallel. What the loop saves is memory: a phone or a small GPU holds 32 layers, not 128.
Chapter 5
The same weights, run in twelve different orders, and four models trained four ways
If more layer calls help, why stop at 128? A loop can, in principle, run any number of passes at test time. The paper asks two separate questions, and it matters to keep them apart.
The number of layer calls for any schedule is H × (L + 1) × 16. H3L3, for example, is 3 cycles of 4 calls: 192 layer calls. Look at H1L1 first. It is one L call and one H call, 32 layer calls in all: a plain model with no loop. In Table 3 it scores exactly like Transformer-VL 1B from Chapter 4, because it is that model.
Rows are cycles (H), columns are L calls per cycle. Before you tap a cell, guess its score. Then switch to the models trained on each schedule and compare.
One model, run 12 ways: Table 4 (the H2L3 checkpoint, no retraining). Four models: Table 3 (each trained from initialization with its own schedule); the other eight cells were not trained.
Three lessons come out of the grid.
Trained loops get better with more calls. In Table 3, MMStar climbs from 55.3 (H1L1, 32 calls) to 58.1 and 60.8 (64 calls) to 63.5 (H2L3, 128 calls). Every step up in trained depth pays.
The order matters, not just the count. H1L3 and H2L1 both make 64 layer calls. Trained, H2L1 beats H1L3 on all five benchmarks: 64.6 against 59.6 on RealWorldQA. Two cycles of a short refinement beat one cycle of a long one. At test time the effect is huge: H2L3 and H4L1 both make 128 calls, and MMStar is 63.5 for one and 24.7 for the other.
A trained loop does not like a new schedule. Run the H2L3 model for fewer cycles, and it collapses. With one cycle (H1L1 to H1L3) it scores almost zero: its answers mostly cannot even be parsed. Run it for more cycles, and it does not improve: H3L3 makes 192 calls and scores 58.2, H4L3 makes 256 and scores 24.0. The model does best on exactly the schedule it was trained with.
The collapse at one cycle has a simple reason, and Chapter 8 measures it. The answer forms almost entirely inside the last H call. Stop early, and the head reads a state that was never meant to be read.
The drop past two cycles is the overthinking from Chapter 1. Extra cycles push the state to places it never reached in training, and the language head no longer knows what to do with it. LoopVL makes no claim of test-time scaling. Its loop is a fixed recipe, not a dial.
Chapter 6
Four numbers that turn a heatmap into something you can plot over 128 layer calls
The hero showed a heatmap: how much attention each image patch gets. A heatmap is good to look at but hard to compare across 128 layer calls. We need a few numbers that say how the attention is spread. The paper uses four. We will build each one by hand.
Start with one query token and one attention head. The query gives weight ai to each image token i, and some more weight to text tokens. To study only the image, rescale the image weights so they add up to 1:
Now a worked case with four image tokens. Say the raw weights are a = (0.40, 0.10, 0.05, 0.05), and the other 0.40 of the attention goes to text. The image total is 0.60, so p = (0.667, 0.167, 0.083, 0.083).
−(0.667 ln 0.667 + 0.167 ln 0.167 + 2 × 0.083 ln 0.083) = 0.98; 0.98 / ln 4 = 0.71
For our four tokens, the gaps over the six unordered pairs are 0.50, 0.583, 0.583, 0.083, 0.083 and 0. They add up to 1.83. Counting each pair in both orders doubles it to 3.67.
3.67 / (2 × 4) = 0.46 (the top value for 4 tokens is 0.75)
0.667 + 0.167 = 0.83 ≥ 0.8 after 2 tokens: C0.8 = 2 / 4 = 0.50
The first three numbers describe the shape of the attention over the image. Mass describes the amount. They are independent, and that is the trap. Halve every image weight and the mass halves, but p, the entropy, the Gini and the coverage stay exactly the same. Try it in the device.
An 8 × 8 image grid. Tap or drag on cells to give them more attention. The four numbers update as you paint. The mass slider moves attention between the image and the text without changing its shape.
The four numbers are computed live from the grid with the paper's Equations 4 to 8 (natural logs, τ = 0.8). The paper averages them over queries, 12 heads and 32 samples; here there is one query.
One more caution before we use these numbers. They describe where attention goes. They do not prove that the attended patches hold the evidence, or that the answer depends on them. The paper treats attention as a diagnostic of allocation, not as an explanation. Chapter 8 adds tests that change the state directly.
Chapter 7
The same layer, run twice, reads the picture in two different ways
Now we can measure the hero. The authors take 32 diagnostic examples: four kinds of picture question, eight of each. They run each one through LoopVL and record the attention of every head at every one of the 128 layer calls. Then they compute entropy and Gini at each layer call and average them.
In the first cycle, the curves wander. Entropy stays fairly high: attention is spread across the image. Just before the cycle ends, at layer call 63, the Gini drops to 0.405. One layer call later, the first layer of the second cycle, it jumps to 0.858. The attention goes from spread out to concentrated in a single step.
Look at which layer makes the jump. Layer call 64 runs the first layer of the L stack: the same weights that ran at calls 0, 16 and 32. Those earlier runs spread their attention far more widely. The weights are identical. What changed is the state they receive: the first cycle's work, plus the new goal from the H call. The authors name this jump the Visual Aha Moment.
The slider picks one physical layer of the cycle. Left: that layer's attention in cycle 1. Right: the same layer, with the same weights, in cycle 2. Below, the curves over all 128 calls mark both moments.
Curves: Figures 6 and 7 (means over 32 samples), traced by eye; the 0.405 and 0.858 at calls 63 and 64 are printed in the paper. The heatmaps are toys built to follow those curves, after Figure 8 (b, c, h). The heatmap numbers and the promotion rate are computed live from the toy maps with the paper's rules.
A higher Gini could mean two things. Either the model focuses harder on the patches it already liked, or it moves to new patches. The paper tests this with promoted tokens.
Rank the image tokens by attention at the end of each cycle. A token is promoted if it climbs from the bottom half to the top quarter. At the end of cycle 1 it sat at or below the 50th percentile. At the end of cycle 2 it sits at or above the 75th.
Turn on Mark promoted patches at layer 63 of the cycle, the paper's comparison point (calls 63 and 127). In the paper's examples, promoted tokens land inside the answer region: the cyan squares, the code in the box, the arrow. Patches the model ignored in the first cycle are among its favourites in the second. That is a change of where it looks, not just how hard.
The paper also measures matched endpoints: the last layer of each L call and of the H call, paired across the two cycles. At every one of the four, entropy is lower in cycle 2. At the final H layer, coverage at 80% falls from about 0.42 to about 0.09. In cycle 2, under a tenth of the image tokens hold four fifths of the attention.
Each pair is one physical point of the computation, measured in cycle 1 and in cycle 2. Choose a measure.
Full-set means from Figure 9, read by eye to about ±0.02. Entropy is measured at L1, L2, L3 and H; coverage and Gini only at the final H layer.
Two parts of LoopVL could fake a jump at the cycle boundary. LoopVL adds the visual anchor back exactly there. And training sends gradients through only the last few calls (Chapter 9), which could make the boundary special. So the authors pretrain two more models from scratch: one with gradients through every call, and one with no anchor and no gate. Both show the same jump at the same place.
A last probe asks which side changed. Attention compares a query (from the reader) with keys (from the image tokens). The authors mix early and late queries with early and late image keys. Swapping in the late image keys changes the attention more than swapping in the late queries. So the shift comes mostly from the image tokens themselves: the picture's own state has been rewritten.
Chapter 8
Stop the image tokens from changing and the model gets worse; read the answer early and there is none
Chapter 7 showed that attention moves. Attention is only how the model reads. This chapter looks at what it reads: the image tokens themselves. Do they keep changing in the second cycle, and does that matter for the answer?
The cleanest test is to stop the change and see what breaks. The authors compare four conditions. The authors apply two of them only at test time, to the normal model. One is a separate model trained with the restriction from the start.
In all four, the text tokens keep updating and the image stays readable. Only the image's own state is held back.
Choose a benchmark. Each bar is one condition. The first three use the same trained weights; the fourth is a model trained with frozen cycle-2 image tokens.
Bar heights read from Figure 3 by eye, to about ±1 point; the paper prints no numbers for this figure.
The normal model wins on all four benchmarks. Freezing the image tokens hurts most: on RealWorldQA, accuracy falls from about 72 to about 38. Training the model to live with the freeze helps, but it still stays far below normal: about 58. So the gain of the second cycle is not only "read the same picture again". The picture's own tokens must keep being rewritten.
Next, the authors measure the change directly. For each image token, compare its vector before and after one call, and take the length of the difference. Average over tokens and samples:
If the second cycle only copied the first, its updates would be close to zero. They are not. Updates in cycle 2 are about as large as in cycle 1: between about 48 and 66 per call in both halves. These are lengths of vector differences, in the units of the model's 1,536-number hidden vectors.
The paper also compares directions with cosine similarity, the cosine of the angle between two vectors: 1 means they point the same way. It finds the image vectors keep turning. A linear probe is a single trained linear map; here it tries to recover the original image features from each state. How well it works varies from call to call, and at the end it works worse than right after the projector. The image tokens drift away from "what the encoder saw" toward something else.
The last probe reads the answer early. Take the state after any layer call. Apply the stack's final normalisation and the shared language head, and you get a guess at the first answer token. This is the logit lens. Compare each early guess with the final one using the KL divergence, a measure of how far one distribution is from another:
A small case with three answers: the final distribution is (0.90, 0.05, 0.05) and an early guess is (0.30, 0.40, 0.30).
0.90 ln(0.90/0.30) + 0.05 ln(0.05/0.40) + 0.05 ln(0.05/0.30) = 0.99 − 0.10 − 0.09 = 0.80 nats
For scale: a flat guess over 100,000 tokens sits about 11 nats from a confident answer. So a KL of 11 or 12 means the early readout knows almost nothing yet.
Two views of the same 32 samples. Choose a view, then drag the playhead through the 128 layer calls.
KL curve: Figure 11 (mean of 32 samples, first answer token), traced by eye. Update sizes: Figure 10 (mean L2 change per 16-layer call), read by eye. The answer bars under the curve are a toy distribution built so that its KL from the final one follows the traced curve.
The distance stays high, around 9 to 12 nats, for most of the run. It dips a little during each H call. Then, inside the very last H call, it falls to zero within about 16 layer calls. The model forms the answer at the end, in the final H call, and almost nowhere else.
That explains the collapse in Chapter 5. Run only one cycle, and the head reads a state from before the final H call. By this curve, that state has not formed an answer yet. The near-zero scores were not a mystery: the model was asked before it had decided.
Chapter 9
Five stages, 0.14 trillion tokens, and gradients that skip the first three calls
The authors build LoopVL in five stages. Each stage decides which parts may learn and which stay frozen. The vision encoder never learns: it stays frozen in every stage.
Choose a stage. The picture shows which parts learn (solid) and which are frozen (hatched), and how many tokens the stage uses on a log scale.
Stages, data and settings: Sections 3 and 4 and Figure 4. Token totals as printed; they add up to about 141 billion, the 0.14T of Tables 1 and 2.
Three choices in the ladder are worth a second look.
The loop learns language first. The authors pretrain their own loop on 75 billion text tokens with the open HRM-Text code and data, plus 15 billion tokens of their own. They do not use HRM-Text's released checkpoint. Every later stage starts from that text loop.
Alignment trains only the bridge. In Stage 1 both the loop and the encoder are frozen. Only the projector learns, from 559,000 image-caption pairs. Its job is narrow: map encoder features into a form the frozen loop already understands. Before the main run, the authors swap images in and out and check that the answers change. That confirms the bridge carries image information.
Long answers are dropped, not cut. A few text-heavy samples have answers longer than the sequence budget. Stage 2 removes them instead of truncating them, so the model never learns from an answer that stops halfway.
Supervised stages use the standard next-token loss, but only on the answer. The image and the question are context; the model is not graded on them. The model factors the answer y = (y1, …, yT) one token at a time, each conditioned on the context x and the answer so far. The loss averages over every valid answer token in the batch:
A batch of two: sample 1 has three answer tokens with log-probabilities −0.1, −0.5 and −0.2; sample 2 has one, at −1.2.
−(−0.1 − 0.5 − 0.2 − 1.2) / (3 + 1) = 2.0 / 4 = 0.5 (a per-sample mean would give (0.27 + 1.2) / 2 = 0.73)
The final stage is reinforcement learning with GRPO, group relative policy optimisation. For each image and question, the model writes a group of answers. A correct final answer earns a positive reward and anything else zero. Each answer's advantage is its reward measured against its own group. So GRPO needs no value network, the second model that other RL methods train to predict the expected reward.
During this stage the authors raise the maximum answer length from 2,048 to 4,096 tokens. The model first learns short correct reasoning, then gets room for long derivations. As the length limit rises, the group size shrinks to keep the cost level.
The forward pass always runs all eight calls. The backward pass does not. During training, gradients flow back only through the last few calls. Early in training it is the last two, L6 and H2. During a warm-up the window grows to the last five: L4, L5, L6, H1 and H2. The first three calls, L1 to L3, never send gradients back.
The eight calls, forward left to right. Move the warm-up slider to grow the backward window. Watch which stored stack still learns.
The window, from the last two calls to the last five, is from Section 2.1. The paper gives no reason for it. Saving memory and keeping long gradient paths stable are the usual reasons for such schemes; that part is our reading.
Here is the subtle part. L1, L2 and L3 send no gradients, but the L stack still learns. It shares its weights with L4, L5 and L6, which are inside the window. Every gradient through L4 to L6 updates the same weights that L1 to L3 use. Skipping the early calls saves the memory of storing their activations for the backward pass. It costs no weights that never learn.
A note on evaluation. All headline scores use greedy decoding and a direct answer: an option letter, yes or no, or a short reply, usually within 32 tokens. The model is not asked to reason step by step. The authors report that step-by-step prompting helps a little on maths and hurts on short-answer tasks.
Chapter 10
A 1B loop against eleven compact models trained on 21 to 260 times more data
Chapter 4 compared LoopVL with models trained on the same data. The real world is less fair. Most compact VLMs today are trained on 3 to 36 trillion tokens. LoopVL saw 0.14 trillion. The authors compare it with eleven of them anyway, across 16 benchmark families: general understanding, hallucination, maths and science.
Choose a benchmark, or switch to the data view: training tokens against the average of three reasoning benchmarks.
Scores and training tokens: Table 2 (in-house evaluations combined with selected published results). The three-benchmark average (LogicVista, MMMU-Pro, MathVision) is our mean from Table 2, the measure of Figure 1.
LoopVL is strongest on visual reasoning. On LogicVista it scores 45.2, ahead of the next model (InternVL3.5-4B) by 8.3 points. On MMK12-Math it leads with 61.6 against 54.8. It also tops MMMU-Pro (38.7), MMStar (63.5), RealWorldQA (71.0) and MathVision (38.5, a hair over Qwen3.5-2B's 38.4).
It is weaker where reading fine text and charts matters. On ChartQA it scores 74.5 against a best of 86.0. On AI2D, a diagram benchmark, 75.5 against 82.8. On HallusionBench, which tests whether a model invents things, 44.7 against 51.6. On MMEval-Pro, 24.6 against 42.8. A model that saw 0.14 trillion tokens has seen far fewer charts and documents than one that saw 36 trillion.
Switch to the data view. On the three-benchmark average, LoopVL sits at about 40.8 with 0.14T tokens. The next best, InternVL3.5-4B, sits at about 32.4 with 36T. Every other model is to the right of LoopVL by a factor of 21 to 260 in data, and below it.
Chapter 11
The cheat sheet, the code, and where LoopVL sits in the field
You can now explain each part of LoopVL, from the loop itself to the place where the answer forms. Let's lock it in.
LoopVL is a 1B-scale vision-language model whose language backbone is a loop. It stores two 16-layer stacks, L and H. It runs them in the order L L L H L L L H: 128 layer calls from 32 stored layers. A frozen Penguin encoder and a small projector feed one token per image patch, calibrated to word-token size, with 2D positions and a PrefixLM mask.
Trained on 0.14 trillion tokens, it beats its non-looped twin by 6 to 23 points and roughly matches 4B dense models on the same data. Its attention over the image jumps from spread to concentrated at the cycle boundary, the Visual Aha Moment. Its image tokens keep changing in the second cycle, and freezing them hurts. Its answer forms in the final H call.
| Quantity | Value | Why it matters |
|---|---|---|
| Stored layers | 32 (L: 16, H: 16), width 1,536 | The weights of a 1B model |
| Schedule | H2L3: L L L H L L L H | 128 layer calls per forward pass |
| Image tokens | 256 to 1,024 per image, one per patch | A stable grid for the diagnostics |
| Loop vs its twin | MMStar 55.3 → 63.5, ChartQA 51.1 → 74.5 | Same 32 layers, same data |
| Training FLOPs | 2.47 vs 0.92 (twin), 2.89 and 2.99 (4B), × 1021 | The loop saves weights, not compute |
| Equal depth, different order | H2L3 63.5 vs H4L1 24.7 (MMStar, 128 calls each) | Order matters, not just count |
| Off-schedule | 1 cycle ≈ 0; H3L3 58.2; H4L3 24.0 | Best only on its trained schedule |
| Visual Aha Moment | Gini 0.405 → 0.858 at calls 63 → 64 | Same weights, new reading of the image |
| Answer formation | KL ≈ 7.5 → 0 nats inside the last H call | Why early readouts fail |
| Training data | 75B text + 0.33B + 49B + 16.81B + 0.006B ≈ 0.14T | 21 to 260 times less than its rivals |
python# LoopVL in pseudo-Python. Frozen: penguin. Trained: projector, L_stack, H_stack, gate, scales, head. def embed(image, prompt): f = penguin(image) # N_img x 1024, one per kept patch (256 to 1024) v = projector(f) # 1024 -> 1536 -> 1536, GELU between v = rms_normalize(v) * median_rms(word_emb) # same size as a word token x = build_sequence(header, v, prompt) # image: 2D RoPE (row, col); text: (t, t) return x, v def forward(image, prompt, H=2, L=3, window=5): x, v = embed(image, prompt) zH = embed_scale * x # the H state starts from the input zL = zeros_like(x) # the L state starts at zero, kept across cycles calls = H * (L + 1); i = 0 for cycle in range(H): # Model-Loop zH[img] += scale[cycle] * gate(zH, prompt) * v # re-inject the visual anchor for _ in range(L): # Module-Loop with grad_if(i >= calls - window): zL = L_stack(combine(zL, zH)) # 16 layers i += 1 with grad_if(i >= calls - window): zH = H_stack(combine(zH, zL)) # 16 layers i += 1 return lm_head(norm(zH)) # the answer is read from the final H state only def train_step(batch, step): w = 2 + round(3 * warmup(step)) # backward window grows from 2 calls to 5 logits = forward(batch.image, batch.prompt, window=w) # PrefixLM mask inside every layer return cross_entropy(logits[answer], batch.answer) # Eq. 3: mean over valid answer tokens
Now press Present or Teach and explain, out loud and from memory, why the same layer can look at different patches on its second run. If you can, you own this paper. Then go back to the picture and drag the layer call across 63 and 64.