LoopVL

A vision-language model that stores 32 layers but runs 128, by looping the same layers. Looping lifts its MMStar score from 55.3 to 63.5 with no extra layers. Halfway through, its attention snaps onto the part of the picture that holds the answer.

Press Play and watch where the model looks while the loop runs. Then we build, piece by piece, why the same weights look somewhere new the second time.

You need what a neural network layer is and the idea that a transformer reads a list of tokens. We build the rest from zero.

Watch it look twice

Ready
0

The scenes are toys after Figure 8 (b, c, h). The heatmaps are illustrative, built so that their concentration follows the Gini curve of Figure 7, traced by eye (the jump from 0.405 to 0.858 between calls 63 and 64 is printed in the paper). The answer bars follow the shape of Figure 11: the answer forms in the last H call. The 55.3 and 63.5 are Table 1.

Chapter 0

Looking once

Why a model that reads a picture in one pass can run out of steps, and what a loop offers instead

Show a model a picture of a cluttered table and ask: how many cyan squares are there? A person does not answer in one glance. You scan the picture, find the shapes, then look again at the cyan ones and count them.

A vision-language model (VLM) answers questions about pictures. It cuts the picture into small patches and turns each patch into a list of numbers called a token. It puts those image tokens in a row with the tokens of the question. Then a stack of layers processes the row. Each layer updates every token a little, and after the last layer the model writes its answer.

Inside each layer, attention lets a token read from the other tokens. The token gives every other token a weight, the weights add up to 1, and it takes in a weighted mix of them. When the answer token puts large weights on a few image tokens, we say the model is looking at those patches.

In an ordinary model, every layer runs once. A 32-layer model gets exactly 32 updates to find the cyan squares, count them and write the answer. If a question needs more steps, the usual fix is a deeper model. But every extra layer has its own weights, so depth costs memory and training data.

A looped transformer takes another route. It stores a few layers and runs them again and again, each time on its own last output. Depth now costs time instead of weights. The idea is not new: the Universal Transformer did it in 2019. It has come back for language models that need more steps of thought.

Stack it or loop it

Two ways to get the same number of layer calls. Each coloured block is one stored layer. Drag the slider and compare what each model must store.

128

The weight counts are our arithmetic, not the paper's: about 12 × 1,536² numbers per layer (attention plus a SwiGLU feed-forward), at LoopVL's width of 1,536. Embeddings and the vision encoder are left out. The real LoopVL loops two different 16-layer stacks; the drawing shows the stored layers as one set of 32.

There is a catch, and it is the reason this paper exists. In a loop, the weights never change. If the input to the loop never changed either, every pass would compute the same thing, and the loop would be useless. A loop only helps because the state it works on changes from pass to pass.

For pictures, that raises a new question. After one pass, the model has read the question and the image together. Perhaps it now knows that the cyan squares matter and the big pink square does not. On the next pass, can the same layers look at different patches?

LoopVL (Qian, Gong, Xu and colleagues, 2026) tests exactly that. It is a VLM at the 1B scale whose language part is a loop. The loop has two stacks of 16 layers, run in a fixed order for 128 layer calls. The authors pretrain that loop from scratch and train the whole model on 0.14 trillion tokens. Then they measure where the model looks on each pass.

Two findings carry the paper. First, the loop works. With the same 32 stored layers, looping lifts MMStar from 55.3 to 63.5 (Table 1). MMStar is a benchmark: a fixed set of picture questions, scored as the percentage answered correctly. The paper uses about 16 of them, and every score in this lesson is a percentage of that kind. The 1B model also beats several 2B to 4B models trained on 79 to 260 times more tokens.

Second, the loop looks again. At the boundary between the first and the second cycle, the attention on the image jumps from spread out to concentrated. The authors call this the Visual Aha Moment.

The idea in one line: same weights, new state, new place to look. A loop stores its layers once and runs them many times. It only computes something new on each pass because its input, the state, has changed. For a picture, that means the second pass can read different patches than the first.

Here is the road, in order. Each chapter explains one part of the instrument in the hero.

  1. One block, many passes: how fixed weights compute something new on every pass.
  2. Fast and slow: LoopVL's two stacks, L and H, and the order they run in. This is the loop track in the hero.
  3. Pictures into the loop: how image patches become tokens, and who may read whom.
  4. Why loop at all: the same data and four model shapes, compared fairly.
  5. The schedule matters: the same weights, run in twelve different orders.
  6. Measuring where it looks: entropy, Gini, coverage and mass, built by hand. This is the curve in the hero.
  7. The Visual Aha Moment: the heart of the paper, layer by layer, with promoted patches.
  8. Inside the state: freezing the image tokens, and where the answer forms. This is the answer bar in the hero.
  9. Training it, what it buys, then connections.
A looped transformer runs the same weights on every pass. Why can a second pass compute something the first pass did not?

Chapter 1

One block, many passes

How a function with fixed weights can do new work on every pass

Start with the smallest possible loop. You have one block of layers, a function with weights θ. It takes a state and the input, and it returns a new state. Run it, then feed its output back in, and run it again.

h(k+1)
The state after pass k + 1. In a transformer it is one vector per token, so a whole table of numbers.
fθ
The shared block. In LoopVL it is a stack of 16 transformer layers. The weights θ are the same on every pass.
h(k)
The state the block receives on this pass: its own output from the pass before.
x
The input, here the image and the question. Some loops add it back on every pass, so the state never drifts far from the evidence.

Let's run one with numbers, in one dimension. Take the block f(h, x) = 0.5 h + 0.5 x, the input x = 2, and the start h = 0. Each pass moves the state halfway to 2.

0 → 1 → 1.5 → 1.75 → 1.875 (the same rule each time; the steps are 1, 0.5, 0.25, 0.125)

The rule never changed, but every pass did something different: it moved the state by a different amount. That is the whole trick. A shared block is a rule, and a rule applied to a new state gives a new result. The device shows the same idea in two dimensions.

Same weights, new state

The faint arrows are one fixed rule: from any state, the arrow shows where one pass moves it. Add passes and watch the state travel. Then change the question and run the same rule again.

3

A toy rule in two dimensions, not the paper's model: each pass moves the state 35% of the way toward a point set by the question, plus a small turn. A real block works on thousands of numbers per token, but the logic is the same.

Notice two things. First, the steps shrink as the state nears the point the rule leads to. The same weights do a big job on pass 1 and a small one on pass 8. Second, a new question moves the point, so the same rule sends the same start somewhere else.

Why would a model want this? Some problems are naturally step by step: counting, following a path, checking a guess and fixing it. A loop gives the model more steps for them without more weights. On such tasks, earlier work found that a shallow shared block, looped, comes close to a much deeper model (Saunshi and colleagues, 2025).

The cost moves elsewhere. A loop of 16 layers run 8 times does as much arithmetic as 128 separate layers. It just stores fewer of them. So a loop trades memory for time: it is small to hold but not cheaper to run. Chapter 4 counts this cost for LoopVL.

One more property matters later. A loop can run more passes at test time than it saw in training, or fewer. Does that help?

For some loops, running longer keeps improving the answer. For others, extra passes push the state somewhere it never went in training, and the answer gets worse. Researchers call this overthinking. Chapter 5 shows which kind LoopVL is.

With f(h, x) = 0.5 h + 0.5 x, x = 2 and h = 0 at the start, what is the state after the third pass?

Chapter 2

Fast and slow

LoopVL loops two stacks at two speeds, and the order is fixed

LoopVL does not loop one block. It loops two, and it runs them at two speeds. The design comes from the Hierarchical Reasoning Model (HRM) and its language version, HRM-Text. Their idea comes from the brain: fast local work happens under a slower, high-level goal that changes less often.

A picture question needs both kinds of work. Some of it is fine detail: a colour, a letter, an edge, which shape sits left of which. Some of it is the big picture: what the question asks, and what the scene is. LoopVL gives each kind its own stack of 16 layers.

One call means one run of a whole 16-layer stack. One cycle is three L calls then one H call. LoopVL runs two cycles. Written out, the order is the paper's Equation 1:

LLLH|LLLH

The paper names this schedule H2L3: two H calls (two cycles), three L calls per cycle. Repeating L inside a cycle is the Module-Loop. Repeating the whole cycle is the Model-Loop. Count the layer calls: 2 cycles × 4 calls × 16 layers = 128. The model stores only 2 × 16 = 32 layers.

Each stack carries its own state. The H state starts as the input itself: the image and text tokens, scaled. The L state starts at zero and is kept from one cycle to the next. An L call reads the L state together with the current H state, the goal it works under. The H state stays fixed during the three L calls. Then the H call reads the refined L state and writes a new H state, the goal for the next cycle.

Step through the schedule

Two lanes, one per state. Step through the eight calls and watch which stack runs, what it reads and what it writes. The counter compares layer calls with stored layers.

The order, the state rules and the re-injection follow Section 2.1. How exactly two states are combined (HRM adds them) is not spelled out for LoopVL; the arrows show what each call reads, not the arithmetic.

Toggle Re-inject the picture and look at the start of each cycle. LoopVL adds back a fixed copy of the projected image tokens into the image positions of the H state. This copy is the visual anchor. A gate looks at the question and sets a strength for each image token, and a learned scale per cycle sets the overall strength. The anchor keeps the original picture within reach, however far the state has moved.

The final H state goes to the language-model head, the layer that turns a vector into a probability for every word in the vocabulary. The answer is read out there, and only there.

Inside each stack, every one of the 16 layers has the same shape. It is a standard transformer layer with two changes worth knowing. First, recall how attention works inside a layer. Each token makes three vectors: a query (what it looks for), a key (what it offers) and a value (what it passes on). Each query is compared with every key, the matches set the attention weights, and the token takes in the values mixed by those weights.

RMSNorm
rescale each token vector to a standard size
↓
Gated self-attention
one linear map makes the queries, keys, values and a gate g; positions enter by rotation (RoPE, Chapter 3); the attention output is multiplied by σ(g), the sigmoid of the gate, a number between 0 and 1, before the output map
↓ + residual
RMSNorm, then SwiGLU feed-forward
a gated two-path MLP applied to each token on its own
↓ + residual   (× 16 layers per stack)

The gate lets each layer turn its attention output down, token by token. This helps in a loop. A layer that runs six times may have nothing to add on some passes, and the gate lets it say so.

Why two speeds and not one long loop? Running L three times under a fixed goal gives dense, focused refinement. Updating the goal only once per cycle lets that refinement settle before the plan changes. Chapter 5 tests this: two schedules with the same number of layer calls score very differently.
Under H2L3, how many times does the L stack run, and how many layer calls does one forward pass make?

Chapter 3

Pictures into the loop

How image patches become tokens, where they sit, and who may read whom

The authors pretrained the loop on text. To read pictures it needs an entrance: a way to turn pixels into tokens that look, to the loop, like the word tokens it already knows. LoopVL builds that entrance in three parts.

Penguin Vision Encoder (frozen)
a pretrained vision transformer; each kept patch of the image becomes 1,024 numbers
↓
Projector (trained)
two layers with a GELU between them: 1,024 → 1,536 → 1,536, the width of the loop
↓
Calibration (no weights)
rescale each image token to the typical size of a word token
↓
Into the image slots of the sequence
header, image, question, answer

The encoder stays frozen through all of training: its weights never change. Only the projector learns to map its features into the loop's space. One patch gives one token. LoopVL does not merge or drop image tokens, so the image keeps a stable grid. That grid is what lets the authors draw attention back onto the picture in Chapter 7.

The image size decides the number of tokens: between 256 and 1,024 per image. Each must fit in the sequence budget, 2,048 tokens in the first training stage and 3,072 in the second. A 1,024-token image plus a 60-token question leaves 964 tokens of room in the first stage.

The calibration step deserves a number. The size of a vector here is its RMS, the square root of the mean of its squared entries. Projected image features can be much larger or smaller than word embeddings. If an image token arrived four times louder than a word, attention would treat it very differently. So LoopVL first normalises each image token to RMS 1. Then it multiplies by the median RMS of the loop's word embeddings.

RMS 3.2 → ÷ 3.2 → RMS 1 → × 0.8 → RMS 0.8 (illustrative numbers: the image token now matches a typical word)

Where each token sits

A transformer needs to know where each token is. LoopVL uses RoPE, rotary position embedding: it rotates parts of the query and key vectors by an angle that grows with the position. Two tokens then compare their positions through the difference of their angles.

Image tokens get two positions, a row and a column, so a patch knows its place on the grid: this is 2D RoPE. Text tokens also get two coordinates, but both advance together: (5, 5), (6, 6), (7, 7). With equal coordinates, the 2D rotation acts exactly like the ordinary 1D rotation the loop learned on text. So the text side of the pretrained loop sees nothing new.

Who may read whom

The last piece is the attention mask: a table of which token may read which. A language model is usually causal: each token reads only the tokens before it. LoopVL uses a PrefixLM mask instead. The header, the image and the question form the prefix, and inside the prefix every token reads every other, in both directions. The answer tokens read the whole prefix and the answer tokens before them. No prefix token may read the answer.

Who can read whom

A tiny sequence: two header tokens, a 2 × 3 image, three question tokens and three answer tokens. Each row is a reader, each column a token it may read. Tap a row, then switch the mask.

The mask rules and the position scheme follow Section 2.1 and Section 6.3. The token counts are shrunk to fit; a real image has 256 to 1,024 tokens.

Why does the mask matter for a loop? Under a causal mask, the first image patch could never read the question, because the question comes later. Under PrefixLM, every image token reads the question on every pass. So an image token can change because of what was asked. That is the precondition for the whole paper: the image state can only evolve toward the question if it can see the question.

Why does LoopVL give text tokens two equal coordinates, (t, t), instead of a new 2D scheme?

Chapter 4

Why loop at all

The same data and four model shapes, compared fairly

A loop makes a model deeper without more weights. But is a loop the best way to spend that depth? Perhaps you would do better to store more layers, or wider ones. The fair test holds everything else fixed: the same data, the same training budget, the same layer design. Only the shape changes.

The authors train four models on the same 0.14 trillion tokens. All use the same transformer layer as LoopVL. They differ in how many layers they store, how wide the layers are, and whether they loop. The last column counts training cost in FLOPs, floating-point operations: the total number of multiplications and additions the training run performed.

ModelStored layersWidthLayer callsTraining FLOPs
LoopVL (H2L3)32 (16 L + 16 H)1,5361282.47 × 1021
Transformer-VL 1B321,536320.92 × 1021
Transformer-VL 4B, deep781,792782.89 × 1021
Transformer-VL 4B, wide322,816322.99 × 1021

Read the first two rows together. LoopVL and the 1B model store the same 32 layers at the same width. The only difference is that LoopVL runs them four times as often. The other two rows spend about three times the weights, one on depth and one on width.

Same data, four shapes

Choose a benchmark. Each bar is one model, all trained on the same 0.14 trillion tokens. The labels under the bars show what each model stores and what it costs to train.

Scores and FLOPs: Table 1. The average is our mean of all eight benchmarks in Table 1 (also VMCBench, AI2D and VisuLogic). Weights in the layers are our estimate at 12 × width² per layer.

Against its twin, the loop wins every benchmark, by 6 to 23 points. MMStar goes from 55.3 to 63.5, RealWorldQA from 55.3 to 71.0, ChartQA from 51.1 to 74.5. Since the stored layers are identical, the gain comes from running them again.

Against the 4B models, LoopVL wins five of eight benchmarks and loses three (VMCBench, AI2D, ChartQA) by at most 1.1 points. On the eight-benchmark average it leads both. So a model with about a third of the stored weights matches or beats models three times its size, on the same data.

What the loop costs

Now the honest part. Looping is not free. LoopVL used 2.7 times the training compute of its 1B twin (2.47 against 0.92, in units of 1021 FLOPs). It is the stored weights that stay small, not the arithmetic. Against the 4B models, LoopVL used 14 to 17% fewer FLOPs, so it is slightly cheaper there too, but in the same range.

The same holds when the model answers. Each answer token runs 128 layer calls, as many as a 128-layer model. The calls must happen in order, one after another, so a loop cannot run its passes in parallel. What the loop saves is memory: a phone or a small GPU holds 32 layers, not 128.

What the comparison proves, and what it does not. Proved: at a fixed data budget, looping 32 layers beats storing 32 layers, and roughly matches storing three times more. Not proved: that a loop is the cheapest way to reach a given score per FLOP. The loop wins on weights; on compute it is roughly even with the 4B models.
LoopVL and Transformer-VL 1B store the same 32 layers at the same width. Which statement about their costs is true?

Chapter 5

The schedule matters

The same weights, run in twelve different orders, and four models trained four ways

If more layer calls help, why stop at 128? A loop can, in principle, run any number of passes at test time. The paper asks two separate questions, and it matters to keep them apart.

  1. Train it differently. Train four separate models from scratch, each with its own schedule: H1L1, H1L3, H2L1, H2L3. Each keeps its schedule through every stage of training. (Table 3)
  2. Run it differently. Take the one trained H2L3 model and only change how it runs at test time: 1 to 4 cycles, 1 to 3 L calls per cycle. Twelve schedules, the same weights. (Table 4)

The number of layer calls for any schedule is H × (L + 1) × 16. H3L3, for example, is 3 cycles of 4 calls: 192 layer calls. Look at H1L1 first. It is one L call and one H call, 32 layer calls in all: a plain model with no loop. In Table 3 it scores exactly like Transformer-VL 1B from Chapter 4, because it is that model.

Run the same weights differently

Rows are cycles (H), columns are L calls per cycle. Before you tap a cell, guess its score. Then switch to the models trained on each schedule and compare.

One model, run 12 ways: Table 4 (the H2L3 checkpoint, no retraining). Four models: Table 3 (each trained from initialization with its own schedule); the other eight cells were not trained.

Three lessons come out of the grid.

Trained loops get better with more calls. In Table 3, MMStar climbs from 55.3 (H1L1, 32 calls) to 58.1 and 60.8 (64 calls) to 63.5 (H2L3, 128 calls). Every step up in trained depth pays.

The order matters, not just the count. H1L3 and H2L1 both make 64 layer calls. Trained, H2L1 beats H1L3 on all five benchmarks: 64.6 against 59.6 on RealWorldQA. Two cycles of a short refinement beat one cycle of a long one. At test time the effect is huge: H2L3 and H4L1 both make 128 calls, and MMStar is 63.5 for one and 24.7 for the other.

A trained loop does not like a new schedule. Run the H2L3 model for fewer cycles, and it collapses. With one cycle (H1L1 to H1L3) it scores almost zero: its answers mostly cannot even be parsed. Run it for more cycles, and it does not improve: H3L3 makes 192 calls and scores 58.2, H4L3 makes 256 and scores 24.0. The model does best on exactly the schedule it was trained with.

The collapse at one cycle has a simple reason, and Chapter 8 measures it. The answer forms almost entirely inside the last H call. Stop early, and the head reads a state that was never meant to be read.

The drop past two cycles is the overthinking from Chapter 1. Extra cycles push the state to places it never reached in training, and the language head no longer knows what to do with it. LoopVL makes no claim of test-time scaling. Its loop is a fixed recipe, not a dial.

Where this leaves "more thinking at test time". Some looped language models are trained so that more passes keep helping. LoopVL is not, and the paper says so: the sweep "does not establish length extrapolation, an adaptive stopping rule, or unlimited test-time scaling." Recurrent language models that do improve with more passes get there through different training: Huginn, for example, draws the number of passes at random for each training step.
H2L3 and H4L1 both make 128 layer calls with the same trained weights. Why do they score 63.5 and 24.7 on MMStar?

Chapter 6

Measuring where it looks

Four numbers that turn a heatmap into something you can plot over 128 layer calls

The hero showed a heatmap: how much attention each image patch gets. A heatmap is good to look at but hard to compare across 128 layer calls. We need a few numbers that say how the attention is spread. The paper uses four. We will build each one by hand.

Start with one query token and one attention head. The query gives weight ai to each image token i, and some more weight to text tokens. To study only the image, rescale the image weights so they add up to 1:

pi
The share of the image attention that goes to image token i. The p values add up to 1 over the image.
ai
The raw attention weight on image token i, from one query in one head.
∑j aj
The total raw weight on all N image tokens. Below we call it the mass, M.

Now a worked case with four image tokens. Say the raw weights are a = (0.40, 0.10, 0.05, 0.05), and the other 0.40 of the attention goes to text. The image total is 0.60, so p = (0.667, 0.167, 0.083, 0.083).

1. Spread: normalised entropy

E(p)
How spread out the attention is, from 0 (all on one token) to 1 (equal on all tokens).
1 / log N
Divides by the largest possible entropy, so images with different token counts compare fairly.
∑ pi log pi
The usual entropy sum, with natural logs. A token with p = 0 adds 0.

−(0.667 ln 0.667 + 0.167 ln 0.167 + 2 × 0.083 ln 0.083) = 0.98;   0.98 / ln 4 = 0.71

2. Inequality: the Gini concentration

G(r)
How unequal the shares are: 0 if all tokens get the same share, close to 1 if one token gets almost all.
1 / 2N
The scale that makes the sum over every pair a fraction. The top value is (N − 1) / N.
∑ |ri − rj|
The gap between the shares of every ordered pair of tokens. r is a share distribution like p.

For our four tokens, the gaps over the six unordered pairs are 0.50, 0.583, 0.583, 0.083, 0.083 and 0. They add up to 1.83. Counting each pair in both orders doubles it to 3.67.

3.67 / (2 × 4) = 0.46 (the top value for 4 tokens is 0.75)

3. Footprint: coverage

Cτ(p)
The fraction of image tokens needed to hold a share τ of the attention. Lower means a smaller footprint.
kτ
Sort the tokens from most to least attention. kτ is how many it takes, from the top, to reach a total share of at least τ. The paper uses τ = 0.8.
N
The number of image tokens.

0.667 + 0.167 = 0.83 ≥ 0.8 after 2 tokens:   C0.8 = 2 / 4 = 0.50

4. Amount: the mass

M
The total share of the query's attention that goes to the image at all, before rescaling. Here M = 0.60.
∑ ai
The raw image weights, summed. The rest of the attention, 1 − M, goes to text tokens.

The first three numbers describe the shape of the attention over the image. Mass describes the amount. They are independent, and that is the trap. Halve every image weight and the mass halves, but p, the entropy, the Gini and the coverage stay exactly the same. Try it in the device.

Paint the attention

An 8 × 8 image grid. Tap or drag on cells to give them more attention. The four numbers update as you paint. The mass slider moves attention between the image and the text without changing its shape.

1.00entropy E
0.00Gini G
0.80coverage C0.8
0.60mass M
0.60

The four numbers are computed live from the grid with the paper's Equations 4 to 8 (natural logs, τ = 0.8). The paper averages them over queries, 12 heads and 32 samples; here there is one query.

One more caution before we use these numbers. They describe where attention goes. They do not prove that the attended patches hold the evidence, or that the answer depends on them. The paper treats attention as a diagnostic of allocation, not as an explanation. Chapter 8 adds tests that change the state directly.

You multiply every image attention weight by 0.5. What happens to the four numbers?

Chapter 7

The Visual Aha Moment

The same layer, run twice, reads the picture in two different ways

Now we can measure the hero. The authors take 32 diagnostic examples: four kinds of picture question, eight of each. They run each one through LoopVL and record the attention of every head at every one of the 128 layer calls. Then they compute entropy and Gini at each layer call and average them.

In the first cycle, the curves wander. Entropy stays fairly high: attention is spread across the image. Just before the cycle ends, at layer call 63, the Gini drops to 0.405. One layer call later, the first layer of the second cycle, it jumps to 0.858. The attention goes from spread out to concentrated in a single step.

Look at which layer makes the jump. Layer call 64 runs the first layer of the L stack: the same weights that ran at calls 0, 16 and 32. Those earlier runs spread their attention far more widely. The weights are identical. What changed is the state they receive: the first cycle's work, plus the new goal from the H call. The authors name this jump the Visual Aha Moment.

The same layer, twice

The slider picks one physical layer of the cycle. Left: that layer's attention in cycle 1. Right: the same layer, with the same weights, in cycle 2. Below, the curves over all 128 calls mark both moments.

63

Curves: Figures 6 and 7 (means over 32 samples), traced by eye; the 0.405 and 0.858 at calls 63 and 64 are printed in the paper. The heatmaps are toys built to follow those curves, after Figure 8 (b, c, h). The heatmap numbers and the promotion rate are computed live from the toy maps with the paper's rules.

Is it the same patches, just brighter?

A higher Gini could mean two things. Either the model focuses harder on the patches it already liked, or it moves to new patches. The paper tests this with promoted tokens.

Rank the image tokens by attention at the end of each cycle. A token is promoted if it climbs from the bottom half to the top quarter. At the end of cycle 1 it sat at or below the 50th percentile. At the end of cycle 2 it sits at or above the 75th.

Turn on Mark promoted patches at layer 63 of the cycle, the paper's comparison point (calls 63 and 127). In the paper's examples, promoted tokens land inside the answer region: the cyan squares, the code in the box, the arrow. Patches the model ignored in the first cycle are among its favourites in the second. That is a change of where it looks, not just how hard.

The paper also measures matched endpoints: the last layer of each L call and of the H call, paired across the two cycles. At every one of the four, entropy is lower in cycle 2. At the final H layer, coverage at 80% falls from about 0.42 to about 0.09. In cycle 2, under a tenth of the image tokens hold four fifths of the attention.

Matched endpoints

Each pair is one physical point of the computation, measured in cycle 1 and in cycle 2. Choose a measure.

Full-set means from Figure 9, read by eye to about ±0.02. Entropy is measured at L1, L2, L3 and H; coverage and Gini only at the final H layer.

Could it be an artefact?

Two parts of LoopVL could fake a jump at the cycle boundary. LoopVL adds the visual anchor back exactly there. And training sends gradients through only the last few calls (Chapter 9), which could make the boundary special. So the authors pretrain two more models from scratch: one with gradients through every call, and one with no anchor and no gate. Both show the same jump at the same place.

A last probe asks which side changed. Attention compares a query (from the reader) with keys (from the image tokens). The authors mix early and late queries with early and late image keys. Swapping in the late image keys changes the attention more than swapping in the late queries. So the shift comes mostly from the image tokens themselves: the picture's own state has been rewritten.

Why vision makes this visible. In a language model, a token in the middle of a loop has no fixed place in the world. An image token does: it is a patch at a row and a column. That grid lets the authors pin the same physical layer, run it twice, and draw both readings on the same picture. The picture becomes a window into the loop.
What makes the Visual Aha Moment surprising?

Chapter 8

Inside the state

Stop the image tokens from changing and the model gets worse; read the answer early and there is none

Chapter 7 showed that attention moves. Attention is only how the model reads. This chapter looks at what it reads: the image tokens themselves. Do they keep changing in the second cycle, and does that matter for the answer?

Freeze the picture

The cleanest test is to stop the change and see what breaks. The authors compare four conditions. The authors apply two of them only at test time, to the normal model. One is a separate model trained with the restriction from the start.

In all four, the text tokens keep updating and the image stays readable. Only the image's own state is held back.

Freeze the picture

Choose a benchmark. Each bar is one condition. The first three use the same trained weights; the fourth is a model trained with frozen cycle-2 image tokens.

Bar heights read from Figure 3 by eye, to about ±1 point; the paper prints no numbers for this figure.

The normal model wins on all four benchmarks. Freezing the image tokens hurts most: on RealWorldQA, accuracy falls from about 72 to about 38. Training the model to live with the freeze helps, but it still stays far below normal: about 58. So the gain of the second cycle is not only "read the same picture again". The picture's own tokens must keep being rewritten.

How much do the image tokens change?

Next, the authors measure the change directly. For each image token, compare its vector before and after one call, and take the length of the difference. Average over tokens and samples:

Uc
The mean size of the update that call c makes to an image token.
1 / N
Average over the N image tokens (then over 32 samples).
yi,c
Image token i after call c.
xi,c
Image token i before call c.

If the second cycle only copied the first, its updates would be close to zero. They are not. Updates in cycle 2 are about as large as in cycle 1: between about 48 and 66 per call in both halves. These are lengths of vector differences, in the units of the model's 1,536-number hidden vectors.

The paper also compares directions with cosine similarity, the cosine of the angle between two vectors: 1 means they point the same way. It finds the image vectors keep turning. A linear probe is a single trained linear map; here it tries to recover the original image features from each state. How well it works varies from call to call, and at the end it works worse than right after the projector. The image tokens drift away from "what the encoder saw" toward something else.

Where does the answer form?

The last probe reads the answer early. Take the state after any layer call. Apply the stack's final normalisation and the shared language head, and you get a guess at the first answer token. This is the logit lens. Compare each early guess with the final one using the KL divergence, a measure of how far one distribution is from another:

Kℓ
How far the guess at layer call ℓ is from the final answer distribution, in nats (natural-log units). Zero means identical.
∑v
A sum over every token v in the vocabulary.
PT(v)
The final probability of token v, after all 128 calls. It is the reference.
Pℓ(v)
The early guess: the probability of v read out after layer call ℓ.

A small case with three answers: the final distribution is (0.90, 0.05, 0.05) and an early guess is (0.30, 0.40, 0.30).

0.90 ln(0.90/0.30) + 0.05 ln(0.05/0.40) + 0.05 ln(0.05/0.30) = 0.99 − 0.10 − 0.09 = 0.80 nats

For scale: a flat guess over 100,000 tokens sits about 11 nats from a confident answer. So a KL of 11 or 12 means the early readout knows almost nothing yet.

Read the answer early

Two views of the same 32 samples. Choose a view, then drag the playhead through the 128 layer calls.

100

KL curve: Figure 11 (mean of 32 samples, first answer token), traced by eye. Update sizes: Figure 10 (mean L2 change per 16-layer call), read by eye. The answer bars under the curve are a toy distribution built so that its KL from the final one follows the traced curve.

The distance stays high, around 9 to 12 nats, for most of the run. It dips a little during each H call. Then, inside the very last H call, it falls to zero within about 16 layer calls. The model forms the answer at the end, in the final H call, and almost nowhere else.

That explains the collapse in Chapter 5. Run only one cycle, and the head reads a state from before the final H call. By this curve, that state has not formed an answer yet. The near-zero scores were not a mystery: the model was asked before it had decided.

Which result shows that the image tokens must keep changing during the second cycle?

Chapter 9

Training it

Five stages, 0.14 trillion tokens, and gradients that skip the first three calls

The authors build LoopVL in five stages. Each stage decides which parts may learn and which stay frozen. The vision encoder never learns: it stays frozen in every stage.

The training ladder

Choose a stage. The picture shows which parts learn (solid) and which are frozen (hatched), and how many tokens the stage uses on a log scale.

Stages, data and settings: Sections 3 and 4 and Figure 4. Token totals as printed; they add up to about 141 billion, the 0.14T of Tables 1 and 2.

Three choices in the ladder are worth a second look.

The loop learns language first. The authors pretrain their own loop on 75 billion text tokens with the open HRM-Text code and data, plus 15 billion tokens of their own. They do not use HRM-Text's released checkpoint. Every later stage starts from that text loop.

Alignment trains only the bridge. In Stage 1 both the loop and the encoder are frozen. Only the projector learns, from 559,000 image-caption pairs. Its job is narrow: map encoder features into a form the frozen loop already understands. Before the main run, the authors swap images in and out and check that the answers change. That confirms the bridge carries image information.

Long answers are dropped, not cut. A few text-heavy samples have answers longer than the sequence budget. Stage 2 removes them instead of truncating them, so the model never learns from an answer that stops halfway.

The loss

Supervised stages use the standard next-token loss, but only on the answer. The image and the question are context; the model is not graded on them. The model factors the answer y = (y1, …, yT) one token at a time, each conditioned on the context x and the answer so far. The loss averages over every valid answer token in the batch:

ℒSFT(θ)
The supervised loss for the trainable weights θ.
∑i ∑t
A sum over the B samples in the batch and over the Ti valid answer tokens of each. Padding is masked out.
pθ(yi,t | …)
The probability the model gives the correct next answer token, given the image, the question and the earlier answer tokens. During training the earlier tokens are the true ones (teacher forcing).
∑i Ti
The total number of answer tokens in the batch. Dividing by it makes this a mean per token, not per sample.

A batch of two: sample 1 has three answer tokens with log-probabilities −0.1, −0.5 and −0.2; sample 2 has one, at −1.2.

−(−0.1 − 0.5 − 0.2 − 1.2) / (3 + 1) = 2.0 / 4 = 0.5 (a per-sample mean would give (0.27 + 1.2) / 2 = 0.73)

The final stage is reinforcement learning with GRPO, group relative policy optimisation. For each image and question, the model writes a group of answers. A correct final answer earns a positive reward and anything else zero. Each answer's advantage is its reward measured against its own group. So GRPO needs no value network, the second model that other RL methods train to predict the expected reward.

During this stage the authors raise the maximum answer length from 2,048 to 4,096 tokens. The model first learns short correct reasoning, then gets room for long derivations. As the length limit rises, the group size shrinks to keep the cost level.

Gradients that skip the start

The forward pass always runs all eight calls. The backward pass does not. During training, gradients flow back only through the last few calls. Early in training it is the last two, L6 and H2. During a warm-up the window grows to the last five: L4, L5, L6, H1 and H2. The first three calls, L1 to L3, never send gradients back.

Where the gradient flows

The eight calls, forward left to right. Move the warm-up slider to grow the backward window. Watch which stored stack still learns.

done

The window, from the last two calls to the last five, is from Section 2.1. The paper gives no reason for it. Saving memory and keeping long gradient paths stable are the usual reasons for such schemes; that part is our reading.

Here is the subtle part. L1, L2 and L3 send no gradients, but the L stack still learns. It shares its weights with L4, L5 and L6, which are inside the window. Every gradient through L4 to L6 updates the same weights that L1 to L3 use. Skipping the early calls saves the memory of storing their activations for the backward pass. It costs no weights that never learn.

A note on evaluation. All headline scores use greedy decoding and a direct answer: an option letter, yes or no, or a short reply, usually within 32 tokens. The model is not asked to reason step by step. The authors report that step-by-step prompting helps a little on maths and hurts on short-answer tasks.

Gradients never flow back through L1, L2 or L3. Why does the L stack still learn?

Chapter 10

What it buys

A 1B loop against eleven compact models trained on 21 to 260 times more data

Chapter 4 compared LoopVL with models trained on the same data. The real world is less fair. Most compact VLMs today are trained on 3 to 36 trillion tokens. LoopVL saw 0.14 trillion. The authors compare it with eleven of them anyway, across 16 benchmark families: general understanding, hallucination, maths and science.

Results bench

Choose a benchmark, or switch to the data view: training tokens against the average of three reasoning benchmarks.

Scores and training tokens: Table 2 (in-house evaluations combined with selected published results). The three-benchmark average (LogicVista, MMMU-Pro, MathVision) is our mean from Table 2, the measure of Figure 1.

LoopVL is strongest on visual reasoning. On LogicVista it scores 45.2, ahead of the next model (InternVL3.5-4B) by 8.3 points. On MMK12-Math it leads with 61.6 against 54.8. It also tops MMMU-Pro (38.7), MMStar (63.5), RealWorldQA (71.0) and MathVision (38.5, a hair over Qwen3.5-2B's 38.4).

It is weaker where reading fine text and charts matters. On ChartQA it scores 74.5 against a best of 86.0. On AI2D, a diagram benchmark, 75.5 against 82.8. On HallusionBench, which tests whether a model invents things, 44.7 against 51.6. On MMEval-Pro, 24.6 against 42.8. A model that saw 0.14 trillion tokens has seen far fewer charts and documents than one that saw 36 trillion.

Switch to the data view. On the three-benchmark average, LoopVL sits at about 40.8 with 0.14T tokens. The next best, InternVL3.5-4B, sits at about 32.4 with 36T. Every other model is to the right of LoopVL by a factor of 21 to 260 in data, and below it.

The limits, in the authors' words

Where is LoopVL weakest against the compact models in Table 2?

Chapter 11

Connections

The cheat sheet, the code, and where LoopVL sits in the field

You can now explain each part of LoopVL, from the loop itself to the place where the answer forms. Let's lock it in.

The one-paragraph summary

LoopVL is a 1B-scale vision-language model whose language backbone is a loop. It stores two 16-layer stacks, L and H. It runs them in the order L L L H L L L H: 128 layer calls from 32 stored layers. A frozen Penguin encoder and a small projector feed one token per image patch, calibrated to word-token size, with 2D positions and a PrefixLM mask.

Trained on 0.14 trillion tokens, it beats its non-looped twin by 6 to 23 points and roughly matches 4B dense models on the same data. Its attention over the image jumps from spread to concentrated at the cycle boundary, the Visual Aha Moment. Its image tokens keep changing in the second cycle, and freezing them hurts. Its answer forms in the final H call.

The numbers that matter

QuantityValueWhy it matters
Stored layers32 (L: 16, H: 16), width 1,536The weights of a 1B model
ScheduleH2L3: L L L H L L L H128 layer calls per forward pass
Image tokens256 to 1,024 per image, one per patchA stable grid for the diagnostics
Loop vs its twinMMStar 55.3 → 63.5, ChartQA 51.1 → 74.5Same 32 layers, same data
Training FLOPs2.47 vs 0.92 (twin), 2.89 and 2.99 (4B), × 1021The loop saves weights, not compute
Equal depth, different orderH2L3 63.5 vs H4L1 24.7 (MMStar, 128 calls each)Order matters, not just count
Off-schedule1 cycle ≈ 0; H3L3 58.2; H4L3 24.0Best only on its trained schedule
Visual Aha MomentGini 0.405 → 0.858 at calls 63 → 64Same weights, new reading of the image
Answer formationKL ≈ 7.5 → 0 nats inside the last H callWhy early readouts fail
Training data75B text + 0.33B + 49B + 16.81B + 0.006B ≈ 0.14T21 to 260 times less than its rivals

One forward pass and one training step, as code

python# LoopVL in pseudo-Python. Frozen: penguin. Trained: projector, L_stack, H_stack, gate, scales, head.
def embed(image, prompt):
    f = penguin(image)                            # N_img x 1024, one per kept patch (256 to 1024)
    v = projector(f)                              # 1024 -> 1536 -> 1536, GELU between
    v = rms_normalize(v) * median_rms(word_emb)   # same size as a word token
    x = build_sequence(header, v, prompt)         # image: 2D RoPE (row, col); text: (t, t)
    return x, v

def forward(image, prompt, H=2, L=3, window=5):
    x, v = embed(image, prompt)
    zH = embed_scale * x                          # the H state starts from the input
    zL = zeros_like(x)                            # the L state starts at zero, kept across cycles
    calls = H * (L + 1); i = 0
    for cycle in range(H):                        # Model-Loop
        zH[img] += scale[cycle] * gate(zH, prompt) * v   # re-inject the visual anchor
        for _ in range(L):                        # Module-Loop
            with grad_if(i >= calls - window): zL = L_stack(combine(zL, zH))   # 16 layers
            i += 1
        with grad_if(i >= calls - window): zH = H_stack(combine(zH, zL))       # 16 layers
        i += 1
    return lm_head(norm(zH))                      # the answer is read from the final H state only

def train_step(batch, step):
    w = 2 + round(3 * warmup(step))                # backward window grows from 2 calls to 5
    logits = forward(batch.image, batch.prompt, window=w)   # PrefixLM mask inside every layer
    return cross_entropy(logits[answer], batch.answer)     # Eq. 3: mean over valid answer tokens

Where it sits in the field

Keep going

The takeaway. A loop gets depth from time instead of weights, but it only earns that depth if its state keeps changing. In a vision-language model, the state includes the picture. LoopVL shows that the picture's tokens keep being rewritten in the second cycle, that the same layer then reads a different part of the image, and that stopping either change costs accuracy.

Now press Present or Teach and explain, out loud and from memory, why the same layer can look at different patches on its second run. If you can, you own this paper. Then go back to the picture and drag the layer call across 63 and 64.

Based on "LoopVL: Recurrent Visual Intelligence" by Zhe Qian, Ziyang Gong, Zhongxing Xu, Hehan Li, Zhonghua Wang, Fei Luo, Mingxuan Wang, Xue Yang, Shiwei Liu, Yanbiao Ma, Junchi Yan and Jungong Han (2026)
Read the paper · Back to Veanors