Adam Hung, Bardienus P. Duisterhof, Deva Ramanan, Jeffrey Ichnowski (Carnegie Mellon University) · arXiv:2609.17524, September 2026

Dream in Structure, Then Act

Most world-action models imagine the future as video, paying for lighting and texture a robot never needed. ModAR imagines it one modality at a time (where things move, what they are, how far away they sit) and only then decides what the arms should do. A 30.1M-parameter model trained from scratch keeps pace with a 6B video-pretrained one (75% vs 72% observed).

Hung et al. 2026 Carnegie Mellon University Robot learning · world-action models arXiv:2609.17524 →
Prerequisites: what a denoising generative model does (flow matching, DiT) + what a robot policy outputs (diffusion policy, action chunking). World-action models, every modality, the mask, the loss and every experiment are built from zero here.
10
Chapters
30.1M
Parameters
75 vs 72
% vs Flex-π
~20×
Fewer training FLOPs

Chapter 0: The Robot That Dreams in Pixels

Two robot arms sit on opposite sides of a table. Between them are two cups, and the job is to stack one inside the other. A camera watches from a fixed spot across the table. Every so often the robot's policy has to answer one concrete question: what should my fourteen joint values be for each of the next sixteen control steps?

(A note on sourcing before we start. Every number, table value, dataset, and design detail in this lesson comes from the paper: Hung, Duisterhof, Ramanan and Ichnowski, Modality-Autoregressive World-Action Models, arXiv:2609.17524. When we use toy numbers to make a mechanism visible, or do our own arithmetic on the paper's numbers, the text says so plainly.)

A fast-growing family of robot models answers that question in two moves. First it imagines: it predicts what the scene will look like a moment from now. Then it acts: it picks the joint commands that make that imagined future happen. The imagining is not decoration. The paper's first sentence calls future-observation prediction "a powerful objective for learning rich representations of the world's dynamics and semantics."

Here is the uncomfortable part. When most of these models imagine, they imagine video. They predict the future as RGB images, usually as compressed image latents produced by a frozen video autoencoder. To do that well, the model has to commit capacity to the exact brown of the table, the glint of the lamp on the cup rim, the soft shadow the forearm throws. In a scene like this, none of those details changes what the arm should do next.

Think about how you would plan the same stack. You would not rehearse the colors. You would think: this cup goes up and over; it is a cup, not a bowl; it is about a hand's width in front of the other one. Motion, identity, distance. That is the whole idea of this paper, stated as a hunch. The rest of the lesson is the machinery and the evidence.

The paper's diagnosis, in its own words. Predicting future RGB "provides a convenient interface to pretrained video generators," but "reconstructing visual appearance does not explicitly prioritize the geometric, physical, and semantic structures most relevant to manipulation tasks." Convenience is why RGB became the default. It is not evidence that RGB is the right thing to imagine.

A model that imagines and acts

Start with an analogy. A climber on a wall looks at the next three holds and, before moving, plays the sequence in their head: hand here, weight shifts there, foot comes up. The mental rehearsal and the movement are one skill, not two. A climber who rehearses badly also climbs badly, and practicing the rehearsal makes the climbing better.

The robot version of that climber has a name. A world-action model (WAM) is a single network that jointly models future observations and the corresponding robot actions. "World" is the part that predicts what the scene will look like; "action" is the part that predicts what the robot should do. The two share one set of weights, so whatever the network learns about how the world moves is available when it chooses how to move.

A WAM is different from two neighbors you may know. A plain policy (for example a diffusion policy) maps the current observation straight to actions and never predicts the future. A plain world model (see the World Models gleam) predicts the future given actions you supply, and leaves the choice of action to a planner. A WAM predicts both, in one model, conditioned on what the robot sees right now.

Why bother imagining? The data argument

The deepest reason WAMs exist is about data, not about elegance. To learn actions, you need action-labeled demonstrations: recordings where every frame comes with the exact joint commands the robot executed. Those come from teleoperation. Someone has to drive the robot, one demonstration at a time. They are slow and expensive to collect.

To learn what the future looks like, you need only video. An actionless demonstration is a recording of the task being done with no action labels attached: a person stacking cups with their own hands, for instance. You cannot train an action head on it, because there are no actions to copy. But you can train a future-prediction head on it, because the future frames are right there in the recording.

The paper names two ways that extra future supervision can help action prediction. First, as a training-time auxiliary objective: the future-prediction loss shapes the shared representation, even if the model never generates the future at deployment (the Fast-WAM line of work). Second, by letting the policy condition its action generation on predicted visual futures (the UniPi and VERA line of work). ModAR is firmly in the second camp: the paper calls it the first WAM to denoise several future modalities one after another before it predicts the actions.

Why this matters for the whole paper. If futures are learnable from actionless data, then a WAM's quality could keep improving as you pour in more actionless data. That makes "how much better do you get with more actionless data?" one of the paper's central measurements. In Chapter 6 you will see ModAR gain 10 percentage points from extra actionless demonstrations while a standard joint-generation WAM gains 1.

Four ways to describe the future

If RGB is not the only way to imagine the future, what are the alternatives? The paper uses the word modality in a deliberately broad sense: "a representation of the future, including both sensory signals and derived features." Depth is a sensory signal. DINO features are a derived feature. Both count as modalities here.

ModAR works with exactly four of them. The paper writes the set as M = {rgb, depth, dino, tracks}. Each one keeps some facts about the scene and throws others away, and the paper argues that each contains "unique inductive biases that capture manipulation-relevant features." An inductive bias is a built-in preference: a representation that makes some patterns easy to express and others impossible.

ModalityWhat one future frame holdsWhat it makes explicit (paper)How ModAR gets it
RGBColor pixelsFull appearance: color, texture, lighting, and everything elsePatchified image
DepthDistance from the camera per pixelScene geometry; direct supervision for the spatial reasoning 3D actions needPatchified depth map (stereo depth on the real robot)
DINOA feature vector per image patchObject semantics and scene structure; robust to appearance variationSpatial patch tokens of a frozen DINOv2 encoder
Point tracksWhere each tracked point moved, and whether it is visibleScene motion and correspondence, independent of appearanceCoTracker3 tracking a grid of query points

Read the third column slowly, because it is the argument in miniature. Point tracks say how task-relevant parts of the scene move through time. A track does not care whether the cup is red or blue. DINO features, the patch embeddings of a self-supervised vision transformer (DINOv2), say "this patch is part of a cup" in a way that survives a change of lighting. Depth says how far away every surface is, which is exactly what a 3D reach needs. RGB says all of that too, but buried under appearance.

The simulation below makes the difference physical. It renders one toy tabletop scene four ways. The toy renderer is ours: the colors in the "DINO" panel are a cartoon of semantic grouping (real DINO features are 384-number vectors, not colors), and the scene is drawn from simple shapes. But the behavior it demonstrates, which panels react to appearance changes and which do not, is exactly the property the paper leans on.

Sim 1 · one future, four modalities

A gripper carries a cup to the right. Scrub the future step to move from now (t) to t+16. Then drag lighting to move the lamp and change its strength, and swap the table texture. Watch the readout: appearance changes rewrite the RGB future but leave the motion (tracks), the geometry (depth), and the cartoon semantics (DINO) essentially untouched. Hollow dots in the tracks panel are points the cup has covered, which a real track marks as not visible.

Future step
Lighting

Two things should stand out after a minute of play.

First, the lighting slider moves a lot of RGB "mass" while the future event (a cup moving right) stays identical. A model trained to predict RGB is graded on that mass. If the lamp in the training videos flickers, or the tablecloth changes between demonstrations, the RGB loss punishes the model for not guessing it, and some of the model's capacity goes to guessing it. The paper's later hypothesis is exactly this: future RGB "introduces high-variance appearance details while adding little information beyond the more structured targets."

Second, the three structured modalities each hold a different part of the event. Tracks alone tell you the cup moved right and slightly up, but not what a cup is. DINO tells you where the cup is, but not how far away. Depth tells you it is closer than the wall, but not which pixels belong together as one object. That complementarity is why the paper asks its real question.

The question this paper asks

Other groups had already tried alternatives to RGB, either in place of it (DINO-WM, EgoWAM) or alongside it (point-track WAMs, ST-WAM, WAM4D). Concurrent work Flex-π found that jointly predicting several future modalities (RGB, 3D pointmaps, and DINO features) can beat predicting future RGB alone. So "use more than RGB" was in the air.

What was not settled is the question the paper states in one line: How should WAMs combine multiple modalities? That breaks into two design axes, and the paper studies both.

Axis 1 · representation
Which futures should the model predict? RGB, depth, DINO, tracks, or some subset? Does each one add something, or do they overlap?
↓ and, independently…
Axis 2 · formulation
How should the predicted futures couple to action prediction? All at once? Independently? Futures first, then actions? And if futures come first, in what order?

ModAR's answer, in one breath

ModAR stands for modality-autoregressive. You probably know autoregressive from language models: generate token 1, then token 2 conditioned on token 1, and so on. ModAR is autoregressive over modalities, not over words or time steps. It fully generates one future modality, then generates the next one conditioned on the finished first, and so on, and it generates the robot's actions last.

Condition
Current camera observation + robot joint configuration + task embedding
↓
Block 1 · point tracks
Denoise how every tracked point will move (motion)
↓ finished tracks become context
Block 2 · DINO features
Denoise the future semantic features (semantics)
↓ tracks + DINO become context
Block 3 · depth
Denoise the future depth maps (geometry)
↓ optionally, block 4 · RGB (appearance)
Last · actions
Denoise the 16-step joint-command chunk, conditioned on the observation and the entire imagined future

Inside each block the model is a denoising generator: it starts from random noise and refines it into a clean prediction, exactly like an image diffusion model refines noise into a picture. Across blocks it is autoregressive. Chapters 4 and 5 unpack both halves.

The "scratchpad" hypothesis. Why go in that order? The paper borrows an idea from image generation (Latent Forcing), where generating a compact semantic latent before the full-detail pixels lets the latent act as a "scratchpad." ModAR's order runs from compact, structured targets that are easier to predict (tracks, DINO) toward high-dimensional, detailed ones (depth, RGB), and ends with actions. Chapter 7 tests it: reversing the order drops average success from 75% to 65%.

Why train everything from scratch

Most existing WAMs start from a pretrained video-generation model. The paper calls that initialization "highly effective," and then points at its cost for science: it "makes it difficult to isolate the effects of WAM formulation, target representations, and actionless-data scale." If a model inherits billions of parameters of video knowledge, you cannot tell whether a design choice helped or whether the pretraining carried it.

So the paper does something unusual. It trains every compared model from scratch, with the same backbone, the same action-labeled data, and the same optimization budget. The only thing that changes between two rows of a results table is the thing being studied. Then, separately, it runs one system-level comparison against a large video-pretrained WAM (Flex-π, 6B parameters), to see whether a small from-scratch model is even in the same league.

What the paper finds, as a preview

Here is the scoreboard you are going to earn, chapter by chapter. Hold these numbers loosely for now; each one gets its full context later.

ClaimEvidence in the paperChapter
Sequential generation beats other WAM formulationsHighest average success at every data scale in simulation: 66%, 75%, 76% at 50, 250, 1,250 demonstrations per task6
It uses actionless data best+10 percentage points (66% to 76%) from 1,200 extra actionless demonstrations per task, versus +1 point for Unified6
Tracks, DINO, depth each help; RGB does not add consistentlyRemoving tracks, DINO, depth drops 75% to 61%, 65%, 70%; removing RGB leaves 75%7
Small from-scratch can match big pretrained75% vs Flex-π's 72%, with about 200× fewer parameters (30.1M vs 6B) and about 20× fewer training FLOPs8
It works on real robots and learns from humans83.3% real-world success vs 66.7% (Unified) and 52.2% (Action-only); human videos lift ModAR from 70.0% to 81.1% to 83.3%8
Read the headline the way the authors do. The paper itself calls the Flex-π result "a slightly higher observed average success rate" and says it is "a system-level comparison rather than a controlled architectural comparison, since the models differ in scale, pretraining, target modalities, and training recipe." The strong claim of the paper lives in the controlled study, where only one factor changes at a time.
Who wrote this, and what it builds on. The authors are Adam Hung, Bardienus P. Duisterhof, Deva Ramanan and Jeffrey Ichnowski at Carnegie Mellon University (project page: adamhung60.github.io/ModAR). The paper cites two earlier 2026 papers by overlapping authors: 3PoinTr (Hung, Duisterhof, Ichnowski), on 3D point tracks from unconstrained human videos, and Modality Forcing (Duisterhof, Ramanan, Ichnowski, with Johnson and Park), on generating depth alongside images. For the idea of ordering the generation, the paper points first to Latent Forcing (Baade et al.) and then to Modality Forcing. Our own observation, not a claim the paper makes: ModAR can be read as joining two threads, tracks as a representation and ordered multimodal generation, inside a world-action model.

The road map

ChapterWhat you will be able to do afterwards
1 · Four CouplingsDraw the Unified, Independent-noise, Disjoint, Action-only and ModAR formulations and say exactly how they differ
2 · Targets & TokensWrite the problem definition and count every token ModAR predicts
3 · The Shared TrunkSketch the DiT with shared blocks, modality experts and adaLN, and budget its parameters
4 · Block-Causal OrderDerive the factorization, build the two-copy training sequence and its mask
5 · Flow & Context NoiseImplement the x-prediction flow loss, the sampler, and context noise
6 · Formulations × ScaleRead Table I and the two control experiments that defend it
7 · Which Futures MatterRead Table II and the ablations, and explain why RGB was dropped
8 · Flex-π & Real RobotsCheck the compute comparison and the real-world results yourself
9 · Limits & Next ReadsState what the paper does not show and where to go next
✎Predict firstWhich single future would you bet on?▸

Before reading on, write down which one modality you think gives the best policy when predicted alone with 50 demonstrations per task: RGB, depth, point tracks, or DINO. Then write down which one you think hurts most to remove from the full set. Chapter 7 grades both guesses, and the two answers turn out to be different modalities.

One control cycle, end to end

Before the details, here is what happens every time the deployed real-robot ModAR is asked what to do, using only numbers the paper states. Keep this picture in mind; every later chapter zooms into one line of it.

1 · Sense
The fixed ZED stereo camera gives an image and stereo depth; the arms report their 14 joint and gripper values; the task ID is known.
↓
2 · Tokenize
The view is cut into a 12×16 grid of 14×14 patches; DINO features come from a frozen DINOv2 encoder; everything is projected to 384-wide tokens.
↓
3 · Imagine, in order
Denoise future tracks at t+8 and t+16 (8 steps), then future DINO features (8 steps), then future depth (8 steps). Each finished block is cached as context.
↓
4 · Act
Denoise a 16-step chunk of 14-D joint targets (8 steps), reading the whole imagined future.
↓
5 · Execute and repeat
Send the 16 targets one per control step; at t+16 the robot looks again and the cycle restarts.

That is 32 network passes per cycle in the real-robot configuration (8 per stream, four streams), by our count. The only latency the paper reports is for the four-future simulation configuration: 147.9 ms for 40 passes on one RTX 5090.

The paper's opening figure, in words

The paper opens with a figure titled "Predict the future one modality at a time, then act." It shows the three real-robot tasks (stack cups, fold towel, place in drawer) as rows, and four numbered columns as the generation order: 1 predicted tracks (motion), 2 predicted DINO (semantics), 3 predicted depth (geometry), and 4 predicted actions, shown on the live RGB camera view.

Look at what is not in that figure: there is no predicted RGB column. The real-robot models never imagine pixels at all. They imagine arrows (where each point goes), a patchwork of feature vectors (what each patch is), and a heat map of distance (how far each patch is), and then they move. By the end of Chapter 7 you will know exactly which experiment justified leaving the pixels out.

The column labels are worth memorizing because they are the paper's whole vocabulary for why each modality exists: tracks carry motion, DINO carries semantics, depth carries geometry. RGB, when present, carries appearance, and appearance is the part the paper found least useful to predict in its setting (Chapter 7 has the numbers).

Three misconceptions to drop before Chapter 1

"A world-action model has to generate video." Most do, because that is how they inherit knowledge from pretrained video generators. But the definition only requires jointly modeling future observations and actions, and the paper's "modality" is any representation of the future. A WAM that never outputs a single pixel is still a WAM.

"Autoregressive means one token at a time." ModAR is autoregressive across four or five blocks, and inside each block all tokens are generated together by denoising. The sequential part is short (at most five stages); the parallel part is wide (384 tokens per visual block).

"More modalities is always better." The paper's own data contradicts this. Adding future RGB to tracks, DINO and depth changes average success by +0.03, 0.00 and −0.01 at the three data scales. Which modalities help is an empirical question, and the answer is task-dependent.

Check yourself before moving on

1. The data argument. In one sentence each, say what an action-labeled demonstration can train that an actionless one cannot, and what an actionless one can still train. Check: actionless data cannot supervise the action chunk; it can still supervise every future modality it has targets for.

2. The two axes. Name the two design axes the paper studies and give one example choice on each. Check: representation (which futures: e.g. tracks versus RGB) and formulation (how futures couple to actions: e.g. joint versus sequential).

3. A prediction to grade later. Write down whether you expect a sequential WAM or a joint WAM to benefit more from extra actionless data, and why. Chapter 6 answers with a number.

According to the paper, what is the core problem with using future RGB as the thing a world-action model imagines?

Chapter 1: Four Ways to Couple Imagination and Action

Imagine two WAMs that predict exactly the same futures, from exactly the same data, with exactly the same network. One reaches 75% average success in simulation. The other reaches 34%. Nothing about what they predict differs. The only difference is when each piece gets generated and which pieces are allowed to look at which.

That is not a hypothetical. It is one column pair of the paper's Table I, at 250 demonstrations per task (ModAR at 75%, Independent-noise at 34%). This chapter builds the vocabulary to see why such a small-sounding choice can be worth forty points.

A one-paragraph primer on denoising generation

Every model in this chapter generates its outputs by denoising. Picture a sculptor who starts with a block of static and, in a handful of passes, chips it into a statue. Each pass looks at the current rough block and predicts the finished statue, and then moves the block a little toward that prediction. After enough passes, the block is the statue.

The formal name for the version ModAR uses is flow matching (Chapter 5 builds it from zero). The progress of one pass is tracked by a number the paper calls the flow timestep τ. In the paper's convention, τ = 0 means pure noise and τ = 1 means the clean result. Generation walks τ from 0 to 1 in a fixed number of steps. ModAR uses 8 Euler steps per generated stream.

One more word you will see constantly: a stream is one set of tokens being generated together. The future point tracks are one stream. The future DINO features are another. The action chunk is another. A WAM that predicts four future modalities plus actions has five streams to generate. The formulations below are five different answers to "how do those five streams relate while they are being generated?"

Formulation (b): Unified, or joint generation

The most common answer is to generate everything at once. All streams start as noise together, share one flow timestep, and are refined in lockstep. At every refinement step, every stream can attend to every other stream, so the action tokens can read the half-finished future tokens and vice versa.

Analogy: a comic artist sketching all six panels of a page simultaneously, glancing across the page while each panel is still a rough pencil sketch. The paper labels this Unified and uses it to represent joint-generation WAMs such as DreamZero and Cosmos Policy, which "denoise future-observation and action streams together" with cross-stream attention. In the paper's controlled implementation, Unified uses one shared flow timestep drawn from a logit-normal distribution with (μ, σ) = (−1, 1).

Notice the subtle consequence. At step 3 of 8, the action stream is being refined while looking at futures that are themselves only three-eighths clean. The action prediction conditions on partially noisy predicted futures. The paper returns to this asymmetry with a clever control experiment in Chapter 6.

Independent-noise: same inference, different training

A close cousin of Unified changes only the training procedure. During training, instead of drawing one shared τ for all streams, it draws a separate τ for each stream, independently. One training example might have the tracks at τ = 0.9 (almost clean) while the actions sit at τ = 0.2 (mostly noise). At inference it does exactly what Unified does: all streams denoised simultaneously on one shared schedule.

The paper attributes this design to Unified World Models and Flex-π. One way to see its appeal (this is our gloss, not a claim from the paper): a stream sampled near τ = 1 behaves almost like clean conditioning for the others, so a single network gets exposed to many different "who conditions on whom" situations during training.

The paper's from-scratch experiments, however, find Independent-noise "performs poorly." Its explanation is a train/test mismatch. At test time, all streams are always at the same noise level. During training, independently sampled noise levels "rarely match the synchronized test-time denoising schedule," so the model "receives insufficient training signal near its inference regime." It spends most of its practice on situations it will never face.

Formulation (c): Disjoint, as in Fast-WAM

The third answer cuts the connection. Disjoint predicts each target independently, "without attention between future-observation and action targets." The futures and the actions share the backbone that reads the current observation, but the action tokens never look at the predicted future tokens, and the future tokens never look at the action tokens.

Why would anyone want that? Because then the future prediction is purely a training-time auxiliary objective. At deployment you can skip generating the future entirely and save its cost. This represents Fast-WAM, whose paper title asks the question directly: do world action models need test-time future imagination? The future-prediction loss still shapes the shared representation during training. It just never feeds the action at run time.

Formulation (d): Action-only

The baseline predicts actions directly from observations and predicts no future at all. The paper describes it as "a typical flow-matching behavior cloning policy." Because it has no future-prediction objective, it has nothing to learn from actionless demonstrations, so it trains only on action-labeled data. That is why, in the results tables, Action-only has a number only at the smallest data scale: adding actionless data changes nothing for it.

Formulation (a): ModAR, futures first and one at a time

A separate lineage (UniPi and VERA) already reversed the coupling: they "fully denoise visual futures and then infer actions from the resulting frames, rather than co-denoising both streams." ModAR builds on that futures-then-actions ordering and extends it across modalities. It denoises future modalities one at a time, each fully, each conditioned on the ones already finished, and denoises actions last.

That gives ModAR's final step a clean interpretation. An inverse dynamics model (IDM) answers a detective's question: given what the world looks like now and what it will look like next, which action caused the change? If you see a drawer closed and then open, the answer is "someone pulled it." The paper notes that ModAR's final action-prediction step "acts as an inverse dynamics model," mapping the fully generated future to the actions that would produce it.

FormulationTraining noiseInference orderActions see futures?Uses actionless data?Represents
ModAROwn τ per stream, plus noisy context copiesTracks → DINO → depth → RGB → actions, 8 steps eachYes, fully clean onesYesThis paper
UnifiedOne shared τAll streams together, 8 stepsYes, partially noisy onesYesDreamZero, Cosmos Policy
Independent-noiseIndependent τ per streamSame as UnifiedYes, partially noisy onesYesUnified World Models, Flex-π
DisjointNot specified (targets predicted independently)Futures skipped at inference; actions onlyNo attention between themYes (as auxiliary loss)Fast-WAM
Action-onlyActions onlyActions only, 8 stepsNo futures existNoFlow-matching behavior cloning

The simulation below animates the table. Each lane is one stream. A lane fills from left (noise) to right (clean) as it is denoised. The arrows show which streams the currently active stream can read. Switch to the training view to see how each formulation samples noise levels during training, and resample a few times to feel why Independent-noise so rarely practices the synchronized situation it meets at test time.

Sim 2 · the formulation explorer

Pick a formulation. Inference view: the clock runs through every Euler step; ModAR takes 40 (8 per stream), the others 8 in total. Drag the clock to scrub. Training view: each dot is one stream's sampled noise level τ for one training example. Press Resample repeatedly. The shaded band marks "all streams within 0.1 of each other," which is what inference always looks like for the simultaneous formulations.

Inference clock

Spend a moment on the training view for Independent-noise. With five streams each drawing its own τ, the chance that all five land close together is small. Chapter 6 puts a toy number on "small." The point to carry forward: the inference procedure of Unified and Independent-noise is identical, so any gap between them in the results is caused purely by how their training noise was sampled.

What stays fixed across all five

A comparison is only as good as its controls. The paper implements every formulation "within one controlled implementation": all methods "share the same backbone, action-labeled data, and optimization budget." The same six simulation tasks, the same demonstration counts, the same 1.2M optimizer steps. Only the coupling changes.

An asymmetry you should be suspicious of. ModAR takes 8 Euler steps per stream, so 40 steps across four futures and actions. Unified, Independent-noise and Disjoint take 8 steps in total. A skeptic would say ModAR wins because it thinks five times longer. The paper anticipates this and re-runs the three simultaneous baselines at 40 steps. Chapter 6 shows the gap does not close (Unified actually drops from 67% to 59%).
A second asymmetry, and the second control. ModAR's action step reads fully clean futures; Unified's reads partially noisy ones. Maybe ModAR's futures are no better, and it only wins because its action head gets cleaner inputs. The paper tests this by throwing away both models' native action predictions and feeding their predicted futures to the same separately trained IDM. ModAR still wins at every data scale (Chapter 6 has the numbers).

Why ordering could matter at all

It is worth pausing on why generating streams in sequence might help, before the evidence arrives. When you generate everything jointly, every stream starts from noise with nothing solid to lean on. When you generate in sequence, the second stream starts with a finished first stream in hand.

The paper cites two image-generation results that motivate this. Latent Forcing generates image latents before pixels, "allowing the latents to serve as a semantic scratchpad for generating fine-grained appearance." Modality Forcing adapts a pretrained image generator to jointly generate depth and finds that image models contain priors that improve depth accuracy. The paper's summary of the lesson: "generating easier-to-predict modalities earlier can expose structure that simplifies subsequent predictions."

There is also a plain probabilistic reason, which Chapter 4 makes exact. Any joint distribution over several variables can be written as a product of conditionals, in any order, with no loss. So sequential generation does not restrict what the model can represent. It only changes what each generation step has to figure out on its own, and that is precisely the thing the order controls.

⇆Cross-domain bridge
You already do this when you draw
Art instruction almost universally says: gesture lines first, then block in the big shapes, then values, then color and texture. Nobody starts a portrait by rendering the eyelashes. The gesture is compact and easy to get right, and once it is on the page every later stage has something solid to agree with. ModAR's order (motion, then semantics, then geometry, then appearance, then act) is the same studio habit, applied to a robot's imagination.

One decision, traced through two samplers

Abstract tables hide the thing that matters, so let us trace one decision. The robot is at decision time t. It must output 16 actions. Both models below predict the same futures (tracks, DINO, depth, RGB) plus actions, and both use 8 Euler steps per unit of work.

In Unified, one shared clock ticks 8 times. At tick k (k = 0 through 7) every stream sits at τ = k/8 and is moved to τ = (k+1)/8. (We assume evenly spaced steps throughout this lesson; the paper specifies eight Euler steps but not their spacing.) So when the action stream takes its first step, the futures it reads are pure noise (τ = 0). When it takes its last step, the futures it reads are at τ = 7/8 = 0.875, still carrying one eighth of their noise.

In ModAR, the clock ticks 8 times for tracks alone, then the finished tracks are frozen and re-embedded as context. Then 8 ticks for DINO, reading clean tracks the whole time. Then depth, then RGB. Only after 32 ticks does the action stream start its own 8 ticks, and on every one of them it reads four finished futures.

MomentUnified: futures seen by the action streamModAR: futures seen by the action stream
First action stepAll four at τ = 0.000 (pure noise)All four at τ = 1 (finished)
Middle action step (k = 4)All four at τ = 4/8 = 0.500All four at τ = 1
Last action step (k = 7)All four at τ = 7/8 = 0.875All four at τ = 1
Total network passes88 × 5 = 40

Here are the two sampling loops side by side, as pseudo-PyTorch. It is a sketch of the control flow the paper describes, not the authors' code. The key line is the one that decides what the network is allowed to read.

# ---- Unified: one shared clock, every stream refined together ----
streams = {m: torch.randn(shape[m]) for m in ["tracks", "dino", "depth", "rgb", "act"]}
for k in range(8):
    tau = k / 8
    x_hat = net(obs, streams, tau={m: tau for m in streams})   # joint attention, all noisy
    for m in streams:
        v = (x_hat[m] - streams[m]) / (1 - tau)
        streams[m] = streams[m] + v / 8

# ---- ModAR: one stream at a time, finished streams become context ----
context = {}
for m in ["tracks", "dino", "depth", "rgb", "act"]:
    y = torch.randn(shape[m])
    for k in range(8):
        tau = k / 8
        x_hat = net(obs, context, target=(m, y), tau=tau)  # reads only FINISHED blocks
        y = y + (x_hat - y) / (1 - tau) / 8
    context[m] = y          # re-embedded as clean context for every later block
actions = context["act"]

The ModAR loop is the paper's inference paragraph turned into code: "We generate one modality at a time... The resulting clean tokens become context for generating the next modality, and we condition the final action chunk on the complete generated future." The paper adds one engineering detail the sketch leaves out: keys and values for the current observation and all previously generated modalities are cached, so they are not recomputed at every denoising step (Chapter 4 explains why that is safe).

What each formulation does with a human video

Return to the data argument from Chapter 0. A human video of cup stacking has future frames but no robot actions. Each formulation handles that example differently, and the difference is the reason data scaling behaves so differently across them in Chapter 6.

FormulationWhat an actionless example trainsHow that reaches the action at run time
ModAREvery future block it has targets for; action prediction is omitted for that exampleDirectly: the action step reads every finished future, alongside the observation, configuration and task
Unified / Independent-noiseThe future streams; the action loss is maskedThrough joint attention, but reading partially noisy futures
DisjointThe future streams, as an auxiliary lossOnly indirectly, through the shared representation; futures are not generated at run time
Action-onlyNothing; it cannot use the exampleNot at all

The paper states the ModAR rule plainly: "For actionless examples, we omit action prediction and supervise only the available future targets." And the loss has an explicit switch for it: the action term is included "only for examples with action labels."

What happens when an imagined future is wrong

Every formulation has a failure mode, and they are different.

In Disjoint, a wrong future cannot hurt the action at run time, because the future is never generated at run time. The flip side: a right future cannot help it either. The paper also sees a subtler cost: as more actionless data is added, Disjoint's performance falls, and the authors suggest the shared representation becomes "increasingly shaped by the future-observation prediction objective, causing negative transfer to the action prediction objective."

In Unified, streams are refined together, so no stream ever commits early. That limits how much any single early mistake can dominate. But it also means the action never gets to read a committed future.

In ModAR, commitment is the point, and so is the danger. If the tracks come out wrong, DINO is generated to agree with wrong tracks, depth with both, and the action with all three. The paper names this directly: modality-autoregressive generation "is susceptible to cascading error, as imperfections in earlier generations become inputs to all later predictions." Its fix, context noise, gets a full treatment in Chapter 5, and removing it is the single most damaging ablation in Chapter 7 after removing tracks.

The cost ModAR pays and admits. Sequential generation is slower. The paper's limitations section says so: it "increases inference latency relative to simultaneous or action-only generation." The measured number is 147.9 ms end to end (6.76 Hz) on a single RTX 5090 when generating all four future modalities and actions. Chapter 8 unpacks it.

The formulations as answers to two questions

A compact way to hold all five in your head is to ask each of them two questions: are futures generated before actions or alongside them? and does the action ever read a predicted future?

 Action reads predicted futuresAction never reads predicted futures
Futures alongside actionsUnified, Independent-noise (reads them while still noisy)Disjoint (futures are a training-time loss only)
Futures before actionsModAR (reads them finished); UniPi and VERA (one visual stream)(no sensible design lives here)
No futures at all—Action-only

Read the table as a map of what each design bets on. Disjoint bets that futures are useful only as a teacher for the shared representation. Unified bets that futures and actions should be negotiated together. ModAR bets that the future should be settled first, in a particular order, and that the action should then be chosen with that settled future in view.

One way to test that you really own the map: pick any cell and say what the model would do with a human video of cup stacking. Unified and Independent-noise learn the future streams from it and hope joint attention carries the lesson to the actions. Disjoint learns the future streams from it and then never generates them at run time. ModAR learns its future blocks from it, and at run time its action step reads blocks of exactly that kind. Action-only cannot use the clip at all.

✎CheckClassify three hypothetical designs▸

(a) A model that denoises depth and actions together, then throws the depth away at run time and never lets actions attend to it. (b) A model that samples one noise level for all streams during training and denoises everything together at test time. (c) A model that fully generates DINO features, then fully generates actions conditioned on them.

Answers: (a) is Disjoint in spirit (future prediction as an auxiliary loss, no future-to-action attention, futures omitted at deployment). (b) is Unified. (c) is a one-modality ModAR (K = 1 in Equation (1)), and it is exactly the "DINO" column of Table II.

What is the only difference between the paper's Unified and Independent-noise formulations?

Chapter 2: Targets, Horizons, and Tokens

Before any architecture, pin down the contract. At one decision moment, what exactly goes into ModAR, and what exactly comes out? If you cannot write both down with shapes, you cannot build it. This chapter writes both down, using the paper's own notation and its own implementation numbers.

What the model is given

The paper's goal is "to jointly model future multimodal observations and robot actions conditioned on the current observation, robot configuration, and task label." Each of those three conditioning pieces has a symbol.

ct = ( ot , qt , g )

Here t is the decision time, the control step at which the robot stops and plans. ot is the multimodal observation at that moment: what the single camera sees, in the modalities the model uses. qt is the robot's proprioceptive configuration, its sense of its own body: "joint positions and gripper opening." g is "the learned embedding associated with a discrete task label," a trainable vector looked up from a task ID such as stack bowls.

Two details are worth underlining now, because they come back as limitations. There is one camera, not several: models "condition on a single-camera 168×224 observation." And the task is a discrete label, not a sentence: this is not a language-conditioned model. The paper lists that second point among its limitations.

What the model must predict: futures

For every modality m in the set M = {rgb, depth, dino, tracks}, the model predicts a short list of future snapshots.

Ytm = ( ymt+Δ , … , ymt+JΔ ) ,    H = JΔ

Read it symbol by symbol. ymt+Δ is one future frame of modality m, taken Δ control steps after the decision. Δ is the dynamics stride: how far apart, in control steps, consecutive predicted frames sit. J is how many future frames are predicted. H is the prediction horizon, the last predicted moment, and it equals J times Δ by construction.

What the model must predict: actions

Actions are predicted differently. Not every Δ steps, but at every single control step up to the horizon.

At = ( at , at+1 , … , at+H−1 )

at is the command sent to the robot at step t. At is a block of H consecutive commands, called an action chunk: rather than choosing one command, looking again, and choosing the next, the policy commits to a whole short sequence at once (the idea popularized by ACT). The paper spells out how the two timelines line up: "the final prediction target corresponds to the next decision point t+H, while At contains the H actions executed between t and t+H."

In other words, the model imagines where the world will be at the moment it next gets to think, and it chooses the commands that fill the gap.

Sparse futures, dense actions. The futures are predicted only at a few moments; the actions at every step. The paper states the settings but does not explain the asymmetry. A likely reason (ours) is budget: a future frame is hundreds of tokens (you will count them in a moment), while an action is 14 numbers. Predicting a future at every one of 16 steps would multiply the most expensive part of the model by eight, for frames that probably differ little from their neighbors.

Worked example 1: the timeline, with the paper's numbers

The implementation section gives the values: "horizon H = 16 and dynamics stride Δ = 8, yielding J = 2 sparse targets at t+8 and t+16, together with a dense H-step action chunk; policies replan every H steps." Let us check every piece of that sentence, with nothing skipped.

Step 1. Horizon from stride and count: H = J × Δ = 2 × 8 = 16.
Step 2. First future target: t + 1×Δ = t + 1×8 = t + 8.
Step 3. Second (last) future target: t + 2×Δ = t + 2×8 = t + 16 = t + H. (matches "the final prediction target corresponds to the next decision point")
Step 4. Action indices: at through at+H−1 = at through at+15. Count = 15 − 0 + 1 = 16 actions.
Step 5. Each action is a 14-dimensional absolute dual-arm joint configuration, so one chunk holds 16 × 14 = 224 numbers.
Step 6. Concrete decision at t = 32: futures at 32 + 8 = 40 and 32 + 16 = 48; actions a32 … a47; next decision at 32 + 16 = 48.
Step 7. Replanning every H steps means decisions at t = 0, 16, 32, 48, … so over 160 control steps the model is queried 160 / 16 = 10 times.

Step 7 is our arithmetic, but it matters later: ModAR's measured latency (147.9 ms per query) is paid once per chunk, not once per control step. The paper does not state the control frequency, so we will not convert that into a duty cycle.

One more fact about the actions: they are absolute joint configurations, not changes. Each at says "put the joints here," which is also the form of qt. So the conditioning and the targets live in the same 14-dimensional space.

Turning each modality into tokens

A transformer consumes a sequence of vectors, called tokens, all of the same width. Four very different kinds of data have to become that one kind of thing. The paper's "Modality tokenization" paragraph does it in five moves.

RGB and depth
"We patchify RGB and depth maps." Cut the image into a 12×16 grid of 14×14 patches and flatten each patch into one vector.
↓
DINO
"We extract DINO tokens from the spatial patch tokens of a frozen DINOv2 encoder" (ViT-S/14). One feature vector per patch.
↓
Point tracks
Initialize a 2D grid of queries at patch centers, track them with CoTracker3, and make one token per point-time pair encoding "its displacement from its initial position and its visibility."
↓
Common width
"We linearly project each modality into the common token dimension."
↓
Identity and position
Add a learned modality embedding, and apply axial rotary position embeddings over time and the two spatial dimensions for visual tokens, and over time for action tokens.

Two terms need definitions. To patchify is to chop an image into non-overlapping squares and treat each square as one token, the same trick a Vision Transformer uses. A point track is the path of one physical point through a video: where it is in each frame, and whether it is visible or hidden behind something. CoTracker3 is the tracker that produces those paths here.

Notice the neat alignment. The paper puts RGB, depth and the track queries on one 12×16 grid, with the track queries at the patch centers. DINOv2 ViT-S/14 also works in 14×14 patches. If it is run on the same 168×224 view (our inference; the paper does not state DINO's input size), it too yields a 12×16 grid. In that case all four visual modalities describe the same 192 locations, and token (row 5, column 9) means the same patch of table in every modality. The token counts below assume that alignment.

Worked example 2: counting every token

The paper gives the grid and the horizon. The counts below are our arithmetic from those numbers. The one place where we have to read between the lines is the exact layout of a track token; the paper says it encodes displacement and visibility, and we read that as two displacement numbers plus one visibility number.

Step 1. Patch rows: 168 / 14 = 12. Patch columns: 224 / 14 = 16. Patches per frame: 12 × 16 = 192.
Step 2. Raw numbers in one RGB patch: 14 × 14 × 3 channels = 196 × 3 = 588.
Step 3. Raw numbers in one depth patch: 14 × 14 × 1 channel = 196.
Step 4. One DINOv2 ViT-S/14 patch token has the ViT-S width: 384 numbers.
Step 5. One track token: displacement (dx, dy) + visibility = 2 + 1 = 3 numbers (our reading).
Step 6. Future tokens per visual modality: 192 patches × J = 192 × 2 = 384 tokens.
Step 7. All four future modalities: 4 × 384 = 1,536 tokens. Without RGB (the real-robot setting): 3 × 384 = 1,152 tokens.
Step 8. Action tokens: one per step, H = 16 tokens, each starting as 14 numbers.
Step 9. Raw size of the predicted future, per modality: RGB 384 × 588 = 225,792; depth 384 × 196 = 75,264; DINO 384 × 384 = 147,456; tracks 384 × 3 = 1,152; actions 16 × 14 = 224.

Step 9 is the quiet argument of the whole paper, in numbers. Future RGB is 225,792 raw values to get right, about 196 times the size of the track future (225,792 / 1,152 = 196). The tracks, which Chapter 7 will show are the most costly modality to remove, are the smallest target of all. Size and usefulness are not the same thing.

Once projected, every token has the same width, 384 (the shared transformer width you will meet in Chapter 3). If each projection is a single linear layer with a bias, as "we linearly project" suggests, then the RGB projection alone holds 588 × 384 + 384 = 225,792 + 384 = 226,176 parameters, and the track projection only 3 × 384 + 384 = 1,152 + 384 = 1,536.

Sim 3 · the horizon and the token budget

Top: the timeline of one decision. Warm diamonds are future targets (every Δ steps up to H), green dots are the dense action chunk. Bottom left: the 12×16 patch grid; tap a patch to see its axial position ids. Bottom right: tokens per stream. Move Δ and J, toggle modalities, or load a preset. The Flex-π preset only sets the timing the paper used for that baseline (targets at t+4, 8, 12, 16); Flex-π tokenizes differently, so the bars stay ModAR-style counts.

Stride Δ
Targets J

Try the Flex-π timing and watch the visual token budget double while the action chunk stays at 16. Halving the stride doubles J, and every visual stream grows with J. This is one reason the future side of a WAM, not the action side, dominates its compute.

Position: axial rotary embeddings

A transformer is blind to order unless you tell it where each token sits. ModAR tells it with rotary position embeddings (RoPE): before attention compares a query with a key, both are rotated by an angle proportional to their position. Because the rotations compose, the similarity between two tokens ends up depending on how far apart they are, not on where they both are. (The RoFormer lesson derives this in full.)

Here is the smallest possible example, with illustrative numbers. Take a 2-number slice of a query, (1, 0), and a frequency of 0.5 radians per time step. A token at time 1 is rotated by 1 × 0.5 = 0.5 rad. A token at time 2 is rotated by 2 × 0.5 = 1.0 rad. Their dot product is cos(1.0 − 0.5) = cos(0.5) = 0.8776. Shift both to times 5 and 6 and the angles become 2.5 and 3.0 rad, but the dot product is still cos(3.0 − 2.5) = cos(0.5) = 0.8776. Only the gap survives.

Axial RoPE applies that idea along several axes at once by splitting each attention head's dimensions into groups, one group per axis. For ModAR's visual tokens the axes are time, row and column, so a token at (future frame 2, row 5, column 9) is rotated by three independent sets of angles. Action tokens only have a time axis. The paper does not say how the dimensions are divided among the axes; the idea is what matters here.

RoPE says where a token is. It does not say what kind of token it is: a DINO token and a depth token at the same (time, row, column) get identical rotations. That is the job of the learned modality embedding, a trainable vector per modality added to every token of that modality, like a colored tag on each card in the deck.

Frozen versus computed offline versus trained

Three different things happen to the three "helper" networks, and it is worth being precise.

ComponentStatusWhy that choice makes sense
DINOv2 ViT-S/14FrozenIt defines the DINO target space. If it trained along with ModAR, the targets would move while the model chased them (our reasoning; the paper states only that the encoder is frozen)
CoTracker3Used to produce track targets from the demonstration videosPoint tracking is a hard problem with its own training recipe; ModAR only needs its outputs as labels
Projections, embeddings, DiT, headsTrained from scratchThe paper's controlled study requires "no pretraining" for every compared model

On the real robot, a fixed third-person ZED stereo camera "provides RGB observations and stereo depth for both robot and actionless human demonstrations." So depth there is measured by stereo matching rather than rendered. A stereo estimate has its own errors, and whatever errors it has become part of the depth targets; the paper does not quantify them.

What the tokens do when the world misbehaves

Each representation degrades differently when the input gets hard, and the paper's design already encodes one of those cases.

Occlusion. When a gripper passes in front of a tracked point, the point's position becomes unknowable from the image. That is exactly why the track token carries a visibility value alongside its displacement: the model can say "this point went behind something" instead of being forced to invent a position.

Lighting and texture changes. These rewrite the RGB future, as Sim 1 showed, while the paper describes DINO features as "robust to appearance variation" and tracks as encoding motion "independently of appearance."

A single viewpoint. With one camera, depth is the only modality that carries distance explicitly. That is the paper's stated reason depth helps: it "makes scene geometry explicit and provides direct supervision for the spatial reasoning required for predicting 3D actions."

The tokenizer, as code

Here is the whole tokenization for one training example as pseudo-PyTorch, with every shape annotated. The shapes follow the paper's implementation numbers; the function names, and the 3-number track layout, are ours.

# One example, batch B. Shapes from the paper: 168x224 image, 12x16 grid of 14x14 patches,
# J = 2 future frames (t+8, t+16), H = 16 actions, 14-D joint configurations.
obs_rgb    # (B, 3, 168, 224)          current camera image
obs_depth  # (B, 1, 168, 224)          current depth
fut_rgb    # (B, 2, 3, 168, 224)       frames at t+8 and t+16
fut_depth  # (B, 2, 1, 168, 224)
q          # (B, 14)                   joint positions + gripper opening, both arms
task_id    # (B,)                      discrete task label
actions    # (B, 16, 14)               absolute joint targets; absent for actionless data

def patchify(x, p=14):                       # (B, T, C, 168, 224) -> (B, T, 12, 16, C*p*p)
    B, T, C, Hh, Ww = x.shape
    x = x.reshape(B, T, C, Hh // p, p, Ww // p, p)
    return x.permute(0, 1, 3, 5, 2, 4, 6).reshape(B, T, Hh // p, Ww // p, C * p * p)

dino = load_dinov2_vits14().eval().requires_grad_(False)   # FROZEN target extractor
with torch.no_grad():
    fut_dino = dino.patch_tokens(fut_rgb.flatten(0, 1)).view(B, 2, 12, 16, 384)

# Tracks: CoTracker3 follows the 192 patch-centre queries through the demo video.
disp, vis = cotracker3_targets(video, queries=patch_centres(12, 16), times=[8, 16])
fut_tracks = torch.cat([disp, vis], dim=-1)          # (B, 2, 12, 16, 3): dx, dy, visible

D = 384                                                  # common token width
proj = nn.ModuleDict({"rgb": nn.Linear(588, D), "depth": nn.Linear(196, D),
                      "dino": nn.Linear(384, D), "tracks": nn.Linear(3, D),
                      "act": nn.Linear(14, D)})
mod_emb = nn.Embedding(5, D)                             # learned "which modality" tag

def visual_tokens(name, grid):                         # grid: (B, 2, 12, 16, c)
    tok = proj[name](grid) + mod_emb(ID[name])          # (B, 2, 12, 16, 384)
    pos = axial_ids(t=2, rows=12, cols=16)             # (384, 3): (time, row, col) for RoPE
    return tok.flatten(1, 3), pos                          # (B, 384, 384)

tok_rgb, pos_rgb = visual_tokens("rgb", patchify(fut_rgb))
tok_act = proj["act"](actions) + mod_emb(ID["act"])       # (B, 16, 384)
pos_act = torch.arange(16)                                 # time-only RoPE for actions
What you can now write down without looking. Input: one 168×224 view, a 14-number configuration, a task ID. Output: 384 tokens for each future visual modality (192 patches at two future times) and a 16-step chunk of 14-number joint targets. Everything in between is Chapters 3 through 5.

A single track token, made concrete

Take the patch at row 5, column 9. Its pixel square spans y = 70 to 83 and x = 126 to 139, so its center, where the track query sits, is at

ycenter = 5 × 14 + 7 = 70 + 7 = 77,   xcenter = 9 × 14 + 7 = 126 + 7 = 133.

Now suppose (illustrative numbers) CoTracker3 reports that the physical point under that query is at x = 151, y = 70 at t+8 and is visible, and at t+16 it has passed behind the gripper. The two track tokens for that query would then carry:

t+8:  dx = 151 − 133 = 18, dy = 70 − 77 = −7, visible = 1
t+16: displacement as reported by the tracker, visible = 0

Two honest gaps: the paper does not say whether displacements are normalized (for example divided by the image size) and does not say what displacement value is stored for an occluded point. A reimplementation has to choose. What the paper does fix is the meaning: displacement "from its initial position" plus "its visibility."

What one training example contains, by data source

FieldRobot demonstrationActionless demonstration (simulated or human)
Observation otYesYes
Future targets YtmYes, all predicted modalitiesYes, "the available future targets"
Action chunk AtYes (16 × 14)No: the action term is dropped from the loss
Configuration qtYes (14 numbers)In simulation the actionless demos are robot demos with their actions withheld, so qt presumably still exists; for human video the paper does not say what is used
Task label gYesPresumably the matching task's label; the EgoDex categories are chosen per task ("stack/unstack cups" for cup stacking, and so on), but the paper does not spell out the labeling
🔍Debug itTracks that describe the wrong object▸

A reimplementation initializes its track queries at the top-left corner of each 14×14 patch instead of the center. Training runs, but the tracks stream is noticeably harder to learn near object edges. Why, using only this chapter?

Answer (our reasoning): the design relies on every modality describing the same 192 locations. A corner query sits on the boundary between four patches, so near an object edge it often tracks a point that belongs to a neighboring patch's object. The track token for patch (r, c) then describes motion that the DINO and depth tokens for (r, c) do not show, and the cross-modal agreement the shared trunk is supposed to exploit is broken. Center queries keep the four modalities aligned.

With the paper's settings (H = 16, Δ = 8), which statement is exactly right?

Chapter 3: One Trunk, Five Specialists

You now have five kinds of tokens: tracks, DINO, depth, RGB and actions, all 384 numbers wide. They need to talk to each other, because the whole point is that tracks inform DINO, DINO informs depth, and all of them inform the action. But they also need room to be themselves, because a depth patch and a joint command are very different objects. ModAR's backbone is designed around exactly that tension.

The newsroom analogy

Picture a newsroom putting out tomorrow's paper. First, everyone sits in one editorial meeting: the sports desk hears what politics is running, the photo desk hears what the lead story needs. Then each desk goes back to its own room and writes its own section, in its own style, without the other desks talking over it. Finally each section goes to its own layout template.

ModAR's network is that newsroom. The shared meeting is a stack of transformer blocks every stream attends through. The desks are small modality-specific expert stacks, where attention is restricted to one stream. The layout templates are per-modality linear output heads. The paper's one-sentence summary: "cross-modal information is fused in the shared blocks before within-modality specialization in the expert blocks."

The diffusion transformer, briefly

The backbone is a diffusion transformer (DiT), the architecture Peebles and Xie introduced for image diffusion (the DiT gleam builds it from zero). A DiT is an ordinary transformer that takes noisy tokens in and predicts something about their clean version, with the noise level injected into every layer. Nothing in it is specific to images, which is why it transfers so easily to tracks, depth and actions.

Each transformer block does two things, each wrapped in a normalization and a residual connection. Self-attention lets every token gather information from the tokens it is allowed to see: each token emits a query, every visible token offers a key and a value, and the token receives a weighted mix of values, weighted by how well its query matches each key. Then an MLP (a small two-layer network) transforms each token on its own.

The paper's exact backbone

The implementation section lists it in one sentence, which we unpack line by line below. "The shared DiT contains six width-384 transformer blocks with six attention heads (ViT-S), followed by two width-384 expert blocks for DINO, depth, and RGB, one width-128 block for tracks, and two width-128 blocks for actions."

StageBlocksWidthAttention reach
Shared trunk6384, 6 heads (ViT-S sizing)All causally available modality tokens
DINO expert2384DINO stream only
Depth expert2384Depth stream only
RGB expert2384RGB stream only
Tracks expert1128Tracks stream only
Action expert2128Action stream only
Output heads1 linear layer eachto the modality's raw sizen/a

We read "two width-384 expert blocks for DINO, depth, and RGB" as two blocks per modality, since the paper describes the expert stacks as modality-specific. With six heads on a width of 384, each head works in 384 / 6 = 64 dimensions.

Why are the tracks and action experts narrower? The paper does not say, but the token sizes from Chapter 2 suggest a reason: a track token starts life as about three numbers and an action token as fourteen. A 384-wide specialist for a 3-number target would be mostly empty space. Narrow experts spend parameters where the targets actually have detail.

"Causally available" in the shared-trunk row is doing a lot of work. It means a stream can only attend to what would exist at that point of generation: the current observation, the finished earlier modalities, and itself. Chapter 4 turns that phrase into an exact attention mask.

Conditioning every layer: adaLN

The network also has to know three things that are not tokens: the robot's joint configuration qt, the task embedding g, and how noisy each stream currently is (its flow timestep τ). The paper feeds all three through adaLN: "The robot configuration, task embedding, and flow timesteps for each modality condition the transformer through adaLN."

Start with plain layer normalization. Given a token vector x, subtract its mean, divide by its standard deviation, then multiply by a learned scale γ and add a learned shift β.

LN(x) = γ · (x − mean(x)) / std(x) + β

In plain layer norm, γ and β are fixed learned constants. In adaptive layer norm (adaLN), a small network computes γ and β from the conditioning. The analogy: a mixing engineer's fader. The song (the token) is the same, but the fader position (γ, β) is set by who is listening, how noisy the room is, and which song this is.

Worked example: adaLN on a 4-number token

The numbers here are illustrative. The point is to see one normalization and two different conditionings, with every step written out.

Token x = [2, 4, 6, 8].
Step 1. Mean = (2 + 4 + 6 + 8) / 4 = 20 / 4 = 5.
Step 2. Deviations = [2−5, 4−5, 6−5, 8−5] = [−3, −1, 1, 3].
Step 3. Variance = (9 + 1 + 1 + 9) / 4 = 20 / 4 = 5. Std = √5 = 2.2361.
Step 4. Normalized = [−3, −1, 1, 3] / 2.2361 = [−1.3416, −0.4472, 0.4472, 1.3416].
Step 5a. Conditioning A (a very noisy stream, τ = 0.1) makes the small network output γ = 0.5, β = 0.2:
   −1.3416×0.5 + 0.2 = −0.4708; −0.4472×0.5 + 0.2 = −0.0236; 0.4472×0.5 + 0.2 = 0.4236; 1.3416×0.5 + 0.2 = 0.8708.
   Output A = [−0.4708, −0.0236, 0.4236, 0.8708].
Step 5b. Conditioning B (an almost clean stream, τ = 0.9) gives γ = 1.5, β = −0.1:
   −1.3416×1.5 − 0.1 = −2.1124; −0.4472×1.5 − 0.1 = −0.7708; 0.4472×1.5 − 0.1 = 0.5708; 1.3416×1.5 − 0.1 = 1.9124.
   Output B = [−2.1124, −0.7708, 0.5708, 1.9124].

Same token, same weights, different behavior. In this toy, the conditioning turns the response down for the noisy stream and up for the clean one; a trained adaLN learns whatever mapping helps, and the paper does not describe what its learned mapping looks like. The general point stands either way: this is how one set of weights can serve every noise level and every task.

One subtlety deserves a flag. ModAR's streams are at different noise levels in the same forward pass (finished context is clean, the block being generated is noisy). "Flow timesteps for each modality" suggests the timestep part of the conditioning is applied per stream, so tokens of different streams get different γ and β in the same layer. That is our reading of the sentence; the paper does not show the adaLN wiring. The robot configuration and task embedding are described as "global adaLN conditioning to each DiT layer."

Sim 4 · who talks to whom, and what it costs

Pick a query stream and a layer type. Arrows show which streams that stream's tokens may attend to, during generation of that stream. In a shared block it reads the observation, every finished earlier modality, and itself; in its expert block it reads only itself. The τ slider drives a toy adaLN: watch the same normalized token get scaled and shifted differently. The bar at the bottom is the transformer-block parameter budget from the worked example below.

Stream τ

Worked example 3: where do 30.1 million parameters live?

First, which model does that number describe? The paper gives 30.1M once, for one specific variant: "our trained-from-scratch tracks–DINO–depth ModAR variant," the one it compares with Flex-π. That variant predicts no RGB, so it has no RGB expert. The paper does not report a parameter count for the four-modality model.

The paper gives the total and the block layout. It does not itemize the rest. So here is a back-of-envelope budget for the tracks–DINO–depth variant, using the standard rule of thumb that one transformer block with a 4× MLP holds about 12d2 parameters (4d2 in the four attention projections, 8d2 in the MLP), ignoring biases and norms. The 4× MLP ratio is our assumption (it is the ViT-S standard); the paper does not state it.

Step 1. Width-384 block: d2 = 384 × 384 = 147,456. 12d2 = 12 × 147,456 = 1,769,472.
Step 2. Width-128 block: d2 = 128 × 128 = 16,384. 12d2 = 12 × 16,384 = 196,608.
Step 3. Shared trunk: 6 × 1,769,472 = 10,616,832.
Step 4. DINO and depth experts: 2 modalities × 2 blocks = 4 blocks; 4 × 1,769,472 = 7,077,888. (No RGB expert in this variant.)
Step 5. Tracks expert: 1 × 196,608 = 196,608.
Step 6. Action expert: 2 × 196,608 = 393,216.
Step 7. Blocks total: 10,616,832 + 7,077,888 = 17,694,720; + 196,608 = 17,891,328; + 393,216 = 18,284,544 ≈ 18.3M.
Step 8. Paper total minus blocks: 30,100,000 − 18,284,544 = 11,815,456, about 11.8M in everything the rule of thumb does not itemize.

What could fill that 11.8M? Some pieces we can size exactly from Chapter 2, and the next worked example does. The rest belongs to parts the paper does not size: the adaLN conditioning networks, the task-embedding table, and whatever connects the width-128 experts to the width-384 trunk.

The leftover is large enough to be interesting. The original DiT computes a separate scale-and-shift regression in every block; at full width that costs roughly 6d2 per block. For this variant's ten width-384 blocks that is 10 × 6 × 147,456 = 8,847,360, and for its three width-128 blocks 3 × 6 × 16,384 = 294,912, together 8,847,360 + 294,912 = 9,142,272, about 9.1M. That fits inside the 11.8M, so the budget is at least consistent with a DiT-style per-block adaLN regression. It does not prove one: the paper does not show the wiring, and the lesson will not pretend otherwise.

Worked example: the projections and heads, sized exactly

Under the single-linear-layer reading of "we linearly project each modality" and "a linear output head," and assuming the track and action heads read from their width-128 experts, the in and out layers can be counted exactly (weights plus biases; our arithmetic). The table lists all five streams for reference; the sums below it are for the tracks–DINO–depth variant, which has no RGB head.

StreamInput projectionOutput head
RGB588 × 384 + 384 = 226,176384 × 588 + 588 = 226,380
Depth196 × 384 + 384 = 75,648384 × 196 + 196 = 75,460
DINO384 × 384 + 384 = 147,840384 × 384 + 384 = 147,840
Tracks3 × 384 + 384 = 1,536128 × 3 + 3 = 387
Actions14 × 384 + 384 = 5,760128 × 14 + 14 = 1,806
Sum, no RGB230,784225,493
Sum, all five456,960451,873
Depth + DINO + tracks + actions, projections: 75,648 + 147,840 + 1,536 + 5,760 = 230,784.
The same four streams, heads: 75,460 + 147,840 + 387 + 1,806 = 225,493.
Projections + heads for the tracks–DINO–depth variant = 230,784 + 225,493 = 456,277 ≈ 0.46M.
(If that simulated variant still observes RGB, which the paper does not say, add the RGB input projection: 456,277 + 226,176 = 682,453 ≈ 0.68M.)
Unitemized remainder from Worked example 3: 11,815,456 − 456,277 = 11,359,179 ≈ 11.4M for adaLN conditioning, embeddings, width adapters, and the biases and norms the 12d2 rule ignores.
Biases and norms are small: roughly 13d per block (our rule of thumb), so 13 × 384 = 4,992 per width-384 block and 13 × 128 = 1,664 per width-128 block; 10 × 4,992 + 3 × 1,664 = 49,920 + 4,992 = 54,912.

So most of the unexplained 11.4M must sit in the conditioning pathway and the adapters, the two parts the paper describes only in words. If you reimplement the tracks–DINO–depth ModAR and land far from 30.1M, those are the places to look.

✎CheckDraw the backbone from memory▸

Sketch the shared trunk, every expert stack with its depth and width, the heads, and the three conditioning signals. Then say where cross-modal attention happens and where it is forbidden.

Answer: 6 shared blocks at width 384 with 6 heads; experts of 2 blocks at width 384 for DINO, depth and RGB, 1 block at width 128 for tracks, 2 blocks at width 128 for actions; one linear head per stream; adaLN from robot configuration, task embedding and per-stream flow timesteps. Cross-modal attention happens only in the shared blocks (and only toward the observation and earlier finished blocks); expert blocks restrict attention to their own stream.

Our estimate: about two fifths of the block parameters are specialists. In the tracks–DINO–depth variant, the 12d2 rule puts 10.6M in the shared trunk and 7.7M in experts (7,077,888 + 196,608 + 393,216 = 7,667,712, which is 42% of the 18.3M in blocks). ModAR does not just bolt tiny heads onto a big shared model; it gives each detailed visual modality a real two-block stack of its own. That is the "within-modality specialization" half of the design, and it is sized like it matters. (Both figures rest on our 4× MLP assumption.)

The four-modality model is larger (our estimate)

Since 30.1M belongs to the tracks–DINO–depth variant, the four-modality model of Tables I and II must be bigger: adding an RGB target adds an RGB expert and an RGB output head. The paper does not report that model's size, but the same rule of thumb estimates the difference.

RGB expert: 2 × 1,769,472 = 3,538,944. RGB output head: 226,380.
Added ≈ 3,538,944 + 226,380 = 3,765,324 ≈ 3.8M, or 3,765,324 + 226,176 = 3,991,500 ≈ 4.0M if the RGB input projection is also new.
So the four-modality model is plausibly about 30.1M + 4M ≈ 34M by this rule (our estimate, not a paper number).

The real-world WAMs "observe and predict only tracks, DINO, and depth," the same modality set as the 30.1M variant. So the real-robot model is plausibly close to 30.1M too, though the paper does not report its size either.

Not a mixture of experts in the router sense

"Experts" can mislead. In large language models, a mixture of experts uses a learned router that sends each token to a few of many experts. ModAR has no router. Every DINO token goes to the DINO expert, always; the assignment is fixed by modality. This is closer in spirit to the separate action expert in π0, where a smaller set of weights is dedicated to one kind of token, than to a routed sparse model.

What is trained, and why nothing is pretrained

Every weight in the trunk, the experts, the adaLN pathway, the embeddings, the projections and the heads starts from random initialization. The only pretrained network anywhere in the pipeline is the frozen DINOv2 encoder that defines DINO targets (plus CoTracker3 producing track targets). That is not a limitation of engineering effort; it is the experimental design. With no pretrained trunk, a gap between two ModAR variants can only come from the variants.

It also means the model is small enough to be trained many times. The paper trains one multitask model per method and per data scale, for 1.2M optimizer steps each (Chapter 5). Presumably (our inference; the paper gives no reason) that cost is also why the 6B Flex-π appears in Chapter 8 as a single system-level comparison rather than across every configuration.

What the split buys when a stream goes wrong

Suppose the depth stream is currently a mess, early in its denoising. In the shared blocks, depth tokens are being generated and can read the observation, tracks and DINO, but nothing generated later can read them yet. In the depth expert blocks, the mess stays inside the depth stream. So during depth generation, no other stream's representation is being perturbed by half-finished depth. The block-causal mask (Chapter 4) guarantees the first property; the expert restriction guarantees the second.

The price of this isolation is that the only place modalities meet is the six shared blocks. If cross-modal reasoning needed more depth than that, the design would bottleneck there. The results in Chapters 6 and 7 suggest six was enough for these tasks; the paper does not ablate the trunk depth.

⇆Cross-domain bridge
Multitask vision models made the same bet
The paper places ModAR in a line of work that predicts several representations of one signal from a shared backbone, citing MultiMAE, 4M, Unified-IO 2 and Cosmos 3. The shared-then-specialized shape is the common thread: a common trunk learns the structure every task needs, and each output keeps a little private capacity for its own details. The paper's twist is to put a generation order on top of that shape.

Tracing one DINO token through the network

Concept plus realization: here is the path a single future-DINO token takes, with shapes, while ModAR generates the DINO block at inference (four-modality setting, batch size B). The shapes of the input and output are fixed by Chapter 2; the attention sizes use our assumed 576 observation tokens.

Input
Noisy DINO block Ỹ of shape (B, 384, 384): 384 tokens (2 frames × 12 × 16), each 384 numbers. Linear projection to width 384, plus the DINO modality embedding.
↓
6 shared blocks
Queries: the 384 DINO tokens. Keys and values: 576 observation + 384 finished tracks + 384 DINO = 1,344 tokens, the first 960 read from the KV cache. Attention scores per head: 384 × 1,344 = 516,096. adaLN scales and shifts using qt, g, and the DINO stream's current τ.
↓
2 DINO expert blocks (width 384)
Attention restricted to the DINO stream: 384 × 384 = 147,456 scores per head.
↓
Linear head
Width 384 to the DINO feature size (384): the clean guess Ŷ of shape (B, 384, 384), reshaped to (B, 2, 12, 16, 384).
↓
Sampler
v = (Ŷ − Ỹ) / (1 − τ); Ỹ ← Ỹ + v / 8. After 8 steps the block is finished and re-embedded as context.

The action stream takes the same path with different sizes: 16 tokens of 14 numbers are projected to width 384; in the shared blocks those 16 queries read 576 + 1,536 + 16 = 2,128 keys; then the width-128 action expert runs, and a head maps back to 14 numbers per step. Going from width 384 to 128 requires some projection between the trunk and the narrow experts. The paper does not describe it, so treat that part of the diagram as a necessary but unspecified adapter.

What each conditioning signal tells the network

SignalShapeWhat the network can do with itEnters through
qt14 numbersKnow where its own arms are, so an inverse dynamics step can compute how far to moveGlobal adaLN
glearned vector per taskKnow which of the multitask behaviors to imagine (six tasks in simulation, three on the robot)Global adaLN
τ per streamone number per streamKnow how much to trust its input: rough structure at low τ, fine detail at high τadaLN
Observation tokenstokensSee the scene right nowAttention (not adaLN)

Worked example: attention work per query, ModAR versus Unified

Sequential generation needs five times more network passes than Unified. Does it need five times more work? Count the attention scores computed per head in the shared blocks, summed over all denoising steps of one query (our arithmetic, four-modality setting, 576 observation tokens assumed, observation and finished blocks cached).

ModAR, queries × visible keys per step:
  tracks 384 × 960 = 368,640 · DINO 384 × 1,344 = 516,096 · depth 384 × 1,728 = 663,552
  RGB 384 × 2,112 = 811,008 · actions 16 × 2,128 = 34,048
  Sum per step-set = 368,640 + 516,096 + 663,552 + 811,008 + 34,048 = 2,393,344; × 8 steps = 19,146,752.
Unified, every step: all 576 + 4 × 384 + 16 = 2,128 tokens attend to all 2,128.
  2,128 × 2,128 = 4,528,384 per step; × 8 steps = 36,227,072.
Ratio ModAR / Unified = 19,146,752 / 36,227,072 = 0.53.

By this proxy, ModAR computes roughly half the shared-block attention of a Unified query, because each of its many passes only pushes one block's queries through the network while the rest sits in the cache. So why does the paper say sequential generation increases latency? Because 40 small passes must run one after another, while 8 larger passes can use the GPU's parallelism more fully; for a 30M-parameter model, per-pass overhead is plausibly a large share of the time. That explanation is ours. The paper's statement, and its measured 147.9 ms, are the facts.

Why not simply one bigger shared trunk?

The paper does not ablate this, so what follows is design reasoning rather than evidence. A fully shared stack forces every layer to serve every modality at once: a depth map, a feature map and a joint-command chunk all compete for the same weights all the way to the output. Experts give each target a few private layers to turn the shared, fused representation into its own peculiar output format. The block-causal mask means cross-modal fusion must happen where attention can reach earlier blocks, which is exactly the shared trunk, so the design puts its sharing where the conditioning happens and its specialization where the formats diverge.

🔍Debug itContext that the network thinks is noisy▸

In a reimplementation, the finished context blocks are fed into adaLN with the same τ as the block currently being denoised (for example τ = 0.125 at the second step). What would you expect, and what is the fix?

Answer (our reasoning): adaLN is how the network learns how much to trust a stream. Labeling finished context as "mostly noise" tells the trunk to treat solid information as unreliable, so later blocks lean less on earlier ones, which undoes the scratchpad effect. Each stream needs its own timestep signal ("flow timesteps for each modality"). What value the paper assigns to context at inference is not stated; a value consistent with how context was noised in training (between 1 − β and 1) is the natural choice.

In ModAR's backbone, where does information move between modalities?

Chapter 4: Generating in Blocks, in Order

A novelist writing a murder mystery does not write every page at once. They decide who did it first. Then the clues get written to agree with that decision. Then the red herrings, which only make sense once the real clues exist. The last chapter, the reveal, is written knowing everything. Each stage is easier because the earlier ones are settled.

ModAR generates the future the same way, and this chapter makes that precise. First with probability, where the ordering turns out to cost nothing. Then with an attention mask, which is how the ordering is enforced inside a transformer. Then with a training trick that supervises every stage in a single forward pass.

The chain rule: ordering is free

Any joint probability of two things can be split into "the first thing" times "the second thing given the first."

p(a, b) = p(a) · p(b | a) = p(b) · p(a | b)

This is the chain rule of probability, and it holds exactly, for any variables and in either order. Nothing is approximated. So a model that generates a first and then b given a can represent exactly the same joint distribution as a model that generates both at once.

A tiny example with illustrative numbers shows why the order still matters in practice. Let M be "the cup moves this chunk" (a fact tracks would reveal) and S be "the cup's patch region shifts" (a fact DINO would reveal). Suppose the joint probabilities are p(M=1, S=1) = 0.45, p(M=1, S=0) = 0.05, p(M=0, S=1) = 0.05, p(M=0, S=0) = 0.45.

Step 1. p(M=1) = p(M=1,S=1) + p(M=1,S=0) = 0.45 + 0.05 = 0.50.
Step 2. p(S=1 | M=1) = p(M=1,S=1) / p(M=1) = 0.45 / 0.50 = 0.90.
Step 3. Chain rule check: p(M=1) × p(S=1 | M=1) = 0.50 × 0.90 = 0.45 = p(M=1, S=1). Exact.
Step 4. Without knowing M, the second variable is a coin flip: p(S=1) = 0.45 + 0.05 = 0.50.
Step 5. With M in hand, the same variable is 90% predictable.

The joint is identical either way. What changes is the difficulty of each step. Once the motion is settled, predicting where the cup's patches go is nearly deterministic. That is the whole "scratchpad" intuition in five lines of arithmetic.

The paper's factorization

Now scale the idea up. Pick an ordering (m1, …, mK) of the modalities you predict (all four, or a subset), and write Yt for their future targets in that order. The paper models the joint distribution of futures and actions as its Equation (1):

pθ(Yt, At | ct) = pθ(At | ct, Yt) · ∏k=1..K pθ(Ytmk | ct, Yt<k)

Symbol by symbol. pθ is the model's distribution, with all network weights collected in θ; there is one network, used for every factor. Ytmk is the future of the k-th modality in the order. Yt<k is shorthand for all the modalities before it, (Ym1, …, Ymk−1); for k = 1 it is empty. The product ∏ multiplies the K future factors, and the leading factor is the action, conditioned on the observation and on every predicted future.

Written out for the three-modality setting the paper uses on real robots, with the order tracks, DINO, depth:

p(Ytr, Ydino, Ydep, A | c) = p(Ytr | c) · p(Ydino | c, Ytr) · p(Ydep | c, Ytr, Ydino) · p(A | c, Ytr, Ydino, Ydep)

Read the last factor as a job description. Given where things were (c), where they will be (Y), and what they will look like, output the commands that make it happen. That is the inverse dynamics model from Chapter 1, and the paper uses the same words: the final action-prediction step "acts as an inverse dynamics model (IDM), mapping the generated future observations to the actions that induce the predicted transitions."

One plausible reading of the IDM framing (ours, not the paper's). A policy must decide what should happen. An IDM is handed a proposed future and asked how to make it happen; we would argue that is often the narrower question. Read that way, ModAR moves much of the "what should happen" decision into the future modalities, which can learn from actionless demonstrations, while the action step still reads the observation, configuration and task alongside every finished future. This may help explain (our hypothesis) why ModAR gains the most from actionless data in Chapter 6. The paper reports that gain but offers no mechanism for it; the only explanation it gives for ModAR's advantage is the scratchpad hypothesis.

Blocks, not tokens

Language models are autoregressive token by token: word 1, then word 2, then word 3. ModAR is autoregressive block by block: "we generate all tokens of one modality jointly before moving to the next modality." Within a block, all 384 tokens of, say, the DINO future are denoised together, in parallel, over 8 flow steps. Across blocks, the order is strict.

Why not token by token? The paper does not argue this, but the numbers make it obvious: four modalities at 384 tokens each would be 1,536 sequential generation steps, each a network pass. Block-wise generation needs 8 passes per block instead, and the within-block structure (every patch of a depth map depends on its neighbors) is handled by the denoiser, which is good at exactly that.

The order: tracks, DINO, depth, RGB

The paper fixes one order for all main experiments: tracks → DINO → depth → RGB, with actions last. The stated intuition: "this orders the targets from compact, structured representations that are easier to predict toward increasingly high-dimensional and detailed representations."

Compare that with the raw target sizes from Chapter 2: tracks 1,152 numbers, DINO 147,456, depth 75,264, RGB 225,792. Depth has fewer raw numbers than DINO but comes after it. So the order is not "smallest first" in a literal sense; it is "most structured first." A DINO future is a map of what is where, coarse and semantic; a depth future asks for metric detail on every patch.

The paper is candid about how much it tested this. It compares the chosen order with its reverse (Chapter 7), and "We leave a full systematic comparison of modality orderings to future work." Its limitations section adds that the best order "may depend on the task or specific scenario."

Inference: denoise, re-embed, move on

At test time the paper's procedure is three verbs: "we therefore denoise one block, re-embed the completed prediction as context, and then denoise the next block."

Re-embed means the finished prediction is treated like an input: it goes through the same linear projection and modality embedding as a clean target would, and becomes a context block. Every later block attends to that context.

The paper adds a speed trick: "We cache keys and values for the current observation and all previously generated modalities to avoid recomputing them at every denoising step." A key-value (KV) cache stores each token's attention keys and values so they need not be recomputed. It is exact here, not approximate, because of the block-causal structure: an earlier block never attends to a later one, so nothing that happens during later denoising can change an earlier block's keys and values.

Worked example 4: what the cache saves

For the arithmetic we need a guess about the observation tokens. Figure 2 of the paper shows the current observation embedded from DINO, depth and RGB, so we assume 3 × 192 = 576 observation tokens in the four-modality simulation setting. Every other count comes from Chapter 2. This is our arithmetic, counting "tokens pushed through the network," which is a rough proxy for compute.

Stage sizes (tokens visible while generating each block, without caching):
  tracks: 576 obs + 384 own = 960
  DINO: 576 + 384 (tracks ctx) + 384 own = 1,344
  depth: 576 + 768 (2 ctx) + 384 own = 1,728
  RGB: 576 + 1,152 (3 ctx) + 384 own = 2,112
  actions: 576 + 1,536 (4 ctx) + 16 own = 2,128
Without a cache, every one of the 8 steps per stage reprocesses everything visible:
  8 × (960 + 1,344 + 1,728 + 2,112 + 2,128) = 8 × 8,272 = 66,176 token passes.
With a cache, each step processes only the block being generated:
  8 × (384 + 384 + 384 + 384 + 16) = 8 × 1,552 = 12,416, plus each context block embedded once: 576 + 4 × 384 = 2,112.
  Total = 12,416 + 2,112 = 14,528 token passes.
Ratio: 66,176 / 14,528 = about 4.6× fewer token passes with the cache.

The exact wall-clock saving depends on details this proxy ignores (attention cost grows with the number of visible tokens; the experts only process their own stream). The direction is unambiguous, and it is why sequential generation is affordable at all: ModAR's measured 147.9 ms includes 40 denoising steps.

Training: every stage in one pass

Training has a problem inference does not. At inference you generate the blocks one after another. In training you have the true futures for every block, and you want to teach all K + 1 conditionals of Equation (1) at once, from one forward pass, without letting any block cheat.

The paper's solution: "we supervise every stage in one pass by creating a clean context copy and a noisy prediction copy of each target modality." Every target modality appears twice in the training sequence. The context copy holds the true future (clean, apart from the context noise of Chapter 5), playing the role a finished block plays at inference. The prediction copy holds a noised version of the same future, which the model must denoise.

Then one block-causal mask decides who reads whom. In the paper's words, it "lets the prediction copy for mk attend only to the observation and context copies of m1, …, mk−1, preventing target leakage." (Each prediction copy also attends within itself, since its own tokens are denoised jointly.)

Target leakage is what happens when a model can see the answer it is graded on. If the noisy DINO copy could read the clean DINO copy, the network would learn to copy it, the loss would drop toward zero, and nothing about predicting DINO would be learned. If it could read the clean depth copy (a later modality), it would learn to rely on information that will not exist yet at inference. The mask forbids both.

What do the context copies themselves attend to? The natural reading, mirroring inference where each finished block was embedded after its predecessors, is that a context copy reads the observation, earlier context copies, and itself. The paper states the rule only for prediction copies, so treat that as our reading.

Worked example 5: one pass versus five

Training sequence with two copies (four-modality setting, our assumed 576 observation tokens):
  observation 576 + context copies 4 × 384 = 1,536 + prediction copies 4 × 384 = 1,536 + noisy actions 16
  = 576 + 1,536 + 1,536 + 16 = 3,664 tokens.
(Actions need no context copy: nothing is generated after them.)
Supervising each stage with its own forward pass instead would cost the five stage sizes from Worked example 4:
  960 + 1,344 + 1,728 + 2,112 + 2,128 = 8,272 tokens.
Saving: 1 − 3,664 / 8,272 = 1 − 0.443 = about 56% fewer tokens per training example.

And beyond compute, the single pass has a quieter benefit: all K + 1 losses in Equation (4) come from the same example at the same time, so the gradient for "predict DINO from tracks" and "predict actions from everything" are computed on identical context.

Sim 5 · the block-causal mask

Rows are the blocks doing the reading (queries); columns are the blocks being read (keys). Filled cells are allowed. Tap any cell for the reason. Switch between ModAR's training sequence, ModAR at inference (scrub the stage), Unified and Disjoint. Then press Introduce a leak and read what would go wrong.

Three patterns are worth finding in the matrix. The prediction-copy rows have a staircase of allowed context cells: each one reads one more context block than the row above. The context-copy columns are read by later prediction copies only, never by the prediction copy of the same modality. And the action row reads every context copy, which is Equation (1)'s last factor drawn as a mask.

The mask, as code

ORDER = ["tracks", "dino", "depth", "rgb", "act"]
SIZE  = {"obs": 576, "tracks": 384, "dino": 384, "depth": 384, "rgb": 384, "act": 16}

# the training sequence: observation, clean context copies, noisy prediction copies
blocks = [("obs", "obs")] + [("ctx", m) for m in ORDER[:-1]] \
                         + [("noisy", m) for m in ORDER]

def allowed(q, k):
    (qk, qm), (kk, km) = q, k
    if kk == "obs":   return True                  # everyone reads the current observation
    if qk == "obs":   return False                 # the observation reads only itself
    if q == k:        return True                  # a block's own tokens are denoised jointly
    if kk == "ctx":   # clean context: only STRICTLY EARLIER modalities
        return ORDER.index(km) < ORDER.index(qm)
    return False                                   # never read another block's noisy copy

def token_mask(blocks):                            # (L, L) boolean, L = 3,664 here
    spans, s = [], 0
    for b in blocks:
        spans.append((s, s + SIZE[b[1]])); s += SIZE[b[1]]
    M = torch.zeros(s, s, dtype=torch.bool)
    for i, q in enumerate(blocks):
        for j, k in enumerate(blocks):
            if allowed(q, k):
                M[spans[i][0]:spans[i][1], spans[j][0]:spans[j][1]] = True
    return M                                        # used in the SHARED blocks; experts see only their own span

Two implementation notes connect this to the rest of the model. The mask applies in the shared blocks; in the expert blocks, each stream's attention is already restricted to its own stream (and, for context versus prediction copies, the same no-leak rule must hold there too). For an actionless example, the action block has no target, so the action loss term is simply dropped, as the paper specifies.

Why the observation reads only itself. The observation is a given, not a prediction. Letting it read the future blocks would make its representation depend on things that do not exist at inference, and it would break the KV cache, since the cached observation keys would then have to change as futures are generated. (The paper does not print the observation row of the mask; this is the rule that keeps its caching statement true.)
🔍Debug itA training loss that looks too good▸

A reimplementation of ModAR shows the DINO loss collapsing to almost zero within a few thousand steps, while closed-loop success stays at the Action-only level. Using only this chapter, name the most likely bug and the one-line check that confirms it.

Answer: the mask lets the noisy DINO prediction copy read the clean DINO context copy (target leakage). The network copies the answer instead of predicting it, so at inference, where no clean DINO exists, it has learned nothing useful. Check: assert that allowed(("noisy","dino"), ("ctx","dino")) is False, and more generally that no prediction copy can read a context copy at or after its own position in the order.

Two sequences, one set of weights

There is a quiet correctness property hiding in this chapter, and it is worth stating plainly. At inference, the DINO block is computed while reading exactly {observation, finished tracks}. In training, the noisy DINO copy is computed while reading exactly {observation, tracks context copy}. Same set of blocks, same attention pattern, same weights. The only differences are that training context is a (lightly noised) ground truth instead of a generation, and that training handles every stage in one pass.

This is the block-level version of teacher forcing, the standard trick for training autoregressive models: during training, feed the model the true previous outputs rather than its own guesses, so every step can be supervised in parallel. Teacher forcing has a well-known weakness: at test time the model sees its own imperfect outputs, which it never practiced on. Context noise (Chapter 5) is ModAR's answer to exactly that weakness. The mask makes training and inference structurally identical; context noise makes them statistically closer.

One training example, stage by stage

Put the mask and Equation (1) side by side for a single four-modality training example. Every row below is computed in the same forward pass.

Prediction copyReads (besides itself)Factor of Equation (1) it trains
noisy tracksobservationp(Ytr | c)
noisy DINOobservation, ctx tracksp(Ydino | c, Ytr)
noisy depthobservation, ctx tracks, ctx DINOp(Ydep | c, Ytr, Ydino)
noisy RGBobservation, ctx tracks, DINO, depthp(Yrgb | c, Ytr, Ydino, Ydep)
noisy actionsobservation, all four context copiesp(A | c, Y)  (the IDM; only if the example has action labels)

For an actionless example, the last row simply contributes no loss. Everything above it still trains, which is how a human video teaches "given this motion, what will the features and depth look like?" without ever saying anything about joint commands.

Notice what the table implies for a labeled example: all five rows train at once, and each reads a different amount of context. The tracks row learns to imagine from the observation alone, the hardest possible starting point, while the action row sees the most. If you mask the table differently (for example, letting the depth row read nothing but the observation), you are no longer training Equation (1), and the inference loop, which always hands depth the finished tracks and DINO, would meet inputs the network never learned to use.

The reverse order, written out

The ablation in Chapter 7 reverses the future order to RGB → depth → DINO → tracks, with actions still last. Its factorization is equally exact:

p(Y, A | c) = p(Yrgb | c) · p(Ydep | c, Yrgb) · p(Ydino | c, Yrgb, Ydep) · p(Ytr | c, Yrgb, Ydep, Ydino) · p(A | c, Y)

The difference is the first job. In the forward order, the first block to be generated from nothing but the observation is the compact track field (1,152 raw numbers). In the reverse order, it is a full RGB future (225,792 raw numbers), with every lighting and texture detail, and every later block conditions on whatever mistakes that hardest-first step made. Same joint distribution on paper; very different sequence of difficulties in practice. The measured cost is 10 points (75% to 65%).

✎DeriveThe two-modality model from Table II▸

Table II includes a "Tracks + DINO" model. Write its factorization, list its training-sequence blocks, and count its tokens with the assumed 576 observation tokens.

Answer: p(Ytr, Ydino, A | c) = p(Ytr | c) · p(Ydino | c, Ytr) · p(A | c, Ytr, Ydino). Blocks: observation, ctx tracks, ctx DINO, noisy tracks, noisy DINO, noisy actions. Tokens: 576 + 2 × 384 + 2 × 384 + 16 = 576 + 768 + 768 + 16 = 2,128. At inference it takes 8 × 3 = 24 network passes instead of 40.

✎CheckWhy is the KV cache exact, not approximate?▸

Explain in two sentences why ModAR can cache the observation's and finished blocks' keys and values across all later denoising steps without changing the result, and name one mask change that would break this.

Answer: a token's keys and values depend only on the tokens it attends to, and under the block-causal mask the observation and finished blocks never attend to anything generated later, so their keys and values are fixed once computed. Letting the observation (or a finished block) attend to a later, still-changing block would break it: its keys and values would change at every step and the cache would go stale.

Why does ModAR's training sequence contain both a clean context copy and a noisy prediction copy of each target modality?

Chapter 5: The Flow Objective and Context Noise

Chapter 4 said each block is "denoised over 8 flow steps" and moved on. This chapter opens that box. By the end you will be able to write ModAR's training loss, its sampler, and its defense against cascading errors, with every symbol accounted for.

Mixing noise into a target: the interpolant

Take one future modality's true target Y (say, the DINO future: 384 tokens of 384 numbers). Draw a same-shaped block of Gaussian noise ε. Pick a mixing level τ between 0 and 1. The paper forms the noisy version with a straight-line blend, its Equation (2):

Ỹtm = τm Ytm + (1 − τm) εm

Ytm is the clean target of modality m. εm is noise drawn from a standard normal, N(0, I), independently for each modality. τm is the flow timestep, also drawn independently per modality in ModAR. Ỹtm (read "Y tilde") is the noisy input the network sees. At τ = 1 it is the clean target; at τ = 0 it is pure noise. This straight-line path is the linear flow-matching interpolant of Lipman et al. (the flow matching gleam derives it).

One number, to make it concrete (illustrative): a target value Y = 0.8, a noise draw ε = −1.2, and τ = 0.25 give Ỹ = 0.25 × 0.8 + 0.75 × (−1.2) = 0.2 − 0.9 = −0.7. Three quarters noise, one quarter signal.

Three things the network could predict

Given Ỹ and τ, a denoiser can be trained to output one of three equivalent-looking quantities.

TargetMeaningFormula
ε-prediction"Which noise was added?"ε
v-prediction"Which way, and how fast, does the path move?"v = dỸ/dτ = Y − ε
x-prediction"What is the clean answer?"Y

Given Ỹ and τ, any one of the three determines the other two, so they look interchangeable. They are not interchangeable for learning. ModAR trains "every output stream with a JiT-style x-prediction objective": the output head directly predicts the clean target, written Ŷtm (read "Y hat"). JiT is the paper "Back to Basics: Let Denoising Generative Models Denoise" by Li and He.

The paper gives both an observation and a reason. The observation: "replacing x-prediction with velocity (v) prediction is often unstable and can cause training to diverge." The reason: "Velocity prediction can struggle in high-dimensional spaces; predicting the clean sample is well-conditioned when the data lie on a low-dimensional manifold, as with images, depth, and other visual representations."

Unpack "low-dimensional manifold" with an analogy. Every possible 384-token DINO map is a point in an enormous space, but the maps that real scenes produce occupy a thin, structured sheet inside it, the way all real faces occupy a tiny corner of the space of all pixel grids. A clean target always lies on that sheet, so predicting it asks the network to land on something structured. Noise and velocity are spread across the whole space, so predicting them asks the network to reproduce full-dimensional randomness. The paper applies the same x-prediction parameterization to robot actions, citing prior VLA and WAM work that does likewise.

From a clean guess to a velocity

A sampler needs a velocity to move along, but the network outputs a clean guess. The paper converts: vθ = (Ŷ − Ỹ) / (1 − τ). Here is where that comes from, in three lines.

Step 1. If the network's clean guess is Ŷ, the noise consistent with it satisfies Ỹ = τŶ + (1 − τ)ε̂, so ε̂ = (Ỹ − τŶ) / (1 − τ).
Step 2. The velocity is clean minus noise: v̂ = Ŷ − ε̂ = Ŷ − (Ỹ − τŶ) / (1 − τ).
Step 3. Common denominator: v̂ = [(1 − τ)Ŷ − Ỹ + τŶ] / (1 − τ) = (Ŷ − Ỹ) / (1 − τ). This is the paper's vθ.

The loss, and why it has that odd denominator

For each supervised modality the paper minimizes its Equation (3):

Lm(θ) = E [ ‖Ŷtm − Ytm‖22 / max(1 − τm, δm)2 ]

E is the average over training examples, noise draws and timesteps. The numerator is the squared error of the clean guess. The denominator divides by (1 − τ) squared, except that it never lets that quantity fall below δm, which the paper says "stabilizes the loss near the clean endpoint." The paper sets δm = δact = 0.05.

Why divide at all? Because that denominator turns the clean-guess error into a velocity error. Here is the derivation (ours, but it follows directly from the paper's two formulas).

Step 1. True velocity: v = Y − ε. From Equation (2), Y − Ỹ = Y − τY − (1 − τ)ε = (1 − τ)(Y − ε), so v = (Y − Ỹ) / (1 − τ).
Step 2. Predicted velocity: v̂ = (Ŷ − Ỹ) / (1 − τ).
Step 3. Difference: v̂ − v = (Ŷ − Ỹ − Y + Ỹ) / (1 − τ) = (Ŷ − Y) / (1 − τ).
Step 4. Squared: ‖v̂ − v‖2 = ‖Ŷ − Y‖2 / (1 − τ)2, which is Equation (3) without the δ clamp.

So the network outputs a clean guess (the well-conditioned target), but it is graded in velocity units, which is what the sampler actually uses. That combination, x-prediction with a velocity-space loss, is the JiT recipe the paper cites. The δ clamp exists because as τ approaches 1 the factor 1/(1 − τ)2 explodes, and tiny errors on almost-clean inputs would dominate every batch.

Worked example 6: how the weight moves with τ

Weight w(τ) = 1 / max(1 − τ, 0.05)2.
τ = 0.00: 1 − τ = 1.00 → w = 1 / 1.002 = 1
τ = 0.50: 1 − τ = 0.50 → w = 1 / 0.25 = 4
τ = 0.90: 1 − τ = 0.10 → w = 1 / 0.01 = 100
τ = 0.95: 1 − τ = 0.05 = δ → w = 1 / 0.0025 = 400
τ = 0.99: 1 − τ = 0.01 < δ, so the clamp uses 0.05 → w = 400 (unclamped it would be 1 / 0.0001 = 10,000)
Example: a clean-guess error of 0.1 on one value at τ = 0.9 costs 0.12 × 100 = 0.01 × 100 = 1.0; the same error at τ = 0 costs 0.01 × 1 = 0.01.

Near the clean end, the loss insists on precision; near the noisy end, it forgives. That matches intuition: when the input is almost clean, a good model has no excuse for being wrong.

The total loss

Each modality's loss and the action loss are simply added, the paper's Equation (4), with unit weights:

L(θ) = ∑k=1..K Lmk + Lact

Lact has the same form as Equation (3), with the action chunk At, its prediction, its own timestep τact and δact in place of the modality versions. The action term is "include[d] only for examples with action labels," which is how actionless human or simulated demonstrations contribute: they supply every future term and nothing else.

Where on the path to practice: logit-normal timesteps

Training needs a τ for every stream in every example. Uniform sampling would spend equal effort everywhere on the path. ModAR instead samples logit-normal timesteps: draw n from a normal distribution N(μ, σ), then squash it, τ = 1 / (1 + e−n). μ slides where the practice concentrates; σ sets how spread out it is. The paper uses a different (μ, σ) for each stream.

Stream(μ, σ)Median τ = sigmoid(μ)Middle ~68% of τ (sigmoid(μ ± σ))
DINO(−2, 1)0.1190.047 to 0.269
Tracks(0, 1)0.5000.269 to 0.731
Depth(−1, 1.6)0.2690.069 to 0.646
RGB(−1, 1)0.2690.119 to 0.500
Actions(−1, 1)0.2690.119 to 0.500

The last two columns are our arithmetic. One of them in full: sigmoid(−2) = 1 / (1 + e2) = 1 / (1 + 7.389) = 1 / 8.389 = 0.1192. So half of all DINO training examples sit at τ below 0.12: the DINO stream practices mostly in the noisy regime, where the big structural decisions happen. Tracks practice centered on the middle of the path, and depth's larger σ spreads its practice widely. The paper reports these settings without explaining how they were chosen, so we will not invent a rationale.

For the baselines: "Unified uses one shared flow timestep with (μ, σ) = (−1, 1); Independent-noise samples each stream's modality-specific distribution independently." That second sentence is the concrete version of the mismatch from Chapter 1: at inference all streams share one τ, but in training a DINO stream near 0.12 and a tracks stream near 0.5 are the typical case.

The sampler: eight Euler steps per stream

At inference, each stream starts as pure noise at τ = 0 and is integrated to τ = 1 "using the velocity implied by its clean prediction." Euler integration is the simplest way to follow a velocity: take the current point, add velocity times step size, repeat. With 8 steps the step size is 1/8 = 0.125:

yk+1 = yk + 0.125 · (Ŷk − yk) / (1 − τk),    τk = k / 8

Worked example 7: one value, eight steps

Illustrative numbers: the true clean value is 0.8 and the starting noise is −1.2. First, a perfect network (it always guesses 0.8).

k = 0: τ = 0; v = (0.8 − (−1.2)) / (1 − 0) = 2.0 / 1 = 2.0; y = −1.2 + 0.125 × 2.0 = −1.2 + 0.25 = −0.95
k = 1: τ = 0.125; v = (0.8 − (−0.95)) / 0.875 = 1.75 / 0.875 = 2.0; y = −0.95 + 0.25 = −0.70
k = 2 … 6: the same arithmetic gives v = 2.0 every time, so y climbs by 0.25 per step: −0.45, −0.20, 0.05, 0.30, 0.55
k = 7: τ = 0.875; v = (0.8 − 0.55) / 0.125 = 0.25 / 0.125 = 2.0; y = 0.55 + 0.25 = 0.80 exactly the target

A perfect x-predictor walks a straight line at constant speed, which is what "linear interpolant" promised. Now an imperfect network whose guess is biased high by 0.3 × (1 − τ), a toy error that shrinks as the input gets cleaner.

kτguess Ŷcurrent yv = (Ŷ − y)/(1 − τ)next y
00.0001.1000−1.20002.3000 / 1.000 = 2.3000−0.9125
10.1251.0625−0.91251.9750 / 0.875 = 2.2571−0.6304
20.2501.0250−0.63041.6554 / 0.750 = 2.2071−0.3545
30.3750.9875−0.35451.3420 / 0.625 = 2.1471−0.0861
40.5000.9500−0.08611.0361 / 0.500 = 2.07210.1730
50.6250.91250.17300.7395 / 0.375 = 1.97210.4195
60.7500.87500.41950.4555 / 0.250 = 1.82210.6472
70.8750.83750.64720.1903 / 0.125 = 1.52210.8375

Look at the last row. With evenly spaced steps, the final Euler step multiplies (Ŷ − y) by 0.125 / (1 − 0.875) = 1, so the output lands exactly on the network's last clean guess, 0.8375. The final error (0.0375) is whatever the network gets wrong at τ = 0.875, which is precisely where Equation (3) weights errors 64 times more heavily than at τ = 0 (1 / 0.1252 = 64). The loss and the sampler are pulling in the same direction.

Cascading error, and the context-noise fix

Now the danger from Chapter 1, in this chapter's language. During training, the context copies are the true futures. During inference, the context blocks are the model's own generations, carrying errors like the 0.0375 above. A model trained only on perfect context has never seen imperfect context, so it may trust a slightly wrong track block completely and build a wrong DINO block on top of it, then a wrong depth block, then wrong actions. The paper: "imperfections in earlier generations become inputs to all later predictions."

The fix, "following Latent Forcing," is to make the training context imperfect on purpose. When predicting mk, for every context block mj and every training example, independently sample

εjctx ~ N(0, I),   τjctx ~ U(1 − β, 1),   Ỹt,ctxmj = τjctx Ytmj + (1 − τjctx) εjctx

U(1 − β, 1) is a uniform draw between 1 − β and 1. β is the maximum context-noise level; the paper sets β = 0.5, so every context block is somewhere between half noise (τ = 0.5) and perfectly clean (τ = 1), with an average mixing level of 0.75. This noise is applied "only during training, not during inference."

Worked example 8 (illustrative values). A context value Y = 0.8, noise draw ε = −1.2, and τctx = 0.7:
Ỹctx = 0.7 × 0.8 + (1 − 0.7) × (−1.2) = 0.56 + 0.3 × (−1.2) = 0.56 − 0.36 = 0.20.
The later block reads 0.20 where the truth is 0.8. To keep its own loss low, it must learn to use context as evidence, cross-checked against the observation, not as gospel.
Average corruption: E[1 − τctx] = 1 − (0.5 + 1)/2 = 1 − 0.75 = 0.25 of the context is noise, on average.
This one trick is worth twelve points. Remove context noise and ModAR's average success at 250 demonstrations falls from 75% to 63% (Figure 6, Chapter 7). For comparison, the Unified baseline at the same scale is 67%. Without context noise, the modality-autoregressive design would actually lose to the joint one it was built to beat.
Sim 6 · the flow lab

Denoise: one row of a toy depth future is integrated from noise over 8 Euler steps with an x-predicting network; raise the predictor error and scrub the steps. Cascade: a toy chain of four stages where an error in the first block propagates; flip context-noise training on and off. Sampler: the paper's five logit-normal timestep distributions and the loss weight curve; drag δ and draw a batch. The paper's own numbers are labeled; everything else is a toy.

Euler step Predictor error

The optimization recipe

Everything above is trained with a conventional, fully specified recipe. Every item in this table is from the paper's implementation section; the right column's arithmetic is ours.

SettingValueWhat it works out to
OptimizerAdamW, learning rate 10−4, betas (0.9, 0.95), weight decay 0.1Standard transformer settings
Gradient clipping1.0Caps the gradient norm per step
Global batch48Equal-sized batches from the action-labeled and actionless pools, losses weighted equally (presumably 24 + 24)
Precisionbfloat16Half-width floats with full dynamic range
EMAdecay 0.999Weights averaged over roughly 1 / (1 − 0.999) = 1,000 recent steps
Warmup48,000 samples, linear, then constant LR48,000 / 48 = 1,000 optimizer steps
Length1.2M optimizer steps, all methods1,200,000 × 48 = 57.6M training samples
EvaluationEvery 100,000 steps; report best checkpoint12 checkpoints per run, on 50 held-out initial conditions per task

An EMA (exponential moving average) keeps a slowly updated copy of the weights, each step moving it 0.1% toward the live weights; evaluating the EMA copy smooths out step-to-step noise. The equal-sized batch rule matters more than it looks: at D = 1,250, actionless demonstrations outnumber labeled ones 1,200 to 50, yet each optimizer step still sees half labeled data, so the action loss is never starved.

A reading note on "best checkpoint." Every from-scratch method in Tables I and II is scored by its best of twelve checkpoints, on the same held-out initial conditions it is reported on. Within that controlled study the rule is identical for all methods, so comparisons between them stay fair, but absolute numbers are best-case for each run. One exception to remember: Flex-π in Chapter 8 is reported at its final checkpoint, not its best one. The paper states both protocols plainly; we flag them so you read the numbers with them in mind.

Training and sampling, as code

def logit_normal(mu, sigma, n):
    return torch.sigmoid(mu + sigma * torch.randn(n))

TS    = {"dino": (-2, 1), "tracks": (0, 1), "depth": (-1, 1.6), "rgb": (-1, 1), "act": (-1, 1)}
BETA, DELTA = 0.5, 0.05

def training_step(batch):
    B = batch.size
    ctx, noisy, taus = {}, {}, {}
    for m in ORDER:                                          # tracks, dino, depth, rgb, act
        Y = batch.target[m]                                    # (B, N_m, c_m) clean future
        if m != "act":                                          # context copy + CONTEXT NOISE
            t_c = torch.empty(B, 1, 1).uniform_(1 - BETA, 1)
            ctx[m] = t_c * Y + (1 - t_c) * torch.randn_like(Y)
        tau = logit_normal(*TS[m], B).view(B, 1, 1)          # independent per stream
        noisy[m] = tau * Y + (1 - tau) * torch.randn_like(Y)        # Eq. (2)
        taus[m] = tau

    x_hat = net(batch.obs, batch.q, batch.task, ctx, noisy, taus,
                mask=BLOCK_CAUSAL)                                 # one pass, every stage
    loss = 0.0
    for m in ORDER:
        w = 1.0 / (1 - taus[m]).clamp_min(DELTA) ** 2              # Eq. (3) weight
        err = ((x_hat[m] - batch.target[m]) ** 2).sum((1, 2))
        per_ex = w.view(-1) * err
        if m == "act":
            per_ex = per_ex * batch.has_actions                    # actionless rows: no action term
        loss = loss + per_ex.mean()                              # Eq. (4), unit weights
    return loss

@torch.no_grad()
def act(obs, q, task, steps=8):
    cache = net.encode_context(obs, q, task)                    # KV for the observation
    for m in ORDER:
        y = torch.randn(SHAPE[m])                                   # tau = 0: pure noise
        for k in range(steps):
            tau = k / steps
            x_hat = net.denoise(m, y, tau, cache)                   # reads cached earlier blocks
            y = y + (x_hat - y) / (1 - tau) / steps              # Euler on v = (x_hat - y)/(1 - tau)
        if m == "act":
            return y                                                 # (16, 14) joint targets
        cache = net.append_context(cache, m, y)                     # re-embed, no context noise

Two details in that sketch are ours and two are the paper's. Ours: the shapes of the noise tensors and the exact cache API. The paper's: context noise is sampled independently per context block and per example, and it is switched off at inference.

What problem does ModAR's context noise solve, and when is it applied?

Chapter 6: Formulations × Data Scale

Five formulations, three data scales, six tasks, fifty trials each. That is 5 × 3 × 6 × 50 = 4,500 potential evaluation episodes behind one figure; Action-only exists only at the smallest scale, which removes 2 × 6 × 50 = 600 of them, leaving 3,900. This chapter reads that figure, its per-task table, and the two control experiments the authors ran to defend it against the obvious objections.

The simulation testbed

The benchmark is RoboTwin 2.0, a simulated bimanual manipulation benchmark with strong domain randomization. The paper picks "six representative RoboTwin tasks": dump bin, pick bottles, place bread, put bottles, stack bowls, and turn switch.

The data design is the clever part. Each task uses D ∈ {50, 250, 1250} total training demonstrations. Exactly 50 of them keep their action labels. The remaining D − 50 are used as actionless demonstrations: their actions are withheld, so they can only supervise futures. The number of action-labeled demonstrations therefore never changes. Only the amount of "video-like" data grows.

Total demos per task DAction-labeledActionlessRatio (actionless : labeled)
505000 : 1
250502004 : 1
1,250501,20024 : 1

"For each method and data scale, we train one multitask model for all six tasks," and every model is evaluated "on 50 held-out initial conditions per task." So each overall number is an average over 6 × 50 = 300 episodes. Because the labeled count is fixed, any improvement from D = 50 to D = 1,250 is attributable to actionless data. Action-only cannot use actionless data, so it has a result only at D = 50 (the figure draws it as a flat dashed line at 0.46).

Table I, in full

Success rates by formulation and scale, copied from the paper. Bold marks the highest value in each row.

TaskDAction-onlyIndep.-noiseDisjointUnifiedModAR
dump bin500.820.640.900.880.92
250—0.680.920.920.94
1250—0.640.780.940.98
pick bottles500.300.400.440.580.48
250—0.380.420.600.60
1250—0.380.380.620.58
place bread500.240.200.460.440.52
250—0.140.480.500.74
1250—0.280.220.440.72
put bottles500.320.200.360.520.64
250—0.200.360.620.82
1250—0.180.420.560.82
stack bowls500.640.360.580.760.76
250—0.180.620.760.68
1250—0.040.620.680.80
turn switch500.440.440.580.600.66
250—0.480.480.600.74
1250—0.460.480.580.64
overall500.460.370.550.630.66
250—0.340.550.670.75
1250—0.330.480.640.76

Notice what the paper claims and what it does not. It claims ModAR "achieves the highest average success rate at every data scale," which the overall rows confirm. It does not claim ModAR wins every task: Unified is ahead on pick bottles at 50 and 1,250 and on stack bowls at 250. The biggest ModAR margins are on place bread (0.74 vs 0.50 at 250) and put bottles (0.82 vs 0.62 at 250).

Worked example 9: rebuilding the overall row

The overall number is the plain mean of the six task numbers. Checking two cells by hand confirms the table is internally consistent and shows how much each task moves the average.

ModAR, D = 250:
  Sum = 0.94 + 0.60 + 0.74 + 0.82 + 0.68 + 0.74
      = (0.94 + 0.60) + (0.74 + 0.82) + (0.68 + 0.74) = 1.54 + 1.56 + 1.42 = 4.52
  Mean = 4.52 / 6 = 0.7533 → reported 0.75 ✓
Unified, D = 250:
  Sum = 0.92 + 0.60 + 0.50 + 0.62 + 0.76 + 0.60 = 1.52 + 1.12 + 1.36 = 4.00
  Mean = 4.00 / 6 = 0.6667 → reported 0.67 ✓
Gap = 0.7533 − 0.6667 = 0.0867, about 8.7 points. Of the 0.52 difference in sums (4.52 − 4.00), place bread contributes 0.74 − 0.50 = 0.24 and put bottles 0.82 − 0.62 = 0.20, so those two tasks supply 0.44 / 0.52 = 85% of the gap.

That last line is the kind of reading the paper's summary sentence hides. ModAR's average advantage is real, but it is concentrated in tasks where, plausibly, a precise sense of where objects will end up matters most; the paper itself makes the general point that "the best modality set naturally varies across tasks."

Reading the controls as a skeptic

A good habit with any controlled study is to list every difference between the compared systems and check which ones were equalized. Here is that list for ModAR versus Unified (the controlled items are from the paper; the "not separately tested" items are our observations).

Possible confoundStatus
Backbone, labeled data, optimization budgetEqualized by design
Sampling budget (40 vs 8 passes)Tested: baselines at 40 steps do not close the gap
Action predictor reading clean vs noisy futuresTested: a shared separate IDM preserves ModAR's lead
Checkpoint selectionSame rule for all (best of the evaluated checkpoints)
Timestep distributions (ModAR per-modality; Unified one shared (−1, 1))Not separately tested
Context noise (a ModAR-only ingredient)Part of the method; its removal is ablated within ModAR (75% to 63%)

The untested row is worth noticing without over-weighting: per-modality timestep distributions are also used by Independent-noise, which performs worst, so they are clearly not sufficient on their own for good results.

Why the labeled count is held at 50

Holding the action-labeled demonstrations fixed is what makes Table I a clean test of actionless data. If labeled data grew with D, any formulation would improve simply from more action supervision, and the scaling column would mix two effects. With 50 labeled demonstrations everywhere, the only way a formulation can improve with D is by turning actionless futures into better actions, which is precisely the WAM promise under test.

One consistency check links the two tables: ModAR's row in Table I (0.66, 0.75, 0.76) is identical to the last column of Table II, because both are the same four-modality model. When two tables in a paper agree like this, you can use either as a cross-check on the other.

Finding 1: ModAR is best at every scale

Overall: 0.66, 0.75, 0.76 at D = 50, 250, 1,250, against Unified's 0.63, 0.67, 0.64. The paper's explanation is the scratchpad hypothesis: "early modalities act as scratchpads for later ones: generating coarser or easier targets first provides structured context for more detailed targets."

Finding 2: ModAR turns actionless data into success

"ModAR benefits most from additional actionless data: adding 1,200 actionless demonstrations improves the average success rate from 66% to 76% (10 percentage points), compared with a mere 1% improvement for Unified." From the table: ModAR 0.66 → 0.76 is +0.10; Unified 0.63 → 0.64 is +0.01. Most of ModAR's gain arrives by D = 250 (+0.09); the next 1,000 actionless demonstrations add one more point.

The paper reports this gain without explaining it. One plausible reading (ours, not the paper's) connects it to the IDM framing from Chapter 4. Actionless data trains the future blocks, and in ModAR the action step reads the observation, configuration and task plus every finished future block. If actionless data makes those blocks better, the action step has better evidence to work from, while Unified's action stream only ever reads partially noisy futures. Control 2 below is consistent with that reading, since it shows ModAR's futures are themselves more useful for acting, but it does not test the mechanism directly.

Finding 3: Disjoint gets worse with more data

Disjoint beats Action-only on average (0.55 vs 0.46 at D = 50), suggesting that future prediction can help even as an auxiliary loss (by the rough error bars later in this chapter, a 9-point gap on 300 episodes is about two standard errors). But then it declines: 0.55, 0.55, 0.48. The paper's hypothesis is negative transfer: "with more actionless data, the shared representation becomes increasingly shaped by the future-observation prediction objective," to the detriment of the action objective, which in Disjoint never gets to read the futures that representation was shaped for. The steepest drops are place bread (0.48 to 0.22) and dump bin (0.92 to 0.78).

Finding 4: Independent-noise collapses

Independent-noise is the worst formulation at every scale (0.37, 0.34, 0.33), below even Action-only, and it also gets worse with more data. On stack bowls it falls from 0.36 to 0.04. Recall the paper's hypothesis: independently sampled noise levels rarely match the synchronized schedule used at test time, so the model gets "insufficient training signal near its inference regime."

How rarely? Here is a toy estimate (ours, not the paper's). Call a training example "inference-like" if all five streams' timesteps fall within 0.1 of each other. For five independent uniform draws, the probability that max − min ≤ w is 5w4 − 4w5.

With w = 0.1: 5 × 0.14 − 4 × 0.15 = 5 × 0.0001 − 4 × 0.00001 = 0.0005 − 0.00004 = 0.00046, about 1 example in 2,200.
With the paper's per-stream logit-normal distributions instead of uniform (our Monte Carlo, 2 million draws): about 0.0019, roughly 1 in 530.
With Unified's single shared timestep: every example is inference-like, probability 1.

The paper does not quantify this, and "within 0.1" is our arbitrary threshold. But the orders of magnitude make the hypothesis concrete: the Independent-noise model almost never practices the exact situation it faces at every inference step. The training view of Sim 2 in Chapter 1 shows the same thing one draw at a time.

Sim 7 · the Table I explorer

Lines show success versus total demonstrations (50 labeled in every case). Pick a task or the overall average. The bars below compare ModAR with Unified per task at the selected scale. Overlays (overall only): 40 Euler steps re-runs the simultaneous baselines with ModAR's sampling budget at D = 250; Separate IDM swaps both ModAR's and Unified's action predictors for one shared, separately trained inverse dynamics model. All values are the paper's. Tap a point to read it.

Bars at D =

Control 1: is ModAR winning only because it samples longer?

ModAR uses 8 Euler steps per stream: 40 in total across four futures and actions. Unified, Independent-noise and Disjoint use 8 in total. So the paper re-evaluated each of those three at D = 250 with 40 Euler steps, "matching ModAR's total number of sampling steps."

Formulation (D = 250)8 steps40 stepsChange
Unified67%59%−8
Disjoint55%54%−1
Independent-noise34%37%+3
ModAR (reference, 8 per stream)75%—

"Additional steps do not close the gap." Unified actually gets worse with more steps. The paper does not explain that drop; a finer integration grid means the model is evaluated at τ values it saw with different frequency during training, which is at least one candidate explanation, but it is our speculation. The conclusion stands either way: ModAR's advantage is not a sampling-budget artifact.

Control 2: are ModAR's futures actually better?

The second objection is subtler. ModAR's action step reads finished futures; Unified's reads half-finished ones. Perhaps ModAR's futures are no better than Unified's, and it only wins because its action predictor has the easier job.

The test: train a separate IDM "to map ground-truth future observations to actions," with context noise on its inputs during training. Then take the best ModAR and Unified checkpoints, let each generate its futures and actions normally at every prediction step, "discard its native action prediction, pass its predicted futures to the separate IDM, and execute the resulting actions." Now both systems use the same action predictor, and the only difference is the quality of the futures they hand it.

D (total demos)Unified + separate IDMModAR + separate IDMGap
500.550.59+0.04
2500.630.69+0.06
1,2500.610.70+0.09

(Values read from the bar labels of the paper's Figure 5.) "ModAR still outperforms Unified with the separate IDM, suggesting that its predicted futures are themselves more useful for predicting actions." Notice also that the gap widens with actionless data, echoing Finding 2. Both systems score a little lower with the separate IDM than with their native action heads (for ModAR, 0.69 vs 0.75 at D = 250). Our explanation, not the paper's: the native heads were trained jointly with the futures they read, while the separate IDM was trained on ground-truth futures.

What the two controls support together. ModAR's win survives equalizing the sampling budget and survives equalizing the action predictor. What remains is the thing the paper set out to test, and the paper's own word for the conclusion is "suggesting": generating futures in sequence, with each block conditioned on finished earlier blocks, appears to produce futures that are more useful for acting.

How much should you trust a 3-point or 8-point gap?

The paper reports no confidence intervals, so here is a rough binomial estimate of our own, treating each of the 300 episodes as an independent coin flip (it ignores task structure and best-checkpoint selection, so read it as a scale, not a verdict).

Standard error of one rate p over n = 300 episodes: √(p(1 − p)/n).
p = 0.75: √(0.75 × 0.25 / 300) = √(0.1875 / 300) = √0.000625 = 0.025.
ModAR 0.75 vs Unified 0.67: p(1−p)/n for Unified = 0.67 × 0.33 / 300 = 0.2211 / 300 = 0.000737.
  SE of the difference = √(0.000625 + 0.000737) = √0.001362 = 0.037; gap 0.08 is about 2.2 SE.
ModAR 0.75 vs Flex-π 0.72 (Chapter 8): 0.72 × 0.28 / 300 = 0.2016 / 300 = 0.000672.
  SE of the difference = √(0.000625 + 0.000672) = √0.001297 = 0.036; gap 0.03 is about 0.8 SE.

This is exactly why the authors word the Flex-π result as "slightly higher observed," and why the formulation claim rests on consistency across three scales, two controls and a real-robot replication rather than on any single cell.

Six tasks, six stories

The overall row is an average of six quite different curves. Reading them one by one (numbers from Table I, interpretations ours) shows where each formulation's character comes from.

Dump bin is the easy task: even Action-only reaches 0.82. ModAR climbs to 0.98 at D = 1,250. Disjoint is the only formulation that drops sharply here (0.92 to 0.78), a first sign of its negative-transfer problem.

Pick bottles is Unified's task: 0.58, 0.60, 0.62 against ModAR's 0.48, 0.60, 0.58. It is the clearest case in the table where joint generation is at least as good.

Place bread is ModAR's biggest win: 0.74 and 0.72 at the larger scales, against 0.50 and 0.44 for Unified. Placing an object precisely is a plausible place for a finished geometric future to pay off, and Chapter 7 shows depth-containing models do especially well here.

Put bottles is the cleanest data-scaling story: ModAR goes 0.64, 0.82, 0.82, while Independent-noise sits at 0.20, 0.20, 0.18.

Stack bowls is where Independent-noise collapses (0.36, 0.18, 0.04) and where Unified edges ModAR at D = 250 (0.76 vs 0.68) before ModAR retakes the lead at 1,250 (0.80 vs 0.68).

Turn switch shows ModAR peaking at D = 250 (0.74) and falling back at 1,250 (0.64), a reminder that "more actionless data" is not guaranteed to help on every task even for the best formulation.

Worked example: what a collapse looks like in episodes

Each task is evaluated on 50 held-out initial conditions, so a rate r means r × 50 successful episodes.
Independent-noise on stack bowls: 0.36 × 50 = 18 → 0.18 × 50 = 9 → 0.04 × 50 = 2 successes.
Disjoint on place bread: 0.48 × 50 = 24 → 0.22 × 50 = 11 successes from D = 250 to D = 1,250.
ModAR on put bottles: 0.64 × 50 = 32 → 0.82 × 50 = 41 successes from D = 50 to D = 250.

These are the per-task realities behind the averages: for two of the baselines, feeding in more actionless data makes the robot fail more often, while for ModAR it mostly makes the robot succeed more often.

Keep the same caution you applied to the averages. With 50 episodes per task, one rate carries a standard error of up to √(0.5 × 0.5 / 50) = √0.005 = 0.071, about 7 points. A 16-success collapse (Independent-noise on stack bowls, 18 to 2) is far outside that; a 2-point wobble on a single task is not. Read the per-task stories as patterns to check, not as verdicts.

A caution about generalizing the Independent-noise result

The paper is careful to say that Independent-noise "performs poorly in our from-scratch experiments." Keep that qualifier. Flex-π, which also samples noise levels independently per stream during training, reaches 72% on the same data in Chapter 8, starting from a pretrained video model. So the collapse in Table I is evidence about training this formulation from scratch at this scale, not a verdict on independent noise sampling in general. The train-test mismatch hypothesis may simply matter much less when the backbone already knows how to denoise video.

Figure 4(a), in words

The paper's figure draws each formulation as a colored bar at each of the three scales, with Action-only as a dashed horizontal line at 0.46. The visual story: the ModAR bars rise from left to right and are the tallest in every group; Unified is flat; Disjoint sags at the right; Independent-noise sits below the dashed line throughout. Sim 7 above redraws the same numbers as lines so the trends are easier to compare.

✎CheckThree numbers from memory▸

(1) ModAR's average success at D = 50, 250, 1,250. (2) What happened to Unified when given 40 Euler steps. (3) The separate-IDM result at D = 1,250.

Answers: (1) 0.66, 0.75, 0.76. (2) It dropped from 67% to 59%. (3) ModAR + IDM 0.70 versus Unified + IDM 0.61.

⚙DesignA lab with hours of video and a handful of teleop demos▸

A lab has 50 teleoperated demonstrations per task and a large pile of task videos without actions. Using only Table I, which formulation should it train, and which should it avoid?

One defensible answer: ModAR, because it is the only formulation whose average keeps rising as actionless data grows (0.66 → 0.75 → 0.76) and the controls show its gain is not a sampling-budget or action-head artifact. Avoid Independent-noise (0.37 → 0.33) and be wary of Disjoint (0.55 → 0.48): in this study both got worse as actionless data was added. Unified is a reasonable fallback but gains almost nothing from the videos (0.63 → 0.64).

Which experiment rules out the explanation "ModAR wins only because its action head reads fully denoised futures while Unified's reads partially noisy ones"?

Chapter 7: Which Futures Matter

Chapter 6 fixed the modalities and varied the formulation. This chapter does the opposite: it fixes the formulation (ModAR) and varies what it predicts. It is the chapter that decides whether the headline "RGB is not the right thing to imagine" survives contact with data, and whether each structured modality earns its place.

Go back to the prediction you wrote at the end of Chapter 0. Which single modality gives the best policy alone at 50 demonstrations, and which one hurts most to remove? Both answers are in this chapter.

Two ways to ask "does this modality help?"

There are two natural experiments, and the paper runs both.

The first is additive: start from one modality and add others one at a time, following the generation order (tracks, then DINO, then depth, then RGB). Alongside, train a variant for each modality predicted alone. That is Figure 4(b) and Table II, run at all three data scales.

The second is leave-one-out: start from the full four-modality model and remove exactly one modality. That is part of Figure 6, run at D = 250. The two experiments can disagree, and when they do, the disagreement is informative.

Table II, in full

Every row is one task and one scale; every column is a set of predicted modalities, all trained as ModAR. Bold marks the best value in each row. The RGB column is, in the paper's words, "the typical WAM setting of predicting only future RGB and actions."

TaskDAction-onlyRGBDepthTracksDINOT+DINOT+DINO+DepthT+DINO+Depth+RGB
dump bin500.820.800.900.900.960.920.920.92
250—0.940.920.920.940.940.960.94
1250—0.980.880.900.961.000.940.98
pick bottles500.300.180.440.120.440.580.480.48
250—0.320.480.480.620.620.580.60
1250—0.300.400.420.620.700.660.58
place bread500.240.260.380.100.440.520.540.52
250—0.340.640.260.560.640.760.74
1250—0.360.600.380.540.640.720.72
put bottles500.320.200.120.300.640.480.480.64
250—0.380.480.480.600.660.740.82
1250—0.540.420.740.640.660.900.82
stack bowls500.640.580.540.740.660.700.720.76
250—0.720.740.640.660.860.780.68
1250—0.660.640.660.840.740.800.80
turn switch500.440.580.360.300.460.580.620.66
250—0.560.540.420.540.580.680.74
1250—0.640.400.560.480.600.580.64
overall500.460.430.460.410.600.630.630.66
250—0.540.630.530.650.720.750.75
1250—0.580.560.610.680.720.770.76

The last column equals ModAR's column in Table I (0.66, 0.75, 0.76), as it must: it is the same four-modality model.

Reading the single-modality columns

Answer to the first prediction: at D = 50, DINO is the best single modality by a wide margin (0.60, against 0.46 for depth, 0.43 for RGB and 0.41 for tracks). It stays the best single modality at 250 (0.65) and 1,250 (0.68). A semantic representation of "what is where" seems to be the most useful single thing to imagine.

RGB alone is weak: 0.43 at D = 50, which is below Action-only (0.46). Imagining the future in pixels, with this small a model and no pretraining, does not help at all at the smallest scale; it only pulls ahead of Action-only once actionless data is added (0.54, then 0.58).

Tracks alone are the weakest single modality at D = 50 (0.41), with some striking task-level failures: 0.12 on pick bottles, 0.10 on place bread. Keep that in mind for the leave-one-out result below.

Worked example 10: the value of each addition

Following the cumulative columns and subtracting neighbors (our arithmetic on the overall rows):

D = 50:   tracks 0.41 → +DINO 0.63 (+0.22) → +depth 0.63 (+0.00) → +RGB 0.66 (+0.03)
D = 250:  tracks 0.53 → +DINO 0.72 (+0.19) → +depth 0.75 (+0.03) → +RGB 0.75 (+0.00)
D = 1250: tracks 0.61 → +DINO 0.72 (+0.11) → +depth 0.77 (+0.05) → +RGB 0.76 (−0.01)
Against the typical RGB-only WAM, the three structured modalities (T+DINO+Depth) add:
  D = 50: 0.63 − 0.43 = +0.20; D = 250: 0.75 − 0.54 = +0.21; D = 1250: 0.77 − 0.58 = +0.19.

Three things jump out. Adding DINO to tracks is the biggest single step at every scale. Depth's contribution grows with data (0.00, 0.03, 0.05). And RGB's contribution is +0.03, +0.00, −0.01: in the paper's words, "additionally predicting future RGB on top of the first three modalities provides no consistent gain."

The paper's summary of the additive experiment: "Across data scales, performance generally holds or increases with each added modality, showing that ModAR can effectively combine the benefits of predicting multiple modalities." And it is careful about the exceptions: "the best modality set naturally varies across tasks because different features define each task, and some modalities represent those features better than others; as a result, additional modalities do not always help."

You can see those exceptions in the table. Put bottles at 1,250 peaks at 0.90 with tracks, DINO and depth, and drops to 0.82 when RGB is added. Stack bowls at 250 peaks at 0.86 with only tracks and DINO. Dump bin at 1,250 reaches 1.00 with only tracks and DINO.

Why would RGB not help?

The paper's hypothesis is short: "predicting future RGB introduces high-variance appearance details while adding little information beyond the more structured targets in our setting." Unpack the two halves.

High variance. RoboTwin 2.0's own title advertises "strong domain randomization." The paper does not say which randomization settings its six tasks used, but wherever lighting, texture or clutter varies between episodes, the RGB target is where that variation lands, and the loss charges the model for failing to guess it. Little additional information. By the time RGB is generated, the model already has motion, semantics and geometry. What RGB adds on top is mostly appearance, which, on the paper's hypothesis, adds little beyond the structured targets for these tasks. Recall from Chapter 2 that RGB is also the largest raw target (225,792 numbers per example), so it is expensive to predict and, per this table, not worth it for these tasks.

The cost argument that decided the real-robot design. "Because adding RGB provided no consistent benefit in simulation while increasing training and generation costs, we train all real-world WAMs to observe and predict only tracks, DINO, and depth." The ablation was not just a scientific finding; it changed the deployed system.

The ablations: Figure 6

Figure 6 takes the full four-modality ModAR at D = 250 (average success 0.75) and changes one thing at a time. Values are read from the figure's bar labels and match the numbers stated in the text.

Variant (D = 250)SuccessChange vs fullWhat it tests
ModAR (full)0.75—Reference
No context noise0.63−0.12Robustness to its own imperfect earlier blocks (Chapter 5)
Reverse order (RGB → depth → DINO → tracks, actions still last)0.65−0.10The scratchpad ordering
w/o tracks0.61−0.14Value of point tracks
w/o DINO0.65−0.10Value of DINO features
w/o depth0.70−0.05Value of depth
w/o RGB0.750.00Value of RGB

Answer to the second prediction: removing tracks hurts most (0.75 to 0.61), followed by DINO (0.65), depth (0.70), and RGB (no change). The paper's conclusion: "predicting RGB contributes the least of the four modalities to success rate in this setting."

On order, the paper says reversing "lowers the average success rate to 65%, supporting our choice to generate compact, structured representations before increasingly detailed ones." On context noise: "context noise during training is critical for preventing errors from compounding across successive modalities."

Worked example: which ablations are clearly outside the noise?

Figure 6 reports six-task averages over 300 episodes each. Converting to episodes and applying the same rough binomial check as Chapter 6 (our arithmetic):

Full model: 0.75 × 300 = 225 successes; variance 0.75 × 0.25 / 300 = 0.000625.
w/o tracks: 0.61 × 300 = 183 (−42). Variance 0.61 × 0.39 / 300 = 0.2379 / 300 = 0.000793. SE of difference √(0.000625 + 0.000793) = √0.001418 = 0.0377. Gap 0.14 → 3.7 SE.
No context noise: 0.63 × 300 = 189 (−36). Variance 0.2331 / 300 = 0.000777. SE √0.001402 = 0.0374. Gap 0.12 → 3.2 SE.
w/o DINO and reverse order: 0.65 × 300 = 195 (−30). Variance 0.2275 / 300 = 0.000758. SE √0.001383 = 0.0372. Gap 0.10 → 2.7 SE.
w/o depth: 0.70 × 300 = 210 (−15). Variance 0.21 / 300 = 0.0007. SE √0.001325 = 0.0364. Gap 0.05 → 1.4 SE.
w/o RGB: 225 (0). Gap 0 → 0 SE.

On this rough scale, removing tracks, removing context noise, removing DINO and reversing the order are all well outside run-to-run noise; removing depth is suggestive but weaker; removing RGB changes nothing. That ranking matches the paper's prose exactly, including its choice to call context noise "critical."

The disagreement, and what it suggests

Put the two experiments side by side for tracks. Alone, tracks are the weakest modality (0.53 at D = 250). Removed from the full set, they cause the largest drop (−0.14 at D = 250).

One reading, consistent with the paper's scratchpad hypothesis, is that tracks are valuable less as a final answer and more as a first draft: compact motion context that makes every later block easier. A caveat the paper does not discuss, and which you should keep: removing tracks also changes which modality goes first (DINO becomes the first block), so the w/o-tracks ablation mixes "no tracks" with "a different first block." The reverse-order ablation points the same direction (putting tracks last costs 10 points), but neither experiment isolates position from content. The paper leaves "a full systematic comparison of modality orderings to future work."

Sim 8 · the modality lab

Toggle which futures ModAR predicts, pick a data scale, and optionally reverse the order or remove context noise. If the paper evaluated that exact configuration, you see its real numbers (per task from Table II, or the overall bar from Figure 6); if not, the lab says so and shows the closest evaluated neighbors. The leaderboard at the bottom ranks every configuration the paper evaluated at that scale.

D =

What the modality study does not say

It is worth being as careful as the paper is. The finding is that RGB gives "no consistent benefit" in this setting: small from-scratch models, six simulated tasks, a single camera. It is not a claim that RGB futures are useless in general. A model initialized from a video generator that already understands appearance might extract more from an RGB target, and the paper explicitly frames its conclusion as "a promising alternative to relying solely on future RGB prediction," not a replacement for it.

Similarly, the per-task variation is large. Anyone deploying this on a new task should expect the best subset to differ, and should budget for a small ablation of their own.

⚙DesignPick the modality set for a new robot▸

You are deploying a ModAR-style policy on a new single-camera arm with a tight latency budget. Using only this chapter's numbers, argue for a modality set and an order, and name the one experiment you would run first on your own tasks.

One defensible answer: tracks → DINO → depth, no RGB. Tracks + DINO is the largest additive step at every scale (+0.22, +0.19, +0.11); depth adds more as data grows (up to +0.05) and is the only explicit geometry with one camera; RGB adds +0.03 / 0.00 / −0.01 while being the largest raw target, which is exactly the paper's real-robot choice. First experiment: the w/o-depth ablation on your tasks, because depth's value (−0.05 when removed) is the least certain of the three you kept.

Per-task patterns worth noticing

The paper's general remark, that different tasks are defined by different features, is easy to see in specific cells (numbers from Table II; the interpretations are ours).

Put bottles at D = 50: DINO alone reaches 0.64 while depth alone reaches 0.12. With little data, knowing what is where seems far more useful than knowing how far for this task.

Place bread at D = 250: depth alone reaches 0.64 while tracks alone reach 0.26, and the best model (0.76) is the one that adds depth to tracks and DINO. Placement seems to reward explicit geometry.

Pick bottles: tracks alone go from 0.12 to 0.48 to 0.42 as data grows. Motion prediction on its own looks data-hungry.

Stack bowls at D = 1,250: DINO alone reaches 0.84, the best value in that row, above every multi-modality model.

Worked example: is "RGB hurts on put bottles" real?

At D = 1,250 on put bottles, adding RGB takes the model from 0.90 to 0.82. Tempting to conclude RGB hurts. A quick binomial check (ours) with 50 episodes per rate:

Successes: 0.90 × 50 = 45 versus 0.82 × 50 = 41, a difference of 4 episodes.
Variance of each rate: 0.90 × 0.10 / 50 = 0.0018; 0.82 × 0.18 / 50 = 0.1476 / 50 = 0.002952.
SE of the difference = √(0.0018 + 0.002952) = √0.004752 = 0.069.
Gap = 0.08, which is 0.08 / 0.069 = 1.16 SE.

A single-task gap of about one standard error is weak evidence either way. That is exactly why the paper phrases its RGB finding as "no consistent gain" across tasks and scales, rather than "RGB hurts." The honest reading is about the pattern: +0.03, 0.00, −0.01 on average, and zero change in the leave-one-out ablation.

Gains are real, but not strictly additive

Compare DINO's value in the two experiments at D = 250. Added on top of tracks, DINO is worth +0.19 (0.53 to 0.72). Removed from the full model, it costs 0.10 (0.75 to 0.65). If modality contributions simply added up, those two numbers would match. They do not, because what a modality contributes depends on what else is already predicted; plausibly, in the full model, depth and RGB partly cover for a missing DINO block. The paper's summary that tracks, DINO and depth "provide additive gains" is best read as "each one adds something on average," not as a claim that their values sum exactly.

Figure 4(b), in words

The paper's figure groups bars by scale. Within each group: single-modality models (RGB, depth, tracks, DINO), then the cumulative sets (tracks+DINO, +depth, +RGB), with Action-only as a dashed line at 0.46. The caption singles out the RGB bar as the one that "matches the typical WAM setting of predicting only future RGB and actions." In every group, the cumulative bars stand clearly above the RGB bar.

The scratchpad hypothesis: an evidence ledger

EvidenceDirectionStrength
Reverse order drops 0.75 to 0.65Supports structured-first orderingOne comparison, one scale
ModAR beats Unified at every scale, also with a shared IDMSupports sequential conditioning in generalConsistent across scales and a control
Adding DINO after tracks is the largest additive step (+0.22, +0.19, +0.11)Consistent with earlier blocks helping later onesAlso consistent with DINO simply being a strong target
Tracks weak alone, yet most costly to removeConsistent with tracks acting as contextConfounded with block position
Other orders never testedUnknownPaper leaves it to future work

The ledger is a fair summary of how the paper words it: the results "support" the hypothesis; they do not prove the chosen order is optimal.

⇆Cross-domain bridge
The tracks paradox is an old friend from feature selection
In classical machine learning, a feature can score poorly on its own yet cause a large drop when removed from a full model, because its value lies in interactions with other features. Univariate importance and leave-one-out importance answer different questions. Tracks in Table II and Figure 6 behave exactly like such a feature: weak as a lone target, important as part of the set. The lesson carries over directly: never judge a modality by its standalone column alone.
🔍Debug itA video WAM that loses to plain behavior cloning▸

A team trains a small from-scratch WAM that predicts only future RGB plus actions, with 50 demos per task, and it scores slightly below their action-only policy. Their conclusion: "world modeling does not help." What does Table II suggest instead?

Answer: the paper sees the same thing (RGB 0.43 vs Action-only 0.46 at D = 50), but DINO alone reaches 0.60 and tracks + DINO + depth 0.63 in the same setting. The failure is plausibly the choice of target, not world modeling itself; and RGB-only does pull ahead once actionless data is added (0.54, 0.58).

✎CheckThree quick readings of Table II▸

(1) At which data scale does depth first add something on top of tracks + DINO? (2) What is the best single modality at every scale? (3) Which column is the "typical WAM" and how far behind the tracks + DINO + depth model is it at D = 1,250?

Answers: (1) D = 250 (+0.03; at D = 50 it adds 0.00). (2) DINO (0.60, 0.65, 0.68). (3) The RGB column; 0.77 − 0.58 = 0.19 behind.

Two design details of the modality study

The cumulative sets follow the generation order. The additive columns grow as tracks, then tracks + DINO, then + depth, then + RGB: the same order the model generates them in. So each cumulative model is the previous one with one more block appended at the end of the sequence, just before the actions. That keeps the experiment clean: adding a modality never changes the position of the ones already there.

The leave-one-out ablations are run at D = 250 only. Figure 6 is explicitly "with 250 total demonstrations per task and 50 action-labeled demonstrations," the scale at which the full model reaches 75%. The paper does not say why that scale was chosen; the additive study in Table II is the one that covers all three scales.

What to carry into Chapter 8. Tracks, DINO and depth each earn their place; RGB does not in this setting; the order and context noise both matter. The real-robot system is built on those conclusions: it observes and predicts only tracks, DINO and depth, so the order is tracks → DINO → depth → actions with no RGB anywhere in the transformer (and, presumably, context noise as in simulation; the paper does not restate it for the real robots).

One last habit from this chapter: whenever a paper reports "modality X helps," ask measured how. Alone, added, or removed? At which data scale? Averaged over which tasks? Table II and Figure 6 answer all three, which is why their conclusions can be stated so precisely.

At D = 250, which statement matches the paper's modality results?

Chapter 8: A 6B Rival, and Real Arms

Everything so far compared ModAR with its own siblings: same size, same data, same budget. Two questions remain that a practitioner will ask immediately. Is a 30-million-parameter model trained from scratch even competitive with the large video-pretrained WAMs people actually use? And does any of this survive a real robot, real lighting, and human videos filmed by people rather than rendered by a simulator?

The rival: Flex-π

Flex-π is concurrent work by G. Yan, J. Liu, Y. Fan, L. Cai, M. Liao, J. Zhang and D. Fox, titled "A Multi-Stream World-Action Model with Compute Flexibility." It represents the pretrained-video-backbone approach Chapter 0 described, extended (as Chapter 0 also mentioned) to predict several future modalities, RGB latents, 3D pointmaps and DINO features, jointly with actions. It has about 6B parameters.

The paper fine-tunes it on the same RoboTwin data at D = 250, following Flex-π's own recipe as closely as the setting allows.

AspectFlex-π as run in this paperModAR (tracks–DINO–depth variant)
InitializationVideo backbone from Wan2.2-TI2V-5B; weights interpolated to a smaller hidden size for the action expertRandom (no pretraining)
What trainsAll trainable components, full fine-tuningEverything
What is frozenVAE, text encoder, DINOv3 encoderDINOv2 target encoder
Future targetsRGB latents, pointmaps, DINO features at t+4, 8, 12, 16Tracks, DINO, depth at t+8, 16
Action horizonH = 16, single head-camera settingH = 16, single camera
Actionless demosKeep every visual objective, mask the action lossSupervise available futures, drop the action term
Batch × steps288 × 30,000 (final checkpoint)48 × 1.2M
InferenceJoint denoising of all visual streams and actions (full multi-stream mode)Sequential, 8 Euler steps per stream
Parameters6B30.1M
Average success, D = 25072%75%

The paper's reading: "on these in-distribution tasks, a compact WAM trained from scratch can perform as well or better than a much larger video-model-initialized WAM using substantially less training compute." And, in the same paragraph, the caveat: "this is a system-level comparison rather than a controlled architectural comparison, since the models differ in scale, pretraining, target modalities, and training recipe." Chapter 6's binomial aside puts the 3-point gap at under one standard error, which is why "as well or better" is the right phrase.

One more asymmetry, which the table's "final checkpoint" note hints at: ModAR's 75% is its best evaluated checkpoint (the protocol from Chapter 5, best of twelve), while Flex-π's 72% is its final checkpoint after 30,000 steps. Picking the best of several evaluations can only raise a score, so this difference tilts a 3-point comparison toward ModAR. The paper states both protocols; it does not say how Flex-π would score under best-checkpoint selection.

Note also the phrase in-distribution. The evaluation uses held-out initial conditions of the same six tasks the models were trained on. It does not test the thing video pretraining is usually credited for, generalization to new objects, scenes and instructions. The paper's limitations section says as much.

Worked example 11: checking "about 20× fewer FLOPs"

The paper states the ratio (approximately 20×) but not its accounting. We can still check that it is plausible from numbers the paper does give. A FLOP is one floating-point operation; training compute is commonly estimated as roughly 6 × (parameters) × (tokens processed), the "6ND" rule of thumb.

Step 1. Samples ModAR trains on: 1,200,000 steps × 48 per batch = 57,600,000.
Step 2. Samples Flex-π trains on: 30,000 steps × 288 per batch = 8,640,000.
Step 3. ModAR sees 57,600,000 / 8,640,000 = 6.67× more samples.
Step 4. Parameter ratio: 6,000,000,000 / 30,100,000 = 6,000 / 30.1 = 199.3× (the paper's "approximately 200×").
Step 5. Under 6ND, compute ratio (Flex-π / ModAR) = parameter ratio × sample ratio × r, where r = Flex-π tokens per sample / ModAR tokens per sample:
  = 199.3 × (8.64 / 57.6) × r = 199.3 × 0.15 × r = 29.9 × r.
Step 6. If both models processed equally many tokens per sample (r = 1), the ratio would be about 30×.
Step 7. The paper's reported ~20× corresponds to r = 20 / 29.9 = about 0.67 under this crude rule.

So the reported ratio is the right order of magnitude, and the remaining factor is exactly what the rule of thumb cannot see: how many tokens each model processes per sample, that frozen components (VAE, text encoder, DINOv3) only cost a forward pass, and that Flex-π's action expert runs at a smaller width than its video backbone. We cannot reconstruct the paper's exact accounting, and this example does not pretend to. What it does show is the shape of the trade: ModAR trains on almost seven times more samples, but each sample costs roughly two hundred times fewer parameter-operations.

The compute story in one sentence. A video-pretrained WAM spends its compute on a huge model seeing relatively few samples; ModAR spends a twentieth of that compute on a tiny model seeing many samples, and on these tasks, with these targets, the tiny model keeps up.

Inference speed

"On a single NVIDIA GeForce RTX 5090 GPU, end-to-end inference for ModAR takes 147.9 ms (6.76 Hz) when generating all four future-observation modalities and actions."

Frequency: 1,000 ms / 147.9 ms = 6.761 Hz (the paper's 6.76 Hz).
Network passes per query: 8 Euler steps × 5 streams = 40.
Average time per pass including all overhead: 147.9 / 40 = 3.70 ms (our division; it lumps in embedding and caching work).
Each query produces a 16-step action chunk, so this cost is paid once per chunk, not once per control step.

The paper does not report the latency of the three-modality real-robot configuration or the robot's control rate, so we stop there. The honest summary is the paper's own: sequential generation "increases inference latency relative to simultaneous or action-only generation."

The real-world setup

Three "challenging tabletop tasks," performed with bimanual YAM arms: stacking cups, folding a crumpled towel, and placing an object in a drawer and closing the drawer. For each task the authors collect three kinds of data.

Data (per task)CountHow it was collectedWhat it can supervise
Robot demonstrations100Teleoperation with paired teacher arms, using the RAIDEN toolkitFutures and actions
In-domain human demonstrations200People performing the task in the same setting, filmed by the same fixed cameraFutures only (actionless)
EgoDex demonstrations1,000From the EgoDex dataset, categories "stack/unstack cups," "basic fold," "insert/remove drawer"Futures only (actionless, out of domain)

A fixed third-person ZED stereo camera "provides RGB observations and stereo depth for both robot and actionless human demonstrations." EgoDex is a large-scale egocentric video dataset of human dexterous manipulation; the paper calls its demonstrations "out-of-domain" relative to the robot's setting. As decided in Chapter 7, all real-world WAMs "observe and predict only tracks, DINO, and depth." Evaluation uses "30 rollouts per task, with varied initial object poses," and again one multitask model per method and data mixture.

Two practical details the paper does not spell out, which you would need to decide in a reimplementation: what the robot-configuration input qt is for human videos (which have no robot), and how depth is obtained for the EgoDex clips. We flag them rather than guess.

It helps to see what one human clip contributes, in the terms of Chapter 4. Its frames give an observation and targets at t+8 and t+16, so every future row it has targets for receives a loss (the paper's phrase is "the available future targets": tracks, DINO, and depth wherever depth exists for that clip). The action row receives nothing, because there is no action chunk to compare against. So every human clip is pure future supervision, and whatever it teaches has to reach the robot's arms through the futures the action step later reads.

Why vary the initial object poses across the 30 rollouts? Because a policy evaluated from one fixed start could succeed by replaying a memorized trajectory. Varying where the cups, towel or object begin forces the model to use what it sees, which is exactly where imagined motion, semantics and geometry should matter. (The paper states the protocol; the rationale is the standard one, and ours to spell out.)

Result 1: formulations on real hardware

Figure 7(a) compares Action-only (100 robot demos per task, since it cannot use actionless data) with Unified and ModAR (both with the 100 robot demos plus 200 in-domain human and 1,000 EgoDex demos). Per-task values are read from the figure's bar labels.

TaskAction-onlyUnifiedModAR
fold towel0.500.700.83
place in drawer0.570.700.83
stack cups0.500.600.83
overall52.2%66.7%83.3%

"ModAR achieves the highest success rate on all three tasks," and the ordering matches simulation: ModAR above Unified above Action-only.

Result 2: learning from human videos

Figure 7(b) trains three separate ModAR models: robot demos only; plus the 200 in-domain human demos; plus those and the 1,000 EgoDex demos. Human data supervises "future-observation predictions, but not action prediction."

TaskModAR (robot only)+ in-domain human+ in-domain + EgoDex
fold towel0.670.800.83
place in drawer0.730.800.83
stack cups0.700.830.83
overall70.0%81.1%83.3%

The paper's conclusion: ModAR "can benefit from actionless data collected across embodiments and domains."

Worked example 12: from rates back to rollouts

With 30 rollouts per task, every rate is a count out of 30. Recovering the counts (our arithmetic) makes the results tangible and checks the overall percentages exactly.

Action-only: 0.50 × 30 = 15; 17 / 30 = 0.567 (shown 0.57); 15. Total 15 + 17 + 15 = 47. 47 / 90 = 0.522 = 52.2%
Unified: 21 (0.70), 21 (0.70), 18 (0.60). Total 60. 60 / 90 = 0.667 = 66.7%
ModAR, full data: 25 / 30 = 0.833 on each task. Total 75. 75 / 90 = 0.833 = 83.3%
ModAR, robot only: 20 (0.67), 22 (0.73), 21 (0.70). Total 63. 63 / 90 = 0.700 = 70.0%
ModAR, + in-domain human: 24 (0.80), 24 (0.80), 25 (0.83). Total 73. 73 / 90 = 0.811 = 81.1%
Human data effect: 200 in-domain human demos per task buy 73 − 63 = 10 more successes out of 90; 1,000 EgoDex demos per task buy 75 − 73 = 2 more.

Two comparisons across the panels are worth making (both numbers are from the paper, the juxtaposition is ours). First, the like-for-like data comparison with Action-only: both trained on robot demos only, ModAR reaches 70.0% against 52.2%, a gap of 17.8 points without any human data at all. Second, ModAR without human data (70.0%) already edges out Unified with all of it (66.7%).

And the diminishing return from EgoDex (+2 successes from 1,000 out-of-domain demos, versus +10 from 200 in-domain ones) is a useful calibration for anyone planning data collection. With 90 trials per configuration, a 2-success difference is well within noise, so the fair reading is "in-domain human video helped clearly; out-of-domain video did not hurt and may have helped a little."

Sim 9 · compute and real rollouts

Compute: the 6ND sanity check from Worked example 11. Drag the token ratio r and watch the implied compute ratio move past the paper's reported ~20×. Real robots: pick a formulation and a data mixture; each dot is one of 30 rollouts per task. Only the combinations the paper ran are available. Dot order is arbitrary; only the counts come from the paper.

Token ratio r

Why the real-robot model does not even observe RGB patches

The paper's wording is precise: the real-world WAMs "observe and predict only tracks, DINO, and depth." So RGB is dropped on both sides, as a target and as an observation stream. The camera image is still used, because DINO features are computed from it by the frozen encoder; it simply never enters the transformer as raw patches. This saves the largest input projection and 192 observation tokens per query, and it follows directly from the simulation finding that RGB added no consistent benefit "while increasing training and generation costs."

A scale marker for training compute

The paper reports only the ratio (~20×), not absolute FLOPs. For a feel of the magnitude, suppose (our assumption) that the compared ModAR variant processes about 3,000 tokens per training example. Then the 6ND rule gives:

6 × 30.1M × 3,000 × 57.6M = 1.806 × 108 × 3,000 × 5.76 × 107
= 5.418 × 1011 × 5.76 × 107 ≈ 3.1 × 1019 FLOPs for ModAR,
and about 20 × that, ~6 × 1020, for Flex-π if the paper's ratio is applied.

Both absolute figures rest on our token assumption and the crude rule; only the ratio is the paper's.

A latency sketch for the real-robot configuration

The paper reports 147.9 ms for 40 passes (four futures plus actions). The real-robot model runs 32 passes (three futures plus actions). A crude proportional estimate, ours, and ignoring that the dropped RGB block is one of the larger ones:

147.9 ms × 32 / 40 = 147.9 × 0.8 = about 118 ms per 16-step chunk, if time scaled with the number of passes.

Treat that as a ballpark only; the paper measured one configuration on one GPU, and it lists the added latency of sequential generation among its limitations.

What would change the real-world conclusion

Three things the study does not rule out, worth keeping in mind (ours): the ranking could shift on tasks where appearance itself matters (reading a label, sorting by color), since the real-robot models do not predict RGB at all; a larger number of rollouts could shrink or widen the Unified gap, which is about 2.6 standard errors on our rough scale; and the human-video benefit was measured only for ModAR, so it is unknown how much Unified would gain from the same videos in isolation.

What the real-world study does and does not show

It shows that the simulation ranking of formulations reproduces on hardware, on three contact-rich bimanual tasks. It shows that in-domain human video clearly improves ModAR (70.0% to 81.1%). And it shows that adding 1,000 out-of-domain EgoDex clips per task gave a further 2 points (83.3%), which the paper reads as benefit "across embodiments and domains" and which, by our rough estimate below, is within noise.

It does not show generalization to unseen tasks or objects (three tasks, discrete labels, varied initial poses only), and each configuration rests on 90 rollouts. The paper does not report how Unified or Action-only would fare with human data ablated in the same way, beyond the configurations listed. These are exactly the limits Chapter 9 collects.

⇆Cross-domain bridge
Why human video suits this design in particular
A human hand and a robot gripper do not share an action space, which is why action-only policies cannot learn from people directly. Our reasoning, not the paper's: they do largely share the outcome. A stacked cup or a closed drawer is mostly embodiment-independent in tracks, DINO features and depth, though the hand or gripper itself still appears in all three. ModAR routes human video into the part of the model that predicts futures, and leaves the action step to the robot data, which is the only data with action labels. 3PoinTr, earlier work by three of the authors, is also about learning manipulation from unconstrained human videos, via 3D point tracks (as its title says).

Per task: where ModAR's real-world margin comes from

Using the recovered counts (out of 30 rollouts per task), the gaps become concrete.

TaskAction-onlyUnifiedModARModAR over Unified
fold towel152125+4 rollouts
place in drawer172125+4 rollouts
stack cups151825+7 rollouts

Stacking cups is where Unified struggles most and ModAR's margin is largest. And the human-video gains per task, from Figure 7(b): fold towel 20 → 24 → 25, place in drawer 22 → 24 → 25, stack cups 21 → 25 → 25. In-domain human video helps every task; EgoDex adds at most one more success on any task.

Worked example: how solid are the real-world gaps?

A rough binomial check with 90 rollouts per configuration (our estimate, treating rollouts as independent):

SE of ModAR's 0.833: √(0.833 × 0.167 / 90) = √(0.1391 / 90) = √0.001546 = 0.039.
SE of Unified's 0.667: √(0.667 × 0.333 / 90) = √(0.2221 / 90) = √0.002468 = 0.050.
ModAR vs Unified: gap 0.167; SE of difference √(0.001546 + 0.002468) = √0.004014 = 0.063; ratio 2.6 SE.
SE of Action-only's 0.522: √(0.522 × 0.478 / 90) = √(0.2495 / 90) = √0.002772 = 0.053.
ModAR vs Action-only: gap 0.311; SE of difference √(0.001546 + 0.002772) = √0.004318 = 0.066; ratio 4.7 SE.
EgoDex step (0.811 to 0.833): SE of 0.811 = √(0.811 × 0.189 / 90) = √0.001703 = 0.041; SE of difference √(0.001703 + 0.001546) = √0.003249 = 0.057; gap 0.022 is 0.4 SE.

So, on this rough scale: ModAR over Action-only is very clear, ModAR over Unified is reasonably clear, and the EgoDex increment is indistinguishable from noise. The paper's wording matches: the formulation result "mirror[s] the simulation results," and human data "improves" ModAR, with the big step coming from in-domain video.

What "interpolating to a smaller hidden size" means

Flex-π's recipe initializes its action expert from the video backbone even though the action expert is narrower. The paper describes this as interpolating the pretrained weights "to a smaller hidden size for the action expert." The general idea is to resample each pretrained weight matrix down to the narrower width so the action expert starts from something video-shaped rather than from random numbers. The exact procedure belongs to Flex-π's own paper; the point for us is the contrast: Flex-π's video backbone and action expert start from pretrained video weights, while every parameter of ModAR starts from scratch. (The paper does not say whether Flex-π's smaller input and output layers, for example for pointmaps or actions, are also pretrained.)

What the Flex-π comparison can and cannot tell you

It can tell youIt cannot tell you
A 30.1M from-scratch WAM can match a 6B video-initialized WAM on these six in-distribution tasks at D = 250Whether sequential generation beats joint generation when both are pretrained
The compact model got there with roughly 20× less training computeHow either model behaves on unseen tasks, objects or scenes
Structured targets without RGB are enough to compete with a model that also predicts RGB latentsWhich of scale, pretraining, targets or recipe explains the 3-point gap (all four differ)

One more difference hides in the table above: ModAR's DINO targets come from a frozen DINOv2 encoder, while Flex-π keeps a frozen DINOv3 encoder. The two models are not even predicting the same feature space, which is one more reason the paper calls this a system-level comparison.

✎CheckRebuild the real-world numbers▸

Without scrolling: (1) the three real tasks, (2) the three data sources and their counts per task, (3) the overall success of Action-only, Unified and ModAR, (4) ModAR's overall success with robot data only, with in-domain human video, and with EgoDex added.

Answers: (1) stacking cups, folding a crumpled towel, placing an object in a drawer and closing it. (2) 100 teleoperated robot demos, 200 in-domain human demos, 1,000 EgoDex demos. (3) 52.2%, 66.7%, 83.3%. (4) 70.0%, 81.1%, 83.3%.

⚙DesignSpend a data-collection week▸

You have one week for data collection on a new bimanual task and already have 100 teleop demos. Options: record 200 in-domain human demos, or download 1,000 matching clips from a public egocentric dataset. Using only Figure 7(b), which do you prioritize?

One defensible answer: the in-domain human demos. In the paper they moved ModAR from 63 to 73 successes out of 90; the 1,000 EgoDex demos added 2 more on top, which is within noise. Out-of-domain video is cheap and did not hurt, so add it if it is free, but the in-domain recordings are where the measured gain was.

In the real-world experiments, what does adding human demonstrations do, and how are they used?

Chapter 9: Limits, Connections, and What to Read Next

You started with a robot that imagined its future in pixels, paying for lamp glints and tablecloth patterns. You now know a design that imagines motion first, then meaning, then distance, and only then acts, and you have checked its evidence line by line. This last chapter draws the boundary of what that evidence covers, connects ModAR to the rest of the site, and points to what to read next.

The whole argument, in six lines

1 · Problem
WAMs usually imagine the future as RGB, which does not prioritize the geometric, physical and semantic structure manipulation needs.
↓
2 · Idea
Predict several structured futures (tracks, DINO, depth, optionally RGB), one block at a time, each conditioned on the finished ones, then actions as an inverse dynamics step.
↓
3 · Machinery
A small DiT (30.1M parameters in its tracks–DINO–depth form) with shared blocks and modality experts, adaLN conditioning, a two-copy block-causal training sequence, JiT-style x-prediction, and context noise (β = 0.5).
↓
4 · Controlled evidence
Best average success at every data scale (66 / 75 / 76%), biggest gain from actionless data (+10 points), robust to equal sampling budgets and to a shared IDM.
↓
5 · Modality evidence
Tracks, DINO and depth each help (removing them costs 14, 10, 5 points); RGB adds no consistent benefit; order and context noise matter (−10, −12).
↓
6 · Beyond the sandbox
75% vs 72% against a 6B video-pretrained WAM at about 20× less training compute; 83.3% on three real bimanual tasks, improved by human video.

The limitations the paper states

The paper's limitations section is short and specific. Here it is, point by point, with what each one means for how far you can carry the results.

Limitation (paper)What it means in practice
"A limited number of tasks"Six simulated tasks and three real ones. ModAR has the best six-task average at every scale and the best score on all three real tasks, but Unified wins some individual simulated cells (pick bottles, stack bowls); nine tasks are not a population.
"Discrete task labels rather than language instructions"The task embedding g is a lookup table. How the design interacts with language conditioning, which most generalist policies use, is untested.
No "broad generalization across tasks, objects, or scenes"Evaluation varies initial conditions and object poses within trained tasks. The Flex-π comparison is explicitly "in-distribution."
Sequential generation "increases inference latency"40 network passes per query in the four-modality setting (147.9 ms on an RTX 5090) versus 8 for simultaneous formulations.

And the future work it proposes: "larger-scale experiments with broader task diversity, generalization evaluation, and training on heterogeneous internet-scale actionless data," plus exploring "the best order for generating modalities, which may depend on the task or specific scenario."

Caveats stated elsewhere in the paper

Three more boundaries appear in the body text rather than the limitations section, and they matter just as much.

The Flex-π comparison is system-level. The models "differ in scale, pretraining, target modalities, and training recipe," so the comparison says a compact from-scratch model can keep up on these tasks, not that sequential generation beats a pretrained backbone in general.

Order was tested only against its reverse. "We leave a full systematic comparison of modality orderings to future work." The chosen order beat its reverse by 10 points; whether some third order would beat both is unknown.

The RGB conclusion is scoped. RGB gave "no consistent benefit in our experiments," with from-scratch models. The conclusion frames the result as "a promising alternative to relying solely on future RGB prediction."

Open questions we would ask next

These are ours, not the paper's. Each follows from something the paper leaves unmeasured.

Where ModAR sits among world models

FamilyImaginesActs howExample on this site
Policies (no imagination)NothingObservation → action chunkDiffusion Policy, π0, ACT
Action-conditioned world modelsFuture video given actionsA planner or evaluator chooses actionsDreamX-Phi, Genie, World-In-World
Video-first WAMsFuture RGB (latents), jointly or causally with actionsCo-denoised or causal action headsCausal World Modeling for Robot Control (LingBot-VA)
Structured-future models3D points, tracks, featuresVariesPointWorld
ModARTracks → DINO → depth (→ RGB), in sequenceInverse dynamics on the finished futureThis lesson

The Causal World Modeling paper in that table (LingBot-VA on this site) is one of the works ModAR cites as predicting future RGB as image latents from frozen video VAEs, which makes it a good contrast read.

How the simulations in this lesson relate to the paper

SimPaper elementReal numbersToy parts
1 · four modalitiesSec. I, II-A motivation; Fig. 1NoneScene, renderer, DINO colors
2 · formulationsFig. 3; IV-A baselines; IV-E timestep settingsStep counts; logit-normal parametersAnimation timing
3 · horizon and tokensIII-A; IV-E inputs and outputsH, Δ, J, grid, patch size, action sizeTrack token layout (3 numbers)
4 · routing and budgetFig. 2; IV-E architectureBlock counts, widths, 30.1MadaLN mapping; 12d2 estimate
5 · maskIII-B block-causal generationThe prediction-copy rule; orderObservation and context-copy rows; 576 observation tokens
6 · flow labIII-B training and inference; IV-E flow settings; Fig. 6Equations, δ, β, timestep parameters, 75% vs 63%Signal, predictor error, cascade gains
7 · Table I explorerTable I; Figs. 4(a), 5; sampling-step controlAll plotted valuesNone
8 · modality labTable II; Figs. 4(b), 6All plotted valuesNone
9 · compute and rolloutsIV-B Flex-π; Fig. 7Parameters, steps, batch sizes, ~20×, success ratesToken ratio r; dot order

Where to go next on this site, by goal

Lessons on this site that connect

Papers to read next (from ModAR's reference list)

PaperWhy read it after ModARLink
Latent Forcing (Baade et al., 2026)The source of the scratchpad ordering and of context noisearXiv:2602.11401
Modality Forcing for Scalable Spatial Generation (Duisterhof, Ramanan, Ichnowski, Johnson, Park, 2026)Joint image and depth generation from an image prior; shares three authors with ModARarXiv:2606.13676
Flex-π (Yan et al., 2026)The 6B multi-stream rival from Chapter 8arXiv:2608.10860
Fast-WAM (Yuan et al., 2026)The Disjoint formulation: do WAMs need test-time imagination?arXiv:2603.16666
World Action Models are Zero-shot Policies (DreamZero; Ye et al., 2026)A joint-generation WAM the Unified baseline representsarXiv:2602.15922
Turning Video Models into Generalist Robot Policies (VERA; Li et al., 2026)Futures-then-actions with a video modelarXiv:2605.27817
Point Tracking Improves World Action Models (Guan et al., 2026)Tracks alongside RGB in a WAMarXiv:2605.23856
3PoinTr (Hung, Duisterhof, Ichnowski, 2026)3D point tracks for learning from unconstrained human videosarXiv:2603.08485
WAM4D (Li et al., 2026)Depth and 4D structure in a fast WAMarXiv:2606.14048
ST-WAM (Wang et al., 2026)Semantic-temporal WAM robust to visual distribution shiftarXiv:2607.28993
EgoWAM (Li et al., 2026)WAMs beyond pixels with in-the-wild egocentric human dataarXiv:2607.08436
RoboTwin 2.0 (Chen et al., 2025)The simulation benchmark behind Tables I and IIarXiv:2506.18088
Back to Basics: Let Denoising Generative Models Denoise (JiT; Li and He, CVPR 2026)The x-prediction objective ModAR trains withCVPR 2026
Unified World Models (Zhu et al., RSS 2025)Coupled video and action diffusion with independent noise levelsRSS 2025
EgoDex (Hoque et al., ICLR 2026)The out-of-domain human video used in Chapter 8ICLR 2026

Where each part of the paper lives in this lesson

Paper sectionLesson chapter
Abstract, I Introduction, Fig. 10
II-A World-action models, Fig. 31
III-A Problem definition; III-B Modality tokenization; IV-E Inputs and outputs2
III-B Shared world-action backbone, Fig. 2; IV-E Architecture3
II-B Multimodal generation; III-B Block-causal modality generation, Equation (1), Inference4
III-B Injecting context noise, Training, Equations (2)–(4); IV-E Optimization, Flow and sampling5
IV-A Simulation setup, Baselines; IV-B Formulation comparison, Scaling, Separate IDM, Sampling steps; Table I, Figs. 4(a), 56
IV-B Modality comparison; IV-C Ablations; Table II, Figs. 4(b), 67
IV-B Flex-π comparison, Inference latency; IV-A Real world; IV-D; IV-E Real-world data collection; Fig. 78
V Conclusion, VI Limitations, related work9

Questions to ask the next world-action model paper you read

ModAR's controlled design gives you a checklist for reading its neighbors critically.

Common mistakes when explaining ModAR

MistakeCorrection
"It predicts tracks for all 16 future steps."Futures are sparse: J = 2 frames, at t+8 and t+16. Only actions are dense (16 steps).
"It is autoregressive token by token."It is autoregressive across modality blocks; each block's tokens are denoised jointly.
"Context noise is added at test time for robustness."It is applied only during training.
"The DINO encoder is fine-tuned."DINOv2 is frozen; it defines the target space.
"The paper shows RGB hurts."It shows RGB gives no consistent benefit in its from-scratch setting.
"ModAR beats Flex-π."It is slightly higher (75% vs 72%, best checkpoint against final checkpoint) in a system-level, in-distribution comparison, with far less compute.
"ModAR wins because it samples five times longer."Baselines given 40 steps do not close the gap.
"Human videos train the action head."They supervise futures only; the action term is dropped for actionless examples.

ModAR in three sentences, for three readers

For a practitioner: if you have few robot demonstrations and lots of task video, a small WAM that predicts point tracks, DINO features and depth in that order, then actions, is a strong, cheap recipe; skip RGB unless your task depends on appearance, and do not skip context noise.

For a researcher: the paper cleanly separates formulation, target representation and actionless-data scale in a from-scratch study, finds sequential ("modality-autoregressive") generation best at every scale with two controls, and leaves open ordering search, pretrained sequential models and broad generalization.

For a student: a robot does better when it first imagines where things will move, what they are and how far away they are, one step at a time, and only then decides how to move its arms.

Experiments the paper invites (ours)

QuestionMinimal experiment
Is the order optimal?Train all 24 orderings of the four futures at D = 250 (or a sensible subset) with everything else fixed
Position or content?Keep tracks but move them to the second or last slot; compare with w/o tracks
How much context noise?Sweep β in {0.25, 0.5, 0.75, 1.0}
Pretrained and sequential?Initialize a sequential WAM from a video model and compare with Flex-π under matched compute
Language instead of labels?Replace the task-embedding table with a text encoder and evaluate on held-out instructions

Reimplementation checklist

If you set out to rebuild ModAR from this lesson, here is every specification, split by where it comes from. The right column is honest about what you would have to decide yourself.

ComponentSpecified by the paperYou must choose (not in the paper)
InputsSingle camera, 168×224; qt = 14-D absolute dual-arm joint configuration; learned task embeddingqt for human videos
TargetsH = 16, Δ = 8, J = 2 (t+8, t+16); 16-step, 14-D action chunk; replan every H steps—
Tokens12×16 grid of 14×14 patches; patchified RGB and depth; frozen DINOv2 ViT-S/14 patch tokens; CoTracker3 tracks of patch-center queries with displacement + visibility; linear projection to a common width; learned modality embedding; axial RoPE (time, row, column) for visual, time for actionsTrack-token normalization and occluded values; RoPE dimension split; exact observation-token layout
BackboneDiT: 6 shared width-384 blocks, 6 heads; experts 2×384 (DINO, depth, RGB), 1×128 (tracks), 2×128 (actions); linear heads; adaLN from qt, g, per-stream τ; 30.1M total for the tracks–DINO–depth variantMLP ratio; adaLN wiring; the 384-to-128 adapters; the four-modality parameter count
Ordering and masktracks → DINO → depth → RGB → actions; clean context copy + noisy prediction copy; block-causal mask (prediction copy of mk reads observation + context of m<k)Rows of the mask for context copies and the observation
ObjectiveLinear interpolant; JiT-style x-prediction; loss ‖Ŷ − Y‖2 / max(1 − τ, δ)2, δ = 0.05; unit loss weights; action term only with labelsLoss reduction (sum versus mean over tokens)
TimestepsLogit-normal (μ, σ): DINO (−2, 1), tracks (0, 1), depth (−1, 1.6), RGB (−1, 1), actions (−1, 1)—
Context noiseτctx ~ U(1 − β, 1), β = 0.5, independent per context block and example; training onlyτ signal given to context at inference
OptimizationAdamW, LR 10−4, betas (0.9, 0.95), WD 0.1, clip 1.0, batch 48 (equal labeled/actionless), bfloat16, EMA 0.999, 48,000-sample warmup then constant, 1.2M steps; evaluate every 100k, report best—
Sampling8 Euler steps per stream; v = (Ŷ − Ỹ)/(1 − τ); KV cache for observation and finished blocksStep spacing
Real robotYAM arms; RAIDEN teleop; ZED stereo RGB + depth; tracks, DINO, depth only; 100 robot + 200 in-domain human + 1,000 EgoDex demos per taskDepth source for EgoDex clips

The right column is short, which is a compliment to the paper: almost everything that determines the results is written down.

Glossary

TermMeaning in this lesson
WAMWorld-action model: one network that jointly models future observations and robot actions
ModalityAny representation of the future: RGB, depth, DINO features, point tracks
Actionless demonstrationA demonstration without action labels; supervises futures only
StreamOne set of tokens generated together (one modality's future, or the action chunk)
Flow timestep τHow clean a stream is: 0 = pure noise, 1 = clean
x-predictionThe network outputs its guess of the clean target; the sampler converts it to a velocity
Block-causal maskAttention rule letting each block read only the observation and earlier finished blocks
Context copy / prediction copyThe clean (lightly noised) and noisy versions of each target, both present in one training sequence
Context noiseTraining-only corruption of context copies, so later blocks tolerate imperfect earlier generations
Inverse dynamics model (IDM)A map from (now, future) to the actions that cause the change; ModAR's final step
Unified / Independent-noise / Disjoint / Action-onlyThe paper's baseline formulations (joint; joint with per-stream training noise; no future-action attention; no futures)
DTotal demonstrations per task in simulation (50 labeled + D − 50 actionless)

Where the field is heading, per the paper's related work

The paper situates itself in a fast-moving 2026 literature, and its reference list is a good map of the open directions.

Richer structured futures. WAM4D uses spatial register tokens for fast 4D prediction; 3PoinTr lifts tracks to 3D; ST-WAM targets semantic-temporal features for robustness to visual distribution shift. Each is a different bet on which structure a policy should imagine.

Human video at scale. EgoWAM trains beyond pixels on in-the-wild egocentric human data; EgoDex supplies large-scale egocentric manipulation video. ModAR's small EgoDex gain is one data point in a question these works are actively probing.

How much imagination is needed at run time. Fast-WAM asks whether test-time future imagination is needed at all. ModAR's results suggest that, in its setting, generating the right futures in the right order improves success, at a latency cost the paper lists as a limitation.

Pretrained generators. DreamZero, Cosmos Policy, VERA and Flex-π all build on video models. The natural synthesis, a pretrained model that also generates structured modalities in sequence, is the experiment this paper leaves for someone to run.

A final self-test

Try these three without looking back; each takes one or two sentences.

1. A colleague proposes generating RGB first "because it contains everything." Which experiment in the paper speaks to this, and what did it find? Check: the reverse-order ablation (RGB → depth → DINO → tracks) dropped success from 75% to 65%.

2. Why can ModAR learn from a human video but Action-only cannot? Check: ModAR supervises its future blocks on actionless data and drops only the action term; Action-only has no future objective, so an actionless example gives it nothing to learn.

3. What single number best summarizes why context noise matters? Check: without it, success at D = 250 falls from 75% to 63%, below Unified's 67%.

Exit gate: teach it back before you leave.

Without scrolling up: (1) draw the five formulations and say what differs between Unified and Independent-noise; (2) write ct, Ytm and At with the paper's H, Δ and J, and count the future tokens per modality; (3) write Equation (1) and explain why the last factor is an inverse dynamics model; (4) derive v = (Ŷ − Ỹ)/(1 − τ) and explain the δ clamp in Equation (3); (5) explain what context noise protects against and quote its ablation number; (6) name the two controls that defend Table I, and the RGB result that shaped the real-robot model. If any of the six stalls, its chapter is one tap away.

The takeaway to keep. What a robot imagines matters, and so does the order in which it imagines it. In this paper's setting, imagining motion, then meaning, then distance, and acting last on the finished picture beat the four alternative couplings the paper tested, gained the most from extra actionless demonstrations in simulation, improved with human video on real robots, and let a 30.1M model keep pace with a 6B one.
Which statement correctly describes a limitation the paper itself acknowledges?