Most world-action models imagine the future as video, paying for lighting and texture a robot never needed. ModAR imagines it one modality at a time (where things move, what they are, how far away they sit) and only then decides what the arms should do. A 30.1M-parameter model trained from scratch keeps pace with a 6B video-pretrained one (75% vs 72% observed).
Two robot arms sit on opposite sides of a table. Between them are two cups, and the job is to stack one inside the other. A camera watches from a fixed spot across the table. Every so often the robot's policy has to answer one concrete question: what should my fourteen joint values be for each of the next sixteen control steps?
(A note on sourcing before we start. Every number, table value, dataset, and design detail in this lesson comes from the paper: Hung, Duisterhof, Ramanan and Ichnowski, Modality-Autoregressive World-Action Models, arXiv:2609.17524. When we use toy numbers to make a mechanism visible, or do our own arithmetic on the paper's numbers, the text says so plainly.)
A fast-growing family of robot models answers that question in two moves. First it imagines: it predicts what the scene will look like a moment from now. Then it acts: it picks the joint commands that make that imagined future happen. The imagining is not decoration. The paper's first sentence calls future-observation prediction "a powerful objective for learning rich representations of the world's dynamics and semantics."
Here is the uncomfortable part. When most of these models imagine, they imagine video. They predict the future as RGB images, usually as compressed image latents produced by a frozen video autoencoder. To do that well, the model has to commit capacity to the exact brown of the table, the glint of the lamp on the cup rim, the soft shadow the forearm throws. In a scene like this, none of those details changes what the arm should do next.
Think about how you would plan the same stack. You would not rehearse the colors. You would think: this cup goes up and over; it is a cup, not a bowl; it is about a hand's width in front of the other one. Motion, identity, distance. That is the whole idea of this paper, stated as a hunch. The rest of the lesson is the machinery and the evidence.
Start with an analogy. A climber on a wall looks at the next three holds and, before moving, plays the sequence in their head: hand here, weight shifts there, foot comes up. The mental rehearsal and the movement are one skill, not two. A climber who rehearses badly also climbs badly, and practicing the rehearsal makes the climbing better.
The robot version of that climber has a name. A world-action model (WAM) is a single network that jointly models future observations and the corresponding robot actions. "World" is the part that predicts what the scene will look like; "action" is the part that predicts what the robot should do. The two share one set of weights, so whatever the network learns about how the world moves is available when it chooses how to move.
A WAM is different from two neighbors you may know. A plain policy (for example a diffusion policy) maps the current observation straight to actions and never predicts the future. A plain world model (see the World Models gleam) predicts the future given actions you supply, and leaves the choice of action to a planner. A WAM predicts both, in one model, conditioned on what the robot sees right now.
The deepest reason WAMs exist is about data, not about elegance. To learn actions, you need action-labeled demonstrations: recordings where every frame comes with the exact joint commands the robot executed. Those come from teleoperation. Someone has to drive the robot, one demonstration at a time. They are slow and expensive to collect.
To learn what the future looks like, you need only video. An actionless demonstration is a recording of the task being done with no action labels attached: a person stacking cups with their own hands, for instance. You cannot train an action head on it, because there are no actions to copy. But you can train a future-prediction head on it, because the future frames are right there in the recording.
The paper names two ways that extra future supervision can help action prediction. First, as a training-time auxiliary objective: the future-prediction loss shapes the shared representation, even if the model never generates the future at deployment (the Fast-WAM line of work). Second, by letting the policy condition its action generation on predicted visual futures (the UniPi and VERA line of work). ModAR is firmly in the second camp: the paper calls it the first WAM to denoise several future modalities one after another before it predicts the actions.
If RGB is not the only way to imagine the future, what are the alternatives? The paper uses the word modality in a deliberately broad sense: "a representation of the future, including both sensory signals and derived features." Depth is a sensory signal. DINO features are a derived feature. Both count as modalities here.
ModAR works with exactly four of them. The paper writes the set as M = {rgb, depth, dino, tracks}. Each one keeps some facts about the scene and throws others away, and the paper argues that each contains "unique inductive biases that capture manipulation-relevant features." An inductive bias is a built-in preference: a representation that makes some patterns easy to express and others impossible.
| Modality | What one future frame holds | What it makes explicit (paper) | How ModAR gets it |
|---|---|---|---|
| RGB | Color pixels | Full appearance: color, texture, lighting, and everything else | Patchified image |
| Depth | Distance from the camera per pixel | Scene geometry; direct supervision for the spatial reasoning 3D actions need | Patchified depth map (stereo depth on the real robot) |
| DINO | A feature vector per image patch | Object semantics and scene structure; robust to appearance variation | Spatial patch tokens of a frozen DINOv2 encoder |
| Point tracks | Where each tracked point moved, and whether it is visible | Scene motion and correspondence, independent of appearance | CoTracker3 tracking a grid of query points |
Read the third column slowly, because it is the argument in miniature. Point tracks say how task-relevant parts of the scene move through time. A track does not care whether the cup is red or blue. DINO features, the patch embeddings of a self-supervised vision transformer (DINOv2), say "this patch is part of a cup" in a way that survives a change of lighting. Depth says how far away every surface is, which is exactly what a 3D reach needs. RGB says all of that too, but buried under appearance.
The simulation below makes the difference physical. It renders one toy tabletop scene four ways. The toy renderer is ours: the colors in the "DINO" panel are a cartoon of semantic grouping (real DINO features are 384-number vectors, not colors), and the scene is drawn from simple shapes. But the behavior it demonstrates, which panels react to appearance changes and which do not, is exactly the property the paper leans on.
A gripper carries a cup to the right. Scrub the future step to move from now (t) to t+16. Then drag lighting to move the lamp and change its strength, and swap the table texture. Watch the readout: appearance changes rewrite the RGB future but leave the motion (tracks), the geometry (depth), and the cartoon semantics (DINO) essentially untouched. Hollow dots in the tracks panel are points the cup has covered, which a real track marks as not visible.
Two things should stand out after a minute of play.
First, the lighting slider moves a lot of RGB "mass" while the future event (a cup moving right) stays identical. A model trained to predict RGB is graded on that mass. If the lamp in the training videos flickers, or the tablecloth changes between demonstrations, the RGB loss punishes the model for not guessing it, and some of the model's capacity goes to guessing it. The paper's later hypothesis is exactly this: future RGB "introduces high-variance appearance details while adding little information beyond the more structured targets."
Second, the three structured modalities each hold a different part of the event. Tracks alone tell you the cup moved right and slightly up, but not what a cup is. DINO tells you where the cup is, but not how far away. Depth tells you it is closer than the wall, but not which pixels belong together as one object. That complementarity is why the paper asks its real question.
Other groups had already tried alternatives to RGB, either in place of it (DINO-WM, EgoWAM) or alongside it (point-track WAMs, ST-WAM, WAM4D). Concurrent work Flex-π found that jointly predicting several future modalities (RGB, 3D pointmaps, and DINO features) can beat predicting future RGB alone. So "use more than RGB" was in the air.
What was not settled is the question the paper states in one line: How should WAMs combine multiple modalities? That breaks into two design axes, and the paper studies both.
ModAR stands for modality-autoregressive. You probably know autoregressive from language models: generate token 1, then token 2 conditioned on token 1, and so on. ModAR is autoregressive over modalities, not over words or time steps. It fully generates one future modality, then generates the next one conditioned on the finished first, and so on, and it generates the robot's actions last.
Inside each block the model is a denoising generator: it starts from random noise and refines it into a clean prediction, exactly like an image diffusion model refines noise into a picture. Across blocks it is autoregressive. Chapters 4 and 5 unpack both halves.
Most existing WAMs start from a pretrained video-generation model. The paper calls that initialization "highly effective," and then points at its cost for science: it "makes it difficult to isolate the effects of WAM formulation, target representations, and actionless-data scale." If a model inherits billions of parameters of video knowledge, you cannot tell whether a design choice helped or whether the pretraining carried it.
So the paper does something unusual. It trains every compared model from scratch, with the same backbone, the same action-labeled data, and the same optimization budget. The only thing that changes between two rows of a results table is the thing being studied. Then, separately, it runs one system-level comparison against a large video-pretrained WAM (Flex-π, 6B parameters), to see whether a small from-scratch model is even in the same league.
Here is the scoreboard you are going to earn, chapter by chapter. Hold these numbers loosely for now; each one gets its full context later.
| Claim | Evidence in the paper | Chapter |
|---|---|---|
| Sequential generation beats other WAM formulations | Highest average success at every data scale in simulation: 66%, 75%, 76% at 50, 250, 1,250 demonstrations per task | 6 |
| It uses actionless data best | +10 percentage points (66% to 76%) from 1,200 extra actionless demonstrations per task, versus +1 point for Unified | 6 |
| Tracks, DINO, depth each help; RGB does not add consistently | Removing tracks, DINO, depth drops 75% to 61%, 65%, 70%; removing RGB leaves 75% | 7 |
| Small from-scratch can match big pretrained | 75% vs Flex-π's 72%, with about 200× fewer parameters (30.1M vs 6B) and about 20× fewer training FLOPs | 8 |
| It works on real robots and learns from humans | 83.3% real-world success vs 66.7% (Unified) and 52.2% (Action-only); human videos lift ModAR from 70.0% to 81.1% to 83.3% | 8 |
| Chapter | What you will be able to do afterwards |
|---|---|
| 1 · Four Couplings | Draw the Unified, Independent-noise, Disjoint, Action-only and ModAR formulations and say exactly how they differ |
| 2 · Targets & Tokens | Write the problem definition and count every token ModAR predicts |
| 3 · The Shared Trunk | Sketch the DiT with shared blocks, modality experts and adaLN, and budget its parameters |
| 4 · Block-Causal Order | Derive the factorization, build the two-copy training sequence and its mask |
| 5 · Flow & Context Noise | Implement the x-prediction flow loss, the sampler, and context noise |
| 6 · Formulations × Scale | Read Table I and the two control experiments that defend it |
| 7 · Which Futures Matter | Read Table II and the ablations, and explain why RGB was dropped |
| 8 · Flex-π & Real Robots | Check the compute comparison and the real-world results yourself |
| 9 · Limits & Next Reads | State what the paper does not show and where to go next |
Before reading on, write down which one modality you think gives the best policy when predicted alone with 50 demonstrations per task: RGB, depth, point tracks, or DINO. Then write down which one you think hurts most to remove from the full set. Chapter 7 grades both guesses, and the two answers turn out to be different modalities.
Before the details, here is what happens every time the deployed real-robot ModAR is asked what to do, using only numbers the paper states. Keep this picture in mind; every later chapter zooms into one line of it.
That is 32 network passes per cycle in the real-robot configuration (8 per stream, four streams), by our count. The only latency the paper reports is for the four-future simulation configuration: 147.9 ms for 40 passes on one RTX 5090.
The paper opens with a figure titled "Predict the future one modality at a time, then act." It shows the three real-robot tasks (stack cups, fold towel, place in drawer) as rows, and four numbered columns as the generation order: 1 predicted tracks (motion), 2 predicted DINO (semantics), 3 predicted depth (geometry), and 4 predicted actions, shown on the live RGB camera view.
Look at what is not in that figure: there is no predicted RGB column. The real-robot models never imagine pixels at all. They imagine arrows (where each point goes), a patchwork of feature vectors (what each patch is), and a heat map of distance (how far each patch is), and then they move. By the end of Chapter 7 you will know exactly which experiment justified leaving the pixels out.
The column labels are worth memorizing because they are the paper's whole vocabulary for why each modality exists: tracks carry motion, DINO carries semantics, depth carries geometry. RGB, when present, carries appearance, and appearance is the part the paper found least useful to predict in its setting (Chapter 7 has the numbers).
"A world-action model has to generate video." Most do, because that is how they inherit knowledge from pretrained video generators. But the definition only requires jointly modeling future observations and actions, and the paper's "modality" is any representation of the future. A WAM that never outputs a single pixel is still a WAM.
"Autoregressive means one token at a time." ModAR is autoregressive across four or five blocks, and inside each block all tokens are generated together by denoising. The sequential part is short (at most five stages); the parallel part is wide (384 tokens per visual block).
"More modalities is always better." The paper's own data contradicts this. Adding future RGB to tracks, DINO and depth changes average success by +0.03, 0.00 and −0.01 at the three data scales. Which modalities help is an empirical question, and the answer is task-dependent.
1. The data argument. In one sentence each, say what an action-labeled demonstration can train that an actionless one cannot, and what an actionless one can still train. Check: actionless data cannot supervise the action chunk; it can still supervise every future modality it has targets for.
2. The two axes. Name the two design axes the paper studies and give one example choice on each. Check: representation (which futures: e.g. tracks versus RGB) and formulation (how futures couple to actions: e.g. joint versus sequential).
3. A prediction to grade later. Write down whether you expect a sequential WAM or a joint WAM to benefit more from extra actionless data, and why. Chapter 6 answers with a number.
Imagine two WAMs that predict exactly the same futures, from exactly the same data, with exactly the same network. One reaches 75% average success in simulation. The other reaches 34%. Nothing about what they predict differs. The only difference is when each piece gets generated and which pieces are allowed to look at which.
That is not a hypothetical. It is one column pair of the paper's Table I, at 250 demonstrations per task (ModAR at 75%, Independent-noise at 34%). This chapter builds the vocabulary to see why such a small-sounding choice can be worth forty points.
Every model in this chapter generates its outputs by denoising. Picture a sculptor who starts with a block of static and, in a handful of passes, chips it into a statue. Each pass looks at the current rough block and predicts the finished statue, and then moves the block a little toward that prediction. After enough passes, the block is the statue.
The formal name for the version ModAR uses is flow matching (Chapter 5 builds it from zero). The progress of one pass is tracked by a number the paper calls the flow timestep τ. In the paper's convention, τ = 0 means pure noise and τ = 1 means the clean result. Generation walks τ from 0 to 1 in a fixed number of steps. ModAR uses 8 Euler steps per generated stream.
One more word you will see constantly: a stream is one set of tokens being generated together. The future point tracks are one stream. The future DINO features are another. The action chunk is another. A WAM that predicts four future modalities plus actions has five streams to generate. The formulations below are five different answers to "how do those five streams relate while they are being generated?"
The most common answer is to generate everything at once. All streams start as noise together, share one flow timestep, and are refined in lockstep. At every refinement step, every stream can attend to every other stream, so the action tokens can read the half-finished future tokens and vice versa.
Analogy: a comic artist sketching all six panels of a page simultaneously, glancing across the page while each panel is still a rough pencil sketch. The paper labels this Unified and uses it to represent joint-generation WAMs such as DreamZero and Cosmos Policy, which "denoise future-observation and action streams together" with cross-stream attention. In the paper's controlled implementation, Unified uses one shared flow timestep drawn from a logit-normal distribution with (μ, σ) = (−1, 1).
Notice the subtle consequence. At step 3 of 8, the action stream is being refined while looking at futures that are themselves only three-eighths clean. The action prediction conditions on partially noisy predicted futures. The paper returns to this asymmetry with a clever control experiment in Chapter 6.
A close cousin of Unified changes only the training procedure. During training, instead of drawing one shared τ for all streams, it draws a separate τ for each stream, independently. One training example might have the tracks at τ = 0.9 (almost clean) while the actions sit at τ = 0.2 (mostly noise). At inference it does exactly what Unified does: all streams denoised simultaneously on one shared schedule.
The paper attributes this design to Unified World Models and Flex-π. One way to see its appeal (this is our gloss, not a claim from the paper): a stream sampled near τ = 1 behaves almost like clean conditioning for the others, so a single network gets exposed to many different "who conditions on whom" situations during training.
The paper's from-scratch experiments, however, find Independent-noise "performs poorly." Its explanation is a train/test mismatch. At test time, all streams are always at the same noise level. During training, independently sampled noise levels "rarely match the synchronized test-time denoising schedule," so the model "receives insufficient training signal near its inference regime." It spends most of its practice on situations it will never face.
The third answer cuts the connection. Disjoint predicts each target independently, "without attention between future-observation and action targets." The futures and the actions share the backbone that reads the current observation, but the action tokens never look at the predicted future tokens, and the future tokens never look at the action tokens.
Why would anyone want that? Because then the future prediction is purely a training-time auxiliary objective. At deployment you can skip generating the future entirely and save its cost. This represents Fast-WAM, whose paper title asks the question directly: do world action models need test-time future imagination? The future-prediction loss still shapes the shared representation during training. It just never feeds the action at run time.
The baseline predicts actions directly from observations and predicts no future at all. The paper describes it as "a typical flow-matching behavior cloning policy." Because it has no future-prediction objective, it has nothing to learn from actionless demonstrations, so it trains only on action-labeled data. That is why, in the results tables, Action-only has a number only at the smallest data scale: adding actionless data changes nothing for it.
A separate lineage (UniPi and VERA) already reversed the coupling: they "fully denoise visual futures and then infer actions from the resulting frames, rather than co-denoising both streams." ModAR builds on that futures-then-actions ordering and extends it across modalities. It denoises future modalities one at a time, each fully, each conditioned on the ones already finished, and denoises actions last.
That gives ModAR's final step a clean interpretation. An inverse dynamics model (IDM) answers a detective's question: given what the world looks like now and what it will look like next, which action caused the change? If you see a drawer closed and then open, the answer is "someone pulled it." The paper notes that ModAR's final action-prediction step "acts as an inverse dynamics model," mapping the fully generated future to the actions that would produce it.
| Formulation | Training noise | Inference order | Actions see futures? | Uses actionless data? | Represents |
|---|---|---|---|---|---|
| ModAR | Own τ per stream, plus noisy context copies | Tracks → DINO → depth → RGB → actions, 8 steps each | Yes, fully clean ones | Yes | This paper |
| Unified | One shared τ | All streams together, 8 steps | Yes, partially noisy ones | Yes | DreamZero, Cosmos Policy |
| Independent-noise | Independent τ per stream | Same as Unified | Yes, partially noisy ones | Yes | Unified World Models, Flex-π |
| Disjoint | Not specified (targets predicted independently) | Futures skipped at inference; actions only | No attention between them | Yes (as auxiliary loss) | Fast-WAM |
| Action-only | Actions only | Actions only, 8 steps | No futures exist | No | Flow-matching behavior cloning |
The simulation below animates the table. Each lane is one stream. A lane fills from left (noise) to right (clean) as it is denoised. The arrows show which streams the currently active stream can read. Switch to the training view to see how each formulation samples noise levels during training, and resample a few times to feel why Independent-noise so rarely practices the synchronized situation it meets at test time.
Pick a formulation. Inference view: the clock runs through every Euler step; ModAR takes 40 (8 per stream), the others 8 in total. Drag the clock to scrub. Training view: each dot is one stream's sampled noise level τ for one training example. Press Resample repeatedly. The shaded band marks "all streams within 0.1 of each other," which is what inference always looks like for the simultaneous formulations.
Spend a moment on the training view for Independent-noise. With five streams each drawing its own τ, the chance that all five land close together is small. Chapter 6 puts a toy number on "small." The point to carry forward: the inference procedure of Unified and Independent-noise is identical, so any gap between them in the results is caused purely by how their training noise was sampled.
A comparison is only as good as its controls. The paper implements every formulation "within one controlled implementation": all methods "share the same backbone, action-labeled data, and optimization budget." The same six simulation tasks, the same demonstration counts, the same 1.2M optimizer steps. Only the coupling changes.
It is worth pausing on why generating streams in sequence might help, before the evidence arrives. When you generate everything jointly, every stream starts from noise with nothing solid to lean on. When you generate in sequence, the second stream starts with a finished first stream in hand.
The paper cites two image-generation results that motivate this. Latent Forcing generates image latents before pixels, "allowing the latents to serve as a semantic scratchpad for generating fine-grained appearance." Modality Forcing adapts a pretrained image generator to jointly generate depth and finds that image models contain priors that improve depth accuracy. The paper's summary of the lesson: "generating easier-to-predict modalities earlier can expose structure that simplifies subsequent predictions."
There is also a plain probabilistic reason, which Chapter 4 makes exact. Any joint distribution over several variables can be written as a product of conditionals, in any order, with no loss. So sequential generation does not restrict what the model can represent. It only changes what each generation step has to figure out on its own, and that is precisely the thing the order controls.
Abstract tables hide the thing that matters, so let us trace one decision. The robot is at decision time t. It must output 16 actions. Both models below predict the same futures (tracks, DINO, depth, RGB) plus actions, and both use 8 Euler steps per unit of work.
In Unified, one shared clock ticks 8 times. At tick k (k = 0 through 7) every stream sits at τ = k/8 and is moved to τ = (k+1)/8. (We assume evenly spaced steps throughout this lesson; the paper specifies eight Euler steps but not their spacing.) So when the action stream takes its first step, the futures it reads are pure noise (τ = 0). When it takes its last step, the futures it reads are at τ = 7/8 = 0.875, still carrying one eighth of their noise.
In ModAR, the clock ticks 8 times for tracks alone, then the finished tracks are frozen and re-embedded as context. Then 8 ticks for DINO, reading clean tracks the whole time. Then depth, then RGB. Only after 32 ticks does the action stream start its own 8 ticks, and on every one of them it reads four finished futures.
| Moment | Unified: futures seen by the action stream | ModAR: futures seen by the action stream |
|---|---|---|
| First action step | All four at τ = 0.000 (pure noise) | All four at τ = 1 (finished) |
| Middle action step (k = 4) | All four at τ = 4/8 = 0.500 | All four at τ = 1 |
| Last action step (k = 7) | All four at τ = 7/8 = 0.875 | All four at τ = 1 |
| Total network passes | 8 | 8 × 5 = 40 |
Here are the two sampling loops side by side, as pseudo-PyTorch. It is a sketch of the control flow the paper describes, not the authors' code. The key line is the one that decides what the network is allowed to read.
# ---- Unified: one shared clock, every stream refined together ---- streams = {m: torch.randn(shape[m]) for m in ["tracks", "dino", "depth", "rgb", "act"]} for k in range(8): tau = k / 8 x_hat = net(obs, streams, tau={m: tau for m in streams}) # joint attention, all noisy for m in streams: v = (x_hat[m] - streams[m]) / (1 - tau) streams[m] = streams[m] + v / 8 # ---- ModAR: one stream at a time, finished streams become context ---- context = {} for m in ["tracks", "dino", "depth", "rgb", "act"]: y = torch.randn(shape[m]) for k in range(8): tau = k / 8 x_hat = net(obs, context, target=(m, y), tau=tau) # reads only FINISHED blocks y = y + (x_hat - y) / (1 - tau) / 8 context[m] = y # re-embedded as clean context for every later block actions = context["act"]
The ModAR loop is the paper's inference paragraph turned into code: "We generate one modality at a time... The resulting clean tokens become context for generating the next modality, and we condition the final action chunk on the complete generated future." The paper adds one engineering detail the sketch leaves out: keys and values for the current observation and all previously generated modalities are cached, so they are not recomputed at every denoising step (Chapter 4 explains why that is safe).
Return to the data argument from Chapter 0. A human video of cup stacking has future frames but no robot actions. Each formulation handles that example differently, and the difference is the reason data scaling behaves so differently across them in Chapter 6.
| Formulation | What an actionless example trains | How that reaches the action at run time |
|---|---|---|
| ModAR | Every future block it has targets for; action prediction is omitted for that example | Directly: the action step reads every finished future, alongside the observation, configuration and task |
| Unified / Independent-noise | The future streams; the action loss is masked | Through joint attention, but reading partially noisy futures |
| Disjoint | The future streams, as an auxiliary loss | Only indirectly, through the shared representation; futures are not generated at run time |
| Action-only | Nothing; it cannot use the example | Not at all |
The paper states the ModAR rule plainly: "For actionless examples, we omit action prediction and supervise only the available future targets." And the loss has an explicit switch for it: the action term is included "only for examples with action labels."
Every formulation has a failure mode, and they are different.
In Disjoint, a wrong future cannot hurt the action at run time, because the future is never generated at run time. The flip side: a right future cannot help it either. The paper also sees a subtler cost: as more actionless data is added, Disjoint's performance falls, and the authors suggest the shared representation becomes "increasingly shaped by the future-observation prediction objective, causing negative transfer to the action prediction objective."
In Unified, streams are refined together, so no stream ever commits early. That limits how much any single early mistake can dominate. But it also means the action never gets to read a committed future.
In ModAR, commitment is the point, and so is the danger. If the tracks come out wrong, DINO is generated to agree with wrong tracks, depth with both, and the action with all three. The paper names this directly: modality-autoregressive generation "is susceptible to cascading error, as imperfections in earlier generations become inputs to all later predictions." Its fix, context noise, gets a full treatment in Chapter 5, and removing it is the single most damaging ablation in Chapter 7 after removing tracks.
A compact way to hold all five in your head is to ask each of them two questions: are futures generated before actions or alongside them? and does the action ever read a predicted future?
| Action reads predicted futures | Action never reads predicted futures | |
|---|---|---|
| Futures alongside actions | Unified, Independent-noise (reads them while still noisy) | Disjoint (futures are a training-time loss only) |
| Futures before actions | ModAR (reads them finished); UniPi and VERA (one visual stream) | (no sensible design lives here) |
| No futures at all | — | Action-only |
Read the table as a map of what each design bets on. Disjoint bets that futures are useful only as a teacher for the shared representation. Unified bets that futures and actions should be negotiated together. ModAR bets that the future should be settled first, in a particular order, and that the action should then be chosen with that settled future in view.
One way to test that you really own the map: pick any cell and say what the model would do with a human video of cup stacking. Unified and Independent-noise learn the future streams from it and hope joint attention carries the lesson to the actions. Disjoint learns the future streams from it and then never generates them at run time. ModAR learns its future blocks from it, and at run time its action step reads blocks of exactly that kind. Action-only cannot use the clip at all.
(a) A model that denoises depth and actions together, then throws the depth away at run time and never lets actions attend to it. (b) A model that samples one noise level for all streams during training and denoises everything together at test time. (c) A model that fully generates DINO features, then fully generates actions conditioned on them.
Answers: (a) is Disjoint in spirit (future prediction as an auxiliary loss, no future-to-action attention, futures omitted at deployment). (b) is Unified. (c) is a one-modality ModAR (K = 1 in Equation (1)), and it is exactly the "DINO" column of Table II.
Before any architecture, pin down the contract. At one decision moment, what exactly goes into ModAR, and what exactly comes out? If you cannot write both down with shapes, you cannot build it. This chapter writes both down, using the paper's own notation and its own implementation numbers.
The paper's goal is "to jointly model future multimodal observations and robot actions conditioned on the current observation, robot configuration, and task label." Each of those three conditioning pieces has a symbol.
Here t is the decision time, the control step at which the robot stops and plans. ot is the multimodal observation at that moment: what the single camera sees, in the modalities the model uses. qt is the robot's proprioceptive configuration, its sense of its own body: "joint positions and gripper opening." g is "the learned embedding associated with a discrete task label," a trainable vector looked up from a task ID such as stack bowls.
Two details are worth underlining now, because they come back as limitations. There is one camera, not several: models "condition on a single-camera 168×224 observation." And the task is a discrete label, not a sentence: this is not a language-conditioned model. The paper lists that second point among its limitations.
For every modality m in the set M = {rgb, depth, dino, tracks}, the model predicts a short list of future snapshots.
Read it symbol by symbol. ymt+Δ is one future frame of modality m, taken Δ control steps after the decision. Δ is the dynamics stride: how far apart, in control steps, consecutive predicted frames sit. J is how many future frames are predicted. H is the prediction horizon, the last predicted moment, and it equals J times Δ by construction.
Actions are predicted differently. Not every Δ steps, but at every single control step up to the horizon.
at is the command sent to the robot at step t. At is a block of H consecutive commands, called an action chunk: rather than choosing one command, looking again, and choosing the next, the policy commits to a whole short sequence at once (the idea popularized by ACT). The paper spells out how the two timelines line up: "the final prediction target corresponds to the next decision point t+H, while At contains the H actions executed between t and t+H."
In other words, the model imagines where the world will be at the moment it next gets to think, and it chooses the commands that fill the gap.
The implementation section gives the values: "horizon H = 16 and dynamics stride Δ = 8, yielding J = 2 sparse targets at t+8 and t+16, together with a dense H-step action chunk; policies replan every H steps." Let us check every piece of that sentence, with nothing skipped.
Step 7 is our arithmetic, but it matters later: ModAR's measured latency (147.9 ms per query) is paid once per chunk, not once per control step. The paper does not state the control frequency, so we will not convert that into a duty cycle.
One more fact about the actions: they are absolute joint configurations, not changes. Each at says "put the joints here," which is also the form of qt. So the conditioning and the targets live in the same 14-dimensional space.
A transformer consumes a sequence of vectors, called tokens, all of the same width. Four very different kinds of data have to become that one kind of thing. The paper's "Modality tokenization" paragraph does it in five moves.
Two terms need definitions. To patchify is to chop an image into non-overlapping squares and treat each square as one token, the same trick a Vision Transformer uses. A point track is the path of one physical point through a video: where it is in each frame, and whether it is visible or hidden behind something. CoTracker3 is the tracker that produces those paths here.
Notice the neat alignment. The paper puts RGB, depth and the track queries on one 12×16 grid, with the track queries at the patch centers. DINOv2 ViT-S/14 also works in 14×14 patches. If it is run on the same 168×224 view (our inference; the paper does not state DINO's input size), it too yields a 12×16 grid. In that case all four visual modalities describe the same 192 locations, and token (row 5, column 9) means the same patch of table in every modality. The token counts below assume that alignment.
The paper gives the grid and the horizon. The counts below are our arithmetic from those numbers. The one place where we have to read between the lines is the exact layout of a track token; the paper says it encodes displacement and visibility, and we read that as two displacement numbers plus one visibility number.
Step 9 is the quiet argument of the whole paper, in numbers. Future RGB is 225,792 raw values to get right, about 196 times the size of the track future (225,792 / 1,152 = 196). The tracks, which Chapter 7 will show are the most costly modality to remove, are the smallest target of all. Size and usefulness are not the same thing.
Once projected, every token has the same width, 384 (the shared transformer width you will meet in Chapter 3). If each projection is a single linear layer with a bias, as "we linearly project" suggests, then the RGB projection alone holds 588 × 384 + 384 = 225,792 + 384 = 226,176 parameters, and the track projection only 3 × 384 + 384 = 1,152 + 384 = 1,536.
Top: the timeline of one decision. Warm diamonds are future targets (every Δ steps up to H), green dots are the dense action chunk. Bottom left: the 12×16 patch grid; tap a patch to see its axial position ids. Bottom right: tokens per stream. Move Δ and J, toggle modalities, or load a preset. The Flex-π preset only sets the timing the paper used for that baseline (targets at t+4, 8, 12, 16); Flex-π tokenizes differently, so the bars stay ModAR-style counts.
Try the Flex-π timing and watch the visual token budget double while the action chunk stays at 16. Halving the stride doubles J, and every visual stream grows with J. This is one reason the future side of a WAM, not the action side, dominates its compute.
A transformer is blind to order unless you tell it where each token sits. ModAR tells it with rotary position embeddings (RoPE): before attention compares a query with a key, both are rotated by an angle proportional to their position. Because the rotations compose, the similarity between two tokens ends up depending on how far apart they are, not on where they both are. (The RoFormer lesson derives this in full.)
Here is the smallest possible example, with illustrative numbers. Take a 2-number slice of a query, (1, 0), and a frequency of 0.5 radians per time step. A token at time 1 is rotated by 1 × 0.5 = 0.5 rad. A token at time 2 is rotated by 2 × 0.5 = 1.0 rad. Their dot product is cos(1.0 − 0.5) = cos(0.5) = 0.8776. Shift both to times 5 and 6 and the angles become 2.5 and 3.0 rad, but the dot product is still cos(3.0 − 2.5) = cos(0.5) = 0.8776. Only the gap survives.
Axial RoPE applies that idea along several axes at once by splitting each attention head's dimensions into groups, one group per axis. For ModAR's visual tokens the axes are time, row and column, so a token at (future frame 2, row 5, column 9) is rotated by three independent sets of angles. Action tokens only have a time axis. The paper does not say how the dimensions are divided among the axes; the idea is what matters here.
RoPE says where a token is. It does not say what kind of token it is: a DINO token and a depth token at the same (time, row, column) get identical rotations. That is the job of the learned modality embedding, a trainable vector per modality added to every token of that modality, like a colored tag on each card in the deck.
Three different things happen to the three "helper" networks, and it is worth being precise.
| Component | Status | Why that choice makes sense |
|---|---|---|
| DINOv2 ViT-S/14 | Frozen | It defines the DINO target space. If it trained along with ModAR, the targets would move while the model chased them (our reasoning; the paper states only that the encoder is frozen) |
| CoTracker3 | Used to produce track targets from the demonstration videos | Point tracking is a hard problem with its own training recipe; ModAR only needs its outputs as labels |
| Projections, embeddings, DiT, heads | Trained from scratch | The paper's controlled study requires "no pretraining" for every compared model |
On the real robot, a fixed third-person ZED stereo camera "provides RGB observations and stereo depth for both robot and actionless human demonstrations." So depth there is measured by stereo matching rather than rendered. A stereo estimate has its own errors, and whatever errors it has become part of the depth targets; the paper does not quantify them.
Each representation degrades differently when the input gets hard, and the paper's design already encodes one of those cases.
Occlusion. When a gripper passes in front of a tracked point, the point's position becomes unknowable from the image. That is exactly why the track token carries a visibility value alongside its displacement: the model can say "this point went behind something" instead of being forced to invent a position.
Lighting and texture changes. These rewrite the RGB future, as Sim 1 showed, while the paper describes DINO features as "robust to appearance variation" and tracks as encoding motion "independently of appearance."
A single viewpoint. With one camera, depth is the only modality that carries distance explicitly. That is the paper's stated reason depth helps: it "makes scene geometry explicit and provides direct supervision for the spatial reasoning required for predicting 3D actions."
Here is the whole tokenization for one training example as pseudo-PyTorch, with every shape annotated. The shapes follow the paper's implementation numbers; the function names, and the 3-number track layout, are ours.
# One example, batch B. Shapes from the paper: 168x224 image, 12x16 grid of 14x14 patches, # J = 2 future frames (t+8, t+16), H = 16 actions, 14-D joint configurations. obs_rgb # (B, 3, 168, 224) current camera image obs_depth # (B, 1, 168, 224) current depth fut_rgb # (B, 2, 3, 168, 224) frames at t+8 and t+16 fut_depth # (B, 2, 1, 168, 224) q # (B, 14) joint positions + gripper opening, both arms task_id # (B,) discrete task label actions # (B, 16, 14) absolute joint targets; absent for actionless data def patchify(x, p=14): # (B, T, C, 168, 224) -> (B, T, 12, 16, C*p*p) B, T, C, Hh, Ww = x.shape x = x.reshape(B, T, C, Hh // p, p, Ww // p, p) return x.permute(0, 1, 3, 5, 2, 4, 6).reshape(B, T, Hh // p, Ww // p, C * p * p) dino = load_dinov2_vits14().eval().requires_grad_(False) # FROZEN target extractor with torch.no_grad(): fut_dino = dino.patch_tokens(fut_rgb.flatten(0, 1)).view(B, 2, 12, 16, 384) # Tracks: CoTracker3 follows the 192 patch-centre queries through the demo video. disp, vis = cotracker3_targets(video, queries=patch_centres(12, 16), times=[8, 16]) fut_tracks = torch.cat([disp, vis], dim=-1) # (B, 2, 12, 16, 3): dx, dy, visible D = 384 # common token width proj = nn.ModuleDict({"rgb": nn.Linear(588, D), "depth": nn.Linear(196, D), "dino": nn.Linear(384, D), "tracks": nn.Linear(3, D), "act": nn.Linear(14, D)}) mod_emb = nn.Embedding(5, D) # learned "which modality" tag def visual_tokens(name, grid): # grid: (B, 2, 12, 16, c) tok = proj[name](grid) + mod_emb(ID[name]) # (B, 2, 12, 16, 384) pos = axial_ids(t=2, rows=12, cols=16) # (384, 3): (time, row, col) for RoPE return tok.flatten(1, 3), pos # (B, 384, 384) tok_rgb, pos_rgb = visual_tokens("rgb", patchify(fut_rgb)) tok_act = proj["act"](actions) + mod_emb(ID["act"]) # (B, 16, 384) pos_act = torch.arange(16) # time-only RoPE for actions
Take the patch at row 5, column 9. Its pixel square spans y = 70 to 83 and x = 126 to 139, so its center, where the track query sits, is at
Now suppose (illustrative numbers) CoTracker3 reports that the physical point under that query is at x = 151, y = 70 at t+8 and is visible, and at t+16 it has passed behind the gripper. The two track tokens for that query would then carry:
Two honest gaps: the paper does not say whether displacements are normalized (for example divided by the image size) and does not say what displacement value is stored for an occluded point. A reimplementation has to choose. What the paper does fix is the meaning: displacement "from its initial position" plus "its visibility."
| Field | Robot demonstration | Actionless demonstration (simulated or human) |
|---|---|---|
| Observation ot | Yes | Yes |
| Future targets Ytm | Yes, all predicted modalities | Yes, "the available future targets" |
| Action chunk At | Yes (16 × 14) | No: the action term is dropped from the loss |
| Configuration qt | Yes (14 numbers) | In simulation the actionless demos are robot demos with their actions withheld, so qt presumably still exists; for human video the paper does not say what is used |
| Task label g | Yes | Presumably the matching task's label; the EgoDex categories are chosen per task ("stack/unstack cups" for cup stacking, and so on), but the paper does not spell out the labeling |
A reimplementation initializes its track queries at the top-left corner of each 14×14 patch instead of the center. Training runs, but the tracks stream is noticeably harder to learn near object edges. Why, using only this chapter?
Answer (our reasoning): the design relies on every modality describing the same 192 locations. A corner query sits on the boundary between four patches, so near an object edge it often tracks a point that belongs to a neighboring patch's object. The track token for patch (r, c) then describes motion that the DINO and depth tokens for (r, c) do not show, and the cross-modal agreement the shared trunk is supposed to exploit is broken. Center queries keep the four modalities aligned.
You now have five kinds of tokens: tracks, DINO, depth, RGB and actions, all 384 numbers wide. They need to talk to each other, because the whole point is that tracks inform DINO, DINO informs depth, and all of them inform the action. But they also need room to be themselves, because a depth patch and a joint command are very different objects. ModAR's backbone is designed around exactly that tension.
Picture a newsroom putting out tomorrow's paper. First, everyone sits in one editorial meeting: the sports desk hears what politics is running, the photo desk hears what the lead story needs. Then each desk goes back to its own room and writes its own section, in its own style, without the other desks talking over it. Finally each section goes to its own layout template.
ModAR's network is that newsroom. The shared meeting is a stack of transformer blocks every stream attends through. The desks are small modality-specific expert stacks, where attention is restricted to one stream. The layout templates are per-modality linear output heads. The paper's one-sentence summary: "cross-modal information is fused in the shared blocks before within-modality specialization in the expert blocks."
The backbone is a diffusion transformer (DiT), the architecture Peebles and Xie introduced for image diffusion (the DiT gleam builds it from zero). A DiT is an ordinary transformer that takes noisy tokens in and predicts something about their clean version, with the noise level injected into every layer. Nothing in it is specific to images, which is why it transfers so easily to tracks, depth and actions.
Each transformer block does two things, each wrapped in a normalization and a residual connection. Self-attention lets every token gather information from the tokens it is allowed to see: each token emits a query, every visible token offers a key and a value, and the token receives a weighted mix of values, weighted by how well its query matches each key. Then an MLP (a small two-layer network) transforms each token on its own.
The implementation section lists it in one sentence, which we unpack line by line below. "The shared DiT contains six width-384 transformer blocks with six attention heads (ViT-S), followed by two width-384 expert blocks for DINO, depth, and RGB, one width-128 block for tracks, and two width-128 blocks for actions."
| Stage | Blocks | Width | Attention reach |
|---|---|---|---|
| Shared trunk | 6 | 384, 6 heads (ViT-S sizing) | All causally available modality tokens |
| DINO expert | 2 | 384 | DINO stream only |
| Depth expert | 2 | 384 | Depth stream only |
| RGB expert | 2 | 384 | RGB stream only |
| Tracks expert | 1 | 128 | Tracks stream only |
| Action expert | 2 | 128 | Action stream only |
| Output heads | 1 linear layer each | to the modality's raw size | n/a |
We read "two width-384 expert blocks for DINO, depth, and RGB" as two blocks per modality, since the paper describes the expert stacks as modality-specific. With six heads on a width of 384, each head works in 384 / 6 = 64 dimensions.
Why are the tracks and action experts narrower? The paper does not say, but the token sizes from Chapter 2 suggest a reason: a track token starts life as about three numbers and an action token as fourteen. A 384-wide specialist for a 3-number target would be mostly empty space. Narrow experts spend parameters where the targets actually have detail.
"Causally available" in the shared-trunk row is doing a lot of work. It means a stream can only attend to what would exist at that point of generation: the current observation, the finished earlier modalities, and itself. Chapter 4 turns that phrase into an exact attention mask.
The network also has to know three things that are not tokens: the robot's joint configuration qt, the task embedding g, and how noisy each stream currently is (its flow timestep τ). The paper feeds all three through adaLN: "The robot configuration, task embedding, and flow timesteps for each modality condition the transformer through adaLN."
Start with plain layer normalization. Given a token vector x, subtract its mean, divide by its standard deviation, then multiply by a learned scale γ and add a learned shift β.
In plain layer norm, γ and β are fixed learned constants. In adaptive layer norm (adaLN), a small network computes γ and β from the conditioning. The analogy: a mixing engineer's fader. The song (the token) is the same, but the fader position (γ, β) is set by who is listening, how noisy the room is, and which song this is.
The numbers here are illustrative. The point is to see one normalization and two different conditionings, with every step written out.
Same token, same weights, different behavior. In this toy, the conditioning turns the response down for the noisy stream and up for the clean one; a trained adaLN learns whatever mapping helps, and the paper does not describe what its learned mapping looks like. The general point stands either way: this is how one set of weights can serve every noise level and every task.
One subtlety deserves a flag. ModAR's streams are at different noise levels in the same forward pass (finished context is clean, the block being generated is noisy). "Flow timesteps for each modality" suggests the timestep part of the conditioning is applied per stream, so tokens of different streams get different γ and β in the same layer. That is our reading of the sentence; the paper does not show the adaLN wiring. The robot configuration and task embedding are described as "global adaLN conditioning to each DiT layer."
Pick a query stream and a layer type. Arrows show which streams that stream's tokens may attend to, during generation of that stream. In a shared block it reads the observation, every finished earlier modality, and itself; in its expert block it reads only itself. The τ slider drives a toy adaLN: watch the same normalized token get scaled and shifted differently. The bar at the bottom is the transformer-block parameter budget from the worked example below.
First, which model does that number describe? The paper gives 30.1M once, for one specific variant: "our trained-from-scratch tracks–DINO–depth ModAR variant," the one it compares with Flex-π. That variant predicts no RGB, so it has no RGB expert. The paper does not report a parameter count for the four-modality model.
The paper gives the total and the block layout. It does not itemize the rest. So here is a back-of-envelope budget for the tracks–DINO–depth variant, using the standard rule of thumb that one transformer block with a 4× MLP holds about 12d2 parameters (4d2 in the four attention projections, 8d2 in the MLP), ignoring biases and norms. The 4× MLP ratio is our assumption (it is the ViT-S standard); the paper does not state it.
What could fill that 11.8M? Some pieces we can size exactly from Chapter 2, and the next worked example does. The rest belongs to parts the paper does not size: the adaLN conditioning networks, the task-embedding table, and whatever connects the width-128 experts to the width-384 trunk.
The leftover is large enough to be interesting. The original DiT computes a separate scale-and-shift regression in every block; at full width that costs roughly 6d2 per block. For this variant's ten width-384 blocks that is 10 × 6 × 147,456 = 8,847,360, and for its three width-128 blocks 3 × 6 × 16,384 = 294,912, together 8,847,360 + 294,912 = 9,142,272, about 9.1M. That fits inside the 11.8M, so the budget is at least consistent with a DiT-style per-block adaLN regression. It does not prove one: the paper does not show the wiring, and the lesson will not pretend otherwise.
Under the single-linear-layer reading of "we linearly project each modality" and "a linear output head," and assuming the track and action heads read from their width-128 experts, the in and out layers can be counted exactly (weights plus biases; our arithmetic). The table lists all five streams for reference; the sums below it are for the tracks–DINO–depth variant, which has no RGB head.
| Stream | Input projection | Output head |
|---|---|---|
| RGB | 588 × 384 + 384 = 226,176 | 384 × 588 + 588 = 226,380 |
| Depth | 196 × 384 + 384 = 75,648 | 384 × 196 + 196 = 75,460 |
| DINO | 384 × 384 + 384 = 147,840 | 384 × 384 + 384 = 147,840 |
| Tracks | 3 × 384 + 384 = 1,536 | 128 × 3 + 3 = 387 |
| Actions | 14 × 384 + 384 = 5,760 | 128 × 14 + 14 = 1,806 |
| Sum, no RGB | 230,784 | 225,493 |
| Sum, all five | 456,960 | 451,873 |
So most of the unexplained 11.4M must sit in the conditioning pathway and the adapters, the two parts the paper describes only in words. If you reimplement the tracks–DINO–depth ModAR and land far from 30.1M, those are the places to look.
Sketch the shared trunk, every expert stack with its depth and width, the heads, and the three conditioning signals. Then say where cross-modal attention happens and where it is forbidden.
Answer: 6 shared blocks at width 384 with 6 heads; experts of 2 blocks at width 384 for DINO, depth and RGB, 1 block at width 128 for tracks, 2 blocks at width 128 for actions; one linear head per stream; adaLN from robot configuration, task embedding and per-stream flow timesteps. Cross-modal attention happens only in the shared blocks (and only toward the observation and earlier finished blocks); expert blocks restrict attention to their own stream.
Since 30.1M belongs to the tracks–DINO–depth variant, the four-modality model of Tables I and II must be bigger: adding an RGB target adds an RGB expert and an RGB output head. The paper does not report that model's size, but the same rule of thumb estimates the difference.
The real-world WAMs "observe and predict only tracks, DINO, and depth," the same modality set as the 30.1M variant. So the real-robot model is plausibly close to 30.1M too, though the paper does not report its size either.
"Experts" can mislead. In large language models, a mixture of experts uses a learned router that sends each token to a few of many experts. ModAR has no router. Every DINO token goes to the DINO expert, always; the assignment is fixed by modality. This is closer in spirit to the separate action expert in π0, where a smaller set of weights is dedicated to one kind of token, than to a routed sparse model.
Every weight in the trunk, the experts, the adaLN pathway, the embeddings, the projections and the heads starts from random initialization. The only pretrained network anywhere in the pipeline is the frozen DINOv2 encoder that defines DINO targets (plus CoTracker3 producing track targets). That is not a limitation of engineering effort; it is the experimental design. With no pretrained trunk, a gap between two ModAR variants can only come from the variants.
It also means the model is small enough to be trained many times. The paper trains one multitask model per method and per data scale, for 1.2M optimizer steps each (Chapter 5). Presumably (our inference; the paper gives no reason) that cost is also why the 6B Flex-π appears in Chapter 8 as a single system-level comparison rather than across every configuration.
Suppose the depth stream is currently a mess, early in its denoising. In the shared blocks, depth tokens are being generated and can read the observation, tracks and DINO, but nothing generated later can read them yet. In the depth expert blocks, the mess stays inside the depth stream. So during depth generation, no other stream's representation is being perturbed by half-finished depth. The block-causal mask (Chapter 4) guarantees the first property; the expert restriction guarantees the second.
The price of this isolation is that the only place modalities meet is the six shared blocks. If cross-modal reasoning needed more depth than that, the design would bottleneck there. The results in Chapters 6 and 7 suggest six was enough for these tasks; the paper does not ablate the trunk depth.
Concept plus realization: here is the path a single future-DINO token takes, with shapes, while ModAR generates the DINO block at inference (four-modality setting, batch size B). The shapes of the input and output are fixed by Chapter 2; the attention sizes use our assumed 576 observation tokens.
The action stream takes the same path with different sizes: 16 tokens of 14 numbers are projected to width 384; in the shared blocks those 16 queries read 576 + 1,536 + 16 = 2,128 keys; then the width-128 action expert runs, and a head maps back to 14 numbers per step. Going from width 384 to 128 requires some projection between the trunk and the narrow experts. The paper does not describe it, so treat that part of the diagram as a necessary but unspecified adapter.
| Signal | Shape | What the network can do with it | Enters through |
|---|---|---|---|
| qt | 14 numbers | Know where its own arms are, so an inverse dynamics step can compute how far to move | Global adaLN |
| g | learned vector per task | Know which of the multitask behaviors to imagine (six tasks in simulation, three on the robot) | Global adaLN |
| τ per stream | one number per stream | Know how much to trust its input: rough structure at low τ, fine detail at high τ | adaLN |
| Observation tokens | tokens | See the scene right now | Attention (not adaLN) |
Sequential generation needs five times more network passes than Unified. Does it need five times more work? Count the attention scores computed per head in the shared blocks, summed over all denoising steps of one query (our arithmetic, four-modality setting, 576 observation tokens assumed, observation and finished blocks cached).
By this proxy, ModAR computes roughly half the shared-block attention of a Unified query, because each of its many passes only pushes one block's queries through the network while the rest sits in the cache. So why does the paper say sequential generation increases latency? Because 40 small passes must run one after another, while 8 larger passes can use the GPU's parallelism more fully; for a 30M-parameter model, per-pass overhead is plausibly a large share of the time. That explanation is ours. The paper's statement, and its measured 147.9 ms, are the facts.
The paper does not ablate this, so what follows is design reasoning rather than evidence. A fully shared stack forces every layer to serve every modality at once: a depth map, a feature map and a joint-command chunk all compete for the same weights all the way to the output. Experts give each target a few private layers to turn the shared, fused representation into its own peculiar output format. The block-causal mask means cross-modal fusion must happen where attention can reach earlier blocks, which is exactly the shared trunk, so the design puts its sharing where the conditioning happens and its specialization where the formats diverge.
In a reimplementation, the finished context blocks are fed into adaLN with the same τ as the block currently being denoised (for example τ = 0.125 at the second step). What would you expect, and what is the fix?
Answer (our reasoning): adaLN is how the network learns how much to trust a stream. Labeling finished context as "mostly noise" tells the trunk to treat solid information as unreliable, so later blocks lean less on earlier ones, which undoes the scratchpad effect. Each stream needs its own timestep signal ("flow timesteps for each modality"). What value the paper assigns to context at inference is not stated; a value consistent with how context was noised in training (between 1 − β and 1) is the natural choice.
A novelist writing a murder mystery does not write every page at once. They decide who did it first. Then the clues get written to agree with that decision. Then the red herrings, which only make sense once the real clues exist. The last chapter, the reveal, is written knowing everything. Each stage is easier because the earlier ones are settled.
ModAR generates the future the same way, and this chapter makes that precise. First with probability, where the ordering turns out to cost nothing. Then with an attention mask, which is how the ordering is enforced inside a transformer. Then with a training trick that supervises every stage in a single forward pass.
Any joint probability of two things can be split into "the first thing" times "the second thing given the first."
This is the chain rule of probability, and it holds exactly, for any variables and in either order. Nothing is approximated. So a model that generates a first and then b given a can represent exactly the same joint distribution as a model that generates both at once.
A tiny example with illustrative numbers shows why the order still matters in practice. Let M be "the cup moves this chunk" (a fact tracks would reveal) and S be "the cup's patch region shifts" (a fact DINO would reveal). Suppose the joint probabilities are p(M=1, S=1) = 0.45, p(M=1, S=0) = 0.05, p(M=0, S=1) = 0.05, p(M=0, S=0) = 0.45.
The joint is identical either way. What changes is the difficulty of each step. Once the motion is settled, predicting where the cup's patches go is nearly deterministic. That is the whole "scratchpad" intuition in five lines of arithmetic.
Now scale the idea up. Pick an ordering (m1, …, mK) of the modalities you predict (all four, or a subset), and write Yt for their future targets in that order. The paper models the joint distribution of futures and actions as its Equation (1):
Symbol by symbol. pθ is the model's distribution, with all network weights collected in θ; there is one network, used for every factor. Ytmk is the future of the k-th modality in the order. Yt<k is shorthand for all the modalities before it, (Ym1, …, Ymk−1); for k = 1 it is empty. The product ∏ multiplies the K future factors, and the leading factor is the action, conditioned on the observation and on every predicted future.
Written out for the three-modality setting the paper uses on real robots, with the order tracks, DINO, depth:
Read the last factor as a job description. Given where things were (c), where they will be (Y), and what they will look like, output the commands that make it happen. That is the inverse dynamics model from Chapter 1, and the paper uses the same words: the final action-prediction step "acts as an inverse dynamics model (IDM), mapping the generated future observations to the actions that induce the predicted transitions."
Language models are autoregressive token by token: word 1, then word 2, then word 3. ModAR is autoregressive block by block: "we generate all tokens of one modality jointly before moving to the next modality." Within a block, all 384 tokens of, say, the DINO future are denoised together, in parallel, over 8 flow steps. Across blocks, the order is strict.
Why not token by token? The paper does not argue this, but the numbers make it obvious: four modalities at 384 tokens each would be 1,536 sequential generation steps, each a network pass. Block-wise generation needs 8 passes per block instead, and the within-block structure (every patch of a depth map depends on its neighbors) is handled by the denoiser, which is good at exactly that.
The paper fixes one order for all main experiments: tracks → DINO → depth → RGB, with actions last. The stated intuition: "this orders the targets from compact, structured representations that are easier to predict toward increasingly high-dimensional and detailed representations."
Compare that with the raw target sizes from Chapter 2: tracks 1,152 numbers, DINO 147,456, depth 75,264, RGB 225,792. Depth has fewer raw numbers than DINO but comes after it. So the order is not "smallest first" in a literal sense; it is "most structured first." A DINO future is a map of what is where, coarse and semantic; a depth future asks for metric detail on every patch.
The paper is candid about how much it tested this. It compares the chosen order with its reverse (Chapter 7), and "We leave a full systematic comparison of modality orderings to future work." Its limitations section adds that the best order "may depend on the task or specific scenario."
At test time the paper's procedure is three verbs: "we therefore denoise one block, re-embed the completed prediction as context, and then denoise the next block."
Re-embed means the finished prediction is treated like an input: it goes through the same linear projection and modality embedding as a clean target would, and becomes a context block. Every later block attends to that context.
The paper adds a speed trick: "We cache keys and values for the current observation and all previously generated modalities to avoid recomputing them at every denoising step." A key-value (KV) cache stores each token's attention keys and values so they need not be recomputed. It is exact here, not approximate, because of the block-causal structure: an earlier block never attends to a later one, so nothing that happens during later denoising can change an earlier block's keys and values.
For the arithmetic we need a guess about the observation tokens. Figure 2 of the paper shows the current observation embedded from DINO, depth and RGB, so we assume 3 × 192 = 576 observation tokens in the four-modality simulation setting. Every other count comes from Chapter 2. This is our arithmetic, counting "tokens pushed through the network," which is a rough proxy for compute.
The exact wall-clock saving depends on details this proxy ignores (attention cost grows with the number of visible tokens; the experts only process their own stream). The direction is unambiguous, and it is why sequential generation is affordable at all: ModAR's measured 147.9 ms includes 40 denoising steps.
Training has a problem inference does not. At inference you generate the blocks one after another. In training you have the true futures for every block, and you want to teach all K + 1 conditionals of Equation (1) at once, from one forward pass, without letting any block cheat.
The paper's solution: "we supervise every stage in one pass by creating a clean context copy and a noisy prediction copy of each target modality." Every target modality appears twice in the training sequence. The context copy holds the true future (clean, apart from the context noise of Chapter 5), playing the role a finished block plays at inference. The prediction copy holds a noised version of the same future, which the model must denoise.
Then one block-causal mask decides who reads whom. In the paper's words, it "lets the prediction copy for mk attend only to the observation and context copies of m1, …, mk−1, preventing target leakage." (Each prediction copy also attends within itself, since its own tokens are denoised jointly.)
Target leakage is what happens when a model can see the answer it is graded on. If the noisy DINO copy could read the clean DINO copy, the network would learn to copy it, the loss would drop toward zero, and nothing about predicting DINO would be learned. If it could read the clean depth copy (a later modality), it would learn to rely on information that will not exist yet at inference. The mask forbids both.
What do the context copies themselves attend to? The natural reading, mirroring inference where each finished block was embedded after its predecessors, is that a context copy reads the observation, earlier context copies, and itself. The paper states the rule only for prediction copies, so treat that as our reading.
And beyond compute, the single pass has a quieter benefit: all K + 1 losses in Equation (4) come from the same example at the same time, so the gradient for "predict DINO from tracks" and "predict actions from everything" are computed on identical context.
Rows are the blocks doing the reading (queries); columns are the blocks being read (keys). Filled cells are allowed. Tap any cell for the reason. Switch between ModAR's training sequence, ModAR at inference (scrub the stage), Unified and Disjoint. Then press Introduce a leak and read what would go wrong.
Three patterns are worth finding in the matrix. The prediction-copy rows have a staircase of allowed context cells: each one reads one more context block than the row above. The context-copy columns are read by later prediction copies only, never by the prediction copy of the same modality. And the action row reads every context copy, which is Equation (1)'s last factor drawn as a mask.
ORDER = ["tracks", "dino", "depth", "rgb", "act"] SIZE = {"obs": 576, "tracks": 384, "dino": 384, "depth": 384, "rgb": 384, "act": 16} # the training sequence: observation, clean context copies, noisy prediction copies blocks = [("obs", "obs")] + [("ctx", m) for m in ORDER[:-1]] \ + [("noisy", m) for m in ORDER] def allowed(q, k): (qk, qm), (kk, km) = q, k if kk == "obs": return True # everyone reads the current observation if qk == "obs": return False # the observation reads only itself if q == k: return True # a block's own tokens are denoised jointly if kk == "ctx": # clean context: only STRICTLY EARLIER modalities return ORDER.index(km) < ORDER.index(qm) return False # never read another block's noisy copy def token_mask(blocks): # (L, L) boolean, L = 3,664 here spans, s = [], 0 for b in blocks: spans.append((s, s + SIZE[b[1]])); s += SIZE[b[1]] M = torch.zeros(s, s, dtype=torch.bool) for i, q in enumerate(blocks): for j, k in enumerate(blocks): if allowed(q, k): M[spans[i][0]:spans[i][1], spans[j][0]:spans[j][1]] = True return M # used in the SHARED blocks; experts see only their own span
Two implementation notes connect this to the rest of the model. The mask applies in the shared blocks; in the expert blocks, each stream's attention is already restricted to its own stream (and, for context versus prediction copies, the same no-leak rule must hold there too). For an actionless example, the action block has no target, so the action loss term is simply dropped, as the paper specifies.
A reimplementation of ModAR shows the DINO loss collapsing to almost zero within a few thousand steps, while closed-loop success stays at the Action-only level. Using only this chapter, name the most likely bug and the one-line check that confirms it.
Answer: the mask lets the noisy DINO prediction copy read the clean DINO context copy (target leakage). The network copies the answer instead of predicting it, so at inference, where no clean DINO exists, it has learned nothing useful.
Check: assert that allowed(("noisy","dino"), ("ctx","dino")) is False, and more generally that no prediction copy can read a context copy at or after its own position in the order.
There is a quiet correctness property hiding in this chapter, and it is worth stating plainly. At inference, the DINO block is computed while reading exactly {observation, finished tracks}. In training, the noisy DINO copy is computed while reading exactly {observation, tracks context copy}. Same set of blocks, same attention pattern, same weights. The only differences are that training context is a (lightly noised) ground truth instead of a generation, and that training handles every stage in one pass.
This is the block-level version of teacher forcing, the standard trick for training autoregressive models: during training, feed the model the true previous outputs rather than its own guesses, so every step can be supervised in parallel. Teacher forcing has a well-known weakness: at test time the model sees its own imperfect outputs, which it never practiced on. Context noise (Chapter 5) is ModAR's answer to exactly that weakness. The mask makes training and inference structurally identical; context noise makes them statistically closer.
Put the mask and Equation (1) side by side for a single four-modality training example. Every row below is computed in the same forward pass.
| Prediction copy | Reads (besides itself) | Factor of Equation (1) it trains |
|---|---|---|
| noisy tracks | observation | p(Ytr | c) |
| noisy DINO | observation, ctx tracks | p(Ydino | c, Ytr) |
| noisy depth | observation, ctx tracks, ctx DINO | p(Ydep | c, Ytr, Ydino) |
| noisy RGB | observation, ctx tracks, DINO, depth | p(Yrgb | c, Ytr, Ydino, Ydep) |
| noisy actions | observation, all four context copies | p(A | c, Y) (the IDM; only if the example has action labels) |
For an actionless example, the last row simply contributes no loss. Everything above it still trains, which is how a human video teaches "given this motion, what will the features and depth look like?" without ever saying anything about joint commands.
Notice what the table implies for a labeled example: all five rows train at once, and each reads a different amount of context. The tracks row learns to imagine from the observation alone, the hardest possible starting point, while the action row sees the most. If you mask the table differently (for example, letting the depth row read nothing but the observation), you are no longer training Equation (1), and the inference loop, which always hands depth the finished tracks and DINO, would meet inputs the network never learned to use.
The ablation in Chapter 7 reverses the future order to RGB → depth → DINO → tracks, with actions still last. Its factorization is equally exact:
The difference is the first job. In the forward order, the first block to be generated from nothing but the observation is the compact track field (1,152 raw numbers). In the reverse order, it is a full RGB future (225,792 raw numbers), with every lighting and texture detail, and every later block conditions on whatever mistakes that hardest-first step made. Same joint distribution on paper; very different sequence of difficulties in practice. The measured cost is 10 points (75% to 65%).
Table II includes a "Tracks + DINO" model. Write its factorization, list its training-sequence blocks, and count its tokens with the assumed 576 observation tokens.
Answer: p(Ytr, Ydino, A | c) = p(Ytr | c) · p(Ydino | c, Ytr) · p(A | c, Ytr, Ydino). Blocks: observation, ctx tracks, ctx DINO, noisy tracks, noisy DINO, noisy actions. Tokens: 576 + 2 × 384 + 2 × 384 + 16 = 576 + 768 + 768 + 16 = 2,128. At inference it takes 8 × 3 = 24 network passes instead of 40.
Explain in two sentences why ModAR can cache the observation's and finished blocks' keys and values across all later denoising steps without changing the result, and name one mask change that would break this.
Answer: a token's keys and values depend only on the tokens it attends to, and under the block-causal mask the observation and finished blocks never attend to anything generated later, so their keys and values are fixed once computed. Letting the observation (or a finished block) attend to a later, still-changing block would break it: its keys and values would change at every step and the cache would go stale.
Chapter 4 said each block is "denoised over 8 flow steps" and moved on. This chapter opens that box. By the end you will be able to write ModAR's training loss, its sampler, and its defense against cascading errors, with every symbol accounted for.
Take one future modality's true target Y (say, the DINO future: 384 tokens of 384 numbers). Draw a same-shaped block of Gaussian noise ε. Pick a mixing level τ between 0 and 1. The paper forms the noisy version with a straight-line blend, its Equation (2):
Ytm is the clean target of modality m. εm is noise drawn from a standard normal, N(0, I), independently for each modality. τm is the flow timestep, also drawn independently per modality in ModAR. Ỹtm (read "Y tilde") is the noisy input the network sees. At τ = 1 it is the clean target; at τ = 0 it is pure noise. This straight-line path is the linear flow-matching interpolant of Lipman et al. (the flow matching gleam derives it).
One number, to make it concrete (illustrative): a target value Y = 0.8, a noise draw ε = −1.2, and τ = 0.25 give Ỹ = 0.25 × 0.8 + 0.75 × (−1.2) = 0.2 − 0.9 = −0.7. Three quarters noise, one quarter signal.
Given Ỹ and τ, a denoiser can be trained to output one of three equivalent-looking quantities.
| Target | Meaning | Formula |
|---|---|---|
| ε-prediction | "Which noise was added?" | ε |
| v-prediction | "Which way, and how fast, does the path move?" | v = dỸ/dτ = Y − ε |
| x-prediction | "What is the clean answer?" | Y |
Given Ỹ and τ, any one of the three determines the other two, so they look interchangeable. They are not interchangeable for learning. ModAR trains "every output stream with a JiT-style x-prediction objective": the output head directly predicts the clean target, written Ŷtm (read "Y hat"). JiT is the paper "Back to Basics: Let Denoising Generative Models Denoise" by Li and He.
The paper gives both an observation and a reason. The observation: "replacing x-prediction with velocity (v) prediction is often unstable and can cause training to diverge." The reason: "Velocity prediction can struggle in high-dimensional spaces; predicting the clean sample is well-conditioned when the data lie on a low-dimensional manifold, as with images, depth, and other visual representations."
Unpack "low-dimensional manifold" with an analogy. Every possible 384-token DINO map is a point in an enormous space, but the maps that real scenes produce occupy a thin, structured sheet inside it, the way all real faces occupy a tiny corner of the space of all pixel grids. A clean target always lies on that sheet, so predicting it asks the network to land on something structured. Noise and velocity are spread across the whole space, so predicting them asks the network to reproduce full-dimensional randomness. The paper applies the same x-prediction parameterization to robot actions, citing prior VLA and WAM work that does likewise.
A sampler needs a velocity to move along, but the network outputs a clean guess. The paper converts: vθ = (Ŷ − Ỹ) / (1 − τ). Here is where that comes from, in three lines.
For each supervised modality the paper minimizes its Equation (3):
E is the average over training examples, noise draws and timesteps. The numerator is the squared error of the clean guess. The denominator divides by (1 − τ) squared, except that it never lets that quantity fall below δm, which the paper says "stabilizes the loss near the clean endpoint." The paper sets δm = δact = 0.05.
Why divide at all? Because that denominator turns the clean-guess error into a velocity error. Here is the derivation (ours, but it follows directly from the paper's two formulas).
So the network outputs a clean guess (the well-conditioned target), but it is graded in velocity units, which is what the sampler actually uses. That combination, x-prediction with a velocity-space loss, is the JiT recipe the paper cites. The δ clamp exists because as τ approaches 1 the factor 1/(1 − τ)2 explodes, and tiny errors on almost-clean inputs would dominate every batch.
Near the clean end, the loss insists on precision; near the noisy end, it forgives. That matches intuition: when the input is almost clean, a good model has no excuse for being wrong.
Each modality's loss and the action loss are simply added, the paper's Equation (4), with unit weights:
Lact has the same form as Equation (3), with the action chunk At, its prediction, its own timestep τact and δact in place of the modality versions. The action term is "include[d] only for examples with action labels," which is how actionless human or simulated demonstrations contribute: they supply every future term and nothing else.
Training needs a τ for every stream in every example. Uniform sampling would spend equal effort everywhere on the path. ModAR instead samples logit-normal timesteps: draw n from a normal distribution N(μ, σ), then squash it, τ = 1 / (1 + e−n). μ slides where the practice concentrates; σ sets how spread out it is. The paper uses a different (μ, σ) for each stream.
| Stream | (μ, σ) | Median τ = sigmoid(μ) | Middle ~68% of τ (sigmoid(μ ± σ)) |
|---|---|---|---|
| DINO | (−2, 1) | 0.119 | 0.047 to 0.269 |
| Tracks | (0, 1) | 0.500 | 0.269 to 0.731 |
| Depth | (−1, 1.6) | 0.269 | 0.069 to 0.646 |
| RGB | (−1, 1) | 0.269 | 0.119 to 0.500 |
| Actions | (−1, 1) | 0.269 | 0.119 to 0.500 |
The last two columns are our arithmetic. One of them in full: sigmoid(−2) = 1 / (1 + e2) = 1 / (1 + 7.389) = 1 / 8.389 = 0.1192. So half of all DINO training examples sit at τ below 0.12: the DINO stream practices mostly in the noisy regime, where the big structural decisions happen. Tracks practice centered on the middle of the path, and depth's larger σ spreads its practice widely. The paper reports these settings without explaining how they were chosen, so we will not invent a rationale.
For the baselines: "Unified uses one shared flow timestep with (μ, σ) = (−1, 1); Independent-noise samples each stream's modality-specific distribution independently." That second sentence is the concrete version of the mismatch from Chapter 1: at inference all streams share one τ, but in training a DINO stream near 0.12 and a tracks stream near 0.5 are the typical case.
At inference, each stream starts as pure noise at τ = 0 and is integrated to τ = 1 "using the velocity implied by its clean prediction." Euler integration is the simplest way to follow a velocity: take the current point, add velocity times step size, repeat. With 8 steps the step size is 1/8 = 0.125:
Illustrative numbers: the true clean value is 0.8 and the starting noise is −1.2. First, a perfect network (it always guesses 0.8).
A perfect x-predictor walks a straight line at constant speed, which is what "linear interpolant" promised. Now an imperfect network whose guess is biased high by 0.3 × (1 − τ), a toy error that shrinks as the input gets cleaner.
| k | τ | guess Ŷ | current y | v = (Ŷ − y)/(1 − τ) | next y |
|---|---|---|---|---|---|
| 0 | 0.000 | 1.1000 | −1.2000 | 2.3000 / 1.000 = 2.3000 | −0.9125 |
| 1 | 0.125 | 1.0625 | −0.9125 | 1.9750 / 0.875 = 2.2571 | −0.6304 |
| 2 | 0.250 | 1.0250 | −0.6304 | 1.6554 / 0.750 = 2.2071 | −0.3545 |
| 3 | 0.375 | 0.9875 | −0.3545 | 1.3420 / 0.625 = 2.1471 | −0.0861 |
| 4 | 0.500 | 0.9500 | −0.0861 | 1.0361 / 0.500 = 2.0721 | 0.1730 |
| 5 | 0.625 | 0.9125 | 0.1730 | 0.7395 / 0.375 = 1.9721 | 0.4195 |
| 6 | 0.750 | 0.8750 | 0.4195 | 0.4555 / 0.250 = 1.8221 | 0.6472 |
| 7 | 0.875 | 0.8375 | 0.6472 | 0.1903 / 0.125 = 1.5221 | 0.8375 |
Look at the last row. With evenly spaced steps, the final Euler step multiplies (Ŷ − y) by 0.125 / (1 − 0.875) = 1, so the output lands exactly on the network's last clean guess, 0.8375. The final error (0.0375) is whatever the network gets wrong at τ = 0.875, which is precisely where Equation (3) weights errors 64 times more heavily than at τ = 0 (1 / 0.1252 = 64). The loss and the sampler are pulling in the same direction.
Now the danger from Chapter 1, in this chapter's language. During training, the context copies are the true futures. During inference, the context blocks are the model's own generations, carrying errors like the 0.0375 above. A model trained only on perfect context has never seen imperfect context, so it may trust a slightly wrong track block completely and build a wrong DINO block on top of it, then a wrong depth block, then wrong actions. The paper: "imperfections in earlier generations become inputs to all later predictions."
The fix, "following Latent Forcing," is to make the training context imperfect on purpose. When predicting mk, for every context block mj and every training example, independently sample
U(1 − β, 1) is a uniform draw between 1 − β and 1. β is the maximum context-noise level; the paper sets β = 0.5, so every context block is somewhere between half noise (τ = 0.5) and perfectly clean (τ = 1), with an average mixing level of 0.75. This noise is applied "only during training, not during inference."
Denoise: one row of a toy depth future is integrated from noise over 8 Euler steps with an x-predicting network; raise the predictor error and scrub the steps. Cascade: a toy chain of four stages where an error in the first block propagates; flip context-noise training on and off. Sampler: the paper's five logit-normal timestep distributions and the loss weight curve; drag δ and draw a batch. The paper's own numbers are labeled; everything else is a toy.
Everything above is trained with a conventional, fully specified recipe. Every item in this table is from the paper's implementation section; the right column's arithmetic is ours.
| Setting | Value | What it works out to |
|---|---|---|
| Optimizer | AdamW, learning rate 10−4, betas (0.9, 0.95), weight decay 0.1 | Standard transformer settings |
| Gradient clipping | 1.0 | Caps the gradient norm per step |
| Global batch | 48 | Equal-sized batches from the action-labeled and actionless pools, losses weighted equally (presumably 24 + 24) |
| Precision | bfloat16 | Half-width floats with full dynamic range |
| EMA | decay 0.999 | Weights averaged over roughly 1 / (1 − 0.999) = 1,000 recent steps |
| Warmup | 48,000 samples, linear, then constant LR | 48,000 / 48 = 1,000 optimizer steps |
| Length | 1.2M optimizer steps, all methods | 1,200,000 × 48 = 57.6M training samples |
| Evaluation | Every 100,000 steps; report best checkpoint | 12 checkpoints per run, on 50 held-out initial conditions per task |
An EMA (exponential moving average) keeps a slowly updated copy of the weights, each step moving it 0.1% toward the live weights; evaluating the EMA copy smooths out step-to-step noise. The equal-sized batch rule matters more than it looks: at D = 1,250, actionless demonstrations outnumber labeled ones 1,200 to 50, yet each optimizer step still sees half labeled data, so the action loss is never starved.
def logit_normal(mu, sigma, n): return torch.sigmoid(mu + sigma * torch.randn(n)) TS = {"dino": (-2, 1), "tracks": (0, 1), "depth": (-1, 1.6), "rgb": (-1, 1), "act": (-1, 1)} BETA, DELTA = 0.5, 0.05 def training_step(batch): B = batch.size ctx, noisy, taus = {}, {}, {} for m in ORDER: # tracks, dino, depth, rgb, act Y = batch.target[m] # (B, N_m, c_m) clean future if m != "act": # context copy + CONTEXT NOISE t_c = torch.empty(B, 1, 1).uniform_(1 - BETA, 1) ctx[m] = t_c * Y + (1 - t_c) * torch.randn_like(Y) tau = logit_normal(*TS[m], B).view(B, 1, 1) # independent per stream noisy[m] = tau * Y + (1 - tau) * torch.randn_like(Y) # Eq. (2) taus[m] = tau x_hat = net(batch.obs, batch.q, batch.task, ctx, noisy, taus, mask=BLOCK_CAUSAL) # one pass, every stage loss = 0.0 for m in ORDER: w = 1.0 / (1 - taus[m]).clamp_min(DELTA) ** 2 # Eq. (3) weight err = ((x_hat[m] - batch.target[m]) ** 2).sum((1, 2)) per_ex = w.view(-1) * err if m == "act": per_ex = per_ex * batch.has_actions # actionless rows: no action term loss = loss + per_ex.mean() # Eq. (4), unit weights return loss @torch.no_grad() def act(obs, q, task, steps=8): cache = net.encode_context(obs, q, task) # KV for the observation for m in ORDER: y = torch.randn(SHAPE[m]) # tau = 0: pure noise for k in range(steps): tau = k / steps x_hat = net.denoise(m, y, tau, cache) # reads cached earlier blocks y = y + (x_hat - y) / (1 - tau) / steps # Euler on v = (x_hat - y)/(1 - tau) if m == "act": return y # (16, 14) joint targets cache = net.append_context(cache, m, y) # re-embed, no context noise
Two details in that sketch are ours and two are the paper's. Ours: the shapes of the noise tensors and the exact cache API. The paper's: context noise is sampled independently per context block and per example, and it is switched off at inference.
Five formulations, three data scales, six tasks, fifty trials each. That is 5 × 3 × 6 × 50 = 4,500 potential evaluation episodes behind one figure; Action-only exists only at the smallest scale, which removes 2 × 6 × 50 = 600 of them, leaving 3,900. This chapter reads that figure, its per-task table, and the two control experiments the authors ran to defend it against the obvious objections.
The benchmark is RoboTwin 2.0, a simulated bimanual manipulation benchmark with strong domain randomization. The paper picks "six representative RoboTwin tasks": dump bin, pick bottles, place bread, put bottles, stack bowls, and turn switch.
The data design is the clever part. Each task uses D ∈ {50, 250, 1250} total training demonstrations. Exactly 50 of them keep their action labels. The remaining D − 50 are used as actionless demonstrations: their actions are withheld, so they can only supervise futures. The number of action-labeled demonstrations therefore never changes. Only the amount of "video-like" data grows.
| Total demos per task D | Action-labeled | Actionless | Ratio (actionless : labeled) |
|---|---|---|---|
| 50 | 50 | 0 | 0 : 1 |
| 250 | 50 | 200 | 4 : 1 |
| 1,250 | 50 | 1,200 | 24 : 1 |
"For each method and data scale, we train one multitask model for all six tasks," and every model is evaluated "on 50 held-out initial conditions per task." So each overall number is an average over 6 × 50 = 300 episodes. Because the labeled count is fixed, any improvement from D = 50 to D = 1,250 is attributable to actionless data. Action-only cannot use actionless data, so it has a result only at D = 50 (the figure draws it as a flat dashed line at 0.46).
Success rates by formulation and scale, copied from the paper. Bold marks the highest value in each row.
| Task | D | Action-only | Indep.-noise | Disjoint | Unified | ModAR |
|---|---|---|---|---|---|---|
| dump bin | 50 | 0.82 | 0.64 | 0.90 | 0.88 | 0.92 |
| 250 | — | 0.68 | 0.92 | 0.92 | 0.94 | |
| 1250 | — | 0.64 | 0.78 | 0.94 | 0.98 | |
| pick bottles | 50 | 0.30 | 0.40 | 0.44 | 0.58 | 0.48 |
| 250 | — | 0.38 | 0.42 | 0.60 | 0.60 | |
| 1250 | — | 0.38 | 0.38 | 0.62 | 0.58 | |
| place bread | 50 | 0.24 | 0.20 | 0.46 | 0.44 | 0.52 |
| 250 | — | 0.14 | 0.48 | 0.50 | 0.74 | |
| 1250 | — | 0.28 | 0.22 | 0.44 | 0.72 | |
| put bottles | 50 | 0.32 | 0.20 | 0.36 | 0.52 | 0.64 |
| 250 | — | 0.20 | 0.36 | 0.62 | 0.82 | |
| 1250 | — | 0.18 | 0.42 | 0.56 | 0.82 | |
| stack bowls | 50 | 0.64 | 0.36 | 0.58 | 0.76 | 0.76 |
| 250 | — | 0.18 | 0.62 | 0.76 | 0.68 | |
| 1250 | — | 0.04 | 0.62 | 0.68 | 0.80 | |
| turn switch | 50 | 0.44 | 0.44 | 0.58 | 0.60 | 0.66 |
| 250 | — | 0.48 | 0.48 | 0.60 | 0.74 | |
| 1250 | — | 0.46 | 0.48 | 0.58 | 0.64 | |
| overall | 50 | 0.46 | 0.37 | 0.55 | 0.63 | 0.66 |
| 250 | — | 0.34 | 0.55 | 0.67 | 0.75 | |
| 1250 | — | 0.33 | 0.48 | 0.64 | 0.76 |
Notice what the paper claims and what it does not. It claims ModAR "achieves the highest average success rate at every data scale," which the overall rows confirm. It does not claim ModAR wins every task: Unified is ahead on pick bottles at 50 and 1,250 and on stack bowls at 250. The biggest ModAR margins are on place bread (0.74 vs 0.50 at 250) and put bottles (0.82 vs 0.62 at 250).
The overall number is the plain mean of the six task numbers. Checking two cells by hand confirms the table is internally consistent and shows how much each task moves the average.
That last line is the kind of reading the paper's summary sentence hides. ModAR's average advantage is real, but it is concentrated in tasks where, plausibly, a precise sense of where objects will end up matters most; the paper itself makes the general point that "the best modality set naturally varies across tasks."
A good habit with any controlled study is to list every difference between the compared systems and check which ones were equalized. Here is that list for ModAR versus Unified (the controlled items are from the paper; the "not separately tested" items are our observations).
| Possible confound | Status |
|---|---|
| Backbone, labeled data, optimization budget | Equalized by design |
| Sampling budget (40 vs 8 passes) | Tested: baselines at 40 steps do not close the gap |
| Action predictor reading clean vs noisy futures | Tested: a shared separate IDM preserves ModAR's lead |
| Checkpoint selection | Same rule for all (best of the evaluated checkpoints) |
| Timestep distributions (ModAR per-modality; Unified one shared (−1, 1)) | Not separately tested |
| Context noise (a ModAR-only ingredient) | Part of the method; its removal is ablated within ModAR (75% to 63%) |
The untested row is worth noticing without over-weighting: per-modality timestep distributions are also used by Independent-noise, which performs worst, so they are clearly not sufficient on their own for good results.
Holding the action-labeled demonstrations fixed is what makes Table I a clean test of actionless data. If labeled data grew with D, any formulation would improve simply from more action supervision, and the scaling column would mix two effects. With 50 labeled demonstrations everywhere, the only way a formulation can improve with D is by turning actionless futures into better actions, which is precisely the WAM promise under test.
One consistency check links the two tables: ModAR's row in Table I (0.66, 0.75, 0.76) is identical to the last column of Table II, because both are the same four-modality model. When two tables in a paper agree like this, you can use either as a cross-check on the other.
Overall: 0.66, 0.75, 0.76 at D = 50, 250, 1,250, against Unified's 0.63, 0.67, 0.64. The paper's explanation is the scratchpad hypothesis: "early modalities act as scratchpads for later ones: generating coarser or easier targets first provides structured context for more detailed targets."
"ModAR benefits most from additional actionless data: adding 1,200 actionless demonstrations improves the average success rate from 66% to 76% (10 percentage points), compared with a mere 1% improvement for Unified." From the table: ModAR 0.66 → 0.76 is +0.10; Unified 0.63 → 0.64 is +0.01. Most of ModAR's gain arrives by D = 250 (+0.09); the next 1,000 actionless demonstrations add one more point.
The paper reports this gain without explaining it. One plausible reading (ours, not the paper's) connects it to the IDM framing from Chapter 4. Actionless data trains the future blocks, and in ModAR the action step reads the observation, configuration and task plus every finished future block. If actionless data makes those blocks better, the action step has better evidence to work from, while Unified's action stream only ever reads partially noisy futures. Control 2 below is consistent with that reading, since it shows ModAR's futures are themselves more useful for acting, but it does not test the mechanism directly.
Disjoint beats Action-only on average (0.55 vs 0.46 at D = 50), suggesting that future prediction can help even as an auxiliary loss (by the rough error bars later in this chapter, a 9-point gap on 300 episodes is about two standard errors). But then it declines: 0.55, 0.55, 0.48. The paper's hypothesis is negative transfer: "with more actionless data, the shared representation becomes increasingly shaped by the future-observation prediction objective," to the detriment of the action objective, which in Disjoint never gets to read the futures that representation was shaped for. The steepest drops are place bread (0.48 to 0.22) and dump bin (0.92 to 0.78).
Independent-noise is the worst formulation at every scale (0.37, 0.34, 0.33), below even Action-only, and it also gets worse with more data. On stack bowls it falls from 0.36 to 0.04. Recall the paper's hypothesis: independently sampled noise levels rarely match the synchronized schedule used at test time, so the model gets "insufficient training signal near its inference regime."
How rarely? Here is a toy estimate (ours, not the paper's). Call a training example "inference-like" if all five streams' timesteps fall within 0.1 of each other. For five independent uniform draws, the probability that max − min ≤ w is 5w4 − 4w5.
The paper does not quantify this, and "within 0.1" is our arbitrary threshold. But the orders of magnitude make the hypothesis concrete: the Independent-noise model almost never practices the exact situation it faces at every inference step. The training view of Sim 2 in Chapter 1 shows the same thing one draw at a time.
Lines show success versus total demonstrations (50 labeled in every case). Pick a task or the overall average. The bars below compare ModAR with Unified per task at the selected scale. Overlays (overall only): 40 Euler steps re-runs the simultaneous baselines with ModAR's sampling budget at D = 250; Separate IDM swaps both ModAR's and Unified's action predictors for one shared, separately trained inverse dynamics model. All values are the paper's. Tap a point to read it.
ModAR uses 8 Euler steps per stream: 40 in total across four futures and actions. Unified, Independent-noise and Disjoint use 8 in total. So the paper re-evaluated each of those three at D = 250 with 40 Euler steps, "matching ModAR's total number of sampling steps."
| Formulation (D = 250) | 8 steps | 40 steps | Change |
|---|---|---|---|
| Unified | 67% | 59% | −8 |
| Disjoint | 55% | 54% | −1 |
| Independent-noise | 34% | 37% | +3 |
| ModAR (reference, 8 per stream) | 75% | — | |
"Additional steps do not close the gap." Unified actually gets worse with more steps. The paper does not explain that drop; a finer integration grid means the model is evaluated at τ values it saw with different frequency during training, which is at least one candidate explanation, but it is our speculation. The conclusion stands either way: ModAR's advantage is not a sampling-budget artifact.
The second objection is subtler. ModAR's action step reads finished futures; Unified's reads half-finished ones. Perhaps ModAR's futures are no better than Unified's, and it only wins because its action predictor has the easier job.
The test: train a separate IDM "to map ground-truth future observations to actions," with context noise on its inputs during training. Then take the best ModAR and Unified checkpoints, let each generate its futures and actions normally at every prediction step, "discard its native action prediction, pass its predicted futures to the separate IDM, and execute the resulting actions." Now both systems use the same action predictor, and the only difference is the quality of the futures they hand it.
| D (total demos) | Unified + separate IDM | ModAR + separate IDM | Gap |
|---|---|---|---|
| 50 | 0.55 | 0.59 | +0.04 |
| 250 | 0.63 | 0.69 | +0.06 |
| 1,250 | 0.61 | 0.70 | +0.09 |
(Values read from the bar labels of the paper's Figure 5.) "ModAR still outperforms Unified with the separate IDM, suggesting that its predicted futures are themselves more useful for predicting actions." Notice also that the gap widens with actionless data, echoing Finding 2. Both systems score a little lower with the separate IDM than with their native action heads (for ModAR, 0.69 vs 0.75 at D = 250). Our explanation, not the paper's: the native heads were trained jointly with the futures they read, while the separate IDM was trained on ground-truth futures.
The paper reports no confidence intervals, so here is a rough binomial estimate of our own, treating each of the 300 episodes as an independent coin flip (it ignores task structure and best-checkpoint selection, so read it as a scale, not a verdict).
This is exactly why the authors word the Flex-π result as "slightly higher observed," and why the formulation claim rests on consistency across three scales, two controls and a real-robot replication rather than on any single cell.
The overall row is an average of six quite different curves. Reading them one by one (numbers from Table I, interpretations ours) shows where each formulation's character comes from.
Dump bin is the easy task: even Action-only reaches 0.82. ModAR climbs to 0.98 at D = 1,250. Disjoint is the only formulation that drops sharply here (0.92 to 0.78), a first sign of its negative-transfer problem.
Pick bottles is Unified's task: 0.58, 0.60, 0.62 against ModAR's 0.48, 0.60, 0.58. It is the clearest case in the table where joint generation is at least as good.
Place bread is ModAR's biggest win: 0.74 and 0.72 at the larger scales, against 0.50 and 0.44 for Unified. Placing an object precisely is a plausible place for a finished geometric future to pay off, and Chapter 7 shows depth-containing models do especially well here.
Put bottles is the cleanest data-scaling story: ModAR goes 0.64, 0.82, 0.82, while Independent-noise sits at 0.20, 0.20, 0.18.
Stack bowls is where Independent-noise collapses (0.36, 0.18, 0.04) and where Unified edges ModAR at D = 250 (0.76 vs 0.68) before ModAR retakes the lead at 1,250 (0.80 vs 0.68).
Turn switch shows ModAR peaking at D = 250 (0.74) and falling back at 1,250 (0.64), a reminder that "more actionless data" is not guaranteed to help on every task even for the best formulation.
These are the per-task realities behind the averages: for two of the baselines, feeding in more actionless data makes the robot fail more often, while for ModAR it mostly makes the robot succeed more often.
Keep the same caution you applied to the averages. With 50 episodes per task, one rate carries a standard error of up to √(0.5 × 0.5 / 50) = √0.005 = 0.071, about 7 points. A 16-success collapse (Independent-noise on stack bowls, 18 to 2) is far outside that; a 2-point wobble on a single task is not. Read the per-task stories as patterns to check, not as verdicts.
The paper is careful to say that Independent-noise "performs poorly in our from-scratch experiments." Keep that qualifier. Flex-π, which also samples noise levels independently per stream during training, reaches 72% on the same data in Chapter 8, starting from a pretrained video model. So the collapse in Table I is evidence about training this formulation from scratch at this scale, not a verdict on independent noise sampling in general. The train-test mismatch hypothesis may simply matter much less when the backbone already knows how to denoise video.
The paper's figure draws each formulation as a colored bar at each of the three scales, with Action-only as a dashed horizontal line at 0.46. The visual story: the ModAR bars rise from left to right and are the tallest in every group; Unified is flat; Disjoint sags at the right; Independent-noise sits below the dashed line throughout. Sim 7 above redraws the same numbers as lines so the trends are easier to compare.
(1) ModAR's average success at D = 50, 250, 1,250. (2) What happened to Unified when given 40 Euler steps. (3) The separate-IDM result at D = 1,250.
Answers: (1) 0.66, 0.75, 0.76. (2) It dropped from 67% to 59%. (3) ModAR + IDM 0.70 versus Unified + IDM 0.61.
A lab has 50 teleoperated demonstrations per task and a large pile of task videos without actions. Using only Table I, which formulation should it train, and which should it avoid?
One defensible answer: ModAR, because it is the only formulation whose average keeps rising as actionless data grows (0.66 → 0.75 → 0.76) and the controls show its gain is not a sampling-budget or action-head artifact. Avoid Independent-noise (0.37 → 0.33) and be wary of Disjoint (0.55 → 0.48): in this study both got worse as actionless data was added. Unified is a reasonable fallback but gains almost nothing from the videos (0.63 → 0.64).
Chapter 6 fixed the modalities and varied the formulation. This chapter does the opposite: it fixes the formulation (ModAR) and varies what it predicts. It is the chapter that decides whether the headline "RGB is not the right thing to imagine" survives contact with data, and whether each structured modality earns its place.
Go back to the prediction you wrote at the end of Chapter 0. Which single modality gives the best policy alone at 50 demonstrations, and which one hurts most to remove? Both answers are in this chapter.
There are two natural experiments, and the paper runs both.
The first is additive: start from one modality and add others one at a time, following the generation order (tracks, then DINO, then depth, then RGB). Alongside, train a variant for each modality predicted alone. That is Figure 4(b) and Table II, run at all three data scales.
The second is leave-one-out: start from the full four-modality model and remove exactly one modality. That is part of Figure 6, run at D = 250. The two experiments can disagree, and when they do, the disagreement is informative.
Every row is one task and one scale; every column is a set of predicted modalities, all trained as ModAR. Bold marks the best value in each row. The RGB column is, in the paper's words, "the typical WAM setting of predicting only future RGB and actions."
| Task | D | Action-only | RGB | Depth | Tracks | DINO | T+DINO | T+DINO+Depth | T+DINO+Depth+RGB |
|---|---|---|---|---|---|---|---|---|---|
| dump bin | 50 | 0.82 | 0.80 | 0.90 | 0.90 | 0.96 | 0.92 | 0.92 | 0.92 |
| 250 | — | 0.94 | 0.92 | 0.92 | 0.94 | 0.94 | 0.96 | 0.94 | |
| 1250 | — | 0.98 | 0.88 | 0.90 | 0.96 | 1.00 | 0.94 | 0.98 | |
| pick bottles | 50 | 0.30 | 0.18 | 0.44 | 0.12 | 0.44 | 0.58 | 0.48 | 0.48 |
| 250 | — | 0.32 | 0.48 | 0.48 | 0.62 | 0.62 | 0.58 | 0.60 | |
| 1250 | — | 0.30 | 0.40 | 0.42 | 0.62 | 0.70 | 0.66 | 0.58 | |
| place bread | 50 | 0.24 | 0.26 | 0.38 | 0.10 | 0.44 | 0.52 | 0.54 | 0.52 |
| 250 | — | 0.34 | 0.64 | 0.26 | 0.56 | 0.64 | 0.76 | 0.74 | |
| 1250 | — | 0.36 | 0.60 | 0.38 | 0.54 | 0.64 | 0.72 | 0.72 | |
| put bottles | 50 | 0.32 | 0.20 | 0.12 | 0.30 | 0.64 | 0.48 | 0.48 | 0.64 |
| 250 | — | 0.38 | 0.48 | 0.48 | 0.60 | 0.66 | 0.74 | 0.82 | |
| 1250 | — | 0.54 | 0.42 | 0.74 | 0.64 | 0.66 | 0.90 | 0.82 | |
| stack bowls | 50 | 0.64 | 0.58 | 0.54 | 0.74 | 0.66 | 0.70 | 0.72 | 0.76 |
| 250 | — | 0.72 | 0.74 | 0.64 | 0.66 | 0.86 | 0.78 | 0.68 | |
| 1250 | — | 0.66 | 0.64 | 0.66 | 0.84 | 0.74 | 0.80 | 0.80 | |
| turn switch | 50 | 0.44 | 0.58 | 0.36 | 0.30 | 0.46 | 0.58 | 0.62 | 0.66 |
| 250 | — | 0.56 | 0.54 | 0.42 | 0.54 | 0.58 | 0.68 | 0.74 | |
| 1250 | — | 0.64 | 0.40 | 0.56 | 0.48 | 0.60 | 0.58 | 0.64 | |
| overall | 50 | 0.46 | 0.43 | 0.46 | 0.41 | 0.60 | 0.63 | 0.63 | 0.66 |
| 250 | — | 0.54 | 0.63 | 0.53 | 0.65 | 0.72 | 0.75 | 0.75 | |
| 1250 | — | 0.58 | 0.56 | 0.61 | 0.68 | 0.72 | 0.77 | 0.76 |
The last column equals ModAR's column in Table I (0.66, 0.75, 0.76), as it must: it is the same four-modality model.
Answer to the first prediction: at D = 50, DINO is the best single modality by a wide margin (0.60, against 0.46 for depth, 0.43 for RGB and 0.41 for tracks). It stays the best single modality at 250 (0.65) and 1,250 (0.68). A semantic representation of "what is where" seems to be the most useful single thing to imagine.
RGB alone is weak: 0.43 at D = 50, which is below Action-only (0.46). Imagining the future in pixels, with this small a model and no pretraining, does not help at all at the smallest scale; it only pulls ahead of Action-only once actionless data is added (0.54, then 0.58).
Tracks alone are the weakest single modality at D = 50 (0.41), with some striking task-level failures: 0.12 on pick bottles, 0.10 on place bread. Keep that in mind for the leave-one-out result below.
Following the cumulative columns and subtracting neighbors (our arithmetic on the overall rows):
Three things jump out. Adding DINO to tracks is the biggest single step at every scale. Depth's contribution grows with data (0.00, 0.03, 0.05). And RGB's contribution is +0.03, +0.00, −0.01: in the paper's words, "additionally predicting future RGB on top of the first three modalities provides no consistent gain."
The paper's summary of the additive experiment: "Across data scales, performance generally holds or increases with each added modality, showing that ModAR can effectively combine the benefits of predicting multiple modalities." And it is careful about the exceptions: "the best modality set naturally varies across tasks because different features define each task, and some modalities represent those features better than others; as a result, additional modalities do not always help."
You can see those exceptions in the table. Put bottles at 1,250 peaks at 0.90 with tracks, DINO and depth, and drops to 0.82 when RGB is added. Stack bowls at 250 peaks at 0.86 with only tracks and DINO. Dump bin at 1,250 reaches 1.00 with only tracks and DINO.
The paper's hypothesis is short: "predicting future RGB introduces high-variance appearance details while adding little information beyond the more structured targets in our setting." Unpack the two halves.
High variance. RoboTwin 2.0's own title advertises "strong domain randomization." The paper does not say which randomization settings its six tasks used, but wherever lighting, texture or clutter varies between episodes, the RGB target is where that variation lands, and the loss charges the model for failing to guess it. Little additional information. By the time RGB is generated, the model already has motion, semantics and geometry. What RGB adds on top is mostly appearance, which, on the paper's hypothesis, adds little beyond the structured targets for these tasks. Recall from Chapter 2 that RGB is also the largest raw target (225,792 numbers per example), so it is expensive to predict and, per this table, not worth it for these tasks.
Figure 6 takes the full four-modality ModAR at D = 250 (average success 0.75) and changes one thing at a time. Values are read from the figure's bar labels and match the numbers stated in the text.
| Variant (D = 250) | Success | Change vs full | What it tests |
|---|---|---|---|
| ModAR (full) | 0.75 | — | Reference |
| No context noise | 0.63 | −0.12 | Robustness to its own imperfect earlier blocks (Chapter 5) |
| Reverse order (RGB → depth → DINO → tracks, actions still last) | 0.65 | −0.10 | The scratchpad ordering |
| w/o tracks | 0.61 | −0.14 | Value of point tracks |
| w/o DINO | 0.65 | −0.10 | Value of DINO features |
| w/o depth | 0.70 | −0.05 | Value of depth |
| w/o RGB | 0.75 | 0.00 | Value of RGB |
Answer to the second prediction: removing tracks hurts most (0.75 to 0.61), followed by DINO (0.65), depth (0.70), and RGB (no change). The paper's conclusion: "predicting RGB contributes the least of the four modalities to success rate in this setting."
On order, the paper says reversing "lowers the average success rate to 65%, supporting our choice to generate compact, structured representations before increasingly detailed ones." On context noise: "context noise during training is critical for preventing errors from compounding across successive modalities."
Figure 6 reports six-task averages over 300 episodes each. Converting to episodes and applying the same rough binomial check as Chapter 6 (our arithmetic):
On this rough scale, removing tracks, removing context noise, removing DINO and reversing the order are all well outside run-to-run noise; removing depth is suggestive but weaker; removing RGB changes nothing. That ranking matches the paper's prose exactly, including its choice to call context noise "critical."
Put the two experiments side by side for tracks. Alone, tracks are the weakest modality (0.53 at D = 250). Removed from the full set, they cause the largest drop (−0.14 at D = 250).
One reading, consistent with the paper's scratchpad hypothesis, is that tracks are valuable less as a final answer and more as a first draft: compact motion context that makes every later block easier. A caveat the paper does not discuss, and which you should keep: removing tracks also changes which modality goes first (DINO becomes the first block), so the w/o-tracks ablation mixes "no tracks" with "a different first block." The reverse-order ablation points the same direction (putting tracks last costs 10 points), but neither experiment isolates position from content. The paper leaves "a full systematic comparison of modality orderings to future work."
Toggle which futures ModAR predicts, pick a data scale, and optionally reverse the order or remove context noise. If the paper evaluated that exact configuration, you see its real numbers (per task from Table II, or the overall bar from Figure 6); if not, the lab says so and shows the closest evaluated neighbors. The leaderboard at the bottom ranks every configuration the paper evaluated at that scale.
It is worth being as careful as the paper is. The finding is that RGB gives "no consistent benefit" in this setting: small from-scratch models, six simulated tasks, a single camera. It is not a claim that RGB futures are useless in general. A model initialized from a video generator that already understands appearance might extract more from an RGB target, and the paper explicitly frames its conclusion as "a promising alternative to relying solely on future RGB prediction," not a replacement for it.
Similarly, the per-task variation is large. Anyone deploying this on a new task should expect the best subset to differ, and should budget for a small ablation of their own.
You are deploying a ModAR-style policy on a new single-camera arm with a tight latency budget. Using only this chapter's numbers, argue for a modality set and an order, and name the one experiment you would run first on your own tasks.
One defensible answer: tracks → DINO → depth, no RGB. Tracks + DINO is the largest additive step at every scale (+0.22, +0.19, +0.11); depth adds more as data grows (up to +0.05) and is the only explicit geometry with one camera; RGB adds +0.03 / 0.00 / −0.01 while being the largest raw target, which is exactly the paper's real-robot choice. First experiment: the w/o-depth ablation on your tasks, because depth's value (−0.05 when removed) is the least certain of the three you kept.
The paper's general remark, that different tasks are defined by different features, is easy to see in specific cells (numbers from Table II; the interpretations are ours).
Put bottles at D = 50: DINO alone reaches 0.64 while depth alone reaches 0.12. With little data, knowing what is where seems far more useful than knowing how far for this task.
Place bread at D = 250: depth alone reaches 0.64 while tracks alone reach 0.26, and the best model (0.76) is the one that adds depth to tracks and DINO. Placement seems to reward explicit geometry.
Pick bottles: tracks alone go from 0.12 to 0.48 to 0.42 as data grows. Motion prediction on its own looks data-hungry.
Stack bowls at D = 1,250: DINO alone reaches 0.84, the best value in that row, above every multi-modality model.
At D = 1,250 on put bottles, adding RGB takes the model from 0.90 to 0.82. Tempting to conclude RGB hurts. A quick binomial check (ours) with 50 episodes per rate:
A single-task gap of about one standard error is weak evidence either way. That is exactly why the paper phrases its RGB finding as "no consistent gain" across tasks and scales, rather than "RGB hurts." The honest reading is about the pattern: +0.03, 0.00, −0.01 on average, and zero change in the leave-one-out ablation.
Compare DINO's value in the two experiments at D = 250. Added on top of tracks, DINO is worth +0.19 (0.53 to 0.72). Removed from the full model, it costs 0.10 (0.75 to 0.65). If modality contributions simply added up, those two numbers would match. They do not, because what a modality contributes depends on what else is already predicted; plausibly, in the full model, depth and RGB partly cover for a missing DINO block. The paper's summary that tracks, DINO and depth "provide additive gains" is best read as "each one adds something on average," not as a claim that their values sum exactly.
The paper's figure groups bars by scale. Within each group: single-modality models (RGB, depth, tracks, DINO), then the cumulative sets (tracks+DINO, +depth, +RGB), with Action-only as a dashed line at 0.46. The caption singles out the RGB bar as the one that "matches the typical WAM setting of predicting only future RGB and actions." In every group, the cumulative bars stand clearly above the RGB bar.
| Evidence | Direction | Strength |
|---|---|---|
| Reverse order drops 0.75 to 0.65 | Supports structured-first ordering | One comparison, one scale |
| ModAR beats Unified at every scale, also with a shared IDM | Supports sequential conditioning in general | Consistent across scales and a control |
| Adding DINO after tracks is the largest additive step (+0.22, +0.19, +0.11) | Consistent with earlier blocks helping later ones | Also consistent with DINO simply being a strong target |
| Tracks weak alone, yet most costly to remove | Consistent with tracks acting as context | Confounded with block position |
| Other orders never tested | Unknown | Paper leaves it to future work |
The ledger is a fair summary of how the paper words it: the results "support" the hypothesis; they do not prove the chosen order is optimal.
A team trains a small from-scratch WAM that predicts only future RGB plus actions, with 50 demos per task, and it scores slightly below their action-only policy. Their conclusion: "world modeling does not help." What does Table II suggest instead?
Answer: the paper sees the same thing (RGB 0.43 vs Action-only 0.46 at D = 50), but DINO alone reaches 0.60 and tracks + DINO + depth 0.63 in the same setting. The failure is plausibly the choice of target, not world modeling itself; and RGB-only does pull ahead once actionless data is added (0.54, 0.58).
(1) At which data scale does depth first add something on top of tracks + DINO? (2) What is the best single modality at every scale? (3) Which column is the "typical WAM" and how far behind the tracks + DINO + depth model is it at D = 1,250?
Answers: (1) D = 250 (+0.03; at D = 50 it adds 0.00). (2) DINO (0.60, 0.65, 0.68). (3) The RGB column; 0.77 − 0.58 = 0.19 behind.
The cumulative sets follow the generation order. The additive columns grow as tracks, then tracks + DINO, then + depth, then + RGB: the same order the model generates them in. So each cumulative model is the previous one with one more block appended at the end of the sequence, just before the actions. That keeps the experiment clean: adding a modality never changes the position of the ones already there.
The leave-one-out ablations are run at D = 250 only. Figure 6 is explicitly "with 250 total demonstrations per task and 50 action-labeled demonstrations," the scale at which the full model reaches 75%. The paper does not say why that scale was chosen; the additive study in Table II is the one that covers all three scales.
One last habit from this chapter: whenever a paper reports "modality X helps," ask measured how. Alone, added, or removed? At which data scale? Averaged over which tasks? Table II and Figure 6 answer all three, which is why their conclusions can be stated so precisely.
Everything so far compared ModAR with its own siblings: same size, same data, same budget. Two questions remain that a practitioner will ask immediately. Is a 30-million-parameter model trained from scratch even competitive with the large video-pretrained WAMs people actually use? And does any of this survive a real robot, real lighting, and human videos filmed by people rather than rendered by a simulator?
Flex-π is concurrent work by G. Yan, J. Liu, Y. Fan, L. Cai, M. Liao, J. Zhang and D. Fox, titled "A Multi-Stream World-Action Model with Compute Flexibility." It represents the pretrained-video-backbone approach Chapter 0 described, extended (as Chapter 0 also mentioned) to predict several future modalities, RGB latents, 3D pointmaps and DINO features, jointly with actions. It has about 6B parameters.
The paper fine-tunes it on the same RoboTwin data at D = 250, following Flex-π's own recipe as closely as the setting allows.
| Aspect | Flex-π as run in this paper | ModAR (tracks–DINO–depth variant) |
|---|---|---|
| Initialization | Video backbone from Wan2.2-TI2V-5B; weights interpolated to a smaller hidden size for the action expert | Random (no pretraining) |
| What trains | All trainable components, full fine-tuning | Everything |
| What is frozen | VAE, text encoder, DINOv3 encoder | DINOv2 target encoder |
| Future targets | RGB latents, pointmaps, DINO features at t+4, 8, 12, 16 | Tracks, DINO, depth at t+8, 16 |
| Action horizon | H = 16, single head-camera setting | H = 16, single camera |
| Actionless demos | Keep every visual objective, mask the action loss | Supervise available futures, drop the action term |
| Batch × steps | 288 × 30,000 (final checkpoint) | 48 × 1.2M |
| Inference | Joint denoising of all visual streams and actions (full multi-stream mode) | Sequential, 8 Euler steps per stream |
| Parameters | 6B | 30.1M |
| Average success, D = 250 | 72% | 75% |
The paper's reading: "on these in-distribution tasks, a compact WAM trained from scratch can perform as well or better than a much larger video-model-initialized WAM using substantially less training compute." And, in the same paragraph, the caveat: "this is a system-level comparison rather than a controlled architectural comparison, since the models differ in scale, pretraining, target modalities, and training recipe." Chapter 6's binomial aside puts the 3-point gap at under one standard error, which is why "as well or better" is the right phrase.
One more asymmetry, which the table's "final checkpoint" note hints at: ModAR's 75% is its best evaluated checkpoint (the protocol from Chapter 5, best of twelve), while Flex-π's 72% is its final checkpoint after 30,000 steps. Picking the best of several evaluations can only raise a score, so this difference tilts a 3-point comparison toward ModAR. The paper states both protocols; it does not say how Flex-π would score under best-checkpoint selection.
Note also the phrase in-distribution. The evaluation uses held-out initial conditions of the same six tasks the models were trained on. It does not test the thing video pretraining is usually credited for, generalization to new objects, scenes and instructions. The paper's limitations section says as much.
The paper states the ratio (approximately 20×) but not its accounting. We can still check that it is plausible from numbers the paper does give. A FLOP is one floating-point operation; training compute is commonly estimated as roughly 6 × (parameters) × (tokens processed), the "6ND" rule of thumb.
So the reported ratio is the right order of magnitude, and the remaining factor is exactly what the rule of thumb cannot see: how many tokens each model processes per sample, that frozen components (VAE, text encoder, DINOv3) only cost a forward pass, and that Flex-π's action expert runs at a smaller width than its video backbone. We cannot reconstruct the paper's exact accounting, and this example does not pretend to. What it does show is the shape of the trade: ModAR trains on almost seven times more samples, but each sample costs roughly two hundred times fewer parameter-operations.
"On a single NVIDIA GeForce RTX 5090 GPU, end-to-end inference for ModAR takes 147.9 ms (6.76 Hz) when generating all four future-observation modalities and actions."
The paper does not report the latency of the three-modality real-robot configuration or the robot's control rate, so we stop there. The honest summary is the paper's own: sequential generation "increases inference latency relative to simultaneous or action-only generation."
Three "challenging tabletop tasks," performed with bimanual YAM arms: stacking cups, folding a crumpled towel, and placing an object in a drawer and closing the drawer. For each task the authors collect three kinds of data.
| Data (per task) | Count | How it was collected | What it can supervise |
|---|---|---|---|
| Robot demonstrations | 100 | Teleoperation with paired teacher arms, using the RAIDEN toolkit | Futures and actions |
| In-domain human demonstrations | 200 | People performing the task in the same setting, filmed by the same fixed camera | Futures only (actionless) |
| EgoDex demonstrations | 1,000 | From the EgoDex dataset, categories "stack/unstack cups," "basic fold," "insert/remove drawer" | Futures only (actionless, out of domain) |
A fixed third-person ZED stereo camera "provides RGB observations and stereo depth for both robot and actionless human demonstrations." EgoDex is a large-scale egocentric video dataset of human dexterous manipulation; the paper calls its demonstrations "out-of-domain" relative to the robot's setting. As decided in Chapter 7, all real-world WAMs "observe and predict only tracks, DINO, and depth." Evaluation uses "30 rollouts per task, with varied initial object poses," and again one multitask model per method and data mixture.
Two practical details the paper does not spell out, which you would need to decide in a reimplementation: what the robot-configuration input qt is for human videos (which have no robot), and how depth is obtained for the EgoDex clips. We flag them rather than guess.
It helps to see what one human clip contributes, in the terms of Chapter 4. Its frames give an observation and targets at t+8 and t+16, so every future row it has targets for receives a loss (the paper's phrase is "the available future targets": tracks, DINO, and depth wherever depth exists for that clip). The action row receives nothing, because there is no action chunk to compare against. So every human clip is pure future supervision, and whatever it teaches has to reach the robot's arms through the futures the action step later reads.
Why vary the initial object poses across the 30 rollouts? Because a policy evaluated from one fixed start could succeed by replaying a memorized trajectory. Varying where the cups, towel or object begin forces the model to use what it sees, which is exactly where imagined motion, semantics and geometry should matter. (The paper states the protocol; the rationale is the standard one, and ours to spell out.)
Figure 7(a) compares Action-only (100 robot demos per task, since it cannot use actionless data) with Unified and ModAR (both with the 100 robot demos plus 200 in-domain human and 1,000 EgoDex demos). Per-task values are read from the figure's bar labels.
| Task | Action-only | Unified | ModAR |
|---|---|---|---|
| fold towel | 0.50 | 0.70 | 0.83 |
| place in drawer | 0.57 | 0.70 | 0.83 |
| stack cups | 0.50 | 0.60 | 0.83 |
| overall | 52.2% | 66.7% | 83.3% |
"ModAR achieves the highest success rate on all three tasks," and the ordering matches simulation: ModAR above Unified above Action-only.
Figure 7(b) trains three separate ModAR models: robot demos only; plus the 200 in-domain human demos; plus those and the 1,000 EgoDex demos. Human data supervises "future-observation predictions, but not action prediction."
| Task | ModAR (robot only) | + in-domain human | + in-domain + EgoDex |
|---|---|---|---|
| fold towel | 0.67 | 0.80 | 0.83 |
| place in drawer | 0.73 | 0.80 | 0.83 |
| stack cups | 0.70 | 0.83 | 0.83 |
| overall | 70.0% | 81.1% | 83.3% |
The paper's conclusion: ModAR "can benefit from actionless data collected across embodiments and domains."
With 30 rollouts per task, every rate is a count out of 30. Recovering the counts (our arithmetic) makes the results tangible and checks the overall percentages exactly.
Two comparisons across the panels are worth making (both numbers are from the paper, the juxtaposition is ours). First, the like-for-like data comparison with Action-only: both trained on robot demos only, ModAR reaches 70.0% against 52.2%, a gap of 17.8 points without any human data at all. Second, ModAR without human data (70.0%) already edges out Unified with all of it (66.7%).
And the diminishing return from EgoDex (+2 successes from 1,000 out-of-domain demos, versus +10 from 200 in-domain ones) is a useful calibration for anyone planning data collection. With 90 trials per configuration, a 2-success difference is well within noise, so the fair reading is "in-domain human video helped clearly; out-of-domain video did not hurt and may have helped a little."
Compute: the 6ND sanity check from Worked example 11. Drag the token ratio r and watch the implied compute ratio move past the paper's reported ~20×. Real robots: pick a formulation and a data mixture; each dot is one of 30 rollouts per task. Only the combinations the paper ran are available. Dot order is arbitrary; only the counts come from the paper.
The paper's wording is precise: the real-world WAMs "observe and predict only tracks, DINO, and depth." So RGB is dropped on both sides, as a target and as an observation stream. The camera image is still used, because DINO features are computed from it by the frozen encoder; it simply never enters the transformer as raw patches. This saves the largest input projection and 192 observation tokens per query, and it follows directly from the simulation finding that RGB added no consistent benefit "while increasing training and generation costs."
The paper reports only the ratio (~20×), not absolute FLOPs. For a feel of the magnitude, suppose (our assumption) that the compared ModAR variant processes about 3,000 tokens per training example. Then the 6ND rule gives:
Both absolute figures rest on our token assumption and the crude rule; only the ratio is the paper's.
The paper reports 147.9 ms for 40 passes (four futures plus actions). The real-robot model runs 32 passes (three futures plus actions). A crude proportional estimate, ours, and ignoring that the dropped RGB block is one of the larger ones:
Treat that as a ballpark only; the paper measured one configuration on one GPU, and it lists the added latency of sequential generation among its limitations.
Three things the study does not rule out, worth keeping in mind (ours): the ranking could shift on tasks where appearance itself matters (reading a label, sorting by color), since the real-robot models do not predict RGB at all; a larger number of rollouts could shrink or widen the Unified gap, which is about 2.6 standard errors on our rough scale; and the human-video benefit was measured only for ModAR, so it is unknown how much Unified would gain from the same videos in isolation.
It shows that the simulation ranking of formulations reproduces on hardware, on three contact-rich bimanual tasks. It shows that in-domain human video clearly improves ModAR (70.0% to 81.1%). And it shows that adding 1,000 out-of-domain EgoDex clips per task gave a further 2 points (83.3%), which the paper reads as benefit "across embodiments and domains" and which, by our rough estimate below, is within noise.
It does not show generalization to unseen tasks or objects (three tasks, discrete labels, varied initial poses only), and each configuration rests on 90 rollouts. The paper does not report how Unified or Action-only would fare with human data ablated in the same way, beyond the configurations listed. These are exactly the limits Chapter 9 collects.
Using the recovered counts (out of 30 rollouts per task), the gaps become concrete.
| Task | Action-only | Unified | ModAR | ModAR over Unified |
|---|---|---|---|---|
| fold towel | 15 | 21 | 25 | +4 rollouts |
| place in drawer | 17 | 21 | 25 | +4 rollouts |
| stack cups | 15 | 18 | 25 | +7 rollouts |
Stacking cups is where Unified struggles most and ModAR's margin is largest. And the human-video gains per task, from Figure 7(b): fold towel 20 → 24 → 25, place in drawer 22 → 24 → 25, stack cups 21 → 25 → 25. In-domain human video helps every task; EgoDex adds at most one more success on any task.
A rough binomial check with 90 rollouts per configuration (our estimate, treating rollouts as independent):
So, on this rough scale: ModAR over Action-only is very clear, ModAR over Unified is reasonably clear, and the EgoDex increment is indistinguishable from noise. The paper's wording matches: the formulation result "mirror[s] the simulation results," and human data "improves" ModAR, with the big step coming from in-domain video.
Flex-π's recipe initializes its action expert from the video backbone even though the action expert is narrower. The paper describes this as interpolating the pretrained weights "to a smaller hidden size for the action expert." The general idea is to resample each pretrained weight matrix down to the narrower width so the action expert starts from something video-shaped rather than from random numbers. The exact procedure belongs to Flex-π's own paper; the point for us is the contrast: Flex-π's video backbone and action expert start from pretrained video weights, while every parameter of ModAR starts from scratch. (The paper does not say whether Flex-π's smaller input and output layers, for example for pointmaps or actions, are also pretrained.)
| It can tell you | It cannot tell you |
|---|---|
| A 30.1M from-scratch WAM can match a 6B video-initialized WAM on these six in-distribution tasks at D = 250 | Whether sequential generation beats joint generation when both are pretrained |
| The compact model got there with roughly 20× less training compute | How either model behaves on unseen tasks, objects or scenes |
| Structured targets without RGB are enough to compete with a model that also predicts RGB latents | Which of scale, pretraining, targets or recipe explains the 3-point gap (all four differ) |
One more difference hides in the table above: ModAR's DINO targets come from a frozen DINOv2 encoder, while Flex-π keeps a frozen DINOv3 encoder. The two models are not even predicting the same feature space, which is one more reason the paper calls this a system-level comparison.
Without scrolling: (1) the three real tasks, (2) the three data sources and their counts per task, (3) the overall success of Action-only, Unified and ModAR, (4) ModAR's overall success with robot data only, with in-domain human video, and with EgoDex added.
Answers: (1) stacking cups, folding a crumpled towel, placing an object in a drawer and closing it. (2) 100 teleoperated robot demos, 200 in-domain human demos, 1,000 EgoDex demos. (3) 52.2%, 66.7%, 83.3%. (4) 70.0%, 81.1%, 83.3%.
You have one week for data collection on a new bimanual task and already have 100 teleop demos. Options: record 200 in-domain human demos, or download 1,000 matching clips from a public egocentric dataset. Using only Figure 7(b), which do you prioritize?
One defensible answer: the in-domain human demos. In the paper they moved ModAR from 63 to 73 successes out of 90; the 1,000 EgoDex demos added 2 more on top, which is within noise. Out-of-domain video is cheap and did not hurt, so add it if it is free, but the in-domain recordings are where the measured gain was.
You started with a robot that imagined its future in pixels, paying for lamp glints and tablecloth patterns. You now know a design that imagines motion first, then meaning, then distance, and only then acts, and you have checked its evidence line by line. This last chapter draws the boundary of what that evidence covers, connects ModAR to the rest of the site, and points to what to read next.
The paper's limitations section is short and specific. Here it is, point by point, with what each one means for how far you can carry the results.
| Limitation (paper) | What it means in practice |
|---|---|
| "A limited number of tasks" | Six simulated tasks and three real ones. ModAR has the best six-task average at every scale and the best score on all three real tasks, but Unified wins some individual simulated cells (pick bottles, stack bowls); nine tasks are not a population. |
| "Discrete task labels rather than language instructions" | The task embedding g is a lookup table. How the design interacts with language conditioning, which most generalist policies use, is untested. |
| No "broad generalization across tasks, objects, or scenes" | Evaluation varies initial conditions and object poses within trained tasks. The Flex-π comparison is explicitly "in-distribution." |
| Sequential generation "increases inference latency" | 40 network passes per query in the four-modality setting (147.9 ms on an RTX 5090) versus 8 for simultaneous formulations. |
And the future work it proposes: "larger-scale experiments with broader task diversity, generalization evaluation, and training on heterogeneous internet-scale actionless data," plus exploring "the best order for generating modalities, which may depend on the task or specific scenario."
Three more boundaries appear in the body text rather than the limitations section, and they matter just as much.
The Flex-π comparison is system-level. The models "differ in scale, pretraining, target modalities, and training recipe," so the comparison says a compact from-scratch model can keep up on these tasks, not that sequential generation beats a pretrained backbone in general.
Order was tested only against its reverse. "We leave a full systematic comparison of modality orderings to future work." The chosen order beat its reverse by 10 points; whether some third order would beat both is unknown.
The RGB conclusion is scoped. RGB gave "no consistent benefit in our experiments," with from-scratch models. The conclusion frames the result as "a promising alternative to relying solely on future RGB prediction."
These are ours, not the paper's. Each follows from something the paper leaves unmeasured.
| Family | Imagines | Acts how | Example on this site |
|---|---|---|---|
| Policies (no imagination) | Nothing | Observation → action chunk | Diffusion Policy, π0, ACT |
| Action-conditioned world models | Future video given actions | A planner or evaluator chooses actions | DreamX-Phi, Genie, World-In-World |
| Video-first WAMs | Future RGB (latents), jointly or causally with actions | Co-denoised or causal action heads | Causal World Modeling for Robot Control (LingBot-VA) |
| Structured-future models | 3D points, tracks, features | Varies | PointWorld |
| ModAR | Tracks → DINO → depth (→ RGB), in sequence | Inverse dynamics on the finished future | This lesson |
The Causal World Modeling paper in that table (LingBot-VA on this site) is one of the works ModAR cites as predicting future RGB as image latents from frozen video VAEs, which makes it a good contrast read.
| Sim | Paper element | Real numbers | Toy parts |
|---|---|---|---|
| 1 · four modalities | Sec. I, II-A motivation; Fig. 1 | None | Scene, renderer, DINO colors |
| 2 · formulations | Fig. 3; IV-A baselines; IV-E timestep settings | Step counts; logit-normal parameters | Animation timing |
| 3 · horizon and tokens | III-A; IV-E inputs and outputs | H, Δ, J, grid, patch size, action size | Track token layout (3 numbers) |
| 4 · routing and budget | Fig. 2; IV-E architecture | Block counts, widths, 30.1M | adaLN mapping; 12d2 estimate |
| 5 · mask | III-B block-causal generation | The prediction-copy rule; order | Observation and context-copy rows; 576 observation tokens |
| 6 · flow lab | III-B training and inference; IV-E flow settings; Fig. 6 | Equations, δ, β, timestep parameters, 75% vs 63% | Signal, predictor error, cascade gains |
| 7 · Table I explorer | Table I; Figs. 4(a), 5; sampling-step control | All plotted values | None |
| 8 · modality lab | Table II; Figs. 4(b), 6 | All plotted values | None |
| 9 · compute and rollouts | IV-B Flex-π; Fig. 7 | Parameters, steps, batch sizes, ~20×, success rates | Token ratio r; dot order |
| Paper | Why read it after ModAR | Link |
|---|---|---|
| Latent Forcing (Baade et al., 2026) | The source of the scratchpad ordering and of context noise | arXiv:2602.11401 |
| Modality Forcing for Scalable Spatial Generation (Duisterhof, Ramanan, Ichnowski, Johnson, Park, 2026) | Joint image and depth generation from an image prior; shares three authors with ModAR | arXiv:2606.13676 |
| Flex-π (Yan et al., 2026) | The 6B multi-stream rival from Chapter 8 | arXiv:2608.10860 |
| Fast-WAM (Yuan et al., 2026) | The Disjoint formulation: do WAMs need test-time imagination? | arXiv:2603.16666 |
| World Action Models are Zero-shot Policies (DreamZero; Ye et al., 2026) | A joint-generation WAM the Unified baseline represents | arXiv:2602.15922 |
| Turning Video Models into Generalist Robot Policies (VERA; Li et al., 2026) | Futures-then-actions with a video model | arXiv:2605.27817 |
| Point Tracking Improves World Action Models (Guan et al., 2026) | Tracks alongside RGB in a WAM | arXiv:2605.23856 |
| 3PoinTr (Hung, Duisterhof, Ichnowski, 2026) | 3D point tracks for learning from unconstrained human videos | arXiv:2603.08485 |
| WAM4D (Li et al., 2026) | Depth and 4D structure in a fast WAM | arXiv:2606.14048 |
| ST-WAM (Wang et al., 2026) | Semantic-temporal WAM robust to visual distribution shift | arXiv:2607.28993 |
| EgoWAM (Li et al., 2026) | WAMs beyond pixels with in-the-wild egocentric human data | arXiv:2607.08436 |
| RoboTwin 2.0 (Chen et al., 2025) | The simulation benchmark behind Tables I and II | arXiv:2506.18088 |
| Back to Basics: Let Denoising Generative Models Denoise (JiT; Li and He, CVPR 2026) | The x-prediction objective ModAR trains with | CVPR 2026 |
| Unified World Models (Zhu et al., RSS 2025) | Coupled video and action diffusion with independent noise levels | RSS 2025 |
| EgoDex (Hoque et al., ICLR 2026) | The out-of-domain human video used in Chapter 8 | ICLR 2026 |
| Paper section | Lesson chapter |
|---|---|
| Abstract, I Introduction, Fig. 1 | 0 |
| II-A World-action models, Fig. 3 | 1 |
| III-A Problem definition; III-B Modality tokenization; IV-E Inputs and outputs | 2 |
| III-B Shared world-action backbone, Fig. 2; IV-E Architecture | 3 |
| II-B Multimodal generation; III-B Block-causal modality generation, Equation (1), Inference | 4 |
| III-B Injecting context noise, Training, Equations (2)–(4); IV-E Optimization, Flow and sampling | 5 |
| IV-A Simulation setup, Baselines; IV-B Formulation comparison, Scaling, Separate IDM, Sampling steps; Table I, Figs. 4(a), 5 | 6 |
| IV-B Modality comparison; IV-C Ablations; Table II, Figs. 4(b), 6 | 7 |
| IV-B Flex-π comparison, Inference latency; IV-A Real world; IV-D; IV-E Real-world data collection; Fig. 7 | 8 |
| V Conclusion, VI Limitations, related work | 9 |
ModAR's controlled design gives you a checklist for reading its neighbors critically.
| Mistake | Correction |
|---|---|
| "It predicts tracks for all 16 future steps." | Futures are sparse: J = 2 frames, at t+8 and t+16. Only actions are dense (16 steps). |
| "It is autoregressive token by token." | It is autoregressive across modality blocks; each block's tokens are denoised jointly. |
| "Context noise is added at test time for robustness." | It is applied only during training. |
| "The DINO encoder is fine-tuned." | DINOv2 is frozen; it defines the target space. |
| "The paper shows RGB hurts." | It shows RGB gives no consistent benefit in its from-scratch setting. |
| "ModAR beats Flex-π." | It is slightly higher (75% vs 72%, best checkpoint against final checkpoint) in a system-level, in-distribution comparison, with far less compute. |
| "ModAR wins because it samples five times longer." | Baselines given 40 steps do not close the gap. |
| "Human videos train the action head." | They supervise futures only; the action term is dropped for actionless examples. |
For a practitioner: if you have few robot demonstrations and lots of task video, a small WAM that predicts point tracks, DINO features and depth in that order, then actions, is a strong, cheap recipe; skip RGB unless your task depends on appearance, and do not skip context noise.
For a researcher: the paper cleanly separates formulation, target representation and actionless-data scale in a from-scratch study, finds sequential ("modality-autoregressive") generation best at every scale with two controls, and leaves open ordering search, pretrained sequential models and broad generalization.
For a student: a robot does better when it first imagines where things will move, what they are and how far away they are, one step at a time, and only then decides how to move its arms.
| Question | Minimal experiment |
|---|---|
| Is the order optimal? | Train all 24 orderings of the four futures at D = 250 (or a sensible subset) with everything else fixed |
| Position or content? | Keep tracks but move them to the second or last slot; compare with w/o tracks |
| How much context noise? | Sweep β in {0.25, 0.5, 0.75, 1.0} |
| Pretrained and sequential? | Initialize a sequential WAM from a video model and compare with Flex-π under matched compute |
| Language instead of labels? | Replace the task-embedding table with a text encoder and evaluate on held-out instructions |
If you set out to rebuild ModAR from this lesson, here is every specification, split by where it comes from. The right column is honest about what you would have to decide yourself.
| Component | Specified by the paper | You must choose (not in the paper) |
|---|---|---|
| Inputs | Single camera, 168×224; qt = 14-D absolute dual-arm joint configuration; learned task embedding | qt for human videos |
| Targets | H = 16, Δ = 8, J = 2 (t+8, t+16); 16-step, 14-D action chunk; replan every H steps | — |
| Tokens | 12×16 grid of 14×14 patches; patchified RGB and depth; frozen DINOv2 ViT-S/14 patch tokens; CoTracker3 tracks of patch-center queries with displacement + visibility; linear projection to a common width; learned modality embedding; axial RoPE (time, row, column) for visual, time for actions | Track-token normalization and occluded values; RoPE dimension split; exact observation-token layout |
| Backbone | DiT: 6 shared width-384 blocks, 6 heads; experts 2×384 (DINO, depth, RGB), 1×128 (tracks), 2×128 (actions); linear heads; adaLN from qt, g, per-stream τ; 30.1M total for the tracks–DINO–depth variant | MLP ratio; adaLN wiring; the 384-to-128 adapters; the four-modality parameter count |
| Ordering and mask | tracks → DINO → depth → RGB → actions; clean context copy + noisy prediction copy; block-causal mask (prediction copy of mk reads observation + context of m<k) | Rows of the mask for context copies and the observation |
| Objective | Linear interpolant; JiT-style x-prediction; loss ‖Ŷ − Y‖2 / max(1 − τ, δ)2, δ = 0.05; unit loss weights; action term only with labels | Loss reduction (sum versus mean over tokens) |
| Timesteps | Logit-normal (μ, σ): DINO (−2, 1), tracks (0, 1), depth (−1, 1.6), RGB (−1, 1), actions (−1, 1) | — |
| Context noise | τctx ~ U(1 − β, 1), β = 0.5, independent per context block and example; training only | τ signal given to context at inference |
| Optimization | AdamW, LR 10−4, betas (0.9, 0.95), WD 0.1, clip 1.0, batch 48 (equal labeled/actionless), bfloat16, EMA 0.999, 48,000-sample warmup then constant, 1.2M steps; evaluate every 100k, report best | — |
| Sampling | 8 Euler steps per stream; v = (Ŷ − Ỹ)/(1 − τ); KV cache for observation and finished blocks | Step spacing |
| Real robot | YAM arms; RAIDEN teleop; ZED stereo RGB + depth; tracks, DINO, depth only; 100 robot + 200 in-domain human + 1,000 EgoDex demos per task | Depth source for EgoDex clips |
The right column is short, which is a compliment to the paper: almost everything that determines the results is written down.
| Term | Meaning in this lesson |
|---|---|
| WAM | World-action model: one network that jointly models future observations and robot actions |
| Modality | Any representation of the future: RGB, depth, DINO features, point tracks |
| Actionless demonstration | A demonstration without action labels; supervises futures only |
| Stream | One set of tokens generated together (one modality's future, or the action chunk) |
| Flow timestep τ | How clean a stream is: 0 = pure noise, 1 = clean |
| x-prediction | The network outputs its guess of the clean target; the sampler converts it to a velocity |
| Block-causal mask | Attention rule letting each block read only the observation and earlier finished blocks |
| Context copy / prediction copy | The clean (lightly noised) and noisy versions of each target, both present in one training sequence |
| Context noise | Training-only corruption of context copies, so later blocks tolerate imperfect earlier generations |
| Inverse dynamics model (IDM) | A map from (now, future) to the actions that cause the change; ModAR's final step |
| Unified / Independent-noise / Disjoint / Action-only | The paper's baseline formulations (joint; joint with per-stream training noise; no future-action attention; no futures) |
| D | Total demonstrations per task in simulation (50 labeled + D − 50 actionless) |
The paper situates itself in a fast-moving 2026 literature, and its reference list is a good map of the open directions.
Richer structured futures. WAM4D uses spatial register tokens for fast 4D prediction; 3PoinTr lifts tracks to 3D; ST-WAM targets semantic-temporal features for robustness to visual distribution shift. Each is a different bet on which structure a policy should imagine.
Human video at scale. EgoWAM trains beyond pixels on in-the-wild egocentric human data; EgoDex supplies large-scale egocentric manipulation video. ModAR's small EgoDex gain is one data point in a question these works are actively probing.
How much imagination is needed at run time. Fast-WAM asks whether test-time future imagination is needed at all. ModAR's results suggest that, in its setting, generating the right futures in the right order improves success, at a latency cost the paper lists as a limitation.
Pretrained generators. DreamZero, Cosmos Policy, VERA and Flex-π all build on video models. The natural synthesis, a pretrained model that also generates structured modalities in sequence, is the experiment this paper leaves for someone to run.
Try these three without looking back; each takes one or two sentences.
1. A colleague proposes generating RGB first "because it contains everything." Which experiment in the paper speaks to this, and what did it find? Check: the reverse-order ablation (RGB → depth → DINO → tracks) dropped success from 75% to 65%.
2. Why can ModAR learn from a human video but Action-only cannot? Check: ModAR supervises its future blocks on actionless data and drops only the action term; Action-only has no future objective, so an actionless example gives it nothing to learn.
3. What single number best summarizes why context noise matters? Check: without it, success at D = 250 falls from 75% to 63%, below Unified's 67%.
Without scrolling up: (1) draw the five formulations and say what differs between Unified and Independent-noise; (2) write ct, Ytm and At with the paper's H, Δ and J, and count the future tokens per modality; (3) write Equation (1) and explain why the last factor is an inverse dynamics model; (4) derive v = (Ŷ − Ỹ)/(1 − τ) and explain the δ clamp in Equation (3); (5) explain what context noise protects against and quote its ablation number; (6) name the two controls that defend Table I, and the RGB result that shaped the real-robot model. If any of the six stalls, its chapter is one tap away.