A robot policy that learned by imagining the future three ways: as colour video, as 3D geometry, and as what each object is. Switch the imagining off and it still beats every tested baseline, deciding in 60 milliseconds.
Pick a policy and watch it rack plates. Then we build, piece by piece, why imagining helps even when it is switched off.
You need what a neural network is and the idea that a model can turn noise into a sample. We build the rest from zero.
Each policy racks the same plate. Watch the pauses between moves: that is the policy thinking.
Task completion on Put Plate on Rack from Figure 6; per-call latency on one RTX 5090 from Figure 15. The scene is a toy: one gripper on an overhead gantry stands in for the robot's arms. Each plate's stages succeed at random at the paper's rate, so a short run wobbles around the paper's number and a long run converges on it.
Chapter 0
Why a robot that predicts the future acts better, and what that has cost until now
A two-armed robot sits at a kitchen table. On the table: a plate, a dish rack, a camera looking down and one on each wrist. You say "put the plate on the rack." Every 1.07 seconds the robot gets one snapshot from its cameras, its own joint readings, and your sentence, and it has to decide its next 32 motor commands.
The program that makes that decision is a policy: a function from what the robot senses to what it does. The most common kind today is a vision-language-action model (VLA): a large network that reads the image and the instruction and writes out actions. It learns from demonstrations, recordings of a person steering the robot through the task, and the only thing it is graded on is whether its actions match the person's.
That is a thin signal. A demonstration of racking a plate is a few thousand images, but the action labels are only 32 numbers per step. The network is told what to do and never why. It must discover on its own that the plate is a rigid disc, that the rack has slots, that the slot is 3 centimetres farther away than it looks.
A world-action model (WAM) adds a second job. Besides the actions, it must predict what the cameras will see next. That prediction is graded pixel by pixel, which is a far richer signal: to predict the next frame, the network has to model how the plate moves when the gripper pushes it. It also means the network can start from a video-generation model, one pre-trained on huge amounts of internet video, which already knows a great deal about how things move. WAMs are known for two things: they need fewer demonstrations, and they generalise better to scenes they have not seen.
Here is the catch. Nearly every generalist WAM predicts the future as RGB latents: compressed codes of colour images, produced by the video model's own encoder. That encoder was trained for one thing, rebuilding pixels. So its codes spend much of their capacity on what the scene looks like: the lamp, the shadows, the pattern on the tablecloth. A manipulation policy needs something else. It needs geometry (where things are in 3D) and object semantics (which blob is the plate and which is the rack). Try it: change the task-relevant thing and the irrelevant thing, and watch which kind of picture moves.
One tabletop, seen three ways. Move the cup (that matters to the task), dim the lamp or swap the tablecloth (that does not). The meters show how much each picture changed from where it started.
A toy renderer. The meters are the real mean pixel change between the current picture and the starting one. The DINO view is idealised as ignoring light and texture entirely; real DINO features are robust to appearance changes, not perfectly blind to them.
The RGB view jumps for every change, relevant or not. The 3D view and the object view move only when the cup moves. A network trained to predict the RGB future is being paid, in large part, to predict lamps and tablecloths.
So why not also ask the WAM to predict the 3D future and the object future? Because, until this paper, that came with three costs:
Flex-π (Yan, Liu, Fan and colleagues at the University of Washington and the Allen Institute for AI, 2026) is a 6-billion-parameter WAM that pays none of the three. The 3D and object streams are computed from the same RGB image by frozen off-the-shelf models, so no new sensor is needed. The 3D stream is encoded by the same frozen video encoder as the colour stream, which turns out to work almost perfectly, so no new prior is needed. And training randomly hides whole streams, so at deployment you may switch any of them off, so no latency is forced on you. The hero above shows the result: the full model imagining everything decides in 193 milliseconds; the same weights imagining nothing decide in 60, faster than π0.5, and still rack more plates than any baseline.
Here is the road, in order. Each chapter explains one part of the instrument in the hero.
Chapter 1
Flow matching, the machine that turns noise into actions, and why one step lands in the middle
Before any of the streams, we need the engine that produces them. Flex-π does not compute its action with one pass of a network the way a classifier computes a label. It generates the action, the way an image generator produces a picture: it starts from random noise and repeatedly nudges that noise toward something that looks like a real action.
Why go to that trouble? Because good actions are often not unique. There is a mug between the gripper and the plate. Going around it on the left is fine. Going around it on the right is fine. A network trained to output a single best guess learns the average of the demonstrations, which here is straight through the mug. A generative policy instead learns the whole spread of good actions and samples one of them. This is the idea behind diffusion policies, and Flex-π uses its modern form, flow matching.
Flow matching draws a straight line from a noise sample to a data sample and teaches a network to point along it. Write the clean thing we want (an action chunk, or a future image code) as z, and a sample of pure Gaussian noise as ε. For a flow time τ between 0 and 1, the point on the line is:
Along that line the velocity is constant: every point moves in the direction z − ε at the same speed. So the network's job is simple to state. Given the noisy point, the flow time, and everything the robot knows, predict that direction. The paper's Equation 1 is the squared error of that guess:
Let's check it with numbers in one dimension. Say the clean value is z = 0.8 and the noise draw is ε = −1.2. At τ = 0.25 the noisy point is 0.25 × 0.8 + 0.75 × (−1.2) = 0.2 − 0.9 = −0.7. The target direction is 0.8 − (−1.2) = 2.0. If the network gets that exactly right, one step covering the remaining flow time of 0.75 lands at −0.7 + 0.75 × 2.0 = 0.8: the clean value.
−0.7 + 0.75 × 2.0 = 0.8 (noisy point + remaining time × velocity)
At deployment there is no clean value to aim at. The policy draws noise, asks the network for a direction, takes a small step, asks again, and so on. Each step is one Euler step, and the number of steps is K. Flex-π uses K = 4. The question is why it cannot use K = 1, which would be four times cheaper. The device answers it.
Each dot is one sampled gripper target. The demonstrations went around the mug on the left or on the right, never through it. Choose how many Euler steps the policy takes, then press Flow and watch the noise travel.
The velocity field here is exact, not learned: for data made of two Gaussian clusters, the best possible vθ has a closed form (the average of z − ε over every data point that could have produced the current noisy point). So everything you see is what a perfectly trained network would do.
With K = 1 every sample lands between the two clusters, on the mug. Here is why. At τ = 0 the network sees pure noise, which says nothing about which side the sample should end up on. The best it can do is predict the average direction over all the data, and one full-length step along the average direction lands on the average of the data. That average is exactly the collision we introduced generative policies to avoid.
With two or more steps, the first step only moves partway. By the second question the point is already a little closer to one cluster, so the network can tell which side it is on and the remaining steps commit to it. The paper's own sweep shows the same cliff. On RoboTwin, in action-only mode, success is 51.0% at K = 1, 93.7% at K = 2, 94.5% at K = 4, and 93.5% at K = 10 (Table 11). Past two steps, more steps buy almost nothing and cost latency.
One engine, many streams. Everything Flex-π generates goes through this same machine: the chunk of 32 actions, and, when they are switched on, the future RGB codes, the future 3D codes, and the future object features. They are all denoised together, in one sequence, sharing the same K steps. The paper puts it in one sentence: sampling from the policy means integrating the flow over whichever output streams are active. The rest of this lesson is about which streams those are, how they are encoded, and who gets to look at whom while they are denoised.
Chapter 2
A video encoder trained only on pictures turns out to encode 3D geometry almost losslessly
Flex-π wants to predict the future in 3D. The first question is what "3D" means as data, and the answer is pleasingly concrete: a pointmap. A pointmap is an image in which each pixel, instead of a colour, stores the 3D position of whatever surface that pixel sees: three numbers, X, Y and Z, in metres, in the camera's frame. It has exactly the shape of a colour image, height by width by 3. Only the meaning of the three channels has changed.
You build one from a depth map (a distance Z for every pixel) and the camera's intrinsics, the focal lengths fx, fy and the image centre cx, cy. A pixel at column u and row v with depth Z sits at:
With numbers: a 384 × 320 camera with fx = fy = 400 and centre (192, 160) sees a surface at pixel (300, 180) at depth 0.8 m. Then X = (300 − 192) × 0.8 / 400 = 0.216 m and Y = (180 − 160) × 0.8 / 400 = 0.04 m. That pixel of the pointmap stores (0.216, 0.040, 0.800). Only the intrinsics are needed; no calibration between cameras.
Where does the depth come from? Here the paper refuses the first cost from Chapter 0. The pre-training data, about 500 hours from AgiBot World, has RGB from a head camera and two wrist cameras, and sensed depth only on the head camera, full of holes. So the authors ignore the sensor and run Depth Anything 3, a model that estimates metric depth from a single RGB image, on all three views, offline. In simulation (LIBERO) they render true depth from the simulator. On their real robot the ZED stereo cameras supply depth. Pointmaps are clipped at 2 metres, well beyond a tabletop.
Now the encoding. Every stream the transformer sees must first be squashed into compact codes, because raw pixels are far too many. For colour video Flex-π uses the VAE (variational auto-encoder) of the Wan 2.2 5-billion-parameter video generator: an encoder that compresses video into a small grid of numbers, and a decoder that expands it back. The encoder turns each 16 × 16 block of pixels, across 4 frames, into 48 numbers; the transformer then packs a 2 × 2 patch of those into one token of 192 numbers. That VAE is frozen: its weights never change, in pre-training or fine-tuning.
The paper's surprise, and its free lunch, is what happens when you feed that same frozen VAE a pointmap instead of a photo. It was trained only on RGB video. It has never seen an image whose channels mean X, Y, Z. And yet the reconstruction comes back nearly perfect: PSNR 38, mean squared error 0.0001 (Figure 3). The pointmap can live in the same latent space as the colour video, encoded by the same weights, with no pointmap-specific training at all.
Why might that work? The paper does not say; it reports it as surprising. Here is one intuition you can test yourself. A pointmap of a tabletop is made of a few smooth pieces: the table is a plane, so its X, Y and Z change linearly across the image; a box is a few more planes. A photo of the same table has wood grain, stripes, shadows and highlights. An encoder built to keep the important structure of natural images will certainly keep smooth ramps and sharp object edges.
The same scene as a colour photo and as a pointmap (X, Y, Z shown as red, green, blue). Both are crushed by the same rough compressor: average every block, then stretch it back. Grow the blocks and compare what survives.
The compressor here is block averaging plus bilinear upsampling, a crude stand-in for the VAE, and the scene is synthetic; the PSNR values are computed live on these pixels. The paper's measurement is the real frozen Wan VAE on a real pointmap: PSNR 38, MSE 0.0001.
At every block size the pointmap keeps several times more of its signal than the photo does, measured in decibels of PSNR (peak signal-to-noise ratio: higher means a closer reconstruction; every 10 dB is ten times less squared error). Smooth, piecewise-planar images are simply easy to compress.
The payoff is larger than convenience. Because pointmap tokens are made by the video model's own encoder, they share its latent space, and the transformer that was pre-trained to predict future video codes already knows how codes in that space evolve. The 3D stream inherits the video prior. Its adapter in and out of the transformer is a private copy of the video stream's patch layers, initialised from Wan 2.2, which then specialises during training.
Chapter 3
Every token the model sees and makes, counted
We now have two visual streams encoded by the same VAE. The third brings object meaning. DINO models are image encoders trained without labels to produce features in which the same object looks the same from different angles and in different light. Flex-π uses a frozen DINOv3 (the ViT-B/16 size): it reads a colour image and returns one 768-number feature per 16 × 16 patch. Like the pointmap, it is computed from the RGB image, so again no new sensor.
Here is the full list of what the model reads at time t (Equation 2):
| Input | Made by | Carries |
|---|---|---|
| RGB tokens zot | frozen Wan VAE on the image ot | appearance; inherits the video prior |
| Pointmap tokens zpt | the same frozen VAE on the pointmap pt | 3D geometry, in the video latent space |
| DINO tokens dt | frozen DINOv3 on ot | object semantics |
| Proprioception st | the robot's joint sensors | where the arms are (Chapter 8) |
| Language l | frozen umT5 text encoder | the task |
The instruction and the joint state are not streams to be generated; they are global conditioning. The instruction becomes 128 text tokens, the joint state becomes one more token appended to them, and every other token reads these 129 through cross-attention: a layer in which tokens look up information in a separate set of tokens.
Now the counting, which is where the architecture becomes real. The robot has three cameras, and rather than encode them separately, Flex-π tiles them into one composite canvas of 384 × 320 pixels: the overhead view across the top, the two wrist views side by side below. One VAE pass covers all three. With 32 × 32 pixels per token (16 from the VAE, 2 from the patch), that canvas becomes a 12 × 10 grid: 120 tokens per frame.
Time adds a dimension. Each training sample spans 33 timesteps at 30 Hz. The visual streams are sampled every 4th step, which gives 9 frames, and the VAE's 4× compression in time keeps the first frame alone and folds the other 8 into two: 3 latent frames, the present and two futures. So the RGB stream is 3 × 120 = 360 tokens, and so is the pointmap stream.
DINO needs care. DINOv3 at 224 pixels gives a 14 × 14 grid of 768-number patches per view: 196 per view, 588 for three views, 1,764 for three timestamps. That is more tokens than everything else combined. Flex-π applies a fold (the same rearrangement as PyTorch's PixelUnshuffle): each 2 × 2 block of neighbouring patches is stacked into one token of 4 × 768 = 3,072 numbers. The grid becomes 7 × 7 per view, 147 per timestamp, 441 in all. Nothing is thrown away; the fold is exactly invertible. The model simply predicts each 2 × 2 neighbourhood as one token instead of four.
360 RGB + 441 DINO + 360 3D + 32 actions = 1,193 tokens (387 present-frame anchors + 806 to denoise)
The present frame is known, so its tokens are anchors: 120 + 147 + 120 = 387 of them, clean, never noised. The rest are what the model generates: 240 + 294 + 240 future visual tokens and 32 action tokens, 806 in all, denoised together over the K flow steps of Chapter 1. Switch the futures on and off below.
Top: the composite canvas and its 12 × 10 token grid. Below: the whole sequence, drawn to scale, anchors first. Choose which futures are generated and whether DINO is folded.
Token counts from the paper's token accounting (Appendix I) and Sections A.2 and A.3. "Pairs" counts how many query-key pairs one denoising step's attention scores, with the anchors cached; it is a size, not a latency.
The accountant shows why switching futures off is such a strong lever. With every future on, each denoising step scores roughly a million query-key pairs. With only the action on, the 387 anchors can be computed once and cached, and each step handles just the 32 action tokens against 419 others: about 13 thousand pairs, some 70 times fewer. Chapter 9 turns those sizes into milliseconds, and finds that the milliseconds do not scale the way the pairs do.
Chapter 4
A Mixture of Transformers: shared attention, separate weights
We have 1,193 tokens of four kinds. The obvious move is to put them all into one big transformer. The obvious move has two problems. The three visual streams are alike: all image-shaped, two of them in the very latent space the video model was pre-trained on. The action stream is not: it is 32 short vectors of joint targets, a different kind of object that deserves its own weights. And a transformer as wide as the video model for 32 action tokens would waste a great deal of computation on the smallest part of the sequence.
Flex-π's answer is a Mixture of Transformers (MoT). In a MoT every modality has its own non-attention weights: its own layer norms, its own feed-forward layers, its own query, key and value projections. What they share is the attention itself. At each block, every stream computes its queries, keys and values with its own weights, the results are concatenated into one long sequence, one attention is run over all of it, and each stream takes its slice of the output back through its own weights.
5B · width 3,072 · 30 blocks
Shared by the RGB, DINO and 3D streams, initialised straight from Wan 2.2 5B, so it arrives knowing how video evolves. Each stream enters and leaves through its own small adapter, and has its own prediction head.
~1B · width 1,024 · 30 blocks
Its own queries, keys, values, feed-forward layers and norms, all narrower: residual width 1,024 against 3,072, feed-forward width 4,096 against 14,336. It reads the visual streams through the shared attention.
For the concatenation to work, the two experts must agree on the shape of attention even though their widths differ. They do: both use 24 heads of 128 dimensions, which is 24 × 128 = 3,072 numbers of queries, keys and values per token, and both have 30 blocks. An action token's 1,024-number state is projected up to 3,072 for attention and back down afterwards. A visual token's 3,072-number state is projected to 3,072 directly. After projection they are the same kind of thing, and one softmax can compare them.
The streams do not mix everywhere. In the first and last blocks, each stream attends only to itself, so early layers can encode each modality in its own terms and late layers can decode it back. Cross-stream attention runs only in the middle 16 of the 30 blocks. That is where RGB, geometry, semantics and action are fused. Walk the stack and see.
The thirty blocks run left to right: the wide visual trunk above, the narrow action expert below. Drag through them. The graph shows who can attend to whom in the block you are on.
The paper states that cross-stream attention is applied in the middle 16 of the 30 blocks; we draw it as blocks 8 to 23, the symmetric reading. Widths are drawn to scale (3,072 against 1,024).
The trunk can copy Wan 2.2 directly. The action expert cannot: its matrices are a third the width, so there is no weight to copy. Starting it from random numbers would throw away what the video model knows about attention patterns and depth. Flex-π instead resamples each Wan tensor to the smaller shape: one axis at a time, it reads the weights as a signal and resamples it by 1D linear interpolation, like shrinking an image.
Resampling has a side effect on the signal. A neuron's output is a sum over its inputs, so its size grows with the number of inputs, the fan-in. Shrink the fan-in from 3,072 to 1,024 and each output sums a third as many terms: its variance drops to a third, and deep in a 30-block network that compounds. The fix is to scale every weight whose fan-in shrank by:
Why the square root: the output's variance is (fan-in) × (weight variance) × (input variance). Cutting fan-in by 3 cuts output variance by 3; multiplying each weight by √3 multiplies weight variance by 3, which cancels it. The device does this for real on 256 neurons.
Top: one row of a 3,072-wide weight matrix, and the same row resampled to 1,024. Bottom: the spread of the layer's outputs for random inputs, before resampling, after it, and after the √3 rescale.
Real arithmetic on random weights, not Wan's. Resampling uses the usual interpolation convention (sample centres aligned), which for an exact 3× shrink lands every new weight on an original one, so only the fan-in changes.
Only two pieces of the action expert start from scratch: the layer that embeds an action into a token, and the head that turns a token back into an action, because nothing in a video model corresponds to them. Everything else, in both experts, begins as Wan 2.2.
Chapter 5
Two masks turn one checkpoint into many policies
This is the heart of Flex-π. The hero offered one set of weights run two ways, imagining everything or imagining nothing. What makes that possible is not a trick at deployment. It is a rule, applied during training, about which tokens may attend to which. Two small masks set that rule for every training example.
The input presence mask min has one bit per visual stream: which streams are observed at time t. A stream with its bit at 0 has its present-frame tokens zeroed, and no one may attend to them. The output attention mask mout also has one bit per visual stream: which futures the action reads. Written as the paper's attention rules (Section A.5), with obs meaning the present-frame tokens of the observed streams:
The anchors (the present) attend only to one another, and the instruction and joint state reach every token through cross-attention whatever the masks say. These rules hold in the middle 16 blocks of Chapter 4, where the streams meet. Before reading on, play with the masks and find the setting where the action reads nothing.
Left: the attention matrix. Each row is a group of query tokens, each column a group of key tokens; a filled cell means the row may attend to the column. Right: the same rules as a graph, like the paper's Figure 4. Set the two masks yourself, or pick one of the paper's four regimes.
Rules from Section A.5 and Figure 4. At least one stream must be observed; the device will not let you hide all three, and neither does training (Chapter 6).
Look at Action only in training. All three futures are still there and still denoised; their mout bits are all 0, so they form one group and see one another, but the action reads none of them. Now switch to At deployment. The three futures vanish from the sequence. Nothing the action reads has changed, because the action never read them and nothing it does read could see them. That is the whole speed-up: 806 tokens to denoise become 32.
Rule 4 is exact in the same way. Because no token attends to the action, the visual futures never depend on it, so the imagined future is the same whether or not the action is being generated alongside it. Try Forcing: the 3D input is hidden, yet the 3D future is still generated and still read by the action. The model must imagine geometry it was not shown. Chapter 6 is about exactly that.
Chapter 6
Stream dropout, cross-modality forcing, and why hiding inputs makes a better policy
Chapter 5 gave us masks. Training needs a way to choose them. Flex-π draws both masks at random, independently, for every training example. Each bit of min is a coin flip: with probability 0.5 the stream is observed. Each bit of mout is another coin flip: with probability 0.5 the action reads that future. The one constraint is that at least one visual stream must be observed; if the coins hide all three, the draw is thrown away and repeated. This is stream dropout.
A hidden input is zeroed, not removed, so every example in a batch keeps the same tensor shapes. And here is the crucial part: every future is still denoised and still pays its loss, observed or not. If the 3D input is hidden, the model must still produce the 3D future, from RGB and DINO alone. The paper calls this cross-modality forcing, and applies it to all three streams.
Let's count what that asks. Of the 8 equally likely patterns for min, one (all hidden) is rejected, leaving 7. A given stream is observed in 4 of them, so it is hidden with probability 3/7 ≈ 43%. On average each training example hides 3 × 3/7 = 9/7 ≈ 1.29 streams and asks the model to imagine their futures anyway. Only 1 example in 7 shows all three.
Each row is one training example. Left: which streams it observes (hollow means hidden). Right: the futures it must produce, always all three; a star marks a future it has to imagine without seeing that stream, and an arrow marks the futures the action reads. Draw a batch, then draw ten thousand.
The sampler is the paper's (Section A.5, Table 12): independent Bernoulli(0.5) bits for both masks, rejection sampling so at least one stream is observed.
This is the same move that made masked language models and masked autoencoders work: hide part of the input and demand it back. A model that can produce future geometry from colour and semantics, and future semantics from colour and geometry, has been pushed to build one internal picture in which the three are mutually predictive. That shared picture is what the action reads.
The paper tests the claim directly. Two models on RoboTwin observe all three streams at test time and differ only in the training rule. With cross-modality forcing the average success is about 66%; without it, trained to predict only the futures it was shown, about 45% (Figure 12). Forcing is not only insurance against a missing sensor. It makes a better policy even when nothing is missing.
It also pays for the insurance. On the real robot, the 3D input is the only one of the three that needs a depth sensor at deployment. Withholding it costs almost nothing: on Put Plate on Rack in full joint generation, task completion is 95.0% with the depth input and 91.7% without (Figure 18). The model supplies the geometry it was not given, and the paper's Figure 13 shows the generated 3D scene still looks right with the pointmap withheld.
The stream still has to exist in training, though. Remove the pointmap stream from training entirely and RoboTwin success drops by 20 points (Figure 11a): adding DINO to video lifts success by 6.8 points, and adding pointmaps on top lifts it by about 20 more. The depth sensor becomes optional; the 3D supervision does not.
Chapter 7
Summing the flow losses, and why the DINO head predicts the answer, not the direction
With the streams, the trunk and the masks in place, the objective is almost an anticlimax. Each stream has its own prediction head, a small layer that reads that stream's tokens off the trunk, and each head is graded by the flow-matching loss of Chapter 1. The total loss adds them up (Equation 3):
Notice what is not in the equation: mout. The output mask shapes attention, never the loss. A future the action ignores in this example is still graded. That is what keeps cross-modality forcing well-posed: every head always has a target.
One head is different, and the reason is a nice piece of recent theory. Chapter 3 folded each DINO token to 3,072 numbers, which is exactly the trunk's width. At that ratio, Li and He (2025, "Back to Basics: Let Denoising Generative Models Denoise") found that asking the network for the velocity works worse than asking it for the clean target itself, an approach called x-prediction. Flex-π's DINO head therefore predicts the clean future features d̂t+1, and converts them to a velocity afterwards.
Why would it matter which one the network outputs? Look at the two targets. The clean features are structured: real images produce DINO features that sit on a thin, low-dimensional surface inside the 3,072-dimensional space. The velocity is d − ε: it contains the whole noise vector ε, a random point spread across all 3,072 directions. A network whose width equals the token size has no spare room to carry that noise through to its output. The other streams are far from this regime: RGB and pointmap tokens are 192 numbers against a width of 3,072, a ratio of about 0.06; action tokens are 32 numbers against 1,024.
Top: each stream's token size against its expert's width. Bottom left: sampled targets for a toy 2-D stream whose data lies on a curve; switch between the clean targets and the velocity targets. Bottom right: one x-prediction, converted to a velocity; drag the flow time.
Ratios from Section A.4 (the paper gives action tokens as 0.01 to 0.03 of the action width). The curve, the noise and the 2-D picture are a toy.
The conversion is one line of algebra, and it is worth doing by hand. The noisy point is dτ = τd + (1 − τ)ε. Subtract it from d: d − dτ = (1 − τ)d − (1 − τ)ε = (1 − τ)(d − ε). The velocity is d − ε, so dividing by (1 − τ) recovers it. With the network's estimate in place of the true d (Equation 4):
In numbers: clean value 0.6, noise −0.4, τ = 0.75, so the noisy point is 0.45 − 0.1 = 0.35. The head guesses 0.58. Then v̂ = (0.58 − 0.35) / 0.25 = 0.92, against the true velocity 0.6 − (−0.4) = 1.0. Equations 1 and 3 are untouched; only the way the head's output is read changes.
Two more settings finish the recipe (Table 12). The visual streams use a flow-time shift of 6.0, a standard video-model setting that spends more of training on the noisier part of the path, while actions use no shift. And the frozen parts stay frozen throughout, in pre-training and fine-tuning: the Wan VAE, the umT5 text encoder, and DINOv3. Everything else trains: the trunk, the action expert, the adapters, and the four heads.
Chapter 8
How the robot's state goes in, how its actions come out, and why every move is measured from one anchor
So far the action has been "32 action tokens". Now we open one. Flex-π is pre-trained on one robot and fine-tuned on others: a simulated bimanual robot in RoboTwin, a single simulated arm in LIBERO, and a real two-armed YAM. They have different numbers of joints and grippers. For one network to serve them all, their states must look alike. Flex-π's answer is a canonical layout: every robot's state is a 32-number vector with fixed meaning per slot.
Each arm owns 16 slots: where its hand is (3 numbers), which way it points (6 numbers, a 6D rotation, two columns of the rotation matrix, which avoids the jumps of angles), its gripper (1), and its six joint angles (6). The layout is grouped by field rather than by arm, so both hands' positions come first. A robot with fewer channels scatters its numbers into the matching slots, fills the rest with zeros, and carries a padding mask. The YAM fills all 32; RoboTwin maps its 14 numbers in; LIBERO uses 8.
The state reaches the model in the cheapest possible way. One linear layer maps the 32 numbers to 4,096, the width of the umT5 text tokens, and the result is appended to the instruction as its 129th token. Every stream reads it by cross-attention exactly as it reads language. Only the present state is used: no history. Non-rotation channels are scaled to the range −1 to 1 using the 1st and 99th percentiles of each robot's data; rotations are left alone and cleaned up (orthonormalised) when decoded.
Actions reuse the same 32 slots, so a slot means the same quantity whether it is being read as state or written as action. What differs is the reference frame. On the real robot every step of a chunk is expressed relative to one anchor, the state at the chunk's first step: hand targets become poses in the anchor's frame, joint angles become displacements from the anchor, and only the grippers stay absolute. The alternative would be to express each step relative to the step before it. The device shows why the paper does not.
The grey curve is the 32-step hand path in a demonstration, seen from above. Both predictions make the same small error on every step. Chained predictions add each step to the one before; anchored predictions measure each step from the start. Raise the error and resample.
A toy with independent Gaussian errors per step. The anchored and relative forms are the paper's (Section A.2); the error sizes are illustrative.
Chained errors add up like a random walk: after n steps the typical drift is the per-step error times √n. At step 32 that is √32 ≈ 5.7 times the per-step error, and it points in a random direction. Anchored errors do not accumulate at all. The paper names both benefits: anchoring keeps targets from "accumulating integration error along the chunk", and it makes an action independent of where in the workspace the motion started. For a task like Self-Repair Gripper, with insertions at ±0.25 to 0.5 mm of clearance, the difference between those two curves is the difference between a screw that goes in and one that does not.
0.3 mm × √32 ≈ 1.7 mm of drift (chained) against 0.3 mm (anchored)
Two more details close the loop. Each of the 32 steps becomes one token through a single linear layer, and a single linear layer turns each denoised token back into 32 numbers; there is no discretisation into bins. And the chunk timing: a training sample spans 33 timesteps, the chunk is H = 32 actions, and at 30 Hz that is 1.07 seconds of motion. On the real robot all 32 are executed without looking again (open loop), and the arms hold still while the next chunk is computed. In LIBERO the policy re-plans after 10 of the 32. How often to re-plan is a deployment choice, not part of the representation.
Chapter 9
Where the 60 and the 193 milliseconds go, and the one knob that shortens imagination
Every number in the hero's latency column is one policy call producing one chunk, measured on a single RTX 5090 at batch size 1, averaged over 20 calls after 3 warm-ups. Every training-free trick the authors used is in Appendix I, row by row, and reading it teaches a lot about where time really goes in a model like this.
Start with the shape of the cost. A call does some work once (encode the three cameras, run DINO, set up the anchors, host-side bookkeeping) and some work K times (one pass of the 30 blocks per Euler step). So the latency of a call is, to a good approximation, a straight line in K:
The paper fits that line to each configuration. With every stream generated, on the TensorRT engine, the fit is 42.9 ms fixed and 52.8 ms per step: at K = 4, 42.9 + 4 × 52.8 = 254 ms, and the measurement is 252. In action-only mode, compiled, it is 44.6 fixed and 4.0 per step: 44.6 + 16 = 60.6, measured 60.3. Look at the per-step columns. Generating 806 tokens costs 52.8 ms a step; generating 32 costs 4.0. Imagination is almost entirely per-step cost, which is why cutting K is the one lever that shortens the joint path without giving up a stream.
Pick a mode and a step count, then climb the ladder of optimisations, each row adding one component to the row above. Top: the measured latency. Bottom: what the robot does with it, 1.07 s of motion and then a pause while it thinks.
Every latency is measured, from Tables 9 and 10 (RTX 5090, batch 1, three cameras, 384 × 320). Success at each K is from Table 11 (RoboTwin, action only). The robot timeline assumes the real-robot deployment: 32 steps at 30 Hz, arms holding still during the call.
Climb the action-only ladder and one row towers over the rest: loop-scope compile, which captures the whole denoising loop, including the anchor prefill, as one compiled graph. It cuts the per-step cost from 14.2 to 4.1 ms, 3.5×. Better kernels could not have done that, because a step that processes 32 tokens against a cached prefix is not limited by arithmetic. It is limited by the overhead of launching many tiny operations on the GPU, and compiling the loop removes the launches.
The joint path is the opposite. There, the big lever is exporting the 30-block denoising core to a TensorRT engine, which cuts per-step cost by 48%. Then comes the prefill/decode split, borrowed from how language models serve text. The 387 anchor tokens attend only to one another and never change across denoising steps (the paper checked: their outputs are bit-identical from step to step). So their keys and values are computed once per call, and every step processes only the 806 noisy tokens against that cache. Per-step cost falls another 27%, from 52.5 to 38.5 ms, at the price of a one-time prefill of about 20 ms. That trade pays from K = 2 upward.
What imagination costs, in the end, is easiest to see at the two ends of K. At K = 1 the two modes nearly meet: 49.0 against 73.1 ms. At K = 10 they are 5× apart: 85.2 against 428.1 ms. At the paper's K = 4 they are 60 and 193. And the failures are as instructive as the wins: 8-bit floating point inside TensorRT was a 1.4% wash, raising TensorRT's optimisation level from 3 to 5 changed per-step cost by 0.5%, and a second-order multistep solver, mathematically sound, collapsed in rollouts to 20.5% success against 94.7% for plain Euler at six steps. These checkpoints were tuned for Euler.
Chapter 10
Real robots, shifted scenes, half the data, and two benchmarks
The paper asks four questions: does Flex-π work on precise, contact-rich, long real tasks; does it hold up under new objects and less data; how does it compare with many VLAs and WAMs across the speed range; and how much do the extra streams matter. We have answered the last one along the way. Here are the other three.
The real robot is a stationary two-armed YAM with an overhead ZED 2i and a ZED Mini on each wrist. Five tasks, chosen to cover dexterity, precision, long horizons and sustained contact: Put Plate on Rack, Sort Utensils, Kitchen Organization (four skills in one episode, including a hand-over between arms), Self-Repair Gripper (the robot screws its own replacement gripper back on over eight stages, with insertions at ±0.25 to 0.5 mm of clearance), and Soft-Bag Zipping (open a fabric pencil case, put pens in, zip it shut). Each baseline was trained by the authors on the identical demonstrations, from 152 episodes (1.2 hours) for Sort Utensils to 802 plus 570 corrections (17.4 hours) for Self-Repair Gripper.
Scoring is partial credit: each task has a rubric of stages, and task completion is the points earned as a fraction of the maximum. Put Plate on Rack, the hero's task, gives 0.5 for grasping the plate from a tilted edge, 0.25 for getting it into the rack in any condition, and 0.25 more for seating it cleanly in an empty slot. Binary success, the fraction of rollouts that clear every stage, is reported too.
Choose an experiment. Bars are task completion (or success, for RoboTwin); switch policies on and off to compare.
Real robot: Figure 6 (task completion; Fast-WAM was not run on the two hardest tasks). Unseen scenes and half the data: Figure 9, bar heights read from the figure, changes as printed. Fewer demos: Figure 10, RoboTwin, 50-task average success, domain-randomised.
The margin grows with difficulty. On Kitchen Organization, full joint generation leads the strongest baseline by 5.0 points; on Soft-Bag Zipping by 27.2; on Self-Repair Gripper by 42.7. Averaged over the five tasks, task completion is 52% for π0.5, 58% for ManiFlow, 76% for Flex-π acting only and 83% imagining all; binary success is 18%, 27%, 50% and 63%, which is where "3.5× π0.5 and 2.3× ManiFlow" comes from. The action-only variant, the cheapest policy tested, beats every baseline on every task.
Under shift, with unseen objects and heavy clutter, Flex-π gives up 4.7 points on average imagining all and 4.1 acting only. ManiFlow, the strongest baseline in distribution, gives up 26.7 even though it sees depth; π0.5 loses 25.6 on the unseen soft bag. That ManiFlow sees 3D and still falls matters: it suggests the robustness comes from the representation learned by joint prediction, not from having geometry as an input. With half the real demonstrations on Put Plate on Rack, full joint Flex-π still beats every baseline trained on all of them, and action-only on half matches π0.5 on all.
In simulation the picture repeats. On RoboTwin's 50 bimanual tasks, Flex-π reaches 94.6% in both modes, ahead of the strongest VLA (Qwen-RobotManip-Context, 93.9%), which was pre-trained on about 38,100 hours, 76 times more robot data than Flex-π's 500. The two modes being equal there suggests the benchmark is close to saturated. With only 50 demonstrations per task the gaps open: 78.8% against 41.9% for the best baseline. On LIBERO, which has no split between training and test scenes and so measures fitting, one checkpoint reaches 98.4% to 98.5%, and 99.2% without stream dropout, tying the best reported. On LIBERO-Plus, 10,030 perturbed tasks, Flex-π scores 88.6% in both modes against 85.7% for π0.5.
Chapter 11
The cheat sheet, the limits, and where Flex-π sits in the field
You can now read the Flex-π paper and explain every design decision: why pointmaps, why the same VAE, why DINO is folded and predicted as x, why two experts, why the output mask groups futures, why every future is always graded, why actions are anchored, and where each millisecond goes. Let's lock it in.
Flex-π is a 6B world-action model that predicts future RGB, 3D pointmaps and DINO features alongside a chunk of 32 actions. The 3D and DINO streams are computed from the RGB image by frozen Depth Anything 3 and DINOv3, and the pointmap is encoded by the same frozen Wan 2.2 VAE as the video, which reconstructs it at PSNR 38. All streams are denoised jointly by flow matching in a Mixture of Transformers: a 5B visual trunk from Wan 2.2 and a 1B action expert, fused in the middle 16 of 30 blocks. Two random masks per example hide inputs and choose which futures the action reads; every future is always graded, so the model learns to imagine what it cannot see, and the grouping rule makes unread futures exactly removable. One checkpoint then runs from action-only at 60 ms to full imagination at 193 ms, beating π0.5, ManiFlow and Fast-WAM on real bimanual tasks in both modes.
| Quantity | Value | Why it matters |
|---|---|---|
| Parameters | 5B visual trunk + ~1B action expert | Trunk copied from Wan 2.2; expert resampled from it |
| Pointmap through the frozen VAE | PSNR 38, MSE 0.0001 | The free lunch: no new encoder |
| Tokens, full joint | 1,193 = 387 anchors + 806 noisy | 360 RGB, 441 DINO (folded), 360 3D, 32 action |
| DINO fold | 2 × 2: 768 → 3,072 numbers, 4× fewer tokens | Lossless; why DINO uses x-prediction |
| Cross-stream blocks | middle 16 of 30 | Encode and decode per stream, fuse in the trunk |
| Stream dropout | 0.5 per bit, both masks; ≥ 1 input observed | Each stream hidden 3/7 of the time |
| Chunk | H = 32 at 30 Hz = 1.07 s, anchored | Open loop on the real robot |
| Euler steps | K = 4 (K = 1 collapses to 51%) | Past 2 steps, more buys little |
| Latency per call | 60 ms action only, 193 ms full joint | π0.5 66, Fast-WAM 86, ManiFlow 103 |
| Real-world average | 76% (acting only), 83% (imagining all) | π0.5 52%, ManiFlow 58% |
| Pre-training data | ~500 h, 100 AgiBot World tasks | 76× less than Qwen-RobotManip |
python# Flex-π in pseudo-Python. Frozen: vae, dino, umt5. Trained: trunk, expert, adapters, heads. def encode(obs): z_o = vae.enc(obs.rgb_composite) # 3 latent frames x 120 tokens x 192 z_p = vae.enc(obs.pointmap_composite) # same frozen VAE, same shape d = fold2x2(dino(obs.rgb_views)) # 3 x 147 tokens x 3072, lossless ctx = concat(umt5(obs.instruction), linear32(obs.state)) # 128 + 1 tokens return z_o, z_p, d, ctx def train_step(batch): streams = encode(batch.obs) m_in = bernoulli(0.5, size=3, reject_all_zero=True) # which inputs are observed m_out = bernoulli(0.5, size=3) # which futures the action reads tau = sample_flow_time(shift=6.0) # shift 1.0 for the action stream eps = noise_like(targets(batch)) noisy = tau * targets(batch) + (1 - tau) * eps # all 3 futures + 32 actions out = mot(anchors=zero_hidden(present(streams), m_in), noisy=noisy, ctx=streams.ctx, mask=attention_rules(m_in, m_out)) # Section A.5, middle 16 blocks v_dino = (out.dino_x - noisy.dino) / (1 - tau) # x-prediction, Eq. 4 return sum(mse(v, targets(batch)[k] - eps[k]) # Eq. 3, every stream, every sample for k, v in out.velocities(dino=v_dino).items()) def act(obs, read=('rgb', 'dino', 'pts'), K=4): streams = encode(obs) # inputs present at deployment (any non-empty subset) cache = mot.prefill(present(streams)) # 387 anchors, once per call x = noise(futures=read, actions=32) # unread futures are not in the sequence at all for k in range(K): # K Euler steps over the active streams x = x + (1 / K) * mot.decode(x, cache, ctx=streams.ctx, tau=k / K) return anchor_to_absolute(x.actions, obs.state) # 32 moves, 1.07 s at 30 Hz
Now press Present or Teach and explain, out loud and from memory, why the action-only policy is better than a policy trained without the extra streams. If you can, you own this paper. Then go back to the rack and run twenty plates with each policy.