Dong, Hung, Sadigh & Finn · Stanford · 2026

Plan with the past.
Correct with the present.

Real-Time EXPO-FT: let a large vision-language-action model propose, then use a fresh observation to edit and select the next action chunk.

Start with real-time chunking, π0.5 and basic actor-critic learning.
12
Chapters
12
Interactive mechanisms
116/120
Reported real-task successes

Four tested tasks, 30 trials each. Compare against 100/120 for EXPO-FT with RTC. Numerical toys below are labeled separately from paper results.

Chapter 0: The world kept moving

A blue block travels around a rotating plate. The robot sees it, begins computing a grasp, and finishes after the block has moved. The proposal can be sensible for the old picture and wrong for the present. More accurate imitation of that old picture does not remove its age.

Two clocks, one robot

The experiments send Cartesian and gripper velocity commands at 30 Hz: one command slot lasts about 33.3 milliseconds. Their base VLA takes roughly 67 milliseconds to infer on the authors’ server. The world therefore advances by about two command periods during this computation. A model update rate, a command rate and a camera rate are different measurements; never label all three simply “frequency”.

The proposed system executes eight commands before selecting another chunk. At 30 Hz that is about 266.7 milliseconds, or 3.75 chunk selections per second under ideal scheduling. The lightweight editor acts at that boundary. The paper does not demonstrate a fresh editor call at every command slot.

Why a running queue helps

A synchronous controller can stop issuing useful new commands while waiting for its next prediction. An asynchronous controller instead keeps consuming an existing queue while the model prepares a continuation. That removes a computation gap when the continuation is ready on time.

However, overlap alone does not make the model’s input current. The prepared chunk is still based on an earlier observation. This is the opening for Real-Time EXPO-FT: retain the expensive proposal, but let a fast decision use another observation just before the selected chunk is executed.

Explore: Latency

Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.

QuantityComputed value

Work it out by hand

  1. One command period: 1 second divided by 30 gives 0.03333 seconds.
  2. Nominal 67 ms inference spans 0.067 × 30 = 2.01 command periods.
  3. Dynamic Picking uses a scheduled lead of d=3 steps: 3/30 seconds = 100 ms.
  4. Eight executed commands span 8/30 seconds = 266.67 ms.
  5. These are arithmetic consequences of the stated configuration, not measured end-to-end control jitter.
Element Value or shape Interpretation
Command rate 30 Hz Commands sent to the robot
Base inference Approximately 67 ms Measured on the authors’ server
Picking lead 3 command steps 100 ms of scheduled overlap
Replan window 8 commands About 266.7 ms at 30 Hz
Extra latency 100 ms on three tasks Not added to Dynamic Picking
command period = 1 / f; chunk period = C / f; scheduled lead = d / f

What the demonstration measures

The rotating block below is an explicit kinematic toy. Its angle changes at the speed you select; the distance between the old and new positions is computed from that circle. It is not a robot rollout, a learned policy, or a reproduction of the paper’s success rates.

The slider deliberately lets you separate elapsed time from object speed. Double speed at fixed latency and the angular mismatch doubles. Set speed to zero and the stale observation can remain geometrically useful. That does not prove a stationary real task is easy; it isolates one failure mechanism.

Latency is not a universal explanation

Sensor delay, occlusion, contact uncertainty and a poor task prior can all hurt a controller. This paper isolates a particularly important issue for dynamic tasks: an action is generated before the state in which it will be used. Its experimental evidence does not establish that every VLA failure is caused by latency.

Use three timestamps in debugging: when the observation was captured, when inference completed, and when the command was applied. An inference timer alone cannot tell you the age of the information reaching the robot.

Run the core calculation

This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.


from math import cos, sin, hypot

frequency = 30.0
latency_ms = 67.0
angular_speed = 1.2  # radians/second; teaching choice
radius = 1.0        # normalized teaching coordinates

def position(seconds):
    angle = angular_speed * seconds
    return radius * cos(angle), radius * sin(angle)

old_position = position(0.0)
now_position = position(latency_ms / 1000.0)
x0, y0 = old_position
x1, y1 = now_position
geometric_error = hypot(x1 - x0, y1 - y0)

command_period = 1.0 / frequency
replan_period = 8.0 / frequency
scheduled_lead = 3.0 / frequency

print(old_position, now_position)
print(geometric_error)
print(command_period, replan_period, scheduled_lead)
assert replan_period > scheduled_lead
# This script integrates no robot dynamics.
# It measures geometric observation mismatch only.

Symbols and data you should be able to name

Name Meaning Concrete referent
f Command frequency 30 Hz
C Executed commands before replanning 8
d Scheduled inference lead in command steps 3 or 5
Observation age Time since sensor capture Distinct from queue length
Explain the failure case

If the queue never empties, have we proved that the executed action uses a fresh observation?

No. Queue continuity answers whether a command is available. Observation freshness asks when the information used to choose that command was captured. An asynchronous base proposal can remain stale even during perfectly smooth execution.

Check your understanding

Which quantity is approximately 3.75 Hz in the reported real setup?

Evidence: PDF pp. 2, 5, 7; Sections I, IV-A, V-C; Table IV p. 16. Paper PDF · Pinned implementation

Chapter 1: Commit the prefix

While the VLA computes, the robot cannot take back the commands already in flight. A new plan must account for that commitment. Start by treating a chunk as a row of numbered commands, then ask which cells belong to the past by the time inference finishes.

Keep three lengths separate

The predicted horizon H is the number of rows generated by the base. The executed horizon C is the number selected for the next rollout segment. The delay d is how many rows are already committed during inference. The new execution window starts after those committed rows, so it is the half-open slice [d:d+C].

In the real setup H=16 and C=8. With d=5, the first five positions supply an inpainted prefix, positions five through twelve are retained, and the final three positions are not executed by that selection. With d=3, the retained window instead starts at position three.

Why condition instead of just discard?

Throwing away early output rows does not tell the model what motion occurred while it was computing. Prefix conditioning provides those committed actions as clean input. The model can propose a continuation consistent with the commands that will actually precede it.

This is the training-time real-time chunking idea used by the paper. During generation, the prefix is clamped; it is not another set of actions the model may freely rewrite. The later editor changes the retained future segment, not commands that have already been executed.

Explore: Queue

Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.

QuantityComputed value

Work it out by hand

  1. Use a small teaching chunk [10,11,20,21,22,23], with H=6, d=2 and C=3.
  2. The committed prefix is [10,11]. These commands cannot be rewritten.
  3. The retained execution slice [2:5] is [20,21,22].
  4. The last two executed commands are [21,22]. They become the next committed prefix.
  5. Check H ≥ d+C. The official real inference code additionally enforces H ≥ 2C.
Element Value or shape Interpretation
Full prediction H × 7 16 × 7 in real tasks
Committed prefix d × 7 3 or 5 command rows
Retained window C × 7 8 × 7
Future prefix Last d executed rows Includes edits
Boot regime d=0 No preceding queue
candidate = base(old observation, committed prefix, noise); execute candidate[d:d+C]

The prefix comes from execution

Suppose the last selected chunk was edited before execution. Its tail, not the corresponding unedited base proposal, is the next committed prefix. Reusing the base tail would condition on a motion the controller did not actually command.

The official AsyncChunkSampler saves the corrected normalized and padded execution window. When exactly d actions remain, it launches background sampling using the current observation and the saved tail. At the next boundary it passes the resulting candidates to the fast selector with a new observation.

Boot and deadline behavior

The first chunk has no preceding committed tail, so it uses a boot path with delay zero. Later chunks have an existing queue and can overlap computation. The simple path assumes d does not exceed C and enough generated rows remain after the prefix.

If a background computation is unfinished at the boundary, the implementation waits for its future. Asynchrony is a scheduling strategy, not a proof that deadlines can never be missed. The widget reports an invalid window instead of silently truncating when H is too short.

Run the core calculation

This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.


def retain_window(candidate, delay, execute):
    if not 0 <= delay <= execute:
        raise ValueError("delay must fit in one execution window")
    if len(candidate) < delay + execute:
        raise ValueError("prediction is too short")
    retained = candidate[delay:delay + execute]
    return retained

candidate = [10, 11, 20, 21, 22, 23]
delay = 2
execute = 3
window = retain_window(candidate, delay, execute)
assert window == [20, 21, 22]

# Assume these are the commands after all edits.
executed_window = list(window)
next_prefix = executed_window[execute-delay:execute]
assert next_prefix == [21, 22]

# Empty prefix at episode boot.
boot = retain_window(candidate, 0, execute)
assert boot == [10, 11, 20]

# Half-open slices make the lengths explicit.
assert len(window) == execute
assert len(next_prefix) == delay

Symbols and data you should be able to name

Name Meaning Concrete referent
H Predicted horizon 16 in real experiments
Prefix Committed actions during inference Clamped input
Postfix Actions following committed rows Denoised output
Half-open slice Includes start, excludes end [2:5] has 3 rows
Explain the failure case

Why is retaining candidate[d:d+C] alone insufficient to implement real-time chunking?

It removes already elapsed positions but does not tell the generator which actions are committed. Prefix-conditioned generation uses the actual committed commands, and the next prefix must come from corrected commands that were executed.

Check your understanding

Which two commands form the next prefix in the worked example?

Evidence: PDF pp. 3–5; Eq. 4; Figure 2; official loop_utils.py. Paper PDF · Pinned implementation

Chapter 2: A good prior needs a local correction

A pretrained VLA can propose plausible ways to grasp a moving block. Asking a tiny policy to invent all that behavior from a few minutes of reward is a different problem. EXPO separates proposing a behavior from learning how to improve a proposal.

The editor has an easier input

The small policy receives both an observation and a base action chunk. It does not begin with a blank action space. It predicts a residual that is added to a candidate, letting the large model supply the behavioral prior while the small model learns local, value-seeking adjustments.

The residual distribution is a tanh-squashed Gaussian. A Gaussian sample can be arbitrarily large; tanh maps it between minus one and one. Multiplication by an edit scale limits each normalized residual coordinate. This is a parameterization of a local search, not a physical safety guarantee.

A value function supplies direction

The critic estimates discounted task return for an observation and a candidate. The editor is trained to produce residuals that make the combined action score better, with an entropy term that discourages collapsing its sampled distribution too early. The gradient changes the editor’s weights.

That value gradient does not run directly through the large VLA backbone. The base continues learning through a separate prefix- conditioned behavior-cloning objective. In the real experiments that supervised update uses demonstrations and successful online episodes, while the critic can learn from failed experience too.

Explore: Editor

Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.

QuantityComputed value

Work it out by hand

  1. Base action 0.3 has squared distance (0.3−0.5)² = 0.04.
  2. The editor samples a raw Gaussian value; tanh of that sample is 0.8 in this worked case.
  3. With scale 0.1, the residual is 0.08.
  4. The edited action becomes 0.38, giving squared distance 0.0144.
  5. A critic that uses negative squared distance ranks −0.0144 above −0.04.
Element Value or shape Interpretation
Base 0 0.10 Q = −0.1600
Base 1 0.30 Q = −0.0400
Base 2 0.70 Q = −0.0400
Edited 1 0.38 Q = −0.0144
Edit dimension 56 8 commands × 7 coordinates
edit = β tanh(z); edited action = base + edit; editor loss = E[α log π(edit) − Q(current, edited)]

Calculate one correction

The widget uses a deliberately transparent teaching critic: score equals the negative squared distance from a desired normalized scalar action. Desired action is 0.5. Three base proposals are 0.1, 0.3 and 0.7; their scores are minus 0.16, minus 0.04 and minus 0.04.

Edits of 0.05, 0.08 and minus 0.05 produce 0.15, 0.38 and 0.65. The best of these is 0.38 with score minus 0.0144. This is not a neural network or a measured robot result. It exposes every arithmetic step that a learned critic would otherwise hide.

A small bound can both help and hurt

A narrow edit range keeps behavior near the prior. If the useful command is far away, the editor cannot reach it from a poor base candidate. More diverse base proposals can help cover different useful regions; they do not guarantee coverage.

A large bound allows bigger corrections but also allows a mistaken critic to favor more destructive changes. The paper’s normalized edit scales are task choices: 0.05 for kicking and 0.1 for the other tasks. Do not reinterpret them as centimeters or degrees.

Run the core calculation

This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.


from math import tanh, atanh

base = [0.1, 0.3, 0.7]
scale = 0.1
desired = 0.5
raw_samples = [atanh(0.5), atanh(0.8), atanh(-0.5)]

def teaching_q(action):
    # Transparent teaching rule, not a trained critic.
    error = action - desired
    return -(error * error)

edits = []
edited = []
for action, raw in zip(base, raw_samples):
    residual = scale * tanh(raw)
    edits.append(residual)
    edited.append(action + residual)

assert all(abs(e) <= scale for e in edits)
all_candidates = base + edited
scores = [teaching_q(a) for a in all_candidates]
choice = max(range(len(scores)), key=scores.__getitem__)
print(edits, scores, choice)
assert abs(all_candidates[choice] - 0.38) < 1e-12

Symbols and data you should be able to name

Name Meaning Concrete referent
β Residual scale 0.05 or 0.1 in reported tasks
z Gaussian sample before tanh Unbounded sample
Q Estimated future task return Learned, not guaranteed correct
α Entropy temperature Controls editor exploration tradeoff
Explain the failure case

Why can a bounded editor fail even if its critic ranks every available candidate correctly?

The useful action may lie outside the residual range of every base candidate. Perfect ranking cannot select an action absent from the candidate set. A diverse and competent prior still matters.

Check your understanding

What does the edit scale bound directly?

Evidence: PDF pp. 3, 5, 14; Equations 1 and 8; Table IV. Paper PDF · Pinned implementation

Chapter 3: Observe again, edit, then choose

Two candidate grasps are ready, but the block has moved toward one and away from the other. Which observation should decide the final choice? The important computation happens after the expensive proposals are available: read the latest state, correct candidates, then rank them.

Replay Figure 2 as an algorithm

Start background generation when d commands remain in the current queue. The base receives that moment’s observation and the committed prefix. It proposes multiple full action chunks. While it works, the queue advances and the physical state changes.

At the next chunk boundary, retain the future C-command segment of each candidate. The lightweight editor receives a fresh observation and each segment. It produces one residual per candidate in the reported real rollout configuration. The critic evaluates originals and edited alternatives at the fresh observation.

The base remains eligible

With 32 base candidates and one edit for each, selection considers 64 alternatives. It is a union, not an average. A candidate can win without editing. Keeping originals avoids forcing every action through a correction that may be unnecessary or mistaken.

The toy below makes that visible. Move the desired action to a base candidate, then set a large incorrect edit. The base can still win. If you remove originals from the selection pool, that fallback disappears. This explains an algorithmic choice without claiming the fallback guarantees success.

Explore: Showcase

Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.

QuantityComputed value

Work it out by hand

  1. Generate the same three scalar teaching candidates: 0.1, 0.3 and 0.7.
  2. Hold the old desired value at 0.2, then move the current desired value to 0.5.
  3. Produce bounded residuals using the chosen editor observation.
  4. For each original and edited action, compute the two teaching critic scores and their minimum.
  5. Pick the largest minimum. Inspect whether the winning action is original or edited before revealing its current error.
Element Value or shape Interpretation
Generation input Old observation + prefix Expensive asynchronous path
Edit input Fresh observation + retained candidate Small synchronous path
Rollout pool 32 originals + 32 edited 64 alternatives
Score aggregation Minimum of two sampled target critics Then argmax
Execution One 8 × 7 chunk No averaging across winners
chosen = argmax over {base candidates, edited candidates} of min(Q′₁(current,a), Q′₂(current,a))

Two critics before one winner

The detailed implementation has ten Q-networks. At selection, it samples two target networks, takes their minimum value for each candidate, and then chooses the candidate with the largest resulting score. Taking a minimum discourages accepting a proposal supported only by one optimistic estimate.

This does not form a formal uncertainty bound. Two critics can agree and both be wrong. The widget uses a fixed second illustrative estimate that is slightly lower than the first, so you can inspect the minimum without treating it as a calibrated uncertainty estimate.

The observation timestamp is the treatment

Toggle fresh versus stale editing and scoring while keeping the proposals fixed. This isolates the intervention point: the fast path can respond to changes that happened after generation began. It is different from merely conditioning the base on the action prefix.

The paper compares against EXPO-FT with RTC as well as without RTC. The stronger comparison helps separate continuous asynchronous execution from the additional benefit of fresh-observation corrections. Do not present RTC as incapable of ever being combined with reinforcement learning.

Run the core calculation

This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.


base = [0.1, 0.3, 0.7]
current_target = 0.5
scale = 0.1

def bounded_delta(action, target):
    return max(-scale, min(scale, target-action))

def q1(action):
    return -(action-current_target)**2

def q2(action):
    return q1(action) - 0.01 * abs(action)

edited = [a + bounded_delta(a, current_target) for a in base]
pool = base + edited
pair_scores = [(q1(a), q2(a)) for a in pool]
conservative = [min(pair) for pair in pair_scores]
selected = max(range(len(pool)), key=conservative.__getitem__)

for index, (action, scores) in enumerate(zip(pool, pair_scores)):
    kind = "base" if index < len(base) else "edited"
    print(index, kind, action, scores)
print("selected", selected, pool[selected])
# Handcrafted editor/critics demonstrate the selection rule.
# They are not the paper's trained networks.

Symbols and data you should be able to name

Name Meaning Concrete referent
N Number of base candidates 32
Candidate union Original and edited alternatives 64 in real rollout
Target critic Slowly updated value network Subsample two of ten
Argmax Choose one best-scoring index Not a weighted average
Explain the failure case

If an edited candidate has the highest score from one critic, must it be selected?

No. The actual selector first takes the minimum across the sampled critic pair for each candidate. Another candidate can have a larger minimum, including an original base candidate.

Check your understanding

When does the fast path receive its new observation?

Evidence: PDF pp. 4–5, 14; Figure 2 and Equations 4–6; _jitted_fast_select. Paper PDF · Pinned implementation

Chapter 4: Train for the queue you will run

At deployment the first few commands are already fixed, but ordinary flow training starts from noise everywhere. That mismatch is avoidable. Give training examples a clean action prefix and ask the network to predict only the remaining commands.

First solve one scalar interpolation

Let a demonstrated action be 5 and a sampled noise value be 2. In the official code’s convention, flow time zero means clean action and time one means noise. At time one quarter, the interpolated input is one quarter of 2 plus three quarters of 5, which gives 4.25.

The derivative with respect to this flow time is noise minus action, or minus 3. Sampling moves time from one toward zero, so an Euler update uses a negative time step. A negative step multiplied by minus 3 moves the sample toward 5. Flow time is not robot time.

Now make the prefix clean

For a chunk, every position gets its own time input. Prefix positions receive clean time zero and their ground-truth action values. Postfix positions receive the sampled noisy time. The loss mask is zero on the prefix and one afterward.

The network can attend to the clean prefix to infer a coherent continuation. It is not rewarded or penalized for a velocity prediction on a committed row. During sampling those rows remain clamped after every Euler step, so numerical integration cannot accidentally move them.

Explore: Masked flow

Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.

QuantityComputed value

Work it out by hand

  1. Use four two-coordinate action rows: [1,2], [3,4], [5,6], [7,8].
  2. Use noise rows [0,0], [1,1], [2,2], [3,3] and make the first two rows committed.
  3. At noisy time 0.25, the input is [1,2], [3,4], [4.25,5], [6,6.75].
  4. The final two target velocities are [−3,−4] and [−4,−5].
  5. Predictions [−2,−5] and [−4,−4] have coordinate-averaged errors 1 and 0.5. Mean over the two active rows is 0.75.
Element Value or shape Interpretation
Rows 0–1 Clean prefix Loss weight 0
Row 2 input [4.25, 5.00] Target velocity [−3, −4]
Row 3 input [6.00, 6.75] Target velocity [−4, −5]
Per-row squared errors 1.00 and 0.50 Mean over coordinates first
Masked mean loss 0.75 Prefix errors excluded
xᵤ = u ε + (1−u) A; target velocity = ε−A; loss = mean over uncommitted rows of squared velocity error

A source convention needs care

Printed Equation 7 uses an interpolation that is clean at time one but retains a target written as noise minus action. That target is opposite to the derivative of the printed interpolation. The official code instead uses the internally consistent clean-at-zero convention explained here.

The core idea is unchanged: a clean prefix conditions a noisy postfix and the loss excludes committed rows. The lesson uses the released implementation’s convention for runnable arithmetic and labels the paper mismatch, rather than combining two incompatible time directions.

Offline and online prefixes differ

During initial supervised training, the prefix length is sampled per example over a range. This exposes the model to both boot and delayed execution. In reported real online training, prefix length is zero for first-chunk transitions and fixed to the deployment delay later.

The simulation procedure instead resamples prefix length during online base fine-tuning and includes all rollout data. Real-world base updates use successful episodes. Neither difference should be erased by one generic statement that “the base trains on all data”.

Run the core calculation

This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.


actions = [[1.,2.],[3.,4.],[5.,6.],[7.,8.]]
noise = [[0.,0.],[1.,1.],[2.,2.],[3.,3.]]
prefix = 2
u = 0.25
mask = [int(k >= prefix) for k in range(4)]

def interpolate(k, j):
    if k < prefix:
        return actions[k][j]
    return u*noise[k][j] + (1-u)*actions[k][j]

x = [[interpolate(k,j) for j in range(2)] for k in range(4)]
target = [[noise[k][j]-actions[k][j] for j in range(2)] for k in range(4)]
predicted = [[999.,999.],[999.,999.],[-2.,-5.],[-4.,-4.]]
row_loss = []
for prediction, truth in zip(predicted, target):
    row_loss.append(sum((p-y)**2 for p,y in zip(prediction,truth))/2)
loss = sum(w*e for w,e in zip(mask,row_loss))/max(sum(mask),1)
assert x[:prefix] == actions[:prefix]
assert abs(loss-0.75) < 1e-12
print(x, target, row_loss, loss)
# Huge errors in committed rows do not change the masked loss.

Symbols and data you should be able to name

Name Meaning Concrete referent
u Flow interpolation time Clean at zero in code
ε Sampled noise matrix Not sensor noise
m_d Postfix loss mask Zero on committed rows
v Predicted flow velocity Integrated from noisy time to clean time
Explain the failure case

Why can a prediction of 999 on a committed row leave this example’s loss unchanged?

That row has loss weight zero and is clamped as clean conditioning input. The model must predict the uncommitted future; the loss is averaged over active postfix rows only.

Check your understanding

In the code convention used here, which flow time denotes a clean prefix?

Evidence: PDF p. 5 Eq. 7; Appendix E pp. 14–16; official pi05.py lines 258–306. Paper PDF · Pinned implementation

Chapter 5: Give the critic the right transition

A critic must learn from what the robot actually executed. If the selected action is an eight-command chunk, the next decision state arrives after eight commands. Treating that as an ordinary one- command transition confuses both reward timing and the next action’s information.

Regroup the return

Start with the ordinary discounted return. Collect the first C rewards into one chunk reward. Each later reward in that chunk has one additional factor of the per-command discount. After those C commands, the remaining future value is multiplied by the discount raised to C.

The target therefore combines the discounted rewards actually observed during the chunk and a discounted estimate at the next chunk boundary. A continuation mask makes the bootstrap zero if the episode has terminated. The replay buffer must preserve where each reward and terminal flag occurred.

Calculate terminal and nonterminal cases

Use a small three-command example with per-command discount 0.9. For a nonterminal chunk with all rewards zero and next value 0.6, the target is 0.9 cubed times 0.6, or 0.4374. There is no reward term to add.

Now suppose success is detected on the third command and terminates the episode. The reward sequence is zero, zero, one. Its discounted sum is 0.81 and the next-value contribution is zero. These are different cases: the real success detector does not emit terminal success and also allow an ordinary future bootstrap.

Explore: Chunk target

Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.

QuantityComputed value

Work it out by hand

  1. For C=3, compute 0.9 × 0.9 × 0.9 = 0.729.
  2. With no rewards and a nonterminal next value of 0.6, multiply 0.729 × 0.6 = 0.4374.
  3. For terminal success on the third command, the reward contributes 0.9² × 1 = 0.81.
  4. The terminal continuation mask is zero, so even a large next Q estimate contributes nothing.
  5. For real C=8 and per-command discount 0.99, the continuation factor is 0.99⁸ ≈ 0.922745.
Element Value or shape Interpretation
Nonterminal rewards [0,0,0] Next Q = 0.6
Nonterminal multiplier 0.9³ = 0.729 Target 0.4374
Terminal rewards [0,0,1] Success ends episode
Terminal target 0.81 No next-value bootstrap
Actual window discount 0.99⁸ ≈ 0.9227 For ordinary task discount setting
y = Σ from i=0 to C−1 of γⁱ rᵢ + γᶜ × continuation × Q′(next state, selected next chunk)

The next base action also has a past

At the next boundary, the backup should emulate rollout selection. The base candidate is generated from the observation d steps before that boundary, conditioned on committed actions. Editing and action- value selection then use the next-boundary observation.

Keep current state, next state and delayed next-generation state as separate records. Substituting the next state into the slow base would give the target an information advantage unavailable during deployment. Substituting the stale state into the editor would remove the proposed correction mechanism.

Paper shorthand and implemented target

Equation 9 writes a chunk-level target using a reward and one discount symbol. The released code explicitly accumulates discounted rewards across the chunk and uses the per-command discount raised to the execution length, together with a continuation mask.

You can treat a single discount symbol as chunk-level shorthand if it is defined that way. You cannot silently take a per-command value such as 0.99 and apply it only once to an eight-command transition. In that case the implemented multiplier is approximately 0.9227, not 0.99.

Run the core calculation

This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.


def chunk_target(rewards, masks, gamma, next_value):
    total = 0.0
    continuation = 1.0
    for index, (reward, mask) in enumerate(zip(rewards, masks)):
        total += gamma**index * continuation * reward
        continuation *= mask
    bootstrap = gamma**len(rewards) * continuation * next_value
    return total + bootstrap

nonterminal = chunk_target(
    rewards=[0,0,0],
    masks=[1,1,1],
    gamma=0.9,
    next_value=0.6,
)
terminal = chunk_target(
    rewards=[0,0,1],
    masks=[1,1,0],
    gamma=0.9,
    next_value=100.0,
)
assert abs(nonterminal-0.4374) < 1e-12
assert abs(terminal-0.81) < 1e-12
print(nonterminal, terminal, 0.99**8)

Symbols and data you should be able to name

Name Meaning Concrete referent
γ Per-command discount 0.99 in ordinary real-task settings
C-step reward Discounted rewards inside selected chunk Computed from replay
Continuation Whether future bootstrap is valid Zero after terminal
Delayed next state Snapshot used to generate next candidate Earlier than next-boundary state
Explain the failure case

Why is the slow base’s observation in the next-action backup earlier than the observation used by the next critic?

The backup reproduces the deployment information sequence. Slow generation starts before the boundary; the fast editor and selector see the boundary observation. Collapsing them into one timestamp changes the policy being evaluated.

Check your understanding

A terminal success occurs on the third command in the worked example. What is the target?

Evidence: PDF p. 5 Eq. 9; Appendix E/F; replay_buffer.py lines 516–530; realtime_expo_ft.py line 1067. Paper PDF · Pinned implementation

Chapter 6: Reject noise before decoding

Building a critic target repeatedly can be more expensive than choosing one real action. If each target samples thirty-two full flow trajectories, much of training compute is spent decoding candidates that will immediately be discarded. Can we reject some candidates before paying that cost?

A seed determines a sampled candidate

A generative policy receives an observation and a random noise seed, then numerically denoises it into an action chunk. Under fixed model and observation inputs, the seed identifies a sampled proposal. A small network can learn to predict the action value associated with that seed.

The noise-ranking critic is not a physical world model and does not observe the future. It learns a supervised relationship between a seed, the relevant observations and a target action-value estimate. Its ranking can be wrong, especially while the action critic or base generator is changing.

Read the right half of Figure 2

During a Bellman backup, draw thirty-two raw noise seeds and score them cheaply. Keep the highest-scoring seed, decode it once, sample one edit, then let the target action critic select between that base and edited candidate. The reported backup therefore decodes one survivor rather than thirty-two.

The outer action-value comparison remains. Noise ranking chooses which base candidate to instantiate; it does not replace the editor or the final target critic. This sequence is useful precisely because a cheap approximate first stage is followed by a more direct action-space evaluation.

Explore: Noise filter

Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.

QuantityComputed value

Work it out by hand

  1. A small teaching pool has filter scores [0.2,0.9,0.1,0.5].
  2. Argmax keeps seed index 1 with score 0.9.
  3. Decode that seed, make one edit, and compare the two action candidates with the target critic.
  4. With base target Q=0.6, code-style filter regression gives (0.9−0.6)² = 0.09. Paper Eq. 11’s edited-winner target of 0.75 would instead give 0.0225.
  5. With thirty-two seeds and ten Euler steps, display 320 versus 10 denoising sample-steps, explicitly excluding other compute.
Element Value or shape Interpretation
Backup seed pool 32 Cheap filter scores
Decoded survivors 1 Full VLA decode
Edited backup alternatives 1 Base + one edited action
Deployment alternatives 64 32 originals + 32 edits
Filter objective Squared error to stopped target Q Separate supervised ranker
best seed = argmax Q_filter(observations, seed); filter loss = (predicted seed value − stop_gradient(target action value))²

Teach what the compute count means

Ten Euler steps for each of thirty-two candidates gives three hundred and twenty denoising sample-steps. Decoding one survivor gives ten. Those counts help explain the source of savings, but they are not measured wall-clock times or proof of a thirty-two-fold end- to-end speedup.

Batching, cached vision-language features, memory traffic, filtering overhead and the remaining updates affect actual performance. Figure 7 compares learning against training compute on four simulation tasks. Its curves support improved compute efficiency; they do not provide a general deployment latency multiplier.

Keep training and deployment separate

The real rollout path still draws thirty-two base candidates and thirty-two edited candidates before choosing its action. The optional filtering mechanism changes the backup path, not that deployment pool. Conflating these paths removes an important part of the method.

Equation 11 writes regression toward the final selected action value. The pinned code instead uses the decoded, unedited base candidate’s target Q as the filter target, while separately selecting the best base-or-edited action for the Bellman backup. The detailed filter conditions on both the next/current observation features and the earlier observation used by the base. Equation 10 and Equation 11 use different state indices in compressed notation; the code keeps the two observation roles explicit. The widget uses simple named seed scores to make this bookkeeping inspectable.

Run the core calculation

This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.


seed_scores = [0.2, 0.9, 0.1, 0.5]
selected_seed = max(range(len(seed_scores)), key=seed_scores.__getitem__)
assert selected_seed == 1

# Only this seed would enter the expensive decoder.
def teaching_decode(index):
    return [0.1, 0.3, 0.7, 0.9][index]

base = teaching_decode(selected_seed)
edited = base + 0.08
pool = [base, edited]
target_values = [0.6, 0.75]  # specified teaching values
selected_action = max(range(2), key=target_values.__getitem__)
regression_target = target_values[0]  # pinned code: unedited base
paper_winner_target = target_values[selected_action]
filter_loss = (seed_scores[selected_seed]-regression_target)**2

assert abs(filter_loss-0.09) < 1e-12
N = 32
steps = 10
print("without filtering", N*steps, "sample-steps")
print("with filtering", steps, "sample-steps")
print("filter squared error", filter_loss)
# No measured runtime or trained network is claimed here.

Symbols and data you should be able to name

Name Meaning Concrete referent
Q_filter Noise-space ranking critic Supervised from target action value
Seed Initial noise matrix H × 32 padded space
Stop-gradient Treat target as a fixed number in this regression No backprop through target
Sample-step One denoising step for one candidate Not elapsed milliseconds
Explain the failure case

Why can filtering discard the ultimately best action even when it saves computation?

The cheap seed ranker is an approximation. If it assigns a low score to the seed that would decode to the best action, that action never reaches the outer action critic. The final critic can only choose among surviving candidates.

Check your understanding

Which path uses the 32-seeds-to-1-decode filter in the paper?

Evidence: PDF p. 6 Equations 10–11; Appendix A; Figure 7 p. 13; Appendix E. Paper PDF · Pinned implementation

Chapter 7: Follow one real tensor

A rectangle labeled “policy” hides the engineering choices that make this work. Follow one observation through the real implementation: images become features, candidates become windows, windows become edit vectors, and one chosen vector becomes physical commands.

Images for the base and the fast path

Most tasks use side and wrist RGB images resized to 224 by 224. The VLA processes its visual input with its pretrained encoder. The fast critic/editor path uses a separate residual visual encoder. Stacking the two views along channels gives a 224 by 224 by 6 tensor.

The detailed appendix describes preactivation basic residual blocks with stage depths three, four, six and three, GroupNorm and a 512-dimensional image representation. The main text’s ResNet-50 label does not match canonical bottleneck ResNet-50 details. Use the detailed configuration, rather than assuming a standard library architecture from the name.

State features carry specific information

A state embedding contributes 64 more features. Ball Balancing adds detector-derived plate center, ball position and velocity; Soccer Kicking adds keeper position and velocity. Those features feed the critic, editor and noise filter while the base VLA’s state input remains unchanged.

Ball Balancing also uses three exterior-camera frames instead of the ordinary exterior-plus-wrist pair. Its channel stack has nine channels. The paper removes the vertical proprioceptive dimension from that critic because monotonic drift might reveal episode time. These task-specific choices matter to what the learning system actually observes.

Explore: Data flow

Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.

QuantityComputed value

Work it out by hand

  1. Two 224 × 224 RGB cameras contain six stacked image channels.
  2. The critic encoder projects its representation to 512 image features; state contributes 64.
  3. A real candidate window contains 8 × 7 = 56 command coordinates.
  4. Thirty-two base windows plus thirty-two edited windows create sixty- four candidates.
  5. Argmax returns one 8 × 7 window; retain its corrected tail for the next prefix.
Element Value or shape Interpretation
Normal visual stack 224 × 224 × 6 Side + wrist
Balancing stack 224 × 224 × 9 Three exterior frames
Feature widths 512 image + 64 state Fast critic/editor path
Noise space 16 × 32 per seed Padded model dimensions
Execution and edit 8 × 7 = 56 Physical command coordinates
images → 512 features; state → 64 features; retained actions [8,7] → editor [56] → candidate selection → commands [8,7]

From padded noise to real commands

The π0.5 configuration has a sixteen-step horizon and padded action width thirty-two. The physical output has seven dimensions: translation, rotation and gripper. Padding is a model-interface choice, not thirty-two independent actuators.

After candidate generation and the delay slice, each execution window has eight rows and seven physical coordinates. Flattening gives fifty-six numbers for the Gaussian editor and action critic. One edited version of each of thirty-two base windows yields sixty- four candidate windows. Selection returns one eight-by-seven window, which is unnormalized before execution.

Which parameters are allowed to move?

The base is initialized from a π0.5 configuration with LoRA adapters in the language and action-expert stack. Its non-LoRA language-stack parameters are frozen by the parameter filter. Parameters outside that frozen set, including vision and projection components, are described as trainable; permitted adapters also update.

The critic and editor are separate learned components. The editor shares the fast visual encoding, while direct value learning does not backpropagate through the full base action generator. A wrapper flag named freeze_pi05_encoder also controls cached inference behavior; it is not sufficient evidence that the entire VLA is frozen.

Run the core calculation

This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.


H = 16
C = 8
physical_dim = 7
padded_dim = 32
N = 32

noise_shape = (N, H, padded_dim)
retained_shape = (N, C, physical_dim)
edit_shape = (N, C*physical_dim)
selection_shape = (2*N, C, physical_dim)
assert edit_shape == (32, 56)
assert selection_shape == (64, 8, 7)

# Dynamic Picking permits xyz and gripper residuals.
coordinate_mask = [1,1,1,0,0,0,1]
flat_mask = coordinate_mask*C
raw_edit = [0.1]*(C*physical_dim)
masked = [e*m for e,m in zip(raw_edit,flat_mask)]
for command in range(C):
    row = masked[command*7:(command+1)*7]
    assert row[3:6] == [0,0,0]
print(noise_shape, retained_shape, edit_shape)
print(masked[:7])
# These are shape and mask checks, not neural inference.

Symbols and data you should be able to name

Name Meaning Concrete referent
Physical width Translation 3 + rotation 3 + gripper 1 7
Padded width Model interface width 32
Image embedding Fast visual representation 512
Privileged detector features Task-derived measurements available to fast path Not new physical sensors
Explain the failure case

Why should the Ball Balancing observation configuration be visible in a results explanation?

Its fast path receives a three-frame exterior stack and detector- derived ball information. Those inputs help estimate motion. Omitting them makes the reported result sound like the same unaugmented two-image configuration used elsewhere.

Check your understanding

How many physical coordinates are corrected in a complete eight-command window before task-specific masking?

Evidence: PDF Appendix D/E pp. 13–16; Tables III–IV; official model configuration and selector. Paper PDF · Pinned implementation

Chapter 8: Two clocks for learning and acting

The robot can execute a queued command while the VLA prepares another chunk. That is one kind of concurrency. The learner can also accumulate experience and update parameters later. Those are separate clocks, and the ten-minute headline refers to only one of them.

Begin with a useful prior

The real experiments begin with task demonstrations and supervised prefix-conditioned fine-tuning of π0.5. The paper describes starting online learning once this policy has roughly thirty percent success or better. Pretraining and demonstrations are part of the starting resources, not counted as online robot minutes.

The replay buffer can be seeded with demonstrations, or demonstrations can be sampled as a fixed fraction of a critic batch. Soccer Kicking uses a fifty-percent prior-data fraction; the other reported task settings seed the buffer. These choices affect what the learner sees early in adaptation.

Learning from success and failure

The critic learns to distinguish promising and unpromising chunks through reward-backed targets. Failed experience therefore contains useful value information. The editor’s training objective conditions on replay actions; those are not necessarily freshly sampled base proposals. At rollout it edits generated base candidates. The real- world base behavior-cloning update uses demonstrations and successful online episodes.

This division prevents a simplistic description that all networks imitate every action the agent takes. The base’s supervised signal and the editor’s value-seeking signal have different roles. In simulation, the base’s behavior-cloning loss is applied to all rollout and demonstration data, an explicitly different setting.

Explore: Training

Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.

QuantityComputed value

Work it out by hand

  1. Choose 300 collected transitions and a teaching update interval of K=30.
  2. That accumulates 300/30 = 10 update calls, which may wait for the episode boundary.
  3. Ten calls produce 200 critic steps.
  4. They produce 10 base steps, 10 editor steps and 10 temperature steps.
  5. The task interaction clock counts environment data; it does not measure those gradient steps or initial demonstration time.
Element Value or shape Interpretation
Dynamic Picking K=25, C=8, d=3 Table IV settings
Soccer Kicking K=20, C=8, d=5 50% demonstration batches
Ball Balancing K=30, C=8, d=5 Demonstrations seeded
Object Passing K=30, C=8, d=5 Demonstrations seeded
One update call 20 critic + 1 base + 1 edit + 1 temperature Not 20 VLA updates
calls = floor(collected transitions / K); critic steps = 20 × calls; base/editor/temperature steps = calls

Count updates carefully

Each reported real update call uses twenty critic gradient steps, followed by one base step, one editor step and one temperature step. Calls are accumulated according to a task-specific number of collected transitions and flushed at episode boundaries. Training begins after ten completed episodes.

The widget lets you select collected transitions and the transition interval. Its counts are a bookkeeping illustration of those settings. It does not simulate learning curves or estimate the minutes required on a GPU. The exact public example script also differs from Table IV in some task values, so the lesson labels table settings separately.

What entropy does and does not do

The learned temperature weights entropy in the edit-policy objective. It changes the preference for diverse residual samples relative to value-seeking. The target entropy is minus half the flattened edit dimension; with fifty-six coordinates its value is minus twenty-eight.

Entropy is not included in the released Bellman backup. It is an optimization term for the small residual policy, not an additional environment success reward. Keeping that distinction makes it possible to trace which loss changes which component.

Run the core calculation

This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.


def update_counts(collected, interval, episodes):
    accumulated = collected // interval
    active = episodes >= 10
    calls = accumulated if active else 0
    return {
        "eligible_calls": calls,
        "critic_steps": 20*calls,
        "base_steps": calls,
        "edit_steps": calls,
        "temperature_steps": calls,
    }

counts = update_counts(300, 30, episodes=10)
assert counts["critic_steps"] == 200
assert counts["base_steps"] == 10
assert update_counts(300, 30, episodes=9)["eligible_calls"] == 0
print(counts)

# Gradient work is flushed at episode boundaries.
# This calculator does not model actual GPU time.
# Demonstrations and supervised initialization happened earlier.
D = 8*7
print("target entropy", -D/2)
assert -D/2 == -28

Symbols and data you should be able to name

Name Meaning Concrete referent
K Transitions per update call in Table IV Different from sampling-count notation elsewhere
UTD Update-to-data scheduling parameter Interpret with the stated loop
Replay Reused past experience Includes successes and failures
Temperature Learned entropy coefficient Editor objective only
Explain the failure case

Would doubling the number of critic updates necessarily double the robot data used?

No. Updates reuse replay data. Gradient compute and new environment interaction are different budgets. More updates can change learning behavior, but are not automatically new robot trials or guaranteed improvements.

Check your understanding

Which data trains the real-world base behavior-cloning update?

Evidence: PDF Section IV-C; Appendix E pp. 15–16; Tables III–IV; Appendix C for simulation differences. Paper PDF · Pinned implementation

Chapter 9: Read the simulation table precisely

A method can have the strongest average without being the strict winner on every task. The Kinetix table is a useful place to practice reading a result more carefully than its headline. Every displayed number in this chapter comes from Table II.

What was actually simulated?

The experiments use ten symbolic-state Kinetix environments and a pretrained state-based flow policy. No vision-language-action model is involved in this simulation setting. The test asks whether the proposed delayed generative-policy mechanism works across dynamic control problems.

The base is pretrained using one million transitions, then fine- tuned online for one hundred thousand environment steps per run. The delayed flow-policy methods incur four steps of inference delay and replan every four steps. RLPD’s small Gaussian actor runs without delay.

Separate evaluation protocols

Reported RL scores average four random seeds and one hundred evaluation episodes per seed. The BC and RTC reference evaluations use five hundred and twelve episodes. These protocols should remain attached to the result display; they are not identical trial budgets.

The paper provides rounded per-task percentages and an average row, not all raw seed outcomes in the table. The interactive view therefore computes differences from those rounded values and does not manufacture confidence intervals or reconstruct precise training trajectories.

Explore: Simulation

Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.

QuantityComputed value

Work it out by hand

  1. Sum the proposed method’s ten rounded scores: 962.
  2. Divide by ten to obtain the reported average of 96.2 percent.
  3. Subtract EXPO-FT with RTC’s 81.7 to obtain 14.5 percentage points.
  4. On Chain Lander the best RL score is 96; 95% of 96 is 91.2.
  5. The proposed score of 92 exceeds 91.2 but remains below 96. Near- best and strict winner are different statements.
Element Value or shape Interpretation
Car Launch 98 Highest reported score
Cartpole 98 EXPO-FT 99; BC no-delay 100
Lunar Lander 94 Some baselines 95
Chain Lander 92 DSRL with RTC 96
Ten-task average 96.2 Table II rounded values
average = sum(task scores) / 10; near-best criterion = score ≥ 0.95 × best RL score on that task

A defensible aggregate

Real-Time EXPO-FT averages 96.2 percent across the ten tasks. EXPO- FT with RTC averages 81.7 percent, and no-delay RLPD averages 81.4 percent. The difference from the stronger delayed EXPO-FT with RTC baseline is 14.5 percentage points in the reported average row.

These are results for the tested configurations, not a theorem that a delayed large-policy system always beats a small real-time actor. Actor structure, demonstrations, optimization schedules and task priors differ between some baseline families.

What “best in ten” can hide

The abstract describes best performance in ten of ten environments, but the table does not support a strict-first-place interpretation. Cartpole shows 98 for the proposed method versus 99 for EXPO-FT and 100 for no-delay BC. Hard Lunar Lander shows 94 versus 95. Chain Lander shows 92 versus 96.

The table’s boldface criterion is within 95 percent of the best RL score on a task, excluding BC and RTC references. The proposed method meets that criterion in all ten. The lesson keeps both the strong average and the three non-winning rows visible, rather than rewriting the table to fit a slogan.

Run the core calculation

This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.


ours = [98,98,93,97,99,94,98,94,92,99]
expo_rtc = [64,96,47,85,87,90,79,83,94,92]
assert sum(ours)/10 == 96.2
assert sum(expo_rtc)/10 == 81.7

chain_ours = 92
chain_best_rl = 96
threshold = 0.95*chain_best_rl
near_best = chain_ours >= threshold
strict_best = chain_ours >= chain_best_rl
assert near_best and not strict_best

non_winners = {
    "Cartpole": (98,99),
    "Hard Lunar Lander": (94,95),
    "Chain Lander": (92,96),
}
for task, (score,best) in non_winners.items():
    print(task, score-best, "percentage points")
print("average", sum(ours)/len(ours))
# Table values are measured; this code only recomputes arithmetic.
# It does not run Kinetix or estimate learning curves.

Symbols and data you should be able to name

Name Meaning Concrete referent
Percentage point Difference of two percentages 96.2−81.7=14.5
Seed Independent randomized training run Four for RL
Reference BC Pretrained flow without online fine-tuning 512 evaluation episodes
Near-best At least 95% of best RL score Not statistical equivalence
Explain the failure case

Why would a bar chart saying “VLA wins ten robot tasks” be inaccurate here?

Kinetix uses a symbolic-state flow policy without a VLA; its environments are simulations. Also, Table II has three rows where the proposed method is below another reported score. The strong supported claim is its average and near-best RL criterion.

Check your understanding

Which statement agrees with Table II?

Evidence: PDF Section V-B; Appendix B/C; Figure 3; Table II p. 13. Paper PDF · Pinned implementation

Chapter 10: Four tasks, thirty trials each

The real robot learns to catch an offered object, balance a ball, grasp a moving block and kick past a moving defender. These are compelling demonstrations. Their success counts become more informative when the starting policy, comparison method and data budget stay visible.

Keep the four environments distinct

Dynamic Picking uses a block on a rotating plate and the original roughly 67-millisecond base latency. Ball Balancing, Object Passing and Soccer Kicking add 100 milliseconds to emulate a slower model or constrained compute, yielding approximately 167 milliseconds. Their scheduled lead is five commands; Picking uses three.

Each task has its own initial-state variation and success detector. Picking uses gripper state and end-effector height sustained for several steps. Balancing requires the ball to remain near the plate center across consecutive frames. Detector definitions turn a physical outcome into the sparse binary reward used for learning.

Compute the headline with its denominator

The initial SFT policy obtains 19, 8, 10 and 13 successes out of 30, respectively. That totals 50 successes out of 120 trials, or 41.67 percent. The proposed method obtains 30, 28, 30 and 28, totaling 116 out of 120, or 96.67 percent.

Rounded, that is the paper’s 42-to-97-percent headline. It describes these four task evaluations, not a general reliability rate for arbitrary robot work. Two tasks have two failures each; thirty successes in another task are finite observations, not proof that failure is impossible.

Explore: Real evidence

Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.

QuantityComputed value

Work it out by hand

  1. SFT totals: 19+8+10+13 = 50 successes over 4×30 = 120 trials.
  2. Proposed totals: 30+28+30+28 = 116 successes over the same displayed total.
  3. The exact aggregate rates are 41.67% and 96.67%; rounded headline values are 42% and 97%.
  4. EXPO-FT with RTC totals 24+23+27+26 = 100 successes, or 83.33%.
  5. Difference from this stronger baseline: (116−100)/120 ×100 = 13.33 percentage points.
Element Value or shape Interpretation
Dynamic Picking 30/30 EXPO-FT + RTC: 24/30
Ball Balancing 28/30 EXPO-FT + RTC: 23/30
Object Passing 30/30 EXPO-FT + RTC: 27/30
Soccer Kicking 28/30 EXPO-FT + RTC: 26/30
Aggregate 116/120 Stronger baseline: 100/120
aggregate rate = total successes / total displayed trials; comparison gain = (116−100)/120 ×100 = 13.33 percentage points

The stronger baseline belongs beside it

SFT with RTC totals 72 of 120, or 60 percent. EXPO-FT with RTC totals 100 of 120, or 83.33 percent. Comparing the proposed method with the latter gives 16 additional successes in the displayed evaluations, or 13.33 percentage points.

The widget defaults to that stronger baseline. Switch to SFT to recover the headline calculation, but keep the baseline’s name attached to its value. A true number with the wrong comparison label can tell a misleading story.

What ten minutes and no intervention mean

Training is capped at ten minutes of online robot interaction per task, or stopped according to the comparison’s 30-of-30 evaluation rule. It excludes the pretrained foundation policy, demonstrations, initial supervised training, compute and resets. The Object Passing run also has a shorter reported transition budget than the roughly eighteen thousand steps listed for three other tasks.

No corrective human interventions are used during online policy rollouts. Some environment resets still require humans, and evaluation success is independently checked by a person. The authors explicitly identify reset labor and task-specific success detectors as limitations. Both qualifications belong in an honest account of rapid adaptation.

Run the core calculation

This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.


sft = [19,8,10,13]
sft_rtc = [22,12,22,16]
expo_rtc = [24,23,27,26]
ours = [30,28,30,28]
trials_per_task = 30
trials = len(ours)*trials_per_task

def percent(values):
    return 100*sum(values)/trials

assert sum(sft) == 50
assert sum(sft_rtc) == 72
assert sum(expo_rtc) == 100
assert sum(ours) == 116

for name, values in [
    ("SFT",sft),
    ("SFT + RTC",sft_rtc),
    ("EXPO-FT + RTC",expo_rtc),
    ("Real-Time EXPO-FT",ours),
]:
    print(name, sum(values), trials, percent(values))
print("gain in percentage points", percent(ours)-percent(expo_rtc))
# Reported trials, not new robot experiments.

Symbols and data you should be able to name

Name Meaning Concrete referent
Trial One evaluated task attempt 30 per method/task
Aggregate Sum across four tested tasks 120 displayed trials
Intervention Corrective human action during rollout Not the same as reset labor
Detector Task-specific success signal Sparse reward and termination
Explain the failure case

Can “no intervention” be shortened to “no humans needed” in this study?

No. It means no corrective interventions during the online policy rollouts. Demonstrations, some resets and human verification of evaluation success remain part of the experimental process.

Check your understanding

What does the ten-minute cap measure?

Evidence: PDF pp. 7–9; Figures 4–6; Table I; Appendix D/E; Discussion. Paper PDF · Pinned implementation

Chapter 11: When the fast editor is not enough

The block makes an abrupt movement just after a chunk has been selected. A fresh observation at selection cannot reveal a future surprise. Understanding the method means knowing where its feedback ends, what its candidate set cannot express, and which experiment would test the remaining gap.

A bounded correction still has a horizon

After selection, the robot executes the chosen commands until the next chunk boundary. The editor’s latest observation is fresh at that boundary, not throughout the whole future window. Shorter execution windows could change the tradeoff, but the paper’s reported real setting is eight commands.

The edit range can also be too small to repair every base candidate. Increasing the range is not a free solution: a wrong value estimate can favor a large bad correction. The stress widget isolates these mechanisms using an explicitly constructed scalar target and queue, not invented real failure percentages.

Keep the hidden state in mind

The paper motivates the method through delayed observations and Markovian credit assignment. However, a proposal distribution can still depend on an old observation and committed actions. Images may also hide velocity, contact or occluded objects. A current image input alone is not a proof that an arbitrary physical task is fully observed.

For implementation, retain the action queue and relevant observation history as explicit state. For scientific claims, distinguish the paper’s motivation from a universal theorem about Markov restoration or safe control. The experimental evidence is about these tested policies and tasks.

Explore: Limits

Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.

QuantityComputed value

Work it out by hand

  1. Place the desired scalar action at 0.8, all base candidates at or below 0.3, and the edit bound at 0.1.
  2. The largest reachable candidate is at most 0.4; selection cannot produce 0.8.
  3. Now make one candidate reach 0.8 but assign it a falsely low critic value. Coverage alone no longer solves selection.
  4. Now move the target after the selection boundary. Even a correct fresh-observation edit was based on information before that future change.
  5. These isolate coverage, value error and within-chunk surprise. None is a measured failure rate from the paper.
Element Value or shape Interpretation
Insufficient candidate coverage Correct action absent Improve prior or explore alternatives
Wrong critic ranking Correct candidate discarded Measure value estimation and feedback
Late background inference Boundary waits Measure complete timing distribution
Within-chunk surprise Observation becomes old again Evaluate execution-horizon tradeoff
Detector/reset burden Operational adaptation limit Measure actual human and reward work
reachable envelope = union of bounded neighborhoods; selectable set = original candidates plus sampled edits

Different intervention points

RTC keeps chunks coherent while computation overlaps execution. Original EXPO-FT supplies value-guided action editing. Real-Time EXPO-FT applies that fast editing and selection using the latest observation after slow candidate generation. DSRL instead learns to choose noise for a generative policy.

The companion ARLI paper also works in latent noise space: its small policy intervenes before action-expert denoising and can use an intermediate observation plus committed actions. Its frozen-base and timing design differs from this paper’s successful-episode base updates and post-generation action edits. Shared concern about latency does not make the algorithms interchangeable.

What to measure next

A useful deployment study would log observation age, background inference completion, edit-and-select time, command deadlines and actual failure conditions. It would also count resets, reward- detector errors and compute spent per unit of new interaction. These are proposed measurements, not experiments reported here.

Read the paper’s result as a strong demonstration of one way to combine a capable prior with fast learned corrections. Preserve the exact baselines, task scope and implementation conventions. Then you can ask a sharper question than whether the system is simply “real time”: which information is available at each decision, and what can the policy still change?

Run the core calculation

This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.


bases = [0.1,0.2,0.3]
bound = 0.1
target = 0.8
reachable_max = max(bases)+bound
assert reachable_max < target

# Distinguish missing coverage from an incorrect ranking.
covered = [0.2,0.4,0.8]
wrong_scores = [0.9,0.7,0.1]
selected = max(range(len(covered)), key=wrong_scores.__getitem__)
assert covered[selected] != target

# A future change can invalidate a correct current choice.
selected_for_now = 0.8
future_target = 0.1
future_error = abs(selected_for_now-future_target)
print(reachable_max, covered[selected], future_error)

# An experiment should record these separately:
metrics = ["observation_age", "inference_deadline", "edit_latency",
           "candidate_coverage", "critic_ranking", "success_detector_error"]
print(metrics)
# This is a failure decomposition, not a robot simulator.

Symbols and data you should be able to name

Name Meaning Concrete referent
Coverage Which useful actions the candidate set contains Distinct from ranking
Partial observation Sensor input omits state information History may still matter
Boundary Moment a new chunk is selected Not every command
Operational cost Reset, sensing and reward work Not captured by interaction minutes alone
Explain the failure case

Why does adding a current observation not automatically prove the complete policy is memoryless in the physical state?

The candidate set can still depend on the earlier observation and committed action queue, and the observation may not reveal all physical state. Explicit temporal bookkeeping and empirical evaluation remain necessary.

Check your understanding

Which is the most precise summary of the contribution?

Evidence: PDF Discussion p. 9; Equations 4–11; Appendix D/E; companion ARLI source audit. Paper PDF · Pinned implementation