Real-Time EXPO-FT: let a large vision-language-action model propose, then use a fresh observation to edit and select the next action chunk.
Four tested tasks, 30 trials each. Compare against 100/120 for EXPO-FT with RTC. Numerical toys below are labeled separately from paper results.
A blue block travels around a rotating plate. The robot sees it, begins computing a grasp, and finishes after the block has moved. The proposal can be sensible for the old picture and wrong for the present. More accurate imitation of that old picture does not remove its age.
The experiments send Cartesian and gripper velocity commands at 30 Hz: one command slot lasts about 33.3 milliseconds. Their base VLA takes roughly 67 milliseconds to infer on the authors’ server. The world therefore advances by about two command periods during this computation. A model update rate, a command rate and a camera rate are different measurements; never label all three simply “frequency”.
The proposed system executes eight commands before selecting another chunk. At 30 Hz that is about 266.7 milliseconds, or 3.75 chunk selections per second under ideal scheduling. The lightweight editor acts at that boundary. The paper does not demonstrate a fresh editor call at every command slot.
A synchronous controller can stop issuing useful new commands while waiting for its next prediction. An asynchronous controller instead keeps consuming an existing queue while the model prepares a continuation. That removes a computation gap when the continuation is ready on time.
However, overlap alone does not make the model’s input current. The prepared chunk is still based on an earlier observation. This is the opening for Real-Time EXPO-FT: retain the expensive proposal, but let a fast decision use another observation just before the selected chunk is executed.
Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.
| Quantity | Computed value |
|---|
| Element | Value or shape | Interpretation |
|---|---|---|
| Command rate | 30 Hz | Commands sent to the robot |
| Base inference | Approximately 67 ms | Measured on the authors’ server |
| Picking lead | 3 command steps | 100 ms of scheduled overlap |
| Replan window | 8 commands | About 266.7 ms at 30 Hz |
| Extra latency | 100 ms on three tasks | Not added to Dynamic Picking |
The rotating block below is an explicit kinematic toy. Its angle changes at the speed you select; the distance between the old and new positions is computed from that circle. It is not a robot rollout, a learned policy, or a reproduction of the paper’s success rates.
The slider deliberately lets you separate elapsed time from object speed. Double speed at fixed latency and the angular mismatch doubles. Set speed to zero and the stale observation can remain geometrically useful. That does not prove a stationary real task is easy; it isolates one failure mechanism.
Sensor delay, occlusion, contact uncertainty and a poor task prior can all hurt a controller. This paper isolates a particularly important issue for dynamic tasks: an action is generated before the state in which it will be used. Its experimental evidence does not establish that every VLA failure is caused by latency.
Use three timestamps in debugging: when the observation was captured, when inference completed, and when the command was applied. An inference timer alone cannot tell you the age of the information reaching the robot.
This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.
from math import cos, sin, hypot
frequency = 30.0
latency_ms = 67.0
angular_speed = 1.2 # radians/second; teaching choice
radius = 1.0 # normalized teaching coordinates
def position(seconds):
angle = angular_speed * seconds
return radius * cos(angle), radius * sin(angle)
old_position = position(0.0)
now_position = position(latency_ms / 1000.0)
x0, y0 = old_position
x1, y1 = now_position
geometric_error = hypot(x1 - x0, y1 - y0)
command_period = 1.0 / frequency
replan_period = 8.0 / frequency
scheduled_lead = 3.0 / frequency
print(old_position, now_position)
print(geometric_error)
print(command_period, replan_period, scheduled_lead)
assert replan_period > scheduled_lead
# This script integrates no robot dynamics.
# It measures geometric observation mismatch only.
| Name | Meaning | Concrete referent |
|---|---|---|
| f | Command frequency | 30 Hz |
| C | Executed commands before replanning | 8 |
| d | Scheduled inference lead in command steps | 3 or 5 |
| Observation age | Time since sensor capture | Distinct from queue length |
If the queue never empties, have we proved that the executed action uses a fresh observation?
No. Queue continuity answers whether a command is available. Observation freshness asks when the information used to choose that command was captured. An asynchronous base proposal can remain stale even during perfectly smooth execution.
Which quantity is approximately 3.75 Hz in the reported real setup?
Evidence: PDF pp. 2, 5, 7; Sections I, IV-A, V-C; Table IV p. 16. Paper PDF · Pinned implementation
While the VLA computes, the robot cannot take back the commands already in flight. A new plan must account for that commitment. Start by treating a chunk as a row of numbered commands, then ask which cells belong to the past by the time inference finishes.
The predicted horizon H is the number of rows generated by the base. The executed horizon C is the number selected for the next rollout segment. The delay d is how many rows are already committed during inference. The new execution window starts after those committed rows, so it is the half-open slice [d:d+C].
In the real setup H=16 and C=8. With d=5, the first five positions supply an inpainted prefix, positions five through twelve are retained, and the final three positions are not executed by that selection. With d=3, the retained window instead starts at position three.
Throwing away early output rows does not tell the model what motion occurred while it was computing. Prefix conditioning provides those committed actions as clean input. The model can propose a continuation consistent with the commands that will actually precede it.
This is the training-time real-time chunking idea used by the paper. During generation, the prefix is clamped; it is not another set of actions the model may freely rewrite. The later editor changes the retained future segment, not commands that have already been executed.
Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.
| Quantity | Computed value |
|---|
| Element | Value or shape | Interpretation |
|---|---|---|
| Full prediction | H × 7 | 16 × 7 in real tasks |
| Committed prefix | d × 7 | 3 or 5 command rows |
| Retained window | C × 7 | 8 × 7 |
| Future prefix | Last d executed rows | Includes edits |
| Boot regime | d=0 | No preceding queue |
Suppose the last selected chunk was edited before execution. Its tail, not the corresponding unedited base proposal, is the next committed prefix. Reusing the base tail would condition on a motion the controller did not actually command.
The official AsyncChunkSampler saves the corrected normalized and padded execution window. When exactly d actions remain, it launches background sampling using the current observation and the saved tail. At the next boundary it passes the resulting candidates to the fast selector with a new observation.
The first chunk has no preceding committed tail, so it uses a boot path with delay zero. Later chunks have an existing queue and can overlap computation. The simple path assumes d does not exceed C and enough generated rows remain after the prefix.
If a background computation is unfinished at the boundary, the implementation waits for its future. Asynchrony is a scheduling strategy, not a proof that deadlines can never be missed. The widget reports an invalid window instead of silently truncating when H is too short.
This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.
def retain_window(candidate, delay, execute):
if not 0 <= delay <= execute:
raise ValueError("delay must fit in one execution window")
if len(candidate) < delay + execute:
raise ValueError("prediction is too short")
retained = candidate[delay:delay + execute]
return retained
candidate = [10, 11, 20, 21, 22, 23]
delay = 2
execute = 3
window = retain_window(candidate, delay, execute)
assert window == [20, 21, 22]
# Assume these are the commands after all edits.
executed_window = list(window)
next_prefix = executed_window[execute-delay:execute]
assert next_prefix == [21, 22]
# Empty prefix at episode boot.
boot = retain_window(candidate, 0, execute)
assert boot == [10, 11, 20]
# Half-open slices make the lengths explicit.
assert len(window) == execute
assert len(next_prefix) == delay
| Name | Meaning | Concrete referent |
|---|---|---|
| H | Predicted horizon | 16 in real experiments |
| Prefix | Committed actions during inference | Clamped input |
| Postfix | Actions following committed rows | Denoised output |
| Half-open slice | Includes start, excludes end | [2:5] has 3 rows |
Why is retaining candidate[d:d+C] alone insufficient to implement real-time chunking?
It removes already elapsed positions but does not tell the generator which actions are committed. Prefix-conditioned generation uses the actual committed commands, and the next prefix must come from corrected commands that were executed.
Which two commands form the next prefix in the worked example?
Evidence: PDF pp. 3–5; Eq. 4; Figure 2; official loop_utils.py. Paper PDF · Pinned implementation
A pretrained VLA can propose plausible ways to grasp a moving block. Asking a tiny policy to invent all that behavior from a few minutes of reward is a different problem. EXPO separates proposing a behavior from learning how to improve a proposal.
The small policy receives both an observation and a base action chunk. It does not begin with a blank action space. It predicts a residual that is added to a candidate, letting the large model supply the behavioral prior while the small model learns local, value-seeking adjustments.
The residual distribution is a tanh-squashed Gaussian. A Gaussian sample can be arbitrarily large; tanh maps it between minus one and one. Multiplication by an edit scale limits each normalized residual coordinate. This is a parameterization of a local search, not a physical safety guarantee.
The critic estimates discounted task return for an observation and a candidate. The editor is trained to produce residuals that make the combined action score better, with an entropy term that discourages collapsing its sampled distribution too early. The gradient changes the editor’s weights.
That value gradient does not run directly through the large VLA backbone. The base continues learning through a separate prefix- conditioned behavior-cloning objective. In the real experiments that supervised update uses demonstrations and successful online episodes, while the critic can learn from failed experience too.
Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.
| Quantity | Computed value |
|---|
| Element | Value or shape | Interpretation |
|---|---|---|
| Base 0 | 0.10 | Q = −0.1600 |
| Base 1 | 0.30 | Q = −0.0400 |
| Base 2 | 0.70 | Q = −0.0400 |
| Edited 1 | 0.38 | Q = −0.0144 |
| Edit dimension | 56 | 8 commands × 7 coordinates |
The widget uses a deliberately transparent teaching critic: score equals the negative squared distance from a desired normalized scalar action. Desired action is 0.5. Three base proposals are 0.1, 0.3 and 0.7; their scores are minus 0.16, minus 0.04 and minus 0.04.
Edits of 0.05, 0.08 and minus 0.05 produce 0.15, 0.38 and 0.65. The best of these is 0.38 with score minus 0.0144. This is not a neural network or a measured robot result. It exposes every arithmetic step that a learned critic would otherwise hide.
A narrow edit range keeps behavior near the prior. If the useful command is far away, the editor cannot reach it from a poor base candidate. More diverse base proposals can help cover different useful regions; they do not guarantee coverage.
A large bound allows bigger corrections but also allows a mistaken critic to favor more destructive changes. The paper’s normalized edit scales are task choices: 0.05 for kicking and 0.1 for the other tasks. Do not reinterpret them as centimeters or degrees.
This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.
from math import tanh, atanh
base = [0.1, 0.3, 0.7]
scale = 0.1
desired = 0.5
raw_samples = [atanh(0.5), atanh(0.8), atanh(-0.5)]
def teaching_q(action):
# Transparent teaching rule, not a trained critic.
error = action - desired
return -(error * error)
edits = []
edited = []
for action, raw in zip(base, raw_samples):
residual = scale * tanh(raw)
edits.append(residual)
edited.append(action + residual)
assert all(abs(e) <= scale for e in edits)
all_candidates = base + edited
scores = [teaching_q(a) for a in all_candidates]
choice = max(range(len(scores)), key=scores.__getitem__)
print(edits, scores, choice)
assert abs(all_candidates[choice] - 0.38) < 1e-12
| Name | Meaning | Concrete referent |
|---|---|---|
| β | Residual scale | 0.05 or 0.1 in reported tasks |
| z | Gaussian sample before tanh | Unbounded sample |
| Q | Estimated future task return | Learned, not guaranteed correct |
| α | Entropy temperature | Controls editor exploration tradeoff |
Why can a bounded editor fail even if its critic ranks every available candidate correctly?
The useful action may lie outside the residual range of every base candidate. Perfect ranking cannot select an action absent from the candidate set. A diverse and competent prior still matters.
What does the edit scale bound directly?
Evidence: PDF pp. 3, 5, 14; Equations 1 and 8; Table IV. Paper PDF · Pinned implementation
Two candidate grasps are ready, but the block has moved toward one and away from the other. Which observation should decide the final choice? The important computation happens after the expensive proposals are available: read the latest state, correct candidates, then rank them.
Start background generation when d commands remain in the current queue. The base receives that moment’s observation and the committed prefix. It proposes multiple full action chunks. While it works, the queue advances and the physical state changes.
At the next chunk boundary, retain the future C-command segment of each candidate. The lightweight editor receives a fresh observation and each segment. It produces one residual per candidate in the reported real rollout configuration. The critic evaluates originals and edited alternatives at the fresh observation.
With 32 base candidates and one edit for each, selection considers 64 alternatives. It is a union, not an average. A candidate can win without editing. Keeping originals avoids forcing every action through a correction that may be unnecessary or mistaken.
The toy below makes that visible. Move the desired action to a base candidate, then set a large incorrect edit. The base can still win. If you remove originals from the selection pool, that fallback disappears. This explains an algorithmic choice without claiming the fallback guarantees success.
Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.
| Quantity | Computed value |
|---|
| Element | Value or shape | Interpretation |
|---|---|---|
| Generation input | Old observation + prefix | Expensive asynchronous path |
| Edit input | Fresh observation + retained candidate | Small synchronous path |
| Rollout pool | 32 originals + 32 edited | 64 alternatives |
| Score aggregation | Minimum of two sampled target critics | Then argmax |
| Execution | One 8 × 7 chunk | No averaging across winners |
The detailed implementation has ten Q-networks. At selection, it samples two target networks, takes their minimum value for each candidate, and then chooses the candidate with the largest resulting score. Taking a minimum discourages accepting a proposal supported only by one optimistic estimate.
This does not form a formal uncertainty bound. Two critics can agree and both be wrong. The widget uses a fixed second illustrative estimate that is slightly lower than the first, so you can inspect the minimum without treating it as a calibrated uncertainty estimate.
Toggle fresh versus stale editing and scoring while keeping the proposals fixed. This isolates the intervention point: the fast path can respond to changes that happened after generation began. It is different from merely conditioning the base on the action prefix.
The paper compares against EXPO-FT with RTC as well as without RTC. The stronger comparison helps separate continuous asynchronous execution from the additional benefit of fresh-observation corrections. Do not present RTC as incapable of ever being combined with reinforcement learning.
This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.
base = [0.1, 0.3, 0.7]
current_target = 0.5
scale = 0.1
def bounded_delta(action, target):
return max(-scale, min(scale, target-action))
def q1(action):
return -(action-current_target)**2
def q2(action):
return q1(action) - 0.01 * abs(action)
edited = [a + bounded_delta(a, current_target) for a in base]
pool = base + edited
pair_scores = [(q1(a), q2(a)) for a in pool]
conservative = [min(pair) for pair in pair_scores]
selected = max(range(len(pool)), key=conservative.__getitem__)
for index, (action, scores) in enumerate(zip(pool, pair_scores)):
kind = "base" if index < len(base) else "edited"
print(index, kind, action, scores)
print("selected", selected, pool[selected])
# Handcrafted editor/critics demonstrate the selection rule.
# They are not the paper's trained networks.
| Name | Meaning | Concrete referent |
|---|---|---|
| N | Number of base candidates | 32 |
| Candidate union | Original and edited alternatives | 64 in real rollout |
| Target critic | Slowly updated value network | Subsample two of ten |
| Argmax | Choose one best-scoring index | Not a weighted average |
If an edited candidate has the highest score from one critic, must it be selected?
No. The actual selector first takes the minimum across the sampled critic pair for each candidate. Another candidate can have a larger minimum, including an original base candidate.
When does the fast path receive its new observation?
Evidence: PDF pp. 4–5, 14; Figure 2 and Equations 4–6; _jitted_fast_select. Paper PDF · Pinned implementation
At deployment the first few commands are already fixed, but ordinary flow training starts from noise everywhere. That mismatch is avoidable. Give training examples a clean action prefix and ask the network to predict only the remaining commands.
Let a demonstrated action be 5 and a sampled noise value be 2. In the official code’s convention, flow time zero means clean action and time one means noise. At time one quarter, the interpolated input is one quarter of 2 plus three quarters of 5, which gives 4.25.
The derivative with respect to this flow time is noise minus action, or minus 3. Sampling moves time from one toward zero, so an Euler update uses a negative time step. A negative step multiplied by minus 3 moves the sample toward 5. Flow time is not robot time.
For a chunk, every position gets its own time input. Prefix positions receive clean time zero and their ground-truth action values. Postfix positions receive the sampled noisy time. The loss mask is zero on the prefix and one afterward.
The network can attend to the clean prefix to infer a coherent continuation. It is not rewarded or penalized for a velocity prediction on a committed row. During sampling those rows remain clamped after every Euler step, so numerical integration cannot accidentally move them.
Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.
| Quantity | Computed value |
|---|
| Element | Value or shape | Interpretation |
|---|---|---|
| Rows 0–1 | Clean prefix | Loss weight 0 |
| Row 2 input | [4.25, 5.00] | Target velocity [−3, −4] |
| Row 3 input | [6.00, 6.75] | Target velocity [−4, −5] |
| Per-row squared errors | 1.00 and 0.50 | Mean over coordinates first |
| Masked mean loss | 0.75 | Prefix errors excluded |
Printed Equation 7 uses an interpolation that is clean at time one but retains a target written as noise minus action. That target is opposite to the derivative of the printed interpolation. The official code instead uses the internally consistent clean-at-zero convention explained here.
The core idea is unchanged: a clean prefix conditions a noisy postfix and the loss excludes committed rows. The lesson uses the released implementation’s convention for runnable arithmetic and labels the paper mismatch, rather than combining two incompatible time directions.
During initial supervised training, the prefix length is sampled per example over a range. This exposes the model to both boot and delayed execution. In reported real online training, prefix length is zero for first-chunk transitions and fixed to the deployment delay later.
The simulation procedure instead resamples prefix length during online base fine-tuning and includes all rollout data. Real-world base updates use successful episodes. Neither difference should be erased by one generic statement that “the base trains on all data”.
This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.
actions = [[1.,2.],[3.,4.],[5.,6.],[7.,8.]]
noise = [[0.,0.],[1.,1.],[2.,2.],[3.,3.]]
prefix = 2
u = 0.25
mask = [int(k >= prefix) for k in range(4)]
def interpolate(k, j):
if k < prefix:
return actions[k][j]
return u*noise[k][j] + (1-u)*actions[k][j]
x = [[interpolate(k,j) for j in range(2)] for k in range(4)]
target = [[noise[k][j]-actions[k][j] for j in range(2)] for k in range(4)]
predicted = [[999.,999.],[999.,999.],[-2.,-5.],[-4.,-4.]]
row_loss = []
for prediction, truth in zip(predicted, target):
row_loss.append(sum((p-y)**2 for p,y in zip(prediction,truth))/2)
loss = sum(w*e for w,e in zip(mask,row_loss))/max(sum(mask),1)
assert x[:prefix] == actions[:prefix]
assert abs(loss-0.75) < 1e-12
print(x, target, row_loss, loss)
# Huge errors in committed rows do not change the masked loss.
| Name | Meaning | Concrete referent |
|---|---|---|
| u | Flow interpolation time | Clean at zero in code |
| ε | Sampled noise matrix | Not sensor noise |
| m_d | Postfix loss mask | Zero on committed rows |
| v | Predicted flow velocity | Integrated from noisy time to clean time |
Why can a prediction of 999 on a committed row leave this example’s loss unchanged?
That row has loss weight zero and is clamped as clean conditioning input. The model must predict the uncommitted future; the loss is averaged over active postfix rows only.
In the code convention used here, which flow time denotes a clean prefix?
Evidence: PDF p. 5 Eq. 7; Appendix E pp. 14–16; official pi05.py lines 258–306. Paper PDF · Pinned implementation
A critic must learn from what the robot actually executed. If the selected action is an eight-command chunk, the next decision state arrives after eight commands. Treating that as an ordinary one- command transition confuses both reward timing and the next action’s information.
Start with the ordinary discounted return. Collect the first C rewards into one chunk reward. Each later reward in that chunk has one additional factor of the per-command discount. After those C commands, the remaining future value is multiplied by the discount raised to C.
The target therefore combines the discounted rewards actually observed during the chunk and a discounted estimate at the next chunk boundary. A continuation mask makes the bootstrap zero if the episode has terminated. The replay buffer must preserve where each reward and terminal flag occurred.
Use a small three-command example with per-command discount 0.9. For a nonterminal chunk with all rewards zero and next value 0.6, the target is 0.9 cubed times 0.6, or 0.4374. There is no reward term to add.
Now suppose success is detected on the third command and terminates the episode. The reward sequence is zero, zero, one. Its discounted sum is 0.81 and the next-value contribution is zero. These are different cases: the real success detector does not emit terminal success and also allow an ordinary future bootstrap.
Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.
| Quantity | Computed value |
|---|
| Element | Value or shape | Interpretation |
|---|---|---|
| Nonterminal rewards | [0,0,0] | Next Q = 0.6 |
| Nonterminal multiplier | 0.9³ = 0.729 | Target 0.4374 |
| Terminal rewards | [0,0,1] | Success ends episode |
| Terminal target | 0.81 | No next-value bootstrap |
| Actual window discount | 0.99⁸ ≈ 0.9227 | For ordinary task discount setting |
At the next boundary, the backup should emulate rollout selection. The base candidate is generated from the observation d steps before that boundary, conditioned on committed actions. Editing and action- value selection then use the next-boundary observation.
Keep current state, next state and delayed next-generation state as separate records. Substituting the next state into the slow base would give the target an information advantage unavailable during deployment. Substituting the stale state into the editor would remove the proposed correction mechanism.
Equation 9 writes a chunk-level target using a reward and one discount symbol. The released code explicitly accumulates discounted rewards across the chunk and uses the per-command discount raised to the execution length, together with a continuation mask.
You can treat a single discount symbol as chunk-level shorthand if it is defined that way. You cannot silently take a per-command value such as 0.99 and apply it only once to an eight-command transition. In that case the implemented multiplier is approximately 0.9227, not 0.99.
This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.
def chunk_target(rewards, masks, gamma, next_value):
total = 0.0
continuation = 1.0
for index, (reward, mask) in enumerate(zip(rewards, masks)):
total += gamma**index * continuation * reward
continuation *= mask
bootstrap = gamma**len(rewards) * continuation * next_value
return total + bootstrap
nonterminal = chunk_target(
rewards=[0,0,0],
masks=[1,1,1],
gamma=0.9,
next_value=0.6,
)
terminal = chunk_target(
rewards=[0,0,1],
masks=[1,1,0],
gamma=0.9,
next_value=100.0,
)
assert abs(nonterminal-0.4374) < 1e-12
assert abs(terminal-0.81) < 1e-12
print(nonterminal, terminal, 0.99**8)
| Name | Meaning | Concrete referent |
|---|---|---|
| γ | Per-command discount | 0.99 in ordinary real-task settings |
| C-step reward | Discounted rewards inside selected chunk | Computed from replay |
| Continuation | Whether future bootstrap is valid | Zero after terminal |
| Delayed next state | Snapshot used to generate next candidate | Earlier than next-boundary state |
Why is the slow base’s observation in the next-action backup earlier than the observation used by the next critic?
The backup reproduces the deployment information sequence. Slow generation starts before the boundary; the fast editor and selector see the boundary observation. Collapsing them into one timestamp changes the policy being evaluated.
A terminal success occurs on the third command in the worked example. What is the target?
Evidence: PDF p. 5 Eq. 9; Appendix E/F; replay_buffer.py lines 516–530; realtime_expo_ft.py line 1067. Paper PDF · Pinned implementation
Building a critic target repeatedly can be more expensive than choosing one real action. If each target samples thirty-two full flow trajectories, much of training compute is spent decoding candidates that will immediately be discarded. Can we reject some candidates before paying that cost?
A generative policy receives an observation and a random noise seed, then numerically denoises it into an action chunk. Under fixed model and observation inputs, the seed identifies a sampled proposal. A small network can learn to predict the action value associated with that seed.
The noise-ranking critic is not a physical world model and does not observe the future. It learns a supervised relationship between a seed, the relevant observations and a target action-value estimate. Its ranking can be wrong, especially while the action critic or base generator is changing.
During a Bellman backup, draw thirty-two raw noise seeds and score them cheaply. Keep the highest-scoring seed, decode it once, sample one edit, then let the target action critic select between that base and edited candidate. The reported backup therefore decodes one survivor rather than thirty-two.
The outer action-value comparison remains. Noise ranking chooses which base candidate to instantiate; it does not replace the editor or the final target critic. This sequence is useful precisely because a cheap approximate first stage is followed by a more direct action-space evaluation.
Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.
| Quantity | Computed value |
|---|
| Element | Value or shape | Interpretation |
|---|---|---|
| Backup seed pool | 32 | Cheap filter scores |
| Decoded survivors | 1 | Full VLA decode |
| Edited backup alternatives | 1 | Base + one edited action |
| Deployment alternatives | 64 | 32 originals + 32 edits |
| Filter objective | Squared error to stopped target Q | Separate supervised ranker |
Ten Euler steps for each of thirty-two candidates gives three hundred and twenty denoising sample-steps. Decoding one survivor gives ten. Those counts help explain the source of savings, but they are not measured wall-clock times or proof of a thirty-two-fold end- to-end speedup.
Batching, cached vision-language features, memory traffic, filtering overhead and the remaining updates affect actual performance. Figure 7 compares learning against training compute on four simulation tasks. Its curves support improved compute efficiency; they do not provide a general deployment latency multiplier.
The real rollout path still draws thirty-two base candidates and thirty-two edited candidates before choosing its action. The optional filtering mechanism changes the backup path, not that deployment pool. Conflating these paths removes an important part of the method.
Equation 11 writes regression toward the final selected action value. The pinned code instead uses the decoded, unedited base candidate’s target Q as the filter target, while separately selecting the best base-or-edited action for the Bellman backup. The detailed filter conditions on both the next/current observation features and the earlier observation used by the base. Equation 10 and Equation 11 use different state indices in compressed notation; the code keeps the two observation roles explicit. The widget uses simple named seed scores to make this bookkeeping inspectable.
This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.
seed_scores = [0.2, 0.9, 0.1, 0.5]
selected_seed = max(range(len(seed_scores)), key=seed_scores.__getitem__)
assert selected_seed == 1
# Only this seed would enter the expensive decoder.
def teaching_decode(index):
return [0.1, 0.3, 0.7, 0.9][index]
base = teaching_decode(selected_seed)
edited = base + 0.08
pool = [base, edited]
target_values = [0.6, 0.75] # specified teaching values
selected_action = max(range(2), key=target_values.__getitem__)
regression_target = target_values[0] # pinned code: unedited base
paper_winner_target = target_values[selected_action]
filter_loss = (seed_scores[selected_seed]-regression_target)**2
assert abs(filter_loss-0.09) < 1e-12
N = 32
steps = 10
print("without filtering", N*steps, "sample-steps")
print("with filtering", steps, "sample-steps")
print("filter squared error", filter_loss)
# No measured runtime or trained network is claimed here.
| Name | Meaning | Concrete referent |
|---|---|---|
| Q_filter | Noise-space ranking critic | Supervised from target action value |
| Seed | Initial noise matrix | H × 32 padded space |
| Stop-gradient | Treat target as a fixed number in this regression | No backprop through target |
| Sample-step | One denoising step for one candidate | Not elapsed milliseconds |
Why can filtering discard the ultimately best action even when it saves computation?
The cheap seed ranker is an approximation. If it assigns a low score to the seed that would decode to the best action, that action never reaches the outer action critic. The final critic can only choose among surviving candidates.
Which path uses the 32-seeds-to-1-decode filter in the paper?
Evidence: PDF p. 6 Equations 10–11; Appendix A; Figure 7 p. 13; Appendix E. Paper PDF · Pinned implementation
A rectangle labeled “policy” hides the engineering choices that make this work. Follow one observation through the real implementation: images become features, candidates become windows, windows become edit vectors, and one chosen vector becomes physical commands.
Most tasks use side and wrist RGB images resized to 224 by 224. The VLA processes its visual input with its pretrained encoder. The fast critic/editor path uses a separate residual visual encoder. Stacking the two views along channels gives a 224 by 224 by 6 tensor.
The detailed appendix describes preactivation basic residual blocks with stage depths three, four, six and three, GroupNorm and a 512-dimensional image representation. The main text’s ResNet-50 label does not match canonical bottleneck ResNet-50 details. Use the detailed configuration, rather than assuming a standard library architecture from the name.
A state embedding contributes 64 more features. Ball Balancing adds detector-derived plate center, ball position and velocity; Soccer Kicking adds keeper position and velocity. Those features feed the critic, editor and noise filter while the base VLA’s state input remains unchanged.
Ball Balancing also uses three exterior-camera frames instead of the ordinary exterior-plus-wrist pair. Its channel stack has nine channels. The paper removes the vertical proprioceptive dimension from that critic because monotonic drift might reveal episode time. These task-specific choices matter to what the learning system actually observes.
Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.
| Quantity | Computed value |
|---|
| Element | Value or shape | Interpretation |
|---|---|---|
| Normal visual stack | 224 × 224 × 6 | Side + wrist |
| Balancing stack | 224 × 224 × 9 | Three exterior frames |
| Feature widths | 512 image + 64 state | Fast critic/editor path |
| Noise space | 16 × 32 per seed | Padded model dimensions |
| Execution and edit | 8 × 7 = 56 | Physical command coordinates |
The π0.5 configuration has a sixteen-step horizon and padded action width thirty-two. The physical output has seven dimensions: translation, rotation and gripper. Padding is a model-interface choice, not thirty-two independent actuators.
After candidate generation and the delay slice, each execution window has eight rows and seven physical coordinates. Flattening gives fifty-six numbers for the Gaussian editor and action critic. One edited version of each of thirty-two base windows yields sixty- four candidate windows. Selection returns one eight-by-seven window, which is unnormalized before execution.
The base is initialized from a π0.5 configuration with LoRA adapters in the language and action-expert stack. Its non-LoRA language-stack parameters are frozen by the parameter filter. Parameters outside that frozen set, including vision and projection components, are described as trainable; permitted adapters also update.
The critic and editor are separate learned components. The editor shares the fast visual encoding, while direct value learning does not backpropagate through the full base action generator. A wrapper flag named freeze_pi05_encoder also controls cached inference behavior; it is not sufficient evidence that the entire VLA is frozen.
This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.
H = 16
C = 8
physical_dim = 7
padded_dim = 32
N = 32
noise_shape = (N, H, padded_dim)
retained_shape = (N, C, physical_dim)
edit_shape = (N, C*physical_dim)
selection_shape = (2*N, C, physical_dim)
assert edit_shape == (32, 56)
assert selection_shape == (64, 8, 7)
# Dynamic Picking permits xyz and gripper residuals.
coordinate_mask = [1,1,1,0,0,0,1]
flat_mask = coordinate_mask*C
raw_edit = [0.1]*(C*physical_dim)
masked = [e*m for e,m in zip(raw_edit,flat_mask)]
for command in range(C):
row = masked[command*7:(command+1)*7]
assert row[3:6] == [0,0,0]
print(noise_shape, retained_shape, edit_shape)
print(masked[:7])
# These are shape and mask checks, not neural inference.
| Name | Meaning | Concrete referent |
|---|---|---|
| Physical width | Translation 3 + rotation 3 + gripper 1 | 7 |
| Padded width | Model interface width | 32 |
| Image embedding | Fast visual representation | 512 |
| Privileged detector features | Task-derived measurements available to fast path | Not new physical sensors |
Why should the Ball Balancing observation configuration be visible in a results explanation?
Its fast path receives a three-frame exterior stack and detector- derived ball information. Those inputs help estimate motion. Omitting them makes the reported result sound like the same unaugmented two-image configuration used elsewhere.
How many physical coordinates are corrected in a complete eight-command window before task-specific masking?
Evidence: PDF Appendix D/E pp. 13–16; Tables III–IV; official model configuration and selector. Paper PDF · Pinned implementation
The robot can execute a queued command while the VLA prepares another chunk. That is one kind of concurrency. The learner can also accumulate experience and update parameters later. Those are separate clocks, and the ten-minute headline refers to only one of them.
The real experiments begin with task demonstrations and supervised prefix-conditioned fine-tuning of π0.5. The paper describes starting online learning once this policy has roughly thirty percent success or better. Pretraining and demonstrations are part of the starting resources, not counted as online robot minutes.
The replay buffer can be seeded with demonstrations, or demonstrations can be sampled as a fixed fraction of a critic batch. Soccer Kicking uses a fifty-percent prior-data fraction; the other reported task settings seed the buffer. These choices affect what the learner sees early in adaptation.
The critic learns to distinguish promising and unpromising chunks through reward-backed targets. Failed experience therefore contains useful value information. The editor’s training objective conditions on replay actions; those are not necessarily freshly sampled base proposals. At rollout it edits generated base candidates. The real- world base behavior-cloning update uses demonstrations and successful online episodes.
This division prevents a simplistic description that all networks imitate every action the agent takes. The base’s supervised signal and the editor’s value-seeking signal have different roles. In simulation, the base’s behavior-cloning loss is applied to all rollout and demonstration data, an explicitly different setting.
Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.
| Quantity | Computed value |
|---|
| Element | Value or shape | Interpretation |
|---|---|---|
| Dynamic Picking | K=25, C=8, d=3 | Table IV settings |
| Soccer Kicking | K=20, C=8, d=5 | 50% demonstration batches |
| Ball Balancing | K=30, C=8, d=5 | Demonstrations seeded |
| Object Passing | K=30, C=8, d=5 | Demonstrations seeded |
| One update call | 20 critic + 1 base + 1 edit + 1 temperature | Not 20 VLA updates |
Each reported real update call uses twenty critic gradient steps, followed by one base step, one editor step and one temperature step. Calls are accumulated according to a task-specific number of collected transitions and flushed at episode boundaries. Training begins after ten completed episodes.
The widget lets you select collected transitions and the transition interval. Its counts are a bookkeeping illustration of those settings. It does not simulate learning curves or estimate the minutes required on a GPU. The exact public example script also differs from Table IV in some task values, so the lesson labels table settings separately.
The learned temperature weights entropy in the edit-policy objective. It changes the preference for diverse residual samples relative to value-seeking. The target entropy is minus half the flattened edit dimension; with fifty-six coordinates its value is minus twenty-eight.
Entropy is not included in the released Bellman backup. It is an optimization term for the small residual policy, not an additional environment success reward. Keeping that distinction makes it possible to trace which loss changes which component.
This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.
def update_counts(collected, interval, episodes):
accumulated = collected // interval
active = episodes >= 10
calls = accumulated if active else 0
return {
"eligible_calls": calls,
"critic_steps": 20*calls,
"base_steps": calls,
"edit_steps": calls,
"temperature_steps": calls,
}
counts = update_counts(300, 30, episodes=10)
assert counts["critic_steps"] == 200
assert counts["base_steps"] == 10
assert update_counts(300, 30, episodes=9)["eligible_calls"] == 0
print(counts)
# Gradient work is flushed at episode boundaries.
# This calculator does not model actual GPU time.
# Demonstrations and supervised initialization happened earlier.
D = 8*7
print("target entropy", -D/2)
assert -D/2 == -28
| Name | Meaning | Concrete referent |
|---|---|---|
| K | Transitions per update call in Table IV | Different from sampling-count notation elsewhere |
| UTD | Update-to-data scheduling parameter | Interpret with the stated loop |
| Replay | Reused past experience | Includes successes and failures |
| Temperature | Learned entropy coefficient | Editor objective only |
Would doubling the number of critic updates necessarily double the robot data used?
No. Updates reuse replay data. Gradient compute and new environment interaction are different budgets. More updates can change learning behavior, but are not automatically new robot trials or guaranteed improvements.
Which data trains the real-world base behavior-cloning update?
Evidence: PDF Section IV-C; Appendix E pp. 15–16; Tables III–IV; Appendix C for simulation differences. Paper PDF · Pinned implementation
A method can have the strongest average without being the strict winner on every task. The Kinetix table is a useful place to practice reading a result more carefully than its headline. Every displayed number in this chapter comes from Table II.
The experiments use ten symbolic-state Kinetix environments and a pretrained state-based flow policy. No vision-language-action model is involved in this simulation setting. The test asks whether the proposed delayed generative-policy mechanism works across dynamic control problems.
The base is pretrained using one million transitions, then fine- tuned online for one hundred thousand environment steps per run. The delayed flow-policy methods incur four steps of inference delay and replan every four steps. RLPD’s small Gaussian actor runs without delay.
Reported RL scores average four random seeds and one hundred evaluation episodes per seed. The BC and RTC reference evaluations use five hundred and twelve episodes. These protocols should remain attached to the result display; they are not identical trial budgets.
The paper provides rounded per-task percentages and an average row, not all raw seed outcomes in the table. The interactive view therefore computes differences from those rounded values and does not manufacture confidence intervals or reconstruct precise training trajectories.
Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.
| Quantity | Computed value |
|---|
| Element | Value or shape | Interpretation |
|---|---|---|
| Car Launch | 98 | Highest reported score |
| Cartpole | 98 | EXPO-FT 99; BC no-delay 100 |
| Lunar Lander | 94 | Some baselines 95 |
| Chain Lander | 92 | DSRL with RTC 96 |
| Ten-task average | 96.2 | Table II rounded values |
Real-Time EXPO-FT averages 96.2 percent across the ten tasks. EXPO- FT with RTC averages 81.7 percent, and no-delay RLPD averages 81.4 percent. The difference from the stronger delayed EXPO-FT with RTC baseline is 14.5 percentage points in the reported average row.
These are results for the tested configurations, not a theorem that a delayed large-policy system always beats a small real-time actor. Actor structure, demonstrations, optimization schedules and task priors differ between some baseline families.
The abstract describes best performance in ten of ten environments, but the table does not support a strict-first-place interpretation. Cartpole shows 98 for the proposed method versus 99 for EXPO-FT and 100 for no-delay BC. Hard Lunar Lander shows 94 versus 95. Chain Lander shows 92 versus 96.
The table’s boldface criterion is within 95 percent of the best RL score on a task, excluding BC and RTC references. The proposed method meets that criterion in all ten. The lesson keeps both the strong average and the three non-winning rows visible, rather than rewriting the table to fit a slogan.
This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.
ours = [98,98,93,97,99,94,98,94,92,99]
expo_rtc = [64,96,47,85,87,90,79,83,94,92]
assert sum(ours)/10 == 96.2
assert sum(expo_rtc)/10 == 81.7
chain_ours = 92
chain_best_rl = 96
threshold = 0.95*chain_best_rl
near_best = chain_ours >= threshold
strict_best = chain_ours >= chain_best_rl
assert near_best and not strict_best
non_winners = {
"Cartpole": (98,99),
"Hard Lunar Lander": (94,95),
"Chain Lander": (92,96),
}
for task, (score,best) in non_winners.items():
print(task, score-best, "percentage points")
print("average", sum(ours)/len(ours))
# Table values are measured; this code only recomputes arithmetic.
# It does not run Kinetix or estimate learning curves.
| Name | Meaning | Concrete referent |
|---|---|---|
| Percentage point | Difference of two percentages | 96.2−81.7=14.5 |
| Seed | Independent randomized training run | Four for RL |
| Reference BC | Pretrained flow without online fine-tuning | 512 evaluation episodes |
| Near-best | At least 95% of best RL score | Not statistical equivalence |
Why would a bar chart saying “VLA wins ten robot tasks” be inaccurate here?
Kinetix uses a symbolic-state flow policy without a VLA; its environments are simulations. Also, Table II has three rows where the proposed method is below another reported score. The strong supported claim is its average and near-best RL criterion.
Which statement agrees with Table II?
Evidence: PDF Section V-B; Appendix B/C; Figure 3; Table II p. 13. Paper PDF · Pinned implementation
The real robot learns to catch an offered object, balance a ball, grasp a moving block and kick past a moving defender. These are compelling demonstrations. Their success counts become more informative when the starting policy, comparison method and data budget stay visible.
Dynamic Picking uses a block on a rotating plate and the original roughly 67-millisecond base latency. Ball Balancing, Object Passing and Soccer Kicking add 100 milliseconds to emulate a slower model or constrained compute, yielding approximately 167 milliseconds. Their scheduled lead is five commands; Picking uses three.
Each task has its own initial-state variation and success detector. Picking uses gripper state and end-effector height sustained for several steps. Balancing requires the ball to remain near the plate center across consecutive frames. Detector definitions turn a physical outcome into the sparse binary reward used for learning.
The initial SFT policy obtains 19, 8, 10 and 13 successes out of 30, respectively. That totals 50 successes out of 120 trials, or 41.67 percent. The proposed method obtains 30, 28, 30 and 28, totaling 116 out of 120, or 96.67 percent.
Rounded, that is the paper’s 42-to-97-percent headline. It describes these four task evaluations, not a general reliability rate for arbitrary robot work. Two tasks have two failures each; thirty successes in another task are finite observations, not proof that failure is impossible.
Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.
| Quantity | Computed value |
|---|
| Element | Value or shape | Interpretation |
|---|---|---|
| Dynamic Picking | 30/30 | EXPO-FT + RTC: 24/30 |
| Ball Balancing | 28/30 | EXPO-FT + RTC: 23/30 |
| Object Passing | 30/30 | EXPO-FT + RTC: 27/30 |
| Soccer Kicking | 28/30 | EXPO-FT + RTC: 26/30 |
| Aggregate | 116/120 | Stronger baseline: 100/120 |
SFT with RTC totals 72 of 120, or 60 percent. EXPO-FT with RTC totals 100 of 120, or 83.33 percent. Comparing the proposed method with the latter gives 16 additional successes in the displayed evaluations, or 13.33 percentage points.
The widget defaults to that stronger baseline. Switch to SFT to recover the headline calculation, but keep the baseline’s name attached to its value. A true number with the wrong comparison label can tell a misleading story.
Training is capped at ten minutes of online robot interaction per task, or stopped according to the comparison’s 30-of-30 evaluation rule. It excludes the pretrained foundation policy, demonstrations, initial supervised training, compute and resets. The Object Passing run also has a shorter reported transition budget than the roughly eighteen thousand steps listed for three other tasks.
No corrective human interventions are used during online policy rollouts. Some environment resets still require humans, and evaluation success is independently checked by a person. The authors explicitly identify reset labor and task-specific success detectors as limitations. Both qualifications belong in an honest account of rapid adaptation.
This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.
sft = [19,8,10,13]
sft_rtc = [22,12,22,16]
expo_rtc = [24,23,27,26]
ours = [30,28,30,28]
trials_per_task = 30
trials = len(ours)*trials_per_task
def percent(values):
return 100*sum(values)/trials
assert sum(sft) == 50
assert sum(sft_rtc) == 72
assert sum(expo_rtc) == 100
assert sum(ours) == 116
for name, values in [
("SFT",sft),
("SFT + RTC",sft_rtc),
("EXPO-FT + RTC",expo_rtc),
("Real-Time EXPO-FT",ours),
]:
print(name, sum(values), trials, percent(values))
print("gain in percentage points", percent(ours)-percent(expo_rtc))
# Reported trials, not new robot experiments.
| Name | Meaning | Concrete referent |
|---|---|---|
| Trial | One evaluated task attempt | 30 per method/task |
| Aggregate | Sum across four tested tasks | 120 displayed trials |
| Intervention | Corrective human action during rollout | Not the same as reset labor |
| Detector | Task-specific success signal | Sparse reward and termination |
Can “no intervention” be shortened to “no humans needed” in this study?
No. It means no corrective interventions during the online policy rollouts. Demonstrations, some resets and human verification of evaluation success remain part of the experimental process.
What does the ten-minute cap measure?
Evidence: PDF pp. 7–9; Figures 4–6; Table I; Appendix D/E; Discussion. Paper PDF · Pinned implementation
The block makes an abrupt movement just after a chunk has been selected. A fresh observation at selection cannot reveal a future surprise. Understanding the method means knowing where its feedback ends, what its candidate set cannot express, and which experiment would test the remaining gap.
After selection, the robot executes the chosen commands until the next chunk boundary. The editor’s latest observation is fresh at that boundary, not throughout the whole future window. Shorter execution windows could change the tradeoff, but the paper’s reported real setting is eight commands.
The edit range can also be too small to repair every base candidate. Increasing the range is not a free solution: a wrong value estimate can favor a large bad correction. The stress widget isolates these mechanisms using an explicitly constructed scalar target and queue, not invented real failure percentages.
The paper motivates the method through delayed observations and Markovian credit assignment. However, a proposal distribution can still depend on an old observation and committed actions. Images may also hide velocity, contact or occluded objects. A current image input alone is not a proof that an arbitrary physical task is fully observed.
For implementation, retain the action queue and relevant observation history as explicit state. For scientific claims, distinguish the paper’s motivation from a universal theorem about Markov restoration or safe control. The experimental evidence is about these tested policies and tasks.
Change one input, inspect the intermediate values, and explain the effect before changing another. Reset restores the worked reference case.
| Quantity | Computed value |
|---|
| Element | Value or shape | Interpretation |
|---|---|---|
| Insufficient candidate coverage | Correct action absent | Improve prior or explore alternatives |
| Wrong critic ranking | Correct candidate discarded | Measure value estimation and feedback |
| Late background inference | Boundary waits | Measure complete timing distribution |
| Within-chunk surprise | Observation becomes old again | Evaluate execution-horizon tradeoff |
| Detector/reset burden | Operational adaptation limit | Measure actual human and reward work |
RTC keeps chunks coherent while computation overlaps execution. Original EXPO-FT supplies value-guided action editing. Real-Time EXPO-FT applies that fast editing and selection using the latest observation after slow candidate generation. DSRL instead learns to choose noise for a generative policy.
The companion ARLI paper also works in latent noise space: its small policy intervenes before action-expert denoising and can use an intermediate observation plus committed actions. Its frozen-base and timing design differs from this paper’s successful-episode base updates and post-generation action edits. Shared concern about latency does not make the algorithms interchangeable.
A useful deployment study would log observation age, background inference completion, edit-and-select time, command deadlines and actual failure conditions. It would also count resets, reward- detector errors and compute spent per unit of new interaction. These are proposed measurements, not experiments reported here.
Read the paper’s result as a strong demonstration of one way to combine a capable prior with fast learned corrections. Preserve the exact baselines, task scope and implementation conventions. Then you can ask a sharper question than whether the system is simply “real time”: which information is available at each decision, and what can the policy still change?
This self-contained Python example uses the standard library. It implements the chapter’s arithmetic or mechanism, without model weights, robot hardware or external services.
bases = [0.1,0.2,0.3]
bound = 0.1
target = 0.8
reachable_max = max(bases)+bound
assert reachable_max < target
# Distinguish missing coverage from an incorrect ranking.
covered = [0.2,0.4,0.8]
wrong_scores = [0.9,0.7,0.1]
selected = max(range(len(covered)), key=wrong_scores.__getitem__)
assert covered[selected] != target
# A future change can invalidate a correct current choice.
selected_for_now = 0.8
future_target = 0.1
future_error = abs(selected_for_now-future_target)
print(reachable_max, covered[selected], future_error)
# An experiment should record these separately:
metrics = ["observation_age", "inference_deadline", "edit_latency",
"candidate_coverage", "critic_ranking", "success_detector_error"]
print(metrics)
# This is a failure decomposition, not a robot simulator.
| Name | Meaning | Concrete referent |
|---|---|---|
| Coverage | Which useful actions the candidate set contains | Distinct from ranking |
| Partial observation | Sensor input omits state information | History may still matter |
| Boundary | Moment a new chunk is selected | Not every command |
| Operational cost | Reset, sensing and reward work | Not captured by interaction minutes alone |
Why does adding a current observation not automatically prove the complete policy is memoryless in the physical state?
The candidate set can still depend on the earlier observation and committed action queue, and the observation may not reveal all physical state. Explicit temporal bookkeeping and empirical evaluation remain necessary.
Which is the most precise summary of the contribution?
Evidence: PDF Discussion p. 9; Equations 4–11; Appendix D/E; companion ARLI source audit. Paper PDF · Pinned implementation