The robot keeps moving while its large model thinks. ARLI gives a small learning policy the missing timeline: committed actions and a fresher observation before noise steers the next chunk.
A low-level controller does not wait for an abstract notion of intelligence. It needs the next command at the next scheduled instant. If a large model takes several control intervals to answer, execution must either pause or continue with commands that were decided earlier. These choices change the observations and transitions that an online learner experiences.
A chunk is a predicted sequence, not a promise to execute the whole sequence. The prediction horizon k tells us how much output exists. The execution horizon n tells us how much of it is followed before replanning. Delay d tells us how many control intervals pass during inference. Keeping three names prevents a common mistake: assuming that a model producing fifty commands must wait fifty steps before looking again.
Interactive teaching model and exact setting arithmetic. Synthetic trajectories are illustrative; no browser widget is a trained ARLI policy or a reported learning curve.
The timeline below is a clock calculation. It does not simulate a UR5e arm or reproduce a reported success rate. Change the delay while holding the execution horizon fixed. Synchronous operation spends time paused; asynchronous operation overlaps computation with existing commands. Both can use an observation that is older than the first command produced from it.
A faster GPU could reduce d, but that is not the contribution studied here. ARLI changes what the small learning controller observes and when it acts. This is useful even when computation time cannot be reduced enough to fit inside one control interval.
| Quantity | Meaning or value | Interpretation |
|---|---|---|
| k | Predicted chunk length | 50 commands in the real experiment |
| n | Executed horizon | 20 usable commands before replanning |
| d | Inference delay | 10 control intervals |
| f | Control frequency | 60 commands per second |
| Delta | Control period | 1000 / 60 = 16.667 ms |
def timing(hz, delay_steps, predicted, executed):
assert hz > 0
assert 0 <= delay_steps < predicted
assert 0 < executed <= predicted - delay_steps
period = 1.0 / hz
return {
"period_seconds": period,
"inference_seconds": delay_steps * period,
"predicted_seconds": predicted * period,
"executed_seconds": executed * period,
}
real = timing(60, 10, 50, 20)
assert abs(real["inference_seconds"] - 1/6) < 1e-12
print(real)
# Arithmetic from the reported schedule, not a latency benchmark.
A robot is holding a shoe above a narrow bag. In its last camera image, the opening is directly underneath. The robot asks its large model for the next movement. While the model computes, the robot and the bag keep moving. When the answer arrives, it describes a world that has already changed.
That delay creates two different problems. A synchronous controller can stop moving whenever it needs another answer. This preserves a simple decision boundary, but repeatedly interrupts the motion. An asynchronous controller starts computing before the old movement finishes. It keeps moving smoothly, but must plan from an earlier observation. Removing the pause has not removed the age of the information.
The paper studies how this affects reinforcement learning. A controller learns from the consequences of its decisions. If the information attached to a decision leaves out what happens before execution, very similar inputs can lead to very different outcomes. The learner may struggle even when its underlying robot policy already knows a useful skill.
The real experiments use sixty control steps per second. Ten delayed steps take about one hundred sixty seven milliseconds. The model predicts fifty actions, but executes twenty before replanning. Prediction length, execution horizon, and inference delay are three separate quantities.
A R L I stands for Asynchronous Reinforcement Learning with Intermediate Information. Its central question is concrete: what information can a small learning controller receive while the large model is still working? We will first locate that small controller, then follow the pending actions and the extra observation that let it make a better decision.
Keep the real command rate at sixty hertz. If delay falls from ten steps to five while the execution horizon stays twenty, which durations change? Decide before moving the delay control. Then compare the observation age with the duration of the twenty-command segment.
The inference interval falls from about 166.7 milliseconds to 83.3 milliseconds. The execution horizon still covers about 333.3 milliseconds. Overlap changes where computation happens in the schedule; it does not make a five-step-old observation current. The predicted fifty-command output continues to cover about 833.3 milliseconds in command-time units. These are conversions of configured counts, not measured GPU benchmark results.
The illustration assumes a stable integer delay that fits inside the execution horizon. A production scheduler must also handle jitter and missed deadlines. The paper studies latency-aware learning under specified delay settings; this little timeline does not establish a worst-case scheduling guarantee.
Source: ARLI v2, PDF pages 1–4, 8. See the source for original plots, uncertainty bands and implementation details.
Write the pretrained generator as G(observation, noise). A conventional rollout samples noise from a standard Gaussian. DSRL replaces that input distribution with the output of a small reinforcement learning actor. The generator converts the sampled latent input into a trajectory using its existing learned structure.
This intervention changes the distribution of generated actions without updating the generator's weights through the RL objective. Gradients train the small actor and critics. The real base policy was previously adapted to each task with LoRA on demonstrations; freezing during online RL does not mean the model was never trained on the task.
Interactive teaching model and exact setting arithmetic. Synthetic trajectories are illustrative; no browser widget is a trained ARLI policy or a reported learning curve.
The paper's implementation chooses a single-step noise vector and repeats it over the action chunk axis. If the illustrative vector is [0.2, -0.4], a three-step chunk receives three copies. A temporal generator can transform those identical latent rows into different action rows because position, conditioning and computation also matter.
The widget uses a fixed toy generator whose trajectory bends as one latent value changes. It is a mechanism illustration: there is no trained VLA inside the browser. The useful invariant is the intervention point. The learner chooses the input to the generator, not a correction pasted onto its output.
| Quantity | Meaning or value | Interpretation |
|---|---|---|
| actor input | Augmented observation | Includes timing-relevant context |
| single latent | [B, D] | One learned noise decision |
| repeated latent | [B, k, D] | Same vector copied over time |
| generated output | [B, k, D_model] | Changing trajectory after denoising |
| physical action | Robot-specific adapter | Padded latent dimensions need not equal joints |
single = [0.2, -0.4]
k = 3
noise = [single.copy() for _ in range(k)]
assert noise == [[0.2, -0.4], [0.2, -0.4], [0.2, -0.4]]
# A fixed illustrative temporal generator, not the paper's policy.
def toy_generator(noise_rows):
return [[(i + 1) * 0.1 + w[0], w[1] / (i + 1)]
for i, w in enumerate(noise_rows)]
actions = toy_generator(noise)
assert actions[0] != actions[1]
print("noise", noise)
print("actions", actions)
Suppose the robot's previously trained policy already knows several plausible ways to approach the bag. We want reinforcement learning to favor the useful movements without rebuilding that entire skill from scratch. A R L I builds on a method called diffusion steering through reinforcement learning, or D S R L.
A diffusion or flow policy starts its action generation from a noise vector. Guided by the observation, a learned generator turns that vector into a sequence of robot commands. Different initial noise can produce different coherent movements. D S R L trains a small actor to choose that initial noise using reward. The actor learns which starting points tend to produce useful behavior through the existing generator.
The large vision and language backbone stays frozen during this reinforcement learning stage. The action expert that converts noise into actions also stays frozen. The small actor and its value estimators are the parts being trained. The real base policy's earlier supervised adaptation is a separate stage.
In these experiments, the actor selects one latent vector and repeats it along the action chunk's time dimension. Repeating the noise does not mean repeating one robot command. The action expert still uses its learned temporal structure to generate a changing sequence of movements.
This is also different from residual control. A residual controller adds a correction to an action that already exists. Latent steering chooses an input before the action generator runs. It inherits the generator's useful structure, but its intervention must arrive before that generator needs the noise. That deadline will determine how fresh the actor's information can be.
Change the latent steering value from positive to negative. Predict which part of the illustrative path moves and which endpoints remain fixed. Then ask whether identical latent rows force identical action rows. Use the Python generator to check the second question independently of the canvas.
The sine-shaped bend reverses direction, while both endpoints remain fixed because the sine term is zero there. That endpoint property belongs to this chosen toy equation. In the Python example, each row also depends on its temporal position, so the generated rows differ despite repeated noise. The real generator has learned temporal structure and observation conditioning; it is not restricted to this sine family.
An arbitrary noise vector cannot guarantee an arbitrary desired robot action. The pretrained generator defines which behaviors latent steering can reach. Freezing it preserves an existing skill prior while also retaining limitations in that prior's behavioral support.
Source: ARLI v2, PDF pages 4, 15–17. See the source for original plots, uncertainty bands and implementation details.
Consider a position x that changes by adding each command. Two episodes begin with the same old position x = 0. In one, the pending queue is [+1,+1]. In the other, it is [-1,-1]. The first new correction begins after those commands, so the execution positions are +2 and -2.
To return to zero in one step, the required corrections are -2 and +2. If the actor sees only the old position, both experiences have the same input. It cannot choose both correct outputs from that input without additional information. A mixture is not the same as knowing which episode needs which action.
Interactive teaching model and exact setting arithmetic. Synthetic trajectories are illustrative; no browser widget is a trained ARLI policy or a reported learning curve.
This is a constructed counterexample to the sufficiency of a chosen state representation. It does not prove that every delayed process is impossible to learn. A representation that includes the relevant queue can distinguish these cases. In a stochastic or partially observed scene, even that augmented representation may still leave uncertainty.
The critic is affected too. A value target associated with one pending queue can differ from the target under another. Combining them under one stale input may obscure the return structure that the actor needs. The problem is causal bookkeeping before it is network capacity.
| Quantity | Meaning or value | Interpretation |
|---|---|---|
| Episode A | Old position 0 | Queue +1,+1 |
| Episode B | Old position 0 | Queue -1,-1 |
| Execution A | 0 + 1 + 1 = +2 | Required correction -2 |
| Execution B | 0 - 1 - 1 = -2 | Required correction +2 |
| Aliased input | Both old observations equal 0 | Different consequences remain hidden |
def execute(start, queue):
trace = [start]
for command in queue:
trace.append(trace[-1] + command)
return trace
for queue in ([1, 1], [-1, -1]):
states = execute(0, queue)
correction = -states[-1]
final = states[-1] + correction
assert final == 0
print(queue, states, correction)
# Without the queue, both actors receive the same stale scalar.
# This example isolates representation aliasing, not learning speed.
Imagine a much simpler robot that moves along a line. Its old observed position is zero. In the first episode, it has already committed to two movements to the right, each of size one. In the second episode, it has committed to two equally large movements to the left. At the instant a new correction can begin, the first robot will be at positive two and the second at negative two.
Both old observations say zero. But returning to the center requires opposite corrections. The first robot needs a correction of negative two. The second needs positive two. A learning policy that sees only the old observation cannot distinguish those cases. Averaging the corrections would produce zero, which solves neither one.
This illustrates a state aliasing problem. Different decision situations have been compressed into the same input. In a Markov decision model, the state should contain the information needed to describe how the next outcome depends on the chosen action. The physical environment may satisfy that description while the small controller's chosen input does not. An old image alone can leave out the action queue that is already shaping the future.
This line robot is a constructed teaching example, not a reported benchmark. Its simple dynamics expose bookkeeping that is harder with a moving gripper and a soft bag.
Notice that the missing information is not secret. The controller has already selected those pending movements. It can read them from its queue. The first modification in A R L I is therefore to give that queue to the learner.
Set each pending command to zero. What happens to the contrast between the two episodes? Next restore a nonzero queue and move the target away from zero. Decide whether moving the shared target removes the need to distinguish the queues.
With zero commands, both toy execution positions equal the old position and require the same correction. With a nonzero command magnitude, the two positions remain separated by four times that magnitude. Moving the target adds the same offset to both corrections, so their difference remains. For command magnitude one and target one, the corrections are minus one and plus three. The queue ambiguity survives the new target.
This construction demonstrates that the stale scalar alone is insufficient in this example. It is not a sample-complexity theorem or proof that every observation-only controller must fail in every environment. A richer history could also encode information about the pending queue.
Source: ARLI v2, PDF pages 2–5. See the source for original plots, uncertainty bands and implementation details.
The missing queue is available at inference time. The controller already chose the commands that will be sent while the next chunk is computed. ARLI adds this intermediate action sequence to the small policy's observation, so the learner can associate different pending motion with different useful latent choices.
Use a half-open slice: a[t-d:t] includes actions at t-d through t-1 and excludes the first new action at t. This convention removes an off-by-one ambiguity. The printed paper has an abbreviated action index in one later tuple, but the preceding definition and the main diagram identify the entire intermediate sequence.
Interactive teaching model and exact setting arithmetic. Synthetic trajectories are illustrative; no browser widget is a trained ARLI policy or a reported learning curve.
Flattening the queue preserves its ordered coordinates when positions in the flattened vector correspond consistently to time and action dimension. It does not mean averaging commands into a bag of values. Reversing the order can change the physical outcome in nonlinear dynamics, so order must remain recoverable by the network.
Our scalar widget adds a disturbance after part of the pending queue. The queue-aware forecast knows the planned commands but cannot anticipate that external event. This deliberately exposes the next limitation: more complete old information is not a substitute for a newer observation.
| Quantity | Meaning or value | Interpretation |
|---|---|---|
| slice start | t-d | First action during inference |
| slice stop | t, excluded | First new action is not pending |
| queue shape | [d, D_action] | Time and coordinate axes |
| flattened shape | [d × D_action] | Order must be preserved |
| prediction limit | Unknown disturbances | Commands alone cannot reveal them |
old = 0.0
pending = [[1.0, 0.0], [1.0, 0.5]]
flat = [coordinate for command in pending for coordinate in command]
assert flat == [1.0, 0.0, 1.0, 0.5]
# One-dimensional teaching dynamics.
forecast = old + sum(command[0] for command in pending)
disturbance = -0.75
actual = forecast + disturbance
print("forecast", forecast) # 2.0
print("actual", actual) # 1.25
print("unexplained", actual - forecast)
# Flattening adds information; it does not observe the disturbance.
Return to the two line robots. This time, the actor receives both the old position and the pending action sequence. It can distinguish zero followed by two rightward movements from zero followed by two leftward movements. Those different inputs can now support different corrections.
A R L I applies this idea to the real inference window. The small actor receives the actions from the previous chunk that will be executed while the large model computes the next chunk. These commands are already committed. Including them does not let the actor change the past or replace commands after they have been sent to the robot.
The implementation described in the paper flattens the intermediate action array and joins it to the other features before the small network's final layers. The actor learns how those commands affect the circumstances in which its noise will be used, without a separate world model.
There is an important distinction from real time chunking, or R T C. R T C also uses committed actions, but primarily to encourage continuity between the old and new motion. Giving the same information to a reward trained actor lets it learn how the pending motion affects task success. Continuity and reward directed adaptation are related goals, but they are not the same computation.
Now break our simple example. Keep the pending commands unchanged, but imagine someone moving the target while inference is running. The queue can explain the robot's planned motion. It cannot reveal a disturbance that was absent from the old observation. In a real scene, objects can slip, contact can change, and another object can enter the path. To react to those events, the learner needs genuinely newer information, not only a better interpretation of the old frame.
Hold the pending commands fixed and sweep the disturbance from negative to positive. Decide whether the queue forecast should move. Then consider two different command orders under a nonlinear transition rule: would flattening their sum preserve enough information?
The queue forecast stays fixed because its known inputs have not changed. The actual position moves with the disturbance, and the forecast error equals that disturbance under the additive toy dynamics. Here addition makes order irrelevant to the final scalar, but that convenience does not extend to general robot dynamics. Applying a rotation then a translation can differ from translating then rotating. A consistent flattening retains every ordered coordinate; a sum generally discards that order.
The paper supplies the intermediate action sequence to the small learner. It does not require a separately trained forward dynamics model that predicts every intervening state perfectly. Our forecast simply makes the information content visible in an exactly solvable case.
Source: ARLI v2, PDF pages 4–6, 16. See the source for original plots, uncertainty bands and implementation details.
Figure 2 places the useful intervention inside the inference window. The vision-language backbone starts from the old observation. Its output features feed an action expert. The latent noise is needed by the expert, so the small actor can run later than the start of the backbone.
Define t as the scheduled first new action. The old observation arrives at t-d. The intermediate observation arrives at t-d_rl, where d_rl includes both the small actor and the action expert computation still required. In the real preset d is 10 and d_rl is 7. Relative to old observation time zero, the intermediate snapshot is at step 3 and the new action at step 10.
Interactive teaching model and exact setting arithmetic. Synthetic trajectories are illustrative; no browser widget is a trained ARLI policy or a reported learning curve.
At 60 Hz, step 3 is 50 ms and step 10 is 166.667 ms. The remaining blind interval is 7 steps, or 116.667 ms. That is a three-step improvement in observation age. It is not a claim that the actor itself takes seven steps, and it does not equal the separate approximate timing graphic on page 1.
The showcase reconstructs the causal structure of Figures 2 and 3. It uses a scalar target disturbance so that success or failure can be calculated exactly. The actor's displayed correction is an ideal toy rule, not a trained ARLI policy. Toggle the action queue and mid-observation inputs separately to see what each reveals. A disturbance after the late snapshot remains invisible to this choice.
| Quantity | Meaning or value | Interpretation |
|---|---|---|
| old snapshot | Step 0 | 0 ms in the local clock |
| mid snapshot | Step d-d_rl = 3 | 50 ms at 60 Hz |
| first new command | Step d = 10 | 166.667 ms |
| remaining blindness | d_rl = 7 steps | 116.667 ms |
| VLM features | Computed from old observation | Not recomputed from the mid snapshot |
def snapshot_times(hz, delay, remaining):
assert 0 <= remaining <= delay
scale = 1000.0 / hz
return 0.0, (delay-remaining)*scale, delay*scale
old, mid, handoff = snapshot_times(60, 10, 7)
assert mid == 50.0
assert abs(handoff - 166.6666666667) < 1e-7
def can_see_event(event_step, delay=10, remaining=7):
# Explicit convention: an event at the snapshot tick is visible.
return event_step <= delay-remaining
assert can_see_event(2)
assert not can_see_event(4)
The useful opportunity is inside the large model's computation. Many vision, language, and action models have a vision and language backbone followed by a smaller action expert. The backbone processes the image and instruction. The expert uses the resulting features and the noise input to generate actions.
Noise is needed when the action expert starts. A R L I therefore runs its small actor later, with another observation, the original observation, and committed actions. This adds information without a second expensive backbone pass.
Be precise about what updates. The intermediate image enters the small actor. It does not replace the earlier features already produced by the frozen vision and language backbone. The actor uses the new information to choose noise that steers the existing action expert.
Here is the real setup expressed on one clock. Start the large inference call at zero milliseconds. The first new action is scheduled about one hundred sixty seven milliseconds later. The remaining delay for the small actor and action expert together is seven control steps, or about one hundred seventeen milliseconds. That places the extra observation fifty milliseconds after the first one. It is fresher, but still not current when execution begins.
Move a disturbance to just before that extra observation in the interactive timeline. Move the disturbance just after the observation, and this noise choice cannot respond to it. The method has shortened an information gap, not abolished causality. Its benefit depends on which events become visible before the intervention deadline and whether the frozen generator can express a useful response.
Use delay ten and remaining delay seven. Compare a target change at step two with one at step four while both information inputs are enabled. Next disable the intermediate snapshot, leaving the committed queue enabled. Predict which change becomes invisible to the actor.
The intermediate snapshot is at step three. A step-two event is visible there; a step-four event happens afterward and remains unseen. With only the old step-zero snapshot, both events are invisible. The queue still predicts the handoff position under these deterministic commands. An event exactly at step three is visible by the explicit convention in this widget; changing that convention would change this boundary case, so timestamps must be documented.
The toy applies an ideal scalar correction instantly at handoff to expose observation error. A real action expert produces a constrained trajectory, and the robot cannot necessarily reach the target immediately. Zero toy error is a bookkeeping result, not an ARLI success-rate prediction.
Source: ARLI v2, PDF pages 4–6, 8. See the source for original plots, uncertainty bands and implementation details.
A newly generated trajectory can conflict with the motion that is already underway. RTC addresses that boundary by conditioning generation on the committed prefix and applying inpainting guidance. ARLI can select the initial noise while retaining this continuity mechanism.
Two uses of intermediate actions must remain distinct. RTC uses them to constrain the relationship between neighboring chunks. ARLI supplies them to a learned actor and critic so reward can determine which latent choices work in the resulting state. A continuity constraint is not a substitute for the actor's decision context.
Interactive teaching model and exact setting arithmetic. Synthetic trajectories are illustrative; no browser widget is a trained ARLI policy or a reported learning curve.
The paper's control delay is measured in robot action intervals. Denoising has its own numerical integration steps. The real configuration lists ten control steps of delay and ten denoising iterations, but those happen to be equal numbers with different meanings. Kinetix uses a different pair. Never use a control-time index as a flow-time integration index.
The browser illustration blends a continuation with a committed trajectory using a toy soft mask. It demonstrates the boundary tradeoff, not the exact RTC guidance implementation. Stronger agreement reduces a boundary jump but can constrain a sudden correction. The authors explicitly note that RTC may hurt reactivity by forcing older actions.
| Quantity | Meaning or value | Interpretation |
|---|---|---|
| predicted indices | 0 through 49 | 50 entries exist |
| elapsed prefix | 0 through 9 | Already covered by old execution |
| executed suffix | 10 through 29 | 20 actions followed |
| unused predicted tail | 30 through 49 | May be replaced by later replanning |
| solver iterations | 10 in real setup | Different axis from command indices |
def usable_commands(chunk, delay, execute):
assert 0 <= delay < len(chunk)
assert 0 < execute <= len(chunk)-delay
return chunk[delay:delay+execute]
chunk = list(range(50))
used = usable_commands(chunk, 10, 20)
assert used == list(range(10, 30))
assert len(used) == 20
# A toy interpolation, not the RTC solver's guidance equation.
old, proposed, weight = 0.2, 0.9, 0.75
blended = weight*old + (1-weight)*proposed
assert abs(blended-0.375) < 1e-12
print(used, blended)
A useful correction can still produce an awkward handoff between chunks. The old plan moves the gripper in one direction. A newly generated plan begins somewhere incompatible. Abruptly switching between them can create a jump in the commanded motion.
Real time chunking addresses this continuity problem. During generation, it encourages the beginning of the new chunk to agree with the actions that are already committed. A softer overlap region can connect that fixed prefix to a freely generated continuation. A R L I can combine this guidance with the noise selected by its learned actor.
Keep two clocks separate. The prefix consists of robot commands indexed by physical control time. The action expert also takes numerical steps while it turns noise into a trajectory. Ten committed robot commands are not necessarily ten denoising iterations. The paper lists inference delay and denoising step count as different settings. Confusing them produces the wrong mask and the wrong execution schedule.
When a chunk is ready, the controller skips the part whose physical time has already passed. It does not execute those prefix commands again. In the real preset, the prediction contains fifty actions. Ten correspond to the elapsed prefix, and the controller executes twenty actions from the usable continuation before replanning. The remaining predicted tail is not automatically executed.
Continuity guidance also creates a tradeoff. If old actions are strongly enforced, the new policy has less freedom to respond immediately. This can improve smoothness while limiting reactivity. The authors explicitly identify that limitation. Their experiments often benefit from R T C, but that does not make it universally helpful under every disturbance.
The learned actor and continuity guidance have distinct roles. Next, follow how their combined experience becomes training data.
Set the toy continuity weight to zero, then to one. Calculate the first continuation command and its jump from the old boundary command in both cases. Separately check whether an execution request for forty-one commands is valid with prediction length fifty and delay ten.
At weight zero, the first continuation equals the proposal, 0.9, with a boundary jump of 0.7. At weight one, it equals the old command, 0.2, with zero boundary jump. An execution request for forty-one commands is invalid: only forty predicted entries remain after skipping ten. This indexing constraint is separate from how strongly generation is guided toward the committed prefix.
The interpolation is deliberately transparent and is not RTC's exact inpainting objective. A lower jump in this scalar illustration does not prove lower jerk for a physical robot. Robot dynamics, action parameterization, and the rest of the generated trajectory all matter.
Source: ARLI v2, PDF pages 4, 6, 9, 17. See the source for original plots, uncertainty bands and implementation details.
The actor decides once per noise sample, while the robot executes many low-level commands. The replay record should pair that latent action with the information available when it was selected and the consequences produced by the associated executed horizon. Assigning a fresh latent action to every servo step would misdescribe the behavior policy.
The paper describes a CNN for pixel observations followed by an MLP. Old and intermediate images concatenate along channels; joint states concatenate separately. For three RGB cameras and two snapshots, a channels-last teaching tensor has shape [B,64,64,18]. This is derived arithmetic, not a claim about the exact public runtime layout. The paper does not specify every convolution stride or the robot's unpadded physical action dimension.
Interactive teaching model and exact setting arithmetic. Synthetic trajectories are illustrative; no browser widget is a trained ARLI policy or a reported learning curve.
The latent actor is trained with SAC. A critic estimates return from an augmented observation and latent choice. The actor trades value against an entropy term. Ten critics appear in the experiment configuration; an ensemble does not by itself establish calibrated uncertainty. The exact reduction and entropy-backup behavior need an actual implementation configuration before being stated as ARLI facts.
For reference, the official upstream DSRL code stores the per-step discount raised to the query interval, then uses a terminal mask in the critic target. Its synchronous loop was inspected at a fixed commit. ARLI modifies that code for asynchronous execution; no author-released ARLI training repository was linked by the project page at the research date. The example below is explicit arithmetic, not a full reproduction.
| Quantity | Meaning or value | Interpretation |
|---|---|---|
| camera inputs | 3 views × 3 channels × 2 times | 18 channels in teaching convention |
| CNN | Four layers, 32 output channels each | Kernel and stride not specified here |
| MLP | 3 layers, width 128 for real/Aloha | Width 256 for Kinetix |
| critics | 10 | Separate value estimators |
| warmup | 1,200 latent transitions; 24,000 updates | Before real online RL |
def target(reward, gamma, elapsed_steps, next_q, terminal):
discount = gamma ** elapsed_steps
mask = 0.0 if terminal else 1.0
return reward + discount * mask * next_q
live = target(-1.0, 0.999, 20, 4.0, False)
done = target(-1.0, 0.999, 20, 4.0, True)
assert abs(live - 2.920755459) < 1e-8
assert done == -1.0
views, color_channels, snapshots = 3, 3, 2
channels = views * color_channels * snapshots
assert channels == 18
# Reward values here are illustrative; not an ARLI task specification.
The robot sends many low level commands during one action chunk, but its noise actor makes one decision for that chunk. Training data must preserve this distinction. Otherwise the learner could be credited with decisions it never made, or paired with observations that belong to the wrong point in the timeline.
The paper treats sampling and executing a chunk as one transition in a latent decision process. The action stored for the small controller is its chosen noise. The environment still evolves at the faster physical control rate. When the simulation plots show environment steps, those are the original control steps, not the number of entries in the latent replay buffer.
The observation pipeline is concrete. The real setup has a front camera and one wrist camera on each arm. The small network uses images resized to sixty four pixels square. Three color views contribute nine channels for one time point. Joining old and intermediate views gives eighteen channels under a channels last description. Joint measurements are joined separately. Pending actions are flattened and added before the final multilayer network.
An actor proposes latent noise. A group of ten critics estimates its value. The actor is rewarded for noise that receives higher value while retaining an entropy incentive for exploration. The critics learn from reward and a discounted estimate of what follows. Terminal masks prevent a finished episode from borrowing value from a nonexistent future.
Replay discount must match the decision interval. Upstream D S R L raises its per step discount to the query horizon. That synchronous reference is not the unreleased A R L I integration. The theorem also uses a different stride from the real execution horizon. Keep these settings distinct.
Switch the same transition between nonterminal and terminal. Why should the target stop depending on the next-state value in the terminal case? Then increase the elapsed steps while keeping the displayed reward and next-state value fixed.
The terminal mask removes the bootstrap term because the recorded episode has no future continuation. The target therefore equals minus one in this example. For a nonterminal transition, increasing elapsed steps decreases the positive future term because the per-step discount is below one. Keeping the reward fixed isolates discount arithmetic; a real longer transition may also accumulate a different reward. The temporal scope of a replay record must agree with both quantities.
The displayed backup follows inspected upstream discount and mask conventions, with explicitly illustrative reward values. Without the unreleased asynchronous integration, this calculation cannot establish the exact storage layout, reward aggregation, or entropy-backup configuration used in every ARLI experiment.
Source: ARLI v2, PDF pages 15–18. See the source for original plots, uncertainty bands and implementation details.
The appendix studies action-chunk values under a delayed observation and a committed prefix. A candidate continuation is selected by maximizing the learned chunk Q function while leaving that prefix fixed. The reference value uses an ideal action-chunking policy, not an omnipotent zero-latency servo controller.
The delayed oracle gap omega_d bounds the absolute discrepancy between an oracle chunk value and a delayed backup. Both the main definition and Appendix Definition 2 have absolute-value bars visible in the PDF. The appendix prints an undefined h in one discount exponent; the surrounding recurrence consistently uses k. We use k and state that normalization rather than silently inventing another horizon.
Interactive teaching model and exact setting arithmetic. Synthetic trajectories are illustrative; no browser widget is a trained ARLI policy or a reported learning curve.
To turn a local discrepancy into a global bound, define beta = gamma^(k-d). Insert and subtract the oracle continuation value inside the delayed backup. The triangle inequality bounds the difference by omega_d plus beta times the largest future value difference. Taking a supremum produces E ≤ omega_d + beta E. Subtract beta E and divide by 1-beta.
The conclusion depends on support coverage. The dataset must cover optimal chunks and optimal completions for in-support prefixes; its trajectories must obey the physical transition law and the required open-loop consistency. The theorem is not a finite-sample guarantee for a neural SAC implementation. Its stride k-d also differs from the separately chosen execution horizon n in several experiments. The real preset has k-d=40 but n=20.
| Quantity | Meaning or value | Interpretation |
|---|---|---|
| gamma | Per-step discount | 0 ≤ gamma < 1 |
| k | Theoretical full chunk length | Must exceed d |
| d | Delayed prefix length | Already committed |
| beta | gamma^(k-d) | Discount over theoretical backup stride |
| omega_d | Assumed local oracle mismatch bound | Not estimated from reported success curves |
| E | Largest in-support value error | Not a robot failure probability |
def delayed_bound(gamma, k, d, omega):
assert 0 <= gamma < 1
assert 0 <= d < k
assert omega >= 0
stride = k-d
beta = gamma**stride
denominator = 1-beta
return omega/denominator
for delay in (1, 4, 7):
print(delay, delayed_bound(.99, 8, delay, .02))
# omega=.02 is invented for arithmetic, not estimated from ARLI.
# This does not certify finite data, neural approximation, or deployment.
# Runtime n need not equal the theorem's k-d stride.
Suppose a critic could evaluate every relevant action chunk perfectly. There would still be a cost to deciding from delayed information. An unexpected event may make the best continuation different from the one chosen earlier. The theory asks how that cost accumulates when each delayed decision has a bounded mismatch from an ideal reference.
The policy in the appendix preserves the actions already committed and chooses the best remaining completion according to a chunk value function. Its next backup advances by the number of actions left after the delayed prefix. That specific interval determines how much the next error is discounted.
Imagine the largest value error at one decision. It is bounded by a local mismatch plus a discounted copy of the future error. Repeating the argument gives a geometric series. If the discount is below one and the completion advances time, the series has a finite sum. A larger local mismatch produces a larger bound. A discount closer to one accumulates more of the future mismatch.
The assumptions do substantial work. The data must obey the environment's transition dynamics. It must have the required open loop consistency. Optimal chunks must lie within the data's support, and so must optimal continuations for the committed prefixes being considered. A critic cannot be declared correct for an unseen completion simply because its neural network produces a confident number.
The paper's theoretical schedule is not identical to every practical setting. Fifty predicted actions minus ten delayed actions leaves forty, while the real controller executes twenty before replanning. That is why the bound should not be presented as a numerical certificate for the deployed actor. It also does not account for every finite data or optimization error. Its useful message is conditional: if delay creates only a controlled oracle mismatch and the learning assumptions hold, that mismatch need not grow without bound.
With gamma 0.99 and prediction length eight, compare delay one with delay seven while holding the assumed mismatch at 0.02. Which case has the smaller denominator? Then set the mismatch to zero and explain what the resulting zero bound assumes.
Delay one leaves a stride of seven and gives a bound about 0.294401. Delay seven leaves a stride of one and gives a bound of two. The second denominator is smaller. A zero assumed oracle mismatch makes the displayed bound zero, but only within the theorem's coverage and dynamics assumptions. It does not show that a trained finite neural critic has zero mismatch. The input is an assumed property, not a measured training diagnostic.
Changing delay while holding mismatch fixed is a controlled algebraic comparison. In an actual task, the mismatch itself may depend on delay and available information. The graph should not be interpreted as a measured degradation curve or a probability of robot failure.
Source: ARLI v2, PDF pages 6, 23–25. See the source for original plots, uncertainty bands and implementation details.
The main simulation figure compares learning curves across three Kinetix tasks and AlohaTransferCube. Kinetix gives the small learner a symbolic vector of scene properties. Aloha gives it pixels and joint state and uses a 3.3B-parameter pi-zero model. Mixing the two observation settings into one generic VLA benchmark would misstate the evidence.
In the main Kinetix runs, a chunk has eight actions, four are executed, total delay is four steps and remaining actor-plus-expert delay is one. Aloha uses fifty predicted actions, twenty-five executed, delay twenty and remaining delay ten. The experiment inspector below compares these exact settings. It does not interpolate unreported success numbers between plotted conditions.
Interactive teaching model and exact setting arithmetic. Synthetic trajectories are illustrative; no browser widget is a trained ARLI policy or a reported learning curve.
Figure 4 generally favors ARLI over naive asynchronous DSRL, with RTC often improving efficiency and reliability. The car-launch result is a useful limitation on a dramatic headline: several methods perform strongly there. The paper demonstrates important failure regimes, not a theorem that standard RL always fails whenever latency is positive.
Simulation curves average three seeds and shade one standard error. Each evaluation averages fifty rollouts. A band across training seeds is not identical to a binomial confidence interval on fifty evaluation outcomes. Read the method, uncertainty measure, training-step unit and base checkpoint together before comparing two endpoints.
| Quantity | Meaning or value | Interpretation |
|---|---|---|
| Kinetix base | BC flow checkpoint, epoch 5 | Symbolic observations |
| Aloha base | pi-zero Aloha simulation checkpoint | Pixels and joint state |
| Training seeds | 3 | Mean and one standard error |
| Evaluation rollouts | 50 per evaluation | Not 50 independent training runs |
| Residual baseline | Current state/action with per-step residual | Different intervention and update settings |
from statistics import mean, stdev
from math import sqrt
# Invented values to teach the reported uncertainty convention.
seed_scores = [0.7, 0.8, 0.9]
center = mean(seed_scores)
sem = stdev(seed_scores)/sqrt(len(seed_scores))
assert abs(center-0.8) < 1e-12
assert abs(sem-0.0577350269) < 1e-9
print(center, sem)
# Fifty evaluation rollouts are a different sampling level.
# Do not combine this SEM with a fictitious exact digitized curve.
The simulation study tests whether these information changes help reinforcement learning recover under delay. It includes three dynamic tasks from a physics benchmark and a simulated bimanual cube transfer task. The physics tasks use compact symbolic observations and a previously trained flow policy. The cube transfer experiment uses images and joint state with a much larger pie zero policy.
The main comparisons include naive asynchronous D S R L, the same method with real time chunking, A R L I with and without real time chunking, and residual reinforcement learning. These methods intervene at different places and receive different information. A fair comparison must keep those differences visible.
In the reported curves, A R L I generally learns more effectively under delay than naive latent steering. Adding continuity guidance helps in several cases. The car launch task is an important exception: several methods perform strongly there. The evidence does not say ordinary reinforcement learning fails in every environment that has latency.
The simulation curves average three random seeds. Their shaded uncertainty corresponds to one standard error. Evaluations use fifty rollouts. Those are different sources of counting: training seeds, evaluation attempts, and original environment steps along the horizontal axis. A smooth line is not an infinite sample, and its final plotted point should not be read as an exact universal performance level.
Our interactive comparison uses the reported experimental settings and qualitative curve evidence. It does not manufacture missing raw measurements or display a synthetic learning curve as if it came from the authors. The next question is whether the same idea helps on physical hardware, where a bag can bend and a connector can catch on its pins.
Switch between the three experiment presets and record predicted length, executed horizon, total delay, and remaining delay. Which configuration uses symbolic observations? Which two use pretrained vision-language-action policies? Explain why equal action counts need not mean equal wall-clock durations.
Kinetix uses symbolic state and a behavior-cloned flow policy, with the displayed main-run counts eight, four, four, and one. AlohaTransferCube uses a vision-based pie zero policy with counts fifty, twenty-five, twenty, and ten. The real setup uses task-adapted pie zero point five and counts fifty, twenty, ten, and seven. Only a control frequency converts action counts into seconds. The real setup's sixty-hertz conversion must not be silently applied to every simulation.
This comparison faithfully reconstructs configuration values. It does not recreate the original learning curves from guessed coordinates. Consult the source's plotted means and uncertainty bands to assess training trends, and retain the distinction between simulation evaluations and real training windows.
Source: ARLI v2, PDF pages 7, 16–18. See the source for original plots, uncertainty bands and implementation details.
The real platform is a bimanual UR5e cell with one front camera and two wrist cameras. The tasks involve connector assembly, placing a shoe into a narrow bag, and lifting/reorienting a bag using both arms. These are specific tests with task-adapted base policies, not a random sample of all household manipulation.
The authors describe improvement from around forty percent starting success to near-perfect windows after roughly one hundred to one hundred twenty-five online episodes. Figure 5 uses the most recent ten episodes for each success point. Individual dashed baselines differ and later points fluctuate. Preserve that moving-window denominator when describing the result.
Interactive teaching model and exact setting arithmetic. Synthetic trajectories are illustrative; no browser widget is a trained ARLI policy or a reported learning curve.
The initial investment includes human demonstration data and offline RL warmup. Assembly uses sixty base-training demonstrations, Shoe-in-Bag uses one hundred twenty-three, and Bag-Placement uses fifty-one. The real RL setup collects twelve hundred latent transitions and runs twenty-four thousand offline updates. Online episode efficiency is valuable, but it is not total learning cost.
Throughput is estimated using successful durations from the last twenty episodes and a failure duration assumed to be twice the mean successful duration. The calculator below uses invented values to make this assumption visible. With probability p and successful duration T, mean episode duration is pT+(1-p)2T. Divide expected successes per attempt by expected time per attempt, then convert seconds to an hour.
| Quantity | Meaning or value | Interpretation |
|---|---|---|
| Assembly demonstrations | 60 episodes | 84,821 recorded steps |
| Shoe demonstrations | 123 episodes | 474,301 recorded steps; also clothing examples |
| Bag demonstrations | 51 episodes | 50,041 recorded steps |
| Real online graph | Trailing 10 episodes | Windows can fluctuate |
| Throughput reference | Last 20 episodes | Failure timeout assumed 2× successful duration |
def estimated_throughput(p, success_seconds, failure_multiple=2):
assert 0 <= p <= 1
assert success_seconds > 0 and failure_multiple > 0
mean_time = (p*success_seconds
+ (1-p)*failure_multiple*success_seconds)
return 3600*p/mean_time
assert abs(estimated_throughput(.8, 12)-200) < 1e-10
assert estimated_throughput(1, 12) == 300
assert estimated_throughput(0, 12) == 0
print(estimated_throughput(.4, 12)) # 75/hour
# These inputs are illustrative, not Figure 5 bar values.
The real experiments use a bimanual robot with two arms. One task inserts a power connector into an inverter box. Another places a shoe into a narrow bag. The third lifts and positions a bag with both arms. Each task begins from a pie zero point five policy adapted using human demonstrations.
The paper reports improvement from roughly forty percent starting success to near perfect success windows after around one hundred to one hundred twenty five online episodes across these tasks. Read that statement together with the graph. Each point summarizes the most recent ten training episodes. Individual task baselines differ, and later windows can dip. Near perfect windows are evidence of improvement, not a guarantee that all future attempts succeed.
There is also work before those online episodes. The base policy receives supervised adaptation from task demonstrations. The real reinforcement learning setup begins with twelve hundred latent transitions and twenty four thousand offline training updates. Comparing only the online episode count can hide that initial investment.
Success rate is not the only practical measurement. A slow policy that succeeds often may complete fewer useful tasks than a faster one with similar reliability. The authors estimate successes per hour using successful episode durations from the last twenty episodes. They assume a failed attempt takes twice the average successful duration before failure is declared.
Try a constructed example. If successful attempts take twelve seconds and success probability is eighty percent, that timeout assumption gives an average attempt length of fourteen point four seconds. The resulting estimate is two hundred successes per hour. Those are teaching inputs, not a measured A R L I throughput result. Change the timeout assumption and the estimate changes.
Keep a successful attempt at twelve seconds and the failure timeout at twice that duration. Compare success probabilities 0.4, 0.8, and 1.0. Decide whether doubling success probability must exactly double the estimated number of successes per hour.
The estimates are seventy-five, two hundred, and three hundred successes per hour. At probability 0.4, expected time per attempt is 19.2 seconds; at probability 0.8 it is 14.4 seconds. Higher success probability both increases the numerator and reduces time spent in longer failed attempts, so the estimate grows by more than a factor of two. At perfect success the failure timeout drops out entirely.
Those numbers are synthetic inputs to the paper-style estimator. The real result uses recent successful-duration and success measurements with a stated timeout assumption. A rolling window can improve or deteriorate as new episodes arrive; its estimate is not a promise about an uninterrupted hour of future operation.
Source: ARLI v2, PDF pages 8, 16–18. See the source for original plots, uncertainty bands and implementation details.
An ablation is most useful when the removed component has a clear information role. Removing actions hides some already committed motion. Removing the intermediate observation hides events after the old snapshot. Removing RTC changes how strongly generation preserves continuity. These interventions are not three interchangeable regularizers.
Figures 6, 9 and 10 compare the information branches with and without RTC. Full information generally gives the most reliable performance, but the size of the advantage varies. On the real Shoe-in-Bag task, actions-only is competitive; Bag-Placement provides a stronger example where the intermediate observation matters. Do not state that every single branch is always necessary in every episode.
Interactive teaching model and exact setting arithmetic. Synthetic trajectories are illustrative; no browser widget is a trained ARLI policy or a reported learning curve.
The swimmer freshness sweep fixes total delay at four and changes the remaining delay from zero to four. Lower remaining delay means a newer snapshot. The appendix describes substantial gains for remaining delay below three with RTC and below two without RTC. The main page's greater-than-half phrase conflicts with that timing and graph; this lesson follows the plotted direction and appendix.
Figure 12 is a different kind of evidence. Four held-out shoe trajectories are replayed through saved actions-only models, tracking predicted noise mean and variance. The heatmaps show localized and distributed changes, not a universal one-number rescaling. They do not establish a success/failure classifier. The toy Gaussian widget teaches distributional drift; its numbers are not recovered from the paper's heatmaps.
| Quantity | Meaning or value | Interpretation |
|---|---|---|
| Actions-only | Old observation + committed queue | No fresh disturbance observation |
| State-only | Old + intermediate observation | No explicit full pending queue |
| Full ARLI | Both information sources | Their contributions can differ by task |
| Fig12 model | Actions-only | Missing intermediate snapshots in replayed data |
| Fig12 sample | 4 held-out shoe trajectories | 2 successful and 2 failed; descriptive analysis |
from math import log
def gaussian_kl(mu0, var0, mu1, var1):
assert var0 > 0 and var1 > 0
displacement = (mu0-mu1)**2
return .5*(log(var1/var0)+(var0+displacement)/var1-1)
assert gaussian_kl(0, 1, 0, 1) == 0
forward = gaussian_kl(0, 1, 1, .25)
reverse = gaussian_kl(1, .25, 0, 1)
print(forward, reverse, forward+reverse)
# We define the symmetric sum explicitly; no hidden averaging.
# This is not the source's unspecified effective-dimension formula.
Ablations ask which parts of the mechanism actually matter. The authors remove the intermediate action information, remove the fresher observation, and compare different amounts of remaining delay. The pattern is informative without being uniform across tasks.
Using both kinds of intermediate information generally works better than using only one in the tested settings. In the real shoe task, the actions only variant is competitive with the full method. In bag placement, that reduced variant struggles to maintain equally strong learning. This suggests that predictable committed motion and newly observed events can contribute differently depending on the task.
The swimmer experiments also move the intermediate observation through the inference window. A smaller remaining delay gives the actor newer information. The plots and appendix show that this generally helps. Without real time chunking, the observation must be fresher to produce the same clear learning improvement. The main text contains an inconsistent phrase about the direction of this effect; our explanation follows the timing definition, the plotted conditions, and the appendix.
A separate diagnostic examines how the learned noise distribution changes during training. It replays four held out shoe episodes, two successful and two failed, through saved model checkpoints. This analysis uses the actions only model because those saved trajectories did not contain the intermediate observation needed by the full method.
The resulting heatmaps show changes that depend on the trajectory and phase of the movement. Failed examples show more displaced and sometimes narrower noise distributions around difficult alignment portions. The authors treat this conservatively. Four replayed trajectories do not establish a population level failure detector. They provide a way to inspect what the noise steering network changes, rather than assuming it merely rescales randomness by the same amount everywhere.
Set the later toy Gaussian mean to zero and variance to one. Both directions of KL divergence should vanish. Then restore mean one and variance one quarter. Why can the two directed divergences disagree even though they compare the same pair of distributions?
Matching distributions give zero in both directions. With mean one and variance one quarter, the directed values are approximately 2.806853 and 0.818147; their symmetric sum is 3.625. KL is an expectation under its first distribution, so swapping the arguments changes which regions receive more weight. This widget explicitly uses the sum of the two directions, avoiding ambiguity with conventions that divide that sum by two.
The curves are synthetic Gaussians, not digitized paper data. The source's four held-out actions-only trajectories illustrate learned noise behavior but cannot establish a population classifier. A separation in a small diagnostic plot needs broader validation before supporting an operational detection rule.
Source: ARLI v2, PDF pages 9, 19–22. See the source for original plots, uncertainty bands and implementation details.
A timing diagram is a useful taxonomy. Base RTC constrains generation using committed motion. ARLI selects latent noise before a frozen action expert finishes. A stepwise residual controller adds a small action correction after generation. The parallel Real-Time EXPO-FT paper edits and selects from completed action proposals at a replanning boundary and also updates its base by behavior cloning on successful experience.
These are different intervention points, not a single speed ranking. ARLI must choose noise before the expert runs, so it cannot use observations that arrive after that deadline. A later editor can see later information, but it acts in a different space and has its own computation and training design. The papers do not supply a direct head-to-head benchmark, so their separate headline results should not be compared as one scoreboard.
Interactive teaching model and exact setting arithmetic. Synthetic trajectories are illustrative; no browser widget is a trained ARLI policy or a reported learning curve.
Without a separable backbone and action expert, ARLI cannot defer its noise input in the same way. It can still condition the learner on pending actions. If the generator's reachable behaviors do not include a useful correction, selecting better noise may remain insufficient. More information improves a decision problem; it does not automatically create every missing capability.
To implement the core idea, begin with the queue, timestamps and observation contract. Then define the actor's exact intervention point and make training reproduce that schedule. Verify what is frozen, which values enter replay, and which prefix cannot be changed. Only after this bookkeeping is coherent should the model capacity or optimization settings become the main question.
| Quantity | Meaning or value | Interpretation |
|---|---|---|
| RTC | During chunk generation | Continuity guidance; no reward-trained actor by itself |
| ARLI | Before action-expert denoising | Small learned noise actor; frozen base during RL |
| Stepwise residual RL | At individual action steps | Adds a direct action correction |
| Real-Time EXPO-FT | At proposal/replan boundary | Latest-observation edit and Q selection; base also learns with BC |
| Unsupported conclusion | Universal winner | Different tasks and setups prevent direct headline ranking |
def available_arli_features(split_expert, queue_available,
intermediate_observation_available):
features = ["old_observation"]
if queue_available:
features.append("committed_actions")
if split_expert and intermediate_observation_available:
features.append("intermediate_observation")
return features
assert available_arli_features(False, True, True) == [
"old_observation", "committed_actions"]
assert len(available_arli_features(True, True, True)) == 3
# This checks information availability, not predicted task success.
We can now place several robot learning methods on one timeline. Real time chunking maintains continuity while a base policy computes asynchronously. A R L I chooses latent noise before the action expert runs, using information that arrives later than the original backbone observation. A stepwise residual controller changes individual actions after generation. These intervention points come with different information deadlines and different constraints.
The parallel paper on Real Time E X P O fine tuning intervenes after action proposals have been generated. It edits and selects from generated action proposals using a more recent observation at a replanning boundary. Its base policy also receives behavior cloning updates from successful experience. That is a different training design from A R L I's frozen generator and latent steering. The two papers use different tasks and experimental setups, so their separate headline numbers do not establish a direct winner.
For a new system, first ask when the needed input becomes available. If the generator requires its noise immediately and cannot separate vision computation from action generation, A R L I cannot obtain the same late observation opportunity. Its committed action conditioning can still be useful. If a disturbance arrives after the actor has selected noise, that particular decision cannot react to it. If the frozen generator cannot express a successful continuation, better information alone may not create one.
The contribution is a precise correction to the learning problem: tell the small controller what is already happening and give it the freshest observation its intervention deadline permits. Then train under that same schedule.
When you explain A R L I to someone else, draw the queue before the network. Mark the old image, the later image, the noise deadline, and the first new action. Once those times are visible, the architecture has a reason to exist. The robot can keep moving while it learns what waiting means.
Place ARLI, RTC, a stepwise residual controller, and Real-Time EXPO-FT along the proposal pipeline. For each, identify the object changed by the intervention. Then disable late noise access for ARLI and ask which timing opportunity disappears.
ARLI chooses latent input before the frozen action expert. RTC guides generation toward a compatible committed prefix. The cited stepwise residual baseline corrects an action during execution. Real-Time EXPO-FT edits completed proposals and uses a critic to choose at a replan boundary; its base also receives successful-experience behavior-cloning updates. Without a late noise interface, ARLI cannot use the same intermediate snapshot opportunity, though committed-action conditioning can still be relevant.
This is a comparison of intervention points and information paths. Different environments, command rates, delay settings, and training protocols prevent a fair ranking from unrelated headline results. A shared benchmark would be needed to compare effectiveness under matched latency and compute budgets.
Source: ARLI v2, PDF pages 9, 15; related paper 2609.18207. See the source for original plots, uncertainty bands and implementation details.