GPT-Policy

An off-the-shelf commercial vision-language model, GPT-6 Astra, is given robot tasks with different kinds of context in its input. One human video lifts towel pickup from 0 of 3 trials to 2 of 3, and recorded robot actions take bottle opening to 3 of 3. Its weights never change.

Pick a task, then change what the robot is shown and watch the success dots, the decisions and the minutes move. Then we build, piece by piece, the loop that turns the model's JSON into checked arm motion.

You need what a vision-language model is and what inverse kinematics does for a robot arm. We build the rest from zero.

Teach it in context

Ready

Choose a context. The same frozen model attempts the same task each time.

Success counts, average decisions and average minutes are Table 1 (GPT-6 Astra on real robots, three trials per condition; the averages include failed trials). The arm's motion is a cartoon of one run, not a replay: the paper publishes counts and averages, not per-trial motion. The counters run to the paper's averages. The two-handed towel grasp is the one the paper writes out for its text-only towel variant.

Chapter 0

A towel it has never seen

Why a robot has to learn at the moment of use, and what it means to learn with frozen weights

A red towel lies flat on a table. A two-armed robot looks down at it through three cameras: one fixed above the table and one on each wrist. You type a single sentence: "Pick up the red towel with the robot's right hand." No demonstration of this towel, this table or this lighting was ever given to the robot's controller. Can it do it?

For a person the task is easy, and the reason it is easy is interesting. A towel lying flat has no edge standing up to pinch. Grab it from above and your fingertips close on air and slide across the cloth. So you press one hand on the towel to hold it still, slide the other hand's fingers under its edge, and only then grip. Most people learned that trick by watching someone else do it once.

Robots have mostly learned the other way: by training. A policy is the program that maps what a robot senses to what it does next. The strongest modern policies are vision-language-action models (VLAs): large neural networks, such as π0 and π0.5, trained on many thousands of recorded demonstrations. When they meet a situation their data never covered, the usual fix is to collect more demonstrations and train again. Training changes the network's weights, the enormous list of numbers that defines what the network computes.

The paper opens with the uncomfortable truth behind that fix: no finite collection of demonstrations can cover every task and situation a robot will meet. There will always be a new towel, a stiffer notebook, a bottle whose cap sticks, a person who wants something slightly different today. The ability to adapt at deployment, from whatever the robot is shown right then, is not a luxury feature. It is the only way past the edge of the training data.

Language models already do something like this. Put three examples of English sentences with their French translations in a prompt, and GPT-3 translates a fourth sentence, even though not one of its weights moved. That ability is called in-context learning (ICL): learning from examples placed in the model's input, at the moment of use, with the model itself left exactly as it was. The examples shape the answer the way a worked example on a whiteboard shapes a student's next attempt, without rewiring the student.

This paper, by Dongzhou Cheng, Taoran Yi, Ye Fang, Xingwu Zhang, Fan Feng, Yixuan Li and Gengxiong Zhuang (co-first authors) with colleagues at Morphi Robot, the Shanghai Innovation Institute and eight universities, asks whether that trick carries over to a robot. Its definition is precise. Robotic in-context learning is the ability to adapt behaviour based on demonstrations, examples or interaction experience provided at test time, without gradient updates or persistent task-specific parameter changes.

Two phrases in that definition do all the work. A gradient update is one step of training: compute how wrong the network was, and nudge every weight a little in the direction that makes it less wrong. A persistent task-specific parameter change is any saved modification for this task, even a small adapter bolted onto a frozen network. The definition forbids both. Everything the robot learns about the towel must live in its input, and it evaporates when the episode ends.

Here are the two ways to teach a robot a new task, side by side:

θ′
New weights after one training step. The left formula is fine-tuning: the demonstrations D change the model itself.
η
The learning rate: how large a nudge each step takes.
∇θ L(θ; D)
The gradient of the training loss on the demonstrations: the direction in which every weight should move to fit D better.
a
The robot's next action. On the right, nothing is trained: the model simply produces an action.
πθ
The policy with its weights θ. On the right, θ is the same before and after the episode.
D, T, o
The demonstrations, the task instruction and the current observation, all passed in as input. D teaches by being read, not by being fitted.

The learner in this paper is not a robot model at all. It is GPT-6 Astra, a commercial vision-language model (VLM): a large language model that also reads images, and a general-purpose model that supports reasoning, tool use and multi-step task execution. The authors wrap it in a framework they call GPT-Policy, which has three parts: a context compiler that selects what goes into the model's input, the VLM itself proposing robot-tool actions, and a constrained controller that checks each proposed action, executes it, and reports the outcome back.

What can go into the context? The paper tests five families, and each pins down a different part of a task. A human video shows a person doing the task, a procedure with no robot numbers in it. A robot video of a teleoperated demonstration shows the same kind of procedure on the robot itself, optionally with the recorded actions logged alongside it, which show the motion between frames. A goal image shows the desired end state. Self-interaction history records what the robot has already seen and changed. Online human interaction carries a partner's moves, gestures and intent. Toggle them below (recorded actions get their own switch) and see which questions about a task each one answers.

What does the robot still have to guess?

Left: the kinds of context the paper tests. Right: five questions any manipulation task raises. Switch sources on and watch which questions get answered by context and which stay the model's to infer from the instruction and the live cameras. The presets load the paper's own conditions.

The mapping from sources to questions follows the paper's own wording (Sections 1 and 3.2): goal images specify outcomes; human and robot videos illustrate procedures and end with the outcome; aligned actions provide motion references; history records earlier observations and outcomes; human feedback clarifies intent or rules. Results quoted are Table 1. The five questions are our framing.

Notice what the "instruction only" preset leaves: every question is the model's to answer from one sentence and three camera images. In Table 1, every task that was tested in that condition failed all three trials: towel, notebook, bottle and plug, 0 of 3 each. Add a human video to the towel and two of three trials succeed. Section 4.2 describes the no-video condition as the same instruction and observations without the video, but Appendix D.4 shows that the video condition's instruction also tells the model to imitate the demonstrated grasp. Success criteria and termination rules were identical. So the comparison is video plus a pointer to it, against neither, with three trials each.

Why ask a general chat model at all, when robot foundation models are starting to do in-context learning themselves? The paper names several: GEN-1.5 reports one-shot skill adaptation from physical prompts, S1 uses video demonstrations to specify new tasks, and Zero-WAM trains a video-action model to follow human video guidance on unseen tasks. Those are trained as robot policies. The paper's question is complementary: how far can an off-the-shelf general model go without being trained as a robot policy? The answer separates the task understanding a general model already has from the abilities that genuinely need specialised embodied learning.

The answer has two halves, and the paper reports both honestly. The first half is the hero above: context helps, sometimes dramatically, and a video of a human hand can steer a robot gripper even though it contains no robot numbers at all. The second half is the gap: better plans do not guarantee precise contact, reliable verification of success, or physical safety. The same trials exposed slow and expensive decisions, and repeated collisions between the robot's own two arms.

The idea in one line: freeze the model, move the learning into the input. GPT-Policy never trains anything. It decides what the model reads (the compiler), constrains what the model can do (the tools), and checks every request for reachability, IK residuals, joint limits and timing (but not collisions) before a motor turns (the controller). Whatever adaptation happens, happens between the prompt and the next JSON reply.

Here is the road. Each chapter explains one part of the instrument in the hero.

  1. One decision at a time: Equation 1, the loop behind the decision counter, and why "done" is not the same as success.
  2. What goes in the window: the context compiler behind the context selector, keyframes, references and history.
  3. Speaking in poses: the JSON tool requests the model writes at every tick of the counter.
  4. From a pose to joint angles: the blue path, sampled and checked by inverse kinematics within 2 millimetres (YAM and ARX X5; 3 on the Morphi Kino).
  5. Slow down until it is safe: time scaling, synchronised arms and the feedback that closes the loop.
  6. Watch a person once: the towel and notebook stops, where a human video lifts 0 of 3 to 2 of 3.
  7. Between the keyframes: the bottle and plug stops, where recorded actions gave the best observed success.
  8. Goals, memories and a partner: goal images, self history and a human across the table.
  9. What each decision costs: three models on the same towel, and the minutes behind every move.
  10. Good plans, tricky execution: the limits, the six directions and the whole loop as code.
What makes GPT-Policy's adaptation in-context learning rather than fine-tuning?

Chapter 1

One decision at a time

Equation 1 turns a frozen vision-language model into a robot policy that asks, acts and hears back

A chat model answers a message with a message. A robot needs something different: a stream of commands, each chosen after seeing what the previous one actually did to the world. GPT-Policy bridges the two with a loop in which every turn of the conversation is one physical decision.

Start with the pieces the paper names. The task instruction T is the sentence you typed: "Pick up the red towel with the robot's right hand." The observation ot is what the robot senses at decision step t, and it has two parts. It is the latest set of camera images, each one labelled by the camera that took it: top, left wrist, right wrist. st is the robot state: joint angles and velocities, the pose of each gripper's tip, and how far open each gripper is, both as commanded and as measured.

The context ct is everything extra the robot has been given or has gathered by step t. Depending on the experiment it can hold a goal image G, demonstration videos V represented by selected keyframes, recorded action sequences A with any measured robot states that came with them, and the online interaction history Ht. Chapter 2 is entirely about how ct is assembled.

The model's output is a tool request at = (ut, vt): a tool name ut, such as move_to, plus its arguments vt, a JSON object such as a target pose for the right gripper. That is the entire vocabulary of the policy. Pick a tool, fill in its arguments. Chapter 3 lists the tools.

Equation 1 of the paper puts the pieces in a loop:

at
The tool request at decision t: a tool name ut and a JSON arguments object vt.
πθ
The VLM acting as the policy. Its parameters θ remain fixed during task execution: nothing is trained, ever.
T
The task instruction, one sentence of text, the same at every step.
ct
The context: task references (goal image, demonstration keyframes, recorded actions) and the interaction history so far.
ot
The latest observation (It, st): view-labelled camera images and the measured robot state.
ft−1
The result of the previous request: returned information, execution progress, or an error. Absent at the first decision.
ℰ
The robot-tool interface: the adapter that checks, times and executes a request on the real robot.
ft
What ℰ reports back about this request: executed with an endpoint error, rejected with a reason, or a completion review's verdict.

Read the first line as a sentence: the model looks at the instruction, the context, the latest observation and the result of its previous request, and produces the next request. The symbol ∼ means "is sampled from". The model generates text token by token, so two runs from identical inputs can produce different requests. A policy that samples is not a bug here; it is simply what a language model is.

Read the second line as the world answering back. The interface ℰ takes the request and the current observation, does something (moves an arm, or refuses to), and returns two things: the next observation ot+1 and a tool result ft. After a completed or a rejected request, the next decision receives fresh observations and that result. At the very first decision there is no preceding result, and the paper's observation envelope simply omits the field.

Now the most important subscript on the page: θ. These are the model's parameters, and the paper states that they remain fixed during task execution. Nothing in the loop computes a loss, takes a gradient or saves a weight. Every bit of adaptation must travel through the inputs on the right of the bar: T, ct, ot and ft−1. That is exactly what makes this in-context learning and not training.

Why does the previous result ft−1 get its own slot, separate from the observation? Because some of the most useful information is not visible in any camera image. When the controller refuses a request because the arm cannot reach the target precisely enough, nothing moves. The images before and after are identical. Only the tool result can tell the model that its request failed, and why. Step through a short episode and watch where each piece of information travels.

Step the loop, one decision at a time

A scripted episode of Remove and Reinsert Plug. Press Next decision: the dot travels from the inputs to the frozen model, out as a JSON request, through the adapter, and back as feedback. Watch the rejected request, the path check that moves nothing, and the first "done". Then switch the completion review off and replay.

An illustrative episode: the tool names, the argument fields, the review behaviour and the rules it shows follow the paper's Appendix D (F1, F6, F7, P3, P4) and Appendix B; the coordinates, residuals, endpoint errors and feedback wording are made up to show the mechanism, and the decision tally assumes every request counts once, which the paper does not specify. The review in a separate model session belongs to one configuration, the execution wrapper of the ARX Claude Messages adapter; the inspected YAM branch omits it, and the paper does not say which configuration each Table 1 trial used.

Three things in that episode deserve names. First, a rejected request still costs a turn of the loop. The model spent a reply, the arm stayed still, and the reason came back in ft. The prompts tell the model what to do next: "After IK rejection, first compare approach positions, heights, and permitted axial rotations while retaining the required tool direction. Do not tilt the grasp or replace the manipulation merely to make IK pass."

Second, the check_path tool lets the model ask "would this work?" without moving. The adapter runs the full inverse kinematics, joint limits and timing on the proposed path and reports feasibility, but sends no command. The paper is careful about what acceptance means: it does not certify collision clearance or physical tracking. A path can be kinematically perfect and still drive a camera housing into the table.

Third, done is itself a tool call, and a tool call is a claim. Model-declared completion remains distinct from physical task success, and the paper repeats this in the method, in the appendix and in the prompts. In one configuration, the ARX Claude Messages adapter, an execution wrapper reviews every completion request in a separate model session and returns unverified ones to the control loop. The inspected YAM branch omits that review, and the paper does not say which configuration each Table 1 trial used. Either way, the success counts in Table 1 come from a human judging the final scene against task-specific geometric and semantic criteria.

What counts as one decision? Under a task-specific convention, each generated action target or action block counts as one decision. The paper does not say, task by task, whether a move_eef_chunk through five waypoints counts once or once per waypoint, or whether rejected requests and check_path calls count. The tally in the device above simply counts every request once, as an illustration. This matters because decision count is one of the paper's two cost metrics, and the prompts push the model toward chunks: "Prefer move_eef_chunk for clear, contact-free paths that need no intermediate observation." They also push back the other way: stop and observe at contact, at gripper changes, at occlusion. Efficiency never justifies skipping verification.

Let's put a tempo on the loop, using the paper's own numbers. On Pick Red Towel with no demonstration, GPT-6 Astra averaged 96.3 decisions and 24.6 minutes per trial (Table 1). Convert the minutes to seconds and divide:

24.6 × 60 = 1,476 s   ÷   96.3 = 15.3 s per decision (no demonstration)

18.9 × 60 = 1,134 s   ÷   76.7 = 14.8 s per decision (with the human video)

That is elapsed time per decision: the model reading its input and writing a reply, the controller planning, the arm moving and settling, and fresh images coming back. The paper does not split those parts, so neither will we. But the order of magnitude is the point. This policy thinks and acts about four times a minute. A trained VLA issues motor commands many times per second. Chapter 9 returns to what that means.

The loop has more exits than done. Every episode starts from a reset scene and ends when the task succeeds, when the execution budget runs out, or when a safety termination condition triggers. The model can call give_up, but only "when evidence shows that the task cannot be completed or further attempts would violate safety constraints"; the prompts forbid giving up while reasonable safe strategies remain. Faults and operator interruption can also end a trial. All conditions share the same success criteria and termination rules, and every experiment is repeated three times.

The feedback is the only teacher. In a trained policy, a mistake teaches through a gradient. Here there is no gradient, so the only channel from a failure to a better next decision is ft and the fresh images: plain text and pixels in the next input. That is why the paper's feedback format asks for a concrete rejection reason and requested-versus-measured endpoint errors (something like "residual 3.4 mm at sample 18", an illustrative message; the paper does not print the format) rather than a bare "failed".
The controller refuses a move_to because the arm cannot reach the target within tolerance, so the arm never moves. Through which part of Equation 1 does the model learn this?

Chapter 2

What goes in the window

The context compiler keeps the transitions that matter, pins the teacher, and lets old chatter go

Everything GPT-6 Astra knows about the current task, it knows because it is in the input of the current decision. So the question "what can the robot learn in context?" becomes a very practical one: what, exactly, is in that input, in what order, with what labels? The part of GPT-Policy that answers it is the context compiler, and the paper's abstract describes its job in one phrase: it preserves task-relevant visual transitions.

Think of it as an editor preparing a briefing for a busy expert. The expert (the model) is capable but reads only what is on the desk. A 24-second video is hundreds of frames, and most of them show a hand hovering, or nothing changing at all. The editor's job is to pull out the handful of frames where something happens, caption them, mark them clearly as history, and keep the desk from drowning in yesterday's paperwork as the task goes on.

Task references: the teacher's material

The paper calls the offline material task references, and there are three kinds. A goal image G is a visual reference for the target arrangement or outcome, "without prescribing intermediate actions". Demonstration videos V show object interactions and the order of actions, either from a person doing the task (a human demonstration) or from a human operator steering the robot (a teleoperated robot demonstration). Recorded actions A optionally supplement a robot video with the commands and measured states logged during that demonstration.

One rule governs all three: these reference records remain distinct from the model's own tool requests at in the current trial. A recorded action from yesterday's demonstration is evidence, not a command waiting to be executed. The prompts enforce the boundary with plain text markers. Every reference opens with the words "HISTORICAL DEMONSTRATION" and closes with "END HISTORICAL DEMONSTRATION. Use the current task and live observations below."

Inside those markers the instructions are strict: "Images and actions below describe a previous episode, not the current scene or pending commands. Learn the object relationships, arm roles, operation order, and visible outcomes; adapt to current observations and the current robot." And later: "A demonstration's outcome describes that historical episode only." A model that finished the demonstration's stages has not thereby finished the current task.

How does a video become model input? As keyframes: a small set of selected moments, each carried as an image block plus a text record. For a human video the record is simply the time, "Historical human demonstration, t=5.190s", followed by the image. For a robot video it is a small JSON record with the time, a short stage name, what the image shows, the visible result, and which camera images belong to it. The paper prints one from the bottle demonstration: at t = 7.53 s, stage "left body grasp", observation "Left gripper surrounds the upright bottle body while the right arm remains clear", result "Body support is established before the right arm approaches."

Choosing the keyframes

Who picks the frames? Another model call, before the robot moves. In the first pass (prompt P0a), a vision model looks at candidate frames within overlapping windows of the video and selects "the smallest set that conveys the initial state, key actions, state changes, and final outcome". It is told to preserve before and after evidence for grasping, release and handoffs, to keep the order of repeated twists, and to remember that similar start and end poses do not imply no motion: a cap twisted a full turn looks the same before and after. It is also told something every annotator should hear: "gripper closure alone does not prove a grasp." A closed gripper in a frame might be holding the object or holding air.

In the second pass (P0b), the model reviews the whole demonstration at once and consolidates. It merges redundant holds and window boundaries while keeping the initial state, preparation, contact, grasp verification, arm-role changes, release and final outcome. It "usually" retains 12 to 16 frames, and it gets a warning of its own about the wrist close-ups: "do not mislabel empty-gripper pressing as regrasping."

Then come hard limits (Appendix C.3). A reference contains at most 24 keyframes and at most 48 images, each resized to fit within 1,280 pixels per dimension, never upscaled. Here is a worked example of how those two caps interact. The plug demonstration uses three camera views at every keyframe (Table 5):

14 keyframes × 3 views = 42 images (under the 48-image cap)

48 images ÷ 3 views = 16 keyframes (the real ceiling for a three-view demo, not 24)

So for a three-camera demonstration the image cap binds long before the keyframe cap does, and 16 keyframes is the most it can carry. The bottle demonstration, recorded from the overhead view only, carries 13 keyframes as 13 images. The towel's human video carries 8 frames from a single view (Table 4). None of these comes close to "the whole video".

The compiler also has to keep comparisons honest. The keyframe annotations themselves can describe the procedure in words, so if the Video and Video + Action conditions used different frames or different captions, any difference in results could come from the captions. The paper states the requirement plainly: comparisons of recorded-action inputs "must hold the selected images and non-action text fixed". The two conditions share the selected images (Appendix C.2), and only Video + Action adds the measured states and action segments (Table 5). Chapter 7 looks at how far that control goes.

Interaction history: the growing part

References are fixed for the whole episode. The online interaction history Ht is not: it grows by one exchange at every decision. It records observations, tool requests, results and operator feedback, separately from the references. Left alone it would swamp everything else, and a little arithmetic shows how fast. With three live cameras, one decision adds three images. On the towel with the human video, a trial averaged 76.7 decisions:

76.7 decisions × 3 live images = about 230 images (against 8 reference images)

The paper's answer is that provider adapters manage the history, the small layer of code that talks to each model provider's API. They retain the reference inputs, limit older live images, and either retain the accumulated text or replace older exchanges with host-generated summaries. The two implementations the authors inspected differ exactly on this choice. The Codex refresh keeps the reference content and all accumulated text while omitting older live images. The ARX Claude Messages adapter pins the initial input, limits recent live images, and summarises older exchanges. Either way, the ongoing trial and its decision count are preserved.

Live images themselves arrive at 640 × 480 pixels, JPEG quality 85, on the YAM and ARX X5 arms, and at 1,280 × 720 on the Morphi Kino (Table 3). Each decision's input is assembled in a fixed order, the paper's format F1: a JSON envelope with the instruction, per-camera metadata, the measured state of each arm, a step index, and the previous result (omitted at the first step), followed by the left wrist, right wrist and top image blocks. Compile a window yourself.

Compile the context window

Each column is one past decision: three live images above, the request and the feedback below. Choose a reference, slide the decision count, and switch between the history policies. Watch what is pinned, what is dropped and what is folded into a summary.

30

Reference sizes are Tables 4 and 5; the caps (24 keyframes, 48 images) are Appendix C.3; the three policies follow Appendix B's description of the Codex refresh and the ARX Claude Messages adapter. How many recent decisions keep their images (3 here) and when summarising starts (after 8 turns here) are not given in the paper and are illustrative.

Two design choices are visible in the compiler. The first is asymmetry: references are pinned and never evicted, while live history is aggressively trimmed. The teacher's material is small, fixed and expensive to have collected; old live frames are cheap to regenerate (the robot can always look again) and usually stale. The second is that trimming touches images first. A text record such as "move_to rejected: residual 3.4 mm at sample 18" (an illustrative message: the paper specifies a concrete rejection reason but does not print its format) is a few dozen tokens and stays useful for a long time; an old camera frame costs far more and describes a scene that no longer exists.

What is lost when older exchanges become a summary? Detail, and possibly the exact failure reason from twenty decisions ago. The paper does not measure the effect of the two history policies, and it lists "context quality" among its open constraints. The host generates the summary, which means the harness decides what the model will remember. That is a design lever worth noticing: in-context learning is only as good as the context someone chose to keep.

For self-history tasks (Chapter 8), the history is the whole point: the prompt format F5 lays out past observations with their images, the selected tool and arguments, the tool outcome and measured feedback, and the next observation, turn after turn, before the current live observation. For online human interaction, it also carries the human's moves or gestures and the robot's responses. Chapter 8 plays one of those games with you.

Pin the teacher, trim the chatter. The compiler's single most important decision is what never leaves the window. References (goal images, keyframes, recorded actions) are pinned for the whole episode; live images age out; old text may be summarised. The model can always look at the scene again, but it can never re-watch a demonstration that was evicted.
A demonstration is recorded with three cameras. The reference allows at most 24 keyframes and 48 images. How many keyframes can it carry?

Chapter 3

Speaking in poses

A handful of tools and one JSON contract let a language model command an arm without ever naming a joint angle

A language model writes text. A robot arm wants joint angles, a hundred times a second. Somewhere between the two there has to be a language both sides can speak. Ask the model for joint angles and it must solve the arm's geometry in its head. Ask it for a pixel to touch and the robot must guess the depth. GPT-Policy chooses a middle ground: the model names where the tip of the gripper should be, and how it should be turned, in the arm's own coordinate frame.

That tip has a name. The tool center point (TCP) is a calibrated reference frame on the end effector, usually the spot between the fingertips where an object is held. On the YAM arms it is a calibrated grasp_site; on the ARX X5 it is the inner-fingertip midpoint; on the Morphi Kino it is a link near each hand (Table 3). The system prompt insists: "Use the calibrated grasp reference; do not substitute a different fingertip or flange origin."

Every arm has its own base frame: +x points forward, +y to the left, +z up. The gripper has its own axes too: tool +z points from the wrist to the fingertips, and tool +y is the axis along which the jaws open. A pose is then two things: where the TCP is (a position p, three numbers in metres) and how the gripper's axes are turned relative to the base axes (an orientation R, a rotation). The paper's first rule of the interface is short: "Control the robot using absolute calibrated TCP targets, not joint-angle commands."

Writing a rotation with four numbers

Positions are easy: x, y, z in metres. Rotations need a trick. Any rotation in 3D can be described as turning by some angle φ about some axis n (a unit vector). A quaternion packs that into four numbers: the axis scaled by sin(φ/2), and cos(φ/2). The half angle looks odd, but it is what makes the arithmetic of composing rotations work out, and the rotations lesson derives it. GPT-Policy writes quaternions in xyzw order: the three axis parts first, the scalar last.

So a full target, which the paper calls pose_xyzquat, is seven numbers: [x, y, z, qx, qy, qz, qw]. The prompt's own example is worth decoding by hand. The quaternion xyzw = [1, 0, 0, 0] has axis part (1, 0, 0) and scalar 0. So cos(φ/2) = 0, which means φ/2 = 90° and φ = 180°, about the x axis. Spinning 180° about x leaves x alone and flips both y and z. Tool +z therefore points along base −z: the fingers point straight down at the table, and tool +x still points forward. The paper says exactly that, and adds a caution: it is "an axis example, not a universally reachable or collision-free target."

A rotation must be a unit quaternion, length exactly 1. Models write rounded numbers, so the adapter repairs them: "Finite, nonzero quaternions are normalized."

q̂
The unit quaternion the adapter actually uses as the target orientation.
q
The four numbers the model wrote, in xyzw order, possibly rounded.
‖q‖
Its length. Must be finite and nonzero: a zero quaternion points nowhere and cannot be repaired.
qw
The scalar part, cos(φ/2). Written last in xyzw order.

Here is a worked example with real numbers from the paper. The first recorded action row of the bottle demonstration (format F3b) gives the left gripper a target of [0.212, −0.024, 0.057, 0.080, 0.854, −0.065, 0.509]. The last four numbers are the quaternion, rounded to three decimals. Square them: 0.0064 + 0.7293 + 0.0042 + 0.2591 = 0.99902. The length is √0.99902 = 0.99951, a hair short of 1. Divide each part by it: [0.08004, 0.85442, −0.06503, 0.50925]. The prompt tells the model to do exactly this with rounded reference quaternions before issuing targets.

0.854 ÷ 0.99951 = 0.85442 (qy after normalisation)

What orientation is that? The rotation angle is 2 × arccos(0.50925) = 118.8°, about an axis close to base +y (the "left" axis). Push the tool's +z axis through that rotation and it comes out as roughly (0.86, −0.19, −0.47): pointing forward and about 28° below horizontal. If the recorded end-effector frame follows the same tool axes as the system prompt's TCP convention (Appendix D's P2 tells the model to check exactly that), this points the left fingers forward and about 28° down, at the very start of the demonstration, while both grippers are still clear of the bottle. Four numbers, decoded into a direction you can picture.

The tools

Every decision returns exactly one tool selection, as one JSON object, {"name": "<tool name>", "arguments": <tool arguments>}, with no Markdown and no text outside it. Provider adapters normalise either a structured JSON reply or a provider's native tool call into that same name and arguments pair. These are the tools the paper's prompts define:

Every motion request also carries a note: a short statement of the current evidence and the purpose of the move. The note is not decoration. It forces the model to say what it sees before it acts, and in reference-guided tasks it must "identify the relevant stage" of the demonstration being followed. Chapter 1's episode showed notes like "Plug visible in top view; approach above it, contact-free."

Nulls, completeness and the gripper

A two-armed request names a target for each arm, and either may be null. A null entry holds that arm's preceding pose; an arm whose entries are all null receives no new trajectory at all. This is how the model says "left arm, stay exactly where you are" without having to copy seven numbers it might get slightly wrong.

What the adapter will not do is fill in half a pose: "Individual coordinates are not filled independently." A target is complete or it is not a target. This rule looks pedantic until you imagine the alternative: the model writes only a new z to lower the gripper, the adapter silently copies the old orientation from the measured state, and the measured state has drifted a degree under load. The prompt warns about exactly that drift: "Never replace a held orientation with load-induced measured drift."

Finally, the gripper. Cartesian motion never changes the gripper's opening; only set_gripper does. Why split them? Because closing a gripper is the moment a grasp succeeds or fails, and the prompts want the model to look first: "Approach and align, inspect fresh images and state, then call set_gripper separately." Results always distinguish what was requested, what was submitted and what was measured. For illustration: commanding 0.0 on a YAM gripper and measuring 0.31 would mean the jaws stopped on something about 0.31 × 0.095 = 0.029 m wide under the nominal conversion, using the 0.095 m opening width from Table 3 (the paper calls these widths nominal conversions of normalised readings). Build a request and watch the adapter read it.

Build a tool request

Choose a tool, a target position and an orientation preset, and switch the left arm between null and a target. The drawing shows the gripper in the right arm's base frame (+x forward, +y left, +z up); the JSON below is exactly what the model would write. Try the zero quaternion, the incomplete pose and the "wxyz by mistake" preset.

0.300 m
−0.050 m
0.120 m
0.00

The request formats, the xyzw order, the normalisation, the null and completeness rules and the gripper range are the paper's (Section 3.3, Appendix A.1, F6a). The "recorded, rounded" preset is the left-arm quaternion from the bottle's first recorded action row (F3b). The drawing is schematic, and nothing here checks reachability: that is Chapter 4's job.

The "identity written wxyz" preset is the trap every robotics programmer falls into once. Many libraries write quaternions scalar first, [w, x, y, z]. The identity rotation (no turn at all) is then [1, 0, 0, 0]. Hand those same four numbers to an adapter that reads xyzw and it sees a 180° flip about x: the gripper that was meant to point up now points down. No error is raised, because both readings are perfectly valid rotations. This is why the Video + Action prompt tells the model to check "the source embodiment, base frame, TCP reference, quaternion order, and current object alignment" before using any recorded number.

Why is this interface a good idea? It splits the work along the line of competence. The model is good at reading images and text and at choosing what should happen next in space: "put the fingers just beyond the towel's edge". It is poor at trigonometry under a time limit. The adapter is the opposite: it knows the arm's geometry exactly and has no idea what a towel is. Absolute Cartesian targets are the narrowest message that lets each side do only what it is good at.

The interface also makes every request checkable. A seven-number pose with a known frame and a known order is something a controller can validate before any metal moves, which is the subject of the next two chapters. A free-form instruction like "gently grab the towel" is not.

Space, not joints. The model never names a joint angle: the system prompt forbids it, and joint readings arrive only as measured feedback. That single choice is what lets the same kind of request drive three different robots (six-joint YAM and ARX X5 arms, seven-joint Morphi Kino arms) through embodiment-specific adapters.
A bimanual move_eef_chunk gives the right arm three waypoints and sets every left-arm entry to null. What does the left arm do?

Chapter 4

From a pose to joint angles

The adapter draws a straight, shortest-turn path, solves the arm at every sample, and refuses any sample it cannot hit within 2 millimetres (3 on the Morphi Kino)

The model has just written: "right fingertips here, turned like this." The arm has six joints (seven on the Morphi Kino), and it does not understand "here". Two problems stand between the request and motion. Between where the gripper is and where it should be there are infinitely many paths, so one has to be chosen. And for every point on that path, the adapter must find joint angles that put the gripper there, or admit that none exist.

Figure 3(b) of the paper describes the adapter's pipeline in five verbs: resolve the targets, sample the pose path, check inverse kinematics residuals, time the joint references, execute. This chapter covers the first three; Chapter 5 covers the last two. Everything here happens before the arm moves, and every box can say no.

Where the path starts

The starting pose is not what the model thinks it is; it is what the robot measures. The planner reads the measured joint angles and runs forward kinematics (FK): the arm's geometry equations, which turn joint angles into the pose of the TCP. That pose, (p0, R0), is the first point of the path. Each target the model sent then goes through a calibrated transform from the TCP frame into the end-effector frame the IK solver works in. None of this is visible to the model, which is the point: it speaks TCP poses, the adapter does the conversions.

The path: straight lines and the shortest turn

The paper's choice of path is the simplest reasonable one. Positions move along a straight line. Orientations turn at a constant rate about a single fixed axis, the shortest way round. Equation 2 writes both for one segment, from target j to target j + 1, as a function of a progress variable s that runs from 0 to 1:

pj(s)
The TCP position partway along segment j: a straight line from pj to pj+1.
s
Progress along the segment, from 0 (start) to 1 (end). No time yet: timing comes in Chapter 5.
Log(Rj⊤Rj+1)
The relative rotation from the start orientation to the end one, written as a single axis and an angle (the angle is the vector's length).
exp( s · )
Turn a fraction s of that angle about that same axis, and convert back into a rotation.
Rj(s)
The orientation partway through: the start orientation followed by the partial turn. This is SLERP, spherical linear interpolation.

The rotation half deserves a slower read. Rj⊤Rj+1 is "the turn that takes the start orientation to the end one". Log converts a rotation into an axis-angle vector: a direction to spin about, scaled by how far to spin. Multiply that vector by s and you have "the same spin, only s of the way". exp converts it back into a rotation matrix. The result is SLERP (Shoemake, 1985): the gripper turns at a steady angular speed about one fixed axis, never wobbling. The Lie groups lesson builds Log and exp from scratch.

Now a worked example. Say the segment runs from pj = (0.30, 0.00, 0.20) m to pj+1 = (0.40, 0.10, 0.10) m, and the orientation must turn 90° about the vertical. At s = 0.25, the position is 0.75 × (0.30, 0.00, 0.20) + 0.25 × (0.40, 0.10, 0.10) = (0.225 + 0.100, 0 + 0.025, 0.150 + 0.025) = (0.325, 0.025, 0.175) m. The rotation has turned 0.25 × 90° = 22.5° about that same vertical axis. A quarter of the way along, in both senses.

0.75 × 0.30 + 0.25 × 0.40 = 0.325 m (x at s = 0.25)   ·   0.25 × 90° = 22.5°

What does "along the shorter rotation arc" add? Every rotation can be reached two ways round: turning 20° one way or 340° the other. With quaternions this shows up as a curious fact: q and −q describe the same rotation. The fix is a sign check. If the dot product of the two quaternions is negative, flip the sign of one, and the interpolation takes the short way. Example: a gripper yawed to 170° and a target yawed to −170° about z. As quaternions, (0, 0, 0.996, 0.087) and (0, 0, −0.996, 0.087): the dot product is −0.996 × 0.996 + 0.087 × 0.087 = −0.985. Negative, so flip: now the dot product is +0.985 = cos(10°), a half-angle of 10°, a 20° turn instead of 340°.

One more corner case appears in Appendix A.3: "nearly coincident orientations use normalized linear interpolation". When the start and end orientations are almost equal, SLERP's formula divides by the sine of a tiny angle, which is numerically fragile. Averaging the quaternions and renormalising gives essentially the same answer there, safely.

How many samples?

The adapter samples this path before solving IK. Table 3 lists the sampling settings for the YAM and ARX X5 arms as 0.005 m in translation and 0.035 rad in rotation. Read as the largest allowed step between neighbouring samples, they decide the count. A move of 0.10 m with a 90° turn (1.571 rad) needs 0.10 / 0.005 = 20 steps for the translation, but 1.571 / 0.035 = 44.9, so 45 steps, for the rotation. The rotation dominates, and the path gets about 45 samples. (The Morphi Kino samples differently: 0.003 m per axis and 0.02 per component of its six-number rotation representation.)

Inverse kinematics, one sample at a time

Inverse kinematics (IK) runs FK backwards: given a desired TCP pose, find joint angles that produce it. It is harder than FK because the answer can be missing (out of reach), unique, or one of several (elbow up or elbow down). GPT-Policy solves it at every sample, and seeds each solve with the previous one (Equation 3):

qk
The joint angles for sample k: one reference for the arm to track.
p̂k, R̂k
Sample k's target position and orientation, already converted into the solver's frame.
qk−1
The seed: the previous sample's solution. The first sample is seeded with the joints measured at planning time.

Why seed from the neighbour? Because neighbouring samples are 5 millimetres apart, so their solutions should be close too, and starting close keeps the solver on the same branch (the elbow does not flip halfway through a move). The paper notes the limit of this: "Sequential seeding does not impose a hard bound on the joint displacement between samples." On the YAM and ARX X5 arms, nothing forbids a sudden jump. The Morphi Kino adds that check: inter-sample joint changes must stay below 0.15 rad, about 8.6°.

Every solution is then graded. Take the joints the solver returned, run FK on them, and compare with the target (Equation 5):

ep,k
The position residual: where FK says the TCP really is, minus where it was asked to be. A 3-vector in metres.
eR,k
The orientation residual: the leftover rotation as a rotation vector (∨ turns the skew-symmetric matrix into a 3-vector). Its length is the leftover angle.
εp
Execution tolerance on position: 0.002 m on the YAM and ARX X5, 0.003 m on the Morphi Kino.
εR
Execution tolerance on orientation: about 1° on the YAM and ARX X5, 0.02 rad (1.15°) on the Morphi Kino.

A worked example with illustrative numbers. The target is (0.400, 0.080) in some plane and the solver lands at (0.4012, 0.0815). The residual vector is (1.2, 1.5) mm, and its length is √(1.2² + 1.5²) = √(1.44 + 2.25) = √3.69 = 1.92 mm. That is under 2 mm: accepted on every platform. Now suppose it lands (1.8, 1.2) mm off: √(3.24 + 1.44) = √4.68 = 2.16 mm. Rejected on the YAM and ARX X5, accepted on the Morphi Kino with its 3 mm tolerance.

√(1.8² + 1.2²) = 2.16 mm > 2 mm (YAM, ARX X5: reject)   < 3 mm (Morphi Kino: accept)

There are two different tolerances in play, and mixing them up is easy. The numerical stopping tolerances, 10−4 m and 5 × 10−4 rad, tell the solver when to stop iterating. The execution tolerances, 0.002 m and about 1°, decide whether the result is good enough to send to motors. The first is 20 times tighter than the second in position. A solver can stop early (hit its iteration limit, fail to "converge") and still be accepted if the execution check passes; the paper says neither adapter requires the solver's convergence flag in that case.

The backends differ. The ARX X5 uses its SDK's IK followed by damped least-squares (DLS) refinement: each iteration nudges the joints by Δq = J⊤(JJ⊤ + λ²I)−1e, where J is the Jacobian (how the TCP pose moves per unit of each joint) and λ is a damping constant. The damping keeps steps small near a singularity, a pose such as a fully stretched arm where some direction of motion becomes impossible. The price is accuracy: near a singularity DLS creeps and may leave a residual above tolerance. Joint bounds are supplied to the solver and clip each update. The YAM uses I2RT's kinematics and additionally checks the solution against the SDK's joint bounds. The Morphi Kino uses native numerical IK with an analytic fallback. Plan some moves and watch the gate.

Plan a path through the IK gate

A side view of a three-joint arm. The warm cross is the target the model asked for; its arrow is the fingers' direction. Tap or drag on the drawing (or use the sliders) to move it, then press Plan and move: the path is sampled, IK is solved at every sample from the previous one, and each residual is checked against the platform's tolerance. Try the far edge of the reach, a gripper pointing up, and the long way round.

0.400 m
0.060 m
−90°

A toy planar arm with illustrative link lengths and joint limits, solved live by damped least squares seeded from the previous sample, with the paper's stopping tolerances (10−4 m, 5 × 10−4 rad), sampling steps (YAM and ARX X5: 0.005 m, 0.035 rad; Morphi Kino: 0.003 m per axis, and 0.02 rad standing in for its 0.02 per rot6d component), execution tolerances and TCP speed limits (0.08 m/s on YAM and ARX X5, 0.03 m/s per axis on the Kino) from Appendix A.2 and Table 3. The real arms have six or seven joints in 3D; the gate logic is the same.

Three lessons fall out of the device. The edge of the reach is not a line: near full stretch, DLS converges slowly and leaves millimetres of residual, so a target that is technically reachable can still be refused. Orientation matters as much as position: a point comfortably inside the workspace may be unreachable with the fingers pointing a particular way, which is why the prompt warns that "an in-range point may still be unreachable at the required orientation." And the long way round sweeps the gripper through orientations the wrist cannot make; the shorter arc is not just shorter, it is often the only feasible one.

When any sample fails, the whole request is refused before anything is submitted, and the reason goes back to the model as ft: "IK rejection returns feedback before submission." The model's instructions for what to do next are specific: keep the required tool direction, compare approach positions, heights and permitted axial rotations, use check_path for uncertain alternatives, and do not tilt the gripper merely to make IK pass. A tilted grasp that the solver accepts can still be the wrong grasp for the task.

Check the forward kinematics, not the solver. FK of the returned joints is compared with the target, and on the ARX and YAM adapters a solution that passes this check is accepted even without the solver's convergence flag. The grade is kinematic: where the arm's model says the fingers would be, not a measurement of where they end up. The Morphi Kino applies its own thresholds (3 mm, 0.02 rad) plus a joint-step bound. Grading every backend by what its answer does, rather than by what its solver reports, is what lets different IK libraries sit behind one interface.
On an ARX X5 arm, one sample's IK solution lands 2.4 mm from its target with 0.5° of orientation error. What happens?

Chapter 5

Slow down until it is safe

Ruckig times the path, one stretch factor enforces three limits at once, and the arm's report closes the loop

After Chapter 4 the adapter holds a list of joint configurations, one per sample, every one verified. What it does not have is a clock. Should sample 17 happen 0.1 seconds after sample 16, or a full second after? Too fast and a motor is asked for more speed, acceleration or jerk than it can deliver. Too slow and the robot wastes minutes it has already been shown to spend freely. Timing is a separate problem from geometry, and the paper solves it in two stages.

Stage one: a progress profile

Picture a single number, the progress, that runs from 0 at the start of the path to 1 at the end. Every sample sits at some progress value. If you decide how progress changes over time, you have timed every sample at once, and all joints move in lockstep along the verified path. The paper uses Ruckig (Berscheid and Kröger, 2021), an open-source library for jerk-limited trajectory generation, to shape that progress profile: "Ruckig assigns timestamps through a scalar progress profile with zero endpoint velocity and acceleration." The move starts from rest, speeds up smoothly, slows down smoothly, and stops at rest.

Why jerk? Jerk is the rate of change of acceleration. A trajectory with unlimited jerk slams from zero acceleration to full acceleration instantly, which feels like a kick to the gearbox and makes a held object swing. Limiting jerk rounds off those corners. Velocity, acceleration and jerk are the first, second and third time-derivatives of position, and a careful controller limits all three.

Stage two: stretch until every joint is legal

The progress profile does not know about individual joints. A gentle, steady motion of the fingertips can demand a fast whirl of one joint, especially near the stretched-out poses of Chapter 4. So the adapter checks. It takes the timed joint samples and estimates each joint's velocity, acceleration and jerk by finite differences: the change between neighbouring samples divided by the time between them, then the change of that, then the change of that. It divides the largest absolute value of each derivative, over all samples and all joints, by its limit, giving three ratios rv, ra, rj. Then it stretches time (Equation 6):

rv
Largest joint speed divided by the speed limit. Above 1 means some joint is too fast somewhere.
ra
The same ratio for acceleration. It enters as a square root, for a reason derived below.
rj
The same ratio for jerk. It enters as a cube root.
α0
The smallest stretch that brings all three within limits. If nothing is over, it is 1 and nothing changes.
α
The stretch actually applied: a margin of 0.1% on top of α0 whenever stretching is needed.
τ′k = α τk
Every sample's timestamp is multiplied by α. The joint samples themselves do not change.

Where do the square root and the cube root come from? Stretching time by α means the new trajectory reaches at time αt whatever the old one reached at time t. Differentiate once and every velocity is divided by α. Differentiate again and every acceleration is divided by α², because the chain rule brings out one factor of 1/α per derivative. Jerk is divided by α³. So to pull a velocity ratio of rv down to 1 you need α ≥ rv; for acceleration you need α² ≥ ra, so α ≥ √ra; for jerk α³ ≥ rj, so α ≥ ∛rj. The smallest α that satisfies all three is their maximum. That is Equation 6, derived.

Now a worked example against the YAM arm's limits from Table 3: 0.6 rad/s, 2 rad/s² and 12 rad/s³ per joint. Suppose the timed samples peak at 0.9 rad/s, 3.2 rad/s² and 20 rad/s³ (illustrative numbers).

rv = 0.9 / 0.6 = 1.500   √ra = √(3.2 / 2) = √1.6 = 1.265   ∛rj = ∛(20 / 12) = ∛1.667 = 1.186

α0 = max{1, 1.500, 1.265, 1.186} = 1.500   →   α = 1.001 × 1.500 = 1.5015

Velocity binds. After stretching, the peaks become 0.9 / 1.5015 = 0.5994 rad/s (just under 0.6), 3.2 / 1.5015² = 3.2 / 2.2545 = 1.419 rad/s², and 20 / 1.5015³ = 20 / 3.385 = 5.91 rad/s³. A move that was timed at 2.0 s now takes 2.0 × 1.5015 = 3.00 s. Why the extra 0.1%? The paper does not say. A natural reading is a sliver of margin, so the binding joint lands strictly under its limit rather than exactly on it, safe from rounding.

Two honest caveats come straight from the paper. First, this "checks the sampled reference, not continuous-time physical jerk": the finite differences only see the samples, and whatever the motor does between them is not examined. Second, the Morphi Kino's Ruckig configuration shapes the timing but does not impose explicit joint acceleration or jerk limits after IK; it caps joint velocity at the smaller of 0.2 rad/s and each joint's URDF limit, and caps both the reference duration and actual playback at 30 s. Try the stretch yourself.

Stretch time until every limit holds

One joint's move, sampled at 100 Hz: position, then velocity, acceleration and jerk from finite differences. Dashed lines are the YAM limits. Shorten the nominal move or lengthen the joint's travel until a limit is broken, then watch Equation 6 stretch the clock. Switch the stretch off to see the violation it prevents.

1.20 rad
2.0 s

Limits are the YAM's from Table 3 (0.6 rad/s, 2 rad/s², 12 rad/s³); the finite differences, ratios and stretch follow Equation 6 exactly. The nominal profile is a smooth quintic that starts and ends at rest, standing in for Ruckig's own profile, which the paper does not print.

In code the whole rule is a few lines. This is the stretch of Equation 6 written from scratch; q holds the timed joint samples as rows, one column per joint.

pythonimport numpy as np

def time_scale(q, dt, v_max=0.6, a_max=2.0, j_max=12.0):
    """Equation 6. q: (N, joints) samples, dt: seconds between samples."""
    v = np.diff(q, n=1, axis=0) / dt            # finite differences
    a = np.diff(q, n=2, axis=0) / dt**2
    j = np.diff(q, n=3, axis=0) / dt**3
    r_v = np.abs(v).max() / v_max                  # worst sample, worst joint
    r_a = np.abs(a).max() / a_max
    r_j = np.abs(j).max() / j_max
    alpha0 = max(1.0, r_v, r_a ** 0.5, r_j ** (1 / 3))
    alpha = 1.0 if alpha0 == 1.0 else 1.001 * alpha0
    return alpha * dt                              # same samples, longer clock

Two arms, one clock

A bimanual request plans both arms before either is submitted. Their corresponding segments are then synchronised to the longer duration "by slowing the faster trajectory", so the arms arrive together. If the left arm's segment needs 3.0 s and the right arm's 1.2 s, both run for 3.0 s. The prompt tells the model the same thing in its own words: "Bimanual arrival times are synchronized by slowing the faster arm."

There are more speed limits than the joint ones. On the YAM and ARX X5 the TCP itself is capped at 0.08 m/s and 0.5 rad/s. A worked lower bound: a 0.20 m reach at 0.08 m/s takes at least 0.20 / 0.08 = 2.5 s before any acceleration or deceleration. On the Morphi Kino, 0.003 m per axis per sample at its 10 Hz reference rate implies 0.03 m/s per coordinate, so the same 0.20 m along one axis takes at least 0.20 / 0.03 = 6.7 s. These numbers are small, but a trial with 96 decisions pays them over and over.

Dispatch, settling and the report

The timed references then go to the robot. The ARX X5 submits them as a timestamped SDK trajectory. The YAM interpolates them for 100 Hz streaming. The Morphi Kino streams joints robot-locally at 10 Hz, one arm per call. On the YAM and ARX X5, each arm starts with a 0.1 s delay on its own clock (Table 3 marks this N/A for the Kino, which streams one arm per call), and every move ends with a final reference hold (0.12 s on YAM and ARX X5, 1 s on Kino).

Then the adapter waits for the arm to settle, and reports it separately from submission. On the YAM, settling means the joints stay within 0.03 rad of the reference over a window of at least 0.3 s and 10 distinct samples, with the encoder span at most 0.002 rad and span over time at most 0.05 rad/s. If that has not happened within 3 s, the result is reported as unsettled. On the Morphi Kino, a stable miss within a 0.08 rad tracking guard can come back as target_incomplete, for the model to assess.

All of that becomes ft, the paper's format F7: the previous result or a concrete rejection reason, the requested-versus-measured endpoint errors, per-arm execution and settling diagnostics, the fresh measured state, and fresh wrist and overhead images. The prompt tells the model how to read it: "Read a rejection before choosing another target. After execution, compare requested and measured pose, gripper command and measured opening, and visible object motion."

Three kinds of "accepted", and none of them is success. IK accepted means the fingers can reach every sample. Timing accepted means no joint exceeds its limits. Settled means the joints stopped moving. The paper is blunt about the gap: "reference acceptance or measured settling does not establish task success." Only fresh images, judged against the task, can say whether the towel is actually in the gripper.
Timed joint samples give rv = 0.8, ra = 2.25 and rj = 3.375. By Equation 6, what stretch α is applied?

Chapter 6

Watch a person once

One human video lifts two pickups from 0 of 3 to 2 of 3, without a single robot action label

Now the machine is built, and the experiments begin. The first family asks the most striking question in the paper: can a robot learn a strategy by watching a person? Not a robot, not a simulation, not a trajectory with numbers attached. Just a short first-person or third-person RGB video of human hands doing the job.

The two tasks are Pick Red Towel and Pick Up Notebook. Both are harder than they sound, for the reason Chapter 0 opened with. A towel is deformable: it has no fixed shape and lies flat, so there is no edge to pinch. A notebook is rigid but thin and flat, so there is no easy side to grip either. Both reward a trick. Section 4.2 describes the condition called None as the same instruction and the same live observations without the video; Human Video adds the video to the standard inputs. (The exact wording differs slightly, as we will see: the video condition's instruction also points the model at the video.)

What the video contains

Appendix C.1 is precise about what a human demonstration is here: a first-person or third-person RGB video of a person doing the task. Selected frames show the approach, the object interaction and the outcome, "without numerical robot states or actions". Each reference uses one camera view. Table 4 lists exactly which moments were kept.

DemonstrationKeyframesSelected timestamps (s)
Pick Red Towel80.000, 5.190, 7.257, 14.488, 18.622, 19.655, 21.722, 23.755
Pick Up Notebook60.000, 1.967, 3.433, 4.400, 5.867, 8.800
Remove Glue Cap (extra reference, not in Table 1)70.000, 2.100, 6.267, 7.833, 8.900, 10.467, 12.000

Look at the gaps. The towel's frames are 5.190, 2.067, 7.231, 4.134, 1.033, 2.067 and 2.033 seconds apart; they sum to 23.755, the span from the first selected frame to the last (the paper gives the selected times, not the length of the recording). The longest silence, from 7.257 s to 14.488 s, is 7.231 seconds in which the model sees nothing at all: 7.231 / 23.755 = 30.4% of the selected span. Whatever the hands did in that stretch, the model must infer from the frame before and the frame after. For the notebook the longest gap is the last one, 2.933 s of an 8.8 s selected span, 33.3%.

The model receives these frames in the paper's format F2: a line of text, "Historical human demonstration, t=7.257s", then the image, then the next line and the next image, closed by "End historical demonstration. Use the live scene for the current task." The towel instruction asks it to "watch the historical human demonstration frames and imitate the demonstrated grasping method to pick up the red towel with the robot's right hand in the current scene." The comparison condition's instruction, which the appendix calls the prepared goal-only comparison, is simply "Pick up the red towel with the robot's right hand." So the two conditions differ in two ways at once: the frames, and an instruction that tells the model to imitate the grasp in them.

The results

Human Video achieves 2 of 3 on both tasks, against 0 of 3 under None (Table 1). And the successful condition is also the cheaper one. Here are the changes, worked from the table:

Towel: 96.3 → 76.7 decisions = −20.4%   24.6 → 18.9 min = −23.2%

Notebook: 94.0 → 66.7 decisions = −29.0%   24.6 → 16.1 min = −34.6%

Each percentage is the drop divided by the starting value: for the towel's decisions, (96.3 − 76.7) / 96.3 = 19.6 / 96.3 = 0.204. Now divide minutes by decisions, as in Chapter 1: the towel goes from 15.3 to 14.8 seconds per decision, the notebook from 15.7 to 14.5. Each decision took about as long as before. The savings came from fewer decisions, not faster ones. The paper reads it the same way: "The combination of higher success and lower average execution costs suggests that the demonstration provides useful procedural guidance."

One more way to count, which the paper's averages make exact: because each condition's averages are over all three trials, total robot time is three times the average. The towel with the video spent 3 × 18.9 = 56.7 minutes and produced 2 successes, 28.4 minutes per success. Without it, 3 × 24.6 = 73.8 minutes produced none. Minutes per success is the number a factory would care about, and without context it cannot even be computed: no success was observed in three trials.

Where do the numbers come from?

Here is the remarkable part. The human video contains no robot numbers: no poses, no joint angles, no gripper openings. A human hand is not even the same kind of object as a parallel gripper. So every number in every tool request the model sent (every x, y, z, every quaternion, every gripper value) came from the robot's own live cameras and measured state. The paper states it plainly: "The human demonstration supplies no robot action labels; the agent generates robot-specific motion targets from current observations."

What the video supplies instead is strategy: which hand does what, in which order, with which kind of grasp. The paper's Figure 8 shows four frames of the towel demonstration labelled Start, Stabilize, Grasp and Lift: one hand steadies the towel, the other grasps it, then it comes up. The paper's Figure 4 shows "grasping methods and interaction sequences that may help constrain the agent's choice of strategy." Its conclusion is carefully worded: "These results are consistent with transferring an interaction strategy across embodiments." Cross-embodiment transfer means learning from a body different from your own, here from human hands to robot grippers. Try being the model: choose the tool calls yourself, with and without the demonstration.

Pick up the towel yourself

You are the policy. Each button is one tool call. First try without the demonstration; then switch the human video on and follow its strip. Every number your calls would carry comes from the live scene, which the adapter fills in here.

A toy with hand-written contact rules, not a physics simulation. The two-handed strategy (block with one hand, slide the other underneath, grip, lift) is the grasping method the paper writes out for its text-only towel variant (Appendix D.4); the four-frame strip is a cartoon of it, not the real keyframes, which are in Figures 4 and 8 (Figure 8 labels them Start, Stabilize, Grasp, Lift). Decision counts on the real robot are Table 1's.

Five calls do it, once you know the order: press, low beside the edge, slide under, close, lift. Without the strip, most people try the pinch first, the move that fails on a flat cloth. The demonstration does not tell you where the edge is; the live scene does. It tells you that the edge is the thing to go for, and that the other hand has a job. That division of labour, strategy from the reference and numbers from the scene, is the heart of cross-embodiment in-context learning.

Why did the real robot need about 77 decisions and not 5? The paper does not itemise them, but its prompts suggest much of it: each human-level step becomes several verified moves on hardware. The prompts insist on it: approach and align, inspect fresh images and state, then call set_gripper separately; stop to observe at contact, gripper changes and occlusion; after release, observe before withdrawing. A human-level step becomes a small conversation with the controller.

How strong is the evidence? Honestly, it is small. Three trials per condition means each trial is worth 33 percentage points, and 2 of 3 against 0 of 3 is a difference of two trials. The paper says so itself in its limitations: "The current evidence covers small task series under selected platform, model, and context conditions." It also notes a third towel instruction, a text-only procedural variant ("Block the towel with one hand, slide the other hand underneath it, and then grip the towel to lift it"), recorded but not reported in Table 1. So we cannot tell from this paper whether a written recipe would help as much as the video.

Chapter 9 shows two other models attempting the same towel with the same video, one run each. They reached 30% and 20% progress. The paper has no run of either model without the video, so we cannot say how much the video helped them.

Strategy transfers across bodies; numbers do not need to. A person's hands and a robot's grippers share almost nothing mechanically, yet the video was followed by better results, consistent with what it carried being the plan: steady here, grasp there, then lift. The controller and the live cameras supply all the metric detail. That split is why a short RGB video of hands can guide a robot at all.
In the Human Video condition, where do the coordinates in GPT-6 Astra's move_to requests come from?

Chapter 7

Between the keyframes

Recorded robot actions fill the gaps a video leaves, and give the highest observed success on both contact tasks

Picking up a towel forgives a lot. Unscrewing a bottle cap does not. The second family of experiments moves to tasks where millimetres and angles decide the outcome. Unscrew Bottle Cap "requires leaving the opened bottle standing securely": one arm must hold the bottle while the other twists, and a tilt at the wrong moment tips it over or leaves the cap on. Remove and Reinsert Plug "requires removing the plug and reinserting it into its original socket so that it remains fully seated after gripper release": prongs must line up with holes, and resting on the socket is not the same as being in it.

Here the demonstrations are robot demonstrations: a human operator teleoperated the robot while everything was recorded, including timestamped images, measured joint and end-effector states, and the motion and gripper commands. The bottle was recorded with the overhead camera only; the plug adds the two wrist views. Recording covers the whole manipulation through release and withdrawal, so the final outcome is visible too.

That enables a three-way comparison on the same demonstration. None has no demonstration. Robot Video gives the selected keyframes and their annotations. Robot Video + Action gives the same keyframes and annotations plus "time-aligned end-effector poses, gripper states, and action commands". Chapter 2 explained why the paper wants the images and non-action text held fixed across the last two. The two conditions share the selected images (Appendix C.2); only Video + Action adds measured states and action segments. For the bottle the recorded instruction is the same in both (a mode-specific prompt restricts Video to images and annotations). For the plug the recorded instruction wording also differs (Appendix D.4). And with three trials per condition, any difference is suggestive, not proof.

The results

TaskContextS/TDecisionsMinutesMinutes per success
Unscrew Bottle CapNone0 / 371.016.1none succeeded
Robot Video2 / 374.315.245.6 ÷ 2 = 22.8
Robot Video + Action3 / 354.717.953.7 ÷ 3 = 17.9
Remove and Reinsert PlugNone0 / 324.05.3none succeeded
Robot Video0 / 333.77.9none succeeded
Robot Video + Action2 / 348.310.832.4 ÷ 2 = 16.2

The first five columns are Table 1. The last is ours, worked the same way as in Chapter 6: three trials times the average minutes, divided by the number of successes. For the bottle with video alone, 3 × 15.2 = 45.6 minutes bought 2 successes, 22.8 minutes each. With actions, 3 × 17.9 = 53.7 minutes bought 3, 17.9 minutes each.

Action references gave the highest observed success on both tasks. On the plug, they were the only condition that succeeded at all. But look at the cost columns: they do not simply fall. From video alone to video with actions, the bottle's average time rose from 15.2 to 17.9 minutes and the plug's from 7.9 to 10.8. Against no demonstration at all, the plug's time roughly doubled (5.3 to 10.8 minutes) and its decisions climbed from 24.0 through 33.7 to 48.3. The paper flags this: action references "do not consistently reduce costs; these averages include failed trials."

Why would a failing condition look cheaper? One possible reason (the paper does not report per-trial costs): a failed trial can end early. A plug trial that goes wrong early and ends costs few decisions. A trial that gets the plug out, back over the socket, in, released, verified and pressed home costs many. It is not the whole story, though: on the bottle, Robot Video (2 of 3) averaged more decisions, 74.3, than Video + Action (3 of 3), 54.7. What the averages do guarantee is that they mix successes and failures, which is why we added the last column.

Why the numbers help

The paper's explanation starts with an observation from selected bottle runs (Figure 7): executions using robot video with action references "more closely match the demonstration in requested orientations and measured support posture than those using video alone". Figure 7 measures two things: the tilt of the gripper that supports the bottle over the course of each run, and the difference between the requested orientation and the demonstration's at two keyframes: the left-hand grasp (KF 1), and the bottle tilt with the right hand's approach to the cap (KF 3). Smaller differences with actions.

Then the hypothesis: "the denser temporal information in action references reduces ambiguity about motion between video keyframes." Table 5 gives the densities: 205 retained action samples for 13 video keyframes in bottle opening, 131 samples for 14 keyframes in plug reinsertion. A keyframe says where the gripper was at one moment. Between two keyframes, many different motions could connect them: a straight line, an arc, a pause and a twist. "Whereas sparse keyframes leave intervening motion to be inferred, action references supply intermediate commanded poses and gripper transitions that help constrain this inference."

Worked from Table 5: the bottle's 13 keyframes bound 12 segments, so 205 / 12 = 17.1 action samples per segment. The plug's 14 keyframes bound 13 segments: 131 / 13 = 10.1 per segment. Every one of those samples is another pinned point on the path the model has to imagine. See what that does to the space of possible paths.

Collapse the gaps between keyframes

A demonstration's gripper height over time. Purple marks are keyframes; the thin lines are motions the model could reasonably infer from the evidence it has. With video alone, keyframe heights are judged from images, so each carries doubt. Switch to Video + Action and slide in the recorded samples: every sample pins the path.

205

Keyframe and sample counts are Table 5 (205 samples, 13 keyframes, 12 segments for the bottle; 131, 14, 13 for the plug). The trajectory, the keyframe timing, the image-only doubt and the candidate paths are illustrative: we draw smooth paths through the evidence with room to vary in proportion to each gap. The paper's own evidence for the effect is Figure 7, on selected runs.

How the records are built

The alignment is careful, and Appendix C.4 explains it. Camera images, measured joint and end-effector states, gripper openings and issued commands all go on one shared timestamp axis. Each keyframe takes the nearest measured state within 0.1 s; other camera views are matched to the overhead image within the same tolerance. "Missing measurements remain absent", and the prompt repeats it: "Null means missing, never zero." The paper is candid that this is "timestamp matching, not hardware-synchronized exposure."

Each keyframe image at time ti is paired with its measured state and with the action segment leading to the next keyframe at ti+1. For the plug, that links the grasp image and the gripper state to the withdrawal commands that follow. Within a segment the compiler keeps the endpoints, about one action sample per second, and both sides of every gripper-command change, so the exact moment of closing or opening is never lost between samples. For the plug that sampling cut the raw record from 1,405 samples to 131, keeping 9.3%; for the bottle, from 212 to 205, keeping 96.7% (Table 5).

The format (F3b) is compact. A header lists the columns: sample index, video frame index, times, then for each arm the joint targets, the end-effector target in xyzw order, the gripper command and the measured gripper, and small time offsets from the overhead capture. In each row, "=" repeats the previous row's value. Chapter 3's normalisation example came from the first of these rows.

And the instructions for using them are strict. In Video + Action mode the model may "use recorded positions, orientations, and gripper events as numerical planning references after checking the source embodiment, base frame, TCP reference, quaternion order, and current object alignment." It must normalise rounded quaternions, remember that "commanded poses are not measured poses", and never "stream old absolute joint commands". The bottle may sit two centimetres from where it stood in the demonstration. The records guide reasoning; they are "not direct trajectory replay". In Video mode, by contrast, the model is told to infer geometry from images and "not invent recorded numerical poses."

Cheaper per success, not per trial. Recorded actions raised the bottle from 2 of 3 to 3 of 3 and the plug from 0 of 3 to 2 of 3, while the average trial got longer. Judge them by minutes per success (22.8 down to 17.9 on the bottle, and from no observed success to 16.2 on the plug), and in these three trials they were cheaper per success, before counting the cost of recording a teleoperated demonstration.
A robot video shows every stage of the bottle task in 13 keyframes. Why can adding the recorded actions still help?

Chapter 8

Goals, memories and a partner

A picture of the end, a record of what already happened, and a person across the table all count as context

Videos show a procedure. The last three context families show something else. A goal image shows only the end. Self-interaction history shows the robot its own past. Online human interaction brings a partner whose moves and gestures change what the task even is. The paper tests each on two tasks, with one condition per task, and every one of the six scored 3 of 3.

One caution before the details: none of these six tasks was run without its context. There is no None row for them in Table 1, so they demonstrate that the system can use these kinds of context, not how much each one helped. For goal images the paper adds only a hedge: "In preliminary qualitative comparisons, we observe closer matches to the desired layout than in runs without a target image."

A goal image: the end, not the way

Arrange T Shape asks the robot to arrange coloured blocks on the table into a T. Arrange Fruit asks it to arrange four fruits to match a layout. Each gets a single goal image, an overhead screenshot or an operator's photograph of the desired arrangement, placed in the input before the live observation (format F4). The instruction reads: "Observe the target image and arrange the blocks on the table into the T shape shown in the image. Match the target image as closely as possible in color, relative position, and spacing."

Why a picture rather than a sentence? Try writing the sentence. "Put this block horizontally at the top centre, that one vertically below it touching its middle, the third continuing the stem, with gaps of about a finger's width," and then say which block is which. Every object needs an identity, a position relative to the others, an orientation and a spacing. The paper's point is that one image "jointly" specifies object identity, relative position and spacing, conveying spatial requirements "that can be cumbersome to describe precisely in words." The image does not prescribe a single action; the model still plans every move. Both tasks: 3 of 3, averaging 66.7 decisions and 15.8 minutes for the T and 49.0 decisions and 12.4 minutes for the fruit.

Self history: remembering what you did

Lemon to Pink Plate: "Explore the current scene, locate the pink plate, and place the lemon from the table onto the plate." The catch is that the plate is not visible at the start. Movable Exploration: continue searching for a Sprite bottle "by changing the observation position or safely moving obstacles", then place it in the yellow basket that holds a strawberry toy. This is the mobile task, the one the paper calls mobile exploration; of its three platforms, only the Morphi Kino has a mobile base, along with an articulated waist and head.

Both run under Self History: the context retains the agent's earlier observations, actions and execution outcomes within the task (format F5). When exploration changes what is visible (a moved obstacle, a new viewpoint), past views and outcomes become the memory of what was found where. Both scored 3 of 3.

The behaviours surprised the authors. "Surprisingly, the agent autonomously removes the towel to uncover and locate the pink plate before placing the lemon." Nobody told it the plate was under a towel, or to move the towel. During mobile exploration "it also actively avoids obstacles along its route while searching for the target." The paper reads both as "consistent with high-level reasoning about intermediate subgoals: changing the scene to obtain missing information and choosing a feasible route to continue the search."

The costs tell a second story. Divide minutes by decisions again:

Lemon: 8.1 × 60 = 486 s ÷ 35.3 = 13.8 s per decision   Movable: 25.53 × 60 = 1,531.8 s ÷ 40.33 = 38.0 s per decision

Similar numbers of decisions, but each mobile decision took nearly three times as long. The paper does not break the time down; the obvious difference is that a mobile robot drives its base and looks again, which takes time that choosing a tool call does not. The paper draws the lesson: this highlights "the distinction between decision count and elapsed time." Counting decisions counts the requests the model made; counting minutes counts the elapsed time, including everything the robot and the world had to do between requests.

A partner across the table

The last family puts a person in the loop. In Pointed Fruit Pickup, the human points at a fruit, and the robot must "repeatedly pick up whichever fruit the human points to and place it onto the plate", wait when no gesture has been made, and stop only at an OK hand gesture. In Tic-Tac-Toe, the robot plays Green and moves first against an on-site human playing Blue, and must "move only when it is Green's turn and the human's hand has left the board."

The tic-tac-toe prompt also spells out a strategy: "Prioritize an immediate win, then block an immediate opponent win, then choose a move that creates a double threat or at least preserves a draw." A double threat (a fork) is a move that creates two lines of two at once, so the opponent can block only one. With perfect play, tic-tac-toe is a draw, so "at least preserves a draw" is a real guarantee. The paper counts both wins and draws as success, and reports that "in the observed Tic-Tac-Toe games, the agent also selects optimal moves for the current board state." Both tasks: 3 of 3.

What makes this in-context learning is the interaction history. Section 3.2's record of observations, tool requests, results and operator feedback provides "context for tracking what the human and robot have each done, whose turn it is, and how far the task has progressed." Whose turn it is cannot be read from a single photograph if the human's hand is still hovering over the board. Play a game against the robot's rules.

Play Blue against the robot

The robot (Green in the paper, marked G) has moved first. Tap a square to place your piece (Blue in the paper, marked B): your hand is now over the board. The robot will not move until you press Take your hand away. The history panel is the interaction record the robot reads, with its reason for each move.


      

The move priorities (win, block, double threat, otherwise keep at least a draw), the turn rule and the history's contents follow the paper's prompts (Appendix D.4, format F5). The robot's choice here is computed by an exact game search, standing in for GPT-6 Astra's reasoning; the paper reports that the agent chose optimal moves in its observed games. The physical pick-and-place of each piece is not shown.

Two rules are working at once, and the paper separates them on purpose: "The game instruction separates legal turn-taking from move selection; a favorable board position does not override the requirement to wait for the human hand to leave." Turn-taking is a safety and etiquette rule, grounded in the history (count the pieces each side has placed) and in the live image (is the hand gone?). Move selection is strategy. A robot that finds a winning move while a hand is still on the board must wait.

A worked lower bound shows how much physical work hides behind each move. A game averaged 69.7 decisions (Table 1). Green moves first, so in any game Green places at most 5 pieces. Therefore, on average, a game spent at least 69.7 / 5 = 13.9 decisions per placement, counting everything in between: waiting for the human's hand to leave, observing the board, finding a piece, approaching, verifying, closing, lifting, moving, lowering, releasing, withdrawing. The paper does not break the decisions down, but the strategic choice itself is a small part of that budget.

Memory is just more context. The robot's own past is fed back exactly like a demonstration: text records and images in the window. The model has no other memory. That is why the history policies of Chapter 2 matter: what the adapter keeps is what the robot remembers, whether that is where the pink plate turned out to be or whose turn it is.
Movable Exploration used 40.33 decisions in 25.53 minutes; Lemon to Pink Plate used 35.3 decisions in 8.1 minutes. What does that show?

Chapter 9

What each decision costs

Three models attempt the same towel, and every move costs about a quarter minute

Every chapter so far has used one model. A fair question is whether GPT-Policy's results belong to the framework or to GPT-6 Astra. The related-work section names several general-purpose models that support reasoning, tool use and multi-step execution: GPT-6 Astra, Claude Fable, Kimi K3 and GLM-5.3. The discussion compares three of them on the task we know best, Pick Red Towel, with the same human video, through GPT-Policy's shared tool interface (provider adapters differ in how they manage history, Appendix B).

The comparison is Table 2, and it is honest about its size. Each row is an individual run, not an average over trials. It reports three things per run: task progress, the degree of completion reached (not a success rate); run time in minutes; and estimated token usage, in millions of tokens, where a token is the unit of text (and image) a model reads or writes, roughly a short word.

ModelDemonstrationTask progressRun time (min)Tokens (M)
GPT-6 AstraNone55%24.6312.047
GPT-6 AstraHuman Video100%15.854.956
Fable 5.1Human Video30%13.272.386
Kimi K3Human Video20%11.731.735

Start with GPT-6 Astra's two rows. The video took progress from 55% to 100%, and the paper summarises the savings as "approximately 35.6% shorter run time and 58.9% lower estimated token usage". Let's reproduce both.

(24.63 − 15.85) ÷ 24.63 = 8.78 ÷ 24.63 = 0.356   → 35.6% shorter

(12.047 − 4.956) ÷ 12.047 = 7.091 ÷ 12.047 = 0.589   → 58.9% fewer tokens

Notice that tokens fell much further than time. The paper does not say why. A likely reason (ours, not the paper's): every decision re-reads the whole window, meaning the instructions and tool schemas, the pinned references, the retained history, and three fresh camera images. If so, total tokens grow roughly with the number of decisions times the size of the window, and the window itself grows as history accumulates. A run that wanders for more decisions pays twice: more reads, and longer reads. The demonstration's own frames sit in the window at every decision too, yet in this pair of runs the one with them used far fewer tokens overall.

We can probe that reading with the paper's own numbers, as a derived rate rather than anything the paper reports. Divide each run's estimated tokens by its run time:

12.047 M ÷ 24.63 min = 0.489 M per minute (no demonstration)

4.956 M ÷ 15.85 min = 0.313 M per minute (with the human video)

(0.489 − 0.313) ÷ 0.489 = 0.176 ÷ 0.489 = 0.36   check: (1 − 0.589) ÷ (1 − 0.356) = 0.411 ÷ 0.644 = 0.64

So the video run was not only shorter: each of its minutes also consumed about 36% fewer estimated tokens. The check line gets there a second way, from the paper's two percentages alone: keeping 41.1% of the tokens over 64.4% of the time leaves 64% of the rate, a 36% drop. That is consistent with a smaller window, fewer wandering exchanges piling up in the history for the compiler to carry. It is not proof. These are two single runs, the paper does not say what the estimate includes, and the minutes include arm motion and settling during which the model reads nothing at all.

Now the other two models. Fable 5.1 and Kimi K3, given the same video, "use fewer resources but reach only 30% and 20% progress, respectively." Their runs were shorter and used fewer estimated tokens than GPT-6 Astra's video run (13.27 and 11.73 minutes against 15.85; 2.386 and 1.735 million tokens against 4.956), and they got less of the task done. It is tempting to divide one by the other and crown a winner. Resist it. Progress is a completion degree, not a linear scale on which 20% is a fifth of the work, and one run per model says little about the model. The paper puts it directly: "These examples do not establish a reliable model ranking."

What the paper does offer is qualitative. Figure 8 lines up selected frames from each run with a one-word label under each frame. The demonstration reads Start, Stabilize, Grasp, Lift, and GPT-6 Astra's video run reads the same four words, ending with the towel in the air. GPT-6 Astra's run without context reads Start, Contact, Adjust, Adjust, Adjust, Adjust, Reorient, No lift. Fable 5.1's run reads Contact, Adjust, Adjust, Withdraw, Retry, Reorient, No lift after its start; Kimi K3's reads Approach, Contact, Adjust, Adjust, Withdraw, Retry, No lift. The caption draws the lesson: "Human video context helps GPT-6 Astra complete the task with fewer unnecessary intermediate actions compared with no context and other models." And notice the run scored at 55% progress: its strip still ends in No lift. Progress measured how far the attempt got; the towel stayed on the table.

Four runs on one towel

Table 2's four individual runs. Choose what to compare: progress, run time, estimated tokens, or tokens per minute of run time. Switch on Table 1's averages to see how single runs sit against three-trial averages for the same model.

Runs are Table 2: individual runs, progress is completion degree, tokens are estimates. Table 1 averages are over three trials of GPT-6 Astra (24.6 and 18.9 minutes; success 0 of 3 and 2 of 3). Tokens per minute is our derived rate, not the paper's; it compares runs, not models. The paper says these runs "do not establish a reliable model ranking."

Three cautions keep this comparison in its place. First, these are single runs. GPT-6 Astra's video run here took 15.85 minutes, while its three-trial average in Table 1 was 18.9; runs vary, and one run per model cannot rank models. Second, task progress is not success. The run that reached 100% progress belongs to a condition that succeeded in 2 of 3 trials, not 3 of 3. Third, the harness around each model need not be identical. Appendix B says history handling is provider-dependent: the Codex refresh keeps all accumulated text but omits older live images, while the ARX Claude Messages adapter summarises older exchanges, applies its own motion rejection rules and reviews completion in a separate session. The paper does not say which adapter each Table 2 run used.

What would a fair ranking need? The paper's own framing points the way. Several trials per model and condition, so that 2 of 3 and 3 of 3 can be told apart from luck; the same execution budget and termination rules for all; and success rates next to progress, tokens and minutes, since a model that stops early will always look cheap. Table 1 has that structure for GPT-6 Astra alone. Table 2 is a first look across models, and it is labelled as one.

It is also worth being precise about what "estimated token usage" covers, and here the paper is silent: it does not define what the estimate includes. Most likely it covers what the model read and wrote across the whole run. If so, in a loop like Equation 1, where each reply is one short JSON object, most of it is likely reading: the same instructions, schemas and references re-read at every decision, plus the growing history and fresh images. That is why the compiler's trimming policy from Chapter 2 is not a detail. Every image it keeps is paid for again at every later decision.

Moves take time

Step back to the tempo of the whole system. Across Table 1, elapsed time per decision sits between about 12 and 20 seconds: 11.7 s for tic-tac-toe, 12.3 s for the bottle with video, 15.3 s for the towel with no demonstration, 19.6 s for the bottle with actions. The one outlier is mobile exploration at 38.0 s. That is roughly four decisions a minute for a tabletop task.

The paper's third takeaway names the consequence: "VLM Agents can already generate useful actions, but each decision can still be slow and expensive. Specialized VLA/WAM models may therefore have an edge in fast, low-level control, while VLM Agents focus on reasoning, adaptation, and replanning." A WAM (world-action model) is a trained policy that also predicts future observations, like Flex-π. Such models issue many actions per second from a network running on a local GPU. A general VLM called through an API every quarter minute cannot catch a slipping towel.

The natural architecture is a split, and it is one of the paper's six future directions. "Hierarchical systems such as Hi Robot separate contextual reasoning from low-level execution." A deliberative System 2 agent, like GPT-Policy's VLM, would choose goals, strategies and recoveries; a fast System 1 controller, "such as a VLA policy", would handle pose refinement and bimanual coordination. The names borrow a psychologist's distinction between slow deliberate thought and fast automatic response. The open question the paper poses is precise: "when local feedback should trigger replanning."

Cost has a second dimension beyond time: money and compute. Twelve million estimated tokens for a single towel attempt with no demonstration is a large bill for one pickup. In this one pair of towel runs, the video run used 58.9% fewer estimated tokens. That is one run pair, not a general cost rule: Table 1 shows that context does not consistently reduce cost, since the bottle and plug trials with recorded actions took longer on average than those with video alone. For a system that pays per token, the compiler's choices of what to pin and what to drop are budget decisions as much as learning decisions.

Episode length drives the bill. In this one pair of towel runs, the video's frames sat in the window at every decision, yet the video run used 58.9% fewer estimated tokens overall and, by our derived rate, about 36% fewer per minute. If the window is re-read at every decision, the number of decisions and the history they pile up weigh more than any single reference. Whether that holds across tasks, two single runs cannot say.
GPT-6 Astra's towel run used 12.047 M tokens without the video and 4.956 M with it. What does the paper's 58.9% describe?

Chapter 10

Good plans, tricky execution

We gather what the paper shows, weigh what it admits, see where it points next, and write the whole loop as code

You can now read the GPT-Policy paper and explain every part of it: what robotic in-context learning means, how Equation 1 closes the loop, what the compiler keeps, why the model speaks in absolute TCP poses, how each request is sampled, solved, checked and timed, and what each family of context bought on real robots. Let's gather it, then look hard at what it did not solve.

The one-paragraph summary

GPT-Policy connects an off-the-shelf vision-language model, GPT-6 Astra, with frozen weights, to robot arms through a shared loop: a context compiler assembles the instruction, labelled camera images, measured state, task references (goal images, demonstration keyframes, recorded actions) and interaction history; the model returns one JSON tool request per decision; a constrained controller resolves the target, samples a straight-line and SLERP path, solves inverse kinematics at every sample with residuals checked against 2 mm and about 1° (YAM and ARX X5; 3 mm and 0.02 rad on the Morphi Kino), stretches the timing until joint velocity, acceleration and jerk are within limits, dispatches it, and reports errors, settling and fresh images back. On real robots, one human video lifted towel and notebook pickup from 0 of 3 to 2 of 3 with fewer decisions and minutes; aligned action references lifted bottle opening to 3 of 3 and plug reinsertion to 2 of 3; goal images, self history and live human interaction each supported 3 of 3 on two tasks. The same trials exposed slow, expensive decisions, unreliable verification of success, and repeated collisions between the two arms.

The paper's three takeaways

The collision problem

The most concrete failure in the paper is also the most sobering: "We repeatedly observed collisions between the two arms during manipulation." Chapters 4 and 5 showed why nothing stopped them. The Cartesian planner checks reachability, residuals, joint bounds and timing, and synchronises the two arms' durations, but "the planner does not check collisions". The prompts warn that "IK acceptance does not certify clearance between camera housings, arms, or objects", and Appendix B adds that "runtime diagnostics and provider-specific checks do not establish collision-free motion or physical task success." Each arm's path is verified alone, and the Cartesian planner does not check the two paths against each other. See what a check across both would catch.

Check both arms together

A top view of two arms reaching across a shared table. Each path alone passes every check the paper's planner makes. Slide the right arm's target toward the left arm's side, or delay its start, and press Plan and run. The strip underneath is the distance between the arms over time. Then switch on the safety layer the paper proposes.

0.00 m
0.0 s

The collisions are the paper's observation (Section 5), and the safety layer is its first future direction; the paper reports no distances. The arm geometry, link thickness, 4 cm clearance and timing here are illustrative, and the arms move with joint interpolation synchronised to the longer duration, as Appendix A.3 describes.

The paper's proposal is exactly this, plus more: "a dedicated safety layer that checks both arms' planned trajectories together, monitors separation and contact during execution, and can interrupt unsafe commands independently of the VLM." Independence matters. A model that reasons about a task every fifteen seconds cannot be the thing that stops two arms from meeting in the next half second. The authors also ask that collisions, near misses and safety interventions be reported explicitly, next to task success.

Six directions

  1. Physical safety for VLM-driven manipulation. The layer above, plus collision-aware planning and uncertainty-aware execution limits, evaluated alongside success.
  2. Contact-aware harnesses for reliable grasping. Expose grasp stability and object motion after contact; slip detection, force-aware limits and local recovery, "without requiring the agent to reason through every correction."
  3. System 1 / System 2 for fine-grained action. Pair the deliberative agent with a fast controller such as a VLA; decide when local feedback should trigger replanning (Chapter 9).
  4. Mobile manipulation through active perception. Building on SayCan's affordance-grounded skill selection: choose where to look and stand as well as how to grasp, with persistent spatial memory and coordinated base and arm control.
  5. Compositional context for long-horizon tasks. Building on trajectory prompting as in ICRT: compose subskills from several demonstrations in a new order, keep completed subgoals, discard obsolete context, rather than replaying a whole trajectory.
  6. In-context adaptation to physical dynamics. Use recent interactions to update predictions of friction, compliance and object response before contact failures pile up, starting from tactile dynamics models and video-conditioned world-action models such as Zero-WAM.

The limits, in the authors' words

The numbers that matter

QuantityValueWhy it matters
Weights updatednone, θ fixedAll adaptation travels through the input
Towel, human video0/3 → 2/3; 96.3 → 76.7 decisions; 24.6 → 18.9 minConsistent with strategy transfer; no robot action labels
Notebook, human video0/3 → 2/3; 94.0 → 66.7; 24.6 → 16.1 minFewer decisions, not faster ones
Bottle: none, video, + actions0/3, 2/3, 3/3Dense references for contact tasks
Plug: none, video, + actions0/3, 0/3, 2/3Only the action references succeeded
Action samples per keyframes205 for 13 (bottle), 131 for 14 (plug)Pins the motion between frames
Goal image, self history, human interaction3/3 on each of six tasksCapability shown; no no-context control
Reference caps24 keyframes, 48 images, 1,280 pxThree views cap a demo at 16 keyframes
IK execution tolerance0.002 m, about 1° (YAM, ARX X5); 0.003 m, 0.02 rad (Kino)Refuse before moving, explain why
YAM joint limits0.6 rad/s, 2 rad/s², 12 rad/s³Equation 6 stretches time to meet them
Time per decisionabout 12 to 20 s (38 s mobile)Why a fast System 1 is still needed
Towel run, video vs none35.6% shorter, 58.9% fewer tokensOne run pair; not a general cost rule

The whole loop, as code

Everything above, compressed into a from-scratch sketch of one episode: compile the context, call the frozen model, parse one tool request, sample the path, check IK residuals, time it, dispatch, and feed the result back. The robot, camera and model calls are placeholders; the geometry and the checks are real.

pythonimport json, numpy as np
# GPT-Policy's closed loop in plain Python. Frozen: the VLM. Nothing below trains anything.
EPS_P, EPS_R = 0.002, np.deg2rad(1.0)          # execution tolerances (YAM, ARX X5)
STEP_P, STEP_R = 0.005, 0.035                    # sampling steps: metres, radians

def compile_context(task, refs, history, obs, prev):
    msgs = [SYSTEM_P1, TOOL_SCHEMAS]                      # frames, xyzw order, one JSON reply
    if refs:                                              # pinned for the whole episode
        msgs += ["HISTORICAL DEMONSTRATION."] + refs.interleaved() + ["END HISTORICAL DEMONSTRATION."]
    msgs += history.trimmed(keep_images=3)                # old images dropped, old text kept or summarised
    envelope = {"instruction": task, "state": obs.state, "extra": {"env_step": obs.step}}
    if prev is not None: envelope["previous_result"] = prev  # absent at the first decision
    return msgs + [json.dumps(envelope)] + obs.images      # left wrist, right wrist, top

def unit(q):
    q = np.asarray(q, float); n = np.linalg.norm(q)
    if not np.isfinite(n) or n < 1e-9: raise ValueError("quaternion cannot be normalised")
    return q / n

def slerp(q0, q1, s):                                  # Eq. 2, xyzw, shorter arc
    d = q0 @ q1
    if d < 0: q1, d = -q1, -d
    if d > 0.9995: return unit((1 - s) * q0 + s * q1)      # nearly equal: normalised lerp
    th = np.arccos(d)
    return (np.sin((1 - s) * th) * q0 + np.sin(s * th) * q1) / np.sin(th)

def sample_path(p0, q0, p1, q1):
    ang = 2 * np.arccos(min(1.0, abs(q0 @ q1)))
    n = max(1, int(np.ceil(max(np.linalg.norm(p1 - p0) / STEP_P, ang / STEP_R))))
    return [((1 - k / n) * p0 + k / n * p1, slerp(q0, q1, k / n)) for k in range(1, n + 1)]

def ik_checked(arm, samples, q_seed):                   # Eqs. 3 to 5
    qs = []
    for k, (p, q) in enumerate(samples):
        q_seed = arm.ik(p, q, seed=q_seed)                 # seeded by the previous sample
        p_fk, q_fk = arm.fk(q_seed)
        e_p = np.linalg.norm(p_fk - p)
        e_r = 2 * np.arccos(min(1.0, abs(q @ q_fk)))         # leftover rotation angle
        if e_p > EPS_P or e_r > EPS_R:
            return None, f"IK residual {e_p*1000:.1f} mm, {np.degrees(e_r):.2f} deg at sample {k+1}"
        qs.append(q_seed)
    return np.array(qs), None

def adapter(req, robot):                                 # the robot-tool interface E
    name, args = req["name"], req["arguments"]
    if name == "set_gripper":
        return robot.gripper(args["positions"])             # requested vs measured opening
    plans = {}
    for side, target in robot.targets(name, args).items():   # None means hold
        if target is None: continue
        x = target["pose_xyzquat"]                          # complete pose or nothing
        p1, q1 = np.array(x[:3]), unit(x[3:])
        p0, q0 = robot.fk_measured(side)                   # the path starts where the arm is
        qs, why = ik_checked(robot.arm(side), sample_path(p0, q0, p1, q1), robot.joints(side))
        if why: return {"rejected": why}                    # before anything is submitted
        plans[side] = qs
    if name == "check_path": return {"feasible": True}      # no command sent, no clearance claim
    timed = robot.time_and_sync(plans)                     # Ruckig progress + Eq. 6 stretch, slow the faster arm
    return robot.dispatch(timed)                           # endpoint errors, settling, fresh images

def run_episode(task, refs, robot, vlm, budget=150):
    history, prev = History(), None
    for t in range(budget):
        obs = robot.observe()
        req = json.loads(vlm(compile_context(task, refs, history, obs, prev)))   # a_t = (u_t, v_t)
        if req["name"] in ("done", "give_up"):
            prev = robot.review(req, obs)                  # optional separate-session review
            if prev["accepted"]: return req, t + 1           # a human still labels success
        else:
            prev = adapter(req, robot)                     # f_t: executed, rejected, or feasible
        history.append(obs, req, prev)
    return {"name": "budget_exhausted"}, budget

Where it sits in the field

Why did GPT-Policy's controller not prevent the inter-arm collisions the authors observed?

Keep going

The takeaway. On small task series, a frozen vision-language model used context to choose better strategies, but only inside a harness that chooses what it reads, restricts what it can say, and checks every request for reachability, IK residuals, joint limits and timing (but not collisions) before a motor turns. GPT-Policy shows the first half working: a video of human hands, or a few hundred recorded samples, was followed by better strategies and more successes. It is just as clear about the second half: turning a good plan into precise contact, verified success and safe motion is still the controller's problem, not the model's.

Now press Present or Teach and explain, out loud and from memory, why a human video with no robot numbers could still raise towel pickup from 0 of 3 to 2 of 3, and why recorded actions were the only condition that succeeded on the plug. If you can, you own this paper. Then go back to the robot and give it each kind of context in turn.

Based on "In-Context Robot Learning with VLM Agents" by Dongzhou Cheng, Taoran Yi, Ye Fang, Xingwu Zhang, Fan Feng, Yixuan Li, Gengxiong Zhuang, Rongze Wang, Shuai Yang, Wei Song, Weizhi Xue, Minyan Wu, Jie Gui, Jiaqi Wang and Tong Wu (Morphi Robot and collaborators, 2026)
Read the paper · Project page · Code · Back to Veanors