Do as I Do

Rebuild a hand and object from an ordinary video, then search a physics simulator for a matching robot hand motion. On 655 everyday clips it tracks the object in simulation 71% of the time; the prior best managed 25%.

Learn how one algorithm turns an everyday video of a hand using an object into a robot hand motion that tracks the human's object in simulation and has been played back on real hardware.

Pick a rung of the ladder and watch the robot hand copy the clip. Then we build, piece by piece, what each rung adds and why the videos make it hard.

You need what a 3D coordinate is and the idea of a simulator that steps physics forward in time. We build the rest from zero.

Copy the clip

Ready

Each rung copies the same kind of clip. Watch how the attempts fail, and which failures go away as you climb.

Success rates and mean position/rotation errors are Table 3 of the paper (a clip succeeds when its mean position error is under 0.1 m and its mean rotation error under 0.5 rad); the paper measures these in simulation against the reference, not on the physical robot. The scene is a toy: one hand seen from the side, one pick, carry and place. Each clip succeeds at random at the paper's rate for the chosen rung, so a short run wobbles around the paper's number. How a clip fails is illustrative, chosen from the failure each component was built to fix (Section 3.2, Figure 4).

Chapter 0

Watching is not doing

Shows why robot hands starve for data, and why the videos we already have are so hard to use

A cooking video starts. A hand picks up a whisk, steadies a bowl, and beats three eggs in fifteen seconds. You watched it once, and if someone handed you a whisk you could do roughly the same. A child does this all day long: watch, then copy.

A robot cannot. Give the same video to a robot with a five-fingered hand and it has nothing it can execute. The video is pixels. The robot needs a list of motor commands, one number for each of its joints, many times a second, that makes a real whisk move the way the video's whisk moved.

The paper gives that gap a name. Data we collect by watching someone else is observational: it records what happened, from the outside. Data a learner collects by doing, its own actions together with what those actions caused, is experiential. Children close the gap between the two without effort. For today's robots it is close to an insurmountable barrier, because they learn almost entirely from experiential data.

Where does experiential data come from today? From two places. In teleoperation a person steers the robot directly, through a glove, a joystick, or a scaled copy of the robot's arm, and every command is recorded. In simulation the robot explores a virtual world and learns from a reward, a number that says how well it did. The paper names the bottleneck of each. Teleoperation is limited by the operator's expertise, by the cost of running the rig, and by its mechanical transparency: how faithfully the operator's hand motion reaches the robot's hand. Simulation is limited by the work of designing many different environments and a reward function for every task.

Both bottlenecks bite hardest on dexterous manipulation: manipulation with a multi-fingered hand, the kind that holds a pen in a writing grip or squeezes a sponge. The robot hand used throughout this paper, the Sharpa Wave, has 22 degrees of freedom: 22 independent numbers a controller must set at every instant. Teleoperating 22 joints transparently is hard. Hand-writing a reward for "whisk like a person" is harder.

Meanwhile the internet holds an enormous number of videos of human hands doing exactly these things. Almost all of them are monocular RGB: one ordinary colour camera, with no depth sensor, no second viewpoint, no gloves and no markers. If those videos could be turned into robot commands, the data problem would change shape. That is the question the paper asks: how do we convert the accessible, observational data of humans into experiential data for robots?

The question is old. The paper traces "Do as I do" back to a 1970 MIT "copy-demo", to Kang and Ikeuchi's robot instruction by human demonstration (1997), and to Efros, Berg, Mori and Malik's action synthesis and retargeting (2003), which used that very phrase. Across fifty years the problem has split into the same two parts:

  1. Reconstruction. Recover what the human did, in 3D: where the hand was, how its fingers bent, which object it held, what shape that object is, and how it moved.
  2. Retargeting. Turn that human behaviour into commands for a robot whose body is different, sometimes very different.

Past successes made the problem tractable by assuming things. Only pick-and-place behaviours. Only objects that had been 3D-scanned in advance. Depth cameras, 3D sensors or tracked hand keypoints from specialised hardware. Each assumption helps, and each one shrinks the pool of usable video. The general case, a single RGB video of a hand doing anything with any object, is where most of the world's data sits.

Two recent advances make that general case approachable, and the whole paper stands on them. The first is 3D vision foundation models: large networks, trained on vast data, that take one RGB image and return its depth (MoGe), a complete 3D object (SAM 3D), or a 3D hand. Run them on every frame and you get a 4D estimate, 3D plus time, of the hand and the object. The second is GPU-parallel physics simulators such as MuJoCo Warp and Isaac, which step thousands of copies of a scene at once. That is fast enough to find a robot's motion by trying thousands of candidate motions and keeping what works; the paper reports inferring dexterous hand actions from 4D hand-object states in minutes.

Do as I Do (Paliwal, Etukuru, Liang, Abbeel, Shafiullah and Malik, UC Berkeley, 2026) chains the two. Step one reconstructs the hand and the object from monocular RGB and tracks them through time. Step two retargets them onto a robot hand by optimization inside a physics simulator. Neither step assumes a grasp type, a behaviour class or an object category; any rigid object will do. Follow one clip through it.

Follow one clip through the pipeline

One frame of a clip, a hand lifting a mug. Step through the ten stages. Each stage says what goes in, what comes out, and whether anything is trained.

Stage order, models and numbers (13 pose numbers, 25 candidates, 20 tracked points, 22 hand joints, 200 Hz simulation, 50 Hz robot) are from the paper's Sections 3 and 4 and Appendices A, B and D. The MANO sizes (778 vertices, 21 keypoints) describe the standard hand model the paper cites, not a number printed in the paper. The drawing is a toy frame.

Look at what came in and what went out. In: a clip of T frames, each an H × W × 3 array of colour bytes, filmed by a person (egocentric) or at a person (exocentric). Out: for every moment of the clip, the robot hand's pose and 22 joint targets that, played in a physics simulator, move a copy of the object along the path the human's object took, and then, through inverse kinematics, the arm joint angles that put that hand there on a real robot.

The paper describes no training of its own. Every model in step one (SAM 3, MoGe, HaWoR, SAM 3D, BootsTAPIR, GeoCalib) is an existing released model. The new ideas are how SAM 3D is steered at inference time, how the pieces are aligned, and how the search in step two is made robust to noise. Step two is optimization, not learning: it finds one motion for one clip. That is why the method can be pointed at a video that was uploaded this morning.

The paper claims four contributions, and the chapters follow them. A two-step algorithm that reconstructs and retargets monocular RGB video onto multi-fingered hands. A hand-object reconstruction that beats the state of the art on standard metrics and handles ego and exo video, from internet clips to the output of video generators. A retargeting procedure that improves on existing scalable, physics-aware retargeting with three new components built for noisy references. And robot data that plays on a real dexterous hand and arm: to the authors' knowledge, the first pipeline that goes from an internet video to real dexterous rollouts.

Here is the road. Each chapter explains one part of the instrument above.

  1. Four data sources, one ladder: where the paper sits among its neighbours, and why an internet clip is the hardest source of all.
  2. The hand is the ruler: a metric 3D hand from one camera, and the alignment that puts the object at the hand's scale. The hero's blue ghost path is built here and in the next chapter.
  3. Tracking with a generator: a single-image 3D generator turned into a video tracker by steering its denoising.
  4. Scoring a reconstruction: F-scores, Chamfer distance, Table 2 and 450 human votes.
  5. Across the embodiment gap: retargeting as sampling-based optimization in simulation. This is the hero's first rung, Annealed sampling.
  6. Three fixes for a noisy reference: the hero's other three rungs, Warmup, Perturbation and the Transition reward.
  7. From internet clip to robot hand: the hero's dataset switch, and the real two-armed robot.
  8. Most clips are not data: the filtering playbook, 2,000 clips in, 83 out.
  9. What it adds up to: the numbers, the core loop as code, and the limits the authors admit.
Even a perfect cooking video cannot be used directly as robot training data. Why?

Chapter 1

Four data sources, one ladder

Places the paper among its neighbours by what each method reconstructs, how it retargets, and which videos it can use

Do as I Do is not the first attempt to learn dexterity from human video. Before we build the method, it pays to meet the neighbours, because the way each one is assembled decides which videos it can use. The paper's related work answers three questions: what do people do with human video, how do they rebuild the hand and object, and how do they move the result onto a robot?

Three ways to use a human video

The first family extracts a prior by pretraining. A prior is knowledge a learner brings before it sees its own task. Some works pretrain a visual representation, an image encoder whose features make later robot learning easier (VIP, LIV, R3M). Some pretrain dexterous policies on large human video collections (VideoDex, EgoScale, EgoVLA, Being-H0, EgoVerse). Some pretrain a forward dynamics model, a network that predicts what happens next, also called a world model (DreamDojo, and a world model for hand-object interactions from Goswami and colleagues).

The second family computes structured priors: an affordance (where and how an object can be grasped, as in ZeroMimic and dexterous functional grasping) or a flow (how points in the image are going to move, as in Track2Act and MimicPlay).

The third family builds 3D reconstructions for retargeting: rebuild the hand and the object in 3D, then map the motion onto a robot (DexMV, H2Sim2Robot, VideoManip, DexMan, DexImit and others). Do as I Do belongs here.

The difference between the families is what comes out. A representation or a prior is a head start: the robot still has to finish learning with data of its own. A retargeted reconstruction is robot data: a trajectory the robot can execute, with the object moving as it should. That is the stronger product, and it is also the hardest to make, because it needs a complete 3D description of what happened.

What does "complete" mean? Three things per frame. The object's shape, usually as a mesh: a list of 3D points (vertices) joined into small triangles (faces) that tile the surface. The object's pose: where it is and which way it faces, six numbers in all, three for position and three for orientation. That is why trackers are called 6-DoF (six degrees of freedom). And the hand's pose: where the wrist is and how every finger joint is bent. Every method in the third family gets these three things somehow. The somehow is where they differ.

Table 1, read as a ladder

The paper's Table 1 lists four close neighbours and itself. Read each row left to right: how it gets the object's shape and pose, how it retargets, and which data sources it was shown on.

The four data-source columns are the interesting part. The paper lists the table "in rough order of difficulty"; read left to right, these four columns climb the same way. Self is data the authors filmed themselves. Gen is video from a generative video model. Ego is egocentric video, filmed from the person's own head. Internet is anything found online. Climb the ladder and watch who is still standing.

Climb the data ladder

Rows are the five methods of Table 1; columns are the four data sources, easiest on the left. Pick a source to see what it takes away from you. Pick a method to see how it is built.

Checkmarks and components are Table 1 as printed. What each source takes away is our summary of the paper's Sections 1, 2 and 4.5. An empty cell means the method was not shown on that source in its paper, not that it could never work there.

Two things stand out. First, the object's mesh moves from something you must own (a LiDAR scan of the physical object) to something you infer from pixels (MeshyAI, TRELLIS, SAM 3D). A generated video has no physical object behind it, and an internet clip's object is in someone else's kitchen, so only a mesh from pixels can reach those columns. Second, four of the five methods track with FoundationPose or its successor. That matters because, as the paper argues and shows in Chapter 4, model-based trackers "tend to lose pose lock, drift, or fail to re-acquire the object once visual evidence degrades", which is exactly what happens when a hand closes around an object in a blurry, shaky clip.

How people rebuild hands and objects

Reconstructing 4D hand-object interaction (3D shape and pose, through time) from one RGB camera splits along three axes: the hand, the object, and the two together.

Hands are the easy half, relatively. Human hands share one structure: the same bones, the same joints, a narrow range of sizes. The paper cites the standard parametric hand model for this (Romero, Tzionas and Black's "Embodied hands", the MANO model), and that shared structure is what lets modern hand trackers (HaMeR, WiLoR, HaWoR) stay robust to motion blur, occlusion and low resolution. Chapter 2 is about this half.

Objects are the hard half, because everyday objects are endlessly diverse. Progress has come from two directions. Image-conditioned 3D generative models (One-2-3-45, TRELLIS, SAM 3D) produce a whole object from one picture, with strong priors over shape and pose that let them cope with occlusion and low resolution. Model-based 6-DoF trackers (FoundationPose, Any6D) follow a known or jointly reconstructed mesh from frame to frame. The paper's criticism of both: they were largely validated on clean lab video and struggle on noisy in-the-wild video.

Then there are joint methods that reason about hand and object together: in the lab (HO) and in the wild (IHOI, HORSE, MCC-HO). The recent video versions are narrow: some work only on egocentric video (WHOLE), and some are category-conditioned (diffusion-guided reconstruction, G-HOP), which means they assume the object comes from a fixed list of categories. A whisk that is not on the list is a problem.

Do as I Do deliberately goes the other way, with a modular decomposition: HaWoR for the hand, SAM 3D for the object's mesh from a single image, and a new SAM 3D-based tracker for the object's pose over time. One way to see the appeal: each module is a foundation model trained on the largest data available for its own sub-problem, whereas a joint model needs paired 3D hand-and-object data, which is scarce. The price is that the modules disagree about scale, and Chapter 2 shows how they are reconciled.

How people move the result onto a robot

Retargeting maps human hand motion onto a robot hand with a different geometry: different finger lengths, different joints, sometimes a different number of fingers. It comes in two kinds.

Kinematic retargeting solves a geometric problem: find robot joint angles whose fingertips (or joints, or task-space points) land where the human's did (mink, PyRoki, AnyTeleop, geometric retargeting). It works on the robot's configuration alone and ignores forces between hand and object, so its results can put a finger through the object, slide a fingertip along it, or hold it in a grasp that would slip.

Dynamics-aware retargeting adds physics. It searches for hand motions whose simulated consequences match the reference: the simulated object must actually move the way the human's did. It is done either by RL, training a policy to track a reference-based reward (ManipTrans, DexMachina, H2Sim2Robot, Dexplore), or by sampling-based optimization, trying many candidate motions in parallel and keeping the good ones (Yang and colleagues, ExoStart, and SPIDER).

Here is the assumption almost all of them share: a clean reference. Their human hand-object trajectories come from motion capture (MoCap), multi-camera or marker-based systems that deliver clean, ground-truth poses. A reference rebuilt from one internet video is nothing like that. It jitters, it has temporal discontinuities where a tracker jumped, and the hand and object can be badly misaligned, a finger floating a centimetre from the object it is supposed to hold.

The new problem is retargeting a reference you cannot trust. Reconstruction from monocular video will always be noisy, so the retargeter has to be robust to that noise, not just faithful to the reference. The hero measures the gap directly: the same annealed-sampling retargeter succeeds on 72% of clean OakInk2 motion-capture tasks and on 25% of references rebuilt from everyday video (Table 3). Chapters 5 and 6 are about closing it.

So the ladder has two rails. On the reconstruction rail, the question is whether you can get a mesh and a pose track from pixels alone, through blur and occlusion. On the retargeting rail, the question is whether you can turn a noisy, misaligned reference into physically valid robot motion. The next three chapters climb the first rail; Chapters 5 and 6 climb the second.

Which assumption, shared by most prior dynamics-aware retargeting methods, does Do as I Do give up?

Chapter 2

The hand is the ruler

Recovers a near-metric hand from one camera, then uses it to put the object at the right distance

Hold your thumb out at arm's length and you can cover the Moon. One eye, or one camera, cannot tell a small thing nearby from a big thing far away. Everything in this chapter follows from that one fact, so let's make it exact.

An ordinary camera is well described as a pinhole camera: light from a point in the scene travels in a straight line through one point, the camera's centre, and lands on the image. Put the camera at the origin, looking along the Z axis. A point at lateral position X and depth Z lands at pixel column u:

u
The pixel column where the point appears in the image.
f
The focal length, in pixels: how strongly the camera magnifies. Part of the camera's intrinsics.
X
How far the point is to the right of the camera's axis, in metres.
Z
The point's depth: its distance along the camera's axis, in metres.
cx
The principal point, the column where the camera's axis meets the image, roughly the image's centre.

X and Z appear only as a ratio. Double both and u does not move. With numbers: a camera with f = 600 px sees a mug 8 cm wide at 0.60 m. Its width in the image is 600 × 0.08 / 0.60 = 80 px. A mug twice as wide, 16 cm, at twice the distance, 1.20 m, spans 600 × 0.16 / 1.20 = 80 px. Same pixels, same picture. This is scale ambiguity, and no amount of cleverness applied to a single image removes it. Something outside the image has to supply a ruler.

600 × 0.08 / 0.60 = 600 × 0.16 / 1.20 = 80 px (half the size at half the distance looks identical)

Three models, three kinds of guess

Do as I Do runs three preprocessing steps on every clip before any tracking.

  1. Segmentation with SAM 3. SAM 3 is a segmentation model: for each frame it returns a mask, a binary image marking which pixels belong to a thing. Here it produces two per frame, one for the hand and one for the object.
  2. Depth and intrinsics with MoGe. MoGe is a monocular geometry model: from one image it predicts a pointmap, an image in which every pixel stores the 3D point (X, Y, Z) it sees, along with the camera's intrinsics. The paper chose MoGe for a specific reason: SAM 3D was trained on pointmaps from MoGe, so feeding it MoGe pointmaps at inference keeps it on familiar ground.
  3. An object mesh with SAM 3D. From a single frame and the object's mask, SAM 3D generates the object's 3D mesh. Chapter 3 is about this model.

MoGe learns to guess scale from context: counters, doors and cups have typical sizes, and a network trained on enough images absorbs them. Its authors call its output metric. But a guess from context can be off by a good fraction, and for a robot that has to close its fingers around a mug, a mug placed 10 cm too far away is a missed grasp. So the paper does not take MoGe's scale at face value. It uses a second opinion that is much harder to fool: the hand.

The hand

The hand tracker is HaWoR, a model for world-space hand motion reconstruction from egocentric video. The paper found it robust enough on noisy internet footage to use as it is.

What does a hand tracker output? Modern trackers describe the hand with a parametric hand model, and the standard one, which the paper cites, is MANO. MANO is a small function. It takes a handful of shape parameters (how long and thick this person's fingers are) and pose parameters (the rotation of the wrist and of each finger joint), and returns a complete hand surface: a mesh of 778 vertices, plus 21 keypoints (the wrist and four points along each finger, ending at the fingertip). So per frame the hand estimate is a few dozen numbers of pose and shape, a global position and rotation, and from them a 778 × 3 array of vertex positions and a 21 × 3 array of keypoints.

"World-space" is the other half of HaWoR's name. In egocentric video the camera rides on a moving head. A naive tracker reports the hand relative to the camera, so a still hand appears to move whenever the person glances around. HaWoR separates the camera's motion from the hand's, and reports the hand in a fixed world frame.

Now the key property. Human hands vary in size far less than objects do. A mug can be an espresso cup or a beer stein; a whisk can be a travel whisk or a baker's balloon whisk. An adult hand is always roughly a hand. A hand model therefore carries a built-in ruler: if the fingers span so many pixels, and fingers are about so long, the hand must be about so far away. That is the pinhole equation run backwards, Z = f · (real size) / (size in pixels). The paper puts it plainly: it treats the hand reconstruction's scale as ground truth, and calls the hand's space near metric.

Play with the two opinions. The camera below sees a hand and a mug. MoGe's depth is off by a scale you choose. HaWoR's hand is sized like a real hand.

Borrow the hand's ruler

Left: the scene from above, camera on the left, depth to the right. Right: what the camera sees. Change MoGe's depth scale and watch the camera's picture stay exactly the same. Then switch the alignment on and off.

×1.20
12 cm
0%

A toy scene. Here MoGe's error is a single global scale, so alignment removes it exactly and the only error left is the hand's own. Real depth errors are not a pure scale, which is why the paper solves the object's translation by least squares, frame by frame. The alignment formulas are the paper's (Section 3.1, Appendix A).

Notice two things. The camera's picture never changes, whatever MoGe's scale: every point MoGe misplaces slides along its own ray from the camera, which is exactly why the image cannot catch the mistake. And alignment works by trusting the ratio MoGe got right while replacing the scale it got wrong. MoGe is good at "the mug is a bit behind and to the right of the hand". It is unreliable about "the hand is 0.72 m away". The hand model is reliable about exactly that.

The alignment, step by step

Here is the paper's procedure, run on every frame. It starts with three centroids, a centroid being the average position of a set of points.

  1. The visible hand, in hand space. For each pixel in the hand's mask, cast a ray from the camera and record where it first hits the HaWoR hand mesh. Average those hits: cHhand. Only the visible side counts, because MoGe can only see the visible side too; we want to compare like with like.
  2. The same hand, in MoGe space. Average MoGe's 3D points at exactly those pixels (the rays that struck the mesh): cMhand.
  3. The object, in MoGe space. Average MoGe's 3D points over every pixel of the object's mask: cMobj.

The scale ratio compares the two opinions about the hand's depth, and the object's target is the hand's position plus MoGe's hand-to-object offset, rescaled:

k
The scale ratio: how much MoGe's distances must be stretched (k > 1) or shrunk (k < 1) to agree with the hand.
zHhand
The depth of the visible hand's centroid according to HaWoR, trusted as near metric.
zMhand
The depth of the same hand pixels according to MoGe's pointmap.
objtarget
Where the object's visible centroid should be, in the hand's near-metric space.
cHhand
The visible hand's centroid from HaWoR: the anchor everything is measured from.
cMobj
The object's centroid from MoGe's points over the object mask.
cMhand
The hand's centroid from MoGe's points over the same pixels as cHhand. The difference in brackets is MoGe's hand-to-object offset.

The last step moves SAM 3D's mesh to that target. SAM 3D places the mesh in its own camera-frame units, with a translation t. The paper keeps the mesh's orientation fixed and optimizes a single number, a translation scale s, so the mesh's visible centroid cmesh + st lands as close as possible to the target. Here cmesh is the centroid of the mesh vertices that project inside the object mask, with the translation left out. Minimising the squared distance over s has a closed form:

s*
The best translation scale: how far to slide the mesh along its viewing ray.
t
SAM 3D's camera-frame translation of the mesh. Scaling it moves the object along the ray from the camera through it.
cmesh
The centroid of the visible mesh vertices, without the translation. ⊤ means a dot product: multiply matching coordinates and add.

This is ordinary one-variable least squares: the error is a parabola in s, and its bottom sits where the error vector is perpendicular to t. The paper describes the effect in words: it "slides the object along its viewing ray to the depth that best matches the target while preserving its recovered orientation".

Let's run it with numbers (illustrative, in metres). HaWoR puts the visible hand's centroid at cHhand = (0.020, 0.050, 0.600). MoGe puts the same pixels at cMhand = (0.016, 0.040, 0.480) and the mug at cMobj = (0.060, 0.070, 0.520).

  1. Scale: k = 0.600 / 0.480 = 1.25. MoGe has everything 20% too close.
  2. MoGe's offset from hand to mug: (0.060 − 0.016, 0.070 − 0.040, 0.520 − 0.480) = (0.044, 0.030, 0.040).
  3. Rescaled: 1.25 × (0.044, 0.030, 0.040) = (0.055, 0.0375, 0.050).
  4. Target: (0.020 + 0.055, 0.050 + 0.0375, 0.600 + 0.050) = (0.075, 0.0875, 0.650).
  5. SAM 3D's mesh has t = (0.05, 0.06, 0.50) and cmesh = (0.01, 0.01, 0.02). The target minus cmesh is (0.065, 0.0775, 0.630).
  6. Numerator: 0.05 × 0.065 + 0.06 × 0.0775 + 0.50 × 0.630 = 0.00325 + 0.00465 + 0.315 = 0.3229.
  7. Denominator: 0.05² + 0.06² + 0.50² = 0.0025 + 0.0036 + 0.25 = 0.2561.
  8. s* = 0.3229 / 0.2561 = 1.261. The mug's centroid lands at cmesh + 1.261 t = (0.073, 0.086, 0.650), within 2 mm of the target on every axis.

s* = 0.3229 / 0.2561 = 1.261 (slide the mesh 26% farther along its ray)

It cannot hit the target exactly because it has only one degree of freedom: it can move along the ray, not sideways. That is a feature. Sideways position is what the image measures well, so the alignment leaves it to the image and fixes only depth, the one thing the image cannot see.

Finally the trajectory is rotated so gravity points down, using GeoCalib, a single-image calibration method that estimates the direction of gravity by geometric optimization. This matters because the retargeting happens in a physics simulator, where gravity pulls along a fixed axis. A trajectory tilted by the camera's pitch would ask the robot to set a cup down on a sloping table. After all of this, the output is an aligned, near-metric 4D hand-object trajectory: for each frame, the hand's MANO pose and the object's mesh pose, in one consistent space.

When does the ruler fail? When there is no hand to measure: the playbook in Chapter 8 drops 41 of 187 candidate clips because the hand or the object leaves the frame. When MoGe's relative arrangement is wrong, the object lands in the wrong place next to the hand; the paper's Figure 4 shows retargeting succeeding anyway on a reference whose incorrect depth estimation caused poor alignment, which is one reason the retargeter must tolerate noise. And in one of the limitations the authors list: a single camera cannot always tell touching from in front of. A finger that occludes a mug looks the same whether it presses on it or hovers a centimetre closer to the lens. The reconstruction cannot settle contact, so contact is left for physics to settle, in Chapter 5.

Why does the alignment take the object's depth from the hand rather than from MoGe's pointmap alone?

Chapter 3

Tracking with a generator

Turns a single-image 3D generator into a video tracker by steering each denoising step toward the previous frame

We have a hand in every frame. For the object we need two things: its shape, once, and its pose in every frame. The obvious tool is a model-based tracker such as FoundationPose: give it the mesh and it follows the object frame to frame. The paper's complaint, which Chapter 4 backs with numbers, is that such trackers lose the object when a hand closes over it or the camera blurs, and do not find it again.

SAM 3D has the opposite temperament. It was trained to rebuild objects from single images with occlusion and at low resolution, so a half-hidden, blurry mug is exactly its kind of input. The trouble is that it is a generative model: it does not compute one answer, it draws a sample from a probability distribution over answers.

Concretely, SAM 3D learns the joint distribution pθ(xs, xp | c) over an object's shape xs and its pose xp, given a conditioning input c (one image and the object's mask). "Joint" means shape and pose are drawn together, as one sample. The pose block is small: 13 numbers (the paper's Dp = 13), which include a translation and a unit quaternion, four numbers of length one that encode a 3D rotation.

Run SAM 3D independently on every frame and you get a different mesh in every frame, and a pose sequence with no temporal coherence: each frame's pose is a fresh draw, free to jump. A tracker needs the opposite: one shape, and poses that change only as much as the object really moved.

How a flow model draws a sample

SAM 3D generates with flow matching. Start with pure noise x0 drawn from a standard Gaussian. The model learned a velocity field vθ(xt, t, c): at any point and any time t between 0 and 1, which way to move. Follow it from t = 0 to t = 1 and the noise turns into a sample. The path it was trained on is a straight line from noise to data:

xt
The partly denoised sample at flow time t: shape and pose blocks together.
t
Flow time, 0 (pure noise) to 1 (clean sample). It has nothing to do with the video's frame index.
x0
The starting noise, drawn from a Gaussian with zero mean and identity covariance.
x1
A clean sample: one shape and one pose consistent with the image.

In practice the integration is done in small Euler steps of size Δ: x ← x + Δ · vθ. The paper's appendix counts T = 25 such steps per sample. Each step is one pass through SAM 3D's large network.

The key observation: fix the shape, steer the pose

Because shape and pose live in one shared latent and are denoised together, the paper observes that you can fix the shape block and obtain an updated pose entirely at inference time. Pick an anchor frame, generate the shape there, and call it x̄s. For every later frame k, we want a pose from pθ(xpk | xs = x̄s, ck), biased toward the previous frame's pose xpk−1, since real objects do not teleport.

Conditioning a big generator exactly on part of its output is intractable: it would mean integrating over every possible pose in a continuous 6-DoF space with a very large network. The paper instead borrows a trick from diffusion inpainting (RePaint) and score-based models. In a flow model, the noisy version of a known target at time t is simply its interpolant: the point on the straight line between the sample's own starting noise and the target. So if we wish the pose block would end at xpk−1, we know where it "should" be at every t along the way, and we can nudge it there, step by step.

Each Euler step becomes two moves: the model's own free Euler update, then a blend toward the target interpolant. This is the paper's Equation 1, written for the pose block (the shape block is identical, with x̄s and αs):

xpt
The pose block after this step, at flow time t.
αp
The guidance strength, between 0 and 1: 0 lets the model denoise freely, 1 forces the pose onto the target's path.
x + Δ v
The free Euler update: where the model alone would move the pose, using its velocity evaluated at the previous point and time.
zpref(t)
The target interpolant: where the pose would be at time t if it were heading straight for the previous frame's pose from this sample's own noise.
εp
The pose block's initial noise, the same noise this sample started from.
xpk−1
The pose chosen for the previous video frame.

Why blend toward the interpolant and not toward the previous pose itself? Because the network only knows how to denoise inputs that look like its own partly noisy samples. At t = 0.25 a real sample is still three quarters noise. Pasting in a clean pose there would hand the network something it has never seen. The interpolant carries the right amount of noise for its time, so the nudge keeps the sample on the model's own road.

Let's run four steps by hand, in one dimension (illustrative numbers). The pose starts at its noise, ε = −1.0. The previous frame's pose is 0.6. The model, looking at this frame's image, is sure the pose is 0.9, so its velocity always points straight there: v = (0.9 − x) / (1 − t). Steps have Δ = 0.25 and αp = 0.7.

  1. t from 0 to 0.25. v = (0.9 + 1.0) / 1 = 1.900. Free update: −1.0 + 0.25 × 1.900 = −0.525. Interpolant: 0.75 × (−1.0) + 0.25 × 0.6 = −0.600. Blend: 0.3 × (−0.525) + 0.7 × (−0.600) = −0.1575 − 0.420 = −0.578.
  2. 0.25 to 0.5. v = (0.9 + 0.578) / 0.75 = 1.970. Free: −0.578 + 0.25 × 1.970 = −0.085. Interpolant: 0.5 × (−1.0) + 0.5 × 0.6 = −0.200. Blend: 0.3 × (−0.085) + 0.7 × (−0.200) = −0.166.
  3. 0.5 to 0.75. v = (0.9 + 0.166) / 0.5 = 2.131. Free: −0.166 + 0.533 = 0.367. Interpolant: 0.25 × (−1.0) + 0.75 × 0.6 = 0.200. Blend: 0.3 × 0.367 + 0.7 × 0.200 = 0.250.
  4. 0.75 to 1. v = (0.9 − 0.250) / 0.25 = 2.599. Free: 0.250 + 0.650 = 0.900. Interpolant: 0 × (−1.0) + 1 × 0.6 = 0.600. Blend: 0.3 × 0.900 + 0.7 × 0.600 = 0.690.

0.3 × 0.9 + 0.7 × 0.6 = 0.69 (the image's opinion, weighed against the last frame)

The answer lands between what the image says (0.9) and where the object was (0.6), and in this toy it is exactly their weighted average: at the last step the free update lands on the model's answer and the interpolant lands on the previous pose. So αp is a trust slider, the same idea as a Kalman filter's gain: how much to believe the past over the new evidence. In a real model the early steps matter as well. They decide which answer the model commits to, and when the picture is ambiguous (a symmetric mug seen side on could be facing either way), a nudge toward the last frame in those early steps is what stops the pose from flipping.

How strong should the nudge be?

For shape the answer is easy. The paper handles rigid objects, whose shape never changes, and reports that any fixed αs in [0.9, 1] works well: pin the shape almost completely.

For pose, a fixed number is a bad compromise. Too strong, and the pose cannot keep up when the object really turns, the paper's "over-rigidity". Too weak, and every ambiguous frame is a chance for a "spurious flip". The paper's answer is to let the video decide, frame by frame, using point tracking.

A point tracker follows individual physical points across frames: give it a pixel on the mug's handle in frame k and it reports where that same bit of handle is in frame k + 1. The paper uses BootsTAPIR. For each pair of consecutive frames it samples 20 points inside the object's mask, tracks them, keeps the ones that stay inside the next frame's mask, and fits a 2D rigid transform (a rotation plus a translation, no stretching) to them in closed form by SVD, the standard least-squares fit of one point set onto another. Out comes the object's in-plane rotation Δθk between the frames, and its centroid's translation. Then:

αp(k)
The pose guidance strength used for frame k.
0.1
The floor: even during a fast turn, keep a little pull toward the last pose.
0.7
The ceiling: an object that is not turning is held firmly to its last pose.
0.09
How fast trust in the past falls per unit of rotation. The paper does not print the unit of Δθ.
Δθk
The in-plane rotation between frames k − 1 and k, estimated from the tracked points.

With numbers: a still object, Δθ = 0, gets 0.7 − 0 = 0.70. A rotation of 3 gets 0.7 − 0.27 = 0.43. A rotation of 5 gets 0.7 − 0.45 = 0.25. A rotation of 8 gives 0.7 − 0.72 = −0.02, below the floor, so it is clamped to 0.10. The paper notes this costs one extra offline tracking pass per video, and its ablation (Chapter 4) shows it consistently helps.

Try it on a clip with a still stretch, a fast turn, another still stretch and a slow turn.

Steer the tracker

The dashed line is the object's true angle over 40 frames. Each grey dot is one sampled pose; the warm line is the pose the tracker keeps. A symmetric object looks identical turned half a circle, so the model sometimes proposes the flipped pose (upper band). Change the guidance, the way one pose is picked per frame, and how degraded the video is.

50%

A toy model standing in for SAM 3D: for each frame it proposes the true angle or its flipped twin, plus scattered outliers, more of both as the video degrades. The sampler, the blend of Equation 1 over 25 Euler steps, the adaptive rule αp = max(0.1, 0.7 − 0.09|Δθ|) with Δθ read as degrees, N = 25 candidates and clustering by distance with outlier rejection follow the paper. The errors shown are this toy's, not the paper's.

Four lessons are in that device. With no guidance, poses scatter and flip. Fixed 0.7 is smooth but over-rigid: when the object turns 8 degrees a frame, the pose falls behind and may never catch up. Fixed 0.1 follows the turn but wobbles on the still stretches. Adaptive guidance is firm when the points say "still" and loose when they say "turning". And push the degradation slider to the end: when the flipped pose becomes as likely as the true one, even guidance can lock onto the wrong one, and then faithfully tracks the wrong answer. Guidance keeps a good track good; it cannot rescue a bad start.

Picking one pose out of 25

Guided sampling is still sampling, so for each frame the paper draws N candidates, all sharing the fixed shape, and must choose one. The principled choice is the model's own log-density of each candidate: how probable the model thinks it is. For a flow model that density is exactly computable by the instantaneous change-of-variables formula (the paper's Equation 2): run the ODE backwards from the candidate to find the noise it came from, and add up how much the velocity field expands or compresses space along the way (the trace of its Jacobian).

The paper prices this out. The trace needs one backward pass per pose coordinate at every Euler step. With Dp = 13 pose numbers, T = 25 steps and N = 25 candidates:

N · T · (1 + Dp) = 25 × 25 × 14 = 8,750 forward and backward passes per frame

Generating those 25 candidates took N · T = 25 × 25 = 625 forward passes. Scoring, each of its 8,750 passes also carrying a backward pass, the paper estimates at roughly two orders of magnitude above generation, per frame, which is prohibitive for video.

So it clusters. It measures the distance between two candidate poses with a weighted distance on SE(3), the space of 3D rigid poses (a position plus an orientation):

d
The distance between two candidate poses.
wt, wr
Weights that trade metres against radians. The paper does not print their values.
‖ti − tj‖
The straight-line distance between the two translations.
2 arccos |⟨qi, qj⟩|
The geodesic angle: the smallest rotation that takes one orientation to the other. The absolute value is there because q and −q encode the same rotation.

A worked angle: qi = (1, 0, 0, 0) is "no rotation", and qj = (cos 15°, sin 15°, 0, 0) is a 30° turn about the x axis. Their dot product is cos 15° = 0.966, and 2 arccos(0.966) = 2 × 15° = 30°, or 0.524 rad. The formula recovers exactly the turn between them.

Candidates closer than a threshold join a cluster. Clusters below a minimum size are thrown away as outliers, and the remaining clusters are ranked by the 2D silhouette IoU: render the posed mesh's outline, and compute the intersection over union with SAM 3's mask, the overlapping area divided by the combined area, 1.0 for a perfect match. The paper's empirical justification: confident samples concentrate on the same pose, while estimator noise scatters across SE(3). A crowd of agreeing samples is a strong signal, and it is free: nothing is re-run through the diffusion backbone.

A generator becomes a tracker without training. Nothing in SAM 3D was changed. The shape is pinned by blending, the pose is pulled toward the last frame by an amount that point tracks decide, and a cheap vote picks one pose from 25. Every piece works on the model's sampling loop, not its weights, which is why a stronger SAM 3D would drop straight in.
The point tracks show the object turning fast between two frames. What does the adaptive rule do, and why?

Chapter 4

Scoring a reconstruction

Measures a rebuilt object with F-scores and Chamfer distance, then reads Table 2, the ablation and 450 human votes

Chapter 3 built a tracker. Is it any good? To grade a reconstruction you need the right answer, and the right answer for an object's 3D shape and pose only exists where someone measured it with special equipment. So the first test is on two lab datasets with ground truth, and the second, on everyday video where no ground truth exists, uses people's eyes.

The lab datasets are DexYCB, videos of hands grasping objects from the YCB set (a standard collection of everyday objects used across robotics), and HOI4D, egocentric videos of hand-object interaction with 4D annotations. Following the setup of the prior work it compares against, the paper evaluates on 160 annotated DexYCB videos and 12 annotated HOI4D videos. And it hands every method the ground-truth hands, so that only the object's reconstruction and tracking are being graded.

Comparing two clouds of points

How do you compare a rebuilt object with the real one? Sample points on both surfaces, in their posed positions. You now have two point clouds: the prediction P and the ground truth G. For every point, find its nearest neighbour in the other cloud, and the distance to it. Everything else is built from those distances.

Precision at a threshold τ is the fraction of predicted points that lie within τ of some true point: how much of what you built is really there. Recall at τ is the fraction of true points that lie within τ of some predicted point: how much of the real object you built. A reconstruction can cheat either one alone. A single perfect point has precision 1 and terrible recall; a huge blob covering the table has recall 1 and terrible precision. The F-score combines them so that neither can be cheated:

F@τ
The F-score at distance threshold τ, from 0 to 1, higher is better. F-5 and F-10 in Table 2 use two thresholds; in the prior work whose setup the paper follows, those are 5 mm and 10 mm.
Prec@τ
Fraction of predicted points within τ of the true surface.
Rec@τ
Fraction of true points within τ of the predicted surface.

This is the harmonic mean, which is dragged down by whichever of the two is smaller. Precision 1.0 with recall 0.1 gives an F of 0.18, not the 0.55 an ordinary average would give.

The second metric has no threshold at all. Chamfer distance (CD) averages the squared nearest-neighbour distances, in both directions, and adds the two averages:

CD
Chamfer distance, lower is better. Papers differ in details (squared or not, sum or average of the two terms); Table 2 does not restate its variant or unit, so compare its CD values only with each other.
mean over P
For every predicted point, the squared distance to the closest true point, averaged: penalises stray predicted points.
mean over G
For every true point, the squared distance to the closest predicted point, averaged: penalises missing parts.

A worked example makes the difference between the two metrics concrete (illustrative, in millimetres, in 2D). The true surface is three points: (0, 0), (10, 0), (20, 0). The prediction is four points: (0, 3), (10, 6), (26, 0), and a stray point at (40, 0).

  1. Predicted to true. (0, 3) is 3 from (0, 0). (10, 6) is 6 from (10, 0). (26, 0) is 6 from (20, 0). The stray (40, 0) is 20 from (20, 0). Distances: 3, 6, 6, 20.
  2. True to predicted. (0, 0) is 3 from (0, 3). (10, 0) is 6 from (10, 6). (20, 0) is 6 from (26, 0). Distances: 3, 6, 6.
  3. At τ = 5 mm. Precision: only the 3 is within 5, so 1 of 4 = 0.25. Recall: 1 of 3 = 0.333. F-5 = 2 × 0.25 × 0.333 / (0.25 + 0.333) = 0.1667 / 0.5833 = 0.286.
  4. At τ = 10 mm. Precision: 3, 6, 6 pass, the 20 fails, so 3 of 4 = 0.75. Recall: all 3 pass, 1.0. F-10 = 2 × 0.75 × 1.0 / 1.75 = 0.857.
  5. Chamfer. Predicted side: (9 + 36 + 36 + 400) / 4 = 481 / 4 = 120.25. True side: (9 + 36 + 36) / 3 = 81 / 3 = 27. CD = 120.25 + 27 = 147.25 mm², or 1.47 cm².

Without the stray point, F-10 would be a perfect 1.0 and the Chamfer distance would be 27 + 27 = 54 mm². One stray point cost F-10 a quarter of its precision, but it nearly tripled the Chamfer distance, because its 20 mm error was squared into 400. F-scores count mistakes; Chamfer distance weighs them by size, squared. Reading both tells you whether a method is sloppy everywhere or mostly right with a few wild points.

Grade a rebuilt mug

Hollow rings are the true surface of a mug, seen from above; dots are a reconstruction. A dot is warm when it is within the threshold of the true surface and red when it is not. Shift it, scale it, add stray points, and change the threshold.

3 mm
0%
0

Illustrative 2D clouds: 104 points on a mug seen from above (90 on a 40 mm radius rim, 14 on the handle), and a copy that you distort (plus a fixed jitter under 1 mm). Precision, recall, F and Chamfer are computed live, with the Chamfer variant defined above. Real evaluations sample thousands of points on 3D surfaces.

Two habits of the metrics show up at once. F-5 is strict: a sideways shift of 7 mm, about the thickness of a pencil, takes it from 1.00 to about 0.46, while F-10 stays at 1.00. And stray points barely move the F-scores but send the Chamfer distance soaring. Keep that in mind when reading Table 2.

Who is compared

Two groups of baselines. The first reconstructs the hand and object jointly: image-based methods (HO, IHOI, HORSE) and video-based ones (MCC-HO, G-HOP). The numbers for HO, IHOI, HORSE and MCC-HO are taken from Wu and colleagues, the authors of MCC-HO. For G-HOP on HOI4D the authors evaluated its released per-clip checkpoints; for DexYCB, where G-HOP reported no video results, they ran its test-time optimization themselves on all 160 clips with the released prior. The second group is the two strongest object trackers, FoundationPose and Any6D. These are slotted into the Do as I Do pipeline in place of its tracker, with every other component unchanged. That is a controlled experiment: any difference is due to tracking alone.

Read the reconstruction results

Choose a view. Table 2 compares every method; the ablation compares versions of the tracker; the raters view shows the in-the-wild vote. For the two tables, pick the dataset and the metric.

Every value is printed in the paper: Table 2 (reconstruction), Table 5 (object tracking ablation) and Appendix C (150 in-the-wild videos, 3 raters each). IHOI reports no DexYCB numbers. Bars for CD are drawn so that shorter is better.

What Table 2 says

On DexYCB the new tracker leads on all three metrics: F-5 0.71 against 0.69 for both FoundationPose and Any6D, F-10 0.93 against 0.89 and 0.88, and Chamfer distance 0.66 against 0.89 and 0.97. That last gap is the largest in relative terms: (0.89 − 0.66) / 0.89 = 0.26, a Chamfer distance about a quarter lower than FoundationPose's. Since Chamfer punishes large errors hardest, a lower CD at nearly the same F-5 suggests fewer wild frames, not a uniformly tighter fit.

(0.89 − 0.66) / 0.89 = 26% lower Chamfer distance than FoundationPose on DexYCB

On HOI4D the race is tight: 0.72, 0.91 and 0.49 against FoundationPose's 0.71, 0.91 and 0.49. That is a one-point lead on F-5 and a tie on the other two. G-HOP, the strongest joint method there, scores 0.69, 0.91 and 0.63. The joint image-based methods trail far behind on both datasets; on DexYCB, HORSE's Chamfer distance (6.97) is more than ten times the new tracker's. The paper's claim of a new state of the art on both datasets holds, but on clean lab video it is a narrow one.

What people saw on everyday video

The lab datasets are clean: controlled lighting, calibrated cameras, objects kept in view, and ground-truth annotations. The paper's real target is the opposite. So the authors collected 150 videos from in-the-wild internet footage, egocentric datasets and generated videos, where no ground truth exists, and asked people. Each rater saw the original video beside two reprojections, the posed mesh from each tracker drawn back onto the frames, and chose the one that tracks the object's true motion more consistently. Left and right were randomised per video to avoid position bias. Each video got three ratings from three different raters out of a pool of five: 450 judgements.

Raters preferred the new tracker 67% of the time, FoundationPose 18%, and called 15% ties. Among decisive votes, that is a win rate of 67 / (67 + 18) = 0.788, the paper's 79%. On 75% of videos all three raters agreed. The paper reports Fleiss' kappa of 0.65: kappa measures how much raters agree beyond what chance alone would produce (0 is chance, 1 is perfect agreement), and the paper calls 0.65 substantial. Qualitatively (its Figure 5), FoundationPose loses the object under mild motion blur and occlusion, while the new tracker recovers consistent translations and rotations across the whole clip.

67 / (67 + 18) = 0.79 win rate among non-tied judgements

Clean benchmarks hide the failure this tracker was built for. On HOI4D, FoundationPose ties the new tracker on F-10 and Chamfer distance. On everyday video, people prefer the new tracker about 3.7 times as often (67% against 18%). The difference between the two tests is exactly what Chapter 3 targeted: occlusion, blur and ambiguity, which lab video mostly lacks.

What the ablation says

Table 5 takes the tracker apart along the two design axes of Chapter 3. With clustering held fixed, adaptive pose guidance beats fixed guidance on both datasets: DexYCB F-5 0.71 against 0.70 and Chamfer distance 0.66 against 0.74; HOI4D F-5 0.72 against 0.69. With adaptive guidance held fixed, picking a candidate at random is clearly worse on HOI4D (F-5 0.62, CD 0.66) than clustering (0.72, 0.49), and clustering matches the principled log-likelihood ranking almost exactly (0.71 against 0.72 F-5 on DexYCB, identical on HOI4D) while being up to 30 times faster. The expensive, exact answer and the cheap vote pick the same poses.

A reconstruction matches the true surface everywhere to within about 7 mm, with no stray points. What should you expect?

Chapter 5

Across the embodiment gap

Recasts copying a human hand as a search for robot controls that make a simulated object move the same way

At the end of Chapter 3 we hold a 4D reference: for every frame, the human hand's pose and the object's 6-DoF pose, in one near-metric, gravity-aligned space. Now it has to become robot motion. The robot is a Sharpa Wave hand with 22 degrees of freedom. Its fingers are not the human's fingers: different lengths, different joint placements, different proportions. The paper states the problem in one sentence: the reference is incomplete, because human and robot morphologies differ, and because contact information and forces are absent from the kinematic signal. The video showed where things were, never how hard anything pressed.

That mismatch of bodies is the embodiment gap. The simplest way to cross it would be to copy the human's joint angles onto the robot. The next simplest is to copy where the fingertips went. The device shows why neither is enough.

Put a robot finger where a human finger was

A human finger (blue, thin) presses on the top of a box. The robot finger (warm) has longer links. Choose how to map one onto the other, curl the finger, and add the reconstruction's error in how deep the fingertip sits.

50°
+2 mm

Illustrative planar fingers: human links 45, 25 and 20 mm; robot links 52, 34 and 26 mm (not the Sharpa Wave's real dimensions). "Match the fingertip" solves inverse kinematics numerically. The contact stiffness in "simulate" (0.5 N per mm) is a toy value; the paper's simulator is MuJoCo Warp.

Copying joint angles fails at once: the same bends on longer bones put the fingertip about 2 cm from where the human's was at a medium curl, sometimes deep inside the object, sometimes off it entirely. Matching the fingertip is the classic fix, and it is what the paper does first. Using mink, an inverse-kinematics library built on MuJoCo, it kinematically retargets the human hand onto the robot hand, solving for the robot joint angles whose fingertips land on the human's fingertips in every frame. Inverse kinematics (IK) is exactly this search: given where a point on the robot should be, find the joint angles that put it there.

Because IK for a many-jointed hand has many local solutions, the appendix adds a cheap robustness trick: compute several kinematic retargetings from different random initial poses, which avoids getting stuck in a poor local minimum (the paper does not say how the one that is kept gets picked). The result is the reference trajectory for the robot: its hand pose and 22 joint angles per frame.

But look at what the depth-error slider did. Kinematic retargeting copies the reconstruction's mistakes faithfully. If the reconstructed fingertip is 2 mm inside the object, the robot's fingertip is 2 mm inside the object, which no real finger can be. If it is 2 mm short, the robot never touches it, and an object that is not touched cannot be lifted. The related work names these failures: penetration, fingertip sliding and grasp instability. Only physics can tell a firm grasp from a finger resting beside an object. So the reference is not the answer. It is the starting point of a search.

Search in simulation

Dynamics-aware retargeting asks a different question. Not "which robot pose looks like the human's?" but "which robot controls, played in a physics simulator, make the object move the way the human's object moved?" The hand is allowed to differ from the human's hand, as long as the object does the right thing.

The simulator is MuJoCo Warp, the GPU version of the MuJoCo physics engine, stepping at sim_dt = 0.005 s, 200 steps per second. Two preparation steps make the object simulate well. Collision checking is fast and stable between convex shapes (shapes with no dents: a line between any two points inside stays inside), so each object mesh is split into convex pieces with CoACD, an approximate convex decomposition method. And to stabilise interactions with many contacts, the meshes are thickened and dilated by 2 mm. One more trick handles objects that do not lie flat, such as a spoon standing in a jar: a flat base plate is added under the object that touches only the floor, never the robot, so any starting pose can rest stably without reconstructing the rest of the scene.

The search itself builds on SPIDER (Pan and colleagues), the previous state of the art for dexterous retargeting at scale: an MPPI-style sampling-based optimization. MPPI, model predictive path integral control, is an idea simple enough to state in one line: try many noisy versions of your current plan, and move the plan toward the ones that scored well. Here is the loop, with the paper's Table 4 values.

  1. Represent the plan by knots. Instead of a control for every one of the 200 simulation steps a second, the plan stores knots, control values every knot_dt = 0.2 s, and interpolates between them. Over the planning horizon of 3.0 s that is 3.0 / 0.2 = 15 intervals, so 16 knots. Reading Table 4's separate noise scales for position (0.01), rotation (0.01) and joints (0.1), each knot holds the hand's base position and orientation plus its finger joint targets: 6 + 22 = 28 numbers, 16 × 28 = 448 per plan. The arm is not in the search; it is added afterwards by IK (Chapter 7).
  2. Sample. Draw num_samples = 1024 perturbed copies of the plan: every knot plus Gaussian noise.
  3. Roll out. Play every copy in the simulator for the full 3 s horizon: 3.0 / 0.005 = 600 simulation steps each, all 1024 in parallel on the GPU.
  4. Score. Compute each rollout's reward Ri: how well the simulated object and hand tracked the reference, minus penalties.
  5. Reweight. Replace the plan by an average of the samples, each weighted by how well it scored.
  6. Anneal and repeat. Shrink the noise and go again, max_num_iterations = 32 times.
  7. Execute and slide. Commit the first ctrl_dt = 0.5 s of the plan (100 simulation steps), move the 3 s window forward by 0.5 s, and plan again, starting from the previous plan with the next chunk of reference appended at its end.

The reweighting step is where MPPI gets its name. Each sample's weight is an exponential of its reward:

wi
Sample i's weight. All weights are positive and add up to 1.
Ri
The reward of sample i's simulated rollout. Subtracting the best reward keeps the exponentials from overflowing; it cancels in the ratio.
λ
The temperature: small λ gives nearly all the weight to the best few samples, large λ averages broadly. The paper does not print its value.
u
The plan: all 16 knots of controls.
δi
Sample i's noise, one Gaussian draw per knot and per control.

A three-sample example (illustrative). Rewards −2.0, −2.5 and −4.0, with λ = 0.5. Subtract the best, −2.0: exponents 0, −1 and −4, so unnormalised weights 1, e−1 = 0.368 and e−4 = 0.018. They sum to 1.386, giving weights 0.721, 0.265 and 0.013. If the three samples nudged one knot by +3, −1 and +8 cm, the plan moves by 0.721 × 3 + 0.265 × (−1) + 0.013 × 8 = 2.163 − 0.265 + 0.104 = +2.0 cm. Mostly toward the best sample, a little toward the second, and the worst one barely counts.

The noise is where the "annealing" happens. The paper describes a kernel "annealed across both iterations and the prediction horizon, which shifts from broad exploration to local refinement". Table 4 lists the ends, which we read as: noise multiplied by first_ctrl_noise_scale = 1.0 at the first knot of the horizon and last_ctrl_noise_scale = 4.0 at the last, and shrinking across the 32 iterations toward final_noise_scale = 0.01. Under that reading, the controls about to be executed are perturbed gently, the far future is explored boldly (it will be refined in later windows, when it comes closer), and each planning step starts broad and ends fine. The exact curve between those ends, and which of the two axes each named parameter controls, is not printed.

What the search is scored on

The paper rewards rollouts for tracking the object (position and orientation) and the hand (position, orientation and finger joints), with a penalty on excessive penetration, and the transition reward of Chapter 6. The penetration penalty exists because optimizers are ruthless: a simulator lets fingers sink slightly into objects, and a search that is scored only on tracking will happily discover that sinking a finger deep into an object produces enormous contact forces that hold it in place. The penalty, scaled by 3000, makes that trick expensive. With the Table 4 scales, schematically:

object terms
Tracking the object's position (scale 1.0) and orientation (0.3). Each r is larger when the error is smaller; the exact functions are not printed.
hand terms
Tracking the hand's base position (0.1), base orientation (0.03) and finger joints (0.01).
terminal
A term at the end of the horizon, scaled by 10.0.
penetration
A penalty on excessive interpenetration, scaled by 3000, so the search cannot exploit the simulator.
transition
A constant penalty per failed rest or in-hand step, scaled by 0.5 (Chapter 6).

Read the scales as a statement of priorities. Suppose the terms are plain distances (an assumption; the paper gives scales, not formulas). Then 1 cm of object position error costs 1.0 × 0.01 = 0.01, and so does 10 cm of hand position error, 0.1 × 0.1 = 0.01. The finger joints weigh a hundredth of the object. The object's motion is the goal; the human's hand is a hint about how to achieve it.

The object is the target; the hand is a suggestion. Kinematic retargeting tries to make the robot hand look like the human hand. Dynamics-aware retargeting lets the robot hand go wherever it must, within physics, to make the object move like the human's object did. That is how a reference whose fingers float 2 cm from the mug can still yield a working grasp.

Watch that trade happen. Below is a one-dimensional version of the search: a hand that must pick an object up, lift it, set it down and let go. The reconstructed hand path floats 2 cm too high at the pickup, like the misaligned references in the paper's Figure 4.

Search for the grasp

Heights over a 3 s horizon. Dashed: the reference hand (blue) and object (pink). Solid: the simulated rollout of the current plan; faint lines are some of the samples; dots are the plan's knots. Press Optimize and watch 32 iterations, or step one at a time.

A toy: one height and one grip instead of 28 controls, a spring-driven hand instead of MuJoCo Warp, and a grasp that holds only if the grip closes within 6 mm of the object's grasp height. The loop follows the paper's Table 4: knots every 0.2 s over a 3.0 s horizon (16 by our count), 600 steps of 0.005 s, up to 1024 samples and 32 iterations, and reward weights 1.0 (object) and 0.1 (hand); the noise running 1.0 to 4.0 along the horizon and annealed toward 0.01 across iterations is our reading of Table 4's noise-scale parameters. The temperature and the decay curve are our choices.

Three behaviours to notice. Starting from the kinematic reference, the hand closes 2 cm above the grasp and lifts nothing. With annealed noise, early broad samples stumble on plans that dip lower and catch the object; their high rewards pull the plan down, and later narrow samples polish it. The final plan tracks the object closely while the hand sits below its own reference, the trade the reward weights asked for. With fixed small noise the search never strays far enough from the reference to find the grasp. With fixed large noise it finds it, but cannot settle precisely.

What does this cost? One planning step evaluates 1024 samples in each of 32 iterations: 1024 × 32 = 32,768 rollouts. Each rollout is 600 simulation steps, so a planning step simulates 32,768 × 600 = 19,660,800 steps, to commit 0.5 s (100 steps) of motion. A 10 s clip (illustrative) needs 10 / 0.5 = 20 planning steps: 655,360 rollouts and about 393 million simulation steps, before the warmup of Chapter 6. That is why a GPU-parallel simulator is not a convenience here but the enabling technology: it turns hundreds of millions of physics steps into minutes.

1024 × 32 × 600 = 19,660,800 simulated steps per 0.5 s of robot motion

One last detail from the appendix, about the "execute and slide" step. Each new planning window starts from the previous plan's controls, with the next chunk of reference appended. Pasting a raw reference onto an optimized plan produces a jump where they meet, and jerky motion. So the optimized controls are blended into the reference by interpolation, the paper's reference blending.

In the search, the final plan's hand sits well below the reference hand at the pickup. Why is that the right outcome?

Chapter 6

Three fixes for a noisy reference

Explains warmup, random pushes and the transition penalty, the three parts that lift success from 25 to 71 percent

Chapter 5's search is SPIDER's, and SPIDER is good. On OakInk2, a large motion-capture dataset of 1,352 clean two-handed human-object task trajectories, it retargets 72% of tasks successfully. On the paper's 655 references rebuilt from everyday video, the same search succeeds on 25%. Same optimizer, same robot, same simulator. The difference is the reference.

Success has a precise meaning here. Following recent retargeting work, a trajectory counts as successful when the simulated object's mean position error stays under 0.1 m and its mean rotation error under 0.5 rad (about 29 degrees), both measured against the reference over the whole clip. It is a forgiving bar for precision and an unforgiving one for disasters: drop the object, or never pick it up, and the average error blows through it.

The paper's Figure 4 shows three ways a noisy reference defeats the baseline, and the method adds one component for each. Table 3 adds them one at a time, on both datasets.

MethodRebuilt: successPos (m)Rot (rad)OakInk2: successPos (m)Rot (rad)
Annealed sampling0.250.080.400.720.080.32
+ Warmup0.660.060.280.770.060.25
+ Perturbation0.670.060.300.790.030.14
+ Transition reward0.710.050.280.810.030.15

In counts: 25% of 655 rebuilt clips is about 164 successes; 71% is about 465. Roughly three hundred clips that were unusable become usable robot data. On OakInk2, 72% of 1,352 is about 973 and 81% is about 1,095.

0.71 × 655 − 0.25 × 655 = 465 − 164 ≈ 301 more usable clips

Fix one: warmup

The paper finds two problems at the very start of a trajectory, in its first horizon of H steps.

The first is the noisy first frame. The simulation starts from the reference's first frame: the hand where the reconstruction put it, the object where the reconstruction put it. If the clip opens with the object already in the hand, say a whisk held in mid-air, and the reconstruction placed the fingers a centimetre off, then at t = 0 the simulated whisk is simply not held. Gravity takes it before the optimizer can do anything. The paper calls such states impossible to recover from.

The second is subtler, and it comes from the way the search slides its window. Planning windows start every 0.5 s and look 3 s ahead. A moment t seconds into the rollout is planned by every window that started in the previous 3 s, so by min(6, ⌊t / 0.5⌋ + 1) windows. A moment late in the clip is planned six times: first when it is 2.5 to 3 s away at the far end of a window, where the noise is largest, and then again and again as it approaches, with ever finer noise. A moment at t = 0.2 s is planned once, in the first window, sitting near the window's start where the noise is smallest. Broad exploration never reaches it. That is the paper's "annealed sampling does not fully explore these H steps since they appear only at the start of the rollout horizon".

windows(t) = min(6, ⌊t / 0.5⌋ + 1):   t = 0.2 s → 1,   t = 1.2 s → 3,   t ≥ 2.5 s → 6

One fix handles both. Warmup prepends H extra steps to the reference. During warmup the object is held in place, for instance in mid-air where the first frame put it, by a weld, a simulator constraint that pins it rigidly to the world, while the robot hand is free to move. When warmup ends the weld is dropped and simulation proceeds normally. So the hand gets a whole horizon to find a stable grip on an object that cannot fall, and the reference's first frames move from the edge of the plan into the middle of it, where every window explores them. Reading H as one 3 s horizon, warmup adds 600 simulation steps and 6 planning windows per clip.

Crucially, the paper stresses, warmup assumes no grasp sampling and no grasp heuristics. It does not tell the robot how to hold a whisk. It just gives the same optimizer the time and the safety to discover a grasp by itself. The results say this is the big one: on rebuilt references success jumps from 0.25 to 0.66, and the errors fall from 0.08 m and 0.40 rad to 0.06 m and 0.28 rad. The paper's explanation: warmup discovers initial states that are much more stable and natural than the noisy initial frame, which leads to successful tracking in the steps that follow. On clean OakInk2 references, whose first frames are already good, the gain is smaller: 0.72 to 0.77.

Fix two: random pushes

The second failure is a local minimum: a solution better than everything near it but far from the best. Here it looks like an object balanced on the fingertips. Such an interaction can track the reference for a while, so it scores well, and small changes to it score worse, so the search stays. But it cannot recover from anything: the first disturbance and the object falls.

The fix is borrowed from sim-to-real training, where robots trained in simulation are hardened for the real world by random disturbances (the paper cites OpenAI's Rubik's cube hand and Rudin and colleagues' legged robots). Do as I Do applies random forces (and, per Table 4, torques) in the sampled rollouts. A grasp that balances on fingertips now fails in the rollouts where a push arrives; an enveloping grasp survives them. Averaged over pushes, the robust grasp scores higher, and the weights of Chapter 5 move the plan toward it. The paper notes that this is general-purpose: unlike alternatives such as SPIDER's contact guidance, it does not assume the reference is accurate enough to say where contacts should be.

Table 4 lists the knobs: num_perturb_samples = 4, perturb_force_scale = 0.5, perturb_torque_scale = 0.5, perturb_prob = 0.05 and perturb_continue_prob = 0.95. The paper does not spell out the exact process, but the names suggest a simple one: at each step a push starts with probability 0.05, and a push in progress continues with probability 0.95. Under that reading a push lasts on average 1 / (1 − 0.95) = 20 steps, and in the long run a push is active 0.05 / (0.05 + 0.05) = 50% of the time. The rollouts are, deliberately, a rough ride.

mean push length = 1 / (1 − 0.95) = 20 steps,   share of time pushed = 0.05 / (0.05 + 0.05) = 50%

The numbers show a mixed quantitative effect. On rebuilt references success moves from 0.66 to 0.67 and rotation error slightly up, 0.28 to 0.30. On OakInk2 success goes from 0.77 to 0.79 and the errors roughly halve: 0.06 to 0.03 m, 0.25 to 0.14 rad. The paper's own reading is that perturbation noticeably improves the qualitative results, natural grasps, while marginally affecting the quantitative metrics. A grasp can be tidier without its average error being lower.

Fix three: the transition reward

The third failure is the missed pickup. An object's life in a clip has transitions: from "rest" (sitting on the table) to "in-hand" (held) and back. These are step functions. Either the fingers closed on the mug or they did not. But a tracking reward is soft: a hand that closes 1 cm above the mug and a hand that grasps it can have nearly the same hand-tracking error, and with a noisy object reference even their object errors can be close. The search gets no clear signal to prefer the grasp.

So the paper adds a hard signal. It labels each reference step as rest or in-hand by measuring the reference hand-object distance: under a threshold ε means in-hand. Then it adds a constant penalty for each failed transition step, of two kinds:

trans
The count of steps where the rollout contradicts the reference's stage.
rest steps
Reference steps where the hand is farther than ε from the object: the object should be on the floor.
in-hand steps
Reference steps where the hand is within ε of the object: the hand should be touching it.
0.5
transition_penalty_scale in Table 4. The value of ε is not printed.

The effect is small but real: 0.67 to 0.71 on rebuilt references, 0.79 to 0.81 on OakInk2. The paper says it encourages successful picks and places for trajectories that would otherwise have missed the object at the crucial transition steps. Explore all three fixes below.

Take the three fixes apart

Pick a fix. Warmup: every bar is one planning window, darker where its noise is larger; slide the reference time and count the windows that plan it. Pushes: two grasps, four pushed rollouts each. Transition: two candidate rollouts scored with and without the penalty.

0.2 s

Warmup: window timing from Table 4 (ctrl_dt 0.5 s, horizon 3.0 s, noise 1.0 to 4.0 along the horizon, drawn as a straight ramp). Pushes: the start and continue probabilities of Table 4 read as a per-step process at 0.02 s steps; the two grasps' tolerances and tracking errors are illustrative. Transition: the 0.5 penalty is Table 4; the reference, the rollouts, the tracking scores and the temperature are illustrative.

Put the three side by side and a pattern appears. None of the fixes makes the reference more accurate. Each one changes what the search can find: warmup gives it a safe place and enough exploration to find a grasp, pushes make fragile solutions score as badly as they deserve, and the transition penalty turns a soft preference into a hard one exactly where the reference is least trustworthy. That is the right way to design for noise you cannot remove.

The appendix adds three lessons that are easy to miss and useful to anyone building this. Reference blending (Chapter 5) smooths the seam between one window's optimized controls and the next chunk of reference. Robust kinematic retargeting runs the cheap fingertip IK from several random starts, because the sampler searches around the reference and needs a reasonable one to find feasible controls. And the object base plate lets an object start in any pose, a spoon upright in a jar, without reconstructing the jar.

Warmup lifts success on rebuilt references from 0.25 to 0.66, far more than on OakInk2 (0.72 to 0.77). What best explains the difference?

Chapter 7

From internet clip to robot hand

Follows the clips from the wild into simulation, then onto two real arms and two real hands

So far every result has lived in a dataset or a simulator. The paper's fourth claim is that the output is robot-complete: trajectories that play on a real dexterous hand and arm. This chapter follows the data from where it comes from to where it ends up, and it is honest about which parts are measured and which are shown.

Where the clips come from

The paper draws on three kinds of video, and each is hard in its own way.

Egocentric clips are filmed from the person's own head, usually with a head-mounted camera or smart glasses. The hands are large and close, which helps a hand tracker, but the working hand often covers the object it holds, and the camera moves with every glance, so a still object seems to move. That is the reason the hand tracker had to be a world-space one (Chapter 2).

Exocentric clips are filmed by someone or something else, from outside: a cooking channel, a tutorial, a phone propped on a shelf. The viewpoint is arbitrary, the hands can be small in the frame, and edited internet footage adds cuts, zooms and pans.

Generated clips come from video generation models. They have a property no other source has: there is no physical object anywhere. Nothing can be scanned, so only a pipeline that builds its mesh from pixels can use them at all. The paper treats them as one more source of human behaviour to reconstruct.

Across the paper these sources feed three collections: the 150-video human-preference benchmark of Chapter 4 (internet, egocentric datasets and generated videos), the 655 rebuilt references used to test retargeting in Chapter 6, and a final set of 500 high-quality, human-verified dexterous manipulation trajectories. That last set is split 53% internet, 31% egocentric and 16% generated:

0.53 × 500 = 265 internet,   0.31 × 500 = 155 egocentric,   0.16 × 500 = 80 generated

The behaviours are broad. The paper's Figure 3 shows 20 distinct verbs coming out of the pipeline: placing, picking, scrubbing, spreading, squeezing, ironing, painting, dusting, digging, erasing, pouring, writing, whisking, stirring, poking, tamping, drilling, hammering, cutting and basting. None of them needed a verb-specific rule. That is the payoff of the method's refusal to assume a grasp type or an object category.

The robot

Every experiment uses the 22-DoF Sharpa Wave hand. The real-world setup is bimanual, two arms side by side: two Universal Robots UR3e arms, each carrying a Sharpa Wave hand, with arms and hands both commanded at 50 Hz, fifty new joint targets every second.

To show the data's quality, the authors ran a representative set of trajectories on this robot: 10 motions chosen to cover different object shapes and different grasp classes from Feix and colleagues' taxonomy of human grasps. They name four: the writing tripod (the way you hold a pen, between thumb, index and middle finger), the power grasp (the whole hand wrapped around a handle), and the ventral and parallel extension grasps. The ten tasks are whisking, pouring, dusting, squeezing, tamping, erasing, stirring, hammering, spreading and picking.

From the simulator to the table

The retargeted trajectory of Chapter 5 is a free-floating hand: a hand base pose plus 22 joint angles, with no arm. It lives in the coordinates of the original video's camera. Getting it onto the real robot takes five steps, described in the paper's Section 4.4 and Appendix D.

  1. Gravity. Already done in reconstruction by GeoCalib: "down" in the trajectory is down in the world. That fixes two of the three rotation angles, roll and pitch.
  2. Place it in the workspace, by hand. Even after gravity alignment the reconstruction follows the video's camera coordinates: nothing says where on the robot's table the action happened or which way it faced. So the authors manually align the initial pose, four numbers, x, y, z and yaw (the heading around the vertical), with the robot's workspace in simulation.
  3. Add the arm by inverse kinematics. Using mink again, they solve for the arm joint angles that put the wrist where the retargeted hand base is, at every step.
  4. Check in a digital twin. A digital twin is a simulated replica of the real setup, here in MuJoCo and visualised with Viser. Every trajectory is watched there first for self-collisions, contacts with the table and similar problems.
  5. Roll out. Validated trajectories run on the real arms and hands at roughly half speed, both commanded at 50 Hz.

Why only four numbers in step 2? A rigid placement in 3D has six: three for position, three for rotation. Gravity alignment has already fixed roll and pitch, since the table must be level. What remains is where the motion happens (x, y, z) and which way it points around the vertical (yaw). Those four the video cannot know, because they depend on your robot, not on the person in the video.

The digital twin in step 4 is doing real work, not a formality. A trajectory that looks fine in the abstract, a hand closing around a mug and lifting it, can still be physically wrong once real geometry is involved: the retargeted fingers can interpenetrate each other or the palm as they curl, a bimanual pour can swing one arm's wrist through the other arm's forearm on the way to the target, or an object that cleared the table in the video's camera view can, once replayed at the robot's actual height, end up with its base a centimetre inside the tabletop. None of this shows up in the Chapter 5 reward, which only scores how well the object and hand tracked their reference; it shows up only by watching the replica move in MuJoCo before any current reaches a real motor, which is exactly why the paper checks every trajectory there first.

Place the clip on the robot's table

Top: the table from above, with two arm bases and the zone each can reach. The warm path is a retargeted pouring motion, still in its camera's coordinates. Level it, then slide and turn it until every waypoint is reachable by the right arm. Bottom left: the same path from the side.

−24 cm
+14 cm
60°

An illustrative layout: the reach zone (15 to 50 cm from each base), the table and the path are toy values, not the paper's. The four aligned numbers (x, y, z and yaw, with height fixed here), the gravity step, the 50 Hz command rate and half-speed playback are the paper's (Section 4.4, Appendix D).

Two small computations tie the rates together. The retargeting simulator steps at 200 Hz and the robot takes commands at 50 Hz, so one robot command spans 200 / 50 = 4 simulation steps. And playing at half speed doubles the number of commands per motion: a 4 s retargeted segment (illustrative) becomes 8 s on the robot, 8 × 50 = 400 commands instead of 200.

200 Hz / 50 Hz = 4 sim steps per command,   4 s × 2 × 50 Hz = 400 commands at half speed

Four numbers are the last manual step. Everything from the pixels to the joint angles is automatic, except placing the clip in the robot's world (x, y, z and yaw) and a human check in the digital twin. That boundary is informative: it is exactly the information a video of someone else's kitchen cannot contain.

What is measured, and what is shown

Be precise about the evidence. The reconstruction results (Chapter 4) and the retargeting results (Chapter 6) are measured: tables of numbers on hundreds of clips. The real-robot results are demonstrated: film strips in the paper's Figures 6 and 12 (spreading, whisking, dusting, pouring, erasing, picking and more) and videos on the project page. The paper does not report a real-world success rate, and it describes playing validated trajectories back, not a controller that watches the object and corrects itself. The authors' own limitation applies directly: simulators model real dynamics only approximately, which bounds how well a trajectory that worked in simulation can work on real hardware.

What the demonstration does establish is the claim in the contributions: to the authors' knowledge, this is the first pipeline that goes from an internet video to real dexterous hand rollouts. In Table 1, no earlier method was shown on internet video at all.

Before deployment, the authors manually align x, y, z and yaw of each trajectory. Why these four numbers and not all six?

Chapter 8

Most clips are not data

Counts how many of 2,000 pre-filtered internet clips survive, and turns every loss into a check you can run

Human video is having a moment in robotics. The paper cites large efforts to scale robot learning with human data, egocentric collections among them (EgoScale, EgoVerse, EgoDex), and asks a sobering practical question: of the human video that is out there, how much can actually become robot data? Its Section 4.5, the human data filtering playbook, answers with a count.

The authors noticed recurring quality problems while analysing common datasets, and present the analysis on one of them: 100DOH, "Understanding human hands in contact at internet scale", a large collection of internet video of hands. The important detail is that 100DOH has already been filtered for hand-object interaction. This is not raw YouTube. The authors expect the lessons to apply more generally.

They sampled 2,000 clips of 10 seconds each. Before reading on, make a guess.

Guess the survivors

Out of 2,000 ten-second clips, already filtered for hand-object interaction, how many end up usable for reconstruction? Set your guess and press Reveal. Then press any filter to see what it catches.

50%

Every count is Section 4.5 of the paper (100DOH, 2,000 clips of 10 s). "Better future models" applies the paper's own best case, in which camera-motion and SAM 3D failures are fixed. The small clip drawings are illustrations of each failure, not frames from 100DOH.

The funnel, gate by gate

Only 187 of the 2,000 clips, 9%, have meaningful hand-object interaction present. That is the first and largest cut: 1,813 clips gone before any 3D model runs, in a dataset that was already selected for hands touching things. A hand resting on a counter, a hand that brushes past a cup, a presenter gesturing beside a product: in contact, perhaps, but not manipulation that a robot could copy.

Of those 187 candidates:

  1. 41 have the hand or the object outside the video's boundary. Chapter 2 needs the hand to measure scale and Chapter 3 needs the object's mask to track it; a hand that leaves the frame takes the ruler with it.
  2. 29 have no activity, or activity that spans a shot boundary, a cut from one camera shot to another. Across a cut the "previous frame" of Chapter 3 is a different scene, the point tracks break, and a 10-second clip is really two unrelated clips.
  3. 14 fail because of camera motion: a moving camera mixes its own motion into the object's apparent motion.
  4. 10 fail because SAM 3D cannot produce a usable object.
  5. 10 are lost for other reasons.

Subtract step by step: 187 − 41 = 146, − 29 = 117, − 14 = 103, − 10 = 93, − 10 = 83. Out of 2,000 sampled clips, 83 survive the quality check for the reconstruction pass: 83 / 2,000 = 4.15%, the paper's 4%.

187 − 41 − 29 − 14 − 10 − 10 = 83 of 2,000, about 4%

The best case, and the 20× penalty

Two of those losses are marked by the paper as fixable: camera motion and SAM 3D failures "may be fixed in the future with better models". Give those 24 clips back and the best case is 83 + 14 + 10 = 107 clips, 107 / 2,000 = 5.35%, roughly 5% of the data directly relevant for learning dexterous manipulation.

The paper turns that into a rule of thumb: a 20× penalty for not properly preprocessing and filtering internet video for robot learning. If only about 1 clip in 20 is usable, then a pipeline that ingests clips blindly spends 95% of its effort on material that cannot become robot data, or, looked at the other way, a dataset needs about 20 times more raw video than its useful size suggests.

107 / 2,000 = 5.35% ≈ 5%,   1 / 0.05 = 20×

Turn the rate around and it becomes a budget. Suppose you want 1,000 clips that survive the reconstruction check. At today's 4.15%, you need 1,000 / 0.0415 ≈ 24,100 clips sampled from a source like 100DOH. In the best case of 5.35%, you still need 1,000 / 0.0535 ≈ 18,700. Either way, the raw pool is an order of magnitude larger than the dataset you end with, and that is after 100DOH's own filtering for hands in contact.

Nine clips in ten fail before any model runs. Of the 1,917 clips that did not make it, 1,813 were lost at the first gate, for lacking real manipulation, and only 24 were lost to limitations the authors attribute to today's models. For anyone building human-video datasets for robots, curation is a bigger lever than a better reconstruction model.

The playbook as a checklist

Turn the gates into questions, and you can screen a clip in seconds, before spending minutes of GPU time on it. Ask them in this order, because the early ones are cheap and catch the most.

  1. Is a hand actually moving an object? Contact is not enough. The pipeline needs a manipulation whose effect on the object can be tracked.
  2. Are the hand and the object in frame for the whole clip? The hand is the scale reference and the mask is the tracker's anchor.
  3. Is it one continuous shot, with activity throughout? Trim at cuts. Drop idle stretches.
  4. Is the camera steady, or its motion recoverable? Handheld pans and zooms are the most common fixable failure.
  5. Is the object rigid, and a thing a single-image 3D model can rebuild? The method assumes rigid objects (a limitation the authors state), and its object pipeline is only as good as SAM 3D's mesh.

Walk one illustrative clip through it. Say a 10-second cooking video opens on a cutting board, cuts at the 6-second mark to a close-up of a pan, then keeps rolling. Question 1 passes: a hand is clearly chopping something in the first half. Question 2 passes too, hand and knife both stay in frame. Question 3 is where this clip dies: it is not one continuous shot, it is two shots joined by an edit, so the reconstruction's "previous frame" reasoning would suddenly compare a cutting board to a pan and see a spurious jump. The fix costs nothing extra: trim the clip to the first 6 seconds, drop the second shot, and only the surviving single-shot segment goes on to questions 4 and 5. That trim is the entire cost of catching this failure, against the 187,500 wasted passes through SAM 3D the pipeline would otherwise have spent trying, and failing, to track an object across a cut it never saw coming.

Why bother screening by eye or with a cheap classifier, instead of letting the pipeline find out? Because the pipeline is expensive per frame. Chapter 3's tracker draws 25 candidates of 25 Euler steps for every frame, 625 passes through SAM 3D's large network. For a 10-second clip filmed at 30 frames per second (an illustrative rate; the paper does not give one), that is 300 frames × 625 = 187,500 passes for the object's pose alone, before any retargeting. Spending that on a clip whose hand leaves the frame halfway through is pure waste. The playbook moves the rejection to the front, where a glance at the clip replaces all of that compute.

Notice how every check corresponds to an assumption we met in an earlier chapter. The ruler of Chapter 2 needs a visible hand. The guided tracker of Chapter 3 needs frame-to-frame continuity and a visible object. The retargeter of Chapter 6 tolerates a noisy reference but not a reference of the wrong thing. The playbook is the method's assumptions, written as questions about raw video.

It also puts the rest of the paper in proportion. The retargeting success rates of Chapter 6 are measured on references that were already reconstructed, not on raw video. A clip that fails this checklist never reaches the retargeter, so the 71% says how well usable clips become robot data, not how much of the internet is usable.

Which losses in the 100DOH analysis does the paper expect better models to recover, and what best case does that give?

Chapter 9

What it adds up to

Collects the numbers, writes the core loop as code, and marks the limits the authors admit

You can now read the Do as I Do paper and explain every design decision: why the hand is the ruler, why SAM 3D is steered instead of retrained, why the guidance strength comes from point tracks, why one pose is chosen by a vote, why the retargeter searches in physics instead of copying joints, and why warmup, pushes and a transition penalty are what a noisy reference needs. Let's lock it in.

The one-paragraph summary

Do as I Do turns monocular RGB video of a human hand using a rigid object into robot-executable dexterous manipulation data, in two steps and with no training of its own described in the paper. Reconstruction segments the hand and object with SAM 3, estimates depth and intrinsics with MoGe, tracks a near-metric hand with HaWoR, builds the object's mesh with SAM 3D, and tracks the object's pose by guided flow sampling: the shape is pinned, and the pose is blended toward the previous frame with a strength set by point-tracked rotation, 25 candidates per frame, chosen by clustering. The object is then placed at the hand's scale and gravity is aligned. Retargeting maps the hand onto a 22-DoF Sharpa Wave hand by fingertip IK, then searches in MuJoCo Warp with annealed MPPI (1024 samples, 32 iterations, 3 s horizon), made robust to noisy references by warmup, random pushes and a transition penalty. Retargeting success on 655 rebuilt clips rises from 25% to 71%, reconstruction sets a new state of the art on DexYCB and HOI4D, people prefer its tracking 67% to 18% on everyday video, and of 500 human-verified trajectories, 10 representative ones were played on two UR3e arms with Sharpa Wave hands.

The numbers that matter

QuantityValueWhy it matters
Models, all used as releasedSAM 3, MoGe, HaWoR, SAM 3D, BootsTAPIR, GeoCalibNothing is trained; everything new happens at inference
Pose candidates per frameN = 25, T = 25 Euler steps, 13 numbers eachExact log-density would cost 8,750 passes per frame
Guidance strengthαs in [0.9, 1]; αp = max(0.1, 0.7 − 0.09|Δθ|)Firm when still, loose when turning; 20 tracked points per frame pair
DexYCB (160 videos)F-5 0.71, F-10 0.93, CD 0.66FoundationPose 0.69, 0.89, 0.89
HOI4D (12 videos)F-5 0.72, F-10 0.91, CD 0.49FoundationPose 0.71, 0.91, 0.49: a near tie on clean video
Everyday video (150 clips, 450 votes)67% prefer it, 18% FoundationPose, 15% ties79% of decisive votes; 75% unanimous; Fleiss' κ 0.65
Retargeting search1024 samples × 32 iterations; 3.0 s horizon; 0.5 s control; 0.2 s knots; 200 Hz sim32,768 rollouts per planning step
Success, 655 rebuilt clips25% → 66% → 67% → 71%Warmup is the big step
Success, 1,352 OakInk2 tasks72% → 77% → 79% → 81%The fixes help even on clean motion capture
Verified trajectories500: 53% internet, 31% egocentric, 16% generated20 verbs, 10 run on real hardware
Robot2 × UR3e + 2 × Sharpa Wave (22 DoF), 50 Hz, about half speedManual x, y, z, yaw placement; digital twin check
100DOH filtering2,000 → 187 → 83 clips (4%); best case 107 (5%)The 20× penalty for not filtering

The core, as code

Two loops carry the paper: the guided pose tracker of Chapter 3 and the sampling-based retargeter of Chapters 5 and 6. Here they are in plain Python with NumPy. Helper functions stand for the released models and the simulator.

pythonimport numpy as np

# ── Step 1: track the object's pose with a steered SAM 3D (Section 3.1, Appendix A) ──
def track_pose(sam3d, frames, masks, shape_bar, pose0, alpha_s=0.95, N=25, T=25):
    poses, prev = [pose0], pose0
    for k in range(1, len(frames)):
        tracks = bootstapir(frames[k - 1], frames[k], masks[k - 1], n_points=20)
        dtheta = rigid_fit_2d(tracks, keep_inside=masks[k]).rotation       # closed form, by SVD
        alpha_p = max(0.1, 0.7 - 0.09 * abs(dtheta))                       # adaptive guidance
        cands = []
        for _ in range(N):
            eps_s, eps_p = np.random.randn(*shape_bar.shape), np.random.randn(13)
            xs, xp, dt = eps_s, eps_p, 1.0 / T
            for i in range(T):                                             # Euler steps, noise to sample
                t, tn = i * dt, (i + 1) * dt
                vs, vp = sam3d.velocity(xs, xp, t, frames[k], masks[k])
                xs = (1 - alpha_s) * (xs + dt * vs) + alpha_s * ((1 - tn) * eps_s + tn * shape_bar)
                xp = (1 - alpha_p) * (xp + dt * vp) + alpha_p * ((1 - tn) * eps_p + tn * prev)   # Eq. 1
            cands.append(xp)
        prev = pick_by_clustering(cands, masks[k])     # SE(3) clusters, drop small ones, best mask IoU
        poses.append(prev)
    return poses

# ── Step 2: retarget in simulation by annealed MPPI (Section 3.2, Appendix B, Table 4) ──
NUM_SAMPLES, ITERS = 1024, 32
SIM_DT, CTRL_DT, HORIZON, KNOT_DT = 0.005, 0.5, 3.0, 0.2
K, STEPS, DIM = int(HORIZON / KNOT_DT) + 1, int(HORIZON / SIM_DT), 6 + 22   # 16 knots, 600 steps, base + fingers
SIGMA = np.r_[np.full(3, 0.01), np.full(3, 0.01), np.full(22, 0.1)]     # pos, rot, joint noise
W = dict(obj_pos=1.0, obj_rot=0.3, base_pos=0.1, base_rot=0.03, joint=0.01,
         terminal=10.0, penetration=3000.0, transition=0.5)
PUSH = dict(n=4, force=0.5, torque=0.5, p_start=0.05, p_continue=0.95)

def noise(it):
    ramp = np.linspace(1.0, 4.0, K)[:, None]            # gentle now, bold far ahead
    decay = 0.01 ** (it / (ITERS - 1))                  # broad first, fine last (our curve)
    return ramp * decay * SIGMA                          # shape (16, 28)

def score(roll, ref, eps_contact):
    r = -(W['obj_pos'] * roll.obj_pos_err + W['obj_rot'] * roll.obj_rot_err
          + W['base_pos'] * roll.base_pos_err + W['base_rot'] * roll.base_rot_err
          + W['joint'] * roll.joint_err).sum(-1)          # tracking, summed over time
    r -= W['terminal'] * roll.final_obj_err
    r -= W['penetration'] * roll.excess_penetration.sum(-1)  # no exploiting the simulator
    in_hand = ref.hand_obj_dist < eps_contact                  # reference stage per step
    failed = np.where(in_hand, ~roll.hand_obj_contact, ~roll.obj_floor_contact)
    return r - W['transition'] * failed.sum(-1)            # the transition reward

def plan(sim, state, mean, ref, lam, eps_contact):
    for it in range(ITERS):
        samples = mean[None] + np.random.randn(NUM_SAMPLES, K, DIM) * noise(it)   # (1024, 16, 28)
        pushes = random_pushes(PUSH, NUM_SAMPLES, STEPS)                        # forces and torques
        roll = sim.rollout(state, samples, steps=STEPS, pushes=pushes)          # all 1024 in parallel
        R = score(roll, ref, eps_contact)                                       # (1024,)
        w = np.exp((R - R.max()) / lam); w /= w.sum()                          # MPPI weights
        mean = np.einsum('n,nkd->kd', w, samples)                            # weighted average
    return mean

def retarget(sim, ref, lam=1.0, eps_contact=0.02):     # lam and eps are not printed in the paper
    ref = prepend_warmup(ref, seconds=HORIZON)              # object welded in place, hand free
    mean = best_of(kinematic_retarget(ref, restarts=8))[:K]   # fingertip IK with mink, random starts
    state, out = sim.initial_state(ref), []
    for t0 in np.arange(0.0, ref.duration, CTRL_DT):
        mean = plan(sim, state, mean, ref.window(t0, t0 + HORIZON), lam, eps_contact)
        state = sim.execute(state, mean, seconds=CTRL_DT)    # commit 0.5 s = 100 sim steps
        if t0 + CTRL_DT >= ref.warmup_end:
            sim.release_weld(state)                          # warmup over: the object is on its own
        out.append(mean[: int(CTRL_DT / KNOT_DT)])
        mean = shift_and_blend(mean, ref, t0 + CTRL_DT)       # reference blending
    return out                                              # then arm IK, digital twin, 50 Hz

The code's structure mirrors the lesson. Nothing in it updates a network's weights: track_pose steers a frozen generator, and retarget searches a simulator. The constants are Table 4. The values we had to choose (the temperature, the contact threshold, the number of IK restarts and the decay curve) are marked, because the paper does not print them.

Where it sits in the field

The limits, in the authors' words

Three further gaps are visible in the paper without being listed as limitations: each trajectory is placed in the robot's workspace by hand (four numbers), the real-robot results are demonstrations without a reported success rate, and no policy is trained on the 500 trajectories. Each is a natural next paper.

Keep going

A colleague summarises the paper as "it turns 71% of internet videos into robot data". What is the most accurate correction?
The takeaway. Do as I Do never makes the video less noisy. It makes every stage tolerate the noise of the stage before: the hand fixes the depth's scale, point tracks decide how much to trust the last frame, a vote replaces an expensive likelihood, and physics, with a safe warmup, random pushes and a hard penalty for missed pickups, decides what the reference could not. That is how a clip filmed for people becomes data for a robot.

Now press Present or Teach and explain, out loud and from memory, why warmup lifts success from 25% to 66% on rebuilt references but only from 72% to 77% on motion capture. If you can, you own this paper. Then go back to the ladder and run twenty clips on every rung.

Based on "Do as I Do: Dexterous Manipulation Data from Everyday Human Videos" by Bhawna Paliwal, Haritheja Etukuru, William Liang, Pieter Abbeel, Nur Muhammad Mahi Shafiullah and Jitendra Malik (UC Berkeley, 2026)
Read the paper · Project page · Back to Veanors