Point trackers ask where every pixel went. AgentSTAR asks which structured 3D object, with which joints and which pose at each moment, could have produced the frames, and lets a vision-language agent steer a numerical optimiser through a render-and-compare loop until the answer fits. It tracks through occlusion, symmetry flips and glass, and cuts the best 3D tracker's error on ARCTIC from 7.65 to 5.59 centimetres.
Someone closes a book in front of a camera. For the first half of the motion you can see both covers and most of the pages. Then the pages slide behind the front cover, the text disappears, and by the last frame the book is a closed slab. Now ask the question every dynamic-reconstruction system has to answer: where did the pages go?
There are two ways to answer it, and the difference between them is the whole paper. The bottom-up way, which nearly every modern method takes, first recovers low-level visual evidence: scene flow, or the 3D trajectory of every pixel you can see. Then it tries to infer what object produced that evidence. The top-down way first recognises the object and its structure, a book with two hinged covers and a stack of pages, and uses that structure to recover the state over time: one hinge angle, one base pose, per frame.
Bottom-up has two failures the authors care about. First, dense correspondence is itself hard exactly when it matters: under occlusion and low visual overlap, which is the normal condition when a hand is manipulating something. A page you cannot see has no trajectory. Second, even when it works, the output is unstructured. A cloud of moving points tells you where visual evidence went, not what the parts, joints and state variables are. That structure has to be inferred afterwards from imperfect geometry, which is brittle, and it is precisely what a robot needs: an explicit object model it can load into a simulator and manipulate.
A book with one hinge, seen from the side. Drag the hinge angle. In bottom-up mode, tracked points live only where pixels are visible; watch them vanish as the front cover swings over the pages and the count of surviving tracks falls. In top-down mode, the same frames are explained by a single hinge angle, and every point on every page, visible or not, has a predicted position. The predicted page tips are drawn hollow where nothing can see them.
The paper's answer is to directly ask which structured 3D object and state sequence could have generated the video. That is analysis-by-synthesis: propose an object, render it, compare with the frames, adjust. Once the structure is known, the permissible configurations collapse from millions of independent point motions to a compact vector: a 6-DoF base pose plus one number per joint. The states are interpretable too. “The lid is at 64 degrees” is something a planner can use; a scene-flow field is not.
Analysis-by-synthesis is an old idea. Why does it work now, on casually captured video of scissors and ketchup bottles, when it did not before? Because two very different kinds of estimator have become available, and each is bad at exactly what the other is good at.
A large vision-language model, the kind that sits inside modern coding agents, is excellent at coarse judgement. Show it a render next to a photo and it will tell you the scissors are upside down, that the loop is on the wrong handle, that the lid looks too thick. It is also good at proposing rough 3D shape. But it is a poor state estimator: it cannot recover a pose to within a few degrees, and it cannot reliably produce continuous numbers. Ask it “what is the yaw?” and you get a guess.
Classical test-time optimisation, bundle adjustment and its relatives, is the opposite. Give it a differentiable or at least scoreable objective and it will find a local optimum to arbitrary precision. But the objective for “does this render match this mask” is riddled with local optima. A ketchup bottle rendered cap-up scores almost as well as cap-down. A box flipped by 180 degrees has the same silhouette. Start the optimiser in the wrong basin and it converges, precisely, to the wrong answer.
The silhouette overlap between a rendered ketchup bottle and the target mask, as a function of one pose variable, the in-plane yaw. The true yaw is marked. Because the body is symmetric and only the cap breaks the symmetry, there are several peaks that look almost as good numerically. Drop a local optimiser anywhere with start here and it climbs to the nearest peak. Then ask the VLM: it cannot say the yaw is 37 degrees, but it can look at the current render and say “the cap is at the wrong end, flip it,” which teleports the optimiser into the correct basin, where a second climb finishes the job.
Notice what the VLM's move is not. It is not a gradient. It is not a number. It is a discrete, semantic decision, “wrong basin, try the flip,” of the kind these models make reliably, applied to a problem where that decision was the expensive part. Everything else in the paper is the engineering that lets a coding agent make such decisions safely: a representation it can edit as code, a score it can read, tools that search regions it names, and diagnostics that tell it when time has gone wrong.
If an agent is going to propose an object, render it and edit it, the object has to live in a form the agent can write. AgentSTAR's answer is the most direct one available to a coding agent: the object is a Python script. Geometry is expressed with Blender shape primitives, which is freeform enough to build arbitrary topologies: a loop for a scissor handle, a lid on a cylinder, a stack of thin slabs for pages. The kinematic structure lives in the same script: which part is the base, which parts are joints, where each joint sits, what its axis is, and what its motion limits are.
Two things follow from that choice. First, every shape edit is a code diff, which is exactly the operation a coding agent performs best and can reason about in text. Adding a missing handle, moving a hinge, changing a joint from revolute to prismatic, each is a few lines. Second, the representation is hierarchical for free: a cabinet with three drawers and a door is a tree of parts and joints, and Python expresses trees natively. Feed-forward networks that emit a fixed number of parts cannot.
The script has one rule. Everything to do with pose, the base transform and the joint states, is stored under predefined variable names. That lets the harness decouple pose from shape mechanically: a pose step can update those variables without touching the geometry, and a shape step can rewrite geometry while the pose variables are re-read from disk. The agent's loop alternates between the two, and which kind of step to take next is the agent's own scheduling decision.
A laptop as the harness would store it: geometry as primitives, a hinge joint with limits, and the pose fields under their reserved names. The sliders are a pose step: they rewrite only the reserved variables and the render follows. The button is a shape step: it adds a primitive the observation history says is missing, and the diff is what the agent would actually commit. Notice that the hinge is clamped to its limits no matter what you ask for.
Why not a mesh, a neural field or a point cloud? All three can represent the geometry. None of them can be edited by intent. “The loop is misplaced, move it to the other handle” is one line in a script and an ill-posed optimisation over a mesh. And none of them carries joints: a point cloud does not know it has a hinge. The lineage the authors place themselves in is 3D-as-code, from DeepSVG and DeepCAD through SceneScript to Real2Code, which used code for articulated asset generation. AgentSTAR is the first to make the code the thing that gets tracked through a video.
An optimiser needs a number to push on. AgentSTAR's number is deliberately simple. Render the current object at the current pose into camera i, take its silhouette, and measure the Intersection-over-Union with the target object's mask: the area both cover divided by the area either covers. Perfect overlap is 1, no overlap is 0. The lineage is shape-from-silhouette, thirty years old.
One modification makes it usable on manipulation video. Hands cover objects. If the mask says “object here” and the render says “object here” but a thumb is in the way, a naive IoU punishes a correct pose for a pixel it could never have matched. So the harness segments the hands and removes the hand region from both sides before computing the overlap:
where Rs is the silhouette render, Mi the object mask, and Hi the hand mask. The authors are candid that this signal is inherently local: it says nothing about a pose that does not overlap at all, and it has many local optima. That is not a bug to be fixed with a fancier score. It is the reason the VLM is in the loop: the numerical signal is precise near the answer, and the agent's job is to get near the answer.
A target mask (grey) of a mug seen from the side, a hand blob across it, and your rendered candidate (warm outline). Slide the candidate's pose toward the truth and watch two numbers: the naive IoU and the paper's hand-excluded IoU. Then press flat cut-out: the candidate becomes a 2D sheet shaped exactly like the mask. Its score is perfect. The turntable inset on the right shows what the agent's novel-view tool would reveal.
Could the score include more? Pixel matching from a dense correspondence network (RAFT, MASt3R, RoMa) is the obvious addition, and the authors say it might be viable. They leave it out on purpose: a noisy signal means the agent must also reason about when that signal is wrong, and a silhouette is hard to get wrong. There is one sanctioned extension. When depth is available, a depth term enters the score with a blending coefficient; on the RGB-D iTACO benchmark the variant uses an L1 depth residual. Everywhere else in the paper, depth supervision is off.
Here is a detail that looks like bookkeeping and turns out to be about language. The agent is going to ask for rotation corrections in words: “tilt it a little to the left,” or “flip it around its long axis.” Those two sentences describe rotations in different coordinate frames, and because rotations do not commute, applying the same 30-degree turn in the wrong frame gives a different object.
Write the current rotation as R. A camera-centric update rotates the object about the camera's axes, the ones the agent sees on screen: left-multiply, R′ = ΔR · R. That is what “tilt it left as I look at it” means. An object-centric update rotates about the object's own axes: right-multiply, R′ = R · ΔR. That is what “spin it around its own handle” means, and it is how you express a symmetry flip: 180 degrees about the object's own axis, whatever the camera is doing.
The increment itself is built for interpretability. The harness composes three rotations about the camera's Z, Y and X axes in a fixed order, so an agent that says “yaw 15, pitch minus 5” is describing something it can picture. The paper's stated reason is that VLM agents “excel at reasoning in terms such as the object has to be slightly tilted to the left,” and the parameterisation should meet them there.
An L-shaped bracket, deliberately asymmetric, already rotated by a fixed R so its own axes (coloured) no longer line up with the camera's (grey). Pick an axis and an angle. The left result applies your turn about the camera's axis (left-multiply); the right applies it about the object's axis of the same name (right-multiply). The readout is the angle between the two outcomes: zero only when the two frames happen to agree. Try the 180-degree flip preset and see why a symmetry flip must be object-centric.
Now the core loop, the one in the paper's Figure 3. The shape is frozen. The object is rendered at its current pose into the keyframe. The agent looks at the render next to the photo and, instead of guessing numbers, names a region of pose space to search: “vary yaw from minus 15 to 15 degrees and the hinge joint from 24 to 64 degrees.” Or, for a suspected flip: “search the symmetry flips about the long axis.” The searchable variables are interpretable on purpose: yaw, pitch, roll, translation in the image plane, depth, and each articulation state.
A gradient-free optimiser then searches that region. The paper uses differential evolution, and it is worth knowing how simple it is. Keep a population of candidate poses inside the region. For each candidate a, pick three others b, c, d and build a mutant b + F(c − d): a step whose size and direction come from the population's own spread. Mix a few coordinates of the mutant into a, render the result, score it, and keep whichever of the two scores higher. Repeat for some generations. No gradient is ever needed, which matters because a Blender silhouette is not differentiable with respect to pose, and the region is small enough that a few hundred renders cover it.
The optimiser does not return one answer. It returns the K highest-scoring candidates, each with its render and its score, and hands them back to the agent. The agent inspects them visually and picks the best-looking one. That candidate becomes the starting pose for the next iteration. Which brings us to the sentence that separates this from ordinary optimisation:
The paper's own example: scissors with one purple handle. The silhouette score cannot see colour, so the mirror pose, handles swapped, scores exactly as well as the truth; only the photo shows which side is purple. The grey mask is the target keyframe with its purple handle marked; the warm outline is the current pose. Set a search region for yaw and for the hinge (centre ± half-width), optionally ask for the symmetry-flip basin, and run the optimiser: differential evolution renders a population inside your region and returns the three best, with their silhouette IoU. Then do the VLM's job: click the candidate that is actually right. It becomes the new current pose, and you iterate. The hidden truth is revealed in the readout so you can see when a high score is lying.
Two engineering consequences fall out of freezing the shape during pose steps. First, pose refinement parallelises across the video: with the object fixed, keyframe 7 and keyframe 12 can be fitted at the same time. The appendix describes the schema: split the sequence into non-overlapping chunks, spawn one pose sub-agent per chunk with a short free-form brief (“the current poses are in the right basin” or “these need a flip”), and let the orchestrator handle the seams between chunks with the temporal tool of Chapter 6. Image inspection is the expensive operation for these agents, and most can look at only a limited number of images, so spreading the looking across sub-agents is what makes a 15-keyframe sequence tractable.
Second, the shape step is the opposite kind of decision. Shape and kinematics depend on every frame at once, so they stay centralised: one coding agent edits the script, conditioned on the whole observation history, including what the pose steps found. If every pose search keeps hitting a wall on the left handle, that is evidence the left handle is modelled wrong, and the next iteration should be a shape step, not another pose step. The scheduling is the agent's call.
Fit each keyframe on its own and you get a sequence of poses that are individually plausible and collectively wrong. The classic symptom is the symmetry flip: keyframe 6 converges to the scissors pointing left, keyframe 7 to the mirror pose pointing right, both with excellent silhouettes, and every 3D point on the object teleports across the frame between them. For a downstream consumer such as a 3D point-tracking metric, that teleport is a huge error even though each frame's silhouette is perfect.
The textbook fix is a smoothness prior: add a penalty between consecutive poses so the estimator prefers small changes, the way factor-graph state estimators do. The paper rejects it for a reason that HOT3D makes vivid: real hands move objects fast. Between consecutive evaluation keyframes in that dataset, orientation can change by as much as 180 degrees. A prior strong enough to suppress a flip will also suppress a genuine swing, and you will have replaced one error with another that is harder to notice.
Fifteen keyframes of one pose variable, yaw. The dashed line is the truth, which contains a genuine fast swing in the middle. The warm line is a per-frame fit: good silhouettes, but two symmetry flips of 180 degrees. Below it, the diagnostic the agent would read: angular velocity per step, with the flagged frames marked. Press fixed prior to smooth everything with one constant, and watch it shave the swing while only softening the flips. Press agentic to correct only the flagged frames. The readout is the mean error against the truth in degrees.
What exactly is in the diagnostic? The tool measures velocities and accelerations component by component rather than one distance over the whole configuration. For rotation it first removes the camera's own motion, forming the object's world orientation as the camera rotation times the object rotation, then takes the geodesic angle between consecutive world orientations, in degrees. For translation there is a subtlety worth understanding. A dynamic object in a monocular video has its own scale gauge, independent of the camera's, so the object's translation is reparameterised as scale times a unit-free offset, and the residuals compare those offsets rotated into world-aligned axes. The camera's translation is left out entirely, because the SLAM system's translation units need not be compatible with the object's reconstruction scale, and mixing them would manufacture fake accelerations.
Accelerations are second differences of those offsets. There is also a radial component, the velocity and acceleration along the camera-to-object direction, singled out because depth is the least constrained direction for a monocular method: a silhouette barely changes when an object slides toward the camera a little. Joint residuals are first and second differences of each joint state in its native unit, degrees for a hinge, object units for a slider. All of it is per step, and all of it goes to the agent as numbers it can read.
Strip away the results and what the paper actually built is a harness: a set of tools that let a general-purpose coding agent behave like an optimiser. The agent is not fine-tuned. The same prompts and the same tools run every experiment. The list, in the authors' own order, omitting infrastructure:
The default runs use the Codex harness with the paper's tool suite and GPT-5.6 Sol at medium reasoning effort, with a ten-hour budget per optimisation loop; some agents declare themselves done earlier. Swapping the underlying model for Claude Fable 5 with the Claude Code harness and the same tools lowered ARCTIC tracking error from 5.59 to 4.95 cm, which tells you the harness, not the specific model, is the contribution, and that the ceiling is still rising with the models.
Left: the paper's Figure 5, tracking error against elapsed hours since the agent launched, redrawn from its reported endpoints (the dashed line is the no-harness baseline, 11.26 cm; the curve settles at 5.59). Scrub the hours. Right: how the appendix's parallelisation schedules the work. Shape and kinematics are centralised decisions that depend on all frames, so shape steps run serially; once the shape is coherent, the keyframes are split into M chunks and a pose sub-agent owns each, with the orchestrator checking the seams with the temporal tool. Slide M and watch the wall clock. The curve shape and the schedule are illustrative; the endpoints and the roles are the paper's.
The ablation in Table 2 is the cleanest reading of what each piece buys, all on ARCTIC 3D tracking error in centimetres:
| Variant | 3D EPE (cm) | What it removes |
|---|---|---|
| No harness | 11.26 | the same agent, same inputs and output format, no tools at all |
| IoU objective only | 151.46 | everything but the score: the agent hacks it with flat silhouette geometry |
| GT mesh + VLM-free optimisation | 14.60 | the VLM's basin choice: perfect geometry, but differential evolution over the whole pose space |
| No temporal tool | 6.15 | sequence-level reasoning |
| Full (GPT-5.6 Sol) | 5.59 | — |
| Full (Claude Fable 5) | 4.95 | same tools, different agent |
Two more facts about the loop's behaviour. Its error decreases steadily over the run, the convergence profile of a conventional iterative optimiser rather than a lucky guess. And although agentic pipelines are stochastic, three independent runs over the whole dataset gave a standard deviation of 0.06 cm on the 5.59 figure: small enough that every comparison in the next chapter is real.
Three benchmarks, chosen so that each tests a different claim. ARCTIC is people manipulating articulated objects with a motion-capture rig providing ground-truth geometry and motion: the articulation claim. HOT3D is rigid objects moved fast under an egocentric camera: the rigid-tracking claim, against methods that are given the CAD model. iTACO is a simulated RGB-D benchmark for articulated digital twins: the kinematics claim, against methods built for exactly that.
ARCTIC protocol. The egocentric sequences of subject s1, 24 in all. The first 100 frames of each are discarded (capture setup, little motion, underexposure), the rest split into 300-frame chunks, and every tenth frame is a keyframe; every method is scored on the same keyframes. Because monocular methods recover geometry only up to scale, one shared scale per chunk is fixed by aligning predicted and ground-truth depth through the median of their log ratio, robustly, for every method without metric depth. The tracking metric is 3D end-point error for query pixels inside the object mask in the first keyframe. The baselines are pixel-anchored, so their answer is a trajectory per pixel. AgentSTAR's answer is a mesh, so each query pixel is matched to the nearest projected mesh point in the first keyframe and that point is followed through the sequence.
Every headline table as bars. Lower is better on every metric here. Methods marked * were given the ground-truth CAD model. The ablation's 151 cm bar is cut off; the number is real.
On ARCTIC the numbers are 5.59 cm against 7.65 for V-DPM, 8.77 for OpenD4RT and 10.79 for SpatialTrackerV2, and the reconstructed geometry is closer too: Chamfer distance 3.36 cm against V-DPM's 4.72, averaged over keyframes, mesh against mesh for AgentSTAR and point cloud against mesh for the baseline. These are methods that estimate dense 3D motion directly, beaten by one that never estimates a correspondence.
HOT3D protocol. The validation split, keeping sequences with a single dynamic target: 93 sequences of 150 frames, every tenth frame a keyframe, so 15 per sequence. The video baselines may process all 150 frames; AgentSTAR sees only the keyframes; everyone is scored on the keyframes. Two rules make this harder than it looks. Symmetries are not quotiented out frame by frame: a prediction that switches between symmetry-equivalent poses over time is wrong, because the switch changes the 3D point trajectory. And for model-free methods the canonical object frame is arbitrary, so one global rotation and translation is fitted between predicted and true trajectories and then frozen for the whole sequence, with one scale from depth alignment.
| Method | Trans. mean (cm) | Trans. median | Rot. mean (°) | Rot. median |
|---|---|---|---|---|
| ProxyPose | – | – | 52.7 | 43.4 |
| GigaPose * | 11.25 | 8.16 | 62.8 | 50.9 |
| GigaPose * + GoTrack | 16.46 | 9.92 | 55.4 | 37.9 |
| FoundationPose * + VGGT | 7.56 | 5.57 | 54.6 | 49.1 |
| SAM3D-Tracker | 6.41 | 5.30 | 51.6 | 43.0 |
| AgentSTAR | 3.04 | 2.42 | 37.6 | 26.3 |
The rotation column deserves an honest reading. Every method's mean rotation error is large, AgentSTAR's included, because HOT3D's objects can turn by up to 180 degrees between consecutive keyframes. The paper says so plainly: this regime is much harder than the smooth-motion settings rigid trackers are usually measured on. The claim is not that rotation is solved; it is that with no CAD model, AgentSTAR beats the methods that had one.
The paper names its main limitation in one sentence: inference-time computational cost. Ten hours of agent time per video, with hundreds of Blender renders per pose search and a language model reading images throughout, is not a tracker you run on a robot. It is a tracker you run on a cluster, once, to find out what happened.
The authors' proposal for what to do with that is the most interesting paragraph in the conclusion. A system this expensive and this accurate is a natural teacher. Run it over a corpus of videos, collect the structured reconstructions and trajectories it produces, and train the next feed-forward perception model on them. Then use that faster model to initialise the agent, run it again, and repeat. “Alternating between agentic optimisation and feed-forward learning could provide a path toward iterative self-improvement.” That is a proposal, not a result; the paper reports no student model.
Every ARCTIC method on one plane: tracking error against the compute it takes to process one sequence. The errors are the paper's; the timings are order-of-magnitude placements (feed-forward trackers run in seconds to minutes, the agent has a ten-hour budget) and are labelled as such. Toggle the proposal to see the teacher-student loop the conclusion sketches: the slider is how many videos you would label with the agent, and the hypothetical student point is exactly that, hypothetical.
Other limits the paper states or implies, worth carrying with you. The method needs an object mask per frame and known camera motion; both come from other systems (SAM 3, a SLAM or feed-forward reconstruction such as VGGT), and their failures are inherited. Transparent objects work precisely because a segmentation mask exists where texture does not, so the method is only as good as the mask on glass. Rotation error on fast rigid motion is still 37.6 degrees on average. And the kinematic prior can over-explain: the two iTACO sequences with an invented extra axis are the shape of that failure.
| Piece | What it is | Why that choice |
|---|---|---|
| Object O | a Python script of Blender primitives plus joints, limits and axes | editable by intent as code diffs; hierarchical for free; renderable from any view |
| Generalised pose Ti | 6-DoF base (quaternion + translation) + one scalar per joint, one shared scale | compact, interpretable, constrained by the mechanism |
| Score S(O,T) | IoU of the silhouette render and the mask, hands removed from both | precise near the answer, cheap, hard to get wrong; local optima are the VLM's job |
| Camera-centric update | left-multiply: ΔR · R | “tilt it left as I see it” |
| Object-centric update | right-multiply: R · ΔR | “flip it about its own axis”; symmetry flips |
| Pose tool | agent names a bounded region; differential evolution returns top-K renders + scores; agent picks | the VLM chooses the basin, the optimiser refines; the pick need not be the numerical best |
| Temporal residuals | per-component velocities and accelerations, camera motion removed, camera translation excluded | diagnostics, not a prior: real fast motion survives, flips get flagged |
| Sub-agents | keyframes split into chunks, one pose agent each; shape stays central | image inspection is the bottleneck; pose is parallel once shape is fixed |