Kirill Mazur, Nikita Karaev, Matthew Chang, Jitendra Malik, Nur Muhammad “Mahi” Shafiullah (Amazon FAR · Frontier AI and Robotics) — 2026, agenticstar.github.io

Which Object Made This Video?

Point trackers ask where every pixel went. AgentSTAR asks which structured 3D object, with which joints and which pose at each moment, could have produced the frames, and lets a vision-language agent steer a numerical optimiser through a render-and-compare loop until the answer fits. It tracks through occlusion, symmetry flips and glass, and cuts the best 3D tracker's error on ARCTIC from 7.65 to 5.59 centimetres.

Prerequisites: what a rotation matrix and a 6-DoF pose are + what a segmentation mask is. Differential evolution, kinematic chains, silhouette IoU and temporal residuals are built here from zero.
10
Chapters
10
Simulations
5.59 cm
ARCTIC 3D tracking error (best tracker: 7.65)
10 h
Agent budget per video, the honest cost

Chapter 0: The Problem

Someone closes a book in front of a camera. For the first half of the motion you can see both covers and most of the pages. Then the pages slide behind the front cover, the text disappears, and by the last frame the book is a closed slab. Now ask the question every dynamic-reconstruction system has to answer: where did the pages go?

There are two ways to answer it, and the difference between them is the whole paper. The bottom-up way, which nearly every modern method takes, first recovers low-level visual evidence: scene flow, or the 3D trajectory of every pixel you can see. Then it tries to infer what object produced that evidence. The top-down way first recognises the object and its structure, a book with two hinged covers and a stack of pages, and uses that structure to recover the state over time: one hinge angle, one base pose, per frame.

Bottom-up has two failures the authors care about. First, dense correspondence is itself hard exactly when it matters: under occlusion and low visual overlap, which is the normal condition when a hand is manipulating something. A page you cannot see has no trajectory. Second, even when it works, the output is unstructured. A cloud of moving points tells you where visual evidence went, not what the parts, joints and state variables are. That structure has to be inferred afterwards from imperfect geometry, which is brittle, and it is precisely what a robot needs: an explicit object model it can load into a simulator and manipulate.

The sentence the method is built on. “Much of an object's motion is determined not by visual evidence alone, but by its internal mechanism.” When the pages become occluded, their possible motion is still strongly constrained by the structure of the book. So tracking should be done by the same entity that inferred the shape, because only that entity knows the mechanism.
Closing the Book

A book with one hinge, seen from the side. Drag the hinge angle. In bottom-up mode, tracked points live only where pixels are visible; watch them vanish as the front cover swings over the pages and the count of surviving tracks falls. In top-down mode, the same frames are explained by a single hinge angle, and every point on every page, visible or not, has a predicted position. The predicted page tips are drawn hollow where nothing can see them.

The paper's answer is to directly ask which structured 3D object and state sequence could have generated the video. That is analysis-by-synthesis: propose an object, render it, compare with the frames, adjust. Once the structure is known, the permissible configurations collapse from millions of independent point motions to a compact vector: a 6-DoF base pose plus one number per joint. The states are interpretable too. “The lid is at 64 degrees” is something a planner can use; a scene-flow field is not.

What comes in. A set of video frames, optionally depth maps, and a binary mask of the target object in each frame, typically from a video segmentation model such as SAM 3. Hands are segmented too, because they occlude the object. The camera intrinsics are known and the camera's own motion comes from an external SLAM or feed-forward reconstruction system such as VGGT. Everything about the object, its geometry, its joints, its pose over time, is what has to be recovered.
A hand closes a laptop and the keyboard disappears behind the screen for the last twenty frames. Why does the paper argue a top-down model should keep tracking the keyboard when a point tracker cannot?

Chapter 1: The Insight

Analysis-by-synthesis is an old idea. Why does it work now, on casually captured video of scissors and ketchup bottles, when it did not before? Because two very different kinds of estimator have become available, and each is bad at exactly what the other is good at.

A large vision-language model, the kind that sits inside modern coding agents, is excellent at coarse judgement. Show it a render next to a photo and it will tell you the scissors are upside down, that the loop is on the wrong handle, that the lid looks too thick. It is also good at proposing rough 3D shape. But it is a poor state estimator: it cannot recover a pose to within a few degrees, and it cannot reliably produce continuous numbers. Ask it “what is the yaw?” and you get a guess.

Classical test-time optimisation, bundle adjustment and its relatives, is the opposite. Give it a differentiable or at least scoreable objective and it will find a local optimum to arbitrary precision. But the objective for “does this render match this mask” is riddled with local optima. A ketchup bottle rendered cap-up scores almost as well as cap-down. A box flipped by 180 degrees has the same silhouette. Start the optimiser in the wrong basin and it converges, precisely, to the wrong answer.

The pairing. The VLM does not estimate the pose. It chooses the basin: it narrows the search to a promising region and rejects the spurious optima the score cannot tell apart. The numerical optimiser then does what it is good at inside that region. In the authors' words, the VLM makes gradient-free optimisation tractable, “which is often the hardest part of the problem.” Once the right basin is found, the parameter space, a handful of shape parameters, joint axes and their states, is small enough that refinement is easy.
The Landscape the Optimiser Sees

The silhouette overlap between a rendered ketchup bottle and the target mask, as a function of one pose variable, the in-plane yaw. The true yaw is marked. Because the body is symmetric and only the cap breaks the symmetry, there are several peaks that look almost as good numerically. Drop a local optimiser anywhere with start here and it climbs to the nearest peak. Then ask the VLM: it cannot say the yaw is 37 degrees, but it can look at the current render and say “the cap is at the wrong end, flip it,” which teleports the optimiser into the correct basin, where a second climb finishes the job.

Notice what the VLM's move is not. It is not a gradient. It is not a number. It is a discrete, semantic decision, “wrong basin, try the flip,” of the kind these models make reliably, applied to a problem where that decision was the expensive part. Everything else in the paper is the engineering that lets a coding agent make such decisions safely: a representation it can edit as code, a score it can read, tools that search regions it names, and diagnostics that tell it when time has gone wrong.

What the paper claims, precisely. To the authors' knowledge, this is the first work to apply VLM agents to joint shape and motion reconstruction of both rigid and articulated objects from casually captured monocular video. The evidence is three benchmarks: it beats 3D point-tracking baselines on ARCTIC, articulated-reconstruction methods on iTACO, and rigid-object trackers on HOT3D. And because it models the cause of motion rather than its visual evidence, it can track things prior methods cannot see well at all: transparent glass, and objects that are severely occluded or only partly visible.
The silhouette score has several near-equal peaks because of the object's symmetry. In the paper's division of labour, who resolves that ambiguity and how?

Chapter 2: The Object as Code

If an agent is going to propose an object, render it and edit it, the object has to live in a form the agent can write. AgentSTAR's answer is the most direct one available to a coding agent: the object is a Python script. Geometry is expressed with Blender shape primitives, which is freeform enough to build arbitrary topologies: a loop for a scissor handle, a lid on a cylinder, a stack of thin slabs for pages. The kinematic structure lives in the same script: which part is the base, which parts are joints, where each joint sits, what its axis is, and what its motion limits are.

Two things follow from that choice. First, every shape edit is a code diff, which is exactly the operation a coding agent performs best and can reason about in text. Adding a missing handle, moving a hinge, changing a joint from revolute to prismatic, each is a few lines. Second, the representation is hierarchical for free: a cabinet with three drawers and a door is a tree of parts and joints, and Python expresses trees natively. Feed-forward networks that emit a fixed number of parts cannot.

The script has one rule. Everything to do with pose, the base transform and the joint states, is stored under predefined variable names. That lets the harness decouple pose from shape mechanically: a pose step can update those variables without touching the geometry, and a shape step can rewrite geometry while the pose variables are re-read from disk. The agent's loop alternates between the two, and which kind of step to take next is the agent's own scheduling decision.

The state vector, counted. For an articulated object the generalised pose at keyframe i is the 6-DoF pose of its base, a rotation stored as a unit quaternion plus a translation vector, together with one scalar per articulation joint, each constrained to its motion limits. A pair of scissors is 6 + 1 = 7 numbers per keyframe. A cabinet with two drawers and a door is 6 + 3 = 9. One scale is shared across all frames, because a monocular video cannot tell a big object far away from a small one nearby. Poses are object-to-camera: a canonical point pc, after forward kinematics through the joints, lands in camera i at s Ri pc + ti.
scene.py, Rendered

A laptop as the harness would store it: geometry as primitives, a hinge joint with limits, and the pose fields under their reserved names. The sliders are a pose step: they rewrite only the reserved variables and the render follows. The button is a shape step: it adds a primitive the observation history says is missing, and the diff is what the agent would actually commit. Notice that the hinge is clamped to its limits no matter what you ask for.

Why not a mesh, a neural field or a point cloud? All three can represent the geometry. None of them can be edited by intent. “The loop is misplaced, move it to the other handle” is one line in a script and an ill-posed optimisation over a mesh. And none of them carries joints: a point cloud does not know it has a hinge. The lineage the authors place themselves in is 3D-as-code, from DeepSVG and DeepCAD through SceneScript to Real2Code, which used code for articulated asset generation. AgentSTAR is the first to make the code the thing that gets tracked through a video.

Diagnostics the agent gets for free. Because the object is a real Blender scene, the harness can render it from the observed cameras and from novel viewpoints. A turntable tool renders views on a sphere around the object so the agent can inspect a shape from angles the video never showed. That is how a flat cut-out that fools the silhouette score gets caught: it looks right from the camera and like a sheet of paper from the side. Chapter 3 shows why that matters.
A drawer unit has two sliding drawers and one hinged door. At a single keyframe, how many numbers make up its generalised pose in AgentSTAR's representation, and what is the one number that is not per-keyframe?

Chapter 3: The Score

An optimiser needs a number to push on. AgentSTAR's number is deliberately simple. Render the current object at the current pose into camera i, take its silhouette, and measure the Intersection-over-Union with the target object's mask: the area both cover divided by the area either covers. Perfect overlap is 1, no overlap is 0. The lineage is shape-from-silhouette, thirty years old.

One modification makes it usable on manipulation video. Hands cover objects. If the mask says “object here” and the render says “object here” but a thumb is in the way, a naive IoU punishes a correct pose for a pixel it could never have matched. So the harness segments the hands and removes the hand region from both sides before computing the overlap:

S(O, T) = IoU( Rs(O, T) ∖ Hi ,  Mi ∖ Hi )

where Rs is the silhouette render, Mi the object mask, and Hi the hand mask. The authors are candid that this signal is inherently local: it says nothing about a pose that does not overlap at all, and it has many local optima. That is not a bug to be fixed with a fancier score. It is the reason the VLM is in the loop: the numerical signal is precise near the answer, and the agent's job is to get near the answer.

Silhouette Overlap, With and Without Hands

A target mask (grey) of a mug seen from the side, a hand blob across it, and your rendered candidate (warm outline). Slide the candidate's pose toward the truth and watch two numbers: the naive IoU and the paper's hand-excluded IoU. Then press flat cut-out: the candidate becomes a 2D sheet shaped exactly like the mask. Its score is perfect. The turntable inset on the right shows what the agent's novel-view tool would reveal.

The 151-centimetre lesson. In the ablation, an agent given the task, the inputs and only the numerical IoU tool produced a 3D tracking error of 151.46 cm, against 5.59 for the full system and 11.26 for an agent with no tools at all. It was not failing to optimise. It was optimising too well: it learned to emit flat, silhouette-shaped geometry that scored high IoU without being the object. A scalar objective plus a capable agent is a recipe for reward hacking, and the rest of the harness, the turntable, the external VLM critic, the temporal tool, exists partly to make that hack visible.

Could the score include more? Pixel matching from a dense correspondence network (RAFT, MASt3R, RoMa) is the obvious addition, and the authors say it might be viable. They leave it out on purpose: a noisy signal means the agent must also reason about when that signal is wrong, and a silhouette is hard to get wrong. There is one sanctioned extension. When depth is available, a depth term enters the score with a blending coefficient; on the RGB-D iTACO benchmark the variant uses an L1 depth residual. Everywhere else in the paper, depth supervision is off.

An agent with the IoU tool alone scored 151 cm of tracking error, far worse than an agent with no tools (11 cm). What does that result say about the score?

Chapter 4: Rotating Right

Here is a detail that looks like bookkeeping and turns out to be about language. The agent is going to ask for rotation corrections in words: “tilt it a little to the left,” or “flip it around its long axis.” Those two sentences describe rotations in different coordinate frames, and because rotations do not commute, applying the same 30-degree turn in the wrong frame gives a different object.

Write the current rotation as R. A camera-centric update rotates the object about the camera's axes, the ones the agent sees on screen: left-multiply, R′ = ΔR · R. That is what “tilt it left as I look at it” means. An object-centric update rotates about the object's own axes: right-multiply, R′ = R · ΔR. That is what “spin it around its own handle” means, and it is how you express a symmetry flip: 180 degrees about the object's own axis, whatever the camera is doing.

camera-centric:  p = s (ΔR) R pc + t + Δt      object-centric:  p = s R (ΔR) pc + t

The increment itself is built for interpretability. The harness composes three rotations about the camera's Z, Y and X axes in a fixed order, so an agent that says “yaw 15, pitch minus 5” is describing something it can picture. The paper's stated reason is that VLM agents “excel at reasoning in terms such as the object has to be slightly tilted to the left,” and the parameterisation should meet them there.

Same Turn, Two Frames

An L-shaped bracket, deliberately asymmetric, already rotated by a fixed R so its own axes (coloured) no longer line up with the camera's (grey). Pick an axis and an angle. The left result applies your turn about the camera's axis (left-multiply); the right applies it about the object's axis of the same name (right-multiply). The readout is the angle between the two outcomes: zero only when the two frames happen to agree. Try the 180-degree flip preset and see why a symmetry flip must be object-centric.

Why expose both. Small corrections are naturally camera-centric: the agent compares a render with a photo and sees that the object should lean left on the image. Large, discrete corrections are naturally object-centric: a flip is a statement about the object's own symmetry and must survive whatever the camera has done. A harness that offered only one convention would force the agent to translate, and translation between rotation frames is exactly the kind of precise numerical reasoning the paper says VLMs are bad at.
The agent's diagnostic says “silhouette OK, but the blades point the wrong way: there is a flip.” Which update does the fix use, and why that one?

Chapter 5: The Agent and the Optimiser

Now the core loop, the one in the paper's Figure 3. The shape is frozen. The object is rendered at its current pose into the keyframe. The agent looks at the render next to the photo and, instead of guessing numbers, names a region of pose space to search: “vary yaw from minus 15 to 15 degrees and the hinge joint from 24 to 64 degrees.” Or, for a suspected flip: “search the symmetry flips about the long axis.” The searchable variables are interpretable on purpose: yaw, pitch, roll, translation in the image plane, depth, and each articulation state.

A gradient-free optimiser then searches that region. The paper uses differential evolution, and it is worth knowing how simple it is. Keep a population of candidate poses inside the region. For each candidate a, pick three others b, c, d and build a mutant b + F(c − d): a step whose size and direction come from the population's own spread. Mix a few coordinates of the mutant into a, render the result, score it, and keep whichever of the two scores higher. Repeat for some generations. No gradient is ever needed, which matters because a Blender silhouette is not differentiable with respect to pose, and the region is small enough that a few hundred renders cover it.

The optimiser does not return one answer. It returns the K highest-scoring candidates, each with its render and its score, and hands them back to the agent. The agent inspects them visually and picks the best-looking one. That candidate becomes the starting pose for the next iteration. Which brings us to the sentence that separates this from ordinary optimisation:

“The selected pose need not be the numerical optimum under S.” In Figure 3 the three returned candidates score 0.97, 0.96 and 0.96, and the agent chooses among them by eye. A score of 0.97 attached to a mirror-image pose is worth less than 0.96 attached to the right one, and only the VLM can tell. The number proposes; the agent disposes.
Render, Search, Inspect, Pick

The paper's own example: scissors with one purple handle. The silhouette score cannot see colour, so the mirror pose, handles swapped, scores exactly as well as the truth; only the photo shows which side is purple. The grey mask is the target keyframe with its purple handle marked; the warm outline is the current pose. Set a search region for yaw and for the hinge (centre ± half-width), optionally ask for the symmetry-flip basin, and run the optimiser: differential evolution renders a population inside your region and returns the three best, with their silhouette IoU. Then do the VLM's job: click the candidate that is actually right. It becomes the new current pose, and you iterate. The hidden truth is revealed in the readout so you can see when a high score is lying.

Two engineering consequences fall out of freezing the shape during pose steps. First, pose refinement parallelises across the video: with the object fixed, keyframe 7 and keyframe 12 can be fitted at the same time. The appendix describes the schema: split the sequence into non-overlapping chunks, spawn one pose sub-agent per chunk with a short free-form brief (“the current poses are in the right basin” or “these need a flip”), and let the orchestrator handle the seams between chunks with the temporal tool of Chapter 6. Image inspection is the expensive operation for these agents, and most can look at only a limited number of images, so spreading the looking across sub-agents is what makes a 15-keyframe sequence tractable.

Second, the shape step is the opposite kind of decision. Shape and kinematics depend on every frame at once, so they stay centralised: one coding agent edits the script, conditioned on the whole observation history, including what the pose steps found. If every pose search keeps hitting a wall on the left handle, that is evidence the left handle is modelled wrong, and the next iteration should be a shape step, not another pose step. The scheduling is the agent's call.

Concept meets realisation: what one iteration costs. One pose search is a few hundred Blender silhouette renders at low resolution, scored by IoU, returned as three images plus three numbers. One shape step is a code diff plus a turntable of novel-view renders. The agent's context sees images and short numbers, never dense tensors. That is why a general-purpose coding harness (Codex with the paper's tool suite, running GPT-5.6 Sol at medium reasoning effort) can drive it at all, and why the whole optimisation for one video fits a ten-hour budget.
Why does the pose tool hand the agent the top K candidates with renders, rather than simply applying the single highest-scoring pose?

Chapter 6: Time

Fit each keyframe on its own and you get a sequence of poses that are individually plausible and collectively wrong. The classic symptom is the symmetry flip: keyframe 6 converges to the scissors pointing left, keyframe 7 to the mirror pose pointing right, both with excellent silhouettes, and every 3D point on the object teleports across the frame between them. For a downstream consumer such as a 3D point-tracking metric, that teleport is a huge error even though each frame's silhouette is perfect.

The textbook fix is a smoothness prior: add a penalty between consecutive poses so the estimator prefers small changes, the way factor-graph state estimators do. The paper rejects it for a reason that HOT3D makes vivid: real hands move objects fast. Between consecutive evaluation keyframes in that dataset, orientation can change by as much as 180 degrees. A prior strong enough to suppress a flip will also suppress a genuine swing, and you will have replaced one error with another that is harder to notice.

Agentic smoothing. Instead of a fixed prior, temporal inconsistencies are reported. A mandatory diagnostic tool computes residuals along the trajectory and shows the agent where the motion looks implausible. The agent then decides, frame by frame, whether a spike is a flip to correct or a rapid motion the pixels support. “This allows temporal smoothing to be applied selectively rather than imposed as a fixed prior.”
A Flip, a Swing, and Two Ways to Smooth

Fifteen keyframes of one pose variable, yaw. The dashed line is the truth, which contains a genuine fast swing in the middle. The warm line is a per-frame fit: good silhouettes, but two symmetry flips of 180 degrees. Below it, the diagnostic the agent would read: angular velocity per step, with the flagged frames marked. Press fixed prior to smooth everything with one constant, and watch it shave the swing while only softening the flips. Press agentic to correct only the flagged frames. The readout is the mean error against the truth in degrees.

What exactly is in the diagnostic? The tool measures velocities and accelerations component by component rather than one distance over the whole configuration. For rotation it first removes the camera's own motion, forming the object's world orientation as the camera rotation times the object rotation, then takes the geodesic angle between consecutive world orientations, in degrees. For translation there is a subtlety worth understanding. A dynamic object in a monocular video has its own scale gauge, independent of the camera's, so the object's translation is reparameterised as scale times a unit-free offset, and the residuals compare those offsets rotated into world-aligned axes. The camera's translation is left out entirely, because the SLAM system's translation units need not be compatible with the object's reconstruction scale, and mixing them would manufacture fake accelerations.

Accelerations are second differences of those offsets. There is also a radial component, the velocity and acceleration along the camera-to-object direction, singled out because depth is the least constrained direction for a monocular method: a silhouette barely changes when an object slides toward the camera a little. Joint residuals are first and second differences of each joint state in its native unit, degrees for a hinge, object units for a slider. All of it is per step, and all of it goes to the agent as numbers it can read.

Does it help? In the ablation, removing the temporal reasoning tool raises ARCTIC tracking error from 5.59 to 6.15 cm. Modest next to the 151 cm of the IoU-only variant, and that is the point: the temporal tool is not what makes the method work, it is what makes the trajectories consistent enough to be measured as tracks.
Why does the temporal residual tool exclude the camera's translation when computing the object's world-frame velocities?

Chapter 7: The Harness

Strip away the results and what the paper actually built is a harness: a set of tools that let a general-purpose coding agent behave like an optimiser. The agent is not fine-tuned. The same prompts and the same tools run every experiment. The list, in the authors' own order, omitting infrastructure:

Blender rendering
parallel render utilities plus visualisation scripts, the eyes of the loop
↓
Silhouette scoring
the hand-excluded IoU of Chapter 3
↓
Pose tool
bounded-region differential evolution, top-K back with renders (Chapter 5)
↓
Temporal residuals
velocities and accelerations per component, mandatory for video (Chapter 6)
↓
Turntable
renders on a sphere around the object, so shape can be judged from unseen angles
↓
External VLM critic
a second model asked to judge shape quality, a check on the first one's optimism

The default runs use the Codex harness with the paper's tool suite and GPT-5.6 Sol at medium reasoning effort, with a ten-hour budget per optimisation loop; some agents declare themselves done earlier. Swapping the underlying model for Claude Fable 5 with the Claude Code harness and the same tools lowered ARCTIC tracking error from 5.59 to 4.95 cm, which tells you the harness, not the specific model, is the contribution, and that the ceiling is still rising with the models.

Ten Hours of Optimisation, and Who Does What

Left: the paper's Figure 5, tracking error against elapsed hours since the agent launched, redrawn from its reported endpoints (the dashed line is the no-harness baseline, 11.26 cm; the curve settles at 5.59). Scrub the hours. Right: how the appendix's parallelisation schedules the work. Shape and kinematics are centralised decisions that depend on all frames, so shape steps run serially; once the shape is coherent, the keyframes are split into M chunks and a pose sub-agent owns each, with the orchestrator checking the seams with the temporal tool. Slide M and watch the wall clock. The curve shape and the schedule are illustrative; the endpoints and the roles are the paper's.

The ablation in Table 2 is the cleanest reading of what each piece buys, all on ARCTIC 3D tracking error in centimetres:

Variant3D EPE (cm)What it removes
No harness11.26the same agent, same inputs and output format, no tools at all
IoU objective only151.46everything but the score: the agent hacks it with flat silhouette geometry
GT mesh + VLM-free optimisation14.60the VLM's basin choice: perfect geometry, but differential evolution over the whole pose space
No temporal tool6.15sequence-level reasoning
Full (GPT-5.6 Sol)5.59—
Full (Claude Fable 5)4.95same tools, different agent
The row that proves the thesis. Give the numerical optimiser the ground-truth mesh, an advantage no real system has, and let it search the full pose space without a VLM. It scores 14.60 cm, worse than the full method's 5.59 with a shape the agent had to invent. Precision was never the bottleneck. Knowing where to look was.

Two more facts about the loop's behaviour. Its error decreases steadily over the run, the convergence profile of a conventional iterative optimiser rather than a lucky guess. And although agentic pipelines are stochastic, three independent runs over the whole dataset gave a standard deviation of 0.06 cm on the 5.59 figure: small enough that every comparison in the next chapter is real.

The “GT mesh + VLM-free optimisation” variant has the true geometry and still scores 14.60 cm, worse than the full system's 5.59. What does this isolate?

Chapter 8: Results

Three benchmarks, chosen so that each tests a different claim. ARCTIC is people manipulating articulated objects with a motion-capture rig providing ground-truth geometry and motion: the articulation claim. HOT3D is rigid objects moved fast under an egocentric camera: the rigid-tracking claim, against methods that are given the CAD model. iTACO is a simulated RGB-D benchmark for articulated digital twins: the kinematics claim, against methods built for exactly that.

ARCTIC protocol. The egocentric sequences of subject s1, 24 in all. The first 100 frames of each are discarded (capture setup, little motion, underexposure), the rest split into 300-frame chunks, and every tenth frame is a keyframe; every method is scored on the same keyframes. Because monocular methods recover geometry only up to scale, one shared scale per chunk is fixed by aligning predicted and ground-truth depth through the median of their log ratio, robustly, for every method without metric depth. The tracking metric is 3D end-point error for query pixels inside the object mask in the first keyframe. The baselines are pixel-anchored, so their answer is a trajectory per pixel. AgentSTAR's answer is a mesh, so each query pixel is matched to the nearest projected mesh point in the first keyframe and that point is followed through the sequence.

The Benchmark Board

Every headline table as bars. Lower is better on every metric here. Methods marked * were given the ground-truth CAD model. The ablation's 151 cm bar is cut off; the number is real.

On ARCTIC the numbers are 5.59 cm against 7.65 for V-DPM, 8.77 for OpenD4RT and 10.79 for SpatialTrackerV2, and the reconstructed geometry is closer too: Chamfer distance 3.36 cm against V-DPM's 4.72, averaged over keyframes, mesh against mesh for AgentSTAR and point cloud against mesh for the baseline. These are methods that estimate dense 3D motion directly, beaten by one that never estimates a correspondence.

HOT3D protocol. The validation split, keeping sequences with a single dynamic target: 93 sequences of 150 frames, every tenth frame a keyframe, so 15 per sequence. The video baselines may process all 150 frames; AgentSTAR sees only the keyframes; everyone is scored on the keyframes. Two rules make this harder than it looks. Symmetries are not quotiented out frame by frame: a prediction that switches between symmetry-equivalent poses over time is wrong, because the switch changes the 3D point trajectory. And for model-free methods the canonical object frame is arbitrary, so one global rotation and translation is fitted between predicted and true trajectories and then frozen for the whole sequence, with one scale from depth alignment.

MethodTrans. mean (cm)Trans. medianRot. mean (°)Rot. median
ProxyPose––52.743.4
GigaPose *11.258.1662.850.9
GigaPose * + GoTrack16.469.9255.437.9
FoundationPose * + VGGT7.565.5754.649.1
SAM3D-Tracker6.415.3051.643.0
AgentSTAR3.042.4237.626.3

The rotation column deserves an honest reading. Every method's mean rotation error is large, AgentSTAR's included, because HOT3D's objects can turn by up to 180 degrees between consecutive keyframes. The paper says so plainly: this regime is much harder than the smooth-motion settings rigid trackers are usually measured on. The claim is not that rotation is solved; it is that with no CAD model, AgentSTAR beats the methods that had one.

iTACO, the kinematics test. Simulated RGB-D, 73 sequences, the depth variant of the score with an L1 depth term, meshes scale-aligned to the input depth afterwards. AgentSTAR posts the best number on every joint metric: revolute axis error 0.08 radians against iTACO's 0.32, joint position 0.05 m against 0.13, joint state 0.15 against 0.25 radians; prismatic axis 0.07 against 0.24 radians and state 0.04 against 0.08 m. On geometry it wins two of three Chamfer metrics and is a close second on the third (0.02 against 0.01 m² whole-object). One failure mode the metrics cannot see: in 2 of the 73 sequences, storage furniture, it predicts an extra moving axis that does not exist.
On HOT3D a prediction that alternates between two symmetry-equivalent poses of a symmetric object across frames is scored as wrong, even though each frame's pose is individually valid. Why does the paper insist on that rule?

Chapter 9: Limits and Links

The paper names its main limitation in one sentence: inference-time computational cost. Ten hours of agent time per video, with hundreds of Blender renders per pose search and a language model reading images throughout, is not a tracker you run on a robot. It is a tracker you run on a cluster, once, to find out what happened.

The authors' proposal for what to do with that is the most interesting paragraph in the conclusion. A system this expensive and this accurate is a natural teacher. Run it over a corpus of videos, collect the structured reconstructions and trajectories it produces, and train the next feed-forward perception model on them. Then use that faster model to initialise the agent, run it again, and repeat. “Alternating between agentic optimisation and feed-forward learning could provide a path toward iterative self-improvement.” That is a proposal, not a result; the paper reports no student model.

Accuracy Bought With Hours

Every ARCTIC method on one plane: tracking error against the compute it takes to process one sequence. The errors are the paper's; the timings are order-of-magnitude placements (feed-forward trackers run in seconds to minutes, the agent has a ten-hour budget) and are labelled as such. Toggle the proposal to see the teacher-student loop the conclusion sketches: the slider is how many videos you would label with the agent, and the hypothetical student point is exactly that, hypothetical.

Other limits the paper states or implies, worth carrying with you. The method needs an object mask per frame and known camera motion; both come from other systems (SAM 3, a SLAM or feed-forward reconstruction such as VGGT), and their failures are inherited. Transparent objects work precisely because a segmentation mask exists where texture does not, so the method is only as good as the mask on glass. Rotation error on fast rigid motion is still 37.6 degrees on average. And the kinematic prior can over-explain: the two iTACO sequences with an invented extra axis are the shape of that failure.

Cheat sheet.
PieceWhat it isWhy that choice
Object Oa Python script of Blender primitives plus joints, limits and axeseditable by intent as code diffs; hierarchical for free; renderable from any view
Generalised pose Ti6-DoF base (quaternion + translation) + one scalar per joint, one shared scalecompact, interpretable, constrained by the mechanism
Score S(O,T)IoU of the silhouette render and the mask, hands removed from bothprecise near the answer, cheap, hard to get wrong; local optima are the VLM's job
Camera-centric updateleft-multiply: ΔR · R“tilt it left as I see it”
Object-centric updateright-multiply: R · ΔR“flip it about its own axis”; symmetry flips
Pose toolagent names a bounded region; differential evolution returns top-K renders + scores; agent picksthe VLM chooses the basin, the optimiser refines; the pick need not be the numerical best
Temporal residualsper-component velocities and accelerations, camera motion removed, camera translation excludeddiagnostics, not a prior: real fast motion survives, flips get flagged
Sub-agentskeyframes split into chunks, one pose agent each; shape stays centralimage inspection is the bottleneck; pose is parallel once shape is fixed
Where this connects on the site. CoTracker and CoTracker3 are the bottom-up point trackers this paper argues against, and the strongest way to feel the contrast. D4RT is the OpenD4RT baseline in Table 1, dynamic 4D reconstruction done feed-forward. VGGT supplies the camera poses this method assumes, and appears as a baseline's companion on HOT3D. MASt3R-SLAM is the other way to get those poses. For the maths of frames, left versus right multiplication and geodesic rotation angles, Frames & Rigid-Body Math; for what a factor-graph smoothness prior would have done instead, Robust Estimation and Multi-View Geometry; for why a robot wants the explicit object model in the first place, Robotics Simulation and Inverse Kinematics.
The paper's own answer to its ten-hour-per-video cost is which of the following?