Film a room, then ask where every visible speck of it goes, frame after frame, in 3D (each "point" below is one of those specks). Old trackers copied the whole scene every frame and ran out of memory near 96 frames; TrackEverything keeps one shared scene instead, and follows videos over 1,000 frames long.
Learn how one shared 3D scene, merged cube by cube, lets a tracker follow every visible point of a long video without its memory growing with every frame.
Drag the number of frames and switch between the two ways of remembering a video, and watch which one fills up. Then we build, piece by piece, every part of the tracker that makes the second way work.
You need what a pixel and an (x, y) coordinate are and a rough idea of what a neural network does. We build the rest from zero.
The camera walks through the first room, loops around it, visits the second room and comes back.
The 96-frame and 400-frame numbers are the paper's (Figure 1, Sections 1 and 5); Chapter 4 works through exactly where 13,400 comes from. The two rooms and the walk here are illustrative.
Chapter 0
Learn what a track is, and the four choices that decide which points a tracker follows
You film two cows in a field with your phone. One cow is in the picture from the very first frame. The second walks in from the left edge a second later. The video is a few hundred frames: a few hundred still pictures shown one after another.
Now pick one tiny speck on the first cow's back, say a white patch the size of a fingernail, and ask a simple question: where is that speck in every frame of the video? In frame 0 it is at one spot on the screen. In frame 1 the cow has taken a step and the speck has moved a few pixels to the right. In frame 80 the cow's head has swung in front of it and you cannot see it at all, but it is still somewhere.
The answer to that question, written down frame by frame, is called a track: the list of positions of one physical point over time. Next to each position a track also stores whether the point can be seen in that frame. When something passes in front of it, the point is occluded (hidden), and its track keeps a best guess of where it is anyway. A program that produces tracks from a video is a point tracker.
Why would anyone want this? Because a track is the most direct record of how things move. It tells you the cow's leg swung forward while its back barely rose. The TrackEverything paper points at three kinds of models that already benefit from knowing where things are in 3D: models that answer questions about images, robot control programs, and video generators. Its authors argue that dense tracks over long videos are what would let those models move from still scenes to scenes where things move.
Every point tracker answers four questions, and the answers decide what it can and cannot do. The first is where the positions live. A 2D track gives each position as a pixel: a column and a row in that frame's picture. A 3D track gives each position as a point in space, three numbers X, Y and Z, measured in the room or field where the video was filmed. Chapter 2 shows why 3D turns out to be the easier place to track, even though it sounds harder.
The second is which points to follow. A sparse tracker follows only the points you hand it. Each one is a query point: a frame number and a pixel, meaning "follow whatever is at this pixel in this frame." A dense tracker follows every point, a full grid of them, so nobody has to choose.
The third is which frames points may start in. Many dense trackers follow every point that is visible in the first frame, and nothing else. We will call them first-frame dense. They cannot follow the second cow, because it was not in frame 0. They also miss something subtler: every patch of grass the first cow uncovers as it walks. Surfaces that come into view after being hidden are called disoccluded, and a first-frame tracker has no track for any of them. An all-frame dense tracker follows every point that is visible in any frame, from the frame where it first appears.
The fourth is how long a video it can handle. That sounds like a detail. It turns out to be the whole story of this paper. Try the first three choices here.
A toy video of a field. Dots are points on surfaces; a coloured dot has a track, a grey one does not. Pick a kind of tracker, then scrub through the video or press Play.
A toy, drawn in 2D so the choices are easy to see: the field, the cows and the point counts are illustrative. The situation is the paper's own example (Figure 4: only one cow is visible in the first frame and the second emerges later).
Scrub to frame 40 with the first-frame tracker. The second cow is covered in grey dots: nobody is following it. Behind the first cow there is a strip of grey ground it has uncovered. Switch to sparse and almost everything is grey: eight points is a sample, not a picture of the scene. Only the all-frame tracker colours everything, because it starts a new track the moment a new surface appears.
So why not always use an all-frame dense 3D tracker? Because until this paper, none could survive a long video. The earlier all-frame trackers ran out of memory after roughly 96 frames, about three seconds of phone video at 30 frames per second. Sparse trackers could go much longer, but only for the few points someone chose. The hero at the top of the page shows the reason in one picture: a tracker that keeps a copy of the scene for every frame needs memory that grows with every frame.
TrackEverything (Jain, Paruchuri, Gupta and colleagues at Carnegie Mellon University and Meta, 2026) is the first 3D tracker, as far as its authors know, that follows every visible point across videos longer than 1,000 frames, within 40 GB of graphics-card memory. It does it by changing what it stores. Instead of one set of points per frame, it keeps one set of points per piece of the world, in 3D, and merges points that turn out to be the same piece.
Here is the road, in order. Each chapter explains one part of that idea, and together they explain every part of the instrument in the hero.
Chapter 1
Count the positions a tracker must predict, watch the memory fill, and see why earlier trackers had to choose
Let's count the work. A track of one point through a video of T frames is T positions, one per frame. In 3D each position is three numbers. So a tracker's output is, at minimum, the number of points it follows times the number of frames it follows them through.
For a sparse tracker that is small. Hand it 384 query points (the number the paper uses in its speed tests) and a 200-frame video, and it owes you 384 × 200 = 76,800 positions. Easy. The trouble starts when you want every point.
A first-frame dense tracker follows every pixel of frame 0. For a video 512 pixels wide and 512 tall, that is 512 × 512 = 262,144 points, each placed in all 200 frames: 262,144 × 200 = 52,428,800 positions. Big, but it grows only in step with the video: twice the frames, twice the work.
An all-frame dense tracker must also follow every point that appears later, so any pixel of any frame can be the birth of a new track. If each frame is treated as its own grid of new points, the video contains 200 × 262,144 candidate points, and each must be placed in all 200 frames. That is where the paper's own number comes from.
262,144 pixels × 200 birth frames × 200 frames = 10,485,760,000 (about 1010 predictions, Section 2)
Notice the frames appear twice. For an all-frame tracker that treats every frame as new, the work grows with the square of the video length: double the video and the work goes up four times. Drag the length and watch the three rows pull apart.
Three trackers on the same 512 × 512 video. Each bar is the number of 3D positions the tracker must produce, on a scale where each step is 1,000 times the last. Choose a row to see its arithmetic, then drag the length.
512 × 512 pixels and 200 frames are the paper's example (Section 2: "on the order of 1010 predictions"); 384 query points is the setting of its Figure 3. The all-frame row counts every pixel of every frame as a separate point, the way frame-by-frame trackers represent a video, which is exactly the redundancy the paper removes.
Most of those ten billion positions are wasted. The same patch of wall appears in frame 1, frame 2, frame 3 and so on, and a frame-by-frame all-frame tracker gives it a brand-new point every time, then tracks every copy through the whole video. The paper puts it plainly: the same surface is represented independently in every frame.
Counting positions is only half the story. The other half is where the work happens. Neural networks run on a GPU (a graphics processing unit, the chip on a graphics card), which is fast because it works on huge batches of numbers at once. Everything the network is working on at a given moment must sit in the GPU's own memory. The paper's tests ran on one NVIDIA L40S card with 46 GB of it. When the working state does not fit, the program simply stops. People write this as OOM, for "out of memory."
For a tracker, that working state is mostly tokens: small vectors of numbers, one per patch of an image, that the network reads and rewrites. A frame-by-frame tracker adds a full grid of tokens for every new frame, whether that frame shows anything new or not. So its memory climbs with the video until it hits the card's limit.
Sparse trackers pay in a different currency. The standard recipe, from trackers like PIPs and CoTracker3, compares each query point's appearance with the dense feature maps of the video frames (feature maps: the same per-patch appearance vectors Chapter 3 will call features, laid out one per pixel instead of one per token). The similarity scores form a 4D correlation volume (in RAFT, the optical-flow method the paper traces this to, that means a similarity score for every pair of locations across two frames; a point tracker builds the analogous table for every query point). That is affordable for a few hundred queries. It grows with every query you add, which is why sparse trackers stay sparse. Chapter 6 shows how TrackEverything avoids building one at all.
Could a sparse tracker simply be handed more queries until it covers everything? The paper tried the first step of that. On 120-frame videos it raised the number of query points toward 10,000 and timed each tracker. SpatialTracker-v2 ran out of memory at 1,000 points. DeltaV2's sparse mode and CoTracker3 started out faster than TrackEverything, but became slower than it past about 750 and about 2,000 points respectively, and kept slowing as points were added. TrackEverything's time stayed nearly flat, because it tracks every point anyway and extra queries add no correlation work. A 512 × 512 frame has 262,144 pixels; 10,000 queries is under 4% of one frame.
The paper measured all of this directly. It ran six trackers on videos of growing length from PointOdyssey, a synthetic dataset with long videos, on the same 46 GB card, and recorded the peak memory each one used. Here are its measurements.
Each bar is one tracker's peak memory on a 46 GB card, at the video length you choose. Each name is labelled with which points it follows. Drag the length past a few hundred frames.
Peak memory read by eye off the plotted points of Figure 3c (single L40S, 384 query points for the trackers that take queries), so each value is approximate. Between measured lengths the bar is interpolated; before a tracker's first and after its last measured length it holds the nearest measured value, up to the out-of-memory length stated in Section 4.3: Any4D past 96 frames, SpatialTracker-v2 and DeltaV2's dense mode at about 200, DeltaV2's sparse mode at about 500, CoTracker3 at about 600.
Drag to 300 frames. Half the trackers are gone. Any4D, a first-frame dense tracker that processes all frames jointly, runs out of memory beyond 96 frames; SpatialTracker-v2 and DeltaV2's dense mode follow at about 200. The sparse trackers last longer, but at 600 frames even CoTracker3 has run out, and it was only following 384 points. TrackEverything's bar barely moves: about 8 GB at the start and about 15 GB at 900 frames on the plot (the text states "under 30 GB even at 900 frames"), while following every point.
Now the earlier choices make sense. You could track few points for a long time (sparse). You could track every point of the first frame for a while (first-frame dense). Or you could track every point of every frame, but only for short clips: the all-frame trackers VDPM and D4RT were evaluated on 48- or 64-frame clips, and the paper reports prior all-frame dense trackers exhausting memory beyond roughly 96 frames. One other method, Omnimotion, does track every point of every frame, but it fits a separate model to each video, which takes 8 or more hours per video.
The paper's claim is that this is a false choice, created by one habit: representing the video frame by frame. The next chapter replaces the habit.
Chapter 2
Depth, camera poses and pointmaps turn every pixel into a place, and a still wall stops moving
Walk past a bookshelf while filming it. On your screen the shelf slides to the left. Did the shelf move? Of course not. You moved. But a tracker that only sees pixels cannot tell the difference: in pixel coordinates, the shelf's motion and a rolling ball's motion are the same kind of thing, a change of column and row.
So pixel motion mixes two causes, the camera moving and the thing moving. The central observation of the paper is that a video is not really a stack of flat pictures. It is a record of one 3D world, seen through a camera that moves through it. If we can put every pixel back where it came from in the world, most of the scene will simply stand still.
A camera maps the 3D world onto a flat picture along straight rays of light that all pass through one point, the pinhole. Two numbers describe that mapping for a given camera, and together they are called its intrinsics: the focal length f, measured in pixels (bigger means more zoomed in), and the image centre (cx, cy), the pixel the camera's straight-ahead axis passes through.
A pixel on its own only tells you a direction: the ray it came along. To get a point you also need how far along that ray the surface is. That distance, measured along the camera's straight-ahead axis, is the depth Z. With it, you can walk back from pixel (u, v) to the 3D point that produced it. This is called unprojection:
With numbers, a shelf we'll follow for the rest of this chapter. In frame 20, once the camera has already slid 1 m to the right of where it started (more on that slide in a moment), a 640 × 480 camera with f = 500 and centre (320, 240) sees the shelf at pixel column 400, row 260, at depth 2.5 m. Then X = (400 − 320) × 2.5 / 500 = 0.40 m and Y = (260 − 240) × 2.5 / 500 = 0.10 m. In frame 20's own camera coordinates, the shelf point is at (0.40, 0.10, 2.50). These numbers are illustrative; the formula is the one every 3D tracker uses.
Those coordinates are relative to the camera, and the camera moves. So we need one more ingredient: the camera's pose, where it stands and which way it faces at each frame. A pose is a rotation R (which way the camera is turned) and a position t (where it is). Turning a camera-relative point into a world coordinate, a position measured in one fixed frame of reference for the whole video, is one line:
In this chapter's example the camera only slides, so R does nothing (turning by zero degrees leaves every coordinate exactly as it was), and R · pcamera is just pcamera. But R genuinely mixes coordinates when the camera turns. Suppose frame 20's camera had also turned 90 degrees to its right: multiplying the very same camera-relative point (0.40, 0.10, 2.50) by that turn's R would swap and flip two of its numbers, giving something like (2.50, 0.10, −0.40): the 2.50 m of depth becomes a 2.50 m sideways distance, because "straight ahead" now points where "off to the side" used to. That is what R is for: it re-expresses a point in whichever direction the camera happens to be facing, before t slides it to the camera's actual position.
Which frame of reference? TrackEverything uses the first camera. Its pose is "no rotation, at the origin," so frame 0's camera coordinates are the world coordinates, and every later frame is expressed in that same frame of reference.
Back to the sliding-only example. At frame 20, R does nothing and t = (1, 0, 0) is that 1 m slide, so the shelf point's world position is (0.40 + 1, 0.10, 2.50) = (1.40, 0.10, 2.50). In frame 30 the camera has slid 1.5 m. Where does the shelf point appear in the picture now? Its camera coordinates are the world position minus the camera's: (1.40 − 1.50, 0.10, 2.50) = (−0.10, 0.10, 2.50). Projecting forward, u = 500 × (−0.10) / 2.5 + 320 = 300.
pixel column 400 → 300, world position (1.40, 0.10, 2.50) → (1.40, 0.10, 2.50)
The shelf point jumped 100 pixels across the picture, and did not move at all in the world. See it happen, next to something that really does move.
Top: the picture, one row of pixels. Below: the room seen from above, with the camera sliding right. The shelf point never moves; the ball rolls. Drag the frame, then switch between where they appear in the picture and where they are in the world.
An illustrative room, drawn from above so height is left out: a 640-pixel-wide camera with f = 500 slides 0.05 m to the right per frame without turning, exactly the worked example above. The projections are computed live with the pinhole formula.
In the picture, both dots leave long trails, and nothing about the trails tells you that one of them is standing still. In the world, the shelf point is a single dot. Only the ball draws a line. That is the whole reason to track in 3D world coordinates: anything that does not move has a constant position, and tracking it is free once you know where it is.
Do the unprojection for every pixel of a frame and put the result back into the pixel's slot, and you get a pointmap: an image in which each pixel stores, instead of a colour, the 3D world position (X, Y, Z) of the surface it sees. A pointmap has exactly the shape of a colour image, height by width by 3. For a whole video, the paper's input is the colour video and one pointmap per frame, both of shape T × H × W × 3, the pointmaps already in world coordinates.
TrackEverything does not compute these pointmaps itself. They come from outside, in one of two ways. A device with depth sensors and a known camera path can measure them directly. Otherwise a feedforward reconstruction model, a neural network that looks at the frames once and outputs depth and camera poses for all of them, can estimate them. The paper's default is VGGT-Ω; it also tests Pi3, another feedforward geometry model. To see how a network can predict pointmaps straight from images, the DUSt3R and VGGT lessons build one from zero.
Two details matter later. The whole scene is rescaled by one number, fixed once for the entire video, so that it has a standard size (the paper follows DUSt3R here). That means the "units" of every 3D quantity in this lesson, including the size of the cubes in Chapter 4, are scene units, not metres. And because geometry is an input, better geometry makes a better tracker with no retraining. Chapter 10 shows how much: swapping Pi3 for VGGT-Ω raises the score by more than 6 points, and real sensor depth raises it further.
So here is the job, restated in 3D. Given the colour video and a world-coordinate pointmap for every frame, output, for every unique point in the scene, its 3D position in every frame (the paper writes these as an array of shape N × T × 3), whether it is visible in each frame (N × T), and one extra label per point: static or dynamic (N). The next chapter starts building the machine that produces them.
Chapter 3
Cut the video into 16-frame windows and turn each one into a cloud of features pinned to 3D places
Think of reading a long book with a small desk. You cannot spread all 1,000 pages out at once, so you read a few pages, write down what matters on an index card, clear the desk, and read the next few pages with the card beside you. The card grows only when the book tells you something new.
TrackEverything reads a video the same way. It cuts the video into windows of L = 16 consecutive frames that do not overlap: frames 0 to 15, then 16 to 31, and so on. It processes one window completely, carries a compact summary forward, and moves on. Processing a long input in consecutive slices like this is called sliding-window inference, and it is the standard way point trackers handle long videos. What is new here is what goes on the index card: not a stack of frames, but a set of points in 3D.
First, each of the 16 frames is turned into features. A feature is a list of numbers, here 384 of them, that describes what a small patch of the image looks like: its texture, its edges, what kind of thing it seems to belong to. Two patches that show the same surface should have similar features, even from different angles. Features are how a tracker recognises a point again.
The features come from two networks in a row. The first is DINOv3-small, a pretrained image network from the DINO family; the paper's reference for its predecessor is titled "learning robust visual features without supervision," which is the family's whole idea (the DINOv2 lesson explains how such networks are trained). It is frozen: its 20 million weights are used as they are and never change during training. That saves memory and compute, and it keeps the general-purpose knowledge the network already has. The second is a ViT Adapter head, a small network trained from scratch that turns DINOv3's output into a dense grid of features at a quarter of the image's resolution. For a window, the result has shape (L, H/4, W/4, 384): one 384-number feature for every 4 × 4 block of pixels of every frame.
Second, each feature is pinned to a place. Chapter 2's pointmaps give every pixel a 3D world position; shrunk to the same quarter resolution, they give every feature one too. Each frame's grid becomes a set of tokens, and each token is now a pair: what it looks like (384 numbers) and where it is (X, Y, Z in world coordinates). Scatter all of them into 3D and you have what the paper calls a feature cloud: a cloud of points in space, each carrying a description of its own appearance.
Some arithmetic, with an illustrative input size since the paper does not state one. A 512 × 512 frame gives 128 × 128 = 16,384 tokens. A window of 16 frames gives 16 × 16,384 = 262,144 tokens. That is already a lot to hand to a transformer (an attention-based network; a later lesson and Chapter 5 build one from zero), so the next step trims it.
Pixels are not spread evenly over the world. A nearby sofa covers hundreds of pixels; a far wall of the same size covers a few dozen. So in 3D, the tokens of one frame are crowded on near surfaces, many of them only millimetres apart and describing nearly the same patch.
The fix is a voxel, a small cube of 3D space, the 3D version of a pixel. Chop space into a grid of cubes of side v. A point's cube is found by dividing each coordinate by v and rounding down, which is called the floor, written ⌊ ⌋:
Tokens of the same frame that share a cube are merged by mean-pooling: their features are averaged into one feature, and their positions into one position. One cube, one token. The device shows it on a room seen from above.
A camera looks into the corner of a room with a sofa in front. Each faint ray is one column of tokens, landing where the pointmap says its surface is. Grow the cubes and watch tokens merge; then switch to the whole 16-frame window.
A toy seen from above: 48 token columns per frame, a camera that drifts slightly across the window, cube sizes in toy units. The rule is the paper's (Section 3.1): quantise positions with ⌊x / v⌋, merge tokens sharing a cube by averaging, within each frame only, then stack the frames of the window.
In one frame, the tokens that pile up on the near sofa collapse into a few cubes, while the sparse tokens on the far wall mostly stay single. Switch to the whole window and something else appears: the same cubes are occupied again and again, once per frame, because the camera keeps seeing the same sofa and the same wall. Those repeats are not merged yet.
That is deliberate. Merging across frames at this stage would be a mistake, because two different things can occupy the same place at different times. Picture a dog running through the spot where a ball was resting a moment earlier: in world coordinates, a dog token from frame 9 and a ball token from frame 2 could land in the same cube, and averaging them would create a thing that is neither. The paper voxelizes per frame during encoding for exactly this reason. Merging across frames waits until every point has been moved to the same moment in time, which is the next chapter.
The per-frame trimmed tokens of all 16 frames are stacked into one list, the window's space-time feature cloud, of shape (N, 384). Each point in it keeps one more label: its source timestep tsrc, the frame of the window in which it was first observed, a number from 0 to 15.
And the index card? Every window after the first also receives the points carried over from all earlier windows: the compact, merged scene memory that the next chapter builds. They join the same list, with the label tsrc = −1, meaning "known from before this window." So the tracker's working set for window w is simply the old scene plus whatever the 16 new frames show.
| Stage | What goes in | What comes out |
|---|---|---|
| Backbone | 16 colour frames | frozen DINOv3-small features |
| Adapter | backbone features | (16, H/4, W/4, 384) feature maps, trained |
| Pin | feature maps + pointmaps | tokens with a 3D world position each |
| Trim | each frame's tokens | one token per occupied cube, per frame |
| Stack | 16 trimmed frames + carried points | the window's cloud, (N, 384), each with tsrc |
One more property of this design matters for memory. The big 2D feature maps are needed only while their window is being processed. Once a window is finished, the paper releases them from GPU memory and keeps only the compact 3D points (and a small table for rebuilding full tracks, Chapter 4). The frames do not accumulate. Only the scene does.
Chapter 4
At each window's end, points that arrive at the same small cube become one, so memory follows space instead of time
Fast-forward to the end of a window. The tracker (the next three chapters build it) has done its work: for every point in the window's cloud, it has predicted where that point is at the window's last frame, frame 15. A patch of sofa first seen in frame 2, the same patch seen again in frame 9, and the same patch carried over from an earlier window now all have a predicted position at one and the same moment.
And at one moment, a rule from everyday life applies: two things cannot occupy the same place at the same time. If three points arrive at the same tiny region of space at frame 15, they are not three things. They are one piece of the world, recorded three times. So TrackEverything merges them. This is the step the paper's title calls de-duplicating the 3D scene representation, and it is why the hero's shared-scene line stays flat.
Recall why this was forbidden during encoding (Chapter 3): then, points came from different moments, and a dog in frame 9 could sit where a ball sat in frame 2. Now every point has been carried to the same frame. The dog and the ball cannot both be in that cube at frame 15. The merge is safe precisely because of the timing.
The merge uses the same small cubes as before, now across every point of the window. Each predicted position is turned into a cube index by dividing by the cube size and rounding down:
Points that share an index are merged by mean-pooling, as in Chapter 3: their 384-number features are averaged into one feature, and (in this lesson's reading of "mean-pooling co-located tokens") their positions are averaged into one position. The merged point is one canonical track, the single representative of that piece of surface from now on.
Let's do one by hand with the paper's cube size, v = 0.02, and four illustrative points that end the window close together:
| Point | Where it came from | Position at frame 15 | ⌊ x / 0.02 ⌋ |
|---|---|---|---|
| A | first seen in frame 2 | (0.113, 0.047, 0.508) | (5, 2, 25) |
| B | first seen in frame 9 | (0.118, 0.041, 0.513) | (5, 2, 25) |
| C | first seen in frame 14 | (0.121, 0.049, 0.502) | (6, 2, 25) |
| D | carried from an earlier window | (0.109, 0.044, 0.517) | (5, 2, 25) |
Take A's first coordinate: 0.113 / 0.02 = 5.65, rounded down to 5. Its others: 0.047 / 0.02 = 2.35 → 2, and 0.508 / 0.02 = 25.4 → 25. B and D land in the same cube, (5, 2, 25), so A, B and D merge. Their mean position is ((0.113 + 0.118 + 0.109) / 3, (0.047 + 0.041 + 0.044) / 3, (0.508 + 0.513 + 0.517) / 3) = (0.1133, 0.0440, 0.5127).
4 points → 2 tracks (A + B + D merged, C alone)
Look at C. It is only 0.003 away from B in the first coordinate, yet 0.121 / 0.02 = 6.05 rounds down to 6, so it falls into the next cube and stays separate. Cube borders are hard lines: two points very close together can land on opposite sides of one. That costs a little memory (C survives as a near-duplicate) but never accuracy, and the next window gets another chance to merge them.
Now the whole window at once, on a toy scene. The coloured rings are points carried from earlier windows; the grey dots are new this window; the pink dots are on a ball that moved.
Every dot is a point's predicted position at the window's last frame, seen from above. Press Merge to average every cube's points into one; change the cube size and merge again. Turn on the receipts to see what is stored for undoing the merge.
A toy scene in toy units, drawn from above (the real merge is in 3D, with cubes). The merge rule, the averaging and the receipts (a point-to-cube table and one offset per point) are the paper's (Section 3.3 and Appendix A.4). In the paper's own example window (Figure 5), 46k points arrive at the boundary and 16k remain after merging.
Two things to notice. First, the ball's points merge with each other but never with the floor or the wall, because at frame 15 they are somewhere else. Second, the carried points (the rings) and the new points of the same surfaces fall into the same cubes. Seeing a surface again does not add memory. It adds evidence, averaged into the point that was already there.
In the paper's own example (its Figure 5), a window ends with 46,000 points and merging leaves 16,000: about 35% kept, some 30,000 duplicates removed. Those 16,000 become the scene memory that the next window starts from.
When the next 16 frames arrive, they go through Chapter 3's encoder: features, pinned to 3D positions from the geometry model (or sensor depth), trimmed per frame. The new points are stacked together with the carried scene memory, with tsrc = −1 marking the carried ones, and that combined set is what the tracker follows through the new window. At its end, everything is merged again.
So the number of points the tracker carries grows only when points land in cubes nobody occupied before, which is to say, when the camera sees new space. Walk back into a room you have already filmed, and its points merge into the existing ones. That is the flat stretch of the hero's warm line, and the step when the camera enters room 2.
How is the grid stored when a scene can be any size? Not as a giant 3D array. The cube indices are fed through a hash, a function that turns a key like (5, 2, 25) into a slot in a lookup table, so only occupied cubes ever get an entry. The paper stresses the consequence: no bounding box, no clipping, no fixed room size. A camera that walks down a street simply keeps adding cubes.
Merging seems to throw information away. If A, B and D are now one track, where did A's own track go? The paper keeps two small receipts at every merge. The first is a point-to-cube table p: for each point before the merge, the number of the merged track it went into. The second is a residual r for each point: its offset from the merged track's position, r = x − xmerged.
To rebuild a full trajectory for an original point, follow the tables backwards and add the offset: Xorig(t) = Xmerged[p](t) + r. The offset is the same at every frame, so each original surface element gets its own trajectory, the merged track shifted by its own small offset. With our numbers, A's offset is (0.113 − 0.1133, 0.047 − 0.0440, 0.508 − 0.5127) = (−0.0003, 0.0030, −0.0047). If the merged track later sits at (0.300, 0.050, 0.600), A's rebuilt position there is (0.2997, 0.0530, 0.5953).
One more receipt is kept inside the merged point itself: its birth record. When points merge, the merged track keeps the source frame tsrc and starting position xsrc of the earliest of them (a carried point beats any new one). A track's birth is when the surface was first seen, not when it was last re-seen. And because of these receipts, the large 2D feature maps of a finished window can be freed at once: the scene memory plus the tables are all that is needed to export every original pixel's track.
Here is the merge as code, from scratch in NumPy. It is short because the idea is short.
pythonimport numpy as np def dedup(xyz, feat, t_src, x_src, v=0.02): """Merge points that land in the same cube at the window's last frame. xyz: (N,3) predicted positions at t_tgt; feat: (N,D); t_src: (N,), -1 = carried.""" k = np.floor(xyz / v).astype(np.int64) # cube index per point keys, p = np.unique(k, axis=0, return_inverse=True) p = p.ravel() # point -> merged track (receipt 1) M = len(keys) cnt = np.bincount(p, minlength=M)[:, None] xyz_m = np.zeros((M, 3)); np.add.at(xyz_m, p, xyz); xyz_m /= cnt # mean position feat_m = np.zeros((M, feat.shape[1])); np.add.at(feat_m, p, feat); feat_m /= cnt r = xyz - xyz_m[p] # per-point offset (receipt 2) order = np.lexsort((t_src, p)) # group by track, earliest birth first first = order[np.r_[True, p[order][1:] != p[order][:-1]]] return xyz_m, feat_m, t_src[first], x_src[first], p, r # birth record of the earliest
Chapter 5
Predict where each point ends up and whether it moves, then spend the expensive work only on the movers
A window's cloud can hold tens of thousands of points: 46,000 in the paper's example window. The obvious way to track them is to predict a full path for each one, 16 positions, one per frame of the window. That is 46,000 × 16 = 736,000 positions per window, and each of them would have to be refined by a network.
Two observations make most of that work unnecessary. The first comes from Chapter 2: in world coordinates, a point that does not move has the same position in every frame. Walls, floors, tables and parked cars are most of most scenes. Their "path" is one position repeated 16 times. The second comes from Chapter 4: to hand points to the next window and merge them, the tracker needs only one thing, each point's position at the window's last frame. Not the path. The endpoint.
So the paper splits tracking in two, a decomposition it calls "endpoints before trajectories." An endpoint refiner runs on every point and answers three questions: where is it at the target frame, can it be seen there, and does it move at all? Then a trajectory refiner (Chapter 7) decodes full 16-frame paths, but only for the points labelled as moving. The static ones simply keep their position for the whole window.
Each point enters with the 384-number feature fi it got in Chapter 3, which says what it looks like. The refiner also needs to know where and when the point comes from, and where and when it is being asked about. Four facts, each a number or a position:
A network reads raw coordinates badly: 0.513 and 0.517 look almost identical to it, yet the difference may matter. The standard fix, borrowed here from the neural rendering method NeRF, is a sinusoidal embedding: turn each number into a long list of sines and cosines of that number at many different frequencies. Slow waves tell the network roughly where the point is; fast waves tell it precisely.
Here is why that helps, with those same two numbers. As raw inputs, 0.513 and 0.517 are only 0.004 apart, so close a network can barely tell them apart. Run both through a slow wave, one that completes a full cycle only every 10 units, and they still land almost on top of each other. Now run both through a fast wave, one that completes dozens of cycles over that same range, and the two numbers can land on very different parts of the wave, one near a peak and the other near a trough, even though they started 0.004 apart. Stack many such waves, slow ones and fast ones together, and somewhere in that long list is a number that changes sharply between 0.513 and 0.517, however identical the raw coordinates looked.
Each of the four facts gets its own embedding, each is squeezed to 384 numbers by its own small learned layer E, and all four are simply added onto the point's feature:
Two more ingredients are glued on before the network runs, a sample of the image where the point was born and a sample where the guess currently lands. They are the subject of the next chapter, so for now just note that the actual input for each point is three pieces side by side: [fi; gsrc,i; gtgt,i].
The inputs go through a stack of self-attention layers over the whole feature cloud. Self-attention is the core operation of a transformer: every point looks at every other point, scores how relevant each one is to itself, and pulls in a weighted mix of their information. Here that means a point on a swinging arm can learn from all the other points on the arm, and a point on the floor can learn that its neighbours are all standing still. These layers (the paper does not say whether this means just the endpoint refiner's, just the trajectory refiner's in Chapter 7, or both) start from the last layers of DINOv2-small, a separate and older pretrained image network from the same DINO family as Chapter 3's DINOv3-small (it only donates its starting weights here; it is not the frame encoder itself), rather than from random numbers.
A small MLP (a few plain fully connected layers) then reads each point's updated feature and outputs three things:
Where does the first guess come from? The standard choice in point tracking, which the paper follows, is a zero-velocity start: assume nothing moved. A new point's first endpoint guess is simply where it was born, x(0)tgt = xsrc. A carried point starts from where the previous window left it. For every static point, that first guess is already right in world coordinates, which is the whole trick.
Now look at how little the expensive part has to do. The device counts the work in one window of a toy scene with a person walking through a still room.
Each dot is one point of the window's cloud: still ones and moving ones. A small square marks a point whose 16-frame path gets decoded. Choose a strategy, then change how much of the scene moves.
The 240-point scene and the moving share are illustrative; the work counts are exact for it (one endpoint per point, 16 path positions per decoded point). The measured saving on real videos is the paper's Table 2b: with the split, 0.05 s per frame and 9.1 GB; decoding every point's path, 0.24 s and 22.4 GB; 31.2 APD-P either way.
At 10% moving, decoding only the movers cuts path work tenfold. The paper measured the real version: forcing every point through the trajectory decoder left accuracy exactly the same (31.2 on its main score, Chapter 9) but made tracking 4.8 times slower (0.24 against 0.05 seconds per frame) and used 2.5 times the peak memory (22.4 against 9.1 GB). The static points' paths were never worth computing.
It is good. On the synthetic datasets where the truth is known, it scores 93.1 on Dynamic Replica and 92.5 on PointOdyssey on the F1 score, a 0-to-100 measure that is high only when a classifier both catches most of the moving points and rarely calls a still point moving. On the paper's own full 1,000-plus-frame PointOdyssey test, the classifier's two halves split apart: it catches stillness very reliably (92.5 F1) but is noticeably weaker at catching motion (78.5 F1), so a real mistake is more likely to be a mover called still than the reverse.
And its mistakes are cheap by design. If a still point is wrongly called moving, it goes through the trajectory refiner, which was trained on all kinds of motion including none, and simply predicts a path that stays put: a little wasted compute, no harm to accuracy. If a moving point is wrongly called still, its path is held fixed for this one window. It is judged again in every later window, and once its motion is clear, it resumes being tracked as a mover, though the paper does not promise this happens in the very next one.
One detail about the target frame. At inference, and for every window but the last during training, the endpoint refiner is always asked about frame 15, because boundary positions are what the sliding window needs. On the last window of each training video, the target frame is instead drawn at random from 0 to 15. That teaches the model to predict where a point is at any frame, not just the boundary, and it is why the target frame needs its own embedding at all.
Chapter 6
Project the guess into the image, sample what is there, and let the mismatch drive the next guess
A friend shows you a photo of a white patch on a cow's back and says: "In the next photo, I think it's right here," pointing at a spot. How do you check? You look at that spot in the next photo. If it shows a white patch of hide, the guess is probably right. If it shows grass, the guess is wrong, and the cow went somewhere else.
That is the whole idea of this chapter, and it is how the endpoint refiner gets its evidence. For every point, take the current guess of where it is at the target frame, look at the target frame's image there, and hand the network both what the point looked like when it was born and what the image shows at the guess. Similar means "keep it." Different means "move it."
Earlier trackers answered "where did it go?" by comparing features at scale. They compared the point's feature with the dense feature maps of the frames and stored the similarity scores, the 4D correlation volume of Chapter 1, indexed by point, frame, row and column. Built densely, one score for every location, at an illustrative 128 × 128 grid of features, one window of 46,000 points would need 46,000 × 16 × 16,384 ≈ 12 billion scores. That is why correlation-based trackers stay sparse.
A recent optical flow method showed that the brute force is unnecessary. Optical flow is the per-pixel motion between two video frames, the 2D ancestor of the 3D tracking this lesson teaches. The method is called WAFT, short for warp-aligned feature transform (Wang and Deng, 2026). Instead of a correlation volume, it warps: it fetches the feature at the currently estimated location and lets attention layers work out the correspondence. The TrackEverything paper sums up WAFT's finding: self-attention over features built this way is enough to find correspondences. TrackEverything carries the idea into 3D and calls it 3D WAFT. (For the 2D ancestors, see optical flow and RAFT, the method that popularised correlation volumes.)
Chapter 2's formula ran from a pixel to a 3D point. Now we run it forwards. The point's current guess xtgt is a 3D world position. Move it into the target frame's camera coordinates by undoing that camera's pose, p = RT(xtgt − t), then project it onto the picture with the pinhole rule:
Then read the target frame's feature map at (u, v). The features live on a grid (one per 4 × 4 pixels), and the guess almost never lands exactly on a grid cell, so the reading is a bilinear sample: a blend of the four nearest cells, each weighted by how close it is. Worked example: the guess lands at (10.3, 4.6) in feature-grid units. Its four neighbours get these weights:
cell (10, 4): 0.7 × 0.4 = 0.28
cell (11, 4): 0.3 × 0.4 = 0.12
cell (10, 5): 0.7 × 0.6 = 0.42
cell (11, 5): 0.3 × 0.6 = 0.18 (total 1.00)
The weights add up to 1, and the sampled feature gtgt is 0.28 of the first cell's 384 numbers plus 0.12 of the second's, and so on. The closer cell counts more. Because the blend changes smoothly as the guess moves, a network can learn fine, sub-cell corrections from it.
One more sample is taken, once per point: gsrc, the feature at the point's birth location in its birth frame. It is the reference photo of the white patch, kept so that, in the paper's words, the model cannot "forget" the original appearance. The refiner's input for each point is the three side by side, [fi; gsrc,i; gtgt,i]. That pairing is an implicit template match: when the two samples agree, the guess is probably right; when they disagree, it needs to move.
Left: the point where it was born, in frame 0. Right: frame 15, with the current guess and the true spot. The bars compare what each sample sees. Pick a point, then press Refine once a few times.
A toy image with its own colours; the camera's own motion is taken as already undone by the projection, as in the paper's Figure 6. The samples are real bilinear samples of the toy image, and the match is computed from them. The update is a stand-in for the learned refiner: it moves the guess further when the samples disagree more. The story (a still point already right at iteration 0, a moving point locking on by iteration 2) follows Figure 6.
The still point is already right before any refinement. Its zero-velocity guess, projected with the target camera's pose, lands on the pebble, the two samples agree, and nothing needs to change. That is Chapter 2's payoff arriving in the image: projection handles the camera's motion, so a still point needs no work.
The moving point starts on bare ground, because the block moved while the guess assumed nothing did. The samples disagree, so the refiner moves the guess; after one round it sits much closer, and after two it has locked onto the stripe. This is exactly the behaviour the paper shows on real images in its Figure 6: at iteration 0 the projected guess for a moving object lands on the background, by iteration 1 it is substantially closer, and by iteration 2 the sampled point locks onto the correct surface.
That is why refinement is iterative. Within each window, the loop runs K times. Each round, the endpoint refiner reads fresh samples at the newly updated guesses, the position embeddings of Chapter 5 are refreshed with the newest estimates, the refiner outputs a correction Δx, and then the trajectory refiner (next chapter) updates the moving points' paths. Every round gives the network better evidence than the last.
The paper points out a neat consequence. Chaining windows uses exactly the same mechanism: the first guess in a new window comes from the previous window's answer instead of from the previous round. One recurrent machine handles both refining a guess and carrying a track through a 1,000-frame video.
Two practical details. First, when the camera swings fast, some guesses project outside the picture. Their coordinates are clamped into an enlarged frame, from −2W to 2W across and −2H to 2H down, and the sample repeats the edge values of the image. Second, when a point is hidden behind something, or out of view, its target sample shows something else, and that mismatch between gsrc and gtgt is exactly the signal the visibility head uses to say "not visible here."
What is it worth? Removing 3D WAFT lowers the paper's main score from 31.2 to 29.3 (Table 2c), and in its qualitative comparison, tracks without it drift and leave long stray trails on fast motion. Removing iterative trajectory refinement costs more, 31.2 down to 26.4. And because WAFT never builds a correlation volume, adding more query points barely changes TrackEverything's speed: in the paper's Figure 3, its latency stays nearly flat from a few hundred to 10,000 extra query points, while correlation-based trackers slow down.
Chapter 7
Start every moving point on a straight line, let attention bend it to the truth frame by frame, then run the whole tracker
A ball is tossed across the room during one window. The endpoint refiner has done its job: it knows where the ball's points start, where they are at frame 15, and that they move. But the ball did not travel in a straight line. It rose and fell. For every frame in between, the tracker still owes a position, and for moving points that is the job of the trajectory refiner.
Before any network runs, each moving point gets a first path by constant-velocity interpolation: a straight line from its source position to its endpoint, travelled at a steady speed. With illustrative numbers: a point starts at (0.20, 0.40, 1.00) at frame 0 and the endpoint refiner puts it at (0.50, 0.10, 1.30) at frame 15. It moves (0.30, −0.30, 0.30) in 15 steps, which is (0.02, −0.02, 0.02) per frame, so its first guess for frame 5 is
(0.20, 0.40, 1.00) + 5 × (0.02, −0.02, 0.02) = (0.30, 0.30, 1.10)
For a tossed ball that is wrong in the middle and right at both ends, which is a good place to start. On later refinement rounds, the path is not rebuilt from scratch: the previous path is kept, and only its last position is replaced by the endpoint refiner's newest estimate.
Each moving point's path becomes a short sequence of tokens, one per frame of the window. The token for frame t glues together three things:
So each token says, in effect, "this is what I looked like when I was born, and this is what the image shows where I am supposed to be at frame t." To the 16 frame tokens, the paper adds one extra token, [cls], a summary of the whole track, started from the point's cloud feature. The sequence has 17 tokens, and all the moving points together form an array of shape (Ndyn, 17, 384), where Ndyn is the number of moving points.
The decoder then alternates two kinds of attention, layer after layer:
Here is a small, illustrative case of that borrowing at work. Suppose a ball's true path puts it at position 0.40 at frame 8 and 0.46 at frame 10, a steady rightward drift, but at frame 9 a hand passes in front of it, so 3D WAFT's image sample there is unreliable and the raw guess for that one frame lands off the path, at 0.30. Temporal self-attention lets frame 9's token look at frames 8 and 10, see that they agree on a smooth, steady drift, and pull its own estimate back toward their average, to something like 0.43, close to where an unoccluded frame 9 would actually sit. No single frame decides this alone; the correction comes from the pattern across the whole 17-token sequence.
Finally a small MLP reads each of the 16 frame tokens (the summary is discarded) and outputs a 3D correction to that frame's position and a visibility logit for that frame. That is Ndyn × 16 outputs per round. Repeat for K rounds, and the straight line bends toward the real path.
The true path of one moving point across a 16-frame window, seen from the side, and the tracker's guess, which starts as a straight line. Pick a motion and press Refine once. The row below is the track as the decoder sees it: 16 frame tokens and the summary.
Toy paths in toy units. The straight-line start is the paper's constant-velocity initialisation; the refinement step is a stand-in for the learned decoder (each frame's position moves 60% of the way to the truth, then is smoothed with its neighbours, the way temporal attention lets neighbouring frames inform each other). The error printed is the average distance between guess and truth over the 16 frames.
What does the finished output look like? The paper's figures show it on real videos: a breakdancer, a dancer jumping, hands manipulating objects, animals in a field. Every visible point gets a track, including regions that only come into view after the first frame, such as a second cow that walks in later. For readability the figures draw trails only for the moving points and show the still content as a 3D point cloud, but the method tracks every point and accounts for the camera's motion. The baselines in the same figures either follow only the queried or first-frame points, or leave large areas untracked.
The ends are right from the start, because the endpoint refiner and the source position pin them. All the work happens in the middle, and each round removes most of what is left. That is why the paper can afford to decode paths only for moving points: the decoder's job is small per track, and the tracks are few.
How much does refinement matter? In the paper's ablation, removing iterative trajectory refinement lowers the main score from 31.2 to 26.4, the larger of the two drops in its component ablation (Table 2c; removing 3D WAFT gave 29.3), and its qualitative figure shows the tracks becoming noisier and less coherent without it.
You now have every part. Here is how they fit together in one window, in order: encode the 16 new frames into tokens pinned to 3D (Chapter 3); stack them with the carried scene memory; for K rounds, sample the images where each guess lands (Chapter 6), update every endpoint and static-or-dynamic label (Chapter 5), and update the paths of the movers (this chapter); then merge everything that lands in the same cube at frame 15 (Chapter 4), free the window's image features, and move on. Watch it run on a toy room with a camera and a dog.
Seen from above, a camera walks around a room while a dog runs around a table. Every 16 frames the window closes and its points merge. Press Play or step one window at a time; switch merging or the static-and-dynamic split off to see what each one saves.
A toy: one point per small patch of surface, one token per patch per frame that sees it (the per-frame trim of Chapter 3), 16-frame windows, and a merge that keeps one point per patch (static) or per patch of the dog. The counts are exact for this toy and are not the paper's measurements; its measured effects are in Chapters 4, 5 and 10.
Let it run for a few windows. With merging on, the number of carried points climbs while the camera discovers the room and then stops growing: every window still brings a few hundred new tokens, but the merge folds almost all of them back into points that already exist. Switch merging off and the line climbs by a full window's worth of tokens every 16 frames, forever, which is the frame-by-frame tracker of Chapter 1 in miniature. Switch the split off and the path counter jumps from the dog's points to every point in the room.
Chapter 8
Meet the three synthetic datasets, the four losses, the rule for what counts as moving, and a trick that trains on long videos without storing them
To learn, a tracker needs videos where the right answer is known: for some points, their true 3D position in every frame. Nobody can label that by hand for a real video, least of all in 3D. So, like most trackers, TrackEverything learns from synthetic videos, rendered by a computer that knows exactly where every object is.
It uses three such datasets. Kubric is a generator of synthetic scenes. PointOdyssey is a large synthetic dataset built for long-term point tracking, with videos of 1,000 frames and more. Dynamic Replica comes from work on consistent depth for dynamic stereo videos. Each provides 3D tracks, but only for a sparse set of points, not for every pixel. So the losses are computed only on the labelled points, and everything else in the scene is tracked without a direct grade.
A loss is a number that measures how wrong the model's outputs are on a training example; training adjusts the weights to make it smaller. TrackEverything's loss, computed window by window, adds up four terms (the paper's Equation 4):
The confidence term and the λ weights are two honest gaps: the paper names them but never spells out what they do or what values it used, so treat them as loose ends rather than something to derive.
To save compute, the full 16-frame paths of the trajectory refiner are decoded during training for only a subset of the moving points. Chapter 7's independence between tracks is what makes that safe: a track decoded alone is decoded exactly as it would be in company.
The static-or-dynamic loss needs a true label for every labelled track, and the paper derives it with a simple geometric rule. Take the track's true positions across the frames of the window where it is labelled. Draw the smallest box, aligned with the axes, that contains all of them. If the box's diagonal is longer than a threshold, the track is dynamic:
Worked example with illustrative numbers. A point on a wall, jiggled by labelling noise, spans X from 0.31 to 0.35, Y from 0.20 to 0.21 and Z from 0.90 to 0.92. Its box is 0.04 by 0.01 by 0.02, and the diagonal is √(0.04² + 0.01² + 0.02²) = √0.0021 = 0.046. That is under 0.05: static. A tossed ball spans a box of 0.12 by 0.05 by 0.03, with diagonal √(0.0144 + 0.0025 + 0.0009) = √0.0178 = 0.133: dynamic. Try your own.
One track's 16 positions in a window, seen from above, with its bounding box. The bar at the right is the box's third side, its height. Pick a kind of track and stretch how far it travels; the rule decides.
The tracks are illustrative, in normalised scene units, with a little labelling noise added to every one. The rule and the 0.05 threshold are the paper's (Appendix A.4, Equation 2).
Notice the swaying branch. Stretch it and it flips to dynamic, even though it ends the window close to where it started. The rule measures the whole extent of the motion within the window, not the distance from start to finish, so a point that goes somewhere and comes back is still a mover. That matters, because its path in between is exactly what the trajectory refiner exists to decode.
First, how any network learns, in one paragraph. For every weight, training works out how much a tiny change to that weight would lower the loss. That number is the weight's gradient, and computing all of them, by tracing the loss backwards through every step that produced it, is the backward pass. Then every weight is nudged a little in its helpful direction; one nudge of all the weights is one training step, and the size of the nudge is the learning rate.
The tracker is recurrent: each window starts from the previous window's output. The textbook way to train such a system is backpropagation through time: keep every intermediate result of every window in memory so that the error at window 4 can be traced back through windows 3, 2 and 1. That is exactly the kind of memory that grows with video length, the thing this whole paper avoids.
So TrackEverything cuts the thread. At every window boundary, the carried state is detached: treated as a fixed input for the next window rather than as something whose own history can be corrected. Each window computes its loss, runs its own backward pass, adds its gradients to the model's, and releases its intermediate results before the next window starts.
That creates a problem for the one trainable part that serves every window, the ViT Adapter that makes the 2D feature maps. The goal is simple to state even before the mechanism: nudge the adapter as if it had seen every window's feedback at once, without actually keeping every window's computation graph alive to prove it. The paper's fix is a neat trick that does exactly that. First, compute the feature maps F for the whole training video once. Then make a detached copy Fparam that collects gradients, and let every window read from the copy instead of from F itself. Each window's backward pass deposits its own gradient into Fparam.grad, then throws away everything else about that window. When all windows are done, one last backward pass sends the accumulated gradient into the adapter through the scalar 〈Fparam.grad, F〉, the sum of the element-wise product of the two; by the chain rule, its gradient with respect to the adapter's weights works out to exactly "accumulated gradient times how the features depend on those weights," the same answer the adapter would have gotten had every window's graph been kept around. Then the optimiser takes one step.
python# One training step on a video of several 16-frame windows (PyTorch-style sketch). F = adapter(dino_frozen(frames)) # (T, H/4, W/4, 384); graph kept back to the adapter F_param = F.detach().requires_grad_(True) # a leaf that collects every window's gradient state = initial_state() for w in range(num_windows): out = tracker(F_param[w * 16:(w + 1) * 16], pointmaps[w], state) loss = coord_l2(out, gt[w]) + lam_vis * ce(out.vis, gt[w]) + lam_dyn * ce(out.dyn, gt[w]) + lam_conf * conf(out) loss.backward() # grads into the tracker and into F_param.grad state = out.carry.detach() # cut the thread: no backprop through time (F_param.grad * F).sum().backward() # one pass: accumulated gradient into the adapter optimizer.step(); optimizer.zero_grad()
| Item | Value |
|---|---|
| Parameters | 61M in total: 41M trained, plus the frozen 20M DINOv3-small |
| Starting weights | all trained from scratch, except the transformer's attention layers (the paper does not say which refiner, or both), taken from the last layers of DINOv2-small |
| Learning rate | 1 × 10−4 |
| Stage 1 | 100k steps on 32-frame videos (2 windows) |
| Stage 2 | 300k steps on a mix of 32- and 64-frame videos (2 or 4 windows) |
| Scale | point clouds scaled to unit norm (an overall size of 1), following DUSt3R |
| Augmentation | random edits to the training videos so the model learns to shrug them off: colour jitter, image and depth blur, random erasing, following TAPIP-3D |
One last detail is a lesson in itself. During training the authors saw sudden spikes in the loss, and traced them to errors in PointOdyssey's labels. In 16 sequences, about 25,000 frames, the camera pose stored for frame t was the one that rendered frame t − 1: an off-by-one. The mismatch is a few pixels when the camera moves slowly and over 600 pixels in fast pans, up to 59 degrees per frame. They re-indexed the cameras, dropped supervision on the 502 frames where the repair moved tracks by 8 pixels or more (2% of the affected sequences, at most 3.8% of any one), and excluded one sequence whose jitter could not be fixed by any whole-frame shift. The loss spikes shrank, and they plan to release the corrections. Even the answer key needed checking.
Chapter 9
Score a tracker by hand, then read how TrackEverything compares on short clips, on long videos and in speed
Your tracker says a point is at a certain spot in frame 30. The truth is 8 centimetres away. Good or bad? It depends on how strict you are. If "close enough" means within 10 cm, it passed. If it means within 1 cm, it failed. A fair score should not hang on one arbitrary choice of strictness.
The main score of this field, APD (from the TAPVid-3D benchmark), handles that by asking the question at several strictness levels and averaging. For each distance threshold, it counts the fraction of predicted positions that fall within that distance of the truth. Then it averages those fractions over all the thresholds. Scores are written as percentages: 100 means every prediction was within even the tightest threshold.
Papers use two sets of thresholds, and the paper reports both. APD-M (for metric) uses fixed distances: 0.1, 0.3, 0.5 and 1.0 metres. APD-P (for pixel) starts from pixel distances, 1, 2, 4, 8 and 16 pixels, and turns each into a 3D distance using the camera's intrinsics. By Chapter 2's formula, a sideways step of δ pixels at depth Z covers δ · Z / f in 3D. So a near point must be placed precisely, and a far one gets more slack, the way an error would look on screen. APD-P is the benchmark's own metric, used by the sparse trackers, and it is typically the stricter one.
Worked example with six illustrative predictions and a camera with f = 500. The last column is how far off each prediction looks on screen, e · f / Z pixels:
| # | Error e | Depth Z | On screen |
|---|---|---|---|
| 1 | 0.02 m | 2 m | 5 px |
| 2 | 0.05 m | 4 m | 6.25 px |
| 3 | 0.08 m | 1.5 m | 26.7 px |
| 4 | 0.15 m | 5 m | 15 px |
| 5 | 0.40 m | 3 m | 66.7 px |
| 6 | 1.20 m | 8 m | 75 px |
APD-M. Within 0.1 m: predictions 1, 2 and 3, so 3 of 6 = 50%. Within 0.3 m: add 4, 4 of 6 = 66.7%. Within 0.5 m: add 5, 83.3%. Within 1.0 m: still 83.3%, since 6 is 1.2 m off. The average is (50 + 66.7 + 83.3 + 83.3) / 4 = 70.8.
APD-P. Comparing a threshold of δ · Z / f in 3D is the same as comparing the on-screen error with δ pixels, so use the last column. Within 1, 2 and 4 pixels: none, 0% each. Within 8: predictions 1 and 2, 33.3%. Within 16: add 4, 50%. The average is (0 + 0 + 0 + 33.3 + 50) / 5 = 16.7.
same six predictions: APD-M 70.8, APD-P 16.7
Same predictions, very different numbers. Prediction 3 is only 8 cm off, a pass for the metric thresholds, but at 1.5 m away that is almost 27 pixels on screen, a miss for every pixel threshold. Scale the errors yourself.
Each row is one threshold; each dot is one of the six predictions, filled when it is within that threshold. The score is the average of the rows. Switch the metric and make the tracker better or worse.
The six predictions are the illustrative ones in the table above, with every error multiplied by the slider. The thresholds are the paper's (Section 4.1): 0.1, 0.3, 0.5 and 1.0 m for APD-M; 1, 2, 4, 8 and 16 pixels, turned into 3D with the intrinsics, for APD-P. The benchmark's full definition (for example, how hidden frames are handled) is in the TAPVid-3D paper.
The appendix also reports three companion scores: Average Jaccard (AJ), a benchmark score that takes the visibility calls into account as well as the positions (its exact definition is in the TAPVid-3D paper); occlusion accuracy (OA), how often the visible-or-hidden call is right; and end-point error (EPE), the distance between prediction and truth, where lower is better. The paper notes that its conclusions do not change with them.
TAPVid-3D packs three kinds of real video: Aria Digital Twin (ADT), indoor videos from a head-worn camera, about 300 frames; DriveTrack, videos from cars, 25 to 300 frames; and Panoptic Studio (PStudio), recordings in a lab with many cameras, about 150 frames. Labelling every point of real dynamic scenes in 3D is not feasible, so the benchmark labels a sparse set of query points, which an all-point tracker simply includes among everything it tracks. Following prior work, the paper evaluates on the official "minival" split, 50 clips from each dataset, and hands the same VGGT-Ω geometry to TrackEverything and to most baselines.
Because earlier dense trackers only ever ran on short clips, the paper scores every method on the first 48 frames of each clip, and then, for those that can run at all, on the full-length videos.
The paper's Table 1. Bars are coloured by kind of tracker: sparse, first-frame dense, all-frame dense, and TrackEverything. Choose the clip length, the metric and the dataset.
All values from Table 1 of the paper, except ST4RTrack, which reports only two numbers and is left out (minival split, 50 clips per dataset; Ω = given VGGT-Ω geometry, native = the tracker's own geometry). D4RT's numbers are the ones its authors reported: it is closed-source, so the paper excludes it from ranking. "Cannot run" means the method does not process full-length videos; "out of memory" is SpatialTracker-v2 with its own geometry.
Three findings stand out. First, against the only other open-source all-frame dense tracker, VDPM, TrackEverything wins by a wide margin on the short clips: 34.7 against 10.9 average APD-P, the "more than 20% APD" of the abstract. D4RT, another all-frame dense tracker but closed-source, reports higher numbers on DriveTrack and PStudio (but lower on ADT, 40.8 against 45.7); it is about 20 times larger, trained with private data, has no public code, and was only evaluated on 48-frame clips.
Second, against the first-frame dense DeltaV2, given the same geometry, it is ahead on average and on ADT, and slightly behind on DriveTrack and PStudio, while solving the harder problem: DeltaV2's dense mode only follows points from frame 0.
Third, on full-length videos, only TrackEverything among the all-frame trackers runs at all, and it has the best average APD-P of every method that does: 31.2, ahead of the first-frame dense DeltaV2 at 29.9 and the best sparse tracker, TAPIP-3D, at 29.8, which only follows the queried points. On APD-M, CoTracker3 edges it (72.1 against 72.0). "Competitive" is the paper's own word, and it is the right one.
Where every point's truth is known, in synthetic data, the gap opens up. On 48-frame clips with estimated geometry (from Pi3), TrackEverything and DeltaV2 are close: 36.6 against 36.2 APD-P on Dynamic Replica, 24.9 against 23.8 on PointOdyssey. With the simulator's true geometry, TrackEverything jumps far ahead: 85.2 against 67.3, and 73.2 against 33.2 (Table 3). One caveat the paper notes elsewhere: TrackEverything trains on held-out splits of these two datasets, while DeltaV2 trained only on Kubric. And on the full 1,000+ frame PointOdyssey videos, scored on all 10,000+ labelled points, it reaches 31.8 APD-P and 74.5 APD-M, with 91.2% of static-or-dynamic calls correct. No baseline can run in that setting, so there is nothing to compare against.
Speed follows the size of the scene, not the length of the video. With an extra option described in the next chapter, the tracker alone runs at 23.1 frames per second on average over the three datasets: 41.5 on the compact PStudio lab, 19.5 on DriveTrack, and 8.4 on ADT's multi-room indoor scenes, which pile up the most unique surface. Including VGGT-Ω to estimate the geometry, the whole pipeline averages 7.2 frames per second.
Chapter 10
Measure what each part buys, see why the authors chose their settings, and learn where the method still breaks
Every design choice in this lesson has a knob. How big are the cubes? How long is a window? What if you remove 3D WAFT, or refinement, or the static-and-dynamic split? What if the geometry is worse, or better? The authors turned each knob and measured the result, in what papers call ablations: experiments that remove or change one part at a time to see what it was worth.
Read them with one caution. The cube-size sweep was run on PointOdyssey; the window-length and component ablations report TAPVid-3D's full-length average; the geometry knob uses the first 48 frames of PStudio and ADT. The numbers inside one knob compare cleanly; numbers across knobs do not.
Pick a knob, then a setting. The bars show the paper's measurements for that setting: accuracy (APD-P, higher is better), time per frame and peak memory (lower is better).
Every number is from the paper's Table 2. Cube size: PointOdyssey. Window length and parts: TAPVid-3D, full length. Geometry: APD-P on the first 48 frames of PStudio and ADT, without retraining, with D4RT's reported numbers for reference. "Not reported" means the paper gives no value for that setting.
Without any merging, the tracker runs out of memory: it keeps as many tokens as a frame-by-frame representation, which is the whole problem of Chapter 1. With cubes of 0.005, 0.01 and 0.02, accuracy barely moves: 31.4, 31.3 and 30.9 APD-P. Beyond that it falls: 28.5 at 0.05, 25.8 at 0.1, 23.5 at 0.2. Presumably, large cubes start merging surfaces that are genuinely different, a finger with the cup it touches, and detail is lost for good (the paper reports the accuracy drop but not this specific cause).
Speed and memory go the other way. Going from 0.005 to 0.02 cuts time per frame from 0.23 to 0.11 seconds and peak memory from 24.0 to 15.0 GB. The paper picks 0.02, in its words an effective trade-off that keeps peak accuracy while halving latency and memory. (Latency does halve; memory drops by about 38%.)
Windows of 8 and 16 frames score almost the same (31.3 and 31.2), and 24 frames drop to 29.8 while using the most memory (11.3 GB). The 8-frame window is even slightly faster at test time. The authors chose 16 anyway, because training with 16-frame windows is 26% faster per step. A reminder that the right setting depends on which cost you care about.
Why would a longer window hurt accuracy at all? The paper doesn't say, but the timing gives a plausible reason. A moving point only gets the safety net of merging, and a fresh zero-velocity restart, at a window's boundary. Stretch the window from 16 frames to 24 and a mover has to be tracked through more frames before that reset arrives, so whatever small errors the endpoint and trajectory refiners make have more distance to compound before the next window cleans things up. Shrinking the window the other way, to 8, gives more frequent resets but more windows overall to process, which is the extra memory and compute Chapter 5's endpoint-and-merge machinery has to repeat.
Three removals, all on the same benchmark. Removing iterative trajectory refinement costs the most accuracy, 31.2 down to 26.4. Removing 3D WAFT costs 31.2 down to 29.3. Removing the static-and-dynamic split costs no accuracy at all, 31.2 either way, but multiplies the time per frame by 4.8 and the memory by 2.5 (Chapter 5). One part buys accuracy, one buys accuracy and robustness on fast motion, and one buys pure efficiency. The paper reports no speed or memory numbers for the first two removals, only accuracy: the static-and-dynamic split is fundamentally about how much work happens (whether a path gets decoded at all), while 3D WAFT and iterative refinement change what the tracker predicts without changing how many points it processes, so the paper isolates their effect on accuracy alone.
Because the pointmaps are an input, you can swap their source without retraining. On PStudio and ADT, Pi3 geometry gives 23.5 and 39.3 APD-P; VGGT-Ω gives 30.1 and 45.7 (gains of 6.6 and 6.4); sensor depth gives 71.7 and 46.5. With sensor geometry, TrackEverything passes D4RT's reported 49.6 and 40.8 on both datasets, with about 20 times fewer parameters. D4RT cannot take external geometry at all. TrackEverything inherits every future improvement in geometry models and depth sensors for free.
The appendix describes an optional shortcut. The static-or-dynamic call settles after the very first refinement round, so after round 0, points labelled static whose predicted endpoints share a cube can be merged right away instead of at the window's end; they then skip every later round and the trajectory decoder. Moving points and query points are kept one for one. On full-length TAPVid-3D this left accuracy unchanged (31.3 against 31.2 APD-P) while raising the tracker's own throughput from 12.6 to 23.1 frames per second (1.8 times), and the whole pipeline with VGGT-Ω from 5.6 to 7.2 (1.3 times). Those are the speeds quoted in Chapter 9.
It is worth turning those rates into time per frame, because it shows where the time now goes. At 12.6 frames per second the tracker needs about 79 milliseconds a frame; at 23.1 it needs about 43. The whole pipeline goes from about 179 milliseconds a frame (5.6 per second) to about 139 (7.2 per second). If the two stages simply add up (our rough arithmetic, since the paper averages rates over datasets), estimating the geometry with VGGT-Ω takes roughly 100 milliseconds a frame either way. Once the tracker is this fast, the geometry model is the slower half. The paper notes the other side of that coin: when depth comes from sensors or pointmaps are precomputed, only the tracker runs, at 23.1 frames per second on average.
The paper is direct about four limitations.
The third limitation is the subtle one, so here it is in motion. A hand reaches past a cup just as a window ends.
Seen from above: a still cup and a moving hand. At frame 15 the hand's predicted endpoint lands near the cup; the grid shows the cubes. Change how close the prediction lands, and try protecting the cup's point.
An illustrative toy with cubes of 0.05. What happens to a merged track afterwards is not specified in the paper beyond "collapse into a single canonical track"; here it follows the hand, so the cup's own track is the one lost. Query points being excluded from merging is the paper's (Section 3.4).
The paper offers one remedy that exists today: points you care about can be given as query points, which are never merged, and so can never be swallowed. Improving the merge itself, perhaps by looking at appearance as well as position before fusing two tracks, is left as future work.
Chapter 11
Lock in the cheat sheet, read the whole loop as code, and place TrackEverything in the field
You can now explain every part of the instrument at the top of the page: why keeping a copy of the scene per frame runs out of memory, why a video is really a camera moving through a 3D world, how a window becomes a feature cloud, why points that land in the same cube at the same moment can be merged, why most points never need a path, how a guess is checked by looking at the image, and what each part is worth. Let's lock it in.
TrackEverything tracks every visible point of a video in 3D world coordinates by keeping one shared scene instead of one copy per frame. It reads the video in non-overlapping 16-frame windows. Each frame is encoded by a frozen DINOv3-small and a trained ViT Adapter, and its features are pinned to 3D with pointmaps from VGGT-Ω (or Pi3, or sensors), trimmed per frame in cubes of 0.02, and stacked with the points carried from earlier windows. For K rounds, an endpoint refiner (self-attention over the whole cloud, fed with 3D WAFT samples of the image where each guess lands) updates every point's position at the window's last frame, its visibility and whether it moves; a trajectory refiner decodes 16-frame paths only for the movers, starting from straight lines. At the window's end, points that land in the same cube are averaged into one, with receipts kept to rebuild every original track. Memory follows unique scene content, so videos over 1,000 frames fit in 40 GB, where earlier all-frame dense trackers ran out near 96 frames.
| Quantity | Value | Why it matters |
|---|---|---|
| Window | L = 16 frames, no overlap | 8 scores the same; 24 drops to 29.8; 16 trains 26% faster than 8 |
| Encoder | frozen DINOv3-small (20M) + trained ViT Adapter | features of shape (L, H/4, W/4, 384) |
| Parameters | 61M total, 41M trained | D4RT is about 20 times larger |
| Training hardware | 8 NVIDIA L40S GPUs (46 GB each) | 1 video per GPU per training step |
| Cube size | v = 0.02 scene units | peak accuracy, half the time of 0.005; no merging runs out of memory |
| One window's merge | 46k → 16k points (Figure 5) | about 35% kept |
| Frame copies vs one scene | 4.86M vs 13.4k after 400 frames | 362 times fewer (Figure 1) |
| Long videos | 1,000+ frames within 40 GB | earlier all-frame dense trackers: out of memory near 96 |
| Static/dynamic split | 0.05 vs 0.24 s per frame, 9.1 vs 22.4 GB | 4.8× faster, 2.5× less memory, same 31.2 APD-P |
| TAPVid-3D, 48 frames | 34.7 average APD-P | VDPM 10.9; DeltaV2 34.4; D4RT reports 43.8 (closed) |
| TAPVid-3D, full length | 31.2 average APD-P | best of every method that can run; VDPM and D4RT cannot |
| Speed | 23.1 fps tracker only, 7.2 with VGGT-Ω | with the early static merge; 41.5 fps on compact PStudio |
Here is the whole tracker in NumPy-flavoured Python. The geometry helpers are complete; the three learned modules (the adapter inside encoder, endpoint_refiner, trajectory_refiner) are the networks of Chapters 3, 5 and 7, and dedup is Chapter 4's function.
pythonimport numpy as np L, V = 16, 0.02 # window length and cube size (the paper's defaults) def unproject(depth, f, cx, cy, R, t): # Chapter 2: pixels + depth -> world points v, u = np.mgrid[0:depth.shape[0], 0:depth.shape[1]] cam = np.stack([(u - cx) * depth / f, (v - cy) * depth / f, depth], -1) return cam @ R.T + t # (H, W, 3) in camera 0's frame def project(x, f, cx, cy, R, t): # Chapter 6: world points -> pixels p = (x - t) @ R # R^T (x - t), one row per point return np.stack([f * p[:, 0] / p[:, 2] + cx, f * p[:, 1] / p[:, 2] + cy], -1) def bilinear(fmap, uv): # blend the 4 nearest feature cells h, w, _ = fmap.shape u = np.clip(uv[:, 0], 0, w - 1.001); v = np.clip(uv[:, 1], 0, h - 1.001) u0, v0 = u.astype(int), v.astype(int); a, b = (u - u0)[:, None], (v - v0)[:, None] return ((1 - a) * (1 - b) * fmap[v0, u0] + a * (1 - b) * fmap[v0, u0 + 1] + (1 - a) * b * fmap[v0 + 1, u0] + a * b * fmap[v0 + 1, u0 + 1]) def track_video(frames, depths, cams, encoder, endpoint_refiner, trajectory_refiner, K): scene = Points.empty() # carried memory: xyz, feat, t_src = -1, lineage for w0 in range(0, len(frames), L): idx = list(range(w0, min(w0 + L, len(frames)))) fmaps = [encoder(frames[i]) for i in idx] # frozen DINOv3 + adapter: (H/4, W/4, 384) new = [] for j, i in enumerate(idx): # pin each frame's tokens to 3D, trim per frame xyz = unproject(depths[i], *cams[i])[::4, ::4].reshape(-1, 3) new.append(Points.voxel_mean(xyz, fmaps[j].reshape(-1, 384), V, t_src=j)) pts = Points.concat([scene] + new) # the window's space-time cloud, (N, 384) pts.x_tgt = pts.x_now.copy() # zero-velocity start: birth spot, or last window's endpoint t_tgt = len(idx) - 1 # the window's last frame g_src = pts.sample_birth(fmaps, cams, idx) # once per point (carried points keep theirs) for k in range(K): # K refinement rounds g_tgt = bilinear(fmaps[t_tgt], project(pts.x_tgt, *cams[idx[t_tgt]]) / 4) dx, vis, dyn = endpoint_refiner(pts.feat, g_src, g_tgt, pts.embed(t_tgt)) pts.x_tgt += dx; pts.vis_end = vis; moving = dyn > 0 pts.paths[moving] = trajectory_refiner(pts, moving, g_src, fmaps, cams, idx, first=(k == 0)) pts.paths[~moving] = pts.x_tgt[~moving, None, :] # still points: endpoint repeated yield idx, pts.paths, pts.vis_end # this window's dense 3D tracks scene = Points(*dedup(pts.x_tgt, pts.feat, pts.t_src, pts.x_src, v=V)) # Chapter 4; queries never merged del fmaps # 2D features freed; only the 3D scene is carried
Two honest gaps in that sketch. The paper does not state K, the number of refinement rounds, so it is left as an argument. And the four position embeddings, the extended-canvas clamping for guesses that fall outside the picture, the early static merge, and the lineage tables that rebuild every original pixel's track (Xorig(t) = Xmerged[p](t) + r) are folded into the helper object Points.
Now press Present or Teach and explain, out loud and from memory, why merging points at the end of a window is safe while merging them during encoding is not. If you can, you own this paper. Then go back to the two rooms, switch to one copy per frame, and watch the memory run out at frame 96.