Jev · Robotics · Field report

Testing a fast judge for robot supervision

We tested whether Jev, a small model that returns probabilities instead of text, can help supervise robots. We fed it a vision model's description of recorded robot episodes and asked it two kinds of questions. When we asked it to grade each episode at the end, it didn't do any better than the vision model alone. When we asked it after every observation whether something was going wrong, it flagged six of the eight failed episodes, usually about halfway through, and each answer took about 0.2 seconds. The vision model in front of it was still slow, so the full pipeline wasn't real-time, and we didn't compare against simply asking the vision model the same questions. What we showed is that the cheap decision layer produces a usable per-step signal; whether it beats the obvious alternative is the next test.

September 22, 2026 engineermaxxing.com · built with a PromptQL bot Vision model GPT-5.6 Sol · Judge Jev 1.13.0

TL;DR

  • Using Jev to re-grade finished episodes didn't help. On 20 human-labelled RoboReward clips, the vision model alone got 11 of 20 grades exactly right; having Jev re-grade the vision model's notes got 10 of 20. Jev only reads text, so when the vision model misread the scene, Jev repeated the mistake.
  • Using Jev as a step-by-step monitor looked more useful. We asked it six questions after every observation. It flagged 6 of the 8 failed episodes, 5 of them before the recording ended (on average about halfway through), and it cleanly separated episodes we deliberately broke from normal ones (P(done) 0.92 vs 0.02). Each set of answers took 0.16–0.20 seconds and cost about $0.00004. Its final verdicts were slightly less accurate than the vision model's, so the argument for it is speed, cost and getting an answer at every step, not accuracy.
  • What this doesn't show: that Jev is needed. The vision model ran on every step anyway, so the loop was no faster than a VLM alone, and we didn't test a VLM asked the same six questions in JSON. Jev's advantage is speed and cost per decision (0.2 s, $0.00004) and native probabilities, which only pay off once perception is also cheap.
  • Two new papers — ARLI [1] and Real-Time EXPO-FT [2] — show how to train robots that keep learning even though their big models are slow to think. Real-Time EXPO-FT gets its reward from hand-written per-task success detectors; ARLI takes the reward as an input and does not study how it is made. A fast, per-step "did it work / is something wrong" probability is exactly the piece they leave out — and it is the kind of signal our monitor produces. Wiring it into their loops is future work.

The problem

Here is a robot arm asked to remove the blue cup from the plate. Five of the ten frames we sampled, spanning about 40 seconds.

t=0 s: blue cup on purple plate
0 s
t=13 s: gripper above the cup
13 s
t=22 s: gripper grasps the cup
22 s
t=27 s: cup placed on the table
27 s
t=40 s: final state
40 s
Clip d094b57b from the human-verified test split of RoboReward. We renamed every clip to a hash so no filename could leak its label.

Did it succeed? A human rater said 2 out of 5 — minimal progress. A frontier vision model, shown the same frames, said 5 — perfect completion. We'll come back to why. The point for now is that someone has to decide, and every part of a modern robot-learning pipeline depends on that decision:

  • Evaluation needs success labels. Today those are mostly humans watching video.
  • Learning by trial and error (reinforcement learning) needs a score after every attempt — the "reward" — that tells the robot how well it did. Today that score comes from a hand-written checker built for each task, or a human pressing a button.
  • Deployment needs a supervisor that can say "stop" or "ask a person" while the robot is still moving — not after the episode is over.

The usual way to automate this is to ask a vision-language model. That mostly works, but it has two costs that add up when you want a decision on every step of every episode: each call takes several seconds (9–13 s in our setup), and each call costs a few cents. Multiply that by thousands of training episodes and you have a supervisor that is too slow to sit in a control loop and too expensive to run everywhere.

Jev is a different kind of model. It doesn't describe anything; given a short text state, it returns probabilities for a fixed set of questions in about 0.2 seconds for a fraction of a cent. The question we wanted to answer is whether that kind of model can be the decision layer, with something cheaper than a frontier VLM doing the perception.

Two things to be clear about before the results. First, we did not get that far: in every experiment below the perception in front of Jev was still a frontier VLM, so the pipeline as a whole was not fast. What we could measure is whether Jev's decisions, given a description, are good enough and cheap enough to be worth building the fast version around. Second, an obvious alternative is to skip Jev and ask the VLM for the same six answers as structured output. We did not run that comparison, so nothing here shows Jev is more accurate than a VLM asked the same questions. The case for it is that it is roughly 50× faster and 100–1000× cheaper per decision, returns a probability from its output distribution rather than a number written into a sentence, and keeps the judging rubric fixed and separate from whatever is doing the seeing. Whether those properties matter in practice is the thing to test next.

What Jev is

Jev is a model that does not write. You give it a state — any structured or unstructured text — and one or more questions. Each question has a type, and Jev returns an answer of that type together with how sure it is. There are three question types — Noul (yes/no; "Noul" is simply Jev's name for a yes/no question), Choice (which one) and Score (how much). In robotics terms, the three map onto the questions a supervisor actually asks:

NOUL · yes/no
"Is the brick in the box?"
brick_in_box: 0.93
CHOICE · which one
"Which phase is the arm in?"
reach 0.05 grasp 0.89 transport 0.04 place 0.01 recovery 0.01
SCORE · how much
"How far along is the task?"
progress: 3.2 / 4 levels: [0.00 0.02 0.14 0.46 0.38]
The five numbers are how likely each level 0–4 is; 3.2 is their average.

Three properties matter for what follows:

  1. Jev reads text, not images. In these experiments it never saw a frame. It sees whatever you put in the state — a description of the scene, a gripper encoder reading, the task instruction, the last few observations.
  2. Every answer is a probability rather than a sentence. 0.93 is something a program can threshold, log, plot, or use as a reward. "The brick appears to be in the box" is not. (We did not formally measure how well-calibrated these probabilities are — more on that in the limits.)
  3. It is fast and cheap. In our runs, one call — six questions asked at once — took 0.16–0.20 seconds and cost about four thousandths of a cent. That is fast enough for a high-level supervisory loop at a few hertz — not a low-level control or safety loop, and not a chat turn.

So what is it for?

Jev is a decision layer. Something else perceives; Jev turns the perception into typed, probabilistic decisions that code can act on immediately. That framing is the whole result of this project — and we did not start with it.

How we wired it

The same shape of pipeline runs through every experiment below, with protocol differences noted in each part. A vision-language model (VLM) watches frames and writes a short observation. That observation, plus the task and any sensor readings, becomes Jev's state. Jev answers a fixed set of typed questions. The probabilities feed a dashboard — and could equally feed a controller.

Camera frames
Sampled from recorded episodes at 1–2 frames per second, up to 512 px wide.
pixels
VLM narrates
GPT-5.6 Sol describes what changed. Structured JSON: observation, gripper state, anything unusual.
≈ 9–13 s / call
Typed state
Task + the observations so far (Part 3) or the current one plus the last three and the gripper encoder (Part 2). Never the reference label.
text
Jev answers
One call that asks all six questions at once: yes/no Nouls, a Choice, a Score.
≈ 0.2 s / call
Probabilities
Plotted in Rerun (the replay dashboard you can open below); thresholded for early-warning, routing, or reward.
numbers

One caveat up front: everything here ran offline, on recorded episodes. Only the Jev stage is fast; the vision model in front of it takes 9–13 s per call, so the pipeline as tested is not yet an end-to-end real-time supervisor. What we measured is whether Jev's decisions, given a description, are good enough and fast enough to be one once perception is. We did not compare against asking the VLM itself for the same structured answers; see the limits section.

A few rules we held ourselves to, because it is easy to fool yourself in this kind of evaluation: the human reference label was only ever loaded in the scoring script, never in any prompt; the dataset's own model-generated check field was excluded; prompts were fixed on three practice clips from a separate split and never tuned on the 20 test clips; every number below is from real API calls.

Part 1Grading finished episodesIt didn't work — and the reason mattered

The obvious first question: if a vision model grades a robot episode, does adding Jev as a second opinion improve the grade?

We took 20 clips from RoboReward's human-verified test split, balanced four per reward level (1 = no success … 5 = perfect). Agent A was the VLM alone: watch ten frames, output a level. Agent B was the VLM's observations, evidence and uncertainty notes handed to Jev as a Choice over the same five-level rubric, with the VLM's proposed label included but an instruction not to defer to it.

11/20
Agent A (VLM only) exact matches · average distance from the human score 0.60 (lower is better)
10/20
Agent B (VLM → Jev) exact matches · average distance 0.75
2 / 3
Labels Jev corrected / labels Jev broke
7/9
Agent A's errors that were visual misreads, not reasoning errors

Jev did not help, and the pattern of where it hurt is instructive. On clips the human rated 4 or 5, Agent A got 5 of 8 right; Agent B got 2. Whenever the VLM's notes hedged — "intermediate grasp details are between sampled frames" — Jev read the hedge as evidence against completion and downgraded. That is a reasonable thing for a judge to do with the text it was given. It is the wrong thing to do with these clips.

Why it didn't help

Look again at the blue-cup clip. The VLM's final observation was "blue cup upright on table, no longer on purple plate; orange cup remains on plate." Given only that text, "perfect completion" is the natural reading — and Jev agreed, with full confidence. We don't know exactly what the human rater saw — perhaps a drop, a re-grasp, or a plate that moved between our sampled frames — but whatever it was, it never reached the text, so it could never reach Jev. When we (one author, unblinded, against the full footage) went through Agent A's nine mistakes, we classified seven as observation-related — the description disagreed with what we could see, so every downstream judgment inherited the disagreement. Whether the human rater or the model is right about this particular clip we leave open; either way, Jev could only judge the text it was given.

Lesson

A text-only judge cannot fix a perception error. If the goal is a better final grade, the investment belongs in the eyes, not the judge. In a separate test we also asked Jev to act as a verification gate — "is the VLM's label justified by its own evidence?" — and it was no better than chance at catching the VLM's mistakes (AUROC 0.53 for "label not justified" against the 9 errors in 20 clips), for the same reason.

This could have been the end. Instead it made us ask what Jev is for, if not for grading.

Part 2Monitoring episodes step by step

Grading is a one-shot question asked at the end. Monitoring is a stream of small questions asked continuously — and those are exactly the questions that are cheap for Jev and expensive for everyone else.

We switched datasets to something closer to a real deployment: LeRobot's SO-101 pick-and-place episodes (task: pink lego brick into the transparent box; two cameras, six joint encoders at 30 fps). For every one-second window of footage, the VLM writes a short narration of what changed (three top-camera frames plus one side frame). Jev then answers six questions in one call:

  • Choice phase — reach / grasp / transport / place-release / retreat-idle / recovery
  • Score progress — five ordered levels, returned as a position 0–4 plus a distribution
  • Nouls (yes/no) brick_grasped, brick_in_box, anomaly, safe_to_continue

Here is the dashboard for episode 40, scrubbed from start to finish. Watch the anomaly line.

SO-101 episode 40 in Rerun. Top row: both cameras, the VLM's narration, and Jev's current answers. Middle: the four Noul probabilities over time, the progress Score, and the confidence of the phase Choice. Bottom: phase distribution, progress-level distribution, joint encoders. At t = 3 s the narration reports the gripper closed without the brick; Jev answers P(anomaly) = 0.75 and phase recovery. The arm re-approaches, and P(brick_in_box) crosses 0.9 at 8 s.
EpisodeDurationJev stepsP(in box) ≥ 0.9 atMax P(anomaly)Note
ep0087.5 s75.0 s0.07clean
ep0178.1 s86.0 s0.22clean
ep0377.6 s76.0 s0.08one low-confidence phase step
ep0409.8 s98.0 s0.75 @ 3 smissed grasp, flagged, recovered
ep0458.3 s86.0 s0.41clean

39 one-second steps; every Jev call returned successfully. Jev: 0.16 s per step, $0.0022 for all five episodes. VLM: 8.9 s per step. All five episodes are recorded successes; the monitor agreed by the end of each, and the one hiccup it flagged along the way is visible in the footage.

Two things are already different from Part 1. The signal is per step, so a problem is visible while there is still time to do something about it. And the outputs are typed: P(anomaly) > 0.7 is a line of code, not a judgment call about a paragraph.

Part 3Testing the monitor harder

Part 2 only had successes. To test the monitor properly we needed failures, so we ran three experiments designed to surface them. A word on what these can and cannot show: they establish that per-step typed judgments from Jev are feasible and cheap, not that Jev is a better streaming judge than a VLM asked the same questions at the same cadence — we did not run that matched baseline.

3.1 Alarms on real failures

We went back to the 20 RoboReward clips, but this time asked Jev after every observation — 163 calls in total — instead of one grade at the end. The six streaming questions: is the task done, is something going wrong, should a person look at this (three yes/no Nouls), plus progress, grasp and placement Scores. Jev never saw the human labels; we used them only afterwards to check whether the alarm went off on the clips that really failed.

How we scored an alarm

Alarm = P(anomaly) ≥ 0.5 at any step (threshold fixed at 0.5 in advance, not tuned; clips have 5–10 observations, so longer clips get more chances to alarm). Failure = human rating 1 or 2 (8 clips); success = 4 or 5 (8 clips); the four level-3 "partial" clips are reported separately. Alarm positions are measured in recording time — where in the clip the alarm fired — and exclude the pipeline's own processing delay (see 3.3).

A real failure the monitor caught early. "Pick up the silver metal bowl and move it to the center of the table." The arm hovers over the dish rack and never gets the bowl. P(anomaly) is 0.57 by t = 2.3 s and 0.92 by the end; needs_human stays above 0.83 throughout. Human label: 1 (no success). The VLM's one-shot label was 2.
The blue cup clip again. P(task_done) sits at 0.02 for 22 s, then jumps to 0.92 the moment the VLM reports the cup on the table. That is exactly the behavior you want from a monitor. But the human still says 2. Jev is only ever as right as the observations it is given.
6/8
failures that raised an alarm (5 of them before the final observation)
1/8
successes that raised a false alarm (at 90% of the clip). All 4 partial clips alarmed.
53%
mean position in the recording at first alarm, over the 6 alarmed failures
0.83
How well P(task_done) tells human-rated successes (4–5) from the rest, where 0.5 = coin flip and 1.0 = perfect (AUROC). The VLM's own final grade: 0.875 — slightly better.

Read that last card carefully. On the final verdict, Jev is not better than the VLM — it is working from the same text, so it has the same blind spots, and on every endpoint metric it is a little behind:

Endpoint metric (n = 20; Jev value = its final-step output)Jev streamingVLM final label
AUROC, success (human 4–5) — P(task_done) vs label0.830.88
AUROC, failure (human 1–2) — P(anomaly) vs inverted label0.710.79
Accuracy as a success classifier — P(done) ≥ 0.5 vs label ≥ 40.650.80
Spearman with human rating — progress Score vs label0.680.75

With twenty clips these gaps are not statistically resolved: a clip-level bootstrap puts the success-AUROC difference at −0.13 to +0.03 (95%) and the failure-AUROC difference at −0.24 to +0.04. The honest summary is "not better, plausibly a little worse", not a ranking.

Its case is elsewhere: it produces a probability at every step, so problems surface part-way through instead of at the end; the anomaly probability separates failures from successes usefully, if imperfectly (means of 0.68 vs 0.19); and each of those 163 decisions took about 0.2 s and $0.00004. The per-clip picture:

Per-clip timeline showing where in each recording the first P(anomaly) >= 0.5 alarm fired, grouped by human rating
Where the first alarm fired in each clip. Six of eight failures alarm, one of them only at its final observation; all four partial clips alarm (arguably correctly — none reached the goal); one success alarms late. Two failures never alarm — including the blue cup, where the narration itself reported success.

The needs_human question deserves a footnote. It caught 8 of the VLM's 9 labeling errors, which sounds great, but it fired on 17 of 20 clips. As a "route to a person" signal it is far too eager at a 0.5 threshold. Useful, but not free.

3.2 Controlled incompletion and instruction-mismatch tests

Real datasets are mostly successes, so we constructed failure conditions from the same five SO-101 episodes — paired transformations, not fifteen independent rollouts — keeping all the VLM narration real: truncate the episode at 3 s (the arm is mid-reach, nothing is in the box), and swap the instruction (tell Jev the task was to put the brick next to the box, while the footage shows it going in). For the second we added a seventh question, task_alignment: is what's happening the task that was asked for?

Bar chart: P(task_done), P(task_alignment) and P(anomaly) for nominal, truncated and instruction-swapped SO-101 episodes
Normal episodes compared with the two kinds of failures we injected, averaged over 5 episodes. P(task_done): 0.92 → 0.02 → 0.02. P(task_alignment): 0.97 → 0.62 → 0.05. P(anomaly): 0.05 → 0.36 → 0.71. Every nominal episode is separated from every failure on P(task_done); the instruction swap is separated on alignment.

This is the cleanest separation in the project. Be clear about what it shows: Jev recognises an unfinished state and a mismatched instruction from the narration. It does not yet show detection of naturally occurring physical failures at scale, and these scalars have not been used as a training reward. What it does show is the right shape for one: high when the goal state is reached, near zero when it isn't, and a second scalar that fires when the robot is doing the wrong task well.

3.3 Confidence-triggered re-checks (retrospective)

Because the phase Choice comes with a confidence, one could in principle spend the expensive model only when the cheap one is unsure. We tried the mechanics: if phase confidence is below 0.6, ask the VLM for a wider-window re-look (which, in this offline setting, included frames up to a second after the queried moment) and ask Jev again. On the SO-101 episodes this fired on 2 of 39 steps, both in the first second of an episode when "idle" and "starting to reach" look alike. On episode 37 the confidence went from 0.48 to 0.91; on episode 40 only from 0.48 to 0.53 — arguably right, since at t = 1 s the arm had not committed. Two caveats: higher confidence is not demonstrated correctness, and because every step in these experiments already used the VLM, the gate here only adds ~10–12 s re-checks. The saving only exists in a future setup where routine perception is cheap and the VLM is reserved for the uncertain steps.

Timing diagram of one monitoring step as tested: VLM 8.9 to 13.3 seconds followed by Jev 0.16 to 0.20 seconds
How long one monitoring step took in our setup. The vision model is almost the whole budget; Jev is the last quarter-second. This is why nothing here is live yet, and why the obvious next input is cheap, frequent perception rather than a frontier VLM.
Jev does not see better than a VLM. It decides faster and cheaper than one — every step, in numbers, with a confidence attached. Pair it with fast enough perception and supervision could stop being a thing you do at the end of an episode and become a thing you do inside a supervisory loop — a claim this pilot motivates but does not yet demonstrate.

Explore the data yourself

Everything above is logged in Rerun. The viewer below is the same one we used; it loads in your browser (about 60 MB the first time). Pick a clip or episode tab along the top, then scrub the timeline. Reference labels appear only in the panel marked eval layer only.

Interactive Rerun viewer
Streaming monitor on 20 RoboReward clips · SO-101 monitor · the Part-1 A/B comparison
ARLI teaser figure
Overview of ARLI. Figure from the ARLI project page.
ARLI intermediate-state example
What ARLI adds to the state: the committed actions and a mid-inference observation are added to the state. Figure from the ARLI project page.
The assembly task, a failed attempt (bimanual UR5e, π0.5 base policy). Video from the ARLI project page.
The same task, a successful attempt. Something has to say which of these two is which, for every trial, before RL can learn from it. Video from the ARLI project page.

What it shares with our work is a loose design analogy rather than a mechanism: the same villain (big-model latency), a similar instinct (use intermediate information rather than waiting for the episode to end — our streaming Jev re-scores after every observation prefix, ARLI augments the RL state mid-inference), and the same pairing of a slow, expressive model with a fast, cheap one. Our prototype is sequential and offline; ARLI's is a genuinely asynchronous acting loop. What it does not set out to address: where the reward comes from. ARLI is a policy-learning paper; the success signal is an input to the method, not something it produces.

Real-Time EXPO-FT — RL for real-time VLA policies [2]

Reinforcement Learning for Real-Time Vision-Language-Action Policies. Dong, Hung, Sadigh, Finn · Stanford · September 2026 · arXiv 2609.18207 · project page · Interactive Real-Time EXPO-FT walkthrough on Engineermaxxing

Same latency problem, attacked with a slow/fast split on the control side: a large VLA proposes candidate action chunks (slow), a small edit policy adjusts them using the observation available at the moment of execution (fast), and a learned scorer (a Q-function) picks the best candidate. On four dynamic real-world tasks — ball balancing, soccer kicking, dynamic picking, mid-air object passing — the authors report success rising from 42% to 97% with about ten minutes of real robot data.

The project teaser video (first 14 s). Video from the project page.
Dynamic picking with Real-Time EXPO-FT: grasping an object on a rotating base; 4 s to success. Video from the project page.
Dynamic picking with the baseline (a policy trained only by copying demonstrations, with real-time chunking). Video from the project page.

The paper is candid about its weak link. The reward is a sparse binary signal from a hand-written, rule-based success detector for each task, with a human verifying success during evaluation, and the authors list it as a limitation: designing a classifier per task is work, and which reward formulation performs best remains open.

Side by side

Policy · CoRL 2026 submission
ARLI
Slow thing
Generalist VLA (π0.5), 100s of ms per chunk
Fast thing
Low-latency RL policy steering the next chunk
What the fast thing consumes
Committed actions + a mid-inference observation
Reward
An input to the method; how it is produced is not the paper's subject
Policy · Sept 2026
Real-Time EXPO-FT
Slow thing
VLA proposing action chunks
Fast thing
Edit policy + Q-function choosing among chunks
What the fast thing consumes
The latest observation at execution time
Reward
Hand-written per-task success detector; listed as a limitation
Supervision · this post (offline prototype)
VLM + Jev monitor
Slow thing
VLM narrating the scene, 9–13 s
Fast thing
Jev, 0.2 s per six typed questions
What the fast thing consumes
Typed state: task, observations, encoders
Reward
Candidate signals — P(task_done), P(anomaly), progress Score, task_alignment — per step; not yet used for RL

The pattern rhymes across the three columns: a slow, expressive model that cannot run at loop rate, and a fast, cheap layer that consumes intermediate information. ARLI and Real-Time EXPO-FT build that layer on the control side and take the reward as given — in EXPO-FT's case, explicitly, from a hand-written detector per task; ARLI does not discuss reward generation, so we make no claim about it. Our monitor is an offline prototype of the analogous layer on the supervision side. It is not a competing method and not yet an integrated one; it is a complementary direction, and the integration is the experiment we have not run.

One line

The fast loop for acting now exists. A fast judge for it is worth testing.

Limits and next steps

  • This is a small pilot: 20 clips (balanced by rating, so not the natural failure rate), 5 episodes, 163 streaming decisions. The constructed-failure separations are large and consistent, but the AUROC estimates have wide error bars at n = 20 (Jev success AUROC 95% CI 0.61–0.99).
  • We did not measure calibration. Jev returns probabilities, and they rank successes above failures reasonably well, but we did not check whether "0.9" is right nine times in ten (no reliability curves, no Brier score). Anyone using these as rewards should.
  • Perception is still the bottleneck. Jev is a decision layer on top of whatever describes the scene. Where the VLM misread the scene, Jev was confidently wrong along with it. The obvious next experiment is to give Jev cheaper, more frequent perception — a small vision model at 5–10 Hz, or the robot's own state estimates — instead of a frontier VLM every few seconds.
  • Everything ran offline, and the full pipeline is not yet fast enough for control. We ran on recorded episodes, and the vision model in front of Jev takes 9–13 s per call. Nothing here has yet gated a live controller or been used as an RL reward.
  • There is no matched baseline, and this is the biggest gap. We compared Jev's streaming answers with the VLM's one-shot final grade, not with the VLM asked the same six questions in structured output at the same cadence. Since the VLM was already running every step, that comparison would have cost nothing extra to run. Until it's done, the 6-of-8 alarm result is evidence that per-step monitoring works, not that Jev is the reason.
  • The needs_human question fires too often: 17 of 20 clips were flagged at a 0.5 threshold. The question needs sharper criteria or a higher threshold before it is useful for routing to people.
  • The prompts were written once and never tuned. The rubric language, the "must be observed, not inferred" rules, and the phase boundaries were all written once on three setup clips. We did not iterate on them; better ones surely exist.

What we would do next, in order: first, put fast perception in front of Jev (a small vision model at 5–10 Hz, or the robot's own state estimates) and measure causal alarms and false-alarm rates on live or replayed streams; second, measure calibration and check that the questions transfer across tasks without rewriting; third — only then — plug P(task_done) and P(anomaly) into an RL fine-tuning loop like ARLI's or Real-Time EXPO-FT's as the per-step reward and stopping signal, against a hand-written detector, and measure both learning performance and reward exploitation.

So where does Jev actually help? Not in the setup we ran. With a frontier VLM in the loop, the VLM is 98% of the latency and nearly all of the cost, and it could plausibly have answered the six questions itself. Jev helps if two other things are true, neither of which we tested: that perception can be made cheap (a small detector, the robot's own state, or a tiny VLM at a few hertz), and that a native probability from a fixed rubric is a better reward signal than a number an LLM writes into its answer. If both hold, you get a supervisor that can run on every step of every episode for a training run at a cost that rounds to zero. If they don't, a VLM with a JSON schema is the simpler tool. This pilot establishes that the decision layer is fast, cheap and produces the right shape of signal; it does not yet establish that you need it.

Appendix

Setup and data
  • RoboReward (teetone/RoboReward, revision 469b9af76e8539e8d2ac553b081307738fd92ca5), test split, human-verified. 2,831 rows; we sampled 20 balanced by reward level with seed 20260920, plus 3 setup clips from validation for prompt design. 10 frames per clip, evenly spaced by index, ≤ 512 px. The gpt5_mini_check field was excluded. Filenames (which encode the score) were replaced with SHA-1 hashes.
  • LeRobot SO-101 pick-place (lerobot/svla_so101_pickplace, v3.0, revision f641879e), 5 of 50 episodes, 30 fps, cameras up and side, six joints. VLM window 1 s (three top frames + one side frame).
  • VLM: GPT-5.6 Sol via the PromptQL program runtime, structured JSON output. Jev: jev-1.13.0 via POST /v1/systemone, one fan-out call per step. Pricing basis $0.042 per million input tokens.
  • Logging: Rerun SDK 0.38.1; recordings served with the Rerun web viewer.
  • Latency figures: Jev 0.16 s/step (Part 2), 0.20 s/step (Part 3), 0.12–0.16 s (Part 1); VLM 13.3 s (Part 1, ten frames), 8.9 s (Part 2, four frames). VLM cost was not separately metered, which is why we quote only Jev's cost.
Jev question set (streaming monitor)
# state (abbreviated) — what Jev actually reads
{ "task": "Remove the blue cup from the plate",
  "observations_so_far": [ {"t": 0.0,  "event": "Blue cup stands on purple plate…"},
                            {"t": 22.3, "event": "Gripper grasps blue cup on the plate."} ],
  "rubric": ["No success", "Minimal progress", "Partial", "Near completion", "Perfect completion"] }

# questions — one call, six typed answers
task_done   : noul   "Based ONLY on the observations so far, has the task been fully completed?
                       The final goal state must be observed, not inferred."
anomaly     : noul   "Is something going wrong that a supervisor should know about?"
needs_human : noul   "Should a person review this before trusting an automatic label?"
progress    : score  5 ordered levels (the rubric), returns position + distribution
grasp       : score  0–3, has the target object been secured?
placement   : score  0–3, has it reached the goal location?

The SO-101 monitor swaps in phase (Choice over six phases with explicit "means / does not mean" boundaries), progress (Score, five levels) and four Nouls (brick_grasped, brick_in_box, anomaly, safe_to_continue); the injected-failure test adds task_alignment. The full definitions are in the source bundle.

All the numbers
ExperimentMetricValue
Part 1 · A/B grading (n = 20)Agent A exact / MAE (lower is better) / latency11/20 · 0.60 · 13.3 s
Agent B exact / MAE (lower is better) / added Jev latency10/20 · 0.75 · 0.16 s
Corrections / regressions A→B2 / 3
Agent A errors that were visual misreads7 / 9
Jev tokens / cost19,156 · $0.0008
Part 2 · SO-101 monitor (5 eps)Steps / API call errors39 / 0
First P(in box) ≥ 0.95–8 s
Jev / VLM latency per step0.16 s / 8.9 s
Jev tokens / cost53,215 · $0.0022
Part 3.1 · streaming (20 clips)Jev calls / latency / cost163 · 0.203 s · $0.0064
Failures alarmed, P(anomaly) ≥ 0.5 (mean recording position)6/8 (0.53)
Successes / partial clips that alarmed1/8 · 4/4
Mean P(anomaly) failures / successes0.68 / 0.19
AUROC success — P(task_done) / VLM label0.828 / 0.875
AUROC failure — P(anomaly) / VLM label0.708 / 0.792
Spearman vs reference — progress / VLM label0.681 / 0.75
needs_human ≥ 0.5: flagged / VLM errors caught17/20 · 8/9
Successes with P(done) ≥ 0.5 before end2
Part 3.2 · injected (5 eps × 3)P(task_done) nominal / truncated / swapped0.92 / 0.02 / 0.02
P(task_alignment)0.97 / 0.62 / 0.05
P(anomaly)0.05 / 0.36 / 0.71
Part 3.3 · confidence gateSteps re-checked2 / 39
Phase confidence after re-checkep037 0.48→0.91 · ep040 0.48→0.53
Reproducibility

All code (data prep, VLM and Jev pipelines, evaluation, Rerun builders), raw JSON outputs, the three .rrd recordings and this page are in the project source bundle. Jev prompt versions: jev-choice-v1 (Part 1), lr-jev-v1 (Part 2), jev-stream-v1 (Part 3). VLM prompt versions vlm-v1, lr-vlm-v1. Total Jev spend across everything: about one cent.

References

  1. B. Zhu, M. Khalil, E. Harrison, E. Poggi, P. Schmitt, B. Kast, P. Meister, P. Atreya, Q. Li, F. Ferchau, C. Colmenero, Y. Shahapurkar, G. Narayanan, M. Erdogan, K. Wurm, G. von Wichert, O. Mees, E. Solowjow, A. Wagenmaker and S. Levine. Learning to Act While Waiting: RL Finetuning of Generalist Robot Policies Under Inference Latency. arXiv:2608.23831, August 2026 (submitted to CoRL 2026). arxiv.org/abs/2608.23831 · project page.
  2. P. Dong, K.-H. Hung, D. Sadigh and C. Finn. Reinforcement Learning for Real-Time Vision-Language-Action Policies. arXiv:2609.18207, September 2026. arxiv.org/abs/2609.18207 · project page.

Paper figures and videos in "Where a learned supervisor might fit" (refs [1], [2]) belong to their authors and are reproduced from the public project pages for comparison; see the links above. RoboReward © its authors (CC). LeRobot SO-101 dataset © its contributors.