The problem
Here is a robot arm asked to remove the blue cup from the plate. Five of the ten frames we sampled, spanning about 40 seconds.





d094b57b from the human-verified test split of RoboReward. We renamed every clip to a hash so no filename could leak its label.Did it succeed? A human rater said 2 out of 5 — minimal progress. A frontier vision model, shown the same frames, said 5 — perfect completion. We'll come back to why. The point for now is that someone has to decide, and every part of a modern robot-learning pipeline depends on that decision:
- Evaluation needs success labels. Today those are mostly humans watching video.
- Learning by trial and error (reinforcement learning) needs a score after every attempt — the "reward" — that tells the robot how well it did. Today that score comes from a hand-written checker built for each task, or a human pressing a button.
- Deployment needs a supervisor that can say "stop" or "ask a person" while the robot is still moving — not after the episode is over.
The usual way to automate this is to ask a vision-language model. That mostly works, but it has two costs that add up when you want a decision on every step of every episode: each call takes several seconds (9–13 s in our setup), and each call costs a few cents. Multiply that by thousands of training episodes and you have a supervisor that is too slow to sit in a control loop and too expensive to run everywhere.
Jev is a different kind of model. It doesn't describe anything; given a short text state, it returns probabilities for a fixed set of questions in about 0.2 seconds for a fraction of a cent. The question we wanted to answer is whether that kind of model can be the decision layer, with something cheaper than a frontier VLM doing the perception.
Two things to be clear about before the results. First, we did not get that far: in every experiment below the perception in front of Jev was still a frontier VLM, so the pipeline as a whole was not fast. What we could measure is whether Jev's decisions, given a description, are good enough and cheap enough to be worth building the fast version around. Second, an obvious alternative is to skip Jev and ask the VLM for the same six answers as structured output. We did not run that comparison, so nothing here shows Jev is more accurate than a VLM asked the same questions. The case for it is that it is roughly 50× faster and 100–1000× cheaper per decision, returns a probability from its output distribution rather than a number written into a sentence, and keeps the judging rubric fixed and separate from whatever is doing the seeing. Whether those properties matter in practice is the thing to test next.
What Jev is
Jev is a model that does not write. You give it a state — any structured or unstructured text — and one or more questions. Each question has a type, and Jev returns an answer of that type together with how sure it is. There are three question types — Noul (yes/no; "Noul" is simply Jev's name for a yes/no question), Choice (which one) and Score (how much). In robotics terms, the three map onto the questions a supervisor actually asks:
Three properties matter for what follows:
- Jev reads text, not images. In these experiments it never saw a frame. It sees whatever you put in the state — a description of the scene, a gripper encoder reading, the task instruction, the last few observations.
- Every answer is a probability rather than a sentence.
0.93is something a program can threshold, log, plot, or use as a reward. "The brick appears to be in the box" is not. (We did not formally measure how well-calibrated these probabilities are — more on that in the limits.) - It is fast and cheap. In our runs, one call — six questions asked at once — took 0.16–0.20 seconds and cost about four thousandths of a cent. That is fast enough for a high-level supervisory loop at a few hertz — not a low-level control or safety loop, and not a chat turn.
So what is it for?
Jev is a decision layer. Something else perceives; Jev turns the perception into typed, probabilistic decisions that code can act on immediately. That framing is the whole result of this project — and we did not start with it.
How we wired it
The same shape of pipeline runs through every experiment below, with protocol differences noted in each part. A vision-language model (VLM) watches frames and writes a short observation. That observation, plus the task and any sensor readings, becomes Jev's state. Jev answers a fixed set of typed questions. The probabilities feed a dashboard — and could equally feed a controller.
One caveat up front: everything here ran offline, on recorded episodes. Only the Jev stage is fast; the vision model in front of it takes 9–13 s per call, so the pipeline as tested is not yet an end-to-end real-time supervisor. What we measured is whether Jev's decisions, given a description, are good enough and fast enough to be one once perception is. We did not compare against asking the VLM itself for the same structured answers; see the limits section.
A few rules we held ourselves to, because it is easy to fool yourself in this kind of evaluation: the human reference label was only ever loaded in the scoring script, never in any prompt; the dataset's own model-generated check field was excluded; prompts were fixed on three practice clips from a separate split and never tuned on the 20 test clips; every number below is from real API calls.
Part 1Grading finished episodesIt didn't work — and the reason mattered
The obvious first question: if a vision model grades a robot episode, does adding Jev as a second opinion improve the grade?
We took 20 clips from RoboReward's human-verified test split, balanced four per reward level (1 = no success … 5 = perfect). Agent A was the VLM alone: watch ten frames, output a level. Agent B was the VLM's observations, evidence and uncertainty notes handed to Jev as a Choice over the same five-level rubric, with the VLM's proposed label included but an instruction not to defer to it.
Jev did not help, and the pattern of where it hurt is instructive. On clips the human rated 4 or 5, Agent A got 5 of 8 right; Agent B got 2. Whenever the VLM's notes hedged — "intermediate grasp details are between sampled frames" — Jev read the hedge as evidence against completion and downgraded. That is a reasonable thing for a judge to do with the text it was given. It is the wrong thing to do with these clips.
Why it didn't help
Look again at the blue-cup clip. The VLM's final observation was "blue cup upright on table, no longer on purple plate; orange cup remains on plate." Given only that text, "perfect completion" is the natural reading — and Jev agreed, with full confidence. We don't know exactly what the human rater saw — perhaps a drop, a re-grasp, or a plate that moved between our sampled frames — but whatever it was, it never reached the text, so it could never reach Jev. When we (one author, unblinded, against the full footage) went through Agent A's nine mistakes, we classified seven as observation-related — the description disagreed with what we could see, so every downstream judgment inherited the disagreement. Whether the human rater or the model is right about this particular clip we leave open; either way, Jev could only judge the text it was given.
Lesson
A text-only judge cannot fix a perception error. If the goal is a better final grade, the investment belongs in the eyes, not the judge. In a separate test we also asked Jev to act as a verification gate — "is the VLM's label justified by its own evidence?" — and it was no better than chance at catching the VLM's mistakes (AUROC 0.53 for "label not justified" against the 9 errors in 20 clips), for the same reason.
This could have been the end. Instead it made us ask what Jev is for, if not for grading.
Part 2Monitoring episodes step by step
Grading is a one-shot question asked at the end. Monitoring is a stream of small questions asked continuously — and those are exactly the questions that are cheap for Jev and expensive for everyone else.
We switched datasets to something closer to a real deployment: LeRobot's SO-101 pick-and-place episodes (task: pink lego brick into the transparent box; two cameras, six joint encoders at 30 fps). For every one-second window of footage, the VLM writes a short narration of what changed (three top-camera frames plus one side frame). Jev then answers six questions in one call:
- Choice
phase— reach / grasp / transport / place-release / retreat-idle / recovery - Score
progress— five ordered levels, returned as a position 0–4 plus a distribution - Nouls (yes/no)
brick_grasped,brick_in_box,anomaly,safe_to_continue
Here is the dashboard for episode 40, scrubbed from start to finish. Watch the anomaly line.
recovery. The arm re-approaches, and P(brick_in_box) crosses 0.9 at 8 s.| Episode | Duration | Jev steps | P(in box) ≥ 0.9 at | Max P(anomaly) | Note |
|---|---|---|---|---|---|
| ep008 | 7.5 s | 7 | 5.0 s | 0.07 | clean |
| ep017 | 8.1 s | 8 | 6.0 s | 0.22 | clean |
| ep037 | 7.6 s | 7 | 6.0 s | 0.08 | one low-confidence phase step |
| ep040 | 9.8 s | 9 | 8.0 s | 0.75 @ 3 s | missed grasp, flagged, recovered |
| ep045 | 8.3 s | 8 | 6.0 s | 0.41 | clean |
39 one-second steps; every Jev call returned successfully. Jev: 0.16 s per step, $0.0022 for all five episodes. VLM: 8.9 s per step. All five episodes are recorded successes; the monitor agreed by the end of each, and the one hiccup it flagged along the way is visible in the footage.
Two things are already different from Part 1. The signal is per step, so a problem is visible while there is still time to do something about it. And the outputs are typed: P(anomaly) > 0.7 is a line of code, not a judgment call about a paragraph.
Part 3Testing the monitor harder
Part 2 only had successes. To test the monitor properly we needed failures, so we ran three experiments designed to surface them. A word on what these can and cannot show: they establish that per-step typed judgments from Jev are feasible and cheap, not that Jev is a better streaming judge than a VLM asked the same questions at the same cadence — we did not run that matched baseline.
3.1 Alarms on real failures
We went back to the 20 RoboReward clips, but this time asked Jev after every observation — 163 calls in total — instead of one grade at the end. The six streaming questions: is the task done, is something going wrong, should a person look at this (three yes/no Nouls), plus progress, grasp and placement Scores. Jev never saw the human labels; we used them only afterwards to check whether the alarm went off on the clips that really failed.
How we scored an alarm
Alarm = P(anomaly) ≥ 0.5 at any step (threshold fixed at 0.5 in advance, not tuned; clips have 5–10 observations, so longer clips get more chances to alarm). Failure = human rating 1 or 2 (8 clips); success = 4 or 5 (8 clips); the four level-3 "partial" clips are reported separately. Alarm positions are measured in recording time — where in the clip the alarm fired — and exclude the pipeline's own processing delay (see 3.3).
needs_human stays above 0.83 throughout. Human label: 1 (no success). The VLM's one-shot label was 2.Read that last card carefully. On the final verdict, Jev is not better than the VLM — it is working from the same text, so it has the same blind spots, and on every endpoint metric it is a little behind:
| Endpoint metric (n = 20; Jev value = its final-step output) | Jev streaming | VLM final label |
|---|---|---|
| AUROC, success (human 4–5) — P(task_done) vs label | 0.83 | 0.88 |
| AUROC, failure (human 1–2) — P(anomaly) vs inverted label | 0.71 | 0.79 |
| Accuracy as a success classifier — P(done) ≥ 0.5 vs label ≥ 4 | 0.65 | 0.80 |
| Spearman with human rating — progress Score vs label | 0.68 | 0.75 |
With twenty clips these gaps are not statistically resolved: a clip-level bootstrap puts the success-AUROC difference at −0.13 to +0.03 (95%) and the failure-AUROC difference at −0.24 to +0.04. The honest summary is "not better, plausibly a little worse", not a ranking.
Its case is elsewhere: it produces a probability at every step, so problems surface part-way through instead of at the end; the anomaly probability separates failures from successes usefully, if imperfectly (means of 0.68 vs 0.19); and each of those 163 decisions took about 0.2 s and $0.00004. The per-clip picture:
The needs_human question deserves a footnote. It caught 8 of the VLM's 9 labeling errors, which sounds great, but it fired on 17 of 20 clips. As a "route to a person" signal it is far too eager at a 0.5 threshold. Useful, but not free.
3.2 Controlled incompletion and instruction-mismatch tests
Real datasets are mostly successes, so we constructed failure conditions from the same five SO-101 episodes — paired transformations, not fifteen independent rollouts — keeping all the VLM narration real: truncate the episode at 3 s (the arm is mid-reach, nothing is in the box), and swap the instruction (tell Jev the task was to put the brick next to the box, while the footage shows it going in). For the second we added a seventh question, task_alignment: is what's happening the task that was asked for?
This is the cleanest separation in the project. Be clear about what it shows: Jev recognises an unfinished state and a mismatched instruction from the narration. It does not yet show detection of naturally occurring physical failures at scale, and these scalars have not been used as a training reward. What it does show is the right shape for one: high when the goal state is reached, near zero when it isn't, and a second scalar that fires when the robot is doing the wrong task well.
3.3 Confidence-triggered re-checks (retrospective)
Because the phase Choice comes with a confidence, one could in principle spend the expensive model only when the cheap one is unsure. We tried the mechanics: if phase confidence is below 0.6, ask the VLM for a wider-window re-look (which, in this offline setting, included frames up to a second after the queried moment) and ask Jev again. On the SO-101 episodes this fired on 2 of 39 steps, both in the first second of an episode when "idle" and "starting to reach" look alike. On episode 37 the confidence went from 0.48 to 0.91; on episode 40 only from 0.48 to 0.53 — arguably right, since at t = 1 s the arm had not committed. Two caveats: higher confidence is not demonstrated correctness, and because every step in these experiments already used the VLM, the gate here only adds ~10–12 s re-checks. The saving only exists in a future setup where routine perception is cheap and the VLM is reserved for the uncertain steps.
Explore the data yourself
Everything above is logged in Rerun. The viewer below is the same one we used; it loads in your browser (about 60 MB the first time). Pick a clip or episode tab along the top, then scrub the timeline. Reference labels appear only in the panel marked eval layer only.


What it shares with our work is a loose design analogy rather than a mechanism: the same villain (big-model latency), a similar instinct (use intermediate information rather than waiting for the episode to end — our streaming Jev re-scores after every observation prefix, ARLI augments the RL state mid-inference), and the same pairing of a slow, expressive model with a fast, cheap one. Our prototype is sequential and offline; ARLI's is a genuinely asynchronous acting loop. What it does not set out to address: where the reward comes from. ARLI is a policy-learning paper; the success signal is an input to the method, not something it produces.
Real-Time EXPO-FT — RL for real-time VLA policies [2]
Reinforcement Learning for Real-Time Vision-Language-Action Policies. Dong, Hung, Sadigh, Finn · Stanford · September 2026 · arXiv 2609.18207 · project page · Interactive Real-Time EXPO-FT walkthrough on Engineermaxxing
Same latency problem, attacked with a slow/fast split on the control side: a large VLA proposes candidate action chunks (slow), a small edit policy adjusts them using the observation available at the moment of execution (fast), and a learned scorer (a Q-function) picks the best candidate. On four dynamic real-world tasks — ball balancing, soccer kicking, dynamic picking, mid-air object passing — the authors report success rising from 42% to 97% with about ten minutes of real robot data.
The paper is candid about its weak link. The reward is a sparse binary signal from a hand-written, rule-based success detector for each task, with a human verifying success during evaluation, and the authors list it as a limitation: designing a classifier per task is work, and which reward formulation performs best remains open.
Side by side
- Slow thing
- Generalist VLA (π0.5), 100s of ms per chunk
- Fast thing
- Low-latency RL policy steering the next chunk
- What the fast thing consumes
- Committed actions + a mid-inference observation
- Reward
- An input to the method; how it is produced is not the paper's subject
- Slow thing
- VLA proposing action chunks
- Fast thing
- Edit policy + Q-function choosing among chunks
- What the fast thing consumes
- The latest observation at execution time
- Reward
- Hand-written per-task success detector; listed as a limitation
- Slow thing
- VLM narrating the scene, 9–13 s
- Fast thing
- Jev, 0.2 s per six typed questions
- What the fast thing consumes
- Typed state: task, observations, encoders
- Reward
- Candidate signals — P(task_done), P(anomaly), progress Score, task_alignment — per step; not yet used for RL
The pattern rhymes across the three columns: a slow, expressive model that cannot run at loop rate, and a fast, cheap layer that consumes intermediate information. ARLI and Real-Time EXPO-FT build that layer on the control side and take the reward as given — in EXPO-FT's case, explicitly, from a hand-written detector per task; ARLI does not discuss reward generation, so we make no claim about it. Our monitor is an offline prototype of the analogous layer on the supervision side. It is not a competing method and not yet an integrated one; it is a complementary direction, and the integration is the experiment we have not run.
One line
The fast loop for acting now exists. A fast judge for it is worth testing.
Limits and next steps
- This is a small pilot: 20 clips (balanced by rating, so not the natural failure rate), 5 episodes, 163 streaming decisions. The constructed-failure separations are large and consistent, but the AUROC estimates have wide error bars at n = 20 (Jev success AUROC 95% CI 0.61–0.99).
- We did not measure calibration. Jev returns probabilities, and they rank successes above failures reasonably well, but we did not check whether "0.9" is right nine times in ten (no reliability curves, no Brier score). Anyone using these as rewards should.
- Perception is still the bottleneck. Jev is a decision layer on top of whatever describes the scene. Where the VLM misread the scene, Jev was confidently wrong along with it. The obvious next experiment is to give Jev cheaper, more frequent perception — a small vision model at 5–10 Hz, or the robot's own state estimates — instead of a frontier VLM every few seconds.
- Everything ran offline, and the full pipeline is not yet fast enough for control. We ran on recorded episodes, and the vision model in front of Jev takes 9–13 s per call. Nothing here has yet gated a live controller or been used as an RL reward.
- There is no matched baseline, and this is the biggest gap. We compared Jev's streaming answers with the VLM's one-shot final grade, not with the VLM asked the same six questions in structured output at the same cadence. Since the VLM was already running every step, that comparison would have cost nothing extra to run. Until it's done, the 6-of-8 alarm result is evidence that per-step monitoring works, not that Jev is the reason.
- The
needs_humanquestion fires too often: 17 of 20 clips were flagged at a 0.5 threshold. The question needs sharper criteria or a higher threshold before it is useful for routing to people. - The prompts were written once and never tuned. The rubric language, the "must be observed, not inferred" rules, and the phase boundaries were all written once on three setup clips. We did not iterate on them; better ones surely exist.
What we would do next, in order: first, put fast perception in front of Jev (a small vision model at 5–10 Hz, or the robot's own state estimates) and measure causal alarms and false-alarm rates on live or replayed streams; second, measure calibration and check that the questions transfer across tasks without rewriting; third — only then — plug P(task_done) and P(anomaly) into an RL fine-tuning loop like ARLI's or Real-Time EXPO-FT's as the per-step reward and stopping signal, against a hand-written detector, and measure both learning performance and reward exploitation.
So where does Jev actually help? Not in the setup we ran. With a frontier VLM in the loop, the VLM is 98% of the latency and nearly all of the cost, and it could plausibly have answered the six questions itself. Jev helps if two other things are true, neither of which we tested: that perception can be made cheap (a small detector, the robot's own state, or a tiny VLM at a few hertz), and that a native probability from a fixed rubric is a better reward signal than a number an LLM writes into its answer. If both hold, you get a supervisor that can run on every step of every episode for a training run at a cost that rounds to zero. If they don't, a VLM with a JSON schema is the simpler tool. This pilot establishes that the decision layer is fast, cheap and produces the right shape of signal; it does not yet establish that you need it.
Appendix
Setup and data
- RoboReward (
teetone/RoboReward, revision469b9af76e8539e8d2ac553b081307738fd92ca5), test split, human-verified. 2,831 rows; we sampled 20 balanced by reward level with seed 20260920, plus 3 setup clips from validation for prompt design. 10 frames per clip, evenly spaced by index, ≤ 512 px. Thegpt5_mini_checkfield was excluded. Filenames (which encode the score) were replaced with SHA-1 hashes. - LeRobot SO-101 pick-place (
lerobot/svla_so101_pickplace, v3.0, revisionf641879e), 5 of 50 episodes, 30 fps, camerasupandside, six joints. VLM window 1 s (three top frames + one side frame). - VLM: GPT-5.6 Sol via the PromptQL program runtime, structured JSON output. Jev:
jev-1.13.0viaPOST /v1/systemone, one fan-out call per step. Pricing basis $0.042 per million input tokens. - Logging: Rerun SDK 0.38.1; recordings served with the Rerun web viewer.
- Latency figures: Jev 0.16 s/step (Part 2), 0.20 s/step (Part 3), 0.12–0.16 s (Part 1); VLM 13.3 s (Part 1, ten frames), 8.9 s (Part 2, four frames). VLM cost was not separately metered, which is why we quote only Jev's cost.
Jev question set (streaming monitor)
# state (abbreviated) — what Jev actually reads
{ "task": "Remove the blue cup from the plate",
"observations_so_far": [ {"t": 0.0, "event": "Blue cup stands on purple plate…"},
{"t": 22.3, "event": "Gripper grasps blue cup on the plate."} ],
"rubric": ["No success", "Minimal progress", "Partial", "Near completion", "Perfect completion"] }
# questions — one call, six typed answers
task_done : noul "Based ONLY on the observations so far, has the task been fully completed?
The final goal state must be observed, not inferred."
anomaly : noul "Is something going wrong that a supervisor should know about?"
needs_human : noul "Should a person review this before trusting an automatic label?"
progress : score 5 ordered levels (the rubric), returns position + distribution
grasp : score 0–3, has the target object been secured?
placement : score 0–3, has it reached the goal location?
The SO-101 monitor swaps in phase (Choice over six phases with explicit "means / does not mean" boundaries), progress (Score, five levels) and four Nouls (brick_grasped, brick_in_box, anomaly, safe_to_continue); the injected-failure test adds task_alignment. The full definitions are in the source bundle.
All the numbers
| Experiment | Metric | Value |
|---|---|---|
| Part 1 · A/B grading (n = 20) | Agent A exact / MAE (lower is better) / latency | 11/20 · 0.60 · 13.3 s |
| Agent B exact / MAE (lower is better) / added Jev latency | 10/20 · 0.75 · 0.16 s | |
| Corrections / regressions A→B | 2 / 3 | |
| Agent A errors that were visual misreads | 7 / 9 | |
| Jev tokens / cost | 19,156 · $0.0008 | |
| Part 2 · SO-101 monitor (5 eps) | Steps / API call errors | 39 / 0 |
| First P(in box) ≥ 0.9 | 5–8 s | |
| Jev / VLM latency per step | 0.16 s / 8.9 s | |
| Jev tokens / cost | 53,215 · $0.0022 | |
| Part 3.1 · streaming (20 clips) | Jev calls / latency / cost | 163 · 0.203 s · $0.0064 |
| Failures alarmed, P(anomaly) ≥ 0.5 (mean recording position) | 6/8 (0.53) | |
| Successes / partial clips that alarmed | 1/8 · 4/4 | |
| Mean P(anomaly) failures / successes | 0.68 / 0.19 | |
| AUROC success — P(task_done) / VLM label | 0.828 / 0.875 | |
| AUROC failure — P(anomaly) / VLM label | 0.708 / 0.792 | |
| Spearman vs reference — progress / VLM label | 0.681 / 0.75 | |
| needs_human ≥ 0.5: flagged / VLM errors caught | 17/20 · 8/9 | |
| Successes with P(done) ≥ 0.5 before end | 2 | |
| Part 3.2 · injected (5 eps × 3) | P(task_done) nominal / truncated / swapped | 0.92 / 0.02 / 0.02 |
| P(task_alignment) | 0.97 / 0.62 / 0.05 | |
| P(anomaly) | 0.05 / 0.36 / 0.71 | |
| Part 3.3 · confidence gate | Steps re-checked | 2 / 39 |
| Phase confidence after re-check | ep037 0.48→0.91 · ep040 0.48→0.53 |
Reproducibility
All code (data prep, VLM and Jev pipelines, evaluation, Rerun builders), raw JSON outputs, the three .rrd recordings and this page are in the project source bundle. Jev prompt versions: jev-choice-v1 (Part 1), lr-jev-v1 (Part 2), jev-stream-v1 (Part 3). VLM prompt versions vlm-v1, lr-vlm-v1. Total Jev spend across everything: about one cent.
References
- B. Zhu, M. Khalil, E. Harrison, E. Poggi, P. Schmitt, B. Kast, P. Meister, P. Atreya, Q. Li, F. Ferchau, C. Colmenero, Y. Shahapurkar, G. Narayanan, M. Erdogan, K. Wurm, G. von Wichert, O. Mees, E. Solowjow, A. Wagenmaker and S. Levine. Learning to Act While Waiting: RL Finetuning of Generalist Robot Policies Under Inference Latency. arXiv:2608.23831, August 2026 (submitted to CoRL 2026). arxiv.org/abs/2608.23831 · project page.
- P. Dong, K.-H. Hung, D. Sadigh and C. Finn. Reinforcement Learning for Real-Time Vision-Language-Action Policies. arXiv:2609.18207, September 2026. arxiv.org/abs/2609.18207 · project page.