PRIMO R1 turns a video MLLM from a passive Observer into an active Critic. Outcome-based reinforcement learning incentivizes explicit Chain-of-Thought for progress estimation, reducing the mean absolute error of specialized baselines by 50%.
A two-stage post-training corpus with Chain-of-Thought annotations: a 116k-sample SFT set and a 182k-sample RL set, paired with PRIMO Bench to evaluate in-domain and out-of-domain generalization across cross-task and cross-environment settings.
Anchoring the video sequence between initial and current state images turns generic temporal perception into structured state-alignment verification, and yields strong zero-shot failure detection: a state-of-the-art 67.0% on RoboFail.
Accurate process supervision remains a critical challenge for long-horizon robotic manipulation. A primary bottleneck is that current video MLLMs, trained primarily under a Supervised Fine-Tuning (SFT) paradigm, function as passive “Observers” that recognize ongoing events rather than evaluating the current state relative to the final task goal. In this paper, we introduce PRIMO R1 (Process Reasoning Induced MOnitoring), a 7B framework that transforms video MLLMs into active “Critics”. We leverage outcome-based Reinforcement Learning to incentivize explicit Chain-of-Thought generation for progress estimation. Furthermore, our architecture constructs a structured temporal input by explicitly anchoring the video sequence between initial and current state images. Supported by the proposed PRIMO Dataset and Benchmark, extensive experiments across diverse in-domain environments and out-of-domain real-world humanoid scenarios demonstrate that PRIMO R1 achieves state-of-the-art performance. Quantitatively, our 7B model achieves a 50% reduction in the mean absolute error of specialized reasoning baselines, demonstrating significant relative accuracy improvements over 72B-scale general MLLMs. Furthermore, PRIMO R1 exhibits strong zero-shot generalization on difficult failure detection tasks. We establish state-of-the-art performance on the RoboFail benchmark with 67.0% accuracy, surpassing closed-source models like OpenAI o1 by 6.0%.
We formalize robotic process supervision as a multi-modal state estimation problem. Given an initial state image Iinit, a process video sequence Vseq, a current state image Icurr, and a language instruction ℬ specifying the task goal, the model outputs a scalar progress indicator y ∈ [0, 100], where y=0 denotes the initial state and y=100 signifies success. In the standard paradigm, video MLLMs act as passive Observers, treating progress estimation as direct regression or classification through SFT, which isolates visual features at a surface level and bypasses the causal structure of state transitions.
To build an active Critic, we reformulate prediction from direct scalar regression into a multi-step generative reasoning task. A policy πθ sequentially generates a latent reasoning chain (Chain-of-Thought) followed by the final progress estimate. Rather than supervising the intermediate reasoning with dense annotations, we optimize the policy with Reinforcement Learning against a reward computed solely from the accuracy of the final prediction. Conditioning this reasoning on diverse natural language task goals connects the semantic objective to the visual execution logic, exploiting the linguistic generalization of foundational MLLMs.
We employ Group Relative Policy Optimization (GRPO). For each task tuple, we sample a group of
outputs and estimate the advantage by normalizing each reward against the group distribution,
avoiding a separate value network whose memory overhead is prohibitive for video MLLMs. A
composite rule-based reward combines a format reward, which enforces the
<think>reasoning</think><answer>prediction</answer>
structure and prevents collapse into direct guessing, with a bounded linear-decay
accuracy reward that provides dense feedback for numerical reasoning. A KL penalty to
the reference policy guards against reward hacking and language degeneration. PRIMO R1 thereby
learns that detailed causal reasoning is the most reliable strategy for maximizing the accuracy
reward, emerging as a robust Critic.
Across four environments, PRIMO R1 attains the highest average Mean Relative Accuracy (MRA 82.90) and the lowest average Mean Absolute Error (MAE 15.52). Despite a 7B foundation, it surpasses Qwen2.5-VL-72B by +9.10 absolute MRA points and roughly halves the error of specialized reasoning and video MLLMs.
| Model | AgiBot | Behavior | RoboTwin | Real Humanoid | Average | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| MRA ↑ | MAE ↓ | MRA ↑ | MAE ↓ | MRA ↑ | MAE ↓ | MRA ↑ | MAE ↓ | MRA ↑ | MAE ↓ | |
| Closed-Source Models | ||||||||||
| GPT-5 mini | 74.27 | 24.81 | 79.60 | 20.08 | 80.52 | 18.34 | 67.14 | 32.59 | 75.38 | 23.96 |
| GPT-4o | 81.01 | 18.99 | 79.92 | 20.08 | 81.73 | 18.27 | 74.65 | 25.35 | 79.33 | 20.67 |
| Gemini 2.5 Flash | 73.54 | 26.41 | 78.01 | 21.99 | 81.04 | 18.58 | 67.37 | 32.63 | 74.99 | 24.90 |
| Claude-Haiku-4.5 | 74.40 | 25.59 | 70.93 | 29.07 | 74.13 | 25.87 | 72.68 | 27.32 | 73.04 | 26.96 |
| Open-Source General MLLMs | ||||||||||
| Qwen2.5-VL-7B | 77.43 | 22.56 | 69.91 | 30.06 | 67.37 | 32.62 | 56.46 | 34.73 | 67.79 | 29.99 |
| InternVL 3.5 8B | 78.52 | 21.47 | 72.19 | 27.15 | 70.81 | 29.18 | 65.44 | 34.55 | 71.74 | 28.09 |
| Qwen2.5-VL-72B | 78.00 | 22.00 | 79.49 | 20.50 | 75.41 | 24.59 | 62.29 | 28.10 | 73.80 | 23.80 |
| Reasoning & Video MLLMs | ||||||||||
| ProgressLM-3B-RL | 42.90 | 30.83 | 46.43 | 28.23 | 31.84 | 36.48 | 32.00 | 34.90 | 38.29 | 32.61 |
| Video R1 7B | 72.42 | 27.58 | 70.63 | 29.37 | 70.86 | 29.14 | 53.57 | 31.87 | 66.87 | 29.49 |
| Robobrain 7B | 72.99 | 25.91 | 72.52 | 26.97 | 70.41 | 28.85 | 55.83 | 28.51 | 67.94 | 27.56 |
| Cosmos-Reasoning 7B | 72.48 | 27.01 | 67.06 | 32.35 | 73.14 | 25.85 | 59.39 | 31.41 | 66.52 | 29.12 |
| Specialized Progress & Reward Models (own input configuration) | ||||||||||
| VLAC | 84.92 | 15.08 | 76.86 | 23.14 | 67.04 | 32.96 | 70.77 | 29.23 | 74.90 | 25.10 |
| GVL (Gemini 2.5 Flash) | 56.81 | 35.84 | 58.11 | 32.31 | 61.18 | 27.91 | 48.77 | 38.37 | 56.22 | 33.61 |
| Robo-Dopamine-3B | 81.18 | 18.82 | 57.78 | 42.22 | 82.96 | 17.04 | 68.43 | 31.57 | 72.59 | 27.41 |
| Robo-Dopamine-4B | 83.88 | 16.12 | 65.30 | 34.70 | 67.22 | 32.78 | 73.89 | 26.11 | 72.57 | 27.43 |
| ProgressLM | 81.13 | 18.79 | 80.89 | 15.97 | 67.03 | 32.97 | 84.24 | 15.76 | 78.32 | 20.87 |
| PRIMO R1 (Ours) | 87.67 | 12.33 | 87.08 | 12.90 | 84.52 | 15.48 | 72.32 | 21.37 | 82.90 | 15.52 |
MRA ↑ higher is better; MAE ↓ lower is better. Best results in bold. Specialized progress and reward models (VLAC, GVL, Robo-Dopamine, ProgressLM) are evaluated under their own recommended input configurations and are therefore not strictly comparable cell-for-cell with models run under our unified protocol; the ProgressLM-3B-RL row in the Reasoning & Video group is the same model re-evaluated under our unified setting.
Optimizing a policy for continuous progress reasoning implicitly constructs the temporal context needed for discrete failure detection. On the entirely unseen RoboFail benchmark, PRIMO R1 reaches a state-of-the-art 67.0% accuracy — matching closed-source Gemini 2.0 Flash and surpassing GPT-4o (63.0%), OpenAI o1 (61.0%), and Cosmos-Reason1-56B (66.2%). SFT alone regresses to 51.0% through format overfitting; GRPO recovers genuine reasoning and lifts accuracy to 67.0%.
| Closed-Source | RoboFail ↑ | Open-Source | RoboFail ↑ | Ours | RoboFail ↑ |
|---|---|---|---|---|---|
| Gemini 2.0 Flash | 67.0 | Qwen2.5-VL-7B | 57.6 | PRIMO (SFT) | 51.0 |
| GPT-4o | 63.0 | Nemotron-H-56B | 64.0 | PRIMO (RL) | 63.0 |
| OpenAI o1 | 61.0 | Cosmos-Reason1-7B | 60.0 | PRIMO R1 | 67.0 |
| Claude-haiku-4.5 | 59.0 | Cosmos-Reason1-56B | 66.2 |
Accuracy (%) on the RoboFail benchmark, measuring the capability to detect and quantify task execution failures. Best result in bold.
SFT improves overall performance but primarily overfits the training distribution; RL-only struggles to discover the correct output format from scratch. The full SFT+RL pipeline creates a synergy that pushes in-domain accuracy near 90% and drastically improves out-of-domain generalization, lifting cross-environment (Real Humanoid) accuracy to 72.32%.
| Model | In-Domain — Seen Tasks | OOD — Cross-Task | OOD — Cross-Env | Avg. | ||||
|---|---|---|---|---|---|---|---|---|
| AgiBot | Behavior | RoboTwin | AgiBot | Behavior | RoboTwin | Real Humanoid | ||
| Qwen2.5-VL-7B (Base) | 70.83 | 69.13 | 71.19 | 74.45 | 77.47 | 61.01 | 48.12 | 67.46 |
| Our Model (SFT) | 83.37 | 80.38 | 80.63 | 82.02 | 82.61 | 79.13 | 67.30 | 79.35 |
| Our Model (RL) | 86.05 | 85.82 | 73.27 | 82.95 | 81.39 | 75.43 | 52.12 | 76.72 |
| PRIMO R1 (SFT+RL) | 87.83 | 89.42 | 88.15 | 87.67 | 87.08 | 84.52 | 72.32 | 85.28 |
All metrics are Mean Relative Accuracy (MRA ↑). Best in bold.
Varying the input modalities shows that temporal context is essential. Relying on the current state alone yields the highest error (Avg. MAE 59.50). The final PRIMO R1 configuration integrates all three signals — Iinit, Vseq, and Icurr — which is necessary for tracking progress over long horizons: on the long-horizon Behavior dataset it reduces MAE to 22.73 and raises Acc@10 to 31.83.
| Input Modality | AgiBot | Behavior | RoboTwin | Avg | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Iinit | Vseq | Icurr | MAE ↓ | Acc@10 ↑ | MAE ↓ | Acc@10 ↑ | MAE ↓ | Acc@10 ↑ | MAE ↓ | Acc@10 ↑ |
| ✓ | 59.64 | 0.00 | 51.91 | 8.17 | 66.94 | 0.77 | 59.50 | 2.98 | ||
| ✓ | ✓ | 43.97 | 18.52 | 49.59 | 9.73 | 45.93 | 11.45 | 46.50 | 13.23 | |
| ✓ | 27.58 | 31.34 | 34.85 | 18.33 | 47.41 | 8.07 | 36.61 | 19.25 | ||
| ✓ | ✓ | 25.04 | 35.29 | 27.59 | 29.21 | 40.24 | 17.37 | 30.96 | 27.29 | |
| ✓ | ✓ | 24.94 | 33.98 | 32.55 | 23.01 | 45.37 | 11.97 | 34.29 | 22.99 | |
| ✓ | ✓ | ✓ | 29.39 | 27.15 | 22.73 | 31.83 | 42.16 | 27.69 | 31.43 | 28.89 |
Iinit: initial state image; Vseq: process video clip; Icurr: current state image. Best in bold.
The PRIMO Dataset supports a two-stage post-training paradigm. Unlike standard video QA datasets, it features fine-grained progress indicators annotated with Chain-of-Thought reasoning paths, aggregating trajectories from a real-world environment (AgiBot) and two high-fidelity simulations (BEHAVIOR-1k and RoboTwin), augmented with general video reasoning data. It is partitioned into a 116k-sample SFT set (PRIMO-R1-CoT-116k) and a 182k-sample RL set (PRIMO-R1-182k).
The accompanying PRIMO Bench evaluates robustness under distribution shift via two splits: an In-Domain (Same Task) split over seen task categories in the three training environments, and an Out-of-Domain split covering both Cross-Task (unseen tasks in familiar environments) and Cross-Environment (real-world trajectories teleoperated on a different physical humanoid, Leju KUAVO-MY, in unstructured factory and service scenarios) generalization.
Given the pace of AI research these days, it is extremely challenging to keep up with all of the work around robot task progress estimation using foundation models. We list below a few key approaches developed concurrently with PRIMO R1. VanceChu's Awesome-Progress-Reward-Model paper list is also a useful entry point for understanding this direction.
@article{liu2026passive,
title={From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation},
author={Liu, Yibin and Lyu, Yaxing and Gao, Daqi and Liang, Zhixuan and Tang, Weiliang and Mu, Shilong and Yang, Xiaokang and Mu, Yao},
journal={arXiv preprint arXiv:2603.15600},
year={2026}
}