RealBench: Standardized Real-World Benchmark for Robot Manipulation Models

Neurips 2025

RoboBench is a comprehensive benchmark to evaluate embodied multi-modal language models (MLLMs) with Q&A framework spanning 4 levels and over 25 subcategories. It emulates the execution process of embodied tasks: starting with the comprehension of instructions, followed by scene perception, strategic planning, and culminating in summarization, reflection, and iterative improvement (top left). The primary metric is further divided into subcategories, demonstrating both the diversity and complexity of RoboBench (top right; bottom). Tasks are color-coded to aid in distinguishing between task types.

Abstract

Recently, Multi-modal Large Language Models (MLLMs) have shown great potential to be the “brain” of embodied agents. However, existing benchmarks fail to evaluate the full capabilities essential for embodied tasks in MLLMs, including instruction comprehension, reasoning in complex environments, decision-making, affordance reasoning, and reflection. To bridge these gaps, we introduce RoboBench, a novel benchmark designed to systematically evaluate the capabilities of MLLMs as the brain of embodied agents. The evaluation formulation is structured around 5 key skills and 20 sub-dimensions, featuring specially designed tasks such as cross-embodiment planning, cross-object perception, and cross-view reasoning to ensure a fine-grained and realistic assessment. To assess the robustness of task planning, we introduce a DAG-based LLM evaluation framework that incorporates variations in execution order and task granularity. Through large-scale evaluations of closed-source and open-source MLLMs, we identify several critical limitations of previous benchmarks. Specifically, we observe that current MLLMs struggle with understanding implicit demand instructions, cross-embodiment planning, object interaction characteristic analysis, and fine-grained error reflection. RoboBench aims to drive progress in the field by guiding the development of next-generation MLLMs with enhanced embodied capabilities. Our code and dataset will be released soon.

Leaderboard

Dataset Construction Pipeline

RoboBench consolidates datasets from over 20 recent studies. These datasets have been processed through our three-category conversion workflow, which is color-coded as follows: green for perception, blue for planning, and red for reflection. The processing pipeline generally involves preprocessing (employing quality filtering or basic binary classification), followed by steps with a Vision-Language Model (VLM), detection model, or human experts, culminating data standardization into metadata summaries. These metadata are then utilized to automate Q&A questions across various categories. The final Q&A format comprises standardized binary classification, multiple choice, and multi-step multiple choice formats, making it adaptable for both open-source and closed-source MLLM models.

Poster

BibTeX

@misc{luo2025robobench,
        title={Robobench: A Comprehensive Evaluation Benchmark for Perception, Planning and Reflection of Embodied Multimodal Large Language Models}, 
        author={Yulin Luo, Chun-Kai Fan, Menghang Dong, Mengdi Zhao, Bo-Wen Zhang, Jiayu Shi, Jiaming Liu, Gaole Dai, Rongyu Zhang, Ruichuan An, Kun Wu, Zhengping Che, Pengwei Wang, Guang Liu, Zhongyuan Wang, Tiejun Huang, Shanghang Zhang},
        year={2025},
        eprint={xxx},
        archivePrefix={arXiv},
        primaryClass={xxx},
        url={https://arxiv.org/abs/xxx}, 
  }