Recently, Multi-modal Large Language Models (MLLMs) have shown great potential to be the “brain” of embodied agents. However, existing benchmarks fail to evaluate the full capabilities essential for embodied tasks in MLLMs, including instruction comprehension, reasoning in complex environments, decision-making, affordance reasoning, and reflection. To bridge these gaps, we introduce RoboBench, a novel benchmark designed to systematically evaluate the capabilities of MLLMs as the brain of embodied agents. The evaluation formulation is structured around 5 key skills and 20 sub-dimensions, featuring specially designed tasks such as cross-embodiment planning, cross-object perception, and cross-view reasoning to ensure a fine-grained and realistic assessment. To assess the robustness of task planning, we introduce a DAG-based LLM evaluation framework that incorporates variations in execution order and task granularity. Through large-scale evaluations of closed-source and open-source MLLMs, we identify several critical limitations of previous benchmarks. Specifically, we observe that current MLLMs struggle with understanding implicit demand instructions, cross-embodiment planning, object interaction characteristic analysis, and fine-grained error reflection. RoboBench aims to drive progress in the field by guiding the development of next-generation MLLMs with enhanced embodied capabilities. Our code and dataset will be released soon.
@misc{luo2025robobench,
title={Robobench: A Comprehensive Evaluation Benchmark for Perception, Planning and Reflection of Embodied Multimodal Large Language Models},
author={Yulin Luo, Chun-Kai Fan, Menghang Dong, Mengdi Zhao, Bo-Wen Zhang, Jiayu Shi, Jiaming Liu, Gaole Dai, Rongyu Zhang, Ruichuan An, Kun Wu, Zhengping Che, Pengwei Wang, Guang Liu, Zhongyuan Wang, Tiejun Huang, Shanghang Zhang},
year={2025},
eprint={xxx},
archivePrefix={arXiv},
primaryClass={xxx},
url={https://arxiv.org/abs/xxx},
}