RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies

Abstract

Memory is critical for long-horizon and history-dependent robotic manipulation. Such tasks often involve counting repeated actions or manipulating objects that become temporarily occluded. Recent vision-language-action (VLA) models have begun to incorporate memory mechanisms; however, their evaluations remain confined to narrow, non-standardized settings. This limits their systematic understanding, comparison, and progress measurement. To address these challenges, we introduce RoboMME: a large-scale standardized benchmark for evaluating and advancing VLA models in long-horizon, history-dependent scenarios. Our benchmark comprises 16 manipulation tasks constructed under a carefully designed taxonomy that evaluates temporal, spatial, object, and procedural memory.

Publication
ICML (Oral)
Yinpei Dai
Yinpei Dai
Ph.D. Candidate
Hongze Fu
Hongze Fu
Graduate Research Assistant
Jayjun Lee
Jayjun Lee
Graduate Research Assistant