Skip to content

Repository files navigation

What Are We Actually Benchmarking in Robot Manipulation?

CoRL 2026 (to appear)
IROS 2026 RGMCW workshop

Paper · Project website · Datasets and checkpoints

TL;DR: High scores on commonly used benchmarks (LIBERO, SimplerEnv) are not proof of broader manipulation capability.

Findings

  • LIBERO: A 0.09B probe nears top scores without language encoding or robotics pretraining.
  • Significance: Only 19.8% / 19.7% of LIBERO / SimplerEnv gains are provably significant.
  • CALVIN: Resampling block poses within the training range lowers performance for every tested policy.
  • SimplerEnv: 22M policies trained near the test reach 94.8%, versus 95.8% for 0.9B X-VLA.
  • Most-reported benchmarks fail more diagnostics: LIBERO, CALVIN, and SimplerEnv fare worse than RoboCasa and RoboTwin 2.0.

Four diagnostics

  • Shortcut solvability: Can a policy reach a high score without the capabilities that score is taken to demonstrate?
  • Statistical significance: Does the reported evidence show that an improvement exceeds what evaluation noise could explain?
  • Creeping overfitting: Have policies become too tuned to a benchmark's narrow test conditions or its particular test examples?
  • Data-source dependence: Does a high score reflect generalization from different training conditions, or training data collected close to the test conditions?

What's included

Contact

About

Code and diagnostics for the manipulation benchmark audit paper.

Resources

Stars

19 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages