CoRL 2026 (to appear)
IROS 2026 RGMCW workshop
Paper · Project website · Datasets and checkpoints
TL;DR: High scores on commonly used benchmarks (LIBERO, SimplerEnv) are not proof of broader manipulation capability.
- LIBERO: A 0.09B probe nears top scores without language encoding or robotics pretraining.
- Significance: Only 19.8% / 19.7% of LIBERO / SimplerEnv gains are provably significant.
- CALVIN: Resampling block poses within the training range lowers performance for every tested policy.
- SimplerEnv: 22M policies trained near the test reach 94.8%, versus 95.8% for 0.9B X-VLA.
- Most-reported benchmarks fail more diagnostics: LIBERO, CALVIN, and SimplerEnv fare worse than RoboCasa and RoboTwin 2.0.
- Shortcut solvability: Can a policy reach a high score without the capabilities that score is taken to demonstrate?
- Statistical significance: Does the reported evidence show that an improvement exceeds what evaluation noise could explain?
- Creeping overfitting: Have policies become too tuned to a benchmark's narrow test conditions or its particular test examples?
- Data-source dependence: Does a high score reflect generalization from different training conditions, or training data collected close to the test conditions?
- Shortcut solvability: LIBERO/CALVIN training, evaluation, and results.
- Statistical significance: shared-test outcomes and leaderboard comparisons.
- Creeping overfitting: results, custom settings, assets, and starting states.
- Data-source dependence: scripted WidowX collection, training, evaluation, and results.
- Analysis: regenerate selected paper figures and tables on CPU.
- Google Drive: demonstration datasets, selected model weights, and evaluation inputs.
- Reproduction details: setup, validation, and exclusions.
- Claim-to-artifact guide: evidence behind the results.