EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models
Introduces RL finetuning for pretrained VLA policies using online interaction, Q-guided sampling, residual edit policies, and human corrections.
Verified eventStanford UniversityMay 25, 2026
Evidence notes
- Reports 30/30 success across evaluated real-robot manipulation tasks with about 19 minutes of online robot data per task.
- Project page and arXiv paper are live; code is explicitly marked coming soon, so this is tracked as a project-page/paper artifact for now.
Company context
Research university; ToddlerBot humanoid research uses ROBOTIS DYNAMIXEL actuators per public bundle and paper materials.