EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models

May 25, 2026 · Research Publication · Stanford University · Development Platform

Introduces RL finetuning for pretrained VLA policies using online interaction, Q-guided sampling, residual edit policies, and human corrections.

Verified eventStanford UniversityMay 25, 2026
  • Reports 30/30 success across evaluated real-robot manipulation tasks with about 19 minutes of online robot data per task.
  • Project page and arXiv paper are live; code is explicitly marked coming soon, so this is tracked as a project-page/paper artifact for now.

Research university; ToddlerBot humanoid research uses ROBOTIS DYNAMIXEL actuators per public bundle and paper materials.