Correct Me If I’m Wrong: Language Guided Exploration for VLAs

Published in Reinforcement Learning for Vision-Language-Action Models Workshop, 2026

Reinforcement learning (RL) is a promising framework for improving vision language action (VLA) models through interaction, but it often struggles when the initial policy repeatedly fails in the same way. Our key observation is that these models already offer a simple way for humans to change what the policy attempts through language. In real world robot training, humans often need to stay nearby to monitor rollouts, reset environments, and check progress. This makes prompt updates a natural and lightweight source of guidance. While watching the robot, a human can redirect exploration by simply revising the instruction, without providing low level actions or directly controlling the robot. We introduce language-guided exploration, an interactive post-training framework that uses human prompt updates to help VLAs collect more useful experience. On real world physical robots, our method significantly improves sample efficiency over standard RL, making robot learning more practical under limited data regime. We further show that interactive post training makes VLAs more responsive to prompt updates at test time, enabling users to guide robot behavior through natural language during execution.

Recommended citation: Sunshine Jiang, John Marangola, David Zhang, Nitish Dashora, Pulkit Agrawal, Zhang-Wei Hong. (2026). "Correct Me If I’m Wrong: Language Guided Exploration for VLAs." Reinforcement Learning for Vision-Language-Action Models Workshop.
Download Paper