Kavli Affiliate: Zheng Zhu | First 5 Authors: Angen Ye, Angen Ye, , , | Summary: Vision-Language-Action (VLA) models aim to unify perception, language understanding, and action generation, offering strong cross-task and cross-scene generalization with broad impact on embodied AI. However, current VLA models often lack explicit step-by-step reasoning, instead emitting final actions without considering […]
Continue.. VLA-R1: Enhancing Reasoning in Vision-Language-Action Models