TaskAnchor: Grounding Task State in Reactive VLAs for Long-Horizon Manipulation

Kavli Affiliate: Wei Gao
| Summary:
Reactive vision–language–action (VLA) models struggle with long-horizon manipulation when visually similar observations can correspond to different actions depending on the task stage or interaction history. We refer to this ambiguity as task-state aliasing and introduce TaskAnchor, a lightweight adapter that grounds pretrained VLAs in execution history. TaskAnchor combines history-conditioned visual refinement with a milestone-supervised task-state coordinate, a scalar representing the semantic stage of execution. These signals are injected through the native visual and language interfaces, respectively, without introducing an explicit planner or modifying the action-generation mechanism. On RMBench, TaskAnchor achieves approximately 4.9–5.5$times$ the average success rates of the published $π_0.5$ and X-VLA baselines, with consistent gains on RoboMemArena and real robots. The added latency is only 2.08,ms per action chunk for $π_0.5$.
| Search Query:arXiv Query: search_query=au:”Gao Wei”&id_list=&start=0&max_results=10
Read More