Kavli Affiliate: Wei Gao| Summary:End-to-end manipulation policies, combined with web-scale pretrained Vision-Language Models (VLMs), show the promise for generalizable and dexterous robotic manipulation. However, they inherit two key limitations from 2D foundation models: 1) the reliance on 2D RGB inputs that ignores the intrinsically 3D nature of manipulation; and 2) the lack of spatial 3D […]
Continue.. Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning