Kavli Affiliate: Wei Gao| Summary:Open-vocabulary 3D scene understanding is commonly achieved by embedding 2D vision-language features such as CLIP into a 3D Gaussian Splatting scene, turning it into a text-queryable semantic field. However, attaching a high-dimensional feature to each of millions of Gaussians inflates a single scene to gigabytes, which makes storage and deployment the […]
Continue.. CoSAG: Compact Semantic Anchor Gaussians via Training-Free Rate-Distortion Coding