Kavli Affiliate: Wei Gao| Summary:In the past year, video-based large language models (Video LLMs) have achieved impressive progress, particularly in their ability to process long videos through extremely extended context lengths. However, this comes at the cost of significantly increased computational overhead due to the massive number of visual tokens, making efficiency a major bottleneck. […]
Continue.. PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance