Kavli Affiliate: Hsiao-Mei (Sherry) Cho| First 5 Authors: [#item_custom_name[1, [#item_custom_name[2, [#item_custom_name[3, [#item_custom_name[4, [#item_custom_name[5| Summary:Recent advances in text-to-video diffusion models have enabled high-quality video synthesis, but controllable generation remains challenging, particularly under limited data and compute. Existing fine-tuning methods for conditional generation often rely on external encoders or architectural modifications, which demand large datasets and are […]
Continue.. Temporal In-Context Fine-Tuning with Temporal Reasoning for Versatile Control of Video Diffusion Models