Video Generation with LDM | | | | Turtles AI
Video Generation with LDM
DukeRem23 April 2023
A team of researchers, namely Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, Karsten Kreis (most of them from #NVIDIA) has developed a new approach to high-resolution #video #generation that promises to revolutionize the way we create and simulate driving scenarios, as well as generate creative content with text-to-video modelling. The approach, called #Latent #Diffusion #Models (#LDM), enables the synthesis of high-quality images while avoiding excessive computing demands.
The LDM approach works by training a diffusion model in a compressed lower-dimensional latent space. The team first pre-trained an LDM on images only, and then turned the image generator into a video generator by introducing a temporal dimension to the latent space diffusion model and fine-tuning on encoded image sequences, or videos. They also temporally aligned diffusion model upsamplers, turning them into temporally consistent video super-resolution models.
The team focused on two relevant real-world applications: simulation of in-the-wild driving data and creative content creation with text-to-video modelling. They validated their Video LDM on real driving videos of resolution 512 × 1024, achieving state-of-the-art performance. Furthermore, their approach can easily leverage off-the-shelf pre-trained image LDMs, as they only need to train a temporal alignment model in that case. By doing so, they turned the publicly available, state-of-the-art text-to-image LDM Stable Diffusion into an efficient and expressive text-to-video model with a resolution of up to 1280 × 2048.
The team showed that the temporal layers trained in this way generalize to different fine-tuned text-to-image LDMs. Utilizing this property, they demonstrated the first results for personalized text-to-video generation, opening exciting directions for future content creation.
In summary, the Video Latent Diffusion Models developed by the team offer an efficient and effective way to generate high-resolution videos, particularly in the context of driving scenarios and creative content creation. The team's key design choice was to build on pre-trained image diffusion models and turn them into video generators by temporally video fine-tuning them with temporal alignment layers. They hope that their work can benefit simulators in the context of autonomous driving research and help democratize high-quality video content creation.
