JVID: Joint Video-Image Diffusion for Visual-Quality and Temporal-Consistency in Video Generation

Hadrien Reynaud; Matthew Baugh; Mischa Dombrowski; Sarah Cechnicka; Qingjie Meng; Bernhard Kainz

We introduce the Joint Video-Image Diffusion model (JVID), a novel approach to generating high-quality and temporally coherent videos. We achieve this by integrating two diffusion models: a Latent Image Diffusion Model (LIDM) trained on images and a Latent Video Diffusion Model (LVDM) trained on video data. Our method combines these models in the reverse diffusion process, where the LIDM enhances image quality and the LVDM ensures temporal consistency. This unique combination allows us to effectively handle the complex spatio-temporal dynamics in video generation. Our results demonstrate quantitative and qualitative improvements in producing realistic and coherent videos.

updated: Fri Sep 27 2024 10:32:29 GMT+0000 (UTC)

published: Sat Sep 21 2024 13:59:50 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト