Multimodal Image Synthesis and Editing: A Survey

Fangneng Zhan; Yingchen Yu; Rongliang Wu; Jiahui Zhang; Shijian Lu; Lingjie Liu; Adam Kortylewski; Christian Theobalt; Eric Xing

マルチモーダル画像の合成と編集：調査

情報は実世界のさまざまなモダリティに存在するため、マルチモーダル情報間の効果的な相互作用と融合は、コンピュータビジョンと深層学習の研究におけるマルチモーダルデータの作成と認識に重要な役割を果たします。マルチモーダル情報間の相互作用をモデル化する優れた能力により、マルチモーダル画像の合成と編集は、近年、ホットな研究トピックになっています。マルチモーダルガイダンスは、ネットワークトレーニングの明示的なガイダンスを提供する代わりに、画像の合成と編集のための直感的で柔軟な手段を提供します。一方、この分野では、特徴と固有のモダリティギャップの調整、高解像度画像の合成、忠実な評価指標など、いくつかの課題にも直面しています。この調査では、最近のマルチモーダル画像合成の進歩を包括的にコンテキスト化します。データモダリティとモデルアーキテクチャに従って分類法を編集および定式化します。まず、画像の合成と編集におけるさまざまなタイプのガイダンスモダリティの紹介から始めます。次に、生成的敵対的ネットワーク（GAN）、自己回帰モデル、拡散モデル、神経放射輝度フィールド（NeRF）などの詳細なフレームワークを使用して、マルチモーダル画像合成および編集アプローチについて詳しく説明します。続いて、マルチモーダル画像の合成と編集で広く採用されているベンチマークデータセットと対応する評価指標の包括的な説明、およびそれぞれの利点と制限の分析を伴うさまざまな合成方法の詳細な比較が続きます。最後に、現在の研究課題と将来の研究の可能な方向性についての洞察を提供します。この調査が、マルチモーダル画像の合成と編集の将来の開発のための健全で価値のある基盤となることを願っています。この調査に関連するプロジェクトは、https：//github.com/fnzhan/MISEで入手できます。

As information exists in various modalities in real world, effective interaction and fusion among multimodal information plays a key role for the creation and perception of multimodal data in computer vision and deep learning research. With superb power in modelling the interaction among multimodal information, multimodal image synthesis and editing has become a hot research topic in recent years. Instead of providing explicit guidance for network training, multimodal guidance offers intuitive and flexible means for image synthesis and editing. On the other hand, this field is also facing several challenges in alignment of features with inherent modality gaps, synthesis of high-resolution images, faithful evaluation metrics, etc. In this survey, we comprehensively contextualize the advance of the recent multimodal image synthesis and editing and formulate taxonomies according to data modality and model architectures. We start with an introduction to different types of guidance modalities in image synthesis and editing. We then describe multimodal image synthesis and editing approaches extensively with detailed frameworks including Generative Adversarial Networks (GANs), Auto-regressive models, Diffusion models, Neural Radiance Fields (NeRF) and other methods. This is followed by a comprehensive description of benchmark datasets and corresponding evaluation metrics as widely adopted in multimodal image synthesis and editing, as well as detailed comparisons of various synthesis methods with analysis of respective advantages and limitations. Finally, we provide insights about the current research challenges and possible directions for future research. We hope this survey could lay a sound and valuable foundation for future development of multimodal image synthesis and editing. A project associated with this survey is available at https://github.com/fnzhan/MISE.

updated: Sun Jul 24 2022 15:54:48 GMT+0000 (UTC)

published: Mon Dec 27 2021 10:00:16 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト