Versatile Diffusion: Text, Images and Variations All in One Diffusion Model

Xingqian Xu; Zhangyang Wang; Eric Zhang; Kai Wang; Humphrey Shi

多彩な拡散: テキスト、画像、バリエーションのすべてを 1 つの拡散モデルに統合

拡散モデルの最近の進歩は、多くの生成タスクで印象的なマイルストーンを設定しており、DALL-E2、Imagen、Stable Diffusion などの流行の研究は大きな関心を集めています。ランドスケープが急速に変化しているにもかかわらず、最近の新しいアプローチは容量ではなく拡張機能とパフォーマンスに重点を置いているため、個別のタスクには個別のモデルが必要です。この作業では、既存の単一フローの拡散パイプラインをマルチタスクのマルチモーダルネットワークに拡張します。このネットワークは Versatile Diffusion (VD) と呼ばれ、テキストから画像、画像からテキスト、およびバリエーションの複数のフローを 1 つの統合されたネットワークで処理します。モデル。 VD のパイプライン設計は、画像やテキストを超えたクロスモーダルな一般性を可能にする共有可能で交換可能なレイヤーモジュールで構成される、統合されたマルチフロー拡散フレームワークをインスタンス化します。大規模な実験を通じて、VD が次のことを首尾よく達成することを実証します。 b) VD は、スタイルとセマンティクスのもつれの解消、デュアルおよびマルチコンテキストのブレンディングなどの新しい拡張を可能にします。 c) 画像とテキストに対するマルチフローマルチモーダルフレームワークの成功は、さらなる拡散ベースのユニバーサル AI 研究を刺激する可能性があります。私たちのコードとモデルは、https://github.com/SHI-Labs/Versatile-Diffusion でオープンソース化されています。

Recent advances in diffusion models have set an impressive milestone in many generation tasks, and trending works such as DALL-E2, Imagen, and Stable Diffusion have attracted great interest. Despite the rapid landscape changes, recent new approaches focus on extensions and performance rather than capacity, thus requiring separate models for separate tasks. In this work, we expand the existing single-flow diffusion pipeline into a multi-task multimodal network, dubbed Versatile Diffusion (VD), that handles multiple flows of text-to-image, image-to-text, and variations in one unified model. The pipeline design of VD instantiates a unified multi-flow diffusion framework, consisting of sharable and swappable layer modules that enable the crossmodal generality beyond images and text. Through extensive experiments, we demonstrate that VD successfully achieves the following: a) VD outperforms the baseline approaches and handles all its base tasks with competitive quality; b) VD enables novel extensions such as disentanglement of style and semantics, dual- and multi-context blending, etc.; c) The success of our multi-flow multimodal framework over images and text may inspire further diffusion-based universal AI research. Our code and models are open-sourced at https://github.com/SHI-Labs/Versatile-Diffusion.

updated: Thu Mar 23 2023 07:13:25 GMT+0000 (UTC)

published: Tue Nov 15 2022 17:44:05 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト