Deep Equilibrium Approaches to Diffusion Models

Ashwini Pokle; Zhengyang Geng; Zico Kolter

拡散モデルへの深い均衡アプローチ

拡散ベースの生成モデルは、高品質の画像を生成するのに非常に効果的であり、生成されたサンプルは、多くの場合、いくつかのメトリックの下で他のモデルによって生成された品質を上回っています。ただし、これらのモデルの際立った特徴の 1 つは、通常、忠実度の高い画像を生成するために長いサンプリングチェーンが必要になることです。これは、サンプリング時間のレンズからだけでなく、モデルの反転などのタスクを達成するためにこれらのチェーンを介して逆伝播すること、つまり既知の画像を生成する潜在状態を近似的に見つけることの固有の難しさからも課題を提示します。この論文では、(深い) 均衡 (DEQ) 固定点モデルという別の視点から拡散モデルを見ていきます。具体的には、最近のノイズ除去拡散陰解モデル (DDIM; Song et al. 2020) を拡張し、サンプリングチェーン全体をジョイントの多変量固定小数点システムとしてモデル化します。このセットアップは、拡散モデルと平衡モデルの洗練された統合を提供し、1) 単一の画像サンプリングで利点を示します。 2) モデル反転。DEQ 設定で高速勾配を利用して、特定の画像を生成するノイズをより迅速に見つけることができます。このアプローチは直交的であるため、サンプリング時間を短縮したり、モデルの反転を改善したりするために使用される他の方法を補完します。 CIFAR10、CelebA、LSUN の寝室と教会を含むいくつかのデータセットで、この方法の強力なパフォーマンスを実証します。

Diffusion-based generative models are extremely effective in generating high-quality images, with generated samples often surpassing the quality of those produced by other models under several metrics. One distinguishing feature of these models, however, is that they typically require long sampling chains to produce high-fidelity images. This presents a challenge not only from the lenses of sampling time, but also from the inherent difficulty in backpropagating through these chains in order to accomplish tasks such as model inversion, i.e. approximately finding latent states that generate known images. In this paper, we look at diffusion models through a different perspective, that of a (deep) equilibrium (DEQ) fixed point model. Specifically, we extend the recent denoising diffusion implicit model (DDIM; Song et al. 2020), and model the entire sampling chain as a joint, multivariate fixed point system. This setup provides an elegant unification of diffusion and equilibrium models, and shows benefits in 1) single image sampling, as it replaces the fully-serial typical sampling process with a parallel one; and 2) model inversion, where we can leverage fast gradients in the DEQ setting to much more quickly find the noise that generates a given image. The approach is also orthogonal and thus complementary to other methods used to reduce the sampling time, or improve model inversion. We demonstrate our method's strong performance across several datasets, including CIFAR10, CelebA, and LSUN Bedrooms and Churches.

updated: Sun Oct 23 2022 22:02:19 GMT+0000 (UTC)

published: Sun Oct 23 2022 22:02:19 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト