Encoding in Style: a StyleGAN Encoder for Image-to-Image Translation

Elad Richardson; Yuval Alaluf; Or Patashnik; Yotam Nitzan; Yaniv Azar; Stav Shapiro; Daniel Cohen-Or

スタイルでのエンコーディング：画像から画像への変換のためのStyleGANエンコーダ

一般的な画像から画像への変換フレームワーク、pixel2style2pixel（pSp）を紹介します。私たちのpSpフレームワークは、事前にトレーニングされたStyleGANジェネレーターに供給される一連のスタイルベクトルを直接生成する新しいエンコーダーネットワークに基づいており、拡張されたW +潜在空間を形成します。最初に、エンコーダーが追加の最適化なしで実際の画像をW +に直接埋め込むことができることを示します。次に、エンコーダーを使用して画像から画像への変換タスクを直接解決し、それらをある入力ドメインから潜在ドメインへのエンコードの問題として定義することを提案します。最初に標準の反転から逸脱し、以前のStyleGANエンコーダーで使用されていた後の方法を編集することで、入力画像がStyleGANドメインで表されていない場合でも、このアプローチでさまざまなタスクを処理できます。 StyleGANを介して翻訳タスクを解決すると、敵が必要ないため、トレーニングプロセスが大幅に簡素化され、ピクセル間の対応なしでタスクを解決するためのサポートが向上し、スタイルのリサンプリングによるマルチモーダル合成が本質的にサポートされることを示します。最後に、単一のタスク用に特別に設計された最先端のソリューションと比較した場合でも、さまざまな顔の画像から画像への翻訳タスクでのフレームワークの可能性を示し、さらにそれを超えて拡張できることを示します人間の顔の領域。

We present a generic image-to-image translation framework, pixel2style2pixel (pSp). Our pSp framework is based on a novel encoder network that directly generates a series of style vectors which are fed into a pretrained StyleGAN generator, forming the extended W+ latent space. We first show that our encoder can directly embed real images into W+, with no additional optimization. Next, we propose utilizing our encoder to directly solve image-to-image translation tasks, defining them as encoding problems from some input domain into the latent domain. By deviating from the standard invert first, edit later methodology used with previous StyleGAN encoders, our approach can handle a variety of tasks even when the input image is not represented in the StyleGAN domain. We show that solving translation tasks through StyleGAN significantly simplifies the training process, as no adversary is required, has better support for solving tasks without pixel-to-pixel correspondence, and inherently supports multi-modal synthesis via the resampling of styles. Finally, we demonstrate the potential of our framework on a variety of facial image-to-image translation tasks, even when compared to state-of-the-art solutions designed specifically for a single task, and further show that it can be extended beyond the human facial domain.

updated: Wed Apr 21 2021 12:53:36 GMT+0000 (UTC)

published: Mon Aug 03 2020 15:30:38 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト