High Resolution Face Editing with Masked GAN Latent Code Optimization

Martin Pernuš; Vitomir Štruc; Simon Dobrišek

マスクされたGAN潜在コード最適化による高解像度の顔編集

顔の編集は、顔の画像の特定の特性を編集することを目的とした、コンピュータビジョンコミュニティで人気のある研究トピックです。最近提案された方法は、条件付きエンコーダーデコーダー生成的敵対的ネットワーク（GAN）をエンドツーエンドでトレーニングするか、事前にトレーニングされたバニラGANジェネレーターモデルの潜在空間で操作を定義することに基づいています。ただし、これらの方法は、ある程度の視覚的劣化を示し、編集された画像のもつれを解く特性を欠いています。さらに、それらは通常、より低い画像解像度で動作します。この論文では、空間的および意味的制約を伴うGAN埋め込み最適化手順を提案します。顔データセットで事前トレーニングされたGANの潜在コードを最適化して、画像の固定領域を埋め込み、顔解析および属性分類ネットワークを使用して、修復された領域に制約を課します。潜在コードの最適化により、GANモデルで定義されているように、画像の確率分布に従うように結果を制約します。このようなフレームワークを使用して、高画質の顔編集を作成します。導入された空間的制約により、編集された画像は、他の方法よりも、目的の顔の属性と画像の残りの部分との間でより高度なもつれを解きます。このアプローチは、3つのデータセットでの実験で、4つの最先端のアプローチと比較して検証されています。結果は、提案されたアプローチが、前例のない画質でいくつかの顔の属性に関して顔の画像を編集できる一方で、望ましくない変動の要因を解きほぐすことができることを示しています。コードが利用可能になります。

Face editing is a popular research topic in the computer vision community that aims to edit a specific characteristic of a face image. Recent proposed methods are based on either training a conditional encoder-decoder Generative Adversarial Network (GAN) in an end-to-end fashion or on defining an operation in the latent space of a pre-trained vanilla GAN generator model. However, these methods exhibit a certain degree of visual degradation and lack disentanglement properties in the edited images. Moreover, they usually operate on lower image resolution. In this paper, we propose a GAN embedding optimization procedure with spatial and semantic constraints. We optimize a latent code of a GAN, pre-trained on face dataset, to embed a fixed region of the image, while imposing constraints on the inpainted regions with face parsing and attribute classification networks. By latent code optimization, we constrain the result to follow an image probability distribution, as defined by the GAN model. We use such framework to produce high image quality face edits. Due to the spatial constraints introduced, the edited images exhibit higher degree of disentanglement between the desired facial attributes and the rest of the image than other methods. The approach is validated in experiments on three datasets and in comparison with four state-of-the-art approaches. The results demonstrate that the proposed approach is able to edit face images with respect to several facial attributes with unprecedented image quality, while disentangling the undesired factors of variation. Code will be made available.

updated: Sat Mar 20 2021 08:39:41 GMT+0000 (UTC)

published: Sat Mar 20 2021 08:39:41 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト