Semantic Text-to-Face GAN -ST^2FG

Manan Oza; Sukalpa Chanda; David Doermann

セマンティックテキストツーフェイスGAN-ST ^ 2FG

生成的敵対的ネットワーク（GAN）を使用して生成された顔は、前例のないリアリズムに達しました。「ディープフェイク」としても知られるこれらの顔は、ピクセルレベルの歪みがほとんどないリアルな写真として表示されます。一部の作業では、被験者の特定のプロパティの生成につながるモデルのトレーニングが可能になりましたが、自然言語の説明に基づいて顔の画像を生成することは十分に検討されていません。セキュリティと犯罪者の識別には、スケッチアーティストのように機能するGANベースのシステムを提供する機能が非常に役立ちます。この論文では、セマンティックテキスト記述から顔画像を生成するための新しいアプローチを提示します。学習したモデルには、モデルがフィーチャをスケッチするために使用するテキストの説明と顔のタイプのアウトラインが提供されます。私たちのモデルは、アフィン結合モジュール（ACM）メカニズムを使用してトレーニングされ、自己注意マトリックスを使用して、BERTからのテキスト埋め込みとGAN潜在空間を結合します。これにより、テキストの埋め込みと潜在ベクトルが単純に連結されている場合に発生する可能性がある不適切な「注意」による機能の損失が回避されます。私たちのアプローチは、顔の多くの細部の特徴を備えた顔の徹底的なテキスト記述に非常に正確に位置合わせされた画像を生成することができ、より良い画像を生成するのに役立ちます。提案された方法は、追加のテキストによる説明または文が提供されている場合、以前に生成された画像に増分変更を加えることもできます。

Faces generated using generative adversarial networks (GANs) have reached unprecedented realism. These faces, also known as "Deep Fakes", appear as realistic photographs with very little pixel-level distortions. While some work has enabled the training of models that lead to the generation of specific properties of the subject, generating a facial image based on a natural language description has not been fully explored. For security and criminal identification, the ability to provide a GAN-based system that works like a sketch artist would be incredibly useful. In this paper, we present a novel approach to generate facial images from semantic text descriptions. The learned model is provided with a text description and an outline of the type of face, which the model uses to sketch the features. Our models are trained using an Affine Combination Module (ACM) mechanism to combine the text embedding from BERT and the GAN latent space using a self-attention matrix. This avoids the loss of features due to inadequate "attention", which may happen if text embedding and latent vector are simply concatenated. Our approach is capable of generating images that are very accurately aligned to the exhaustive textual descriptions of faces with many fine detail features of the face and helps in generating better images. The proposed method is also capable of making incremental changes to a previously generated image if it is provided with additional textual descriptions or sentences.

updated: Fri Aug 26 2022 12:51:49 GMT+0000 (UTC)

published: Thu Jul 22 2021 15:42:25 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト