SAC-GAN: Structure-Aware Image Composition

Hang Zhou; Rui Ma; Ling-Xiao Zhang; Lin Gao; Ali Mahdavi-Amiri; Hao Zhang

SAC-GAN: 構造を意識した画像合成

画像から画像への合成のためのエンドツーエンドの学習フレームワークを導入し、オブジェクト画像からトリミングされたパッチとして表されるオブジェクトを背景シーン画像にもっともらしく合成することを目指しています。私たちのアプローチは、ピクセルレベルのRGB精度ではなく、構成された画像のセマンティックおよび構造的な一貫性を重視しているため、構造認識機能を使用してネットワークの入力と出力を調整し、それに応じてネットワーク損失を設計します。オブジェクトのトリミングによる自己管理設定。具体的には、ネットワークは、入力シーン画像からセマンティックレイアウト機能、入力オブジェクトパッチのエッジとシルエットからエンコードされた機能、潜在コードを入力として取得し、変換とスケーリングを定義する 2D 空間アフィン変換を生成します。オブジェクトパッチ。学習したパラメーターはさらに微分可能な空間変換ネットワークに供給され、オブジェクトパッチがターゲットイメージに変換されます。ここで、モデルはアフィン変換弁別器とレイアウト弁別器を使用して敵対的にトレーニングされます。合成画像の品質、合成可能性、および一般化可能性の観点から、さまざまな画像合成シナリオについて、SAC-GAN と名付けられたネットワークを評価します。インスタンス挿入、ST-GAN、CompGAN、PlaceNet などの最先端の代替手段との比較が行われ、私たちの方法の優位性が確認されました。

We introduce an end-to-end learning framework for image-to-image composition, aiming to plausibly compose an object represented as a cropped patch from an object image into a background scene image. As our approach emphasizes more on semantic and structural coherence of the composed images, rather than their pixel-level RGB accuracies, we tailor the input and output of our network with structure-aware features and design our network losses accordingly, with ground truth established in a self-supervised setting through the object cropping. Specifically, our network takes the semantic layout features from the input scene image, features encoded from the edges and silhouette in the input object patch, as well as a latent code as inputs, and generates a 2D spatial affine transform defining the translation and scaling of the object patch. The learned parameters are further fed into a differentiable spatial transformer network to transform the object patch into the target image, where our model is trained adversarially using an affine transform discriminator and a layout discriminator. We evaluate our network, coined SAC-GAN, for various image composition scenarios in terms of quality, composability, and generalizability of the composite images. Comparisons are made to state-of-the-art alternatives, including Instance Insertion, ST-GAN, CompGAN and PlaceNet, confirming superiority of our method.

updated: Fri Dec 02 2022 09:27:41 GMT+0000 (UTC)

published: Mon Dec 13 2021 12:24:50 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト