SAC-GAN: Structure-Aware Image Composition

Hang Zhou; Rui Ma; Lingxiao Zhang; Lin Gao; Ali Mahdavi-Amiri; Hao Zhang

SAC-GAN：構造を意識した画像構成

オブジェクト画像から背景シーン画像にトリミングされたパッチとして表されるオブジェクトをシームレスに構成することを目的とした、画像から画像への合成のためのエンドツーエンドの学習フレームワークを紹介します。私たちのアプローチは、ピクセルレベルのRGB精度ではなく、構成された画像のセマンティックおよび構造の一貫性に重点を置いているため、構造認識機能を使用してネットワークの入力と出力を調整し、それに応じてネットワーク損失を設計します。オブジェクトのトリミングによる自己監視設定。具体的には、私たちのネットワークは、入力シーン画像からセマンティックレイアウト機能、入力オブジェクトパッチのエッジとシルエットからエンコードされた機能、および入力としての潜在コードを取得し、の変換とスケーリングを定義する2D空間アフィン変換を生成します。オブジェクトパッチ。学習したパラメーターは、微分可能な空間トランスフォーマーネットワークにさらに供給され、オブジェクトパッチをターゲット画像に変換します。ここで、モデルは、アフィン変換弁別子とレイアウト弁別子を使用して敵対的にトレーニングされます。合成画像の品質、構成可能性、および一般化可能性の観点から、さまざまな画像構成シナリオについて、ネットワークである造語SAC-GANを評価します。インスタンス挿入、ST-GAN、CompGAN、PlaceNetなどの最先端の代替手段と比較され、この方法の優位性が確認されました。

We introduce an end-to-end learning framework for image-to-image composition, aiming to seamlessly compose an object represented as a cropped patch from an object image into a background scene image. As our approach emphasizes more on semantic and structural coherence of the composed images, rather than their pixel-level RGB accuracies, we tailor the input and output of our network with structure-aware features and design our network losses accordingly, with ground truth established in a self-supervised setting through the object cropping. Specifically, our network takes the semantic layout features from the input scene image, features encoded from the edges and silhouette in the input object patch, as well as a latent code as inputs, and generates a 2D spatial affine transform defining the translation and scaling of the object patch. The learned parameters are further fed into a differentiable spatial transformer network to transform the object patch into the target image, where our model is trained adversarially using an affine transform discriminator and a layout discriminator. We evaluate our network, coined SAC-GAN, for various image composition scenarios in terms of quality, composability, and generalizability of the composite images. Comparisons are made to state-of-the-art alternatives, including Instance Insertion, ST-GAN, CompGAN and PlaceNet, confirming superiority of our method.

updated: Tue Jul 05 2022 10:07:40 GMT+0000 (UTC)

published: Mon Dec 13 2021 12:24:50 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト