SAC-GAN: Structure-Aware Image-to-Image Composition for Self-Driving

Hang Zhou; Ali Mahdavi-Amiri; Rui Ma; Hao Zhang

SAC-GAN：自動運転のための構造を意識した画像から画像への構成

自動運転アプリケーションの画像拡張への構成的アプローチを提示します。これは、オブジェクト画像から背景シーン画像にトリミングされたパッチとして表されるオブジェクト（車両や歩行者など）をシームレスに構成するようにトレーニングされたエンドツーエンドのニューラルネットワークです。私たちのアプローチは、ピクセルレベルのRGB精度ではなく、構成された画像のセマンティックおよび構造の一貫性に重点を置いているため、構造認識機能を使用してネットワークの入力と出力を調整し、それに応じてネットワーク損失を設計します。具体的には、私たちのネットワークは、入力シーン画像からセマンティックレイアウト機能、入力オブジェクトパッチのエッジとシルエットからエンコードされた機能、および入力としての潜在コードを取得し、の変換とスケーリングを定義する2D空間アフィン変換を生成します。オブジェクトパッチ。学習したパラメーターは、微分可能な空間変換ネットワークにさらに供給され、オブジェクトパッチをターゲット画像に変換します。ここで、モデルは、アフィン変換弁別子とレイアウト弁別子を使用して敵対的にトレーニングされます。合成画像の品質、構成可能性、および一般化可能性の観点から、著名な自動運転データセットで、構造認識構成のために造られたSAC-GANというネットワークを評価します。最先端の代替案との比較が行われ、私たちの方法の優位性が確認されています。

We present a compositional approach to image augmentation for self-driving applications. It is an end-to-end neural network that is trained to seamlessly compose an object (e.g., a vehicle or pedestrian) represented as a cropped patch from an object image, into a background scene image. As our approach emphasizes more on semantic and structural coherence of the composed images, rather than their pixel-level RGB accuracies, we tailor the input and output of our network with structure-aware features and design our network losses accordingly. Specifically, our network takes the semantic layout features from the input scene image, features encoded from the edges and silhouette in the input object patch, as well as a latent code as inputs, and generates a 2D spatial affine transform defining the translation and scaling of the object patch. The learned parameters are further fed into a differentiable spatial transformer network to transform the object patch into the target image, where our model is trained adversarially using an affine transform discriminator and a layout discriminator. We evaluate our network, coined SAC-GAN for structure-aware composition, on prominent self-driving datasets in terms of quality, composability, and generalizability of the composite images. Comparisons are made to state-of-the-art alternatives, confirming superiority of our method.

updated: Sat Jan 08 2022 04:10:44 GMT+0000 (UTC)

published: Mon Dec 13 2021 12:24:50 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト