SAT: 2D Semantics Assisted Training for 3D Visual Grounding

Zhengyuan Yang; Songyang Zhang; Liwei Wang; Jiebo Luo

SAT：3Dビジュアルグラウンディングのための2Dセマンティクス支援トレーニング

3Dビジュアルグラウンディングは、通常3Dポイントクラウドの形式で表される3Dシーンに関する自然言語の記述をターゲットオブジェクト領域にグラウンディングすることを目的としています。点群はまばらでノイズが多く、2D画像と比較して限られたセマンティック情報が含まれています。これらの固有の制限により、3Dビジュアルグラウンディング問題はより困難になります。本研究では、トレーニング段階で2D画像セマンティクスを利用して点群と言語の共同表現学習を容易にし、3D視覚的接地を支援する2Dセマンティクス支援トレーニング（SAT）を提案します。主なアイデアは、リッチでクリーンな2Dオブジェクト表現と、3Dシーン内の対応するオブジェクトまたは言及されたエンティティとの間の補助的な配置を学習することです。 SATは、トレーニングの追加入力として2Dオブジェクトのセマンティクス、つまりオブジェクトラベル、画像の特徴、2Dの幾何学的特徴を取りますが、推論中にそのような入力を必要としません。トレーニングで2Dセマンティクスを効果的に利用することにより、私たちのアプローチはNr3Dデータセットの精度を37.7％から49.2％に高めます。これは、同一のネットワークアーキテクチャと推論入力で非SATベースラインを大幅に上回ります。私たちのアプローチは、複数の3Dビジュアルグラウンディングデータセットで最先端の技術を大幅に上回っています。つまり、Nr3Dで+ 10.4％、Sr3Dで+ 9.9％、ScanRefで+ 5.6％です。

3D visual grounding aims at grounding a natural language description about a 3D scene, usually represented in the form of 3D point clouds, to the targeted object region. Point clouds are sparse, noisy, and contain limited semantic information compared with 2D images. These inherent limitations make the 3D visual grounding problem more challenging. In this study, we propose 2D Semantics Assisted Training (SAT) that utilizes 2D image semantics in the training stage to ease point-cloud-language joint representation learning and assist 3D visual grounding. The main idea is to learn auxiliary alignments between rich, clean 2D object representations and the corresponding objects or mentioned entities in 3D scenes. SAT takes 2D object semantics, i.e., object label, image feature, and 2D geometric feature, as the extra input in training but does not require such inputs during inference. By effectively utilizing 2D semantics in training, our approach boosts the accuracy on the Nr3D dataset from 37.7% to 49.2%, which significantly surpasses the non-SAT baseline with the identical network architecture and inference input. Our approach outperforms the state of the art by large margins on multiple 3D visual grounding datasets, i.e., +10.4% absolute accuracy on Nr3D, +9.9% on Sr3D, and +5.6% on ScanRef.

updated: Wed Sep 22 2021 04:35:38 GMT+0000 (UTC)

published: Mon May 24 2021 17:58:36 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト