Exploring Deep 3D Spatial Encodings for Large-Scale 3D Scene Understanding

Saqib Ali Khan; Yilei Shi; Muhammad Shahzad; Xiao Xiang Zhu

大規模な3Dシーンを理解するためのディープ3D空間エンコーディングの調査

生の3D点群のセマンティックセグメンテーションは、3Dシーン分析に不可欠なコンポーネントですが、主に3D点群の非ユークリッド性のために、いくつかの課題があります。このタスクに対処するためにいくつかの深層学習ベースのアプローチが提案されていますが、それらのほとんどすべてが、従来の畳み込みニューラルネットワーク（CNN）からの潜在的な（グローバル）特徴表現の使用に重点を置いており、空間情報の深刻な損失をもたらし、 3Dシーンのリモートセンシングで重要な役割を果たす、基になる3Dオブジェクトのジオメトリをモデル化します。この手紙では、生の3D点群の空間的特徴を無向対称グラフモデルにエンコードすることにより、CNNベースのアプローチの制限を克服するための代替アプローチを提案しました。次に、これらのエンコーディングは、従来のCNNから抽出された高次元の特徴ベクトルと組み合わされて、必要な3Dセグメンテーションマップを出力するローカライズされたグラフ畳み込み演算子になります。 2つの標準ベンチマークデータセット（屋外の空中リモートセンシングデータセットと屋内の合成データセットを含む）で実験を行いました。提案された方法は、改善されたトレーニング時間とモデルの安定性を備えた最先端の精度を達成し、したがって、3Dシーン理解のための一般化された最先端の方法に向けたさらなる研究の強力な可能性を示しています。

Semantic segmentation of raw 3D point clouds is an essential component in 3D scene analysis, but it poses several challenges, primarily due to the non-Euclidean nature of 3D point clouds. Although, several deep learning based approaches have been proposed to address this task, but almost all of them emphasized on using the latent (global) feature representations from traditional convolutional neural networks (CNN), resulting in severe loss of spatial information, thus failing to model the geometry of the underlying 3D objects, that plays an important role in remote sensing 3D scenes. In this letter, we have proposed an alternative approach to overcome the limitations of CNN based approaches by encoding the spatial features of raw 3D point clouds into undirected symmetrical graph models. These encodings are then combined with a high-dimensional feature vector extracted from a traditional CNN into a localized graph convolution operator that outputs the required 3D segmentation map. We have performed experiments on two standard benchmark datasets (including an outdoor aerial remote sensing dataset and an indoor synthetic dataset). The proposed method achieves on par state-of-the-art accuracy with improved training time and model stability thus indicating strong potential for further research towards a generalized state-of-the-art method for 3D scene understanding.

updated: Sun Nov 29 2020 12:56:19 GMT+0000 (UTC)

published: Sun Nov 29 2020 12:56:19 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト