3D Siamese Transformer Network for Single Object Tracking on Point Clouds

Le Hui; Lingpeng Wang; Linghua Tang; Kaihao Lan; Jin Xie; Jian Yang

点群での単一オブジェクト追跡のための3Dシャムトランスフォーマーネットワーク

シャムネットワークベースのトラッカーは、テンプレートのポイントフィーチャと検索領域の間の相互相関学習として、3D単一オブジェクトトラッキングを定式化します。追跡中のテンプレートと検索領域の間の外観のばらつきが大きいため、検索領域内の潜在的なターゲットを特定するために、それらの間の堅牢な相互相関をどのように学習するかは、依然として困難な問題です。このホワイトペーパーでは、Transformerを明示的に使用して3Dシャムトランスフォーマーネットワークを形成し、テンプレートと点群の検索領域との間のロバストな相互相関を学習します。具体的には、ターゲットの形状コンテキスト情報を学習するために、シャムポイントトランスフォーマーネットワークを開発します。そのエンコーダーは、自己注意を使用してポイントクラウドの非ローカル情報をキャプチャし、オブジェクトの形状情報を特徴付けます。デコーダーは、クロスアテンションを使用して、識別可能なポイントの特徴をアップサンプリングします。その後、テンプレートと検索領域の間のロバストな相互相関を学習するために、反復的な粗から細への相関ネットワークを開発します。これは、クロスアテンションを介してテンプレートを検索領域内の潜在的なターゲットに関連付けるためのクロス機能拡張を定式化します。潜在的なターゲットをさらに強化するために、特徴空間のローカルk-NNグラフに自己注意を適用してターゲット特徴を集約する自我特徴拡張を採用しています。 KITTI、nuScenes、およびWaymoデータセットでの実験は、私たちの方法が3D単一オブジェクト追跡タスクで最先端のパフォーマンスを達成することを示しています。

Siamese network based trackers formulate 3D single object tracking as cross-correlation learning between point features of a template and a search area. Due to the large appearance variation between the template and search area during tracking, how to learn the robust cross correlation between them for identifying the potential target in the search area is still a challenging problem. In this paper, we explicitly use Transformer to form a 3D Siamese Transformer network for learning robust cross correlation between the template and the search area of point clouds. Specifically, we develop a Siamese point Transformer network to learn shape context information of the target. Its encoder uses self-attention to capture non-local information of point clouds to characterize the shape information of the object, and the decoder utilizes cross-attention to upsample discriminative point features. After that, we develop an iterative coarse-to-fine correlation network to learn the robust cross correlation between the template and the search area. It formulates the cross-feature augmentation to associate the template with the potential target in the search area via cross attention. To further enhance the potential target, it employs the ego-feature augmentation that applies self-attention to the local k-NN graph of the feature space to aggregate target features. Experiments on the KITTI, nuScenes, and Waymo datasets show that our method achieves state-of-the-art performance on the 3D single object tracking task.

updated: Mon Jul 25 2022 09:08:30 GMT+0000 (UTC)

published: Mon Jul 25 2022 09:08:30 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト