Temporal-Channel Transformer for 3D Lidar-Based Video Object Detection in Autonomous Driving

Zhenxun Yuan; Xiao Song; Lei Bai; Wengang Zhou; Zhe Wang; Wanli Ouyang

自動運転における3Dライダーベースのビデオオブジェクト検出のための時間チャネルトランスフォーマー

業界における自動運転の強い需要は、3Dオブジェクト検出への強い関心につながり、多くの優れた3Dオブジェクト検出アルゴリズムをもたらしました。ただし、アルゴリズムの大部分は、データシーケンスの時間情報を無視して、単一フレームデータのみをモデル化します。この作業では、Lidarデータからビデオオブジェクトを検出するための時空間ドメインとチャネルドメインの関係をモデル化するために、Temporal-ChannelTransformerと呼ばれる新しいトランスフォーマーを提案します。このトランスフォーマーの特別な設計として、エンコーダーでエンコードされる情報はデコーダーの情報とは異なります。つまり、エンコーダーは複数のフレームの時間チャネル情報をエンコードし、デコーダーはボクセル内の現在のフレームの空間チャネル情報をデコードします。賢明な方法。具体的には、トランスの時間チャネルエンコーダは、異なるチャネルおよびフレームからの特徴間の相関を利用することによって、異なるチャネルおよびフレームの情報をエンコードするように設計されている。一方、トランスフォーマーの空間デコーダーは、現在のフレームの各位置の情報をデコードします。検出ヘッドで物体検出を行う前に、現在のフレームの特徴を再較正するためのゲートメカニズムが展開されます。これは、アップサンプリングプロセスとともにターゲットフレームの表現を繰り返し改良することにより、オブジェクトに関係のない情報を除外します。実験結果は、nuScenesベンチマークでグリッドボクセルベースの3Dオブジェクト検出で最先端のパフォーマンスを達成することを示しています。

The strong demand of autonomous driving in the industry has lead to strong interest in 3D object detection and resulted in many excellent 3D object detection algorithms. However, the vast majority of algorithms only model single-frame data, ignoring the temporal information of the sequence of data. In this work, we propose a new transformer, called Temporal-Channel Transformer, to model the spatial-temporal domain and channel domain relationships for video object detecting from Lidar data. As a special design of this transformer, the information encoded in the encoder is different from that in the decoder, i.e. the encoder encodes temporal-channel information of multiple frames while the decoder decodes the spatial-channel information for the current frame in a voxel-wise manner. Specifically, the temporal-channel encoder of the transformer is designed to encode the information of different channels and frames by utilizing the correlation among features from different channels and frames. On the other hand, the spatial decoder of the transformer will decode the information for each location of the current frame. Before conducting the object detection with detection head, the gate mechanism is deployed for re-calibrating the features of current frame, which filters out the object irrelevant information by repetitively refine the representation of target frame along with the up-sampling process. Experimental results show that we achieve the state-of-the-art performance in grid voxel-based 3D object detection on the nuScenes benchmark.

updated: Fri Nov 27 2020 09:35:39 GMT+0000 (UTC)

published: Fri Nov 27 2020 09:35:39 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト