Applying Plain Transformers to Real-World Point Clouds

Lanxiao Li; Michael Heizmann

プレーントランスフォーマーを実世界の点群に適用する

誘導性バイアスがないため、変圧器ベースのモデルには通常、大量のトレーニングデータが必要です。 3D データは取得して注釈を付けるのが難しいため、この問題は特に 3D ビジョンに関係しています。この問題を克服するために、以前の研究では、変圧器のアーキテクチャを変更して、局所的な注意やダウンサンプリングなどを適用することで誘導バイアスを組み込みました。彼らは有望な結果を達成しましたが、点群の変換に関する以前の研究には 2 つの問題がありました。まず、単純な変圧器の能力はまだ十分に研究されていません。第二に、複雑な実世界の点群ではなく、単純で小さな点群に焦点を当てています。この作業は、現実世界の点群の理解における単純な変換器を再検討します。最初に、効率とパフォーマンスの両方について、単純なトランスフォーマーのいくつかの基本的なコンポーネント (パッチ適用子や位置埋め込みなど) を詳しく見ていきます。誘導バイアスと注釈付きデータの欠如によるパフォーマンスのギャップを埋めるために、マスクされたオートエンコーダー (MAE) を使用した自己教師付き事前トレーニングを調査します。具体的には、情報漏えいを防ぎ、MAEの効果を大幅に向上させるドロップパッチを提案します。私たちのモデルは、より低い計算コストで、S3DIS データセットのセマンティックセグメンテーションと ScanNet データセットのオブジェクト検出で SOTA の結果を達成します。私たちの仕事は、点群の変換器に関する将来の研究の新しいベースラインを提供します。

Due to the lack of inductive bias, transformer-based models usually require a large amount of training data. The problem is especially concerning in 3D vision, as 3D data are harder to acquire and annotate. To overcome this problem, previous works modify the architecture of transformers to incorporate inductive biases by applying, e.g., local attention and down-sampling. Although they have achieved promising results, earlier works on transformers for point clouds have two issues. First, the power of plain transformers is still under-explored. Second, they focus on simple and small point clouds instead of complex real-world ones. This work revisits the plain transformers in real-world point cloud understanding. We first take a closer look at some fundamental components of plain transformers, e.g., patchifier and positional embedding, for both efficiency and performance. To close the performance gap due to the lack of inductive bias and annotated data, we investigate self-supervised pre-training with masked autoencoder (MAE). Specifically, we propose drop patch, which prevents information leakage and significantly improves the effectiveness of MAE. Our models achieve SOTA results in semantic segmentation on the S3DIS dataset and object detection on the ScanNet dataset with lower computational costs. Our work provides a new baseline for future research on transformers for point clouds.

updated: Sat Mar 04 2023 13:07:24 GMT+0000 (UTC)

published: Tue Feb 28 2023 21:06:36 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト