Enhancing Object Detection for Autonomous Driving by Optimizing Anchor Generation and Addressing Class Imbalance

Manuel Carranza-García; Pedro Lara-Benítez; Jorge García-Gutiérrez; José C. Riquelme

アンカー生成を最適化し、クラスの不均衡に対処することにより、自動運転のためのオブジェクト検出を強化

物体検出は、過去数年間、コンピュータビジョンで最も活発なトピックの1つです。最近の作品は、主に汎用COCOベンチマークで最先端技術を推進することに焦点を当てています。ただし、自動運転などの特定のアプリケーションでこのような検出フレームワークを使用することは、まだ対処すべき領域です。この研究では、自動運転車のコンテキストにより適したFasterR-CNNに基づく拡張2Dオブジェクト検出器を紹介します。 2つの主要な側面が改善されました。アンカー生成手順と少数派クラスのパフォーマンスの低下です。デフォルトの均一アンカー構成は、車両カメラの透視投影のため、このシナリオには適していません。したがって、クラスタリングを介して画像を主要な領域に分割し、進化的アルゴリズムを使用して各領域のベースアンカーを最適化する、パースペクティブを意識した方法論を提案します。さらに、第1段階で提案された候補領域の空間情報を含めることにより、第2段階のヘッダーネットワークの精度を高めるモジュールを追加します。また、前景と前景のクラスの不均衡に対処するためのさまざまな再重み付け戦略を検討し、焦点損失の削減バージョンを使用すると、2段階の検出器で困難で過小評価されたオブジェクトの検出を大幅に改善できることを示します。最後に、さまざまな学習戦略の長所を組み合わせるためのアンサンブルモデルを設計します。私たちの提案は、最新で最も広範で多様なWaymo OpenDatasetを使用して評価されます。結果は、最良の単一モデルを使用した場合の平均精度が6.13％mAP向上し、アンサンブルを使用した場合の平均精度が9.69％向上したことを示しています。 Faster R-CNNに対して提案された変更は、計算コストを増加させず、他のアンカーベースの検出フレームワークを最適化するために簡単に拡張できます。

Object detection has been one of the most active topics in computer vision for the past years. Recent works have mainly focused on pushing the state-of-the-art in the general-purpose COCO benchmark. However, the use of such detection frameworks in specific applications such as autonomous driving is yet an area to be addressed. This study presents an enhanced 2D object detector based on Faster R-CNN that is better suited for the context of autonomous vehicles. Two main aspects are improved: the anchor generation procedure and the performance drop in minority classes. The default uniform anchor configuration is not suitable in this scenario due to the perspective projection of the vehicle cameras. Therefore, we propose a perspective-aware methodology that divides the image into key regions via clustering and uses evolutionary algorithms to optimize the base anchors for each of them. Furthermore, we add a module that enhances the precision of the second-stage header network by including the spatial information of the candidate regions proposed in the first stage. We also explore different re-weighting strategies to address the foreground-foreground class imbalance, showing that the use of a reduced version of focal loss can significantly improve the detection of difficult and underrepresented objects in two-stage detectors. Finally, we design an ensemble model to combine the strengths of the different learning strategies. Our proposal is evaluated with the Waymo Open Dataset, which is the most extensive and diverse up to date. The results demonstrate an average accuracy improvement of 6.13% mAP when using the best single model, and of 9.69% mAP with the ensemble. The proposed modifications over the Faster R-CNN do not increase computational cost and can easily be extended to optimize other anchor-based detection frameworks.

updated: Thu Apr 08 2021 16:58:31 GMT+0000 (UTC)

published: Thu Apr 08 2021 16:58:31 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト