Forward-Looking Sonar Patch Matching: Modern CNNs, Ensembling, and Uncertainty

Arka Mallick; Paul Plöger; Matias Valdenegro-Toro

将来を見据えたソナーパッチマッチング：最新のCNN、アンサンブル、および不確実性

水中ロボットのアプリケーションは増加しており、それらのほとんどは水中視覚をソナーに依存していますが、強力な知覚能力の欠如がこのタスクでそれらを制限しています。ソナー知覚の重要な問題は、画像パッチのマッチングです。これにより、ローカリゼーション、変更検出、マッピングなどの他の手法が可能になります。カラー画像にはこの問題に関する豊富な文献がありますが、音響画像の場合、これらの画像を生成する物理学のために不足しています。この論文では、この問題の以前の結果（Valdenegro-Toro et al、2017）を改善し、特徴を手動でモデル化する代わりに、畳み込みニューラルネットワーク（CNN）が類似性関数を学習し、2つの入力ソナー画像が類似しているかどうかを予測します。ソナー画像マッチングの問題をさらに改善する目的で、3つの最先端のCNNアーキテクチャ、つまりDenseNetとVGGが、シャムまたは2チャネルアーキテクチャと対照的な損失で海洋ゴミデータセットで評価されます。各ネットワークの公正な評価を確実にするために、徹底的なハイパーパラメータ最適化が実行されます。最高のパフォーマンスを発揮するモデルは、0.955AUCのDenseNet2チャネルネットワーク、0.949AUCの対照的な損失のVGG-Siamese、および0.921AUCのDenseNetSiameseであることがわかります。最高のパフォーマンスを発揮するDenseNet2チャネルモデルとDenseNet-Siameseモデルをアンサンブルすることにより、得られる全体的な最高の予測精度は0.978 AUCであり、最先端の0.91AUCを大幅に上回っています。

Application of underwater robots are on the rise, most of them are dependent on sonar for underwater vision, but the lack of strong perception capabilities limits them in this task. An important issue in sonar perception is matching image patches, which can enable other techniques like localization, change detection, and mapping. There is a rich literature for this problem in color images, but for acoustic images, it is lacking, due to the physics that produce these images. In this paper we improve on our previous results for this problem (Valdenegro-Toro et al, 2017), instead of modeling features manually, a Convolutional Neural Network (CNN) learns a similarity function and predicts if two input sonar images are similar or not. With the objective of improving the sonar image matching problem further, three state of the art CNN architectures are evaluated on the Marine Debris dataset, namely DenseNet, and VGG, with a siamese or two-channel architecture, and contrastive loss. To ensure a fair evaluation of each network, thorough hyper-parameter optimization is executed. We find that the best performing models are DenseNet Two-Channel network with 0.955 AUC, VGG-Siamese with contrastive loss at 0.949 AUC and DenseNet Siamese with 0.921 AUC. By ensembling the top performing DenseNet two-channel and DenseNet-Siamese models overall highest prediction accuracy obtained is 0.978 AUC, showing a large improvement over the 0.91 AUC in the state of the art.

updated: Mon Aug 02 2021 17:49:56 GMT+0000 (UTC)

published: Mon Aug 02 2021 17:49:56 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト