Deep Dual-resolution Networks for Real-time and Accurate Semantic Segmentation of Road Scenes

Yuanduo Hong; Huihui Pan; Weichao Sun; Senior Member; IEEE; Yisong Jia

道路シーンのリアルタイムで正確なセマンティックセグメンテーションのためのディープデュアルレゾリューションネットワーク

セマンティックセグメンテーションは、自動運転車が周囲のシーンを理解するための重要なテクノロジーです。実際の自動運転車の場合、高精度のセグメンテーション結果を達成するためにかなりの量の推論時間を費やすことは望ましくありません。最近の方法では、軽量アーキテクチャ（エンコーダ-デコーダまたは2パスウェイ）または低解像度画像の推論を使用して、単一の1080TiGPUで100FPSを超える速度で実行される非常に高速なシーン解析を実現しています。ただし、これらのリアルタイムの方法と拡張バックボーンに基づくモデルの間には、パフォーマンスに明らかなギャップがあります。この問題に取り組むために、道路シーンのリアルタイムセマンティックセグメンテーションのための新しいディープデュアルレゾリューションネットワーク（DDRNets）を提案します。さらに、Deep Aggregation Pyramid Pooling Module（DAPPM）という名前の新しいコンテキスト情報抽出機能を設計して、効果的な受容野を拡大し、マルチスケールコンテキストを融合します。私たちの方法は、CityscapesとCamVidデータセットの両方で精度と速度の間の新しい最先端のトレードオフを実現します。特に、単一の2080Ti GPUでは、DDRNet-23-slimはCityscapesテストセットで109 FPSで77.4％mIoU、CamVidテストセットで230 FPSで74.4％mIoUを生成します。注意メカニズム、より大きなセマンティックセグメンテーションデータセットの事前トレーニング、または推論アクセラレーションを利用せずに、DDRNet-39は都市の景観で23 FPSで80.4％のテストmIoUを達成します。広く使用されているテスト拡張により、私たちの方法は、ほとんどの最先端モデルよりも優れており、計算がはるかに少なくて済みます。コードとトレーニング済みモデルは公開されます。

Semantic segmentation is a critical technology for autonomous vehicles to understand surrounding scenes. For practical autonomous vehicles, it is undesirable to spend a considerable amount of inference time to achieve high-accuracy segmentation results. Using light-weight architectures (encoder-decoder or two-pathway) or reasoning on low-resolution images, recent methods realize very fast scene parsing which even run at more than 100 FPS on single 1080Ti GPU. However, there are still evident gaps in performance between these real-time methods and models based on dilation backbones. To tackle this problem, we propose novel deep dual-resolution networks (DDRNets) for real-time semantic segmentation of road scenes. Besides, we design a new contextual information extractor named Deep Aggregation Pyramid Pooling Module (DAPPM) to enlarge effective receptive fields and fuse multi-scale context. Our method achieves new state-of-the-art trade-off between accuracy and speed on both Cityscapes and CamVid dataset. Specially, on single 2080Ti GPU, DDRNet-23-slim yields 77.4% mIoU at 109 FPS on Cityscapes test set and 74.4% mIoU at 230 FPS on CamVid test set. Without utilizing attention mechanism, pre-training on larger semantic segmentation dataset or inference acceleration, DDRNet-39 attains 80.4% test mIoU at 23 FPS on Cityscapes. With widely used test augmentation, our method is still superior to most state-of-the-art models, requiring much less computation. Codes and trained models will be made publicly available.

updated: Fri Jan 15 2021 12:56:18 GMT+0000 (UTC)

published: Fri Jan 15 2021 12:56:18 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト