Deep Dual-resolution Networks for Real-time and Accurate Semantic Segmentation of Road Scenes

Yuanduo Hong; Huihui Pan; Weichao Sun; Yisong Jia

道路シーンのリアルタイムで正確なセマンティックセグメンテーションのためのディープデュアルレゾリューションネットワーク

セマンティックセグメンテーションは、自動運転車が周囲のシーンを理解するための重要なテクノロジーです。現代のモデルの魅力的なパフォーマンスは、通常、重い計算と長い推論時間を犠牲にしてもたらされます。これは、自動運転には耐えられません。最近の方法では、軽量アーキテクチャ（エンコーダ-デコーダまたは2パスウェイ）または低解像度画像の推論を使用して、単一の1080TiGPUで100FPSを超える速度で実行しても、非常に高速なシーン解析を実現します。ただし、これらのリアルタイムの方法と拡張バックボーンに基づくモデルの間には、依然としてパフォーマンスに大きなギャップがあります。この問題に取り組むために、リアルタイムのセマンティックセグメンテーション用に特別に設計された効率的なバックボーンのファミリーを提案しました。提案されているディープデュアルレゾリューションネットワーク（DDRNet）は、2つのディープブランチで構成されており、その間で複数のバイラテラルフュージョンが実行されます。さらに、Deep Aggregation Pyramid Pooling Module（DAPPM）という名前の新しいコンテキスト情報抽出機能を設計して、効果的な受容野を拡大し、低解像度の特徴マップに基づいてマルチスケールコンテキストを融合します。私たちの方法は、CityscapesとCamVidデータセットの両方で精度と速度の間の新しい最先端のトレードオフを実現します。特に、単一の2080Ti GPUでは、DDRNet-23-slimはCityscapesテストセットで102 FPSで77.4％mIoU、CamVidテストセットで230 FPSで74.7％mIoUを生成します。広く使用されているテスト拡張により、私たちの方法はほとんどの最先端モデルよりも優れており、必要な計算がはるかに少なくなります。コードとトレーニング済みモデルはオンラインで入手できます。

Semantic segmentation is a key technology for autonomous vehicles to understand the surrounding scenes. The appealing performances of contemporary models usually come at the expense of heavy computations and lengthy inference time, which is intolerable for self-driving. Using light-weight architectures (encoder-decoder or two-pathway) or reasoning on low-resolution images, recent methods realize very fast scene parsing, even running at more than 100 FPS on a single 1080Ti GPU. However, there is still a significant gap in performance between these real-time methods and the models based on dilation backbones. To tackle this problem, we proposed a family of efficient backbones specially designed for real-time semantic segmentation. The proposed deep dual-resolution networks (DDRNets) are composed of two deep branches between which multiple bilateral fusions are performed. Additionally, we design a new contextual information extractor named Deep Aggregation Pyramid Pooling Module (DAPPM) to enlarge effective receptive fields and fuse multi-scale context based on low-resolution feature maps. Our method achieves a new state-of-the-art trade-off between accuracy and speed on both Cityscapes and CamVid dataset. In particular, on a single 2080Ti GPU, DDRNet-23-slim yields 77.4% mIoU at 102 FPS on Cityscapes test set and 74.7% mIoU at 230 FPS on CamVid test set. With widely used test augmentation, our method is superior to most state-of-the-art models and requires much less computation. Codes and trained models are available online.

updated: Wed Sep 01 2021 08:08:55 GMT+0000 (UTC)

published: Fri Jan 15 2021 12:56:18 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト