ROOD-MRI: Benchmarking the robustness of deep learning segmentation models to out-of-distribution and corrupted data in MRI

Lyndon Boone; Mahdi Biparva; Parisa Mojiri Forooshani; Joel Ramirez; Mario Masellis; Robert Bartha; Sean Symons; Stephen Strother; Sandra E. Black; Chris Heyn; Anne L. Martel; Richard H. Swartz; Maged Goubran

ROOD-MRI：深層学習セグメンテーションモデルのMRIでの分布外および破損したデータに対する堅牢性のベンチマーク

深層人工ニューラルネットワーク（DNN）は、分類、セグメンテーション、および検出の課題に成功したため、医療画像分析の最前線に移動しました。ニューロイメージ分析におけるDNNの大規模な展開における主な課題は、スキャナーと取得プロトコルの違いによる、信号対ノイズ比、コントラスト、解像度、およびサイト間のアーティファクトの存在の変化の可能性です。 DNNは、コンピュータビジョンにおけるこれらの分布の変化の影響を受けやすいことで有名です。現在、MRIの特定の分布シフトに対する新規および既存のモデルの堅牢性を評価するためのベンチマークプラットフォームまたはフレームワークはなく、アクセス可能なマルチサイトベンチマークデータセットはまだ不足しているか、タスク固有です。これらの制限に対処するために、ROOD-MRIを提案します。これは、MRIの分布外（OOD）データ、破損、およびアーティファクトに対するDNNの堅牢性をベンチマークするためのプラットフォームです。このプラットフォームは、MRIで分布シフトをモデル化する変換を使用してベンチマークデータセットを生成するためのモジュール、画像セグメンテーション用に新しく導出されたベンチマークメトリックの実装、および新しいモデルとタスクで方法論を使用するための例を提供します。いくつかの大規模な研究で、海馬、心室、白質の高信号セグメンテーションに方法論を適用し、海馬のデータセットを公開されているベンチマークとして提供します。これらのデータセットで最新のDNNを評価することにより、MRIでの分布の変化や破損の影響を非常に受けやすいことを示しています。データ拡張戦略は解剖学的セグメンテーションタスクのOODデータに対する堅牢性を大幅に向上させることができますが、拡張を使用する最新のDNNは、より困難な病変ベースのセグメンテーションタスクでは依然として堅牢性に欠けることを示します。最後に、U-Netとトランスベースのモデルのベンチマークを行い、アーキテクチャ間で特定のクラスのトランスフォームに対する堅牢性に一貫した違いがあることを確認しました。

Deep artificial neural networks (DNNs) have moved to the forefront of medical image analysis due to their success in classification, segmentation, and detection challenges. A principal challenge in large-scale deployment of DNNs in neuroimage analysis is the potential for shifts in signal-to-noise ratio, contrast, resolution, and presence of artifacts from site to site due to variances in scanners and acquisition protocols. DNNs are famously susceptible to these distribution shifts in computer vision. Currently, there are no benchmarking platforms or frameworks to assess the robustness of new and existing models to specific distribution shifts in MRI, and accessible multi-site benchmarking datasets are still scarce or task-specific. To address these limitations, we propose ROOD-MRI: a platform for benchmarking the Robustness of DNNs to Out-Of-Distribution (OOD) data, corruptions, and artifacts in MRI. The platform provides modules for generating benchmarking datasets using transforms that model distribution shifts in MRI, implementations of newly derived benchmarking metrics for image segmentation, and examples for using the methodology with new models and tasks. We apply our methodology to hippocampus, ventricle, and white matter hyperintensity segmentation in several large studies, providing the hippocampus dataset as a publicly available benchmark. By evaluating modern DNNs on these datasets, we demonstrate that they are highly susceptible to distribution shifts and corruptions in MRI. We show that while data augmentation strategies can substantially improve robustness to OOD data for anatomical segmentation tasks, modern DNNs using augmentation still lack robustness in more challenging lesion-based segmentation tasks. We finally benchmark U-Nets and transformer-based models, finding consistent differences in robustness to particular classes of transforms across architectures.

updated: Fri Mar 11 2022 16:34:15 GMT+0000 (UTC)

published: Fri Mar 11 2022 16:34:15 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト