Self-supervised 3D anatomy segmentation using self-distilled masked image transformer (SMIT)

Jue Jiang; Neelam Tyagi; Kathryn Tringale; Christopher Crane; Harini Veeraraghavan

自己蒸留マスク画像変換器（SMIT）を使用した自己監視3D解剖学的セグメンテーション

ビジョントランスフォーマーは、長距離コンテキストをより効率的にモデル化する機能を備えており、セグメンテーションを含むいくつかのコンピュータービジョンおよび医療画像分析タスクで印象的な精度の向上を示しています。ただし、このような方法では、トレーニング用に大きなラベル付きデータセットが必要であり、医療画像分析では取得が困難です。自己監視学習（SSL）は、畳み込みネットワークを使用した医療画像セグメンテーションで成功を収めています。この作業では、CTおよびMRIからの3D多臓器セグメンテーションに適用されるビジョントランスフォーマー（SMIT）のSSLを実行するために、マスクされた画像モデリング方法を使用した自己蒸留学習を開発しました。私たちの貢献は、マスクされた画像予測と呼ばれるマスクされたパッチ内の高密度のピクセル単位の回帰です。これを、ビジョントランスフォーマーを事前トレーニングするための口実タスクとしてマスクされたパッチトークンの蒸留と組み合わせました。私たちのアプローチは、他の口実タスクよりも正確で、必要なデータセットの微調整が少ないことを示しています。疾患部位から生じる画像セットと標的タスクに対応する画像モダリティを通常使用していた以前の医用画像法とは異なり、頭頸部癌、肺癌、腎臓癌、およびCOVID-19から生じる3,643 CTスキャン（602,708画像）を使用しました。事前トレーニングのために、MRI膵臓癌患者からの腹部臓器セグメンテーション、およびCTからの公的に利用可能な13の異なる腹部臓器セグメンテーションに適用しました。私たちの方法は、一般的に使用される口実タスクよりもデータセットを微調整する必要性が減少し、明確な精度の向上（MRIからの平均DSCが0.875、CTからの平均DSCが0.878）を示しました。複数の現在のSSLメソッドとの広範な比較が行われました。コードは、公開が承認されると利用可能になります。

Vision transformers, with their ability to more efficiently model long-range context, have demonstrated impressive accuracy gains in several computer vision and medical image analysis tasks including segmentation. However, such methods need large labeled datasets for training, which is hard to obtain for medical image analysis. Self-supervised learning (SSL) has demonstrated success in medical image segmentation using convolutional networks. In this work, we developed a self-distillation learning with masked image modeling method to perform SSL for vision transformers (SMIT) applied to 3D multi-organ segmentation from CT and MRI. Our contribution is a dense pixel-wise regression within masked patches called masked image prediction, which we combined with masked patch token distillation as pretext task to pre-train vision transformers. We show our approach is more accurate and requires fewer fine tuning datasets than other pretext tasks. Unlike prior medical image methods, which typically used image sets arising from disease sites and imaging modalities corresponding to the target tasks, we used 3,643 CT scans (602,708 images) arising from head and neck, lung, and kidney cancers as well as COVID-19 for pre-training and applied it to abdominal organs segmentation from MRI pancreatic cancer patients as well as publicly available 13 different abdominal organs segmentation from CT. Our method showed clear accuracy improvement (average DSC of 0.875 from MRI and 0.878 from CT) with reduced requirement for fine-tuning datasets over commonly used pretext tasks. Extensive comparisons against multiple current SSL methods were done. Code will be made available upon acceptance for publication.

updated: Fri May 20 2022 17:55:14 GMT+0000 (UTC)

published: Fri May 20 2022 17:55:14 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト