Medical Image Segmentation Using Squeeze-and-Expansion Transformers

Shaohua Li; Xiuchao Sui; Xiangde Luo; Xinxing Xu; Yong Liu; Rick Goh

スクイーズアンドエクスパンショントランスフォーマーを使用した医療画像のセグメンテーション

医療画像のセグメンテーションは、コンピューター支援診断にとって重要です。優れたセグメンテーションでは、モデルが全体像と詳細を同時に見る必要があります。つまり、高い空間解像度を維持しながら、大きなコンテキストを組み込んだ画像の特徴を学習する必要があります。この目標に近づくために、最も広く使用されている方法 - U-Net とそのバリアントは、マルチスケールの特徴を抽出して融合します。ただし、融合された機能には、ローカル画像の手がかりに焦点を当てた小さな「有効受容野」がまだあり、パフォーマンスが制限されています。この作業では、トランスフォーマーに基づく代替セグメンテーションフレームワークである Segtran を提案します。これは、高解像度でも無制限の「有効受容野」を持つものです。 Segtran のコアは、新しい Squeeze-and-Expansion トランスフォーマーです。絞ったアテンションブロックはトランスフォーマーのセルフアテンションを正規化し、エクスパンションブロックは多様な表現を学習します。さらに、画像に連続性誘導バイアスを課す、トランスフォーマーの新しい位置エンコード方式を提案します。 2D および 3D 医用画像セグメンテーションタスクで実験が行われました。眼底画像の視神経乳頭/カップのセグメンテーション (REFUGE'20 チャレンジ)、結腸内視鏡画像のポリープセグメンテーション、MRI スキャンの脳腫瘍のセグメンテーション (BraTS'19 チャレンジ)代表的な既存の方法と比較して、Segtran は一貫して最高のセグメンテーション精度を達成し、優れたクロスドメイン汎化機能を示しました。 Segtran のソースコードは https://github.com/askerlee/segtran で公開されています。

Medical image segmentation is important for computer-aided diagnosis. Good segmentation demands the model to see the big picture and fine details simultaneously, i.e., to learn image features that incorporate large context while keep high spatial resolutions. To approach this goal, the most widely used methods -- U-Net and variants, extract and fuse multi-scale features. However, the fused features still have small "effective receptive fields" with a focus on local image cues, limiting their performance. In this work, we propose Segtran, an alternative segmentation framework based on transformers, which have unlimited "effective receptive fields" even at high feature resolutions. The core of Segtran is a novel Squeeze-and-Expansion transformer: a squeezed attention block regularizes the self attention of transformers, and an expansion block learns diversified representations. Additionally, we propose a new positional encoding scheme for transformers, imposing a continuity inductive bias for images. Experiments were performed on 2D and 3D medical image segmentation tasks: optic disc/cup segmentation in fundus images (REFUGE'20 challenge), polyp segmentation in colonoscopy images, and brain tumor segmentation in MRI scans (BraTS'19 challenge). Compared with representative existing methods, Segtran consistently achieved the highest segmentation accuracy, and exhibited good cross-domain generalization capabilities. The source code of Segtran is released at https://github.com/askerlee/segtran.

updated: Wed Jun 02 2021 02:42:19 GMT+0000 (UTC)

published: Thu May 20 2021 04:45:47 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト