Intelligent Masking: Deep Q-Learning for Context Encoding in Medical Image Analysis

Mojtaba Bahrami; Mahsa Ghorbani; Nassir Navab

インテリジェントマスキング：医療画像分析におけるコンテキストエンコーディングのための深いQ学習

教師あり設定での大量のラベル付きデータの必要性により、最近の研究では、教師あり学習を利用して、ラベルなしデータを使用してディープニューラルネットワークを事前トレーニングするようになりました。多くの自己監視型トレーニング戦略は、特に医療データセットについて調査されており、ラベルのないデータがはるかに少ない場合に利用可能な情報を活用しています。画像ベースの自己監視の基本的な戦略の1つは、コンテキスト予測です。このアプローチでは、モデルは、周囲に基づいて画像の任意の欠落領域のコンテンツを再構築するようにトレーニングされます。ただし、既存の方法では、画像のすべての領域に均一に焦点を合わせることにより、ランダムでブラインドマスキングのアプローチを採用しています。このアプローチでは、多くの不要なネットワーク更新が発生し、モデルが抽出された豊富な特徴を忘れてしまいます。この作業では、事前トレーニング手順を改善するためにターゲット領域を閉塞する新しい自己監視アプローチを開発します。この目的のために、我々は、深いQ学習を通じて入力画像をインテリジェントにマスクすることを学習する強化学習ベースのエージェントを提案します。予測モデルに対してエージェントをトレーニングすると、ダウンストリーム分類タスク用に抽出されたセマンティック機能を大幅に改善できることを示します。超音波画像で乳がんを診断し、MR画像で低悪性度神経膠腫を検出するための2つの公開データセットで実験を行います。私たちの実験では、新しいマスキング戦略が、精度、マクロF1、およびAUROCの観点から、分類タスクのパフォーマンスに応じて学習した特徴を前進させることを示しています。

The need for a large amount of labeled data in the supervised setting has led recent studies to utilize self-supervised learning to pre-train deep neural networks using unlabeled data. Many self-supervised training strategies have been investigated especially for medical datasets to leverage the information available in the much fewer unlabeled data. One of the fundamental strategies in image-based self-supervision is context prediction. In this approach, a model is trained to reconstruct the contents of an arbitrary missing region of an image based on its surroundings. However, the existing methods adopt a random and blind masking approach by focusing uniformly on all regions of the images. This approach results in a lot of unnecessary network updates that cause the model to forget the rich extracted features. In this work, we develop a novel self-supervised approach that occludes targeted regions to improve the pre-training procedure. To this end, we propose a reinforcement learning-based agent which learns to intelligently mask input images through deep Q-learning. We show that training the agent against the prediction model can significantly improve the semantic features extracted for downstream classification tasks. We perform our experiments on two public datasets for diagnosing breast cancer in the ultrasound images and detecting lower-grade glioma with MR images. In our experiments, we show that our novel masking strategy advances the learned features according to the performance on the classification task in terms of accuracy, macro F1, and AUROC.

updated: Fri Mar 25 2022 19:05:06 GMT+0000 (UTC)

published: Fri Mar 25 2022 19:05:06 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト