Is Medical Chest X-ray Data Anonymous?

Kai Packhäuser; Sebastian Gündel; Nicolas Münster; Christopher Syben; Vincent Christlein; Andreas Maier

医療用胸部 X 線データは匿名ですか?

近年の深層学習技術の台頭とますますの可能性により、公的に利用可能な医療データセットは、医療分野での診断アルゴリズムの再現可能な開発を可能にする重要な要素になりました。医療データには患者関連の機密情報が含まれているため、通常、公開前に患者の名前などの患者の識別子を削除して匿名化します。私たちの知る限り、よく訓練された深層学習システムが胸部 X 線データから患者の ID を回復できることを初めて示したのは私たちです。これは、30,805 人の患者からの 112,120 の正面図胸部 X 線画像のコレクションである、公開されている大規模な胸部 X 線 14 データセットを使用してこれを示しています。私たちの検証システムは、AUC 0.9940、分類精度 95.55% で、2 つの正面胸部 X 線画像が同一人物のものであるかどうかを識別できます。さらに、提案されたシステムが最初のスキャンから 10 年以上経っても同じ人物を明らかにできることを強調しています。検索アプローチを追求すると、0.9748 の mAP@R と 0.9963 のprecision@1 が観察されます。さらに、CheXpert と COVID-19 画像データコレクションでトレーニングされたネットワークを評価した場合、最大 0.9870 の AUC と最大 0.9444 の Precision@1 を達成しています。この高い識別率に基づいて、潜在的な攻撃者は患者関連の情報を漏えいし、さらに多くの情報を入手するために画像を相互参照する可能性があります。したがって、機密コンテンツが許可されていない手に渡ったり、関係する患者の意思に反して拡散されたりするリスクが非常に高くなります。特にCOVID-19のパンデミックの間、研究を進めるために多数の胸部X線データセットが公開されました。したがって、そのようなデータは、深層学習ベースの再識別アルゴリズムによる潜在的な攻撃に対して脆弱である可能性があります。

With the rise and ever-increasing potential of deep learning techniques in recent years, publicly available medical datasets became a key factor to enable reproducible development of diagnostic algorithms in the medical domain. Medical data contains sensitive patient-related information and is therefore usually anonymized by removing patient identifiers, e.g., patient names before publication. To the best of our knowledge, we are the first to show that a well-trained deep learning system is able to recover the patient identity from chest X-ray data. We demonstrate this using the publicly available large-scale ChestX-ray14 dataset, a collection of 112,120 frontal-view chest X-ray images from 30,805 unique patients. Our verification system is able to identify whether two frontal chest X-ray images are from the same person with an AUC of 0.9940 and a classification accuracy of 95.55%. We further highlight that the proposed system is able to reveal the same person even ten and more years after the initial scan. When pursuing a retrieval approach, we observe an mAP@R of 0.9748 and a precision@1 of 0.9963. Furthermore, we achieve an AUC of up to 0.9870 and a precision@1 of up to 0.9444 when evaluating our trained networks on CheXpert and the COVID-19 Image Data Collection. Based on this high identification rate, a potential attacker may leak patient-related information and additionally cross-reference images to obtain more information. Thus, there is a great risk of sensitive content falling into unauthorized hands or being disseminated against the will of the concerned patients. Especially during the COVID-19 pandemic, numerous chest X-ray datasets have been published to advance research. Therefore, such data may be vulnerable to potential attacks by deep learning-based re-identification algorithms.

updated: Mon May 31 2021 17:22:04 GMT+0000 (UTC)

published: Mon Mar 15 2021 17:26:43 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト