Exposing Backdoors in Robust Machine Learning Models

Ezekiel Soremekun; Sakshi Udeshi; Sudipta Chattopadhyay

堅牢な機械学習モデルでバックドアを公開する

堅牢な最適化の導入により、敵対的攻撃に対する防御の最先端が推進されました。ただし、このような最適化の動作は、バックドアと呼ばれる根本的に異なるクラスの攻撃に照らして研究されていません。このホワイトペーパーでは、敵対的に堅牢なモデルがバックドア攻撃の影響を受けやすいことを示します。その後、バックドアがそのようなモデルの特徴表現に反映されていることを確認しました。次に、この観察結果を利用して、AEGIS と呼ばれる検出手法を使用して、バックドアに感染したモデルを検出します。具体的には、AEGIS は機能クラスタリングを使用して、バックドアに感染した堅牢なディープニューラルネットワーク (DNN) を効果的に検出します。 CIFAR-10、MNIST、FMNIST データセットを使用した主要な分類タスクでのいくつかの可視および非表示のバックドアトリガーの評価において、AEGIS はバックドアに感染した堅牢な DNN を効果的に検出します。 AEGIS は、誤検知なしで、91.6% の精度でバックドアに感染したモデルを検出します。さらに、AEGIS は、バックドアに感染したモデルの標的クラスを、かなり低い (11.1%) 偽陽性率で検出します。私たちの調査により、敵対的に堅牢な DNN の顕著な特徴が、バックドア攻撃のステルス性を打破することが明らかになりました。

The introduction of robust optimisation has pushed the state-of-the-art in defending against adversarial attacks. However, the behaviour of such optimisation has not been studied in the light of a fundamentally different class of attacks called backdoors. In this paper, we demonstrate that adversarially robust models are susceptible to backdoor attacks. Subsequently, we observe that backdoors are reflected in the feature representation of such models. Then, this observation is leveraged to detect backdoor-infected models via a detection technique called AEGIS. Specifically, AEGIS uses feature clustering to effectively detect backdoor-infected robust Deep Neural Networks (DNNs). In our evaluation of several visible and hidden backdoor triggers on major classification tasks using CIFAR-10, MNIST and FMNIST datasets, AEGIS effectively detects robust DNNs infected with backdoors. AEGIS detects a backdoor-infected model with 91.6% accuracy, without any false positives. Furthermore, AEGIS detects the targeted class in the backdoor-infected model with a reasonably low (11.1%) false positive rate. Our investigation reveals that salient features of adversarially robust DNNs break the stealthy nature of backdoor attacks.

updated: Thu Jun 03 2021 07:02:14 GMT+0000 (UTC)

published: Tue Feb 25 2020 04:45:26 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト