Towards Backdoor Attacks and Defense in Robust Machine Learning Models

Ezekiel Soremekun; Sakshi Udeshi; Sudipta Chattopadhyay

堅牢な機械学習モデルにおけるバックドア攻撃と防御に向けて

堅牢な最適化の導入により、敵対的攻撃に対する防御の最先端が押し上げられました。特に、最先端の射影勾配降下 (PGD) ベースのトレーニング方法は、敵対的な入力に対する防御において普遍的かつ確実に効果的であることが示されています。この堅牢性アプローチでは、PGD を信頼できる普遍的な「一次敵対者」として使用します。ただし、このような最適化の動作は、バックドアと呼ばれる根本的に異なるクラスの攻撃に照らして研究されていません。このホワイトペーパーでは、PGD ベースのロバストな最適化を使用してトレーニングされたロバストなモデルに対して、バックドア攻撃を挿入して防御する方法について説明します。これらのモデルがバックドア攻撃を受けやすいことを示しています。その後、バックドアがそのようなモデルの機能表現に反映されていることがわかりました。次に、この観察結果を活用して、AEGIS と呼ばれる検出技術を使用して、このようなバックドアに感染したモデルを検出します。具体的には、PGD ベースの一次敵対的トレーニングアプローチを使用してトレーニングされた堅牢なディープニューラルネットワーク (DNN) が与えられた場合、AEGIS は特徴クラスタリングを使用して、そのような DNN がバックドアに感染しているかクリーンであるかを効果的に検出します。 CIFAR-10、MNIST、および FMNIST データセットを使用した主要な分類タスクでのいくつかの可視および非表示のバックドアトリガーの評価では、AEGIS は、バックドアに感染した PGD トレーニング済みの堅牢な DNN を効果的に検出します。 AEGIS は、このようなバックドアに感染したモデルを 91.6% の精度 (テストされた 12 モデル中 11 モデル) で検出し、誤検知はありません。さらに、AEGIS は、バックドアに感染したモデルの標的クラスを、かなり低い (11.1%) 誤検知率で検出します。私たちの調査により、敵対的に堅牢な DNN の顕著な特徴が、バックドア攻撃のステルス性を打ち破る可能性があることが明らかになりました。

The introduction of robust optimisation has pushed the state-of-the-art in defending against adversarial attacks. Notably, the state-of-the-art projected gradient descent (PGD)-based training method has been shown to be universally and reliably effective in defending against adversarial inputs. This robustness approach uses PGD as a reliable and universal "first-order adversary". However, the behaviour of such optimisation has not been studied in the light of a fundamentally different class of attacks called backdoors. In this paper, we study how to inject and defend against backdoor attacks for robust models trained using PGD-based robust optimisation. We demonstrate that these models are susceptible to backdoor attacks. Subsequently, we observe that backdoors are reflected in the feature representation of such models. Then, this observation is leveraged to detect such backdoor-infected models via a detection technique called AEGIS. Specifically, given a robust Deep Neural Network (DNN) that is trained using PGD-based first-order adversarial training approach, AEGIS uses feature clustering to effectively detect whether such DNNs are backdoor-infected or clean. In our evaluation of several visible and hidden backdoor triggers on major classification tasks using CIFAR-10, MNIST and FMNIST datasets, AEGIS effectively detects PGD-trained robust DNNs infected with backdoors. AEGIS detects such backdoor-infected models with 91.6% accuracy (11 out of 12 tested models), without any false positives. Furthermore, AEGIS detects the targeted class in the backdoor-infected model with a reasonably low (11.1%) false positive rate. Our investigation reveals that salient features of adversarially robust DNNs could be promising to break the stealthy nature of backdoor attacks.

updated: Wed Jan 11 2023 05:00:29 GMT+0000 (UTC)

published: Tue Feb 25 2020 04:45:26 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト