Odyssey: Creation, Analysis and Detection of Trojan Models

Marzieh Edraki; Nazmul Karim; Nazanin Rahnavard; Ajmal Mian; Mubarak Shah

オデッセイ：トロイの木馬モデルの作成、分析、検出

ディープニューラルネットワーク（DNN）モデルの成功とともに、これらのモデルの整合性に対する脅威が高まります。最近の脅威は、攻撃者が一部のトレーニングサンプルにトリガーを挿入することでトレーニングパイプラインに干渉し、トリガーを含むサンプルに対してのみ悪意を持って動作するようにモデルをトレーニングするトロイの木馬攻撃です。トリガーの知識は攻撃者に知られているため、トロイの木馬ネットワークの検出は困難です。既存のトロイの木馬検出器は、トリガーと攻撃の種類について強い仮定を立てています。固有のDNNプロパティの分析に基づく検出器を提案します。トロイの木馬プロセスによって影響を受けます。包括的な分析のために、3,000を超えるクリーンなトロイの木馬モデルを備えたこれまでで最も多様なデータセットであるOdysseusを開発します。オデュッセウスは広範囲の攻撃をカバーしています。トリガーデザインとソースからターゲットへのクラスマッピングの多様性を活用して生成されます。私たちの分析結果は、トロイの木馬攻撃が、クリーンなデータの多様体の周りの分類器のマージンと決定境界の形状に影響を与えることを示しています。これら2つの要素を悪用して、攻撃の知識がなくても動作し、既存の方法を大幅に上回る効率的なトロイの木馬検出器を提案します。包括的な一連の実験を通じて、クロスモデルアーキテクチャ、見えないトリガー、および正規化されたモデルに対する検出器の有効性を示します。

Along with the success of deep neural network (DNN) models, rise the threats to the integrity of these models. A recent threat is the Trojan attack where an attacker interferes with the training pipeline by inserting triggers into some of the training samples and trains the model to act maliciously only for samples that contain the trigger. Since the knowledge of triggers is privy to the attacker, detection of Trojan networks is challenging. Existing Trojan detectors make strong assumptions about the types of triggers and attacks. We propose a detector that is based on the analysis of the intrinsic DNN properties; that are affected due to the Trojaning process. For a comprehensive analysis, we develop Odysseus, the most diverse dataset to date with over 3,000 clean and Trojan models. Odysseus covers a large spectrum of attacks; generated by leveraging the versatility in trigger designs and source to target class mappings. Our analysis results show that Trojan attacks affect the classifier margin and shape of decision boundary around the manifold of clean data. Exploiting these two factors, we propose an efficient Trojan detector that operates without any knowledge of the attack and significantly outperforms existing methods. Through a comprehensive set of experiments we demonstrate the efficacy of the detector on cross model architectures, unseen Triggers and regularized models.

updated: Tue Dec 08 2020 08:09:51 GMT+0000 (UTC)

published: Thu Jul 16 2020 06:55:00 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト