Analyzing Adversarial Robustness of Vision Transformers against Spatial and Spectral Attacks

Gihyun Kim; Jong-Seok Lee

空間およびスペクトル攻撃に対するビジョントランスフォーマーの敵対的ロバスト性の分析

ビジョントランスフォーマーは、画像分類タスクで畳み込みニューラルネットワーク (CNN) を凌駕できる強力なアーキテクチャとして登場しました。敵対的攻撃に対する Transformer の堅牢性を理解するためにいくつかの試みが行われましたが、既存の研究では一貫性のない結果が得られています。つまり、Transformer は CNN よりも堅牢であると結論付けているものもあれば、同程度の堅牢性を持っていると結論付けているものもあります。このホワイトペーパーでは、トランスフォーマーの敵対的ロバスト性を調べる既存の研究では調査されていない 2 つの問題に対処します。まず、敵対的ロバスト性を評価する際に画質を同時に考慮する必要があると主張します。堅牢性の観点から、あるアーキテクチャが別のアーキテクチャよりも優れているかどうかは、攻撃された画像の品質によって表される攻撃の強さに応じて変化する可能性があることがわかりました。次に、トランスフォーマーと CNN が画像内のさまざまな種類の情報に依存していることに注目することで、柔軟な攻撃を実装するためのツールとして、フーリエ攻撃と呼ばれる攻撃フレームワークを策定します。空間ドメイン。この攻撃は、特定の周波数成分の振幅と位相情報を選択的に乱します。大規模な実験を通じて、トランスフォーマーは CNN よりも位相情報と低周波情報に依存する傾向があるため、周波数選択攻撃に対してさらに脆弱になる場合があることがわかりました。この作業が、トランスフォーマーの特性と敵対的堅牢性を理解する上で新しい視点を提供することを願っています。

Vision Transformers have emerged as a powerful architecture that can outperform convolutional neural networks (CNNs) in image classification tasks. Several attempts have been made to understand robustness of Transformers against adversarial attacks, but existing studies draw inconsistent results, i.e., some conclude that Transformers are more robust than CNNs, while some others find that they have similar degrees of robustness. In this paper, we address two issues unexplored in the existing studies examining adversarial robustness of Transformers. First, we argue that the image quality should be simultaneously considered in evaluating adversarial robustness. We find that the superiority of one architecture to another in terms of robustness can change depending on the attack strength expressed by the quality of the attacked images. Second, by noting that Transformers and CNNs rely on different types of information in images, we formulate an attack framework, called Fourier attack, as a tool for implementing flexible attacks, where an image can be attacked in the spectral domain as well as in the spatial domain. This attack perturbs the magnitude and phase information of particular frequency components selectively. Through extensive experiments, we find that Transformers tend to rely more on phase information and low frequency information than CNNs, and thus sometimes they are even more vulnerable under frequency-selective attacks. It is our hope that this work provides new perspectives in understanding the properties and adversarial robustness of Transformers.

updated: Sat Aug 20 2022 04:14:27 GMT+0000 (UTC)

published: Sat Aug 20 2022 04:14:27 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト