Efficient Robustness Assessment via Adversarial Spatial-Temporal Focus on Videos

Wei Xingxing; Wang Songping; Yan Huanqian

ビデオに対する敵対的時空間フォーカスによる効率的な堅牢性評価

ビデオ認識モデルの敵対的ロバスト性評価は、安全性が重要なタスクに幅広く適用されているため、懸念が生じています。画像と比較して、ビデオは非常に高次元であるため、敵対的なビデオを生成する際に膨大な計算コストが発生します。これは、脅威モデルの勾配推定が通常利用されるクエリベースのブラックボックス攻撃では特に深刻であり、高次元は多数のクエリにつながります。この問題を軽減するために、ビデオ内の時間的および空間的な冗長性を同時に排除して、削減された検索スペースで効果的かつ効率的な勾配推定を実現することを提案します。これにより、クエリ数が減少する可能性があります。このアイデアを実装するために、ビデオに対する新しい敵対的時空間フォーカス (AstFocus) 攻撃を設計します。これは、ビデオのフレーム間およびフレーム内から同時にフォーカスされたキーフレームとキー領域に対して攻撃を実行します。 AstFocus 攻撃は、協調型マルチエージェント強化学習 (MARL) フレームワークに基づいています。 1 つのエージェントがキーフレームの選択を担当し、別のエージェントがキー領域の選択を担当します。これら 2 つのエージェントは、ブラックボックスの脅威モデルから受け取った共通の報酬によって共同でトレーニングされ、協調的な予測を実行します。継続的にクエリを実行することで、キーフレームとキー領域で構成される縮小された検索空間が正確になり、全体のクエリ数が元のビデオよりも少なくなります。 4 つの主流のビデオ認識モデルと 3 つの広く使用されているアクション認識データセットに関する広範な実験は、提案された AstFocus 攻撃が SOTA メソッドよりも優れていることを示しています。

Adversarial robustness assessment for video recognition models has raised concerns owing to their wide applications on safety-critical tasks. Compared with images, videos have much high dimension, which brings huge computational costs when generating adversarial videos. This is especially serious for the query-based black-box attacks where gradient estimation for the threat models is usually utilized, and high dimensions will lead to a large number of queries. To mitigate this issue, we propose to simultaneously eliminate the temporal and spatial redundancy within the video to achieve an effective and efficient gradient estimation on the reduced searching space, and thus query number could decrease. To implement this idea, we design the novel Adversarial spatial-temporal Focus (AstFocus) attack on videos, which performs attacks on the simultaneously focused key frames and key regions from the inter-frames and intra-frames in the video. AstFocus attack is based on the cooperative Multi-Agent Reinforcement Learning (MARL) framework. One agent is responsible for selecting key frames, and another agent is responsible for selecting key regions. These two agents are jointly trained by the common rewards received from the black-box threat models to perform a cooperative prediction. By continuously querying, the reduced searching space composed of key frames and key regions is becoming precise, and the whole query number becomes less than that on the original video. Extensive experiments on four mainstream video recognition models and three widely used action recognition datasets demonstrate that the proposed AstFocus attack outperforms the SOTA methods, which is prevenient in fooling rate, query number, time, and perturbation magnitude at the same.

updated: Mon Mar 27 2023 01:57:56 GMT+0000 (UTC)

published: Tue Jan 03 2023 00:28:57 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト