IA-RED^2: Interpretability-Aware Redundancy Reduction for Vision Transformers

Bowen Pan; Rameswar Panda; Yifan Jiang; Zhangyang Wang; Rogerio Feris; Aude Oliva

IA-RED ^ 2：ビジョントランスフォーマーの解釈可能性を意識した冗長性の削減

自己注意ベースのモデルであるトランスフォーマーは、最近、コンピュータービジョンの分野における主要なバックボーンになりつつあります。さまざまなビジョンタスクでトランスフォーマーが目覚ましい成功を収めたにもかかわらず、それでも大量の計算と大量のメモリコストに悩まされています。この制限に対処するために、このペーパーでは、解釈可能性を意識した冗長性REDuctionフレームワーク（IA-RED ^ 2）を紹介します。まず、主に無相関の入力パッチに費やされる大量の冗長な計算を観察し、次にこれらの冗長なパッチを動的かつ適切に削除するための解釈可能なモジュールを導入します。次に、この新しいフレームワークは階層構造に拡張され、さまざまな段階で相関のないトークンが徐々に削除されるため、計算コストが大幅に削減されます。画像とビデオの両方のタスクに関する広範な実験が含まれています。この方法では、0.7％未満の精度を犠牲にするだけで、DeiTやTimeSformerなどの最先端モデルで最大1.4倍の速度向上を実現できます。さらに重要なことに、他の加速アプローチとは対照的に、私たちの方法は本質的に実質的な視覚的証拠で解釈可能であり、軽量でありながらビジョントランスフォーマーをより人間が理解できるアーキテクチャに近づけます。私たちのフレームワークで自然に出現した解釈可能性は、定性的および定量的結果の両方で、元の視覚トランスフォーマーによって学習された生の注意、および既製の解釈方法によって生成されたものよりも優れていることを示します。プロジェクトページ：http：//people.csail.mit.edu/bpan/ia-red/。

The self-attention-based model, transformer, is recently becoming the leading backbone in the field of computer vision. In spite of the impressive success made by transformers in a variety of vision tasks, it still suffers from heavy computation and intensive memory costs. To address this limitation, this paper presents an Interpretability-Aware REDundancy REDuction framework (IA-RED^2). We start by observing a large amount of redundant computation, mainly spent on uncorrelated input patches, and then introduce an interpretable module to dynamically and gracefully drop these redundant patches. This novel framework is then extended to a hierarchical structure, where uncorrelated tokens at different stages are gradually removed, resulting in a considerable shrinkage of computational cost. We include extensive experiments on both image and video tasks, where our method could deliver up to 1.4x speed-up for state-of-the-art models like DeiT and TimeSformer, by only sacrificing less than 0.7% accuracy. More importantly, contrary to other acceleration approaches, our method is inherently interpretable with substantial visual evidence, making vision transformer closer to a more human-understandable architecture while being lighter. We demonstrate that the interpretability that naturally emerged in our framework can outperform the raw attention learned by the original visual transformer, as well as those generated by off-the-shelf interpretation methods, with both qualitative and quantitative results. Project Page: http://people.csail.mit.edu/bpan/ia-red/.

updated: Tue Oct 26 2021 22:41:23 GMT+0000 (UTC)

published: Wed Jun 23 2021 18:29:23 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト