GLiT: Neural Architecture Search for Global and Local Image Transformer

Boyu Chen; Peixia Li; Chuming Li; Baopu Li; Lei Bai; Chen Lin; Ming Sun; Junjie yan; Wanli Ouyang

GLiT：グローバルおよびローカルImageTransformerのニューラルアーキテクチャ検索

画像認識のためのより良いトランスフォーマーアーキテクチャを見つけるための最初のニューラルアーキテクチャ検索（NAS）メソッドを紹介します。最近、CNNベースのバックボーンを持たないトランスフォーマーが画像認識の優れたパフォーマンスを実現することがわかりました。ただし、トランスフォーマーはNLPタスク用に設計されているため、画像認識に直接使用すると最適ではない可能性があります。変圧器の視覚的表現能力を向上させるために、新しい探索空間と探索アルゴリズムを提案します。具体的には、より少ない計算コストで画像の局所相関を明示的にモデル化する局所性モジュールを導入します。ローカリティモジュールを使用すると、検索アルゴリズムがグローバル情報とローカル情報の間で自由にトレードオフできるように検索スペースが定義され、各モジュールの低レベルの設計選択が最適化されます。巨大な探索空間によって引き起こされる問題に取り組むために、進化的アルゴリズムを用いて2つのレベルから別々に最適なビジョントランスフォーマーを探索するための階層的ニューラルアーキテクチャ探索法が提案されています。 ImageNetデータセットでの広範な実験は、私たちの方法がResNetファミリー（ResNet101など）や画像分類のベースラインViTよりも識別力があり効率的なトランスフォーマーバリアントを見つけることができることを示しています。

We introduce the first Neural Architecture Search (NAS) method to find a better transformer architecture for image recognition. Recently, transformers without CNN-based backbones are found to achieve impressive performance for image recognition. However, the transformer is designed for NLP tasks and thus could be sub-optimal when directly used for image recognition. In order to improve the visual representation ability for transformers, we propose a new search space and searching algorithm. Specifically, we introduce a locality module that models the local correlations in images explicitly with fewer computational cost. With the locality module, our search space is defined to let the search algorithm freely trade off between global and local information as well as optimizing the low-level design choice in each module. To tackle the problem caused by huge search space, a hierarchical neural architecture search method is proposed to search the optimal vision transformer from two levels separately with the evolutionary algorithm. Extensive experiments on the ImageNet dataset demonstrate that our method can find more discriminative and efficient transformer variants than the ResNet family (e.g., ResNet101) and the baseline ViT for image classification.

updated: Sat Aug 07 2021 10:34:23 GMT+0000 (UTC)

published: Wed Jul 07 2021 00:48:09 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト