Pattern Attention Transformer with Doughnut Kernel

WenYuan Sheng

ドーナツカーネルを使用したパターンアテンショントランスフォーマー

この論文では、新しいドーナツカーネルで構成される新しいアーキテクチャであるパターンアテンショントランスフォーマー (PAT) を紹介します。 NLP 分野のトークンと比較して、コンピュータービジョンの Transformer には、画像内のピクセルの高解像度を処理するという問題があります。 ViT では、画像が正方形のパッチに切り出されます。 ViT のフォローアップとして、Swin Transformer は、固定境界の存在を減らすためにシフトする追加のステップを提案しています。これにより、モデルの最小単位として「2 つの接続された Swin Transformer ブロック」も発生します。パッチ/ウィンドウのアイデアを継承するドーナツカーネルは、パッチのデザインをさらに強化します。これは、ラインカット境界を 2 種類の領域に置き換えます。センサーと、セルフアテンション (QKVA グリッドと呼ばれる) の理解に基づく更新です。ドーナツの核は、正方形を超えた核の形状に関する新しいトピックももたらします。画像分類におけるパフォーマンスを検証するために、PAT は正八角形のドーナツカーネルの Transformer ブロックを使用して設計されています。そのアーキテクチャは軽量です。最小のパターンアテンションレイヤーは各ステージに対して 1 つだけです。同様の複雑な計算条件下で、ImageNet 1K のパフォーマンスはより高いスループット (+10%) に達し、Swin Transformer (+0.8 acc1) を上回ります。

We present in this paper a new architecture, the Pattern Attention Transformer (PAT), that is composed of the new doughnut kernel. Compared with tokens in the NLP field, Transformer in computer vision has the problem of handling the high resolution of pixels in images. In ViT, an image is cut into square-shaped patches. As the follow-up of ViT, Swin Transformer proposes an additional step of shifting to decrease the existence of fixed boundaries, which also incurs 'two connected Swin Transformer blocks' as the minimum unit of the model. Inheriting the patch/window idea, our doughnut kernel enhances the design of patches further. It replaces the line-cut boundaries with two types of areas: sensor and updating, which is based on the comprehension of self-attention (named QKVA grid). The doughnut kernel also brings a new topic about the shape of kernels beyond square. To verify its performance on image classification, PAT is designed with Transformer blocks of regular octagon shape doughnut kernels. Its architecture is lighter: the minimum pattern attention layer is only one for each stage. Under similar complexity of computation, its performances on ImageNet 1K reach higher throughput (+10%) and surpass Swin Transformer (+0.8 acc1).

updated: Sun Sep 17 2023 13:43:53 GMT+0000 (UTC)

published: Wed Nov 30 2022 13:11:46 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト