Pattern Attention Transformer with Doughnut Kernel

WenYuan Sheng

ドーナツカーネルを使用したパターンアテンショントランスフォーマー

この論文では、新しいドーナツカーネルで構成される新しいアーキテクチャ、パターンアテンショントランスフォーマー (PAT) を紹介します。 NLP 分野のトークンと比較すると、コンピュータービジョンの Transformer は、画像内のピクセルの高解像度を処理するという問題があります。 ViT では、画像は正方形のパッチにカットされます。 ViT のフォローアップとして、Swin Transformer は、モデルの最小単位として「2 つの接続された Swin Transformer ブロック」も発生させる、固定境界の存在を減らすためにシフトする追加のステップを提案します。パッチ/ウィンドウのアイデアを継承したドーナツカーネルは、パッチの設計をさらに強化します。ラインカットの境界を、自己注意の理解に基づくセンサーと更新の 2 種類の領域 (QKVA グリッドと呼ばれる) に置き換えます。ドーナツカーネルは、正方形を超えたカーネルの形状に関する新しいトピックももたらします。画像分類のパフォーマンスを検証するために、PAT は正八角形のドーナツカーネルの Transformer ブロックを使用して設計されています。そのアーキテクチャは軽量です。最小のパターンアテンションレイヤーは各ステージに 1 つだけです。同様の計算の複雑さの下で、ImageNet 1K でのパフォーマンスはより高いスループット (+10%) に達し、Swin Transformer (+0.1 acc1) を上回ります。

We present in this paper a new architecture, the Pattern Attention Transformer (PAT), that is composed of the new doughnut kernel. Compared with tokens in the NLP field, Transformer in computer vision has the problem of handling the high resolution of pixels in images. In ViT, an image is cut into square-shaped patches. As the follow-up of ViT, Swin Transformer proposes an additional step of shifting to decrease the existence of fixed boundaries, which also incurs 'two connected Swin Transformer blocks' as the minimum unit of the model. Inheriting the patch/window idea, our doughnut kernel enhances the design of patches further. It replaces the line-cut boundaries with two types of areas: sensor and updating, which is based on the comprehension of self-attention (named QKVA grid). The doughnut kernel also brings a new topic about the shape of kernels beyond square. To verify its performance on image classification, PAT is designed with Transformer blocks of regular octagon shape doughnut kernels. Its architecture is lighter: the minimum pattern attention layer is only one for each stage. Under similar complexity of computation, its performances on ImageNet 1K reach higher throughput (+10%) and surpass Swin Transformer (+0.1 acc1).

updated: Sun Jan 15 2023 15:58:04 GMT+0000 (UTC)

published: Wed Nov 30 2022 13:11:46 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト