Progressive Co-Attention Network for Fine-grained Visual Classification

Tian Zhang; Dongliang Chang; Zhanyu Ma; Jun Guo

きめ細かい視覚分類のための進歩的な共同注意ネットワーク

きめ細かい視覚的分類は、同じカテゴリ内の複数のサブカテゴリに属する画像を認識することを目的としています。非常に混乱したカテゴリー間で本質的に微妙な違いがあるため、これは困難な作業です。ほとんどの既存の方法は、入力として個々の画像のみを取得するため、モデルが異なる画像からの対照的な手がかりを認識する能力が制限される可能性があります。本論文では、この問題に取り組むために、プログレッシブ共注意ネットワーク（PCA-Net）と呼ばれる効果的な方法を提案します。具体的には、同じカテゴリの画像ペア内の特徴チャネル間の相互作用を促進して、共通の識別可能な特徴をキャプチャすることにより、チャネルごとの類似性を計算します。補完的な情報も認識に不可欠であることを考慮して、チャネルの相互作用によって強化された顕著な領域を消去して、ネットワークを他の識別領域に集中させます。提案されたモデルは、CUB-200-2011、Stanford Cars、およびFGVCAircraftの3つのきめ細かい視覚分類ベンチマークデータセットで競争力のある結果を達成しています。

Fine-grained visual classification aims to recognize images belonging to multiple sub-categories within a same category. It is a challenging task due to the inherently subtle variations among highly-confused categories. Most existing methods only take an individual image as input, which may limit the ability of models to recognize contrastive clues from different images. In this paper, we propose an effective method called progressive co-attention network (PCA-Net) to tackle this problem. Specifically, we calculate the channel-wise similarity by encouraging interaction between the feature channels within same-category image pairs to capture the common discriminative features. Considering that complementary information is also crucial for recognition, we erase the prominent areas enhanced by the channel interaction to force the network to focus on other discriminative regions. The proposed model has achieved competitive results on three fine-grained visual classification benchmark datasets: CUB-200-2011, Stanford Cars, and FGVC Aircraft.

updated: Mon Aug 30 2021 16:38:12 GMT+0000 (UTC)

published: Thu Jan 21 2021 10:19:02 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト