Facial Expression Recognition with Visual Transformers and Attentional Selective Fusion

Fuyan Ma; Bin Sun; Shutao Li

視覚変換器と注意選択的融合による顔の表情認識

野生の顔の表情認識（FER）は、拘束されていない条件下でのオクルージョン、さまざまな頭のポーズ、顔の変形、モーションブラーのために非常に困難です。過去数十年で自動FERは大幅に進歩しましたが、以前の研究は主にラボ制御のFER用に設計されていました。現実世界のオクルージョン、異形の頭のポーズ、およびその他の問題は、これらの情報が不足している領域と複雑な背景のために、FERの難易度を確実に高めます。以前の純粋なCNNベースの方法とは異なり、顔の画像を視覚的な単語のシーケンスに変換し、グローバルな視点から表現認識を実行することが実行可能で実用的であると主張します。したがって、2つの主要なステップで野生のFERに取り組むために、機能融合を備えたビジュアルトランスフォーマー（VTFF）を提案します。まず、2分岐CNNによって生成された2種類の特徴マップを活用するための注意選択融合（ASF）を提案します。 ASFは、グローバルローカルの注意を払って複数の機能を融合することにより、識別情報をキャプチャします。次に、融合された特徴マップが平坦化され、一連の視覚的な単語に投影されます。第二に、自然言語処理におけるトランスフォーマーの成功に触発されて、これらの視覚的な単語間の関係をグローバルな自己注意でモデル化することを提案します。提案された方法は、3つの公開されている野生の顔の表情データセット（RAF-DB、FERPlus、およびAffectNet）で評価されます。同じ設定の下で、広範な実験により、私たちの方法が他の方法よりも優れたパフォーマンスを示し、RAF-DBが88.14％、FERPlusが88.81％、AffectNetが61.85％という新しい最先端の設定を示しています。 CK +のクロスデータセット評価は、提案された方法の有望な一般化機能を示しています。

Facial Expression Recognition (FER) in the wild is extremely challenging due to occlusions, variant head poses, face deformation and motion blur under unconstrained conditions. Although substantial progresses have been made in automatic FER in the past few decades, previous studies were mainly designed for lab-controlled FER. Real-world occlusions, variant head poses and other issues definitely increase the difficulty of FER on account of these information-deficient regions and complex backgrounds. Different from previous pure CNNs based methods, we argue that it is feasible and practical to translate facial images into sequences of visual words and perform expression recognition from a global perspective. Therefore, we propose the Visual Transformers with Feature Fusion (VTFF) to tackle FER in the wild by two main steps. First, we propose the attentional selective fusion (ASF) for leveraging two kinds of feature maps generated by two-branch CNNs. The ASF captures discriminative information by fusing multiple features with the global-local attention. The fused feature maps are then flattened and projected into sequences of visual words. Second, inspired by the success of Transformers in natural language processing, we propose to model relationships between these visual words with the global self-attention. The proposed method is evaluated on three public in-the-wild facial expression datasets (RAF-DB, FERPlus and AffectNet). Under the same settings, extensive experiments demonstrate that our method shows superior performance over other methods, setting new state of the art on RAF-DB with 88.14%, FERPlus with 88.81% and AffectNet with 61.85%. The cross-dataset evaluation on CK+ shows the promising generalization capability of the proposed method.

updated: Tue Feb 22 2022 01:51:36 GMT+0000 (UTC)

published: Wed Mar 31 2021 07:07:56 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト