Robust Facial Expression Recognition with Convolutional Visual Transformers

Fuyan Ma; Bin Sun; Shutao Li

畳み込み視覚変換器によるロバストな表情認識

野生の顔の表情認識（FER）は、拘束されていない条件下でのオクルージョン、さまざまな頭のポーズ、顔の変形、モーションブラーのために非常に困難です。過去数十年で自動FERは大幅に進歩しましたが、以前の研究は主に実験室で制御されるFER向けに設計されています。現実世界のオクルージョン、異形の頭のポーズ、およびその他の問題は、これらの情報が不足している領域と複雑な背景のために、FERの難易度を確実に高めます。以前の純粋なCNNベースの方法とは異なり、顔の画像を視覚的な単語のシーケンスに変換し、グローバルな視点から表現認識を実行することが実行可能で実用的であると主張します。したがって、2つの主要なステップで野生のFERに取り組むための畳み込みビジュアルトランスフォーマーを提案します。まず、2分岐CNNによって生成された特徴マップを活用するための注意選択的融合（ASF）を提案します。 ASFは、グローバルローカルの注意を払って複数の機能を融合することにより、識別情報をキャプチャします。次に、融合された特徴マップが平坦化され、一連の視覚的な単語に投影されます。次に、自然言語処理におけるTransformersの成功に触発されて、これらの視覚的な単語間の関係をグローバルな自己注意でモデル化することを提案します。提案された方法は、3つの公開されている野生の表情データセット（RAF-DB、FERPlus、およびAffectNet）で評価されます。同じ設定の下で、広範な実験により、私たちの方法が他の方法よりも優れたパフォーマンスを示し、RAF-DBで88.14％、FERPlusで88.81％、AffectNetで61.85％の新しい最先端を設定していることが示されています。また、CK +のクロスデータセット評価を実施し、提案された方法の一般化能力を示します。

Facial Expression Recognition (FER) in the wild is extremely challenging due to occlusions, variant head poses, face deformation and motion blur under unconstrained conditions. Although substantial progresses have been made in automatic FER in the past few decades, previous studies are mainly designed for lab-controlled FER. Real-world occlusions, variant head poses and other issues definitely increase the difficulty of FER on account of these information-deficient regions and complex backgrounds. Different from previous pure CNNs based methods, we argue that it is feasible and practical to translate facial images into sequences of visual words and perform expression recognition from a global perspective. Therefore, we propose Convolutional Visual Transformers to tackle FER in the wild by two main steps. First, we propose an attentional selective fusion (ASF) for leveraging the feature maps generated by two-branch CNNs. The ASF captures discriminative information by fusing multiple features with global-local attention. The fused feature maps are then flattened and projected into sequences of visual words. Second, inspired by the success of Transformers in natural language processing, we propose to model relationships between these visual words with global self-attention. The proposed method are evaluated on three public in-the-wild facial expression datasets (RAF-DB, FERPlus and AffectNet). Under the same settings, extensive experiments demonstrate that our method shows superior performance over other methods, setting new state of the art on RAF-DB with 88.14%, FERPlus with 88.81% and AffectNet with 61.85%. We also conduct cross-dataset evaluation on CK+ show the generalization capability of the proposed method.

updated: Sun May 23 2021 03:41:03 GMT+0000 (UTC)

published: Wed Mar 31 2021 07:07:56 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト