3D to 4D Facial Expressions Generation Guided by Landmarks

Naima Otberdout; Claudio Ferrari; Mohamed Daoudi; Stefano Berretti; Alberto Del Bimbo

ランドマークに導かれた3Dから4Dの顔の表情の生成

ディープラーニングベースの3D顔生成は最近進歩しましたが、動的3D（4D）顔の表情合成の問題はあまり調査されていません。この論文では、次の質問に対する新しい解決策を提案します。1つの入力3Dニュートラルな顔が与えられた場合、それから動的な3D（4D）の表情を生成できますか？この問題に取り組むために、まず、3Dランドマークのセットを活用してニュートラルな対応物から表現力豊かな3D顔を生成するメッシュエンコーダ-デコーダアーキテクチャ（Expr-ED）を提案します。次に、表情ラベル（Motion3DGAN）から一連の3Dランドマークを生成できる多様体値GANを使用して、顔の表情の時間的ダイナミクスをモデル化することにより、それを4Dに拡張します。生成されたランドマークはメッシュエンコーダ-デコーダに送られ、最終的に一連の3D表現力豊かな顔を生成します。 2つのステップを分離することにより、メッシュの変形とモーションダイナミクスによって引き起こされる非線形性に個別に対処します。 CoMAデータセットの実験結果は、ランドマークによってガイドされるメッシュエンコーダーデコーダーが他のランドマークベースの3Dフィッティングアプローチと比較して大幅な改善をもたらし、高品質の動的な表情を生成できることを示しています。このフレームワークにより、3D発現強度を低強度から高強度まで継続的に適応させることができます。最後に、フレームワークを2D-3D表情転送などの他のタスクに適用できることを示します。

While deep learning-based 3D face generation has made a progress recently, the problem of dynamic 3D (4D) facial expression synthesis is less investigated. In this paper, we propose a novel solution to the following question: given one input 3D neutral face, can we generate dynamic 3D (4D) facial expressions from it? To tackle this problem, we first propose a mesh encoder-decoder architecture (Expr-ED) that exploits a set of 3D landmarks to generate an expressive 3D face from its neutral counterpart. Then, we extend it to 4D by modeling the temporal dynamics of facial expressions using a manifold-valued GAN capable of generating a sequence of 3D landmarks from an expression label (Motion3DGAN). The generated landmarks are fed into the mesh encoder-decoder, ultimately producing a sequence of 3D expressive faces. By decoupling the two steps, we separately address the non-linearity induced by the mesh deformation and motion dynamics. The experimental results on the CoMA dataset show that our mesh encoder-decoder guided by landmarks brings a significant improvement with respect to other landmark-based 3D fitting approaches, and that we can generate high quality dynamic facial expressions. This framework further enables the 3D expression intensity to be continuously adapted from low to high intensity. Finally, we show our framework can be applied to other tasks, such as 2D-3D facial expression transfer.

updated: Sun May 16 2021 15:52:29 GMT+0000 (UTC)

published: Sun May 16 2021 15:52:29 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト