FTFDNet: Learning to Detect Talking Face Video Manipulation with Tri-Modality Interaction

Ganglai Wang; Peng Zhang; Junwen Xiong; Feihan Yang; Wei Huang; Yufei Zha

FTFDNet: トライモダリティインタラクションによる話し顔ビデオ操作の検出方法の学習

DeepFake ベースのデジタル顔偽造は、特に話し顔の生成に口唇操作が使用されている場合、公共メディアのセキュリティを脅かしており、偽ビデオ検出の難易度はさらに向上しています。所定の音声に合わせて唇の形を変えるだけで、このような偽の話している顔のビデオでは、身元の顔の特徴を区別するのが困難になります。事前知識としての音声ストリームへの注意の欠如と合わせて、偽の話し顔ビデオの検出失敗も避けられなくなります。実際のビデオのオプティカルフローが規則的に変化するのに対し、偽の話し顔ビデオのオプティカルフローは特に唇の領域で乱れていることがわかりました。これは、オプティカルフローからの動きの特徴が操作の手がかりを捉えるのに役立つことを意味します。この研究では、効率的なクロスモーダルフュージョン (CMF) モジュールを使用して、視覚、音声、および動きの特徴を組み込むことにより、偽会話顔検出ネットワーク (FTFDNet) を提案します。さらに、より有益な機能を発見するために、新しいオーディオビジュアルアテンションメカニズム (AVAM) が提案されており、モジュール化によって任意のオーディオビジュアル CNN アーキテクチャにシームレスに統合できます。追加の AVAM により、提案された FTFDNet は、確立された偽の話し顔検出データセット (FTFDD) だけでなく、ディープフェイクビデオ検出データセットでも、他の最先端のディープフェイクビデオ検出方法よりも優れた検出パフォーマンスを達成できます。 (DFDC および DF-TIMIT)。

DeepFake based digital facial forgery is threatening public media security, especially when lip manipulation has been used in talking face generation, and the difficulty of fake video detection is further improved. By only changing lip shape to match the given speech, the facial features of identity are hard to be discriminated in such fake talking face videos. Together with the lack of attention on audio stream as the prior knowledge, the detection failure of fake talking face videos also becomes inevitable. It's found that the optical flow of the fake talking face video is disordered especially in the lip region while the optical flow of the real video changes regularly, which means the motion feature from optical flow is useful to capture manipulation cues. In this study, a fake talking face detection network (FTFDNet) is proposed by incorporating visual, audio and motion features using an efficient cross-modal fusion (CMF) module. Furthermore, a novel audio-visual attention mechanism (AVAM) is proposed to discover more informative features, which can be seamlessly integrated into any audio-visual CNN architecture by modularization. With the additional AVAM, the proposed FTFDNet is able to achieve a better detection performance than other state-of-the-art DeepFake video detection methods not only on the established fake talking face detection dataset (FTFDD) but also on the DeepFake video detection datasets (DFDC and DF-TIMIT).

updated: Sat Jul 08 2023 14:45:16 GMT+0000 (UTC)

published: Sat Jul 08 2023 14:45:16 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト