FoleyGAN: Visually Guided Generative Adversarial Network-Based Synchronous Sound Generation in Silent Videos

Sanchita Ghose; John J. Prevost

FoleyGAN：サイレントビデオでの視覚的にガイドされた生成的敵対的ネットワークベースの同期サウンド生成

ディープラーニングベースの視覚から音声への生成システムは、基本的に、特に視覚機能と音声機能の時間との同期性の側面を考慮して開発する必要があります。この研究では、視聴覚モダリティ間の同期特性を適応させる視覚から音への生成タスクのためのビデオ入力の時間的視覚情報を使用して、クラス条件付きの生成的敵対的ネットワークを導く新しいタスクを紹介します。私たちが提案するFoleyGANモデルは、視覚的に整列したリアルなサウンドトラックの生成につながる視覚イベントのアクションシーケンスを調整することができます。以前に提案した自動フォーリーデータセットを拡張してFoleyGANでトレーニングし、注目に値する（平均81％）オーディオビジュアルシンクロニシティパフォーマンスを示す人間の調査を通じて合成サウンドを評価します。私たちのアプローチは、他のベースラインモデルや視聴覚データセットと比較して、統計実験でも優れています。

Deep learning based visual to sound generation systems essentially need to be developed particularly considering the synchronicity aspects of visual and audio features with time. In this research we introduce a novel task of guiding a class conditioned generative adversarial network with the temporal visual information of a video input for visual to sound generation task adapting the synchronicity traits between audio-visual modalities. Our proposed FoleyGAN model is capable of conditioning action sequences of visual events leading towards generating visually aligned realistic sound tracks. We expand our previously proposed Automatic Foley dataset to train with FoleyGAN and evaluate our synthesized sound through human survey that shows noteworthy (on average 81%) audio-visual synchronicity performance. Our approach also outperforms in statistical experiments compared with other baseline models and audio-visual datasets.

updated: Tue Jul 20 2021 04:59:26 GMT+0000 (UTC)

published: Tue Jul 20 2021 04:59:26 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト