DanceIt: Music-inspired Dancing Video Synthesis

Xin Guo; Yifan Zhao; Jia Li

DanceIt：音楽にインスパイアされたダンスビデオ合成

目を閉じて音楽を聴くと、俳優が音楽に合わせてリズミカルに踊るのを簡単に想像できます。これらのダンスの動きは通常、以前に見たダンスの動きで構成されています。この論文では、コンピュータビジョンシステム内に人間のこうした固有の能力を再現することを提案します。提案されたシステムは、3つのモジュールで構成されています。音楽とダンスの動きの関係を調査するために、事前設計された音楽に伴うダンスビデオクリップに焦点を合わせたクロスモーダルアライメントモジュールを提案し、ポーズシーケンスの視覚的特徴と音楽の音響機能。次に、学習したモデルをイマジネーションモジュールで使用して、特定の音楽のポーズシーケンスを選択します。しかしながら、音楽から選択されたそのようなポーズシーケンスは、通常、不連続です。この問題を解決するために、時空間アライメントモジュールで、ダンスの動きの傾向と周期性に基づいて空間アライメントアルゴリズムを開発し、不連続な断片間のダンスの動きを予測します。さらに、選択したポーズシーケンスは、多くの場合、音楽のビートとずれています。この問題を解決するために、音楽とダンスのリズムを調整する時間調整アルゴリズムをさらに開発します。最後に、処理されたポーズシーケンスを使用して、想像力モジュールでリアルなダンスビデオを合成します。生成されたダンスビデオは、音楽の内容とリズムに一致します。実験結果と主観的評価は、提案されたアプローチが音楽を入力することによって有望なダンスビデオを生成する機能を実行できることを示しています。

Close your eyes and listen to music, one can easily imagine an actor dancing rhythmically along with the music. These dance movements are usually made up of dance movements you have seen before. In this paper, we propose to reproduce such an inherent capability of the human-being within a computer vision system. The proposed system consists of three modules. To explore the relationship between music and dance movements, we propose a cross-modal alignment module that focuses on dancing video clips, accompanied on pre-designed music, to learn a system that can judge the consistency between the visual features of pose sequences and the acoustic features of music. The learned model is then used in the imagination module to select a pose sequence for the given music. Such pose sequence selected from the music, however, is usually discontinuous. To solve this problem, in the spatial-temporal alignment module we develop a spatial alignment algorithm based on the tendency and periodicity of dance movements to predict dance movements between discontinuous fragments. In addition, the selected pose sequence is often misaligned with the music beat. To solve this problem, we further develop a temporal alignment algorithm to align the rhythm of music and dance. Finally, the processed pose sequence is used to synthesize realistic dancing videos in the imagination module. The generated dancing videos match the content and rhythm of the music. Experimental results and subjective evaluations show that the proposed approach can perform the function of generating promising dancing videos by inputting music.

updated: Sat Aug 07 2021 09:14:40 GMT+0000 (UTC)

published: Thu Sep 17 2020 02:29:13 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト