It's Time for Artistic Correspondence in Music and Video

Didac Suris; Carl Vondrick; Bryan Russell; Justin Salamon

音楽とビデオの芸術的対応の時が来ました

時間的配置と芸術的レベルでの対応の両方に基づいて、特定のビデオの音楽トラックを推奨するためのアプローチを提示します。その逆も同様です。人間の注釈を必要とせずに、データから直接この対応を学習する自己監視アプローチを提案します。タスクを解決するために必要な高レベルの概念をキャプチャするために、各モダリティのTransformerネットワークを使用して、ビデオ信号と音楽信号の両方の長期的な時間的コンテキストをモデル化することを提案します。実験によると、このアプローチは、時間的コンテキストを利用しない代替案よりも大幅に優れています。私たちの貢献の組み合わせは、従来の最先端技術の最大10倍の検索精度を向上させます。この強力な改善により、幅広い分析とアプリケーションを導入することができます。たとえば、視覚的に定義された属性に基づいて音楽検索を調整できます。

We present an approach for recommending a music track for a given video, and vice versa, based on both their temporal alignment and their correspondence at an artistic level. We propose a self-supervised approach that learns this correspondence directly from data, without any need of human annotations. In order to capture the high-level concepts that are required to solve the task, we propose modeling the long-term temporal context of both the video and the music signals, using Transformer networks for each modality. Experiments show that this approach strongly outperforms alternatives that do not exploit the temporal context. The combination of our contributions improve retrieval accuracy up to 10x over prior state of the art. This strong improvement allows us to introduce a wide range of analyses and applications. For instance, we can condition music retrieval based on visually defined attributes.

updated: Tue Jun 14 2022 20:21:04 GMT+0000 (UTC)

published: Tue Jun 14 2022 20:21:04 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト