What You Say Is What You Show: Visual Narration Detection in Instructional Videos

Kumar Ashutosh; Rohit Girdhar; Lorenzo Torresani; Kristen Grauman

What You Say is What You Show: 教育ビデオにおける視覚的なナレーションの検出

ナレーション付きの「ハウツー」ビデオは、視覚表現の学習からロボットポリシーのトレーニングに至るまで、幅広い学習問題に対する有望なデータソースとして浮上しています。ただし、ナレーションがビデオで示されているアクションを必ずしも説明しているわけではないため、このデータには非常にノイズが含まれています。この問題に対処するために、視覚的なナレーション検出という新しいタスクを導入します。これには、ナレーションがビデオ内のアクションによって視覚的に表現されているかどうかを判断することが含まれます。私たちは、マルチモーダルキューと疑似ラベル付けを利用して、弱くラベル付けされたデータのみで視覚的なナレーションを検出する方法を学習する方法である What You Say is What You Show (WYS^2) を提案します。私たちのモデルは、強力なベースラインを上回る、実際のビデオ内の視覚的なナレーションの検出に成功し、教育ビデオの最先端の要約と時間的調整に対するその影響を実証しました。

Narrated ''how-to'' videos have emerged as a promising data source for a wide range of learning problems, from learning visual representations to training robot policies. However, this data is extremely noisy, as the narrations do not always describe the actions demonstrated in the video. To address this problem we introduce the novel task of visual narration detection, which entails determining whether a narration is visually depicted by the actions in the video. We propose What You Say is What You Show (WYS^2), a method that leverages multi-modal cues and pseudo-labeling to learn to detect visual narrations with only weakly labeled data. Our model successfully detects visual narrations in in-the-wild videos, outperforming strong baselines, and we demonstrate its impact for state-of-the-art summarization and temporal alignment of instructional videos.

updated: Tue Jul 18 2023 17:29:16 GMT+0000 (UTC)

published: Thu Jan 05 2023 21:43:19 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト