Exploiting Prompt Caption for Video Grounding

Hongxiang Li; Meng Cao; Xuxin Cheng; Zhihong Zhu; Yaowei Li; Yuexian Zou

ビデオグラウンディングのためのプロンプトキャプションの悪用

ビデオグラウンディングは、トリミングされていないビデオから、特定のクエリ文に一致する興味のある瞬間を見つけることを目的としています。以前の作品は、データセット内の潜在的なイベントとクエリ文の間のコンテキスト情報を提供できない、ビデオアノテーションの希薄なジレンマを無視しています。この論文では、一般的なアクションを説明する簡単に利用できるキャプション、つまり私たちの論文で定義されたプロンプトキャプション (PC) を利用すると、パフォーマンスが大幅に向上すると主張します。この目的のために、ビデオグラウンディング用のプロンプトキャプションネットワーク (PCNet) を提案します。具体的には、最初に高密度ビデオキャプションを導入して高密度キャプションを生成し、次に非プロンプトキャプション抑制 (NPCS) によってプロンプトキャプションを取得します。プロンプトキャプションの潜在的な情報をキャプチャするために、キャプションガイド付きアテンション (CGA) がプロンプトキャプションとクエリ文の間の意味的関係を時間空間に投影し、それらを視覚的表現に融合することを提案します。プロンプトキャプションとグラウンドトゥルースの間のギャップを考慮して、クロスモーダル相互情報を最大化するために、より多くのネガティブペアを構築するための非対称クロスモーダル対照学習 (ACCL) を提案します。付加機能がなければ、3 つの公開データセット (つまり、ActivityNet Captions、TACoS、および ActivityNet-CG) での広範な実験により、私たちの方法が最先端の方法よりも大幅に優れていることが実証されました。

Video grounding aims to locate a moment of interest matching the given query sentence from an untrimmed video. Previous works ignore the sparsity dilemma in video annotations, which fails to provide the context information between potential events and query sentences in the dataset. In this paper, we contend that exploiting easily available captions which describe general actions i.e. , prompt captions (PC) defined in our paper, will significantly boost the performance. To this end, we propose a Prompt Caption Network (PCNet) for video grounding. Specifically, we first introduce dense video captioning to generate dense captions and then obtain prompt captions by Non-Prompt Caption Suppression (NPCS). To capture the potential information in prompt captions, we propose Caption Guided Attention (CGA) project the semantic relations between prompt captions and query sentences into temporal space and fuse them into visual representations. Considering the gap between prompt captions and ground truth, we propose Asymmetric Cross-modal Contrastive Learning (ACCL) for constructing more negative pairs to maximize cross-modal mutual information. Without bells and whistles, extensive experiments on three public datasets (i.e. , ActivityNet Captions, TACoS and ActivityNet-CG) demonstrate that our method significantly outperforms state-of-the-art methods.

updated: Tue Mar 28 2023 10:49:59 GMT+0000 (UTC)

published: Sun Jan 15 2023 02:04:02 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト