Generating Templated Caption for Video Grounding

Hongxiang Li; Meng Cao; Xuxin Cheng; Zhihong Zhu; Yaowei Li; Yuexian Zou

ビデオグラウンディング用のテンプレート化されたキャプションの生成

ビデオグラウンディングは、トリミングされていないビデオから、特定のクエリ文に一致する興味のある瞬間を見つけることを目的としています。以前の作品は、データセット内の潜在的なイベントとクエリ文の間のコンテキスト情報を提供できない、ビデオアノテーションの希薄なジレンマを無視しています。この論文では、一般的なアクションを説明する簡単に利用できるキャプション、つまりこの論文で定義されたテンプレート化されたキャプションを提供すると、パフォーマンスが大幅に向上すると主張します。この目的のために、ビデオグラウンディング用の Templated Caption Network (TCNet) を提案します。具体的には、最初に高密度のビデオキャプションを導入して高密度のキャプションを生成し、次に、Non-Templated Caption Suppression (NTCS) によってテンプレート化されたキャプションを取得します。テンプレート化されたキャプションをより適切に利用するために、テンプレート化されたキャプションとクエリ文の間の意味的関係を時間空間に投影し、それらを視覚的表現に融合する Caption Guided Attention (CGA) を提案します。テンプレート化されたキャプションとグラウンドトゥルースの間のギャップを考慮して、クロスモーダル相互情報を最大化するために、より多くのネガティブペアを構築するための非対称デュアルマッチング教師あり対照学習 (ADMSCL) を提案します。付加機能がなければ、3 つの公開データセット (つまり、ActivityNet Captions、TACoS、および ActivityNet-CG) での広範な実験により、私たちの方法が最先端の方法よりも大幅に優れていることが実証されました。

Video grounding aims to locate a moment of interest matching the given query sentence from an untrimmed video. Previous works ignore the sparsity dilemma in video annotations, which fails to provide the context information between potential events and query sentences in the dataset. In this paper, we contend that providing easily available captions which describe general actions i.e. , templated captions defined in our paper, will significantly boost the performance. To this end, we propose a Templated Caption Network (TCNet) for video grounding. Specifically, we first introduce dense video captioning to generate dense captions, and then obtain templated captions by Non-Templated Caption Suppression (NTCS). To utilize templated captions better, we propose Caption Guided Attention (CGA) project the semantic relations between templated captions and query sentences into temporal space and fuse them into visual representations. Considering the gap between templated captions and ground truth, we propose Asymmetric Dual Matching Supervised Contrastive Learning (ADMSCL) for constructing more negative pairs to maximize cross-modal mutual information. Without bells and whistles, extensive experiments on three public datasets (i.e. , ActivityNet Captions, TACoS and ActivityNet-CG) demonstrate that our method significantly outperforms state-of-the-art methods.

updated: Sun Jan 15 2023 02:04:02 GMT+0000 (UTC)

published: Sun Jan 15 2023 02:04:02 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト