Temporal Contrastive Graph for Self-supervised Video Representation Learning

Yang Liu; Keze Wang; Haoyuan Lan; Liang Lin

自己教師ありビデオ表現学習のための時間的対照グラフ

この作品は、自己教師ありビデオ表現学習のためのきめ細かい時間的構造とグローバルローカル時系列特性を完全に探求することを試み、ビデオの時間的構造を活用することを詳しく調べ、Temporal ContrastiveGraphという名前の新しい自己教師あり方法をさらに提案します（TCG）。ビデオ内のビデオフレームまたはビデオスニペットをランダムにシャッフルする既存の方法とは対照的に、提案されたTCGは、スニペット間およびスニペット内の時間的関係を時間的表現の自己監視信号と見なすハイブリッドグラフ対照学習戦略に根ざしています。学習。人間の視覚系が局所的および全体的な時間的変化の両方に敏感であるという神経科学研究に触発されて、提案されたTCGは、フレームとスニペットの順序に関する事前知識を時間的対照グラフ構造、つまりスニペット内/スニペット間時間的対照に統合しますグラフモジュール。ビデオフレームセットとスニペット間のローカルおよびグローバルの時間的関係を適切に保持します。エッジをランダムに削除し、スニペット内グラフまたはスニペット間グラフのノード機能をマスキングすることで、TCGはさまざまな相関グラフビューを生成できます。次に、特定の対照的な損失は、異なるビューでのノード埋め込み間の一致を最大化するように設計されています。グローバルコンテキスト表現を学習し、チャネルごとの機能を適応的に再調整するために、適応型ビデオスニペット順序予測モジュールを導入します。これは、ビデオスニペット間の関係知識を活用して、実際のスニペット順序を予測します。広範な実験結果は、大規模な行動認識とビデオ検索のベンチマークにおいて、最先端の方法に対するTCGの優位性を示しています。

Attempt to fully explore the fine-grained temporal structure and global-local chronological characteristics for self-supervised video representation learning, this work takes a closer look at exploiting the temporal structure of videos and further proposes a novel self-supervised method named Temporal Contrastive Graph (TCG). In contrast to the existing methods that randomly shuffle the video frames or video snippets within a video, our proposed TCG roots in a hybrid graph contrastive learning strategy to regard the inter-snippet and intra-snippet temporal relationships as self-supervision signals for temporal representation learning. Inspired by the neuroscience studies that the human visual system is sensitive to both local and global temporal changes, our proposed TCG integrates the prior knowledge about the frame and snippet orders into temporal contrastive graph structures, i.e., the intra-/inter- snippet temporal contrastive graph modules, to well preserve the local and global temporal relationships among video frame-sets and snippets. By randomly removing edges and masking node features of the intra-snippet graphs or inter-snippet graphs, our TCG can generate different correlated graph views. Then, specific contrastive losses are designed to maximize the agreement between node embeddings in different views. To learn the global context representation and recalibrate the channel-wise features adaptively, we introduce an adaptive video snippet order prediction module, which leverages the relational knowledge among video snippets to predict the actual snippet orders. Extensive experimental results demonstrate the superiority of our TCG over the state-of-the-art methods on large-scale action recognition and video retrieval benchmarks.

updated: Tue Jan 26 2021 08:51:36 GMT+0000 (UTC)

published: Mon Jan 04 2021 08:11:39 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト