Temporal Contrastive Graph Learning for Video Action Recognition and Retrieval

Yang Liu; Keze Wang; Haoyuan Lan; Liang Lin

ビデオアクションの認識と検索のための時間的対照グラフ学習

自己教師ありビデオ表現学習の時間的多様性と時系列特性を完全に発見することを試み、この作業はビデオ内の時間的依存性を利用し、さらに時間的対照グラフ学習（TCGL）と呼ばれる新しい自己教師あり方法を提案します。複雑な時間依存性のモデリングを無視する既存の方法とは対照的に、TCGLは、スニペット間およびスニペット内の時間依存性を時間表現学習の自己監視信号と共同で見なすハイブリッドグラフ対照学習戦略に根ざしています。マルチスケールの時間依存性をモデル化するために、TCGLは、フレームとスニペットの順序に関する事前知識をグラフ構造、つまりスニペット内/スニペット間時間対照グラフに統合します。スニペット内グラフまたはスニペット間グラフのエッジをランダムに削除し、ノードをマスキングすることで、TCGLはさまざまな相関グラフビューを生成できます。次に、特定の対照学習モジュールが、異なるビューのノード間の一致を最大化するように設計されています。グローバルコンテキスト表現を適応的に学習し、チャネルごとの機能を再調整するために、適応型ビデオスニペット順序予測モジュールを導入します。これは、ビデオスニペット間の関係知識を活用して実際のスニペット順序を予測します。実験結果は、大規模な行動認識とビデオ検索のベンチマークにおいて、最先端の方法に対するTCGLの優位性を示しています。

Attempt to fully discover the temporal diversity and chronological characteristics for self-supervised video representation learning, this work takes advantage of the temporal dependencies within videos and further proposes a novel self-supervised method named Temporal Contrastive Graph Learning (TCGL). In contrast to the existing methods that ignore modeling elaborate temporal dependencies, our TCGL roots in a hybrid graph contrastive learning strategy to jointly regard the inter-snippet and intra-snippet temporal dependencies as self-supervision signals for temporal representation learning. To model multi-scale temporal dependencies, our TCGL integrates the prior knowledge about the frame and snippet orders into graph structures, i.e., the intra-/inter- snippet temporal contrastive graphs. By randomly removing edges and masking nodes of the intra-snippet graphs or inter-snippet graphs, our TCGL can generate different correlated graph views. Then, specific contrastive learning modules are designed to maximize the agreement between nodes in different views. To adaptively learn the global context representation and recalibrate the channel-wise features, we introduce an adaptive video snippet order prediction module, which leverages the relational knowledge among video snippets to predict the actual snippet orders. Experimental results demonstrate the superiority of our TCGL over the state-of-the-art methods on large-scale action recognition and video retrieval benchmarks.

updated: Wed Mar 17 2021 03:32:52 GMT+0000 (UTC)

published: Mon Jan 04 2021 08:11:39 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト