TCLR: Temporal Contrastive Learning for Video Representation

Ishan Dave; Rohit Gupta; Mamshad Nayeem Rizve; Mubarak Shah

TCLR：ビデオ表現のための時間的対照学習

対照学習は、画像表現の教師あり学習と自己教師あり学習の間のギャップをほぼ埋め、ビデオについても調査されています。ただし、ビデオデータの対照学習に関するこれまでの研究では、時間的次元全体で特徴を明確に区別するように明示的に奨励する効果については調査されていません。既存の対照的な自己教師ありビデオ表現学習方法を改善するために、2つの新しい損失からなる新しい時間的対照学習フレームワークを開発します。ローカルローカル時間コントラスト損失は、同じビデオからの重複しないクリップを区別するタスクを追加しますが、グローバルローカル時間コントラスト損失は、入力クリップの特徴マップのタイムステップを区別して、の時間的多様性を高めることを目的としています。学習した機能。私たちが提案する時間的対照学習フレームワークは、アクション認識、限定ラベルアクション分類、複数のビデオデータセットとバックボーンでの最近傍ビデオ検索など、さまざまなダウンストリームビデオ理解タスクで最先端の結果を大幅に改善します。また、視覚的に類似したクラスのきめ細かいアクション分類が大幅に改善されていることも示しています。一般的に使用されている3DResNet-18アーキテクチャを使用すると、UCF101で82.4％（+ 5.1％増加）、HMDB51アクション分類で52.9％（+ 5.4％増加）、56.2％（+ 11.7）のトップ1精度を達成します。％増加）トップ1UCF101最近傍ビデオ検索のリコール。

Contrastive learning has nearly closed the gap between supervised and self-supervised learning of image representations, and has also been explored for videos. However, prior work on contrastive learning for video data has not explored the effect of explicitly encouraging the features to be distinct across the temporal dimension. We develop a new temporal contrastive learning framework consisting of two novel losses to improve upon existing contrastive self-supervised video representation learning methods. The local-local temporal contrastive loss adds the task of discriminating between non-overlapping clips from the same video, whereas the global-local temporal contrastive aims to discriminate between timesteps of the feature map of an input clip in order to increase the temporal diversity of the learned features. Our proposed temporal contrastive learning framework achieves significant improvement over the state-of-the-art results in various downstream video understanding tasks such as action recognition, limited-label action classification, and nearest-neighbor video retrieval on multiple video datasets and backbones. We also demonstrate significant improvement in fine-grained action classification for visually similar classes. With the commonly used 3D ResNet-18 architecture, we achieve 82.4% (+5.1% increase over the previous best) top-1 accuracy on UCF101 and 52.9% (+5.4% increase) on HMDB51 action classification, and 56.2% (+11.7% increase) Top-1 Recall on UCF101 nearest neighbor video retrieval.

updated: Thu Apr 08 2021 15:39:49 GMT+0000 (UTC)

published: Wed Jan 20 2021 05:38:16 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト