Auxiliary Learning for Self-Supervised Video Representation via Similarity-based Knowledge Distillation

Amirhossein Dadashzadeh; Alan Whone; Majid Mirmehdi

類似性に基づく知識蒸留による自己監視ビデオ表現のための補助学習

ビデオ表現学習のための自己監視型事前トレーニング方法の目覚ましい成功にもかかわらず、事前トレーニング用のラベルなしデータセットが小さい場合、またはソースタスク（事前トレーニング）のラベルなしデータとターゲットタスク（微調整）のラベル付きデータ間のドメインの違いが大きい場合、一般化は不十分です。。これらの問題を軽減するために、知識類似性蒸留（auxSKD）に基づいて、補助的な事前トレーニングフェーズを介して自己監視事前トレーニングを補完する新しいアプローチを提案します。 400。私たちの方法は、ラベルのないビデオデータのセグメント間の類似性情報をキャプチャすることにより、その知識を学生モデルに繰り返し抽出する教師ネットワークを展開します。次に、学生モデルは、この事前知識を活用して口実のタスクを解決します。また、新しい口実タスク、ビデオセグメントペース予測またはVSPPを紹介します。これは、モデルが入力ビデオのランダムに選択されたセグメントの再生速度を予測して、より信頼性の高い自己監視表現を提供することを要求します。私たちの実験結果は、K100で事前トレーニングした場合、UCF101データセットとHMDB51データセットの両方で最先端の結果よりも優れた結果を示しています。さらに、補助関連のauxSKDを、最新の自己監視方式（VideoPaceやRSPNetなど）に追加の事前トレーニングフェーズとして追加すると、UCF101およびHMDB51での結果が向上することを示します。私たちのコードはまもなくリリースされます。

Despite the outstanding success of self-supervised pretraining methods for video representation learning, they generalise poorly when the unlabeled dataset for pretraining is small or the domain difference between unlabelled data in source task (pretraining) and labeled data in target task (finetuning) is significant. To mitigate these issues, we propose a novel approach to complement self-supervised pretraining via an auxiliary pretraining phase, based on knowledge similarity distillation, auxSKD, for better generalisation with a significantly smaller amount of video data, e.g. Kinetics-100 rather than Kinetics-400. Our method deploys a teacher network that iteratively distils its knowledge to the student model by capturing the similarity information between segments of unlabelled video data. The student model then solves a pretext task by exploiting this prior knowledge. We also introduce a novel pretext task, Video Segment Pace Prediction or VSPP, which requires our model to predict the playback speed of a randomly selected segment of the input video to provide more reliable self-supervised representations. Our experimental results show superior results to the state of the art on both UCF101 and HMDB51 datasets when pretraining on K100. Additionally, we show that our auxiliary pertaining, auxSKD, when added as an extra pretraining phase to recent state of the art self-supervised methods (e.g. VideoPace and RSPNet), improves their results on UCF101 and HMDB51. Our code will be released soon.

updated: Tue Dec 07 2021 21:50:40 GMT+0000 (UTC)

published: Tue Dec 07 2021 21:50:40 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト