Egocentric Video-Language Pretraining

Kevin Qinghong Lin; Alex Jinpeng Wang; Mattia Soldan; Michael Wray; Rui Yan; Eric Zhongcong Xu; Difei Gao; Rongcheng Tu; Wenzhe Zhao; Weijie Kong; Chengfei Cai; Hongfa Wang; Dima Damen; Bernard Ghanem; Wei Liu; Mike Zheng Shou

自己中心的なビデオ言語の事前トレーニング

Video-Language Pretraining (VLP) は、幅広いビデオテキストダウンストリームタスクを進めるために転送可能な表現を学習することを目的としており、最近注目を集めています。最高のパフォーマンスを発揮する作品は、HowTo100M などの大規模な 3 人称のビデオテキストデータセットに依存しています。この作業では、最近リリースされた Ego4D データセットを活用して、3 つの方向に沿って Egocentric VLP を開拓します。 (i) EgoClip は、Ego4D から適切に選択された 380 万のクリップテキストペアで構成される一人称ビデオテキスト事前トレーニングデータセットであり、多種多様な人間の日常活動をカバーしています。 (ii) EgoNCE と呼ばれる新しい事前トレーニングの目的を提案します。これは、自己中心的な肯定的および否定的なサンプルをマイニングすることにより、ビデオテキストの対照的な学習を自己中心的なドメインに適応させます。 (iii) EgoClip に近い開発ベンチマークである EgoMCQ を導入したため、EgoClip と EgoNCE での設計決定の効果的な検証と迅速な調査をサポートできます。さらに、3 つのデータセットにわたる 5 つの自己中心的なダウンストリームタスクで強力なパフォーマンスを発揮します。シャレードエゴの行動認識; Ego4D チャレンジベンチマークでの自然言語クエリ、モーメントクエリ、およびオブジェクトの状態変化の分類。データセットとコードは https://github.com/showlab/EgoVLP で入手できます。

Video-Language Pretraining (VLP), which aims to learn transferable representation to advance a wide range of video-text downstream tasks, has recently received increasing attention. Best performing works rely on large-scale, 3rd-person video-text datasets, such as HowTo100M. In this work, we exploit the recently released Ego4D dataset to pioneer Egocentric VLP along three directions. (i) We create EgoClip, a 1st-person video-text pretraining dataset comprising 3.8M clip-text pairs well-chosen from Ego4D, covering a large variety of human daily activities. (ii) We propose a novel pretraining objective, dubbed EgoNCE, which adapts video-text contrastive learning to the egocentric domain by mining egocentric-aware positive and negative samples. (iii) We introduce EgoMCQ, a development benchmark that is close to EgoClip and hence can support effective validation and fast exploration of our design decisions in EgoClip and EgoNCE. Furthermore, we demonstrate strong performance on five egocentric downstream tasks across three datasets: video-text retrieval on EPIC-KITCHENS-100; action recognition on Charades-Ego; natural language query, moment query, and object state change classification on Ego4D challenge benchmarks. The dataset and code are available at https://github.com/showlab/EgoVLP.

updated: Thu Oct 13 2022 03:31:05 GMT+0000 (UTC)

published: Fri Jun 03 2022 16:28:58 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト