Unsupervised Video Summarization with a Convolutional Attentive Adversarial Network

Guoqiang Liang; Yanbing Lv; Shucheng Li; Shizhou Zhang; Yanning Zhang

畳み込み注意深い敵対的ネットワークによる教師なしビデオ要約

ビデオデータの爆発的な増加に伴い、メインストーリーを伝えながらフレームの最小サブセットを探そうとするビデオ要約が最もホットなトピックの1つになっています。今日、特に深層学習の出現後、教師あり学習技術によって実質的な成果が得られています。ただし、大規模なビデオデータセットの人間による注釈を収集することは非常に費用がかかり、困難です。この問題に対処するために、畳み込みの注意深い敵対的ネットワーク（CAAN）を提案します。その重要なアイデアは、教師なしの方法で詳細なサマライザーを構築することです。生成的敵対的ネットワークでは、全体的なフレームワークはジェネレーターとディスクリミネーターで構成されます。前者はビデオのすべてのフレームの重要度スコアを予測し、後者はスコア加重フレームの特徴を元のフレームの特徴と区別しようとします。具体的には、ジェネレータは、完全畳み込みシーケンスネットワークを使用してビデオのグローバル表現を抽出し、注意ベースのネットワークを使用して正規化された重要度スコアを出力します。パラメータを学習するために、目的関数は3つの損失関数で構成されており、フレームレベルの重要度スコアの予測を共同でガイドできます。この提案された方法を検証するために、2つの公開ベンチマークSumMeとTVSumで広範な実験を実施しました。結果は、他の最先端の教師なしアプローチに対する提案された方法の優位性を示しています。私たちの方法は、いくつかの公開された教師ありアプローチよりも優れています。

With the explosive growth of video data, video summarization, which attempts to seek the minimum subset of frames while still conveying the main story, has become one of the hottest topics. Nowadays, substantial achievements have been made by supervised learning techniques, especially after the emergence of deep learning. However, it is extremely expensive and difficult to collect human annotation for large-scale video datasets. To address this problem, we propose a convolutional attentive adversarial network (CAAN), whose key idea is to build a deep summarizer in an unsupervised way. Upon the generative adversarial network, our overall framework consists of a generator and a discriminator. The former predicts importance scores for all frames of a video while the latter tries to distinguish the score-weighted frame features from original frame features. Specifically, the generator employs a fully convolutional sequence network to extract global representation of a video, and an attention-based network to output normalized importance scores. To learn the parameters, our objective function is composed of three loss functions, which can guide the frame-level importance score prediction collaboratively. To validate this proposed method, we have conducted extensive experiments on two public benchmarks SumMe and TVSum. The results show the superiority of our proposed method against other state-of-the-art unsupervised approaches. Our method even outperforms some published supervised approaches.

updated: Mon May 24 2021 07:24:39 GMT+0000 (UTC)

published: Mon May 24 2021 07:24:39 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト