Scalable Temporal Localization of Sensitive Activities in Movies and TV Episodes

Xiang Hao; Jingxiang Chen; Shixing Chen; Ahmed Saad; Raffay Hamid

映画やテレビのエピソードにおける敏感な活動のスケーラブルな時間的ローカリゼーション

顧客がより多くの情報に基づいて視聴を選択できるように、ビデオストリーミングサービスはコンテンツをモデレートし、映画やテレビエピソードのどの部分に年齢に適した素材（ヌード、性別、暴力、薬物使用など）が含まれているかをより明確に把握できるようにします。）。これらの機密性の高いアクティビティをローカライズするための教師ありモデルは、取得が困難な大量のクリップレベルのラベル付きデータを必要としますが、この目的のための弱教師ありモデルは通常、競争力のある精度を提供しません。この課題に対処するために、年齢に適した活動のまばらなクリップレベルのラベルと組み合わせて、簡単に入手できるビデオレベルの弱いラベルを利用するように設計された新しいCoarse2Fineネットワークを提案します。私たちのモデルは、フレームレベルの予測を集約してビデオレベルの分類を行うため、ビデオレベルのラベルとともにスパースクリップレベルのラベルを活用できます。さらに、フレームレベルの予測を階層的に実行することにより、私たちのアプローチは、年齢に適したコンテンツのまれな発生の性質によって引き起こされるラベルの不均衡の問題を克服することができます。 521のサブジャンルと250か国からの41,234本の映画とTVエピソード（約3年間のビデオコンテンツ）を使用したアプローチの比較結果を提示します。これまでに公開されたビデオ。私たちのアプローチは、既存の最先端の活動ローカリゼーションアプローチに比べて107.2％の相対的なmAPの改善（5.5％から11.4％）を提供します。

To help customers make better-informed viewing choices, video-streaming services try to moderate their content and provide more visibility into which portions of their movies and TV episodes contain age-appropriate material (e.g., nudity, sex, violence, or drug-use). Supervised models to localize these sensitive activities require large amounts of clip-level labeled data which is hard to obtain, while weakly-supervised models to this end usually do not offer competitive accuracy. To address this challenge, we propose a novel Coarse2Fine network designed to make use of readily obtainable video-level weak labels in conjunction with sparse clip-level labels of age-appropriate activities. Our model aggregates frame-level predictions to make video-level classifications and is therefore able to leverage sparse clip-level labels along with video-level labels. Furthermore, by performing frame-level predictions in a hierarchical manner, our approach is able to overcome the label-imbalance problem caused due to the rare-occurrence nature of age-appropriate content. We present comparative results of our approach using 41,234 movies and TV episodes (~3 years of video-content) from 521 sub-genres and 250 countries making it by far the largest-scale empirical analysis of age-appropriate activity localization in long-form videos ever published. Our approach offers 107.2% relative mAP improvement (from 5.5% to 11.4%) over existing state-of-the-art activity-localization approaches.

updated: Thu Jun 16 2022 20:16:28 GMT+0000 (UTC)

published: Thu Jun 16 2022 20:16:28 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト