Novel View Video Prediction Using a Dual Representation

Sarah Shiraz; Krishna Regmi; Shruti Vyas; Yogesh S. Rawat; Mubarak Shah

デュアル表現を使用した新規ビュービデオ予測

新規視聴ビデオ予測の問題に取り組みます。単一または複数のビューからの一連の入力ビデオクリップを指定すると、ネットワークは新しいビューからビデオを予測できます。提案されたアプローチは事前確率を必要とせず、視点の小さな変化を予測する最近の研究と比較して、最大 45 度のより広い角度距離からビデオを予測できます。さらに、私たちの方法は、RGB フレームのみに依存して、新しい視点からビデオを生成するために使用されるデュアル表現を学習します。デュアル表現には、ビューに依存する表現と、新しいビューのビデオ予測を可能にする補完的な詳細を組み込んだグローバルな表現が含まれます。 NTU-RGB+D と CMU Panoptic という 2 つの現実世界のデータセットでのフレームワークの有効性を示します。最先端の新規ビュービデオ予測方法との比較では、ターゲットビューからの明示的な事前予測を使用せずに、SSIM で 26.1%、PSNR で 13.6%、FVD スコアで 60% の改善が示されています。

We address the problem of novel view video prediction; given a set of input video clips from a single/multiple views, our network is able to predict the video from a novel view. The proposed approach does not require any priors and is able to predict the video from wider angular distances, upto 45 degree, as compared to the recent studies predicting small variations in viewpoint. Moreover, our method relies only onRGB frames to learn a dual representation which is used to generate the video from a novel viewpoint. The dual representation encompasses a view-dependent and a global representation which incorporates complementary details to enable novel view video prediction. We demonstrate the effectiveness of our framework on two real world datasets: NTU-RGB+D and CMU Panoptic. A comparison with the State-of-the-art novel view video prediction methods shows an improvement of 26.1% in SSIM, 13.6% in PSNR, and 60% inFVD scores without using explicit priors from target views.

updated: Mon Jun 07 2021 20:41:33 GMT+0000 (UTC)

published: Mon Jun 07 2021 20:41:33 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト