β-Multivariational Autoencoder for Entangled Representation Learning in Video Frames

Fatemeh Nouri; Robert Bergevin

ビデオフレームにおけるエンタングルド表現学習のための β-多変量オートエンコーダ

状態と以前の報酬を考慮して一連のアクションが期待される一連の意思決定プロセスを学習しながら、適切な分布からアクションを選択することが重要です。しかし、2 つ以上の潜在変数があり、2 つの変数すべてに共分散値がある場合、データから既知の事前確率を学習することは困難になります。データが大きくて多様な場合、多くの事後推定方法で事後崩壊が発生するためです。この論文では、意思決定プロセスの形で単一のオブジェクト追跡の一部として使用するために、ビデオフレームから多変量ガウス事前分布を学習する β-Multivariational Autoencoder (βMVAE) を提案します。単一のオブジェクト追跡タスクに対処するために、一連の依存パラメーターを使用して、ビデオ内のオブジェクトの動きの新しい定式化を提示します。モーションパラメータの真の値は、トレーニングセットのデータ分析を通じて取得されます。パラメータ母集団は、多変量ガウス分布を持つと仮定されます。 βMVAE は、出力がフレームパッチのオブジェクトマスクであるフレームパッチから、このエンタングルされた事前確率 p = N(μ, Σ) を直接学習するために開発されました。事後パラメータ、すなわち μ'、Σ' を推定するためのボトルネックを考案します。新しい再パラメータ化のトリックにより、入力のオブジェクトマスクとして尤度 p(x|z) を学習します。さらに、βMVAE のニューラルネットワークを U-Net アーキテクチャに変更し、新しいネットワークを βMultivariational U-Net (βMVUnet) と名付けました。私たちのネットワークは、24 (βMVUnet) および 78 (βMVAE) の 100 万ステップの 85,000 ビデオフレームを介してゼロからトレーニングされます。 βMVUnet が、テストセット全体で事後推定とセグメンテーション機能の両方を強化することを示します。私たちのコードと訓練されたネットワークは公開されています。

It is crucial to choose actions from an appropriate distribution while learning a sequential decision-making process in which a set of actions is expected given the states and previous reward. Yet, if there are more than two latent variables and every two variables have a covariance value, learning a known prior from data becomes challenging. Because when the data are big and diverse, many posterior estimate methods experience posterior collapse. In this paper, we propose the β-Multivariational Autoencoder (βMVAE) to learn a Multivariate Gaussian prior from video frames for use as part of a single object-tracking in form of a decision-making process. We present a novel formulation for object motion in videos with a set of dependent parameters to address a single object-tracking task. The true values of the motion parameters are obtained through data analysis on the training set. The parameters population is then assumed to have a Multivariate Gaussian distribution. The βMVAE is developed to learn this entangled prior p = N(μ, Σ) directly from frame patches where the output is the object masks of the frame patches. We devise a bottleneck to estimate the posterior's parameters, i.e. μ', Σ'. Via a new reparameterization trick, we learn the likelihood p(x|z) as the object mask of the input. Furthermore, we alter the neural network of βMVAE with the U-Net architecture and name the new network βMultivariational U-Net (βMVUnet). Our networks are trained from scratch via over 85k video frames for 24 (βMVUnet) and 78 (βMVAE) million steps. We show that βMVUnet enhances both posterior estimation and segmentation functioning over the test set. Our code and the trained networks are publicly released.

updated: Tue Nov 22 2022 23:25:17 GMT+0000 (UTC)

published: Tue Nov 22 2022 23:25:17 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト