Unsupervised motion segmentation in one go: Smooth long-term model over a video

Etienne Meunier; Patrick Bouthemy

Human beings have the ability to continuously analyze a video and immediately extract the main motion components. Motion segmentation methods based on deep learning often proceed frame by frame. We want to go beyond this paradigm, and perform the motion segmentation in series of flow fields of any length, up to the complete video sequence. It will be a prominent added value for downstream computer vision tasks, and could provide a pretext criterion for unsupervised video representation learning. In this perspective, we propose a novel long-term spatio-temporal model operating in a totally unsupervised way. It takes as input the volume of consecutive optical flow (OF) fields, and delivers a volume of segments of coherent motion over the video. More specifically, we have designed a transformer-based network, where we leverage a mathematically well-founded framework, the Evidence Lower Bound (ELBO), to infer the loss function. The loss function combines a flow reconstruction term involving spatio-temporal parametric motion models combining, in a novel way, polynomial (quadratic) motion models for the (x,y)-spatial dimensions and B-splines for the time dimension of the video sequence, and a regularization term enforcing temporal consistency on the masks. We report experiments on four VOS benchmarks with convincing quantitative results. We also highlight through visual results the key contributions on temporal consistency brought by our method.

updated: Sun Jan 28 2024 01:15:50 GMT+0000 (UTC)

published: Mon Oct 02 2023 09:33:54 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト