TTVOS: Lightweight Video Object Segmentation with Adaptive Template Attention Module and Temporal Consistency Loss

Hyojin Park; Ganesh Venkatesh; Nojun Kwak

TTVOS：アダプティブテンプレートアテンションモジュールと時間的整合性損失を備えた軽量ビデオオブジェクトセグメンテーション

半教師ありビデオオブジェクトセグメンテーション（semi-VOS）は、多くのアプリケーションで広く使用されています。このタスクは、特定のターゲットマスクからクラスに依存しないオブジェクトを追跡します。これを行うために、オンライン学習、メモリネットワーク、およびオプティカルフローに基づいてさまざまなアプローチが開発されてきました。これらの方法は高い精度を示しますが、推論時間が遅く、非常に複雑であるため、実際のアプリケーションで利用するのは困難です。この問題を解決するために、テンプレートマッチングの方法は、処理速度を高速化するために考案されていますが、以前のモデルでは多くのパフォーマンスが犠牲になっています。テンプレートマッチング法と時間的整合性損失に基づく新しいセミVOSモデルを紹介し、推論時間を大幅に短縮しながら、重いモデルとのパフォーマンスギャップを減らします。私たちのテンプレートマッチング方法は、短期および長期のマッチングで構成されています。短期間のマッチングはターゲットオブジェクトのローカリゼーションを強化し、長期的なマッチングは細部を改善し、新しく提案された適応型テンプレートアテンションモジュールを通じてオブジェクトの形状変化を処理します。ただし、長期的なマッチングでは、テンプレートの更新時に過去の推定結果が流入するため、エラーが伝播します。この問題を軽減するために、遷移行列の概念を採用することにより、隣接するフレーム間の時間的コヒーレンスを向上させるための時間的一貫性の喪失も提案します。私たちのモデルは、DAVIS16ベンチマークで73.8 FPSの速度で79.5％のJ＆Fスコアを取得します。コードはhttps://github.com/HYOJINPARK/TTVOSで入手できます。

Semi-supervised video object segmentation (semi-VOS) is widely used in many applications. This task is tracking class-agnostic objects from a given target mask. For doing this, various approaches have been developed based on online-learning, memory networks, and optical flow. These methods show high accuracy but are hard to be utilized in real-world applications due to slow inference time and tremendous complexity. To resolve this problem, template matching methods are devised for fast processing speed but sacrificing lots of performance in previous models. We introduce a novel semi-VOS model based on a template matching method and a temporal consistency loss to reduce the performance gap from heavy models while expediting inference time a lot. Our template matching method consists of short-term and long-term matching. The short-term matching enhances target object localization, while long-term matching improves fine details and handles object shape-changing through the newly proposed adaptive template attention module. However, the long-term matching causes error-propagation due to the inflow of the past estimated results when updating the template. To mitigate this problem, we also propose a temporal consistency loss for better temporal coherence between neighboring frames by adopting the concept of a transition matrix. Our model obtains 79.5% J&F score at the speed of 73.8 FPS on the DAVIS16 benchmark. The code is available in https://github.com/HYOJINPARK/TTVOS.

updated: Sun Apr 04 2021 10:02:52 GMT+0000 (UTC)

published: Mon Nov 09 2020 14:09:54 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト