Bootstrapped Self-Supervised Training with Monocular Video for Semantic Segmentation and Depth Estimation

Yihao Zhang; John J. Leonard

セマンティックセグメンテーションと深度推定のための単眼ビデオによるブートストラップ自己教師ありトレーニング

世界に配備されたロボットにとって、初期の事前設定された知識を改善するための自律学習の能力を有することが望ましい。これを、ラベル付きデータセットの教師ありトレーニングでシステムが最初にブートストラップされるブートストラップ自己教師あり学習問題として形式化し、その後、ラベルなしデータのみを使用して教師ありトレーニングベースラインよりもシステムを改善できる自己教師ありトレーニング方法を探します。この作業では、単眼ビデオのフレーム間の時間的一貫性を活用して、このブートストラップされた自己教師ありトレーニングを実行します。よく訓練された最先端のセマンティックセグメンテーションネットワークが、私たちの方法によってさらに改善できることを示します。さらに、ブートストラップされた自己教師ありトレーニングフレームワークは、ネットワークが純粋な教師ありトレーニングや自己教師ありトレーニングよりも深度推定を学習するのに役立つことを示しています。

For a robot deployed in the world, it is desirable to have the ability of autonomous learning to improve its initial pre-set knowledge. We formalize this as a bootstrapped self-supervised learning problem where a system is initially bootstrapped with supervised training on a labeled dataset and we look for a self-supervised training method that can subsequently improve the system over the supervised training baseline using only unlabeled data. In this work, we leverage temporal consistency between frames in monocular video to perform this bootstrapped self-supervised training. We show that a well-trained state-of-the-art semantic segmentation network can be further improved through our method. In addition, we show that the bootstrapped self-supervised training framework can help a network learn depth estimation better than pure supervised training or self-supervised training.

updated: Sun Aug 01 2021 03:50:19 GMT+0000 (UTC)

published: Fri Mar 19 2021 21:28:58 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト