Accelerated Video Annotation driven by Deep Detector and Tracker

Eric Price; Aamir Ahmad

Deep Detector と Tracker による高速化されたビデオアノテーション

ビデオ内のオブジェクトのグラウンドトゥルースに注釈を付けることは、オブジェクトトラッカーのパフォーマンスの評価や画像ベースのオブジェクト検出器のトレーニングなど、ロボットの認識や機械学習におけるいくつかのダウンストリームタスクにとって不可欠です。ビデオ内のすべての画像フレームで、移動オブジェクトの注釈付きインスタンスの精度は非常に重要です。手作業で注釈を付けることは、非常に時間と労力がかかるだけでなく、エラー率が高くなる傾向があります。最先端の注釈方法は、最初のフレームでのみオブジェクトの境界ボックスを手動で初期化することに依存しており、その後、これらの境界ボックスを追跡するために、adaboost やカーネル化された相関フィルターなどの従来の追跡方法を使用します。これらはすぐにドリフトする可能性があるため、面倒な手動による監視が必要になります。この論文では、学習ベースの検出器 (SSD) と学習ベースのトラッカー (RE^3) の組み合わせを活用する新しい注釈方法を提案します。これにより、注釈のドリフトが大幅に減少し、その結果、必要な手動の監視が減少します。提案された注釈方法と一連のドローンビデオフレームに対する既存のベースラインを使用した注釈実験を通じて、アプローチを検証します。アノテーションプログラムの実行方法に関するソースコードと詳細情報は、https://github.com/robot-perception-group/smarter-labelme にあります。

Annotating object ground truth in videos is vital for several downstream tasks in robot perception and machine learning, such as for evaluating the performance of an object tracker or training an image-based object detector. The accuracy of the annotated instances of the moving objects on every image frame in a video is crucially important. Achieving that through manual annotations is not only very time consuming and labor intensive, but is also prone to high error rate. State-of-the-art annotation methods depend on manually initializing the object bounding boxes only in the first frame and then use classical tracking methods, e.g., adaboost, or kernelized correlation filters, to keep track of those bounding boxes. These can quickly drift, thereby requiring tedious manual supervision. In this paper, we propose a new annotation method which leverages a combination of a learning-based detector (SSD) and a learning-based tracker (RE^3). Through this, we significantly reduce annotation drifts, and, consequently, the required manual supervision. We validate our approach through annotation experiments using our proposed annotation method and existing baselines on a set of drone video frames. Source code and detailed information on how to run the annotation program can be found at https://github.com/robot-perception-group/smarter-labelme

updated: Sun Feb 19 2023 15:16:05 GMT+0000 (UTC)

published: Sun Feb 19 2023 15:16:05 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト