AirObject: A Temporally Evolving Graph Embedding for Object Identification

Nikhil Varma Keetha; Chen Wang; Yuheng Qiu; Kuan Xu; Sebastian Scherer

AirObject：オブジェクト識別のための時間的に進化するグラフ埋め込み

オブジェクトのエンコードと識別は、自律的な探索、セマンティックシーンの理解、再ローカリゼーションなどのロボットタスクに不可欠です。以前のアプローチでは、オブジェクトを追跡するか、オブジェクトを識別するための記述子を生成しようとしました。ただし、このようなシステムは、単一の観点からの「固定」部分オブジェクト表現に限定されます。ロボット探索のセットアップでは、ロボットが複数の視点からオブジェクトを観察するときに構築される、時間的に「進化する」グローバルオブジェクト表現が必要です。さらに、現実の世界には未知の新規オブジェクトが広範に分布していることを考えると、オブジェクトの識別プロセスはクラスにとらわれないものでなければなりません。このコンテキストでは、オブジェクトのグローバルキーポイントグラフベースの埋め込みを取得するために、AirObjectと呼ばれる新しい時間3Dオブジェクトエンコーディングアプローチを提案します。具体的には、グローバル3Dオブジェクトの埋め込みは、グラフの注意ベースのエンコード方法から取得された複数のフレームの構造情報にわたる時間畳み込みネットワークを使用して生成されます。 AirObjectは、ビデオオブジェクト識別の最先端のパフォーマンスを実現し、深刻なオクルージョン、知覚エイリアシング、視点シフト、変形、およびスケール変換に対して堅牢であり、最先端のシングルフレームを上回り、シーケンシャル記述子。私たちの知る限り、AirObjectは最初の時間オブジェクトエンコーディングメソッドの1つです。

Object encoding and identification are vital for robotic tasks such as autonomous exploration, semantic scene understanding, and re-localization. Previous approaches have attempted to either track objects or generate descriptors for object identification. However, such systems are limited to a "fixed" partial object representation from a single viewpoint. In a robot exploration setup, there is a requirement for a temporally "evolving" global object representation built as the robot observes the object from multiple viewpoints. Furthermore, given the vast distribution of unknown novel objects in the real world, the object identification process must be class-agnostic. In this context, we propose a novel temporal 3D object encoding approach, dubbed AirObject, to obtain global keypoint graph-based embeddings of objects. Specifically, the global 3D object embeddings are generated using a temporal convolutional network across structural information of multiple frames obtained from a graph attention-based encoding method. We demonstrate that AirObject achieves the state-of-the-art performance for video object identification and is robust to severe occlusion, perceptual aliasing, viewpoint shift, deformation, and scale transform, outperforming the state-of-the-art single-frame and sequential descriptors. To the best of our knowledge, AirObject is one of the first temporal object encoding methods.

updated: Tue Nov 30 2021 06:17:03 GMT+0000 (UTC)

published: Tue Nov 30 2021 06:17:03 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト