SupeRGB-D: Zero-shot Instance Segmentation in Cluttered Indoor Environments

Evin Pınar Örnek; Aravindhan K Krishnan; Shreekant Gayaka; Cheng-Hao Kuo; Arnie Sen; Nassir Navab; Federico Tombari

SupeRGB-D: 雑然とした屋内環境におけるゼロショットインスタンスセグメンテーション

オブジェクトインスタンスのセグメンテーションは、多数の小さなオブジェクトが存在する乱雑な環境を移動する屋内ロボットにとって重要な課題です。 3D センシング機能には制限があるため、考えられるすべてのオブジェクトを検出することが困難になることがよくあります。この問題にはディープラーニングのアプローチが効果的である可能性がありますが、教師あり学習用に 3D データに手動でアノテーションを付けるのは時間がかかります。この研究では、意味論的なカテゴリーに依存しない方法で目に見えないオブジェクトを識別するために、RGB-D データからのゼロショットインスタンスセグメンテーション (ZSIS) を調査します。この研究を可能にするために、Tabletop Objects Dataset (TOD-Z) のゼロショット分割を導入し、アノテーション付きオブジェクトを使用してピクセルの「オブジェクト性」を学習し、雑然とした屋内環境で目に見えないオブジェクトカテゴリに一般化する方法を紹介します。私たちの手法である SupeRGB-D は、幾何学的な手がかりに基づいてピクセルを小さなパッチにグループ化し、深い凝集クラスタリング方式でパッチをマージする方法を学習します。 SupeRGB-D は、目に見えないオブジェクトでは既存のベースラインを上回り、目に見えるオブジェクトでは同様のパフォーマンスを実現します。さらに、実際のデータセット OCID での競合結果を示します。軽量設計 (0.4 MB メモリ要件) により、私たちの方法はモバイルおよびロボットのアプリケーションに非常に適しています。 DINO 機能を追加すると、メモリ要件が増加してもパフォーマンスが向上します。データセットの分割とコードは https://github.com/evinpinar/supergb-d で入手できます。

Object instance segmentation is a key challenge for indoor robots navigating cluttered environments with many small objects. Limitations in 3D sensing capabilities often make it difficult to detect every possible object. While deep learning approaches may be effective for this problem, manually annotating 3D data for supervised learning is time-consuming. In this work, we explore zero-shot instance segmentation (ZSIS) from RGB-D data to identify unseen objects in a semantic category-agnostic manner. We introduce a zero-shot split for Tabletop Objects Dataset (TOD-Z) to enable this study and present a method that uses annotated objects to learn the ``objectness'' of pixels and generalize to unseen object categories in cluttered indoor environments. Our method, SupeRGB-D, groups pixels into small patches based on geometric cues and learns to merge the patches in a deep agglomerative clustering fashion. SupeRGB-D outperforms existing baselines on unseen objects while achieving similar performance on seen objects. We further show competitive results on the real dataset OCID. With its lightweight design (0.4 MB memory requirement), our method is extremely suitable for mobile and robotic applications. Additional DINO features can increase performance with a higher memory requirement. The dataset split and code are available at https://github.com/evinpinar/supergb-d.

updated: Thu May 25 2023 12:25:24 GMT+0000 (UTC)

published: Thu Dec 22 2022 17:59:48 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト