CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory

Nur Muhammad Mahi Shafiullah; Chris Paxton; Lerrel Pinto; Soumith Chintala; Arthur Szlam

CLIP フィールド: ロボットメモリの弱く監視されたセマンティックフィールド

私たちは、セグメンテーション、インスタンスの識別、空間上のセマンティック検索、ビューの位置特定などのさまざまなタスクに使用できる暗黙的なシーンモデルである CLIP-Fields を提案します。 CLIP-Fields は、空間的な位置から意味論的な埋め込みベクトルへのマッピングを学習します。重要なのは、このマッピングは、CLIP、Detic、Sentence-BERT などの Web 画像および Web テキストでトレーニングされたモデルのみからの監視によってトレーニングできることを示しています。したがって、人間による直接の監視は使用されません。 Mask-RCNN のようなベースラインと比較した場合、私たちの方法は、ほんの一部の例を使用した HM3D データセット上の少数ショットのインスタンス識別やセマンティックセグメンテーションで優れたパフォーマンスを発揮します。最後に、CLIP-Fields をシーンメモリとして使用することで、ロボットが現実世界の環境でセマンティックナビゲーションを実行できることを示します。私たちのコードとデモビデオはここから入手できます: https://mahis.life/clip-fields

We propose CLIP-Fields, an implicit scene model that can be used for a variety of tasks, such as segmentation, instance identification, semantic search over space, and view localization. CLIP-Fields learns a mapping from spatial locations to semantic embedding vectors. Importantly, we show that this mapping can be trained with supervision coming only from web-image and web-text trained models such as CLIP, Detic, and Sentence-BERT; and thus uses no direct human supervision. When compared to baselines like Mask-RCNN, our method outperforms on few-shot instance identification or semantic segmentation on the HM3D dataset with only a fraction of the examples. Finally, we show that using CLIP-Fields as a scene memory, robots can perform semantic navigation in real-world environments. Our code and demonstration videos are available here: https://mahis.life/clip-fields

updated: Mon May 22 2023 22:34:29 GMT+0000 (UTC)

published: Tue Oct 11 2022 17:57:10 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト