Deep Learning for Embodied Vision Navigation: A Survey

Fengda Zhu; Yi Zhu; Vincent CS Lee; Xiaodan Liang; Xiaojun Chang

身体化された視覚ナビゲーションのための深層学習：調査

「具体化された視覚的ナビゲーション」の問題は、エージェントが3D環境でナビゲートすることを主に一人称の観察に依存することを要求します。この問題は、自動運転、掃除機、レスキューロボットへの幅広い応用により、近年注目を集めています。ナビゲーションエージェントは、視覚的知覚、マッピング、計画、探索、推論など、さまざまなインテリジェントスキルを備えている必要があります。観察、思考、行動するエージェントを構築することは、実際の知能の鍵です。ディープラーニング手法の驚くべき学習能力により、エージェントは具体化された視覚的ナビゲーションタスクを実行できるようになりました。それにもかかわらず、部分的に観察された視覚入力の認識、見えない領域の探索、見られたシナリオの記憶とモデリング、クロスモーダル命令の理解、新しい環境への適応など、多くの高度なスキルが必要なため、具体化された視覚ナビゲーションはまだ初期段階です。最近、具体化された視覚ナビゲーションがコミュニティの注目を集めており、これらのスキルを学ぶために多くの作品が提案されています。この論文は、包括的な文献調査を提供することにより、具体化された視覚ナビゲーションの分野における現在の研究の概要を確立することを試みる。ベンチマークとメトリックを要約し、さまざまな方法を確認し、課題を分析し、最先端の方法を強調します。最後に、具体化された視覚ナビゲーションの分野における未解決の課題について議論し、将来の研究を追求する上で有望な方向性を示します。

"Embodied visual navigation" problem requires an agent to navigate in a 3D environment mainly rely on its first-person observation. This problem has attracted rising attention in recent years due to its wide application in autonomous driving, vacuum cleaner, and rescue robot. A navigation agent is supposed to have various intelligent skills, such as visual perceiving, mapping, planning, exploring and reasoning, etc. Building such an agent that observes, thinks, and acts is a key to real intelligence. The remarkable learning ability of deep learning methods empowered the agents to accomplish embodied visual navigation tasks. Despite this, embodied visual navigation is still in its infancy since a lot of advanced skills are required, including perceiving partially observed visual input, exploring unseen areas, memorizing and modeling seen scenarios, understanding cross-modal instructions, and adapting to a new environment, etc. Recently, embodied visual navigation has attracted rising attention of the community, and numerous works has been proposed to learn these skills. This paper attempts to establish an outline of the current works in the field of embodied visual navigation by providing a comprehensive literature survey. We summarize the benchmarks and metrics, review different methods, analysis the challenges, and highlight the state-of-the-art methods. Finally, we discuss unresolved challenges in the field of embodied visual navigation and give promising directions in pursuing future research.

updated: Mon Oct 11 2021 08:48:18 GMT+0000 (UTC)

published: Wed Jul 07 2021 12:09:04 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト