From Recognition to Prediction: Analysis of Human Action and Trajectory Prediction in Video

Junwei Liang

認識から予測へ：ビデオにおける人間の行動と軌道予測の分析

コンピュータビジョンディープラーニングの進歩により、システムは、自動運転、社会認識ロボットアシスタント、公安監視などのアプリケーションを可能にするために、ビデオからの前例のない量の豊富な視覚情報を分析できるようになりました。これらのアプリケーションでは、人間の行動を解読して将来の経路/軌道を予測し、ビデオから何をするかを予測することが重要です。ただし、シーンのセマンティクスと人間の意図をモデル化するのは難しいため、人間の軌道予測は依然として困難な作業です。多くのシステムは、歩行者の将来について推論するための高レベルのセマンティック属性を提供していません。この設計は、さまざまなドメインや目に見えないシナリオからのビデオデータの予測パフォーマンスを妨げます。最適な将来の人間の行動予測を可能にするには、システムが人間の活動とシーンのセマンティクスを検出および分析し、コンテキストを理解するために後続の予測モジュールに有益な機能を渡すことができることが重要です。

With the advancement in computer vision deep learning, systems now are able to analyze an unprecedented amount of rich visual information from videos to enable applications such as autonomous driving, socially-aware robot assistant and public safety monitoring. Deciphering human behaviors to predict their future paths/trajectories and what they would do from videos is important in these applications. However, human trajectory prediction still remains a challenging task, as scene semantics and human intent are difficult to model. Many systems do not provide high-level semantic attributes to reason about pedestrian future. This design hinders prediction performance in video data from diverse domains and unseen scenarios. To enable optimal future human behavioral forecasting, it is crucial for the system to be able to detect and analyze human activities as well as scene semantics, passing informative features to the subsequent prediction module for context understanding.

updated: Fri Jul 16 2021 13:45:43 GMT+0000 (UTC)

published: Fri Nov 20 2020 22:23:34 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト