CGAP2: Context and gap aware predictive pose framework for early detection of gestures

Nishant Bhattacharya; Suresh Sundaram

CGAP2：ジェスチャの早期検出のためのコンテキストおよびギャップを意識した予測ポーズフレームワーク

自動運転車の操作への関心が高まるにつれ、人間と車両の相互作用のための効率的な予測ジェスチャ認識システムの必要性も同様に高まっています。既存のジェスチャ認識アルゴリズムは、主に履歴データに制限されていました。この論文では、オンラインでジェスチャーを予測的に認識するための将来のポーズデータを予測する、新しいコンテキストとギャップを意識したポーズ予測フレームワーク（CGAP2）を提案します。 CGAP2は、ポーズ予測モジュールと組み合わせたエンコーダ-デコーダアーキテクチャを実装して、将来のフレームとそれに続く浅い分類器を予測します。 CGAP2ポーズ予測モジュールは3D畳み込み層を使用し、提供されるポーズフレームの数、各ポーズフレーム間の時間差、および予測されるポーズフレームの数に依存します。 CGAP2のパフォーマンスは、MPJPEメトリックを使用してHuman3.6Mデータセットで評価されます。事前に15フレームのポーズ予測を行うと、79.0mmの誤差が発生します。ポーズ予測モジュールはわずか26Mのパラメーターで構成され、NVidia RTXTitanで50FPSで実行できます。さらに、アブレーション研究は、ポーズ予測モジュールにより高いコンテキスト情報を提供することは、予測的認識にとって有害である可能性があることを示しています。 CGAP2は、自動運転車にとって非常に重要な他のジェスチャ認識システムと比較して、1秒の時間的利点があります。

With a growing interest in autonomous vehicles' operation, there is an equally increasing need for efficient anticipatory gesture recognition systems for human-vehicle interaction. Existing gesture-recognition algorithms have been primarily restricted to historical data. In this paper, we propose a novel context and gap aware pose prediction framework(CGAP2), which predicts future pose data for anticipatory recognition of gestures in an online fashion. CGAP2 implements an encoder-decoder architecture paired with a pose prediction module to anticipate future frames followed by a shallow classifier. CGAP2 pose prediction module uses 3D convolutional layers and depends on the number of pose frames supplied, the time difference between each pose frame, and the number of predicted pose frames. The performance of CGAP2 is evaluated on the Human3.6M dataset with the MPJPE metric. For pose prediction of 15 frames in advance, an error of 79.0mm is achieved. The pose prediction module consists of only 26M parameters and can run at 50 FPS on the NVidia RTX Titan. Furthermore, the ablation study indicates supplying higher context information to the pose prediction module can be detrimental for anticipatory recognition. CGAP2 has a 1-second time advantage compared to other gesture recognition systems, which can be crucial for autonomous vehicles.

updated: Wed Nov 18 2020 11:21:04 GMT+0000 (UTC)

published: Wed Nov 18 2020 11:21:04 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト