A Temporal Densely Connected Recurrent Network for Event-based Human Pose Estimation

Zhanpeng Shao; Wen Zhou; Wuzhen Wang; Jianyu Yang; Youfu Li

イベントベースの人間の姿勢推定のための時間密結合再帰型ネットワーク

イベントカメラは、ピクセルごとの明るさの変化を非同期的に報告する、バイオにヒントを得た新しいビジョンセンサーです。高ダイナミックレンジ、高速応答、および低電力バジェットという顕著な利点があり、制御されていない環境でローカルモーションを最適にキャプチャできます。イベントカメラによる人間の姿勢推定はめったに調査されないため、これは人間の姿勢推定のためのイベントカメラの可能性を解き放つ動機になります。ただし、従来のフレームベースのカメラからの新しいパラダイムシフトにより、時間間隔内のイベント信号に含まれる情報は非常に限られています。これは、イベントカメラがキャプチャできるのは動いている身体部分のみであり、それらの静的な身体部分を無視するため、一部の部分が不完全になるためです。または時間間隔で消えさえしました。この論文では、不完全な情報の問題に対処するために、高密度に接続された新しいリカレントアーキテクチャを提案します。このリカレントアーキテクチャにより、時間ステップ全体でシーケンシャルだけでなく非シーケンシャルの幾何学的一貫性も明示的にモデル化して、前のフレームから情報を蓄積し、人体全体を復元して、イベントデータから安定した正確な人間の姿勢推定を実現できます。さらに、モデルをより適切に評価するために、人間の姿勢の注釈を含む大規模なマルチモーダルイベントベースのデータセットを収集します。これは、私たちの知る限り、最も困難なものです。 2 つの公開データセットと私たち自身のデータセットに関する実験結果は、私たちのアプローチの有効性と強みを示しています。コードは、将来の研究を容易にするためにオンラインで入手できます。

Event camera is an emerging bio-inspired vision sensors that report per-pixel brightness changes asynchronously. It holds noticeable advantage of high dynamic range, high speed response, and low power budget that enable it to best capture local motions in uncontrolled environments. This motivates us to unlock the potential of event cameras for human pose estimation, as the human pose estimation with event cameras is rarely explored. Due to the novel paradigm shift from conventional frame-based cameras, however, event signals in a time interval contain very limited information, as event cameras can only capture the moving body parts and ignores those static body parts, resulting in some parts to be incomplete or even disappeared in the time interval. This paper proposes a novel densely connected recurrent architecture to address the problem of incomplete information. By this recurrent architecture, we can explicitly model not only the sequential but also non-sequential geometric consistency across time steps to accumulate information from previous frames to recover the entire human bodies, achieving a stable and accurate human pose estimation from event data. Moreover, to better evaluate our model, we collect a large scale multimodal event-based dataset that comes with human pose annotations, which is by far the most challenging one to the best of our knowledge. The experimental results on two public datasets and our own dataset demonstrate the effectiveness and strength of our approach. Code can be available online for facilitating the future research.

updated: Wed Apr 05 2023 09:36:18 GMT+0000 (UTC)

published: Thu Sep 15 2022 04:08:18 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト