Learning Enhanced Resolution-wise features for Human Pose Estimation

Kun Zhang; Peng He; Ping Yao; Ge Chen; Rui Wu; Min Du; Huimin Li; Li Fu; Tianyao Zheng

人間の姿勢推定のための強化された解像度に関する機能の学習

最近、多重解像度ネットワーク（砂時計、CPN、HRNetなど）は、さまざまな解像度の特徴マップを組み合わせることにより、ポーズ推定で大きなパフォーマンスを達成しています。この論文では、正確な姿勢推定のための強化された解像度ごとの特徴マップを学習するために、解像度ごとの注意モジュール（RAM）と段階的ピラミッド精密化（GPR）を提案します。具体的には、RAMは重みのグループを学習して、解像度間での特徴マップのさまざまな重要性を表します。GPRは、2つの特徴マップごとに、低解像度から高解像度まで徐々にマージして、最終的な人間のキーポイントヒートマップを回帰します。 CNNによって学習された拡張された解像度に関する機能により、より正確な人間のキーポイントの位置を取得します。提案された方法の有効性はMS-COCOデータセットで実証され、追加のヒューマンキーポイントトレーニングデータセットを使用せずに、COCO val2017セットで77.7、test-dev2017セットで77.0の平均精度で最先端のパフォーマンスを達成します。

Recently, multi-resolution networks (such as Hourglass, CPN, HRNet, etc.) have achieved significant performance on pose estimation by combining feature maps of various resolutions. In this paper, we propose a Resolution-wise Attention Module (RAM) and Gradual Pyramid Refinement (GPR), to learn enhanced resolution-wise feature maps for precise pose estimation. Specifically, RAM learns a group of weights to represent the different importance of feature maps across resolutions, and the GPR gradually merges every two feature maps from low to high resolutions to regress final human keypoint heatmaps. With the enhanced resolution-wise features learnt by CNN, we obtain more accurate human keypoint locations. The efficacies of our proposed methods are demonstrated on MS-COCO dataset, achieving state-of-the-art performance with average precision of 77.7 on COCO val2017 set and 77.0 on test-dev2017 set without using extra human keypoint training dataset.

updated: Sun Dec 13 2020 15:22:41 GMT+0000 (UTC)

published: Wed Sep 11 2019 14:46:28 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト