PIFu for the Real World: A Self-supervised Framework to Reconstruct Dressed Human from Single-view Images

Zhangyang Xiong; Dong Du; Yushuang Wu; Jingqi Dong; Di Kang; Linchao Bao; Xiaoguang Han

現実世界の PIFu: 単一ビュー画像から服を着た人間を再構築するための自己教師ありフレームワーク

1 つの画像から、さまざまなポーズや衣服によって引き起こされる洗練された人間の形状を正確に再構築することは非常に困難です。最近、ピクセル整列陰関数 (PIFu) に基づく作品が大きな一歩を踏み出し、画像ベースの 3D ヒューマンデジタル化で最先端の忠実度を達成しました。ただし、PIFu のトレーニングは、高価で限られた 3D グラウンドトゥルースデータ (つまり、合成データ) に大きく依存しているため、より多様な現実世界の画像への一般化が妨げられています。この作業では、SelfPIFu という名前のエンドツーエンドの自己監視型ネットワークを提案して、豊富で多様な野生の画像を利用し、制約のない野生の画像でテストすると、再構成が大幅に改善されます。 SelfPIFu の中核となるのは、深度ガイド付きボリューム/サーフェス認識符号付き距離フィールド (SDF) 学習で、GT メッシュにアクセスせずに PIFu の自己教師あり学習を可能にします。フレームワーク全体は、通常の推定器、深度推定器、および SDF ベースの PIFu で構成され、トレーニング中に追加の深度 GT をより適切に利用します。広範な実験により、自己監視フレームワークの有効性と、深さを入力として使用することの優位性が実証されています。合成データでは、Intersection-Over-Union (IoU) は 93.5% を達成し、PIFuHD と比較して 18% 高くなります。野生の画像については、再構成された結果に対してユーザー調査を行います。結果の選択率は、他の最先端の方法と比較して 68% を超えています。

It is very challenging to accurately reconstruct sophisticated human geometry caused by various poses and garments from a single image. Recently, works based on pixel-aligned implicit function (PIFu) have made a big step and achieved state-of-the-art fidelity on image-based 3D human digitization. However, the training of PIFu relies heavily on expensive and limited 3D ground truth data (i.e. synthetic data), thus hindering its generalization to more diverse real world images. In this work, we propose an end-to-end self-supervised network named SelfPIFu to utilize abundant and diverse in-the-wild images, resulting in largely improved reconstructions when tested on unconstrained in-the-wild images. At the core of SelfPIFu is the depth-guided volume-/surface-aware signed distance fields (SDF) learning, which enables self-supervised learning of a PIFu without access to GT mesh. The whole framework consists of a normal estimator, a depth estimator, and a SDF-based PIFu and better utilizes extra depth GT during training. Extensive experiments demonstrate the effectiveness of our self-supervised framework and the superiority of using depth as input. On synthetic data, our Intersection-Over-Union (IoU) achieves to 93.5%, 18% higher compared with PIFuHD. For in-the-wild images, we conduct user studies on the reconstructed results, the selection rate of our results is over 68% compared with other state-of-the-art methods.

updated: Tue Aug 23 2022 07:00:44 GMT+0000 (UTC)

published: Tue Aug 23 2022 07:00:44 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト