Single Stage Multi-Pose Virtual Try-On

Sen He; Yi-Zhe Song; Tao Xiang

シングルステージマルチポーズバーチャル試着

マルチポーズ仮想試着 (MPVTON) は、対象の衣類を対象のポーズで人に装着することを目的としています。衣服にフィットするがポーズを変更しない従来の仮想試着 (VTON) と比較して、MPVTON はより優れた試着体験を提供しますが、衣服とポーズの編集の目的が 2 つあるため、より挑戦的でもあります。既存の MPVTON メソッドは、ターゲットセマンティックレイアウト予測モジュール、粗い試着イメージジェネレーター、洗練された試着イメージジェネレーターを含む 3 つのばらばらなモジュールで構成されるパイプラインを採用しています。これらのモデルは個別にトレーニングされるため、モデルのトレーニングが最適ではなく、満足のいく結果が得られません。本稿では、MPVTON の新しい単一ステージモデルを提案します。私たちのモデルの鍵は、ターゲットポーズで調整された人物と衣服の両方の画像の流れ場を予測する並列の流れ推定モジュールです。続いて、予測されたフローを使用して人物の外観特徴マップと衣服の画像をワープし、スタイルマップを構築します。次に、このマップを使用して、ターゲットの試着画像を生成するために、ターゲットポーズの特徴マップを調整します。並列フロー推定設計により、モデルは 1 つのステージでエンドツーエンドでトレーニングでき、計算効率が向上し、既存の MPVTON ベンチマークで新しい SOTA パフォーマンスが得られます。さらに、マルチタスクトレーニングを紹介し、モデルが従来の VTON およびポーズ転送タスクにも適用され、両方のタスクで SOTA 特化モデルに匹敵するパフォーマンスを達成できることを示します。

Multi-pose virtual try-on (MPVTON) aims to fit a target garment onto a person at a target pose. Compared to traditional virtual try-on (VTON) that fits the garment but keeps the pose unchanged, MPVTON provides a better try-on experience, but is also more challenging due to the dual garment and pose editing objectives. Existing MPVTON methods adopt a pipeline comprising three disjoint modules including a target semantic layout prediction module, a coarse try-on image generator and a refinement try-on image generator. These models are trained separately, leading to sub-optimal model training and unsatisfactory results. In this paper, we propose a novel single stage model for MPVTON. Key to our model is a parallel flow estimation module that predicts the flow fields for both person and garment images conditioned on the target pose. The predicted flows are subsequently used to warp the appearance feature maps of the person and the garment images to construct a style map. The map is then used to modulate the target pose's feature map for target try-on image generation. With the parallel flow estimation design, our model can be trained end-to-end in a single stage and is more computationally efficient, resulting in new SOTA performance on existing MPVTON benchmarks. We further introduce multi-task training and demonstrate that our model can also be applied for traditional VTON and pose transfer tasks and achieve comparable performance to SOTA specialized models on both tasks.

updated: Sat Nov 19 2022 15:02:11 GMT+0000 (UTC)

published: Sat Nov 19 2022 15:02:11 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト