GAUDI: A Neural Architect for Immersive 3D Scene Generation

Miguel Angel Bautista; Pengsheng Guo; Samira Abnar; Walter Talbott; Alexander Toshev; Zhuoyuan Chen; Laurent Dinh; Shuangfei Zhai; Hanlin Goh; Daniel Ulbricht; Afshin Dehghan; Josh Susskind

GAUDI：没入型3Dシーン生成のためのニューラルアーキテクト

移動するカメラから没入型にレンダリングできる複雑でリアルな3Dシーンの分布をキャプチャできる生成モデルであるGAUDIを紹介します。スケーラブルでありながら強力なアプローチでこの困難な問題に取り組みます。最初に、放射輝度フィールドとカメラポーズを解きほぐす潜在表現を最適化します。次に、この潜在表現を使用して、3Dシーンの無条件および条件付き生成の両方を可能にする生成モデルを学習します。私たちのモデルは、カメラのポーズ分布をサンプル間で共有できるという仮定を取り除くことにより、単一のオブジェクトに焦点を当てた以前の作品を一般化します。 GAUDIは、複数のデータセットにわたる無条件の生成設定で最先端のパフォーマンスを取得し、まばらな画像観測やシーンを説明するテキストなどの条件変数を指定して3Dシーンの条件付き生成を可能にすることを示します。

We introduce GAUDI, a generative model capable of capturing the distribution of complex and realistic 3D scenes that can be rendered immersively from a moving camera. We tackle this challenging problem with a scalable yet powerful approach, where we first optimize a latent representation that disentangles radiance fields and camera poses. This latent representation is then used to learn a generative model that enables both unconditional and conditional generation of 3D scenes. Our model generalizes previous works that focus on single objects by removing the assumption that the camera pose distribution can be shared across samples. We show that GAUDI obtains state-of-the-art performance in the unconditional generative setting across multiple datasets and allows for conditional generation of 3D scenes given conditioning variables like sparse image observations or text that describes the scene.

updated: Wed Jul 27 2022 19:10:32 GMT+0000 (UTC)

published: Wed Jul 27 2022 19:10:32 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト