MINERVAS: Massive INterior EnviRonments VirtuAl Synthesis

Haocheng Ren; Hao Zhang; Jia Zheng; Jiaxiang Zheng; Rui Tang; Rui Wang; Hujun Bao

MINERVAS：大規模な内部環境の仮想合成

データ駆動技術の急速な発展に伴い、データはさまざまなコンピュータビジョンタスクで重要な役割を果たしてきました。さまざまな問題に対処するために、多くの現実的で合成的なデータセットが提案されています。ただし、未解決の課題がたくさんあります。（1）データセットの作成は通常、手動の注釈を使用した面倒なプロセスです。（2）ほとんどのデータセットは単一の特定のタスク用にのみ設計されています。（3）3Dシーンの変更またはランダム化（4）商用3Dデータのリリースは著作権の問題に遭遇する可能性があります。このホワイトペーパーでは、さまざまなビジョンタスクの3Dシーンの変更と2D画像の合成を容易にする、大規模な内部環境の仮想合成システムであるMINERVASについて説明します。特に、ドメイン固有言語を使用してプログラム可能なパイプラインを設計し、ユーザーが（1）商用屋内シーンデータベースからシーンを選択し、（2）カスタマイズされたルールを使用してさまざまなタスクのシーンを合成し、（3）さまざまな画像データをレンダリングできるようにします。視覚的な色、幾何学的構造、セマンティックラベルなど。私たちのシステムは、マルチレベルサンプラーを使用してユーザーが制御可能なランダム性を提供することにより、さまざまなタスクに合わせて膨大な数のシーンをカスタマイズすることの難しさを軽減し、ユーザーがきめ細かいシーン構成を操作することから解放します。最も重要なことは、ユーザーが何百万もの屋内シーンを含む商用シーンデータベースにアクセスできるようにし、3DCADモデルなどのコアデータ資産の著作権を保護することです。合成データを使用してさまざまな種類のコンピュータービジョンタスクのパフォーマンスを向上させることにより、システムの有効性と柔軟性を示します。

With the rapid development of data-driven techniques, data has played an essential role in various computer vision tasks. Many realistic and synthetic datasets have been proposed to address different problems. However, there are lots of unresolved challenges: (1) the creation of dataset is usually a tedious process with manual annotations, (2) most datasets are only designed for a single specific task, (3) the modification or randomization of the 3D scene is difficult, and (4) the release of commercial 3D data may encounter copyright issue. This paper presents MINERVAS, a Massive INterior EnviRonments VirtuAl Synthesis system, to facilitate the 3D scene modification and the 2D image synthesis for various vision tasks. In particular, we design a programmable pipeline with Domain-Specific Language, allowing users to (1) select scenes from the commercial indoor scene database, (2) synthesize scenes for different tasks with customized rules, and (3) render various imagery data, such as visual color, geometric structures, semantic label. Our system eases the difficulty of customizing massive numbers of scenes for different tasks and relieves users from manipulating fine-grained scene configurations by providing user-controllable randomness using multi-level samplers. Most importantly, it empowers users to access commercial scene databases with millions of indoor scenes and protects the copyright of core data assets, e.g., 3D CAD models. We demonstrate the validity and flexibility of our system by using our synthesized data to improve the performance on different kinds of computer vision tasks.

updated: Wed Jul 14 2021 14:21:45 GMT+0000 (UTC)

published: Tue Jul 13 2021 14:53:01 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト