AudioEar: Single-View Ear Reconstruction for Personalized Spatial Audio

Xiaoyang Huang; Yanjun Wang; Yang Liu; Bingbing Ni; Wenjun Zhang; Jinxian Liu; Teng Li

AudioEar: パーソナライズされた空間オーディオのための単一ビューの耳の再構築

臨場感あふれる 3D サウンドレンダリングに重点を置いた空間オーディオは、音響業界で広く適用されています。現在の空間オーディオレンダリング方法の主な問題の 1 つは、正確な音源位置を生成するために不可欠な、個人のさまざまな解剖学的構造に基づくパーソナライズの欠如です。この作業では、学際的な観点からこの問題に対処します。空間オーディオのレンダリングは、人体、特に耳の 3D 形状と強く相関しています。この目的のために、単一ビュー画像を使用して 3D 人間の耳を再構築することにより、パーソナライズされた空間オーディオを実現することを提案します。まず、耳の再構築タスクのベンチマークを行うために、RGB 画像を使用した 112 の点群耳スキャンで構成される高品質の 3D 耳データセットである AudioEar3D を紹介します。再構成モデルを自己監視トレーニングするために、2,000 枚の画像で構成される 2D 耳データセットをさらに収集します。各画像には、AudioEar2D という名前のオクルージョンと 55 のランドマークの手動注釈が付けられています。私たちの知る限り、両方のデータセットは、公共で使用するために、その種類の中で最大の規模と最高の品質を備えています。さらに、耳データ用に調整された 2 つの損失関数を使用して、合成データでトレーニングされた深度推定ネットワークによって導かれる再構成方法である AudioEarM を提案します。最後に、視覚と音響のコミュニティ間のギャップを埋めるために、再構築された耳のメッシュを市販の 3D 人体と統合し、コアであるパーソナライズされた頭部関連伝達関数 (HRTF) をシミュレートするパイプラインを開発します。空間オーディオレンダリングの。コードとデータは、https://github.com/seanywang0408/AudioEar で公開されています。

Spatial audio, which focuses on immersive 3D sound rendering, is widely applied in the acoustic industry. One of the key problems of current spatial audio rendering methods is the lack of personalization based on different anatomies of individuals, which is essential to produce accurate sound source positions. In this work, we address this problem from an interdisciplinary perspective. The rendering of spatial audio is strongly correlated with the 3D shape of human bodies, particularly ears. To this end, we propose to achieve personalized spatial audio by reconstructing 3D human ears with single-view images. First, to benchmark the ear reconstruction task, we introduce AudioEar3D, a high-quality 3D ear dataset consisting of 112 point cloud ear scans with RGB images. To self-supervisedly train a reconstruction model, we further collect a 2D ear dataset composed of 2,000 images, each one with manual annotation of occlusion and 55 landmarks, named AudioEar2D. To our knowledge, both datasets have the largest scale and best quality of their kinds for public use. Further, we propose AudioEarM, a reconstruction method guided by a depth estimation network that is trained on synthetic data, with two loss functions tailored for ear data. Lastly, to fill the gap between the vision and acoustics community, we develop a pipeline to integrate the reconstructed ear mesh with an off-the-shelf 3D human body and simulate a personalized Head-Related Transfer Function (HRTF), which is the core of spatial audio rendering. Code and data are publicly available at https://github.com/seanywang0408/AudioEar.

updated: Mon Jan 30 2023 02:15:50 GMT+0000 (UTC)

published: Mon Jan 30 2023 02:15:50 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト