ViDA-MAN: Visual Dialog with Digital Humans

Tong Shen; Jiawei Zuo; Fan Shi; Jin Zhang; Liqin Jiang; Meng Chen; Zhengchen Zhang; Wei Zhang; Xiaodong He; Tao Mei

ViDA-MAN：デジタルヒューマンとのビジュアルダイアログ

マルチモーダルインタラクション用のデジタルヒューマンエージェントであるViDA-MANのデモを行います。これは、音声の即時問い合わせに対してリアルタイムの視聴覚応答を提供します。従来のテキストまたは音声ベースのシステムと比較して、ViDA-MANは人間のような相互作用（たとえば、鮮やかな声、自然な表情、ボディジェスチャ）を提供します。音声要求が与えられると、デモンストレーションは1秒未満の遅延で高品質のビデオで応答できます。没入型のユーザーエクスペリエンスを提供するために、ViDA-MANは、音響音声認識（ASR）、マルチターンダイアログ、テキスト読み上げ（TTS）、トーキングヘッズビデオ生成などのマルチモーダル技術をシームレスに統合します。 ViDA-MANは、大規模な知識ベースに支えられており、チャット、天気、デバイスコントロール、ニュースの推奨事項、ホテルの予約、構造化された知識による質問への回答など、さまざまなトピックについてユーザーとチャットできます。

We demonstrate ViDA-MAN, a digital-human agent for multi-modal interaction, which offers realtime audio-visual responses to instant speech inquiries. Compared to traditional text or voice-based system, ViDA-MAN offers human-like interactions (e.g, vivid voice, natural facial expression and body gestures). Given a speech request, the demonstration is able to response with high quality videos in sub-second latency. To deliver immersive user experience, ViDA-MAN seamlessly integrates multi-modal techniques including Acoustic Speech Recognition (ASR), multi-turn dialog, Text To Speech (TTS), talking heads video generation. Backed with large knowledge base, ViDA-MAN is able to chat with users on a number of topics including chit-chat, weather, device control, News recommendations, booking hotels, as well as answering questions via structured knowledge.

updated: Tue Oct 26 2021 03:23:51 GMT+0000 (UTC)

published: Tue Oct 26 2021 03:23:51 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト