Streaming Audio-Visual Speech Recognition with Alignment Regularization

Pingchuan Ma; Niko Moritz; Stavros Petridis; Christian Fuegen; Maja Pantic

アラインメント正則化によるオーディオビジュアル音声認識のストリーミング

話された直後に単語を認識することは、現実世界のシナリオにおける自動音声認識 (ASR) システムにとって重要な要件です。その結果、オーディオのみの ASR モデルのストリーミングに関する多数の研究が文献で発表されています。ただし、ストリーミングオーディオビジュアル自動音声認識 (AV-ASR) は、以前の作品ではほとんど注目されていませんでした。この作業では、ハイブリッドコネクショニスト時間分類 (CTC)/注意ニューラルネットワークアーキテクチャに基づくストリーミング AV-ASR システムを提案します。オーディオエンコーダーとビジュアルエンコーダーのニューラルネットワークは、どちらもコンフォマーアーキテクチャに基づいており、チャンク単位の自己注意 (CSA) と因果的畳み込みを使用してストリーミング可能になっています。デコーダニューラルネットワークによるストリーミング認識は、CTC/Attention スコアリングを組み合わせた時間同期デコードを実行する Triggered Attention 手法を使用して実現されます。 CTC などのフレームレベルの ASR 基準の場合、オーディオエンコーダーとビジュアルエンコーダーからの同期応答は、共同 AV 意思決定プロセスにとって重要です。この作業では、オーディオとビジュアルのエンコーダーの同期を促進する新しいアライメント正則化手法を提案します。これにより、ストリーミングおよびオフライン AV-ASR モデルのすべての SNR レベルでワードエラーレート (WER) が向上します。提案された AV-ASR モデルは、オフラインとオンラインのセットアップで、Lip Reading Sentences 3 (LRS3) データセットでそれぞれ 2.0% と 2.6% の WER を達成します。どちらも、外部トレーニングデータがない場合に最先端の結果を示します。使用済み。

Recognizing a word shortly after it is spoken is an important requirement for automatic speech recognition (ASR) systems in real-world scenarios. As a result, a large body of work on streaming audio-only ASR models has been presented in the literature. However, streaming audio-visual automatic speech recognition (AV-ASR) has received little attention in earlier works. In this work, we propose a streaming AV-ASR system based on a hybrid connectionist temporal classification (CTC)/attention neural network architecture. The audio and the visual encoder neural networks are both based on the conformer architecture, which is made streamable using chunk-wise self-attention (CSA) and causal convolution. Streaming recognition with a decoder neural network is realized by using the triggered attention technique, which performs time-synchronous decoding with joint CTC/attention scoring. For frame-level ASR criteria, such as CTC, a synchronized response from the audio and visual encoders is critical for a joint AV decision making process. In this work, we propose a novel alignment regularization technique that promotes synchronization of the audio and visual encoder, which in turn results in better word error rates (WERs) at all SNR levels for streaming and offline AV-ASR models. The proposed AV-ASR model achieves WERs of 2.0% and 2.6% on the Lip Reading Sentences 3 (LRS3) dataset in an offline and online setup, respectively, which both present state-of-the-art results when no external training data are used.

updated: Thu Nov 03 2022 20:20:47 GMT+0000 (UTC)

published: Thu Nov 03 2022 20:20:47 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト