Structured Multimodal Attentions for TextVQA

Chenyu Gao; Qi Zhu; Peng Wang; Hui Li; Yuliang Liu; Anton van den Hengel; Qi Wu

TextVQAの構造化されたマルチモーダルアテンション

この論文では、主に上記の最初の2つの問題を解決するために、エンドツーエンドの構造化マルチモーダル注意（SMA）ニューラルネットワークを提案します。 SMAは、最初に構造グラフ表現を使用して、画像に表示されるオブジェクト-オブジェクト、オブジェクト-テキスト、およびテキスト-テキストの関係をエンコードし、次にマルチモーダルグラフアテンションネットワークを設計してそれを推論します。最後に、上記のモジュールからの出力は、グローバルローカル注意応答モジュールによって処理され、M4Cに従って、OCRと一般語彙の両方からのトークンを繰り返しスプライシングする応答を生成します。提案されたモデルは、事前トレーニングベースのTAPを除くすべてのモデルの中で、TextVQAデータセットのSoTAモデルとST-VQAデータセットの2つのタスクよりも優れています。強力な推論能力を示し、TextVQAチャレンジ2020でも1位を獲得しました。さまざまなOCR手法をいくつかの推論モデルで広範囲にテストし、OCRパフォーマンスが徐々に向上した場合のTextVQAベンチマークへの影響を調査します。 OCRの結果が向上すると、さまざまなモデルでVQAの精度が劇的に向上しますが、私たちのモデルは、強力なテキストと視覚の推論能力に最も恵まれています。私たちの方法に上限を与え、さらなる作業のために公正なテストベースを利用できるようにするために、元のリリースでは提供されなかった、TextVQAデータセットの人間が注釈を付けたグラウンドトゥルースOCR注釈も提供します。 TextVQAデータセットのコードとグラウンドトゥルースOCRアノテーションは、https：//github.com/ChenyuGAO-CS/SMAで入手できます。

In this paper, we propose an end-to-end structured multimodal attention (SMA) neural network to mainly solve the first two issues above. SMA first uses a structural graph representation to encode the object-object, object-text and text-text relationships appearing in the image, and then designs a multimodal graph attention network to reason over it. Finally, the outputs from the above modules are processed by a global-local attentional answering module to produce an answer splicing together tokens from both OCR and general vocabulary iteratively by following M4C. Our proposed model outperforms the SoTA models on TextVQA dataset and two tasks of ST-VQA dataset among all models except pre-training based TAP. Demonstrating strong reasoning ability, it also won first place in TextVQA Challenge 2020. We extensively test different OCR methods on several reasoning models and investigate the impact of gradually increased OCR performance on TextVQA benchmark. With better OCR results, different models share dramatic improvement over the VQA accuracy, but our model benefits most blessed by strong textual-visual reasoning ability. To grant our method an upper bound and make a fair testing base available for further works, we also provide human-annotated ground-truth OCR annotations for the TextVQA dataset, which were not given in the original release. The code and ground-truth OCR annotations for the TextVQA dataset are available at https://github.com/ChenyuGAO-CS/SMA

updated: Fri Nov 26 2021 03:00:58 GMT+0000 (UTC)

published: Mon Jun 01 2020 07:07:36 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト