How Transferable are Reasoning Patterns in VQA?

Corentin Kervadec; Theo Jaunet; Grigory Antipov; Moez Baccouche; Romain Vuillemot; Christian Wolf

VQAの推論パターンはどの程度転送可能ですか？

開始以来、視覚的質問応答（VQA）はタスクとして有名であり、モデルは高レベルの推論を実行する代わりに、データセットのバイアスを利用してショートカットを見つける傾向があります。従来の方法では、トレーニングデータからバイアスを削除するか、モデルにブランチを追加してバイアスを検出して削除することで、これに対処します。この論文では、視覚の不確実性が、視覚と言語の問題における推論の学習の成功を妨げる支配的な要因であると主張します。私たちは視覚的なオラクルを訓練し、大規模な研究で、標準モデルと比較して偽のデータセットバイアスを悪用する傾向がはるかに少ないという実験的証拠を提供します。ビジュアルオラクルで機能している注意メカニズムを研究し、SOTATransformerベースのモデルと比較することを提案します。公開しているオンライン視覚化ツール（https://reasoningpatterns.github.io）で取得した推論パターンの詳細な分析と視覚化を提供します。オラクルからSOTATransformerベースのVQAモデルに推論パターンを転送し、微調整によって標準のノイズの多い視覚入力を取得することで、これらの洞察を活用します。実験では、全体的な精度が高く、各質問タイプのまれな回答の精度が高いことを報告します。これにより、一般化が改善され、データセットのバイアスへの依存度が低下する証拠が得られます。

Since its inception, Visual Question Answering (VQA) is notoriously known as a task, where models are prone to exploit biases in datasets to find shortcuts instead of performing high-level reasoning. Classical methods address this by removing biases from training data, or adding branches to models to detect and remove biases. In this paper, we argue that uncertainty in vision is a dominating factor preventing the successful learning of reasoning in vision and language problems. We train a visual oracle and in a large scale study provide experimental evidence that it is much less prone to exploiting spurious dataset biases compared to standard models. We propose to study the attention mechanisms at work in the visual oracle and compare them with a SOTA Transformer-based model. We provide an in-depth analysis and visualizations of reasoning patterns obtained with an online visualization tool which we make publicly available (https://reasoningpatterns.github.io). We exploit these insights by transferring reasoning patterns from the oracle to a SOTA Transformer-based VQA model taking standard noisy visual inputs via fine-tuning. In experiments we report higher overall accuracy, as well as accuracy on infrequent answers for each question type, which provides evidence for improved generalization and a decrease of the dependency on dataset biases.

updated: Thu Apr 08 2021 10:18:45 GMT+0000 (UTC)

published: Thu Apr 08 2021 10:18:45 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト