Dynamic Graph Attention for Referring Expression Comprehension

Sibei Yang; Guanbin Li; Yizhou Yu

式理解を参照するための動的グラフ注意

参照表現の理解は、画像内の表現を参照する自然言語で記述されたオブジェクトインスタンスを見つけることを目的としています。このタスクは構成的であり、本質的に画像内のオブジェクト間の関係に加えて視覚的な推論を必要とします。一方、視覚的推論プロセスは、参照表現の言語構造によって導かれます。ただし、既存のアプローチでは、オブジェクトを分離して処理するか、式の潜在的な複雑さに合わせずにオブジェクト間の1次関係のみを探索します。したがって、彼らが複雑な参照表現の基礎に適応することは困難です。この論文では、言語駆動の視覚的推論の観点から表現理解を参照する問題を探求し、画像内のオブジェクトと言語構造の両方の関係をモデル化することにより、多段階推論を実行する動的グラフ注意ネットワークを提案します式の。特に、オブジェクトとそれらの関係にそれぞれ対応するノードとエッジを持つ画像のグラフを構築し、言語にガイドされた視覚的推論プロセスを予測する差分アナライザーを提案し、グラフの上で段階的推論を実行して更新しますすべてのノードでの複合オブジェクト表現。実験結果は、提案された方法が、3つの一般的なベンチマークデータセット全体で既存のすべての最先端アルゴリズムを大幅に上回るだけでなく、複雑な言語記述で参照されるオブジェクトを段階的に特定するための解釈可能な視覚的証拠を生成できることを示しています。

Referring expression comprehension aims to locate the object instance described by a natural language referring expression in an image. This task is compositional and inherently requires visual reasoning on top of the relationships among the objects in the image. Meanwhile, the visual reasoning process is guided by the linguistic structure of the referring expression. However, existing approaches treat the objects in isolation or only explore the first-order relationships between objects without being aligned with the potential complexity of the expression. Thus it is hard for them to adapt to the grounding of complex referring expressions. In this paper, we explore the problem of referring expression comprehension from the perspective of language-driven visual reasoning, and propose a dynamic graph attention network to perform multi-step reasoning by modeling both the relationships among the objects in the image and the linguistic structure of the expression. In particular, we construct a graph for the image with the nodes and edges corresponding to the objects and their relationships respectively, propose a differential analyzer to predict a language-guided visual reasoning process, and perform stepwise reasoning on top of the graph to update the compound object representation at every node. Experimental results demonstrate that the proposed method can not only significantly surpass all existing state-of-the-art algorithms across three common benchmark datasets, but also generate interpretable visual evidences for stepwisely locating the objects referred to in complex language descriptions.

updated: Wed Sep 18 2019 01:47:27 GMT+0000 (UTC)

published: Wed Sep 18 2019 01:47:27 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト