Prophet Attention: Predicting Attention with Future Attention for Improved Image Captioning

Fenglin Liu; Xuewei Ma; Xuancheng Ren; Xian Wu; Wei Fan; Yuexian Zou; Xu Sun

Prophet Attention: 改善された画像キャプションのための将来の注意で注意を予測する

最近、注意ベースのモデルは、多くのシーケンスからシーケンスへの学習システムで広く使用されています。特に画像キャプションの場合、注意ベースのモデルは、適切に生成された単語で正しい画像領域を接地することが期待されます。ただし、デコードプロセスのタイムステップごとに、注意ベースのモデルは通常、現在の入力の非表示状態を使用して画像領域に注意を向けます。この設定では、これらの注意モデルには、生成される単語ではなく前の単語に基づいて注意の重みを計算するという「焦点のずれ」の問題があり、グラウンディングとキャプションの両方のパフォーマンスが損なわれます。この論文では、自己監督の形に似た預言者の注意を提案します。トレーニング段階では、このモジュールは将来の情報を利用して、画像領域に対する「理想的な」注意の重みを計算します。これらの計算された「理想的な」重みは、「逸脱した」注意を正則化するためにさらに使用されます。このようにして、画像領域は正しい単語に基づいています。提案された Prophet Attention は、既存の画像キャプションモデルに簡単に組み込むことができ、グラウンディングとキャプションの両方のパフォーマンスを向上させることができます。 Flickr30k Entities と MSCOCO データセットに関する実験は、提案された Prophet Attention が、自動測定基準と人間による評価の両方で一貫してベースラインを上回っていることを示しています。注目に値するのは、2 つのベンチマークデータセットに新しい最先端技術を設定し、オンライン MSCOCO ベンチマークのリーダーボードで、デフォルトのランキングスコア、つまり CIDEr-c40 で 1 位を獲得したことです。

Recently, attention based models have been used extensively in many sequence-to-sequence learning systems. Especially for image captioning, the attention based models are expected to ground correct image regions with proper generated words. However, for each time step in the decoding process, the attention based models usually use the hidden state of the current input to attend to the image regions. Under this setting, these attention models have a "deviated focus" problem that they calculate the attention weights based on previous words instead of the one to be generated, impairing the performance of both grounding and captioning. In this paper, we propose the Prophet Attention, similar to the form of self-supervision. In the training stage, this module utilizes the future information to calculate the "ideal" attention weights towards image regions. These calculated "ideal" weights are further used to regularize the "deviated" attention. In this manner, image regions are grounded with the correct words. The proposed Prophet Attention can be easily incorporated into existing image captioning models to improve their performance of both grounding and captioning. The experiments on the Flickr30k Entities and the MSCOCO datasets show that the proposed Prophet Attention consistently outperforms baselines in both automatic metrics and human evaluations. It is worth noticing that we set new state-of-the-arts on the two benchmark datasets and achieve the 1st place on the leaderboard of the online MSCOCO benchmark in terms of the default ranking score, i.e., CIDEr-c40.

updated: Wed Oct 19 2022 22:29:31 GMT+0000 (UTC)

published: Wed Oct 19 2022 22:29:31 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト